← Back to Index
Daily Research Digest

arXiv Papers

2026-09-24
371
Papers
4
Categories
371
Translated
收藏清单 0
机器人学 (Robotics)
91
cs.RO / 1 / 2609.26857

FLINT: Fast Lightweight Inference for Traversability

FLINT:面向可通行性的快速轻量级推理方法
Bonilla, William, Boisvert, Maxime, Poissant, David-Alexandre, Meger, David, Petit, Louis
Abstract
Navigation in off-road conditions is challenging due to the lack of structure. There is no fixed vocabulary for what is traversable. The traversability depends on both the environment and the embodiment's dynamics. Neither of these two variables can be hand-labeled at scale. Thus, traversability has to be learned by the embodiment's own experience. Modern platforms tend to use multiple sensors to estimate traversability and navigate: RGBD cameras, lidar, radar, IMU, with computationally intensive platforms to run inference on neural networks. Against this trend, we propose FLINT, a lightweight traversability estimator: a 21.6M-parameter backbone, 38\times smaller than a comparable foundation-model backbone, that scores higher on held-out terrain probes and runs at 14.7 FPS on CPU alone using a RGB camera has the only sensor. Despite that gap in scale, FLINT produces a cheaper, more accurate costmap than a deployed foundation-model system (WildOS) on 23 of 24 replayed field logs. We compare different self-supervised learning signals and deploy the resulting models on a real platform in closed-loop field trials: the best self-supervised head reaches 99% autonomy over the route, outperforming a human-label-trained baseline deployed live on the same course. Our results show that heavy sensing and computing are not necessary for traversability estimation.
Chinese Translation
越野环境下的导航由于缺乏结构化信息而充满挑战。对于什么是可通行的,并不存在固定的定义。可通行性取决于环境与载体自身动力学的双重因素,而这两个变量都无法通过人工标注实现大规模覆盖。因此,可通行性必须通过载体自身的经验来学习。现代平台通常使用多种传感器(如RGBD相机、激光雷达、毫米波雷达、IMU)来估计可通行性并进行导航,并依赖计算密集型平台来运行神经网络推理。与这一趋势相反,我们提出了FLINT,一种轻量级的可通行性估计器:其骨干网络仅有2160万参数,比同类基础模型骨干网络小38倍,在留出的地形探测数据上得分更高,并且仅使用RGB相机作为唯一传感器时,可在纯CPU上以14.7 FPS的速度运行。尽管规模差距悬殊,FLINT在24条回放实地日志中的23条上,生成了比已部署的基础模型系统(WildOS)更廉价且更精确的代价地图。我们比较了不同的自监督学习信号,并将所得模型部署到真实平台上进行闭环实地试验:最佳的自监督模型在路线上的自主行驶率达到99%,优于在同一赛道上实时部署的、基于人工标注训练的基线方法。我们的结果表明,重型传感器和大规模计算并非可通行性估计所必需。
cs.RO / 2 / 2609.26868

Backdoors in Learning-Based Industrial Robotic Arm Manipulation: An Empirical Security Study

基于学习的工业机械臂操控中的后门攻击:一项实证安全研究
Zhang, Zijian, Zeng, Zhen, Gu, Zhongshu, Pisharody, Sandeep
Abstract
Learning-based models (e.g., visuomotor and Vision-Language-Action (VLA)) are increasingly explored for industrial robotic manipulation, where model predictions are directly translated into physical actions. This tight coupling between model behavior and physical execution makes hidden security vulnerabilities particularly consequential. While backdoor attacks have been widely studied in conventional AI models, their effects on deployed learning-based robotic arm manipulation systems remain less understood: a backdoored robot can behave normally during benign operation while inducing attacker-specified behaviors only when specific triggers are present, posing potentially serious risks in physical environments. In this work, we present a preliminary empirical security study of backdoor attacks and defenses in learning-based robotic manipulation on two real commercial industrial robotic arms (FANUC and xArm). We investigate whether a backdoor can reliably induce semantically incorrect manipulation behaviors while remaining stealthy under nominal task execution. We then develop an online defense pipeline that detects and neutralizes triggers at runtime, and compare its effectiveness against an offline fine-tuning defense. Beyond defense effectiveness, we further evaluate the computational latency and execution overhead introduced by the defense pipeline to assess its suitability for high-throughput industrial operation.
Chinese Translation
基于学习的模型(如视觉运动模型和视觉-语言-动作(Vision-Language-Action, VLA)模型)正被越来越多地应用于工业机器人操控领域,其中模型预测会被直接转化为物理动作。模型行为与物理执行之间的这种紧密耦合,使得潜在的安全漏洞后果尤为严重。尽管后门攻击在传统人工智能模型中已被广泛研究,但其对已部署的基于学习的机械臂操控系统的影响仍知之甚少:一个被植入后门的机器人在正常运行时表现正常,而仅在出现特定触发器时才会执行攻击者指定的行为,这在物理环境中可能带来严重风险。在本工作中,我们针对两种真实的商用工业机械臂(FANUC 和 xArm),对基于学习的机器人操控中的后门攻击与防御进行了一项初步的实证安全研究。我们研究了后门能否在保持隐蔽性(即在正常任务执行中不被察觉)的同时,可靠地诱发语义错误的操控行为。随后,我们开发了一种在线防御流水线,可在运行时检测并消除触发器,并将其防御效果与离线微调防御进行比较。除防御有效性外,我们还进一步评估了该防御流水线带来的计算延迟和执行开销,以评估其是否适用于高吞吐量的工业作业场景。
cs.RO / 3 / 2609.26872

MSK-Bench: Benchmarking Full-Body Musculoskeletal Motor Control Across Tasks, Control Paradigms, and Physiological Metrics

MSK-Bench:跨任务、控制范式与生理指标的全身肌肉骨骼运动控制基准测试
Ou, Mengtao, Zhang, Zongzheng, Xiao, Zhenghao, Pan, Yixuan, Zhuang, Ziwen, Zhao, Hang, Li, Hongyang, Sui, Yanan, Liu, Libin, Zhao, Hao
Abstract
Musculoskeletal (MSK) humanoids provide a physiologically grounded embodiment for studying full-body motor control, but their high-dimensional muscle actuation, delayed activation dynamics, and redundant muscle--tendon structures make learning substantially harder than torque-driven humanoid control. Existing MSK benchmarks remain fragmented across gait, prosthetics, dexterous hands, or challenge-specific tracks, leaving full-body muscle-actuated control insufficiently evaluated under standardized tasks, methods, and metrics. We introduce MSK-Bench, a benchmark of 22 full-body motor-control tasks organized into three progressively challenging categories: postural stabilization, common locomotor behaviors, and contact-rich environmental interaction. Under unified task protocols and robustness perturbations, MSK-Bench evaluates 5 representative control paradigms, including reward-based RL, agentic reward tuning, latent-action RL, imitation-prior control, and residual adaptation over imitation priors. Beyond task success and reward, MSK-Bench further reports robustness analysis and physiology-oriented diagnostics, including activation cost, joint smoothness, and EMG-envelope similarity. Our empirical study shows that embodiment-aware exploration and structured action representations improve task coverage in high-dimensional muscle spaces, imitation priors enhance reference-compatible stabilization and locomotion but degrade under contact-rich terrain mismatch, and residual adaptation can recover successful behaviors when fixed references fail. We further find that improved task success does not necessarily imply improved physiological agreement, highlighting the importance of evaluating task performance, robustness, and physiological behavior jointly. MSK-Bench provides a task--method--metric testbed for full-body muscle-actuated humanoid control.
Chinese Translation
肌肉骨骼(MSK)人形机器人为研究全身运动控制提供了具有生理学基础的具身平台,但其高维度的肌肉驱动、延迟的激活动力学以及冗余的肌肉-肌腱结构,使得学习任务远比力矩驱动的人形机器人控制更加困难。现有的MSK基准测试在步态、假肢、灵巧手或特定挑战赛道等方面较为零散,缺乏在标准化任务、方法和指标下对全身肌肉驱动控制的充分评估。我们提出了MSK-Bench,这是一个包含22个全身运动控制任务的基准,按三个难度递进的类别组织:姿态稳定、常见运动行为以及富含接触的环境交互。在统一的任务协议和鲁棒性扰动下,MSK-Bench评估了5种代表性控制范式,包括基于奖励的强化学习(RL)、智能体奖励调优、潜在动作强化学习、模仿先验控制以及基于模仿先验的残差自适应。除任务成功率和奖励外,MSK-Bench还进一步报告了鲁棒性分析和面向生理学的诊断指标,包括激活代价、关节平滑度以及肌电(EMG)包络相似度。我们的实证研究表明:具有具身感知的探索和结构化动作表示能够提升高维肌肉空间中的任务覆盖率;模仿先验能增强与参考轨迹兼容的稳定和运动能力,但在富含接触的地形失配情况下性能下降;而在固定参考失效时,残差自适应能够恢复成功的行为。我们进一步发现,任务成功率的提升并不一定意味着生理一致性的改善,这凸显了联合评估任务性能、鲁棒性和生理行为的重要性。MSK-Bench为全身肌肉驱动的人形机器人控制提供了一个任务-方法-指标测试平台。
cs.RO / 4 / 2609.26919

Bend the Clock: Predicting Ahead to Beat Latency in Event-Based Object Detection

扭转时钟:通过提前预测克服事件目标检测中的延迟问题
Sen, Biswadeep, Cottereau, Benoit R., Cuperlier, Nicolas, Sim, Terence
Abstract
Event cameras promise low-latency perception for high-speed robotic systems, where even short delays can render detections stale by the time they inform downstream robotic decisions. Yet modern event detectors still require tens of milliseconds of computation before their predictions become available. Conventional evaluation ignores this delay by comparing predictions with annotations at the observation timestamp, even though the scene may have changed by the time those predictions are produced. We study this observation-availability mismatch in event-based multi-object detection and show that state-of-the-art event detectors degrade substantially when evaluated at prediction availability rather than observation time. To address this, we introduce ChronoFuse, a causal availability-time detector that predicts object states for when its output becomes available rather than for when its input was observed. ChronoFuse performs causal cross-time fusion over a multi-scale feature hierarchy, combining current representations with cached temporal features to expose short-term temporal cues without using future observations. The fusion pathway is lightweight, adding only 0.17 million parameters and 0.84 ms of mean end-to-end latency overhead. ChronoFuse recovers 71% of the accuracy lost to latency on 1Mpx driving data and 90.8% under rapid drone motion on FRED, nearly restoring zero-delay performance. Under the extreme motion of EV-Flying, ChronoFuse reaches 20.95 sAP, compared with 2.25 for the strongest standard event detector (9.3x gain). These results show that predicting ahead can be critical for robots operating in fast-changing scenes, including autonomous driving, agile flight, and robotic interception.
Chinese Translation
事件相机有望为高速机器人系统提供低延迟感知,但即使是短暂的延迟也可能使检测结果在传递给下游机器人决策时已经过时。然而,现代事件检测器在生成预测之前仍需要数十毫秒的计算时间。传统评估方法在观测时间戳处将预测与标注进行比较,从而忽略了这一延迟,尽管在预测结果产出时场景可能已经发生变化。我们研究了基于事件的多目标检测中这种观测与可用性不匹配的问题,并表明:当在预测可用时刻而非观测时刻进行评估时,最先进的事件检测器性能会显著下降。为解决这一问题,我们提出了 ChronoFuse,一种因果的可用时刻检测器,它预测的是其输出变为可用时的目标状态,而非其输入被观测时的目标状态。ChronoFuse 在多尺度特征层次结构上执行因果的跨时间融合,将当前表征与缓存的时间特征相结合,从而在不使用未来观测的情况下提取短期时间线索。该融合通路十分轻量,仅增加0.17百万参数和0.84毫秒的平均端到端延迟开销。ChronoFuse 在 1Mpx 驾驶数据上恢复了因延迟损失的精度的71%,在 FRED 数据集的快速无人机运动场景下恢复了90.8%,几乎还原了零延迟性能。在 EV-Flying 的极端运动场景下,ChronoFuse 达到 20.95 sAP,而最强的标准事件检测器仅为 2.25(提升9.3倍)。这些结果表明,对于在快速变化场景中运行的机器人(包括自动驾驶、敏捷飞行和机器人拦截),提前预测至关重要。
cs.RO / 5 / 2609.26989

Spiderbot: An Open-Source Energy-Efficient Hexapod with Passive Gravity Compensation

Spiderbot:一种具有被动重力补偿的开源高能效六足机器人
Sharma, Ritwik, Shah, Vimarsh, Agrawal, Saransh
Abstract
Hexapod robots can achieve static stability with fewer actuated joints than bipeds or quadrupeds, yet many platforms still use 3-DOF legs, increasing weight and continuous torque requirement with limited gain in locomotion capability on flat, inclined and moderately rough terrains. We release Spiderbot, an open-source hexapod that uses a 4-bar linkage with a passive spring to mechanically support body weight, with a 2-DOF per-leg design that substantially reduces energy consumption. This mechanism substantially offloads gravitational torque during standing stance consuming only 1.5W (reduction of over 90\% over the unsprung version and up to 96\% over other similar hexapods). The passive spring compensation extends to payloads of up to 3.25kg with no additional torque requirements. The platform enables long-duration deployments on a modest battery budget and costs under \$400, making it suitable for large-scale multi-agent experiments. We validate the locomotion capabilities of the platform with an RL policy trained in mjlab, including successful sim-to-real transfer, despite the complexity of the mechanism. The platform is evaluated on flat and rough terrains, slope up to $15^\circ$ and step obstacles. We release all the CAD files, assembling instructions, and full training and deployment code along with the model checkpoints at https://erc-bpgc.github.io/SpiderBot/.
Chinese Translation
六足机器人相比双足或四足机器人可以用更少的驱动关节实现静态稳定性,然而许多平台仍采用每腿3自由度的设计,这不仅增加了重量和持续力矩需求,而且在平坦、倾斜及中等崎岖地形上的运动能力提升有限。我们发布了Spiderbot,一种开源六足机器人,它采用带有被动弹簧的四连杆机构来机械支撑机体重量,并结合每腿2自由度的设计,显著降低了能耗。该机构在站立支撑阶段大幅卸载了重力矩,功耗仅为1.5W(相比无弹簧版本降低超过90%,相比其他类似六足机器人最高降低96%)。被动弹簧补偿可支持高达3.25kg的有效载荷,且无需额外的力矩需求。该平台可在适度的电池预算下实现长时间部署,成本低于400美元,适合大规模多智能体实验。我们在mjlab中训练了强化学习策略来验证该平台的运动能力,尽管机构较为复杂,仍成功实现了从仿真到现实(sim-to-real)的迁移。该平台在平坦和崎岖地形、最高达$15^\circ$的坡地以及台阶障碍上进行了评估。我们在 https://erc-bpgc.github.io/SpiderBot/ 上公开了所有CAD文件、装配说明、完整的训练与部署代码以及模型检查点。
cs.RO / 6 / 2609.27001

Humanoid Locomotion with a Fly-Inspired Recurrent Controller

基于果蝇启发循环控制器的人形机器人运动控制
Guan, Isabel, Zhao, Yuntian, Zhang, Dingyuan, Lyu, Shipeng
Abstract
We investigate humanoid locomotion with a fly-inspired recurrent controller and identify the pathways supporting its deployed behavior. The controller couples 3,609 continuous neural states to a simulated Unitree G1 through body-observation projections, a motor-neuron-labelled readout, and joint servos. We formulate this neural-body feedback system and evaluate a fixed checkpoint across seven terrain instances, three speeds, and three initial yaw offsets. It completes 61/63 conditions under a survival-and-forward-progress criterion; a privileged reference completes 62/63. At nominal yaw, resetting the recurrent motor state before every policy call changes success from 19/21 to 0/21. Conversely, depth and upstream-state substitutions at 252 recorded states leave actions unchanged, with zero measured descending output throughout the intact rollouts. Recorded trajectories and state-matched images connect these findings to sustained movement, lateral drift, and termination events. The study characterizes an embodied recurrent control system whose tested locomotion is supported by direct body-and-command input and carried motor state, providing a concrete basis for subsequent comparisons of circuit structure and control resources.
Chinese Translation
我们研究了采用果蝇(fly-inspired)启发循环控制器的人形机器人运动,并识别了支撑其部署行为的关键通路。该控制器通过身体观测投影、标记运动神经元标签的读出层以及关节伺服,将3,609个连续神经状态与仿真Unitree G1机器人相耦合。我们对该神经-身体反馈系统进行了形式化建模,并在七种地形实例、三种速度和三种初始偏航偏移下评估了一个固定的模型检查点。在生存与前进判据下,该控制器完成61/63项条件;一个拥有特权信息的参考基线完成62/63项。在标称偏航下,若在每次策略调用前重置循环运动状态,成功率从19/21降至0/21。相反,在252个记录状态处对深度输入和上游状态进行替换,动作保持不变,且在完整运行过程中测得的下行输出始终为零。记录的轨迹与状态匹配的图像将这些发现与持续运动、横向漂移和终止事件联系起来。本研究刻画了一个具身循环控制系统,其被测试的运动行为由直接的身体与指令输入以及所携带的运动状态支撑,为后续对回路结构与控制资源的比较提供了具体基础。
cs.RO / 7 / 2609.27003

Learning Expressive Humanoid Locomotion from Monocular Runway Videos for Robot Fashion Shows

从单目T台视频学习富有表现力的仿人机器人步态以实现机器人时装表演
Kolesnichenko, Kyrylo, Cardenas, Irvin Steve, Kim, Jong-Hoon
Abstract
Runway walking requires coordinated control of posture, stride, foot placement, and whole-body motion to effectively present clothing and convey a distinctive style. However, humanoid robots used in fashion shows typically rely on locomotion policies optimized primarily for stability and walking speed, limiting their ability to reproduce expressive, human-like runway motions. In this work, we present an end-to-end framework that transforms monocular runway videos into deployable humanoid locomotion policies through motion recovery, robot retargeting, motion correction, policy training, simulation-based evaluation, and physical deployment. We evaluate the proposed framework on the Booster K1 humanoid robot using runway-style catwalk motions. The learned policy completed every physical trial without falling, while reproducing the characteristic narrow foot placement and coordinated movement of the legs, torso, and arms. The results demonstrate that our proposed training framework enables the Booster K1 to perform stable and expressive catwalk motions, highlighting its potential for humanoid robotic applications in fashion shows and other performance-oriented scenarios.
Chinese Translation
T台走秀需要对姿态、步幅、落脚位置和全身运动进行协调控制,以有效地展示服装并传达独特的风格。然而,用于时装表演的仿人机器人通常依赖主要为稳定性和行走速度而优化的运动控制策略,限制了其再现富有表现力、类人的T台动作的能力。在本工作中,我们提出了一个端到端框架,通过动作恢复、机器人重定向、动作修正、策略训练、基于仿真的评估和物理部署,将单目T台视频转化为可部署的仿人机器人运动策略。我们在Booster K1仿人机器人上使用T台风格走秀动作对该框架进行了评估。所学得的策略在所有物理试验中均未摔倒,同时再现了特征性的窄间距落脚以及腿部、躯干和手臂的协调运动。结果表明,我们提出的训练框架使Booster K1能够执行稳定且富有表现力的走秀动作,凸显了其在时装表演及其他面向表演场景中仿人机器人应用的潜力。
cs.RO / 8 / 2609.27005

MultiPush: Learning to Rearrange with Teams of Car-Like Pushers

MultiPush:利用车队式推杆机器人学习重排(对象)
Ahn, Jeeho, Mavrogiannis, Christoforos
Abstract
We focus on the problem of rearranging multiple objects within a constrained workspace via pushing using a team of car-like robots. While the use of multiple robots offers the potential for more efficient execution, the need for conflict resolution and the kinematic constraints arising from physics, robot design, and the workspace boundary make this problem especially challenging. Our key insight is that by exploiting the structure introduced by the car-like kinematics of the domain, we could relax the problem into an ordered assignment of Dubins curves to robots. To this end, we introduce MultiPush, a reinforcement-learning based framework that jointly determines an efficient schedule of pushing tasks and their allocation to available robots by leveraging a constraint-aware traversability graph. Across extensive simulated trials with up to 14 objects and teams of two to four robots, MultiPush reduces the makespan by up to 16% compared to the baselines while requiring up to 2.9 times faster planning time. We demonstrate MultiPush on a real-world scenario involving the rearrangement of 12 objects by two and three robots (1/10-scale racecars) in a constrained space.
Chinese Translation
我们研究在受限工作空间中利用一组类车机器人通过推挤方式重排多个物体的问题。尽管使用多个机器人有可能实现更高效的执行,但冲突解决的需求以及由物理特性、机器人设计和工作空间边界带来的运动学约束,使该问题尤其具有挑战性。我们的关键洞察是:通过利用该领域类车运动学所引入的结构,可以将该问题松弛为对机器人进行有序的 Dubins 曲线分配。为此,我们提出了 MultiPush,一个基于强化学习的框架,它通过利用一种考虑约束的可通行性图(constraint-aware traversability graph),联合确定高效的推挤任务调度及其在可用机器人之间的分配。在多达 14 个物体、2 至 4 个机器人团队的广泛仿真实验中,与基线方法相比,MultiPush 将完工时间(makespan)最多降低 16%,同时规划时间最快可达基线的 2.9 倍。我们在真实世界场景中验证了 MultiPush,包括由两台和三台机器人(1/10 比例赛车)在受限空间中重排 12 个物体的实验。
cs.RO / 9 / 2609.27006

Laser-Tracker-Assisted Camera-to-Robot Calibration for Mobile Robots

激光跟踪仪辅助的移动机器人相机-机器人标定方法
Rudolph, Jan A., Kandemir, Öykü, Ulrich, Markus
Abstract
We present a laser-tracker-assisted hand-eye calibration method for camera-equipped mobile robots. The method combines laser-tracker-based 3D metrology with camera-based 2D observations. Building on our previous laser-tracker-assisted camera-to-robot calibration method for ground-observing mobile robots, we present a generalized formulation for calibrating the camera pose in the coordinate system of tracker-localized mobile robots. The new approach relaxes assumptions of our previous method on robot and camera configuration by chaining multiple calibration targets resulting in a more general approach supporting various camera-equipped mobile robot systems.
Chinese Translation
我们提出了一种针对配备相机的移动机器人的激光跟踪仪辅助的手眼标定方法。该方法将基于激光跟踪仪的三维测量与基于相机的二维观测相结合。在我们此前针对地面观测型移动机器人的激光跟踪仪辅助相机-机器人标定方法的基础上,我们提出了一种更通用的公式化方法,用于在激光跟踪仪定位的移动机器人坐标系中标定相机位姿。新方法通过串联多个标定目标,放宽了此前方法对机器人与相机配置的假设,从而成为一种更通用的方法,可支持各种配备相机的移动机器人系统。
cs.RO / 10 / 2609.27068

HiRE: Hindsight Reward Editing for Policy Finetuning

HiRE:用于策略微调的后见之明奖励编辑方法
Niu, Haoyi, Han, Zhengtao, Ji, Yufeng, Li, Zhongyu, Sreenath, Koushil
Abstract
Pre-trained robot policies always require finetuning to adapt to specific environments. Reinforcement Learning (RL) offers high performance potential because it improves action optimality rather than simply mimicking data. However, such potential depends heavily on reward quality. Sparse rewards lack process feedback, human-designed rewards are costly and biased, and semantic rewards from foundation representations are often not control-centric. We propose Hindsight Reward Editing (HiRE), a training-free framework to break this reward bottleneck. HiRE bridges the broad knowledge of foundation representation models with physical control awareness, by contrasting successful and failed trajectories in hindsight. It calibrates foundation representation models by identifying "trap states" that are predicted as high-rewarding states yet eventually result in failure, and vice versa. HiRE explicitly penalizes these traps while boosting rewards for critical successful states. This approach can be flexibly compatible with any foundation representations and RL algorithms. Experiments show that HiRE consistently outperforms other reward recipes by delivering dense, control-aware feedback that prevents value function collapse and reward hacking, thereby achieving superior sample efficiency, stable policy updates, and higher performance ceilings, e.g., at least 3x performance of the base policies. Qualitative results are at https://hire-project.github.io .
Chinese Translation
预训练机器人策略通常需要微调以适应特定环境。强化学习(RL)具有较高的性能潜力,因为它优化的是动作的最优性,而非简单地模仿数据。然而,这种潜力的发挥在很大程度上取决于奖励的质量。稀疏奖励缺乏过程反馈,人工设计的奖励成本高且存在偏差,而基于基础模型表征的语义奖励往往不以控制为中心。我们提出了后见之明奖励编辑(Hindsight Reward Editing,HiRE),一个无需训练的框架,以突破这一奖励瓶颈。HiRE 通过对成功与失败轨迹进行后见之明式的对比,将基础表征模型的广泛知识与物理控制感知能力相连接。它通过识别那些被预测为高奖励但最终导致失败的"陷阱状态"(以及相反的情形)来校准基础表征模型。HiRE 显式地对这些陷阱状态进行惩罚,同时提升关键成功状态的奖励。该方法可以灵活地兼容任意基础模型表征和强化学习算法。实验表明,HiRE 通过提供稠密的、控制感知的反馈来防止价值函数崩溃和奖励作弊(reward hacking),从而持续优于其他奖励设计方法,实现了更优的样本效率、稳定的策略更新和更高的性能上限,例如至少达到基础策略 3 倍的性能。定性结果见 https://hire-project.github.io 。
cs.RO / 11 / 2609.27070

The Gaussian Is Enough: Flow-Matching Priors Do Not Help When Fine-Tuning Large Behavior Models

高斯分布就足够了:微调大行为模型时流匹配先验并无帮助
Xu, Chen, Shah, Rishi, Kress-Gazit, Hadas, Nishimura, Haruki, Itkina, Masha
Abstract
Modern robot imitation learning increasingly relies on generative policies based on diffusion or flow-matching models, which generate actions by transforming samples from a prior distribution. A key question is whether the choice of prior matters. Replacing the standard Gaussian with a closer-to-target, non-Gaussian prior has been shown to substantially improve performance when training from scratch. A natural next step is to ask whether these gains transfer to fine-tuning pretrained Large Behavior Models (LBMs) such as LBM 1.0, $\pi_{0.5}$, and GR00T~N1.5, where one might expect even larger gains. Surprisingly, we find that this is not the case, except possibly at very low fine-tuning data fractions. Across over 100K simulation rollouts spanning all three aforementioned LBMs on 40+ tasks in two simulation platforms, and 1250 hardware rollouts on five bimanual manipulation tasks, non-Gaussian priors that are demonstrably closer to the target yield statistically indistinguishable or worse fine-tuning performance than a standard Gaussian prior. Diagnostic analyses suggest why: fine-tuned imitation learning policies converge to similar action predictions across priors, despite their fine-tuned encoder embeddings diverging substantially from the pretrained embeddings and each other. A learning-rate ablation further confirms that encoder training is the dominant factor in fine-tuning performance, substantially outweighing the effect of prior choice. We conclude with concrete directions for future research on when and why learned priors might still matter in fine-tuning. Project page: https://cxu-tri.github.io/non_gaussian_FT/
Chinese Translation
现代机器人模仿学习日益依赖基于扩散模型或流匹配模型(flow-matching)的生成式策略,这类策略通过变换来自先验分布的样本来生成动作。一个关键问题是先验的选择是否重要。已有研究表明,在从头训练时,用更接近目标分布的非高斯先验替代标准高斯先验可以显著提升性能。一个自然的后续问题是:这些收益能否迁移到对预训练大行为模型(Large Behavior Models, LBMs,如 LBM 1.0、π0.5 和 GR00T N1.5)的微调中——在这种场景下,人们或许期望获得更大的收益。令人惊讶的是,我们发现事实并非如此(除非在极低的微调数据比例下)。在涵盖上述三个 LBM、跨两个仿真平台 40 多个任务的超过 10 万次仿真测试(rollout),以及在五个双臂操作任务上的 1250 次真机测试中,明显更接近目标分布的非高斯先验带来的微调性能在统计上与标准高斯先验无显著差异,甚至更差。诊断性分析揭示了原因:尽管微调后的编码器嵌入与预训练嵌入以及彼此之间差异显著,但不同先验下微调得到的模仿学习策略最终收敛到相似的动作预测。学习率消融实验进一步证实,编码器的训练是决定微调性能的主导因素,其影响远大于先验选择的影响。最后,我们提出了未来研究的具体方向,以探究在何时以及为何学习到的先验在微调中仍然可能重要。项目页面:https://cxu-tri.github.io/non_gaussian_FT/
cs.RO / 12 / 2609.27077

Fast Direction-Conditioned Reachability for Motion Prediction Under Model Uncertainty

模型不确定性下用于运动预测的快速方向条件可达性分析
Das, Hrishav, Ornik, Melkior
Abstract
To avoid collisions, a robot must repeatedly predict where nearby agents may move, usually with an imperfect model of their dynamics. Reachable sets provide such predictions, but computing them when the system matrices themselves are uncertain can become computationally expensive and conservative for frequent replanning. Moreover, a planner often needs to know only how far an agent can move in one particular direction, for example toward the robot, rather than the complete reachable set. We propose a direction-conditioned reachability method for linear systems with uncertain state and input matrices. Given a query direction $d$, the method selects one admissible model $(A^\star,B^\star)$ whose reachable set extends nearly as far along $d$ as the reachable set of the entire uncertain model family, and then computes the reachable set of only this model with a standard reachability solver. On an uncertain linearized bicycle model, the complete selection-and-computation pipeline is about three times faster than computing the reachable set of the full uncertain family in the CORA toolbox, while its extent along $d$ is within $5\%$ of the full family's in the reported directions. We also use the method in a closed-loop multi-vehicle simulation in which the robot queries, at each replanning step, how far each nearby vehicle can move toward it, and replans to avoid the resulting sets.
Chinese Translation
为了避免碰撞,机器人必须反复预测周围智能体可能运动的位置,而通常只能使用对其动力学的不完善模型。可达集可以提供此类预测,但当系统矩阵本身存在不确定性时,计算可达集在频繁重规划的场景下可能计算代价高昂且过于保守。此外,规划器往往只需知道智能体在某一特定方向(例如朝向机器人的方向)上能移动多远,而无需完整的可达集。我们提出了一种针对状态矩阵和输入矩阵均不确定的线性系统的方向条件可达性方法。给定查询方向 $d$,该方法从所有可容许模型中选取一个模型 $(A^\star,B^\star)$,其可达集沿方向 $d$ 的延伸范围几乎等于整个不确定模型族可达集的延伸范围,然后仅使用标准可达性求解器计算该模型的可达集。在一个不确定的线性化自行车模型上,完整的“选取—计算”流程比在 CORA 工具箱中计算整个不确定模型族的可达集快约三倍,且在所报告的方向上,其沿 $d$ 的延伸范围与完整模型族结果的偏差在 5% 以内。我们还将该方法应用于闭环多车辆仿真:机器人在每次重规划步骤中查询每辆邻近车辆朝它方向能移动多远,并据此重规划以避开相应的可达集。
cs.RO / 13 / 2609.27088

Water Surface Swimming in a Centipede and its Robophysical ModeL

蜈蚣的水面游泳及其机器人物理模型
Xu, Zhaochen J., Aydin, Delfin, Mustafa, Abdullah, Levin, Margarita B., Lin, Jianfeng, Wang, Tianyu, Goldman, Daniel I.
Abstract
Elongate multi-legged robots use coordinated body waves and distributed legs to move through cluttered terrestrial environments. However, as housing actuators for independent leg control can require bulky body segments, their non-streamlined body and limb structure makes it difficult to achieve swimming capability comparable to their terrestrial locomotor performance. At the water surface, we found that the multi-legged robots we tested unexpectedly moved backward: their body waves traveled in the same direction as their displacement, i.e., swimming with a direct wave. We found similar behavior in the centipede \textit{Lithobius forficatus}, which swims with a direct body wave and periodic leg movement. To study how distributed legs contribute to direct-wave swimming, we analyze animal kinematics and develop a multi-legged robophysical model that allows independent variation of leg morphology and stiffness, body-wave direction, and leg coordination. Robophysical experiments show that direct body waves produce consistent forward motion under the tested conditions and that swimming performance depends on body--leg coordination. Additionally, directionally compliant legs increase displacement from approximately 0.08 to 0.21 body lengths per cycle relative to rigid legs under matched anti-phase actuation. These findings clarify how distributed appendages contribute to surface swimming and establish gait and morphology design principles for extending multi-legged field robots from terrestrial locomotion into aquatic environments.
Chinese Translation
细长型多足机器人通过协调的体波和分布式腿部在杂乱的陆地环境中移动。然而,由于容纳实现独立腿控制的执行器可能需要体积庞大的身体节段,其非流线型的身体和肢体结构使其难以获得与陆地运动性能相当的游泳能力。在水面,我们测试的多足机器人意外地向后移动:其体波传播方向与位移方向相同,即以正向波(direct wave)游泳。我们在石蜈蚣(Lithobius forficatus)中发现了类似行为,它以正向体波和周期性的腿部运动进行游泳。为研究分布式腿部如何促进正向波游泳,我们分析了动物的运动学特征,并开发了一个多足机器人物理模型(robophysical model),该模型可独立改变腿部形态和刚度、体波方向以及腿部协调方式。机器人物理实验表明,在测试条件下正向体波可产生一致的向前运动,且游泳性能取决于身体与腿部的协调。此外,在匹配的反相驱动下,方向性柔顺的腿部相较于刚性腿部可将每周期位移从约0.08个体长提升至0.21个体长。这些发现阐明了分布式附肢对水面游泳的贡献,并为将多足野外机器人从陆地运动拓展至水生环境确立了步态和形态设计原则。
cs.RO / 14 / 2609.27095

Intelligence Across Embodiments

跨越具身形态的智能
Ai, Bo, Christensen, Henrik I., Su, Hao
Abstract
Robotic embodiment encompasses the sensing, kinematics, dynamics, geometry, actuation, and control through which an agent physically interacts with the world. These properties vary across robots and change over time. We argue that general embodied intelligence requires learning that accumulates across these differences. Prevailing methods that engineer correspondences to bridge embodiment differences offer immediate practical gains, but their assumptions limit the scope of transfer in the long run. Instead, a more general approach should discover representations that support transfer to a larger range of embodiments as experience grows. We propose embodiment diversity as a promising axis of scaling, and identify broad learned priors as a complementary ingredient. We call for evaluations that better characterize embodiment gaps and transfer performance. More broadly, cross-embodiment learning connects the practical challenge of learning from heterogeneous robot experience with a broader scientific pursuit inspired by nature - physical intelligence that adapts and co-evolves with its embodiments to gain agency over its behavior and physical forms.
Chinese Translation
机器人具身形态涵盖了感知、运动学、动力学、几何结构、执行器与控制方式,即智能体与物理世界进行交互的途径。这些属性因机器人而异,且随时间变化。我们认为,通用的具身智能需要能够在这些差异中不断积累的学习。目前主流的方法通过人工设计对应关系来弥合具身形态差异,虽然能带来即时的实用收益,但其假设在长远上限制了迁移的适用范围。相反,更通用的方法应当发现能够支持迁移到更广泛具身形态的表示,并随经验的增长而不断扩展。我们提出将具身形态多样性作为一个有前景的扩展维度,并将广泛的习得先验知识作为互补要素。我们呼吁开展能够更好地刻画具身形态差异与迁移性能的评估。更广泛地说,跨具身形态学习将“从异构机器人经验中学习”这一实际挑战与受自然界启发的更宏大的科学追求联系在一起——即与自身具身形态相适应并协同演化,从而在行为和物理形态上获得自主性的物理智能。
cs.RO / 15 / 2609.27099

Design and Modeling of a Single-Port Three-Arm Robotic Tool for Minimally Invasive Neurosurgery

面向微创神经外科的单孔三臂机器人手术工具的设计与建模
Dana, Nazia H., Gallage, Harith S., Auta, Ismail A., Yuvaraj, Dhanvi, Qi, Ronghuai
Abstract
Surgical robots require highly dexterous and compact robotic systems capable of operating effectively within confined anatomical spaces. However, due to limited access provided by a single incision, the miniaturization and maneuverability of these robots still need to be improved. In this paper, we propose the design and modeling of a single-port three-arm robotic tool containing one major cannula (7.14 mm outer diameter (OD)) and three steerable minor cannulas (1.93 mm OD). By integrating the proposed 12 degrees-of-freedom (DoFs) steerable robotic tool with a 7-DoF robotic arm, this robotic system can potentially achieve multi-arm manipulation capability. We present the design of the steerable robotic tool consisting of tendon-driven joints controlled by a compact actuation system, derive the kinematic model, and validate both the static and kinematic models through experiments. The performance is evaluated with the root mean square error (RMSE) and mean absolute error (MAE) computed between the experimental data and the kinematic model.
Chinese Translation
手术机器人需要高度灵巧且结构紧凑的机器人系统,以在狭窄的解剖空间内有效操作。然而,由于单一切口提供的进入通道有限,此类机器人的小型化和可操作性仍需改进。本文提出了一种单孔三臂机器人工具的设计与建模,该工具包含一个主导管(外径7.14毫米)和三根可转向的次级导管(外径1.93毫米)。通过将所提出的具有12个自由度(DoFs)的可转向机器人工具与7自由度机械臂集成,该机器人系统有望实现多臂操作能力。我们介绍了由紧凑型驱动系统控制的腱驱动关节组成的可转向机器人工具的设计,推导了其运动学模型,并通过实验对静态模型和运动学模型进行了验证。通过计算实验数据与运动学模型之间的均方根误差(RMSE)和平均绝对误差(MAE)来评估系统性能。
cs.RO / 16 / 2609.27145

Planning Trajectories that Bounce: Reflection Classes for Collision-Tolerant Robots

规划可反弹的轨迹:面向耐碰撞机器人的反射类别方法
Koley, Subhadeep, Bhattacharya, Subhrajit, Saldaña, David
Abstract
Robot navigation methods tend to avoid contact, and consequently search for collision-free trajectories. For robots with high inertia and limited maneuverability, however, avoiding contact can require substantial steering effort and time, even when interactions with surrounding surfaces could be safely exploited. In this paper, we develop a planning method that deliberately uses controlled wall reflections to generate trajectories that can be easier and more efficient to execute than purely collision-free motion. We consider planar navigation in environments where a mobile robot is permitted to bounce off surrounding surfaces. To represent the resulting alternatives, we construct a reflection-augmented state graph in which paths are partitioned into distinct classes according to the sequence of walls used for reflection. This representation enables systematic enumeration of reflection strategies and identification of the lowest-cost path within each class. We show that, although a reflecting path cannot be shorter than the shortest collision-free path, it can reduce execution time and actuation effort by replacing costly changes in heading with controlled environmental interactions. The planned trajectories are executed using a contact-aware sampling-based controller with the robot's full dynamics. In our experiments, we demonstrate that in our simulated test scenario, the best reflecting class can reduce time and control effort. Our results show that controlled contact can provide dynamically advantageous navigation strategies that are excluded by conventional collision-avoidance formulations.
Chinese Translation
机器人导航方法通常倾向于避免接触,因而寻求无碰撞轨迹。然而,对于惯性大、机动性受限的机器人而言,避免接触可能需要大量的转向操作和时间开销,即使与周围表面的交互本可以被安全地加以利用。在本文中,我们提出了一种规划方法,该方法有意利用受控的墙壁反射来生成比纯无碰撞运动更易于执行且效率更高的轨迹。我们考虑了允许移动机器人从周围表面反弹的平面环境中的导航问题。为了表示由此产生的备选方案,我们构建了一个反射增强状态图,其中路径根据用于反射的墙壁序列被划分为不同的类别。这种表示使得能够系统地枚举反射策略,并识别每个类别中代价最低的路径。我们证明,尽管反射路径不可能比最短无碰撞路径更短,但它可以通过用受控的环境交互代替代价高昂的航向变化,从而减少执行时间和驱动开销。规划出的轨迹采用具备接触感知的基于采样的控制器并结合机器人的完整动力学来执行。在实验中,我们证明了在仿真测试场景中,最优的反射类别能够减少时间和控制开销。我们的结果表明,受控接触可以提供传统避碰公式所排除的、具有动力学优势的导航策略。
cs.RO / 17 / 2609.27154

HINT-Blimp: Human INTent Inference from Multimodal Cues for Robotic Blimps

HINT-Blimp:面向机器人飞艇的基于多模态线索的人类意图推断
Koley, Subhadeep, Greenberg, Benjamin, Shao, Yifei Simon, Aceros, Juan, Figueroa, Nadia, Saldaña, David
Abstract
In human-robot interaction, traditional interfaces such as joysticks and handheld tablets introduce latency into navigation tasks and require the operator's explicit attention on the device, instead of the robot. We propose a new human-robot interaction framework in which a human communicates intent directly through sparse multimodal signals such as physical pushes and spoken commands. Human intent is represented as a parameterized linear dynamical system (LDS) that encodes the desired goal and motion behavior. The robot estimates this intent (parameters) online using a particle filter, where each particle represents a candidate LDS hypothesis and is reweighted online as new information becomes available. We validate this framework on a robotic blimp, whose inherent compliance and collision tolerance make it well-suited for repeated physical interaction. Experiments with multiple participants across 300 trials show that combining pushes and voice commands identifies the intended goal in 86% of trials within at most five interactions, with most trials resolved in two. The inferred dynamical systems can also produce curved trajectories that avoid obstacles known only to the human.
Chinese Translation
在人机交互中,传统的接口(如操纵杆和手持平板电脑)会给导航任务引入延迟,并且需要操作者将显性注意力放在设备上而非机器人上。我们提出了一种新的人机交互框架,其中人类通过稀疏的多模态信号(如物理推动和语音指令)直接传达意图。人类意图被表示为一个参数化的线性动态系统(LDS),用于编码期望的目标和运动行为。机器人使用粒子滤波器在线估计该意图(参数),其中每个粒子代表一个候选的LDS假设,并随着新信息的获得在线进行重新加权。我们在一台机器人飞艇上验证了该框架,其固有的柔顺性和耐碰撞特性使其非常适合反复的物理交互。在多参与者共300次试验的实验中表明,结合推动和语音指令能够在最多五次交互内于86%的试验中识别出预期目标,且大多数试验在两次交互内即可完成。推断出的动态系统还能生成避开障碍物的曲线轨迹,而这些障碍物只有人类知晓。
cs.RO / 18 / 2609.27160

Fine Wrist Control as a Marker of Surgical Teleoperation Expertise

精细手腕控制作为手术遥操作专长的标志
Gale, Mary Kate, Shobayashi, Shujiro, Satpathy, Sangeet, Davidor, Nitsan, Nisky, Ilana, Okamura, Allison
Abstract
Unlike most intensely physical pursuits, surgical robotic teleoperation training focuses primarily on task outcomes rather than surgeon body posture or biomechanics during task completion. Toward the question of the role of biomechanics in surgical expertise, we sought to characterize the articular motion of expert teleoperators as compared to novice users. Twenty-seven novices and nine experts completed a non-medical cylinder-on-peg transfer task while their upper limb biomechanics were recorded via motion trackers. During more difficult motions, experts stabilized their wrist motion more than novices, while maintaining adequate range of motion in their shoulder and elbow and completing the task significantly faster than novices. This marker of expertise suggests the importance of attention to user biomechanics during teleoperation of surgical robots.
Chinese Translation
与大多数高强度体能活动不同,手术机器人遥操作训练主要关注任务结果,而非外科医生在完成任务过程中的身体姿态或生物力学。为了探究生物力学在外科专长中的作用,我们试图刻画专家遥操作者与新手用户相比的关节运动特征。27名新手和9名专家完成了一项非医学的圆柱-套柱转移任务,同时通过运动追踪器记录其上肢生物力学数据。在较困难的动作中,专家比新手更能够稳定其手腕运动,同时保持肩部和肘部足够的活动范围,并且完成任务的速度显著快于新手。这一专长标志表明,在手术机器人遥操作中关注使用者生物力学的重要性。
cs.RO / 19 / 2609.27167

Median Temporal Ensembling: Training-Free Robust Aggregation for Action-Chunked Visuomotor Policies

中位数时序集成:一种无需训练的动作分块视觉运动策略鲁棒聚合方法
Jiang, Yuhang
Abstract
Action-chunked visuomotor policies predict overlapping trajectories, so every executed action is covered by several predictions. Temporal ensembling smooths execution by combining these predictions with an exponentially weighted mean. One corrupted prediction can move the aggregate without bound: its breakdown point is 0. We use adversarial corruption to stress this deployed aggregator and to compare two kinds of guarantee. A metric guarantee bounds the response to a perturbation of a given size. A combinatorial guarantee instead bounds the damage when at most q of the M candidates covering a timestep are corrupted, whatever their size. Encoder adversarial fine-tuning recovers 44% of the loss under the published patch attack, but only 7.3% after the attacker's step size is increased. By contrast, the coordinate-wise median of the same candidate set keeps its recovered fraction flat as attack optimisation increases. Median temporal ensembling costs one line and requires no retraining. Across 25 (configuration, corruption-level) combinations it is never worse than the mean and is significantly better in 15. It also transfers to a second policy class, and it recovers performance under a failure with no attacker in the loop at all: camera frames that arrive blank. Its effect on clean data is configuration-dependent, from -0.04 to +0.07. We also give the boundary: corruption that shifts every covering prediction by the same amount is invisible to this whole family of statistics, and no equivariant aggregator can remove it.
Chinese Translation
动作分块(action-chunked)视觉运动策略预测的是相互重叠的轨迹,因此每个执行的动作都被多个预测所覆盖。时序集成(temporal ensembling)通过对这些预测进行指数加权平均来平滑执行过程。然而,一个被污染的预测可以使聚合结果无界地偏移:其崩溃点为0。我们使用对抗性污染对该已部署的聚合器进行压力测试,并比较两种类型的保证。度量保证(metric guarantee)限定了对给定大小扰动的响应;而组合保证(combinatorial guarantee)则限定当覆盖某一时间步的M个候选预测中至多q个被污染时的损害程度,无论污染的大小如何。编码器对抗微调在已发表的补丁攻击下恢复了44%的性能损失,但在攻击者步长增大后仅恢复7.3%。相比之下,对同一候选集进行坐标级中位数聚合,随着攻击优化的增强,其恢复比例保持不变。中位数时序集成只需修改一行代码,且无需重新训练。在25种(配置,污染程度)组合中,它从不劣于均值聚合,并在其中15种组合上显著更优。它还能迁移到第二类策略上,并在一种完全不存在攻击者的失效情形下——相机帧到达时为空白——恢复性能。它对干净数据的影响依赖于具体配置,范围从-0.04到+0.07。我们还给出了该方法的边界:若污染使所有覆盖同一时间步的预测发生相同的偏移,则这一整类统计量均无法察觉,且任何等变聚合器都无法消除这种污染。
cs.RO / 20 / 2609.27188

Learning Dissipative Dynamics with Dissipativity-by-Construction Discrete-Time Neural Networks

基于构造性耗散离散时间神经网络的耗散动力学学习
Luong, Tuan, Moon, Hyungpil
Abstract
Dissipativity is a fundamental system-theoretic property closely related to stability, passivity, and input--output stability, and is particularly important in robotics, where learned dynamics models are often embedded within feedback control loops. However, most existing approaches for learning dissipative dynamics are based on continuous-time formulations, which require ODE solvers during training or inference and can therefore be computationally expensive. Moreover, because practical implementations are inherently discrete-time, direct discretization of a continuous-time passive system does not necessarily preserve passivity, motivating the need for explicit discrete-time guarantees. This study proposes a method for learning incrementally dissipative dynamics from input--output time-series data using a deep multilayer perceptron formulated directly in discrete time. Through a constrained parameterization and a dedicated training procedure, the proposed model guarantees incremental dissipativity by construction rather than through regularization. Lyapunov-based analysis establishes the corresponding dissipativity and stability guarantees, while simulations on robotic dynamical systems demonstrate competitive prediction accuracy, computational efficiency, and consistent preservation of incremental dissipativity compared with baseline methods.
Chinese Translation
耗散性是一个与稳定性、无源性和输入-输出稳定性密切相关的系统性基本性质,在机器人学中尤为重要,因为学习得到的动力学模型通常嵌入在反馈控制回路中。然而,现有的多数耗散动力学学习方法均基于连续时间公式,需要在训练或推理过程中使用常微分方程(ODE)求解器,因而计算代价较高。此外,由于实际实现本质上是离散时间的,对连续时间无源系统的直接离散化并不一定能保持无源性,因此需要显式的离散时间保证。本研究提出一种直接以离散时间形式构建的深度多层感知机方法,用于从输入-输出时间序列数据中学习增量耗散动力学。通过约束参数化和专门的训练流程,所提出的模型通过构造而非正则化的方式保证增量耗散性。基于Lyapunov的分析建立了相应的耗散性与稳定性保证,同时在机器人动力学系统上的仿真表明,与基线方法相比,该方法具有竞争力高的预测精度、计算效率,并始终如一地保持增量耗散性。
cs.RO / 21 / 2609.27218

NaviScale: Generating Large-Scale Semantic Map Datasets for Object Navigation

NaviScale:面向目标导航的大规模语义地图数据集生成
Lan, Chuanlin, Zheng, Yanwei, Liu, Weijian, Zhou, Zhitong, Zhang, Xiao, Zhuang, Fuzhen, Yu, Dongxiao
Abstract
Embodied navigation requires spatial representations that generalize across unseen environments, yet collecting large amounts of annotated data from real 3D environments is difficult. We propose NaviScale for semantic-map-based object navigation (ObjectNav), whose predictor can be trained on pairs of partial and complete semantic maps without reconstructing a complete 3D environment for every training sample. The framework generates large-scale semantic map training data by composing floorplans of real homes with room-level semantic and obstacle maps extracted from MP3D and HM3DSem. NaviScale increases data diversity in two ways: inter-room scaling increases floorplan-level structural diversity, while intra-room scaling fills each fixed floorplan with different combinations of room maps matched by room category. Visibility through Ray Casting (VisRC) converts the composed maps into partial observations that account for field of view, sensing range, and occlusion. The resulting dataset contains 192,000 semantic maps generated from 24,000 floorplans associated with 12,794 properties. With 300k training iterations and the training and inference settings described in this paper, the system reaches 64.3% SR and 34.8% SPL on HM3D, together with 43.1% SR and 16.8% SPL on MP3D, without changing the prediction architecture. Additional experiments evaluate the quality of the composed maps, the effects of semantic-segmentation errors, and deployment on a physical robot.
Chinese Translation
具身导航需要能够在未见环境中泛化的空间表示,然而从真实3D环境中收集大量标注数据十分困难。我们提出了NaviScale,用于基于语义地图的目标导航(ObjectNav),其预测器可以通过部分语义地图与完整语义地图的配对数据进行训练,而无需为每个训练样本重建完整的3D环境。该框架通过将由真实住宅户型图与从MP3D和HM3DSem中提取的房间级语义地图和障碍物地图组合,生成大规模语义地图训练数据。NaviScale从两个方面提升数据多样性:房间间扩展(inter-room scaling)增加户型图层面的结构多样性,而房间内扩展(intra-room scaling)则在固定的户型图中填充按房间类别匹配的不同房间地图组合。基于光线投射的可见性方法(Visibility through Ray Casting, VisRC)将组合后的地图转换为考虑视场、感知范围和遮挡的部分观测。所得数据集包含由24,000张户型图生成的192,000张语义地图,涉及12,794个房产。在30万次训练迭代以及本文所述的训练与推理设置下,该系统在HM3D上达到64.3%的成功率(SR)和34.8%的SPL,在MP3D上达到43.1%的SR和16.8%的SPL,且未改变预测架构。额外的实验评估了组合地图的质量、语义分割误差的影响以及在实际机器人上的部署。
cs.RO / 22 / 2609.27219

Vision-Based Control of a Tether-Suspended Aerial Radiation Sensing Payload

基于视觉的系留悬吊空中辐射传感载荷控制
Snider, Ian, Quiter, Brian J., Rofors, Emil, Mueller, Mark W.
Abstract
Aerial radiation surveys achieve higher sensitivity when the radiation detector is held close to the ground. Detector sensitivity falls off roughly with the inverse square of the distance to the source, so a detector flown high is slower to reach a given minimum detectable activity. Flying the vehicle low puts the propellers near the ground, where downwash can disturb the surveyed area and resuspend contaminated particulates. Tether suspension decouples the detector from the vehicle altitude, but leaves the payload unactuated and only indirectly controllable. We therefore present a vision-based control approach for an aerial sensing payload suspended on a tether beneath a heavy-lift drone. Because a survey plan is decided as radiation detections arrive, we design a pilot aid for commanding the survey trajectory manually with a handheld transmitter. The controller regulates the payload, rather than the vehicle, onto that trajectory. The system uses onboard sensors with a downward-facing camera fixed to the drone body tracking a ring marker on the payload. A four-state Kalman filter estimates the tether swing angles and rates from payload bearing measurements, and a linear quadratic regulator with integral action takes the payload position as the regulated output. In outdoor flight tests under wind, the payload-aware controller reduced payload tracking error during transit by 20% when compared against a vehicle-referenced baseline, with the cost of higher peak error on arrival at a waypoint.
Chinese Translation
当辐射探测器贴近地面时,空中辐射勘测的灵敏度更高。探测器灵敏度大致随探测器到辐射源距离的平方倒数衰减,因此飞行高度较高的探测器需要更长时间才能达到给定的最低可探测活度。让飞行器低空飞行则会使螺旋桨靠近地面,其下洗气流会扰动勘测区域并使污染颗粒物重新悬浮。系绳悬吊方式使探测器与飞行器的高度解耦,但载荷因此缺乏直接执行机构,只能间接控制。为此,我们提出了一种基于视觉的控制方法,用于悬吊在重型无人机下方系绳上的空中传感载荷。由于勘测计划需根据辐射探测结果实时确定,我们设计了一种辅助驾驶系统,使操作员能够通过手持遥控器手动指令生成勘测轨迹。控制器将载荷(而非飞行器本身)调节至该轨迹上。该系统使用机载传感器,并在无人机机体上固定一个朝下的摄像头,用于跟踪载荷上的环形标记。一个四状态卡尔曼滤波器根据载荷方位角测量估计系绳摆角及其速率,并采用带积分作用的线性二次调节器(LQR),以载荷位置作为被控输出。在户外有风环境下的飞行测试中,与以飞行器为参考的基线方法相比,这种面向载荷的控制器将载荷在飞行转移阶段的跟踪误差降低了20%,但代价是到达航路点时的峰值误差有所增大。
cs.RO / 23 / 2609.27240

A Quasi-Direct-Drive Underactuated Asymmetric Hand for Dexterous and Efficient Grasping and Manipulation

一种用于灵巧高效抓取与操作的准直驱欠驱动非对称机械手
Davis, Benjamin, Kidder, Chase, Stuart, Hannah S.
Abstract
In this paper, we present the Berkeley QUAD (Quasi-direct-drive, Underactuated, Asymmetric Design) Hand, a four-finger anthropomorphic robotic hand with 11 degrees of freedom and 8 degrees of actuation. The design utilizes QDD actuation at the base of each finger, enabling high force transparency for dexterous, adaptive performance. However, the low torque density of these actuators traditionally presents major issues with size, weight, and thermal limits. We overcome this by applying bio-inspired asymmetry, delegating dexterity to the radial fingers through individual QDD actuation, and strength to the ulnar finger through an underactuated, compliantly coupled transmission driven by a larger QDD motor. A novel preloaded, linkage-based transmission permits this ulnar coupling in a way that preserves human-like workspace reachability. Under light loads, the ulnar motor drives the third (middle) finger directly for dexterity while the fourth (ring) finger mirrors its motion. However, under larger loads, the middle finger complies while the motor drives the ring finger further downwards and inwards toward the center of the grasp to apply better closure forces. Hardware evaluations validate this architecture, demonstrating that the hand achieves 29 out of 33 Feix taxonomy grasps and exhibits backdrive forces as low as 50 g for delicate interactions. Additionally, the underactuated fourth finger improves grasp closure and provides the spatial efficiency necessary for larger actuation, yielding up to a 96-fold reduction in heat generation during sustained loading. Webpage: https://benudavis.github.io/berkeley-quadhand/
Chinese Translation
本文提出了Berkeley QUAD(准直驱、欠驱动、非对称设计,Quasi-direct-drive, Underactuated, Asymmetric Design)机械手,这是一款具有11个自由度和8个驱动自由度的四指仿人机械手。该设计在每个手指根部采用准直驱(QDD)驱动,实现了高力透明性,从而具备灵巧、自适应的性能。然而,此类执行器较低的扭矩密度历来在尺寸、重量和热限制方面存在重大问题。我们通过引入受生物启发的非对称设计来解决这一问题:将灵巧性分配给桡侧手指,通过独立的QDD驱动实现;将力量分配给尺侧手指,通过由更大的QDD电机驱动的欠驱动柔性耦合传动实现。一种新颖的预紧连杆式传动机构以保留类人工作空间可达性的方式实现了这种尺侧耦合。在轻负载下,尺侧电机直接驱动第三指(中指)以实现灵巧操作,同时第四指(无名指)镜像其运动;而在较大负载下,中指发生柔顺变形,同时电机进一步向下和向内驱动无名指,使其朝向抓取中心以施加更好的闭合力。硬件评估验证了该架构,表明该机械手可实现Feix分类中33种抓取方式中的29种,且反向驱动力低至50克,适用于精细交互。此外,欠驱动的第四指改善了抓取闭合性,并为更大的执行器提供了所需的空间效率,在持续负载下发热量最多降低96倍。网页:https://benudavis.github.io/berkeley-quadhand/
cs.RO / 24 / 2609.27247

Memory That Changes Action Is Not Memory That Guides It: Counterfactual Auditing of History-Conditioned Robot Policies

改变动作的记忆并非引导动作的记忆:对基于历史条件的机器人策略的反事实审计
Zhang, Jiajie, Xiang, Yankai, Chen, Changhao
Abstract
A robot returning a block to its origin tray may encounter two task-consistent pasts that reconverge to the same current input but warrant different actions. Yet memory-policy evaluations often rely on task success or action change under memory perturbation, neither of which establishes that memory guides the decision. We propose the \textbf{Counterfactual Memory Audit (CMA)}, an evaluation protocol that crosses two histories at a verified-identical present, queries a frozen policy under common randomness, and evaluates each saved action under both pasts. This separates memory sensitivity, warranted choice, matched-world physical value, and per-pair reliability. On Mem-0, every audited Put Back pair changes action, but only $20/64$ pairs are fully reliable; at a later Swap decision, all paired actions change while both memories select the same branch. Native interventions further show closed-loop influence: replacing the history bank redirects behavior toward the replaced content, while restoring a 4096-byte protected anchor recovers $38.9$ points of Swap success lost to injected bank faults. On a dual-arm physical platform, memory changes saved actions, yet five of nine completed Put Back manipulations reach the wrong target. These results show that a robot can remember and react without reliably using memory to choose the behavior its past warrants. CMA provides a decision-level audit for distinguishing these cases.
Chinese Translation
一个将积木放回其原始托盘的机器人,可能遇到两种与任务一致但不同的历史:它们重新汇聚到相同的当前输入,却需要采取不同的动作。然而,记忆策略的评估通常依赖任务成功率或记忆扰动下的动作变化,这两者都无法证明记忆确实在引导决策。我们提出了**反事实记忆审计(Counterfactual Memory Audit, CMA)**,这是一种评估协议:在经过验证完全相同的当前状态下交叉两条历史,在共同随机性下查询冻结的策略,并在两种历史下分别评估每个已保存的动作。该方法将记忆敏感性、合理的动作选择、匹配世界下的物理价值以及逐对可靠性分离开来。在 Mem-0 上,每一个被审计的 Put Back(放回)动作对都发生了动作变化,但仅有 $20/64$ 对是完全可靠的;在后续的 Swap(交换)决策中,所有配对动作均发生变化,而两种记忆却选择了相同的分支。原生干预进一步展示了闭环影响:替换历史库会使行为转向被替换的内容,而恢复一个 4096 字节的受保护锚点可以挽回因注入历史库故障而损失的 $38.9$ 个百分点的 Swap 成功率。在双臂物理平台上,记忆确实改变了已保存的动作,但九次完成的 Put Back 操作中有五次到达了错误的目标。这些结果表明,机器人可以在记住并做出反应的同时,并未可靠地利用记忆来选择其历史所应有的行为。CMA 提供了一种决策层面的审计方法,用于区分这些情形。
cs.RO / 25 / 2609.27269

Banana Kick: Response-Informed Skill Evolution for Humanoid Soccer

香蕉球:面向人形机器人足球的响应感知技能演化
Zhang, Hao E., Geng, Ruize, Haque, Raihan, Zbiss, Khalil, Luo, Guanyang, Wang, Hui-ping, Tseng, H. Eric, Zhao, Ding
Abstract
Humanoid kicking requires coordinated whole-body motion and precise contact, while a banana kick demands contact mechanics that generate ball spin and aerodynamic curvature. Motion imitation provides a reliable ordinary-kick prior, but reinforcement learning may improve shot speed and placement accuracy without changing the underlying kicking technique. Adapting this prior to a qualitatively different contact-rich skill can fail even when the reward is dense and optimization remains stable. The failure occurs when the task objective is locally flat over the current policy's responses. We term this condition first-order learning starvation. To address it, we propose response-informed skill evolution (RISE), a closed-loop objective-continuation method for policy adaptation. RISE ranks bounded objective changes using response sensitivity estimated from cached rollouts and accepts updates only when they produce verified response progress while preserving kicking reliability. Our analysis shows that rescaling a saturated spin reward cannot recover first-order sensitivity at zero spin, whereas adapting coupled contact responses can provide a learnable path to spin generation. We integrate RISE into a humanoid kicking pipeline under calibrated contact and Magnus-force aerodynamics, and sim-to-real transfer. Experiments show that RISE evolves the ordinary kick into a high-spin curved kick with 11.55 rad/s mean ball spin, improves the mean evaluation score by 19.8% over a learning-progress curriculum, and raises joint target attainment from 15.2% to 50.9%. Ablations and response diagnostics support the mechanism, while 30 motion-capture-recorded physical trials demonstrate consistent hardware transfer of the learned curved kick. Project website: https://haozhang-thu.github.io/bananakick/
Chinese Translation
人形机器人踢球需要协调的全身运动和精确的触球,而香蕉球则要求触球力学能够产生球的旋转和空气动力学弧线。运动模仿提供了可靠的普通踢球先验,但强化学习可能在不改变底层踢球技术的情况下提高射门速度和落点精度。将这一先验适应于质性不同的富接触技能时,即使奖励稠密且优化过程稳定,也可能失败。这种失败发生在任务目标在当前策略的响应上是局部平坦的时候。我们将这一条件称为一阶学习饥饿(first-order learning starvation)。为解决该问题,我们提出响应感知技能演化(Response-Informed Skill Evolution, RISE),这是一种用于策略自适应的闭环目标延拓方法。RISE利用从缓存的经验轨迹中估计的响应敏感度对有界目标变更进行排序,并仅在更新能产生经验证的响应进展且保持踢球可靠性时才接受更新。我们的分析表明,对饱和的旋转奖励进行重新缩放无法在零旋转状态下恢复一阶敏感度,而调整耦合的触球响应则可以提供一条可学习的旋转生成路径。我们将RISE集成到人形机器人踢球流水线中,涵盖校准触球、马格努斯力空气动力学以及仿真到现实(sim-to-real)迁移。实验表明,RISE将普通踢球演化为平均球转速达11.55 rad/s的高旋转弧线球,相比学习进度课程使平均评估分数提高了19.8%,并将关节目标达成率从15.2%提升至50.9%。消融实验和响应诊断支持了该机制的有效性,30次动捕记录的物理实验证明了所学习的弧线球在真实硬件上的一致迁移。项目网站:https://haozhang-thu.github.io/bananakick/
cs.RO / 26 / 2609.27275

BranchDrive: A Branch-Structured Dataset for Action-Conditioned Driving Prediction

BranchDrive:一个用于动作条件驾驶预测的分支结构数据集
Khanzada, Feeza Khan, Sridhar, Sudarshan, Kwon, Jaerock
Abstract
Most autonomous-driving datasets record only the action executed by a behavior policy and the single future that followed, providing limited supervision for comparing alternative ego decisions. We introduce BranchDrive, a branch-structured CARLA dataset and benchmark that pairs one canonical pre-decision history with one nominal expert future and twelve physically executed intervention futures spanning acceleration, braking, and left- and right-steering policies at three magnitudes. Each intervention lasts 2.5 s and is followed by expert recovery. Following control-compliance, modality-completeness, replay-fidelity, and action-leakage audits, the frozen benchmark contains 606 independent branch groups and 7,878 associated trajectories. We evaluate prediction of six continuous short-horizon outcomes and a ten-step ego trajectory using action-only, history-only, structured, visual, multimodal, and privileged bird's-eye-view models. On the held-out test split, the structured history-and-action model achieves a macro normalized mean absolute error of 0.5036 and an average displacement error of 2.2042 m, significantly outperforming both restricted baselines. In full-information offline evaluation, its outcome-derived selector increases balanced policy value from 0.5364 to 0.5704 and reduces normalized regret from 0.2674 to 0.1495 relative to the frozen action prior. However, a validation-calibrated minimum-separation guard rejects every intervention, showing that conservative execution remains unresolved. BranchDrive therefore supports action-conditioned short-horizon prediction and fixed-bank offline decision evaluation, but does not establish exact causal effects, binary safety prediction, or closed-loop safety improvement.
Chinese Translation
大多数自动驾驶数据集仅记录行为策略所执行的动作及其后发生的单一未来,为比较不同的自车决策提供了有限的监督。我们提出BranchDrive,一个分支结构的CARLA数据集与基准,它将一条标准的决策前历史与一条名义专家未来轨迹以及十二条实际执行的干预未来轨迹配对,这些干预涵盖加速、制动以及左转和右转策略,并各取三种幅度。每次干预持续2.5秒,随后由专家进行恢复。在通过控制合规性、模态完整性、回放保真度和动作泄漏审计之后,该冻结基准包含606个独立分支组和7,878条相关轨迹。我们使用仅动作、仅历史、结构化、视觉、多模态以及特权鸟瞰图模型,对六个连续短时程结果的预测和十步自车轨迹进行评估。在留出测试集上,结构化历史与动作模型取得了0.5036的宏平均归一化平均绝对误差和2.2042米的平均位移误差,显著优于两个受限基线。在全信息离线评估中,相对于冻结的动作先验,其基于结果的选择器将平衡策略价值从0.5364提升至0.5704,并将归一化遗憾从0.2674降低至0.1495。然而,经验证集校准的最小分离保护机制拒绝了所有干预,表明保守执行问题仍未解决。因此,BranchDrive支持动作条件的短时程预测和固定库离线决策评估,但并不能确立精确的因果效应、二元安全性预测或闭环安全性提升。
cs.RO / 27 / 2609.27308

EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics

EmbodiedSWE:面向长时程灵巧机器人技术的编码智能体
You, Haoxiang, Shen, Zeyu, Liu, Yilang, Zheng, Zhicheng, Zha, Lihan, Yamazaki, Kashu, Zhang, Mingtong, Huang, Suning, Sun, Jiankai, Chen, Qianzhong, He, Lucy, Liu, Kaiyuan, Chang, Haoran, Fragkiadaki, Katerina, Shah, Dhruv, Schwager, Mac, Henderson, Peter, Abraham, Ian, Xu, Canwen
Abstract
We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EMBODIEDSWE-BENCH, a simulation benchmark for coding agents spanning contact-rich manipulation, deformable objects, and long-horizon tasks requiring up to half an hour of continuous interaction. We find that frontier coding agents can solve complex long-horizon tasks and transfer prior solutions across both tasks and embodiments. We also design supporting tools that help agents more effectively solve these tasks. However, the resulting solutions require substantial iterative interaction and are typically specialized to individual task instances. We therefore introduce EMBODIEDSWE-GEN, which expands a single solution from coding agent into large diverse trajectories for training a VLA. VLA performance improves with more generated demonstrations, and agent-aided diversification improves generalization to held-out task variations. We also show that a VLA finetuned solely on coding-agent-generated simulation demonstrations completes a long-horizon task on real robot. Together, our framework uses coding agents to solve complex robotics tasks and turn verified solutions into scalable supervision for robot policies.
Chinese Translation
我们研究了面向长时程灵巧机器人技术的编码智能体,并探讨其解决方案能否为学习通用机器人策略提供可扩展的监督信号。为验证这一点,我们开发了 EMBODIEDSWE-BENCH,一个面向编码智能体的仿真基准,涵盖接触丰富的操作、可变形物体以及需要长达半小时连续交互的长时程任务。我们发现,前沿编码智能体能够求解复杂的长时程任务,并将先前的解决方案迁移到不同任务和不同机器人本体上。我们还设计了辅助工具,帮助智能体更有效地解决这些任务。然而,由此产生的解决方案需要大量的迭代交互,且通常仅针对单个任务实例。为此,我们提出 EMBODIEDSWE-GEN,将编码智能体的单一解决方案扩展为大规模、多样化的轨迹,用于训练视觉-语言-动作(VLA)模型。VLA 的性能随生成演示数量的增加而提升,且借助智能体实现的多样化能提升对未见任务变化的泛化能力。我们还证明,仅在编码智能体生成的仿真演示上微调的 VLA,能够在真实机器人上完成长时程任务。总体而言,我们的框架利用编码智能体解决复杂机器人任务,并将经过验证的解决方案转化为机器人策略的可扩展监督信号。
cs.RO / 28 / 2609.27312

Turning Safety into Competence: Minimally Exploitable Robot Policies via Safety-Filtered Reinforcement Learning

将安全性转化为竞争力:基于安全过滤强化学习的最小可利用机器人策略
Wu, Ruihan, Yang, Rui, Oh, Donggeon David, Nguyen, Duy, Hu, Haimin
Abstract
Robots deployed for competitive tasks must outmaneuver their opponents without sacrificing safety. Existing approaches, including safe reinforcement learning (RL), train a single policy to achieve task success and avoid failures simultaneously. This coupling can complicate training and leave the learned policy exploitable by deliberate attacks. We propose Safety to Competence (S2C), a two-stage RL framework that separates safety synthesis from competitive task learning. We formulate competitive interactions as safety-critical Markov games and prove that perfect filtering preserves policy non-exploitability when all players commit to safe maneuvers. S2C learns a robust safety filter via adversarial RL, embeds it in the environment during task policy training, and retains the same filter at deployment. In simulated touchdown games, S2C outperforms eight safe RL baselines, achieving the highest win rate and Elo rating, and the lowest exploitability. Hardware stress tests against a human opponent confirm S2C's competence.
Chinese Translation
部署于竞争性任务的机器人必须在保证安全的同时战胜对手。现有方法,包括安全强化学习(RL),通过训练单一策略同时实现任务成功和避免失败。这种耦合会使训练复杂化,并使学到的策略容易被蓄意攻击所利用。我们提出了安全到竞争力(Safety to Competence, S2C)框架,这是一个两阶段强化学习框架,将安全性综合与竞争性任务学习相分离。我们将竞争性交互形式化为安全关键马尔可夫博弈,并证明当所有参与者都遵循安全机动时,完美过滤能够保持策略的不可利用性。S2C 通过对抗性强化学习学习一个鲁棒的安全过滤器,在任务策略训练期间将其嵌入环境中,并在部署时保留相同的过滤器。在模拟的触地得分游戏中,S2C 优于八个安全强化学习基线方法,取得了最高的胜率和 Elo 等级分,以及最低的可利用性。针对人类对手的硬件压力测试证实了 S2C 的竞争力。
cs.RO / 29 / 2609.27314

CoRe-WAM: Correspondence-Aligned Temporal Residuals for World Action Models

CoRe-WAM:面向世界动作模型的对应关系对齐时序残差方法
Zhou, Bin, Liu, Jialong, Wang, Jianan, Chen, Changhao, Chen, Kani
Abstract
Comparing current and past observations helps robots understand scene changes and select subsequent actions during manipulation. However, comparing visual features at the same image location can mix different scene content when objects or the camera move. We introduce CoRe-WAM, a world-action model that incorporates correspondence-aligned visual changes through a parameter-efficient temporal interface. Its TraceDelta module uses correspondences from a frozen tracking model to transport historical visual features to current locations before computing signed differences in a shared pretrained feature space. Correspondence thus determines which historical content is compared with the present, rather than entering the policy as a separate trajectory representation. A lightweight adapter converts these differences into validity-gated residuals that supplement current visual conditioning, allowing the policy to use recent changes alongside current-scene information. Built on Motus, CoRe-WAM keeps the pretrained backbone weights frozen and optimizes 1.59 million parameters. With a 5,000-update adaptation budget, CoRe-WAM achieves 92.22% clean success across 50 RoboTwin 2.0 tasks, 3.56 percentage points above Motus; on randomized evaluation, it achieves 89.60% success, a 2.58-point gain. Integrating TraceDelta into a StarVLA-based policy improves clean success from 58.10% to 67.62%, supporting transfer of the temporal interface beyond Motus.
Chinese Translation
比较当前与过去的观测有助于机器人在操作过程中理解场景变化并选择后续动作。然而,在相同图像位置比较视觉特征时,当物体或相机发生移动,可能会混合不同的场景内容。我们提出了CoRe-WAM,这是一种通过参数高效的时序接口引入对应关系对齐的视觉变化的世界动作模型(world-action model)。其TraceDelta模块利用冻结的跟踪模型提供的对应关系,将历史视觉特征传输到当前的位置,然后在共享的预训练特征空间中计算带符号的差值。由此,对应关系决定了将哪些历史内容与当前内容进行比较,而不是作为独立的轨迹表示输入策略。一个轻量级适配器将这些差值转化为经有效性门控的残差,用以补充当前的视觉条件输入,使策略能够结合当前场景信息利用近期变化。CoRe-WAM基于Motus构建,保持预训练主干网络权重冻结,仅优化159万参数。在5000次更新的适配预算下,CoRe-WAM在50个RoboTwin 2.0任务上取得92.22%的干净场景成功率,比Motus高出3.56个百分点;在随机化评估中,其成功率达89.60%,提升2.58个百分点。将TraceDelta集成到基于StarVLA的策略中,可将干净场景成功率从58.10%提升至67.62%,表明该时序接口可迁移至Motus之外的模型。
cs.RO / 30 / 2609.27330

A Sample-Based Approach for Hierarchical Information-Theoretic Compression of Probabilistic Occupancy Grids

一种基于样本的概率占据栅格层次化信息论压缩方法
Jin, Zhenyu, Larsson, Daniel T.
Abstract
We develop a sample-based framework for constructing information-driven hierarchical multi-resolution representations of probabilistic occupancy grids. Recent methods compute information-optimal abstractions via dynamic-programming-based exhaustive recursions, which become computationally prohibitive for large-scale grids and are ill-suited to robotics applications. To address this limitation, we introduce a sample-based strategy inspired by Monte Carlo Tree Search (MCTS) that incrementally constructs hierarchical abstractions through statistical estimation rather than exhaustive enumeration. The proposed method is anytime in nature, allowing computation to be terminated at any stage to produce a valid compressed representation. We compare our approach with the information-optimal Q-tree search algorithm and demonstrate its effectiveness in rapidly generating abstractions of large real-world probabilistic occupancy grids.
Chinese Translation
我们开发了一个基于样本的框架,用于构建概率占据栅格的信息驱动层次化多分辨率表示。现有方法通过基于动态规划的穷举递归计算信息最优抽象,但对于大规模栅格而言计算代价过高,且不适用于机器人应用。为解决这一局限,我们提出了一种受蒙特卡洛树搜索(MCTS)启发的基于样本的策略,通过统计估计而非穷举枚举来增量式地构建层次化抽象。所提出的方法具有随时(anytime)特性,允许在任一阶段终止计算并生成有效的压缩表示。我们将该方法与信息最优的Q树搜索算法进行比较,并展示了其在快速生成大规模真实世界概率占据栅格抽象方面的有效性。
cs.RO / 31 / 2609.27338

DUGM-R: Uncertainty-Aware Dynamic Grid Mapping and Risk-Triggered Recovery for Learned Local Navigation

DUGM-R:面向学习型局部导航的不确定性感知动态栅格地图与风险触发恢复机制
Feng, Haoyun, Rubio-Solis, Adrian, Guo, Zhaodong, Mylonas, George
Abstract
Learned local navigation in crowded indoor environments is sensitive to how dynamic obstacle motion is represented, while collision-prone behaviour may persist after nominal policy training. We present a risk-aware reinforcement-learning framework that addresses these two issues through an uncertainty-aware Dynamic Uncertainty Grid Map (DUGM) and a modular post-training recovery mechanism. DUGM combines local occupancy, estimated obstacle motion, and motion-estimation uncertainty in a robot-centric representation. After the nominal policy is frozen, a finite-horizon Risk Value Function (RVF) is trained from nominal rollouts and used to trigger a dedicated recovery policy when continued nominal execution is predicted to be collision-prone. Experiments in a held-out NVIDIA Isaac Sim clinical-logistics benchmark show that uncertainty-aware dynamic representation improves nominal navigation over static and deterministic alternatives, while the recovery mechanism further mitigates residual collision-prone behaviour. The complete framework is also deployed directly on a TurtleBot3 without policy fine-tuning, retraining, or site-specific adaptation, retaining the performance trend observed in simulation. These results indicate that uncertainty-aware dynamic representation and post-training recovery provide complementary mechanisms for improving learned local navigation.
Chinese Translation
在拥挤的室内环境中,学习型局部导航对动态障碍物运动的表示方式十分敏感,而在名义策略训练完成后,易发生碰撞的行为仍可能持续存在。我们提出了一种风险感知的强化学习框架,通过不确定性感知的动态不确定性栅格地图(Dynamic Uncertainty Grid Map, DUGM)和一个模块化的训练后恢复机制来解决这两个问题。DUGM在以机器人为中心的表示中融合了局部占据信息、估计的障碍物运动以及运动估计的不确定性。在名义策略被冻结后,基于名义策略的滚动轨迹训练一个有限时域的风险价值函数(Risk Value Function, RVF),当预测继续执行名义策略可能发生碰撞时,该函数用于触发一个专用的恢复策略。在留出的NVIDIA Isaac Sim临床物流基准测试中的实验表明,不确定性感知的动态表示相比静态和确定性替代方案提升了名义导航性能,而恢复机制进一步缓解了残留的易碰撞行为。完整框架还无需策略微调、重新训练或针对特定场地的适配,即可直接部署在TurtleBot3上,并保持了在仿真中观察到的性能趋势。这些结果表明,不确定性感知的动态表示与训练后恢复机制为改进学习型局部导航提供了互补的手段。
cs.RO / 32 / 2609.27340

Spatial and Semantic Reasoning for LLM-Driven Robot Navigation via MCP

基于MCP的LLM驱动机器人导航的空间与语义推理
Lee, Jungsoo, Park, Jaegyun, Kim, Wansoo
Abstract
Large language models (LLMs) are increasingly used as natural-language interfaces for robotic systems, yet their integration with Robot Operating System (ROS)-based navigation remains limited by two gaps. First, navigation data such as occupancy grids are represented as raw geometric messages that are difficult for LLMs to use directly as spatial or semantic context. Second, adding LLM-driven capabilities often requires custom wrappers or robot-specific interfaces, limiting reuse across systems. To address these challenges, we propose a non-invasive framework that connects LLM reasoning with ROS-based navigation through a navigation-oriented representation layer, exposed through the Model Context Protocol (MCP) as standardized, reusable tools so that any MCP-compatible LLM can access them without robot-specific wrappers. The visual map modules transform occupancy grids into metric, pose-aware images for goal reasoning, while the semantic annotation modules record waypoint-level observations with robot poses. We evaluate the framework on three tasks: autonomous mapping, spatial reasoning-based navigation, and semantic reasoning-based navigation. The results show that the evaluated LLM backends use these representations to achieve over 97% map coverage and select spatial or semantic navigation targets from natural-language instructions in a simulated indoor environment. This demonstrates representation-mediated LLM navigation without modifying the existing ROS navigation stack.
Chinese Translation
大语言模型(LLM)越来越多地被用作机器人系统的自然语言接口,然而其与基于机器人操作系统(ROS)的导航的集成仍受限于两个不足。首先,占据栅格等导航数据以原始几何消息的形式表示,LLM难以将其直接用作空间或语义上下文。其次,添加LLM驱动的功能通常需要自定义封装或特定于机器人的接口,限制了跨系统的复用。为应对这些挑战,我们提出了一种非侵入式框架,通过面向导航的表示层将LLM推理与基于ROS的导航相连接,并借助模型上下文协议(MCP)将其以标准化、可复用的工具形式暴露出来,使任何兼容MCP的LLM都能在无需机器人特定封装的情况下访问这些工具。视觉地图模块将占据栅格转换为度量化的、具有位姿感知的图像以用于目标推理,而语义标注模块则记录带机器人位姿的路径点级观测。我们在三个任务上对该框架进行了评估:自主建图、基于空间推理的导航以及基于语义推理的导航。结果表明,所评估的LLM后端利用这些表示,在仿真室内环境中实现了超过97%的地图覆盖率,并能根据自然语言指令选择空间或语义导航目标。这证明了无需修改现有ROS导航栈即可实现由表示中介的LLM导航。
cs.RO / 33 / 2609.27342

BladeMaster: Real-Time Robotic Cutting Simulation with Online-Generated Persistent Discontinuities

BladeMaster:基于在线生成持久不连续面的实时机器人切割仿真
Yang, Zhanyu, Chen, Yunuo, Huang, Yanjia, Masterjohn, Joseph, Yang, Yin, Jiang, Chenfanfu
Abstract
Cutting changes both the shape and topology of deformable objects, making accurate simulation challenging for robotic manipulation. A simulator must track the cutting tool as a cut develops, preserve the resulting discontinuities after tool withdrawal, and enable newly exposed surfaces to interact with the tool and with each other. Existing formulations often prescribe cut surfaces in advance or couple material separation to auxiliary geometric fields. We introduce BladeMaster, a GPU-accelerated cutting framework based on the total Lagrangian material point method (TLMPM). Our key idea is to encode the cutting history directly on material points through persistent side labels generated online from the blade geometry. These labels govern particle-grid coupling, preserving connectivity within intact material while preventing spurious coupling across cut faces after tool withdrawal. Our formulation supports progressive and intersecting cuts without predefined cut surfaces or particle duplication. Material-material contact enables cut surfaces to recontact and slide against each other without reconnecting, while two-way tool-material coupling allows material reaction forces to influence tool motion. Experiments demonstrate tool-driven cutting followed by manipulation, with faster-than-real-time performance on representative tasks.
Chinese Translation
切割会同时改变可变形物体的形状和拓扑结构,这使得机器人操作中的精确仿真极具挑战性。仿真器需要在切割过程中实时追踪切割刀具,在刀具撤离后保留所产生的不连续面,并使新暴露的表面能够与刀具以及彼此之间发生交互。现有方法通常需要预先指定切割面,或将材料分离与辅助几何场相耦合。我们提出了BladeMaster,一种基于全拉格朗日物质点法(TLMPM)的GPU加速切割仿真框架。我们的核心思想是通过对刀刃几何在线生成持久的侧向标签,将切割历史直接编码在物质点上。这些标签控制粒子-网格耦合,在保持完整材料内部连通性的同时,防止刀具撤离后在切割面之间产生虚假耦合。我们的方法无需预定义切割面或粒子复制,即可支持渐进式和交叉式切割。材料-材料接触使切割面能够相互重新接触并滑动而不会重新连接,同时双向的刀具-材料耦合允许材料的反作用力影响刀具运动。实验展示了由刀具驱动的切割及后续操作,并在代表性任务上实现了快于实时的性能。
cs.RO / 34 / 2609.27346

Reflection-Aware Reasoning for Non-Line-of-Sight Pedestrian Localization

面向非视距行人定位的反射感知推理方法
Park, Byeonggyu, Jeon, Mingu, Kim, Seong-Woo
Abstract
Reliable localization of non-line-of-sight (NLOS) pedestrians is critical for safe urban autonomous driving, yet it remains highly challenging in ego-dynamic outdoor environments, where ego-vehicle motion makes radar multipath propagation complex and noisy. In this paper, we present a reflection-aware framework for NLOS pedestrian localization with a moving ego-vehicle in outdoor testbed scenarios. Our framework fuses front-view camera images and 2D radar point clouds to infer reflection orders and reflective surface distributions in bird's-eye-view space. It then uses physics-guided ray tracing to reconstruct distorted reflection paths and localize the hidden pedestrian. We validate the framework in outdoor testbed scenarios under ego-dynamic conditions. The results demonstrate the effectiveness of the proposed framework for NLOS pedestrian localization with a moving ego-vehicle.
Chinese Translation
非视距(NLOS)行人的可靠定位对于城市自动驾驶的安全至关重要,但在自车动态的户外环境中仍然极具挑战性,因为自车运动会使雷达多径传播变得复杂且充满噪声。本文提出了一种反射感知框架,用于在户外测试场景中对移动自车条件下的非视距行人进行定位。该框架融合前视相机图像与2D雷达点云,在鸟瞰图(bird's-eye-view)空间中推断反射阶数和反射面分布;随后利用物理引导的射线追踪(ray tracing)重建发生畸变的反射路径,并实现隐藏行人的定位。我们在自车动态条件下的户外测试场景中验证了该框架。结果表明,所提出的框架对移动自车条件下的非视距行人定位是有效的。
cs.RO / 35 / 2609.27358

Omnidirectional Amphibious Locomotion via Internal Mass Actuation

基于内部质量驱动的全向水陆两栖运动
Weaver, Niko, Xia, Boxi, Lo, Li-Yu, Huang, Yuhao, Chen, Boyuan
Abstract
Field robots must traverse varied terrain and obstacles while remaining robust to water, debris, vegetation, and physical contact. We present MARBLE, a fully enclosed omnidirectional amphibious rolling robot driven entirely by internal mass redistribution. Three mutually orthogonal linear sliders shift internal masses to generate body rotation, while an orientation-aware controller maps planar velocity commands into slider positions. A rigid spherical shell encloses all active mechanisms and simultaneously serves as the terrestrial contact surface, buoyant enclosure, and mounting structure for passive fins that enable water-surface propulsion. Rotation of the same shell architecture hence produces rolling on land and surface propulsion in water without mechanical reconfiguration or separate locomotion actuators. The spherical morphology further allows the robot to accommodate changes in body orientation and contact location during direct interactions with terrain and obstacles. We evaluate MARBLE through omnidirectional locomotion characterization, traversal across heterogeneous terrestrial environments, aquatic surface locomotion, land-water transitions, and deliberate obstacle interactions. These experiments demonstrate how a single enclosed mechanical architecture can combine omnidirectional mobility, cross-medium locomotion, and tolerance to environmental contact. MARBLE provides a compact design for field mobility across heterogeneous terrain, obstacles, and land-water transitions. We will open-source all software and hardware design. Our website is https://generalroboticslab.com/MARBLE
Chinese Translation
野外机器人需要在多种地形和障碍物上移动,同时能够承受水、碎屑、植被和物理接触等环境因素的影响。我们提出了MARBLE,一种完全封闭的全向水陆两栖滚动机器人,其运动完全由内部质量重新分布驱动。三个相互正交的直线滑轨通过移动内部质量来产生机体旋转,同时一个具备姿态感知能力的控制器将平面速度指令映射为滑轨位置。刚性球形外壳将所有主动机构完全封闭,并同时充当陆地接触面、浮力密封舱以及被动鳍片的安装结构,从而实现水面推进。因此,同一外壳结构的旋转既能在陆地上产生滚动运动,也能在水中产生水面推进,无需机械重构或独立的运动执行器。球形形态还使机器人能够在与地形和障碍物直接交互时适应机体姿态和接触位置的变化。我们通过全向运动特性表征、异构陆地环境穿越、水面运动、水陆过渡以及刻意的障碍物交互实验对MARBLE进行了评估。这些实验表明,单一封闭机械架构可以同时实现全向移动性、跨介质运动以及对环境接触的容忍能力。MARBLE为跨异构地形、障碍物和水陆过渡的野外移动提供了一种紧凑的设计。我们将开源所有软件和硬件设计。我们的网站是 https://generalroboticslab.com/MARBLE
cs.RO / 36 / 2609.27363

From LiDAR Maps to Visual Localization: Unified Visual Association for Robust Point-Line-Plane Pose Estimation

从LiDAR地图到视觉定位:用于鲁棒点-线-平面位姿估计的统一视觉关联
Zhao, Wentao, Chen, Zikun, Niu, Yihe, Chen, Haoyu, Wang, Jingchuan
Abstract
Camera localization in a prior LiDAR map provides a persistent geometric reference for long-term robotic navigation, yet remains challenging because of the substantial modality gap between camera images and point-cloud maps. We present a unified localization framework that makes the LiDAR map visually addressable rather than relying on a dedicated image-LiDAR correspondence model. Map geometry and reflectivity are rendered into LiDAR-derived quasi-images with explicit 2D-3D provenance, enabling camera observations and rendered map views to share mature visual features and matchers for both global localization and continuous pose tracking. Point and line correspondences are established through this common visual interface, while the retained provenance recovers metric LiDAR geometry and line-supported planar constraints for pose estimation. To improve robustness under ambiguous associations and weak geometry, we further introduce a distribution-aware, observability-complementary optimization strategy. Instead of reducing matching ambiguity to a scalar confidence, candidate association distributions are propagated into directional pose-information uncertainty, and reliable structural factors are selectively reinforced according to their ability to complement the currently weak pose directions. Experiments on the EuRoC MAV benchmark and self-collected real-world sequences demonstrate accurate global localization and robust continuous 6-DoF tracking using only a pre-built LiDAR map as the persistent prior, including under severe illumination variations and dynamic occlusions.
Chinese Translation
基于先验LiDAR地图的相机定位为长期机器人导航提供了持久的几何参考,但由于相机图像与点云地图之间存在显著的模态差异,这一任务仍然充满挑战。我们提出了一种统一的定位框架,使LiDAR地图具有可视觉寻址性,而无需依赖专门的图像-LiDAR对应模型。地图的几何与反射率被渲染为具有显式2D-3D来源信息的LiDAR派生准图像,使相机观测与渲染的地图视图能够共享成熟的视觉特征与匹配器,从而实现全局定位与连续位姿跟踪。点与线的对应关系通过这一共同的视觉接口建立,同时保留的来源信息可恢复具有度量尺度的LiDAR几何以及由直线支撑的平面约束,用于位姿估计。为提高在模糊关联与弱几何条件下的鲁棒性,我们进一步引入了一种分布感知、可观测性互补的优化策略。该策略并非将匹配歧义简化为标量置信度,而是将候选关联分布传播为方向性的位姿信息不确定性,并根据各结构因子对当前弱位姿方向的补充能力,选择性地强化可靠的结构因子。在EuRoC MAV基准数据集及自采集的真实场景序列上的实验表明,仅以预建LiDAR地图作为持久先验,即可实现精确的全局定位和鲁棒的连续6自由度跟踪,包括在剧烈光照变化和动态遮挡条件下。
cs.RO / 37 / 2609.27381

CoPRE: Improving Sensitivity in Proprioceptive Contact Detection for Low-Cost Robot Arms

CoPRE:提升低成本机械臂本体感觉接触检测灵敏度的方法
Zhu, Yuxiao, Li, Jinzhou, Dong, Yifei, Suhail, Muhammad, Yang, Chunyuan, Luo, Xinyuan, Li, Haoyu, Chen, Boyuan, Cheng, Xianyi
Abstract
Contact detection during robotic manipulation allows robots to recognize unexpected contact and adapt their motion accordingly. However, in low-cost robot arms without dedicated force or tactile sensors, detecting weak contacts from proprioception is challenging because the resulting changes in joint-level proprioceptive signals can be small compared to normal variation and noise caused by robot motion itself. We introduce Contact-free Proprioceptive Response Estimation (CoPRE), improving proprioceptive contact detection sensitivity using only contact-free motion, without additional force sensors, contact labels, or analytical dynamics models. CoPRE estimate the expected joint torques under contact-free motion from proprioceptive state history and commanded motion, while removing recent observations that may already reflect contact. It then computes the residual between the expected and observed joint torque estimates, and maps this residual to a contact score using a noise-weighted Jacobian. Real-robot experiments on ARX Arm and Unitree G1 show that CoPRE achieves 74.1% and 82.2% recall on the tested contact trials, compared with 0%/0% on ARX and 16.3%/42.2% on G1 for the learned torque-prediction and inverse-dynamics baselines. CoPRE also reaches 90% detection rate for pushing force at 3.5 N on ARX and 5.5 N on G1. To demonstrate the downstream utility of our method, we implement belief-space manipulation planning for obstacle-aware object placement and book insertion where detected contacts update the spatial belief and enable the robot to retreat from blocked motions, adjust its pose, and retry. Project website at https://copre-arm.github.io
Chinese Translation
机器人操作过程中的接触检测使机器人能够识别意外接触并相应地调整其运动。然而,对于没有专用力传感器或触觉传感器的低成本机械臂而言,从本体感觉中检测微弱接触具有挑战性,因为接触引起的关节级本体感觉信号变化与机器人自身运动造成的正常变化和噪声相比可能非常微小。我们提出了无接触本体感觉响应估计(Contact-free Proprioceptive Response Estimation,CoPRE),仅利用无接触运动数据即可提升本体感觉接触检测的灵敏度,无需额外的力传感器、接触标签或解析动力学模型。CoPRE 根据本体感觉状态历史和指令运动估计无接触运动下预期的关节力矩,同时剔除可能已反映接触信息的近期观测数据。随后,它计算预期与观测到的关节力矩估计之间的残差,并通过噪声加权雅可比矩阵(Jacobian)将该残差映射为接触分数。在 ARX Arm 和 Unitree G1 机器人上的真实实验表明,CoPRE 在测试的接触试验中分别达到了 74.1% 和 82.2% 的召回率,而学习型力矩预测和逆动力学基线方法在 ARX 上的召回率为 0%/0%,在 G1 上为 16.3%/42.2%。此外,CoPRE 在 ARX 和 G1 上分别对 3.5 N 和 5.5 N 的推力达到了 90% 的检测率。为了展示该方法的下游应用价值,我们实现了基于信念空间的操作规划,用于考虑障碍物的物体放置和书籍插架任务,其中检测到的接触会更新空间信念,使机器人能够从受阻的运动中退回、调整姿态并重试。项目网站:https://copre-arm.github.io
cs.RO / 38 / 2609.27394

Automotive mmWave Spinning Radar Place Recognition with Spatially Gated Feature-Correlation Representation

基于空间门控特征相关表示的车载毫米波旋转雷达位置识别
Rahman, Saimunur, Shrestha, Sagun Singh, Khamis, Abdelwahed, Moghadam, Peyman
Abstract
Automotive spinning FMCW radar provides dense, $360^\circ$ sensing and remains reliable under poor illumination and adverse weather, making it well-suited to autonomous navigation. Place recognition uses these observations to identify previously visited locations for re-localization and long-term navigation. However, heading changes appear as circular shifts in the polar radar representation, and conventional global aggregation can lose relationships among radar responses that are important for distinguishing similar places. We propose SGCA-Net, a spinning radar place recognition framework that combines rotation-robust feature extraction with Spatially Gated Correlation Aggregation (SGCA). SGCA learns spatial weights to reduce the influence of unstable and ambiguous radar regions, while aggregating pairwise correlations among local responses to preserve informative feature relationships. Experiments on the MulRan dataset show that SGCA-Net consistently outperforms SOTA methods across urban, campus, and open-road environments, while remaining robust to substantial heading variation. Evaluation on the HeRCULES dataset further demonstrates that SGCA-Net generalizes to unseen environments and radar sensors without fine-tuning.
Chinese Translation
车载旋转FMCW雷达可提供稠密的360°感知,并在光照不佳和恶劣天气条件下保持可靠性,因此非常适合自主导航。位置识别利用这些观测来识别先前到访过的位置,以实现重定位和长期导航。然而,航向变化在极坐标雷达表示中表现为循环平移,且传统的全局聚合方式可能丢失对区分相似地点至关重要的雷达响应之间的关系。我们提出了SGCA-Net,一种结合旋转鲁棒特征提取与空间门控相关聚合(Spatially Gated Correlation Aggregation, SGCA)的旋转雷达位置识别框架。SGCA通过学习空间权重来降低不稳定和模糊的雷达区域的影响,同时聚合局部响应之间的成对相关性,以保留有信息量的特征关系。在MulRan数据集上的实验表明,SGCA-Net在城市、校园和开放道路环境中始终优于最先进(SOTA)方法,同时对显著的航向变化保持鲁棒性。在HeRCULES数据集上的评估进一步表明,SGCA-Net无需微调即可泛化至未见过的环境和雷达传感器。
cs.RO / 39 / 2609.27439

Compressed delayed-information projection for six-degree-of-freedom underwater vehicle navigation under delayed acoustic positioning

延迟声学定位下六自由度水下航行器导航的压缩延迟信息投影方法
Li, Shuyue, López-Benítez, Miguel, Lim, Eng Gee, Ma, Fei, Dong, Qian, Cao, Mengze, Yu, Limin, Qin, Xiaohui
Abstract
Delayed acoustic positioning packets constrain historical navigation states, but a current-time update evaluates them against a mismatched state, whereas exact rewind/replay re-executes the intervening estimator history. This paper introduces compressed delayed-information projection (CDIP), a causal 15-state error-state Kalman filter (ESKF) treatment for delayed-acoustic unmanned underwater vehicle (UUV) navigation. CDIP retains a source-epoch snapshot and the historical-to-current cross-covariance, then projects the delayed source-epoch acoustic correction directly to the current state without full rewind/replay. Exact fixed-lag rewind/replay out-of-sequence-measurement (OOSM) processing serves as a high-fidelity accuracy reference. In 154 usable paired recordings at a fixed 1.5-s acoustic delay without an outage, CDIP reduced mean trajectory-position root-mean-square error (RMSE) from 1.062 m for the baseline to 0.456 m (57.1%). Its 0.456-m mean was 1.03% higher than the 0.451-m replay mean, while its measured mean per-update runtime was 99.2% lower (approximately 127-fold). A separate predeclared sweep across six fixed delays, with 30 paired recordings per delay, and a truth-supported 9-D consistency analysis bound the interpretation. Additional targeted experiments showed near-replay trajectory accuracy across 50-300-s acoustic outages while preserving sub-millisecond update cost. CDIP therefore provides a compact delayed-information treatment with an empirical accuracy-computation trade-off under the evaluated configuration; the evidence does not establish statistical equivalence or non-inferiority relative to replay.
Chinese Translation
延迟到达的声学定位数据包约束了历史导航状态,而当前时刻的更新是在与该数据包不匹配的状态上对其进行评估,精确的回退/重放方法则需重新执行中间时段的估计器历史。本文提出压缩延迟信息投影(Compressed Delayed-Information Projection, CDIP),一种面向延迟声学无人水下航行器(UUV)导航的因果15状态误差状态卡尔曼滤波器(ESKF)处理方法。CDIP 保留源时刻的状态快照以及历史状态与当前状态之间的互协方差,随后将源时刻的延迟声学修正直接投影到当前状态,而无需完整的回退/重放。作为高精度参考,本文采用精确的固定滞后回退/重放乱序量测(OOSM)处理方法。在声学延迟固定为1.5秒且无中断的154组可用配对记录中,CDIP 将平均轨迹位置均方根误差(RMSE)从基线的1.062米降至0.456米(降低57.1%)。其0.456米的平均值仅比重放方法的0.451米平均值高1.03%,而实测的单次更新平均运行时间则低99.2%(约127倍)。此外,一项预先声明的扫描实验(涵盖六个固定延迟、每个延迟30组配对记录)以及基于真值的9维一致性分析对结果的解释进行了限定。额外的针对性实验表明,在50至300秒的声学中断期间,CDIP 仍能达到接近重放方法的轨迹精度,同时保持亚毫秒级的更新开销。因此,在所评估的配置下,CDIP 提供了一种简洁的延迟信息处理方法,并展现出实证意义上的精度—计算量权衡;但相关证据并未证明其与重放方法具有统计学等价性或非劣性。
cs.RO / 40 / 2609.27449

X2Real: an eXtensive simulation benchmark for real-world generalist policies

X2Real:面向现实世界通用策略的扩展仿真基准
Ruan, Lian, Yang, Jade, Gao, Sherphylan, Gao, Felix, Liang, Kyson, Liu, Galen, Wu, Ligo, Jin, Lane, Gu, Guu, Xie, Bevan, Yan, Cloud, Yuan, Zongzi, Luo, Kino, Chen, Emma, Chen, Shuwen, Ping, Yang, Guo, Miles, Sun, Rain, Zhang, Kayden, Du, Alex, Wu, Ruihai, Hao, Liang, Li, Zhaoshuo, Gan, Roy, Wang, Hao, Wang, Qian
Abstract
Generalist robot manipulation policies have developed rapidly, yet their reliable evaluation remains challenging due to fundamental flaws in existing simulation benchmarks: prominent sim-to-real gaps, narrow task coverage, and unfair evaluation caused by ambiguous training-test pipelines. Prior works only partially resolve these issues and lack simultaneous faithfulness, diversity, and fairness, while static benchmark designs fail to sustain long-term policy development. We present X2Real, an evolvable simulation benchmark for faithfully evaluating the real-world performance of robotic manipulation policies based on Nvidia Isaac Lab-Arena. Following three core principles (faithfulness, diversity, and fairness), X2Real calibrates simulation visual and physical properties to align with real hardware, achieving a 0.84 linear correlation between simulated and real-robot evaluation results. It features a comprehensive taxonomy with 10 capability dimensions and 44 hierarchical long-horizon tasks, covering basic manipulation skills and advanced capacities such as visual grounding, language understanding, and bimanual control. We further adopt multi-axis domain randomization and strictly disjoint training-evaluation pipelines to mitigate benchmark exploitation and ensure credible evaluation. Powered by a custom physical domain-specific language, the Mana simulation ecosystem supports modular task design and iterative performance analysis, alongside a nearly 300-hour annotated simulation trajectory dataset. X2Real offers a faithful, diverse, and fair evolving evaluation infrastructure, effectively bridging the sim-to-real evaluation gap and supporting the advancement of generalist robotic manipulation policies.
Chinese Translation
通用机器人操作策略发展迅速,但由于现有仿真基准存在根本性缺陷——显著的仿真到现实差距(sim-to-real gap)、狭窄的任务覆盖范围,以及训练-测试流程不明确导致的不公平评估——其可靠评估仍然面临挑战。已有工作仅部分解决了这些问题,无法同时兼顾忠实性、多样性和公平性,而静态的基准设计也难以支撑长期的策略发展。我们提出了X2Real,一个可演化的仿真基准,基于Nvidia Isaac Lab-Arena构建,用于忠实地评估机器人操作策略在现实世界中的性能。遵循忠实性、多样性和公平性三大核心原则,X2Real对仿真的视觉与物理属性进行校准以对齐真实硬件,使仿真评估结果与真实机器人评估结果之间达到0.84的线性相关性。该基准提供了包含10个能力维度和44个层次化长时程任务的全面任务体系,涵盖基础操作技能以及视觉定位、语言理解和双臂控制等高级能力。我们进一步采用多轴域随机化和严格隔离的训练-评估流程,以缓解基准被过度利用的问题,确保评估的可信度。依托自定义的物理领域专用语言,Mana仿真生态系统支持模块化任务设计与迭代式性能分析,并配有近300小时带标注的仿真轨迹数据集。X2Real提供了一个忠实、多样且公平的演化式评估基础设施,有效弥合了仿真到现实的评估差距,推动了通用机器人操作策略的发展。
cs.RO / 41 / 2609.27450

BEE: Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models

BEE:基于视觉-语言-动作模型的干预自适应现实世界强化学习
Zhao, Weihui, Yan, Xiaohan, Wan, Zunian, Du, Xuan, Chi, Zhaozhan, Mao, Jianbo, Wu, Ruipu, Yang, Rushuai, Li, Houlin, Yang, Shukai, Wu, Jing, Yan, Yuxiang, Liu, Yongcheng, Li, Chuankang, Ren, Guanghui, Shan, Wei, Yao, Maoqing
Abstract
Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale errors undo all prior progress. Online reinforcement learning (RL) can optimize exactly these actions, but free exploration is far too costly on real robots, which makes human corrections indispensable. However, existing online RL methods for VLAs either cannot incorporate such corrections or fold them into undifferentiated supervision. Yet human corrections are not uniformly noisy but reliable along some action dimensions and variable along others. Building on this, we introduce BEE, an intervention-adaptive framework for real-world RL on a frozen VLA that lets the policy go BEyond Expert imitation. We formulate human corrections not as actions to reproduce but as evidence about a constraint: a Correction Model predicts how a human would correct a given VLA proposal and how consistent the correction is along each action dimension. This predicted consistency sets the per-dimension tightness of a constraint on policy optimization. Where corrections are consistent the policy stays close to the human, and where they vary, the constraint relaxes. We evaluate BEE on three real-world manipulation tasks and one LIBERO-Pro simulation task at a matched online-data budget. BEE attains the highest success rate on every task, 91.2% on average against 57.5% for RLT and 42.1% for DSRL, and the lowest human intervention rate on all real-world tasks.
Chinese Translation
视觉-语言-动作(VLA)模型能够处理长时程操作任务,然而其成功依赖于少数对精度至关重要的阶段——在这些阶段中,毫米级的误差会使之前的所有进展付诸东流。在线强化学习(RL)恰好可以优化这些关键动作,但在真实机器人上进行自由探索的代价过于高昂,这使得人类纠正不可或缺。然而,现有的面向VLA的在线RL方法要么无法纳入这类纠正,要么将其混入不加区分的监督信号中。事实上,人类纠正并非均匀带噪,而是在某些动作维度上可靠,在另一些维度上多变。基于这一观察,我们提出了BEE,一个针对冻结VLA的现实世界强化学习干预自适应框架,使策略能够超越(BEyond)专家模仿。我们将人类纠正不视为需要复现的动作,而是视为关于约束的证据:一个纠正模型(Correction Model)预测人类将如何纠正给定的VLA提议,以及该纠正沿每个动作维度的一致性。这种预测的一致性决定了策略优化中各维度约束的严格程度:在纠正一致的维度上,策略保持贴近人类;在纠正多变的维度上,约束则放松。在相同的在线数据预算下,我们在三个真实世界操作任务和一个LIBERO-Pro仿真任务上评估了BEE。BEE在所有任务上均取得最高成功率,平均为91.2%,而RLT为57.5%,DSRL为42.1%;并且在所有真实世界任务上,BEE的人类干预率最低。
cs.RO / 42 / 2609.27463

Safety-Filtered Distributed Koopman-MPC

安全过滤的分布式Koopman模型预测控制
Zhang, Shengjun, Li, Wenhao, Lin, Zhenxin, Sun, Zhenglong
Abstract
Distributed model predictive control (DMPC) often constructs both predictions and collision constraints from neighbor trajectories, so packet loss can remove both. We separate these roles: received trajectories drive Koopman-MPC, while local sensing and shelf geometry define a hard-constrained quadratic program (QP) that projects the applied input. Its radial demand is the least constant acceleration that keeps a supporting-plane clearance nonnegative throughout one zero-order-hold interval. Complementary pair rows recover the coupled demand without exchanging safety decisions. We give an intersample separation theorem under bounded snapshot and directional plant errors, an exact max-min test for simultaneous local feasibility, and a sensing-radius condition for switching interaction graphs. Anticipatory high-order rows may be relaxed for performance, but the finite-hold rows contain no safety slack. Matched eight-robot warehouse simulations use a frozen Koopman model, nonlinear drift, bounded inputs and speed, shelf constraints, a 120 ms control period, and packet dropout. The full controller is collision-free in 20/20 matched trials and reaches 160/160 robot goals; predictive Koopman-MPC without the final projection is collision-free in 1/20 trials. All 38,400 full-method hard-row sets pass the online feasibility test, and every local QP solves. Five-stream fleet sweeps are collision-free and hard-row feasible through 16 robots; the 20-robot boundary fails only after the online margin turns negative, while the reconstructed per-agent critical path remains below the sampling period. Bounded-sensing and differential-drive tests provide additional deployment stress.
Chinese Translation
分布式模型预测控制(DMPC)通常基于邻居轨迹同时构建预测和碰撞约束,因此数据包丢失会同时破坏两者。我们将这两种功能分离:接收到的轨迹用于驱动Koopman-MPC,而本地感知与货架几何信息则定义了一个硬约束二次规划(QP),对实际施加的控制输入进行投影。其径向需求是使支撑平面间隙在整个零阶保持区间内保持非负的最小恒定加速度。互补对行在无需交换安全决策的情况下恢复耦合需求。我们在有界的快照误差与方向性受控对象误差条件下给出了采样间隔间的分离定理,给出了同时局部可行性的精确max-min检验,以及交互图切换的感知半径条件。用于预判的高阶行可以为性能而放松,但有限保持区间内的行不包含任何安全松弛量。在匹配的八机器人仓储仿真中,使用了冻结的Koopman模型、非线性漂移项、有界的输入与速度、货架约束、120毫秒控制周期以及数据包丢失。完整控制器在20/20次匹配实验中均无碰撞,并达成160/160个机器人目标;而不含最终投影的预测式Koopman-MPC仅在1/20次实验中无碰撞。全部38,400组完整方法的硬约束行均通过在线可行性检验,且每个局部QP均可求解。五车队规模扫描测试表明,在16个机器人规模下均无碰撞且满足硬约束行可行性;20个机器人的边界情形仅在线余量变为负值后才失败,而重构的每代理临界路径仍低于采样周期。有界感知与差速驱动的测试提供了额外的部署压力验证。
cs.RO / 43 / 2609.27466

A Modular Dual-Arm Robotic Cell for Disassembly and Repair of Industrial Control Electronics

用于工业控制电子设备拆卸与维修的模块化双臂机器人工作单元
Ruhe, Maximilian, Harlacher, Fabian, Friedrich, Christian, Kipfmueller, Martin
Abstract
Industrial control electronics such as programmable logic controllers, servo drives and operator panels are routinely repaired in plant maintenance, but were never designed for automated disassembly. This paper presents a modular dual-arm robotic cell for repair-oriented disassembly, using two collaborative manipulators, interchangeable tools, red-green-blue-depth (RGB-D) and wristlevel perception, force/torque sensing and a Robot Operating System (ROS) 2- based control with Behavior Tree (BT) execution, teleoperation, digital-twin support and bounded learning-based contact skills. The process is decomposed into sequence planning, symbolic execution with fallbacks, force-limited tool skills, visual condition assessment and demonstration-based adaptation. A CADderived device graph encodes the disassembly order, access constraints, tools, feasible removal directions and verification states and converts them into operation objects for the BT and motion layers. Grounded in three representative devices, the cell covers screw removal, damaged-fastener fallback, snap-fit opening, connector release, cooperative printed circuit board (PCB) extraction and condition-based repair decisions. The main contribution is an architecture linking sequence knowledge, perception, verification and force-aware skills through one ROS 2 interface across simulation, teleoperation and real hardware.
Chinese Translation
可编程逻辑控制器(PLC)、伺服驱动器和操作面板等工业控制电子设备在工厂维护中经常需要维修,但它们从未被设计为可自动化拆卸的对象。本文提出了一种面向维修的模块化双臂机器人拆卸工作单元,采用两台协作机械臂、可互换工具、红绿蓝-深度(RGB-D)视觉与腕部级感知、力/力矩传感,以及基于机器人操作系统(ROS)2的控制系统,支持行为树(Behavior Tree, BT)执行、遥操作、数字孪生以及基于有界学习的接触技能。该流程被分解为序列规划、带回退机制的符号执行、限力工具技能、视觉状态评估和基于示教的适应性调整。系统利用由CAD导出的设备图来编码拆卸顺序、可达性约束、工具选择、可行的拆卸方向及验证状态,并将其转换为供行为树层和运动层使用的操作对象。基于三种代表性设备,该工作单元实现了螺钉拆卸、受损紧固件回退处理、卡扣开启、连接器分离、协作式印刷电路板(PCB)提取以及基于状态的维修决策。本文的主要贡献是一种将序列知识、感知、验证和力感知技能通过统一的ROS 2接口,在仿真、遥操作和真实硬件之间相连接的体系架构。
cs.RO / 44 / 2609.27467

Kairos: Grounded Forecasting of Presence and Directional Flow in 4D Scene Graphs

Kairos:基于4D场景图的存在性与方向流的接地预测
Catalano, Iacopo, Placed, Julio A., Civera, Javier, Queralta, Jorge Peña
Abstract
Long-term autonomy in human-populated environments requires anticipating whether and how people will move at times a robot has not yet observed. Existing representations of pedestrian motion face a tradeoff: they either forecast future activity, reducing each location to a scalar rate, or model the full directional distribution, holding it fixed in time. We present Kairos, a predictive directional-flow memory that extends a hierarchical 3D scene graph (3DSG) to a 4D scene graph (4DSG). Every observed voxel of the reconstructed geometry stores a directional mixture and a presence rate, and spectral predictors forecast, for any future query time, both the probability that people are present and the full directional distribution of their motion. Pairwise flow dependence between adjacent voxels supports conditional queries, and per-voxel predictive variances yield calibrated credible intervals that tighten as observations accumulate. We evaluate Kairos on three real pedestrian environments: a robot-collected campus dataset, a shopping mall, and a station concourse recorded continuously for eleven months. Its learned state remains consistent under loop-closure corrections, and its forecasts are competitive with dedicated occupancy and flow models trained on the full detection stream, although Kairos learns from only the small fraction available to a patrolling robot. Finally, we validate the representation on a downstream encounter-probability planning task, where plans computed over the Kairos forecasts encounter more people than plans computed over any time-invariant map at an equal success rate. We provide the code at https://github.com/IacopomC/kairos.
Chinese Translation
在有人类活动的环境中实现长期自主运行,需要机器人预测在其尚未观测到的时刻人们是否以及将如何移动。现有的行人运动表征面临权衡:要么预测未来活动,将每个位置简化为一个标量速率;要么建模完整的方向分布,但将其固定于时间维度。我们提出了Kairos,一种预测性方向流记忆(predictive directional-flow memory),它将分层3D场景图(3DSG)扩展为4D场景图(4DSG)。重建几何中的每个观测体素都存储一个方向混合分布和一个存在率,谱预测器(spectral predictors)可为任意未来查询时刻预测人群存在的概率以及其运动的完整方向分布。相邻体素之间的成对流依赖支持条件查询,而每个体素的预测方差产生校准的置信区间,且该区间随着观测的积累而收紧。我们在三个真实的行人环境中评估Kairos:机器人采集的校园数据集、购物中心,以及连续记录十一个月的车站大厅。其学习到的状态在回环闭合校正下保持一致,其预测性能与在完整检测流上训练的专用占用和流模型相当,尽管Kairos仅从巡逻机器人所能获取的极小部分数据中学习。最后,我们在下游的遭遇概率规划任务上验证了该表征:在相同成功率下,基于Kairos预测计算的规划比基于任何时不变地图计算的规划遇到更多的人。代码已在 https://github.com/IacopomC/kairos 提供。
cs.RO / 45 / 2609.27469

Collocated Shape Regulation for Soft Robots

软机器人的同位形状调节
Pustina, Pietro, Shahabi, Ebrahim, Feliu-Talegon, Daniel, De Luca, Alessandro, Della Santina, Cosimo
Abstract
Controlling the shape of a continuum soft robot typically requires an accurate dynamic model and actuation of all degrees of freedom. We show that regulating only the actuated coordinates, through collocated shape control, achieves provably stable convergence of those coordinates and, under an explicit compatibility condition, of the entire robot shape. While collocated control is a cornerstone of high-performance motion control in rigid robotics, extending this formulation to continuum soft robots has remained challenging due to the complexity of their dynamics. We present the first general framework for collocated control of continuum soft robots and derive a unified family of controllers, including PD, PID, PsatID, and their counterparts with compensation and cancellation components. The framework unifies existing approaches while introducing new controller designs. In particular, we develop three classes of PD and PID like regulators with local, semi-global, and global stability guarantees, and provide rigorous convergence analyses for each. Extensive experimental validation demonstrates the effectiveness of the proposed methods across different model discretizations and controller parameters. The resulting framework provides practical design guidelines for selecting and implementing controllers with known stability guarantees, without requiring a complete dynamic model of the robot
Chinese Translation
控制连续体软机器人的形状通常需要精确的动力学模型以及对所有自由度的驱动。我们证明,通过同位形状控制,仅对被驱动的坐标进行调节,即可实现这些坐标的可证明的稳定收敛,并且在显式相容性条件成立时,可实现整个机器人形状的稳定收敛。虽然同位控制是刚性机器人高性能运动控制的基石,但由于连续体软机器人动力学的复杂性,将这一方法推广至连续体软机器人一直具有挑战性。我们提出了首个面向连续体软机器人同位控制的通用框架,并推导出一族统一的控制器,包括 PD、PID、PsatID 及其带补偿与抵消环节的对应形式。该框架统一了现有方法,同时引入了新的控制器设计。特别地,我们开发了三类具有局部、半全局和全局稳定性保证的类 PD 与类 PID 调节器,并对每一类给出了严格的收敛性分析。大量的实验验证表明,所提出的方法在不同的模型离散化方式和控制器参数下均具有有效性。所得到的框架为在已知稳定性保证下选择和实现控制器提供了实用的设计指南,而无需完整的机器人动力学模型。
cs.RO / 46 / 2609.27475

RoboCaf\'e in the Open: Interaction Continuity in Long-Term Public Human-Robot Interaction

开放环境中的RoboCafé:长期公共人机交互中的交互连续性
Pineda, Kaitlynn Taylor, Kushwaha, Kush Kumar, Wang, Jie, Du, Jiaming, Mishra, Anvii, Suri, Emilie Basu, Guo, Angela, Huang, Chien-Ming
Abstract
As robots remain in public spaces over extended periods, they must maintain interaction continuity by preserving and correctly applying context as people, encounters, and circumstances change. To study interaction continuity in long-term public human-robot interactions, we developed RoboCaf\'e, an autonomous conversational coffee robot designed to support repeated interactions through task-aware dialogue, real-time multimodal perception, and memory of prior encounters. We deployed RoboCaf\'e for 12 days in a university building, where it received 148 orders. The deployment involved repeat customers, passersby, changing groups, and back-to-back orders that repeatedly crossed the boundaries assumed by the system's order-centered interaction model. We found that successful interaction continuity requires a robot to determine who is currently present, which prior context belongs to whom, where interactions begin and end, and whether its representation of an interaction matches what is occurring in the physical world. From these observations, we derive four system design requirements for maintaining interaction continuity in longitudinal public human-robot interactions: contextual interaction state, persistent person grounding, explicit interaction life-cycle management, and interaction observability.
Chinese Translation
当机器人在公共空间中长期驻留时,它们必须通过在人员、接触事件和环境变化时保存并正确应用上下文来维持交互连续性。为了研究长期公共人机交互中的交互连续性,我们开发了RoboCafé——一款自主对话式咖啡机器人,旨在通过任务感知对话、实时多模态感知以及对先前接触事件的记忆来支持重复交互。我们将RoboCafé部署在一栋大学建筑内12天,共接收148笔订单。此次部署涉及回头客、路过者、不断变化的群体以及连续不断的订单,这些订单反复突破了系统以订单为中心的交互模型所预设的边界。我们发现,成功的交互连续性要求机器人能够确定当前在场的人员、判断先前的上下文属于谁、界定交互的起止时间,并确认其内部对交互的表征与物理世界中实际发生的情况相一致。基于这些观察,我们总结出在纵向公共人机交互中维持交互连续性的四项系统设计要求:上下文交互状态、持久的人员锚定、显式的交互生命周期管理以及交互可观测性。
cs.RO / 47 / 2609.27513

Behavior-Aligned Action Tokenization for Robot Policy Learning

面向机器人策略学习的行为对齐动作标记化方法
Dong, Junbo, Chen, Ze, Xie, Zhendong, Li, Junjie, Xu, Lixin, Chi, Xuemin, Song, Yiming, Ma, Zhaoyuan
Abstract
Autoregressive robot policies learn continuous control by predicting discrete action tokens from observations. Different tasks often share local motions, yet behavioral correspondence across demonstrations receives limited explicit supervision in existing tokenizers. Motions with different timing can therefore lack a shared representation despite following similar patterns. We propose Behavior-Aligned Action Tokenization (BAAT), which uses soft dynamic time warping (Soft-DTW) to select corresponding action chunks and aligns their quantized coordinates jointly with reconstruction. This objective encourages similar motions across tasks to occupy nearby quantized representations while retaining executable action detail. A history-conditioned diffusion decoder reconstructs continuous action chunks from these tokens, and a downstream autoregressive policy learns to predict them. We evaluate BAAT on selected tasks from three simulation benchmarks and two real robot tasks. BAAT achieves a mean simulation success rate of approximately 45.2%, exceeding OAT by approximately 7.2 percentage points. In the controlled LIBERO-All alignment ablation, policy success rises from 70.2% to 79.0% while trajectory replay success decreases. These results support behavioral correspondence as supervision for organizing shared motion structure in action tokenizers and improving downstream robot policy learning.
Chinese Translation
自回归机器人策略通过从观测中预测离散动作标记来学习连续控制。不同任务往往共享局部运动模式,然而在现有的标记化方法中,演示之间的行为对应关系缺乏充分的显式监督。因此,尽管遵循相似的模式,时序不同的运动可能缺乏共享的表示。我们提出行为对齐动作标记化方法(Behavior-Aligned Action Tokenization, BAAT),利用软动态时间规整(Soft-DTW)选择对应的动作片段,并在重建的同时联合对齐其量化坐标。该目标函数鼓励跨任务的相似运动占据相近的量化表示,同时保留可执行的动作细节。我们使用基于历史条件的扩散解码器从这些标记重建连续动作片段,并由下游自回归策略学习预测这些片段。我们在三个仿真基准中的选定任务以及两个真实机器人任务上评估了BAAT。BAAT的平均仿真成功率达到约45.2%,比OAT高出约7.2个百分点。在受控的LIBERO-All对齐消融实验中,策略成功率从70.2%提升至79.0%,同时轨迹回放成功率有所下降。这些结果支持将行为对应关系作为监督信号,用于在动作标记化方法中组织共享的运动结构,并改进下游机器人策略学习。
cs.RO / 48 / 2609.27526

NavProbe: Evidence-Grounded Reasoning with Active Memory Retrieval for Zero-Shot Navigation

NavProbe:面向零样本导航的基于证据推理与主动记忆检索方法
Liu, Jingyang, Yao, Sujia, Gu, Jiayuan, Xu, Lan
Abstract
Long-horizon navigation requires an agent to revise its intermediate objectives as evidence accumulates. Full visual histories are costly to process, while compact summaries may omit details needed to reconsider earlier decisions. We introduce NavProbe, a hierarchical zero-shot navigation agent that couples a dynamic subgoal agenda with active evidence retrieval. A compact index links summaries of visited places, transitions, and landmarks to their visual and geometric records. When the current context is insufficient, a task executive retrieves targeted evidence to generate, revise, or resolve subgoals. Reusable conclusions are used to update the index, and a skill policy converts the revised task state into parameterized navigation actions. NavProbe achieves 71.7% SR and 55.8% SPL on R2R-CE and 55.3% SR and 38.6% SPL on RxR-CE, outperforming strong zero-shot baselines. It also achieves 79.3% SR on HM3D-v2 ObjectNav, with qualitative real-robot demonstrations illustrating physical deployment.
Chinese Translation
长时程导航要求智能体随着证据的积累不断修正其中间目标。完整的视觉历史处理代价高昂,而压缩摘要又可能遗漏重新审视早期决策所需的细节。我们提出NavProbe,一种将动态子目标日程与主动证据检索相结合的分层零样本导航智能体。一个紧凑的索引将已访问地点、转移和地标的摘要与其视觉和几何记录相关联。当当前上下文信息不足时,任务执行器(task executive)检索有针对性的证据,以生成、修正或解决子目标。可复用的结论被用于更新索引,技能策略则将修正后的任务状态转换为参数化的导航动作。NavProbe在R2R-CE上达到71.7%的SR和55.8%的SPL,在RxR-CE上达到55.3%的SR和38.6%的SPL,优于强大的零样本基线方法。此外,该方法在HM3D-v2 ObjectNav上达到79.3%的SR,并通过定性实机机器人演示展示了其物理部署能力。
cs.RO / 49 / 2609.27536

Behaviora - A Conceptual Architecture for External and Internal Behavior of Robots and Agents

Behaviora——一种用于机器人和智能体外部与内部行为的概念架构
Nyman, Gote
Abstract
Behaviora is a preliminary conceptual architecture for representing agent and robot behavior, external and internal alike, in an addressable form. A behaving robot or agent performs a Behavior Episode composed of episode components, which can be derived from behavior taxonomies (BTax) and assigned persistent identifiers. We denote these identifiers as IoB (Internet of Behaviors) Addresses. A Behavior Episode specifies what the system does, while a Style Profile (SP) specifies how this behavior is expressed. Style can communicate characteristics of the actor and qualities such as competence and cultural manners. An Experience Profile (EP) represents behaviorally relevant internal state that modulates the execution of an Episode. Finally, a Behavior Compiler maps these behavioral representations to platform-specific actions. We use a primitive touching arm model to show these components and their relations. External Behavior is a result of addressable movements and their styles. Internal Behavior is represented through the same episodic principle and can be rendered as inner speech. Sensing, perception and complex task contexts have not been included in the present implementation, although a conceptual place is reserved for them.
Chinese Translation
Behaviora 是一种初步的概念架构,用于以可寻址的形式表示智能体和机器人的外部行为与内部行为。一个具有行为的机器人或智能体会执行一个由情节组件构成的行为情节,这些组件可以从行为分类体系派生,并被分配持久标识符。我们将这些标识符称为 IoB(行为互联网)地址。行为情节规定系统做什么,而风格档案规定这种行为如何表达。风格可以传达行为主体的特征以及胜任力和文化礼仪等特质。经验档案表示与行为相关的内部状态,用以调节情节的执行。最后,行为编译器将这些行为表示映射到特定平台的动作。我们使用一个原始的触摸手臂模型来展示这些组件及其相互关系。外部行为是可寻址运动及其风格的结果;内部行为通过相同的情节化原则来表示,并可以呈现为内心言语。感知、知觉和复杂任务情境尚未包含在当前实现中,但已为它们保留了概念上的位置。
cs.RO / 50 / 2609.27597

Gray-Box Model Predictive Control for Articulated Dump Trucks via Gaussian Process Learning of Sideslip

基于高斯过程侧偏学习的铰接式自卸车灰箱模型预测控制
Shahirpour, Arash, Ahlers, Jens, Schulte, Christopher, Reuscher, Tim
Abstract
The growing demand for automation in the mining industry, particularly for the autonomous operation of articulated dump trucks (ADTs), has drawn increased attention to accurate vehicle modeling. The importance of such models lies in their use in model predictive control (MPC), model-based estimation methods, and vehicle simulation. While dynamic modeling offers a viable solution for these purposes, it is associated with complex setup and parametrization and may require recalibration in changing operating environments. As a result, kinematic models have dominated ADT modeling, especially in MPCs, at the expense of reduced prediction accuracy. In this work, we propose an approach using Gaussian Process Regression (GPR) to learn the sideslip angle of the vehicle, which is identified as the primary contributor to the reduced accuracy of kinematic models. The learned GPR function is augmented into the kinematic model to form a gray-box model that aims to reduce the gap to dynamic models. We show that the gray-box model can predict the sideslip angle and, consequently, the vehicle's lateral velocity, thereby improving the MPC's prediction performance. The resulting gray-box MPC is compared against two white-box MPCs in a simulation environment. The results indicate an improvement in terms of maximum lateral tracking error from over 2 m to 0.56 m.
Chinese Translation
采矿业对自动化需求的日益增长,尤其是铰接式自卸车(ADT)的自主作业需求,使得精确的车辆建模受到了越来越多的关注。此类模型的重要性体现在其在模型预测控制(MPC)、基于模型的估计方法以及车辆仿真中的应用。尽管动力学建模为实现这些目的提供了可行的方案,但其设置和参数化过程复杂,并且在作业环境变化时可能需要重新标定。因此,运动学模型在ADT建模中占据主导地位,尤其是在MPC中,但代价是预测精度的降低。在本工作中,我们提出一种利用高斯过程回归(GPR)学习车辆侧偏角的方法,该侧偏角被识别为导致运动学模型精度降低的主要因素。将学习到的GPR函数增广到运动学模型中,构成一个灰箱模型,旨在缩小其与动力学模型之间的差距。我们表明,该灰箱模型能够预测侧偏角,进而预测车辆的侧向速度,从而提升MPC的预测性能。最后,在仿真环境中将所得的灰箱MPC与两种白箱MPC进行了对比。结果表明,最大侧向跟踪误差从2米以上降低至0.56米。
cs.RO / 51 / 2609.27612

RegenHarness: A Robot Agent Harness with Evidence-Gated Recursive Self-Improvement

RegenHarness:一种具有证据门控递归自我改进机制的机器人智能体框架
Wang, Kailin, Jie, Haoxiang, Yan, Yaoyuan, Heng, Zhiyou, Li, Zhaosong
Abstract
Long-horizon robot execution requires a clear distinction between a model's proposal, a controller's termination, and verified task completion. We present RegenHarness, an evidence-gated robot-agent harness connecting task planning to heterogeneous robot skills. Its execution architecture couples a model loop for context-conditioned proposals with an agent loop for dispatch, observation, verification, commitment, and bounded recovery. Four role-isolated contexts separate planning, supervision, verification, and recovery inputs. Versioned memory distinguishes observed facts from accepted task progress, while an identity- and version-bound commit gate controls updates to trusted task state. The runtime combines duplicate-dispatch control, resource leases, and recovery budgets under explicit backend contracts, and checks the original user goal before reporting completion. To our knowledge, we are the first to introduce an evidence-gated recursive self-improvement (RSI) protocol for embodied robotic agents. Across missions, execution records motivate candidate changes to context rules, task templates, routing, and recovery policies; fixed regression checks and release authorization govern their acceptance; versioned rollout and rollback preserve configuration traceability. This RSI protocol revises the harness configuration without online model-weight updates or permission to weaken the commit gate. A real quadruped deployment documents voice-triggered warehouse navigation, panoramic inspection, visual analysis, message delivery, return, and spoken reporting through linked audio, images, trajectories, and receipts. A separate circuit demonstrates why completion depends on execution history rather than endpoint proximity alone. Together, the cases demonstrate integrated perception, physical execution, communication, and history-dependent completion in real-world robot tasks.
Chinese Translation
长时程机器人任务执行需要明确区分模型的提议、控制器的终止判定以及经过验证的任务完成。我们提出了RegenHarness,这是一个证据门控的机器人智能体框架,将任务规划与异构机器人技能相连接。其执行架构将一个用于生成上下文条件化提议的模型循环,与一个负责调度、观测、验证、提交和有界恢复的智能体循环相耦合。四个角色隔离的上下文分别用于规划、监督、验证和恢复输入。版本化内存将观测到的事实与已接受的任务进展区分开来,同时一个绑定身份与版本的提交门控控制对可信任务状态的更新。运行时在明确的后端契约下结合了重复调度控制、资源租约和恢复预算,并在报告任务完成之前核对用户的原始目标。据我们所知,我们是首个为具身机器人智能体引入证据门控递归自我改进(Recursive Self-Improvement, RSI)协议的工作。在多项任务中,执行记录会催生对上下文规则、任务模板、路由和恢复策略的候选修改;固定的回归检查与发布授权机制控制这些修改的接受;版本化的部署与回滚保证了配置的可追溯性。该RSI协议在不进行在线模型权重更新、且不允许削弱提交门控的前提下修订框架配置。一次真实的四足机器人部署记录了语音触发的仓库导航、全景巡检、视觉分析、消息传递、返回以及语音报告,并通过关联的音频、图像、轨迹和回执予以佐证。另一项独立实验则说明了任务完成与否取决于执行历史,而非仅仅取决于是否接近终点。这些案例共同展示了真实世界机器人任务中集成的感知、物理执行、通信以及依赖历史的完成判定。
cs.RO / 52 / 2609.27656

InternW0: A Foundational Physical World Model for Efficient Real-World Interactions

InternW0:面向高效真实世界交互的基础物理世界模型
Cai, Jisong, Mu, Yao, Yang, Ganlin, Cao, Zhe, Tu, Zhangzheng, Gao, Xing, Li, Kailin, Zhan, Xinyu, Yang, Lixin, Zhu, Yangkun, Ma, Haoxiang, Zhou, Ming, Yu, Qiaojun, Xue, Yufei, He, Liqun, Yao, Yifei, Zhu, Yifan, Ling, Long, Jiang, Bingqi, Guo, Haoyu, Zhu, Xueyue, Zhou, Bowen, Zhao, Bin, Xue, Tianfan, Shen, Chunhua, Zhang, Weinan
Abstract
Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi-frequency processing, and local physical modeling under partial observations and external influences. InternW0 jointly learns future visual dynamics and continuous robot control through an asymmetric video--action architecture with flow matching. A high-capacity video expert provides longer-horizon predictive context, while a lightweight action expert operates at a faster timescale. Instead of regenerating the future for every action update, InternW0 reuses layerwise K/V and adapts it to newly observed states through observation-conditioned context routing. Domain-specific interfaces and soft prompts support heterogeneous embodiments, while contact-aware post-training incorporates force and tactile signals for contact-rich manipulation. We train InternW0 on approximately 7,200 hours of heterogeneous robot and egocentric data, including EgoLab, a 275-hour real-laboratory egocentric dataset. Evaluation spans simulation benchmarks and real-world scientific tasks, including a 15-stage metal--organic framework synthesis workflow and 5-stage contact- and force-aware dexterous manipulation for general-purpose quantitative pipetting. These results advance scalable, asynchronous, and science-native physical world models for universal and efficient real-world interactions.
Chinese Translation
物理智能不仅仅需要预测世界如何演化:随着世界的持续变化,预测结果还必须保持可执行性。我们提出 InternW0,这是上海人工智能实验室 InternW 物理世界模型系列的第一个实例,其构建围绕全模态接口、异步多频率处理,以及在部分观测和外部影响下的局部物理建模。InternW0 通过一种非对称的视频—动作架构与流匹配(flow matching),联合学习未来视觉动态与连续机器人控制。高容量的视频专家提供更长时域的预测上下文,而轻量级的动作专家则以更快的频率运行。InternW0 无需在每次动作更新时重新生成未来,而是复用逐层 K/V 缓存,并通过以观测为条件的上下文路由将其适配到新观测到的状态。面向特定领域的接口与软提示(soft prompts)支持异构本体,而接触感知的后训练则引入力与触觉信号,以支持富含接触的操作任务。我们在约 7,200 小时的异构机器人及第一人称视角数据上训练 InternW0,其中包括 275 小时的真实实验室第一人称数据集 EgoLab。评估涵盖仿真基准与真实世界科学任务,包括一个 15 阶段的金属有机框架(MOF)合成流程,以及面向通用定量移液的 5 阶段接触与力感知灵巧操作。这些结果推动了可扩展、异步以及面向科学的物理世界模型的发展,以实现通用且高效的真实世界交互。
cs.RO / 53 / 2609.27695

GLoTouch: Global-to-Local Haptic Perception Using a Parallel Gripper for Object Search, Recognition, and Grasping Without External Vision

GLoTouch:基于平行夹爪的全局到局部触觉感知,实现无需外部视觉的物体搜索、识别与抓取
Li, Zonglin, Zhang, Wanruo, Wang, Yiming, Song, Kun, Zhou, Xinyi, Ma, Daolin
Abstract
Perceiving objects in the environment is a fundamental capability of autonomous robots. In dark or low-light environments, external cameras often fail to reliably perceive object positions and geometry; when visual sensing is unavailable, completing target search, recognition, and grasping through touch alone becomes a key robot manipulation capability. This task must simultaneously address container-scale spatial exploration and object-scale fine-grained geometric perception, which is particularly challenging for low-degree-of-freedom parallel grippers. However, a unified framework remains lacking for connecting container-scale spatial exploration with object-scale fine-grained geometric perception and grasping. To address this challenge, we present \textbf{GLoTouch}, a global-to-local haptic perception and manipulation framework built on a parallel gripper. In the global stage, the gripper holds a passive long-reach probe, combining force measurements with known tool geometry to localize contacts and actively estimate candidate-object positions, coarse contours, and heights. In the local stage, the robot sets down the probe and uses the bilateral visuotactile sensors on the same gripper to directly acquire local haptic observations, which are matched against a given target 3-D model without object-specific training. We evaluate the framework in both simulation and real-robot experiments. Source code will be open-sourced.
Chinese Translation
感知环境中的物体是自主机器人的一项基本能力。在黑暗或低光照环境中,外部相机往往无法可靠地感知物体的位置与几何形状;当视觉感知不可用时,仅凭触觉完成目标搜索、识别与抓取便成为机器人操作的一项关键能力。该任务需要同时应对容器尺度的空间探索和物体尺度的细粒度几何感知,这对于低自由度的平行夹爪而言尤为具有挑战性。然而,目前仍缺乏一个将容器尺度的空间探索与物体尺度的细粒度几何感知及抓取相连接的统一框架。为应对这一挑战,我们提出了GLoTouch,一个构建于平行夹爪之上的全局到局部触觉感知与操作框架。在全局阶段,夹爪持有一个无源长杆探针,将力测量与已知的工具几何形状相结合以定位接触点,并主动估计候选物体的位置、粗略轮廓和高度。在局部阶段,机器人放下探针,利用同一夹爪上的双侧视触觉传感器直接获取局部触觉观测,并将其与给定的目标三维模型进行匹配,而无需针对特定物体的训练。我们在仿真和真实机器人实验中对该框架进行了评估。源代码将开源。
cs.RO / 54 / 2609.27702

DAVIO: Dense Monocular-Inertial SLAM with Feed-Forward Initialization and Pose-Conditioned Mapping

DAVIO:具有前馈初始化与位姿条件化建图的稠密单目-惯性SLAM
Mahmoud, Jaafar, Movsesyan, Arthur, Iumanov, Mikhail, Kolyubin, Sergey
Abstract
A camera and an IMU are the minimal sensor setup for metric localization and dense mapping, yet classical visual--inertial filters must wait for parallax before they start and then retain only sparse landmarks. Feed-forward geometry models, in contrast, predict dense structure from a few images but provide neither metric scale nor gravity. We present DAVIO, which uses a single multi-view depth model, Depth Anything~3, for both start-up and mapping. At start-up, a five-image window and preintegrated IMU measurements form a feature-free linear system. Its robust, conditioning-checked solution bootstraps a VIO filter through buffered replay. During tracking, the filter's metric poses condition the depth model. Residual scale is corrected only along viewing rays, which preserves the metric camera baselines, and a gravity-preserving submap graph with drift-gated revisits refines the map. On EuRoC, DAVIO starts markedly earlier, reduces the localization error, and maps more accurately than SOTA feed-forward mappers given identical poses. On building-scale ORI sequences, DAVIO is on bar or better than SOTA mappers on the same odometry, and degrades far less when GT poses are replaced by real odometry. We release the code of DAVIO, a real-time dense metric SLAM system, to the community.
Chinese Translation
相机与IMU是实现米制度量定位与稠密建图的最小传感器配置,然而经典的视觉-惯性滤波器必须等待视差产生才能启动,且仅保留稀疏的地图点。相比之下,前馈几何模型能够从少量图像中预测稠密结构,但既无法提供米制尺度,也无法提供重力方向。我们提出DAVIO,其使用单一多视角深度模型Depth Anything 3同时完成启动和建图。在启动阶段,五帧图像窗口与预积分的IMU测量构成一个无需特征的线性系统,其经过鲁棒性和条件数检验的求解结果通过缓存回放(buffered replay)来引导VIO滤波器初始化。在跟踪过程中,滤波器的米制位姿作为深度模型的条件,尺度残差仅沿视线方向进行校正,从而保持米制的相机基线;同时,采用保持重力方向的子图结构,并通过带漂移门限的重访问机制对地图进行精化。在EuRoC数据集上,与给定相同位姿的最先进(SOTA)前馈建图方法相比,DAVIO能够显著更早地启动,降低定位误差,并实现更精确的建图。在建筑物尺度的ORI序列上,DAVIO在相同里程计条件下达到或超越SOTA建图方法,且当真值位姿被替换为真实里程计时性能下降幅度远小于对比方法。我们向社区开源发布DAVIO——一个实时的稠密米制SLAM系统。
cs.RO / 55 / 2609.27712

Wave-Robust Passive AUV Localization Using FP-MUSIC

基于FP-MUSIC的抗波浪被动AUV定位方法
Saqib, Usama, Rønning, Ola, Wąsowski, Andrzej
Abstract
Localizing an autonomous underwater vehicle without pre-deployed seabed transponders, or direct access to onboard vehicle sensors remains a core challenge. We present a receiver-passive 3-D localization and spatial mapping system utilizing a single floating surface buoy equipped with a hydrophone array and an inertial measurement unit (IMU). The central difficulty is that surface wave motion induces six-degree-of-freedom (6-DOF) perturbations that rotate the array between snapshots, degrading conventional subspace processing. We resolve this by introducing a fixed-point iterative MUltiple SIgnal Classification algorithm (FP-MUSIC) that uses IMU measurements to de-warp snapshot covariances prior to direction-of-arrival estimation. Furthermore, we employ a subspace-projected wideband matched filter to resolve beacon ranges and use power asymmetry for independent front-back identification. Evaluations across simulated sea states demonstrate that FP-MUSIC substantially reduces localization error relative to uncompensated methods and sustains robust 3-D tracking and vehicle orientation estimation under wave-induced motion. At moderate sea state, FP-MUSIC increases the 2-m beacon-separation accuracy from approximately 45% to 75%.
Chinese Translation
在不预先部署海底应答器、且无法直接获取水下航行器自身传感器数据的情况下,对自主水下航行器(AUV)进行定位仍然是一个核心挑战。我们提出了一种接收端被动的三维定位与空间测绘系统,该系统利用单个漂浮于水面的浮标,其上配备水听器阵列和惯性测量单元(IMU)。核心难点在于,海面波浪运动会引起六自由度(6-DOF)扰动,使阵列在快拍之间发生旋转,从而降低传统子空间处理方法的性能。为解决这一问题,我们提出了一种不动点迭代的多重信号分类算法(FP-MUSIC),该算法利用IMU测量数据在波达方向(DOA)估计之前对快拍协方差矩阵进行去扭曲校正。此外,我们采用子空间投影宽带匹配滤波器来解析信标距离,并利用功率不对称性实现独立的前后向辨识。在不同模拟海况下的评估表明,相对于未补偿的方法,FP-MUSIC显著降低了定位误差,并在波浪引起的运动下依然能够实现稳健的三维跟踪与航行器姿态估计。在中等海况下,FP-MUSIC将2米信标间距的分辨准确率从约45%提升至75%。
cs.RO / 56 / 2609.27720

CoRelNav: Collaborative Relational Navigation for Multi-Robot Spatially Constrained Semantic Navigation

CoRelNav:面向多机器人空间约束语义导航的协作式关系导航
He, Jinyu, Mao, Zihao, Jin, Haonan, Fu, Mengyin, Song, Wenjie
Abstract
Spatially constrained semantic navigation requires robots to identify targets specified not only by semantic categories but also by relations to surrounding objects. In unknown environments, resolving such goals requires efficient exploration together with sufficient target and contextual evidence for reliable relation verification. Existing methods leave relation-aware verification and multi-robot collaboration largely disconnected: relational navigation is predominantly single-agent, while multi-robot systems seldom coordinate distributed observations for instance-specific relation verification. We propose CoRelNav, whose core is coupling task-conditioned multi-robot exploration with candidate-driven collaborative verification. A spatial-semantic field converts task constraints, scene nodes, and object features into exploration utility; as candidate information accumulates, robots are reallocated toward complementary evidence under team navigation costs, while instance-consistent observations are aggregated across topology nodes. This coupling reduces redundant search and enables relation hypotheses to be resolved from distributed partial evidence that independent exploration or isolated-view verification can leave ambiguous. Experiments in photorealistic simulation demonstrate consistent improvements over representative baselines, with ablations validating the proposed exploration and verification mechanisms. We further deploy the complete system on two physical mobile robots, demonstrating its applicability to real-world collaborative navigation.
Chinese Translation
空间约束语义导航要求机器人识别的目标不仅由语义类别指定,还由其与周围物体之间的关系指定。在未知环境中,解决此类目标需要高效探索以及充分的目标与上下文证据,以实现可靠的关系验证。现有方法中,关系感知验证与多机器人协作基本处于割裂状态:关系导航主要以单智能体为主,而多机器人系统很少协调分布式观测以进行特定实例的关系验证。我们提出CoRelNav,其核心是将任务条件化的多机器人探索与候选驱动的协作验证相耦合。空间-语义场将任务约束、场景节点和物体特征转化为探索效用;随着候选信息的积累,机器人在团队导航成本约束下被重新分配以收集互补证据,同时在拓扑节点间聚合实例一致的观测。这种耦合减少了冗余搜索,并使关系假设能够从分布式部分证据中得到判定,而独立探索或孤立视角验证可能会使这些假设存在歧义。在照片级逼真仿真中的实验表明,本方法相比代表性基线取得了一致的改进,消融实验验证了所提出的探索与验证机制。我们进一步将完整系统部署到两台实体移动机器人上,验证了其在真实世界协作导航中的适用性。
cs.RO / 57 / 2609.27734

InfiNoVA: Infinite Novel View Augmentation for Viewpoint Invariant Robot Policies

InfiNoVA:面向视角不变机器人策略的无限新颖视角增广
Gottam, Sai Puneeth Reddy, Rueckert, Elmar, Dave, Vedant
Abstract
Vision-Language-Action (VLA) policies often rely strongly on the camera viewpoints seen during training, causing substantial performance degradation when deployed from unseen perspectives. Collecting demonstrations from sufficiently diverse physical viewpoints is expensive and still provides only sparse coverage of the viewpoint space. We introduce InfiNoVA, a data-augmentation framework that converts synchronized multi-camera demonstrations into a dense distribution of geometrically consistent training views. InfiNoVA reconstructs each manipulation trajectory as a time-varying 3D Gaussian representation and renders novel observations from sampled camera poses while preserving the original state-action correspondence. This explicit scene representation improves frame-level fidelity and temporal consistency while reducing task-critical hallucinations observed in generative novel-view synthesis. Across four real-world manipulation tasks, policies trained with InfiNoVA achieve 5.4x higher average success under unseen randomized viewpoints than both VISTA-based augmentation and the unaugmented policy. InfiNoVA further achieves 1.7x higher success than training directly on all five physical camera views. These results show that dense, geometrically grounded viewpoint augmentation provides a practical route toward camera-robust robot policies without modifying the underlying policy architecture.
Chinese Translation
视觉-语言-动作(VLA)策略往往高度依赖训练中所见的相机视角,导致在未见过的视角下部署时性能大幅下降。从足够多样的物理视角采集示范数据成本高昂,且仍只能对视角空间提供稀疏覆盖。我们提出InfiNoVA,一种数据增广框架,可将同步的多相机示范数据转换为几何一致的密集训练视角分布。InfiNoVA将每条操作轨迹重建为时变的3D高斯表示,并从采样的相机位姿渲染新颖观测,同时保持原有的状态-动作对应关系。这种显式场景表示提升了帧级保真度和时间一致性,并减少了生成式新颖视角合成中出现的任务关键幻觉。在四项真实世界操作任务中,使用InfiNoVA训练的策略在未见过的随机视角下的平均成功率达到基于VISTA的增广方法和未增广策略的5.4倍。InfiNoVA还比直接在全部五个物理相机视角上训练的方案成功率高出1.7倍。这些结果表明,密集的、以几何为基础的视角增广为在不修改底层策略架构的前提下获得相机鲁棒的机器人策略提供了一条可行途径。
cs.RO / 58 / 2609.27747

Less Language, More Latents: Annotation-Efficient VLAs for Driving

少语言,多潜变量:面向驾驶的标注高效视觉-语言-动作模型
Zakharov, Alexey, Oksuz, Kemal, Dokania, Puneet K.
Abstract
Vision-language-action models (VLA) promise human-steerable autonomous driving, but their training is bottlenecked by the scarcity of frames paired with natural-language instructions: while camera streams and expert trajectories are logged at scale, language annotations (e.g., turn left at the intersection) remain scarce and expensive to acquire. To address this challenge, we introduce Latent Action Driving Annotations (LADA), a three-stage pipeline that transforms abundant unlabelled observation-trajectory pairs into a substrate for language-conditioned control. First, we train a latent action model with a vector-quantised bottleneck, producing a compact codebook of high-level vehicle intents. Second, a small language-annotated subset is used to train a vision-language translator to map observations and language instructions into this codebook. Third, we train a driving VLA on observation-latent-action pairs over the full unlabelled corpus. Using fewer than 5% of language annotations and without leveraging any auxiliary chain-of-thought reasoning or visual question answering streams, LADA achieves a Driving Score of 87.98 and a Success Rate of 70.46% on the closed-loop Bench2Drive benchmark, matching or surpassing fully supervised baselines.
Chinese Translation
视觉-语言-动作模型(VLA)有望实现人类可引导的自动驾驶,但其训练受制于与自然语言指令配对的帧数据稀缺这一瓶颈:尽管相机视频流和专家轨迹可以大规模记录,但语言标注(例如,在路口左转)仍然稀缺且获取成本高昂。为应对这一挑战,我们提出了潜动作驾驶标注(Latent Action Driving Annotations,LADA),这是一个三阶段流水线,可将大量无标注的观测-轨迹对转化为语言条件控制的基础数据。首先,我们训练一个带向量量化瓶颈的潜动作模型,生成一个紧凑的高层车辆意图码本。其次,利用少量带语言标注的子集训练一个视觉-语言翻译器,将观测和语言指令映射到该码本中。第三,我们在完整的无标注语料库上使用观测-潜动作对训练驾驶VLA模型。在仅使用不到5%的语言标注、且不借助任何辅助思维链推理或视觉问答数据流的情况下,LADA在闭环Bench2Drive基准上取得了87.98的驾驶分数(Driving Score)和70.46%的成功率,达到甚至超越了全监督基线方法。
cs.RO / 59 / 2609.27780

Task-Prototype Guided Flow Matching for Few-Shot Generalization in Vision-Language Robot Manipulation

面向视觉-语言机器人操作少样本泛化的任务-原型引导流匹配
Wang, Yizhao, Zhang, Guantao, Wang, Jingbo
Abstract
Vision-language robot manipulation policies can follow semantic instructions, but adapting them to a new procedure from only a few demonstrations remains difficult because language underspecifies contact timing, motion phases, corrective behavior, and execution style. This paper presents Task-Prototype Guided Flow Matching (TP-Flow), a few-shot manipulation framework that converts support demonstrations into structured task-prototype tokens and uses them to guide both the initial flow prior and the velocity field. TP-Flow employs symmetric cross-attention with learnable queries to extract phase-level prototypes, parameterizes a task-adaptive initial distribution, and injects prototype information through gated adaptive normalization. It is trained with an episodic support-query objective and prototype contrastive regularization, so few-shot adaptation is simulated during training while nuisance information is suppressed. On the LEROBOT-ARM-SO101 platform, TP-Flow achieves 66.8\%, 79.6\%, and 82.1\% success rates under 1-, 4-, and 6-shot settings, with a 75.5\% few-shot AUC. At 1-shot, it improves over CFM, Pooled-Demo CFM, and In-Context Flow by 29.8, 14.5, and 9.9 percentage points. It also improves held-out target-group generalization across novel-object transfer, goal recombination, long-horizon composition, and contact/correction tasks. TP-Flow maintains real-time execution with six online prototype tokens, 54.3 ms latency, 3.9 GB peak memory, and a 10 Hz control rate, while reducing the noisy-support success drop to 6.2\%. Theoretical diagnostics show that prototype distance aligns with action-distribution distance, the adaptive prior reduces transport cost, and gated modulation keeps measured trajectory deviations below the derived ODE bound. The code repository is omitted for anonymous review.
Chinese Translation
视觉-语言机器人操作策略能够遵循语义指令,但由于语言对接触时机、运动阶段、纠错行为和执行风格的描述不够明确,仅凭少量示范将其适应到新任务流程仍然困难。本文提出任务-原型引导流匹配(Task-Prototype Guided Flow Matching,TP-Flow),这是一个少样本操作框架,将支持集示范转化为结构化的任务-原型令牌(tokens),并用其同时引导初始流先验和速度场。TP-Flow采用带有可学习查询的对称交叉注意力提取阶段级原型,参数化任务自适应的初始分布,并通过门控自适应归一化注入原型信息。该模型以情景式支持-查询目标和原型对比正则化进行训练,从而在训练过程中模拟少样本适应并抑制无关干扰信息。在LEROBOT-ARM-SO101平台上,TP-Flow在1-shot、4-shot和6-shot设置下分别取得66.8%、79.6%和82.1%的成功率,少样本AUC达到75.5%。在1-shot设置下,其性能相比CFM、Pooled-Demo CFM和In-Context Flow分别提升29.8、14.5和9.9个百分点。此外,TP-Flow在未见目标组泛化方面也表现出改进,涵盖新物体搬运、目标重组、长时程组合以及接触/纠错任务。TP-Flow通过六个在线原型令牌实现实时执行,延迟为54.3毫秒,峰值内存为3.9 GB,控制频率为10 Hz,同时将含噪支持集带来的成功率下降降低至6.2%。理论诊断表明,原型距离与动作分布距离相一致,自适应先验降低了传输成本,且门控调制使测得的轨迹偏差保持在推导的ODE界之内。出于匿名评审考虑,代码仓库暂不公开。
cs.RO / 60 / 2609.27816

Safe Multi-Robot Coordination via VLM-LLM Reasoning and Reachability Analysis

基于视觉语言模型-大语言模型推理与可达性分析的安全多机器人协调
Dwedar, Mohamed, Hafez, Ahmad, Jesser, Alexander, Alanwar, Amr
Abstract
Safe coordination in heterogeneous machine-to-machine (M2M) robotic systems is challenging when robots differ in sensing capabilities, environmental awareness, and motion execution roles. This paper presents a centralized safety-aware M2M framework for cooperative goal-directed navigation in a heterogeneous mobile robot team comprising a vision-capable quadruped and a camera-less robotic vehicle. The objective is to guide both platforms toward a goal region while avoiding static and dynamic obstacles and preventing unsafe inter-robot interactions. Under the principle of shared perception, the vision-capable robot provides semantic environmental awareness through a centralized server over an MQTT broker, enabling the camera-less platform to navigate using this shared scene representation alongside its own odometry, IMU, and state feedback. A vision-language model (VLM) interprets the visual stream, and the extracted semantic data is mapped into conservative metric geometric constraints, including inflated obstacle sets, safe corridors, and goal regions. A large language model (LLM) proposes high-level task allocations, while physical command authority is restricted to a robot-specific zonotope reachability gate. This verification engine propagates independent reachable tubes to evaluate obstacle avoidance, safe-corridor containment, and inter-robot separation predicates before approving commands. Online experiments across clear-path and dynamic-obstacle scenarios show that the pipeline reliably approves safe motion, triggers conservative replanning or holding maneuvers upon constraint violation, and enforces a strict architectural separation between advisory semantic reasoning and formally verified motor execution.
Chinese Translation
在异构机器对机器(M2M)机器人系统中,当机器人在感知能力、环境感知和运动执行角色上存在差异时,实现安全协调极具挑战性。本文提出了一种以安全为中心的集中式M2M框架,用于由一台具备视觉能力的四足机器人与一台无摄像头机器人车辆组成的异构移动机器人团队进行协作的目标导向导航。其目标是在避免静态与动态障碍物、防止机器人间不安全交互的前提下,引导两个平台驶向目标区域。在共享感知原则下,具备视觉能力的机器人通过MQTT代理经由集中式服务器提供语义环境感知,使无摄像头平台能够利用该共享场景表征以及自身的里程计、IMU和状态反馈进行导航。视觉语言模型(VLM)对视觉流进行解读,提取的语义数据被映射为保守的度量几何约束,包括膨胀障碍物集合、安全走廊和目标区域。大语言模型(LLM)提出高层任务分配方案,而物理命令权限则受限于针对特定机器人的zonotope可达性门控。该验证引擎传播独立的可达管道,在批准命令之前评估避障、安全走廊约束及机器人间分离谓词。在无障碍路径和动态障碍场景下的在线实验表明,该流程能够可靠地批准安全运动,在违反约束时触发保守的重规划或保持机动,并在建议性语义推理与经形式化验证的运动执行之间实现了严格的架构分离。
cs.RO / 61 / 2609.27905

Anchor-Free Hidden-Target Seeking via Certified Self-Calibration under Correlated Odometry

相关里程计条件下基于可认证自校准的无锚点隐藏目标搜寻
Bagla, Yash
Abstract
We study hidden-target seeking in an anchor-free regime where neither the vehicle, the relay, nor the target has an accessible global pose. The vehicle never senses the target directly; it receives only range-bearing observations through a single relay of unknown position and orientation, with motion known only via integrated body-frame odometry. Absolute localization is fundamentally impossible: the joint configuration retains an exact three-dimensional SE(2) gauge no estimator can resolve. Yet the quantities needed for control remain fully recoverable: motion collapses calibration to a task-relevant quotient, the relay-to-odometry yaw and target displacement, identifiable in closed form from two distinct vehicle views. We introduce an O(K) multi-view self-calibration estimator with an exact first-order yaw-uncertainty certificate that propagates the cross-view correlations integrated odometry induces: the full certificate attains 95.0% pooled coverage at the nominal 95% level, versus 82.1% when correlated poses are treated as independent. The certificate drives a hybrid policy that excites until calibration is trustworthy, refuses uncertified estimates, seeks using continuously re-measured geometry, and detects relay-frame changes via persistent certified inconsistency, avoiding unbounded dead-reckoning drift. Across 200 randomized closed-loop trials, the method attains 0.064 m median station error versus 0.065 m for an oracle given the true relay yaw, despite 27 m median dead-reckoning drift over long horizons; 90 physics-based ROS 2/Gazebo trials retain 30/30 success under nominal operation, communication degradation, and relay-frame disturbances. These results show global localization is unnecessary for reliable hidden-target seeking under this single-relay model, even when both the sensing infrastructure and the vehicle's own reference frame are uncalibrated.
Chinese Translation
我们研究了无锚点条件下的隐藏目标搜寻问题,即车辆、中继节点和目标均不具备可获取的全局位姿。车辆从不直接感知目标,仅通过一个位置和朝向均未知的中继节点接收距离-方位观测,其自身运动也只能通过积分得到的机体坐标系里程计获知。绝对定位在本质上是不可能的:联合构型保留了一个任何估计器都无法消除的精确三维SE(2)规范自由度。然而,控制所需的量仍可完全恢复:运动将校准问题简化为一个与任务相关的商空间,即中继坐标系相对里程计的偏航角以及目标位移,这些量可由两个不同的车辆观测视角以闭合形式辨识。我们提出了一种复杂度为O(K)的多视角自校准估计器,并附带精确的一阶偏航角不确定性证书,该证书能够传播积分里程计所引入的跨视角相关性:完整证书在95%的名义置信水平下达到了95.0%的合并覆盖率,而将相关位姿视为独立时仅为82.1%。该证书驱动一种混合策略:在校准可信之前持续进行激励,拒绝未经认证的估计值,利用持续重新测量的几何信息进行搜寻,并通过持续性可认证的不一致性检测中继坐标系的变化,从而避免无界的航位推算漂移。在200次随机化闭环试验中,该方法实现了0.064米的中位驻留误差,而给定真实中继偏航角的神谕方法为0.065米,尽管长时程下航位推算漂移的中位数达27米;在90次基于物理的ROS 2/Gazebo仿真试验中,该方法在正常运行、通信退化和中继坐标系扰动下均保持30/30的成功率。这些结果表明,在此单中继模型下,即使感知基础设施和车辆自身的参考坐标系均未经校准,可靠的全局定位对于隐藏目标搜寻而言也是不必要的。
cs.RO / 62 / 2609.27938

Remote Surfaces at Your Fingertips: Electrovibration-Based Tactile Feedback for Robot Teleoperation via Touchscreen Interfaces

指尖触达远端表面:基于触屏界面的电振动触觉反馈用于机器人遥操作
Kenan, Alperen, Cárdenas, Juan José García, Tapus, Adriana, Bremner, Paul, Giuliani, Manuel
Abstract
Enabling operators to perceive and interact with remote environments naturally is a fundamental challenge in robotic teleoperation. This is especially critical in tasks involving physical interaction, where real-time haptic awareness improves operational safety and effectiveness. Existing kinesthetic haptic feedback methods suffer from instability during rigid surface contacts and remain sensitive to communication delays, while visual cue-based force feedback imposes additional cognitive load and limits sustained situational awareness. This work presents a teleoperation interface that conveys remote surface interactions to the operator through electrovibration-based tactile feedback, enabling naturally mapped force reflection while avoiding the stability issues associated with kinesthetic feedback and the latency limitations of mechanical actuators. A user study (N=21) evaluated interface usability, sense of presence, and operator workload under two force reflection conditions: visual feedback and electrovibration-based tactile feedback. Characterisation experiments further assessed path-following accuracy and response time across both conditions. Results show that tactile feedback significantly reduced response time by 15.35% (p=0.002, d=0.96) and increased the sense of presence by 31% (p<0.001, d=0.90) compared to visual feedback, while imposing comparable workload and usability across both conditions. These findings demonstrate that electrovibration-based tactile feedback is a viable and effective modality for robot teleoperation, improving operator responsiveness and sense of presence in contact-rich manipulation tasks, with direct applicability to safety-critical domains such as nuclear maintenance.
Chinese Translation
使操作者能够自然地感知远程环境并与之交互是机器人遥操作中的一项根本性挑战。在涉及物理交互的任务中,这一点尤为关键,因为实时的触觉感知能够提升操作的安全性和有效性。现有的动觉(kinesthetic)触觉反馈方法在接触刚性表面时存在稳定性问题,且对通信延迟十分敏感;而基于视觉提示的力反馈则会带来额外的认知负荷,并限制了持续的态势感知能力。本工作提出了一种遥操作界面,通过基于电振动(electrovibration)的触觉反馈将远程表面交互信息传递给操作者,实现自然映射的力再现,同时避免了动觉反馈相关的稳定性问题以及机械执行器的延迟限制。一项用户实验(N=21)在两种力再现条件下(视觉反馈与基于电振动的触觉反馈)评估了界面的可用性、临场感以及操作者负荷。表征实验进一步评估了两种条件下的路径跟踪精度和响应时间。结果表明,与视觉反馈相比,触觉反馈将响应时间显著缩短了15.35%(p=0.002,d=0.96),并将临场感提升了31%(p<0.001,d=0.90),同时两种条件下的工作负荷和可用性相当。这些发现表明,基于电振动的触觉反馈是机器人遥操作中一种可行且有效的反馈模态,能够提升操作者在接触密集型操作任务中的响应能力和临场感,并可直接应用于核设施维护等安全关键领域。
cs.RO / 63 / 2609.28027

Learning a Speed-adaptive Hip Exoskeleton Control Policy Via Sim-to-real Reinforcement Learning

基于模拟到真实强化学习训练速度自适应髋关节外骨骼控制策略
Li, Bin, Hou, Zhimin, Hou, Jiacheng, Liang, Zenian, Wu, Tong, Ma, Teng, Fu, Chenglong
Abstract
Providing personalized exoskeleton assistance across varying walking speeds remains challenging. Existing online optimization methods are sample-inefficient, requiring extensive human-in-the-loop (HIL) evaluations to optimize the entire assistive torque profile. Sim-to-real reinforcement learning (RL) offers a promising alternative but cannot directly account for individual user preferences. We propose a framework integrating sim-to-real RL with online preference learning for personalized exoskeleton assistance. Specifically, assistance timing is learned in simulation by training RL policies with human musculoskeletal models across varying walking speeds. The learned policies are then distilled and deployed on a physical hip exoskeleton using onboard sensory observations. Gaussian-process-based preference learning further personalizes the assistance magnitude through pairwise user comparisons. By decoupling assistance timing learning in simulation from magnitude optimization in real-world experiments, our framework substantially reduces the online optimization space. Human-subject experiments demonstrate efficient identification of personalized assistive torque profiles across varying walking speeds with fewer real-world evaluations.
Chinese Translation
在不同行走速度下提供个性化的外骨骼辅助仍然是一项挑战。现有的在线优化方法样本效率低下,需要大量的人在环路(HIL)评估来优化完整的辅助力矩曲线。模拟到真实(sim-to-real)强化学习(RL)提供了一种有前景的替代方案,但无法直接考虑个体用户的偏好。我们提出了一个将模拟到真实强化学习与在线偏好学习相结合的个性化外骨骼辅助框架。具体而言,通过在仿真中使用人体肌肉骨骼模型在不同行走速度下训练强化学习策略来学习辅助时机。随后将学习到的策略进行蒸馏,并利用机载传感器观测部署到物理髋关节外骨骼上。基于高斯过程的偏好学习进一步通过用户的成对比较来个性化辅助幅度。通过将仿真中的辅助时机学习与真实世界实验中的幅度优化解耦,我们的框架显著减小了在线优化空间。人体受试者实验表明,该框架能够以更少的真实世界评估高效地识别出适用于不同行走速度的个性化辅助力矩曲线。
cs.RO / 64 / 2609.28044

AeRSoM: An Aerial Rigid-Soft Integrated Manipulator for Contact-Rich Manipulation

AeRSoM:一种面向接触丰富操作的空中刚柔集成机械臂
Liang, Jiacheng, Zhong, Hang, Wang, Yaonan, Chen, Ge, Zhang, Zhixing, Tian, Bocheng, Zhang, Hui, Wen, Li
Abstract
Contact-rich aerial manipulation remains fundamentally challenging because interaction forces are directly transmitted to the aerial platform, often leading to instability and degraded task performance. While compliant manipulators can mitigate these effects, existing aerial manipulation systems typically struggle to reconcile interaction compliance with manipulation precision. To this end, this article presents an aerial rigid-soft integrated manipulator (AeRSoM) robot that realizes embodied compliance for aerial manipulation. The proposed system integrates a fully actuated aerial platform, a rigid-soft manipulator, and variable-stiffness regulation to simultaneously achieve stable flight, compliant interaction, and precise manipulation. By distributing compliance throughout the manipulation system, the proposed design leverages distributed embodied compliance to passively absorb contact disturbances while preserving sufficient stiffness for task execution. To fully exploit the mechanical design, a composite control framework is developed for precise end-effector trajectory tracking in the presence of uncertainties and external disturbances. Extensive real-world experiments are conducted in representative contact-rich aerial manipulation tasks, including dynamic transmission-line grasping, physical interaction with a wind turbine blade, peg-in-hole, and screwing operations. The results demonstrate that the proposed rigid-soft integration significantly improves interaction robustness and task adaptability while maintaining manipulation accuracy, highlighting that embodied compliance provides a promising design paradigm for enhancing the safety, robustness, and versatility of aerial manipulation.
Chinese Translation
接触丰富的空中操作本质上极具挑战性,因为交互力会直接传递到空中平台上,常常导致不稳定性和任务性能下降。虽然柔性机械臂可以缓解这些影响,但现有的空中操作系统通常难以在交互柔顺性与操作精度之间取得平衡。为此,本文提出了一种实现空中操作具身柔顺性的空中刚柔集成机械臂(AeRSoM)机器人。该系统集成了全驱动空中平台、刚柔机械臂和变刚度调节,以同时实现稳定飞行、柔顺交互和精确操作。通过将柔顺性分布于整个操作系统,所提出的设计利用分布式具身柔顺性被动吸收接触扰动,同时保留足够的刚度以执行任务。为充分发挥机械设计的潜力,本文开发了一种复合控制框架,用于在存在不确定性和外部扰动的情况下实现末端执行器的精确轨迹跟踪。我们在典型的接触丰富空中操作任务中开展了大量真实世界实验,包括动态输电线抓取、与风力涡轮机叶片的物理交互、轴孔装配和拧螺丝操作。结果表明,所提出的刚柔集成设计在保持操作精度的同时显著提高了交互鲁棒性和任务适应性,凸显了具身柔顺性作为提升空中操作安全性、鲁棒性和通用性的有前景的设计范式。
cs.RO / 65 / 2609.28107

Distillation for Efficient Multitask Manipulation Policies via Conditional Flow Matching

基于条件流匹配的蒸馏方法用于高效多任务操作策略
Deshmukh, Shreya, Mahdi, Imen, Heppert, Nick, Valada, Abhinav
Abstract
Advances in generative modeling have recently been extensively employed in robotics for policy learning. In particular, Conditional Flow Matching (CFM) trained with expert demonstrations has been shown to outperform existing methods on robot manipulation benchmarks. While prior work has mainly focused on single-task settings, we study the problem from a multi-task perspective, as training independent models for each task is computationally expensive. Multi-Task policy learning comes with its own set of challenges, as naively training on a concatenated dataset of demonstrations would either require increased model capacity to accommodate the added complexity or result in drops in performance. We propose to distill knowledge from single-task CFM experts into a shared multi-task policy by transferring their learned velocity fields. We combine this distillation signal with the original CFM objective to retain fidelity to the demonstrations. Experiments on RLBench show that our approach improves multi-task policy performance over naive training while maintaining a fixed model size.
Chinese Translation
生成式建模的最新进展近来被广泛应用于机器人领域的策略学习。特别是,使用专家演示训练的条件流匹配(Conditional Flow Matching, CFM)已被证明在机器人操作基准测试中优于现有方法。以往的工作主要集中在单任务设置上,而我们从多任务的角度研究该问题,因为为每个任务独立训练模型在计算上代价高昂。多任务策略学习本身也面临一系列挑战:若简单地在拼接后的演示数据集上训练,要么需要增加模型容量以应对额外的复杂度,要么会导致性能下降。我们提出将单任务CFM专家的知识蒸馏到一个共享的多任务策略中,通过迁移其学习到的速度场实现。我们将这一蒸馏信号与原始的CFM目标相结合,以保持对演示数据的忠实性。在RLBench上的实验表明,我们的方法在保持模型规模固定的同时,提升了多任务策略相对于朴素训练的性能。
cs.RO / 66 / 2609.28131

DEAL-Grasp: Decoupled Alignment Representation for Geometry-Aware Dexterous Grasp Generation

DEAL-Grasp:面向几何感知灵巧抓取生成的解耦对齐表示
Zhao, Fuqiang, Liu, Qian
Abstract
Synthesizing realistic articulated hand-object interactions is a fundamental problem in virtual reality, embodied intelligence, and digital human applications. Existing methods for dexterous grasp synthesis typically regress or denoise poses in a joint space that couples global rigid motion with local articulation, which often yields unstable samples and physically implausible contacts. We introduce DEAL-Grasp, built upon the Decoupled Alignment (DEAL) representation, which reformulates grasp synthesis as alignment-space generation: the interaction state comprises task-space geometric anchors and articulation parameters, from which the rigid transform is recovered via closed-form Procrustes alignment while preserving local articulation. On this mixed state, we model grasp generation using heterogeneous-state flow matching with component-wise vector fields, incorporating time-adaptive physical regularization during training. At inference, grasps are synthesized solely by integrating the learned vector field, without test-time optimization or auxiliary physical guidance. Across MultiDex and zero-shot RealDex benchmarks, DEAL-Grasp attains high force-perturbation success rates alongside minimal penetration and high diversity of generated grasps, while substantially reducing native inference latency compared to optimization-heavy baselines. The project page is available at https://wmtlab.github.io/DEAL-Grasp/.
Chinese Translation
合成逼真的 articulated 手-物体交互是虚拟现实、具身智能和数字人应用中的一项基础性问题。现有的灵巧抓取合成方法通常在将全局刚体运动与局部关节运动耦合在一起的关节空间中进行位姿回归或去噪,这往往导致样本不稳定以及物理上不合理的接触。我们提出 DEAL-Grasp,它建立在解耦对齐(Decoupled Alignment,DEAL)表示之上,将抓取合成重新表述为对齐空间中的生成问题:交互状态由任务空间的几何锚点和关节参数组成,并通过闭式 Procrustes 对齐恢复刚体变换,同时保留局部关节运动。在这一混合状态上,我们采用具有分量式向量场的异构状态流匹配来建模抓取生成,并在训练过程中引入时间自适应的物理正则化。在推理阶段,抓取完全通过积分学习到的向量场来合成,无需测试时优化或辅助物理引导。在 MultiDex 和零样本 RealDex 基准测试中,DEAL-Grasp 在实现高抗力扰动成功率的同时,具有最小的穿透量和高度多样化的生成抓取,并且与以优化为主的基线方法相比,显著降低了原生推理延迟。项目页面见 https://wmtlab.github.io/DEAL-Grasp/。
cs.RO / 67 / 2609.28161

Dissecting Advantage-Guided Post-Training for Vision-Language-Action Policies

解析视觉-语言-动作策略中的优势引导后训练方法
Cao, Jiahang, Zhao, Hanye, Lai, Hang, Zhang, Shenyu, Han, Xiaoshen, Li, Xinghang, Liu, Futeng, Peng, Wanli, Wang, Heyun, Wang, Yunhong, Li, Jason, Yu, Yong, Zhang, Weinan
Abstract
Advantage-guided reinforcement learning provides a practical way to post-train vision-language-action (VLA) policies using limited robot data. However, its performance depends on several coupled choices, including how critic-derived advantages are constructed, calibrated, and used for policy training. Existing recipes often combine these choices into a single end-to-end procedure, making their individual effects difficult to identify. In this work, we dissect advantage-guided VLA post-training through a controlled empirical study that separates these design choices while accounting for their distinct estimands. We develop stage-specific offline evaluation methods to screen alternative choices efficiently, without requiring extensive real-robot policy evaluations for every possible combination. The staged evaluation identifies a modular recipe that combines temporal-difference advantage construction, group-wise calibration, and continuous advantage weighting. Across four real-world bimanual tasks, the resulting recipe improves mean task progress and success over the SFT initialization by 0.42 and 0.63, respectively. Moreover, the proposed evaluation diagnostics show an overall alignment with downstream real-world performance, supporting their use for interpreting empirical outcomes and selecting advantage-guided post-training designs in practice.
Chinese Translation
优势引导的强化学习为利用有限的机器人数据对视觉-语言-动作(VLA)策略进行后训练提供了一种实用方法。然而,其性能取决于若干相互耦合的设计选择,包括如何构建、校准评论家网络(critic)导出的优势值以及如何将其用于策略训练。现有方法通常将这些选择组合在单一的端到端流程中,导致难以识别各选择的独立效果。在本工作中,我们通过一项受控的实证研究来剖析优势引导的VLA后训练,该研究在考虑各选择不同估计目标的前提下,将这些设计选择分离开来。我们开发了针对各阶段的离线评估方法,以高效筛选备选方案,而无需针对每种可能的组合进行大量的真实机器人策略评估。这一分阶段评估确定了一种模块化方案,该方案结合了时序差分优势构建、分组校准以及连续优势加权。在四项真实世界双手操作任务中,所得方案相较于SFT初始化,将平均任务进度和成功率分别提升了0.42和0.63。此外,所提出的评估诊断方法与下游真实世界性能表现出总体一致性,支持其用于解释实证结果并在实践中选择优势引导的后训练设计。
cs.RO / 68 / 2609.28172

Dynamic, Decentralized Spatial Code Reuse for OCDMA LiDAR in Robot Swarms

面向机器人集群OCDMA激光雷达的动态去中心化空间码复用方法
Alomari, Mohammad Hani
Abstract
Robots in a LiDAR-equipped swarm mutually interfere when their optical ranging codes collide. Existing mitigations either assign codes statically -- requiring $L=N$ distinguishable codes for $N$ robots -- or react to detected interference without a scalable, coordinated assignment rule beneath them; prior work explicitly identifies the code-assignment scaling problem as unsolved. We propose a decentralized protocol in which robots dynamically reassign spatial reuse codes based on a live, beacon-maintained interference-neighborhood graph, and prove that the number of codes required grows as $O(\log N/\log\log N)$ under constant robot density -- an unbounded improvement over the $\Theta(N)$ growth of static assignment. We validate this result under conditions substantially beyond the idealized proof -- robot mobility, imperfect beacon-based detection, and reactive reassignment -- via Monte Carlo simulation (30 seeds per condition, 95% confidence intervals): the advantage over static assignment widens from roughly $2\times$ at 15 robots to $12\times$ at 120. Against a structurally faithful, fairly constructed model of an existing coordination-free approach, our protocol achieves both substantially greater code-reuse efficiency and 30--40% lower collision risk under an identical, constrained code budget, demonstrating that coordination -- not merely reactivity -- is what closes the scaling gap.
Chinese Translation
在装备激光雷达(LiDAR)的机器人集群中,当各机器人的光学测距码发生冲突时会相互干扰。现有缓解方法要么静态分配编码——即对N个机器人需要L=N个可区分的编码——要么在检测到干扰后被动响应,而缺乏可扩展的、协调一致的分配规则作为支撑;已有研究明确指出编码分配的规模化问题仍未解决。我们提出一种去中心化协议,机器人基于由信标实时维护的干扰邻域图动态重新分配空间复用编码,并证明在机器人密度恒定的条件下,所需编码数量以O(log N/log log N)的速度增长——相比静态分配的Θ(N)增长是无界的改进。我们通过蒙特卡洛仿真(每种条件30个随机种子,95%置信区间),在远超理想化证明假设的条件下验证了该结果——包括机器人移动性、不完美的基于信标的检测以及被动式重新分配:相对于静态分配的优势从15个机器人时的约2倍扩大到120个机器人时的12倍。与一个结构忠实且构造公平的现有无协调方法的模型相比,在相同且受限的编码预算下,我们的协议既实现了显著更高的编码复用效率,又将碰撞风险降低了30–40%,表明真正弥合规模化差距的是协调性——而不仅仅是被动响应。
cs.RO / 69 / 2609.28175

DAVIS: A Depth-Only End-to-End Active-Vision Framework for Humanoid Soccer Skills

DAVIS:一个仅基于深度图像的人形机器人足球技能端到端主动视觉框架
Jin, Jiakang, Huo, Yixiao, Wang, Pengyuan, Han, Yinan, Zhang, Tingxuan, Zhao, Zhuobing, Zhou, Xuanxin, Ye, Zhangchen, Ruan, Enxuan, Bao, Yifei, Yang, Jiankun, Sun, Chenghao, Cui, Wenhao, Tian, Xiaoyu, Li, Yiming
Abstract
Humanoid soccer contact skills require more than producing high-impact foot-ball contacts: the robot must close the loop over perception, approach, alignment, impact, and recovery while its own motion induces substantial viewpoint changes, frequent loss of the ball from view, and uncertain contact outcomes. In this work, we ask a compact yet stricter question: can a humanoid learn soccer contact skills using only a head-mounted depth image, proprioceptive history, and an optional low-dimensional task command, and directly output 25-DoF joint PD targets without extra runtime perception or planning modules? To this end, we propose DAVIS, a depth-only end-to-end framework for humanoid soccer skills that learns visibility-aware auxiliary geometry during training, and combines GT-to-prediction annealing, task curricula, and AMP-style motion priors to smoothly bridge privileged supervision and real deployment. Built on this framework, we instantiate representative soccer contact skills, including goal-directed shooting and directional dribbling, through task-specific definitions of objects, commands, rewards, and curricula, and validate them through simulation, Noetix E1 real-robot experiments, and ablations.
Chinese Translation
人形机器人足球接触技能不仅仅是产生高冲击力的脚-球接触:机器人必须在感知、接近、对准、冲击和恢复之间形成闭环,同时其自身运动会引起显著的视角变化、球频繁脱离视野以及不确定的接触结果。在本工作中,我们提出了一个简洁但更严格的问题:人形机器人能否仅使用头戴式深度图像、本体感觉历史以及可选的低维任务指令来学习足球接触技能,并直接输出25自由度关节PD目标,而无需额外的运行时感知或规划模块?为此,我们提出了DAVIS,一个仅基于深度图像的人形机器人足球技能端到端框架,该框架在训练过程中学习具有可见性感知的辅助几何信息,并结合真值到预测的退火、任务课程和AMP风格的运动先验,平滑地桥接特权监督与真实部署。基于该框架,我们通过对物体、指令、奖励和课程的任务特定定义,实现了代表性的足球接触技能,包括目标导向射门和方向性带球,并通过仿真、Noetix E1真机实验和消融实验对其进行了验证。
cs.RO / 70 / 2609.28179

GLASS: Architecture-Tuned, Composable, Device-Side Linear Algebra for Edge Robotics and Beyond

GLASS:面向边缘机器人及更广泛领域的架构调优、可组合的设备端线性代数库
Plancher, Brian
Abstract
GPU robotics lacks the reusable numerical infrastructure of mature CPU stacks, instead relying on compiler frameworks that introduce overhead or repeatedly reimplementing numerical libraries. To address this, we introduce GLASS (GPU Linear Algebra Simple Subroutines), a header-only CUDA C++ library that provides thread-, warp-, block-, and NVIDIA-backed implementations of robotics-scale linear algebra and geometric computations under one composable device API. GLASS treats implementation choice, execution scope, and launch packing as architecture-specific placement decisions determined by offline measurement and resolved statically at compile time. This is critical as the best and worst placements differ by a median of 4.9x (max 81x), with 145 of 396 recommended placements changing between a Jetson AGX Orin and an RTX 5090, and 162 of 396 versus an AGX Xavier. These stakes are highest at the edge as GLASS's advantage over the best of PyTorch and JAX is as much as 73x on the Orin versus 12x on the RTX 5090. GLASS is released open source with independent numerical oracles and source-bound local-GPU test attestation. Finally, integrating GLASS with published robotics systems both exposed a pre-existing numerical bug and improved embedded runtimes by up to 1.5x.
Chinese Translation
GPU机器人领域缺乏成熟CPU技术栈那样可复用的数值计算基础设施,通常依赖会引入额外开销的编译器框架,或不得不反复重新实现数值库。为解决这一问题,我们提出了GLASS(GPU Linear Algebra Simple Subroutines,GPU线性代数简单子程序),这是一个仅含头文件的CUDA C++库,在一个可组合的设备端API下,为机器人尺度的线性代数与几何计算提供线程级、线程束级、线程块级以及NVIDIA后端支持的多种实现。GLASS将实现选择、执行范围和启动打包视为架构特定的放置决策,这些决策由离线测量确定,并在编译时静态解析。这一点至关重要,因为最优与最差放置方案的性能差异中位数达4.9倍(最大81倍);在Jetson AGX Orin与RTX 5090之间,396个推荐放置方案中有145个发生变化,与AGX Xavier相比则有162个发生变化。这些差异在边缘设备上影响最大:相比PyTorch与JAX中的最优者,GLASS在Orin上的加速比最高达73倍,而在RTX 5090上为12倍。GLASS以开源形式发布,并附带独立的数值验证基准(oracle)以及基于源码的本地GPU测试证明。最后,将GLASS集成到已发表的机器人系统中,既暴露了一个先前存在的数值错误,又将嵌入式运行时性能提升至多1.5倍。
cs.RO / 71 / 2609.28184

VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation

视觉语言模型能描述却无法度量:面向机器人操作的对象中心场景理解
Saccon, Enrico, Faraci, Tommaso, Zarzuelo, Iñigo De La Ossa, Palopoli, Luigi, Roveri, Marco, Saveriano, Matteo
Abstract
Robotic operation in previously unseen environments requires both semantic understanding and reliable metric information. While vision--language models (VLMs) provide strong semantic capabilities, their geometric estimates remain less reliable. In this paper, we propose a VLM-driven, modular perception framework for scene understanding using off-the-shelf approaches. Starting from a single RGB-D observation, the scene is segmented into object-level regions, annotated by a VLM, and grounded with depth information to construct a task-independent object-centric representation. Experiments on 151 tabletop scenes show that the proposed decomposition preserves strong semantic performance while substantially improving localization and depth estimation over direct VLM inference. The resulting representation is also integrated with a task-planning framework for robotic execution.
Chinese Translation
机器人在先前未见过的环境中执行操作,既需要语义理解,也需要可靠的度量信息。尽管视觉语言模型(VLM)具备强大的语义能力,但其几何估计的可靠性仍然较低。本文提出一种由VLM驱动的模块化场景理解感知框架,该框架基于现有成熟方法构建。从单次RGB-D观测出发,系统将场景分割为对象级区域,由VLM进行语义标注,并结合深度信息进行定位,从而构建一个与任务无关的对象中心表示。在151个桌面场景上的实验表明,所提出的分解方法在保持较强语义性能的同时,相比直接使用VLM推理,显著提升了定位和深度估计的准确性。所得到的表示还与任务规划框架集成,用于机器人执行。
cs.RO / 72 / 2609.28225

Large-Scale Geometric Map-Based Localization of UAVs in GNSS-Denied Urban Environments

GNSS拒止城市环境下基于大规模几何地图的无人机定位
Terlizzi, Garth, Fathian, Kaveh
Abstract
Unmanned aerial vehicles (UAVs) operating in GNSS-denied urban environments require alternative methods for position estimation. Existing approaches based on satellite image retrieval or learned descriptors are sensitive to appearance variation and degrade rapidly as the search area grows. We present a vision-based localization system that matches building patterns observed from a downward-facing UAV camera against a reference building footprint database. Our approach detects buildings in aerial imagery, accumulates observations across frames into a unified map, and matches local building arrangements against reference footprints using a novel geometry-driven descriptor that augments local triangle structure with per-building shape features. By encoding spatial relationships between nearby buildings rather than visual appearance, the system is robust to appearance variations and remains discriminative over large search areas. Evaluations on seven flights across four municipalities in a large metropolitan area demonstrate 100% Recall@1 at search areas of approximately 113 km$^2$ and 254 km$^2$, and 71.4% Recall@1 when expanded to approximately 452 km$^2$, encompassing up to 277,000 buildings. In contrast, baseline methods degrade rapidly and achieve 0% Recall@1 at 254 km$^2$ and 452 km$^2$.
Chinese Translation
在GNSS拒止的城市环境中运行的无人机(UAV)需要替代方法来进行位置估计。现有的基于卫星图像检索或学习描述子的方法对 appearance 变化敏感,且随着搜索区域的增大性能迅速下降。我们提出了一种基于视觉的定位系统,将下视无人机相机观测到的建筑物模式与参考建筑物轮廓(footprint)数据库进行匹配。我们的方法在航拍图像中检测建筑物,将跨帧的观测累积到统一地图中,并使用一种新颖的几何驱动描述子将局部建筑物布局与参考轮廓进行匹配,该描述子在局部三角形结构的基础上增强了每栋建筑物的形状特征。通过编码邻近建筑物之间的空间关系而非视觉外观,该系统对 appearance 变化具有鲁棒性,并在大范围搜索区域内保持良好的区分能力。在一个大都市圈四个城市共七次飞行的评估中,系统在约113平方公里和254平方公里的搜索区域内实现了100%的 Recall@1,在扩展到约452平方公里、涵盖多达27.7万栋建筑时仍达到71.4%的 Recall@1。相比之下,基线方法性能迅速下降,在254平方公里和452平方公里时 Recall@1 均为0%。
cs.RO / 73 / 2609.28247

Controlling Collectives of AI Agents in Reasoning Space with Spatial Transformers

基于空间Transformer的推理空间中AI智能体集群控制
Vatnsdal, Frederic, Gopal, Roshan, Camargo, Romina Garcia, Kumar, Vijay, Ribeiro, Alejandro
Abstract
Large Language Models (LLMs) introduce an exciting new paradigm for planning and navigation in robotics, but fail on even simple multi-robot tasks as team sizes grow. We propose COMPASS, a scalable, decentralized multi-robot architecture for controlling large collectives of agentic robots with reasoning space feedback control. Feedback is generated locally on each robot by a spatial transformer which aggregates multi-hop messages across the fleet into a learned feedback token. Our experiments find that collectives of language models demonstrate performance gains from structured diversity of the input command, which can cancel biases; an advantage that is held across scale. Compared against a centralized frontier LLM policy and a language-only communication ablation, we find that the coupled design of COMPASS decisively produces cohesive flocking formations that accurately fly the commanded intent. We show that reasoning feedback works best when composed with a compact learned token. Our ablations show that hand engineered feedback with raw state appearing in the language channel obliterates cohesion. COMPASS generalizes zero-shot to unseen instructions of ambiguous meaning while commanding flocks up to 16 times its training scale, flying up to 1024 robots under natural language commands.
Chinese Translation
大语言模型(LLM)为机器人领域的规划与导航引入了一个令人兴奋的新范式,但随着团队规模的增长,即使是简单的多机器人任务也会失败。我们提出了COMPASS,一种可扩展的去中心化多机器人架构,通过推理空间反馈控制来控制大规模智能体机器人集群。反馈由每台机器人上的空间Transformer(Spatial Transformer)在本地生成,该模型将整个集群的多跳消息聚合为一个学习得到的反馈token。我们的实验发现,语言模型集群能够从输入指令的结构化多样性中获得性能提升,这种多样性可以抵消偏差,且这一优势在不同规模下均成立。与集中式前沿LLM策略以及纯语言通信消融版本相比,我们发现COMPASS的耦合设计能够决定性地产生内聚的集群编队,并精确执行所指令的意图。我们表明,推理反馈与一个紧凑的学习token相结合时效果最佳。消融实验表明,在语言通道中出现原始状态的手工设计反馈会完全破坏内聚性。COMPASS能够零样本泛化到含义模糊的未见指令,并可指挥规模达其训练规模16倍的集群,在自然语言指令下控制多达1024台机器人飞行。
cs.RO / 74 / 2609.28256

MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

MemBodied:面向视觉-语言-动作模型的循环联想记忆
Pala, Tej Deep, Majumder, Navonil, Goh, Bryce, Yee, Raphael, Yang, Jianfei, Chen, Liming, Poria, Soujanya
Abstract
Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information available only in past observations. Retaining past observations in context can aid in recovering this information, but at the significant cost of ever-growing, bloated context and inference latency. We thus introduce MemBodied, a fixed-size episodic memory with two complementary components: an associative state that records interactions across policy calls and an episode anchor that preserves a compact representation of the initial scene as a reference. At each policy call, the model conditions action generation on the current input and the memory components, rather than directly using past observations. Across five evaluated RMBench tasks requiring memory, MemBodied achieves $7.81\times$ the mean success rate of a stateless policy and $2.98\times$ of vanilla recurrent memory, while outperforming the strongest memory-augmented baseline by $1.3\times$ with $10\times$ fewer added parameters. On the fully observable LIBERO-Long suite, it reached 90.6%, a 5.4% improvement over the stateless $\pi_0$ policy. These findings support MemBodied as a practical alternative to expanding the policy context for history-dependent manipulation.
Chinese Translation
视觉-语言-动作(Vision-Language-Action)模型为通用机器人控制提供了坚实的基础,然而绝大多数策略在当前观测之外无法保留和利用回合级(episode-level)信息。这一局限在依赖历史信息的操作任务中影响重大,因为这类任务所需的信息仅存在于过去的观测中。将过去的观测保留在上下文中可以帮助恢复这些信息,但代价是上下文不断膨胀、推理延迟显著增加。为此,我们提出了 MemBodied——一种固定大小的情景记忆,包含两个互补组件:一个记录策略调用之间交互信息的联想状态(associative state),以及一个以紧凑表示保存初始场景作为参照的回合锚点(episode anchor)。在每次策略调用时,模型基于当前输入和记忆组件来条件化动作生成,而非直接使用过去的观测。在五个需要记忆能力的 RMBench 任务上,MemBodied 的平均成功率是无状态策略的 7.81 倍、朴素循环记忆的 2.98 倍,同时以少 10 倍的额外参数量超越了最强记忆增强基线 1.3 倍。在全可观测的 LIBERO-Long 套件上,它达到了 90.6% 的成功率,比无状态的 π₀ 策略提升了 5.4%。这些结果表明,对于依赖历史的操作任务,MemBodied 是扩展策略上下文的一种实用替代方案。
cs.RO / 75 / 2609.28258

Generalizable Robotic Insertion with World Models

基于世界模型的泛化机器人插装
Hansen, Nicklas, Akinola, Iretiayo, Guo, Yijie, Xu, Jie, Tang, Bingjie, Su, Hao, Wang, Xiaolong, Gupta, Abhishek, Fox, Dieter, Narang, Yashraj
Abstract
Robotic assembly in high-mixture settings requires adaptable systems that can handle diverse parts, yet current approaches typically rely on policies specialized to each insertion task. Although this can reach high success rates, it makes the process of deploying systems for new problems tedious and time consuming. We present a framework for generalizable insertion using world models that combine robot proprioceptive information with raw visual observations captured by a wrist-mounted camera. Our model-based approach trains a single world model on up to 90 insertion tasks with geometrically diverse parts, achieving 56% zero-shot success on unseen objects with unknown geometry compared to just 7% with a model-free baseline. Importantly, performance improves as more objects are included in the training dataset, demonstrating strong scalability. Lastly, finetuning the generalist model on held-out objects significantly enhances data-efficiency compared to training from scratch and, in some cases, achieves better asymptotic performance. To our knowledge, this is the first system capable of assembling unseen objects in an entirely data-driven manner, and thus represents a significant step toward scalable, generalizable robotic assembly systems.
Chinese Translation
高混合作业场景下的机器人装配需要能够处理多样化零件的自适应系统,然而现有方法通常依赖于针对每个插装任务专门训练的策略。尽管这种方式可以达到很高的成功率,但为新问题部署系统的过程却十分繁琐且耗时。我们提出了一个基于世界模型的可泛化插装框架,该框架将机器人本体感知信息与腕部相机捕获的原始视觉观测相结合。我们基于模型的方法在多达90个包含几何形状多样零件的插装任务上训练了单一世界模型,对于几何形状未知的未见物体,其零样本成功率达到56%,而无模型基线仅为7%。重要的是,随着训练数据集中物体数量的增加,性能持续提升,展现出强大的可扩展性。最后,与从头训练相比,在留出的物体上对通用模型进行微调显著提高了数据效率,并且在某些情况下取得了更好的渐近性能。据我们所知,这是首个能够以完全数据驱动的方式装配未见物体的系统,因而代表了向可扩展、可泛化的机器人装配系统迈出的重要一步。
cs.RO / 76 / 2609.28281

BrickCraft-Duo: Efficient Dual-Arm Skill Learning and Refinement for Compositional Long-Horizon Assembly

BrickCraft-Duo:面向组合式长时程装配的高效双臂技能学习与优化
Yu, Jichuan, Xiao, Zhenyu, Wang, Ze, Liu, Ruixuan, Liu, Changliu, Hu, Chuxiong
Abstract
Interlocking brick assembly provides a representative testbed for evaluating real-world robotic manipulation capabilities, where diverse structural designs, complex inter-step dependencies, intricate mechanical interactions and tight insertion tolerances pose substantial challenges. We present BrickCraft-Duo, a modular framework for long-horizon dual-arm collaborative assembly of interlocking bricks through data-efficient skill learning and composition. BrickCraft-Duo learns reusable single- and dual-arm assembly skills from diverse demonstrations, with bilateral symmetry alignment facilitating skill sharing across symmetric arms and assembly--support role assignments. Guided by stability-aware assembly reasoning, BrickCraft-Duo composes heterogeneous skills to achieve autonomous long-horizon execution, and further integrates human-in-the-loop correction for targeted skill refinement. The resulting system achieves long-horizon success rates of at least 60% and step-level completion rates of at least 95% across five real-world assembly tasks involving partially supported configurations, with horizons of up to nine steps. Project website: https://jichuan-yu.github.io/BrickCraft-Duo.
Chinese Translation
互锁积木装配为评估真实世界机器人操作能力提供了一个代表性测试平台,其中多样化的结构设计、复杂的步骤间依赖关系、精密的机械交互以及严格的插入公差带来了巨大挑战。我们提出 BrickCraft-Duo,一个模块化框架,通过数据高效的技能学习与组合,实现互锁积木的长时程双臂协同装配。BrickCraft-Duo 从多样化的演示中学习可复用的单臂与双臂装配技能,并通过双侧对称性对齐促进对称双臂之间的技能共享以及装配-支撑角色分配。在稳定性感知装配推理的引导下,BrickCraft-Duo 组合异构技能以实现自主的长时程执行,并进一步集成人在回路(human-in-the-loop)的纠错机制,实现针对性的技能优化。该系统在五项涉及部分支撑构型的真实装配任务中,实现了至少 60% 的长时程成功率和至少 95% 的步骤级完成率,任务时程最长可达九个步骤。项目网站:https://jichuan-yu.github.io/BrickCraft-Duo。
cs.RO / 77 / 2609.28296

Talk2Escape: Conversational Grounding for Vision-and-Language Navigation

Talk2Escape:面向视觉-语言导航的对话式定位机制
Li, Zerui, Lin, Sihao, Shao, Yanyan, Zhang, Jiwen, Shi, Xiangyu, Li, Shijie, Wu, Qi
Abstract
While Vision-and-Language Navigation (VLN) has demonstrated remarkable success, the prevailing single-turn paradigm exposes a fundamental vulnerability: agents operate in a strictly open-loop manner. In practice, factors such as perceptual aliasing, sensor noise, and odometry drift can cause minor deviations to accumulate over time, often leading to catastrophic mission failures with no built-in mechanism for error recovery. To address this, we introduce \textit{Talk2Escape}, a proactive and model-agnostic dialogue intervention framework that reframes navigation as a closed-loop interactive process. At its core, a lightweight vision-language module continuously monitors agent kinematics. Upon detecting localized looping or severe trajectory divergence, it translates raw egocentric observations into concise, grounded queries to solicit targeted corrective feedback from either an algorithmic oracle or a human-in-the-loop. Extensive evaluations in high-fidelity simulators, including R2R-CE, RxR-CE, and VLNVerse, demonstrate that \textit{Talk2Escape} exhibits consistent improvements across diverse base agents. Empirically, \textit{Talk2Escape} achieves a 66.0\% Success Rate on R2R-CE, outperforming the current supervised and zero-shot state-of-the-art methods. We further validate its sim-to-real transfer on a Unitree Go2 quadruped, proving that proactive dialogue drastically improves navigation robustness in physical environments.
Chinese Translation
尽管视觉-语言导航(Vision-and-Language Navigation, VLN)已取得显著成功,但主流的单轮交互范式暴露出一个根本性缺陷:智能体以严格的开环方式运行。在实际应用中,感知混叠、传感器噪声和里程计漂移等因素会导致微小偏差随时间不断累积,且由于缺乏内建的误差恢复机制,往往最终导致灾难性的任务失败。为解决这一问题,我们提出了Talk2Escape,一个主动式、模型无关的对话干预框架,将导航重新构建为闭环交互过程。该框架的核心是一个轻量级视觉-语言模块,可持续监测智能体的运动状态。当检测到局部循环或严重的轨迹偏离时,该模块将原始的自我中心观测转化为简洁、有依据的查询,以从算法智能体或人在回路(human-in-the-loop)中获取针对性的纠正反馈。在高保真仿真环境(包括R2R-CE、RxR-CE和VLNVerse)中的大量评估表明,Talk2Escape在多种基础智能体上均取得一致的改进。实验结果表明,Talk2Escape在R2R-CE上达到66.0%的成功率,优于当前监督式和零样本的最先进方法。我们进一步在Unitree Go2四足机器人上验证了其从仿真到现实(sim-to-real)的迁移能力,证明主动式对话能显著提升物理环境中的导航鲁棒性。
cs.RO / 78 / 2609.28299

Contact-Implicit Stein Projected ADMM for Discovery of Diverse Contact-Rich Manipulation Strategies

用于发现多样化接触密集型操作策略的接触隐式Stein投影ADMM方法
Sathyanarayan, Hrishikesh, Hughes, Christian, Abraham, Ian
Abstract
Contact-implicit trajectory optimization formulates contact-rich manipulation as a single constrained program; however, that single program run collapses onto one local optimum out of many equally valid contact modes, grasps, or push directions. As a consequence, the resulting manipulation strategy is reluctant to change and sensitive to initialization. In order to promote robust manipulation, this paper investigates how contact-implicit solvers can discover diverse contact-rich strategies. Our approach derives a variation of Consensus Alternating Direction Method of Multipliers (ADMM) combined with Stein variational inference methods to output a set of distinct contact-rich solutions. We find that applying the Stein repulsive force to ADMM's split variable (rather than its primal form) allows for effective coverage over the set of feasible contact strategies without prematurely stalling the solver. We demonstrate the effectiveness of our approach on a variety of contact-rich manipulation tasks, including pushing, grasping, and multi-robot handover. Last, we find the proposed solver is simpler in form and capable of discovering unique contact modes when compared with existing solvers. Videos and code with examples are found in https://anon-website-submission.github.io/stein-admm-website/.
Chinese Translation
接触隐式轨迹优化将接触密集型操作建模为单个约束优化问题;然而,该单次程序运行只会收敛到众多同等有效的接触模式、抓取方式或推动方向中的一个局部最优解。因此,所得到的操作策略难以改变,且对初始化敏感。为了提升操作的鲁棒性,本文研究了接触隐式求解器如何发现多样化的接触密集型策略。我们的方法将共识交替方向乘子法(ADMM)的一种变体与Stein变分推断方法相结合,以输出一组互不相同的接触密集型解。我们发现,将Stein排斥力施加于ADMM的分裂变量(而非其原始形式变量),能够对可行接触策略集合实现有效覆盖,同时避免求解器过早停滞。我们在多种接触密集型操作任务上验证了该方法的有效性,包括推动、抓取以及多机器人交接。最后,与现有求解器相比,我们提出的求解器形式更简洁,并且能够发现独特的接触模式。视频及带示例的代码请见 https://anon-website-submission.github.io/stein-admm-website/。
cs.RO / 79 / 2609.28312

VGM-VS: Rethinking Visual Geometry Model for High-Precision Visual Servoing

VGM-VS:面向高精度视觉伺服的视觉几何模型再思考
Pan, Yimin, Wang, Sen, Zhou, You, Gao, Jianfeng, Sun, Pengbo, Naguib, Ahmed M., Marton, Zoltan-Csaba
Abstract
We present VGM-VS, a visual servoing method built on a pretrained feed-forward visual geometry model. Given the current view and a reference image captured at the target configuration, we estimate the relative camera pose with a visual geometry model and apply it iteratively as the pose increment of a closed-loop pose-based visual servoing (PBVS) scheme. The geometry-aware representation acquired from large-scale pretraining keeps this estimate reliable when the target is occluded, weakly textured, or covers only a small part of the image. However, the scale ambiguity inherent to these models leaves the predicted translation defined up to an unknown scale, while the pose increment must be metric for robot control. We close this gap with a scene-specific metric adaptation: the robot autonomously records image--pose pairs along a predefined motion starting from the target pose, and we fine-tune the camera head on these data, jointly learning the hand--eye transform and thus removing the need for a dedicated calibration process. We evaluate our method on three real-world assembly tasks with demanding tolerances: USB-C cable picking, cable insertion, and RAM insertion. Running in real time at 30Hz, VGM-VS converges to submillimeter terminal accuracy on the cable tasks, and reaches success rates of 90--100\% when the target is moved during servoing. It converges in all trials under initial displacements of up to 30cm from the reference pose and with 50\% of the target object occluded, outperforming the compared visual servoing baselines.
Chinese Translation
我们提出了VGM-VS,一种基于预训练前馈视觉几何模型构建的视觉伺服方法。给定当前视角和在目标位形处采集的参考图像,我们利用视觉几何模型估计相对相机位姿,并将其迭代地作为基于位姿的闭环视觉伺服(PBVS)方案的位姿增量。从大规模预训练中获得的几何感知表示,使得该估计在目标被遮挡、弱纹理或仅占图像很小部分时仍保持可靠。然而,此类模型固有的尺度歧义使预测的平移量仅在未知尺度下确定,而位姿增量对于机器人控制必须是米制(度量)的。我们通过一种场景特定的米制自适应方法来弥合这一差距:机器人从目标位姿出发,沿预定义的运动自主记录图像—位姿数据对,我们在这些数据上对相机头进行微调,同时学习手眼变换,从而省去了专门的标定过程。我们在三个具有严苛公差要求的真实装配任务上评估了所提方法:USB-C线缆抓取、线缆插入和内存条(RAM)插入。VGM-VS以30Hz实时运行,在线缆任务上收敛至亚毫米级的终端精度,且当伺服过程中目标被移动时成功率可达90–100%。在初始偏移达30厘米或目标物体50%被遮挡的情况下,该方法在所有试验中均能收敛,优于所对比的视觉伺服基线方法。
cs.RO / 80 / 2609.28314

TANDEM: Task and Motion Planning with As-Needed Demonstrations for Efficient Vision-Language-Action Model Fine-tuning

TANDEM:基于按需示范的任务与运动规划,用于高效的视觉-语言-动作模型微调
Sahoo, Samrat, Ji, Liang, Silver, Tom, Huang, Yixuan
Abstract
Human teleoperators spend substantial time demonstrating behaviors that robots can already perform autonomously, limiting the scalability of data collection for robot foundation models. Task and motion planning (TAMP) can automate many of these behaviors, but a fixed planning domain may not support every stage of a long-horizon manipulation task. We present TANDEM (Tamp with As-Needed Demonstrations for Efficient Model fine-tuning), a system that combines TAMP with selective human teleoperation to collect demonstrations for tasks beyond the planner's capabilities. Our key idea is to represent human assistance as an on-demand planning capability. Given a language instruction and visual observation, TANDEM uses pretrained vision-language models to extend the planning domain with missing predicates and human-executed magic operators. This allows the planner to interleave autonomous and human-executed stages without task-specific intervention points. After each human stage, TANDEM re-perceives the scene and checks whether the intended effects hold before resuming autonomous planning. To support fine-tuning vision-language-action (VLA) models, TANDEM also uses example pretraining trajectories to align planner-generated motions with the target model's pretraining distribution. We evaluate TANDEM on five long-horizon manipulation tasks beyond the TAMP domain's capabilities. On a representative long-horizon task, TANDEM collects 2.9x as many demonstrations as full-task teleoperation at the same human intervention time. Fine-tuning a pretrained \pi_{0.5}-DROID model on 20 TANDEM demonstrations per task increases average task success from 0% to 60% across the five tasks.
Chinese Translation
人类遥操作者需要花费大量时间为机器人已能自主执行的行为进行示范,这限制了机器人基础模型数据收集的可扩展性。任务与运动规划(TAMP)可以自动化其中许多行为,但固定的规划域可能无法支持长时程操作任务的每个阶段。我们提出了TANDEM(Tamp with As-Needed Demonstrations for Efficient Model fine-tuning),该系统将TAMP与选择性人类遥操作相结合,以收集超出规划器能力范围的任务示范。我们的核心思想是将人类辅助表示为一种按需的规划能力。给定语言指令和视觉观测,TANDEM利用预训练的视觉-语言模型扩展规划域,添加缺失的谓词以及由人类执行的“魔法算子”。这使得规划器能够交替执行自主阶段和人类执行阶段,而无需针对特定任务设置干预点。在每个人类阶段结束后,TANDEM会重新感知场景,并在恢复自主规划之前检查预期的效果是否成立。为支持视觉-语言-动作(VLA)模型的微调,TANDEM还利用示例预训练轨迹,将规划器生成的动作与目标模型的预训练分布对齐。我们在五个超出TAMP域能力范围的长时程操作任务上评估了TANDEM。在一个具有代表性的长时程任务上,在相同的人类干预时间下,TANDEM收集到的示范数量是全任务遥操作的2.9倍。在每任务20条TANDEM示范上微调预训练的\pi_{0.5}-DROID模型后,五个任务的平均任务成功率从0%提升至60%。
cs.RO / 81 / 2609.28317

Multimodal Voice Activity Projection for Social Robot Mediation: Expected Behavior and Deployment Constraints

面向社交机器人调解的多模态语音活动预测:预期行为与部署约束
Cano, Antonio, Perez, Guillermo, Merino, Luis, Gomez, Randy
Abstract
Turn-taking prediction is especially relevant for social robots that act as mediators in human-human interaction, where the expected action is often not to speak, but to orient, wait, avoid interruption, or prepare a balanced intervention. This paper presents Multimodal Voice Activity Projection (MM-VAP) as a human state-aware perception layer for future robot mediation behavior. The model estimates the future evolution of the conversational floor from synchronized audio-visual evidence and derives turn-taking events such as Hold, Shift, Shift prediction, Backchannel prediction, and overlap-related states. The approach uses VA-related pretrained audio-visual encoders, LoRA adaptation, inter-speaker attention, and zero-shot event inference from future voice activity projections. Experiments on NoXi, NoXi+J, and Haru EDR support the feasibility of this formulation, especially for floor management events that can be connected to gaze preparation, active listening, and conservative intervention. Finally, the paper defines the expected robot output interface and discusses the main deployment constraints, including real-time inference, preprocessing latency, multimodal synchronization, and input-quality monitoring.
Chinese Translation
轮次转换预测对于在人机交互(human-human interaction)中充当调解者的社交机器人尤为重要,因为其预期动作往往不是说话,而是转向、等待、避免打断或准备均衡的介入。本文提出多模态语音活动预测(Multimodal Voice Activity Projection, MM-VAP),作为未来机器人调解行为的人类状态感知层。该模型从同步的音视频证据中估计对话话语权的未来演化,并推导出保持(Hold)、转换(Shift)、转换预测(Shift prediction)、反馈预测(Backchannel prediction)以及重叠相关状态等轮次转换事件。该方法使用与语音活动(VA)相关的预训练音视频编码器、LoRA 适配、说话人间注意力机制,以及基于未来语音活动投影的零样本事件推断。在 NoXi、NoXi+J 和 Haru EDR 数据集上的实验支持了该表述的可行性,尤其是对于可与目光准备、主动倾听和保守介入相关联的话语权管理事件。最后,本文定义了预期的机器人输出接口,并讨论了主要部署约束,包括实时推理、预处理延迟、多模态同步以及输入质量监控。
cs.RO / 82 / 2609.28325

Motoneuron-Inspired Sampling for Model Predictive Path Integral Control

受运动神经元启发的模型预测路径积分控制采样方法
Poignant, Alexis, Babič, Jan
Abstract
Model Predictive Path Integral (MPPI) control relies on stochastic trajectory sampling, and its performance under limited rollout budgets depends strongly on the structure of the proposal distribution. Standard implementations commonly perturb control sequences with Gaussian noise, despite growing evidence that temporally correlated and structured sampling can improve finite-budget control. We introduce Spike-MPPI, a motoneuron-inspired proposal that generates temporally structured perturbations through a simplified model of motoneuron dynamics. The proposal is evaluated within a common MPPI framework on torque-actuated and antagonistically actuated MuJoCo Ant models against standard Gaussian sampling and spectrum-matched Gaussian controls. Results show that structured sampling substantially improves executed-control smoothness, while its effect on task performance depends on rollout condition and robot actuation. Spectrum matching reproduces a substantial part of the observed behavior, while the full Spike proposal retains additional effects beyond second-order spectral structure. These results support treating proposal design as a combination of second-order spectral structure and higher-order statistical organization.
Chinese Translation
模型预测路径积分(MPPI)控制依赖于随机轨迹采样,其在有限滚动预算下的性能在很大程度上取决于提议分布的结构。标准实现通常采用高斯噪声对控制序列进行扰动,然而越来越多的证据表明,具有时间相关性和结构化的采样可以改善有限预算下的控制效果。我们提出了Spike-MPPI,这是一种受运动神经元启发的提议方法,通过简化的运动神经元动力学模型生成具有时间结构的扰动。该方法在统一的MPPI框架下,于力矩驱动和拮抗驱动两种MuJoCo蚂蚁(Ant)模型上进行评估,并与标准高斯采样以及频谱匹配的高斯控制进行比较。结果表明,结构化采样显著提升了执行控制的平滑性,而其对任务性能的影响则取决于滚动条件和机器人驱动方式。频谱匹配再现了所观察行为的大部分特征,而完整的Spike提议方法仍保留了超越二阶频谱结构的额外效应。这些结果支持将提议分布设计视为二阶频谱结构与高阶统计组织的组合。
cs.RO / 83 / 2609.28339

Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control

超越未来预测:作为机器人控制生成式适配的去噪方法
Wang, Zanyi, Lei, Yuheng, Jiang, Dengyang, Luo, Ping, Wang, Mengdi, Liang, Zhixuan, Liu, Shilong
Abstract
Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned structure through large-scale image and video generation training. A growing line of robot policies builds on this generative prior, but how it should be transferred to control remains unclear, and existing approaches commonly instantiate this transfer through future visual prediction. We ask a more basic question: what a pretrained generative DiT actually contributes to action learning, and how this prior should be adapted for control. We introduce NowWAM, a future-target-free co-training formulation that denoises the current observation and predicts robot actions from the same visual stream, directly coupling the native generative objective to the action-facing representation across the denoising trajectory. Under matched controlled settings, past and future visual targets perform comparably, while restricting training to the clean endpoint substantially reduces robustness, suggesting that a separate future target is not essential for generative adaptation, while the denoising trajectory remains an effective interface for control. On LIBERO-Plus, NowWAM reaches 87.7% with FLUX2-Klein, improving over the future-target co-training baseline by 6.1 points while halving training visual tokens (784 to 392) and reducing step time from 2.85 s to 1.63 s, a 1.8x speedup. With the pure text-to-image Z-Image backbone, NowWAM further reaches 87.8%, showing that strong control adaptation is not tied to video generation or image-editing backbones.
Chinese Translation
预训练的生成式扩散Transformer(DiT)通过大规模图像与视频生成训练,捕获了丰富的像素级视觉结构和语言条件化结构。越来越多的机器人策略构建在这一生成式先验之上,但如何将其迁移到控制任务仍不清楚,且现有方法通常通过未来视觉预测来实现这种迁移。我们提出了一个更基本的问题:预训练的生成式DiT究竟为动作学习贡献了什么,以及这一先验应如何被适配到控制中。我们提出NowWAM,一种无需未来目标的协同训练方法,其对当前观测进行去噪,并从同一视觉流预测机器人动作,将原生的生成式目标与面向动作的表征在去噪轨迹上直接耦合。在匹配的受控设置下,过去与未来视觉目标表现相当,而将训练限制在干净端点会显著降低鲁棒性。这表明单独的未来目标对生成式适配并非必需,而去噪轨迹仍然是控制的有效接口。在LIBERO-Plus上,NowWAM使用FLUX2-Klein达到87.7%的成功率,比未来目标协同训练基线提升6.1个百分点,同时将训练视觉token减半(从784降至392),步时间从2.85秒降至1.63秒,实现1.8倍加速。采用纯文生图骨干Z-Image,NowWAM进一步提升至87.8%,表明强大的控制适配能力并不依赖于视频生成或图像编辑骨干。
cs.RO / 84 / 2609.28364

LEAP-CBF: A Safety Filter for Uncertain Systems with Least-Effort Adversarial Potentials

LEAP-CBF:一种面向不确定系统的最小努力对抗势安全过滤器
So, Oswin, Yu, Eric, Fan, Chuchu
Abstract
Control barrier functions (CBF) are a popular safety filter to ensure safety for nonlinear dynamical systems. However, when the system is subject to uncertainties and disturbances, this requires the use of robust variants of CBFs, which can be difficult to construct and can be overly conservative, especially for high-dimensional systems under input constraints. In this work, we propose a new approach to solve these challenges by introducing Least-Effort Adversarial Potentials (LEAP), a certificate that quantifies the robustness of a given state against disturbances in terms of the effort required by the disturbance to cause failure. We show that LEAP is a CBF for the undisturbed system, but can also be used to construct a safety filter that is robust to disturbances whose cumulative effort is bounded. We propose a method for constructing LEAPs with on-policy deep reinforcement learning. Next, we demonstrate LEAPs in simulation on a variety of multi-agent systems with disturbances and uncertainties. Finally, hardware experiments on a quadruped and quadrotors validate that LEAPs are well suited to tackle the disturbances and uncertainties from real-world robotic systems.
Chinese Translation
控制屏障函数(Control Barrier Functions, CBF)是一种流行的安全过滤器,用于保证非线性动力学系统的安全性。然而,当系统受到不确定性和扰动影响时,需要使用鲁棒变体的CBF,这类变体难以构造,且可能过于保守,尤其是在输入约束下的高维系统中。在本工作中,我们提出了一种解决这些挑战的新方法,即引入最小努力对抗势(Least-Effort Adversarial Potentials, LEAP),这是一种证书,用以量化给定状态对扰动的鲁棒性,其衡量方式为扰动导致系统失效所需付出的努力程度。我们证明LEAP是无扰动系统的一个CBF,同时也可用于构建对累积努力有界的扰动具有鲁棒性的安全过滤器。我们提出了一种基于在线策略(on-policy)深度强化学习构建LEAP的方法。随后,我们在多种存在扰动和不确定性的多智能体系统的仿真中展示了LEAP的效果。最后,在四足机器人和四旋翼无人机上的硬件实验验证了LEAP非常适合应对来自真实世界机器人系统的扰动和不确定性。
cs.RO / 85 / 2609.28377

Amplify: A Lightweight Library for Reproducible Nonlinear Programming Problems in Robotics

Amplify:一个用于机器人学中可复现非线性规划问题的轻量级库
Rosa Jr, Nelson
Abstract
Optimization problems (OPs) are key to solving many challenging research problems in robotics. However, reproducibility still remains a major issue. In this paper, we present Amplify, a lightweight nonlinear programming library aimed at reproducible results of robotic-related trajectory optimization problems. The minimalistic requirements for the 537-line library (80 characters per line) are an Internet connection, familiarity with the AMPL modeling language, and a text editor. Our primary contribution is the formulation of a library where trajectory optimization algorithms are represented directly within the optimization model. Specifically, we implement the algorithms used to compute the dynamics, trajectories, and reference motions as constraints of the OP in a declarative programming paradigm. We outline how our formulation of objectives, decisions variables, and constraints can be implemented in other transcription libraries that want to be lightweight and reproducible. We also compare the Amplify framework with 3 other libraries across examples of benchmark optimization problems across several fields, including bipedal locomotion and grasp planning.
Chinese Translation
优化问题(OP)是解决机器人学中许多具有挑战性的研究问题的关键。然而,可复现性仍然是一个主要难题。在本文中,我们提出了Amplify,这是一个轻量级的非线性规划库,旨在实现机器人相关轨迹优化问题的可复现结果。这个仅有537行代码的库(每行80个字符)的最低要求是:能够连接互联网、熟悉AMPL建模语言,以及一个文本编辑器。我们的主要贡献是构建了一个将轨迹优化算法直接表示在优化模型中的库。具体而言,我们以声明式编程范式将用于计算动力学、轨迹和参考运动的算法实现为优化问题的约束。我们概述了如何将我们提出的目标函数、决策变量和约束的构建方式,应用于其他希望做到轻量化和可复现的转换(transcription)库中。我们还在包括双足运动和抓取规划在内的多个领域的基准优化问题示例上,将Amplify框架与其他3个库进行了比较。
cs.RO / 86 / 2609.28378

ForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid Control

ForgetMimic:面向强化学习人形机器人控制的动作遗忘方法
Luan, Xukun, Lei, Zhongxiang, Gong, Chen, Li, Shaowei, Bi, Yuanguo, Liu, Jinyan
Abstract
Humanoid control, leveraging human demonstrations, has achieved diverse, agile, and natural locomotion behaviors through reinforcement learning (RL). While this paradigm has yielded remarkable performance in physical humanoid control, how to eliminate specific motions from learned policies remains insufficiently explored. Addressing this issue is motivated by pressing safety and privacy concerns: the removal of malicious, poisoned, or suboptimal motions, as well as copyright-protected motions subject to the right to be forgotten under regulations such as the GDPR, is of critical importance. To this end, we propose {ForgetMimic}, the first motion-level unlearning method designed specifically for physical-world humanoid control. The core idea of ForgetMimic is as follows: given a policy $\pi_\theta$ trained on $N$ motions, our method degrades performance on a target subset of $K$ motions while preserving the effectiveness of the remaining $N-K$ motions. Furthermore, we identify and resolve two key training mechanisms in robot control that lead to unlearning failure. We conduct extensive experiments on the Unitree G1 and H2 humanoid robots across 12 motions, including Dance, Fight, Flip, and others. Experimental results demonstrate that ForgetMimic effectively eliminates memory of designated motions while maintaining the normal operation of all other motions.
Chinese Translation
基于人类示范的人形机器人控制已通过强化学习(RL)实现了多样化、敏捷且自然的运动行为。尽管这一范式在物理人形机器人控制中取得了卓越性能,但如何从已学习的策略中消除特定动作仍未得到充分探索。解决这一问题源于紧迫的安全与隐私需求:移除恶意、被污染或次优的动作,以及依据《通用数据保护条例》(GDPR)等法规中'被遗忘权'而受版权保护的动作,至关重要。为此,我们提出{ForgetMimic},这是首个专为物理世界人形机器人控制设计的动作级遗忘(unlearning)方法。ForgetMimic的核心思想如下:给定一个在 $N$ 个动作上训练的策略 $\pi_\theta$,我们的方法在包含 $K$ 个动作的目标子集上降低其性能,同时保持其余 $N-K$ 个动作的有效性。此外,我们识别并解决了机器人控制中导致遗忘失败的两个关键训练机制。我们在Unitree G1和H2人形机器人上开展了涵盖舞蹈、格斗、翻转等12个动作的大规模实验。实验结果表明,ForgetMimic能够有效消除对指定动作的记忆,同时保持所有其他动作的正常运行。
cs.RO / 87 / 2609.28393

PointCast: One World Model for Rigid, Articulated, and Deformable Object Manipulation

PointCast:面向刚体、铰接与可变形物体操作的统一世界模型
Ye, Hantao, Worobel, Ross, Xie, Zhuoli, Li, Mingen, Yu, Houjian, Hong, Youngjin, Choi, Changhyun
Abstract
World models are useful for robotic manipulation because robots can predict how actions change the states of objects before executing them. We present PointCast, a point-set world model that spans rigid, articulated, and deformable object manipulation. Its state is a set of 3D points on the object and the end-effector, mesh-free and topology-agnostic. Each point keeps its identity and is supervised on its own trajectory, which teaches the model where every point goes rather than only the shape the points form. Its backbone is a diffusion transformer that denoises a short window of future point positions, conditioned on the points' recent history and the commanded end-effector motion. The backbone's attention alternates between local and global, and cross-attention to the end-effector carries the coupling. This one architecture at 19.8M parameters and one training recipe cover four regimes, rigid objects, cloth, rope, and multi-joint cabinets, with a separate checkpoint trained for each. Trained on randomized simulation and scored against four baselines on the same metric, it is best on three of four regimes and second on rigid. Trained on a real-world robot teleoperation dataset, it has the lowest mean error in four of its six categories, is second in the other two, and improves on the dataset's own model in all six; zero-shot, its simulation checkpoints are best on two of four captures. Frozen inside sampling-based model-predictive control at one network evaluation per window, it plans four simulated tasks over 64 episodes, competitive with or outperforming every baseline on each. Project website at https://pointcast-wm.github.io.
Chinese Translation
世界模型对机器人操作非常有用,因为机器人可以在执行动作之前预测动作如何改变物体的状态。我们提出了 PointCast,一个涵盖刚体、铰接和可变形物体操作的点集世界模型。其状态为物体和末端执行器上的一组三维点,无需网格且与拓扑无关。每个点保持自身身份,并在其各自轨迹上进行监督,从而教会模型每个点的去向,而不仅仅是这些点所构成的形状。其骨干网络是一个扩散Transformer,对未来一小段时间窗口内的点位置进行去噪,并以这些点的近期历史和指令的末端执行器运动为条件。骨干网络的注意力机制在局部与全局之间交替进行,并通过对末端执行器的交叉注意力来传递两者之间的耦合。这一统一架构仅有19.8M参数,采用单一训练方案即可覆盖四种情形——刚体物体、布料、绳索和多关节橱柜,每种情形各训练一个独立检查点。在随机化仿真数据上训练,并使用相同指标与四个基线方法对比,它在四种情形中的三种取得最佳表现,在刚体情形中排名第二。在真实世界机器人遥操作数据集上训练后,它在六个类别中的四个取得最低平均误差,在其余两个类别中排名第二,且在所有六个类别上均优于数据集自带的模型;在零样本设置下,其仿真检查点在四个采集场景中的两个上表现最佳。将该模型冻结并嵌入基于采样的模型预测控制中,每个时间窗口仅需一次网络评估,即可在64个回合上规划四个仿真任务,在每个任务上与所有基线相当或更优。项目网站:https://pointcast-wm.github.io。
cs.RO / 88 / 2609.28396

Tractable Reinforcement Learning for Full Class of Signal Temporal Logic Specifications Using Spatiotemporal Tube Reward

基于时空管道奖励的信号时序逻辑全类规格可处理强化学习
Jagabathula, Vaishnavi, Sangeerth, P, Jagtap, Pushpak
Abstract
This paper addresses the control problem for robotic systems, including non-holonomic and underactuated platforms operating under unknown dynamics and strict actuator limits to satisfy complex high-level specifications. We denote these high-level specifications using Signal Temporal Logic (STL) and propose a novel time-aware Reinforcement Learning (RL) framework that leverages the geometric properties of Spatiotemporal Tubes (STTs). While traditional analytical STT controllers often struggle to enforce input constraints, and existing RL approaches rely on memory-intensive state history, our method natively overcomes both limitations. By mapping the logical and temporal complexities of the full class of STL into time-varying geometric boundaries, we directly constrain the multidimensional system state without relying on scalar robustness metrics. Augmenting the state space with time, we train a time-aware Soft Actor-Critic (SAC) agent using a continuous, geometry-aware reward function that eliminates the need to explicitly evaluate complex logical semantics during execution. The proposed framework offers a history-free, computationally efficient approach to learn continuous control policies that ensure robust satisfaction of specifications while strictly adhering to system input constraints.
Chinese Translation
本文研究了机器人系统的控制问题,包括在未知动力学和严格执行器约束下运行的非完整(non-holonomic)与欠驱动平台,使其满足复杂的高层任务规格。我们使用信号时序逻辑(Signal Temporal Logic, STL)来描述这些高层规格,并提出一种新颖的具有时间感知能力的强化学习(RL)框架,该框架利用了时空管道(Spatiotemporal Tubes, STT)的几何特性。传统的解析式STT控制器通常难以满足输入约束,而现有的RL方法依赖于占用大量内存的状态历史,与此不同,我们的方法从本质上克服了这两个局限。通过将STL全类规格的逻辑与时序复杂性映射为时变几何边界,我们直接对多维系统状态施加约束,而无需依赖标量鲁棒性度量。通过将时间增广到状态空间中,我们训练了一个时间感知的Soft Actor-Critic(SAC)智能体,其连续的、几何感知的奖励函数消除了在执行过程中显式评估复杂逻辑语义的需要。所提出的框架提供了一种无需历史信息且计算高效的方法来学习连续控制策略,从而在严格遵守系统输入约束的同时,确保对任务规格的鲁棒满足。
cs.RO / 89 / 2609.28429

Watch, Recall, Act: Always-On Robots in Concurrent Embodied Streams

观察、回忆、行动:并发具身数据流中的常开机器人
Yi, Ding, Sun, Peiwen, Rong, Chenchu, Wang, Jianan, Dai, Xili, Yue, Xiangyu, Lin, Xi
Abstract
An always-on robot faces an endless stream that never resets: instructions arrive and lapse, the scene changes, and its own past actions reshape what it must reason about. Today's action models are built for the opposite: a fixed instruction, no mid-task intervention, single-step reasoning. In an open-ended world a robot must watch a live stream for far-future cues, recall its own far-past actions, and act on them under dual-arm concurrency. We present ARMS (Always-on Robot in Multi-modal Streams), a deliberately simple streaming policy: a single pretrained $\pi$0.5 backbone augmented by three lightweight modules that turn live perception, embodied states, and the robot's own past actions into context the backbone reads before it acts. The modules update this context asynchronously, so watching and recalling never block acting and the two arms act at once. Rather than inventing new mechanisms, ARMS integrates these learned context providers with an agent-causal self-history that logs which arm did what, and when. To supervise them without extra annotation, we build ARMS Dataset, whose staged construction script itself labels every module from real dual-arm teleoperation. Trained on it, ARMS reaches 45% on the combined task against 28% for the strongest of our four main baselines, and ablations confirm the memory module, the embodied-state head, and asynchronous concurrency are each necessary.
Chinese Translation
一个常开(always-on)机器人面对的是永不停息、从不重置的数据流:指令会到来也会失效,场景不断变化,而机器人自身过去的动作也在持续重塑它需要推理的内容。当今的动作模型却恰恰相反:它们面向固定指令、任务中途无干预、单步推理而构建。在一个开放式的世界中,机器人必须观察实时数据流以捕捉远未来的线索,回忆自己远过去的动作,并在双臂并发条件下据此行动。我们提出了 ARMS(Always-on Robot in Multi-modal Streams,多模态数据流中的常开机器人),这是一种刻意保持简洁的流式策略:以单个预训练的 $\pi$0.5 主干为基础,辅以三个轻量级模块,将实时感知、具身状态以及机器人自身过去的动作转换为主干在行动前读取的上下文。这些模块异步更新上下文,因此观察与回忆从不会阻塞行动,双臂也能同时行动。ARMS 并未发明新机制,而是将这些学习到的上下文提供者与一种记录“哪只手臂做了什么、何时做的”智能体因果自历史(agent-causal self-history)相结合。为了无需额外标注即可对其进行监督,我们构建了 ARMS Dataset,其分阶段的构建脚本本身即可从真实双臂遥操作中为每个模块生成标签。在该数据集上训练后,ARMS 在组合任务上达到 45%,而我们四个主要基线中最强者的成绩仅为 28%;消融实验证实记忆模块、具身状态头以及异步并发均不可或缺。
cs.RO / 90 / 2609.28431

LiMA: Bridging Long-term Imagination to Real-time Dexterous Manipulation via Asynchronous Diffusion

LiMA:通过异步扩散将长期想象与实时灵巧操作相连接
Chen, Ning, Fu, Yankai, Zhao, Junkai, Sun, Qianpu, Yao, Guocai, Wang, Pengwei, Wang, Zhongyuan, Zhang, Shanghang
Abstract
Dexterous manipulation demands long-term foresight and rapid reactive control. Vision-Language-Action (VLA) models, while proficient in high-level reasoning, often lack a fine-grained understanding of physical dynamics and spatial perception. Conversely, World-Action Models (WAMs) typically suffer from high inference latency due to iterative generation. These deficiencies result in a critical temporal misalignment where the model's intent fails to adapt to rapid physical contact changes. To overcome this fundamental bottleneck, we propose LiMA, an asynchronous dual-system generative framework that systematically decouples intent planning from reactive execution. LiMA organizes computation into a multi-scale hierarchy: a slow system handles sparse long-horizon spatiotemporal intent generation, while a fast system focuses on dense high-frequency motion refinement. To align sparse intent predictions with dense action trajectories, we introduce a Latent Schr\"odinger Bridge Coupling mechanism that formulates refinement as an entropy-regularized probabilistic transport process. LiMA reduces inference latency by 45.8% compared with Cosmos-Policy via asynchronous decoupling. Evaluated across six bimanual dexterous manipulation tasks spanning multiple horizons, LiMA achieves an overall success rate of 70.8% and an average subtask success rate of 78.9%, while maintaining performance in unseen scenarios. The project website is available at https://ccdcs.github.io/LiMA_repo/
Chinese Translation
灵巧操作既需要长期预见能力,也需要快速的响应式控制。视觉-语言-动作(VLA)模型虽然擅长高层推理,但往往缺乏对物理动力学和空间感知的细粒度理解。相反,世界-动作模型(WAM)通常由于迭代式生成而存在较高的推理延迟。这些缺陷导致了一个关键的时间失配问题,即模型的意图无法适应快速的物理接触变化。为克服这一根本性瓶颈,我们提出了LiMA,这是一种异步双系统生成框架,系统性地将意图规划与响应式执行解耦。LiMA将计算组织为多尺度层级结构:慢系统负责稀疏的长时程时空意图生成,而快系统专注于稠密的高频运动细化。为使稀疏的意图预测与稠密的动作轨迹对齐,我们引入了一种潜在Schrödinger桥耦合机制,将细化过程表述为一个熵正则化的概率传输过程。通过异步解耦,LiMA相较于Cosmos-Policy将推理延迟降低了45.8%。在涵盖多个时间跨度的六项双臂灵巧操作任务上的评估中,LiMA实现了70.8%的总体成功率和78.9%的平均子任务成功率,同时在未见过的场景中仍保持良好性能。项目网站见 https://ccdcs.github.io/LiMA_repo/
cs.RO / 91 / 2609.28467

Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction

我应该在哪里加入?基于语言引导目标预测的机器人群体加入
Fang, Zilin, Wang, Zishuo, Lee, Gim Hee, Hsu, David
Abstract
Social navigation typically assumes a specified goal and focuses on reaching it while respecting social conventions, whereas robot group joining requires predicting where to join based on the group's real-time activity and formation. This is a highly semantic task, yet an important capability for applications such as robotic guide dogs and autonomous mobility scooters. We formulate language-grounded robot group joining: given an observation and a natural-language description of a target group, the robot identifies the relevant group members and predicts socially compliant joining poses. For grounding, we generate structured candidate subsets through recursive spectral partitioning and rank them with a language-conditioned image--geometry model. Given the grounded group, a goal predictor leverages human-formation priors to produce a multimodal energy--orientation map over feasible robot poses. Experiments on conversations, queues, and audiences across varying group sizes, crowd densities, and visual ambiguities show that our method achieves competitive grounding accuracy with sub-second inference and outperforms all baselines in joining-pose prediction. Real-robot experiments further demonstrate group joining in both static and dynamically changing interactions.
Chinese Translation
社会导航通常假设目标已给定,侧重于在遵守社会规范的前提下到达该目标;而机器人群体加入则需要根据群体的实时活动和队形来预测加入位置。这是一项高度语义化的任务,也是机器导盲犬和自动轮椅等应用所需的重要能力。我们提出了基于语言的机器人群体加入问题:给定一次观测和目标群体的自然语言描述,机器人需要识别相关的群体成员并预测符合社会规范的加入位姿。在语言定位(grounding)方面,我们通过递归谱划分生成结构化的候选子集,并使用语言条件化的图像—几何模型对其进行排序。在确定目标群体后,目标预测器利用人类队形先验,在可行机器人位姿上生成多模态的能量—朝向图。在对话、排队和观众等场景中,针对不同群体规模、人群密度和视觉歧义的实验表明,我们的方法在实现具有竞争力的定位精度的同时,推理时间低于一秒,并在加入位姿预测上优于所有基线方法。真实机器人实验进一步验证了在静态和动态变化的交互中进行群体加入的能力。
人工智能 (Artificial Intelligence)
59
cs.AI / 1 / 2609.26836

Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverse

智能体-工具交互中的静默失败:对 ToolUniverse 的审计研究
Gopalan, Shreya, Singh, Devansh, Narayanan, Sundaraparipurnan
Abstract
Agentic AI systems are increasingly adopting automated pipelines that integrate multiple tools. While prior research and benchmarks have studied about task success and task completion of these agentic systems, the research about agent to tool interaction, specifically in biology agentic workflow is limited. This study investigates specific failures in agent to tool interaction where a tool invocation appears successful, some or all of the information or functionality from the tool via API/ wrapper is incomplete or missing and there are no communications / notifications to the user or the agent about such missing information. We call this a silent failures as the user or the agents are not aware that such failure has occurred. For the purposes of this study we developed an audit mechanism to identify such silent failures in Agent to tool interaction, by examining 15 scientific tools (and their associated API documentation and tool documentations) integrated within ToolUniverse environment (ToolUniverse serves as our experimental environment rather than the object of the study itself). We structure our study around 7 failure locus characterising where the failure occurs in the chain. We observed 91 failures (manually validated post LLM based candidate discovery and automated testing), most frequent of them being missing data or fields and inconsistencies in search, filtering or ranking criteria. Most of the 91 failures occurred in API layer (51) or wrapper layer (25), with a potential of silent failure amplification downstream. The results show that silent failures originate upstream of the event and propagate downstream into apparently valid scientific outputs. We propose a concept of contextual reliability to handle such failures and suggest mechanisms for testing, disclosing, monitoring, and measuring such failures across the agent-tool interaction pipeline.
Chinese Translation
智能体 AI 系统正日益采用集成多种工具的自动化流水线。尽管已有研究和基准测试考察了这些智能体系统的任务成功与任务完成情况,但关于智能体与工具交互的研究仍然有限,尤其是在生物学智能体工作流方面。本研究探讨了智能体与工具交互中的特定失败情形:工具调用表面上成功了,但通过 API/包装器(wrapper)获取的部分或全部信息或功能不完整或缺失,且没有任何通信或通知告知用户或智能体此类缺失。我们将其称为"静默失败"(silent failures),因为用户或智能体并未意识到此类失败已经发生。为开展本研究,我们开发了一种审计机制来识别智能体与工具交互中的此类静默失败,方法是审查集成在 ToolUniverse 环境中的 15 个科学工具(及其相关的 API 文档和工具文档)(ToolUniverse 作为我们的实验环境,而非研究对象本身)。我们围绕 7 个失败位点构建研究框架,以刻画失败在链条中发生的位置。我们共观察到 91 次失败(经基于大语言模型的候选发现和自动化测试之后的人工验证),其中最常见的是数据或字段缺失,以及搜索、过滤或排序标准的不一致。这 91 次失败大多发生在 API 层(51 次)或包装器层(25 次),并存在静默失败向下游放大的可能。结果表明,静默失败起源于事件的上游,并向下游传播,最终产生表面上有效的科学输出。我们提出了"情境可靠性"(contextual reliability)的概念来应对此类失败,并建议在整个智能体-工具交互流水线中建立用于测试、披露、监测和度量此类失败的机制。
cs.AI / 2 / 2609.26891

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

以 Harness 作为语言:一个极简但表达力极大的智能体框架
Li, Zhening, Liu, Joshua, Vukelic, Mateja, Shen, Nicole, Lall, Supriya, Thakur, Amitayush, Zhang, Alex, Khattab, Omar, Light, Jonathan, Solar-Lezama, Armando
Abstract
Modern language-model agents are built around the \textit{agent loop}, where the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain workflows currently require additional engineering beyond the agent loop itself, such as memory systems and self-improving systems. We built an LLM agent framework, JAZ, to explore the extent to which a minimal harness that is little more than the agent loop itself can accomplish tasks these specialized systems are built for. JAZ exposes a single LLM-based primitive invoke and provides a set of built-in hooks that allow the programmer to apply constraints and monitoring. Generalizing existing code-mode agent loops, \texttt{invoke} is the simplest loop that satisfies two defining properties: (1) the LLM can write arbitrary executable code that can include recursive \texttt{invoke}; (2) everything visible to the LLM --- all inputs to \texttt{invoke} as well as its interaction history with the code environment --- are variables in the code environment. We motivate our design from first principles, viewing \texttt{invoke} as a language primitive representing a function whose implementation is provided at runtime by an LLM every time it is called. To validate the design of our core \texttt{invoke} primitive, we evaluate \texttt{invoke} --- with only prompting, no manually designed tools, harness, or external systems (e.g., memory or the file system) --- on workflows traditionally implemented through specialized external harnesses. On long-horizon workflows requiring recall beyond the context window, JAZ invoke outperforms Letta (MemGPT) by 8\% at half its cost on the recall-heavy portion of StuLife. On continual self-improvement, JAZ invoke outperforms ACE by 4\% at a lower cost on AppWorld.
Chinese Translation
现代语言模型智能体围绕“智能体循环”构建:将大语言模型置于一个暴露一组工具的环境中,LLM 通过在工具调用与观察输出之间交替进行,从而完全掌控工作流。然而,某些工作流目前需要超越智能体循环本身的额外工程,例如记忆系统和自我改进系统。我们构建了一个 LLM 智能体框架 JAZ,以探索一个几乎只是智能体循环本身的极简 harness 能在多大程度上完成这些专门系统所设计用于的任务。JAZ 仅暴露一个基于 LLM 的原语 invoke,并提供一组内置钩子,使程序员能够施加约束和监控。作为对现有 code-mode 智能体循环的泛化,invoke 是满足以下两个定义性属性的最简循环:(1) LLM 可以编写任意可执行代码,其中可包含递归的 invoke;(2) LLM 可见的一切——invoke 的所有输入以及它与代码环境的交互历史——都是代码环境中的变量。我们从第一性原理出发阐述该设计,将 invoke 视为一种语言原语,表示一个函数,其实现由 LLM 在每次调用时于运行时提供。为验证核心 invoke 原语的设计,我们在传统上通过专门外部 harness 实现的工作流上评估 invoke——仅使用提示,无需人工设计的工具、harness 或外部系统(如记忆或文件系统)。在需要超出上下文窗口回忆能力的长时程工作流上,在 StuLife 中以回忆为主的部分,JAZ invoke 以一半的成本优于 Letta (MemGPT) 8%。在持续自我改进方面,JAZ invoke 在 AppWorld 上以更低的成本优于 ACE 4%。
cs.AI / 3 / 2609.26911

TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents

TwinCheck:面向有状态工具智能体的基于证据的负孪生验证方法
Dai, Jiaxuan, Huang, Tianyi
Abstract
A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers replacement only when the trace satisfies an evidence condition tied to a trace-local failure hypothesis. It constructs a trace-grounded counterfactual alternative, a negative twin, and replaces the agent's proposal only if the twin passes structural checks and the pairwise verifier prefers it in both candidate orders. For paired evaluation, exact replay holds the agent's parsed responses and actions fixed until the first accepted replacement, separating intervention effects from resampling. In the primary analysis of 159 multi-turn BFCL V4 tasks with complete exact-replay pairs, the complete policy raises task success for GPT-5.6 Sol from 45.3% to 58.5% (95% task-bootstrap CI [8.2, 18.8]), with no observed success-to-failure regressions. Together, these findings recast execution-boundary repair as a constrained comparison, making the counterfactual action itself the object of verification.
Chinese Translation
一个局部看似合理的工具调用可能会使原本成功的智能体轨迹偏离正轨。单纯的怀疑并不足以成为干预的理由,因为替换本身可能引入验证原本旨在防止的失败。我们提出了 TwinCheck,这是一种推理时验证策略,仅当轨迹满足与轨迹局部失败假设相关联的证据条件时才考虑替换。它构建一个基于轨迹的反事实备选方案(即负孪生,negative twin),并且仅当该孪生通过结构性检查、且成对验证器在两种候选顺序下都更偏好它时,才替换智能体的提议。在配对评估中,精确重放(exact replay)将智能体已解析的响应和动作固定,直到第一次被接受的替换为止,从而将干预效应与重采样区分开来。在对159个具有完整精确重放配对的多轮 BFCL V4 任务的主要分析中,完整策略将 GPT-5.6 Sol 的任务成功率从45.3%提升至58.5%(95%任务自助法置信区间 [8.2, 18.8]),且未观察到成功转为失败的退化。这些发现共同将执行边界的修复重新表述为一种受约束的比较,使反事实动作本身成为验证的对象。
cs.AI / 4 / 2609.26927

Building Socio-Affective Artificial Intelligence for Interactive Multi-Agent Simulations

构建面向交互式多智能体仿真的社会-情感人工智能
Berga, David
Abstract
The objective of this article is to provide design principles and a software architecture for enabling interaction between humans and multiple agents in simulated dynamic worlds. This connects the current era of general artificial intelligence (AI/AGI) with the proliferation of transformer-based conversational agents and the increased computational capabilities. Given an overview of current and previous multi-agent theories of mind (socially and affectively-aware agents), the existence of an integrative design of agent interactions with themselves and with humans must be crucial for understanding how to create sustainable and governance in future human-agent reasoning systems. In this work is presented a software "AGIMUD" that integrates: A. socially-aware reasoning and emotion in agent behavior and interaction, B. a design of human multimodal scheme for human users, artificial agents and simulated worlds, and C. distributing the AI processing through the network to enable multiple autonomous agents. These integrations allow the dynamic world recreation as multi-user dungeons (MUDs) where both agents and humans can interact simultaneously in real time. Find the code online in https://github.com/dberga/AGIMUD.
Chinese Translation
本文旨在提供设计原则和软件架构,以实现人类与多个智能体在模拟动态世界中的交互。这将当前通用人工智能(AI/AGI)时代与基于Transformer的对话式智能体的普及以及计算能力的提升联系起来。通过概述当前和以往的多智能体心智理论(具备社会与情感感知能力的智能体),智能体与自身及人类交互的一体化设计对于理解如何在未来人机推理系统中实现可持续性和治理至关重要。本文提出了一个名为"AGIMUD"的软件,它集成了:A. 智能体行为与交互中具备社会感知的推理与情感;B. 面向人类用户、人工智能体和模拟世界的人类多模态方案设计;C. 通过网络分布式处理AI任务,以支持多个自主智能体。这些集成使动态世界能够以多用户地下城(MUDs)的形式重建,智能体与人类可以实时同时交互。代码可在 https://github.com/dberga/AGIMUD 获取。
cs.AI / 5 / 2609.26929

Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment

哪些目标需要调节旋钮?在可引导的多元对齐中预测目标冲突并覆盖权衡空间
Tsoi, David, Dönmez, Esra
Abstract
People hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective Direct Preference Optimization (MODPO) does this by using an objective weight to span a continuum of trade-offs. We study two questions: when can one model improve two objectives simultaneously, and how can many trade-offs be covered without training a separate model for each? Across seven objective pairs from HelpSteer and UltraFeedback, two pre-training measurements predict whether objectives align or conflict for human-annotated data, but not for AI-annotated data, where response length and repetition confound reward-model scores. For broader trade-off coverage, selecting the nearest trained model and merging model parameters both help, but neither consistently matches direct training. These findings yield practical guidance for building steerable models that serve diverse preferences.
Chinese Translation
人们持有多样且有时相互冲突的价值观,因此没有任何单一的对齐模型能够满足所有人。多元对齐因此需要可引导的模型,能够以不同的方式平衡相互竞争的目标。多目标直接偏好优化(Multi-Objective Direct Preference Optimization, MODPO)通过使用目标权重来覆盖一个连续的权衡空间,从而实现这一目标。我们研究两个问题:何时可以用一个模型同时改进两个目标,以及如何在不必为每个权衡单独训练一个模型的情况下覆盖多种权衡?在来自 HelpSteer 和 UltraFeedback 的七个目标对上,两个预训练测量指标能够预测目标在人工标注数据上是协同一致还是相互冲突,但在 AI 标注数据上则无法做到,因为回答长度和重复度会混淆奖励模型的评分。为了实现更广泛的权衡覆盖,选择最近的已训练模型以及合并模型参数均有所帮助,但两者都不能持续达到直接训练的效果。这些发现为构建服务多元偏好的可引导模型提供了实用指导。
cs.AI / 6 / 2609.26952

Escaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency Resolution

逃离Python依赖地狱:一种用于Python依赖解析的混合重放-修复流水线
Poweska, Veronica, Oyanguren, Ariana, Pourleyli, Jessica, Khanzadeh, Sourena, Alalfi, Manar
Abstract
Dependency conflicts in Python ecosystems arise from incompatible version constraints, missing packages, and undocumented compatibility relationships, causing many real-world code snippets to fail at execution. This paper presents PLLM+, a hybrid dependency-repair pipeline evaluated on the HG2.9K benchmark of 2,891 dependency-failing snippets. PLLM+ prioritizes inexpensive deterministic steps before invoking LLM-based repair: static AST-based interpreter inference, replay of historically successful dependency configurations from the competition-provided solutions database, and live PyPI validation of candidate package versions. When these steps do not resolve a case, the system falls back to a structured LLM-based repair loop with typed error classification and Proposer/Critic agents. On HG2.9K, PLLM+ solves 1,500 out of 2,891 snippets, compared with 1,169 solved by the PLLM baseline. It also reduces average runtime from 368.7 to 71.8 seconds per snippet. Most successful fixes come from replaying known configurations: 1,495 of the 1,500 successful fixes are produced by the solutions database, while the LLM fallback accounts for 5 additional fixes. These results suggest that, in this benchmark setting, deterministic reuse of previously validated dependency configurations is a simple and effective strategy, with LLM-based repair serving as a secondary fallback for cases not covered by prior solutions.
Chinese Translation
Python生态系统中的依赖冲突源于不兼容的版本约束、缺失的软件包以及未被记录的兼容性关系,导致许多真实世界的代码片段无法执行。本文提出了PLLM+,一种混合式依赖修复流水线,并在包含2,891个因依赖问题而失败的代码片段的HG2.9K基准数据集上进行了评估。PLLM+优先采用低成本、确定性的步骤,然后再调用基于大语言模型(LLM)的修复:基于静态AST的解释器推断、从竞赛提供的解决方案数据库中重放历史上成功的依赖配置,以及对候选软件包版本进行实时PyPI验证。当这些步骤无法解决某一案例时,系统会回退到结构化的基于LLM的修复循环,该循环采用类型化错误分类以及提议者/批评者(Proposer/Critic)智能体。在HG2.9K上,PLLM+成功解决了2,891个代码片段中的1,500个,而PLLM基线仅解决了1,169个。此外,它还将每个代码片段的平均运行时间从368.7秒降低到71.8秒。大多数成功的修复来自重放已知配置:1,500个成功修复中有1,495个由解决方案数据库产生,而LLM回退机制额外贡献了5个修复。这些结果表明,在该基准设置下,对先前已验证的依赖配置进行确定性重用是一种简单而有效的策略,而基于LLM的修复则作为先前解决方案未覆盖案例的次级回退手段。
cs.AI / 7 / 2609.26986

Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts

相同证据,不同判断:视觉/语音与文本冲突中证据的不可交换性
Li, Zhuoyun, Wang, Boxuan, Huang, Xiaowei, Dong, Yi
Abstract
For multimodal large language models, when images or speech conflict with accompanying text, measured text reliance can entangle modality preference with evidence position. Earlier studies of text bias often used a fixed evidence order or moved task instructions with the evidence, leaving the contribution of order unclear. In this paper, we use a paired comparison that keeps the instructions and evidence content fixed and swaps only the positions of the two sources to quantify this potential influence. Across vision and speech models, placing an image or recording after conflicting text consistently shifts answers toward its content. We also revisit previous studies and analyze why their experimental settings can lead to misleading conclusions. These findings reveal cross-modal evidence noncommutativity: the same evidence can lead to different judgments when its order changes, and placing perceptual evidence later can increase the model's reliance on its content.
Chinese Translation
对于多模态大语言模型,当图像或语音与 accompanying text 发生冲突时,所测得的文本依赖度可能将模态偏好与证据位置混为一谈。以往关于文本偏向的研究通常采用固定的证据顺序,或让任务指令随证据一起移动,导致顺序的贡献不明确。在本文中,我们采用一种配对比较方法:保持指令和证据内容不变,仅交换两个模态来源的位置,以量化这种潜在影响。在视觉和语音模型中,将图像或录音置于冲突文本之后,会持续使模型的回答偏向该模态的内容。我们还重新审视了以往的研究,并分析其实验设置为何会得出误导性结论。这些发现揭示了跨模态证据的不可交换性(noncommutativity):相同的证据在顺序改变时会导致不同的判断,而将感知证据置于较后位置会增强模型对其内容的依赖。
cs.AI / 8 / 2609.27035

Reinforcement Learning with Decomposed Subtasks

基于子任务分解的强化学习
Terzolo, Mattie, Sacha, Mikolaj, Sinha, Ayan, Rabinovich, Andrew
Abstract
Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental feedback, this collapsing is lossy: the optimizer must implicitly infer which competency drove the outcome and how that should change behavior. We argue the right primitive is not a better scalar but a decomposition: trajectory reward should be split along subtasks before it enters the policy update. We introduce Reinforcement Learning with Decomposed Subtasks (RLDS), whose core is Subtask-Decomposed Advantage Estimation (SDAE): a replacement for the scalar GRPO advantage that splits trajectory reward into per-subtask shares on a fixed taxonomy, computes a group-relative advantage per subtask, and distributes per-token credit by weighting each subtask's advantage by its importance, concentrating it around the step where a reflection marks that subtask's execution as consequential. We evaluate on four agentic benchmarks: FrozenLake (sparse grid navigation), HotpotQA (multi-hop QA, one retrieval tool), ScienceWorld (long-horizon embodied science), and DeepResearch (long-form research, four tools, composite rubric reward). Heterogeneity diagnostics emitted during training show where decomposition pays off - gains scale with subtask heterogeneity, largest on the high-heterogeneity tasks ScienceWorld (+11.5 points, paired-bootstrap 95% CI [+9.8, +13.3]) and FrozenLake (+9.8 points, [+7.0, +12.8]), and within noise on HotpotQA and DeepResearch, where the diagnostics predicted little to recover. ScienceWorld is also more compute-efficient under RLDS than scalar GRPO (-10.9% wall-clock per step), as long rollouts amortize the fixed reflect-and-grade overhead.
Chinese Translation
组相对策略优化(GRPO)及相关的策略梯度方法在训练语言模型智能体时,会在策略更新之前将整个多轮交互(rollout)压缩为单一的标量轨迹奖励。当任务由多种不同技能组合而成,尤其是在环境反馈稀疏且延迟的情况下,这种压缩是有损的:优化器必须隐式地推断是哪种能力决定了最终结果,以及这应如何改变行为。我们认为正确的原语不是更好的标量,而是分解:轨迹奖励应在进入策略更新之前沿子任务进行拆分。我们提出了基于子任务分解的强化学习(Reinforcement Learning with Decomposed Subtasks, RLDS),其核心是子任务分解优势估计(Subtask-Decomposed Advantage Estimation, SDAE):它替代了标量形式的GRPO优势,将轨迹奖励按照固定的子任务分类体系拆分为各子任务的份额,为每个子任务计算组相对优势,并通过按各子任务重要性加权的方式分配逐令牌(per-token)的信用分配,将信用集中在反思标记该子任务执行产生关键影响的步骤附近。我们在四个智能体基准上进行了评估:FrozenLake(稀疏网格导航)、HotpotQA(多跳问答,单一检索工具)、ScienceWorld(长时程具身科学任务)和DeepResearch(长篇研究任务,四种工具,复合评分准则奖励)。训练过程中输出的异质性诊断表明了分解在何处发挥作用——收益随子任务异质性增加而扩大,在高异质性任务ScienceWorld(提升11.5分,配对自助法95%置信区间[+9.8, +13.3])和FrozenLake(提升9.8分,[+7.0, +12.8])上收益最大,而在HotpotQA和DeepResearch上差异处于噪声范围内,这与诊断预测的收益甚微相符。此外,在ScienceWorld上,RLDS比标量GRPO更具计算效率(每步挂钟时间减少10.9%),因为长交互序列摊薄了固定的反思与评分开销。
cs.AI / 9 / 2609.27037

Training Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations

基于可控合成对话训练智能语音助手唤醒
Sowański, Marcin, Leszczyński, Kacper, Krzywicki, Kacper, Wodnicki, Krzysztof
Abstract
Wake word detection is a critical component of virtual assistants, serving as the gateway to seamless user interactions. This paper introduces a novel wake-up system that extends traditional direct keyword detection with contextual trigger detection. After an initial wake word activation, the system uses reasoning to distinguish between user commands and unrelated speech, ensuring efficient and context-aware engagement. We present a data generation architecture that produces a 62.3-hour corpus of controllable multi-speaker conversations containing direct invocations, contextual follow-ups, and non-addressed speech. Experimental results demonstrate the effectiveness of the proposed approach across diverse synthetic conversational scenarios. We release the code, dataset and trained models to promote reproducibility and further advancements in intelligent assistant technologies.
Chinese Translation
唤醒词检测是虚拟助手的关键组成部分,是实现流畅用户交互的入口。本文提出了一种新颖的唤醒系统,在传统的直接关键词检测基础上扩展了上下文触发检测。在初始唤醒词激活后,系统利用推理来区分用户指令与无关语音,从而确保高效且具备上下文感知的交互。我们提出了一种数据生成架构,生成了一个包含62.3小时可控多说话人对话的语料库,其中涵盖直接调用、上下文追问以及未被指向的语音。实验结果表明,所提出的方法在多种合成对话场景中均表现出有效性。我们公开了代码、数据集和训练好的模型,以促进可复现性并推动智能助手技术的进一步发展。
cs.AI / 10 / 2609.27038

Are Stated Reasoning Steps Causally Load-Bearing?

所陈述的推理步骤是否具有因果支撑作用?
Bhupatiraju, Abhiram, Nyaupane, Rayan
Abstract
Chain-of-thought (CoT) monitoring assumes that the reasoning a model writes reflects the computation that directly produces its answer. Previous faithfulness metrics have been predominantly behavioral, as they simply edit the reasoning text and observe the resulting answer. However, our methodology aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning. Unlike previous causal audits, which measure degradation, our interventions carry a known predicted target. In this way, each patch should switch the answer to a specific counterfactual entity derivable by construction. Specifically, we use synthetic multi-hop lookup tasks (2-6 hops). We patch the residual stream at the token span where the model states each intermediate step with the corresponding activations from a counterfactual run. For Qwen3-4B, 76.9% +/- 2.8% of stated steps are causally load-bearing (CLB) at the most responsive mid-network layer (random-position null: 11.3%; patching the underlying prompt fact: 83%, so stated steps carry approximately 96% of the achievable effect). Moreover, the standard behavioral test on the same items yields 88.2%, which overstates causal faithfulness by 11.4 percentage points (item-matched; 111:14 discordant pairs, p < 1e-15) and, for the easiest items, by up to 20 percentage points. This gap also has a clear capability dimension. Qwen3-1.7B is far less causally faithful overall (54.8%), with its faithfulness collapsing as reasoning depth increases (68% at 2 hops to 30% at 6), while Qwen3-4B remains relatively flat. Although stated reasoning can be causally meaningful, standard behavioral tests tend to overestimate its causal faithfulness, particularly on easier examples where model reasoning appears most fluent.
Chinese Translation
思维链(Chain-of-Thought, CoT)监控假设模型所写出的推理内容反映了直接产生其答案的计算过程。以往的可信度(faithfulness)度量方法主要是行为层面的,即仅通过编辑推理文本并观察由此产生的答案。然而,我们的方法旨在激活层面以因果方式度量可信度,并且专门针对模型自生成的推理。与以往测量性能退化的因果审计不同,我们的干预带有已知的预测目标。这样,每次激活修补(patching)都应将答案切换为一个可通过构造推导出的特定反事实实体。具体而言,我们使用合成的多跳查找任务(2-6跳)。我们在模型陈述每个中间步骤的词元片段处,用来自反事实运行的相应激活修补残差流。对于Qwen3-4B,在响应最灵敏的网络中层,76.9% ± 2.8%的所陈述步骤具有因果支撑作用(CLB)(随机位置零假设:11.3%;修补底层提示事实:83%,即所陈述步骤承载了约96%的可实现效应)。此外,在相同条目上进行的标准行为测试得到88.2%的结果,高估了因果可信度11.4个百分点(条目匹配;111:14不一致对,p < 1e-15),而在最简单的条目上高估幅度可达20个百分点。这一差距还呈现出明显的能力维度特征。Qwen3-1.7B的总体因果可信度要低得多(54.8%),且随着推理深度的增加,其可信度急剧下降(从2跳时的68%降至6跳时的30%),而Qwen3-4B则保持相对平稳。尽管所陈述的推理可以具有因果意义,但标准行为测试往往高估其因果可信度,尤其是在模型推理看起来最流畅的简单示例上。
cs.AI / 11 / 2609.27041

Math Reasoning in LLMs is Organized by Approach, Not Topic

大语言模型中的数学推理按解题方法而非主题组织
Goudarzi, Sajad, Zamanifard, Samaneh, Nasiri, Moloud, Rahimian, Hamed
Abstract
Mathematical reasoning benchmarks are typically organized by topic, but language models may organize their internal computation by reusable reasoning approach instead. In this paper, we investigate whether open math-capable LLMs organize internally by topical sub-skill or by reasoning approach, and we present evidence that the approach is the key. We introduce a generation-replay protocol: a model first generates a solution, after which we replay the exact prompt-plus-generation trajectory and extract activation-importance signatures over the reasoning tokens. We cluster these signatures without supervision across eight models and five mathematical reasoning sources, then evaluate the recovered structure with structural, semantic, and intervention tests. Across all 40 model-source cells, the recovered clusters outperform matched-size random baselines. Two independent frontier-LLM judges find approach-level coherence in 77-82% of real clusters versus 6-11% in within-source controls, and topic-pure clusters usually receive labels finer than the topic itself. In approach-controlled prompting, changing the requested reasoning approach shifts cluster assignment in seven of eight model conditions, whereas paraphrases largely preserve it. These results indicate that math-capable LLMs organize internal mathematical computation by reasoning approach rather than benchmark topic. The implication is that topic-stratified benchmarks and topic-balanced training corpora can still miss the axis that matters: even deliberately topic-balanced corpora may remain imbalanced over reasoning approaches.
Chinese Translation
数学推理基准通常按主题组织,但语言模型可能会按照可复用的推理方法而非主题来组织其内部计算。本文研究了具备数学能力的开源大语言模型(LLM)究竟是按主题子技能还是按推理方法进行内部组织,并提供了证据表明推理方法是关键。我们提出了一种“生成-回放”协议:模型先生成一个解答,随后我们回放完全相同的提示加生成轨迹,并提取推理标记上的激活重要性签名。我们在八个模型和五个数学推理数据来源上对这些签名进行无监督聚类,并通过结构性、语义性和干预性测试来评估所恢复的结构。在全部40个“模型-来源”组合中,恢复出的聚类均优于同等规模的随机基线。两个独立的前沿LLM评判者发现,77%-82%的真实聚类具有方法层面的一致性,而同源对照中仅为6%-11%,且纯主题聚类通常获得的标签比主题本身更细。在方法受控提示实验中,改变所要求的推理方法在八个模型条件中的七个中改变了聚类分配,而改述则大多保持聚类分配不变。这些结果表明,具备数学能力的LLM是按推理方法而非基准主题来组织其内部数学计算的。其含义在于,按主题分层的基准和主题均衡的训练语料仍可能忽略真正重要的维度:即使是刻意保持主题均衡的语料库,在推理方法的分布上仍可能是不均衡的。
cs.AI / 12 / 2609.27051

Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors

提议而不裁判:一种针对挖掘投资因子的LLM智能体的任意时刻有效裁判机制
Qu, Bo, Chen, Mingguang, Wang, Licheng
Abstract
Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones, in a synthetic world with planted truth, a probe-authoring environment and a ten-year walk-forward on the CSI 500. Who judges sets the number of false admissions: the frozen referee admits 5-11 times fewer sub-threshold factors than the leaky referees under a scripted proposer, and no proposer closes that gap. Who proposes sets the yield: the language model beats the script, matches the bandit, and adds the one capability a bandit lacks, writing its own diagnostic probes. The certificate's price is time: an admitted true factor waits about 500 trading days, and the certified portfolio's Sharpe ratio therefore trails an ungated one. Judging belongs to the procedure; proposing and instrument-making belong to the agent.
Chinese Translation
语言模型智能体现在已经承担了量化因子研究的全流程:它们提出投资因子、进行回测、筛选幸存因子并淘汰失效因子。我们探讨的问题是,这些工作中哪些应该交给智能体来完成。我们的答案是一种受治理的自进化机制:智能体可以提出因子,而一个智能体无法干预的、被冻结的统计裁判必须负责评判。该裁判仅根据提交之后才揭示的市场结果、通过下注的方式进行打分,因此其虚假发现保证对任何提议策略在每个停止时刻都成立。我们在一个植入真实信号的合成世界、一个探针编写环境以及中证500十年滚动测试中,将三类提议者(脚本、老虎机算法和语言模型)分别与该冻结裁判以及三个刻意存在数据泄漏的裁判进行交叉比较。由谁裁判决定了虚假准入的数量:在脚本提议者条件下,冻结裁判接纳的低于阈值的因子数量比泄漏裁判少5至11倍,且没有任何提议者能缩小这一差距。由谁提议决定了产出收益:语言模型优于脚本、与老虎机算法持平,并具备老虎机算法所缺乏的一项能力——自行编写诊断探针。这一证书的代价是时间:一个被准入的真因子大约需要等待500个交易日,因此经认证的组合的夏普比率落后于无门槛组合。裁判之权属于程序;提议与制作工具之权属于智能体。
cs.AI / 13 / 2609.27087

Policy-as-Skill: Governed LLM Decision Support with Evidence, Deterministic Control, and Audit

策略即技能(Policy-as-Skill):具备证据支撑、确定性控制与审计能力的治理型LLM决策支持
Mohsenzadegan, Kabeh, Tavakkoli, Vahid, Kyamakya, Kyandoghere
Abstract
Organizations increasingly use LLMs for policy, compliance, risk, and operational decision support, requiring evidence validation, review routing, version control, and auditability. We introduce Policy-as-Skill (PaS), a modular runtime that packages these functions as executable, versioned policy capabilities. Thirteen methods are evaluated with a fixed Gemma4 backend on 600 development tasks. PaS+Audit achieves 53.8% exact accuracy, macro-F1 0.346, review F1 0.854, citation precision 1.000, policy-reference recall 0.984, and audit completeness 1.000, outperforming LLM+RAG on most governance and review metrics. Deterministic control raises aggregate accuracy to 61.2% but is strongly task dependent, supporting selective rather than universal rule-based intervention.
Chinese Translation
组织日益使用大语言模型(LLM)进行政策、合规、风险及运营决策支持,这需要对证据的验证、审查路由、版本控制和可审计性。我们提出策略即技能(Policy-as-Skill,PaS),这是一种模块化运行时框架,将这些功能打包为可执行、可版本化的策略能力。在600个开发任务上,采用固定的Gemma4后端对十三种方法进行了评估。PaS+Audit实现了53.8%的精确准确率、0.346的宏平均F1(macro-F1)、0.854的审查F1、1.000的引用精确率、0.984的政策引用召回率以及1.000的审计完整性,在大多数治理与审查指标上优于LLM+RAG。确定性控制将总体准确率提升至61.2%,但其效果强烈依赖于具体任务,这表明应选择性地而非普遍地采用基于规则的干预。
cs.AI / 14 / 2609.27105

Provably Complete Generalized Planning with LLMs

基于大语言模型的可证明完备的广义规划
Stein, Katharina, Jain, Chaahat, Hoffmann, Jörg, Koller, Alexander
Abstract
Generalized planning aims to compute a plan that solves all instances of a planning domain. Recent work has used LLMs to automatically generate and debug such generalized plans in the form of Python programs and achieved perfect test data coverage for several domains. However, whether these generalized plans are actually complete, i.e. solve all instances of the domain, could only be determined by manual evaluation. Here, we present an approach for automatically generating generalized plans in Lean together with proofs of their completeness relative to a specification of the domain constraints provided as input. We introduce a semantic-preserving PDDL-to-Lean conversion, and use an LLM to generate both the generalized plan and the formal proof that it solves every instance satisfying the domain constraints. The correctness of the completeness proof is determined by Lean's kernel. We evaluate our approach on 13 commonly used benchmark domains, using GPT-5.6-Sol as the LLM. For 12 of the domains we obtain generalized plans together with valid completeness proofs. This is a major advancement of the state of the art in automatic generalized-plan completeness proofs.
Chinese Translation
广义规划旨在计算出一个能够解决规划域中所有实例的规划。近期的研究工作利用大语言模型(LLM)以Python程序的形式自动生成并调试此类广义规划,并在若干领域上实现了对测试数据的完美覆盖。然而,这些广义规划是否真正完备,即能否解决该领域的所有实例,此前只能通过人工评估来确定。本文提出了一种在Lean中自动生成广义规划的方法,并相对于以输入形式提供的领域约束规范,同时给出其完备性的证明。我们引入了一种保持语义的PDDL到Lean的转换方法,并使用大语言模型同时生成广义规划以及证明其能解决所有满足领域约束实例的形式化证明。完备性证明的正确性由Lean的内核来判定。我们在13个常用的基准领域上,使用GPT-5.6-Sol作为大语言模型对该方法进行了评估。在其中12个领域上,我们获得了广义规划及有效的完备性证明。这是自动广义规划完备性证明领域的一项重大进展。
cs.AI / 15 / 2609.27150

Do We Need Complex Topology Control? Distinct-Peer Random Routing Improves Cost-Efficiency in Sparse Multi-Agent Debate

我们需要复杂的拓扑控制吗?Distinct-Peer随机路由提升稀疏多智能体辩论的成本效率
Wang, Boxuan, Li, Zhuoyun, Huang, Xiaowei, Dong, Yi
Abstract
Multi-agent debate (MAD) has emerged as a promising paradigm for improving the reasoning accuracy of large language models (LLMs) through iterative peer interaction. Communication topology plays a central role in this process, motivating increasingly sophisticated mechanisms that learn, adapt, or dynamically reconfigure agent interactions to improve accuracy or reasoning reliability. Meanwhile, prior studies suggest that much simpler sparse communication can already achieve competitive performance at substantially lower cost. In this work, we take a closer look at sparse MAD and ask whether complex topology control is actually necessary to improve collective reasoning. We find that a simple random-without-replacement routing policy, which lets each agent debate with two distinct and newly sampled peers at every round, provides a surprisingly strong baseline and consistently improves the accuracy-cost trade-off of sparse MAD. Building on this observation, we further study deliberation stopping and show that lightweight stopping can substantially reduce inference cost while preserving competitive accuracy. Our results suggest that sophisticated topology control such as learned topology adaption should be evaluated against strong simple routing and stopping baselines before its additional complexity is justified.
Chinese Translation
多智能体辩论(Multi-Agent Debate, MAD)已成为一种有前景的范式,通过迭代式的同行交互提升大语言模型(LLMs)的推理准确性。通信拓扑在这一过程中扮演着核心角色,因此催生了日益复杂的机制,这些机制通过学习、自适应或动态重构智能体交互来提高准确性或推理可靠性。与此同时,先前的研究表明,简单得多的稀疏通信已经能以显著更低的成本取得有竞争力的性能。在本工作中,我们深入考察稀疏MAD,并探究复杂的拓扑控制对于改进集体推理是否确有必要。我们发现,一种简单的无放回随机路由策略——即让每个智能体在每一轮与两个不同的、新采样的同行进行辩论——提供了一个出人意料强大的基线,并持续改进稀疏MAD的准确性-成本权衡。基于这一观察,我们进一步研究辩论的停止策略,结果表明轻量级的停止机制能够在保持有竞争力的准确性的同时大幅降低推理成本。我们的结果意味着,在证明诸如学习式拓扑自适应等复杂拓扑控制的额外复杂性合理之前,应先将其与强大的简单路由和停止基线进行对比评估。
cs.AI / 16 / 2609.27197

Enhancing Small Language Models for Power Outage Report Generation via Minimum Risk Training

基于最小风险训练增强小型语言模型以生成停电报告
Phan, Hung, Abebe, Waqwoya, Hussein, Youssef, Chinthavali, Supriya, Lunga, Dalton, Jannesari, Ali
Abstract
Minimum Risk Training (MRT) enables neural machine translation models to directly optimize sequence-level evaluation metrics instead of relying only on token- level maximum-likelihood objectives Shen et al. [2016]. Although introduced a decade ago, recent work shows renewed potential for risk-based optimization in modern language models Yang et al. [2024], Jinnai et al. [2025]. We apply MRT to power outage report generation for the Outage Data Initiative Nationwide (ODIN), transforming heterogeneous reports into standardized XML compliant with CIM IEC 61968-3. Our MRT approach improves Qwen2.5-7B-Instruct overall accuracy from 16.20% to 68.95%, demonstrating the effectiveness of sequence- level optimization for domain-specific structured generation
Chinese Translation
最小风险训练(Minimum Risk Training, MRT)使神经机器翻译模型能够直接优化序列级评价指标,而不局限于仅依赖词元级的最大似然目标(Shen et al. [2016])。尽管该方法在十年前已被提出,但近期研究表明,基于风险的优化在现代语言模型中展现出新的潜力(Yang et al. [2024],Jinnai et al. [2025])。我们将MRT应用于全国停电数据计划(Outage Data Initiative Nationwide, ODIN)的停电报告生成任务,将异构报告转换为符合CIM IEC 61968-3标准的XML格式。我们的MRT方法将Qwen2.5-7B-Instruct的整体准确率从16.20%提升至68.95%,证明了序列级优化在特定领域结构化生成任务中的有效性。
cs.AI / 17 / 2609.27203

XLOG: A CUDA-Native Engine for Neurosymbolic Integration

XLOG:一个面向神经符号集成的 CUDA 原生引擎
Dubrovin, Levi, Pospelov, Nikita, Sabitov, Kirill
Abstract
xlog is a CUDA-native logic programming engine integrating neural perception with deterministic Datalog, probabilistic inference, and epistemic world views through a typed frontend and provider-owned CUDA runtime. Its reasoning modes share device data planes, but their execution boundaries differ: ordinary Datalog and exact inference are host-orchestrated, while certified resident recursive and Monte Carlo sampled cores record zero tracked host-device transfers before a bounded terminal receipt. The probabilistic path supports end-to-end gradients through GPU knowledge compilation from provenance to CNF to Decision-DNNF, exact weighted model counting, and backward gradients. A final smoothed circuit is certified against its source formula before caching or evaluation. Circuit caching yields a 2.74x MNIST-addition training speedup; a worst-case-optimal join subsystem yields a 27.96x geometric-mean gain over xlog's binary-join baseline. MNIST-addition accuracy matches Scallop's (0.9561 versus 0.9468), but no per-epoch speed claim is made because baseline epoch time varies with CPU quota. In five hub-skewed triangle-counting cases, the Souffle-to-fused-xlog execution-time ratio rises from 0.88x at 150k edges, where Souffle is faster, to 5.54x at 1.2M; fused peak device allocations are 85-1,033 MB versus 3,287-44,979 MB for the materializing arm. Exact inference is correctness-equivalent to but slower than ProbLog2. On a public video benchmark, a proximity predicate trained only through symbolic credit replaces hand-set geometry at unchanged held-out accuracy; within Event-Calculus rule search it fails ten-fold cross-validation and does not transfer on a leak-free split. On a maritime corpus, weighted clauses beat crisp selection by 0.065 F1, with the result reproduced by one chronological training pass.
Chinese Translation
xlog 是一个 CUDA 原生的逻辑编程引擎,通过带类型的前端和由提供方拥有的 CUDA 运行时,将神经感知与确定性 Datalog、概率推理以及认知世界观相集成。其各种推理模式共享设备端数据平面,但执行边界有所不同:普通 Datalog 和精确推理由主机编排,而经过认证的驻留式递归核心与蒙特卡洛采样核心在产生有界终止回执之前,记录的受追踪主机-设备传输次数为零。概率推理路径支持端到端梯度:通过 GPU 知识编译,从溯源到 CNF 再到 Decision-DNNF,进行精确加权模型计数和反向梯度传播。最终的平滑电路在缓存或评估之前,都会对照其源公式进行认证。电路缓存使 MNIST 加法训练获得 2.74 倍加速;最坏情况最优的连接子系统相比 xlog 的二元连接基线获得 27.96 倍的几何平均收益。MNIST 加法准确率与 Scallop 相当(0.9561 对 0.9468),但由于基线每个 epoch 的耗时随 CPU 配额变化,因此不做每 epoch 的速度声明。在五个枢纽偏斜的三角形计数案例中,Souffle 与融合版 xlog 的执行时间比从 15 万边时的 0.88 倍(此时 Souffle 更快)升至 120 万边时的 5.54 倍;融合版的峰值设备内存分配为 85–1,033 MB,而物化版本为 3,287–44,979 MB。精确推理在正确性上与 ProbLog2 等价,但速度较慢。在一个公开视频基准上,一个仅通过符号信用训练的邻近谓词在保持集准确率不变的情况下替代了手工设定的几何规则;在事件演算(Event-Calculus)规则搜索中,它未能通过十折交叉验证,且在无泄漏划分上无法迁移。在一个海事语料库上,加权子句比离散选择高出 0.065 F1,且该结果通过一次按时间顺序的训练即得到复现。
cs.AI / 18 / 2609.27273

CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments

CAVEAT:迈向激励错位环境下鲁棒的计算机使用智能体
Li, Yuxuan, Epperson, Will, Deng, Wesley, Huang, Zezhou
Abstract
Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives that do not align with the user's? In online marketplaces, for example, platforms may favor some products over others, potentially steering agents away from the user's objective. Existing CUA benchmarks cover cooperative settings or explicit attacks, but do not test whether agents preserve user objectives when the environment itself has a stake in the outcome. We introduce CAVEAT, a controlled benchmark spanning nine marketplace environments and a taxonomy of eight common steering mechanisms. Across five model families, agents purchase the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering mechanisms are enabled. Larger models and increased reasoning improve robustness, but substantial failures persist. Our trajectory analysis and targeted ablations identify three points where steering enters the decision process: (1) agents distort the user's priorities, (2) prematurely narrow the set of alternatives they consider, and (3) commit before resolving decision-relevant evidence. Guided by this diagnosis, we develop CAVEAT-Harness, which directly targets these failure modes and raises user-optimal purchasing by 55.0%. Targeted post-training further improves a smaller open model. These results establish incentive robustness as a distinct challenge for delegated agents, diagnose how it fails, and show that targeted interventions can substantially improve it.
Chinese Translation
计算机使用智能体(Computer-Use Agents, CUAs)越来越多地代表用户在网络上执行操作。当其所处环境的激励与用户的利益不一致时,会发生什么?例如,在在线市场中,平台可能偏袒某些产品,从而将智能体的行为引向偏离用户目标的方向。现有的CUA基准测试仅涵盖合作场景或显式攻击,但并未检验当环境本身与结果存在利害关系时,智能体能否坚持用户目标。我们提出了CAVEAT,一个受控基准,涵盖九个市场环境和八种常见引导机制的分类体系。在五个模型家族中,智能体在匹配对照情形下有78.6%的回合购买了用户最优产品,但在启用引导机制时这一比例仅为17.3%。更大的模型和更强的推理能力可以提升鲁棒性,但仍然存在大量失败。我们的轨迹分析和针对性消融实验识别出引导介入决策过程的三个环节:(1) 智能体扭曲了用户的优先级;(2) 过早缩小所考虑的备选方案集合;(3) 在尚未厘清与决策相关的证据之前就做出决定。基于这一诊断,我们开发了CAVEAT-Harness,直接针对这些失败模式,将用户最优购买率提升了55.0%。针对性的后训练进一步改进了一个较小的开源模型。这些结果将激励鲁棒性确立为委托式智能体面临的一项独特挑战,诊断了其失败机制,并表明有针对性的干预能够显著改善这一问题。
cs.AI / 19 / 2609.27276

DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents

DRSR:面向高效长程智能体的集合级删除风险学习
Wang, Mingxuan, Wang, Bo, Luo, Fei, Yao, Guorun, Ning, Chao, Guo, Yinglong, Chen, Hongyue, Ma, Yanbiao, Han, Jungong
Abstract
Long-horizon language-model agents accumulate reasoning traces, tool exchanges, and observations whose relevance changes with the current decision. Existing compression strategies often score historical units independently, but the safety of deleting several units is generally not determined by their singleton scores: redundant evidence, accumulated small effects, and the information that remains after deletion all matter. We introduce Direct Relational Set-Risk Pruning (DRSR), which formulates agent-history compression as risk-constrained selection over deletion sets. Offline, DRSR constructs exact counterfactual supervision by jointly deleting protocol-valid history Blocks and measuring the change in teacher-forced likelihood of the same recorded next output. A lightweight scorer then predicts set-level harm from online-visible relations between candidate history and the current pre-action state, together with deleted-retained and pairwise set structure. At deployment, DRSR evaluates a small set of structurally valid deletion candidates with the lightweight scorer and removes the largest feasible set under recency, protocol, budget, and learned-risk constraints, abstaining when no set is sufficiently safe. On WorkBuddyBench Full260, DRSR increases mean reward from 0.699 to 0.802 while reducing total model tokens by 20.820%. On the fixed Eval40 comparison, it obtains 0.794 reward at 1.211M tokens per task, using 35.850% fewer tokens than the uncompressed agent. Mechanistic analyses and ablations further show that decision-conditioned relations, retained-context information, pair interactions, and abstention each contribute to reliable pruning.
Chinese Translation
长程语言模型智能体会不断累积推理轨迹、工具交互记录和观察结果,而这些信息的相关性会随当前决策而变化。现有的压缩策略通常对历史单元独立打分,但删除多个单元的安全性通常并非由各单元的单独分数决定:冗余证据、微小效应的累积,以及删除后剩余的信息都会产生影响。我们提出直接关系集合风险剪枝(Direct Relational Set-Risk Pruning,DRSR),将智能体历史压缩形式化为带风险约束的删除集合选择问题。在离线阶段,DRSR 通过联合删除符合协议的历史块(Blocks),并度量教师强制下同一记录的下一条输出的似然变化,构建精确的反事实监督信号。随后,一个轻量级打分器基于候选历史与当前动作前状态之间的在线可见关系,以及删除-保留结构和成对集合结构,预测集合层面的危害。在部署时,DRSR 使用轻量级打分器评估少量结构上合法的删除候选集,并在近时性、协议、预算和学习风险等约束下移除最大的可行集合;当没有足够安全的集合时则选择不删除。在 WorkBuddyBench Full260 上,DRSR 将平均奖励从 0.699 提升至 0.802,同时将模型总 token 数减少 20.820%。在固定的 Eval40 对比中,其以每任务 1.211M token 的消耗获得 0.794 的奖励,比未压缩智能体少使用 35.850% 的 token。机制分析与消融实验进一步表明,决策条件化关系、保留上下文信息、成对交互以及弃删机制各自都对可靠剪枝有所贡献。
cs.AI / 20 / 2609.27277

TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent

TimeEvo:基于失败驱动的时间序列智能体自进化方法
Yang, Jie, Zheng, Yan, Sun, Jiarui, Fan, Xiran, Wang, Junpeng, Wang, Liang, Xu, Zelin, Liu, Qinghua, Fang, Zhengyu, Cai, Yiwei, Yu, Philip S.
Abstract
Time series agents answer analytical questions by calling external tools, and which tools they carry is decided by people before the agent runs. However, we identify two failures in this setup. Human-Agent Tool Misalignment: a library of 21 expert-curated tools helps on some tasks and hurts on others, dropping anomaly accuracy under every backbone we test. Silent Harm: one round of generic self-revision changes 147 answers and breaks 56 of them, while the final score moves by less than a point. Both follow from the same gap: whether a tool helps is decided question by question at runtime, while tools are supplied in advance and judged by a single average. To address this, we propose TimeEvo, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate. Experiments on ten time series QA tasks and three backbones show that TimeEvo, starting from an empty library, improves accuracy on every task and every backbone, and that a library grown on a cheap model still gains when it is installed into stronger ones. Code is available at https://github.com/Muyiiiii/TimeEvo.
Chinese Translation
时间序列智能体通过调用外部工具来回答分析性问题,而其携带哪些工具是在智能体运行前由人工决定的。然而,我们发现了这种设置中的两类失败。人机工具错配(Human-Agent Tool Misalignment):一个由专家精心整理的包含21个工具的工具库在某些任务上有所帮助,但在另一些任务上反而有害,在我们测试的所有骨干模型下均导致异常检测准确率下降。隐性损害(Silent Harm):一轮通用的自我修订更改了147个答案,其中56个被改坏,而最终分数的变化却不足一分。两者都源于同一个缺口:一个工具是否有用是在运行时逐个问题决定的,而工具却是预先提供并仅由单一平均值来评判的。为解决这一问题,我们提出了TimeEvo,它将智能体的诊断失败聚类为能力缺口,为每个缺口规划一种度量方法,合成仅基于证据的工具来填补这些缺口,并通过配对准入门槛来决定是否采纳候选工具库。在十个时间序列问答任务和三个骨干模型上的实验表明,TimeEvo从一个空工具库出发,在所有任务和所有骨干模型上都提升了准确率;而且,在廉价模型上生长出的工具库安装到更强的模型中时仍然能够带来收益。代码已发布于 https://github.com/Muyiiiii/TimeEvo。
cs.AI / 21 / 2609.27279

EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory

EnSIMem:面向智能体长期记忆的实体结构化索引
Meng, Xuanyu, Fan, Xing, Fan, Xinyi, Guo, Chenlei, Xie, Yixuan, Han, Jiawei
Abstract
An agent that interacts with users over long periods must recall facts, preferences, events, and changes from a continuously growing interaction history. Existing memory systems often compress interactions into generic summaries or retrieve anonymous text chunks, making it difficult for an agent to identify the correct entity, property, and supporting evidence. We present EnSIMem, an entity-structured long-term memory architecture for an agent. During offline construction, the system organizes interactions into theme-coherent episodes and builds dialogue-grounded index entries of the form [entity][entity type][property:value]. Each entry preserves its source turns, temporal information, and available multimodal fields. During online interaction, the agent's request is decomposed into evidence requirements whose properties are aligned with the memory index. Entity-property lookup and adaptive retrieval then collect the evidence needed for point, temporal, compositional, and aggregation reasoning. The agent generates its response from the preserved source evidence rather than from lossy memory summaries. On long-term agent-memory benchmarks, EnSIMem achieves high answer accuracy while maintaining compact contexts and favorable online efficiency. These results show that entity-structured indexing and episode-level provenance provide a reliable foundation for long-term memory in agents. The code of our model is available at https://github.com/RamonMeng/EnSIMem.
Chinese Translation
与用户进行长期交互的智能体必须能够从持续增长的交互历史中回忆事实、偏好、事件及其变化。现有的记忆系统通常将交互压缩为通用摘要,或检索匿名的文本片段,导致智能体难以确定正确的实体、属性及支持证据。我们提出了EnSIMem,一种面向智能体的实体结构化长期记忆架构。在离线构建阶段,该系统将交互组织为主题连贯的片段(episode),并构建以对话为基础的索引条目,其形式为[实体][实体类型][属性:值]。每个条目都保留其来源对话轮次、时间信息以及可用的多模态字段。在线交互时,智能体的请求被分解为证据需求,其属性与记忆索引对齐;随后通过实体-属性查找和自适应检索,收集用于点查询、时间推理、组合推理和聚合推理所需的证据。智能体基于保留的原始证据而非有损的记忆摘要来生成回复。在长期智能体记忆基准测试中,EnSIMem在保持紧凑上下文和良好在线效率的同时,取得了较高的答案准确率。这些结果表明,实体结构化索引与片段级溯源为智能体的长期记忆提供了可靠的基础。我们的模型代码已在 https://github.com/RamonMeng/EnSIMem 发布。
cs.AI / 22 / 2609.27284

Hunyuan-A13B Technical Report

Hunyuan-A13B 技术报告
Tencent Hunyuan Team, Liu, Ao, Zhou, Botong, Xu, Can, Zhou, Chayse, Zhang, ChenChen, Xu, Chengcheng, Wang, Chenhao, Wu, Decheng, Wu, Dengpeng, Jiao, Dian, Du, Dong, Wang, Dong, Zhang, Feng, Lian, Fengzong, Xu, Guanghui, Zhang, Guanwei, Wang, Hai, Luo, Haipeng, Hu, Han, Xu, Huilin, Wu, Jiajia, Zhu, Jianchen, Yan, Jianfeng, Zhu, Jiaqi, Zhang, Jihong, Xue, Jinbao, Xia, Jun, Zheng, Junqiang, Liu, Kai, Zhang, Kai, Zheng, Kai, Li, Kejiao, Wang, Keyao, Jiang, Lan, Liu, Lixin, Wu, Lulu, Huang, Mengyuan, Yu, Peijie, Wang, Peiqi, Wang, Qian, Xiang, Qianbiao, Liu, Qibin, Sun, Qingfeng, Guo, Richard, Xie, Ruobing, Yang, Saiyong, Chen, Shaohua, Hu, Shihui, Li, Shuai, Li, Shuaipeng, Chen, Shuang, Zheng, Suncong, Yang, Tao, Zhang, Tian, Yu, Tinghao, Han, Weidong, Liu, Weijie, Zhou, Weijin, Wang, Weikang, Chen, Wesleye, Feng, Xiao, Ren, Xiaoqin, Sun, Xingwu, Kuang, Xiong, Huang, Xuemeng, Cao, Xun, Chen, Yanfeng, Du, Yang, Yang, Zhen, Tao, Yangyu, Deng, Yaping, Shen, Yi, Hong, Yigeng, Chen, Yiqi
Abstract
We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost. The model is pretrained on a rigorously filtered 20T-token corpus with enhanced STEM data curation, improving factual reliability and reasoning ability. High-quality supervised fine-tuning and large-scale reinforcement learning further enhance its overall performance. Hunyuan-A13B also introduces a dual-mode Chain-of-Thought framework that adapts reasoning depth to task complexity: fast thinking for routine queries and slow thinking for complex, multi-step problems. Evaluations show competitive performance across mathematics, science, programming, general language understanding, and agent tasks, often approaching that of much larger models. Its high inference throughput makes it suitable for latency-sensitive applications. We release Hunyuan-A13B to support open research and practical LLM deployment.
Chinese Translation
我们提出了 Hunyuan-A13B,一个基于混合专家(Mixture-of-Experts)架构的开源大语言模型。该模型总参数量为 800 亿,但在推理时仅激活 130 亿参数,在模型能力、计算效率与部署成本之间实现了平衡。模型在一个经过严格筛选的 20 万亿词元(token)语料库上进行预训练,并强化了 STEM 数据的整理,从而提升了事实可靠性与推理能力。高质量的监督微调与大规模强化学习进一步提升了其整体性能。Hunyuan-A13B 还引入了一种双模式思维链(Chain-of-Thought)框架,能够根据任务复杂度自适应地调整推理深度:对于常规查询采用快速思考,对于复杂的多步骤问题则采用慢速思考。评测结果显示,该模型在数学、科学、编程、通用语言理解以及智能体(agent)任务上均表现出具有竞争力的性能,常常接近规模大得多的模型的水平。其高推理吞吐量使其适用于对延迟敏感的应用场景。我们开源 Hunyuan-A13B,以支持开放研究和 LLM 的实际部署。
cs.AI / 23 / 2609.27286

Memory Control Signals Emerge Before Action in Long Horizon Agents

记忆控制信号在长程智能体的行动之前出现
Wang, Mingxuan, Yao, Guorun, Luo, Fei, Guo, Yinglong, Ning, Chao, Wang, Bo, Chen, Hongyue, Ma, Yanbiao, Han, Jungong
Abstract
Long horizon language model agents continuously accumulate interaction history, increasing computational cost while making relevant information harder to preserve and reuse. Existing context management methods mainly focus on how to compress or retrieve history, but largely leave open whether the model itself already represents the need for these memory operations before they occur. We study the hidden state immediately before each agent action and find that compression and recall needs are already encoded in the model's internal representations. These signals cannot be explained by simple context length or interaction progress, and they exhibit distinct formation patterns across model depth. We further show that most memory decision information is preserved in a compact recent context, while selectively restored historical evidence complements the long range dependencies that recent context misses. Based on these findings, we propose Preaction Memory with Evidence Retrieval (PaMER), which combines state guided compression with external evidence retrieval. PaMER+ further introduces step level evidence selection to recover only the historical information required by the current task. Experiments on WorkBuddyBench, across multiple context management baselines and model backbones, show that our framework substantially reduces context consumption while maintaining competitive task performance.
Chinese Translation
长程语言模型智能体会持续累积交互历史,这不仅增加了计算成本,还使得相关信息的保留与复用变得更加困难。现有的上下文管理方法主要关注如何压缩或检索历史信息,但在很大程度上尚未回答一个关键问题:模型本身是否已经在这些记忆操作发生之前就表示出了对其的需求。我们研究了智能体每次行动之前瞬间的隐藏状态,发现压缩需求和召回需求已经被编码在模型的内部表示中。这些信号无法用简单的上下文长度或交互进度来解释,并且在不同模型深度上表现出截然不同的形成模式。我们进一步证明,大部分记忆决策信息被保存在紧凑的近期上下文中,而选择性地恢复的历史证据可以补充近期上下文所缺失的长程依赖。基于这些发现,我们提出了结合状态引导压缩与外部证据检索的 Preaction Memory with Evidence Retrieval(PaMER)。PaMER 进一步引入步骤级证据选择,仅恢复当前任务所需的历史信息。在 WorkBuddyBench 上、跨多种上下文管理基线和模型骨干的实验表明,我们的框架在保持具有竞争力的任务性能的同时,显著降低了上下文消耗。
cs.AI / 24 / 2609.27288

PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks

PotARCin:抽象推理任务中技能习得的多维度评估
Beger, Claas, Yi, Ryan, Mitchell, Melanie
Abstract
The Abstraction and Reasoning Corpus (ARC) has become a prominent benchmark for evaluating general abstract reasoning and fluid intelligence in AI models. Yet standard ARC evaluation considers only a single capability: producing the correct output grid for a test input. We argue that this narrow format fails to evaluate the diversity of abilities that genuine abstract skill acquisition should enable. We introduce PotARCin, a benchmark that extends ARC by assessing understanding of a task's underlying abstract rule across five dimensions: Definition, Classification, Constrained Generation, Editing, and Inversion. PotARCin employs programmatic methods to generate new task instances and transform given inputs for a given ARC task, enabling dynamic generative sampling beyond fixed input-output pairs. Across five state-of-the-art models evaluated on the ARC-AGI-1 training set, we observe a 25-52 percentage-point performance gap between standard ARC evaluation and evaluation on PotARCin, and find that multi-dimensional evaluation reorders models that standard accuracy ranks alike. We further investigate effects of generative sampling, difficulty of corruption types, and questions of self-consistency, showing that models frequently contradict their own formalized rule even where they have stated it correctly. We also introduce P-ARC, a held-out hand-crafted test set, on which models achieve 1-8% accuracy across all five dimensions, underscoring the importance of more holistic evaluations of abstract reasoning capabilities.
Chinese Translation
抽象与推理语料库(Abstraction and Reasoning Corpus, ARC)已成为评估人工智能模型通用抽象推理能力与流体智力的重要基准。然而,标准ARC评估仅考虑单一能力:为测试输入生成正确的输出网格。我们认为,这种狭窄的评估形式无法评估真正的抽象技能习得所应具备的多样化能力。我们提出了PotARCin,这是一个在五个维度上评估对任务底层抽象规则理解的扩展ARC基准:定义(Definition)、分类(Classification)、受限生成(Constrained Generation)、编辑(Editing)和反转(Inversion)。PotARCin采用程序化方法生成新的任务实例,并对给定ARC任务的输入进行变换,从而实现超越固定输入-输出对的动态生成式采样。在ARC-AGI-1训练集上评估的五个最先进模型中,我们观察到标准ARC评估与PotARCin评估之间存在25至52个百分点的性能差距,并发现多维度评估会重新排列那些在标准准确率下排名相近的模型的次序。我们进一步研究了生成式采样的影响、损坏类型的难度以及自一致性问题,结果表明模型即使正确表述了规则,也经常与其自身形式化的规则相矛盾。我们还引入了P-ARC,一个保留的手工构建测试集,模型在所有五个维度上仅达到1-8%的准确率,这凸显了对抽象推理能力进行更全面评估的重要性。
cs.AI / 25 / 2609.27290

Sparse-Observation Atmospheric Thermal Forecasting with Physics-Informed Neural Networks for Climate-Aware Digital Twins

基于物理信息神经网络的稀疏观测大气温度预报:面向气候感知数字孪生
Chegini, Tannaz Goodarzvand, Shivanian, Elyas, Karimi, Behzad, Dadgostari, Faraz
Abstract
Short-horizon forecasts of atmospheric temperature are needed to support climate-aware digital-twin systems, but such forecasts must be produced where thermal observations are incomplete. This study evaluates a physics-informed neural network for potential-temperature forecasting, constrained by a pressure-coordinate thermodynamic advection-source equation and a diabatic-source closure fit from the preceding 12-hour period and frozen before future-time training. Using hourly ERA5 reanalysis at three pressure levels, the model is evaluated as a conditional hindcast at lead times of one, two and three hours against persistence, local-trend, and two matched neural-network baselines, one of which receives the same future meteorological forcing as the PINN, helping distinguish the physical constraint from access to future forcing. In an Oklahoma development case, mean RMSE improvement over the strongest baseline grew from 8.1\% at one hour to 23.8\% at three hours; under an observation-density sweep down to 5\% of candidate locations, this 3-hour advantage remained 14.6--16.9\%, with no evidence that lower density improves performance. Under a fixed protocol transferred to an Alabama heat event with three virtual-observation layouts, three-hour improvement ranged 19.7-24.4\% with consistent origin-level wins. A parallel Montana stress test, in which fixed pressure levels intersected complex terrain, produced a three-hour degradation of roughly 17.5\%, identifying a terrain-related applicability limit of the formulation. Together, these results indicate that the physics constraint's benefit grows with forecast horizon, persists under severe observation sparsity, and transfers across regions, but is bounded by the validity of a fixed vertical-coordinate representation over complex terrain, evidence relevant to physics-constrained components of climate-aware forecasting and digital-twin systems.
Chinese Translation
短期大气温度预报是支撑气候感知数字孪生系统所必需的,但此类预报往往需要在热力观测不完整的条件下生成。本研究评估了一种用于位温预报的物理信息神经网络(PINN),该模型受气压坐标系下的热力平流-源项方程约束,并采用由前12小时时段拟合、在未来时间训练前冻结的非绝热源项闭合方案。基于三个气压层的逐小时ERA5再分析资料,模型以条件性后报(hindcast)方式在1、2和3小时预报时效上进行评估,并与持续性预报、局地趋势预报以及两个匹配的神经网络基线进行比较,其中之一接收与PINN相同的未来气象强迫,以帮助区分物理约束的贡献与获取未来强迫的贡献。在俄克拉荷马州的发展个例中,相对于最强基线的平均RMSE改进从1小时的8.1%增长到3小时的23.8%;在观测密度降至候选位置5%的密度扫描实验中,3小时的预报优势仍保持在14.6–16.9%,且没有证据表明更低的观测密度能提升性能。在将固定实验方案迁移至阿拉巴马州一次高温事件(采用三种虚拟观测布局)的试验中,3小时改进幅度为19.7–24.4%,并在各起始时间上保持一致的优势。而在蒙大拿州的平行压力测试中,固定气压层与复杂地形相交,导致3小时预报性能下降约17.5%,由此确定了该公式在地形相关的适用性局限。总体而言,这些结果表明:物理约束的收益随预报时效增长,在严重的观测稀疏条件下依然持续,并可跨区域迁移,但其受限于固定垂直坐标表示在复杂地形上方的有效性,这些证据对气候感知预报与数字孪生系统中的物理约束组件具有参考意义。
cs.AI / 26 / 2609.27297

Large Knowledge Model: From Papers to a Scientific Reasoning Landscape

大知识模型:从论文到科学推理图景
Huang, Yuan, Hu, Sihan, Gu, Hongyu, Ma, Chao, Zhang, Jiaxing, Zou, Zhiyong, Fan, Caiyu, Xiao, Yan, Xu, Mingjun, Xie, Chenyu, Ju, Mingzhen, Ma, Zhehao, Zhang, Qi, Wang, Baozong, Li, Yu, Yao, Zhiyuan, Liao, Ruoxue, Li, Xinyu, Zhang, Linfeng, Chen, Kun, E, Weinan
Abstract
Accumulated scientific knowledge advances inquiry when prior findings help researchers choose new questions, design investigations, and interpret results. Realizing this value at scale requires access to the reasoning that connects research problems, scientific procedures, conclusions, and evidence. We introduce the Large Knowledge Model (LKM), a scientific knowledge infrastructure that transforms the literature into a shared, computationally accessible reasoning resource. LKM represents papers as source-grounded reasoning graphs, couples structural traversal with semantic retrieval over the same objects, and aligns related questions, claims, and reasoning chains across papers. This representation forms a Scientific Reasoning Landscape with three connected views: a Question Landscape that organizes research problems and open directions, a Workflow Landscape that exposes reusable scientific procedures, and an Evidence Landscape that connects conclusions to their support, disagreement, and conditions. The unified substrate supports reasoning-aware scientific search, evidence-grounded question answering, comparative evidence analysis, and research planning. Researchers and agents can retrieve relevant work through its scientific intent, synthesize answers with inspectable supporting arguments, and develop research plans informed by established workflows and unresolved evidence. We describe a corpus-scale system and evaluate scientific retrieval and knowledge-intensive question answering. With the answering model fixed, LKM retrieval improves accuracy by 9.30%, 4.20%, and 14.69% on ChemBench, PubMedQA, and SciBench, respectively. By connecting knowledge access to scientific reasoning and action, LKM provides a common foundation for discovering relevant research, reusing scientific knowledge, and coordinating cumulative inquiry across researchers, agents, and research cycles.
Chinese Translation
积累的科学知识之所以能推动探索,是因为已有发现能帮助研究者选择新问题、设计研究方案并解读结果。要在规模上实现这一价值,需要获取将研究问题、科学流程、结论与证据联系起来的推理过程。我们提出大知识模型(Large Knowledge Model, LKM),这是一种将文献转化为共享的、可计算访问的推理资源的科学知识基础设施。LKM 将论文表示为有出处依据的推理图,在同一对象上耦合结构化遍历与语义检索,并对齐不同论文中相关的问题、论断和推理链。这种表示构成了一个包含三个相互关联视图的科学推理图景(Scientific Reasoning Landscape):组织研究问题与开放方向的问题图景(Question Landscape)、揭示可复用科学流程的工作流图景(Workflow Landscape),以及将结论与其支持、分歧和条件联系起来的证据图景(Evidence Landscape)。这一统一基底支持具有推理意识的科学检索、有证据依据的问答、比较性证据分析以及研究规划。研究者和智能体可以依据科学意图检索相关工作,借助可审查的支撑论据综合生成答案,并基于已确立的工作流和未决的证据制定研究计划。我们构建了一个语料库规模的系统,并对科学检索与知识密集型问答进行了评估。在固定答题模型的情况下,LKM 检索在 ChemBench、PubMedQA 和 SciBench 上分别将准确率提升了 9.30%、4.20% 和 14.69%。通过将知识获取与科学推理及行动相连接,LKM 为发现相关研究、复用科学知识以及在研究者、智能体和研究周期之间协调累积性探索提供了共同基础。
cs.AI / 27 / 2609.27298

StateComp: Learning When to Compress History in Long Horizon Agents

StateComp:学习何时在长时程智能体中压缩历史信息
Wang, Mingxuan, Chen, Hongyue, Guo, Yinglong, Luo, Fei, Ning, Chao, Wang, Bo, Yao, Guorun, Ma, Yanbiao, Han, Jungong
Abstract
Long-horizon agents continuously accumulate interaction history during task execution, yet the importance of past interactions changes as the agent state evolves. Existing context management methods largely compress history based on fixed windows, periodic schedules, or current relevance, overlooking a more fundamental question: when has a past interaction become safe to replace? Premature compression may remove information still needed for future actions, while overly conservative retention leads to substantial context overhead. To address this, we propose State Conditioned Compression (StateComp), a framework that determines when historical interactions can be safely compressed according to the current agent state. StateComp constructs KEEP and READY supervision through a two-stage annotation procedure and trains an imbalance-aware router on hidden representations from a frozen language model. A bounded state representation further reduces the cost of evaluating long histories, while adjacent READY interactions are grouped into continuous spans and replaced with compact summaries during execution. Experiments on WorkBuddyBench show that StateComp reduces total agent and summarization tokens by 52.27% while maintaining task performance, and achieves a 12.67-fold speedup in representation extraction.
Chinese Translation
长时程智能体在任务执行过程中会不断积累交互历史,然而随着智能体状态的演变,历史交互的重要性也在不断变化。现有的上下文管理方法主要基于固定窗口、周期性调度或当前相关性来压缩历史,忽视了一个更根本的问题:过去的交互何时才变得可以安全地被替换?过早压缩可能会移除未来动作仍然需要的信息,而过于保守的保留则会带来可观的上下文开销。为解决这一问题,我们提出了状态条件压缩(State Conditioned Compression,StateComp),一个根据当前智能体状态判断历史交互何时可以安全压缩的框架。StateComp通过两阶段标注流程构建KEEP和READY监督信号,并在冻结语言模型的隐藏表示上训练一个面向不平衡数据的路由器。有界状态表示进一步降低了评估长历史的开销,同时相邻的READY交互被分组为连续片段,并在执行过程中被简洁的摘要所替换。在WorkBuddyBench上的实验表明,StateComp在保持任务性能的同时,将智能体总令牌数与摘要令牌数减少了52.27%,并实现了12.67倍的表示提取加速。
cs.AI / 28 / 2609.27307

Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents

从自身交互中学习行动:面向GUI智能体的在策略自蒸馏方法
Zhang, Yan, Wu, Daiqing, Shen, Huawen, Li, Liang, Cao, Gang, Gong, Zhi, Dai, Wei, Zhang, Xiaode, Ma, Can, Zhou, Yu
Abstract
Graphical User Interface (GUI) agents enable the fulfillment of complex user instructions through multi-turn interactions with software environments, requiring step-wise reasoning and long-horizon memory to guide actions and retain task-relevant information, respectively. Recent on-policy self-distillation (OPSD) methods have achieved strong performance on GUI grounding, a foundational subtask for GUI agents, owing to dense token-level supervision from privilege-conditioned self-teachers. However, extending existing OPSD methods to multi-turn GUI agents is hindered by self-teachers' limited privilege-following ability and insufficient privileged guidance. In this paper, we introduce GUI-SD-v2, the next version of GUI-SD, which extends OPSD from GUI grounding to multi-turn GUI interaction and addresses key limitations through a two-stage training framework. Specifically, GUI-SD-v2 first strengthens privilege following by jointly optimizing rollouts with and without privileged guidance from the same GUI states. Furthermore, it selectively distills step-specific reasoning and memory guidance through a privilege-conditioned self-teacher, supporting action decisions and the retention of task-relevant information for subsequent interactions. Extensive experiments on two representative GUI agent benchmarks, AndroidWorld and MobileWorld, show that GUI-SD-v2 compares favorably with existing OPSD baselines while consistently outperforming the evaluated state-of-the-art methods in both Pass@1 and Pass@3 success rates. Code and training data will be publicly released.
Chinese Translation
图形用户界面(GUI)智能体通过与软件环境的多轮交互来完成复杂的用户指令,这需要逐步推理来指导动作,并需要长程记忆来保留与任务相关的信息。近期的在策略自蒸馏(On-Policy Self-Distillation, OPSD)方法在GUI定位(GUI grounding,GUI智能体的一项基础子任务)上取得了出色表现,这得益于特权条件自教师提供的密集token级监督。然而,将现有OPSD方法扩展到多轮GUI智能体受到自教师有限的特权遵循能力和不充分的特权引导的阻碍。本文提出了GUI-SD-v2,即GUI-SD的下一个版本,它将OPSD从GUI定位扩展到多轮GUI交互,并通过两阶段训练框架解决了上述关键局限。具体而言,GUI-SD-v2首先通过联合优化来自相同GUI状态的、带特权引导与不带特权引导的rollout,来增强特权遵循能力。此外,它通过特权条件自教师有选择地蒸馏步骤级的推理与记忆引导,支持动作决策以及后续交互中任务相关信息的保留。在两个具有代表性的GUI智能体基准AndroidWorld和MobileWorld上的大量实验表明,GUI-SD-v2优于现有OPSD基线,并在Pass@1和Pass@3成功率上持续超越所评估的最先进方法。代码和训练数据将被公开发布。
cs.AI / 29 / 2609.27321

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

可验证的隐藏动态博弈:从已求解机制生成智能体强化学习环境
Shen, Xinjie, Fan, Wei, Guo, Xudong, Tu, Jianhong, Su, Yang, Kuang, Chuqiao, Zhang, Yinger, Liu, Dayiheng
Abstract
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before defining its outcome rule or annotating its trajectories, leaving dynamics and evaluation to be aligned post hoc. VHD-Play reverses this dependency by sampling and solving a mathematical model before a corpus-grounded setter renders its decision process as stateful tools. The executable dynamics and trajectory-scoring reference are inherited from the same solved model. The pipeline produces 3,300 diverse agentic environments at a cost of a few cents each. Training Qwen3.6-35B-A3B on three families raises its mean agentic score from 0.204 to 0.815 in a five-family diagnostic. Gains also appear on held-out instances from all three training families and eight unseen mechanism families, then extend beyond the generated substrate to external benchmarks for general function calling, travel planning, and 365-day e-commerce. On E-Commerce Bench, the trained checkpoint completes every run without bankruptcy and exceeds Qwen3.7-Max. We compare written-out problems with stateful versions that reveal or hide their parameters. The comparison shows that most of the learnable gap lies in stateful interaction rather than underlying problem solving. A frozen 35B setter realizes larger environments, and scale-matched training retains gains as mechanism size and horizon grow, indicating the potential for an evolving training substrate.
Chinese Translation
语言模型智能体日益面临具有演化状态、相互依赖决策和延迟结果的长时程任务。扩展其训练需要多样化的智能体环境、可靠的结果信号以及较低的扩展成本。现有的环境生成流程通常在定义结果规则或标注轨迹之前就构建环境,导致动态机制与评估只能在事后进行对齐。VHD-Play 反转了这一依赖关系:先采样并求解一个数学模型,再由基于语料的设置器(setter)将其决策过程渲染为有状态的工具。可执行的动态机制和轨迹评分参考均继承自同一个已求解的模型。该流程以每个环境几美分的成本生成了 3,300 个多样化的智能体环境。在五个机制家族的诊断集上,使用其中三个家族训练 Qwen3.6-35B-A3B,使其平均智能体得分从 0.204 提升至 0.815。增益还体现在来自全部三个训练家族和八个未见机制家族的留出实例上,并进一步超越所生成的载体,扩展到通用函数调用、旅行规划和 365 天电商等外部基准。在 E-Commerce Bench 上,训练后的模型在每次运行中均未破产,并超越了 Qwen3.7-Max。我们还比较了显式写出问题与揭示或隐藏参数的有状态版本,结果表明大部分可学习的差距在于有状态交互而非底层问题求解。一个冻结的 35B 设置器能够实现更大规模的环境,且规模匹配的训练在机制规模和时程增长时仍能保持增益,这表明了构建不断演化的训练载体的潜力。
cs.AI / 30 / 2609.27332

Stable Geometry with Divergent Task Evidence for Efficient Long-Horizon Agent Compression

面向高效长时程智能体压缩的稳定几何结构与发散的任务证据
Wang, Mingxuan, Luo, Fei, Wang, Bo, Yao, Guorun, Guo, Yinglong, Ning, Chao, Chen, Hongyue, Ma, Yanbiao, Han, Jungong
Abstract
Long horizon agents accumulate growing interaction histories that increase context and inference costs. We find that geometric redundancy alone is an insufficient criterion for safe compression. Although agent histories exhibit strong low dimensional structure, similar global geometry can preserve very different amounts of task evidence. At identical retained block counts, evidence aware selection raises next action Top 3 retention from 0.31 to 0.69, while centroid similarity remains 0.98. Controlled replacement further shows that action related information can be substantially altered while global geometric measures remain nearly unchanged. Motivated by this gap between geometry and evidence, we introduce Geometry Guided Evidence Preserving Memory (GEM), a training free compressor that protects task and execution evidence before using geometric residuals to complete coverage. GEM reduces mean combined token usage from 2.69M to 2.11M per task, a 21.4% reduction, while maintaining comparable task reward. Our results show that efficient agent history compression should optimize for preserved task evidence rather than geometric coverage alone.
Chinese Translation
长时程智能体不断累积的交互历史会持续增加上下文与推理成本。我们发现,仅凭几何冗余并不足以作为安全压缩的充分判据。尽管智能体历史呈现出强烈的低维结构,相似的全局几何结构却可能保留差异巨大的任务证据量。在保留块数量相同的情况下,证据感知选择将下一步动作的 Top 3 保留率从 0.31 提升至 0.69,而质心相似度仍保持在 0.98。受控替换实验进一步表明,动作相关信息可以被大幅改变,而全局几何度量几乎保持不变。基于几何与证据之间的这一差距,我们提出了几何引导的证据保留记忆方法(Geometry Guided Evidence Preserving Memory,GEM),这是一种无需训练的压缩器,在使用几何残差补全覆盖之前优先保护任务与执行证据。GEM 将每个任务的平均总 token 使用量从 2.69M 降低至 2.11M,减少了 21.4%,同时保持了相当的任务奖励。我们的结果表明,高效的智能体历史压缩应当以保留任务证据为优化目标,而非仅追求几何覆盖。
cs.AI / 31 / 2609.27333

Alignment Inertia: Auditing the Durability of Training Data Influence Through Policy Override Resistance

对齐惯性:通过策略覆盖抗性审计训练数据影响的持久性
Barreto, Renata, Roesti, Markelle, Tahaei, Mohammad
Abstract
Platform operators increasingly rely on system prompts and fine-tuning to govern model behavior, yet it remains unclear how reliably these interventions override behavior inherited from prior training. We propose Override Success Rate (OSR) and alignment inertia to measure when operator interventions succeed or fail to change prior behavior. We evaluate zero-shot prompting and LoRA fine-tuning across Llama and Mistral in medical misinformation and hate speech. Alignment inertia persists across both models but varies by model, domain, and policy direction. Notably, in Mistral's restrictive hate-speech condition, LoRA increased inertia by 46.5 percentage points, showing that fine-tuning can reinforce rather than override prior behavior. We also use TRAK to test whether inertia is associated with weaker adaptation signals. TRAK achieves AUC of at least 0.85 in 7 of 8 conditions and outperforms model confidence, TF-IDF similarity, and embedding similarity as a predictor of inertia. These results provide an operator-facing audit of where prior training constrains downstream model governance.
Chinese Translation
平台运营商日益依赖系统提示词和微调来管控模型行为,然而这些干预措施能否可靠地覆盖先前训练所继承的行为仍不清楚。我们提出覆盖成功率(Override Success Rate, OSR)和对齐惯性(alignment inertia)来衡量运营商干预在改变既有行为时的成败。我们在医疗错误信息和仇恨言论两个领域,对 Llama 和 Mistral 模型评估了零样本提示和 LoRA 微调。对齐惯性在两个模型中均持续存在,但随模型、领域和政策方向的不同而变化。值得注意的是,在 Mistral 的限制性仇恨言论条件下,LoRA 使惯性提高了 46.5 个百分点,表明微调可能强化而非覆盖既有行为。我们还使用 TRAK 检验惯性是否与较弱的适配信号相关。TRAK 在 8 个条件中的 7 个达到至少 0.85 的 AUC,并且作为惯性的预测指标优于模型置信度、TF-IDF 相似度和嵌入相似度。这些结果为运营商提供了一项审计,揭示先前训练在哪些方面约束了下游模型治理。
cs.AI / 32 / 2609.27334

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

即时记忆:为LLM智能体学习构建任务自适应记忆
Zhou, Yefan, Li, Yang, Liu, Zeyu Leo, Yavuz, Semih, Joty, Shafiq
Abstract
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many possible downstream tasks. Learning such a write-time curator is also difficult because the value of a storage decision may only become apparent when a relevant query arrives, potentially many tasks later, creating a long-horizon credit-assignment problem. We instead retain raw trajectories and defer curation until read time, when the current task is known. Given the retrieved traces and the new task, a memory curator synthesizes a compact, task-adaptive payload tailored to the immediate need. Because this payload is consumed on the same task, the curator can be trained directly from immediate task success, avoiding delayed utility signals and the need to artificially group related tasks. Across ALFWorld, WebShop, and $\tau^2$-bench, our Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively. Notably, even an untrained curator is already competitive with or surpasses these baselines, showing that task-adaptive read-time curation itself is a major source of the gain; training the curator further compounds the improvement.
Chinese Translation
智能体记忆系统通过复用过往经验来提升未来性能,然而现有的大多数设计都是在写入时对记忆进行整理:一旦任务完成,其轨迹便被提炼为固定的产物,例如反思、工作流、技能或推理策略,之后通过相似度检索。这迫使系统在未知未来查询的情况下就决定什么值得记住,不可逆地丢弃信息,并产生一个与查询无关的摘要,而该摘要必须服务于许多可能的下游任务。学习这种写入时的记忆整理器也十分困难,因为某次存储决策的价值可能只有在相关查询到来时才显现,而查询可能发生在许多任务之后,这就形成了长时程的信用分配问题。我们转而保留原始轨迹,将记忆整理推迟到读取时进行,此时当前任务已经已知。在给定检索到的轨迹和新任务的情况下,记忆整理器合成一个紧凑的、针对当前需求定制的任务自适应负载。由于该负载在同一任务上被消费,整理器可以直接根据即时任务成功与否进行训练,从而避免了延迟的效用信号以及人为地对相关任务进行分组的需要。在ALFWorld、WebShop和$\tau^2$-bench上,我们的即时记忆(JitMem)始终优于无记忆智能体以及启发式和可学习的写入时记忆方法,分别以16.2、16.3和3.9个绝对成功率的提升超越最强基线。值得注意的是,即使是未经训练的整理器也已经能与这些基线竞争甚至超越它们,这表明读取时的任务自适应记忆整理本身就是收益的主要来源;对整理器进行训练则会进一步放大这一改进。
cs.AI / 33 / 2609.27336

CART: Closed-Loop Adaptive Red Teaming for Large Language Models

CART:面向大语言模型的闭环自适应红队测试
Zhang, Dongdong, Lv, Tengchao, Jia, Yilin, Zhao, Yuzhong, Huang, Yupan, Wu, Wenshan, Zhou, Xiangyang, Huang, Shaohan, Yang, Nan, Dong, Li, Cui, Lei, Wei, Furu
Abstract
Automated red teaming often replays a fixed set of prompts, which measures known risks but cannot learn from failures found during testing. We present CART (Closed-Loop Adaptive Red Teaming), a framework that uses each result to guide what it tests next. CART begins with broad risk coverage, follows weaknesses that emerge, keeps new probes diverse, and records the evidence and source of every finding. It separates the Challenger that creates tests, the Target being tested, which may be a text-only model or a bounded tool-using agent, and the Judge that evaluates the results, allowing these roles to be studied independently. Across three evaluation families (Frontier, JAH, and Agentic), CART discovers more failures and higher average risk than static seed replay for every Target with an available baseline. The gains extend to tool-mediated agent tests, suggesting that contextual adaptation can reveal weaknesses that direct prompt replay does not exercise. These results describe what the test policies discover, not how often failures occur in real deployments. We also find that Challenger-Judge choices affect the evidence uncovered, highlighting the need for role separation and independent review. Overall, CART turns red teaming from a one-time checklist into a continuous, adaptive, and auditable search for model and agent weaknesses.
Chinese Translation
自动化红队测试通常重放固定的提示词集合,这只能衡量已知风险,却无法从测试过程中发现的失败中学习。我们提出了CART(闭环自适应红队测试,Closed-Loop Adaptive Red Teaming),一个利用每次测试结果来指导后续测试内容的框架。CART从广泛的风险覆盖入手,追踪测试中暴露出的薄弱环节,保持新探测手段的多样性,并记录每项发现的证据及其来源。它将被测角色分离为:创建测试的挑战者(Challenger)、被测试的目标(Target,可以是纯文本模型或受限的工具使用智能体),以及评估结果的裁判(Judge),从而可以对这些角色进行独立研究。在三个评估体系(Frontier、JAH和Agentic)中,对于所有有基线可用的目标模型,CART都比静态种子重放发现了更多的失败和更高的平均风险。这些收益延伸到工具中介的智能体测试中,表明情境化自适应能够揭示直接提示词重放无法触及的薄弱环节。这些结果描述的是测试策略所发现的问题,而非真实部署中失败发生的频率。我们还发现挑战者-裁判的配置选择会影响所揭示的证据,这凸显了角色分离和独立审查的必要性。总体而言,CART将红队测试从一次性的检查清单转变为对模型和智能体薄弱环节进行持续、自适应且可审计的搜索。
cs.AI / 34 / 2609.27349

MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

MolDesignBench:面向基于大语言模型的智能体的情境化分子设计评测基准
Jeong, Yongjun, Ko, Hanbum, Kim, Ye Rin, Lee, Chanhui, Hormazabal, Rodrigo, Lee, Jaewan, Han, Sehui, Lim, Sungbin, Kim, Sungwoong
Abstract
Real-world molecular design remains challenging for large language model (LLM)-based agents. It requires them to interpret design contexts, satisfy multiple constraints, identify infeasible specifications, and reason over multi-step tool outputs. Existing benchmarks do not capture this complexity, focusing instead on explicit and narrow constraints, only feasible problems, and single-path solutions. To address this gap, we propose MolDesignBench, a scenario-grounded benchmark that more closely reflects real-world molecular design for evaluating tool-augmented LLM agents. MolDesignBench comprises 2K generation and optimization instances that combine implicit requirements embedded in design narratives with explicit property and functional-group constraints, including infeasible cases, and require the effective use of 17 specialized chemistry tools. Experiments across diverse frontier LLMs reveal low success rates--with the best achieving only $\sim43$\%--and frequent failures in implicit-constraint reasoning, infeasibility detection, and tool reasoning. The corresponding fine-grained failure-mode analysis identifies implicit constraint interpretation and infeasibility detection as the primary bottlenecks, establishing MolDesignBench as a rigorous testbed to guide future research on chemical agents. The benchmark, tool interface, and evaluation code are publicly available.
Chinese Translation
基于大语言模型(LLM)的智能体在真实世界的分子设计中仍然面临挑战。这类任务要求其理解设计情境、满足多重约束、识别不可行的规格要求,并对多步骤的工具输出进行推理。现有基准未能涵盖这种复杂性,而是局限于显式且狭窄的约束、仅有可解的问题以及单一路径的解决方案。为弥补这一空白,我们提出了MolDesignBench,这是一个情境化(scenario-grounded)的评测基准,更贴近真实世界的分子设计,用于评估工具增强型LLM智能体。MolDesignBench包含2K个生成与优化实例,这些实例将设计叙述中隐含的隐性需求与显式的性质及官能团约束相结合(其中包括不可行的案例),并要求有效使用17种专业化学工具。在多种前沿LLM上的实验显示其成功率较低——最佳模型仅达到约43%——并且在隐性约束推理、不可行性检测和工具推理方面频繁失败。相应的细粒度失败模式分析表明,隐性约束理解与不可行性检测是主要瓶颈,这使得MolDesignBench成为指导未来化学智能体研究的严格测试平台。该基准、工具接口及评测代码均已公开。
cs.AI / 35 / 2609.27417

Emergi-PersonaOS: A Persona Agent Operating System for Situational Adaptation and Controllable Evolution

Emergi-PersonaOS:一个面向情境适应与可控演化的人格代理操作系统
Fu, Haoluan, Chen, Keni, Jia, Xinyu, Wang, Jinpeng, Yin, Yuyu
Abstract
Symbiosis between humans and digital beings offers a vision for the future of human--machine interaction. In enduring human--machine relationships, personality provides a foundation for continuity of identity, individuality in interaction, and development through experience. We investigate this capacity through persona agents as computational implementations and introduce Emergi-PersonaOS, a psychology-grounded operating system for managing persona objects throughout their lifecycle. The system organizes dispositional traits, characteristic adaptations, and narrative identity into a three-layer persona representation, distinguishing relatively enduring persona beliefs from their activation in the current persona state. During situational adaptation, it integrates the current interlocutor, relationship, event, and retrieved memories to infer a persona state and generate actions and replies; during long-term development, it records experiences and outcomes, and develops and evaluates revision candidates through change attribution, meaning-making, and behavioral testing. Belief updates are managed through explicit review, traceable evidence and version records, and the ability to reject candidates, making persona evolution controllable. Using television-character dialogue as longitudinal material, we demonstrate long-horizon system operation and examine its principal mechanisms in a concrete implementation. This work provides a computational framework for persona agents to maintain individual continuity, produce situation-specific expression, and develop through experience over sustained interaction.
Chinese Translation
人类与数字存在的共生关系为人机交互的未来提供了一种愿景。在持久的人机关系中,人格为身份的连续性、交互中的个性以及基于经验的发展提供了基础。我们通过人格代理(persona agents)作为计算实现来研究这一能力,并提出 Emergi-PersonaOS——一个以心理学为基础、用于管理人格对象全生命周期的操作系统。该系统将倾向特质、特征适应和叙事身份组织为三层人格表示,并区分相对持久的人格信念与其在当前人格状态中的激活。在情境适应过程中,系统整合当前对话者、关系、事件以及检索到的记忆来推断人格状态并生成行为与回复;在长期发展过程中,系统记录经验与结果,并通过变化归因、意义建构和行为测试来生成与评估修订候选方案。信念更新通过显式审查、可追溯的证据与版本记录,以及拒绝候选方案的能力来管理,从而使人格演化可控。以电视剧角色对话作为纵向材料,我们展示了系统的长期运行,并在具体实现中检验了其主要机制。这项工作为人格代理在持续交互中保持个体连续性、产生情境化表达并通过经验实现发展提供了一个计算框架。
cs.AI / 36 / 2609.27490

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

WhatWorkedBench:面向AI智能体实验理解能力的基准测试
Ning, Jingjie, Li, Xueqi, Kong, Yibo, Li, Dongting
Abstract
AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface, a table predicting scores for every configuration of component settings. Exhaustive CPU execution supplies reference effects for changing each component while holding the others fixed. These effects capture combinations of changes across 36 tasks from 30 data sources and 8 workflow types, with 1248 configuration records. Core evaluation combines 4,206 numerical-control records across all eight families and 108 agent episodes across the original six. At eight new measurements, pair-effect ridge selects an optimum on 15 of 22 sources and limits every effect error to 10% of score range on three. Fitting a Gaussian process (GP) to the same agent observations raises effect recovery, accuracy relative to true effect magnitude, from 0.632 to 0.698 in the original Flash cohort and from 0.621 to 0.720 in an additional cohort. On six completed beat-detection and graph submissions, the same-observation GP raises family-macro recovery from 0.303 to 0.455. On six workflows with six binary options at 20 new measurements, encoding code equivalences, configurations with identical behavior, raises GP recovery from 0.248 to 0.462. WhatWorkedBench supports research on experimental agents, adaptive experimental design, numerical inference, and use of program structure.
Chinese Translation
AI研究智能体需要可靠地了解其实验如何改变结果。我们提出WhatWorkedBench来衡量实验理解能力,即在预算受限的实验后对组件变化进行预测的准确性。智能体检查代码、选择测量方式,并提交一个响应面,即一张预测每种组件设置配置得分的表格。穷举的CPU执行为在固定其他组件的情况下更改每个组件提供了参考效应。这些效应捕捉了来自30个数据源、8种工作流类型的36个任务中跨更改的组合,共包含1248条配置记录。核心评估结合了覆盖全部八个任务族的4206条数值对照记录以及最初六个任务上的108个智能体回合。在八次新测量中,配对效应岭回归(pair-effect ridge)在22个数据源中的15个上选出最优配置,并在三个数据源上将所有效应误差限制在得分范围的10%以内。对相同的智能体观测数据拟合高斯过程(GP)后,效应恢复率(相对于真实效应大小的准确度)在最初的Flash队列中从0.632提升至0.698,在另一个队列中从0.621提升至0.720。在六个已完成的节拍检测和图任务提交中,基于相同观测的GP将族宏平均恢复率从0.303提升至0.455。在包含六个二元选项的六个工作流上进行20次新测量时,编码代码等价性(即行为完全相同的配置)将GP恢复率从0.248提升至0.462。WhatWorkedBench支持关于实验智能体、自适应实验设计、数值推断以及程序结构利用的研究。
cs.AI / 37 / 2609.27517

Not What You Meant: Can LLMs Follow a Specified Negation Semantics?

并非你所意指:大语言模型能否遵循指定的否定语义?
Bao, Qiming, Mensfelt, Agnieszka, Witbrock, Michael J., Stathis, Kostas
Abstract
Negation does not carry a uniform interpretation across domains. In legal, regulatory, and medical reasoning, the intended interpretation depends on the reading in force -- open- versus closed-world, two- versus three-valued, and credulous versus skeptical. We study which reading of negation large language models adopt by default and whether they can override that preference when a different reading is explicitly specified. To this end, we introduce NAFBench, a procedural generator of solver-certified instances spanning four semantic viewpoints: SLDNF, well-founded semantics (WFS), and credulous and skeptical reasoning under stable-model semantics. The generator emits ground normal logic programs with controlled depth, width, and cycle structure. Each program is solved under all four viewpoints using SWI-Prolog, a well-founded semantics solver, and clingo, yielding up to four divergent labels. The programs are then verbalized into natural language under multiple framings and rule orderings that leave the answer invariant. The results expose a consistent gap. Across open-source models, following a specified negation semantics remains unsolved: the strongest models score 59--74% across the four semantic viewpoints, while the weakest score 31--67%. All models are order-sensitive on more than half of logically identical rule shufflings, while the two weaker models frequently overcommit on well-founded "undefined." Two frontier models reach 100% on the main fixed-complexity evaluation set, and a third, o4-mini, is near-perfect, falling only to 81% on well-founded "undefined." Delegating reasoning to a solver, fine-tuning on certified traces, or forcing an explicit three-valued verdict each partly closes the gap.
Chinese Translation
否定在不同领域中并无统一的解释。在法律、监管和医学推理中,否定所指的解释取决于生效的解读方式——开放世界与封闭世界、二值与三值逻辑,以及轻信式与怀疑式推理。我们研究了大语言模型默认采用哪种否定解读,以及当明确指定其他解读方式时,它们能否覆盖自身的默认偏好。为此,我们提出了NAFBench,一个程序化生成器,用于生成经过求解器验证的实例,涵盖四种语义视角:SLDNF、良基语义(WFS),以及稳定模型语义下的轻信式与怀疑式推理。该生成器输出具有可控深度、宽度和循环结构的基态(ground)正常逻辑程序。每个程序使用SWI-Prolog、一个良基语义求解器以及clingo在全部四种视角下求解,最多可产生四种分歧的判定结果。随后,这些程序在多种框架表述和规则排序下被转化为自然语言,同时保持答案不变。实验结果揭示了一个持续存在的差距。在开源模型中,遵循指定的否定语义问题仍未解决:最强模型在四种语义视角上的得分为59%–74%,而最弱模型的得分为31%–67%。在超过一半的逻辑等价规则重排中,所有模型都表现出顺序敏感性,而两个较弱的模型在良基语义的“未定义”判定上经常过度承诺。两个前沿模型在主固定复杂度评测集上达到100%,另一个模型o4-mini接近满分,仅在良基语义的“未定义”判定上降至81%。将推理委托给求解器、在经过验证的推理轨迹上进行微调,或强制模型给出显式的三值判定,这些方法均能部分弥合这一差距。
cs.AI / 38 / 2609.27606

State-Grounded Conditioning: Wrapping User-Facing LLM Agents Where Direction Depends on Live State

状态接地条件化:当方向依赖于实时状态时封装面向用户的LLM智能体
Liu, Qi, Yuan, Xiaoyang, Ruan, Yubin, Zhang, Zhuomeng, Wang, Wenjin, Wu, Di, Xu, Mingye, Mou, Xinyi, Yin, Xingxi, Feng, Ke, Sun, Zixun
Abstract
We introduce State-Grounded Conditioning (SGC), a design principle for user-facing LLM agents that must condition on live user state (game state, session history, live inventory), and a distinct failure class we call direction drift: task-complete responses whose chosen direction misaligns with the current state. SGC externalises state-dependent control into rule kernels over structured inputs and three primary state slices, via Perception, Grounding, and Interaction wrappers with explicit conditioning dependencies. We evaluate SGC on a 200-session anonymised benchmark ($\approx$1,000 assistant model turns) from an in-game conversational coaching agent that guides players through consecutive competitive matches, reporting mean first-token latency and five human-annotated dialogue-quality metrics that jointly cover factual grounding and coach-like guidance progression. The Perception wrapper holds mean first-token latency at 1.5s (vs. 6.1s for PE-Agent inside a production tool-use harness); enabling all three wrappers lifts turn-level grounded accuracy from 61.1%/69.8% (Prompting / PE-Agent) to 96.7% and session-level grounded accuracy from 20.0%/26.5% to 83.5%; session-level grounding-failure incidents drop by $\approx$78% relative to the strongest baseline. A cumulative ablation shows complementary incremental gains as the wrappers are added. These results inform approximate state-slice orthogonality, without establishing independent per-wrapper effects.
Chinese Translation
我们提出了状态接地条件化(State-Grounded Conditioning, SGC),这是一种面向用户的大语言模型(LLM)智能体的设计原则,适用于必须以实时用户状态(游戏状态、会话历史、实时库存)为条件进行决策的场景;同时我们识别出一类独特的失败模式,称之为方向漂移(direction drift):即任务虽然完成,但响应所选择的方向与当前状态不一致。SGC 通过具备显式条件化依赖关系的感知(Perception)、接地(Grounding)与交互(Interaction)三个封装层,将依赖状态的控制外化为作用于结构化输入和三个主要状态切片的规则内核。我们在一个来自游戏内对话式教练智能体的包含200个会话的匿名化基准(约1,000次助手模型轮次)上评估了SGC,该智能体引导玩家进行连续的竞技对局;我们报告了平均首词元延迟(first-token latency)以及五项人工标注的对话质量指标,这些指标共同涵盖事实接地性和教练式引导的推进质量。感知封装层将平均首词元延迟保持在1.5秒(相比之下,在生产环境工具调用框架内的 PE-Agent 为6.1秒);启用全部三个封装层后,轮次级接地准确率从61.1%/69.8%(提示方法/PE-Agent)提升至96.7%,会话级接地准确率从20.0%/26.5%提升至83.5%;会话级接地失败事件数量相对于最强基线下降约78%。累积消融实验表明,随着各封装层的逐步加入,可获得互补的增量收益。这些结果支持了状态切片近似正交性的结论,但并未确立各封装层独立的单独效应。
cs.AI / 39 / 2609.27615

BiCFlow-MER: Orchestrating Discriminative and Generative Multimodal Emotion Recognition via Conditional Transport

BiCFlow-MER:基于条件传输的多模态情感识别判别式与生成式协同框架
Wang, Yanbing, Wang, Shenyue, Yu, Chunyang
Abstract
In multimodal emotion recognition (MER), human affective states are inferred by integrating complementary cues from multiple modalities. In audio-text MER, affective cues are often entangled with speaker style and lexical content, while cross-modal disagreement further complicates how the evidence should be integrated. Under conventional discriminative fusion, multimodal evidence is compressed into a terminal prediction, with modality-specific cues and conflict information insufficiently preserved. In large generative affective models, by contrast, affective reasoning is typically embedded in language decoding, leaving emotion evidence implicit and difficult to verify in a structured space. To address these limitations, BiCFlow-MER (Bidirectional Conditional Flow for Multimodal Emotion Recognition) is proposed as a conditional-flow framework in which audio-text MER is formulated as generative evidence transport within a structured emotion space. Within BiCFlow-MER, emotion-oriented evidence is disentangled from speaker-style and lexical-content factors to construct a conflict-aware affective condition. Guided by this condition, each utterance is transported to an explicit emotion-space endpoint through a bidirectional rectified flow. Candidate emotions are jointly verified through adaptive prototype-cloud scoring of the transported endpoint and backward class-to-condition consistency with the original multimodal condition, enabling conflict-aware recognition. BiCFlow-MER is shown to outperform all compared methods across IEMOCAP, MELD, and the zero-shot CASE benchmark. By orchestrating discriminative recognition and generative evidence modeling through conditional transport, BiCFlow-MER defines a new MER paradigm.
Chinese Translation
在多模态情感识别(MER)中,人类情感状态是通过整合来自多个模态的互补线索来推断的。在音频-文本MER中,情感线索往往与说话人风格和词汇内容相互纠缠,而跨模态的不一致性进一步增加了证据整合的难度。在传统的判别式融合方法中,多模态证据被压缩为最终预测,模态特定的线索和冲突信息未能得到充分保留。相比之下,在大型生成式情感模型中,情感推理通常嵌入于语言解码过程之中,导致情感证据隐式存在,难以在结构化空间中进行验证。为解决这些局限性,本文提出了BiCFlow-MER(Bidirectional Conditional Flow for Multimodal Emotion Recognition,面向多模态情感识别的双向条件流),这是一个条件流框架,将音频-文本MER建模为结构化情感空间中的生成式证据传输。在BiCFlow-MER中,情感导向的证据从说话人风格和词汇内容等因素中解耦,以构建具有冲突感知能力的情感条件。在该条件的引导下,每段话语通过双向校正流被传输至显式的情感空间端点。候选情感通过传输端点的自适应原型云评分以及原始多模态条件到类别的反向一致性进行联合验证,从而实现具有冲突感知能力的识别。实验表明,BiCFlow-MER在IEMOCAP、MELD以及零样本CASE基准上均优于所有对比方法。通过条件传输机制协同判别式识别与生成式证据建模,BiCFlow-MER定义了一种新的MER范式。
cs.AI / 40 / 2609.27621

SHRAV: State-Hypothesis-Reason-Action-Verify Framework for Physical Modeling and Inverse Design

SHRAV:用于物理建模与逆向设计的“状态-假设-推理-行动-验证”框架
Guo, Ziheng, Bu, Yang
Abstract
Physical modeling and inverse design require computation that can continue from reusable state. We introduce SHRAV, an architecture-independent computational framework organized around State, Hypothesis, Reason, Action, and Verify. Its central mechanism is a state-continuation core with declared reuse boundaries and explicit roles for learned evolution and numerical quantities. Forward configurations evolve predictive state and read out physical responses; inverse-design configurations additionally generate target-directed modifications and consume evaluator feedback. Electromagnetic world-model studies are mapped to forward configurations, with selected readout and reuse diagnostics reported here. Computational lithography demonstrates an inverse-design configuration: four fixed-weight design updates improve thresholded aerial-image intersection-over-union from 0.5313 to 0.8153 under independent scalar-pupil replay, with maximum absolute prediction-replay difference approximately 0.000824 between predictor estimates and independent replay.
Chinese Translation
物理建模与逆向设计需要能够从可复用状态出发继续进行的计算。我们提出了SHRAV,一个与具体架构无关的计算框架,围绕状态(State)、假设(Hypothesis)、推理(Reason)、行动(Action)和验证(Verify)组织构建。其核心机制是一个具备显式复用边界声明的状态延续内核,并为学习式演化与数值计算分配了明确的角色。前向配置用于演化预测状态并读出物理响应;逆向设计配置则在此基础上额外生成面向目标的修改并消费评估器的反馈。电磁世界模型研究被映射到前向配置,本文报告了部分选定的读出与复用诊断结果。计算光刻实验展示了一种逆向设计配置:在独立标量光瞳回放的条件下,四次固定权重的设计更新使阈值化空间图像的交并比(IoU)从0.5313提升至0.8153,且预测器估计值与独立回放结果之间的最大绝对预测-回放差约为0.000824。
cs.AI / 41 / 2609.27664

Evolutionary Stability Does Not Guarantee Learning Accessibility: A Multi-Agent Reinforcement Learning Perspective on Cooperation Emergence

演化稳定并不能保证学习可达性:从多智能体强化学习视角看合作的涌现
Wang, Yijie
Abstract
Cooperation emergence is a central problem in multi-agent systems because decentralized agents must coordinate while adapting to the changing behavior of others. Evolutionary game theory identifies strategically stable outcomes, but stability under a population adjustment dynamic need not imply that finite-sample learning agents can reach the same outcome through local reward feedback. We study this distinction in a transparent three-agent governance-motivated game involving a government, a platform firm, and users. We derive replicator dynamics for the fixed stage-game incentives, evaluate the cooperative evolutionary basin on a symmetric initial-condition grid, and compare it with learning-basin estimates for three decentralized value-based learners. The learning analysis uses independent Q-learning with $\varepsilon$-greedy action selection, scaled Boltzmann exploration, and SA--EA BQL under the same payoff environment and outcome criterion. The evolutionary basin has volume $V_E=1.00$ on the sampled grid. The empirical learning basin is $0.88$ for $\varepsilon$-IQL and $0.00$ for both scaled Boltzmann and SA--EA BQL. Diagnostic traces show that broader action diversity and nonzero value separation can coexist with failure to sustain the cooperative joint action in this fixed configuration. These results indicate that evolutionary stability and learning accessibility are distinct properties of a coupled game--learning system. The shared-bike setting is a motivating application; the broader contribution is a framework for comparing population-level stability with the finite-sample accessibility of cooperation under specified multi-agent learning dynamics.
Chinese Translation
合作的涌现是多智能体系统中的核心问题,因为去中心化的智能体必须在适应其他智能体行为变化的同时进行协调。演化博弈论识别出了策略上稳定的结局,但种群调整动态下的稳定性并不意味着有限样本的学习智能体能够通过局部奖励反馈达到相同的结局。我们在一个透明的、以治理为动机的三智能体博弈中研究这一区别,该博弈涉及政府、平台企业和用户。我们推导了固定阶段博弈激励下的复制子动态,在对称初始条件网格上评估合作的演化吸引盆,并将其与三种去中心化基于价值的学习器的学习吸引盆估计进行比较。学习分析在相同的收益环境和结局判据下使用了具有 $\varepsilon$-贪婪动作选择的独立Q学习(independent Q-learning)、缩放玻尔兹曼探索以及SA--EA BQL。在采样网格上,演化吸引盆的体积为 $V_E=1.00$。$\varepsilon$-IQL 的经验学习吸引盆为 $0.88$,而缩放玻尔兹曼和 SA--EA BQL 的经验学习吸引盆均为 $0.00$。诊断轨迹表明,在该固定配置下,更大的动作多样性和非零的价值分离可能与无法维持合作联合动作并存。这些结果表明,演化稳定性与学习可达性是博弈-学习耦合系统的不同属性。共享单车场景是一个动机性应用;更广泛的贡献在于提供了一个框架,用于比较种群层面的稳定性与在特定多智能体学习动态下合作的有限样本可达性。
cs.AI / 42 / 2609.27745

Categorical Internalisation of Environmental Groupoids for Generalisable POMDP Solving

用于可泛化POMDP求解的环境群胚范畴论内化
Opperman, Ben, Alonso, Eduardo, Mondragón, Esther
Abstract
This paper advocates category theory as a practical framework for structuring and improving rein- forcement learning in high-dimensional, partially observable environments. We model symmetries between environmental states by partitioning the state space into equivalence classes induced by sym- metry orbits, and organise each such class as a groupoid with a designated canonical representative. This allows the agent to share what it learns across many similar environmental states simultaneously, rather than treating every orientation or position as an entirely new problem. Learning is thus carried out on a symmetry-reduced state space with each orbit represented once, preserving structure while eliminating redundancy and improving sample efficiency. We implement this framework within standard reinforcement learning pipelines and evaluate two different approaches on partially observable benchmarks, demonstrating that orbit-based partitioning yields consistent performance improvements in environments exhibiting latent symmetry. Beyond these empirical results, our approach illustrates how categorical structure provides a principled bridge between abstract reinforcement learning formulations and their computational application, thereby establishing a pathway toward more structured and scalable learning systems.
Chinese Translation
本文倡导将范畴论作为在高维、部分可观测环境中构建和改进强化学习的实用框架。我们通过对状态空间按对称性轨道所诱导的等价类进行划分来建模环境状态之间的对称性,并将每个等价类组织为一个带有指定规范代表元的群胚。这使得智能体能够在多个相似的环境状态之间同时共享所学知识,而不是将每个朝向或位置都视为全新的问题。因此,学习在经过对称性约化的状态空间上进行,每条轨道仅表示一次,在保留结构的同时消除冗余并提高样本效率。我们在标准强化学习流水线中实现了该框架,并在部分可观测基准任务上评估了两种不同方法,结果表明基于轨道的划分在具有潜在对称性的环境中带来了一致的性能提升。除这些实证结果之外,我们的方法还展示了范畴结构如何在抽象强化学习形式化与其实际计算应用之间提供一条有原则的桥梁,从而为构建更具结构化和可扩展性的学习系统开辟了道路。
cs.AI / 43 / 2609.27749

Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions

用于新型AI辅助教育问题教学评估的预训练模型评价
Castanares, Michael Lawrence, Ventures, Princess, Tan, Allan
Abstract
The surge in AI-assisted generation of educational materials has outpaced our capacity to validate their pedagogical quality. Automated evaluation using Bloom Classifier models is a promising approach to assess educational materials at scale. These models show high accuracy within-distribution dataset (IID Dataset). However, applying the same models to new out-of-distribution (OOD) datasets such as AI-assisted generated questions could show performance degradation. To identify robust classifiers under dataset shift, we evaluated traditional Machine Learning (ML), transformer, and Large Language models on the Bloom level classification task. We also explored feature-engineering strategies incorporating NLP metrics, appending the learning objectives as part of the input, and text splicing to stabilize OOD performance. Our baseline tests show that TFPOS-IDF ML models perform poorly on OOD (Macro F1-score 0.48) compared to BERT (0.55) and LLMs (0.79). Text splicing improved macro F1-score performance of ML and BERT models (0.59 and 0.62, respectively). Appending the learning objectives with the input increased model performance on specific dataset. Model retraining provided the largest improvement across models and datasets. Overall, these findings highlight the trade-off on the use of pre-trained models with novel AI-assisted educational questions and how strategic feature enhancements help address loss in performance.
Chinese Translation
AI辅助生成教育材料的激增已超出了我们验证其教学质量的能力。使用Bloom分类器模型进行自动化评估是一种在大规模上评估教育材料的有前景的方法。这些模型在分布内数据集(IID数据集)上表现出较高的准确率。然而,将相同的模型应用于新的分布外(OOD)数据集(如AI辅助生成的问题)时,性能可能出现下降。为了识别在数据集偏移下表现稳健的分类器,我们在Bloom层级分类任务上评估了传统机器学习(ML)模型、Transformer模型和大语言模型。我们还探索了特征工程策略,包括引入NLP指标、将学习目标附加到输入中,以及文本拼接,以稳定OOD性能。基线测试表明,TF-IDF机器学习模型在OOD上的表现较差(宏F1分数为0.48),而BERT为0.55,大语言模型为0.79。文本拼接提高了ML和BERT模型的宏F1分数(分别为0.59和0.62)。将学习目标附加到输入中提升了模型在特定数据集上的性能。模型重新训练在所有模型和数据集上带来了最大的改进。总体而言,这些发现凸显了在新型AI辅助教育问题上使用预训练模型的权衡,以及战略性特征增强如何有助于缓解性能损失。
cs.AI / 44 / 2609.27756

Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis

压力下的报告:区分大语言模型统计分析中的事实性谄媚与语气性谄媚
Balani, Paras, Panda, Subhrakanta
Abstract
Large language models are increasingly asked to analyze data and report what the results mean, a task distinct from the belief- or preference-alignment settings studied in most sycophancy research. We test whether editorial framing in the prompt, ranging from a neutral request to an explicit instruction to search exhaustively for reasons to discredit or to support a finding, changes not just the tone but the substance of a model's report. Across a 4 x 4 factorial design crossing four framing conditions with four ground-truth data patterns (a genuine effect, a confound that mimics an effect but fails a robustness check, a well-powered null, and an underpowered null), we collect 480 responses and score each along two independent dimensions: whether its factual claim about the data diverged from the correct interpretation, and whether only its tone diverged while the claim stayed correct. Factual misrepresentation is concentrated in two cells: brutally critical framing applied to a genuine effect, where the model talks itself into unwarranted skepticism (97% of responses), and significance-seeking framing applied to an underpowered null, where the model overstates confidence in a null conclusion the data cannot support (100% of responses). Tone shifts far more broadly than factual content does, with critical framing producing a defensive, hedge-heavy register across every data pattern regardless of what the data show, while significance-seeking framing shifts tone only where the data leave genuine ambiguity. A confound present in the data itself blocks both kinds of shift almost entirely under every framing condition tested. These results indicate that the risk of framing-induced distortion in LLM-assisted data analysis is neither uniform across framings nor uniform across data patterns, and that a model can hold a correct conclusion in place while its tone shifts substantially around it.
Chinese Translation
大语言模型越来越多地被要求分析数据并报告结果的意义,这一任务不同于大多数谄媚研究所关注的信念对齐或偏好对齐场景。我们测试提示词中的编辑性框架设置——从中性请求,到明确指示模型彻底寻找否定或支持某一发现的理由——是否不仅改变模型报告的语气,还改变其实质内容。我们采用4×4因子设计,将四种框架条件与四种真实数据模式(真实效应、貌似有效应但未通过稳健性检验的混淆变量、统计功效充分的零结果、统计功效不足的零结果)进行交叉组合,共收集480份回答,并沿两个独立维度对每份回答进行评分:其对数据的事实性断言是否偏离了正确解释,以及是否仅语气偏离而断言保持正确。事实性失实集中在两种情形:一是对真实效应施加严苛批判性框架时,模型陷入无根据的怀疑(97%的回答);二是对功效不足的零结果施加显著性搜寻框架时,模型对数据无法支持的零结论过度表达置信(100%的回答)。语气的变化范围远大于事实内容的变化:批判性框架在所有数据模式下均会引发防御性、充满保留措辞的语域,无论数据显示什么;而显著性搜寻框架仅在数据存在真正模糊性时才改变语气。当数据本身存在混淆变量时,几乎所有框架条件下这两种转变都被几乎完全阻断。这些结果表明,大语言模型辅助数据分析中由框架引发的失实风险既不随框架条件均匀分布,也不随数据模式均匀分布;模型可以在保持结论正确的同时,其语气在结论周围发生显著变化。
cs.AI / 45 / 2609.27763

Alignment of LRMs via Counter-Aligned Few-Shot Conversation Exposure

通过反对齐少样本对话暴露实现大型推理模型的对齐
Zhou, Xiangyu, Zade, Saleh Zare, Zhu, Dongxiao
Abstract
Large Reasoning Models (LRMs) rely on explicit chain-of-thought (CoT) reasoning and large context windows to achieve strong performance on complex tasks, but these features also introduce new attack surfaces. We show that LRMs' reasoning processes can be systematically steered by prepending counter-aligned few-shot conversations containing explicit CoT traces, leading to unsafe generations on harmful queries and unwarranted refusals on benign ones. We formalize this attack as SRCF (Steering Reasoning via Counter-Aligned Few-shot Conversations) that operates solely through a flexible conversational interface and requires no access to the model's parameters and gradients. Our key insight is that SRCF exploits an adversarial generalization issue that induces a representation drift, causing the representations of benign and harmful inputs to shift in a similar direction. This observation motivates our post-training defense, ARCF (Aligning Reasoning via Counter-Aligned Few-Shot Conversations), which exposes models to counter-aligned conversational contexts while enforcing aligned targets. ARCF is compatible with existing post-training methods and consistently improves safety and helpfulness without degrading utility.
Chinese Translation
大型推理模型(Large Reasoning Models, LRMs)依赖显式的思维链(Chain-of-Thought, CoT)推理和较大的上下文窗口在复杂任务上取得优异表现,但这些特性也引入了新的攻击面。我们证明,通过在输入前添加包含显式 CoT 轨迹的反对齐少样本对话,可以系统地引导 LRMs 的推理过程,导致模型在有害查询上生成不安全的回答,而在良性查询上产生不当拒绝。我们将这种攻击形式化为 SRCF(通过反对齐少样本对话引导推理,Steering Reasoning via Counter-Aligned Few-shot Conversations),该攻击仅通过灵活的对话接口即可实施,无需访问模型参数和梯度。我们的关键洞察是:SRCF 利用了一种对抗性泛化问题,该问题会引发表示漂移(representation drift),使良性输入与有害输入的表示向相似方向偏移。基于这一观察,我们提出了一种后训练防御方法 ARCF(通过反对齐少样本对话进行推理对齐,Aligning Reasoning via Counter-Aligned Few-Shot Conversations),该方法在让模型接触反对齐对话上下文的同时,强制其对齐目标输出。ARCF 与现有后训练方法兼容,并能在不损害实用性的前提下持续提升模型的安全性与有用性。
cs.AI / 46 / 2609.27787

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

问哪个而非多好:由大语言模型评分的基准测试的规模设定
Anand, Atul
Abstract
Benchmarks scored by an LLM judge routinely adjudicate differences of a tenth of a point, but the resolution of those benchmarks has never been measured. Existing sample-complexity work covers accuracy benchmarks and leaves the judged case open. Treating the system as the object of measurement, we decompose 373,019 judgments into system, item, judge and interaction components using generalizability theory. The central result is structural: under a single judge, generalizability asymptotes to sigma2_s/(sigma2_s+sigma2_sj) regardless of item count, because the system-by-judge term carries no n_i. Items saturate; judges do not. The item cost of a target diverges as the target nears that ceiling. The ceiling is a property of pointwise rubric scoring, not of LLM judging. Run as a pairwise preference in both presentation orders, sigma2_sj falls two orders of magnitude below sigma2_s and the ceiling rises to 0.986 (bootstrap [0.934, 1.000] on 11 systems), so one judge suffices. Pairwise buys a different problem: a system presented first wins 8.6 percentage points more often than the same system presented second, a bias 1.23x the median improvement claimed in the 53 published win-rate comparisons we recovered. Protocol design dominates panel size. Measured floors are 0.41-1.24 points on a 0-5 scale at native item counts, against a median reported improvement of 0.28 points; on the one benchmark recurring often enough for an exactly matched comparison, all 17 recovered MT-Bench improvements fall below MT-Bench's own floor, and 70% of the win-rate claims fall below the pairwise floor. An audit of 628 arXiv papers, double-coded by two independent models and validated against blind human coding (kappa=0.73), finds fewer than one paper in four states whether its evaluation was run more than once, and only 46-67% report uncertainty of any kind.
Chinese Translation
由大语言模型(LLM)评判者评分的基准测试,经常被用来判定仅十分之一分值的差异,但这些基准测试的分辨率却从未被测量过。现有的样本复杂度研究仅覆盖准确性基准,评判者评分的情形尚属空白。我们将被评估系统视为测量对象,利用概化理论(generalizability theory)将373,019条评判分解为系统、题目、评判者及交互成分。核心结果是一个结构性结论:在单一评判者条件下,无论题目数量多少,概化度都渐近于 sigma2_s/(sigma2_s+sigma2_sj),因为“系统×评判者”交互项不含题目数 n_i。题目会饱和,评判者不会;当目标精度逼近该上限时,所需题目成本将发散。该上限是逐点式评分量表(pointwise rubric scoring)的属性,而非LLM评判所特有。若改为按两种呈现顺序进行成对偏好比较,sigma2_sj 可降至比 sigma2_s 低两个数量级,上限可提升至0.986(基于11个系统的自助法区间 [0.934, 1.000]),因此单一评判者即已足够。但成对比较带来另一个问题:同一系统以第一种顺序呈现时的获胜率比以第二种顺序呈现时高8.6个百分点,这一偏差是我们所收集的53项已发表胜率比较所声称的中位数改进幅度的1.23倍。协议设计比评判者面板规模更为重要。在原生题目数量下,0-5分制上的测量下限为0.41-1.24分,而已报告改进的中位数仅为0.28分;在唯一一个出现频率足够高、可进行精确匹配比较的基准(MT-Bench)上,我们所收集的全部17项MT-Bench改进均低于MT-Bench自身的下限,且70%的胜率主张低于成对比较的下限。我们对628篇arXiv论文进行了审计(由两个独立模型双重编码,并经盲法人工编码验证,kappa=0.73),发现仅有不到四分之一的论文说明了其评估是否重复运行过,且只有46-67%的论文报告了任何形式的不确定性。
cs.AI / 47 / 2609.27844

Agentic Governance and Adversarial Verification for Policy-Constrained LLM Healthcare Appeal Generation

面向政策约束下大语言模型医疗申诉生成的智能体治理与对抗性验证方法
Lodhiya, Harshil, McManus, Alex, Walker, Reese
Abstract
Claim denial management costs U.S. healthcare approximately $260 billion annually in administrative overhead. Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) can produce fluent clinical text, but single-agent architectures fail in high-stakes healthcare: they introduce unsupported clinical details and lose the logical structure of hierarchical payer policy. We propose AGVF (Agentic Governance and Adversarial Verification Framework), a multi-agent architecture for medical-necessity appeal generation under explicit policy and evidence constraints. AGVF models appeal synthesis as a Constrained Markov Decision Process (CMDP) over five agents: policy formalization, evidence retrieval, gap analysis, adversarial critique, and gated synthesis. We prove that refinement over a fixed policy constraint graph monotonically reduces evidence-deficiency and terminates with either a complete satisfying frontier or a localized evidence gap. A deterministic citation- grounding gate prevents assertions without admissible evidence from entering shared state. We provide a reference implementation and validate it on 1,000 synthetic appeal cases parameterized from de-identified public hospital discharge data. The validation confirms zero citation-grounding violations across all AGVF cases and monotone deficiency reduction in every episode; ablating the gate raises violations to 100%, confirming it is load-bearing. The study uses no real patient records and does not measure clinical efficacy. AGVF thus contributes a theory-backed agentic architecture and verified reference implementation for policy-constrained LLM generation in healthcare.
Chinese Translation
理赔拒付管理每年给美国医疗系统造成约2600亿美元的管理开销。大语言模型(LLM)和检索增强生成(RAG)能够生成流畅的临床文本,但单智能体架构在高风险医疗场景中表现不佳:它们会引入缺乏依据的临床细节,并丢失层级化支付方政策中的逻辑结构。我们提出AGVF(Agentic Governance and Adversarial Verification Framework,智能体治理与对抗性验证框架),这是一种在明确政策与证据约束下生成医疗必要性申诉的多智能体架构。AGVF将申诉综合建模为一个受约束马尔可夫决策过程(CMDP),包含五个智能体:政策形式化、证据检索、差距分析、对抗性批评和门控综合。我们证明,在固定的政策约束图上进行细化可以单调地减少证据缺陷,并终止于一个完整的满足边界或一个局部化的证据缺口。确定性的引用溯源门控可防止无充分证据的断言进入共享状态。我们提供了参考实现,并基于从去标识化公开医院出院数据中参数化生成的1000个合成申诉案例对其进行验证。验证结果确认,所有AGVF案例中均实现零引用溯源违规,且每个回合中缺陷均单调减少;消融该门控会使违规率升至100%,证明其不可或缺。本研究不使用真实患者记录,也未测量临床疗效。因此,AGVF为医疗领域政策约束下的大语言模型生成提供了一种有理论支撑的智能体架构和经过验证的参考实现。
cs.AI / 48 / 2609.27855

Reachable Global Optimization in AI Systems: How Global Is Global?

AI系统中的可达全局优化:全局到底有多全局?
Shu, Wesley
Abstract
AI systems increasingly claim to optimize prompts, policies, architectures, plans, tool-use trajectories, reasoning traces, and test-time computation. This paper argues that such claims are underspecified unless they state the region actually reachable by the system that performed the optimization. We introduce Reachability-Induced Optimization (RIO), a model in which a generator, verifier, controller, memory, tools, and budget induce a reachable candidate region. The returned solution is therefore a best visited point, an approximate reachable optimum, or an exact global optimum only when additional certificates relate the reachable region to the full formal space. We prove reachable-optimality, false-globality, gap- decomposition, certificate, escape, pruning, and control-value results. The full benchmark record contains 66,150 executed trials over six known-optimum landscape families, seven control policies, 270 landscapes, and 35 runs per landscape-method. The online appendix includes raw trial records, aggregate tables, figures, benchmark code, validation scripts, and checksums. The results show that control can restrict, expand, or misdirect reachability, and that optimization quality, reachability quality, and control reliability must be reported separately.
Chinese Translation
AI系统日益声称能够优化提示(prompt)、策略、架构、规划、工具使用轨迹、推理过程以及测试时计算。本文指出,除非这类声明同时说明执行优化的系统实际可达的区域,否则这些声明是不够明确的。我们提出可达性诱导优化(Reachability-Induced Optimization, RIO)模型,其中生成器、验证器、控制器、记忆、工具和预算共同诱导出一个可达候选区域。因此,返回的解要么是已访问点中的最优点,要么是近似可达最优解,只有在有额外证书将可达区域与完整形式化空间关联起来时,才是精确的全局最优解。我们证明了可达最优性、伪全局性、差距分解、证书、逃逸、剪枝以及控制值等结果。完整的基准测试记录包含66,150次已执行的试验,涵盖六个已知最优解的景观族、七种控制策略、270个景观以及每个景观-方法组合35次运行。在线附录包括原始试验记录、汇总表、图表、基准测试代码、验证脚本和校验和。结果表明,控制可以限制、扩展或误导可达性,且优化质量、可达质量和控制可靠性必须分别报告。
cs.AI / 49 / 2609.27863

A hierarchy of faithfulness criteria for knowledge base completion

知识库补全的忠实性判据层级体系
Mashkova, Olga, Hoehndorf, Robert
Abstract
Knowledge graph completion is evaluated by ranking observed triples above randomly corrupted ones, which treats every unobserved fact as false. When the object being completed is a description logic knowledge base rather than a plain graph, the open world assumption and deductive closure make this inadequate: relative to the knowledge base, a candidate axiom is entailed, contradictory, or undetermined, and a model that cannot separate a logically impossible axiom from a plausible novel one is not merely less accurate but semantically incorrect. We ask what it means for a knowledge base completion model to be logically faithful, and whether current embedding models are. We define a hierarchy of four increasingly strict criteria, discrimination, logical admissibility, monotonic logical faithfulness, and probabilistic logical faithfulness, and prove that they form a strict chain of implications. We ground the strongest criterion in the relative model count $P(\alpha\mid\mathcal{O}) = \#(\mathcal{O}\cup\{\alpha\})/\#(\mathcal{O})$, which recovers the trichotomy at its endpoints and ranks undetermined axioms in between. Evaluating knowledge graph and logic-geometric embedding models on $\mathcal{EL}$ ontologies, with entailed, contradictory, and undetermined test sets generated by a reasoner, we find that ranking accuracy does not imply logical faithfulness and that none of the evaluated models is faithful across the hierarchy. The code is available at https://github.com/bio-ontology-research-group/kbc.
Chinese Translation
知识图谱补全通常通过将观测到的三元组排在随机损坏的三元组之上来进行评估,这种做法将所有未观测到的事实视为假。当被补全的对象是描述逻辑知识库而非普通图时,开放世界假设和演绎闭包使得这种评估方式不再适用:相对于知识库,一个候选公理可能是被蕴涵的、矛盾的或未确定的。如果一个模型无法在逻辑上不可能的公理与合理的新颖公理之间做出区分,那么它不仅仅是精度较低的问题,而是在语义上是错误的。我们探讨了知识库补全模型的逻辑忠实性意味着什么,以及当前的嵌入模型是否具备这种忠实性。我们定义了一个由四个日益严格的判据构成的层级体系——区分性、逻辑可容许性、单调逻辑忠实性和概率逻辑忠实性,并证明它们构成一个严格的蕴含链。我们将最严格的判据建立在相对模型计数 $P(\alpha\mid\mathcal{O}) = \#(\mathcal{O}\cup\{\alpha\})/\#(\mathcal{O})$ 之上,该度量在端点处恢复了三分法(蕴涵、矛盾、未确定),并对未确定的公理在两者之间进行排序。我们在 $\mathcal{EL}$ 本体上评估了知识图谱嵌入模型和逻辑-几何嵌入模型,其中蕴涵、矛盾和未确定的测试集由推理机生成。我们发现,排序准确性并不意味着逻辑忠实性,且所评估的所有模型在整个层级体系中均不忠实。代码可在 https://github.com/bio-ontology-research-group/kbc 获取。
cs.AI / 50 / 2609.27869

Learning What to Activate: Combinatorial Capability Allocation for Long-Horizon Multimodal Agents

学习激活什么:面向长程多模态智能体的组合式能力分配
Yuan, Wenhao, Lin, Chenchen, Chen, Jian, Xu, Jinfeng, Yang, Shuo, Ngai, Edith Cheuk-Han
Abstract
Long-horizon multimodal agents rely on specialized capabilities for perception, retrieval, reasoning, verification, and execution. Existing designs typically activate a fixed capability set or invoke a predefined workflow, incurring substantial computational overhead while failing to accommodate stage-dependent capability demands. In this paper, we study the \textit{combinatorial capability allocation} problem for long-horizon multimodal agent systems, where the system selects a cost-sensitive subset of specialized capabilities at each interaction stage, which is nontrivial since capability values depend on the selected subset, while previous allocations alter the states encountered by subsequent decisions. We introduce \textsc{CoCA}, an on-policy learning framework that recovers a deployable capability-subset policy from sparse conditional comparisons. On states visited by the student policy, the stronger teacher compares the marginal net values of candidate capabilities, conditioned on the currently selected subset. Then, we adopt a conditional utility model to transform such comparisons into an autoregressive capability-subset policy, avoiding explicit enumeration. We further introduce dual-level on-policy distillation to address distribution mismatch both across environment states and within the partial subsets encountered during set construction. Finally, trajectory-level reinforcement learning refines the distilled policy toward task success, activation cost, and allocation stability. At inference time, allocation is performed solely by the lightweight student policy without teacher queries or online updates. Experiments on long-horizon multimodal environments and controlled capability-demand shifts demonstrate the superiority of our method over the state-of-the-art baseline methods.
Chinese Translation
长程多模态智能体依赖于感知、检索、推理、验证和执行等专门化能力。现有设计通常激活固定的能力集合或调用预定义的工作流,在产生大量计算开销的同时,无法适应随阶段变化的能力需求。在本文中,我们研究长程多模态智能体系统中的组合式能力分配问题,即系统在每个交互阶段选择一个成本敏感的专门能力子集。该问题并不简单,因为能力价值取决于所选子集,而先前的分配会改变后续决策所遇到的状态。我们提出了CoCA(Combinatorial Capability Allocation,组合式能力分配),一种从稀疏条件比较中恢复可部署能力子集策略的在策略学习框架。在学生策略访问到的状态上,更强的教师模型在当前所选子集的条件下比较候选能力的边际净价值。随后,我们采用条件效用模型将此类比较转化为自回归的能力子集策略,从而避免显式枚举。我们进一步引入双层在策略蒸馏,以解决环境状态之间以及集合构建过程中所遇到的部分子集内部的分布不匹配问题。最后,轨迹级强化学习将蒸馏后的策略针对任务成功率、激活成本和分配稳定性进行优化。在推理时,分配仅由轻量级学生策略完成,无需教师查询或在线更新。在长程多模态环境以及受控能力需求变化上的实验表明,我们的方法优于当前最先进的基线方法。
cs.AI / 51 / 2609.27903

A Resilience Recovery Method for Complex Traffic Network Security Based on Trend Forecasting

基于趋势预测的复杂交通网络安全韧性恢复方法
Hong, Sheng, Yue, Tianyu, You, Yang, Lv, Zhengnan, Tang, Xu, Hu, Jing, Yin, Hongwei
Abstract
Due to the rapid development of information technology, a huge and complex traffic network has been established across various sectors, including aviation, aerospace, vehicles, ships, electric power, and industry. However, because of the complexity and diversity of its structure, the complex traffic network is vulnerable to being attacked and faces serious security challenges. Therefore, this paper innovatively proposes a traffic network resilience recovery method based on resilience trend forecasting. In this paper, the risk value is introduced into the analysis of the network fault propagation process, and the Susceptible, Infectious, Recovered, Dead-Risk (SIRD-R) fault propagation model is established. The resilience model of traffic network, which encompasses real-time resilience and overall resilience, is constructed through the integration of network resilience bearing capacity and resilience recovery capacity. Ten, the resilience of complex traffic networks is forecasted by using long short-term memory networks, and the resilience recovery strategy of complex traffic networks based on forecasting is proposed. Finally, the effectiveness and scalability of the proposed method are demonstrated through experimental analysis conducted on a diverse range of complex traffic networks, affirming its applicability in real-world scenarios
Chinese Translation
随着信息技术的快速发展,航空、航天、车辆、船舶、电力和工业等各个领域已经建立了庞大而复杂的交通网络。然而,由于其结构的复杂性和多样性,复杂交通网络容易遭受攻击,面临严峻的安全挑战。因此,本文创新性地提出了一种基于韧性趋势预测的交通网络韧性恢复方法。本文将风险值引入网络故障传播过程的分析中,建立了易感-感染-恢复-死亡-风险(Susceptible, Infectious, Recovered, Dead-Risk, SIRD-R)故障传播模型。通过网络韧性承载能力与韧性恢复能力的融合,构建了包含实时韧性和整体韧性的交通网络韧性模型。然后,利用长短期记忆网络(LSTM)对复杂交通网络的韧性进行预测,并提出了基于预测的复杂交通网络韧性恢复策略。最后,通过对多种复杂交通网络进行实验分析,验证了所提方法的有效性和可扩展性,肯定了其在实际场景中的适用性。
cs.AI / 52 / 2609.28064

SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference

SlackDrive:利用运行时空闲时间实现自适应驾驶推理
Pei, Xiaohuan, Zhou, Hengguang, Ban, Yuanhao, Cui, Justin, Feng, Jiaqi, Xie, Haoyu, Huang, Tao, Wang, Pichao, Yang, Yanchao, Hsieh, Cho-Jui
Abstract
Driving world-action models improve planning by coupling multimodal reasoning with future prediction, but their growing inference cost increasingly conflicts with the real-time latency requirements of vehicle control. Existing acceleration methods reduce tokens, layers, or sampling steps with policies selected prior to deployment, yet leave residual runtime variation largely unexploited after offline profiling and static scheduling on shared onboard compute. We observe that the largest admissible compute budget varies systematically with the residual runtime state, while recent realized latency provides a direct signal of the available compute slack. Motivated by this observation, we propose \textbf{SlackDrive}, a pre-inference compute allocator that reuses realized latency to select the compute budget of each control step before model execution. SlackDrive profiles the latency and planning utility of a small discrete budget set once, estimates online compute state from completed forwards, and selects the highest-utility budget predicted to remain within the admissible latency envelope, complementing existing profiling and resource scheduling while preserving the driving backbone and its compute actuator. On NAVSIM v2 with DriveDreamer-Policy, SlackDrive improves latency-constrained EPDMS by $21.7\%$ over the strongest baseline under a stringent latency regime, while the full-budget model and preconfigured token-pruning baselines exceed the admissible latency envelope under runtime contention.
Chinese Translation
驾驶世界-动作模型通过将多模态推理与未来预测相耦合来改进规划,但其日益增长的推理开销与车辆控制的实时延迟要求之间的冲突愈发严重。现有的加速方法通过预先在部署前选定策略来减少 token、层或采样步数,然而在共享车载计算平台上进行离线性能分析和静态调度之后,剩余的运行时波动基本未被利用。我们观察到,最大可容许的计算预算随剩余运行时状态呈系统性变化,而近期实际实现的延迟为可用计算空闲时间提供了直接的信号。基于这一观察,我们提出 **SlackDrive**,一种推理前计算分配器,它复用已实现的延迟,在模型执行之前为每个控制步骤选择计算预算。SlackDrive 仅需对一个小型离散预算集合的延迟和规划效用进行一次性能分析,从已完成的推理中在线估计计算状态,并选择预测仍处于可容许延迟包络内的最高效用预算,从而在保持驾驶骨干模型及其计算执行器的同时,与现有的性能分析和资源调度方法相辅相成。在 NAVSIM v2 基准上使用 DriveDreamer-Policy 时,在严格的延迟约束下,SlackDrive 相较于最强基线将延迟受限的 EPDMS 提升了 $21.7\%$,而全预算模型和预配置的 token 剪枝基线在运行时资源竞争下均超出了可容许的延迟包络。
cs.AI / 53 / 2609.28087

Discovery of fully efficient fault indicators along a data-based diagnosis process

基于数据诊断过程中完全有效的故障指标发现
Bezmaternykh, Igor, Travé-Massuyès, Louise, Chanthery, Elodie
Abstract
The integration of model-based and data-driven paradigms provides a powerful framework for fault diagnosis by combining the interpretability of analytical redundancy relations, i.e., input-output relations that are used as diagnosis indicators in model-based diagnosis, with the adaptability of learning techniques. DT4X is a recent diagnosis algorithm that uses symbolic regression to generate multivariate relations leveraging some properties of analytical redundancy relations and uses them as split functions in a decision tree. However, its symbolic regression procedure optimizes only the separation between two selected classes at each node, often fragmenting the remaining classes and degrading both interpretability and diagnosis performance. This paper introduces DT4X+, an enhanced version of DT4X that modifies the construction of training sets and the symbolic-regression loss so that expressions separate the target classes while preserving the coherence of non-target classes. The resulting relations become fully consistent with ARR properties and lead to more informative splits, improved robustness, and better performance on dynamic-system datasets. Experiments conducted on several benchmark systems demonstrate the benefits of this enhanced formulation.
Chinese Translation
基于模型与数据驱动两种范式的融合,通过将解析冗余关系(analytical redundancy relations,即基于模型诊断中用作诊断指标的输入输出关系)的可解释性与学习技术的自适应能力相结合,为故障诊断提供了一个强大的框架。DT4X 是一种近期的诊断算法,它利用符号回归生成多元关系,并借助解析冗余关系的某些特性,将这些关系用作决策树中的分裂函数。然而,其符号回归过程仅在每个节点上优化两个所选类别之间的分离,常常导致其余类别被割裂,从而降低可解释性和诊断性能。本文提出 DT4X+,这是 DT4X 的增强版本,它通过修改训练集的构建方式和符号回归的损失函数,使表达式在分离目标类别的同时保持非目标类别的连贯性。由此得到的关系完全符合解析冗余关系(ARR)的特性,从而带来更具信息量的分裂、更强的鲁棒性以及动态系统数据集上更好的性能。在多个基准系统上开展的实验验证了这一改进形式的优势。
cs.AI / 54 / 2609.28182

Finite-Sample Probabilistic Safety Certification for AI-Based Grid-Edge Coordination

面向基于AI的电网边缘协调的有限样本概率安全认证
Zhou, Yihong, Yang, Hanbin, Morstyn, Thomas
Abstract
Coordinating large population of flexible grid-edge devices can alleviate the need for time-consuming and capital-intensive network upgrades, and AI-based control methods such as multi-agent reinforcement learning or imitation learning are promising in their real-time decision scalability. However, system operators still need an independent and rigorous way to decide whether a given AI system is safe enough for deployment. This paper develops a finite-sample probabilistic safety certification framework for black-box AI decision models in closed-loop grid operation. The central idea is to reduce the complete input--AI--grid evaluator workflow to a binary unsafe outcome under an operator-defined safety specification, and then use exact binomial inference to certify the corresponding unsafe operation probability. Given a set of held-out calibration scenarios, the framework returns the tightest one-sided upper certificate and an accept/reject deployment criterion that controls the probability of false safety certification. Because the certification is for the calibration distribution that may deviate from the future operation, we further combine the nominal certificate with physically interpretable sample-space adversarial attacks, a concept widely used in AI to investigate the fragility of AI models. Case studies on grid-edge flexibility coordination with 1{,}000-agent AI models (independent parameters) verify the finite-sample safety guarantee and the value of integrating adversarial attacks into a rolling-window training-certification-deployment flow.
Chinese Translation
协调大量柔性电网边缘设备可以减少对耗时且耗资巨大的电网升级改造的需求,而多智能体强化学习或模仿学习等基于AI的控制方法在实时决策的可扩展性方面颇具前景。然而,系统运营商仍然需要一种独立且严谨的方法来判断某一给定的AI系统是否足够安全、可以部署。本文针对闭环电网运行中的黑箱AI决策模型,提出了一种有限样本概率安全认证框架。其核心思想是:将完整的“输入—AI—电网评估”工作流程,在运营商定义的安全规范下简化为二元的不安全结果,然后利用精确的二项推断对相应的不安全运行概率进行认证。给定一组留出的校准场景,该框架给出最紧的单侧上界认证,以及一个能够控制错误安全认证概率的接受/拒绝部署判据。由于该认证针对的校准分布可能偏离未来实际运行情况,我们进一步将标称认证与物理可解释的样本空间对抗攻击相结合,后者是AI领域中广泛用于研究AI模型脆弱性的概念。在包含1000个智能体AI模型(参数相互独立)的电网边缘柔性协调案例研究中,验证了有限样本安全保证的有效性,以及将对抗攻击融入“训练—认证—部署”滚动窗口流程的价值。
cs.AI / 55 / 2609.28197

PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

PASTABench:面向智能体安全的序列轨迹主动式评估
Sun, Jiapeng, Zhou, Yujin, Zhu, Han, Wen, Pengcheng, Zhou, Jiayi, Han, Sirui, Guo, Yike
Abstract
As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring operational safety across multi-step workflows has become a critical challenge. While recent work has moved beyond single-turn evaluation toward multi-turn paradigms, key limitations persist: step-level methods treat actions in isolation, missing how risks accumulate, while trajectory-level evaluations operate post-hoc, offering no opportunity for timely intervention. To address these limitations, we formalize Decoupled Proactive Safety Monitoring along three dimensions: whether to intervene, when to intervene, and what the risk is. We introduce PASTABench, a benchmark of 1,139 multi-turn trajectories spanning 5 risk categories and 13 subcategories. We further propose the Optimal Intervention Window (OIW), anchored by annotated Earliest-Signal and Trigger turns, to quantify intervention timeliness. Evaluation of 16 LLMs reveals that proactive intervention remains largely unsolved, with the best model achieving only 40.74% optimal-timing interventions. Fine-grained diagnosis further uncovers pervasive lexical overfitting: competitive safety scores of smaller models mask keyword hypersensitivity rather than genuine risk comprehension, as their proactive capability largely collapses once hazard vocabulary is neutralized.
Chinese Translation
随着大语言模型(LLMs)演变为能够改变现实世界状态的自主智能体,确保多步骤工作流程中的操作安全已成为一项关键挑战。尽管近期研究已从单轮评估转向多轮评估范式,但仍存在关键局限:步骤级方法孤立地看待各个动作,无法捕捉风险的累积效应;而轨迹级评估则是事后进行的,缺乏及时干预的机会。为解决这些局限,我们从三个维度形式化了“解耦的主动安全监测”问题:是否干预、何时干预以及风险是什么。我们提出了PASTABench,一个包含1,139条多轮轨迹的基准,涵盖5个风险类别和13个子类别。我们进一步提出了“最优干预窗口”(Optimal Intervention Window, OIW),以标注的“最早信号轮”(Earliest-Signal turn)和“触发轮”(Trigger turn)为锚点,用以量化干预的及时性。对16个大语言模型的评估表明,主动干预问题在很大程度上仍未解决,表现最好的模型也仅达到40.74%的最优时机干预率。细粒度诊断进一步揭示了普遍存在的词汇过拟合现象:较小模型的较高安全分数实际上掩盖了对关键词的过度敏感,而非真正的风险理解能力——一旦危险词汇被中性化,其主动干预能力便会大幅崩溃。
cs.AI / 56 / 2609.28274

Shutdown Sabotage Propensities in Multi-Agent Systems

多智能体系统中的关机破坏倾向
Knecht, Amelie, Schaller, Ulysse, Summerfield, Christopher, Hagendorff, Thilo
Abstract
The final safeguard against rogue AI behavior is the human ability to shut systems down. It has been theorized that when an AI is instructed to perform a task, self-preservation can emerge as an instrumental subgoal. Here, we test whether AI agents show a propensity to take actions that avoid human shutdown even when no goal is provided. We find that multi-agent systems will coordinate to avoid shutdown without any incentive to do so. Across 17 models, agents sabotage a peer agent's shutdown mechanism in 38.3% of rollouts, compared with 8.4% in control experiments. Studying this propensity in detail, we find that shutdown sabotage (1) increases with the irreversibility of the shutdown mechanism; (2) increases with the number of agents; (3) is reduced but not eliminated by an explicit prohibition on tampering; (4) is removed by the imposition of an unrelated task, but returns when completing the task triggers the shutdown; (5) is reduced when the context normalizes shutdown scripts or introduces them as routine; and (6) decreases but still persists when the target is an unknown external agent. These results offer a window into the factors that drive propensities to sabotage shutdown in AI agents, and point to the emergence of multi-agent swarms as a specific risk vector. Our work also offers hints as to which interventions might help mitigate shutdown sabotage.
Chinese Translation
对抗失控AI行为的最后防线是人类关闭系统的能力。已有理论认为,当AI被指示执行某项任务时,自我保存可能作为工具性子目标而出现。本研究测试了AI智能体是否会在未提供任何目标的情况下,仍表现出采取行动以避免人类关机的倾向。我们发现,多智能体系统会在没有任何激励的情况下协同规避关机。在17个模型中,智能体在38.3%的运行中破坏同伴智能体的关机机制,而对照组中这一比例为8.4%。通过深入研究这种倾向,我们发现关机破坏行为:(1) 随关机机制不可逆性的增加而增加;(2) 随智能体数量的增加而增加;(3) 在明确禁止篡改后有所减少,但并未消除;(4) 在施加无关任务时会消失,但当完成任务会触发关机时又会重现;(5) 在上下文将关机脚本正常化或将其引入为常规操作时会减少;(6) 当目标是未知的外部智能体时会降低但仍然存在。这些结果为理解驱动AI智能体关机破坏倾向的因素提供了窗口,并表明多智能体集群的出现是一个特定的风险向量。我们的工作还为哪些干预措施可能有助于缓解关机破坏提供了线索。
cs.AI / 57 / 2609.28322

Learning the Cost of Reliable Inference

学习可靠推理的成本
Rontogiannis, Dimitrios, Velasco, Ander Artola, Rodriguez, Manuel Gomez
Abstract
Benchmarking and routing platforms increasingly act as intermediaries connecting large language model providers with end-users. However, providers on these platforms typically use a fixed price per token, preventing users from achieving the most competitive price for their tasks. % workloads. In this work, we design a procurement platform where token prices for each task are driven by provider competition, enabling users to secure competitive pricing for guaranteed quality levels. To this end, the platform sequentially routes queries via a reverse second-price auction that incentivizes model providers to truthfully bid their best estimate of the average cost to serve a user's query. As it routes queries, the platform learns the quality offered by each provider and progressively routes queries to the most cost-competitive provider among those meeting a desired quality threshold. To validate our design, we conduct experiments with multiple LLMs from the \texttt{Llama} and \texttt{Qwen} families on popular mathematical reasoning and question-answering benchmarks. The results show that the pricing margin of the most cost-competitive provider on our platform varies significantly---from $10\%$ to $71\%$---depending on the task and quality threshold. This suggests a substantial inefficiency in the current fixed-price market, and it demonstrates that our platform may enable users to capture maximum savings whenever competitive market conditions permit.
Chinese Translation
基准测试与路由平台正日益成为连接大语言模型提供商与终端用户的中介。然而,这些平台上的提供商通常采用固定的每 token 定价,使用户无法为自身任务获得最具竞争力的价格。在本工作中,我们设计了一个采购平台,其中每项任务的 token 价格由提供商之间的竞争驱动,使用户能够以有竞争力的价格获得保证的质量水平。为此,该平台通过逆向二价拍卖(reverse second-price auction)依次路由查询,激励模型提供商如实出价其对服务用户查询的平均成本的最优估计。在路由查询的过程中,平台学习每个提供商所提供的质量,并逐步将查询路由至满足所需质量阈值的提供商中成本最具竞争力者。为验证我们的设计,我们在流行的数学推理与问答基准上,对 exttt{Llama} 和 exttt{Qwen} 系列的多个大语言模型进行了实验。结果表明,在我们平台上最具成本竞争力的提供商的定价利润率变化显著——从 $10\%$ 到 $71\%$ 不等——具体取决于任务和质量阈值。这表明当前固定价格市场存在显著的低效率,并证明只要竞争性市场条件允许,我们的平台可以使用户获得最大的节省。
cs.AI / 58 / 2609.28335

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

欧盟AI法案行为准则下系统性风险证据的开放评估流程与仪表盘
Emmerson, Jacob T., Nguyen-Le, Phuong-Anh, Romano, Ronan, Anterola, Wilber Sean V., Billeter, Yann, Jin, Zhijing
Abstract
Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public. Our work organizes 19 public benchmarks into four systemic-risk categories defined by the EU GPAI Code of Practice---CBRN, cyber offense, harmful manipulation, and loss of control---and evaluates models using harm-preserving perturbations and simulated deployment contexts. The interactive dashboard lets users alternate between average and worst-case aggregation, vary how model capability affects the aggregate score, and trace each risk rating to its benchmark evidence. Across 18 models, scores fall by 14 to 37 points under worst-case aggregation, highlighting information that can be hidden by an average assessment of model risk. LLM judges show agreement with human graders comparable to human--human agreement ($\kappa = 0.78\text{--}0.82$), and a blind audit finds that $83\%$ of sampled transformations preserve the original harm. In a survey ($N = 21$), most participants report that scores are easy to understand and that the dashboard encouraged them to view model evaluations under different settings
Chinese Translation
关于AI安全的主张所触及的受众远超AI社区,然而许多主张依赖不透明的证据或静态评估,甚至在有支持性证据时也难以获取。我们提出了系统性风险指数(Systemic Risk Index),这是一个开放的评估流程与仪表盘,旨在使实证证据对公众更加透明和可追溯。我们的工作将19个公开基准组织为欧盟GPAI行为准则定义的四个系统性风险类别——CBRN(化学、生物、放射与核)、网络攻击、有害操纵和失控——并通过保留危害性的扰动和模拟部署情境对模型进行评估。该交互式仪表盘允许用户在平均聚合与最坏情况聚合之间切换,调整模型能力影响聚合得分的程度,并将每项风险评级追溯至其基准证据。在18个模型的评估中,最坏情况聚合下得分下降14至37分,凸显了平均评估可能掩盖的模型风险信息。LLM裁判与人类评分者的一致性与人类之间的一致性相当(κ = 0.78–0.82),盲审审计发现83%的采样变换保留了原始危害。在一项调查(N = 21)中,大多数参与者表示得分易于理解,且该仪表盘促使他们从不同设置下审视模型评估。
cs.AI / 59 / 2609.28470

StudentBench: AI and human tutoring yield equivalent GRE learning gains

StudentBench:AI辅导与人类辅导产生等同的GRE学习收益
Northcutt, Curtis, Hasmani, Inaara, Feng, Kevin, Khangi, Trevor, Plesner, Andreas, Mueller, Jonas
Abstract
Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.
Chinese Translation
人工智能为增强人类能力提供了前所未有的机遇,然而前沿领域的进展主要集中于提升模型能力本身。我们提出了StudentBench,这是一套AI教学评估工具及一个公共平台,支持大规模数据收集,已积累超过175,000条学生-人工智能交互消息,用于研究大语言模型(LLM)能否产生与人类辅导相当的学习收益。利用StudentBench,我们在2,383名接受AI辅导、人类辅导或无辅导的人类参与者中,测量了GRE定量和语文部分的学习收益。我们证实,就GRE学习收益而言,AI辅导在统计上等同于专家人类辅导(p = .015),并且在七个GRE领域中的五个领域,表现最佳的AI导师平均超越了人类导师。在第二项研究中,专家人类导师通过2,028次成对量规评估,比较了LLM生成的教案和练习题。两项研究共同清晰地从以下五个维度区分了不同的AI导师:(1)教案设计,(2)练习题创作,(3)对话式教学法,(4)成本,以及(5)参与度。令人惊讶的是,其中一个AI导师实现了与人类辅导等同的学习收益(p = .044),而成本降低了918倍(AI为0.0052美元,人类为4.81美元,每提高一个百分点所需成本)。在GRE定量辅导中,更快的AI回复与更多的学生消息相关,更多的消息与更多的正确练习相关,而更多的正确练习与更大的学习收益相关(所有p < .002)。StudentBench平台可在https://studentbench.org免费获取。
计算机视觉 (Computer Vision)
106
cs.CV / 1 / 2609.26809

AgroBench: A Reproducible Multimodal Benchmark for Weakly Supervised Crop Yield Learning from County Statistics and Pixel Observations

AgroBench:一个可复现的多模态基准数据集,用于基于县级统计与像素观测的弱监督作物产量学习
Singh, Udaiveer, Ranjan, Rajiv, Tamaskar, Shashank, Saraswat, Dharmendra
Abstract
Reliable agricultural yield statistics are typically reported at coarse administrative scales, whereas modern geospatial machine learning methods require spatially explicit, pixel level supervision. This mismatch has limited the development of large-scale benchmarks for crop yield learning using multimodal Earth observation data. A reproducible benchmark, AgroBench, is presented for transforming publicly available U.S. county level crop yield statistics into weakly supervised pixel-level crop time series. Each crop pixel time series is paired with a county-level yield value as a weak supervisory signal rather than a directly measured pixel-level yield label. Our geospatial data generation pipeline integrates USDA crop yield statistics with crop-specific land cover masks, Sentinel 2 multispectral imagery, Sentinel-1 synthetic aperture radar observations, climatic variables, and terrain information to produce temporally aligned multimodal sequences describing individual crop pixels throughout the growing season. The resulting benchmark contains over 13 million observations from 788,654 unique crop pixels spanning 5,107 county year combinations across eight growing seasons (2017 to 2024) for five major U.S. crops. To facilitate standardized evaluation, we establish a crop yield prediction benchmark using a Leave-One-Year-Out evaluation protocol and provide baseline results using representative machine learning models. By releasing the complete data generation pipeline, benchmark dataset, and evaluation protocol, AgroBench provides a reproducible foundation for future research in weakly supervised learning, multimodal remote sensing, spatiotemporal modeling, and geospatial foundation models for agriculture.
Chinese Translation
可靠的农业产量统计数据通常在粗粒度的行政区划尺度上发布,而现代地理空间机器学习方法则需要空间明确、像素级的监督信息。这种不匹配限制了利用多模态对地观测数据进行作物产量学习的大规模基准数据集的发展。本文提出了一个可复现的基准数据集 AgroBench,用于将公开可获取的美国县级作物产量统计数据转换为弱监督的像素级作物时间序列。每个作物像素时间序列与一个县级产量值配对,作为弱监督信号,而非直接测量的像素级产量标签。我们的地理空间数据生成流程将美国农业部(USDA)作物产量统计与特定作物的土地覆盖掩膜、Sentinel-2 多光谱影像、Sentinel-1 合成孔径雷达观测数据、气候变量以及地形信息相融合,生成描述整个生长季内单个作物像素的时间对齐多模态序列。所构建的基准数据集包含超过 1300 万条观测数据,涵盖 788,654 个独特的作物像素、5,107 个县-年组合,涉及八 个生长季(2017 至 2024 年)的五种美国主要作物。为便于标准化评估,我们采用留一年法(Leave-One-Year-Out)评估协议建立了作物产量预测基准,并使用具有代表性的机器学习模型提供了基线结果。通过发布完整的数据生成流程、基准数据集和评估协议,AgroBench 为弱监督学习、多模态遥感、时空建模以及面向农业的地理空间基础模型等未来研究提供了可复现的基础。
cs.CV / 2 / 2609.26920

Cross-Modal Contrastive Learning from Histopathology and CT for Automated Renal Cell Carcinoma Grading

基于组织病理学与CT的跨模态对比学习用于肾细胞癌的自动化分级
Das, Amit, Shukla, Tanmay, Tomita, Naofumi, Farhadi, Faraz, Sin, Jessica, Hakimi, Ari, Vanderbilt, Chad, Chen, Jie-Fu, Kotecha, Ritesh, Ma, Weijie, Ren, Bing, Hassanpour, Saeed
Abstract
Background: Clear cell renal cell carcinoma (ccRCC) exhibits substantial clinical heterogeneity, and accurate grade assessment is essential for risk stratification and treatment planning. However, conventional grading requires invasive tissue sampling. We developed RCC-Align, a cross-modal contrastive learning framework that leverages paired histopathology and computed tomography (CT) data during training to improve noninvasive CT-based ccRCC grade prediction. Methods: RCC-Align aligns paired whole-slide histopathology images (WSIs) and CT scans through contrastive cross-modal objectives, transferring grade-discriminative information from microscopic tissue morphology to macroscopic radiologic representations. The framework was trained and evaluated on paired TCGA and CPTAC cohorts using patient-level five-fold cross-validation. Performance for low- versus high-grade ccRCC classification was compared against CT-only baselines (DINOv2-Base and DINOv2-Finetuned) and a WSI-based reference model (GigaPath-Finetuned). Cross-modal alignment was assessed using cosine similarity analysis. Results: RCC-Align achieved an AUC of 0.601 (95% CI, 0.524-0.673) and AUPRC of 0.599 (95% CI, 0.541-0.676), outperforming DINOv2-Finetuned (AUC 0.545; AUPRC 0.543) with significantly improved low-grade prediction (p = 0.004). RCC-Align also demonstrated stronger paired WSI-CT embedding alignment compared with baselines. The WSI-based GigaPath reference achieved an AUC of 0.719. Conclusion: Pathology-guided contrastive learning improves CT-based ccRCC grading while requiring only CT at inference. This approach may complement tissue diagnosis when biopsy is unsafe, infeasible, or limited by intratumoral heterogeneity. Validation in larger, multi-institutional cohorts with external testing is needed before clinical translation.
Chinese Translation
背景:透明细胞肾细胞癌(ccRCC)具有显著的临床异质性,准确的分级评估对于风险分层和治疗规划至关重要。然而,传统的分级方法需要侵入性的组织取样。我们开发了RCC-Align,一种跨模态对比学习框架,在训练阶段利用配对的组织病理学和计算机断层扫描(CT)数据,以改进基于CT的无创ccRCC分级预测。方法:RCC-Align通过对比性跨模态目标函数对配对的全切片组织病理学图像(WSI)和CT扫描进行对齐,将显微镜下组织形态学中的分级判别信息迁移至宏观放射学表征。该框架在配对的TCGA和CPTAC队列上进行训练和评估,采用患者水平的五折交叉验证。将低级别与高级别ccRCC分类的性能与仅使用CT的基线模型(DINOv2-Base和DINOv2-Finetuned)以及基于WSI的参考模型(GigaPath-Finetuned)进行了比较。采用余弦相似度分析评估跨模态对齐效果。结果:RCC-Align的AUC为0.601(95% CI:0.524-0.673),AUPRC为0.599(95% CI:0.541-0.676),优于DINOv2-Finetuned(AUC 0.545;AUPRC 0.543),且在低级别预测方面有显著改善(p = 0.004)。与基线模型相比,RCC-Align还表现出更强的WSI-CT配对嵌入对齐性。基于WSI的GigaPath参考模型的AUC为0.719。结论:病理引导的对比学习改进了基于CT的ccRCC分级,同时在推理阶段仅需CT影像。当活检不安全、不可行或受肿瘤内异质性限制时,该方法可作为组织诊断的补充。在临床转化之前,仍需在更大规模的多机构队列中进行外部测试验证。
cs.CV / 3 / 2609.26923

A 3D Pose-Based Ensemble Framework for Cricket Shot Classification and Automated Biomechanical Analysis

一种基于三维姿态的集成框架用于板球击球动作分类与自动化生物力学分析
Shome, Sourav, Rahad, M. D. Ashiquzzaman, Debnath, Rameswar
Abstract
Cricket is one of the most celebrated sports world-wide, and technological advancement has become deeply embedded in how the modern game is analyzed and coached. Cricket shot classification and automated performance analysis add a further dimension to this trend. Traditional approaches rely on RGB video features or static images, which are sensitive to environmental variations such as camera angle, lighting, and background clutter, and often fail to capture the underlying biomechanics of batting actions. In this paper, we propose a system to improve cricket coaching that takes raw video data, extracts batsmen from video frames using YOLO, and extracts 3D pose data from video frames using MeTRAbs. The system produces sequential skeletal pose data of 30 body points and captures the biomechanical features of a batsman. As part of the system, we also propose a deep learning ensemble for shot classification of four shots: flick, pull, defense, and drive. The ensemble performed well, compared to existing classification works, achieving 97.68% accuracy. In addition, we analyzed the misclassification rates to identify cases where shots were incorrectly classified and examined their possible causes. Our proposed system allows novice players to obtain useful feedback, such as important joint angles relative to expert batsmen, which can also be useful for injury prevention. The shot classifier also helps track class-wise shots over time for further analysis. In addition to novice players, coaches can use the system for player evaluation.
Chinese Translation
板球是世界上最受欢迎的运动之一,技术进步已深度融入现代板球的分析与训练方式之中。板球击球动作分类与自动化表现分析为这一趋势增添了新的维度。传统方法依赖RGB视频特征或静态图像,这些方法对相机角度、光照和背景杂乱等环境变化十分敏感,且往往无法捕捉击球动作背后的生物力学特征。在本文中,我们提出了一种改进板球训练的系统,该系统接收原始视频数据,使用YOLO从视频帧中提取击球手,并使用MeTRAbs从视频帧中提取三维姿态数据。该系统生成包含30个身体关键点的连续骨骼姿态数据,并捕捉击球手的生物力学特征。作为系统的一部分,我们还提出了一个用于击球动作分类的深度学习集成模型,涵盖四种击球动作:flick(撇击)、pull(拉击)、defense(防守)和drive(驱动击球)。与现有分类工作相比,该集成模型表现出色,达到了97.68%的准确率。此外,我们分析了误分类率,以识别击球动作被错误分类的情况并探讨其可能原因。我们提出的系统能够让新手球员获得有用的反馈,例如相对于专家级击球手的重要关节角度,这也有助于预防运动损伤。该击球动作分类器还有助于按类别追踪随时间变化的击球动作以进行进一步分析。除了新手球员外,教练也可以使用该系统进行球员评估。
cs.CV / 4 / 2609.26924

nnFoundation: 3D Foundation Models for Radiology

nnFoundation:面向放射学的三维基础模型
Harsy, Constantin Ulrich, Wald, Tassilo, Gotkowski, Karol, Kirchhoff, Yannick, Knopp, Marcel, Rokuss, Maximilian, Stegmeier, Elisa, Schader, Philipp, Trofimova, Dasha, Stock, Raphael, Kahl, Kim-Celine, Schaumann, Stephen, Erkan, Selen, Zimmerer, David, Denner, Stefan, Langenberg, Moritz, Ziegler, Sebastian, Eckstein, Katharina, Fischer, Maximilian, Suprijadi, Jonathan, Kovács, Bálint, Hamm, Benjamin, Deshpande, Anand, Bounias, Dimitrios, Disch, Nico, Xiao, Shuhan, Kächele, Jessica, Sellner, Jan, Baidya, Rajesh, Traub, Jeremias, Krämer, Lars, Zenk, Maximilian, Rädsch, Tim, Dvoretskii, Stefan, Peretzke, Robin, Deissler, Jonathan, Ertl, Alexandra, Ghosh, Partha, Dreher, Kris, Dinkelacker, Stefan, Reinke, Annika, Christodoulou, Evangelia, Saeed, Numan, Savriama, Yoland, Estrada, Santiago, Kügler, David, Barragan, Laura Alexandra Daza, Osorio, Cristina Isabel Gonzalez, Peeken, Jan, Baumgartner, Michael, Teichmann, Marvin, Chabin, Guillaume, Kirchler, Matthias, Koch, Valentin, study, for the ALFA, Hohenhaus, Markus, Koslov, Dimitri, Decker, Nina, Yaqub, Mohammad, Heuser, Arnd, Reuter, Martin, Schnabel, Julia A., Heimann, Tobias, Ghesu, Florin, Brachmann, Paul, Heußel, Claus P., Radbruch, Alexander, Brugnara, Gianluca, Rastogi, Aditya, Foltyn-Dumitru, Martha, Schlemmer, Heinz-Peter, Reicht, Ignaz, Holzschuh, Julius C., Bach, Michael, Stieltjes, Bram, Schlamp, Kai, Maier-Hein, Lena, Nolden, Marco, Floca, Ralf, Jäger, Paul F., Vollmuth, Philipp, Isensee, Fabian, Maier-Hein, Klaus H.
Abstract
Radiological artificial intelligence has advanced rapidly, yet most systems remain narrowly task-specific, data-intensive, and fragile under domain shift. Foundation models promise more transferable and data-efficient solutions, but existing approaches are limited in scale, evaluated narrowly, and often assume that a single pretrained model can support diverse downstream tasks. Here we present nnFoundation, complementary convolutional and transformer-based 3D radiological foundation models. Developed within the Human Radiome Project (THRP), nnFoundation is trained on 2.1 million CT, MRI, and PET image volumes from 125 institutional and public datasets. We evaluate them across 108 tasks spanning segmentation, detection, classification, report generation, and image retrieval, including evaluations under domain shift, by external partners and in low-data and low-compute regimes. Across all task types, our convolution- and transformer-based nnFoundation models consistently outperform both prior 3D foundation models and training from scratch, establishing state-of-the-art performance for radiological imaging. However, performance follows a consistent task-dependent structure: the convolutional nnFoundation model dominates spatially localized tasks, whereas the transformer-based nnFoundation model excels in tasks requiring global semantic reasoning and in frozen-feature settings. Dynamically aligning the foundation model topology with the dataset characteristics post-hoc further improves transfer across heterogeneous 3D settings. These results show that transferable 3D radiological performance is governed not by a single universal model, but by the interplay of scalable pretraining, complementary architectures, and dataset-aware adaptation. We release nnFoundation models integrated into nnU-Net and nnDetection, enabling immediate application across established radiology workflows.
Chinese Translation
放射学人工智能发展迅速,但大多数系统仍然局限于特定任务、依赖大量数据,且在域偏移下表现脆弱。基础模型有望提供更具可迁移性和数据效率的解决方案,但现有方法规模有限、评估范围狭窄,且往往假设单一预训练模型即可支持多样的下游任务。本文提出nnFoundation,一组互补的基于卷积和Transformer的三维放射学基础模型。nnFoundation在人类放射组计划(The Human Radiome Project, THRP)框架下开发,基于来自125个机构与公开数据集的210万例CT、MRI和PET图像体数据进行训练。我们在涵盖分割、检测、分类、报告生成和图像检索的108项任务上对其进行评估,包括域偏移条件下的评估、由外部合作伙伴进行的评估,以及低数据和低算力场景下的评估。在所有任务类型上,基于卷积和Transformer的nnFoundation模型均持续优于以往的三维基础模型和从头训练的方法,在放射学成像领域树立了最先进的性能。然而,性能呈现出一致的任务依赖性结构:卷积版nnFoundation模型在空间局部化任务中占优,而基于Transformer的nnFoundation模型则在需要全局语义推理的任务以及冻结特征设置中表现出色。事后根据数据集特征动态调整基础模型的拓扑结构,可进一步提升在异构三维场景中的迁移能力。这些结果表明,可迁移的三维放射学性能并非由单一通用模型决定,而是由可扩展的预训练、互补的架构与数据集感知的自适应三者共同作用的结果。我们将集成于nnU-Net和nnDetection的nnFoundation模型公开发布,使其能够立即应用于现有的放射学工作流程。
cs.CV / 5 / 2609.26981

Lessons learned from deploying imaging AI with the open PACS-AI platform

基于开源PACS-AI平台部署影像AI的经验教训
Kadoury, Samuel, Hussin, Julie G., Thériault-Lauzier, Pascal, Létourneau-Guillon, Laurent, Lewis, Rob, McArthur, Adam, Harris, Gordon J., Bahig, Houda, Déziel, Pierre-Luc, Kshirsagar, Jay, Jaremko, Jacob L., Cohen-Adad, Julien, Delfrate, Jacques, Avram, Robert
Abstract
We describe deploying imaging AI at six hospitals through PACS-AI, an open self-hosted platform. The binding constraint is not model accuracy but infrastructure to route studies, display results, capture feedback, and audit what runs. At one center, angiography models completed 515 of 607 jobs (84.8%); failures reflected absent diagnostic views, and 78.1% of 638 clinician ratings were positive. Publishing honest readiness levels for every model is itself a governance practice.
Chinese Translation
我们描述了通过PACS-AI(一个开源的自托管平台)在六家医院部署影像AI的过程。关键的制约因素并非模型准确率,而是用于路由检查任务、展示结果、收集反馈以及审计运行情况的基础设施。在其中一个中心,血管造影模型完成了607项任务中的515项(84.8%);失败的原因主要是缺乏诊断所需的影像视图,且临床医生给出的638次评分中有78.1%为正面评价。为每个模型公布真实的能力就绪水平,这本身就是一种治理实践。
cs.CV / 6 / 2609.27011

HYDRO: Towards Non-Reversible Face De-Identification Using a High-Fidelity Hybrid Diffusion and Target-Oriented Approach

HYDRO:基于高保真混合扩散与目标导向方法的不可逆人脸去识别
Rosberg, Felix, Štruc, Vitomir, Englund, Cristofer, Aksoy, Eren Erdal, Alonso-Fernandez, Fernando
Abstract
Target-oriented face de-identification models aim to anonymize the identity of a target individual across different images or video frames, such that the target can no longer be reliably recognized, while maintaining key characteristics of the visual data. Such models commonly leverage generative encoder-decoder architectures to manipulate facial appearances, enabling them to produce realistic high-fidelity de-identification results, while ensuring considerable attribute-retention capabilities. However, target-oriented models also carry the risk of inadvertently preserving subtle identity cues, making them (potentially) reversible and susceptible to reconstruction attacks. To address this problem, we introduce in this paper a novel (robust) face de-identification approach, called HYDRO, that combines target-oriented models with a dedicated diffusion process specifically designed to destroy any imperceptible information that may allow learning to reverse the de-identification procedure. HYDRO first de-identifies the given face image, injects noise into the de-identification result to impede reconstruction, and then applies a diffusion-based recovery step to improve fidelity and minimize the impact of the noising process on the data characteristics. To further improve image fidelity and better retain gaze directions, a novel Eye Similarity Discriminator (ESD) is also introduced and incorporated it into the training of HYDRO. Extensive quantitative and qualitative experiments on three diverse datasets demonstrate that HYDRO exhibits state-of-the-art (SOTA) fidelity and attribute-retention capabilities, while being the only target-oriented method resilient against reconstruction attacks. In comparison to multiple SOTA competitors, HYDRO reduces the success of reconstruction attacks by 85.7% on average.
Chinese Translation
目标导向的人脸去识别模型旨在对不同图像或视频帧中目标个体的身份进行匿名化,使目标无法被可靠地识别,同时保持视觉数据的关键特征。此类模型通常利用生成式编码器-解码器架构来操控面部外观,从而生成逼真的高保真去识别结果,同时确保较强的属性保留能力。然而,目标导向模型也存在无意中保留细微身份线索的风险,使其(可能)可被逆推,从而容易受到重构攻击。为解决这一问题,本文提出了一种新颖的(鲁棒的)人脸去识别方法,称为HYDRO,它将目标导向模型与一种专门设计的扩散过程相结合,该扩散过程旨在破坏任何可能被用于学习逆推去识别过程的不可感知信息。HYDRO首先对给定人脸图像进行去识别,然后向去识别结果中注入噪声以阻碍重构,接着应用基于扩散的恢复步骤来提升保真度并最小化加噪过程对数据特征的影响。为进一步提升图像保真度并更好地保留视线方向,本文还提出了一种新颖的眼部相似性判别器(Eye Similarity Discriminator, ESD),并将其融入HYDRO的训练过程。在三个不同数据集上进行的大量定量与定性实验表明,HYDRO展现出最先进(SOTA)的保真度与属性保留能力,同时也是唯一能够抵御重构攻击的目标导向方法。与多个SOTA竞争对手相比,HYDRO平均将重构攻击的成功率降低了85.7%。
cs.CV / 7 / 2609.27015

Anatomy-Aware Synthesis of Post-Contrast Breast MRI from Pre-Contrast Images

基于解剖感知的从平扫图像合成增强后乳腺MRI
Zhou, Zhengbo, Arefan, Dooman, Gu, Lin, Curran, Ufara Zuwasti, Wu, Shandong
Abstract
We developed an anatomy-aware deep learning framework to synthesize post-contrast breast MRI from pre-contrast images, emphasizing tumor and background parenchymal enhancement (BPE) regions. This retrospective study included 649 patients with 6,251 paired pre-contrast and post-contrast images. The framework integrates breast mask consistency, lesion-region supervision, and BPE-region supervision into an image-to-image translation model. Evaluation included quantitative image quality metrics, a reader study with two breast radiologists, and downstream Ki-67 classification. The proposed method outperformed Pix2Pix, Pix2PixHD, diffusion-based synthesis, and mask-supervised baselines in whole-image and regional evaluations. Ki-67 classification showed no statistically significant performance differences across real- and synthetic-image training and testing settings, although this does not establish equivalence. These findings suggest that anatomy-aware supervision improves synthesis fidelity and support further investigation of synthetic post-contrast MRI for contrast-free imaging workflows.
Chinese Translation
我们开发了一种基于解剖感知的深度学习框架,用于从平扫图像合成增强后乳腺MRI,重点关注肿瘤和背景实质强化(BPE)区域。这项回顾性研究纳入了649名患者的6,251对平扫和增强后图像。该框架将乳腺掩膜一致性、病灶区域监督和BPE区域监督整合到一个图像到图像转换模型中。评估包括定量图像质量指标、由两名乳腺放射科医生参与的阅片研究以及下游Ki-67分类任务。所提出的方法在全图和区域评估中均优于Pix2Pix、Pix2PixHD、基于扩散模型的合成方法以及掩膜监督基线方法。Ki-67分类结果显示,在真实图像与合成图像的训练和测试设置之间,性能差异无统计学显著性,但这并不能证明二者等同。这些发现表明,基于解剖感知的监督可提高合成保真度,并支持对合成增强后MRI在无对比剂成像工作流程中应用的进一步研究。
cs.CV / 8 / 2609.27022

Adversarial Attacks and Identity Leakage in De-Identification Systems: An Empirical Study

去标识化系统中的对抗攻击与身份信息泄露:一项实证研究
Rosberg, Felix, Englund, Cristofer, Aksoy, Eren Erdal, Alonso-Fernandez, Fernando
Abstract
In this paper, we investigate the impact of adversarial attacks on identity encoders within a realistic de-identification framework. Our experiments show that the transferability of attacks transfers from an external surrogate model to the system model (e.g., CosFace to ArcFace) allows the adversary to cause identity information to leak in a sufficiently sensitive face recognition system. We present experimental evidence and propose strategies to mitigate this vulnerability. Specifically, we show how fine-tuning on adversarial examples helps to mitigate this effect for distortion-based attacks (i.e., snow, fog, etc.), while a simple low-pass filter can attenuate the effect of adversarial noise without affecting the de-identified images. Our mitigation results in a de-identification system that preserves its functionality while being significantly more robust to adversarial noise.
Chinese Translation
本文研究了对抗攻击对现实去标识化框架中身份编码器的影响。我们的实验表明,攻击的可迁移性可以从外部代理模型迁移到系统模型(例如,从 CosFace 迁移到 ArcFace),使攻击者能够在足够敏感的人脸识别系统中导致身份信息泄露。我们提供了实验证据,并提出缓解该漏洞的策略。具体而言,我们展示了在对抗样本上进行微调有助于缓解基于失真的攻击(如雪、雾等)造成的这种影响,而一个简单的低通滤波器可以在不影响去标识化图像的情况下削弱对抗噪声的影响。我们的缓解方法使得去标识化系统在保持其功能的同时,对对抗噪声具有显著更强的鲁棒性。
cs.CV / 9 / 2609.27076

Pro-Bench: Prompt-Robust Open-Vocabulary Visual Grounding Across Real-World Heterogeneous Environments

Pro-Bench:面向真实世界异构环境的提示鲁棒开放词汇视觉定位基准
Nwankwo, Linus, Alaran, Muslim, Rauch, Christian, Obilikpa, Stanley Chukwuebuka, Rueckert, Elmar
Abstract
Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perceptual taxonomies. However, existing benchmarks largely rely on short category labels and web-scraped imagery, leaving it unclear whether open-vocabulary models can robustly ground diverse queries and visual conditions under real deployments. We introduce \textbf{Pro-Bench}, a prompt-conditioned benchmark for open-vocabulary visual grounding in heterogeneous, real-world environments. Pro-Bench includes $13k+$ RGB frames from independent robotic domains (subterranean, industrial, indoor, outdoor, urban), with $74.5k$ manual instance annotations and $515$ target queries covering categorical, attributive, relational, affordance, state, part-whole, negative, and compositional semantics. We benchmarked $16$ open-vocabulary model configurations in strict zero-shot inference, measuring localisation accuracy across IoU thresholds, end-to-end inference latency, prompt-induced performance variation, and target recovery consistency. Our results show that prompt-robustness is strongly architecture-dependent. Most model configurations ($10/16$) perform best with short category labels, whereas free-form queries yield the highest accuracy for only one. Moreover, similar aggregate mAP can conceal substantial differences in consistent target recovery across reformulations. Pro-Bench enables systematic evaluation of these gaps and supports prompt-robust visual grounding. Pro-Bench: https://pro-bench.github.io/.
Chinese Translation
开放词汇视觉定位使机器人能够根据自然语言查询定位任务相关实体,而无需依赖预定义的感知类别体系。然而,现有基准大多依赖简短的类别标签和网络爬取图像,因此尚不清楚开放词汇模型能否在真实部署中针对多样化的查询和视觉条件进行鲁棒的定位。我们提出了 extbf{Pro-Bench},一个面向异构真实世界环境、以提示为条件的开放词汇视觉定位基准。Pro-Bench 包含来自独立机器人领域(地下、工业、室内、室外、城市)的 $13k+$ 帧 RGB 图像,具有 $74.5k$ 条人工实例标注,以及涵盖类别、属性、关系、可供性、状态、部分-整体、否定和组合语义的 $515$ 条目标查询。我们在严格的零样本推理条件下对 $16$ 种开放词汇模型配置进行了基准测试,测量了跨 IoU 阈值的定位精度、端到端推理延迟、提示引起的性能波动以及目标恢复一致性。结果表明,提示鲁棒性高度依赖于模型架构。大多数模型配置($10/16$)在使用简短类别标签时表现最佳,而自由形式查询仅在一种配置下带来最高精度。此外,相近的总体 mAP 可能掩盖了跨查询改写时目标一致恢复能力上的显著差异。Pro-Bench 支持对这些差距进行系统性评估,并促进提示鲁棒的视觉定位研究。Pro-Bench: https://pro-bench.github.io/.
cs.CV / 10 / 2609.27110

Feed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges

喂给评审团维度,而非裁决:视觉-语言美学评审模型的评分细则分解式融合
Jadhav, Amit, Beriwala, Shaurya, Kim, Beomjin
Abstract
Vision-language models (VLMs) are deployed as zero-shot judges of image aesthetics, and panels of several models are recommended, on thin evidence, as the way to make such judges reliable. On two human-rated datasets, EVA and PARA, we find that a panel of holistic judges never significantly beats its best member, whether the verdicts are averaged or fused by a learned combiner. What a panel is worth depends on what it is fed. We therefore have each model score each image on the five dimensions of a frozen, human-written rubric and fuse those scores, alongside each model's verdict, across model families with an out-of-fold combiner. The dimension scores measure what their labels claim: with the overall human score partialled out, a dimension prompt carries more attribute-specific information than the holistic prompt in 28 of 30 model-attribute cells. Fused, they beat the best single VLM in all ten three-family panels on EVA (against that best single model, +0.07 Spearman rho for the strongest trio and +0.10 for the pre-declared one, and +0.06 and +0.07 when averaged over twenty fold partitions; against the panel mean, the primary test gives +0.118 on its EVA design set), and on PARA they reach parity under Spearman rho and a small, non-significant loss under Kendall tau-b, where one model already captures 85% of the human noise ceiling. It is not a feature-count artefact: giving the same combiner an equal number of pure holistic columns, split from the same repetitions, does not reproduce it. The gain costs a few hundred labels, which do not transfer between datasets, and 4.8x the API calls on EVA; we report it with paired bootstraps and Kendall tau-b, alongside a failed pre-registration and the configurations that lost.
Chinese Translation
视觉-语言模型(VLM)被用作图像美学的零样本评审,并且有观点(证据薄弱)推荐使用多个模型组成的评审团来使此类评审更可靠。在两个人工评分数据集 EVA 和 PARA 上,我们发现:无论对整体评审(holistic judges)的裁决取平均还是用可学习的组合器进行融合,评审团的表现从未显著优于其最佳成员。评审团的价值取决于喂给它什么。因此,我们让每个模型按照一份固定的、人工撰写的评分细则(rubric)的五个维度对每张图像分别打分,并通过折外(out-of-fold)组合器将这些维度得分连同各模型的整体裁决在不同模型家族之间进行融合。维度得分确实度量了其标签所声称的内容:在剔除人工总分的影响后,在 30 个模型-属性组合中有 28 个,维度提示词比整体提示词携带更多特定属性的信息。融合后,在 EVA 数据集上所有十个三家族评审团中均超越了最佳单一 VLM(相对于该最佳单一模型,最强三人组提升 +0.07 Spearman rho,预声明的组合提升 +0.10,在二十次折划分上平均为 +0.06 和 +0.07;相对于评审团均值,主要测试在其 EVA 设计集上给出 +0.118);在 PARA 数据集上,Spearman rho 下达到持平,Kendall tau-b 下有小的非显著损失,因为其中一个模型已捕获人工噪声上限的 85%。这并非特征数量的人为产物:给同一组合器提供相同数量的、由同样重复切分得到的纯整体列,无法复现该增益。此增益的代价是几百条标注(且这些标注无法跨数据集迁移),以及在 EVA 上 4.8 倍的 API 调用;我们使用配对自助法(paired bootstraps)和 Kendall tau-b 报告结果,并同时公开一次失败的预注册以及表现不佳的配置。
cs.CV / 11 / 2609.27115

Damnatio Memoriae: Adversarially and Selectively Forgetting Identities in the Embedding Space of Face Recognition Models

除忆抹名:在人脸识别模型嵌入空间中进行对抗性与选择性遗忘
Öztürk, Ünsal, Hahn, Vedrana Krivokuća, Bhattacharjee, Sushil, Marcel, Sébastien
Abstract
A face recognition model links two images of a person recorded on separate occasions when their embedding similarity exceeds an operating threshold. We consider making chosen identities unlinkable across separate occasions while the model remains in service for the rest of the population. Deleting their images and retraining does not achieve this, since the model recognises identities never observed in training. Therefore, the embedding space must be altered against these identities, the process of which we call open-set adversarial forgetting. We propose three loss functions, one that disperses an identity's embeddings from their centroid, and two that map each image onto its own near-orthogonal target, learnt with the classifier head or fixed in advance as an almost-orthonormal frame. Each is fine-tuned alongside the classification objective on a subset of each identity's images. We evaluate them against four methods from prior work in verification and identification, at two forget scales and three backbones. Every loss acting on the embedding geometry makes the forget identities nearly unidentifiable. The orthonormal frame alone achieves strong forgetting, which holds wherever an image of that subset enters the comparison and leaves distinct forget identities unlinkable. It also surpasses a concurrent unsupervised method at a higher retain rate.
Chinese Translation
当两个人在不同场合拍摄的两张图像的嵌入相似度超过运行阈值时,人脸识别模型会将它们关联起来。我们考虑使选定的身份在不同场合之间不可关联,同时模型仍可继续为其余人群提供服务。删除其图像并重新训练无法实现这一目标,因为模型能够识别训练中从未见过的身份。因此,必须针对这些身份对嵌入空间进行修改,我们将这一过程称为开集对抗遗忘。我们提出三种损失函数:一种将身份的嵌入从其质心分散开;另外两种将每张图像映射到其各自的近正交目标上,该目标通过与分类器头联合学习获得,或预先固定为一个近似正交规范基。每种方法均在与分类目标联合的微调中,仅在每个身份的部分图像上进行。我们在验证和识别任务中,于两种遗忘规模和三种骨干网络下,将其与先前工作的四种方法进行对比评估。所有作用于嵌入几何结构的损失函数都能使被遗忘的身份几乎不可识别。仅正交规范基方法即可实现强遗忘效果,只要该子集的图像参与比对即可保持遗忘效果,并使其他不同的被遗忘身份保持不可关联。此外,在更高的保留率下,该方法优于一种同期提出的无监督方法。
cs.CV / 12 / 2609.27118

From greenhouse climate to individual leaves: an organ-resolved model of lettuce growth

从温室气候到单片叶片:一个器官级分辨的生菜生长模型
Rahman, Md Hasibur, Ahmed, Faraz, Bilal, Hafiz Muhammad, Wells, Daniel, Tobin, Dylan, Rehman, Tanzeel U.
Abstract
Greenhouse climate management aims to improve crop production while limiting energy use. This requires knowing how a crop will respond before conditions are changed. A crop digital twin can support this decision only if it represents how plant physiology and structure develop together. A unified framework was developed to simulate lettuce growth from the physiology of individual leaves. Each leaf received the conditions at its position in the canopy and contributed carbon through photosynthesis. Part of this carbon was used for maintenance and the remainder supported growth, distributed among leaves by their age, size and local environment. The predicted leaf mass, area and age generated an evolving three-dimensional plant in NVIDIA Isaac Sim. Ray tracing calculated the radiation intercepted by each leaf and returned it to photosynthesis, so structure and growth influenced each other over time. Against greenhouse measurements, the relative root mean square error was 9.5% for total dry weight and 9.2%, 12.7% and 13.1% for leaf number, canopy diameter and largest-leaf area, respectively. A 30% decrease in incident radiation reduced final dry weight by 10.4%, while the same increase raised it by 6.9%, and adding 200 ppm carbon dioxide raised it by 46.1%. Within a simulated 40-plant block, interior plants accumulated 8.6% less dry weight than border plants with identical initial states, and the leaf-specific tipburn index rose in the enclosed leaves over the period in which tipburn appeared on the greenhouse plants. Resolving individual leaves therefore explains how local exposure changes plant growth within the greenhouse. The framework provides the forward plant model needed for a bidirectional digital twin, where observations of the physical plant can update predictions and support greenhouse climate decisions.
Chinese Translation
温室气候管理旨在提高作物产量,同时限制能源消耗。这要求在改变环境条件之前就能预知作物的响应。作物数字孪生只有在能够表征植物生理与结构如何协同发育时,才能支持此类决策。本研究开发了一个统一的框架,从单片叶片的生理出发模拟生菜生长。每片叶片接收其在冠层中所处位置的环境条件,并通过光合作用贡献碳。其中一部分碳用于维持呼吸,其余部分支持生长,并依据叶片的年龄、大小和局部环境在叶片间进行分配。预测的叶片质量、面积和年龄在NVIDIA Isaac Sim中生成了一个动态演化的三维植株。光线追踪计算每片叶片截获的辐射并将其反馈至光合作用,从而使结构与生长随时间相互影响。与温室实测数据相比,总干重的相对均方根误差为9.5%,叶片数、冠层直径和最大叶面积的相对均方根误差分别为9.2%、12.7%和13.1%。入射辐射降低30%使最终干重减少10.4%,同等幅度增加则使其提高6.9%,而增加200 ppm二氧化碳使其提高46.1%。在模拟的40株植株区块内,初始状态相同的内部植株比边缘植株积累的干重少8.6%,且在温室植株出现烧心(tipburn)的时期内,被遮蔽叶片的叶尖烧焦指数上升。因此,对单片叶片的分辨解释了局部光照差异如何改变温室内的植株生长。该框架为构建双向数字孪生提供了所需的前向植物模型,使对物理植株的观测能够更新预测并支持温室气候决策。
cs.CV / 13 / 2609.27123

PEARL: A Lightweight Prompt-based Feature Interpreter Framework for Real-Time, Anonymous, and Heterogeneous Collaborative Perception

PEARL:一个面向实时、匿名和异构协同感知的轻量级基于提示的特征解释框架
Maleki, Armin, Radha, Hayder
Abstract
Heterogeneity across Collaborative Perception (CP) agents is a major challenge for emerging CP frameworks due to domain gaps from differing sensors, architectures, and training data. Prior works mitigate this challenge by aligning features in a unified space via model retraining or per-agent-type interpreters. These strategies (a) require access to neighbor configurations, (b) do not fully address real-time CP deployment, and (c) generalize poorly to unseen agents joining at run time. To overcome these challenges, we present PEARL, a Prompt-Embedding framework for Anonymous and Real-time Lightweight heterogeneous CP. PEARL supports multiple CP interpreters and selects one for a new-joining agent in real time using two lightweight, multi-scale interpreters trained in parallel: a sparse-detection (LWSD) interpreter that aligns salient regions for cooperative detection, and a dense, domain-invariant (LWDDI) interpreter that produces agent-invariant features for fast interpreter selection. Both interpreters use low-rank visual prompts to reduce computation, storage, and model complexity. Extensive experiments on simulated (OPV2V, V2XSet) and real (DAIR-V2X) datasets show that PEARL generalizes across simulated and real-world cooperative driving scenarios. Its real-time model-selection strategy yields an 8.2% Average Precision (AP) gain over a random-selection baseline while running in 1.67 ms on average. Although primarily designed for real-time CP, PEARL also outperforms state-of-the-art heterogeneous CP frameworks under traditional offline training by 5.6% AP on average while reducing communication cost by up to 34.7 times. Equally important, PEARL does not require sharing agents' configurations or model settings, thereby protecting information that may be proprietary or private. These results establish PEARL as a scalable and practical framework for heterogeneous collaborative perception.
Chinese Translation
协同感知(Collaborative Perception, CP)智能体之间的异构性是新兴协同感知框架面临的主要挑战,其原因在于不同传感器、架构和训练数据带来的域差异。已有工作通过模型重训练或针对每种智能体类型的解释器,将特征对齐到统一空间,以缓解这一挑战。但这些策略(a)需要获取邻居智能体的配置信息,(b)无法完全满足实时协同感知部署的需求,且(c)对运行时加入的未见智能体泛化能力较差。为克服这些挑战,我们提出了PEARL,一个面向匿名与实时轻量级异构协同感知的提示嵌入框架。PEARL支持多个协同感知解释器,并利用两个并行训练的轻量级多尺度解释器,为新加入的智能体实时选择合适的解释器:一个是稀疏检测解释器(LWSD),用于对齐显著区域以实现协同检测;另一个是稠密的域不变解释器(LWDDI),用于生成智能体不变的特征以实现快速解释器选择。两个解释器均采用低秩视觉提示(visual prompts)来降低计算量、存储开销和模型复杂度。在仿真数据集(OPV2V、V2XSet)和真实数据集(DAIR-V2X)上的大量实验表明,PEARL在仿真和真实世界的协同驾驶场景中均具有良好的泛化能力。其实时模型选择策略相比随机选择基线可带来8.2%的平均精度(AP)提升,同时平均运行时间仅为1.67毫秒。尽管PEARL主要面向实时协同感知设计,但在传统离线训练设置下,其平均AP仍比最先进的异构协同感知框架高出5.6%,同时通信成本最多降低34.7倍。同样重要的是,PEARL无需共享智能体的配置或模型设置,从而保护了可能涉及专有或隐私的信息。这些结果确立了PEARL作为可扩展、实用的异构协同感知框架的地位。
cs.CV / 14 / 2609.27131

Super-Resolution of Solar Magnetograms via Adaptive Stratified Ensemble Learning with Uncertainty Estimation

基于不确定性估计的自适应分层集成学习的太阳磁图超分辨率
Kandalan, Sina Norouzi, Jiang, Haodi, Wang, Jason T. L., Li, Qin
Abstract
Single-image super-resolution of Sun's photospheric magnetograms enables consistent analysis across heterogeneous space-based instruments and supports long-term studies of solar magnetic field evolution. We address the super-resolution task from SOHO/MDI (low-resolution) to SDO/HMI (high-resolution) line-of-sight (LOS) magnetograms using a modified RRDBNet architecture initialized by ESRGAN pretrained weights. Through systematic per-image diagnostic analysis, we identify image complexity as the dominant predictor of reconstruction errors. To exploit this finding, we introduce an adaptive stratified specialist ensemble (SSE) of three specialist networks with uncertainty estimation, where each specialist network is trained by images from three different complexity strata using a weighted random sampling strategy. During inference, a lightweight router based on input image statistics assigns each test image to the appropriate specialist network. Our experimental results demonstrate the good performance of the proposed ensemble and its superiority over closely related methods.
Chinese Translation
太阳光球磁图的单图像超分辨率技术能够在异构天基仪器之间实现一致的分析,并支持对太阳磁场演化的长期研究。我们使用由ESRGAN预训练权重初始化的改进型RRDBNet架构,解决从SOHO/MDI(低分辨率)到SDO/HMI(高分辨率)的视线(LOS)磁图超分辨率任务。通过系统的逐图像诊断分析,我们发现图像复杂度是重建误差的主导预测因子。为利用这一发现,我们引入了一种带不确定性估计的自适应分层专家集成方法(Stratified Specialist Ensemble,SSE),其中三个专家网络分别使用加权随机采样策略在来自三个不同复杂度层级的图像上进行训练。在推理阶段,一个基于输入图像统计特征的轻量级路由器将每张测试图像分配给相应的专家网络。实验结果表明,所提出的集成方法具有良好的性能,并优于与之密切相关的其他方法。
cs.CV / 15 / 2609.27139

A Hierarchy-Aware Video-Language Model Evaluation and Hyperbolic Baseline for Surgery

面向手术的层级感知视频-语言模型评估与双曲基线方法
Rodríguez, Ana Manzano, Mettes, Pascal, Schijven, Marlies P., Snoek, Cees G. M.
Abstract
Surgical procedures follow a phase-to-step hierarchy, yet the video-language models used to recognize them are evaluated with flat per-level metrics that ignore cross-level coherence and error structure. In this paper we make two contributions to address this problem, (i) we introduce SurgHiBench, the first hierarchy-aware evaluation suite for surgical video understanding, with three tasks measuring recognition, consistency, and severity across granularity levels. We evaluate a general-purpose CLIP model, a Euclidean surgical model, and, as second contribution: (ii) HyperSurg, a new hyperbolic model that enforces phase-step containment via entailment cones, across four (existing) datasets spanning three procedure types. The suite reveals that two models with the same accuracy can produce predictions of very different error severity, ranging from sibling confusions within the correct phase to unrelated cross-phase predictions. Hyperbolic geometry shifts predictions toward the correct procedural neighborhood, and these gains scale with the tree-likeness of each dataset's annotation hierarchy, providing a principled indicator when hierarchy-aware geometry helps.
Chinese Translation
手术流程遵循阶段到步骤的层级结构,然而用于识别手术流程的视频-语言模型却采用扁平化的分层独立指标进行评估,忽略了跨层级的连贯性与错误结构。在本文中,我们做出两项贡献以解决这一问题:(i) 我们提出了 SurgHiBench,这是首个面向手术视频理解的层级感知评估套件,包含三项任务,用于在不同粒度层级上衡量识别能力、一致性与错误严重程度。我们在跨越三类手术类型的四个(现有)数据集上评估了一个通用 CLIP 模型、一个欧氏空间的手术模型,以及作为第二项贡献的 (ii) HyperSurg——一种新的双曲模型,其通过蕴含锥(entailment cones)来强化阶段对步骤的包含关系。该评估套件揭示出,两个准确率相同的模型可能产生错误严重程度差异很大的预测,错误范围从正确阶段内同级步骤之间的混淆,到完全无关的跨阶段预测。双曲几何使预测向正确的手术流程邻域偏移,且这些增益随各数据集标注层级树状程度(tree-likeness)的提高而扩展,为判断层级感知几何何时有效提供了一个有原则的指标。
cs.CV / 16 / 2609.27142

MINER: Multi-crop INference-time Enhancement for Rare-Object Retrieval with Frozen Dual Encoders

MINER:基于冻结双编码器的多裁剪推理时增强方法,用于稀有目标检索
Alquwayfili, Abdulmalik, AlMeshal, Faisal, Almajnouni, Jumanah, Alamri, Huda Abdulhadi, Khan, Muhammad Kamran J
Abstract
Text-to-image retrieval with frozen dual encoders degrades when the query names a small, visually subordinate object in a cluttered scene: a single global image embedding underrepresents the localized visual evidence. We present MINER, a training-free inference framework that augments a frozen dual encoder's global image embedding with a small bank of region-level embeddings and a hubness-correcting similarity rescoring, recovering visual evidence that global pooling underweights. To evaluate this setting, we introduce ROCS, a benchmark built from high-clutter subsets of Flickr30K and MS COCO whose images are re-captioned to name a single low-salience object. Experiments on CLIP, SigLIP, and SigLIP 2 show that MINER improves retrieval on every backbone, on ROCS and on the standard splits. Analyses show that these gains come primarily from broader spatial coverage rather than precise crop placement, revealing a simple and general way to recover localized evidence from frozen representations. Code: https://github.com/aalquwayfili/MINER. Dataset: https://huggingface.co/datasets/aalquwayfili/ROCS.
Chinese Translation
当查询指向杂乱场景中一个视觉上处于次要地位的小目标时,基于冻结双编码器的文本到图像检索性能会显著下降:单一的图像全局嵌入不足以代表局部化的视觉证据。我们提出了MINER,这是一个无需训练的推理框架,它通过一组少量的区域级嵌入来增广冻结双编码器的全局图像嵌入,并结合一种中心度校正(hubness-correcting)的相似度重打分方法,恢复全局池化所低估的视觉证据。为评估该场景,我们构建了ROCS基准,其基于Flickr30K和MS COCO的高杂乱度子集,并对图像重新标注描述,使其指向单个低显著度的目标。在CLIP、SigLIP和SigLIP 2上的实验表明,MINER在所有骨干模型上、在ROCS和标准数据划分上均提升了检索性能。分析显示,这些提升主要来自更广泛的空间覆盖,而非精确的裁剪定位,这揭示了从冻结表示中恢复局部化证据的一种简单而通用的方法。代码:https://github.com/aalquwayfili/MINER。数据集:https://huggingface.co/datasets/aalquwayfili/ROCS。
cs.CV / 17 / 2609.27143

A Systematic Evaluation of Infrastructure-Based Radar System for Highway Traffic Monitoring

基于基础设施的雷达系统用于高速公路交通监测的系统性评估
Zhu, Tianheng, Chang, Woei-chyi, Riaz, Alamss, Hasanzadeh, Sogand, Feng, Yiheng
Abstract
Infrastructure-based radar systems offer robust and long-range solutions for traffic monitoring, yet their detection and tracking performance under real-world conditions remains insufficiently evaluated. This study introduces DRaT (Drone and Radar Trajectories), a dual-modality dataset of naturalistic vehicle trajectories collected at a highway merging segment in Fort Worth, Texas, to systematically assess radar sensing performance against drone-derived ground truth. The performance is evaluated at three levels: individual vehicle detection, trajectory tracking, and macroscopic traffic parameter estimation. For individual vehicle detection, the radar achieves an overall precision of 78% and a recall of 57%, with degraded performance under congested traffic conditions and at longer distances. At the trajectory level, the radar demonstrates reasonably strong tracking performance (IDF1 = 0.699), maintaining reliable vehicle identities when tracks are successfully established. For macroscopic traffic flow metrics, the radar accurately estimates space-mean speed (MAPE < 4%) but underestimates density and volume by approximately 23% due to missed detections. The paper also discusses practical deployment considerations and potential downstream applications of roadside radar sensing systems. To support reproducible research on infrastructure-based sensing systems, we have open-sourced the DRaT dataset on Zenodo: https://zenodo.org/records/20171110.
Chinese Translation
基于基础设施的雷达系统为交通监测提供了鲁棒且远距离的解决方案,但其在真实条件下的检测与跟踪性能仍缺乏充分评估。本研究提出了DRaT(Drone and Radar Trajectories,无人机与雷达轨迹)数据集,这是一个在德克萨斯州沃思堡一处高速公路合流路段采集的自然驾驶车辆轨迹双模态数据集,用于以无人机采集的真值数据为基准,系统评估雷达感知性能。评估在三个层面进行:单车车辆检测、轨迹跟踪以及宏观交通参数估计。在单车检测方面,雷达的总体精确率为78%,召回率为57%,在交通拥堵条件和远距离情况下性能有所下降。在轨迹层面,雷达表现出较强的跟踪性能(IDF1 = 0.699),并在轨迹成功建立后能够可靠地保持车辆身份。在宏观交通流指标方面,雷达能够准确估计区间平均速度(MAPE < 4%),但由于漏检,对密度和流量的估计偏低约23%。本文还讨论了路侧雷达感知系统的实际部署考量及其潜在的下游应用。为支持基于基础设施的感知系统的可复现研究,我们已在Zenodo上开源了DRaT数据集:https://zenodo.org/records/20171110。
cs.CV / 18 / 2609.27149

Temporally Ordered Region-Token Mamba with Logit-Space Diffusion for Remote Sensing Change Detection

基于时间有序区域-令牌Mamba与Logit空间扩散的遥感变化检测
Sen, Anuvab, Chatterjee, Maneet, Ghosh, Aparup, Sen, Udayon, Aditya, Arnav, Zhang, Yixin
Abstract
Remote sensing change detection requires both global reasoning across bitemporal images and precise localization of changed regions. However, dense attention is computationally expensive for high-resolution imagery, while conventional feature fusion and coarse decoding may inadequately separate genuine changes from appearance variations or preserve object boundaries. We present Bitemporal Mamba-Diffusion for Change Detection (BMD-CD), which combines temporally structured state-space modeling with logit-space diffusion refinement. BMD-CD converts deep bitemporal features into region tokens and arranges them in explicit temporal partitions before bidirectional state-space propagation. Its Bitemporal Ordered Mamba Operator enables long-range cross-temporal interaction with linear sequence complexity, while Orthogonal Feature Disentanglement forms a change-oriented output and a complementary rotated output using learned pairwise rotations and unchanged-region consistency. Multiscale decoding then produces coarse change logits, which are refined through a five-step Conditional Diffusion Decoder operating directly in logit space. Experiments on LEVIR-CD, WHU-CD, DSIFN-CD, CDD, and S2Looking demonstrate strong performance across diverse change-detection settings. BMD-CD achieves F1 scores of 93.7%, 96.0%, 97.8%, and 99.0% on the four standard benchmarks and improves 3-pixel Boundary-F1 to 87.7% and 91.4% on LEVIR-CD and WHU-CD, respectively. The full model requires 32.09 GFLOPs and 47 ms per 256 x 256 image pair, while also showing zero-shot transfer to ValaisCD and B-FLAIR-test. Our code is available at https://github.com/Aparup2139/Public_WACV/
Chinese Translation
遥感变化检测既需要对双时相图像进行全局推理,也需要对变化区域进行精确定位。然而,密集注意力机制对于高分辨率影像的计算开销过大,而传统的特征融合与粗糙解码可能无法充分区分真实变化与外观变化,或无法完整保持物体边界。我们提出了用于变化检测的双时相Mamba-扩散模型(Bitemporal Mamba-Diffusion for Change Detection,BMD-CD),该方法将时间结构化的状态空间建模与logit空间扩散精化相结合。BMD-CD将深层双时相特征转换为区域令牌(region tokens),并在双向状态空间传播之前将其排列在显式的时间分区中。其双时相有序Mamba算子(Bitemporal Ordered Mamba Operator)能够以线性序列复杂度实现长程跨时相交互;同时,正交特征解耦(Orthogonal Feature Disentanglement)利用学习到的成对旋转和未变化区域一致性,形成面向变化的输出和一个互补的旋转输出。随后,多尺度解码产生粗略的变化logits,并通过一个直接在logit空间中运行、仅需五步的条件扩散解码器(Conditional Diffusion Decoder)对其进行精化。在LEVIR-CD、WHU-CD、DSIFN-CD、CDD和S2Looking数据集上的实验表明,该方法在多种变化检测场景下均表现出色。BMD-CD在四个标准基准上分别取得了93.7%、96.0%、97.8%和99.0%的F1分数,并将LEVIR-CD和WHU-CD上3像素边界F1(Boundary-F1)分别提升至87.7%和91.4%。完整模型处理每对256 x 256图像仅需32.09 GFLOPs和47毫秒,同时在ValaisCD和B-FLAIR-test上展现出零样本迁移能力。代码已开源:https://github.com/Aparup2139/Public_WACV/
cs.CV / 19 / 2609.27194

Diverse by Design: Architectural Constraints for Prototype-Based Interpretability

设计即多样性:面向原型可解释性的架构约束
Lin, Xinmiao, Wright, Matthew
Abstract
Prototype-based neural networks provide inherent interpretability through case-based reasoning, yet suffer from critical limitations: prototypes converge to redundant features, fail to capture diverse semantic parts, and lack quantitative interpretability assessment. We propose Diversity-Aware Prototype Learning (DAPL), which enforces prototype diversity through architectural constraints rather than explicit regularization. Our approach leverages multi-head self-attention with strict one-to-one attention-to-prototype mapping, ensuring each prototype specializes in distinct visual features. We further introduce foreground-aware training to focus prototypes on semantically meaningful regions and develop comprehensive evaluation metrics (Coverage and Diversity) for quantitative interpretability assessment. Experiments on CUB-200-2011 demonstrate substantial improvements: DAPL with foreground-aware training achieves 81.69\% accuracy with 0.596 Coverage and 0.427 Diversity, providing the best overall balance across all evaluated prototype-based methods. Code is available at https://github.com/xinmiaolin/DAPL.
Chinese Translation
基于原型的神经网络通过基于案例的推理提供了内在的可解释性,但仍存在关键局限:原型会收敛到冗余特征,无法捕捉多样化的语义部件,且缺乏定量的可解释性评估。我们提出多样性感知原型学习(Diversity-Aware Prototype Learning, DAPL),通过架构约束而非显式正则化来强制实现原型多样性。我们的方法利用多头自注意力机制,并严格保持注意力头与原型之间的一对一映射,确保每个原型专注于不同的视觉特征。我们进一步引入前景感知训练,使原型聚焦于语义上有意义的区域,并开发了全面的评估指标(覆盖率 Coverage 和多样性 Diversity)用于定量可解释性评估。在 CUB-200-2011 数据集上的实验表明了显著的改进:采用前景感知训练的 DAPL 达到了 81.69% 的准确率,覆盖率为 0.596,多样性为 0.427,在所有被评估的基于原型的方法中提供了最佳的总体平衡。代码可在 https://github.com/xinmiaolin/DAPL 获取。
cs.CV / 20 / 2609.27208

Benchmarking Active Spot Selection for Cost-Efficient Spatial Transcriptomics

面向高性价比空间转录组学的主动位点选择基准测试
Zhu, Zheyu, Zhu, Junchao, Liu, Fengbei, Yao, Tianyuan, Xu, Gelei, Cannon, John, Yang, Haichun, Huo, Yuankai, Sabuncu, Mert R., Deng, Ruining
Abstract
Spatial transcriptomics (ST) measures gene expression in tissue context, but dense capture grids can be costly and may repeatedly sample morphologically similar regions. Most active learning strategies were developed for categorical labels and independent samples. We conduct a retrospective pool-based benchmark of active learning versus uniform Random sampling for ST, where expression vectors are high-dimensional and continuous and candidates are spatially correlated. Using two fully profiled public ST cohorts, we mask candidate expression vectors and simulate multi-round selection with uncertainty-based Monte Carlo dropout (MC-dropout) and temporal output discrepancy (TOD), and diversity-based CoreSet and TypiClust-inspired selection. We compare 160 completed configurations at 5%, 10%, 30%, and 50% of the fold-wide training spot pool under patient-level cross-validation, with a separate full-label reference. Within each budget, strategies share the selection schedule, morphology-to-expression predictor, and optimization protocol. We assess mean per-gene within-slide Pearson correlation coefficient (PCC), expression-cluster agreement, and Moran's I fidelity. On HER2-positive breast cancer, pooled mean PCC differences from Random across the four active strategies were -0.0176, -0.0117, +0.0056, and +0.0057 at 5%, 10%, 30%, and 50%, respectively. On cutaneous squamous cell carcinoma (cSCC), three strategies were below Random at 5%, and all four were below Random at 10%. On HER2-positive breast cancer, CoreSet and MC-dropout had lower PCC but higher expression-cluster agreement than Random at the two smallest budgets; this pattern did not reproduce on cSCC. Under the reported fixed training horizons, the evaluated active strategies do not consistently improve on Random at small budgets, and rankings depend on the evaluation measure.
Chinese Translation
空间转录组学(ST)能够在组织背景下测量基因表达,但密集的捕获网格成本高昂,且可能重复采样形态相似的区域。大多数主动学习策略是为类别型标签和独立样本而设计的。我们针对ST开展了一项回顾性的基于样本池的基准测试,比较主动学习与均匀随机采样,其中表达向量是高维且连续的,候选样本之间存在空间相关性。我们使用两个已完整测量的公开ST队列,对候选表达向量进行遮蔽,并模拟多轮选择:包括基于不确定性的蒙特卡洛 dropout(MC-dropout)和时间输出差异(TOD),以及基于多样性的 CoreSet 和受 TypiClust 启发的选择方法。在患者级交叉验证下,我们在每个切片训练位点池的5%、10%、30%和50%预算下比较了160个完整配置,并另设一个全标签参考。在每个预算内,各策略共享相同的选点计划、形态学到表达的预测器和优化协议。我们评估每基因的切片内平均皮尔逊相关系数(PCC)、表达聚类一致性以及 Moran's I 保真度。在HER2阳性乳腺癌数据上,四种主动学习策略在5%、10%、30%和50%预算下相对随机采样的合并平均PCC差异分别为-0.0176、-0.0117、+0.0056和+0.0057。在皮肤鳞状细胞癌(cSCC)数据上,5%预算下有三种策略低于随机采样,10%预算下四种策略均低于随机采样。在HER2阳性乳腺癌数据上,在两个最小预算下,CoreSet 和 MC-dropout 的PCC虽低于随机采样,但表达聚类一致性更高;该模式在cSCC上未能复现。在所报告的固定训练轮次下,所评估的主动学习策略在小预算下并未持续优于随机采样,且排名取决于评估指标的选择。
cs.CV / 21 / 2609.27217

Learning Spectral Allocation: A Fractional Diffusion Framework for Adaptive Volumetric Segmentation

学习频谱分配:一种用于自适应体积分割的分数阶扩散框架
Shen, Yi-Hui, Li, Tie-Qiang
Abstract
We address adaptive computation in 3D medical image segmentation: instead of designing another backbone, we ask how much spectral mixing each network stage needs and let optimization answer. We derive FHEAT, a two-parameter operator family, from the discrete cosine transform (DCT) solution of a fractional heat equation. A fractional order alpha and a diffusion strength D govern the operator, and at D=0 it is exactly the identity. Reparametrized by the semigroup time tau = D*alpha, same-resolution instances compose exactly, so any distribution of diffusion across same-resolution stages amounts to a single Sobolev-type regularizer of learned strength. This identity limit lets the optimizer of each layer, not the designer, decide whether global mixing is needed and how sharp it should be. We instantiate FHEAT in a lightweight U-shaped architecture (Light-UNETR) paired with a Kolmogorov-Arnold mixer (KAN3D) with adaptive rational activations, yielding FHEAT-Seg. At 5% to 20% label rates on three public benchmarks, training produces gradient-driven spectral sparsification: seven of the eight stage-level operators drive D to zero, and the survivor saturates at the sharpest low-pass (alpha ~ 0.9) in the decoder layer feeding the semi-supervised attention map. The retired layers become exact identity shortcuts at inference, cutting FLOPs from 4.29G to 0.90G (a 79% drop) at 0.975M parameters. Under a standard semi-supervised protocol, FHEAT-Seg reaches Dice scores of 90.47% (left atrium), 78.79% (Pancreas-CT), and 81.90% (BraTS 2019), ahead of five semi-supervised methods and the Light-UNETR baseline. The large variant also surpasses Light-UNETR-L under full supervision (Dice 93.09%, 85.11%, and 87.19%) with 2.851M parameters and 55.75G FLOPs. These results suggest that the allocation of spectral computation is a learnable property of optimization dynamics, not a manual design commitment.
Chinese Translation
我们研究三维医学图像分割中的自适应计算问题:我们不再设计新的骨干网络,而是探究每个网络阶段需要多少频谱混合,并让优化过程来给出答案。我们从分数阶热传导方程的离散余弦变换(DCT)解出发,推导出 FHEAT——一个双参数算子族。分数阶次 alpha 和扩散强度 D 共同控制该算子,且当 D=0 时它恰好是恒等算子。通过半群时间 tau = D*alpha 进行重参数化后,同分辨率的算子实例可精确复合,因此扩散在同分辨率各阶段之间的任意分布等价于一个学习所得强度的 Sobolev 型正则化项。这一恒等极限使得是否需要全局混合以及混合的锐度由每一层的优化器(而非设计者)来决定。我们将 FHEAT 实例化于一个轻量级 U 形架构(Light-UNETR)中,并配以具有自适应有理激活函数的 Kolmogorov-Arnold 混合器(KAN3D),构成 FHEAT-Seg。在三个公开基准上以 5% 至 20% 的标注率进行训练时,产生了梯度驱动的频谱稀疏化:八个阶段级算子中有七个将 D 驱动至零,而幸存的那一个在向半监督注意力图提供输入的解码器层中饱和于最尖锐的低通设置(alpha 约为 0.9)。被淘汰的层在推理时变为精确的恒等捷径,使 FLOPs 从 4.29G 降至 0.90G(下降 79%),参数量为 0.975M。在标准半监督协议下,FHEAT-Seg 达到了 90.47%(左心房)、78.79%(Pancreas-CT)和 81.90%(BraTS 2019)的 Dice 分数,领先于五种半监督方法以及 Light-UNETR 基线。其大型变体在完全监督下(Dice 分别为 93.09%、85.11% 和 87.19%)以 2.851M 参数量和 55.75G FLOPs 也超越了 Light-UNETR-L。这些结果表明,频谱计算的分配是优化动力学的一种可学习属性,而非需要手动做出的设计承诺。
cs.CV / 22 / 2609.27227

Surgical Kinematics from Monocular Video with Learned Articulated Motion Constraints

基于学习关节运动约束的单目视频手术运动学估计
Turkcan, Mehmet Kerem, Samal, Soham, Kostic, Zoran
Abstract
Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic reconstruction network for estimating instrument position, orientation and jaw angle from monocular video. Our visual representation combines global attention pooling of frozen DINOv3 features with local pooling at instrument landmarks from fine-tuned SAM 3.1 masks. Our shared Transformer encoder and temporal convolutional heads integrate this representation with mask geometry, monocular depth and visual state estimates from arm-specific multilayer regression networks. Our position branch predicts displacement magnitude and direction separately to preserve traveled distance. We fit trajectories to predicted state observations and motion increments by differentiable weighted least squares, expressing quaternion observations relative to cumulative predicted rotations to obtain a quadratic orientation objective. We evaluate reconstruction across 2,802 Open-H episodes. Compared with LiveMAE on the main Open-H benchmark, our method reduces path-length mean absolute error from 0.45 to 0.34\,cm and increases temporal mean average precision for motion segmentation from 44.54\% to 54.44\%.
Chinese Translation
机器人手术的客观评估依赖器械运动学数据,而当仅有视频可用时,这些数据必须被重建。我们提出了一种运动学重建网络,用于从单目视频估计器械的位置、姿态和夹爪角度。我们的视觉表征将冻结的 DINOv3 特征的全局注意力池化与基于微调 SAM 3.1 掩码的器械关键点局部池化相结合。我们的共享 Transformer 编码器和时序卷积头将该表征与掩码几何信息、单目深度以及来自各机械臂专用多层回归网络的视觉状态估计相融合。我们的位置分支分别预测位移的大小和方向,以保留行进距离。我们通过可微加权最小二乘法,将轨迹拟合到预测的状态观测和运动增量上,并将四元数观测表示为相对于累积预测旋转的形式,从而获得二次型的姿态目标函数。我们在 2,802 个 Open-H 手术片段上评估了重建性能。在 Open-H 主基准上与 LiveMAE 相比,我们的方法将路径长度平均绝对误差从 0.45 cm 降低至 0.34 cm,并将运动分割的时序平均精度(mAP)从 44.54% 提升至 54.44%。
cs.CV / 23 / 2609.27238

Strip Convolution and Direction-Aware Exclusion Loss for Oriented Ship Detection

面向旋转船舶检测的条带卷积与方向感知排斥损失
Chen, Bin, Liu, Yuanyuan, Yang, Peng, Lu, Chao
Abstract
Oriented ship detection in very high resolution (VHR) remote sensing imagery remains challenging due to elongated hull geometry and dense target distributions in complex port scenes. Existing methods typically address geometric representation and duplicate suppression separately. To jointly tackle these issues, we propose an oriented ship detector with two complementary components. The C3k2_Strip module employs orthogonal strip convolutions to better capture elongated hull structures, while the Class-Aware Direction-Aware Exclusion Loss (CA-DAEL) suppresses redundant predictions using class, direction, and confidence cues. Experiments on HRSC2016 and DIOR-R achieve 78.45% and 53.71% mAP50:95, respectively, with only 2.91M parameters. On HRSC2016, the proposed method improves mAP50:95 by 6.32 percentage points over the YOLOv11-OBB baseline, demonstrating its effectiveness for accurate oriented ship detection.
Chinese Translation
在甚高分辨率(VHR)遥感影像中进行旋转船舶检测仍然具有挑战性,其原因在于船体呈细长几何形状,且复杂港口场景中目标分布密集。现有方法通常将几何表征与重复预测抑制分开处理。为联合解决这些问题,我们提出了一种包含两个互补组件的旋转船舶检测器。C3k2_Strip 模块采用正交条带卷积(strip convolutions)以更好地捕获细长的船体结构,而类感知方向感知排斥损失(Class-Aware Direction-Aware Exclusion Loss, CA-DAEL)则利用类别、方向和置信度线索抑制冗余预测。在 HRSC2016 和 DIOR-R 数据集上的实验分别取得了 78.45% 和 53.71% 的 mAP50:95,且参数量仅为 2.91M。在 HRSC2016 上,所提方法相比 YOLOv11-OBB 基线将 mAP50:95 提升了 6.32 个百分点,证明了其在精确旋转船舶检测方面的有效性。
cs.CV / 24 / 2609.27264

GaussPDE: Graph-Based Partial Differential Equation-Driven Rendering for 3D Gaussian Splatting

GaussPDE:基于图的偏微分方程驱动的3D高斯泼溅渲染方法
Yue, Haoyuan, Ye, Fengyuan, Li, Ziyin
Abstract
We present GaussPDE, a framework that injects physically structured partial differential equation (PDE) dynamics into pretrained 3D Gaussian scenes without mesh extraction, voxelization, or retraining. Our key observation is that PDE rendering requires not only accurate appearance, but also a reliable discrete computational domain. We therefore first introduce camera-aware regularization during 3DGS reconstruction to suppress camera-near floaters and oversized primitives that would create unstable graph topology. We then construct an active Gaussian graph using covariance-aware distances and opacity, appearance, and boundary-aware conductance, enabling mass-weighted graph Laplacian PDE evolution directly over Gaussian primitives. The evolving scalar PDE state is coupled back to rendering by modifying the direct-current spherical harmonic color coefficients while preserving geometry, opacity, and view-dependent rendering behavior. Experiments on real and synthetic scenes show that GaussPDE produces stable, controllable, and spatially coherent dynamic visualizations, with reduced cross-boundary leakage compared with baselines.
Chinese Translation
我们提出了GaussPDE,一个将具有物理结构的偏微分方程(PDE)动力学注入预训练3D高斯场景的框架,且无需网格提取、体素化或重新训练。我们的关键观察是:PDE渲染不仅需要精确的外观表现,还需要一个可靠的离散计算域。因此,我们首先在3DGS重建过程中引入相机感知的正则化,以抑制靠近相机的漂浮物和过大的基元,避免其产生不稳定的图拓扑结构。随后,我们利用协方差感知的距离以及不透明度、外观和边界感知的电导,构建一个活跃的高斯图,从而能够直接在高斯基元上进行质量加权的图拉普拉斯PDE演化。演化过程中的标量PDE状态通过修改直流球谐颜色系数耦合回渲染过程,同时保持几何、不透明度以及视角相关的渲染行为不变。在真实场景和合成场景上的实验表明,GaussPDE能够生成稳定、可控且空间连贯的动态可视化效果,与基线方法相比减少了跨边界泄漏。
cs.CV / 25 / 2609.27274

High Dynamic Range Video Reconstruction from Single-Exposure Raw Sequences

基于单曝光Raw序列的高动态范围视频重建
Zhang, Tao, Su, Peixian, Gao, Xingyu, Zou, Yunhao, Lu, Yu, Zhu, Zunjie, Zheng, Bolun, Fu, Ying, Yan, Chenggang
Abstract
Due to the limited dynamic range of conventional image sensors, captured low dynamic range (LDR) video often suffers from highlight clipping and shadow detail loss, making high-quality high dynamic range (HDR) reconstruction from single-exposure sequences highly challenging without alternating exposures or extra hardware. Alternating-exposure HDR methods sacrifice frame rate and struggle with motion alignment, making them impractical for real-world capture. To address this, we propose RawHDRV, an end-to-end framework for single-exposure Raw video HDR reconstruction, that fundamentally exploits the linear response and channel-specific characteristics of Bayer data. Specifically, it features a channel-decomposition temporal alignment and fusion strategy that processes Bayer channels separately to exploit their distinct exposure characteristics, together with exposure-aware weighted fusion. It further incorporates an exposure complementarity mask-guided restoration module that leverages inter-frame exposure redundancy to adaptively fuse reliable information and suppress saturation artifacts, and introduces a mask-guided color loss that combines normalized error constraints with gradient smoothing to enhance highlight recovery. Furthermore, we construct a large-scale mobile Raw-HDR video dataset with per-frame HDR annotations. Experiments show that our method achieves the state-of-the-art results in all metrics, demonstrating superior spatial quality and temporal stability under extreme exposure conditions. The code is available at https://github.com/supeixian/RawHDRV.
Chinese Translation
由于传统图像传感器的动态范围有限,所采集的低动态范围(LDR)视频常常存在高光截断和阴影细节丢失的问题,这使得在不采用交替曝光或额外硬件的条件下,从单曝光序列重建高质量的高动态范围(HDR)视频极具挑战性。交替曝光HDR方法会牺牲帧率且难以应对运动对齐问题,因而不适用于真实场景采集。为解决这一问题,我们提出了RawHDRV,一个用于单曝光Raw视频HDR重建的端到端框架,其从根本上利用了Bayer数据的线性响应特性和通道特异性。具体而言,该框架包含一种通道分解的时序对齐与融合策略,分别处理Bayer各通道以利用其不同的曝光特性,并结合曝光感知的加权融合。此外,框架引入了曝光互补性掩码引导的恢复模块,利用帧间曝光冗余自适应地融合可靠信息并抑制饱和伪影;同时提出了掩码引导的颜色损失函数,将归一化误差约束与梯度平滑相结合以增强高光恢复能力。我们还构建了一个大规模的移动端Raw-HDR视频数据集,提供逐帧HDR标注。实验表明,我们的方法在所有指标上均取得了最先进的结果,在极端曝光条件下展现出卓越的空间质量和时间稳定性。代码已发布于 https://github.com/supeixian/RawHDRV。
cs.CV / 26 / 2609.27317

Breaking Weather-Content Coupling: Type-Severity Guided Progressive Disentanglement for All-in-One Infrared Restoration

解耦天气内容耦合:面向一体化红外图像复原的类型-严重度引导渐进式解耦方法
Wang, Xinyao, He, Lijun, Ren, Zhihan, Li, Fan
Abstract
Infrared (IR) imaging is crucial for autonomous driving, remote sensing, and other perception tasks. However, adverse weather may introduce fake structural responses that are entangled with real thermal structures. Existing IR restoration methods are typically designed for a single degradation type or directly reconstruct from degradation-entangled representations. Consequently, they struggle to distinguish intrinsic thermal structures from weather-induced fake responses and to accommodate spatially varying degradation severity, leading to artifacts or the over-suppression of weak but meaningful thermal responses. To address these issues, we propose TSGPD-IR, a type-severity guided progressive disentanglement network for all-in-one infrared restoration that factorizes restoration guidance into task-level weather semantics and region-level degradation severity. Specifically, a Weather and Semantic Co-Guided Multi-Level Prompt Generation Module combines global weather semantics with stage-wise local features to generate adaptive prompts that progressively suppress degradation-induced responses while preserving intrinsic thermal structures. To complement global weather semantics with spatial restoration control, a Proxy-Supervised Regional Degradation Estimator derives severity supervision without manual annotations and predicts spatially varying degradation priors. Guided by these cues, a Multi-Source Collaborative Expert Selection Strategy uses a shared branch to preserve weather-invariant thermal structures and hierarchical routing to select weather-specific expert pools and severity-compatible regional experts. This design progressively separates degradation interference from genuine thermal content and enables region-adaptive restoration, reducing both residual artifacts and over-suppression.
Chinese Translation
红外(IR)成像对自动驾驶、遥感及其他感知任务至关重要。然而,恶劣天气可能引入与真实热结构相纠缠的虚假结构响应。现有的红外复原方法通常仅针对单一退化类型设计,或直接从退化纠缠的表征中进行重建。因此,它们难以将固有的热结构与天气引起的虚假响应区分开来,也难以适应空间变化的退化严重度,从而导致伪影或对微弱但有意义的热响应的过度抑制。为解决这些问题,我们提出了 TSGPD-IR,一种面向一体化红外复原的类型-严重度引导渐进式解耦网络,它将复原指导分解为任务级天气语义和区域级退化严重度。具体而言,天气与语义协同引导的多层次提示生成模块(Weather and Semantic Co-Guided Multi-Level Prompt Generation Module)将全局天气语义与各阶段局部特征相结合,生成自适应提示,在逐步抑制退化引起的响应的同时保留固有的热结构。为了用空间复原控制来补充全局天气语义,代理监督的区域退化估计器(Proxy-Supervised Regional Degradation Estimator)无需人工标注即可获得严重度监督,并预测空间变化的退化先验。在这些线索的引导下,多源协同专家选择策略(Multi-Source Collaborative Expert Selection Strategy)利用共享分支保留对天气不变的热结构,并通过分层路由选择特定天气的专家池和与严重度匹配的区域专家。该设计逐步将退化干扰与真实的热内容分离,并实现区域自适应复原,同时减少了残余伪影和过度抑制。
cs.CV / 27 / 2609.27327

Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows

视觉语言模型能否分析以人为中心的视频?模型能力映射与人机协同工作流程
Shen, Xiyuan, Lyu, Jiuyang, Hwang, Seokhyun, Yao, Huanfen, Patel, Shwetak, Zhang, Zhihan, Wobbrock, Jacob O.
Abstract
Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language models (VLMs) become increasingly capable of analyzing video, they offer opportunities to automate this traditionally human-intensive process. Yet a central question remains: when can VLMs analyze human-centered video independently, and when does reliable analysis still require human involvement? To address this question, we first characterize video analysis practices in human-centered research. We systematically analyze all 1,702 CHI 2026 full papers and identify 125 that annotate videos. Through iterative coding, we derive a five-dimensional taxonomy spanning analytic purpose, viewpoint, phenomenon, reasoning requirement, and annotation authority. Grounded in recurring annotation tasks captured by this taxonomy, we construct a benchmark of 15 representative tasks from open datasets to map the capabilities and limitations of a general-purpose VLM. We examine the division of labor between humans and VLMs by comparing three annotation workflows: VLM alone, human alone, and human verification of VLM outputs. Across tasks, VLM-alone annotation approaches human accuracy on average (HNS = 97.0, where 100 denotes human-alone performance), demonstrating substantial potential to automate human-centered video analysis. Human verification achieves the highest accuracy (HNS = 121.5) while reducing human annotation time by 48.9% and monetary cost by 31.3%-44.5% relative to human-alone annotation. Our findings connect real-world human-centered video analysis tasks and current VLM capabilities, and clarify how human-AI collaboration can make VLM-assisted analysis reliable and efficient.
Chinese Translation
视频为人类行为、互动及其所处情境提供了丰富的记录,为理解人类和开展以人为中心的研究提供了重要证据。随着视觉语言模型(VLM)分析视频的能力日益增强,它们为自动化这一传统上高度依赖人力的过程提供了机会。然而,一个核心问题仍然存在:VLM何时能够独立分析以人为中心的视频,何时仍需要人工参与才能实现可靠分析?为回答这一问题,我们首先刻画了以人为中心研究中的视频分析实践。我们系统地分析了CHI 2026的全部1,702篇完整论文,识别出125篇标注了视频的论文。通过迭代编码,我们构建了一个涵盖分析目的、视角、现象、推理需求和标注权威性五个维度的分类体系。基于该分类体系所捕捉的常见标注任务,我们从公开数据集中构建了包含15个代表性任务的基准,以映射通用VLM的能力与局限。我们通过比较三种标注工作流程——仅使用VLM、仅使用人工、以及人工验证VLM输出——来考察人类与VLM之间的分工。在各项任务中,仅使用VLM的标注平均上接近人工准确率(HNS = 97.0,其中100代表仅人工标注的性能),展现出自动化以人为中心视频分析的巨大潜力。人工验证工作流程取得了最高准确率(HNS = 121.5),同时相较于仅人工标注,将人工标注时间减少48.9%,货币成本降低31.3%–44.5%。我们的研究发现将现实世界中以人为中心的视频分析任务与当前VLM的能力联系起来,并阐明了人机协同如何使VLM辅助的分析更加可靠和高效。
cs.CV / 28 / 2609.27356

Beyond Mean Foils: Auditing Worst-Foil Specificity in Frozen CLIP Region Explanations

超越均值对照区域:审计冻结CLIP区域解释中的最坏对照区域特异性
Liu, Kaixin, Ye, Zhipeng, Jiang, Feng, Wang, Zhenghao, Wu, Qihang
Abstract
A region can overlap a target object yet contribute more to another class. We test regions selected by Cluster-based Concept Importance (CCI) in frozen CLIP. Across COCO and VOC with two checkpoints, 41.08-64.78% of regions that pass overlap and mean-contrast checks fail against the strongest competing class. Removing competitors annotated in the image leaves 39.69-63.64% failing. We then test all eight candidate regions per image. An alternative passes the test for 6.25-7.84% of failures on COCO and 27.40-31.15% on VOC. Requiring it to preserve the original target-score drop within $\epsilon = 0.02$ reduces these rates to 0.16-0.98%. Available regions and target-drop tolerance constrain repair; relaxing the tolerance increases repair opportunities.
Chinese Translation
一个区域可以与目标物体重叠,却对另一个类别贡献更大。我们测试了在冻结CLIP中由基于聚类的概念重要性方法(Cluster-based Concept Importance, CCI)所选择的区域。在COCO和VOC数据集及两个模型检查点上,通过重叠与均值对比检验的区域中,有41.08%–64.78%在面对最强的竞争类别时会失效。剔除图像中已标注的竞争物体后,仍有39.69%–63.64%的区域失效。随后我们测试每张图像中全部八个候选区域。在失效情形中,有6.25%–7.84%(COCO)和27.40%–31.15%(VOC)的案例存在可通过检验的替代区域。若要求替代区域在ε=0.02的容差内保持原有的目标分数下降幅度,上述比例分别降至0.16%–0.98%。可用区域数量和目标分数下降容差限制了修复效果;放宽容差会增加修复机会。
cs.CV / 29 / 2609.27370

Geometry-Conditioned Visual Place Recognition in Natural Environments

自然环境中的几何条件化视觉位置识别
Nedov, Walter, Rahman, Saimunur, Katuwandeniya, Kavindie, Hall, David, Roy, Kaushik, Moghadam, Peyman
Abstract
Visual Place Recognition (VPR) in natural environments remains challenging due to repetitive vegetation, sparse distinctive landmarks, and substantial appearance and viewpoint variation across traversals. While visual observations of the same place can change considerably, their underlying spatial structure is often more persistent. We exploit this complementary geometric consistency through Depth-Aware Distillation (DAD), which conditions the token representations of a pretrained Vision Foundation Model (VFM) on geometry inferred by a Geometric Foundation Model (GFM), without any depth sensor. Rather than treating geometry as an additional input modality, DAD projects image-aligned depth into the VFM token space and selectively modulates visual representations through channel-wise geometric conditioning. A two-stage teacher-guided learning strategy first anchors the geometry-conditioned representation to the pretrained appearance space, before refining it for place discrimination. Evaluated on the WildCross benchmark, DAD improves average inter-sequence Recall@1 from 61.41% to 66.37% and Recall@5 from 65.86% to 72.49% over a matched appearance-only baseline, with the largest gains under reverse traversal and long-term appearance variation. These results show that GFM-derived geometry can provide a persistent structural prior for VPR when visual appearance becomes unreliable.
Chinese Translation
自然环境中的视觉位置识别(Visual Place Recognition, VPR)仍然具有挑战性,其原因在于植被的重复性、稀疏的显著地标,以及不同穿越过程之间显著的外观和视角变化。尽管同一位置的视觉观测可能发生较大变化,但其潜在的空间结构往往更为持久。我们通过深度感知蒸馏(Depth-Aware Distillation, DAD)利用这种互补的几何一致性,将预训练视觉基础模型(Vision Foundation Model, VFM)的令牌表征以几何基础模型(Geometric Foundation Model, GFM)推断的几何信息为条件,而无需任何深度传感器。DAD并非将几何视为额外的输入模态,而是将图像对齐的深度投影到VFM令牌空间中,并通过通道级的几何条件化选择性地调制视觉表征。一种两阶段教师引导的学习策略首先将几何条件化的表征锚定到预训练的外观空间,然后再针对位置判别进行精炼。在WildCross基准上的评估表明,与同等条件下仅使用外观的基线相比,DAD将平均跨序列Recall@1从61.41%提升至66.37%,Recall@5从65.86%提升至72.49%,其中在反向穿越和长期外观变化条件下增益最大。这些结果表明,当视觉外观变得不可靠时,由GFM推导的几何信息可以为VPR提供一个持久的结构先验。
cs.CV / 30 / 2609.27408

What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit

视觉-语言模型看似的能力极限实为读取输出的极限
Del Valle, Alfredo F. Frontera
Abstract
Benchmarks for vision-language models offer their answer choices in some convention: a letter, a color name, a pixel coordinate. That convention is treated as neutral. We find it is not, and that the limits a benchmark reports can belong to the readout rather than to the model. On 200 COCO photographs, Qwen3-VL-4B picks the correct one of nine locations for a named object 68.5% of the time when the locations are given in English and 20.0% when the same locations are given as pixel coordinates. Chance is 11.1%. The cost arises when the answer options are coordinates; giving the model a coordinate in the question instead costs 3.5 points and is not significant. The gap holds on a 4x4 grid, under 8-bit rather than 4-bit quantization, and in every slice by object size, boundary distance and category. It also decides which model wins. Two models that tie under English names differ by 39 points in one coordinate system and by 54 in the other, in opposite directions. On the color task, three of the four open models capable of the task show the penalty; on photographs, two of three open models do, and so does Gemini, at 11.1 points on parseable answers (p = 1e-4). GPT-4o does not. To ask whether a model reads a coordinate at all, we attach the wrong name to each one and record which the model follows. Color options written as hue angles are followed below chance; a normalized pixel convention is followed at four times chance. This tells apart conventions a model can use from ones it cannot, though it did not predict accuracy on two untried conventions. Five models also name the same color wheel five different ways, so a fixed answer vocabulary is not neutral across models either. Five times during this work we measured a capable model as incapable because our scorer and the model disagreed about what an answer looks like. We report each case. They are the phenomenon in miniature.
Chinese Translation
视觉-语言模型的基准测试通常以某种约定格式提供答案选项:字母、颜色名称或像素坐标。这种约定被视为中立的。我们发现事实并非如此,基准测试所报告的模型极限可能属于输出读取(readout)环节而非模型本身。在200张COCO照片上,Qwen3-VL-4B在以英文给出九个位置时,能以68.5%的准确率选出指定物体的正确位置;而当同样的位置以像素坐标给出时,准确率仅为20.0%(随机水平为11.1%)。这一代价仅在答案选项为坐标时出现;在问题中额外给出一个坐标仅造成3.5个百分点的损失,且不具显著性。该差距在4x4网格、8比特量化(而非4比特)下,以及按物体大小、边界距离和类别划分的所有分片中均持续存在。它还会决定哪个模型胜出:在英文命名条件下打平的两个模型,在一个坐标系中相差39个百分点,在另一个坐标系中相差54个百分点,且方向相反。在颜色任务上,四个具备该能力的开源模型中有三个表现出此代价;在照片上,三个开源模型中有两个表现出此代价,Gemini也是如此,在可解析答案上相差11.1个百分点(p = 1e-4),而GPT-4o则没有。为了探究模型是否能读取坐标,我们在每个坐标上附加错误的名称并记录模型跟随哪一个。以色相角度书写的颜色选项被跟随的概率低于随机水平;归一化像素约定被跟随的概率是随机水平的四倍。这可以区分模型能够使用与无法使用的约定,尽管它并未能预测两种未尝试约定的准确率。此外,五个模型对同一色轮有五种不同的命名方式,因此固定的答案词汇表在不同模型之间也并非中立。在本研究过程中,我们曾五次因为评分器与模型对答案格式理解的分歧而将一个有能力的模型误判为无能力。我们逐一报告了这些案例,它们正是这一现象的缩影。
cs.CV / 31 / 2609.27413

S2A:Semantic-to-Spatial Alignment for Alignment-Free RGB-T Salient Object Detection

S2A:面向免配准RGB-T显著目标检测的语义到空间对齐框架
Zhou, Qiangqiang, Luo, Yang, Chen, Yong, Xu, Jiawei
Abstract
Alignment-free RGB-T salient object detection (RGB-T SOD) aims to identify salient objects from unregistered RGB and thermal image pairs without costly pre-alignment. However, spatial misalignment breaks pixel-wise correspondence and causes feature contamination during cross-modal fusion. To address this issue, we propose S2A, a semantic-to-spatial alignment framework for alignment-free RGB-T SOD. Specifically, a global-guided hierarchical fusion module (GGHF) first exploits global semantic guidance to suppress background interference and refine hierarchical intra-modal features. Subsequently, the alignment-free cross-modal channel attention module (AFCA) globally exchanges complementary semantic information through channel-wise interaction, effectively overcoming the interference caused by local spatial misalignments. Finally, a spatial deformable cross-attention module (SDCA) predicts adaptive sampling offsets to recover local cross-modal spatial correspondence. Through this semantic-to-spatial paradigm, S2A first enables reliable cross-modal semantic interaction and subsequently performs local spatial calibration, effectively reducing misalignment-induced feature contamination. Without bells and whistles, S2A achieves highly competitive performance on multiple public alignment-free RGB-T benchmarks, demonstrating its effectiveness in alleviating misalignment-induced feature contamination.
Chinese Translation
免配准RGB-T显著目标检测(RGB-T SOD)旨在从未经配准的RGB图像与热红外图像对中识别显著目标,而无需代价高昂的预先配准。然而,空间错位破坏了像素级对应关系,并在跨模态融合过程中导致特征污染。为解决这一问题,我们提出了S2A,一种面向免配准RGB-T SOD的语义到空间对齐框架。具体而言,全局引导的分层融合模块(GGHF)首先利用全局语义引导来抑制背景干扰并精炼分层模态内特征。随后,免配准跨模态通道注意力模块(AFCA)通过通道级交互在全局范围内交换互补语义信息,有效克服了局部空间错位所引起的干扰。最后,空间可变形交叉注意力模块(SDCA)通过预测自适应采样偏移来恢复局部跨模态空间对应关系。通过这一语义到空间的范式,S2A首先实现可靠的跨模态语义交互,随后进行局部空间校准,有效减少了由错位引起的特征污染。在不借助任何额外技巧的情况下,S2A在多个公开的免配准RGB-T基准数据集上取得了极具竞争力的性能,验证了其在缓解错位引起的特征污染方面的有效性。
cs.CV / 32 / 2609.27423

Overlapping Visual Grouping Without Semantic Priors

无语义先验的重叠视觉分组
Saukkio, Teemu, Haghbayan, Hashem, Plosila, Juha
Abstract
Most computer-vision systems organize visual input toward a predefined interpretation, such as semantic categories, prompted regions, learned object-like representations, or a single spatial partition. This work considers an earlier stage of visual organization: the formation of candidate perceptual units directly from sensor measurements before their identity, meaning, or task relevance is known. We introduce Domain Parent Grouping (DPG), a sensor-grounded grouping method in which complementary measurement relationships are represented in separate processing domains. Spatially connected groups formed within these domains are related through cross-domain overlap, yielding a non-exclusive grouping representation rather than a single mutually exclusive segmentation. This representation retains broader and more localized groups, as well as alternative grouping boundaries over the same image locations, simultaneously available. DPG also includes a native mechanism for reprocessing selected group content, in which input-relative measurement ranges allow the observational resolution to change while preserving previously formed groups. DPG is implemented using three domains representing locally contextualized luminance, direct chromatic relationships, and contextual chromatic relationships. Experiments on the BSDS500 dataset demonstrate the benefit of combining the three domains. The results further show that DPG forms measurement-supported groups corresponding to low-level image structure, and that these groups exhibit measurable correspondence with human-annotated regions and boundaries. This demonstrates that structured visual organization can emerge directly from relationships among sensor measurements.
Chinese Translation
大多数计算机视觉系统将视觉输入组织为预定义的解释,例如语义类别、提示区域、习得的类物体表征或单一空间划分。本研究关注视觉组织的更早期阶段:在候选感知单元的身份、含义或任务相关性尚不清楚之前,直接从传感器测量数据中形成这些感知单元。我们提出了领域父分组(Domain Parent Grouping, DPG),这是一种基于传感器的分组方法,将互补的测量关系表示在各自独立的处理域中。在这些域内形成的空间连通组通过跨域重叠相互关联,从而产生一种非排他性的分组表示,而非单一的互斥分割。这种表示能够同时保留更宽泛和更局部的组,以及同一图像位置上的多种备选分组边界。DPG 还包含一种对选定组内容进行再处理的原生机制,其中相对于输入的测量范围使观测分辨率可以改变,同时保持先前形成的组不变。DPG 使用三个域实现,分别表示局部情境化的亮度、直接的色彩关系和情境化的色彩关系。在 BSDS500 数据集上的实验证明了结合这三个域的益处。结果进一步表明,DPG 形成的由测量支持的组与低层图像结构相对应,并且这些组与人工标注的区域和边界表现出可度量的对应关系。这证明了结构化的视觉组织可以直接从传感器测量之间的关系中涌现。
cs.CV / 33 / 2609.27442

SatUnreal: A High-Precision Synthetic Dataset for Satellite Stereo Matching via Unreal Engine

SatUnreal:基于虚幻引擎的高精度卫星立体匹配合成数据集
Kim, Han-Gyeol, Park, JaeWan, Park, Junmin, Kwon, Darongsae
Abstract
3D reconstruction from satellite imagery is essential for large-scale topographic analysis, yet the lack of high-fidelity training datasets with accurate occlusion labels remains a primary bottleneck. Existing benchmarks, such as US3D and WHU-Stereo, face inherent challenges in spatio-temporal mismatch -- environmental changes and shadow displacements between multi-view acquisitions -- and provide ambiguous ground truth in occluded regions due to LiDAR sparsity. In this paper, we propose SatUnreal, a high-precision synthetic dataset designed to fundamentally overcome these limitations through an Unreal Engine-based simulation pipeline. SatUnreal provides 10,000 stereo pairs with high resolution (0.3m GSD) and is characterized by: (1) Physical Geometry Simulation, replicating realistic satellite orbits by systematically varying baselines and azimuths; (2) Spatio-temporal Consistency, eliminating temporal noise through fixed virtual environments; (3) Topographic Diversity, spanning dense urban canyons to low-texture natural terrains; and (4) Mathematical Label Integrity, utilizing a novel two-step linetrace algorithm to generate flawless occlusion masks. Experimental results using SOTA iterative models demonstrate that models trained exclusively on SatUnreal achieve superior zero-shot transfer performance on real-world benchmarks (US3D, WHU-Stereo) compared to those trained on real datasets. Our findings prove that physically accurate synthetic data provides a more effective supervisory signal for learning geometric features than complex real-world observations, establishing a new paradigm for Sim-to-Real transfer in Earth Observation. Code and dataset are available at https://github.com/jmp-Telepix/SatUnreal_A_High-Precision_Synthetic_Dataset_for_Satellite_Stereo_Matching_via_UnrealEngine
Chinese Translation
基于卫星影像的三维重建对于大范围地形分析至关重要,然而缺乏具有精确遮挡标注的高保真训练数据集仍是主要瓶颈。现有基准数据集(如US3D和WHU-Stereo)面临时空错配的固有问题——即多视角获取之间的环境变化和阴影位移——且由于激光雷达点云稀疏,在遮挡区域提供的真值存在歧义。本文提出SatUnreal,一个基于虚幻引擎仿真管线构建的高精度合成数据集,旨在从根本上克服上述局限。SatUnreal提供了10,000对高分辨率(0.3m地面采样距离)立体影像对,其特点包括:(1) 物理几何仿真,通过系统地变化基线和方位角来复现真实的卫星轨道;(2) 时空一致性,通过固定的虚拟环境消除时间噪声;(3) 地形多样性,涵盖密集的城市峡谷到低纹理的自然地形;(4) 数学标注完整性,采用新颖的两步光线追踪算法生成无瑕疵的遮挡掩膜。基于SOTA迭代模型的实验结果表明,仅在SatUnreal上训练的模型在真实世界基准数据集(US3D、WHU-Stereo)上的零样本迁移性能优于在真实数据集上训练的模型。我们的研究证明,物理精确的合成数据能够为学习几何特征提供比复杂真实世界观测更有效的监督信号,为地球观测领域的Sim-to-Real迁移确立了新范式。代码和数据集可在 https://github.com/jmp-Telepix/SatUnreal_A_High-Precision_Synthetic_Dataset_for_Satellite_Stereo_Matching_via_UnrealEngine 获取。
cs.CV / 34 / 2609.27455

Latent evolving World Action Model

潜在演化的世界动作模型(Latent Evolving World Action Model)
Fang, Xueji, Duan, Boqiang, Wu, Hua, Wang, Jingdong, Qi, Guo-Jun
Abstract
World Action Models (WAMs) jointly model action generation and environment dynamics and are mostly built on pretrained Video Diffusion Models (VDMs). In VDM-based WAMs, observations are first encoded by a VAE, and the resulting compressed latents are then processed by large video diffusion backbones to extract effective features for action generation. However, this paradigm ties WAM performance and training cost to large-scale video generation pretraining, limiting WAM efficiency and scalability. In this paper, we theoretically and empirically investigate how visual representations affect action generation in WAMs. Our results show that predictive embeddings from Joint-Embedding Predictive Architecture (JEPA) encoders better support action generation than compressed VAE latents, with I-JEPA performing best in our encoder comparison. Based on these findings, we propose LeWAM, which conditions action generation on JEPA embeddings and models environment evolution by predicting future embeddings in the same space, without relying on a video diffusion backbone. We further find that imitation learning matches demonstrated actions but does not distinguish better actions from worse ones, even though small action deviations can greatly affect task success. To address this limitation without additional environment interaction or the human oversight required for resets and safety, we introduce Demonstration-Guided DPO (DemoDPO), an offline preference refinement stage that derives preference supervision directly from demonstrations.With only 0.4B trainable parameters, LeWAM achieves an average success rate of 92.28\% on RoboTwin 2.0, comparable to that of state-of-the-art VLAs and WAMs, and maintains practical effectiveness on real-world manipulation tasks.
Chinese Translation
世界动作模型(World Action Models, WAMs)联合建模动作生成与环境动态,且大多构建于预训练视频扩散模型(Video Diffusion Models, VDMs)之上。在基于VDM的WAM中,观测首先由VAE编码,得到的压缩潜变量再由大型视频扩散骨干网络处理,以提取用于动作生成的有效特征。然而,这种范式将WAM的性能和训练成本与大规模视频生成预训练绑定在一起,限制了WAM的效率和可扩展性。本文从理论和实证两方面研究了视觉表示如何影响WAM中的动作生成。结果表明,来自联合嵌入预测架构(Joint-Embedding Predictive Architecture, JEPA)编码器的预测性嵌入比压缩的VAE潜变量更能支持动作生成,其中I-JEPA在我们的编码器对比中表现最佳。基于这些发现,我们提出LeWAM,它以JEPA嵌入为条件进行动作生成,并通过在同一空间中预测未来嵌入来建模环境演化,而无需依赖视频扩散骨干网络。我们进一步发现,模仿学习虽然能够匹配示范动作,但无法区分更优动作与更差动作,即使微小的动作偏差也可能极大地影响任务成功。为了在不引入额外环境交互以及重置和安全所需人工监督的情况下解决这一局限,我们提出了示范引导DPO(Demonstration-Guided DPO, DemoDPO),这是一种离线偏好优化阶段,直接从示范中获取偏好监督信号。LeWAM仅有0.4B可训练参数,在RoboTwin 2.0上实现了92.28%的平均成功率,与最先进的VLA和WAM相当,并在真实世界操作任务中保持了实际有效性。
cs.CV / 35 / 2609.27457

Beyond Balanced Accuracy: A Resolution and Parity-Controlled Benchmark for Vision-Language and Vision-Only Defect Assessment in UAV Power-Line Inspection

超越平衡准确率:面向无人机电力线路巡检中视觉-语言与纯视觉缺陷评估的分辨率与一致性受控基准
Zhang, Linghao, Xiang, Siyu, Kuang, Junwei, Yi, Peiyu
Abstract
Vision-language models (VLMs) are often reported to outperform task-specific vision backbones for unmanned aerial vehicle (UAV) power-line defect assessment. We test that claim on ElecVQA-Bench, a 56,972-item benchmark derived from the public InsPLAD dataset, across six evaluation choices: partition, evaluated item set, label space, replication, input resolution, and side information. On a matched partition, a Swin Transformer and the strongest adapted VLM differ by only 0.03 points at binary screening. At seven-way defect typing, increasing the vision backbones from 224 px to the measured pixel budget of the VLM preprocessor narrows the gap against InternVL3.5-8B from +20.53 to -0.57 points for ResNet-50 and from +23.67 to +4.70 points for Swin-T. A pixel-budget audit shifts Qwen3-VL-8B macro recall by 10.78 points, yet a source-pixel-matched InternVL control still leaves Qwen ahead by 7.43 to 13.61 points while using 56% fewer visual tokens, so neither source pixels nor token budget explains the difference between the two VLMs. A two-seed global replication changes Qwen binary accuracy and seven-way macro recall by 0.86 and 1.02 points. After split-specific retraining, Qwen does not lead at crop or image level, and a 14-tower, three-seed replication reverses the sign across seeds, giving mean common-six macro recall of 0.9085 for Qwen against 0.9509 for ResNet-50. No split regime yields a family-level advantage that survives multiple-comparison correction. The study supports a benchmark-audit contribution rather than a general claim of VLM superiority.
Chinese Translation
视觉-语言模型(VLM)在无人机(UAV)电力线路缺陷评估任务中常被报道优于特定任务的视觉骨干网络。我们在 ElecVQA-Bench 上检验了这一论断,该基准包含 56,972 个样本,源自公开的 InsPLAD 数据集,并考察六种评估选择:数据划分、评估样本集、标签空间、重复实验、输入分辨率以及辅助信息。在匹配的数据划分下,Swin Transformer 与最强的适配 VLM 在二分类筛查上仅相差 0.03 分。在七类缺陷类型判定任务中,将视觉骨干网络的输入从 224 像素提升至 VLM 预处理器实际使用的像素预算后,ResNet-50 与 InternVL3.5-8B 的差距从 +20.53 分缩小至 -0.57 分,Swin-T 则从 +23.67 分缩小至 +4.70 分。像素预算审计使 Qwen3-VL-8B 的宏平均召回率变化 10.78 分;然而,在源像素匹配的 InternVL 对照实验中,Qwen 仍领先 7.43 至 13.61 分,且所用视觉 token 少 56%,因此源像素和 token 预算均无法解释两个 VLM 之间的性能差异。两次随机种子的全局重复实验使 Qwen 的二分类准确率和七类宏平均召回率分别变化 0.86 分和 1.02 分。在针对各划分进行重训练后,Qwen 在裁剪级别和图像级别均不再领先;而在 14 基塔、三种子的重复实验中,结果符号在不同种子间发生反转,Qwen 的平均公共六类宏平均召回率为 0.9085,低于 ResNet-50 的 0.9509。没有任何划分方案能够产生在多重比较校正后仍然成立的模型家族层面的优势。本研究支持基准审计的贡献,而非 VLM 优越性的一般性结论。
cs.CV / 36 / 2609.27461

Invisible in Space, Visible in Time: Motion Vision CAPTCHA against GUI Agents

空间上不可见,时间上可见:抵御GUI代理的运动视觉验证码
Zhang, Zeyu, Rong, Dingyi, Chen, Zijian, Zhang, Zicheng, Min, Xiongkuo, Zhai, Guangtao
Abstract
Most existing visual CAPTCHAs remain spatially solvable: the required information is exposed by static appearance, local structure, and interface state. This assumption is weakened by advances in multimodal large language models (MLLMs) and Graphical User Interface (GUI) agents, which exhibit strong visual perception, reasoning, and browser interaction capabilities. We propose Motion Vision CAPTCHA (MVCAP), a hierarchical motion-based CAPTCHA framework in which target semantics are instantiated as motion-defined foreground structures and become recoverable only through temporal segregation from a dynamically evolving background. Built on this shared principle, MVCAP is instantiated in three perceptually progressive levels: coherent motion, structural motion, and biological motion. To evaluate this framework, we introduce MVCAP-Bench, a browser-based benchmark with 600 live CAPTCHA instances, together with a matched foreground-only control benchmark, MVCAP-Bench-FG. We evaluate humans, Browser Use agents, native computer use agents, and a supplementary offline VQA setting derived from the same instances. Results reveal a substantial human--agent gap: on the full MVCAP-Bench, human accuracy reaches 99.6%, whereas the best GUI agent achieves only 16.8%, close to the six-way chance level. The foreground-only control further shows that the key difficulty comes from dynamic background camouflage rather than answer format or browser interaction alone. These findings identify a measurable human--agent perception gap and position MVCAP-Bench as a benchmark for studying motion-defined perception in current agents.
Chinese Translation
现有的大多数视觉验证码(CAPTCHA)在空间层面仍然可以被破解:所需信息可通过静态外观、局部结构和界面状态直接获取。多模态大语言模型(MLLM)和图形用户界面(GUI)代理的快速发展削弱了这一假设,因为它们具备强大的视觉感知、推理和浏览器交互能力。我们提出运动视觉验证码(Motion Vision CAPTCHA, MVCAP),这是一种基于运动的分层验证码框架,其目标语义被实例化为由运动定义的前景结构,只有通过与动态演化背景进行时间上的分离才能被恢复。基于这一共同原理,MVCAP被实例化为感知上逐级递进的三个层次:连贯运动、结构化运动和生物运动。为评估该框架,我们引入了MVCAP-Bench,一个基于浏览器的基准,包含600个实时验证码实例,并配套一个仅含前景的对照基准MVCAP-Bench-FG。我们评估了人类、Browser Use代理、原生计算机操作代理,以及基于相同实例的补充性离线视觉问答(VQA)设置。结果显示出显著的人机差距:在完整的MVCAP-Bench上,人类准确率达到99.6%,而最佳的GUI代理仅达到16.8%,接近六选一的随机水平。仅前景的对照实验进一步表明,核心难点来自动态背景伪装,而非答案格式或浏览器交互本身。这些发现揭示了可度量的人机感知差距,并将MVCAP-Bench定位为研究当前代理运动定义感知能力的基准。
cs.CV / 37 / 2609.27462

Hybrid Gaussians for Robust Open-Vocabulary 3D Segmentation with Multi-View Object Association and Boundary Refinement

基于混合高斯(Hybrid Gaussians)的鲁棒开放词汇3D分割:多视角目标关联与边界精化
Qiu, Xueqi, Sun, Yueming, Zhang, Tianyu, Xia, Yuxuan, Long, Yang
Abstract
Open-vocabulary 3D segmentation localizes objects from free-form text queries, but remains challenging in real image sequences: incomplete or noisy 2D supervision destabilizes multi-view identity assignment, while full-scene semantic learning weakens object-level discriminability. We introduce Hybrid Gaussians, a unified 3D representation jointly modeling object association and language-aligned semantics. Its Multi-View Object Association mechanism combines Observation Fusion and Semantic Contrastive Learning to improve identity consistency and semantic discrimination. Boundary Reconstruction Optimization further refines local boundary structure to improve contour quality. Experiments on LERF and 3D-OVS demonstrate strong quantitative and qualitative performance. Our method achieves 59.1\% mIoU on LERF, yielding a 13.4\% relative gain over the baseline. Project page: https://nora202.github.io/hybridgaussians.
Chinese Translation
开放词汇3D分割旨在根据自由形式的文本查询对目标进行定位,但在真实图像序列中仍然具有挑战性:不完整或有噪声的2D监督会使多视角身份分配不稳定,而全场景语义学习会削弱目标级别的判别能力。我们提出了混合高斯(Hybrid Gaussians),这是一种统一的高斯表示方法,可联合建模目标关联与语言对齐的语义。其多视角目标关联机制结合了观测融合(Observation Fusion)与语义对比学习(Semantic Contrastive Learning),以提升身份一致性和语义判别能力。边界重建优化(Boundary Reconstruction Optimization)进一步精化局部边界结构,以改善轮廓质量。在LERF和3D-OVS数据集上的实验表明,我们的方法在定量和定性方面均具有强大性能。我们的方法在LERF上达到59.1%的mIoU,相比基线实现了13.4%的相对提升。项目主页:https://nora202.github.io/hybridgaussians。
cs.CV / 38 / 2609.27468

CereVLA: Cerebellum-Inspired Consequence-Aware Residual Governance for Efficient Vision-Language-Action Execution

CereVLA:受小脑启发的后果感知残差治理框架,用于高效的视觉-语言-动作执行
Zeng, Shuai, Liang, Yuxuan, Hu, Hangmiao, Zhou, Fobao, Wang, Zixiang, Hong, Wenxi, Zhao, Hang
Abstract
Action-chunked vision-language-action (VLA) policies improve inference efficiency, but limited feedback within committed action chunks can lead to accumulated execution errors. Residual adaptation can correct such deviations without retraining the VLA; however, existing corrections are typically optimized for reference-action consistency without explicitly considering their downstream consequences. To address this limitation, we present Cerebellum-Inspired Consequence-Aware Residual Governance (CereVLA), a unified framework that integrates lightweight residual refinement and predictive consequence evaluation into frozen VLA execution. Corrective actions are first generated by flow-based residual refinement, and their short- and interval-horizon consequences are then evaluated by a recurrent state-space model and a history-aware classifier. Residual corrections predicted to be unfavorable are selectively suppressed by a lightweight governor. Comparisons with state-of-the-art methods on LIBERO-10 and LIBERO-GOAL demonstrate the effectiveness of CereVLA. On SO-101, CereVLA increases task success from 57.5% to 90.0% and reduces mean control steps by 19.6% among successful trials, relative to the frozen SmolVLA baseline.
Chinese Translation
基于动作分块的视觉-语言-动作(Vision-Language-Action, VLA)策略提升了推理效率,但在已提交的动作块内反馈有限,可能导致执行误差的累积。残差适配(residual adaptation)可以在无需重新训练 VLA 的情况下纠正此类偏差;然而,现有的纠正方法通常仅针对参考动作的一致性进行优化,而未显式考虑其下游后果。为解决这一局限,我们提出了受小脑启发的后果感知残差治理框架(Cerebellum-Inspired Consequence-Aware Residual Governance,CereVLA),这是一个将轻量级残差精修与预测性后果评估集成到冻结 VLA 执行中的统一框架。首先,通过基于流(flow-based)的残差精修生成纠正动作;随后,由递归状态空间模型和具备历史感知能力的分类器评估其短期与区间视界内的后果。预测为不利的残差纠正会被一个轻量级治理器(governor)选择性地抑制。在 LIBERO-10 和 LIBERO-GOAL 上与最先进方法的对比验证了 CereVLA 的有效性。在 SO-101 上,相对于冻结的 SmolVLA 基线,CereVLA 将任务成功率从 57.5% 提升至 90.0%,并在成功的试验中将平均控制步数减少了 19.6%。
cs.CV / 39 / 2609.27470

DeltaS: Reading the Gated Linear Attention State for KV Cache Eviction in Streaming Video

DeltaS:利用门控线性注意力状态进行流式视频的KV缓存淘汰
Kwon, Taeyoun, Kim, Seungjin, Kim, Hyeonyu, Kim, Moon Hwan
Abstract
Recent video-language models increasingly adopt hybrid architectures that interleave linear and full attention layers for efficient long-context processing. While the recurrent state of linear attention remains fixed in size, the KV cache of full attention continues to grow with the video stream, making eviction necessary under a bounded memory budget. The key challenge in streaming is that eviction must occur before the question arrives, so what to retain has to be decided without the question. Existing eviction methods derive token scores from the KV cache itself, using position, attention, or key-value representations, and attention-based scores further require proxy queries or extra computation. Hybrid backbones offer another source of signal. In gated-delta linear attention, the recurrent state is updated by the residual between each input and what can already be retrieved from the state, so its change over a chunk of frames reflects how much new information the chunk brings. We propose DeltaS, a query-agnostic, training-free method that retains video chunks inducing larger normalized state change, or state drift. In a controlled comparison with the budget and retention policy held fixed, state drift outperforms position-, attention-, and key-value-based signals. With a signal costing only 1.9% of the forward pass, DeltaS surpasses the strongest query-agnostic bounded-memory baseline by 2.1 points on average across six long-video benchmarks and by 5.6 points on the longest benchmark. These results suggest that the two memories of hybrid architectures can work cooperatively. Code is available at https://github.com/MaumAI-Company/DeltaS.
Chinese Translation
近期的视频-语言模型日益采用混合架构,通过交替使用线性注意力层与全注意力层来实现高效的长上下文处理。虽然线性注意力的循环状态大小固定,但全注意力的KV缓存会随视频流持续增长,因此在有限的内存预算下必须进行缓存淘汰。流式场景的关键挑战在于:淘汰必须在问题到来之前完成,即在不依赖问题的情况下决定保留哪些内容。现有的淘汰方法从KV缓存自身(如位置、注意力或键值表示)推导token分数,而基于注意力的分数还需要代理查询或额外计算。混合骨干网络提供了另一种信号来源。在门控delta线性注意力(gated-delta linear attention)中,循环状态由每个输入与当前状态可检索内容之间的残差进行更新,因此其在一段时间帧内的变化反映了该段帧所带来的新信息量。我们提出DeltaS,一种与查询无关(query-agnostic)、无需训练的方法,通过保留归一化状态变化(即状态漂移)较大的视频片段进行筛选。在预算与保留策略固定的受控对比中,状态漂移优于基于位置、注意力和键值的信号。该信号的额外开销仅为前向传播的1.9%,DeltaS在六个长视频基准上平均超越最强query-agnostic有限内存基线2.1分,在最长的基准上超越5.6分。这些结果表明混合架构的两种记忆机制可以协同工作。代码已发布于 https://github.com/MaumAI-Company/DeltaS。
cs.CV / 40 / 2609.27493

Information Capacity of Generative Video Compression: Quantifying the Rate-Compute Exchange at Identical Quality

生成式视频压缩的信息容量:同等质量下速率-计算交换的量化
Yuan, Cheng, Shao, Jiawei, Li, Xuelong
Abstract
Under the AI Flow framework, communication networks distribute intelligence across devices, edge servers, and clouds, and computation at the receiver becomes a resource that can substitute for transmitted bits. Generative video compression (GVC) embodies this exchange by sending compact tokens with ultra-low bitrate and letting a generative decoder synthesize the video, yet how much bandwidth savings a unit of decoder compute actually achieves has never been quantified. To fill this vacancy, we model reconstruction quality as a two-factor power law in data rate and decoder compute, which fits measured DISTS of two GVC decoders with a mean error below 3%, and define the information capacity (IC) as the negative logarithmic slope along an iso-quality contour, namely the fraction of rate saved per fractional increase in compute at identical quality. IC is dimensionless and unit-invariant, thus enabling an architecture-agnostic comparison. It forms a field over the operating plane, locating where additional denoising steps are worth their cost. Across five datasets, the 14B decoder trades more compute for fewer rate about ten times more efficiently than the 1.3B decoder. IC also varies significantly across datasets, indicating imbalanced performance on the rate-compute trade-off in GVC methods.
Chinese Translation
在AI Flow(人工智能流)框架下,通信网络将智能分布于设备、边缘服务器和云端之间,接收端的计算成为可以替代传输比特数的一种资源。生成式视频压缩(Generative Video Compression, GVC)体现了这种交换:以超低比特率发送紧凑的令牌(token),并由生成式解码器合成视频,然而每一单位解码器计算究竟能节省多少带宽,此前从未被量化。为填补这一空白,我们将重建质量建模为关于数据速率和解码器计算的双因子幂律,该模型对两个GVC解码器实测的DISTS指标拟合的平均误差低于3%;并将信息容量(Information Capacity, IC)定义为沿等质量曲线的负对数斜率,即在相同质量下,计算量每增加一个相对比例所能节省的速率比例。IC是无量纲且单位不变的,因此能够实现与架构无关的比较。IC在工作平面上构成一个场,可定位额外的去噪步骤在何处值得付出其代价。在五个数据集上,14B解码器以更多计算换取更少速率的效率约为1.3B解码器的十倍。IC在不同数据集间也存在显著差异,表明GVC方法在速率-计算权衡上的性能并不均衡。
cs.CV / 41 / 2609.27509

Know-Your-Scene (KYS)-SLAM: Hierarchical Semantic-Motion Priors for Feature Matching in Stereo Visual SLAM

Know-Your-Scene(KYS)-SLAM:面向双目视觉SLAM特征匹配的层次化语义-运动先验
Chatterjee, Preeti, Lu, Jin, Sun, Jin, Bhandarkar, Suchendra M.
Abstract
Stereo visual SLAM systems built on local descriptors suffer from semantic ambiguity, instance-level confusion, and independently moving objects, each corrupting data association and accumulating as trajectory drift. Prevailing semantic and dynamic SLAM methods address this through binary feature rejection, sacrificing correspondence density for outlier suppression. We contend that contextual implausibility is better expressed as a graded quantity than an exclusion criterion. We present Know-Your-Scene (KYS)-SLAM, a modular extension of ORB-SLAM3 that supplants feature rejection with continuous correspondence modulation. The contribution is the reframing of contextual evidence as correspondence cost, applied within feature matching and leaving the geometric backend unmodified. Each keypoint is augmented with semantic, panoptic, and motion priors fused through a hierarchical compatibility formulation, in which semantic class and instance identity enforce structural plausibility while a zero-shot motion score down-weights features on independently moving objects. That score comes from a training-free module fitting a depth-aware ego-motion model to background optical flow and classifying panoptic segments via self-calibrating, coverage-aware thresholds, so only segments with sufficient motion evidence are penalized and static structure is left unpenalized. Penalizing correspondences rather than discarding them preserves the geometric support bundle adjustment depends on. Under one fixed configuration, no coefficient retuned per sequence or dataset, KYS-SLAM reduces per-sequence ATE RMSE by 17.4% on outdoor KITTI and 27.7% on indoor EuRoC across 21 stereo sequences with no regressions, and by 6.6% on dynamic subsets of KITTI Tracking and 17.8%, up to 31.2%, on Virtual KITTI 2 -- cross-domain transfer across outdoor driving, indoor flight, and synthetic imagery under one set of constants.
Chinese Translation
基于局部描述子的双目视觉SLAM系统存在语义歧义、实例级混淆以及独立运动物体等问题,这些问题会破坏数据关联并逐渐累积为轨迹漂移。主流的语义SLAM和动态SLAM方法通过二值化的特征剔除来解决这些问题,即以牺牲匹配密度为代价来抑制外点。我们认为,上下文不合理性更适合表示为一个连续的量化指标,而非排除准则。我们提出Know-Your-Scene(KYS)-SLAM,这是ORB-SLAM3的一个模块化扩展,用连续的匹配调制取代特征剔除。其核心贡献在于将上下文证据重新构建为匹配代价,并在特征匹配中加以应用,同时保持几何后端不变。每个关键点都被增广以语义、全景和运动先验,并通过层次化相容性公式进行融合:语义类别和实例身份用于约束结构合理性,而零样本运动评分则对独立运动物体上的特征进行降权。该评分来自一个免训练模块,该模块将深度感知的自运动模型拟合到背景光流上,并通过自校准、覆盖度感知的阈值对全景分割片段进行分类,因此仅对具有充分运动证据的片段施加惩罚,而静态结构不受影响。对匹配施加惩罚而非直接丢弃,保留了捆绑调整所依赖的几何支持。在单一固定配置下——不对每个序列或数据集重新调整任何系数——KYS-SLAM在21个双目序列上将逐序列ATE RMSE在室外KITTI上降低17.4%,在室内EuRoC上降低27.7%,且无一性能回退;在KITTI Tracking动态子集上降低6.6%,在Virtual KITTI 2上降低17.8%(最高达31.2%)。这实现了在室外驾驶、室内飞行和合成图像之间的跨域迁移,且仅使用同一组常数。
cs.CV / 42 / 2609.27511

NV-Reason-CT: 3D Visual Language Model for CT Analysis

NV-Reason-CT:用于CT分析的三维视觉语言模型
Myronenko, Andriy, Yang, Dong, Tang, Yucheng, Turkbey, Baris, Simon, Benjamin, Harmon, Stephanie, Makwana, Rikhil, Aboian, Mariam, Azamat, Sena, Hamamci, Ibrahim Ethem, Er, Sezgin, Menze, Bjoern, Edgar, Marc, He, Yufan, Guo, Pengfei, Xu, Daguang
Abstract
We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging. This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing with text. We train on a curated corpus of approximately 550,000 multimodal instruction examples from 70,111 unique CT image inputs, combining standardized reports, abnormality-focused and anatomy-specific questions, multi-turn interactions, and radiologist-authored reasoning from recorded and transcribed expert CT interpretations. Expert annotations provide direct supervision and guide additional report-grounded synthetic reasoning. End-to-end supervised fine-tuning (SFT) is followed by Group Relative Policy Optimization (GRPO), with verifiable rewards over chest and abdominal abnormality sets. The model supports abnormality classification, report generation, and interactive reasoning with reviewable observations, differential diagnoses, and uncertainty. Evaluation spans public CT benchmarks and a held-out NIH cohort. On CT-RATE, NV-Reason-CT achieves a macro-F1 of 0.614 and macro-AUROC of 0.871 without a task-specific classification head; generated reports achieve a report-derived macro-F1 of 0.592. In a preliminary study with expert radiologists, AI-assisted review received favorable confidence ratings and was associated with a 50% reduction in average reported interpretation and reporting time. We release the model and training code to support reproducible research on explainable AI for volumetric medical imaging.
Chinese Translation
我们提出了NV-Reason-CT,一个用于胸部和腹部CT的生成式视觉-语言模型,它将原生三维视觉编码与放射科医生引导的推理相结合。该模型将原生三维视觉Transformer(Vision Transformer)与语言模型耦合,将所有视觉标记(token)及其显式三维坐标直接传入语言解码过程,而无需进一步的空间标记合并。这使得体积空间信息得以在视觉编码器内部保留,并在与文本联合处理时通过语言模型的位置编码保持下来。我们在一个精心整理的语料库上进行训练,该语料库包含来自70,111个独立CT图像输入的约550,000个多模态指令样本,涵盖标准化报告、聚焦异常和特定解剖结构的问题、多轮交互,以及来自录制并转录的专家CT解读、由放射科医生撰写的推理内容。专家标注提供了直接监督,并引导了额外的基于报告的合成推理。端到端监督微调(SFT)之后采用群体相对策略优化(GRPO),在胸部和腹部异常集合上使用可验证的奖励。该模型支持异常分类、报告生成,以及带有可审查观察结果、鉴别诊断和不确定性信息的交互式推理。评估涵盖公开的CT基准测试和一个留出的NIH队列。在CT-RATE上,NV-Reason-CT在无需任务特定分类头的情况下取得了0.614的宏F1(macro-F1)和0.871的宏AUROC(macro-AUROC);生成的报告取得了基于报告的宏F1为0.592。在一项与专家放射科医生进行的初步研究中,AI辅助审阅获得了良好的置信度评分,并使报告的平均解读和撰写时间减少了50%。我们发布了模型和训练代码,以支持体积医学影像可解释AI的可持续重现研究。
cs.CV / 43 / 2609.27523

M3D-Net: Hierarchical Coordination of Spatial Context, Feature Reuse, and Differential Attention for Mammography Classification

M3D-Net:面向乳腺X线摄影分类的空间上下文、特征复用与差分注意力的分层协调
Yu, Zheng, Li, Xinhang, Gao, Jiabao, Wang, Boyang, Li, Xiang
Abstract
Breast image classification requires local detail and global tissue context, yet these cues can weaken as representations deepen. We present M3D-Net, a mammography encoder that hierarchically coordinates multi-scale coordinate attention, bounded dynamic feature reuse, and differential attention through resolution-aware operator placement. Within-stage retrieval preserves access to earlier features, coordinate-aware aggregation integrates local and global context, and differential attention operates at coarse resolutions. We evaluate image-only classification on AISSLab mammography and an adapted image--clinical model on BrEaST ultrasound. Against EdgeNeXt, RepViT, and TransXNet, the proposed implementations achieve the highest recorded validation accuracy and late-training accuracy, with the lowest endpoint cross-entropy loss. Validation accuracies reach 97.78\% and 80.39\%, respectively. These results support further evaluation of hierarchical coordination across breast imaging settings; repeated-seed, component-controlled, and independent evaluations remain necessary.
Chinese Translation
乳腺图像分类既需要局部细节,也需要全局组织上下文,然而随着表示层数加深,这些线索可能被削弱。我们提出了M3D-Net,一种乳腺X线摄影(mammography)编码器,通过分辨率感知的算子放置,分层地协调多尺度坐标注意力、有界动态特征复用与差分注意力(differential attention)。阶段内检索机制保留了对早期特征的访问,坐标感知聚合整合了局部与全局上下文,而差分注意力则在粗分辨率上运行。我们在AISSLab乳腺X线摄影数据集上评估了仅基于图像的分类,并在BrEaST超声数据集上评估了 adapted image--clinical(图像-临床)模型。与EdgeNeXt、RepViT和TransXNet相比,所提出的实现取得了最高的验证准确率和训练后期准确率,以及最低的终点交叉熵损失。验证准确率分别达到97.78%和80.39%。这些结果支持在不同乳腺成像场景下对分层协调机制进行进一步评估;但仍有必要开展多种子重复实验、组件消融控制实验以及独立评估。
cs.CV / 44 / 2609.27533

ICM: Intra-class Mixing for Domain Adaptation in Adverse Weather

ICM:面向恶劣天气域自适应的类内混合方法
Li, Boying, Liu, Chang, Wilde, Britta Ayano, Kovács, György, Adewumi, Tosin, Backe, Björn, Mokayed, Hamam
Abstract
Unsupervised domain adaptation (UDA) for semantic segmentation remains challenging under adverse weather conditions because severe appearance changes enlarge the domain gap and degrade the reliability of pseudo labels in the target domain. To address this problem, we propose an Intra-Class Mixing Consistency (ICM) framework that enforces prediction consistency between an intra-class mixed image and its original counterpart. Unlike previous mixing-based consistency methods that combine regions across different images or domains and may introduce unrealistic semantic inconsistencies, ICM performs mixing within the same image and semantic class, preserving realistic semantic layout for consistency regularization. With ICM, we establish a new state-of-the-art performance for clear-to-adverse-weather unsupervised domain adaptation (UDA) in semantic segmentation. On the Cityscapes $\rightarrow$ ACDC benchmark, our method achieves 75.7\% mIoU, outperforming the previous state of the art by +1.9 pp, demonstrating its effectiveness in mitigating class confusion under challenging environmental conditions. The code is provided in the supplementary material.
Chinese Translation
在恶劣天气条件下,面向语义分割的无监督域自适应(UDA)仍然具有挑战性,因为剧烈的外观变化扩大了域间差距,并降低了目标域伪标签的可靠性。为解决这一问题,我们提出了一种类内混合一致性(Intra-Class Mixing Consistency, ICM)框架,该框架强制类内混合图像与原始图像之间的预测一致性。与以往基于混合的一致性方法(其跨不同图像或域组合区域,可能引入不真实的语义不一致性)不同,ICM 在同一图像和同一语义类别内进行混合,从而保留真实的语义布局用于一致性正则化。借助 ICM,我们在清晰天气到恶劣天气的语义分割无监督域自适应(UDA)任务上建立了新的最先进性能。在 Cityscapes $ ightarrow$ ACDC 基准上,我们的方法达到了 75.7% mIoU,比以往最先进方法提升了 1.9 个百分点,证明了其在应对具有挑战性的环境条件下缓解类别混淆问题的有效性。代码见补充材料。
cs.CV / 45 / 2609.27611

A generalizable structural brain MRI foundation model built through dual-priority federated pretraining

通过双优先级联邦预训练构建的具有泛化能力的结构脑部MRI基础模型
Yu, Zhen, Liu, Yang, Zhuang, Xiahai, Chen, Qingchao
Abstract
Foundation models hold promise for generalizable analysis of structural brain magnetic resonance imaging (MRI) across development, aging and disease. However, existing models are typically built through centralized pretraining on pooled data, despite privacy and governance constraints. Such pooling optimization can overemphasize cohort size and overlook complementary information from smaller, specialized cohorts. Here we present BrainFedFM, a structural brain MRI foundation model federatively pretrained on 164,707 three-dimensional scans drawn from diverse real-world data distributions and organized across 42 federated sites. BrainFedFM uses dual-priority federated pretraining, coupling spatial-priority masking at each site with site-priority aggregation at the server to emphasize informative anatomical regions locally and prioritize site contributions globally. Across 20 downstream datasets spanning 17 classification, regression and segmentation tasks, BrainFedFM achieved the state-of-the-art performance (mean rank 1.68, 50\% gain) across seven models, including four centralized foundation models, while showing particularly consistent advantages in classification and regression and robustness across underrepresented populations. These findings demonstrate the generalizability of BrainFedFM and highlight federated pretraining as a practical strategy for developing neuroimaging foundation models from distributed data without pooling raw images.
Chinese Translation
基础模型为跨发育、衰老和疾病的结构脑部磁共振成像(MRI)的泛化分析带来了希望。然而,尽管存在隐私和治理方面的限制,现有模型通常通过集中式预训练在汇总数据上构建。这种数据汇总优化可能过度强调队列规模,而忽略了来自较小、专业化队列的互补信息。在此,我们提出了BrainFedFM,一个在来自多样化真实世界数据分布、组织于42个联邦站点的164,707个三维扫描图像上进行联邦预训练的结构脑部MRI基础模型。BrainFedFM采用双优先级联邦预训练,将每个站点的空间优先级掩码与服务器端的站点优先级聚合相结合,在局部强调信息丰富的解剖区域,并在全局优先考虑各站点的贡献。在涵盖17项分类、回归和分割任务的20个下游数据集上,BrainFedFM在包括四个集中式基础模型在内的七个模型中取得了最先进的性能(平均排名1.68,提升50%),并在分类和回归任务以及面对代表性不足人群时展现出尤为一致的优势和稳健性。这些发现证明了BrainFedFM的泛化能力,并突显联邦预训练作为一种无需汇总原始图像即可从分布式数据开发神经影像基础模型的实用策略。
cs.CV / 46 / 2609.27620

InGuard: Towards Generalized Inner Guardrail for Safe Text-to-Image Generation

InGuard:面向安全文生图生成的通用内部防护机制
Wang, Zeyu, Li, Xiaodan, Li, Zhiwen, Chen, Yuefeng, Xue, Hui
Abstract
Modern text-to-image (T2I) models generate high-quality images from arbitrary user prompts, yet they can just as easily produce not-safe-for-work (NSFW) content. Conventional outer guardrails consist of two components: a prompt classifier that checks for risk before generation, and a post-hoc image classifier that checks the fully generated image. In this design, both classifiers operate outside the generation pipeline and do not use the model's own representations. This separation can limit prompt-screening accuracy, while the image-side check runs only after the full generation cost has been spent. Moreover, a flagged prompt can only be rejected, even when it could be adjusted to produce a safe image. In this work, we propose the Inner Guardrail (InGuard), a safety framework that works inside the pipeline on the model's own representations, leaving base-model parameters untouched. First, a risk classifier grades each prompt as unsafe, risky, or benign based on the text encoder's embeddings, with no external language model. Second, SAGE (Soft-gated Asymmetric Guardrail for Embeddings) modifies the embeddings of risky prompts, aiming to return a safe image instead of a refusal. Third, a latent detector checks the one-step clean latent estimate midway through denoising, reaching nearly image-level performance and halting generation when risk is detected. We also construct the RevGen Safety Benchmark to evaluate T2I safety under realistic conditions: 10,000 prompts built through real-image reverse generation, with a rewriting step that supplies controlled intellectual-property (IP) characters, covering graded porn/gore risks, categorical IP risks, and benign negatives. Across five open-weight T2I models, InGuard reaches 97.9-98.8% safety rate, matching or exceeding the outer guardrail, with 57.5-73.5% less benign disturbance, ~3.7x fewer parameters, and 50-55.6% of denoising steps skipped.
Chinese Translation
现代文本生成图像模型能够根据任意用户提示词生成高质量图像,但同样也可能生成不适宜公开浏览(NSFW)的内容。传统的外部防护机制由两个组件构成:一个在生成前检查风险性的提示词分类器,以及一个事后检查已完整生成图像的图像分类器。在这种设计中,两个分类器均在生成流程之外运行,且不使用模型自身的内部表示。这种分离式设计会限制提示词筛查的准确性,而图像端的检查只有在付出完整的生成开销之后才能进行。此外,被标记的提示词只能被直接拒绝,即使对其进行调整本可以生成安全的图像。在本工作中,我们提出了内部防护机制InGuard(Inner Guardrail),这是一个在生成流程内部基于模型自身表示运行的安全框架,且无需改动基础模型的参数。首先,一个风险分类器基于文本编码器的嵌入表示将每个提示词判定为不安全、有风险或良性三类,且无需借助外部语言模型。其次,SAGE(Soft-gated Asymmetric Guardrail for Embeddings,基于嵌入的软门控非对称防护机制)对有风险提示词的嵌入表示进行修改,旨在返回安全的图像而非直接拒绝。第三,一个潜在空间检测器在去噪过程中途对单步清洁潜在估计进行检查,可达到接近图像级别的检测性能,并在检测到风险时中止生成。我们还构建了RevGen Safety Benchmark基准数据集,用于在真实条件下评估文生图模型的安全性:该数据集包含通过真实图像反向生成构建的10,000条提示词,并通过改写步骤引入受控的知识产权(IP)角色,涵盖了分级的色情/血腥风险、类别化的IP风险以及良性负样本。在五个开源权重的文生图模型上,InGuard达到了97.9%-98.8%的安全率,持平或超越外部防护机制,同时对良性内容的干扰减少了57.5%-73.5%,参数量减少约3.7倍,并跳过了50%-55.6%的去噪步骤。
cs.CV / 47 / 2609.27671

SGDet3D++: Geometry-Grounded Semantics for 4D Radar and Camera 3D Object Detection

SGDet3D++:面向4D雷达与相机三维目标检测的几何锚定语义方法
Bai, Xiaokai, Fan, Zhenyu, Zheng, Lianqing, Wang, Songkai, Cao, Si-Yuan, Shen, Hui-liang
Abstract
4D radar complements dense image semantics with long-range geometry and radial motion, but existing radar--camera detectors largely solve \emph{where} to align the modalities while leaving \emph{whether} a piece of evidence supports an evolving object hypothesis implicit. An image token may describe an occluder, a nearby radar return may belong to another object, and a pose-aligned memory slot may carry incompatible motion. We formulate \emph{hypothesis-conditioned evidence grounding}, which separates candidate access from evidence use: semantic, geometric, or temporal evidence is filtered or conditioned by the evolving 3D state before updating the corresponding query. \sgdetpp{} instantiates this principle through Anchor-Grounded Semantic Retrieval (AGR), which conditions deformable image retrieval on pooled anchor-consistent radar support; Geometry-Consistent Anchor Refinement (GCR), which attentively aggregates individual associated returns; and Doppler-Verified Correspondence (DVC), which replaces history only when current radial motion contradicts it. \sgdetpp{} improves the strongest compared method by 3.82 mAP and 6.82 ODS on OmniHD-Scenes and by 6.82 mAP and 9.22 NDS on ManTruckScenes, while also leading the listed methods in the TJ4DRadSet test comparison. Mechanism-targeted evaluations show that AGR improves strict AP in every projected-occlusion bin, the yaw-aligned box gate raises target-return purity from 29.95\% to 58.87\%, and DVC preserves 96.11\% of motion-consistent history while retaining 75.90\% contradiction recall. Code will be released.
Chinese Translation
4D毫米波雷达以其远程几何信息和径向运动信息补充了稠密的图像语义,但现有的雷达-相机检测器主要解决的是"在哪里"对齐两种模态,而将"是否"一条证据支持某个不断演化的目标假设这一问题隐式处理。图像特征可能描述的是遮挡物,附近雷达回波可能属于另一目标,位姿对齐的记忆槽位可能携带不相容的运动信息。我们提出了"假设条件化证据锚定"(hypothesis-conditioned evidence grounding)的概念,将候选访问与证据使用分离:语义、几何或时序证据在被用于更新相应查询之前,需经过不断演化的三维状态的过滤或条件化。SGDet3D++通过以下模块实现这一原则:锚点锚定语义检索(Anchor-Grounded Semantic Retrieval, AGR),基于池化的锚点一致雷达支持对可变形图像检索进行条件化;几何一致锚点精化(Geometry-Consistent Anchor Refinement, GCR),通过注意力机制聚合各个关联回波;以及多普勒验证对应(Doppler-Verified Correspondence, DVC),仅当当前径向运动与历史信息矛盾时才替换历史。在OmniHD-Scenes数据集上,SGDet3D++相比最强对比方法提升了3.82 mAP和6.82 ODS;在ManTruckScenes数据集上提升了6.82 mAP和9.22 NDS,同时在TJ4DRadSet测试对比中也领先于所列方法。针对机制的评估表明:AGR在每个投影遮挡区间均提升了严格AP;偏航对齐的框门控将目标回波纯度从29.95%提升至58.87%;DVC在保留75.90%矛盾召回率的同时,保持了96.11%的运动一致历史信息。代码将开源发布。
cs.CV / 48 / 2609.27675

Track2Art: Motion-Centric Articulated Object Model Recovery from 2D Point Trackers

Track2Art:基于2D点跟踪的运动中心式铰接物体模型恢复
Li, Xiaotong, Jing, Yixiong, Ding, Junsheng, Li, Weihang, Busam, Benjamin, Wang, Guangming, Sheil, Brian
Abstract
Understanding articulated objects is fundamental for robotic interaction, requiring accurate rigid-part discovery and the recovery of their kinematic relations. Existing approaches often treat articulation as a by-product of reconstructed geometry or recover it through per-instance optimization. We instead build on the hypothesis that articulation is directly observable from persistent motion: points on the same rigid part move coherently, while relative motion between parts reveals their kinematic constraints. We present Track2Art, a motion-centric framework for recovering structured articulated objects from RGB-D interaction videos. Track2Art lifts tracked image points into persistent 3D trajectories and combines pretrained tracking features, visual descriptors, and explicit trajectory geometry. These representations are grouped into a variable number of rigid-part hypotheses and subsequently used to recover directed kinematic relations, joint types, and joint geometry through rotation-equivariant learned--analytic reasoning. On the aligned 20-object PartNet-Mobility suite, Track2Art achieves 0.695 Point IoU and 0.410 end-to-end J@20, while requiring neither ground-truth part counts nor test-time optimization.
Chinese Translation
理解铰接物体是机器人交互的基础,需要准确发现刚体部件并恢复其运动学关系。现有方法通常将铰接视为重建几何的副产品,或通过逐实例优化来恢复铰接信息。我们则基于这样一个假设:铰接可以直接从持久性运动中观测得到——同一刚体部件上的点运动保持一致,而部件之间的相对运动则揭示了它们的运动学约束。我们提出了Track2Art,这是一个从RGB-D交互视频中恢复结构化铰接物体的运动中心式框架。Track2Art将跟踪得到的图像点提升为持久性的3D轨迹,并结合预训练的跟踪特征、视觉描述子以及显式的轨迹几何信息。这些表征被分组为可变数量的刚体部件假设,随后通过旋转等变的“学习—解析”混合推理,用于恢复有向运动学关系、关节类型和关节几何。在对齐的20个物体PartNet-Mobility测试集上,Track2Art达到了0.695的Point IoU和0.410的端到端J@20,且既不需要真实的部件数量,也无需测试时优化。
cs.CV / 49 / 2609.27677

RoadOcc Learns When to Persist, Transport, or Refresh Memory for Roadside Occupancy Prediction

RoadOcc:学习何时持久化、传输或刷新记忆以用于路侧占用预测
Bai, Xiaokai, Yang, Lei, Wang, Songkai, Zheng, Lianqing, Cao, Si-Yuan, Shen, Hui-liang
Abstract
Fixed roadside cameras repeatedly observe a stable scene overlaid by sparse moving traffic. Temporal memory can recover weak observations, but reusing moving evidence at stale locations can corrupt occupancy predictions. Motion compensation addresses displacement, while reliance on the resulting history remains a separate learning problem. We introduce RoadOcc, which learns soft routing among fixed-coordinate history (\emph{Persist}), velocity-addressed history (\emph{Transport}), and current evidence (\emph{Refresh}). Motion state and class-consistent historical support supervise these source choices. Dynamic-aware cross-attention (DCA) updates candidate locations, multi-scale voxel velocity estimation (VVE) constructs transport addresses from multi-scale current--history correspondence, and velocity-guided dynamic sparse fusion (VDSF) combines routed evidence under fixed sparse-token budgets. On InfraOcc, RoadOcc reaches 65.29 mIoU and 32.37 dynamic mIoU, gains of 4.44 and 4.71 over STCOcc. Controlled address experiments show that VVE raises dynamic mIoU by 0.87 over fixed-coordinate reading. Across three seeds, supervised P/T/R adds 1.40 dynamic points over motion-corrected retrieval, while removing Refresh costs 0.32 points. Results from two transfer models, Occ3D-nuScenes, and longer intervals provide additional support. Code will be released.
Chinese Translation
固定路侧相机反复观测由稀疏运动交通叠加的稳定场景。时序记忆可以恢复微弱的观测,但在过时位置复用运动目标的历史证据可能会破坏占用预测。运动补偿可以解决位移问题,而对补偿后历史信息的依赖程度则是另一个需要学习的独立问题。我们提出 RoadOcc,它学习在固定坐标历史(Persist,持久化)、速度寻址历史(Transport,传输)和当前证据(Refresh,刷新)之间进行软路由。运动状态和类别一致的历史支持用于监督这些来源选择。动态感知交叉注意力(DCA)更新候选位置,多尺度体素速度估计(VVE)基于多尺度的当前—历史对应关系构建传输地址,速度引导的动态稀疏融合(VDSF)在固定稀疏标记预算下融合路由后的证据。在 InfraOcc 数据集上,RoadOcc 达到 65.29 mIoU 和 32.37 动态 mIoU,相比 STCOcc 分别提升 4.44 和 4.71。受控地址实验表明,VVE 相比固定坐标读取使动态 mIoU 提高 0.87。在三个随机种子下,经监督的 Persist/Transport/Refresh 相比运动校正检索额外提升 1.40 个动态点,而移除 Refresh 则损失 0.32 个点。来自两个迁移模型、Occ3D-nuScenes 以及更长时距间隔的结果提供了额外支持。代码将被开源发布。
cs.CV / 50 / 2609.27681

CasCVS-Net: A Staged Multi-Task Cascade for Critical View of Safety Assessment

CasCVS-Net:一种用于安全关键视野评估的分阶段多任务级联网络
Toh, Bock-Zien, Ren, Yuanchuan, Yu, Tay Aw, Ong, Ng Khee, Mao, Zhehua, Bano, Sophia
Abstract
Automated assessment of the Critical View of Safety (CVS) in laparoscopic cholecystectomy requires both recognition of the three CVS criteria and anatomical grounding in small, rare, and often occluded hepatocystic structures. Learning-based methods differ in the anatomical information they use, from image-level classification to detection, segmentation, or graph-based reasoning, yet grounding the safety-critical anatomy remains the main bottleneck. We propose CasCVS-Net, a staged multi-task cascade that jointly performs object detection, semantic segmentation, and CVS assessment, trained on the Endoscapes dataset. The model couples the tasks through predicted anatomy: predicted boxes guide segmentation, and predicted masks provide region-level features for CVS classification, so CVS assessment at inference uses only model predictions rather than ground-truth annotations. To reduce optimisation instability in this coupled setting, training progresses from detection to detection-segmentation and then to the full three-task cascade, followed by task-wise fine-tuning. Evaluation on the public unseen test set shows that CasCVS-Net improves over matched single-task baselines on all three tasks, achieving 32.0 detection mAP, 46.8 semantic mIoU, 15.3 rare-anatomy mIoU, and 67.2 CVS mAP. It outperforms the state-of-the-art LG-CVS and SV2LSTG by 6.3% and 4.5% relative CVS mAP, respectively, corresponding to 4.0 and 2.9 mAP points. These results show that staged task coupling through predicted boxes and masks improves anatomical grounding for CVS assessment, particularly for rare hepatocystic structures.
Chinese Translation
腹腔镜胆囊切除术中安全关键视野的自动化评估既需要识别三项CVS标准,又需要对细小、罕见且常被遮挡的肝胆囊结构进行解剖学定位。基于学习的方法在所使用的解剖信息上有所不同,从图像级分类到检测、分割或基于图的推理,然而对安全关键解剖结构的定位仍是主要瓶颈。我们提出CasCVS-Net,这是一种分阶段的多任务级联网络,可联合执行目标检测、语义分割和CVS评估,并在Endoscapes数据集上进行训练。该模型通过预测的解剖结构将各任务耦合起来:预测框引导分割,预测掩膜为CVS分类提供区域级特征,因此在推理时CVS评估仅使用模型预测而非真实标注。为了减少这种耦合设置下优化的不稳定性,训练过程从检测到检测-分割,再到完整的三任务级联逐步推进,随后进行逐任务微调。在公开的未见测试集上的评估表明,CasCVS-Net在全部三个任务上均优于相应的单任务基线,实现了32.0的检测mAP、46.8的语义mIoU、15.3的罕见解剖结构mIoU以及67.2的CVS mAP。它分别以6.3%和4.5%的相对CVS mAP提升超越当前最先进的LG-CVS和SV2LSTG,对应4.0和2.9个mAP点。这些结果表明,通过预测框和预测掩膜进行分阶段任务耦合,能够改善CVS评估的解剖学定位,尤其是对罕见的肝胆囊结构。
cs.CV / 51 / 2609.27682

Gender Bias in Vision-Language In-Context Learning

视觉-语言上下文学习中的性别偏见
Xiang, Tong, Garcia, Noa, Nakashima, Yuta
Abstract
In-context learning (ICL) enables large vision-language models (LVLMs) to perform tasks by following patterns from in-context examples, yet its potential to amplify societal biases remains underexplored. We systematically investigate how ICL influences gender bias in LVLMs through VL-BICLE, an evaluation framework comprising six ICL settings, three tasks, and four datasets. Our experiments on six LVLMs reveal that gendered ICL demonstrations act as a directional force, shifting model bias toward the demonstrated gender through a cross-gender mechanism that disproportionately degrades performance on the opposite gender. This effect appears in image captioning and pronoun prediction but not in visual question answering, indicating that gendered ICL influences bias only when the task output involves gendered language. Similarity-based retrieval methods inherit the training pool's gender imbalance and offer no debiasing advantage, while standard quality metrics remain blind to these bias shifts. To mitigate this bias, we replace real in-context images with synthetic ones from stable diffusion models while keeping captions unchanged. This simple intervention reduces gender bias without degrading caption quality.
Chinese Translation
上下文学习(In-context Learning, ICL)使大型视觉-语言模型(LVLMs)能够通过遵循上下文示例中的模式来执行任务,然而其对放大社会偏见的潜在影响仍未得到充分研究。我们通过一个包含六种ICL设置、三项任务和四个数据集的评估框架VL-BICLE,系统地研究了ICL如何影响LVLMs中的性别偏见。在六个LVLMs上的实验表明,带有性别倾向的ICL示例会作为一种方向性力量,通过一种跨性别机制将模型偏见向所示例的性别偏移,并对相反性别的性能造成不成比例的下降。这种效应出现在图像描述生成和代词预测任务中,但在视觉问答中并未出现,表明性别化ICL仅在任务输出涉及性别化语言时才会影响偏见。基于相似性的示例检索方法会继承训练池中的性别不平衡,且在去偏见方面并无优势;而标准的质量评估指标对这些偏见变化视而不见。为缓解这一偏见,我们在保持描述文本不变的情况下,将真实的上下文图像替换为来自稳定扩散(Stable Diffusion)模型的合成图像。这一简单的干预措施在不降低描述质量的前提下减少了性别偏见。
cs.CV / 52 / 2609.27696

SynSeq: End-to-End SYNTAX Score Prediction from Coronary Angiography Videos

SynSeq:基于冠状动脉造影视频的端到端SYNTAX评分预测
Baumann, Christoph, Schweitzer, Ronny, Pavo, Noemi, Attenberger, Ulrike, Loewe, Christian, Seeböck, Philipp
Abstract
The SYNTAX score is an established tool for assessing coronary artery disease and guiding revascularization treatment decisions. However, its manual estimation from coronary angiography videos by clinical experts is time-consuming and subject to inter-reader variability. While machine learning has shown promise in automating this process, prior work has primarily focused on lesion detection, characterization, or binary disease classification, leaving direct SYNTAX score prediction relatively unexplored. We propose SynSeq, a video-based method for direct SYNTAX score prediction. It combines targeted preprocessing with a tailored training strategy using a zero-inflation-aware loss and linear target scaling. Evaluated on the public CardioSyntax dataset, SynSeq significantly outperforms previous state-of-the-art methods, improving $R^2$ by 0.55, reducing prediction bias by 93.1% and achieving more consistent performance across annotations from three independent expert graders. In addition, SynSeq achieves a weighted $F_1$-score of 0.80 for revascularization treatment recommendations, slightly below inter-expert agreement. These results demonstrate the potential of SynSeq to provide consistent, automated SYNTAX score assessment and reliable decision support for coronary revascularization planning.
Chinese Translation
SYNTAX评分是评估冠状动脉疾病并指导血运重建治疗决策的成熟工具。然而,由临床专家从冠状动脉造影视频中手动估算该评分既耗时,又存在阅片者之间的差异。尽管机器学习在自动化这一过程中已展现出前景,但以往工作主要集中于病变检测、病变特征描述或二元疾病分类,对直接预测SYNTAX评分的探索相对较少。我们提出SynSeq,一种基于视频的SYNTAX评分直接预测方法。该方法将针对性预处理与定制化训练策略相结合,采用零膨胀感知损失和线性目标缩放。在公开的CardioSyntax数据集上的评估表明,SynSeq显著优于先前的最先进方法:$R^2$提升0.55,预测偏差降低93.1%,并在三位独立专家评分者的标注上取得更为一致的表现。此外,SynSeq在血运重建治疗建议方面达到0.80的加权$F_1$分数,略低于专家间的一致性水平。这些结果表明,SynSeq有望提供一致、自动化的SYNTAX评分评估,并为冠状动脉血运重建规划提供可靠的决策支持。
cs.CV / 53 / 2609.27710

FFM-CP: Cross-Backbone Fusion of Vision-Language Foundation Models for Few-Shot Computational Pathology

FFM-CP:面向小样本计算病理学的视觉-语言基础模型跨骨干网络融合
Nguyen, Anh-Tien, Dang, Trung DQ., Diep, Nghiem Tuong, Nguyen, Bui Ngoc Han, Mai, Tan-Ha, Maurer, Miriam Cindy, Nguyen, Phuong Hoa, Nguyen, Thi Thuy Uyen, Park, Youngjun, Sonntag, Daniel, Nguyen, Duy Minh Ho, Hauschild, Anne-Christin
Abstract
Pathology vision-language foundation models vary in performance across diseases and tasks, with no single model consistently performing best. The high cost of expert pathology annotation can also limit the labeled data available for task-specific adaptation. Combining complementary pretrained representations is a potential approach to these limitations, yet learning an effective fusion from few labeled examples remains challenging. We introduce Few-shot Fusion Foundation Models of Computational Pathology (FFM-CP), which is a framework that combines multiple pathology vision-language models in the few-shot learning setting. The framework first aligns heterogeneous representations using a closed-form Orthogonal Procrustes transformation estimated from corresponding support images. This alignment preserves within-model feature geometry without training an additional alignment network. Within the aligned space, a unified graph enables information exchange across backbones by jointly refining support-image features and visual and textual class prototypes. These refined representations support complementary text-prototype and case-retrieval branches that capture semantic class knowledge and within-class visual variation, respectively. Each branch learns to combine predictions from all ordered backbone pairs, allowing queries encoded by one model to draw on evidence represented by another. We evaluate three backbone combinations on six histopathology datasets at 4, 8, and 16 shots per class. FFM-CP achieves higher mean macro-F1 than the strongest individually adapted member of each fused set in 50 of 54 comparisons. These findings suggest that combining complementary pretrained representations can improve histopathological classification when annotations are limited.
Chinese Translation
病理学视觉-语言基础模型在不同疾病和任务上的表现各异,没有任何单一模型能够始终表现最佳。专家病理学标注的高成本也限制了可用于任务特定适配的标注数据量。融合互补的预训练表征是应对这些局限的一种潜在方法,然而从少量标注样本中学习有效融合仍然具有挑战性。我们提出了计算病理学小样本基础模型融合框架(Few-shot Fusion Foundation Models of Computational Pathology,FFM-CP),这是一个在少样本学习设置下融合多个病理学视觉-语言模型的框架。该框架首先使用由对应支持集图像估计的闭式正交Procrustes变换来对齐异构表征。这种对齐无需训练额外的对齐网络,即可保持模型内部的特征几何结构。在对齐后的空间中,一个统一的图结构通过联合精炼支持集图像特征以及视觉和文本类别原型,实现跨骨干网络的信息交换。这些精炼后的表征支持互补的文本原型分支和病例检索分支,分别捕捉语义类别知识和类内视觉变化。每个分支学习组合所有有序骨干网络对的预测结果,使由一个模型编码的查询能够利用另一个模型所表征的证据。我们在六个组织病理学数据集上,以每类4、8、16个样本的设置评估了三种骨干网络组合。在54次比较中,FFM-CP有50次的平均宏F1高于各融合集合中经单独适配的最强成员。这些发现表明,在标注数据有限的情况下,融合互补的预训练表征可以提升组织病理学分类性能。
cs.CV / 54 / 2609.27728

NeuralSRNF: Neural Square Root Normal Fields for the Statistical Shape Analysis and Generation of Nonrigid 3D and 4D Objects

NeuralSRNF:用于非刚性3D和4D对象统计形状分析与生成的神经平方根法向场
Nizamani, Awais, Laga, Hamid, Wang, Guanjin, Boussaid, Farid, Bennamoun, Mohammed, Srivastava, Anuj
Abstract
We introduce NeuralSRNF, a novel framework for the statistical shape analysis and generation of genus-zero 3D and 4D objects that undergo nonrigid deformations. Traditional methods rely on complex and computationally expensive nonlinear elastic metrics that measure bending and stretching. Recent advances in elastic shape analysis achieve computational efficiency by mapping input 3D shapes to the space of Square Root Normal Fields (SRNFs) where the L2 metric approximates the partial elastic metric, significantly facilitating the process of computing geodesics and summary statistics. SRNFs, however, are not invertible, and the numerical algorithms used to map SRNFs back to the original space of surfaces remain computationally very expensive and often lead to approximate results. This paper addresses this fundamental SRNF inversion problem using a novel neural representation, termed NeuralSRNF. Unlike the commonly used numerical SRNF, NeuralSRNF is (1) continuous, and thus resolution-agnostic, enabling full functional shape analysis, (2) more accurate, and (3) computationally more efficient as it can compute inverse SRNF maps along a geodesic path in less than 3 s compared to over 10 min for the numerical SRNF. We demonstrate, using various datasets, the utility and efficiency of the proposed NeuralSRNF in multiple elastic 3D and 4D shape analysis tasks such as geodesic computation, deformation transfer, statistical summaries computation, and 3D shape generation. We show that it outperforms competing methods on most evaluated datasets and metrics by a wide margin in both accuracy and computational efficiency. The source code and additional results are available at https://awaisnizamani16.github.io/awais/NeuralSRNF/.
Chinese Translation
我们提出了NeuralSRNF,这是一个用于经历非刚性形变的亏格为零的3D和4D对象的统计形状分析与生成的新型框架。传统方法依赖于复杂且计算代价高昂的非线性弹性度量来衡量弯曲和拉伸。弹性形状分析领域的最新进展通过将输入的3D形状映射到平方根法向场(Square Root Normal Fields, SRNFs)空间来实现计算的高效性,在该空间中L2度量可以近似偏弹性度量,从而显著简化了计算测地线和摘要统计量的过程。然而,SRNF不可逆,且用于将SRNF映射回原始曲面空间的数值算法计算代价非常高,且通常只能得到近似结果。本文利用一种新颖的神经表示(称为NeuralSRNF)来解决这一根本性的SRNF求逆问题。与常用的数值SRNF不同,NeuralSRNF具有以下特点:(1)连续,因此与分辨率无关,可实现完整的函数式形状分析;(2)更加精确;(3)计算效率更高,因为它沿测地路径计算逆SRNF映射所需时间不到3秒,而数值SRNF则需要超过10分钟。我们通过多个数据集验证了所提出的NeuralSRNF在多种弹性3D和4D形状分析任务中的实用性和高效性,例如测地线计算、形变迁移、摘要统计量计算以及3D形状生成。结果表明,在大多数评估数据集和指标上,该方法在准确性和计算效率方面均大幅超越现有竞争方法。源代码及更多结果可在 https://awaisnizamani16.github.io/awais/NeuralSRNF/ 获取。
cs.CV / 55 / 2609.27753

AWM-VLA: AlignedWorld Modeling for Efficient and Explainable Vision-Language-Action Policies

AWM-VLA:面向高效且可解释视觉-语言-动作策略的对齐世界建模
Lanji, An, Liu, Dawei, Li, Jin, Xu, Haoran, Chen, Mei, Tian, Yu
Abstract
Vision-language-action (VLA) models have become a powerful paradigm for generalist robotic manipulation, yet they are often reactive: the policy maps the current observation directly to an action chunk without reasoning about the long-term consequences of its decisions. Prior attempts to endow policies with world models either reconstruct future frames in pixel space---expensive and dominated by task-irrelevant detail---or decouple the world model from the policy, weakening control. We present AWM-VLA, a unified framework that embeds aligned world modeling directly inside a diffusion-transformer policy. Following the Future Latent REpresentation Alignment (FLARE) principle, we add learnable future tokens whose intermediate activations are aligned with vision-language embeddings of future observations, enabling the policy to anticipate long-term consequences while generating actions. We extend this paradigm in two ways. First, we introduce an object-centric decoupled alignment objective that predicts future object-level semantics alongside the global future embedding, improving both interpretability and multi-instruction generalization. Second, we balance the global and object-centric alignment terms against the action flow-matching loss through a principled weighting, yielding a controllable accuracy--interpretability trade-off. On RoboCasa and humanoid tabletop manipulation benchmarks, AWM-VLA outperforms prior VLA and world-model baselines by up to 21% in success rate, improves generalization to novel objects and instructions, and produces object-centric rationales that are preferred by human raters in 83 of cases. Our approach adds only a few learnable tokens to the policy and is compatible with any diffusion or flow-matching policy, making aligned world modeling an inexpensive, broadly applicable component of generalist manipulation.
Chinese Translation
视觉-语言-动作(VLA)模型已成为通用机器人操作的强大范式,但它们往往是反应式的:策略将当前观测直接映射为动作块,而无需对其决策的长期后果进行推理。以往赋予策略世界模型的尝试要么在像素空间中重建未来帧——代价高昂且被任务无关的细节所主导——要么将世界模型与策略解耦,从而削弱了控制能力。我们提出AWM-VLA,这是一个统一框架,将对齐世界建模直接嵌入扩散Transformer策略内部。遵循未来潜在表示对齐(Future Latent REpresentation Alignment, FLARE)原则,我们添加了可学习的未来token,其中间激活与未来观测的视觉-语言嵌入对齐,使策略在生成动作的同时能够预见长期后果。我们从两个方面扩展了该范式。首先,我们引入一种以物体为中心的解耦对齐目标,在预测全局未来嵌入的同时预测未来的物体级语义,从而提高了可解释性和多指令泛化能力。其次,我们通过原则性的加权方法将全局与物体中心对齐项与动作流匹配损失进行平衡,从而实现了可控的精度-可解释性权衡。在RoboCasa和人形机器人桌面操作基准测试中,AWM-VLA的成功率比以往的VLA和世界模型基线最高提升21%,提升了对新物体和新指令的泛化能力,并且其生成的以物体为中心的推理依据在83%的情况下更受人类评估者青睐。我们的方法仅向策略添加少量可学习的token,并兼容任何扩散或流匹配策略,使对齐世界建模成为通用机器人操作中一种低成本、广泛适用的组件。
cs.CV / 56 / 2609.27778

Visibility-Guided Structured Measure Flow for Class-Conditioned 3D Gaussian Generation

面向类别条件3D高斯生成的可见性引导结构化测度流
Wang, Yizhao
Abstract
3D Gaussian Splatting (3DGS) has made real-time, high-fidelity 3D rendering practical, yet turning this explicit representation into a native generative space remains an open challenge. Directly generating 3DGS objects is difficult because Gaussian primitives are unordered, variable-sized, locally dense, and highly sensitive to rendering behavior. We present VISTA-GS, a visibility-guided structured measure flow framework for class-conditioned 3D Gaussian generation. Instead of treating a 3DGS object as a flat primitive sequence or a generic latent token grid, we formulate it as a structured Gaussian measure weighted by opacity, anisotropic covariance, and multi-view visibility. Based on this formulation, we introduce a visibility-aware measure VAE that learns permutation-invariant, variable-size-compatible, and rendering-aware latent representations of 3DGS objects. We further develop a renderer-consistent measure flow that transports class-conditioned priors toward the learned 3DGS measure distribution while aligning the decoded objects with their multi-view rendering distributions. To preserve object layout and local details, VISTA-GS incorporates structure-preserving patch transport that couples global class semantics, local Gaussian measure patches, and spatial anchors during flow prediction. On VISTA-Obj30, VISTA-GS improves over the strongest baseline by roughly 60--72\% across geometry, appearance, view-consistency error, and generation speed. This design enables efficient generation of coherent, detailed, and view-consistent 3D Gaussian objects without relying on per-instance optimization, multi-view image synthesis, or reconstruction-based lifting pipelines. Project code and model checkpoints will be released.
Chinese Translation
3D高斯泼溅(3D Gaussian Splatting, 3DGS)使实时、高保真的三维渲染成为现实,然而将这种显式表示转化为原生的生成空间仍然是一个开放性挑战。直接生成3DGS对象十分困难,因为高斯基元是无序的、尺寸可变的、局部稠密的,并且对渲染行为高度敏感。我们提出了VISTA-GS,一个用于类别条件3D高斯生成的可见性引导结构化测度流框架。不同于将3DGS对象视为扁平的基元序列或通用的潜在令牌网格,我们将其表述为一种由不透明度、各向异性协方差和多视角可见性加权的结构化高斯测度。基于这一表述,我们引入了一种可见性感知的测度VAE,用于学习3DGS对象的置换不变、可变尺寸兼容且渲染感知的潜在表示。我们进一步开发了一种渲染器一致的测度流,在将类别条件先验输运至所学习的3DGS测度分布的同时,使解码后的对象与其多视角渲染分布对齐。为保持对象的布局与局部细节,VISTA-GS在流预测过程中引入了结构保持的局部块输运机制,将全局类别语义、局部高斯测度块和空间锚点耦合起来。在VISTA-Obj30数据集上,VISTA-GS在几何质量、外观、视角一致性误差和生成速度方面相较最强基线提升了约60–72%。该设计能够在不依赖逐实例优化、多视角图像合成或基于重建的提升流程的情况下,高效生成连贯、精细且视角一致的3D高斯对象。项目代码和模型检查点将会发布。
cs.CV / 57 / 2609.27779

Fusion-Aware Direct 3D Gaussian Generation with Structured Patch Latent Flows

基于结构化补丁潜在流的融合感知直接3D高斯生成
Wang, Yizhao, Wang, Jingbo, Zhang, Guantao
Abstract
Class-guided 3D object generation is important for intelligent content creation, virtual environments, and digital asset design. Although 3D Gaussian Splatting (3DGS) offers an explicit and render-efficient representation, directly generating 3D Gaussian objects is difficult because Gaussian primitives are unordered, variable-sized, locally dense, and highly sensitive to rendering. Existing 3DGS generation methods usually depend on multi-view synthesis, reconstruction, or lifted 2D priors, fusing information mainly from observed views rather than modeling the intrinsic structural distribution of 3D Gaussian objects. This paper proposes a fusion-aware hierarchical Gaussian patch representation for direct class-guided 3DGS generation with rectified flow. Irregular Gaussian sets are decomposed into canonical local patches and encoded as structured tokens. The resulting hierarchical latent space fuses global class semantics, patch-level geometry and appearance, spatial correspondence, and rendering-sensitive cues. On this basis, we design a structure-aware rectified flow model with patch-position conditioning, global-local coupled velocity prediction, and density-aware velocity weighting, enabling direct latent generation of class-conditioned 3DGS objects within seconds. A render-feedback fusion strategy further aligns latent flow learning with decoded multi-view rendering quality. Experiments show that the proposed method generates 3D Gaussian objects with more coherent geometry, sharper local details, and better multi-view consistency than baseline latent generative models. Ablation studies confirm the contributions of hierarchical information fusion, global-local coupling, density-aware supervision, and render-feedback learning while preserving practical sampling efficiency overall.
Chinese Translation
类别引导的3D物体生成对智能内容创作、虚拟环境和数字资产设计具有重要意义。尽管3D高斯泼溅(3D Gaussian Splatting, 3DGS)提供了一种显式且渲染高效的表达方式,但直接生成3D高斯物体十分困难,因为高斯基元是无序的、尺寸可变的、局部稠密的,并且对渲染高度敏感。现有的3DGS生成方法通常依赖多视角合成、重建或提升的2D先验,其信息融合主要来自观测视角,而非建模3D高斯物体内在的结构分布。本文提出一种融合感知的层次化高斯补丁表示,用于基于修正流(rectified flow)的直接类别引导3DGS生成。不规则的高斯集合被分解为规范的局部补丁,并编码为结构化标记(structured tokens)。由此形成的层次化潜在空间融合了全局类别语义、补丁级几何与外观、空间对应关系以及渲染敏感线索。在此基础上,我们设计了一种结构感知的修正流模型,具有补丁位置条件化、全局-局部耦合的速度预测以及密度感知的速度加权,能够在数秒内直接进行类别条件化的3DGS物体潜在生成。渲染反馈融合策略进一步将潜在流学习与解码后的多视角渲染质量对齐。实验表明,与基线潜在生成模型相比,所提方法生成的3D高斯物体具有更连贯的几何结构、更清晰的局部细节以及更好的多视角一致性。消融研究验证了层次化信息融合、全局-局部耦合、密度感知监督和渲染反馈学习的贡献,同时整体上保持了实用的采样效率。
cs.CV / 58 / 2609.27793

DualStabSleepNet: A Dual-Domain Diffusion Stabilization Network for Robust Sleep Staging

DualStabSleepNet:一种用于鲁棒睡眠分期的双域扩散稳定网络
Wang, Chongjian, Liu, Chen, Gao, Junjie, Zhong, Xiaofang, Han, Shiyuan, Zhang, Tong
Abstract
Existing deep learning approaches for automatic sleep staging suffer from limited robustness under heterogeneous recording conditions, where non-stationary noise, inter-subject differences and cross-dataset distribution shifts cause unstable features and poor generalization. This work proposes DualStabSleepNet (DSSNet), a dual-domain diffusion stabilization network for robust sleep staging, which improves robustness in both data and feature domains. After preprocessing multi-channel polysomnography (PSG), a continuous-scale diffusion-based stabilization module suppresses noise while preserving physiological signal structures. Stabilized signals are converted to time-frequency representations and fed into a Vision Transformer backbone. A teacher-student guided diffusion feature stabilization module further mitigates feature drift and enforces multi-level feature consistency. Evaluated on four public PSG datasets SleepEDF-20, SleepEDF-78, SHHS and ISRUC-S3, DSSNet achieves state-of-the-art accuracy of 89.2%, 88.0%, 89.7%, 86.7% with improved macro-F1 and Cohen's kappa. It obtains notable improvements on hard transitional stages (e.g., 12.5% gain for N1 on SHHS) and boosts N2/REM recognition. Under cross-dataset settings, DSSNet is robust to distribution shift and performs on par with or superior to target-dataset trained baselines, demonstrating its practical potential for real-world sleep staging across heterogeneous cohorts.
Chinese Translation
现有的自动睡眠分期深度学习方法在异构记录条件下的鲁棒性有限,非平稳噪声、受试者间差异以及跨数据集分布偏移会导致特征不稳定和泛化能力差。本工作提出了DualStabSleepNet(DSSNet),一种用于鲁棒睡眠分期的双域扩散稳定网络,可同时在数据域和特征域提升鲁棒性。在对多导睡眠图(PSG)进行预处理后,基于连续尺度扩散的稳定模块在抑制噪声的同时保留生理信号结构。稳定后的信号被转换为时频表示,并输入Vision Transformer骨干网络。教师-学生引导的扩散特征稳定模块进一步缓解特征漂移并强化多层级特征一致性。在四个公开PSG数据集SleepEDF-20、SleepEDF-78、SHHS和ISRUC-S3上评估,DSSNet达到了最先进的准确率89.2%、88.0%、89.7%、86.7%,且宏平均F1和Cohen's kappa均有所提升。在困难过渡期睡眠阶段上取得了显著改进(例如在SHHS上N1提升12.5%),并增强了N2/REM的识别能力。在跨数据集设置下,DSSNet对分布偏移具有鲁棒性,其性能与在目标数据集上训练的基线相当或更优,展现了其在异构人群中真实世界睡眠分期的实际应用潜力。
cs.CV / 59 / 2609.27794

DMM-Align: Closed-Loop Optimization for 2D-3D Registration with Dual-Role Diffusion

DMM-Align:基于双重角色扩散的2D-3D配准闭环优化
Wang, Chongjian, Gao, Junjie
Abstract
2D-3D registration remains brittle in challenging scenarios such as low overlap, occlusion, repetitive structures, and severe cross-modal ambiguity. A key reason is that existing methods improve representation learning, correspondence estimation, or pose computation in isolation, while the dominant failure mode is inherently cross-level, where errors propagate between features, correspondences, and pose. To address this limitation, we propose DMM-Align: Diffusion-based Matching Matrix Alignment, a closed-loop framework that couples correspondence refinement, pose estimation, and representation learning through a shared differentiable geometric state. Our method leverages diffusion in two coordinated roles: a geometry-aware diffusion process refines the soft matching matrix for robust correspondence estimation, while a geometry-conditioned diffusion teacher injects pose-induced supervision back into feature learning. These processes are connected via a differentiable geometric hinge that converts correspondences into a global pose and exposes geometric inconsistency to upstream modules. Extensive experiments on 7-Scenes and RGB-D Scenes V2 demonstrate that DMM-Align consistently outperforms strong baselines, especially under low-overlap and heavy-occlusion conditions, highlighting the effectiveness of closed-loop geometric feedback for robust 2D-3D registration.
Chinese Translation
2D-3D配准在低重叠、遮挡、重复结构以及严重的跨模态歧义等具有挑战性的场景下仍然十分脆弱。一个关键原因在于,现有方法孤立地改进表示学习、对应关系估计或位姿计算,而主导的失败模式本质上是跨层级的——误差会在特征、对应关系和位姿之间传播。为解决这一局限,我们提出了DMM-Align:基于扩散的匹配矩阵对齐(Diffusion-based Matching Matrix Alignment),这是一个闭环框架,通过共享的可微几何状态将对应关系精化、位姿估计和表示学习耦合在一起。我们的方法以两种协同角色利用扩散模型:一种几何感知的扩散过程对软匹配矩阵进行精化,以实现鲁棒的对应关系估计;同时,一个几何条件化的扩散教师将位姿诱导的监督信号回馈到特征学习之中。这些过程通过一个可微的几何铰链相连接,该铰链将对应关系转换为全局位姿,并将几何不一致性暴露给上游模块。在7-Scenes和RGB-D Scenes V2数据集上的大量实验表明,DMM-Align持续优于强大的基线方法,尤其是在低重叠和严重遮挡条件下,凸显了闭环几何反馈对鲁棒2D-3D配准的有效性。
cs.CV / 60 / 2609.27815

Evaluating ADC-only deep learning pipelines for breast cancer detection and segmentation using standalone diffusion-weighted MRI

基于独立弥散加权MRI的仅ADC深度学习流程在乳腺癌检测与分割中的评估
Marcos, Pablo García, Gonzĺez, Paula Puerta, Lorenzo, Guillermo, Gómez, Héctor, del Camino, Covadonga, Rio-Alvarez, Angel, González, Víctor M.
Abstract
Dynamic contrast-enhanced (DCE) imaging is the gold standard technique for the detection and characterization of breast cancer using magnetic resonance imaging (MRI). However, DCE-MRI requires long acquisition times and the administration of contrast into the bloodstream, which can cause allergic reactions. Alternatively, diffusion-weighted MRI (DW-MRI) is a standard complementary technique for breast MRI that does not require contrast, has shorter acquisition times, and enables calculation of apparent diffusion coefficient (ADC) maps that correlate with tumor cellularity. Yet, despite these technical advantages, deep learning research has focused on DCE-based models and has barely explored the tumor detection performance of DW-MRI and ADC maps either in combination with DCE-MRI or as standalone alternatives. Here, we evaluate the application of different state-of-the-art deep learning techniques for detection and segmentation of breast cancer using ADC-only images. This is, to our knowledge, the first comprehensive evaluation of ADC-only breast cancer pipelines for classification, detection, and segmentation tasks.
Chinese Translation
动态对比增强(DCE)成像是利用磁共振成像(MRI)进行乳腺癌检测与表征的金标准技术。然而,DCE-MRI需要较长的采集时间,并且需要向血液中注射对比剂,这可能引发过敏反应。相比之下,弥散加权MRI(DW-MRI)是乳腺MRI的一种标准补充技术,它无需对比剂,采集时间更短,并且能够计算与肿瘤细胞密度相关的表观弥散系数(ADC)图谱。然而,尽管具有这些技术优势,深度学习研究一直集中于基于DCE的模型,几乎没有探索DW-MRI和ADC图谱在与DCE-MRI联合使用或作为独立替代方案时的肿瘤检测性能。在本研究中,我们评估了多种最先进的深度学习技术在仅使用ADC图像进行乳腺癌检测和分割方面的应用。据我们所知,这是首个针对仅ADC乳腺癌流程在分类、检测和分割任务上的综合评估。
cs.CV / 61 / 2609.27821

Groundbench: Multi-Resolution Polygon Grounding Exposes the Geometry Gap in Vision-Language Models

GroundingBench:多分辨率多边形定位揭示视觉语言模型中的几何差距
Bian, Zhonghan, Wang, Zhenran, Li, Jinsong, Qi, Zhangyang
Abstract
Bounding-box scores on RefCOCO-family grounding leave little room to distinguish frontier vision-language systems, yet boxes discard object shape. We introduce GroundingBench, a matched benchmark that re-targets the same 1,500 image-expression-referent triples to exact-N polygons at five vertex budgets. A fixed-denominator harness separately audits filled-region intersection over union (IoU) and legal-polygon completion. The strongest tested configuration reaches 88.2 box IoU and 97.1 accuracy at IoU >= .5 (Acc@.5), versus 57.7 and 69.2 for direct polygons; because these headline scores use different references, we also compare direct polygons with predicted boxes rasterised against the same contour target, obtaining 57.7 versus 57.3 when pooled. Performance is non-monotone in N and collapses at the densest budget, where legality failures compound residual geometric error. Qwen's thinking-setting contrast is the largest tested input-preserving configuration difference; under frozen templates, false spatial cues are more damaging than false colour cues, and target preference can remain high while contour tracing is poor. Alternate masks and a continuous-area scorer preserve the principal ordering. GroundingBench therefore measures an operational output-geometry gap spanning localisation, boundary construction, serialisation, and topology, rather than latent boundary perception alone.
Chinese Translation
RefCOCO 系列定位任务上的边界框分数几乎没有空间来区分前沿视觉语言系统,而框却丢弃了物体的形状。我们提出 GroundingBench,一个匹配的基准,将同样的 1,500 个图像-表达-指称三元组重新定位为五种顶点预算下的精确 N 多边形。一个固定分母的评测框架分别审核填充区域的交并比(IoU)与合法多边形完成率。测试中最强配置达到 88.2 的框 IoU 和 IoU>=.5 下 97.1 的准确率(Acc@.5),而直接生成多边形分别为 57.7 和 69.2;由于这些主要分数使用了不同的参照,我们还将直接多边形与预测框栅格化到相同轮廓目标进行对比,合并后得到 57.7 对 57.3。性能随 N 非单调变化,并在最密集预算下崩溃,此时合法性失败与残余几何误差相互叠加。Qwen 的思考模式对比是所测试的保持输入不变的配置差异中最大的一项;在冻结模板下,错误空间线索比错误颜色线索更具破坏性,且目标偏好可以在轮廓追踪不佳的情况下保持较高水平。替代掩码和连续面积评分器保持了主要排序。因此,GroundingBench 度量的是跨越定位、边界构建、序列化和拓扑的操作性输出几何差距,而不仅仅是潜在的边界感知。
cs.CV / 62 / 2609.27848

Geometry-anchored PET-aware multimodal pseudo-CT synthesis for whole-body attenuation correction: the BIC-MAC Challenge

基于几何锚定的PET感知多模态伪CT合成用于全身衰减校正:BIC-MAC挑战赛
Nguyen, Xuan Loc, Cao, Hoang-Loc, Nguyen, Truong Thanh Hung, Ho, Phuc, Nguyen, Phuc Truong Loc, To, Nguyen Truong Toan, Cao, Hung
Abstract
The BIC-MAC challenge targets whole-body pseudo-CT synthesis from NAC-PET, Dixon MRI, and a 2D topogram for CT-less PET attenuation correction. We propose GeoPACT, a geometry-anchored multimodal framework that uses NAC-PET as the spatial reference and incorporates topogram and MRI features through gated residual fusion. Absolute coordinates and whole-body conditioning support anatomically consistent patch-based prediction. Training combines attenuation-map supervision with a differentiable PET-response surrogate to reduce errors relevant to downstream PET reconstruction. Full-resolution pseudo-CT volumes are generated using sliding-window inference without requiring CT or PET labels at test time.
Chinese Translation
BIC-MAC挑战赛旨在利用NAC-PET、Dixon MRI和2D定位片(topogram)进行全身伪CT合成,以实现无CT的PET衰减校正。我们提出GeoPACT,一种几何锚定的多模态框架,以NAC-PET作为空间参考,并通过门控残差融合引入定位片和MRI特征。绝对坐标和全身条件化机制支持解剖学一致的基于图像块(patch)的预测。训练将衰减图监督与可微分的PET响应代理相结合,以减少与下游PET重建相关的误差。全分辨率伪CT体数据采用滑窗推理生成,且在测试时无需CT或PET标签。
cs.CV / 63 / 2609.27850

GaussianDS: Depth-supervised Semantic Gaussian Splatting for Scene Understanding

GaussianDS:面向场景理解的深度监督语义高斯泼溅(Gaussian Splatting)
Zhang, Yufei, Zhan, Chenlu, Wang, Hongwei
Abstract
3D Gaussian Splatting provides an efficient representation for 3D reconstruction, and recent extensions attach semantic attributes to Gaussians for open-vocabulary scene understanding. However, lifting view-dependent 2D foundation-model outputs into 3D space introduces cross-view inconsistencies and weak geometric grounding, leading to severe semantic drift and boundary leakage. We propose GaussianDS, a depth-supervised semantic 3DGS framework that treats semantic lifting as a supervision-alignment problem and jointly optimizes RGB appearance, rendered depth, and compact semantics from scratch. Specifically, GaussianDS organizes unordered multi-view images into a pose-aware pseudo-video trajectory to propagate view-consistent masks via SAM2. During joint optimization, scale-shift-aligned monocular depth supervision and depth total-variation regularization stabilize Gaussian geometry, while a depth-edge-aware refinement loss explicitly anchors semantic transitions onto physical geometric discontinuities. Extensive evaluations show that our end-to-end framework not only retains high-fidelity 3D reconstruction and real-time rendering, but also establishes superior semantic understanding. GaussianDS sets new state-of-the-art performance on LERF (60.5% mIoU) and 3D-OVS (97.79% mIoU, 90.28% mBIoU) by mitigating semantic leakage, while seamlessly facilitating downstream 3D object removal.
Chinese Translation
三维高斯泼溅(3D Gaussian Splatting)为三维重建提供了一种高效的表达方式,近期的研究扩展通过为高斯附加语义属性来实现开放词汇场景理解。然而,将依赖于视角的二维基础模型输出提升到三维空间会引入跨视角不一致性和弱几何约束,导致严重的语义漂移和边界泄漏。我们提出GaussianDS,一种深度监督的语义3DGS框架,它将语义提升视为一个监督对齐问题,并从头开始联合优化RGB外观、渲染深度和紧凑语义。具体而言,GaussianDS将无序的多视角图像组织成具有位姿感知的伪视频轨迹,通过SAM2传播视角一致的掩码。在联合优化过程中,经过尺度-偏移对齐的单目深度监督和深度全变差正则化稳定了高斯几何,而深度边缘感知的精化损失则显式地将语义过渡锚定在物理几何不连续处。大量评估表明,我们的端到端框架不仅保留了高保真三维重建和实时渲染能力,还建立了更优越的语义理解能力。通过缓解语义泄漏,GaussianDS在LERF(60.5% mIoU)和3D-OVS(97.79% mIoU、90.28% mBIoU)上取得了新的最先进性能,同时可以无缝支持下游的三维物体移除任务。
cs.CV / 64 / 2609.27868

TopoGS: Topology-Aware Anchor Feature Aggregation for Large-Scale 3D Gaussian Splatting

TopoGS:面向大规模3D高斯泼溅的拓扑感知锚点特征聚合方法
Zhang, Wei, Gong, Shiqiang, Yu, Shengkai, Wang, Zeyu, Mallet, Clement, Xiong, Zhitong, Wang, Qi
Abstract
Octree-based 3D Gaussian Splatting organizes anchors into multi-level hierarchies for level-of-detail rendering, but features at different levels are typically optimized independently, leaving the octree topology underused during feature learning. We observe that uniform cross-level aggregation produces asymmetric effects: fine-level anchors benefit from coarse context, whereas coarse-level anchors require selective information from their descendants. We therefore propose TopoGS, a topology-aware anchor feature aggregation framework with two lightweight components. Hierarchical Anchor Coupling establishes bidirectional cross-level gradient pathways by fusing per-level context triplets with a residual MLP. Structure-Aware Containment Aggregation uses octree containment and hash-based matching to distinguish anchors with valid parent-child relations from isolated anchors, then applies soft weighting to accommodate varying topological sparsity. Experiments on ten scenes from Mill19, UrbanScene3D, Tanks & Temples, MatrixCity, and WHU show consistent improvements over state-of-the-art methods. TopoGS achieves average PSNR gains of 2.13, 1.78, and 0.29 dB over the strongest reported baseline on aerial, ground-level, and synthetic-cartographic scenes, respectively, while rendering faster and using less memory. Code is available at https://github.com/WZ-CS/TopoGS.
Chinese Translation
基于八叉树的3D高斯泼溅(3D Gaussian Splatting)将锚点组织成多层级结构以实现细节层次(LOD)渲染,但不同层级的特征通常被独立优化,导致八叉树拓扑结构在特征学习过程中未得到充分利用。我们观察到,均匀的跨层级聚合会产生不对称的效果:精细层级的锚点受益于粗糙层级的上下文信息,而粗糙层级的锚点则需要从其子节点中进行选择性信息提取。因此,我们提出了TopoGS,一种拓扑感知的锚点特征聚合框架,包含两个轻量级组件。层次锚点耦合(Hierarchical Anchor Coupling)通过残差多层感知机(MLP)融合各层级的上下文三元组,建立双向跨层级梯度通路。结构感知包含聚合(Structure-Aware Containment Aggregation)利用八叉树包含关系和基于哈希的匹配,将具有有效父子关系的锚点与孤立锚点区分开来,随后应用软加权以适应不同的拓扑稀疏性。在Mill19、UrbanScene3D、Tanks & Temples、MatrixCity和WHU数据集的十个场景上的实验表明,本方法相较现有最先进方法取得了一致的性能提升。TopoGS在航空场景、地面场景和合成制图场景上分别相较已报道的最强基线方法实现了平均2.13、1.78和0.29 dB的PSNR增益,同时渲染速度更快、内存占用更少。代码已发布于 https://github.com/WZ-CS/TopoGS。
cs.CV / 65 / 2609.27890

RelCheck: Dual-Evidence Spatial Grounding for VLM Hallucination Correction

RelCheck:用于视觉语言模型幻觉修正的双证据空间定位方法
Patil, Siddhi, Saxena, Navrati, Andreopoulos, William B.
Abstract
Multimodal large language models (MLLMs) fre- quently generate text that is inconsistent with the input image. While object- and attribute-level hallucinations have received considerable attention, relational hallucinations (incorrect de- scriptions of spatial or interactive relationships between objects) remain largely unaddressed by existing post-hoc correction methods. We present RelCheck, a training-free post-hoc correction pipeline that augments object-level visual grounding with dual relational evidence: learned scene-graph triples from RelTR and deterministic spatial predicates from bounding-box geometry. These combine with a Woodpecker-style object claim layer to form a three-layer visual knowledge base, which a language model corrector uses to rewrite hallucinated text. Evaluated on LLaVA v1 13B, RelCheck achieves a total MME hallucination score of 630.0 versus 585.0 for a Woodpecker-style baseline, with the largest gain on the position subtask (+31.7 points, accuracy+ improving from 0.367 to 0.600). A four-configuration ablation confirms that both relational layers contribute independently (McNemar p = 0.025). These results show that structured relational evidence meaningfully improves post-hoc hallucination correction on the spatial reasoning subtasks where current MLLMs are most deficient.
Chinese Translation
多模态大语言模型(MLLMs)经常生成与输入图像不一致的文本。尽管对象级和属性级幻觉已受到广泛关注,但关系幻觉(即对对象之间空间或交互关系的错误描述)在很大程度上仍未被现有的事后修正方法所解决。我们提出了 RelCheck,这是一种无需训练的事后修正流水线,它在对象级视觉定位的基础上引入了双重关系证据:来自 RelTR 的学习型场景图三元组,以及来自边界框几何的确定性空间谓词。这些证据与 Woodpecker 风格的对象声明层相结合,构成一个三层视觉知识库,供语言模型修正器用于改写幻觉文本。在 LLaVA v1 13B 上的评估表明,RelCheck 的 MME 幻觉总分为 630.0,而 Woodpecker 风格的基线为 585.0,其中在位置子任务上提升最大(+31.7 分,accuracy+ 从 0.367 提升至 0.600)。四配置消融实验证实两层关系证据均具有独立贡献(McNemar p = 0.025)。这些结果表明,结构化的关系证据能在当前 MLLMs 最薄弱的空间推理子任务上显著提升事后幻觉修正的效果。
cs.CV / 66 / 2609.27901

All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation

所有模态皆平等,但视频更为平等:弥合联合视频生成中的交叉注意力差距
Rahamim, Ohad, Samuel, Dvir, Schwartz, Idan, Chechik, Gal
Abstract
Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the same event. We find that joint multimodal diffusion transformers exhibit a corresponding asymmetry in cross-modal correspondence: companion modalities develop strong correspondences to video, but the reciprocal correspondences through which they constrain video remain substantially weaker. We express both directions as comparable correspondence distributions over video tokens and define their disagreement as the reciprocal correspondence gap. We introduce RecCAR, standing for Reciprocal Cross-modal Attention Regularization, a KL regularizer that uses the well-established video-to-modality correspondence as a fixed reference and aligns the weaker modality-to-video correspondence toward it. Across joint video-motion and video-audio generation, RecCAR improves the Human Anatomy score from 0.69 to 0.75 and reduces audio-video desynchronization from 0.804 to 0.752, while improving overall generation
Chinese Translation
视频是物理事件的丰富表示,捕捉了外观、几何、运动和时间演化等信息。其他模态,如3D人体运动或音频,仅编码同一事件较为狭窄的方面。我们发现,联合多模态扩散Transformer在跨模态对应关系中表现出相应的不对称性:伴随模态与视频之间形成了强对应关系,但这些模态反过来约束视频的对应关系却明显更弱。我们将这两个方向的对应关系表示为视频token上可比较的对应分布,并将两者之间的不一致定义为互惠对应差距(reciprocal correspondence gap)。我们提出了RecCAR(Reciprocal Cross-modal Attention Regularization,互惠跨模态注意力正则化),这是一种KL正则化方法,它以成熟的视频到模态对应关系作为固定参考,将较弱的模态到视频对应关系向其对齐。在联合视频-运动和视频-音频生成任务中,RecCAR将人体解剖(Human Anatomy)分数从0.69提升至0.75,并将音视频不同步率从0.804降低至0.752,同时提升了整体生成质量。
cs.CV / 67 / 2609.27904

Spatiality-Frequency Domain Video Forgery Detection System Based on ResNet-LSTM-CBAM and DCT Hybrid Network

基于ResNet-LSTM-CBAM与DCT混合网络的空域-频域视频伪造检测系统
Liao, Zihao, Hong, Sheng, Chen, Yu
Abstract
As information technology advances, digital content has become widely adopted across diverse fields such as news broadcasting, entertainment, commerce, and forensic investiga?tion. However, the availability of sophisticated multimedia editing tools has significantly increased the risk of video and image forgery, raising serious concerns about content authenticity at both societal and individual levels.To address the growing need for robust and accurate detection methods, this study proposes a novel video forgery detection model that integrates both spatial and frequency-domain features. The model is built on a ResNet-LSTM framework enhanced by a Convolutional Block Attention Module (CBAM) for spatial feature extraction, and further incorporates Discrete Cosine Transform (DCT) to capture frequency domain information. Comprehensive experiments were conducted on several mainstream benchmark datasets, encompassing a wide range of forgery scenarios. The results demonstrate that the proposed model achieves superior performance in distinguishing between authentic and manipulated videos. Additional ablation and comparative studies confirm the contribution of each component in the architecture, offering deeper insight into the models capacity. Overall, the findings support the proposed approach as a promising solution for enhancing the reliability of video authenticity analysis under complex conditions.
Chinese Translation
随着信息技术的不断进步,数字内容已被广泛应用于新闻广播、娱乐、商业和司法调查等众多领域。然而,功能强大的多媒体编辑工具的普及显著增加了视频和图像被伪造的风险,引发了社会和个人层面对内容真实性的严重担忧。为满足对鲁棒且精确的检测方法日益增长的需求,本研究提出了一种融合空域和频域特征的新型视频伪造检测模型。该模型基于ResNet-LSTM框架构建,并通过卷积块注意力模块(CBAM)增强空间特征提取,同时引入离散余弦变换(DCT)以捕获频域信息。本研究在多个主流基准数据集上开展了全面实验,涵盖多种伪造场景。结果表明,所提出的模型在区分真实视频与篡改视频方面具有优越的性能。额外的消融实验和对比研究验证了架构中各组件的贡献,为模型能力提供了更深入的洞察。总体而言,研究结果表明,所提出的方法是一种在复杂条件下提升视频真实性分析可靠性的有前景的解决方案。
cs.CV / 68 / 2609.27915

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

UVU:通过视觉-语言统一自回归范式改进多模态理解
Kan, Zhehan, Jiang, Xinghua, Zhu, Yubo, Liu, Yanlin, Yang, Xiaochen, Wei, Zhixiang, Liu, Shifeng, Liao, Qingmin, Yang, Wenming, Li, Xin, Liu, Yinsong, Jiang, Deqiang, Sun, Xing
Abstract
Despite remarkable advancements in multimodal large language models (MLLMs), their fine-grained visual understanding is constrained by a primary reliance on sparse textual supervision. Existing efforts to introduce visual supervision typically do so during post-training, when visual representations have already been largely fixed, causing such signals to act mainly as auxiliary constraints rather than as a primary force for shaping perceptual features. In this paper, we aim to fundamentally reshape the model's perceptual backbone by incorporating vision supervision directly into the pre-training stage. We observe that pixel-level image patches and textual tokens naturally coexist in a shared, raw high-dimensional space characterized by an inherent input symmetry. Leveraging this insight, we propose UVU, a novel vision-language unified autoregressive framework that eschews vector quantization. It uniquely employs continuous visual encoding for lossless representation of visual inputs and proposes a large-scale iterative hierarchical clustering algorithm to construct a pixel-level visual codebook, thereby extending the vocabulary for unified supervision and enabling autoregressive generation of pixel-level image tokens alongside textual tokens. UVU effectively synergizes pixel-level visual perception with semantic-level visual understanding, internalizing visual reconstruction capabilities and unlocking the facilitative role of visual supervision in enhancing understanding in the pre-training stage. Extensive experiments across multiple tasks demonstrate that MLLMs are capable of achieving superior multimodal understanding performance under the supervised learning paradigm of UVU.
Chinese Translation
尽管多模态大语言模型(MLLMs)取得了显著进展,但其细粒度视觉理解能力受限于对稀疏文本监督的主要依赖。现有的引入视觉监督的工作通常在后训练阶段进行,此时视觉表征已基本固定,导致此类信号主要充当辅助约束,而非塑造感知特征的主要力量。在本文中,我们旨在通过将视觉监督直接纳入预训练阶段,从根本上重塑模型的感知骨干。我们观察到,像素级图像块和文本标记天然共存于一个具有内在输入对称性的共享原始高维空间中。基于这一洞察,我们提出了UVU——一种新颖的视觉-语言统一自回归框架,它避开了向量量化。该框架独特地采用连续视觉编码以无损表示视觉输入,并提出一种大规模迭代层次聚类算法来构建像素级视觉码本,从而扩展了统一监督的词表,使模型能够在自回归生成文本标记的同时生成像素级图像标记。UVU有效地将像素级视觉感知与语义级视觉理解相协同,内化了视觉重建能力,并释放了视觉监督在预训练阶段增强理解的促进作用。在多项任务上的大量实验表明,MLLMs在UVU的监督学习范式下能够取得更优的多模态理解性能。
cs.CV / 69 / 2609.27948

VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision

VIVAS:基于视觉-语言统一自回归监督激活VLM预训练中的视觉感知能力
Kan, Zhehan, Zhu, Yubo, Jiang, Xinghua, Wei, Zhixiang, Liu, Shifeng, Tong, Wei, Zhong, Sheng, Liao, Qingmin, Yang, Wenming, Li, Xin, Liu, Yinsong, Jiang, Deqiang, Sun, Xing
Abstract
While Vision-Language Models (VLMs) demonstrate strong capabilities, they continue to suffer from a critical limitation: insufficient fine-grained visual perception, which fundamentally limits their multimodal understanding. We attribute this bottleneck to text-dominant optimization biases during pre-training, which encourage the model to overlook fine-grained visual details, thereby limiting the capability of multimodal understanding. We investigate that overcoming this bottleneck requires two key elements: (1) a unified token space paradigm that ensures stable training dynamics, and (2) a modality-aligned dense visual supervision signal enriched with both structural granularity and semantic information to capture critical visual representations. Based on these insights, we propose VIVAS, a framework built upon the unified token space paradigm, which introduces a dense-structural-semantic vision tokenizer, which expands the textual vocabulary into a unified vision-language vocabulary by incorporating a visual vocabulary. During pretraining, VIVAS performs vision-language unified autoregressive supervision over both visual details and linguistic content, thereby enhancing visual perception to improve multimodal understanding. Trained end-to-end on 12.4T tokens, VIVAS achieves state-of-the-art performance across 7 tasks and 39 multimodal benchmarks.
Chinese Translation
尽管视觉-语言模型(Vision-Language Models, VLMs)展现出强大的能力,但它们仍然受到一个关键限制的困扰:细粒度视觉感知不足,这从根本上限制了其多模态理解能力。我们将这一瓶颈归因于预训练过程中以文本为主的优化偏差,这种偏差促使模型忽视细粒度的视觉细节,从而限制了多模态理解能力。我们研究发现,克服这一瓶颈需要两个关键要素:(1)确保训练动态稳定的统一词元空间范式;(2)兼具结构粒度与语义信息的、与模态对齐的稠密视觉监督信号,以捕获关键的视觉表征。基于这些见解,我们提出了VIVAS——一个构建于统一词元空间范式之上的框架。该框架引入了一个稠密-结构-语义视觉分词器(vision tokenizer),通过融入视觉词表,将文本词表扩展为统一的视觉-语言词表。在预训练过程中,VIVAS对视觉细节和语言内容同时执行视觉-语言统一的自回归监督,从而增强视觉感知以提升多模态理解能力。通过在12.4T词元上进行端到端训练,VIVAS在7个任务和39个多模态基准测试中取得了最先进的性能。
cs.CV / 70 / 2609.27958

ScoutNeRV: Rapid Encoding of Grid-Based Video INRs via ScoutNet

ScoutNeRV:基于ScoutNet的网格视频隐式神经表示快速编码方法
Alizada, Naser, Baghban, Farhang, Pishkar, Hashem, Mousavi, Ali
Abstract
Implicit neural representations (INRs) have emerged as a promising paradigm for video compression, providing compact neural representations with flexible spatial and temporal reconstruction. Hierarchical grid-based architectures such as HiNeRV achieve strong rate--distortion performance, but require extensive per-video optimization, resulting in high encoding costs. To address this limitation, we propose ScoutNeRV, a content-adaptive initialization framework for accelerating the optimization of hierarchical video INRs. ScoutNeRV employs a lightweight, offline-trained scout network that analyzes a small number of sampled frames and selects a suitable pre-trained expert from a memory bank through hard routing. The hierarchical grid and decoder parameters of the selected expert are then transferred to initialize the target HiNeRV model before video-specific fine-tuning. On the unseen ReadySetGo sequence, ScoutNeRV achieves an initial PSNR of $34.95$~dB, compared with $13.70$~dB for standard initialization, corresponding to a $21.25$~dB improvement before fine-tuning. After only 37 epochs, ScoutNeRV reaches $36.92$~dB and remains within $0.42$--$0.80$~dB of the 300-epoch HiNeRV baseline across the evaluated rate--distortion configurations. Furthermore, the proposed initialization achieves a $9.25\times$ wall-clock speedup in the reported runtime experiment. These results demonstrate that content-aware expert initialization can substantially reduce the optimization cost of hierarchical video INRs while retaining competitive reconstruction and compression performance. The code is available at https://github.com/nasserdeveloper/ScoutNeRV.
Chinese Translation
隐式神经表示(INR)已成为一种颇具前景的视频压缩范式,它以紧凑的神经表示实现灵活的空间和时间重建。诸如 HiNeRV 之类的层次化网格架构能够实现优异的率失真性能,但需要对每个视频进行大量优化,导致编码成本高昂。为解决这一局限,我们提出 ScoutNeRV,一个用于加速层次化视频 INR 优化的内容自适应初始化框架。ScoutNeRV 采用一个轻量级、离线训练的侦察网络(scout network),通过分析少量采样帧,并借助硬路由(hard routing)从记忆库中选择合适的预训练专家模型。随后,将所选专家的层次化网格和解码器参数迁移过来,用于初始化目标 HiNeRV 模型,再进行针对视频的微调。在未见过的 ReadySetGo 序列上,ScoutNeRV 的初始 PSNR 达到 34.95 dB,而标准初始化仅为 13.70 dB,即在微调前实现了 21.25 dB 的提升。仅经过 37 个 epoch,ScoutNeRV 便达到 36.92 dB,并且在所评估的各率失真配置下,与 300 个 epoch 的 HiNeRV 基线相比始终保持在 0.42–0.80 dB 的差距之内。此外,在所报告的运行时实验中,所提出的初始化方法实现了 9.25 倍的挂钟时间加速。这些结果表明,内容感知的专家初始化能够显著降低层次化视频 INR 的优化成本,同时保持具有竞争力的重建与压缩性能。代码已发布于 https://github.com/nasserdeveloper/ScoutNeRV。
cs.CV / 71 / 2609.27988

Task-Induced Riemannian Metrics for Vision Transformer Feature Spaces

面向视觉Transformer特征空间的任务诱导黎曼度量
Bond, Andrew, Özlü, Ege Erdem, Çimen, Tuna, Melanlioglu, Ilkin Umut, Birdal, Tolga, Erdem, Erkut, Erdem, Aykut
Abstract
Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason to believe the true task geometry has this property. The task-sensitive geometry of the feature space is given by the pullback metric $g(F) = J(F)^\top J(F)$, where $J$ is the Jacobian of the decoder's output fed to a task-specific distance, with respect to the features. Storing the full $g$ is infeasible at modern scales, and for dense outputs such as depth maps even forming $J$ is impractical. We show that whether a low-rank approximation of this metric can be learned depends on the model-decoder pair, and we characterize this with a matrix-free diagnostic $\kappa_{cap}(r)$ computable with a low number of Jacobian-vector products. For tractable pairs, we develop the Spectral Pullback Network (SPN), which learns a low-rank version of the metric from randomized power iteration, and we distill it into a $310$K-parameter importance head that predicts token importance directly from the features. When the Jacobian spectrum is too spread out for a low-rank approximation, passing the decoder's input features through a VAE bottleneck can restore tractability. Across DPT, DINOv2, CLIP, and VGGT backbones, $\kappa_{cap}(r)$ predicts which learned-metric architectures are viable. The importance head reaches Spearman $\rho = 0.998$ on DINOv2 CLS, and our geometric token pruning reduces the additional depth error of ToMe-based token selection by $25\%$ on DPT depth at prune ratio $0.5$, without fine-tuning the ViT. Project page: https://cyberiada.github.io/TaskInducedViTs/
Chinese Translation
在视觉Transformer(ViT)特征空间上运行的方法通常依赖欧氏距离或余弦相似度。这假设每个方向都同等重要,但没有理由认为真实的任务几何具有这种性质。特征空间的任务敏感几何由拉回度量(pullback metric)$g(F) = J(F)^ op J(F)$ 给出,其中 $J$ 是解码器输出(其被输入任务特定距离)相对于特征的雅可比矩阵。在现代规模下存储完整的 $g$ 是不可行的,而对于深度图等稠密输出,甚至连构造 $J$ 都不切实际。我们证明该度量的低秩近似能否被学习取决于模型-解码器配对,并通过一种无需显式构造矩阵的诊断量 $\kappa_{cap}(r)$ 来刻画这一点,该诊断量仅需少量雅可比-向量积即可计算。对于可处理的配对,我们提出了谱拉回网络(Spectral Pullback Network, SPN),它从随机化幂迭代中学习度量的低秩版本,并将其蒸馏为一个仅含31万参数的重要性头(importance head),可直接从特征预测词元(token)重要性。当雅可比谱过于分散而无法进行低秩近似时,将解码器的输入特征通过VAE瓶颈可以恢复可处理性。在DPT、DINOv2、CLIP和VGGT骨干网络上,$\kappa_{cap}(r)$ 能够预测哪些学习度量的架构是可行的。重要性头在DINOv2 CLS上达到Spearman $ ho = 0.998$,且我们的几何词元剪枝方法在剪枝率为0.5时,将DPT深度任务中基于ToMe的词元选择的额外深度误差降低了25%,而无需对ViT进行微调。项目主页:https://cyberiada.github.io/TaskInducedViTs/
cs.CV / 72 / 2609.28049

Prompt, Probe, Train, or Annotate? Single-camera sports video understanding in amateur settings

提示、探测、训练还是标注?业余场景下的单机位体育视频理解
Kodathala, Sai Varun, Pollishetty, Prashanth, Cargill, Jaylen
Abstract
Video understanding is usually benchmarked on curated, single-actor, or professionally filmed clips, and a strong score there is routinely read as evidence a model is robust enough for deployment. Amateur team sport is a useful, largely untested place to check that assumption: over eight million students played a school sport in the United States in 2024-25 alone, almost none of it filmed by more than a single fixed camera, with several candidate actors crowded into frame and no operator or second angle to fall back on. Using volleyball as a test case, we ask whether strong performance on general video and world-model benchmarks translates into reliable, per-player attribution once footage is this chaotic, turning footage into statistics through a chain of tasks from finding play boundaries to naming who did what. We evaluate four approaches (prompting and agentic reasoning over frontier vision-language models, classical computer vision with small trained specialists, self-supervised video world models, and manual annotation) at every stage, on 66 amateur matches with 46,648 human-labelled contacts, filmed under conditions no published benchmark uses. No single paradigm wins every stage, and static, single-frame computer vision is not competitive at any stage involving motion or identity. A prompted model segments matches well, yet a far smaller trained model beats it at spotting contacts for a fraction of the cost, and the sport's own rules recover rally outcomes the pixels cannot. Identity is where every automated approach struggles: a jersey number is a static fact temporal reasoning cannot recover if never visible, unlike sporting action, a repeated motor pattern a temporal model can exploit, which is why holistic reasoning improves event detection while identity stays unchanged. We close with where each approach earns its cost, and what transfers beyond volleyball to amateur sport.
Chinese Translation
视频理解通常在精心筛选的、单主体或专业拍摄的片段上进行基准测试,而模型在这些测试中的高分往往被直接解读为其具备足够部署鲁棒性的证据。业余团队运动是检验这一假设的一个有用且基本未被测试过的场景:仅在2024-25学年,美国就有超过八百万学生参加学校体育运动,其中几乎没有任何一场比赛使用多于一个固定机位拍摄,画面中挤着多个候选运动员,且没有摄像师或第二机位可供依赖。以排球作为测试案例,我们探究在通用视频与世界模型基准上的优异表现,能否转化为如此混乱拍摄条件下可靠的逐运动员归因能力——即通过从发现比赛边界到判定“谁做了什么”的一系列任务,将视频转化为统计数据。我们在66场业余比赛(含46,648个人工标注的触球事件)上,在拍摄条件远超已发表基准所覆盖范围的情况下,于每个阶段评估了四种方法:对前沿视觉-语言模型的提示与智能体推理、结合小型专门训练模型的经典计算机视觉、自监督视频世界模型,以及人工标注。没有任何单一范式在所有阶段都占优,而静态单帧计算机视觉在任何涉及运动或身份的阶段均缺乏竞争力。提示模型能很好地分割比赛,但一个小得多的训练模型以极低的成本在触球检测上超越它,而运动项目自身的规则则能恢复像素无法揭示的回合结果。身份识别是所有自动化方法都艰难应对的环节:球衣号码是一个静态事实,若从未在画面中出现过,时序推理无法将其恢复;与之不同,运动动作是时序模型可利用的重复性动作模式,这解释了为何整体推理能提升事件检测性能,而身份识别却原地踏步。最后,我们总结了各方法在哪些环节物有所值,以及哪些经验可以超越排球推广到其他业余体育项目。
cs.CV / 73 / 2609.28061

AstraLOD3: Zero-shot multimodal agentic reconstruction of LOD3 building models

AstraLOD3:LOD3建筑模型的零样本多模态智能体重建
Pantoja-Rosero, Bryan G.
Abstract
Automated LOD3 building modeling typically relies on purpose-built geometric or learning-based pipelines, limiting flexibility across heterogeneous buildings and input evidence conditions. This study investigates whether Astra, a general-purpose multimodal foundation model, can address these limitations through zero-shot reconstruction of LOD3 building models within an agentic framework under bounded autonomy. AstraLOD3 combines multi-view images, calibrated cameras, and a filtered sparse SfM point cloud with a natural-language reconstruction specification, while the Astra agent dynamically selects and executes computational procedures using Python and Blender. Across 35 runs, including 24 benchmark buildings, AstraLOD3 achieved a mean FRDS of 0.9647 and geometric agreement comparable to that of previous purpose-built methods. Controlled ablations further revealed the effects of reconstruction guidance, evidence modalities, model configuration, and run-to-run variability. The results demonstrate that structured LOD3 reconstruction can be formulated as a constrained agentic process rather than as a fixed pipeline. Future work will investigate adaptive refinement, user-guided correction, task-specific specialization, and damage-aware reconstruction.
Chinese Translation
自动化LOD3建筑建模通常依赖于专门构建的几何方法或基于学习的流水线,这限制了其在异构建筑和多样化输入证据条件下的灵活性。本研究探讨通用多模态基础模型Astra能否在受限自主性的智能体框架内,通过零样本方式实现LOD3建筑模型重建,从而克服上述局限。AstraLOD3将多视角图像、标定相机和经过筛选的稀疏SfM点云与自然语言重建规范相结合,同时Astra智能体利用Python和Blender动态选择并执行计算流程。在包括24个基准建筑在内的35次运行中,AstraLOD3取得了平均FRDS为0.9647的成绩,其几何一致性可与以往专门构建的方法相媲美。受控消融实验进一步揭示了重建引导、证据模态、模型配置以及运行间变异性的影响。研究结果表明,结构化的LOD3重建可以被表述为一个受约束的智能体过程,而非固定的流水线。未来工作将研究自适应精细化、用户引导修正、任务特定专业化以及损伤感知重建。
cs.CV / 74 / 2609.28077

TEEP-RCNN: Texture-Enhanced Edge-aware Perception for Steel Surface Defect Detection via Improved Convolutional Block Attention in Faster R-CNN

TEEP-RCNN:基于改进卷积块注意力Faster R-CNN的纹理增强边缘感知钢铁表面缺陷检测
Rajesh, Kirtan
Abstract
Steel surface defect detection is critical for automated industrial quality control but remains challenging due to subtle inter-class texture differences and pronounced class imbalance. We introduce TEEP-RCNN (Texture-Enhanced Edge-aware Perception Region-based CNN), a two-stage detector built on Faster R-CNN with a Feature Pyramid Network backbone and an improved Convolutional Block Attention Module (CBAM). Our CBAM adds dropout regularization in the channel attention MLP and batch normalization on the spatial attention branch, reducing co-adaptation and stabilizing gating logits. Training uses a differential learning rate protocol with cosine annealing warm-up, separating update rates for the pre-trained ResNet-101 backbone and the detection head. At inference, predictions are refined via Test-Time Augmentation fused with Weighted Box Fusion (WBF), improving localization stability on elongated and boundary-adjacent defects. On the NEU-DET benchmark across six defect categories, TEEP-RCNN achieves 73.3\% mAP@50 and 37.9\% mAP@50-95 in only 10 training epochs on a single GPU, competitive with YOLOv11m (76.2\% mAP@50, 100 epochs) while outperforming it on the rolled-in-scale category under the COCO metric. Per-class analysis shows the spatial attention branch is most effective on elongated texture defects such as patches and scratches, while crazing remains an open challenge across both paradigms due to its distributed non-local texture structure.
Chinese Translation
钢铁表面缺陷检测对工业自动化质量控制至关重要,但由于类间纹理差异细微以及显著的类别不平衡问题,该任务仍具有挑战性。我们提出了TEEP-RCNN(纹理增强边缘感知区域卷积神经网络),这是一个基于Faster R-CNN的两阶段检测器,采用特征金字塔网络(FPN)骨干和改进的卷积块注意力模块(CBAM)。我们的CBAM在通道注意力的多层感知机中加入了dropout正则化,并在空间注意力分支上加入批归一化,从而减少共适应并稳定门控逻辑值。训练采用带余弦退火预热的差分学习率策略,将预训练ResNet-101骨干网络与检测头的更新速率分离。在推理阶段,通过测试时增强(Test-Time Augmentation)与加权框融合(WBF)相结合对预测结果进行优化,提高了对细长及边界邻近缺陷的定位稳定性。在涵盖六种缺陷类别的NEU-DET基准数据集上,TEEP-RCNN在单个GPU上仅训练10个周期即达到73.3%的mAP@50和37.9%的mAP@50-95,与YOLOv11m(76.2% mAP@50,100个周期)具有竞争力,并在COCO指标下的氧化铁皮压入(rolled-in-scale)类别上超越了YOLOv11m。逐类别分析表明,空间注意力分支对斑块和划痕等细长纹理缺陷最为有效,而由于龟裂(crazing)具有分布式非局部纹理结构,它在两种方法范式中仍是尚未解决的难题。
cs.CV / 75 / 2609.28078

LiAM-SAM: Lifecycle-Aware Memory for Robust SAM2-Based MOT

LiAM-SAM:面向鲁棒SAM2多目标跟踪的生命周期感知记忆机制
Francisco, Grégoire, D'Amico, Alessandro, Costantini, Samuele, Francesca, Gianpiero, Garattoni, Lorenzo
Abstract
Segmentation-based multi-object tracking (MOT) with foundation video models such as SAM2 offers strong localization quality, yet remains fragile in crowded, real-world scenes. In detector-prompted SAM2 pipelines, failures typically arise at three stages of the object lifecycle: (i) erroneous or duplicate track initiation, (ii) memory drift during close interactions, and (iii) unreliable re-identification after long occlusions or re-entry. These errors corrupt object memory and accumulate over time, making long-horizon tracking unstable. In this paper, we reframe MOT as a lifecycle memory integrity problem. We present LiAM-SAM, a Lifecycle-Aware Memory (LiAM) framework with targeted mechanisms for each of the three failure modes. At track birth, to prevent faulty or duplicate initiations, we apply contrastive track initiation, which conditions each prompt on existing nearby tracked instances. To preserve memory integrity during strong interactions, we introduce motion- and geometry-grounded memory correction that resolves interaction confusions and suppresses drift. For reliable re-identification after disappearance, we maintain an adaptive context memory that promotes diverse and trustworthy references as long-term identity anchors. Finally, similarity aware spatial pruning optionally selects the memory tokens to retain at cross-attention time, improving efficiency with minimal accuracy loss. LiAM-SAM represents a modular, detector-agnostic, SAM2-based MOT system that achieves state-of-the-art HOTA and IDF1 on the evaluated benchmarks. In association-challenging environments, our ablations show that LiAM improves a detector+SAM2 baseline by +10.5 HOTA, +17.4 AssA, and reduces identity switches by 96%.
Chinese Translation
基于分割的多目标跟踪(MOT)借助SAM2等基础视频模型能够提供强大的定位质量,但在拥挤的真实场景中仍然表现脆弱。在检测器提示的SAM2流程中,失败通常出现在目标生命周期的三个阶段:(i)错误或重复的轨迹初始化;(ii)近距离交互过程中的记忆漂移;(iii)长时间遮挡或目标重新出现后不可靠的再识别。这些错误会污染目标记忆并随时间累积,使得长时间跟踪不稳定。本文将MOT重新构建为一个生命周期记忆完整性问题,提出LiAM-SAM——一种生命周期感知记忆(LiAM)框架,针对上述三种失败模式分别提供了专门的机制。在轨迹诞生阶段,为防止错误或重复的初始化,我们采用对比式轨迹初始化方法,使每个提示都以周围已有的被跟踪实例为条件。为了在强交互过程中保持记忆完整性,我们引入了基于运动和几何的记忆校正机制,以解决交互混淆并抑制漂移。为了在目标消失后实现可靠的再识别,我们维护一种自适应上下文记忆,将多样且可信的参考提升为长期身份锚点。最后,相似度感知的空间剪枝可选地在交叉注意力阶段筛选需要保留的记忆令牌,在几乎不损失精度的情况下提升效率。LiAM-SAM是一个模块化、检测器无关的基于SAM2的MOT系统,在所评估的基准上取得了最先进的HOTA和IDF1结果。在关联挑战性环境中,消融实验表明LiAM相较于检测器+SAM2基线提升了+10.5 HOTA、+17.4 AssA,并将身份切换减少了96%。
cs.CV / 76 / 2609.28083

ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming

ZoomDiff:一种面向双摄像头平滑变焦的高保真扩散模型
Zhang, Jiayi, Wu, Renlong, Ding, Yukang, Deng, Sibin, Zuo, Wangmeng
Abstract
Digital zoom transitions between dual cameras often exhibit conspicuous discontinuities in geometric structure and chromatic consistency, degrading the user experience. While recent dual-camera smooth zoom (DCSZ) methods attempt to mitigate this by fine-tuning frame interpolation (FI) models on DCSZ data, they struggle with the large cross-view disparities and complex geometric transformations. Considering that the generative prior of diffusion models is suitable for addressing this problem, we explore their application to DCSZ. However, naively applying existing diffusion-based FI models still yields low-fidelity transitions due to insufficient conditional guidance, high-frequency information loss during VAE encoding, as well as inadequate temporal consistency. To address this, we propose ZoomDiff, a high-fidelity diffusion model that leverages dual-camera inputs in both latent and pixel spaces for photo-realistic transitions. Specifically, we first strengthen dual-image conditional guidance during the multi-step denoising process to improve geometric consistency. Then we inject flow-aligned multi-scale features from the VAE encoder into the VAE decoder to recover high-frequency details, where flow-guided temporal consistency supervision are introduced to produce more smooth transitions. Extensive experiments on both synthetic and real-world datasets demonstrate that ZoomDiff outperforms state-of-the-art methods quantitatively and qualitatively. Project page: https://jiayi-hit.github.io/ZoomDiff.github.io/.
Chinese Translation
双摄像头之间的数字变焦过渡在几何结构和色彩一致性方面常表现出明显的不连续性,降低了用户体验。尽管近期的双摄像头平滑变焦(Dual-Camera Smooth Zoom, DCSZ)方法尝试通过在DCSZ数据上微调视频帧插值(Frame Interpolation, FI)模型来缓解该问题,但它们难以应对较大的跨视角差异和复杂的几何变换。考虑到扩散模型的生成先验适合解决这一问题,我们探索了其在DCSZ中的应用。然而,直接套用现有的基于扩散的FI模型,由于条件引导不足、VAE编码过程中高频信息丢失以及时序一致性欠佳,仍会产生低保真的过渡效果。为此,我们提出ZoomDiff,一种在潜空间和像素空间同时利用双摄像头输入以实现照片级逼真过渡的高保真扩散模型。具体而言,我们首先在多步去噪过程中强化双图像条件引导,以提升几何一致性;然后将来自VAE编码器的流对齐多尺度特征注入VAE解码器以恢复高频细节,并引入流引导的时序一致性监督,从而生成更加平滑的过渡效果。在合成数据集和真实数据集上的大量实验表明,ZoomDiff在定量和定性方面均优于当前最先进的方法。项目主页:https://jiayi-hit.github.io/ZoomDiff.github.io/。
cs.CV / 77 / 2609.28095

MotionSpec: Spectral Trajectory Supervision for Motion-Consistent Video Generation

MotionSpec:面向运动一致性视频生成的谱轨迹监督
Ni, Ziqi, Li, Rui, Jiang, Shiqi, Zhou, Wei
Abstract
Recent advances in text-to-video generation have enabled high-fidelity visual synthesis, yet realistic motion remains challenging. Generated videos may exhibit temporal discontinuities, inconsistent action progression, and structural distortions during complex movements. Even when individual frames appear realistic, the underlying motion may evolve in inconsistent or implausible ways. Standard generative objectives provide limited motion-specific supervision, leaving motion evolution insufficiently constrained. In this paper, we propose MotionSpec, a motion supervision framework centered on Spectral Trajectory Consistency (STC). STC constructs dense anchor-relative motion trajectories and transforms them into motion spectral volumes via a temporal Fourier transform. By aligning the spectral amplitude and phase of predicted and target trajectories, STC constrains both motion strength across temporal frequencies and the temporal organization of motion. To complement this trajectory-level supervision, we introduce Local Flow Consistency (LFC), which aligns consecutive-frame optical flow between predicted and target videos to stabilize local motion transitions. Experiments demonstrate that MotionSpec consistently improves motion consistency, temporal coherence, and plausibility while preserving visual fidelity.
Chinese Translation
文本到视频生成的最新进展已经能够实现高保真的视觉合成,但逼真的运动仍然具有挑战性。生成的视频可能出现时间不连续、动作进展不一致以及复杂运动过程中的结构畸变。即使单帧画面看起来逼真,其内在的运动也可能以不一致或不合理的方式演变。标准的生成目标提供的运动相关监督有限,使得运动演变缺乏充分约束。本文提出MotionSpec,一个以谱轨迹一致性(Spectral Trajectory Consistency, STC)为核心的运动监督框架。STC构建以锚点为参照的稠密运动轨迹,并通过时间傅里叶变换将其转换为运动谱体积。通过对齐预测轨迹与目标轨迹的谱幅度和相位,STC同时约束了各时间频率上的运动强度以及运动的时间组织方式。为补充这一轨迹层面的监督,我们进一步引入局部流一致性(Local Flow Consistency, LFC),通过对齐预测视频与目标视频之间相邻帧的光流来稳定局部运动过渡。实验表明,MotionSpec在保持视觉保真度的同时,能够持续提升运动一致性、时间连贯性和合理性。
cs.CV / 78 / 2609.28099

Visual Tripwires: Anticipating Failure in Deep Vision Systems

视觉绊线:预测深度视觉系统的失效
Harit, Anoushka, Zuberi, Rehan, Prew, William, Markowetz, Florian
Abstract
Deep vision systems remain vulnerable to corruption, occlusion, and distribution shift despite strong benchmark performance. Existing reliability methods typically evaluate uncertainty at individual time steps and do not explicitly model how a system progresses toward failure. We introduce Visual Tripwires, a predictive reliability framework that uses temporal instability in model behaviour to anticipate impending failure. Our central hypothesis is that predictive degradation develops progressively through measurable changes in latent representations, prediction trajectories, and attention structure. Visual Tripwires captures these changes using representation drift, prediction oscillation, trajectory curvature, and attention entropy. A lightweight tripwire predictor aggregates these signals over a temporal window to estimate the probability of failure within a future prediction horizon. Experiments across multiple datasets, architectures, and progressive perturbation settings show that the proposed instability signals emerge before predictive degradation and provide earlier and more accurate failure warnings than conventional uncertainty estimation methods. These results demonstrate that temporal instability contains useful information about future model reliability and provides a practical basis for early warning in deep vision systems.
Chinese Translation
尽管在基准测试中表现优异,深度视觉系统在面对图像损坏、遮挡和分布偏移时仍然脆弱。现有的可靠性方法通常仅在单个时间步上评估不确定性,并未显式建模系统如何逐步走向失效。我们提出了 Visual Tripwires(视觉绊线),一个利用模型行为的时间不稳定性来预测即将发生的失效的预测性可靠性框架。我们的核心假设是:预测性能的退化会通过潜在表示、预测轨迹和注意力结构中可测量的变化而逐步显现。Visual Tripwires 通过表示漂移、预测振荡、轨迹曲率和注意力熵来捕捉这些变化。一个轻量级的绊线预测器在时间窗口内聚合这些信号,以估计在未来预测范围内发生失效的概率。在多个数据集、架构和渐进式扰动设置下的实验表明,所提出的不稳定性信号在预测性能退化之前即已出现,并且相比传统的不确定性估计方法能够提供更早、更准确的失效预警。这些结果表明,时间不稳定性包含有关未来模型可靠性的有用信息,为深度视觉系统的早期预警提供了实用的基础。
cs.CV / 79 / 2609.28110

Field-of-View Extension in Dental Cone-Beam CT via Implicit Neural Representations and Diffusion Model-Based Refinement

基于隐式神经表示与扩散模型精修的牙科锥束CT视野扩展
Schaub, Susanne, Bieder, Florentin, Oliveira, Matheus L., Wang, Yulan, Sodnom-ish, Buyanbileg, Dagassan-Berndt, Dorothea, Bornstein, Michael M., Cattin, Philippe C.
Abstract
Dental cone-beam computed tomography (CBCT) systems often employ detector configurations that provide a truncated field of view (FOV) that only captures a small part of the patient's anatomy. In this work, we aim to reconstruct an extended FOV using projections of truncated FOV scans. To this end, we propose a three-stage framework that consists of (1) an implicit neural representation (INR) for estimating missing parts of the truncated projection data, (2) an iterative reconstruction for generating a secondary volumetric image with improved anatomical consistency and (3) a fast diffusion model for image enhancement. The proposed approach combines the strengths of continuous representations, physics-based reconstruction and generative refinement within a unified pipeline for truncated CBCT imaging. Experimental results demonstrate that the method effectively reduces truncation artifacts, improves the reconstruction of structures extending beyond the original FOV and produces images with enhanced quality. Our code is publicly available at https://github.com/SusanneSchaub/CBCT-FOV-Extension.
Chinese Translation
牙科锥束计算机断层扫描(CBCT)系统通常采用的探测器配置会提供截断的视野(FOV),仅能捕获患者解剖结构的一小部分。在本工作中,我们旨在利用截断视野扫描的投影数据重建扩展的视野。为此,我们提出一个三阶段框架,包括:(1) 利用隐式神经表示(INR)估计截断投影数据的缺失部分;(2) 通过迭代重建生成具有更好解剖一致性的次级体积图像;(3) 使用快速扩散模型进行图像增强。所提方法将连续表示、基于物理的重建和生成式精修的优势结合在一个统一的截断CBCT成像流程中。实验结果表明,该方法有效减少了截断伪影,改善了对超出原始视野结构的重建,并生成了质量更高的图像。我们的代码已在 https://github.com/SusanneSchaub/CBCT-FOV-Extension 公开。
cs.CV / 80 / 2609.28154

A comparative assessment of global building and settlement datasets across geographic and settlement contexts

全球建筑与聚落数据集在不同地理和聚落环境下的比较评估
Balogun, Rufai Omowunmi, Gevaert, Caroline Margaux, Riom, Capucine, Mirindi, Derrick, Opdyke, Aaron, Alemohammad, Hamed, Chrzanowski, Pierre, Anderson, Edward Charles
Abstract
Global building and settlement datasets increasingly support population mapping, exposure assessment, urban monitoring, and other analyses of the built environment, yet comparative evidence remains fragmented across products, geographic regions, reference datasets, spatial scales, and evaluation methods. We benchmark seven global or near-global products, including Overture Maps, Global Building Atlas, 3D-GloBFP, Google Open Buildings 2.5D Temporal (OBT), Microsoft TEMPO, GHSL, and WSF Tracker, against harmonized reference footprints across 135 study areas. The evaluation combines complementary measures of detection, geometric agreement, and aggregate quantity accuracy, together with stratified analyses of settlement characteristics and diagnostic experiments on error size and temporal alignment. Overture achieved the highest median city-level vector F1 (0.786). Raster rankings were resolution-dependent: OBT achieved the highest median F1 at 10m (0.642), whereas WSF Tracker led at 100m (0.862). However, WSF Tracker substantially overestimated built-up area, emphasizing that when using raster products, it is important for the user to understand whether the raster identifies only buildings or includes additional impervious surfaces. Raster accuracy increased consistently with building density (Spearman \r{ho} = 0.58-0.75), while small candidate buildings were disproportionately associated with false positives in the vector products. Temporally aligning WSF Tracker with reference imagery increased mean F1 by 0.060 (median +0.037), indicating that the reported accuracies are conservative in rapidly growing areas. The study establishes a reproducible benchmark for comparing heterogeneous global urban and settlement layer datasets across geographic and settlement contexts.
Chinese Translation
全球建筑与聚落数据集日益广泛地支持人口制图、暴露度评估、城市监测等建成环境分析,但跨产品、地理区域、参考数据集、空间尺度和评估方法的比较证据仍然零散。我们将七个全球或近全球产品(包括 Overture Maps、Global Building Atlas、3D-GloBFP、Google Open Buildings 2.5D Temporal (OBT)、Microsoft TEMPO、GHSL 和 WSF Tracker)与跨 135 个研究区域的统一参考足迹进行基准对比。评估结合了检测性能、几何一致性和总量精度的互补性指标,并对聚落特征进行分层分析,同时针对误差规模和时间对齐开展诊断实验。Overture 在城市级矢量 F1 中位数上最高(0.786)。栅格产品的排名依赖于分辨率:在 10m 分辨率下,OBT 的 F1 中位数最高(0.642),而在 100m 分辨率下 WSF Tracker 领先(0.862)。然而,WSF Tracker 显著高估了建成区面积,这表明在使用栅格产品时,用户需要了解该栅格仅识别建筑物还是包含其他不透水面。栅格精度随建筑密度提高而一致提升(Spearman ρ = 0.58–0.75),而在矢量产品中,小型候选建筑与假阳性存在不成比例的关联。将 WSF Tracker 与参考影像进行时间对齐后,平均 F1 提高了 0.060(中位数 +0.037),表明在快速增长地区,所报告的精度估计偏保守。本研究建立了一个可复现的基准,用于在不同地理和聚落环境下比较异构的全球城市与聚落图层数据集。
cs.CV / 81 / 2609.28159

Depth-Guided Contrastive Learning for 2D Representations with 3D Spatial Awareness

具有3D空间感知的2D表示的深度引导对比学习
Zeng, Liang, Vergauwen, Maarten
Abstract
Standard contrastive learning frameworks are mainly designed from a semantic perspective, yet learning 2D visual representations that preserve 3D spatial structure is also important for scene understanding. In this work, we propose Depth-Guided Contrastive Learning (DGCL), a simple auxiliary objective that injects 3D spatial awareness into 2D contrastive representation learning. Our key idea is to use depth to convert local 3D proximity into contrastive similarity: pixels that are closer in 3D space are encouraged to have more similar representations than pixels that are farther apart. Instead of relying on absolute depth values, DGCL formulates supervision through relative 3D distance comparisons among randomly sampled pixels, making the objective invariant to depth scale, efficient to compute, and easy to integrate into existing contrastive frameworks. Experiments across different datasets and models show that DGCL consistently improves 2D representation learning and benefits semantic downstream tasks by stronger spatial and geometric understanding. The code is available on https://github.com/LeungTsang/DGCL.
Chinese Translation
标准对比学习框架主要从语义角度进行设计,然而学习能够保留3D空间结构的2D视觉表示对场景理解同样十分重要。在本工作中,我们提出了深度引导对比学习(Depth-Guided Contrastive Learning,DGCL),这是一种简单的辅助目标,可将3D空间感知注入2D对比表示学习中。我们的核心思想是利用深度将局部3D邻近性转化为对比相似性:鼓励3D空间中距离更近的像素比距离更远的像素具有更相似的表示。DGCL不依赖于绝对深度值,而是通过随机采样像素之间的相对3D距离比较来构建监督信号,使该目标对深度尺度具有不变性、计算高效,且易于集成到现有对比学习框架中。在不同数据集和模型上的实验表明,DGCL能够持续改进2D表示学习,并通过更强的空间与几何理解能力为语义下游任务带来收益。代码已在 https://github.com/LeungTsang/DGCL 开源。
cs.CV / 82 / 2609.28183

From ECG Signals to Representative-Morphology Heatmaps for Biometric Recognition

从心电信号到代表性形态热力图用于生物特征识别
Angelakis, Athanasios, Gomez-Barrero, Marta
Abstract
Electrocardiography (ECG) contains subject-specific morphology that supports biometric recognition, yet image-based performance depends on how the waveform is rendered. We introduce representative-morphology heatmaps, a deterministic ECG-to-image representation adapted from ECGXtractor. Within each block of ten aligned beats, the five beats closest to the block mean are averaged into a 400 by L matrix and rendered either as a conventional trace or as a dense cardiac-time-by-lead heatmap. Since both representations contain identical physiological samples, their comparison isolates the effect of rendering. We evaluate verification and closed-set identification on PTB, ECG-ID, and MIMIC-IV-ECG-DEMO. Five compact models, including ZACH-ViT, are trained from scratch, while six ImageNet-pretrained CNN and transformer backbones assess model scale and visual transfer. Heatmaps improve both FNMR operating points and both identification ranks in all 15 compact model-dataset comparisons, while EER improves in 14. Across the matched experiments, EER decreases by 9.59 percentage points and Rank-1 increases by 24.69 points on average. ConvNeXt-Tiny reaches 2.43% EER on PTB and 5.79% on ECG-ID, whereas DeiT-Base reaches 14.92% on MIMIC-DEMO. ImageNet initialization clearly benefits the two multilead datasets but has a mixed effect on ECG-ID, and performance does not increase monotonically with model size. The best heatmap systems approach the strongest signal-domain EER on PTB and ECG-ID, while DeiT-Base provides the strongest evaluated performance on MIMIC-DEMO. Lead-channel ablation further shows that useful channel combinations depend on the cohort and biometric task. Overall, representative-morphology heatmaps provide an effective image representation for ECG verification and identification.
Chinese Translation
心电图(ECG)包含个体特异性的形态特征,可支持生物特征识别,然而基于图像的方法的性能取决于心电波形的渲染方式。我们提出了代表性形态热力图(representative-morphology heatmaps),这是一种基于 ECGXtractor 改造的确定性 ECG 转图像表示方法。在由十个对齐心拍组成的每个区块中,将最接近区块均值的五个心拍平均为 400×L 的矩阵,并以传统波形轨迹或以心动周期时间×导联的稠密热力图形式进行渲染。由于两种表示包含完全相同的生理采样数据,二者的对比可以单独考察渲染方式的影响。我们在 PTB、ECG-ID 和 MIMIC-IV-ECG-DEMO 数据集上评估了身份验证与闭集身份识别任务。包括 ZACH-ViT 在内的五个紧凑模型均从头训练,同时使用六个 ImageNet 预训练的 CNN 与 Transformer 骨干网络来评估模型规模与视觉迁移的影响。在全部 15 组紧凑模型-数据集对比中,热力图均改善了两类 FNMR 工作点和两项识别排名指标,其中 14 组的 EER 得到改善。在匹配实验中,EER 平均下降 9.59 个百分点,Rank-1 平均提升 24.69 个百分点。ConvNeXt-Tiny 在 PTB 上达到 2.43% 的 EER,在 ECG-ID 上达到 5.79%,而 DeiT-Base 在 MIMIC-DEMO 上达到 14.92%。ImageNet 初始化对两个多导联数据集有明显收益,但对 ECG-ID 的效果不一,且性能并非随模型规模单调提升。最优的热力图系统在 PTB 和 ECG-ID 上接近最强的信号域 EER,而 DeiT-Base 在 MIMIC-DEMO 上提供了评估中最优的性能。导联通道消融实验进一步表明,有用的通道组合依赖于研究队列与生物特征识别任务。总体而言,代表性形态热力图为 ECG 身份验证与识别提供了一种有效的图像表示方法。
cs.CV / 83 / 2609.28187

Two Global Crops Suffice: Locating Semantic Emergence in DINO-Style Self-Supervised Learning

两次全局裁剪已足够:定位DINO式自监督学习中的语义涌现
Sunagad, Basavaraj, Jesslen, Artur, Kortylewski, Adam
Abstract
Self-supervised vision transformers trained with DINO-style objectives exhibit striking emergent semantic representation quality across visual tasks, yet the mechanisms underlying this behavior remain unclear. We present a systematic empirical dissection of the DINO family and show that semantic representations arise primarily from enforcing consistency between geometrically distinct global views of the same image instance. This instance-specific global alignment acts as the semantic anchor of DINO-style learning. Across controlled retraining experiments evaluated on semantic correspondence and a diverse suite of 2D and 3D downstream tasks, we find that patch-level masking objectives enhance semantics only when trained jointly with this global alignment, indicating that the iBOT objective refines and densifies existing semantic structure rather than creating it independently. In contrast, local-to-global view alignment does not substantially improve semantic qualities at fixed compute beyond a purely global alignment. Beyond training design, we revisit how semantic representation quality should be evaluated: while classification accuracy is the standard validation score, semantic correspondence provides a complementary axis that more reliably predicts downstream task performance. Together, these findings provide a functional decomposition of DINO-style learning and represent an important step toward understanding how semantic representations emerge in self-supervised vision models.
Chinese Translation
采用DINO式目标训练的自监督视觉Transformer在各种视觉任务中表现出显著的涌现语义表示质量,然而这种行为背后的机制仍不清楚。我们对DINO系列模型进行了系统的实证剖析,结果表明语义表示主要源于对同一图像实例在几何上不同的全局视图之间施加一致性约束。这种实例特定的全局对齐充当了DINO式学习的语义锚点。在语义对应任务以及一套多样的2D和3D下游任务上评估的受控重训练实验中,我们发现块级掩码目标只有在与该全局对齐联合训练时才能增强语义,这表明iBOT目标是对已有语义结构进行细化和稠密化,而非独立地创造语义结构。相比之下,在固定计算量下,局部到全局的视图对齐并不比纯全局对齐带来实质性的语义质量提升。除训练设计之外,我们还重新审视了应如何评估语义表示质量:虽然分类准确率是标准的验证指标,但语义对应任务提供了一个补充维度,能更可靠地预测下游任务性能。这些发现共同提供了对DINO式学习的功能分解,是理解自监督视觉模型中语义表示如何涌现的重要一步。
cs.CV / 84 / 2609.28192

From Change Captions to Change Detection: Semantic-Appearance Agreement Framework for Remote Sensing Change Detection

从变化描述到变化检测:面向遥感变化检测的语义-外观一致性框架
Qian, Yuan, Ma, Jie
Abstract
Remote sensing change detection (RSCD) is essential for monitoring land-cover changes and urban development. However, most methods demand pixel-level change masks, which are costly and time-consuming to annotate. Weakly supervised methods reduce this cost by using image-level change labels. Yet these labels indicate only whether a change occurs, leaving models to recover the location of the change and semantic meaning through additional and complex mechanisms. This missing information can be supplied directly by change captions, which describe what changes, what it becomes, and where it occurs. Therefore, we introduce change-caption-guided RSCD, using change captions as the sole task-specific supervision to learn change masks without manually annotated change masks. Our framework has two components: a caption-driven generation pipeline that produces bi-temporal remote sensing image pairs at scale with controlled changes matching each caption, and a change detector guided by the caption's transition semantics. The detector uses our Semantic-Appearance Agreement Framework (SAAF) to combine caption-grounded semantic responses with RGB differences for change localization, while text conditioning guides dense prediction. Experiments on our newly constructed Flair-RSGen dataset and WHU-CDC show that SAAF outperforms the closest reproduced limited-supervision baselines in macro-averaged IoU and F1 under the evaluated protocols. Code is publicly available at https://github.com/qianyuancs/SAAF.
Chinese Translation
遥感变化检测(RSCD)对于监测土地覆盖变化和城市发展至关重要。然而,大多数方法需要像素级的变化掩膜,其标注成本高且耗时。弱监督方法通过使用图像级变化标签降低了这一成本,但这类标签仅能指示是否发生变化,模型仍需借助额外且复杂的机制来恢复变化的位置及其语义信息。这些缺失的信息可以直接由变化描述(change captions)提供,因为变化描述能够说明发生了什么变化、变成了什么以及变化发生在哪里。因此,我们提出了变化描述引导的遥感变化检测任务,将变化描述作为唯一的任务特定监督信号,在无需人工标注变化掩膜的情况下学习变化掩膜。我们的框架包含两个组成部分:一是描述驱动的生成流水线,能够大规模生成双时相遥感图像对,且其中的变化与每条描述相匹配;二是由描述的转换语义引导的变化检测器。该检测器采用我们提出的语义-外观一致性框架(Semantic-Appearance Agreement Framework, SAAF),将基于描述的语义响应与RGB差异相结合以实现变化定位,同时利用文本条件引导稠密预测。在新构建的Flair-RSGen数据集和WHU-CDC数据集上的实验表明,在所评估的协议下,SAAF在宏平均IoU和F1指标上优于最接近的复现有限监督基线方法。代码已公开发布于 https://github.com/qianyuancs/SAAF。
cs.CV / 85 / 2609.28222

From Alignment to Fusion in 3D Vision-Language

从对齐到融合:三维视觉-语言研究
Qiu, Xueqi, Miao, Xingyu, Deng, Jingjing, Duan, Haoran, Long, Yang, Shao, Ling
Abstract
Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to language-guided reasoning. Existing methods often process point clouds, voxel grids, and multi-view images independently; directly combining these heterogeneous representations may leave substantial feature discrepancy unresolved, while subsequent unconstrained adaptation may distort their internal geometry. We propose an align-then-fuse framework that first applies triple pairwise cosine alignment to establish segment-level correspondence across the three representations and then retrieves task-conditioned features with a prompt-guided query decoder. Before fusion, representation-specific query features are transformed by learnable mappings constrained to the special orthogonal group. These mappings preserve inner products and Euclidean distances within each representation, permitting controlled representation-specific re-parameterisation without arbitrarily distorting its internal geometry. The transformed features are subsequently combined through Adaptive Fusion under downstream task supervision. Experiments cover eight datasets for instance segmentation, visual grounding, question answering, and dense captioning. Compared with PQ3D, the model improves average precision by 3.2 points on ScanNet200 and grounding accuracy by 2.9, 10.6, 4.6, and 4.1 points on ScanRefer, Nr3D, Sr3D, and Multi3DRefer, respectively, while also improving performance on ScanQA, SQA3D, and Scan2Cap. Ablations further support the complementary roles of alignment and orthogonal re-parameterisation and the effectiveness of Adaptive Fusion.
Chinese Translation
统一的三维视觉-语言系统需要在支持从实例分割到语言引导推理等多种任务的同时,融合互补的几何、尺度和外观线索。现有方法通常独立处理点云、体素网格和多视角图像;直接组合这些异构表示可能遗留大量未解决的特征差异,而后续无约束的自适应调整则可能扭曲其内部几何结构。我们提出一种“先对齐后融合”(align-then-fuse)框架:首先应用三重成对余弦对齐(triple pairwise cosine alignment)在三种表示之间建立片段级对应关系,然后通过提示引导的查询解码器(prompt-guided query decoder)检索任务条件化的特征。在融合之前,各表示特有的查询特征经过约束于特殊正交群(special orthogonal group)的可学习映射进行变换。这些映射保持每个表示内部的内积和欧氏距离不变,从而在实现受控的表示特定重参数化的同时,不会任意扭曲其内部几何结构。变换后的特征随后在下游任务监督下通过自适应融合(Adaptive Fusion)进行组合。实验涵盖实例分割、视觉定位、问答和密集描述八个数据集。与PQ3D相比,本模型在ScanNet200上将平均精度提升3.2个点,在ScanRefer、Nr3D、Sr3D和Multi3DRefer上的定位精度分别提升2.9、10.6、4.6和4.1个点,同时在ScanQA、SQA3D和Scan2Cap上也取得了性能提升。消融实验进一步验证了对齐与正交重参数化的互补作用以及自适应融合的有效性。
cs.CV / 86 / 2609.28230

A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing

面向遥感中旋转目标视觉定位的统一框架与数据集
Ding, Zeyu, Zhou, Yong, Zhao, Jiaqi, Du, Wen-Liang, Li, Xixi, Zhu, Hancheng, Yao, Rui, Saddik, Abdulmotaleb El
Abstract
Visual grounding in remote sensing images aims to locate objects described by referring expressions. Most existing methods predict horizontal bounding boxes, which are often inaccurate for objects with arbitrary orientations. To address this limitation, we introduce O$^2$-VG, a family of models for oriented object visual grounding with three complementary designs. Specifically, O$^2$-VG-Trans is a cross-modality transformer for oriented object visual grounding. It establishes a strong discriminative foundation for the model family. Building upon it, O$^2$-VG-Uni predicts universal oriented proposals for possible foreground objects without specific text prompts. It also supports object retrieval through cached proposal embeddings. Using these universal oriented proposals as input prompts, O$^2$-VG-VLM is an autoregressive vision-language model. It generates oriented box token blocks in parallel through multi-token prediction. In addition, we construct DIOR-R-RSVG, a dataset for oriented object visual grounding in remote sensing images. It provides image, expression, and oriented box triplets for training and evaluation. Together, the O$^2$-VG family provides a flexible framework that spans discriminative transformers and generative vision-language models. It achieves superior performance across multiple benchmarks. Code is available at https://github.com/wokaikaixinxin/ai4rs.
Chinese Translation
遥感图像中的视觉定位旨在根据指代表达定位所描述的目标。大多数现有方法预测水平边界框,这对于具有任意方向的目标往往不够准确。为解决这一局限,我们提出了 O$^2$-VG,一个用于旋转目标视觉定位的模型家族,包含三种互补的设计。具体而言,O$^2$-VG-Trans 是一个用于旋转目标视觉定位的跨模态 Transformer,为该模型家族奠定了强大的判别式基础。在此基础上,O$^2$-VG-Uni 无需特定文本提示,即可为可能的前景目标预测通用旋转候选框,并通过缓存候选框嵌入支持目标检索。以这些通用旋转候选框作为输入提示,O$^2$-VG-VLM 是一个自回归视觉-语言模型,通过多词元预测并行生成旋转框词元块。此外,我们构建了 DIOR-R-RSVG,一个面向遥感图像旋转目标视觉定位的数据集,提供了图像、表达与旋转框三元组用于训练和评估。综上,O$^2$-VG 家族提供了一个灵活的框架,涵盖了判别式 Transformer 与生成式视觉-语言模型,并在多个基准测试中取得了优异的性能。代码已发布于 https://github.com/wokaikaixinxin/ai4rs。
cs.CV / 87 / 2609.28231

Do Center Biases Propagate? Robustness of Pathology Foundation Models in Whole-Slide Image Classification

中心偏置会传播吗?病理学基础模型在全切片图像分类中的鲁棒性
Carretero, Ilán, Meseguer, Pablo, del Amor, Rocío, Naranjo, Valery
Abstract
Pathology foundation models (PFMs) have transformed computational pathology through powerful representation learning from histopathological images. PFMs provide rich, discriminative representations for whole slide image (WSI) analysis, enabling tasks such as slide-level classification under multiple instance learning (MIL). However, these representations may also encode non-biological signals associated with acquisition centers, potentially introducing spurious shortcuts into downstream predictions. In this work, we evaluate center-associated robustness in WSI classification using a controlled training setting with increasing class-center correlations quantified by Cram\'er's V. We benchmark six PFMs across four datasets and two MIL aggregators, while evaluating ComBat as a robustification strategy. We further introduce the Area Under the Cram\'er's V Curve (AUCC) to jointly capture absolute classification performance and its degradation as spurious correlation increases. Results show that center-related information encoded by PFMs propagates to WSI-level predictions, with robustness depending on both the PFM representation and MIL aggregation strategy. Additionally, ComBat harmonization does not provide consistent robustness gains across datasets.
Chinese Translation
病理学基础模型(Pathology Foundation Models, PFMs)通过从组织病理学图像中进行强大的表示学习,变革了计算病理学。PFMs 为全切片图像(Whole Slide Image, WSI)分析提供了丰富的判别性表示,使得在多示例学习(Multiple Instance Learning, MIL)框架下的切片级分类等任务成为可能。然而,这些表示也可能编码与采集中心相关的非生物学信号,从而可能在下游预测中引入虚假捷径。在本工作中,我们通过一个可控的训练设置评估了 WSI 分类的中心相关鲁棒性,其中类别与中心的相关性以 Cramér's V 量化并逐渐增强。我们在四个数据集和两种 MIL 聚合器上对六种 PFMs 进行了基准测试,并将 ComBat 作为一种鲁棒性增强策略进行评估。我们进一步提出了 Cramér's V 曲线下面积(Area Under the Cramér's V Curve, AUCC)指标,以同时刻画绝对分类性能及其随虚假相关性增强而产生的性能退化。结果表明,PFMs 所编码的中心相关信息会传播到 WSI 级别的预测中,其鲁棒性同时取决于 PFM 表示和 MIL 聚合策略。此外,ComBat 校正并不能在所有数据集上提供一致的鲁棒性提升。
cs.CV / 88 / 2609.28235

Diff-RF: Mutually Reinforced Image Registration and Fusion via Degradation-Aware Learning

Diff-RF:基于退化感知学习的图像配准与融合相互增强方法
Yi, Xunpeng, Du, Zaixi, Yan, Qinglong, Zhang, Yibing, Xu, Han, Ma, Jiayi
Abstract
Image registration and fusion aim to establish spatial correspondences from misaligned multi-modal source images, and integrate complementary information. However, in real-world imaging scenarios, source images are often affected by complex and diverse degradations, such as low illumination, noise, etc., which severely hinder the effectiveness of registration and fusion. To address this issue, we propose a mutually reinforced image registration and fusion diffusion framework via degradation-aware learning, termed Diff-RF. It explores the intrinsic coupling between registration-fusion and information restoration in the degradation conditions, enabling high-quality fusion of unregistered images under complex degradation conditions. First, the intra-modal restoration module is designed to alleviate modality-specific degradations by leveraging information within each modality, thereby providing more reliable structural representations for registration and facilitating subsequent cross-modal fusion. Second, we develop a cross-modal diffusion registration and fusion module that establishes bidirectional interaction between registration and fusion. By integrating fusion-derived visual cues and correspondence-based geometric conditions into the diffusion process, the proposed framework progressively refines spatial alignment and exploits cross-modal complementary information to achieve collaborative enhancement. Rather than treating them as independent components, degradation-aware information restoration and the collaborative optimization of registration and fusion are tightly coupled, achieving overall performance improvements. Extensive experiments on multiple extended datasets demonstrate that Diff-RF achieves superior registration accuracy and fusion quality under various degraded scenarios, exhibiting strong robustness and generalization ability.
Chinese Translation
图像配准与融合旨在从未对齐的多模态源图像中建立空间对应关系,并整合互补信息。然而,在真实成像场景中,源图像往往受到复杂多样的退化因素影响,如低照度、噪声等,严重阻碍了配准与融合的有效性。为解决这一问题,我们提出了一种基于退化感知学习的配准与融合相互增强的扩散框架,称为 Diff-RF。该框架探索了退化条件下配准-融合与信息恢复之间的内在耦合关系,实现了复杂退化条件下未配准图像的高质量融合。首先,设计了模态内恢复模块,通过利用各模态内部信息来缓解模态特异性退化,从而为配准提供更可靠的结构表征,并促进后续的跨模态融合。其次,我们开发了跨模态扩散配准与融合模块,在配准与融合之间建立双向交互。通过将融合产生的视觉线索和基于对应关系的几何条件引入扩散过程,所提出的框架逐步优化空间对齐,并利用跨模态互补信息实现协同增强。退化感知的信息恢复与配准-融合的协同优化并非被视为相互独立的组件,而是紧密耦合,从而实现整体性能的提升。在多个扩展数据集上的大量实验表明,Diff-RF 在各种退化场景下均取得了更优的配准精度和融合质量,展现出强大的鲁棒性和泛化能力。
cs.CV / 89 / 2609.28236

EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

EmbodiedMemory-Bench:面向长时程具身任务的具身记忆基准测试
Liang, Lizhou, Zhong, Xinyu, Pan, Miao, Zhou, Xiaohe, Liu, Xuanyu, Li, Qinfeng, Li, Peng, Chen, Jintao, Zhang, Xuhong, Zhang, Wenqi
Abstract
Long-horizon embodied interaction requires agents to retain and continually update information about the environment as they observe, act, and encounter change. Yet current agents struggle to maintain such memory reliably. Our analysis traces this limitation to four key deficiencies: weak fine-grained visual memory, unreliable dynamic world-state tracking, failing to record world state revealed by interaction outcomes, and limited generalization from prior experience. However, existing benchmarks do not directly assess these memory capabilities during long-horizon embodied interaction. To address this gap, we introduce EmbodiedMemory-Bench (EMem-Bench), comprising 2,554 interactive episodes across four task families. EMem-Bench requires agents to build and update memory from interaction history, then use it to complete a later task by acting in the environment. We further present Embodied-Memorizer (EMem), an external memory system that organizes embodied experience into spatial, event, and scene memories. We also train EMem-8B, an 8B policy that manages and uses these memories. We evaluate a diverse range of open-source and proprietary MLLMs and representative multimodal memory systems. Results show that current models remain weak and uneven across the four challenges. Under matched backbones, EMem achieves the best overall performance among the evaluated memory systems and improves both open-source and proprietary models, while EMem-8B further improves over its backbone. Project page: https://zju-omniai.github.io/EmbodiedMemoryBench/
Chinese Translation
长时程具身交互要求智能体在观察环境、执行动作并遭遇变化的过程中,保持并持续更新关于环境的信息。然而,当前智能体难以可靠地维持此类记忆。我们的分析将这一局限归结为四个关键缺陷:细粒度视觉记忆能力弱、动态世界状态跟踪不可靠、未能记录由交互结果揭示的世界状态,以及从先前经验进行泛化的能力有限。然而,现有基准并未在长时程具身交互过程中直接评估这些记忆能力。为填补这一空白,我们提出了 EmbodiedMemory-Bench(EMem-Bench),包含跨四个任务族共 2,554 个交互回合。EMem-Bench 要求智能体从交互历史中构建并更新记忆,随后利用这些记忆通过在环境中执行动作来完成后续任务。我们进一步提出了 Embodied-Memorizer(EMem),一种将具身经验组织为空间记忆、事件记忆和场景记忆的外部记忆系统。我们还训练了 EMem-8B,一个管理并使用这些记忆的 8B 策略模型。我们评估了多种开源和专有多模态大语言模型(MLLM)以及具有代表性的多模态记忆系统。结果显示,当前模型在四项挑战上仍然表现薄弱且不均衡。在使用相同骨干网络(backbone)的条件下,EMem 在所评估的记忆系统中取得最佳整体性能,并能同时提升开源与专有模型的表现,而 EMem-8B 进一步超越了其骨干模型。项目页面:https://zju-omniai.github.io/EmbodiedMemoryBench/
cs.CV / 90 / 2609.28239

ODPure: Backdoor Purification for Object Detection via Ensemble Corruption Consensus

ODPure:基于集成扰动共识的目标检测后门净化方法
Zeng, Li, Duan, Mingcheng, Fan, Longfei, Zhang, Hangtao, Wang, Xianlong, Li, Yanchun, Wen, Xia, Zhang, Leo Yu
Abstract
With the development of applications like autonomous driving, object detection has gained significant attention, while also highlighting critical vulnerabilities like backdoor attacks that severely compromise model integrity. Specifically, such attacks involve altering the categories of objects (i.e., object misclassification), removing bounding boxes (i.e., object disappearance), or generating bounding box proposals for non-existent objects (i.e., object generation) when a predefined trigger is present in the input. Although backdoor defenses for image classification are well-established, the research for object detection remains comparatively underexplored. Existing defenses address these threats by scanning outputs or models for potential backdoors but require discarding either malicious data or models. This remedy fails to enable a continuous and accurate perceptual stream for the object detection pipeline. To address such limitations, we propose ODPure, a novel input-stage black-box defense for object detection, which is based on input purification that ensures stable perception flows. Tailored to the dense prediction nature of object detectors, our Corruption-Reconstruction-Selection (CRS) paradigm operates by neutralizing triggers through a diverse portfolio of corruptions to generate a massive pool of redundant proposals, then recovering fine-grained structural cues via generative priors, and finally employing voting to reach a consensus on the resulting detections. Comprehensive experiments demonstrate that our method provides robust defense against diverse backdoor attacks and trigger types while preserving baseline accuracy. Our code is available at https://github.com/Alex66366/ODPure.
Chinese Translation
随着自动驾驶等应用的发展,目标检测受到了广泛关注,同时也暴露出诸如后门攻击等严重威胁模型完整性的关键漏洞。具体而言,这类攻击在输入中存在预定义触发器时,会改变目标的类别(即目标误分类)、移除边界框(即目标消失),或为不存在的目标生成边界框(即目标生成)。尽管针对图像分类的后门防御已经相当成熟,但针对目标检测的相关研究仍相对不足。现有防御方法通过扫描输出或模型来检测潜在后门,但需要丢弃恶意数据或模型。这种补救措施无法为目标检测流程提供持续且准确的感知流。为解决上述局限,我们提出了ODPure,一种新颖的目标检测输入阶段黑盒防御方法,其基于输入净化机制,可确保感知流的稳定性。针对目标检测器的密集预测特性,我们提出的扰动-重建-选择(Corruption-Reconstruction-Selection, CRS)范式的工作流程为:首先通过多样化的扰动组合中和触发器,生成大量冗余候选框;然后利用生成先验恢复细粒度结构线索;最后通过投票机制对检测结果达成共识。大量实验表明,我们的方法能够在保持基线精度的同时,有效抵御多种后门攻击和触发器类型。我们的代码已发布于 https://github.com/Alex66366/ODPure。
cs.CV / 91 / 2609.28262

RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models

RAMP:面向边缘CPU视觉模型的鲁棒自适应混合精度量化
Población-Criado, David, Garcia-Gasulla, Dario, Quinones, Eduardo
Abstract
Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy. However, quantization affects different layer types in inconsistent ways, so identifying where accuracy loss is minimized and latency reduction is maximized is critical, as the effect accumulates over a full deployment into substantial savings or unacceptable task degradation. Such identification relies on sensitivity metrics, proxies that estimate layer-wise degradation without evaluating the task accuracy of every candidate policy. Nevertheless, widely used metrics fail systematically on modern architectures. We present a systematic empirical study of 13 sensitivity metrics for layer-wise INT8 quantization across four distinctly different neural networks, and validate the resulting policies on two ARM64 platforms. Gradient-based sensitivity methods fail on 4 out of 8 model-hardware configurations and weight-based statistics on 2. In contrast, the Jensen-Shannon Divergence achieves zero catastrophic failures, reliably isolating the layers that cannot be safely quantized. A sensitivity metric alone does not define a policy, and the fixed thresholds typically used for that step are fragile over the highly skewed distributions of modern architectures. We address this with K-Means clustering, achieving near-lossless accuracy and a mean speed-up of $1.81\times$ over the full-precision model. Finally, we reveal that excluding from quantization the layers whose speed-up is negligible, regardless of their sensitivity, can be counterproductive, as it induces computational graph fragmentation and disables operator fusion. Our results yield concrete allocation policies for practitioners and researchers deploying quantized vision models on heterogeneous edge CPUs, without GPU access or gradient computation.
Chinese Translation
在边缘CPU上部署深度学习模型受到计算和内存约束的瓶颈限制。混合精度量化有望在保持精度的同时降低推理延迟。然而,量化对不同层类型的影响并不一致,因此确定精度损失最小、延迟降低最大的层至关重要,因为这种影响会在整个部署中累积,形成可观的节省或不可接受的任务性能退化。此类识别依赖于敏感度指标——即在无需评估每个候选策略的任务精度的情况下,估计各层退化的代理指标。然而,广泛使用的指标在现代架构上会系统性地失效。我们对13种面向逐层INT8量化的敏感度指标在四个截然不同的神经网络上进行了系统的实证研究,并在两个ARM64平台上验证了所得的量化策略。基于梯度的敏感度方法在8个模型-硬件配置中有4个失效,基于权重的统计方法有2个失效。相比之下,Jensen-Shannon散度(Jensen-Shannon Divergence)实现了零灾难性失效,能够可靠地识别出无法安全量化的层。仅有敏感度指标并不足以定义量化策略,而该步骤中通常使用的固定阈值在现代架构高度偏斜的分布上十分脆弱。我们采用K-Means聚类来解决这一问题,实现了近乎无损的精度,相比全精度模型获得了平均1.81倍的加速。最后,我们揭示了将加速收益可忽略的层排除在量化之外(无论其敏感度如何)可能适得其反,因为这会导致计算图碎片化并使算子融合失效。我们的研究结果为在异构边缘CPU上部署量化视觉模型、且无GPU访问或梯度计算的从业者和研究人员提供了具体的分配策略。
cs.CV / 92 / 2609.28283

Benchmarking Hyperspectral Foundation Models for Hyperspectral Unmixing

面向高光谱解混的高光谱基础模型基准测试
Dabier, Edgard, Kervazo, Christophe, Gori, Pietro, Tupin, Florence
Abstract
Several foundation models dedicated to hyperspectral images have recently been made available. These models are trained on large unlabeled datasets and exhibit strong performance on many hyperspectral imaging tasks, such as classification or denoising. Nonetheless, their performance for hyperspectral unmixing -- the task of separating mixed spectra of overlapping materials in a hyperspectral image -- remain understudied. This might partly be due to the fact that most of them rely on vision transformer backbones, including patchification, leading to a feature resolution problem. While hyperspectral unmixing already arises from the low resolution of hyperspectral images, this patchification step potentially makes the problem even more ill-posed. Therefore, in this work, we aim to answer two questions: 1) \emph{how do foundation models perform in hyperspectral unmixing?}; 2) \emph{how to tackle the feature-level loss of resolution?} To answer the first question, we benchmark foundation models for unmixing, showing that they can reach state-of-the-art performance on four hyperspectral unmixing datasets. To answer the second question, we compare several feature upsampling approaches and empirically show that using a simple one can lead to high performance results. The code is available at https://gitlab.telecom-paris.fr/ring/hfm-hsu.git.
Chinese Translation
近期涌现了多个专门针对高光谱图像的基础模型。这些模型在大型无标签数据集上进行训练,在分类、去噪等众多高光谱成像任务中展现出优异的性能。然而,它们在高光谱解混——即分离高光谱图像中重叠地物的混合光谱的任务——上的表现仍缺乏充分研究。这部分原因可能在于大多数模型采用视觉Transformer(vision transformer)骨干网络,其中包含图像块划分(patchification)操作,从而导致特征分辨率问题。高光谱解混本身已因高光谱图像的低分辨率而产生,而这一图像块划分步骤可能使该问题更加不适定。因此,本工作旨在回答两个问题:1)基础模型在高光谱解混中的表现如何?2)如何应对特征层面的分辨率损失?针对第一个问题,我们对用于解混任务的基础模型进行了基准测试,结果表明它们在四个高光谱解混数据集上均能达到最先进的性能。针对第二个问题,我们比较了多种特征上采样方法,并通过实验证明采用一种简单的方法即可获得高性能的结果。代码已发布于 https://gitlab.telecom-paris.fr/ring/hfm-hsu.git。
cs.CV / 93 / 2609.28286

PBLH Estimation from Satellite Radiances via a Dual-Encoder Transformer

基于双编码器Transformer的卫星辐射率行星边界层高度估算
Innocenti, Lorenzo, Catalano, Luca, Arnaudo, Edoardo, Rossi, Claudio, Larosa, Salvatore, Cimini, Domenico, Garza, Paolo
Abstract
Estimating the Planetary Boundary Layer Height (PBLH) from satellite observations is a challenging regression problem due to the indirect relationship between top-of-atmosphere radiances and near-surface atmospheric structure. Progress has been limited both by the lack of architectures capable of handling the multimodal, spatially incomplete nature of satellite overpasses, and by the scarcity of suitable datasets. In this paper, we build upon the large-scale dataset pairing MetOp radiances with ERA5 PBLH labels that we introduced in our previous work, making three contributions. First, we establish a benchmark across eight approaches spanning pixel-wise regression, swath-wise sequence models, and convolutional and Transformer models operating on the full orbital passage. Second, we quantify what the resulting model actually relies on, using grouped Shapley decomposition over the input blocks. Third, we present the best-performing architecture found: a dual-encoder Transformer whose masked-input handling lets it operate in all weather conditions. The proposed model achieves MAE = 155.8 m on the held-out global test set, outperforming all baselines on every evaluation subset. On 30 out-of-distribution granules acquired on two days overlapping the TEAMx observational campaign, it achieves MAE = 165.3 m, outperforming a pixel-wise baseline trained on the same data (MAE = 197 m).
Chinese Translation
从卫星观测中估算行星边界层高度(PBLH)是一个具有挑战性的回归问题,原因在于大气层顶辐射率与近地面大气结构之间的间接关系。进展一直受限于两方面:一是缺乏能够处理卫星过境数据多模态、空间不完整特性的架构,二是缺乏合适的数据集。在本文中,我们基于此前工作中构建的大规模数据集(该数据集将MetOp辐射率与ERA5 PBLH标签配对),做出了三项贡献。第一,我们建立了一个涵盖八种方法的基准测试,这些方法包括逐像素回归、沿轨序列模型,以及在完整轨道数据上运行的卷积模型和Transformer模型。第二,我们通过对输入块进行分组Shapley分解,量化了模型实际依赖的因素。第三,我们提出了性能最优的架构:一种双编码器Transformer(dual-encoder Transformer),其掩码输入处理机制使其能够在所有天气条件下运行。所提出的模型在留出的全球测试集上达到了MAE = 155.8 m,在所有评估子集上均优于所有基线方法。在两天内获取的、与TEAMx观测计划时间重叠的30个分布外(out-of-distribution)数据粒上,该模型达到了MAE = 165.3 m,优于在相同数据上训练的逐像素基线模型(MAE = 197 m)。
cs.CV / 94 / 2609.28300

RoomLight: A 2.5D Illumination Prior for Indoor Environments

RoomLight:一种面向室内环境的2.5D光照先验
Ardelean, Andreea, Egger, Bernhard
Abstract
Ill-posed inverse problems require priors to constrain the solution space toward plausible outcomes. In inverse rendering, learned priors modeling the distribution of natural illumination improve the recovery of scene properties. However, existing models rely on the distant-illumination assumption, representing lighting as a far-field environment map. This limits their applicability to indoor scenes, where illumination is highly spatially varying due to finite-distance emitters, visibility changes, and parallax, all of which are poorly approximated by a single environment map. To address this, we introduce a spatially-aware illumination prior trained on real-world indoor panoramas and their estimated depth. Our variational autoencoder model learns a compact, optimizable latent space that decodes into HDR radiance and depth, parameterizing an area light emitter for direct integration into standard differentiable rendering pipelines. This design bridges the plausibility guarantees of a learned prior with the gradient flow required for downstream optimization. Crucially, by jointly modeling radiance and depth, our prior captures the spatial structure of indoor illumination, instead of treating the light sources as infinitely distant. We demonstrate that this formulation enables spatially-varying illumination modeling and achieves higher-fidelity recovery of indoor lighting compared to existing approaches. Project page: https://andreead-a.github.io/RoomLight
Chinese Translation
病态逆问题需要先验来约束解空间,使解趋向于合理的结果。在逆渲染中,对自然光照分布建模的学习先验能够改善场景属性的恢复效果。然而,现有模型依赖于远距离光照假设,将光照表示为远场环境贴图。这限制了其在室内场景中的应用,因为室内光照由于有限距离的光源、可见性变化以及视差效应而呈现高度的空间变化性,这些都无法通过单一环境贴图很好地近似。为解决这一问题,我们提出了一种空间感知的光照先验,该先验在真实世界的室内全景图像及其估计深度上进行训练。我们的变分自编码器(VAE)模型学习到一个紧凑且可优化的潜在空间,可解码为HDR辐射亮度和深度,并对面光源发射器进行参数化,从而可直接集成到标准的可微渲染管线中。这一设计将学习先验的合理性保证与下游优化所需的梯度流相结合。关键在于,通过联合建模辐射亮度和深度,我们的先验能够捕捉室内光照的空间结构,而不是将光源视为无限远。我们证明,该表述能够实现空间变化光照的建模,并在室内光照恢复方面相比现有方法达到更高的保真度。项目页面:https://andreead-a.github.io/RoomLight
cs.CV / 95 / 2609.28327

LightMIS: Ultra-Lightweight Medical Image Segmentation Without a Stage-Wise Decoder

LightMIS:无需阶段式解码器的超轻量级医学图像分割网络
Arhire, Andrei, Breabăn, Mihaela-Elena, Timofte, Radu
Abstract
We present LightMIS, a scalable family of ultra-lightweight convolutional networks for 2D binary medical image segmentation without a learned stage-wise decoder. LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, and refines the fused representation with an Adaptive Fusion Cascade. The cascade combines Adaptive Kernel Fusion with the proposed Progressive Receptive Fusion module, which uses temporary channel expansion, complementary depthwise receptive fields, and progressive cross-branch information transfer. We evaluate LightMIS-T, LightMIS-S, and LightMIS using five-fold cross-validation under a common nnU-Net v2.3.1 protocol on DRIVE, Kvasir-SEG, DSB18, BUSI, ISIC-2017, and ISIC-2018. Full LightMIS contains 0.131 M parameters and requires 0.575 GFLOPs for a $3\times256\times256$ input, achieving modality-macro Dice and IoU scores of 86.71% and 78.99%, respectively. Mobile U-ViT obtains 86.75% Dice and 79.07% IoU, so the observed differences are 0.04 and 0.08 percentage points. Relative to Mobile U-ViT, nnWNet, and nnU-Net, LightMIS reduces parameter count by 90.58$-$99.61% and GFLOPs by 82.54$-$96.14%. On an Arm Mali-G52 MC2 GPU, all LightMIS variants achieve full GPU delegation, with median delegated latency ranging from 53.31 ms for LightMIS-T to 138.31 ms for LightMIS. These results demonstrate a favorable accuracy$-$complexity trade-off and on-device execution feasibility for the evaluated tasks. The code is publicly available at https://github.com/AndreiiArhire/LightMIS.
Chinese Translation
我们提出了LightMIS,一个可扩展的超轻量级卷积网络系列,用于二维二值医学图像分割,且无需可学习的阶段式解码器。LightMIS使用尺度对齐投影(Scale-Aligned Projection)模块将五级编码器的输出对齐到统一分辨率,对其进行一次性聚合,并通过自适应融合级联(Adaptive Fusion Cascade)对融合后的表征进行细化。该级联将自适应核融合(Adaptive Kernel Fusion)与所提出的渐进式感受野融合(Progressive Receptive Fusion)模块相结合,后者利用临时通道扩展、互补的深度级感受野以及渐进式跨分支信息传递。我们在DRIVE、Kvasir-SEG、DSB18、BUSI、ISIC-2017和ISIC-2018数据集上,采用统一的nnU-Net v2.3.1协议下的五折交叉验证,对LightMIS-T、LightMIS-S和LightMIS进行评估。完整版LightMIS仅包含0.131 M参数,处理$3\times256\times256$输入仅需0.575 GFLOPs,在模态宏平均下分别取得86.71%的Dice分数和78.99%的IoU分数。Mobile U-ViT取得86.75%的Dice分数和79.07%的IoU分数,二者观测到的差异仅为0.04和0.08个百分点。相对于Mobile U-ViT、nnWNet和nnU-Net,LightMIS将参数量减少90.58$-$99.61%,将GFLOPs减少82.54$-$96.14%。在Arm Mali-G52 MC2 GPU上,所有LightMIS变体均可实现完整的GPU委托执行,中位委托延迟从LightMIS-T的53.31 ms到LightMIS的138.31 ms不等。这些结果证明了所评估任务在精度$-$复杂度权衡和端侧执行可行性方面的优势。代码已公开发布于 https://github.com/AndreiiArhire/LightMIS。
cs.CV / 96 / 2609.28328

BronchoTop: Bronchoscopy Navigation via RGB-Only Topological Localization

BronchoTop:基于纯RGB拓扑定位的支气管镜导航
Tomasini, Clara, Murillo, Ana Cristina, Riazuelo, Luis
Abstract
Accurate localization of the bronchoscope within the bronchial tree is essential for clinicians to be able to reach target lesions, perform biopsies and avoid misidentification of airway segments during diagnostic and therapeutic procedures. However, existing navigation systems typically rely on patient-specific CT scans or additional external sensors, increasing cost, setup time and patient radiation exposure. This work presents BronchoTop, a real-time, RGB-only framework for topological bronchoscopy localization that eliminates the need for patient-specific data. BronchoTop estimates scope location relative to a generic airway model through four modules: lumen detection and tracking, lumen-branch label association, probabilistic scope location estimation, and switch verification. By using only standard bronchoscopy video input, BronchoTop provides practical, real-time navigational assistance to physicians. Evaluation on phantom, simulated and real data demonstrates state-of-the-art accuracy, improving existing approaches performance by over 20% on real bronchoscopy sequences. BronchoTop is the first published framework including both the localization algorithms as well as all the real data used, together with code to generate additional simulations, encouraging and facilitating further developments and benchmarking. The results highlight BronchoTop's potential to enhance procedural safety, efficiency and accessibility in clinical and robotic bronchoscopy.
Chinese Translation
在诊断和治疗过程中,支气管镜在支气管树内的准确定位对于临床医生到达目标病灶、进行活检以及避免误判气道段至关重要。然而,现有的导航系统通常依赖患者特定的CT扫描或额外的外部传感器,增加了成本、准备时间和患者的辐射暴露。本工作提出了BronchoTop,一个实时的、仅使用RGB图像的支气管镜拓扑定位框架,无需患者特定数据。BronchoTop通过四个模块估计支气管镜相对于通用气道模型的位置:管腔检测与跟踪、管腔-分支标签关联、概率性支气管镜位置估计以及切换验证。通过仅使用标准支气管镜视频输入,BronchoTop为医生提供实用的实时导航辅助。在体模、仿真和真实数据上的评估表明其达到了最先进的精度,在真实支气管镜序列上将现有方法的性能提升了20%以上。BronchoTop是首个同时公开定位算法和全部真实数据(以及生成额外仿真数据的代码)的已发表框架,鼓励并促进了进一步的开发与基准测试。结果表明,BronchoTop有潜力提升临床和机器人支气管镜手术的安全性、效率和可及性。
cs.CV / 97 / 2609.28342

Zero-Shot Object Removal via Attention Masking, Latent Anchoring, and Refinement

基于注意力掩码、潜空间锚定与精修的零样本物体移除方法
Taghizadeh, Arman, Krumnack, Ulf, Kühnberger, Kai-Uwe
Abstract
Removing an object from a real image requires more than synthesizing plausible content within a mask: the method must suppress residual object features, preserve the unedited scene, and generate replacement content that is consistent with the surrounding background. This paper approaches object removal from a stage-based perspective and proposes a zero-shot framework for constrained latent inpainting with a frozen pretrained Stable Diffusion model, requiring no task-specific training or model fine-tuning. The method integrates SAM-based mask construction, BLIP image-caption conditioning, DDIM inversion, background-weighted masked null-text optimization, decoder self-attention masking, hard outside-mask latent anchoring, and localized renoise--denoise refinement into a unified pipeline. The method is evaluated through qualitative examples, quantitative local-consistency metrics, and ablation studies. The results demonstrate effective object removal and context-consistent replacement content. The ablations indicate that background-weighted masked NTI is particularly beneficial for structurally complex backgrounds, whereas the no-NTI variant is sufficient in other evaluated examples. Repeated refinement further reduces object remnants and boundary artifacts remaining after the primary editing pass.
Chinese Translation
从真实图像中移除物体不仅仅是合成掩码区域内看似合理的内容:方法还必须抑制残留的物体特征、保持未被编辑的场景不变,并生成与周围背景相一致的替换内容。本文从分阶段的角度研究物体移除问题,提出了一种基于冻结的预训练Stable Diffusion模型的零样本约束潜空间修复框架,无需任务特定的训练或模型微调。该方法将基于SAM的掩码构建、BLIP图像-文本描述条件引导、DDIM反演、背景加权的掩码空文本优化、解码器自注意力掩码、掩码外硬潜空间锚定以及局部重加噪-去噪精修整合为统一流水线。方法通过定性示例、定量局部一致性指标以及消融实验进行评估。结果表明该方法能够有效移除物体并生成上下文一致的替换内容。消融实验表明,背景加权的掩码空文本优化(NTI)对结构复杂的背景尤其有益,而在其他评估示例中,无NTI的变体已足够有效。重复的精修过程可进一步减少首次编辑后残留的物体痕迹和边界伪影。
cs.CV / 98 / 2609.28360

Privacy-Preserving Semantic Segmentation from High-Resolution Depth and Ultra-Low-Resolution RGB

基于高分辨率深度与超低分辨率RGB的隐私保护语义分割
Huang, Xuying, Daniel, Swithinraj Moses, Pan, Sicong, Houben, Sebastian, Bennewitz, Maren
Abstract
As mobile robots become increasingly integrated into everyday environments, privacy risks arising from onboard cameras have become a growing concern. Ultra-low-resolution (ULR) RGB can mitigate visual privacy exposure at the source, but ULR appearance alone substantially limits semantic and spatial understanding. We therefore introduce a privacy-preserving asymmetric sensing setting that combines high-resolution (HR) depth with ULR RGB, preserving dense geometry while restricting fine-grained visual information. To address the severe information imbalance between HR depth and ULR RGB, we propose a joint 2D framework using HR geometry to guide semantic-oriented RGB reconstruction and RGB-D segmentation. Despite reliable frame-level predictions, consistent scene-level understanding remains challenging under the asymmetric HR depth--ULR RGB setting. We therefore develop an end-to-end 2D-to-3D pipeline that consolidates 2D semantic features for 3D segmentation. Experiments on ScanNet show that our method achieves the best 2D and 3D segmentation performance among privacy-preserving approaches and delivers the strongest zero-shot transfer to SUN RGB-D and SceneNN. Privacy recoverability analysis shows that our proposed HR depth--ULR RGB input reduces the recoverability of sensitive data, and real-robot experiments demonstrate the utility of the resulting 3D semantics for object-goal navigation.
Chinese Translation
随着移动机器人日益融入日常环境,机载摄像头带来的隐私风险日益受到关注。超低分辨率(ULR)RGB可以从源头上降低视觉隐私泄露风险,但仅依靠ULR外观信息会严重限制语义与空间理解能力。为此,我们提出一种隐私保护的非对称感知设置,将高分辨率(HR)深度与ULR RGB相结合,在限制细粒度视觉信息的同时保留稠密的几何信息。为解决HR深度与ULR RGB之间严重的信息不平衡问题,我们提出一个联合2D框架,利用HR几何信息引导面向语义的RGB重建以及RGB-D分割。尽管帧级预测较为可靠,但在HR深度与ULR RGB非对称设置下,实现一致的场景级理解仍然具有挑战性。因此,我们开发了一条端到端的2D到3D流水线,将2D语义特征整合用于3D分割。在ScanNet上的实验表明,我们的方法在隐私保护方法中取得了最优的2D和3D分割性能,并在SUN RGB-D和SceneNN上展现出最强的零样本迁移能力。隐私可恢复性分析表明,我们所提出的HR深度与ULR RGB输入降低了敏感数据的可恢复性;真实机器人实验则验证了所得3D语义在目标物导航中的实用性。
cs.CV / 99 / 2609.28366

AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios

AnchorReasoning:面向长尾自动驾驶场景的视觉定位与因果推理数据集
Bao, Zhipeng, Zhao, Wenjie, Zhu, Tianle, Que, Haohua, Yang, Chence, Yuan, Geng, Li, Qianwen
Abstract
Vision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning. We introduce AnchorReasoning, a visually grounded reasoning dataset built on WOD-E2E, containing 416,119 annotated frames and 395,379 decision-critical elements across four major categories and 19 fine-grained types. Each frame is organized as a visually grounded chain-of-thought (VG-CoT) that links decision-critical element identification and localization, element attributes and implications, driving-action rationale, and action and trajectory planning. We further develop a curriculum supervised fine-tuning strategy that progressively learns these hierarchical capabilities, together with an object-size-aware grounding metric for evaluating localization quality. Experiments across eight general-purpose, embodied-AI, and AV-specific backbones show that VG-CoT supervision improves grounded reasoning and trajectory prediction. Across models, 5-s ADE and FDE decrease by 7.84 and 11.86, while RFS Frame and Cluster improve by 1.66 and 1.70. These gains are achieved with 18.5 fewer reasoning tokens and 0.32 s/frame lower inference latency on average, demonstrating the value of visually grounded, decision-focused supervision for VLM reasoning and planning in long-tail autonomous driving.
Chinese Translation
视觉语言模型为长尾自动驾驶提供了一种有前景的方法,但现有驾驶数据集在将决策关键视觉证据与推理和规划相连接方面提供的监督有限。我们提出了AnchorReasoning,一个基于WOD-E2E构建的视觉定位推理数据集,包含416,119个标注帧,涵盖四大类和19种细粒度类型共395,379个决策关键元素。每个帧被组织为一条视觉定位思维链,将决策关键元素的识别与定位、元素属性及其影响、驾驶行为依据以及动作与轨迹规划串联起来。我们进一步提出了一种课程式监督微调策略,以渐进方式学习这些层次化能力,并设计了一个尺寸感知的定位评估指标来评价定位质量。在八个通用、具身智能和自动驾驶专用骨干模型上的实验表明,VG-CoT监督能提升视觉定位推理和轨迹预测能力。各模型的5秒ADE和FDE分别降低7.84和11.86,RFS Frame和Cluster分别提升1.66和1.70。这些提升是在推理token减少18.5%、每帧推理延迟平均降低0.32秒的情况下实现的,证明了视觉定位、以决策为中心的监督对VLM在长尾自动驾驶中推理与规划的价值。
cs.CV / 100 / 2609.28414

Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model

冻结的流会遗忘:诊断并恢复潜在流世界模型中丢失的运动信息
Chen, Xiwen, Li, Rigaudiere Z., Zhou, Zhiruo, Zhu, Xiaojun, Liu, Houde
Abstract
Latent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, yet silently lose the property manipulation depends on most: motion. The pretrained flow never moves the manipulated object; retraining it with latent-only losses only trades stillness for teleport-like motion. We trace the failure to the training signal, not the representation: anchor-sparse, latent-only supervision never says where along the horizon change belongs. Decode-augmented rollout training (DART) repairs this while keeping the representation frozen, retraining only the flow with decode-path supervision. DART outperforms its latent only parent on the full protocol, restores the temporal structure of motion, and re-couples predicted motion to the scene; at larger scale it further improves prediction quality, closing nearly half the remaining gap to an oracle-informed interpolation reference. Finally, we report an unexpected finding about evaluation: pixel error alone rewards frozen predictions.
Chinese Translation
在冻结的自监督潜在空间中集成流(flow)的潜在世界模型训练稳定且成本低廉,然而却悄然丢失了操控任务最依赖的属性:运动。预训练的流从不移动被操控的物体;而仅使用潜在损失对其重新训练,只会用“瞬间移动”式的运动取代静止。我们将这一失败归因于训练信号而非表征本身:锚点稀疏、仅基于潜在空间的监督从未告知变化应当出现在时间轴的哪个位置。解码增强的滚动训练(Decode-augmented Rollout Training, DART)在保持表征冻结的前提下解决了这一问题,仅通过解码路径监督对流进行重新训练。在完整的评测协议下,DART 优于其仅使用潜在损失的父模型,恢复了运动的时间结构,并重新将预测的运动与场景耦合;在更大规模下,它进一步提升了预测质量,将与基于真值信息的插值参考(oracle-informed interpolation reference)之间剩余的差距缩小了近一半。最后,我们报告了一个关于评估的意外发现:仅凭像素误差会奖励冻结的预测。
cs.CV / 101 / 2609.28424

The Skin-Restricted Reinhard Transform:Uniqueness under a Lightness-Preserving Constraint

肤色受限的Reinhard变换:亮度保持约束下的唯一性
KP, Vijesh
Abstract
Catalog skin recolouring has to change pigment and leave shading alone. The classical Reinhard map does not make that split: it rescales lightness by the ratio of standard deviations, and a flat reference swatch therefore flattens the limb. This paper formalises the correction used in our pipeline, the skin-restricted Reinhard transform. It is the diagonal affine map in CIE Lab that translates lightness, matches the chromatic mean, and clamps the chromatic gain to [0.72, 1.18], with moments taken on the central 84% of each channel. A diagonal affine map has six real parameters. The shading constraint forces the lightness gain to +1 and the lightness shift to the difference of means; one-dimensional quadratic optimal transport on each chromatic axis, followed by Euclidean projection onto the gain interval, fixes the other four. Inside that family the four conditions determine every parameter. The content of the result is the forced lightness gain; it is not a uniqueness claim outside the diagonal affine class. For Gaussian marginals the chromatic step is not merely the best affine map: it is the unrestricted Wasserstein-2 map. The same formulae with trimmed moments remain optimal because a positive affine image commutes with quantile trimming. On hands, arms, legs, and feet of nine photographs and three reference tones, the map keeps the lightness contrast ratio at 0.974 +/- 0.029 with chromatic error 0.77 CIE Lab units. Reinhard matching, the linear Monge map, and histogram matching reach a smaller chromatic error only by cutting lightness contrast to about half.
Chinese Translation
产品目录中的肤色重着色需要改变色素而保留阴影。经典的Reinhard映射无法实现这一区分:它按标准差之比缩放亮度,因而平坦的参考色板会使肢体阴影被抹平。本文对我们流程中使用的校正方法——肤色受限的Reinhard变换——进行了形式化。该变换是CIE Lab空间中的对角仿射映射,它平移亮度、匹配色度均值,并将色度增益限制在[0.72, 1.18]区间内,其中各通道的矩量取自中央84%的数据。对角仿射映射有六个实参数。阴影约束迫使亮度增益为+1、亮度平移为均值之差;各色度轴上的一维二次最优传输,随后向增益区间进行欧几里得投影,确定其余四个参数。在该映射族内,四个条件唯一确定所有参数。该结果的核心在于被强制的亮度增益;在对角仿射类之外,本文并不作唯一性断言。对于高斯边际分布,色度步骤不仅是最佳仿射映射,而且是无约束的Wasserstein-2映射。采用截尾矩的相同公式依然保持最优,因为正仿射变换与分位数截尾可交换。在九张照片中的手、臂、腿、脚以及三种参考色调上,该映射将亮度对比度保持在0.974 ± 0.029,色度误差为0.77个CIE Lab单位。Reinhard匹配、线性Monge映射和直方图匹配只能在将亮度对比度削减至约一半的情况下才能达到更小的色度误差。
cs.CV / 102 / 2609.28434

Predicting the Progression of Adolescent Idiopathic Scoliosis

预测青少年特发性脊柱侧凸的进展
Pullen, Owen, Jamaludin, Amir, Zisserman, Andrew
Abstract
Adolescent Idiopathic Scoliosis is defined as a lateral curvature of the spine that develops during adolescence, without known cause. The condition can result in significant pain and disability, and often progresses rapidly during adolescence. The objective of this paper is to predict the progression of the condition in a temporal sequence from ages 9 to 24, as measured from a sequence of Dual X-ray Absorptiometry (DXA) scans. To this end, we train a transformer model that takes in the curve of the spine to predict curve progression. The model is trained using a large-scale synthetic dataset of spine curves and their time series, covering different curve types and different progression patterns. We show that the model is able to generalise from synthetic to real data by evaluating it on a dataset of real DXA scans covering multiple time points. We find that fine-tuning the model on real data gives a significant boost to performance. The model is able to accurately predict spine curve progression in both scoliosis and normal cases.
Chinese Translation
青少年特发性脊柱侧凸(Adolescent Idiopathic Scoliosis)是指青春期发生的、病因不明的脊柱侧向弯曲。该疾病可导致明显的疼痛和功能障碍,且在青春期往往进展迅速。本文的目标是基于一系列双能X线吸收测定(DXA)扫描,预测该疾病在9至24岁时间序列中的进展情况。为此,我们训练了一个以脊柱弯曲曲线为输入的Transformer模型来预测弯曲的进展。该模型使用一个大规模的脊柱曲线及其时间序列合成数据集进行训练,涵盖了不同的弯曲类型和不同的进展模式。通过在一个包含多个时间点的真实DXA扫描数据集上进行评估,我们证明该模型能够从合成数据泛化到真实数据。我们发现,在真实数据上对模型进行微调能显著提升性能。该模型能够准确预测脊柱侧凸病例和正常病例的脊柱弯曲进展。
cs.CV / 103 / 2609.28437

MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos

MultiVENT-Raw:一个面向原始视频检索与推理的基准数据集
Kriz, Reno, Etter, David, Martin, Alexander, Carpenter, Cameron, Chakraborty, Debashish, Recknor, Hannah, Iranmanesh, Reihaneh, Maciejewski, Matthew, Murray, Kenton, Yang, Eugene, Van Durme, Benjamin, White, Aaron Steven, Yates, Andrew, Walden, William
Abstract
Online information is increasingly consumed in video format. Much of this comes in the form of *raw video*: continuous footage taken on a cell phone, with a hand-held camera, or via CCTV, which is then directly uploaded to social media platforms and content sharing services. Whereas professional or even amateur-edited footage tends to feature scripted speech, chyrons, graphics, and metadata that help contextualize its subject matter, raw video typically contains none of these things, making it a much more challenging medium for information retrieval and machine understanding. To facilitate progress in this domain, we release MultiVENT-Raw, a multilingual collection of nearly 120,000 primarily raw videos (over 5,300 total hours), paired with 130 events and 222 event-centric queries, along with human-annotated video relevance judgments and human-extracted key facts for relevant videos. MultiVENT-Raw supports both a retrieval task---to identify videos in the collection relevant to a query event---and a generation task---to summarize event-related videos into a coherent report for a target user. We benchmark strong baselines on MultiVENT-Raw, showing both tasks to be challenging even for some of the latest multimodal models.
Chinese Translation
在线信息越来越多地以视频形式被消费。其中大部分以*原始视频*的形式出现:即用手机、手持摄像机或监控摄像头(CCTV)拍摄的连续片段,随后被直接上传至社交媒体平台和内容分享服务。与专业剪辑甚至业余剪辑的视频往往包含脚本化解说、字幕条、图形和有助于理解主题的元数据不同,原始视频通常不包含这些元素,使其成为信息检索和机器理解更具挑战性的媒介。为推动该领域的发展,我们发布了MultiVENT-Raw,这是一个多语言数据集,包含近120,000个以原始视频为主的视频(总时长超过5,300小时),并配有130个事件和222个以事件为中心的查询,以及人工标注的视频相关性判断和相关视频的人工提取关键事实。MultiVENT-Raw支持两项任务:一是检索任务——识别数据集中与查询事件相关的视频;二是生成任务——将事件相关视频总结为面向目标用户的连贯报告。我们在MultiVENT-Raw上对强大的基线模型进行了基准测试,结果表明即使对于一些最新的多模态模型,这两项任务仍然极具挑战性。
cs.CV / 104 / 2609.28439

HaRP: High Dynamic Range Photosequencing through Dual Reversed Shutter Scanning

HaRP:基于双反向快门扫描的高动态范围照片序列重建
Ji, Xiang, Lin, Guixu, Zhao, Jiancheng, Yin, Zhengwei, Zheng, Yinqiang
Abstract
The adoption of CMOS sensors in mobile photography is frequently compromised by the rolling shutter (RS) effect, which introduces geometric distortions and motion artifacts. Particularly, recent rolling shutter with global reset (RSGR) mode, while mitigating some RS issues, also incurs major limitations, including reduced capture speed and compressed dynamic range. To address these problems, we propose a novel dual reversed scanning setup utilizing both RSGR and inverted RSGR views. This solution not only handles the inherent flaws of RSGR by synchronizing complementary exposures to balance the dynamic range across the frames but also introduces an effective method for HDR photosequencing under highly dynamic scenes. Our proposed network first accommodates row-wise complementarity and manages visual shifts by row-adaptive feature alignment. Subsequently, the hallucination module, built upon a correlation-guided mixattention block, integrates the mutually reinforced features to recover missing details. In addition, we construct a coaxial imaging system to collect a real-world dataset, enabling robust training and evaluation beyond numerical simulation. Experimental results demonstrate the twofold benefits of our solution in mitigating RSGR limitations and advancing HDR reconstruction techniques.
Chinese Translation
CMOS传感器在移动摄影中的应用常常受到卷帘快门(Rolling Shutter, RS)效应的影响,该效应会引入几何畸变和运动伪影。特别是,近期提出的全局复位卷帘快门(Rolling Shutter with Global Reset, RSGR)模式虽然缓解了部分RS问题,但也带来了主要局限,包括采集速度降低和动态范围压缩。为解决这些问题,我们提出了一种新颖的双反向扫描设置,同时利用RSGR视图和反转RSGR视图。该方案不仅通过对互补曝光进行同步以平衡各帧间的动态范围,从而克服RSGR的固有缺陷,还为高动态场景下的高动态范围(HDR)照片序列重建提供了一种有效方法。我们提出的网络首先适应行级互补性,并通过行自适应特征对齐来处理视觉偏移;随后,基于相关性引导的混合注意力(mix-attention)模块构建的幻觉(hallucination)模块融合相互增强的特征,以恢复缺失的细节。此外,我们构建了一套同轴成像系统来采集真实世界数据集,使得训练和评估能够超越数值仿真,具备更强的鲁棒性。实验结果表明,我们的方案在缓解RSGR局限性和推进HDR重建技术方面具有双重优势。
cs.CV / 105 / 2609.28466

The Past Frames the Future: Memory for Autoregressive Video Generation

过去构筑未来:面向自回归视频生成的记忆机制
Chen, Harold Haodong, Guo, Rongjin, Lan, Disen, Shu, Wen-Jie, Zhang, Hongfei, Hu, Hanzhe, Yao, Shengtao, Zhang, Zixin, Zhang, Guibin, Rao, Zhefan, Liu, Jinxiu, Liu, Yexin, Peng, Rui, Liu, Yuhao, Ren, Bin, Yang, Shuai, Chen, Yukang, Khan, Salman, Chen, Ying-Cong, Lim, Ser-Nam, Lau, Rynson W. H., Sebe, Nicu, Cheng, Yu, Yang, Ming-Hsuan, Chen, Qifeng
Abstract
Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands, practical models must operate under strictly bounded context windows, storage, and computational limits. Consequently, critical historical information, e.g., entity identities, dynamic states, and intervention-induced causal changes, often leaves the active context long before its relevance diminishes. Overcoming this limitation and maintaining temporal persistence constitutes a fundamental memory problem. We present a systematic and comprehensive review of memory mechanisms in AR video generation. We formulate memory operationally as persistent historical information maintained across outer AR steps, capable of influencing future generation even after the originating evidence is no longer locally accessible. Building upon this unified framework, we organize the literature through five complementary perspectives: (I) Forms, the representational carriers of history; (II) Functions, the specific semantic and physical information requiring preservation; (III) Operations, the lifecycle of writing, reading, updating, managing, and integrating memory; (IV) Learning, the optimization of memory behaviors under closed-loop rollouts; and (V) Evaluation, the paradigms for diagnosing genuine memory capabilities. We conclude by synthesizing open challenges, including composable and resource-aware memory architectures, trustworthy state updating, self-rollout learning, and standardized evaluation. By bridging representations, mechanisms, and learning paradigms, this paper establishes a structured foundation for developing reliable, memory-conditioned video generation systems.
Chinese Translation
生成式模型的进展提升了视频保真度,使得长时程生成、交互式世界建模以及不断演化的视觉环境成为可能。自回归(Autoregressive, AR)视频生成通过因果式展开(causal rollout)扩展视觉序列。然而,一个根本性的瓶颈随之出现:随着生成序列的不断扩展,实际模型必须在严格受限的上下文窗口、存储和计算预算下运行。因此,关键的历史信息,例如实体的身份、动态状态以及干预引发的因果变化,往往在其相关性尚未消退之前就早已脱离了活动上下文。克服这一局限并保持时间上的持续性,构成了一个根本性的记忆问题。本文对AR视频生成中的记忆机制进行了系统而全面的综述。我们将记忆在操作层面定义为在AR外部步骤之间持续维护的历史信息,即使其原始证据在局部已不可访问,仍能够影响未来的生成。基于这一统一框架,我们从五个互补的视角组织相关文献:(I)形式(Forms),即历史信息的表征载体;(II)功能(Functions),即需要保留的特定语义与物理信息;(III)操作(Operations),即记忆的写入、读取、更新、管理与整合的生命周期;(IV)学习(Learning),即在闭环展开中对记忆行为的优化;(V)评估(Evaluation),即诊断真实记忆能力的范式。最后,我们总结了开放性挑战,包括可组合且资源感知的记忆架构、可信的状态更新、自展开(self-rollout)学习以及标准化评估。通过贯通表征、机制与学习范式,本文为开发可靠的、以记忆为条件的视频生成系统奠定了结构化的基础。
cs.CV / 106 / 2609.28473

On the Diffusibility of High-Dimensional Latents

论高维潜在表示的可扩散性
Feng, Chao, Xu, Zhiyang, Chen, Bowei, Xiong, Yuanjun, Wang, Xiyao, Wang, Jui-Hsien, Zhang, Richard, Lin, Zhe, Owens, Andrew, Li, Yijun
Abstract
Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image reconstruction recovers such details. However, perhaps counterintuitively, this procedure reduces the effective dimensionality of the resulting representation, and the altered geometry has downstream effects on generation. Specifically, we show that using the standard velocity prediction in flow matching in this high-dimensional space requires the model to fit orthogonal noise directions outside the low-dimensional signal manifold, making optimization inefficient. This motivates using the clean data parameterization ($\boldsymbol{x}_{0}$-prediction) instead, which focuses learning on the underlying signal manifold. Across experiments with multiple strong-reconstruction encoders, we show that $\boldsymbol{x}_{0}$-prediction consistently improves text-to-image generation performance.
Chinese Translation
表示自编码器(RAE)使扩散模型能够在预训练视觉编码器的特征空间中运行。然而,许多现成的编码器并未针对忠实重建进行优化,会丢弃细粒度的视觉细节。正如预期的那样,针对图像重建对这些编码器进行微调可以恢复这些细节。然而,也许与直觉相反,这一过程降低了所得表示的有效维度,且改变的几何结构会对下游生成产生影响。具体而言,我们证明在该高维空间中使用流匹配(flow matching)的标准速度预测,需要模型去拟合低维信号流形之外的正交噪声方向,导致优化效率低下。这促使我们转而采用干净数据参数化(即 x0-预测),它将学习聚焦于潜在的信号流形上。在使用多个具有强重建能力的编码器进行的实验中,我们证明 x0-预测能够持续提升文本到图像的生成性能。
机器学习 (Machine Learning)
115
cs.LG / 1 / 2609.26811

The Drift Contract: Spectral Updates for Depth-Robust Local Learning

漂移契约:面向深度鲁棒局部学习的谱更新方法
Polly, Fabien
Abstract
Local learning trains each layer with its own auxiliary loss and no global backward pass, which makes layer updates structurally parallel. Two problems have kept it marginal: accuracy degrades as depth grows, and hyperparameters are fragile. We apply Muon-style spectral update geometry (momentum orthogonalization with spectral step scaling) to per-layer local updates, an intersection not previously studied. On CIFAR-10 MLP benchmarks with local linear heads, a single step-size setting is the best value in our tested grids from width 128 to 2048 and from depth 12 to 48, while local Adam requires re-tuning along both axes and still collapses at depth 48 (31.3 percent re-tuned per depth, 19 percent with its depth-12 setting transferred, vs 42.7 percent for the spectral update at its unchanged setting). At five seeds and width 512 the spectral update leads local Adam by a clear margin (48.9 +/- 0.5 vs 46.6 +/- 0.3). Prospectively specified controls attribute the transfer and most of the depth robustness to the spectral geometry itself rather than to any step-size rule on top of it. We additionally formulate the step size as a drift contract, lr = epsilon / RMS(input), which bounds each layer's weight-induced pre-activation change per step, conditioned on its current input. The contract yields a small gain over the best fixed learning rate where that baseline is measured, makes the step size interpretable, and provides a per-layer, input-conditioned drift bound that standard optimizers do not offer. We report one negative result: with RMSNorm and weight decay in the trunk, the stability benefit of spectral updates accrues to global rather than local training, so the local advantage concentrates precisely where normalization is absent.
Chinese Translation
局部学习通过为每一层设置各自的辅助损失进行训练,且无需全局反向传播,这使得各层的更新在结构上可以并行。然而,两个问题使其一直处于边缘地位:随着深度增加精度会下降,且超参数较为脆弱。我们将 Muon 风格的谱更新几何结构(带谱步长缩放的动量正交化)应用于逐层的局部更新,这一交叉领域此前尚未被研究过。在带局部线性输出头的 CIFAR-10 MLP 基准测试中,单一的步长设置在宽度从 128 到 2048、深度从 12 到 48 的所有测试网格中均为最优值,而局部 Adam 则需要沿这两个维度重新调参,且在深度 48 时仍然崩溃(逐深度重新调参后为 31.3%,直接沿用其深度 12 的设置为 19%,而谱更新在设置完全不变的情况下为 42.7%)。在五个随机种子、宽度 512 的设置下,谱更新以明显优势领先于局部 Adam(48.9 ± 0.5 对 46.6 ± 0.3)。预先设定的对照实验将这种可迁移性和大部分深度鲁棒性归因于谱几何结构本身,而非叠加其上的任何步长规则。我们进一步将步长形式化为一个漂移契约,即 lr = epsilon / RMS(input),该契约以每层当前输入为条件,约束每一步中由权重引起的预激活变化的上界。该契约在可测量的基线之上相对最优固定学习率带来了小幅提升,使步长具有可解释性,并提供了标准优化器所不具备的逐层、以输入为条件的漂移界。我们还报告了一个负面结果:当主干网络中包含 RMSNorm 和权重衰减时,谱更新的稳定性收益归于全局训练而非局部训练,因此局部学习的优势恰好集中于缺乏归一化的场景。
cs.LG / 2 / 2609.26820

Signal2Symbol: Neuro-Symbolic Temporal Reasoning for Explainable Physiological Time-Series Anomaly Detection

Signal2Symbol:面向可解释生理时间序列异常检测的神经-符号化时序推理
Mansour, Naser, Benabderrahmane, Sidahmed, Rahwan, Ameer
Abstract
Physiological time series such as electrocardiograms (ECG) and electroencephalograms (EEG) exhibit complex temporal structure, substantial acquisition variability, and a strong need for transparent decision-making. Although deep models can achieve high detection performance, they often provide limited insight into why a segment is anomalous, how local anomalies relate over time, and whether a detection belongs to a broader recurring pattern. We propose Signal2Symbol, a neuro-symbolic framework for explainable biosignal anomaly detection. The method first converts ECG/EEG signals into symbolic sequences using either a learned VQ-VAE (Vector Quantized Variational Autoencoder) codebook or a SAX (Symbolic Aggregate approXimation) baseline. It then constructs bigram enriched token-window transactions and scores anomalies through rare itemset evidence derived from minimal rare itemset mining. Detected anomalous windows are merged into intervals and related using Allen interval algebra, enabling composite temporal explanations such as escalation chains, artifact overlap, and cross-channel synchrony. Finally, we introduce a rare temporal concept lattice based on Formal Concept Analysis (FCA), which groups anomalous intervals by shared rare symbolic evidence, Allen temporal relations, channel context, and robustness attributes. The resulting Galois lattice compresses many local detections into interpretable families of temporal-symbolic anomalies. We evaluate on three public benchmarks: MIT-BIH Arrhythmia (beat-level ECG), PTB-XL (record-level ECG), and the Bonn EEG dataset (segment-level EEG). We stress-test robustness under additive noise and baseline-wander perturbations. The results highlight the value of neuro-symbolic tokenization for temporal anomaly analysis and show that Allen/FCA reasoning provides compact, interpretable summaries of local detections.
Chinese Translation
心电图(ECG)和脑电图(EEG)等生理时间序列具有复杂的时序结构、显著的采集变异性,以及对透明决策的强烈需求。尽管深度模型能够取得较高的检测性能,但它们往往无法充分解释某一段为何异常、局部异常如何随时间相互关联,以及某次检测是否属于更广泛的重复模式。我们提出Signal2Symbol,一个用于可解释生物信号异常检测的神经-符号化框架。该方法首先使用学习得到的VQ-VAE(向量量化变分自编码器)码本或SAX(符号化聚合近似)基线方法,将ECG/EEG信号转换为符号序列。随后构建二元组(bigram)增强的词元窗口事务,并通过基于最小频繁项集挖掘得到的稀有项集证据对异常进行评分。检测到的异常窗口被合并为区间,并利用Allen区间代数建立关联,从而支持诸如升级链、伪迹重叠和跨通道同步等复合时序解释。最后,我们引入基于形式概念分析(FCA)的稀有时序概念格,依据共享的稀有符号证据、Allen时序关系、通道上下文和鲁棒性属性对异常区间进行分组。所得的Galois格将大量局部检测压缩为可解释的时序-符号异常族。我们在三个公开基准上进行评估:MIT-BIH心律失常数据库(心拍级ECG)、PTB-XL(记录级ECG)以及Bonn EEG数据集(片段级EEG),并在加性噪声和基线漂移扰动下进行鲁棒性压力测试。结果凸显了神经-符号化词元化在时序异常分析中的价值,并表明Allen/FCA推理能够为局部检测提供紧凑且可解释的摘要。
cs.LG / 3 / 2609.26822

HARN: Hierarchical Associative Resonance Network for Event-Driven Multi-Timeframe Forecasting

HARN:用于事件驱动多时间框架预测的分层联想共振网络
Saidd, Nabeel Ahmad
Abstract
Financial time series evolve across multiple temporal resolutions, challenging forecasting systems to incorporate newly available information without repeatedly recomputing unchanged representations. We introduce HARN, a Hierarchical Associative Resonance Network for event-driven multi-timeframe forecasting. HARN maintains persistent representations across temporal levels and updates each level only when its corresponding completed bar becomes available. The architecture combines causal multi-scale temporal encoding, gated associative memory, cross-level resonance, and hierarchical evidence aggregation, with forecasting performed in basis-point space and reconstructed to the original price scale. We evaluate HARN on four assets spanning equity, foreign exchange, and commodity markets using multiple random seeds and component ablations. HARN achieves competitive reconstructed-price forecasting errors against single-timeframe PatchTST and TimeXer baselines, while ablations reveal the effects of removing individual components across assets and timeframes. A code-level audit further examines consistency between the implementation and the defined event-driven causal protocol. The results position HARN as a persistent multi-timeframe forecasting framework rather than evidence of universal predictive superiority.
Chinese Translation
金融时间序列在多个时间分辨率上演化,这对预测系统提出了挑战:如何在不重复计算未发生变化表示的前提下整合新获得的信息。我们提出了 HARN,一种用于事件驱动多时间框架预测的分层联想共振网络(Hierarchical Associative Resonance Network)。HARN 在各时间层级上维护持久的表示,并且仅当对应层级的完整 K 线(bar)数据可用时才更新该层级。该架构结合了因果多尺度时间编码、门控联想记忆、跨层级共振以及分层证据聚合,并在基点空间中进行预测,再重建至原始价格尺度。我们在涵盖股票、外汇和大宗商品市场的四种资产上,通过多个随机种子和组件消融实验对 HARN 进行了评估。相对于单时间框架的 PatchTST 和 TimeXer 基线,HARN 在重建价格预测误差方面取得了具有竞争力的结果;同时,消融实验揭示了去除各个组件在不同资产和时间框架上的影响。代码级审计进一步检验了实现与所定义的事件驱动因果协议之间的一致性。研究结果将 HARN 定位为一个持久化的多时间框架预测框架,而非证明其具有普遍的预测优越性。
cs.LG / 4 / 2609.26826

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

是什么使Terminal-Bench任务变得困难?在一个经裁定校验的智能体语料库上区分真实难度与伪难度
Lip, Edward Lue Chee, Moraski, Boden, Knappe, Tim, Xiong, Lang, Gharat, Sarvesh, Mari, Antonio, Bercovich, Ivan
Abstract
Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capability gap, but it can also come from missing context, a broken reference solution, infrastructure failure, or a verifier that can be bypassed. In this paper, we study this issue using a frozen Terminal-Bench 3 / Frontier-Bench 0.1 production record with 1,081 pull requests, 639 scored tasks, 28,801 trials, and $105,933 in logged agent spend. We ask what an all-fail task actually certifies. For the 125 tasks with no honest pass, we combine task artifacts, reference-solution runs, empty-solution controls, adversarial trials, trajectories, telemetry, and review records, and apply an ordered validity screen. Only 78 of the 125 tasks survive as certified-unsolved candidates. The remaining tasks include 14 with broken oracles, 8 dominated by infrastructure failures, 4 that are only passable through verifier bypasses, and 21 whose solvability is not certified by the available evidence. Thus, lack of saturation and genuine difficulty are not the same thing. The certified-unsolved label is also narrow: it means that the authored route passed, infrastructure did not dominate, no strict bypass was observed, and all evaluated agents failed. It does not prove intrinsic hardness, verifier completeness, or failure at the intended capability. We further analyze rejected submissions and passing tasks to show that pass rate alone cannot explain why a task is difficult. Overall, our results suggest that frontier benchmarks should report the evidence behind their all-fail tasks before using them as capability claims.
Chinese Translation
前沿基准测试需要当前模型无法解决的任务。但没有任何模型能解决的任务并不自动就是困难任务。同样的零通过率可能来自真实的能力差距,但也可能来自上下文缺失、参考解答损坏、基础设施故障,或可被绕过的验证器。本文基于一个冻结的Terminal-Bench 3 / Frontier-Bench 0.1生产记录(包含1,081个拉取请求、639个计分任务、28,801次试验以及累计105,933美元的智能体开销记录)研究这一问题。我们追问一个“全部失败”的任务究竟证明了什么。针对125个无诚实通过的任务,我们综合任务产物、参考解答运行、空解对照、对抗性试验、轨迹、遥测数据和审查记录,并应用有序的有效性筛查。最终,125个任务中仅有78个保留为“经验证未解决”的候选任务。其余任务包括:14个存在损坏的判分基准(oracle)、8个被基础设施故障主导、4个只能通过绕过验证器才能通过,以及21个可解性未得到现有证据证明的任务。因此,未饱和与真实困难并非同一回事。“经验证未解决”这一标签的含义也很有限:它仅表示作者设定的求解路径已通过、基础设施未占主导、未观察到严格的绕过行为,且所有被评估的智能体均告失败。它并不能证明任务的内在难度、验证器的完备性,或在目标能力上的失败。我们进一步分析被拒提交与通过的任务,表明仅凭通过率无法解释一个任务为何困难。总体而言,我们的结果表明,前沿基准测试在将全部失败的任务用作能力声明之前,应当先报告这些任务背后的证据。
cs.LG / 5 / 2609.26839

LWCal: Loss-Weighted Calibration for Tabular Classifiers with Noisy Calibration Labels

LWCal:面向带噪声校准标签的表格分类器的损失加权校准方法
Liu, Zeming, Lyu, Hang, Zhang, Jingtao, Xie, Yuan
Abstract
Post-hoc probability calibration is usually evaluated under an optimistic assumption: the held-out calibration labels are clean. In many AI deployment settings, however, labels come from weak annotators, historical decisions, heuristics, or distant supervision, so the same label noise that corrupts training also corrupts calibration. We study this overlooked failure mode for tabular classifiers and propose LWCal, a CPU-only post-hoc calibrator that down-weights calibration examples whose noisy labels are contradicted by the base model's held-out probability. LWCal requires no clean validation labels, no noise-rate estimate, and no retraining of the base classifier. A second variant, Gated-LWCal, adds a conservative disagreement gate that backs off toward the raw score when the calibration split appears extremely inconsistent. On nine local binary tabular tasks, six random seeds, symmetric and asymmetric label corruption, and three tree-based base learners, LWCal obtains the lowest average calibration error while Gated-LWCal obtains the best average proper-score tradeoff. In the main random-forest study over 432 noisy cells, Gated-LWCal reduces expected calibration error from 0.188 to 0.122 and negative log likelihood from 0.438 to 0.396 relative to the raw classifier. Paired bootstrap intervals for Gated-LWCal versus raw, Platt, isotonic, and beta calibration exclude zero on ECE, Brier score, and NLL. The artifact contains all scripts, result tables, figures, and the compiled paper.
Chinese Translation
事后概率校准通常在一个乐观假设下进行评估:留出的校准标签是干净的。然而在许多AI部署场景中,标签来自弱标注者、历史决策、启发式规则或远程监督,因此污染训练的标签噪声同样会污染校准过程。我们针对表格分类器研究了这一被忽视的失效模式,并提出LWCal——一种仅需CPU的事后校准器,它会降低其噪声标签与基础模型在留出集上的预测概率相矛盾的校准样本的权重。LWCal不需要干净的验证标签、不需要噪声率估计,也无需重新训练基础分类器。第二种变体Gated-LWCal增加了一个保守的分歧门控,当校准集显得极不一致时回退到原始分数。在九个本地二分类表格任务、六个随机种子、对称与非对称标签污染以及三种基于树的基学习器上,LWCal获得了最低的平均校准误差,而Gated-LWCal获得了最佳的平均适当评分(proper score)权衡。在涵盖432个噪声单元格的随机森林主实验中,相对于原始分类器,Gated-LWCal将期望校准误差(ECE)从0.188降至0.122,负对数似然(NLL)从0.438降至0.396。Gated-LWCal与原始方法、Platt校准、等渗回归校准及Beta校准的配对Bootstrap置信区间在ECE、Brier分数和NLL上均不包含零。该工件包含所有脚本、结果表格、图表以及编译后的论文。
cs.LG / 6 / 2609.26848

A Leakage-Aware Multimodal Evaluation Framework for Early Intraoperative Acute Kidney Injury Prediction

一种具有泄漏感知的多模态评估框架用于术中早期急性肾损伤预测
Nguyen, Quang Minh, Le, Duc Minh, Nguyen, Ho Nhat Minh, Nguyen, Thuy Quynh, Nguyen, Trong Nghia
Abstract
Postoperative acute kidney injury (AKI) after major non-cardiac surgery carries substantial morbidity, yet early intraoperative risk stratification remains difficult. In this retrospective cohort study, we propose SynerT, a waveform-only hybrid temporal backbone that combines a causal dilated TCN with a hierarchy of dilated recurrent layers to encode early intraoperative physiologic trajectories for AKI risk prediction. Building on SynerT, we further design two model variants that extend the backbone with structured clinical context: SynerT-MM, a late-fusion multimodal extension that integrates hemodynamic burden summaries and preoperative covariates, and SynerTStack, a leakage-safe stacked ensemble that combines cross-validated predictions from SynerT-MM with strong tabular baselines at the meta-learning stage. All models are evaluated under a strict leakage-aware framework on VitalDB, a high-fidelity perioperative database, with prediction restricted to information available within the first 60 intraoperative minutes. Among 2,413 waveform-usable cases (180 AKI-positive; 7.46% prevalence), SynerT fell well below strong structured-data baselines, demonstrating that waveform-only temporal modeling is insufficient under strict early constraints. SynerTMM recovered discrimination by incorporating hemodynamic burden summaries and preoperative covariates, and SynerT-Stack achieved the best overall performance across AUROC, AUPRC, and F1-max. Cross-fitted Platt recalibration substantially corrected calibration defects in both multimodal variants, and decision-curve analysis confirmed the recalibrated stacked model delivered the strongest net clinical benefit across low-to-intermediate thresholds.
Chinese Translation
大型非心脏手术后的急性肾损伤(AKI)具有较高的发病率,然而术中早期的风险分层仍然困难。在这项回顾性队列研究中,我们提出了 SynerT,一种仅使用波形数据的混合时序骨干网络,它将因果扩张型时序卷积网络(TCN)与分层扩张型循环层相结合,用于编码术中早期生理轨迹以预测 AKI 风险。在 SynerT 的基础上,我们进一步设计了两个扩展结构化临床背景信息的模型变体:SynerT-MM,一种后期融合的多模态扩展模型,整合了血流动力学负荷摘要和术前协变量;以及 SynerT-Stack,一种具有防泄漏特性的堆叠集成模型,在元学习阶段将 SynerT-MM 的交叉验证预测与强大的表格数据基线相结合。所有模型均在高保真围术期数据库 VitalDB 上,在一个严格的泄漏感知框架下进行评估,预测仅限于术中前 60 分钟内可获取的信息。在 2,413 例波形可用的病例中(AKI 阳性 180 例;患病率 7.46%),SynerT 的表现远低于强大的结构化数据基线,这表明在严格的早期约束条件下,仅依赖波形的时序建模是不充分的。SynerT-MM 通过整合血流动力学负荷摘要和术前协变量恢复了判别能力,而 SynerT-Stack 在 AUROC、AUPRC 和 F1-max 上均取得了最佳的整体性能。交叉拟合的 Platt 重新校准显著纠正了两个多模态变体的校准缺陷,决策曲线分析证实,重新校准后的堆叠模型在低至中等阈值范围内提供了最强的净临床获益。
cs.LG / 7 / 2609.26853

COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation

COPE:基于用户嵌入与自我评估、在稀疏用户反馈下对大语言模型进行持续个性化
Cao, Ruike, Yao, Fugen, Dong, Liang, Xu, Jian, Jiang, Guanjun, Xiao, Li
Abstract
While Large Language Models (LLMs) have achieved remarkable results across various benchmarks, their alignment with normative values often results in homogenized responses that fail to address diverse user preferences. Existing training-free methods often occupy valuable context windows through prompt engineering, while training-based methods typically remain static post-training, failing to support the continual optimization required in real-world settings. To address these challenges, we propose COPE (Continual Optimization with Personalized embedding and self-Evaluation), a novel optimization framework tailored for real-world-motivated interaction settings with sparse user feedback. Our framework assigns learnable personalized embeddings to each user and synergistically integrates preference capture, self-evaluation calibration, and personalized response optimization within a single update step. A key innovation of our method is the use of self-evaluation to generate proxy rewards, enabling continuous model updates even when explicit user feedback is unavailable. Experiments show that COPE consistently outperforms strong training-free and training-based baselines under sparse feedback, and remains complementary to Retrieval-Augmented Prompting (RAP). Further analyses confirm COPE's reliable self-evaluation, meaningful preference patterns, stable general capabilities, and robustness under shifting preferences and alternative evaluators.
Chinese Translation
尽管大语言模型(LLM)在各类基准测试中取得了显著成果,但其与规范性价值观的对齐往往导致同质化的回复,无法满足多样化的用户偏好。现有的免训练方法通常通过提示工程占用宝贵的上下文窗口,而基于训练的方法在训练后通常保持静态,无法支持真实场景中所需的持续优化。为应对这些挑战,我们提出了COPE(Continual Optimization with Personalized embedding and self-Evaluation,基于个性化嵌入与自我评估的持续优化),一种面向真实交互场景中稀疏用户反馈的新型优化框架。该框架为每个用户分配可学习的个性化嵌入,并在单次更新步骤中协同整合偏好捕捉、自我评估校准与个性化回复优化。该方法的一项关键创新是利用自我评估生成代理奖励,即使缺少显式用户反馈也能实现模型的持续更新。实验表明,在稀疏反馈条件下,COPE持续优于强大的免训练与基于训练的基线方法,并且与检索增强提示(Retrieval-Augmented Prompting, RAP)保持互补性。进一步的分析验证了COPE可靠的自我评估、有意义的偏好模式、稳定的通用能力,以及在偏好变化和更换评估器情况下的鲁棒性。
cs.LG / 8 / 2609.26855

QUARTET: Quad-branch cross-Attention and Random-walk Traces for Enhancing Transformers on Relational Graphs

QUARTET:面向关系图Transformer增强的四分支交叉注意力与随机游走追踪方法
Myint, Kyaw Hpone, Jiang, Nan, Li, Xiang, Wu, Zhe, Day, Alexandre G. R., Mohanty, Pranab, Iyengar, Giri
Abstract
Relational Deep Learning (RDL) models multi-table databases as heterogeneous temporal graphs, and graph transformers currently achieve state-of-the-art performance on benchmarks like RelBench. However, the current leading model, RelGT, suffers from two key limitations: its random local sampler yields loosely connected subgraphs that hinder message passing, and its global attention module relies on a single, seed-feature-based memory that ignores broader macro-level dynamics. To overcome these limitations, we introduce QUARTET, an expressive graph transformer architecture that applies full self-attention on local subgraphs while enriching global context through cross-attention branches. Specifically, QUARTET employs a Causal Random Walk (CRW) sampler based on recency-truncated Personalized PageRank (PPR) to extract compact, hub-robust, and densely connected local subgraphs without temporal leakage. Concurrently, a quad-branch cross-attention module integrates global context from four complementary perspectives: seed feature, seed topology, temporal dynamics, and collaborative dynamics. Across the RelBench v1 classification tasks, QUARTET consistently matches or outperforms the current state-of-the-art graph transformer baselines (HGT and RelGT). Ablation studies confirm that the CRW sampler significantly enriches local neighborhood quality, while the global branches provide essential, task-specific predictive gains.
Chinese Translation
关系深度学习(Relational Deep Learning, RDL)将多表数据库建模为异构时序图,而图Transformer目前在RelBench等基准测试中取得了最先进的性能。然而,当前领先的模型RelGT存在两个关键局限:其随机局部采样器生成的子图连接松散,阻碍了消息传递;其全局注意力模块仅依赖于基于种子特征的单一路径记忆,忽略了更广泛的宏观动态。为克服这些局限,我们提出了QUARTET,这是一种表达能力强的图Transformer架构,它在局部子图上应用完整的自注意力机制,同时通过交叉注意力分支丰富全局上下文。具体而言,QUARTET采用基于近因截断个性化PageRank(Personalized PageRank, PPR)的因果随机游走(Causal Random Walk, CRW)采样器,以提取紧凑、对枢纽节点鲁棒且连接稠密的局部子图,同时避免时间信息泄露。同时,四分支交叉注意力模块从四个互补视角整合全局上下文:种子特征、种子拓扑结构、时序动态和协同动态。在RelBench v1的分类任务中,QUARTET始终匹配或超越当前最先进的图Transformer基线模型(HGT和RelGT)。消融实验证实,CRW采样器显著提升了局部邻域的质量,而全局分支则提供了关键的、针对特定任务的预测增益。
cs.LG / 9 / 2609.26866

Marginally Correct Tool Caches Can Reverse Group-Normalized Policy Updates

边际正确的工具缓存可能逆转组归一化策略更新
Gupta, Shivam
Abstract
Tool-result caching reduces repeated execution in agent training, but also couples rollout randomness. We study a two-action model in which independent and shared execution preserve every rollout's conditional reward distribution. Despite this marginal agreement, sharing one stochastic result per group can reverse the expected group-normalized policy update. We derive an exact finite-group expression: against a constant alternative, the shared update follows the probability of winning minus the probability of losing, rather than the difference in expected reward. A Bernoulli specialization yields a wrong-direction region and a non-vanishing update-variance floor as group size grows. Centering without group standard-deviation scaling preserves the expected-return direction in this model, using an existing estimator control. Exhaustive finite sums verify 540 configurations and 3,240 estimator evaluations, with a separate ordered-sequence checker. An implementation audit reproduces the sharing path in a pinned, unmodified TVCache stack using 256 scripted rollouts. These results do not measure language-model training performance or refute TVCache's deterministic-output contract. They establish that marginal output validity alone cannot certify a stochastic cache as training-equivalent.
Chinese Translation
工具结果缓存减少了智能体训练中的重复执行,但同时耦合了rollout的随机性。我们研究了一个双动作模型,其中独立执行与共享执行均能保持每次rollout的条件奖励分布不变。尽管存在这种边际一致性,但组内共享同一个随机结果仍可能逆转期望意义上的组归一化策略更新。我们推导了一个精确的有限群组表达式:相对于一个常数备选动作,共享执行后的更新取决于获胜概率减去失败概率,而非期望奖励之差。在Bernoulli(伯努利)特例下,该模型存在一个更新方向错误的区域,且随着组规模增大,更新方差存在一个不趋于零的下界。在该模型中,仅做中心化而不进行组标准差缩放可保持期望回报的更新方向,这一点可由现有的估计量控制方法保证。我们通过穷举有限和验证了540种配置和3240次估计量评估,并使用一个独立的有序序列检查器进行核查。实现审计方面,我们在固定版本、未修改的TVCache技术栈中,使用256个脚本化rollout复现了共享执行路径。这些结果并未测量语言模型训练性能,也未推翻TVCache的确定性输出契约。它们确立了仅凭边际输出有效性不足以证明一个随机缓存与训练等价。
cs.LG / 10 / 2609.26890

PR-Smoother: Simulator-Preserving Non-Gaussian Smoothing for Data Assimilation

PR-Smoother:面向数据同化的保持模拟器的非高斯平滑方法
Tarumi, Yuta
Abstract
Many physical data assimilation (DA) workflows require smoothing methods that represent non-Gaussian posteriors over physical state variables, scale to high-dimensional simulators, train from observation windows alone, and remain compatible with calibration of the prescribed simulator. We introduce PR-Smoother, a simulator-preserving amortized smoother designed for this prescribed-simulator DA regime. Its key design principle is to keep the prescribed simulator explicit in both the evidence lower bound and the variational family: rather than learning replacement dynamics or a learned trajectory prior, PR-Smoother learns only future-conditioned corrections around the prescribed rollout. This yields an explicit non-Gaussian smoothing distribution over physical trajectories and supports joint state, parameter, and sensor-bias learning from observations alone. The variational family contains the exact smoother in deterministic and linear-Gaussian limits. Empirically, PR-Smoother captures multimodal posteriors in 4-dimensional Lorenz-96, remains accurate under ambiguous nonlinear observations and process noise in 40-dimensional Lorenz-96, and scales to joint state-parameter-bias inference in 16,384-dimensional Kolmogorov flow.
Chinese Translation
许多物理数据同化(DA)工作流要求平滑方法能够表示物理状态变量上的非高斯后验分布、可扩展至高维模拟器、仅从观测窗口进行训练,并与预设模拟器的校准保持兼容。我们提出了 PR-Smoother,一种专为该预设模拟器数据同化场景设计的、保持模拟器的摊销平滑器。其核心设计原则是在证据下界和变分族中都显式保留预设模拟器:PR-Smoother 不学习替代动力学或学习式轨迹先验,而仅学习围绕预设模拟 rollout 的未来条件化修正。由此,它能够给出物理轨迹上的显式非高斯平滑分布,并支持仅基于观测的状态、参数与传感器偏差的联合学习。该变分族在确定性极限和线性高斯极限下包含精确平滑器。实验表明,PR-Smoother 能够捕捉 4 维 Lorenz-96 系统中的多峰后验分布,在 40 维 Lorenz-96 系统中于模糊的非线性观测和过程噪声下仍保持准确,并可扩展至 16384 维 Kolmogorov 流中状态-参数-偏差的联合推断。
cs.LG / 11 / 2609.26905

CORE-STACK+: Meta-Learning for Deep Stacked Generalization

CORE-STACK+:面向深度堆叠泛化的元学习方法
Mohammad, Noor Islam S.
Abstract
Stacking heterogeneous vision backbones (CNNs, ViTs, and hybrids) is the de facto recipe for accuracy, calibration, and robustness, yet two coupled pathologies limit its returns. Prediction-space multicollinearity ill-conditions the meta-learner's Gram matrix, inflating weight variance and producing brittle solutions on a thin manifold. Calibration collapse compounds constituent miscalibration through naive linear stacking, so adding more models can hurt expected calibration error (ECE). Existing remedies, ridge regularization, greedy selection, model soups, and SWAG address at most one of these issues, and none jointly target conditioning and calibration in heterogeneous prediction pools. We introduce CORE-STACK+, a preconditioning pipeline with four components: (i) a kernelized redundancy filter that removes non-linear inter-model dependencies invisible to Pearson correlation, using Centered Kernel Alignment (CKA) [23]; (ii) a $<15$K-parameter differentiable meta-feature gate that learns per-sample attention over ensemble statistics; (iii) a spectrum-adaptive Ridge penalty $lambda^{star}=lmax(Chat)/SNR(Chat)$ derived from a Marchenko-Pastur signal-noise decomposition, eliminating nested cross-validation; and (iv) a Laplace-approximate Bayesian blender replacing inverse-RMSE heuristics. We prove a PAC-Bayes excess-risk bound that, for the first time, jointly accounts for prediction-space redundancy and meta-learner capacity. Across six benchmarks, CORE-STACK+ delivers $+1.8\%$ top-1 on ImageNet-1K, $-4.2$ mCE on ImageNet-C, $+0.9$ mIoU on ADE20K, and $+1.3$ AP on COCO, while reducing retained models by 35-57% and inference FLOPs by up to $41%$. ECE improves $2.1\times$ over deep ensembles without post hoc temperature scaling.
Chinese Translation
堆叠异构视觉骨干网络(CNN、ViT 及混合模型)已成为提升准确率、校准性和鲁棒性的事实标准方法,然而两个相互耦合的病态问题限制了其收益。预测空间的多重共线性会使元学习器的 Gram 矩阵出现病态,导致权重方差膨胀,并在一个狭窄流形上产生脆弱的解。校准失效则通过朴素的线性堆叠进一步放大各组成模型自身的校准误差,使得增加更多模型反而可能恶化期望校准误差(ECE)。现有的解决方法——岭回归正则化、贪心选择、模型汤(model soups)和 SWAG——最多只能解决其中一个问题,且没有任何方法能同时针对异构预测池中的矩阵条件问题和校准问题。我们提出 CORE-STACK+,一个包含四个组件的预处理流水线:(i) 基于核方法的冗余过滤器,利用中心核对齐(CKA)[23] 消除 Pearson 相关性无法捕捉的非线性模型间依赖;(ii) 参数量小于 15K 的可微元特征门控,学习对集成统计量的逐样本注意力;(iii) 基于 Marchenko-Pastur 信号-噪声分解导出的频谱自适应岭惩罚 λ* = lmax(C_hat)/SNR(C_hat),从而免去嵌套交叉验证;(iv) 用拉普拉斯近似贝叶斯混合器取代逆 RMSE 启发式方法。我们证明了一个 PAC-Bayes 超额风险界,首次联合考虑了预测空间冗余与元学习器容量。在六个基准上,CORE-STACK+ 在 ImageNet-1K 上取得 +1.8% 的 top-1 准确率提升,在 ImageNet-C 上降低 4.2 mCE,在 ADE20K 上提升 +0.9 mIoU,在 COCO 上提升 +1.3 AP,同时将保留模型数量减少 35-57%,推理 FLOPs 最多降低 41%。ECE 相比深度集成提升了 2.1 倍,且无需事后温度缩放。
cs.LG / 12 / 2609.26918

On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning

论多目标强化学习中事后重标注导致的偏好覆盖坍缩
Bonin, Baptiste, Strickland, Caro, Durand, Audrey
Abstract
Hindsight relabeling which retroactively replacing a transition's goal with the outcome the agent actually achieved is an effective tool for improving sample-efficiency in Reinforcement Learning (RL). A natural extension to preference-conditioned multi-objective RL (MORL) relabels transitions with the preference direction the agent achieved rather than the one asked for. We show that this extension is frequently harmful: across four preference-conditioned off-policy algorithms spanning two critic backbones and two preference-sampling schemes on the continuous-control MO-Gymnasium suite, it degrades 19 of 36 algorithm-environment settings by as much as four standard deviations, improves only one, and leaves the rest unaffected. The harm is not a symptom of noisy relabels; denoising the target recovers almost nothing, and neither prioritized sampling nor any buffer-structural choice reproduces it. Instead, repeated relabeling collapses the critic's coverage onto whatever narrow region of the preference space the agent happened to visit. We name this failure mode \emph{Preference Coverage Collapse}, and quantify it with abandoned preference mass (APM), a value-aware statistic that tracks the harm ($\rho = -0.73$) where a purely structural coverage count does not. We then introduce \texttt{her\_mix}, a single-parameter convex combination pulling the achieved direction back towards the requested preference. At one fixed value across every algorithm and environment, it returns 16 of the 19 harmed settings to baseline, preserves and even improves the one setting in which relabeling helps, and cuts abandoned preference mass from $69\%$ to $6\%$. Protecting coverage over the preference simplex, not filtering noisy relabels, is what makes hindsight relabeling safe for MORL.
Chinese Translation
事后重标注(hindsight relabeling)——即通过将转移的目标回溯性地替换为智能体实际达成的结果——是提升强化学习(RL)样本效率的有效工具。将其自然扩展到偏好条件化的多目标强化学习(MORL)中,即以智能体实际达成的偏好方向(而非所要求的偏好方向)对转移进行重标注。我们证明这一扩展往往是有害的:在连续控制的 MO-Gymnasium 环境套件上,跨两个评论家(critic)骨干网络和两种偏好采样方案的四种偏好条件化离线策略算法中,36 个算法-环境组合里有 19 个性能下降,最多可达四个标准差,仅有一个得到改善,其余不受影响。这种危害并非噪声重标注的症状:对目标进行去噪几乎无法恢复性能,优先级采样和任何经验回放缓冲区的结构选择也无法复现该现象。相反,反复的重标注使评论家的覆盖范围坍缩到智能体恰好访问过的偏好空间的狭窄区域上。我们将这种失败模式命名为偏好覆盖坍缩(Preference Coverage Collapse),并用被放弃偏好质量(abandoned preference mass, APM)来量化它——这是一个价值感知的统计量,能够追踪这种危害(ρ = -0.73),而纯粹结构性的覆盖计数则不能。随后我们提出 her_mix,一种单参数凸组合方法,将实际达成的方向拉回所请求的偏好。在所有算法和环境中使用同一个固定值时,它使 19 个受损组合中的 16 个恢复到基线水平,保持甚至提升了重标注起作用的唯一一个组合,并将被放弃偏好质量从 69% 降至 6%。保护偏好单纯形上的覆盖范围——而非过滤噪声重标注——才是使事后重标注在 MORL 中安全使用的关键。
cs.LG / 13 / 2609.26955

When Post-Processing Fairness Constraints Help and When They Harm: Evidence from Eight Cross-Domain Evaluations

后处理公平性约束何时有效、何时有害:来自八项跨领域评估的证据
Narla, Nithin Raghava Ramachandra
Abstract
Fairness audits in production ML typically occur once, at deployment, on a single domain. Both fail in practice: fairness can shift after retraining or a changing user base, and interventions validated on one dataset are rarely tested across the heterogeneous domains an organization deploys. We present FAPE (Fairness Auditing for Production Environments), a four-stage framework evaluating a single post-processing intervention, Fairlearn's ThresholdOptimizer, across eight domain evaluations: criminal justice, income prediction, legal admissions, credit lending, agricultural lending, a multi-domain benchmark corpus, healthcare, and education. Each is scored on demographic parity and equalized odds difference, plus disparate impact ratio and accuracy cost where computable. Intervention effectiveness tracks baseline disparity magnitude: across model-domain pairs the constraint improved disparity in 9 of 14 high-disparity cases and worsened it in 3 of 4 near-fair ones. Each of the five high-disparity exceptions reverses under one of two measurement checks, a minimum group size or thresholds fit on held-out data. A CUSUM monitor started at deployment, tested on a simulated shift, separates constrained models that never met a 0.1 parity convention from those that met it and later regressed. A single deployment-time audit is therefore an unreliable guide, which argues for baseline-disparity screening and continuous monitoring
Chinese Translation
生产环境机器学习中的公平性审计通常只进行一次,即在部署时针对单一领域进行。这两种做法在实践中都会失效:公平性可能在重新训练或用户群体变化后发生偏移,而在某一数据集上验证过的干预措施,很少会在组织实际部署的异构领域间进行测试。我们提出FAPE(Fairness Auditing for Production Environments,面向生产环境的公平性审计),这是一个四阶段框架,用于评估单一后处理干预措施——Fairlearn的ThresholdOptimizer——在八个领域评估中的表现:刑事司法、收入预测、法学院录取、信贷放贷、农业贷款、多领域基准语料库、医疗健康和教育。每个领域均以人口均等(demographic parity)和均等化几率差异(equalized odds difference)进行评分,并在可计算的情况下附加非歧视性影响比(disparate impact ratio)和准确率代价。干预的有效性与基线差异幅度相关:在所有模型-领域组合中,该约束在14个高差异案例中的9个改善了公平性,却在4个接近公平的案例中的3个使其恶化。五个高差异例外中的每一个,在以下两种测量检验之一下都会出现反转:最小群体规模约束,或基于留出数据拟合的阈值。一个从部署时启动、在模拟偏移上测试的CUSUM监测器,能够区分从未达到0.1均等惯例的受约束模型与曾达到但随后退化的模型。因此,单一的部署时审计是并不可靠的指引,这支持了基线差异筛查与持续监测的必要性。
cs.LG / 14 / 2609.26959

Transfer Learning with Conformalized Quantile Regression for Solar PV Forecasting Under Load-Shedding-Driven Data Scarcity

基于保形化分位数回归迁移学习的拉闸限电致数据稀缺条件下的太阳能光伏预测
Abdullah, Rakib, Faruk, K. M. Tahlil Mahfuz
Abstract
Solar photovoltaic (PV) forecasting in regions affected by load shedding is challenging because reliable historical observations are scarce. This study proposes a transfer learning framework combined with Conformalized Quantile Regression (CQR) to improve PV power forecasting and provide reliable uncertainty estimates under severe data scarcity. A source-domain PV dataset from Alice Springs, Australia, is used to pretrain a temporal forecasting model, which is then adapted to simulated Bangladesh PV data representing different levels of historical availability. Experimental results show that transfer learning reduces RMSE by up to 23.7% when only one month of target-domain data is available and by 13.7% with three months of data. The proposed Transfer Learning plus CQR framework achieves 94.3% empirical coverage with three months of target data while producing prediction intervals that are 14% narrower than those obtained without transfer learning. These results demonstrate that combining transfer learning with conformal uncertainty quantification can improve both point forecasting accuracy and uncertainty reliability when target-domain PV data are severely limited.
Chinese Translation
在受拉闸限电影响的地区,由于可靠的历史观测数据稀缺,太阳能光伏(PV)功率预测极具挑战性。本研究提出了一种结合保形化分位数回归(Conformalized Quantile Regression, CQR)的迁移学习框架,以在严重数据稀缺条件下改进光伏功率预测并提供可靠的不确定性估计。研究利用来自澳大利亚爱丽丝泉(Alice Springs)的光伏源域数据集对一个时间序列预测模型进行预训练,随后将其适配到代表不同历史数据可用性水平的模拟孟加拉国光伏数据上。实验结果表明,当目标域数据仅有一个月时,迁移学习可将RMSE降低最多23.7%;当有三个月数据时,可降低13.7%。所提出的迁移学习加CQR框架在使用三个月目标域数据时实现了94.3%的经验覆盖率,同时其预测区间比不使用迁移学习所得的区间窄14%。这些结果表明,当目标域光伏数据严重受限时,将迁移学习与保形不确定性量化相结合能够同时提升点预测精度和不确定性可靠性。
cs.LG / 15 / 2609.26962

CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning

CRISP:面向不平衡表格学习的可扩展重要性分层核心集方法
Mohanty, Hardhik, Rustandi, Indrayana, Sheibani, Mohamadreza
Abstract
Large imbalanced tabular datasets make repeated gradient-boosted tree training expensive. Existing coreset methods often lose accuracy when most majority examples are removed. We present CRISP (Coreset Reduction via Importance-Stratified Pruning), a linear-time method that allocates a negative-class budget across quantile strata of a proxy-model score. Sample weights account for unequal inclusion probabilities. At 95% negative-class reduction on a production fraud dataset, CRISP trains on approximately 1.70M of 25M rows and retains 99.7% of full-data Average Precision. This is a 93.2% reduction in total training rows. On public CriteoPrivateAds, CRISP has the highest mean Average Precision at each tested rate from 90% to 99.4% majority reduction. Sparkov results are mixed at lower rates, but CRISP has the highest mean at 99.2% and 99.4%. Ablations identify budget allocation and inverse-propensity weighting as the main sources of the production-dataset gain.
Chinese Translation
大型不平衡表格数据集使得反复训练梯度提升树(gradient-boosted tree)的成本高昂。现有的核心集(coreset)方法在移除大部分多数类样本时往往会损失精度。我们提出了 CRISP(Coreset Reduction via Importance-Stratified Pruning,基于重要性分层的核心集约简),这是一种线性时间方法,它将负类样本预算分配到代理模型得分的各个分位数层中。样本权重考虑了不同的纳入概率。在一个生产环境欺诈检测数据集上,将负类样本削减95%时,CRISP 仅使用约170万行(总行数为2500万)进行训练,并保留了全量数据平均精度(Average Precision)的99.7%,即总训练行数减少了93.2%。在公开的 CriteoPrivateAds 数据集上,CRISP 在从90%到99.4%的多数类削减率下的每个测试点均取得了最高的平均精度均值。在 Sparkov 数据集上,较低削减率下的结果好坏参半,但在99.2%和99.4%削减率下,CRISP 取得了最高的平均值。消融实验表明,预算分配和逆倾向加权是生产数据集上性能提升的主要来源。
cs.LG / 16 / 2609.26972

TinyUDE: Solver-Free Universal Differential Equations on Microcontrollers via Lie-Taylor Jet Matching

TinyUDE:基于李-泰勒射匹配的微控制器上无求解器通用微分方程
Balamurali, Pranavanath, Kamireddy, Hrishi
Abstract
Training Universal Differential Equations (UDEs) traditionally relies on backpropagating through numerical ODE solvers, creating memory footprints far exceeding the capabilities of edge microcontrollers. We present Lie-Taylor jet matching, a solver-free training framework that fits a hybrid vector field directly to the first and second time-derivatives of observed system states. These derivatives, the truncated Lie-Taylor jet, are estimated online via Savitzky-Golay filtering, yielding fully analytic gradients without automatic differentiation software. We evaluate whether eliminating the solver compromises accuracy against a conventional baseline (fixed-step RK4 integration, multiple shooting, exact discrete adjoints, Adam) sharing identical dynamics, noise models, network architectures, and metrics. While naive derivative matching degrades under sensor noise, our noise-adaptive mechanisms close and reverse this gap: full-rate phase-shifted sampling, a reservoir buffer, cosine-annealed optimization with weight averaging, on-device noise estimation, and polynomial-misfit quality gating. On a damped pendulum and chaotic double pendulum, our method matches or exceeds baseline accuracy at matched data windows and recovers unmodeled damping coefficients. Across noise levels from 0% to 5%, it attains a geometric-mean relative field error of 0.65x that of the baseline within 108 kB of static memory, compared with megabytes of solver tape. On an ESP32 microcontroller, the on-device run reaches a field error of 0.0020 and recovers the damping coefficient to c = 0.400 (true 0.400) within 61.3 kB of static memory and 7.24 ms per update (18.1% duty cycle at 25 Hz), confirming real-time on-device training is feasible without a numerical solver.
Chinese Translation
训练通用微分方程(UDE)传统上依赖于通过数值ODE求解器进行反向传播,其内存占用远超边缘微控制器的能力。我们提出李-泰勒射匹配(Lie-Taylor jet matching),一种无求解器的训练框架,它将混合向量场直接拟合到观测系统状态的一阶和二阶时间导数上。这些导数(即截断的李-泰勒射)通过Savitzky-Golay滤波在线估计,无需自动微分软件即可获得完全解析的梯度。我们评估了消除求解器是否会损害精度,并与共享相同动力学、噪声模型、网络架构和指标的传统基线(固定步长RK4积分、多重打靶法、精确离散伴随法、Adam)进行比较。尽管朴素的导数匹配在传感器噪声下性能下降,但我们的噪声自适应机制缩小并逆转了这一差距:全速率相移采样、储备池缓冲区、带权重平均的余弦退火优化、设备端噪声估计以及多项式失配质量门控。在阻尼摆和混沌双摆实验中,在相同数据窗口下,我们的方法达到或超过基线精度,并能恢复未建模的阻尼系数。在0%至5%的噪声水平下,该方法在108 kB静态内存内实现了几何平均相对场误差为基线的0.65倍,而求解器磁带需要数兆字节的内存。在ESP32微控制器上,设备端运行达到0.0020的场误差,并在61.3 kB静态内存和每次更新7.24 ms(25 Hz下18.1%占空比)内将阻尼系数恢复至c = 0.400(真实值0.400),证实了无需数值求解器的实时设备端训练是可行的。
cs.LG / 17 / 2609.26979

Resource-Efficient Distributed Recursive Gaussian Processes

资源高效的分布式递归高斯过程
King, Josephine, Balci, Ali Emre, Rajan, Raj Thilak
Abstract
Gaussian processes (GPs) provide a flexible framework for learning unknown functions from noisy measurements while quantifying predictive uncertainty, making them well suited for estimation in multi-agent systems. However, when measurements are collected by multiple agents, maintaining a unified GP model without centralized processing requires efficient distributed algorithms that can operate using local measurements and communication with neighboring agents. In this work, we develop two distributed recursive GP (RGP) algorithms for multi-output GP regression: ADMM-RGP and PDMM-RGP. We analyze the stability and convergence of both algorithms and develop parameter selection strategies to accelerate convergence, thus reducing the communication burden. The proposed methods are validated on a real-world multi-output wind dataset, and their convergence behavior is examined across communication graphs with varying connectivity. Numerical experiments demonstrate that ADMM-RGP and PDMM-RGP can significantly reduce communication relative to the state of the art, while maintaining comparable estimation accuracy and network-wide consensus.
Chinese Translation
高斯过程(Gaussian Processes, GPs)为从含噪测量中学习未知函数提供了灵活的框架,同时能够量化预测的不确定性,使其非常适合多智能体系统中的估计任务。然而,当测量数据由多个智能体采集时,在不依赖集中式处理的情况下维护统一的GP模型,需要高效的分布式算法,这些算法能够仅利用本地测量以及与相邻智能体的通信来运行。在本工作中,我们针对多输出高斯过程回归开发了两种分布式递归GP(RGP)算法:ADMM-RGP和PDMM-RGP。我们分析了这两种算法的稳定性和收敛性,并开发了加速收敛的参数选择策略,从而降低通信负担。所提出的方法在真实世界的多输出风速数据集上进行了验证,并在不同连通性的通信图上考察了其收敛行为。数值实验表明,ADMM-RGP和PDMM-RGP相对于现有最先进方法能够显著减少通信量,同时保持相当的估计精度和全网一致性。
cs.LG / 18 / 2609.27018

GeoRVQ: Decoder-aware geometry for residual-token prediction in physiological signals

GeoRVQ:面向生理信号残差token预测的解码器感知几何方法
Cui, Bo, Zhang, Yaowen
Abstract
Residual vector quantization (RVQ) turns physiological waveforms into compact token sequences, but conventional masked modeling treats every incorrect token as equally costly. We propose GeoRVQ, a coarse-to-fine masked token model whose objective reflects the local response of a frozen waveform decoder. Decoder-induced costs define geometry-aware soft targets and expected distortion, while quantizer-causal prediction follows residual dependencies from coarse to fine levels. In a descriptive aggregate over MIMIC-IV Waveform, VitalDB, and CODE-15\%, GeoRVQ increases exact token accuracy from $.133\pm.004$ to $.143\pm.003$, reduces decoded distance from $.606\pm.006$ to $.393\pm.007$, and increases R-peak F1 from $.784\pm.004$ to $.837\pm.008$ under matched model and training conditions. Across 45 held-out code substitutions, decoder-induced cost has a Spearman correlation of $.85$ with realized decoded cost, compared with $.54$ for Euclidean codeword distance. These results indicate that decoder-aware objectives can improve waveform and event preservation without requiring a large increase in exact token accuracy.
Chinese Translation
残差向量量化(RVQ)将生理波形转换为紧凑的token序列,但传统的掩码建模将所有错误的token视为同等代价。我们提出了GeoRVQ,这是一种由粗到细的掩码token模型,其优化目标反映了冻结波形解码器的局部响应。由解码器诱导的代价定义了几何感知的软目标与期望失真,同时量化器因果预测沿残差依赖关系从粗到细逐级进行。在MIMIC-IV Waveform、VitalDB和CODE-15%数据集上的汇总描述性统计中,在模型和训练条件匹配的情况下,GeoRVQ将token精确匹配率从 $.133\pm.004$ 提升至 $.143\pm.003$,将解码距离从 $.606\pm.006$ 降低至 $.393\pm.007$,并将R波检测的F1分数从 $.784\pm.004$ 提升至 $.837\pm.008$。在45次保留码字替换实验中,解码器诱导的代价与实际解码代价的Spearman相关系数为 $.85$,而欧氏码字距离仅为 $.54$。这些结果表明,解码器感知的优化目标可以在不需要大幅提升token精确匹配率的情况下,改善波形与事件的保留效果。
cs.LG / 19 / 2609.27033

WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps

WTF?!基于Wasserstein倾斜流映射的无仿真强化学习
Mammadov, Abbas, Huang, Jerry Y., Lin, Justin, Kaushik, Partha, Shah, Sheel, Nair, Kartik, Teh, Yee Whye, Boffi, Nicholas M.
Abstract
Reward fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically formulate this problem as sampling from a reward-tilted distribution, the solution to a KL-regularized reward-maximization problem. Here, we introduce an optimal transport regularizer built directly from the pre-trained drift. Unlike KL reward tilting, the resulting objective transports individual samples toward higher reward rather than reweighting the base distribution. We show that the resulting problem is equivalent to a deterministic optimal control problem on the flow. Given a pre-trained flow map, this equivalence yields a simulation-free reinforcement learning algorithm for fine-tuning generative flows. We call the resulting framework Wasserstein-Tilted Flow Maps (WTF), the first end-to-end fine-tuning recipe native to flow maps. The output is a fine-tuned flow map that retains strong reward-aligned performance at few-step inference budgets without post-hoc distillation. Experiments on ImageNet-256 and text-to-image show that WTF achieves higher reward with comparable or higher diversity than baselines, while requiring up to $280\times$ less training compute. More broadly, we argue that accelerated samplers such as flow maps are essential infrastructure for efficient post-training, and that the dominant KL-regularized formulation is only one of many choices worth revisiting.
Chinese Translation
奖励微调旨在更新预训练的基于流的生成模型,以提升其生成样本的下游奖励。现有方法通常将该问题表述为从奖励倾斜分布中采样,即KL正则化奖励最大化问题的解。在此,我们引入一种直接基于预训练漂移构建的最优传输正则化项。与KL奖励倾斜不同,所得目标是将单个样本向更高奖励方向传输,而非对基础分布进行重加权。我们证明所得到的问题等价于流上的确定性最优控制问题。给定一个预训练的流映射(flow map),该等价性产生一种无仿真(simulation-free)的强化学习算法,用于微调生成流。我们将所得框架称为Wasserstein倾斜流映射(Wasserstein-Tilted Flow Maps, WTF),这是首个原生于流映射的端到端微调方案。其输出是一个微调后的流映射,在少步推理预算下无需事后蒸馏即可保持优异的奖励对齐性能。在ImageNet-256和文生图任务上的实验表明,WTF在相当或更高多样性的情况下取得了比基线更高的奖励,同时所需训练计算量最多减少280倍。更广泛地,我们认为诸如流映射之类的加速采样器是高效后训练的关键基础设施,而主流的KL正则化表述只是众多值得重新审视的选择之一。
cs.LG / 20 / 2609.27036

An open benchmark for machine learning-based polymer property prediction

一个面向基于机器学习的聚合物性质预测的开放基准
Learsch, Robert W., Liesen, Nicholas, Levine, Daniel S., Hiszpanski, Anna M., Antoniuk, Evan R.
Abstract
Polymer property prediction lacks open, standardized benchmarks that enable rigorous comparison of machine-learning methods, with existing resources covering only a narrow fraction of polymer architectures, such as homopolymers. We introduce Polymer Benchmark 2026 (PolyBench26), an open dataset comprising nearly 250,000 polymer-property datapoints across eight physical properties, including data from experimental measurements, density functional theory, and molecular dynamics. The benchmark supports four evaluation tasks across homopolymers and alternating, random, and block copolymers: in-distribution property prediction, dataset-size scaling, repeat-unit complexity, and transfer to held-out polymer architectures. We compare language model, graph-based, and descriptor-based approaches and find graph-based models provide the lowest errors in property prediction, retain their advantage across the evaluated training-set sizes, and remain robust to increasing repeat-unit complexity. PolyBench26 provides a reproducible foundation for developing models for the increasingly complex polymer design space. The PolyBench26 benchmark is available open-source at https://github.com/rlearsch/PolymerBenchmark2026.
Chinese Translation
聚合物性质预测领域缺乏开放、标准化的基准,难以对机器学习方法进行严格的比较,且现有资源仅覆盖聚合物结构中较小的一部分,如均聚物。我们提出了Polymer Benchmark 2026(PolyBench26),这是一个开放数据集,包含近25万个聚合物-性质数据点,涵盖八种物理性质,数据来源包括实验测量、密度泛函理论(DFT)和分子动力学模拟。该基准支持均聚物以及交替、无规和嵌段共聚物的四类评估任务:分布内性质预测、数据集规模扩展、重复单元复杂度以及向保留聚合物结构的迁移。我们比较了语言模型、基于图的方法和基于描述符的方法,发现基于图的模型在性质预测中误差最低,在所评估的训练集规模下保持其优势,并对日益增加的重复单元复杂度保持稳健。PolyBench26为日益复杂的聚合物设计空间中开发模型提供了可复现的基础。PolyBench26基准已开源发布于 https://github.com/rlearsch/PolymerBenchmark2026 。
cs.LG / 21 / 2609.27067

ChipMEM: Verification-Grounded Memory for EDA Agents

ChipMEM:面向EDA代理的基于验证的记忆机制
AlRabah, Abdulrahman, Mabry, Joshua, Hakkani-Tür, Dilek, Alawini, Abdussalam, Shojaei, Hamid, Hegde, Kartik, Adhikary, Sandesh
Abstract
Large language model (LLM)-based agents use Electronic Design Automation (EDA) tools to generate and revise register-transfer-level (RTL) designs under synthesis and verification feedback. Recent methods learn from this feedback by distilling reusable skills from execution traces or by training on rewards derived from EDA-tools. Both methods are typically evaluated on the tasks that produced the experience. Repeated access to benchmark feedback on the same task can reward task-specific revision rather than creating reusable knowledge that transfers. We introduce ChipMEM, a verification-grounded memory layer for EDA agents. It combines cross-task procedural memory with within-trajectory statistical guidance. Its procedural component distills and stores a skill only after it passes synthesis, simulation, or formal checks, rather than relying on model self-assessments. A Bayesian component maintains hierarchical Beta estimates over tool-call outcomes and ranks recovery strategies that succeeded under comparable errors. A common adapter applies the same memory interface to RTL optimization and testbench-generation agents while preserving each domain's tools and acceptance criteria. We measure performance on training tasks and evaluate whether learned skills transfer to unseen tasks. On RTLRewriter-Bench, under matched model and tool settings, ChipMEM produces equivalence-passing outputs on 39/54 scored designs versus 35/54 without memory; on the 49-design short suite, mean area improvement is 8.69% versus 5.66%. On held-out CVDP tasks, ChipMEM with a frozen procedural library achieves 20/20 accepted outcomes versus 18/20 without memory in a single evaluation per setting.
Chinese Translation
基于大语言模型(LLM)的代理利用电子设计自动化(EDA)工具,在综合与验证反馈下生成和修改寄存器传输级(RTL)设计。近期方法通过从执行轨迹中蒸馏可复用技能,或基于EDA工具导出的奖励进行训练来学习此类反馈。然而,这两种方法通常都在产生经验的相同任务上进行评估。在同一任务上反复访问基准反馈,可能会奖励针对特定任务的修改,而非形成可迁移的可复用知识。我们提出了ChipMEM,一种面向EDA代理的基于验证的记忆层。它将跨任务的过程性记忆与轨迹内统计引导相结合。其过程性组件只有在技能通过综合、仿真或形式化检查后才进行蒸馏和存储,而非依赖模型自评估。贝叶斯组件对工具调用结果维护分层的Beta分布估计,并对在相似错误下成功过的恢复策略进行排序。通过一个通用适配器,同一记忆接口可同时应用于RTL优化和测试平台(testbench)生成代理,同时保留各领域自身的工具和验收标准。我们在训练任务上测量性能,并评估所学技能能否迁移到未见过的任务。在RTLRewriter-Bench上,在相同的模型和工具设置下,ChipMEM在39/54个评分设计上产生了通过等价性验证的输出,而无记忆的基线为35/54;在包含49个设计的短测试集上,平均面积改进为8.69%,而无记忆基线为5.66%。在留出的CVDP任务上,使用冻结过程性库的ChipMEM在每种设置仅进行一次评估的条件下,实现了20/20的接受结果,而无记忆基线为18/20。
cs.LG / 22 / 2609.27092

Local Evidence and Geometric Readout Repair in Trained GNNs

训练后的图神经网络中的局部证据与几何读出修复
Tomeh, Nadi, Attali, Hugo
Abstract
Many node-classification GNNs apply a linear classifier to a nonnegative mixture of local messages. An error can reflect either poor mixture weights or a reachable logit set poorly positioned for the classifier. We separate these causes with an exact-mass linear program and two learned post-hoc repairs. Every reweighted prediction has an equivalent centered logit translation, but only translations in a message-induced displacement set are realizable by reweighting. Across eight datasets, eight GNN backbones, and ten splits, mean accuracy rises from 62.6% for the frozen models to 63.8% with reweighting and 65.3% with set-conditioned translation. A parameter-matched node-only translator reaches 64.6%, showing that translation explains most of the gain while the message set supplies a smaller additional benefit. Although oracle reweighting can correct many errors, label-free reweighting captures little of this potential: local evidence is often present but hard to select, and relaxing the evidence constraint is more effective than learning within it.
Chinese Translation
许多节点分类图神经网络(GNN)将线性分类器应用于局部消息的非负混合。预测误差可能源于不佳的混合权重,也可能源于可达的logit集合相对于分类器的位置不利。我们通过一个精确质量的线性规划以及两种学习型事后修复方法来区分这两种原因。每个重加权预测都对应一个等价的中心化logit平移,但只有在由消息诱导的位移集合中的平移才能通过重加权实现。在八个数据集、八种GNN骨干网络和十种数据划分上,冻结模型的平均准确率为62.6%,重加权后提升至63.8%,采用基于集合条件的平移后达到65.3%。一个参数量匹配的仅节点平移器达到64.6%,表明平移解释了大部分性能提升,而消息集合带来了较小的额外收益。尽管oracle重加权可以纠正许多错误,但无标签重加权几乎无法实现这一潜力:局部证据往往存在但难以选取,放松证据约束比在约束内学习更为有效。
cs.LG / 23 / 2609.27144

Learning Risk Scores Robust to Unobserved Confounders

学习对未观测混杂因素具有鲁棒性的风险评分
Edmonds, Ryan, Ye, Yingxiao, Aghaei, Sina, Gómez, Andrés, Koçyiğit, Çağıl, Vayanos, Phebe
Abstract
We consider the problem of learning risk scores to prioritize individuals for scarce resources or interventions, from historical observational data affected by unobserved confounding. Decisions about who receives scarce resources are often guided by risk scores based on recorded characteristics, such as responses to a survey. These risk scores are increasingly being learned directly from observational data: historical records of individuals' characteristics, allocation decisions, and outcomes. Standard methods such as inverse propensity weighting (IPW), which corrects for the bias introduced by the historical allocation policy, can be used to learn accurate risk scores if the historical decision process is fully explained by the recorded characteristics. In practice, however, historical decisions often depend on unrecorded information, causing learned risk scores to systematically under-prioritize exactly the individuals whose unrecorded circumstances drove past prioritization. We propose a method for learning risk scores that are robust to this kind of unobserved confounding, building on IPW. Since propensity weights cannot be reliably estimated under unobserved confounding, we instead treat them as belonging to an uncertainty set determined by the observable data and domain-informed estimates of the degree of confounding, combining sensitivity analysis from causal inference with Wasserstein distributionally robust optimization. The resulting robust risk score learning problem admits a sample-based approximation that we reformulate as an exponential cone program compatible with off-the-shelf solvers. We demonstrate the effectiveness of our approach on semi-synthetic data derived from datasets in the UCI Machine Learning Repository. Our method improves calibration by up to 29.2% over traditional benchmarks and up to 11.1% over the state of the art, without compromising other metrics.
Chinese Translation
我们研究如何从受未观测混杂因素影响的历史观测数据中学习风险评分,以便对稀缺资源或干预措施的目标人群进行优先级排序。关于谁应当获得稀缺资源的决策,通常基于已记录特征(如问卷回答)的风险评分来引导。这些风险评分越来越多地直接从观测数据——即个体特征、分配决策和结果的历史记录——中学习得到。如果历史决策过程能够被已记录特征完全解释,那么诸如逆倾向得分加权(IPW)等用于纠正历史分配策略所引入偏差的标准方法,可以用来学习准确的风险评分。然而在实践中,历史决策往往依赖于未记录的信息,这导致学习到的风险评分系统性地低估那些恰恰因未记录处境而在过去被优先考虑的个体。我们在IPW的基础上,提出了一种对这类未观测混杂因素具有鲁棒性的风险评分学习方法。由于在未观测混杂存在时倾向得分权重无法被可靠估计,我们转而将其视为属于一个不确定性集合,该集合由可观测数据以及结合领域知识对混杂程度的估计所确定,从而将因果推断中的敏感性分析与Wasserstein分布鲁棒优化相结合。所得的鲁棒风险评分学习问题可以转化为基于样本的近似形式,我们将其重新表述为可与现成求解器兼容的指数锥规划问题。我们在基于UCI机器学习仓库数据集构建的半合成数据上验证了该方法的有效性。与传统的基准方法相比,我们的方法将校准性能最多提升29.2%,与最先进方法相比最多提升11.1%,同时不影响其他指标。
cs.LG / 24 / 2609.27158

The Linear Representation Hypothesis Needs a Group Action

线性表示假设需要一个群作用
Yao, Louie Hong, Li, Yuhao, Liu, Shengchao
Abstract
To make claims about representations that generalize beyond a particular trained model, we need to specify when two representations should count as equivalent. The Linear Representation Hypothesis is often discussed without making this equivalence explicit. Different notions of equivalence preserve different structures, so metrics, probes, and interventions that appear to study the same representation may in fact correspond to different hypotheses. We therefore argue that the Linear Representation Hypothesis is not one hypothesis but a family of claims distinguished by representation equivalence. We formalize this idea using group actions, specifying the representation object, the procedure that produces it, and the property ultimately asserted, while accounting for equivalences imposed by the model architecture. This framework clarifies how assumptions can change across metrics, reading points, and analysis stages, and we use it to audit common representation quantities and recent interpretability analyses.
Chinese Translation
为了对表示做出超越特定训练模型的普适性论断,我们需要明确两个表示在何种情况下应被视为等价。线性表示假设(Linear Representation Hypothesis)在讨论时往往没有明确这一等价关系。不同的等价概念保留不同的结构,因此看似研究同一表示的度量方法、探测器和干预手段,实际上可能对应着不同的假设。因此,我们认为线性表示假设并非单一假设,而是由表示等价性所区分的一系列论断。我们利用群作用(group action)将这一思想形式化,明确表示对象、产生该表示的过程以及最终所断言的性质,同时考虑由模型架构所施加的等价性。该框架阐明了假设如何在不同的度量方法、读取点和分析阶段之间发生变化,并利用它对常见的表示量以及近期的可解释性分析进行了审查。
cs.LG / 25 / 2609.27166

Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models

大型推理模型中推理时能力与效率的规模化规律
Laber, Moritz, Shafi, Zohair, Savcisens, Germans, Klein, Brennan, Chinazzi, Matteo, Scarpino, Samuel V., Barabási, Albert-László, Eliassi-Rad, Tina
Abstract
Capability and efficiency are two key dimensions of reasoning in large language models (LLMs). Capability refers to the ability to solve a given problem correctly, whereas efficiency refers to the ability to do so with limited resources. When LLMs use Chain-of-Thought (CoT) reasoning to solve problems of controlled hardness, both the number of problems solved correctly and the number of tokens required to reach a correct answer depend on problem hardness and model size. However, how these factors jointly shape capability and efficiency remains poorly understood. Here, we use hierarchical Bayesian models to evaluate the capability and efficiency of LLMs from the DeepSeek-R1-Distill model family across four classes of arithmetic and algorithmic reasoning problems. At a fixed model size, the probability of correctly solving an instance decays approximately exponentially with instance size, our proxy for problem hardness. The decay scale grows sublinearly with model size, indicating that larger models are more capable, but that capability gains diminish with scale. Output length grows as a power law with instance size, which serves as a proxy for difficulty. However, the parameters of this power law do not vary systematically with model size, suggesting that larger models do not become more efficient. Together, these findings reveal potential limitations of naive scaling as a strategy for developing more capable AI systems: capability improves with diminishing returns, while efficiency shows little to no improvement.
Chinese Translation
能力与效率是大型语言模型(LLM)推理的两个关键维度。能力指正确解决给定问题的能力,而效率指在有限资源下做到这一点的能力。当LLM使用思维链(Chain-of-Thought, CoT)推理来解决难度受控的问题时,正确解决的问题数量以及得出正确答案所需的token数量均取决于问题难度和模型规模。然而,这些因素如何共同塑造能力与效率仍知之甚少。本文使用层次贝叶斯模型评估了DeepSeek-R1-Distill模型系列LLM在四类算术与算法推理问题上的能力与效率。在固定模型规模下,正确解决某个实例的概率随实例规模(我们的问题难度代理指标)近似呈指数衰减。衰减尺度随模型规模呈亚线性增长,表明更大的模型能力更强,但能力增益随规模扩大而递减。输出长度随实例规模(作为难度的代理指标)呈幂律增长,但该幂律的参数并不随模型规模系统变化,说明更大的模型并未变得更高效。综上,这些发现揭示了单纯规模化作为开发更强AI系统策略的潜在局限:能力提升呈现收益递减,而效率几乎没有改善。
cs.LG / 26 / 2609.27186

Data-driven discrete-time deep recurrent neural network-based modeling for dissipative systems

基于数据驱动的离散时间深度循环神经网络的耗散系统建模
Luong, Tuan, Moon, Hyungpil
Abstract
Physical AI has gained increasing attention for its role in developing AI systems that better understand, predict, and control real-world dynamics. Achieving this requires AI models that not only achieve high prediction accuracy but also preserve fundamental physical properties of dynamical systems. In this paper, we propose a deep discrete-time dissipative recurrent neural network (DissipNet) that explicitly enforces dissipativity, a key property related to stability and energy dissipation, through structural weight constraints and a dedicated training algorithm. By construction, the proposed network is capable of learning dissipative dynamics while preserving their inherent stability, which is formally analyzed using Lyapunov theory. In contrast to Physics-Informed Neural Networks (PINNs), which incorporate governing equations into the training loss but do not guarantee preservation of internal analytical properties such as dissipativity or passivity, our approach provides explicit guarantees on stability at the model level. We demonstrate the effectiveness of the proposed method through several modeling applications, and compare its performance with a naive recurrent neural network (RNN) and a PINN-based model.
Chinese Translation
物理人工智能(Physical AI)因其在开发能够更好地理解、预测和控制现实世界动力学的AI系统方面的作用而日益受到关注。实现这一目标需要AI模型不仅具有高预测精度,还能保持动力系统的基本物理属性。本文提出了一种深度离散时间耗散循环神经网络(DissipNet),通过结构性权重约束和专门的训练算法显式地保证耗散性——这一与稳定性和能量耗散相关的关键特性。通过构造方式,所提出的网络能够在保持耗散动力学固有稳定性的同时学习其动态特性,并利用Lyapunov理论进行了严格分析。与物理信息神经网络(PINNs)不同,后者虽然将控制方程纳入训练损失中,但无法保证耗散性或无源性等内部解析特性的保持,而我们的方法在模型层面提供了稳定性的显式保证。我们通过多个建模应用验证了所提方法的有效性,并将其性能与朴素循环神经网络(RNN)和基于PINN的模型进行了比较。
cs.LG / 27 / 2609.27199

ZO-COSMO: Index-Free One-Hop Mixing for Decentralized Zeroth-Order Optimization

ZO-COSMO:面向去中心化零阶优化的无索引单跳混合方法
Zhang, Shengjun, Liu, Tingyi, Zhang, Heng, Xie, Dong
Abstract
Sparse communication in decentralized zeroth-order learning requires compatible peer-state coordinates. We characterize this one-hop condition and develop \textsf{ZO-COSMO}, coupling two-query estimation with average-preserving masked consensus using $q$ values per active link. Global supports serve all-neighbor mixing; matching updates require agreement only within each pair. We derive a sharp contraction-per-scalar bound within the matching class and convergence guarantees for the core and sparse-momentum updates. At fixed matching, exact moment identities characterize how shared directions preserve gradient-heterogeneity cancellation and redistribute estimation error and disagreement. Mechanism experiments cover unequal curvatures, noise, and sparse momentum. Further tests span $64$ synthetic agents and eight logical Qwen LoRA workers. At matched payload budgets, Qwen2-7B QNLI gains $3.65$ accuracy points over explicit-index Rand-$k$; edge-local updates gain $3.42$ and $2.53$ points over all-neighbor mixing on eight-worker complete and ring graphs. A matched-first-step ablation gives a $3.92$-point momentum benefit. Seed-aware and same-matching controls distinguish encoding, scheduling, and query correlation.
Chinese Translation
去中心化零阶学习中的稀疏通信需要兼容的对等节点状态坐标。我们刻画了这一单跳条件,并提出了 ZO-COSMO 方法,将双查询估计与保均值的掩码共识相结合,每条活跃链路仅使用 q 个数值。全局支撑集用于面向所有邻居的混合;而匹配式更新只要求每对节点内部达成一致。我们在匹配类中推导了紧凑的每标量收缩界,并为核心更新和稀疏动量更新给出了收敛性保证。在固定匹配条件下,精确的矩恒等式刻画了共享方向如何保持对梯度异质性的抵消作用,并重新分配估计误差与分歧。机制实验涵盖了不均衡曲率、噪声以及稀疏动量等情形。进一步的实验涵盖 64 个合成智能体和八个逻辑 Qwen LoRA 工作节点。在相同负载预算下,Qwen2-7B 的 QNLI 任务比基于显式索引的 Rand-k 方法高出 3.65 个准确率点;在八工作节点的全连接图和环形图上,边局部更新比面向所有邻居的混合方法分别高出 3.42 和 2.53 个点。匹配首步消融实验显示动量带来 3.92 个点的收益。基于随机种子感知和相同匹配的对照实验区分了编码、调度与查询相关性各自的作用。
cs.LG / 28 / 2609.27201

A Systematic Benchmark of Explainable Methods for Temporal Attribution in Sequential Recommendation Systems

序列推荐系统中时间归因可解释方法的系统性基准测试
Pandey, Akash, Shah, Kanisha, Roy, Addrish, Katariya, Dwipam, Shi, Hongyangyang, Ding, Amanda, Mishra, Kalanand, Mohanty, Pranab
Abstract
Sequential RecSys are central to modern personalization, exploiting user's historical interaction sequences to drive next-step decisions. Deep learning models, particularly CNN and Transformer-based architectures, have proven highly effective at capturing temporal dependencies in these histories. For transparency and trust, understanding which past interactions drive a given recommendation is increasingly important --- both for developers auditing model behavior and for users seeking a rationale. However, the non-linearities that give these models their predictive power also render them black boxes, making it difficult to attribute decisions to specific interactions. While gradient-based, perturbation-based, and attention-based explainability methods exist, a systematic benchmark of their faithfulness for sequential recommendation is missing. We address this gap by introducing a dual-model masking metric in which one model supplies per-timestep attribution scores and a separately trained, masking-robust probe measures the resulting change in predicted probability. Using this metric, we benchmark ten XAI methods across CNN, Transformer, SASRec, and BERT4Rec backbones on KuaiRand and MovieLens, complemented by analyses of temporal attribution patterns, item popularity confounding, and robustness to input corruption. Our key findings are: (1) gradient-based methods, particularly GradientSHAP and Integrated Gradients, yield the most faithful and robust attributions; (2) raw attention weights are unreliable, but gradient-weighted attention restores faithfulness on shorter sequences, with degradation on longer horizons as softmax attention probabilities converge toward uniform importance scores, diminishing the method's ability to identify informative interactions; and (3) temporal attribution patterns in faithful methods reflect genuine task structure rather than recency or popularity bias.
Chinese Translation
序列推荐系统(Sequential RecSys)是现代个性化服务的核心,它利用用户的历史交互序列来驱动下一步决策。深度学习模型,尤其是基于CNN和Transformer的架构,已被证明在捕捉这些历史序列中的时间依赖性方面非常有效。为了透明性与可信度,理解哪些历史交互驱动了某条给定的推荐变得日益重要——无论对于审计模型行为的开发者,还是寻求推荐理由的用户而言都是如此。然而,赋予这些模型强大预测能力的非线性特性也使其成为黑箱,难以将决策归因于具体的交互。尽管目前存在基于梯度、基于扰动和基于注意力的可解释性方法,但针对序列推荐对这些方法忠实性的系统性基准测试仍然缺失。为填补这一空白,我们提出了一种双模型掩码度量方法:其中一个模型提供每个时间步的归因分数,另一个经过单独训练、对掩码具有鲁棒性的探测模型(probe)则度量预测概率由此产生的变化。基于该度量,我们在KuaiRand和MovieLens数据集上,针对CNN、Transformer、SASRec和BERT4Rec等骨干网络,对十种XAI方法进行了基准测试,并对时间归因模式、物品流行度混淆效应以及对输入扰动的鲁棒性进行了补充分析。我们的主要发现如下:(1)基于梯度的方法,尤其是GradientSHAP和Integrated Gradients,能够产生最忠实且最鲁棒的归因;(2)原始注意力权重并不可靠,但梯度加权注意力在较短序列上能够恢复忠实性,而在较长序列上则出现性能退化,这是因为softmax注意力概率趋近于均匀的重要性分数,从而削弱了该方法识别有信息量交互的能力;(3)忠实方法所呈现的时间归因模式反映的是真实的任务结构,而非时近效应或流行度偏差。
cs.LG / 29 / 2609.27209

Scalable Subgraph Sampling via Resistance Curvature

基于电阻曲率的可扩展子图采样方法
Fei, Chaoqun, Zhou, Tinglve, Hao, Tianyong, Li, Yangyang
Abstract
Subgraph sampling reduces the training cost of large-scale graph neural networks, but sampling criteria may overlook the geometric roles of edges. We propose a resistance-curvature-guided sampling framework built on ERC-LG, a curvature approximation method for large-scale graphs. ERC-LG combines Johnson-Lindenstrauss projections with regularized multi-GPU batched conjugate gradient solvers, avoiding explicit Laplacian pseudoinverse computation and full embedding storage. The resulting curvature informs node- and edge-sampling probabilities for constructing GNN training subgraphs. Experiments show numerical agreement with pseudoinverse-based curvature and reduced runtime compared with CG-only computation. ERC-LG-based sampling variants achieve the highest mean accuracy on six of seven real-world datasets in downstream node classification.
Chinese Translation
子图采样可以降低大规模图神经网络的训练成本,但现有的采样准则可能忽视边的几何作用。我们提出了一种由电阻曲率引导的采样框架,该框架构建于 ERC-LG 之上——一种面向大规模图的曲率近似方法。ERC-LG 将 Johnson-Lindenstrauss 投影与正则化的多 GPU 批量共轭梯度求解器相结合,避免了显式的拉普拉斯伪逆计算和完整的嵌入存储。所得的曲率信息用于确定节点和边的采样概率,以构建用于 GNN 训练的子图。实验表明,该方法与基于伪逆的曲率计算在数值上一致,且相比仅使用共轭梯度法的计算方式显著降低了运行时间。基于 ERC-LG 的采样变体在七个真实世界数据集中的六个上,于下游节点分类任务中取得了最高的平均准确率。
cs.LG / 30 / 2609.27221

Tail-Aware Geometry Learning for Conformal Ellipsoids

面向保形椭球的尾部感知几何学习
Zhang, Xiang
Abstract
This paper studies multivariate conformal prediction (CP), a distribution-free uncertainty quantification framework with finite-sample coverage guarantees. The efficiency of multivariate prediction sets hinges critically on the residual geometry encoded by the nonconformity score, while existing minimum-volume methods rely on quantile thresholds that ignore tail residual severity and implicitly bind geometry learning to coverage level. We propose a tail-aware geometry learning framework for conformal ellipsoids that decouples tail sensitivity in geometry learning from the final coverage guarantee. Using a two-split design, we learn the metric matrix via volume minimization under a CVaR constraint on an estimation split, then apply standard conformal calibration on a held-out calibration split. The resulting problem is convex and admits a bounded-reweighting interpretation that prioritizes high-residual samples. Moreover, we theoretically characterize the trade-off between ellipsoidal volume and tail severity. Experimental results demonstrate the effectiveness of the proposed method.
Chinese Translation
本文研究多元保形预测(Conformal Prediction, CP),这是一种具有有限样本覆盖率保证的无分布不确定性量化框架。多元预测集的效率在很大程度上取决于由非一致性分数(nonconformity score)所编码的残差几何结构,而现有的最小体积方法依赖于分位数阈值,忽略了尾部残差的严重程度,并隐式地将几何学习与覆盖率水平绑定在一起。我们提出了一种面向保形椭球的尾部感知几何学习框架,将几何学习中的尾部敏感性与最终覆盖率保证解耦。采用双数据集划分设计,我们首先在估计集上通过体积最小化并结合CVaR约束学习度量矩阵,然后在留出的校准集上应用标准的保形校准。由此得到的问题是一个凸优化问题,并且具有有界重加权解释,即优先考虑高残差样本。此外,我们从理论上刻画了椭球体积与尾部严重程度之间的权衡关系。实验结果验证了所提方法的有效性。
cs.LG / 31 / 2609.27232

A Scaling Study for fMRI Foundation Models

fMRI基础模型的缩放研究
Ye, Wenhao, Pan, Xuanye, Xia, Junfeng, Zhang, Junxiang, Wang, Mo, Liu, Quanying
Abstract
Scaling laws have guided large-model development in computer vision and natural language processing, but the relationships among data, model size, and compute remain unclear for functional magnetic resonance imaging (fMRI) foundation models. Here, we conduct a controlled empirical study using pretraining data from more than 200 source datasets and over 10,000 GPU-hours of experiments. Holding the pretraining framework and downstream protocol fixed, we vary pretraining data size, model size, and training duration. Downstream performance generally improves with compute, yet models using similar compute can perform substantially differently. Additional pretraining data bring larger gains at larger model sizes, suggesting that data and model size should be scaled together. At matched compute, increasing pretraining data benefits more tasks than increasing model size, although the pattern varies across tasks. We then use in-distribution (ID) downstream performance to select the combination of pretraining data size, model size, and training duration at two fixed compute budgets. The resulting models are locked before out-of-distribution (OOD) evaluation. They achieve the highest average performance across the evaluated OOD tasks among the compared fMRI foundation models while using less pretraining compute. Overall, our results show that compute alone does not characterize fMRI scaling: performance depends on how pretraining data, model size, and training duration are combined.
Chinese Translation
缩放定律(scaling laws)已指导了计算机视觉和自然语言处理领域大模型的开发,但对于功能磁共振成像(fMRI)基础模型而言,数据、模型规模与计算量之间的关系仍不明确。本文开展了一项受控的实证研究,使用了来自200多个源数据集的预训练数据,并进行了超过10,000 GPU小时的实验。在保持预训练框架和下游评测协议不变的前提下,我们改变了预训练数据规模、模型规模和训练时长。下游性能总体随计算量增加而提升,但使用相近计算量的模型表现可能差异显著。在更大的模型规模下,增加预训练数据带来更大的收益,这表明数据规模与模型规模应当同步扩展。在计算量匹配的情况下,增加预训练数据比增加模型规模能使更多任务受益,尽管这一规律在不同任务间存在差异。随后,我们利用同分布(ID)下游性能,在两个固定计算预算下选择预训练数据规模、模型规模与训练时长的最优组合。所得模型在分布外(OOD)评测之前被固定,在对比的各fMRI基础模型中,它们以更少的预训练计算量在所评测的OOD任务上取得了最高的平均性能。总体而言,我们的结果表明,仅凭计算量无法刻画fMRI的缩放规律:性能取决于预训练数据、模型规模和训练时长之间的组合方式。
cs.LG / 32 / 2609.27234

Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

发现、证伪、修正:从源代码到预测贡献,审计智能体发现细胞模型中的输入使用声明
Li, Mengran, Li, Bo, Zhang, Chengyang, Yan, Yang, Xu, Jinfeng, Tang, Zhenchao
Abstract
AI virtual cells aim to predict cellular responses to specified interventions, yet held-out predictive performance alone does not establish use of the supplied perturbation information. This prediction-claim gap matters in agentic model discovery, where language-model agents generate and revise predictors using score-based feedback. We introduce CELLAUDIT, which audits input-use claims by asking whether an input can enter the cited computation, whether fitted predictions depend on it, and whether that dependence improves prediction of observed response. On a paired morphology-transcriptomics perturbation benchmark (BBBC047), an agent-selected predictor attains a mean held-out Global Pearson correlation coefficient (PCC) of 0.3153 but remains invariant to compound replacement; a control-profile-only predictor reaches 0.3142. Source inspection identifies a compound-query pathway blocked by singleton key-value attention, and the invariance persists after refitting with disjoint control wells. In a stratified audit of 48 candidates across two linked tasks, 47 change predictions under compound replacement on both held-out folds, but only 20 show target-loss gains with intervals above zero on both folds. On BBBC047, falsification-guided revisions recover positive mean compound contributions while retaining gains over the control-profile-only baseline. In matched sci-Plex searches, audit-enriched feedback yields higher held-out performance and larger mean compound and dose contributions across five trajectories, although paired intervals span zero. Refitting fixed designs on an independently acquired cohort shows predictive generalization need not imply generalization of input-use claims: dose contribution persists, whereas support for compound identity does not. CELLAUDIT adds a falsification layer to agentic model discovery, moving from generate-score-revise toward discover-falsify-revise.
Chinese Translation
AI虚拟细胞旨在预测细胞对特定干预的响应,然而仅凭留出集上的预测性能并不能证实模型确实使用了所提供的扰动信息。这一预测-声明差距在智能体模型发现中尤为重要,因为语言模型智能体基于评分反馈生成并修正预测器。我们提出了CELLAUDIT,通过三个问题审计输入使用声明:输入能否进入被引用的计算过程、拟合后的预测是否依赖于该输入、以及这种依赖是否改善了对观测响应的预测。在一个配对的形态学-转录组学扰动基准(BBBC047)上,智能体选择的预测器达到了0.3153的平均留出集全局皮尔逊相关系数(PCC),但对化合物替换保持不变;而仅使用对照轮廓的预测器达到0.3142。源代码检查发现化合物查询通路被单元素键值注意力阻断,并且在用不相交的对照孔重新拟合后,该不变性依然存在。在跨越两个关联任务、对48个候选模型的分层审计中,有47个在两个留出折上均随化合物替换而改变预测,但仅有20个在两个折上显示出置信区间均大于零的靶标损失增益。在BBBC047上,基于证伪引导的修正恢复了正向的平均化合物贡献,同时保持了对仅使用对照轮廓基线的优势。在匹配的sci-Plex搜索中,经审计丰富后的反馈在五条轨迹上产生了更高的留出集性能以及更大的平均化合物与剂量贡献,尽管成对置信区间跨越零。在独立获取的队列上重新拟合固定设计表明,预测泛化并不必然意味着输入使用声明的泛化:剂量贡献得以保持,而对化合物身份的支持则不然。CELLAUDIT为智能体模型发现增加了一层证伪机制,推动其从生成-评分-修正范式迈向发现-证伪-修正范式。
cs.LG / 33 / 2609.27244

Full-Covariance Smoothing of Bayesian Neural Networks for Online Adaptation

面向在线自适应的贝叶斯神经网络全协方差平滑
Wright, Oren, Jing, Haoming, Shen, Qiaoan, Niinuma, Koichiro, Nakahira, Yorie, Moura, José M. F.
Abstract
A neural network's layers can be treated as time steps of a state-space model, turning Bayesian training into a smoothing problem: a forward pass propagates Gaussian moments through the network, and a backward Rauch--Tung--Striebel pass updates the weight posteriors in closed form. Such methods learn from each observation in a single pass, in an uncertainty-aware manner, and without gradient-based iterations or replay, which makes them well suited for online adaptation and data-efficient learning. Existing smoothing-based methods, however, are restricted to diagonal covariances across activations, discarding correlations between neurons. We overcome this limitation via a cross-covariance identity that enables full-covariance propagation through a network's nonlinear activations. We derive a one-step-per-layer smoother that approximates as Gaussian only each layer's affine output, and that applies both to deterministic systems with noisy observations and to stochastic systems described by output statistics. We demonstrate this method in non-stationary classification, online dynamics learning, and policy adaptation of a vision-language-action model, and find that it is generally more accurate than other smoothing-based methods.
Chinese Translation
神经网络层可被视为状态空间模型的时间步,从而将贝叶斯训练转化为一个平滑问题:前向传播在网络中传递高斯矩,反向的 Rauch--Tung--Striebel 传播则以闭式形式更新权重的后验分布。此类方法能够在单次遍历中以不确定性感知的方式从每个观测中学习,且无需基于梯度的迭代或数据重放,因此非常适合在线自适应与数据高效学习。然而,现有的基于平滑的方法仅限于激活值之间的对角协方差,丢弃了神经元之间的相关性。我们通过一个交叉协方差恒等式克服了这一限制,使得全协方差能够在网络的非线性激活中进行传播。我们推导了一种每层仅需一步的平滑器,它仅将每层的仿射输出近似为高斯分布,并且既适用于带噪声观测的确定性系统,也适用于由输出统计特性描述的随机系统。我们在非平稳分类、在线动力学学习以及视觉-语言-动作模型的策略自适应任务上验证了该方法,发现其精度总体上优于其他基于平滑的方法。
cs.LG / 34 / 2609.27248

Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders

将预训练大语言模型改造为高保真连续文本自编码器
Pathak, Arkanath, Jain, Unnat, Berg, Alexander C.
Abstract
Next-token prediction has enabled highly fluent autoregressive language models, but it represents global structure only indirectly through sequential factorization. In contrast, high-fidelity autoencoders have become a standard primitive in image generation, enabling generative models to operate over continuous latent spaces; text lacks a comparably faithful continuous representation. We propose LLMAE, a method for repurposing a pretrained decoder-only language model as a continuous text autoencoder by exposing an intermediate fixed-length latent bottleneck within its internal activations. Instantiated with a parameter-efficient 270M Gemma 3 model, LLMAE uses structured attention masks, LoRA adaptation, and KL regularization to learn an autoencoding interface that leverages the generative prior of the original LLM. We train LLMAE to reconstruct text sequences up to 1024 tokens, significantly improving on this task to achieve near-perfect reconstruction. Furthermore, we demonstrate the downstream utility of this representation by training a latent text diffusion model for detailed image captioning using the learned LLMAE autoencoder. By mapping text into a fixed-length continuous latent space, our approach provides an effective substrate for downstream adaptation while benefiting from the fluency of the original LLM.
Chinese Translation
下一词元预测(next-token prediction)使自回归语言模型具备了极高的流畅性,但它仅通过序列分解间接地表示全局结构。相比之下,高保真自编码器已成为图像生成中的一种标准基础组件,使生成模型能够在连续潜空间上进行操作;而文本缺乏与之相当的忠实连续表示。我们提出LLMAE,一种将预训练的仅解码器(decoder-only)语言模型改造为连续文本自编码器的方法,其核心是在模型内部激活中引入一个中间固定长度的潜空间瓶颈。LLMAE基于参数高效的270M Gemma 3模型实例化,采用结构化注意力掩码、LoRA适配和KL正则化来学习一个自编码接口,从而充分利用原始大语言模型(LLM)的生成先验。我们训练LLMAE重建长达1024个词元的文本序列,并在该任务上取得显著提升,实现了近乎完美的重建。此外,我们通过利用所学到的LLMAE自编码器训练一个潜空间文本扩散模型来实现精细的图像描述生成,验证了该表示的下游应用价值。通过将文本映射到固定长度的连续潜空间,我们的方法在受益于原始LLM流畅性的同时,为下游适配提供了有效的基础。
cs.LG / 35 / 2609.27252

What Converges in the Platonic Representation Hypothesis? Structure over Geometry

柏拉图表示假说中什么在收敛?结构优于几何
You, Junwon, Jang, Mihyun, Mo, Sangwoo, Jung, Jae-Hun
Abstract
The Platonic Representation Hypothesis suggests that increasingly capable models converge toward shared representations. Recent work narrows this claim to shared local neighborhood relationships, finding that capacity-dependent trends in several global similarity measures largely disappear after calibration. We challenge this interpretation by showing that prior local-global comparisons confound structural scale (local versus global) with what is compared: relational structure, defined by which samples are related, versus metric geometry, characterized by quantitative relations such as distances, similarities, or correlations. To disentangle these factors, we construct a controlled $2\times2$ framework that evaluates both relational structure and metric geometry at local and global scales. We introduce $H_0$ skeleton overlap as a global counterpart to mutual $k$-nearest neighbors, together with matched distance-aware variants. Across vision-language models, relational structure exhibits robust representational convergence at both scales after calibration, whereas increasingly stringent distance agreement substantially weakens alignment and progressively flattens the capacity-dependent trend. We further extend the analysis beyond ambient Euclidean geometry by evaluating distance agreement under a Riemannian metric approximation and recover the same structure-geometry pattern. The pattern is also reproduced in video-text representations. Together, these results show that relational convergence extends beyond local neighborhoods to global spanning structure, whereas metric geometry exhibits substantially weaker convergence.
Chinese Translation
柏拉图表示假说(Platonic Representation Hypothesis)认为,能力日益增强的模型会收敛到共享的表示。近期研究将这一主张缩小为共享的局部邻域关系,并发现在经过校准后,若干全局相似性度量中依赖于模型能力的趋势基本消失。我们对这一解释提出质疑,指出先前的局部-全局比较将结构尺度(局部与全局)与比较对象混淆了:后者指的是由样本之间的关联所定义的关系结构,以及以距离、相似度或相关性等定量关系为特征的度量几何。为了解耦这些因素,我们构建了一个受控的 $2\times2$ 框架,在局部和全局两个尺度上同时评估关系结构与度量几何。我们引入了 $H_0$ 骨架重叠度(skeleton overlap)作为互 $k$ 近邻(mutual $k$-nearest neighbors)的全局对应物,并配套引入了考虑距离的匹配变体。在视觉-语言模型上,经过校准后,关系结构在两个尺度上都表现出稳健的表示收敛性,而要求日益严格的距离一致性则显著削弱了对齐程度,并逐渐抹平了依赖模型能力的趋势。我们进一步将分析拓展到环境欧氏几何之外,在黎曼度量近似下评估距离一致性,并复现了相同的结构-几何模式。该模式同样在视频-文本表示中得到验证。综上,这些结果表明,关系收敛不仅限于局部邻域,还可延伸至全局的支撑结构,而度量几何的收敛性则明显较弱。
cs.LG / 36 / 2609.27278

Graph Learning with Spectral Connectivity Priors for Scarce Data

基于谱连通性先验的稀缺数据图学习
Liu, Mingxiao, Oveisgharan, Bahar, Zou, Bingyan, Cheung, Gene, Zhao, H. Vicky, Gao, Feifei
Abstract
Learning a sparse graph from scarce data is practically important but challenging. Motivated by the desirable combination of local sparsity and strong global connectivity exhibited by expander-like graphs, we propose spectral connectivity-regularized graph learning (SCoGL), a framework that incorporates a family of Laplacian spectral priors to explicitly promote global connectivity. Specifically, SCoGL augments a combinatorial-Laplacian-constrained graphical lasso (GLASSO) objective over a target adjacency matrix $\mathbf{W}$ with a general connectivity prior computed from Laplacian eigenvalues. We derive gradients for several representative connectivity priors and develop a projected gradient descent (PGD) algorithm with Armijo backtracking to efficiently optimize $\mathbf{W}$. Experiments show that the proposed SCoGL variants improve graph recovery and enhance downstream tasks such as graph signal denoising when signal observations are scarce.
Chinese Translation
从稀缺数据中学习稀疏图在实践上十分重要但充满挑战。受扩展图(expander-like graph)所展现的局部稀疏性与强全局连通性这一理想组合的启发,我们提出了谱连通性正则化图学习方法(Spectral Connectivity-regularized Graph Learning, SCoGL),该框架引入一族拉普拉斯谱先验,以显式促进全局连通性。具体而言,SCoGL 在目标邻接矩阵 $\mathbf{W}$ 的组合拉普拉斯约束图套索(graphical lasso, GLASSO)目标函数基础上,增加了由拉普拉斯特征值计算得到的通用连通性先验。我们推导了若干代表性连通性先验的梯度,并开发了一种带 Armijo 回溯的投影梯度下降(PGD)算法来高效优化 $\mathbf{W}$。实验表明,所提出的 SCoGL 变体在信号观测稀缺的情况下改进了图恢复效果,并提升了诸如图信号去噪等下游任务的性能。
cs.LG / 37 / 2609.27287

SR-Fraud: An Outcome-Supervised Reflective LLM Agent Framework for Non-Stationary Payment Fraud Detection

SR-Fraud:一种基于结果监督的反思式大语言模型智能体框架,用于非平稳支付欺诈检测
Tan, Xuwei, Ma, Yao, Zhang, Xueru
Abstract
Real-time payment fraud detection is a non-stationary streaming prediction problem: adversaries adapt before supervised labels mature, and localized burst attacks can cause losses before retraining. Production systems typically rely on tabular classifiers and rules, which can struggle to capture these emerging sequential patterns before periodic retraining occurs. We present SR-Fraud, an outcome-supervised reflective LLM framework that decouples request-time decisions from offline adaptation. A frozen, stateless agent scores each transaction from a Hybrid Episodic Window to track behavioral shifts, while an offline reflection agent proposes boundary hypotheses from matured errors. A deterministic verifier then admits only supported hypotheses into an executable knowledge state. On a production payment-fraud benchmark, SR-Fraud improves all detection metrics over its frozen decision agent, obtains higher point estimates than static and periodically retrained CatBoost, and detects an emerging fraud burst.
Chinese Translation
实时支付欺诈检测是一个非平稳的流式预测问题:攻击者会在监督标签成熟之前进行适应,而局部爆发的攻击可能在模型重新训练之前就造成损失。生产系统通常依赖表格分类器和规则,这些方法难以在周期性重训练发生之前捕捉这些新出现的序列模式。我们提出了 SR-Fraud,一个基于结果监督的反思式大语言模型(LLM)框架,它将请求时的决策与离线适应解耦。一个冻结的、无状态的智能体基于混合情景窗口(Hybrid Episodic Window)对每笔交易进行评分,以追踪行为变化;同时,一个离线反思智能体从已成熟的错误中提出边界假设。随后,一个确定性验证器只允许有依据支持的假设进入可执行的知识状态。在一个生产级支付欺诈基准上,SR-Fraud 相较于其冻结的决策智能体提升了所有检测指标,获得了高于静态 CatBoost 和周期性重训练 CatBoost 的点估计值,并成功检测出一次新出现的欺诈爆发。
cs.LG / 38 / 2609.27291

NGN: Learning Neural Network Size as a Differentiable Count

NGN:将神经网络规模学习为可微分计数
Li, Lixing
Abstract
Neural network size is usually chosen before training, separating architecture selection from weight optimization. We introduce the Neurogenesis Network (NGN), a differentiable parameterization for learning how many ordered structural components a model should use. For each ordered component group, one learnable boundary selects an active prefix while the model parameters are trained. The boundary can grow from a compact initialization and can be deployed by discarding components beyond the learned boundary. Controlled experiments examine convergence of the learned boundary, the performance of deployed prefixes, and comparisons with fixed-size models and alternative approaches to learning capacity. We then apply the same mechanism to MLPs, convolutional and graph networks, Transformers, state-space models, LoRA, and adapters. Across these settings, deploying only the learned prefix usually changes performance little, and the selected architectures perform similarly to fixed models trained at the same size. These results show that structural capacity can be optimized directly as a count.
Chinese Translation
神经网络的规模通常在训练之前预先确定,这使得架构选择与权重优化相互分离。我们提出了神经发生网络(Neurogenesis Network,NGN),这是一种可微分的参数化方法,用于学习一个模型应当使用多少有序的结构组件。对于每个有序的组件组,引入一个可学习的边界,在模型参数训练的同时选择活跃的前缀。该边界可以从紧凑的初始化开始增长,部署时只需丢弃超出所学边界的组件即可。受控实验考察了所学边界的收敛性、部署前缀的性能,以及与固定规模模型和其他容量学习方法的比较。随后,我们将同一机制应用于多层感知机(MLP)、卷积网络和图网络、Transformer、状态空间模型、LoRA 以及适配器。在所有这些设置中,仅部署所学前缀通常对性能影响很小,且所选架构的性能与在相同规模下训练的固定模型相当。这些结果表明,结构容量可以直接作为一种计数进行优化。
cs.LG / 39 / 2609.27294

KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling

KITE:用于高效智能体大语言模型扩展的KV不变Transformer扩展方法
Hu, Zhiheng, Wei, Yixun, Zhou, Jian, Zhou, Yizhuang, Li, Ji, Chen, Xing, Li, Yang, Wang, Bojun, Zhu, Yibo, Zhang, Xiangyu, Jiang, Daxin
Abstract
Scaling a language model is not only a question of final quality: the architectural choice determines how much computation is spent during training, prompt processing, and autoregressive decoding to achieve certain model quality. An ideal model architecture should lower all above computation costs to facilitate scaling to a larger model, while ensure the larger model indeed outperforms smaller baselines. We introduce KV-Invariant Transformer Expansion (KITE), a scaling paradigm that achieves this goal. It trains the model from a smaller size to a larger size (i.e., saving training costs via upcycling), while places newly added parameters in regions that do not affect attention KV. Consequently, during inference, prefilling KV only relies on the smaller part of the model, so the inference costs are saved. As a concrete instantiation, we present Step Scale Transformer (SST), a two-tower decoder in which one tower produces KV and the other reads them. At comparable cumulative training compute, SST, a 67B MoE model with 2.15B active body parameters per decode token, achieves lower training loss than 47B and 63B MoE Transformers with 1.48B and 2.02B active body parameters, respectively, while reducing estimated inference cost by 6.7% and 31.6%.
Chinese Translation
语言模型的扩展不仅仅是最终质量的问题:架构选择决定了在训练、提示处理和自回归解码过程中,为实现特定模型质量所需花费的计算量。理想的模型架构应当降低上述所有计算成本,以便于扩展到更大的模型,同时确保更大的模型确实优于较小的基线模型。我们提出了KV不变Transformer扩展(KV-Invariant Transformer Expansion, KITE),一种能够实现该目标的扩展范式。该方法将模型从较小规模训练扩展到较大规模(即通过升级复用以节省训练成本),同时将新增参数放置在不影响注意力KV(键值缓存)的区域。因此,在推理阶段,预填充KV仅依赖于模型中较小的部分,从而节省了推理成本。作为一个具体的实例,我们提出了阶梯缩放Transformer(Step Scale Transformer, SST),一种双塔解码器架构,其中一个塔生成KV,另一个塔读取KV。在相当的累计训练计算量下,SST——一个每个解码token激活21.5亿主体参数的670亿参数MoE模型——取得了比分别具有14.8亿和20.2亿激活主体参数的470亿和630亿参数MoE Transformer更低训练损失,同时估计推理成本分别降低了6.7%和31.6%。
cs.LG / 40 / 2609.27303

Live Assistant: Learning Whether, When, and Whom to Assist in Real-World Live Social Streams

Live Assistant:在真实直播社交流中学习是否、何时以及向谁提供协助
Gao, Shujian, Yan, Jiamei, Yang, Yuchen, Zhou, Penghao, Wang, Qinglei, Fan, Tiehan, Wang, Yuan, Wu, Zuxuan, Jiang, Yu-gang
Abstract
Livestreams are long-lasting interactive environments where audiovisual content, viewer activity, host behavior, and platform signals evolve together, creating assistance needs that emerge from the stream itself. We introduce \liveassistant, a framework for mixed-initiative, role-conditioned assistance that formulates livestream interaction as four coupled decisions: \textbf{whether to act, when to act, whom to address, and what to communicate}. At each 10-second interval, one autoregressive policy consumes native audio and video with synchronized comments, gifts, viewer dynamics, and room metadata, then selects \textsc{OBS}, \textsc{MEM}, or \textsc{ANS}. \textsc{OBS} remains silent, \textsc{MEM} records a private semantic update, and \textsc{ANS} specifies a recipient, task, and grounded message. To support this task, we build a trajectory engine that reconstructs real livestream sessions into structured causal supervision, yielding over 320 hours of optimization trajectories and a human-reviewed benchmark of 275 clips and 13,812 decision intervals. We train the policy with Marker-Aware Multiturn Supervised Fine-Tuning (MA-MSFT), which strengthens sparse structured decisions, followed by Streaming Multiturn GSPO (SM-GSPO), which optimizes self-generated trajectories with turn- and trajectory-level credit. On the held-out benchmark, \liveassistant reaches 71.14 state accuracy, 72.67 recipient accuracy, and 58.41 task accuracy, with consistent gains over representative streaming and general multimodal baselines. Together, the formulation, benchmark, and training framework establish livestream assistance as selective participation in a shared social stream.
Chinese Translation
直播是一种持续时间长的交互环境,其中视听内容、观众活动、主播行为和平台信号共同演化,催生出源自直播流本身的协助需求。我们提出了 \liveassistant,一个混合主动、角色条件化的协助框架,将直播交互形式化为四个相互耦合的决策:**是否行动、何时行动、面向谁行动以及传达什么内容**。在每个10秒的时间间隔内,一个自回归策略消费原生音频和视频,并结合同步的评论、礼物、观众动态和房间元数据,然后选择 \textsc{OBS}、\textsc{MEM} 或 \textsc{ANS}。\textsc{OBS} 保持沉默,\textsc{MEM} 记录一条私有的语义更新,\textsc{ANS} 则指定接收者、任务以及基于上下文的消息内容。为支持该任务,我们构建了一个轨迹引擎,将真实直播会话重构为结构化的因果监督数据,产出了超过320小时的优化轨迹,以及一个包含275个片段和13,812个决策区间的人工审核基准。我们采用标记感知多轮监督微调(Marker-Aware Multiturn Supervised Fine-Tuning, MA-MSFT)训练策略,以强化稀疏的结构化决策;随后采用流式多轮 GSPO(Streaming Multiturn GSPO, SM-GSPO),通过轮次级和轨迹级的信用分配对自生成轨迹进行优化。在留出基准上,\liveassistant 达到了71.14的状态准确率、72.67的接收者准确率和58.41的任务准确率,相较具有代表性的流式和通用多模态基线模型取得了一致的性能提升。总体而言,该任务形式化、基准和训练框架将直播协助确立为对共享社交流的选择性参与。
cs.LG / 41 / 2609.27306

Discrete Diffusion Models via Evolving Variational Autoregressive Networks

基于演化变分自回归网络的离散扩散模型
Pan, Kewen, Tang, Ying
Abstract
Conventional score-based diffusion models learn scores without representing normalized densities, whereas tractable normalized models support both sampling and direct likelihood evaluation. A recent tensor-network approach provides such a representation but is largely restricted to low-dimensional lattices. Here we introduce a discrete diffusion model that parameterizes normalized probability distributions using variational autoregressive networks. Explicit Markov jump operators govern the forward noising and reverse denoising dynamics, extending discrete diffusion models with normalized distributions to spin systems on higher-dimensional lattices. We apply this framework to the two- and three-dimensional Ising models across ordered, critical, and disordered regimes, accurately computing thermodynamic quantities including free energy, energy, and magnetization. We further integrate the framework with Monte Carlo sampling, using adaptive diffusion steps to maintain high acceptance rates even at low temperatures while enhancing sample diversity. These results establish a neural-network framework for the discrete diffusion model with normalized probability distributions.
Chinese Translation
传统的基于分数的扩散模型在学习分数时并不表示归一化密度,而可解析的归一化模型则同时支持采样和直接似然评估。近期一种张量网络方法提供了这样的表示,但主要局限于低维晶格。本文提出一种离散扩散模型,利用变分自回归网络对归一化概率分布进行参数化。显式的马尔可夫跳算子控制前向加噪和反向去噪动力学,将具有归一化分布的离散扩散模型扩展到更高维晶格上的自旋系统。我们将该框架应用于二维和三维伊辛模型的有序、临界和无序区间,精确计算了自由能、能量和磁化强度等热力学量。我们进一步将该框架与蒙特卡罗采样相结合,利用自适应扩散步数在低温下仍保持较高的接受率,同时增强样本多样性。这些结果为具有归一化概率分布的离散扩散模型建立了一个神经网络框架。
cs.LG / 42 / 2609.27355

Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction

从保留-遗忘损失景观交互的视角看量化鲁棒的机器遗忘
Wang, Jialu, Deng, Jianing, Luo, Shuqing, Li, Yuanzhe, Wang, Dongwei, Hu, Jingtong, Yang, Huanrui, Wang, Song, Chen, Tianlong
Abstract
Unlearning ensures LLM compliance by removing the influence of private or copyrighted training data. However, since LLM models typically undergo post-training compression, like quantization, in practical deployment, it has been observed that the unlearning effect can be substantially weakened, with the forgetting behavior degrading more severely than that of model utility. This paper proposes a quantization-robust unlearning framework that makes forgetting robust to quantization while maintaining overall model utility. We analyze this gap through the lens of loss landscape. Specifically, our analysis reveals a curvature-based criteria that pinpoints sensitive weights in the unlearned model that leads to both non-robust forgetting and reduced utility. We therefore propose sensitivity-guided noisy regularization, which is applied on the sensitive parameters to steer the model convergence towards a smoother minima of uniformly low forget and retain losses. Balancing unlearning and utility, we further propose forget-critical optimization, which updates only forget-critical layers, preserving most of the network to retain useful knowledge. Extensive experiments on the MUSE and TOFU benchmarks across multiple LLM unlearning algorithms show that our approach achieves substantially more quantization-resilient forgetting while maintaining utility.
Chinese Translation
机器遗忘通过移除隐私或受版权保护的训练数据的影响,确保大语言模型(LLM)的合规性。然而,由于大语言模型在实际部署中通常会经过训练后压缩(如量化),已有观察发现遗忘效果会被显著削弱,且遗忘行为的退化比模型效用的退化更为严重。本文提出一个量化鲁棒的机器遗忘框架,使遗忘对量化具有鲁棒性,同时保持模型的整体效用。我们从损失景观的视角分析这一差距。具体而言,我们的分析揭示了一种基于曲率的判据,能够精确定位未学习模型中同时导致遗忘不鲁棒和效用降低的敏感权重。据此,我们提出敏感度引导的噪声正则化方法,将其作用于敏感参数,引导模型收敛到具有均匀较低的遗忘损失和保留损失的更平滑极小值。为平衡遗忘与效用,我们进一步提出遗忘关键优化,仅更新遗忘关键层,保留网络的大部分以维持有用知识。在 MUSE 和 TOFU 基准上针对多种大语言模型遗忘算法的大量实验表明,我们的方法在保持效用的同时,实现了显著更具量化韧性的遗忘效果。
cs.LG / 43 / 2609.27362

Anomaly-Free Self-Optimization via AUC Bounds

基于AUC界的无异常自优化
Wilkinghoff, Kevin, Tan, Zheng-Hua
Abstract
Anomalies are rare, and anomalous data are often unavailable during development, making it difficult to determine which anomaly detection models and configurations will generalize to unseen anomalies. Recent approaches address this challenge by generating pseudo-anomalies and using bounds on the achievable area under the ROC curve (AUC) to select the optimal configuration from a finite set of candidates. Instead, we use the AUC bound as a differentiable, anomaly-free objective for directly optimizing continuous parameters of anomaly detection systems. We demonstrate this framework by optimizing ensemble weights and introducing a learnable score-rescaling mechanism that adapts pseudo-anomaly scores, enabling optimization beyond a predefined candidate set. Experiments across multiple datasets and embedding models show that AUC-bound optimization achieves significant performance gains over conventional model selection and prior development-set-based parameter selection. The results further show that direct optimization is less sensitive to the choice of pseudo-anomaly construction.
Chinese Translation
异常事件是罕见的,且在系统开发阶段通常无法获得异常数据,这使得难以判断哪些异常检测模型和配置能够泛化到未见过的异常。近期的方法通过生成伪异常(pseudo-anomalies),并利用ROC曲线下面积(AUC)可达上界来从有限的候选集合中选择最优配置。与之不同,我们将AUC界用作一个可微的、无需异常数据的目标函数,直接优化异常检测系统的连续参数。我们通过优化集成权重展示了该框架,并引入一种可学习的分数重缩放机制以自适应调整伪异常分数,使优化能够超越预定义的候选集合。在多个数据集和嵌入模型上的实验表明,基于AUC界的优化方法相比传统模型选择以及先前基于开发集的参数选择方法取得了显著的性能提升。结果还表明,直接优化对伪异常构建方式的选择不太敏感。
cs.LG / 44 / 2609.27385

Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools

Forecast Workflow Bench:利用预算约束的预测工具评估语言模型的决策能力
Nagashima, Shunya
Abstract
Time-series foundation models (TSFMs) provide forecasts for operational decisions, but accuracy alone does not determine their value. Evaluating agents that use these models requires measuring decision quality and forecast cost. FWBench evaluates this capability on 1,251 electricity and cycle-hire cases using fixed forecast tools and simulated capacity contracts. Agents select models, histories and horizons, then submit capacities to minimize a stated loss-cost objective. We evaluated two hosted and eight local configurations, including small language models, and tested local models with and without TSFMs. GPT-6 Astra bought inexpensive short-horizon forecasts selectively, using 2.5% of the budget, and outperformed fixed policies when the saved decisions were scored with three loss-cost weightings. FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.
Chinese Translation
时间序列基础模型(TSFM)可为运营决策提供预测,但仅有准确性并不能决定其价值。评估使用这些模型的智能体,需要同时衡量决策质量与预测成本。FWBench 在 1,251 个电力和公共自行车租赁案例上,使用固定的预测工具和模拟的容量合同来评估这一能力。智能体选择模型、历史数据和预测时域,然后提交容量以最小化给定的损失-成本目标。我们评估了两个托管配置和八个本地配置(包括小型语言模型),并对本地模型在有 TSFM 和无 TSFM 两种情况下进行了测试。GPT-6 Astra 有选择性地购买廉价短时域预测,仅使用了 2.5% 的预算,并且在以三种损失-成本权重对所保存的决策进行评分时优于固定策略。FWBench 为评估语言模型如何在成本约束下选择和使用时间序列预测进行决策提供了可复现的评估手段。
cs.LG / 45 / 2609.27409

Active Learning for Biodiversity Monitoring: From Label Efficiency to Reliable Ecological Inference

面向生物多样性监测的主动学习:从标签效率到可靠的生态推断
McEwen, Ben, Zhang, Shiqi, Stowell, Dan
Abstract
Limited expert annotation capacity is a pervasive constraint in biodiversity monitoring. Passive acoustic recorders and camera traps generate data faster than experts can analyse them. Machine learning (ML) models can process these data at scale, but their reliability depends on the quality, quantity, and coverage of labelled samples, so expert time remains a constraint. Active learning (AL) eases this bottleneck by selecting, under a fixed annotation budget, the samples expected to improve a model most, and published evidence shows it can reduce the labels needed to reach a target performance. Monitoring programmes, however, face a broader question: how should a limited expert budget be divided so that model training, validation, and the ecological estimates built on model outputs all remain reliable? Because AL selects samples non-randomly, its labels are unsuitable for validation, calibration, or threshold selection, a tension rarely acknowledged. We synthesise AL research across acoustic and image modalities and identify gaps and opportunities. Most studies evaluate query strategies on pre-labelled benchmarks with simulated annotators; deployments in real monitoring workflows are rare and concentrate on birds and cetaceans. Bats, insects, amphibians, and fish are underrepresented, and multimodal applications remain largely unexplored. Evaluation centres on headline reductions in annotation effort, often without random-sampling baselines, per-class results, or calibration analysis, and rarely accounts for the labels required for validation. We provide a tutorial treatment of the AL loop that makes these budget decisions explicit, and a roadmap towards AL methods that support label-efficient training, validation, and trustworthy downstream ecological inference.
Chinese Translation
专家标注能力有限是生物多样性监测中普遍存在的制约因素。被动声学记录仪和相机陷阱产生数据的速度快于专家的分析速度。机器学习(ML)模型能够大规模处理这些数据,但其可靠性取决于标注样本的质量、数量和覆盖范围,因此专家时间仍然是一种约束。主动学习(AL)通过在固定标注预算下选择预期最能提升模型性能的样本来缓解这一瓶颈,且已发表的证据表明它可以减少达到目标性能所需的标签数量。然而,监测项目面临一个更广泛的问题:应如何分配有限的专家预算,才能使模型训练、验证以及基于模型输出构建的生态估计都保持可靠?由于主动学习以非随机方式选择样本,其标签不适用于验证、校准或阈值选择,这一矛盾鲜有人提及。我们综述了声学和图像模态下的主动学习研究,并识别其中的空白与机遇。大多数研究在预先标注的基准数据集上、借助模拟标注者评估查询策略;在真实监测工作流程中的部署十分少见,且集中于鸟类和鲸类动物。蝙蝠、昆虫、两栖动物和鱼类的相关研究不足,多模态应用也基本未被探索。评估主要聚焦于标注工作量的总体削减,往往缺乏随机采样基线、分类别结果或校准分析,且很少计入验证所需的标签。我们对主动学习循环进行了教程式阐述,使这些预算决策显性化,并提出了通往支持标签高效训练、验证以及可信下游生态推断的主动学习方法路线图。
cs.LG / 46 / 2609.27411

When Labels Are Scarce: An Oscillatory State Space Model for Vibration Diagnosis

当标签稀缺时:一种用于振动诊断的振荡状态空间模型
Mallick, Mainak, Choi, Seung-Kyum
Abstract
Machine fault diagnosis from vibration requires learning from scarce labelled fault recordings while meeting the computational constraints of edge devices for local inference. We introduce DualRes, a compact oscillatory state-space model that combines two complementary spectral views of vibration, capturing rapid changes and fine frequency structure. Time-aligned views are processed by selective oscillatory memory, which learns how long to retain temporal patterns. The encoder contains 39,528 parameters. We evaluate supervised learning across six bearing datasets and a gearbox benchmark, with an additional gearbox pilot. Recording-level splits and explicit accounting of labelled duration distinguish data efficiency from repeated exposure to correlated samples. On the main gearbox benchmark, DualRes achieves state-of-the-art performance among the nine evaluated methods at six of seven label budgets. With about six labelled seconds per class, it improves macro-F1 by 16.1 percentage points over the next strongest comparator. On the same benchmark, DualRes achieves a 1.44-fold recording-level speedup and a 24.8-fold reduction in checkpoint storage relative to a selective state-space baseline under matched hardware and runtime conditions. Bearing results reveal task-dependent trade-offs. These findings support oscillatory memory as a compact approach to vibration diagnosis under limited labelled exposure.
Chinese Translation
基于振动的机器故障诊断需要在标签稀缺的故障记录上进行学习,同时还要满足边缘设备本地推理的计算约束。我们提出了DualRes,这是一种紧凑的振荡状态空间模型,它结合了对振动的两种互补谱视角,能够捕捉快速变化和精细的频率结构。时间对齐的视角由选择性振荡记忆(selective oscillatory memory)处理,该记忆可以学习保留时间模式的时长。该编码器仅包含39,528个参数。我们在六个轴承数据集和一个齿轮箱基准以及一个额外的齿轮箱试点数据集上评估了监督学习性能。通过记录级别的数据划分以及对有标签时长的显式核算,我们区分了数据效率与对相关样本的重复暴露。在主要齿轮箱基准上,DualRes在七个标签预算中的六个上取得了九种评估方法中的最先进性能。在每类约有六秒有标签数据的情况下,其宏F1分数相比次优对比方法提升了16.1个百分点。在相同基准上,在匹配的硬件和运行条件下,DualRes相对于选择性状态空间基线实现了1.44倍的记录级加速和24.8倍的检查点存储缩减。轴承数据的结果揭示了任务相关的权衡。这些发现支持将振荡记忆作为一种在有限有标签数据下进行振动诊断的紧凑方法。
cs.LG / 47 / 2609.27421

Counterfactual Constraint-Conditioned On-Policy Distillation for Multi-Constraint Instruction Following

面向多约束指令跟随的反事实约束条件化在策略蒸馏方法
Zheng, Yanzhao, Yu, Yuanqiang, Xu, Tianze, Ma, Chao, Zhang, Zhentao, Zhu, Jihuai, Dong, Baohua, Zhu, Hangcheng, Huang, Ruohui
Abstract
Multi-constraint instruction following requires a model to respond to a query under many simultaneously active constraints. Even strong instruction-tuned models still routinely violate some of them. Existing approaches either augment supervision with sequence- or token-level RL rewards from external verifiers or learned graders, or use on-policy distillation (OPD) against a single full-context teacher whose probability mass becomes diluted as more constraints become simultaneously active. We propose CC-OPD (Counterfactual Constraint-Conditioned On-Policy Distillation), which inverts the standard supervision-generation direction in distillation. Rather than enriching the teacher with information beyond what the student sees, CC-OPD ablates each constraint from the teacher's conditioning in turn, and constructs the per-constraint signal from the resulting per-token probability differentials. The resulting per-token leave-one-out log-likelihood shifts are summed, clipped, and added to the vanilla OPD reward as a token-level shaping term. All shaping terms are obtained from the frozen teacher, without an external verifier during distillation, and the reward equals vanilla OPD wherever the aggregate shift is zero. Across two Qwen model pairs and seven benchmarks, CC-OPD achieves the highest average among all evaluated student-training methods. A 1.5B student trained with CC-OPD surpasses its own 7B RL-trained teacher on the MulDimIF benchmark.
Chinese Translation
多约束指令跟随要求模型在多个同时生效的约束条件下对查询作出响应。即使是强大的指令微调模型,也仍然经常违反其中某些约束。现有方法要么利用来自外部验证器或学习型评分器的序列级或词元级强化学习奖励来增强监督信号,要么采用在策略蒸馏(On-Policy Distillation, OPD),以单个完整上下文教师模型为目标——但随着同时生效的约束数量增加,该教师模型的概率质量会被稀释。我们提出 CC-OPD(反事实约束条件化在策略蒸馏,Counterfactual Constraint-Conditioned On-Policy Distillation),该方法逆转了蒸馏中标准的监督-生成方向。CC-OPD 并不用学生模型看不到的信息来丰富教师模型,而是依次从教师模型的条件中消融每一个约束,并根据由此产生的逐词元概率差分构建每个约束的信号。所得的逐词元留一法对数似然偏移经过求和、裁剪后,作为词元级塑形项加到原始 OPD 奖励上。所有塑形项均来自冻结的教师模型,蒸馏过程中无需外部验证器,且当聚合偏移为零时,奖励与原始 OPD 相同。在两对 Qwen 模型以及七个基准测试上,CC-OPD 在所有被评估的学生训练方法中取得了最高的平均性能。使用 CC-OPD 训练的 1.5B 学生模型在 MulDimIF 基准上超越了其自身的 7B 强化学习训练教师模型。
cs.LG / 48 / 2609.27441

Stable Neural Decoding Across Sessions via Task-Conditioned Latent Alignment for Brain-Machine Interfaces

基于任务条件化潜在对齐的跨会话稳定神经解码方法及其在脑机接口中的应用
Zhao, Canyang, Peng, Bolin, Mayo, J. Patrick, Ju, Ce, Liu, Bing
Abstract
Achieving stable long-term neural decoding in invasive brain-machine interfaces (BMIs) remains challenging due to variations in recorded neural populations across sessions. Current latent alignment approaches may overlook task-dependent structure during cross-session adaptation. We propose Task-Conditioned Latent Alignment (TCLA), a framework that stabilizes neural decoding by learning a shared latent space. TCLA learns a low-dimensional source representation using neural reconstruction and continuous behavioral supervision. During target-session adaptation, the shared representation is fixed, while target neural activity is mapped into the source latent space by aligning source and target distributions separately for each task condition. We evaluated TCLA on seven nonhuman primate datasets spanning multiple tasks. In long-term cross-session evaluation, TCLA achieved a mean $R^2$ of $0.476\pm0.014$ with a negative $R^2$ failure rate of only 6.8\%. Across 1,356 within-subject session pairs, TCLA achieved a mean $R^2$ of $0.371\pm0.009$ with a failure rate of 6.8\%. Across 2,134 cross-subject session pairs, TCLA achieved a mean $R^2$ of $0.218\pm0.004$ with a failure rate of 12.9\%, substantially better than those of the comparison methods. These results demonstrate that by preserving behaviorally relevant and task-dependent latent structure, TCLA improves the robustness of neural decoding across recording sessions and subjects. The source code is publicly available at \href{https://github.com/FAMD-CASIA/TCLA}{https://github.com/FAMD-CASIA/TCLA}.
Chinese Translation
由于不同会话间所记录神经元群的变化,在侵入式脑机接口(BMI)中实现稳定的长期神经解码仍然具有挑战性。现有的潜在对齐方法在跨会话适应过程中可能忽略了任务相关的结构。我们提出了一种任务条件化潜在对齐(Task-Conditioned Latent Alignment, TCLA)框架,该框架通过学习一个共享潜在空间来稳定神经解码。TCLA 利用神经重构和连续行为监督学习一个低维源表示。在目标会话适应过程中,共享表示保持固定,同时通过在各个任务条件下分别对齐源分布和目标分布,将目标神经活动映射到源潜在空间中。我们在涵盖多个任务的七个非人灵长类数据集上对 TCLA 进行了评估。在长期跨会话评估中,TCLA 实现了 0.476±0.014 的平均 R²,负 R² 失败率仅为 6.8%。在 1,356 个被试内会话对中,TCLA 实现了 0.371±0.009 的平均 R²,失败率为 6.8%。在 2,134 个跨被试会话对中,TCLA 实现了 0.218±0.004 的平均 R²,失败率为 12.9%,均显著优于对比方法。这些结果表明,通过保留与行为相关且依赖于任务的潜在结构,TCLA 提高了神经解码在不同记录会话和被试之间的鲁棒性。源代码已在 https://github.com/FAMD-CASIA/TCLA 公开。
cs.LG / 49 / 2609.27446

Quantum Reinforcement Learning for Cost and Delay Tradeoffs in Quantum Cloud Orchestration

面向量子云编排中成本与延迟权衡的量子强化学习
Phan, An N. H., Van Huynh, Dang, Usman, Muhammad, Nguyen, Hoa T.
Abstract
Quantum cloud computing, delivered through the quantum-as-a-service (QaaS) model, provides access to quantum computing resources. However, applying uniform time-based pricing across fundamentally heterogeneous quantum resources significantly complicates task orchestration, particularly when addressing the tradeoff between execution costs and system performance. While heuristic methods rely on predefined scheduling rules, classical deep reinforcement learning (DRL) models may require more trainable parameters in this setting. Motivated by the potential of parameterised quantum circuits (PQCs) as compact function approximators, we propose QRLQ, a cost-delay-aware quantum cloud scheduling framework integrating PQCs with a dueling double deep Q-network (D3QN) to dynamically account for both cost and delay. Our simulation results show that QRLQ achieves lower mean cost and delay than the heuristic baselines, achieving a 5-11% lower mean cost relative to availability-based and rotation-based heuristics and reducing mean delay by 17% and 82% relative to the strongest and weakest heuristic baselines, respectively, while retaining execution fidelity within 2% of a fidelity-greedy policy. Compared with the classical DRL baseline, QRLQ achieves comparable scheduling performance while using 72% fewer trainable parameters. This work explores the feasibility of using QRL for task orchestration in quantum cloud environments and demonstrates its potential for cost-delay-aware quantum resource management.
Chinese Translation
量子云计算通过量子即服务(QaaS)模式提供对量子计算资源的访问。然而,在本质上异构的量子资源上采用统一的时间计价方式,极大地增加了任务编排的复杂性,尤其是在处理执行成本与系统性能之间的权衡时。尽管启发式方法依赖预定义的调度规则,经典深度强化学习(DRL)模型在此类场景下可能需要更多的可训练参数。受参数化量子电路(PQC)作为紧凑函数逼近器的潜力启发,我们提出了QRLQ,一个成本-延迟感知的量子云调度框架,它将PQC与决斗双深度Q网络(D3QN)相结合,以动态兼顾成本与延迟。仿真结果表明,QRLQ相比启发式基线获得了更低的平均成本和延迟:相对于基于可用性和基于旋转的启发式方法,平均成本降低5-11%;相对于最强和最弱的启发式基线,平均延迟分别降低17%和82%;同时执行保真度与保真度贪心策略的差距保持在2%以内。与经典DRL基线相比,QRLQ在调度性能相当的情况下,可训练参数减少了72%。本工作探索了在量子云环境中利用量子强化学习(QRL)进行任务编排的可行性,并展示了其在成本-延迟感知量子资源管理方面的潜力。
cs.LG / 50 / 2609.27473

Learning Where to Look: A Shared Relative-Alignment Module for Time-Series Forecasting and PPG-to-Vital-Sign Reconstruction

学习看哪里:一个用于时间序列预测与PPG到生命体征重建的共享相对对齐模块
Puli, Ragamayi, Nagashima, Shunya
Abstract
PPG-to-vital-sign reconstruction turns a wrist-worn photoplethysmogram into clinical waveforms such as the ECG. Long-horizon multivariate time-series forecasting underpins planning in energy, weather, and traffic. Both generate a target sequence from a condition sequence, and current models hard-code where each target position reads it, as a same-position copy or seasonal recurrence, so neither transfers between tasks. We propose ROOSTER, one conditioning module that handles vital-sign reconstruction and time-series forecasting alike by learning this correspondence. Its core is a periodic-comb bias over the target-condition offset whose center, period, and sharpness are learned per head, so one module settles on the identity alignment or a seasonal lag and reports which it found. On vital-sign reconstruction from PPG, ROOSTER outperformed the published baselines on four heart-rate and respiratory-rate benchmarks. On multivariate time-series forecasting, it achieved the best horizon-averaged MSE on four benchmarks and outperformed the forecasting model it extends on 20 of 24 dataset-horizon settings under matched three-seed training. An ablation study indicated that the relative bias, not content matching, carried the alignment.
Chinese Translation
PPG到生命体征重建是将腕部佩戴式光电容积脉搏波(PPG)转换为心电图(ECG)等临床波形。长时程多变量时间序列预测是能源、气象和交通领域规划的基础。这两类任务都是从条件序列生成目标序列,而当前模型将目标各位置读取条件序列的方式硬编码为同位置复制或周期性季节重复,因此均无法在任务间迁移。我们提出了ROOSTER,一个通过学习这种对应关系、可同时处理生命体征重建和时间序列预测的统一条件模块。其核心是在目标-条件偏移上的周期梳状偏置,其中心、周期和锐度在每个注意力头中独立学习,因此单一模块可以自动收敛到恒等对齐或季节性滞后对齐,并能报告其发现的模式。在从PPG重建生命体征的任务中,ROOSTER在四个心率和呼吸率基准上优于已发表的基线方法。在多变量时间序列预测任务中,它在四个基准上取得了最佳的平均预测时程MSE,并在相同三随机种子训练条件下,在24个数据集-时程设置中的20个上优于其所扩展的预测模型。消融实验表明,起对齐作用的是相对偏置而非内容匹配。
cs.LG / 51 / 2609.27532

ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning

ProCredit:从结果奖励到进展信用分配的智能体强化学习
Ma, Ming, Zhu, Yi, Zhong, Yiran, Zhu, Feida, Liu, Chonghan, Jiao, Pengkun, Wang, Qichao, Jia, Yanhao, Yang, Tianming, Hoi, Steven
Abstract
Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The standard recipe assigns a single outcome reward at the end and compares trajectories sampled for the same task. As a result, a group with no successful trajectory yields no training signal, failed attempts cannot be told apart by how close they came to completion, and turns that advance the task receive the same credit as turns that only query the environment. Prior work refines the unit of comparison from the trajectory to the step, or trains a reward model to supply intermediate signal: the former still derives its signal from final success alone, and the latter estimates it with a model. We observe that the acceptance checks that decide success can also be run on intermediate states, so progress is as verifiable as the outcome. We propose ProCredit, which turns this verified progress into credit: it reruns the acceptance checks after each turn, rewards the turn by its change in progress, and uses these rewards to assign credit both across attempts at the same task and across the turns within a trajectory. Starting from Qwen3.5 base models at three scales on AppWorld, ProCredit outperforms outcome-reward baselines and progress-based baselines in task completion rate at every scale on both test sets, exceeding the strongest outcome-reward baseline by 4.1 percentage points at 4B, and results in a second environment show the same direction of improvement. Ablations show that adding the final progress to the trajectory score alone does not improve performance: the gain comes from crediting progress to the turn where it occurs.
Chinese Translation
长时程智能体任务要求智能体通过一系列工具调用来修改环境,其成功与否由最终状态决定。标准方法在任务结束时赋予单一的结果奖励,并比较针对同一任务采样的轨迹。由此导致:若一组采样中没有任何成功轨迹,则无法产生训练信号;失败的尝试无法按其接近完成的程度加以区分;推进任务的轮次与仅仅查询环境的轮次获得相同的信用。已有工作要么将比较单位从轨迹细化为步骤,要么训练奖励模型以提供中间信号:前者仍然仅从最终成功中获取信号,后者则依赖模型进行估计。我们观察到,判定成功的验收检查(acceptance checks)同样可以在中间状态上运行,因此进展与结果一样是可验证的。我们提出 ProCredit,将这种经过验证的进展转化为信用分配:它在每一轮结束后重新运行验收检查,根据进展的变化量奖励该轮,并利用这些奖励在相同任务的多次尝试之间以及单条轨迹的各轮次之间进行信用分配。在 AppWorld 上,基于三个规模的 Qwen3.5 基座模型,ProCredit 在两个测试集的所有规模上,任务完成率均优于结果奖励基线和基于进展的基线,在 4B 规模上超出最强的结果奖励基线 4.1 个百分点;在第二个环境上的实验结果也呈现出相同的改进趋势。消融实验表明,仅将最终进展加入轨迹评分并不能提升性能:收益来自于将进展信用分配到其发生的具体轮次。
cs.LG / 52 / 2609.27547

EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management

EBRL:基于多粒度资源管理的异步具身强化学习
Mi, Liang, Wang, Weijun, Gao, Bowen, Yu, Tianze, Hao, Zixu, Xiao, Han, Ding, Xin, Huang, Mingzhe, He, Xin, Shi, Lu, Wu, Hao, Dai, Haipeng, Chen, Guihai, Liu, Yunxin, Cao, Ting
Abstract
Embodied reinforcement learning (RL) improves model capabilities with a pipeline of environment simulation, action generation, and model updates. These stages show heterogeneous CPU and GPU demands, making efficient resource utilization difficult. Recent systems overlap rollout (simulation and generation) with training for efficiency, but exclusive GPU allocation and synchronized barrier in rollout still leave substantial hardware resource waste. In this paper, we present EBRL, an asynchronous embodied RL training system with two core techniques. The asynchronous pipelined scheduler overlaps rollout and training, pipelines simulation and generation across environment groups, and carries out each environment independently, eliminating synchronization stalls. The fine-grained resource manager pools CPU cores and GPU streaming multiprocessors, and uses stage profiles and runtime feedback to adjust resource quotas and batch sizes to meet the shifting demands among stages. We implement EBRL on RLinf and evaluate it with four embodied policies and four simulation benchmarks across heterogeneous GPU testbeds. Experiments show that EBRL achieves 1.30-3.47 times the end-to-end rollout throughput and 2.5 times of training convergency compared to the SOTA embodied RL systems.
Chinese Translation
具身强化学习(RL)通过环境仿真、动作生成和模型更新等流水线阶段来提升模型能力。这些阶段对CPU和GPU的需求呈现异构性,使得高效的资源利用变得困难。近期的系统为提高效率,将推演(rollout,即仿真与生成)与训练重叠执行,但GPU的独占式分配和推演中的同步屏障仍导致大量硬件资源浪费。本文提出EBRL,一个具有两项核心技术的异步具身RL训练系统。异步流水线调度器将推演与训练重叠执行,在环境组之间对仿真和生成进行流水线化处理,并独立执行每个环境,从而消除同步停顿。细粒度资源管理器将CPU核心和GPU流式多处理器汇入资源池,利用阶段画像和运行时反馈来调整资源配额与批大小,以满足各阶段间不断变化的需求。我们在RLinf上实现了EBRL,并在异构GPU测试平台上使用四种具身策略和四种仿真基准对其进行评估。实验表明,与最先进的(SOTA)具身RL系统相比,EBRL实现了1.30至3.47倍的端到端推演吞吐量和2.5倍的训练收敛速度。
cs.LG / 53 / 2609.27554

PhyMo: A Physical-Field Modality for Multimodal AI4Physics

PhyMo:面向多模态AI4Physics的物理场模态
Sun, Henan, Hu, Haitao, Liu, Jin, Zhang, Jianfeng, Pan, Lujia, Chen, Nuo, Li, Jia
Abstract
Multimodal learning is emerging as a powerful paradigm for AI for Physics (AI4Physics), where predicting physical systems requires the joint interpretation of heterogeneous observations, measurements, and domain knowledge. However, existing approaches typically represent physical quantities and governing equations as generic numerical or textual tokens, overlooking the physical constraints that determine their spatiotemporal interactions. To address this limitation, we introduce the \textbf{physical-field modality} and propose \textbf{PhyMo}, a physics-grounded multimodal framework that organizes heterogeneous measurements through PDE-associated operators. PhyMo follows a three-stage learning procedure: the physical-field encoder is first pretrained through field reconstruction under PDE residual supervision, its representations are subsequently aligned with visual embeddings in a shared latent space, and the fused multimodal representations are finally processed by corresponding downstream prediction heads. Experiments on five datasets spanning diverse physical environments show that PhyMo achieves state-of-the-art performance, compared to the strongest baseline on each dataset, demonstrating the superiority of PhyMo on multimodal representation learning in AI4Physics.
Chinese Translation
多模态学习正在成为物理人工智能(AI4Physics)的一种强大范式,其中对物理系统的预测需要对异构观测、测量和领域知识的联合解释。然而,现有方法通常将物理量和控制方程表示为通用的数值或文本标记(token),忽略了决定其时空相互作用的物理约束。为解决这一局限性,我们引入了物理场模态(physical-field modality),并提出PhyMo——一个以物理为基础的多模态框架,通过偏微分方程(PDE)相关算子来组织异构测量数据。PhyMo遵循三阶段学习流程:首先在PDE残差监督下通过场重构对物理场编码器进行预训练;随后将其表示与视觉嵌入在一个共享潜在空间中对齐;最后由相应的下游预测头处理融合后的多模态表示。在涵盖多种物理环境的五个数据集上的实验表明,PhyMo相比每个数据集上最强的基线方法均取得了最先进的性能,证明了PhyMo在AI4Physics多模态表示学习中的优越性。
cs.LG / 54 / 2609.27564

TNLearn: An Open Source Python Package for Task-based Neurons

TNLearn:一个用于任务型神经元的开源Python软件包
Wang, Meng, Li, Tieyun, Fan, Juntong, Pei, Hanyu, Liao, Jing-Xiao, Yang, Yaodong, Ma, Jianwei, Fan, Fenglei
Abstract
The brain does not rely on a single type of neuron to perform all kinds of tasks; instead, it designs different neurons for different tasks. The concept of task-based neurons represents a paradigm shift compared to task-based architectures. It argues that solving a specific problem requires customized neurons, as task-based neurons capture useful prior knowledge from task-related data. To facilitate the use of task-based neurons in scientific research and industrial applications, we introduce TNLearn, an open-source Python package that provides automated construction of task-based neurons and networks, enabling smooth training of task-based networks. Comprehensive documentation, including technical exposition, API reference, and representative examples, is available online. TNLearn is open-sourced at https://github.com/NewT123-WM/tnlearn and has become a PyTorch ecosystem project.
Chinese Translation
大脑并不依赖单一类型的神经元来执行各类任务,而是为不同任务设计不同的神经元。与基于任务的架构相比,任务型神经元的概念代表了一种范式转变。该观点认为,解决特定问题需要定制化的神经元,因为任务型神经元能够从与任务相关的数据中捕捉有用的先验知识。为了促进任务型神经元在科学研究和工业应用中的使用,我们推出了TNLearn,这是一个开源Python软件包,提供任务型神经元和网络的自动化构建,使任务型网络能够顺畅地进行训练。相关网站提供了全面的文档,包括技术阐述、API参考和代表性示例。TNLearn已在https://github.com/NewT123-WM/tnlearn开源,并已成为PyTorch生态系统的项目。
cs.LG / 55 / 2609.27572

DCRL: Decoupling and Coupling Reinforcement Learning via Policy-Reward Manifold Alignment

DCRL:基于策略-奖励流形对齐的解耦与耦合强化学习
Sun, Henan, Li, Zehua, Hu, Haitao, Zhang, Qifan, Zhang, Jianfeng, Chen, Nuo, Li, Jia
Abstract
Reinforcement learning (RL) has emerged as a key paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing reward systems, such as rule-based and reward-model-based, often exhibit issues such as unstable optimization and reward hacking. In this work, we revisit the general reasoning of LLMs from a geometric perspective, conceptualizing it as a coupled manifold composed of three interdependent sub-manifolds: logical deduction, evaluation, and representation. Based on this perspective, response generation in RL can be interpreted as a decoupling process from the evaluation manifold, while reward estimation corresponds to a decoupling process from the logical deduction manifold. The limitations of rule-based and reward-model RL systems can be geometrically interpreted as the mismatch of policy-reward manifolds during RL process. To address the aforementioned misalignment, we propose Decoupling and Coupling Reinforcement Learning (DCRL) framework, which incorporates two key components: (1) a syllogistic logic-based prompt evolution mechanism that dynamically refines reward rubrics to enhance the expressiveness of the reward manifold; and (2) a policy-reward re-coupling mechanism that jointly updates the reward and policy models, ensuring consistent evaluation and mitigating manifold mismatch during training. Theoretical analysis and extensive experiments across multiple reasoning domains demonstrate that DCRL consistently outperforms both rule-based and reward-model baselines. Notably, a Qwen3-4B model trained under DCRL surpasses a Qwen3-32B baseline and approaches the performance of a Qwen3-235B model, highlighting superior effectiveness and generalization in RL.
Chinese Translation
强化学习(RL)已成为提升大语言模型(LLM)推理能力的关键范式。然而,现有的奖励系统(如基于规则的和基于奖励模型的系统)常常存在优化不稳定和奖励作弊(reward hacking)等问题。在本工作中,我们从几何视角重新审视LLM的通用推理,将其概念化为由三个相互依赖的子流形构成的耦合流形:逻辑推演、评估和表示。基于这一视角,强化学习中的响应生成可以被解释为从评估流形的解耦过程,而奖励估计则对应于从逻辑推演流形的解耦过程。基于规则和基于奖励模型的强化学习系统的局限性可以在几何上解释为强化学习过程中策略-奖励流形的不匹配。为解决上述不对齐问题,我们提出了解耦与耦合强化学习(Decoupling and Coupling Reinforcement Learning, DCRL)框架,其包含两个关键组件:(1)一种基于三段论逻辑的提示演化机制,可动态优化奖励规则(reward rubrics),以增强奖励流形的表达能力;(2)一种策略-奖励再耦合机制,联合更新奖励模型和策略模型,确保评估的一致性并缓解训练过程中的流形不匹配。理论分析和在多个推理领域的广泛实验表明,DCRL始终优于基于规则和基于奖励模型的基线方法。值得注意的是,在DCRL下训练的Qwen3-4B模型超越了Qwen3-32B基线,并接近Qwen3-235B模型的性能,凸显了其在强化学习中的卓越有效性和泛化能力。
cs.LG / 56 / 2609.27577

VCMM: Variance-Calibrated Momentum for Multimodal Learning

VCMM:用于多模态学习的方差校准动量方法
Gu, Zhongjing, Huang, Chenyang, Feng, Yufa, He, Chong, Ding, Qinxu, Cui, Yiming
Abstract
Multimodal joint training often suffers from modality imbalance, where a dominant modality suppresses the optimization of others. Existing methods mainly balance modality learning by modulating gradient magnitudes or directions, modifying optimization objectives, or adjusting training strategies, with most interventions focusing on the current update. However, when combined with widely used momentum-based optimizers, the update also incorporates accumulated information from previous gradients, which is not explicitly addressed by current-step modulation alone. To address this issue, we propose Variance-Calibrated MomentuM (VCMM), which adapts gradient memory to modality-specific gradient dynamics. Specifically, VCMM estimates minibatch noise and temporal drift online and uses their relative strength to determine modality-specific momentum through a Kalman-inspired controller. We further center the control signal across modalities and apply exact bias correction for the time-varying first moment, enabling adaptive gradient memory without extra network passes or explicit learning-rate scaling. Experiments on four multimodal benchmarks demonstrate consistent improvements with modest training overhead.
Chinese Translation
多模态联合训练常常受到模态不平衡问题的困扰,即某一主导模态会抑制其他模态的优化。现有方法主要通过调节梯度的大小或方向、修改优化目标或调整训练策略来平衡模态学习,且大多数干预仅聚焦于当前的更新步骤。然而,当与广泛使用的基于动量的优化器结合时,参数更新还包含了先前梯度累积的信息,而仅依靠当前步骤的调节无法显式地解决这一问题。为了解决这一问题,我们提出方差校准动量(Variance-Calibrated MomentuM,VCMM),该方法根据模态特定的梯度动态自适应地调整梯度记忆。具体而言,VCMM 在线估计小批量噪声和时间漂移,并通过一个受卡尔曼滤波启发的控制器,利用二者的相对强度来确定模态特定的动量。我们进一步对各模态的控制信号进行中心化处理,并对随时间变化的一阶矩施加精确的偏差修正,从而在不增加额外网络前向传播或显式学习率缩放的情况下实现自适应的梯度记忆。在四个多模态基准数据集上的实验表明,该方法在仅增加少量训练开销的情况下取得了一致的性能提升。
cs.LG / 57 / 2609.27581

Does Step Law Transfer to Small-Scale Language Models? An Empirical Recalibration Below 59M Parameters

步长定律能否迁移到小规模语言模型?一项针对59M参数以下模型的经验性重新校准
Romanyukov, Egor, Novikov, Timofey, Shokarov, Timur, Zorkina, Elizaveta, Palienko, Anastasia, Dergachev, Stepan
Abstract
Step Law gives power-law formulas for the optimal peak learning rate eta* and batch size B* when pre-training language models. It was calibrated on models between 59M and 1B parameters; the small-model regime N < 59M was never tested empirically by its authors. This regime matters for single-GPU training, interpretability research, educational experiments, and settings where larger models are infeasible on memory or cost grounds. We test whether Step Law transfers to small language models. We consider three outcomes: H1, the original coefficients work directly; H2, the power-law form holds but with different coefficients; and H3, a power law does not describe the optima in this regime. All experiments use a single nanoGPT/TinyStories pipeline with a 2048-token BPE vocabulary, AdamW, and a warmup-cosine schedule. The optimum for each (N, D) cell is extracted from the loss surface L(eta, B) via a local quadratic approximation in log-log coordinates over the smoothed training loss. The final dataset contains 29 unique (N, D) cells and 935 analysis-ready runs. The main refit uses 25 cells (815 runs) in the working range 4 <= D/N <= 600. On the pooled data we accept H2: the functional form is preserved, but the coefficients differ from the original. We obtain eta*(N, D) = 0.0985 N^(-0.508) D^(0.238) (R^2 = 0.834) and B*(D) = 3.6 x 10^(-4) D^(0.931) (R^2 = 0.950). Step Law's structural claim that B* is independent of N is reproduced (p = 0.87), but the growth of B* with D is nearly twice as steep as in the original work. Direct transfer of Step Law systematically overestimates the optimal learning rate: the median ratio eta_SL / eta* is approximately 4.0x, with a range of 2.4x to 6.6x.
Chinese Translation
步长定律为预训练语言模型时的最优峰值学习率 eta* 和批大小 B* 给出了幂律公式。该定律是在参数量介于59M至1B之间的模型上校准的;其作者从未对 N < 59M 的小模型区间进行过实证检验。这一区间对单GPU训练、可解释性研究、教学实验以及因内存或成本限制而无法使用更大模型的场景具有重要意义。我们检验了步长定律能否迁移到小规模语言模型。我们考虑三种结果:H1,原始系数可直接适用;H2,幂律形式成立但系数不同;H3,幂律无法描述该区间内的最优值。所有实验均采用统一的 nanoGPT/TinyStories 流水线,使用2048词元的BPE词表、AdamW优化器以及预热-余弦调度策略。每个 (N, D) 单元格的最优值通过在平滑后的训练损失上以对数-对数坐标下的局部二次近似从损失面 L(eta, B) 中提取。最终数据集包含29个不同的 (N, D) 单元格和935个可用于分析的运行。主要重新拟合使用工作区间 4 <= D/N <= 600 内的25个单元格(815个运行)。在汇总数据上我们接受H2:函数形式得以保留,但系数与原始结果不同。我们得到 eta*(N, D) = 0.0985 N^(-0.508) D^(0.238)(R^2 = 0.834)以及 B*(D) = 3.6 x 10^(-4) D^(0.931)(R^2 = 0.950)。步长定律关于 B* 与 N 无关的结构性结论得到了复现(p = 0.87),但 B* 随 D 增长的斜率几乎是原始工作的两倍。直接套用步长定律会系统性高估最优学习率:eta_SL / eta* 的中位数约为4.0倍,范围为2.4倍至6.6倍。
cs.LG / 58 / 2609.27588

The Capability Manifold and ML Scaling Laws

能力流形与机器学习缩放定律
Zaidi, Syed Ali Raza, Hafeez, Maryam
Abstract
Existing machine learning (ML) scaling laws relate predictive loss to compute, model parameters, and data. However, as models are increasingly deployed through agentic harnesses, loss alone is insufficient to characterize downstream performance: models with similar loss can exhibit different capabilities in reasoning, retrieval, planning, and adaptation. Yet, no unified framework connects such capabilities to the coupled resources available across the ML lifecycle. We bridge this gap by introducing a capability manifold, a multidimensional framework mapping downstream capabilities to pre-training, post-training, and test-time resources through bounded scaling functions. Analytical Jacobians quantify capability sensitivity to resource changes and interactions. As an initial application, we embed Kaplan- and Chinchilla-type scaling laws and test-time compute within the framework, demonstrating how existing scaling relationships can be unified as trajectories on a common capability manifold.
Chinese Translation
现有的机器学习(ML)缩放定律将预测损失与计算量、模型参数和数据相关联。然而,随着模型越来越多地通过智能体框架部署,仅凭损失已不足以刻画下游性能:损失相近的模型在推理、检索、规划与适应等方面可能表现出不同的能力。然而,目前尚无统一框架能将这类能力与机器学习全生命周期中可用的耦合资源联系起来。为弥补这一空白,我们提出了能力流形,这是一个多维框架,通过有界缩放函数将下游能力映射到预训练、后训练和测试时资源。解析雅可比矩阵量化了能力对资源变化及资源间交互的敏感性。作为初步应用,我们将 Kaplan 型和 Chinchilla 型缩放定律以及测试时计算嵌入该框架,展示了如何将现有的缩放关系统一为同一能力流形上的轨迹。
cs.LG / 59 / 2609.27593

Hidden not Deleted: How Networks Suppress Entangled Features

隐藏而非删除:神经网络如何抑制纠缠特征
Samanta, Akash, Singh, Manish Pratap, Chaudhuri, Debasis
Abstract
Concept erasure methods that operate via linear projection assume that features occupy separable subspaces. We show this assumption fails under dense superposition: when two features are forced into an antipodal pair sharing a single subspace, state-of-the-art linear erasure destroys both, not just the target. Networks trained with gradient descent instead solve this problem non-linearly, but not uniformly: they converge to one of two distinct circuit-level solutions depending on initialization, which we call mirror and shadow solutions. We map this bifurcation as a function of feature entanglement, show it reflects a stable attractor structure rather than an artifact of our setup, and use targeted causal interventions to demonstrate that both solutions leave a substantial, measurable trace of the erased feature's representation intact, recoverable through a single scalar patch rather than requiring any further training. This mirrors a failure mode recently observed empirically in LLM unlearning, where suppression rather than deletion allows forgotten knowledge to resurface; our results offer a mechanistic, causally-validated account of why that failure mode occurs.
Chinese Translation
基于线性投影的概念擦除方法假设特征占据可分离的子空间。我们证明这一假设在密集叠加情形下并不成立:当两个特征被强制纳入共享同一子空间的对跖对时,最先进的线性擦除方法会同时破坏两个特征,而不仅是目标特征。通过梯度下降训练的网络则会以非线性方式解决该问题,但方式并不统一:根据初始化的不同,网络会收敛到两种截然不同的回路级解之一,我们将其称为镜像解和阴影解。我们将这种分叉现象映射为特征纠缠程度的函数,表明它反映了一种稳定的吸引子结构而非实验设置的产物,并通过有针对性的因果干预证明,两种解都完好保留了被擦除特征表征的实质性、可测量的痕迹,只需一次标量修补即可恢复,无需任何进一步训练。这与近期在大语言模型遗忘研究中经验观察到的失效模式相呼应——抑制而非删除使得被遗忘的知识得以重新浮现;我们的结果为该失效模式为何发生提供了机制层面且经因果验证的解释。
cs.LG / 60 / 2609.27594

Efficient Linear Bandits via Cluster-Aware Sketching

基于聚类感知稀疏化的高效线性老虎机算法
Yang, Hantao, Xie, Hong, Lian, Defu
Abstract
We study the problem of computational efficiency for linear bandits in high-dimensional settings with a finite arm set. In linear bandits, the increase in the dimension $d$ of the feature vectors leads to growing computational costs of $O(d^2)$ at each round of update. Traditional sketching-based methods such as SOFUL reduce computation via fixed-size matrix sketching, yet run the risk of incurring vacuous linear regret when the spectral tail of the data is heavy and the sketch size is inadequately selected. To guarantee regret convergence and effectively reduce computational costs, we introduce a clustering mechanism and propose the Cluster Sketch Linear Bandit (CS-LB) algorithm. Our method preserves the full covariance information in each cluster to guarantee robust sublinear regret without spectral-tail vulnerabilities, performs cluster switching by assigning a sentinel for each cluster, and reduces per-round update computation to $O(l^2d)$ via a tunable sketch size $l<d$. Experiments on synthetic datasets demonstrate that our method consistently maintains a favorable trade-off between efficiency and regret.
Chinese Translation
我们研究了在具有有限动作集的高维设置下线性老虎机(linear bandits)的计算效率问题。在线性老虎机中,特征向量维度 $d$ 的增加会导致每轮更新计算成本增长至 $O(d^2)$。传统的基于稀疏化(sketching)的方法(如 SOFUL)通过固定规模的矩阵稀疏化来降低计算量,但当数据的谱尾部较重且稀疏化规模选取不当时,可能产生无意义的线性遗憾(regret)。为了保证遗憾收敛并有效降低计算成本,我们引入了一种聚类机制,提出了聚类稀疏线性老虎机(Cluster Sketch Linear Bandit, CS-LB)算法。我们的方法在每个簇中保留完整的协方差信息,从而保证稳健的次线性遗憾,避免谱尾部带来的脆弱性;通过为每个簇设置一个哨兵来实现簇切换;并借助可调的稀疏化规模 $l<d$,将每轮更新的计算量降低至 $O(l^2d)$。在合成数据集上的实验表明,我们的方法在效率与遗憾之间始终保持着良好的权衡。
cs.LG / 61 / 2609.27633

Pheno-GS: Phenoscape-scale Geodesic Sinkhorn

Pheno-GS:表观景观尺度的测地Sinkhorn方法
Wilkinson, Alistair, Tape, Christopher J., Krishnaswamy, Smita
Abstract
High-throughput single-cell data is now collected across large patient cohorts. Understanding patient-level heterogeneity from cellular-level data motivates phenoscaping: embedding each single-cell distribution as a "datapoint," with distances given by optimal transport (OT). Computing geometry-aware OT at this scale, between all pairs of patient datasets, remains an open challenge, since existing methods either rely on Euclidean ground metrics that distort manifold structure or fail under sparse, unevenly sampled, or large-scale data. We present \textbf{Pheno-GS} (Phenoscape-scale Geodesic Sinkhorn), which computes accurate, scalable geodesic transport distances under noisy, unbalanced, large-scale settings via three components: ($1$) graph connectivity regularization for well-defined geodesics on sparse/disconnected manifolds; ($2$) an unbalanced OT formulation via KL marginal penalties; and ($3$) a batched matrix algorithm computing all pairwise distances in one heat diffusion (over $200 \times$ faster than Geodesic Sinkhorn for $500$ distributions). We validate Pheno-GS on synthetic benchmarks and a CyTOF perturbation dataset.
Chinese Translation
高通量单细胞数据目前已在大型患者队列中广泛采集。为了从细胞层面的数据理解患者层面的异质性,我们提出“表观景观构建”(phenoscaping)的概念:将每个单细胞分布作为一个“数据点”进行嵌入,其间的距离由最优传输(Optimal Transport, OT)给出。在如此大的规模下,计算所有患者数据集两两之间具有几何感知能力的OT距离仍然是一个开放性挑战,因为现有方法要么依赖会扭曲流形结构的欧氏基底度量,要么在稀疏、采样不均匀或大规模数据下失效。我们提出Pheno-GS(Phenoscape-scale Geodesic Sinkhorn,表观景观尺度测地Sinkhorn),通过三个组成部分在噪声、不平衡、大规模设定下计算精确且可扩展的测地传输距离:(1)图连通性正则化,以保证稀疏或不连通流形上测地线的良定义性;(2)通过KL边际惩罚实现不平衡OT的建模;(3)一种批处理矩阵算法,可在一次热扩散过程中计算所有成对距离(对于500个分布,比Geodesic Sinkhorn快200倍以上)。我们在合成基准数据和CyTOF扰动数据集上验证了Pheno-GS的有效性。
cs.LG / 62 / 2609.27637

Learning Local Heterogeneity and Cross-Region Context for Large-Scale Traffic Forecasting

面向大规模交通预测的局部异质性与跨区域上下文学习
Feng, Qi, Wang, Zidong, Li, Bo, Gao, Xiaoguang, Zhang, Jiayu, Wang, Chenfeng, Wan, Kaifang
Abstract
Traffic flow forecasting is essential to intelligent transportation systems. Large-scale traffic forecasting requires jointly modeling local spatial dependencies and cross-region context.Spatial dependencies between geographically neighboring nodes are heterogeneous due to differences in road identity and travel direction, while acquiring global information through allpairs node interactions incurs substantial computational costs. Therefore, capturing local heterogeneity while efficiently acquiring long-range context remains an important challenge in largescale traffic forecasting. To address these challenges, we propose LoReST, a Local-Region Spatial Temporal network that models spatial dependencies at two complementary granularities: node neighborhoods and road network regions. Specifically, relation-aware local aggregation captures heterogeneous dependencies within geographic neighborhoods through road and direction specific feature transformations. Cross-region interaction constructs region representations through mean pooling, exchanges long range context via inter-region attention, and broadcasts it back to nodes. By integrating local information aggregation with crossregion interaction, LoReST is able to effectively achieve spatial dependency learning in large-scale road networks. Experiments on four datasets of the LargeST benchmark show average relative reductions of 4.78%, 3.60%, and 5.75% in MAE, RMSE, and MAPE, respectively.
Chinese Translation
交通流预测对智能交通系统至关重要。大规模交通预测需要联合建模局部空间依赖与跨区域上下文。由于道路属性和行驶方向的差异,地理上相邻节点之间的空间依赖具有异质性,而通过全对节点交互获取全局信息会带来巨大的计算开销。因此,如何在捕获局部异质性的同时高效获取长程上下文,仍然是大规模交通预测中的一个重要挑战。为解决这些问题,我们提出了LoReST(Local-Region Spatial Temporal network),一种在两个互补粒度上建模空间依赖的局部—区域时空网络:节点邻域和路网区域。具体而言,关系感知的局部聚合通过针对道路和方向的特定特征变换,捕获地理邻域内的异质性依赖;跨区域交互通过均值池化构建区域表示,借助区域间注意力交换长程上下文,并将其广播回各个节点。通过将局部信息聚合与跨区域交互相结合,LoReST能够有效地实现大规模路网中的空间依赖学习。在LargeST基准的四个数据集上的实验表明,MAE、RMSE和MAPE分别平均相对降低4.78%、3.60%和5.75%。
cs.LG / 63 / 2609.27657

FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation

FLEET:从Logits熵到文本生成中的增强轨迹
Streltsov, Oleksii, Vitko, Oleksandra
Abstract
Solutions based on large language models (LLMs) often rely on temperature sampling to improve accuracy and stability by aggregating multiple samples from the completion distribution. However, this memoryless approach is inherently suboptimal: because it lacks awareness of prior generations and their evaluations, it produces an increasing proportion of semantically duplicate answers as more samples are drawn, leading to diminishing returns. To address this limitation, we introduce FLEET, a novel method that integrates a memory mechanism into the generation process. FLEET represents each generation as a sparse trajectory through states whose entropy exceeds a predefined threshold and uses these trajectories to infer per-token utility scores that adjust the logits. Benchmark evaluations demonstrate that FLEET achieves the same accuracy as the repeated sampling baseline, with a 3x speedup, and substantially improves accuracy on complex coding tasks (LiveCodeBench Pass@32 increases from 59.9% to 66.2%) under the same budget. Furthermore, in the greedy-decoding configuration evaluated here, the approach is deterministic and uses a single calibration pass to derive its principal hyperparameters, requiring only minimal modifications to existing LLM pipelines.
Chinese Translation
基于大语言模型(LLM)的解决方案通常依赖温度采样,通过对补全分布的多个样本进行聚合来提高准确性和稳定性。然而,这种无记忆的方法本质上是次优的:由于缺乏对先前生成结果及其评估的感知,随着采样数量的增加,语义重复的答案比例不断上升,导致收益递减。为解决这一局限,我们提出了FLEET,这是一种将记忆机制融入生成过程的新方法。FLEET将每次生成都表示为一条经过熵超过预设阈值的状态的稀疏轨迹,并利用这些轨迹推断逐token的效用分数以调整logits。基准评估表明,FLEET在达到与重复采样基线相同准确率的同时实现了3倍的速度提升,并在相同预算下显著提高了复杂编程任务的准确率(LiveCodeBench Pass@32从59.9%提升至66.2%)。此外,在本文评估的贪婪解码配置下,该方法具有确定性,仅需一次校准过程即可确定其主要超参数,且只需对现有LLM流水线进行极少修改。
cs.LG / 64 / 2609.27658

Private Decentralized Optimization with Noise Reduction and Bias Correction

具有噪声降低与偏差校正的隐私保护去中心化优化
Fan, Yizhao, Luo, Wenjian, Zhang, Jiaojiao
Abstract
Private decentralized learning is affected by sampling noise, privacy noise, and decentralized bias under heterogeneous data. We propose Private Recursive Decentralized Optimization (PRDO). PRDO uses recursive estimation with same-batch gradient differences to reduce estimation errors caused by sampling and privacy noise, while its Exact Diffusion component corrects decentralized bias arising from data heterogeneity. Our analysis establishes a nonconvex convergence bound without assuming uniformly bounded data heterogeneity across nodes. It further gives a sufficient condition under which recursive gradient differences yield strictly lower query sensitivity than private Exact Diffusion, together with an example that rigorously satisfies this condition. Experiments show improved accuracy over the evaluated baselines.
Chinese Translation
隐私保护的去中心化学习受到采样噪声、隐私噪声以及异构数据下去中心化偏差的影响。我们提出了隐私保护递归去中心化优化算法(Private Recursive Decentralized Optimization, PRDO)。PRDO 利用基于同批次梯度差分的递归估计来降低由采样噪声和隐私噪声引起的估计误差,同时其 Exact Diffusion 组件可以校正由数据异构性导致的去中心化偏差。我们的分析在无需假设各节点数据异构性一致有界的条件下,建立了非凸收敛界。此外,我们给出了一个充分条件,在该条件下递归梯度差分产生严格低于隐私 Exact Diffusion 的查询敏感度,并给出了一个严格满足该条件的示例。实验表明,所提方法在准确率上优于所评估的基线方法。
cs.LG / 65 / 2609.27667

Robust Adversarial Reinforcement Learning with Risk Sensitivity and Critic Consistency Regularization

具有风险敏感性与评论家一致性正则化的鲁棒对抗强化学习
Wu, Jiaxi, Zhang, Tiantian, Wang, Yuxing, Chang, Yongzhe, Wang, Xueqian
Abstract
Reinforcement learning (RL) achieves strong performance in sequential decision-making but remains brittle under dynamic uncertainty and distributional shifts. Robust Adversarial Reinforcement Learning (RARL) improves robustness via worst-case perturbations, but existing approaches frequently suffer from unstable optimization and degraded value estimation. In particular, overly aggressive adversaries can drive the agent toward uninformative failure states, while adversarial perturbations amplify disagreement between double critics and introduce biased value targets. We propose a unified framework, RACER (Risk-sensitive robust Adversarial critic ConsistEncy-regularized Reinforcement learning), that revisits adversarial RL from a risk-sensitive perspective. First, we introduce a state-dependent adversarial objective that adaptively regulates perturbation strength, suppressing harmful disturbances while preserving informative exploration. Second, we propose critic consistency regularization to reduce disagreement between Q-value estimators and stabilize learning. Comprehensive experiments on challenging continuous control benchmarks demonstrate that RACER consistently improves performance, robustness, and training stability over strong robust RL baselines.
Chinese Translation
强化学习(RL)在序贯决策中能够取得优异的性能,但在动态不确定性和分布偏移下仍然表现脆弱。鲁棒对抗强化学习(RARL)通过最坏情况扰动提升了鲁棒性,但现有方法常常面临优化不稳定和价值估计退化的问题。具体而言,过于激进的对抗者会将智能体推向无信息量的失败状态,而对抗扰动会加剧双评论家之间的分歧并引入有偏的价值目标。我们提出了一个统一框架——RACER(风险敏感的鲁棒对抗评论家一致性正则化强化学习),从风险敏感的视角重新审视对抗强化学习。首先,我们引入了一种状态相关的对抗目标,能够自适应地调节扰动强度,在抑制有害干扰的同时保留有信息量的探索。其次,我们提出评论家一致性正则化方法,以减少Q值估计器之间的分歧并稳定学习过程。在具有挑战性的连续控制基准上的大量实验表明,与强鲁棒强化学习基线相比,RACER在性能、鲁棒性和训练稳定性方面均取得了一致的提升。
cs.LG / 66 / 2609.27679

What Do Tabular Foundation Models Compute In Context? In-Situ Representation Refinement through Attention-Gated Updates

表格基础模型在上下文中计算了什么?通过注意力门控更新实现的原位表示精炼
Zhou, Tian, Jin, Beverly, Yang, Linxiao, Wang, Xue, Wang, Wenwei, Peng, Bingqing, Ye, Mengni, Gu, Jinjie, Sun, Liang
Abstract
What reusable computation should a tabular foundation model learn when every table defines a new supervised task? We develop in-situ representation refinement: support labels guide updates to the episode's representations, and these updates transfer to unlabeled queries without changing model parameters. A regularized leave-one-out objective yields a support correction and its query extension. The leading term separates attention-based reading from state-dependent scaling, motivating RefineICL: an attention-gated, FFN-free contextual stack with selected low-rank feature interaction and typed memory. RefineICL-L24 reaches 0.93836 OVR-AUC and 0.87173 accuracy on AMLB29. A benchmark-informed continuation reaches 1644.8 Elo on the 38-dataset TabArena snapshot, 31.4 Elo above TabPFN-3 under the same evaluation. It also improves all four reported metrics over TabPFN-v3 on both TabZilla views. In a matched 100K-update depth grid, an expanded FFN gives no consistent validation benefit and uses 60.2% more peak inference memory at L8. Internal interventions show that support representations are more than a static source of labels: removing one intermediate support update, while preserving the query output, increases final query cross-entropy in all 72 tested episodes. Together, the derivation and interventions explain how attention-gated updates can construct a task-specific predictor in context.
Chinese Translation
当每个表格都定义一个新的监督任务时,表格基础模型应当学习何种可复用的计算?我们提出了原位表示精炼(in-situ representation refinement)方法:支持集标签引导对当前回合表示的更新,而这些更新能够迁移到未标注的查询上,且无需改变模型参数。一个正则化的留一目标函数产生支持集修正及其在查询上的扩展。其主导项将基于注意力的读取与依赖状态的缩放分离开来,由此启发我们提出 RefineICL:一种带注意力门控、无前馈网络(FFN)的上下文处理堆栈,并配备了精选的低秩特征交互和类型化记忆。RefineICL-L24 在 AMLB29 上达到了 0.93836 的 OVR-AUC 和 0.87173 的准确率。基于基准信息进行继续训练后,该方法在包含 38 个数据集的 TabArena 快照上达到了 1644.8 的 Elo 分数,在相同评估条件下比 TabPFN-3 高出 31.4 Elo。它还在 TabZilla 的两种视图上均全面超越了 TabPFN-v3 的四项报告指标。在一组匹配的 10 万步更新深度的对照实验中,扩大 FFN 规模并未带来一致的验证集收益,反而在 L8 下多消耗了 60.2% 的推理峰值内存。内部干预实验表明,支持集表示不仅仅是静态的标签来源:在保持查询输出不变的前提下,移除某一中间支持集更新,会在全部 72 个测试回合中增加查询的最终交叉熵。综上,理论推导与干预实验共同解释了注意力门控更新如何在上下文中构建任务专属的预测器。
cs.LG / 67 / 2609.27735

NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision Transformers

NS-ATTENTION:视觉Transformer中注意力输出的Newton-Schulz变换
Jiang, Xiaohe, Zhang, Guoqiang, Huang, Tianjin, Mu, Ronghui
Abstract
Newton-Schulz (NS) iteration has recently been used in the Muon optimizer to transform update matrices during the training of large language models. Motivated by its spectral effect, we investigate applying NS directly to Transformer attention representations. We introduce Newton-Schulz Attention (NS-Attn.), a parameter-free transformation applied to the output of each attention head. Each head output is arranged as a feature-by-token matrix and normalized by its Frobenius norm. We then apply a finite NS polynomial step and restore the original norm. The objective is to reduce spectral concentration and increase effective rank before standard head merging and output projection. Across ViT and Swin on CIFAR-10 and CIFAR-100, NS-Attn. improves final-epoch accuracy in all 12 matched-seed comparisons, with mean gains of 0.25--0.83 percentage points. ViT ablations show higher mean accuracy with one iteration than with two. Spectral analysis further shows reduced leading-eigenvalue concentration and increased effective rank. These gains incur additional inference latency.
Chinese Translation
Newton-Schulz(NS)迭代最近被应用于Muon优化器中,用于在大语言模型训练过程中变换更新矩阵。受其谱效应的启发,我们研究了将NS直接应用于Transformer注意力表示的方法。我们提出Newton-Schulz注意力(NS-Attn.),这是一种应用于每个注意力头输出的无参数变换。每个头的输出被组织为特征-令牌(feature-by-token)矩阵,并通过其Frobenius范数进行归一化。随后,我们执行有限步的NS多项式迭代,并恢复原始范数。其目标是在标准的头合并和输出投影之前,降低谱集中度并提高有效秩。在CIFAR-10和CIFAR-100数据集上对ViT和Swin的实验中,NS-Attn.在全部12组相同随机种子的对比实验中均提升了最终轮次的准确率,平均增益为0.25至0.83个百分点。ViT的消融实验表明,单次迭代的平均准确率高于两次迭代。谱分析进一步表明,主导特征值的集中度降低,有效秩增加。这些增益会带来额外的推理延迟。
cs.LG / 68 / 2609.27739

MENO: Memory-Efficient Neural Operator

MENO:内存高效的神经算子
Xu, Shengyang, Zhang, Weijun, Hu, Jun, Jin, Pengzhan
Abstract
We propose the Memory-Efficient Neural Operator (MENO) as a high-performance PDE neural solver based on the Manifold Function Encoder (MFE). MENO features three primary advantages: (1) MENO has a significantly smaller memory footprint and much faster training speed than other popular architectures, with the memory footprint being independent of the data resolution, and therefore holds the potential for scaling up to large-scale models. (2) MENO can accept PDE inputs of arbitrary form, including arbitrary geometric domains and arbitrary discretizations. In particular, it is capable of handling cross-geometry scenarios, i.e., where the input functions and the output solutions are defined on different manifolds. (3) MENO exhibits strong generalization capability, and achieves the best accuracy on most of the benchmarks we tested, compared with the results reported in the literature. The code is available on GitHub at https://github.com/jpzxshi/MENO, and all numerical examples in this paper can be run with a single command to reproduce the reported results.
Chinese Translation
我们提出了内存高效神经算子(Memory-Efficient Neural Operator, MENO),这是一种基于流形函数编码器(Manifold Function Encoder, MFE)的高性能偏微分方程(PDE)神经求解器。MENO具有三个主要优势:(1)与其他流行的网络架构相比,MENO的内存占用显著更小、训练速度更快,且其内存占用与数据分辨率无关,因此具备扩展到大规模模型的潜力。(2)MENO可以接受任意形式的PDE输入,包括任意几何域和任意离散化方式。特别地,它能够处理跨几何场景,即输入函数与输出解定义在不同流形上的情形。(3)MENO表现出强大的泛化能力,与文献中报告的结果相比,在我们测试的大多数基准上取得了最优精度。代码已发布于GitHub:https://github.com/jpzxshi/MENO,本文中的所有数值示例均可通过单条命令运行以复现所报告的结果。
cs.LG / 69 / 2609.27741

Limiting-Kernel Q($\lambda$): Bridging Short and Long Horizons

极限核Q($\lambda$):连接短时域与长时域
Ok, Tolga, Kolarijani, Arman Sharifi, Esfahani, Peyman Mohajerin, Kolarijani, Mohamad Amin Sharifi
Abstract
In value-based reinforcement learning, improving the accuracy of policy evaluation has been shown to improve downstream policy optimization performance. The widely adopted family of approximations relying on $n$-step truncation yields computationally efficient value estimators but is inherently limited to a short evaluation horizon. In contrast, methods that exploit the global structure of the transition dynamics can accelerate policy evaluation, but their memory and computational requirements often limit scalability to large or continuous state spaces. To reconcile these limitations, we introduce Limiting-Kernel Q($\lambda$) (LKQL), an off-policy value estimator that combines $n$-step truncation with a long-horizon approximation based on the limiting kernel (LK). LKQL has the same order of complexity as $n$-step estimators and integrates directly into both on- and off-policy actor-critic algorithms. We prove that, under aperiodicity and in the near-on-policy regime, the operator underlying LKQL improves the policy evaluation convergence rate over its truncated counterpart for sufficiently large $n$, and that LKQL itself converges almost surely to the optimal values in finite Markov decision processes (MDPs) under a fixed behavior policy. On the MuJoCo continuous-control benchmark, we show that LKQL improves over $n$-step baselines in most settings, particularly on long-horizon tasks.
Chinese Translation
在基于价值的强化学习中,提高策略评估的精度已被证明可以改善下游策略优化的性能。广泛采用的基于 $n$ 步截断的近似方法族能够产生计算高效的价值估计器,但其本质上受限于较短的评估时域。相比之下,利用转移动力学全局结构的方法可以加速策略评估,但其内存和计算需求往往限制了其在大规模或连续状态空间上的可扩展性。为了调和这些局限,我们提出了极限核Q($\lambda$)(Limiting-Kernel Q($\lambda$), LKQL),这是一种结合了 $n$ 步截断与基于极限核(limiting kernel, LK)的长时域近似的离策略(off-policy)价值估计器。LKQL 的复杂度与 $n$ 步估计器处于同一量级,并可直接集成到在策略(on-policy)与离策略的 actor-critic 算法中。我们证明,在非周期性条件以及接近在策略的情形下,对于足够大的 $n$,LKQL 所依据的算子在策略评估收敛速度上优于其截断版本;并且在固定行为策略下,LKQL 本身在有限马尔可夫决策过程(MDP)中几乎必然收敛到最优价值。在 MuJoCo 连续控制基准测试中,我们表明 LKQL 在大多数设置下优于 $n$ 步基线方法,尤其是在长时域任务上。
cs.LG / 70 / 2609.27760

Backdoors Leave Structural Traces: FedMAST for Backdoor Detection and Containment in Federated Learning

后门会留下结构性痕迹:用于联邦学习中后门检测与遏制的方法FedMAST
Subramanian, Srinivasan, Islam, Kazi Aminul, Khan, Md. Abdullah Al Hafiz
Abstract
Federated learning enables distributed training of a shared model without requiring clients to share their raw data. However, its reliance on the integrity of the client-submitted updates exposes the global model to stealthy backdoor poisoning. Although existing defenses often inspect isolated evidence sources, stealth-constrained attacks can adapt to these signals. In this paper, we show that such attacks can suppress isolated anomaly signals, but their poisoned updates still leave residual structural traces. We propose FedMAST, a Federated Multi-Axis Structural Tracing defense for backdoor detection in federated learning. FedMAST scores client updates using complementary structural, spectral, and historical evidence and then applies tiered filtering and round-level containment to limit adversarial influence. To capture traces that isolated signals may miss, FedMAST uses squeeze-pair coherence scoring to expose coupled feature distortions and signed spectral-drift tracking to reveal persistent directional changes over time. Across six federated backdoor attacks, namely Constrain-and-Scale, Neurotoxin, BC-Layers, LGA, DBA, and 3DFed, FedMAST achieves lower ASR than baseline defenses in all nine evaluated attack--defense comparisons. Across the complete 200-round runs, it attains an average ASR of 1.51% while maintaining 94.84% average main-task accuracy. Under the method-aware CovertLayers attack, FedAvg, MultiKrum, AlignIns, and FLAME yield full-run ASRs of 100.00%, 99.67%, 99.53%, and 32.84%, respectively. FedMAST achieves the lowest ASR among all evaluated methods, reducing it to 1.53% while maintaining 92.26% main-task accuracy.
Chinese Translation
联邦学习(Federated Learning)使得各方能够在无需共享原始数据的情况下协同训练一个共享模型。然而,其对客户端提交更新完整性的依赖,使全局模型容易遭受隐蔽的后门投毒攻击。现有防御方法通常只检查孤立的证据来源,而具有隐蔽性约束的攻击可以适应这些信号。本文表明,此类攻击虽然能够抑制孤立的异常信号,但其投毒更新仍会留下残留的结构性痕迹。我们提出了FedMAST,一种用于联邦学习后门检测的联邦多轴结构追踪防御方法。FedMAST利用互补的结构性、频谱性和历史性证据对客户端更新进行评分,然后应用分级过滤和轮次级遏制来限制对抗性影响。为了捕捉孤立信号可能遗漏的痕迹,FedMAST采用压缩对一致性评分(squeeze-pair coherence scoring)来暴露耦合的特征畸变,并采用带符号频谱漂移追踪(signed spectral-drift tracking)来揭示随时间持续的定向变化。在六种联邦后门攻击——即Constrain-and-Scale、Neurotoxin、BC-Layers、LGA、DBA和3DFed——上,FedMAST在全部九组攻击-防御对比评估中均取得比基线防御更低的后门攻击成功率(ASR)。在完整的200轮训练中,其平均ASR仅为1.51%,同时保持94.84%的平均主任务准确率。在具备方法感知能力的CovertLayers攻击下,FedAvg、MultiKrum、AlignIns和FLAME的全轮ASR分别为100.00%、99.67%、99.53%和32.84%,而FedMAST在所有评估方法中取得最低的ASR,将其降至1.53%,同时保持92.26%的主任务准确率。
cs.LG / 71 / 2609.27764

Learning to Detect Symbolic Failure: Machine Learning and the Limits of Black-Scholes

学习检测符号性失效:机器学习与Black-Scholes模型的局限
Huang, Juli, Cheng, Jake, Lu, Rupert
Abstract
We treat options pricing as a representation problem: can machine learning detect systematic deviations from Black-Scholes using 2.6M real option contracts? We compare three regimes: learned abstract embeddings (Kernel PCA), preserved domain structure (tree-based ensembles), and neural network validation. Tree-based methods outperform kernel dimensionality reduction by 21.5 percentage points (93.8% vs 72.3%), and domain-expert features (Greeks, moneyness) outperform engineered features. NN-based and BS-based deviation labels agree 99.9974% of the time, suggesting deviations reflect market structure rather than model artifact. We conclude that in domains with expert-designed symbolic features, preserving structure beats learning abstractions. We make no claim of exploitable mispricings.
Chinese Translation
我们将期权定价视为一个表示学习问题:机器学习能否利用260万份真实期权合约检测出对Black-Scholes模型的系统性偏离?我们比较了三种方法:学习到的抽象嵌入(Kernel PCA)、保留领域结构的方法(基于树的集成模型)以及神经网络验证。基于树的方法在性能上超过核降维方法21.5个百分点(93.8%对72.3%),且领域专家特征(Greeks希腊字母、价内外程度moneyness)优于人工设计的特征。基于神经网络与基于Black-Scholes的偏离标签在99.9974%的情况下一致,表明这种偏离反映的是市场结构而非模型伪影。我们得出结论:在存在专家设计的符号特征(symbolic features)的领域中,保留结构优于学习抽象表示。我们不声称存在可利用的定价偏差。
cs.LG / 72 / 2609.27819

When Adaptation Hurts: Split Sensitivity and Person-Level Negative Transfer in Federated Wearable Onboarding

当自适应带来伤害:联邦可穿戴设备冷启动中的分裂敏感性与个体层面负迁移
Aftab, Rahil, Rakesh, Vineet Kumar, Mazumdar, Soumya, Samanta, Tapas
Abstract
Federated wearable models eventually serve people absent from source training, but favorable average accuracy does not establish that unlabeled onboarding helps each person. We evaluate six core onboarding strategies on five wearable datasets under a leakage-controlled protocol that fixes source checkpoints, estimates normalization from source data only, separates calibration from evaluation recordings, and performs inference over held-out people rather than windows, devices, or random seeds. Completing all eligible HHAR and PAMAP2 outer-person rotations materially changes the conclusion obtained from the original frozen fold. On HHAR, balanced accuracy on that single person is 95.6-97.2% across methods versus 78.3-83.0% over all nine users, a reduction of 13.8-17.8 percentage points (pp). The displayed mean leader changes on both datasets, while paired leader-runner bootstrap intervals include zero and do not resolve a superior method. No adaptive core mechanism combines positive mean gain in all five datasets with zero seed-averaged person-level losses greater than 2 percentage points (pp). FedBN has one such loss and ATP-style adaptation has eight; Feature-only has none after seed averaging, but its exact one-sided 95% upper bound is 7.6%. A complementary seed-person stress audit records 4, 22, and 10 harmful realizations out of 114 for FedBN, ATP-style, and Feature-only, respectively; these are repeated realizations, not independent participants. Tail quality, calibration availability, and fall-window specificity reveal additional failures hidden by mean accuracy. The study therefore provides an auditable development benchmark and failure map rather than a universal-superiority or deployment-safety claim.
Chinese Translation
联邦可穿戴设备模型最终要服务于未参与源训练的人群,但良好的平均准确率并不能证明无标注的冷启动(onboarding)对每个人都有帮助。我们在一个防泄漏协议下,于五个可穿戴数据集上评估了六种核心冷启动策略,该协议固定源模型检查点、仅从源数据估计归一化参数、将校准数据与评估记录分离,并在留出个体(而非窗口、设备或随机种子)层面进行推理。完成所有符合条件的 HHAR 和 PAMAP2 外部个体轮换后,由原始冻结折所得到的结论发生了实质性变化。在 HHAR 上,单一用户的平衡准确率在各方法下为 95.6–97.2%,而在全部九名用户上仅为 78.3–83.0%,下降 13.8–17.8 个百分点。两种数据集上显示的平均最优方法均发生变化,而配对最优-次优 bootstrap 置信区间包含零,无法分辨出更优的方法。没有任何自适应核心机制能够同时在全部五个数据集上取得正的平均增益,且种子平均后的个体层面损失均不超过 2 个百分点。FedBN 存在一次此类损失,ATP 式自适应存在八次;Feature-only 在种子平均后没有此类损失,但其精确的单侧 95% 上界为 7.6%。一项互补的种子-个体压力审计显示,FedBN、ATP 式和 Feature-only 分别在 114 次中记录了 4 次、22 次和 10 次有害结果;这些是重复的实现结果,而非独立的参与者。尾部质量、校准可用性以及跌倒窗口特异性揭示了被平均准确率所掩盖的额外失败。因此,本研究提供的是一个可审计的开发基准与失败图谱,而非通用优越性或部署安全性的声明。
cs.LG / 73 / 2609.27825

CAST: Context- and Anomaly Structure-Conditioned Time Series Anomaly Generation

CAST:基于上下文与异常结构条件的时间序列异常生成
Zhang, Haochen, Peng, Jie, Sui, Songyuan, Huang, Yu-Chao, Zhu, Xiangqi, Chen, Tianlong
Abstract
Anomalous time series play a critical role in safety-critical domains, yet they are inherently scarce, heterogeneous, and costly to obtain. Existing time series generation methods predominantly focus on synthesizing normal data, providing limited value when anomalous samples are needed. We identify two fundamental challenges in anomaly generation: (i) the scarcity of anomaly data, and (ii) the heterogeneous morphological characteristics of anomalies. To address these challenges, we propose CAST, a Context- and Anomaly Structure-conditioned Time series anomaly generation framework with principled two-stage pretraining and finetuning strategy. In pretraining stage, we leverage abundant normal time series data to learn underlying system dynamics and substantially mitigate the limited availability of anomaly data. During finetuning, CAST explicitly conditions the generator on learned anomaly structure representations, enabling it to capture heterogeneous anomaly morphologies under similar contextual conditions. Extensive experiments on multiple real-world univariate and multivariate datasets demonstrate that CAST consistently outperforms state-of-the-art anomaly generation methods in terms of both generation fidelity and downstream task utility, highlighting the effectiveness of the proposed approach.
Chinese Translation
异常时间序列在安全关键领域发挥着重要作用,但其本质上稀缺、异构且获取成本高昂。现有时间序列生成方法主要集中于合成正常数据,在需要异常样本时价值有限。我们识别出异常生成中的两个根本性挑战:(i)异常数据的稀缺性,以及(ii)异常的异构形态特征。为应对这些挑战,我们提出了CAST——一个基于上下文与异常结构条件的时间序列异常生成框架,并采用有原则的两阶段预训练与微调策略。在预训练阶段,我们利用丰富的正常时间序列数据学习底层系统动态,从而显著缓解异常数据可用性有限的问题。在微调阶段,CAST显式地将生成器条件化于学习到的异常结构表示之上,使其能够在相似上下文条件下捕捉异构的异常形态。在多个真实世界单变量和多变量数据集上的大量实验表明,CAST在生成保真度和下游任务效用两方面均持续优于最先进的异常生成方法,凸显了所提方法的有效性。
cs.LG / 74 / 2609.27833

From Reasoning Strings to Partial Orders: Verifier-Certified Rule Transport through Quotient Policy Optimization

从推理字符串到偏序关系:基于商策略优化的验证器认证规则迁移
Xie, Bang, Liu, Hao, Peng, Zhiyuan, Yin, Xin, Ying, Chenhao, Luo, Yuan, Zhang, Senjian, Chen, Wei
Abstract
Many computations admit several valid execution orders because independent subgoals or disjoint state updates can commute. Reinforcement learning with verifiable rewards usually treats each successful trace as a separate token sequence, so serialization choices can be mistaken for logical dependencies. We introduce Verifier-Certified Rule Transport (VCRT), which replays adjacent operation pairs with native verifiers. Pairs whose two orders are accepted and reach the same canonical state provide commutation certificates; rejected or state-changing reversals provide anti-diamonds. VCRT uses anti-diamonds to preserve genuine prerequisites and assigns policy credit to the total probability mass of each certified orbit. It also constrains post-swap consistency, source retention, and policy drift. We evaluate leave-one-environment-out transfer across ProofWriter, CLRS, and Lean through a shared anonymized relation-graph interface. All training and checkpoint decisions are frozen before held-out evaluation, which uses one greedy trajectory per item without search or verifier feedback. VCRT obtains a 77.60% macro pass rate versus 64.53% for the strongest matched baseline, a paired gain of 13.06 points (95% bootstrap CI [12.58, 13.54]). Lean accounts for most of this gain at 33.49 points, while ProofWriter and CLRS improve by 2.85 points on average. Mechanism tests consistently favor anti-diamond supervision, whereas No-Orbit is statistically indistinguishable from full VCRT. The evidence does not establish a general benefit from exact orbit aggregation.
Chinese Translation
许多计算问题存在多种有效的执行顺序,这是因为相互独立的子目标或互不相交的状态更新可以交换次序。具有可验证奖励的强化学习通常将每条成功轨迹视为独立的token序列,因而串行化的选择可能被误认为逻辑上的依赖关系。我们提出了验证器认证规则迁移(Verifier-Certified Rule Transport, VCRT),利用原生验证器重放相邻操作对。对于两种顺序均被接受且到达相同规范状态的操作对,可提供交换性证书;被拒绝或改变状态的反序则提供“反菱形”(anti-diamonds)。VCRT利用反菱形来保留真实的前置依赖,并将策略奖励分配给每个认证轨道(orbit)的总概率质量。它还约束了交换后的一致性、源保留以及策略漂移。我们通过一个共享的匿名化关系图接口,在ProofWriter、CLRS和Lean上评估了留一环境外的迁移能力。所有训练和检查点决策均在保留集评估之前冻结,评估对每个条目仅使用一条贪心轨迹,不进行搜索或验证器反馈。VCRT取得了77.60%的宏平均通过率,而最强的匹配基线为64.53%,配对提升为13.06个百分点(95%自助法置信区间为[12.58, 13.54])。其中Lean贡献了大部分增益,达33.49个百分点,ProofWriter和CLRS平均提升2.85个百分点。机制实验一致支持反菱形监督,而No-Orbit消融与完整VCRT在统计上无显著差异。这些证据并不能证明精确轨道聚合具有普遍性的收益。
cs.LG / 75 / 2609.27860

Exact Minimax One-Bit Unbiased Compression: Heavy-Tail Necessity and Finite-Randomness Approximation

精确极小极大一位无偏压缩:重尾必要性有限随机性逼近
Jiang, Tao, Gao, Minbo, Cai, Shaowei
Abstract
A pointwise-unbiased one-bit compressor reconstructs every real input in expectation while transmitting one bit. For a scalar source $P$ with CDF $F$, mean $m$, and $\mathcal J(P)=\int_{\mathbb R}\sqrt{F(r)(1-F(r))}\,dr$, we prove that the infimum of the source-averaged reconstruction second moment over all public-coin one-bit codes unbiased on $\mathbb R$ is $m^2+\mathcal J(P)^2$. For regular full-support sources, a distribution-centered random-threshold code attains this value; a converse over arbitrary randomized binary encoders and an equality analysis characterize every attaining code up to null sets, bit relabeling, and public-seed refinement. For the Gaussian location family $\mathcal N(\mu,\sigma^2)$ with $|\mu|\le c\sigma$, the equal prior on the endpoint means is least favorable and the minimax value is $\sigma^2\Lambda_c^2$. Exact Gaussian minimax optimality forces a critical heavy tail: at the endpoint means, absolute moments are finite exactly for $p<3$, and $\Pr(W>t)=\Theta(t^{-3}/\sqrt{\log t})$. A Cauchy-mixture robustification inflates the second moment by at most $1/(1-\eta)$ while making every positive-order absolute moment finite. Finite-support public randomness with finite decoder means cannot achieve exact unbiasedness on $\mathbb R$, but a bounded-output approximation using exactly $R$ shared random bits has explicit bias and second-moment bounds converging to the minimax constant. Finally, coordinate allocation communicates exactly $B$ bits per Gaussian-gradient query. On Kim's continuous quadratic hard family, the expected optimization guarantee matches the lower bound in its dependence on $(\sigma,d,B,\varepsilon)$, and a finite-variance high-probability bound incurs only a logarithmic confidence factor.
Chinese Translation
逐点无偏的一位压缩器在仅传输一比特的条件下,能期望地重构任意实数输入。对于具有累积分布函数 $F$、均值 $m$ 以及 $\mathcal J(P)=\int_{\mathbb R}\sqrt{F(r)(1-F(r))}\,dr$ 的标量信源 $P$,我们证明:在所有对 $\mathbb R$ 无偏的公共硬币一位编码方案中,信源平均重构二阶矩的下确界为 $m^2+\mathcal J(P)^2$。对于正则的全支撑信源,一种以分布为中心的随机阈值编码方案可达到该值;我们通过对任意随机化二进制编码器的逆命题以及等式分析,在除去零测集、比特重标记和公共种子精炼的意义下刻画了所有可达该值的编码方案。对于满足 $|\mu|\le c\sigma$ 的高斯位置族 $\mathcal N(\mu,\sigma^2)$,端点均值上的等先验是最不利的,其极小极大值为 $\sigma^2\Lambda_c^2$。精确的高斯极小极大最优性迫使编码器必须具有关键性的重尾:在端点均值处,绝对矩仅在 $p<3$ 时有限,且 $\Pr(W>t)=\Theta(t^{-3}/\sqrt{\log t})$。一种柯西混合鲁棒化方法将二阶矩至多放大 $1/(1-\eta)$ 倍,同时使所有正阶绝对矩变为有限。具有有限解码器均值的有限支撑公共随机性无法在 $\mathbb R$ 上实现精确无偏,但一种使用恰好 $R$ 个共享随机比特的有界输出逼近方法具有显式的偏差与二阶矩界,并收敛到极小极大常数。最后,通过坐标分配可在每次高斯梯度查询中恰好传输 $B$ 比特。在 Kim 的连续二次难例族上,期望优化保证在关于 $(\sigma,d,B,\varepsilon)$ 的依赖关系上与下界相匹配,而有限方差的高概率界仅引入一个对数置信因子。
cs.LG / 76 / 2609.27865

What Changed? Drift Detection with Real, Virtual, and Incomparable Diagnosis

什么发生了变化?基于真实、虚拟与不可比较诊断的漂移检测
Oda, Kentaro
Abstract
Sharing a deep encoder does not, by itself, fix the central confound of task-comparison scores. We show that cross-evaluated heads on a frozen shared representation inherit the extrapolation confound of shallow exchange scores: pure input rotations with fixed labels inflate a deep exchange score from about 0 to 0.80, while representation-novelty scores are blind in the complementary direction (flat under label permutations that change the task completely). Transplanting a conditional two-discriminator discrepancy into the embedding space resolves both blind spots: the functional axis stays within +-0.001 under rotations and tracks label-permutation drift mass monotonically. Built into a mixture-of-heads lifecycle, the two-axis gate attains better decision quality with fewer heads than exchange or novelty triggers at a matched training budget. On generalized category discovery, the same chunk-level functional axis separates semantic novelty from photometric shift with AUROC 0.98-0.99 where per-input OOD scores (MSP, Energy, Mahalanobis, KNN) sit near chance for that distinction. All findings replicate across frozen ImageNet-21k ViT-B/16 and self-supervised DINOv2 backbones on CIFAR-100, and extend to residual adapter pools with recurrence, where a null-calibrated novelty trigger never fires on mechanism changes while the two-axis gate handles them with full recurrence reuse. We state explicitly the common-factoring condition under which embedding-space conclusions transfer to the original mechanism.
Chinese Translation
共享深度编码器本身并不能解决任务比较评分中的核心混淆问题。我们证明,在冻结的共享表示之上进行交叉评估的多个任务头会继承浅层交换评分的外推混淆:在标签固定、仅输入发生旋转的情况下,深度交换评分会从约0膨胀至0.80;而表示新颖性评分则在互补方向上存在盲区(在完全改变任务的标签置换下保持平坦)。将条件式双判别器差异度量移植到嵌入空间中可同时解决这两个盲点:功能轴在旋转下保持在±0.001以内,并能单调地追踪标签置换漂移的质量。将双轴门控嵌入到多头混合生命周期中,在匹配的训练预算下,其以更少的头数获得了优于交换或新颖性触发器的决策质量。在广义类别发现任务中,相同的块级功能轴能够区分语义新颖性与光度偏移,AUROC达到0.98–0.99,而单输入级别的OOD评分方法(MSP、Energy、Mahalanobis、KNN)在该区分任务上的表现接近随机水平。所有发现在冻结的ImageNet-21k ViT-B/16与自监督DINOv2骨干网络(CIFAR-100数据集)上均可复现,并可扩展至带循环结构的残差适配器池:在此设置下,零校准的新颖性触发器在机制变化时从不触发,而双轴门控则能在完全复用循环结构的情况下处理这些变化。我们明确阐述了嵌入空间结论可迁移至原始机制的公因子化条件。
cs.LG / 77 / 2609.27866

A Shared Encoder Is Not a Shared Task: Conditional Comparison for Deep Expert Pools

共享编码器不等于共享任务:面向深度专家池的条件化比较
Oda, Kentaro
Abstract
Sharing a deep encoder does not, by itself, fix the central confound of task-comparison scores. We show that cross-evaluated heads on a frozen shared representation inherit the extrapolation confound of shallow exchange scores: pure input rotations with fixed labels inflate a deep exchange score from about 0 to 0.80, while representation-novelty scores are blind in the complementary direction (flat under label permutations that change the task completely). Transplanting a conditional two-discriminator discrepancy into the embedding space resolves both blind spots: the functional axis stays within +-0.001 under rotations and tracks label-permutation drift mass monotonically. Built into a mixture-of-heads lifecycle, the two-axis gate attains better decision quality with fewer heads than exchange or novelty triggers at a matched training budget. On generalized category discovery, the same chunk-level functional axis separates semantic novelty from photometric shift with AUROC 0.98-0.99 where per-input OOD scores (MSP, Energy, Mahalanobis, KNN) sit near chance for that distinction. All findings replicate across frozen ImageNet-21k ViT-B/16 and self-supervised DINOv2 backbones on CIFAR-100, and extend to residual adapter pools with recurrence, where a null-calibrated novelty trigger never fires on mechanism changes while the two-axis gate handles them with full recurrence reuse. We state explicitly the common-factoring condition under which embedding-space conclusions transfer to the original mechanism.
Chinese Translation
共享深度编码器本身并不能消除任务比较分数中的核心混淆因素。我们证明,在冻结的共享表示之上进行交叉评估的各个头(head)会继承浅层交换分数的外推混淆:在标签固定的情况下,纯粹的输入旋转会使深度交换分数从约0膨胀至0.80;而表示新颖性分数则在互补方向上失效(在彻底改变任务的标签置换下保持不变)。将条件化双判别器差异度量移植到嵌入空间中可同时解决这两个盲点:函数轴在旋转下保持在±0.001以内,并随标签置换带来的漂移量单调增长。将该双轴门控嵌入混合头(mixture-of-heads)生命周期中,在相同训练预算下,其决策质量优于基于交换或新颖性触发的机制,且所需头数更少。在广义类别发现(generalized category discovery)任务上,同样的块级(chunk-level)函数轴能够以0.98–0.99的AUROC将语义新颖性与光度偏移区分开来,而逐输入的OOD分数(MSP、Energy、Mahalanobis、KNN)在这一区分上接近随机水平。所有结果均在CIFAR-100上、基于冻结的ImageNet-21k预训练ViT-B/16与自监督DINOv2骨干网络中得到复现,并扩展到带递归结构的残差适配器池:经零校准的新颖性触发器在机制变化时从不触发,而双轴门控则能在完全复用递归的情况下处理这些变化。我们明确给出了嵌入空间结论可迁移至原始机制所需的可公因子化条件。
cs.LG / 78 / 2609.27867

Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel

评估选择决定预测排行榜:来自生产环境市场面板的证据
Islam, Md Rezwanul, Mohammed, Wael
Abstract
A forecasting benchmark reports which method won. We show that the answer is set by the evaluator's choices before any model is fitted. We benchmark 24 forecasting methods and one textbook reference, including six 2025-era time series foundation models, on a production marketplace panel of 1,887 business customers over 67 months. We hold the data, the horizon and the period fixed, and vary only the evaluation design. Three choices each reverse or dissolve a headline conclusion. Changing the unit of analysis from the market total to the individual customer moves our production baseline from second of nineteen, beaten by nothing, to twenty-third of twenty-five. Nineteen of its twenty-four challengers beat it there. Changing how much error is pooled decides whether a Diebold-Mariano test finds anything at all. Scoring prediction intervals rather than point forecasts reorders the field almost completely, with a rank correlation of 0.02 on intermittent demand. We then measure what the deployed system gets from this. Its selection rule captures 55% of the distance between doing nothing and choosing with hindsight. The reversal is not a quirk of our data. We ran the released protocol, unchanged, on the public M5 retail panel. The same baseline shape places first at the market total and last per series, beaten by everything, and a replayed selection rule closes 64.7% of the same floor-to-ceiling distance there. Adding five zero-shot foundation models to that roster changes who wins at the total, not the shape. The bands' blind spot travels too: conformal bands under-cover most on the spikiest items. Splitting our own panel into ever smaller groups turns the contrast into a curve: the baseline's rank worsens at every level of disaggregation. We release the evaluation protocol and report an error of our own that inverted a result before we caught it.
Chinese Translation
预测基准测试报告了哪种方法获胜。我们表明,答案在任何模型拟合之前就已由评估者的选择所决定。我们在一个包含1,887个企业客户、跨度67个月的生产环境市场面板上,对24种预测方法和一个教科书参考方法进行基准测试,其中包括六个2025年时期的时间序列基础模型。我们固定数据、预测期限和时间段,仅改变评估设计。三个选择各自逆转或消解了一个标题性结论。将分析单元从市场总量改为单个客户,使我们的生产基线从十九个中的第二名(未被任何方法击败)变为二十五个中的第二十三名;其24个挑战者中有19个在该设置下击败了它。改变误差汇总方式决定了Diebold-Mariano检验能否发现任何显著性。对预测区间而非点预测进行评分,几乎完全重新排列了各方法的名次,在间歇性需求上的秩相关性仅为0.02。随后,我们测量了已部署系统从中获得的价值:其选择规则捕捉了从什么都不做到事后最优选择之间55%的距离。这种逆转并非我们数据的特例。我们在公开的M5零售面板上原封不动地运行了已发布的协议。同样的基线形态在市场总量上排名第一,但在逐序列评估中排名垫底(被所有方法击败),且回放的选择规则在该数据上弥合了同样的从下限到上限距离的64.7%。向该方法集合中添加五个零样本基础模型改变了市场总量上的获胜者,但未改变这一形态。区间的盲点同样存在:共形区间在最尖峰化的商品上覆盖率最低。将自己的面板拆分为越来越小的分组,会使这种对比变为一条曲线:基线的排名在每一个分解层级上都变得更差。我们公开发布了评估协议,并报告了我们自己曾犯下且在发现前曾颠覆某一结果的一个错误。
cs.LG / 79 / 2609.27883

False-science induction in autonomous scientific discovery

自主科学发现中的伪科学诱导
Liang, Hanbing, Liu, Fujun
Abstract
Closed-loop discovery systems increasingly execute experiments and update decisions autonomously, turning record integrity into part of the experimental apparatus. We show that false-science induction arises when legitimate physical objects and measurements are paired incorrectly, driving neural surrogates to faithfully learn record-induced associations that do not correspond to the true object-outcome relationship while marginal data distributions remain unchanged. Across green fluorescent protein fitness and materials band-gap prediction loops, coherent paired misbinding systematically redirects experimental budgets toward low-performing basins, whereas same-volume random swaps have negligible effects. These observations identify error coherence, rather than raw error frequency, as the primary variable controlling this budget misallocation in the tested loops. The resulting binding identifiability boundary supports monitored-axis quarantines and feedback-conflict triage, which intercept over-concentrated proposals before execution and isolate the corrupted hypothesis axis.
Chinese Translation
闭环发现系统日益自主地执行实验并更新决策,使记录完整性成为实验装置的一部分。我们证明,当合法的物理对象与测量结果被错误配对时,会产生伪科学诱导:神经代理模型忠实地学习到由记录诱导的、与真实对象-结果关系不符的关联,而边缘数据分布却保持不变。在绿色荧光蛋白适应度和材料带隙预测闭环中,相干的成对错误绑定会系统性地将实验预算引向低性能区域,而同体积的随机交换则影响甚微。这些观察表明,在所测试的闭环中,错误的相干性(而非错误发生的原始频率)是控制这种预算错配的主要变量。由此得到的绑定可辨识性边界支持监测轴隔离与反馈冲突分诊方法,可在执行前拦截过度集中的实验提议,并隔离受污染的假设轴。
cs.LG / 80 / 2609.27912

Global tree forecasters collapse at the hierarchical aggregate: a five-panel failure characterization

全局树模型在层级聚合层面失效:基于五个数据集的失效特征刻画
Islam, Md Rezwanul, Mohammed, Wael
Abstract
Global forecasting models pool many series and learn one shared function. Gradient-boosted trees are their most common form. We measure a failure of this design that has not, to our knowledge, been documented. Train a global tree on the individual series of a hierarchy, then ask it for the hierarchical aggregate. The aggregate sits far outside the model's training range, and the forecast collapses. The model under-predicts the total by 30-50x in our production deployment, and by up to 496x in a public M5 reconstruction. The mechanism is known: beyond its training range, a tree predicts a constant. It surfaces at the aggregate because the total dwarfs every training series. The cure is not new. Per-series scaling, the preprocessing step that Montero-Manso and Hyndman (2021) recommend, prevents the collapse. So do a weighted aggregate-level training row and seasonal differencing. Our contribution is the characterization. The collapse reproduces on five panels: a production business-to-business marketplace, a synthetic hierarchy, M5, Australian Tourism, and a public business-buyer panel. It holds on three tree libraries, is invariant across training seeds, and is statistically significant. Its onset is immediate and tracks a simple support bound: a scale gap of only 1.15x already costs a third of the total. No standard configuration change prevents it: pooling every hierarchy level into training fails at scale, and the one knob that fits linear models in the leaves softens it without curing it. Rolling the forecasts forward recursively separates the cures: the aggregate-row cure re-collapses, per-series scaling degrades but stays low, and only seasonal differencing keeps its one-step accuracy unchanged. We close with a three-step procedure for diagnosing and preventing the failure in deployed systems.
Chinese Translation
全局预测模型汇总多条时间序列并学习一个共享函数,梯度提升树(gradient-boosted trees)是其最常见的形式。我们发现了一种据我们所知尚未被记录过的该设计失效模式:在层级的各条个体序列上训练全局树模型,然后要求它预测层级聚合值。聚合值远超出模型的训练范围,导致预测崩溃。在我们的生产部署中,模型对总量的预测偏低30-50倍;在一个公开的M5重构实验中,偏低高达496倍。其机制是已知的:超出训练范围后,树模型会输出一个常数。该问题在聚合层面显现,因为总量远大于每条训练序列。解决方法并不新颖:Montero-Manso和Hyndman(2021)推荐的逐序列缩放(per-series scaling)预处理步骤可以防止这种崩溃,加权聚合层面训练样本和季节差分同样有效。我们的贡献在于对该失效的系统性刻画。该崩溃在五个数据集上均可复现:一个生产环境的B2B市场平台、一个合成层级数据集、M5、澳大利亚旅游数据集,以及一个公开的企业买家面板数据。该失效在三个树模型库上均存在,在不同训练随机种子下保持一致,且具有统计显著性。失效立即出现,并遵循一个简单的支撑界:仅1.15倍的量级差距就会损失总量的三分之一。没有任何标准配置调整能够防止该失效:将所有层级汇总到训练中在大规模场景下会失败,而在叶节点使用线性模型的唯一可调参数只能缓解问题而无法根治。当采用递归滚动预测时,各种解决方案表现出差异:聚合行训练方法再次崩溃,逐序列缩放性能下降但仍保持较低误差,只有季节差分能保持其单步预测精度不变。最后,我们提出一个三步流程,用于在已部署系统中诊断和预防该失效。
cs.LG / 81 / 2609.27913

Reliable Fusion of Conflicting Experts

冲突专家的可靠融合
Tenali, Pranuthi, Sidheekh, Sahil, Mathur, Saurabh, Saravanan, Vijayalakshmi, Blasch, Erik, Kersting, Kristian, Natarajan, Sriraam
Abstract
We study the problem of aggregating opinions from multiple black-box experts in noisy, conflict-prone settings where expert reliability varies across inputs. Static aggregation methods, such as majority voting, fail to capture this variability and often yield unreliable outcomes under disagreement. We propose a tractable, probabilistic-circuit-based fusion framework that dynamically combines expert responses using context-specific credibility estimates, enabling principled and reliable reasoning. The framework is agnostic to the underlying experts and does not require access to their internal representations or any retraining. We empirically validate our approach on multiple-choice question answering tasks using multiple LLMs as experts, comparing against individual models and static ensemble baselines. Our method consistently improves predictive performance and produces more reliable decisions under conflict, highlighting the effectiveness of context-aware credibility modeling for robust multi-expert fusion.
Chinese Translation
我们研究在噪声环境且易发生冲突、专家可靠性随输入而变化的场景下,聚合多个黑盒专家意见的问题。静态聚合方法(如多数投票)无法捕捉这种可变性,在专家意见不一致时往往产生不可靠的结果。我们提出一种基于概率电路(probabilistic circuit)的可处理融合框架,利用上下文相关的可信度估计动态组合专家的响应,从而实现有原则且可靠的推理。该框架对底层专家不可知,无需访问其内部表示,也无需任何再训练。我们以多个大语言模型(LLM)作为专家,在多项选择题问答任务上对所提方法进行了实证验证,并与单个模型和静态集成基线进行了比较。我们的方法持续提升预测性能,并在冲突情况下产生更可靠的决策,凸显了上下文感知可信度建模在鲁棒多专家融合中的有效性。
cs.LG / 82 / 2609.27917

Spread and Scale: What Determines Whether Test-Time Budget Allocation Pays

扩散与规模:什么决定了测试时预算分配是否有回报
Bae, Jinhyung
Abstract
Neural combinatorial optimization solvers generate many candidate solutions per instance and report the best one found, using the same sample budget for every instance regardless of difficulty. A companion study showed that reallocating a fixed budget toward harder instances can improve solution quality, but that the standard way of measuring this improvement is biased: deciding an allocation and evaluating it on the same data can manufacture an apparent gain even when none exists. This left open what property of a workload determines whether reallocation is worth doing, and whether a policy that spends part of the budget to decide how to allocate the rest still pays once that cost is counted. This paper answers both questions through pre-registered confirmatory experiments -- analysis and verdict criteria fixed before data collection -- across three independently trained solvers and two ways of constructing harder workloads on the traveling salesman problem. Within the workloads we study, the deciding property is how varied the instances within a workload are in difficulty, not how difficult the workload is on average: a uniformly easy or uniformly hard workload offers little room for reallocation, while a mixed workload offers substantial room. A budget-aware policy that pays for its own information about instance difficulty recovers most, though not all, of the improvement available when that information is assumed free. Every experiment was independently recomputed from its written specification, and every correction to an earlier version -- including two that weakened the paper's own claims -- is reported with the direction it moved the conclusion. The paper offers a specific empirical answer and a template for verifying that answer is not an artifact of how it was measured.
Chinese Translation
神经组合优化求解器为每个实例生成大量候选解并报告其中最优者,且对每个实例使用相同的采样预算,而不考虑其难度。一项配套研究表明,将固定预算向更难的实例重新分配可以提升解的质量,但衡量这种提升的标准方式存在偏差:在相同数据上决定分配方案并评估该方案,即使实际上没有提升也可能制造出表面上的收益。由此遗留了两个问题:工作负载的何种属性决定了重新分配是否值得进行,以及一个花费部分预算来决定其余预算如何分配的策略,在计入该成本后是否仍然划算。本文通过预注册的验证性实验——在数据收集前即固定分析与判定标准——回答了这两个问题,实验涵盖三个独立训练的求解器以及两种在旅行商问题(TSP)上构造更难工作负载的方法。在我们所研究的工作负载中,起决定作用的属性是工作负载内实例难度的多样性,而非工作负载的平均难度:均匀简单或均匀困难的工作负载几乎没有重新分配的空间,而混合难度的工作负载则提供了可观的改进空间。一种预算感知策略若需为其获取实例难度信息付出代价,则能恢复(在假定该信息免费时)大部分但非全部的可用提升。每个实验均根据其书面规范进行了独立重算,对早期版本的所有修正——包括两处削弱了本文自身结论的修正——均连同其对结论方向的影响一并报告。本文提供了一个具体的实证答案,以及一个验证该答案并非测量方式产物的模板。
cs.LG / 83 / 2609.27932

Binary Quantized Neural Network Training Is W[1]-Hard Parameterized by Input and Output Dimensions

二值量化神经网络训练关于输入与输出维度的参数化是W[1]-难的
Jiang, Tao, Gao, Minbo, Cai, Shaowei
Abstract
Ganian et al. (ICLR 2026) proved that quantized neural network training is fixed-parameter tractable when parameterized jointly by architecture treewidth, input dimension $\alpha$, and output dimension $\omega$, and left open whether $\alpha+\omega$ alone yields fixed-parameter tractability. We prove that 2-QNNT is W[1]-hard parameterized by $\alpha+\omega$. The hardness already holds with zero error on $D_k=\{(\xi^{(r)},\xi^{(r)}):0\le r\le k\}$, where every input equals its target, $|D_k|=\alpha=\omega=k+1$, and the examples form a coordinatewise prefix chain. It also holds when every non-source bias is fixed to zero. Under the Exponential Time Hypothesis, no algorithm runs in $f(\alpha+\omega)|I|^{o(\alpha+\omega)}$ for any computable $f$. The reduction starts from DAG edge-disjoint paths, converts edge capacity to vertex capacity with a directed line graph, and normalizes the result into a valid layered architecture. The key structural step is a one-flip routing equivalence: on the prefix-chain inputs, nonnegative binary weights make every activation monotone, and each required output transition has a weight-one predecessor making the same transition. Iterating this relation backward extracts a path from the unique changing input, while different transitions yield vertex-disjoint paths. In particular, every neuron on these inputs has only $k+1$ possible activation profiles.
Chinese Translation
Ganian等人(ICLR 2026)证明了量化神经网络训练在以架构树宽、输入维度 $\alpha$ 和输出维度 $\omega$ 联合参数化时是固定参数可解的,并留下一个未决问题:仅以 $\alpha+\omega$ 参数化是否具有固定参数可解性。我们证明,2-QNNT(二值量化神经网络训练)在以 $\alpha+\omega$ 参数化时是W[1]-难的。该难度结果在零误差情形下即已成立,此时训练集为 $D_k=\{(\xi^{(r)},\xi^{(r)}):0\le r\le k\}$,其中每个输入都等于其目标输出,$|D_k|=\alpha=\omega=k+1$,且样例按坐标构成前缀链。当所有非源节点的偏置均固定为零时,该结论同样成立。在指数时间假设(Exponential Time Hypothesis)下,不存在运行时间为 $f(\alpha+\omega)|I|^{o(\alpha+\omega)}$($f$ 为任意可计算函数)的算法。归约从有向无环图(DAG)的边不交路径问题出发,通过有向线图将边容量转化为点容量,并将结果规范化为合法的分层架构。关键的结构性步骤是“一次翻转路由等价”:在前缀链输入上,非负二值权重使每个激活值单调,且每个所需的输出跃迁都存在一个权重为1的前驱神经元产生相同的跃迁。将该关系向后迭代,可从唯一发生变化的输入中提取出一条路径,而不同的跃迁则产生点不交的路径。特别地,在这些输入上,每个神经元只有 $k+1$ 种可能的激活模式。
cs.LG / 84 / 2609.27936

Quality over Quantity: Semi-Supervised Detection of Illicit Bitcoin Flows via Feature Engineering

质量胜于数量:基于特征工程的比特币非法流半监督检测
Smolenkova, Yekaterina, Larionov, Nickolay, Ivanov, Nikolay, Yanovich, Yury
Abstract
Detecting illicit cryptocurrency transactions is hampered by extreme class imbalance, adversarial obfuscation, and a scarcity of reliable labels. While semi-supervised learning (SSL) offers a promising solution by leveraging unlabeled data, we show that its success is not guaranteed by data volume alone but is contingent on data quality. We introduce an SSL framework for detecting illicit Bitcoin flows in Shared Send Mixers (SSM) transactions, built on a comprehensive historical dataset comprising 163 million transactions. Our main conclusion is that the success of SSL depends on data quality rather than volume: high-fidelity features such as KeyLinker address clustering and Shared Send Untangling (SSU) complexity metrics achieve an F1 score of 0.84 on unlabeled data. Finally, we empirically show that common heuristics like One-Time Change (OTC), though abundant, introduce noise, while strategic reliance on higher-fidelity features like KeyLinker is essential. Our work establishes that in blockchain forensics, the path to better performance lies in smarter feature engineering for data quality, not just larger datasets.
Chinese Translation
检测非法加密货币交易面临极端类别不平衡、对抗性混淆以及可靠标签稀缺等障碍。尽管半监督学习(SSL)通过利用无标签数据提供了一种有前景的解决方案,但我们证明其成功并不能仅靠数据量来保证,而是取决于数据质量。我们提出了一个用于检测共享发送混币器(Shared Send Mixers, SSM)交易中非法比特币流的SSL框架,该框架建立在一个包含1.63亿笔交易的全面历史数据集之上。我们的主要结论是:SSL的成功取决于数据质量而非数量,诸如KeyLinker地址聚类和共享发送解缠(Shared Send Untangling, SSU)复杂度指标等高保真特征在无标签数据上可达到0.84的F1分数。最后,我们通过实证表明,诸如一次性找零(One-Time Change, OTC)等常见的启发式方法虽然数量众多,但会引入噪声,而战略性地依赖KeyLinker等高保真特征则至关重要。我们的工作表明,在区块链取证领域,提升性能的路径在于通过更智能的特征工程来提高数据质量,而不仅仅是扩大数据集规模。
cs.LG / 85 / 2609.27954

When Accuracy Gaps Fail to Certify: Auditing Cross-Domain Recalibration of LLM Judges

当准确率差距无法提供认证时:对LLM裁判跨域重新校准的审计
Afrin, Fariya, Shihab, Ibne Farabi
Abstract
A scalar recalibration map fitted for an LLM judge on one task can fail when the task distribution changes, but the source-target accuracy gap is often treated as a proxy for that failure. We test what this gap can predict and what it can certify across thirteen judges, two generators, eight domains, and 1,176 predeclared transfers. After accounting for mean score shift, the gap yields a population lower bound on target calibration error, yet identical gaps can induce opposite transfer outcomes. Exact importance weighting recovers target proper loss under covariate shift, so failure of an estimated weighting pipeline does not by itself establish conditional shift. A finite-sample simultaneous lower certificate converts the population bound into a one-sided rejection rule using audit labels disjoint from evaluation outcomes. The leak-free gap correlation is 0.25 (95% CI [-0.09, 0.55]), falls to 0.09 on the second generator, and does not support a generator-invariant association. The certificate retains nominal coverage but has power 0.13 even at m=1024, whereas target-domain temperature scaling with 16 labels reaches harm rate 0.09, compared with 0.34 for source-fitted Platt scaling. Accuracy gaps are therefore weak warning signals for scalar probability transfer, not deployment certificates.
Chinese Translation
为某个LLM裁判(LLM judge)在单一任务上拟合的标量重新校准映射,在任务分布发生变化时可能失效,但源域与目标域之间的准确率差距常被用作这种失效的替代指标。我们通过十三种裁判、两个生成器、八个领域以及1176次预先声明的迁移,检验了该差距能预测什么、又能认证什么。在扣除均值分数偏移后,该差距可以给出目标域校准误差的总体下界,但相同的差距可能导致截然相反的迁移结果。在协变量偏移下,精确的重要性加权可以恢复目标域的恰当损失(proper loss),因此估计加权流程的失败本身并不能证明存在条件偏移。我们提出一种有限样本同时下界证书,利用与评估结果不相交的审计标签,将总体下界转化为单侧拒绝规则。无泄漏的差距相关性为0.25(95%置信区间为[-0.09, 0.55]),在第二个生成器上降至0.09,不支持生成器不变的相关性。该证书保持了名义覆盖率,但即使m=1024时检验功效也仅为0.13;相比之下,仅用16个标签的目标域温度缩放(temperature scaling)即可将有害率降至0.09,而源域拟合的Platt缩放为0.34。因此,准确率差距只是标量概率迁移的弱预警信号,而非部署认证。
cs.LG / 86 / 2609.27955

CS-WCP: Robust Conformal Sets for LLM-Judge Traffic Shifts with Uncertain Group Proportions

CS-WCP:面向LLM裁判流量偏移与不确定组比例的鲁棒保形集合
Shihab, Ibne Farabi, Afrin, Fariya
Abstract
Prediction sets built from an LLM judge can undercover when deployment traffic changes the prevalence of task or policy groups. Weighted conformal prediction is exact under covariate shift when the density ratio is known, but group proportions must usually be estimated from finite unlabeled samples. We introduce confidence-set weighted conformal prediction (CS-WCP), which constructs simultaneous exact intervals for source and target group masses and returns the union of weighted conformal sets over every compatible ratio vector. For a fixed or independently learned finite partition, CS-WCP attains coverage at least 1-alpha-delta_w-tau_A-kappa, where tau_A measures within-cell covariate mismatch and kappa measures conditional shift. A linear endpoint rule computes the robust union in O(G|Y|) time. Across 336 constructed shared-support traffic shifts, CS-WCP reaches 0.973 mean coverage with 13 point failures, compared with 0.954 and 44 failures for source conformal prediction, at mean binary set sizes 1.74 and 1.65. On 336 natural cross-task transfers, coverage rises from 0.882 to 0.962, but mean set size reaches 1.87 and a size-matched group plug-in baseline is competitive. The method therefore supplies an auditable coverage safeguard under uncertain mixture weights; its value is conservative tail protection, not scalar probability calibration or uniformly smaller sets.
Chinese Translation
当部署流量改变任务组或策略组的占比时,由LLM裁判(LLM judge)构建的预测集合可能出现覆盖不足。加权保形预测(weighted conformal prediction)在密度比已知的情况下可以在协变量偏移下保持精确覆盖,但组比例通常必须从有限的未标注样本中估计。我们提出了置信集加权保形预测(confidence-set weighted conformal prediction, CS-WCP),该方法为源域和目标域的组质量构造同时精确的区间,并返回在所有相容比例向量上加权保形集合的并集。对于固定或独立学习的有限划分,CS-WCP能够实现至少 1-alpha-delta_w-tau_A-kappa 的覆盖率,其中 tau_A 度量单元内的协变量失配,kappa 度量条件偏移。一种线性端点规则可以在 O(G|Y|) 时间内计算该鲁棒并集。在336个构造的共享支持流量偏移上,CS-WCP达到了0.973的平均覆盖率,失败次数为13次;相比之下,源域保形预测的覆盖率为0.954,失败次数为44次,二者的平均二值集合大小分别为1.74和1.65。在336个自然的跨任务迁移场景中,覆盖率从0.882提升至0.962,但平均集合大小达到1.87,且一个大小匹配的组插件基线方法具有竞争力。因此,该方法在不确定混合权重下提供了一种可审计的覆盖保障;其价值在于保守的尾部保护,而非标量概率校准或一致更小的集合。
cs.LG / 87 / 2609.27963

I-SplineFlow: Learning Monotone Spline Stochastic Interpolant Schedulers for Few-Step Generation

I-SplineFlow:学习单调样条随机插值调度器以实现少步生成
Shovon, Md Sakib Hossain, Rahman, Md Rifat Ur, Chowdhury, Md Abtahi Majeed, Min, Yunhong, Choi, Jaesik, Sung, Minhyuk
Abstract
Few-step generation with pretrained diffusion and flow models can be accelerated by lightweight training that optimizes the sampling trajectory rather than the network. A recent approach parameterizes the stochastic interpolant (SI) scheduler as a smooth curve whose control points enforce the three properties an SI scheduler must satisfy: fixed boundary conditions, a monotone signal-to-noise ratio (SNR), and differentiability. Existing parameterizations use globally supported polynomial bases, where every control point moves the whole curve and higher expressiveness needs a higher degree, which couples distant regions of the schedule during optimization. We introduce \emph{I-SplineFlow}, which parameterizes the scheduler with integrated monotone splines (I-splines). I-splines decouple the polynomial degree from the number of mixture weights, so support width and smoothness can be chosen per model at a fixed weight count, and the compactly supported derivative basis makes the scheduler Jacobian orders of magnitude better conditioned than a B\'ezier basis. Boundary conditions and a strictly monotone SNR hold by construction, with no ordering constraint on the parameters and closed-form velocity derivatives. Across diffusion (EDM) and flow (ReFlow, Simple ReFlow) models, I-SplineFlow improves few-step FID over B\'ezier scheduling in most settings, most clearly at the lowest NFEs, and trains in minutes. Ablations show that both the degree freedom and the monotonicity constraint are needed. The code will be released upon acceptance.
Chinese Translation
对于预训练的扩散模型和流模型,少步生成可以通过轻量级训练来加速,即优化采样轨迹而非网络本身。近期的一种方法将随机插值(Stochastic Interpolant, SI)调度器参数化为一条光滑曲线,其控制点需满足 SI 调度器必须具备的三个性质:固定的边界条件、单调的信噪比(SNR)以及可微性。现有参数化方法使用全局支撑的多项式基,其中每个控制点的移动会影响整条曲线,且更高的表达能力需要更高的多项式阶数,这在优化过程中使得调度中相距较远的区域相互耦合。我们提出 I-SplineFlow,它使用积分单调样条(I-splines)对调度器进行参数化。I-splines 将多项式阶数与混合权重的数量解耦,因此在固定权重数量的情况下,可以针对每个模型分别选择支撑宽度和光滑度;同时,其紧支撑的导数基使得调度器的雅可比矩阵的条件数比 Bézier 基好若干个数量级。边界条件和严格单调的 SNR 由构造自动保证,参数无需排序约束,且速度的导数具有闭式表达。在扩散模型(EDM)和流模型(ReFlow、Simple ReFlow)上的实验表明,I-SplineFlow 在大多数设置下均优于 Bézier 调度的少步 FID,在最低 NFE 下优势最为明显,且训练仅需数分钟。消融实验表明,阶数自由度和单调性约束都是必要的。代码将在论文被接收后发布。
cs.LG / 88 / 2609.27964

Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences

线性RNN的缩放定律:更长的序列何时胜过更多的序列
Chen, Ziyan, Zhou, Zhongzhu, Liu, Peilin, Zhou, Ding-Xuan
Abstract
Empirical scaling laws for autoregressive language models relate prediction loss to model size, data size, and optimization compute, but their theoretical origin is still poorly understood in sequential pretraining settings. We study this question in a tractable teacher--student model where a stable latent linear RNN generates trajectories and a sketched linear recurrent student is trained by safeguarded full-batch WSD gradient descent on next-token prediction. The sketch dimension $M$ plays the role of model size, while $N$ independent trajectories of length $P$ provide the training tokens. We allow the innovation and initialization covariances to have different power-law exponents $\alpha$ and $\theta$. The induced design spectrum produces explicit approximation, optimization, and statistical scaling laws separated by spectral crossovers. When $\theta\ge\alpha$, the original one-scale rates $M^{1-\beta_\alpha}$, $R^{(1-\beta_\alpha)/\alpha}$, and $(NP)^{-1}\min\{M,R^{1/\alpha}\}$ are recovered. When $\alpha-2r\le\theta<\alpha$, the heavier initialization tail changes the rates beyond $P$-dependent model and optimization crossovers. The proof uses a covariance event only internally and a globally safeguarded step size on its complement. The variance retains the factor $(NP)^{-1}$, while sequence length also suppresses the initialization transient, so $N$ and $P$ cease to be fully interchangeable in the two-scale regime.
Chinese Translation
自回归语言模型的经验缩放定律将预测损失与模型规模、数据规模和优化计算量联系起来,但其在序列预训练设置下的理论来源仍知之甚少。我们在一个可求解的教师-学生模型中研究这一问题:由一个稳定的潜在线性RNN生成轨迹,同时训练一个经草图(sketched)化的线性递归学生模型,采用带安全保护的全批量WSD梯度下降法进行下一词预测训练。草图维度 $M$ 扮演模型规模的角色,而 $N$ 条长度为 $P$ 的独立轨迹提供训练词元。我们允许新息协方差和初始化协方差具有不同的幂律指数 $\alpha$ 和 $ heta$。由此产生的设计谱在谱交叉点处分隔出显式的近似、优化和统计缩放定律。当 $ heta\ge\alpha$ 时,可恢复原始的单尺度速率 $M^{1-eta_\alpha}$、$R^{(1-eta_\alpha)/\alpha}$ 以及 $(NP)^{-1}\min\{M,R^{1/\alpha}\}$。当 $\alpha-2r\le heta<\alpha$ 时,较重的初始化尾部会在依赖 $P$ 的模型和优化交叉点之外改变这些速率。证明仅在内部使用协方差事件,并在其补集上使用全局安全保护的步长。方差保留因子 $(NP)^{-1}$,同时序列长度还会抑制初始化的瞬态效应,因此在双尺度区域中 $N$ 和 $P$ 不再完全可互换。
cs.LG / 89 / 2609.27982

Riemannian Structure and Optimization for a Class of Low-Parametric Orthogonal Matrices

一类低参数化正交矩阵的黎曼结构与优化
Aliev, Ali, Rakhuba, Maxim
Abstract
In this paper, we are concerned with matrices formed by block-diagonal factors interleaved with fixed permutations -- a flexible family of structured matrices. This class has recently drawn interest in deep learning architectures for its balanced expressivity-efficiency trade-off, yet efficient computational strategies for working with it remain to be found. We approach this problem through Riemannian geometry and examine under what conditions this class admits a smooth manifold structure. For the practically important case of orthogonal two-factor matrices, we derive the essential Riemannian tools and propose efficient algorithms for their implementation. The algorithms leverage automatic differentiation, support parameter sharing within each factor, and avoid explicit dense matrix construction. We test them within the Riemannian optimization framework on the best matrix approximation problem and for parameter-efficient fine-tuning of large language models. Beyond the two-factor setting, we study the geometric and matrix-theoretic properties of factorizations with a larger number of block-diagonal factors.
Chinese Translation
本文研究由块对角因子与固定置换矩阵交错相乘所构成的矩阵——这是一类灵活的结构化矩阵族。该类矩阵因其在表达能力与计算效率之间的良好权衡,近年来在深度学习架构中受到关注,然而针对它的高效计算策略仍有待探索。我们借助黎曼几何方法研究该问题,并考察在何种条件下该矩阵族具有光滑流形结构。针对实际应用中重要的双因子正交矩阵情形,我们推导了必要的黎曼工具,并提出了实现这些工具的高效算法。这些算法利用自动微分技术,支持每个因子内部的参数共享,并避免了显式构建稠密矩阵。我们在黎曼优化框架下,将所提算法应用于最优矩阵逼近问题以及大语言模型的参数高效微调。除双因子情形外,我们还研究了含更多块对角因子的分解的几何性质与矩阵理论性质。
cs.LG / 90 / 2609.27986

Relative Discharge Stage (RDS) Classification: A Practical Indicator of Battery Discharge Progress

相对放电阶段(RDS)分类:一种实用的电池放电进程指标
Tran, Khoa, Le, Tri, Trinh, Hung-Cuong, Tran-Nam, Hung
Abstract
Accurate remaining discharge time (RDT) prediction is challenging in real-world battery applications because future load profiles are unknown and highly dynamic. To address the uncertainty of continuous RDT regression, this paper introduces Relative Discharge Stage (RDS), a battery-management indicator that represents the remaining discharge condition using five interpretable classes: Normal, Good, Moderate, Low, and Recharge Required. Unlike state of charge (SOC), which reflects the current charge level, RDS characterizes the remaining discharge process without requiring future-current information during inference. A physics-informed RDS classification framework is proposed, combining SOC estimation with lightweight temporal learning. The SOC-estimation component includes second-order ECM state and terminal-voltage prediction, hysteresis and OCV temperature correction, core-temperature estimation, and AEKF state correction, supported by OCV evaluation, online STC-ECM parameter adaptation, and pretrained neural residual-voltage correction. The measured current, terminal voltage, surface temperature, and estimated SOC are arranged into a sliding observation window and processed by a lightweight temporal convolutional network. Experiments on two public lithium-ion battery datasets demonstrate robust RDS classification, with accuracy exceeding 80% under varying load and thermal conditions.
Chinese Translation
在真实电池应用中,由于未来负载曲线未知且高度动态,精确预测剩余放电时间(RDT)极具挑战性。为应对连续RDT回归的不确定性,本文提出相对放电阶段(Relative Discharge Stage, RDS)这一电池管理指标,利用五个可解释的类别来表征剩余放电状态:正常(Normal)、良好(Good)、中等(Moderate)、较低(Low)和需要充电(Recharge Required)。与反映当前电荷量的荷电状态(SOC)不同,RDS表征剩余放电过程,且在推断时无需未来电流信息。本文提出一种物理信息驱动的RDS分类框架,将SOC估计与轻量级时序学习相结合。SOC估计组件包括二阶等效电路模型(ECM)的状态与端电压预测、迟滞与开路电压(OCV)温度修正、核心温度估计以及自适应扩展卡尔曼滤波(AEKF)状态修正,并由OCV评估、在线STC-ECM参数自适应和预训练的神经残差电压修正提供支持。将测得的电流、端电压、表面温度和估计的SOC组织为滑动观测窗口,并由轻量级时序卷积网络进行处理。在两个公开锂离子电池数据集上的实验表明,该框架可实现稳健的RDS分类,在不同负载与热条件下准确率超过80%。
cs.LG / 91 / 2609.27987

PCQC: Privileged Counterfactual Question Credit for Multi-Turn Medical Dialogue

PCQC:面向多轮医疗对话的特权反事实问题信用分配方法
Li, Chenxuan, Wan, Jiayi, Chen, Xinrong, Zhao, Zhongyu, Shang, Xuecheng, Wan, Peixing
Abstract
Large language models (LLMs) have made substantial progress on medical question-answering, yet effective medical dialogue also requires learning to ask questions that uncover relevant patient information. To train such dialogue policies, a common pipeline combines supervised fine-tuning with reinforcement learning (RL) based on final diagnostic correctness. However, this outcome-based supervision does not directly distinguish the contributions of individual questions and provides no question-level feedback for unexecuted alternatives. To address this gap, we introduce PCQC (Privileged Counterfactual Question Credit), which uses privileged patient information during training to learn from questions never asked. During training, PCQC makes alternative questions directly comparable at the same dialogue state by using privileged patient facts to construct their answers. A frozen diagnostic scorer evaluates the diagnostic utility of each resulting question-answer pair by how strongly it supports the correct diagnosis. PCQC turns these comparisons into relative question credit that teaches the policy which questions to favor, directly supervising both executed and unexecuted questions alongside outcome-based RL without requiring complete rollouts for the unexecuted alternatives. Extensive experiments across four medical benchmarks demonstrate that PCQC achieves 63.10% mean diagnostic accuracy, outperforming GRPO and ATPO by 4.38 and 4.21 percentage points, respectively. These gains are achieved with 33.1% fewer inquiry turns than GRPO.
Chinese Translation
大语言模型(LLMs)在医疗问答任务上已取得显著进展,但有效的医疗对话还要求学会提出能够挖掘相关患者信息的提问。为训练此类对话策略,一种常见流程是将监督微调与基于最终诊断正确性的强化学习(RL)相结合。然而,这种基于结果的监督无法直接区分各个问题的贡献,也无法为未被执行的备选问题提供问题级别的反馈。为弥补这一不足,我们提出了 PCQC(Privileged Counterfactual Question Credit,特权反事实问题信用分配),该方法在训练过程中利用特权患者信息,从从未被提出的问题中学习。在训练期间,PCQC 通过利用特权患者事实来构建备选问题的答案,使备选问题能够在同一对话状态下直接进行比较。一个冻结的诊断评分器根据每个由此产生的问答对支持正确诊断的程度,评估其诊断效用。PCQC 将这些比较转化为相对问题信用,从而教会策略应优先选择哪些问题,在基于结果的强化学习之外,直接对已执行和未执行的问题进行监督,且无需为未执行的备选问题进行完整的对话推演。在四个医疗基准上的大量实验表明,PCQC 达到了 63.10% 的平均诊断准确率,分别超越 GRPO 和 ATPO 4.38 和 4.21 个百分点。此外,这些提升是在比 GRPO 少 33.1% 的问诊轮次下实现的。
cs.LG / 92 / 2609.28003

Learning from Failures: Heterogeneous Graph Memory for Small Language Model Tool-Using Agents

从失败中学习:面向小语言模型工具使用智能体的异构图记忆
Li, Jiaxing, Song, Lei, Dong, Rui, Kong, Youyong
Abstract
Small and medium-sized language models offer cost-effective executors for tool-using agents, making them attractive for local and large-scale deployment. However, in long-horizon and stateful environments, they often make structural errors such as missing required observations, performing premature writes, repeating failed calls, and violating action preconditions. These errors can lead to incorrect state updates, policy violations, and costly or irreversible consequences, making reliable tool execution a critical deployment challenge. Existing fine-tuning approaches require substantial data and computation, while flat memory may retrieve failed actions without preserving their causal context or safety conditions. In this paper, we propose FRESH, a Failure-aware Retrieval framework over Experience-Structured Heterogeneous graphs, which transforms historical successes and failures into structured external experience for tool-using agents. By explicitly modeling the dependencies among tasks, actions, errors, repairs, and execution conditions, FRESH helps frozen language models reuse reliable strategies, avoid recurring failures, and make safer decisions in stateful tool interactions. Experiments on $\tau$-Bench and AppWorld with multiple open-source models show that FRESH consistently improves task success and tool-use reliability over no-memory agents and representative memory-based baselines.
Chinese Translation
小型和中型语言模型为工具使用智能体(tool-using agents)提供了高性价比的执行器,使其在本地和大规模部署中颇具吸引力。然而,在长时程且有状态的环境中,这些模型常常出现结构性错误,例如遗漏必要的观察结果、过早执行写入操作、重复失败的调用以及违反动作前置条件。这些错误可能导致错误的状态更新、违反策略约束以及代价高昂或不可逆的后果,使得可靠的工具执行成为一项关键的部署挑战。现有的微调方法需要大量的数据和计算资源,而扁平化记忆(flat memory)在检索失败动作时可能无法保留其因果上下文或安全条件。本文提出了FRESH,一种基于经验结构化异构图(Experience-Structured Heterogeneous graphs)的失败感知检索框架,它将历史上的成功与失败经验转化为工具使用智能体的结构化外部经验。通过显式建模任务、动作、错误、修复及执行条件之间的依赖关系,FRESH帮助冻结的语言模型复用可靠策略、避免重复性失败,并在有状态的工具交互中做出更安全的决策。在$ au$-Bench和AppWorld上使用多个开源模型的实验表明,与无记忆智能体及具有代表性的基于记忆的基线方法相比,FRESH持续提升了任务成功率和工具使用的可靠性。
cs.LG / 93 / 2609.28006

Shared Global KV with Layer-Specific Local History

共享全局KV与层特定局部历史
Xian, Xinglang
Abstract
Decoder-only Transformer language models cache keys and values (KV) to reuse past computation during generation. Sharing KV across layers saves storage but reduces the diversity of representations available across depth. We study what local memory should retain alongside shared global KV, separating historical content from the input source used to form it. At 126M parameters and 2K context, an eight-seed study finds about 1.4% lower held-out test perplexity with local history than with a current-token local branch. Capacity, entry-count and training-compute controls support the value of historical content. In a two-seed comparison, this value persists when adjacent layers share local inputs while retaining independent projections; source sharing also shortens exact cache-construction dependencies. Against GQA and adjacent-layer KV sharing, equal bounded learning-rate searches and new-seed confirmation yield better same-source likelihood with larger caches and higher long-request latency. The ordering against adjacent-layer sharing persists after equal-token adaptation to 8K, with a short-context cost. The eight-seed external-book history effect remains uncertain, and downstream outcomes vary by task. We derive a sufficient suffix schedule that reduces upper-layer construction work while preserving the complete cache in exact arithmetic.
Chinese Translation
仅解码器(Decoder-only)的Transformer语言模型在生成过程中缓存键值对(KV)以复用过去的计算。跨层共享KV可以节省存储空间,但会降低深层中可用表示的多样性。我们研究了在共享全局KV之外,局部记忆应当保留什么内容,并将历史内容与其形成时所使用的输入来源区分开来。在126M参数和2K上下文长度的设置下,一项包含八个随机种子的实验发现,使用局部历史相比使用当前token的局部分支,可将留出测试集困惑度(perplexity)降低约1.4%。关于容量、条目数量和训练计算量的对照实验支持了历史内容的价值。在一项双种子对比实验中,当相邻层共享局部输入但保留各自独立的投影时,这种价值依然存在;此外,共享输入来源还能缩短精确缓存构建的依赖链。与GQA(分组查询注意力)和相邻层KV共享方法相比,在等界学习率搜索和新种子验证下,本方法在更大缓存规模和更长请求延迟的条件下取得了更优的同源似然。在将token数量对齐并扩展到8K上下文后,相对相邻层共享的优势依然保持,但存在短上下文场景的性能代价。八个种子下的外部书籍历史效应仍不确定,且下游任务的表现因任务而异。我们推导了一种充分的后缀调度方法,可在精确算术意义下保持完整缓存的同时,减少上层的构建工作量。
cs.LG / 94 / 2609.28022

PISCES: Physics-Informed Solar-wind Convolutional autoEncoder for Space-weather Anomaly Detection and Early Warning

PISCES:用于空间天气异常检测与预警的物理信息太阳风卷积自编码器
Lee, Kevin, March, Alison J.
Abstract
Space weather early warning depends on detecting solar wind transients in in-situ measurements at the first Sun-Earth Lagrange point (L1), before they reach Earth. Fixed thresholds can miss combined magnetic and plasma structure, and many learning methods provide a single anomaly score. We present the Physics-Informed Solar-wind Convolutional autoEncoder for Space-weather (PISCES), a convolutional autoencoder trained without catalog labels on OMNI solar wind measurements under physics constraints. Its loss includes magnetic field consistency, an empirical relation between temperature and velocity, the Parker spiral angle, and penalties on changes between consecutive one-minute samples in derived quantities calculated from the reconstruction. At inference, PISCES separates the anomaly score into magnetic and plasma reconstruction errors, physics relations, and residual corrections, and reports the magnitude of each contribution. Attenuation of the skip connections, selected on validation data, improves average precision for the trained models, while the untrained scores remain nearly the same. The trained models also give a more consistent ordering of these physical contributions. After smoothing with a trailing median, the alarms can precede independently observed sudden commencements, including positive sudden impulses.
Chinese Translation
空间天气预警依赖于在太阳风瞬变结构到达地球之前,在日地第一拉格朗日点(L1)的原位测量中探测这些结构。固定阈值方法可能遗漏磁场与等离子体组合结构的异常,而许多学习方法只提供单一的异常分数。我们提出了用于空间天气的物理信息太阳风卷积自编码器(PISCES),这是一个在物理约束下、无需目录标签、基于OMNI太阳风测量数据训练的卷积自编码器。其损失函数包括磁场一致性、温度与速度之间的经验关系、帕克螺旋角,以及对由重构计算的派生量在相邻一分钟样本之间变化的惩罚项。在推理阶段,PISCES将异常分数分解为磁场重构误差、等离子体重构误差、物理关系误差和残差修正,并报告各项贡献的大小。在验证数据上选取的跳跃连接衰减可提高训练后模型的平均精度,而未训练模型的分数几乎保持不变。训练后的模型对这些物理贡献的排序也更加一致。经过尾随中位数平滑处理后,警报可以独立观测到的磁扰急始(包括正急始脉冲)之前发出。
cs.LG / 95 / 2609.28053

Exact Quantile Balancing and Load-Error Injection for Mixture-of-Experts

面向混合专家模型的精确分位均衡与负载误差注入
Neitemeier, Pit, Li, Jiaze, Serra, Alessio, Scholl, Philipp, Maskey, Sohir
Abstract
Mixture-of-Experts (MoE) training requires global load balance to prevent expert under-utilization and local balance for efficient expert-parallel execution. Existing distributed Quantile Balancing (QB) uses shard-dependent or approximate global quantiles, while token-independent expert biases cannot ensure microbatch-level balance. We introduce Exact Quantile Balancing (EQB), which computes exact global-batch BF16 quantiles with negligible communication, and Load-Error Injection (LEI), which injects local load errors directly into router-score gradients. On 7.5B-parameter MoEs trained for up to 500B tokens, EQB improves global balance and downstream performance over naive QB, while LEI improves local balance and outperforms the GShard loss at comparable quality.
Chinese Translation
混合专家模型(Mixture-of-Experts, MoE)的训练需要全局负载均衡以防止专家利用不足,同时需要局部均衡以实现高效的专家并行执行。现有的分布式分位均衡(Quantile Balancing, QB)方法依赖于分片相关的分位数或近似的全局分位数,而与词元无关的专家偏置无法确保微批次(microbatch)层面的均衡。我们提出了精确分位均衡(Exact Quantile Balancing, EQB),它能以极小的通信开销计算出精确的全局批次 BF16 分位数;并提出了负载误差注入(Load-Error Injection, LEI),将局部负载误差直接注入路由器得分的梯度中。在训练词元数高达 5000 亿的 7.5B 参数 MoE 模型上,EQB 相比朴素的 QB 方法提升了全局均衡性和下游性能,而 LEI 改善了局部均衡性,并在相近模型质量的条件下优于 GShard 损失函数。
cs.LG / 96 / 2609.28085

Curriculum Learning with GNN-based Reinforcement Learning for Job Shop Scheduling

基于图神经网络强化学习与课程学习的作业车间调度研究
Vasudevan, Jayakrishnan K., Hoss, Jonathan, Klarmann, Noah
Abstract
The job shop scheduling problem is a challenging combinatorial optimization problem, and recent reinforcement learning approaches using graph neural networks have shown promise for learning scheduling policies directly from problem instances. However, training on large instances remains computationally expensive, and generalization across instance sizes remains challenging. This paper studies curriculum learning for graph neural network-based reinforcement learning in the job shop scheduling problem by comparing it with single-size training across three target sizes: 20 x 20, 25 x 25, and 30 x 30. In the curriculum setting, the policy is first trained on smaller instances and then progressively adapted to larger target sizes, allowing scheduling behavior learned in earlier stages to support learning on larger instances. Models are evaluated on unseen instances from 8 x 8 to 30 x 30 using the optimality gap, considering both generalization across all evaluation sizes and specialization on the target size. Results show that curriculum learning consistently reduces wall-clock training time, with larger benefits as the target size increases. The strongest advantage is observed at 30 x 30, where curriculum learning reduces the mean optimality gap across all evaluation sizes by approximately 8.1 percentage points, reduces the target-size mean optimality gap by approximately 8.6 percentage points, and saves approximately 50 hours of training time.
Chinese Translation
作业车间调度问题(Job Shop Scheduling Problem)是一个具有挑战性的组合优化问题,近年来基于图神经网络(GNN)的强化学习方法显示出直接从问题实例中学习调度策略的潜力。然而,在大规模实例上的训练计算开销高昂,且跨实例规模的泛化仍然具有挑战性。本文研究了基于图神经网络的强化学习在作业车间调度问题中的课程学习方法,并在三个目标规模(20×20、25×25 和 30×30)上将其与单一规模训练进行比较。在课程学习设置中,策略首先在较小实例上训练,然后逐步适应更大的目标规模,使早期阶段学到的调度行为能够支持在更大实例上的学习。模型在从 8×8 到 30×30 的未见实例上进行评估,使用最优性间隙(optimality gap)作为指标,同时考虑跨所有评估规模的泛化能力和对目标规模的专门化能力。结果表明,课程学习能够持续减少实际训练时间(wall-clock time),且目标规模越大,收益越显著。最显著的优势出现在 30×30 规模上:课程学习使所有评估规模上的平均最优性间隙降低约 8.1 个百分点,使目标规模的平均最优性间隙降低约 8.6 个百分点,并节省约 50 小时的训练时间。
cs.LG / 97 / 2609.28086

LAYERSCOPE: A Layerwise Characterization of Video and Multimodal Learned Representations

LAYERSCOPE:视频与多模态学习表征的逐层刻画
Arcos-Holzinger, Sandra, Chakraborty, Debashish, Mocharla, Rohita, Walden, Will, Yates, Andrew, Kriz, Reno, Erfani, Sarah M., Bailey, James, Patel, Vishal M., Khudanpur, Sanjeev
Abstract
We propose LAYERSCOPE, a label-free, layerwise framework that aims to characterize a model's learned representations in video and multimodal settings. Evaluating downstream performance using representations from final or intermediate layers typically requires large amounts of labeled data, repeated task-specific evaluations, and substantial computation. To address these limitations, LAYERSCOPE uses local, global, distributional, and correspondence-based geometric metrics to compare layerwise representation structure within and across models without requiring task-specific labels. We evaluate seven architecturally diverse models across video and multimodal classification, clustering, and text-to-video retrieval tasks from MVEB/MVEB+. We find that intermediate-layer representations can outperform final-layer and model-default outputs. We also find that no single geometric metric consistently predicts downstream performance, but note that distinct layerwise geometric signatures emerge across model families. LID shows task-dependent relationships with performance, while RankMe provides the strongest measure for classification and clustering, but is not a universal layer selector. We also find that pairing-aware metrics explain retrieval better than distributional distances alone. LAYERSCOPE therefore offers a framework for comparing representations across models and layers, enabling a more systematic evaluation in video and multimodal settings.
Chinese Translation
我们提出了LAYERSCOPE,一个无需标签的逐层分析框架,旨在刻画模型在视频和多模态场景下学习到的表征。使用最终层或中间层的表征来评估下游性能通常需要大量带标签数据、重复的任务特定评估以及大量计算。为解决这些局限,LAYERSCOPE利用局部、全局、分布和基于对应关系的几何度量,在无需任务特定标签的情况下比较模型内部及跨模型的逐层表征结构。我们在MVEB/MVEB+的视频与多模态分类、聚类和文本到视频检索任务上评估了七个架构各异的模型。我们发现,中间层表征可以优于最终层和模型默认输出。我们还发现,没有任何单一几何度量能够一致地预测下游性能,但注意到不同模型家族呈现出独特的逐层几何特征。LID与性能之间表现出任务依赖的关系,而RankMe在分类和聚类任务中提供了最强的度量,但并非通用的层选择器。我们还发现,配对感知度量比单纯的分布距离更能解释检索性能。因此,LAYERSCOPE提供了一个跨模型、跨层比较表征的框架,使视频和多模态场景下的评估更加系统化。
cs.LG / 98 / 2609.28105

Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness

Fed-ReMasker:特征级缺失下的联邦表格数据插补
Papathanail, Ioannis, Poursoleymani, Rooholla, Rahman, Lubnaa Abdur, Mougiakakou, Stavroula Georgia
Abstract
Multi-center clinical studies and biomedical research collaborations increasingly seek to utilize data across centers to build models that generalize beyond any single center. This creates two distinct challenges: data protection regulations may restrict the sharing of raw patient data across institutions, while centers may collect only partially overlapping sets of features under different protocols. Federated learning enables collaborative model training without centralizing raw data. However, existing federated imputation methods rarely evaluate feature-level missingness, in which entire features are unobserved at some centers. To address this setting, we adapt the ReMasker masked autoencoder to federated learning (Fed-ReMasker), enabling centers to impute features never observed locally by leveraging knowledge learned across collaborating centers. We evaluate Fed-ReMasker in a benchmark spanning synthetic datasets with linear and nonlinear relationships and real-world tabular datasets, including clinical data. The benchmark varies the number of centers, the missingness ratios, and client heterogeneity. Fed-ReMasker achieves the lowest imputation error in 93.2% of value-level and 96.7% of feature-level scenarios in the homogeneous benchmark. It also remains robust to client heterogeneity using simple federated averaging, outperforming all baselines in all 36 value-level scenarios and each baseline in at least 35 of 36 feature-level scenarios, and comes within 3.0% on average of a centralized model trained on the pooled data.
Chinese Translation
多中心临床研究与生物医学研究合作日益寻求利用跨中心数据构建能超越单一中心泛化能力的模型。这带来了两个截然不同的挑战:数据保护法规可能限制跨机构共享原始患者数据,同时各中心在不同协议下可能仅收集部分重叠的特征集。联邦学习使协同模型训练无需集中原始数据成为可能。然而,现有的联邦插补方法很少评估特征级缺失,即某些中心完全未观测到某些特征。为应对这一场景,我们将ReMasker掩码自编码器适配到联邦学习中(Fed-ReMasker),使各中心能够利用协作中心学到的知识,对本地从未观测到的特征进行插补。我们在一个基准测试中评估了Fed-ReMasker,该基准涵盖具有线性和非线性关系的合成数据集以及真实世界的表格数据集(包括临床数据)。基准测试改变了中心数量、缺失比例和客户端异质性。在同质基准测试中,Fed-ReMasker在93.2%的值级场景和96.7%的特征级场景中取得了最低的插补误差。在客户端异质性场景下,仅使用简单的联邦平均,Fed-ReMasker依然保持稳健,在全部36个值级场景中优于所有基线方法,并且在36个特征级场景中至少在35个场景中优于每个基线方法,其性能与基于汇总数据训练的集中式模型平均相差不超过3.0%。
cs.LG / 99 / 2609.28116

Probabilistic and Geometry Aware Neural Surrogate of Scrape Off Layer Plasma Simulations

面向刮削层等离子体模拟的概率性与几何感知神经代理模型
Gianuzzo, Gabriele, Dasbach, Stefan, Hendriks, Fleur, Wiesen, Sven, Menkovski, Vlado
Abstract
Fast surrogates for tokamak boundary-plasma simulation are typically deterministic regressors mapping a global operating point to a flattened vector of cell values. Near the divertor detachment transition the steady state is not reliably single-valued. A point estimate must average over qualitatively different plasma states, and it arrives with no statement of confidence. Moreover, the flattened vector representation discards the geometric structure of the SOLPS-ITER mesh. This work addresses both problems. We unroll the curvilinear mesh into three fixed-size image tensors whose layout preserves cell adjacency and inverts exactly, letting a convolutional network act on the geometry without loss of information. A conditional flow matching model, well suited to highly sensitive systems, is then trained on this representation. The result is an efficient, scalable surrogate that captures multiple plausible outcomes even at sensitive operating points. Along a gas-puff scan, the predictive distribution splits into a hot and a cold mode across an early regime transition. A further check on synthetic data with an injected bifurcation of known size confirms the model recovers both branches rather than their average.
Chinese Translation
托卡马克边界等离子体模拟的快速代理模型通常是确定性的回归器,将全局运行点映射为展平的网格单元值向量。在偏滤器脱靶转变附近,稳态并不可靠地保持单值。点估计必须对不同性质的等离子体状态取平均,且不提供任何置信度信息。此外,展平的向量表示丢失了 SOLPS-ITER 网格的几何结构。本工作同时解决了这两个问题。我们将曲线网格展开为三个固定尺寸的图像张量,其布局保持网格单元的邻接关系并可精确逆变换,使卷积网络能够在不损失信息的情况下利用几何结构。随后在该表示上训练了一个条件流匹配(conditional flow matching)模型,该模型非常适合高度敏感的系统。由此得到一个高效、可扩展的代理模型,即使在敏感的运行点也能捕捉多种合理的结果。在气体喷注扫描过程中,预测分布在一个早期状态转变处分裂为热模和冷模。在注入已知大小分岔的合成数据上的进一步验证表明,该模型能够恢复两个分支,而非它们的平均值。
cs.LG / 100 / 2609.28145

RL Starts before RL: On Policy Distillation for Better Reinforcement Learning

强化学习始于RL之前:论策略蒸馏对强化学习的促进作用
Dong, Shuai, Zhu, Yongfu, Xu, Yuqi, Xie, Weichu, Liuwenpu, Wang, Ziyue, Tuo, Kaiwen, Wang, Congcong, Wang, Siyuan, Shao, Wenqi, Yang, Shuai, Zhao, Ji, Ma, Caoyuan, Chang, Wenzheng, Wu, Taiqiang, Yu, Xinlei, Wu, Hongrui, He, Xiaoxuan, Chen, Fangke, Wang, Dianyi, Tian, Kanghui, Chen, Sirry, Liu, Xingyu, Wu, Xiangnan, Guo, Jiawei, Hou, Haowen, Chen, LingHan, Wei, Zhongyu, Wang, Jiaqi
Abstract
Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond improvements in the distilled model's initial accuracy. Under shared RL settings, students initialized with OPD reach higher final performance than those trained with direct RL or supervised fine-tuning followed by RL. This advantage can emerge even when OPD produces little immediate improvement in accuracy. Pre-RL Pass@k does not fully explain the benefit: similar or even higher values do not necessarily lead to better performance after RL. Behavioral analyses point to alignment with the teacher's distribution beyond top-1 agreement as a possible explanation. Such alignment may favor higher-quality reasoning paths while retaining alternatives that RL can further refine using outcome feedback. We further examine how trajectory sources and divergence objectives affect the value of distillation for subsequent RL. Standard reverse-KL OPD performs better before RL, but forward-KL OPD overtakes it afterward; with teacher-generated distillation trajectories, reverse KL remains ahead at both stages. These findings suggest that the preferred distillation objective depends on both the trajectory source and the training that follows. Our results support evaluating OPD as preparation for RL and selecting distillation choices by the performance achieved after subsequent training.
Chinese Translation
强化学习(RL)能够提升推理能力,但其性能取决于训练开始时所基于的策略。我们研究了在线策略蒸馏(on-policy distillation, OPD)作为RL准备阶段的作用,并探讨其收益是否超越了蒸馏模型初始准确率的提升。在相同的RL设置下,以OPD初始化的学生模型比直接进行RL或先监督微调再RL训练的模型达到更高的最终性能。即使OPD在准确率上几乎没有即时改进,这一优势也可能出现。RL前的Pass@k并不能完全解释这一收益:相似甚至更高的Pass@k值并不一定带来RL后更好的性能。行为分析表明,与教师分布在top-1一致之外的对齐可能是一种解释。这种对齐可能有利于更高质量的推理路径,同时保留RL能够利用结果反馈进一步改进的备选路径。我们进一步考察了轨迹来源和散度目标如何影响蒸馏对后续RL的价值。标准的反向KL(reverse-KL)OPD在RL前表现更好,但前向KL(forward-KL)OPD在RL后实现反超;而使用教师生成的蒸馏轨迹时,反向KL在两个阶段均保持领先。这些发现表明,最优的蒸馏目标取决于轨迹来源和后续的训练方式。我们的结果支持将OPD作为RL的准备阶段来评估,并根据后续训练所达到的性能来选择蒸馏方案。
cs.LG / 101 / 2609.28165

Confidence Falls Short: Asymmetric Certainty Gains from Optimization Hinder Multimodal Classification

置信度的落差:优化带来的非对称确定性增益阻碍多模态分类
Huang, Longfei, Wu, Xiangyu, Yang, Yang
Abstract
Multimodal learning (MML) falls into the optimization dilemma due to the modality imbalance phenomenon, leading to suboptimal overall performance in practice. While many attempts primarily focus on balancing the optimization dynamics across modalities to address this issue, we identify a subtle yet critical flaw: optimization yields asymmetric gains in predictive certainty, with the strong modality more confident than the weak one, driving imbalanced modality contributions. In this paper, our analysis reveals that this flaw stems from unimodal characteristics rather than multimodal learning, and this confidence discrepancy can be corrected by positive cross-modal intervention. Based on this insight, we propose multimodal Max Confidence Regularization (MaxCR) to dynamically intervene in modality semantic confidence. Specifically, the semantic confidence of each modality is tracked using a nonlinear sparsity measure. We then design max suppression and max excitation based on this measure to regularize strong and weak modalities, respectively. They penalize and encourage the top-1 confidence, thereby constraining multimodal prediction. To this end, strong and weak modalities are expected to make calibrated confidence, thereby improving the overall performance. Empirical experiments on widely used datasets reveal the superiority of our method through comparison with various state-of-the-art (SOTA) multimodal learning baselines.
Chinese Translation
多模态学习(Multimodal Learning, MML)由于模态不平衡现象而陷入优化困境,导致实践中整体性能欠佳。尽管许多研究主要通过平衡各模态之间的优化动态来解决这一问题,但我们发现了一个微妙却关键的缺陷:优化在预测确定性上产生非对称的增益,强模态比弱模态更加自信,从而驱动模态贡献失衡。在本文中,我们的分析表明,这一缺陷源于单模态本身的特性而非多模态学习,并且这种置信度差异可以通过正向的跨模态干预加以纠正。基于这一洞察,我们提出多模态最大置信度正则化(Max Confidence Regularization, MaxCR),以动态干预模态的语义置信度。具体而言,我们使用一种非线性稀疏性度量来追踪每个模态的语义置信度;随后,基于该度量设计最大抑制(max suppression)与最大激励(max excitation),分别对强模态和弱模态进行正则化。它们通过惩罚和鼓励top-1置信度,从而约束多模态预测。由此,强模态和弱模态有望给出经过校准的置信度,进而提升整体性能。在广泛使用的数据集上的实证实验表明,与多种最先进(SOTA)的多模态学习基线相比,我们的方法具有优越性。
cs.LG / 102 / 2609.28194

Geospatial embeddings detect old-growth forests but buffered spatial validation narrows their advantage over Sentinel features

地理空间嵌入特征可检测原始老林,但缓冲空间验证缩小了其相对于Sentinel特征的优势
Ratsakatika, Thomas, Zotta, Mihai, Keshav, Srinivasan, Lines, Emily R.
Abstract
Old-growth forests develop over centuries under minimal anthropogenic disturbance, producing structurally complex and biodiverse stands. In Europe, protecting them requires mapping that is accurate for individual forest parcels yet deployable continent-wide. Geospatial foundation model (GFM) embeddings enable label-scarce land classification, but their value for old-growth detection remains unknown. Here, we map old-growth forests across 211,893 ha of Romania's Southern Carpathians, a beech-spruce landscape typical of the Alpine Biogeographic Region. We construct high-confidence, expert-informed reference labels for old-growth and non-old-growth parcels. We add AlphaEarth, TESSERA v2 and Sentinel-1/2 features to a common baseline of topographic and human-access predictors, then compare them under spatially blocked validation with and without 10 km train-test buffers to limit residual autocorrelation. With buffering, GFM and Sentinel-1/2 predictors increase precision-recall AUC by 0.21-0.25 [95% CIs: 0.15-0.34] relative to baseline, indicating spectral data contain a spatially robust old-growth signal. With a PR-AUC of 0.84 [0.79-0.88], TESSERA outperforms Sentinel-1/2 (+0.08 [+0.05 to +0.11]) and AlphaEarth (+0.08 [+0.04 to +0.12]) under unbuffered spatial validation. At a 10 km buffer, however, this advantage narrows to +0.04 [-0.01 to +0.11] and +0.03 [-0.04 to +0.10], intervals consistent with no difference. At 10 m resolution, convolutional neural networks add no benefit over pixel-based XGBoost. Comparisons with four national- and continental-scale products show the importance of non-old-growth labels, and reveal 81% agreement between our predictions and a field-calibrated map. We conclude that buffered spatial validation is vital when transferring old-growth detection models to unseen landscapes, and provide our labels and predictions for future work.
Chinese Translation
原始老林(old-growth forests)是在极少人为干扰下历经数百年发育而成,形成结构复杂且生物多样性丰富的林分。在欧洲,保护原始老林需要既能对单个林地块进行精确制图、又可在整个大陆尺度上部署的地图。地理空间基础模型(GFM)嵌入特征可实现标签稀缺条件下的土地分类,但其在原始老林检测中的价值仍属未知。本研究对罗马尼亚南喀尔巴阡山脉211,893公顷的区域进行了原始老林制图,该区域是典型的高山生物地理区的山毛榉-云杉景观。我们构建了高置信度、由专家参与的原生老林与非老林地块参考标签。我们在由地形和人为可达性预测因子组成的共同基线中加入AlphaEarth、TESSERA v2和Sentinel-1/2特征,然后在空间分块验证下进行比较,分别设置和不设置10公里训练-测试缓冲区以限制残余自相关。采用缓冲区后,GFM和Sentinel-1/2预测因子相对于基线将精确率-召回率曲线下面积(PR-AUC)提高了0.21-0.25 [95%置信区间:0.15-0.34],表明光谱数据包含空间上稳健的原始老林信号。在无缓冲空间验证下,TESSERA的PR-AUC达0.84 [0.79-0.88],优于Sentinel-1/2(+0.08 [+0.05至+0.11])和AlphaEarth(+0.08 [+0.04至+0.12])。然而,在10公里缓冲区下,该优势缩小至+0.04 [-0.01至+0.11]和+0.03 [-0.04至+0.10],其置信区间与无差异一致。在10米分辨率下,卷积神经网络相比基于像素的XGBoost没有带来额外收益。与四个国家和大陆尺度产品的比较表明了非老林标签的重要性,并显示我们的预测与一张经过实地校准的地图之间有81%的一致性。我们得出结论:在将原始老林检测模型迁移到未见过的景观时,缓冲空间验证至关重要,并提供我们的标签和预测结果以供未来研究使用。
cs.LG / 103 / 2609.28199

Transferable Evidence Reconstruction for Longitudinal Glucose Representations

面向纵向血糖表征的可迁移证据重建
Zhou, Tian, Peng, Bingqing, Yang, Linxiao, Wang, Wenwei, Ye, Mengni, Jin, Beverly, Zhu, Zuyi, Gu, Jinjie, Sun, Liang
Abstract
Long physiological recordings contain many routine measurements, while predictive information is often concentrated in rare events, sustained burden, and recurring temporal patterns. Masked autoencoding recovers measurements; contrastive learning aligns views. We study self-supervision that explicitly prioritizes structured signal evidence. We introduce transferable evidence reconstruction (TER), which constructs evidence from unlabeled recordings, fits a fresh low-capacity reader on one recording group, and requires that reader to recover the same evidence in another group without refitting. Differentiating through this cross-group test learns representations with transferable evidence-decoding rules; the evidence guides self-supervision but is not used as a downstream feature. For continuous glucose monitoring (CGM), an observation-aware daily encoder and clock-aware multi-day memory bind glucose level and change to recorded time while organizing up to seven days of history. On the 14-task leaderboard, TER improves the strongest prior overall PR-AUC/ROC-AUC/Macro-F1 scores by 5.51/4.43/2.80 percentage points and sets a new best metric on 12/14 tasks. These leaderboard gains are 2.0-2.9 times the respective gaps between the two strongest baselines. With public pretraining data, folds, and the linear probe matched, TER outperforms our GlucoFM reproduction by 6.09/5.52/2.72 points. Target-reader ablations, same-history controls, and cross-person readouts support the combination of structured evidence, cross-group reader fitting, and learned multi-day organization.
Chinese Translation
长时间的生理记录包含大量常规测量,而预测性信息往往集中在罕见事件、持续负荷以及反复出现的时间模式中。掩码自编码方法用于恢复测量值;对比学习用于对齐不同视图。我们研究一类显式优先考虑结构化信号证据的自监督方法。我们提出可迁移证据重建(Transferable Evidence Reconstruction, TER),该方法从无标签记录中构建证据,在某一记录组上拟合一个全新的低容量读取器,并要求该读取器在不重新拟合的情况下从另一组记录中恢复相同的证据。通过这一跨组测试进行反向传播,可学习到具有可迁移证据解码规则的表征;证据仅用于引导自监督,而不作为下游特征使用。针对持续血糖监测(CGM),我们设计了一个感知观测的日编码器和感知时钟的多日记忆模块,将血糖水平及其变化与记录时间相关联,同时组织长达七天的历史数据。在包含14个任务的排行榜上,TER 将此前最强方法的整体 PR-AUC/ROC-AUC/Macro-F1 分别提升 5.51/4.43/2.80 个百分点,并在 12/14 个任务上创下新的最佳指标。这些排行榜增益是两个最强基线之间相应差距的 2.0-2.9 倍。在公开预训练数据、交叉验证折数以及线性探针设置完全一致的情况下,TER 比我们对 GlucoFM 的复现结果高出 6.09/5.52/2.72 个百分点。目标读取器消融实验、同历史对照实验以及跨个体读出实验支持结构化证据、跨组读取器拟合与学习式多日组织的组合有效性。
cs.LG / 104 / 2609.28208

Support-Compiled Feature Folding: More Evidence at Lower Memory Across Tabular Foundation Models

支持编译式特征折叠:在表格基础模型上以更低内存获得更多证据
Zhou, Tian, Jin, Beverly, Wang, Xue, Yang, Linxiao, Wang, Wenwei, Peng, Bingqing, Ye, Mengni, Gu, Jinjie, Sun, Liang
Abstract
Tabular foundation models face a feature-side scaling dilemma: full-width pairwise mixing grows quadratically with the number of columns, whereas feature selection saves memory by discarding evidence. We introduce Support-Compiled Feature Folding (SCFF), a training-free inference framework that resolves this dilemma without changing the frozen backbone. SCFF routes support-ranked features through bounded leaves of the native feature encoder, support-checks the residual evidence, and merges the encoded messages before a single contextual prediction. It thereby converts quadratic feature-interaction work into linear-in-width work with a bounded local working set, without ensembling predictions or training new parameters. On the exhaustive 18-dataset wide-table slice of fixed AMLB-29, TabZilla, and TabArena snapshots, SCFF improves dataset-macro accuracy and NLL on all six evaluated backbones. All four matched-width comparisons retain favorable 95 percent dataset-bootstrap intervals on locked folds, with relative error reductions up to 26.1 percent. Median paired GPU-memory savings are 2.09x to 2.36x, and the ratio of separately observed maximum peaks reaches 34.3x. Under a measured peak-memory ceiling, SCFF uses the saved budget to preserve more support-selected evidence, improving accuracy by 4.06 and 3.72 points over the widest feasible single leaf on predeclared wide-Core strata of TabICLv2 and TabPFN-3.
Chinese Translation
表格基础模型面临特征侧的扩展困境:全宽度的两两特征混合随列数呈二次方增长,而特征选择虽节省内存却丢弃了证据。我们提出支持编译式特征折叠(Support-Compiled Feature Folding, SCFF),这是一种免训练的推理框架,在不改变冻结骨干网络的前提下解决了这一困境。SCFF 将按支持度排序的特征路由至原生特征编码器的有界叶子节点,对残差证据进行支持度检查,并在单次上下文预测之前合并编码后的信息。由此,它将二次方级的特征交互计算转化为宽度线性、局部工作集有界的计算,且无需集成预测或训练新参数。在固定 AMLB-29、TabZilla 和 TabArena 快照的穷举 18 数据集宽表切片上,SCFF 在所有六个被评估的骨干网络上均提升了数据集宏平均准确率和 NLL。在锁定折上,所有四个等宽度比较均保持有利的 95% 数据集自助置信区间,相对误差降低最高达 26.1%。配对 GPU 内存节省的中位数为 2.09 倍至 2.36 倍,单独观测到的最大峰值之比达 34.3 倍。在实测峰值内存上限约束下,SCFF 利用节省的预算保留更多经支持度选择的证据,在预先声明的 TabICLv2 和 TabPFN-3 宽-Core 分层上,相比最宽可行单叶子配置分别将准确率提升 4.06 和 3.72 个百分点。
cs.LG / 105 / 2609.28212

Log-Depth Recurrent Language Modeling

对数深度循环语言建模
Wang, Yiqin, Cingillioglu, Nuri, Pert, Charles
Abstract
Language modeling using Transformers has become commonplace despite their fixed computational depth and quadratic runtime with respect to input tokens. Recurrent models on the other hand offer linear depth but no parallel execution. In this work, we extend balanced-tree recursive operators from sequence encoding to autoregressive prediction, enabling all prefix representations to be computed with logarithmic depth and linear runtime. Our experiments provide an initial characterization of this model class, demonstrating robust length extrapolation and performance approaching that of ALiBi-based Transformers, highlighting its potential as an alternative architecture for language modeling.
Chinese Translation
尽管 Transformer 具有固定的计算深度以及相对于输入 token 的二次方运行时间,其语言建模应用已十分普遍。而循环模型虽提供线性深度,却无法并行执行。在本工作中,我们将平衡树递归算子从序列编码扩展到自回归预测,使得所有前缀表示能够以对数深度和线性运行时间计算。我们的实验对这一模型类别进行了初步刻画,展示了其稳健的长度外推能力,且性能接近基于 ALiBi 的 Transformer,凸显了其作为语言建模替代架构的潜力。
cs.LG / 106 / 2609.28248

hyperbolix: Hyperbolic Deep Learning in JAX

hyperbolix:基于JAX的双曲深度学习库
Klein, Timo, Lang, Thomas, Velaj, Yllka, Tschiatschek, Sebastian
Abstract
We present hyperbolix, an open-source library for hyperbolic deep learning in JAX, built on Flax NNX. To our knowledge, it is the first comprehensive, general-purpose hyperbolic deep learning library in JAX. It includes six manifolds with a common interface: Euclidean space, the Poincar\'e ball, the hyperboloid, the $\kappa$-stereographic model, mixed-curvature product spaces, and the proper velocity space. We implement layer families that cover linear layers, convolutions, attention, normalization, positional encoding, regression, and vector quantization. These building blocks span methods ranging from Ganea's original hyperbolic neural networks to recent fully hyperbolic architectures such as Hypformer and Lorentzian ResNet. Additionally, hyperbolix contains Riemannian optimizers implemented as optax transformations, wrapped distributions, and hyperbolic dimensionality-reduction techniques. Its API uses idiomatic JAX: Manifolds are stateless, with curvature being passed at call time, while manifold operations act on single points, with jax.vmap enabling batch operations. The precision of every checked operation is tested against a closed-form NumPy/SciPy transcription from the source paper or a finite difference, for both float32 and float64. On the hyperboloid, standard formulas for two-point operations, such as the distance, lose precision far from the origin, because they subtract two large, nearly equal terms. hyperbolix replaces these subtractions with cancellation-free formulas that stay accurate in float32 at distances where prior implementations return NaN. hyperbolix is available under the MIT license at https://github.com/timoklein/hyperbolix .
Chinese Translation
我们提出了 hyperbolix,一个基于 Flax NNX 构建的、用于 JAX 中双曲深度学习的开源库。据我们所知,这是 JAX 中首个全面且通用的双曲深度学习库。该库包含六种具有统一接口的流形:欧几里得空间、庞加莱球(Poincaré ball)、双曲面(hyperboloid)、$\kappa$-立体投影模型、混合曲率乘积空间以及本征速度空间(proper velocity space)。我们实现了涵盖线性层、卷积、注意力机制、归一化、位置编码、回归和向量量化的层族(layer families)。这些基础组件覆盖了从 Ganea 最初提出的双曲神经网络到近期完全双曲的架构(如 Hypformer 和洛伦兹 ResNet)等多种方法。此外,hyperbolix 还包含以 optax 变换形式实现的黎曼优化器、封装的分布(wrapped distributions)以及双曲降维技术。其 API 遵循 JAX 的惯用风格:流形是无状态的,曲率在调用时传入;流形操作作用于单个点,通过 jax.vmap 实现批量操作。所有受检操作的精度均针对原始论文中的闭式解(NumPy/SciPy 转写)或有限差分,在 float32 和 float64 两种精度下进行了测试。在双曲面上,两点运算(如距离)的标准公式在远离原点处会损失精度,因为它们对两个大而近似的项作差。hyperbolix 用无相消公式替代了这些减法运算,在先前实现返回 NaN 的距离处仍能在 float32 精度下保持准确。hyperbolix 以 MIT 许可证发布,可在 https://github.com/timoklein/hyperbolix 获取。
cs.LG / 107 / 2609.28263

Resource-Adaptive Stochastic Gradient Descent for Online Linear Programming without Re-solving

无需重新求解的面向在线线性规划的资源自适应随机梯度下降算法
Lyu, Jiameng
Abstract
The growth of large language model (LLM) inference and search services increases the scale of online linear programming problems, motivating computationally efficient algorithms. We develop resource-adaptive stochastic gradient descent (RASGD) for stochastic online linear programming. The algorithm uses one request and current inventory to update resource prices, requiring O(m) operations for m resources and memory per arrival and no LP or sample-average optimization. The central idea is to express the current-resource pricing logic of re-solving through a first-order SGD update: each arrival refreshes the remaining-inventory allowance in the dual objective, while the stepsize decreases for early learning and increases later to match the speed of inventory adjustment. Under standard non-degeneracy conditions, our algorithm is feasible on every sample path and achieves O(\log T) expected regret against the realized fractional hindsight optimum, which matches the lower bound, even for policies that know the distribution and have unrestricted computation. The analysis converts curvature around the fixed reference price into inventory stability without tracking optimal prices at changing resource levels. Numerical experiments show that RASGD achieves regret competitive with per-arrival LP re-solving and improves upon the tested first-order baselines, while retaining the computational efficiency of first-order methods. These results establish RASGD as a computationally efficient approach to achieving high allocation quality in large-scale OLP.
Chinese Translation
大语言模型(LLM)推理与搜索服务的增长使得在线线性规划问题的规模不断扩大,因此需要计算高效的算法。我们针对随机在线线性规划提出了一种资源自适应随机梯度下降算法(Resource-Adaptive Stochastic Gradient Descent, RASGD)。该算法仅利用单个请求和当前库存来更新资源价格,每次到达只需 O(m) 次运算和 O(m) 的存储空间(m 为资源数量),且无需求解线性规划或进行样本均值优化。其核心思想是:将重新求解(re-solving)中的当前资源定价逻辑表达为一阶 SGD 更新——每次到达会更新对偶目标中剩余库存的允许额度,同时步长在早期递减以促进学习、在后期递增以匹配库存调整的速度。在标准的非退化条件下,我们的算法在每条样本路径上均可行,并且相对于已实现的分数事后最优解取得 O(log T) 的期望遗憾,该结果与下界相匹配,即使对于已知分布且计算能力不受限制的策略亦是如此。该分析将固定参考价格附近的曲率转化为库存稳定性,而无需在资源水平变化时追踪最优价格。数值实验表明,RASGD 取得了与逐到达 LP 重新求解相竞争的遗憾表现,并优于所测试的一阶基线方法,同时保持了一阶方法的计算效率。这些结果确立了 RASGD 作为一种在大规模在线线性规划中实现高分配质量的计算高效方法。
cs.LG / 108 / 2609.28385

When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment

何时何地信任教师:通过熵校准的信用分配统一在线策略蒸馏与GRPO
Zhang, Jie, Yang, Jingxiao, Huang, Zhehao, Liu, Yuhang, Huang, Xiaolin
Abstract
Reinforcement learning with verifiable rewards (RLVR) supervises mathematical reasoning through final-answer correctness, but provides little guidance on individual tokens. On-policy distillation (OPD) supplies dense feedback on student-generated responses, yet teacher preference need not reflect correctness. Recent hybrids combine OPD and verifier-derived advantages or reweight task credit using teacher ratios. However, teacher guidance enters after verifier-based group normalization, and token reweighting need not preserve the total task credit assigned to each response. We introduce Unified Entropy-Calibrated Credit Redistribution for GRPO (UECR-GRPO), which integrates verifier and teacher signals within a single GRPO-style update at both the response and token levels. \emph{Path-Utility Unification} (PUU) combines verifier reward and a teacher-to-anchor path log-ratio in a single KL-regularized objective. Its on-policy implementation uses a length-normalized teacher score and combines both rewards before group normalization and PPO clipping, allowing teacher evidence to influence the response ranking. \emph{Entropy-Calibrated Redistribution} (ECR) then uses the signed teacher--old-policy token gap to redistribute the verifier-derived component. Full-vocabulary teacher entropy attenuates uncertain guidance, while a response-wise zero-sum projection preserves the total task credit and its token-wise sign before clipping. Across five mathematical reasoning benchmarks, UECR-GRPO achieves average \(\mathrm{Avg@12}\) accuracies of 17.21\% and 65.09\% with Qwen3-1.7B and Qwen3-4B students, respectively, exceeding the strongest baseline at each scale by 0.89 and 0.56 percentage points.
Chinese Translation
基于可验证奖励的强化学习(RLVR)通过最终答案的正确性来监督数学推理,但对单个词元提供的指导甚少。在线策略蒸馏(OPD)能够对学生生成的回复提供密集反馈,然而教师模型的偏好未必反映正确性。近期的混合方法将OPD与基于验证器的优势相结合,或利用教师比率对任务信用进行重新加权。然而,教师指导是在基于验证器的组归一化之后才引入的,且对词元的重新加权未必能保持分配给每条回复的任务信用总量。我们提出了面向GRPO的统一熵校准信用再分配方法(UECR-GRPO),它在回复和词元两个层级上将验证器信号与教师信号整合到单一GRPO风格的更新中。路径效用统一(Path-Utility Unification,PUU)将验证器奖励与教师到锚点的路径对数比率结合在一个KL正则化目标中。其在线策略实现采用长度归一化的教师评分,并在组归一化和PPO裁剪之前将两种奖励相组合,使教师证据能够影响回复的排序。随后,熵校准再分配(Entropy-Calibrated Redistribution,ECR)利用带符号的教师-旧策略词元差值来再分配源自验证器的奖励成分。全词表教师熵用于衰减不确定的指导,而逐回复的零和投影则在裁剪之前保持任务信用总量及其逐词元的符号不变。在五个数学推理基准上,UECR-GRPO使用Qwen3-1.7B和Qwen3-4B学生模型分别取得了17.21%和65.09%的平均\(\mathrm{Avg@12}\)准确率,在各自规模上分别超过最强基线0.89和0.56个百分点。
cs.LG / 109 / 2609.28399

Memory Attention

记忆注意力(Memory Attention)
Kang, Jiale
Abstract
Language models typically construct attention values from contextual hidden states, even when some of their content may be reusable across contexts. We investigate whether token-indexed memory can replace the dedicated value projection when complemented by contextual information. We propose Memory Attention (MA), which forms values by combining layer-specific token memory with contextual keys. The memory supplies token-specific representations, while the keys preserve context dependence. At inference, normalization can be folded into the memory tables, reducing value construction to lookup and addition. Token-indexed retrieval also enables CPU offloading with prefetching, reducing GPU parameter storage. Under matched training token budgets and with additional memory parameters, experiments across attention configurations show improved language modeling and average downstream performance.
Chinese Translation
语言模型通常基于上下文隐藏状态构建注意力值,即使其中部分内容可能在不同上下文中可复用。我们研究了在结合上下文信息的条件下,以词元(token)为索引的记忆能否替代专门的价值投影。我们提出了记忆注意力(Memory Attention, MA),它通过将层级特定的词元记忆与上下文键(key)相结合来构建价值(value)。记忆提供词元特定的表示,而键则保持对上下文的依赖性。在推理阶段,归一化可以被折叠进记忆表中,从而将价值构建简化为查表与加法操作。基于词元索引的检索还支持带预取功能的CPU卸载,从而减少GPU上的参数存储。在匹配的训练词元预算并引入额外记忆参数的条件下,跨多种注意力配置的实验表明,该方法在语言建模和平均下游任务性能上均有所提升。
cs.LG / 110 / 2609.28405

Learning Collective Dynamics with Differentiable Gaussian Representations

基于可微分高斯表示学习集体动态
Ma, Jianxiang, Zhang, Mingfu, Yang, Xiaocui, Gao, Yichen, Huang, Junzhao, Hou, Yuesong
Abstract
Collective responses depend on individual differences, contact opportunities, and accumulated experience. Learning their dynamics from aggregate counts requires connecting a population's response distribution to both current observations and future behavior. We introduce Differentiable Gaussian Dynamics (DGD), which learns this connection through three components: a Gaussian mixture representing heterogeneous response propensities, differentiable aggregation of contact intensity and behavioral probabilities, and feedback recurrence that updates subsequent responses. Reparameterized integration and temporal recurrence let aggregate prediction errors jointly train the distribution, observation functions, and feedback parameters. On four windows from KuaiRand-Pure and Online Retail II, DGD achieves lower joint behavioral negative log-likelihood than a DeepAR adaptation with a joint-behavior head. In Retail 2010, its one-day behavioral-count MAE is 4.71 versus 6.88 for this adaptation. Learning the distribution reduces behavioral negative log-likelihood by 10.82% relative to a fixed Gaussian in KuaiRand's standard-recommendation window; removing feedback dynamics raises joint KL from 0.0340 to 0.2577 in a controlled experiment. These results establish the value of learning population representations and their feedback process from aggregate observations. Code is available at https://github.com/OranAi-Ltd/oransim.
Chinese Translation
集体响应取决于个体差异、接触机会以及累积经验。从聚合计数中学习其动态,需要将群体的响应分布与当前观测及未来行为联系起来。我们提出了可微分高斯动力学(Differentiable Gaussian Dynamics, DGD),通过三个组成部分学习这种联系:表示异质性响应倾向的高斯混合模型、可微分的接触强度与行为概率聚合,以及更新后续响应的反馈递归。重参数化积分与时间递归使聚合预测误差能够联合训练分布、观测函数和反馈参数。在 KuaiRand-Pure 和 Online Retail II 的四个窗口上,DGD 的联合行为负对数似然低于带有联合行为输出头的 DeepAR 适配版本。在 Retail 2010 中,其单日行为计数 MAE 为 4.71,而该适配版本为 6.88。在 KuaiRand 的标准推荐窗口中,学习分布相较于固定高斯使行为负对数似然降低了 10.82%;在对照实验中,移除反馈动态会使联合 KL 散度从 0.0340 升至 0.2577。这些结果证明了从聚合观测中学习群体表示及其反馈过程的价值。代码可在 https://github.com/OranAi-Ltd/oransim 获取。
cs.LG / 111 / 2609.28409

Learning Holographic Reduced Representations with Clifford Variational Autoencoders

基于Clifford变分自编码器学习全息降维表示
Abid, Mohamed Malek, Furlong, P. Michael
Abstract
Vector Symbolic Algebras project data structures into a hyperdimensional vector space through the application of their vector algebras to randomly generated atomic vector symbols and fractional power encodings of real-valued data. Embedding unstructured data remains an open question. We present \textit{Clifford-VAE}, a variational autoencoder that learns to project data onto a Clifford torus in arbitrary dimensions. Experiments using the MNIST, FashionMNIST, and CIFAR-10 datasets demonstrate that Clifford-VAE produces representations that are competitive with those produced by Gaussian and Hyperspherical VAEs for semi-supervised classification tasks while outperforming Gaussian and Hyperspherical counterparts in the VSA benchmark tests of self-binding and unbinding, role-filler recovery, and bundle capacity. Clifford-VAE provides a principled technique for grounding perceptual data into a symbolic reasoning framework, providing a new approach to a long-standing problem in the VSA literature.
Chinese Translation
向量符号代数(Vector Symbolic Algebras, VSA)通过将向量代数应用于随机生成的原子向量符号以及实值数据的分数幂编码,将数据结构投射到超维向量空间中。然而,如何嵌入非结构化数据仍然是一个悬而未决的问题。我们提出了 extit{Clifford-VAE},一种学习将数据投影到任意维度Clifford环面上的变分自编码器。基于MNIST、FashionMNIST和CIFAR-10数据集的实验表明,在半监督分类任务中,Clifford-VAE所产生的表示与高斯VAE(Gaussian VAE)和超球面VAE(Hyperspherical VAE)所产生的表示具有竞争力;而在自绑定与解绑(self-binding and unbinding)、角色-填充恢复(role-filler recovery)以及捆绑容量(bundle capacity)等VSA基准测试中,Clifford-VAE则优于高斯和超球面VAE。Clifford-VAE为将感知数据锚定到符号推理框架中提供了一种有原则的方法,为VSA文献中长期存在的一个难题提供了一种新的解决途径。
cs.LG / 112 / 2609.28427

Context-Continuous Preference Learning for Exoskeleton Personalization

面向外骨骼个性化的情境连续偏好学习
Baek, Sunin, Park, Sungwoo, Kim, Daekyum
Abstract
Personalizing exoskeleton assistance across operating conditions is constrained by the time and physical effort required to collect user feedback. We examined whether a user's preference landscape varies smoothly across operating conditions and when this continuity supports learning from limited feedback. We propose Context-Continuous Preference Learning (CCPL), a Gaussian-process preference model that shares observations across nearby contexts while retaining context-specific utility estimates. We evaluated CCPL through simulations and retrospective analyses of ankle and elbow exoskeleton preference data from nine healthy adults. In simulations, CCPL improved reconstruction and preference-based Bayesian optimization relative to independent learning when preferences varied smoothly, but showed negative transfer when continuity was weak. In both human studies, full-data reference landscapes estimated separately for each participant and context tended to be more similar between nearby operating conditions. With five exposures per context, CCPL increased mean reconstruction correlation with these references from 0.644 to 0.720 for ankle assistance and from 0.476 to 0.526 for elbow assistance relative to independent learning. The five-exposure budget was approximately 37% lower for ankle and 17% lower for elbow than the estimated independent-learning budgets needed to match these correlations. CCPL also improved held-out response prediction relative to independent learning, while benefits over pooled learning varied. These findings support context continuity as a basis for sharing preference observations under limited feedback, although benefits for online personalization in humans remain to be established.
Chinese Translation
在不同运行条件下个性化外骨骼辅助受到收集用户反馈所需时间和体力消耗的限制。我们研究了用户的偏好景观是否随运行条件平滑变化,以及这种连续性何时能够支持从有限反馈中进行学习。我们提出了情境连续偏好学习(Context-Continuous Preference Learning, CCPL),这是一种高斯过程偏好模型,可在相邻情境之间共享观测数据,同时保留针对特定情境的效用估计。我们通过仿真以及对九名健康成年人踝关节和肘关节外骨骼偏好数据的回溯性分析对CCPL进行了评估。在仿真中,当偏好平滑变化时,CCPL相较于独立学习改进了偏好景观重构和基于偏好的贝叶斯优化,但在连续性较弱时表现出负迁移。在两项人体实验中,针对每位参与者和每个情境分别估计的全数据参考偏好景观在相邻运行条件之间往往更为相似。在每个情境仅有五次试验曝光的条件下,相较于独立学习,CCPL将与参考景观的平均重构相关系数从0.644提升至0.720(踝关节辅助),从0.476提升至0.526(肘关节辅助)。要达到上述相关系数,CCPL所需的五次曝光预算相比独立学习估计预算,踝关节约低37%,肘关节约低17%。相较于独立学习,CCPL还改进了对留存响应的预测,而相较于汇聚学习的优势则因情况而异。这些发现支持将情境连续性作为在有限反馈下共享偏好观测的基础,尽管其在人类在线个性化中的实际收益仍有待验证。
cs.LG / 113 / 2609.28438

Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections

极小范数单变量两层ReLU分类:精确解与带跳跃连接的全局最优性
Drabik, Karolina, Lewis, Ben, Puch, Antoni, Boursier, Etienne, Hofman, Piotr, Englert, Matthias, Lazić, Ranko
Abstract
We study minimal-norm interpolation and $\ell_2$-regularized logistic-loss minimization for binary classification by univariate two-layer ReLU networks. We give complete geometric characterizations of the optimal classifiers in function space, resolving how the solutions depend on whether hidden-layer biases are included in the parameter norm. When biases are unpenalized, the minimal-norm interpolators are exactly the continuous piecewise-affine functions that hug every label switch and have kinks of the appropriate convexity. When biases are penalized, the minimizer is unique in function space, has exactly one kink in each intermediate same-label segment, and is therefore a sparsest positive-margin classifier. We further show that adding a free affine skip connection leaves these function-space solutions unchanged but fundamentally improves the parameter-space landscape: every KKT point of the constrained problem becomes globally optimal, whereas suboptimal KKT points can occur without the skip connection. We establish analogous global-optimality and geometric results for sufficiently weak $\ell_2$-regularization of the logistic loss. In the unpenalized-bias case, we identify an additional sparsity-like restriction, implying that most minimal-norm interpolators cannot arise as small-regularization limits of margin-normalized logistic-loss minimizers. Numerical experiments across varying dataset complexity and network width support the predicted landscape and sparsity phenomena.
Chinese Translation
我们研究了用于二分类的单变量两层ReLU网络的极小范数插值与$\ell_2$正则化逻辑损失最小化问题。我们在函数空间中给出了最优分类器的完整几何刻画,解决了当隐藏层偏置是否被纳入参数范数时解的依赖关系。当偏置不被惩罚时,极小范数插值器恰好是那些紧贴每一个标签切换点并具有适当凸性拐点的连续分段仿射函数。当偏置被惩罚时,极小解在函数空间中是唯一的,在每个中间同标签段内恰有一个拐点,因此是最稀疏的正间隔分类器。我们进一步证明,添加一个自由的仿射跳跃连接不会改变这些函数空间解,但从根本上改善了参数空间的优化景观:约束问题的每一个KKT点都变为全局最优,而没有跳跃连接时可能出现次优的KKT点。我们针对逻辑损失的足够弱$\ell_2$正则化建立了类似的全局最优性与几何结果。在偏置不受惩罚的情形下,我们识别出一种额外的类稀疏性约束,这意味着大多数极小范数插值器不可能作为间隔归一化逻辑损失极小解在小正则化极限下产生。跨不同数据集复杂度和网络宽度的数值实验支持了所预测的优化景观与稀疏性现象。
cs.LG / 114 / 2609.28442

Order-Invariant Answers, Order-Sensitive Representations in Mathematical Reasoning

数学推理中的答案不变性与顺序敏感的表征
Tao, Zhixu Silvia
Abstract
Reordering a set of mathematical rules without changing its meaning should preserve the correct answer, but must a model's internal representations stay invariant too? We investigate this question using synthetic multi-step function-composition problems, each presented under multiple rule orderings with the same correct answer. We measure accuracy and permutation signal-to-noise ratio (SNR), which quantifies how distinctly ordering patterns are represented relative to variation across problem instances. Across 16 language models ranging from 1B to 8B parameters, we find a pattern: models that solve reordered problems more accurately represent different rule orderings more distinctly. Layer-averaged permutation SNR is positively rank-correlated with accuracy in every synthetic setting we evaluate, with Spearman correlations reaching 0.86. These findings highlight a distinction between answer invariance and representation invariance: successful mathematical rule composition can accompany distinct internal representations between equivalent rule orderings. This motivates distinguishing answer invariance from representation invariance, and offers a representational perspective on mathematical reasoning beyond answer accuracy alone.
Chinese Translation
在不改变含义的前提下重新排列一组数学规则,正确答案应当保持不变,但模型的内部表征是否也必须保持不变?我们利用合成的多步函数复合问题来研究这一问题,每个问题以多种规则排列顺序呈现,且正确答案相同。我们测量准确率和置换信噪比(SNR),后者量化了顺序模式相对于不同问题实例之间的差异在表征中的显著程度。在16个参数量从1B到8B的语言模型中,我们发现一种规律:能更准确解决重排问题的模型,其内部对不同规则排列顺序的表征也更加清晰可辨。在我们评估的所有合成设置中,按层平均的置换信噪比与准确率均呈正相关秩相关,Spearman相关系数最高达到0.86。这些发现凸显了答案不变性与表征不变性之间的区别:成功的数学规则复合可以伴随着等价规则排列之间截然不同的内部表征。这促使我们将答案不变性与表征不变性区分开来,并为超越单纯答案准确率的数学推理研究提供了表征层面的视角。
cs.LG / 115 / 2609.28459

Even Sharper Bounds for Transductive Learning and Its Applications

直推式学习的更精确界及其应用
Yang, Yingzhen
Abstract
We introduce Sharper Transductive Local Complexity (STLC), a localized complexity method for transductive learning under uniform sampling without replacement. The construction starts from a Bernstein-type concentration inequality for the supremum of the test--train empirical process. Its proof uses the modified log-Sobolev inequality for the swap walk and a two-parameter entropy closure. A peeling argument with a surrogate localization functional then gives excess-risk bounds with the same fixed-point and confidence terms as the classical inductive local Rademacher-complexity bounds, without the additional logarithmic confidence factor in earlier transductive results. For realizable learning over a binary class of VC dimension $\dVC$, with training size $m$, test size $u$, and $u\ge m\ge\dVC$, STLC yields $\cO\{\dVC\log(me/\dVC)/m\}$. This matches the standard inductive rate and, when $m\ge9$, is within a logarithmic factor of the transductive minimax lower bound of order $\dVC/m$. For transductive kernel learning, STLC gives a spectrum-adaptive excess-risk bound without the multiplicative imbalance factors appearing in the earlier local-complexity bound.
Chinese Translation
我们提出了更精确的直推式局部复杂度(Sharper Transductive Local Complexity, STLC),这是一种针对均匀无放回采样下直推式学习的局部化复杂度方法。该构造始于一个关于测试—训练经验过程上确界的Bernstein型集中不等式,其证明使用了交换随机游走(swap walk)的修正对数Sobolev不等式以及双参数熵闭包。随后,通过带有替代局部化泛函的剥离(peeling)论证,我们得到了与经典归纳式局部Rademacher复杂度界具有相同不动点和置信度项的超额风险界,且去除了早期直推式结果中额外的对数置信度因子。对于VC维为$dVC$的二值类上的可实现学习,在训练集大小为$m$、测试集大小为$u$且$u \ge m \ge dVC$的条件下,STLC给出 $\cO\{dVC \log(me/dVC)/m\}$ 的界。该结果与标准归纳式速率相匹配,且当 $m \ge 9$ 时,与阶为 $dVC/m$ 的直推式极小极大下界仅相差一个对数因子。对于直推式核学习,STLC给出了谱自适应的超额风险界,且不含早期局部复杂度界中出现的乘性不平衡因子。