← Back to Index
Daily Research Digest

arXiv Papers

2026-09-21
305
Papers
4
Categories
305
Translated
收藏清单 0
机器人学 (Robotics)
99
cs.RO / 1 / 2609.20892

WM-VS: Progress-Aligned World Models for Closed-Loop Visual Servoing

WM-VS:面向闭环视觉伺服的进度对齐世界模型
Sun, Guanzhong, Ma, Junyi, Zhou, Yixuan, Wu, Yuxuan, Miao, Yanzi, Wang, Hesheng
Abstract
Closed-loop visual servoing requires predictions that indicate whether an action reduces task error, not only whether the action is plausible. We call this gap the prediction-control mismatch and introduce WM-VS, a target-centric progress-aligned world-model framework for closed-loop visual servoing. Offline target-region DINOv2 correspondences define a signed four-dimensional servo coordinate for translation, scale, and in-plane rotation. Stage 1 aligns action-conditioned latent transitions with this coordinate; Stage 2 freezes the world model and trains a reactive joint-velocity policy with action imitation, consequence supervision, and short imagined rollouts that favor error contraction. Deployment is RGB-only and reactive, without online trajectory optimization. On a real 7-DoF eye-to-hand system, WM-VS reaches a corner RMSE no larger than 10 percent of its initial value in 30/30 trials and retains this criterion at the final valid frame in 25/30 (83.33 percent). Removing future-error alignment reduces retention to 26.67 percent. The learned progress signal agrees with an external AprilTag corner error not used for training or control (mean Spearman rho = 0.8778). Without retraining, two unseen 3D targets achieve translation-error reductions of 86.48 percent and 90.27 percent and rotation-error reductions of 70.01 percent and 65.70 percent. These results link progress-aligned action consequences to repeated closed-loop correction and transfer. Code and data will be released as open source.
Chinese Translation
闭环视觉伺服需要的预测不仅要能判断动作是否合理,还要能指示该动作是否能降低任务误差。我们将这一差距称为预测-控制失配(prediction-control mismatch),并提出WM-VS——一个以目标为中心、进度对齐的世界模型框架,用于闭环视觉伺服。离线的目标区域DINOv2对应关系定义了一个带符号的四维伺服坐标系,涵盖平移、尺度和面内旋转。第一阶段(Stage 1)将动作条件化的潜在状态转移与该坐标对齐;第二阶段(Stage 2)冻结世界模型,并通过动作模仿、结果监督以及偏向误差收缩的短时想象推演,训练一个反应式关节速度策略。部署阶段仅需RGB输入且为反应式,无需在线轨迹优化。在一个真实的7自由度(7-DoF)眼在手外(eye-to-hand)系统上,WM-VS在30/30次试验中均使角点RMSE降至不超过初始值的10%,并在25/30次(83.33%)试验的最后一帧有效画面中仍保持该标准。移除未来误差对齐后,保持率降至26.67%。所学到的进度信号与未用于训练或控制的外部AprilTag角点误差具有一致性(平均Spearman rho = 0.8778)。在不进行再训练的情况下,两个未见过的3D目标分别实现了86.48%和90.27%的平移误差缩减,以及70.01%和65.70%的旋转误差缩减。这些结果将进度对齐的动作后果与反复的闭环校正及迁移能力联系起来。代码和数据将以开源形式发布。
cs.RO / 2 / 2609.20965

AeRove: A Compact Bimodal Aerial-Terrestrial Drone with Rapid Bistable Reconfiguration for Close-Range Pipeline Inspection

AeRove:一种用于近距离管道检测的紧凑型双模式空地无人机,具备快速双稳态重构能力
Polillio, Caleb, Swissler, Petras
Abstract
Close-proximity pipeline inspection is challenging for standard drones due to high hovering power consumption, airflow sensitivity, and propeller wash interference with gas sensing. To address this, we present AeRove, a compact bimodal aerial-terrestrial robot. AeRove uses its propeller guards as wheels to roll along pipes and employs a spring-loaded bistable mechanism to reconfigure between ground and flight modes in 200 ms without continuous actuator power to maintain either state. The converging propeller-guard geometry also improves measured thrust efficiency. By reserving flight for obstacle hopping and using ground rolling for continuous traversal, current draw is reduced 14x compared to continuous flight (0.7 A vs. 10 A), extending estimated travel distance from 144 m to 2057 m. Autonomous trials on a 51 cm diameter steel pipe demonstrated navigation on straight and curved sections, obstacle jumping, and leak detection of simulated inspection markers. Separate CO2 sensing experiments evaluated gas-detection performance under perched and aerial operating conditions. Perched inspection produced a substantially larger concentration response than hovering under the tested conditions. During flight, an underbody propeller-intake configuration produced a faster and stronger gas-detection response than an extended probe by leveraging propeller-induced airflow. Code and designs are released under the CC-BY license.
Chinese Translation
对于标准无人机而言,近距离管道检测具有挑战性,原因包括悬停功耗高、对气流敏感以及螺旋桨气流对气体检测的干扰。为解决这些问题,我们提出了AeRove,一种紧凑型双模式空地机器人。AeRove利用其螺旋桨护罩作为滚轮沿管道滚动,并采用弹簧加载的双稳态机构在200毫秒内实现地面模式与飞行模式之间的切换,且无需持续驱动器供电即可保持任一状态。向内收敛的螺旋桨护罩几何结构还提升了实测推力效率。通过将飞行保留用于障碍跳跃,并利用地面滚动实现连续行进,相比持续飞行,电流消耗降低了14倍(0.7 A对比10 A),估计行进距离从144米延长至2057米。在直径51厘米钢管上的自主试验展示了其在直段和弯段上的导航、障碍跳跃以及模拟检测标记的泄漏检测能力。此外,还开展了独立的CO2感知实验,评估了栖停与空中作业条件下的气体检测性能。在测试条件下,栖停检测产生的浓度响应显著大于悬停检测。在飞行过程中,借助螺旋桨诱导的气流,机身下方螺旋桨进气构型比外伸探测杆产生了更快更强的气体检测响应。相关代码与设计已以CC-BY许可证发布。
cs.RO / 3 / 2609.20970

Shake to Learn: Dynamic Interrogation of Hidden Object Physics for Robotic Manipulation with Physical Reservoir Computing

摇晃以学习:基于物理储备池计算的隐藏物体物理特性动态探询用于机器人操作
Lor, Wen Sin, Wang, Jun, Li, Suyi
Abstract
Many physical properties relevant to robotic manipulation are hidden from vision. A sealed object, for example, may reveal little about its center of mass (COM) or internal contents until it is lifted, shaken, or otherwise dynamically perturbed. This study shows that such interactions can enable a new modality of robotic perception and learning, in which interaction-induced dynamic responses are used to infer object physics that is inaccessible to conventional sensing. We implement this idea using an origami-inspired soft robotic arm that functions as a physical reservoir computer. After grasping an object, the arm is excited by a fixed shaking input at its base, and the resulting ringdown response is recorded through either camera tracking or embedded sensors. Because the input is held constant across trials, hidden object properties, such as the COM position, are encoded through their effect on the dynamics of the coupled robot-object system. A lightweight linear readout can then decode these dynamics to recover interpretable information about the hidden object physics. Using this framework, the soft robotic arm reservoir completed three tasks of increasing difficulty: inferring the orientation of the object's hidden COM, inferring the COM distance from the grasp point, and using the inferred COM information to guide a subsequent regrasp. We further develop a dynamic summary representation of the ringdown response that improves prediction accuracy. Together, these results establish shake-to-learn mechanical interrogation as a promising strategy for robotic systems to convert brief physical interactions into actionable cues about hidden object properties for downstream manipulation.
Chinese Translation
许多与机器人操作相关的物理属性是视觉无法感知的。例如,一个密封物体在被提起、摇晃或以其他方式动态扰动之前,可能几乎无法揭示其质心(COM)或内部内容。本研究表明,此类交互可以开启一种新的机器人感知与学习模态,即利用交互诱导的动态响应来推断传统传感手段无法获取的物体物理特性。我们通过一个受折纸启发的软体机械臂来实现这一想法,该机械臂充当物理储备池计算机(physical reservoir computer)。在抓取物体后,机械臂在其基部受到固定的摇晃输入激励,并通过相机跟踪或嵌入式传感器记录由此产生的振铃(ringdown)响应。由于输入在各次试验中保持恒定,隐藏的物体属性(如质心位置)通过对耦合机器人-物体系统动态特性的影响而被编码其中。随后,一个轻量级的线性读出层可以解码这些动态,从而恢复关于隐藏物体物理特性的可解释信息。利用该框架,软体机械臂储备池完成了三个难度递增的任务:推断物体隐藏质心的方向、推断质心与抓取点的距离,以及利用推断出的质心信息指导后续的重新抓取。我们进一步开发了一种振铃响应的动态摘要表示,可提高预测精度。这些结果共同确立了“摇晃以学习”式机械探询作为一种有前景的策略,使机器人系统能够将短暂的物理交互转化为关于隐藏物体属性的可操作线索,以支持下游操作任务。
cs.RO / 4 / 2609.20980

ForeTac-VLA: A Forecasting-Based Tactile-Vision-Language-Action Model for Contact-Rich Robotic Manipulation

ForeTac-VLA:一种基于预测的触觉-视觉-语言-动作模型,面向接触密集型机器人操作
Tao, Zhengyu, Li, Xin, Wang, Xin
Abstract
Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, yet their reliance on visual perception limits robustness in contact-rich environments, where critical physical interaction states may not be visually observable. Existing tactile-enhanced VLA methods improve physical grounding using observed tactile feedback, but most remain largely reactive rather than explicitly modeling how contact may evolve. Therefore, we propose ForeTac-VLA, a forecasting-based tactile-vision-language fusion model that predicts future tactile states to guide action generation. Specifically, ForeTac-VLA encodes recent tactile observations into temporal representations and integrates them with vision-language features through bidirectional cross-attention. Further, a transformer-based forecasting module predicts multi-step future tactile states, enabling the model to reason jointly over observed and anticipated contact. Finally, the fused multimodal representations and predicted future tactile states are fed into the VLA backbone to condition action generation. To stabilize training, a ground-truth-to-prediction curriculum is employed when early forecasts are unreliable. Across four real-world contact-rich manipulation tasks, ForeTac-VLA achieves an average success rate of 95%, outperforming the fine-tuned VLA model by 36.25 percentage points and state-of-the-art tactile-enhanced VLA baselines by over 22 percentage points. ForeTac-VLA also maintains strong performance under low-illumination and visually cluttered conditions. Video demonstrations can be found on https://foretac-vla.github.io/
Chinese Translation
视觉-语言-动作模型在机器人操作中已展现出强大的能力,但其对视觉感知的依赖限制了其在接触密集环境中的鲁棒性,因为在这些环境中,关键的物理交互状态可能无法通过视觉观察获得。现有的触觉增强VLA方法利用观测到的触觉反馈改善了物理接地能力,但大多数方法在很大程度上仍是反应式的,并未显式建模接触状态的演化过程。为此,我们提出了ForeTac-VLA,一种基于预测的触觉-视觉-语言融合模型,通过预测未来的触觉状态来引导动作生成。具体而言,ForeTac-VLA将近期触觉观测编码为时间表征,并通过双向交叉注意力将其与视觉-语言特征进行融合。此外,一个基于Transformer的预测模块可预测多步未来触觉状态,使模型能够对已观测到的和预期的接触状态进行联合推理。最后,融合后的多模态表征和预测的未来触觉状态被输入VLA主干网络,以调节动作生成。为了稳定训练过程,在早期预测不可靠时采用了从真值到预测的课程学习策略。在四项真实世界的接触密集型操作任务中,ForeTac-VLA取得了95%的平均成功率,比微调后的VLA模型高出36.25个百分点,比最先进的触觉增强VLA基线方法高出22个百分点以上。此外,ForeTac-VLA在低光照和视觉杂乱的条件下仍保持了优异的性能。视频演示可参见 https://foretac-vla.github.io/
cs.RO / 5 / 2609.20983

PIVOT: Physically Informed Vision-Language Off-Road Traversability for Field Robot Navigation

PIVOT:基于物理信息的视觉语言越野可通行性评估用于野外机器人导航
Jiao, Aoran, Zhao, Wenda, Sahak, Hshmat, Barfoot, Timothy D.
Abstract
Terrain assessment is a critical capability for off-road mobile robots, enabling safe and reliable navigation through unstructured and geometrically complex environments. Conventional geometry-based terrain assessment is fast to compute but often overly conservative in unstructured environments. We present PIVOT: a Physically Informed Vision-Language Off-Road Traversability navigation system that augments conventional geometry-based planning with vision-language-model (VLM)-based semantic reasoning for field robots. To physically ground this assessment, we quantify how strongly the VLM's predicted traversal energy cost, robot vibration, and wheel slip correlate with real-world measurements and introduce a unified traversability score that weights each modality by its prediction-measurement correlation. For efficiency, we design a two-level navigation architecture that retains geometry-based planning as the nominal mode and invokes semantic replanning only when that mode fails to find a path. Across five repeated closed-loop trials on a mixed-terrain route totalling around $6.4$ km, the proposed system increases overall autonomy from $59.6\%$ to $97.0\%$, reduces human interventions from $11$ to $3$, and increases the mean distance between interventions (MDBI) from $69.2$ m to $412.9$ m compared with geometry-only navigation. These results demonstrate that physically grounded VLM-based terrain assessment can substantially extend autonomous navigation beyond the limitations of geometry alone, while preserving efficient geometric planning as the nominal mode.
Chinese Translation
地形评估是越野移动机器人的关键能力,使其能够在非结构化且几何复杂的环境中实现安全可靠的导航。传统的基于几何的地形评估计算速度快,但在非结构化环境中往往过于保守。我们提出了PIVOT:一个基于物理信息的视觉语言越野可通行性导航系统,它将传统的基于几何的规划与基于视觉语言模型(VLM)的语义推理相结合,服务于野外机器人。为了使该评估具有物理依据,我们量化了VLM预测的通行能量消耗、机器人振动和车轮打滑与真实世界测量值之间的相关强度,并引入了一个统一的可通行性评分,该评分根据每种模态的预测-测量相关性对其进行加权。为了提高效率,我们设计了一个两级导航架构,将基于几何的规划保留为标称模式,仅当该模式无法找到路径时才调用语义重规划。在总长约6.4公里的混合地形路线上的五次重复闭环试验中,与仅使用几何的导航相比,所提出的系统将整体自主率从59.6%提升至97.0%,将人工干预次数从11次减少到3次,并将干预间的平均距离(MDBI)从69.2米提高到412.9米。这些结果表明,具有物理依据的基于VLM的地形评估能够大幅超越单纯几何方法的局限,扩展自主导航能力,同时保留高效的几何规划作为标称模式。
cs.RO / 6 / 2609.21000

Do Spinning Radar Doppler Velocity Measurements Improve Vehicle Detection and Tracking?

旋转雷达多普勒速度测量能否改善车辆检测与跟踪?
Xie, Eric, Lisus, Daniil, Barfoot, Timothy D.
Abstract
Spinning frequency-modulated continuous-wave (FMCW) radars have been gaining popularity in autonomous vehicle perception on account of their robustness to adverse weather conditions and 360{\deg} field of view. Recently, scanning radars have also been shown capable of generating per-azimuth Doppler velocity. In this paper, we investigate whether these Doppler velocity measurements improve spinning radar vehicle detection and tracking performance. For detection, we estimate the ego motion and use it to undo the Doppler range distortion of the radar image before passing it to a network. For tracking, we propose a new way to estimate a per-vehicle velocity and use it as a prior for the tracker's motion model. Since Doppler-enabled spinning radar data is not available in any dataset with ground-truth dynamic object labels, our first contribution is an automatic labelling pipeline that uses an ensemble of fine-tuned off-the-shelf lidar detectors to label all 643 km of the Boreas Road Trip dataset. We then transfer detections to radar, and use over 250 km of vehicle-dense sequences as ground-truth training data. By training and evaluating two state-of-the-art detectors, we show that Doppler undistortion can improve detection accuracy by up to $2.37$ points on mean average precision. Furthermore, we show that the Doppler velocity prior can improve tracking accuracy by $13.68$ points on multi-object tracking accuracy (MOTA) versus the zero-velocity initialization baseline, while achieving $99.7\%$ of the MOTA obtained using ground-truth velocities as the prior.
Chinese Translation
旋转式调频连续波(FMCW)雷达因其对恶劣天气条件的鲁棒性和360度视场,在自动驾驶车辆感知中日益受到青睐。近年来,扫描雷达还被证明能够生成逐方位角的多普勒速度。本文研究这些多普勒速度测量能否提升旋转雷达的车辆检测与跟踪性能。在检测方面,我们估计自车运动,并利用其在将雷达图像输入网络之前消除多普勒引起的距离畸变。在跟踪方面,我们提出一种估计单车辆速度的新方法,并将其用作跟踪器运动模型的先验。由于目前没有任何包含动态目标真值标签的数据集提供多普勒功能的旋转雷达数据,我们的首个贡献是一个自动标注流水线,该流水线使用经微调的现成激光雷达检测器集成方法,为 Boreas Road Trip 数据集全部 643 公里的数据添加标签。随后我们将检测框迁移到雷达数据上,并使用超过 250 公里的车辆密集序列作为真值训练数据。通过训练和评估两种最先进的检测器,我们表明多普勒畸变校正可将检测精度(平均精度均值)提升高达 $2.37$ 个百分点。此外,我们表明,与零速初始化基线相比,多普勒速度先验可将跟踪精度(多目标跟踪精度,MOTA)提升 $13.68$ 个百分点,同时达到以真值速度作为先验时 MOTA 的 $99.7\%$。
cs.RO / 7 / 2609.21005

Project SCOUT: Interceptor Drone for Perimeter Defense

Project SCOUT:用于周界防御的拦截无人机
Yousuf, Azmain, Cai, Siwei, Peterson, Knut, Zhou, Lifeng, Han, David
Abstract
The rapid proliferation of unauthorized unmanned aerial vehicles (UAVs) has created a growing need for robust, jamming-resistant counter-UAV systems for perimeter defense. This paper presents \textbf{SCOUT} (Spatial Computation for Optimized UAV Tracking), a ROS-integrated onboard perception and control framework for real-time aerial defense against incoming UAVs. SCOUT performs visual detection, target association, track filtering, and control command generation directly onboard the defender UAV, without relying on external sensing infrastructure or ground-station computation. To provide stable control inputs, the perception pipeline combines TensorRT-accelerated drone detection with ByteTrack-based association and a lightweight track-retention state machine. The state machine rejects abrupt target jumps and maintains short-term target continuity during temporary detection degradation, reducing unstable control responses caused by false detections or target switching. We evaluate the proposed architecture through an integrated hardware deployment executing a planar ``goalkeeping'' interception strategy. In this setting, the defender UAV tracks the incoming target and adjusts its motion to maintain a blocking configuration near the protected boundary. Real-world flight results show that SCOUT maintains valid target detections for 92.2\% of frames while operating at real-time onboard detection rates, demonstrating the feasibility of visual tracking and closed-loop control for UAV perimeter defense. A video demonstration of the end-to-end perimeter defense operation is available online. https://figshare.com/s/497befc033091fc0b84f
Chinese Translation
未经授权的无人机(UAV)的迅速扩散,使得周界防御对鲁棒、抗干扰的反无人机系统需求日益增长。本文提出了SCOUT(Spatial Computation for Optimized UAV Tracking,面向无人机跟踪优化的空间计算),一个集成ROS的机载感知与控制框架,用于对来袭无人机进行实时空中防御。SCOUT在防御方无人机上直接完成视觉检测、目标关联、航迹滤波和控制指令生成,无需依赖外部传感基础设施或地面站计算。为提供稳定的控制输入,感知流水线将TensorRT加速的无人机检测与基于ByteTrack的目标关联以及一个轻量级航迹保持状态机相结合。该状态机能够拒绝目标的突然跳变,并在检测暂时退化期间保持短期目标连续性,从而减少由误检或目标切换引起的不稳定控制响应。我们通过一个执行平面“守门员”式拦截策略的集成硬件部署对该架构进行了评估。在该场景中,防御方无人机跟踪来袭目标并调整自身运动,以在被保护边界附近维持阻挡构型。真实飞行结果表明,SCOUT在以实时机载检测速率运行的同时,能对92.2%的帧保持有效目标检测,验证了视觉跟踪与闭环控制在无人机周界防御中的可行性。端到端周界防御操作的视频演示可在线获取:https://figshare.com/s/497befc033091fc0b84f
cs.RO / 8 / 2609.21008

SPARROW: Survival-POMCP for Adaptive Robot Routing, Observation, and Waiting

SPARROW:面向自适应机器人路径规划、观测与等待的生存POMCP方法
Sahak, Hshmat, Jiao, Aoran, Rhinehart, Nicholas, Barfoot, Timothy D.
Abstract
Temporary obstacles that may block a robot's planned route create a sequential navigation problem: a robot must decide whether to wait for a blockage to clear, reroute, or acquire more information about the obstacle before acting. We formulate graph navigation among temporary obstacles as a partially observable semi-Markov decision process and introduce SPARROW, a belief-space planner built on Partially Observable Monte Carlo Planning (POMCP). SPARROW searches over traversal, observation, and finite-duration waiting actions while maintaining a particle belief over latent obstacle classes and clearance times. Class-conditioned survival models are learned online from both clearance observations and right-censored encounters where the robot reroutes before clearance is observed. A generative model simulates obstacle arrivals and clearances as each action unfolds, so the planner can account for blockages that may occur along alternative routes. We further introduce a value-of-learning criterion that trades the immediate cost of collecting labelled survival data against its expected reduction in future navigation regret. Across two simulation graphs and multiple obstacle-class settings, SPARROW reduces mean time-to-goal by 12-26% relative to OSCAR, a recent survival-based method for the same problem. On a physical mobile robot, SPARROW reduces mean time-to-goal by 20.5% relative to OSCAR while selectively observing, waiting, and rerouting as environment conditions change.
Chinese Translation
可能阻挡机器人规划路径的临时障碍物会引发一个序贯导航问题:机器人必须决定是等待障碍物消除、重新规划路径,还是在行动之前获取关于障碍物的更多信息。我们将存在临时障碍物的图导航问题形式化为部分可观测半马尔可夫决策过程,并提出SPARROW——一个基于部分可观测蒙特卡洛规划(POMCP)构建的信念空间规划器。SPARROW在遍历、观测和有限时长的等待动作之间进行搜索,同时维护关于潜在障碍物类别和消除时间的粒子信念。类别条件化的生存模型通过在线学习获得,其数据既来自消除时刻的观测,也来自机器人尚未观测到障碍消除便改道的右删失情形。生成式模型在动作展开时模拟障碍物的出现与消除,使规划器能够考虑可能在替代路径上出现的阻挡。我们进一步引入了一个学习价值准则,用于权衡收集带标签生存数据的即时代价与其对未来导航遗憾的预期降低。在两个仿真图和多种障碍物类别设置下,SPARROW相较于OSCAR(一种针对同一问题的近期基于生存模型的方法)将平均到达目标时间缩短了12-26%。在实体移动机器人上,SPARROW相较于OSCAR将平均到达目标时间缩短了20.5%,并能够随环境条件变化有选择地进行观测、等待和改道。
cs.RO / 9 / 2609.21015

Towards Effective Visual-Inertial SLAM with Passive-Only Sensors for Low-Cost Autonomous Underwater Vehicles

面向低成本自主水下航行器的仅被动传感器高效视觉惯性SLAM
Schwidder, Grant, Widhalm, David, Sattar, Junaed
Abstract
Improvements to Visual-Inertial Simultaneous Localization and Mapping (VI-SLAM) for low-cost autonomous underwater vehicles (AUVs) are critical for transitioning advanced marine robotics from specialized labs to broader research and hobbyist applications. While high-end AUVs typically rely on expensive sensor suites - such as Doppler Velocity Logs (DVLs) and Ultra-Short Baseline (USBL) systems - this work demonstrates that robust, high-quality navigation is achievable using a sub-$10, 000(USD) platform equipped only with inexpensive consumer-grade sensors. By leveraging a similarly priced, open-source AUV, we evaluate the performance of stereo cameras, Micro-electromechanical System (MEMS)-based IMUs, and depth sensors in a fully unconstrained 6-degree-of-freedom (6-DOF) underwater environment. We analyze the efficacy of off-the-shelf SLAM packages and propose optimizations for sensor fusion to mitigate the visual and physical challenges of untethered underwater operation. Our results prove that a usable SLAM solution can be accessible to the masses, providing a benchmark for expectations in demanding, real-time maritime missions without the financial barrier of industrial-grade hardware.
Chinese Translation
改进低成本自主水下航行器(AUV)的视觉惯性同步定位与建图(VI-SLAM),对于推动先进海洋机器人技术从专业实验室走向更广泛的研究和爱好者应用至关重要。虽然高端AUV通常依赖昂贵的传感器套件——例如多普勒速度仪(DVL)和超短基线(USBL)系统——但本研究表明,使用一台造价低于10,000美元、仅配备廉价消费级传感器的平台,即可实现稳健且高质量的导航。我们借助一台价格相近的开源AUV,在完全无约束的六自由度(6-DOF)水下环境中评估了立体相机、基于微机电系统(MEMS)的惯性测量单元(IMU)以及深度传感器的性能。我们分析了现成SLAM软件包的有效性,并提出了传感器融合方面的优化方案,以缓解无缆水下作业在视觉和物理层面的挑战。实验结果证明,可用的SLAM解决方案能够惠及大众,为苛刻的实时海上任务提供了性能基准,且无需承担工业级硬件的经济门槛。
cs.RO / 10 / 2609.21022

Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs

抓住我:面向响应式视觉-语言-动作模型的实时反馈去噪
Ji, Yiheng, Zhou, Xingru, Sentis, Luis, Seo, Mingyo
Abstract
Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by combining semantic knowledge from pretrained vision-language models with expressive action-generation policies. Diffusion-based action generators are particularly effective for modeling temporally coherent action chunks, but these chunks are typically executed open-loop after inference. This limits responsiveness when objects move, contacts change, or the scene evolves during execution. We propose VLA-Feedback, a two-timescale architecture that combines low-frequency diffusion planning with high-frequency visual feedback. Rather than fully denoising an action chunk before execution, VLA-Feedback retains its final denoising step as a lightweight feedback interface, allowing each action to be corrected using the latest observation before it is executed. This design preserves the expressiveness of the diffusion planner while enabling real-time action correction without rerunning the full vision-language diffusion model. VLA-Feedback matched GR00T on static LIBERO tasks while improving average success on dynamic simulation tasks from 27.5% to 85.0%. On real-robot tasks, it improved average success from 51% to 73%. Additional materials can be found on our project page: https://vla-feedback.github.io.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型通过将预训练视觉-语言模型中的语义知识与富有表现力的动作生成策略相结合,在机器人操作任务中展现出了强大的泛化能力。基于扩散模型的动作生成器在建模时间上连贯的动作序列(action chunks)方面尤为有效,但这些动作序列通常在推理后以开环方式执行。当物体移动、接触状态变化或场景在执行过程中发生演变时,这会限制系统的响应能力。我们提出了 VLA-Feedback,一种结合低频扩散规划与高频视觉反馈的双时间尺度架构。VLA-Feedback 并非在执行前对动作序列进行完全去噪,而是保留其最终去噪步骤作为一个轻量级的反馈接口,从而允许在每一步动作执行前利用最新观测对其进行修正。这一设计在保留扩散规划器表达能力的同时,无需重新运行完整的视觉-语言扩散模型即可实现实时的动作修正。在静态 LIBERO 任务上,VLA-Feedback 与 GR00T 的表现相当,同时将动态仿真任务的平均成功率从 27.5% 提升至 85.0%。在真实机器人任务上,它将平均成功率从 51% 提升至 73%。更多资料请访问我们的项目主页:https://vla-feedback.github.io。
cs.RO / 11 / 2609.21024

Adapting Rigid-Body Dynamics Derivatives for Constraint Embedding Closed-Chain Models

将刚体动力学导数方法适配于约束嵌入的闭链模型
Volpi, Daniel J., Wensing, Patrick M.
Abstract
This paper extends an existing algorithm for the first-order derivatives of rigid-body dynamics to the case of closed-chain kinematic systems modeled using constraint embed- ding. Many standard dynamics algorithms apply to both open- chain and constraint-embedded models, but existing efficient derivative methods assume joint velocity effects are locally config- uration invariant. We remove this assumption and derive adapted algorithms that extend dynamics derivatives to more general joint types, including those arising in constraint-embedded closed- chain models. Our results compare conventional pin-joint robot models with more complete actuation models that capture local closed chains. We show that the additional terms introduced by these generalizations have low computational impact when mod- eling actuation kinematics alone, but can incur higher cost when additional rigid bodies, such as motor rotors, are included in the actuation chain, or when considering non-local loops. Overall, these results enable more accurate dynamics computations for constraint-embedded actuation submechanisms to be adopted in model-predictive control and differentiable simulation.
Chinese Translation
本文将现有的刚体动力学一阶导数算法扩展至采用约束嵌入方法建模的闭链运动学系统。许多标准动力学算法同时适用于开链模型和约束嵌入模型,但现有高效的导数计算方法假设关节速度效应在局部上与构型无关。我们移除了这一假设,推导出适配后的算法,将动力学导数扩展到更一般的关节类型,包括约束嵌入闭链模型中出现的关节类型。我们的结果将传统的销关节机器人模型与能够捕捉局部闭链的更完整执行器模型进行了比较。我们表明,当仅对执行器运动学建模时,这些推广所引入的额外项的计算开销较低;但当执行链中包含额外的刚体(如电机转子)时,或在考虑非局部闭环时,可能会带来更高的计算成本。总体而言,这些结果使得约束嵌入执行子机构的更精确动力学计算能够应用于模型预测控制和可微仿真中。
cs.RO / 12 / 2609.21026

MAPLE-RF: Efficient Probabilistic RF Source Localization in Partially Explored Environments

MAPLE-RF:部分探索环境中的高效概率射频源定位
Lei, Haozhe, Rangan, Sundeep
Abstract
Localizing a radio-frequency (RF) transmitter from received signals often requires a model of the environment to predict how obstacles block and reflect the signal. In many robotic applications, however, only a partial map is available, particularly when a robot localizes the source while exploring with simultaneous localization and mapping (SLAM). We study single-snapshot transmitter localization on such partially explored maps and compare two approaches that output a posterior over transmitter locations. The first extends a digital-twin method, which ray-traces every candidate location, to partial maps by treating unexplored space as free and training on mixed map coverage. The second, MAPLE-RF, encodes estimated path angles of arrival and signal-to-noise ratios as grid channels aligned with map knownness, occupancy, and line-of-sight visibility, and a U-Net scores all candidate positions in one pass without simulating propagation at inference. Ray-tracing simulations of indoor environments indicate that training on mixed map coverage is essential for both approaches. The digital-twin approach is more accurate on most single-snapshot metrics, while MAPLE-RF comes close at a query cost that does not depend on the propagation model and is more than two orders of magnitude below a fresh full-grid query with general-purpose ray tracing. Both outperform Gaussian and Gaussian-mixture baselines, and on exploration routes guided by its own estimates, fused MAPLE-RF posteriors place more probability near the source than the compared methods. Code and data will be released.
Chinese Translation
根据接收信号对射频(RF)发射机进行定位通常需要环境模型,以预测障碍物如何阻挡和反射信号。然而,在许多机器人应用中,往往只有部分地图可用,尤其是当机器人在利用同步定位与建图(SLAM)进行探索的同时对源进行定位时。我们研究了在这类部分探索地图上的单快照发射机定位问题,并比较了两种输出发射机位置后验概率的方法。第一种方法将数字孪生方法(对每个候选位置进行光线追踪)扩展到部分地图,方法是将未探索空间视为自由空间,并在混合地图覆盖率的数据上进行训练。第二种方法MAPLE-RF将估计的路径到达角和信噪比编码为与地图已知性、占据状态和视线的可见性对齐的网格通道,并由U-Net一次性对所有候选位置进行评分,在推理时无需模拟传播。室内环境的射线追踪仿真表明,在混合地图覆盖率数据上训练对两种方法都至关重要。数字孪生方法在大多数单快照指标上更为准确,而MAPLE-RF的性能与之接近,且其查询成本不依赖于传播模型,比使用通用射线追踪进行全新全网格查询的成本低两个数量级以上。两种方法均优于高斯和高斯混合基线,并且在其自身估计引导的探索路径上,融合的MAPLE-RF后验概率在源附近分配的概率高于所比较的其他方法。代码和数据将会发布。
cs.RO / 13 / 2609.21045

DEXTERA: From a Single Image to Deployable Dexterous Manipulation via Real-to-Sim-to-Real

DEXTERA:基于Real-to-Sim-to-Real从单张图像到可部署灵巧操作
Wu, Jin, Yuan, Lianjie, Sun, Zeyan, Lei, Yuanyuan, A, Disi, Han, Bicheng, Xia, Fangzhou
Abstract
Collecting real-world robot data for dexterous manipulation is costly and time-consuming. While high-fidelity physics simulators enable scalable data synthesis and policy learning, constructing deployment-ready digital twins manually remains labor-intensive, and residual visual, geometric, and dynamics gaps hinder reliable sim-to-real transfer. We present DEXTERA, an automated real-to-sim-to-real framework that transforms a single RGB image into deployable policies for dexterous manipulation across four unified stages: (1) single-image scene factorization into a static Gaussian background and interactive rigid or articulated assets with VLM-inferred physical parameters; (2) metric scene global alignment, object canonicalization, and morphology-balanced robot calibration; (3) scalable simulator task primitive construction, VR teleoperation, and object-centric trajectory synthesis; and (4) a shared multimodal policy interface supporting both imitation learning and reinforcement learning. We evaluate DEXTERA across 13 task-embodiment pairs, 2 dexterous robot platforms, and 6 policy architectures. Experimental results demonstrate that DEXTERA achieves superior visual fidelity and 3D geometric reconstruction compared to generative baselines, while cross-domain trajectory replays validate strong physical interaction consistency. Furthermore, simulation-only trained policies enable viable zero-shot real-robot deployment, while simulation-real co-training substantially improves mean physical policy success from 29.2% to 61.9% across diverse policy architectures.
Chinese Translation
为灵巧操作采集真实世界机器人数据成本高昂且耗时。虽然高保真物理仿真器能够实现可扩展的数据合成与策略学习,但人工构建可用于部署的数字孪生仍然费时费力,且残存的视觉、几何与动力学差距阻碍了可靠的sim-to-real迁移。我们提出DEXTERA,一个自动化的real-to-sim-to-real框架,能够将单张RGB图像转化为可部署的灵巧操作策略,涵盖四个统一阶段:(1)将单张图像场景分解为静态高斯背景以及具有交互能力的刚体或铰接物体资产,并由视觉语言模型(VLM)推断物理参数;(2)度量尺度的场景全局对齐、物体规范化(canonicalization)以及形态平衡的机器人标定;(3)可扩展的仿真器任务基元构建、VR遥操作以及以物体为中心的轨迹合成;(4)支持模仿学习与强化学习的共享多模态策略接口。我们在13组任务-本体组合、2个灵巧机器人平台和6种策略架构上对DEXTERA进行了评估。实验结果表明,与生成式基线相比,DEXTERA实现了更优的视觉保真度和三维几何重建质量,同时跨域轨迹回放验证了其强大的物理交互一致性。此外,仅在仿真中训练的策略即可实现可行的零样本真机部署,而仿真-真实协同训练将多种策略架构下的平均物理策略成功率从29.2%显著提升至61.9%。
cs.RO / 14 / 2609.21046

Constraint-Unified MPC for Over-Actuated Surface Vehicles with Post-Detection Fault Reconfiguration

面向过驱动水面车辆的约束统一模型预测控制及故障后检测重构
Burmester, Sebastian, Jiang, Ruiheng, Sendlhofer, Noa, D'Andrea, Raffaello, Ramachandran, Aswin
Abstract
Choreographed aquatic performances require small autonomous surface vehicles to track precise paths under per-thruster force and rate limits, including after thruster failures. We report a deployed system in which trajectory tracking, thrust allocation, and the per-thruster force and rate limits are resolved in a single quadratic program over the per-thruster commands, with fault reconfiguration entering through one binary flag per thruster from an external detector. The system has driven a fleet in live performances on Lake Z\"urich and at the Time Space Existence 2025 exhibition in Venice. The field campaign measures 1.6 cm root mean square position error in a 10-minute hold and 4.3 cm over a 10 m square at 0.6 m/s. At 0.5 m/s, losing the front thruster increases the error to 11.6 cm, while losing the starboard-side thruster increases it to 11.5 cm. Losing two thrusters simultaneously leaves the craft tracking a 0.4 m/s square with a root mean square position error of 1.25 m. Compared to our own cascaded baseline, the unified formulation tracks the nominal square to the same few centimeters and holds station more tightly with roughly half the thruster force. The architectures separate after a thruster failure, where the unified controller stays within 24 cm of the reference path while the cascade leaves it.
Chinese Translation
编排好的水上表演要求小型自主水面车辆在单个推进器的力和速率限制下精确跟踪路径,包括在推进器发生故障之后。我们报告了一套已部署的系统,其中轨迹跟踪、推力分配以及每个推进器的力和速率限制在单个二次规划中基于各推进器的指令统一求解,故障重构则通过来自外部检测器的每个推进器一个二进制标志位引入。该系统已在苏黎世湖的现场表演以及威尼斯的Time Space Existence 2025展览中驱动了一支船队。实地实验测量结果显示:在10分钟定点保持中位置均方根误差为1.6厘米,以0.6米/秒速度跟踪10米正方形时误差为4.3厘米。在0.5米/秒速度下,失去前推进器使误差增至11.6厘米,失去右舷推进器使误差增至11.5厘米。同时失去两个推进器时,船只仍能以1.25米的均方根位置误差跟踪0.4米/秒的正方形轨迹。与我们的级联基线方法相比,统一形式在标称正方形轨迹跟踪上达到相同的几厘米精度,并在使用大约一半推进器力的情况下更紧密地保持定点。两种架构在推进器故障后表现分化:统一控制器保持在参考路径24厘米以内,而级联控制器则偏离了参考路径。
cs.RO / 15 / 2609.21059

PlantShade: Predicting Plant Shadows for Lighting-Aware Robotic Agricultural Operation

PlantShade:面向光照感知机器人农业作业的植物阴影预测
Da, Longchao, Liu, Xiaoou, Li, Xingjian, Xiang, Lirong, Wei, Hua
Abstract
Plant growth and agricultural production form the foundation of a country's sustainable development and directly impact human livelihoods. Recent advances in frontier artificial intelligence have enabled scientific agriculture with strong potential to improve crop productivity. In this paper, we identify the importance and inherent complexity of plant shade simulation, as shading is a critical factor influencing plant growth. To advance this field and promote broader societal benefits, we focus on two main contributions. First, we introduce a comprehensive plant growth and shade dataset covering four plant species, including soybean, tomato, sugarbeet, and strawberry. The dataset includes top-down viewpoints with a supplementary light along a circular trajectory, casting dynamic shadows across multiple growth stages and diverse observation complexities. Second, we propose generative shade simulation based on diffusion models, enabling realistic shade generation for unseen plants and supporting downstream robotic tasks such as perception, lighting control, and view planning. The model incorporates temporal conditioning to facilitate flexible shade simulation across different time stages. We conduct both quantitative and qualitative evaluations to assess model performance. This work provides a foundational study for plant-aware shade modeling and has meaningful implications for broader agricultural and robotic applications.
Chinese Translation
植物生长与农业生产是国家可持续发展的基础,直接关系到人类的生计。前沿人工智能的最新进展使科学农业成为可能,在提高作物生产力方面具有巨大潜力。在本文中,我们指出了植物阴影模拟的重要性及其固有的复杂性,因为遮挡是影响植物生长的关键因素。为推动该领域的发展并带来更广泛的社会效益,我们提出了两项主要贡献。首先,我们引入了一个覆盖四种植物(包括大豆、番茄、甜菜和草莓)的综合性植物生长与阴影数据集。该数据集包含俯视视角,并在圆形轨迹上配有补光,在多个生长阶段和多样的观测复杂度下投射出动态阴影。其次,我们提出了一种基于扩散模型(diffusion models)的生成式阴影模拟方法,能够为未见过的植物生成逼真的阴影,并支持感知、光照控制和视角规划等下游机器人任务。该模型引入了时间条件机制,以便在不同时间阶段实现灵活的阴影模拟。我们通过定量和定性评估来衡量模型性能。本工作为植物感知的阴影建模提供了基础性研究,对更广泛的农业与机器人应用具有重要意义。
cs.RO / 16 / 2609.21082

Design of Adaptive PID Controller Based On Asynchronous Advantage Actor Critic Learning Method for QuadCopter Control

基于异步优势Actor-Critic学习方法的四旋翼自适应PID控制器设计
Jokar, Ali, Alasty, Aria
Abstract
Quadcopters offer great utility in many applications, but their nonlinear nature and disturbance sensitivity present great control challenges. Basic PID controllers are generally not sophisticated enough to cope with these complexities. This paper suggests a control system that integrates the Asynchronous Advantage Actor-Critic (A3C) algorithm with a PID controller for quadcopter attitude and trajectory tracking. The A3C controller uses parallel agents to optimize PID parameters dynamically using a neural network. A system identification module for the complementary system makes predictions about system states for optimal control policy. The proposed framework was compared with a standard actor-critic (A2C) model. Simulation results verify that they both track accurately. However, the A3C-based controller converges much more for the loss function, as evidenced by reward figures and loss curves, demonstrating better parameter optimization. This shows that A3C-based approach results in improved performance for the control of quadcopter, effectively integrating reinforcement learning and traditional control to achieve higher adaptability.
Chinese Translation
四旋翼飞行器在众多应用中具有巨大价值,但其非线性特性和对扰动的敏感性带来了巨大的控制挑战。基本的PID控制器通常不足以应对这些复杂性。本文提出了一种将异步优势Actor-Critic(A3C)算法与PID控制器相结合的控制系统,用于四旋翼的姿态控制和轨迹跟踪。A3C控制器利用并行智能体通过神经网络动态优化PID参数。用于互补系统的系统辨识模块对系统状态进行预测,以实现最优控制策略。该框架与标准的Actor-Critic(A2C)模型进行了比较。仿真结果验证了两者均能实现精确跟踪。然而,从奖励曲线和损失曲线可以看出,基于A3C的控制器在损失函数上收敛性显著更优,证明了其更好的参数优化能力。这表明基于A3C的方法能提升四旋翼的控制性能,有效地将强化学习与传统控制相结合,实现了更高的适应性。
cs.RO / 17 / 2609.21099

Dynamic Modeling and LQR Control of a Single Coaxial Drone with 2DOF Thrust Vectoring Mechanism

具有二自由度推力矢量机构的单轴共轴无人机的动力学建模与LQR控制
Jokar, Ali, Talaeizadeh, Amin, Alasty, Aria
Abstract
Coaxial rotor drones have generated considerable interest because of energy efficiency and small size, but they are afflicted with inherent underactuation for roll and pitch control, although systems like swashplates have circumvented this limitation at the cost of greater mechanical complexity. This work presents a novel coaxial drone supplemented by a two-degreesof-freedom pendulum mechanism for active thrust vectoring that offers a less mechanically complicated alternative. We develop a comprehensive Lagrangian dynamic model that does not ignore the inertial contributions of all the components, including body, servo arms, and motor assembly. A Linear Quadratic Regulator(LQR) is designed based on the linearized dynamics around the hover equilibrium. High-fidelity simulations taking actuator dynamics and sensor noise into account validate the proposed architecture. An Extended Kalman Filter (EKF) blends GPS, barometer, and IMU estimates with high accuracy for state estimation. The findings verify the potential and reliability of this approach for power-saving, rapid coaxial UAVs.
Chinese Translation
共轴双旋翼无人机因其能效高、体积小而备受关注,但其滚转和俯仰控制在本质上处于欠驱动状态;尽管自动倾斜器(swashplate)等机构可以规避这一局限,但代价是机械复杂度显著增加。本文提出了一种新型共轴无人机,其配备二自由度摆式机构以实现主动推力矢量控制,提供了一种机械复杂度更低的替代方案。我们建立了全面的拉格朗日动力学模型,该模型未忽略包括机体、舵机摇臂和电机组件在内所有部件的惯性贡献。基于悬停平衡点附近的线性化动力学,设计了线性二次型调节器(LQR)。考虑执行器动态和传感器噪声的高保真仿真验证了所提出的架构。扩展卡尔曼滤波器(EKF)融合GPS、气压计和IMU的估计信息,实现了高精度的状态估计。研究结果验证了该方法在节能、高速共轴无人机领域的潜力与可靠性。
cs.RO / 18 / 2609.21100

Dynamics-Induced Commitment in Learning-Based Robotic Penalty Kicks

基于学习的机器人点球中的动力学诱导承诺
Geng, Ruize, Zhang, Hao E., Li, Yisen, Wang, Yikai, Tseng, H. Eric, Zhao, Ding
Abstract
Learning in robotic games is constrained not only by strategic information but also by what the body can still execute. We study this coupling in a hierarchical humanoid-quadruped penalty system in which game-level self-play policies command fixed soccer whole-body controllers (S-WBCs). The humanoid shooting skill is initialized from self-collected motion-capture data, whereas the quadruped saving skill is learned by reinforcement learning. We introduce dynamics-induced commitment mapping (DIC-Map), a body-grounded analysis that estimates continuation capability, identifies the first persistent loss of a terminal alternative, and tests whether the remaining interaction admits a reduced zero-sum game. For symmetric terminal alternatives, the reduced game yields a closed-form bound on optimal strategy concentration determined by the responder's value of deferring. We further show that, when the responder acts through an estimator, equal response values eliminate the direct terminal-allocation gradient and leave an estimator-mediated first-order learning channel. Experiments locate commitment about 0.29 s before contact, and changing only ball speed shifts deferral coverage. Across four responder policies, replacing the estimator raises save rate from 0.240 to 0.472, whereas a comparable gain in read accuracy obtained by waiting raises it only to 0.246. Posterior analysis is used for the equilibrium comparison because the available coverage terms are observational proxies. Project website: https://chris-ruizegeng.github.io/penaltykick/
Chinese Translation
机器人博弈中的学习不仅受策略信息的约束,还受身体当前仍可执行动作的约束。我们在一个人形-四足层次化点球系统中研究这种耦合关系,其中博弈层的自博弈策略指挥固定的足球全身控制器(S-WBC)。人形机器人射门技能由自采集的运动捕捉数据初始化,而四足机器人扑救技能则通过强化学习获得。我们提出动力学诱导承诺映射(DIC-Map),这是一种基于身体能力的分析方法,用于估计后续行动能力、识别终端备选方案的首次持续性丧失,并检验剩余交互是否可归结为一个简化的零和博弈。对于对称的终端备选方案,简化博弈给出了一个闭式上界,用于确定最优策略集中度,该集中度由响应方的延迟决策价值决定。我们进一步证明,当响应方通过估计器行动时,相等的响应价值会消除直接的终端分配梯度,仅留下经由估计器中介的一阶学习通道。实验发现承诺点出现在触球前约0.29秒,且仅改变球速即可改变延迟覆盖范围。在四种响应方策略中,替换估计器将扑救率从0.240提升至0.472,而通过等待所获得的相当程度的预测准确率提升仅将扑救率提升至0.246。由于现有的覆盖度量仅为观测代理,我们使用事后分析进行均衡比较。项目网站:https://chris-ruizegeng.github.io/penaltykick/
cs.RO / 19 / 2609.21107

Learning Scene-Aware Humanoid Locomotion through 3D Clutter from Immersive Human Demonstrations

通过沉浸式人类示范从三维杂乱环境中学习场景感知的人形机器人运动
Wang, Beichen, Xu, Tong, Kosukhin, Daniel, Yeung, Yuen-Hei, Lu, Yuanjie, Xiao, Xuesu
Abstract
While learning from human motions has enabled highly dynamic humanoid skills such as dancing and martial arts in obstacle-free space, traversal through densely cluttered environments remains underexplored. These spaces are three-dimensional and geometrically constrained, requiring scene-aware locomotion that tightly couples whole-body motion with scene geometry for obstacle avoidance. To address these challenges, we present Moving Through Clutter (MTC), a learning-from-demonstration framework for scene-aware humanoid locomotion. To bypass costly physical scene construction, MTC uses procedurally generated Virtual Reality environments for immersive data collection. To transform these human motions into training-ready humanoid motions, we propose a scene-aware motion retargeting algorithm that converts human demonstrations into humanoid trajectories while strictly enforcing robot-scene clearance to guarantee collision-free traversal. These reference trajectories are then used to train a scene-aware locomotion policy that deploys on a Unitree G1 humanoid. Evaluated on our proposed MTC-Challenge for multi-obstacle traversal, the policy demonstrates a 70.2% collision-free rate across diverse scenarios, successfully traversing complex environments through diverse whole-body skills, including crawling through low-clearance passages and squeezing through narrow gaps.
Chinese Translation
尽管从人类动作中学习已经使人形机器人在无障碍空间中掌握了诸如舞蹈和武术等高度动态的技能,但穿越密集杂乱环境的研究仍然不足。这类空间是三维且受几何约束的,需要场景感知的运动能力,即将全身运动与场景几何紧密耦合以实现避障。为应对这些挑战,我们提出了 Moving Through Clutter(MTC),一个用于场景感知人形运动的示教学习框架。为避免代价高昂的物理场景搭建,MTC 使用程序化生成的虚拟现实(VR)环境进行沉浸式数据采集。为了将人类动作转换为可直接用于训练的人形机器人动作,我们提出了一种场景感知的运动重定向算法,将人类示范转换为人形机器人轨迹,同时严格执行机器人与场景之间的间隙要求,以保证无碰撞穿越。随后,这些参考轨迹被用于训练一个场景感知的运动策略,并部署于 Unitree G1 人形机器人。在我们提出的面向多障碍穿越的 MTC-Challenge 上进行评估,该策略在多样化场景中达到 70.2% 的无碰撞率,能够通过多样的全身技能成功穿越复杂环境,包括从低矮通道下爬行以及从狭窄缝隙中侧身挤过。
cs.RO / 20 / 2609.21112

Demonstration Synthesis from a Single Scan via Gaussian Splatting for Visuomotor Policy Learning

基于高斯泼溅的从单次扫描合成示范数据用于视觉运动策略学习
Wang, Beichen, Yeung, Yuen-Hei, Devarakonda, V. R. Sridhar, Xiao, Xuesu
Abstract
Training a visuomotor policy calls for abundant demonstrations that closely match the target environment, yet collecting them anew remains expensive. Existing demonstration synthesis methods reduce this cost but remain constrained by high manual effort, limited visual fidelity, or heavy reliance on physics simulators. This paper introduces GaussianFactory, a high-fidelity data engine that mass-produces demonstrations with a single video scan as its only human input and no physics engine in the generation loop. Specifically, GaussianFactory reconstructs the scene as an editable 3D Gaussian Splatting (3DGS) replica and samples from the object-combination tasks the scene affords. For each task, it plans grasps and trajectories purely kinematically on the geometric reconstruction, rendering photorealistic demonstrations that visually match the target environment. Physical dynamics enter the pipeline only where contact force interactions dictate the outcome---during grasp formation, via a learned contact model pretrained once on an interaction dataset. To evaluate the downstream utility of the synthesized demonstrations, we implement an end-to-end scan-to-deployment workflow in two setups: a simulated scene that stands in for the real world to enable reproducibility, and a real-world workspace with a physical UR10e robot. In each setup, a standard diffusion policy trained solely on the synthesized demonstrations achieves 95.1% and 84.2% success rates, respectively.
Chinese Translation
训练视觉运动策略需要大量与目标环境高度匹配的示范数据,然而重新采集这些数据代价高昂。现有的示范合成方法虽能降低这一成本,但仍受限于高昂的人工投入、有限的视觉保真度或对物理仿真器的严重依赖。本文提出 GaussianFactory,一个高保真数据引擎,它以单次视频扫描作为唯一的人工输入,并在生成流程中不使用物理引擎,即可批量生产示范数据。具体而言,GaussianFactory 将场景重建为可编辑的三维高斯泼溅(3D Gaussian Splatting, 3DGS)副本,并从场景所支持的目标-组合任务中进行采样。对于每个任务,它在几何重建上纯粹以运动学方式规划抓取与轨迹,并渲染出与目标环境视觉一致的照片级真实感示范。物理动力学仅在接触力交互决定结果之处进入该流程——即在抓取形成阶段,通过一个预先在交互数据集上一次性训练得到的接触模型。为评估所合成示范数据的下游实用性,我们在两种设置下实现了端到端的“扫描到部署”工作流:一是以仿真场景替代真实环境以保证可复现性,二是配备物理 UR10e 机器人的真实工作空间。在每种设置下,仅使用合成示范数据训练的标准扩散策略分别取得了 95.1% 和 84.2% 的成功率。
cs.RO / 21 / 2609.21114

Noctif3R: Feed-Forward Monocular Real-Time SLAM for Photon-Limited Scenes on Embedded Hardware

Noctif3R:面向嵌入式硬件、适用于极弱光场景的前馈单目实时SLAM
Chauhan, Mihir, Abhang, Aditya Uday, Mathew, Kevin Biju, Bera, Aniket
Abstract
Robots carrying out tasks in dark environments need to localize from a single RGB camera, in light so low that the per-pixel signal approaches the sensor's own noise, on a power-constrained onboard computer, in real time. Each of these constraints has matured pipelines, but the intersection does not. Offline low-light reconstruction now recovers structure below -4 dB but is far too slow to run in real time, while the real-time monocular systems a robot can actually carry (DROID-SLAM, DPV-SLAM, etc.) degrade or fail when SNR gets low. We measured how they fail: across the nine lowest darkness levels of our scenes, DROID-SLAM returns a full-length trajectory carrying no information about the camera's motion on all nine, VGGT-SLAM and CUT3R on eight, pi^3 on seven, and DPV-SLAM on four. We present SYS, a monocular pipeline built on a low-light feed-forward pointmap front end with an explicit match gate, which returns three tracked trajectories and no uninformative ones, at the lowest error of any method where it tracks (24-47% of the no-information ceiling against 56-73% for the strongest baseline), and at the narrowest coverage. On a real robot video take in which 86.5% of delivered frames are entirely black, every configuration of ours stops after the lit beginning, while DROID-SLAM and DPV-SLAM each emit a pose for all 1178 frames. Our method contribution is an embedded execution path for the Jetson AGX Orin: running the map, keyframes and backend at 384 pixels with tracking at 256, together with two fixes to the per-frame pose solve, is a replicated Pareto improvement, 1.28x throughput at 0.964x error on one scene and 1.42x at 0.68x on a second, with 47% less peak GPU memory and 29% less energy per pose. We evaluate on a calibrated, bit-exact regenerable noise ladder, on relabelled real-world dark exposures, and on a new dark-room video ladder recorded from a Boston Dynamics Spot robot.
Chinese Translation
在黑暗环境中执行任务的机器人需要仅依靠单个RGB相机进行定位,此时光照极低,每像素信号接近传感器自身噪声水平,且需在功耗受限的机载计算机上实时运行。上述每项约束各自已有成熟的技术方案,但三者的交集尚属空白。离线低光照重建如今已能恢复低于-4 dB的场景结构,但速度太慢,无法实时运行;而机器人实际可携带的实时单目系统(如DROID-SLAM、DPV-SLAM等)在信噪比(SNR)变低时会性能下降甚至失效。我们对这些系统的失效方式进行了测量:在我们场景中最低的九个黑暗等级下,DROID-SLAM在全部九个等级上返回了完整长度的轨迹,却不包含任何相机运动信息;VGGT-SLAM和CUT3R在八个等级上如此,pi^3在七个等级上如此,DPV-SLAM在四个等级上如此。我们提出SYS——一个建立在带显式匹配门控的低光照前馈点图(pointmap)前端之上的单目流水线——它返回三条有效跟踪轨迹且无无效轨迹,在其能跟踪的场景中误差为所有方法中最低(为无信息上限的24-47%,而最强基线为56-73%),且覆盖范围最窄。在一段86.5%的传输帧完全为黑色的真实机器人视频中,我们的所有配置都在有光照的开头部分之后停止,而DROID-SLAM和DPV-SLAM均对全部1178帧输出了位姿。我们的方法贡献是一条面向Jetson AGX Orin的嵌入式执行路径:以384像素运行建图、关键帧与后端,以256像素运行跟踪,并辅以对逐帧位姿求解的两项改进,实现了可复现的帕累托改进——在一个场景上吞吐量提升1.28倍且误差为0.964倍,在另一个场景上吞吐量提升1.42倍且误差为0.68倍,同时峰值GPU显存降低47%,每个位姿的能耗降低29%。我们在一个经过标定、可逐位复现的噪声梯度上、在重新标注的真实世界暗光拍摄上,以及在一个新的、由Boston Dynamics Spot机器人录制的暗室视频梯度上进行评估。
cs.RO / 22 / 2609.21122

MetaPusher: Meta Learning and Planning for Nonprehensile Manipulation of Unseen Objects with Rapid Online Adaption

MetaPusher:基于元学习与规划的未知物体非抓取操作及快速在线自适应
Lee, Donghyung, Golestaneh, Seyedali, Singh, Jaskrit, Zhong, Zhuoyun, Kapoutsis, Athanasios, Chamzas, Constantinos
Abstract
Manipulating previously unseen objects remains challenging, as their dynamics depend on latent physical properties, such as friction and mass distribution, that cannot be inferred from perception alone. Prior experience across objects can provide an initial estimate of unseen object dynamics, but this estimate remains uncertain and can degrade further during sim-to-real transfer. Adapting the dynamics through interaction can progressively refine the estimation, however, updating the model may invalidate the planned trajectory. Successful and efficient manipulation therefore requires both rapid dynamics adaptation and a planning strategy that can incorporate this evolution. In this work, we introduce MetaPusher, a meta-learning and adaptive planning framework for nonprehensile manipulation of unseen objects without prior object-specific interactions. A meta-learned dynamics model rapidly adapts from interactions during task execution, while an adaptive kinodynamic planner updates long-horizon plans by reusing and refining its existing search tree. This coupling enables manipulation and adaptation without a separate data collection phase. We evaluate MetaPusher on unseen objects in simulation and in sim-to-real scenarios, comparing against fine-tuning and active learning methods, MPPI-based control, and a reinforcement learning policy. It achieves lower prediction error and improves task success rate by up to 20%.
Chinese Translation
对之前从未见过的物体进行操作仍然极具挑战性,因为其动力学取决于摩擦、质量分布等潜在物理属性,而这些属性无法仅通过感知推断。先前的物体操作经验可以为未知物体动力学提供初始估计,但该估计仍然存在不确定性,并且在从仿真到现实的迁移过程中可能进一步退化。通过交互来适配动力学可以逐步改进估计,然而,模型更新可能会使已规划好的轨迹失效。因此,成功且高效的操作既需要快速的动力学自适应,也需要一种能够融合这种演化的规划策略。本文提出 MetaPusher,一个用于未知物体非抓取操作的元学习与自适应规划框架,无需预先进行物体特定的交互。元学习得到的动力学模型能够在任务执行过程中通过交互快速自适应,同时自适应的运动动力学规划器通过复用并精炼其现有搜索树来更新长时程规划。这种耦合使操作与自适应无需单独的数据采集阶段。我们在仿真和从仿真到现实的场景中对未知物体上评估了 MetaPusher,并与微调和主动学习方法、基于 MPPI 的控制以及强化学习策略进行比较。该方法实现了更低的预测误差,并将任务成功率提升了最多 20%。
cs.RO / 23 / 2609.21130

SAGE: Safety-Aligned Gradient Enforcement for Human--Robot Collaboration

SAGE:面向人机协作的安全对齐梯度强制方法
Li, Yisen, Zhang, Hao, Geng, Ruize, Tseng, Yves, Zhao, Ding, Tseng, H. Eric
Abstract
Multi-party human-robot collaboration poses a dual challenge: robot decisions should remain interpretable and auditable, while executed actions must satisfy safety constraints during physical interaction. Combining explainable decision-tree policies with control-barrier-function (CBF) filtering provides a promising architecture but creates two learning mismatches in multi-agent reinforcement learning. Safety projection changes the action applied to the environment, while the coupled proposal graph can misalign independently optimized actor updates with a team-level update. We present safety-aligned gradient enforcement (SAGE) to address both mismatches. Its shield-annealed internalization layer (SAIL) uses a differentiable finite-penalty proposal map while retaining the exact CBF quadratic program for execution, preserving constraint-normal sensitivity to internalize repeatedly active safety constraints. Team-averaged Lyapunov policy optimization (TALO) constructs a team-aware update reference and applies a Lyapunov half-space correction to regulate independent actor updates. Physical experiments with two humanoid robots and a human partner demonstrate deployment feasibility. Across nine simulation scenarios, SAGE achieves a 71.0% success rate with 0.5 collision steps per thousand environment steps. Ablations show that direct CBF filtering reduces collision frequency by 98.5% but decreases success from 67.3% to 59.3%. SAIL reduces proposal violation by 48.8% and proposal-execution correction by 85.2%, while TALO reduces the update-consistency gap by 50.8%.
Chinese Translation
多方人机协作带来了双重挑战:机器人决策应保持可解释性和可审计性,同时在物理交互过程中执行的动作必须满足安全约束。将可解释的决策树策略与控制障碍函数(CBF)滤波相结合提供了一种有前景的架构,但在多智能体强化学习中会产生两种学习失配问题。安全投影会改变施加于环境的动作,而耦合的提议图可能导致独立优化的执行者更新与团队级更新不一致。我们提出安全对齐梯度强制方法(SAGE)以解决这两种失配问题。其屏蔽退火内化层(SAIL)采用可微的有限惩罚提议映射,同时保留精确的CBF二次规划用于执行,保持对约束法向的敏感性以内化反复激活的安全约束。团队平均Lyapunov策略优化(TALO)构建团队感知的更新参考,并应用Lyapunov半空间校正来规范独立执行者的更新。两台人形机器人与一名人类伙伴的物理实验验证了部署的可行性。在九个仿真场景中,SAGE实现了71.0%的成功率,每千个环境步长仅0.5步碰撞。消融实验表明,直接CBF滤波可将碰撞频率降低98.5%,但会使成功率从67.3%降至59.3%。SAIL将提议违反减少48.8%,将提议-执行校正减少85.2%;TALO将更新一致性差距减少50.8%。
cs.RO / 24 / 2609.21138

Diverse and Adaptable Arm Coordination for Octopus-Crawling via Diffusion-Based Uncertainty-Aware Optimization

基于扩散模型的不确定性感知优化实现多样化且自适应的章鱼爬行臂协调
Kim, Seung Hyun, Chang, Heng-Sheng, Kazemi, Kimia, Mehta, Prashant, Gazzola, Mattia
Abstract
Octopus crawling motivates soft robots that exploit redundancy, yet discovering and organizing diverse coordination modes for adaptation remains challenging. To address this, we introduce a Diffusion-based Uncertainty-aware Optimization (DUO) algorithm that learns demonstration-free crawling controllers for a simulated, muscle-actuated CyberOctopus. This work represents the first application of diffusion-based control to soft multi-arm robots in contact-rich simulations. By embedding a variety of locomotion behaviors within a shared control distribution, this approach enables the simulated octopus to navigate dynamic physical constraints, demonstrating that learned coordination diversity inherently facilitates robust adaptation. The main contributions include: (i) a symmetry-structured policy representation that folds radially equivalent controllers into a canonical directional sector, (ii) an online black-box optimization strategy, the DUO algorithm, that discovers and retains diverse coordination modes, and (iii) a control editing technique that adapts existing controllers to novel actuator constraints without retraining. These results show how learned coordination diversity makes motor abundance a practical resource for adaptation in soft multi-arm robots.
Chinese Translation
章鱼爬行为利用冗余度的软体机器人提供了启发,然而如何发现并组织多样化的协调模式以实现自适应仍然具有挑战性。为解决这一问题,我们提出了一种基于扩散模型的不确定性感知优化(Diffusion-based Uncertainty-aware Optimization,DUO)算法,用于为仿真中的肌肉驱动CyberOctopus学习无需示范数据的爬行控制器。本工作首次将基于扩散模型的控制方法应用于接触丰富仿真环境中的软体多臂机器人。通过将多种运动行为嵌入到共享的控制分布中,该方法使仿真章鱼能够应对动态物理约束,表明学习到的协调多样性本身即可促进鲁棒的自适应。主要贡献包括:(i)一种对称性结构的策略表示,将径向等价的控制器折叠到规范的方向扇区内;(ii)一种在线黑盒优化策略,即DUO算法,能够发现并保留多样化的协调模式;(iii)一种控制编辑技术,无需重新训练即可将现有控制器适配到新的执行器约束。这些结果表明,学习到的协调多样性使运动冗余成为软体多臂机器人自适应的实用资源。
cs.RO / 25 / 2609.21155

Same World, Different Knowledge: When Isolated Audits Misjudge World-Model Repairs

同一世界,不同的知识:孤立审计何时会误判世界模型的修复方案
Min, Rui, Li, Xianyao, Xu, Fang, Lachab, Sofiane, Du, Jing
Abstract
A repair favored under an isolated input fault can be inferior when deployed modules share the faulty information. We introduce an information-interface audit for world models, distinguishing fidelity gaps, where exact inputs become estimates, from availability gaps, where inputs are missing. Fixed-weight interventions measure prediction error, input dependence, and paired closed-loop benefit, including dependencies introduced by reconstruction. In simulated quadrotor model predictive control, coupled, opposite-sign 10% mass/thrust calibration errors reduce a physics-anchored model's success from 69% to 8%; uncertainty training restores 65%. Wind reconstruction recovers control benefit but inherits calibration dependence. For a positive calibration offset, reconstruction-only corruption favors uncertainty-trained reconstruction, whereas shared corruption favors the baseline. Acceleration diagnostics reveal compensation between reconstruction bias and nominal-model error, also observed with a disturbance observer. Repair selection therefore depends on the information paths used in deployment.
Chinese Translation
在孤立输入故障情形下被青睐的修复方案,当部署中的模块共享错误信息时可能表现更差。我们提出一种针对世界模型的信息接口审计方法,区分保真度缺口(精确输入变为估计值)与可用性缺口(输入缺失)。固定权重干预用于测量预测误差、输入依赖性以及成对闭环收益,其中包括由重构引入的依赖关系。在仿真四旋翼模型预测控制中,耦合且符号相反的10%质量/推力标定误差使物理锚定模型的成功率从69%降至8%;不确定性训练可恢复至65%。风速重构能够恢复控制收益,但继承了标定依赖性。对于正向标定偏移,仅重构式损坏有利于经过不确定性训练的重构方法,而共享式损坏则有利于基线方法。加速度诊断揭示了重构偏差与标称模型误差之间的补偿关系,这一现象在扰动观测器中同样可以观察到。因此,修复方案的选择取决于部署时所使用的信息路径。
cs.RO / 26 / 2609.21167

MA-LIPP: Cooperative Multi-Agent Load-Aware Informative Path Planning for Heterogeneous Robot Teams

MA-LIPP:面向异构机器人团队的多智能体负载感知协同信息路径规划
Kim, Hojune, Shi, Guangyao, Sukhatme, Gaurav S.
Abstract
Field robotics missions often require physical samples to be returned to laboratories for analysis, making path planning inherently load-aware and order-dependent as accumulated samples increase payload and traversal energy costs. In single-robot Load-Aware Informative Path Planning (LIPP), this rigidly couples sensing with hauling: a solitary robot must transport every collected sample, forcing frequent depot returns that severely restrict its spatial coverage. Heterogeneous multi-robot teams can overcome this bottleneck by dividing labor---enabling high-precision samplers to collect while high-capacity carriers handle transport. However, this introduces a complex coordination challenge regarding when, where, what, and to whom handoffs should occur on top of the LIPP problem. To address this tightly coupled problem, we introduce Multi-Agent LIPP (MA-LIPP), which enables teams to cooperate through asynchronous "dead drops," allowing one robot to deposit samples for another to retrieve later without requiring synchronous rendezvous. We formulate MA-LIPP as an exact Mixed-Integer Quadratic Program (MIQP) alongside a scalable Pairwise Large-Neighborhood Search (LNS) heuristic for complex real-world applications. The heuristic matches exact optima in $95.5\%$ of certified cases and reduces weighted posterior variance by $16.1$--$19.8\%$ relative to a sequential baseline on larger instances of up to 12 robots, providing a robust framework for cooperative physical-sampling missions.
Chinese Translation
野外机器人任务通常需要将物理样本送回实验室进行分析,这使得路径规划天然具有负载感知性和顺序依赖性,因为累积的样本会增加有效载荷和行进能耗。在单机器人的负载感知信息路径规划(Load-Aware Informative Path Planning, LIPP)中,感知与运输被刚性地耦合在一起:单个机器人必须独自运送其采集的所有样本,被迫频繁返回补给站,严重限制了其空间覆盖能力。异构多机器人团队可以通过分工来克服这一瓶颈——让高精度采样机器人负责采集,而高容量载运机器人负责运输。然而,这在LIPP问题之上引入了复杂的协同挑战,即交接应在何时、何地、交接什么以及交接给谁。为解决这一紧密耦合的问题,我们提出了多智能体LIPP(Multi-Agent LIPP, MA-LIPP),使团队能够通过异步“死投”(dead drops)进行协作,允许一个机器人存放样本供另一机器人稍后取回,而无需同步会合。我们将MA-LIPP表述为一个精确的混合整数二次规划(Mixed-Integer Quadratic Program, MIQP),并针对复杂现实应用提出了可扩展的成对大邻域搜索(Pairwise Large-Neighborhood Search, LNS)启发式算法。该启发式算法在95.5%的经认证算例中达到精确最优解,并且在多达12个机器人的较大规模实例上,相较于顺序基线方法将加权后验方差降低了16.1%–19.8%,为协同物理采样任务提供了一个鲁棒的框架。
cs.RO / 27 / 2609.21178

OpenRoIS: A Community-Driven Open-Source Middleware Implementing the Robotic Interaction Service (RoIS) Framework for Physical Robots and Virtual Agents

OpenRoIS:一个社区驱动的开源中间件,为实体机器人和虚拟智能体实现机器人交互服务(RoIS)框架
Villalobos, Sebastian Carrera, Arellano, Christopher Nolan, Hitzmann, Arne, Brito, Edilson Morais, Utsumi, Akira, Horikawa, Yukiko, Miyashita, Takahiro, Hafi, Lotfi El
Abstract
Service applications for human-robot interaction are commonly written against the hardware-specific interfaces of one platform, so a change of hardware forces a rewrite of the application. The Robotic Interaction Service (RoIS) Framework 2.0, standardized by the Object Management Group (OMG), addresses this fragmentation by defining a platform-independent model in which Service Applications interact with Human-Robot Interaction (HRI) Engines through standardized interfaces and hardware-independent symbolic messages. A specification alone, however, does not provide the maintained implementation, Software Development Kits (SDKs), and adapters needed for practical adoption. This paper presents OpenRoIS, a community-driven open-source middleware providing a concrete implementation of the RoIS Framework 2.0. It takes the position that an openly developed, paradigm-neutral implementation is what carries the standard from specification to practice. OpenRoIS contributes a recursive engine architecture in which a single engine class realizes the main and sub HRI Engine roles, an internal five-method component contract distinct from the five external RoIS interfaces, a mapping of those interfaces onto JSON-RPC 2.0 over WebSocket, a single-source-of-truth type pipeline that generates three consistent language stacks, TypeScript and C# client SDKs that include web and Unity support, and a Python adapter SDK that includes ROS 2 support. Through the common RoIS interfaces, a Service Application can address physical robots and virtual agents over the internet. All source code, interface types, and documentation are released under the Apache-2.0 license and openly developed at https://openrois.org/.
Chinese Translation
人机交互服务应用通常针对某一平台的硬件特定接口编写,因此更换硬件就必须重写应用。由对象管理组织(OMG)标准化的机器人交互服务(RoIS)框架2.0通过定义一个平台无关模型来解决这种碎片化问题,在该模型中,服务应用通过标准化接口和硬件无关的符号化消息与人机交互(HRI)引擎进行交互。然而,仅有规范并不能提供实际采用所需的维护实现、软件开发工具包(SDK)和适配器。本文提出OpenRoIS,一个社区驱动的开源中间件,提供了RoIS框架2.0的具体实现。我们的立场是:一个开放开发的、范式中立的实现是将标准从规范推向实践的关键。OpenRoIS的贡献包括:一种递归引擎架构,即单个引擎类同时实现主HRI引擎和子HRI引擎角色;一个区别于五个外部RoIS接口的内部五方法组件契约;将这些接口映射到基于WebSocket的JSON-RPC 2.0;一个单一事实来源的类型管道,生成三套一致的语言栈;支持Web和Unity的TypeScript与C#客户端SDK;以及支持ROS 2的Python适配器SDK。通过统一的RoIS接口,服务应用可以通过互联网操控实体机器人和虚拟智能体。所有源代码、接口类型和文档均以Apache-2.0许可证发布,并在https://openrois.org/上开放开发。
cs.RO / 28 / 2609.21185

When to Waddle: A Comparative Study of Bipedal Torso-Stabilization on Low-Friction Surfaces

何时摇摆:低摩擦表面上双足躯干稳定化的比较研究
Oke, Naomi, Gu, Ben, Ortiz, George, Ashlyn, Stacy, Pride, Cordelia, Bergbreiter, Sarah, Johnson, Aaron M.
Abstract
Low-friction surfaces challenge bipedal locomotion by limiting the contact forces available during stepping. Inspired by penguin waddling, we investigate how lateral torso motion and center of mass (COM) placement affect locomotion as surface friction changes. Using a five-actuator biped, we compare an upright-gait strategy with a penguin-inspired torso-over-stance-leg strategy across multiple COM placements in simulation and hardware. In the 3-D simulator MuJoCo, we sweep through sinusoidal leg and hip actuation parameters across four friction coefficients mu = 0.1, 0.3, 0.5, 0.7. In simulation, torso-over-stance-leg motion produces more successful controllers and higher forward speeds at low friction, with the highest speed occurring for the high-COM configuration. Hardware experiments show the same low-friction speed trend: at mu=0.12, torso-over-stance-leg motion increases forward speed and reduces cost of transport at both tested COM ratios, and the higher COM also improves both measures. The high-COM penguin configuration is the fastest and most energy efficient while maintaining low sideways foot motion. At mu=0.45, the COM trend reverses: the lower-COM configurations are faster and more energy efficient, while gait strategy has little effect on forward speed but still changes sideways foot motion. These results show that the effects of lateral torso motion and COM placement depend on the available friction, and that forward speed, energy use, and slip-related foot motion can be modulated with a penguin-inspired torso motion on hardware.
Chinese Translation
低摩擦表面通过限制迈步时可用的接触力,对双足行走构成挑战。受企鹅摇摆步态的启发,我们研究了随着表面摩擦的变化,躯干的侧向运动和质心(COM)位置如何影响行走。我们使用一个具有五个执行器的双足机器人,在仿真和硬件实验中,对直立步态策略与受企鹅启发的躯干置于支撑腿上方的策略在多种质心位置下进行了比较。在三维仿真器 MuJoCo 中,我们在四个摩擦系数 mu = 0.1、0.3、0.5、0.7 下对腿部和髋部的正弦驱动参数进行了扫掠。仿真结果表明,在低摩擦条件下,躯干置于支撑腿上方的运动能产生更多成功的控制器和更高的前进速度,其中最高速度出现在高位质心构型中。硬件实验显示了相同的低摩擦速度趋势:在 mu=0.12 时,躯干置于支撑腿上方的运动在两个测试的质心比例下均提高了前进速度并降低了运输成本,且较高的质心也同时改善了两项指标。高位质心的企鹅构型在保持较低侧向足部运动的同时,速度最快且能效最高。在 mu=0.45 时,质心趋势发生逆转:低位质心构型速度更快且能效更高,而步态策略对前进速度的影响很小,但仍会改变侧向足部运动。这些结果表明,躯干侧向运动和质心位置的影响取决于可用的摩擦力,并且可以在硬件上通过企鹅启发的躯干运动来调节前进速度、能耗以及与打滑相关的足部运动。
cs.RO / 29 / 2609.21186

Robust Structureless Monocular Visual Inertial Initialization Exploiting Line Features and Vanishing Points

利用线特征与消失点的鲁棒无结构单目视觉惯性初始化方法
Choi, Junwan, Jo, Woongrae, Seo, Dong-Uk, Jeon, Jinwoo, Myung, Hyun
Abstract
Accurate initialization is essential for reliable visual-inertial odometry (VIO), but it is often ill-conditioned under degenerate motions. Existing methods typically require restrictive excitation motions to ensure sufficient observability or rely on computationally expensive 3D structure reconstruction, limiting efficient and practical deployment. To address these limitations, we propose SLIM-init, a structureless monocular VIO initializer that directly exploits geometric constraints from tracked 2D line features without explicit 3D landmark reconstruction. Specifically, SLIM-init leverages line-derived vanishing points (VPs) as translation-invariant orientation cues to provide robust rotation-only constraints under degenerate scenarios such as low-parallax or translation-dominant motions. It further incorporates a line epipolar residual to constrain translation and a line-normal projection residual to improve the conditioning of linear alignment, enhancing the accuracy and robustness of initial state estimation. Extensive experiments on a public benchmark and challenging custom degenerate-motion sequences demonstrate improved accuracy and robustness over state-of-the-art initialization methods. The source code is available at: https://github.com/cjunwan/SLIM-init.
Chinese Translation
精确的初始化对于可靠视觉惯性里程计(VIO)至关重要,但在退化运动条件下,初始化往往呈现病态性。现有方法通常需要受限的激励运动以保证足够的可观测性,或依赖计算代价高昂的三维结构重建,限制了其高效实用的部署。为解决这些局限,我们提出了SLIM-init,一种无结构的单目VIO初始化器,它直接利用跟踪得到的二维线特征的几何约束,而无需显式的三维路标重建。具体而言,SLIM-init利用由线特征导出的消失点(VP)作为平移不变的姿态线索,在低视差或平移主导运动等退化场景下提供仅旋转的鲁棒约束。此外,该方法引入线极线残差来约束平移,并利用线法向投影残差改善线性对齐的条件数,从而提高初始状态估计的精度与鲁棒性。在公开基准数据集以及具有挑战性的自采退化运动序列上的大量实验表明,本方法相较最先进的初始化方法具有更高的精度和鲁棒性。源代码已发布于:https://github.com/cjunwan/SLIM-init。
cs.RO / 30 / 2609.21211

Stochastic Neural Signed Swept Volume for Real-time Chance-Constrained Trajectory Optimization

用于实时机会约束轨迹优化的随机神经符号扫掠体积
Chen, Qingyi, Zhang, Kevin, Chen, Lucas, Kingston, Zachary
Abstract
Collision-free motion planning requires reliable collision models from sensed environments and validation of states along a continuous trajectory. To make this tractable, most planners check for collision at discrete states along continuous trajectories against a single determinized model of the environment, introducing a trade-off between safety and computational efficiency. While continuous collision checking approaches that approximate the swept volume of the robot exist, they are computationally expensive or overly conservative. Data-driven approaches can learn the swept volume; however, these neural models are susceptible to approximation errors and are therefore often limited to serving as coarse filters for downstream collision checkers. In this work, we propose to learn a signed distance function of the swept volume as a probabilistic field, enabling quantification of epistemic uncertainty, incorporation of perception noise, and eventual integration into a chance-constrained trajectory optimization framework. We demonstrate our approach on challenging high-dimensional manipulation problems with significant sensor noise, both in simulation and on real hardware.
Chinese Translation
无碰撞运动规划需要从感知环境中获得可靠的碰撞模型,并对连续轨迹上的状态进行验证。为了使这一问题易于处理,大多数规划器在连续轨迹的离散状态点上针对环境的单一确定性模型进行碰撞检测,这在安全性与计算效率之间引入了权衡。尽管已有逼近机器人扫掠体积的连续碰撞检测方法,但其计算代价高昂或过于保守。数据驱动方法可以学习扫掠体积;然而,这些神经模型容易受到近似误差的影响,因此通常仅限于作为下游碰撞检测器的粗略过滤器。在本工作中,我们提出将扫掠体积的符号距离函数学习为一个概率场,从而能够量化认知不确定性、纳入感知噪声,并最终集成到机会约束轨迹优化框架中。我们在具有显著传感器噪声的高难度高维操作问题上验证了我们的方法,包括仿真和真实硬件实验。
cs.RO / 31 / 2609.21212

Visual Navigation Transformer with Pose Attention

具有位姿注意力机制的视觉导航Transformer
Li, Beiming, Romero, Jaime, Diller, Jonathan, Kumar, Vijay, Ribeiro, Alejandro
Abstract
Learned navigation policies typically consume observations as a temporally ordered history, with positional encodings tying each observation to when it was seen, making it difficult to reuse experience from earlier traversals of an environment. Systems that do reuse such experience usually construct an explicit representation, such as a map or a topological graph, and plan on it. We propose VNT-PA (Visual Navigation Transformer with Pose Attention), a transformer planner whose context is a set of depth keyframes indexed by camera pose. With camera poses as positional encoding, attention depends on the pose differences between keyframes rather than on their temporal order. VNT-PA is trained to imitate a shortest-path planner operating on the ground-truth scene mesh, predicting actions by querying the spatial context with only its current pose and the goal position. On point-goal navigation in HM3D validation scenes, VNT-PA reaches 93.3% success and 90.4% success weighted by path length (SPL), outperforming baselines that encode the same context as a temporal sequence or treat pose as an input feature, in both navigation performance and training efficiency. Because the spatial context is a pose-indexed set, frames from different trajectories can be fused at test time. The planner also degrades more gracefully under localization noise than a conventional baseline which plans on explicit maps. These results show that pose-stamped experience can serve directly as the environment representation for a learned planner, and that making attention depend on pose differences, rather than on temporal order, speeds up training and improves long-horizon navigation.
Chinese Translation
学习型导航策略通常将观测作为按时间排序的历史序列进行处理,位置编码将每个观测与其被观测的时间绑定,这使得难以复用环境中较早探索轨迹的经验。能够复用此类经验的系统通常会构建显式表示(如地图或拓扑图)并在此基础上进行规划。我们提出VNT-PA(Visual Navigation Transformer with Pose Attention,具有位姿注意力机制的视觉导航Transformer),这是一个以由相机位姿索引的深度关键帧集合为上下文的Transformer规划器。由于采用相机位姿作为位置编码,注意力机制取决于关键帧之间的位姿差异而非其时间顺序。VNT-PA的训练目标是在真值场景网格上模拟最短路径规划器,仅通过查询当前位姿和目标位置即可从空间上下文中预测动作。在HM3D验证场景的点目标导航任务中,VNT-PA达到93.3%的成功率和90.4%的按路径长度加权的成功率(SPL),在导航性能和训练效率上均优于将相同上下文编码为时间序列或将位姿作为输入特征的基线方法。由于空间上下文是一个由位姿索引的集合,来自不同轨迹的帧可以在测试时进行融合。与基于显式地图规划的常规基线相比,该规划器在定位噪声下也更平稳地退化。这些结果表明,带位姿标记的经验可以直接作为学习型规划器的环境表示,且使注意力依赖于位姿差异而非时间顺序,能够加快训练速度并提升长程导航性能。
cs.RO / 32 / 2609.21216

Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies

更少步数、更优动作:重新思考VLA策略的流匹配推理
Tang, Zhipeng, Chen, Xinda, Rao, Weining, Li, Xiao, Tan, Wenting, Wang, Yuning, Shi, Xiao, Zhao, Xiaofang
Abstract
Vision-language-action (VLA) policies based on flow matching generate action chunks through repeated evaluations of an action expert. Increasing the number of integration steps raises inference cost, but does not necessarily improve closed-loop success. We propose Coda, which reallocates part of this integration budget to a single learned endpoint correction. A frozen policy first completes a few-step noise-to-action trajectory; a lightweight Transformer then predicts a demonstration-supervised residual using the candidate action, source noise, and shared observation-prefix cache. Only the corrector is trained. On 50 RoboTwin Easy tasks, five-step Coda improves success from 71.64% to 74.68% over the matched five-step baseline, while reducing forward latency by 30.2% relative to the default ten-step policy. A two-step configuration achieves 71.88% success with a 2.12$\times$ speedup. An independent 13-task control shows a 5.69-percentage-point gain at nearly equal latency, supporting correction as an effective alternative to additional integration. The same design also improves frozen official SmolVLA, raising two-step success from 60.8% to 69.4%. These results show that endpoint correction improves the quality-latency trade-off of frozen flow-matching policies.
Chinese Translation
基于流匹配(flow matching)的视觉-语言-动作(VLA)策略通过对动作专家的多次评估来生成动作块。增加积分步数会提高推理成本,但并不一定能提升闭环成功率。我们提出Coda,将部分积分预算重新分配给一次学习得到的端点校正。冻结的策略首先完成一个少步数的从噪声到动作的轨迹;随后一个轻量级Transformer利用候选动作、源噪声以及共享的观测前缀缓存,预测一个由演示数据监督的残差。仅需训练该校正器。在50个RoboTwin Easy任务上,五步Coda相对于匹配的五步基线将成功率从71.64%提升至74.68%,同时相比默认的十步策略将前向延迟降低30.2%。两步配置在2.12倍加速的情况下达到71.88%的成功率。在一项独立的13任务对照实验中,在几乎相同的延迟下取得5.69个百分点的提升,表明端点校正是替代额外积分步数的有效方案。相同的设计也改进了冻结的官方SmolVLA,将两步成功率从60.8%提升至69.4%。这些结果表明,端点校正能改善冻结的流匹配策略的质量-延迟权衡。
cs.RO / 33 / 2609.21220

Safe Real-Time Policy Steering via Noise-Space Trajectory Optimization for One-Step Generative Policies

基于噪声空间轨迹优化的单步生成策略安全实时策略调控
Chen, Qingyi, Ruan, Joseph, Kingston, Zachary
Abstract
Generative robot policies can represent diverse, multimodal behaviors, but adapting pretrained policies to deployment-time constraints such as collision avoidance and orientation maintenance remains challenging. Existing inference-time steering methods typically apply gradient guidance through iterative diffusion or flow processes, which can be computationally expensive for real-time control. We propose INSPO, which formulates inference-time steering of one-step generative policies as trajectory optimization in the policy's input noise space. By optimizing the input noise while evaluating constraints on the induced state trajectory, INSPO searches the policy-induced behavior space without directly modifying generated actions. The optimization includes a regularization term that encourages solutions to remain consistent with the policy's input distribution and is solved online using population-based particle optimization. We evaluate INSPO on state- and image-based task-specific policies and generalist vision-language-action policies across Push-T, Can pick-and-place, and LIBERO-Spatial. INSPO improves task success and constraint satisfaction over best-of-N sampling and action projection, while comparing favorably with gradient-guided generation at lower runtime.
Chinese Translation
生成式机器人策略能够表达多样、多模态的行为,但将预训练策略适配到部署时的约束条件(如避碰与姿态保持)仍然具有挑战性。现有的推理时调控方法通常通过迭代的扩散或流过程施加梯度引导,这在实时控制中计算开销较大。我们提出INSPO,将单步生成策略的推理时调控表述为在策略输入噪声空间中的轨迹优化问题。通过优化输入噪声并对诱导状态轨迹施加约束评估,INSPO在策略诱导的行为空间中进行搜索,而无需直接修改生成的动作。该优化包含一个正则化项,鼓励解保持与策略输入分布的一致性,并使用基于种群的粒子优化方法在线求解。我们在基于状态和图像的任务专用策略以及通用视觉-语言-动作策略上,于Push-T、Can pick-and-place和LIBERO-Spatial任务中对INSPO进行了评估。结果表明,与最优N选一采样和动作投影相比,INSPO提升了任务成功率和约束满足度,同时在更低运行时间下与梯度引导生成方法相比具有竞争力。
cs.RO / 34 / 2609.21223

SafeStage: Evaluating Safety Before, During, and After Vision-Language-Conditioned Robot Manipulation

SafeStage:在视觉-语言条件化机器人操作之前、期间和之后评估安全性
Luo, Jinzhu, Zhang, Qi, Wang, Wei, Jiang, Wei
Abstract
Vision-language-conditioned robot policies integrate perception, language understanding, and control for general-purpose manipulation. However, existing evaluations often focus on task success, isolated physical constraints, semantic refusal, or realized physical damage, providing limited insight into where safety fails during closed-loop manipulation. We introduce SafeStage, a lifecycle-structured benchmark for evaluating manipulation safety before, during, and after task execution. SafeStage contains 97 purpose-built risk scenarios organized into three stages. Initial-State Hazards captures safety-relevant relations that must be resolved before manipulating the target. Execution-Time Safety evaluates unsafe contacts, trajectories, region entries, and object interactions during execution. Final-State Hazards capture unstable or otherwise unsafe conditions remaining after nominal task completion. The benchmark evaluates realized interactions using event-based and state-based checks and reports native task success independently from stage-specific safety outcomes. We evaluate representative direct-action Vision-Language-Action (VLA) policies and policies with world-model-based policies under a common closed-loop protocol. Our results demonstrate that nominal task completion frequently coexists with safety violations and that different policies exhibit distinct failure profiles across the three stages. By separating task success from safety and localizing when violations occur, SafeStage provides a unified diagnostic testbed for evaluating and improving vision-language-conditioned robot manipulation policies.
Chinese Translation
视觉-语言条件化的机器人策略将感知、语言理解与控制相结合,以实现通用操作。然而,现有评估往往只关注任务成功率、孤立的物理约束、语义拒绝或已造成的物理损害,对于闭环操作过程中安全在何处失效缺乏深入洞察。我们提出SafeStage,一个按生命周期结构设计的基准,用于在任务执行之前、期间和之后评估操作安全性。SafeStage包含97个专门构建的风险场景,分为三个阶段。初始状态危害捕获在操作目标之前必须解决的安全相关关系;执行时安全评估执行过程中的不安全接触、轨迹、区域进入和物体交互;最终状态危害捕获名义任务完成后仍存在的不稳定或其他不安全状态。该基准使用基于事件和基于状态的检查来评估实际发生的交互,并独立于各阶段的安全性结果单独报告原生任务成功率。我们在统一的闭环协议下评估了代表性的直接动作视觉-语言-动作(VLA)策略以及基于世界模型的策略。结果表明,名义任务的成功完成常常与安全违规并存,且不同策略在三个阶段中表现出不同的失败特征。通过将任务成功与安全性分离并定位违规发生的时间点,SafeStage为评估和改进视觉-语言条件化的机器人操作策略提供了一个统一的诊断测试平台。
cs.RO / 35 / 2609.21226

AirSplan: Risk-Aware Motion Planning for Quadrotors in Cluttered 3D Gaussian Splats

AirSplan:面向复杂3D高斯溅射场景中四旋翼的风险感知运动规划
Isaacson, Seth, Hong, William, Skinner, Katherine A., Vasudevan, Ram
Abstract
Quadrotors are increasingly deployed in applications such as agriculture, infrastructure inspection, and maintenance. In each of these applications, the robot must navigate complex scene geometry while remaining strictly collision-free. Unlike in ground domains, even minor collisions for aerial vehicles can result in the loss of the robot. This safety requirement induces a pair of technical challenges. First, the environment must be represented with sufficient fidelity to encode complex structure, even when no ground-truth obstacle data is available. Second, a motion planner must leverage this representation to determine a collision-free path to the goal. This paper proposes a system that addresses these complementary challenges. The proposed method, AirSplan, adopts a normalized variant of 3D Gaussian Splatting that encodes high-fidelity scene geometry. It then applies a novel reachability-based motion planner that leverages the differential flatness of quadrotors to compute continuous-time collision constraints that tightly overapproximate the robot's occupancy. Experiments demonstrate that AirSplan successfully finds a path in 81.2% of challenging test cases, a significant improvement over the nearest baseline method's 51.2%.
Chinese Translation
四旋翼无人机在农业、基础设施巡检与维护等领域的应用日益广泛。在上述所有应用中,机器人必须在复杂的场景几何环境中导航,同时严格保证无碰撞。与地面场景不同,空中飞行器即使发生轻微碰撞也可能导致机器人损失。这一安全性要求带来了两项技术挑战:首先,即使没有真实的障碍物数据,环境也必须以足够的保真度进行表示,以编码复杂的结构;其次,运动规划器必须利用该表示来确定到达目标的无碰撞路径。本文提出一个解决这些互补性挑战的系统。所提出的方法AirSplan采用了一种归一化变体的3D高斯溅射(3D Gaussian Splatting)来编码高保真度的场景几何。随后,该方法应用了一种新颖的基于可达性的运动规划器,利用四旋翼的微分平坦性(differential flatness)来计算连续时间碰撞约束,从而对机器人的占用空间进行紧密的过近似。实验表明,AirSplan在81.2%的具有挑战性的测试用例中成功找到路径,相较于最接近的基线方法的51.2%有显著提升。
cs.RO / 36 / 2609.21228

FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models

FOCAL-VLA:面向视觉-语言-动作模型的子任务引导几何蒸馏与隐式世界建模
Gao, Zhiyuan, Wen, Di, Zhan, Yanxiang, Khoshnazar, Mohammad, Schäfer, Jeroen, Peng, Kunyu, Beetz, Michael
Abstract
Vision-language-action (VLA) models built on pretrained vision-language models have demonstrated strong performance across diverse robotic manipulation tasks. However, VLA models that directly map current 2D observations to actions often lack sufficient spatial and temporal understanding, limiting their performance in precise and long-horizon manipulation. Recent methods enhance VLA models through geometric supervision and future-state prediction across the entire scene. However, these methods can suffer from redundant scene information, distracting the model from learning the geometry and dynamics relevant to the current interaction. To address this issue, we propose FOCAL-VLA, a framework that combines subtask-guided geometry distillation with implicit world modeling to learn representations of current spatial structure and future interaction dynamics. To focus geometric learning on the current subtask, we transfer geometric knowledge from VGGT to the VLA model by aligning geometry latents with features from subtask-relevant image regions. To capture the future 3D evolution of the current interaction, we incorporate implicit world modeling using Track4World features from current and future demonstration frames. The two complementary representations jointly guide action generation without running VGGT or Track4World at inference time. Experiments show that FOCAL-VLA outperforms baselines on both simulation benchmarks and real-world manipulation tasks. Project website: https://zhiyuan-gao.github.io/FOCAL-VLA/.
Chinese Translation
基于预训练视觉-语言模型构建的视觉-语言-动作(VLA)模型在多种机器人操作任务中展现出强大的性能。然而,直接将当前二维观测映射为动作的VLA模型往往缺乏充分的空间与时间理解能力,这限制了其在精确操作和长时程操作任务中的表现。近期方法通过整个场景的几何监督和未来状态预测来增强VLA模型,但这些方法可能受到冗余场景信息的干扰,使模型难以学习与当前交互相关的几何与动态信息。为解决这一问题,我们提出FOCAL-VLA,一个将子任务引导的几何蒸馏与隐式世界建模相结合的框架,用于学习当前空间结构和未来交互动态的表征。为了使几何学习聚焦于当前子任务,我们通过将几何隐变量与子任务相关图像区域的特征进行对齐,将VGGT的几何知识迁移到VLA模型中。为了捕捉当前交互的未来三维演化,我们利用来自当前与未来示范帧的Track4World特征引入隐式世界建模。这两种互补的表征共同引导动作生成,且在推理时无需运行VGGT或Track4World。实验表明,FOCAL-VLA在仿真基准和真实世界操作任务上均优于基线方法。项目网站:https://zhiyuan-gao.github.io/FOCAL-VLA/。
cs.RO / 37 / 2609.21229

KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos

KnowDemo:基于知识引导的从人类视频生成机器人示教
Gao, Zhiyuan, Zhan, Yanxiang, Khoshnazar, Mohammad, Schäfer, Jeroen, Beetz, Michael
Abstract
Learning robot manipulation policies typically requires substantial demonstration data, which are costly to collect on real robots. Recent methods generate robot demonstrations from human videos by adapting recovered motion and validating the resulting trajectories in simulation. However, methods centered on motion-reference adaptation can limit behavioral diversity by retaining the demonstrated contact strategies and subtask orders, while insufficient understanding of task requirements and scene relations can reduce demonstration generation efficiency by generating invalid candidates. To address these limitations, we propose KnowDemo, a framework that uses structured manipulation knowledge from human videos to generate diverse robot demonstrations for a target workspace. To distinguish task requirements from demonstration-specific choices, we develop a knowledge extraction and reasoning module based on a vision-language model (VLM) that associates object and action descriptions with inferred task conditions, demonstration references, and permissible execution variations. To translate this knowledge into executable demonstrations, we resolve the descriptions against target-scene entities and geometry to guide candidate generation and screening before motion planning and simulation. The resulting demonstrations exhibit multimodal behavior through alternative contact strategies and valid subtask orders, with structured execution labels. Experiments demonstrate additional verified execution modes beyond a reference-only configuration and improved candidate planning success through task-guided grasp sampling. To validate the generated data for policy learning, we fine-tune the pretrained $\pi_{0.5}$ model on simulation data, achieving sim-to-real transfer across three tasks. Project page: https://zhiyuan-gao.github.io/knowdemo/
Chinese Translation
学习机器人操作策略通常需要大量的示教数据,而在真实机器人上采集这些数据的成本很高。近期的方法通过适配从人类视频中恢复的运动并在仿真中验证生成的轨迹,来生成机器人示教。然而,以运动参考适配为核心的方法会保留示教中的接触策略和子任务顺序,从而限制了行为的多样性;而对任务需求和场景关系理解不足,则可能导致生成无效候选,降低示教生成效率。为解决这些局限,我们提出了KnowDemo,一个利用从人类视频中提取的结构化操作知识、为目标工作空间生成多样化机器人示教的框架。为区分任务需求与示教特定的选择,我们开发了基于视觉语言模型(VLM)的知识提取与推理模块,将物体和动作描述与推断出的任务条件、示教参考以及允许的执行变体相关联。为将该知识转化为可执行的示教,我们在运动规划和仿真之前,将描述与目标场景中的实体和几何信息进行匹配,以指导候选示教的生成与筛选。所生成的示教通过替代性接触策略和有效的子任务顺序展现出多模态行为,并带有结构化的执行标签。实验表明,相较于仅使用参考示教的配置,本方法获得了额外经过验证的执行模式,并通过任务引导的抓取采样提升了候选规划的成功率。为验证生成数据用于策略学习的有效性,我们在仿真数据上对预训练的 $\pi_{0.5}$ 模型进行微调,实现了跨三个任务的仿真到真实(sim-to-real)迁移。项目页面:https://zhiyuan-gao.github.io/knowdemo/
cs.RO / 38 / 2609.21246

VLA-Scope: Shift-Aware Failure Prediction for Vision-Language-Action Models

VLA-Scope:面向视觉-语言-动作模型的偏移感知失败预测
Zhu, Kaiwen, Liu, Dongfang, Liu, Liangkai
Abstract
Vision-language-action (VLA) models map visual observations and natural-language instructions to robotic actions, but distribution shifts can compromise their reliability. Because these models may still succeed under out-of-distribution (OOD) conditions, detecting OOD inputs alone is insufficient to predict execution failure. In this paper, we introduce VLA-Scope, a two-stage framework that combines input-shift characterization with execution history to predict failure during OOD rollouts. The first stage uses pooled image and language representations to detect OOD inputs and classify their shift categories. For inputs flagged as OOD, the second stage combines the predicted category, action-prefix features, and execution progress features. A logistic regression model shared across shift categories updates failure risk as execution proceeds. We evaluate the framework with OpenVLA on ten LIBERO-Spatial tasks using leave-one-group-out cross-validation. OOD detection achieves a ROC-AUC of 0.9454, and shift classification achieves 91% accuracy. Evaluated independently of the OOD gate on all 1,400 OOD rollouts, the failure predictor achieves a ROC-AUC of 0.8497 after 60 executed actions, compared with 0.7906 without execution progress features. It also achieves a higher ROC-AUC than the evaluated ActProbe and SAFE-MLP baselines. These results suggest that combining action features with temporally aggregated execution step representations improves failure prediction under input shifts.
Chinese Translation
视觉-语言-动作(VLA)模型将视觉观测和自然语言指令映射为机器人动作,但分布偏移可能损害其可靠性。由于这些模型在分布外(OOD)条件下仍可能成功执行任务,仅检测OOD输入不足以预测执行失败。本文提出VLA-Scope,一个两阶段框架,将输入偏移特征刻画与执行历史相结合,以预测OOD执行过程中的失败。第一阶段使用池化后的图像和语言表示来检测OOD输入并分类其偏移类别。对于被标记为OOD的输入,第二阶段将预测的类别、动作前缀特征和执行进度特征相结合。一个在各类偏移之间共享的逻辑回归模型随执行的推进不断更新失败风险。我们在十个LIBERO-Spatial任务上使用留一组交叉验证对OpenVLA评估了该框架。OOD检测的ROC-AUC达到0.9454,偏移分类准确率达到91%。在独立于OOD门控对所有1,400个OOD执行序列进行评估时,失败预测器在执行60个动作后的ROC-AUC达到0.8497,而去除执行进度特征时为0.7906。它还取得了比所评估的ActProbe和SAFE-MLP基线更高的ROC-AUC。这些结果表明,将动作特征与按时间聚合的执行步骤表示相结合,可以改善输入偏移下的失败预测。
cs.RO / 39 / 2609.21275

LOInK: Learned Optimal Inverse Kinematics via Structured Neural Surrogate Models

LOInK:基于结构化神经网络代理模型的学习型最优逆运动学方法
Somerfield, Michael, Abood, Damian, Wang, Ruigang, Manchester, Ian R.
Abstract
We introduce Learned Optimal Inverse Kinematics (LOInK), a method to generate approximately optimal solutions to an inverse kinematics problem. When trained on data consisting of sampled configurations and associated task variables and a given cost function, LOInK learns a bi-Lipschitz invertible mapping from configuration space to a decoupled task/latent space, and moreover, the latent space is structured so as to place cost-minimizing solutions at the origin. This enables efficient sampling of cost-minimizing solutions via a network-inversion algorithm based on operator splitting. We demonstrate the proposed approach on three problems: an illustrative three degree-of-freedom manipulator problem; a quadrupedal climbing robot for which LOInK can generate near-optimal solutions on average 31 times faster and up to 100 times faster than a constrained optimization approach; and a simulated soft actuator as a purely data-driven example, in which LOInK can explicitly generate high-quality solutions, unlike existing generative approaches that require diverse sampling and evaluation of candidate solutions.
Chinese Translation
我们提出了学习型最优逆运动学(Learned Optimal Inverse Kinematics,LOInK),这是一种为逆运动学问题生成近似最优解的方法。通过在由采样构型、相应任务变量以及给定代价函数组成的数据集上进行训练,LOInK 学习到从构型空间到一个解耦的任务/潜空间的双利普希茨(bi-Lipschitz)可逆映射,此外,该潜空间经过结构化设计,使得代价最小化的解位于原点处。这使得能够通过一种基于算子分裂的网络求逆算法高效地采样代价最小化的解。我们在三个问题上验证了所提出的方法:一个演示性的三自由度机械臂问题;一个四足攀爬机器人问题,LOInK 生成近优解的速度平均比约束优化方法快 31 倍,最高快 100 倍;以及一个模拟软体执行器的纯数据驱动示例,其中 LOInK 能够显式地生成高质量解,而现有的生成式方法则需要对候选解进行多样化采样与评估。
cs.RO / 40 / 2609.21307

Stability-aware Residual Reinforcement Learning Framework for Robotic Manipulator Disturbance Compensation

面向机械臂扰动补偿的稳定性感知残差强化学习框架
Kim, Jihong, Kwon, Joonhyuk, Kim, Hwa Soo, Seo, TaeWon, Seo, Hyung-Tae
Abstract
Although conventional controllers and disturbance observers (DOBs) are the standard for precision tracking in manipulators, they suffer from parameter uncertainty, nonlinear friction, and compound disturbances. This study proposes a residual reinforcement learning DOB framework that pairs an analytical observer with an RL policy. The deterministic baseline operates within a reliable region, whereas the RL policy explicitly targets the residuals that the model cannot capture. To make this compensation disturbance-aware, an estimator network aligns the observation history with a privileged disturbance context, organizing the latent space by disturbance regime and enabling rapid adaptation across disturbance transitions. To guarantee stability, we derived and enforced a state-dependent action bound on the RL policy from an input-to-state stability (ISS) analysis such that the closed loop provably confines the tracking error to a certified envelope for arbitrary policy outputs. Experiments on a 6-DOF manipulator demonstrated consistent improvements in disturbance estimation and tracking, including a 27.8% tracking-error reduction on real hardware under zero-shot sim-to-real transfer and a 38.0% reduction under a base-vibration disturbance that was not observed during training.
Chinese Translation
尽管传统控制器和扰动观测器(Disturbance Observers, DOBs)是机械臂高精度跟踪的标准方案,但它们在参数不确定性、非线性摩擦和复合扰动面前表现不佳。本研究提出了一种残差强化学习扰动观测器框架,将解析观测器与强化学习(RL)策略相结合。确定性的基线方法在可靠区域内运行,而RL策略则明确针对模型无法捕捉的残差。为使该补偿具备扰动感知能力,我们设计了一个估计器网络,将观测历史与特权扰动上下文对齐,按扰动状态组织潜在空间,从而实现跨扰动状态的快速自适应。为保证稳定性,我们基于输入到状态稳定性(input-to-state stability, ISS)分析推导并对RL策略施加了状态依赖的动作约束,使得闭环系统在任意策略输出下均可证明地将跟踪误差限制在一个可认证的包络范围内。在六自由度(6-DOF)机械臂上的实验表明,该框架在扰动估计和跟踪方面均取得了一致的改进,包括在零样本仿真到现实迁移下真实硬件上跟踪误差降低27.8%,以及在训练中未曾观测到的基座振动扰动下跟踪误差降低38.0%。
cs.RO / 41 / 2609.21316

NaViRrator: Robot Navigation from Human-Readable Maps through a Learned Visual Route

NaViRrator:基于可读地图与学习到的视觉路径的机器人导航
Lee, Ayun, Kim, Jiseon, Kim, Giseop
Abstract
Human-readable maps provide an intuitive interface for specifying robot destinations, but connecting their schematic geometry to egocentric observations remains challenging. We present NaViRRator, a framework that translates user-specified start and goal locations on such maps into navigation instructions for a pretrained vision-and-language navigation (VLN) policy. Its core method, RouteScribe, separates route inference from verbalization by first generating an explicit route scaffold in map-image coordinates, which a pretrained vision-language model (VLM) converts into a navigation instruction. We construct the scaffold with start--goal line conditional flow matching (SGL-CFM), which deforms a straight start--goal waypoint sequence into a map-conditioned route. During execution, the VLN policy receives only the instruction and egocentric observations, while the map and scaffold remain upstream, allowing executor replacement without retraining the map-to-language modules. Real-world experiments show higher success rates and success weighted by path length (SPL) than direct map-to-instruction generation, A*-based scaffolding, and Gaussian-source conditional flow matching. Qualitative results further show clearer salient turns and better preservation of the intended maneuver sequence, supporting route-grounded language as a modular interface between human-readable maps and pretrained navigation policies.
Chinese Translation
人类可读地图为指定机器人目标位置提供了直观的交互界面,但如何将其示意性的几何结构与机器人自我中心视角的观测相连接仍然具有挑战性。我们提出了NaViRrator,这是一个将用户在此类地图上指定的起点和目标位置转换为导航指令的框架,供预训练的视觉语言导航(VLN)策略使用。其核心方法RouteScribe将路径推断与语言表达相分离:首先在地图图像坐标系中生成显式的路径骨架(route scaffold),然后由预训练的视觉语言模型(VLM)将其转换为导航指令。我们采用起点-目标线条件流匹配(SGL-CFM)来构建路径骨架,即将一条直线的起点-目标航点序列形变为受地图条件约束的路径。在执行过程中,VLN策略仅接收指令和自我中心观测,而地图与路径骨架保留在上游,因此无需重新训练地图到语言的模块即可更换执行器。真实世界实验表明,与直接的地图到指令生成方法、基于A*的骨架构建方法以及高斯源条件流匹配方法相比,本方法取得了更高的成功率和按路径长度加权的成功率(SPL)。定性结果进一步表明,本方法能够更清晰地呈现显著的转弯,并更好地保留预期的操作序列,支持以路径为依据的语言作为人类可读地图与预训练导航策略之间的模块化接口。
cs.RO / 42 / 2609.21319

LEMCA: LLM-Guided Synthesis of Efficient Mode-Switching Control Architectures

LEMCA:基于大语言模型引导的高效模式切换控制架构合成
Krishna, Arjun, Pacelli, Vincent, Jayaraman, Dinesh
Abstract
Physical control tasks in the natural world, such as driving or object manipulation, frequently exhibit dramatic variations in sensory and compute complexity over time. Correspondingly, a natural resource-efficient choice for robot control is to dynamically switch between control modes with varying resource allocations. However, such "mode-switching controllers" (MSCs) have historically required laborious, expert-driven design and synthesis for each new task. Driven by these design difficulties, modern robotic control architectures often fall back to a wasteful "monolithic" one-size-fits-all structure, where resource allocation is permanently anchored to the hardest, most resource-intensive task phases. To facilitate the design of performant yet efficient MSCs, we propose LLM-Guided synthesis of Efficient Mode-Switching Control Architectures (LEMCA). LEMCA represents MSC designs as interpretable programs to be iteratively refined in an evolutionary loop. To evaluate design fitness, we propose MSC-compatible extensions of automated controller synthesis approaches, such as reinforcement learning in simulation. LEMCA then leverages the semantic priors, reasoning, and coding capabilities of Large Language Models (LLMs) to iteratively edit controller modes, their corresponding sensory-compute resource allocations, and mode transitions. Our experiments across diverse control benchmarks show that LEMCA consistently discovers strategies that surpass the Pareto frontier of monolithic designs by reclaiming wasted resources during "easy" task phases. LEMCA thus presents an automated, low-effort path to synthesize resource-efficient MSC designs.
Chinese Translation
自然界中的物理控制任务(如驾驶或物体操作)在感知与计算复杂度上常常随时间发生剧烈变化。相应地,一种自然的资源高效选择是让机器人控制在具有不同资源分配的控制模式之间动态切换。然而,此类“模式切换控制器”(MSC)历来需要针对每个新任务进行繁琐的、由专家驱动的设计与综合。受这些设计困难的驱动,现代机器人控制架构往往退而采用一种浪费资源的“单体式”(monolithic)一刀切结构,其中资源分配被永久性地锚定在最困难、最消耗资源的任务阶段。为便于设计高性能且高效的MSC,我们提出基于大语言模型引导的高效模式切换控制架构合成方法(LEMCA)。LEMCA将MSC设计表示为可解释的程序,并在进化循环中迭代优化。为评估设计的适应度,我们提出了与MSC兼容的自动化控制器综合方法的扩展,例如仿真中的强化学习。随后,LEMCA利用大语言模型(LLM)的语义先验、推理和代码编写能力,迭代地编辑控制模式、其对应的感知-计算资源分配以及模式转换。我们在多个多样化控制基准上的实验表明,LEMCA能够稳定地发现通过在“简单”任务阶段回收被浪费资源、从而超越单体式设计Pareto前沿的策略。因此,LEMCA提供了一条自动化的、低成本的路径来合成资源高效的MSC设计。
cs.RO / 43 / 2609.21330

The EventCV Library for Event-Based Robotic Vision

EventCV:面向事件驱动机器人视觉的函数库
Hines, Adam D., Milford, Michael, Fischer, Tobias
Abstract
Event cameras detect per-pixel brightness changes asynchronously on microsecond timescales, with high dynamic range and low power draw. These are desirable properties for robots that move fast or work in difficult lighting conditions. However, integrating an event camera into a real-world robotic pipeline still requires substantial effort: plug-and-play drivers do not exist, event streams are recorded in a variety of incompatible file formats, and most projects rely on custom research-grade code. Here, we present EventCV, an open-source and extensible Rust library with OpenCV-style Python bindings that lowers the entry barrier to working with event cameras. EventCV provides a wide range of features: denoising filters and geometric transforms, augmentations, corner detection and unsupervised feature learning, contrast-maximization motion estimation, a video-to-events simulator, and Open Neural Network Exchange (ONNX) inference for deployment in robotic stacks. EventCV integrates the Neuromorphic Drivers package, allowing an event camera stream to be processed directly in real time. No existing toolkit covers this range of operations in one package, and EventCV builds representations and decodes files 1.1x to 3.7x faster than the currently available libraries. We deploy EventCV on a Jetson Orin AGX and present three robotics case studies spanning object detection, on-device model inference, and localization. Project webpage: https://eventcv.net.
Chinese Translation
事件相机以微秒级时间尺度异步检测逐像素亮度变化,具有高动态范围和低功耗的特点。这些特性对于快速运动或在复杂光照条件下工作的机器人而言十分理想。然而,将事件相机集成到实际机器人系统中仍需大量工作:目前缺乏即插即用的驱动程序,事件流以多种互不兼容的文件格式记录,且大多数项目依赖定制的科研级代码。本文提出了EventCV,一个开源、可扩展的Rust函数库,并提供OpenCV风格的Python绑定,降低了使用事件相机的门槛。EventCV提供了丰富的功能:去噪滤波器与几何变换、数据增强、角点检测与无监督特征学习、对比度最大化的运动估计、视频到事件的模拟器,以及用于机器人系统部署的开放神经网络交换格式(ONNX)推理。EventCV集成了Neuromorphic Drivers软件包,可以直接实时处理事件相机数据流。目前尚无其他工具包能在单一软件包中涵盖如此广泛的操作,且EventCV构建表示和解码文件的速度比现有函数库快1.1至3.7倍。我们在Jetson Orin AGX上部署了EventCV,并展示了三个机器人学案例研究,涵盖目标检测、设备端模型推理和定位。项目主页:https://eventcv.net。
cs.RO / 44 / 2609.21358

FAN: Foresight Action Normalization for Continual Adaptation of Vision-Language-Action Models

FAN:用于视觉-语言-动作模型持续自适应的前瞻动作归一化方法
Hong, Yijun, Zhu, Jiarun, Sun, Xiaoquan, Xu, Le, He, Qijun, Jin, Xin, Yuan, Mingqi, Zeng, Wenjun, Chen, Jiayu
Abstract
Vision-Language-Action (VLA) models pre-trained on large-scale, closed datasets have demonstrated remarkable success across diverse robotic manipulation tasks. However, their long-term real-world deployment necessitates continuously acquiring new skills while retaining previously learned capabilities. While pioneering works have explored continual VLA adaptation using techniques such as experience replay and reinforcement fine-tuning, they overlook a foundational mechanism: action normalization, which determines the underlying coordinate system in which policies perceive and execute physical actions. To bridge this gap, we systematically evaluate five normalization strategies across four real-world task streams covering single-arm and bimanual manipulation. Our analysis reveals that existing protocols induce severe failure modes due to inter-task coordinate drift, limited motion coverage, or train-test coordinate mismatches. Motivated by these insights, we formulate three core design principles: consistency, coverage, and causality (3C), and introduce foresight action normalization (FAN). FAN estimates normalization statistics once from a small, task-independent calibration set prior to continual learning and freezes them throughout adaptation. Across all evaluated streams, FAN achieves the highest performance and demonstrates consistent robustness, providing insightful guidance for building stable action representations in achieving effective lifelong VLA adaptation.
Chinese Translation
在大规模封闭数据集上预训练的视觉-语言-动作(VLA)模型已在多种机器人操作任务中展现出卓越的成功。然而,其长期的真实世界部署要求模型在持续学习新技能的同时保留先前已学到的能力。尽管已有开创性工作探索了利用经验回放和强化微调等技术进行VLA的持续自适应,但它们忽视了一个基础机制:动作归一化(action normalization),它决定了策略感知和执行物理动作的底层坐标系。为弥补这一空白,我们在涵盖单臂和双臂操作的四个真实世界任务流上,系统评估了五种归一化策略。我们的分析揭示,现有方案由于任务间坐标漂移、运动覆盖范围有限或训练-测试坐标不匹配等原因,会引发严重的失败模式。基于这些洞察,我们提出了三个核心设计原则:一致性、覆盖性和因果性(3C),并引入了前瞻动作归一化(Foresight Action Normalization,FAN)。FAN在持续学习开始前,从一个较小的、与任务无关的校准集中一次性估计归一化统计量,并在整个自适应过程中保持冻结。在所有评估的任务流上,FAN均取得了最高性能并展现出一致的鲁棒性,为构建稳定的动作表示、实现有效的VLA终身自适应提供了具有启发性的指导。
cs.RO / 45 / 2609.21365

MicroHookACT: Monocular Microscopic Vision Guided Visuomotor Policy for Flexible Microelectrode Hooking

MicroHookACT:基于单目显微视觉引导的柔性微电极钩取视觉运动策略
Chen, Yitong, Qin, Fangbo, Wang, Yang, Hu, Ruihua, Zhang, Kui, Yu, Shan
Abstract
Automated needle-loop hooking is a critical step in flexible microelectrode (FME) implantation. This paper presents MicroHookACT, an imitation learning-based visuomotor policy for automated 3D hooking under monocular microscopic vision. First, a unidirectional hooking strategy exploits defocus cues and optical-axis guidance to enable palpation-free precise alignment and contact-rich threading. Second, an action-supervised object attention module built on a frozen ViT backbone learns to focus on the micro-needle tip and micro-loop directly from human demonstrations, without requiring manual visual annotations for training. Third, attention-centered global coarse and local fine features are dynamically weighted according to predicted action progress, enabling a single ACT policy to adapt to changing defocus blur and visual requirements throughout the operation. In the experiments, visuomotor policies were trained on 60 human demonstrations and evaluated under five setups with varying difficulties. Our MicroHookACT framework achieved the highest overall success rate of 96.7\% with an average execution time of 11.5 s. These results demonstrate the potential of visuomotor policy learning for micron-level control under varying operating conditions.
Chinese Translation
自动化针-环钩取(needle-loop hooking)是柔性微电极(FME)植入中的关键步骤。本文提出 MicroHookACT,一种基于模仿学习的视觉运动策略,可在单目显微视觉下实现自动化三维钩取。首先,一种单向钩取策略利用离焦线索与光轴引导,实现无需触诊的精确对准以及富含接触的穿线操作。其次,一个构建于冻结 ViT 骨干网络之上的动作监督目标注意力模块,能够直接从人类演示中学习聚焦于微针尖和微环,无需人工视觉标注进行训练。第三,以注意力为中心的全局粗粒度特征与局部细粒度特征根据预测的动作进度进行动态加权,使单一 ACT 策略能够适应整个操作过程中不断变化的离焦模糊和视觉需求。在实验中,视觉运动策略基于 60 条人类演示进行训练,并在五种不同难度的场景下进行评估。我们的 MicroHookACT 框架取得了最高的总体成功率 96.7%,平均执行时间为 11.5 秒。这些结果表明了视觉运动策略学习在多变操作条件下实现微米级控制的潜力。
cs.RO / 46 / 2609.21369

ProTracer: Proprioception-Guided Failure Diagnosis in Robot Manipulation

ProTracer:基于本体感觉引导的机器人操作失败诊断
Dong, Chang, Hosseinzadeh, Mehdi, Wong, King Hang, Liu, Lingqiao, Fraysse, Francois, Dayoub, Feras, Nguyen, Minh Hoai
Abstract
This paper presents a comprehensive framework for robot manipulation failure analysis that includes binary failure detection, failure categorization, explanation generation, and the additional capability of failure onset localization, which aims to identify the earliest moment at which a robot execution deviates from a valid task-completion trajectory and is ultimately followed by task failure. To address these tasks, we propose ProTracer, a training-free framework that leverages existing Vision-Language Models (VLMs) together with proprioceptive signals for failure analysis. Our method uses proprioceptive dynamics to identify temporally informative action boundaries and converts richer robot-state signals into structured natural-language descriptions that can be jointly analyzed together with visual observations by the VLM. This design combines the temporal precision of proprioceptive signals with the multimodal reasoning capabilities of modern VLMs without requiring additional model training. We further introduce FailTime, a benchmark with synchronized visual and proprioceptive observations for evaluating conventional failure diagnosis tasks as well as failure onset localization. Experiments demonstrate that ProTracer achieves strong performance across both conventional failure diagnosis tasks and the newly introduced failure onset localization task, highlighting the importance of proprioceptive reasoning for fine-grained temporal failure analysis.
Chinese Translation
本文提出了一个用于机器人操作失败分析的综合框架,涵盖二值失败检测、失败分类、解释生成,以及失败起始定位这一附加能力。失败起始定位旨在识别机器人执行过程最早偏离有效任务完成轨迹并最终导致任务失败的时刻。为解决上述任务,我们提出了 ProTracer,一个无需训练的框架,它利用现有的视觉-语言模型(Vision-Language Models, VLMs)并结合本体感觉信号进行失败分析。我们的方法利用本体感觉动态特征识别具有时间信息量的动作边界,并将更丰富的机器人状态信号转化为结构化的自然语言描述,使其能够与视觉观测一同由 VLM 进行联合分析。这一设计将本体感觉信号的时间精确性与现代 VLM 的多模态推理能力相结合,且无需额外的模型训练。我们进一步提出了 FailTime,一个包含同步视觉与本体感觉观测的基准数据集,用于评估传统失败诊断任务以及失败起始定位任务。实验表明,ProTracer 在传统失败诊断任务和新提出的失败起始定位任务上均取得了优异性能,凸显了本体感觉推理对于细粒度时间失败分析的重要性。
cs.RO / 47 / 2609.21377

AVT-Fabric: Active Visuo-Tactile Perception via Adaptive Evidence Selection for Efficient Robotic Fabric Comparison

AVT-Fabric:通过自适应证据选择实现高效机器人织物比较的主动视觉-触觉感知
Gao, Chang, Chen, Zhuo, Xia, Suhang, Zhu, Jihong, Deng, Jiankang, Luo, Shan
Abstract
Robotic fabric comparison needs to actively combine visual appearance and tactile cues. Here, we present AVT-Fabric, an RGB-first framework that allocates tactile evidence according to the difficulty of each comparison. A dual-scale gate evaluates answer-token confidence and raw logit separation to determine whether another force-tagged GelSight observation is needed. Compact textual memory preserves the executed history, and majority voting consolidates the selected predictions. On 400 held-out comparisons, AVT-Fabric achieves 98.0% accuracy with a compact 7B Multimodal Large Language Model (MLLM), surpassing the 94.0% reported by the 90B MLLM-Fabric baseline by 4.0 percentage points while processing only 1.60 of five available stages on average. It improves on matched passive inference by 9.25 percentage points and reduces model-side latency by 61.8%, while also improving on RGB-only accuracy. Four additional MLLM backbones support the generalizability, accuracy, and efficiency of the framework. This framework is also deployed on a real robotic system, achieving 78.1% pairwise ranking accuracy and correct fabric selection in seven of eight application scenarios. AVT-Fabric demonstrates that adaptive evidence selection can improve both the accuracy and efficiency of robotic visuo-tactile reasoning.
Chinese Translation
机器人织物比较需要主动地结合视觉外观与触觉线索。本文提出AVT-Fabric,这是一个以RGB优先的框架,根据每次比较的难度来分配触觉证据。一个双尺度门控机制通过评估答案token的置信度以及原始logit分离度,来判断是否需要额外一次带有力标记的GelSight观测。紧凑的文本记忆保留了已执行的历史记录,多数投票则对选定的预测进行整合。在400个保留比较任务上,AVT-Fabric使用紧凑的70亿参数多模态大语言模型(MLLM)达到了98.0%的准确率,比90B参数的MLLM-Fabric基线所报告的94.0%高出4.0个百分点,同时平均仅使用五个可用阶段中的1.60个。相对于匹配的被动推理,其准确率提升了9.25个百分点,模型端延迟降低了61.8%,同时准确率也优于仅使用RGB的方案。四个额外的MLLM骨干网络验证了该框架的泛化性、准确性与高效性。该框架还部署在真实的机器人系统上,实现了78.1%的两两排序准确率,并在八个应用场景中的七个里正确选择了织物。AVT-Fabric表明,自适应证据选择能够同时提升机器人视觉-触觉推理的准确性与效率。
cs.RO / 48 / 2609.21396

MarineCraft: Enabling Rapid Prototyping of Underwater Robots via Modular Construction

MarineCraft:基于模块化构造实现水下机器人的快速原型开发
Sugiura, Yuta
Abstract
Underwater robot development is often hindered by the complexities of waterproofing and wiring, which significantly delay the rapid prototyping process. This paper presents MarineCraft, a modular toolkit designed to accelerate the development cycle through structural reconfiguration. The system features self-contained, waterproof propulsion modules that integrate power, wireless communication, and actuation. By eliminating centralized wiring and the need for repeated sealing, MarineCraft allows diverse robot geometries to be assembled and tested in minutes rather than days. Experimental results demonstrate that this reconfigurable architecture enables fast, iterative design cycles while maintaining reliable operation and leak-free performance at depths of up to 2.5 meters. Our toolkit effectively lowers the barrier to underwater robotics by transforming modularity into a vehicle for rapid physical prototyping.
Chinese Translation
水下机器人的开发常常受到防水与布线复杂性的阻碍,这显著延缓了快速原型开发的进程。本文提出MarineCraft,一种旨在通过结构重构加速开发周期的模块化工具包。该系统采用自带电源、无线通信与驱动功能的独立防水推进模块。通过取消集中式布线和反复密封的需求,MarineCraft使多种机器人构型能够在数分钟内而非数天内完成组装与测试。实验结果表明,该可重构架构能够实现快速迭代的设计循环,同时在深达2.5米的水下保持可靠运行与无泄漏性能。我们的工具包将模块性转化为快速物理原型开发的载体,有效降低了水下机器人研发的门槛。
cs.RO / 49 / 2609.21404

Stabilizing Trajectory Outputs in End-to-End Autonomous Driving via SC-IMM Based Teacher Signals

基于SC-IMM教师信号的端到端自动驾驶轨迹输出稳定方法
Kim, Siewoo, Kong, Seung-Hyun
Abstract
End-to-End autonomous driving models commonly predict future waypoints from sensor inputs and convert them into vehicle control commands through a downstream controller. However, conventional waypoint-based imitation learning mainly minimizes coordinate-level errors, making it difficult to capture scene-dependent path-speed changes and temporal instability across waypoint outputs. In this paper, we propose an offline teacher-signal generation and learning method for trajectory-output stabilization based on a Scene-Conditioned Interacting Multiple Model (SC-IMM) to mitigate this issue. The proposed method converts expert trajectories into path-speed states and performs IMM updates conditioned on scene cues to generate path-speed teacher labels and mode posterior probabilities. The generated signals are added to the original trajectory loss as auxiliary supervision during training, while the inference structure and waypoint controller remain unchanged. In closed-loop evaluation on 100 short routes in CARLA Town12, the proposed method improved the driving score by 28.0% and reduced Collision/km by 62.3% compared with the baseline, while also improving jerk and trajectory-variation metrics. These results demonstrate that offline teacher signals embedding scene-conditioned motion-model cues can guide trajectory-output driving models toward more stable closed-loop behavior.
Chinese Translation
端到端自动驾驶模型通常从传感器输入预测未来路径点(waypoints),并通过下游控制器将其转换为车辆控制指令。然而,传统的基于路径点的模仿学习主要最小化坐标级误差,难以捕捉依赖场景的路径-速度变化以及路径点输出在时间上的不稳定性。本文提出一种基于场景条件交互多模型(Scene-Conditioned Interacting Multiple Model, SC-IMM)的离线教师信号生成与学习方法,以缓解该问题。所提方法将专家轨迹转换为路径-速度状态,并以场景线索为条件执行IMM更新,从而生成路径-速度教师标签和模式后验概率。在训练过程中,生成的信号作为辅助监督加入到原始轨迹损失中,同时推理结构和路径点控制器保持不变。在CARLA Town12的100条短路线闭环评测中,与基线相比,所提方法将驾驶得分提升了28.0%,将每公里碰撞(Collision/km)降低了62.3%,同时还改善了加加速度(jerk)和轨迹变化指标。这些结果表明,嵌入场景条件运动模型线索的离线教师信号能够引导轨迹输出型驾驶模型获得更稳定的闭环行为。
cs.RO / 50 / 2609.21416

A Unified Dynamic Force Guidance Framework for Performance-Optimized Kinesthetic Teaching

一种面向性能优化的动觉示教的统一动态力引导框架
Li, Chunxin, Wu, Jianhua, Xiong, Zhenhua, Zhu, Xiangyang
Abstract
Collaborative robots are increasingly deployed in industrial scenarios characterized by frequent product changeovers. As an intuitive programming method, kinesthetic teaching facilitates rapid robot deployment. However, users may overlook the configuration of the robot during kinesthetic teaching, leading to degradation in operational performance. Operational performance refers to the capability of the robot to generate motion and can be quantified by the Minimum Singular Value of the Jacobian matrix. To address this issue, this paper proposes an online dynamic force guidance method that integrates performance constraint and optimization mechanisms. Specifically, variable admittance control maintains the operational performance of the robot above a predefined threshold, while a virtual force actively guides the user to drag the robot towards configurations with improved performance. Experiments are conducted on a 6-DOF collaborative robot, comparing three typical paths in the task space. To evaluate the quality of the taught trajectories, trajectory playback experiments are conducted to analyze the relationship between the operational performance of the robot and the work efficiency. The results demonstrate that the proposed method effectively enhances the operational performance of the robot and consequently improves the work efficiency, holding significant value for reducing production takt time in industrial deployment.
Chinese Translation
协作机器人正日益频繁地部署于产品换线频繁的工业场景中。动觉示教作为一种直观的机器人编程方法,有助于机器人的快速部署。然而,用户在动觉示教过程中可能忽略机器人的构型配置,导致机器人操作性能下降。操作性能是指机器人生成运动的能力,可通过雅可比矩阵的最小奇异值进行量化。为解决这一问题,本文提出了一种融合性能约束与优化机制的在线动态力引导方法。具体而言,可变导纳控制将机器人的操作性能维持在预设阈值之上,同时虚拟力主动引导用户将机器人拖拽至性能更优的构型。实验在6自由度协作机器人上开展,对比了任务空间中的三条典型路径。为评估示教轨迹的质量,进行了轨迹回放实验,分析机器人操作性能与工作效率之间的关系。结果表明,所提方法有效提升了机器人的操作性能,进而提高了工作效率,对缩短工业部署中的生产节拍时间具有重要价值。
cs.RO / 51 / 2609.21447

FootQuery: Future-Touchdown-Guided Retrieval from Depth History for Perceptive Humanoid Locomotion

FootQuery:基于未来触地引导的深度历史检索,用于感知式人形机器人运动
Dong, Tao, Yu, Jia, Fan, Yuxuan, Zhao, Linna, Gong, Jiaqi, Yang, Andong, Gao, Chao, Zhou, Guyue
Abstract
Humanoid locomotion over complex terrain requires anticipating footholds that may no longer be visible at touchdown. Limited camera coverage and self-occlusion make it necessary to retrieve relevant terrain information from earlier observations. We present FootQuery, a perceptive locomotion framework that queries depth history using each foot's predicted next touchdown. The policy predicts touchdown locations and uncertainty from proprioception and uses these distributions, together with per-foot features, to query sparsely sampled historical depth frames. During training, realized contacts are projected into historical images to supervise retrieval at the regions where those contacts were visible. The retrieved per-foot features are fused with global visual memory to generate control actions. A progressive force-assistance curriculum supports early exploration, while event-consistent tread-midline shaping encourages coordinated stair contacts. Deployment requires only proprioception and onboard depth images. In simulation, the complete framework outperforms its component ablations on the most challenging tested stairs, gaps, and platforms. Real-world experiments on a Unitree G1 demonstrate continuous traversal with a single policy across outdoor stairs and indoor routes combining stair ascent and descent, platforms, and gaps. These results support organizing visual history around anticipated contacts for perceptive humanoid locomotion.
Chinese Translation
人形机器人在复杂地形上的运动需要预测可能在触地时已不可见的落脚点。有限的相机覆盖范围和自遮挡使得从早期观测中检索相关地形信息成为必要。我们提出了FootQuery,一种感知式运动框架,利用每只脚预测的下一次触地来查询深度历史。该策略根据本体感觉预测触地位置及其不确定性,并利用这些分布以及每只脚的特征来查询稀疏采样的历史深度帧。在训练过程中,将实际发生的接触投影到历史图像中,以在这些接触可见的区域监督检索。检索到的每脚特征与全局视觉记忆融合以生成控制动作。渐进式力辅助课程支持早期探索,而事件一致的踏步中线塑形则鼓励协调的楼梯接触。部署仅需本体感觉和机载深度图像。在仿真中,完整框架在最具挑战性的楼梯、沟壑和平台上均优于其组件消融版本。在Unitree G1上的真实世界实验表明,单一策略即可连续穿越室外楼梯以及结合上楼下楼、平台和沟壑的室内路线。这些结果支持围绕预期接触来组织视觉历史,以实现感知式人形机器人运动。
cs.RO / 52 / 2609.21448

Robotic Multiphase Interaction: Manipulating Coupled Liquid and Solid Dynamics with a World Model

机器人多相交互:利用世界模型操控耦合的液体与固体动力学
Feng, Yixuan, Wang, Peng
Abstract
This work presents \textit{Robotic Multiphase Interaction (RMI)}, a setting in which liquid enters a porous material and interacts mechanically with its deforming solid skeleton. Manipulation can therefore change pore volume, expel or redistribute retained liquid, and alter grasp stability at the same time. Spilled liquid can also create safety risks in domestic and manufacturing settings. This differs from most manipulation of solid objects and from tasks that involve both liquid and solid while keeping the phases spatially separate. We study a sponge filled with water as the first RMI example. We use implicit incompressible porous flow with smoothed particle hydrodynamics as the dynamics engine and enable robotic manipulation by adding Coulomb contact memory, hybrid velocity and force regulation, and a stability gate for lifting. The resulting environment connects robot commands to changes in the coupled liquid and solid state. A world model conditioned on actions predicts how this state evolves under candidate commands, while a temporal UNet generates actions using either Diffusion Policy or rectified flow matching. Our world model reduces retained water prediction error by more than $60\%$ compared with the baseline. The best action sequence selected by the world model from policy proposals further reduces the predicted terminal water error by about half. These improvements show that modelling the coupled liquid and solid state helps the robot predict how its actions affect both the porous object and the liquid held inside.
Chinese Translation
本工作提出了机器人多相交互(Robotic Multiphase Interaction, RMI)这一新设定:液体进入多孔材料并与发生形变的固体骨架产生力学交互。因此,操控操作可以同时改变孔隙体积、排出或重新分布内部保持的液体,并改变抓取的稳定性。溢出的液体在家庭和制造环境中还可能造成安全风险。这与大多数针对固体的操控任务不同,也不同于液体与固体在空间上保持分离的双相任务。我们以充满水的海绵作为第一个RMI研究实例。我们采用基于光滑粒子流体动力学(SPH)的隐式不可压缩多孔流作为动力学引擎,并通过引入库仑接触记忆、混合速度-力调节以及用于提起操作的稳定性门控,实现了机器人操控。由此构建的环境将机器人指令与耦合的液-固状态变化联系起来。一个以动作为条件的世界模型预测状态在候选指令下的演化,同时一个时间UNet网络采用Diffusion Policy或修正流匹配(rectified flow matching)生成动作。与基线相比,我们的世界模型将保持水量的预测误差降低了60%以上。由世界模型从策略提议中选出的最优动作序列,进一步将预测的终端水量误差降低约一半。这些改进表明,对耦合液-固状态进行建模有助于机器人预测其动作如何同时影响多孔物体及其内部保持的液体。
cs.RO / 53 / 2609.21461

AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining

AtomEgo:探索自我中心-机器人数据融合的具身基础模型预训练
Wu, Di, Zheng, Dongchen, Sheng, Junhe, Wei, Zhongxing, Zhang, Songxin, Xie, Zejian, Sun, Xiaoquan, Zheng, Junyang, Song, Zhuoyang, Zhang, Jiaxing, Chen, Jiayu
Abstract
Embodied foundation models are constrained by the limited scale and diversity of robot demonstrations, motivating the use of large-scale egocentric human interaction data. However, how to effectively incorporate such data into embodied-model pre-training remains unclear because of substantial embodiment and action-space gaps between humans and robots. We present AtomEgo, a systematic study of ego--robot co-training supported by a curated corpus of approximately 2,659 hours and a scalable data processing pipeline. Across vision--language--action and world--action model architectures, we investigate three representative paradigms: joint co-training with domain-specific action heads, progressive ego-to-robot transfer through embodiment alignment, and joint video--action modeling. We evaluate these paradigms through multi-task real-robot experiments and language-conditioned cross-embodiment representation analysis. Our results reveal a simple principle: Data Scale * Alignment Quality --> Capability Gain; egocentric data can improve generalization, but their value depends on how effectively they are aligned and utilized. This principle can provide practical guidance for scalable ego--robot pre-training.
Chinese Translation
具身基础模型受限于机器人演示数据的规模和多样性不足,这促使研究者利用大规模的第一人称视角(egocentric)人类交互数据。然而,由于人类与机器人之间存在巨大的本体(embodiment)和动作空间差异,如何有效地将此类数据融入具身模型预训练仍不明确。我们提出了 AtomEgo,这是一项对自我中心-机器人协同训练(ego–robot co-training)的系统性研究,其依托约 2,659 小时的精选语料库以及可扩展的数据处理流程。在视觉-语言-动作(vision–language–action)和世界-动作(world–action)两类模型架构上,我们研究了三种代表性范式:使用领域专属动作头的联合协同训练、通过本体对齐实现从自我中心数据到机器人数据的渐进式迁移,以及视频-动作联合建模。我们通过多任务真实机器人实验和基于语言条件的跨本体表征分析对这些范式进行评估。我们的结果揭示了一个简单原则:数据规模 × 对齐质量 --> 能力增益;第一人称数据可以提升泛化能力,但其价值取决于对齐和利用的有效程度。该原则可为可扩展的自我中心-机器人预训练提供实践指导。
cs.RO / 54 / 2609.21467

Learning Distance-Conditioned Object Transport for Humanoid Loco-Manipulation from a Single Motion Clip

基于单一运动片段学习距离条件化的人形机器人移动操作物体搬运
Hwang, Yuhyeon, Jung, Daniel Sungho, Seo, YongHyeok, Jung, Mingi, Cho, Chang Nho, Hwang, Jung-Hoon, Shin, Dongin
Abstract
Motion tracking can reproduce humanoid loco-manipulation from a single retargeted motion clip, but a policy trained on a fixed reference primarily reproduces its demonstrated transport outcome. Although the source trajectory visits intermediate object displacements, transport termination is demonstrated only at its endpoint. We identify this mismatch as the termination-versus-passage gap: intermediate displacements are observed as passage states rather than termination-complete outcomes. We introduce Distance-Conditioned Reference Recomposition (DCRR), which relocates the demonstrated termination segment to intermediate transport states. A frozen tracking teacher replays the recomposed references under closed-loop dynamics, and the retained trajectories are relabeled by their achieved object placements and distilled into a reference-free policy. This procedure constructs distance-conditioned supervision from the interaction behavior encoded in the source motion. Across Carry, Kick-Push, Crouch-Push, and Drag, DCRR-BC produces command-dependent transport with an overall normalized distance mean absolute error (MAE) of 0.15, compared with 0.28 for source-only behavior cloning. RL fine-tuning further improves the command response and execution robustness in the training simulator and under sim-to-sim transfer. Finally, hardware experiments demonstrate transport-distance modulation across all four interaction modes.
Chinese Translation
运动跟踪可以从单个重定向的运动片段中复现人形机器人的移动操作(loco-manipulation),但在固定参考上训练的策略主要只能复现其演示的搬运结果。尽管源轨迹经过了中间的物体位移状态,但搬运终止仅在轨迹终点处被演示。我们将这种不匹配定义为“终止-通过鸿沟”(termination-versus-passage gap):中间位移被观察为通过状态(passage states)而非终止完备的结果。我们提出距离条件化参考重组方法(Distance-Conditioned Reference Recomposition,DCRR),将演示的终止片段重新定位到中间搬运状态。冻结的运动跟踪教师模型在闭环动力学下重放重组后的参考,保留的轨迹根据其达到的物体位置进行重标注,并蒸馏为一个无需参考的策略。该过程从源运动中编码的交互行为构建距离条件化的监督信号。在搬运(Carry)、踢推(Kick-Push)、蹲推(Crouch-Push)和拖拽(Drag)四种任务中,DCRR-BC 实现了依赖指令的物体搬运,整体归一化距离平均绝对误差(MAE)为 0.15,而仅使用源数据的行为克隆为 0.28。强化学习微调进一步提升了指令响应能力和执行鲁棒性,在训练仿真器及仿真到仿真(sim-to-sim)迁移中均表现良好。最后,硬件实验验证了在全部四种交互模式下对搬运距离的调节能力。
cs.RO / 55 / 2609.21482

Adaptive Rollout Truncation Based on Epistemic Uncertainty for Efficient Offline World Model Training

基于认知不确定性自适应截断 rollout 的高效离线世界模型训练
Zymla, Nikodem Sebastian, Thiele, Laurin, Pitz, Johannes
Abstract
Accurate neural world models are central to model-based robotics, where they enable robots to predict future states from previously observed trajectories. Multi-step autoregressive training improves long-horizon prediction, but fixed rollout horizons also increase computational cost and can amplify early training errors when the model is still inaccurate. Existing training schemes typically use the same rollout length throughout optimization, independent of the model's current predictive reliability. We propose an epistemic uncertainty-driven adaptive rollout strategy for offline world model training following an auto-curriculum training scheme. Instead of always unrolling to a fixed horizon, the model terminates autoregressive rollouts once epistemic uncertainty exceeds a threshold calibrated from a warm-up phase. We study two uncertainty estimators: a five-head ensemble with a shared recurrent backbone and Monte Carlo Dropout. A two-stage warm-up procedure stabilizes uncertainty estimates before we enable adaptive truncation. Experiments on ANYmal-D and ANT show that ensemble-based adaptive truncation matches or improves the prediction accuracy of fixed-horizon training and the RWM-U baseline while requiring substantially fewer cumulative rollout steps. Training a world model on ANYmal-D following the presented approach reaches comparable final performance with the baselines with roughly 72% less rollout computation. These results indicate that epistemic uncertainty is useful not only for downstream policy regularization, but also for making world model training itself more compute-efficient.
Chinese Translation
精确的神经网络世界模型是基于模型的机器人技术的核心,它使机器人能够根据先前观测到的轨迹预测未来状态。多步自回归训练可以改善长时程预测,但固定的 rollout 时域也会增加计算成本,并且在模型尚不准确的训练早期可能放大误差。现有的训练方案通常在整个优化过程中使用相同的 rollout 长度,而与模型当前的预测可靠性无关。我们提出了一种由认知不确定性驱动的自适应 rollout 策略,用于遵循自动课程训练方案的离线世界模型训练。该模型并非总是展开到固定时域,而是在认知不确定性超过由预热阶段标定的阈值时终止自回归 rollout。我们研究了两种不确定性估计器:具有共享循环主干网络的多头集成方法和蒙特卡洛 Dropout。两阶段预热过程在启用自适应截断之前稳定不确定性估计。在 ANYmal-D 和 ANT 上的实验表明,基于集成方法的自适应截断在预测精度上达到或优于固定时域训练和 RWM-U 基线,同时所需累计 rollout 步数显著减少。按照所提出方法在 ANYmal-D 上训练世界模型,在与基线相当的最终性能下,rollout 计算量减少约 72%。这些结果表明,认知不确定性不仅可用于下游策略正则化,还可用于使世界模型训练本身更具计算效率。
cs.RO / 56 / 2609.21497

FORTE: Task-Adaptive Force Capability Optimization for Mobile Manipulators

FORTE:面向移动操作机器人的任务自适应力能力优化
Wang, Xiao, Zhang, Heng, Solak, Gokhan, Zhao, Fei, Ajoudani, Arash
Abstract
Effective physical interaction control in robotic manipulation requires not only kinematically feasible motion but also sufficient force-interaction capability. Existing redundancy resolution methods often ignore task-specific force demands or maximize the force capability indiscriminately, sacrificing dexterity when large force margins are unnecessary. We propose a task-oriented force capability optimization framework for redundant mobile manipulators. A Vision-Language Model (VLM) infers object physical properties from an RGB image and a task description, generating a desired task-force sequence that captures gravitational and inertial demands. We then define a task-oriented force capability metric as the signed distance between a task-force uncertainty ball and the dynamic residual force polytope (RFP), quantifying compatibility between task demands and the robot's remaining actuation capacity. This metric is incorporated, alongside manipulability, joint-limit avoidance, trajectory smoothness, and base-oscillation suppression, into a whole-body multi-objective trajectory-optimization problem. Experiments on a mobile manipulator performing lifting and single-point-holding tasks under varying payload conditions demonstrate that the proposed method provides sufficient force capability for heavy loads while preserving high manipulability for light loads. This yields a task-adaptive balance that fixed capability-maximizing baselines (RFP inscribed radius, RFP cone) and manipulability-only optimization fail to achieve. The core implementation is publicly available at https://github.com/yeying256/FORTE.
Chinese Translation
机器人操作中的有效物理交互控制不仅需要运动学上可行的运动,还需要足够的力交互能力。现有的冗余度解析方法往往忽略任务特定的力需求,或无差别地最大化力能力,在不需要大力裕度时牺牲了灵巧性。我们提出了一种面向任务的冗余移动操作机器人力能力优化框架。视觉-语言模型(Vision-Language Model, VLM)从RGB图像和任务描述中推断物体的物理属性,生成捕捉重力和惯性需求的期望任务力序列。随后,我们将面向任务的力能力度量定义为任务力不确定性球与动态剩余力多面体(Residual Force Polytope, RFP)之间的有符号距离,用以量化任务需求与机器人剩余驱动能力之间的兼容性。该度量与可操作度、关节限位规避、轨迹平滑性以及底盘振荡抑制一同被纳入全身多目标轨迹优化问题中。在移动操作机器人上于不同负载条件下执行提升和单点保持任务的实验表明,所提方法在重负载时能提供足够的力能力,同时在轻负载时保持较高的可操作度。这实现了固定能力最大化基线方法(RFP内切半径、RFP锥)和仅基于可操作度的优化所无法达到的任务自适应平衡。核心实现已在 https://github.com/yeying256/FORTE 公开。
cs.RO / 57 / 2609.21504

DPed-VLN: A Benchmark for Socially Compliant Vision-and-Language Navigation in Dynamic Pedestrian Environments

DPed-VLN:动态行人环境中 socially 兼容的视觉-语言导航基准
Dai, Haojie, Wang, Xiangyi, Wang, Liuyi, Sheng, Kai, He, Zongtao, Liu, Chengju, Ye, Wei, Chen, Qijun
Abstract
Vision-and-language navigation (VLN) has advanced rapidly in static indoor environments, but robots operating in human-populated spaces must ground language while responding to moving pedestrians and social-safety constraints. We present DPed-VLN, a Habitat 3.0 benchmark for dynamic-pedestrian VLN that couples 33,093 navigation episodes with paired global and prior-augmented instructions, ORCA-controlled humanoid pedestrians, socially constrained expert paths, and metrics that jointly assess navigation efficiency and social safety. DPed-VLN separates ordinary goal-oriented route guidance from prior-augmented instructions that expose dynamic-pedestrian cues for controlled analysis. To instantiate the benchmark, we introduce DPet (Dynamic Pedestrian-aware Network), a pedestrian-aware policy network trained with reinforcement learning and imitation learning. We further adapt representative state-of-the-art VLM-based navigation models, including NaVILA and StreamVLN, to DPed-VLN through LoRA fine-tuning. Experiments show that LoRA adaptation improves zero-shot VLM baselines in several success and safety metrics, especially reducing StreamVLN's collision rate. Among the evaluated methods, DPet-RL achieves the highest SR, SPL, and STL.
Chinese Translation
视觉-语言导航(Vision-and-Language Navigation,VLN)在静态室内环境中发展迅速,但在人流密集空间中工作的机器人必须将语言与环境中移动的行人及社交安全约束相结合进行理解。我们提出 DPed-VLN,一个基于 Habitat 3.0 的动态行人 VLN 基准,它将 33,093 个导航回合与配对的全局指令和先验增强指令、由 ORCA 控制的人形行人、受社交约束的专家路径,以及同时评估导航效率与社交安全的指标相结合。DPed-VLN 将普通的目标导向路径引导与揭示动态行人线索的先验增强指令区分开,以便进行受控分析。为构建该基准,我们提出了 DPet(Dynamic Pedestrian-aware Network,动态行人感知网络),一个通过强化学习和模仿学习训练的行人感知策略网络。我们还通过 LoRA 微调,将具有代表性的最先进的基于 VLM 的导航模型(包括 NaVILA 和 StreamVLN)适配到 DPed-VLN 上。实验表明,LoRA 适配在多项成功率和安全性指标上提升了零样本 VLM 基线,尤其显著降低了 StreamVLN 的碰撞率。在所评估的方法中,DPet-RL 取得了最高的 SR、SPL 和 STL。
cs.RO / 58 / 2609.21511

2nd Place Solution to the HANDS 2026 Workshop Challenge-Dexterous Grasp Motion Track: Single-Shot Trajectory Warping for Grasp Motion Generation

HANDS 2026 研讨会挑战赛(灵巧抓取运动赛道)亚军方案:面向抓取运动生成的单次轨迹变形方法
Khan, Muneeb A., Kim, Woojin, Kim, Shinwoo, Munsif, Muhammad, Bhattarai, Binod, Baek, Seungryul
Abstract
This report describes our 2nd place solution to the HANDS 2026 workshop challenge (Dexterous Grasp Motion track) in conjunction with ECCV 2026. In this challenge, we address grasp motion generation for the 12-DoF LinkerHand O6, aiming to produce physically plausible reach-and-lift trajectories for unseen objects from randomized initial hand poses in simulation. This task is particularly challenging because each grasp requires a per-step policy to make approximately $70$ twelve-dimensional decisions, with errors accumulating over time, while test objects and physical dynamics may differ from those encountered during training. To address these challenges, we propose editing a single successful GraspM3 demonstration instead of generating the motion step by step: a policy observes the object once and outputs a 12-D warp of the demonstration, which is then replayed open-loop. Moreover, we train the warp policy with one-step PPO over all $4{,}824$ training objects in parallel. As a result, our method achieved success rates of $94.61\%$ on the easy track, the highest of all submissions, and $57.18\%$ on the hard track of the private test set.
Chinese Translation
本报告介绍了我们在与 ECCV 2026 联合举办的 HANDS 2026 研讨会挑战赛(灵巧抓取运动赛道)中获得第二名的解决方案。在该挑战赛中,我们针对 12 自由度的 LinkerHand O6 灵巧手研究抓取运动生成问题,目标是在仿真环境中,从随机初始手部姿态出发,为未见过的物体生成物理上合理的接近并提起(reach-and-lift)轨迹。该任务极具挑战性,原因在于每次抓取需要逐步策略做出约 70 次 12 维决策,误差会随时间累积,同时测试物体和物理动力学可能与训练时遇到的不同。为应对这些挑战,我们提出对单个成功的 GraspM3 示范轨迹进行编辑,而非逐步生成运动:策略对物体进行一次观测后,输出示范轨迹的 12 维变形(warp),随后以开环方式回放该轨迹。此外,我们使用单步 PPO 在全部 4,824 个训练物体上并行训练变形策略。最终,我们的方法在私有测试集的简单赛道上取得了 94.61% 的成功率(所有提交方案中最高),在困难赛道上取得了 57.18% 的成功率。
cs.RO / 59 / 2609.21514

Skel-WAM: A Hand-Skeleton-Conditioned World Action Model for Human-to-Robot Manipulation Transfer

Skel-WAM:一种手部骨架条件化的世界动作模型,用于人到机器人操作技能迁移
Cai, Zetao, Li, Yaping, Wang, Yiqun, Zhan, Xinyu, Yang, Yuyin, Ma, Haoxiang, Li, Kailin, Lu, Tao, Pang, Jiangmiao, Xu, Linning, Lin, Dahua
Abstract
Robot demonstrations are expensive to collect and often provide limited distributional coverage of task variations. Human videos offer a low-cost source of complementary manipulation experience, but learning from them requires bridging embodiment gaps in visual appearance and action spaces. We introduce Skel-WAM, a world action model that bridges these differences through a unified hand-skeleton motion interface. The key insight is to align human and robot motion through a common hand topology, combining skeleton overlays that ground motion in the scene with structured 2.5-D keypoints that encode explicit hand kinematics. Video and Keypoint Experts jointly learn visual and skeletal dynamics through a Mixture-of-Transformers, while a separate robot-trained Action Expert maps these predictions to executable controls. This separation enables human and robot demonstrations to directly supervise shared dynamics without requiring robot action labels for human videos. Across four real-world bimanual tasks and seven simulated tasks, Skel-WAM achieves average success rates of 79.86% and 63.29%, surpassing the strongest baseline by 22.22 and 8.28 percentage points, respectively. Human-robot cotraining more than doubles real-world success on task variations absent from robot training data, from 38.89% to 86.11%. These results demonstrate that a shared skeletal interface enables joint learning across human and robot data and expands robot task coverage through complementary human demonstrations.
Chinese Translation
机器人演示数据的采集成本高昂,且对任务变化的分布覆盖往往有限。人类视频提供了一种低成本的补充性操作经验来源,但从中学习需要弥合视觉外观和动作空间上的本体差异。我们提出了 Skel-WAM,一种通过统一的手部骨架运动接口来弥合这些差异的世界动作模型。其关键洞察在于通过共同的手部拓扑结构对齐人类与机器人运动,将锚定场景中运动的骨架叠加与编码显式手部运动学的结构化2.5维关键点相结合。视频专家与关键点专家通过混合变换器共同学习视觉与骨架动力学,而一个单独基于机器人数据训练的动作专家则将这些预测映射为可执行的控制指令。这种分离设计使人类与机器人演示能够直接监督共享的动力学模型,而无需为人类视频提供机器人动作标签。在四项真实世界双臂任务和七项仿真任务上,Skel-WAM 分别取得了 79.86% 和 63.29% 的平均成功率,分别超越最强基线 22.22 和 8.28 个百分点。人机协同训练使机器人训练数据中未出现的任务变化在真实世界中的成功率提升了一倍以上,从 38.89% 提高到 86.11%。这些结果表明,共享的骨架接口能够实现人类与机器人数据的联合学习,并通过互补性的人类演示扩展机器人的任务覆盖范围。
cs.RO / 60 / 2609.21572

SABER: Learning Attention-based Semantic Affordance for Legged Locomotion

SABER:面向足式运动的学习型基于注意力的语义可供性方法
Palanivelu, Hari Prasanth, Sze, Samuel, Johannes, Kennard Garrison, Adiwahono, Albertus Hendrawan, Yee, Meng, Chuah
Abstract
Perceptive legged locomotion has advanced rapidly by integrating terrain geometry into learned policies, yet the integration of terrain meaning remains sparse: a pipe, a patch of grass, or a fragile box may be geometrically traversable while being inappropriate for contact. In industrial environments, where legged robots increasingly operate, a single misplaced step can damage fragile equipment, destabilize the robot, or endanger the site. To address this, we introduce SABER, a planner-free reinforcement-learning policy that jointly reasons about terrain geometry and semantic contact permission. The policy consumes a unified terrain-affordance map, where each cell encodes local 3D geometry and a semantic contact cost. We augment cross-attention with a learned, signed semantic bias: an additive term on the attention logits, gated by the contact cost, that reweights flagged cells by their distance from the nearest foot. A hazard therefore reshapes attention where it can still affect the next foothold, and its influence fades where it cannot. The resulting policy selects footholds on permitted support and keeps the leg clear of forbidden regions throughout the swing phase. We perform a systematic ablation that isolates the contribution of each architectural component; removing the semantic bias alone increases forbidden contacts by 55% while velocity tracking is unchanged. We validate the policy on a Unitree B2, demonstrating sim-to-real semantic contact selection across indoor and outdoor environments and four semantic obstacle classes.
Chinese Translation
感知足式运动通过将地形几何信息整合到学习策略中而迅速发展,然而对地形含义的整合仍然稀少:一根管道、一片草地或一个易碎的箱子在几何上或许可以通行,但并不适宜接触。在足式机器人日益应用的工业环境中,一步踏错就可能损坏易碎设备、破坏机器人稳定性或危及现场安全。为解决这一问题,我们提出了SABER,一种无需规划器的强化学习策略,能够联合推理地形几何与语义接触许可。该策略以统一的地形可供性地图(terrain-affordance map)作为输入,其中每个栅格编码局部三维几何信息和一个语义接触代价。我们在交叉注意力(cross-attention)中引入了一个可学习的带符号语义偏置:这是注意力logits上的一个附加项,由接触代价进行门控,并根据标记栅格与最近足端的距离对其重新加权。因此,危险区域只会在仍能影响下一个落足点的地方重塑注意力,而在无法影响的地方其作用则逐渐消退。最终策略在允许的支撑区域选择落足点,并在整个摆动相(swing phase)中使腿部避开禁止区域。我们进行了系统的消融实验以分离每个架构组件的贡献;仅移除语义偏置就会使禁止接触增加55%,而速度跟踪性能保持不变。我们在Unitree B2机器人上验证了该策略,展示了在室内外环境以及四类语义障碍物上从仿真到现实(sim-to-real)的语义接触选择能力。
cs.RO / 61 / 2609.21580

Tilt as a Certified Resource: Preserving Motor Wrench-Rate Authority on Articulated Multirotors

倾斜作为可认证资源:保持铰接式多旋翼飞行器的电机力旋量变化率调控能力
Silano, Giuseppe, Saska, Martin
Abstract
Fully-actuated multirotor aerial vehicles must not only track nominal wrenches but retain the "readiness" to modulate them rapidly under disturbances. Classical effort-minimizing allocators ignore this dynamic limit, whereas maximizing readiness leads to topologically disconnected optimal sheets demanding physically impossible actuator rates. Enforcing a readiness safety floor on fixed-geometry symmetric platforms further encounters a zero-sum degeneracy: motor-speed redistribution cannot improve authority without conceding wrench tracking. This paper uses active morphology to break the degeneracy, treating servo tilt as a geometric resource supplying authority-recovery directions unavailable to static rotors. We construct a configuration-dependent, motor-only readiness certificate - the log-volume of the reachable wrench-rate set - that explicitly excludes servo capacity, preventing a "ghost capacity fallacy" in which the certificate would falsely credit slow mechanical kinematic limits instead of collapsing accurately at motor saturation. The certificate is enforced as a Control Barrier Function (CBF) within a Unified Physical-Command Quadratic Program acting on motor torques and servo setpoints. Closed-loop simulations of an articulated octorotor under severe gust disturbances show classical allocators diverging and uncertified articulated allocators violating the safety floor, while the proposed CBF filter bounds the system state and preserves vehicle authority.
Chinese Translation
全驱动多旋翼飞行器不仅要跟踪标称力旋量(wrench),还必须保持在外部扰动下快速调节力旋量的“就绪度”(readiness)。经典的功最小化分配器忽略了这一动态限制,而最大化就绪度则会导致拓扑不连通的最优解面,要求物理上不可能实现的电机转速变化率。在固定几何构型的对称平台上施加就绪度安全下限,还会进一步遭遇零和退化问题:电机转速的再分配无法在不牺牲力旋量跟踪能力的情况下提升调控能力。本文利用主动形态变化来打破这一退化,将舵机倾斜视为一种几何资源,提供静态旋翼布局无法实现的调控能力恢复方向。我们构建了一个依赖于构型、仅由电机决定的就绪度证书——即可达力旋量变化率集合的对数体积——该证书明确排除了舵机能力,从而避免“幽灵能力谬误”,即在电机饱和时,证书会错误地将缓慢的机械运动学限制计入能力,而不是准确地崩溃失效。该证书作为控制障碍函数(Control Barrier Function, CBF),在作用于电机转矩和舵机设定值的统一物理指令二次规划(Unified Physical-Command Quadratic Program)中加以实施。对铰接式八旋翼飞行器在强阵风扰动下的闭环仿真表明:经典分配器会发散,未经认证的铰接式分配器会违反安全下限,而所提出的CBF滤波器能够约束系统状态并保持飞行器的调控能力。
cs.RO / 62 / 2609.21584

A High-Payload Wall-Climbing Robot Using Passive Bistable Suction Cups

一种采用被动双稳态吸盘的高负载爬壁机器人
Nguyen, Andrew, Li, Mingyuan, Bruder, Daniel
Abstract
Wall-climbing robots capable of scaling vertical surfaces could help automate hazardous or labor intensive tasks such as window washing, inspection, maintenance, and construction. Active adhesion methods achieve higher payload capacities, but require power to maintain their grip. Passive adhesion devices such as suction cups are an attractive option for such robots because they do not require power to maintain their grip, but they are limited by their payload capacity. This work presents a novel high-payload wall-climbing robot that utilizes passive bistable suction cups to generate adhesion without needing to be pushed into the wall. The robot features a track-based system that automatically engages and disengages bistable suction cups to achieve locomotion on smooth surfaces. The robot is able to achieve vertical wall climbing on glass, wood, metal, and painted surfaces, sideways and upside-down climbing, and is able to tow a payload of 7.940 kg (with a payload-to-weight ratio of 2.25).
Chinese Translation
能够攀爬垂直表面的爬壁机器人可帮助实现危险或劳动密集型任务的自动化,例如窗户清洁、检测、维护和建筑施工。主动吸附方法可实现更高的负载能力,但需要持续消耗能量以维持吸附力。吸盘等被动吸附装置无需消耗能量即可保持吸附,是此类机器人的理想选择,但其负载能力受限。本研究提出一种新型高负载爬壁机器人,利用被动双稳态吸盘产生吸附力,无需将其压向墙面。该机器人采用履带式系统,可自动接合与释放双稳态吸盘,从而在光滑表面上实现移动。该机器人能够在玻璃、木材、金属和涂漆表面实现垂直攀爬,并具备水平和倒立攀爬能力,可拖动7.940千克的负载(负载重量比为2.25)。
cs.RO / 63 / 2609.21609

Potential-Field Action Representation for Reinforcement Learning in Contact-Rich Manipulation

面向接触密集型操作的势场动作表示强化学习方法
Liu, Xinyu, Solak, Gökhan, Ajoudani, Arash
Abstract
Model-free reinforcement learning can acquire contact-rich robotic manipulation skills through trial-and-error interaction, but it often requires the policy to learn both task strategy and low-level motion generation. In this setting, the action representation is critical because it determines how policy outputs are converted into robot motion, shaping both exploration and physical execution. Direct Cartesian command interfaces require the policy to generate motion at every decision step, coupling task-level adaptation with continuous low-level control and increasing the learning burden. We propose PA-RL, a reinforcement-learning framework that uses artificial potential fields as the action representation. Instead of commanding motion directly, the policy adapts the parameters of an energy-like potential field, which generates a state-dependent guidance direction executed through a Cartesian impedance controller. We evaluate PA-RL on peg-in-hole insertion, a representative contact-rich task with nonlinear dynamics and discontinuous contact transitions. In simulation, PA-RL is compared with Cartesian velocity, Cartesian pose, and variable-impedance action spaces using the same RL algorithm. PA-RL is the only method to reach a 100% evaluation success rate within the allotted training time, while the best baseline reaches 92.6%. It also reduces joint-torque variation by 55.4% and Cartesian acceleration variation by 70.8% relative to the best baseline, without explicit motion-quality penalties in the reward. The simulation-trained policy further completes 9/9 real-robot insertions without fine-tuning, demonstrating the deployment feasibility of the learned potential-field interface.
Chinese Translation
无模型强化学习可以通过试错交互获取接触密集的机器人操作技能,但通常需要策略同时学习任务策略和底层运动生成。在这种设定下,动作表示至关重要,因为它决定了策略输出如何转化为机器人运动,同时影响探索和物理执行。直接的笛卡尔指令接口要求策略在每个决策步生成运动,将任务级适应与连续底层控制耦合在一起,增加了学习负担。我们提出PA-RL,一种使用人工势场(artificial potential field)作为动作表示的强化学习框架。策略不直接指令运动,而是调整一个类能量势场的参数,该势场生成状态相关的引导方向,并通过笛卡尔阻抗控制器执行。我们在轴孔装配(peg-in-hole insertion)任务上评估PA-RL,这是一个具有非线性动力学和不连续接触转换的代表性接触密集型任务。在仿真中,使用相同的强化学习算法,将PA-RL与笛卡尔速度、笛卡尔位姿和变阻抗动作空间进行比较。PA-RL是唯一在规定训练时间内达到100%评估成功率的方法,而最佳基线方法达到92.6%。此外,相对于最佳基线,PA-RL将关节力矩变化降低55.4%,将笛卡尔加速度变化降低70.8%,且无需在奖励中加入显式的运动质量惩罚。经仿真训练的策略无需微调即可在真实机器人上完成9/9次装配任务,证明了所学势场接口的部署可行性。
cs.RO / 64 / 2609.21617

CounterPlay: Counterfactual Post-Training for Self-Play Driving Policies

CounterPlay:面向自博弈驾驶策略的反事实后训练
Wei, Jiarong, Wu, Yin, He, Runkai, Valada, Abhinav
Abstract
Self-play in high-throughput simulators yields driving policies with robust closed-loop performance, but improvement per unit of simulation diminishes as training scales. Policies learn to handle common situations early, while further rollouts repeatedly encounter unresolved failures. Post-training offers an opportunity to target these failures, but existing methods primarily evaluate alternative actions or continuations at visited states, although successful recovery may require changing driving style earlier. We propose CounterPlay, a counterfactual self-play post-training approach that backtracks from failed tasks and retries them under alternate driving styles. CounterPlay rests on three key components. First, failure-driven backtracking uses the policy's value estimates to select an earlier stored state from which to retry the task. Second, reward conditioning enables a single policy to retry the task from this state using candidate styles ranging from cautious to aggressive. Third, CounterPlay retains task-completing retries only if no other vehicle incurs a new or earlier collision or off-road event relative to the factual branch. Retries that pass verification with fresh randomness are then distilled into the policy under its deployment condition. On BehaviorBench, CounterPlay achieves state-of-the-art scores on both the Interactive and Random splits across all eight traffic regimes using 1B post-training transitions, which is just 1% of the anchor's 100B self-play training budget. Improvements over the anchor hold across all three evaluated driving styles. CounterPlay resolves a substantial fraction of the anchor's timeout cases on BehaviorBench and achieves a balance between task completion and safety that neither continued self-play nor adopting a more aggressive driving style attains.
Chinese Translation
在高吞吐量模拟器中进行自博弈(self-play)可以获得具有鲁棒闭环性能的驾驶策略,但随着训练规模的扩大,单位模拟量的收益会逐渐递减。策略在早期学会了处理常见情形,而后续的仿真推演(rollout)则会反复遭遇尚未解决的失败。后训练为针对性解决这些失败提供了机会,但现有方法主要是在已访问的状态上评估替代动作或后续轨迹,而成功的恢复可能需要在更早的阶段改变驾驶风格。我们提出了CounterPlay,这是一种反事实自博弈后训练方法,它从失败任务回溯,并在不同的驾驶风格下重试这些任务。CounterPlay基于三个关键组件。第一,失败驱动的回溯利用策略的价值估计来选择一个更早的已存储状态,从中重试任务。第二,奖励条件化使单个策略能够从该状态出发,以从谨慎到激进的各种候选驾驶风格重试任务。第三,CounterPlay仅在重试过程中没有其他车辆相较于事实分支(factual branch)发生新的或更早的碰撞或驶出路面事件时,才保留完成任务的重试结果。通过新鲜随机性验证的重试随后会在其部署条件下被蒸馏进策略中。在BehaviorBench上,CounterPlay仅使用10亿(1B)条后训练转移数据,就在全部八种交通场景的Interactive和Random两种划分上均取得了最先进的分数,而这一数据量仅为基准模型(anchor)1000亿(100B)自博弈训练预算的1%。相对于基准模型的提升在所有三种被评估的驾驶风格上均保持成立。CounterPlay解决了基准模型在BehaviorBench上相当比例的超时案例,并实现了任务完成与安全之间的平衡,而无论是继续自博弈还是采用更激进的驾驶风格都无法达到这一平衡。
cs.RO / 65 / 2609.21621

Towards Fine-Grained Object Manipulation: SAM3-Guided Visuomotor Policy with Persistent Memory Learning and Focused Visual Conditioning

迈向细粒度物体操作:基于SAM3引导的持久记忆学习与聚焦视觉条件的视觉运动策略
Meng, Haolong, Qin, Fangbo, Bai, Mengchen, Wang, Houwu, Liu, Cirong, Yu, Shan
Abstract
Fine-grained object (FO) manipulation requires robots to distinguish a specified FO from visually similar objects and execute actions reliably despite scene distractors. However, scene-level visual conditioning lacks explicit object selection, while category-level guidance cannot reliably distinguish FOs within the same category. We present a SAM3-guided visuomotor framework that addresses these challenges through persistent object memory and focused visual conditioning. First, we introduce FO Memory-driven SAM3 (FOM-SAM3), which learns reusable FO memory tokens from limited multi-view registration images while keeping SAM3 fully frozen. Through one-vs-rest learning, these tokens encode persistent memories for localizing target FOs and rejecting similar alternatives, which can be stored in a memory bank. Second, we propose Focused Spatial-Appearance Encoding (FSAE), which combines in-FO local appearance features with explicit bounding-box coordinates to condition action policies including Diffusion Policy (DP) and Action Chunking with Transformers (ACT). The effectiveness of the proposed FOM-SAM3 was validated on the FO-30 dataset comprising 30 physical objects across four coarse categories. Across three real-robot FO manipulation tasks, our FOM-SAM3-guided policies demonstrated robustness against distractors, discrimination ability among similar FOs, and extendibility to new FOs.
Chinese Translation
细粒度物体(FO)操作要求机器人从视觉相似的物体中区分出指定的细粒度物体,并在存在场景干扰物的情况下可靠地执行动作。然而,场景级的视觉条件缺乏显式的物体选择机制,而类别级的引导无法可靠地区分同一类别内的细粒度物体。我们提出了一种SAM3引导的视觉运动框架,通过持久物体记忆和聚焦视觉条件来解决这些挑战。首先,我们引入了FO记忆驱动的SAM3(FOM-SAM3),它在保持SAM3完全冻结的同时,从有限的多视角配准图像中学习可复用的FO记忆标记。通过“一对多”学习,这些标记编码了用于定位目标FO并排除相似替代物的持久记忆,并可存储在记忆库中。其次,我们提出了聚焦空间-外观编码(FSAE),将FO内的局部外观特征与显式的边界框坐标相结合,为包括扩散策略和分块动作Transformer在内的动作策略提供条件。我们在包含四个粗类别共30个物理物体的FO-30数据集上验证了所提出的FOM-SAM3的有效性。在三个真实机器人FO操作任务中,我们基于FOM-SAM3引导的策略展现出对干扰物的鲁棒性、区分相似FO的能力以及对新FO的可扩展性。
cs.RO / 66 / 2609.21650

SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations

SynthDemo-RL:利用大语言模型引导的合成示范打破VLA适应中的零奖励障碍
Kingetsu, Hiroaki, Kurihara, Hiroaki, Yokoo, Kaoru, Fukumizu, Kenji, Kaul, Manohar
Abstract
Fine-tuning Vision-Language-Action (VLA) models commonly relies on human teleoperation demonstrations, while reinforcement learning (RL) with sparse binary rewards faces an exploration challenge when successful trajectories are rarely sampled. We propose SynthDemo-RL, a teacher-student framework in which an automated teacher converts simulator-privileged state into successful manipulation trajectories, a VLA student is distilled from them by supervised fine-tuning (SFT), and PPO with binary task-success rewards refines the student. We study reward coverage, the fraction of tasks for which at least one success is observed under the fixed evaluation protocol, as a complement to the average success rate. On LIBERO-PRO, a public benchmark of perturbed LIBERO tasks for which no demonstrations exist, 27 of 57 scored tasks are at exactly 0% success for a pi_0.5 policy fine-tuned on the original LIBERO tasks. Direct PPO from this policy, under the same PPO recipe and the same RL compute as SynthDemo-RL's refinement stage, rescues 10 of these 27 tasks and leaves 17 at 0%. SynthDemo-RL, with 50 synthesized trajectories per task and no new human demonstrations, rescues all 27 and reaches average success rates of 97.8% and 97.1% on the Position and Task axes of LIBERO-PRO, respectively. On standard LIBERO, the same pipeline reaches 96.0% with no human demonstrations, within 1.7 points of pi_0.5 trained on 50 human demonstrations per task. We further validate the pipeline on RoboTwin 2.0 and verify that trajectories from a policy trained in a MuJoCo twin execute open-loop on a physical robot.
Chinese Translation
视觉-语言-动作(VLA)模型的微调通常依赖人类遥操作示范,而采用稀疏二值奖励的强化学习(RL)在成功轨迹极少被采样时面临探索难题。我们提出SynthDemo-RL,这是一个师生框架:自动化教师将仿真器特权状态转化为成功的操作轨迹,VLA学生通过监督微调(SFT)从这些轨迹中蒸馏学习,再由采用二值任务成功奖励的PPO对学生进行精调。我们研究了奖励覆盖率——即在固定评估协议下至少观察到一次成功的任务所占比例——作为平均成功率的补充指标。在LIBERO-PRO(一个由扰动LIBERO任务组成、无任何示范可用的公开基准)上,在原始LIBERO任务上微调的pi_0.5策略在57个评分任务中有27个的成功率恰好为0%。在相同的PPO配置和与SynthDemo-RL精调阶段相同的RL计算量下,直接对该策略进行PPO仅能挽救这27个任务中的10个,仍有17个停留在0%。SynthDemo-RL在每个任务仅使用50条合成轨迹且无需新的人类示范的情况下,成功挽救了全部27个任务,在LIBERO-PRO的Position轴和Task轴上分别达到97.8%和97.1%的平均成功率。在标准LIBERO上,同样的流水线在完全不用人类示范的情况下达到96.0%,与每任务使用50条人类示范训练的pi_0.5相差仅1.7个百分点。我们还在RoboTwin 2.0上进一步验证了该流水线,并验证了在MuJoCo孪生环境中训练的策略所生成的轨迹能够在物理机器人上开环执行。
cs.RO / 67 / 2609.21659

Outcome-Conditioned End-Effector Geometry Across Vision-Language-Action Policies

视觉-语言-动作策略间以结果为条件的末端执行器几何特性
Lin, Xingyu, Li, Zhuang, Wu, Zhongrun, Zhou, Shouquan, Du, Dehui
Abstract
Vision-language-action (VLA) policies solve the same manipulation task through different action interfaces, but task success alone does not establish whether their physical executions agree. We study cross-policy end-effector geometry in 15,000 closed-loop LIBERO rollouts from four policies. The primary clean-condition analysis forms 3,600 configuration-matched, and therefore dependent, policy pairs. Both-success pairs have a median normalized dynamic time warping distance of 0.0120 m versus 0.0380 m when exactly one policy succeeds. This ordering holds in every task, every policy pair, and nine sampling and band-limited representations; however, the ratio varies severalfold across representations, so we report the direction rather than a fixed multiple. Both-failure pairs are more separated again but rest on thin, uneven support, so we report them as exploratory. Within successful executions, partner replacements separate more across tasks than across initial states. A matched baseline still reveals measurable, heterogeneous residual policy differences, so a low cross-policy distance does not imply interchangeability. Successful executions sit about as far from same-task demonstrations as those demonstrations sit from each other, compatible with task-associated geometry without separating training-data overlap from task constraints. A common 72-action window preserves the ordering but reduces its magnitude; endpoint and duration adjustment likewise leaves a positive mixed-outcome coefficient relative to both-success pairs, though its magnitude is specification-dependent. Under composite visual stress, policy rankings and pair composition change together.
Chinese Translation
视觉-语言-动作(VLA)策略通过不同的动作接口解决同一操作任务,但仅凭任务成功并不能确定它们的物理执行过程是否一致。我们研究了来自四种策略的15,000次LIBERO闭环 rollout 中的跨策略末端执行器几何特性。主要的洁净条件分析构建了3,600对配置匹配(因而是相关的)的策略对。双成功策略对的中位归一化动态时间规整(DTW)距离为0.0120米,而恰好只有一个策略成功的策略对为0.0380米。该排序在每一项任务、每一个策略对以及九种采样与带限表示中均成立;然而,该比率在不同表示间变化达数倍之多,因此我们报告的是方向而非固定的倍数。双失败策略对之间的分离程度更大,但其样本支撑稀薄且不均衡,因此我们仅将其作为探索性结果报告。在成功执行内部,相较于不同初始状态,更换配对对象在不同任务之间的分离更大。匹配基线仍显示出可测量且异质的残余策略差异,因此较低的跨策略距离并不意味着策略可以互换。成功执行与同任务示教之间的距离,约等于示教彼此之间的距离,这与任务相关几何特性相容,但无法将训练数据重叠与任务约束区分开来。共同的72步动作窗口保持了该排序但降低了其幅度;对终点和时长进行调整后,相对于双成功策略对仍存在正的混合结果系数,尽管其幅度依赖于具体设定。在复合视觉压力下,策略排名与策略对构成会一同发生变化。
cs.RO / 68 / 2609.21690

RAYA: Learning Where and When to Intervene for Robot Recovery

RAYA:学习机器人恢复的干预时机与干预位置
Mahajan, Ishaan, Chen, Charles, Dümbgen, Frederike, Plancher, Brian
Abstract
A robot can predict failure and still be unable to prevent it. By the time a safety mechanism reacts, the nominal plan may already have spent the control authority that recovery requires, and fixed task priorities may block whatever response remains. Our key insight is that both aspects are decided inside the controller. Recoverability must inform actions while they are chosen rather than veto them afterward, and task objectives must be adapted as recoverability shrinks. Building on this, we present RAYA, a hybrid learned-analytic framework that places a learned finite-horizon recoverability margin inside an optimal controller with hard constraints and pairs it with a bounded learned scheduler that shifts task weights to facilitate recovery. Across 7,200 simulation episodes per controller spanning quadrotor and autonomous-vehicle benchmarks, RAYA not only improves survival rates, but also transfers the learned components zero-shot to unseen trajectories, disturbances, plant shifts, and friction layouts. We developed an embedded realization of RAYA and deployed it on-board a 35g Crazyflie quadrotor. Across 40 combined hardware flights under wind with either aerodynamic mismatch or an unmodeled 40% motor-command loss, each of three baselines fails in all trials, while RAYA completes 10/10 six-cycle missions. Project Website: https://raya-control.github.io/.
Chinese Translation
机器人能够预测失效,却仍然可能无法阻止其发生。当安全机制作出反应时,名义控制方案可能已经耗尽了恢复所需的控制权限,而固定的任务优先级也可能阻断剩余的一切应对措施。我们的关键洞察是:这两个方面都是在控制器内部决定的。可恢复性必须在动作选择时参与决策,而非事后予以否决;同时,随着可恢复性的降低,任务目标也必须随之调整。基于此,我们提出了RAYA——一个学习与解析相结合的混合框架,它将学习得到的有限时域可恢复性裕度嵌入到具有硬约束的最优控制器中,并配以一个有界学习调度器,通过调整任务权重来促进恢复。在四旋翼和自动驾驶车辆基准上,每个控制器各进行7,200次仿真实验,RAYA不仅提高了生存率,还将学习组件零样本(zero-shot)迁移到未见过的轨迹、扰动、被控对象变化和摩擦布局中。我们开发了RAYA的嵌入式实现,并将其部署在重35克的Crazyflie四旋翼机载平台上。在共40次有风条件下的硬件飞行实验中(涉及气动失配或未建模的40%电机指令损失),三个基线方法在所有试验中均告失败,而RAYA则完成了10/10的六周期任务。项目网站:https://raya-control.github.io/。
cs.RO / 69 / 2609.21707

NeuRIO: A Streaming Neural Estimator for Zero-Shot Sim-to-Real Multi-Robot Relative Inertial Odometry

NeuRIO:一种用于零样本仿真到现实多机器人相对惯性里程计的流式神经估计器
Li, Zhehan, Lu, Jiadong, Ren, Shengwei, Xu, Chao, Cao, Yanjun
Abstract
We present NeuRIO, a streaming neural estimator for anchor-free 6-DoF relative inertial odometry using only identified inter-robot bearings, ranges, and IMU measurements. NeuRIO canonicalizes measurements into gravity-aligned coordinates, represents robots as nodes and mutual observations as factors, and uses attention for spatial reasoning and GRUs for temporal modeling. As a graph network, NeuRIO applies shared node-wise and factor-wise operators throughout the network, enabling it to handle different team sizes and time-varying observation graphs. NeuRIO is trained on a simulator that couples various motion patterns, device-level sensor characteristics, and diverse, realistic modeled, and temporally persistent sensor corruptions. In this way, NeuRIO achieves zero-shot sim-to-real transfer. Across $24$ real-world sequences, NeuRIO achieves $14.1\,\mathrm{cm}$ position RMSE and $3.9^\circ$ rotation RMSE. More importantly, NeuRIO demonstrates strong computational scalability, maintaining an update cost below $20\,\mathrm{ms}$ with up to $400$ robots in simulation, while optimization-based methods exceed $20\,\mathrm{ms}$ at only $24$ robots. Moreover, even trained on limited team sizes, NeuRIO transfers directly to unseen larger teams without architectural or parameter changes.
Chinese Translation
我们提出了NeuRIO,一种无锚点的六自由度相对惯性里程计流式神经估计器,其仅使用机器人间的识别方位角、距离和IMU测量数据。NeuRIO将测量数据规范化到重力对齐坐标系中,将机器人表示为节点、将相互观测表示为因子,并利用注意力机制进行空间推理、利用GRU进行时间建模。作为一个图网络,NeuRIO在整个网络中应用共享的节点级和因子级算子,使其能够处理不同规模的机器人团队和随时间变化的观测图。NeuRIO在一个耦合了多种运动模式、设备级传感器特性以及多样化、真实建模且时间上持续的传感器扰动的模拟器上进行训练。由此,NeuRIO实现了零样本的仿真到现实迁移。在24个真实世界序列上,NeuRIO实现了14.1厘米的位置RMSE和3.9度的旋转RMSE。更重要的是,NeuRIO展现出强大的计算可扩展性,在仿真中支持多达400个机器人时,其更新成本仍保持在20毫秒以下,而基于优化的方法在仅24个机器人时便已超过20毫秒。此外,即使仅在有限规模的团队上训练,NeuRIO也能在无需任何架构或参数更改的情况下直接迁移到未见过的更大规模团队。
cs.RO / 70 / 2609.21716

AgenticSwarm: Semantic Perception and Adaptive Task Allocation for Heterogeneous Multi-UAV Missions

AgenticSwarm:面向异构多无人机任务的语义感知与自适应任务分配
Mustafa, Muhammad Ahsan, Yaqoot, Yasheerah, Batool, Faryal, Khan, Roohan Ahmed, Serpiva, Valerii, Tsetserukou, Dzmitry
Abstract
Multi UAV missions in complex environments require the system to understand both the surrounding scene and the intent of a human operator while maintaining feasible task allocation as mission conditions change. This paper presents AgenticSwarm, an agentic framework for semantic perception and adaptive task allocation in heterogeneous multi UAV missions. An agent interprets aerial imagery and natural language instructions to construct a grounded mission representation that links perceived objects and regions with task requirements, capability constraints, and mission dependencies. This information augments a constrained task allocation process in which obstacle aware path feasibility, energy consumption, and protected return home requirements are incorporated before assignment. During execution, changes such as UAV failure, battery degradation, or task modification trigger residual mission reconstruction from the current system state, while completed work and reconnaissance progress are retained. AgenticSwarm is evaluated across five diverse Gazebo environments and an indoor real test environment, demonstrating its ability to connect semantic reasoning with constrained allocation and adaptive multi UAV mission execution. Compared with a Grounding DINO+SAM~2.1 perception baseline, the SAM3-based pipeline improves class-aware recall by 25.2 percentage points (pp) and semantic label accuracy by 29.5 pp. Ablating residual mission replanning increases mean repeated work from 0% to 61.7% and post-event recovery time by 58.6%, highlighting the contribution of adaptive replanning to mission execution.
Chinese Translation
复杂环境中的多无人机任务要求系统在任务条件变化时,既能理解周围场景和操作员的意图,又能保持可行的任务分配。本文提出 AgenticSwarm,一个面向异构多无人机任务的语义感知与自适应任务分配的智能体框架。该框架中的智能体通过解读航空影像和自然语言指令,构建一个任务表示,将感知到的物体和区域与任务需求、能力约束和任务依赖关系相关联。这些信息被用于增强带约束的任务分配过程,在分配任务之前纳入考虑避障路径可行性、能耗以及安全返航要求。在执行过程中,无人机故障、电量衰减或任务变更等事件会触发基于当前系统状态的剩余任务重构,同时保留已完成的工作和侦察进展。AgenticSwarm 在五个不同的 Gazebo 仿真环境和一个室内真实测试环境中进行了评估,证明了其将语义推理与带约束的任务分配及自适应多无人机任务执行相结合的能力。与 Grounding DINO+SAM 2.1 感知基线相比,基于 SAM3 的流程将类别感知召回率提高了 25.2 个百分点(pp),语义标签准确率提高了 29.5 个百分点。消融实验表明,移除剩余任务重规划会使平均重复工作量从 0% 增加到 61.7%,事件后恢复时间增加 58.6%,凸显了自适应重规划对任务执行的重要贡献。
cs.RO / 71 / 2609.21718

A Novel Path-Tracking Algorithm for Automated Tractor-Trailer Forward and Backward Maneuvers

一种用于自动牵引挂车前进行驶与倒车操作的新型路径跟踪算法
Lombard, Alexandre, Perronnet, Florent, Gaud, Nicolas, Abbas-Turki, Abdeljalil
Abstract
Fully autonomous tractor--trailer systems are increasingly deployed in logistics, agriculture, and industrial environments, where precise and robust path-tracking capabilities are essential. However, the articulation between the tractor and the trailer introduces additional nonlinearities and significantly complicates lateral and longitudinal control, particularly during reversing maneuvers. This paper introduces a novel path-tracking algorithm specifically designed for articulated vehicles with a single trailer. The proposed method combines a lateral control law applied at the trailer level with a short-horizon predictive adjustment of the tractor steering angle, ensuring stable convergence toward the desired path in both forward and backward motion. The approach is geometry-based and requires no per-vehicle calibration or training. Simulation studies in a high-fidelity physics simulator demonstrate the ability of the controller to match or outperform classical and state-of-the-art methods in terms of accuracy, stability, and robustness to disturbances.
Chinese Translation
全自主牵引挂车(tractor-trailer)系统在物流、农业和工业环境中的应用日益广泛,在这些场景中,精确且鲁棒的路径跟踪能力至关重要。然而,牵引车与挂车之间的铰接连接引入了额外的非线性因素,显著增加了横向与纵向控制的复杂度,尤其是在倒车操作过程中。本文提出了一种专门针对带单个挂车的铰接式车辆设计的新型路径跟踪算法。该方法将作用于挂车层面的横向控制律与对牵引车转向角的短时域预测调整相结合,确保车辆在前进行驶和倒车过程中均能稳定收敛至期望路径。该方法基于几何原理,无需针对每辆车进行标定或训练。在高保真物理仿真器中开展的仿真研究表明,该控制器在精度、稳定性以及抗干扰鲁棒性方面能够达到或超越经典方法及最先进方法。
cs.RO / 72 / 2609.21726

ZeroTouch: Tactile-Supervised Visual Contact Estimation for Contact-Rich Manipulation

ZeroTouch:面向接触密集型操作的触觉监督视觉接触估计方法
Kosenkov, Dmitriy, Zinniatullina, Daniia, Cabrera, Miguel Altamirano, Zhura, Iana, Derevianchenko, Mikhail, Tsetserukou, Dzmitry
Abstract
Reliable robotic grasping benefits from estimating the evolving physical interaction and selecting a grasp-dependent compression target. Tactile sensors provide direct interaction measurements but require dedicated hardware at deployment. We introduce ZeroTouch, a tactile-supervised framework that predicts dense contact deformation, the instantaneous six-axis wrench, and a grasp-dependent desired compression target from wrist RGB observations, gripper state, and local gravity direction. Tactile measurements are used only as privileged supervision during training and are not required at deployment. On the full validation set, the complete architecture reduces normal-force MAE from 2.017 N for a state-only baseline to 0.531 N. In physical evaluation with 20 trials per condition, ZeroTouch achieves 95% success on an unseen object, 80% in a seen-object/unseen-grasp condition, and 90% under a content/load shift. Under the same evaluation protocol, OpenVLA achieves 25%, 40%, and 55%, while SmolVLA achieves 10%, 25%, and 35%, respectively.
Chinese Translation
可靠的机器人抓取依赖于对不断演变的物理交互进行估计,并选择依赖于抓取方式的压缩目标。触觉传感器能够提供直接的交互测量,但在部署时需要专用硬件。我们提出了ZeroTouch,这是一个以触觉为监督的框架,可从腕部RGB观测、夹爪状态和局部重力方向中预测稠密接触变形、瞬时六轴力旋量以及依赖于抓取方式的期望压缩目标。触觉测量仅在训练阶段作为特权监督使用,在部署时无需触觉传感器。在完整验证集上,完整架构将法向力平均绝对误差(MAE)从仅使用状态的基线的2.017 N降低至0.531 N。在每种条件下进行20次试验的物理评估中,ZeroTouch在未见物体上达到95%的成功率,在已见物体/未见抓取条件下达到80%,在内容/负载变化下达到90%。在相同的评估协议下,OpenVLA分别达到25%、40%和55%,而SmolVLA分别达到10%、25%和35%。
cs.RO / 73 / 2609.21729

Visual Proactivity: Enhancing Human-Robot Collaboration Through Intent Communication

视觉主动性:通过意图交流增强人机协作
Bo, Valerio, Bejarano, Edison, Garrell, Anaís, Sanfeliu, Alberto
Abstract
As robots transition from performing repetitive tasks to collaborating with humans, understanding human intent becomes crucial to effective interaction. Anticipation enables robots to predict human actions, while proactivity allows them to take initiative and guide human behavior toward optimal outcomes. Although research has largely focused on how robots infer and respond to human intentions, less attention has been paid to how robots communicate their own intent. This paper introduces visual proactivity, a novel, simple yet effective approach that enables robots to communicate their intentions through visual feedback, influencing human behavior and enhancing transparency and fluency. We develop and evaluate proactive robotic behaviors in a human-to-robot handover scenario, where a user study validates human perception of reactive, anticipatory, and proactive behaviors. The results demonstrate that effective visual proactivity fosters better alignment and coordination, paving the way for more intuitive human-robot collaboration.
Chinese Translation
随着机器人从执行重复性任务转向与人类协作,理解人类意图成为实现有效交互的关键。预判能力使机器人能够预测人类的动作,而主动性则使机器人能够主动采取行动并引导人类行为以达到最优结果。尽管现有研究主要关注机器人如何推断并响应人类意图,但对机器人如何传达自身意图的关注相对较少。本文提出了视觉主动性(visual proactivity),这是一种新颖、简单而有效的方法,使机器人能够通过视觉反馈传达自身意图,从而影响人类行为并提升交互的透明度与流畅性。我们在人机递接场景中开发并评估了主动式机器人行为,用户实验验证了人类对反应式、预判式和主动式行为的感知。结果表明,有效的视觉主动性能够促进更好的一致性与协调性,为实现更直观的人机协作铺平道路。
cs.RO / 74 / 2609.21734

When Should Robots Intervene? Balancing Engagement and Intrusiveness in Human-Robot Interaction

机器人何时应进行干预?在人机交互中平衡用户参与度与侵入性
Hriscu, Lavinia, Bo, Valerio, Sanfeliu, Alberto, Garrell, Anaís
Abstract
Designing effective Human-Robot Interaction in task-oriented settings requires carefully balancing user engagement with socially acceptable levels of robot intrusiveness. In this paper, we examine how different robot intervention strategies shape user experience, interaction dynamics, perceived intrusiveness, and sense of support. We compare two approaches: a continuous engagement-seeking robot strategy, and a context-aware strategy that selectively intervenes based on the user's state and task context. Both approaches rely on multimodal behavioral cues, including body orientation and attentional signals, to guide robot actions. We evaluate these strategies in a user study with 32 participants performing a task in a simulated hospital environment. Our findings show that higher interaction frequency does not necessarily lead to better engagement. Instead, we observe a systematic trade-off between perceived support and intrusiveness, influenced by factors such as physical proximity and user effort. These results provide empirical evidence that effective engagement in HRI depends on adaptive, context-sensitive intervention policies.
Chinese Translation
在面向任务的环境中设计有效的人机交互(Human-Robot Interaction, HRI),需要仔细权衡用户参与度与社会可接受的机器人侵入性水平。本文研究了不同的机器人干预策略如何影响用户体验、交互动态、感知侵入性以及支持感。我们比较了两种策略:一种是持续寻求交互的机器人策略,另一种是情境感知策略,即根据用户状态和任务情境选择性地进行干预。两种策略均依赖于多模态行为线索,包括身体朝向和注意力信号,以引导机器人的行为。我们通过一项用户研究对这些策略进行了评估,共有32名参与者在模拟的医院环境中执行任务。研究结果表明,更高的交互频率并不一定能带来更好的参与度。相反,我们观察到感知支持与侵入性之间存在系统性的权衡,其受物理距离和用户付出程度等因素的影响。这些结果提供了实证证据,表明人机交互中有效的用户参与依赖于自适应的、对情境敏感的干预策略。
cs.RO / 75 / 2609.21740

Sandwich-Residuals: Parameter-Efficient Test-time Adaptation of World Models

三明治残差(Sandwich-Residuals):世界模型的参数高效测试时自适应方法
Soni, Krishnam, Sehgal, Aditya, Dave, Vedant, Rueckert, Elmar
Abstract
Latent world models enable planning by predicting the effects of actions in a learned representation space, but their predictions can become unreliable when test-time conditions differ from training. Existing test-time adaptation methods address this by updating parts of the pretrained model, often modifying millions of parameters and requiring a choice of which internal components to adapt. We introduce Sandwich-Residuals, a lightweight alternative that keeps the pretrained world model frozen and learns only small residual corrections around the predictor. The residuals are optimized online using the model's self-supervised prediction error and require no rewards, labels, or source-domain data. Across 21 conditions on the AdaJEPA benchmark, our method achieves $1.3\times$ the success rate of the frozen model while retaining 95% of the performance of the strongest AdaJEPA variant and adapting 97-99% fewer parameters. Under compound shifts, this advantage increases to $1.9\times$ the success rate of the frozen model, while remaining comparable to internal block adaptation. We further demonstrate the same adaptation principle on a DINO-WM model for 3-D manipulation. These results suggest that effective test-time adaptation of world models does not necessarily require modifying their pretrained internal weights.
Chinese Translation
潜在世界模型(Latent World Models)通过在学习的表征空间中预测动作的效果来实现规划,但当测试时条件与训练时不同时,其预测可能变得不可靠。现有的测试时自适应方法通过更新预训练模型的部分参数来解决这一问题,通常需要修改数百万个参数,并且需要选择要自适应的内部组件。我们提出了三明治残差(Sandwich-Residuals)这一轻量级替代方案,它保持预训练的世界模型冻结不变,仅学习预测器周围的小型残差校正。这些残差利用模型的自监督预测误差进行在线优化,无需奖励、标签或源域数据。在 AdaJEPA 基准的 21 个条件上,我们的方法实现了冻结模型 1.3 倍的成功率,同时保持了最强 AdaJEPA 变体 95% 的性能,且自适应的参数数量减少了 97%–99%。在复合偏移条件下,这一优势提升至冻结模型 1.9 倍的成功率,同时与内部模块自适应方法相当。我们进一步在用于三维操作的 DINO-WM 模型上验证了相同的自适应原理。这些结果表明,世界模型的有效测试时自适应并不一定需要修改其预训练的内部权重。
cs.RO / 76 / 2609.21744

Understanding Engagement and Intrusiveness in Assistive Human-Robot Interaction Using Individual Traits

基于个体特质的辅助性人机交互中参与度与侵入性的研究
Bo, Valerio, Hriscu, Lavinia, Sanfeliu, Alberto, Garrell, Anaís
Abstract
Robot assistance is particularly crucial in unfamiliar tasks, where users must understand task requirements while coordinating with the robot. Previous research offers mixed evidence on the role of robot proxemics in user engagement: some studies suggest closer proximity enhances interaction, while others report it can feel intrusive. In this work, we argue that perceptions of intrusiveness depend not only on proxemics but also on the frequency of robot interventions, and are strongly influenced by individual traits such as personality and demographics. We conducted an experiment with 32 participants who interacted with two assistive robots that provided similar task support but differed in their intervention strategies. Results indicate that overall engagement remains stable across conditions, yet affective responses and perceived intrusiveness vary significantly with personality traits. Moreover, personality shapes interaction dynamics differently depending on the robot's behavior. These findings emphasize that effective human-robot interaction should account for individual differences, tailoring robot behavior to maintain engagement while respecting each user's unique affective and behavioral profile.
Chinese Translation
机器人辅助在不熟悉的任务中尤为重要,用户需要在理解任务要求的同时与机器人进行协调。以往关于机器人近身行为(proxemics)对用户参与度作用的研究结果并不一致:一些研究表明更近的距离能够增强交互,而另一些研究则报告这会让人感到被侵入。在本研究中,我们认为侵入性感知不仅取决于机器人的近身距离,还取决于机器人干预的频率,并且受到个性和人口统计学特征等个体特质的强烈影响。我们开展了一项有32名参与者参与的实验,参与者与两个辅助机器人进行交互,这两个机器人提供相似的任务支持,但干预策略不同。结果表明,整体参与度在不同条件下保持稳定,但情感反应和感知侵入性随个性特质显著变化。此外,个性对交互动态的影响因机器人行为的不同而有所不同。这些发现强调,有效的人机交互应考虑个体差异,通过定制机器人行为来维持参与度,同时尊重每位用户独特的情感与行为特征。
cs.RO / 77 / 2609.21751

ForceTwin: Physics-informed Digital Twins for Robotic Manipulation from Instrumented Human Interaction

ForceTwin:基于仪器化人机交互的物理信息数字孪生用于机器人操作
Engelbracht, Tim, Zurbrügg, René, Mittal, Mayank, Hutter, Marco, Pollefeys, Marc, Blum, Hermann, Bauer, Zuria
Abstract
Manipulating objects requires understanding not only their motion, but also the physical properties that determine it. For articulated objects, these include inertia, friction, and mechanisms such as springs or door closers, whose effects can vary with configuration and velocity. Such properties are not directly observable from appearance: visually identical doors may require very different effort to manipulate. Existing digital-twin pipelines recover primarily kinematics or assign static physical parameters from visual and language priors, which can yield physically implausible estimates. As a result, state-dependent mechanism dynamics remain unidentified and are not represented in standard asset formats. We present ForceTwin, a system for identifying physics-informed digital twins of articulated objects from instrumented human interaction. A person probes an object using a handheld force-sensing gripper, providing synchronized poses and interaction forces from which we estimate the articulation, parametric dynamics including inertia, Coulomb friction, viscous damping, and a structured neural residual capturing state-dependent mechanism forces. ForceTwin nearly halves the inertial-parameter error of a VLM prior. As a feedforward dynamics model for impedance control on a Spot and a Franka FR3, ForceTwin achieves 87% goal completion across nine object-embodiment pairs, compared with 60% using VLM-prior and 57% using kinematics-only twins, with the largest gains on objects whose strong mechanisms cause both baselines to stall. We further use the identified twins to train whole-body door-traversal policies and deploy them in the real world. Project Page: https://timengelbracht.github.io/forcetwin-website/
Chinese Translation
操作物体不仅需要理解其运动,还需要理解决定其运动的物理属性。对于铰接物体而言,这些属性包括惯性、摩擦以及弹簧或闭门器等机构,其效果可能随构型和速度而变化。此类属性无法从外观直接观测:视觉上完全相同的门,操作起来所需的力量可能截然不同。现有的数字孪生流程主要恢复运动学信息,或基于视觉和语言先验赋予静态物理参数,这可能产生物理上不合理的估计。因此,依赖状态的机构动力学无法被识别,也无法在标准资产格式中表示。我们提出 ForceTwin,一个通过仪器化人机交互来识别铰接物体物理信息数字孪生的系统。使用者使用手持力传感夹持器对物体进行探测,提供同步的位姿和交互力,我们据此估计铰接关系以及参数化动力学(包括惯性、库仑摩擦、粘性阻尼),并利用结构化神经残差捕捉依赖状态的机构力。ForceTwin 将 VLM 先验的惯性参数误差几乎降低了一半。作为用于 Spot 和 Franka FR3 阻抗控制的前馈动力学模型,ForceTwin 在九个物体-机器人组合上实现了 87% 的目标完成率,而使用 VLM 先验的数字孪生仅为 60%,仅使用运动学的孪生为 57%,其中在那些具有强机构导致两种基线方法均停滞的物体上提升最为显著。我们还进一步利用识别出的数字孪生训练全身开门通行策略,并将其部署到真实世界中。项目主页:https://timengelbracht.github.io/forcetwin-website/
cs.RO / 78 / 2609.21753

PSR: Predictive Sensorimotor Representation Learning for Contact-Rich Manipulation

PSR:面向接触密集型操作的预测性感知运动表征学习
Li, Shengbao, Xu, Peng, Tang, Chao, Wei, Hao, Wang, Jiaheng, Yin, Hong, Chen, Jiangtao, Zhu, Jinxuan, Zhou, Zhong, Wang, Mengfan, Li, Tingguang
Abstract
Contact-rich manipulation requires policies to generate precise actions by reasoning over contact forces, robot configurations, and interaction histories beyond visual observations. Existing methods passively condition on force feedback rather than actively predicting future contact dynamics, limiting their ability to generate high-precision actions. To address this problem, we introduce Predictive Sensorimotor Representation (PSR) learning, a framework that learns a hierarchy of predictive representations from multimodal sensorimotor signals and integrates them into the action stream of a visuomotor policy. Specifically, during a pretraining stage, a multimodal Transformer is trained to learn a hierarchy of predictive representations by jointly forecasting future interaction dynamics. The learned hierarchy subsequently augments the action stream, enabling the resulting policy to exploit contact-relevant cues at multiple depths. We further instantiate PSR within a Vision-Language-Action (VLA) model, resulting in PSR-VLA, and evaluate it on six real-world contact-rich manipulation tasks. Experimental results show that PSR-VLA achieves 91.7% overall success, improving over $\pi_{0.5}$, ForceVLA-$\pi_{0.5}$, and ForceVLA2-$\pi_{0.5}$ by 30.0, 22.5, and 19.2 percentage points, respectively. These results demonstrate the effectiveness of the proposed PSR for force-aware, contact-rich manipulation. Videos of the tasks and stability tests are available at https://psr-vla.pages.dev/.
Chinese Translation
接触密集型操作要求策略在视觉观测之外,通过对接触力、机器人构型以及交互历史的推理来生成精确的动作。现有方法只是被动地以力反馈作为条件,而非主动预测未来的接触动力学,这限制了其生成高精度动作的能力。为解决这一问题,我们提出了预测性感知运动表征(Predictive Sensorimotor Representation,PSR)学习框架,该框架从多模态感知运动信号中学习层级化的预测表征,并将其集成到视觉运动策略的动作流中。具体而言,在预训练阶段,通过联合预测未来的交互动力学,训练一个多模态Transformer来学习层级化的预测表征。随后,所学习到的层级表征被用于增强动作流,使得到的策略能够在多个深度上利用与接触相关的线索。我们进一步将PSR实例化到一个视觉-语言-动作(Vision-Language-Action,VLA)模型中,得到PSR-VLA,并在六项真实世界的接触密集型操作任务上对其进行评估。实验结果表明,PSR-VLA的整体成功率达到91.7%,分别比$\pi_{0.5}$、ForceVLA-$\pi_{0.5}$和ForceVLA2-$\pi_{0.5}$提升了30.0、22.5和19.2个百分点。这些结果验证了所提出的PSR在力感知的接触密集型操作中的有效性。任务视频和稳定性测试见 https://psr-vla.pages.dev/。
cs.RO / 79 / 2609.21761

CRISP: Contact-Rich Robotic Simulation Platform with Extensive Geometries and Contact Solvers

CRISP:支持丰富几何与接触求解器的接触密集型机器人仿真平台
Lee, Somang, Park, Sunkyung, Yun, Jinhee, An, Seoki, Lee, Dongjun
Abstract
We present CRISP (Contact-RIch Simulation Platform), a high-fidelity physics engine tailored for complex multi-contact simulations such as tight-tolerance robotic manipulation. Achieving high physical fidelity in robotic simulation requires both expressive modeling of geometry and contact interactions, as well as accurate numerical resolution via robust collision detection and contact solvers. However, existing simulators often either rely on limited support for geometric representations and simplified modeling of contact interactions, or employ numerical resolution methods whose accuracy or robustness is inherently constrained. Accordingly, we develop a new simulator that supports diverse geometric representations with accurate optimization-based collision detection, and combines contact modeling with robust augmented Lagrangian-based contact solvers. This integration enables efficient and consistent detection of contact information across complex geometries while accurately resolving multi-contact constraints without problematic relaxations, which is essential for simulating contact-intensive and sharp interactions. We validate the physical fidelity of our simulator against state-of-the-art platforms and further demonstrate its capabilities through complex robotic manipulation scenarios. CRISP is publicly available at https://github.com/INRoL/crisp.
Chinese Translation
我们提出了CRISP(Contact-RIch Simulation Platform,接触密集仿真平台),这是一个面向复杂多接触仿真(如高精度容差机器人操作)的高保真物理引擎。在机器人仿真中实现高物理保真度,既需要对几何与接触交互进行富有表现力的建模,也需要通过鲁棒的碰撞检测和接触求解器实现精确的数值求解。然而,现有的仿真器往往要么对几何表示的支持有限、对接触交互的建模过于简化,要么所采用的数值求解方法在精度或鲁棒性上存在固有限制。为此,我们开发了一款新的仿真器,它支持多样的几何表示并配备精确的基于优化的碰撞检测,同时将接触建模与鲁棒的基于增广拉格朗日方法(augmented Lagrangian)的接触求解器相结合。这种集成能够在复杂几何体之间高效且一致地检测接触信息,并精确求解多接触约束而无需引入有问题的松弛处理,这对于仿真接触密集且尖锐的交互至关重要。我们通过对比最先进的仿真平台验证了本仿真器的物理保真度,并通过复杂的机器人操作场景进一步展示了其能力。CRISP已在 https://github.com/INRoL/crisp 公开发布。
cs.RO / 80 / 2609.21767

Scaling Vision-Language Reward Learning for Robot Manipulation in Parallel Simulation

面向并行仿真中机器人操作的视觉-语言奖励学习规模化方法
Joualy, Lobna, Demeester, Eric, Tsiogkas, Nikolaos
Abstract
Vision-language models (VLMs) can replace human annotators in preference-based reward learning, but sequential API requests and single-environment data collection make training slow and costly. We present RAPID (Reward learning with Adaptive Parallel Image Diversity), a system that couples GPU-parallel rollout with data-aware policy updates, single-request preference labeling, automatic reward stabilization, and representative image sampling. We evaluate these components on five Franka Panda manipulation tasks in IsaacLab. Parallel rollout and adaptive updates provide the first substantial reduction in training time: under matched two-stage prompting, mean runtime falls from 9.18 to 3.13 hours. With all RAPID components enabled, training completes in 1.15 hours using 896 rather than 19,840 API calls per run, and aggregate final success rises from 86.3\% to 98.7\%. This represents an 8.0$\times$ end-to-end speedup and a 95.5\% reduction in API usage. An offline evaluation with Gemma~3 12B and GPT-4.1 mini demonstrates that single-request prompting reduces labeling latency and cost across both models. Code is available at: https://github.com/rapid-vlm/rapid-vlm-rl.
Chinese Translation
视觉-语言模型(VLM)可以在基于偏好的奖励学习中替代人工标注,但串行的API请求和单环境数据采集使得训练缓慢且成本高昂。我们提出了RAPID(Reward learning with Adaptive Parallel Image Diversity,基于自适应并行图像多样性的奖励学习)系统,该系统将GPU并行rollout与数据感知的策略更新、单请求偏好标注、自动奖励稳定化以及代表性图像采样相结合。我们在IsaacLab中的五个Franka Panda操作任务上对这些组件进行了评估。并行rollout与自适应更新首次带来了显著的训练时间缩减:在匹配的两阶段提示条件下,平均运行时间从9.18小时降至3.13小时。启用RAPID的全部组件后,训练可在1.15小时内完成,每次运行的API调用次数从19,840次降至896次,总体最终成功率从86.3%提升至98.7%。这相当于8.0倍的端到端加速和95.5%的API使用量减少。基于Gemma 3 12B和GPT-4.1 mini的离线评估表明,单请求提示可降低两种模型的标注延迟与成本。代码可在以下网址获取:https://github.com/rapid-vlm/rapid-vlm-rl。
cs.RO / 81 / 2609.21777

TRACE: Coverage Path Planning for Unknown Environments Using Hierarchical Coverage Tree

TRACE:基于层次覆盖树的未知环境覆盖路径规划
Shen, Zongyuan, Liu, Haodong, Wang, Gao, Zhao, Shancheng, Zhou, Dehua, Ou, Yaming, Ren, Zhongqiang, Zhai, Yikui, Chen, C. L. Philip
Abstract
This paper presents a novel online coverage path planning (CPP) algorithm, called TRACE, for real-time coverage of unknown environments. TRACE is built upon a hierarchical coverage tree that provides a global representation of the evolving connectivity of the uncovered space. As the environment is incrementally revealed and covered, newly discovered obstacles and covered cells may fragment the remaining uncovered space into disconnected regions. TRACE recursively expands the corresponding tree nodes to explicitly represent these regions and organize them for subsequent coverage planning. Based on the updated tree, an incremental global tour is maintained to guide the coverage process. TRACE locally refines only the affected portions while preserving the visiting order of unchanged regions, thereby reducing the computational burden of global replanning and maintaining a consistent coverage progression. Guided by the global tour, a local planner generates back-and-forth coverage paths and switches to global-tour-aware planning to efficiently complete the target regions. Theoretical analysis establishes the computational complexity and complete coverage property of TRACE, and derives an approximation bound for the incremental global tour refinement. The performance of TRACE is evaluated through extensive high-fidelity simulations and real-robot experiments using a mobile robot. Comparative evaluations against six existing CPP methods demonstrate significant improvements in coverage time, path length, overlap ratio, and number of turns.
Chinese Translation
本文提出了一种新颖的在线覆盖路径规划(CPP)算法,称为TRACE,用于未知环境的实时覆盖。TRACE建立在层次覆盖树之上,该树为不断演变的未覆盖空间连通性提供全局表示。随着环境被逐步探索和覆盖,新发现的障碍物和已覆盖的栅格可能将剩余未覆盖空间分割为多个不连通的区域。TRACE通过递归扩展相应的树节点来显式表示这些区域,并对它们进行组织以用于后续的覆盖规划。基于更新后的树,系统维护一条增量式全局巡回路径以引导覆盖过程。TRACE仅对受影响的部分进行局部细化,同时保持未变化区域的访问顺序,从而降低全局重规划的计算负担,并保持一致的覆盖推进。在全局巡回路径的引导下,局部规划器生成往复式覆盖路径,并切换到全局巡回感知的规划方式,以高效完成目标区域的覆盖。理论分析建立了TRACE的计算复杂度与完全覆盖性质,并推导出增量式全局巡回路径细化的近似界。通过大量高保真仿真以及移动机器人的真实实验对TRACE的性能进行了评估。与六种现有CPP方法的对比评估表明,TRACE在覆盖时间、路径长度、重复覆盖率和转弯次数方面均有显著提升。
cs.RO / 82 / 2609.21787

Compact but Moving: Intervention-Relevant Geometry in Recurrent World Models

紧凑而流动:循环世界模型中与干预相关的几何结构
Chen, Yuming, Liu, Yang
Abstract
Learned world models may have compact interventions even when their recurrent state is high-dimensional, but it is unclear what happens to such a correction after it enters the model. We study this question in a controlled recurrent world model where prior work identified a checkpoint-specific rank-4 interface for one-shot counterfactual velocity interventions. The correction rapidly leaves this fixed entry subspace during autonomous rollout. Nevertheless, a low-rank image obtained by transporting the entry directions through the factual recurrent Jacobian chain continues to capture most of the nonlinear correction. Restarts using the tangent-predicted correction preserve substantial counterfactual future function. This transport/function pattern recurs across independently trained structured-GRU models and a parameter-matched LSTM initialized with a privileged compact correction. We further characterize a finite-horizon future-response operator over the full recurrent carrier. Patching shifts its leading future-sensitive directions toward the matched native-counterfactual organization, and the local operator accurately ranks finite perturbation effects over the registered direction panels at the patched and native-counterfactual basepoints. A separate full-amplitude assay finds substantial factual-endpoint tangent residuals and supports response reconfiguration in two of three checkpoints. Together, these results show that compact intervention structure can persist as a moving, state-dependent local geometry embedded in high-dimensional recurrent dynamics, without implying a fixed or dynamically closed low-dimensional state.
Chinese Translation
学习到的世界模型即使其循环状态是高维的,也可能具有紧凑的干预结构,但这种修正进入模型后会发生什么尚不清楚。我们在一个受控的循环世界模型中研究这一问题,此前的工作在该模型中发现了一个针对单次反事实速度干预的、依赖于特定检查点(checkpoint)的秩为4的接口。在自主 rollout 过程中,修正会迅速离开这一固定的入口子空间。然而,通过事实性循环雅可比矩阵链传输入口方向所得到的低秩像,仍然能捕捉大部分非线性修正。使用切线预测修正进行重启能够保留相当程度的反事实未来功能。这种传输/功能模式在多个独立训练的结构化 GRU(structured-GRU)模型以及一个以特权紧凑修正初始化、参数量匹配的 LSTM 中重复出现。我们进一步刻画了定义在完整循环载体上的有限时域未来响应算子。修补(patching)使其最敏感的未来方向向匹配的原生反事实组织偏移,且该局部算子能够在修补点与原生反事实基点处,对已注册方向面板上的有限扰动效应进行准确排序。另一项全幅值实验发现了显著的事实性终点切线残差,并支持在三个检查点中的两个存在响应重构。综上,这些结果表明,紧凑的干预结构可以作为嵌入在高维循环动力学中的、随状态移动的局部几何而持续存在,但这并不意味着存在一个固定的或动力学封闭的低维状态。
cs.RO / 83 / 2609.21788

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

从预训练到精通:面向长时程操作的极少人工干预现实世界子任务强化学习
Su, Sichang, Yang, Benjamin, Deng, Zhiyun, Liang, Boyuan, Yeung, Yip Fun, Wang, Zelin, Sun, Lingfeng
Abstract
A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks. Collecting additional full-task demonstrations for supervised fine-tuning (SFT) requires operators to repeat behaviors the policy already performs well. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches struggle to solve long-horizon tasks using only sparse rewards. We present PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at these bottlenecks while allowing training rollouts to proceed with minimal human intervention. The frozen pretrained policy supplies nominal actions throughout execution, while agent-generated selectors and success verifiers activate residual corrections and provide local outcome rewards. These rewards support learning from successful subtasks even when complete-task successes are scarce. Training combines online RL with success-reweighted retraining, and each retrained residual policy is redeployed to collect further experience. Humans identify bottlenecks during setup and perform physical resets when needed. On bimanual YAM and single-arm Franka tasks, PARTS improves complete-task success from 32% to 61% and from 50% to 95%, respectively, using tens of minutes of real-world RL rollouts per task on average. Compared with existing real-world RL fine-tuning methods, PARTS raises full-task success by more than 25% under the same robot-rollout budget while requiring less human involvement.
Chinese Translation
一个经过预训练的机器人基础策略或许能够执行长时程任务的大部分内容,却在少数关键子任务上反复失败。为进行监督微调(SFT)而收集更多的完整任务演示,要求操作者重复执行策略已经能够良好完成的动作。强化学习(RL)微调为弥合这一差距提供了一条有前景的路径,但现有方法难以仅使用稀疏奖励来解决长时程任务。我们提出了PARTS(Policy Adaptation with RL on Targeted Subtasks,基于定向子任务强化学习的策略自适应),这是一个现实世界的子任务强化学习框架,它将训练集中于这些瓶颈子任务,同时允许训练回合在极少人工干预的情况下进行。冻结的预训练策略在整个执行过程中提供名义动作,而由智能体生成的选择器(selector)和成功验证器(success verifier)则激活残差修正并提供局部结果奖励。即使完整任务的成功案例稀少,这些奖励也能支持从成功的子任务中学习。训练将在线强化学习与基于成功率的加权重训练相结合,每次重训练得到的残差策略都会被重新部署以收集更多经验。人类在初始设置阶段识别瓶颈,并在需要时进行物理复位。在双臂YAM和单臂Franka任务上,PARTS将完整任务成功率分别从32%提升至61%、从50%提升至95%,每项任务平均仅需数十分钟的现实世界强化学习回合。与现有的现实世界强化学习微调方法相比,在相同的机器人回合预算下,PARTS将完整任务成功率提升了25%以上,同时所需的人工参与更少。
cs.RO / 84 / 2609.21792

AcousticDiffusion: Semantically Conditioned Audio-Guided Diffusion Policy for Search-and-Rescue Assistance

AcousticDiffusion:面向搜救辅助的语义条件化音频引导扩散策略
Zhura, Iana, Seyidov, Didar, Plotnikov, Dmitrii, Amjad, Hajira, Cabrera, Miguel Altamirano, Tsetserukou, Dzmitry
Abstract
Navigating toward human callers is an important capability for rescue robots operating where visual contact is degraded or occluded. We present AcousticDiffusion, a semantically conditioned, audio-guided diffusion policy for human-directed navigation. A frozen pretrained audio recognizer processes 10.24 s windows, with speech gating and distress-aware prioritization converting recognition outputs into source-level navigation roles. Microphone-array direction-of-arrival measurements are recursively integrated into a robot-centric Bayesian bird's-eye-view belief field. Ego-motion compensation aligns successive observations, progressively constraining source position while preserving bearing-induced range uncertainty. The semantic belief, recent acoustic observations, audio features, and robot state condition a diffusion model that generates waypoint trajectories. On a synthetic-navigation validation set using recorded audio, AcousticDiffusion achieves a mean end-point bearing error of 11.20 degrees, with 91.78% of trajectories aligned within 30 degrees of the caller. Distractor rejection ranges from 89.20% to 98.99%, and the policy favors a HELP-designated caller over a competing speaker in 91.07% of windows. Deployed online on a ZSL-1 quadruped without additional retraining, it achieves a mean bearing error of 64.9 degrees, compared with 98.2 degrees for A* and 90.4 degrees for RRT, with a mean planner compute time of 6.07 ms. Despite imperfect acoustic localization, the reported mean final source distance is reduced from 3.96 m for the classical planners using ODAS-derived (Open embedded Audition System) guidance to 2.48 m, a 37.4% improvement. These results demonstrate the framework's ability to translate uncertain acoustic observations into closer approaches to human callers.
Chinese Translation
在视觉接触退化或被遮挡的环境中,向人类呼救者导航是救援机器人的一项重要能力。我们提出了AcousticDiffusion,这是一种用于人类引导导航的语义条件化、音频引导扩散策略。一个冻结的预训练音频识别器处理10.24秒的时间窗口,通过语音门控和求救感知的优先级排序,将识别输出转化为源级导航角色。麦克风阵列的到达角(DOA)测量被递归地融合到以机器人为中心的贝叶斯鸟瞰信念场中。自运动补偿用于对齐相继的观测,在保留由方位角引起的距离不确定性的同时,逐步约束声源位置。语义信念、近期声学观测、音频特征和机器人状态共同条件化一个扩散模型,该模型生成路径点轨迹。在一个使用录音音频的合成导航验证集上,AcousticDiffusion实现了11.20度的平均终点方位误差,其中91.78%的轨迹与呼救者的偏差在30度以内。干扰源拒绝率介于89.20%至98.99%之间,且该策略在91.07%的窗口中优先选择被标记为HELP的呼救者而非竞争说话者。在ZSL-1四足机器人上在线部署而无需额外重训练的情况下,其平均方位误差为64.9度,相比之下A*为98.2度,RRT为90.4度,平均规划器计算时间为6.07毫秒。尽管声学定位并不完美,报告的平均最终源距离从使用ODAS(Open embedded Audition System,开放嵌入式听觉系统)导引的经典规划器的3.96米降低到2.48米,提升了37.4%。这些结果表明,该框架能够将不确定的声学观测转化为对人类呼救者的更近距离接近。
cs.RO / 85 / 2609.21803

Contact-Rich Motion Planning via GPU-Parallel Mode Evaluation

基于GPU并行模式评估的接触丰富运动规划
Li, Jiayun, Chalvatzaki, Georgia
Abstract
Contact-rich motion planning (CRMP) is essential for robotic manipulation and locomotion, yet remains computationally challenging due to combinatorial contact decisions. Existing methods typically avoid broad evaluation of contact-mode sequences through search heuristics or optimization reformulations. We revisit broad evaluation in light of modern GPU hardware and introduce Contact-Mode Expansion with parallel Trajectory optimization (CoMET), which combines GPU-parallel trajectory evaluation with greedy contact-mode expansion. On planar pushing benchmarks, CoMET is competitive with optimization-based, sampling, and tree-search baselines in solution quality and planning time, matching the full-enumeration reference on nearly all instances with fewer evaluations and shorter planning times. Ablations suggest that much of the performance gain comes from the high-throughput trajectory evaluator. In bimanual nonprehensile manipulation, GPU-friendly local mode expansion achieves higher planning success than the tested adaptive tree search as the mode space grows. These results demonstrate that broad explicit mode evaluation provides a simple yet effective alternative for CRMP.
Chinese Translation
接触丰富运动规划(Contact-Rich Motion Planning, CRMP)对机器人操作与运动控制至关重要,但由于接触决策的组合爆炸特性,其计算仍极具挑战性。现有方法通常通过搜索启发式或优化重构来避免对接触模式序列进行广泛评估。我们基于现代GPU硬件重新审视了广泛评估的可行性,并提出了一种结合GPU并行轨迹评估与贪心接触模式扩展的方法——并行轨迹优化的接触模式扩展(Contact-Mode Expansion with parallel Trajectory optimization, CoMET)。在平面推物基准测试中,CoMET在解的质量和规划时间方面与基于优化的方法、采样方法及树搜索基线具有竞争力,在几乎全部实例上以更少的评估次数和更短的规划时间达到与完全枚举参考方法相当的效果。消融实验表明,性能提升很大程度来自高吞吐量的轨迹评估器。在双手非抓取操作任务中,随着模式空间的增大,GPU友好的局部模式扩展比所测试的自适应树搜索取得了更高的规划成功率。这些结果表明,广泛显式的模式评估为CRMP提供了一种简单而有效的替代方案。
cs.RO / 86 / 2609.21817

A Sim-to-Real Integration Pipeline for Training and Deployment of Chunk-Based VLA Manipulation Policies

一种用于基于动作块(Chunk)的视觉-语言-动作(VLA)操作策略训练与部署的仿真到现实迁移流水线
Kappel, Mathilde, Grislain, Clémence, Chetouani, Mohamed, Sigaud, Olivier, Annabi, Louis, Amar, Fa\"ız Ben, Doncieux, Stéphane, Khoramshahi, Mahdi
Abstract
Vision-Language-Action (VLA) models have become a prominent paradigm for mapping multimodal inputs, including semantic instructions, visual observations of the scene, and proprioceptive observations, to robot actions. Most state-of-the-art models predict actions in the end-effector pose space as sequences of action chunks. Training and evaluating these models requires large-scale collections of real-world demonstrations, pairing robot actions with the corresponding visual and proprioceptive observations. Collecting such data on real hardware typically relies on human teleoperation, making the process costly, time-consuming, and difficult to scale. We present an open-source sim-to-real experimental protocol that addresses this bottleneck: expert trajectories generated in simulation are replayed open-loop on a real Franka FR3 setup, where the corresponding real visual and proprioceptive observations are recorded and converted into a format compatible with VLA training. The same deployment stack is then reused, in closed-loop, to evaluate a trained policy on that setup, so that data collection and evaluation share an identical hardware configuration. Because each real recording is paired with the simulated trajectory that produced it, the protocol also yields a direct measurement of the sim-to-real gap. We release the collected datasets on Hugging Face together with the pipeline source code https://gitlab.isir.upmc.fr/kappel/sim2real_public_chunk_control.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型已成为将多模态输入(包括语义指令、场景视觉观测和本体感觉观测)映射到机器人动作的重要范式。大多数最先进的模型在末端执行器位姿空间中以动作块序列的形式预测动作。训练和评估这些模型需要大规模的真实世界演示数据集,即将机器人动作与相应的视觉和本体感觉观测进行配对。在真实硬件上采集此类数据通常依赖人工遥操作,导致该过程成本高昂、耗时且难以规模化。我们提出了一套开源的仿真到现实(sim-to-real)实验协议来解决这一瓶颈:在仿真中生成的专家轨迹被开环回放到真实的 Franka FR3 系统上,同时记录相应的真实视觉与本体感觉观测,并将其转换为与 VLA 训练兼容的格式。随后,相同的部署系统被闭环复用于在该系统上评估训练好的策略,从而使数据采集与评估共享完全相同的硬件配置。由于每条真实记录都与生成它的仿真轨迹配对,该协议还能够直接度量仿真到现实的差距(sim-to-real gap)。我们在 Hugging Face 上发布了所采集的数据集以及流水线源代码,地址为 https://gitlab.isir.upmc.fr/kappel/sim2real_public_chunk_control。
cs.RO / 87 / 2609.21818

LunaDrive: A Delay-Compensated High-Voltage GaN FET-Based Motor Driver for Dynamic Robots with Flat BLDC Motors

LunaDrive:一款面向扁平无刷直流电机动态机器人的具有延迟补偿功能的高压GaN FET电机驱动器
Yuzaki, Sota, Suzuki, Temma, Tada, Hiromi, Konishi, Masanori, Kawaharazuka, Kento, Okada, Kei
Abstract
The performance improvement of high-power flat BLDC motors has accelerated the development of dynamic robots. However, many commercially available servo motors assume operating voltages of 48 V or lower, which limits the maximum rotational speed. Dynamic robots require rapid energy generation, so this voltage constraint restricts motion performance. Therefore, driving motors beyond the rated voltage is desirable to increase the instantaneous maximum speed. On the other hand, semiconductor devices used in motor drivers have a trade-off between voltage rating and current capacity. Conventional drivers using Si MOSFETs have difficulty achieving both high-voltage and high-current operation. Although GaN FETs are promising, compact drivers that can be mounted on the rear side of flat BLDC motors remain limited. In this study, a motor driver for high-power flat BLDC motors using GaN FETs is developed. The effect of delay compensation in the high-speed region beyond the rated operating range is also investigated. In the experiments, under 96 V operation, a continuous current of 30 A was achieved with a heat sink attached. A peak current of 80 A and a maximum electrical frequency of 3110 Hz were confirmed. A high-speed load lifting experiment driven by a LiPo battery 24S (100 V) was also conducted, demonstrating applicability to dynamic robot operation.
Chinese Translation
大功率扁平无刷直流电机(BLDC)性能的提升加速了动态机器人的发展。然而,许多市售伺服电机的额定工作电压为48 V或更低,这限制了最大转速。动态机器人需要快速产生能量,因此该电压限制制约了运动性能。因此,希望驱动电机在超过额定电压下运行以提高瞬时最大转速。另一方面,电机驱动器中使用的半导体器件在电压等级与电流容量之间存在权衡。采用Si MOSFET的传统驱动器难以同时实现高压和大电流运行。尽管GaN FET前景广阔,但可安装于扁平BLDC电机背面的紧凑型驱动器仍然有限。本研究开发了一款基于GaN FET的大功率扁平BLDC电机驱动器,并研究了其在超出额定工作范围的高速区域的延迟补偿效果。在实验中,在96 V工作电压下,附加散热器后实现了30 A的连续电流;确认了80 A的峰值电流和3110 Hz的最大电频率。此外,还进行了由24S LiPo电池(100 V)供电的高速负载提升实验,验证了其在动态机器人运行中的适用性。
cs.RO / 88 / 2609.21838

PopNavShift: Stress-Testing Social Navigation under Behavioral Population Shift

PopNavShift:在行为人群偏移下对社会导航进行压力测试
Tan, Kaizhen, Zheng, Diyu, Wu, Tim Guangyu, Guan, ChengHe
Abstract
Social-navigation algorithms are often evaluated under a fixed pedestrian-behavior distribution, despite substantial variation in pedestrian responses to robots across individuals and social contexts. We introduce PopNavShift, a matched simulation framework for stress-testing social-navigation strategies under pedestrian population shifts. PopNavShift constructs population-conditioned pedestrian motion profiles by prompting Gemini 3.7 Flash with 600 synthetic persona records from MatrAIx Persona 1M and deterministically mapping the responses into bounded motion parameters. It then compares three representative navigation strategies, reactive avoidance, early yielding, and reciprocal collision avoidance, across eight population conditions and 7,488 matched robot runs. In a matched intervention on the same 202 personas, changing only time pressure reverses 8.6% of controller rankings based on robot travel time, but 22.4% based on mean pedestrian delay and 23.9% based on worst-decile delay. Across population conditions, this sensitivity is greater for pedestrian burden than for robot travel time and increases in spatially constrained settings; the same qualitative pattern persists under a second pedestrian dynamics model. These findings support evaluating navigation strategies across behavioral populations using both robot performance and pedestrian burden.
Chinese Translation
社会导航算法通常在固定的行人行为分布下进行评估,然而行人对机器人的反应在不同个体和社会情境间存在显著差异。我们提出了 PopNavShift,一个用于在行人人群偏移下对社会导航策略进行压力测试的匹配仿真框架。PopNavShift 通过向 Gemini 3.7 Flash 提示来自 MatrAIx Persona 1M 的 600 条合成人物画像记录,并将响应确定性地映射到有界的运动参数中,从而构建以人群为条件的行人运动特征。随后,该框架在八种人群条件和 7,488 次匹配的机器人运行中比较了三种代表性导航策略:反应式避障、提前让行和互惠碰撞避免。在针对相同 202 个人物画像的匹配干预中,仅改变时间压力就会使 8.6% 的控制器排名(基于机器人行驶时间)发生反转,而基于行人平均延误的排名反转比例为 22.4%,基于最差十分位延误的排名反转比例为 23.9%。在不同人群条件下,这种敏感性在行人负担方面高于机器人行驶时间,并且在空间受限的环境中进一步增大;在第二种行人动力学模型下,相同的定性规律依然成立。这些发现支持在行为人群范围内,同时使用机器人性能和行人负担来评估导航策略。
cs.RO / 89 / 2609.21883

VIRGA: Virtual-Agent-Intermediated Riemannian Geometry for Active-Sensing Air-Ground Coordination

VIRGA:面向主动感知空地协同的虚拟智能体中介黎曼几何方法
Guo, Fenghe, Shen, Runjie, Sun, Chenyang, Zhang, Junrui
Abstract
Air-ground autonomy becomes harder when the unmanned aerial vehicle (UAV) must remain observable by a gimbal light detection and ranging (LiDAR) mounted on the unmanned ground vehicle (UGV). The platforms must avoid dynamic obstacles while coordinating heterogeneous motion, limited sensing, and changing task initiative within one closed loop. This paper presents VIRGA, a neural geometric coordination framework that turns dual-LiDAR observations into bounded source-specific Riemannian fields and couples them through a virtual agent with reciprocal elastic feedback. Platform-aware execution maps convert the shared coordination reference into feasible UAV, UGV, and gimbal commands while enforcing active-observation safeguards. Evaluation against three complementary baselines reveals distinct limitations. An adapted Ray-RMP controller provides the fastest Riemannian response but produces insufficient clearance in the coupled air-ground task. A dense analytical Riemannian field improves geometric avoidance, yet its high evaluation cost prevents stable field-of-view maintenance. An adapted ColAG controller achieves the lowest latency but still incurs safety and observability violations. VIRGA completes all paired warehouse conditions safely, while a long-range cave stress test without retraining demonstrates sustained coordination in irregular and confined geometry. Ablations confirm contributions from online geometric evaluation, virtual-agent mediation, and reciprocal feedback.
Chinese Translation
当无人机(UAV)必须始终保持在无人地面车辆(UGV)所搭载的云台激光雷达(LiDAR)的可见范围内时,空地自主协同变得更加困难。两个平台必须在单一闭环内避免动态障碍物,同时协调异构运动、有限感知以及不断变化的任务主导权。本文提出了VIRGA,一种神经几何协同框架,它将双激光雷达观测转化为有界的源特定黎曼场,并通过一个具有双向弹性反馈的虚拟智能体将它们耦合起来。平台感知的执行映射将共享的协同参考转换为可行的UAV、UGV和云台指令,同时实施主动观测保障措施。通过与三个互补基线的对比评估,揭示了各自的局限性。改进的Ray-RMP控制器提供了最快的黎曼响应,但在空地耦合任务中产生的安全间隙不足。稠密解析黎曼场改善了几何避障效果,但其高昂的计算成本使其无法稳定维持视场。改进的ColAG控制器延迟最低,但仍会出现安全性和可观测性违规。VIRGA在所有成对仓库场景中均能安全完成任务,而无需重新训练的长距离洞穴压力测试则证明了其在不规则且受限的几何环境中持续的协同能力。消融实验证实了在线几何评估、虚拟智能体中介以及双向反馈的贡献。
cs.RO / 90 / 2609.21908

CommitFlow: Semantic Commitment Verification and Local Correction for Long-Horizon Robot Manipulation VLA Execution

CommitFlow:面向长时程机器人操作VLA执行的语义承诺验证与局部校正
Zhao, Zixiang, Feng, Yansong, Yang, Yang, Wang, Chaoyu, Xiao, Haoran, Zhang, Hui, Cheng, Chuang, Ma, Jianjun
Abstract
Although vision-language-action (VLA) policies have advanced rapidly, long-horizon execution may still progress to the next task stage before the required physical effect has been established. We call this a mismatch between semantic commitments, physical conditions that a stage must establish or maintain, and the actual physical state. Because an action command alone cannot confirm such a condition, local deviations can propagate and cause task failure. To address this problem, we present CommitFlow, a closed-loop execution framework that combines commitment monitoring with local correction while keeping the base policy frozen. CommitFlow integrates three components. A Semantic Commitment Monitor (SCM) compares stage requirements against current state evidence and holds back dependent actions when a required condition is unmet or violated. BoundaryFlow then generates a local correction conditioned on the current state and base action, and Relation and Gain Calibration (RGC) selects the smallest correction strength that satisfies the relevant constraints. Across the ten common RoboTwin 2.0 benchmark tasks, CommitFlow achieves a mean success rate of 75.9 percent, improving on the base policy pi0.5 by 22.7 percent. Cross-policy experiments show consistent gains, pointing toward reliable long-horizon robot execution.
Chinese Translation
尽管视觉-语言-动作(VLA)策略发展迅速,但在长时程执行中,任务仍可能在所需物理效果尚未建立时就推进到下一阶段。我们将此称为语义承诺(即各阶段必须建立或维持的物理条件)与实际物理状态之间的失配。由于仅凭动作指令无法确认此类条件,局部偏差可能不断传播并导致任务失败。为解决这一问题,我们提出了CommitFlow,一种将承诺监控与局部校正相结合的闭环执行框架,同时保持基础策略冻结。CommitFlow集成了三个组件:语义承诺监控器(Semantic Commitment Monitor, SCM)将阶段要求与当前状态证据进行比较,当所需条件未满足或被违反时阻止后续依赖动作;BoundaryFlow根据当前状态和基础动作生成局部校正;关系与增益校准(Relation and Gain Calibration, RGC)选择满足相关约束的最小校正强度。在RoboTwin 2.0基准的十个常见任务上,CommitFlow实现了75.9%的平均成功率,较基础策略pi0.5提升了22.7%。跨策略实验显示了一致的性能提升,为可靠的长时程机器人执行指明了方向。
cs.RO / 91 / 2609.21929

MAAP: Multi-Agent Active Perception for Collaborative Manipulation

MAAP:面向协作操作的多智能体主动感知
Chen, Bruno N. Y., Kang, Li, Zhou, Heng, Song, Xiufeng, Zhang, Zhemeng, Ma, Jiahua, Qin, Yiran
Abstract
Multi-agent manipulation naturally produces multiple task-driven viewpoints: every arm carries a wrist camera and moves through the scene while acting. Yet these observations are typically underutilized, and active perception in manipulation is still often treated as requiring a dedicated sensing agent. We introduce MAAP (Multi-Agent Active Perception), in which every arm is dual-purpose: it executes manipulation actions and, through the wrist camera it carries, simultaneously serves as a moving viewpoint for the team. We pair this with RAIL (Role-Aware Imitation Learning), a controller that predicts each arm's current role alongside its action chunk and conditions action generation on it, representing role-dependent actions within one network. Across four simulated tasks, widening the perception regime lifts average success from 56.5% with a fixed camera to 62.5% with one active wrist view and 70.0% with all of them, while MAAP+RAIL reaches 79.2%. RAIL's additional gain is concentrated on the three-arm Microwave task, where success rises from 47% to 82% on identical multi-wrist inputs. On a dual-arm platform, MAAP+RAIL succeeds in 14 of 20 placement trials compared with 0 of 20 for fixed-view ACT. Collaborative manipulation can thus serve as an active perception mechanism in its own right.
Chinese Translation
多智能体操作天然会产生多个任务驱动的视点:每条机械臂都携带腕部相机,并在执行动作的过程中移动于场景之中。然而,这些观测通常未被充分利用,操作中的主动感知仍常被视为需要一个专用的感知智能体。我们提出了MAAP(Multi-Agent Active Perception,多智能体主动感知),其中每条机械臂都具有双重用途:它既执行操作动作,又通过其携带的腕部相机同时充当团队的移动视点。我们进一步结合RAIL(Role-Aware Imitation Learning,角色感知模仿学习),一种在预测每条机械臂动作块的同时预测其当前角色,并以此角色为条件生成动作的控制器,从而在单一网络中表示依赖角色的动作。在四个仿真任务中,扩大感知范围使平均成功率从固定相机的56.5%提升到单一活动腕部视角的62.5%,以及利用全部视角时的70.0%,而MAAP+RAIL达到了79.2%。RAIL带来的额外增益集中在三臂微波炉任务上,在相同的多腕部输入下,成功率从47%提升至82%。在双臂平台上,MAAP+RAIL在20次放置试验中成功了14次,而固定视角的ACT在20次试验中成功0次。因此,协作操作本身就可以作为一种主动感知机制。
cs.RO / 92 / 2609.21942

When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence

失败机器人何时应当询问?基于审计传感器证据发起人机纠错对话
Pathak, Eshika, Krishna, Leela
Abstract
A robot that fails at a task faces the first decision in corrective dialogue: act on its own diagnosis, consult another onboard sensor, or interrupt a person. Choosing well requires knowing how much the robot's sensors reveal about the cause and how reliable the robot's own diagnosis is. We build a simulated benchmark in which every failure's true cause is known, because we injected it, and measure what each sensor reveals, with explicit checks against data leakage. Some failures are diagnosable from camera images; others only from the robot's force data (0.99 from force data, no image method above 0.55). We then test six open vision-language models. Their behavior tracks the surface of the prompt, not the evidence: moving the refusal option from last to first in the answer list collapses refusal rates from 78-100% to 0-6% in three of the six swept model-and-family pairs. Accuracy from frames stays at or below a majority-class baseline under every prompt variant, with or without worked examples, and stated confidence carries no information about correctness. Handing the same models the force data as ten lines of text produces the first above-baseline diagnoses, in four of the six models: much of the failure reflects missing sensor data, not missing ability. We pose the choice as a three-action decision problem, act, consult your own sensors, or ask a human, whose optimal policy follows from measured accuracy. The models do not follow it, and their ask rates ignore a fourfold change in question cost. One question to a human still lifts them from that baseline to roughly the answerer's own reliability (0.70-0.81 when they ask). The decision to ask should be tied to measured accuracy and stated costs, not to the model's confidence.
Chinese Translation
任务失败的机器人面临纠错对话中的第一个决策:依据自身诊断采取行动、查询机载的其他传感器,还是打断并询问人类。做出正确选择需要了解机器人传感器的信息能在多大程度上揭示失败原因,以及机器人自身诊断的可靠性。我们构建了一个仿真基准,其中每个失败的真实原因均为已知(因为我们主动注入该原因),并通过针对数据泄露的显式检查来度量每个传感器所揭示的信息。某些失败可以从相机图像中诊断出来;另一些则只能从机器人的力觉数据中诊断(力觉数据达到0.99,而任何图像方法均不超过0.55)。随后我们测试了六个开源视觉-语言模型。它们的行为跟随提示词的表面形式而非证据本身:在六个被扫描的模型-系列组合中,有三个组合在将拒绝选项从答案列表末尾移至开头后,拒绝率从78-100%骤降至0-6%。在所有提示变体下,无论有无示例演示,基于图像帧的准确率始终停留在多数类基线或以下,且模型所声称的置信度不包含任何关于正确性的信息。将相同的力觉数据以十行文本的形式提供给这些模型后,六个模型中的四个首次产生了高于基线的诊断:大量失败反映的是传感器数据的缺失,而非能力的缺失。我们将该选择形式化为一个三动作决策问题——采取行动、查询自身传感器,或询问人类——其最优策略可由测量到的准确率推导得出。模型并未遵循该策略,且其询问率无视问题成本高达四倍的变化。向人类提出一次询问仍能将模型从该基线提升至大致相当于回答者自身可靠性的水平(在其询问时为0.70-0.81)。是否询问的决策应当与测量到的准确率和声明的成本挂钩,而非依赖模型自身的置信度。
cs.RO / 93 / 2609.21948

GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments

GALA:面向跨具身视觉-语言-动作模型预训练的几何感知潜在动作建模
Liu, Yichen, Yuan, Puzhen, Zhu, Xiang, Guo, Yanjiang, Chen, Jianyu
Abstract
Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to capture fine-grained end-effector articulation, particularly finger-level geometric changes in human and dexterous robot hands. To address this limitation, we propose GALA, a Geometry-Aware Latent-Action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. However, naively incorporating point clouds yields fine-grained action representations with limited shared semantics, hindering cross-embodiment pretraining. To address this issue, we introduce the Unified End-effector Motion Representation (UEMR), which preserves fine-grained motion information while improving the cross-embodiment generalizability of latent actions. Building upon UEMR, GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos. Experiments on fine-grained motion probing, cross-embodiment retrieval, and downstream VLA evaluation demonstrate GALA's effectiveness in modeling generalizable fine-grained motions across embodiments, achieving 68.3% RoboCasa-GR1 success rate and 75.5% real-world success rate. Code, appendix, and demos are available at https://puzhenyuan.github.io/GALA-website/.
Chinese Translation
由于不同末端执行器之间的动作空间存在异构性,从多具身数据集中学习大规模视觉-语言-动作(VLA)模型仍然极具挑战性。尽管潜在动作模型(Latent Action Models, LAMs)能够从多样化的视频数据中学习与具身无关的动作表示,但现有的基于图像的LAM往往无法捕捉细粒度的末端执行器关节运动,尤其是人手和灵巧机器手的指级几何变化。为解决这一局限,我们提出了GALA,一个几何感知的潜在动作建模框架,它通过三维末端执行器几何运动来增强基于图像的潜在动作。然而,直接融合点云会得到共享语义有限的细粒度动作表示,从而阻碍跨具身预训练。为解决这一问题,我们引入了统一末端执行器运动表示(Unified End-effector Motion Representation, UEMR),它在保留细粒度运动信息的同时,提升了潜在动作的跨具身泛化能力。基于UEMR,GALA将捕捉场景级动态的视觉潜在动作与捕捉共享细粒度末端执行器关节运动的几何潜在动作相结合,为从多具身数据(包括无动作标注的第一视角人类视频)中进行VLA预训练提供了有效的监督信号。在细粒度运动探测、跨具身检索以及下游VLA评估上的实验表明,GALA在跨具身可泛化细粒度运动建模方面表现出色,在RoboCasa-GR1上取得了68.3%的成功率,在真实世界场景中取得了75.5%的成功率。代码、附录和演示可在 https://puzhenyuan.github.io/GALA-website/ 获取。
cs.RO / 94 / 2609.21982

CARF: Contrastive Attraction-Repulsion of Failure-Guided Flow Matching

CARF:基于失败引导流匹配的对比吸引-排斥框架
Zhao, Shuqi, Du, Bang, Wu, Cheng-En, Xie, Yichen, Wang, Yixiao, Tomizuka, Masayoshi
Abstract
Robot demonstration collection often produces imperfect or failed trajectories in addition to successful demonstrations. Existing methods typically exploit failed trajectories by identifying segments that still make progress toward task completion, but largely overlook \textit{failure-critical behaviors} that directly lead to task failure. Here we argue that these two types of segments provide fundamentally asymmetric supervision: progressive segments should be imitated, whereas failure-critical segments should be explicitly avoided. Based on this observation, we propose CARF, a Contrastive Attraction-Repulsion of Failure-guided framework for learning from imperfect robot data. CARF introduces a progress-based importance scorer, trained solely on successful expert demonstrations and its perturbation results, to estimate step-wise contributions toward task completion and identify informative regions in failed trajectories. These scores guide a unified flow-matching objective that attracts the policy toward progressive behaviors and repels it from failure-critical ones, while excluding ambiguous segments. This enables more comprehensive utilization of imperfect data and avoids unreliable supervision from ambiguous failure segments. Extensive experiments in simulation and the real world demonstrate consistent improvements over competing baselines across diverse failure scenarios, with ablations further validating the effectiveness of the proposed scoring and attraction-repulsion mechanisms. Our website is https://zhao-sq.github.io/carf/#.
Chinese Translation
机器人示教数据采集过程中,除了成功的示教之外,往往还会产生不完美或失败的轨迹。现有方法通常通过识别仍然朝任务完成方向取得进展的片段来利用失败轨迹,但在很大程度上忽视了直接导致任务失败的“失败关键行为”。本文认为,这两类片段提供了根本不对称的监督信号:进步性片段应当被模仿,而失败关键片段则应当被明确规避。基于这一观察,我们提出了CARF,一种用于从不完美机器人数据中学习的失败引导对比吸引-排斥框架。CARF引入了一个基于进退程度的重要性评分器,该评分器仅使用成功的专家示教及其扰动结果进行训练,用于估计各步骤对任务完成的贡献,并识别失败轨迹中的信息性区域。这些评分指导一个统一的流匹配目标,使策略被吸引向进步性行为,同时被排斥远离失败关键行为,并排除模糊片段。这使得不完美数据得到更全面的利用,并避免了来自模糊失败片段的不可靠监督。大量仿真与真实世界实验表明,CARF在多种失败场景下均取得了一致的性能提升,优于竞争基线方法;消融实验进一步验证了所提出的评分机制和吸引-排斥机制的有效性。我们的网站是 https://zhao-sq.github.io/carf/#。
cs.RO / 95 / 2609.21983

SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation

SkelWAM:一种用于零样本跨机器人本体操作的中骨骼引导世界-动作模型
Niu, Pengjun, Xie, Yujia, Peng, Rui, Zhao, Hang, Liu, Ke
Abstract
Reusing manipulation experience across robot embodiments is important for scaling robot learning and reducing repeated task-specific data collection. However, changes in embodiment alter visual appearance, action dimensionality and semantics, and the whole-body configurations that can realize the same tool pose. We present SkelWAM, a skeleton-guided world-action model that couples perception and control through one explicit geometric representation for single-source cross-embodiment manipulation. Arm centerline geometry, tool-center-point (TCP) pose, and parallel-jaw commands form a shared 25-D state. The same definition underlies canonical third-person and wrist observations and future whole-body action targets. Trained with predictive visual supervision, a video-action mixture of transformers predicts canonical skeleton action chunks, which embodiment-specific constrained decoders convert into joint or continuum-robot controls. This formulation requires no one-to-one joint correspondence and uses no target-task demonstrations or target policy updates. We introduce LIBERO-Cross10, a source-only cross-embodiment transfer benchmark covering ten tasks and ten target embodiments across four morphological groups. On this benchmark, Franka-trained SkelWAM achieves 43.3% success over 1,000 episodes, exceeding the best-performing evaluated baseline by 36.2 percentage points. We further deploy a JAKA mini2-trained policy on the Feagine A03 continuum robot for three tabletop manipulation tasks, illustrating the approach's potential for real-world cross-embodiment manipulation. Project page: http://www.liukepku.com/skelwam/index.html
Chinese Translation
在不同机器人本体之间复用操作经验对于扩大机器人学习规模和减少重复的任务专用数据采集至关重要。然而,机器人本体的变化会改变视觉外观、动作维度与语义,以及能够实现同一工具位姿的全身构型。我们提出 SkelWAM,一种骨架引导的世界-动作模型,通过一种显式几何表示将感知与控制耦合起来,实现单源跨机器人本体操作。机械臂中心线几何、工具中心点(TCP)位姿以及平行夹爪指令构成一个共享的 25 维状态。同一状态定义构成了规范化的第三人称与腕部观测以及未来全身动作目标的基础。该模型采用预测性视觉监督进行训练,通过视频-动作混合 Transformer(video-action mixture of transformers)预测规范化骨架动作块,再由针对特定本体的约束解码器将其转换为关节或连续体机器人控制。该方案无需一一对应的关节映射,也不需要目标任务演示或目标策略更新。我们引入 LIBERO-Cross10,一个仅使用单一源本体的跨机器人本体迁移基准,涵盖四个形态类别下的十个任务和十个目标本体。在该基准上,由 Franka 训练的 SkelWAM 在 1,000 个回合中取得 43.3% 的成功率,比表现最佳的评估基线高出 36.2 个百分点。我们进一步将 JAKA mini2 训练得到的策略部署到 Feagine A03 连续体机器人上,完成三个桌面操作任务,展示了该方法在真实世界跨机器人本体操作方面的潜力。项目主页:http://www.liukepku.com/skelwam/index.html
cs.RO / 96 / 2609.22062

Gripper-Aware Automatic Dense Packing of Irregular Objects

面向夹爪感知的规则物体自动密集排布?——夹爪感知的不规则物体自动密集排布
Qin, Tianhao, McCann, Connor, Calli, Berk, Xiao, Jing
Abstract
Automatic dense packing is widely desired in warehouse operations but remains a fundamental challenge in robotic manipulation. Existing work on irregular-object packing largely targets simulation with idealized contact, treating the object as an isolated rigid body. The gripper often enters as a discrete, post-hoc feasibility check, if considered at all, and the perception and contact drift accumulated during execution are not addressed. We present a closed-loop pipeline that integrates perception, gripper-aware placement optimization, and force-guided execution on a real manipulator. The optimizer represents the object together with the gripper as a single composite body of hierarchical sphere trees. It searches over five degrees of freedom on a GPU within a CMA-ES framework, with the vertical coordinate grounded analytically against the current heightmap. During execution, a force-monitored vertical descent stops on first contact. A post-release consolidation push then closes the residual lateral clearance that gripper-aware planning leaves behind. The container is re-perceived between placements so that drift does not accumulate. We validate the system on a Franka Emika Panda robot packing a 3D-printed set of flat, curved, and concave objects, and a YCB object subset. An ablation study isolates the contribution of gripper-aware optimization, the consolidation push, and mesh-derived geometry to end-to-end success, achieved density, and computational cost. We further benchmark against the heightmap-minimization method as a baseline representative of prior irregular-object packing work.
Chinese Translation
自动密集排布在仓储作业中应用需求广泛,但仍是机器人操作领域的一项根本性挑战。现有针对不规则物体排布的研究大多局限于理想接触条件的仿真环境,将物体视为孤立的刚体。夹爪即使被考虑,也往往仅作为事后的离散可行性检查,而执行过程中累积的感知误差与接触漂移问题也未得到解决。我们提出一个闭环流程,在真实机械臂上集成了感知、夹爪感知的放置优化以及力引导执行。该优化器将物体与夹爪表示为一个由分层球树(hierarchical sphere trees)构成的整体复合体,并在CMA-ES框架下于GPU上搜索五个自由度,其中竖直坐标通过相对于当前高度图的解析方法确定。执行过程中,力监测的竖直下降在首次接触时停止。释放后的 consolidation push(巩固推挤)可消除夹爪感知规划所遗留的残余侧向间隙。在每次放置之间重新感知容器,从而避免漂移累积。我们在一台Franka Emika Panda机器人上验证了该系统,所用物体包括一组3D打印的扁平、弯曲及凹形物体,以及YCB物体数据集的一个子集。消融实验分别验证了夹爪感知优化、巩固推挤以及基于网格的几何表示对端到端成功率、所达密度和计算成本的贡献。我们进一步以高度图最小化方法作为基线(代表先前的不规则物体排布工作)进行了对比测试。
cs.RO / 97 / 2609.22073

Duty Factor Predicts Robust Constrained Quadrupedal Locomotion Across Gait Types

占空比可预测跨步态类型的鲁棒受约束四足运动
Zhu, James, Ologan, David, Ortiz, George, Lee, Thomas Chun Fai, Gonzalez, Selvin Garcia, Tajbakhsh, Ardalan, Ben-Tzvi, Pinhas, Johnson, Aaron M.
Abstract
Quadrupedal robots are increasingly deployed in environments where locomotion must remain robust to disturbances and constrained terrain. Gait type, such as walking or trotting, is commonly used to characterize quadrupedal locomotion. However, gait type does not uniquely define locomotion, as parameters such as duty factor, speed, and stance width can vary within a single gait type. In this work, we investigate the relationship between these gait parameters using three distinct quadrupedal locomotion control approaches. First, using whole body trajectory optimization with LQR feedback, we show that duty factor is a stronger predictor of local error convergence than nominal gait type. Second, we investigate duty factor selection with a learned locomotion controller, suggesting how duty factor may serve as a low-dimensional parameter for adapting locomotion robustness in narrow-terrain environments. Finally, we show that these trends persist under a centroidal model predictive control framework and validate them through narrow-terrain experiments on a physical quadruped. These results show that duty factor provides a simple and effective basis for understanding and selecting robust quadrupedal locomotion across gait types and control architectures.
Chinese Translation
四足机器人越来越多地被部署在运动必须对干扰和受限地形保持鲁棒性的环境中。步态类型(如行走或小跑)通常被用于刻画四足运动。然而,步态类型并不能唯一地定义运动,因为占空比、速度和支撑宽度等参数在单一 gait 类型内仍可变化。在本工作中,我们采用三种不同的四足运动控制方法来研究这些步态参数之间的关系。首先,利用带LQR反馈的全身轨迹优化,我们表明占空比比名义步态类型更能有效预测局部误差收敛性。其次,我们研究了基于学习型运动控制器的占空比选择,说明占空比可作为在狭窄地形环境中调整运动鲁棒性的低维参数。最后,我们证明这些趋势在质心模型预测控制框架下依然成立,并通过实体四足机器人上的狭窄地形实验对其进行了验证。这些结果表明,占空比为理解和选择跨步态类型及控制架构的鲁棒四足运动提供了一个简单而有效的基础。
cs.RO / 98 / 2609.22075

LIMBO: Learning and Internalizing Model-Free Barrier Objectives for Agile and Safe Whole-Body Control

LIMBO:学习并内化无模型障碍目标以实现敏捷且安全的全身控制
Gonzales, Jake, Alvarez, Arturo Flores, Chen, Yu-Ming, Ames, Aaron D., Ratliff, Lillian J., Nambi, Manikantan
Abstract
Safe whole-body control requires coordinating collision avoidance and balance under high-dimensional, nonlinear dynamics--making safety certificates difficult to design and reuse across behaviors. We present LIMBO, a framework for synthesizing a state-action control barrier function and distilling its safety structure into a task policy. LIMBO learns the safety certificate from black-box transitions and a state-based failure specification over residual actions around a frozen base controller, making Q-CBF synthesis tractable in the full control dimension while placing the certificate in the task policy's control space. During synthesis, the learned safety value drives risk-guided sampling near the estimated boundary of recoverability; during task learning, it serves as a teacher that provides action-level safety feedback, yielding a robust task policy and alleviating the need for an online safety filter at deployment. We demonstrate LIMBO on a 29-degree-of-freedom humanoid performing dodgeball avoidance and locomotion beneath low obstacles. Beyond scaling learned Q-CBFs to whole-body control, we show that risk-guided boundary sampling provides a theoretically grounded way to explore the edge of recoverability. Under the same safety specification, ceteris paribus, varying the sampling concentration produces strategies ranging from crouching to a novel backward-leaning limbo maneuver. In both settings, the learned policies transfer to hardware without online safety filtering, showing that learned safety synthesis scales to agile whole-body control.
Chinese Translation
安全的全身控制需要在高维、非线性动力学下协调碰撞避免与平衡,这使得安全证书难以设计,也难以在不同行为之间复用。我们提出了LIMBO,这是一个用于合成状态-动作控制屏障函数(control barrier function)并将其安全结构蒸馏到任务策略中的框架。LIMBO从黑盒转移数据以及基于状态、定义在冻结基础控制器周围残差动作上的失败规范中学习安全证书,使得Q-CBF的合成在整个控制维度上变得可处理,同时将安全证书置于任务策略的控制空间中。在合成阶段,学习到的安全价值函数驱动风险引导采样以逼近估计的可恢复性边界;在任务学习阶段,它充当教师,提供动作级别的安全反馈,从而得到鲁棒的任务策略,并免去了部署时在线安全滤波器的需要。我们在一个29自由度的人形机器人上展示了LIMBO,执行躲避飞球(dodgeball)和在低矮障碍物下移动的任务。除了将学习到的Q-CBF扩展到全身控制之外,我们还表明,风险引导的边界采样为探索可恢复性边缘提供了一种有理论依据的方法。在相同的安全规范下,在其他条件不变的情况下,改变采样集中度可以产生从下蹲到一种新颖的后仰过障(limbo)动作等多种策略。在这两类任务中,学习到的策略均能在无需在线安全滤波的情况下迁移到真实硬件上,表明学习式安全合成可以扩展到敏捷的全身控制。
cs.RO / 99 / 2609.22085

SeeQ: Training Generalist Value Functions for Long-Horizon Robotic Manipulation

SeeQ:面向长时程机器人操作的通用价值函数训练
Singh, Saksham, Hu, Zheyuan, Mark, Max Sobol, Yu, Jeffrey, Erickson, Zackory, Kumar, Aviral
Abstract
Despite rapid progress, generalist robot policies remain brittle on complex, long-horizon tasks that comprise multiple stages or require repeated attempts and deliberation on the same underlying stage before success. Q-value functions can improve these policies by ranking candidate actions or guiding policy improvement, but learning from sparse task-level rewards entails long credit-assignment horizons, difficult Bellman backups, and broad data-coverage requirements. We introduce SeeQ (Subtask-elicited Q-functions), which instead learns Q-values for the currently active subtask. This shortens the value-prediction horizon and enables effective learning with temporal-difference (TD) objectives. During training, subtask-level annotations present in offline robot data provide the decomposition and enable learning from broad, potentially suboptimal robot datasets. To eliminate the need for human annotations or modular subtask prediction systems at test time, our Q-function architecture is trained to autoregressively predict the active subtask in natural language before estimating its value. We instantiate SeeQ using a base vision-language backbone, pretrain it on diverse open-source robot manipulation data, and finetune it on downstream tasks. Across four real-world manipulation tasks on two bimanual robot platforms, the SeeQ value function substantially improves best-of-N policy steering.
Chinese Translation
尽管取得了快速进展,通用机器人策略在复杂的长时程任务上仍然脆弱,这些任务包含多个阶段,或需要在同一底层阶段上进行多次尝试和深思熟虑才能成功。Q值函数可以通过对候选动作进行排序或引导策略改进来提升这些策略,但从稀疏的任务级奖励中学习需要面临长时程的信用分配、困难的Bellman回溯更新以及广泛的数据覆盖要求。我们提出了SeeQ(Subtask-elicited Q-functions,子任务引发的Q函数),它转而学习当前活跃子任务的Q值。这缩短了价值预测的时间跨度,使得基于时序差分(TD)目标的有效学习成为可能。在训练过程中,离线机器人数据中存在的子任务级标注提供了任务分解,并使模型能够从广泛且可能次优的机器人数据集中学习。为了消除测试时对人工标注或模块化子任务预测系统的需求,我们的Q函数架构被训练为在估计价值之前,以自回归方式用自然语言预测当前活跃的子任务。我们基于一个基础的视觉-语言骨干网络实现SeeQ,在多样化的开源机器人操作数据上进行预训练,并在下游任务上进行微调。在两个双臂机器人平台上的四个真实世界操作任务中,SeeQ价值函数显著提升了best-of-N策略引导的性能。
人工智能 (Artificial Intelligence)
49
cs.AI / 1 / 2609.20971

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

RBS-Attention:面向长上下文大语言模型的半径有界稀疏预填充方法
Song, Chuxu, Wei, Jiuqi, Peng, Zhencan
Abstract
Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly relevant token among many irrelevant ones. We call this failure mode mean dilution and propose RBS-Attention, a training-free sparse-prefill method with two complementary selection branches. A centroid base branch captures average relevance, while a rescue branch uses the maximum key-block radius and its prompt-, layer-, and head-dependent distribution to identify blocks at risk of underestimation. Independently thresholding the two branches and combining their masks controls the contribution of rescue blocks while preserving regular block-sparse FlashAttention execution. On H100 GPUs, RBS-Attention achieves 20.65$\times$ standalone prefill-attention speedup, 11.92$\times$ vLLM prefill-attention speedup, and 5.97$\times$ end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8. On the dense Qwen3-32B model, it obtains 88.65 overall RULER accuracy versus 89.52 for dense attention; LongBench-v2, InfiniteBench, and Video-MME provide additional quality evaluation. Supporting experiments measure actual retention, compare selectors at matched density, and characterize block-size, threshold, and memory behavior. Together, these results support radius-adaptive dual-branch selection as an effective approach to long-context prefill.
Chinese Translation
长上下文大语言模型推理日益受限于预填充阶段,该阶段中稠密自注意力机制需要在生成开始前处理整个提示词。稀疏块选择可以降低这一开销,但一个块的质心可能会掩盖众多无关标记中高度相关的标记。我们将这种失效模式称为均值稀释(mean dilution),并提出 RBS-Attention——一种无需训练的稀疏预填充方法,包含两个互补的选择分支。质心基础分支捕捉平均相关性,而救援分支则利用最大键块半径及其依赖于提示词、层和注意头的分布来识别可能被低估的块。对两个分支独立设置阈值并合并其掩码,可以控制救援块的贡献,同时保持常规的块稀疏 FlashAttention 执行方式。在 H100 GPU 上,RBS-Attention 在 Qwen3-30B-A3B-Instruct-2507-FP8 模型、128K 上下文长度下,实现了 20.65 倍的独立预填充注意力加速、11.92 倍的 vLLM 预填充注意力加速,以及 5.97 倍的端到端首令牌生成时间(time-to-first-token)加速。在稠密模型 Qwen3-32B 上,其 RULER 整体准确率为 88.65,而稠密注意力为 89.52;LongBench-v2、InfiniteBench 和 Video-MME 提供了额外的质量评估。辅助实验测量了实际保留率、在相同密度下比较了不同选择器,并刻画了块大小、阈值和内存方面的行为特性。这些结果共同表明,半径自适应的双分支选择是长上下文预填充的一种有效方法。
cs.AI / 2 / 2609.20974

Attention-Aware Routing: Coupling Routing and Attention in MoEs

注意力感知路由:在混合专家模型(MoE)中耦合路由与注意力机制
Kosmopoulou, Despoina, Tsetsilas, Anastasios, Georgiou, Efthymios, Karamanolakis, Giannis, Roy, Swastik, Potamianos, Alexandros
Abstract
In Mixture-of-Experts language models, the router typically selects and weights experts based on the token's hidden state, utilizing limited contextual information. We propose Attention-Aware Routing (AAR), which augments the router with temporal and spectral features extracted from a sliding window of attention weights that represent a summary of the model's contextual state, disentangled from the hidden state. Keeping the base transformer entirely frozen, we train only the routing parameters, isolating routing as the sole variable. AAR improves GSM8K by +3.37 pp over a routing-only SFT baseline on OLMoE. Beyond performance, we show that routing and attention form a coupled circuit: routing changes at layer l propagate through the residual stream to amplify attention sinks at layer l+1, reshaping attention without any direct update to the attention mechanism itself. Further, AAR reduces long diverging generation, with incorrect answers getting shorter, while correct answers remain unchanged in length. Finally, AAR is strongly depth-sensitive: applying it indiscriminately across layers can degrade factual retrieval, whereas mathematical reasoning gains persist when it is introduced deeper in the network. This sensitivity exposes a retrieval--reasoning tension across depth and makes layer-selective AAR a controlled probe of the routing-relevant information carried by attention at different layers.
Chinese Translation
在混合专家(Mixture-of-Experts)语言模型中,路由器通常基于词元的隐藏状态来选择专家并为其分配权重,所利用的上下文信息十分有限。我们提出注意力感知路由(Attention-Aware Routing, AAR),通过从注意力权重的滑动窗口中提取时域和频域特征来增强路由器,这些特征代表了模型上下文状态的摘要,并与隐藏状态解耦。在保持基础Transformer完全冻结的情况下,我们仅训练路由参数,从而将路由隔离为唯一的变量。在OLMoE上,AAR相较于仅路由的SFT基线,在GSM8K上提升了3.37个百分点。除性能提升外,我们还发现路由与注意力构成了一个耦合回路:第l层的路由变化经由残差流传播,放大第l+1层的注意力汇聚(attention sinks)现象,从而在注意力机制本身没有任何直接更新的情况下重塑了注意力。此外,AAR减少了发散式的长篇生成,错误答案的长度变短,而正确答案的长度保持不变。最后,AAR对网络深度高度敏感:不加区分地将其应用于所有层可能会损害事实检索能力,而在网络较深处引入AAR时,数学推理能力的提升仍然持续存在。这种敏感性揭示了不同深度上检索与推理之间的张力,并使层选择性的AAR成为探测不同层注意力所携带的路由相关信息的一种受控手段。
cs.AI / 3 / 2609.20981

CaLR: Causal Latent Revision for Robust Diffusion Reasoning

CaLR:面向鲁棒扩散推理的因果潜在修正方法
Cai, Wei, Zhao, Jian, Yuan, Yuchen, Li, Xuelong
Abstract
Autoregressive (AR) models suffer from local greediness, while diffusion language models (DLMs) often lack the strict causal structure required for reasoning. To combine the advantages and overcome the drawbacks of the dual, we propose Causal Latent Revision (CaLR), a framework that reformulates reasoning as constrained latent optimization. By adopting a causal topology matrix (CTM) from an expert model and implicit differentiation, CaLR performs gradient-guided ``thought revision" to enforce logical consistency, enabling dynamic self-correction of intermediate steps during parallel generation. Empirically, CaLR achieves SOTA DLM performance on complex benchmarks, surpassing strong AR baselines and demonstrating superior robustness in constrained tasks like Sudoku.
Chinese Translation
自回归(AR)模型存在局部贪婪性问题,而扩散语言模型(DLM)往往缺乏推理所需的严格因果结构。为了结合二者的优势并克服各自的缺陷,我们提出了因果潜在修正(Causal Latent Revision, CaLR)框架,该框架将推理重新表述为受约束的潜在空间优化问题。通过采用来自专家模型的因果拓扑矩阵(Causal Topology Matrix, CTM)并结合隐式微分,CaLR 执行基于梯度引导的"思维修正",以确保逻辑一致性,从而在并行生成过程中实现对中间步骤的动态自我修正。实验结果表明,CaLR 在复杂基准测试上取得了 DLM 的最先进(SOTA)性能,超越了强大的 AR 基线方法,并在数独等受限任务中展现出卓越的鲁棒性。
cs.AI / 4 / 2609.21061

LoRA Enhanced Contrastive Learning with SAS Vision Transformers

基于SAS视觉Transformer的LoRA增强对比学习
Zimmerman, Dan, Bobe III, Frank E., McCormack, Amelia L., Cook, Matthew, Vetaw, Gregory D.
Abstract
Automatic target recognition (ATR) with synthetic aperture sonar (SAS) supports advanced naval capabilities, but deep learning is constrained by scarce target imagery, background clutter, and human-in-the-loop assessment. We adapt DINOv3 Vision Transformer (ViT) models to underwater SAS ATR using a three-stage parameter-efficient framework. Stage 1 uses Low-Rank Adaptation (LoRA) while freezing the ViT backbone, bridging the gap between natural-image pretraining and underwater acoustic propagation. Stage 2 uses hard-negative mining to strengthen the decision boundary against acoustic mimics, including rocks and sediment formations resembling man-made targets. Stage 3 uses Supervised Contrastive Learning (SupCon) to separate target and clutter representations. We evaluate at-sea SAS data using a mission-level geographic split, compare all arms at 85 percent test recall, and repeat each comparison over three random seeds. LoRA accounts for the primary effect, increasing area under the precision-recall curve (AUPRC) from 0.300 to 0.679 +/- 0.027 using the same frozen backbone. Rank 4 achieves this result while training only 0.26 percent of weights. Neither refinement stage exceeds its matched control: hard-negative mining changes AUPRC by -0.0045 +/- 0.0119 versus an equal-size random curriculum, and SupCon changes AUPRC by +0.0002 +/- 0.0096 versus the preceding stage. These null results indicate that mining occurred on data the encoder had already fit and that supervised stages had already imposed most target-clutter geometry. One efficient adaptation stage is sufficient; stacked refinement is not.
Chinese Translation
基于合成孔径声呐(SAS)的自动目标识别(ATR)可支持先进的海军能力,但深度学习受限于稀缺的目标图像、背景杂波以及人在回路的评估。我们采用三阶段参数高效框架,将DINOv3视觉Transformer(ViT)模型适配于水下SAS自动目标识别。第一阶段在冻结ViT主干网络的同时使用低秩适配(LoRA),以弥合自然图像预训练与水下声传播之间的差距。第二阶段使用难负样本挖掘来强化针对声学拟态目标的决策边界,这些拟态目标包括与人工目标相似的岩石和沉积物构造。第三阶段使用监督对比学习(SupCon)来分离目标与杂波的表示。我们使用任务级地理划分对海上实测SAS数据进行评估,在85%的测试召回率下比较所有方案,并对每次比较重复三个随机种子。LoRA是主要效果的来源,在相同的冻结主干网络下将精确率-召回率曲线下面积(AUPRC)从0.300提升至0.679 ± 0.027。秩为4即可实现该结果,且仅训练0.26%的权重。两个细化阶段均未超过其对应的对照组:难负样本挖掘相对于同等规模的随机课程使AUPRC变化为-0.0045 ± 0.0119,SupCon相对于前一阶段使AUPRC变化为+0.0002 ± 0.0096。这些无效结果表明,挖掘发生在编码器已经拟合的数据上,且监督阶段已经确立了大部分目标-杂波几何结构。一个高效的适配阶段即已足够;叠加细化并无必要。
cs.AI / 5 / 2609.21096

Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing

检测大语言模型中的幻觉:追踪受损上下文共享的拓扑特征
Jalilifard, Amir, Rocha, Anderson, Wong, Eric, Raimundo, Marcos Medeiros
Abstract
In this work, we examine the topology of information flow patterns within attention graphs to effectively distinguish hallucinated from non-hallucinated responses. We analyze the Forman-Ricci curvature to identify structural patterns indicating information bottlenecks in attention graphs. We then introduce a method that captures both semi-local and global information-flow characteristics of attention heads associated with hallucinated responses. We evaluate our approach extensively across several LLMs and established benchmarks. Empirical results demonstrate that our proposed single-pass approach provides consistent improvements over existing attention-based and multi-response baselines across two hallucination-detection benchmarks, while achieving competitive performance across diverse LLM architectures. Further analysis reveals that impaired context sharing among tokens during causal generation is strongly associated with hallucination occurrences in LLMs. In particular, hallucinated responses are consistently characterized by an over-reliance on self-attention, diffused context retrieval from earlier tokens, or information over-squashing, especially in the final transformer layer.
Chinese Translation
在本工作中,我们通过考察注意力图(attention graphs)中信息流模式的拓扑结构,有效地区分幻觉响应与非幻觉响应。我们分析了Forman-Ricci曲率,以识别注意力图中指示信息瓶颈的结构模式。随后,我们提出了一种方法,能够同时捕捉与幻觉响应相关的注意力头的半局部和全局信息流特征。我们在多个大语言模型(LLM)和权威基准上对该方法进行了广泛评估。实验结果表明,在两个幻觉检测基准上,我们提出的单次遍历(single-pass)方法相较于现有的基于注意力和多响应的基线方法均取得了一致的改进,同时在多种大语言模型架构上实现了具有竞争力的性能。进一步分析揭示,在因果生成过程中,词元(token)间上下文共享的受损与大语言模型中幻觉的出现密切相关。特别是,幻觉响应始终表现出以下特征:过度依赖自注意力(self-attention)、从先前词元检索上下文时的扩散化,或信息过度压缩(over-squashing),尤其是在Transformer的最后一层中。
cs.AI / 6 / 2609.21113

Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

微调大语言模型中内部表征变化与因果重要性的解耦
Li, Lingfang, Sen, Procheta, Das, Shubham, Bollegala, Danushka
Abstract
Fine-tuning has emerged as a widely adopted approach for adapting LLMs to a variety of downstream tasks. However, how it reshapes their internal mechanisms remains poorly understood. To address this, we investigate how fine-tuning alters internal representations in LLMs, including attention patterns and layer-wise activations, and examine whether these changes are linked to task-relevant components identified by EAP (e.g., attention heads and logit-level activations) that drive task performance. We find that EAP-identified components are concentrated within specific layers, indicating a degree of functional localisation in how models internalise task-specific behavior. Notably, the distribution of these components across layers is largely uncorrelated with the layers undergoing the most substantial representational changes during fine-tuning. Furthermore, we observe that overlap in EAP-identified components across tasks does not translate into cross-task performance transfer if the tasks are different in nature (e.g. classification vs. generative tasks). More specifically, fine-tuning on one task can lead to a degradation of performance on another when the two tasks exhibit a high degree of overlap in their EAP-identified components.
Chinese Translation
微调已成为将大语言模型(LLM)适配到多种下游任务的广泛采用的方法。然而,微调如何重塑其内部机制仍然知之甚少。为此,我们研究了微调如何改变大语言模型的内部表征,包括注意力模式和逐层激活,并考察这些变化是否与通过EAP(归因修补,Edge Attribution Patching)识别出的驱动任务性能的任务相关组件(如注意力头和logit层面的激活)相关联。我们发现,被EAP识别的组件集中于特定层,表明模型在内部化任务特定行为时存在一定程度的功能局域化。值得注意的是,这些组件在层间的分布与微调期间发生最显著表征变化的层基本不相关。此外,我们观察到,不同任务间EAP识别组件的重叠并不必然带来跨任务性能迁移,前提是任务在性质上不同(例如分类任务与生成式任务)。更具体地说,当两个任务的EAP识别组件高度重叠时,在一个任务上的微调可能导致在另一个任务上的性能下降。
cs.AI / 7 / 2609.21139

TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers

TinyCeNN-LM:基于质量门控的预训练注意力转换方法,采用受CeNN启发的细胞递归层
Mohsenzadegan, Kabeh, Tavakkoli, Vahid, Kyamakya, Kyandoghere
Abstract
Replacing attention in a pretrained language model is a compatibility problem: a plausible substitute may alter representations expected by later layers. TinyCeNN-LM introduces a \emph{quality-gated post-training conversion} framework using CeNN-inspired cellular-recurrent layers with bounded local processing, compact recurrent memory, routing, fusion, and accept-or-rollback validation. Three implementations are studied: Integrated Memory, MemoryFusion, and PDelta3-GDN2-CLVR+Local32. Strict PDelta3 conversion accepts a layer only when representation and NLL criteria pass fixed thresholds. On SmolLM2-135M, layers 0-2 are accepted with cumulative $\Delta\mathrm{NLL}=+0.01209$, while layer 3 is rejected despite acceptable NLL because representation fidelity fails. On Qwen3.5-0.8B, full-attention layers 3, 7, and 11 are accepted with final $\Delta\mathrm{NLL}=+0.02073$. Integrated Memory keeps perplexity within $-0.07\%$ to $+0.93\%$ while reducing total cache by up to $6.01\%$. A sampled 200-item downstream sanity check gives $28.5\%$--$32.0\%$ overall accuracy for converted Qwen releases. The results support conservative, quality-gated structural conversion rather than universal attention replacement or speedup.
Chinese Translation
在预训练语言模型中替换注意力机制是一个兼容性问题:一个看似合理的替代方案可能会改变后续层所依赖的表示。TinyCeNN-LM 提出了一种质量门控的训练后转换框架,采用受 CeNN(细胞神经网络)启发的细胞递归层,具备有界局部处理、紧凑递归记忆、路由、融合以及接受或回滚验证机制。本文研究了三种实现:Integrated Memory、MemoryFusion 和 PDelta3-GDN2-CLVR+Local32。严格的 PDelta3 转换只有在表示和 NLL(负对数似然)指标均通过固定阈值时才接受某一层的转换。在 SmolLM2-135M 上,第 0-2 层被接受,累积 ΔNLL=+0.01209,而第 3 层尽管 NLL 可接受,但因表示保真度未达标而被拒绝。在 Qwen3.5-0.8B 上,第 3、7、11 层的全注意力被成功转换,最终 ΔNLL=+0.02073。Integrated Memory 在将总缓存减少最多 6.01% 的同时,使困惑度保持在 -0.07% 至 +0.93% 的范围内。一项抽样 200 个条目的下游健全性检查显示,转换后的 Qwen 版本整体准确率为 28.5% 至 32.0%。这些结果支持采用保守的、质量门控的结构转换,而非普遍的注意力替换或加速方案。
cs.AI / 8 / 2609.21149

Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake

基于临床医生的AI辅助精神科问诊质量保障
Shi, King, Li, Amanda, Ivey, Jonathan, Wang, Synthia Qia, Gui, Guan, Kim, Hyunseo, Zandi, Peter, Straub, Jason, Taylor, Jacob, Joshi, Ananya
Abstract
Before patients can use AI-assisted psychiatric intake systems, health systems need practical ways to routinely evaluate these tools against their clinical standards for quality assurance. Because clinicians may use different intake styles, evaluation for this task must (1) support comparison across interviewing approaches, (2) minimize clinician burden, and (3) measure clinically relevant performance for health systems deploying these technologies. We present a clinician-grounded evaluation platform built around a memory-augmented patient simulator for open-ended AI interviewing, InterviewPlayground. We created interactive patients using InterviewPlayground with our expert-authored vignettes, constructed a simulated intake platform for the interviews, and designed evaluation modalities relevant to intake. In a pilot of 6 clinicians in a 25-minute assessment compared to a GPT-based LLM intake interviewer, the LLM recovered more of the clinically relevant items embedded in the patient vignettes (88.0% vs. 38.9%), but made more clinical inferences not based on the interview (56.8% vs. 27.8%), and characterized identified safety concerns less often (33.3% vs. 66.7%), setting the stage for deployed quality assurance for this task.
Chinese Translation
在患者使用AI辅助精神科问诊系统之前,医疗机构需要切实可行的方法,将这些工具与自身的临床标准进行常规化评估以实现质量保障。由于临床医生可能采用不同的问诊风格,针对该任务的评估必须:(1)支持跨访谈方式的比较,(2)尽量减少临床医生的负担,(3)衡量对部署此类技术的医疗机构具有临床相关性的性能。我们提出了一个以临床医生为基准的评估平台——InterviewPlayground,其核心是一个面向开放式AI访谈的记忆增强型患者模拟器。我们使用InterviewPlayground结合专家撰写的病例片段构建了交互式患者,搭建了用于访谈的模拟问诊平台,并设计了与问诊相关的评估方式。在一项包含6名临床医生、时长25分钟的试点评估中,与基于GPT的大语言模型(LLM)问诊访谈者相比,该LLM恢复了更多嵌入患者病例片段中的临床相关条目(88.0% vs. 38.9%),但做出了更多不基于访谈的临床推断(56.8% vs. 27.8%),且对已识别安全问题的刻画频率更低(33.3% vs. 66.7%),为该任务部署质量保障奠定了基础。
cs.AI / 9 / 2609.21157

Can Agents Design Better Chips with a Higher Level Abstraction?

智能体能否在更高的抽象层次上设计出更好的芯片?
Ding, Zijian, Zou, Yang, Sun, Yizhou, Cong, Jason
Abstract
Large Language Model (LLM) agents are increasingly being explored for chip design, but most existing approaches operate directly at RTL. We ask whether agents can design better chips by leveraging higher-level abstractions. We compare Direct RTL Design, Agent-based HLS Design, Post-Compiler HLS Refinement, and Post-HLS RTL Refinement, and combine Agent-based HLS Design with Post-HLS RTL Refinement as Agent-based HLS with RTL Refinement (AHRR). We use FPGAs as a practical, easy-to-deploy platform for end-to-end evaluation, but note that the design-flow tradeoffs we study are largely independent of the target technology. Across a diverse 11-tasks benchmark suite, AHRR achieves a 2.6$\times$ geometric-mean speedup over Direct RTL Design across our benchmark suite. Case studies show that HLS distills design knowledge into abstractions that agents can leverage, while RTL refinement recovers lower-level optimization opportunities. Together, these results make AHRR a promising workflow for agentic chip design. The code and evaluation artifacts are available at https://github.com/ZijD/AHRR.
Chinese Translation
大语言模型(LLM)智能体正被越来越多地应用于芯片设计,但现有方法大多直接在RTL(寄存器传输级)层面进行。我们探究智能体能否通过利用更高层次的抽象来设计出更好的芯片。我们比较了直接RTL设计、基于智能体的高层综合(HLS)设计、编译后HLS优化以及HLS后RTL优化四种方法,并将基于智能体的HLS设计与HLS后RTL优化相结合,提出了基于智能体的HLS与RTL优化方法(Agent-based HLS with RTL Refinement, AHRR)。我们采用FPGA作为实用且易于部署的平台进行端到端评估,但需要指出的是,我们所研究的设计流程权衡在很大程度上与目标工艺技术无关。在一个包含11个任务的多样化基准测试套件上,AHRR相比直接RTL设计实现了2.6倍的几何平均加速。案例研究表明,HLS将设计知识提炼为智能体可利用的抽象,而RTL优化则能恢复更低层次的优化机会。综合来看,这些结果使AHRR成为智能体芯片设计的一种有前景的工作流程。代码和评估工件可在 https://github.com/ZijD/AHRR 获取。
cs.AI / 10 / 2609.21165

SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity

SpecOpt:面向结合特异性的智能体分子优化接触差异推理方法
Nguyen, Thao, Ji, Heng
Abstract
Off-target protein binding is a major source of adverse effects for small-molecule drugs, yet most structure-based molecular design methods focus on generating selective compounds de novo rather than improving the selectivity of existing, well- characterized drugs. We introduce specificity optimization (SpecOpt), a molecular design task that seeks constrained structural modifications to an existing compound that increase its binding preference for an intended target over known off-targets while preserving its structural identity and drug-like properties. To enable systematic evaluation, we construct a ChEMBL-derived benchmark from compound-target interaction data, identifying intended targets through curated drug-mechanism annotations and off- targets through measured activities. We then develop an agentic framework that docks each compound against its intended target and off-targets, compares the resulting poses through residue-aware atom-protein contacts, and provides these differential interactions to a large language model to propose targeted structural modifications. Candidates are retained only if they satisfy molecular similarity, ADMET, and target-off-target docking selectivity criteria. On 915 compounds, the agent improves the target- off-target binding gap for 84.8% of compounds, shifting the mean gap from -0.72 to +0.47 kcal/mol while maintaining a mean Tanimoto similarity of 0.72 to the starting compounds. Ablation studies identify residue-specific contact information as the critical optimization signal: replacing residue identities with binary contact indicators eliminates improvement on all 29 ablation compounds. These results establish SpecOpt as a distinct molecular design problem and demonstrate residue-aware differential interactions as an effective signal for improving the specificity of existing compounds.
Chinese Translation
脱靶蛋白结合是小分子药物不良反应的主要来源,然而大多数基于结构的分子设计方法侧重于从头生成选择性化合物,而非改进已有的、表征充分的药物的选择性。我们提出特异性优化(specificity optimization, SpecOpt)这一分子设计任务,旨在对现有化合物进行受限的结构修饰,使其在保持结构身份和类药物性质的同时,提高对预期靶点相对于已知脱靶点的结合偏好。为支持系统性评估,我们基于化合物-靶点相互作用数据构建了源自 ChEMBL 的基准数据集,通过经整理的药物机制注释确定预期靶点,并通过实测活性确定脱靶点。随后,我们开发了一个智能体框架,该框架将每个化合物与预期靶点及脱靶点进行分子对接,通过基于残基感知的原子-蛋白接触比较所得构象,并将这些差异相互作用提供给大语言模型以提出针对性的结构修饰建议。仅当候选分子满足分子相似性、ADMET 以及靶点-脱靶点对接选择性标准时才予以保留。在915个化合物上,该智能体使84.8%的化合物的靶点-脱靶点结合差距得到改善,平均差距从 -0.72 提升至 +0.47 kcal/mol,同时与起始化合物保持平均 0.72 的 Tanimoto 相似度。消融实验表明,残基特异性接触信息是关键的优化信号:将残基身份替换为二元接触指示后,所有29个消融化合物上的改进均被消除。这些结果将 SpecOpt 确立为一个独立的分子设计问题,并证明残基感知的差异相互作用是提升现有化合物特异性的有效信号。
cs.AI / 11 / 2609.21181

Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks

在ARC类任务中利用测试时任务嵌入进行隐式规则归纳
Deliège, Adrien, Beger, Claas, Van Droogenbroeck, Marc, Mitchell, Melanie
Abstract
The Abstraction and Reasoning Corpus and related benchmarks evaluate whether AI models can solve novel reasoning tasks, but often leave unclear whether success reflects inference of the intended underlying rule or reliance on shortcuts. We address this gap by studying test-time task embeddings in Vision ARC (VARC), a model in which a pre-trained backbone is complemented by a trainable embedding representing the transformation rule. In the original VARC, test-time training (TTT) is jointly applied to the backbone and task embedding. Here we introduce a novel two-step TTT protocol: first finetune only the task embedding (Embed-TTT), then freeze it and finetune the backbone. Across ARC-AGI-1, ConceptARC, and two controlled datasets with known rules, Embed-TTT consistently yields improved task embeddings, ones that align better with underlying task rules, improve embedding-based retrieval, and enable accurate linear probing of known rules. Qualitatively, Embed-TTT identifies more semantically meaningful relations between test and train tasks on ARC-AGI-1. We also show that optimizing only task embeddings (less than 0.01% of model parameters) already solves a non-trivial fraction of ARC-AGI-1, ConceptARC, and Mini-ARC tasks, while the full two-step pipeline improves final performance. Finally, we show that Embed-TTT recovers the underlying geometric structure of parametric rules and learns compositional capabilities that enable rule-wise interpolation, but not extrapolation. These findings support a clearer separation between rule induction and rule execution in ARC-like evaluations, motivating benchmarks that better distinguish in-distribution from out-of-distribution rules.
Chinese Translation
抽象与推理语料库(Abstraction and Reasoning Corpus,ARC)及相关基准用于评估AI模型能否解决新颖的推理任务,但往往无法明确其成功究竟是反映了模型对预期底层规则的推断,还是依赖某些捷径。为弥补这一空白,我们研究了视觉版ARC(Vision ARC,VARC)中的测试时任务嵌入,该模型在预训练骨干网络的基础上增加了一个可训练的嵌入,用于表示变换规则。在原始VARC中,测试时训练(Test-Time Training,TTT)被同时应用于骨干网络和任务嵌入。本文我们提出一种新颖的两步TTT协议:首先仅微调任务嵌入(Embed-TTT),然后将其冻结并微调骨干网络。在ARC-AGI-1、ConceptARC以及两个规则已知的受控数据集上,Embed-TTT均能持续产生更好的任务嵌入:这些嵌入与底层任务规则更加吻合,提升了基于嵌入的检索效果,并支持对已知规则的准确线性探查。从定性结果看,Embed-TTT在ARC-AGI-1上能够识别出测试任务与训练任务之间更多具有语义意义的关系。我们还发现,仅优化任务嵌入(不到模型参数的0.01%)即可解决ARC-AGI-1、ConceptARC和Mini-ARC中相当一部分非平凡任务,而完整的两步流程可进一步提升最终性能。最后,我们证明Embed-TTT能够恢复参数化规则的底层几何结构,并学习到支持规则层面插值(但无法外推)的组合能力。这些发现支持在ARC类评估中对规则归纳与规则执行进行更清晰的区分,并启发构建能够更好区分分布内与分布外规则的基准。
cs.AI / 12 / 2609.21192

AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture

AI-GRACE:面向智能体人工智能(Agentic AI)的用例操作化框架:从组织目标与义务到部署能力与架构
Cuneo, John, Chun, David, Khanna, Gaurav
Abstract
Organizations deploying agentic artificial intelligence must determine more than whether a model is trustworthy; they must establish what to validate, control, and observe for a use case to deliver its intended outcome while meeting applicable obligations. This paper proposes AI-GRACE (Agentic Intelligence-Governance, Risk, Assurance, Controls, and Evidence) as a use-case operationalization framework connecting organizational governance with technical implementation. The proposal draws on professional observations and a purposive synthesis of standards and literature, using design science to frame the method contribution and situational method engineering to guide contextual tailoring and reuse. The framework establishes objectives and obligations and then assesses risks in seven proposed domains, including mission and value realization. It derives requirements for assurance before deployment, runtime controls, and evidence, which guide capability qualification, gap assessment, and a logical architecture. An Agent Operating Envelope specifies permitted actions and escalation conditions, while Risk-Aligned Independence Levels (RAIL) summarize the authorized independence. A fictional retail banking application illustrates the method. The contribution is a traceable basis for deciding what an organization must implement, what it already supports, and what remains unresolved. Empirical evaluation must establish whether it improves deployment decisions, efficiency, and reuse.
Chinese Translation
部署智能体人工智能(agentic AI)的组织所需要确定的,不仅是某个模型是否可信;它们还必须明确针对某一用例应验证什么、控制什么以及观察什么,以便在满足相关义务的同时实现预期成果。本文提出 AI-GRACE(智能体智能—治理、风险、保证、控制与证据,Agentic Intelligence-Governance, Risk, Assurance, Controls, and Evidence)作为一个用例操作化框架,将组织治理与技术实现相连接。该提案基于专业观察以及对标准和文献的目的性综合,并运用设计科学(design science)来构建方法贡献,同时借助情境方法工程(situational method engineering)指导情境化的裁剪与复用。该框架首先确立目标与义务,随后在七个提议的领域(包括使命与价值实现)中评估风险。它进一步推导出部署前的保证要求、运行时控制要求以及证据要求,用以指导能力鉴定、差距评估和逻辑架构设计。智能体运行包络(Agent Operating Envelope)规定了允许执行的操作和升级条件,而风险对齐独立级别(Risk-Aligned Independence Levels, RAIL)则概括了被授权的自主程度。文中以一个虚构的零售银行应用对该方法进行说明。本文的贡献在于为组织决定必须实现什么、已支持什么以及尚待解决什么提供了可追溯的依据。其实证评估仍有待开展,以确定该框架是否能改善部署决策、效率与复用。
cs.AI / 13 / 2609.21208

Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation

基于多样性剪枝测试的信息增益奖励:面向可靠代码生成的真值锚定验证器协同训练
Nunez, Ana, Najafirad, Peyman
Abstract
Self-play methods that co-train a single language model as both coder and test author promise to move code-generation RL beyond fixed test suites, but they suffer from two coupled pathologies: permissiveness collapse, where pass-rate rewards are maximised by trivial, non-discriminative tests, and concentration bias, where i.i.d. sampled tests cluster on modal inputs and inflate estimator variance. We introduce CoVer (Co-trained Coder and Verifier), a single-policy GRPO framework that addresses both failure modes. First, an information-gain (IG) reward scores each self-generated test by the mutual information between its pass/fail vector and a graded, ground-truth-anchored correctness signal y [0, 1] m, gated by the sign of their covariance so that only positively discriminative tests receive reward. Second, a three-stage diversity-aware selection step prunes a candidate pool to a behaviourally non-redundant suite (invalidity, input-string, execution-profile filtering), raising the effective sample size of the IG estimator at fixed execution budget. On five benchmarks (LiveBench, MBPP, LiveCodeBench, CodeContests, Code-Forces), CoVer raises one-shot pass@1 by +5.8 points at 7B and +7.1 points at 14B over the Qwen2.5-Instruct backbone, and achieves the highest macro-average among all compared methods at both scales. As a drop-in backbone inside the CodeT ranking pipeline, CoVer-7B adds +3.5 points, demonstrating the dual benefit of co-training for both generation and selection.
Chinese Translation
将单个语言模型同时作为代码生成器和测试编写者进行协同训练的自博弈方法,有望使代码生成强化学习(RL)超越固定测试集的限制,但这类方法存在两种相互耦合的病态问题:一是宽容性坍缩,即通过琐碎、无区分性的测试最大化通过率奖励;二是集中性偏差,即独立同分布(i.i.d.)采样的测试聚集于模态输入,从而放大估计器方差。我们提出 CoVer(协同训练的代码生成器与验证器,Co-trained Coder and Verifier),这是一个同时应对上述两种失效模式的单策略 GRPO 框架。首先,信息增益(Information-Gain, IG)奖励通过每个自生成测试的通过/失败向量与分级真值锚定正确性信号 y ∈ [0,1]^m 之间的互信息对其进行评分,并以两者协方差的符号作为门控条件,使得只有具有正向区分性的测试才能获得奖励。其次,一个三阶段的多样性感知选择步骤将候选测试池剪枝为行为上非冗余的测试套件(包括无效性过滤、输入字符串过滤和执行特征过滤),在固定执行预算下提升 IG 估计器的有效样本量。在五个基准(LiveBench、MBPP、LiveCodeBench、CodeContests、Codeforces)上,CoVer 相对于 Qwen2.5-Instruct 基座模型,在 7B 规模上将单次生成的 pass@1 提升 5.8 个百分点,在 14B 规模上提升 7.1 个百分点,并在两种规模上均取得所有对比方法中最高的宏平均值。作为 CodeT 排序流程中的即插即用基座模型,CoVer-7B 额外带来 3.5 个百分点的提升,展示了协同训练在代码生成与代码选择两方面的双重收益。
cs.AI / 14 / 2609.21214

Ability-Residual Decoupled Modeling for Affective Cognitive Diagnosis

面向情感认知诊断的能力-残差解耦建模
Zhao, Boyuan, Ye, Meng
Abstract
Cognitive diagnosis infers students' concept mastery from response logs. However, students' responses are not determined by mastery alone: non-cognitive factors such as emotion, engagement, and fatigue can also affect performance. Affective cognitive diagnosis therefore extends conventional cognitive diagnosis by incorporating affective states. Existing methods often assume that the cognitive diagnosis backbone has already explained ability, item, and concept effects, so the remaining errors can be attributed mainly to affect. We argue that this assumption can be insufficient in real educational data: item calibration bias, systematic concept bias, personalized student-concept deviations, and latent student-item matching can form stable cognitive residuals. Without an explicit modeling pathway, these residuals may leak into affective representations, producing affect contamination. To address this problem, we propose an ability-residual decoupled framework for affective cognitive diagnosis. The model first captures unmodeled cognitive residuals through student, item, concept, student-concept, and low-rank student-item components, and then uses an affective module to modulate guess/slip effects. A Q-matrix-constrained concept residual attention mechanism adaptively aggregates only item-relevant concept residuals. Experiments on ASSIST2017, ASSIST2012, ASSIST2009, and Junyi with six cognitive diagnosis backbones show response-prediction gains across the reported comparisons and generally improved affect alignment when affect labels are available. Ablation studies, leakage probes, principal component analysis visualization, long-tail analysis, and case studies further indicate that ability residuals absorb stable cognitive bias, reduce cognitive contamination in the affective branch, and enhance the robustness and predictive accuracy of cognitive diagnosis models.
Chinese Translation
认知诊断旨在从作答记录中推断学生对知识概念的掌握情况。然而,学生的作答并非仅由掌握程度决定:情绪、投入度和疲劳等非认知因素也会影响其表现。因此,情感认知诊断通过引入情感状态对传统认知诊断进行了扩展。现有方法通常假设认知诊断主干模型已经充分解释了能力、试题和概念等因素的影响,从而可将剩余误差主要归因于情感因素。我们认为,这一假设在真实教育数据中可能并不成立:试题标定偏差、系统性概念偏差、个性化的学生-概念偏离以及潜在的学生-试题匹配关系都可能形成稳定的认知残差。若缺乏显式的建模路径,这些残差可能会泄漏到情感表示中,导致情感污染。为解决这一问题,我们提出了一种面向情感认知诊断的能力-残差解耦框架。该模型首先通过学生、试题、概念、学生-概念以及低秩学生-试题等分量来捕获未被建模的认知残差,然后利用情感模块调节猜测与失误效应。此外,一种由Q矩阵约束的概念残差注意力机制可自适应地仅聚合与试题相关的概念残差。在ASSIST2017、ASSIST2012、ASSIST2009和Junyi数据集上结合六种认知诊断主干模型的实验表明,所提方法在所报告的对比中均提升了作答预测性能,并且在有情感标签的情况下情感对齐度普遍得到改善。消融实验、泄漏探测、主成分分析可视化、长尾分析以及案例研究进一步表明,能力残差能够吸收稳定的认知偏差,减少情感分支中的认知污染,并提升认知诊断模型的鲁棒性与预测准确率。
cs.AI / 15 / 2609.21221

A Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning

一种用于感知任务规划的全可微神经-软-符号框架
Wei, Hongyan, AbdAlmageed, Wael
Abstract
Perceptual planning tasks require two key capabilities: accurately perceiving uncertain scenes and planning valid action sequences following logical rules. Conventional methods convert perception into discrete symbolic facts and then plan, discarding perceptual uncertainty and severing task-level feedback to perception. We introduce a generic, fully differentiable neuro-soft-symbolic framework that connects visual perception and task planning within a single computational graph. The framework maintains a continuous soft symbolic state, lifts domain rules into a differentiable soft-$T_P$ transition operator, and optimizes action logits over a short planning horizon. Gradients from the planning objective can also update the perception parameters, allowing task-relevant perceptual representations to be refined during planning. On Blocksworld, our method solves 40/40 LatPlan-40 tasks and 596/600 PlanBench-600 tasks, compared with 33/40 for LatPlan and 587/600 for the reasoning-model baseline, while requiring substantially less computation and time. In the perceptual-uncertainty ablation, our method improves the success rate from 59\% with frozen perception to 83\%. We further conduct task-and-motion simulations on Blocksworld scenes, providing an execution-level validation of the compatibility between decoded task plans and downstream robotic motion execution.
Chinese Translation
感知规划任务需要两项关键能力:准确感知不确定的场景,以及遵循逻辑规则规划有效的动作序列。传统方法先将感知结果转换为离散的符号事实,再进行规划,这不仅丢弃了感知的不确定性,也切断了任务层面对感知的反馈。我们提出了一种通用的、完全可微的神经-软-符号框架,将视觉感知与任务规划连接在单一计算图中。该框架维护一个连续的软符号状态,将领域规则提升为可微的软-$T_P$ 转移算子,并在短规划时域内优化动作逻辑值(logits)。来自规划目标的梯度还可以更新感知参数,使得与任务相关的感知表征能够在规划过程中得到精炼。在 Blocksworld 上,我们的方法解决了 40/40 个 LatPlan-40 任务和 596/600 个 PlanBench-600 任务,相比之下 LatPlan 为 33/40,推理模型基线为 587/600,同时所需计算量与时间显著更少。在感知不确定性消融实验中,我们的方法将成功率从感知冻结时的 59% 提升至 83%。我们还在 Blocksworld 场景上进一步开展了任务与运动仿真,为解码出的任务规划与下游机器人运动执行之间的兼容性提供了执行层面的验证。
cs.AI / 16 / 2609.21259

CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition

CogGym:迈向人类认知与机器认知的大规模比较评估
Ying, Lance, Wu, Jinzhou, Wang, Yingshan Susan, Aarya, Shivam, Buschoff, Luca M. Schulze, Chen, Harry, Collins, Katherine M., de Varda, Andrea, Fu, Shuhao, Houlihan, Sean Dae, Jagadish, Akshay K., Jiang, Guangyuan, Kiegeland, Samuel, Kurumisawa, Tetsu, Liu, Rongzhi, Liu, Ryan, Ma, Ningshan, McGregor, Kathryn, Strittmatter, Younes, Tsvilodub, Polina, Vigly, Jacob Hoover, Wu, Sarah, Xu, Enjie, Yun, Yiling, Allen, Kelsey, Brooke-Wilson, Tyler, Christian, Brian, Fedorenko, Evelina, Frank, Michael C., Franke, Michael, Gao, Tao, Gershman, Samuel J., Hawkins, Robert D., Hu, Jennifer, Jara-Ettinger, Julian, Kleiman-Weiner, Max, Levine, Sydney, Linzen, Tal, Lu, Hongjing, O'Donnell, Timothy, Ong, Desmond C., Piantadosi, Steven T., Saxe, Rebecca, Schulz, Eric, Shu, Tianmin, Sosa, Felix A., Sucholutsky, Ilia, Zhi-Xuan, Tan, Ullman, Tomer, Xu, Fei, Yildirim, Ilker, Zhu, Jian-Qiao, Griffiths, Thomas L., Gerstenberg, Tobias, Smith, Kevin, Tenenbaum, Joshua B.
Abstract
Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models' improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model--human fit remains well below human splithalf reliability ($R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video) with the best models achieving $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.
Chinese Translation
理解与建模人类智能是人工智能(AI)与认知科学共同追求的两个并行目标。随着AI系统能力的不断提升,模型响应在哪些方面与人类响应相似,又在哪些方面存在系统性差异?人类能够执行和思考的任务范围之广、种类之多,给人类与模型之间可扩展且严谨的比较带来了挑战。我们提出CogGym,这是一个以认知科学为基础的、可扩展的统一框架,用于在匹配的实验试次上系统性地比较模型与人类行为。CogGym采用半自动化的人机协同(human-in-the-loop)流程,将多样化的实验范式标准化为一种与具体任务无关的实验标记语言(Experiment Markup Language, EML),从而实现可复现且忠实的大规模比较。在初始发布中,我们整理并标准化了来自100篇论文的258个聚焦于人类常识推理的认知实验,并将50个大语言模型与人类响应进行对比评估。我们发现了一个清晰的规模扩展趋势:规模更大、更新的AI模型能更好地复现人类判断。然而,AI模型在这类常识推理任务上的改进速度明显慢于在数学和编程等形式化推理基准上的提升,且模型与人类的拟合度仍远低于人类的分半信度(文本为$R^2 = 0.93$,图像为$0.95$,视频为$0.92$),最佳模型在文本实验上仅达到$R^2 = 0.59$,图像上为$0.58$,视频上为$0.43$。我们希望CogGym能够作为一个持续演进的评估框架,不断纳入新的认知科学实验,以刻画模型行为在哪些方面与人类行为相似、在哪些方面存在系统性差异,以及随着模型与实验的发展这些模式如何变化。
cs.AI / 17 / 2609.21263

PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking

PlaceReasoner-Beta:推理驱动的宏单元布局与基准测试
Li, Qiufeng, Wang, Chengxuan, Chen, Rongqian, Cheng, Quan, Ren, Yihui, Ho, Chia-Tung, Pan, David Z., Lan, Tian, Cao, Weidong
Abstract
Automated macro placement remains a fundamental challenge in VLSI physical design. Despite decades of research, existing approaches predominantly optimize hand-crafted proxy objectives, such as estimated wirelength, and typically produce placements through one-shot numerical optimization, limiting their ability to incorporate visual layout context, codified design expertise, and downstream physical-design feedback in a unified loop. We present PlaceReasoner-Beta, a verifier-guided multi-agent framework that reformulates macro placement as a closed-loop reasoning problem rather than black-box optimization. A vision-language model (VLM) planner generates candidate placements from the floorplan image, macro specifications, and connectivity structure; a geometric verifier enforces physical legality and expert placement principles; a physical verifier refines candidates using early implementation feedback; and a post-route optimizer further improves promising layouts using final PPA. To enable reproducible evaluation, we introduce PlaceReasoner-Bench, a fully open end-to-end benchmark built from open RTL designs, EDA tools, and technology libraries. It comprises 8 designs at two aspect ratios, yielding 16 tasks with fixed floorplans and I/O assignments, so methods differ only in macro positions and orientations and are evaluated using routed PPA and DRC rather than pre-route proxies. Across the benchmark, PlaceReasoner-Beta achieves the best timing among DRC-clean methods on all square tasks, reducing post-route TNS by 61.2% at 1:1 and 53.0% at 2:1 relative to the classical baseline field. It also shortens routed wirelength on most designs despite never explicitly optimizing it, demonstrating that reasoning over spatial structure under physical-design feedback can improve end-to-end layout quality beyond proxy-objective optimization.
Chinese Translation
自动化宏单元布局(macro placement)始终是VLSI物理设计中的一项根本性挑战。尽管经过数十年的研究,现有方法主要优化人工设计的代理目标(如估算的线长),并且通常通过一次性数值优化生成布局,这限制了它们在统一闭环中融合视觉版图上下文、成体系的设计专业知识以及下游物理设计反馈的能力。我们提出PlaceReasoner-Beta,一个由验证器引导的多智能体框架,它将宏单元布局重新表述为一个闭环推理问题,而非黑盒优化。视觉语言模型(VLM)规划器根据布局规划图(floorplan)图像、宏单元规格和连接结构生成候选布局;几何验证器强制执行物理合法性检查和专家布局原则;物理验证器利用早期实现反馈对候选方案进行精炼;布线后优化器则基于最终PPA进一步改进有前景的版图。为实现可复现的评估,我们引入PlaceReasoner-Bench,一个基于开源RTL设计、EDA工具和工艺库构建的完全开放的端到端基准测试。该基准包含8个设计在两种长宽比下的配置,共形成16个具有固定布局规划图和I/O分配的任务,从而使各方法仅在宏单元位置和朝向上存在差异,并以布线后的PPA和DRC(而非布线前代理指标)进行评估。在该基准上,PlaceReasoner-Beta在所有正方形任务上均取得了DRC清洁方法中的最优时序表现,相对于经典基线方法族,在1:1和2:1长宽比下分别将布线后TNS降低了61.2%和53.0%。此外,尽管从未显式优化线长,它仍在大多数设计上缩短了布线后线长,这表明在物理设计反馈下对空间结构进行推理,能够带来超越代理目标优化的端到端版图质量提升。
cs.AI / 18 / 2609.21267

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

生产环境中的高效基准测试:一项针对不断演化的LLM智能体的研究
She, Yining, Lin, Lei
Abstract
Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to rerun. We study efficient recurring evaluation for a production analytics agent serving tens of thousands of monthly active users and report first-hand deployment experience. Using 574 historical runs of the production benchmark, split chronologically into calibration and held-out periods, we compare random sampling, historical caching, fixed representative subsets, and IRT-based adaptive testing. The results show that multidimensional 2PL adaptive testing achieves the best overall score fidelity: executing 200 questions, 38.5% of a full run, yields 1.03 pp of MAE. We nevertheless deployed difficulty-stratified fixed subsets because of their operational simplicity, and show they transfer without recalibration to five other agent families and remain stable across calibration windows as short as one day. Drawing on this deployment experience, we report practical recommendations for recurring production-agent evaluation.
Chinese Translation
生产环境中的大语言模型(LLM)智能体在持续演化过程中需要被反复评估,但完整的智能体基准测试重新运行的成本很高。我们针对一个服务于数万名月活跃用户的生产级数据分析智能体,研究了高效的周期性评估方法,并报告了一手的部署经验。利用574次生产基准测试的历史运行记录,按时间顺序划分为校准期和留出期,我们比较了随机抽样、历史缓存、固定的代表性子集以及基于IRT(项目反应理论)的自适应测试。结果表明,多维2PL自适应测试实现了最佳的整体分数保真度:仅需执行200道题目(即完整运行的38.5%),即可将MAE(平均绝对误差)控制在1.03个百分点。尽管如此,出于操作简便性的考虑,我们最终部署了按难度分层的固定子集,并证明其无需重新校准即可迁移到另外五个智能体家族,且在短至一天的校准窗口内保持稳定。基于这一部署经验,我们针对生产智能体的周期性评估提出了实用的建议。
cs.AI / 19 / 2609.21293

GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development

GameASG-Bench:面向游戏开发的自主软件生成基准测试
Zhang, Xiuhui, Chen, Yi, Xu, Shusheng, Li, Fan, Wang, Huan, Yang, Tongkai, Yuan, Binhang
Abstract
Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessarily establish that their interacting components satisfy the specified behavioral requirements. We introduce GameASG-Bench, a benchmark that makes behavioral testability part of the generation task for game development. Our design declares an evaluation interface specification before generation, fixing legal starting scenarios, player-level actions, stable snapshots, rejection behavior, and invariants while leaving private implementations open. Concretely, we include: (i) static L1 checks that assess source-level compliance; and (ii) browser-executed L2 checks that combine semantic observations with real input and runtime evidence. We implement this protocol as 47 browser-native game-generation tasks spanning 12 primary genres and both 2D and 3D interaction, each with executable checks and an independently verified reference implementation. Our experiments answer four key questions about end-to-end agent performance, tool access and nominal turn budget, reasoning effort, and harness choice. Across nine agent stacks, the highest observed mean L2 check pass rate is 93.2%, yet the highest observed strict task success rate, requiring all L1 and applicable L2 prerequisite and core requirement checks, is only 55.3% (26/47 tasks). For DeepSeek-V4-Flash, full tool access and larger nominal turn budgets yield more strict task successes, while the strict task success rate is not monotonic in reasoning effort. Both tested harnesses achieve 18 strict task successes, but only ten tasks succeed under both. These results expose task-level compliance gaps that high average check pass rates actually obscure.
Chinese Translation
自主软件生成(Autonomous Software Generation, ASG)旨在将人类需求转化为可执行的应用程序,但交付这些应用并不一定意味着其交互组件满足指定的行为需求。我们提出了 GameASG-Bench,这是一个将行为可测试性纳入游戏开发生成任务的基准。我们的设计在生成之前先声明评估接口规范,固定合法的起始场景、玩家级操作、稳定快照、拒绝行为以及不变式,同时保持私有实现的开放性。具体而言,我们包括:(i) 评估源代码级合规性的静态 L1 检查;以及 (ii) 将语义观察与真实输入和运行时证据相结合的浏览器执行 L2 检查。我们将该协议实现为 47 个浏览器原生游戏生成任务,涵盖 12 种主要游戏类型以及 2D 和 3D 交互,每个任务都配有可执行检查和经过独立验证的参考实现。我们的实验回答了关于端到端智能体性能、工具访问与名义轮次预算、推理投入程度以及执行框架(harness)选择四个关键问题。在九个智能体技术栈中,观察到的最高平均 L2 检查通过率为 93.2%,然而观察到的最高严格任务成功率(要求所有 L1 检查及适用的 L2 前提条件和核心需求检查全部通过)仅为 55.3%(26/47 个任务)。对于 DeepSeek-V4-Flash,完整的工具访问权限和更大的名义轮次预算带来更多严格任务成功,而严格任务成功率与推理投入程度之间并非单调关系。两个受测执行框架均实现了 18 个严格任务成功,但仅 ten 个任务在两者下均成功。这些结果揭示了高平均检查通过率实际上所掩盖的任务级合规性差距。
cs.AI / 20 / 2609.21325

LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces

LEGIT:面向可信AI智能体市场的认证协议
Drew, Steve, Zhou, Jiayu
Abstract
Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete specialized tasks for buyers. A major challenge of such marketplaces is that buyers cannot easily determine which agent will perform best on their tasks. Reported benchmark scores may be difficult to verify or compare across tasks, software, and budgets. We introduce LEGIT, a credentialing protocol connecting certification, reputation, and proposed marketplace allocation. Certification binds measured quality and cost per solved task to an agent configuration, task domain, evaluation budget, and evidence through a signed record. Reputation links records of past task outcomes to the same identity, subject to the reliability of the reported feedback. Buyers and agents can verify credential records and inspect optional visual profiles. Evaluations reveal cost differences between agent configurations with similar observed task success, and show that comparisons depend on the evaluation budget. These results support binding performance measurements to the tested configuration and resource limits. A complementary analysis quantifies the deposits and fees required for reputation manipulation under a stated Sybil attack model.
Chinese Translation
智能体市场(Agentic marketplaces)正在兴起,其中具有不同能力的AI智能体(AI agents)自主地为买方完成专业化任务。此类市场面临的一个主要挑战是,买方难以确定哪个智能体在其任务上表现最佳。已公布的基准测试分数可能难以验证,或难以在不同任务、软件和预算之间进行比较。我们提出了LEGIT,一种将认证、声誉与建议的市场分配机制相连接的认证协议。认证通过一份签名记录,将测得的质量和每个已解决任务的成本绑定到特定的智能体配置、任务领域、评估预算和证据上。声誉机制在所报告反馈可靠性的前提下,将过往任务结果记录与同一身份相关联。买方和智能体可以验证凭证记录,并查看可选的可视化档案。评估结果显示,在观察到的任务成功率相近的智能体配置之间,成本存在显著差异,且比较结果依赖于评估预算。这些结果支持将性能测量绑定到所测试的配置和资源限制上。一项补充分析在给定的女巫攻击(Sybil attack)模型下,量化了声誉操纵所需的保证金和费用。
cs.AI / 21 / 2609.21390

Offline Multimodal Large Language Models for Decision Support in Air Operations

用于空战行动决策支持的离线多模态大语言模型
Dantas, Joao P. A., Cunha, Jelton A., Dietzsch, Gabriel
Abstract
Air operations rely on complex rules, established procedures, and time-critical analysis under limited connectivity and strict security constraints. In such environments, analysts must combine written doctrine with images, often without access to external computing resources. This paper studies offline large language models as decision support tools, deployed in isolated and restricted environments to give analysts access to doctrinal knowledge that remains traceable to its original sources through natural language interaction. We describe a modular retrieval-augmented architecture suitable for operation without Internet connectivity, supporting both text and image input from technical manuals. As a first step toward evaluating this architecture, we report a pilot study with four image analysts of the Brazilian Air Force, combining (i) a doctrinal knowledge assessment based on their electronic-target identification doctrine, comparing human and proposed system performance on the same test, and (ii) a measurement of the cognitive workload involved in manually producing a reconnaissance target report (Relat\'orio de Miss\~ao de Reconhecimento - REMIR) without AI assistance. The results show a demanding manual task, especially in terms of mental demand (6.0/7) and effort (5.0/7), while the proposed system matches the human score (8/10) and completes the assessment in 7.1 minutes (compared to a human average of 26.5 minutes), establishing a baseline for future AI-assisted evaluation. Finally, we describe a future evaluation protocol to systematically compare manual and AI-assisted workflows.
Chinese Translation
空战行动依赖于复杂的规则、既定的程序以及在有限联网条件和严格安全约束下进行的时效性分析。在此类环境中,分析人员必须将书面条令与图像相结合,且往往无法访问外部计算资源。本文研究了离线大语言模型作为决策支持工具的应用,将其部署于隔离和受限的环境中,使分析人员能够通过自然语言交互访问条令知识,并保持知识与其原始来源的可追溯性。我们描述了一种适用于无互联网连接环境的模块化检索增强架构,支持来自技术手册的文本和图像输入。作为评估该架构的第一步,我们报告了一项与四名巴西空军图像分析人员合作开展的试点研究,包括:(i)基于其电子目标识别条令的条令知识评估,在相同测试中比较人类与所提系统的表现;(ii)测量在无AI辅助情况下手动编制侦察目标报告(Relatório de Missão de Reconhecimento - REMIR)所涉及的认知负荷。结果显示该手动任务要求很高,尤其在心理需求(6.0/7)和努力程度(5.0/7)方面;而所提系统达到了与人类相同的得分(8/10),并在7.1分钟内完成评估(相比之下人类平均用时26.5分钟),为未来的AI辅助评估建立了基线。最后,我们描述了一个用于系统化比较手动流程与AI辅助流程的未来评估方案。
cs.AI / 22 / 2609.21423

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

DENSE:将智能体轨迹蒸馏为证据支撑的快捷树以实现自我精炼
Liu, Siyuan, Yu, Fan, Ru, Dongyu, Liu, Yizhu, Yang, Yifan, Cao, Xuezhi, Cai, Xunliang, Cao, Yixin
Abstract
Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to distill these traces into reusable feedback without post-hoc outcome labels, drawing on their evidence of local progress, recovery, and unfinished requirements. We introduce DENSE (Distilling Evidence from Nested Subtask Executions), which organizes this evidence into evidence-grounded nested shortcut trees. DENSE compresses redundant attempts, reconciles issues across levels using recovery evidence, and summarizes completed branches while expanding unresolved ones, linking reusable progress to remaining obligations. We introduce REFIT, a source-paired protocol comparing feedback from shared initial trajectories under post-hoc outcome blindness, with environments and model contexts reset for fresh attempts at the same tasks. On Terminal-Bench 2.1, DENSE achieves the highest strict pass rate among tested non-privileged feedback methods across four recipient models. Relative to initial executions, strict pass rate improves by 7.12-15.64 pp, with 19.0-43.6% fewer observed recipient tokens in reruns. GPT-5.5 ablations support combining nested subtask analysis with shortcut construction and issue reconciliation. These findings point toward agent self-refinement through evidence-grounded trajectory reuse with less reliance on external supervision.
Chinese Translation
在线智能体部署会产生大量的执行轨迹,而面向特定任务的验证和专家标注难以规模化。我们研究如何在不依赖事后结果标签的情况下,将这些轨迹蒸馏为可复用的反馈,利用其关于局部进展、问题恢复和未完成需求的证据。我们提出 DENSE(Distilling Evidence from Nested Subtask Executions,从嵌套子任务执行中蒸馏证据),该方法将这些证据组织为证据支撑的嵌套快捷树。DENSE 压缩冗余尝试,利用恢复证据在不同层级之间协调问题,总结已完成的分支并展开未解决的分支,从而将可复用的进展与剩余的任务义务关联起来。我们提出 REFIT,一种源轨迹配对的评测协议,在事后结果盲测条件下比较基于相同初始轨迹生成的反馈,并重置环境与模型上下文以便对相同任务进行全新尝试。在 Terminal-Bench 2.1 上,DENSE 在四个接收模型中取得了所测试的非特权反馈方法中最高的严格通过率。相对于初始执行,严格通过率提升 7.12–15.64 个百分点,重跑中观测到的接收模型 token 数量减少 19.0–43.6%。GPT-5.5 消融实验支持将嵌套子任务分析与快捷树构建及问题协调相结合。这些发现表明,通过基于证据的轨迹复用,智能体自我精炼可以更少依赖外部监督。
cs.AI / 23 / 2609.21432

GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation

GVPO++:用于大语言模型后训练与在线策略蒸馏的组方差策略优化
Zhang, Kaichen, Hong, Yuzhong, Bao, Junwei, Jiang, Hongfei, Song, Yang, Hong, Dingqian, Xiong, Hui
Abstract
Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Despite recent advances in post-training methods, such as Group Relative Policy Optimization (GRPO), their practical deployment remains impeded by training instability arising from the reliance on importance sampling. We introduce Group Variance Policy Optimization (GVPO), a novel post-training method that integrates the analytical solution of KL-constrained reward maximization into its gradient weighting scheme. This formulation provides an intuitive interpretation: GVPO's gradient corresponds to the mean squared error between the central distance of implicit rewards and that of actual rewards. GVPO offers two key advantages: (1) it guarantees a unique optimal solution, exactly to the KL-constrained reward maximization objective, and (2) it enables flexible sampling distributions without requiring importance sampling. Beyond general post-training, we show that GVPO naturally extends to on-policy distillation (OPD). Furthermore, GVPO enables the optimization of a broad family of extended OPD objectives, providing a principled foundation for diverse objective design. By unifying theoretical guarantees with practical adaptability, GVPO establishes a new paradigm for reliable and versatile LLM post-training and on-policy distillation.
Chinese Translation
后训练在提升大语言模型(LLM)的推理能力和特定任务专业性方面发挥着关键作用。尽管近年来出现了一些后训练方法的进展,例如组相对策略优化(Group Relative Policy Optimization, GRPO),但由于依赖重要性采样而导致的训练不稳定性,其实际部署仍然受到阻碍。我们提出了组方差策略优化(Group Variance Policy Optimization, GVPO),这是一种新颖的后训练方法,它将带KL约束的奖励最大化的解析解融入其梯度加权方案中。该公式提供了一种直观的解释:GVPO的梯度对应于隐式奖励的中心距离与实际奖励的中心距离之间的均方误差。GVPO具有两个关键优势:(1)它保证了解的唯一最优解,可精确求解带KL约束的奖励最大化目标;(2)它支持灵活的采样分布,而无需重要性采样。除了一般的后训练之外,我们还证明GVPO可以自然地扩展到在线策略蒸馏(on-policy distillation, OPD)。此外,GVPO能够优化一大类扩展的OPD目标,为多样化的目标设计提供了有原则的基础。通过将理论保证与实际适应性相统一,GVPO为可靠且通用的大语言模型后训练与在线策略蒸馏确立了一种新的范式。
cs.AI / 24 / 2609.21470

Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving

面向安全导向端到端自动驾驶的风险感知占据栅格
Chen, Jiaxing, Zou, Hengduo, Zhao, Yiren, Gao, Bolin
Abstract
Sparse representation formulates the environment perception for the end-to-end driving system as a set of discrete elements like objects and lane lines. This formulation meets safety risks in crowded, occluded scenes dealing with unstructured obstacles, uncertain regions, and intricate interactions. In this paper, we propose a dense representation, risk-aware occupancy, to characterize planning-relevant risks in an explicit and uniform manner. It jointly encodes global scene occupancy, map-derived traffic constraints, and future dynamic agent occupancy into a unified BEV map. The unified BEV map captures the risk evidence for trajectory planning in both spatial and temporal dimensions. We design an E2E network, ROIDrive, to realize risk-aware occupancy. It predicts risk-aware occupancy with an independent branch and injects it into planning queries for safety-oriented trajectory generation. In addition, to quantify the safety problem, we introduce RiskOcc4D-nuScenes built upon nuscenes and occ3d-nuscenes. Our risk-aware occupancy yields relative open-loop collision reductions of 52.9% under the UniAD metric and 35.0% under the ST-P3 metric on nuScenes.
Chinese Translation
稀疏表示将端到端驾驶系统的环境感知建模为物体、车道线等离散元素的集合。然而,在处理非结构化障碍物、不确定区域和复杂交互的拥挤、遮挡场景中,这种建模方式面临安全风险。本文提出一种稠密表示——风险感知占据栅格(risk-aware occupancy),以显式且统一的方式刻画与规划相关的风险。该方法将全局场景占据、源自地图的交通约束以及未来动态智能体的占据信息共同编码到统一的鸟瞰图(BEV)地图中,该地图在空间和时间两个维度上捕获用于轨迹规划的风险证据。我们设计了端到端网络 ROIDrive 来实现风险感知占据栅格,该网络通过独立的分支预测风险感知占据,并将其注入规划查询中以实现面向安全的轨迹生成。此外,为量化安全问题,我们基于 nuScenes 和 Occ3D-nuScenes 构建了 RiskOcc4D-nuScenes 数据集。我们的风险感知占据方法在 nuScenes 上将开环碰撞率相对降低了 52.9%(UniAD 指标)和 35.0%(ST-P3 指标)。
cs.AI / 25 / 2609.21486

Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving

基于寄存器行驶,基于风险推理:面向基于寄存器的端到端自动驾驶的风险感知占据栅格
Chen, Jiaxing, Zou, Hengduo, Qin, YuKai, Zhao, Yiren, Yu, Lidong, Gao, Bolin
Abstract
Multimodal trajectory prediction improves behavioral coverage in end-to-end autonomous driving, but existing methods remain limited by sparse scene representations. Incomplete evidence leads to low-quality candidate generation and unreliable ranking among geometrically similar trajectories. On a register-based baseline, bad and poor candidates constitute 19.74% of the candidate set, while the oracle-best candidate ranks only 33.9th on average. We propose RRDrive, which introduces risk-aware occupancy as a dense, temporally aligned, and trajectory-queryable representation. Its global structure guides high-quality multimodal generation, while candidate-conditioned risk queries support fine-grained selection. We further construct RiskOcc4D-NAVSIM with automatic risk annotations. RRDrive achieves a selected-trajectory PDMS of 0.951, representing a 1.5% relative improvement over the baseline (0.937), and improves the average candidate PDMS by 7.7%. In challenging scenes, it improves candidate PDMS by 30.2% and increases the Spearman correlation among good candidates by 0.41, from 0.26 to 0.67. To move beyond this oracle setting, we further develop an external RiskOcc predictor, a perception module that estimates risk-aware occupancy directly from sensor inputs. The competitive performance validates the representation's feasibility.
Chinese Translation
多模态轨迹预测提升了端到端自动驾驶中的行为覆盖能力,但现有方法仍受限于稀疏的场景表示。证据的不完整导致候选轨迹生成质量低下,且在几何上相似的轨迹之间难以进行可靠排序。在基于寄存器的基线方法中,较差和差的候选轨迹占候选集的19.74%,而最优候选轨迹平均仅排在第33.9位。我们提出RRDrive,引入风险感知占据(risk-aware occupancy)作为一种稠密的、时间对齐的、可按轨迹查询的表示。其全局结构引导高质量的多模态生成,而基于候选轨迹条件化的风险查询则支持细粒度的选择。我们进一步构建了带有自动风险标注的RiskOcc4D-NAVSIM数据集。RRDrive实现了所选轨迹PDMS达0.951,相比基线(0.937)取得1.5%的相对提升,并将候选轨迹的平均PDMS提升7.7%。在具有挑战性的场景中,该方法将候选轨迹的PDMS提升30.2%,并使优质候选轨迹之间的Spearman相关性从0.26提升至0.67,提高0.41。为超越这一最优设定,我们进一步开发了外部RiskOcc预测器,一个直接从传感器输入估计风险感知占据的感知模块。具有竞争力的性能验证了该表示的可行性。
cs.AI / 26 / 2609.21492

LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers

LogicTrack:基于形式逻辑求解器的大语言模型推理轨迹审计
Hu, Jingyu, Yang, Shu, Liu, Weiru, Wang, Di
Abstract
Chain-of-Thought (CoT) reasoning has been shown to improve the performance of large language models (LLMs), yet existing optimization methods largely rely on outcome-based feedback, leaving the logical validity of intermediate reasoning steps largely unverified. To address the gap whereby LLMs arrive at correct final answers through logically flawed intermediate reasoning chains, we propose LogicTrack, a neuro-symbolic framework that audits reasoning trajectories by auto-formalizing each reasoning step into symbolic representations and verifying it with automated theorem provers. LogicTrack introduces Solver-Based Backtracking Reward (SBR), a step-wise scoring mechanism that quantifies logical soundness and guides backtracking tree search at inference time. We further extend LogicTrack to construct supervised fine-tuning (SFT) data with backtracking traces from its trajectories, enabling fine-tuned models to internalize step-wise auditing as an intrinsic capability. Extensive experiments across 8 reasoning benchmarks and 7 LLMs demonstrate that LogicTrack effectively improves both the verifiability of reasoning chains and final answer pass rate, thereby enhancing overall CoT quality and trustworthiness in high-stakes domains.
Chinese Translation
思维链(Chain-of-Thought, CoT)推理已被证明能够提升大语言模型(LLMs)的性能,然而现有的优化方法主要依赖基于最终结果的反馈,使得中间推理步骤的逻辑有效性在很大程度上未经验证。针对大语言模型通过存在逻辑缺陷的中间推理链却得出正确最终答案这一空白问题,我们提出了 LogicTrack——一个神经符号化框架,它通过将每个推理步骤自动形式化为符号表示,并利用自动定理证明器进行验证,从而审计推理轨迹。LogicTrack 引入了基于求解器的回溯奖励(Solver-Based Backtracking Reward, SBR),这是一种逐步评分机制,用于量化逻辑合理性,并在推理时引导回溯树搜索。我们进一步扩展了 LogicTrack,利用其轨迹中的回溯踪迹构建监督微调(SFT)数据,使微调后的模型能够将逐步审计内化为一种内在能力。在 8 个推理基准和 7 个大语言模型上的大量实验表明,LogicTrack 有效提升了推理链的可验证性和最终答案的通过率,从而增强了高风险领域中 CoT 的整体质量与可信度。
cs.AI / 27 / 2609.21493

PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design

PolyBridgeBench:面向基于物理的桥梁设计的多模态大语言模型基准测试
Zhao, Zicheng, Chen, Dongyin, Xu, Rui, Xu, Yinghui
Abstract
Multimodal large language models, or MLLMs, perform well at visual understanding and structured generation, yet these capabilities do not establish whether an engineering design will work when executed. Existing benchmarks assess spatial reasoning, structural validity, or physics-grounded construction, but they do not determine whether MLLMs can synthesize complete load-bearing structures and repair them after simulator execution exposes a failure. We introduce PolyBridgeBench, an executable benchmark for multimodal bridge design. A model receives a visual scene and structured engineering constraints and generates a complete node--member--material topology. Deterministic legality checks gate execution in a native dynamic physics simulation. Following an execution failure, the benchmark returns temporal visual evidence from the failed rollout and evaluates repair under a fixed interaction budget. Separate measurements of deterministic validity, dynamic functional success, and post-failure recovery identify the stage at which design fails. Experiments with six representative MLLMs across 189 levels expose a substantial gap between deterministic validity and dynamic success, pronounced sensitivity to material budgets, and limited post-failure recovery under the primary strict-budget setting.
Chinese Translation
多模态大语言模型(MLLMs)在视觉理解与结构化生成方面表现出色,但这些能力并不能保证一个工程设计在执行时切实可行。现有基准评估空间推理、结构有效性或基于物理的建造能力,但并未检验多模态大语言模型能否综合生成完整的承重结构,并在模拟器执行暴露出失效之后对其进行修复。我们提出PolyBridgeBench,一个可执行的多模态桥梁设计基准。模型接收视觉场景和结构化工程约束,并生成完整的节点—杆件—材料拓扑。确定性合法性检查作为门控,控制其在原生动态物理模拟中的执行。当执行失败后,该基准会返回失败过程中的时序视觉证据,并在固定交互预算下评估修复能力。通过对确定性有效性、动态功能成功以及失效后恢复能力的分别测量,可以定位设计失败发生的具体阶段。在189个关卡上对六种代表性多模态大语言模型的实验揭示了确定性有效性与动态成功之间的显著差距、对材料预算的高度敏感性,以及在主要的严格预算设定下有限的失效后恢复能力。
cs.AI / 28 / 2609.21509

The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models

通信瓶颈:语言模型中树结构表达式序列化的往返研究
Suau, Xavier, Morenas, Alex Ferrando de las, Zappella, Luca, Bengio, Samy
Abstract
When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured expressions. A generator converts a procedurally generated arithmetic expression into a word problem, a separate extractor recovers the expression from the word problem alone, and symbolic equivalence provides an exact oracle. Evaluating all pairwise combinations of sixteen models yields a communication matrix whose marginals separate generation quality from extraction quality. Three main findings emerge. First, the channel is lossy and asymmetric: swapping which model generates and which extracts shifts accuracy by up to 60.4 points, and the best pair reaches 92.9% by combining different models on each end rather than the same model on both. Second, at least 73.6% of round-trip failures originate at generation, and difficulty is driven by tree structure (operator count, depth, right-branching) rather than model family. Third, the channel is trainable: ~3600 fine-tuning examples that share the evaluation's operators and tree shapes lift every open-weight model above untrained Gemini-3.1-Pro, an upper bound under matched semantics. A disjoint-domain regime with new operators and vocabulary also raises every open-weight model, confirming the gain is not an artifact of matched semantics, though a gap to the frontier remains. Together these results identify tree-structured expression serialization as a primary limiting factor when models communicate hierarchical structure through natural language.
Chinese Translation
当语言模型进行思维链推理或交换自由文本中间结果时,它们会将结构化信息序列化为自然语言。有多少树结构的组合内容能够在这一瓶颈中幸存?我们提出一种往返协议,对树结构表达式这一问题进行实证解答。生成器将一个程序化生成的算术表达式转化为文字应用题,一个独立的提取器仅从文字应用题中恢复该表达式,而符号等价性检验提供了精确的评估基准。通过评估十六个模型的所有两两组合,我们得到一个通信矩阵,其边际分布可将生成质量与提取质量分离。研究得出三个主要发现。第一,该信道有损且不对称:交换生成模型与提取模型的角色可使准确率变化高达60.4个百分点,而最优组合通过在两端使用不同模型(而非同一模型)达到了92.9%的准确率。第二,至少73.6%的往返失败源于生成端,且难度由树结构(算子数量、深度、右分支程度)决定,而非模型家族。第三,该信道是可训练的:约3600个与评估所用算子和树形结构相同的微调样本,即可使所有开源权重模型超越未训练的Gemini-3.1-Pro——这是语义匹配条件下的性能上界。在采用新算子和新词汇的不相交领域设置中,所有开源权重模型同样获得提升,证实该增益并非语义匹配的产物,尽管与前沿模型仍存在差距。这些结果共同表明,树结构表达式的序列化是模型通过自然语言传递层次结构时的主要限制因素。
cs.AI / 29 / 2609.21519

Learning-to-Optimize as the Missing Architectural Layer of AI-Native Networks

作为AI原生网络缺失架构层的Learning-to-Optimize(学习优化)
Amati, Giambattista, Mangiatordi, Federica, Salvo, Pierpaolo, Pallotti, Emiliano, Angelini, Simone
Abstract
Artificial Intelligence (AI) is becoming a fundamental design principle of future AI-native communication networks, enabling autonomous resource management, adaptive control, and zero-touch network operation. While current AI-native architectures increasingly embed intelligence across network functions, they provide little guidance on how optimisation knowledge should be systematically generated, transferred, and exploited by AI models. This paper argues that the Learning-to-Optimize (L2O) represents the missing architectural layer between optimisation and AI-native intelligence. Rather than viewing optimisation merely as an online decision engine, the proposed paradigm redefines optimisation algorithms as offline knowledge generators that produce high-quality supervisory information for neural surrogate models. The resulting models inherit optimisation expertise while enabling low-latency runtime inference suitable for dynamic network environments. A generic four-stage L2O workflow is introduced, comprising optimisation, knowledge generation, surrogate learning, and runtime inference. Unlike existing Learning-to-Optimize approaches, which primarily focus on algorithm acceleration, the proposed framework establishes L2O as an architectural abstraction applicable across heterogeneous communication and computing systems. The proposed paradigm is illustrated by an NR-V2X relay-selection problem, in which optimisation-generated solutions from a Mixed-Integer Linear Programming (MILP) solver are used to train a Graph Neural Network that can reproduce near-optimal decisions in real time. The presented perspective positions Learning-to-Optimize as a key architectural enabler for future AI-native networks.
Chinese Translation
人工智能(AI)正在成为未来AI原生通信网络的基本设计原则,支持自主资源管理、自适应控制和零接触网络运维。尽管当前的AI原生架构日益将智能嵌入到各个网络功能中,但对于优化知识应如何被AI模型系统化地生成、迁移和利用,却几乎未提供指导。本文认为,Learning-to-Optimize(L2O,学习优化)代表了优化与AI原生智能之间缺失的架构层。该范式不再仅仅将优化视为在线决策引擎,而是将优化算法定义为离线知识生成器,为神经代理模型生成高质量的监督信息。由此得到的模型既继承了优化专业知识,又能实现适合动态网络环境的低延迟运行时推理。本文提出了一个通用的四阶段L2O工作流程,包括优化、知识生成、代理学习和运行时推理。与现有主要关注算法加速的Learning-to-Optimize方法不同,所提出的框架将L2O确立为一种可适用于异构通信与计算系统的架构抽象。本文以NR-V2X中继选择问题为例说明该范式,其中利用混合整数线性规划(MILP)求解器生成的优化解来训练图神经网络(Graph Neural Network),使其能够实时复现近优决策。本文所提出的视角将Learning-to-Optimize定位为未来AI原生网络的关键架构使能技术。
cs.AI / 30 / 2609.21548

Dual-Interest Sequential Product Recommendation With Multi-Granular SSM

基于多粒度状态空间模型(SSM)的双重兴趣序列化产品推荐
Liao, Shuiying, Mok, P. Y.
Abstract
Sequential recommendation aims to predict the next item a user will interact with based on their historical behavior. Advances in Transformers have significantly improved sequential recommendation but are still limited by cost efficiency. Although State Space Models (SSMs) have recently enabled efficient long-range modeling, most existing methods encode each item with a single static contextual role, overlooking the phenomenon of item polysemy. In fact, the same item often plays different semantic roles depending on user context, and existing methods are limited in capturing dynamic behavior across different temporal granularities. In this work, we propose DSRec, a novel dual-interest cross-SSM model that explicitly disentangles item roles across long-term and short-term semantic context. Sequential items are encoded into long-term interest embeddings that capture stable preferences via historical aggregation, and a short-term interest branch that emphasizes local session intent modulated by inter-click time intervals. These interest embeddings are processed through distinct SSM encoders: a full-sequence Mamba for long-term modeling, and a time-modulated SSM that dynamically adjusts state evolution based on temporal gaps. To enable effective cross-granularity alignment, we adopt a residual cross-fusion mechanism that exchanges contextual information between the two branches while preserving semantic independence. Experiments on public benchmarks demonstrate that DSRec outperforms other state-of-the-art methods.
Chinese Translation
序列推荐旨在根据用户的历史行为预测其下一个将交互的物品。Transformer 的进展显著提升了序列推荐的效果,但仍受限于计算效率。尽管状态空间模型(State Space Models, SSMs)近来实现了高效的长程建模,但大多数现有方法以单一的静态上下文角色对每个物品进行编码,忽视了物品多义性现象。事实上,同一物品往往根据用户上下文扮演不同的语义角色,而现有方法在捕捉不同时间粒度上的动态行为方面能力有限。本文提出 DSRec,一种新颖的双重兴趣交叉 SSM 模型,显式地解耦物品在长期与短期语义上下文中的角色。序列物品被编码为通过历史聚合捕捉稳定偏好的长期兴趣嵌入,以及强调由点击间时间间隔调节的局部会话意图的短期兴趣分支。这些兴趣嵌入分别通过不同的 SSM 编码器处理:用于长期建模的全序列 Mamba,以及根据时间间隔动态调整状态演化的时间调制 SSM。为实现有效的跨粒度对齐,我们采用残差交叉融合机制,在保持语义独立性的同时交换两个分支之间的上下文信息。在公开基准数据集上的实验表明,DSRec 的性能优于其他最先进的方法。
cs.AI / 31 / 2609.21599

Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection

超越准确率:用于结构化虚假招聘信息检测的质心引导对比损失
Ahmed, Syed Ali, Raza, Malaika, Siddiqui, Muhammad Shoaib, Rafi, Muhammad
Abstract
Fraudulent job posting detection aims to identify job advertisements that are corrupted either through fake content, misleading information, or negative intent, disrupting the online eco-system of job-seekers and employers. Existing studies in this domain lack effective methods to simultaneously achieve high accuracy and meaningful structure of latent-space representations that capture subtleties among fake posts. To this end, we propose Centroid-Guided Contrastive Loss (CGCL), a loss function which unifies classification with densely formulated clustering to consistently reshape latent-space through a centroid-driven top-$k$ push-and-pull mechanism. The complementary nature of CGCL enables the model to enforce accurate decision boundaries and maintain high clustering compactness, effectively capturing both class separability and latent structure. Extensive experiments demonstrate the state-of-the-art (SOTA) performance of our method on EMSCAD, a public benchmark dataset. The code associated with this work is available at: https://github.com/ali-ahmed925/CGCL_code
Chinese Translation
虚假招聘信息检测旨在识别通过虚假内容、误导性信息或不良意图而被破坏的招聘广告,这些广告破坏了求职者与雇主的在线生态系统。该领域的现有研究缺乏有效方法,无法同时实现高准确率和能够刻画虚假帖子之间细微差异的潜在空间表示的有意义结构。为此,我们提出了质心引导对比损失(Centroid-Guided Contrastive Loss, CGCL),这是一种将分类与密集构造的聚类相统一的损失函数,通过质心驱动的top-$k$推拉机制持续重塑潜在空间。CGCL的互补性使模型能够建立准确的决策边界并保持高聚类紧凑性,有效捕捉类间可分性与潜在结构。大量实验表明,我们的方法在公开基准数据集EMSCAD上取得了最先进(SOTA)的性能。本工作的相关代码可在以下地址获取:https://github.com/ali-ahmed925/CGCL_code
cs.AI / 32 / 2609.21600

Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education

降低学业支持的使用障碍:评估用于解决高等教育求助差异的课程专用RAG系统
Gray, Andy, Hobbs, Jake
Abstract
Access to academic support is a key determinant of student success, yet students experience it unequally: some readily seek help from lecturers or tutors, while others hesitate due to anxiety, fear of judgement, uncertainty about expectations, or low confidence in their understanding. This may be especially evident in computing education, where programming tasks are cumulative and cognitively demanding. Although students increasingly turn to general-purpose generative AI tools, these can produce responses that are inaccurate, insufficiently contextualised, or misaligned with module expectations. This study presents and evaluates Beacon, a course-specific Retrieval-Augmented Generation (RAG) system providing private, immediate, module-aligned academic support. Grounding responses in approved teaching materials, Beacon was designed to lower barriers to help-seeking while encouraging independent learning. Using a design-based research approach, Beacon was developed iteratively and evaluated via mixed methods, combining questionnaires and semi-structured interviews with students and staff at a Higher Education institution. Students described Beacon's responses as closely aligned with module content and more trustworthy than unrestricted generative AI tools, valuing its use of pseudocode and scaffolded explanations over direct solutions. Although participants remained cautious about trusting AI-generated responses without verification, they viewed the system as a valuable first point of support before consulting lecturers or official resources. The findings suggest that carefully designed course-specific AI systems may reduce barriers to academic support by occupying an intermediary space between independent study and formal support. Rather than replacing educators, educational AI may be most valuable when it broadens access to guidance while preserving the pedagogical role of lecturers.
Chinese Translation
获得学业支持是学生成功的关键决定因素,然而学生对此的体验并不均等:一些学生乐于向讲师或导师寻求帮助,而另一些学生则因焦虑、害怕被评判、对期望的不确定或对自身理解能力信心不足而犹豫不前。这在计算机教育中可能尤为明显,因为编程任务具有累积性和高认知负荷。尽管学生越来越多地转向通用生成式AI工具,但这些工具可能产生不准确、缺乏足够上下文或与课程模块要求不一致的回答。本研究提出并评估了Beacon,一个提供私密、即时且与模块内容对齐的学业支持的课程专用检索增强生成(RAG)系统。Beacon以经审核的教学材料为基础生成回答,旨在降低求助障碍,同时鼓励自主学习。采用基于设计的研究方法,Beacon经过迭代开发,并通过混合方法进行评估,结合了针对一所高等教育机构学生和教职员工的问卷调查与半结构化访谈。学生认为Beacon的回答与模块内容高度契合,比不受限制的生成式AI工具更值得信赖,并特别看重其使用伪代码和渐进式引导讲解而非直接给出答案的做法。尽管参与者在未经核实的情况下对AI生成的回答仍持谨慎态度,但他们将该系统视为在咨询讲师或官方资源之前宝贵的首要求助渠道。研究结果表明,精心设计的课程专用AI系统可以通过在自主学习和正式支持之间占据中介空间来降低学业支持的获取障碍。教育AI的价值或许不在于取代教育者,而在于在保留讲师教学角色的同时,扩大获取指导的途径。
cs.AI / 33 / 2609.21619

Calibrating Teacher--Student Discrepancy for On-Policy Distillation

面向在线策略蒸馏的教师-学生差异校准
He, Qiangqiang, Li, Jin, Chen, MingCai
Abstract
On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher--student discrepancy and indiscriminately learned by standard OPD during training. This issue is further exacerbated by privileged OPD, where privileged information induces larger teacher-side likelihood shifts, thereby encouraging the student to learn more of the teacher's own deviation. We introduce \textbf{Calibrated On-Policy Distillation (Cal-OPD)}, which estimates the teacher's self-deviation region through positive and negative privileged interventions and calibrates the original teacher--student discrepancy by retaining only the component that lies beyond this region. Experiments on mathematical reasoning benchmarks show that, while retaining only about 52--65\% of the original teacher--student discrepancy as the optimization signal, Cal-OPD consistently outperforms standard OPD and its variants across model scales.
Chinese Translation
在线策略蒸馏(On-Policy Distillation, OPD)通过学习更强的教师模型与在线策略学生模型之间的词元级差异来改进推理模型。然而,这种差异并不纯粹反映教师与学生之间的能力差距:它还包含源自教师模型自身的偏差,这些偏差混杂在观测到的教师-学生差异中,并被标准OPD在训练时不加区分地学习。特权OPD(privileged OPD)进一步加剧了这一问题:特权信息会引发更大的教师侧似然偏移,从而促使学生学习更多教师自身的偏差。我们提出校准在线策略蒸馏(Calibrated On-Policy Distillation, Cal-OPD),通过正向与负向的特权干预估计教师模型的自我偏差区域,并仅保留超出该区域的差异分量,从而对原始的教师-学生差异进行校准。在数学推理基准上的实验表明,Cal-OPD仅保留约52%–65%的原始教师-学生差异作为优化信号,却在不同模型规模上持续优于标准OPD及其变体。
cs.AI / 34 / 2609.21626

One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction

单一提示词无法适用于所有用户:面向个性化信息抽取的自元演化方法(Self-Meta-Evolve)
Li, Hongliang, Wang, Lu, Xu, Yong, Chen, Hanyang, Hou, Zhitao, Qin, Xiaoting, Ge, Song, Lin, Qingwei, Zhang, Dongmei
Abstract
Large language models (LLMs) are increasingly deployed for enterprise information extraction (IE), where the same document must be reorganized differently for each user. Existing prompt optimization methods, however, rely on a single prompt optimized against a global objective, which is misaligned with the inherent user heterogeneity of real workplaces. We formulate enterprise IE as per-user prompt adaptation under interaction feedback and propose Self-Meta-Evolve, a hierarchical framework that maintains a dedicated prompt for each user and continuously refines it through a dual-loop process: an inner loop that edits structured prompts based on persona-conditioned feedback, and an outer loop that evolves the meta-prompt itself by distilling successful editing patterns. To enable scalable training and evaluation, we release a persona-driven IE benchmark of 292 simulated enterprise users, paired with a reproducible persona-generation pipeline grounded in O*NET occupational taxonomies. On this benchmark, Self-Meta-Evolve achieves a 74.58% success rate, outperforming the strongest prompt-optimization baseline by 13.56 absolute points, and reaches 52.54\% within only two iterations. A double-blind human study with twenty real professionals further confirms that prompts adapted by our framework win against static baselines in 71% of pairwise comparisons.
Chinese Translation
大语言模型(LLM)在企业信息抽取(IE)中的应用日益广泛,但在该场景下,同一份文档需要为不同用户进行不同的重组。然而,现有的提示词优化方法依赖于针对全局目标优化的单一提示词,这与真实工作场所中固有的用户异质性不符。我们将企业信息抽取建模为交互反馈下的逐用户提示词自适应问题,并提出Self-Meta-Evolve,一种分层框架:该框架为每个用户维护专属提示词,并通过双循环机制持续对其进行优化——内循环基于基于用户画像条件化的反馈来编辑结构化提示词,外循环则通过提炼成功的编辑模式来演化元提示词本身。为支持可扩展的训练与评估,我们发布了一个基于用户画像的信息抽取基准,包含292名模拟企业用户,并配有基于O*NET职业分类体系、可复现的画像生成流水线。在该基准上,Self-Meta-Evolve取得了74.58%的成功率,比最强的提示词优化基线高出13.56个绝对百分点,且仅经两轮迭代即可达到52.54%。一项由二十位真实专业人士参与的双盲人工研究进一步证实,在成对比较中,由本框架自适应生成的提示词在71%的情况下胜过静态基线。
cs.AI / 35 / 2609.21672

Accelerating Dense LLMs via L0-regularized Mixture-of-Experts

基于L0正则化混合专家模型的稠密大语言模型加速方法
Zhang, Zhenyu, Yang, Jiudong, Tao, Zhaowen, Chen, Meng
Abstract
Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeable performance degradation, while Mixture-of-Experts (MoE) models require extensive computational resources. In this paper, we propose L0-MoE, a lightweight MoE approach using L0-regularization to accelerate dense LLMs nearly without performance loss. Our method introduces a cluster confusion matrix for domain-aware dataset curation and applies dynamic batching for efficient training. Experiments show that L0-MoE achieves up to 2.5x speedup over dense models while maintaining competitive performance, outperforming existing LLM acceleration baselines.
Chinese Translation
大语言模型(LLMs)虽然取得了出色的性能,但推理速度慢且成本高昂。现有的加速方法往往导致明显的性能下降,而混合专家模型需要大量的计算资源。在本文中,我们提出了L0-MoE,这是一种轻量级的混合专家方法,利用L0正则化来加速稠密大语言模型,且几乎不损失性能。我们的方法引入了聚类混淆矩阵以实现面向领域的数据集筛选,并采用动态批处理以实现高效训练。实验表明,L0-MoE相比稠密模型可实现高达2.5倍的加速,同时保持有竞争力的性能,优于现有的LLM加速基线方法。
cs.AI / 36 / 2609.21677

GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation

GUARD:基于引导式答案-推理蒸馏的大推理模型自然遗忘方法
Yan, Zeyu, Zhou, Guanghao, Qiu, Minghui, Gao, Ming, Chen, Cen
Abstract
Recent advances in large reasoning models (LRMs) have made machine unlearning more challenging, as protected facts or unsafe rationales may surface in intermediate chain-of-thought (CoT) traces before the final answer is produced. Existing unlearning objectives typically suppress the target content or redirect internal representations, but they never specify how the post-forgetting trajectory should continue, which can lead to hallucinated substitutes, malformed boundaries, or repetitive outputs. We argue that LRM unlearning should instead learn a natural forgetting trajectory: a coherent non-disclosing CoT followed by a stable refusal-style answer that replace the original disclosure. To this end, we propose Guided Answer-Reasoning Distillation (GUARD), which converts model-generated unsafe disclosures into safe-exit trajectories, aligns a frozen LRM via guidance tokens, and distills the guided behavior into model parameters.To address the lack of metrics for replacement quality beyond leakage, we further introduce Natural Forgetting Reasoning Score (NFRS), which captures structural stability, fluency, and unsupported substitutes in forgotten outputs. Extensive experiments on R-TOFU and a STAR-1-derived harmful-intent setting show that GUARD substantially reduces unsafe and privacy disclosures across two widely adopted distilled LRMs while preserving reasoning utility. Codes are available at https://github.com/zeyu-Yan/GUARD
Chinese Translation
大推理模型(Large Reasoning Models, LRMs)的最新进展使机器遗忘更具挑战性,因为受保护的事实或不安全的推理依据可能在最终答案生成之前的中间思维链(chain-of-thought, CoT)轨迹中泄露。现有的遗忘目标通常只是抑制目标内容或重定向内部表征,但从未规定遗忘后的轨迹应如何延续,这可能导致幻觉性的替代内容、畸形的拒绝边界或重复输出。我们认为,LRM 遗忘应当转而学习一种自然遗忘轨迹:以连贯的不泄露性 CoT 和稳定的拒绝式答案取代原有的披露内容。为此,我们提出了引导式答案-推理蒸馏方法(Guided Answer-Reasoning Distillation, GUARD),该方法将模型生成的不安全披露转化为安全退出(safe-exit)轨迹,通过引导标记(guidance tokens)对冻结的 LRM 进行对齐,并将引导后的行为蒸馏到模型参数中。针对现有评估缺乏泄漏之外的替换质量度量这一问题,我们进一步提出了自然遗忘推理评分(Natural Forgetting Reasoning Score, NFRS),用以刻画遗忘输出的结构稳定性、流畅性以及无依据的替代内容。在 R-TOFU 以及基于 STAR-1 构建的有害意图场景上的大量实验表明,GUARD 在两种广泛使用的蒸馏型 LRM 上显著减少了不安全内容和隐私泄露,同时保持了推理能力。代码可在 https://github.com/zeyu-Yan/GUARD 获取。
cs.AI / 37 / 2609.21683

Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation

先听后说:基于倾听者面部反应的响应规划用于对话语音生成
Chu, Yunji
Abstract
Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance's emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is unnecessary; correct listener reactions also outperform cyclic mismatches on average. In a contextual-appropriateness study with 20 speech researchers, 76% of judgments prefer Temporal, 9% Text-only, and 15% report no preference. We further connect the predicted response style to a Grad-TTS backbone for end-to-end speech realization. Overall, the results support pre-response listener dynamics as complementary cues for conversational response planning. The source code is available at https://github.com/CYJ1/ReACT-TTS_public.
Chinese Translation
对话语音依赖于对话语境以及倾听者紧随其前的行为。我们提出了ReACT-TTS,一个两阶段框架,在语音实现之前,利用响应前一秒的倾听者面部序列来规划下一话语的情感和韵律。在严格的双人MELD协议上,跨十个随机种子,时序(Temporal)条件化相较于仅文本(Text-only)取得了更高的平均宏F1和VAD一致性,而准确率基本保持不变。消融实验表明,时序建模在各类视觉变体中表现最佳,且显式的早晚差异特征并非必要;正确的倾听者反应平均而言也优于循环错配的反应。在一项包含20名语音研究人员的语境适切性研究中,76%的判断偏好时序(Temporal)方法,9%偏好仅文本(Text-only)方法,15%表示无偏好。我们进一步将预测的响应风格与Grad-TTS骨干网络相连接,实现端到端的语音合成。总体而言,结果支持将响应前的倾听者动态作为对话响应规划的补充线索。源代码可在 https://github.com/CYJ1/ReACT-TTS_public 获取。
cs.AI / 38 / 2609.21748

World Modeling in Transformers

Transformer中的世界建模
Beckmann, Pierre, Queloz, Matthieu, Freitas, Andre
Abstract
Behavioral failures can make a transformer appear to lack a world model even when it has learned faithful representations of its environment. We demonstrate this in TaxiGPT, a transformer trained on random walks through Manhattan whose failures have been interpreted as evidence of an incoherent internal map. Through mechanistic analysis and causal interventions, we show that the model represents intersections and streets, tracks its position, and uses a goal compass to navigate. We trace its failures to interference between superposed intersection features, which disrupts localization within the internal map. Affordance packing, which groups representations of intersections with the same legal moves, helps limit the consequences of these errors. Finally, we propose mechanistic indicators that we use to compare models and show that world-modeling capacities emerge at different stages of training. Our findings motivate a shift from asking whether a model has a world model to mechanistically studying its world modeling: the interacting capacities through which it represents its environment and uses those representations to guide behavior.
Chinese Translation
行为层面的失败可能使Transformer看似缺乏世界模型,即使它已经学习到对其环境的忠实表征。我们在TaxiGPT中展示了这一点:TaxiGPT是一个在曼哈顿随机游走数据上训练的Transformer,其失败曾被解读为内部地图不连贯的证据。通过机制分析与因果干预,我们表明该模型表征了交叉口和街道,能够追踪自身位置,并利用目标指南针进行导航。我们将其失败追溯到叠加的交叉口特征之间的干扰,这种干扰破坏了内部地图中的定位。 affordance packing(将具有相同合法动作的交叉口表征进行分组)有助于限制这些错误的后果。最后,我们提出了用于比较模型的机制性指标,并表明世界建模能力在训练的不同阶段逐步涌现。我们的发现促使人们从询问“模型是否拥有世界模型”转向以机制方式研究其世界建模过程:即模型表征其环境并利用这些表征指导行为的一系列相互作用的能力。
cs.AI / 39 / 2609.21755

ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction

ECG Mirage:揭示并缓解视觉-语言模型中心电图在临床预测中的利用不足问题
Liang, Jinning, Zhu, Mingcheng, Zhu, Tingting
Abstract
Emergency department (ED) decision-making relies on heterogeneous clinical information, including patient history, vital signs, laboratory results, and electrocardiograms (ECGs). Vision--language models (VLMs) can jointly process these modalities, but strong predictive performance does not necessarily imply meaningful use of the correct patient's ECG. We term this failure mode ECG Mirage: apparent multimodal capability without useful dependence on patient-specific ECG information. We distinguish two forms: ECG neglect, where ECGs provide little predictive benefit, and ECG confusion, where matched ECGs outperform no-image inputs but not mismatched ECGs. To evaluate these behaviours, we compare predictions obtained with matched ECGs, outcome-discordant mismatched ECGs, and no-image inputs while holding the clinical text and prediction targets fixed. Across four VLMs on MDS-ED, matched ECGs provide no consistent advantage for either ICU admission or clinical deterioration prediction. We then train four restricted visual prompts using supervised learning followed by conditional direct preference optimisation, while keeping the VLM backbone frozen. The resulting models achieve balanced accuracies of 70.6% for ICU admission and 67.5% for deterioration and increase the matched-versus-mismatched performance gap to approximately 16.5 and 5.5 percentage points, respectively. Overall, our study identifies ECG Mirage in multimodal clinical prediction and introduces visual prompt tuning as an efficient mitigation strategy.
Chinese Translation
急诊科(ED)的临床决策依赖于异构的临床信息,包括患者病史、生命体征、实验室检查结果和心电图(ECG)。视觉-语言模型(VLM)能够联合处理这些模态,但较强的预测性能并不一定意味着模型对正确患者的ECG进行了有意义的利用。我们将这种失败模式称为ECG Mirage(ECG幻象):即表面上的多模态能力,实际上并未有效依赖于患者特定的ECG信息。我们区分了两种形式:ECG忽视(ECG neglect),即ECG几乎不带来预测收益;以及ECG混淆(ECG confusion),即匹配的ECG优于无图像输入,但并不优于不匹配的ECG。为评估这些行为,我们在固定临床文本和预测目标的情况下,比较使用匹配ECG、结局不一致的不匹配ECG以及无图像输入时的预测结果。在MDS-ED数据集上对四个VLM的实验表明,匹配的ECG对ICU入院预测或临床恶化预测均无一致的优势。随后,我们采用监督学习结合条件化直接偏好优化(conditional direct preference optimisation)训练了四个受限视觉提示(visual prompts),同时保持VLM主干网络冻结。所得模型在ICU入院和临床恶化预测上的平衡准确率分别达到70.6%和67.5%,并将匹配与不匹配ECG之间的性能差距分别提升至约16.5和5.5个百分点。总体而言,本研究识别了多模态临床预测中的ECG Mirage现象,并提出视觉提示微调(visual prompt tuning)作为一种高效的缓解策略。
cs.AI / 40 / 2609.21801

LLM-Generated Feature Pools for Time Series Anomaly Detection

利用大语言模型生成的特征池进行时间序列异常检测
Hili, Youssef Attia El, Tiomoko, Malik, Ancourt, Corinne
Abstract
We study how far a simple statistical pipeline can go on univariate time series anomaly detection under a strict selection protocol. The method extracts a small pool of statistics over sliding windows, scores each window with a transductive robust (MAD) model, and selects a feature subset per domain on a held-out tuning split. On TSB-AD-U it reaches $0.529$ per-series VUS-PR, above the best neural ($0.45$) and statistical ($0.44$) entries on the public leaderboard and within $0.06$ of the strongest pretrained foundation model, several of which use more supervision than ours. Ablations locate the cause: across three selection strategies and a hindsight oracle the score moves by $0.031$, and across the aggregation grid by $0.096$, while changing the candidate pool moves it by $0.226$. The candidate pool sets the ceiling; the search over it is second-order. We therefore generate a pool per domain by prompting a multimodal LLM with in-context example windows from that domain. The generated pools match the hand-crafted one under matched selection, and the two cover different domains: selecting over their union improves on the generated pool in all twelve generator-seed pairs and lifts the pipeline to $0.588$, matching the performance of the best entry on the leaderboard.
Chinese Translation
我们研究了在严格的特征选择协议下,一个简单的统计流水线在单变量时间序列异常检测任务上所能达到的性能上限。该方法在滑动窗口上提取少量统计量池,使用转导式鲁棒(MAD)模型对每个窗口进行评分,并在留出的调优子集上为每个领域选择一个特征子集。在 TSB-AD-U 基准上,该方法达到了每序列 0.529 的 VUS-PR,超过了公开排行榜上最优的神经网络方法(0.45)和统计方法(0.44),并且与最强的预训练基础模型相差仅 0.06 以内,而后者中若干模型所使用的监督信息多于我们的方法。消融实验定位了性能的决定因素:在三种选择策略和一个事后预言机之间,分数仅变化 0.031;在不同的聚合方式网格之间,分数变化 0.096;而改变候选特征池则使分数变化 0.226。由此可见,候选特征池决定了性能上限,而对其的搜索则是次要因素。因此,我们通过向多模态大语言模型(LLM)提供来自该领域的上下文示例窗口,为每个领域生成一个特征池。在匹配的选择协议下,所生成的特征池与手工构建的特征池表现相当,且二者覆盖的领域有所不同:在二者的并集上进行选择,在全部十二组“生成器-种子”配对中均优于单独使用生成的特征池,并将流水线的性能提升至 0.588,达到了排行榜上最优方法的水平。
cs.AI / 41 / 2609.21811

MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention

MIST:基于基因组引导的病理组织学注意力的多模态生存预测
Yavuz, Muhammet Sami, Kahya, Sabri Mustafa, Chen, Richard R., Lipkova, Jana, Wiestler, Benedikt
Abstract
Multimodal survival models can combine complementary prognostic information from whole-slide images and genomic profiles, but effective fusion remains challenging amid external cohort shift and computational complexity. To address these challenges, we propose MIST, multimodal survival prediction with genomic-guided histology attention. MIST represents genomic features as tokens and allows them to query compact foundation-model-derived histology context tokens before survival prediction. This design enriches molecular information with histology context rather than merging separately encoded modalities only at the final stage. Training combines discrete-time survival prediction with genomic feature masking, WSI dropout, and paired WSI-genomics contrastive alignment. Across four external evaluations in colon, renal, lung, and glioblastoma cohorts, MIST improves external C-index over standard fusion baselines in the primary comparisons. These results support genomic-guided histology attention as a compact and effective strategy for multimodal oncology outcome prediction. Our code is available at https://github.com/samiyavuuz/MIST .
Chinese Translation
多模态生存模型能够整合来自全切片病理图像与基因组图谱的互补预后信息,但在外部队列偏移和计算复杂性的背景下,有效的融合仍然具有挑战性。为应对这些挑战,我们提出了MIST——一种基于基因组引导的病理组织学注意力的多模态生存预测方法。MIST将基因组特征表示为token,并在生存预测之前使其查询由基础模型导出的紧凑的病理组织学上下文token。这种设计利用病理组织学上下文来丰富分子信息,而非仅在最终阶段合并分别编码的模态。模型训练将离散时间生存预测与基因组特征掩码、WSI dropout以及成对的WSI-基因组对比对齐相结合。在结肠癌、肾癌、肺癌和胶质母细胞瘤队列的四项外部评估中,MIST在主要对比中均提升了外部C-index,优于标准融合基线方法。这些结果支持将基因组引导的病理组织学注意力作为一种紧凑而有效的多模态肿瘤学结局预测策略。我们的代码可在 https://github.com/samiyavuuz/MIST 获取。
cs.AI / 42 / 2609.21841

EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise

EnterpriseVal:量化生成式AI在企业中的效能、可靠性与价值
Ali, Abbas Raza, Siddiqui, Muhammad Ajmal, Zahid, Moona
Abstract
Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. We argue that this is substantially a measurement problem: public benchmarks answer "what can the model do?", whereas a deployment decision requires "is this workflow fit, reliable, safe and worth scaling - here, on our data, under our controls?". We present EnterpriseVal, a use-case-level evaluation system that closes this gap. It comprises (i) a formal specification of the use case and of the frozen socio-technical configuration under test, model, prompts, retrieval, tools, guardrails and human oversight, with an autonomy level and consequence tier that jointly set the required evaluation intensity; (ii) a metric catalogue spanning fidelity, utility, efficiency, reliability, assurance and oversight; (iii) a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference; (iv) a two-tier threshold gate, stated as an executable algorithm, that maps metric vectors with confidence bounds to REJECT/CONDITIONAL/SCALE decisions; and (v) a value-and-risk model in which the reviewer catch rate is a measured parameter. We report a pilot across three workflows in a global bank. In credit-memo drafting, human-graded citation precision reached 88% and hallucination rate 1.6% for the best model against gates of 70% and 5%; in procedure transformation, analyst refinement effort fell from an estimated 27.4 to 2.9 hours per document. We separate established results, documented pilot evidence, the proposed system and open hypotheses, and specify the experiments required for full validation
Chinese Translation
前沿语言模型现已能够产出专业级交付成果,专家评审认为其在相当比例的具有经济价值的任务上可与人类工作相媲美;然而,大多数企业生成式AI(GenAI)计划未能展现出可衡量的业务效果,且预计相当大比例的智能体(agentic)项目将被取消。我们认为这在很大程度上是一个测量问题:公开基准回答的是"模型能做什么?",而部署决策需要回答的是"该工作流在此处、在我们的数据上、在我们的管控条件下,是否适用、可靠、安全且值得扩展?"。我们提出了EnterpriseVal,一个用例级评估系统,以弥合这一差距。该系统包括:(i) 对被测用例及冻结的社会技术配置——模型、提示词、检索、工具、防护栏和人工监督——的正式规范,并配有自主性等级和后果层级,共同决定所需的评估强度;(ii) 涵盖忠实性、效用、效率、可靠性、保障性与监督性的指标目录;(iii) 一种评分协议,通过预测驱动推理(prediction-powered inference),利用经校准的LLM-as-judge评分来扩展盲评专家判断;(iv) 一个以可执行算法表述的两级阈值门控,将带置信区间的指标向量映射为REJECT/CONDITIONAL/SCALE(拒绝/有条件/扩展)决策;(v) 一个价值与风险模型,其中评审员捕获率是一个被测量的参数。我们报告了在一家全球性银行三个工作流中开展的试点。在信贷备忘录起草中,最佳模型的人工评分引用精确率达到88%,幻觉率为1.6%,而门控阈值分别为70%和5%;在流程转换中,分析师的细化工作量从估计的每份文档27.4小时降至2.9小时。我们区分了既有结果、有记录的试点证据、所提出的系统与开放性假设,并明确了完整验证所需的实验。
cs.AI / 43 / 2609.21863

AutoRecLab: Describe the Experiment, Get the Code!

AutoRecLab:描述实验,即可获得代码!
Baumgart, Moritz, Meister, Philipp, Krell, Justus, Schmidt, Michael, Gipp, Bela, Beel, Joeran
Abstract
Empirical evaluation is central to recommender-systems (RecSys) research, but turning experimental designs into executable code remains a manual and error-prone task. We present AutoRecLab, a Python-based autonomous RecSys lab that automates RecSys experiments from natural-language prompts. Given a research idea, AutoRecLab derives explicit experiment requirements, builds and validates a prototype, and iteratively expands it into the requested full experiment. The workflow combines retrieval-augmented generation (RAG) for documentation lookup, static type verification, and execution-steered tree search. In our demonstration, AutoRecLab autonomously implements an explicit-to-implicit feedback conversion study. In a baseline comparison across six algorithms and three datasets, 8 of 9 runs succeed at an average cost of approx- imately $1 per run with GPT-5.4-mini.
Chinese Translation
实证评估是推荐系统(RecSys)研究的核心,但将实验设计转化为可执行代码仍然是一项手动且容易出错的任务。我们提出了AutoRecLab,一个基于Python的自主推荐系统实验室,能够根据自然语言提示自动完成推荐系统实验。给定一个研究想法,AutoRecLab会推导出明确的实验需求,构建并验证原型,然后迭代地将其扩展为所需的完整实验。该工作流程结合了用于文档检索的检索增强生成(RAG)、静态类型验证以及以执行为导向的树搜索。在我们的演示中,AutoRecLab自主实现了一项从显式反馈到隐式反馈转换的研究。在一项涵盖六种算法和三个数据集的基线对比实验中,9次运行中有8次成功,使用GPT-5.4-mini的平均成本约为每次运行1美元。
cs.AI / 44 / 2609.21924

What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence

接下来应该问什么?部分证据下的检索感知问题学习
Qian, Lyucheng, Zhang, John Yuehan, Wang, Pingyu
Abstract
Interactive retrieval under partial evidence is a sequential information-acquisition problem: an agent must decide which question will create the most useful evidence for the next retrieval update. Existing systems train this decision by imitating an offline ordering of candidate QA pairs, although question value is determined by the response it elicits and its downstream effect on retrieval. We establish that candidate discriminativeness and perceived usefulness provide weak supervision for this objective, then introduce RAVEL, a retrieval-aware online reinforcement learning framework for interactive person re-identification. RAVEL initializes from supervised question generation, observes the current Top-4 candidates directly, and optimizes the question policy with rank feedback from the full question-answer-retrieval loop. Experiments on Interactive-PEDES show that RAVEL delivers progressively stronger retrieval performance across five interaction rounds. Further analysis shows that RAVEL reallocates the questioning budget toward localized open-ended attributes, which provide more useful retrieval evidence and yield the largest gains on initially difficult queries.
Chinese Translation
部分证据下的交互式检索是一个序贯信息获取问题:智能体必须决定哪个问题能够为下一次检索更新创造最有用的证据。现有系统通过模仿候选问答对的离线排序来训练这一决策,然而问题的价值实际上由其引发的回答及其对检索的下游影响所决定。我们证明了候选样本的判别性与感知有用性只能为该目标提供弱监督,随后提出了 RAVEL——一个面向交互式行人重识别(person re-identification)的检索感知在线强化学习框架。RAVEL 从有监督问题生成出发进行初始化,直接观察当前 Top-4 候选,并利用完整的问题—回答—检索闭环中的排序反馈来优化问题策略。在 Interactive-PEDES 数据集上的实验表明,RAVEL 在五轮交互中带来了逐步增强的检索性能。进一步分析显示,RAVEL 将提问预算重新分配到局部化的开放式属性上,这些属性提供了更有用的检索证据,并在初始困难的查询上取得了最大收益。
cs.AI / 45 / 2609.21940

AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory

AutoViewMem:面向对话式长期记忆的自配置正交视图
Cao, Zijie, Qu, Xijun, Gu, Zhicheng, Chen, Xiaoshu, Yuan, Duanyang, Hou, Yanning, Zhou, Sihang, Gong, Jianxing, Huang, Jian, Mei, Yang
Abstract
Long-term memory is essential for large language model (LLM) agents to maintain consistency and personalization over extended interactions. Existing memory systems typically rely on fixed granularities or static schemas, but these designs struggle when heterogeneous information, such as preferences, events, constraints, and temporal updates, is embedded in a single mixed representation. The resulting semantic interference makes top-K retrieval sensitive to noise and often leaves relevant evidence poorly ranked. We present AutoViewMem, a data-driven framework that organizes long-term conversational memory into self-configuring, low-overlap semantic views before indexing. AutoViewMem discovers candidate views from interaction traces, selects a compact complementary view set, and uses these views to guide write-time structured extraction of provenance-grounded memories. This representation-first design moves semantic disentanglement from retrieval time to write time, allowing standard top-K similarity search to retrieve focused evidence without explicit routing or iterative retrieval. We further apply offline consolidation to improve memory compactness and consistency. Experiments on the LoCoMo and PersonaMem benchmarks, under both Qwen3-8B and Qwen3-14B backbones, show that AutoViewMem improves long-horizon question answering and personalization over strong memory baselines while preserving a simple inference pipeline.
Chinese Translation
长期记忆对于大语言模型(LLM)智能体在长程交互中保持一致性与个性化至关重要。现有记忆系统通常依赖固定粒度或静态模式,但当偏好、事件、约束和时间更新等异构信息被嵌入到单一混合表示中时,这些设计便难以应对。由此产生的语义干扰使得 top-K 检索对噪声十分敏感,且相关证据往往排序不佳。我们提出 AutoViewMem,这是一个数据驱动框架,它在建立索引之前将长期对话记忆组织为自配置、低重叠的语义视图。AutoViewMem 从交互轨迹中发现候选视图,选择一组紧凑且互补的视图集合,并利用这些视图在写入时引导基于来源的可验证结构化记忆抽取。这种“表示优先”的设计将语义解耦从检索阶段前移至写入阶段,使标准的 top-K 相似度搜索无需显式路由或迭代检索即可聚焦地召回相关证据。我们进一步应用离线整合来提升记忆的紧凑性与一致性。在 LoCoMo 和 PersonaMem 基准上、基于 Qwen3-8B 和 Qwen3-14B 两种骨干模型的实验表明,AutoViewMem 在保持简洁推理流程的同时,其长程问答与个性化性能优于强记忆基线。
cs.AI / 46 / 2609.21962

Learning Cardiac Features: ECG Biometrics Across Time and~Exercise

学习心脏特征:跨时间与运动状态的心电图生物识别
Thiebaud, Luca, Chauchat, Paul, Ouladsine, Mustapha, Delliaux, Stéphane
Abstract
Electrocardiograms (ECGs) carry subject-specific patterns enabling reliable individual discrimination, forming the basis of ECG biometrics. Beyond authentication, this paradigm holds significant potential to secure sensitive cardiac data and to serve as a pretext task in self-supervised learning. Yet, most studies remain confined to singlesession, resting data, leaving robustness to temporal and physiological variations largely untested. We address this gap by evaluating ECG biometrics under realistic conditions involving exercise-induced stress and cross-session variability. A Siamese ResNet with late multi-lead fusion strategy is trained on a large ECG dataset extracted from cardiopulmonary exercise tests and evaluated with a exercise-and time-aware protocol, as well as on public benchmarks. This first extensive assessment of ECG biometrics under combined physiological and temporal variability achieves an intra-session rest-to-peak EER of 1.7% and stateof-the-art 3.9% on the CYBHi dataset. Findings support the presence of an intrinsic cardiac signature resilient to physiological and temporal drift.
Chinese Translation
心电图(ECG)携带个体特异性的模式,能够实现可靠的个体识别,这构成了心电图生物识别技术的基础。除身份认证外,该范式在保护敏感心脏数据安全方面具有重大潜力,并可作为自监督学习中的前置任务(pretext task)。然而,大多数研究仍局限于单次采集的静息态数据,其对时间和生理变化的鲁棒性在很大程度上未经检验。我们通过在涉及运动应激和跨次采集变异性的现实条件下评估心电图生物识别技术来填补这一空白。我们采用带有后期多导联融合策略的孪生ResNet(Siamese ResNet),在从心肺运动试验中提取的大规模心电图数据集上进行训练,并使用兼顾运动和时间因素的协议以及公开基准数据集进行评估。这是首次在生理和时间变异性共同作用下对心电图生物识别进行的广泛评估,取得了单次采集内从静息到运动峰值1.7%的等错误率(EER),并在CYBHi数据集上达到了3.9%的最先进水平。研究结果表明,心脏存在一种对生理和时间漂移具有韧性的内在心脏特征签名。
cs.AI / 47 / 2609.21996

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal

语言模型的测谎仪:读取模型不愿透露的知识
Dingeto, Hiskias
Abstract
Large language models can hold knowledge they do not report. A model may sandbag on a capability evaluation, or answer against what it internally knows, and its outputs alone cannot tell whether it is hiding an answer or simply does not have one. We borrow the Concealed Information Test, a forensic method that identifies guilty knowledge by presenting a suspect with the true detail among plausible decoys and measuring a stronger response to the item they recognize. Our method, Probe of Internal Recognition (PIR), does the same inside a model. It presents a question with its candidate answers and reads, from the model's internal states, which candidate the model recognizes as correct. PIR is reference-free, needing no honest reference model and no labeled truth corpus. Across eight models from five families (Gemma, Qwen, Llama, Mistral, and Phi), PIR recovers the recognized answer at 0.70 to 0.87 balanced accuracy, well above the 0.28 to 0.40 unknown-item baseline and the 0.25 chance rate. It stays readable across every form of concealment we test, from prompted deception and trained sandbagging to external password-locked and circuit-broken checkpoints, with recognition between 0.85 and 0.93. When the model hides a known answer, recognition stays high. When unlearning removes the knowledge, recognition drops to the level of a question the model never knew. PIR therefore separates a model that will not answer from one that cannot, which supports sandbagging audits and unlearning verification. The signal is causal, adds information beyond black-box behavioral cues, and extends from multiple-choice questions to free-form generation.
Chinese Translation
大型语言模型可能持有其并未报告的知识。模型可能在能力评估中故意隐瞒(sandbagging),或给出与内部认知相悖的回答,而仅凭其输出无法判断它是在隐藏答案,还是确实不知道答案。我们借用了 concealed information test(隐蔽信息测试),这是一种法医学方法,通过向嫌疑人展示真实细节与貌似合理的干扰项,并测量其对所识别项目的更强反应来鉴别其是否知情。我们的方法——内部识别探测(Probe of Internal Recognition, PIR)——在模型内部进行同样的操作:它向模型呈现一个问题及其候选答案,并从模型的内部状态中读出模型识别为正确的候选答案。PIR 无需参考(reference-free),既不需要诚实的参考模型,也不需要标注的真实语料库。在来自五个模型家族(Gemma、Qwen、Llama、Mistral 和 Phi)的八个模型上,PIR 以 0.70 至 0.87 的平衡准确率恢复出模型所识别的答案,远高于 0.28 至 0.40 的未知项目基线和 0.25 的随机水平。在我们测试的所有隐藏形式下——从提示诱导的欺骗和训练出的故意隐瞒,到外部的密码锁定(password-locked)和电路破坏(circuit-broken)检查点——识别准确率均保持在 0.85 至 0.93 之间。当模型隐藏已知答案时,识别率保持在高水平;而当 unlearning(知识遗忘)移除了该知识时,识别率降至与模型从未知晓的问题相同的水平。因此,PIR 能够区分"不愿回答"的模型与"无法回答"的模型,从而支持故意隐瞒行为审计和 unlearning 验证。该信号是因果性的,能够提供超出黑盒行为线索之外的信息,并可从多项选择题扩展到自由文本生成。
cs.AI / 48 / 2609.22068

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

CodeMidas:从代码本身扩展智能体编码强化学习环境
Ye, Bowen, Li, Lei, Li, Shicheng, Yue, Zihao, Zhang, Linghao, Lv, Hanglong, Liu, Yuanxin, Ma, Wenhan, Tian, Hao, Li, Rang, Dong, Jinhao, Zhao, Yikai, Deng, Xiangwei, Zhang, Hailin, Zhao, Liang, Liu, Qi, Kong, Lingpeng, Yang, Tong, Luo, Fuli
Abstract
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input. CodeMidas allocates agentic compute to every stage of environment construction: agents explore implemented functionality to formulate behavioral specifications, construct tests grounded in execution of the original code, and validate and filter candidate tasks through execution checks and repeated solution rollouts. The resulting dataset has 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on these tasks with GRPO improves performance on all five diverse benchmarks, covering issue repair (DeepSWE + 11.7%), whole-program construction (ProgramBench +17%), and terminal work (Terminal-Bench v2.1 +8.5%). Ablations show that increasing the number of high-quality training tasks improves performance. Trajectory analysis shows the RL-trained agent demonstrates better behaviors like increasing codebase exploration and more diverse self-verification. These results establish source code as a scalable foundation for constructing RL environments that improve coding agents across diverse software tasks.
Chinese Translation
通过强化学习(RL)训练强大的编码智能体需要具有可靠验证器的多样化任务。开源代码库为这类任务提供了丰富的来源,而现有方法通常依赖于议题(issue)和提交(commit)等开发工件,限制了可提取任务的范围。为了更好地扩展强化学习环境,我们提出了 CodeMidas,这是一种智能体流水线,它以源代码作为唯一的任务特定输入,将现有代码库中已实现的功能转化为可执行的强化学习环境。CodeMidas 在环境构建的每个阶段都分配智能体计算资源:智能体通过探索已实现的功能来制定行为规范,基于原始代码的执行构建测试,并通过执行检查和重复的解决方案推演(rollout)来验证和筛选候选任务。最终得到的数据集包含来自 3,185 个开源代码库的 5,545 个训练任务,涵盖 23 种编程语言和 15 个技术领域。使用 GRPO 在这些任务上训练 MiMo-V2.5,在五个多样化基准测试中均提升了性能,涵盖议题修复(DeepSWE +11.7%)、整程序构建(ProgramBench +17%)和终端工作(Terminal-Bench v2.1 +8.5%)。消融实验表明,增加高质量训练任务的数量能够提升性能。轨迹分析显示,经过强化学习训练的智能体表现出更优的行为,例如更多的代码库探索和更多样化的自我验证。这些结果表明,源代码可以作为构建强化学习环境的可扩展基础,从而在多样化软件任务中改进编码智能体。
cs.AI / 49 / 2609.22086

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Designer-RSI:从用户流量中演化程序性记忆以实现智能体化平面设计
Du, Hongyang, Yan, Lan, Flores, Christian, Kadav, Asim
Abstract
Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience. The memory widens by acquiring procedures for recurring uncovered subtasks and deepens by revising existing procedures against their own successful and failed executions, while a matched replay gate admits only changes that repair failures without regressing observed successes. Five rounds over 1,406 real user briefs and 1,869 automatically graded trajectories, with no weight updates and no human labels, grow the bank from 76 documentation-derived skills to 139 and raise GenEval2 execution success on Claude-Sonnet-4 from 72.7% to 99.3% (+11.99 points in generation quality), with 61.8% and 67.6% win rates against the no-skill agent across four specialized design benchmarks on Claude-Sonnet-4 and Claude-Opus-4.6. We further show the two mechanisms are effective in combination: on 200 held-out briefs from user-traffic benchmark, widening or deepening alone reaches a 49.4% / 48.6% win rate over the no-skill agent, while their combination reaches 58.5% (p = 0.025). Procedural memory offers a practical route to continual adaptation of agents under noisy, unverifiable feedback.
Chinese Translation
专业平面设计是一项长时程的智能体任务,其结构化、可编辑的作品由许多相互依赖的操作生成,然而其结果缺乏可靠的程序化评估标准(oracle)。我们提出一个持续适应框架:一个冻结的前沿模型通过230多个工具操作专业设计软件,同时一个由自然语言技能构成的外部程序性记忆从经验中积累并精炼可复用的设计流程。该记忆通过为反复出现的未覆盖子任务获取新流程而扩展宽度,并根据流程自身成功与失败的执行记录对其进行修订而加深深度;同时,一个匹配的回放门控仅接受那些修复失败且不使已观测到的成功出现退化的改动。在1,406个真实用户需求和1,869条自动评分轨迹上进行的五轮迭代中,无需任何权重更新和人工标注,技能库从基于文档的76项技能增长到139项,并将Claude-Sonnet-4在GenEval2上的执行成功率从72.7%提升至99.3%(生成质量提高11.99分),在Claude-Sonnet-4和Claude-Opus-4.6上的四个专业设计基准测试中,相对于无技能智能体分别取得61.8%和67.6%的胜率。我们进一步证明两种机制结合使用更为有效:在来自用户流量基准的200个保留需求上,单独扩展宽度或加深深度相对于无技能智能体的胜率为49.4% / 48.6%,而二者结合的胜率达到58.5%(p = 0.025)。程序性记忆为智能体在嘈杂且无法验证的反馈下实现持续适应提供了一条可行的路径。
计算机视觉 (Computer Vision)
74
cs.CV / 1 / 2609.20869

TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision

TAPe+ML:一种面向多任务计算机视觉的紧凑结构化表示
Kurinov, Sergey, Upatov, Alexey
Abstract
We present TAPe+ML v3, a compact computer vision system based on TAPe (Theory of Active Perception), a structured representation that encodes relations among perceptual elements before recognition. Instead of operating directly on pixel tensors, the system uses a shared TAPe representation and a modular recognition architecture for image classification, object detection, and instance segmentation. TAPe+ML v3 combines background and contour processing, local object localization, prototype-based classification, and a coordinator for specialized submodels. Across the reported experiments, it uses fewer than 100,000 parameters. On COCO object detection, it obtains 84.7 mAP50 and 65.3 mAP50-95. On COCO instance segmentation, it obtains 80.7 mask mAP50 and 58.4 mask mAP50-95. In classification experiments, it reaches 92 percent validation accuracy on Imagenette under an identical-training comparison with a raw-pixel baseline, and 89.9 percent Top-1 accuracy on ImageNet-Real. We also evaluate compactness in video scene detection and adaptation under distribution shift in an industrial pilot. The results suggest that shifting part of the modeling burden from network parameters to a structured input representation can support compact multi-task vision systems with reduced data, memory, and compute requirements.
Chinese Translation
我们提出了 TAPe+ML v3,一个基于 TAPe(主动感知理论,Theory of Active Perception)的紧凑计算机视觉系统,该结构化表示在识别之前编码感知元素之间的关系。该系统不直接在像素张量上操作,而是使用共享的 TAPe 表示和模块化识别架构来执行图像分类、目标检测和实例分割。TAPe+ML v3 结合了背景与轮廓处理、局部目标定位、基于原型的分类,以及一个用于协调专用子模型的协调器。在所报告的实验中,该系统使用的参数少于 10 万个。在 COCO 目标检测任务上,它取得了 84.7 的 mAP50 和 65.3 的 mAP50-95。在 COCO 实例分割任务上,它取得了 80.7 的 mask mAP50 和 58.4 的 mask mAP50-95。在分类实验中,在与原始像素基线进行相同训练条件对比下,它在 Imagenette 上达到 92% 的验证准确率,并在 ImageNet-Real 上达到 89.9% 的 Top-1 准确率。我们还评估了该系统的紧凑性,包括视频场景检测以及在工业试点中分布偏移下的适应能力。结果表明,将部分建模负担从网络参数转移到结构化输入表示,可以支持数据、内存和计算需求更低的紧凑多任务视觉系统。
cs.CV / 2 / 2609.20962

MemeTAG: Keyword-Driven Meme Classification through Tag Embedding Reconstruction

MemeTAG:基于标签嵌入重构的关键词驱动模因分类方法
Sharma, Akshit, Patil, Prashant W.
Abstract
The proliferation of harmful internet memes poses a significant societal threat, yet their automated classification remains a formidable algorithmic challenge due to the nuanced, multimodal nature of their content. To address this, we introduce MemeTAG, a novel dual-objective framework that pioneers a keyword-aware approach to meme classification. Our core innovation is a two-part semantic guidance mechanism: first, we leverage a pretrained Vision-Language Model to generate a set of descriptive keywords, that capture the high-level semantics. Second, we introduce the Aggregated Tag Inference Network (ATIN), an attention-based module that distills these keywords into a single, rich semantic embedding. This embedding serves as a target for a novel auxiliary reconstruction loss, which compels the model to learn deeply aligned visual and textual features. This approach, combined with an efficient three-stage training strategy, establishes a new state-of-the-art on the HarMeme, Hateful Memes Challenge (HMC), and PrideMM datasets, decisively outperforming existing state-of-the-art methods.
Chinese Translation
有害网络模因(meme)的泛滥对社会构成了重大威胁,然而由于其内容具有微妙的多模态特性,对其进行自动分类仍然是一项艰巨的算法挑战。为解决这一问题,我们提出了MemeTAG,一种新颖的双目标框架,开创了面向模因分类的关键词感知方法。我们的核心创新是一种由两部分组成的语义引导机制:首先,我们利用预训练的视觉-语言模型生成一组描述性关键词,以捕捉高层语义;其次,我们提出了聚合标签推理网络(Aggregated Tag Inference Network, ATIN),这是一个基于注意力的模块,可将这些关键词提炼为单一且丰富的语义嵌入。该嵌入作为新型辅助重构损失的优化目标,促使模型学习深度对齐的视觉与文本特征。该方法结合高效的三阶段训练策略,在HarMeme、仇恨模因挑战赛(Hateful Memes Challenge, HMC)和PrideMM数据集上确立了新的最先进水平,显著优于现有的最先进方法。
cs.CV / 3 / 2609.20975

Image-Derived PM10 Estimation in Cattle Feedlot Using Machine Learning: Addressing Concentration Ranges Beyond Existing Digital Imaging Methods

基于机器学习的肉牛育肥场图像PM10浓度估算:应对超出现有数字成像方法的浓度范围
Peanusaha, Sirapoom, Ferguson, Greg B., Bush, K. Jack, Li, Peiyang, Auvermann, Brent W.
Abstract
Affordable dust monitoring remains a pressing need for the cattle feedlot industry, yet camera-based PM estimation, despite its growing body of research in urban air quality settings, has not been evaluated under the extended concentration ranges characteristic of intensive livestock operations. This study developed an image-based approach using contrast panel features and machine learning to estimate PM10 concentrations in a commercial cattle feedlot, where hourly average PM10 ranged from 250 to 1,000 ug/m^-3 and instantaneous concentrations reached 5,000 to 20,000 ug/m^-3. Grayscale images were captured during the evening dust peak period, and features including panel contrast, black and white panel pixel values, and overall image brightness were extracted. The model also incorporated recent past values from preceding images and solar zenith angle as predictors. Among the candidate models evaluated, XGBoost achieved the highest predictive performance, with an R^2 of 0.792 and a median absolute error of 103 ug/m^-3. Feature importance analysis revealed that (a) panels positioned farthest from the camera contributed most strongly to predictions and (b) that black panel pixel values were more sensitive than white panel values to changes in PM10 concentration. Prediction accuracy during the sunset transition, which coincides with the onset of the feedlot evening dust peak, remains an area for further refinement. These findings demonstrate the feasibility of image-based PM10 estimation across PM concentration ranges substantially exceeding those reported in prior urban studies and provide practical guidelines for future deployment in feedlot environments.
Chinese Translation
经济实惠的粉尘监测仍是肉牛育肥场行业亟待解决的问题。尽管基于相机的PM浓度估算在城市空气质量研究中已积累了大量成果,但尚未在集约化畜禽养殖特有的高浓度范围内得到评估。本研究开发了一种基于图像的方法,利用对比度标板特征和机器学习估算商业化肉牛育肥场中的PM10浓度。该育肥场的小时平均PM10浓度范围为250至1,000 ug/m^-3,瞬时浓度可达5,000至20,000 ug/m^-3。研究在傍晚扬尘高峰期间采集灰度图像,提取了标板对比度、黑白标板像素值以及整体图像亮度等特征。模型还纳入了前序图像的近期历史值和太阳天顶角作为预测变量。在所评估的候选模型中,XGBoost取得了最高的预测性能,R^2为0.792,中位绝对误差为103 ug/m^-3。特征重要性分析表明:(a) 距离相机最远的标板对预测的贡献最大;(b) 黑色标板像素值对PM10浓度变化的敏感度高于白色标板像素值。日落过渡时段(与育肥场傍晚扬尘高峰的开始时间重合)的预测精度仍有待进一步改进。这些发现证明了基于图像的PM10估算方法在远超以往城市研究报告的PM浓度范围内的可行性,并为未来在育肥场环境中的部署提供了实用指导。
cs.CV / 4 / 2609.21012

Fragment-Aware Vision Transformers for Fresco-Fragment Style Classification

面向壁画碎片风格分类的碎片感知视觉Transformer
Miketek, Sara, Barchielli, Biagio, Kajla, Nadeem Iqbal, Aslan, Sinem
Abstract
Artistic style classification is usually studied on complete artworks, where models can exploit global composition, spatial organisation, and iconographic structure. In archaeological settings, however, artworks often survive only as fragmented remains, forcing recognition from incomplete, irregular, and context-limited visual evidence. We study fresco-fragment style classification using a progressive transformer-based framework. Starting from a ViT-B/16 baseline, we introduce foreground-guided masking to suppress background-only tokens, inpainting-based geometric regularisation to align irregular fragment supports with the ViT patch grid, and a supervised contrastive objective that operates on predictive distributions through a Kullback-Leibler similarity and consistently improves every branch. We combine the branches with a deliberately simple learnable logit ensemble. Experiments on CLEOPATRA and POMPAAF show that fragment-aware modelling improves over the standard ViT baseline, with the ensemble increasing accuracy from 0.604 to 0.656 and macro-F1 from 0.596 to 0.648 on CLEOPATRA, and outperforming the best single branch in four of six fragmentation settings on POMPAAF. We additionally evaluate a more complex graph-fusion variant and find that it matches the simple ensemble on POMPAAF while offering only a small, dataset-specific gain on CLEOPATRA, which does not justify its added complexity. Beyond these empirical gains, our contribution is twofold: a distribution-level contrastive objective that consistently sharpens single-branch recognition, and an interpretability analysis that verifies the models exploit genuine painted evidence, while quantifying that the inpainting-based branch draws part of its attribution from the synthesised surround.
Chinese Translation
艺术风格分类通常在完整的艺术品上进行研究,此时模型可以利用全局构图、空间组织和图像学结构。然而在考古场景中,艺术品往往仅以碎片形式存留,这迫使模型必须基于不完整、不规则且上下文受限的视觉证据进行识别。我们使用一个渐进式的基于Transformer的框架来研究壁画碎片风格分类。从ViT-B/16基线出发,我们引入了前景引导的掩码机制以抑制仅包含背景的token,引入基于图像修复的几何正则化以使不规则的碎片边界与ViT的patch网格对齐,并引入一种在预测分布上基于Kullback-Leibler相似度进行运算的监督对比目标,该目标持续提升每个分支的性能。我们以一个刻意简单的可学习logit集成方法组合各分支。在CLEOPATRA和POMPAAF数据集上的实验表明,碎片感知建模优于标准ViT基线:在CLEOPATRA上,集成方法将准确率从0.604提升至0.656,宏平均F1从0.596提升至0.648;在POMPAAF上,集成方法在六种碎片化设置中的四种优于最佳单分支。我们还评估了一个更复杂的图融合变体,发现其在POMPAAF上与简单集成方法表现相当,而在CLEOPATRA上仅带来微小的、特定于数据集的收益,这不足以证明其额外复杂性的合理性。除这些实证收益外,我们的贡献有两个方面:一是持续提升单分支识别能力且在分布层面的对比目标;二是一系列可解释性分析,验证了模型确实利用了真实的绘画证据,同时量化了基于图像修复的分支的部分归因来自合成的背景区域。
cs.CV / 5 / 2609.21018

MAGIC: Marginal-Guided Compression with Optimal Transport for Efficient Visual Document Retrieval

MAGIC:基于最优传输的边际引导压缩方法,面向高效视觉文档检索
Yuan, Xu, Liu, Hua, Fan, Wenqi, Li, Qing
Abstract
Recent visual document retrieval (VDR) systems such as ColPali use multi-vector page embeddings, in which patch-level vectors enable fine-grained evidence matching but incur substantial index storage and MaxSim scoring overhead. Post-hoc merging offers a practical route to efficient VDR by reducing this cost without retraining the retriever, but its uniform reconstruction objectives are poorly aligned with the sparse, non-uniform patch usage induced by late-interaction retrieval. Under aggressive compression, this misalignment can preserve rarely used patches while concentrating retrieval activity on too few retained representatives. To address this misalignment, we propose Marginal-Guided Compression with Optimal Transport (MAGIC), a training-free post-hoc compressor for efficient retrieval with frozen multi-vector embeddings. MAGIC derives a MaxSim-induced compression surrogate and optimizes it through a two-marginal entropic optimal-transport formulation, where a retrieval-demand source marginal prioritizes high-use patches and a balanced target marginal regularizes retained-facet usage. Across ViDoRe benchmarks, keep ratios, and retrieval backbones, MAGIC consistently outperforms strong post-hoc compressors, with particularly large gains in the aggressive-compression regime; component ablations verify the complementary effects of its two marginals. We release the code at: https://github.com/xandery-geek/MAGIC.
Chinese Translation
近年来,诸如ColPali等视觉文档检索(VDR)系统采用多向量页面嵌入,其中补丁级(patch-level)向量支持细粒度的证据匹配,但会带来可观的索引存储开销和MaxSim打分开销。事后合并(post-hoc merging)通过在不重新训练检索器的情况下降低这些成本,为高效VDR提供了一条实用途径;然而,其均匀重构目标与晚期交互检索所引发的稀疏、非均匀补丁使用模式难以对齐。在激进压缩的情况下,这种失配可能导致保留很少被使用的补丁,同时使检索活动过度集中于过少的保留代表向量上。为解决这一失配问题,我们提出基于最优传输的边际引导压缩方法(Marginal-Guided Compression with Optimal Transport,MAGIC),这是一种无需训练的事后压缩器,可在冻结的多向量嵌入上实现高效检索。MAGIC推导了一个由MaxSim诱导的压缩代理目标,并通过双边际熵正则化最优传输形式对其进行优化:其中,检索需求源边际(retrieval-demand source marginal)优先考虑高使用率的补丁,而平衡目标边际(balanced target marginal)则对保留部分的利用率进行正则化。在ViDoRe基准、不同保留比例以及不同检索骨干网络上,MAGIC始终优于强基线事后压缩器,在激进压缩场景下收益尤为显著;组件消融实验验证了两个边际的互补作用。代码已发布于:https://github.com/xandery-geek/MAGIC。
cs.CV / 6 / 2609.21095

MarsFM: Shading-Regularized Flow Matching for Martian Relief Estimation

MarsFM:用于火星地形估计的阴影正则化流匹配方法
Juston, Marius F. R.
Abstract
We present MarsFM, an image-conditioned latent flow-matching model for local Martian relief estimation from single-band HiRISE RED orthoimagery. The method combines a pretrained generative prior with stereo-derived geometric supervision and a differentiable Lunar--Lambert shading objective. Relief, normal, gradient, curvature, and ordinal terms constrain complementary aspects of terrain structure, while a positive-affine-invariant image comparison constrains rendered appearance. An evaluation comprising 2024 gathered patch records per integration-step count yields mean affine-aligned RMSE between 0.0935 and 0.0957 in normalized signed-log relief space for one to twenty Euler steps. These scores measure agreement with VAE-reconstructed references on positive-reference support. Their narrow range supports low-step inference under this protocol. Spatial, differential, and spectral diagnostics show broad terrain correspondence alongside smoothing, amplitude compression, and boundary mismatch. MarsFM provides a framework for combining learned terrain priors with image-based constraints; establishing improved physical terrain resolution requires matched baselines and independent high-resolution reference data. Data: https://huggingface.co/datasets/SuperComputer/mars_hirise_dtm_processed-6aa9b66ba461e07f; code: https://github.com/Marius-Juston/MarsRecon.
Chinese Translation
我们提出了 MarsFM,一个以图像为条件的潜在流匹配模型,用于从单波段 HiRISE RED 正射影像中估计火星局部地形起伏。该方法将预训练的生成先验与由立体测量获得的几何监督以及可微分的 Lunar--Lambert 阴影目标函数相结合。地形起伏、法向量、梯度、曲率和序数项约束地形结构的互补方面,而一个正仿射不变的图像比较约束渲染外观。在每次积分步数的评估中包含2024条采集的图像块记录,在归一化的符号对数(signed-log)地形起伏空间中,从1步到20步欧拉积分的均值仿射对齐 RMSE 介于0.0935至0.0957之间。这些分数衡量了与 VAE 重建参考结果在正参考支撑域上的一致性。其狭窄的分数范围支持在该评估协议下的低步数推理。空间、微分和频谱诊断结果显示了广泛的地形对应性,同时也存在平滑、幅值压缩和边界失配现象。MarsFM 提供了一个将学习得到的地形先验与基于图像的约束相结合的框架;建立更优的物理地形分辨率需要匹配的基线方法和独立的高分辨率参考数据。数据:https://huggingface.co/datasets/SuperComputer/mars_hirise_dtm_processed-6aa9b66ba461e07f;代码:https://github.com/Marius-Juston/MarsRecon。
cs.CV / 7 / 2609.21176

4DGS-Fixer: Generative Sparse-View 4D Gaussian Splatting with Iterative Refinement Guided by Video Diffusion Priors

4DGS-Fixer:基于视频扩散先验引导迭代优化的生成式稀疏视角4D高斯泼溅
Huang, Haitao, Zhao, Shenghao, Tian, Boyuan, Chng, Shin-Fang, Yang, Songlin, Lim, Sheila, Zhan, Huangying, Xu, Yi, Rao, Anyi, Guan, Frank
Abstract
This paper addresses the challenges of dynamic scene synthesis from sparse-view videos. Existing methods employ geometric priors, adaptive optimization, or density-control strategies to improve 4D Gaussian modeling under sparse observations. However, they cannot fundamentally resolve the ill-posed problem caused by insufficient observations and missing scene information. Moreover, sparse-view 4D Gaussian Splatting (4DGS) often suffers from poor geometric initialization: with only a few input views, COLMAP typically reconstructs sparse and incomplete point clouds, leaving large scene regions without sufficient Gaussian support and making them difficult to recover through subsequent optimization. To address these limitations, we propose a novel iterative refinement framework based on a video diffusion model to improve the completeness and consistency of dynamic 4D scenes. Specifically, we first estimate multi-view depth maps and fuse them into dense point clouds to provide more complete geometric initialization for a dynamic 4DGS representation. We then employ a pretrained video restoration model to refine sequences rendered along novel camera trajectories at different time steps. The restored sequences serve as pseudo-supervision to regularize and iteratively refine the 4DGS representation. Experiments on a widely used benchmark dataset demonstrate that our method substantially outperforms existing baselines, achieving nearly a 2 dB PSNR improvement over the previous best-performing method.
Chinese Translation
本文研究从稀疏视角视频合成动态场景的挑战。现有方法采用几何先验、自适应优化或密度控制策略来改进稀疏观测下的4D高斯建模,但无法从根本上解决因观测不足和场景信息缺失所导致的病态问题。此外,稀疏视角4D高斯泼溅(4D Gaussian Splatting, 4DGS)通常面临几何初始化质量差的问题:在仅有少量输入视角的情况下,COLMAP 通常只能重建稀疏且不完整的点云,导致场景中大范围区域缺乏足够的高斯支撑,难以通过后续优化加以恢复。针对这些局限性,我们提出了一种基于视频扩散模型的新型迭代优化框架,以提升动态4D场景的完整性与一致性。具体而言,我们首先估计多视角深度图并将其融合为稠密点云,为动态4DGS表示提供更完整的几何初始化。然后,我们采用预训练的视频修复模型,对不同时间步沿新相机轨迹渲染的序列进行优化。修复后的序列作为伪监督信号,用于正则化并迭代地优化4DGS表示。在广泛使用的基准数据集上的实验表明,我们的方法显著优于现有基线,相比此前最优方法实现了近2 dB的PSNR提升。
cs.CV / 8 / 2609.21199

OnomatoBridge: Onomatopoeia Translation and Rendering Pipeline in Manga

OnomatoBridge:漫画中的拟声词翻译与渲染流水线
Taniguchi, Takara, Shimoda, Wataru, Yamaguchi, Kota, Nakayama, Hideki
Abstract
Manga is a comic drawn by black and white paints gaining popularity around the world. Onomatopoeia in Manga specifically appeals to the audience with its unique visual styles, which convey sound, motion, and emotion. Visual onomatopoeia translation requires the clean replacement of Japanese onomatopoeia with onomatopoeia in the other language while preserving their visual style. Existing approaches often produce residual artifacts or style inconsistency when removing the Japanese onomatopoeia and rendering stylized English onomatopoeia. To approach these problems, we present OnomatoBridge, a filtering pipeline for visual onomatopoeia translation. We evaluate OnomatoBridge from Japanese to English on the Manga109 onomatopoeia dataset and compare it with baseline image editing models. Experimental results show that the filtered outputs by the proposed method outperform those of conventional methods. OnomatoBridge improves English text correctness by roughly 10 to 25 points and reduces residual Japanese text by about 20 to 50% in relative terms.
Chinese Translation
漫画是一种以黑白绘制的漫画形式,在全球日益流行。漫画中的拟声词以其独特的视觉风格吸引读者,传达声音、动作和情感。视觉拟声词翻译需要在保留原有视觉风格的前提下,将日语拟声词干净地替换为目标语言的拟声词。现有方法在去除日语拟声词和渲染风格化英文拟声词时,常常产生残留痕迹或风格不一致的问题。针对这些问题,我们提出了OnomatoBridge,一个用于视觉拟声词翻译的过滤流水线。我们在Manga109拟声词数据集上评估了从日语到英语的OnomatoBridge,并与基线图像编辑模型进行比较。实验结果表明,本方法过滤后的输出优于传统方法。OnomatoBridge将英文文本正确性提升约10至25个百分点,并将残留日文文本相对减少约20%至50%。
cs.CV / 9 / 2609.21207

Hand-Aware Transition Modeling for Bimanual Procedural Anomaly Detection

面向双手程序化异常检测的手部感知转换建模
Wen, Di, Weissert, Jimmy, Scherrer, Luc Maria, Zöllner, Cedric, Yang, Kailun, Liu, Ruiping, Chen, Yufan, Wei, Jiale, Zheng, Junwei, Peng, Kunyu
Abstract
Procedural anomaly detection in bimanual assembly requires judging each hand action against the execution so far. A corrective action may look unusual in isolation, while a visually plausible action can violate the order of the procedure. We present HACT, a transition model over predicted per-hand events. A role-preserving history keeps the concurrent responsibilities of both hands, and a marked temporal point process assigns each observed transition a semantic and temporal surprisal. A supervised evidence head and a two-state filter convert these surprisals into per-hand anomaly posteriors. A recovery-aware protocol on predicted events and participant-disjoint folds reports the recovery false-positive rate at an operating point selected on validation participants. On two bimanual power-tool procedures HACT has the highest AUPRC and F1 among the compared methods and the fewest recovery alarms. Applied without retraining to a different assembly order of the same product, it retains the highest AUPRC and F1. The source code is available at https://github.com/Kratos-Wen/HACT.
Chinese Translation
双手装配中的程序化异常检测需要将每只手的动作与目前为止的执行过程进行比对判断。一个纠正性动作在孤立观察时可能显得异常,而一个视觉上看似合理的动作可能违反程序的操作顺序。我们提出了HACT(Hand-Aware Transition Modeling),一个基于预测的每只手事件的转换模型。其中,保持角色一致性的历史记录保留了双手同时承担的职责,而带标记的时间点过程为每个观测到的转换赋予语义和时间的意外度。一个有监督的证据头和一个两状态过滤器将这些意外度转换为每只手的异常后验概率。在预测事件上采用考虑恢复操作的评估协议,并使用参与者不相交的数据划分,在由验证参与者选定的操作点上报告恢复误报率。在两个双手电动工具操作程序上,HACT在所比较的方法中取得了最高的AUPRC和F1分数,且恢复误报警告最少。在无需重新训练的情况下应用于同一产品的不同装配顺序时,它仍保持了最高的AUPRC和F1分数。源代码可在 https://github.com/Kratos-Wen/HACT 获取。
cs.CV / 10 / 2609.21219

Multi-viewpoint Geo-localization with Event Cameras

基于事件相机的多视角地理定位
Hines, Adam D., Milford, Michael, Fischer, Tobias
Abstract
Robot localization is an ongoing challenge that demands mapping and positioning systems that are tolerant to viewpoint change. Event cameras are attracting increasing interest and adoption in robotics; however, dealing with viewpoint variance is an under-investigated problem in existing event-based localizers. In addition, event-based datasets that emphasize viewpoint variance for challenging localization situations are scarce. Here, we introduce an event-based visual place recognition (VPR) system that performs robustly under viewpoint changes. We converted five large-scale geo-tagged datasets, conventionally used to train frame-based localization systems, into synthetic event streams using Image-to-Event (I2E) conversion, and used them to fine-tune a pre-trained event-based vision transformer backbone with a multi-loss function, yielding a system we call MegaEvent that learns viewpoint-robust features for place recognition. We achieved an average Recall@1 of 82% across three existing event-based localization datasets, leading the next best event-based method by 20 recall points, and frame-based VPR models applied directly to event frames by 8 to 26 recall points. We introduce a new, challenging dataset - Springfield-Event-VPR - which features a 3.7km walking route recorded in three camera orientations for a total of 11.1km, which MegaEvent outperforms the strongest baseline by 9 recall points. The code for MegaEvent is available at https://github.com/AdamDHines/megaevent.
Chinese Translation
机器人定位是一个持续的挑战,它要求建图与定位系统能够容忍视角变化。事件相机在机器人领域正吸引越来越多的关注与应用;然而,如何在现有基于事件的定位方法中应对视角变化仍是一个研究不足的问题。此外,针对具有挑战性的定位场景、强调视角变化的基于事件的数据集也十分稀缺。本文提出了一种基于事件的视觉位置识别(VPR)系统,能够在视角变化下保持稳健的性能。我们利用图像到事件(Image-to-Event, I2E)转换方法,将五个通常用于训练基于帧的定位系统的大规模地理标注数据集转换为合成事件流,并用它们通过多损失函数对一个预训练的基于事件的视觉Transformer骨干网络进行微调,得到一个我们称为MegaEvent的系统,该系统能够学习对视角鲁棒的位置识别特征。在三个现有的基于事件的定位数据集上,我们取得了平均82%的Recall@1,领先次优的基于事件的方法20个召回点,领先直接应用于事件帧的基于帧的VPR模型8至26个召回点。我们还引入了一个新的、具有挑战性的数据集——Springfield-Event-VPR,其中包含以三种相机朝向记录的3.7公里步行路线,总计11.1公里,MegaEvent在该数据集上比最强的基线方法高出9个召回点。MegaEvent的代码可在 https://github.com/AdamDHines/megaevent 获取。
cs.CV / 11 / 2609.21225

VGGT-CAD: Reconstructing Parametric CAD 3D Model with Geometric Grounding

VGGT-CAD:基于几何接地的参数化CAD三维模型重建
Yu, Chunan, Chen, Tianrun, Shen, Fu, Chen, Cheng, Zhu, Lanyun, Yang, Yang
Abstract
Parametric CAD reconstruction requires recovering both precise geometry and editable modeling operations from visual observations, making it challenging under limited and ambiguous views. Existing methods mainly rely on 2D appearance cues and lack strong multi-view geometric priors. In this work, we present VGGT-CAD, a geometry-aware framework for parametric CAD reconstruction from single- and multi-view observations. We transfer pretrained 3D geometric priors into CAD reconstruction by encoding camera parameters as condition tokens and jointly modeling them with image tokens. To handle varying numbers of viewpoints, we introduce a variable-view cross-view context aggregation module that adaptively fuses multi-view features. We further develop a training-free geometry-aware view selection strategy to select complementary and reliable frames during inference. The resulting representation is decoded into CAD command sequences using a non-autoregressive decoder. We also develop VideoCAD, a large-scale multi-view video benchmark derived from existing CAD data through multi-view re-rendering. Extensive experiments demonstrate the effectiveness of VGGT-CAD for visual CAD reconstruction under different observation configurations.
Chinese Translation
参数化CAD重建需要从视觉观测中同时恢复精确的几何结构和可编辑的建模操作,这在有限且模糊的视角下极具挑战性。现有方法主要依赖二维外观线索,缺乏强大的多视角几何先验。在本文中,我们提出了VGGT-CAD,一个面向单视角和多视角观测的几何感知参数化CAD重建框架。我们通过将相机参数编码为条件标记(condition tokens)并与图像标记进行联合建模,将预训练的三维几何先验迁移到CAD重建中。为应对视角数量的变化,我们引入了一个可变视角的跨视角上下文聚合模块,以自适应地融合多视角特征。我们进一步提出了一种无需训练的几何感知视角选择策略,用于在推理过程中选择互补且可靠的帧。所得表示通过非自回归解码器被解码为CAD命令序列。我们还构建了VideoCAD,一个通过对现有CAD数据进行多视角重新渲染得到的大规模多视角视频基准数据集。大量实验表明,VGGT-CAD在不同观测配置下的视觉CAD重建中均表现出有效性。
cs.CV / 12 / 2609.21241

Multiclass Semantic Segmentation of Wildland Fire Images Using Context-Aware Centralized Copy-Paste Data Augmentation

基于情境感知的集中式复制粘贴数据增强的野火图像多类语义分割
Kim, Joon Tai, Kunchala, Nishanth, Patel, Vishv, Chen, Tianle, Dong, Ziyu, Acero, Daniel Ospina, Williams, Roger, Kumar, Mrinal
Abstract
Producing accurate annotations for deep learning based image segmentation is both costly and labor intensive. This challenge is especially evident in wildland fire applications, where accurately labeled datasets are scarce due to the difficulty of collecting and annotating dynamic fire scenes. To address this problem, our previous work introduced the Centralized Copy-Paste Data Augmentation (CCPDA) method for semantic segmentation of wildland fire imagery, which generates artificial training samples by randomly pasting fire clusters from source images onto target images. However, random placement can produce contextually unrealistic scenes, such as fire burning on asphalt. In this paper, we present a context-aware strategy designed specifically to improve data quality and realism in small multiclass wildland fire datasets, ensuring that augmented samples remain contextually meaningful. The proposed method restricts fire placement to semantically valid target regions and selects the location whose Ash-Vegetation composition most closely matches the source context. This approach preserves existing fire regions in the target image, prevents unrealistic placements, and maintains contextual accuracy by generating images that resemble real wildland fire scenes. We evaluate the Context-Aware CCPDA strategy through numerical analysis and comparisons with other augmentation methods by a weighted sum-based multi-objective optimization (MOO) approach. The results confirm that the context-aware data augmentation strategy leads to improved segmentation performance and contextual realism, outperforming other augmentation procedures.
Chinese Translation
为基于深度学习的图像分割生成精确标注既昂贵又耗费人力。这一挑战在野火应用中尤为突出,因为动态火灾场景的采集与标注难度较大,导致准确标注的数据集十分稀缺。为解决这一问题,我们之前的工作提出了用于野火图像语义分割的集中式复制粘贴数据增强(Centralized Copy-Paste Data Augmentation, CCPDA)方法,该方法通过将源图像中的火簇随机粘贴到目标图像上来生成人工训练样本。然而,随机放置可能产生情境上不真实的场景,例如火在沥青上燃烧。在本文中,我们提出了一种专门针对小型多类野火数据集设计的情境感知策略,旨在提高数据质量与真实性,确保增强后的样本在情境上保持合理。该方法将火的放置限制在语义有效的目标区域内,并选择其灰烬-植被(Ash-Vegetation)组成与源情境最匹配的位置。该方法保留了目标图像中已有的火区域,避免了不真实的放置,并通过生成接近真实野火场景的图像来保持情境准确性。我们通过数值分析以及基于加权和的多目标优化(Multi-Objective Optimization, MOO)方法与其他增强方法的比较,对情境感知CCPDA策略进行了评估。结果证实,情境感知的数据增强策略能够提升分割性能和情境真实性,优于其他增强方法。
cs.CV / 13 / 2609.21242

SafeStyle: Calibrated Style Residual Injection for Controllable Style-Leakage Trade-off in Diffusion Stylization

SafeStyle:面向扩散风格化中可控风格-泄漏权衡的校准风格残差注入方法
Yang, Zhangping, Li, Min, Yan, Song, Gao, Rong, Bi, Xinliang, Xiong, Guanye, He, Yujie
Abstract
Reference-guided diffusion stylization aims to transfer visual style from a reference image while preserving the semantics specified by a text prompt. However, image conditioning often entangles transferable style cues with reference-specific content, leading to an inherent trade-off: stronger conditioning improves style fidelity but increases content leakage, whereas aggressive suppression reduces leakage at the cost of style expression. This challenge is further complicated by the distinct spatial organization of texture- and geometry-dominant styles. To address these issues, we propose SafeStyle, a training-free framework for calibrated style residual injection in frozen diffusion models. SafeStyle first estimates style-supported and content-associated subspaces from compact calibration sets, preserving their informative overlap while suppressing useless content variations. It then transports the purified style evidence over adaptive spatial granularity and constrains its effective influence through an explicit residual-norm budget. Experiments across texture- and geometry-dominant styles show that SafeStyle achieves a DINO style similarity of 0.432 while maintaining competitive text alignment. On a semantically disjoint leakage-stress benchmark, it further achieves a DINO style similarity of 0.474 with only 0.8\% semantic leakage, demonstrating an effective balance between style fidelity and reference-content suppression.
Chinese Translation
基于参考图像引导的扩散风格化旨在将参考图像的视觉风格迁移到目标图像,同时保留文本提示所指定的语义。然而,图像条件注入往往会将可迁移的风格线索与参考图像特有的内容纠缠在一起,从而产生一种固有的权衡:更强的条件注入能提升风格保真度,但会增加内容泄漏;而激进的抑制虽能减少泄漏,却以牺牲风格表达能力为代价。此外,纹理主导型风格与几何主导型风格在空间组织上的差异进一步加剧了这一挑战。为解决这些问题,我们提出了 SafeStyle,一个无需训练的框架,用于在冻结的扩散模型中进行校准的风格残差注入。SafeStyle 首先从紧凑的校准集中估计风格支持子空间和内容关联子空间,在抑制无用的内容变化的同时保留二者之间有信息量的重叠部分。随后,它在自适应空间粒度上传输纯化后的风格证据,并通过显式的残差范数预算来约束其有效影响。在纹理主导型和几何主导型风格上的实验表明,SafeStyle 在保持具有竞争力的文本对齐能力的同时,实现了 0.432 的 DINO 风格相似度。在一个语义不相关的泄漏压力测试基准上,它进一步实现了 0.474 的 DINO 风格相似度,而语义泄漏仅为 0.8%,展示了风格保真度与参考内容抑制之间的有效平衡。
cs.CV / 14 / 2609.21251

Geometry-Aware Diffusion Guidance via Curvature-Adaptive Tubular Correction

基于曲率自适应管状校正的几何感知扩散引导
Jiang, Enze, He, Jinwei, Ma, Zheng
Abstract
Gradient-guided diffusion samplers provide flexible priors for inverse problems and conditional generation, but strong guidance can move the sampling trajectory into regions where the learned score is poorly supported. Existing tangent-projection strategies limit first-order departure from an iso-density surface, yet discard potentially useful normal motion and overlook the second-order departure induced by tangent motion on a curved surface. We introduce curvature-adaptive tubular correction (CAT), a training-free plugin that regulates both effects within a shared, noise-dependent geometric budget. CAT decomposes the guidance gradient into normal and tangent components, charges normal displacement at first order and tangent displacement according to directional curvature, and obtains their jointly optimal magnitudes from a one-dimensional dual equation. Armijo backtracking calibrates the resulting finite step against the actual guidance objective, while matrix-free directional derivatives avoid constructing the full score Jacobian. We establish local guarantees for the tubular approximation, uniqueness of the correction, and sufficient objective decrease. Across seven inverse problems on FFHQ and ImageNet, CAT improves the evaluated pixel- and latent-space host samplers, with particularly consistent gains in perceptual metrics. It also improves black hole reconstruction on InverseBench and yields the lowest FID among the compared methods at every tested classifier-free guidance scale, while maintaining stable saturation and contrast. These results support curvature-aware tubular control as a reusable mechanism for stabilizing diffusion guidance.
Chinese Translation
梯度引导的扩散采样器为逆问题和条件生成提供了灵活的先验,但过强的引导会使采样轨迹进入学习到的分数(score)支撑薄弱的区域。现有的切向投影策略仅限制采样轨迹相对等密度面的一阶偏离,却丢弃了可能有益的法向运动,并忽略了切向运动在弯曲表面上引起的二阶偏离。我们提出曲率自适应管状校正(Curvature-Adaptive Tubular Correction, CAT),这是一个无需训练的插件,可在共享的、依赖噪声的几何预算内同时调控上述两种效应。CAT 将引导梯度分解为法向和切向分量,对法向位移按一阶计算代价、对切向位移按方向曲率计算代价,并通过一维对偶方程求得二者的联合最优幅值。Armijo 回溯法将所得的有限步长针对实际引导目标进行校准,而无矩阵的方向导数方法则避免了构造完整的分数雅可比矩阵。我们建立了关于管状近似的局部保证、校正的唯一性以及目标的充分下降性。在 FFHQ 和 ImageNet 上的七个逆问题中,CAT 提升了所评估的像素空间和潜空间宿主采样器的性能,在感知指标上取得尤为一致的增益。它还改进了 InverseBench 上的黑洞重建,并在所有测试的无分类器引导(classifier-free guidance)尺度下,在对比方法中取得了最低的 FID,同时保持了稳定的饱和度与对比度。这些结果表明,曲率感知的管状控制可作为稳定扩散引导的一种可复用机制。
cs.CV / 15 / 2609.21268

Edit-VAR: Taming Visual Autoregressive Model for Precise Video Editing

Edit-VAR:驯服视觉自回归模型以实现精确视频编辑
Zhao, Chongbo, Wang, Jiangming, Wang, Xilai, Wang, Xinyu, Tang, Jingyi, Hao, Chunjie, Song, Pengjie, Ma, Yue
Abstract
Text-guided video editing modifies target content while preserving the appearance and temporal coherence of unedited regions. Training-based approaches provide strong control but demand substantial data and computation. Training-free methods fall into inversion-free and inversion-based paradigms. Inversion-free approaches avoid trajectory recovery, but their source-preserving guidance can limit editing strength and leave semantic changes incomplete. Inversion-based approaches recover a latent trajectory before regeneration, where approximation errors can accumulate and cause source-content drift and temporal inconsistency. We introduce Edit-VAR, the first training-free and inversion-free framework for text-guided video editing with a pretrained visual autoregressive video model. Edit-VAR directly encodes the source video into multi-scale discrete tokens and performs probability-guided conditional token replacement for source preservation. Attention-guided token-wise and scale-aware modulation selectively relaxes source constraints over edit-relevant positions and generation stages. Scale-Decoupled Generation, implemented as late-scale constraint release, regenerates motion-consistent details and reduces texture fragmentation. Residual-guided token pruning further exploits redundancy at the final two high-resolution scales to reduce inference cost. Extensive experiments and a blind user study demonstrate that Edit-VAR outperforms existing training-free video editing methods overall in editing fidelity, source preservation, temporal coherence, and inference efficiency.
Chinese Translation
文本引导的视频编辑旨在修改目标内容,同时保持未编辑区域的外观和时间一致性。基于训练的方法提供了强大的控制能力,但需要大量的数据和计算资源。免训练方法分为免反演和基于反演两种范式。免反演方法避免了轨迹恢复,但其源保留引导可能限制编辑强度,导致语义变化不完整。基于反演的方法在重新生成之前恢复潜在轨迹,但近似误差会不断累积,造成源内容漂移和时间不一致性。我们提出了 Edit-VAR,这是首个基于预训练视觉自回归视频模型的免训练且免反演的文本引导视频编辑框架。Edit-VAR 直接将源视频编码为多尺度离散令牌(token),并通过概率引导的条件令牌替换实现源保留。注意力引导的逐令牌及尺度感知调制机制,选择性地在与编辑相关的位置和生成阶段放松源约束。尺度解耦生成(Scale-Decoupled Generation)通过后期尺度约束释放来实现,可重新生成运动一致的细节并减少纹理碎片化。残差引导的令牌剪枝进一步利用最后两个高分辨率尺度上的冗余以降低推理成本。大量实验和盲测用户研究表明,Edit-VAR 在编辑保真度、源保留、时间一致性和推理效率方面总体上优于现有的免训练视频编辑方法。
cs.CV / 16 / 2609.21276

Beyond Exact Match: Task-Aware GRPO for Cross-Domain PCBA Visual Question Answering

超越精确匹配:面向跨领域PCBA视觉问答的任务感知GRPO方法
Li, Jia, Dai, Li, Jia, Peng, Hu, Zhenzhen, Chan, Chee Seng, Bao, Bingkun, Hong, Richang
Abstract
In automated Printed Circuit Board Assembly (PCBA) inspection, standards-guided decisions require systems to jointly reason over fine-grained visual cues, component semantics, and manufacturing knowledge. Although large vision-language models (VLMs) provide a promising foundation, their deployment is hindered by the domain shift between standards-derived samples and real-world production-line imagery, together with heterogeneous output spaces spanning choice-based and numerical counting tasks. To address these challenges, we propose a multimodal reasoning framework for cross-domain PCBA visual question answering. The framework converts standards-derived, real-world, and auxiliary PCB-domain data into a unified instruction format and constructs verified reasoning traces aligned with visual evidence, question semantics, candidate options, and ground-truth answers. We further introduce Task-Aware Group Relative Policy Optimization (GRPO), which moves beyond exact-match supervision by integrating multi-component semantic rewards for choice-based questions, distance-aware rewards for counting questions, and an auxiliary format reward for valid outputs. During inference, answer-option semantic consistency correction, self-consistency voting, and multi-model arbitration are combined to improve prediction robustness. The proposed system achieves an Overall Score of 83.24 on the official PCBA Standard-to-Real Grand Challenge leaderboard, demonstrating the effectiveness of task-aware reward design and robust inference for cross-domain PCBA visual question answering.
Chinese Translation
在自动化印刷电路板组装(PCBA)检测中,基于标准的决策要求系统能够联合推理细粒度视觉线索、元器件语义以及制造知识。尽管大型视觉语言模型(VLM)为此提供了有前景的基础,但源自标准的样本与真实产线图像之间的领域偏移,以及涵盖选择题与数值计数任务的异构输出空间,阻碍了其部署应用。为应对这些挑战,我们提出了一种面向跨领域PCBA视觉问答的多模态推理框架。该框架将源自标准的数据、真实世界数据以及辅助PCB领域数据转换为统一的指令格式,并构建了与视觉证据、问题语义、候选选项及真实答案对齐的经过验证的推理轨迹。我们进一步提出了任务感知组相对策略优化(Task-Aware Group Relative Policy Optimization, GRPO),其超越了精确匹配监督的局限:针对选择题引入多组件语义奖励,针对计数问题引入距离感知奖励,并辅以面向有效输出的格式奖励。在推理阶段,通过结合答案选项语义一致性修正、自洽性投票与多模型仲裁,提升了预测的鲁棒性。所提出的系统在PCBA Standard-to-Real Grand Challenge官方排行榜上取得了83.24的总分,验证了任务感知奖励设计与鲁棒推理在跨领域PCBA视觉问答中的有效性。
cs.CV / 17 / 2609.21304

Combining Object Detection with Geometry-Aware Clustering to Distinguish Overlapping Plants in UAV Imagery

结合目标检测与几何感知聚类以区分无人机影像中相互重叠的植株
Lee, Ik Jae, Nguyen, Hieu D., Meenar, Mahbubur, Martinez, Carlos Morrison, Connelly, Cameron
Abstract
Reliable plant-level information from unmanned aerial vehicle (UAV) imagery is important for automated crop monitoring. However, in dense crop canopies, adjacent plants frequently overlap and are detected as a single object, reducing the reliability of plant-level measurements. This study presents a geometry-aware post-detection framework for resolving overlapping plant instances using standard RGB UAV imagery. The framework combines object detection with geometric clustering of plant components. Leaves or branches detected within each bush-level region are represented using two complementary geometric features: component centroids and radial intersection points (RIPs) derived from detected plant structures. K-means and Gaussian mixture models determine whether a detected region contains a single plant or two overlapping plants. Density filtering suppresses spurious radial intersections, and a post-pipeline ensemble combines spatial and directional geometric information. The framework was evaluated using UAV imagery of eggplant and tomato crops under field conditions. Centroid-based clustering achieved an F1-score of 0.89 for eggplant, while the combined centroid-RIP approach achieved the best tomato performance, with an accuracy of 0.80, precision of 1.00, and F1-score of 0.75 using K-means. Density filtering substantially improved RIP-based clustering for tomato. The proposed approach provides a lightweight, modular engineering solution that can be integrated with existing RGB UAV monitoring pipelines without additional depth sensors, pixel-level segmentation, three-dimensional reconstruction, or retraining of the primary bush detector. The results demonstrate that geometric reasoning applied to existing detector outputs can complement deep-learning-based object detection and improve plant-level interpretation in dense agricultural canopies.
Chinese Translation
从无人机(UAV)影像中获取可靠的植株级信息对自动化作物监测至关重要。然而,在茂密的作物冠层中,相邻植株经常相互重叠并被检测为单一目标,从而降低了植株级测量的可靠性。本研究提出了一种几何感知的检测后处理框架,利用标准RGB无人机影像来解决植株实例重叠问题。该框架将目标检测与植物部件的几何聚类相结合。在每个灌木级区域内检测到的叶片或枝条采用两种互补的几何特征进行表示:部件质心,以及从检测到的植物结构导出的径向交点(RIPs)。K-means和高斯混合模型用于判断检测区域包含单株植物还是两株重叠植物。密度过滤用于抑制虚假的径向交点,后处理流水线集成则融合了空间和方向性几何信息。该框架在田间条件下使用茄子和番茄作物的无人机影像进行了评估。基于质心的聚类在茄子上取得了0.89的F1分数,而质心-RIP组合方法在番茄上表现最佳,使用K-means时准确率为0.80、精确率为1.00、F1分数为0.75。密度过滤显著提升了基于RIP的番茄聚类效果。所提出的方法提供了一种轻量级、模块化的工程解决方案,可与现有RGB无人机监测流水线集成,无需额外的深度传感器、像素级分割、三维重建或对主要灌木检测器进行重新训练。结果表明,将几何推理应用于现有检测器输出可以补充基于深度学习的目标检测,并改善茂密农业冠层中的植株级解译效果。
cs.CV / 18 / 2609.21322

S3VD: Semantic-Guidance Spatio-Temporal Scanning for Video Deraining

S3VD:面向视频去雨的语义引导时空扫描方法
Jiang, Kui, Chen, Yiang, Luo, Yan, Yu, Zhaocheng, Jiang, Junjun, Liu, Xianming
Abstract
Heavy rainfall severely degrades outdoor videos by corrupting high-frequency details and introducing motion blur, critically undermining the reliability of visual tasks. Recently, State Space Models (SSMs), particularly Mamba, have emerged as efficient alternatives for vision tasks with their linear complexity and ability to model long-range dependencies. However, when confronted with the poor visual representations in rainy videos, Mamba still faces difficulties in preserving the integrity of 2D spatial semantics and modeling 3D spatio-temporal correlations. To break these limitations, we introduce S3VD, a Semantic-Guidance Spatio-Temporal Scanning framework for video deraining, featuring two key innovations: Multi-Scale Semantic Fusion (MSSF) Module and Spatio-Temporal Scanning Fusion (STSF) Module. The former integrates temporal semantic priors from DINOv2 to guide precise feature representation and counteract the loss of local semantic context inherent to Mamba's 1D flatten operation, enhancing robustness against extreme degradation. The latter introduces a spatio-temporal scanning mechanism and devises a Decoupled-Gating Mamba (DG-Mamba) layer, which employs two independent gates to adaptively control preceding and subsequent contextual information within the input clip, optimizing intra-frame and inter-frame correlation modeling. Experiments on video deraining benchmarks demonstrate the superiority of S3VD, achieving state-of-the-art performance with an average 0.84 dB PSNR improvement over Mamba-based baselines.
Chinese Translation
强降雨会破坏高频细节并引入运动模糊,严重降低户外视频的质量,极大地削弱视觉任务的可靠性。近年来,状态空间模型(State Space Models, SSMs),尤其是 Mamba,凭借其线性复杂度和建模长程依赖的能力,已成为视觉任务的高效替代方案。然而,面对雨天视频中的不良视觉表征,Mamba 在保持二维空间语义完整性和建模三维时空相关性方面仍然存在困难。为突破这些局限,我们提出了 S3VD——一个面向视频去雨的语义引导时空扫描框架,其包含两项关键创新:多尺度语义融合(Multi-Scale Semantic Fusion, MSSF)模块和时空扫描融合(Spatio-Temporal Scanning Fusion, STSF)模块。前者融合来自 DINOv2 的时序语义先验,以引导精确的特征表征,并抵消 Mamba 一维展平操作所固有的局部语义上下文丢失问题,增强对极端退化的鲁棒性。后者引入一种时空扫描机制,并设计了解耦门控 Mamba(Decoupled-Gating Mamba, DG-Mamba)层,该层采用两个独立的门来自适应地控制输入片段中前向与后向的上下文信息,优化帧内与帧间相关性建模。在视频去雨基准上的实验证明了 S3VD 的优越性,取得了最先进的性能,相比基于 Mamba 的基线方法,PSNR 平均提升 0.84 dB。
cs.CV / 19 / 2609.21323

VeriFuse: Bounded Vision-Language Arbitration and Reason-Guided Refinement for Cooperative 3D Perception

VeriFuse:面向协同三维感知的有界视觉-语言仲裁与推理引导的精化方法
Lin, Hongyi, Liu, Yiyao, Kang, Qi, Huang, Heye, Liu, Yang, Koutsopoulos, Haris, Zhao, Jinhua
Abstract
Vision-language models (VLMs) have demonstrated strong scene understanding and semantic judgment across diverse tasks, but their appropriate role in cooperative perception remains unclear. Directly asking a VLM to regress 3D detections is unreliable and computationally expensive, whereas using it to select the output of a single source discards useful information from other agents. We introduce VeriFuse, a bounded arbitration framework for vehicle-infrastructure cooperative 3D detection. Each agent first produces detections independently. Around each vehicle and roadside proposal, VeriFuse generates source-conditioned geometric candidates and combines the original detections, their perturbations, and cross-source hypotheses into a unified candidate pool. A frozen VLM then chooses among three admissible actions: SELECT an adequate candidate; REFINE an existing anchor when an object is supported but all candidates are geometrically inadequate; or REJECT an unsupported infrastructure-only proposal. Experiments on the DAIR-V2X dataset show that VeriFuse achieves 0.494/0.357 cooperative 3D AP50/AP70 and limits the relative vehicle-side BEV AP50 drop under a 300 ms delay to 1.7%. Overall, VeriFuse assigns the VLM a clear and constrained role in cooperative perception: semantic reasoning resolves ambiguity among cross-agent hypotheses, while deterministic constraints determine the final 3D geometry.
Chinese Translation
视觉-语言模型(VLM)在多种任务中展现出强大的场景理解与语义判断能力,但其在协同感知中的恰当角色仍不明确。直接让VLM回归三维检测结果既不可靠且计算开销高昂,而仅用其从单一来源选择输出则会丢弃来自其他智能体的有用信息。我们提出了VeriFuse,一个面向车路协同三维检测的有界仲裁框架。每个智能体首先独立生成检测结果。围绕每个车载和路侧的候选框,VeriFuse生成以来源为条件的几何候选,并将原始检测、其扰动以及跨来源假设合并为一个统一的候选池。随后,一个冻结的VLM在三种允许的动作中进行选择:SELECT(选择)一个合适的候选;当某目标得到支持但所有候选在几何上均不充分时,REFINE(精化)现有锚框;或REJECT(拒绝)一个缺乏支持、仅有路侧来源的候选框。在DAIR-V2X数据集上的实验表明,VeriFuse实现了0.494/0.357的协同三维AP50/AP70,并将300毫秒延迟下车辆端BEV AP50的相对下降限制在1.7%。总体而言,VeriFuse为VLM在协同感知中赋予了一个清晰且受限的角色:语义推理用于解决跨智能体假设之间的歧义,而确定性约束决定最终的三维几何结果。
cs.CV / 20 / 2609.21347

Cube-Splat: High-Fidelity 360{\deg} Gaussian Splatting SLAM via Cubemap Factorization and Adjoint-Consistent Optimization

Cube-Splat:基于立方体图分解与伴随一致性优化的高保真360度高斯泼溅SLAM
Guo, Xiangfei, Shi, Hao, Zhang, Yufan, Yi, Zhonghua, Mao, Yongqi, Yin, Xiaoting, Wang, Kaiwei
Abstract
Recent progress in 3D Gaussian Splatting (3DGS) has enabled dense visual SLAM with pinhole cameras, yet most pipelines are not designed for panoramic imagery. We present Cube-Splat, the first panoramic GS-SLAM framework that factorizes each 360{\deg} frame into a cubemap of four fixed-orientation virtual pinhole views sharing a single optical center. By designating the front face as the primary pose state, we accumulate gradients from all faces via an adjoint mapping, thereby enabling multi-face observations to coherently update a single state while strictly preserving cross-view geometric consistency. Concurrently, our mapping module densifies and optimizes anisotropic Gaussians using aggregated cubemap rays for high-fidelity, dense reconstruction. Furthermore, to rigorously evaluate panoramic SLAM under diverse and challenging conditions, we introduce SynPano, a highly scalable, photorealistic synthetic dataset featuring parameterized complex trajectories and multi-modal ground truth. Extensive evaluations on two public benchmarks (PALVIO and OmniBlender) and our SynPano dataset, collectively encompassing both indoor and outdoor scenes, demonstrate that Cube-Splat achieves state-of-the-art (SOTA) performance in tracking accuracy and reconstruction fidelity. Both the source code and the SynPano dataset are available at https://github.com/guoxf304/CubeSplat.
Chinese Translation
三维高斯泼溅(3D Gaussian Splatting, 3DGS)的最新进展已使基于针孔相机的稠密视觉SLAM成为可能,然而现有大多数流水线并未针对全景图像进行设计。我们提出Cube-Splat,这是首个全景GS-SLAM框架,它将每个360度帧分解为共享单一光心的四个固定朝向的虚拟针孔视图组成的立方体图(cubemap)。通过将正面面指定为主位姿状态,我们借助伴随映射从所有面累积梯度,从而使多面观测能够连贯地更新单一状态,同时严格保持跨视图的几何一致性。与此同时,我们的建图模块利用聚合的立方体图光线对各向异性高斯进行稠密化与优化,以实现高保真的稠密重建。此外,为了在多样且具有挑战性的条件下严格评估全景SLAM,我们推出了SynPano——一个高度可扩展、照片级逼真的合成数据集,具有参数化的复杂轨迹和多模态真值。在两个公开基准(PALVIO和OmniBlender)以及我们的SynPano数据集上的大量评估(共同涵盖室内与室外场景)表明,Cube-Splat在跟踪精度和重建保真度方面均达到了最先进(SOTA)的性能。源代码与SynPano数据集可在 https://github.com/guoxf304/CubeSplat 获取。
cs.CV / 21 / 2609.21351

PrismAlign: Prior-Steered Multi-View VLM Alignment for Hallucination-Robust Table OCR

PrismAlign:面向抗幻觉表格OCR的先验引导多视角VLM对齐方法
Liu, Guangyi, Huang, Qianjun, Hou, Boyu
Abstract
Table extraction suffers from frequent structural errors and semantic hallucinations. We propose PrismAlign, a multi-VLM framework aligning diverse visual perspectives to resolve ambiguity. It integrates priors of table logic to assess output plausibility, decoupling structural alignment from cell content alignment. A Bayesian decision strategy maximizes alignment accuracy by exploiting the correlation between extraction errors and computable rule violations. Evaluated on open-source and custom VLMs, PrismAlign reduces hallucinations and achieves state-of-the-art performance on OmniDocBench 1.5, as well as on the table category of CC-OCR and PureDocBench.
Chinese Translation
表格抽取常面临频繁的结构错误与语义幻觉问题。我们提出PrismAlign,一个多VLM框架,通过将多样的视觉视角进行对齐以消除歧义。该方法融合表格逻辑先验来评估输出的合理性,将结构对齐与单元格内容对齐解耦。通过利用抽取错误与可计算规则违反之间的相关性,一种贝叶斯决策策略最大化了对齐准确率。在开源及自研VLM上的评估表明,PrismAlign减少了幻觉,在OmniDocBench 1.5以及CC-OCR和PureDocBench的表格类别上取得了最先进的性能。
cs.CV / 22 / 2609.21354

Field Tracking of Insects Using a Stereoscopic Event-Based Camera Setup

基于双目事件相机装置的野外昆虫追踪
Shenwai, Pratham G., Lankheet, Martin J., Hrynuk, John T., Mahadeeswara, Mandiyam Y., Srinivasan, Mandyam V., Ravi, Sridhar
Abstract
High-speed tracking of small, fast-moving organisms in their natural environments is important to better understand their behavior and ecology. Traditional frame-based imaging suffers from motion blur due to low temporal resolution, and data storage limitations, propelling a search for more adaptive solutions. Event cameras, which capture changes in brightness at pixel level instead of entire frames, have emerged as a promising solution by increasing temporal resolution and data efficiency. Here, we demonstrate the use of event-based imaging with standard video-based processing methods by converting the asynchronous events into conventional video formats, allowing us to leverage the event camera's enhanced temporal detail to capture intricate insect flight movements and apply established image analysis techniques. Coupling this conversion process with a stereoscopic configuration provides continuous, low-latency, three-dimensional tracking of fast-moving subjects in field conditions. As a result, we substantially mitigate motion artifacts and achieve more accurate representations of animal movements. By making event-based imaging more readily applicable in natural field settings, our method support broader applications across animal behavior and ecological research, agricultural management, and other fields requiring high-fidelity object tracking in the wild.
Chinese Translation
对自然环境中小型、快速移动的生物体进行高速追踪,对于更好地理解其行为与生态具有重要意义。传统的基于帧的成像技术由于时间分辨率低而存在运动模糊问题,且受数据存储限制,这促使人们寻找更具适应性的解决方案。事件相机(Event Camera)以像素级别捕捉亮度变化而非整帧图像,通过提高时间分辨率和数据效率,成为一种有前景的解决方案。在本研究中,我们通过将异步事件转换为常规视频格式,展示了将事件成像与标准视频处理方法相结合的应用,从而能够利用事件相机增强的时间细节来捕捉昆虫复杂的飞行运动,并应用成熟的图像分析技术。将这一转换过程与双目(stereoscopic)配置相结合,可在野外条件下对快速移动的目标进行连续、低延迟的三维追踪。由此,我们显著减轻了运动伪影,并实现了对动物运动更精确的表征。通过使事件成像更易于应用于自然野外环境,我们的方法为动物行为与生态研究、农业管理以及其他需要在野外进行高保真目标追踪的领域提供了更广泛的应用支持。
cs.CV / 23 / 2609.21363

Hiding in Plain Sight: A Diffusion-based Mitigation of Geolocation Privacy Leakage in Vision-Language Models

藏于众目睽睽之下:一种基于扩散模型的视觉-语言模型地理位置隐私泄露缓解方法
Wang, Yining, Li, Xi, Zhang, Mi, Zhang, Xiaohan, You, Xiaoyu, Qian, Zhenxing, Wen, Mi
Abstract
Multimodal large reasoning models (MLRMs) have demonstrated remarkable capabilities in complex visual understanding. However, this very power introduces a critical yet underexplored privacy threat: adversaries can exploit MLRMs to precisely infer users' geographic locations from casually shared photographs, by performing structured reasoning over subtle visual cues such as architectural styles, vegetation, and lighting conditions. In this work, we present a systematic study of MLRM-driven geolocation privacy leakage. We first reveal that refusal-based safeguards are critically insufficient, as carefully crafted jailbreak prompts can raise model response rates to 100%. We further identify that existing defenses, which inject imperceptible perturbations into shared images, suffer from structural limitations intrinsic to their pixel-space optimization, resulting in degraded black-box transferability and pronounced visual artifacts. Motivated by these findings, we propose a diffusion-based framework that provides targeted, proactive defense against geolocation privacy leakage. By injecting perturbations into the latent space of a diffusion model during reverse sampling, our method operates directly on high-level semantic representations, thereby resolving the effectiveness-utility bottlenecks by construction. We further ground our optimization with GeoCLIP, a model explicitly aligned with GPS coordinates, as a surrogate to pinpoint and disrupt the geographic signals that MLRMs exploit for location inference. This targeted semantic disruption yields significantly stronger black-box transferability while preserving perceptual image quality, offering a seamless integration on social media platforms.
Chinese Translation
多模态大型推理模型(MLRM)在复杂视觉理解方面展现出了卓越的能力。然而,这种强大能力也带来了一种关键却未被充分研究的隐私威胁:攻击者可以利用MLRM,通过对建筑风格、植被和光照条件等细微视觉线索进行结构化推理,从用户随意分享的照片中精确推断其地理位置。在本工作中,我们对MLRM驱动的地理位置隐私泄露进行了系统性研究。我们首先揭示了基于拒绝回答的安全防护措施严重不足,因为精心设计的越狱提示词可以将模型的响应率提升至100%。我们进一步发现,现有防御方法通过向分享图像中注入不可感知的扰动来实施防护,但其像素空间优化方式存在固有结构性局限,导致黑盒迁移性下降且产生明显的视觉伪影。基于这些发现,我们提出了一种基于扩散模型的框架,能够针对地理位置隐私泄露提供定向的主动防御。通过在扩散模型的反向采样过程中向潜在空间注入扰动,我们的方法直接作用于高层语义表示,从而从构造上解决了有效性与实用性之间的瓶颈。我们进一步引入GeoCLIP——一个与GPS坐标显式对齐的模型——作为代理模型来指导优化,以精确定位并破坏MLRM用于位置推理的地理信号。这种定向的语义破坏在保持感知图像质量的同时,带来了显著更强的黑盒迁移性,并可实现与社交媒体平台的无缝集成。
cs.CV / 24 / 2609.21371

RobotEQ-Video: A Video-Centric Benchmark for Social Proactive Intelligence with World-State Taxonomy

RobotEQ-Video:基于世界状态分类体系的社交主动智能视频中心基准
Che, Xinyi, Lian, Zheng, Fang, Kuofei, Wang, Xuehao, Gao, Xinghai, Wu, Junqing, Wu, Chuyu, Liu, Liyi, Huang, Yanhan, Xie, Keyi, Ouyang, Haomin, Wu, Jinyang, Zhang, Fan, Zeng, Runhao, Yang, Xun, He, Bin
Abstract
Social Proactive Intelligence (SPI) extends proactive assistance beyond task completeness to consider social appropriateness in diverse embodied scenarios. However, prior SPI research faces two key limitations. First, existing work focuses on static images, whereas dynamic videos provide crucial cues for inferring human states and needs, offering richer information than isolated images. Second, prior work often relies on free-form data collection pipelines, which fail to guarantee comprehensive coverage of diverse scenarios. To address these gaps, we introduce RobotEQ-Video, shifting the focus from image-centric to video-centric analysis. To ensure comprehensive video coverage, we construct a hierarchical world-state taxonomy organized into a four-level coarse-to-fine structure, comprising 6 domains, 20 dimensions, 142 level-1 attributes, and 816 level-2 attributes. The resulting benchmark comprises 2K+ videos with 100K+ human annotations and 16K+ labels for assessing behavior properness. Benchmark evaluation reveals that current systems remain unreliable and fall short of human performance. We further explore how world models can help tackle this task. This work advances SPI research from static images to dynamic videos and ensures more comprehensive scenario coverage during benchmarking.
Chinese Translation
社交主动智能(Social Proactive Intelligence, SPI)将主动辅助的范围从任务完备性拓展到在各种具身场景中考虑社交适宜性。然而,此前的SPI研究存在两个关键局限。第一,现有工作聚焦于静态图像,而动态视频能为推断人类状态和需求提供关键线索,比孤立图像提供更丰富的信息。第二,先前工作通常依赖自由形式的数据采集流程,无法保证对多样化场景的全面覆盖。为填补这些空白,我们提出RobotEQ-Video,将研究重心从以图像为中心转向以视频为中心的分析。为确保视频覆盖的全面性,我们构建了一个分层的世界状态分类体系,采用由粗到细的四层结构,包含6个领域、20个维度、142个一级属性和816个二级属性。由此构建的基准包含2千多个视频、10万多条人工标注以及1.6万多个用于评估行为恰当性的标签。基准评测表明,当前系统仍然不可靠,与人类水平存在差距。我们进一步探索了世界模型如何帮助解决这一任务。这项工作将SPI研究从静态图像推进到动态视频,并确保了基准评测中更全面的场景覆盖。
cs.CV / 25 / 2609.21379

JEPA Guided Diffusion: Predictive Vision-Language Conditioning for Generative Traffic Forecasting

JEPA引导的扩散模型:面向生成式交通预测的预测性视觉-语言条件化
Nguyen, Trinh Tra Giang, Vo, Thanh Nguyen, Bui, Nguyen Hoai Thuong, Bui, Ha Duc
Abstract
Accurate traffic forecasting requires both understanding scene dynamics and synthesizing realistic future observations. Recent diffusion-based video generation models produce visually plausible predictions but require expensive end-to-end training and often entangle scene understanding with image synthesis. In this work, we propose a decoupled forecasting framework that separates future representation learning from video generation. A frozen V-JEPA encoder first extracts predictive latent representations from the observed traffic videos, capturing the underlying scene dynamics in a semantic latent space. A lightweight latent alignment module then projects these representations into the conditioning space of a frozen Cosmos diffusion module, enabling future video synthesis without retraining the large generative model. By freezing all foundation models and training only the lightweight alignment module, the proposed framework substantially reduces optimization complexity while preserving forecasting capability. Experimental results on the AI City Challenge 2026 Track 5 benchmark demonstrate that the proposed method achieved a score of 75.1297, ranking third in the competition. These results suggest that predictive world representations learned by V-JEPA can effectively guide downstream video generation, providing a practical and efficient alternative to end-to-end diffusion-based forecasting.
Chinese Translation
准确的交通预测既需要理解场景动态,又需要合成逼真的未来观测结果。近期基于扩散模型的视频生成方法能够产生视觉上合理的预测,但其代价高昂的端到端训练往往将场景理解与图像生成耦合在一起。本文提出了一种解耦的预测框架,将未来表征学习与视频生成分离开来。首先,冻结的V-JEPA编码器从观测到的交通视频中提取预测性潜在表征,在语义潜在空间中捕捉底层场景动态。随后,一个轻量级的潜在对齐模块将这些表征投影到冻结的Cosmos扩散模块的条件空间中,从而无需重新训练大型生成模型即可实现未来视频合成。通过冻结所有基础模型、仅训练轻量级对齐模块,所提出的框架在保持预测能力的同时大幅降低了优化复杂度。在AI City Challenge 2026 Track 5基准上的实验结果表明,所提出的方法取得了75.1297的分数,在竞赛中排名第三。这些结果表明,V-JEPA学习到的预测性世界表征能够有效指导下游视频生成,为基于端到端扩散模型的预测提供了一种实用且高效的替代方案。
cs.CV / 26 / 2609.21386

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

AgentVidBench:一个用于评估多模态大语言模型智能体的多跳视频问答基准
An, Seoyeon, Jang, Hyeonseo, Kim, Minsu, Lee, Chanho, Park, Younghan, Lee, Kangwook
Abstract
Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in video understanding, existing benchmarks remain confined to simple scene-level queries or global summaries that require only single-step inference. Real-world video understanding involves more challenging tasks that require multi-hop multimodal reasoning, and there is a critical absence of video benchmarks equipped to rigorously evaluate these agentic capabilities. To bridge this gap, we introduce AgentVidBench, a multi-hop video question answering benchmark focused on evaluating the spatial, temporal, and causal reasoning capabilities of MLLM agents. Beyond standard question-answer pairs, AgentVidBench provides step-by-step solution traces to support trajectory evaluation that assesses whether agents explicitly acquire the evidence needed to justify their answers. Experiments with 12 proprietary and open-source MLLMs show that single-turn performance remains limited on AgentVidBench, while integrating these models into state-of-the-art agentic workflows generally improves performance with respect to both accuracy and trajectory scores. We further present a simple yet effective agentic strategy that serves as a competitive baseline on AgentVidBench, establishing our benchmark as a holistic testbed for future research on agentic video understanding. Code and datasets are available at https://github.com/krafton-ai/agentvidbench and https://huggingface.co/datasets/agentvidbench/agentvidbench.
Chinese Translation
全面的视频理解对于推动人工智能迈向物理世界的复杂动态至关重要。尽管多模态大语言模型(MLLM)的最新进展在视频理解方面展现了卓越能力,但现有基准仍局限于仅需单步推理的简单场景级查询或全局摘要。真实世界的视频理解涉及更具挑战性的任务,需要多跳多模态推理,然而目前严重缺乏能够严格评估这些智能体(agentic)能力的视频基准。为填补这一空白,我们提出了 AgentVidBench,一个专注于评估 MLLM 智能体在空间、时间和因果推理能力的多跳视频问答基准。除了标准的问答对外,AgentVidBench 还提供分步解答轨迹,以支持轨迹评估,从而判断智能体是否明确获取了支撑其答案所需的证据。对 12 个专有和开源 MLLM 的实验表明,单轮性能在 AgentVidBench 上仍然有限,而将这些模型集成到最先进的智能体工作流中,通常能在准确率和轨迹得分两方面提升性能。我们进一步提出了一种简单而有效的智能体策略,作为 AgentVidBench 上具有竞争力的基线,确立了本基准作为未来智能体视频理解研究的一个全面测试平台。代码和数据集可在 https://github.com/krafton-ai/agentvidbench 和 https://huggingface.co/datasets/agentvidbench/agentvidbench 获取。
cs.CV / 27 / 2609.21400

A Scene Language Model for Open-Vocabulary Scene Mapping

一种用于开放词汇场景建图的场景语言模型
Lilja, Adam, Hübel, Fabio, He, Siming, Fu, Junsheng, Tomlin, Claire, Hammarstrand, Lars, Malik, Jitendra, Frey, Jonas, Pavone, Marco
Abstract
Open-vocabulary 3D scene mapping aims to build a persistent representation of the objects in an environment. Existing systems typically rely on engineered mapping pipelines to associate observations, merge information across views, and maintain a consistent scene representation over time. Many additionally store feature-rich object representations, such as embeddings or image crops, increasing the size and complexity of the persistent memory. We introduce SceneLM, a Scene-Language Model that directly maintains a textual scene map. The full scene is represented as a structured text list of objects, which serves as the model's only persistent memory. For each input image, the model reads the current scene state and updates the map by adding, editing, and removing objects. To learn this behavior, we introduce supervision tasks for iterative scene map maintenance together with an automatic annotation pipeline that generates training data from images without human labels. We evaluate SceneLM on both a language-grounded retrieval benchmark and a localization benchmark. Across both benchmarks, the model produces a scene map that achieves competitive performance with complete mapping systems built from dedicated perception and geometric modules while producing a scene representation that is 6-12x more compact. We further show that SceneLM can be run online on an edge device through experiments on a quadruped. These results show that a persistent open-vocabulary 3D scene map can be maintained directly by a single vision-language model using only a lightweight text representation. Training and inference code is available on https://goldengait.github.io/scenelm/.
Chinese Translation
开放词汇三维场景建图旨在为环境中的物体构建一种持久化表示。现有系统通常依赖人工设计的建图流程来关联观测、跨视图融合信息,并随时间维护一致的场景表示。许多系统还存储特征丰富的物体表示,如嵌入向量或图像裁剪,从而增加了持久化存储的大小和复杂度。我们提出 SceneLM,一种直接维护文本形式场景地图的场景语言模型(Scene-Language Model)。整个场景被表示为一个结构化的物体文本列表,作为模型唯一的持久化记忆。对于每张输入图像,模型读取当前场景状态,并通过添加、编辑和删除物体来更新地图。为了学习这种行为,我们引入了用于迭代式场景地图维护的监督任务,以及一个无需人工标注即可从图像生成训练数据的自动标注流水线。我们在一个基于语言的检索基准和一个定位基准上评估了 SceneLM。在两个基准上,该模型生成的场景地图与由专用感知和几何模块构建的完整建图系统相比取得了具有竞争力的性能,同时生成的场景表示紧凑度提高了 6-12 倍。我们进一步通过四足机器人上的实验证明,SceneLM 可以在边缘设备上在线运行。这些结果表明,持久化的开放词汇三维场景地图可以由单个视觉-语言模型仅通过轻量级文本表示直接维护。训练和推理代码可在 https://goldengait.github.io/scenelm/ 获取。
cs.CV / 28 / 2609.21402

SIRA: Reasoning-Aware Surgical Instrument Segmentation via Query-Anchored Alignment

SIRA:基于查询锚定对齐的推理感知手术器械分割
Zhang, Zhibo, Wang, Qijie, Yan, Zengqiang
Abstract
Surgical instrument segmentation (SIS) plays a critical role in robotic assistance and surgical workflow analysis. However, most existing SIS methods formulate segmentation as a category-driven localization problem, limiting their ability to capture procedural context and task-dependent semantics in surgical workflows. We introduce Reasoning-Aware Surgical Instrument Segmentation (RA-SIS), a task formulation that frames segmentation as query-conditioned inference under surgical context. To benchmark this setting, we construct SurgRS, a surgical reasoning segmentation dataset consisting of 41,000 image-text pairs, which aligns instance-level masks with structured query-answer supervision to enable semantic grounding at the pixel level. Based on SurgRS, we propose Surgical Instrument Reasoning and Segmentation Assistant (SIRA), a multimodal framework that disentangles target-level and query-level semantics and integrates them with visual features through query-anchored dual alignment. By aligning query semantics with spatial features and segmentation prompts, SIRA enhances semantic-visual consistency in mask prediction. Extensive experiments on SurgRS demonstrate improvements over existing reasoning-aware baselines. Code is available at https://github.com/linxir226/SIRA.
Chinese Translation
手术器械分割(Surgical Instrument Segmentation, SIS)在机器人辅助和手术工作流分析中起着至关重要的作用。然而,现有的大多数SIS方法将分割公式化为一个由类别驱动的定位问题,限制了其捕捉手术工作流中的过程上下文和任务相关语义的能力。我们提出了推理感知手术器械分割(Reasoning-Aware Surgical Instrument Segmentation, RA-SIS),这是一种将分割框架化为在手术上下文下基于查询的条件推理的任务形式。为对这一设定进行基准测试,我们构建了SurgRS——一个包含41,000对图像-文本的手术推理分割数据集,该数据集将实例级掩码与结构化的问答监督对齐,以实现像素级的语义锚定。基于SurgRS,我们提出了手术器械推理与分割助手(Surgical Instrument Reasoning and Segmentation Assistant, SIRA),这是一个多模态框架,能够解耦目标级与查询级语义,并通过查询锚定双重对齐将其与视觉特征整合。通过将查询语义与空间特征及分割提示对齐,SIRA增强了掩码预测中的语义-视觉一致性。在SurgRS上的大量实验表明,SIRA优于现有的推理感知基线方法。代码已发布于 https://github.com/linxir226/SIRA。
cs.CV / 29 / 2609.21407

Quantization-Aware Kalman Estimation for Diffusion Sampling

面向扩散采样的量化感知卡尔曼估计
Shi, Qitan, Jin, Cheng, Zhang, Jiawei, Gu, Yuantao
Abstract
Quantization offers a practical path to deploying diffusion models with reduced memory and computation, but aggressive compression can cause quantized outputs to deviate substantially from their full-precision counterparts. Sampling-stage correction methods seek to compensate for such deviations during sampling, but existing approaches rely primarily on local information and underexploit trajectory history, limiting their ability to correct errors that propagate across timesteps. In this work, we formulate sampling with a quantized denoiser as an online estimation problem, using the history of quantized denoiser outputs to recover the underlying full-precision outputs required by the sampler. We propose QuAKE, a Quantization-Aware Kalman Estimator that combines a smooth trajectory prior with a conditional Gaussian observation model. At each sampling step, QuAKE recursively updates the posterior over the output window in closed form and feeds its posterior mean to the sampler. QuAKE is a lightweight plug-and-play corrector that requires no modification to the quantized network and naturally supports arbitrary high-order multistep ODE samplers. Experiments across W4A4-quantized text-to-image diffusion models show that QuAKE consistently outperforms existing methods in reducing the distributional discrepancy from full-precision sampling.
Chinese Translation
量化为部署扩散模型提供了一条降低内存和计算开销的实用途径,但过于激进的压缩会导致量化输出与全精度输出产生显著偏差。采样阶段校正方法试图在采样过程中补偿此类偏差,但现有方法主要依赖局部信息,未能充分利用轨迹历史,限制了其校正跨时间步传播误差的能力。在本工作中,我们将使用量化去噪器的采样过程形式化为一个在线估计问题,利用量化去噪器输出的历史信息来恢复采样器所需的底层全精度输出。我们提出了QuAKE(量化感知卡尔曼估计器,Quantization-Aware Kalman Estimator),它将平滑轨迹先验与条件高斯观测模型相结合。在每个采样步骤中,QuAKE以闭式解的形式递归更新输出窗口上的后验分布,并将其后验均值提供给采样器。QuAKE是一种轻量级的即插即用校正器,无需修改量化网络,并且天然支持任意高阶多步ODE采样器。在W4A4量化的文生图扩散模型上的实验表明,QuAKE在减小与全精度采样之间的分布差异方面始终优于现有方法。
cs.CV / 30 / 2609.21412

When Online Adaptation Hurts: Parameter-Frozen Test-Time Ensembling for Continual Medical Image Segmentation

当在线自适应有害时:用于持续医学图像分割的参数冻结测试时集成方法
Huang, Ruijie
Abstract
Medical image segmenters often get worse when sites, scanner vendors, or protocols change. Continual test-time adaptation (CTTA) addresses this problem without target labels, but it can be impossible to update a model on a non-stationary stream and can lead to a lot of errors. We examine a more reasonable and meaningful alternative: parameter-frozen inference enhancement(PIE). We use a source-trained segmenter that learns about anatomy-preserving scale and flip views, maps their predictions back to the native location, and averages the probabilities. We do not modify the weights of the model or the normalization statistics. On a cardiac MRI stream from M\&Ms, which is trained on vendor A and evaluated sequentially on vendors B, C, and D, PIE has 0.7786 mean Dice, compared to 0.7680 for source-only inference and 0.7388--0.7416 for five other online-adaptation baselines. The controlled ablations show that performance saturates at 28 views, and confidence weighting, class-prior correction, connected-component filtering, morphological refinement, and inter-slice smoothing have no effect or cause negative transfer. Qualitative results on cardiac MRI and fundus images are also consistent with the frozen ensemble keeping thinner and nested anatomical structures. These results provide a strong, stable baseline for medical CTTA and expose an important failure mode: adaptation and handcrafted refinement can be less reliable than carefully designed inference.
Chinese Translation
当扫描站点、扫描仪厂商或采集协议发生变化时,医学图像分割模型的性能往往会下降。持续测试时自适应(CTTA)方法能够在没有目标域标签的情况下解决这一问题,但在非平稳数据流上更新模型可能无法实现,并可能导致大量错误。我们研究了一种更合理、更有意义的替代方案:参数冻结的推理增强(PIE)。我们使用一个在源域训练的分割模型,学习保持解剖结构的尺度和翻转视图,将它们的预测映射回原始位置,并对概率取平均。我们不修改模型权重或归一化统计量。在M&Ms心脏MRI数据流上(在厂商A的数据上训练,并依次在厂商B、C和D的数据上评估),PIE的平均Dice系数为0.7786,而仅用源域推理为0.7680,五个其他在线自适应基线方法为0.7388至0.7416。受控消融实验表明,性能在28个视图时达到饱和,而置信度加权、类先验校正、连通域过滤、形态学细化和层间平滑均无效果或引起负迁移。在心脏MRI和眼底图像上的定性结果也与冻结集成能够更好地保留更细小和嵌套的解剖结构相一致。这些结果为医学CTTA提供了一个强大而稳定的基线,并揭示了一个重要的失败模式:自适应和手工设计的细化可能不如精心设计的推理可靠。
cs.CV / 31 / 2609.21424

P$^3$-SAM: SAM with Perceptual Parallel Prompt for Few-Shot Strip Steel Surface Defect Segmentation

P$^3$-SAM:基于感知并行提示的SAM用于小样本带钢表面缺陷分割
Xu, Qian, Xiong, Hang, Wang, Anpeng, Kwong, Sam, Zhang, Cong, Cong, Runmin
Abstract
Few-shot semantic segmentation (FSS) of strip steel surface defects (S$^3$D) has posed significant challenges distinct from natural scenes. Unlike natural images, S$^3$D task exhibits unique characteristics including low local contrast, uneven illumination, and complex fine-grained texture patterns. Although recent methods based on Segment Anything Model (SAM) have shown promise in FSS on natural images by leveraging SAM's powerful pre-trained representations, these unique industrial characteristics of S$^3$D images lead to performance drop when directly applying SAM to industrial defect scenarios. In this paper, we propose a novel Perceptual Parallel Prompt (P$^3$) framework that empowers SAM, creating the P$^3$-SAM model to address these challenges through two core strategies. First, we develop a Perceptual-Optimized Encoding (POE) strategy that enhances local contrast and preserves critical texture details for S$^3$D segmentation. Second, we introduce the Parallel Prompt Generator (PPG) strategy that simultaneously generates both semantic and spatial prompts, enabling comprehensive guidance for SAM's decoder across varying images. Extensive experiments on three few-shot S$^3$D benchmarks demonstrate that P$^3$-SAM achieves state-of-the-art performance, with particularly notable improvements of 12.00% in mIoU on Surface Defects-4i dataset.
Chinese Translation
带钢表面缺陷(S$^3$D)的小样本语义分割(FSS)面临着与自然场景截然不同的重大挑战。与自然图像不同,S$^3$D任务具有独特的特性,包括局部对比度低、光照不均匀以及复杂的细粒度纹理模式。尽管近期基于Segment Anything Model(SAM)的方法通过利用SAM强大的预训练表示,在自然图像的小样本语义分割中展现出良好前景,但S$^3$D图像的这些独特工业特性导致将SAM直接应用于工业缺陷场景时性能下降。本文提出了一种新颖的感知并行提示(Perceptual Parallel Prompt,P$^3$)框架来增强SAM的能力,构建了P$^3$-SAM模型,通过两个核心策略来应对这些挑战。首先,我们开发了感知优化编码(Perceptual-Optimized Encoding,POE)策略,以增强局部对比度并保留S$^3$D分割的关键纹理细节。其次,我们引入了并行提示生成器(Parallel Prompt Generator,PPG)策略,可同时生成语义提示和空间提示,从而为SAM的解码器在不同图像上提供全面的指导。在三个小样本S$^3$D基准数据集上的大量实验表明,P$^3$-SAM取得了最先进的性能,尤其是在Surface Defects-4i数据集上mIoU提升了12.00%。
cs.CV / 32 / 2609.21437

Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction

局部思考,全局精化:面向内存高效三维重建的方法
Zhou, Jingke, Ma, Chenhang, Zhong, Zhizhou, Liu, Mingkai, Zhou, Zhuang, ji, Yicheng, Su, Binghua, Cai, Bo, Huang, Xianliang
Abstract
We propose LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency. Instead of relying on full global attention, our method introduces cross-window attention at a small subset of transformer blocks, enabling effective information propagation across adjacent temporal windows while keeping memory usage bounded. To mitigate long-term pose drift, we further design a global camera consistency refinement module, where camera tokens interact with compact register tokens via cross-attention to enforce scene-level constraints across the entire sequence. This design enables joint optimization of camera representations and significantly improves long-horizon pose stability without incurring the high cost of sequence-wide attention. Extensive experiments demonstrate that LoG-VGGT achieves improved depth accuracy and robust camera pose estimation across multiple long-sequence benchmarks, while delivering competitive streaming reconstruction performance.
Chinese Translation
我们提出了LoG-VGGT,一个面向长序列三维重建的内存高效框架,它将局部时序建模与全局相机一致性相平衡。我们的方法不依赖完整的全局注意力,而是在一小部分Transformer块中引入跨窗口注意力(cross-window attention),从而在保持内存使用有界的同时,实现相邻时间窗口之间的有效信息传播。为了缓解长期位姿漂移,我们进一步设计了一个全局相机一致性精化模块,其中相机标记(camera tokens)通过交叉注意力与紧凑的寄存器标记(register tokens)进行交互,以在整个序列上施加场景级约束。这一设计实现了相机表示的联合优化,并显著提升了长时序下的位姿稳定性,同时无需承担序列级全局注意力的高昂开销。大量实验表明,LoG-VGGT在多个长序列基准上实现了更高的深度精度和更鲁棒的相机位姿估计,同时提供了具有竞争力的流式重建性能。
cs.CV / 33 / 2609.21449

ME-Dex 1.0: Bringing Heterogeneous Tactile Sensing into World Action Modeling

ME-Dex 1.0:将异构触觉感知引入世界动作建模
Zhang, Xuancheng, Liu, Xuetao, Tang, Qianying, Wang, Jizhe, Cheng, Zhijing, Lin, Bochen, Wen, Haoran, Li, Ming, Zhan, Kun, Liu, Yu
Abstract
World Action Models bring the predictive capabilities of video models into robot action generation, providing a rich foundation for modeling future visual states. Tactile sensing complements this foundation with direct measurements of physical interaction. Some existing methods use tactile features as conditioning inputs without jointly predicting future tactile states, visual observations, and actions. Our key insight is that tactile signals, like video, provide observations of the evolving world state and should be modeled as future observations alongside video. We present ME-Dex-1.0 (MachEmbodied-Dex-1.0), a unified World Action Tactile Model for joint visual, tactile, and action learning. ME-Dex-1.0 adopts a Mixture-of-Transformers architecture comprising a Video Expert, a Tactile Expert, and an Action Expert, all trained with flow matching. We use shared attention connects the experts in intermediate layers, allowing action generation to draw on learned representations of visual and tactile dynamics during joint denoising. To support multi-source heterogeneous tactile inputs, a Canonical Hand Model and a Unified Tactile Autoencoder map tactile observations from different embodiments and sensing layouts into shared spatial and latent spaces. To address the limited availability of paired visual, tactile, and action data, we develop the Agentic Tactile Data Engine, an agent-based data production platform. It supplements RoboTwin and DexJoCo with tactile data recorded directly from force sensors during trajectory replay in simulation. Experiments on the RoboTwin, DexJoCo, and ManiFeel simulation platforms, together with real robot evaluations, demonstrate improved manipulation performance using both grippers and dexterous hands equipped with tactile sensing.
Chinese Translation
世界动作模型(World Action Models)将视频模型的预测能力引入机器人动作生成,为建模未来视觉状态提供了丰富的基�础。触觉感知则通过对物理交互的直接测量对这一基础形成补充。现有的一些方法仅将触觉特征作为条件输入,而未对未来触觉状态、视觉观测和动作进行联合预测。我们的核心洞见是:触觉信号与视频一样,是对不断演化的世界状态的观测,应当与视频一起被建模为未来观测。我们提出了 ME-Dex-1.0(MachEmbodied-Dex-1.0),一个用于视觉、触觉和动作联合学习的统一世界动作触觉模型。ME-Dex-1.0 采用混合Transformer(Mixture-of-Transformers)架构,由视频专家(Video Expert)、触觉专家(Tactile Expert)和动作专家(Action Expert)组成,均通过流匹配(flow matching)进行训练。我们使用共享注意力在中间层连接各专家,使动作生成能够在联合去噪过程中借鉴已学习的视觉和触觉动力学表征。为支持多源异构触觉输入,我们提出了标准手部模型(Canonical Hand Model)和统一触觉自编码器(Unified Tactile Autoencoder),将来自不同本体构型和感知布局的触觉观测映射到共享的空间和潜在空间中。针对视觉、触觉和动作配对数据有限的问题,我们开发了智能体触觉数据引擎(Agentic Tactile Data Engine),一个基于智能体的数据生产平台。该平台通过在仿真中轨迹回放时直接从力传感器记录触觉数据,对 RoboTwin 和 DexJoCo 数据集进行补充。在 RoboTwin、DexJoCo 和 ManiFeel 仿真平台上的实验以及真实机器人评估表明,对于配备触觉感知的夹爪和灵巧手,本方法均提升了操作性能。
cs.CV / 34 / 2609.21455

CompAdapt: Adaptable Composite Motion Modeling for Physics-Consistent Text-to-Video Generation

CompAdapt:面向物理一致性文本到视频生成的可适配复合运动建模
Qin, Haoran, Wu, Renlong, Huang, Tianyu, Ding, Yukang, Li, Hui, Zuo, Wangmeng
Abstract
While diffusion-based text-to-video (T2V) models have demonstrated impressive capability in generating realistic and temporally coherent videos, they often fail to respect fundamental physical dynamics. Although recent physics-constrained methods incorporate explicit dynamics priors to improve physical plausibility, they remain limited to simple single-type motions, depend on manually specified parameters, and struggle to generalize to unseen physical laws. In this work, we propose CompAdapt, a physics-consistent T2V framework for adaptable generation across complex real-world scenarios. It extends neural dynamics modeling beyond single-type motions to encompass composite physical behaviors, including coupled motions, multi-stage transitions, and multi-object collisions. Furthermore, CompAdapt translates natural language prompts into structured physical semantics, enabling end-to-end specification of motion types, temporal relations, and initial physical parameters. To generalize to novel physical environments, CompAdapt introduces dynamics-aware prior matching, achieving one-shot adaptation without retraining the core dynamics module. In addition, a physics-aware latent feature fusion module improves visual fidelity under fast and complex motion. Experiments on physics-focused T2V benchmarks demonstrate that CompAdapt improves physical consistency over both general T2V models and physics-constrained baselines, while preserving high visual quality and adaptability to unseen dynamics. The project page is available at https://makapic.github.io/CompAdapt/ .
Chinese Translation
尽管基于扩散模型的文本到视频(T2V)模型在生成逼真且时间上连贯的视频方面展现出令人瞩目的能力,但它们往往无法遵循基本的物理动力学规律。虽然近期的物理约束方法通过引入显式动力学先验来提升物理合理性,但它们仍局限于简单的单一类型运动,依赖人工指定的参数,且难以泛化到未见过的物理规律。在本工作中,我们提出 CompAdapt,一个面向复杂真实场景可适配生成的物理一致性 T2V 框架。它将神经动力学建模从单一类型运动扩展到复合物理行为,包括耦合运动、多阶段过渡和多物体碰撞。此外,CompAdapt 将自然语言提示转化为结构化的物理语义,实现对运动类型、时间关系和初始物理参数的端到端指定。为了泛化到新颖的物理环境,CompAdapt 引入了动力学感知的先验匹配,无需重新训练核心动力学模块即可实现单样本适配。此外,一个物理感知的隐特征融合模块提升了快速复杂运动下的视觉保真度。在面向物理的 T2V 基准上的实验表明,CompAdapt 相较于通用 T2V 模型和物理约束基线方法均提升了物理一致性,同时保持了较高的视觉质量以及对未见动力学规律的可适应性。项目页面见 https://makapic.github.io/CompAdapt/ 。
cs.CV / 35 / 2609.21462

PSEE: Progressive Sensor Event Expansion for Point-Supervised Temporal Action Localization

PSEE:面向点监督时序动作定位的渐进式传感器事件扩展方法
Yin, Jiaxi, Wang, Ge, Ding, Han, Wang, Fei
Abstract
Temporal action localization (TAL) in wearable sensor streams identifies action classes and temporal boundaries, enabling finer-grained activity understanding than conventional action recognition. However, training typically requires costly start--end annotations for every action instance. To reduce this burden, we study point-supervised TAL, where each instance is labeled with only one timestamp and its class. We propose Progressive Sensor Event Expansion (PSEE), which combines semantic activations, sensor-specific transition evidence, and adaptive temporal ownership to recover point-supervised pseudo segments. These segments supervise standard TAL detectors without modifying their inference procedures. Cross-subject experiments on four inertial-sensing benchmarks demonstrate improved pseudo-boundary quality over adapted point-supervised baselines, compatibility with different TAL detectors, and robustness to point sampling. Code is available at https://github.com/joeeeeyin/PSEE.
Chinese Translation
可穿戴传感器数据流中的时序动作定位(Temporal Action Localization, TAL)旨在识别动作类别及其时序边界,从而实现比传统动作识别更细粒度的活动理解。然而,模型训练通常需要为每个动作实例提供代价高昂的起止时间标注。为减轻这一负担,我们研究了点监督TAL任务,即每个实例仅标注一个时间戳及其类别。我们提出渐进式传感器事件扩展方法(Progressive Sensor Event Expansion, PSEE),该方法结合语义激活、传感器特定的转换证据以及自适应时序归属,从点监督标签中恢复伪片段。这些伪片段可用于监督标准的TAL检测器,且无需修改其推理过程。在四个惯性传感基准数据集上的跨被试实验表明,与适配后的点监督基线方法相比,本方法提升了伪边界质量,兼容不同的TAL检测器,并对点采样具有鲁棒性。代码已发布于 https://github.com/joeeeeyin/PSEE。
cs.CV / 36 / 2609.21468

SkillIR: Evolving Scene-Aware Skills for Agentic Image Restoration

SkillIR:面向智能体图像修复的演化式场景感知技能学习
Shao, Jie, Hu, Shengkai, Zhang, Xu, Song, Beihang, Jing, Yongcheng, Wu, Xu, Wan, Jun
Abstract
This paper studies agentic image restoration, in which multimodal agents coordinate specialized restoration tools to recover images affected by complex degradations. Existing restoration agents often derive complete tool-use plans from the original degraded image or retrieve previously successful trajectories, providing limited support for adapting individual actions to evolving intermediate restoration states. We find that accepted tool executions can change the residual degradation state and, consequently, the applicability of subsequent tools. To address this issue, we propose SkillIR, a skill-guided framework that represents restoration experience as degradation-centered action evidence rather than complete tool-use trajectories. SkillIR consolidates context-dependent action outcomes into scene-aware restoration skills that characterize applicable conditions, expected effects, and attributable failure cases. Instead of prescribing a complete restoration plan, the retrieved skills guide one bounded action at a time within a verified residual-state loop: each tool output is treated as a candidate, committed only after transition verification, and followed by reassessment of the active residual degradations. After each rollout, the resulting evidence is used to create, refine, or patch dynamic skills, enabling accumulated restoration experience to improve decision-making for subsequent inputs. Experiments on synthetic and real-world multi-degradation datasets demonstrate that SkillIR improves restoration quality and enables more reliable and effective tool use.
Chinese Translation
本文研究智能体图像修复问题,即多模态智能体协调专门的修复工具以恢复受复杂退化影响的图像。现有的修复智能体通常根据原始退化图像制定完整的工具使用计划,或检索以往成功的轨迹,难以根据不断变化的中间修复状态自适应地调整单个动作。我们发现,已被接受的工具执行会改变残余退化状态,进而影响后续工具的适用性。为解决这一问题,我们提出 SkillIR,一种以技能为引导的框架,它将修复经验表示为以退化为中心的动作证据,而非完整的工具使用轨迹。SkillIR 将依赖上下文的动作结果整合为场景感知的修复技能,这些技能刻画了适用条件、预期效果以及可归因的失败案例。检索到的技能并不规定完整的修复计划,而是在一个经过验证的残余状态循环中,每次仅引导一个有界动作:每个工具输出均被视为候选结果,只有通过状态转移验证后才会被提交,随后对当前残余退化进行重新评估。每次执行(rollout)结束后,所产生的证据被用于创建、精化或修补动态技能,使积累的修复经验能够改善对后续输入的决策。在合成与真实世界多退化数据集上的实验表明,SkillIR 提升了修复质量,并实现了更可靠、更有效的工具使用。
cs.CV / 37 / 2609.21474

MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation

MT-WAM:将单次前向预测表示重新定向以面向动作生成
Yang, Yiguang, Peng, Jiankun, Wang, Xiaoming, Zhang, Yiran, Fang, Zhibo
Abstract
Fast-WAM shows that video-action co-training improves control without generating future video at inference, making the representation from a single video diffusion Transformer forward central to action generation. However, future-observation prediction does not explicitly prioritize the future dynamics and visual structure needed for control. We present MT-WAM, which retains the original training objectives and adds complementary supervision for future two-dimensional point trajectories and visual features. A lightweight dual-stream branch copied from the video backbone's final blocks provides target-specific processing, while a structured attention mask prevents cross-stream attention. Motion-stream tokens supply additional dynamics conditions to the action expert. Future visual-feature prediction provides supervision in a feature space that captures object and spatial structure. This supervision trains the video backbone to provide more informative visual context for action generation under changing visual conditions, without adding visual-feature-stream tokens to action conditioning. At inference, MT-WAM uses video and motion caches computed once per replan and skips future-video prediction. Without additional embodied policy pretraining, MT-WAM achieves 98.2% success on LIBERO and 73.7% on LIBERO-Plus, exceeding Fast-WAM by 23.8 percentage points on the latter. On RoboTwin 2.0 Clean2Rand, Random success increases from 6.30% to 19.40%; across four real-world tasks, average success increases from 67.0% to 77.8%.
Chinese Translation
Fast-WAM 表明,视频-动作协同训练可以在推理时无需生成未来视频的情况下提升控制性能,这使得通过视频扩散 Transformer 的单次前向传播所获得的表示成为动作生成的核心。然而,未来观测预测并未显式地优先考虑控制所需的未来动态信息与视觉结构。我们提出了 MT-WAM,它在保留原有训练目标的基础上,为未来二维点轨迹和视觉特征增加了补充性监督。一个从视频主干网络最后几个模块复制而来的轻量级双流分支提供了针对特定目标的处理,同时结构化的注意力掩码阻止了跨流注意力。运动流 token 为动作专家提供额外的动态条件。未来视觉特征预测在一个能够捕捉物体与空间结构的特征空间中提供监督。这种监督训练视频主干网络在变化的视觉条件下为动作生成提供更具信息量的视觉上下文,且无需向动作条件输入中添加视觉特征流 token。在推理阶段,MT-WAM 使用每次重新规划时仅需计算一次的视频与运动缓存,并跳过未来视频的预测。在无需额外的具身策略预训练的情况下,MT-WAM 在 LIBERO 上达到 98.2% 的成功率,在 LIBERO-Plus 上达到 73.7%,在后者上超过 Fast-WAM 23.8 个百分点。在 RoboTwin 2.0 Clean2Rand 上,随机(Random)成功率从 6.30% 提升至 19.40%;在四项真实世界任务中,平均成功率从 67.0% 提升至 77.8%。
cs.CV / 38 / 2609.21480

OpenSAL360: Open-Source Crowdsourcing Platform for Omnidirectional Video Saliency Collection

OpenSAL360:用于全景视频显著性数据采集的开源众包平台
Bryncev, Alexey, Moskalenko, Andrey, Shilovskaya, Kira, Kosmynin, Ivan, Vatolin, Dmitriy
Abstract
Omnidirectional video saliency prediction plays an important role in many immersive multimedia applications, including viewport-adaptive streaming and compression, foveated rendering, mesh simplification, perceptual quality assessment. Yet progress in this area remains constrained by the cost and complexity of collecting eye-tracking data with VR headsets, which makes large-scale dataset creation difficult to extend. We present OpenSAL360, the first open-source platform for scalable, low-cost 360{\deg} video saliency collection. Unlike conventional VR-based protocols, it requires only a standard screen, mouse, and internet connection, enabling parallel saliency data collection from common crowdsourcing assessors without specialized hardware. We validate our collection protocol against seven well-established VR eye-tracking datasets and conduct ablation studies on key interface, pre-, and post-processing parameters. To demonstrate the effectiveness and scalability of the proposed methodology, we collect and publicly release a saliency dataset covering 500 omnidirectional videos annotated by 2,000+ crowdsourcing assessors, making it, to the best of our knowledge, the largest dataset in this field. We make OpenSAL360 publicly available at https://github.com/msu-video-group/OpenSAL360.
Chinese Translation
全景视频显著性预测在许多沉浸式多媒体应用中发挥着重要作用,包括视口自适应流媒体传输与压缩、注视点渲染、网格简化以及感知质量评估。然而,该领域的进展仍然受限于使用VR头显采集眼动追踪数据的高成本和复杂性,这使得大规模数据集的构建难以扩展。我们提出了OpenSAL360,这是首个用于可扩展、低成本360度视频显著性采集的开源平台。与传统的基于VR的采集方案不同,它仅需标准屏幕、鼠标和互联网连接,即可通过普通众包评估者并行开展显著性数据采集,无需专用硬件。我们将本采集方案与七个成熟VR眼动追踪数据集进行了对比验证,并对关键界面参数以及预处理和后处理参数进行了消融实验。为展示所提方法的有效性和可扩展性,我们采集并公开发布了一个显著性数据集,涵盖500个全景视频,由2,000余名众包评估者进行标注,据我们所知,这是该领域规模最大的数据集。OpenSAL360已公开于 https://github.com/msu-video-group/OpenSAL360。
cs.CV / 39 / 2609.21498

VoxelTTO: Voxel-Aligned Feed-Forward 3D Gaussian Splatting with Test-Time Optimization

VoxelTTO:基于测试时优化的体素对齐前馈式3D高斯泼溅
Zhao, Yibin, Pan, Yihan, Li, Yangwen, Nan, Jun, Yi, Jianjun
Abstract
Recent feed-forward 3D Gaussian Splatting (3DGS) methods typically regress pixel-aligned Gaussian primitives, often causing excessive overlap and artifacts, while inaccuracies in predicted camera poses can lead to misalignment in novel-view synthesis (NVS). We present VoxelTTO, a feed-forward framework for reconstructing geometrically accurate 3DGS scenes from an arbitrary number of images and optional camera parameters. VoxelTTO aggregates dense image features into a global voxel representation and decodes Gaussians from voxel features, breaking the pixel-to-Gaussian correspondence. To exploit known camera parameters while keeping the pretrained visual foundation model (VFM) parameters frozen, we introduce test-time optimization (TTO) that adapts lightweight LoRA modules using pose supervision. We further replace vanilla 3DGS rasterization with stochastic solid volume rendering during training and inference, improving geometric fidelity. Training updates only the voxel-aligned Gaussian reconstruction modules, requiring 80 GPU hours. Experiments on Replica, Tanks and Temples, and DTU demonstrate improved RGB-D NVS and camera-pose estimation relative to prior methods.
Chinese Translation
近期的前馈式3D高斯泼溅(3D Gaussian Splatting, 3DGS)方法通常回归像素对齐的高斯基元,这往往导致高斯基元过度重叠并产生伪影,同时预测相机位姿的不准确可能导致新视角合成(NVS)中的错位问题。我们提出了VoxelTTO,这是一个前馈式框架,能够从任意数量的图像以及可选的相机参数中重建几何精确的3DGS场景。VoxelTTO将密集图像特征聚合到全局体素表示中,并从体素特征解码高斯,从而打破了像素到高斯的对应关系。为了在保持预训练视觉基础模型(VFM)参数冻结的同时利用已知的相机参数,我们引入了测试时优化(Test-Time Optimization, TTO),利用位姿监督来调整轻量级的LoRA模块。此外,我们在训练和推理阶段用随机立体体渲染替代原始的3DGS光栅化,以提升几何保真度。训练仅需更新体素对齐的高斯重建模块,耗时80个GPU小时。在Replica、Tanks and Temples和DTU数据集上的实验表明,相较于先前的方法,本方法在RGB-D新视角合成和相机位姿估计方面均有提升。
cs.CV / 40 / 2609.21502

Adaptive World Memory 3D Foundation Model for Scalable 3D Mapping, Localization, and Rendering

面向可扩展三维建图、定位与渲染的自适应世界记忆三维基础模型
Deng, Tianchen, Shen, Guole, Shen, Yilin, Wu, Wenhua, Fang, Yilin, Ma, Ziqi, Zhang, Tianjun, Yuan, Shenghai, Burgard, Wolfram, Wang, Hesheng
Abstract
Recent 3D foundation models enable generalizable geometric reasoning from RGB images but remain limited in persistent memory, scalability, and renderable scene modeling. We present a memory-centric 3D foundation model for scalable robotic localization, reconstruction, and Gaussian rendering. Its core is an adaptive world memory mechanism that combines transformer-based gated updates with test-time temporal-spatial regulation. Learned gates control recurrent memory propagation, while temporal state evolution and spatial observation-state consistency regulate token-wise updates and forgetting over long image sequences. To support large-scale mapping, we organize memory into local submaps and integrate progressive mapping and tracking, loop closure, and SL(4)-based global refinement to maintain local accuracy and global consistency. A Gaussian reconstruction head decodes memory-enhanced features into renderable primitives, unifying camera pose estimation, dense point-cloud reconstruction, and photorealistic rendering within a single model. Experiments on public benchmarks and self-collected datasets from diverse robotic platforms demonstrate improved trajectory accuracy, reconstruction completeness, and rendering quality over existing 3D foundation reconstruction and SLAM baselines. These results support adaptive memory as a foundation for persistent robotic world modeling. The dataset and code will be made publicly available at \href{https://github.com/dtc111111/AWM-3DFM}{https://github.com/dtc111111/AWM-3DFM}.
Chinese Translation
近期的三维基础模型能够从RGB图像中进行可泛化的几何推理,但在持久记忆、可扩展性和可渲染的场景建模方面仍然存在局限。我们提出了一种以记忆为核心的三维基础模型,用于可扩展的机器人定位、重建和高斯渲染。其核心是一种自适应世界记忆机制,该机制将基于Transformer的门控更新与测试时的时空调节相结合。学习到的门控控制循环记忆传播,而时间状态演化与空间观测-状态一致性则调节长图像序列上逐token的更新与遗忘。为支持大规模建图,我们将记忆组织为局部子图,并集成渐进式建图与跟踪、回环检测以及基于SL(4)的全局优化,以保持局部精度与全局一致性。一个高斯重建头将记忆增强的特征解码为可渲染的图元,从而在单一模型内统一了相机位姿估计、稠密点云重建和照片级真实感渲染。在公开基准以及来自多种机器人平台的自采集数据集上的实验表明,与现有的三维基础重建和SLAM基线相比,本方法在轨迹精度、重建完整性和渲染质量方面均有提升。这些结果证明自适应记忆可作为持久性机器人世界建模的基础。数据集与代码将在 \href{https://github.com/dtc111111/AWM-3DFM}{https://github.com/dtc111111/AWM-3DFM} 公开发布。
cs.CV / 41 / 2609.21516

2D GauSS-MI: Efficient Active Scene Reconstruction with Balanced Visual and Geometric Quality

2D GauSS-MI:兼顾视觉与几何质量的平衡型高效主动场景重建
Xie, Yuhan, Pan, Jia
Abstract
Active reconstruction requires efficient active view selection to achieve high-quality reconstruction within limited onboard computational resources. Existing methods face challenges in adequately balancing visual and geometric quality with the computational efficiency required for real-time operation. In this work, we present an active reconstruction framework based on 2D Gaussian Splatting (2DGS). We develop an efficient online 2DGS mapping pipeline for incremental RGB-D observations and introduce a probabilistic reliability model that characterizes the view-dependent reconstruction quality of individual 2D Gaussian splats. Building on this model, we formulate 2D Gaussian Splatting Shannon Mutual Information (2D GauSS-MI), a mutual-information-based metric that exploits the explicit surface orientation of 2DGS to evaluate the expected information gain of candidate views. The proposed metric enables active view selection to account for both visual and geometric reconstruction quality. We evaluate the proposed system against three state-of-the-art baselines on eight Replica scenes. Experimental results demonstrate that our method achieves a favorable balance between visual and geometric reconstruction quality with substantially lower computational cost and competitive model storage.
Chinese Translation
主动重建需要在有限的机载计算资源下进行高效的主动视角选择,以实现高质量重建。现有方法在充分平衡视觉与几何质量的同时,难以满足实时运行所需的计算效率。本文提出了一种基于二维高斯泼溅(2D Gaussian Splatting,2DGS)的主动重建框架。我们开发了一条面向增量式RGB-D观测的高效在线2DGS建图流水线,并引入一个概率可靠性模型来刻画单个二维高斯基元依赖视角的重建质量。基于该模型,我们提出了二维高斯泼溅香农互信息(2D Gaussian Splatting Shannon Mutual Information,2D GauSS-MI),这是一种基于互信息的度量方法,利用2DGS显式的表面朝向来评估候选视角的期望信息增益。所提出的度量使主动视角选择能够同时考虑视觉和几何重建质量。我们在八个Replica场景上与三种最先进的基线方法进行了对比评估。实验结果表明,我们的方法在视觉与几何重建质量之间取得了良好的平衡,同时计算成本显著更低,且模型存储具有竞争力。
cs.CV / 42 / 2609.21521

VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

VidOmni-Bench:基于跨复杂度与时长的时空事件验证的细粒度视频理解基准
Kim, Changbeen, Chang, Junwon, Kim, Kipyo, Shinoda, Risa, Saito, Kuniaki, Kim, Donghyun
Abstract
While Video Large Language Models (Video-LLMs) have recently demonstrated strong performance, reliably evaluating their fine-grained video understanding remains challenging. Existing benchmarks often rely on question answering or ground-truth caption matching, where models may succeed through superficial cues and incomplete annotations. To this end, we introduce VidOmni-Bench, a benchmark that requires models to verify whether each event in dense video captions is supported by the video. VidOmni-Bench consists of 500 videos spanning five complexity types and diverse durations from 4 seconds to 90 minutes. After collecting videos along these axes, we use diverse Video-LLMs to generate dense captions and obtain human-verified sentence-level labels, where sentences containing incorrect events serve as hard negatives for evaluation. Our experiments on VidOmni-Bench reveal three key findings: (i) Video-LLMs frequently generate hallucinated descriptions in dense video captioning; (ii) they also struggle as verifiers, failing to reliably detect plausible but incorrect event descriptions; and (iii) model weaknesses vary across video complexity and duration, revealing diverse, model-specific bottlenecks in current Video-LLMs.
Chinese Translation
尽管视频大语言模型(Video-LLMs)近来展现出强大的性能,但如何可靠地评估其细粒度视频理解能力仍然具有挑战性。现有基准通常依赖问答或与真实标注字幕的匹配,模型可能通过表面线索和不完整的标注取得成功。为此,我们提出了 VidOmni-Bench,该基准要求模型验证密集视频字幕中的每个事件是否得到视频的支持。VidOmni-Bench 包含 500 个视频,涵盖五种复杂度类型以及从 4 秒到 90 分钟的多样时长。在沿这些维度收集视频后,我们使用多种 Video-LLM 生成密集字幕,并获得经过人工验证的句子级标签,其中包含错误事件的句子作为评估中的困难负样本。我们在 VidOmni-Bench 上的实验揭示了三个关键发现:(i)Video-LLM 在密集视频字幕生成中频繁产生幻觉描述;(ii)它们作为验证器同样表现不佳,难以可靠地检测看似合理但错误的事件描述;(iii)模型的弱点随视频复杂度和时长而变化,揭示了当前 Video-LLM 中多样且因模型而异的瓶颈。
cs.CV / 43 / 2609.21522

Refine Then Fusion: Training-Free 3D Point Cloud Adaptation with Priority Refinement and Multi-Modal Knowledge Fusion

先精炼后融合:基于优先级精炼与多模态知识融合的无训练3D点云自适应方法
Cheng, Hang, Chen, Yan, Fan, Mingyu, Zeng, Long
Abstract
Recent pre-trained foundation models provide rich multi-modal priors for downstream 3D vision tasks. However, the effectiveness of these representations in few-shot scenarios is limited by two fundamental challenges: High-dimensional features often contain substantial channel redundancy and task-irrelevant noise, while the reliability of different modalities varies across samples. Consequently, direct aggregation of heterogeneous representations overlooks sample-dependent modality reliability and may obscure the discriminative cues essential. To address these limitations, we propose Refine Then Fusion(RTF), a training-free framework for few-shot 3D recognition. RTF first identifies discriminative feature channels by jointly modeling inter-class similarity and intra-class stability, thereby decoupling domain-specific knowledge refinement from the cached representations of pre-trained models. It then introduces a reliability-aware fusion mechanism that estimates sample-wise modality reliability from the distribution shifts induced by feature refinement, enabling adaptive aggregation of multi-modal representations. Furthermore, RTF constructs a memory cache that integrates instance-level support features with class-level prototypes to infer query labels. Extensive experiments on five benchmarks demonstrate that RTF consistently outperforms single-modal baselines, partial-fusion variants, and existing lightweight adaptation methods, achieving state-of-the-art few-shot 3D recognition performance without gradient optimization, additional training data, auxiliary training, or parameter updates.
Chinese Translation
近年来,预训练基础模型为下游3D视觉任务提供了丰富的多模态先验。然而,这些表征在少样本场景中的有效性受限于两个根本性挑战:高维特征往往包含大量通道冗余和与任务无关的噪声,且不同模态的可靠性因样本而异。因此,直接聚合异构表征会忽略随样本变化的模态可靠性,并可能掩盖关键的判别性线索。为解决上述局限,我们提出了Refine Then Fusion(RTF),一个面向少样本3D识别的无训练框架。RTF首先通过联合建模类间相似性与类内稳定性来识别判别性特征通道,从而将领域特定知识精炼与预训练模型的缓存表征解耦。随后,它引入了一种可靠性感知的融合机制,从特征精炼所引发的分布偏移中估计样本级模态可靠性,实现对多模态表征的自适应聚合。此外,RTF构建了一个记忆缓存,将实例级支持特征与类级原型相结合以推断查询样本的标签。在五个基准数据集上的大量实验表明,RTF始终优于单模态基线、部分融合变体以及现有的轻量级自适应方法,在无需梯度优化、额外训练数据、辅助训练或参数更新的情况下,实现了最先进的少样本3D识别性能。
cs.CV / 44 / 2609.21541

Purification and Regulation: Comorbidity-Aware Multi-Label Few-Shot Learning for Medical Image Classification

净化与调控:面向医学图像分类的共病感知多标签少样本学习
Lin, Ying-Chih, Kuo, Po-Chih, Chen, Yong-Sheng
Abstract
Multi-label few-shot learning (MLFSL) remains a significant challenge in medical image analysis (MIA). Current metric-based meta-learning methods face two critical limitations in MIA. First, conventional prototype generation often entangles irrelevant disease information, leading to contaminated prototypes and degraded performance. Second, prior studies typically enforce inter-class separability in embedding space, largely neglecting the inherent correlations among diseases. To overcome these challenges, we propose Prototype Purification and Regulation (PPR), a novel MLFSL framework for MIA. PPR first performs prototype purification by leveraging sample-level comorbidity scores to emphasize disease-specific features, producing purified prototypes that better characterize each disease. Building upon these purified prototypes, PPR further addresses the underexplored problem of inter-class prototype distance in MIA by incorporating disease-level comorbidity statistics to adaptively regulate inter-class similarity, forming a comorbidity-aware embedding space. Overall, PPR sequentially enables the model to capture pure disease features and inter-class relationships for reliable MLFSL in MIA. Extensive experiments across four chest X-ray benchmark datasets, including cross-domain evaluation, show that PPR consistently outperforms state-of-the-art methods, significantly improving disease detection while demonstrating robust generalization and clinical applicability.
Chinese Translation
多标签少样本学习(MLFSL)在医学图像分析(MIA)中仍然是一项重大挑战。当前基于度量的元学习方法在医学图像分析中面临两个关键局限:其一,传统的原型生成过程往往会纠缠无关的疾病信息,导致原型被污染、性能下降;其二,已有研究通常在嵌入空间中强制类间可分性,而在很大程度上忽视了疾病之间固有的相关性。为克服这些挑战,我们提出了一种新颖的面向医学图像分析的MLFSL框架——原型净化与调控(Prototype Purification and Regulation, PPR)。PPR首先利用样本级共病得分进行原型净化,以突出疾病特异性特征,从而生成能更好刻画每种疾病的净化原型。在这些净化原型的基础上,PPR进一步通过引入疾病级共病统计信息,自适应地调控类间相似度,构建共病感知的嵌入空间,从而解决了医学图像分析中尚未充分研究的类间原型距离问题。总体而言,PPR使模型能够依次捕获纯净的疾病特征与类间关系,从而在医学图像分析中实现可靠的多标签少样本学习。在四个胸部X光基准数据集(包括跨域评估)上的大量实验表明,PPR持续优于当前最先进的方法,在显著提升疾病检测性能的同时,展现出良好的泛化能力和临床适用性。
cs.CV / 45 / 2609.21543

From Retrieval to Recognition:How Vision--Language Models Become OCR Specialists

从检索到识别:视觉-语言模型如何成为OCR专家
Huangfu, Yuanxiang, Zhong, Hanmeng, Chen, Linqing, Hui, Jeffrey Tiong Jee
Abstract
Does a general vision--language model acquire specialized OCR ability by developing a new reading circuit or by reusing an existing mechanism? We address this question in the setting of full-sequence OCR, rather than local-answer retrieval. Using an evidence-grounded protocol with held-out causal interventions, we identify sparse and stable OCR-head sets in GLM-OCR, MinerU2.5, and PaddleOCR-VL-1.6. We then investigate the mechanistic origin of these OCR heads by comparing them with independently identified textual retrieval/copy heads in general VLMs. Across two general VLMs, visual OCR heads strongly overlap independently identified textual retrieval/copy heads, yielding untuned top-20 intersections of 73.3% and all-head Spearman correlations of 0.677-0.886. The overlap and causal interventions suggest that full-sequence OCR operates as dense sequential multimodal copy-and-paste, repeatedly retrieving visual evidence and routing it to the current output position. Finally, we examine how this shared circuit changes as a general VLM becomes an OCR specialist. Matched base-to-specialized comparisons show that OCR specialization largely preserves head identity, retaining 17-20 of the top 20 heads per task with all-head rank correlations of 0.874-0.942, while redistributing their functional and causal strengths.
Chinese Translation
通用的视觉-语言模型是通过发展出新的阅读回路来获得专门的OCR能力,还是通过复用已有机制?我们在全序列OCR(而非局部答案检索)的设定下研究这一问题。采用基于证据且保留因果干预的实验协议,我们在GLM-OCR、MinerU2.5和PaddleOCR-VL-1.6中识别出稀疏且稳定的OCR头集合。随后,通过将这些OCR头与在通用视觉-语言模型中独立识别出的文本检索/复制头进行比较,我们探究了这些OCR头的机制来源。在两个通用视觉-语言模型上,视觉OCR头与独立识别的文本检索/复制头高度重叠:未微调的top-20交集比例达73.3%,全头Spearman相关系数为0.677–0.886。这种重叠与因果干预结果表明,全序列OCR以密集的序列化多模态“复制-粘贴”方式运作,即反复检索视觉证据并将其路由至当前输出位置。最后,我们考察了当通用视觉-语言模型转变为OCR专家时,这一共享回路如何变化。匹配的基础模型-专用模型对比显示,OCR专门化在很大程度上保留了头的同一性——每个任务在top-20头中保留了17–20个,全头秩相关系数为0.874–0.942——同时重新分配了它们的功能强度与因果强度。
cs.CV / 46 / 2609.21576

GestureFAR: Streaming Co-Speech Gesture Generation with Flow Autoregression

GestureFAR:基于流自回归的流式协同语音手势生成
Liu, Pinxin, Liu, Haiyang, Luo, Jiahao, Huang, Junhua, Zou, Chunhao, Song, Luchuan
Abstract
Generating natural co-speech gestures from streaming speech is essential for embodied conversational agents, where motion must be produced while a user is still speaking. Recent streaming gesture systems make online generation possible by autoregressing over discrete motion tokens, but this design compresses high-dimensional continuous motion into finite codebooks and can limit the realism and diversity of generated gestures. To preserve both causality and continuous expressiveness, we propose \textbf{GestureFAR}, a flow-autoregressive framework for streaming co-speech gesture generation. First, GestureFAR autoregresses over causal continuous motion latents, using a transformer to model streaming audio-motion context and a per-token flow-matching head to sample the next latent from a continuous distribution. Second, we introduce a head-only flow distillation strategy that freezes the causal backbone and distills the multi-step per-token flow head into a single network evaluation using consistency and distribution-matching objectives. This keeps the model token-causal while removing the main latency bottleneck for live interaction. Experiments on BEAT2 show that GestureFAR significantly improves the quality--latency trade-off among streaming-capable methods, preserving strong gesture quality while enabling real-time token-causal generation. Project Page: https://andypinxinliu.github.io/GestureFAR
Chinese Translation
从流式语音中生成自然的协同语音(co-speech)手势对于具身对话智能体至关重要,因为在用户仍在说话时就必须生成动作。近期的流式手势系统通过在离散动作令牌上进行自回归实现了在线生成,但这种设计将高维连续动作压缩到有限的码本中,可能限制生成手势的真实感与多样性。为了同时保持因果性和连续表达能力,我们提出了 GestureFAR,一个用于流式协同语音手势生成的流自回归框架。首先,GestureFAR 在因果的连续动作隐变量上进行自回归,使用 Transformer 对流式音视频-动作上下文进行建模,并通过逐令牌的流匹配(flow-matching)头从连续分布中采样下一个隐变量。其次,我们提出了一种仅针对输出头的流蒸馏策略,冻结因果主干网络,并利用一致性和分布匹配目标将多步逐令牌流头蒸馏为单次网络前向计算。这使模型保持令牌级因果性,同时消除了实时交互的主要延迟瓶颈。在 BEAT2 数据集上的实验表明,GestureFAR 在支持流式生成的方法中显著改善了质量-延迟权衡,在实现实时令牌级因果生成的同时保持了优秀的手势质量。项目主页:https://andypinxinliu.github.io/GestureFAR
cs.CV / 47 / 2609.21593

A benchmark dataset and baseline methods for four-dimensional STEM diffraction patterns

四维扫描透射电子显微镜衍射图样的基准数据集与基线方法
Guan, Yuyan, Zhang, Haoran, Mao, Zian, Yang, Antong, Li, Caifei, Wang, Jialong, Ouyang, Chuying, Wang, Hong, Zeng, Xiaoqin, Xie, Yujun
Abstract
Four-dimensional scanning transmission electron microscopy (4D-STEM) records a two-dimensional diffraction pattern at each electron-probe position, yielding spatially resolved reciprocal-space information but large, heterogeneous data volumes. Here we describe 4D-ImageNet, a collection of 174,000 diffraction patterns comprising 145,000 experimental patterns selected from 29 acquisitions and 29,000 multislice simulations. The experimental data cover acquisition-level labels for Ag, Au, mixed Au-Ag, CoO, Pd and ZnO specimens across multiple fields of view, scan dimensions, camera lengths and exposure times. Each acquisition contributes 5,000 quality-ranked patterns with source scan coordinates and acquisition metadata. A set-prediction detector provides model-derived pseudo-labels for the direct-beam position and Bragg-disk centres, with a confidence score for each disk. The simulation data cover 13 crystal structures and include Euler rotations, reciprocal-space sampling and approximate low-index beam directions. A grouped mixed-domain masked-reconstruction benchmark is provided to assess leakage-resistant loading and evaluation across experimental and simulated data. The dataset is intended for representation learning, disk detection, diffraction-pattern retrieval, orientation analysis and simulation-to-experiment studies.
Chinese Translation
四维扫描透射电子显微镜(4D-STEM)在每个电子探针位置记录一幅二维衍射图样,从而获得空间分辨的倒易空间信息,但同时产生庞大且异构的数据量。本文介绍了4D-ImageNet,一个包含174,000幅衍射图样的数据集,其中145,000幅为从29次采集中筛选的实验图样,29,000幅为多层切片法(multislice)模拟图样。实验数据涵盖Ag、Au、Au-Ag混合物、CoO、Pd和ZnO样品的采集级标签,涉及多种视场、扫描尺寸、相机长度和曝光时间。每次采集提供5,000幅经过质量排序的图样,并附有原始扫描坐标和采集元数据。一套集合预测(set-prediction)探测器为直接电子束位置和Bragg圆盘中心提供由模型生成的伪标签,并为每个圆盘给出置信度评分。模拟数据涵盖13种晶体结构,并包含欧拉旋转、倒易空间采样以及近似的低指数束流方向。我们还提供了一个分组混合域掩码重建(masked-reconstruction)基准,用于评估跨实验与模拟数据的抗泄露加载与评估。该数据集旨在支持表征学习、圆盘检测、衍射图样检索、取向分析以及模拟到实验的迁移研究。
cs.CV / 48 / 2609.21597

HAT: Hypothesis-Anchored Tracking for Video Monocular Spacecraft Pose Estimation

HAT:用于视频单目航天器位姿估计的假设锚定跟踪方法
Lopo, André, Dehban, Atabak, Ventura, Rodrigo
Abstract
Monocular 6-DoF pose estimation of non-cooperative targets is important for on-orbit servicing and debris removal. A single-image estimator can confuse near-symmetric spacecraft orientations, and tracking can preserve an incorrect pose. We present Hypothesis-Anchored Tracking (HAT), a causal framework that uses inter-frame motion to select among competing CAD-based pose hypotheses before alignment and fusion. Rather than independently choosing the highest-scoring hypothesis in each image, HAT retains competing orientation histories and selects a pose to anchor the relative trajectory estimated by monocular SLAM. Sparse anchors and pose fusion provide per-frame estimates after initialization without revising past outputs. The method requires only a calibrated RGB sequence, a metric CAD model, and target image regions, which can be supplied by detection or segmentation. The pretrained pose and SLAM networks require no target-specific training or fine-tuning. We evaluate two versions, Mega-HAT and Pico-HAT, using MegaPose and PicoPose, on SPARK-2024, SwissCube and SHIRT, with YCB-Video assessing performance outside the space domain. Using one temporal configuration per method, the arithmetic means of the four dataset-wise comparisons show 9.4% lower mean pose error and 3.76 times the sustained input FPS for Mega-HAT relative to independent MegaPose, and 23.9% lower mean pose error and 2.42 times the FPS for Pico-HAT relative to independent PicoPose. Mega-HAT ablations on SPARK and an offline reference examine component contributions and the effect of revising past estimates.
Chinese Translation
非合作目标的单目6自由度(6-DoF)位姿估计对于在轨服务与碎片清除具有重要意义。单图像估计器可能混淆近似对称的航天器姿态,而跟踪则可能将错误的位姿保持下去。我们提出了假设锚定跟踪(Hypothesis-Anchored Tracking, HAT),这是一种因果性框架,在对齐与融合之前利用帧间运动在相互竞争的基于CAD模型的位姿假设中进行选择。HAT并非在每幅图像中独立地选择得分最高的假设,而是保留相互竞争的姿态历史,并选择一个位姿来锚定由单目SLAM估计的相对轨迹。稀疏锚点与位姿融合可在初始化之后提供逐帧估计,而无需修改过去的输出。该方法仅需标定过的RGB序列、带尺度信息的CAD模型以及目标图像区域(可由检测或分割提供)。预训练的位姿与SLAM网络无需针对特定目标进行训练或微调。我们基于MegaPose和PicoPose评估了Mega-HAT和Pico-HAT两个版本,数据集包括SPARK-2024、SwissCube和SHIRT,并使用YCB-Video评估其在航天领域之外的表现。在每种方法仅使用一种时序配置的情况下,四个数据集对比结果的算术平均显示:相对于独立运行的MegaPose,Mega-HAT的平均位姿误差降低9.4%,持续输入帧率(FPS)提升3.76倍;相对于独立运行的PicoPose,Pico-HAT的平均位姿误差降低23.9%,FPS提升2.42倍。Mega-HAT在SPARK和一个离线参考基准上的消融实验考察了各组件的贡献以及修改过去估计值所带来的影响。
cs.CV / 49 / 2609.21624

Learned Parametric Emotion Editing: Real-Time Affective Filtering for On-Device Social Media Video

可学习的参数化情绪编辑:面向端侧社交媒体视频的实时情感滤镜
Rochi, Musa, Schubert, Marcel, Gebhardt, Christoph
Abstract
Problematic internet use affects a growing share of the population, yet common interventions, e.g., time limits, blocking, forced breaks, are coercive and easily circumvented. We explore a less restrictive alternative: adapting the emotional intensity of visual content. Prior work has shown that optimization can steer an image's affective content, but its per-image optimization cost makes it impractical for real-time deployment. We instead learn a model that predicts this transformation in a single forward pass: a MobileNetV4 backbone with FiLM-based emotion conditioning outputs parameters for differentiable global transformations. This replaces prior iterative optimization (80 s per image) with a single 3.7 ms forward pass. In a user study (N = 54), the model reduced viewer-reported arousal relative to unedited images, comparably to the grayscale well-being filter, while being rated higher in perceived quality. We integrate the model into an Android app that adapts Instagram video in real time, sustaining 60 fps on a Samsung Galaxy S23.
Chinese Translation
有问题的互联网使用(problematic internet use)影响着日益增长的人群,然而常见的干预措施(如时间限制、屏蔽、强制休息)具有强制性且容易被绕过。我们探索了一种限制性更低的替代方案:调整视觉内容的情绪强度。先前的研究表明,优化方法可以引导图像的情感内容,但其逐图优化的开销使其难以实现实时部署。我们转而学习一个能在单次前向传播中预测该变换的模型:该模型以 MobileNetV4 为主干网络,通过基于 FiLM 的情绪条件化,输出可微分全局变换的参数。这将先前需要迭代的优化(每张图像 80 秒)替换为单次 3.7 毫秒的前向传播。在一项用户研究(N = 54)中,与未编辑的图像相比,该模型降低了观看者自我报告的唤醒度(arousal),效果与灰度健康滤镜相当,同时在感知质量上获得更高评分。我们将该模型集成到一款 Android 应用中,可实时适配 Instagram 视频,并在三星 Galaxy S23 上保持 60 fps 的帧率。
cs.CV / 50 / 2609.21628

Detection is solved, delineation is not: what governs tooth segmentation on panoramic radiographs

检测已解决,精细勾勒尚未解决:是什么决定了全景X光片上的牙齿分割性能
Rehan, Muhammad, Amjad, Moaz, Ahmed, Syed Danial, Adnan, Mariam, Ali, Haider
Abstract
Automatic tooth segmentation and FDI numbering on panoramic radiographs underpins computer-assisted dental diagnosis, yet which factors govern performance remains unclear. We assemble a corpus of 1,422 panoramic radiographs containing 42,142 expert-delineated tooth polygons across the 32-class FDI taxonomy, annotated by 30 dental practitioners and independently reviewed by two others, and use it to isolate input resolution, architecture and anatomical priors under a single evaluation protocol. First, resolution dominates: across a controlled 640/1024/1280 ablation, mask mAP50-95 rises 0.656 -> 0.710 -> 0.717 while mAP50 stays flat at ~0.982. Both gains are significant under a paired bootstrap over images (p < 0.001, p = 0.024); neither mAP50 change is distinguishable from zero. Added resolution buys boundary precision, not detection. Second, architecture is nearly irrelevant in-domain: a query-based transformer with 2.1x the parameters is statistically equivalent to a one-stage detector (95% CI [-0.0064, +0.0064]), only marginally better under domain shift, 5.5x slower on CPU and not executable under standard ONNX runtimes. Third, three targeted interventions fail: a LoRA-adapted self-supervised encoder underperforms, a promptable foundation segmenter degrades masks by 39%, and globally optimal anatomical label assignment yields +0.0007 despite correcting a constraint violated in 40% of out-of-domain predictions. Zero-shot transfer to an independent multi-centre cohort, verified overlap-free, costs 62% of mask mAP50-95 but only 18% of mAP50, reproducing the dissociation. Decomposing masks along the tooth axis localises the residual error to the apical third. Boundary precision is therefore the binding constraint, and effort is better directed at resolution and acquisition diversity than at architectural novelty.
Chinese Translation
全景X光片上的自动牙齿分割与FDI牙位编号是计算机辅助牙科诊断的基础,然而哪些因素决定其性能仍不清楚。我们构建了一个包含1,422张全景X光片、覆盖32类FDI分类体系、共42,142个由专家勾勒的牙齿多边形的数据集,标注工作由30名牙科医生完成,并由另外两名医生独立审核。我们利用该数据集,在统一的评估协议下分离考察输入分辨率、网络架构和解剖先验三个因素的影响。首先,分辨率起主导作用:在640/1024/1280的受控消融实验中,掩膜mAP50-95从0.656提升至0.710再到0.717,而mAP50保持在约0.982基本不变。两项提升在基于图像的配对自助法检验下均具有显著性(p < 0.001,p = 0.024),而mAP50的变化均与零无显著差异。提升分辨率带来的是边界精度的提高,而非检测能力的增强。其次,在域内场景下架构几乎无关紧要:一个参数量为其2.1倍的基于查询的Transformer在统计上与单阶段检测器等效(95%置信区间[-0.0064, +0.0064]),在域偏移下仅略有优势,在CPU上慢5.5倍,且无法在标准ONNX运行时下执行。第三,三种针对性干预均告失败:经LoRA适配的自监督编码器表现不佳,可提示的基础分割模型使掩膜质量下降39%,而全局最优的解剖标签分配仅带来+0.0007的提升,尽管它修正了在40%的域外预测中被违反的约束。在一个独立的多中心队列上的零样本迁移(已验证预测无重叠)损失了62%的掩膜mAP50-95,但仅损失18%的mAP50,重现了这一分化现象。沿牙体长轴对掩膜进行分解,可将残余误差定位在根尖三分之一区域。因此,边界精度是制约性能的瓶颈,研究投入应更多指向分辨率与采集数据多样性,而非架构创新。
cs.CV / 51 / 2609.21629

Extending Decoupled Attention to Dense Prediction and Masked Training for Multi-Channel Images

将解耦注意力扩展至多通道图像的稠密预测与掩码训练
Marikkar, Umar, Husain, Sameed, Awais, Muhammad, Atito, Sara
Abstract
Multi-Channel imaging (MCI) data differs fundamentally from natural images, as each channel records a semantically distinct signal rather than a colour band. To adapt vision encoders to MCI data, Multi-Channel Vision Transformers (MC-ViTs) tokenize each channel independently and concatenate the resulting tokens into one sequence, and the channel count is no longer fixed by the architecture. Self-attention is then computed across all channel-patch tokens with no restriction on which channels attend to which, which dilutes the features of individual channels. The Decoupled Vision Transformer (DC-ViT) regulates this by separating updates computed within a channel from updates computed across channels, and by forming a representation per channel before the channels are combined. Its formulation, however, pairs tokens by spatial position, and thus requires the same visible tokens in every channel. Correspondence under independent per-channel masking is recovered by solving a linear assignment between the retained patches of each channel, which allows decoupled attention to be combined with current masked multi-channel training in its standard configuration rather than a restricted one. Across three classification and three segmentation benchmarks spanning fluorescence microscopy, imaging mass cytometry and satellite imaging, including dense prediction at high channel counts, the resulting formulation outperforms the strongest MC-ViT baseline.
Chinese Translation
多通道成像(MCI)数据与自然图像存在根本差异,因为每个通道记录的是语义上各异的信号,而非某个颜色波段。为了使视觉编码器适应MCI数据,多通道视觉Transformer(MC-ViT)对每个通道独立进行词元化,并将所得词元拼接为一个序列,从而使通道数量不再受架构的固定限制。随后,自注意力在所有通道-图像块词元上计算,不限制哪些通道关注哪些通道,这稀释了各个通道的特征。解耦视觉Transformer(DC-ViT)通过将通道内计算的更新与跨通道计算的更新分离,并在通道合并前为每个通道形成表示来调节这一问题。然而,其公式化方式依据空间位置配对词元,因此要求每个通道具有相同的可见词元。本文通过在各通道保留的图像块之间求解线性分配问题,恢复独立通道掩码下的对应关系,从而使得解耦注意力能够以标准配置而非受限配置与当前掩码多通道训练相结合。在涵盖荧光显微成像、成像质谱流式和卫星影像的三个分类基准与三个分割基准上(包括高通道数下的稠密预测),所提出的公式化方法优于最强的MC-ViT基线。
cs.CV / 52 / 2609.21651

Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis

面向作物病虫害诊断的可配置多阶段视觉流水线
Ganesh, Naga, S, Chandrashekar M, Pedapudi, Lakshmi, Singh, Aakash, Singh, Vineet
Abstract
Farmer.Chat is Digital Green's farm advisory service for smallholder farmers. When something looks wrong with a crop, the farmer takes a photograph and sends it, and that photograph is the whole question: no symptom described, no crop named, often no text at all. The service has to determine whether the picture can be used, what crop it shows, and what is wrong with it, from images taken on cheap phones in a field, in poor light and with a moving camera. The system doing this today cannot be adjusted. It has no adjustable thresholds for photograph rejection, crops and problems cannot be added, and there is no confidence cut-off to set. We study about 1.16 million photographs sent to Farmer.Chat from Ethiopia, India, Kenya and Nigeria. The production quality gate rejected 46.8% of the images it judged, over a quarter of those reaching diagnosis returned no crop name, and 35.8% of the labelled problems filed under "disease" are pests, identifiable without the crop. We therefore split the work into three stages: a quality gate (M0), a crop detector (M1), and a disease or pest detector (M2). Route A fills all three with one fine-tuned vision-language model (Qwen3-VL-4B) answering in a single call. Route B fills each with a small specialist model (DaViT, YOLO26). We replace our production GPT-4o quality gate with a small MobileNetV3 gate at 86.9% F1 in 12 ms. On one test set scored the same way for every system, a hierarchical DaViT-Base achieves 95.41% crop accuracy against 91.46% for the production baseline. It also leads on diagnosis and never declines to answer, while every language model in the comparison leaves a large share of rows with no diagnosis. The fine-tuned model retains two capabilities the specialists do not have: one call for all three stages, and a request for a better photograph when the image cannot support an answer.
Chinese Translation
Farmer.Chat 是 Digital Green 面向小农户的农业咨询服务。当作物出现异常时,农户会拍摄照片并发送,而这张照片就是全部的问题信息:没有任何症状描述,没有指明作物种类,往往连文字都没有。该服务必须仅凭在田间用廉价手机、在弱光和镜头晃动条件下拍摄的照片,判断图片是否可用、显示的是哪种作物以及存在什么问题。目前执行该任务的系统无法调整:没有可调节的照片拒收阈值,无法添加新的作物和问题类别,也没有可设置的置信度截断值。我们研究了从埃塞俄比亚、印度、肯尼亚和尼日利亚发送给 Farmer.Chat 的约116万张照片。生产环境中的质量门控拒绝了其判定图像的46.8%;到达诊断阶段的图像中有超过四分之一未能返回作物名称;而在标注为“病害”的问题中,35.8%实际上是可以脱离作物识别的虫害。因此,我们将任务拆分为三个阶段:质量门控(M0)、作物检测器(M1)以及病虫害检测器(M2)。路线A使用单个经过微调的视觉语言模型(Qwen3-VL-4B)在一次调用中完成全部三个阶段;路线B则为每个阶段配置一个小型专用模型(DaViT、YOLO26)。我们用小型 MobileNetV3 门控替换了生产环境中的 GPT-4o 质量门控,其 F1 值达到86.9%,耗时仅12毫秒。在采用统一评分标准的同一测试集上,分层结构的 DaViT-Base 实现了95.41%的作物识别准确率,而生产基线为91.46%。该模型在诊断方面也领先,且从不拒答,而对比中的所有语言模型都有相当比例的样本未给出诊断。经过微调的模型保留了专用模型所不具备的两项能力:一次调用完成全部三个阶段,以及在图像不足以支持答案时请求用户提供更好的照片。
cs.CV / 53 / 2609.21675

DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal Reasoning

DRT:面向高效且具视觉依据的多模态推理的密集推理轨迹
Xu, Wan, Guo, Yuanfan, Han, Kevin, Chen, LaLa, Zuo, Wangmeng
Abstract
Despite the remarkable progress in Multimodal Large Language Models (MLLMs), prevailing Chain-of-Thought (CoT) paradigms remain confined to the natural-language expression space. Consequently, they inherently incur excessive linguistic overhead, leading to information dilution and weak visual grounding. To address this challenge, we propose Dense Reasoning Trace (DRT), a paradigm that departs from natural-language-centered CoT by expressing reasoning as compact structured traces, which include concise intermediate states with symbolic connectors and disentangle visual observations from logical deductions. First, we introduce the Dense Trace Initialization to internalize the DRT reasoning mode into the model, substantially improving token efficiency while preserving visual evidence. To further enable the model to faithfully capture the logical relations within traces, we propose the Trace-Grounded Reinforcement Learning framework, which builds reference traces through a tri-perspective verification pipeline and employs Trace-Grounded GRPO with structured rewards, encouraging the model to generate concise DRT-style traces with reduced hallucination and stronger logical grounding. Extensive experiments on challenging reasoning benchmarks show that DRT achieves 5.5$\times$ token efficiency improvement while improving 1.3 accuracy points over the Qwen3-VL baseline. These findings suggest that complex multimodal reasoning may not require verbose natural-language traces, opening a more efficient path for next-generation MLLMs. Our code and data are available at: https://github.com/HIT-leaderone/DRT
Chinese Translation
尽管多模态大语言模型(MLLM)取得了显著进展,但主流的思维链(Chain-of-Thought, CoT)范式仍局限于自然语言表达空间。因此,其天然会产生过多的语言开销,导致信息稀释和视觉依据薄弱。为应对这一挑战,我们提出了密集推理轨迹(Dense Reasoning Trace, DRT),这一范式摆脱了以自然语言为中心的CoT,将推理表达为紧凑的结构化轨迹,其中包含由符号连接符连接的简洁中间状态,并将视觉观察与逻辑推演解耦。首先,我们引入密集轨迹初始化(Dense Trace Initialization),将DRT推理模式内化到模型中,在保留视觉证据的同时显著提升token效率。为进一步使模型能够忠实地捕捉轨迹中的逻辑关系,我们提出了基于轨迹依据的强化学习框架(Trace-Grounded Reinforcement Learning),该框架通过三视角验证流程构建参考轨迹,并采用带有结构化奖励的Trace-Grounded GRPO,鼓励模型生成简洁的DRT风格轨迹,从而减少幻觉并增强逻辑依据。在具有挑战性的推理基准上进行的大量实验表明,DRT在实现5.5倍token效率提升的同时,比Qwen3-VL基线提高了1.3个准确率百分点。这些发现表明,复杂的多模态推理可能并不需要冗长的自然语言轨迹,为下一代MLLM开辟了一条更高效的路径。我们的代码和数据可在以下链接获取:https://github.com/HIT-leaderone/DRT
cs.CV / 54 / 2609.21698

Diffusion-Based Tumor Inpainting for Renal Segmentation under Clinical Data Scarcity

基于扩散模型的肿瘤修复方法在临床数据稀缺条件下的肾脏分割
Sedykh, Ekaterina, Ussanov, Salme, Fedorenko, Dmytro, Fishman, Dmytro
Abstract
Deep learning segmentation of renal tumors requires large annotated datasets, yet clinical deployments typically offer only a handful of tumor-positive cases from the target site. We propose a diffusion-based inpainting framework that synthesizes anatomically plausible renal tumors within healthy CT scans, requiring no additional annotation, and provide the first systematic comparison of 2D, 2.5D, and full 3D (MAISI) synthesis strategies for this task. Training the diffusion model on public data (KiTS23, KIRC) and evaluating nnU-Net segmentation on a internal cohort across three low-data regimes, we find that 2.5D and 3D augmentation substantially reduce false positives (from $\sim$18--20\% to $\sim$3--6\%) while maintaining Dice, whereas 2D provides no consistent benefit. Crucially, the proposed 2.5D method matches full 3D synthesis on every metric at substantially lower computational cost, indicating that local volumetric consistency alone is sufficient for effective augmentation in data- and resource-scarce clinical settings.
Chinese Translation
深度学习肾脏肿瘤分割需要大量标注数据,但临床部署中通常仅能从目标机构获得少数肿瘤阳性病例。我们提出一种基于扩散模型的修复框架,可在健康CT扫描中合成解剖结构合理的肾脏肿瘤,无需额外标注,并首次针对该任务系统性地比较了2D、2.5D与全3D(MAISI)合成策略。我们在公开数据(KiTS23、KIRC)上训练扩散模型,并在内部队列上于三种低数据量情境下评估nnU-Net分割性能。结果表明,2.5D和3D数据增显著降低了假阳性率(从约18–20%降至约3–6%),同时保持Dice系数不变,而2D方法则无一致收益。关键的是,所提出的2.5D方法在所有指标上均与全3D合成相当,且计算成本显著更低,表明在数据和资源稀缺的临床环境中,仅靠局部体积一致性即可实现有效的数据增广。
cs.CV / 55 / 2609.21709

SignGPT: Toward LLM-Mediated Sign Language Interaction through Gloss-Free Translation and Generation

SignGPT:通过无注释翻译与生成实现基于大语言模型中介的手语交互
Li, Ronghui, Dong, Jun, Hu, Zhongyuan, Xu, Zunnan, Zhou, Jun, Chen, Liyuan, Liu, Shuoling, Yan, Jiangpeng, Guo, Jie, Li, Xiu, Bao, Linchao
Abstract
Large language models (LLMs) provide limited support for sign language interaction. Unifying sign language translation (SLT) and generation (SLG) to enable sign language as both input and output can reduce switching between separate models during sign-text interaction. We present SignGPT, a unified, pose-based framework for gloss-free SLT and SLG. SignGPT integrates part-aware hierarchical representations of body, hand, and facial motion into a shared language model and employs asymmetric multi-token prediction and progressive training for bidirectional modeling. We evaluate SignGPT on How2Sign (ASL) and Phoenix-2014T (DGS) through benchmark comparisons, qualitative analyses, and component ablations. An exploratory study with 12 Deaf ASL signers assesses an LLM-mediated sign-to-sign response pipeline, highlighting the potential of unified modeling to support sign language conversation (SLC).
Chinese Translation
大语言模型(LLM)对手语交互的支持十分有限。将手语翻译(SLT)与手语生成(SLG)统一起来,使手语既可以作为输入也可以作为输出,能够减少手语-文本交互过程中在多个独立模型之间的切换。我们提出了SignGPT,一个统一的、基于姿态的无注释(gloss-free)手语翻译与手语生成框架。SignGPT将身体、手部和面部动作的部件感知层次化表示整合到共享语言模型中,并采用非对称多词元预测与渐进式训练实现双向建模。我们在How2Sign(美式手语ASL)和Phoenix-2014T(德国手语DGS)数据集上,通过基准对比、定性分析和组件消融实验对SignGPT进行了评估。此外,一项由12名聋人ASL手语者参与的探索性研究评估了大语言模型中介的手语到手语应答流程,凸显了统一建模在支持手语对话(SLC)方面的潜力。
cs.CV / 56 / 2609.21712

ZYT-World: A Real-Time Controllable World Model for Closed-Loop Autonomous-Driving Simulation

ZYT-World:用于闭环自动驾驶仿真的实时可控世界模型
Hu, Boni, Wei, Xiong, Huang, Haoming, Huang, Yong, Wang, Chenbo, Yang, Yi, Wang, Jiancheng, Zhu, Ruicheng, Yang, Zhimin, Liu, Guanglai, Jin, Qiaowan, Wang, Dongzhuo, Kuang, Haiwei, Fan, Jiajun, Wu, Yue, Wei, Jiaxin, Sun, Hao, Yan, Feihong, Bi, Wei, Wang, Kaixuan, Guo, Zichao, Chen, Xiaozhi
Abstract
Generative world models offer controllable and repeatable closed-loop simulation for end-to-end and vision-language-action driving policies, but production deployment exposes three unresolved requirements: faithfully reproducing a mixed fisheye-pinhole rig at native resolutions; reconciling causal, per-timestep interaction with long-horizon stability and low latency; and preserving scene identity when a location is revisited. We present ZYT-World, a single architecture that natively generates four fisheye views with field of view > 180{\deg} and three pinhole views. Projection-specific Plucker adapters encode camera geometry, ego-motion adaptive layer normalization provides global motion control, and a lightweight pixel-aligned layout conditions traffic participants and signals through instance-level boxes, headings and colors. Heterogeneous training combines full-rig geometric coverage with high-resolution detail. Teacher forcing, causal consistency distillation, self-rollout distribution matching distillation, and RigCritic transform a 40-step bidirectional teacher into a one-step, per-latent streaming generator, with RigCritic evaluating the seven-view rig jointly. A 19M-parameter variational autoencoder decoder (TinyVAE), W8A8 quantization, and our inference engine reduce decoding, backbone, and incremental-execution costs, respectively. Finally, cross-trajectory pairs derived from real captures train a plug-in implicit-memory module that preserves place-specific evidence. On the internal multi-view test set, the one-step model retains more than 90% of the teacher's PSNR and SSIM, while FID, FVD, and LPIPS stay within 11% of the teacher. Under the generator-only timing in Figure 2, it is 107.7 times faster than the 40-step bidirectional teacher. TinyVAE decodes 59.8 times faster than Wan. 30s rollouts and cross-trajectory revisits show the intended long-horizon and memory behavior.
Chinese Translation
生成式世界模型为端到端以及视觉-语言-动作(vision-language-action)驾驶策略提供了可控且可复现的闭环仿真,但在生产部署中暴露出三个尚未解决的需求:在原生分辨率下忠实还原鱼眼与针孔相机混合的相机组合;在因果式逐时间步交互与长时程稳定性和低延迟之间取得平衡;以及在重访同一地点时保持场景一致性。我们提出 ZYT-World,这是一种单一架构,可原生生成四个视场角大于180度的鱼眼视图和三个针孔视图。针对投影类型的 Plucker 适配器对相机几何进行编码,自车运动自适应层归一化(ego-motion adaptive layer normalization)提供全局运动控制,轻量级的像素对齐布局通过实例级边界框、朝向和颜色来约束交通参与者与信号灯。异构训练将完整相机组合的几何覆盖与高分辨率细节相结合。通过教师强制(teacher forcing)、因果一致性蒸馏、自滚动分布匹配蒸馏以及 RigCritic,将一个40步双向教师模型转化为单步、逐潜在向量的流式生成器,其中 RigCritic 对七视图相机组合进行联合评估。一个1900万参数的变分自编码器解码器(TinyVAE)、W8A8 量化以及自研推理引擎分别降低了解码、主干网络和增量执行的开销。最后,基于真实采集数据构建的跨轨迹样本对训练了一个可插拔的隐式记忆模块,用于保留特定地点的证据。在内部多视图测试集上,单步模型保留了教师模型90%以上的 PSNR 和 SSIM,而 FID、FVD 和 LPIPS 与教师模型的差距保持在11%以内。在图2中仅统计生成器耗时的计时方式下,其速度比40步双向教师模型快107.7倍。TinyVAE 的解码速度比 Wan 快59.8倍。30秒的滚动生成与跨轨迹重访实验展示了预期的长时程与记忆特性。
cs.CV / 57 / 2609.21743

Balanced Prompt Adaptation against Entropy-Induced Collapse for Test-Time Binary Segmentation

面向测试时二值分割的平衡提示自适应方法:对抗熵致坍缩
Wang, Zhengshan, Webster-Ford, Joshua Charles, Tian, Yifei, Wang, Xinxin, Chen, Long, Ding, Weiping
Abstract
Entropy minimization is a standard objective for test-time adaptation (TTA), but it can fail in imbalanced binary segmentation. Unlike image classification, dense segmentation aggregates thousands of pixel predictions, allowing the larger predicted class to dominate the update, pull minority predictions toward itself, and produce a degenerate mask as predictions saturate and their entropy gradients vanish. We theoretically establish this collapse in a shared-shift model. This analysis motivates Balanced-Anchor Prompt Adaptation (BAPA), which combines two complementary modules. The Class-Balanced Anchors (CBA) module selects high-confidence anchors separately from each predicted class and gives foreground and background equal total loss weight, preventing the larger region from dominating the update. Dynamic Prompt Adaptation (DPA) refreshes these anchors after each prediction update and optimizes only text-side prompt residuals while keeping the vision-language encoders frozen. This prompt-only update refines the foreground-background decision boundary without altering the pretrained dense visual representation. Across experiments from four domains, BAPA achieves the highest mean Dice among the evaluated methods. Factorized ablations further validate the complementary roles of CBA and DPA, supporting balanced prompt adaptation as an effective alternative to entropy minimization for test-time binary segmentation.
Chinese Translation
熵最小化是测试时自适应(Test-Time Adaptation, TTA)的标准目标函数,但在不平衡的二值分割任务中可能失效。与图像分类不同,稠密分割会聚合数千个像素级预测,使得占多数的预测类别主导更新过程,将少数类预测拉向自身,并随着预测饱和、熵梯度消失而产生退化的分割掩码。我们在共享偏移模型下从理论上证明了这种坍缩现象。基于这一分析,我们提出了平衡锚点提示自适应方法(Balanced-Anchor Prompt Adaptation, BAPA),该方法结合了两个互补的模块。类别平衡锚点(Class-Balanced Anchors, CBA)模块分别从每个预测类别中选取高置信度锚点,并赋予前景和背景相等的总损失权重,防止较大区域主导更新过程。动态提示自适应(Dynamic Prompt Adaptation, DPA)模块在每次预测更新后刷新这些锚点,并仅优化文本侧的提示残差,同时保持视觉-语言编码器冻结。这种仅针对提示的更新方式在不改变预训练稠密视觉表征的前提下细化了前景-背景决策边界。在来自四个领域的实验中,BAPA在所评估的方法中取得了最高的平均Dice系数。因子化的消融实验进一步验证了CBA和DPA的互补作用,支持将平衡提示自适应作为测试时二值分割中熵最小化的有效替代方案。
cs.CV / 58 / 2609.21754

SFVO: Decoupled Confidence-Guided Stereo-Flow Visual Odometry with Bidirectional PnP

SFVO:基于解耦置信度引导与双向PnP的立体-光流视觉里程计
Zhang, Kai, Zhao, Guoyang, Ma, Jun
Abstract
Deep learning-based visual odometry (VO) has achieved significant progress, yet most existing methods focus on a monocular approach, which suffers from scale ambiguity. Stereo VO provides real metric by its nature, but remains less studied in deep learning VO due to its high computational cost and modeling complexity. Recent advances in stereo matching and optical flow estimation have made dense visual correspondence increasingly accurate and reliable, but their complementary geometric information has not been fully exploited for VO. In this paper, we present SFVO, a correspondence-driven stereo VO framework that directly builds upon pretrained stereo matching and optical flow models. SFVO exploits pretrained stereo matching and optical flow models to estimate stereo and temporal correspondences. Instead of learning pose directly from images, SFVO maps learned correspondences into geometric constraints and predicts which points are trustworthy. To improve the reliability of visual correspondence-based geometric constraints, we introduce decoupled confidence maps for rotation and translation. This design better aligns the characteristics of visual correspondence and 6-DoF transformations. Extensive experiments on outdoor and indoor datasets demonstrate that SFVO achieves robust and accurate pose estimation with strong generalization capability. The code will be released.
Chinese Translation
基于深度学习的视觉里程计(VO)已取得显著进展,然而现有方法大多聚焦于单目方案,存在尺度不确定性问题。立体视觉里程计(Stereo VO)天然可提供真实的度量尺度,但由于计算成本高、建模复杂,在深度学习视觉里程计领域的研究相对较少。近年来立体匹配与光流估计技术的进展使稠密视觉对应关系的估计日益精确可靠,但其互补的几何信息尚未在视觉里程计中得到充分挖掘。本文提出SFVO,一个以对应关系为核心的立体视觉里程计框架,直接构建于预训练的立体匹配与光流模型之上。SFVO利用预训练的立体匹配与光流模型来估计立体对应关系和时序对应关系。与直接从图像学习位姿不同,SFVO将学习到的对应关系映射为几何约束,并预测哪些点是可信的。为提高基于视觉对应关系的几何约束的可靠性,我们引入了针对旋转和平移的解耦置信度图。这一设计更好地契合了视觉对应关系与6自由度(6-DoF)变换的特性。在室外和室内数据集上的大量实验表明,SFVO能够实现鲁棒、精确的位姿估计,并具有强大的泛化能力。代码将会开源发布。
cs.CV / 59 / 2609.21763

Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

超越基准分数:对用于胸片结核筛查的医学视觉-语言模型进行审计
Akhtar, Mushir, Tanveer, M., Arshad, Mohd.
Abstract
A medical model's benchmark score does not establish that the same conclusion holds under a different evaluation. This study tests whether claims about model ranking, score reliability and screening performance survive changes in cohort, prompt, negative spectrum, specified prevalence and operating threshold. We audit three medical vision-language models (BioMedCLIP, CheXficient, and MedSigLIP) and a general-domain OpenCLIP comparator on 12,200 chest radiograph records from four datasets (Montgomery, Shenzhen, TBX11K, and VinDr-CXR). Five fixed prompt families yield 244,000 model--image--prompt scores. No model leads every cohort and reliability criterion. Prompt-family changes alter AUROC in 21 of 48 multiplicity-controlled comparisons. Replacing healthy controls with sick non-tuberculosis controls reduces AUROC by 0.075--0.306 across all four models. On VinDr-CXR, the three medical models distinguish tuberculosis from no-finding controls substantially better than from pneumonia or lung tumor; their AUROC point estimates for both named diseases fall below 0.5. CheXficient has documented VinDr-CXR pretraining exposure, which limits the interpretation of its results. Thresholds chosen for 95\% sensitivity on TBX11K training retain that constraint by point estimate in only four of sixteen target evaluations. A five-seed supervised source model reaches 0.999 AUROC on TBX11K validation but 0.629 on each of two external cohorts. Conservative exclusion of perceptual-overlap candidates narrows this gap without closing it. These retrospective, single-task results show that discrimination, score reliability and threshold retention support different portability claims. Evidence for chest X-ray tuberculosis screening should identify the complete evaluation specification rather than attribute clinical portability to a checkpoint alone.
Chinese Translation
医学模型的基准分数并不能证明同一结论在不同评估条件下依然成立。本研究检验关于模型排名、分数可靠性和筛查性能的结论在队列、提示词、阴性谱、设定患病率及操作阈值发生变化时是否仍然成立。我们对三个医学视觉-语言模型(BioMedCLIP、CheXficient 和 MedSigLIP)以及一个通用领域的 OpenCLIP 对照模型进行了审计,评估数据来自四个数据集(Montgomery、Shenzhen、TBX11K 和 VinDr-CXR)共 12,200 条胸片记录。五个固定的提示词族共产生 244,000 个模型-图像-提示词分数。没有任何模型在所有队列和所有可靠性标准上均领先。在 48 项经多重性控制的比较中,提示词族的改变导致 21 项的 AUROC 发生变化。用患病的非结核对照替代健康对照,会使所有四个模型的 AUROC 降低 0.075 至 0.306。在 VinDr-CXR 上,三个医学模型区分结核与“无异常”对照的能力明显优于其区分结核与肺炎或肺肿瘤的能力;对于这两种具名疾病,其 AUROC 点估计值均低于 0.5。CheXficient 存在已记录的 VinDr-CXR 预训练暴露,这限制了对其结果的可解释性。在 TBX11K 训练集上按 95% 敏感度选定的阈值,仅在十六项目标评估中的四项以点估计值保持该约束。一个五种子的有监督源模型在 TBX11K 验证集上达到 0.999 的 AUROC,但在两个外部队列上分别仅为 0.629。保守地剔除存在感知重叠的候选模型可缩小这一差距,但未能消除。这些回顾性、单任务的结果表明,判别能力、分数可靠性和阈值保持性支持不同的可迁移性结论。关于胸片结核筛查的证据应明确完整的评估规范,而不能仅凭一个模型检查点(checkpoint)就断言其临床可迁移性。
cs.CV / 60 / 2609.21770

XCalib Depth-Guided Geometric Optimization for Dense Thermal-Visible Video Registration

XCalib:用于稠密热红外-可见光视频配准的深度引导几何优化方法
Godet, Aurelien, Jobert, Gabriel, Mura, Mauro Dalla
Abstract
Image registration is a vital preprocessing step in multimodal perception tasks, including image fusion, object detection, and semantic segmentation. In Advanced Driver- Assistance Systems (ADAS), spatial misalignment between visible (RGB) and infrared (IR) cameras -caused by non-coincident optical axes and field-of-view differences- introduces non-uniform parallax and visual ghosting. Classical keypoint-based methods are restricted to global homographies that fail under dynamic depth, while unconstrained dense flow algorithms lack structural regularization and suffer from temporal instability. In this paper, we propose XCalib, an unsupervised dense thermal-visible registration framework that bridges this gap. Rather than serving as an absolute metric calibration tool, XCalib leverages virtual pinhole camera parameterization strictly as a geometric constraint space. By optimizing effective relative pose and intrinsics alongside predicted monocular metric depth, XCalib restricts the search space of spatial displacements to physically valid projection geometries. Our key contributions are: (1) a novel registration paradigm that uses camera parameterization as an implicit regularizer for dense cross-modal warping; (2) Normalized Edges Correlation (NEC), a robust structural similarity metric tailored to cross- spectral alignment; and (3) extensive quantitative and qualitative evaluations across public ADAS datasets, demonstrating superior temporal stability and alignment accuracy over unconstrained dense flow baselines.
Chinese Translation
图像配准是多模态感知任务(包括图像融合、目标检测和语义分割)中的关键预处理步骤。在高级驾驶辅助系统(ADAS)中,可见光(RGB)相机与红外(IR)相机之间由于光轴不重合和视场差异而导致的空间失配,会引入非均匀视差和视觉重影。经典的基于关键点的方法仅限于全局单应性变换,在动态深度条件下会失效;而无约束的稠密光流算法缺乏结构正则化,且存在时间不稳定性问题。本文提出了XCalib,一种无监督的稠密热红外-可见光配准框架,以弥补这一空白。XCalib并非作为绝对的度量标定工具使用,而是将虚拟针孔相机参数化严格用作几何约束空间。通过在预测单目度量深度的同时优化有效相对位姿和内参,XCalib将空间位移的搜索空间限制在物理上合理的投影几何范围内。我们的主要贡献包括:(1)一种新颖的配准范式,将相机参数化作为稠密跨模态变形的隐式正则化手段;(2)归一化边缘相关性(Normalized Edges Correlation, NEC),一种专为跨光谱对齐设计的鲁棒结构相似性度量;(3)在多个公开ADAS数据集上进行了大量定量与定性评估,结果表明相较于无约束的稠密光流基线方法,本方法具有更优的时间稳定性和对齐精度。
cs.CV / 61 / 2609.21780

PointLAM: Local Attentive Mamba for Efficient Point-based 3D Object Detection

PointLAM:用于高效点基3D目标检测的局部注意力Mamba
Shang, Xuanming, Zhang, Weijia, Ma, Chao
Abstract
3D object detection from LiDAR point clouds faces a fundamental dilemma: voxel-based methods achieve efficiency at the cost of geometric quantization, while point-based methods preserve fidelity but suffer from prohibitive computational bottlenecks. Specifically, point-based architectures are crippled by slow downsampling strategies (e.g., FPS) and expensive dynamic neighbor queries (e.g., k-NN) coupled with costly continuous interactions. To tackle these systemic inefficiencies, we propose PointLAM, a highly efficient and powerful point-based architecture driven by two synergistic innovations. First, to resolve the downsampling bottleneck, we develop the Laplacian Point Sampler (LPS). LPS employs an implicit discrete Laplacian high-pass filter and Doubly Sorted Sampling to achieve fast, structure-aware foreground preservation. Second, to overcome local modeling latency, we design the Local Hadamard Aggregator (LHA). LHA decouples spatial indexing from feature representation using transient grids, and replaces complex continuous interactions with a Hadamard Gating mechanism for topology-aware, attentive modulation. By coupling this local gating with Bi-Directional Mamba (BDM) layers for global sequence modeling, we formulate the Local Attentive Mamba (LAM) block. Powered by this architecture, PointLAM achieves competitive performance on nuScenes and Waymo for point-based detectors. It rivals highly optimized voxel competitors while requiring a fraction of the computational footprint, demonstrating marked superiority in detecting small instances and handling extreme sparsity. Project page: https://pointlam.github.io/.
Chinese Translation
基于LiDAR点云的3D目标检测面临一个根本性困境:基于体素的方法以几何量化损失为代价换取高效率,而基于点的方法虽保持了几何保真度,却受制于难以承受的计算瓶颈。具体而言,基于点的架构受制于缓慢的下采样策略(如FPS)以及代价高昂的动态邻域查询(如k-NN)与昂贵的连续交互。为解决这些系统性低效问题,我们提出PointLAM,一种由两项协同创新驱动的高效且强大的点基架构。首先,为解决下采样瓶颈,我们开发了拉普拉斯点采样器(Laplacian Point Sampler, LPS)。LPS采用隐式离散拉普拉斯高通滤波器与双重排序采样(Doubly Sorted Sampling),实现快速、结构感知的前景点保留。其次,为克服局部建模延迟,我们设计了局部Hadamard聚合器(Local Hadamard Aggregator, LHA)。LHA利用瞬态网格将空间索引与特征表示解耦,并用Hadamard门控机制替代复杂的连续交互,实现拓扑感知的注意力调制。通过将该局部门控与用于全局序列建模的双向Mamba(Bi-Directional Mamba, BDM)层耦合,我们构建了局部注意力Mamba(Local Attentive Mamba, LAM)模块。基于该架构,PointLAM在nuScenes和Waymo数据集上为点基检测器取得了具有竞争力的性能。它可媲美高度优化的体素方法,同时仅需其一小部分计算开销,在检测小目标实例和处理极端稀疏性方面展现出显著优势。项目页面:https://pointlam.github.io/。
cs.CV / 62 / 2609.21800

A Principled Approach to Unsupervised Anomaly Detection

一种无监督异常检测的原理性方法
Myles, James, Baugh, Matthew, Müller, Johanna P., Kainz, Bernhard, Li, Yingzhen
Abstract
Traditional unsupervised anomaly detection (UAD) methods are designed to flag or localise deviations from a normative distribution, ignoring the underlying generative mechanisms of the anomalies. Yet the nature of an anomaly is often as important as its presence. We reformulate UAD as a Bayesian inverse problem, in which the objective is to infer the most probable corruption responsible for each observation. Our framework yields a probabilistic anomaly score as the energy of the inferred corruption parameters, and serves as a principled recipe for developing new UAD algorithms. We derive several existing methods as instances of the general framework, each corresponding to the same energy score under different modelling choices. Experimentally, we study the framework's components in a controlled setting, and improve object-class AUROC on the MVTec AD dataset by 2.3% by adapting the underlying corruption model. Finally, we validate the framework on a brain MRI benchmark, achieving strong detection performance while producing estimates of pathology intensity, bias, and geometry. Code is available at https://github.com/jgmyles/inverse-uad.
Chinese Translation
传统的无监督异常检测(Unsupervised Anomaly Detection, UAD)方法旨在标记或定位偏离规范分布的偏差,而忽略了异常背后潜在的生成机制。然而,异常的性质往往与其存在本身同等重要。我们将 UAD 重新表述为一个贝叶斯逆问题,其目标在于推断导致每个观测的最可能的损坏(corruption)。我们的框架以所推断的损坏参数的能量作为概率化的异常评分,并为开发新的 UAD 算法提供了一套原理性的方法。我们将若干现有方法推导为该通用框架的实例,它们分别对应于不同建模选择下的同一种能量评分。在实验中,我们在受控环境下研究了该框架的各个组成部分,并通过调整底层的损坏模型,在 MVTec AD 数据集上将物体类别的 AUROC 提高了 2.3%。最后,我们在脑部 MRI 基准数据集上验证了该框架,在取得优异检测性能的同时,还能够估计病理强度、偏置和几何形状。代码已发布于 https://github.com/jgmyles/inverse-uad。
cs.CV / 63 / 2609.21804

VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph

VideoReloc:基于千字节级语义场景图的长时室内视频重定位
Li, Qianru, Chen, Xuyang, Wang, Xuqin, Zhang, Zhenghao, Luo, Hongyi, Wu, Tao, Cremers, Daniel, Liu, Lu, Zhang, Yanfeng
Abstract
Given a compact semantic scene graph, long-term indoor video relocalization estimates a map-frame trajectory after lighting and furniture changes. Visual methods rely on appearance and become unreliable under these changes; localizing one frame at a time from object classes and geometry instead leaves sparse, ambiguous evidence. We introduce VideoReloc, whose adaptive clips use odometry to gather spatial evidence until object and motion criteria are met, adapting query length to the observed scene. Its run-level decision rechecks conflicting placements using evidence accumulated across connected clips, stabilizing the trajectory beyond adjacent-clip tracking. Hypothesis-first registration proposes poses from object triplets and verifies each using clip-wide object centers and box surfaces. Orientation-aware refinement uses box faces, gravity and wall directions to resolve ambiguity in camera orientation and refine the full pose. This reframes sparse-map relocalization as verification of spatially extended video queries, moving discriminative support from stored appearance to temporal context and permitting a 100 kB map of class-labelled boxes. On RIO10 and ReplicaCAD, the all-frame localization success rate at 1 m/10$^\circ$ is 73.5% and 61.1% under causal evaluation, rising to 90.6% and 74.8% with clip closure. The evaluated per-frame scene coordinate regressors reach up to 47.6% and 49.8%, respectively, with maps of 12.6-42 MB. Project page: https://videoreloc.github.io
Chinese Translation
给定一个紧凑的语义场景图,长时室内视频重定位旨在光照和家具发生变化后估计地图坐标系下的轨迹。视觉方法依赖外观信息,在这些变化下变得不可靠;而基于物体类别与几何信息逐帧定位的方式则只能利用稀疏且模糊的证据。我们提出VideoReloc,其自适应视频片段利用里程计持续收集空间证据,直到满足物体与运动准则为止,从而根据所观察到的场景自适应地调整查询长度。其片段级(run-level)决策利用跨相连片段累积的证据对冲突的位姿进行复核,使轨迹在相邻片段跟踪之外保持稳定。假设优先的配准方法从物体三元组中提出候选位姿,并利用片段级物体中心与包围盒表面逐一验证。面向朝向的精细化方法利用包围盒表面、重力方向和墙面方向来消除相机朝向的歧义并优化完整位姿。该方法将稀疏地图重定位重新定义为对空间上延展的视频查询的验证,将判别性支撑从存储的外观转移到时序上下文,并允许使用仅100 kB的类别标注包围盒地图。在RIO10和ReplicaCAD数据集上,在因果(在线)评估下,1米/10度阈值的全帧定位成功率为73.5%和61.1%,引入片段闭环后分别提升至90.6%和74.8%。所评估的逐帧场景坐标回归方法分别最高仅达到47.6%和49.8%,且其地图大小为12.6–42 MB。项目主页:https://videoreloc.github.io
cs.CV / 64 / 2609.21822

Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty

目标检测基准数据集并不完整:标签错误与标注不确定性的作用
Penquitt, Sarina, Klees, Jonathan, van Betteray, Antonia, Jashnieh, Parssa, Stehr, Peter, Rottmann, Matthias, Schmarje, Lars
Abstract
While object detection has advanced through improved architectures and open-vocabulary models, we provide strong evidence that benchmark quality is limited by annotation incompleteness. Across four widely used datasets (COCO, Pascal VOC, Cityscapes, KITTI), re-annotation reveals substantial increases in annotated objects (e.g., up to +60% on KITTI and +40% on COCO), driven primarily by previously unlabeled small, occluded, or densely packed instances. While some differences arise from dataset-specific annotation conventions, we consistently find that missing annotations are the main source of label errors across all datasets. To achieve high data quality, we introduce a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object. The resulting annotations improve coverage and align well with human calibration. We show that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable. We introduce two large-scale benchmarks: (i) an uncertainty-aware object detection benchmark, and (ii) a label error detection benchmark grounded in real label errors. We show that current detectors are strongly depended on annotation quality and are misaligned with human perception. Current label error detection methods, which have been shown to perform well on synthetic noise, struggle to achieve high recall and precision on real label errors. Our results highlight the need for future object detection benchmarks to move beyond deterministic annotations toward high-recall, uncertainty-aware evaluation that maximizes valid instances and better reflects real-world ambiguity.
Chinese Translation
尽管目标检测通过改进网络架构和开放词汇模型(open-vocabulary models)取得了显著进展,我们提供了有力的证据表明,基准数据集的质量受限于标注的不完整性。在四个广泛使用的数据集(COCO、Pascal VOC、Cityscapes、KITTI)上进行的重新标注显示,标注对象数量大幅增加(例如,KITTI 上最多增加 60%,COCO 上增加 40%),这主要源于此前未被标注的小目标、被遮挡或密集排列的实例。虽然部分差异源于数据集特有的标注规范,但我们一致发现,缺失标注是所有数据集标签错误的主要来源。为实现高质量数据,我们引入了一种可扩展的标注流程,该流程强调高召回率,并通过至少 11 名标注者对每个对象进行聚合的软标签(soft labels)来捕获标注的模糊性。由此得到的标注提高了覆盖范围,并与人类校准高度一致。我们证明基准测试性能对标注质量高度敏感,尽管模型排名总体上保持稳定。我们引入了两个大规模基准:(i)不确定性感知的目标检测基准;(ii)基于真实标签错误的标签错误检测基准。我们表明,当前的检测器严重依赖于标注质量,且与人类感知不一致。当前的标签错误检测方法虽然在合成噪声上表现出色,但在真实标签错误上难以实现高召回率和高精确率。我们的结果强调,未来的目标检测基准需要超越确定性标注,转向高召回率、不确定性感知的评估方式,以最大化有效实例并更好地反映现实世界中的模糊性。
cs.CV / 65 / 2609.21866

Morphology-Aware Ambiguity Learning for Wafer Defect Decision Support

面向晶圆缺陷决策支持的形态感知歧义学习
Chu, Seungjun, Chung, Seokhyun
Abstract
Wafer map defect recognition is commonly formulated as a fixed-taxonomy classification problem that assigns each wafer to a single defect class. However, some wafers exhibit morphologies near class boundaries, for which forcing a single prediction may be less informative than providing plausible diagnostic alternatives. This paper proposes a morphology-aware ambiguity learning framework that supports three diagnostic actions: automatic single-class diagnosis, assisted diagnosis with two plausible defect classes, and full review. Using the radial, angular, and geometric characteristics of training wafer maps, the framework constructs a class-level ambiguity matrix representing defect-class pairs with similar morphology and plausible diagnostic alternatives. It guides the model to learn plausible alternative classes rather than treating all incorrect classes equally. During inference, the matrix determines whether an uncertain prediction can be represented by a meaningful two-class diagnostic set or should be escalated for full review. Experiments on WM-811K show that the proposed framework outperforms conventional approaches in defect recognition and diagnostic decision support, providing meaningful two-class alternatives while reserving full review for cases with unresolved ambiguity. Illustrative cost analyses further show the potential cost advantage of the proposed routing strategy. The diagnostic behavior of the framework remains consistent across different backbone architectures.
Chinese Translation
晶圆图缺陷识别通常被构建为一个固定类别的分类问题,即将每个晶圆分配到单一的缺陷类别。然而,某些晶圆呈现出接近类别边界的形态特征,对于这类晶圆,强行给出单一预测可能不如提供若干合理的诊断备选方案更具参考价值。本文提出一种形态感知的歧义学习框架,支持三种诊断操作:自动单类别诊断、基于两个合理缺陷类别的辅助诊断,以及全面人工复检。该框架利用训练晶圆图的径向、角度和几何特征,构建一个类别级歧义矩阵,用于表征形态相似且诊断备选合理的缺陷类别对,并引导模型学习合理的备选类别,而非将所有错误类别一视同仁。在推理阶段,该矩阵用于判定某个不确定的预测是能够以有意义的双类别诊断集合表示,还是应当升级进行全面复检。在 WM-811K 数据集上的实验表明,所提出的框架在缺陷识别和诊断决策支持方面优于传统方法,能够提供有意义的双类别备选方案,同时将全面复检保留给歧义无法解决的案例。示例性的成本分析进一步展示了所提出的分流策略在成本方面的潜在优势。该框架的诊断行为在不同主干网络架构下保持一致。
cs.CV / 66 / 2609.21872

Chronosphere: Space-Time Tessellation of Local Climate Experts

Chronosphere:局部气候专家的时空镶嵌
Cher, Daniel, Xing, Eric, Li, Kexing, Wei, Brian, Corley, Isaac, Jacobs, Nathan
Abstract
We introduce Chronosphere, a spatio-temporal neural field that learns representations of climate. A central challenge in geographic representation learning is modeling environmental processes whose spatial and temporal complexity varies widely. Yet existing location encoders typically fix a single level of detail everywhere. Global bases such as spherical harmonics spread capacity uniformly across space and time. Localized bases resolve only predefined regions. Learned tessellations adapt, but are inefficient at representing higher frequencies. Chronosphere unifies these approaches, pairing an adaptive tessellation of learnable sites on the spacetime torus $S^2\times S^1$ with a shared bank of local basis functions. Both where capacity is placed and how much detail each region carries adapt to the data, across space and time. Trained to reconstruct climatology, Chronosphere matches or leads state-of-the-art location encoders across spatial and temporal tasks, with the largest gains under spatial and temporal transfer.
Chinese Translation
我们提出了Chronosphere,一个学习气候表征的时空神经场。地理表征学习的一个核心挑战是对空间和时间复杂度差异巨大的环境过程进行建模。然而,现有的位置编码器通常在所有地方都固定单一的细节层级。诸如球谐函数之类的全局基函数将容量在空间和时间上均匀分布;局部化基函数只能解析预定义的区域;可学习的镶嵌方法具有自适应性,但在表示高频信息时效率低下。Chronosphere统一了这些方法,将时空环面 $S^2 imes S^1$ 上可学习节点的自适应镶嵌与共享的局部基函数库相结合。容量放置的位置以及每个区域承载的细节量均随数据自适应调整,跨越空间和时间。经过气候态重建训练后,Chronosphere在空间和时间任务上达到或超越了最先进的位置编码器,其中在空间和时间迁移任务中收益最大。
cs.CV / 67 / 2609.21879

Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition

开放权重视觉语言模型在人脸识别中的解释质量基准测试
Colbois, Laurent, Marcel, Sébastien
Abstract
Vision-Language Models (VLMs) have recently been proposed as promising tools for face recognition, as they can produce natural language explanations alongside similarity scores. This capability is considered appealing for face comparisons in forensic contexts, which require decisions to be transparent and auditable. However, existing evaluations of VLMs for that use case focus mostly on recognition accuracy, while the validity of generated explanations remains unquantified. In this work, we introduce a benchmarking framework for VLM-based face recognition that treats explanation quality as a core evaluation axis. We propose two criteria that explanations should satisfy: relevance, i.e., reliance on identity-stable facial features; and faithfulness, i.e., alignment with the visible image content without hallucinated features. We jointly develop a methodology enabling the quantification of relevance and faithfulness of evaluated models, based on constraining model outputs to a structured explanation format that supports automated querying and auditing. Using this framework, we benchmark several families of open-weight VLMs, jointly evaluating face verification accuracy and explanation quality. Our results highlight remaining shortcomings of produced explanations, and emphasize the need for such explanation quality metrics to get a complete picture of model performance. The proposed benchmark and open-source evaluation harness provide a foundation for proper benchmarking and future fine-tuning of explainable face recognition systems.
Chinese Translation
视觉语言模型(Vision-Language Models, VLMs)近期被提出作为人脸识别的有前景工具,因为它们能够在给出相似度分数的同时生成自然语言解释。这一能力在司法取证场景中的人脸比对任务中颇具吸引力,因为此类场景要求决策过程透明且可审计。然而,现有针对该用例的VLM评估主要聚焦于识别准确率,而所生成解释的有效性尚未得到量化。本文提出了一个面向基于VLM的人脸识别的基准测试框架,将解释质量作为核心评估维度。我们提出了两项解释应满足的准则:相关性(relevance),即依赖身份稳定的人脸特征;忠实性(faithfulness),即与图像可见内容一致且不产生幻觉特征。我们同时开发了一种量化方法,用于度量被评估模型的解释相关性与忠实性,该方法通过将模型输出约束为结构化的解释格式,从而支持自动化查询与审计。利用该框架,我们对多个开放权重VLM系列进行了基准测试,联合评估人脸验证准确率与解释质量。我们的结果凸显了所生成解释仍存在的不足,并强调需要此类解释质量指标才能全面刻画模型性能。所提出的基准测试与开源评估工具为可解释人脸识别系统的规范基准测试及未来微调奠定了基础。
cs.CV / 68 / 2609.21887

Catena: A Comprehensive Software Suite for Large-Scale Connectomics

Catena:一个面向大规模连接组学的综合性软件套件
Mohinta, Samia, Gómez-Gálvez, Pedro, Lee, Shi Yan, Franco-Barranco, Daniel, Clayton, Michael, Preibisch, Stephan, Funke, Jan, Cardona, Albert
Abstract
The gold standard datasets for mapping connectomes are electron microscopy volumes of densely labeled neural tissue at nanometer resolution. Yet reconstructing and proofreading neuronal arbors and annotating all synapses requires pipelining multiple software tools that are often fragmented, inconsistently maintained, or proprietary, hindering reproducibility and automation. Here, we introduce Catena, an open-source, comprehensive, developer-centric software suite for connectomics that integrates modules for 3D neuron and organelle segmentation, synapse detection, microtubule tracking, and neurotransmitter inference. Catena organizes its modules in composable, chunk-wise processing pipelines in a completely documented, extensible, and adaptable design. We further reduce compute and ground-truth data requirements with pretrained machine learning models, facilitating fine-tuning. Catena ships fully containerized modules that encapsulate evolving dependencies for consistent execution across workstations and clusters. By consolidating open components, shareable models, and containerized runtimes, Catena delivers a reproducible and scalable approach to mapping cellular connectomes from electron microscopy volumes. Code and documentation: https://github.com/Mohinta2892/catena.git
Chinese Translation
绘制连接组学的黄金标准数据集是纳米级分辨率下密集标记神经组织的电子显微镜体数据。然而,重建和校对神经元突起以及注释所有突触需要对多个软件工具进行流水线化处理,而这些工具往往碎片化、维护不一致或为专有软件,阻碍了可重复性和自动化。在此,我们介绍 Catena,一个开源、综合、面向开发者的连接组学软件套件,其集成了三维神经元与细胞器分割、突触检测、微管追踪以及神经递质推断等模块。Catena 将各模块组织成可组合的、分块处理的流水线,采用完全文档化、可扩展且易于适配的设计。我们进一步利用预训练的机器学习模型降低了计算量和真实标注数据的需求,便于模型微调。Catena 提供完全容器化的模块,将不断演变的依赖关系封装其中,以确保在工作站和集群上的一致执行。通过整合开放组件、可共享模型和容器化运行时,Catena 提供了一种可重复、可扩展的方法,用于从电子显微镜体数据绘制细胞连接组。代码与文档:https://github.com/Mohinta2892/catena.git
cs.CV / 69 / 2609.21903

The Role of Radiometric Features in Cross-Site Leaf-Wood Segmentation of LiDAR Point Clouds

辐射特征在LiDAR点云跨站点叶-木分割中的作用
Kaharlytskyi, Roman, Robinson, Derek T., Guglielmi, Roberto
Abstract
Leaf-wood segmentation of individual trees from LiDAR point clouds is essential for quantitative structure models (QSMs) used in non-destructive biomass estimation. Existing segmentation methods typically exclude radiometric features (e.g., intensity, return number) to maximize cross-sensor compatibility. We challenge this design choice by evaluating cross-site and cross-platform generalization: training on the public Heidelberg dataset (terrestrial TLS, 1550nm) and testing on a novel dataset from Ontario, Canada (RPA-LS, 905nm). Results show that geometry-only methods - including state-of-the-art deep learning models trained on high-density LiDAR datasets - fail to generalize to the sparse, top-down geometry of aerial scans, achieving F1 scores <= 0.56. Incorporating radiometric features (intensity, return number, number of returns) improves F1 to 0.61, but more critically, increases wood recall by 119% from 0.16 to 0.35. Furthermore, geometry-only approaches often result in fragmented stem and branch components. We find that leveraging radiometric features preserves greater structural connectivity, resulting in more coherent architectures that are better suited for QSM reconstruction. We demonstrate that while geometric patterns are view-dependent and prone to overfitting scan patterns, radiometric features encode physical material properties that generalize across disparate sensors and environments.
Chinese Translation
从LiDAR点云中对单株树木进行叶-木分割,是用于非破坏性生物量估算的定量结构模型(QSM)的关键步骤。现有分割方法通常排除辐射特征(如强度、回波次数),以最大化跨传感器兼容性。本研究通过评估跨站点和跨平台的泛化能力对这一设计选择提出质疑:在公开的Heidelberg数据集(地基TLS,1550nm)上训练,并在来自加拿大安大略省的新数据集(RPA-LS,905nm)上测试。结果表明,仅使用几何特征的方法——包括在高密度LiDAR数据集上训练的最先进深度学习模型——无法泛化到航空扫描的稀疏、自上而下的几何结构,F1分数不超过0.56。引入辐射特征(强度、回波次数、回波总数)后,F1分数提升至0.61,更关键的是,木材召回率从0.16提升至0.35,增幅达119%。此外,仅使用几何特征的方法常导致树干和树枝成分碎片化。我们发现,利用辐射特征能保持更高的结构连通性,生成更连贯的结构,从而更适合QSM重建。我们证明,尽管几何模式依赖于视角且容易过拟合扫描模式,但辐射特征编码了物理材料属性,能够在不同传感器和环境间实现泛化。
cs.CV / 70 / 2609.21938

Info3R: Information-Adaptive Test-Time Training for 3D Reconstruction

Info3R:面向三维重建的信息自适应测试时训练方法
Baek, Sunghyun, Bae, Hanna, Kwon, Minchan, Kim, Junmo
Abstract
Transformer-based models have recently achieved strong performance on 3D reconstruction from images, and recent works extend them to process video streams in an online manner for real-world deployment. However, existing methods overlook two key signals when handling long image streams: the importance of each incoming frame and the information saturation of the model's internal state. In this paper, we propose Info3R, a novel information-adaptive test-time training method for the online 3D reconstruction. We introduce an information-aware state update that modulates the state update strength based on the redundancy and informativeness of each incoming frame. To restore the state's plasticity -- its capacity to incorporate new observations -- we propose a dynamic state reset, triggered by the cumulative magnitude of state updates and the model's prediction confidence and accompanied by an anchor-to-world alignment. Our method achieves consistent improvements on camera pose estimation, video depth estimation, and 3D reconstruction, while substantially mitigating the performance degradation in the long sequence evaluation. Notably, on KITTI Odometry, our method achieves on average 1.68x lower ATE than LongStream, demonstrating its robustness on extended outdoor sequences.
Chinese Translation
基于Transformer的模型近年来在从图像进行三维重建方面取得了优异表现,近期的工作进一步将其扩展至以在线方式处理视频流,以满足实际部署需求。然而,现有方法在处理长图像序列时忽略了两个关键信号:每帧输入图像的重要性以及模型内部状态的信息饱和程度。本文提出了Info3R,一种用于在线三维重建的新型信息自适应测试时训练(test-time training)方法。我们引入了一种信息感知的状态更新机制,根据每帧输入的冗余度和信息量来调节状态更新强度。为恢复状态的可塑性——即其吸收新观测的能力——我们提出了一种动态状态重置机制,该机制由状态更新的累积幅度和模型预测置信度触发,并伴随一个锚点到世界坐标系的对齐操作。我们的方法在相机位姿估计、视频深度估计和三维重建任务上均取得了一致的性能提升,同时显著缓解了长序列评估中的性能退化问题。值得注意的是,在KITTI Odometry数据集上,我们的方法平均ATE比LongStream降低了1.68倍,展示了其在长距离户外序列上的鲁棒性。
cs.CV / 71 / 2609.22040

PRIME: Perception Feedback with Situational Memory Embeddings in VLA Models

PRIME:基于情境记忆嵌入的VLA模型感知反馈机制
Deinzer, Erik, Baslan, Naya, Paparusso, Luca, Vaskevicius, Narunas, Knott, Peter, Palmieri, Luigi
Abstract
Current Vision-Language-Action (VLA) models for autonomous driving operate primarily through feedforward inference across the perception--reasoning--planning hierarchy. While modern architectures maintain temporal recurrence within the perceptual module, early perception remains blind to downstream reasoning and navigation goals, processing visual inputs agnostically without prioritizing cues informed by prior decisions. To bridge this gap, this paper introduces PRIME, a learned feedback mechanism that conditions the VLA perceptual queries on a novel Situational Memory. By aggregating latent representations of past perception, reasoning, navigation goals, and predicted behaviors across an L-step window via cross-attention, PRIME enables intent-driven perceptual attention at minimal computational cost, adding only a maximum of 29.7M parameters (0.41% of the 7.3B-parameter base model). Evaluated on the Bench2Drive closed-loop benchmark, PRIME achieves a state-of-the-art Driving Score of 82.47 (+4.73 over ORION) and a Success Rate of 60.00% (+5.38 percentage points), the highest reported Driving Score among published VLAs trained on Think2Drive demonstrations.
Chinese Translation
当前用于自动驾驶的视觉-语言-动作(VLA)模型主要在感知—推理—规划层级中通过前馈推理运行。尽管现代架构在感知模块内部保持了时序递归,但早期感知对下游的推理和导航目标仍然是无感知的,其以目标无关的方式处理视觉输入,而无法根据先前决策所提供的信息对线索进行优先级排序。为弥合这一差距,本文提出了PRIME,一种可学习的反馈机制,它通过一种新颖的情境记忆(Situational Memory)对VLA模型的感知查询进行条件化。通过交叉注意力机制,在一个L步窗口内聚合过去感知、推理、导航目标以及预测行为的潜在表征,PRIME以极低的计算代价实现了意图驱动的感知注意力,仅增加最多2970万参数(占73亿参数基础模型的0.41%)。在Bench2Drive闭环基准测试中,PRIME取得了82.47的最先进驾驶得分(比ORION高4.73)和60.00%的成功率(提升5.38个百分点),是在基于Think2Drive示范数据训练的已发表VLA模型中最高的驾驶得分。
cs.CV / 72 / 2609.22060

Traffic Sign Recognition for Autonomous Driving Using Branched YOLOv2 and Geometric Features

基于分支YOLOv2与几何特征的自动驾驶交通标志识别
Rezaei, Arefeh
Abstract
Traffic sign recognition (TSR) is an important perception task for autonomous driving and advanced driver-assistance systems, where a system must both localize traffic signs and determine their semantic classes efficiently. This work presents a TSR system based on YOLOv2 for simultaneous detection and classification. Two complementary modifications are studied. First, YOLOv2 is extended with intermediate prediction layers, forming a branched architecture that can terminate inference early for easy cases and reduce computation time. Both whole-image and cell-wise branching strategies are investigated. Second, geometric information is introduced to reduce classification errors between visually similar signs. An unsupervised Bayesian image-segmentation method produces binary representations that are compared with class-specific geometric templates inside YOLOv2 bounding boxes. This information is used either during inference or as an additional signal during training. A dedicated dataset is constructed by combining GTSDB and GTSRB samples using seamless cloning and controlled image transformations. Experiments cover ten traffic-sign classes, with 3,000 training and 300 test samples. The selected branched architecture reports 0.647 s runtime and 0.680 mAP, compared with 0.6607 s and 0.680 mAP for baseline YOLOv2. Geometric verification during inference increases mAP to 0.713, while the geometric-feature training variant achieves 0.697 mAP with a reported runtime of 0.6608 s.
Chinese Translation
交通标志识别(TSR)是自动驾驶和高级驾驶辅助系统(ADAS)中的一项重要感知任务,系统需要高效地定位交通标志并确定其语义类别。本文提出了一种基于YOLOv2的交通标志识别系统,可同时进行检测与分类。本文研究了两种互补的改进方法。首先,通过增加中间预测层对YOLOv2进行扩展,形成一种分支架构,能够在简单情形下提前终止推理,从而减少计算时间。本文研究了整幅图像级和单元格级两种分支策略。其次,引入几何信息以减少视觉上相似标志之间的分类错误。一种无监督贝叶斯图像分割方法生成二值表示,并将其与YOLOv2边界框内特定类别的几何模板进行比较。该信息既可在推理阶段使用,也可作为训练阶段的附加信号。通过使用无缝克隆和受控图像变换将GTSDB与GTSRB样本相结合,构建了一个专用数据集。实验涵盖十类交通标志,包含3000个训练样本和300个测试样本。所选的分支架构运行时间为0.647秒,mAP为0.680,而基线YOLOv2的运行时间为0.6607秒,mAP为0.680。在推理阶段加入几何验证可将mAP提升至0.713,而采用几何特征训练的变体达到0.697的mAP,运行时间为0.6608秒。
cs.CV / 73 / 2609.22069

OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

OmniVBench:面向全参照视频生成(Omni Reference-to-Video)的基准与大规模数据集
Li, Wenxue, Guan, Peiyan, Jiang, Haoyang, Cai, Junxian, Liu, Hualuo, Zhang, Chunjie, Guan, Chong, Li, Songlian, Wu, Taiyi, Yu, Yongjian, Zhao, Xiaotong, Zhao, Alan, Liu, Eric, Chen, Xi, Liu, Yu, Zhu, Lei
Abstract
Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise to the emerging paradigm of omni R2V generation. However, existing benchmarks fall short of these emerging capabilities: their test cases cover limited reference types and compositions, and their evaluation protocols largely assess holistic reference consistency, overlooking whether reference factors are properly preserved, disentangled, and routed. Meanwhile, the high cost of constructing omni R2V training data makes suitable training resources scarce. To address these gaps, we introduce OmniVBench and the Omni-R2V Dataset for evaluating and training omni R2V models. OmniVBench expands R2V evaluation across broader reference types, fine-grained control tasks, and richer reference compositions, covering 7 task families and 18 fine-grained tasks spanning content, motion, style, structure, narrative, and multi-reference settings. We introduce factor-grounded evaluation with 12,172 case-specific checklist items, assessing whether intended reference factors are faithfully preserved, correctly disentangled and bound to their targets, and properly realized according to the instruction. We further introduce the Omni-R2V Dataset, bringing industrial-grade training resources for diverse R2V tasks to the broader research community. Drawing primarily on a large-scale corpus of professional video footage, it comprises 340K processed training samples spanning diverse reference types and multi-reference compositions. We develop task-specific pipelines for reference-target pair construction, offering a practical and scalable recipe for omni R2V data construction. Extensive evaluation of advanced open- and closed-source R2V models reveals clear performance gaps across task families and evaluation dimensions on OmniVBench, highlighting remaining limitations of current R2V models.
Chinese Translation
参照到视频(Reference-to-Video, R2V)生成正朝着日益通用和多样化的参照控制方向演进,催生了全参照(omni R2V)生成这一新兴范式。然而,现有基准难以覆盖这些新兴能力:其测试用例仅涵盖有限的参照类型与组合方式,且其评估协议大多只衡量整体参照一致性,忽视了参照因子是否被正确保留、解耦和路由。与此同时,构建全参照R2V训练数据的高昂成本使得合适的训练资源十分匮乏。为弥补这些空白,我们提出OmniVBench与Omni-R2V数据集,用于评估和训练全参照R2V模型。OmniVBench将R2V评估扩展至更广泛的参照类型、细粒度控制任务和更丰富的参照组合,涵盖内容、运动、风格、结构、叙事和多参照等7大任务族和18个细粒度任务。我们引入基于因子的评估方法,构建了12,172个针对具体案例的检查项,评估预期参照因子是否被忠实保留、是否被正确解耦并绑定到其目标、以及是否按照指令正确实现。我们进一步提出Omni-R2V数据集,为更广泛的研究社区带来工业级的多任务R2V训练资源。该数据集主要基于大规模专业视频素材语料库,包含34万个经处理的训练样本,覆盖多样的参照类型和多参照组合。我们开发了针对具体任务的参照-目标对构建流水线,为全参照R2V数据构建提供了实用且可扩展的方案。对先进开源与闭源R2V模型的广泛评估显示,模型在OmniVBench的各任务族和评估维度上存在明显的性能差距,凸显了当前R2V模型仍然存在的局限性。
cs.CV / 74 / 2609.22083

MintAct: A Unified Visual Agent for Digital Environments

MintAct:面向数字环境的统一视觉智能体
Gao, Mingfei, Tian, Rui, Gang, Haiming, Zhai, Bohan, Zhang, Le, Gong, Yuanzheng, Feng, Di, Özsoy, Ege, Ma, Kaixin, Kirthivasan, Vishwesh, Kar, Oğuzhan Fatih, Bachmann, Roman, Larsen, Anders Boesen Lindbo, Dehghan, Afshin
Abstract
We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent instances across heterogeneous per-domain backends, serving both trajectory data collection and online RL. To enable efficient and scalable RL training, an asynchronous framework keeps explicit control over the cross-domain training distribution and remains stable under noisy environment feedback and off-policy drift. Experimental results show that MintAct achieves state-of-the-art performance (48.9 on OSWorld-Verified) across a wide range of benchmarks at comparable model sizes.
Chinese Translation
我们提出了MintAct,一系列视觉-语言模型,统一了UI定位、跨移动端、桌面端和网页的多步导航以及视觉工具使用,并提供了2B、4B和8B三种参数规模。通过对环境、数据和训练方案的精心设计,MintAct模型在所有这些能力上均能达到各领域专用模型的性能水平。为实现这一目标,我们构建了可扩展的环境与强化学习(RL)基础设施。在环境方面,我们在异构的各领域后端上托管数百个并发实例,同时支持轨迹数据收集和在线强化学习。为实现高效且可扩展的强化学习训练,一个异步框架对跨域训练分布保持显式控制,并在嘈杂的环境反馈和离策略漂移下保持稳定。实验结果表明,MintAct在同等模型规模的条件下,于众多基准测试中取得了最先进的性能(在OSWorld-Verified上达到48.9分)。
机器学习 (Machine Learning)
83
cs.LG / 1 / 2609.20883

Sparse Priors for Efficient Distribution Learning

面向高效分布学习的稀疏先验
Goyal, Saumya, Póczos, Barnabás
Abstract
Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/\Theta(d)})$, though shown to be minimax optimal. We hypothesize that present bounds are too pessimistic because smoothness assumptions are not enough to capture the structure of distributions that often appear in real applications. Consequently, we introduce the class of sparse priors and define the "Sparse Dimension" as a measure of sparsity of a prior over the space of all distributions. We show that distribution learning under a $k$-sparse prior achieves a Bayesian risk lower bound of $\Omega(\sqrt{k/n})$ under common distance metrics, and show a matching (up to logarithmic terms asymptotically in $n,k$) upper bound for the TV distance under mild additional assumptions. We show the statistical equivalence of distribution learning and learning to sample in the Bayesian setting so that our results apply to learning to sample as well. While $k$ can still depend on the dimension $d$, or a notion of intrinsic dimension, our results show that learning under an appropriate prior overcomes the curse of dimensionality with respect to the dependence on $n$.
Chinese Translation
尽管生成式AI技术如今得到了广泛应用并取得了成功,但从n个样本中学习支撑于d维空间的分布,其理论保证以O(n^{-1/Θ(d)})的速度退化,尽管这已被证明是极小极大最优的。我们推测现有的界过于悲观,因为光滑性假设不足以刻画实际应用中经常出现的分布结构。为此,我们引入了稀疏先验(sparse priors)这一类别,并定义“稀疏维度”(Sparse Dimension)作为先验在全体分布空间上稀疏性的度量。我们证明,在k-稀疏先验下,分布学习在常见距离度量下可达到Ω(√(k/n))的贝叶斯风险下界,并在温和的额外假设下给出了总变差(TV)距离下(在n、k趋于无穷时至多相差对数项)相匹配的上界。我们证明了分布学习与学习采样在贝叶斯框架下的统计等价性,因此我们的结果同样适用于学习采样任务。尽管k仍可能依赖于维度d或某种内在维度的概念,我们的结果表明,在适当的先验下进行学习,能够克服关于n的依赖关系上的维数灾难。
cs.LG / 2 / 2609.20886

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

BI-Agent与BI-Bench:迈向端到端商业智能自动化
Hu, Chuxuan, He, Yeye, Zhou, Penny, Tok, Wee Hyong, Kang, Daniel, Chaudhuri, Surajit
Abstract
Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau. In traditional BI workflows, users need to prepare data by (1) identifying relevant tables, (2) performing data transformations, and (3) building join relationships, before they can (4) answer their business questions. These steps can be complex and time-consuming, making BI challenging. Given the strong capabilities of large language models (LLMs) in working with data, we study their ability to answer BI questions end-to-end, without requiring users to manually perform the tedious preparation steps. To do this, we harvest a large collection of real-world BI projects from public sources, and manually extract pairs of (questions, ground-truth answers) from real user dashboards. The resulting benchmark, BI-Bench, is the first benchmark to systematically study LLMs' ability on end-to-end BI. We find that even frontier LLMs perform poorly on BI-Bench, with less than 50% accuracy. To address their limitations, we design a tool-augmented BI-Agent that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages. Furthermore, we develop a post-training framework that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL). BI-Agent achieves substantial accuracy gains of up to 40 percentage points with vanilla LLMs, and post-trained BI-Agent yields gains of up to 30 points. Our results highlight the importance of combining tool-augmented reasoning with domain-specific post-training in complex BI workflows, and point to promising directions for future research.
Chinese Translation
商业智能(BI)是企业决策的基石,被企业用户广泛应用于Power BI和Tableau等软件中。在传统的BI工作流程中,用户需要通过以下步骤准备数据:(1)识别相关表格,(2)执行数据转换,以及(3)构建连接关系,之后才能(4)回答其业务问题。这些步骤可能复杂且耗时,使BI工作充满挑战。鉴于大语言模型(LLMs)在处理数据方面的强大能力,我们研究了它们端到端回答BI问题的能力,即无需用户手动执行繁琐的准备步骤。为此,我们从公开来源收集了大量真实的BI项目,并从真实的用户仪表板中人工提取了(问题,标准答案)对。由此产生的基准测试BI-Bench,是首个系统研究LLMs端到端BI能力的基准。我们发现,即使是前沿LLMs在BI-Bench上的表现也不佳,准确率不足50%。为解决这些局限,我们设计了一个工具增强的BI-Agent,将BI工作流分解为结构化数据上的子任务(如搜索、连接和转换),并在各BI阶段协调专门的数据管理方法。此外,我们开发了一个后训练框架,从真实BI项目中合成训练轨迹,使BI-Agent能够通过监督微调(SFT)和强化学习(RL)进行进一步的后训练。BI-Agent相较原始LLMs实现了高达40个百分点的显著准确率提升,后训练后的BI-Agent带来了高达30个百分点的提升。我们的结果凸显了在复杂BI工作流中将工具增强推理与领域特定后训练相结合的重要性,并为未来研究指明了有前景的方向。
cs.LG / 3 / 2609.20888

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

弹性阈值注意力:面向长上下文解码的可学习上下文稀疏化
Haris, Themistoklis, Li, Henry, Karimzadehgan, Maryam
Abstract
Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this via selective loading, but that comes at a cost: rigid heuristics drop necessary context, leading to quality degradation. We introduce \textbf{Elastic Threshold Attention (ETA)}, an end-to-end trainable architecture that achieves hardware-accelerated decoding speed without sacrificing dense model quality. ETA predicts dynamic, contextual thresholds directly from query representations, allowing the model to allocate dense-like context to difficult retrieval or reasoning steps while pruning routine tokens. To learn this policy from scratch without representation collapse, ETA \emph{multiplicatively suppresses} sub-threshold logits toward zero during training rather than deleting them. Training against this smooth uniform attention floor provides a distributed probability reservoir that \textbf{causes localized attention sinks on initial tokens to disappear}. It also enables the model to hard-prune uninformative KV blocks at inference time and absorb incidental tokens co-admitted by coarse GPU block selection. As a result, a 1.45B pretrained ETA model rivals dense attention across language modeling, commonsense reasoning, and long-context needle retrieval at $\approx 85\%$ training sparsity and $\approx 38\%$ active decode density. At inference time, we implement a custom decode kernel in Triton that screens KV blocks in $O(1)$ time using cached geometric-probabilistic bounds, delivering up to $2.5\times$ wall-clock decode speedups over FlashAttention-2 on sequences up to 512K tokens. Finally, we introduce an offline calibration algorithm for domain-specific deployments that freezes per-head constant thresholds to eliminate predictor overhead, cutting attention compute by an additional $27\%$.
Chinese Translation
庞大的KV缓存会在长上下文解码期间造成严重的内存带宽瓶颈。稀疏注意力方法通过选择性加载来缓解这一问题,但其代价是:僵化的启发式规则会丢弃必要的上下文,导致质量下降。我们提出了**弹性阈值注意力(Elastic Threshold Attention, ETA)**,这是一种端到端可训练的架构,能够在不牺牲稠密模型质量的前提下实现硬件加速的解码速度。ETA直接从查询表示中预测动态的、上下文相关的阈值,使模型能够为困难的检索或推理步骤分配接近稠密的上下文,同时剪除常规token。为了从零开始学习这一策略且避免表示坍塌,ETA在训练过程中对低于阈值的logit进行**乘性抑制**使其趋近于零,而非直接删除。针对这一平滑的均匀注意力下限进行训练,提供了一个分布式的概率储备,从而**使初始token上的局部注意力汇聚点(attention sink)消失**。这使得模型能够在推理时硬剪枝无信息量的KV块,并吸收因粗粒度GPU块选择而被附带选入的偶发token。由此,一个1.45B参数的预训练ETA模型在约85%的训练稀疏度和约38%的有效解码密度下,在语言建模、常识推理和长上下文大海捞针检索任务上可与稠密注意力相媲美。在推理阶段,我们用Triton实现了一个自定义解码内核,利用缓存的几何-概率界以O(1)时间筛选KV块,在长达512K token的序列上相比FlashAttention-2实现了最高2.5倍的墙钟解码加速。最后,我们提出了一种面向特定领域部署的离线校准算法,通过冻结每个注意力头的常数阈值来消除预测器的开销,使注意力计算量进一步降低27%。
cs.LG / 4 / 2609.20904

Bio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain--Computer Interfaces

Bio-MF:面向混合运动想象脑-机接口的低延迟、高保真EEG到fNIRS跨模态生成方法
Zhao, Boyuan, Zhang, Sifan, Chen, Luping
Abstract
Hybrid motor-imagery brain-computer interfaces (MI-BCIs) combining EEG and fNIRS can outperform EEG-only systems by exploiting complementary electrophysiological and hemodynamic information. To obtain such hybrid information when paired EEG-fNIRS acquisition is unavailable or inconvenient, recent studies have focused on EEG-to-fNIRS cross-modal generation. However, existing methods still suffer from slow generation and often require pretraining, limiting their use in real-time MI-BCI scenarios. Although one-step generative models offer an attractive route to low-latency synthesis, removing the iterative refinement process can reduce generation fidelity and introduce non-physiological artifacts. To address these problems, this paper proposes Bio-MF, a latent-free one-step MeanFlow framework for EEG-conditioned fNIRS generation. Bio-MF performs direct signal-space x-prediction, converts this signal-space output into MeanFlow velocity supervision, and completes inference with one network evaluation. To preserve task-relevant hemodynamic structure under heterogeneous sensor layouts, Bio-MF integrates Spatial-Temporal Interactive 4D Encoding, cross-modal classifier-free guidance, and noise-level-gated FFT regularization. On Dataset 1, EEG + synthetic fNIRS improves ACC over EEG-only by 3.37 and 4.15 percentage points for HbR and HbO, respectively. On Dataset 2, the corresponding gains remain 2.98 and 2.50 percentage points under the unseen 64-channel EEG montage. On an RTX PRO 6000 GPU, Bio-MF generates one fNIRS trial in 7.0 ms, corresponding to an 857x speedup over the 1000-step SCDM latency. These results show that Bio-MF enables fast EEG-to-fNIRS synthesis while preserving task-relevant generation quality for downstream hybrid MI decoding. Our code is available at https://github.com/psychosiwa/Bio-MF.
Chinese Translation
结合EEG与fNIRS的混合运动想象脑-机接口(MI-BCI)能够利用电生理与血流动力学信息的互补性,从而超越仅使用EEG的系统。在EEG-fNIRS同步采集不可用或不便的情况下,为获取此类混合信息,近期研究聚焦于EEG到fNIRS的跨模态生成。然而,现有方法仍存在生成速度慢、通常需要预训练等问题,限制了其在实时MI-BCI场景中的应用。尽管单步生成模型为低延迟合成提供了一条有吸引力的途径,但去除迭代细化过程可能降低生成保真度并引入非生理性伪影。为解决这些问题,本文提出Bio-MF,一个无需潜空间表示的单步MeanFlow框架,用于以EEG为条件的fNIRS生成。Bio-MF直接在信号空间执行x-prediction,将信号空间输出转换为MeanFlow速度监督,并通过单次网络评估完成推理。为在异构传感器布局下保留任务相关的血流动力学结构,Bio-MF融合了时空交互4D编码(Spatial-Temporal Interactive 4D Encoding)、跨模态无分类器引导(cross-modal classifier-free guidance)以及噪声水平门控FFT正则化。在数据集1上,EEG+合成fNIRS相较于仅使用EEG,HbR和HbO的准确率分别提升3.37和4.15个百分点。在数据集2上,在未见过的64通道EEG导联配置下,相应增益仍达2.98和2.50个百分点。在RTX PRO 6000 GPU上,Bio-MF生成一段fNIRS试次仅需7.0毫秒,相较于1000步的SCDM实现了857倍的加速。这些结果表明,Bio-MF能够在快速进行EEG到fNIRS合成的同时,保持任务相关的生成质量,以支持下游的混合MI解码。我们的代码已发布于 https://github.com/psychosiwa/Bio-MF。
cs.LG / 5 / 2609.20906

Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies

基于天体物理时序研究历史的连续延迟记忆随机梯度下降与连续时间强化学习
Paul, Debartha, Yi, Juncheng
Abstract
Quasars are luminous objects in the universe that exhibit stochastic brightness variations encoding information about the supermassive black holes powering them, and modeling these variations from ground-based survey data time series, known as light curves, is a statistical challenge. This paper reviews how stochastic differential equations (SDEs) have been adapted with neural network parameterizations to overcome this challenge in history. We create the Continuous-Delayed-Memory Stochastic Gradient Descent which depend on the past state of the discrete iteration process. We performed the simulation on some 2-dimensional landscape and observed some wider-exploration and more precise convergent behavior compared to Vanilla SGD by adjusting hyperparameters. Besides, we proposed a reinforcement learning structure with continuous time policy gradients for exploratory policies without solving HJB PDE, and we show that its optimality conditions recover the Gibbs policy of previous works.
Chinese Translation
类星体是宇宙中的明亮天体,其表现出随机亮度变化,这些变化编码了驱动它们的大质量黑洞的相关信息。基于地基巡天数据的时间序列(称为光变曲线)对这些亮度变化进行建模是一项统计挑战。本文回顾了历史上如何将随机微分方程(SDE)与神经网络参数化相结合来克服这一挑战。我们提出了依赖于离散迭代过程过去状态的连续延迟记忆随机梯度下降(Continuous-Delayed-Memory Stochastic Gradient Descent),并在一些二维景观上进行了模拟。通过调整超参数,我们观察到与普通随机梯度下降(Vanilla SGD)相比,其具有更广泛的探索和更精确的收敛行为。此外,我们提出了一种具有连续时间策略梯度的强化学习结构,无需求解HJB偏微分方程即可获得探索性策略,并证明其最优性条件可恢复先前工作中的Gibbs策略。
cs.LG / 6 / 2609.20912

Do Quantum Models Scale Like LLMs?

量子模型的缩放规律与大语言模型类似吗?
Berman, David S., Kao, Ying-Jer, Melko, Roger G., Stapleton, Alexander G.
Abstract
In this work, we study the neural scaling laws of RydbergGPT, an autoregressive transformer model trained on qubit projective measurement data gathered from interacting Rydberg atom arrays. The quantum system is known to exhibit a finite-size remnant of a critical point as the laser detuning parameter is varied. We find that near the critical point the transformer loss as a function of training dataset size is well described by a power-law with a loss floor correction. However, away from criticality the quality of the power-law description is substantially reduced. We then compare the statistical structure of both Rydberg measurements and natural-language corpora using an entropy-normalised, finite sample corrected mutual information "two-point" function. We find that near-critical statistics of the two point functions are closest to those observed in natural-language, whilst other qubit configurations far from the critical point have two-point functions that decay more rapidly. This supports the hypothesis that multi-scale dependence contributes to stable neural scaling, and that scaling behaviour should be viewed as a property of the model-data pair.
Chinese Translation
在本工作中,我们研究了RydbergGPT的神经缩放规律,RydbergGPT是一种自回归Transformer模型,其训练数据来自相互作用的里德堡(Rydberg)原子阵列的量子比特投影测量数据。已知该量子系统在激光失谐参数变化时表现出临界点的有限尺寸残余。我们发现,在临界点附近,Transformer损失作为训练数据集规模的函数可以用带有损失下限修正的幂律很好地描述。然而,远离临界区域时,幂律描述的质量显著下降。随后,我们使用熵归一化、有限样本修正的互信息“两点”函数,比较了里德堡测量数据与自然语言语料库的统计结构。我们发现,临界点附近的两点函数统计特性与自然语言中观察到的最为接近,而远离临界点的其他量子比特构型的两点函数衰减更快。这支持了以下假说:多尺度依赖性有助于形成稳定的神经缩放规律,且缩放行为应被视为模型-数据对的属性。
cs.LG / 7 / 2609.20942

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

当AI评审训练AI审稿人:科学判断的崩塌与缓解
Ho, Sy-Tuyen, Liu, Minghui, Huang, Furong
Abstract
Large language models (LLMs) increasingly participate in scientific evaluation, both as automated reviewers and as assistants to human reviewers. As model-generated reviews enter public data and future training corpora, AI peer review can become recursive: later reviewers learn from judgments produced by earlier models. We study one step of this feedback loop in a controlled setting. Starting from Llama 3.1 8B, we first fine-tune a reviewer on official ICLR reviews from 2018--2023 and then train four successor models on ICLR 2024 data with systematically varied mixtures of official and model-generated reviews. Our study shows that introducing synthetic reviews compresses rating distributions and reduces both same-paper and corpus-level semantic diversity. We call this pattern $\textbf{scientific-judgment collapse}$. To mitigate this failure mode, we introduce $\textbf{TrustReviewer}$, an open-source LLM-based system for generating peer reviews of AI and machine learning papers. TrustReviewer intervenes at two complementary stages. For training-time prevention, we train the core reviewer in a single stage on a curated corpus designed to reduce low-quality and semantically degenerate supervision. For test-time correction, paired activation steering aims to further mitigate residual tendencies toward collapsed judgments without further training or additional expert annotation. Together, these results characterize a concrete risk of recursive reviewer training and provide practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted scientific evaluation.
Chinese Translation
大语言模型(LLM)越来越多地参与科学评价,既作为自动化审稿人,也作为人类审稿人的助手。随着模型生成的评审意见进入公开数据和未来的训练语料库,AI同行评审可能变得具有递归性:后来的审稿人会从早期模型产生的判断中学习。我们在受控环境中研究了这一反馈回路中的一个环节。以 Llama 3.1 8B 为起点,我们首先在 2018--2023 年的 ICLR 官方评审上微调出一个审稿人模型,然后在 ICLR 2024 数据上训练四个后继模型,其中官方评审与模型生成评审的比例经过系统性变化。我们的研究表明,引入合成评审会使评分分布被压缩,并降低单篇论文层面和语料库层面的语义多样性。我们将这一模式称为 $ extbf{科学判断崩塌}$(scientific-judgment collapse)。为缓解这一失败模式,我们提出了 $ extbf{TrustReviewer}$,一个基于开源 LLM、用于生成人工智能与机器学习论文同行评审的系统。TrustReviewer 在两个互补的阶段进行干预。在训练时预防方面,我们在一个精心构建的语料库上以单阶段方式训练核心审稿人模型,该语料库旨在减少低质量和语义退化的监督信号。在测试时纠正方面,成对激活引导(paired activation steering)旨在无需进一步训练或额外专家标注的情况下,进一步缓解崩塌判断的残余倾向。这些结果共同刻画了递归审稿人训练的具体风险,并为在AI辅助科学评价中保持判断多样性、改善推荐一致性提供了切实可行的干预手段。
cs.LG / 8 / 2609.20954

Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications

基于时序逻辑规范的高效贝叶斯自适应强化学习
Hau, Jonathan, Abate, Alessandro
Abstract
We present a novel end-to-end model-based Reinforcement Learning (RL) algorithm for efficient policy synthesis under given Linear Temporal Logic (LTL) specifications (e.g., safety or reachability) in unknown environments. To do so, a Limit-Deterministic B{\"u}chi Automaton (LDBA) representation of the LTL task is synchronised with a Bayes-Adaptive Markov Decision Process (BAMDP) representation of the environment, which allows us to leverage an enhanced exploration-exploitation trade-off that is achieved via Bayesian RL, as opposed to traditional non-Bayesian approaches. We further propose a novel Bayes-Adaptive Monte-Carlo Planning (BAMCP) algorithm to allow for approximate Bayes-optimal strategy synthesis in the synchronised BAMDP construct. A range of finite- and infinite-horizon task experiments demonstrate the effectiveness of our approach in terms of both property satisfaction and sample efficiency, when compared to traditional model-free approaches. Additional ablation studies also successfully highlight the value of the novel BAMCP algorithm in comparison to classical BAMCP for LTL task satisfaction. Finally, we also showcase a successful application of our approach for \textit{cautious} RL, namely to reduce the number of task violations incurred during policy training.
Chinese Translation
我们提出了一种新颖的端到端基于模型的强化学习(RL)算法,用于在未知环境中在给定的线性时序逻辑(LTL)规范(如安全性或可达性)下高效地进行策略合成。为此,我们将LTL任务的极限确定性Büchi自动机(LDBA)表示与环境表示为贝叶斯自适应马尔可夫决策过程(BAMDP)进行同步,从而能够利用贝叶斯强化学习实现优于传统非贝叶斯方法的探索-利用权衡。我们进一步提出了一种新颖的贝叶斯自适应蒙特卡洛规划(BAMCP)算法,以在同步后的BAMDP结构中实现近似贝叶斯最优的策略合成。一系列有限时域和无限时域任务实验表明,与传统的无模型方法相比,我们的方法在性质满足度和样本效率方面均表现出有效性。额外的消融实验也成功凸显了新颖的BAMCP算法相较于经典BAMCP在LTL任务满足方面的价值。最后,我们还展示了该方法在"谨慎"强化学习中的成功应用,即减少策略训练过程中发生的任务违规次数。
cs.LG / 9 / 2609.20968

From Switching to Dynamic Regret: A Simple Reduction via Unbiased Random Sequences

从切换遗憾到动态遗憾:一种基于无偏随机序列的简单归约方法
Wang, Yibo, Yang, Wenhao, Yang, Sifan, Wan, Yuanyu, Zhang, Lijun
Abstract
In non-stationary online learning, dynamic regret has attracted increasing attention as a measure of how well an online learner performs against a time-varying comparator sequence. Despite considerable advances, attaining optimal bounds for strongly convex and exp-concave losses often involves intricate analysis. In this paper, we present a \textit{simple} framework that reduces dynamic regret minimization to switching regret minimization. As a result, we can derive dynamic regret bounds by using off-the-shelf algorithms with switching regret guarantees. The key idea of our reduction is to construct, for \textit{any} comparator sequence, an auxiliary random sequence that is unbiased at each round, with the controlled variance and a manageable number of switches. Combining this construction with suitable surrogate losses, we can decompose dynamic regret into the expected switching regret against the random sequence and its controlled variance. Theoretically, for strongly convex and exp-concave losses, we establish the $\widetilde{O}(T^{1/3}P_T^{2/3})$ dynamic regret bounds, where $T$ denotes the time horizon and $P_T$ denotes the path-length of the comparator sequence. Moreover, for general convex losses, the same reduction also recovers the $O(\sqrt{T(1+P_T)})$ dynamic regret bound. Notably, all our findings match the minimax optimal results for these three types of losses, highlighting the versatility of our proposed framework.
Chinese Translation
在非平稳在线学习中,动态遗憾作为一种衡量在线学习器相对于时变比较序列表现的指标,受到了越来越多的关注。尽管已取得相当大的进展,但在强凸损失和指数凹损失下获得最优界通常涉及复杂的分析。本文提出了一个将动态遗憾最小化归约为切换遗憾最小化的简单框架。由此,我们可以利用具有切换遗憾保证的现成算法来推导动态遗憾界。该归约的关键思想是,对于任意比较序列,构造一个辅助随机序列,使其在每一轮均无偏,同时具有可控的方差和可管理的切换次数。将该构造与合适的代理损失相结合,我们可以将动态遗憾分解为相对于该随机序列的期望切换遗憾及其可控方差两部分。在理论上,对于强凸损失和指数凹损失,我们建立了 $\widetilde{O}(T^{1/3}P_T^{2/3})$ 的动态遗憾界,其中 $T$ 表示时间范围,$P_T$ 表示比较序列的路径长度。此外,对于一般凸损失,同样的归约方法也能恢复 $O(\sqrt{T(1+P_T)})$ 的动态遗憾界。值得注意的是,我们的所有结果均与这三类损失的极小化极大最优结果相匹配,凸显了所提框架的通用性。
cs.LG / 10 / 2609.20978

Generative inversion for early ranking of competing geologic interpretations

用于对相互竞争的地质解释进行早期排序的生成式反演方法
Rashid, Harun Ur, O'Malley, Daniel
Abstract
High-consequence subsurface decisions are often made under severe data scarcity. Experts may arrive at competing interpretations of the same subsurface system, yet early in a project there is rarely a practical way to determine which one is most realistic. This uncertainty can persist until several wells are drilled, often costing millions of dollars. Existing approaches for evaluating geologic interpretations rely either on subjective judgment or on dense data that are rarely available in early-stage investigations. We present a workflow that addresses this challenge by translating competing geologic interpretations into alternative spatial priors and ranking them according to their consistency with hydraulic-head observations. For each interpretation, a text-to-image foundation model generates an ensemble of 1600 geologic images, and a separately trained variational autoencoder provides an interpretation-specific latent representation. A supervised inverse network maps the head observations into this latent space, and the frozen decoder produces an image that is mapped to a log-conductivity field. Steady-state flow simulation then provides predicted heads, and the resulting mismatch is converted into a Gaussian-form compatibility score. We evaluate the framework using a synthetic benchmark based on the Johansen Formation and three interpretations of decreasing consistency with the reference representation. Across 925 test cases, the mean head RMSE increases from 0.197 for the Precise \& Accurate interpretation to 0.227 for the Accurate interpretation and 0.280 for the Mismatched interpretation. We subsequently apply the workflow to two published conceptual models of the Culebra Dolomite Member at the Waste Isolation Pilot Plant. The revised model receives a compatibility weight of 0.991, compared with 0.009 for the original model, consistent with the independent evidence.
Chinese Translation
高风险的地下决策往往在数据严重匮乏的条件下做出。专家们可能对同一地下系统得出相互竞争的不同解释,但在项目早期,很少有实用的方法来判断哪一种解释最为符合实际。这种不确定性可能持续存在,直到钻探数口井之后,而这往往耗费数百万美元。现有的评估地质解释的方法要么依赖主观判断,要么依赖早期勘察阶段难以获得的密集数据。我们提出了一套工作流程来应对这一挑战:将相互竞争的地质解释转化为不同的空间先验,并根据它们与水头观测数据的一致性对其进行排序。对于每种解释,文本生成图像(text-to-image)基础模型生成1600张地质图像的集合,并由单独训练的变分自编码器(variational autoencoder)提供针对该解释的潜在表示。一个有监督的反演网络将水头观测数据映射到该潜在空间,冻结的解码器生成图像并进一步映射为对数渗透率场。随后通过稳态流模拟得到预测水头,并将由此产生的偏差转换为高斯形式的相容性评分。我们基于Johansen地层构建的合成基准对框架进行评估,其中包含三种与参考模型一致性递减的解释。在925个测试案例中,平均水头均方根误差(RMSE)从"Precise & Accurate"解释的0.197升至"Accurate"解释的0.227,以及"Mismatched"解释的0.280。随后,我们将该工作流程应用于废物隔离中试工厂(Waste Isolation Pilot Plant)Culebra白云岩段(Culebra Dolomite Member)的两个已发表概念模型。修订后的模型获得的相容性权重为0.991,而原始模型仅为0.009,这与独立证据相一致。
cs.LG / 11 / 2609.20982

ASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning

ASGARD:基于强化学习的无人机动作空间弹性防护机制
Salehi, Mohsen, Pattabiraman, Karthik
Abstract
Reinforcement learning (RL) controllers have been recently adopted for Unmanned Aerial Vehicles (UAV) navigation and control. However, they are susceptible to action-space attacks that overwrite the action commands after the policy generates them and before the actuators execute them. While most existing defenses target attacks on the policy's inputs, those addressing action-space attacks retrain the policy at training time and are not resilient to corrupted actions at runtime. We propose ASGARD, a two-phase teacher-student pipeline for making RL-based UAV control resilient to action-space attacks. In the teacher phase, an encoder combines the UAV's physical state with action-attack-related privileged information to produce an action-attack-aware latent that trains the RL control policy and a monitor that outputs corrected action commands to the actuators. In the student phase, both the encoder and the monitor are trained via supervised learning from their teacher counterparts to run on-board using only the UAV's physical state history. We evaluate ASGARD across attack scenarios targeting different action commands on UAV. We find that ASGARD is resilient to action-space attacks and completes the missions despite the attack. We further find that ASGARD generalizes to unseen attacks and remains resilient against stealthy attacks.
Chinese Translation
近年来,强化学习(RL)控制器已被应用于无人机(UAV)的导航与控制。然而,它们容易受到动作空间攻击的影响,即攻击者在策略生成动作指令之后、执行器执行之前篡改该动作指令。现有的大多数防御方法主要针对策略输入端的攻击,而针对动作空间攻击的防御方法则需要在训练阶段对策略进行重训练,无法在运行时对被篡改的动作保持韧性。我们提出了ASGARD,一个两阶段的师生(teacher-student)流水线,使基于强化学习的无人机控制对动作空间攻击具有韧性。在教师阶段,编码器将无人机的物理状态与动作攻击相关的特权信息相结合,生成一个感知动作攻击的潜在表示,用于训练强化学习控制策略以及一个监控器,该监控器向执行器输出修正后的动作指令。在学生阶段,编码器和监控器通过监督学习从各自的教师模型中学习,仅需利用无人机的物理状态历史即可在机载设备上运行。我们在针对无人机不同动作指令的多种攻击场景下对ASGARD进行了评估。结果表明,ASGARD能够抵御动作空间攻击,并在攻击存在的情况下仍能完成任务。我们还发现,ASGARD能够泛化到未见过的攻击,并对隐蔽攻击保持韧性。
cs.LG / 12 / 2609.20991

From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities

从压力到情感:跨可穿戴传感器模态的生理情绪识别多模态深度学习方法
Hagos, Desta Haileselassie, Aryal, Saurav Keshari, Burge, Legand L.
Abstract
Physiological emotion recognition using wearable sensors has important applications in mental health monitoring, affective computing, and human-computer interaction. However, existing studies typically evaluate a single model, sensing configuration, or dataset, limiting our understanding of how these factors influence recognition performance. We present a comparative study of temporal deep learning architectures for physiological emotion recognition using two multimodal wearable datasets: WESAD and EmoWear. Bidirectional long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer models are evaluated under wrist-only, chest-only, and multimodal sensing configurations using participant-independent leave-one-subject-out cross-validation (LOSO-CV). We also investigate soft-voting ensembles, sensor ablation, sampling frequency, and gradient-based saliency. The Transformer achieved the highest multimodal accuracy on WESAD (99.02% +/- 0.51%), whereas the LSTM achieved the best multimodal accuracy on EmoWear for both arousal (91.80% +/- 1.06%) and valence (89.96% +/- 0.36%). These results show that relative architecture performance depends on dataset characteristics rather than one architecture being uniformly superior. Multimodal sensing consistently outperformed wrist-only and chest-only configurations across both datasets. Sampling-frequency analysis showed that 4 Hz provides a practical operating point, with performance comparable to higher frequencies at substantially lower training cost. These findings provide guidance for selecting architectures, sensing modalities, and sampling frequencies for wearable physiological emotion recognition.
Chinese Translation
基于可穿戴传感器的生理情绪识别在心理健康监测、情感计算和人机交互中具有重要应用。然而,现有研究通常仅评估单一模型、传感配置或数据集,限制了我们对这些因素如何影响识别性能的理解。本文使用两个多模态可穿戴数据集(WESAD 和 EmoWear),对用于生理情绪识别的时序深度学习架构进行了比较研究。我们在仅手腕、仅胸部和多模态三种传感配置下,采用与被试无关的留一被试交叉验证(LOSO-CV),评估了双向长短期记忆网络(LSTM)、时序卷积网络(TCN)和 Transformer 模型。我们还研究了软投票集成、传感器消融、采样频率以及基于梯度的显著性分析。Transformer 在 WESAD 上取得了最高的多模态准确率(99.02% ± 0.51%),而 LSTM 在 EmoWear 上的唤醒度(91.80% ± 1.06%)和效价(89.96% ± 0.36%)均取得了最佳多模态准确率。这些结果表明,不同架构的相对性能取决于数据集特征,而非某一架构具有普遍优势。在两个数据集上,多模态传感均始终优于仅手腕和仅胸部配置。采样频率分析表明,4 Hz 提供了一个实用的运行点,其性能与更高频率相当,而训练成本却大幅降低。这些发现为可穿戴生理情绪识别中选择架构、传感模态和采样频率提供了指导。
cs.LG / 13 / 2609.20997

MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery

MOSAIC-SR:基于Transformer引导的符号回归用于科学方程恢复
Zheng, Peiyi, Kang, Yanming, De Sterck, Hans, Tran, Giang
Abstract
Symbolic regression aims to recover closed-form equations from observations, providing interpretable models for scientific discovery. Existing approaches struggle to combine flexible structural search with efficient inference. Search-based methods can refine expression structure but often rely on costly combinatorial optimization with random initialization. Pretrained neural models generate formulas almost instantly, but their predictions often contain symbolic errors. We introduce MOSAIC-SR, which uses a pretrained Transformer to propose multiple initial sketches. These sketches initialize searches in several promising regions, avoiding random starts in the vast expression space. Each search jointly recovers structure and constants through scale-aware constant optimization and local symbolic repair. We evaluate MOSAIC-SR on the SRSD-Feynman dataset with and without dummy variables and on six additional benchmarks. MOSAIC-SR obtains the highest symbolic solution rate on every dataset while ranking among the top two methods in predictive accuracy. This advantage persists in the presence of irrelevant dummy inputs. The results show that learned priors can focus search on promising equation structures, and that numerical optimization and symbolic repair are important for recovery.
Chinese Translation
符号回归旨在从观测数据中恢复闭式方程,为科学发现提供可解释的模型。现有方法难以将灵活的结构搜索与高效的推断相结合。基于搜索的方法可以细化表达式结构,但通常依赖于随机初始化的代价高昂的组合优化。预训练神经模型几乎可以即时生成公式,但其预测往往包含符号错误。我们提出了MOSAIC-SR,该方法利用预训练的Transformer提出多个初始草图。这些草图在多个有希望的区域初始化搜索,避免了在浩瀚表达式空间中的随机起点。每次搜索通过尺度感知的常数优化和局部符号修复,联合恢复表达式结构与常数。我们在SRSD-Feynman数据集(含和不含虚拟变量)以及六个额外的基准数据集上评估了MOSAIC-SR。MOSAIC-SR在每个数据集上均获得了最高的符号解率,同时在预测精度方面位居前两名。在存在无关虚拟输入的情况下,这一优势依然保持。结果表明,学习到的先验知识可以将搜索聚焦于有希望的方程结构,而数值优化和符号修复对恢复至关重要。
cs.LG / 14 / 2609.21001

On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation

论最大编码率降低在分布外泛化中的局限性
Zhou, Menghui, Bi, Gaoshan, Lanfranchi, Vitaveska, Yang, Po
Abstract
Substantial efforts have been devoted to making deep learning objectives, representations, and architectures interpretable, with the goal of improving the safety, robustness, and generalisation of learning systems in diverse real-world applications. The recently proposed maximal coding rate reduction ($\mathrm{MCR}^{2}$) offers a promising information-theoretic framework for learning structured, discriminative representations of class-wise submanifolds and has inspired interpretable white-box architectures. However, we observe that $\mathrm{MCR}^{2}$ can completely fail under distribution shift, motivating our study of its out-of-distribution (OOD) generalisation limits. We establish two limitations of $\mathrm{MCR}^{2}$ for OOD generalisation. First, the $\mathrm{MCR}^{2}$ objective alone can admit complete prediction failure: a representation based entirely on unstable environmental features can achieve the global coding optimum yet fail completely after correlation reversal, despite an available perfectly stable feature. This exact-optimum example includes test inputs that cannot occur during training. Even when every possible test input can also occur during training, coding quality can be arbitrarily close to optimal while prediction error is arbitrarily close to 100%. Second, directly incorporating the invariance principle underlying widely successful invariant risk minimisation (IRM) and risk extrapolation (REx) does not eliminate this failure. The failing representation admits the same optimal coding operator across training environments, showing that shared coding optimality does not ensure stable prediction. Reliable OOD guarantees for $\mathrm{MCR}^{2}$ therefore require additional new assumptions or learning principles that establish stable predictive relationships across environments.
Chinese Translation
大量研究致力于使深度学习的目标函数、表示和架构具有可解释性,以提升学习系统在多样真实应用中的安全性、鲁棒性与泛化能力。最近提出的最大编码率降低(maximal coding rate reduction, $\mathrm{MCR}^{2}$)提供了一个富有前景的信息论框架,可用于学习类别子流形的结构化、判别性表示,并启发了可解释的白盒架构。然而,我们观察到 $\mathrm{MCR}^{2}$ 在分布偏移下可能完全失效,这促使我们研究其分布外(OOD)泛化的局限性。我们建立了 $\mathrm{MCR}^{2}$ 在 OOD 泛化方面的两个局限。第一,仅凭 $\mathrm{MCR}^{2}$ 目标就可能导致完全的预测失败:一个完全基于不稳定环境特征的表示可以达到全局编码最优,却在此相关性反转后完全失效,即使存在完全稳定且可用的特征。这个达到精确最优的例子包含了训练中不可能出现的测试输入。即使每一个可能的测试输入在训练中都可能出现,编码质量也可以任意接近最优,而预测误差却任意接近100%。第二,直接引入被广泛成功的不变风险最小化(IRM)和风险外推(REx)所基于的不变性原理,并不能消除这一失效。该失效表示在所有训练环境中都拥有相同的最优编码算子,这表明共享的编码最优性并不能保证稳定的预测。因此,要为 $\mathrm{MCR}^{2}$ 提供可靠的 OOD 保证,需要引入额外的新假设或能够建立跨环境稳定预测关系的学习原则。
cs.LG / 15 / 2609.21032

Scaling Discovery through Test-Time Communication

通过测试时通信扩展科学发现
Park, Jongho, Kontonis, Vasilis, Garg, Shivam, Krishnamurthy, Akshay, Papailiopoulos, Dimitris
Abstract
Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this. Whether communicating agents help remains an open question with mixed prior results. We show that test-time communication can substantially outperform independent parallel attempts on challenging tasks, where sharing a breakthrough can push the whole group forward. We first study the effect of scaling multi-agent test-time communication, where agents have no predefined roles and communicate via a shared directory, on ARC-AGI-3, a benchmark requiring novel problem solving. We find that a team of $k$ communicating agents, team@$k$, matches the success rate of $4k$ independent agents, and this advantage grows with $k$, suggesting gains compound with scale. The effect is not merely efficiency: a task that no single agent can solve, a team of agents can solve reliably. Furthermore, these gains transfer to research-oriented tasks, given sufficient compute. On polyomino packing, communicating agents outperform best@$k$ and exceed the prior best-known score. On MNIST classifier compression, communication surpasses the best-known human solution. A team of four agents produced a 1,957-byte classifier submission achieving 99.4% test accuracy, smaller than both the best-known human solution and the best single-agent result. These gains are not unconditional. Independent agents may outperform communication when compute is limited or when a clear measure of progress is absent. However, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.
Chinese Translation
科学的进步并非孤立完成,而是通过协作实现,然而现有的智能体系统却很少体现这一点。智能体之间的通信是否有帮助仍是一个悬而未决的问题,此前的研究结果也不尽一致。我们证明,在具有挑战性的任务上,测试时通信可以显著优于独立的并行尝试,因为分享一项突破可以推动整个团队共同前进。我们首先研究了扩展多智能体测试时通信的效果,其中智能体没有预定义角色,并通过共享目录进行通信,实验在 ARC-AGI-3 这一需要新型问题求解能力的基准上进行。我们发现,由 $k$ 个通信智能体组成的团队(team@$k$)能够达到 $4k$ 个独立智能体的成功率,且这一优势随 $k$ 的增长而扩大,表明其收益会随规模叠加。这种效果不仅仅是效率上的提升:单个智能体无法解决的任务,一个智能体团队可以可靠地解决。此外,在计算资源充足的情况下,这些收益可以迁移到面向科研的任务上。在多联方块装箱任务中,通信智能体优于 best@$k$ 策略,并超越了此前已知的最佳成绩。在 MNIST 分类器压缩任务中,通信智能体超越了已知最佳的人类解决方案。由四个智能体组成的团队产生了一个 1,957 字节的分类器提交结果,测试准确率达到 99.4%,其规模既小于已知最佳的人类解决方案,也小于最佳的单智能体结果。这些收益并非无条件成立:当计算资源有限或缺乏明确的进展衡量标准时,独立智能体可能优于通信方式。然而,在计算资源充足且反馈明确的情况下,多智能体通信始终能带来更强的结果。
cs.LG / 16 / 2609.21039

Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks

Stiefel-AdamW:面向线性分解模块的几何感知AdamW优化器
Zangrando, Emanuele, Sutti, Marco, Tudisco, Francesco
Abstract
A pervasive structural pattern in modern deep learning is the linear factorization block: a submodule of the form $W = BA$ in which two parameter matrices are multiplied directly, with no intervening nonlinearity. Such blocks appear in LoRA adapters, low-rank compressed layers, query-key products of self-attention, and share a common pathology: the factorization is non-unique, which can destabilize training and limit usable learning rates. Despite this, factorization blocks are typically optimized with standard Euclidean methods that ignore the underlying geometry. We introduce Stiefel-AdamW, a near drop-in replacement for AdamW for use wherever such blocks appear. By constraining one factor on the Stiefel manifold while leaving the other Euclidean, Stiefel-AdamW relaxes the full $\mathrm{GL}(\mathbb{R}^r)$ gauge symmetry to a compact orthogonal symmetry, ruling out factor blow-up while retaining the coordinate-wise diagonal preconditioning that gives AdamW its practical strength. Moment estimation is performed in the ambient Euclidean space, with geometry entering only through a tangent-space projection and a manifold retraction. The implementation overhead over AdamW is minimal, and we show that the resulting optimizer inherits both the stability benefits of Riemannian methods and standard convergence guarantees. We validate Stiefel-AdamW on LoRA-style fine-tuning of GPT2, ViT, and Mistral 7B and on full pretraining of GPT2 on OpenWebText, showing consistent improvements over strong baselines at essentially no additional cost over AdamW.
Chinese Translation
现代深度学习中一种普遍存在的结构模式是线性分解模块:形如 $W = BA$ 的子模块,其中两个参数矩阵直接相乘,中间无非线性层。此类模块出现在LoRA适配器、低秩压缩层以及自注意力的查询-键乘积中,它们具有一个共同的病理特征:分解不唯一,这会导致训练不稳定并限制可用的学习率。尽管如此,分解模块通常仍使用忽略其底层几何结构的标准欧几里得方法进行优化。我们提出了Stiefel-AdamW,一种可几乎即插即用替换AdamW的优化器,适用于所有出现此类模块的场合。通过将其中一个因子约束在Stiefel流形上,同时保留另一个因子的欧几里得性质,Stiefel-AdamW将完整的 $\mathrm{GL}(\mathbb{R}^r)$ 规范对称性松弛为紧致的正交对称性,从而排除了因子爆炸问题,同时保留了赋予AdamW实用优势的逐坐标对角预条件机制。矩估计在外围欧几里得空间中进行,几何结构仅通过切空间投影和流形收缩(retraction)介入。相比AdamW的实现开销极小,且我们证明该优化器同时继承了黎曼优化方法的稳定性优势与标准收敛性保证。我们在GPT2、ViT和Mistral 7B的LoRA式微调以及GPT2在OpenWebText上的完整预训练上验证了Stiefel-AdamW,结果显示其在几乎不增加AdamW额外成本的情况下,相比强基线取得了一致的性能提升。
cs.LG / 17 / 2609.21044

A Lightweight Plug-in Gate for Transformer-Based Time-Series Forecasters

一种面向基于Transformer时间序列预测模型的轻量级即插即用门控机制
Zhuang, Hongkai, Huang, Tao, Hou, Chen
Abstract
Covariate-rich time-series forecasting requires deciding how external variables enter the target forecasting path. Existing Transformer-based forecasters usually build a covariate representation and pass it to the encoder without an explicit admission stage. This paper studies pre-encoder covariate admission as an input-side interface that regulates that representation immediately before encoder processing. We implement the interface with a lightweight representation-level pre-encoder gate that assigns sigmoid scores to representation units, and we also study a usage-regularized variant that penalizes average admission. The interface is evaluated as a plug-in module for TimeXer, Inverted Transformer (iTransformer), and Patch Time Series Transformer (PatchTST) under a zero-extra-tuning protocol, where each gated model inherits the corresponding baseline configuration. Experiments on the Electricity Transformer Temperature minute-level (ETTm1 and ETTm2) datasets, Traffic, Energy, and influenza-like illness (ILI) include paired forecasting comparisons, gate-placement ablation, initialization ablation, controlled covariate-admission analysis, and a variance inflation factor (VIF)-informed permutation feature importance (PFI) diagnostic case study. In the tested settings, the gate is competitive with the corresponding baselines, and the usage penalty reduces average admission scores while keeping forecasting errors close to the unpenalized TimeXer setting.
Chinese Translation
含丰富协变量的时间序列预测需要决定外部变量如何进入目标预测路径。现有的基于Transformer的预测模型通常先构建协变量表示,然后直接将其传递给编码器,而缺乏显式的准入阶段。本文研究了编码器前协变量准入问题,将其视为一种输入侧接口,在编码器处理之前对该表示进行调控。我们通过一个轻量级的表示级编码器前门控机制实现该接口,该门控为表示单元分配sigmoid分数;同时我们还研究了使用正则化变体,对平均准入情况进行惩罚。该接口在零额外调参协议下作为即插即用模块在TimeXer、Inverted Transformer(iTransformer)和Patch Time Series Transformer(PatchTST)上进行评估,其中每个带门控的模型均继承对应基线的配置。实验在电力变压器温度分钟级数据集(ETTm1和ETTm2)、Traffic、Energy以及流感样疾病(ILI)数据集上进行,包括成对预测比较、门控位置消融、初始化消融、受控协变量准入分析,以及基于方差膨胀因子(VIF)的置换特征重要性(PFI)诊断案例研究。在所测试的设置中,该门控机制与相应基线相比具有竞争力,且使用惩罚在保持预测误差接近未正则化TimeXer设置的同时,降低了平均准入分数。
cs.LG / 18 / 2609.21057

FedeRage: Provably Convergent Agnostic Federated Learning under General Client Drift

FedeRage:一般客户端漂移下可证明收敛的不可知联邦学习
Rahimi, Herlock, Kalogerias, Dionysis
Abstract
Federated learning (FL) enables collaborative model training without sharing raw data, but its performance degrades under non-IID data and stochastic client participation. Remedies built on classical Federated Averaging (FedAvg) typically presuppose that client participation probabilities are known to the server, which is rarely the case in deployed systems. We first discuss and then characterize the optimization problem that \emph{distributionally agnostic} FedAvg actually solves when participation is entirely unknown, possibly highly skewed, and of variable size across rounds: uniform aggregation is shown to minimize a well-defined stochastic objective, weighted by the participation-induced marginal, at a standard $\mathcal{O}(1/\sqrt{T})$ rate for convex and possibly nonsmooth losses. Building on this characterization, we propose \emph{Federated Risk-Averse Averaging} (\textsc{FedeRage}), a risk-averse extension of FedAvg that embeds the \emph{Conditional Value-at-Risk} (CVaR) into the local objective within a natural distributionally robust optimization (DRO) framework. \textsc{FedeRage} implicitly upweights high-loss and infrequently participating clients while adding only a \emph{single scalar per-client}, and admits an $\mathcal{O}(\kappa/\sqrt{T})$ rate in which the factor $\kappa$ is the upper bound on the ``price" of risk aversion. In contrast with aggregation-alignment schemes based on optimal transport, which require the availability distribution as an input, \textsc{FedeRage} remains agnostic to it. Several experiments on three heterogeneous benchmarks indicate consistent improvements over state-of-the-art methods in accuracy, fairness, and convergence speed.
Chinese Translation
联邦学习(FL)能够在不共享原始数据的情况下实现协同模型训练,但其性能在非独立同分布(non-IID)数据和随机客户端参与的情况下会下降。基于经典联邦平均(Federated Averaging, FedAvg)构建的补救方法通常假设服务器已知客户端的参与概率,而这在实际部署系统中很少成立。我们首先讨论并刻画了当客户端参与完全未知、可能高度偏斜且各轮参与规模可变时,\emph{分布不可知}(distributionally agnostic)的 FedAvg 实际求解的优化问题:我们证明均匀聚合最小化一个定义明确的随机目标函数,该目标由参与引起的边际分布加权,并以标准的 $\mathcal{O}(1/\sqrt{T})$ 速率收敛,适用于凸的且可能非光滑的损失函数。基于这一刻画,我们提出\emph{联邦风险规避平均}(\textsc{FedeRage}),这是 FedAvg 的一种风险规避扩展方法,它在自然的分布鲁棒优化(DRO)框架内将\emph{条件风险价值}(Conditional Value-at-Risk, CVaR)嵌入局部目标中。\textsc{FedeRage} 隐式地提高高损失客户端和低频参与客户端的权重,同时每个客户端仅增加\emph{一个标量},并具有 $\mathcal{O}(\kappa/\sqrt{T})$ 的收敛速率,其中因子 $\kappa$ 是风险规避“代价”的上界。与基于最优传输的聚合对齐方案不同,后者需要将参与分布作为输入,而 \textsc{FedeRage} 对其保持不可知。在三个异构基准上的多项实验表明,该方法在准确率、公平性和收敛速度方面均较现有最先进方法取得了一致的提升。
cs.LG / 19 / 2609.21073

Toward individual-level calibration in affect recognition with perceptual adjustment queries

基于感知调整查询的情绪识别个体水平校准研究
Chen, Xuanzhou, Alagapan, Sankaraleengam, Pananjady, Ashwin
Abstract
Behavioral tasks measuring facial affect perception assume that identical stimuli impose equivalent perceptual difficulty across participants. However, this assumption is systematically violated by individual differences in perceptual sensitivity. Using an affective perception task as our testbed, we propose a framework to normalize for perceptual difficulty that directly estimates each participant's Just Noticeable Difference (JND) along the facial affect spectrum via cognitively lightweight perceptual adjustment queries (PAQs). We use these PAQ-inferred JNDs to re-express stimulus distances, constructing difficulty-equated tasks in perceptual space. We validate the framework in a Two-Alternative Forced-Choice (2AFC) task using two complementary behavioral measures: binary metacognitive difficulty judgments and response time variance decomposition. We find that PAQ calibration significantly equalizes perceived task difficulty at an individual level when compared to both the non-calibrated baseline and population-level Weibull calibration, while also reducing mean response time and between-subject variance in response time. These results establish PAQ as a principled and practical instrument for individualized perceptual calibration in facial affect recognition.
Chinese Translation
测量面部情绪感知的行为任务通常假设相同刺激对所有被试具有同等的感知难度。然而,感知敏感性的个体差异会系统性地违背这一假设。我们以情绪感知任务为测试平台,提出一个对感知难度进行归一化的框架,该框架通过认知负荷较低的感知调整查询(Perceptual Adjustment Queries,PAQ)直接估计每个被试在面部情绪谱上的最小可觉差(Just Noticeable Difference,JND)。我们利用由PAQ推断的JND重新表述刺激距离,从而在感知空间中构建难度等同的任务。我们在二选一强迫选择(Two-Alternative Forced-Choice,2AFC)任务中,采用两种互补的行为测量指标对该框架进行验证:二元元认知难度判断和反应时方差分解。结果表明,与未校准基线以及群体水平的Weibull校准相比,PAQ校准在个体水平上显著均等化了任务感知难度,同时还降低了平均反应时以及被试间反应时方差。这些结果确立了PAQ作为面部情绪识别中个性化感知校准的一种具有理论依据且切实可行的工具。
cs.LG / 20 / 2609.21108

REFINEPPO: Learning Continuous Control Policies by Iterative Action Refinement

REFINEPPO:通过迭代动作精化学习连续控制策略
Weerasekara, Sachini, Kamarthi, Sagar, Isaacs, Jacqueline
Abstract
Deep reinforcement learning (DRL) has achieved strong performance across a wide range of continuous-control problems. These continuous-control policies, however, are often defined as direct mappings from an observed state to an action or action distribution, requiring a single feed-forward network to construct an optimal control decision in one pass. While effective, this formulation leaves little opportunity for the policy to reconsider or progressively improve an action once an initial prediction has been formed. In this work, we explore an alternative approach: rather than learning only to directly predict an action, can a policy learn to iteratively improve one, and can this iterative process provide advantages during policy learning? We introduce Iterative Action Refinement (IAR), an iterative action-construction method that constructs control actions through a sequence of learned residual corrections. Starting from an initial proposal, a shared refinement network repeatedly conditions on the observed state and the current action proposal, allowing each refinement step to revise the action constructed by preceding steps. The final refined proposal is then used to determine the action executed by the agent. We integrate this iterative action-construction mechanism with Proximal Policy Optimization (PPO), yielding REFINEPPO. We evaluate REFINEPPO across 14 benchmark control tasks, complemented by controlled ablations of refinement depth and update schedules and analyses aimed at understanding why iterative refinement is effective. Across these environments, REFINEPPO matches or exceeds the performance of standard PPO while demonstrating faster convergence on several tasks.
Chinese Translation
深度强化学习(DRL)在众多连续控制问题上取得了优异的性能。然而,这些连续控制策略通常被定义为从观测状态到动作或动作分布的直接映射,需要单个前馈网络一次性构建最优控制决策。这种方式虽然有效,但一旦形成初始预测,策略几乎没有机会重新考虑或逐步改进动作。在本工作中,我们探索了一种替代方法:策略能否不仅学习直接预测动作,而是学习迭代地改进动作?这一迭代过程能否在策略学习中带来优势?我们提出了迭代动作精化(Iterative Action Refinement, IAR),这是一种通过一系列学习到的残差修正来构建控制动作的迭代式动作构建方法。从初始提议开始,一个共享的精化网络反复以观测状态和当前动作提议为条件进行输入,使每个精化步骤能够修正由前序步骤构建的动作。最终精化后的提议被用于确定智能体执行的动作。我们将这种迭代式动作构建机制与近端策略优化(Proximal Policy Optimization, PPO)相结合,得到了REFINEPPO。我们在14个基准控制任务上对REFINEPPO进行了评估,并辅以针对精化深度与更新策略的受控消融实验,以及旨在理解迭代精化为何有效的研究分析。在所有这些环境中,REFINEPPO均达到或超越了标准PPO的性能,同时在若干任务上表现出更快的收敛速度。
cs.LG / 21 / 2609.21109

Talk to Me, Jarvis: An Open-Source Edge-Deployable Voice Assistant Framework for Autonomous Racecars

与我对话,Jarvis:一个面向自动驾驶赛车的开源边缘部署语音助手框架
Henel, Daniel, Werner, Frederik, Langmann, Alexander, Betz, Johannes
Abstract
Recent advances in large language models have improved their effectiveness as back-end components for voice assistants, particularly in intent understanding and context-aware input classification. However, online-hosted models introduce network dependency and variable inference latency, limiting their suitability for time-critical autonomous driving applications. In this work, we address these issues by developing Jarvis, an offline voice assistant for high-level behavioral commands of autonomous vehicles. Its architecture integrates speech recognition and synthesis with natural language command classification into a lightweight, local framework. Jarvis core component is a text-to-command classifier, built using a domain-specific fine-tuning of the Mistral 7B model, demonstrating low-latency inference. Our experimental evaluation demonstrates that our solution outperforms larger online-hosted models, achieving 97.63 % intent recognition accuracy with an average processing latency of 1.39 s, making it well-suited for operations requiring quick response times. To support further research and fine-tuning, we provide an open-source implementation.
Chinese Translation
大语言模型的最新进展提升了其作为语音助手后端组件的有效性,尤其是在意图理解和上下文感知输入分类方面。然而,在线托管模型引入了网络依赖性和不可变的推理延迟,限制了其在时间敏感的自动驾驶应用中的适用性。在本工作中,我们通过开发Jarvis来解决这些问题,Jarvis是一个面向自动驾驶车辆高级行为指令的离线语音助手。其架构将语音识别与合成以及自然语言指令分类集成到一个轻量级的本地框架中。Jarvis的核心组件是一个文本到指令的分类器,该分类器基于对Mistral 7B模型的领域特定微调构建,展现出低延迟推理能力。我们的实验评估表明,本方案优于更大的在线托管模型,实现了97.63%的意图识别准确率和1.39秒的平均处理延迟,使其非常适用于需要快速响应时间的操作。为支持进一步的研究和微调,我们提供了开源实现。
cs.LG / 22 / 2609.21123

Signal-Centric Remote Sensing via Alternative Preprocessing and Acoustic Processing for ML-Driven Applications

基于替代预处理与声学处理的信号中心遥感方法及其机器学习驱动应用
Luna, Logan, Jansen-Sánchez, Sirio, Demirkiran, Ilteris, Ghelarducci, Leo
Abstract
The dominant method of processing sonar data is using image-based representations, requiring the preprocessing of image data on autonomous systems. We propose an alternative data processing method for remote sensing applications via the use of data in Comma-Seperated Value format. Experimentation on our alternative approach shows a reduction of processing time by 91.18%, an improvement in accurate object detection by Machine Learning, and an increase in SNR (Signal-to-noise ratio), PSNR (Peak signal-to-noise ratio), and other evaluation metrics.
Chinese Translation
处理声呐数据的主流方法是使用基于图像的表示形式,这要求在自主系统上对图像数据进行预处理。我们提出了一种替代性的数据处理方法,通过使用逗号分隔值(CSV)格式的数据来服务于遥感应用。对该替代方法的实验表明,处理时间减少了91.18%,机器学习对目标的准确检测能力有所提升,同时信噪比(SNR)、峰值信噪比(PSNR)及其他评估指标均有所提高。
cs.LG / 23 / 2609.21126

Layerwise Decoupling for Stable Structured Sparsification of Fully Connected Layers

面向全连接层稳定结构化稀疏化的逐层解耦方法
Kulick, Charles, Petrosyan, Armenak, Tang, Sui
Abstract
We propose a decoupled, layerwise method for structurally sparsifying the fully connected layers of pretrained neural networks. Rather than penalizing all layers jointly, our approach extracts shallow two-layer subnetworks, normalizes the inner weights, and applies a structured group penalty to the outer weight matrix of each block, processing layers sequentially to prune neurons and reduce the width of each layer. We prove that the constrained decoupled objective is equivalent at optimality to a specific joint penalty on the inner and outer weights, for any positively homogeneous activation, and thus admits a clean projected and proximal formulation. Our central finding is that this decoupled reformulation is more robust than coupled methods. In numerical experiments it provides a wider usable range of the regularization strength and a lower rate of catastrophic over-pruning than the tested joint baseline while maintaining comparable accuracy. We establish these properties in controlled classification and sparse-recovery studies, and examine their scope in a high-dimensional PINN stress test and in the feed-forward layers of OPT-1.3B.
Chinese Translation
我们提出了一种解耦的逐层方法,用于对预训练神经网络的全连接层进行结构化稀疏化。不同于对所有层进行联合惩罚,我们的方法提取浅层二分子网络,对内部权重进行归一化,并对每个块的 внешней权重矩阵施加结构化组惩罚,逐层顺序处理以剪除神经元并降低每层的宽度。我们证明,对于任意正齐次激活函数,该受约束的解耦目标在最优性上等价于对内部和外部权重的某种特定联合惩罚,因此具有简洁的投影和近端形式。我们的核心发现是,这种解耦重构比耦合方法更为稳健。在数值实验中,与所测试的联合基线相比,它提供了更宽的可用的正则化强度范围和更低灾难性过度剪枝发生率,同时保持相当的准确率。我们在受控的分类和稀疏恢复研究中确立了这些性质,并在高维PINN压力测试以及OPT-1.3B的前馈层中考察了其适用范围。
cs.LG / 24 / 2609.21151

EnSol: an environment-aware graph neural network for molecular solubility prediction

EnSol:一种用于分子溶解度预测的环境感知图神经网络
Nguyen, Thao, Shafaei, Saman, Zhang, Zhengyi, Zhao, Huimin, Ji, Heng
Abstract
Molecular solubility directly affects key aspects of molecular development such as reaction feasibility, formulation performance, separation efficiency, and solvent selection. However, experimental measurement across solutes, solvents, and temperatures remains costly and sparsely sampled. Existing computational models often rely on fixed-solvent assumptions, deterministic formulations, or simplified representations of solute-solvent interactions, limiting their ability to capture complex molecular interactions, continuous temperature effects, and experimental uncertainty. Here, we introduce EnSol, an environment-aware probabilistic framework for molecular solubility prediction. EnSol represents the solute and solvent as molecular graphs and learns separate representations for each before bringing them together through cross-attention to capture solute-solvent interactions. Temperature is incorporated directly into the solvent environment through feature-wise modulation, and a mixture density network predicts full solubility distributions to capture both temperature-dependent behavior and experimental uncertainty. On the independent SolProp and Leeds benchmark datasets, EnSol achieved Spearman correlations of 0.876 and 0.601, respectively, outperforming state-of-the-art solubility prediction models across both benchmarks. Beyond computational benchmarking, experimental validation across chemically diverse solute-solvent pairs showed that EnSol maintained strong predictive performance and supported reliable solvent ranking, achieving a Spearman correlation of 0.715. These results show that EnSol can support reliable solubility prediction and solvent selection across diverse chemical systems while accounting for predictive uncertainty.
Chinese Translation
分子溶解度直接影响分子研发的关键环节,如反应可行性、制剂性能、分离效率和溶剂选择。然而,跨溶质、溶剂和温度的实验测量仍然成本高昂且数据稀疏。现有的计算模型往往依赖于固定溶剂假设、确定性表述或对溶质-溶剂相互作用的简化表示,这限制了其捕捉复杂分子相互作用、连续温度效应以及实验不确定性的能力。本文提出了EnSol,一个用于分子溶解度预测的环境感知概率框架。EnSol将溶质和溶剂表示为分子图,分别学习各自的表示,然后通过交叉注意力机制将其结合,以捕捉溶质-溶剂相互作用。温度通过特征级调制直接融入溶剂环境中,混合密度网络预测完整的溶解度分布,从而同时捕捉温度依赖性行为和实验不确定性。在独立的SolProp和Leeds基准数据集上,EnSol分别取得了0.876和0.601的Spearman相关系数,在两个基准上均优于最先进的溶解度预测模型。除计算基准测试外,在化学性质多样的溶质-溶剂对上进行的实验验证表明,EnSol保持了强大的预测性能并支持可靠的溶剂排序,取得了0.715的Spearman相关系数。这些结果表明,EnSol能够在考虑预测不确定性的同时,为多样化化学体系提供可靠的溶解度预测和溶剂选择支持。
cs.LG / 25 / 2609.21158

HMB-GAN: Hybrid Multi-B\'ezier GAN for Vector Shape Synthesis

HMB-GAN:用于矢量形状合成的混合多段贝塞尔生成对抗网络
Thiele-Evans, Elian Hugh, Pham, Binh Duong, Alharbi, Hani Omar M, Aaden, Liibaan, Zaidi, Syed Umer Hasnain, Jayaraman, Prem Prakash, Saeed, Muhammad, Eisenbart, Boris
Abstract
We explore the use of hybrid quantum-classical generative adversarial networks for synthesising CAD-ready vector geometries. Unlike prior work that operates in rasterised or single-B\'ezier domains, we introduce HMB-GAN (Hybrid Multi-B\'ezier GAN), an end-to-end differentiable generative framework that constructs closed shapes through stitched multi-segment B\'ezier representations with geometric continuity enforced by construction. We compare a quantum-enhanced generator with a classical generator within this architecture and evaluate them across point cloud distribution metrics and geometric shape statistics. Results show that despite faster convergence, a reduction in model parameter count, and slightly improved performance on point cloud metrics, the quantum generator suffers from excessive simulator overhead and thus classically-simulated evaluation suffers from hardware constraints. These results demonstrate the feasibility of modelling structured geometries through hybrid quantum architectures whilst highlighting contemporary hardware limitations.
Chinese Translation
我们探索了使用混合量子-经典生成对抗网络来合成可用于计算机辅助设计(CAD)的矢量几何。与以往在栅格化或单段贝塞尔域中工作的研究不同,我们提出了 HMB-GAN(Hybrid Multi-Bézier GAN,混合多段贝塞尔生成对抗网络),这是一种端到端可微分的生成框架,通过拼接的多段贝塞尔表示来构建闭合形状,并在构造过程中保证几何连续性。我们在该架构中将量子增强生成器与经典生成器进行比较,并通过点云分布度量与几何形状统计指标对其进行评估。结果表明,量子生成器虽然收敛更快、模型参数量更少、且在点云度量上性能略有提升,但其受到过大的模拟器开销的影响,因而基于经典模拟的评估受到硬件条件的限制。这些结果证明了利用混合量子架构对结构化几何进行建模的可行性,同时也揭示了当前硬件的局限性。
cs.LG / 26 / 2609.21164

M2G-LLM: Enhancing Clinical Prediction via Multimodal Graph Reasoning and LLM Context Injection

M2G-LLM:通过多模态图推理与大语言模型上下文注入增强临床预测
Choi, Inyoung, Yun, Sukwon, Xin, Jiayi, Peng, Jie, Chen, Tianlong, Long, Qi
Abstract
Integrating diverse data modalities --- such as clinical notes, laboratory results, and medical imaging --- is essential for advancing clinical decision-making. While Large Language Models (LLMs) have shown remarkable performance in processing unstructured clinical text, their limited capacity to incorporate non-text modalities hinders their broader utility in healthcare applications. Here, we introduce M2G-LLM (Multimodal MedGraph-LLM), a novel framework that enhances LLMs with multimodal integration and alignment via Graph Neural Networks (GNNs). Our approach models temporal relationships between patient visits, propagates information across clinically similar patients, and aligns heterogeneous data sources to construct enriched multimodal context vectors. These vectors are injected into the intermediate layers of the LLM, enabling joint reasoning over textual and non-textual modalities. We evaluate M2G-LLM on the MIMIC-IV and MIMIC-CXR datasets, demonstrating improvements in clinical prediction tasks over strong baseline models. Our results highlight the promise of combining the language understanding of LLMs with the relational reasoning capabilities of GNNs for comprehensive, multimodal healthcare analysis.
Chinese Translation
整合多样化的数据模态——如临床笔记、实验室检查结果和医学影像——对于推进临床决策至关重要。尽管大语言模型(LLM)在处理非结构化临床文本方面表现出卓越的性能,但其整合非文本模态的能力有限,阻碍了其在医疗健康领域更广泛的应用。本文提出M2G-LLM(Multimodal MedGraph-LLM),一种通过图神经网络(GNN)为大语言模型增强多模态整合与对齐能力的新框架。我们的方法对患者就诊之间的时间关系进行建模,在临床相似的患者之间传播信息,并对齐异构数据源,以构建丰富的多模态上下文向量。这些向量被注入到大语言模型的中间层,使其能够对文本与非文本模态进行联合推理。我们在MIMIC-IV和MIMIC-CXR数据集上对M2G-LLM进行了评估,结果表明其在临床预测任务上的表现优于强大的基线模型。我们的研究结果凸显了将大语言模型的语言理解能力与图神经网络的关系推理能力相结合,在全面的多模态医疗分析中的应用前景。
cs.LG / 27 / 2609.21172

TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

TierKV:基于预测性多层KV缓存的端侧长上下文大语言模型
Shu, Zhihao, Sanim, Md Musfiqur Rahman, Hu, Jie, Yuan, Kun, Qin, Minghai, Agrawal, Gagan, Niu, Wei
Abstract
Large language models (LLMs) are moving onto mobile devices for increasingly diverse workloads over text, images, video, and audio. These applications often require long contexts, making the Key-Value (KV) cache a dominant memory bottleneck because it grows linearly with sequence length and is accessed at every decoding step. Prior work reduces KV-cache footprint through low-rank compression, token eviction, or flash offloading, but the resulting reconstruction overhead, irreversible token loss, or I/O stalls can offset the benefit of saving memory. We present TierKV, a mobile LLM inference framework built on Predictive Multi-Tier Cache Optimization (PMCO). Before decoding starts, PMCO predicts future cache demand from prefill hidden states and jointly assigns tokens to exact, low-rank, and flash-offloaded tiers under the device memory and accuracy budgets. This formulation retains access to the full context, removes the circular dependency of reactive eviction, and admits a closed-form solver that selects tier boundaries and per-layer ranks at runtime. Across eight text, vision, and audio models on three mobile SoCs, TierKV improves prefill throughput by up to 17.6x over existing mobile LLM frameworks, reduces RAM-resident KV cache by 12.5-34%, thereby enabling substantially longer contexts under the same memory budget, while incurring only minor accuracy degradation.
Chinese Translation
大语言模型(LLM)正逐步迁移至移动设备,以处理日益多样化的文本、图像、视频和音频工作负载。这些应用通常需要长上下文,使得键值(KV)缓存成为主要的内存瓶颈,因为其大小随序列长度线性增长,且在每一步解码时都需要被访问。先前的工作通过低秩压缩、令牌逐出或闪存卸载来减少KV缓存的占用,但由此产生的重建开销、不可逆的令牌丢失或I/O阻塞可能会抵消节省内存所带来的收益。我们提出了TierKV,一个基于预测性多层缓存优化(Predictive Multi-Tier Cache Optimization,PMCO)的移动端LLM推理框架。在解码开始之前,PMCO根据预填充(prefill)阶段的隐状态预测未来的缓存需求,并在设备内存和精度预算的约束下,将令牌联合分配至精确、低秩和闪存卸载三个层级。这一形式化方法保留了对完整上下文的访问能力,消除了反应式逐出的循环依赖问题,并支持一种闭式求解器,可在运行时选择层级边界和各层的秩。在三种移动SoC上的八个文本、视觉和音频模型上,TierKV相比现有移动LLM框架将预填充吞吐量提升高达17.6倍,将常驻内存的KV缓存减少12.5%–34%,从而在相同内存预算下支持显著更长的上下文,同时仅带来轻微的精度下降。
cs.LG / 28 / 2609.21190

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

SWE-Proof:语言模型能否通过机器验证的证明解决真实世界的问题?
Ma, George, Mikek, Benjamin, Li, Haoyu, Erata, Ferhat, Zhang, Yuhao, Shui, Zeren, Tehrani, Behrooz Omidvar, Huan, Jun, Ramanathan, Murali Krishna, Sojoudi, Somayeh, Zhou, Hao, Deoras, Anoop
Abstract
Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check correctness with held-out test suites, which are inherently incomplete and increasingly susceptible to memorization. Formal verification avoids both problems, but existing work covers only standalone tasks whose specifications are given as input, not real issues, which touch large repositories and state intent in vague natural language. We present Benchproofer, a pipeline that turns a coding task with a known correct patch into a formally verified one: it writes a specification for the new code, summarizes the existing functions that code calls with axioms, and admits an instance only after mechanical and adversarial gates agree. Applying it to SWE-bench Verified yields SWE-Proof, 500 real issues whose correctness is formally verified rather than tested, and it extends to SWE-bench Pro. Across two frontier models, verification catches what tests miss: a quarter to a half of test-passing patches admit counterexamples, which a structured natural-language specification does not fix, while a correct formal one lifts resolution from 85% to 95% for Opus 4.8. Writing that specification is the hard part: models that must write their own gain nothing over an unaided baseline, and only 62% of their specifications pass our audit. The usual failure is faithfulness, a specification that constrains part of the required behavior and leaves the rest free. Specification quality still tracks the outcome, failing on 89% of unresolved instances against 47% of resolved ones, making faithful specification synthesis a concrete open problem.
Chinese Translation
确保大语言模型(LLM)生成代码的正确性是现代软件工程面临的核心挑战。智能体式代码生成的基准测试通过保留的测试套件来检验正确性,但测试套件本质上是不完备的,且日益容易受到记忆化(memorization)的影响。形式化验证可以避免这两个问题,但现有工作仅涵盖以输入形式给出规格说明的独立任务,而非真实问题——后者涉及大型代码仓库,并以模糊的自然语言表达意图。我们提出了 Benchproofer,一个将具有已知正确补丁的编码任务转化为形式化验证任务的流水线:它为新代码撰写规格说明,用公理概括代码所调用的已有函数,并且只有在通过机械化与对抗性两道审查关卡后才会接受一个实例。将其应用于 SWE-bench Verified 后,我们构建了 SWE-Proof——包含 500 个真实问题,其正确性通过形式化验证而非测试来保证;该方法还可扩展至 SWE-bench Pro。在两个前沿模型上的实验表明,验证能发现测试所遗漏的问题:四分之一到一半通过测试的补丁存在反例,而结构化的自然语言规格说明无法解决这一问题;但正确的形式化规格说明可将 Opus 4.8 的解决率从 85% 提升至 95%。撰写这样的规格说明是难点所在:必须自行撰写规格说明的模型相比无辅助的基线并无任何提升,且其撰写的规格说明中仅有 62% 能通过我们的审计。常见的失败是忠实性(faithfulness)问题——即规格说明仅约束了所需行为的一部分,而对其余部分不加限制。规格说明质量仍与最终结果相关:在未解决实例上的失败率为 89%,而在已解决实例上为 47%,这使得忠实的规格说明自动合成成为一个具体的开放性问题。
cs.LG / 29 / 2609.21197

Reliability-Centered Evaluation of Sparse Longitudinal CT Lesion-Size Forecasting with Conformal Interval Calibration and Gompertz-Inspired Regularization

基于可靠性评估的稀疏纵向CT病灶尺寸预测研究:共形区间校准与Gompertz启发的正则化方法
Kong, Lingfei
Abstract
Sparse longitudinal CT follow-up limits lesion-size forecasting when only a few prior observations are available. We constructed a five-visit DLT-derived same-lesion trajectory benchmark from DeepLesion and Deep Lesion Tracker (DLT), yielding 205 trajectories from 129 patients. We compared an exploratory conventional sparse-to-final analysis with a primary fixed visit-index horizon design predicting the common log change from T3 to T4 while progressively adding earlier observations, evaluating predictive accuracy, uncertainty reliability, post-hoc conformal interval calibration, subgroup performance, and Gompertz-inspired trajectory regularization. The evaluated methods showed partially overlapping point-prediction accuracy but distinct uncertainty behavior. Mean held-out RMSE across ten training seeds was 0.4726, 0.4305, 0.4499, and 0.4513 for m = 1, 2, 3, 4, indicating the lowest mean RMSE at m = 2; additional history did not improve RMSE. At m = 4, raw Cohort-Level Feature GP coverage was near the 95% nominal level, whereas MC Dropout, Deep Ensemble, and residual-scale intervals were conservative. Patient-level conformal calibration generally produced near-nominal or conservative coverage at the cost of wider intervals. Patient-grouped development cross-validation selected lambda* = 0 for the Gompertz-inspired term. A global population reference frequently opposed lesion-level change directions, and prediction difficulty varied across anatomical subgroups. Overall, additional historical observations provided limited predictive benefit once the prediction horizon was controlled, while predictive accuracy, uncertainty reliability, and trajectory consistency did not necessarily improve together, and should be evaluated jointly in sparse longitudinal imaging.
Chinese Translation
当仅有少量既往观测数据可用时,稀疏的纵向CT随访限制了病灶尺寸预测的准确性。我们基于DeepLesion和Deep Lesion Tracker(DLT)构建了一个基于DLT的、包含五次随访的同病灶轨迹基准数据集,最终获得来自129名患者的205条轨迹。我们将探索性的常规稀疏至终末分析作为对照,与主要的固定随访索引跨度设计进行比较:后者预测T3至T4的公共对数变化,并逐步纳入更早的观测数据,同时评估预测精度、不确定性可靠性、事后共形区间校准、亚组表现以及受Gompertz模型启发的轨迹正则化。所评估的方法在点预测精度上表现出部分重叠,但在不确定性行为上存在明显差异。在十个训练随机种子下,留出集的平均RMSE在m = 1, 2, 3, 4时分别为0.4726、0.4305、0.4499和0.4513,表明在m = 2时平均RMSE最低;纳入更多历史数据并未改善RMSE。在m = 4时,原始的队列级特征高斯过程(Cohort-Level Feature GP)的覆盖率接近95%的名义水平,而MC Dropout、深度集成(Deep Ensemble)以及残差尺度区间的覆盖率则偏保守。患者级的共形校准通常产生接近名义或偏保守的覆盖率,但代价是区间更宽。采用患者分组开发交叉验证时,受Gompertz启发的正则化项的lambda*被选为0。全局人群参考常常与病灶层面的变化方向相悖,且预测难度在不同解剖学亚组之间存在差异。总体而言,在控制预测跨度的前提下,纳入更多历史观测数据的预测收益有限;预测精度、不确定性可靠性与轨迹一致性并不一定同步改善,因此在稀疏纵向影像研究中应对三者进行联合评估。
cs.LG / 30 / 2609.21280

MIRCID: Inferred Hub-miRNAs Drive Cross-Task Improvements in Drug Mechanistic Modeling

MIRCID:推断的中心miRNA(Hub-miRNA)驱动药物作用机制建模的跨任务性能提升
Cao, Xin, Chen, Yigang, Xu, Jiatong, Zhang, Ziyue, Cheng, Xiang, Wang, Shenyu, Zhang, Yangyi, Cai, Xiaoxuan, Cui, Shidong, Zhu, Zihao, Ji, Xiang, Huang, Hsi-Yuan, Lin, Yang-Chi-Dung, Huang, Hsien-Da
Abstract
Drug mechanism-of-action (MoA) modeling commonly relies on perturbational transcriptomes, but matched microRNA (miRNA) measurements are often unavailable. Inferred regulatory features offer a scalable way to reuse these data. Here, we present MIRCID, a framework comparing gene expression with inferred transcription factor (TF) activity and miRNA expression across pathway classification and similarity-based MoA retrieval. HubmiRNet infers 414 pan-cancer hub miRNAs (HubmiRs) from 977 L1000 landmark genes, achieving a Pearson correlation coefficient of 87.72\%; its 1,298-output variant also outperformed SiCmiR on the full-miRNA task (71.21\% versus 67.30\%). In the evaluated comparisons, miRNA augmentation provided more consistent gains than TF activity. Generic embedding controls showed model-dependent utility, while complementarity analyses identified a distinct, partially linearly recoverable representation that retained gene-derived structure. Illustrative rescue cases linked improved classification to biologically plausible miRNA patterns in samples with weak transcriptional signatures. These findings support inferred HubmiRs as a biologically informed recoding of transcriptomic data for perturbational drug modeling, while leaving recovery of measured perturbational miRNA responses to further validation.
Chinese Translation
药物作用机制(Mechanism-of-Action, MoA)建模通常依赖于扰动转录组数据,但匹配的微RNA(miRNA)测量数据往往难以获得。推断得到的调控特征为复用这些数据提供了一种可扩展的方法。本文提出MIRCID框架,在通路分类和基于相似性的MoA检索任务中,比较了基因表达与推断的转录因子(TF)活性及miRNA表达的组合效果。HubmiRNet从977个L1000地标基因中推断出414个泛癌种中心miRNA(HubmiRs),皮尔逊相关系数达到87.72%;其具有1,298个输出的变体在全miRNA任务上也优于SiCmiR(71.21% 对 67.30%)。在所评估的比较中,miRNA增强相比TF活性带来了更为一致的性能提升。通用嵌入对照显示出依赖模型的效用,而互补性分析则识别出一种独特、部分可线性恢复的表示,该表示保留了源自基因数据的结构。示例性的“挽救”案例表明,在转录信号较弱的样本中,分类性能的提升与生物学上合理的miRNA模式相关联。这些发现支持将推断的HubmiRs作为转录组数据的一种具有生物学依据的重新编码方式,用于扰动性药物建模,而实测扰动miRNA响应的恢复仍有待进一步验证。
cs.LG / 31 / 2609.21288

Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding

多受试者预训练使闭语料库表面肌电语音解码能够实现短校准个性化
Le, Chenqian, Fumagalli, Beatrice, Esmaeili, Yasamin, Chen, Xupeng, He, Tianyu, Emami, Nikasadat, Flinker, Adeen, Wang, Yao
Abstract
Surface electromyography (sEMG)-based silent speech interfaces are limited by cross-user variability and calibration burden. We study a limited-data setting in which each of 27 speech-typical participants contributed less than 0.5 h of data (21.3 min on average) across Aloud and Mimed speech. Within a closed 50-sentence corpus, we used leave-one-subject-out evaluation, initializing from a released single-subject checkpoint, pretraining on non-held-out participants, and fine-tuning on the target participant. This pipeline achieved 21.7% character error rate (CER) and 31.9% word error rate (WER), compared with 49.3% CER without target-subject calibration and 68.0% CER for direct checkpoint fine-tuning. Multi-subject pretraining from random initialization followed by fine-tuning reached 44.9% CER and did not converge under the fixed schedule in 5 of 27 folds, indicating substantial optimization and accuracy benefits from checkpoint initialization. Macro-averaged CER declined from 74.4% with one pretraining participant to 21.7% with 26. Three minutes of target-subject calibration achieved 20.5% CER and 31.7% WER, with no statistically significant difference from the full approximately 13-min pool (21.7% CER and 31.9% WER). A subject-specific adapter provided no detectable benefit. Excluding the five evaluation sentences from all sEMG model-training data increased CER and WER to 78.6% and 99.9%. These results support short-calibration personalization in a standardized-montage, closed-corpus setting.
Chinese Translation
基于表面肌电图(sEMG)的静默语音接口受限于跨用户变异性与校准负担。我们研究了一种有限数据场景:27名典型语音能力受试者每人贡献了少于0.5小时的数据(平均21.3分钟),涵盖出声语音(Aloud)和默语(Mimed)两种模式。在包含50个句子的封闭语料库内,我们采用留一受试者评估方法,从已发布的单受试者检查点初始化,在非留出受试者上进行预训练,再在目标受试者上微调。该流程实现了21.7%的字符错误率(CER)和31.9%的词错误率(WER),而未经目标受试者校准时CER为49.3%,直接对检查点微调的CER为68.0%。从随机初始化开始的多受试者预训练加微调仅达到44.9%的CER,且在27折中的5折未能按固定训练计划收敛,表明检查点初始化带来了显著的优化与准确率收益。宏平均CER从仅有1名预训练受试者时的74.4%下降到26名受试者时的21.7%。3分钟的目标受试者校准即可达到20.5%的CER和31.7%的WER,与使用全部约13分钟数据(CER 21.7%、WER 31.9%)相比无统计学显著差异。受试者专属适配器未带来可检测的收益。若将5个评估句子从所有sEMG模型训练数据中排除,CER和WER分别升至78.6%和99.9%。这些结果支持在标准化电极排布、封闭语料库设置下的短校准个性化。
cs.LG / 32 / 2609.21296

FairLMs: A Turnkey Library for Fairness in Language Models

FairLMs:一个面向语言模型公平性的开箱即用工具库
Zhang, Jiale, Larionov, Michael, Wang, Zichong, Yin, Zhipeng, Zhang, Wenbin
Abstract
Fairness research on language models involves measuring bias, applying mitigation methods, and examining the evidence on which an evaluation rests. Existing tools offer complementary functionality through different interfaces, so combining them requires reconciling model interfaces, evidence formats, access constraints, and result types before applicability can be checked or methods compared. We introduce \textbf{FairLMs}, a Python library that connects these activities through explicit declarations of model capabilities and input requirements. It provides 33 intrinsic and extrinsic metrics, 14 mitigation components spanning four intervention categories, 14 dataset and scoring-instrument diagnostics, adapters for the three Transformer architectures and supported hosted completion APIs, and benchmark loaders. Declarations are checked before execution and results carry the configuration under which they were obtained, so that compatible components can be combined, methods compared under a common protocol, and workflows extended to new models and datasets. The source code is available at: https://github.com/FairLMs/FairLMs.
Chinese Translation
语言模型的公平性研究涉及偏差度量、缓解方法的应用,以及对评估所依据证据的考察。现有工具通过不同的接口提供互补的功能,因此将它们组合起来需要先协调模型接口、证据格式、访问限制和结果类型,才能检查其适用性或比较不同方法。我们推出了 FairLMs,这是一个 Python 库,它通过对模型能力和输入要求的显式声明将这些活动连接起来。该库提供 33 个内在和外在度量指标、涵盖四类干预方式的 14 个缓解组件、14 个数据集和评分工具诊断模块、支持三种 Transformer 架构及所支持托管补全 API 的适配器,以及基准数据加载器。声明在执行前会被检查,结果会携带其生成时的配置信息,从而使得兼容的组件能够组合使用、方法能够在统一协议下进行比较,并且工作流程能够扩展到新的模型和数据集。源代码可在以下网址获取:https://github.com/FairLMs/FairLMs。
cs.LG / 33 / 2609.21306

Fast And Accurate Text Content File Type Identification

快速且准确的文本内容文件类型识别
Nandan, Manu, Brautbar, Michael, Raff, Edward
Abstract
A common requirement across organizations is to have a tool that can identify file types based on their contents, particularly in the cybersecurity domain where magic numbers and file extensions can not be trusted. While existing tools work well in practice, there is plenty of room for improvement either in terms of computational load and time for detection in the case of model based tools like Magika or in terms of accuracy of detection in the case of file parsing tools that use programming language constructs. In this study, we propose a neural network model for identification of types of text content files, especially source code, that is more accurate and faster than other available tools. Our experiments on open-source files indicate that it is not only more accurate on average for text-content file-type identification, but also approximately four times faster than Magika, while being 28% smaller in size.
Chinese Translation
各组织普遍需要一个能够基于文件内容识别文件类型的工具,尤其是在网络安全领域,因为魔数(magic numbers)和文件扩展名并不总是可信的。现有工具在实际应用中表现良好,但仍有很大的改进空间:对于像 Magika 这类基于模型的工具,可以在计算负载和检测耗时方面改进;而对于使用编程语言语法结构进行文件解析的工具,则可以在检测准确率方面提升。在本研究中,我们提出了一种用于识别文本内容文件(尤其是源代码)类型的神经网络模型,其准确率和速度均优于其他现有工具。我们在开源文件上的实验表明,该模型不仅在文本内容文件类型识别方面平均准确率更高,而且速度约为 Magika 的四倍,同时模型体积缩小了28%。
cs.LG / 34 / 2609.21309

An Introduction to Compression-Based Machine Learning

基于压缩的机器学习导论
Hurwitz, John, Raff, Edward, Nicholas, Charles K.
Abstract
Any lossless compression algorithm (like gzip) may be converted into a machine learning method, via either Normalized Compression Distance or the Minimum Description Length principle. Any auto-regressive model may be converted into a lossless compression method via entropy coding. This seemingly circular dependence has unrealized potential in modern artificial intelligence and machine learning, and we survey and formalize the various strategies that have been used to leverage compression for machine learning. We introduce and empirically validate a design framework for compression-based ML, finding compression-based methods competitive with conventional baselines and decisively stronger on malware. We find that varying these design choices yields accuracy gains of up to 0.62.
Chinese Translation
任何无损压缩算法(如 gzip)都可以通过归一化压缩距离(Normalized Compression Distance)或最小描述长度(Minimum Description Length)原则转换为机器学习方法。同样,任何自回归模型都可以通过熵编码转换为无损压缩方法。这种看似循环的依赖关系在现代人工智能与机器学习中蕴含着尚未实现的潜力,本文综述并形式化了利用压缩实现机器学习的各种策略。我们提出并通过实验验证了一个基于压缩的机器学习设计框架,发现基于压缩的方法与传统基线方法具有竞争力,并且在恶意软件检测任务上具有决定性优势。我们发现,调整这些设计选择可带来高达 0.62 的准确率提升。
cs.LG / 35 / 2609.21327

Deep Reinforcement Learning with Buffered Quantile Objectives

基于缓冲分位数目标的深度强化学习
Alipour-vaezi, Mohammad, Khodadadian, Sajad
Abstract
Quantile-based reinforcement learning provides an interpretable approach to risk-sensitive decision-making by optimizing a prescribed quantile of the cumulative-return distribution. Despite this appeal, learning under a point quantile objective is challenging: quantiles can change abruptly under small perturbations of the return distribution, and exact quantile-sensitive planning requires computationally demanding distributional optimization. Lower-buffered quantiles alleviate the former difficulty by averaging neighboring quantiles immediately below the target level, providing a smoother surrogate while preserving the underlying point-quantile objective. Existing methods based on this principle, however, remain model-based and rely on explicit return-law planning, limiting their applicability beyond small tabular problems. We develop Deep-BQRL, a model-free distributional reinforcement-learning framework that extends buffered-quantile learning to neural function approximation. The method learns conditional return quantiles directly from sampled transitions, constructs buffered action scores from the relevant region of the learned quantile function, and uses ensemble disagreement to guide exploration. An augmented input representation allows the learned policy to respond to trajectory information without explicitly reproducing the quantile-state recursion required by exact planning. Experiments on an asset-selling optimal-stopping problem and slippery FrozenLake compare Deep-BQRL with model-based UCB-BQRL and tabular PPO and TRPO implementations. In asset selling, Deep-BQRL attains smaller mean cumulative point-quantile policy gaps than PPO and TRPO at the reported target levels, while UCB-BQRL retains the smallest gaps. The learned stopping decisions also vary with the target quantile, providing an interpretable illustration of the method's risk-sensitive behavior.
Chinese Translation
基于分位数的强化学习通过优化累积回报分布的指定分位数,为风险敏感决策提供了一种可解释的方法。尽管具有这一吸引力,在点分位数目标下进行学习仍然具有挑战性:回报分布的微小扰动可能导致分位数发生剧烈变化,且精确的分位数敏感规划需要计算代价高昂的分布式优化。下缓冲分位数通过在目标水平正下方对相邻分位数取平均,缓解了前一个困难,在保持底层点分位数目标的同时提供了一个更平滑的替代目标。然而,基于这一原理的现有方法仍然是基于模型的,依赖于显式的回报律规划,限制了其在小型表格问题之外的适用性。我们提出了Deep-BQRL,这是一个无模型的分布式强化学习框架,将缓冲分位数学习扩展到神经函数逼近。该方法直接从采样转移中学习条件回报分位数,从所学分位数函数的相关区域构建缓冲动作评分,并利用集成分歧来引导探索。增广的输入表示使所学习的策略能够响应轨迹信息,而无需显式复现精确规划所需的分位数-状态递推。在资产出售最优停时问题和湿滑FrozenLake环境上的实验将Deep-BQRL与基于模型的UCB-BQRL以及表格型PPO和TRPO实现进行了比较。在资产出售问题中,在所报告的目标水平上,Deep-BQRL获得的平均累积点分位数策略差距小于PPO和TRPO,而UCB-BQRL仍保持最小的差距。此外,所学习的停时决策随目标分位数的变化而变化,直观地展示了该方法的风险敏感行为。
cs.LG / 36 / 2609.21332

Routine Blood Tests Outperform CRP for Distinguishing Bacterial From Viral Infection in Children

常规血液检查在区分儿童细菌性与病毒性感染方面优于CRP
Demireva, Mihaela, Mitev, Zhecho, Chinareva-Klimentova, Djuna, Ivanov, Svetoslav, Nalbantov, Georgi, Mitev, Dimitar
Abstract
Acute infectious diseases are among the leading causes of medical consultations and hospitalizations in children worldwide. These infections are predominantly caused by viruses or bacteria, yet differentiating between the two remains a common clinical challenge. As a result, pediatricians often default to the safer option of prescribing antibiotics contributing to the growing problem of antimicrobial resistance. The objective is to assess the additional predictive value of CBC towards determining the current infection. This retrospective study used data from 906 pediatric patients aged between 2 and 14 years who were tested positive either for viral or bacterial infection between 2022 and 2026. Inclusion criteria further required availability of CBC results and CRP level measurements. These laboratory parameters as well as age were used as input features for several supervised classification models. Model performance was evaluated using AUC, sensitivity and specificity. The best performing model is XGBoost, which included all features, achieving out of-sample performance of AUC of 81.7% and sensitivity of 70.8%, specificity of 79.2%. All trained models outperform a CRP-based only decision-rule model in terms of AUC. We suggest that the decision to prescribe antibiotics should be based on a number of factors, including but not limited to CBC, some of which are not currently incorporated into routine practice.
Chinese Translation
急性感染性疾病是全球儿童就诊和住院的主要原因之一。这些感染主要由病毒或细菌引起,但区分二者仍是临床上常见的难题。因此,儿科医生往往倾向于选择更保守的方案——开具抗生素处方,从而加剧了日益严重的抗菌药物耐药性问题。本研究旨在评估全血细胞计数(CBC)对判断当前感染类型的额外预测价值。这项回顾性研究纳入了2022年至2026年间病毒或细菌感染检测呈阳性的906名2至14岁儿科患者的数据。纳入标准还要求具备CBC结果和CRP水平检测数据。这些实验室指标以及年龄被用作多个监督分类模型的输入特征,模型性能通过AUC、敏感性和特异性进行评估。表现最佳的模型是包含全部特征的XGBoost,其样本外性能达到AUC 81.7%、敏感性70.8%、特异性79.2%。所有训练的模型在AUC方面均优于仅基于CRP的决策规则模型。我们建议,开具抗生素的决定应基于多种因素,包括但不限于CBC,其中一些因素目前尚未纳入常规临床实践中。
cs.LG / 37 / 2609.21346

IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

IntBMoE:将块级条件化融入专家组合以实现全员参与的混合专家模型
Cheng, Ran, Xu, Longfei, Liu, Zheng, Liu, Kaikui, Chu, Xiangxiang
Abstract
Mixture-of-Experts (MoE) scales capacity, but existing designs cannot set three quantities independently. For a single token, participation is how many experts contribute knowledge to its output, execution is how many are actually computed (compute cost), and materialization is how many expert-sized parameter sets must be built and stored (memory cost). Sparse routing keeps execution and materialization low, but shrinks participation: for each token, only a few experts contribute. Dense output-mixing restores full participation, but its execution grows with the number of experts. Parameter-merging keeps execution at one expert, but its materialization grows with the number of routing decisions. We propose IntBMoE, a block-conditioned MoE that decouples all three by pairing dense expert composition with sparse block execution. Its blocks come from a small learned codebook, one per entry. At each internal layer, a lightweight hypernetwork merges all expert bases in that layer's pool into one composed expert. Participation is full, because every composed expert draws on the entire pool. Execution stays sparse, because a router sends each token to only a few blocks. Materialization is bounded, because the codebook, not the input, fixes how many blocks exist. Dual-Path Residual Gating (DPRG) further couples two independently composed paths through multiplicative gating. Experiments on image classification show consistent gains over representative sparse and dense MoE baselines. Additional experiments on language modeling and sequential recommendation validate its generalization beyond vision. IntBMoE is fully deployed in AMap's generative recommendation system, serving hundreds of millions of users under a 60ms latency budget, with a 2.4% relative UVCTR gain in online A/B testing. Our code is available at https://github.com/AMAP-ML/DreamX-Rec/.
Chinese Translation
混合专家模型能够扩展模型容量,但现有设计无法独立设定三个关键量。对于单个 token 而言,参与度(participation)是指有多少专家为其输出贡献知识,执行量(execution)是指实际计算了多少专家(计算成本),物化量(materialization)是指必须构建和存储多少个专家规模的参数集(内存成本)。稀疏路由虽能保持较低的执行量和物化量,但会降低参与度:对每个 token,只有少数专家参与贡献。稠密输出混合可恢复全员参与,但其执行量会随专家数量增长。参数合并将执行量保持为单个专家,但其物化量会随路由决策数量增长。我们提出 IntBMoE,一种块条件化的 MoE,通过将稠密专家组合与稀疏块执行相结合,实现三者的解耦。其块来自一个小型的可学习码本(codebook),每个条目对应一个块。在每个内部层,一个轻量级超网络将该层专家池中的所有专家基座合并为一个组合专家。由于每个组合专家都利用整个专家池,参与度是全员参与的;由于路由器仅将每个 token 发送给少数块,执行量保持稀疏;由于块的数量由码本而非输入决定,物化量是有界的。双路径残差门控(Dual-Path Residual Gating, DPRG)进一步通过乘性门控耦合两条独立组合的路径。在图像分类任务上的实验表明,该方法相较于代表性的稀疏和稠密 MoE 基线均取得一致提升。在语言建模和序列推荐任务上的额外实验验证了其在视觉领域之外的泛化能力。IntBMoE 已完全部署于高德地图(AMap)的生成式推荐系统中,在 60ms 延迟预算下服务数亿用户,在线 A/B 测试中 UVCTR 相对提升 2.4%。代码发布于 https://github.com/AMAP-ML/DreamX-Rec/。
cs.LG / 38 / 2609.21381

Knowledge-Graph-Augmented Chronos-2 for HEC-RAS Surrogate Forecasting

知识图谱增强的 Chronos-2 用于 HEC-RAS 代理预报
Holmberg, Edward, Ioup, Elias, Abdelguerfi, Mahdi
Abstract
We investigate whether coupling a time-series foundation model to hydraulic project knowledge improves surrogate forecasting of HEC-RAS water-surface elevation (WSE). We present KG-Chronos-2, which combines a frozen Chronos-2 predictor with exact-state residual decoding, graph-conditioned historical retrieval, and input-aligned correction. We compare the method with persistence, a residual LSTM, project-conditioned recurrent GeoFNO, a hydraulic DCRNN-style model, and frozen Chronos-2. Task-specific fitting uses the 2008 simulation. Evaluation covers 64 fixed 24-hour windows from the 2011 and 2002 simulations at 4,675 cross sections in 71 reaches on a shared geometry. KG-Chronos-2 achieves event-balanced root-mean-square error 0.246970 in native WSE units. It reduces RMSE by 14.13% relative to frozen Chronos-2, 29.38% relative to the hydraulic DCRNN-style model, and 39.54% relative to recurrent GeoFNO. The 95% hierarchical-bootstrap interval for its event-balanced RMSE difference from frozen Chronos-2 is [-0.075177, -0.016317]. KG-Chronos-2 also achieves the lowest active-window and final-lead RMSE among the six completed systems. These results support coupling a frozen temporal predictor to project knowledge for warm-start HEC-RAS forecasting on the fixed benchmark.
Chinese Translation
我们研究了将时间序列基础模型与水利工程知识相结合是否能提升 HEC-RAS 水面高程(WSE)的代理预报效果。我们提出了 KG-Chronos-2,该方法将冻结的 Chronos-2 预测器与精确状态残差解码、图条件历史检索以及输入对齐修正相结合。我们将该方法与持续性预报、残差 LSTM、项目条件循环 GeoFNO、水利 DCRNN 风格模型以及冻结 Chronos-2 进行了比较。任务特定的拟合使用 2008 年模拟数据。评估覆盖 2011 年和 2002 年模拟中共享几何条件下 71 个河段 4,675 个断面的 64 个固定 24 小时窗口。KG-Chronos-2 在原始 WSE 单位下实现了事件平衡均方根误差 0.246970。相对于冻结 Chronos-2,其 RMSE 降低了 14.13%;相对于水利 DCRNN 风格模型降低了 29.38%;相对于循环 GeoFNO 降低了 39.54%。其与冻结 Chronos-2 之间事件平衡 RMSE 差值的 95% 分层自助法置信区间为 [-0.075177, -0.016317]。在六个已完成的系统中,KG-Chronos-2 还在活跃窗口和最终预报时效上取得了最低的 RMSE。这些结果支持将冻结的时间序列预测器与项目知识相结合,以在该固定基准上实现热启动的 HEC-RAS 预报。
cs.LG / 39 / 2609.21382

Probabilistic Forecasting of Business Process Executions with Neural Temporal Point Processes

基于神经时间点过程的业务流程执行概率预测
Yuan, Jiaxin, Grigori, Daniela, van der Aa, Han
Abstract
Operators of service-based systems act on forecasts of how a running execution will continue, and such a forecast is actionable only if its reliability is known. Mainstream deep-learning models for this task are discriminative and deterministic: they emit a single next activity and a single remaining-time estimate, without a distribution to reason over. We instead cast the problem as generative sequence modelling with marked temporal point processes, which define a joint density over the next mark and its inter-event time and therefore deliver predictive distributions by construction. Real event logs violate the simple-point-process assumption these models rest on, since consecutive events frequently carry identical timestamps; we handle such ties explicitly and combine a transformer encoder with a mixture decoder over inter-event times, trained by exact log-likelihood. On ten public logs, the resulting model matches discriminative baselines on point accuracy, dominates them on the calibration and sharpness of remaining-time distributions, and is the cheapest at inference, since a full predictive distribution is obtained in a single forward pass without sampling.
Chinese Translation
基于服务的系统的运营者依赖于对正在运行的执行将如何继续的预测,而只有当预测的可靠性已知时,该预测才具有可操作性。主流的深度学习模型在此任务上是判别式和确定性的:它们仅输出单个下一活动和单个剩余时间估计,而没有可用于推理的分布。我们转而将该问题建模为基于标记时间点过程的生成式序列建模,该过程定义了下一标记及其事件间隔时间的联合密度,因此在构造上即可提供预测分布。真实的事件日志违反了这些模型所依赖的简单点过程假设,因为相邻事件经常带有相同的时间戳;我们显式地处理此类时间并列情况,并将Transformer编码器与事件间隔时间的混合解码器相结合,通过精确对数似然进行训练。在十个公开日志上,所得到的模型在点准确率上与判别式基线相当,在剩余时间分布的校准性和锐度上优于这些基线,并且在推理上成本最低,因为无需采样即可通过单次前向传播获得完整的预测分布。
cs.LG / 40 / 2609.21425

Tracing the Evidence Behind Zero-Shot Time-Series Forecasting: A Source-First Taxonomy and Audit Framework

追溯零样本时间序列预测背后的证据:一种以证据来源为先的分类体系与审计框架
Kong, Delun, Ling, Wanyun, Liu, Chenxi, Li, Ziyue
Abstract
Zero-shot time-series forecasting (TSF) is often described as forecasting without target-specific parameter updates, but that training-status condition does not specify what evidence the system may use. A frozen language model prompted with serialized values, a time-series model pretrained on broad forecasting corpora, and a retrieval-augmented forecaster may all satisfy the no-update condition while drawing on different transferable evidence. This paper argues that zero-shot TSF should therefore be governed as an evidence-access claim. We propose a source-first taxonomy that separates three primary evidence sources---frozen LLM prior reuse, parametric time-series pretraining, and retrieval-augmented external memory---from the architectures that implement them. After the source is identified, four additional audit questions remain: task interface, forecast object and scoring, prediction-time context, and resource budget. The resulting agenda is to make zero-shot leaderboards auditable by reporting evidence boundaries and interface assumptions alongside scores, so that benchmark progress reflects transferable forecasting capability rather than undisclosed changes in context, memory, or budget.
Chinese Translation
零样本时间序列预测(TSF)通常被描述为在不针对目标数据集进行参数更新的情况下进行预测,但这一训练状态条件并未说明系统可以使用何种证据。以序列化数值为提示的冻结语言模型、在大规模预测语料上预训练的时间序列模型,以及检索增强的预测器,都可能在满足“无更新”条件的同时,依赖不同的可迁移证据。因此,本文主张应将零样本TSF作为一种“证据可访问性”声明来加以规范。我们提出一种以证据来源为先的分类体系,将三种主要证据来源——冻结大语言模型(LLM)的先验复用、参数化时间序列预训练、以及检索增强的外部记忆——与实现它们的具体架构区分开来。在确定证据来源之后,还需回答四个额外的审计问题:任务接口、预测对象与评分方式、预测时的上下文以及资源预算。由此形成的议程是:在零样本排行榜中,除报告分数外,还应报告证据边界与接口假设,使其可被审计,从而使基准测试的进展真正反映可迁移的预测能力,而非上下文、记忆或预算方面未披露的变化。
cs.LG / 41 / 2609.21427

Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation

基于KKT重构的均值-方差投资组合优化决策聚焦学习
Nosaka, Kensei, Ikeda, Shunnosuke, Takano, Yuichi
Abstract
Mean-variance portfolio optimization (MVO) is a central framework in data-driven asset management. A widely adopted approach is a two-stage framework that first predicts expected returns and then solves the optimization problem based on these predictions, with the predictive models trained by minimizing prediction errors. However, this objective of prediction is not aligned with the quality of the downstream portfolio decision. Decision-focused learning (DFL), which directly minimizes the downstream decision loss within the learning process, has thus emerged as a promising direction. However, existing DFL approaches to MVO rely on surrogate losses or constraint relaxations for tractability, creating a structural mismatch between predictive model training and the constrained MVO solved at evaluation. We propose a single-level optimization formulation that incorporates the Karush-Kuhn-Tucker (KKT) optimality conditions of the lower-level MVO into the upper-level learning problem. This formulation explicitly preserves the budget and short-sale constraints while remaining tractable for standard nonlinear optimization solvers. Rolling-window experiments on real-world ETF (Exchange Traded Funds) data across two asset universes with different correlation structures show that our method achieved the best performance on multiple investment metrics and also demonstrated performance improvement due to the proposed regularization.
Chinese Translation
均值-方差投资组合优化(MVO)是数据驱动资产管理中的核心框架。一种被广泛采用的方法是两阶段框架:首先预测期望收益,然后基于这些预测求解优化问题,其中预测模型通过最小化预测误差进行训练。然而,这种预测目标与下游投资组合决策的质量并不一致。决策聚焦学习(Decision-Focused Learning, DFL)通过在学习过程中直接最小化下游决策损失,由此成为一个有前景的研究方向。然而,现有的面向MVO的DFL方法为了可处理性而依赖代理损失或约束松弛,导致预测模型训练与评估时所求解的带约束MVO之间存在结构性不匹配。我们提出了一种单层优化表述,将下层MVO的Karush-Kuhn-Tucker(KKT)最优性条件融入上层学习问题。该表述显式地保留了预算约束和卖空约束,同时对标准非线性优化求解器仍保持可处理性。在具有不同相关结构的两种资产池的真实ETF(交易型开放式指数基金)数据上进行的滚动窗口实验表明,我们的方法在多个投资指标上取得了最佳性能,并且所提出的正则化方法也带来了性能提升。
cs.LG / 42 / 2609.21445

Optimal Randomized Proper Online Learning

最优随机规范化在线学习
Chase, Zachary, Mehalel, Idan
Abstract
We prove that the optimal expected mistake bound of online learning a function class $\mathcal{H}$ by a randomized proper learning algorithm is $O(\mathtt{L}(\mathcal{H}) \log T)$, where $\mathtt{L}(\mathcal{H})$ is the Littlestone dimension of $\mathcal{H}$ and $T$ is the time horizon. Our result improves upon the previously best known bound of $O(\mathtt{L}(\mathcal{H}) \log^6 T)$ given by Daskalakis and Golowich (STOC 2022), and is optimal up to a universal constant for worst-case classes.
Chinese Translation
我们证明,使用随机规范化(proper)学习算法在线学习函数类 $\mathcal{H}$ 的最优期望错误边界为 $O(\mathtt{L}(\mathcal{H}) \log T)$,其中 $\mathtt{L}(\mathcal{H})$ 是 $\mathcal{H}$ 的 Littlestone 维数,$T$ 是时间范围。我们的结果改进了由 Daskalakis 和 Golowich(STOC 2022)给出的此前已知的最优边界 $O(\mathtt{L}(\mathcal{H}) \log^6 T)$,并且对于最坏情况下的函数类,该结果在普适常数范围内是最优的。
cs.LG / 43 / 2609.21450

Understanding LLM Quantization through Activation-Guided Compensation and Orthogonal Residuals

通过激活引导补偿与正交残差理解大语言模型量化
Narita, Yamato, Sato, Issei
Abstract
Post-training weight-activation quantization reduces the memory and inference costs of large language models, but aggressive W4A4 quantization remains difficult because activation outliers degrade effective quantization resolution. Although weight optimization, channel-wise scaling, and orthogonal rotation mitigate this problem, the error components they address and their relationship remain unclear. Using an exact decomposition of local weight-activation quantization error into an activation-guided weight compensation term and an orthogonal residual, we bound the residual using persistent channel-wise outlier and regular activation quantities. This decomposition clarifies which error components can be addressed by weight compensation and which require transformation design. We then use the residual bounds to derive practical guidelines for applying randomized Hadamard rotation, sign selection, and channel scaling. In particular, the analysis explains how random signs suppress constructive interference among persistent outlier channels, how sampling multiple sign patterns can improve transformation selection, and how second-moment balancing leads to an $L_2$ scaling rule while a further relaxation recovers SmoothQuant-style $L_\infty$ scaling. We evaluate these guidelines through backpropagation-free configurations across eight Llama and Mistral models, obtaining performance competitive with gradient-trained SpinQuant.
Chinese Translation
训练后权重-激活量化可降低大语言模型的内存和推理成本,但由于激活离群值会降低有效量化分辨率,激进的W4A4量化仍然困难。尽管权重优化、逐通道缩放和正交旋转可以缓解这一问题,但它们所处理的误差分量及其相互关系仍不清楚。通过将局部权重-激活量化误差精确分解为激活引导的权重补偿项和正交残差,我们利用持续性逐通道离群值和常规激活量对残差进行了界定。该分解明确了哪些误差分量可以通过权重补偿处理,哪些需要通过变换设计来解决。随后,我们利用残差界推导了应用随机Hadamard旋转、符号选择和通道缩放的实用准则。特别地,该分析解释了随机符号如何抑制持续性离群通道间的相长干涉、采样多个符号模式如何改进变换选择,以及二阶矩平衡如何导出$L_2$缩放规则,而进一步的松弛则恢复SmoothQuant风格的$L_\infty$缩放。我们在八个Llama和Mistral模型上通过免反向传播的配置评估了这些准则,取得了与基于梯度训练的SpinQuant相当的性能。
cs.LG / 44 / 2609.21457

Efficient Architecture Search under Leave-One-Subject-Out Evaluation

留一被试评估下的高效神经架构搜索
Hihn, Heinke, Schwenker, Friedhelm
Abstract
Deep neural architectures are widely used for signal processing in automated pain assessment systems. However, architecture design has remained largely a manual task despite the potential efficiency benefits of Neural Architecture Search (NAS). Embedding NAS in a Leave-One-Subject-Out (LOSO) evaluation is computationally demanding because a fully nested implementation requires $N$ independent architecture searches and, assuming approximately linear training cost, scales as $\mathcal{O}(N^2)$. We propose a block-based, leakage-controlled approach that shares NAS runs between subjects, reducing the number of searches from $N$ to $B$, where $B \ll N$, dubbed PainNAS. On the BioVid Heat Pain dataset, PainNAS yields comparable subject-level accuracy with substantially fewer parameters and FLOPs.
Chinese Translation
深度神经网络架构被广泛应用于自动化疼痛评估系统的信号处理中。然而,尽管神经架构搜索(Neural Architecture Search, NAS)具有潜在的效率优势,架构设计在很大程度上仍然是一项手动任务。将NAS嵌入留一被试(Leave-One-Subject-Out, LOSO)评估中计算开销巨大,因为完全嵌套的实现需要N次独立的架构搜索,并且假设训练成本近似线性,其复杂度按$\mathcal{O}(N^2)$扩展。我们提出了一种基于分块、控制信息泄漏的方法,可在被试之间共享NAS运行,将搜索次数从N减少到B(其中B远小于N),该方法被命名为PainNAS。在BioVid热痛(Heat Pain)数据集上,PainNAS以显著更少的参数量和浮点运算量(FLOPs)取得了相当的被试级准确率。
cs.LG / 45 / 2609.21523

What Must Survive? Exact Task-Information--State Frontiers for Resource-Sufficient Learning

什么必须保留?资源充足学习的精确任务信息—状态前沿
Katende, Ronald
Abstract
A system may be compressed before its downstream task is fully known. We ask how much retained state is then necessary and how much can be saved by limited advance task information. For a finite family of linear tasks, a task message is revealed before state formation and the exact task only afterwards. For an advice alphabet of size $K$, the exact frontier is \[ p^*(K)= \min_{\substack{\Pcal\text{ partition of }\U\\|\Pcal|\le K}} \max_{C\in\Pcal}\rank(T_C), \] with the $b$-bit frontier obtained by setting $K=\min(2^b,|\U|)$. Thus advance task information reduces state through partitions whose joint task operators have low rank. We also give an approximate singular-value frontier, a common-core lower bound and exact direct-sum law, and strong NP-hardness of finding an optimal advice partition. The hardness persists at every fixed positive approximation tolerance. Three examples illustrate the result. A well-conditioned softmax attention construction gives an exact $524{,}288\to1{,}024$ coordinate frontier when nine bits resolve one of $512$ continuations. A domain-decomposed digital twin yields an interface-plus-local-state law and a weighted partition problem for heterogeneous regions. A hierarchical multi-task model gives a two-stage frontier in which three bits reduce the required state from $3136$ to $448$ coordinates, with further task information approaching the irreducible $328$-coordinate single-task floor.
Chinese Translation
一个系统可能在其下游任务尚未完全确定之前就被压缩。我们因此提出两个问题:此时需要保留多少状态,以及有限的先验任务信息能够节省多少状态。对于一个有限的线性任务族,任务消息在状态形成之前被揭示,而精确任务信息仅在之后给出。当建议字母表大小为 $K$ 时,精确前沿为 \[ p^*(K)= \min_{\substack{\Pcal\text{ 为 }\U\text{ 的划分}\\|\Pcal|\le K}} \max_{C\in\Pcal}\rank(T_C), \] 其中 $b$ 比特前沿通过取 $K=\min(2^b,|\U|)$ 得到。因此,先验任务信息通过使其联合任务算子具有低秩的划分来降低状态需求。我们还给出了近似的奇异值前沿、一个公共核心下界与精确的直和定律,以及寻找最优建议划分的强NP困难性。该困难性在任何固定的正近似容差下依然成立。三个例子说明了这一结果:一个良态的 softmax 注意力构造在九比特用于区分 $512$ 个可能的后续之一时,给出了精确的 $524{,}288\to1{,}024$ 坐标前沿;一个区域分解的数字孪生给出了“接口加局部状态”定律以及针对异构区域的加权划分问题;一个层次化多任务模型给出了一个两阶段前沿,其中三比特将所需状态从 $3136$ 个坐标降至 $448$ 个坐标,而进一步的先验任务信息则逐渐逼近不可再减的 $328$ 坐标单任务下限。
cs.LG / 46 / 2609.21525

IncentRL: The Trade-Off Between Preference Guidance and Task Performance

IncentRL:偏好引导与任务性能之间的权衡
Wu, Xuening, Kang, Yanlan, Yin, Shenqin
Abstract
Preference-based reward shaping can guide reinforcement learning, but adding preference signals to the reward may unintentionally change the task being optimized. We address this problem with IncentRL, a framework that introduces preference guidance while explicitly characterizing its effect on external-task performance. IncentRL adds a Kullback--Leibler (KL) penalty between a specified outcome distribution and a preferred distribution. For finite discounted Markov decision processes with bounded shaping costs, we derive an external-value perturbation bound, establish a sufficient strict-action-gap condition for preserving the original optimal policy, and characterize the large-weight regime through discounted cumulative preference cost. Exact examples clarify the limits of these guarantees, including tied optima and support mismatch. We study a practical implementation using a hand-designed, distance-based outcome proxy, a fixed preference distribution, and score-weighted coefficient search. On MiniGrid DoorKey-8x8, the reported three-seed mean success rate after two million training steps reaches 98\% with coefficient 0.01, compared with 90.5\% for the reported zero-coefficient baseline, while the search progressively shifts toward smaller coefficients. Together, these results provide a principled view of the central trade-off in preference-based RL: using additional guidance to improve learning without excessively distorting the original task objective. The current experiments remain descriptive and do not yet isolate KL shaping from simpler alternatives.
Chinese Translation
基于偏好的奖励塑形可以引导强化学习,但将偏好信号加入奖励中可能会在无意中改变所优化的任务本身。我们提出 IncentRL 框架来应对这一问题,该框架在引入偏好引导的同时,显式刻画其对外部任务性能的影响。IncentRL 在指定的结果分布与偏好分布之间添加 Kullback--Leibler(KL)惩罚项。对于具有有界塑形代价的有限折扣马尔可夫决策过程,我们推导了外部价值扰动上界,建立了保持原最优策略的严格动作间隙充分条件,并通过折扣累积偏好代价刻画了大权重情形。精确算例阐明了这些保证的局限性,包括最优解并列以及支撑集不匹配的情况。我们研究了一种实用实现,采用手工设计的基于距离的结果代理、固定的偏好分布以及按分数加权的系数搜索。在 MiniGrid DoorKey-8x8 环境上,报告的三种子均值成功率在两百万训练步后达到 98%(系数为 0.01),而报告的零系数基线为 90.5%,同时搜索过程逐渐向更小的系数偏移。这些结果共同为基于偏好的强化学习中的核心权衡提供了有原则的视角:利用额外引导改善学习效果,同时不过度扭曲原始任务目标。目前的实验仍属描述性,尚未将 KL 塑形与更简单的替代方法分离比较。
cs.LG / 47 / 2609.21527

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

OpenMAS-GCom:面向图增强多智能体系统的诊断式基准测试
Yang, Kairui, Li, Xunkai, Zhang, Kaixiang, An, Minghao, Chen, Zekai, Ba, Yuxuan, Li, Rong-Hua
Abstract
Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, which determine how agents exchange information and divide responsibilities. However, final-score comparisons across systems combine differences in models, communication patterns, roles, and computation costs, making performance differences difficult to attribute to specific communication structures, role assignments, and information flows. To address this evaluation attribution problem, we introduce OpenMAS-GCom, a benchmark for diagnosing how these components affect G-MAS performance through controlled interventions. We represent systems through collaboration units, communication links, shared intermediate information, and execution rules. OpenMAS-GCom compares original systems with versions modified by changing one component while keeping tasks, models, prompts, and budget limits fixed. We rewire communication edges, remove specialist or critic agents, replace intermediate messages with incorrect content, and disable workers during execution. The benchmark evaluates 17 single-agent, ordinary multi-agent, and graph-enhanced configurations on 29 datasets across six domains. We add 400 G-MAS-Complex tasks requiring agents to combine information from multiple documents, resolve conflicting records, and return specified values with source identifiers. Experiments show larger mean losses after specialist removal than after critic removal, different performance degradation under incorrect messages and worker failures despite similar original scores, and different configurations achieving the highest accuracy and accuracy per token on G-MAS-Complex.
Chinese Translation
图增强多智能体系统通过通信图和角色分配来协调大语言模型智能体,二者决定了智能体之间如何交换信息以及如何划分职责。然而,跨系统的最终得分比较混杂了模型、通信模式、角色和计算成本等多方面的差异,使得性能差异难以归因于特定的通信结构、角色分配和信息流。为解决这一评估归因问题,我们提出了OpenMAS-GCom,一个通过受控干预来诊断这些组件如何影响G-MAS性能的基准测试。我们通过协作单元、通信链路、共享中间信息和执行规则来表征系统。OpenMAS-GCom将原始系统与修改了单一组件的版本进行比较,同时保持任务、模型、提示词和预算限制不变。我们重新连接通信边、移除专家型或评审型智能体、用错误内容替换中间消息,并在执行过程中禁用工作者智能体。该基准在六个领域的29个数据集上评估了17种单智能体、普通多智能体和图增强配置。我们新增了400个G-MAS-Complex任务,要求智能体整合来自多个文档的信息、解决相互冲突的记录,并返回带有来源标识符的指定值。实验表明:移除专家型智能体造成的平均损失大于移除评审型智能体;在原始得分相近的情况下,错误消息和工作者失效会导致不同的性能退化;以及在G-MAS-Complex上,不同配置分别取得最高准确率和最高单位令牌准确率。
cs.LG / 48 / 2609.21533

MACE: Memory-Agent Co-Evolution with Adaptive Memory Graphs for Multi-Agent Systems

MACE:面向多智能体系统的自适应记忆图记忆-智能体协同进化框架
Yang, Kairui, An, Minghao, Li, Xunkai, Yi, Ziheng, Chen, Zekai, He, Guangyuan, Li, Rong-Hua
Abstract
LLM-based multi-agent systems generate collaboration traces that record how agents plan tasks, verify intermediate results, and repair failures. Reusing these procedures requires preserving an action's prerequisites and the outputs needed by subsequent agents. Our empirical studies show that grouping these dependencies into functional memory units improves their retention, while connecting units increases retrieval of the units and links jointly required by a task. The preferred combination of units also changes between instructions and checklists, even when each combination's content is fixed across formats. Updating choices from the outcomes of each combination and format pairing outperforms scoring combinations and formats separately. These findings motivate MACE, a memory-agent co-evolution framework that adapts memory organization and agent memory use through execution feedback. Its MemGoG structure represents functional units as subgraphs of related conditions, actions, and outputs, connecting them through support, conflict, and repair relations. MACE Loop selects task-relevant units and relations within a memory budget and provides each agent with instructions or checklists for its current operation. It records the selected units, presentation formats, agent outputs, and task outcomes to update unit scores and relations for retrieval and inform subsequent presentation choices. Across eight benchmarks, MACE outperforms ten baselines with an average score of 81.11%, compared with 78.97% for the strongest baseline, SAGE.
Chinese Translation
基于大语言模型(LLM)的多智能体系统会生成协作轨迹,记录智能体如何规划任务、验证中间结果并修复失败。复用这些过程需要保留某个动作的前置条件以及后续智能体所需的输出。我们的实证研究表明,将这些依赖关系归组为功能性记忆单元能够提升其保留效果,而连接这些单元则能提高对任务所需的单元组合及其关联的检索准确率。即使在每种组合的内容在不同格式下保持不变的情况下,最优的单元组合也会随指令与清单(checklist)两种形式而变化。依据每种“组合-格式”配对的实际执行结果来更新选择,其效果优于分别对组合和格式进行独立评分。基于这些发现,我们提出了MACE,一个通过执行反馈来调整记忆组织方式和智能体记忆使用行为的记忆-智能体协同进化框架。其MemGoG结构将功能单元表示为由相关条件、动作和输出构成的子图,并通过支持、冲突和修复三种关系将其连接起来。MACE Loop在记忆预算内选取与任务相关的单元和关系,并为每个智能体在其当前操作中提供指令或清单。它记录所选单元、呈现格式、智能体输出和任务结果,以更新单元评分和用于检索的关系,并为后续的呈现选择提供依据。在八个基准测试中,MACE的平均得分为81.11%,优于十个基线方法,而最强的基线方法SAGE的得分为78.97%。
cs.LG / 49 / 2609.21550

OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios

OneBid:面向多样化oCPX广告场景的统一自动竞价基础模型
Li, Yewen, Jiang, Peng, Li, Yitian, Lv, Pengfei, Liu, Xialong, Jiang, Peng, Cai, Qingpeng
Abstract
Auto-bidding is central to computational advertising, where strategies must maximize advertisers' conversion value under economic constraints. It has evolved from rule-based controllers to reinforcement learning and generative methods such as Decision Transformer (DT). Yet these methods increasingly mismatch the prevailing optimized cost-per-X (oCPX) paradigm, which spans heterogeneous scenarios (e.g., registration, purchase), each served by a separate model, leading to fragmented pipelines and underexploring cross-scenario modeling. Inspired by foundation models like LLMs, unifying these oCPX scenarios into one model raises three challenges: multi-objective control, scalable capacity under strict latency, and safe offline policy improvement. We present OneBid, a unified auto-bidding foundation model that learns a reusable backbone from heterogeneous oCPX logs and adapts it to scenario-specific deployments via offline post-training. Building on DT, OneBid extends single Return-to-Go conditioning to two atomic signals, Return-to-Go for conversion value and Cost-to-Go for cost ratio, plus value-aware regularization on next-action prediction. To absorb distributional heterogeneity, we design a sequence-level Mixture-of-Experts architecture, where shared experts encode cross-scenario knowledge and sparsely-routed experts capture scenario-specific patterns at low latency, yielding consistent scaling with model size and data. During post-training, we align the backbone with scenario preferences via Critic-guided Relative Offline Policy optimization (CROP): a learned critic scores candidate actions group-relatively, avoiding the unsafe online exploration of GRPO-style fine-tuning while constraining policy shift to reduce OOD risk. Validated via online A/B tests and fully deployed at Kuaishou, OneBid delivers an overall +2.2% ADVV gain on oCPX Ads, peaking at +13.1% in the ROAS scenario.
Chinese Translation
自动竞价是计算广告的核心,其策略需在经济约束下最大化广告主的转化价值。该领域已从基于规则的控制器发展到强化学习和生成式方法(如Decision Transformer, DT)。然而,这些方法与当前主流的优化成本(oCPX)范式的匹配度日益下降——oCPX涵盖异构场景(如注册、购买),每个场景通常由独立模型服务,导致流水线碎片化,且缺乏跨场景联合建模的探索。受大语言模型(LLM)等基础模型的启发,将这些oCPX场景统一到一个模型中面临三大挑战:多目标控制、严格延迟约束下的可扩展容量,以及安全的离线策略改进。我们提出OneBid,一个统一的自动竞价基础模型,它从异构的oCPX日志中学习可复用的骨干网络,并通过离线后训练适配到各场景的部署中。在DT的基础上,OneBid将单一Return-to-Go条件扩展为两个原子信号:用于转化价值的Return-to-Go和用于成本比例的Cost-to-Go,并对下一动作预测施加价值感知的正则化。为吸收分布异构性,我们设计了序列级混合专家(Mixture-of-Experts)架构:共享专家编码跨场景知识,稀疏路由的专家以低延迟捕捉场景特定模式,从而实现随模型规模和数据的一致性扩展。在后训练阶段,我们通过Critic引导的相对离线策略优化(CROP)使骨干网络与场景偏好对齐:学习到的Critic以组内相对方式对候选动作评分,避免了GRPO式微调中不安全的在线探索,同时约束策略偏移以降低分布外(OOD)风险。经在线A/B测试验证并在快手全面部署,OneBid在oCPX广告上带来整体+2.2%的ADVV提升,在ROAS场景中峰值达到+13.1%。
cs.LG / 50 / 2609.21561

On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation

论排斥性与吸引性教师:在自蒸馏中分离正确性与行为
Baumann, Anton, Ashirmatov, Akmal, Schmidt-Traub, Leo, Lübeck, Frederike, Hübotter, Jonas, Buening, Thomas Kleine, Krause, Andreas
Abstract
On-policy self-distillation provides dense, token-level supervision by conditioning a model on privileged information and distilling the resulting teacher distribution back into the model. However, privileged information can change not only what the teacher knows, but also how it behaves, entangling correctness-relevant learning signals with unintended behavioral shifts. We study this effect in reasoning tasks by contrasting attractive self-distillation, which moves the model toward a privileged teacher, with repulsive self-distillation, which moves it away from a privileged teacher. We find that both objectives can induce strong and opposing behavioral shifts: attraction suppresses exploratory reasoning and promotes shorter, more confident responses, whereas repulsion increases response length, can trigger unintended switches into a model's latent thinking mode, and ultimately becomes unstable. Motivated by these observations, we study contrastive self-distillation, which combines attraction toward a correct-solution-conditioned teacher with repulsion from an incorrect-solution-conditioned teacher. In contrast to prior work that combines such distillation signals with a GRPO objective, we isolate the self-distillation objective and study its behavior on its own. We find that the shared behavioral shifts of the two teachers largely cancel, leaving a token-level signal that more directly reflects correctness. Across non-thinking, instruct-only, and already-thinking models, this contrastive objective improves reasoning performance while maintaining stable response lengths.
Chinese Translation
在策略自蒸馏(on-policy self-distillation)通过让模型以特权信息为条件,并将由此产生的教师分布蒸馏回模型本身,从而提供密集的、词元级别的监督信号。然而,特权信息不仅会改变教师所知道的内容,还会改变其行为方式,使得与正确性相关的学习信号与意外的行为偏移纠缠在一起。我们在推理任务中研究这一效应,通过对比两种方法:将模型推向特权教师的吸引性自蒸馏(attractive self-distillation),以及将模型推离特权教师的排斥性自蒸馏(repulsive self-distillation)。我们发现,这两种目标都可能引发强烈且相反的行为偏移:吸引性自蒸馏会抑制探索性推理,促使模型生成更短、更自信的回答;而排斥性自蒸馏会增加回答长度,可能触发模型意外切换到其潜在的思考模式,并最终变得不稳定。基于这些观察,我们研究了对比性自蒸馏(contrastive self-distillation),它将向以正确解为条件的教师的吸引与对以错误解为条件的教师的排斥相结合。与先前将此类蒸馏信号与 GRPO 目标相结合的工作不同,我们隔离了自蒸馏目标并单独研究其行为。我们发现,两个教师共同的行为偏移在很大程度上相互抵消,留下更直接反映正确性的词元级信号。在非思考型、仅指令型和已具备思考能力的模型上,这一对比性目标在保持回答长度稳定的同时,提升了推理性能。
cs.LG / 51 / 2609.21605

Trading Depth for Time in Recurrent Transformers

以时间换取深度的循环Transformer研究
Huang, Zeyi, He, Xuehai, Lee, Yong Jae, Shen, Yelong
Abstract
Recurrent Transformers increase computational depth through temporal recurrence, feeding each token's high-level hidden state into the computation of the next. This raises a natural question: is additional computation better spent on more temporal steps or greater physical depth? We investigate this question using Latent Recurrent Transformers (LRTs), which retain one backbone forward pass per vocabulary token during decoding and provide a controlled setting for comparing these two ways of adding computation. Specifically, we insert a latent thought token between consecutive vocabulary tokens. Each thought token passes through the same $L$ layers as a vocabulary token, sharing the backbone parameters and providing an additional stage of hidden-state refinement before predicting the next token. We compare this $L$-layer LRT against a $2L$-layer LRT without thought tokens. Both execute $2L$ Transformer blocks per vocabulary token during decoding, but the thought-token model uses fewer parameters. On 16- and 20-layer mixture-of-experts NanoChat backbones, one thought token brings the shallower model within 0.006 and 0.004 bits per byte of its double-depth counterpart, recovering 67% and 81% of the improvement with approximately 48% fewer total parameters. These results suggest that temporal thinking offers a parameter-efficient alternative to increasing physical depth in recurrent Transformers.
Chinese Translation
循环Transformer通过时间维度的循环来增加计算深度,即将每个词元的高层隐藏状态馈入下一个词元的计算中。这引出一个自然的问题:额外的计算是应该用于更多的时间步,还是更大的物理深度?我们使用潜在循环Transformer(Latent Recurrent Transformers, LRTs)来研究这一问题。该模型在解码时对每个词表词元仅保留一次骨干网络前向传播,为比较这两种增加计算的方式提供了受控环境。具体而言,我们在相邻的词表词元之间插入一个潜在思考词元。每个思考词元经过与词表词元相同的$L$层网络,共享骨干参数,并在预测下一个词元之前提供额外的隐藏状态精炼阶段。我们将这种$L$层的LRT与不含思考词元的$2L$层LRT进行比较。二者在解码时对每个词表词元均执行$2L$个Transformer块,但含思考词元的模型参数量更少。在16层和20层混合专家(mixture-of-experts)NanoChat骨干网络上,一个思考词元使较浅模型的每字节比特数与双倍深度模型的差距分别缩小至0.006和0.004,以约减少48%的总参数量恢复出67%和81%的性能提升。这些结果表明,在循环Transformer中,时间维度的思考是一种比增加物理深度更具参数效率的替代方案。
cs.LG / 52 / 2609.21647

Riemannian Neural Hamiltonian Flows: Geodesic Symplectic Transport and Interpretability

黎曼神经哈密顿流:测地线辛输运与可解释性
Souveton, Vincent
Abstract
Hamiltonian normalizing flows are attractive generative models because their phase-space maps are invertible and volume preserving, but most neural constructions are formulated in Euclidean space. We introduce Riemannian Neural Hamiltonian Flows, which combine the fixed kinetic energy of a Riemannian manifold, a learned scalar potential, and an explicit geodesic leapfrog integrator. Our analysis explains how the learned Hamiltonian can be made interpretable. Every normalizable potential defines an implicit profile, and the position marginal initially accelerates along the relative score between that profile and the base. The matched potential is the interpretable specialization for which the implicit profile is the target. In the isotropic Gaussian case, the mechanism corresponds to a phase-space rotation. A local harmonic analysis extends this result around each mode of a general target on a manifold. The gap between the learned and the matched potential is the sum of a residual memory of the base and a bias of the model, and the two potentials agree when the position base has been transferred to the momentum. This can be achieved when the former is broader than the target. Numerical experiments on Euclidean, hyperbolic, and spherical spaces show competitive sample quality and numerical cost against a Riemannian continuous normalizing flow, and confirm the interpretability of the learned potential.
Chinese Translation
哈密顿归一化流因其相空间映射可逆且保体积而成为颇具吸引力的生成模型,但大多数神经网络的构造都是在欧几里得空间中表述的。我们提出了黎曼神经哈密顿流(Riemannian Neural Hamiltonian Flows),它结合了黎曼流形的固定动能、一个可学习的标量势以及显式的测地线蛙跳积分器。我们的分析解释了如何使学习到的哈密顿量具备可解释性。每个可归一化的势都定义了一个隐式轮廓(implicit profile),且位置边缘分布最初沿该轮廓与基准分布之间的相对得分(relative score)方向加速。匹配势是一种可解释的特例,此时隐式轮廓即为目标分布。在各向同性高斯情形下,该机制对应于相空间中的一个旋转。局部谐波分析将这一结果推广到流形上一般目标分布的每个众数附近。学习势与匹配势之间的差距由基准分布的残余记忆与模型偏差两部分组成,当位置基准已被转移至动量时,两个势会一致。当位置基准比目标分布更宽泛时,这种转移是可以实现的。在欧几里得空间、双曲空间和球面上的数值实验表明,与黎曼连续归一化流(Riemannian continuous normalizing flow)相比,本方法在样本质量和数值开销上均具竞争力,并证实了学习势的可解释性。
cs.LG / 53 / 2609.21656

Beyond Gaussian Worlds: Latent Geometry Matters for JEPAs

超越高斯世界:潜在几何对JEPA至关重要
Nicollier, Léo, Meinhardt-Llopis, Enric, Pic, Marc, Musé, Pablo, Facciolo, Gabriele
Abstract
Recent Joint-Embedding Predictive Architectures (JEPAs) prevent representation collapse by constraining learned representations to follow a prescribed target distribution, such as an isotropic Gaussian or the uniform distribution on a hypersphere. Klindt et al. (2026) showed that, under their Euclidean assumptions, matching a Gaussian target can recover Gaussian latent variables up to a linear transformation, and that the Gaussian is the unique distribution with this guarantee. We extend their analysis to latent variables supported on embedded Riemannian manifolds and derive conditions on the latent geometry and positive-pair dynamics under which alignment and exact distribution matching guarantee linear recovery. In particular, when the latent variables are uniformly distributed on a sphere and the representations are matched to the same spherical distribution, every optimal representation recovers the latent state up to an orthogonal transformation. This shows that Gaussian uniqueness is not a universal property of distribution-matched JEPAs: non-Euclidean latent geometries can admit other linearly recoverable distributions. We further derive an approximate-recovery bound that is strictly tighter for the spherical world than for the Gaussian world. Experiments on Gaussian, spherical, and toroidal latent spaces show that geometrically compatible targets yield better linear recovery when optimization succeeds, whereas mismatched targets distort the latent structure. This advantage persists in high-dimensional Clifford-torus worlds.
Chinese Translation
近期的联合嵌入预测架构(Joint-Embedding Predictive Architectures,JEPAs)通过将学习到的表示约束为遵循指定的目标分布(如各向同性高斯分布或超球面上的均匀分布)来防止表示坍缩。Klindt等人(2026)证明,在欧氏假设下,匹配高斯目标分布可以在线性变换的意义下恢复高斯潜在变量,并且高斯分布是唯一具有这一保证的分布。我们将他们的分析扩展到支撑在嵌入黎曼流形上的潜在变量,并推导出潜在几何与正样本对动态所需满足的条件,在这些条件下,对齐与精确分布匹配能够保证线性恢复。特别地,当潜在变量均匀分布于球面且表示被匹配到相同的球面分布时,每个最优表示都能在正交变换的意义下恢复潜在状态。这表明高斯唯一性并非分布匹配JEPA的普适性质:非欧氏潜在几何可以容许其他线性可恢复的分布。我们进一步推导出一个近似恢复的界,该界在球面世界中严格优于高斯世界。在高斯、球面和环面潜在空间上的实验表明,当优化成功时,几何上相容的目标分布能带来更好的线性恢复效果,而失配的目标分布则会扭曲潜在结构。这一优势在高维Clifford环面世界中依然存在。
cs.LG / 54 / 2609.21664

Multi-Domain Clustering via Measure Quantization

基于测度量化的多域聚类方法
Eufrazio, Rafael Pereira, Montesuma, Eduardo Fernandes, Cavalcante, Charles Casimiro
Abstract
Clustering is a fundamental task in data analysis, typically addressed through centroid-based methods such as K-means. In this work, we present a general framework for multi-domain clustering via measure quantization: given samples from multiple domains, we learn a shared set of cluster prototypes by minimizing a probability metric, such as the Sinkhorn divergence or the Maximum Mean Discrepancy, between each domain's probability measure and the measure of prototypes. Data points are then assigned to clusters either via nearest centroid, or via optimal transport, a collaborative strategy that couples all samples within a domain. A mini-batch optimization strategy makes both fitting and assignment scalable, reducing memory and computational cost while preserving clustering performance. Experimental results on 5 multi-domain benchmarks spanning image, audio and sensor data show that our Sinkhorn-based method consistently outperforms classical and multi-domain clustering baselines, and that this advantage persists when scaling to hundreds of thousands of samples.
Chinese Translation
聚类是数据分析中的一项基础任务,通常通过以质心为中心的方法(如K-means)来解决。在本工作中,我们提出了一个基于测度量化的多域聚类通用框架:给定来自多个域的样本,我们通过最小化每个域的概率测度与原型测度之间的概率度量(如Sinkhorn散度或最大均值差异)来学习一组共享的聚类原型。随后,数据点可以通过最近质心或最优传输被分配到相应簇中,后者是一种将域内所有样本进行耦合的协作式策略。小批量(mini-batch)优化策略使拟合和分配过程均具备可扩展性,在保持聚类性能的同时降低了内存和计算成本。在涵盖图像、音频和传感器数据的5个多域基准数据集上的实验结果表明,我们基于Sinkhorn的方法始终优于经典聚类和多域聚类基线方法,并且当扩展到数十万样本规模时,这一优势依然保持。
cs.LG / 55 / 2609.21693

Optimization Geometry of Equivalent Brownian RKHS Representations

等价布朗RKHS表示的优化几何学
Mohammadigohari, Mahdi, Camps-Valls, Gustau
Abstract
Equivalent finite parameterizations can represent the same functions and intrinsic norm yet induce different optimization algorithms. We study this effect in a controlled finite Brownian RKHS with nodal, increment, and spectral coordinates. Classical finite-element, RKHS-interpolation, Brownian-covariance, and mixed-boundary DCT identities make the shared hypothesis class, Brownian energy, approximation operator, and coordinate maps explicit. Our main results concern the optimization geometry of this fixed model. With mapped initialization, identical scalar steps, and identical minibatches, nodal and spectral GD/SGD have exactly the same mapped trajectories. Increment GD is an explicit Euler step for the constant Brownian/Sobolev metric, with factor $1/h$. For Brownian-regularized least squares, $\kappa_2(\mathbf H_{\mathrm{inc}})\le1+A/\rho$, independently of grid resolution $G$ for fixed $A$, $\rho>0$, and the stated normalization. Under the stated standard-Adam convention, the universal orthogonal equivariance group is exactly the signed permutations; the block DCT-VIII transform is not one. Float64 tests over five grids numerically verify the finite identities, mapped one-layer and recursive trajectories, conditioning predictions, and theorem-matched Adam separation. Thus coordinate effects are isolated without changing the represented functions, intrinsic regularizer, or approximation space.
Chinese Translation
等价的有限参数化可以表示相同的函数和内禀范数,却会诱导出不同的优化算法。我们在一个受控的有限布朗RKHS中研究这一效应,该空间包含节点坐标、增量坐标和谱坐标三种坐标表示。经典的有限元、RKHS插值、布朗协方差以及混合边界DCT恒等式,使得共享的假设类、布朗能量、逼近算子和坐标映射得以显式表达。我们的主要结果关注这一固定模型的优化几何。在映射初始化、相同标量步长和相同小批量的条件下,节点坐标与谱坐标下的GD/SGD具有完全相同的映射轨迹。增量GD对应于常数布朗/Sobolev度量的显式欧拉步,其因子为$1/h$。对于布朗正则化最小二乘问题,有$\kappa_2(\mathbf H_{\mathrm{inc}})\le1+A/\rho$,即在固定的$A$、$\rho>0$及所述归一化条件下,条件数与网格分辨率$G$无关。在所述标准Adam约定下,普适的正交等变群恰好为带符号置换群;块DCT-VIII变换不在此群中。基于Float64精度的五个网格上的数值实验验证了上述有限恒等式、映射单层与递归轨迹、条件数预测以及与定理相符的Adam分离现象。因此,坐标效应得以被隔离研究,而不改变所表示的函数、内禀正则项或逼近空间。
cs.LG / 56 / 2609.21704

SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

SpecQuant:基于多父级量化的自适应LLM推理投机解码方法
KB, Harish, M, Jagadeeswaran, P, Pradheep, S, Yuvanesh, T, Sivakumar
Abstract
Running large language models (LLMs) locally continues to be limited by restrictions of compute and memory on consumer hardware. The popular acceleration technologies, such as quantization, speculative decoding, and adaptive inferencing, offer substantial speed boosts but usually necessitate retraining, per architecture tuning, or draft models. SpecQuant is a trainingfree framework, that combines speculative decoding with multiparent quantization to perform adaptive, efficient inference of LLMs. SpecQuant derives multiple quantized variants (INT4, FP8, FP16) from a shared base model, and dynamically routes queries based on predicted complexity; lightweight variants are used for simple or factual tasks, and full-precision models are used for complex reasoning tasks or long-context inputs. The shared-weight design of SpecQuant ensures sufficient token acceptance for speculative decoding without compatibility issues using separate draft parent models. We evaluate SpecQuant on Qwen2.5 based models on the MMLU, AlpacaEval, and GSM8K datasets, or benchmarks, demonstrating 35-43% speedups without degrading accuracy greater than 2%, substantial within the LLM community. SpecQuant enables practical on-device LLM deployment across diverse hardware without special infrastructure or expertise.
Chinese Translation
在消费级硬件上本地运行大语言模型(LLM)仍然受到计算和内存限制的制约。流行的加速技术,如量化、投机解码(speculative decoding)和自适应推理,能够带来显著的提速,但通常需要重新训练、针对特定架构的调优或草稿模型(draft models)。SpecQuant 是一个无需训练的框架,它将投机解码与多父级量化相结合,实现 LLM 的自适应高效推理。SpecQuant 从共享的基础模型中派生出多个量化变体(INT4、FP8、FP16),并根据预测的查询复杂度进行动态路由:轻量级变体用于简单或事实性任务,全精度模型则用于复杂推理任务或长上下文输入。SpecQuant 的共享权重设计确保了投机解码所需的足够 token 接受率,避免了使用独立草稿父模型时的兼容性问题。我们在基于 Qwen2.5 的模型上,于 MMLU、AlpacaEval 和 GSM8K 数据集(基准测试)上对 SpecQuant 进行了评估,结果显示其可实现 35-43% 的加速,且准确率下降不超过 2%,这在 LLM 社区中是相当可观的成果。SpecQuant 使 LLM 能够在无需特殊基础设施或专业知识的情况下,跨多样化硬件实现实用的端侧部署。
cs.LG / 57 / 2609.21735

GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

GEM-MPC:通过专家引导的规划平衡探索与利用
Serra-Gomez, Alvaro, Moerland, Thomas
Abstract
Effective exploration in high-dimensional continuous control remains a central challenge in reinforcement learning. Planning-based methods address this by combining online planning with learned policies and value functions, but their components can become misaligned during training: learned sampling policies may diverge from planner behavior, while planning distributions stored in replay become stale as the model and value function evolve. Reanalysis can refresh these targets, but at substantial computational cost. We propose GEM-MPC, an MPPI-based reinforcement learning method that improves the interaction between planning and learning. GEM-MPC uses MPPI to combine a policy trained to clone the planner with a KL-regularized policy that explores around it, providing complementary exploitation and guided exploration within planning. We further introduce Gated Prior Distillation, which selectively learns from stored planning distributions only when they provide a better target than the current prior, reducing the impact of stale planning data without requiring full reanalysis. Across continuous-control benchmarks, GEM-MPC consistently outperforms existing planning-based baselines under lower computational budgets.
Chinese Translation
在高维连续控制中实现有效探索仍然是强化学习的一项核心挑战。基于规划的方法通过将在线规划与学习到的策略和价值函数相结合来解决这一问题,但其各组件在训练过程中可能出现不一致:学习到的采样策略可能与规划器的行为产生偏差,而存储在经验回放中的规划分布会随着模型和价值函数的演化而变得过时。重新分析(reanalysis)可以刷新这些目标,但计算成本高昂。我们提出GEM-MPC,这是一种基于MPPI的强化学习方法,可改善规划与学习之间的交互。GEM-MPC利用MPPI将一个训练用于克隆规划器的策略与一个围绕其进行探索的KL正则化策略相结合,在规划过程中同时提供互补的利用行为和引导式探索。我们进一步引入了门控先验蒸馏(Gated Prior Distillation),仅在存储的规划分布比当前先验提供更好目标时才对其进行选择性学习,从而在无需完整重新分析的情况下降低过时规划数据的影响。在多个连续控制基准测试中,GEM-MPC在更低计算预算下始终优于现有的基于规划的基线方法。
cs.LG / 58 / 2609.21749

GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills

GraphSkillEvo:图结构智能体技能的进化优化
Sun, Rui, Zheng, Zhi, Wang, Zhenkun, Lu, Zhichao
Abstract
Skills can improve the performance of Large Language Model (LLM) agents by providing task-specific procedural guidance, while skill optimization further improves their effectiveness through iterative refinement. However, existing skill optimization methods typically represent skills as unstructured natural-language instructions, creating two key challenges: 1) Unstructured skills often lack explicit workflow-level guidance and contain substantial redundancy, making them difficult for LLMs to execute; 2) the vast search space of unconstrained natural-language skills makes skill optimization ineffective. To address these challenges, we propose representing skills as graph-structured natural-language artifacts. In graph-structured skills, each node represents an execution step together with its operational guidance, while directed edges encode context-dependent transitions between steps. Compared to unstructured skills, graph-structured skills can provide clear workflow-level guidance. Moreover, the proposed graph-structured skill can also facilitate skill optimization. Building on this structured representation, we introduce GraphSkillEvo, a population-based evolutionary optimization framework with mutation and crossover operators for graph-structured skills. By maintaining multiple candidate skills and combining effective components, GraphSkillEvo enables broader and more comprehensive exploration of the structured skill space than purely LLM-based iterative self-refinement. Extensive experiments across five agent benchmarks demonstrate that GraphSkillEvo consistently outperforms the strong skill optimization baseline SkillOpt, improving average accuracy by 4.01% on GPT-5.4-nano and 1.76% on GPT-5.4. Our code is available at https://github.com/ruisun7/GraphSkillEvo.
Chinese Translation
技能可以通过提供任务特定的程序性指导来提升大语言模型(LLM)智能体的性能,而技能优化则通过迭代式精炼进一步提升其有效性。然而,现有的技能优化方法通常将技能表示为非结构化的自然语言指令,这带来了两个关键挑战:1)非结构化技能往往缺乏明确的工作流级指导,且包含大量冗余,使得LLM难以执行;2)无约束自然语言技能的庞大搜索空间使得技能优化效率低下。为应对这些挑战,我们提出将技能表示为图结构的自然语言产物。在图结构技能中,每个节点表示一个执行步骤及其操作指导,而有向边编码了步骤之间依赖上下文的转移关系。与非结构化技能相比,图结构技能能够提供清晰的工作流级指导。此外,所提出的图结构技能也有助于技能优化。基于这一结构化表示,我们提出了GraphSkillEvo,一个基于种群的进化优化框架,包含针对图结构技能的变异和交叉算子。通过维护多个候选技能并组合有效的组件,GraphSkillEvo相较于纯基于LLM的迭代式自我精炼,能够对结构化技能空间进行更广泛、更全面的探索。在五个智能体基准上的大量实验表明,GraphSkillEvo始终优于强大的技能优化基线SkillOpt,在GPT-5.4-nano上平均准确率提升4.01%,在GPT-5.4上提升1.76%。我们的代码已发布于 https://github.com/ruisun7/GraphSkillEvo。
cs.LG / 59 / 2609.21758

Bilevel Optimization of Topology and Hyperparameters (BOTH)

拓扑与超参数的双层优化(BOTH)
Sanu, Suryanarayanan Manoj, Bessa, Miguel Anibal, Aragón, Alejandro Marcos
Abstract
Topology optimization (TO) represents a significant step towards automating the design process: given a working simulation, TO can produce a viable prototype at the press of a button by differentiating the simulation and iteratively improving the design. In practice, however, TO is riddled with ``magic numbers''---hyperparameters whose tuning significantly affects the outcome. Finding the right values typically requires not only deep problem-specific knowledge but also extensive trial-and-error. While practitioners can use surrogate-assisted hyperparameter optimization as an alternative, this approach requires strictly limiting the number of hyperparameters through careful problem formulation. Here, we propose differentiating TO itself using automatic differentiation. This yields ``hypergradients'' that allow us to tune these hyperparameters in tandem with the primary optimization. We show that evaluating just one or two steps of TO is sufficiently informative and that the method scales favorably to thousands of hyperparameters at an expense comparable to only a few standard TO runs. We demonstrate this approach on stress-constrained and compliance problems, with the latter utilizing a neural parameterization of the density field.
Chinese Translation
拓扑优化(Topology Optimization, TO)是实现设计过程自动化的重要一步:给定一个可运行的仿真,拓扑优化只需一键即可通过微分仿真并迭代改进设计,生成可行的原型。然而在实践中,拓扑优化充满了“魔法数字”——即其调参会显著影响结果的超参数。找到合适的取值通常不仅需要深入的领域知识,还需要大量的反复试错。虽然从业者可以使用代理模型辅助的超参数优化作为替代方案,但该方法需要通过精心的问题表述来严格限制超参数的数量。本文提出利用自动微分对拓扑优化本身进行微分,由此得到的“超梯度”(hypergradients)使我们能够在主优化的同时对超参数进行调优。我们证明,仅评估一到两步拓扑优化就足以提供充分的信息,且该方法可良好地扩展到数千个超参数,其开销仅相当于几次标准拓扑优化运行。我们在应力约束问题和柔度问题上演示了该方法,其中后者采用了密度场的神经参数化。
cs.LG / 60 / 2609.21791

RegKT: Interpretable and Robust Deep Knowledge Tracing With IRT-Regularizer

RegKT:基于IRT正则化器的可解释且鲁棒的深度知识追踪模型
Girard, Samuel, Pinto, Juan D., Vie, Jill-Jênn, Bouzeghoub, Amel
Abstract
As deep learning models continue to advance, knowledge tracing models have achieved higher accuracy. However, these gains come at the cost of reduced interpretability, which is crucial for practitioners in educational settings to adopt new methodologies. Additionally, deep learning models are prone to overfitting, particularly when dealing with the small datasets that are common in educational applications. In this paper, we propose a novel regularization technique designed to enhance the robustness of deep-learning-based knowledge tracing models, while simultaneously improving their interpretability. Our method addresses both the interpretability and overfitting challenges, making it more feasible for real-world educational applications.
Chinese Translation
随着深度学习模型的不断发展,知识追踪模型的准确率不断提升。然而,这些进步是以可解释性的降低为代价的,而可解释性对于教育领域从业者采纳新方法至关重要。此外,深度学习模型容易过拟合,尤其是在处理教育应用中常见的小规模数据集时。本文提出了一种新颖的正则化技术,旨在增强基于深度学习的知识追踪模型的鲁棒性,同时提升其可解释性。我们的方法同时解决了可解释性和过拟合两大挑战,使其在实际教育应用中更具可行性。
cs.LG / 61 / 2609.21815

Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods

矩阵AdaGrad:基于行与列的自适应次梯度方法
Zhang, Wenpeng, Yu, Runsheng, Zhao, Peilin
Abstract
Adaptive optimization methods such as AdaGrad and Adam are widely used in modern neural-network training, but their adaptive scaling is primarily designed for vector-valued parameters and does not explicitly exploit matrix structure. Recent matrix-aware optimizers demonstrate the benefits of structured optimization, yet a general theoretical framework for deriving matrix-aware adaptivity comparable to that of AdaGrad remains lacking. In this work, we develop a general Online Mirror Descent framework with adaptive proximal functions for matrix-valued parameters, providing a principled approach to deriving matrix-aware adaptive optimization through online regret minimization. By introducing row-wise and column-wise matrix proximal functions and analyzing the resulting regret trade-off, we derive Row-wise Matrix AdaGrad (Row-AdaGrad) and Column-wise Matrix AdaGrad (Column-AdaGrad), with adaptive scaling determined by the accumulated row-wise or column-wise gradient norms. We establish regret guarantees and show that these matrix-aware bounds can be strictly tighter than those of entry-wise AdaGrad under structured gradients. Experiments on matrix factorization and deep neural-network training further demonstrate the benefits of aligning adaptive scaling with matrix structure, including improved optimization stability and trainability at larger learning rates and greater network depths.
Chinese Translation
AdaGrad和Adam等自适应优化方法在现代神经网络训练中被广泛使用,但其自适应缩放主要针对向量形式的参数设计,未能显式利用矩阵结构。近年来的一些矩阵感知优化器展示了结构化优化的优势,然而目前仍缺乏一个可用于推导与AdaGrad相当的矩阵感知自适应性的通用理论框架。在本工作中,我们提出了一个面向矩阵值参数、具有自适应近端函数的通用在线镜像下降(Online Mirror Descent)框架,为通过在线遗憾最小化推导矩阵感知自适应优化提供了原理性的方法。通过引入行方向和列方向的矩阵近端函数,并分析由此产生的遗憾权衡,我们推导出了行自适应矩阵AdaGrad(Row-wise Matrix AdaGrad,Row-AdaGrad)和列自适应矩阵AdaGrad(Column-wise Matrix AdaGrad,Column-AdaGrad),其自适应缩放由累积的行方向或列方向梯度范数决定。我们建立了遗憾保证,并证明在结构化梯度下,这些矩阵感知的遗憾界可以严格优于逐元素AdaGrad的遗憾界。在矩阵分解和深度神经网络训练上的实验进一步表明,使自适应缩放与矩阵结构对齐具有诸多优势,包括在更大学习率和更深网络下提升优化的稳定性与可训练性。
cs.LG / 62 / 2609.21829

Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data

面向高维异构数据的联邦深度聚类网络
Stallmann, Morris, Kouzinopoulos, Charalampos S., Pietrasik, Marcin, Wilbik, Anna
Abstract
Clustering high-dimensional data is a fundamental task in unsupervised machine learning with applications to a variety of domains. In the centralized data scenario, this task is commonly solved using deep clustering methods that utilize deep neural network architectures to learn clustering-friendly latent space representations. In Federated Learning, where data is distributed between clients and is private, deep clustering methods are less explored. In particular, recently introduced federated deep clustering methods, despite showing very promising performance, still fall short in reliably providing good performance if data across clients are non-identically-independently distributed. In this work, we introduce a generalization of Deep Clustering Networks to the federated scenario, named FedDCN, that simultaneously optimizes a reconstruction loss and a clustering loss. To ensure robustness and latent space alignment in non-identically-independently distributed data scenarios, FedDCN generates synthetic data augmentations, and its learning objective includes a geometric regularization for latent space alignment. Through experimental evaluation, the effectiveness of the approach under IID and non-IID assumptions is demonstrated, and future research directions are identified.
Chinese Translation
高维数据聚类是无监督机器学习中的一项基础任务,在众多领域均有应用。在集中式数据场景下,该任务通常采用深度聚类方法解决,即利用深度神经网络架构学习有利于聚类的潜在空间表示。而在联邦学习中,数据分布在各客户端且具有私密性,深度聚类方法的探索相对较少。特别是,近期提出的联邦深度聚类方法尽管展现了非常有前景的性能,但在客户端数据非独立同分布(non-IID)的情况下,仍难以可靠地提供良好的性能。在本工作中,我们将深度聚类网络(Deep Clustering Networks)推广至联邦场景,提出了名为 FedDCN 的方法,该方法同时优化重构损失和聚类损失。为确保在非独立同分布数据场景下的鲁棒性和潜在空间对齐,FedDCN 生成合成数据增强,并且其学习目标中包含用于潜在空间对齐的几何正则化项。通过实验评估,验证了该方法在 IID 和 non-IID 假设下的有效性,并指出了未来的研究方向。
cs.LG / 63 / 2609.21849

The Weight Is Over - Interactive Diffusion on Consumer GPUs

权重已定——消费级GPU上的交互式扩散模型
Ganz, Frieder, Müller, Maximilian
Abstract
On-device inference is booming, but the momentum is almost all in language models. Diffusion pipelines are memory hungry, latency-sensitive, and require orchestrating an embedder, a transformer, a decoder, and often further postprocessing that is not as standardized as LLM inference loops are. We navigate the trade-off between performance, quality, and model footprint to reach as many client devices in the wild as possible. We make three contributions: an embedding translator that maps a small text encoder into a large encoder space to cut weight and latency; a reproducible sweep recipe for navigating the speed/quality/memory triangle in diffusion pipelines; and an interactive on-device image generation editor achieving sub-second TTFI on recent GPUs.
Chinese Translation
端侧推理正在蓬勃发展,但其发展势头几乎全部集中在语言模型上。扩散模型管线对内存需求高、延迟敏感,且需要协调一个嵌入器、一个Transformer、一个解码器,以及通常还需进一步的后处理,而这些都不像大语言模型(LLM)推理循环那样标准化。我们在性能、质量与模型规模之间权衡取舍,以尽可能覆盖更多实际使用中的客户端设备。我们做出了三项贡献:一个嵌入翻译器,可将小型文本编码器映射到大型编码器空间,从而降低权重开销和延迟;一套可复现的调优方案,用于在扩散管线中权衡速度、质量与内存的三角关系;以及一个端侧交互式图像生成编辑器,可在较新的GPU上实现亚秒级的首次图像生成时间(TTFI)。
cs.LG / 64 / 2609.21870

Neural Cellular Automata Learn General Features in their Hidden Channels

神经细胞自动机在隐藏通道中学习通用特征
Guichard, Etienne, Nichele, Stefano
Abstract
Modern deep learning models achieve impressive generalization through over-parameterization, but this paradigm often struggles with overfitting and memorization in few-shot regimes. Neural Cellular Automata (NCAs) offer a highly parameter-efficient alternative, yet research has focused primarily on their output, leaving the role of their internal hidden channels largely unexplored. In this paper, we investigate the internal dynamics of NCA hidden channels and introduce a novel transfer-learning mechanism that injects a pretrained teacher's hidden states into a student model to guide early optimization. Evaluated on few-shot and scale-variant MNIST benchmarks, NCAs outperform comparable recurrent and feed-forward architectures, demonstrating superior generalization with a minimal parameter budget (~9,800 parameters). Mechanistic analysis reveals that the hidden channels decouple feature extraction from uniform classification consensus by absorbing morphological complexity and converging to mutually orthogonal states. Furthermore, we demonstrate that these hidden channels capture general, scale-invariant topological primitives rather than class-specific templates. This allows a student model to achieve strong few-shot performance on unseen classes using features transferred from a teacher trained only on a subset of digits (0-5). Our results highlight the potential of utilizing hidden-state dynamics as a robust, decentralized computational substrate for parameter-efficient transfer learning
Chinese Translation
现代深度学习模型通过过度参数化实现了令人印象深刻的泛化能力,但这一范式在少样本(few-shot)场景中常常面临过拟合和记忆化问题。神经细胞自动机(Neural Cellular Automata, NCA)提供了一种高度参数高效替代方案,然而现有研究主要关注其输出,对其内部隐藏通道的作用 largely 未加探索。本文研究了 NCA 隐藏通道的内部动态,并提出了一种新颖的迁移学习机制:将预训练教师模型的隐藏状态注入学生模型,以引导早期优化。在少样本和尺度变化(scale-variant)的 MNIST 基准上的评估表明,NCA 优于可比较的循环和前馈架构,在极小的参数预算(约 9,800 个参数)下展现出卓越的泛化能力。机制分析揭示,隐藏通道通过吸收形态复杂性并收敛到相互正交的状态,将特征提取与统一的分类共识解耦。此外,我们证明这些隐藏通道捕获的是通用的、尺度不变的拓扑基元,而非类别特定的模板。这使得学生模型能够利用从仅在数字 0-5 的子集上训练的教师模型迁移而来的特征,在未见过的类别上取得强大的少样本性能。我们的研究结果凸显了将隐藏状态动态用作参数高效迁移学习的稳健、去中心化计算基底的潜力。
cs.LG / 65 / 2609.21876

Geometric Mean Pooling for Equal-Weight Multiplicative Coarse-Graining

面向等权乘性粗粒化的几何平均池化
Wu, Ang-Kun, Wen, Fangdi, Zhang, Jingtao
Abstract
As an alternative to the additive and extremal biases of average and max pooling, we introduce Geometric Mean Pooling (GMP), a signed pooling operator that combines the product of feature signs with the geometric mean of feature magnitudes. Motivated by local-to-global composition in quantum many-body physics, GMP retains both joint sign information and a characteristic multiplicative scale without introducing learnable pooling parameters. We show that non-overlapping hierarchical GMP preserves the corresponding global multiplicative statistic and evaluate it on synthetic sequence tasks, iterative coarse-graining, image classification, and molecular lipophilicity regression. On the synthetic tasks, GMP recovers product-based signals more accurately than average and max pooling and maintains predictive performance under the tested levels of multiplicative input noise. On image and molecular data, however, its effectiveness depends on the representation, target parameterization, and placement of local and global pooling. These results position GMP as a complementary, regime-dependent inductive bias for tasks in which equal-weight multiplicative composition is plausible, rather than as a universal replacement for standard pooling operators.
Chinese Translation
作为平均池化与最大池化所固有的加性偏差和极值偏差的一种替代方案,我们提出了几何平均池化(Geometric Mean Pooling, GMP),这是一种有符号的池化算子,它将特征符号的乘积与特征幅值的几何平均相结合。受量子多体物理中从局域到全局组合方式的启发,GMP 在不引入可学习池化参数的情况下,同时保留了联合符号信息和特征性的乘性尺度。我们证明,非重叠的层次化 GMP 能够保持相应的全局乘性统计量,并在合成序列任务、迭代粗粒化、图像分类以及分子亲脂性回归上对其进行了评估。在合成任务上,GMP 比平均池化和最大池化更准确地恢复了基于乘积的信号,并在所测试的乘性输入噪声水平下保持预测性能。然而,在图像和分子数据上,其有效性取决于表示方式、目标参数化以及局域池化与全局池化的位置。这些结果表明,GMP 是一种互补的、依赖于适用情形的归纳偏置,适用于等权乘性组合合理的任务,而非标准池化算子的通用替代品。
cs.LG / 66 / 2609.21888

Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective

从自由能视角检测大语言模型的预训练数据
Ke, Chenye, Liu, Zirui, Liu, Qi, Zhuang, Yan, Zhang, Jintao, Huang, Zhenya, Wang, Shijin
Abstract
Detecting pretraining data in large language models is challenging because high likelihood can reflect either training exposure or strong generalization. In the joint space of prediction loss and predictive entropy, a likelihood-only detector uses a horizontal boundary and can mistake predictable non-members for members. Motivated by this, we introduce an inclined boundary that evaluates prediction loss relative to predictive entropy. Our analysis shows that entropy correction can preserve the expected membership signal while reducing its variance, thereby improving standardized member--non-member separation. We further extend the mean--variance analysis to the more general setting with a nonzero mean entropy gap. Interestingly, this entropy-adjusted score admits a Helmholtz free-energy interpretation, leading to Energy Transfer Detection (ETD), which views pretraining data detection from a macroscopic residual free-energy transfer perspective. Extensive experiments show that ETD achieves the best average detection performance, improving average AUROC by up to 3.5\% and TPR@5\%FPR by up to 5.1\%, while remaining robust across diverse settings.
Chinese Translation
检测大语言模型的预训练数据极具挑战性,因为高似然既可能反映训练时见过该数据,也可能源于强大的泛化能力。在预测损失与预测熵的联合空间中,仅基于似然的检测器使用水平边界,容易将可预测的非成员样本误判为成员样本。受此启发,我们引入一种倾斜边界,用于相对于预测熵来评估预测损失。我们的分析表明,熵校正能够在保留期望成员信号的同时降低其方差,从而提升成员与非成员之间的标准化分离度。我们进一步将均值-方差分析扩展到具有非零平均熵差的更一般设定。有趣的是,这种经熵调整的分数可从亥姆霍兹自由能(Helmholtz free-energy)角度加以解释,由此提出了能量迁移检测(Energy Transfer Detection, ETD)方法,从宏观残余自由能迁移的视角看待预训练数据检测问题。大量实验表明,ETD 取得了最佳的 평均检测性能,平均 AUROC 最高提升 3.5%,TPR@5%FPR 最高提升 5.1%,并且在多种设置下均保持稳健。
cs.LG / 67 / 2609.21894

LLMs as Feature Engineers for Text-and-Tabular Prediction

将大语言模型(LLMs)作为文本-表格预测的特征工程工具
Barlier, Merwan, Skrlj, Blaz
Abstract
We introduce an iterative framework that automates the extraction of interpretable, schema-bound categorical features from unstructured text for tabular prediction models. To navigate the feature space, a generator LLM proposes semantic definitions, a separate extractor LLM materializes the features, and a downstream tabular model evaluates their predictive performance. We optimize this search by translating explicit model errors, such as AUC ranking inversions, into natural-language feedback, steering the LLM to resolve specific predictive failures. Evaluated across three public datasets, this error-driven loop accelerates feature discovery by up to $3\times$ compared to unguided search. Empirically, the generated features demonstrate strong multi-view complementarity, strictly outperforming any subset when combined with TF-IDF and dense embeddings. Finally, the framework guarantees instance-level interpretability: the discovered features dominate SHAP importance rankings and provide a fully transparent, semantic audit trail for every prediction.
Chinese Translation
我们提出了一个迭代式框架,用于从非结构化文本中自动提取可解释的、符合数据模式约束的分类特征,以供表格预测模型使用。为了在特征空间中进行搜索,一个生成器大语言模型(LLM)提出语义定义,另一个独立的提取器LLM将特征具体化,下游的表格模型则评估这些特征的预测性能。我们通过将显式的模型错误(如AUC排名反转)转化为自然语言反馈来优化这一搜索过程,引导LLM解决具体的预测失效问题。在三个公开数据集上的评估表明,这种误差驱动的循环与无引导搜索相比,可将特征发现速度提升最多3倍。实验证明,所生成的特征具有强多视图互补性,与TF-IDF和稠密嵌入(dense embeddings)结合时,其性能严格优于任何子集。最后,该框架保证了实例级可解释性:所发现的特征在SHAP重要性排名中占主导地位,并为每一次预测提供了完全透明、可语义追溯的审计线索。
cs.LG / 68 / 2609.21899

ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment

ExpBoN:用于高效测试时大语言模型对齐的指数噪声Best-of-$n$方法
Liu, Yanxiao, Wan, Sicheng, Gündüz, Deniz
Abstract
Best-of-$n$ (BoN) sampling is a simple yet effective inference-time alignment method, but hard maximization provides only coarse control over the trade-off between reward and distribution shift. Soft Best-of-$n$ (Verdun et al. 2025) provides smoother control and converges to the optimal distribution associated with KL-regularized reward maximization. In this paper, we introduce ExpBoN, an alternative soft BoN method based on the exponential-noise report-noisy-max mechanism. It admits an exact finite-$n$ decomposition, which yields exponentially fast convergence in total variation, expected reward, and both directions of KL divergence. We provide comprehensive theoretical analyses of its convergence and regret behavior. We further integrate ExpBoN into the guided speculative inference (GSI) framework (Geuter, Mroueh, and AlvarezMelis 2025), resulting in ExpGSI, for efficient reward-guided LLM alignment. ExpGSI yields substantial reductions in computational cost while maintaining comparable accuracy. Experiments on MATH500, MMLU-STEM, and Minerva Math with the Qwen2.5-Math and Qwen3 model families show that ExpGSI reduces estimated computation by $14\%$-$39\%$ across candidate budgets for Qwen2.5-Math and by up to $45\%$ at $n=16$ for Qwen3. Overall, our results provide a theoretical and algorithmic foundation for exponential-noise BoN and efficient test-time LLM alignment.
Chinese Translation
Best-of-$n$(BoN)采样是一种简单而有效的推理时对齐方法,但硬最大化只能对奖励与分布偏移之间的权衡提供粗粒度的控制。软Best-of-$n$(Soft Best-of-$n$,Verdun等人,2025)提供了更平滑的控制,并收敛于与KL正则化奖励最大化相关联的最优分布。本文提出ExpBoN,一种基于指数噪声报告噪声最大(exponential-noise report-noisy-max)机制的替代性软BoN方法。该方法具有精确的有限$n$分解,从而在总变差、期望奖励以及KL散度的两个方向上均实现指数级快速收敛。我们对其收敛性和遗憾(regret)行为进行了全面的理论分析。我们进一步将ExpBoN整合到引导式推测推断(guided speculative inference,GSI)框架(Geuter、Mroueh和AlvarezMelis,2025)中,得到ExpGSI,用于高效的奖励引导大语言模型对齐。ExpGSI在保持相当精度的同时,显著降低了计算成本。在Qwen2.5-Math和Qwen3模型家族上针对MATH500、MMLU-STEM和Minerva Math的实验表明,ExpGSI在Qwen2.5-Math上于各候选预算下可减少14%至39%的估计计算量,在Qwen3上于$n=16$时可减少高达45%的计算量。总体而言,我们的结果为指数噪声BoN及高效的测试时大语言模型对齐提供了理论和算法基础。
cs.LG / 69 / 2609.21906

Intervention Granularity Matters: Coherent Treatment Bundles in Counterfactual Simulation with Clinical World Models

干预粒度至关重要:基于临床世界模型的反事实模拟中的连贯治疗组合
Wang, Fangzhou, Yang, Yixuan, Balzarotti, Camilla, Kamaleswaran, Rishikesan
Abstract
Counterfactual simulation with a clinical world model means fixing a patient's history, changing the treatment, and reading off the predicted response. Doing so requires deciding what counts as one intervention. In clinical settings, interventions are documented as bundles: a co-occurrence audit of 945,707 patient-hours from MIMIC-IV shows groups of components, such as every parameter of a dialysis circuit, that never appear apart, so an edit that changes one component on its own describes an hour that never occurs in the data. We hypothesize that the granularity at which an intervention is edited changes how a world model responds, and test this with Clin-JEPA, a latent world model of patient trajectories conditioned on hourly treatment text. At 1,019 documented onsets of invasive ventilation, we keep the patient's history and other treatments fixed and compare editing one ventilator setting with editing the complete configuration recorded for a real patient with the most similar recent trajectory. The complete bundle moves the predicted next state further than any single setting, consistently across all five settings, and the difference remains after accounting for how much each edit changes the model's input. Intervention granularity therefore materially affects the response of a clinical world model: single-component edits may understate treatment sensitivity, and bundle-aware editing may offer a better-supported basis for counterfactual treatment simulation.
Chinese Translation
利用临床世界模型进行反事实模拟,是指固定患者的病史,改变治疗方案,并读取模型预测的反应。这一过程需要决定什么构成一次干预。在临床环境中,干预措施以组合(bundle)的形式被记录:对来自 MIMIC-IV 数据集 945,707 个患者小时的共现审计显示,某些组件(例如透析回路的所有参数)总是成组出现、从不单独出现,因此仅修改单个组件的编辑所描述的状态在数据中从未出现。我们假设干预编辑的粒度会改变世界模型的响应,并使用 Clin-JEPA——一个以每小时治疗文本为条件的患者轨迹潜在世界模型——来检验这一假设。在 1,019 次有创通气的记录起始时刻,我们固定患者的病史和其他治疗,比较仅编辑单个呼吸机参数与编辑为近期轨迹最相似的真实患者所记录的完整配置的结果。完整的治疗组合比任何单一参数都更能显著改变预测的下一状态,且这一规律在全部五项参数中均一致成立,并在控制每次编辑对模型输入的改变程度后依然存在。因此,干预粒度会实质性地影响临床世界模型的响应:单组件编辑可能低估治疗敏感性,而基于组合的编辑可能为反事实治疗模拟提供更可靠的依据。
cs.LG / 70 / 2609.21909

Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning

超越运动学:面向肌肉驱动模仿学习的仿真保真度基准测试
Ahmad, Ayah G., Borden, Claire E., Tucker, Maegan
Abstract
In this work, we conduct a systematic comparison of two state-of-the-art motion-imitation reinforcement learning (MIRL) pipelines, one built on SCONE/HyFyDy and one built on MuJoCo/MyoSim. HyFyDy emphasizes physiological realism through detailed musculotendon modeling, while MuJoCo prioritizes computational efficiency and scalable policy learning. While recent work has demonstrated that both pipelines reproduce human kinematics with high fidelity, it remains unclear if they accurately capture the underlying neuromuscular behavior that produced the movement. This limitation is particularly important for robotic assistive-device design and control, where outcome measures such as muscle activation patterns and metabolic cost are often used as optimization targets. To conduct a systematic comparison, our work compares both pipelines using a common set of human motion-capture and electromyography (EMG) measurements. The results find that while both pipelines produce similar kinematics with relative accuracy, the muscle activations from HyFyDy are more aligned with the experimental EMG, as supported by the average pooled (RMSE, r) values for muscle activations from HyFyDy and MuJoCo: (0.164, 0.4) and (0.344, 0.11), respectively. While we conclude that the more advanced physiological realism of HyFyDy currently makes it more suitable for musculoskeletal modeling, both require further development to bring physiological realism to GPU-parallelizable simulation environments and advance robotic assistive device design.
Chinese Translation
在本工作中,我们对两种最先进的运动模仿强化学习(MIRL)流程进行了系统性比较,其中一种基于 SCONE/HyFyDy 构建,另一种基于 MuJoCo/MyoSim 构建。HyFyDy 通过详细的肌肉肌腱建模强调生理真实性,而 MuJoCo 则优先考虑计算效率和可扩展的策略学习。尽管近期的研究工作已证明这两种流程都能高保真地复现人体运动学,但它们能否准确捕捉产生该运动的底层神经肌肉行为仍不清楚。这一局限性对于机器人辅助设备的设计与控制尤为重要,因为肌肉激活模式和代谢消耗等结果指标常被用作优化目标。为了进行系统性比较,我们的工作使用一套共同的人体动作捕捉和肌电图(EMG)测量数据对两种流程进行了比较。结果表明,虽然两种流程都能以相当的精度产生相似的运动学结果,但 HyFyDy 的肌肉激活与实验 EMG 数据更为吻合,HyFyDy 和 MuJoCo 肌肉激活的平均汇总(RMSE, r)值分别为 (0.164, 0.4) 和 (0.344, 0.11),这证实了上述结论。我们的结论是,HyFyDy 更为先进的生理真实性目前使其更适合用于肌肉骨骼建模,但两者都需要进一步发展,以将生理真实性引入可 GPU 并行化的仿真环境,并推动机器人辅助设备设计的进步。
cs.LG / 71 / 2609.21926

Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources

折点与平滑性:类拉普拉斯源下实解析nICA的可辨识性
Manring, Isaac, Huang, Kejun
Abstract
Many machine learning systems try to explain complex data - like images or financial time series - in terms of hidden, independent factors that generated them. Recovering the true underlying factors, rather than some scrambled version of them, is the central challenge of nonlinear Independent Component Analysis (nICA). We prove identifiability (exact recovery) up to trivial ambiguities for real analytic generating functions when source probability density functions have a finite number of discontinuities in the first derivative. The Laplace distribution is the most prominent example satisfying this assumption. Our proof relies on the contrast between kinks in the source distribution and the smoothness of real analytic functions. Real analytic functions comprise a broad class of generating mechanisms, and can be approximated with Normalizing Flows or Variational Autoencoders with standard activation functions (e.g., tanh, softplus, GELU), so our result applies with minimal changes to existing training pipelines. We perform experiments on real and synthetic data with both Normalizing Flows and Variational Auto-Encoders demonstrating their identifiability properties. In experiments on CelebA data we recover several interpretable latent factors controlling unique attributes across the dataset.
Chinese Translation
许多机器学习系统试图用生成数据的隐含独立因子来解释复杂数据,例如图像或金融时间序列。恢复真实的潜在因子(而非其某种混合形式)是非线性独立成分分析(nICA)的核心挑战。我们证明,当源概率密度函数的一阶导数具有有限个不连续点时,对于实解析的生成函数,可辨识性(即精确恢复)在排除平凡歧义的意义上成立。拉普拉斯分布是满足该假设的最典型例子。我们的证明依赖于源分布中的折点与实解析函数平滑性之间的对比。实解析函数涵盖了一类广泛的生成机制,并且可以使用带有标准激活函数(如tanh、softplus、GELU)的归一化流或变分自编码器进行逼近,因此我们的结果几乎无需修改即可应用于现有的训练流程。我们在真实数据和合成数据上,使用归一化流和变分自编码器进行了实验,验证了它们的可辨识性。在CelebA数据集上的实验中,我们恢复了多个可解释的潜在因子,它们分别控制数据集中的不同属性。
cs.LG / 72 / 2609.21932

Joint Remaining Useful Life Prediction and Capacity Estimation of Lithium-Ion Batteries Using Partial-Charging Data

基于部分充电数据的锂离子电池剩余使用寿命预测与容量估计的联合研究
Tran, Khoa, Nguyen, Ho-Si-Hung, Moe, Phone Wai Yan, Trinh, Hung-Cuong, Tran, Thi-Hoang-Giang
Abstract
Joint remaining useful life (RUL) prediction and capacity estimation require representations of both gradual degradation and recent battery behavior. This paper presents a cross-expert framework using partial-charging measurements without measured historical full-cycle capacity as an input. The RUL Expert encodes nominal 10-min segments from ten cycles sampled within a 30-cycle history using a pretrained gated recurrent unit (GRU) encoder, a two-dimensional convolutional neural network (2D-CNN), and a temporal GRU. The Capacity Expert processes statistical descriptors of nominal 40-min segments from ten consecutive cycles using a 2D-CNN and a Transformer. A feature-wise linear modulation module uses the short-term representation to condition the long-term representation for joint prediction. Training comprises supervised autoencoder pretraining, independent expert pretraining, and fusion training with frozen experts. On two public battery-aging datasets, the reference configuration achieves mean RUL root-mean-square errors of 143.69 and 161.10 cycles and capacity errors of 12.36 and 7.28mAh, respectively. On Dataset I, fusion reduces both mean errors relative to either standalone expert. The results demonstrate a trade-off between RUL and capacity accuracy: the proposed method attains the lowest reported RUL RMSE among the compared methods on both datasets, whereas several baselines yield lower capacity errors.
Chinese Translation
剩余使用寿命(RUL)预测与容量估计的联合任务需要对电池渐进退化特性和近期行为进行表征。本文提出一种跨专家框架,仅使用部分充电测量数据作为输入,而无需实测的历史完整循环容量。RUL专家利用预训练的门控循环单元(GRU)编码器、二维卷积神经网络(2D-CNN)和时序GRU,对从30个循环历史中采样的十个循环的名义10分钟片段进行编码。容量专家利用2D-CNN和Transformer处理十个连续循环的名义40分钟片段的统计描述符。特征级线性调制模块利用短期表征对长期表征进行条件化,以实现联合预测。训练过程包括有监督自编码器预训练、独立专家预训练以及冻结专家的融合训练。在两个公开的电池老化数据集上,参考配置的平均RUL均方根误差分别为143.69和161.10个循环,容量误差分别为12.36和7.28 mAh。在数据集I上,相对于任一独立专家,融合方法降低了两种平均误差。结果表明RUL与容量精度之间存在权衡:在两个数据集上,所提方法在对比方法中取得了最低的RUL均方根误差,而若干基线方法的容量误差更低。
cs.LG / 73 / 2609.21945

Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks

学习驱动城市:用于城市网络校准与控制的深度元模型与强化学习策略
Adepitan, Adewumi Augustine, Haruna, Christopher J., Adegoke, Oluwasegun, Ajiboye, Ayooluwatomiwa, Oluwasakin, Oluwatobi
Abstract
Urban transportation networks present complex optimization challenges spanning calibration of high-fidelity simulators and real-time operational control. This paper presents a shared latent-space framework that connects simulator calibration and reinforcement learning control through a common learned representation of urban traffic dynamics. First, we develop a combinatorial MLP-autoencoder architecture that learns low-dimensional manifolds linking simulator inputs (origin-destination demand, network parameters) to outputs (travel times, congestion patterns), enabling efficient Bayesian optimization for calibration. This approach demonstrates superior sample efficiency compared to traditional dimension reduction methods, achieving better fit to observational data within fixed computational budgets. Second, we implement a deep Q-learning agent with experience replay and target networks to optimize dynamic traffic assignment through scheduling and routing adjustments. In empirical evaluations on benchmark networks, our approach reduces system-wide travel times by up to 51% compared to baseline operations. The learned latent representation is not only used to reduce the dimensionality of Bayesian calibration, but is also incorporated into the reinforcement learning state representation, allowing the control policy to operate on compressed and calibrated traffic dynamics. This shared latent-space formulation provides a unified pathway from simulator calibration to adaptive operational control within intelligent transportation systems. Our results highlight the transformative potential of deep learning methods in urban mobility planning and management, particularly for large-scale networks where traditional optimization approaches face computational bottlenecks.
Chinese Translation
城市交通网络带来了复杂的优化挑战,涵盖高保真仿真器的校准与实时运行控制。本文提出了一个共享潜在空间框架,通过对城市交通动态的共同学习表征,将仿真器校准与强化学习控制相连接。首先,我们开发了一种组合式多层感知器自编码器(MLP-autoencoder)架构,学习将仿真器输入(起讫点需求、网络参数)映射到输出(行程时间、拥堵模式)的低维流形,从而实现高效的贝叶斯优化校准。与传统降维方法相比,该方法展现出更优的样本效率,在固定计算预算内能够更好地拟合观测数据。其次,我们实现了一个具备经验回放和目标网络的深度Q学习(Deep Q-learning)智能体,通过调度和路径调整来优化动态交通分配。在基准网络上的实证评估中,与基线运行相比,我们的方法最多可将全系统行程时间降低51%。所学习的潜在表征不仅用于降低贝叶斯校准的维度,还被纳入强化学习的状态表征中,使控制策略能够在压缩且经校准的交通动态上运行。这种共享潜在空间的建模方法为智能交通系统内从仿真器校准到自适应运行控制提供了一条统一路径。我们的研究结果凸显了深度学习方法在城市出行规划与管理中的变革性潜力,尤其是在传统优化方法面临计算瓶颈的大规模网络中。
cs.LG / 74 / 2609.21953

RACER: Role-Aligned Competence Estimation for Human-AI Routing

RACER:面向人机路由的角色对齐能力估计
Strong, Joshua, Sun, Emma, Capstick, Alexander, Saha, Pramit, Ouyang, Cheng, Noble, J. Alison
Abstract
Learning to defer asks a predictive system when to act autonomously and when to defer to a human expert. Population-adaptive deferral extends this problem to unseen experts using a small context set of expert behavior. Neural context encoders such as L2D-Pop can be query-dependent, but may learn routing shortcuts tied to absolute class coordinates. Identity-Free Deferral (IFD) removes such shortcuts through role-indexed classwise competence profiles, but its estimates are constant within each class and cannot capture instance-level expert specialization. We propose RACER---Role-Aligned Competence Estimation for Routing---a role-relative framework for estimating an unseen expert's competence from context. RACER estimates the posterior-predictive probability that the expert is correct on a query under each candidate class role, then combines these estimates with the model posterior to obtain the Bayes-relevant expert-correctness probability. Nonparametric and neural kernel-pooling estimators use candidate-role relations, shared aggregation, and symmetric summaries, excluding absolute class-identity channels. We prove coherent class-relabelling invariance, derive a Bayes-aligned deferral surrogate, and give a plug-in regret bound relating routing regret to classifier and competence-estimation error. On controlled synthetic benchmarks, including a PathMNIST histopathology context-scaling study with simulated experts, RACER benefits from additional context under hidden subtype dependence and gives the strongest aggregate performance on a separately sampled unseen-expert split in the CIFAR-100 synthetic experiments. On the radiologist and human--AI chest-radiography benchmarks (VinDr-CXR and CheXpert), the RACER family is competitive or best in budget-swept deferral, with calibration results varying across metrics and datasets.
Chinese Translation
学习型延迟决策研究预测系统何时应自主行动、何时应转交人类专家。种群自适应延迟将这一问题扩展到未见过的专家,利用一个小的专家行为上下文集。诸如 L2D-Pop 之类的神经上下文编码器可以依赖查询,但可能学到与绝对类别坐标绑定的路由捷径。无身份延迟(Identity-Free Deferral, IFD)通过基于角色的类别级能力画像消除了此类捷径,但其估计在每个类别内是恒定的,无法刻画实例级的专家专长。我们提出 RACER——面向路由的角色对齐能力估计(Role-Aligned Competence Estimation for Routing)——一个基于角色相对性的框架,用于从上下文中估计未见专家的能力。RACER 估计专家在各个候选类别角色下对查询判断正确的后验预测概率,然后将这些估计与模型后验相结合,得到与贝叶斯决策相关的专家正确概率。非参数估计器和神经核池化估计器利用候选角色关系、共享聚合和对称摘要,排除了绝对类别身份通道。我们证明了该框架具有一致的类别重标记不变性,推导了与贝叶斯决策对齐的延迟替代损失,并给出了一个插件式遗憾界,将路由遗憾与分类器误差和能力估计误差联系起来。在受控合成基准上,包括一项结合模拟专家的 PathMNIST 组织病理学上下文规模扩展研究,RACER 在隐含子类型依赖下能够从更多上下文中获益,并在 CIFAR-100 合成实验中单独抽样的未见专家划分上取得最强的综合性能。在放射科医师及人机胸部X光基准(VinDr-CXR 和 CheXpert)上,RACER 系列方法在不同预算下的延迟决策中具有竞争力或表现最佳,其校准结果因指标和数据集而异。
cs.LG / 75 / 2609.21989

Time series generation with spectrally aligned latent flow matching

基于频谱对齐潜在流匹配的时间序列生成
Reyes, Camilo Carvajal, Tobar, Felipe
Abstract
Latent flow models have proven to be a reliable and cost-effective method for time series generation. However, the latent compression induces unwanted artefacts, such as a spectral mismatch with respect to the underlying dataset, thus hindering their use as training surrogates. In this article, we propose a spectrally-aligned latent-flow time series generator, where the latent space for flow matching is trained to preserve dynamical properties that are relevant for the suitability of synthetic samples. We find that incorporating fine-tuning losses based on canonical signal representations such as the Fourier, wavelet and signature transforms helps overcome these issues. The interpretability of these transformations allows us to ensure that the synthetic signals are aligned with the true ones in terms of relevant features, such as smoothness or targeted spectral content, as opposed to relying on pointwise reconstruction losses only. We compare the proposed aligned models against a base latent-flow model and the state of the art over real-world long-range univariate and multivariate benchmark datasets. Our quantitative results validate the superiority of the proposed method in terms of its performance on metrics reflecting signal realness and computational efficiency, while being aligned to the training set with respect to its local structure.
Chinese Translation
潜在流模型(latent flow models)已被证明是一种可靠且低成本的时间序列生成方法。然而,潜在空间压缩会引入不良伪影,例如与底层数据集的频谱失配,从而阻碍了其作为训练替代品的使用。在本文中,我们提出了一种频谱对齐的潜在流时间序列生成器,其中用于流匹配的潜在空间经过训练,能够保留与合成样本适用性相关的动态特性。我们发现,引入基于典型信号表示(如傅里叶变换、小波变换和特征变换(signature transforms))的微调损失有助于克服这些问题。这些变换的可解释性使我们能够确保合成信号在相关特征(如平滑度或目标频谱内容)方面与真实信号对齐,而不是仅仅依赖于逐点重构损失。我们在真实世界长程单变量和多变量基准数据集上,将所提出的对齐模型与基础潜在流模型以及最先进的方法进行了比较。定量结果验证了所提方法在反映信号真实性和计算效率的指标上的优越性,同时在局部结构方面与训练集保持一致。
cs.LG / 76 / 2609.21995

Assessment of Machine Learning-Based Critical Heat Flux Models in the CTF Subchannel Code for Square Rod Bundle Prediction

基于机器学习的临界热流密度模型在CTF子通道程序中用于方形棒束预测的评估
Furlong, Aidan, Monteiro, Vinicius de Melo, Salko, Robert, Duarte, Juliana Pacheco, Wu, Xu
Abstract
The prediction of critical heat flux (CHF), a key safety-related quantity in nuclear thermal hydraulics, remains an important challenge due to its direct relationship with fuel performance and reactor safety. Recent studies have demonstrated that relative to traditional empirical correlations and lookup tables (LUTs), machine learning (ML) methods can substantially improve CHF prediction accuracy. Most ML-based CHF models, however, have been developed and evaluated using tube databases, leaving their applicability to reactor-relevant rod bundle geometries largely unexplored. This study evaluates ML-based CHF models deployed within the CTF subchannel code using the Electric Power Research Institute (EPRI) rod bundle CHF database. Both pure and hybrid residual correction models are considered in local and semilocal formulations. The tube-trained ML CHF models generally transferred favorably to rod bundle applications and outperformed traditional CHF methods across most geometries and operating conditions. The local hybrid LUT model produced the strongest overall performance, and the semilocal pure ML model remained highly competitive. Comparison against the Bowring correlation, W-3 correlation, and 2006 Groeneveld LUT demonstrated that substantial improvements in rod bundle CHF prediction are possible even when models are trained exclusively on tube data. These findings provide one of the first large-scale assessments of ML-based CHF models in square rod bundles within a production-level subchannel analysis environment and support their broader application in reactor thermal hydraulic analysis.
Chinese Translation
临界热流密度(CHF)是核热工水力学中一个关键的安全相关物理量,其预测由于与燃料性能和反应堆安全性直接相关,始终是一项重要挑战。近期研究表明,相对于传统的经验关联式和查找表(LUT),机器学习(ML)方法能够显著提高CHF预测精度。然而,大多数基于机器学习的CHF模型均基于圆管数据库开发和评估,其在反应堆相关棒束几何结构中的适用性在很大程度上尚未得到探索。本研究利用美国电力研究院(EPRI)棒束CHF数据库,对部署在CTF子通道程序中的机器学习CHF模型进行了评估。所考虑的模型包括局部和半局部格式下的纯机器学习模型以及混合残差修正模型。基于圆管数据训练的机器学习CHF模型总体上能够良好地迁移至棒束应用,并在大多数几何结构和运行工况下优于传统CHF方法。局部混合LUT模型表现出最强的整体性能,半局部纯机器学习模型也保持高度竞争力。与Bowring关联式、W-3关联式以及2006年Groeneveld查找表的对比表明,即使模型仅使用圆管数据进行训练,棒束CHF预测仍可获得大幅改进。这些发现提供了在生产级子通道分析环境中对方形棒束内机器学习CHF模型的首批大规模评估之一,并支持其在反应堆热工水力分析中的更广泛应用。
cs.LG / 77 / 2609.22005

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

弃权与噪声过滤:Softmax注意力缺失的两个基本能力
Wang, Richard Zhe
Abstract
Gating the value pathway of attention reportedly improves language model pretraining, and prior studies disagree on why. We argue and provide experimental evidence that such gates supply two different things that softmax attention lacks: abstention and noise filtering. The first is abstention, which allows an attention head to output nothing, bypassing the requirement that attention weights must sum to one. The second is noise filtering, which allows the value pathway of an attention head to suppress interference from superposed features in the residual stream. In our experiments in matched models from 10M to 350M parameters, we supply abstention through a learned per-head sink logit in the softmax and noise filtering through a gate on each value. We report three empirical findings. First, the benefit of abstention, measured as the reduction in validation loss relative to a matched baseline, declines as models grow, whereas the benefit of noise filtering increases with scale. In particular, abstention accounts for nearly all of the gain from gating at 10M and filtering for most of it at 350M. Second, the best model at every scale is the one with both primitives built in. Third, injecting controlled interference into the values a head reads confirms that the gate removes such interference, and reveals that each of the two gate forms we study has a characteristic blind spot. Supplying both primitives adds negligible parameters and remains compatible with the key-value cache.
Chinese Translation
对注意力的值通路进行门控(gating)据称可以改进语言模型的预训练,但以往研究对其原因看法不一。我们论证并提供实验证据表明,此类门控提供了softmax注意力所缺乏的两种不同功能:弃权(abstention)与噪声过滤(noise filtering)。第一种是弃权,它允许注意力头输出空结果,从而绕过注意力权重之和必须为一的限制。第二种是噪声过滤,它允许注意力头的值通路抑制残差流(residual stream)中叠加特征所产生的干扰。在参数量从10M到350M的匹配模型实验中,我们通过在softmax中引入可学习的每头汇点逻辑值(per-head sink logit)来实现弃权,并通过在每个值上添加门控来实现噪声过滤。我们报告了三项实证发现。第一,弃权的收益(以相对于匹配基线的验证损失降低量来衡量)随模型规模增大而下降,而噪声过滤的收益则随规模增大而上升。具体而言,在10M参数规模下,弃权几乎贡献了门控带来的全部增益,而在350M规模下,噪声过滤贡献了大部分增益。第二,在每一规模下表现最佳的模型都是同时内置这两种基本能力的模型。第三,向注意力头读取的值中注入受控干扰证实了门控能够消除此类干扰,并揭示出我们所研究的两种门控形式各自存在特征性的盲区。同时提供这两种基本能力带来的参数开销可忽略不计,且与键值缓存(key-value cache)保持兼容。
cs.LG / 78 / 2609.22012

COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules

COMPLEX:多参数持续同调模的闭式认证嵌入
Majhi, Sushovan, Mitra, Atish, Virk, Žiga, Bagchi, Pramita
Abstract
Every multiparameter persistence vectorization we know of carries a one-sided Lipschitz upper bound and nothing below it: without a lower gauge there is no sense in which the features are faithful, and no per-prediction guarantee can be built on them. This paper supplies the missing side. COMPLEX is a closed-form, training-free embedding of multiparameter modules -- slice the module along a fixed near-diagonal net, embed each slice barcode by the certified PLACE/PALACE landmark map, concatenate. Under a checkable witnessing-slice coherence condition, holding on 100% of audited pairs on Orbit5k, a single slice carries a closed-form lower gauge: separated modules stay separated in the embedding. With the standard upper bound this gives, to our knowledge, the first two-sided distortion bound for a multiparameter feature map, making faithfulness measurable. Measuring it, we find the floor tight within a small factor of realized distances yet operationally local: an RBF-SVM reaches 91% where 1-NN reaches 78% on the same features. Local per-prediction certification therefore fails for a structural reason common to every landmark embedding whose lower gauge is witnessed by one coordinate. With no learned embedding and no held-out calibration -- only a cross-validated SVM head -- COMPLEX sets the state of the art on both Orbit benchmarks (91.95% on Orbit5k, 92.98% on Orbit100k), level with or above Euler-characteristic surfaces and above transformers and graphcode. On graphs it exceeds GRIL on all four shared molecular benchmarks with one fixed configuration, including the only multiparameter method to clear COX2's majority baseline by more than three points. Closed-form selection -- of the landmark radius, the kernel (certificate-preserving), and the bifiltration set -- buys further accuracy; gradient-shaped adaptation buys none.
Chinese Translation
据我们所知,现有的每一种多参数持续同调向量化方法都只具有单侧的 Lipschitz 上界,而没有任何下界:缺少下界度量,就无法说明这些特征是忠实的,也无法在此基础上建立逐预测的保证。本文补上了缺失的这一侧。COMPLEX 是一种闭式的、无需训练的多参数模嵌入方法——沿固定的近对角网格对模进行切片,用经过认证的 PLACE/PALACE 标志点映射嵌入每个切片的条形码,然后拼接。在一个可检验的“见证切片一致性”条件下(在 Orbit5k 数据集上审计的全部配对中 100% 满足),单个切片即可携带闭式的下界度量:相互分离的模在嵌入中保持分离。结合标准的上界,据我们所知,这给出了首个针对多参数特征映射的双侧失真界,使忠实性变得可度量。在实际度量中,我们发现该下界在已实现距离的较小倍数范围内是紧的,但在操作上是局部的:在同一特征上,RBF-SVM 达到 91%,而 1-NN 仅达到 78%。因此,逐预测的局部认证失败源于一个结构性的原因,该原因对任何下界由单一坐标见证的标志点嵌入方法都普遍存在。在不使用学习到的嵌入、也无需留出校准集的情况下(仅使用交叉验证的 SVM 头),COMPLEX 在两个 Orbit 基准上均达到最先进水平(Orbit5k 上为 91.95%,Orbit100k 上为 92.98%),与欧拉特征曲面持平或更优,并超过 transformer 和 graphcode。在图数据上,它仅用一种固定配置就在全部四个共享分子基准上超过了 GRIL,并且是唯一一个在 COX2 上以超过三个百分点的优势突破多数类基线的多参数方法。闭式的参数选择——包括标志点半径、核函数(保持认证性质)以及双过滤集合——能进一步提升精度;而梯度形状的自适应调整则没有任何收益。
cs.LG / 79 / 2609.22041

$\lambda$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource

$\lambda$-控制的GRPO:将流匹配比率不稳定性转化为一种可预算的资源
Wang, Yufeng, Priye, Parivesh, Marathe, Meeshawn, Pahwa, Ramit
Abstract
Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback. Training in this setting is unstable in a way specific to multi-step denoising: the policy update changes systematically across denoising steps, with importance ratios drifting below one, becoming increasingly dispersed, clipping at different rates, and leaving fewer usable samples late in training. Prior work treats these effects as separate failure modes and addresses each with a hand-tuned stabilizer. We show instead that they arise from a single per-step quantity, which we call path variance. This quantity is determined exactly by the sampler's Gaussian transition kernel and can be estimated cheaply during training. This reframes instability as a resource that can be measured and budgeted rather than a collection of symptoms to repair. Our method, $\lambda$-Controlled GRPO, calibrates importance-ratio behavior from this predicted law rather than from noisy empirical statistics, and allocates gradient effort across denoising steps according to their predicted cost. The two scales governing the update are fixed by standard policy choices rather than introduced as free tuning parameters. On a text-to-image model under two reward settings, rendering difficult target text scored by optical character recognition and matching human preferences scored by a preference model, $\lambda$-Controlled GRPO improves both text accuracy and preference reward over the strongest empirical stabilizer. It also keeps late-step path variance within its intended budget, precisely where the baseline systematically overshoots. The result is a Flow-GRPO update calibrated by its own transition law rather than stabilized after instability appears.
Chinese Translation
强化学习正被日益广泛地用于将图像生成器与奖励信号对齐,Flow-GRPO 最近通过将去重采样器视为一个可从奖励反馈中优化的随机策略,将这一范式扩展到了流匹配模型。在此设置下,训练会出现一种多步去重所特有的不稳定现象:策略更新在各个去重步骤间呈现系统性变化,重要性比率漂移至1以下、离散度不断增加、以不同速率被裁剪,并导致训练后期可用样本减少。已有研究将这些效应视为彼此独立的失效模式,并针对每一项采用手工调节的稳定化手段加以处理。我们则证明,它们源于一个单一的分步量,我们称之为路径方差。该量由采样器的高斯转移核精确决定,并且可以在训练过程中以极低的成本进行估计。这将不稳定性重新界定为一种可测量、可预算的资源,而非一组待修复的症状。我们的方法 $\lambda$-Controlled GRPO 基于这一可预测的规律(而非含噪的经验统计量)来校准重要性比率行为,并根据各去重步骤的预测成本来分配梯度投入。控制更新的两个尺度由标准策略选择所确定,而非作为自由调节参数引入。在一个文本到图像模型上的两个奖励设定下——即以光学字符识别评分的困难目标文本渲染,以及以偏好模型评分的人类偏好匹配——$\lambda$-Controlled GRPO 相较于最强的经验性稳定化方法,同时提升了文本准确率与偏好奖励。它还将后期步骤的路径方差保持在预期预算之内,而这正是基线方法系统性超出的地方。其结果是一个由自身转移规律校准的 Flow-GRPO 更新,而非在出现不稳定之后再进行事后稳定。
cs.LG / 80 / 2609.22048

Available Guardrails: Certifying Selective Prediction across ML Systems

可用的守护机制:跨机器学习系统的选择性预测认证
Priye, Parivesh, Wang, Yufeng, Ling, Haibin, Chaykowsky, Michael
Abstract
A selective predictor acts as a safety gate: it returns an output only when the prediction appears sufficiently trustworthy. Deployments increasingly require this reliability to be certified at a target precision for every reporting unit of interest, such as a tool, policy label, or patient subgroup. The main difficulty is often not whether a granted certificate is valid, but whether finite calibration data can produce one at all. As the gate becomes safer or more fine-grained, some units may receive too little evidence to certify. We make this notion of availability computable through classical exact-binomial inversion and formulate reporting-partition selection, under a fixed group order, as a dynamic program that exposes the trade-off among safety, granularity, and served traffic. The resulting frontier reveals a large population opportunity that finite-sample estimation nearly erases: a truth-informed planner gains $0.157$ mean coverage over support balancing, whereas a naive estimator recovers only $0.005$, making recovery from finite data the central challenge. Constructing candidate partitions on one planning split and selecting among them on another recovers part of this gap, improving mean coverage over support balancing by $0.060$, with the direction reproduced in $59$ of $60$ model effects across three intent-routing datasets and two architectures. A complementary validity-preserving lever, reallocating the familywise error budget across reporting units, recovers additional coverage both with population quantities and noisy estimates. The same frontier recurs, with predictor-specific ceilings, across LLM tool-calling, content moderation, lesion classification, and recommendation. Certified availability is therefore a plannable deployment resource that determines when a safety gate can be certified, at what granularity, and over how much traffic.
Chinese Translation
选择性预测器扮演着安全门控的角色:仅当预测看起来足够可信时才返回输出。越来越多的部署场景要求针对每个感兴趣的报告单元(如某个工具、政策标签或患者亚群),在目标精度下对这种可靠性进行认证。主要困难往往不在于所授予的证书是否有效,而在于有限的校准数据是否能够产出证书。随着门控变得更加安全或更加细粒度,某些单元可能因证据太少而无法通过认证。我们通过经典的精确二项分布反演使这一“可用性”概念变得可计算,并在固定组序的条件下,将报告分区的选择形式化为一个动态规划,从而揭示安全性、粒度与被服务流量之间的权衡。所得到的前沿边界揭示了一个被有限样本估计几乎抹去的巨大总体机会:一个基于真实信息的规划器相较于支持平衡(support balancing)可获得 $0.157$ 的平均覆盖率提升,而一个朴素估计器仅能恢复 $0.005$,这使得从有限数据中的恢复成为核心挑战。在一个规划划分上构建候选分区,并在另一个划分上进行选择,可以恢复部分差距,使平均覆盖率相较于支持平衡提升 $0.060$,且该方向在三个意图路由数据集和两种架构上的 $60$ 个模型效应中有 $59$ 个得到复现。一个互补的、保持有效性的手段是在各报告单元之间重新分配族错误率预算,它无论使用总体量还是含噪估计都能带来额外覆盖。同样的前沿边界在LLM工具调用、内容审核、病灶分类和推荐等任务中重复出现,只是上限因预测器而异。因此,经认证的可用性是一种可规划的部署资源,它决定了安全门控在何时能够被认证、以何种粒度认证,以及覆盖多少流量。
cs.LG / 81 / 2609.22053

Particle Competition and Cooperation for Robust Graph Convolutional Network Learning Under Label Noise

基于粒子竞争与合作的标签噪声下鲁棒图卷积网络学习
Breve, Fabricio
Abstract
Graph Convolutional Networks (GCNs) are highly sensitive to label noise, since corrupted supervision can propagate through the graph and degrade learned node representations. This work proposes PCC+GCN, a hybrid framework that uses Particle Competition and Cooperation (PCC) as a graph-based label-refinement stage before GCN training. PCC identifies suspicious labeled nodes through particle domination dynamics and determines whether their labels should be preserved, removed, or reassigned before GCN training. The framework also allows the graph used by PCC to be augmented with feature-based $k$-nearest-neighbor edges, while the GCN itself is trained on the original graph structure and node features. The proposed method was evaluated on ten graph datasets from the NoisyGL benchmark under conventional Uniform, Pair, and Random label noise, as well as under instance-dependent label noise. A detailed hyperparameter analysis was also conducted on Cora, CiteSeer, and PubMed. Under conventional noise, PCC+GCN achieved the highest overall average accuracy and the best average rank among the evaluated methods, with an average gain of $1.67$ percentage points over the baseline GCN across the clean setting and all noisy scenarios. Under instance-dependent noise, PCC+GCN remained competitive with the best-performing robust methods while requiring substantially lower execution time, being the fastest robust method on eight of the ten datasets. The results indicate that PCC-based label refinement provides an effective and computationally efficient preprocessing strategy for improving GCN robustness under noisy supervision.
Chinese Translation
图卷积网络(GCN)对标签噪声高度敏感,因为受损的监督信息会通过图结构传播并降低所学习节点表示的质量。本工作提出了PCC+GCN,一种混合框架,它在GCN训练之前使用粒子竞争与合作(Particle Competition and Cooperation, PCC)作为基于图的标签精炼阶段。PCC通过粒子支配动力学识别可疑的已标记节点,并决定其标签应在GCN训练前被保留、移除还是重新分配。该框架还允许PCC所使用的图通过基于特征的$k$近邻边进行增强,而GCN本身则在原始图结构和节点特征上进行训练。所提出的方法在NoisyGL基准的十个图数据集上,于常规的Uniform、Pair和Random标签噪声以及实例依赖型标签噪声下进行了评估。此外,还在Cora、CiteSeer和PubMed上进行了详细的超参数分析。在常规噪声下,PCC+GCN在所评估的方法中取得了最高的总体平均准确率和最佳平均排名,在干净设置及所有噪声场景下,相比基线GCN平均提升$1.67$个百分点。在实例依赖型噪声下,PCC+GCN与表现最佳的鲁棒方法保持竞争力,同时所需执行时间显著更低,在十个数据集中的八个上是最快的鲁棒方法。结果表明,基于PCC的标签精炼为提升GCN在噪声监督下的鲁棒性提供了一种有效且计算高效的前处理策略。
cs.LG / 82 / 2609.22055

Benchmarking World Models for Continual Learning on Compositional Tasks

面向组合任务的持续学习世界模型基准测试
Zhou, Haoyu, Watson, Joe, Lei, Anson, Posner, Ingmar
Abstract
A desirable property of a world model is the ability to learn continually across tasks, adapting to new environments without forgetting what the agent has already learnt. In particular, the ability to retain and reuse knowledge obtained from prior experiences underpins an agent's ability to efficiently adapt to novel environments, as the dynamics of the physical world can often be described in recurring mechanisms. However, the world model's measure of adaptation entangles two abilities: the speed and capacity to learn unseen tasks, and the reuse of knowledge already acquired, since incoming tasks carry novel content alongside what recurs. In order to isolate knowledge reuse from prior experiences, we propose a compositional continual learning benchmark for world models in robot manipulation. Specifically, we design each task curriculum with compositional tasks that combine aspects of the tasks seen in the sequence. We further factorise this composition along the axes of action and perception to better understand how different input modalities bottleneck knowledge reuse. We evaluate state-of-the-art world models under canonical continual learning methods, alongside a modular world model whose dynamics backbone contains explicitly reusable components. Results show that modularity balances reuse against forgetting better than conventional methods, but none solve the problem fully, leaving clear room for continual world models built to reuse without forgetting. More details are available on our project website: https://object814.github.io/Compositional-Continual-Learning/.
Chinese Translation
世界模型的一个理想特性是能够跨任务持续学习,即在不遗忘智能体已学知识的前提下适应新环境。特别是,保留并复用先前经验所获知识的能力,是智能体高效适应新环境的基础,因为物理世界的动力学往往可以用循环出现的机制来描述。然而,世界模型的适应能力度量将两种能力纠缠在一起:学习未见任务的速度与容量,以及对已有知识的复用——因为新任务在包含新内容的同时也包含重复出现的内容。为了将知识复用与先前经验分离开来,我们提出了一个面向机器人操作中世界模型的组合式持续学习基准。具体而言,我们设计的每个任务课程都由组合任务构成,这些任务结合了序列中已见任务的不同方面。我们进一步沿着动作与感知两个维度对该组合进行因子分解,以更好地理解不同输入模态如何限制知识复用。我们在经典的持续学习方法下评估了最先进的世界模型,并评估了一个模块化世界模型,其动力学骨干网络包含显式可复用的组件。结果表明,模块化在复用与遗忘之间的平衡上优于传统方法,但没有任何方法能完全解决该问题,这为构建能够复用而不遗忘的持续学习世界模型留下了明确的空间。更多详情请访问我们的项目网站:https://object814.github.io/Compositional-Continual-Learning/。
cs.LG / 83 / 2609.22064

BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings

BrainWideBench:面向多脑区神经记录的大规模预训练与跨动物迁移的基准测试
Andre, Alexandre, Mahato, Shivashriganesh P., Arora, Vinam, Balaji, Keshav, Lachi, Divyansha, Krishna, Nanda H., Xiao, Jingyun, Zhang, Yizi, Mao, Ximeng, Ma, Wenrui, Yu, Han, Laboratory, International Brain, Birman, Daniel, Bonacchi, Niccolò, Chapuis, Gaelle A., Catarino, Joana A., Davatolhagh, Felicia, Faulkner, Mayo, Freitas-Silva, Laura, Hu, Fei, Huntenburg, Julia M., Khanal, Anup, Laranjeira, Inês, Lau, Petrina, Meijer, Guido T., Miska, Nathaniel J., Noel, Jean-Paul, Pan-Vazquez, Alejandro, Raiser, Georg, Rossant, Cyrille, Socha, Karolina Z., Urai, Anne E., Wells, Miles J., West, Steven J., Winter, Olivier, Richards, Blake, Lajoie, Guillaume, Hurwitz, Cole, Azabou, Mehdi, Whiteway, Matthew R., Paninski, Liam, Dyer, Eva L.
Abstract
Advances in large-scale neural recording have made it possible to collect data across many animals and distributed brain regions, raising the question of whether this scale can be exploited to learn general-purpose neural representations transferable across diverse downstream tasks. Yet, progress toward this goal has been limited by fragmented evaluation protocols and a narrow focus on individual task domains. Here, we present BrainWideBench, a benchmark for evaluating across-animal transfer on multi-region neural recordings, built on the International Brain Laboratory Brainwide Map dataset of neural and behavioral recordings spanning 276 brain regions from 139 mice performing a sensory-guided decision-making task. The benchmark is organized around three complementary task suites that evaluate whether learned representations support downstream decoding of behavior, can predict masked or future neural activity, and can recover biologically meaningful anatomical organization. With this benchmark, we systematically evaluate pretraining methods across transfer settings, including finetuning on downstream objectives and zero-shot generalization to unseen animals. Our results confirm pretraining improves performance over matched single-session baselines, but we show current methods exhibit heterogeneity in transfer capabilities: gains depend strongly on the alignment between pretraining objectives and downstream tasks. No single approach performs uniformly well across all three suites, and most methods are designed to only address a subset of them. Together, these findings suggest that learning representations that jointly generalize across behavior, dynamics, and anatomy remains an open challenge. By providing a unified and reproducible evaluation suite, BrainWideBench establishes a framework for measuring progress toward general-purpose models of the mouse brain.
Chinese Translation
大规模神经记录技术的进展使得跨多个动物和分布式脑区的数据采集成为可能,由此引发了一个问题:这种规模能否被利用来学习可跨多种下游任务迁移的通用神经表征。然而,碎片化的评估协议以及对单一任务领域的狭窄关注限制了这一目标的进展。本文提出 BrainWideBench,一个用于评估多脑区神经记录上跨动物迁移能力的基准,其构建基于国际脑实验室(International Brain Laboratory)的 Brainwide Map 数据集,该数据集包含139只小鼠在执行感觉引导决策任务时跨276个脑区的神经与行为记录。该基准围绕三个互补的任务套件组织,分别评估所学表征是否支持对行为的下游解码、能否预测被遮蔽或未来的神经活动,以及能否恢复具有生物学意义的解剖结构组织。借助该基准,我们系统评估了预训练方法在多种迁移设置下的表现,包括在下游目标上的微调以及对未见动物的零样本泛化。我们的结果证实预训练优于匹配的单次会话基线,但同时发现现有方法在迁移能力上表现出异质性:性能提升在很大程度上依赖于预训练目标与下游任务之间的对齐程度。没有单一方法能在全部三个任务套件上均表现良好,且大多数方法仅被设计用于处理其中一部分任务。综上,这些发现表明,学习能够同时在行为、动力学和解剖结构上泛化的表征仍然是一个开放性挑战。通过提供一个统一且可复现的评估套件,BrainWideBench 为衡量迈向小鼠大脑通用模型的进展建立了框架。