← Back to Index
Daily Research Digest

arXiv Papers

2026-09-22
755
Papers
4
Categories
716
Translated
收藏清单 0
机器人学 (Robotics)
182
cs.RO / 1 / 2609.22274

CHOREO: Every Humanoid Skill as a Trajectory

CHOREO:将每种人形机器人技能表示为轨迹
Sun, Ziyi, Chen, Jingwen, Wang, Yuxi, Xia, Xiuze, Cheng, Long, Zhang, Zhaoxiang, Dong, Yujun
Abstract
Recent advances in humanoid robotics have produced diverse skills through reinforcement learning, motion imitation, and generative modeling. Yet these capabilities remain siloed because they are built around incompatible representations, interfaces, and controllers. We present CHOREO, a framework for training-free composition of heterogeneous humanoid skills. Our key observation is that, regardless of how a skill is learned, it can ultimately be expressed as an executable motion trajectory. Based on this observation, CHOREO converts each capability into SkillMotion, a unified representation that combines motion states, contacts, semantics, and boundary conditions. Skills are composed through direct continuation, cubic Hermite blending, or validated bridge motions, without retraining source models or updating models at test time. On Unitree G1 in MuJoCo, CHOREO organizes 2,950 admitted SkillMotion assets derived from heterogeneous sources and achieves 95.4\% sequence success across 130 multi-action tasks, including 93.8\% success on eight-action sequences. These results demonstrate that executable trajectories provide a scalable interface for accumulating and composing pretrained humanoid capabilities.
Chinese Translation
人形机器人技术的最新进展通过强化学习、动作模仿和生成式建模产生了多样化的技能。然而,这些能力由于建立在互不兼容的表示、接口和控制器之上,仍然彼此割裂。我们提出了CHOREO,一个无需训练即可组合异构人形机器人技能的框架。我们的关键观察是:无论技能是如何学得的,它最终都可以被表达为一条可执行的运动轨迹。基于这一观察,CHOREO将每种能力转换为SkillMotion——一种融合运动状态、接触、语义和边界条件的统一表示。技能通过直接延续、三次Hermite样条混合或经过验证的桥接动作进行组合,无需重新训练源模型或在测试时更新模型。在MuJoCo仿真环境中的Unitree G1平台上,CHOREO组织了源自异构来源的2,950个通过的SkillMotion资产,在130个多动作任务中实现了95.4%的序列成功率,其中包括八动作序列上93.8%的成功率。这些结果表明,可执行轨迹为积累和组合预训练的人形机器人能力提供了一个可扩展的接口。
cs.RO / 2 / 2609.22276

D3DWA: Adaptive Weight and Prediction-Horizon for Dynamic Window Approach via Dueling Double Deep Q-Network

D3DWA:基于Dueling双深度Q网络的动态窗口法自适应权重与预测时域方法
Jooyandeh, Zahra, Kobayashi, Masato, Uranishi, Yuki
Abstract
The Dynamic Window Approach (DWA) is widely used for local navigation, but its performance depends strongly on parameters that are typically fixed before navigation. In particular, the appropriate prediction horizon can vary with local free space: longer horizons support efficient motion in open areas, whereas shorter horizons help preserve feasible motions in narrow or cluttered regions. This paper proposes D3DWA, an adaptive DWA framework based on a Dueling Double Deep Q-Network (D3QN), which jointly selects the DWA evaluation weights and prediction horizon from a continuous navigation state at every control step while retaining DWA's trajectory generation and collision checking. In eight simulated environments, including unseen layouts, D3DWA reached every goal. Real-robot experiments further showed that D3DWA completed all three tested configurations, including a constrained case in which the weights-only variant timed out. These results demonstrate the benefit of jointly adapting the evaluation weights and prediction horizon. Additional material is available at https://mertcookimg.github.io/d3dwa/
Chinese Translation
动态窗口法(DWA)被广泛应用于局部导航,但其性能在很大程度上依赖于通常在导航前固定设置的参数。特别是,合适的预测时域会随局部自由空间的变化而变化:较长的时域有利于在开阔区域实现高效运动,而较短的时域则有助于在狭窄或杂乱区域保留可行的运动方案。本文提出D3DWA,一种基于Dueling双深度Q网络(D3QN)的自适应DWA框架,该框架在每个控制步中根据连续的导航状态同时选择DWA评价权重与预测时域,同时保留DWA的轨迹生成与碰撞检测机制。在包括未见过的布局在内的八个仿真环境中,D3DWA均成功到达所有目标点。真实机器人实验进一步表明,D3DWA完成了全部三种测试配置,其中包括仅调整权重的方法超时的一种受限场景。这些结果证明了同时自适应调整评价权重与预测时域的优越性。补充材料可参见 https://mertcookimg.github.io/d3dwa/
cs.RO / 3 / 2609.22278

DeViGrasp: Robust Visual Mobile Grasping for Quadruped Manipulators under Degraded Perception

DeViGrasp:退化感知下四足操作臂的鲁棒视觉移动抓取
Zhou, Liang, Su, Jiaming, Wei, Yancong, Dong, Kangkang, Liu, Houde
Abstract
Quadruped manipulators enable mobile grasping in complex environments, yet their whole-body control policies remain vulnerable to unreliable onboard visual perception. Existing methods are typically developed under relatively reliable observations and have not systematically examined how occlusion, segmentation-mask dropout, depth noise, and target-localization jitter affect grasp reasoning and target tracking. To address this gap, we introduce DeViGrasp-Bench, a benchmark for mobile grasping under degraded vision that incorporates controlled visual degradations, seen and unseen objects, multiple difficulty levels, and complex terrains, and evaluates task success, execution efficiency, and action smoothness. We further propose DeViGrasp-Net, a teacher--student framework that combines state-conditioned grasp reasoning with reliability-aware temporal target estimation. The privileged teacher attends to offline grasp candidates conditioned on object, robot, end-effector, and task states, while the deployable student fuses dual-view segmented-depth observations with current, memory, and recovery target hypotheses through Target Hold Memory and Temporal Memory Attention. DeViGrasp-Net outperforms VBC across degradation levels, unseen objects, and complex terrains, and surpasses an adapted DQ-Net across all evaluated degradation levels. Under the Difficult setting, it achieves a success rate of 62.3\%, improving upon VBC and DQ-Net by 16.1 and 4.3 percentage points, respectively; under the Hard setting, its margin over DQ-Net increases to 10.5 percentage points. Ablation studies confirm the complementary benefits of grasp-aware supervision and reliability-aware temporal memory.
Chinese Translation
四足操作臂(quadruped manipulators)能够在复杂环境中实现移动抓取,然而其全身控制策略仍然容易受到机载视觉感知不可靠性的影响。现有方法通常是在相对可靠的观测条件下开发的,尚未系统性地考察遮挡、分割掩码缺失(segmentation-mask dropout)、深度噪声以及目标定位抖动如何影响抓取推理和目标跟踪。为填补这一空白,我们提出了 DeViGrasp-Bench,一个面向退化视觉条件下移动抓取的基准,其中包含可控的视觉退化、已见与未见物体、多个难度级别以及复杂地形,并评估任务成功率、执行效率与动作平滑度。我们进一步提出了 DeViGrasp-Net,这是一个结合了状态条件化抓取推理与可靠性感知时序目标估计的师生(teacher--student)框架。拥有特权信息的教师模型以物体、机器人、末端执行器和任务状态为条件,关注离线抓取候选;可部署的学生模型则通过目标保持记忆(Target Hold Memory)与时序记忆注意力(Temporal Memory Attention),将双视角分割深度观测与当前、记忆和恢复目标假设相融合。DeViGrasp-Net 在退化程度、未见物体和复杂地形等各方面均优于 VBC,并在所有评估的退化程度上超越经过适配的 DQ-Net。在困难(Difficult)设置下,其成功率达到 62.3%,分别比 VBC 和 DQ-Net 高出 16.1 和 4.3 个百分点;在极难(Hard)设置下,其相对 DQ-Net 的优势扩大至 10.5 个百分点。消融实验证实了抓取感知监督与可靠性感知时序记忆的互补优势。
cs.RO / 4 / 2609.22285

ORDER: A Fictitious-World Benchmark for Domain-Adaptive Embodied AI

ORDER:一个面向领域自适应具身智能的虚构世界基准
Sathi, Sai Krishna Reddy, Tiwari, Anuj
Abstract
Adapting language models to new domains via continual pre-training raises a basic evaluation problem: if the training corpus overlaps with what the model already knows, performance gains cannot be cleanly attributed to new learning rather than pre-existing knowledge. This matters most for knowledge-intensive, task-light (KHTL) robot deployments - pharmaceutical dispensing, hazardous-material handling, facility-specific protocols, where the physical task is simple but the governing rules are proprietary and safety-critical, and where extensive live testing is costly or unsafe. We introduce ORDER (Ontology-driven Decision-making for Embodied Reasoning), a benchmark built on a fictitious world: a 342,069-token synthetic corpus defining a self-consistent physics that cannot appear in any model's pre-training data. ORDER pairs a 500-question knowledge test (ORDER-BENCH) with a harder compositional task, ORDER-SPATIAL: ordering objects for safe manipulation across both familiar and entirely novel scenes. GPT-4.1 without adaptation scores below chance on ORDER-SPATIAL (Kendall's tau = 0.441), showing its priors actively conflict with the invented physics. After continual pre-training, small models improve substantially on both familiar and novel scenes alike evidence of genuine world-model induction rather than memorization. We then carry this through to a robot pipeline: models that answer the knowledge test well often cannot produce valid, executable plans without a further skill-adaptation stage, after which small, fully offline models outperform GPT-4.1 even when GPT-4.1 is given retrieval access to the same rules (Kendall's tau = 0.848 vs. 0.606), on a full perception-to-execution loop demonstrated on a simulated iiwa7 arm with human-in-the-loop correction. Throughout, ORDER-SPATIAL performance, not knowledge-test accuracy is what predicts real plan quality.
Chinese Translation
通过持续预训练将语言模型适配到新领域时,存在一个基本的评估问题:如果训练语料与模型已知内容重叠,性能提升就无法清晰地归因于新知识的学习而非已有知识。这一问题对于知识密集、任务轻量(KHTL)的机器人部署场景最为关键——例如药品配发、危险品处理和机构专用规程——在这些场景中,物理任务本身简单,但约束规则是专有的且关乎安全,而大量的实机测试成本高昂或存在危险。我们提出了ORDER(Ontology-driven Decision-making for Embodied Reasoning,面向具身推理的本体驱动决策),一个建立在虚构世界之上的基准:其包含一个342,069词元的合成语料库,定义了一套自洽的物理体系,不可能出现在任何模型的预训练数据中。ORDER将一个500题的知识测试(ORDER-BENCH)与一个更具挑战性的组合任务ORDER-SPATIAL相结合:在熟悉场景和全新场景中进行面向安全操作物体排序。未经适配的GPT-4.1在ORDER-SPATIAL上的得分低于随机水平(Kendall's tau = 0.441),表明其先验知识与所发明的物理体系存在主动冲突。经过持续预训练后,小模型在熟悉场景和全新场景上均获得显著提升,这证明模型是真正归纳出了世界模型而非死记硬背。随后,我们将这一流程延伸至机器人管线:在知识测试中表现良好的模型,往往仍无法在不经过进一步技能适配阶段的情况下生成有效、可执行的计划;经过技能适配后,小型的完全离线模型超越了GPT-4.1——即使为GPT-4.1提供对相同规则的检索访问(Kendall's tau = 0.848 对 0.606)。整个流程在一个带有遥操作人类在环校正的仿真iiwa7机械臂上,通过完整的感知到执行闭环进行了验证。贯穿全文的核心发现是:真正能预测真实计划质量的是ORDER-SPATIAL上的表现,而非知识测试的准确率。
cs.RO / 5 / 2609.22289

OJOx: Specification-Conditioned Demonstrations for Embodied AI in Construction

OJOx:面向建筑具身智能的规格条件化示范
Dawod, Mohamed
Abstract
Large-scale egocentric and whole-body human demonstrations are becoming a primary source of data for embodied intelligence. They record what people perceive and do, but rarely the external specification that gave an action its purpose. In construction that omission is consequential: skilled work is directed at project-specific configurations defined in a design model - configurations not yet present in the environment being observed. A mason's transferable competence is not the geometry of one wall but the ability to realise a new geometry from a specification. We introduce the specification-conditioned demonstration: a synchronised record of the physical state a demonstrator perceives, the intended state supplied to them by an external design, and the behaviour connecting the two. We present OJOx, a capture interface that realises this for construction - delivering design geometry to a headset, anchoring it in the physical workspace, rendering it into a demonstrator's stereo passthrough view, and recording that view synchronously with whole-body and hand motion. We report one fully instrumented session - a 33-component wall laid against a specification that changes while the work proceeds - and check the record against the physical scene through an external camera registered independently of the capture. Recorded sessions remain compatible with existing humanoid retargeting infrastructure and replay onto a Unitree G1 in simulation. The result is a data interface for testing whether embodied policies can learn not merely to imitate demonstrated actions, but to act toward specifications absent from their training experience.
Chinese Translation
大规模自我中心视角与全身人体示范正日益成为具身智能的主要数据来源。这些数据记录了人们的感知与行为,却很少记录赋予动作目的的外部规格说明。在建筑领域,这种缺失的影响尤为重大:熟练的工作是针对设计模型中定义的项目特定配置展开的——而这些配置在被观察的环境中尚不存在。砌砖工人的可迁移能力并非某一面墙的几何形状,而是能够依据规格说明实现新的几何形状的能力。我们提出了规格条件化示范(specification-conditioned demonstration):一种同步记录,包含示范者感知到的物理状态、由外部设计提供给示范者的目标状态,以及连接二者的行为。我们开发了OJOx——一个面向建筑领域实现这一理念的采集接口:将设计几何信息传输至头戴式设备,将其锚定在物理工作空间中,渲染到示范者的立体透视视图(stereo passthrough view)中,并将该视图与全身及手部运动同步记录。我们报告了一次完整采集的实验环节——砌筑一面包含33个构件的墙体,且规格说明在施工过程中发生变更——并通过一台与采集系统独立配准的外部相机,将记录数据与物理场景进行核对。所记录的环节与现有人形机器人动作重定向基础设施保持兼容,并可在仿真中回放到Unitree G1机器人上。这一成果提供了一个数据接口,用于检验具身策略能否学会不仅是模仿示范动作,而是朝向其训练经验中未曾出现的规格说明去行动。
cs.RO / 6 / 2609.22299

When Does Test-Time Physical Diagnosis Pay? A Frozen Policy Buys Evidence It Never Reads

测试时物理诊断何时有价值?冻结策略获取了它从未读取的证据
Zhang, Zhengshu
Abstract
When a robot faces unfamiliar physical conditions, a common approach is to collect evidence about what changed and adapt. For such diagnosis to improve behavior, six ordered empirical conditions must hold: a meaningful reference, identifiability of the physical condition, use of the acquired evidence, decision value, selection value over a fixed alternative, and safe realization. We test this chain in controlled and public environments. It holds end to end in our controlled environments. After transfer to unseen mechanisms, however, it breaks at evidence use. On decisions requiring the full trace, the frozen decoder does not change its choice. A linear model using only trace increments recovers the correct choice on mechanisms excluded from fitting, showing that the trace is informative but unused. The failure is concentrated at the richest evidence level: those decisions fall to chance, while decisions settled with lower-cost evidence remain correct, a split hidden by aggregate accuracy. The same chain can fail at other links in public environments. Successful physical identification therefore guarantees neither evidence use nor useful adaptation; evaluation should identify where the chain breaks rather than rely on recovery accuracy or aggregate performance alone.
Chinese Translation
当机器人面对陌生的物理条件时,一种常见的方法是收集关于发生了什么变化的证据并进行适应。要使这种诊断能够改善行为,必须满足六个有序的实证条件:有意义的参照基准、物理条件的可辨识性、对所获证据的使用、决策价值、相对于固定备选方案的选择价值,以及安全的实现方式。我们在受控环境和公开环境中检验了这一条件链。在我们的受控环境中,该链条端到端成立。然而,在迁移到未见过的机构之后,链条在证据使用环节断裂。在需要完整证据轨迹的决策上,冻结的解码器不会改变其选择。仅使用轨迹增量的线性模型在拟合时未包含的机构上恢复了正确选择,这表明轨迹蕴含信息但未被使用。失败集中在信息最丰富的证据层面:这些决策降至随机水平,而依靠低成本证据即可判定的决策仍保持正确——这一分化被总体准确率所掩盖。在公开环境中,同样的链条也可能在其他环节断裂。因此,成功的物理辨识既不能保证证据被使用,也不能保证产生有用的适应;评估应当识别链条在何处断裂,而不是仅仅依赖恢复准确率或总体性能。
cs.RO / 7 / 2609.22317

EditWM: Event-Decomposed World Modeling with Incremental Correction for End-to-End Autonomous Driving

EditWM:面向端到端自动驾驶的事件分解世界模型与增量修正方法
Yang, Junjie, Zeng, Qingwei, Li, Youyou, Ding, Zicheng, Shi, Ziyi, Shen, Shuqi, Lu, Hongliang, Yang, Hai
Abstract
World models support autonomous driving by predicting the scene evolution associated with candidate trajectories. Driving dynamics differ in predictability, motivating a distinction between regular evolution and event-induced deviations that call for selective correction. We propose EditWM, a world model that decomposes future prediction into normal evolution and event-driven incremental correction in compact visual feature space. A trajectory-conditioned normal predictor provides the base forecast and is then frozen for correction learning. A correction decoder compares this forecast with observation history and planned actions, producing a bounded feature update whose contribution is regulated by a learned gate. The corrected future features condition trajectory scoring through candidate-specific cross-attention, linking world modeling to plan selection. At inference, EditWM uses only past and current observations, ego state, and candidate trajectories. Across all 12,146 NAVSIM navtest scenes, expert-trajectory-conditioned evaluation shows a 5.35\% reduction in future-feature MSE over Normal, with improvements in 83.54\% of scenes. The system achieves 91.05 EPDMS on a 100-point scale using the official EPDMS evaluator. These results demonstrate improved future-feature prediction and competitive trajectory selection when corrected future representations are integrated into planning.
Chinese Translation
世界模型通过预测与候选轨迹相关联的场景演化来支持自动驾驶。驾驶动态的可预测性各不相同,因此有必要区分规则演化与事件引起的偏差,并对后者进行选择性修正。我们提出EditWM,一种将未来预测分解为正常演化和事件驱动的增量修正的世界模型,其在紧凑的视觉特征空间中运行。轨迹条件化的正常预测器提供基础预测,随后被冻结用于修正学习。修正解码器将该预测与观测历史和规划动作进行比较,生成一个有界的特征更新,其贡献由一个可学习的门控机制调节。修正后的未来特征通过针对各候选轨迹的交叉注意力机制来条件化轨迹评分,从而将世界建模与规划选择相连接。在推理阶段,EditWM仅使用过去和当前的观测、自车状态以及候选轨迹。在NAVSIM navtest的全部12,146个场景中,专家轨迹条件化评估显示,未来特征MSE相比Normal降低了5.35%,且在83.54%的场景中均有改进。使用官方EPDMS评估器,该系统在满分100分的标准下达到91.05 EPDMS。这些结果表明,当修正后的未来表征被整合到规划中时,未来特征预测得到改进,且轨迹选择具有竞争力。
cs.RO / 8 / 2609.22319

Embedding Physics Priors in Robot Learning: A Survey

在机器人学习中嵌入物理先验:综述
Piccinini, Mattia, Schulze, Lucas, Plebe, Alice, Saveriano, Matteo, Beckers, Thomas, Gao, Yuan, Arenz, Oleg, Zarrouki, Baha, Wang, Dingrui, Schäfer, Finn Rasmus, Peters, Jan, Betz, Johannes, Papini, Gastone Pietro Rosati
Abstract
The rapid progress of artificial intelligence is reshaping robotics and accelerating the adoption of learning-based approaches. While purely data-driven methods have achieved remarkable success in computer vision and natural language processing, robotics remains constrained by limited data, complex real-world interactions, and the need for reliable operation. These challenges have motivated the exploration of physics-embedded robot learning, which embeds physics priors into learning algorithms. By encoding the underlying physical laws and constraints, physics priors can complement limited data with robotics-specific inductive biases, potentially improving generalization, interpretability, and sample efficiency. However, the literature on physics-embedded robot learning remains fragmented across terminology, methodologies, and application domains, making it difficult to assess this growing body of work. This survey reviews physics-embedded robot learning across a broad range of physics priors, robotics applications, and machine learning models, from single-layer perceptrons to generative foundation models. We adopt a unified taxonomy that classifies existing approaches according to their physics embedding: physics-guided inputs, data, and representations; physics-encoded model architectures; and physics-informed training loss functions. Building on this taxonomy, we review methods for robot dynamics learning, trajectory planning, prediction, control, and estimation, together with the corresponding open-source software ecosystem. We identify key open challenges, and outline promising future research directions. Overall, we argue that physics priors provide a particularly relevant robotics-specific inductive bias, complementing rather than replacing data-driven learning, and paving the way toward more generalizable, data-efficient, and trustworthy robotic systems.
Chinese Translation
人工智能的快速发展正在重塑机器人学,并加速基于学习的方法的采用。尽管纯数据驱动的方法在计算机视觉和自然语言处理领域取得了显著成功,但机器人学仍然受到数据有限、复杂真实世界交互以及对可靠运行需求的制约。这些挑战促使研究者探索物理嵌入的机器人学习,即将物理先验嵌入到学习算法中。通过编码潜在的物理定律和约束,物理先验可以为有限的数据补充机器人领域特有的归纳偏置,从而有望提升泛化能力、可解释性和样本效率。然而,关于物理嵌入机器人学习的文献在术语、方法和应用领域上仍较为分散,使得评估这一不断增长的研究体系变得困难。本综述从广泛的物理先验、机器人应用和机器学习模型(从单层感知机到生成式基础模型)出发,回顾了物理嵌入的机器人学习。我们采用统一的分类体系,根据物理嵌入方式对现有方法进行分类:物理引导的输入、数据与表示;物理编码的模型架构;以及物理 informed 的训练损失函数。基于该分类体系,我们回顾了机器人动力学学习、轨迹规划、预测、控制与估计的方法,以及相应的开源软件生态。我们指出了关键的开放性挑战,并展望了有前景的未来研究方向。总体而言,我们认为物理先验提供了一种与机器人学尤为相关的领域特有归纳偏置,它是对数据驱动学习的补充而非替代,并为构建更具泛化能力、数据高效且可信的机器人系统铺平了道路。
cs.RO / 9 / 2609.22325

ReliCAD: From Uncertain LLM Generation to Reliable Parametric CAD Modeling

ReliCAD:从不确定的大语言模型生成到可靠的参数化CAD建模
Zheng, Peng, Dong, Xintong, Li, Chuanyang, Jing, Jiaxin, Han, Chuqi, Shen, Hailong, Song, Yanzhi, Yang, Zhouwang
Abstract
Large language models have shown considerable potential for natural-language-driven parametric CAD modeling. However, a fundamental contradiction exists between their probabilistic generation and the deterministic requirements of CAD modeling, resulting in limitations in reliability, design-intent preservation, and geometric validity. Existing methods typically rely on large-scale annotated datasets, lack explicit modeling of design intent, and underutilize the deterministic capabilities of CAD kernels. To address these limitations, we propose ReliCAD, a unified framework that transforms uncertain LLM generation into reliable parametric CAD modeling. Through explicit design-intent modeling, ReliCAD converts user requirements into structured design specifications and explicitly models geometric relations, topological dependencies, and feature construction order. It then generates constraint-aware parametric instructions and invokes the CAD kernel through an Agent-ready API to perform geometric construction and constraint solving. ReliCAD further records runtime evidence and employs a verification-feedback mechanism to assess consistency between the generated model and the design specifications, enabling error localization and iterative repair. Experiments on the public HistCAD generation dataset and our multi-granularity CAD editing dataset demonstrate that ReliCAD significantly outperforms baseline methods, achieving 99.8\% validity rate and 0.8753 IoU. ReliCAD provides a verifiable, repairable, and generalizable approach to natural-language-interactive CAD modeling.
Chinese Translation
大语言模型在自然语言驱动的参数化CAD建模方面展现出巨大潜力。然而,其概率化生成方式与CAD建模的确定性要求之间存在根本性矛盾,导致在可靠性、设计意图保持和几何有效性方面存在局限。现有方法通常依赖大规模标注数据集,缺乏对设计意图的显式建模,并且未能充分利用CAD内核的确定性能力。为解决这些局限,我们提出了ReliCAD,这是一个将不确定的大语言模型生成转化为可靠参数化CAD建模的统一框架。通过显式的设计意图建模,ReliCAD将用户需求转化为结构化的设计规格,并显式建模几何关系、拓扑依赖关系和特征构建顺序。随后,它生成具备约束感知能力的参数化指令,并通过面向Agent的API调用CAD内核执行几何构建与约束求解。ReliCAD还记录运行时证据,并采用验证-反馈机制评估生成模型与设计规格之间的一致性,从而实现错误定位与迭代修复。在公开的HistCAD生成数据集和我们构建的多粒度CAD编辑数据集上的实验表明,ReliCAD显著优于基线方法,达到99.8%的有效率和0.8753的IoU。ReliCAD为自然语言交互式CAD建模提供了一种可验证、可修复且可泛化的方法。
cs.RO / 10 / 2609.22332

AffordanceWAM: Affordance-Aware Joint World-Action Modeling for Robot Manipulation

AffordanceWAM:面向机器人操作的可供性感知世界-动作联合建模
You, Jiadi, Yu, Qize, Chen, Yue, Cai, Minghong, Zhong, Zhide, Wang, Yuran, Ping, Bowen, Liang, Jiaqi, Shen, Zhenhao, Yan, Haodong, Li, Yinchuan, Wu, Ruihai, Qi, Xiaojuan, Chen, Yingcong
Abstract
Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act. Action-labeled robot videos directly supervise control but are costly and limited in diversity, whereas egocentric human videos capture diverse interactions but lack robot actions and differ in embodiment and appearance. We introduce AffordanceWAM, an affordance-aware generative World Action Model that represents object-centric spatiotemporal affordance through Scalar Affordance and Affordance Heatmap, within the generated future World. This representation grounds visual prediction in task-relevant objects and interaction regions for action generation, and provides shared interaction targets across human and robot videos. Built on a pretrained video diffusion Transformer, AffordanceWAM uses separately parameterized World and Action Experts, coupled through Masked Joint Self-Attention, to jointly predict future RGB observations, Scalar Affordance fields, Affordance Heatmaps, and continuous robot actions under a unified flow-matching objective. Human videos supervise all three future-World streams, whereas robot trajectories additionally provide action supervision, enabling transfer without human action labels or retargeting. Experiments on RoboCasa, CALVIN ABC$\rightarrow$D, and real-world manipulation demonstrate consistent gains over RGB-only and robot-data-only baselines. Under fixed robot supervision, RoboCasa performance improves monotonically as affordance-annotated human video scales. These results support affordance as an effective interface for both vision-language-action learning and human-to-robot transfer.
Chinese Translation
可泛化的机器人操作要求预测场景将如何演化、识别交互可行的位置,并确定如何行动。带动作标签的机器人视频可直接监督控制,但成本高昂且多样性有限;而以自我为中心的人类视频捕捉了多样的交互,但缺乏机器人动作,且在身体形态和外观上存在差异。我们提出 AffordanceWAM,一种可供性(Affordance)感知的生成式世界动作模型,它在生成的未来世界中通过标量可供性(Scalar Affordance)和可供性热力图(Affordance Heatmap)来表示以物体为中心的时空可供性。该表示将视觉预测锚定在与任务相关的物体和交互区域上,从而支持动作生成,并为人类视频和机器人视频提供共享的交互目标。AffordanceWAM 构建于预训练的视频扩散 Transformer 之上,采用分别参数化的世界专家(World Expert)和动作专家(Action Expert),通过掩码联合自注意力(Masked Joint Self-Attention)耦合,在统一的流匹配目标下联合预测未来 RGB 观测、标量可供性场、可供性热力图以及连续机器人动作。人类视频监督全部三个未来世界流,而机器人轨迹额外提供动作监督,从而无需人类动作标签或重定向即可实现迁移。在 RoboCasa、CALVIN ABC→D 以及真实世界操作上的实验表明,该方法相较仅使用 RGB 和仅使用机器人数据的基线方法取得了一致的性能提升。在机器人监督固定的条件下,RoboCasa 性能随带可供性标注的人类视频规模的增加而单调提升。这些结果表明,可供性可作为视觉-语言-动作学习和人机迁移的有效接口。
cs.RO / 11 / 2609.22385

Active Spatial Inspection for Effective and Efficient Embodied Exploration

面向高效具身探索的主动空间检查方法
Wang, Wenbin, Bai, Xiang, Wang, Yizhao, Sun, Hang, Ren, Dong, Qin, Jie, Li, Qingquan, Wang, Bing
Abstract
Achieving high task success and efficiency remains a central pursuit in embodied exploration. Existing frameworks typically guide agent behavior through a spatially coarse and indirect assessment of suggestive cues and directions, yet such designs may struggle to judge cue sufficiency and the need for further inspection, leading to early termination or excessive continuation and ultimately reducing task success and efficiency. This work rethinks embodied exploration from a spatially explicit standpoint and introduces the spatial-inspection-guided ACtive Exploration (ACE) framework coupling evidence-grounded perception with exposure-informed movement and establishing a spatially resolved decision paradigm. Evidence-grounded perception strengthens inferential validity by integrating granular localization with focused verification, turning suggestive cues into precise decisive visual support. Exposure-informed movement promotes directional judgment through prospective prioritization and retrospective suppression, selecting promising directions for efficient spatial progress. ACE alleviates the tension between preventing early termination and avoiding unnecessary continuation. Extensive experiments demonstrate that ACE achieves 18.0% higher navigation task success and 10.3% higher exploration efficiency for question answering than prior state-of-the-art baselines, advancing effective and efficient embodied exploration.
Chinese Translation
实现高任务成功率与高效率始终是具身探索的核心追求。现有框架通常通过对提示性线索和方向进行空间上粗略且间接的评估来引导智能体行为,然而此类设计可能难以判断线索是否充分以及是否需要进一步检查,从而导致过早终止或过度延续,最终降低任务成功率与效率。本工作从空间显式的角度重新思考具身探索,提出了空间检查引导的主动探索(spatial-inspection-guided ACtive Exploration,ACE)框架,该框架将基于证据的感知与基于暴露信息的移动相结合,建立了一种空间分辨的决策范式。基于证据的感知通过将细粒度定位与聚焦验证相融合来增强推理有效性,将提示性线索转化为精确且具有决定性的视觉支持。基于暴露信息的移动通过前瞻性优先排序与回溯性抑制来促进方向判断,选择有前景的方向以实现高效的空间推进。ACE缓解了防止过早终止与避免不必要延续之间的矛盾。大量实验表明,ACE的导航任务成功率比此前最先进的基线方法高出18.0%,问答探索效率高出10.3%,推动了高效具身探索的发展。
cs.RO / 12 / 2609.22404

Behavior Trees for Robotic Systems: An Empirical Study on Practices and Experiences

机器人系统中的行为树:关于实践与经验的一项实证研究
Ghzouli, Razan, Horkoff, Jennifer, Strüber, Daniel, Wohlrab, Rebekka
Abstract
Over the last decade, behavior trees (BT) have become one of the dominant behavior models for coordinating missions of robotic systems. Yet empirical evidence on BT adoption in real-world contexts remains limited, especially regarding practitioners experiences and practices. Without practitioner-grounded evidence from both academic and industrial settings, the research community risks developing guidelines and tools that are plausible in principle but only partly aligned with real challenges. This scarcity of evidence leads to ad-hoc practices, impeding software reuse, maintenance, and evolution. This paper reports a mixed-methods study combining a technical action research investigation at an automotive company with a survey of 34 robotics practitioners. Our results indicate that BTs improve team communication and the understandability of robotic decision-making logic, reflecting BTs practical value beyond mission coordination. At the same time, practitioners face multiple non-trivial design and integration decisions, complicated by the lack of adequate guidelines and tool support. These decisions span architectural and language choices when implementing BTs and integrating them within ROS, for which we report observed patterns. Additionally, practitioners faced multi-factor granularity decisions and mixed experiences with current libraries. For the most used BT libraries, they reported implementation challenges and documentation gaps. Regarding planning algorithms, practitioners find deciding optimal BT node ordering challenging. Views on scalability remained inconclusive, and limitations in current libraries and practices hinder broader adoption. We conclude with cross-cutting observations and implications for both practitioners adopting BTs in robotic systems and researchers aiming to advance empirical understanding of BT adoption.
Chinese Translation
在过去十年中,行为树(Behavior Trees, BT)已成为协调机器人系统任务的主要行为模型之一。然而,关于BT在真实环境中应用的实证证据仍然有限,尤其是关于从业者的经验与实践方面。如果缺乏来自学术和工业环境、以从业者为基础的证据,研究界就可能制定出原则上看似合理、但与实际挑战仅部分契合的指南和工具。这种证据的匮乏导致了临时性的实践方式,阻碍了软件的复用、维护和演化。本文报告了一项混合方法研究,将某汽车公司的技术行动研究调查与对34名机器人领域从业者的问卷调查相结合。我们的结果表明,BT改善了团队沟通以及机器人决策逻辑的可理解性,这体现了BT在任务协调之外的实用价值。与此同时,从业者在设计和集成方面面临多重非平凡决策,而缺乏充分的指南和工具支持使这些决策更加复杂。这些决策涉及实现BT以及将其集成到ROS(Robot Operating System)时的架构和语言选择,我们报告了所观察到的相关模式。此外,从业者还面临多因素影响下的粒度决策问题,并且对现有库的体验褒贬不一。对于使用最广泛的BT库,从业者报告了实现方面的挑战和文档缺失问题。关于规划算法,从业者认为确定最优的BT节点排序具有挑战性。对可扩展性的看法尚无定论,而现有库和实践的局限性阻碍了BT的更广泛应用。最后,我们给出了跨领域的观察结论,对在机器人系统中采用BT的从业者以及旨在深化对BT应用的实证理解的研究者均具有借鉴意义。
cs.RO / 13 / 2609.22462

VLPSA: Vision-Language-Poisson-Safe Actions for Full-Body Safety of Learned Policies

VLPSA:面向学习策略全身安全的视觉-语言-泊松安全动作
Wilkinson, Meg, Fourney, Emily, Burdick, Joel W., Ames, Aaron D.
Abstract
Vision-language-action (VLA) models enable increasingly general-purpose robotic manipulation, but such learned policies do not provide safety guarantees for collision avoidance---especially in environments outside of training distributions. This work presents Vision-Language-Poisson-Safe Actions (VLPSA), a safety filtering framework that provides full-body safety for VLA policies in cluttered and dynamic environments without retraining. VLPSA synthesizes Poisson Safety Functions (PSF) online from perception data, yielding a Control Barrier Function (CBF) that is enforced through a CBF-QP safety filter over the full body and any grasped object, treated as an extension of the final robot link. To enable real-time deployment while maintaining fine spatial resolution in critical task regions, VLPSA combines dual resolutions of this PSF using Boolean CBF compositions. We evaluate VLPSA on SafeLIBERO against safety-filtering baselines, where it achieves the highest collision avoidance rate among the evaluated methods, increasing collision avoidance from 23.1% for the base $\pi_{0.5}$ policy to 91.2% while surpassing its task success rate. We further deploy VLPSA on a Franka FR3 in cluttered scenes with dynamic obstacles and human interference, demonstrating real-time full-body safety during manipulation tasks.
Chinese Translation
视觉-语言-动作(Vision-Language-Action,VLA)模型使机器人操作日益通用化,但此类学习策略无法为碰撞规避提供安全保障——尤其是在训练分布之外的环境中。本工作提出了视觉-语言-泊松安全动作(Vision-Language-Poisson-Safe Actions,VLPSA),这是一个安全过滤框架,无需重新训练即可为 VLA 策略在杂乱和动态环境中提供全身安全保障。VLPSA 根据感知数据在线合成泊松安全函数(Poisson Safety Function,PSF),得到一个控制屏障函数(Control Barrier Function,CBF),并通过 CBF-QP 安全滤波器对机器人全身及任何被抓握物体(视为末端连杆的延伸)实施该约束。为了在保持关键任务区域高空间分辨率的同时实现实时部署,VLPSA 采用布尔 CBF 组合来融合该 PSF 的双分辨率表示。我们在 SafeLIBERO 上将 VLPSA 与安全过滤基线方法进行比较,VLPSA 在所评估的方法中达到最高的碰撞规避率,将基础 π₀.₅ 策略的碰撞规避率从 23.1% 提升至 91.2%,同时超越了其任务成功率。我们进一步将 VLPSA 部署在 Franka FR3 上,在包含动态障碍物和人为干扰的杂乱场景中进行测试,验证了其在操作任务中的实时全身安全性。
cs.RO / 14 / 2609.22483

Tracker-Free Robotic Ultrasound Calibration with a Spherical-Marker Phantom and Threshold-Free Center Localization

基于球形标记体模与免阈值中心定位的无跟踪机器人超声标定
Koo, Kyoungmo, Ma, Guangshen, Wang, Xueding, Draelos, Mark
Abstract
Robotic ultrasound (US) calibration is essential for accurately relating US images to the robot coordinate system, but accurate and automated calibration remains challenging because existing methods often require complex phantoms or external 3D trackers. In this work, we develop a tracker-free robotic US calibration framework using a spherical-marker phantom and a highly automated perception pipeline for sphere-center localization. The proposed threshold-free image-processing method localizes the spherical feature based on each US image's intensity distribution, eliminating hand-tuned intensity thresholds and improving robustness across imaging settings without per-system retuning. Multi-pose observations of the spherical fiducial are then used to estimate the US-to-EE transformation without external tracking or prior localization of the sphere center in the robot base frame. We validate the proposed framework on a robotic US platform through repeated sphere scans and further assess the calibrated system using geometrically distinct phantoms with known CAD models. Across 12 cross-validation folds, single-marker calibration achieved sphere-center accuracy and precision of $1.75\pm0.51$ mm and $0.85\pm0.11$ mm, respectively, compared with $1.72\pm0.51$ mm and $0.84\pm0.12$ mm for three-marker calibration. Across the three reconstruction sets, the single-marker calibration yielded pooled post-registration point-to-surface MAE$\pm$SD values of $0.41\pm0.35$ mm for the cone and $0.50\pm0.41$ mm for the triangular prism, closely matching the three-marker results of $0.40\pm0.35$ mm and $0.48\pm0.40$ mm, respectively.
Chinese Translation
机器人超声(US)标定对于准确建立超声图像与机器人坐标系之间的关系至关重要,但由于现有方法通常需要复杂的体模或外部三维跟踪器,实现精确且自动化的标定仍具有挑战性。在本工作中,我们开发了一种无跟踪的机器人超声标定框架,采用球形标记体模以及高度自动化的球心定位感知流程。所提出的免阈值图像处理方法基于每幅超声图像的灰度分布来定位球形特征,从而无需人工调节的灰度阈值,提高了在不同成像条件下的鲁棒性,且无需针对每个系统重新调参。随后,利用球形基准标志的多位姿观测来估计超声探头到机器人末端执行器(US-to-EE)的变换,无需外部跟踪,也无需在机器人基座坐标系中预先定位球心。我们在机器人超声平台上通过多次重复扫描球形标志对所提框架进行了验证,并进一步使用具有已知CAD模型的几何形状不同体模评估标定后的系统。在12个交叉验证折中,单标记标定达到了1.75±0.51 mm的球心精度(accuracy)和0.85±0.11 mm的精确度(precision),而三标记标定为1.72±0.51 mm和0.84±0.12 mm。在三个重建数据集中,单标记标定得到的配准后点到面MAE±SD分别为圆锥的0.41±0.35 mm和三棱柱的0.50±0.41 mm,与三标记标定的0.40±0.35 mm和0.48±0.40 mm结果高度一致。
cs.RO / 15 / 2609.22493

Layered e-skin for Shear Sensing

用于剪切力感知的分层电子皮肤
Cong, Qingzheng, Devillard, Alexis W. M., Dawood, Abu Bakar, Zhang, Xinxin, Fan, Wen, Dei, Neri Niccolò, Suulker, Cem, Althoefer, Kaspar, Burdet, Etienne, Zhang, Dandan
Abstract
This paper presents a stacked two-layer force-sensing resistor (FSR) array designed for robotic fingertips that combines high-resolution pressure mapping with shear-force estimation. A compliant lattice elastomer spacer converts shear loading into a measurable inter-layer displacement, producing relative center-of-pressure (CoP) shifts between layers. A physics-based moment balance links inter-layer CoP displacement to shear force, while an end-to-end CNN--GRU model captures nonlinear effects from load-dependent compression and contact redistribution. This model, with both layers as input, achieves coefficients of determination $R^2 = 0.914$ for $F_x$ and $R^2 = 0.944$ for $F_y$, consistently outperforming single-layer baselines for shear-force estimation. Robotic manipulation experiments show that, for contact-motion tracking, the deep layer tracks the translation and rotation imposed by the robot arm, whereas the superficial layer tracks the slip at the contact surface. Transient changes in the difference between the total pressure responses of the two layers provide the best slip-event detection performance among the tested cues. These results demonstrate that two-layer FSR arrays can provide three-axis force estimation, contact-motion tracking, and slip-event detection beyond conventional normal-force sensing.
Chinese Translation
本文提出了一种堆叠式双层力敏电阻(FSR)阵列,专为机器人指尖设计,兼具高分辨率压力映射与剪切力估计功能。柔性网格状弹性体间隔层将剪切载荷转换为可测量的层间位移,从而在两层之间产生相对压力中心(CoP)偏移。基于物理力矩平衡模型将层间压力中心位移与剪切力相关联,同时采用端到端的 CNN–GRU 模型捕捉载荷相关压缩和接触重新分布所导致的非线性效应。以两层信号作为输入,该模型在 $F_x$ 上达到决定系数 $R^2 = 0.914$,在 $F_y$ 上达到 $R^2 = 0.944$,在剪切力估计方面持续优于单层基线方法。机器人操作实验表明,在接触运动跟踪中,深层传感器能够跟踪机械臂施加的平移和旋转,而浅层传感器则跟踪接触表面的滑移。在所测试的各类线索中,两层总压力响应之差的瞬态变化提供了最佳的滑移事件检测性能。这些结果表明,双层 FSR 阵列不仅能够实现传统法向力感知之外的三轴力估计,还可进行接触运动跟踪与滑移事件检测。
cs.RO / 16 / 2609.22521

Latent Policy Steering: An Efficient and Flexible Framework for Cross-Embodiment Transfer

潜在策略引导:一种高效灵活的跨本体迁移框架
Wang, Yiqi, Verghese, Mrinal, Schneider, Jeff
Abstract
The performance of learned robot visuomotor policies depends heavily on the size and quality of their training data, yet collecting high-quality demonstrations remains costly for robots in the real world. Although large-scale robot and human datasets are increasingly available, embodiment gaps and mismatched action spaces make them difficult to leverage directly. Cross-embodiment transfer, reusing experience from other embodiments to improve learning on a target embodiment, is therefore crucial for scaling robot learning beyond per-robot data collection. In this work, we find that efficient transfer can be achieved by learning from what is shared across embodiments, the visual dynamics of how the world responds to motion, and by effectively exploiting the scarce target-embodiment data at test time. The proposed framework, called Latent Policy Steering (LPS), implements an embodiment-agnostic pretraining phase, which trains an image-based World Model (WM) with optical flow across diverse embodiments. The resulting WM is finetuned on the target embodiment with robot actions. It then steers the base policy toward better actions by searching in the WM's latent space for plans that stay close to the finetuning data. LPS is a policy-agnostic framework: it can flexibly accommodate different policies without having to retrain them. In Robomimic and real-world evaluations, LPS improves the average performance of Diffusion Policy relatively by 16% and 62%, and Pi0.5 by 8% and 14%, with only 50 demonstrations on an unseen target embodiment.
Chinese Translation
学习型机器人视觉运动策略的性能在很大程度上取决于训练数据的规模和质量,然而在真实世界中为机器人收集高质量演示数据的成本依然高昂。尽管大规模机器人和人类数据集日益增多,但本体差异和不匹配的动作空间使其难以被直接利用。因此,跨本体迁移,即复用其他本体的经验来改进目标本体的学习,对于突破单机器人数据采集的限制、扩展机器人学习至关重要。在本工作中,我们发现,可以通过学习各本体之间的共性,即世界如何响应运动的视觉动力学,并在测试时有效利用稀缺的目标本体数据,来实现高效迁移。我们提出的框架称为潜在策略引导(Latent Policy Steering, LPS),其实现了一个本体无关的预训练阶段,利用光流在多种本体上训练基于图像的世界模型(World Model, WM)。随后,该世界模型在目标本体上使用机器人动作进行微调,并通过在世界模型的潜在空间中搜索与微调数据保持接近的规划,来引导基础策略趋向更优的动作。LPS 是一个策略无关的框架:它可以灵活适配不同的策略,而无需重新训练。在 Robomimic 和真实世界的评估中,在仅提供 50 条未见过的目标本体演示的情况下,LPS 使 Diffusion Policy 的平均性能分别相对提升了 16% 和 62%,使 Pi0.5 分别提升了 8% 和 14%。
cs.RO / 17 / 2609.22538

FRAMES: Failure Recovery And Monitoring of Embodied Skills for Humanoid Loco-Manipulation

FRAMES:面向人形机器人移动操作的具身技能失败恢复与监控
Periasami, Ajay Vikram, Luo, Xinyuan, Li, Haoyu, Cheng, Xianyi
Abstract
Large language model (LLM) planners can decompose natural-language instructions and select reusable robot skills, but choosing the correct skill does not guarantee successful physical execution. This gap is especially important in humanoid loco-manipulation, where errors during approach, grasping, transport, or placement can invalidate the remainder of a long-horizon plan. We present FRAMES, a failure-aware supervisory framework for the Unitree G1 humanoid that operates above the CEER whole-body controller. A Planner Agent selects subtasks through parameterized mid-level skills, while a vision-language-model-based Monitor Agent evaluates each skill using temporal multi-view observations and structured robot and contact evidence. Detected failures stop the active skill and provide grounded feedback to a Recovery Agent. The framework further includes a Memory Module for reusing prior skill experience, and geometric grounding via depth and segmentation. We independently evaluate the monitoring module of the framework in MuJoCo using 100 trials comprising 50 failed and 50 successful executions across five tasks. The monitor detects 48 of 50 failures, correctly accepts 46 of 50 successful executions, and achieves 94.0% overall accuracy. These results provide initial evidence for the monitoring component, while end-to-end evaluation of the complete recovery loop remains ongoing.
Chinese Translation
大语言模型(LLM)规划器能够分解自然语言指令并选择可复用的机器人技能,但选择正确的技能并不能保证物理执行的成功。这一差距在人形机器人移动操作(loco-manipulation)中尤为重要,因为在接近、抓取、搬运或放置过程中出现的错误可能会使长时程规划的剩余部分失效。我们提出了FRAMES,一个面向Unitree G1人形机器人的失败感知监督框架,其运行于CEER全身控制器之上。规划智能体(Planner Agent)通过参数化的中层技能选择子任务,而基于视觉语言模型的监控智能体(Monitor Agent)则利用时序多视角观测以及结构化的机器人状态与接触证据对每项技能进行评估。检测到失败时,系统会停止当前技能,并向恢复智能体(Recovery Agent)提供有依据的反馈。该框架还包含一个用于复用先前技能经验的记忆模块(Memory Module),以及基于深度和分割的几何定位。我们在MuJoCo中独立评估了该框架的监控模块,实验包含五项任务共100次试验,其中50次为失败执行、50次为成功执行。监控器检测出了50次失败中的48次,正确接受了50次成功执行中的46次,整体准确率达到94.0%。这些结果为监控组件提供了初步证据,而完整恢复回路的端到端评估仍在进行中。
cs.RO / 18 / 2609.22587

React When You Need To: Event-Triggered Asynchronous Inference for VLA Policies

按需触发:面向视觉-语言-动作(VLA)策略的事件触发异步推理
Wu, Yansong, Li, Huaqing, Hou, Tianding, Chen, Lingyun, Knoll, Alois
Abstract
Vision-Language-Action (VLA) models commonly predict action chunks, limiting their ability to react to environmental changes during execution. Existing asynchronous inference methods improve reactivity but typically rely on a fixed inference gap. In this paper, we propose an event-guided dynamic inference strategy that adapts the inference gap according to scene changes observed since the previous inference. Thereby, it simultaneously preserves motion consistency and prompt reactivity. Across static and dynamic real-world settings, our method consistently performs best, averaging 95% success and exceeding the strongest baseline by 55 percentage points. The code will be made publicly available upon acceptance. The project page is available at https://react-when-you-need-to.github.io/.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型通常预测动作块(action chunks),这限制了其在执行过程中对环境变化做出反应的能力。现有的异步推理方法虽然提升了反应性,但通常依赖于固定的推理间隔。在本文中,我们提出了一种事件引导的动态推理策略,根据自上次推理以来观测到的场景变化来调整推理间隔,从而同时保持运动一致性和及时的反应性。在静态和动态的真实世界场景中,我们的方法始终表现最佳,平均成功率达95%,超过最强基线55个百分点。代码将在论文录用后公开发布。项目页面见 https://react-when-you-need-to.github.io/。
cs.RO / 19 / 2609.22591

REBOOT: From Failure to Recovery - A Dataset and Benchmark for Precision Assembly

REBOOT:从失败到恢复——面向精密装配的数据集与基准
Ofori-Ampofo, Nana Yaw Owusu, Kahou, Samira Ebrahimi, Thekinen, Joseph
Abstract
Robot learning policies fail in characteristic ways: they stall in uncertain states, drift during contact-rich alignment, and miss targets by millimetres in precision tasks. Yet training datasets consist largely of successful demonstrations, while real-world benchmarks often reduce performance to binary success. This limits both supervision for recovery and analysis of where failures occur. We introduce REBOOT (Recovery Episode Benchmark for Off-nominal Trajectories), the first robot manipulation benchmark designed around failure as a first-class signal. REBOOT contains 2,160 demonstrations across 18 precision assembly tasks, each decomposed into five shared phases: Align(pick), Engage(pick), Transport, Align(place), and Engage(place), enabling phase-level evaluation beyond terminal success. Failures are introduced across phases and paired with expert recovery trajectories that return the system to a valid continuation state. Tasks are annotated with rotational symmetry, engagement-clearance precision tier, and assembly direction through matched install-remove pairs. Failure episodes are labeled by phase and categorical failure mode, enabling attribution to kinematic stage and tolerance violation. Data includes synchronized RGB-D observations from four viewpoints and grounded natural-language descriptions of phase-level success and failure conditions. Half the dataset contains expert demonstrations; the other half contains recovery demonstrations sampled to reflect failures observed in imitation-learned policy rollouts. We benchmark action-chunked transformer, diffusion, and $\pi_0$-FAST policies using phase-level completion rates, revealing model-specific failure points hidden by binary evaluation. Dataset and code: https://nanayawoa.github.io/REBOOT
Chinese Translation
机器人学习策略会以特定方式失败:它们在不确定状态下停滞,在接触密集的对齐过程中漂移,并在精密任务中偏离目标数毫米。然而,训练数据集主要由成功的演示构成,而现实世界的基准测试往往将性能简化为二元成功与否。这既限制了对恢复行为的监督,也限制了对失败发生位置的分析。我们提出了REBOOT(非正常轨迹恢复片段基准,Recovery Episode Benchmark for Off-nominal Trajectories),这是首个将失败作为一等信号来设计的机器人操作基准。REBOOT包含18个精密装配任务共2,160条演示,每个任务被分解为五个共享阶段:Align(pick)、Engage(pick)、Transport、Align(place)和Engage(place),从而实现超越终端成功的阶段级评估。失败被引入到各个阶段,并配有使系统返回到有效续行状态的专家恢复轨迹。任务通过匹配的安装-拆卸对,标注了旋转对称性、配合间隙精度等级和装配方向。失败片段按阶段和分类失败模式进行标注,从而可将失败归因于运动学阶段和公差违规。数据包括来自四个视角的同步RGB-D观测,以及关于阶段级成功与失败条件的接地自然语言描述。数据集中一半为专家演示;另一半为恢复演示,其采样方式反映了模仿学习策略 rollout 中观察到的失败。我们使用阶段级完成率对动作分块的Transformer、扩散模型以及π0-FAST策略进行基准测试,揭示了被二元评估所掩盖的模型特定失败点。数据集与代码:https://nanayawoa.github.io/REBOOT
cs.RO / 20 / 2609.22594

Characterizing Wildlife Response to Biomimetic and Conventional Underwater Vehicles

野生生物对仿生与传统水下航行器响应的特征研究
Pham, Huy, Cai, Levi, Girdhar, Yogesh, Rus, Daniela, Patterson, Zach J.
Abstract
Autonomous underwater vehicles (AUVs) are a promising alternative to divers for scalable collection of natural ocean ecology data. However, these robots may disturb local fauna and cause drastic behavioral differences compared to other monitoring techniques, decreasing their value as scientific tools. A promising prospect is to make AUVs that are more biomimetic, with the hope that taking on the form and behavior of a non-predatory animal may reduce adverse responses. We present the first dataset comparing fish disturbance in response to a conventional thruster-driven AUV, a sea turtle inspired flipper-driven AUV, and a diver. Experiments were conducted at a biodiversity hotspot in a Caribbean reef, and images of the scene were analyzed using computer vision to localize and study fish behavior change. Both AUVs cause measurable changes in fish behavior. Although no between-robot differences remain significant after correction for multiple comparisons, point estimates generally favor the biomimetic AUV, motivating larger studies capable of resolving modest effects and further design changes to optimize for disturbance. We also observe larger responses during diver transects than during AUV transects, although this exploratory comparison is based on a small diver sample with several confounds. While conclusions must be taken as preliminary due to operational and experimental limitations, this study provides the community with a first known dataset and benchmarks to quantify the behavioral impacts of biomimetic robots for ecological monitoring in the wild.
Chinese Translation
自主水下航行器(AUV)是替代潜水员进行可扩展的海洋自然生态数据采集的一种有前景的方案。然而,这些机器人可能会干扰当地动物群,并与其他监测技术相比引起显著的行为差异,从而降低其作为科学工具的价值。一个有前景的方向是使AUV更具仿生性,希望借助模仿非捕食性动物的形态和行为来减少不良反应。我们提出了首个比较鱼类对传统螺旋桨驱动AUV、受海龟启发的鳍状肢驱动AUV以及潜水员的干扰响应的数据集。实验在加勒比海珊瑚礁的一个生物多样性热点地区进行,并利用计算机视觉对场景图像进行分析,以定位并研究鱼类行为的变化。两种AUV均引起了鱼类可测量的行为变化。尽管经过多重比较校正后机器人之间不再存在显著差异,但点估计总体上偏向仿生AUV,这为开展能够解析中等效应的更大规模研究以及进一步优化以减少干扰的设计改进提供了动力。我们还观察到潜水员 transect 期间的鱼类响应大于AUV transect 期间,但这一探索性比较基于较小的潜水员样本且存在多种混杂因素。由于操作和实验条件的限制,本研究的结论应被视为初步性的,但该研究为学界提供了首个可用于量化仿生机器人在野外生态监测中行为影响的数据集和基准。
cs.RO / 21 / 2609.22606

VT-Bridge: Bridging Pretrained Foundation VLAs to VTLAs via Lightweight Residual Adaptation

VT-Bridge:通过轻量级残差适配将预训练视觉-语言-动作基础模型桥接至视觉-触觉-语言-动作模型
Wu, Yansong, Yang, Tuo, Zhao, Rongping, Chen, Lingyun, Chen, Xiao, Li, Junnan, Wu, Fan, Knoll, Alois
Abstract
Vision-Tactile-Language-Action (VTLA) models have demonstrated clear advantages over Vision-Language-Action (VLA) models in contact-rich manipulation. However, developing VTLA models is severely constrained by the massive amounts of vision-tactile data and computational resources required. To address this bottleneck, we propose VT-Bridge, a lightweight residual adaptation strategy that bridges pretrained foundation VLAs to VTLAs. Rather than training a VTLA model from scratch or modifying the original architecture of a pretrained VLA, VT-Bridge employs an identical lightweight residual-adapter architecture across VLA backbones and uses backbone-specific weights to refine actions at the robot execution frequency. This design substantially lowers the data and training barriers. Specifically, it requires up to 50 vision-tactile demonstrations per task to fine-tune a VLA backbone and train a 0.98M-parameter residual adapter. Experiments with three representative VLA backbones ($\pi_0$, $\pi_{0.5}$, and SmolVLA) across four contact-rich manipulation tasks further demonstrate its consistent effectiveness across VLA architectures. On average, VT-Bridge raises the task completion rate from 11.7% with task-level VLA fine-tuning alone to 62.9%. Together, these findings demonstrate the broad applicability, effectiveness, and accessibility of VT-Bridge for contact-rich manipulation. The project page is available at https://hoxnocha.github.io/vt-bridge-web/.
Chinese Translation
视觉-触觉-语言-动作(Vision-Tactile-Language-Action, VTLA)模型在接触密集型操作任务中相比视觉-语言-动作(Vision-Language-Action, VLA)模型展现出明显优势。然而,开发VTLA模型受到海量视觉-触觉数据和计算资源需求的严重制约。为解决这一瓶颈,我们提出VT-Bridge,一种轻量级残差适配策略,可将预训练的基础VLA模型桥接至VTLA。VT-Bridge既不从零训练VTLA模型,也不修改预训练VLA的原始架构,而是在不同VLA主干上采用一致的轻量级残差适配器架构,并使用针对特定主干的权重,以机器人执行频率对动作进行细化。该设计大幅降低了数据和训练门槛。具体而言,每个任务仅需最多50条视觉-触觉演示数据即可微调一个VLA主干并训练一个0.98M参数的残差适配器。在三个代表性VLA主干($\pi_0$、$\pi_{0.5}$和SmolVLA)以及四个接触密集型操作任务上的实验进一步证明了该方法在不同VLA架构上的一致有效性。平均而言,VT-Bridge将任务完成率从仅使用任务级VLA微调时的11.7%提升至62.9%。这些结果表明,VT-Bridge在接触密集型操作中具有广泛的适用性、有效性和易用性。项目页面见https://hoxnocha.github.io/vt-bridge-web/。
cs.RO / 22 / 2609.22608

Becoming a Fruit Ninja: Real-Time Probabilistic Kinodynamic Planning for Manipulator Projectile Interception

成为水果忍者:面向机械臂抛射物拦截的实时概率运动动力学规划
Chen, Lucas, Garrett, Austin, Niu, Andrew, Kingston, Zachary
Abstract
Projectile interception is a challenging dynamic manipulation problem. Intercepting a thrown object with a robot arm requires reaching a point on the object's path as the object passes through it. Slicing also fixes the blade's velocity and orientation at contact. The goal is therefore a subset of the states of the robot and arrival times that moves as the object falls, and the arm must reach it within its actuator limits in milliseconds. We present FRUITNINJA, an anytime sampling-based planner that grows a tree on the GPU in batches toward the interception manifold. Each edge is an exact cubic whose travel time is found by a parallel search against the arm's dynamics, so every edge satisfies the actuator limits. Plans are ranked by a risk-aware objective over the uncertainty in the object's position and the arm's arrival time. We evaluate on a Franka Research 3 against six baselines in a calibrated real-time simulator, where FRUITNINJA cuts 96.7% of tosses in the open and 68.3% among five obstacles, versus the best baseline's 68.3% and 35.0% respectively.
Chinese Translation
抛射物拦截是一个具有挑战性的动态操作问题。用机械臂拦截抛掷物需要在该物体经过其路径上的某一点时抵达该点。切割还要求刀刃在接触时刻具有特定的速度和姿态。因此,目标是机器人状态和到达时间的一个子集,且该子集随物体下落而移动,机械臂必须在毫秒级时间内、在其执行器限制范围内抵达目标。我们提出了FRUITNINJA,这是一种任意时间(anytime)的基于采样的规划器,它在GPU上以批次方式朝拦截流形方向生长一棵树。每条边是一条精确的三次曲线,其行进时间通过与机械臂动力学进行并行搜索来确定,因此每条边都满足执行器限制。规划方案根据一个风险感知目标进行排序,该目标考虑了物体位置和机械臂到达时间的不确定性。我们在一个经过校准的实时仿真器中,使用Franka Research 3机械臂与六个基线方法进行了对比评估。结果显示,FRUITNINJA在开阔环境中成功切割了96.7%的抛掷物,在五个障碍物环境中成功切割了68.3%,而最佳基线方法分别仅达到68.3%和35.0%。
cs.RO / 23 / 2609.22609

From Documented Strengths to Force Limits: Material-Informed Robotic Insertion for Construction Assembly

从文献记载的强度到力限制:面向建筑装配的材料信息驱动的机器人插装
He, Lin, Chen, Yanyi, Sun, Haofei, Li, Lingyao, Deng, Min
Abstract
Insertion is a fundamental operation in robotic construction assembly, where variations in material properties and assembly conditions make it difficult to select contact forces that complete the task without exceeding the assembly's capacity. Although construction documents encode engineering knowledge about materials and their conditions, translating this knowledge into load limits for a specific assembly remains difficult. This paper presents SAGE (Source-grounded Assembly Gating and Execution), a system that converts documented material evidence into capacity estimates for robotic insertion. SAGE restricts a large language model (LLM) to extracting tensile and compressive strengths from retrieved passages and tables and records their sources. A response model then interpolates offline finite element (FE) solutions to convert these strengths and the assembly conditions into axial load capacity. For fits with positive clearance, the estimated capacity sets the policy's axial force limit; for interference fits, it is compared with measured support demand to determine admission. On the primary benchmark, SAGE reduces mean capacity error from 80.65\% for direct LLM estimates based on the same evidence to 10.74\%. Without refitting, the mean error remains 8.00\% on 16 additional geometries. Under the assigned support release model, SAGE correctly classifies 59 of 62 scored simulation runs, with only conservative errors. In recorded xArm6 demonstrations, SAGE takes material documents as input and completes physical insertion in 9 of 13 trials. These results show that assigning document interpretation to the LLM and force calculation to an explicit mechanical model produces accurate capacity estimates and traceable insertion decisions.
Chinese Translation
插装是机器人建筑装配中的一项基础操作,材料属性和装配条件的变化使得选择既能在不超过装配承载能力的情况下完成任务的接触力变得困难。尽管建筑文件中编码了关于材料及其状态的工程知识,但将这些知识转化为特定装配的载荷限制仍然困难重重。本文提出了SAGE(基于文献的装配门控与执行系统,Source-grounded Assembly Gating and Execution),该系统将文献记载的材料证据转换为用于机器人插装的承载能力估计。SAGE限制大语言模型(LLM)仅从检索到的段落和表格中提取抗拉强度和抗压强度,并记录其来源。随后,响应模型对离线有限元(FE)解进行插值,将这些强度和装配条件转换为轴向载荷承载能力。对于具有正间隙的配合,估计的承载能力设定策略的轴向力限制;对于过盈配合,则将其与测得的支撑需求进行比较以确定是否准许执行。在主要基准测试中,SAGE将平均承载能力误差从基于相同证据的直接LLM估计的80.65%降至10.74%。在无需重新拟合的情况下,在16个额外几何结构上的平均误差仍保持为8.00%。在指定的支撑释放模型下,SAGE在62次评分仿真运行中正确分类了59次,且仅有保守性误差。在xArm6的实录演示中,SAGE以材料文档作为输入,在13次试验中有9次完成了物理插装。这些结果表明,将文档解读任务分配给LLM、将力计算任务分配给显式力学模型,能够产生准确的承载能力估计和可追溯的插装决策。
cs.RO / 24 / 2609.22611

HIGenNTO: Scalable Humanoid Interaction Generation via Noise-Space Trajectory Optimization

HIGenNTO:基于噪声空间轨迹优化的可扩展人形机器人交互生成
Jayanti, Lalit, Yamazaki, Kashu, Shibata, Yuto, Amaya, Kotaro, Fragkiadaki, Katerina
Abstract
Humanoid robots can acquire complex skills by imitating kinematic humanoid motion references, yet reliable references for contact-rich interactions remain difficult to obtain: motion capture deteriorates under occlusion and close physical contact, while retargeting introduces additional contact and geometric inconsistencies. We present HIGenNTO, a framework that synthesizes humanoid-scene interaction motion references by optimizing the initial noise of a pretrained text-conditioned motion model under sparse spatiotemporal and scene constraints. The same formulation satisfies desired contacts, avoids collisions, and maintains stable support while retaining the prior's realism and temporal coherence, generating interaction motions from scratch and composing long-horizon behaviors stage-wise. Across robot-environment and robot-object tasks, HIGenNTO produces motions that can be executed by tracking policies in simulation and used to train depth-conditioned visuomotor policies operating solely from onboard sensing. We deploy these policies on a Unitree G1 across four contact-rich tasks. Finally, the task specifications themselves can be written by a coding agent, which proposes interaction tasks and compiles them into prompt, constraint, and scene programs, authoring three of our eight evaluated tasks and four further behaviors. Together, these results establish a scalable path from high-level task descriptions to physically executable humanoid interactions.
Chinese Translation
人形机器人可以通过模仿人形运动学运动参考来获取复杂技能,然而接触密集型交互的可靠参考仍然难以获得:动作捕捉在遮挡和近距离物理接触下性能下降,而动作重定向(retargeting)会引入额外的接触与几何不一致问题。我们提出 HIGenNTO,一个通过在稀疏时空与场景约束下优化预训练文本条件运动模型的初始噪声,来合成人形机器人-场景交互运动参考的框架。同一优化表述能够实现期望的接触、避免碰撞并保持稳定的支撑,同时保留先验模型的真实性与时间连贯性,既能从零生成交互动作,也能分阶段组合长时程行为。在机器人-环境和机器人-物体任务上,HIGenNTO 生成的动作可以被仿真中的跟踪策略执行,并可用于训练仅依赖机载传感的深度条件视觉运动策略。我们将这些策略部署在 Unitree G1 机器人上,完成了四个接触密集型任务。最后,任务规范本身可以由代码智能体(coding agent)撰写,它提出交互任务并将其编译为提示、约束和场景程序,在我们评估的八个任务中撰写了其中三个,以及另外四种行为。这些结果共同确立了一条从高层任务描述到物理可执行的人形机器人交互的可扩展路径。
cs.RO / 25 / 2609.22630

Simultaneous Forward and Inverse Human-in-the-Loop Optimization

同步正向与逆向人在回路优化
Park, Kyeongwon, Collins, Steven H.
Abstract
Subjective user experience is important to human-robot interaction, but the outcomes users value, and how those preferences vary across individuals and contexts, are often unknown. While inverse learning approaches using human data can help identify user rewards, in many assistive settings the experimental costs of executing a control policy, measuring biomechanical or physiological outcomes, and collecting user feedback often limit the number of queries and optimization iterations. Here, we present Simultaneous Forward and Inverse Human-In-the-Loop Optimization (SFIHILO), which efficiently infers individual-specific reward functions from preferences over human outcomes and identifies a final control policy that maximizes the learned reward. SFIHILO bootstraps a forward model to predict user outcomes from control policies, uses this model for active querying to accelerate inverse reward learning, and then optimizes the final control policy without additional user trials. We score candidate policies by their expected reduction in uncertainty across both forward and inverse beliefs, targeting a regional preference boundary to robustly inform this simultaneous learning process. In simulation, we show that SFIHILO was effective across user heterogeneity, outcome dimensionalities, precision requirements, noise levels, and nonstationarity; compared with mutual information approaches, the proposed active querying strategy significantly improved sample efficiency in inverse learning while preserving forward model accuracy. This approach demonstrates the potential to infer latent human goals, enabling more transferable and effective human-robot interaction.
Chinese Translation
主观用户体验对人机交互至关重要,但用户所重视的结果,以及这些偏好在个体和情境之间的差异,往往是未知的。尽管利用人类数据的逆向学习方法有助于识别用户奖励,但在许多辅助应用场景中,执行控制策略、测量生物力学或生理结果以及收集用户反馈的实验成本,常常限制了查询次数和优化迭代的数量。本文提出了同步正向与逆向人在回路优化(SFIHILO),该方法能够根据用户对结果偏好的反馈高效推断个体特定的奖励函数,并确定最大化所学奖励的最终控制策略。SFIHILO首先引导训练一个正向模型以根据控制策略预测用户结果,然后利用该模型进行主动查询以加速逆向奖励学习,最后在无需额外用户试验的情况下优化最终控制策略。我们通过候选策略在正向与逆向两方面不确定性期望降低的程度对其进行评分,并针对区域偏好边界,以稳健地支持这一同步学习过程。仿真实验表明,SFIHILO在用户异质性、结果维度、精度要求、噪声水平以及非平稳性等不同条件下均表现有效;与基于互信息的方法相比,所提出的主动查询策略在保持正向模型精度的同时,显著提高了逆向学习的样本效率。该方法展示了推断潜在人类目标的潜力,有望实现更具可迁移性和更有效的人机交互。
cs.RO / 26 / 2609.22668

Density-Driven Area Coverage for Nonholonomic Multi-Robot Systems with Safety Guarantee

具有安全保障的非完整多机器人系统密度驱动区域覆盖
Martinez, Julian, Lee, Kooktae
Abstract
Density-Driven Optimal Control (D2OC) provides a principled approach to distributing multi-robot teams over non-uniform spatial distributions. Applying D2OC to nonholonomic robots, however, creates a gap between safety constraints imposed on a reference motion and the physical inputs that determine the actual robot motion. We address this issue by enforcing the safety constraint directly on the robot's physical inputs while preserving the density-driven coverage objective. The proposed framework combines D2OC with a control barrier function safety filter through a feedback-linearizing look-ahead point, allowing safety and actuator limits to be considered together during control. We further derive a safety margin that accounts for the look-ahead geometry, robot footprint, and motion during each control interval. Simulation results show that the proposed method maintains the required physical separation while achieving coverage performance comparable to a conventional reference-tracking approach, which can satisfy safety on the reference motion yet violate the corresponding physical clearance. Experiments on multiple nonholonomic robots in the Robotarium further demonstrate safe execution while driving the robots toward the desired spatial distribution. These results show that enforcing safety directly on the physical inputs can eliminate the mismatch between safety certification and physical robot motion in density-driven multi-robot coverage.
Chinese Translation
密度驱动最优控制(Density-Driven Optimal Control, D2OC)为将多机器人团队分布于非均匀空间分布提供了一种系统化的方法。然而,将 D2OC 应用于非完整机器人时,施加在参考运动上的安全约束与决定机器人实际运动的物理输入之间存在差距。为解决这一问题,我们直接在机器人物理输入上施加安全约束,同时保持密度驱动的覆盖目标。所提出的框架通过前视点反馈线性化,将 D2OC 与控制屏障函数(Control Barrier Function, CBF)安全滤波器相结合,使安全约束和执行器限制能够在控制过程中被统一考虑。我们进一步推导了一个安全裕度,该裕度考虑了前视点几何、机器人足印以及每个控制间隔内的运动。仿真结果表明,所提出的方法在维持所需物理间距的同时,实现了与常规参考跟踪方法相当的覆盖性能,而后者虽能在参考运动上满足安全性,却可能违反相应的物理间距。在 Robotarium 平台上对多个非完整机器人进行的实验进一步证明了该方法能够驱动机器人趋向期望的空间分布,同时保证安全执行。这些结果表明,直接在物理输入上施加安全约束能够消除密度驱动多机器人覆盖中安全认证与物理机器人运动之间的不匹配。
cs.RO / 27 / 2609.22670

AquaWorld: Structure-Consistent Underwater World Generation for Robot Simulation

AquaWorld:面向机器人仿真的结构一致水下世界生成
Peng, Bin, Zhang, Tiandong, Wang, Ruidong, Luo, Min, Wang, Shuo
Abstract
Underwater robot simulation requires diverse environments in which terrain, scene composition, tasks, and currents remain mutually consistent. We present AquaWorld, a world-generation framework that preserves these relationships through shared terrain structure. A language-conditioned plan generates 3D terrain and shared structural references that guide asset placement, task definition, and inflow specification under stochastic variation. The framework incorporates over 10,000 underwater-compatible assets, predicts reusable terrain-conditioned mean-flow fields through a CFD-supervised residual model, and supports conventional underwater vehicles and bio-inspired robotic fish. On 24 paired terrains, structure-consistent randomization produces substantially better cross-factor consistency than independent randomization. In a matched-budget policy-training comparison, structurally coherent randomization achieves a validation success rate 21% higher than independent randomization. In separate physical experiments, a simulation-trained visual navigation policy succeeds in 95% of physical tank trials without updating its perception or control modules. Overall, AquaWorld provides a practical way to generate varied underwater environments while retaining the structural relationships needed for flow simulation and robot learning.
Chinese Translation
水下机器人仿真需要多样化的环境,其中地形、场景构成、任务与水流保持相互一致。我们提出AquaWorld,一个通过共享地形结构来保持这些关系的世界生成框架。语言条件化的规划生成三维地形和共享结构参考,用于在随机变化下指导资产放置、任务定义和入流设定。该框架包含超过10,000个水下兼容资产,通过CFD监督的残差模型预测可复用的地形条件化平均流场,并支持传统水下航行器和仿生机器鱼。在24个配对地形上,结构一致的随机化相比独立随机化产生了显著更好的跨因子一致性。在匹配训练预算的策略训练比较中,结构连贯的随机化比独立随机化的验证成功率高21%。在独立的物理实验中,经仿真训练的视觉导航策略在物理水池试验中达到95%的成功率,且无需更新其感知或控制模块。总体而言,AquaWorld提供了一种实用的方法,能够生成多样的水下环境,同时保留流场仿真和机器人学习所需的结构关系。
cs.RO / 28 / 2609.22677

SHAFT: A Slack-Compensating, Helical-Buckling-Attenuating Flexible-Shaft Transmission for Lightweight Multi-DoF Manipulation

SHAFT:一种用于轻量化多自由度操作的slack补偿、螺旋屈曲衰减柔性轴传动机构
Takahashi, Tomoya, Selvamuthu, Moses Gladson, Tadakuma, Richiro, Tanaka, Kazutoshi
Abstract
Lightweight and slim manipulators enable safe operation in human living environments. Proximal actuation using remote transmission mechanisms, such as wire-driven or Bowden cables, effectively reduces inertia and arm size by relocating motors near the base and transmitting torque to distal joints. Existing approaches either increase mass through additional components, such as pulleys for direction changes, or suffer from reduced transmission efficiency due to friction losses. Flexible shaft transmission avoids both mass increase and excessive friction losses, but faces increasing angular transmission error due to helical buckling caused by slack generated at joint bending. To address this problem, we propose SHAFT: a Slack-compensating, Helical-buckling-Attenuating Flexible- shaft Transmission mechanism. This mechanism compensates for slack through a proximal tensioner, improving the angular transmission error and efficiency of flexible shaft transmission without increasing the moving mass of the arm section. In a transmission path containing four 90-degree bends, the proposed mechanism demonstrated approximately 30% higher efficiency and approximately 65% lower angular transmission error compared to a flexible shaft transmission without a tensioner. Using this mechanism, we fabricated a 6-Degree-of-Freedom (DoF) arm with a 1-DoF gripper manipulator consisting of a rotary module housing motors with a tensioner, and a 280 g weight for the arm module. The proposed manipulator represents a novel remote actuation system for achieving lightweight construction with high efficiency, contributing to the acceleration of safe robot deployment in human environments.
Chinese Translation
轻量且纤细的机械臂能够在人类生活环境中实现安全操作。采用远程传动机构(如钢丝驱动或Bowden线缆)的近端驱动方式,通过将电机置于基座附近并将扭矩传递至远端关节,可有效降低惯量并减小臂部尺寸。现有方法要么因增加额外部件(如用于改变方向的滑轮)而增加质量,要么因摩擦损耗导致传动效率下降。柔性轴传动可同时避免质量增加和过大摩擦损耗,但由于关节弯曲时产生的slack(松弛)会引起螺旋屈曲,导致角度传动误差随弯曲程度增加。为解决这一问题,我们提出了SHAFT:一种slack补偿、螺旋屈曲衰减的柔性轴传动机构。该机构通过近端张紧器补偿slack,在不增加臂部运动质量的前提下,提升了柔性轴传动的角度传动精度和效率。在包含四个90度弯折的传动路径中,与无张紧器的柔性轴传动相比,所提机构的传动效率提高了约30%,角度传动误差降低了约65%。利用该机构,我们制作了一个6自由度(DoF)机械臂,配备1自由度夹爪操作器,由一个内置电机和张紧器的旋转模块组成,臂部模块重量仅为280克。所提出的操作器代表了一种新颖的远程驱动系统,可实现高效率的轻量化结构,有助于推动机器人在人类环境中的安全部署与加速普及。
cs.RO / 29 / 2609.22681

A Direct Rigid Transmission 2-DoF Wrist Extension for Tendon-Driven Hand

用于肌腱驱动手的直接刚性传动二自由度腕部扩展装置
Pang, Yujie, Sakib, Sadman, Faruque, Mohammad Abdullah Al
Abstract
Dexterous manipulation in confined spaces requires local control of hand orientation. Without a wrist, a dexterous hand must obtain this local orientation through coordinated motion of the robot arm, often involving several joints and a more complex end-effector path. We present CRAFT-Wrist, a concentric 2-DoF wrist extension that mounts between a robot arm and the CRAFT Hand without modifying the hand. Two XC430-T240BB-T servos drive sideways and front-back rotation through short rigid transmissions. Our initial hardware prototype actuates both axes under load, achieving a demonstrated workspace of $\pm20^{\circ}$ in radial--ulnar deviation (left--right) and a front-back range of $+80^{\circ}$ in flexion and $-18^{\circ}$ in extension. We characterize the resulting finger-motor loading across several wrist postures and demonstrate how the wrist supplies local orientation in grasping, nail hammering, and blackboard wiping. Demonstrations and assembly instructions are available at https://craft-wrist.github.io/.
Chinese Translation
在狭小空间中的灵巧操作需要对手部姿态进行局部控制。如果没有腕部,灵巧手必须通过机械臂的协调运动来获得局部姿态,这通常涉及多个关节和更复杂的末端执行器路径。我们提出了CRAFT-Wrist,一种同心式二自由度腕部扩展装置,安装在机械臂与CRAFT Hand之间,无需对手进行任何修改。两个XC430-T240BB-T舵机通过短刚性传动机构驱动侧向和前后旋转。我们的初始硬件样机在有负载的情况下驱动两个轴,在尺偏-桡偏(左右)方向实现了±20°的已验证工作空间,在前后方向实现了+80°的屈曲范围和-18°的伸展范围。我们表征了多种腕部姿态下手指电机的负载情况,并演示了该腕部如何在抓取、敲钉子和擦黑板任务中提供局部姿态。演示视频和装配说明可在 https://craft-wrist.github.io/ 获取。
cs.RO / 30 / 2609.22684

StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies

StateMem:面向视觉-语言-动作策略的单状态残差记忆与自适应推理方法
Li, Wenzhuo, Shi, Qiongfeng, Zhou, Yi
Abstract
Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA) policies primarily rely on current observations, limiting historical information retention. Memory-augmented VLAs, such as MemoryVLA, address this limitation with external memory banks but require explicit storage and retrieval. To address these limitations, we propose StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes. A training-free controller adjusts the routing threshold online, while fast correction compensates for stale prefix features during cache reuse. We evaluate StateMem on LIBERO, RoboMemArena, and real-world manipulation tasks. On LIBERO, StateMem achieves an average success rate of 97.6% and reduces the average VLM prefix refresh rate by 20.25% relative to full refresh. In the Occlusion category of RoboMemArena, StateMem achieves the best performance among single-VLA methods, reaching 21.8% Task Success Rate (TSR) and 44.3% Cumulative Success Rate (CSR). Across six real-world manipulation tasks, it achieves +21% in average success rate.
Chinese Translation
依赖记忆的机器人操作通常需要后续动作利用早期交互中的信息。现有的视觉-语言-动作(VLA)策略主要依赖当前观测,限制了对历史信息的保留。诸如MemoryVLA等记忆增强型VLA通过外部记忆库解决了这一局限,但需要显式的存储和检索。为解决这些局限,我们提出了StateMem,一种面向VLA策略的单状态残差记忆框架,该框架利用预测误差通过低秩残差更新持久记忆令牌,并对缓存的前缀进行自适应路由。一个免训练的控制器在线调整路由阈值,同时快速校正机制在缓存复用期间补偿过时的前缀特征。我们在LIBERO、RoboMemArena以及真实世界操作任务上对StateMem进行了评估。在LIBERO上,StateMem取得了97.6%的平均成功率,并将VLM前缀的平均刷新率相比全量刷新降低了20.25%。在RoboMemArena的遮挡(Occlusion)类别中,StateMem在单VLA方法中取得了最佳性能,达到21.8%的任务成功率(TSR)和44.3%的累计成功率(CSR)。在六项真实世界操作任务中,其平均成功率提升了21%。
cs.RO / 31 / 2609.22726

Decentralized Multi-Robot Exploration with Probabilistic Peer Intent and Multi-hop Plan Propagation

基于概率化同伴意图与多跳计划传播的分散式多机器人探索
Jamwal, Saurbh Singh, Chebrolu, Nived, Kalyanakrishnan, Shivaram
Abstract
Efficient coordination under limited communication remains a key challenge in decentralized multi-robot exploration. While centralized approaches benefit from global information sharing, they are often impractical in large-scale or communication-constrained environments. Existing Monte Carlo Tree Search (MCTS)-based approaches, such as Decentralized Monte Carlo Exploration (DMCE), enable decentralized planning by taking peer intent into account. This peer intent is obtained by communicating sequences of planned waypoints with robots within direct communication range. In this work, we extend this idea by introducing Probabilistic Peer Intent (PPI), which converts peer trajectories into a continuous spatial representation of predicted intent and incorporates it into local MCTS action evaluation. We additionally study the effects of sharing peer intent beyond direct communication range by propagating plans over multiple hops. Experiments across multiple simulated environments and team sizes show that PPI and Multi-hop propagation can each improve decentralized exploration, with their relative benefits depending on environment structure and team size. We also demonstrate the real-world deployment of our method on three robots operating in different environment types.
Chinese Translation
在有限通信条件下实现高效协调仍是分散式多机器人探索中的关键挑战。集中式方法虽然受益于全局信息共享,但在大规模或通信受限环境中往往不切实际。现有的基于蒙特卡洛树搜索(MCTS)的方法,如分散式蒙特卡洛探索(Decentralized Monte Carlo Exploration, DMCE),通过考虑同伴意图实现了分散式规划。这种同伴意图通过与直接通信范围内的机器人交换规划路径点序列来获取。在本工作中,我们扩展了这一思想,提出了概率化同伴意图(Probabilistic Peer Intent, PPI),它将同伴轨迹转换为预测意图的连续空间表示,并将其纳入局部MCTS动作评估中。此外,我们研究了通过多跳传播计划,将同伴意图共享扩展到直接通信范围之外的效果。在多种仿真环境和不同团队规模下的实验表明,PPI和多跳传播各自都能提升分散式探索性能,其相对优势取决于环境结构和团队规模。我们还在三种不同环境类型中部署了三个机器人,验证了该方法的真实世界应用效果。
cs.RO / 32 / 2609.22730

BEACON: Belief-Enabled Adaptive CONtrol for Imitation Learning under Uncertainty

BEACON:面向不确定性下模仿学习的基于信念的自适应控制
Lee, Moonyoung, Bhattacharya, Soumojit, Kantor, George, Kroemer, Oliver
Abstract
Robot manipulation tasks often involve hidden state information that cannot be directly observed and must be inferred through sequential physical interactions. In such partially observable settings, conditioning an imitation learning policy directly on the recent raw observation history leads to poor performance. This is due to state aliasing, wherein identical observations may arise from different hidden states, and the policy receives conflicting action labels for the same input. To enable history-aware disambiguation capability, we propose conditioning a diffusion policy on a structured representation of the hidden states using Bayesian belief that exposes both the current most likely state estimate and the remaining uncertainty. This representation replaces raw history with a structured, compact input, enabling the policy to implicitly modulate between exploratory and exploitative behaviors based on belief uncertainty, without explicit mode switching or reward shaping. We evaluate across two domains with qualitatively different belief representations: a continuous belief for cornstalk gripper alignment via tactile sensing, and a discrete categorical distribution for latched door opening. In both domains, the belief-conditioned policy substantially outperforms the observation-only baseline and approaches privileged ground-truth performance, with ablations illustrating that the policy adapts its exploration behavior depending on the belief uncertainty at inference time.
Chinese Translation
机器人操作任务通常涉及无法直接观测的隐藏状态信息,必须通过连续的物理交互来推断。在此类部分可观测环境中,直接将模仿学习策略条件化于近期的原始观测历史会导致较差的性能。其原因在于状态混叠(state aliasing),即相同的观测可能来自不同的隐藏状态,导致策略对同一输入收到相互冲突的动作标签。为了赋予策略基于历史的消歧能力,我们提出将扩散策略(diffusion policy)条件化于利用贝叶斯信念(Bayesian belief)构建的隐藏状态结构化表示,该表示同时给出当前最可能的状态估计和剩余的不确定性。这种表示以结构化、紧凑的输入取代了原始历史,使策略能够基于信念不确定性在探索性与利用性行为之间进行隐式调节,而无需显式的模式切换或奖励塑造。我们在两个具有本质不同信念表示的领域中进行评估:一是通过触觉感知进行玉米秸秆夹爪对准的连续信念表示,二是用于开锁拉门任务的离散类别分布。在两个领域中,基于信念条件化的策略均大幅优于仅基于观测的基线方法,并接近具有特权真实状态信息的性能;消融实验表明,策略能够根据推理时的信念不确定性自适应地调整其探索行为。
cs.RO / 33 / 2609.22733

Rapidly-Iterating Grid-Based Near-Optimal Kinodynamic Motion Planning

基于快速迭代栅格的近最优运动学约束运动规划
Moncton, Michael, Frew, Eric
Abstract
This paper develops the Rapidly-iterating kinoDynamic Grid (RDG) algorithm, an asymptotically near-optimal kinodynamic motion planning algorithm that produces high quality solutions through rapid iteration. The algorithm leverages a state space grid decomposition to perform node selection, dynamics propagation, and graph revision in constant time complexity with respect to the number of nodes in the trajectory tree. Through a covering ball sequence induction proof, the algorithm is shown to be asymptotically near-optimal and probabilistically complete. Different subsystems of the algorithm are evaluated against common nearest-neighbor search-based methods at generating exploration bias. The RDG algorithm is evaluated through simulated trials in complex, kinodynamic motion planning problem environments up to 10 DOF relative to similar sparse, kinodynamic planning algorithms with optimality guarantees, SST and DIRT. The RDG algorithm outperforms both SST and DIRT in mean final solution quality by up to 104% and 40% respectively. Additionally, RDG maintained a 100% success rate, even on a 10-DOF test case where both SST and DIRT did not.
Chinese Translation
本文提出了快速迭代运动学栅格(Rapidly-iterating kinoDynamic Grid, RDG)算法,这是一种渐近近最优的运动学约束(kinodynamic)运动规划算法,能够通过快速迭代产生高质量的解。该算法利用状态空间栅格分解来执行节点选择、动力学传播和图修正,其时间复杂度相对于轨迹树中节点数量为常数。通过覆盖球序列的归纳证明,该算法被证明具有渐近近最优性且概率完备。文中针对常见的基于最近邻搜索的探索偏置生成方法,对该算法的不同子系统进行了评估。RDG 算法在最高 10 自由度的复杂运动学约束运动规划问题环境中进行了仿真试验评估,并与具有最优性保证的类似稀疏运动学约束规划算法 SST 和 DIRT 进行了对比。RDG 算法在平均最终解质量上分别超越 SST 和 DIRT 达 104% 和 40%。此外,RDG 保持了 100% 的成功率,即使在一个 10 自由度的测试用例中亦然,而 SST 和 DIRT 在该测试中均未能成功。
cs.RO / 34 / 2609.22786

Distributed Infrastructure Sensors and Cues for Robotic Fire-Fighting and Safety System on a Lunar Base

面向月球基地机器人消防与安全系统的分布式基础设施传感器与标识码
Munguia, Alexis, Rodriguez, Eryc, Russ, Drake, Thangavelautham, Jekan
Abstract
There are renewed efforts to build a lunar base and house a team of astronauts on the Moon for extended periods. However, important challenges remain, such as a rapid-response fire-hazard management system. In the event of a fire or toxic leak inside an early crewed lunar habitat, a single pressurized cylinder must be handled within minutes. However, suited astronauts and radio-delayed support from Earth cannot guarantee such response times. We describe a prototype fire-fighting and emergency response safety technology that embeds intelligence in the habitat itself: using thin ceramic QR codes affixed every few meters, each carries a local floor map, safe sensor limits, and step-by-step hazard responses. A low-power rover decodes a plate in a matter of seconds, polls the adjoining temperature-gas sensor cluster, and acts immediately, eliminating the need for a global map . In a representative layout, the interior is divided into roughly one code per two square meters. Simulations show that rapid-response firefighting both simplifies the task and avoids costly resource use and indeterminate outcomes. Because knowledge is spread across passive plates, damage is localized, procedures can be updated by replacing a single code, and processor demands remain minimal. However, important challenges remain, such as a rapid-response fire-hazard management system. In the event of a fire or toxic leak inside an early crewed lunar habitat, a single pressurized cylinder must be handled within minutes. However, suited astronauts and radio-delayed support from Earth cannot guarantee such response times. We describe a prototype fire-fighting and emergency response safety technology that embeds intelligence in the habitat itself: using thin ceramic QR codes affixed every few meters, each carries a local floor map, safe sensor limits, and step-by-step hazard responses.
Chinese Translation
目前,建造月球基地并让宇航员团队在月球上长期驻留的努力正在重新兴起。然而,仍存在一些重要挑战,例如快速响应的火灾危险管理系统。在早期载人月球居住舱内发生火灾或有毒气体泄漏时,必须在几分钟内处理这个单一增压圆柱体环境。然而,身穿宇航服的宇航员以及来自地球的无线电延迟支援都无法保证这样的响应时间。我们描述了一种将智能嵌入居住舱本身的消防与应急响应安全技术原型:使用每隔几米贴附的薄陶瓷二维码,每个码都携带局部平面图、安全传感器限值以及分步危险应对流程。一台低功耗巡视器(rover)可在几秒钟内解码一个码牌,查询相邻的温度-气体传感器集群,并立即采取行动,从而无需全局地图。在一个代表性布局中,舱内大约每两平方米分布一个码。仿真表明,快速响应消防既简化了任务,又避免了昂贵的资源消耗和不确定的结果。由于知识分散在无源码牌上,损坏影响是局部化的,只需更换单个二维码即可更新处置流程,且对处理器的需求保持在最低水平。然而,仍存在一些重要挑战,例如快速响应的火灾危险管理系统。在早期载人月球居住舱内发生火灾或有毒气体泄漏时,必须在几分钟内处理完毕。然而,身穿宇航服的宇航员以及来自地球的无线电延迟支援都无法保证这样的响应时间。我们描述了一种将智能嵌入居住舱本身的消防与应急响应安全技术原型:使用每隔几米贴附的薄陶瓷二维码,每个码都携带局部平面图、安全传感器限值以及分步危险应对流程。
cs.RO / 35 / 2609.22795

Task-Oriented Co-Design and Optimization of Geared Actuators for Robotic Applications

面向任务的机器人齿轮执行器协同设计与优化
Huang, Xuanyu, Dong, Jianqiang, Zhao, Hang
Abstract
Different tasks performed by legged robots impose distinct torque and speed requirements on actuators. Existing robotic actuators are generally optimized at the component level for metrics such as torque or power density, without explicit task guidance. System-level optimization across components such as motors, gearboxes, and sensors is challenging because of the high computational cost and coupling among mechanical, electrical, and electromagnetic behaviors. Consequently, improvements in individual components may not translate into better robot performance in a specific task. To this end, we present a systematic optimization framework for task-oriented co-design of actuator hardware and control. First, surrogate models are employed to accelerate motor evaluation and support global exploration of the coupled design space. Then, a hierarchical mixed-variable optimization strategy is adopted, combining discrete enumeration with continuous search over dimensions and real-valued indices. These indices are rounded to select admissible values for the remaining discrete choices before each evaluation. Within this search, rated output torque density and task performance are jointly optimized, with Bezier-parameterized joint torque profiles determined for each hardware candidate. Finally, the effectiveness of the proposed framework is validated through actuator fabrication and experiments on a two-degree-of-freedom jumping leg. Based on its measured mass, the fabricated prototype achieves a nominal rated output torque density of 35.7 N m/kg, approximately 60% higher than that of a widely used commercial geared joint actuator, while being 18.6% lighter. Under matched bench conditions, it achieves 12.0% greater jump height at twice-rated torque. Together, these results demonstrate a systematic route from task requirements to actuator design and control.
Chinese Translation
腿式机器人执行的不同任务对执行器提出了不同的扭矩和速度要求。现有的机器人执行器通常在部件层面针对扭矩或功率密度等指标进行优化,缺乏明确的任务导向。由于计算成本高以及机械、电气和电磁行为之间的耦合,跨电机、减速器和传感器等部件的系统级优化极具挑战性。因此,单个部件的改进未必能转化为机器人在特定任务中的性能提升。为此,我们提出了一个系统化的优化框架,用于执行器硬件与控制的面向任务协同设计。首先,采用代理模型加速电机评估,支持对耦合设计空间的全局探索。然后,采用分层混合变量优化策略,将离散枚举与针对尺寸和实值指标的连续搜索相结合。这些指标在每次评估前会被取整,以便为其余离散选项选择可接受的取值。在搜索过程中,额定输出扭矩密度与任务性能被联合优化,并为每个硬件候选方案确定贝塞尔(Bezier)参数化的关节扭矩曲线。最后,通过执行器样机制作和二自由度跳跃腿实验验证了所提框架的有效性。基于实测质量,所制造的样机实现了35.7 N·m/kg的额定输出扭矩密度,比一款广泛商用的齿轮关节执行器高约60%,同时重量减轻18.6%。在相同的台架测试条件下,其在两倍额定扭矩下实现了12.0%更高的跳跃高度。这些结果共同展示了从任务需求到执行器设计与控制的系统化路径。
cs.RO / 36 / 2609.22798

Expert-Play Contouring Control: Faster-than-Demonstration Planning from Slow Expert and Fast Play

专家演示-自由演奏等高线控制:从慢速专家演示与快速自由演奏实现超越演示速度的规划
Cho, Seunghoon, Jung, Wonsuhk, Sangeetha, Sundhar Vinodh, Kousik, Shreyas
Abstract
Expert demonstrations often specify what a robot should do, but not how fast it can do it. Imitation Learning (IL) inherits demonstration timing, while directly accelerating the learned motion can fail when faster execution changes the robot-object dynamics. We study faster-than-demonstration execution as a dynamics-aware control problem and introduce Expert-Play Contouring Control (EPCC), which combines slow expert demonstrations with fast, non-expert play. Expert demonstrations train a latent trajectory generator whose predictions are reparameterized into a time-independent contour of successful task progression, while play trains a world model (WM) of fast-action outcomes. At deployment, our proposed planner uses the WM to optimize actions that makes maximize progress along the expert-derived contour while penalizing deviation from the intended task evolution. Averaged across three visuomotor manipulation tasks, EPCC achieves a $2.0\times$ the throughput of the IL baseline, including $2.2\times$ that of the throughput of the strongest acceleration baseline on a task with interaction-sensitive object dynamics. Our analysis shows that the gains concentrate where faster execution changes robot-object evolution. Together, our results highlight a simple yet effective principle: demonstrations provide task intent, while play data provides the dynamic coverage needed to execute that intent faster.
Chinese Translation
专家演示通常规定了机器人应该做什么,却没有规定它能做得多快。模仿学习(Imitation Learning, IL)继承了演示的时间节奏,而直接加速所学的动作可能在更快的执行改变机器人-物体动力学时导致失败。我们将超越演示速度的执行研究为一个考虑动力学的控制问题,并提出专家演示-自由演奏等高线控制(Expert-Play Contouring Control, EPCC),该方法将慢速的专家演示与快速的非专家自由演奏(play)数据相结合。专家演示用于训练一个潜在轨迹生成器,其预测被重新参数化为与时间无关的成功任务进程等高线(contour);而自由演奏数据用于训练一个关于快速动作结果的世界模型(World Model, WM)。在部署阶段,我们提出的规划器利用该世界模型优化动作,使机器人沿专家导出的等高线最大化任务进度,同时对偏离预期任务演化的行为施加惩罚。在三个视觉运动操作任务上的平均结果显示,EPCC 达到了模仿学习基线 2.0 倍的吞吐量,其中在物体动力学对交互敏感的任务上,达到了最强加速基线 2.2 倍的吞吐量。我们的分析表明,性能提升集中在更快执行会改变机器人-物体演化的场景中。综上所述,我们的结果揭示了一个简单而有效的原则:演示提供任务意图,而自由演奏数据提供了更快执行该意图所需的动力学覆盖。
cs.RO / 37 / 2609.22803

ARCGym: Benchmarking Deep Reinforcement Learning in Autonomous Robotic Colonoscopy

ARCGym:自主机器人结肠镜检查中深度强化学习的基准测试
Ji, Guanglin, Finocchiaro, Martina, Erleben, Kenny, Yin, Hang
Abstract
Simulations for learning-based autonomous colonoscopic navigation focus mainly on fully actuated capsule robots, failing to capture the contact-rich navigation of long and flexible clinical colonoscopes. We present the Autonomous Robotic Colonoscopy Gym (ARCGym), an open-source reinforcement learning environment and benchmark for image-based navigation in clinically derived deformable colon anatomies. ARCGym supports multiple types of colonoscope robots, spanning capsule robots and flexible endoscopes, with this work focusing on flexible endoscopes including magnetic-driven tip actuation and clinically used proximally translational actuation. This work includes five CT-reconstructed colons representing typical clinical scenarios, a set of clinically meaningful navigation subtasks, and unified success metrics. We introduce a reward combining depth-based lumen alignment with a lumen-visibility score to improve learning under occlusions. Experiments across tasks, robots, and anatomies show that autonomous navigation remains challenging for both magnetic-driven and proximal-insertion flexible robots, with proximal-insertion actuation remaining an open problem.
Chinese Translation
现有的基于学习的自主结肠镜导航仿真主要针对全驱动胶囊机器人,未能捕捉细长柔性临床结肠镜在丰富接触环境下的导航特性。我们提出了自主机器人结肠镜检查 gym(Autonomous Robotic Colonoscopy Gym, ARCGym),这是一个开源的强化学习环境与基准,用于在基于临床数据构建的可变形结肠解剖结构中进行基于图像的导航。ARCGym 支持多种类型的结肠镜机器人,涵盖胶囊机器人和柔性内窥镜,本工作重点研究柔性内窥镜,包括磁性驱动的尖端致动以及临床常用的近端平移致动方式。本工作包含五个由 CT 重建的结肠模型,代表典型临床场景,并提供了一组具有临床意义的导航子任务和统一的成功评价指标。我们引入了一种奖励函数,将基于深度的肠腔对齐与肠腔可见性评分相结合,以改善遮挡情况下的学习效果。跨任务、机器人和解剖结构的实验表明,自主导航对于磁性驱动和近端插入式柔性机器人而言仍然具有挑战性,其中近端插入式致动仍是一个有待解决的开放性问题。
cs.RO / 38 / 2609.22809

Kinematic Interface for the Wild: Modular Bimanual Loco-Manipulation Capture from 360$^{\circ}$ Cameras Alone

面向野外环境的运动学接口:仅使用360度相机的模块化双臂移动操作数据采集
Yang, Benjamin, Wang, Weiying, Li, Shenggao, Yan, Keming, Wilkinson, Sasha, Wang, Zelin, Yeung, Yip Fun, Sun, Lingfeng
Abstract
A wrist-mounted camera for UMI-style data collection must do two jobs: record the manipulation and localize in the scene. Most handheld devices localize online from workspace-facing views crowded by hands and objects, or add dedicated tracking hardware. Room-scale bimanual capture therefore still tends to instrument the operator or the scene for accurate localization. We present KIWI (Kinematic Interface for the Wild), a capture kit whose only electronics are off-the-shelf cameras. Our core system splits the two jobs across the two lenses of a 360-degree camera. The rear lens faces the room and builds a shared metric map that registers both hands, and an optional head camera, in one frame without workspace co-visibility; the front lens records the manipulation, and offline IMU fusion bridges front-lens tracking loss. Through our quick-release plate, the camera module attaches to chopstick grippers, parallel-jaw grippers, hand-wrist mounts, or robot flanges. Across six bimanual recordings, combining the rear and front lenses failed to localize only 0.1% of query frames, whereas front-only bimanual feature alignment failed on 24.8% of frames and lost one recording entirely; against evaluation fiducials, localization error stayed within 4.5 mm. KIWI's recovered poses were sufficiently consistent for the four wrist streams alone to reconstruct the scene as a 3D Gaussian splat.
Chinese Translation
用于UMI风格数据采集的腕部安装相机必须同时完成两项任务:记录操作过程并在场景中定位。大多数手持设备依赖面向工作空间的视野进行在线定位,但该视野常被手和物体遮挡,或需要额外的专用跟踪硬件。因此,房间尺度的双臂数据采集通常仍需对操作者或场景进行标记以实现精确定位。我们提出了KIWI(Kinematic Interface for the Wild),一套唯一电子设备仅为现成相机的采集套件。我们的核心系统将两项任务分配给360度相机的两个镜头:后置镜头朝向房间,构建一个共享的度量地图,将双手以及可选的头戴相机统一配准到同一坐标系中,无需工作空间共视;前置镜头记录操作过程,并通过离线IMU融合弥补前置镜头的跟踪丢失。通过快装板,相机模块可安装在筷式夹爪、平行夹爪、手腕支架或机器人法兰上。在六段双臂数据采集中,结合后置与前置镜头的方案仅有0.1%的查询帧定位失败,而仅使用前置镜头的双臂特征对齐方案在24.8%的帧上失败,且完全丢失了一段采集数据;相对于评估基准标记,定位误差保持在4.5毫米以内。KIWI恢复的位姿具有足够的一致性,仅凭四路腕部数据流即可将场景重建为3D高斯泼溅(3D Gaussian Splatting)。
cs.RO / 39 / 2609.22813

Commonsense-Grounded Path Planning from Abstract Instructions

基于常识的抽象指令路径规划
Endo, Masafumi, Honda, Kohei, Yonetani, Ryo
Abstract
We present \emph{commonsense ranked search} (CoRS), a novel path planner that turns an abstract instruction into a route that follows commonsense. While existing methods respect the considerations written down in advance, a robot working among people must follow those left unstated too, as with a wet floor that a worker avoids without being told. CoRS leverages large language models (LLMs) and vision-language models (VLMs) as commonsense knowledge to reason about these latent considerations in its planning. Given an abstract instruction (\emph{e.g.}, ``move carefully'') and visual observations of each region in the environment, CoRS derives a consideration for each region, as in ``this wet floor is slippery and worth a detour.'' It then compares the considerations between regions to see which of the two the robot should avoid more, as in ``the crowd is worse than the wet floor.'' These judgments sort the regions into a commonsense ranking, whose costs drive a conventional search that always returns a valid route. We build a benchmark for planning under latent considerations, with three environments, 1350 problems, and five instructions at three levels of abstraction. Experiments show that CoRS discovers the unstated considerations and goes around the ones worth a detour while crossing the rest, a behavior that recent LLM-based planners do not achieve.
Chinese Translation
我们提出了常识排序搜索(Commonsense Ranked Search,CoRS),这是一种新颖的路径规划器,能够将抽象指令转化为符合常识的行进路线。现有方法仅遵循事先明确写下的考量因素,而在人群中工作的机器人还必须遵守那些未被明说的考量,例如工人无需被告知便会绕开湿滑的地面。CoRS 利用大语言模型(LLM)和视觉-语言模型(VLM)作为常识知识,在规划过程中对这类潜在考量进行推理。给定一条抽象指令(例如"小心移动")以及对环境中每个区域的视觉观测,CoRS 会为每个区域推导出相应的考量因素,如"这块湿地面很滑,值得绕行"。随后,它比较各区域之间的考量,判断机器人更应避开哪一个,如"人群比湿地面的威胁更大"。这些判断将各区域排序为一个常识排名,其代价驱动常规搜索算法,从而始终返回一条有效路线。我们构建了一个针对潜在考量规划问题的基准测试,包含三种环境、1350 个问题以及三种抽象程度下的五条指令。实验表明,CoRS 能够发现未被明说的考量因素,绕开值得规避的区域并穿过其余区域,这是近期基于 LLM 的规划器所无法实现的行为。
cs.RO / 40 / 2609.22829

Whole-Body UMI: Transferring UMI Manipulation Skills to Humanoid Whole-Body Manipulation via Real-Time Motion Generation

全身UMI:通过实时运动生成将UMI操作技能迁移至人形机器人全身操作
Nai, Yuxuan, Chang, Leixin, Yang, Liangjing, Yang, Shuo, Li, Zhongyu
Abstract
Collecting whole-body demonstrations for humanoid manipulation mostly relies on teleoperation, which is costly and hard to scale up. The Universal Manipulation Interface (UMI) provides a scalable data collection paradigm, but end-effector trajectories alone underdetermine humanoid whole-body coordination, which is insufficient for whole-body demonstration collection. Therefore, we introduce Whole-Body UMI (WB-UMI), a task-agnostic, real-time and end-effector conditioned motion generator that decouples whole-body coordination learning from task semantics learning through a shared end-effector interface. A diffusion policy learns from native UMI demonstrations, while WB-UMI learns independently from retargeted motion capture, requiring no body trackers or paired image--whole-body demonstrations during task-specific data collection. In real deployment, an asynchronous hierarchy integrates the diffusion policy, motion generator, and a whole-body controller with latency compensation and measured-state feedback. Real-robot experiments on G1 support real-time closed-loop transfer across four tasks, achieving 90% success in drawer closing, 80% in shelf pick-and-place, 30% in ball toss, and 40% in Loco-PnP, which shows the effectiveness of this hierarchy in transferring native UMI skills to humanoid whole-body manipulation.
Chinese Translation
为人形机器人操作收集全身示范数据主要依赖遥操作,这种方式成本高且难以规模化。通用操作接口(Universal Manipulation Interface, UMI)提供了一种可扩展的数据收集范式,但仅有末端执行器轨迹不足以确定人形机器人的全身协调,因而无法满足全身示范数据收集的需求。为此,我们提出了全身UMI(Whole-Body UMI, WB-UMI),这是一个任务无关、实时且以末端执行器为条件的运动生成器,通过共享的末端执行器接口将全身协调学习与任务语义学习解耦。扩散策略从原生UMI示范数据中学习,而WB-UMI则独立地从重定向后的动作捕捉数据中学习,因此在任务特定的数据收集过程中无需身体跟踪器或配对的图像-全身示范数据。在真实部署中,一个异步分层架构集成了扩散策略、运动生成器和具有延迟补偿与测量状态反馈的全身控制器。在G1机器人上的真机实验支持四个任务中的实时闭环迁移,在关门抽屉任务中达到90%的成功率,货架抓放任务中为80%,抛球任务中为30%,Loco-PnP任务中为40%,证明了该分层架构在将原生UMI技能迁移至人形机器人全身操作方面的有效性。
cs.RO / 41 / 2609.22840

ForceRFT: Refining VLA Actions through Force-Guided Residual Reinforcement Learning

ForceRFT:通过力引导残差强化学习优化VLA动作
Wang, Yichen, Zhang, Chaoyang, Su, Xuqi, Ma, Jun, Zhu, Haiyue, Li, Xiaocong
Abstract
Force-conditioned vision-language-action (VLA) policies can respond to contact, but when trained solely on demonstrations, their recovery behavior may be limited by demonstration coverage, and they do not learn from deployment outcomes. Human corrective imitation provides additional recovery examples, but its objective matches local action targets without explicitly optimizing task return. We present ForceRFT, a force-guided residual reinforcement learning framework that learns contact-dependent corrections from human supervision and autonomous task outcomes. A frozen, demonstration-trained SmolVLA-based prior generates force-conditioned action chunks, while a lightweight residual actor refines individual end-effector pose commands using wrist feedback acquired during chunk execution. The decision-time wrist wrench, its temporal change, and the selected base motion condition both residual correction and value estimation. Human corrections supervise the residual actor, while verified autonomous transitions train the twin critics and support value-guided updates to the same actor. Bootstrapping is restricted to autonomous segments, preventing TD credit from crossing human-intervention boundaries. Real-robot experiments on plug insertion, ring-on-peg assembly, and whiteboard wiping show higher autonomous success rates than the evaluated demonstration-trained and residual-imitation baselines. Comparisons with residual imitation support value-guided residual optimization, while plug-insertion ablations indicate the benefit of direct execution-time wrist feedback.
Chinese Translation
力条件化的视觉-语言-动作(VLA)策略能够对接触作出响应,但当仅通过演示数据进行训练时,其恢复行为可能受限于演示数据的覆盖范围,并且无法从部署结果中学习。人类纠正性模仿提供了额外的恢复示例,但其目标仅匹配局部动作目标,而未显式优化任务回报。我们提出了ForceRFT,这是一个力引导的残差强化学习框架,能够从人类监督和自主任务结果中学习依赖接触的修正。一个冻结的、基于SmolVLA并由演示数据训练的先验模型生成力条件化的动作块(action chunks),而一个轻量级的残差动作器(residual actor)利用在动作块执行期间获取的腕部反馈来细化单个末端执行器位姿指令。决策时刻的腕部力/力矩、其时间变化以及所选的基础运动,共同条件化残差修正与价值估计。人类纠正数据监督残差动作器,而经验证的自主转移数据训练双 Critic(twin critics),并支持对同一动作器进行价值引导的更新。自举(bootstrapping)被限制在自主片段内,从而防止TD信用跨越人类干预边界。在插头插入、套环装配和白板擦拭任务上的真实机器人实验表明,所提方法的自主成功率高于所评估的演示训练基线和残差模仿基线。与残差模仿方法的对比验证了价值引导残差优化的有效性,而插头插入任务的消融实验则表明了在执行时直接利用腕部反馈的收益。
cs.RO / 42 / 2609.22852

Physical-Touch Observability from Wrist Wrench in Granular Scooping

从腕部力觉观测颗粒物料铲掘中的物理接触可观测性
Lin, Hongyi, Zhang, Song, Liu, Xubo, Liu, Yang
Abstract
Mining and earthmoving are important real-world deployment settings for embodied intelligence. Autonomous transport and driving systems have improved substantially, but loading and scooping still often depend on skilled human operators, exposing personnel and equipment to operational risk. For robotic scooping, pre-contact RGB-D sensing reveals surface geometry but not the resistance, compaction, tool engagement, or load transfer that emerge during interaction. We test whether the current scoop's six-axis wrist force/torque (F/T), or wrist wrench, contains information about final collected volume, and whether that information depends on the correctly paired action-terrain interaction. We call this property physical-touch observability. Using 6,700 real-robot scoops across 67 terrains, we evaluate correctly paired current-scoop F/T against pre-contact prediction and correspondence-breaking controls under terrain-held-out testing. At the retrospective 60% sequence boundary, correctly paired F/T reduces mean absolute error by 14.2% relative to Action-only and by 21.9% relative to cross-terrain mismatched F/T. Engineered signal summaries reproduce the result across model architectures. Together, these findings position wrist wrench not merely as a low-level feedback signal, but as a task-level perceptual modality through which embodied robots can infer hidden physical states during interaction, providing a foundation for response-aware autonomy in mining and other contact-rich tasks.
Chinese Translation
采矿与土方作业是具身智能重要的现实部署场景。自主运输与驾驶系统已取得长足进步,但装载与铲掘作业仍常依赖熟练的人类操作员,使人员和设备面临操作风险。对于机器人铲掘而言,接触前的RGB-D感知只能揭示表面几何信息,而无法获取在交互过程中涌现的阻力、密实度、工具接触状态或载荷传递等信息。我们检验当前铲斗的六轴腕部力/力矩(F/T),即腕部旋量(wrist wrench),是否包含关于最终采集体积的信息,以及该信息是否依赖于正确配对的动作-地形交互。我们将这一特性称为物理接触可观测性(physical-touch observability)。基于跨越67种地形的6,700次真实机器人铲掘实验,我们在地形留出测试条件下,将正确配对的当前铲掘力/力矩数据与接触前预测以及破坏对应关系的对照实验进行了对比评估。在回溯性60%序列边界处,正确配对的力/力矩数据相较于仅动作(Action-only)基线将平均绝对误差降低了14.2%,相较于跨地形错误匹配的力/力矩数据降低了21.9%。人工设计的信号特征汇总在不同模型架构下均复现了该结果。综合而言,这些发现表明腕部旋量不仅是一种低层反馈信号,更是一种任务级感知模态,具身机器人可借助它在交互过程中推断隐藏的物理状态,为采矿及其他接触密集型任务中的响应感知自主性奠定了基础。
cs.RO / 43 / 2609.22854

SmoLSTM: A Compact Vision-Language-Action Model with Recurrent Memory that Persists

SmoLSTM:一种具有持久循环记忆的紧凑型视觉-语言-动作模型
Habekost, Jan-Gerrit, Kashani, Parsa Mastouri, Gäde, Connor, Kerzel, Matthias, Allgeuer, Philipp, Weber, Cornelius, Wermter, Stefan, Lee, Jae Hee
Abstract
Vision-language-action models often predict actions from only the current observation, which can leave tasks involving object occlusion or visually identical objects ambiguous without episode history. The usual countermeasure, widening the observation window, turns the horizon into a hyperparameter and lets per-step cost grow with it. We instead capture the episode in a recurrent state. SmoLSTM couples a frozen 256M-parameter SmolVLM backbone to a matrix-memory LSTM control layer in which observation tokens and action queries are unified in a single causal stream that is never reset throughout the entire episode. Recurrent-state storage is therefore O(1) in episode length. A flow-matching action head predicts chunks of 10 end-effector pose deltas and gripper commands at each control step. Our single policy, trained jointly on 7,461 demonstrations across 140 tasks and evaluated on held-out initial states, performs best, reaching 85.1% subgoal coverage and 77.5% full-task success on LIBERO-Mem with 0.04B trainable parameters, surpassing both the benchmark's own object-centric baseline and recent memory-based approaches. Resetting the recurrent state at every control step reduces full-task success to 7.0%, showing that the trained policy relies on context carried between decisions. The same model achieves 79.6% average success on standard LIBERO.
Chinese Translation
视觉-语言-动作(Vision-Language-Action)模型通常仅根据当前观测来预测动作,这使得涉及物体遮挡或视觉上相同物体的任务在缺乏回合历史的情况下变得模糊不清。通常的应对方法是扩大观测窗口,但这将时间跨度变成了超参数,并使每步计算成本随之增长。我们转而用循环状态来捕获回合信息。SmoLSTM 将冻结的 2.56 亿参数 SmolVLM 主干与矩阵记忆 LSTM 控制层相耦合,其中观测 token 和动作查询被统一在单一因果序列中,该序列在整个回合中从不重置。因此,循环状态存储相对于回合长度为 O(1)。流匹配(flow-matching)动作头在每个控制步预测 10 个末端执行器位姿增量和夹爪指令的数据块。我们唯一的策略在 140 个任务的 7,461 条演示数据上联合训练,并在留出的初始状态上评估,表现最佳:以 0.04B 可训练参数在 LIBERO-Mem 上达到 85.1% 的子目标覆盖率和 77.5% 的完整任务成功率,超越了该基准自身的以物体为中心的基线以及近期的基于记忆的方法。在每个控制步重置循环状态会将完整任务成功率降至 7.0%,这表明训练出的策略依赖于在决策之间传递的上下文信息。同一模型在标准 LIBERO 上达到 79.6% 的平均成功率。
cs.RO / 44 / 2609.22858

PileBelief: Persistent Physical State for Interaction-Driven World Modeling

PileBelief:面向交互驱动世界建模的持久物理状态
Lin, Hongyi, Zhang, Song, Liu, Haiquan, Liu, Yang, Zhao, Jinhua, Qu, Xiaobo
Abstract
World models allow robots to anticipate action consequences before execution. This capability is especially valuable in excavation, where each scoop reshapes the terrain and affects subsequent actions. Local observations, however, cannot fully reveal the underlying support and material conditions. We present PileBelief, an interaction-driven persistent world model for partially observed excavation that retains physical evidence beyond the visible surface. It combines an observation-conditioned physical prior with world-addressed deformation memory and physical-response memory. Action-aligned reads and gated residual corrections refine terrain-change and outcome predictions. With deployment weights fixed, completed interactions update measured belief, while hypothetical actions advance a separate imagined state. Compared with a current-observation-only baseline, PileBelief reduces five-step joint prediction error by 10.8% and offline action-selection regret by 65.5%. Experiments on Newton/MPM and real excavation datasets further demonstrate improved terrain-change and bucket-volume prediction. Our method enables multi-step prediction and candidate-action ranking from local observations, even when the underlying soil state is unknown. These results identify persistent physical belief as a useful representation for world models of environments that robots continually reshape.
Chinese Translation
世界模型使机器人能够在执行动作之前预判其后果。这一能力在挖掘作业中尤为宝贵,因为每一次铲挖都会重塑地形并影响后续动作。然而,局部观测无法完全揭示潜在的支持状态与物料条件。我们提出 PileBelief,一个面向部分可观测挖掘任务的交互驱动持久世界模型,它能够保留可见表面之外的物理证据。该方法将基于观测条件的物理先验与世界寻址的形变记忆及物理响应记忆相结合,通过对齐动作的读取和门控残差修正来改进地形变化与作业结果的预测。在部署权重固定的情况下,已完成的交互会更新测得的信念,而假设性动作则推进一个独立的想象状态。与仅使用当前观测的基线相比,PileBelief 将五步联合预测误差降低了 10.8%,将离线动作选择遗憾降低了 65.5%。在 Newton/MPM 仿真与真实挖掘数据集上的实验进一步证明了地形变化和铲斗体积预测精度的提升。我们的方法能够在底层土壤状态未知的情况下,仅凭局部观测实现多步预测与候选动作排序。这些结果表明,持久物理信念是机器人持续重塑环境的世界模型的一种有效表征。
cs.RO / 45 / 2609.22871

A Reconfigurable Dual-Opposition Architecture for Single-Hand Assembly and Manipulation

一种用于单手装配与操作的可重构双对置架构
Su, William, Nakamura, Yunosuke, Wang, Yixiao, Li, Yitong, Yu, Mingrui, Qi, Huanan, Liang, Boyuan, Tomizuka, Masayoshi, Zhou, Jianshu
Abstract
In-hand assembly is constrained by the need to maintain grasps on two separate parts while controlling their relative motion within a single hand. To enable both in-hand assembly and manipulation, we present a reconfigurable dual-opposition architecture. Specifically, to support simultaneous grasping of two parts and coordinated in-hand manipulation, four independently actuated fingers are organized into two virtual finger (VF) oppositions, with their relative configuration controlled by a reconfigurable palm. To describe hand motion and simultaneous two-object grasping configurations, a kinematic model of the fingers and palm and an object-size-conditioned workspace formulation are built. To further evaluate motion performance and assembly capability, finger-joint motion and palm tracking are characterized, and in-hand assembly is demonstrated through tasks involving grasping, alignment, fastening, and pressing. Ablation experiments further demonstrate the importance of finger abduction/adduction and palm reconfiguration for successful in-hand assembly. In simulation, the proposed hand achieves a mean continuous sphere rotation success rate of 98.6% over diameters of 40-230 mm, compared with 73.8% for the LEAP Hand. After policy fine-tuning with external disturbances, the proposed hand achieves 92.8% success under disturbances from multiple directions, compared with 45.2% for the LEAP Hand. Hardware demonstrations further show in-hand rotation of objects of different sizes using policies trained in simulation. Together, these results show that the proposed architecture supports both assembly of two separately held parts and coordinated manipulation of a single object within one hand.
Chinese Translation
手内装配的难点在于需要在单手中保持对两个独立零件的抓持,同时控制它们之间的相对运动。为实现手内装配与操作的兼顾,我们提出了一种可重构的双对置架构。具体而言,为支持对两个零件的同时抓取以及协调的手内操作,四根独立驱动的手指被组织为两个虚拟手指(VF)对置结构,其相对构型由一个可重构手掌控制。为描述手部运动及双物体同时抓取构型,我们建立了手指与手掌的运动学模型,以及以物体尺寸为条件的workspace(工作空间)表达式。为进一步评估运动性能与装配能力,我们对手指关节运动和手掌跟踪特性进行了表征,并通过涉及抓取、对准、紧固和按压的任务演示了手内装配。消融实验进一步证明了手指外展/内收和手掌重构对于成功完成手内装配的重要性。在仿真中,所提出的手在40–230 mm直径范围内的连续球体旋转平均成功率达到98.6%,而LEAP Hand为73.8%。经过外部扰动下的策略微调后,所提出的手在来自多个方向的扰动下达到92.8%的成功率,而LEAP Hand为45.2%。硬件演示进一步展示了利用仿真训练的策略对不同尺寸物体进行手内旋转。这些结果表明,所提出的架构既能支持单手中两个独立被抓持零件的装配,也能支持单手内对单一物体的协调操作。
cs.RO / 46 / 2609.22885

Barrier Certificate Synthesis for Non-Polynomial Robotic Dynamics via Polynomial Lifting

基于多项式提升的非多项式机器人动力学障碍证书综合
Chaubey, Shivam, Verdoja, Francesco, Deka, Shankar, Kyrki, Ville
Abstract
Safe operation of robotic systems requires trajectories to remain within a prescribed safe set under admissible control inputs. Barrier certificates provide such guarantees by certifying a controlled-invariant region within that set. Sum-of-squares optimization offers a systematic way to synthesize such certificates, but its direct application requires polynomial dynamics, excluding common robotic nonlinearities, including trigonometric terms. We address this limitation using exact polynomial lifting, which replaces non-polynomial dynamics with polynomial-augmented dynamics subject to lifting-induced algebraic constraints, preserving nonlinear geometry without approximation. We formulate lifted-domain joint barrier synthesis that computes a certificate with a state-feedback control witness and develop a sampled-data safety filter for zero-order-hold implementation. To assess whether the benefits of lifting persist across synthesis frameworks, we also adapt a sample-guided successive-barrier method to the lifted representation. On coordinated-turn and planar multirotor models, exact lifting improves certified coverage in both methods: at matched sample sizes, lifted successive-barrier synthesis achieves higher coverage with fewer barriers and lower computational cost, while lifted joint barrier synthesis provides higher coverage and lower computational cost than the finest tested piecewise resolution. In closed-loop experiments, the safety filter maintains feasibility and safety across all evaluated trajectories, reduces spatial conservativeness, and requires less intervention for both models.
Chinese Translation
机器人系统的安全运行要求轨迹在允许的控制输入下始终保持在指定的安全集内。障碍证书(Barrier Certificates)通过认证该集合内一个受控不变区域来提供此类保证。平方和(Sum-of-Squares)优化为综合此类证书提供了一种系统化方法,但其直接应用要求动力学为多项式形式,从而排除了包括三角函数项在内的常见机器人非线性因素。我们利用精确多项式提升(exact polynomial lifting)来解决这一局限:通过提升引入的代数约束,用多项式增广动力学替代非多项式动力学,在不进行近似的情况下保留非线性几何结构。我们提出了提升域上的联合障碍综合方法,以计算带有状态反馈控制见证(witness)的证书,并开发了一种用于零阶保持(zero-order-hold)实现的采样数据安全滤波器。为评估提升的优势是否在不同综合框架下依然成立,我们还将一种样本引导的逐次障碍方法(successive-barrier method)适配到提升表示中。在协同转弯模型和平面多旋翼模型上的实验表明,精确提升在两种方法中都提高了认证覆盖率:在相同样本数量下,提升的逐次障碍综合以更少的障碍函数和更低的计算成本实现更高的覆盖率;而提升的联合障碍综合相比所测试的最细分段分辨率,具有更高的覆盖率和更低的计算成本。在闭环实验中,安全滤波器在所有评估轨迹上均保持了可行性与安全性,降低了空间保守性,并且对两种模型所需的干预更少。
cs.RO / 47 / 2609.22888

Stable and Efficient Real-World Online VLA Post-Training via Asynchronous Replay-Anchored Policy Improvement

基于异步回放锚定策略改进的稳定高效真实世界在线VLA后训练
Yang, Jiarui, Zhang, Jiajin, Zhu, Bin, Chen, Jingjing, Jiang, Yu-Gang
Abstract
Online post-training of vision-language-action (VLA) models requires efficient use of robot interaction and reliable policy improvement from continually collected experience. We propose asynchronous Replay-Anchored Policy improvement (RAPolicy), a framework that performs rollout and learning concurrently while grounding both critic and actor updates in replayed behavior. The critic learns chunk-level values from recorded actions and constructs Bellman targets without predicting next actions, reducing computation and dependence on action-value estimates outside replay coverage. The one-step flow actor reuses the initial noise stored during rollout and learns through advantage-weighted conditional likelihood, directly supervising the action mapping used for execution. We evaluate RAPolicy across four single-task settings and one joint five-task setting in the real world, with online training budgets of only 1--2 hours. Starting from policies fine-tuned on just 10 demonstrations per task, RAPolicy rapidly adapts to new single tasks and achieves an average 86.3% success rate. In the joint multi-task setting, RAPolicy improves overall success rate from 52% to 88% while preserving performance on already reliable tasks and improving weaker capabilities. Overall, RAPolicy substantially outperforms the baselines in aggregate task success while requiring fewer human interventions, demonstrating stable policy improvement and high online training efficiency. Project page: https://flyfaerss.github.io/RAPolicy.
Chinese Translation
视觉-语言-动作(VLA)模型的在线后训练需要高效利用机器人交互,并从持续收集的经验中实现可靠的策略改进。我们提出了异步回放锚定策略改进(Replay-Anchored Policy improvement, RAPolicy)框架,该框架在并发执行 rollout 与学习的同时,将评论家(critic)和演员(actor)的更新都锚定在回放的行为数据上。评论家从记录的动作中学习片段级(chunk-level)价值,并在构建 Bellman 目标时无需预测下一动作,从而减少了计算量,也降低了对回放覆盖范围之外动作价值估计的依赖。单步流(one-step flow)演员复用 rollout 期间存储的初始噪声,并通过优势加权的条件似然进行学习,直接监督用于执行的动作映射。我们在真实世界中评估了 RAPolicy,涵盖四个单任务设置和一个五任务联合设置,在线训练预算仅为 1-2 小时。从每任务仅 10 条演示微调得到的策略出发,RAPolicy 能快速适应新的单任务,并达到平均 86.3% 的成功率。在多任务联合设置中,RAPolicy 将总体成功率从 52% 提升至 88%,同时保持了已可靠任务的性能,并改进了较弱的能力。总体而言,RAPolicy 在总体任务成功率上显著优于基线方法,同时所需的人工干预更少,展示了稳定的策略改进能力和较高的在线训练效率。项目主页:https://flyfaerss.github.io/RAPolicy。
cs.RO / 48 / 2609.22895

H-VLA: Hierarchical Vision-Language-Action Model with Key-Action Reasoning and Motion Planning in a Unified Action Space

H-VLA:在统一动作空间中结合关键动作推理与运动规划的分层视觉-语言-动作模型
Peng, Xiongfeng, Xu, Lu, Wang, Yandong, Yu, Jiaqian, Zheng, Zirui, Mao, Yamin, Li, Weiming, Chung, Inseop, Cho, Hyun-woong, Yoo, Jaewook, Lee, Dongwook, Ji, Daehyun, Zhang, Chao
Abstract
Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but many existing methods still rely on direct mappings from language and visual observations to dense actions. This formulation can weaken the semantic reasoning capability inherited from pre-trained Vision-Language Models (VLMs), which are mainly optimized for visual-linguistic understanding rather than low-level control, and becomes fragile under spatial variations, including changes in object positions, scene layouts, robot embodiments, and camera viewpoints. To address these limitations, we propose H-VLA, a hierarchical VLA framework that decouples high-level key-action reasoning from low-level motion generation. H-VLA combines a Key-Action Model for predicting a key-action as the next manipulation subgoal, a Motion Planning Model for generating dense future actions conditioned on the predicted key-action, and a Unified Camera-Centric Action Space for consistent representation across datasets, embodiments, and viewpoints. We further adopt a two-stage training strategy that emphasizes key-action reasoning during pre-training and dense motion generation during fine-tuning. Experiments show that H-VLA achieves strong performance on SimplerEnv, reaching 91% on Google Robot visual matching, 84% on Google Robot variant aggregation, and 81% on WidowX visual matching. On Agilex real-robot tasks, H-VLA improves over the strongest baseline by 10, 47, and 16 percentage points under in-distribution, out-of-distribution position, and out-of-distribution scene/object settings, respectively.
Chinese Translation
视觉-语言-动作(VLA)模型在机器人操作领域展现出强大的潜力,但许多现有方法仍依赖于从语言和视觉观测到稠密动作的直接映射。这种建模方式会削弱从预训练视觉-语言模型(VLM)继承而来的语义推理能力——这类模型主要针对视觉-语言理解而非底层控制进行优化——并且在空间变化(包括物体位置、场景布局、机器人本体和相机视角的变化)下表现脆弱。为解决这些局限,我们提出了 H-VLA,这是一个分层 VLA 框架,将高层关键动作推理与低层运动生成解耦。H-VLA 结合了三个组件:用于预测关键动作作为下一个操作子目标的关键动作模型(Key-Action Model)、以预测的关键动作为条件生成稠密未来动作的运动规划模型(Motion Planning Model),以及用于在数据集、机器人本体和视角之间实现一致表示的统一以相机为中心的动作空间(Unified Camera-Centric Action Space)。我们进一步采用两阶段训练策略,在预训练阶段强调关键动作推理,在微调阶段强调稠密运动生成。实验表明,H-VLA 在 SimplerEnv 上取得了优异性能:在 Google Robot 视觉匹配上达到 91%,在 Google Robot 变体聚合上达到 84%,在 WidowX 视觉匹配上达到 81%。在 Agilex 真实机器人任务上,H-VLA 分别在分布内、分布外位置以及分布外场景/物体设置下,较最强基线提升了 10、47 和 16 个百分点。
cs.RO / 49 / 2609.22925

"Dear LLaVA, Please Drive": A Depth-Aware Vision-Language Agent for Closed-Loop Robotic Control

“亲爱的LLaVA,请开车”:一种用于闭环机器人控制的深度感知视觉-语言智能体
Berger, Sebastian, Winter, Katharina, Flohr, Fabian B.
Abstract
Vision-language models (VLMs) provide a compelling foundation for reasoning-driven mobile navigation, offering rich contextual understanding and strong generalization from large-scale pretraining. Most existing navigation frameworks rely on imitation learning and therefore require substantial labeled trajectory data, limiting their scalability and robustness. In this work, we propose a parameter-efficient approach to fine-tune a pretrained VLM for autonomous navigation using an Imperative Learning paradigm. By optimizing against differentiable geometric cost fields rather than labeled trajectories, our model learns to generate collision-free paths exclusively from stereoscopic depth observations. We introduce a unified end-to-end navigation pipeline for natural-language-driven robotic control. This system leverages a shared VLM backbone with task-specific Low-Rank Adaptation (LoRA) modules, effectively bridging the gap from semantic target selection to low-level trajectory planning. Our approach achieves competitive Success weighted by Path Length (SPL) in unseen environments while updating less than 1% of the model's total parameters. Qualitative real-world experiments validate sim-to-real generalization and stable path planning without fine-tuning on real-world data. These results highlight a practical approach for deploying VLM-based agents on mobile robots, enabling high-level semantic navigation without the prohibitive requirement for large-scale, labeled trajectory data.
Chinese Translation
视觉-语言模型(VLMs)为基于推理的移动导航提供了坚实的基础,凭借大规模预训练带来丰富的上下文理解能力和强大的泛化能力。大多数现有导航框架依赖模仿学习,因此需要大量带标注的轨迹数据,限制了其可扩展性和鲁棒性。本文提出一种参数高效的方法,采用命令式学习范式对预训练的VLM进行微调以实现自主导航。通过优化可微的几何代价场而非带标注的轨迹,我们的模型仅从双目深度观测中学习生成无碰撞路径。我们提出了一个统一的端到端导航流水线,用于自然语言驱动的机器人控制。该系统利用共享的VLM主干网络和任务特定的低秩适应模块,有效弥合了从语义目标选择到底层轨迹规划之间的鸿沟。我们的方法在未见过的环境中取得了具有竞争力的路径加权成功率,且仅更新了模型总参数中不到1%的部分。真实世界的定性实验验证了从仿真到现实的泛化能力,以及无需在真实数据上微调即可实现稳定的路径规划。这些结果表明,这是一种在移动机器人上部署基于VLM的智能体的实用方法,使其能够进行高层语义导航,而无需依赖大规模带标注轨迹数据这一高昂要求。
cs.RO / 50 / 2609.22926

Embodied Snap: Octopus-Inspired Distributed Reach-and-Attach with a Speed-Limited Soft Arm

具身化弹射:基于速度受限软体臂的章鱼启发式分布式触及与吸附
Hou, Linxin, Qin, Zhihang, Zou, Heyang, Wu, Qirui, Wang, Peiyi, Nazeer, Muhammad Sunny, Guo, Yongxin, Laschi, Cecilia
Abstract
Reach-and-attach of soft robotic arms with passive suction requires accurate targeting and sufficient contact speed, yet geared actuators can impose a speed limit that improved trajectory tracking alone cannot overcome. This paper proposes an embodied snap controller that separates slow servo-driven preloading from rapid elastic release, enabling a compliant arm to move beyond its direct tendon-driven speed limit. Octopus biology motivates the controller's section-wise organizational prior, rather than reproduction of the octopus nervous system. A learned policy shared across three sections selects preloads, aim, tendon slack, and release timing, determining where, how, and when to load and release the body. The policy is optimized offline using a hardware-validated recurrent model within experimentally supported bounds. Across five optimization seeds and 400 unseen simulated targets, attachment success is $(73\pm4)\%$ at a $5\text{ cm}$ lateral tolerance, and the shared policy reaches the matched centralized controller's mean final reward after a median $17\%$ of the common evaluation budget. Hardware characterization achieves tip speeds of 1.56-1.64 m/s, at least $108\%$ above direct tendon-driven release. In 18 open-loop hardware trials across six placements, 17 exceed the 1 m/s snap threshold and nine retrieve the object, with successful retrieval at five placements. These results demonstrate a practical division of responsibility in the control problem: learned control prepares the body, and passive body mechanics execute the rapid movement needed for dynamic reach-and-attach.
Chinese Translation
带有被动吸盘的软体机械臂的触及与吸附任务要求精确的目标定位和足够的接触速度,然而齿轮传动执行器所施加的速度限制无法仅通过改进轨迹跟踪来克服。本文提出一种具身化弹射控制器,将慢速的伺服驱动预加载与快速的弹性释放相分离,使柔性臂的运动速度超越其直接的腱驱动速度极限。章鱼的生物学特性启发了该控制器的分段式组织先验,而非对章鱼神经系统的复制。一个在三个节段之间共享的学习策略负责选择预加载量、瞄准、腱松弛量以及释放时机,从而决定在何处、以何种方式以及何时对本体进行加载和释放。该策略在实验支持的边界约束下,利用经过硬件验证的循环模型进行离线优化。在五个优化种子和400个未见过的仿真目标上,在5 cm的横向容差下吸附成功率为 (73±4)%,且该共享策略在仅用共同评估预算的中位数17%时便达到了与集中式控制器相当的平均最终回报。硬件特性测试实现了1.56–1.64 m/s的末端速度,比直接的腱驱动释放至少高出108%。在六个放置位置共18次开环硬件实验中,17次超过1 m/s的弹射速度阈值,九次成功取回物体,其中五个位置实现了成功取回。这些结果表明控制问题中存在一种切实可行的职责划分:学习到的控制负责准备本体姿态,而被动本体力学则执行动态触及与吸附所需的快速运动。
cs.RO / 51 / 2609.22954

MIRA-PRM: Mission-Informed Reusable Roadmap Planning for Mobile Gas-Sensing Inspection

MIRA-PRM:面向移动气体传感检测的任务感知可复用路线图规划
Wang, Gongsen, Wang, Siyuan, Wang, Xinyuan, Wen, Yuhan, Yang, Feng, Tian, Zhen, Lee, Hoi Leong
Abstract
Mobile gas-sensing inspection requires a mobile platform to reliably reach ordered sampling poses. We present MIRA-PRM, a mission-informed probabilistic roadmap that combines geometry-, mission-flow-, and inspection-conditioned sampling with locally adaptive connectivity and evidence-gated refinement. Eight development campaigns comprising 142,448 planner runs evaluated mission scaling, repeated-query adaptation, finite node budgets, target distributions, local map repair, ablation, cross-family comparison, and parameter sensitivity. MIRA-PRM maintained 100% seed-level success as ordered targets increased from two to eight, and achieved 97.67% seed-level success under a 40-node ceiling, compared with 67.17% for PRM and 86.50% for PRM*. Local repair reduced event time by 79.9-81.9% and node additions by 90.8-94.3% relative to rebuilding. On a separate frozen holdout of 12 maps, 48 ordered missions, and three seeds per map-mission unit, MIRA-PRM produced 96/144 validated successes, versus 82/144 for PRM, 89/144 for PRM*, and 70/144 for LE-HG-PRM. On paired common-success runs, MIRA-PRM reduced wall time by 16.89% relative to PRM* and 81.01% relative to LE-HG-PRM, but was 52.23% slower than PRM; paths were 0.54-1.86% longer. All validated paths passed independent full-segment checks and clearance non-inferiority. The results support a reliability-oriented trade-off in structured simulations and motivate external-map and physical mobile-sensing validation.
Chinese Translation
移动气体传感检测要求移动平台可靠地到达有序的采样位姿。我们提出了MIRA-PRM,一种任务感知的概率路线图(probabilistic roadmap),它将基于几何、任务流和检测条件的采样与局部自适应连通性及证据门控的精化相结合。通过包含142,448次规划器运行的八个开发实验,评估了任务规模扩展、重复查询自适应、有限节点预算、目标分布、局部地图修复、消融实验、跨算法族比较以及参数敏感性。当有序目标从两个增加到八个时,MIRA-PRM保持了100%的种子级成功率;在40个节点上限条件下,达到了97.67%的种子级成功率,而PRM为67.17%,PRM*为86.50%。与重建地图相比,局部修复将事件时间减少了79.9%至81.9%,节点增加量减少了90.8%至94.3%。在一个单独的冻结保留集(包含12张地图、48个有序任务、每个地图-任务单元三个随机种子)上,MIRA-PRM取得了96/144的验证成功,而PRM为82/144,PRM*为89/144,LE-HG-PRM为70/144。在配对的共同成功运行中,MIRA-PRM相对PRM*将运行时间(wall time)减少了16.89%,相对LE-HG-PRM减少了81.01%,但比PRM慢52.23%;路径长度长0.54%至1.86%。所有验证通过的路径均通过了独立的分段完整性检查与间隙非劣性检验。这些结果支持在结构化仿真中以可靠性为导向的权衡,并为外部地图验证及物理移动传感验证提供了依据。
cs.RO / 52 / 2609.22963

Prescribed-Time Contracting-Boundary Control of a Tendon-Driven Flexible Arm

腱驱动柔性臂的预设时间收缩边界控制
Lu, Yi, Tang, Chao, Han, Zhiji, Wang, Hongdu
Abstract
This study develops a prescribed-time performance-shaping control method for curvature tracking of a single-segment flexible arm actuated by three antagonistic tendon pairs. A Cartesian curvature representation is introduced to avoid the undefined bending direction at the straight configuration and to establish an explicit six-tendon kinematic mapping. A cubic performance boundary contracts smoothly from an initially admissible error bound to a nonzero terminal accuracy bound within a prescribed time. Based on this boundary, a dual transformation combining static symmetric error scaling and time-varying behavior shaping maps the tracking error into a fixed unit box. The resulting controller guarantees boundary invariance, prescribed-time entry into the terminal accuracy region, and subsequent asymptotic convergence. Numerical evaluations with Python and OpenCR--MuJoCo, together with a supervised reduced-order experiment on a two-section, four-channel platform, provide complementary validation. Across six experimental trials, no violation of the prescribed boundary is observed, and the proposed controller reduces the mean terminal curvature RMSE by 32.5% relative to a matched baseline, with comparable terminal-band entry times. These results support the feasibility of the proposed approach in the reduced-order experimental setting.
Chinese Translation
本研究针对由三对拮抗腱驱动的单段柔性臂的曲率跟踪问题,提出了一种预设时间性能整形控制方法。引入笛卡尔曲率表示以避免直线构型下弯曲方向无定义的问题,并建立显式的六腱运动学映射。一个三次性能边界在预设时间内从初始允许误差界平滑收缩至非零终端精度界。基于该边界,结合静态对称误差缩放与时变行为整形的双重变换将跟踪误差映射至固定单位盒中。所得控制器保证边界不变性、在预设时间内进入终端精度区域及其后的渐近收敛。基于Python与OpenCR–MuJoCo的数值评估,以及在两段四通道平台上开展的有监督降阶实验,提供了互补性验证。在六次实验中,未观察到对预设边界的违背,且所提控制器相对于匹配的基线方法将平均终端曲率RMSE降低32.5%,同时终端频带进入时间相当。这些结果支持了所提方法在降阶实验设定下的可行性。
cs.RO / 53 / 2609.22966

Transferring the Intelligence of VLMs to Robotic Control

将视觉语言模型的智能迁移至机器人控制
Guo, Meng-Hao, Mo, Zhe-Han, Wang, Jia-Jun, Zhang, Yi, Wang, Kejin, Deng, Yi-Xuan, Zhang, Jia-Peng, Rao, Yongming, Hu, Shi-Min
Abstract
Humans can seamlessly adapt to both physical and digital worlds, suggesting that while a digital-to-real gap exists in embodiment, environment and task, human intelligence itself may transfer across this gap. This naturally raises a fundamental question: can the intelligence of vision-language models (VLMs) similarly generalize from the digital world to the physical world for robotic control? We investigate this question through RoboDawn, a human-intuitive interface that exposes robotic control to an agentic VLM through a compact set of discrete translation, rotation, and gripper commands. Using this interface, the VLM controls a robot in a closed loop: it observes the current visual state, reasons about the next action, executes it, and adapts subsequent decisions to the resulting state. Furthermore, we introduce an in-context learning (ICL) scheme that uses a few demonstrations to ground the VLM in both interface usage and task-solving strategies. Experiments on RoboTwin 2.0 C2R and RoboDojo demonstrate that RoboDawn achieves strong performance without task-specific robot training. In the zero-shot setting, RoboDawn outperforms several strong policies trained on benchmarkspecific robot data, while a single in-context demonstration further yields substantial performance gains and establishes state-of-the-art (SOTA) results. On RoboTwin 2.0 C2R, the success rate increases from 53.2% zero-shot to 73.6% one-shot, exceeding the solid baseline {\pi}0.5 (46.0%). Similar gains are observed on RoboDojo, where success rate improves from 35.67% zero-shot to 47.17% one-shot. The same framework also transfers to real-world robots, performing block-in-basket and block stacking on Franka.
Chinese Translation
人类可以无缝地适应物理世界和数字世界,这表明尽管在具身、环境和任务上存在数字到现实的差距,但人类智能本身或许能够跨越这一差距进行迁移。这自然引出一个根本性问题:视觉语言模型(VLM)的智能能否类似地从数字世界泛化到物理世界,以实现机器人控制?我们通过 RoboDawn 来研究这一问题。RoboDawn 是一种符合人类直觉的接口,它通过一组紧凑的离散平移、旋转和夹爪指令,将机器人控制暴露给具有智能体能力的 VLM。借助该接口,VLM 以闭环方式控制机器人:它观察当前的视觉状态,推理下一步动作,执行该动作,并根据执行结果调整后续决策。此外,我们引入了一种上下文学习(ICL)方案,利用少量演示使 VLM 同时掌握接口的使用方法和任务解决策略。在 RoboTwin 2.0 C2R 和 RoboDojo 上的实验表明,RoboDawn 无需任务特定的机器人训练即可取得出色的性能。在零样本设置下,RoboDawn 的表现优于多个在基准特定机器人数据上训练的强策略;而仅一条上下文演示即可带来显著的性能提升,并达到了最先进(SOTA)水平。在 RoboTwin 2.0 C2R 上,成功率从零样本的 53.2% 提升至单样本的 73.6%,超过了稳健的基线方法 {\pi}0.5(46.0%)。在 RoboDojo 上也观察到类似的提升,成功率从零样本的 35.67% 提高到单样本的 47.17%。同一框架还可迁移到真实世界的机器人上,在 Franka 机械臂上完成了将方块放入篮筐和方块堆叠任务。
cs.RO / 54 / 2609.22973

An Empirical Study and Open Testbed for Federated Fine-Tuning of Vision-Language-Action Models

视觉-语言-动作模型联邦微调的实证研究与开放测试平台
Duan, Zhekai, Xie, Kevin Ziyang, Tan, Xinyu, Geng, Shikai, Zhou, Chengxu, Kompella, Ramana, Liu, Gaowen, Lu, Chris Xiaoxuan
Abstract
Adapting a pretrained Vision-Language-Action (VLA) model to a new robot, environment, or task requires demonstrations that are collected locally and often discarded. Federated learning is a promising approach to exploiting such distributed demonstrations by learning a shared policy. However, whether it can adapt large pretrained VLAs remains an open question, and a lack of reproducible benchmarks for pretrained VLAs and reusable training frameworks makes existing results difficult to compare. In this paper, we conduct a systematic study of federated fine-tuning of three modern pretrained VLA policies on the 40 simulated tasks of the LIBERO manipulation benchmark, and on six real-world tasks in two real-robot experiments, with demonstrations collected across two and three sites, respectively. Our study analyzes the key choices in this setting, spanning multiple federated parameter scopes, three aggregation algorithms, and evaluation under distribution shift. Based on the study, we derive a series of lessons, including the dominance of the federated scope over the choice of aggregation algorithm and the difficulty of matching centralized fine-tuning on physical robots, where cross-site heterogeneity is stronger than simulation captures. We also highlight opportunities for federated VLA learning, such as the ability to match centralized fine-tuning on heterogeneous data, to remain at least as robust as centralized fine-tuning under distribution shift, and to personalize, with each client federating part of the policy and keeping the rest local, which helps where the policy's pretraining is weak but leaves no usable global model. We open-source \decentvla{}, the model- and runtime-agnostic testbed behind the study, to facilitate future research and fair comparisons in federated VLA learning.
Chinese Translation
将预训练的视觉-语言-动作(Vision-Language-Action,VLA)模型适配到新的机器人、环境或任务,需要在本地收集且通常随后被丢弃的演示数据。联邦学习(Federated Learning)通过学习一个共享策略,为利用这类分布式演示数据提供了一种有前景的方法。然而,它能否用于适配大型预训练VLA模型仍是一个悬而未决的问题,并且由于缺乏可复现的预训练VLA基准和可复用的训练框架,现有结果难以相互比较。在本文中,我们对三种现代预训练VLA策略的联邦微调进行了系统性研究:实验涵盖LIBERO操作基准的40个仿真任务,以及两个真实机器人实验中的六个真实世界任务,其中演示数据分别收集自两个和三个站点。我们的研究分析了该场景下的关键选择,包括多种联邦参数范围、三种聚合算法,以及分布偏移下的评估。基于该研究,我们总结出一系列经验教训,包括联邦参数范围的重要性胜过聚合算法的选择,以及在物理机器人上难以匹集中式微调的效果——真实场景中的跨站点异质性比仿真所刻画得更为强烈。我们还强调了联邦VLA学习的机遇,例如:能够在异构数据上匹集中式微调的效果;在分布偏移下至少保持与集中式微调同等的鲁棒性;以及实现个性化,即每个客户端将策略的一部分进行联邦化而将其余部分保留在本地——这在策略预训练较弱时有所帮助,但不会留下可用的全局模型。我们开源了该研究所使用的测试平台 \decentvla{},它不依赖于特定模型和运行时环境,以促进联邦VLA学习领域的未来研究和公平比较。
cs.RO / 55 / 2609.22974

A Horizon-slicing Approach to Minimum Obstacle Displacement Planning for Robot Navigation

一种面向机器人导航的最小障碍物位移规划的视界切片方法
Thomas, Antony, Ferro, Giulio, Mastrogiovanni, Fulvio, Robba, Michela, Baglietto, Marco
Abstract
In this paper, we investigate the Minimum Obstacle Displacement Planning problem from a robot motion planning perspective. The problem involves determining a feasible path to a goal location by displacing movable obstacles when no collision-free path initially exists. We show that this problem is computationally challenging and, in particular, NP-hard when obstacles are modeled as polygons in the plane. Besides an exact formulation of the minimum obstacle displacement problem generalizing other problems in the literature, and the associated optimal solution, this paper proposes an approximate solution that is less intensive from a computational standpoint, and differs from the optimal solution by a fraction of the optimal cost, being able to trade-off between path length and amount of obstacle displacements.
Chinese Translation
本文从机器人运动规划的视角研究了最小障碍物位移规划问题。该问题旨在当初始不存在无碰撞路径时,通过移动可移动的障碍物来确定一条到达目标位置的可行路径。我们证明了该问题在计算上具有挑战性,特别是当障碍物被建模为平面上的多边形时,该问题是NP难的。除了给出推广了文献中其他问题的最小障碍物位移问题的精确形式化描述及其相应的最优解之外,本文还提出了一种计算开销较低的近似解法,该解法与最优解的偏差仅为最优代价的一小部分,并能够在路径长度与障碍物位移量之间进行权衡。
cs.RO / 56 / 2609.22983

Connectivity-Aware Exploration of Robotic Grasp Spaces

连接性感知的机器人抓取空间探索
Kazanskii, Maksim A
Abstract
Robotic grasping is typically formulated as the problem of identifying successful actions from a space of candidate grasp poses. However, the organization of successful actions within this space has received less attention. We study the multiscale structure of viable robotic grasps in $SE(3)$ and investigate whether this structure can be exploited for more efficient exploration. Using a large-scale grasp dataset, we show that successful grasp sets exhibit heterogeneous and reproducible connectivity structure across objects. We then introduce a connectivity-aware sampling strategy that incrementally explores the currently observed grasp space by prioritizing potential bridges between components, structural frontiers, boundary extensions, and geometric novelty. In controlled reconstruction experiments, the method recovers the connectivity structure of successful grasp sets substantially more efficiently than random sampling and farthest-point sampling. We further evaluate whether connectivity acquired under hidden grasp viability can improve subsequent grasp discovery, and whether structural experience from previously explored objects can be retrieved and transferred to unseen objects. These results suggest that the spatial organization of viable actions provides information relevant to grasp-space exploration beyond the viability of individual candidate actions. More broadly, they motivate structure-aware exploration as a means of exploiting the geometry of viable action spaces in robotic manipulation.
Chinese Translation
机器人抓取通常被表述为从候选抓取位姿空间中识别成功动作的问题。然而,成功动作在该空间中的组织结构却较少受到关注。我们研究了SE(3)中可行机器人抓取的多尺度结构,并探讨能否利用这一结构实现更高效的探索。基于大规模抓取数据集,我们表明成功抓取集合在不同物体上呈现出异构且可复现的连接性结构。随后,我们提出一种连接性感知的采样策略,通过优先探索潜在桥接、结构边界前沿、边界扩展和几何新颖性,对当前观测到的抓取空间进行增量式探索。在受控重建实验中,该方法恢复成功抓取集合连接性结构的效率显著高于随机采样和最远点采样。我们进一步评估了在抓取可行性隐含条件下习得的连接性能否改进后续抓取发现,以及先前探索物体的结构经验能否被检索并迁移至未见物体。这些结果表明,可行动作的空间组织提供了超越单个候选动作可行性的、与抓取空间探索相关的信息。更广泛地说,这些结果支持以结构感知的探索作为利用机器人操作中可行动作空间几何性质的一种手段。
cs.RO / 57 / 2609.23037

Fast and Robust Temporal Logic Planning via ADMM-based Trajectory Optimization

基于ADMM轨迹优化的快速鲁棒时序逻辑规划
Pries, Lukas, Verhagen, Joris, Arrizabalaga, Jon, Tumova, Jana, Ryll, Markus, Manchester, Zachary
Abstract
We present a fast numerical method for safe continuous-time motion planning under Temporal Logic (TL) specifications. The method generates smooth continuous trajectories that remain collision-free while robustly satisfying temporal and logical task requirements. A central component of our method is the formulation of nonconvex safety and logic constraints as unions of convex sets where associated discrete decisions are encoded in a joint feasibility graph. This graph representation allows Euclidean projection onto the feasible set and proximal robustness maximization to be reformulated as shortest- and widest-path problems, respectively. Building on this structure, we develop a nonconvex splitting method based on the Alternating Direction Method of Multipliers (ADMM), which decouples smooth spatio-temporal trajectory optimization from nonsmooth discrete constraint handling within the optimization. The resulting algorithm exhibits reliable convergence across benchmarks and scales to large-scale motion-planning problems, providing a 4.7x average speedup over the state of the art on discrete and continuous-time logic problems.
Chinese Translation
我们提出了一种在时序逻辑(Temporal Logic, TL)规范约束下进行安全连续时间运动规划的快速数值方法。该方法生成平滑的连续轨迹,在保持无碰撞的同时鲁棒地满足时序与逻辑任务要求。本方法的核心是将非凸的安全约束与逻辑约束表述为凸集的并集,并将相关的离散决策编码在一个联合可行性图中。这种图表示方法使得向可行集的欧几里得投影和近端鲁棒性最大化问题可分别重新表述为最短路径与最宽路径问题。基于该结构,我们开发了一种基于交替方向乘子法(Alternating Direction Method of Multipliers, ADMM)的非凸分裂方法,该方法在优化过程中将平滑的时空轨迹优化与非平滑的离散约束处理解耦。所得到的算法在多个基准测试中表现出可靠的收敛性,并可扩展至大规模运动规划问题,在离散与连续时间逻辑问题上相比现有最先进方法实现了平均4.7倍的加速。
cs.RO / 58 / 2609.23048

Anatomy of a Closed-Loop Collapse: A Causal Case Study of a Compressed VLA Policy

闭环崩溃的解剖:一个压缩VLA策略的因果案例研究
Jia, Fengze
Abstract
Compressed manipulation policies can pass offline evaluation while failing in closed-loop execution; this dissociation is established in prior work and is not our claim. We contribute a causal anatomy of one naturally occurring case. An 8-layer distillation of Octo-Base retains 86% of parameters, passes every offline check we applied (0.996 and 1.000 teacher-ratios on the family's own validation metrics), and collapses in closed loop: 0/72 vs. the teacher's 40/72 on a simulated WidowX pick-and-place task. The collapse is structured, not diffuse: early task stages degrade gradually (the student moves the object at 90% of the teacher's rate and grasps at 55%), while transport-to-target fails categorically, at 0% in every training variant. Paired action-trace forensics isolate the signature: a negative, late-heavy $z$ residual, roughly 10x its post-repair magnitude, and persistent across the base distillation and both continuation branches. Four standard therapies fail under matched controls: continued training and in-domain offline data leave success at zero, even though the latter measurably improves marginal action statistics; command-level compensation recovers nothing at any offset, although the same perturbations degrade healthy policies; clamping the symptom in the command channel preserves grasping, yet success stays at floor. A minimal-pair intervention that substitutes half of the training stream with deployment-distribution teacher rollouts, with every other setting held fixed, restores parity with the teacher (18/36 vs. 17/36 held-out), eliminates that signature, and recovers a teacher-like perturbation-response profile. We claim existence, not universality. Operationally, offline gates, including a family's own validation metrics, are insufficient acceptance tests for compressed policies; a few dozen closed-loop trials sufficed to find what they missed.
Chinese Translation
压缩后的操作策略可能通过离线评估,却在闭环执行中失败;这种脱节已在先前的工作中得到证实,并非本文的主张。我们的贡献是对一个自然发生案例的因果解剖。对Octo-Base的8层蒸馏模型保留了86%的参数,通过了我们应用的所有离线检查(在该模型家族自身的验证指标上教师比率分别为0.996和1.000),但在闭环中崩溃:在模拟的WidowX抓取放置任务上,成功率为0/72,而教师模型为40/72。这种崩溃是结构化的,而非弥散的:早期任务阶段逐渐退化(学生模型移动物体的速率达到教师模型的90%,抓取达到55%),而向目标运输则在所有训练变体中均彻底失败,为0%。配对的动作轨迹取证分离出了特征签名:一个负向、后期加重的$z$残差,约为修复后量级的10倍,且在基础蒸馏和两个继续训练分支中持续存在。四种标准疗法在匹配对照下均告失败:继续训练和域内离线数据使成功率保持为零,尽管后者确实可测量地改善了边缘动作统计;命令级补偿在任何偏移量下均无恢复,尽管相同的扰动会使健康策略退化;在命令通道中钳制该症状虽保住了抓取能力,但成功率仍停留在底部。一项最小配对干预——用部署分布的教师rollout替换一半训练流,其余所有设置保持不变——恢复了与教师模型的对等水平(保留集上18/36对17/36),消除了该特征签名,并恢复了类似教师模型的扰动响应曲线。我们主张的是存在性,而非普遍性。在操作层面,离线门控(包括模型家族自身的验证指标)不足以作为压缩策略的验收测试;只需几十次闭环试验就足以发现它们所遗漏的问题。
cs.RO / 59 / 2609.23060

On the Control of Mobile Ad-Hoc Agent Deployments in Partially Observed Space

部分可观测空间中移动自组织智能体部署的控制研究
Meriaux, Edwin, Langevin, Louis-Roy, Wen, Shuo, Ndiaye, Ndiamé, Dudek, Gregory, Loría, Antonio
Abstract
We study the online deployment of mobile ad hoc networks in unknown orthogonal environments, formalized as the Partially Observable Cooperative Guard Art Gallery Problem. We give a full proof that CADENCE algorithms achieve full coverage while maintaining a connected visibility graph using at most $n/2 + h - 2$ agents in orthogonal worlds with $n$ corners and $h$ holes. We further evaluate deployment-order heuristics that reduce agent count and deployment time in practice.
Chinese Translation
我们研究了移动自组织网络在未知正交环境中的在线部署问题,并将其形式化为部分可观测协作守卫艺术画廊问题(Partially Observable Cooperative Guard Art Gallery Problem)。我们给出了完整的证明:对于拥有 $n$ 个角点和 $h$ 个洞的正交世界,CADENCE 算法能够在保持可见图连通的前提下,使用至多 $n/2 + h - 2$ 个智能体实现完全覆盖。我们进一步评估了在实际中能够减少智能体数量和部署时间的部署顺序启发式方法。
cs.RO / 60 / 2609.23100

Splat-CBF: Safe Next-Best-View Control in 3D Gaussian-Splat Maps

Splat-CBF:基于3D高斯溅射地图的安全下一最优视角控制
Khass, Amirhossein Mollaei, Cosse, Athanasios, Motee, Nader
Abstract
Where to look and how to move? A robot navigating an unmapped environment must do both at once, and the two goals pull against each other. The regions most worth observing are the ones the map knows least about, and those are exactly where the robot cannot trust its collision margins. We resolve this tension by introducing Splat-CBF, an active perception control barrier function that steers the camera toward the next best view while collision avoidance is enforced as a hard constraint. Safety is enforced by a risk-aware control barrier function that turns the Average Value-at-Risk of the Gaussian field into a single smooth hard constraint. Perception is enforced by a second barrier that rewards camera orientations with high expected Fisher information gain near the robot's planned path. The two meet in a quadratic program where safety is hard and perception is soft, with a slack penalty that adapts to how often perception has already been relaxed and how close the robot is to an uncertain region. We verify the method in indoor simulations, a Isaac Kinova manipulator and in experiments on an Ackermann-drive robot. Our results assert that robot navigates faster, gathers more information, and runs faster online than safety-only and perception-only baselines, giving up informative motion only when safety requires it.
Chinese Translation
往哪里看,以及如何移动?在未建图环境中导航的机器人必须同时完成这两项任务,而这两个目标却相互冲突。最值得观测的区域恰恰是地图了解最少的区域,而这些区域正是机器人无法信任其碰撞裕度的地方。我们通过提出 Splat-CBF 来化解这一矛盾——这是一种主动感知控制屏障函数(control barrier function),在将碰撞规避作为硬约束加以强制执行的同时,引导相机朝向下一最优视角。安全性由一个风险感知的控制屏障函数保证,该函数将高斯场的平均风险价值(Average Value-at-Risk)转化为单一的光滑硬约束。感知则由第二个屏障函数实现,该函数对机器人在规划路径附近具有高期望Fisher信息增益的相机朝向给予奖励。二者在一个二次规划(quadratic program)中结合,其中安全为硬约束、感知为软约束,并引入一个松弛惩罚项,其大小根据感知已被放松的频率以及机器人距不确定区域的接近程度自适应调整。我们在室内仿真、Isaac Kinova 机械臂以及阿克曼(Ackermann)驱动移动机器人实验中验证了该方法。结果表明,相较于仅考虑安全性和仅考虑感知的基线方法,本方法使机器人导航更快、收集的信息更多、在线运行速度更快,并且仅在安全需要时才放弃信息性运动。
cs.RO / 61 / 2609.23103

DiagGen: Agentic Generation of Deformable Assets with Sim-based Diagnostics for Robotic Simulation

DiagGen:基于仿真诊断的机器人仿真可变形资产智能体生成方法
Chen, Guanxiong, Qu, Yiduo, Xia, Qianjun, Jing, Pengyu, Cheng, Yixian, Ma, Bole, Yang, Pengzhi, Zhou, Bingyang, Li, Ziming, Suri, Shashwat, Sun, Gongbo, Liu, Chao, Chen, Peter Yichen, Zeng, Ziqiu, Shi, Fan
Abstract
While simulation-ready deformable assets are essential for in-silico robotic manipulation tasks, existing generation frameworks typically assess physical plausibility after generation, leaving an object's simulated response unused as feedback for repairing upstream errors. We present DiagGen, an agentic framework that turns a single in-the-wild image into a simulation-ready deformable asset through a generate--simulate--diagnose--refine loop. DiagGen constructs part-aware geometry and material parameters, then uses a VLM (vision-language model)-based agent to select semantically informative regions, probe them in a physics simulator, observe material responses, and route evidence-backed repair cues to the responsible generation stage. Experiments on 40 assets show that diagnostics provides useful repair cues and can moderately improve the quality of generated deformable assets. Finally, we show that unlike assets generated from visual foundation models which may not be simulatable, DiagGen-generated deformables can be directly dropped into a high-fidelity physical simulator for the planning and simulation of contact-rich pick-and-place tasks. The project's website is https://diaggen.github.io/.
Chinese Translation
尽管可用于仿真的可变形资产对于计算机内机器人操作任务至关重要,现有的生成框架通常在生成之后才评估物理合理性,未能将物体的仿真响应作为反馈来修复上游错误。我们提出了DiagGen,这是一个智能体框架,能够通过“生成—仿真—诊断—优化”的循环,将单张真实世界图像转化为可用于仿真的可变形资产。DiagGen构建具有部件感知的几何与材料参数,随后利用基于视觉语言模型(VLM)的智能体选择语义上有意义的区域,在物理仿真器中对其进行探测,观察材料响应,并将有证据支持的修复线索传递给相应的生成阶段。在40个资产上的实验表明,诊断能够提供有用的修复线索,并能适度提升生成可变形资产的质量。最后,我们展示了与那些可能无法进行仿真的视觉基础模型生成的资产不同,DiagGen生成的可变形资产可以直接放入高保真物理仿真器中,用于接触密集型抓取与放置任务的规划与仿真。项目网站为 https://diaggen.github.io/。
cs.RO / 62 / 2609.23113

Search, Ground, Plan: Functional Sufficiency for Task and Motion Planning under Incomplete Scene Knowledge

搜索、落地、规划:面向不完全场景知识下任务与运动规划的功能充分性
Vijayakumar, Narendhiran, Singhal, Nav, Varma, Girish, Thomas, Antony
Abstract
Foundation models (FMs) have expanded task and motion planning (TAMP) to manipulation problems specified through language and visual observations. However, incomplete scene knowledge leaves a critical gap between understanding what the task requires and knowing whether the physical scene can actually realize it. We introduce GRAB-TAMP, an FM-based TAMP framework that searches for scene entities required for task completion, grounds functional roles to valid physical objects, and plans only after a complete joint assignment establishes functional sufficiency. We represent the task through functional roles, relations, and assignment constraints, and incrementally inspect the scene while requirements remain unresolved, verifying candidate objects through semantic, geometric, and relational checks. We evaluate GRAB-TAMP across 32 scene variants spanning Kitchen, Living Room, and Workshop domains. Across 200 feasible trials, our approach achieves 54.0% end-to-end success with 67.3% plan goal coverage. Compared with three FM-based TAMP frameworks under the same execution setting, GRAB-TAMP improves end-to-end success by 25.7 percentage points over the mean baseline. Implementation and evaluation code: https://github.com/Narendhiranv04/GRAB-TAMP
Chinese Translation
基础模型(Foundation Models, FMs)已将任务与运动规划(Task and Motion Planning, TAMP)扩展至通过语言和视觉观测指定的操作问题。然而,场景知识的不完全性在“理解任务需求”与“知晓物理场景能否真正实现该需求”之间留下了关键鸿沟。我们提出了GRAB-TAMP,一个基于基础模型的TAMP框架,它搜索完成任务所需的场景实体,将功能角色落地到有效的物理对象上,并在完整的联合分配建立功能充分性之后才进行规划。我们通过功能角色、关系和分配约束来表示任务,并在需求尚未满足时增量式地检查场景,通过语义、几何和关系检验来验证候选对象。我们在涵盖厨房(Kitchen)、客厅(Living Room)和工作间(Workshop)三个领域的32个场景变体上对GRAB-TAMP进行了评估。在200次可行试验中,我们的方法取得了54.0%的端到端成功率和67.3%的规划目标覆盖率。与在相同执行设置下的三种基于基础模型的TAMP框架相比,GRAB-TAMP的端到端成功率较基线平均值提升了25.7个百分点。实现与评估代码:https://github.com/Narendhiranv04/GRAB-TAMP
cs.RO / 63 / 2609.23118

Verti-WM: A Physics-Aided Exteroceptive World Model for Off-Road Reinforcement Learning

Verti-WM:一种面向越野强化学习的物理辅助外感受世界模型
Pan, Chenhui, Xu, Tong, Xiao, Xuesu
Abstract
Reinforcement learning for off-road navigation requires extensive vehicle-terrain interaction data, which are costly to collect in high-fidelity simulation. World models offer a promising alternative by replacing simulator roll-outs during policy optimization. However, an off-road world model must condition state transitions on exteroceptive terrain information, which proprioception alone does not provide. This challenge is further amplified by the need to model both rigid and deformable terrain, where data-driven and physics-based approaches offer complementary strengths. We propose Verti-WM, a physics-aided exteroceptive world model that recurrently fuses a frozen Transformer for rigid terrain and a neuro-symbolic terramechanics model for deformable terrain. Elevation and semantic observations queried from a supplied map at each predicted pose condition fusion, enabling six-degree-of-freedom rollouts for policy optimization without further simulator access. Verti-WM reduces prediction error by 34.6% and 21.7% over data-driven and physics-based baselines, respectively. Policies trained entirely within Verti-WM achieve comparable task success rates while reducing computation time by 23.6X relative to direct training in the high-fidelity simulator. We further validate Verti-WM using real-world data, enabling policy optimization within learned real-world kinodynamics and achieving a 80% success rate on the Verti-4-Wheeler platform, compared with 40% for direct sim-to-real transfer.
Chinese Translation
越野导航的强化学习需要大量的车辆-地形交互数据,而在高保真仿真中采集这些数据成本高昂。世界模型通过在策略优化过程中替代仿真器 rollout,提供了一种有前景的替代方案。然而,越野世界模型必须使状态转移依赖于外感受式地形信息,而仅靠本体感知无法提供这些信息。此外,该挑战还因需要同时建模刚性地形和可变形地形而进一步加剧,其中数据驱动方法与基于物理的方法各具互补优势。我们提出了 Verti-WM,这是一种物理辅助的外感受世界模型,它循环地融合一个用于刚性地形的冻结 Transformer 和一个用于可变形地形的神经符号(neuro-symbolic)地面力学模型。在每个预测位姿处从给定地图中查询的高程和语义观测为融合提供条件,从而无需进一步访问仿真器即可进行六自由度 rollout 以优化策略。相比数据驱动基线和物理基线,Verti-WM 分别将预测误差降低了 34.6% 和 21.7%。完全在 Verti-WM 中训练的策略实现了相当的任务成功率,同时相比在高保真仿真器中直接训练,计算时间减少了 23.6 倍。我们进一步使用真实世界数据验证了 Verti-WM,实现了在所学习的真实世界运动动力学内进行策略优化,并在 Verti-4-Wheeler 平台上取得了 80% 的成功率,而直接进行 sim-to-real 迁移的成功率仅为 40%。
cs.RO / 64 / 2609.23131

Selective Commitment for Language-Guided Object Retrieval under Partial Observability

部分可观测性下语言引导目标抓取的选择性承诺机制
Koh, Wonhee, Dinesh, Sushil Samuel, Ko, Hansol, Park, Shinkyu, Lee, Eungjoo
Abstract
Language-guided object retrieval under partial observability requires deciding whether to gather more evidence, interact with the scene, grasp a candidate, or abstain. We present a closed-loop framework that coordinates these decisions for retrieving a target specified in relation to a reference container. The framework maintains a persistent joint belief over target identity, container relation, and presence through tracked-object, unobserved-target, and target-absent hypotheses. View-conditioned categorical VLM observations update this belief; conformal grasp eligibility and robot feasibility govern commitment, while finite-horizon belief-space planning selects information-gathering actions. Across five different scenarios, our proposed method succeeds in 19/25 simulation episodes versus 12/25 for the best-performing task-adapted baseline and is the only evaluated policy to achieve at least one success in each scenario. Ablations show that cross-view memory improves success under partial occlusion, while the full system does not consistently outperform simplified variants. Real-robot trials demonstrate closed-loop re-observation and autonomous recovery from injected grasp failures, while injected viewpoint failures end in false defer. Experimental results demonstrate the feasibility of coordinating evidence gathering and selective grasp commitment within a unified framework for retrieval under partial observability.
Chinese Translation
部分可观测性下的语言引导目标抓取需要决策是否收集更多证据、与场景交互、抓取候选目标或放弃操作。我们提出一个闭环框架,用于协调检索与参考容器相关的指定目标时的这些决策。该框架通过被追踪目标、未观测目标和目标缺失三类假设,维护关于目标身份、容器关系和目标存在性的持久联合信念。视角条件化的类别型VLM观测用于更新该信念;保形抓取可行性与机器人可执行性决定承诺时机,而有限时域的信念空间规划用于选择信息收集动作。在五个不同场景中,我们提出的方法在25次仿真回合中成功19次,而表现最好的任务适配基线仅成功12次;且该方法是在每个场景中均至少取得一次成功的唯一被评估策略。消融实验表明,跨视角记忆可提升部分遮挡下的成功率,而完整系统并不总是优于简化变体。真实机器人实验展示了闭环重观测以及从注入的抓取失败中的自主恢复,而注入的视角失败则以错误推迟告终。实验结果证明了在统一框架内协调证据收集与选择性抓取承诺以实现部分可观测性下目标检索的可行性。
cs.RO / 65 / 2609.23132

Tip Manipulation in Soft Everting Robots via Wall Retraction and Deployable Fingers

基于管壁回缩与可展开手指的软体倒置生长机器人末端操作
Perez, Nelson Badillo, Pagliarani, Niccolo, Cianchetti, Matteo, Howe, Robert D.
Abstract
Soft everting robots can traverse long, confined paths by continuously growing, yet active interaction remains limited to a single tool fixed at or near the tip, unable to be repositioned on demand and difficult to reconcile with the robot's soft body. We introduce a tip-manipulation and multi-tool deployment strategy for soft everting robots based on wall retraction, implemented with a base roller assembly that independently meters membrane flow in the outer wall while a tail spool regulates growth in the internal tail section. Coordinated wall and tail actuation decouples robot length from membrane-material position, enabling membrane-mounted devices to be transported, exposed, and repositioned at selected locations near the distal tip. We pair this capability with ultralight pleated inflatable fingers integrated into the membrane, fabricated from TPU-coated nylon with an internal airtight bladder. The fingers achieve large bending at low pressures (approximately 100 degrees at 50 kPa in high-pleat designs) and generate blocking forces up to 1.9 N, while remaining limp during transport. The system demonstrates adaptive grasping across diverse household objects (21 to 550 g; 16 to 200 mm), three-dimensional object manipulation and stacking, environmentally braced extension, distal camera panning for confined-space inspection, and controlled sequential payload delivery. These results enable embodied and reversible tip manipulation for soft growing robots in cluttered and tortuous environments.
Chinese Translation
软体倒置生长机器人可通过持续生长穿越狭长受限路径,然而其主动交互能力仍局限于固定在末端或末端附近的单一工具,无法按需重新定位,且难以与机器人的软体结构相协调。我们提出了一种基于管壁回缩的软体倒置生长机器人末端操作与多工具部署策略,其实现方式为:基座滚轮组件独立控制外壁膜的流动速率,同时尾部线轴调节内部尾段的生长。管壁与尾部的协调驱动将机器人长度与膜材料位置解耦,使安装在膜上的装置能够被输送、暴露并重新定位到远端末端的选定位置附近。我们将该能力与集成于膜中的超轻褶皱充气手指相结合,该手指由TPU涂层尼龙制成并带有内部气密囊。手指在低压下即可实现大角度弯曲(高褶皱设计在50 kPa下弯曲约100度),并可产生高达1.9 N的阻塞力,同时在输送过程中保持柔软。该系统展示了对多种家居物品的自适应抓取(21至550克;16至200毫米)、三维物体操作与堆叠、环境支撑延伸、用于受限空间检查的远端相机平移,以及受控的顺序载荷投放。这些成果为软体生长机器人在杂乱且曲折的环境中的具身化、可逆末端操作提供了可能。
cs.RO / 66 / 2609.23133

AquaCap: A Training-Free Underwater Embodied Agent with Code-as-Policy

AquaCap:一种基于“代码即策略”的无训练水下具身智能体
Li, Xiaoshi, Xu, Yule, Kong, Chunghiu, Zhou, Yizhou, Liu, Yang, Yang, Hao, Huang, Zihao, Shan, Yunxiao
Abstract
Recent advances in vision-language-action models have stimulated growing interest in underwater embodied intelligence. However, their reliance on large-scale interaction data limits their applicability underwater, where data collection is costly and scarce. To address this challenge, we present AquaCap, a training-free Code-as-Policy framework for autonomous underwater navigation and manipulation. AquaCap employs a dual-layer agent that translates task instructions and environmental observations into condition-aware plans and executable control programs. Structured perception then provides the agent with semantic, geometric, and reliability-aware observations under degraded underwater conditions. A failure-aware memory diagnoses unsuccessful actions and supports closed-loop replanning and code revision. This design enables online adaptation without task-specific training or parameter updates. AquaCap achieves a 66.43% success rate in simulation. Real-world experiments further demonstrate autonomous grasping and object transport with an ROV, including the manipulation of targets displaced by hydrodynamic disturbances.
Chinese Translation
视觉-语言-动作模型的最新进展激发了人们对水下具身智能日益增长的兴趣。然而,此类模型对大规模交互数据的依赖限制了其在水下场景的适用性,因为水下数据采集成本高昂且数据稀缺。为应对这一挑战,我们提出了AquaCap,这是一个无需训练的“代码即策略”框架,用于自主水下导航与操作。AquaCap采用双层智能体,将任务指令和环境观测转化为条件感知的计划以及可执行的控制程序。随后,结构化感知模块在水下退化条件下为智能体提供语义、几何和可靠性感知的观测信息。具备失败感知能力的记忆机制可诊断不成功的动作,并支持闭环重规划与代码修正。这一设计使系统能够在线适应,而无需任务特定的训练或参数更新。AquaCap在仿真中达到了66.43%的成功率。真实世界实验进一步验证了利用ROV(遥控潜水器)实现自主抓取和物体搬运的能力,包括对受水动力扰动而位移的目标物体的操作。
cs.RO / 67 / 2609.23144

Probabilistic Scene Graphs: Hierarchical Representation and Real-time System

概率场景图:分层表示与实时系统
Ali, Waqas, Antonazzi, Michele, Homberger, Timon, Nguyen, Thien-Minh, Schmid, Lukas Rosenberger, Jensfelt, Patric, Cai, Yixi
Abstract
3D scene graphs provide semantically rich and hierarchical representations for robot perception. However, existing systems do not maintain uncertainty as an explicit belief or propagate it through the operations that construct and refine the graph. We introduce Probabilistic Scene Graph (PSG), a generalization of the conventional scene graph that represents a posterior over possible graphs, factorized into a discrete graph structure of entities, relations, and semantic attributes, and continuous states that ground them spatially, with uncertainty maintained over both components. Geometry is carried directly by the nodes rather than selected from a separately constructed metric map, so a metric map, where needed, follows from the graph rather than preceding it. We instantiate PSG's probabilistic spatial grounding with hierarchical graphs of Gaussians (HGG): each object primitive is represented by a full-covariance Gaussian under a Normal-Inverse-Wishart belief, and the same parametrization applied recursively within a node yields a geometry graph that resolves its surface at finer resolution. We then build a mapping pipeline that preserves these beliefs throughout graph construction and refinement: a purely graph-based coarse-to-fine alignment registers observations by comparing node beliefs, while a nested Expectation-Maximization and factor-graph optimization jointly refines poses, object parameters, and internal geometry. Across six datasets spanning indoor RGB-D, outdoor LiDAR, and cross-modality deployment, HGG operates at sensor rate with near-constant memory and achieves state-of-the-art object accuracy and zero-shot graph alignment.
Chinese Translation
3D场景图为机器人感知提供了语义丰富且分层的表示。然而,现有系统并未将不确定性作为显式信念加以维护,也未在构建和细化图的操作中对其进行传播。我们提出概率场景图(Probabilistic Scene Graph, PSG),这是对传统场景图的推广,表示可能图上的后验分布,并将其分解为离散的图结构(包括实体、关系和语义属性)与将其在空间中进行定位(grounding)的连续状态,且对两个组成部分均维护不确定性。几何信息直接由节点承载,而非从单独构建的度量地图中选取;因此在需要时,度量地图由图派生而来,而非先于图存在。我们通过分层高斯图(Hierarchical Graphs of Gaussians, HGG)实现PSG的概率空间定位:每个物体基元由正态-逆威沙特(Normal-Inverse-Wishart)信念下的全协方差高斯分布表示,并且对同一参数化在节点内递归应用,得到一个以更细分辨率解析其表面的几何图。随后,我们构建了一个在整个图构建与细化过程中保留这些信念的建图流水线:一种纯基于图的由粗到精对齐方法通过比较节点信念来配准观测,同时嵌套的期望最大化(Expectation-Maximization)与因子图优化联合地细化位姿、物体参数和内部几何。在涵盖室内RGB-D、室外激光雷达(LiDAR)以及跨模态部署的六个数据集上,HGG以传感器频率运行且内存占用近乎恒定,并实现了最先进的物体精度和零样本图对齐性能。
cs.RO / 68 / 2609.23252

Robot World Models Are Not Invariant to How the Actions Are Written

机器人世界模型对动作的书写方式不具备不变性
Karim, Ahmed, Chlon, Leon
Abstract
A robot policy is trained with one of two action parameterizations: absolute joint targets, or deltas relative to the current state. The choice is a live engineering decision in robot learning, and a world model conditioned on actions inherits it silently. We show the inheritance is catastrophic. A latent dynamics model trained on one parameterization and handed the identical commanded trajectory written in the other collapses: retrieval degrades by 2.6-13.4x across three robot datasets and two morphologies, goal-conditioned action selection falls from 53% to 15%, and on PushT the two beliefs about the same future are near-orthogonal (cos = 0.067, worst case -0.377), so the predictor does not degrade gracefully, it answers a different question. This is not a distribution-shift artifact in the usual sense: the two encodings are mutually reconstructible at R^2 = 0.996 given the joint input, so no information is lost, and we give the test that separates a valid re-parameterization from a lossy summary or a sensor swap. The test rejected three of the four axes we proposed. The defect lives in the action channel, which the invariance literature for visual models does not examine: work there concerns crops, jitter and camera pose, while the parameterization of the commands goes unaudited. The repair is averaging over the two encodings, and where it goes matters. Averaging the objective restores task performance by itself; averaging the outputs, safe for probabilities by concavity, is not available for direction-valued prediction, where the normalized mean can score below every member of the orbit. What objective-averaging leaves behind is the tail: worst-case agreement stays at 0.78, a disagreement penalty closes it to 0.995, and over a latent rollout it is the difference between a worst case that erodes and one that holds. On PushT, averaging alone does not repair the axis.
Chinese Translation
机器人策略通常采用两种动作参数化之一进行训练:绝对关节目标,或相对于当前状态的增量。这一选择是机器人学习中一个实际的工程决策,而以动作为条件的世界模型会默认继承该选择。我们证明这种继承是灾难性的。将在一种参数化上训练的潜在动力学模型输入以另一种参数化书写的完全相同的指令轨迹时,模型会崩溃:在三个机器人数据集和两种机器人形态上,检索性能下降2.6至13.4倍;目标条件下的动作选择率从53%降至15%;在PushT任务上,对同一未来的两种信念近乎正交(cos = 0.067,最差情况为-0.377)。因此预测器并非平缓退化,而是在回答另一个问题。这并非通常意义上的分布偏移伪影:给定关节输入,两种编码可以相互重建(R² = 0.996),即没有信息丢失;我们给出了一个检验方法,用以区分有效的重参数化与有损摘要或传感器替换。该检验否定了我们提出的四个轴中的三个。这一缺陷存在于动作通道中,而视觉模型的不变性文献并未考察该通道:那里的工作关注裁剪、抖动和相机位姿,而指令的参数化却未经审计。修复方法是对两种编码取平均,且在哪里取平均至关重要。对目标(objective)取平均本身即可恢复任务性能;对输出取平均虽然基于凹性对概率是安全的,但不适用于方向值预测——归一化均值可能低于轨道中每一个成员。目标平均所遗留的是尾部问题:最差情况的一致性停留在0.78,引入不一致性惩罚后可提升至0.995,而在潜在空间滚动预测中,这正是最差情况持续恶化与保持稳定之间的差别。在PushT上,仅靠平均无法修复该轴。
cs.RO / 69 / 2609.23263

Scenario MPC with STL Specifications and Pareto-Based Feasibility Repair

基于STL规范与Pareto可行性修复的场景模型预测控制
Wu, Tianhao, Lyu, Yiwei
Abstract
Temporal logic is a formal language for reasoning about system behaviors over time. Signal temporal logic (STL), in particular, has been used to encode spatio-temporal requirements for control synthesis in multi-agent systems, often under the assumption that agents are cooperative and their dynamics are known. However, real-world multi-agent applications, such as autonomous driving, typically involve stochastic and uncontrollable agents. Recent work explored robust control with worst-case or probabilistic formulations, but remains limited in that it either (1) certifies strict satisfaction of STL constraints without addressing feasibility recovery, or (2) relaxes infeasible constraints with ego-centric objectives. In this paper, we propose a model predictive control (MPC) framework that treats feasibility repair as a Pareto optimization problem to explicitly characterize tradeoffs among agent objectives. We further provide a probabilistic certificate on STL violation rate to formally quantify uncertainty under stochastic and uncontrollable agents. The proposed framework is evaluated on two autonomous driving scenarios. Results show that the framework recovers feasible control with demonstrated safe behaviors.
Chinese Translation
时序逻辑是一种用于对系统行为随时间演化进行推理的形式化语言。其中,信号时序逻辑(Signal Temporal Logic, STL)已被用于编码多智能体系统控制综合中的时空需求,但通常假设智能体是合作性的且其动力学已知。然而,现实中的多智能体应用(如自动驾驶)通常涉及随机且不可控的智能体。近期研究探索了基于最坏情况或概率化表述的鲁棒控制,但仍存在局限:要么(1)在不解决可行性恢复的情况下严格保证STL约束的满足,要么(2)以自我中心的目标函数松弛不可行的约束。本文提出一种模型预测控制(MPC)框架,将可行性修复视为Pareto优化问题,以显式刻画各智能体目标之间的权衡。我们进一步提供了关于STL违反率的概率证书,以形式化地量化随机且不可控智能体带来的不确定性。该框架在两个自动驾驶场景中进行了评估。结果表明,该框架能够恢复可行的控制,并展现出安全的行为。
cs.RO / 70 / 2609.23269

Latent Telepathy: Multi-Robot Communication with Self-Supervised Perceptual Latents

潜在心灵感应:基于自监督感知潜向量的多机器人通信
Wang, Howard, Zheng, Han, Wu, Cathy
Abstract
In a decentralized multi-robot team under partial observability, the fact that decides a robot's next action is often visible only to a teammate. Existing decentralized methods communicate kinematic information, such as position or planned trajectory, which cannot convey what the teammate perceives. Learned communication in multi-agent reinforcement learning (MARL) can carry perceptual content, but the resulting messages are task-coupled and opaque. We propose Latent Telepathy. Each robot broadcasts the perceptual latent vector it already computes for its own use, the output of an encoder trained with a self-supervised joint-embedding predictive objective, frozen, and shared across the team. A teammate learns to act on it from task reward alone. Because the encoder already runs for perception, the message costs no additional computation and a single compact vector of bandwidth. Because the encoder is frozen before any policy is trained, the message means the same thing to every robot, and the receiving robot is never told what it means. We evaluate Latent Telepathy with a content-controlled protocol in which bandwidth, latency, topology and receiver are held fixed and only the message content varies. Broadcasting the latent lets a navigator avoid an occluded hazard in 99.7% of episodes, matching a noiseless hand-designed message. Position and trajectory messages remain at chance, and the raw camera image, 186 times wider, is less reliable than the compressed latent. The result holds from a discrete gridworld to rendered pixels under continuous velocity control, and the encoder decodes the hazard from a physical robot's camera in 102 of 102 live decisions. We also identify a requirement for porting MARL communication results to continuous control, that the decision a message informs must remain reachable by exploration, and show how to restore it.
Chinese Translation
在部分可观测的去中心化多机器人团队中,决定机器人下一步动作的信息往往只有队友能够看到。现有的去中心化方法传输运动学信息,如位置或规划轨迹,但无法传达队友的感知内容。多智能体强化学习(MARL)中的学习式通信虽然可以携带感知内容,但由此产生的消息与任务耦合且不可解释。我们提出潜在心灵感应(Latent Telepathy)方法:每个机器人广播它本来就已为自己计算好的感知潜向量——即由自监督联合嵌入预测目标训练得到、经过冻结并在团队内共享的编码器的输出。队友仅需通过任务奖励即可学会基于该潜向量进行决策。由于编码器本身已为感知而运行,该消息不产生额外的计算开销,且仅占用单个紧凑向量的带宽。由于编码器在任何策略训练之前即被冻结,该消息对所有机器人含义一致,且接收方机器人从不需要被告知其含义。我们采用内容受控的评估协议对潜在心灵感应进行评估,在该协议中带宽、延迟、拓扑结构和接收方均保持固定,仅消息内容发生变化。广播潜向量使导航机器人能在99.7%的回合中避开被遮挡的危险物,与无噪声的人工设计消息效果相当。位置和轨迹消息仍停留在随机水平,而宽度大186倍的原始相机图像,其可靠性反而低于压缩后的潜向量。该结果从离散网格世界一直延续到连续速度控制下的渲染像素场景,且编码器在102次实时决策中有102次从实体机器人的相机中成功解码出危险物。我们还识别出将MARL通信成果迁移到连续控制中的一项必要条件:消息所告知的决策必须始终可通过探索到达,并展示了如何恢复这一条件。
cs.RO / 71 / 2609.23271

Steering Through Contact: A Finite-Support Motion Model for Single-Track Center-Articulated Robots

穿越接触的操纵:一种用于单轨中心铰接式机器人的有限支撑运动模型
Ounally, Mohamed Dhia, Samson, Nicolas, Turgeon-Roy, Mathis, Vannini, Veronica, Laconte, Johann, Pomerleau, François
Abstract
Trajectory planning and control in field robotics rely on predicting how propulsion and steering affect vehicle motion when contact points undergo slip. For articulated vehicles, the point-contact kinematic model (PCK) accounts for the linkage geometry but neglects the rotational resistance distributed along the contacts. We propose a finite-support quadratic model (FSQ) for single-track, center-articulated vehicles that incorporates this resistance through a quasi-static balance of lateral slip. Our approach generalizes the standard PCK formulation by relaxing the contact-point assumption. An exact reduction of the quadratic slip cost to contact moments gives a compact closed-form solution for real-time prediction of lateral velocity and yaw rate. We evaluate the proposed method in real-world experiments across asphalt, grass, ice, and mixed routes, using more than 7 km of data. For five-second predictions, FSQ reduces the weighted median translation and yaw errors by 53.5% and 68.9%, respectively, relative to PCK.
Chinese Translation
野外机器人技术中的轨迹规划与控制依赖于预测当接触点发生滑移时,推进与转向如何影响车辆运动。对于铰接式车辆,点接触运动学模型(PCK)考虑了连杆几何结构,但忽略了沿接触分布的旋转阻力。我们提出了一种针对单轨中心铰接式车辆的有限支撑二次模型(FSQ),通过侧向滑移的准静态平衡来纳入该阻力。我们的方法通过放宽接触点假设,推广了标准的PCK公式。通过对二次滑移代价到接触力矩的精确约简,得到了一个紧凑的闭式解,可用于侧向速度和偏航角速率的实时预测。我们在沥青、草地、冰面及混合路线的真实世界实验中评估了所提出的方法,使用了超过7公里的数据。对于五秒预测,FSQ相比PCK将加权中值平移误差和偏航误差分别降低了53.5%和68.9%。
cs.RO / 72 / 2609.23275

SCULPT-VLA: Learning Structured Control through Staged Action Grounding

SCULPT-VLA:通过分阶段动作接地学习结构化控制
Li, Wenbo, Chen, Yiteng, Zhang, Wei, Li, Wenhao, Yang, Jun, Wu, Qingyao
Abstract
Vision-language-action (VLA) policies increasingly incorporate structured intermediate supervision beyond action labels. Yet specifying what an intermediate representation should encode leaves open how action prediction learns to depend on it. We introduce \textbf{SCULPT-VLA}, a policy that learns structured control through staged action grounding. Its action-conditioning state comprises complementary factors for task progression, scene dynamics, and spatial grounding. Training first forms these factors with teacher scaffolds, then grounds coarse action prediction through their composition as scaffold inputs are withdrawn. Direct perceptual access is subsequently restored for continuous refinement, combining the learned state with perceptual detail. The curriculum separates learning to condition actions on structure from refining continuous control. Deployment requires neither teachers nor discrete-action autoregression. SCULPT-VLA achieves higher average success than shared-backbone baselines on LIBERO, SimplerEnv-WidowX, and RoboTwin 2.0 Full. On SimplerEnv-WidowX, final success is 83.5\%, versus 71.3\% when Stage-II action learning directly accesses vision and language. Across four physical robot tasks, average success under the tested distribution shifts reaches 58.1\%, compared with 45.6\% for $\pi_{0.5}$. Training ablations and factor-wise interventions support the staged design and show that the learned state continues to contribute to control after direct perceptual access is restored.
Chinese Translation
视觉-语言-动作(VLA)策略日益纳入超越动作标签的结构化中间监督。然而,指定中间表示应编码什么内容,仍未解决动作预测如何学会依赖于该表示的问题。我们提出SCULPT-VLA,一种通过分阶段动作接地学习结构化控制的策略。其动作条件状态由互补的因子构成,分别对应任务进展、场景动态和空间接地。训练首先借助教师支架形成这些因子,然后随着支架输入的逐步撤除,通过它们的组合来接地粗粒度动作预测。随后恢复直接的感知访问,以结合所学状态与感知细节进行连续优化。该课程将“学习以结构为条件预测动作”与“优化连续控制”分离开来。部署时既不需要教师,也不需要离散动作的自回归过程。SCULPT-VLA在LIBERO、SimplerEnv-WidowX和RoboTwin 2.0 Full上取得了比共享主干基线更高的平均成功率。在SimplerEnv-WidowX上,最终成功率为83.5%,而阶段II动作学习直接访问视觉和语言时仅为71.3%。在四项真实机器人任务中,在所测试的分布偏移下平均成功率达58.1%,相比之下π0.5为45.6%。训练消融实验和按因子干预的结果支持了这种分阶段设计,并表明在恢复直接感知访问后,所学状态仍持续为控制做出贡献。
cs.RO / 73 / 2609.23280

Safety-Critical Control under Uncertainty via Adaptive Conformal Quantile Prediction Intervals

基于自适应保形分位数预测区间的不确定性下安全关键控制
Zhou, Hao, Zhang, Yanze, Lyu, Yiwei, Luo, Wenhao
Abstract
Safety-critical control under uncertainty requires uncertainty representations that are both statistically valid (for certifiable performance) and compatible with enforceable safety constraints. However, existing methods often assume particular distributions of uncertainty for provable safety guarantees or establish symmetric and input-agnostic prediction intervals for robust safety, which can lead to misaligned or overly conservative safety constraints in control synthesis. In this paper, we introduce a novel safe control framework with adaptive uncertainty quantification that constructs calibrated and state-dependent prediction intervals to enable high-probability safety guarantees, while improving constrained control performance. The framework leverages adaptive conformal prediction (ACP) and extends it with conformal quantile regression (CQR) to capture distribution-free, asymmetric uncertainty intervals with certifiable probabilistic coverage, and integrates the resulting uncertainty sets into a probabilistic control barrier function formulation to enforce robust safety with reduced conservativeness. This yields uncertainty-aware safe control constraints that can be incorporated within a model predictive control(MPC) framework to provide provably safe behaviors with high probability. Simulation and theoretical results are provided to demonstrate the effectiveness of our approach.
Chinese Translation
不确定性下的安全关键控制要求不确定性表征既具有统计有效性(以保证可认证的性能),又能与可执行的安全约束相兼容。然而,现有方法通常为了获得可证明的安全保证而假设不确定性服从特定分布,或者为鲁棒安全而建立对称且与输入无关的预测区间,这可能导致控制综合中出现安全约束不匹配或过于保守的问题。本文提出了一种新颖的具有自适应不确定性量化的安全控制框架,该框架通过构建经校准且依赖于状态的预测区间,实现高概率安全保证,同时提升受限控制性能。该框架利用自适应保形预测(ACP),并结合保形分位数回归(CQR)加以扩展,以捕获无分布假设的、非对称的不确定性区间并保证可认证的概率覆盖率,进而将所得的不确定性集合融入概率控制障碍函数(Probabilistic Control Barrier Function)框架中,以在降低保守性的前提下强制实现鲁棒安全。由此得到的不确定性感知安全控制约束可嵌入模型预测控制(MPC)框架中,以高概率提供可证明安全的行为。仿真和理论结果验证了该方法的有效性。
cs.RO / 74 / 2609.23305

Shared Execution-Clock Drifting Policy for Dynamic Precision Manipulation

面向动态精度操控的共享执行时钟漂移策略
Dong, Zhenchen, Wu, Qingran, Fu, Jinna, Wu, Jiaming, Chen, Fulin, Yu, Hongyu, Liu, Yide
Abstract
Manipulation under time constraints requires both accurate actions and an execution rhythm that matches the evolving scene. This becomes critical when a robot must intercept moving objects or complete a sequence of adjustments before a deadline. Although one-step policies reduce generation cost, their directly predicted action sequences leave temporal allocation implicit. We propose Shared Execution-Clock Drifting (SECD), which makes execution rhythm an explicit part of one-step action generation. Conditioned on an observation and a latent sample, the policy jointly predicts a progress-indexed action curve and a shared monotone clock that maps fixed control times to locations on the curve. Demonstration-derived alignment anchors this decomposition, which is trained jointly through drifting on the decoded actions. The resulting policy retains a fixed-rate control interface and requires one network evaluation. We evaluate SECD across four real-robot tasks with inference on NVIDIA Thor. Across 300 trials, it achieves 77.00% task-averaged success and outperforms the evaluated one-step baselines on every task, including 91% success in cup retrieval from a 16 m/min conveyor and 54% in restoring and folding a crumpled shirt within 90 s. A fixed-clock variant reaches 79% on the same conveyor protocol. Complementary state-based RoboMimic experiments, including cross-seed ablations on Transport and Square, further support the joint design of the temporal representation and demonstration alignment. Project page: https://secd-anonymous-ewn.pages.dev/
Chinese Translation
时间约束下的操控任务既需要精确的动作,也需要与不断演化的场景相匹配的执行节奏。当机器人必须拦截运动物体或在截止期限前完成一系列调整时,这一点尤为关键。尽管单步(one-step)策略降低了生成成本,但其直接预测的动作序列使时间分配处于隐式状态。我们提出共享执行时钟漂移(Shared Execution-Clock Drifting, SECD)方法,将执行节奏显式地纳入单步动作生成过程。该策略以观测和一个隐变量样本为条件,联合预测一条以进度为索引的动作曲线以及一个共享的单调时钟,该时钟将固定的控制时刻映射到曲线上的位置。基于演示数据导出的对齐锚定这一分解结构,并通过在解码动作上进行漂移来联合训练。由此得到的策略保留了固定频率的控制接口,且每次仅需一次网络前向推理。我们在四个真实机器人任务上评估了SECD,推理硬件为NVIDIA Thor。在300次试验中,该方法取得了77.00%的任务平均成功率,并在每个任务上都优于所评估的单步基线方法,其中包括从16 m/min传送带上抓取杯子达到91%的成功率,以及在90秒内完成皱缩衬衫的整理与折叠达到54%的成功率。固定时钟变体在相同的传送带协议下达到79%的成功率。基于状态的RoboMimic补充实验,包括在Transport和Square任务上的跨随机种子消融研究,进一步验证了时间表征与演示对齐联合设计的有效性。项目页面:https://secd-anonymous-ewn.pages.dev/
cs.RO / 75 / 2609.23312

Manipulation Feasible Navigation Among Movable Obstacles with Discrete Contact Pushing

基于离散接触推挪的可操纵可行移动障碍物间导航
Wang, Shaohu, Song, Aiguo, Yuan, Yulong, Sun, Zhongyu, Miao, Tianyuan, Ji, Qinjie
Abstract
In environments with large movable obstacles, detour-only navigation can be inefficient or even infeasible, while obstacle interaction requires reasoning about navigation benefit, feasible placement, and executable manipulation. We present a hierarchical navigation among movable obstacles (NAMO) framework for mobile manipulators. At the high level, the planner identifies key blocking obstacles from reference paths and searches for relocation plans that jointly satisfy geometric, manipulation, and downstream navigation constraints. When direct relocation is hindered by other movable objects, a large language model (LLM) is selectively invoked to infer auxiliary manipulation dependencies, which are then verified by deterministic geometric planning. To execute the resulting relocation goals, we define discrete contact modes on the surfaces of box-shaped obstacles and select contact faces and regions online based on position and orientation errors, enabling straight, side, and corner pushing through contact switching. A recurrent reinforcement-learning policy coordinates the mobile base and manipulator to track tool center point (TCP) targets while preserving end-effector reachability during sustained pushing. Simulation and real-robot experiments demonstrate feasible navigation-manipulation in detour, single- and multi-obstacle relocation, and dependency-constrained scenarios, validating the framework for interactive navigation with large non-graspable obstacles. The open-source project is available at https://cloudytosunny.github.io/NAMO_DCPushing/.
Chinese Translation
在存在大型可移动障碍物的环境中,仅绕行的导航方式可能效率低下甚至不可行,而与障碍物交互则需要综合考虑导航收益、可行放置位置以及可执行的操纵动作。我们提出了一种面向移动操纵机器人(mobile manipulator)的分层式可移动障碍物间导航(Navigation Among Movable Obstacles, NAMO)框架。在高层,规划器从参考路径中识别关键阻挡障碍物,并搜索同时满足几何约束、操纵约束与下游导航约束的搬移规划。当直接搬移受到其他可移动物体阻碍时,系统有选择地调用大语言模型(LLM)来推断辅助操纵依赖关系,并通过确定性的几何规划进行验证。为执行所得到的搬移目标,我们在箱形障碍物表面上定义了离散接触模式,并根据位置和姿态误差在线选择接触面与接触区域,通过接触切换实现直线推、侧推和转角推。一个循环强化学习策略协调移动底盘与机械臂以跟踪工具中心点(TCP)目标,同时在持续推动过程中保持末端执行器的可达性。仿真与真实机器人实验在绕行、单障碍物与多障碍物搬移以及依赖约束场景中展示了可行的导航-操纵能力,验证了该框架在与大型不可抓取障碍物进行交互式导航方面的有效性。开源项目见 https://cloudytosunny.github.io/NAMO_DCPushing/ 。
cs.RO / 76 / 2609.23392

Towards Reliable Underwater Diver-Robot Interaction: Gesture Design, Interaction Logic, and Real-World Evaluation

迈向可靠的水下潜水员-机器人交互:手势设计、交互逻辑与真实环境评估
Liu, Yingqi, Yao, Kanzhong, Peng, Zimeng, Bi, Yuanbo, Li, Anran, Sun, Zhe, Li, Xuelong
Abstract
Underwater human--robot interaction requires gesture commands that are both easy for divers to use and reliable for robots to recognize. We investigate these aspects through a closed-loop diver--robot interaction framework integrating a compact seven-gesture vocabulary, lightweight landmark-based recognition, and command-level interaction logic. We evaluate the framework through a user study and underwater robot experiments in a laboratory tank and a swimming pool. The user study supported the reproducibility of the gestures after brief learning. Recognition analysis further showed that visual similarity was associated with gesture confusion, while intermediate poses during gesture formation introduced temporal ambiguity. Command-level processing mitigated the effects of transient recognition errors on robot execution, reducing unintended triggers and premature task interruptions. These findings show that reliable underwater gesture interaction depends on human usability, gesture recognizability, and execution reliability in underwater interaction.
Chinese Translation
水下人机交互需要潜水员易于使用且机器人能够可靠识别的手势命令。我们通过一个闭环潜水员-机器人交互框架来研究这些问题,该框架集成了紧凑的七手势词汇表、轻量级的基于关键点的识别方法以及命令级交互逻辑。我们通过用户研究以及在实验室水箱和游泳池中开展的水下机器人实验对该框架进行了评估。用户研究结果表明,手势经过简单学习后即可被稳定复现。识别分析进一步表明,手势之间的视觉相似性会导致手势混淆,而手势形成过程中的中间姿态会引入时间上的歧义。命令级处理缓解了瞬态识别错误对机器人执行的影响,减少了意外触发和任务过早中断。这些研究结果表明,可靠的水下手势交互依赖于人的易用性、手势的可识别性以及水下交互中的执行可靠性。
cs.RO / 77 / 2609.23414

EmoPose: Vision-Language Model Guided Emotion-Aware Gesture Generation for Humanoid Robots

EmoPose:视觉语言模型引导的面向人形机器人的情感感知手势生成
Peng, Daojie, Wang, Bingtao, Ma, Fulong, Yue, Wenjun, Zhang, Liang, Ma, Jun
Abstract
Socially competent humanoid robots must communicate affect and intent through gesture as well as speech, yet open-ended interaction must become motion that is both expressive and executable on a specific body. This demands semantic flexibility for contextual social intent while preserving deterministic, embodiment-aware robot control. We present EmoPose, a vision-language model (VLM)-guided framework that bridges this gap through an executable semantic interface. Given language, dialogue history, and optional visual context, the VLM selects an ordered gesture plan containing a communicative class, library variant, intensity, and speech anchor. A scalable robot-owned motion library defines the available expressive vocabulary and the source of 14-DoF joint targets. Pose Studio supports automatic trajectory generation, MuJoCo preview, and automatic synchronization of new library entries with the VLM guide; deterministic robot-side modules validate plans, construct trajectories, schedule gestures, and manage queueing and interruption. This division lets the interaction repertoire grow for new social contexts without changing the control interface or delegating raw joint commands to the foundation model. On the EmoPose-Bench, structured GPT-5.5 planning reaches $98.25\pm0.52\%$ on the Easy tier and $76.50\pm0.54\%$ overall, exceeding same-model direct-label prompting. Further tests validate dialogue-context use and ordered multi-action composition. The system completes the nominal MuJoCo suite and realizes all 29 authored variants on the physical Unitree G1. A four-stop laboratory tour demonstrates expressive narration with interruption, camera-grounded dialogue, and navigation.
Chinese Translation
具备社交能力的人形机器人必须通过手势和语音来传达情感与意图,而开放式交互必须转化为既具表现力又能在特定机器人身体上执行的动作。这既要求具备针对情境化社交意图的语义灵活性,又要保持确定性的、具身感知的机器人控制。我们提出EmoPose,一个由视觉语言模型(VLM)引导的框架,通过可执行的语义接口弥合了这一差距。给定语言、对话历史和可选的视觉上下文,VLM会选择一个有序的手势计划,其中包含交流类别、库变体、强度和语音锚点。一个可扩展的、由机器人自主拥有的动作库定义了可用的表达词汇以及14自由度关节目标的来源。Pose Studio支持自动轨迹生成、MuJoCo预览,以及新库条目与VLM指南的自动同步;确定性的机器人端模块负责验证计划、构建轨迹、调度手势,并管理排队与中断。这种分工使交互技能库能够针对新的社交情境进行扩展,而无需更改控制接口,也无需将原始关节指令委派给基础模型。在EmoPose-Bench上,结构化的GPT-5.5规划在简单层级上达到98.25±0.52%,总体达到76.50±0.54%,超过了同模型的直接标签提示方法。进一步的测试验证了对对话上下文的利用以及有序的多动作组合能力。该系统完成了MuJoCo名义测试套件,并在物理Unitree G1机器人上实现了全部29个编排的变体。一个四站式实验室导览展示了带中断的表现性讲解、基于摄像头的对话以及导航功能。
cs.RO / 78 / 2609.23418

HEARTH: An Object-Centric RGB-Thermal-3D Dataset for Temperature-Aware Robot Manipulation

HEARTH:一个用于温度感知机器人操作的对象级RGB-热成像-3D数据集
Su, Yuning, Li, Borui, Shi, Yonghao, Liu, Bofei, Yang, Xing-Dong
Abstract
Language-guided manipulation can depend on physical properties that visible appearance does not reveal. Temperature is one such property, but object datasets for robot learning rarely associate measured temperatures with object appearance and geometry. We present HEARTH, an object-centric RGB-thermal-3D dataset of 90 physical objects from 18 everyday categories, comprising 145 captured object states. Our pipeline maps apparent surface temperatures onto reconstructed meshes through camera calibration and pose transfer. The dataset includes raw temperature measurements, camera parameters, RGB-textured meshes, and thermal textures for simulation. We use these assets to construct three LIBERO-derived tasks and collect 1,200 demonstrations for fine-tuning a pretrained vision-language-action (VLA) model, $\pi_{0.5}$. In an ablation study, adding thermal observations to the VLA increases success on temperature-dependent object-selection tasks from 35.0% for the RGB-only baseline to 75.0%. These results demonstrate the utility of HEARTH for training robot policies to follow temperature-related instructions.
Chinese Translation
语言引导的操作可能依赖于外观无法揭示的物理属性。温度就是其中之一,但面向机器人学习的物体数据集很少将实测温度与物体的外观和几何信息关联起来。我们提出了HEARTH,一个对象级的RGB-热成像-3D数据集,包含来自18个日常类别的90个实物物体,共计145个采集的物体状态。我们的流程通过相机标定与位姿迁移,将表观表面温度映射到重建的网格模型上。该数据集包括原始温度测量数据、相机参数、RGB纹理网格以及用于仿真的热纹理。我们利用这些资产构建了三个基于LIBERO的任务,并收集了1,200条演示数据,用于微调预训练的视觉-语言-动作(VLA)模型 $\pi_{0.5}$。在消融实验中,为VLA模型添加热观测后,其在依赖温度的物体选择任务上的成功率从仅使用RGB的基线的35.0%提升至75.0%。这些结果证明了HEARTH在训练机器人策略以遵循温度相关指令方面的实用价值。
cs.RO / 79 / 2609.23423

RiverVLN: Phase-Grounded Temporal Vision--Language Navigation for Unmanned Surface Vehicles

RiverVLN:面向无人水面艇的相位锚定时序视觉—语言导航
Wu, Jieling, Huang, Yuehao, Lv, Jiajun, Huang, Tao, Liu, Yong, Liu, Weiwei
Abstract
Vision-language navigation (VLN) has largely been developed for indoor and terrestrial robots, where language can often be treated as a static goal and motion is approximated by discrete or near-instantaneous actions. These assumptions break down for unmanned surface vehicles (USVs): river navigation requires continuous motion under inertia and limited maneuverability, while long-horizon instructions must be executed through sparse and visually ambiguous maritime landmarks. We introduce RiverVLN, to our knowledge the first benchmark designed for long-horizon USV VLN under continuous riverine motion, and PGT-NAV, a phase-grounded temporal navigation framework for USVs. Rather than directly mapping an entire instruction to motion, PGT-NAV converts it into an ordered sequence of visually verifiable semantic phases and maintains the active phase online through grounded visual and motion evidence. This explicit semantic progress state is fused with visual-motion history and phase-specific grounding to predict six local SE(2) pose increments. The resulting trajectory is executed in a predict-execute-re-observe loop, where the vessel executes toward W3, updates phase and grounding, and replans through a map-based safety layer. Experiments show that PGT-NAV substantially reduces recursive position and heading drift relative to GNM-style and ViNT-style baselines and achieves an average success rate of 0.79 in Unity-ROS closed-loop navigation. Unseen bridge-opening trials and real-world USV experiments further demonstrate that the phase-grounded representation transfers from controlled evaluation to physical USV deployment.
Chinese Translation
视觉—语言导航(Vision-Language Navigation, VLN)主要针对室内和陆地机器人发展而来,在这类场景中,语言通常可被视为静态目标,运动则通过离散或近瞬时动作来近似。这些假设对无人水面艇(Unmanned Surface Vehicles, USVs)不再成立:河流航行需要在惯性和机动性受限的条件下进行连续运动,同时长时程指令必须借助稀疏且视觉上模糊的海事地标来执行。我们提出了RiverVLN——据我们所知,这是首个面向连续河流运动下长时程无人水面艇VLN的基准数据集,并提出了PGT-NAV,一种面向无人水面艇的相位锚定时序导航框架。PGT-NAV并非将整条指令直接映射为运动,而是将其转换为一系列有序的、可通过视觉验证的语义相位,并通过锚定的视觉与运动证据在线维护当前活跃相位。该显式的语义进度状态与视觉—运动历史及相位特定的锚定信息相融合,用于预测六个局部SE(2)位姿增量。最终得到的轨迹在“预测—执行—再观测”循环中执行:船舶朝向W3执行,更新相位与锚定信息,并通过基于地图的安全层进行重规划。实验表明,相较于GNM和ViNT风格的基线方法,PGT-NAV显著降低了递归位置与航向漂移,并在Unity-ROS闭环导航中取得了0.79的平均成功率。针对未见桥洞场景的试验以及真实世界的无人水面艇实验进一步证明,这种相位锚定的表示能够从受控评估迁移到实际的无人水面艇部署中。
cs.RO / 80 / 2609.23432

RopeFormer: Cross-Trial Adaptation from Interaction History for Dynamic Rope Manipulation

RopeFormer:基于交互历史的跨试验自适应动态绳索操纵
Wu, Menglin, Yao, Kaixiang, Luan, Shangbo, Tomizuka, Masayoshi, Chen, Yuxin
Abstract
Dynamic rope manipulation is highly sensitive to unknown object dynamics: the same robot motion can produce substantially different responses across ropes, while explicitly identifying the relevant physical properties is difficult. We present RopeFormer, a history-conditioned framework that uses prior task interaction as context for subsequent control. The policy retains cross-trial action-response history while keeping its weights fixed and requires no explicit online rope-parameter estimation. In matched simulation evaluations across sustained single-arm rotation, bimanual rotation, and transient whipping, retaining context improves subsequent control relative to resetting the same checkpoint, with the benefit varying across rope dynamics and observation settings. We further deploy the frozen policies on a Unitree H1-2 with previously unseen physical ropes. From T1 to T3, target-acquisition time decreases by 30.9% for Rope Swing and 33.9% for Rope Twirl, while mean Rope Whip target hits increase from 0.2 to 2.3 out of three. These results show that prior interaction can provide effective control context for dynamic deformable-object manipulation. Robot videos, code, and data are available at https://ropeformer.github.io/.
Chinese Translation
动态绳索操纵对未知物体动力学高度敏感:相同的机器人动作在不同绳索上可能产生显著不同的响应,而显式识别相关物理属性十分困难。我们提出了RopeFormer,一个以历史条件为驱动的框架,将先前的任务交互作为后续控制的上下文。该策略在保持网络权重固定的情况下保留跨试验的动作-响应历史,且无需显式的在线绳索参数估计。在涵盖持续单臂旋转、双臂旋转和瞬时甩动的匹配仿真评估中,与重置同一模型检查点相比,保留上下文能够改善后续控制效果,且其收益因绳索动力学和观测设置的不同而有所差异。我们进一步将冻结策略部署到Unitree H1-2机器人上,并使用此前未见过的真实物理绳索进行实验。从试验T1到T3,Rope Swing(绳索摆动)的目标获取时间减少了30.9%,Rope Twirl(绳索旋转)减少了33.9%,而Rope Whip(绳索甩击)的平均目标命中次数从三次中的0.2次提升至2.3次。这些结果表明,先前的交互可以为动态可变形物体操纵提供有效的控制上下文。机器人视频、代码和数据可在 https://ropeformer.github.io/ 获取。
cs.RO / 81 / 2609.23439

Receding-Horizon Pushing with Composable Object-Centric Policies

基于可组合物体中心策略的滚动时域推物方法
Yuan, Zhiyi, Hu, Tianrun, Xiao, Anxing, Deng, Yuhong, Hsu, David, Zhang, Hanbo
Abstract
Non-prehensile manipulation is practical for relocating large, heavy, or geometrically ungraspable objects. Yet, long-horizon pushing of arbitrarily-shaped 3D objects couples three problems: 1) where to push the object so as to approach the target pose, 2) whether each push is stable and reachable, 3) whether subsequent actions remain feasible. We present an object-centric pushing policy within a feedback-guided hierarchical framework. At the low level, a learning-based policy predicts contact actions from a pose- and scale-normalized point cloud, conditioned on a near single-step subgoal. A stability score is applied to evaluate the predicted contacts by a quasi-static sliding-versus-tipping analysis. At the high level, BIT$^*$ first searches for an object path, and the next several subgoals are checked by contact prediction and robot motion planning for future feasibility. Failed motion plans, as feedback, change the local path costs and trigger re-planning. During execution, only the first feasible action is executed. In simulation, we evaluate 22 objects in six different scenes, upon which we also conduct comprehensive ablation studies. Results demonstrate that our method outperforms baselines with a clear margin and can reliably achieve long-horizon object pushing tasks under different situations. We also report quantitative real-robot experiments with a Franka arm and qualitative demonstrations with a mobile manipulator for large and heavy objects, with directly zero-shot sim-to-real transfer.
Chinese Translation
非抓取式操作(non-prehensile manipulation)对于搬运大型、重型或几何上不可抓取的物体十分实用。然而,对任意形状的3D物体进行长时程推物操作涉及三个耦合问题:1)应从何处推动物体以使其趋近目标位姿;2)每次推动是否稳定且可达;3)后续动作是否仍然可行。我们提出了一种物体中心的推物策略,并将其嵌入一个反馈引导的分层框架中。在低层,一个基于学习的策略从经过位姿与尺度归一化的点云中预测接触动作,并以接近单步的子目标为条件。通过准静态的滑动-倾覆分析计算稳定性分数,用以评估预测的接触动作。在高层,BIT$^*$首先搜索物体路径,随后若干子目标则通过接触预测和机器人运动规划来检查其未来可行性。失败的机器人运动规划作为反馈,改变局部路径代价并触发重规划。在执行过程中,仅执行第一个可行动作。在仿真中,我们在六个不同场景下评估了22个物体,并开展了全面的消融实验。结果表明,我们的方法以明显优势超越基线方法,并能在不同情形下可靠地完成长时程推物任务。我们还报告了基于Franka机械臂的定量真机实验,以及使用移动操作臂搬运大型重型物体的定性演示,实现了直接的零样本(zero-shot)仿真到现实(sim-to-real)迁移。
cs.RO / 82 / 2609.23445

BiRoAD: Learning Shared and Role-Adaptive Representations for Bimanual Manipulation

BiRoAD:面向双手操作的学习共享与角色自适应表示
Shen, Yan, Liu, Yuchen, Jiang, Feng, Hu, Hangtian, Li, Xiaoqi, Chen, Shu, Wu, Ruihai, Dong, Hao
Abstract
Bimanual manipulation requires policies that coordinate two arms while adapting their functional roles to scene geometry, object configuration, and task context. Learning such scene-conditioned role adaptation remains challenging, as demonstrations may contain uneven role distributions that limit generalization to underrepresented arm--role configurations. In addition, many bimanual policies predict actions in fixed left- and right-arm action spaces. While this provides a natural parameterization for robot control, it does not explicitly specify how behaviors should transform when functional roles are exchanged across arms. Across different scene initializations, the two arms may follow a similar coordination pattern, but the role-specific behavior assigned to each arm should change with the scene. Therefore, we propose BiRoAD, a Bimanual Role-Adaptive Decomposition framework for learning shared and role-adaptive representations in bimanual policies. Given bimanual trajectory or action-token features, BiRoAD decomposes these features into swap--symmetric and swap--antisymmetric components: the former captures coordination structure invariant to arm exchange, and the latter captures role-specific distinctions that vary consistently with functional role assignment. The two components are then recomposed as residual updates to the original paired arm representations, allowing BiRoAD to serve as a modular feature transformation without changing the policy inputs, imitation-learning objective, or requiring manually defined role labels. Across multiple bimanual manipulation tasks with balanced and imbalanced role distributions, BiRoAD improves robustness across role configurations over corresponding base policies, with notable gains on underrepresented role configurations.
Chinese Translation
双手操作需要能够协调两条机械臂的策略,同时使其功能角色适应场景几何、物体构型和任务上下文。学习这种以场景为条件的角色自适应仍然具有挑战性,因为演示数据中可能包含不均匀的角色分布,从而限制了对代表性不足的机械臂—角色构型的泛化能力。此外,许多双手操作策略在固定的左臂和右臂动作空间中预测动作。虽然这为机器人控制提供了自然的参数化方式,但它并未明确指定当功能角色在两臂之间交换时,行为应如何变换。在不同的场景初始化条件下,两条机械臂可能遵循相似的协调模式,但分配给每条机械臂的角色特定行为应随场景而变化。因此,我们提出 BiRoAD,一种双手角色自适应分解(Bimanual Role-Adaptive Decomposition)框架,用于在双手操作策略中学习共享的和角色自适应的表示。给定双手轨迹或动作令牌(action-token)特征,BiRoAD 将这些特征分解为交换对称(swap-symmetric)和交换反对称(swap-antisymmetric)两个分量:前者捕获对手臂交换不变的协调结构,后者捕获随功能角色分配而一致变化的角色特定差异。随后,这两个分量被重组为对原始成对机械臂表示的残差更新,使 BiRoAD 能够作为一种模块化的特征变换,而无需改变策略输入、模仿学习目标,也无需人工定义的角色标签。在具有均衡与非均衡角色分布的多个双手操作任务上,BiRoAD 相比相应的基础策略提升了跨角色构型的鲁棒性,在代表性不足的角色构型上取得了显著的性能提升。
cs.RO / 83 / 2609.23456

Smartphone GNSS Booster: Centimeter-Level Pedestrian Positioning Using a Portable Signal Re-Radiator

智能手机GNSS增强器:利用便携式信号再辐射器实现厘米级行人定位
Suzuki, Taro
Abstract
High-precision positioning using the global navigation satellite system (GNSS) embedded in smartphones is demanded for sidewalk-level pedestrian navigation and pinpoint location-based applications. However, current smartphone positioning accuracy for pedestrians remains at the meter level. This is mainly because limitations of the compact linearly polarized (LP) antenna in smartphones increase GNSS observation noise and hinder stable carrier phase tracking. In this paper, we propose an external GNSS signal re-radiation system that boosts smartphone GNSS observation quality while still using the smartphone's built-in GNSS receiver and antenna. The system directly connects a compact active helical antenna and a thin passive patch antenna for re-radiation, and is designed to be attached directly to the smartphone. The proposed GNSS booster achieves (1) reduced thermal noise and stable signal tracking via increased received signal strength, (2) improved multipath robustness compared with a LP antenna, and (3) stabilization of antenna phase center variation. Static experiments confirm a substantial improvement in the observation quality of GNSS carrier phase measurements compared with standard smartphone positioning. Furthermore, in pedestrian experiments, the proposed booster enabled stable carrier phase integer ambiguity resolutions, achieving centimeter-level positioning using a smartphone.
Chinese Translation
利用智能手机内置的全球导航卫星系统(GNSS)实现高精度定位,是人行道级行人导航和精确定位类应用的需求。然而,目前智能手机的行人定位精度仍停留在米级。这主要是由于智能手机中紧凑型线极化(LP)天线的局限性增大了GNSS观测噪声,并阻碍了稳定的载波相位跟踪。本文提出了一种外部GNSS信号再辐射系统,在仍使用智能手机内置GNSS接收机和天线的情况下,提升智能手机GNSS观测质量。该系统将一个紧凑的有源螺旋天线与一个薄片无源贴片天线直接相连用于信号再辐射,并设计为可直接附着在智能手机上。所提出的GNSS增强器实现了:(1)通过提高接收信号强度降低热噪声并实现稳定的信号跟踪;(2)相比线极化天线提高了多径鲁棒性;(3)稳定了天线相位中心变化。静态实验证实,与标准智能手机定位相比,GNSS载波相位测量观测质量得到显著改善。此外,在行人实验中,所提出的增强器实现了稳定的载波相位整周模糊度固定,利用智能手机达到了厘米级定位精度。
cs.RO / 84 / 2609.23483

STRIDER: Stepping-Enabled Multi-Gait Hierarchical 3D Loco-Manipulation Framework for Humanoid Robots

STRIDER:面向人形机器人的支持跨步的多步态分层式三维移动操作框架
Li, Yuanzhuo, Zhao, Wen, Yong, Zhe, Meng, Xiang, Han, Gang, Ren, Hengle, Zheng, Xiaoyang, Wang, Zhen, Guo, Yijie
Abstract
Humanoid loco-manipulation faces two prominent limitations: controllers using continuous velocity commands cannot precisely regulate individual footholds, while specialized foothold-tracking modules are difficult to integrate with whole-body manipulation. Furthermore, standard action-based imitation distillation primarily transfers expert actions, without explicitly encouraging a shared representation of heterogeneous skills. This paper introduces STRIDER, a hierarchical multi-gait framework to bridge these gaps. The framework integrates terrain-aware 3D stepping logic, Adversarial Motion Priors (AMP)-based natural walking, and Cartesian upper-body control: its stepping expert selects feasible footholds in the stance-foot frame and generates clearance-aware swing trajectories. To fuse distinct walking and stepping experts into one executable student policy, we propose Latent Distillation Proximal Policy Optimization (LD-PPO), a distillation algorithm augmented with teacher-conditioned latent alignment. By jointly optimizing on-policy reinforcement learning, DAgger-based action reconstruction, and latent alignment, LD-PPO transfers expert actions while encouraging a shared skill representation across heterogeneous modes. Simulation and real-robot evaluations on the TianGong Omni humanoid show that LD-PPO outperforms vanilla distillation-PPO in foothold-tracking and posture-tracking accuracy. Deployed on hardware, STRIDER realizes multi-gait loco-manipulation with accurate foothold and end-effector tracking.
Chinese Translation
人形机器人的移动操作(loco-manipulation)面临两个突出的局限:采用连续速度指令的控制器无法精确调节单个落脚点,而专门的落脚点跟踪模块又难以与全身操作相集成。此外,标准的基于动作的模仿蒸馏主要迁移专家动作,缺乏对异构技能共享表征的显式鼓励。本文提出STRIDER,一个弥合上述不足的分层多步态框架。该框架集成了地形感知的三维跨步逻辑、基于对抗运动先验(Adversarial Motion Priors, AMP)的自然行走以及笛卡尔空间上半身控制:其跨步专家在支撑足坐标系中选择可行的落脚点,并生成具备间隙感知的摆动腿轨迹。为了将不同的行走专家与跨步专家融合为单一可执行的学生策略,我们提出了带教师条件潜在对齐增强的蒸馏算法——潜在蒸馏近端策略优化(Latent Distillation Proximal Policy Optimization, LD-PPO)。通过联合优化在线强化学习、基于DAgger的动作重构以及潜在对齐,LD-PPO在迁移专家动作的同时,促进跨异构模式的共享技能表征。在天工Omni(TianGong Omni)人形机器人上的仿真与真机评估表明,LD-PPO在落脚点跟踪和姿态跟踪精度上优于原始的蒸馏-PPO方法。在硬件部署中,STRIDER实现了具有精确落脚点与末端执行器跟踪的多步态移动操作。
cs.RO / 85 / 2609.23486

Cognitive Action Reasoning for Proactive Robots from Human-Centered Multimodal Observations

基于以人为中心的多模态观察的主动机器人认知动作推理
Gu, Zhihao, Zhu, Kechao, Wu, Yuanfeng, Liu, Mohan, Shaw, Ankit Kumar, Hong, ChenDong, Chen, Xuanyu, Mei, Dengchen, Tianyi, Xu, Wang, Lin
Abstract
Robots operating in human-centered environments are typically designed to execute explicit instructions, and most robot-learning datasets likewise pair observations with task instructions or low-level actions. Although recent work has begun to explore proactive embodied assistance, existing resources target different settings and action levels, leaving real-world human-centered multimodal decision-making underexplored. We formulate this problem as \textit{Proactive Robot Action Reasoning} (\textit{ProRobo}), an upstream cognitive decision problem in which a robot must determine which action to take based on multimodal human and environmental cues without explicit action instructions. To support ProRobo, we introduce \textit{ProAction}, a real-world multimodal dataset containing 10K samples of visual observations, audio signals, and text inputs across 12 daily-life scenarios in five common scenes. To construct cognitively grounded high-level action supervision, we develop a two-stage human-in-the-loop pipeline that combines appraisal-guided candidate generation with Affective Theory-of-Mind-guided human refinement, explicitly incorporating contextual judgment about human states, urgency, feasibility, and potential risk into action annotation. Based on this supervision, we benchmark representative Multimodal Large Language Models (MLLMs) and introduce \textit{MMC2Act}, a reference model that implicitly learns the mapping from multimodal observations to cognitively grounded high-level actions. Experiments across modality settings, subject-disjoint generalization, cross-dataset transfer, and human evaluation show that general-purpose MLLMs struggle with proactively reasoning high-level actions from multimodal cues, whereas training on \textit{ProAction} substantially improves performance.
Chinese Translation
在以人为中心的环境中运行的机器人通常被设计为执行明确的指令,而且大多数机器人学习数据集也同样是把观察与任务指令或底层动作配对。尽管最近已有工作开始探索主动具身辅助,但现有资源针对的是不同的设置和动作层级,使得真实世界中以人为中心的多模态决策问题仍未得到充分研究。我们将该问题形式化为“主动机器人动作推理”(Proactive Robot Action Reasoning,ProRobo),这是一个上游认知决策问题:机器人必须基于多模态的人类和环境线索,在没有明确动作指令的情况下确定应采取何种动作。为支持 ProRobo,我们提出了 ProAction,一个真实世界的多模态数据集,包含 1 万条样本,涵盖五个常见场景中 12 个日常生活情境下的视觉观察、音频信号和文本输入。为构建具有认知依据的高层动作监督,我们开发了一个两阶段人机协同(human-in-the-loop)流程,将评价驱动的候选生成与情感心智理论(Affective Theory-of-Mind)指导的人工修正相结合,在动作标注中显式融入关于人类状态、紧迫性、可行性和潜在风险的情境判断。基于该监督,我们对代表性的多模态大语言模型(MLLMs)进行了基准测试,并提出了 MMC2Act 这一参考模型,它隐式地学习从多模态观察到具有认知依据的高层动作的映射。在多种模态设置、受试者不相交泛化、跨数据集迁移以及人类评估上的实验表明,通用 MLLM 难以从多模态线索中主动推理高层动作,而在 ProAction 上训练则显著提升了性能。
cs.RO / 86 / 2609.23487

Design and Control of a Cable-Driven Switchable Actuator with Torque/Tension Dual Modes for Exoskeletons

面向外骨骼的具有力矩/张力双模式的绳驱可切换执行器的设计与控制
Ji, YuanLong, Liu, Xu, Cai, Xinyuan, Ye, Qihan, Xie, Xiangyu, Jiang, Ruizhe, Xiang, Shuhan, Liu, Wenjing, Wang, Qijun, Chen, Yang, Yang, Xingbang
Abstract
Existing wearable exoskeleton architectures are typically constrained by a single mechanical output modality, providing either joint torque around an anatomical joint or linear traction along a limb-training-oriented direction, which limits adaptability to diverse training scenarios. This letter presents a cable-driven switchable actuator (CDSA) that can rapidly switch between torque and tension modes while centralizing all sensing and actuation components at the proximal drive unit. A Coupled Movable Pulley Mechanism (CMPM) provides tension amplification at the distal end-effector, while a bidirectional Cable-Driven Ratchet Mechanism (CDRM) enables mode switching and preload regulation. To eliminate the need for distal instrumentation, multi-source proximal sensors are integrated with a data-driven fusion model to estimate distal output forces. An adaptive dual-mode force control strategy based on iterative learning control (ILC) is further developed. Platform experiments demonstrate transmission efficiencies of $(92.4 \pm 2.0)\%$ and $(96.5 \pm 3.3)\%$ in the torque and tension modes, respectively, along with a tension amplification ratio of $2.77 \pm 0.10$ under tension mode. Tracking tests on simulated knee-joint gait trajectories and short-stroke tension profiles yield stable control, with RMSEs of $(4.52 \pm 0.51)\%$ and $(3.15 \pm 0.19)\%$ of the uncontrolled peak value, respectively. Finally, seated human-coupled experiments validate the system's controllable force generation in both joint-torque and linear-traction application modes.
Chinese Translation
现有的可穿戴外骨骼架构通常受限于单一的机械输出模式,即仅能提供围绕解剖关节的关节力矩或沿肢体训练方向的线性牵引力,这限制了对多样化训练场景的适应性。本文提出了一种绳驱可切换执行器(Cable-Driven Switchable Actuator, CDSA),能够在力矩模式与张力模式之间快速切换,同时将所有传感与驱动组件集中部署于近端驱动单元。所设计的耦合动滑轮机构(Coupled Movable Pulley Mechanism, CMPM)在远端末端执行器处提供张力放大,而双向绳驱棘轮机构(Cable-Driven Ratchet Mechanism, CDRM)实现模式切换与预紧力调节。为免除远端传感装置的需求,系统集成了多源近端传感器,并结合数据驱动的融合模型来估计远端输出力。此外,进一步提出了一种基于迭代学习控制(Iterative Learning Control, ILC)的自适应双模式力控制策略。平台实验表明,力矩模式与张力模式下的传动效率分别为 (92.4 ± 2.0)% 和 (96.5 ± 3.3)%,张力模式下的张力放大比为 2.77 ± 0.10。在模拟膝关节步态轨迹和短行程张力曲线的跟踪测试中,系统实现了稳定控制,未受控峰值对应的均方根误差(RMSE)分别为 (4.52 ± 0.51)% 和 (3.15 ± 0.19)%。最后,坐姿人机耦合实验验证了该系统在关节力矩与线性牵引两种应用模式下均可实现可控的力输出。
cs.RO / 87 / 2609.23488

FeasibleFlow: One-Step Joint Transport of Configuration Feasibility and Trajectories for End-to-End Driving

FeasibleFlow:面向端到端驾驶的构型可行性与轨迹一步式联合传输
Li, Xiang, Wang, Bikun, Xu, Qing, Wang, Jianjun
Abstract
End-to-end autonomous driving maps current observations directly to future trajectories, yet those trajectories must remain valid as the scene evolves. Future state modeling aims to address this temporal mismatch, but general representations often contain information unrelated to ego planning and affect trajectory generation only through auxiliary supervision, static conditioning, or proposal evaluation. We propose FeasibleFlow, a one-step end-to-end generative framework that jointly transports a configuration-space feasibility field and multimodal ego trajectories. Our Asymmetric Joint MeanFlow uses the pathwise Jacobian-vector product in the MeanFlow identity to incorporate field evolution into trajectory transport. Because safety feedback is sparser than progress feedback, we further introduce the Anchor-relative ranker (ARR) and Pareto-ReinFlow to balance safety and progress in candidate selection and generation, respectively. Experiments on the NAVSIM benchmark demonstrate the strong performance of FeasibleFlow and validate both the joint transport of feasibility and trajectories and the proposed safety-first mechanisms.
Chinese Translation
端到端自动驾驶将当前观测直接映射为未来轨迹,然而随着场景演化,这些轨迹必须保持有效。未来状态建模旨在解决这一时序失配问题,但通用的表征往往包含与自车规划无关的信息,且仅通过辅助监督、静态条件约束或候选轨迹评估来影响轨迹生成。我们提出 FeasibleFlow,这是一种一步式端到端生成框架,可对构型空间(configuration-space)可行性场与多模态自车轨迹进行联合传输。我们的非对称联合 MeanFlow(Asymmetric Joint MeanFlow)利用 MeanFlow 恒等式中的路径雅可比-向量积,将可行性场的演化融入轨迹传输过程。由于安全性反馈比行驶进度反馈更为稀疏,我们进一步引入锚点相对排序器(Anchor-relative ranker, ARR)和 Pareto-ReinFlow,分别在候选轨迹选择与生成过程中平衡安全性与行驶进度。在 NAVSIM 基准上的实验表明,FeasibleFlow 具有出色的性能,同时验证了可行性与轨迹联合传输机制以及所提出的安全优先机制的有效性。
cs.RO / 88 / 2609.23491

Elevator-VIGS: Separating Elevator Motion from Robot Motion in Visual-Inertial Gaussian Splatting SLAM

Elevator-VIGS:在视觉-惯性高斯泼溅SLAM中分离电梯运动与机器人运动
Zhou, Rui, Zhu, Zihan, Zhang, Wei, Luo, Zizhou, Haala, Norbert, Pollefeys, Marc
Abstract
We present Elevator-VIGS, a visual-inertial 3D Gaussian Splatting SLAM system that keeps tracking and mapping through elevator rides. Inside a moving elevator, the two sensors are in conflict. The camera sees only the robot's motion relative to the elevator, while the IMU senses that motion plus the elevator's motion relative to the world. This conflict is challenging for existing visual-inertial estimators. If vision dominates, the estimator tracks only the robot's motion within the elevator and misses the elevator's rise, and if the conflict remains, the estimator diverges. We observe that the conflict comes from forcing both observations into a single coordinate frame. We instead estimate the robot's pose in the elevator's coordinate frame, and the elevator's motion relative to the world as a per-keyframe transport state, the elevator's rise and vertical velocity, within dense visual-inertial bundle adjustment. Elevator-VIGS detects rides zero-shot with a vision-language model and a depth network, and constrains the transport state at the departure and the arrival. We record real-world and simulated elevator sequences. On these sequences, Elevator-VIGS achieves state-of-the-art tracking and rendering performance. On four elevator-free public benchmarks it keeps the state-of-the-art performance of VIGS-SLAM. Project page: https://ruizhou-cn.github.io/elevator-vigs/.
Chinese Translation
我们提出了Elevator-VIGS,一个能够在乘坐电梯过程中持续进行跟踪与建图的视觉-惯性三维高斯泼溅(3D Gaussian Splatting)SLAM系统。在运动的电梯内部,两种传感器之间存在冲突:相机只能观测到机器人相对于电梯的运动,而IMU感知到的则是该运动加上电梯相对于世界的运动。这种冲突对现有的视觉-惯性估计器而言极具挑战性。如果视觉占主导地位,估计器仅能跟踪机器人在电梯内部的运动,而遗漏电梯的上升运动;若冲突持续存在,估计器则会发散。我们观察到,这一冲突源于将两种观测强行纳入同一坐标系。因此,我们改为在电梯坐标系中估计机器人的位姿,并将电梯相对于世界的运动建模为逐关键帧的运输状态(即电梯的上升量和垂直速度),在稠密视觉-惯性光束法平差(bundle adjustment)中一同进行估计。Elevator-VIGS利用视觉-语言模型和深度网络实现零样本的电梯乘坐检测,并在电梯出发和到达时刻对运输状态施加约束。我们录制了真实世界和仿真环境中的电梯序列。在这些序列上,Elevator-VIGS实现了最先进的跟踪与渲染性能。在四个无电梯的公开基准数据集上,它保持了VIGS-SLAM的最先进性能。项目页面:https://ruizhou-cn.github.io/elevator-vigs/。
cs.RO / 89 / 2609.23504

Imagine then Verify: Affordance-Targeted Active Perception for Task-Oriented Grasping in Cluttered Scenes

先想象再验证:面向杂乱场景中任务导向抓取的功能可供性目标主动感知
Cui, Jingzhi, Liu, Xuefeng, Han, Feng, Liu, Xinyu, Chen, Wei, Niu, Jianwei
Abstract
Task-oriented grasping (TOG) requires robots to grasp functional parts of objects (e.g., the handle of a mug for pouring), yet these affordance regions are frequently occluded in cluttered scenes. Active perception via next-best-view (NBV) planning can resolve such occlusions by moving the camera for more informative observations. However, existing NBV methods typically optimize viewpoints for grasping the target object as a whole without distinguishing which part is task-relevant. A naive adaptation, fully scanning the target object before predicting the affordance, wastes most of the viewpoint budget on task-irrelevant surfaces (e.g., the mug body for pouring). To address this, we propose ATAP, an Affordance-Targeted Active Perception framework that shifts viewpoint planning from exhaustive target scanning to targeted affordance verification. ATAP hypothesizes the occluded target geometry via a generative shape prior and predicts the affordance distribution over the imagined complete surface. In cluttered scenes, severe occlusion can make the location of the hidden affordance ambiguous, leaving multiple locations plausible given the partial observation. ATAP therefore introduces an uncertainty-aware viewpoint planner that jointly optimizes expected entropy reduction over these competing hypotheses and expected affordance verification gain from real observations. This process iterates until the affordance is sufficiently verified for grasp execution. Experiments in simulation and real-world cluttered scenes show that ATAP substantially improves the functional grasp success rate over fixed-view TOG baselines, and outperforms reconstruction-based active perception with over 57% fewer NBV steps.
Chinese Translation
任务导向抓取(Task-oriented grasping, TOG)要求机器人抓取物体的功能性部位(例如,为了倒水而抓取马克杯的把手),然而这些功能可供性区域在杂乱场景中经常被遮挡。通过下一最优视角(next-best-view, NBV)规划实现的主动感知可以通过移动相机获取更具信息量的观测来消除此类遮挡。然而,现有的NBV方法通常针对将目标物体作为一个整体进行抓取来优化视角,而未区分哪个部位与任务相关。一种朴素的适应方法是在预测功能可供性之前对目标物体进行全面扫描,这会将大部分视角预算浪费在与任务无关的表面上(例如,倒水任务中的杯身)。为解决这一问题,我们提出了ATAP(Affordance-Targeted Active Perception,功能可供性目标主动感知)框架,该框架将视角规划从穷尽式目标扫描转变为针对性的功能可供性验证。ATAP通过生成式形状先验对被遮挡的目标几何结构进行假设,并在想象出的完整表面上预测功能可供性分布。在杂乱场景中,严重的遮挡可能使隐藏的功能可供性位置变得模糊,导致在仅有部分观测的情况下存在多个看似合理的位置。因此,ATAP引入了一种不确定性感知的视角规划器,该规划器联合优化对这些竞争性假设的预期熵减少量以及从真实观测中获得的预期功能可供性验证增益。这一过程不断迭代,直到功能可供性得到充分验证以执行抓取。在仿真和真实世界杂乱场景中的实验表明,ATAP相比固定视角的TOG基线方法大幅提升了功能性抓取的成功率,并且优于基于重建的主动感知方法,同时NBV步数减少了超过57%。
cs.RO / 90 / 2609.23554

PINGU: Extending Air-Bearing Spacecraft Emulators with Open-Source Actuators and Learned Control for Contact-Rich Proximity Operations

PINGU:通过开源执行器与学习控制扩展气浮航天器模拟平台,面向接触密集型近距离操作
Castan, Ricard Marsal I, Uchida, Akiyoshi, Arora, Aman, Lima, Pedro, El-Hariry, Matteo, Orsula, Anrej, Grella, Francesco, Richard, Antoine, Pradalier, Cedric, Olivarez-Mendez, Miguel A.
Abstract
Low-cost planar air-bearing testbeds have matured into a standard proxy for free-flying spacecraft GNC, but they remain largely thruster-only and are rarely equipped for contact-rich, inertia-coupled manipulation. Building on the open-source ATMOS testbed, we contribute a reaction wheel and two force/torque-sensed robotic arms (LEVION) with interchangeable end-effectors, integrated as first-class control actuators through a unified ROS 2 abstraction layer. On top of the software stack we build a reinforcement-learning training environment and digital twin, and a controller that exploits these added degrees of freedom, letting classical optimal controllers and learned policies be swapped on the same hardware without modification. We validate the integrated system, PINGU, across four benchmark tasks: point-to-pose navigation (classical LQR vs. sim-to-real PPO), dynamic disturbance rejection under arm-induced center-of-mass shifts, reaction-wheel momentum stabilization, and force-controlled docking. The results show that these additions extend an ATMOS-class emulator into the contact-rich regime and bridge classical optimal control and reinforcement learning on one reproducible platform.
Chinese Translation
低成本平面气浮试验台已发展成为自由飞行航天器制导、导航与控制(GNC)研究的标准替代平台,但它们大多仅配备推进器,很少能够进行接触密集、惯性耦合的操控任务。在开源ATMOS试验台的基础上,我们贡献了一个反作用轮和两个带可更换末端执行器的力/扭矩感知机械臂(LEVION),并通过统一的ROS 2抽象层将其作为一等控制执行器集成到系统中。在软件栈之上,我们构建了强化学习训练环境与数字孪生,以及一个利用这些新增自由度的控制器,使经典最优控制器与学习策略无需修改即可在同一硬件上互换运行。我们在四个基准任务上对集成系统PINGU进行了验证:位姿点导航(经典LQR对比仿真到真实迁移的PPO)、机械臂引起质心偏移下的动态扰动抑制、反作用轮动量稳定以及力控对接。结果表明,这些扩展将ATMOS级别的模拟平台拓展至接触密集领域,并在一个可复现的平台上架起了经典最优控制与强化学习之间的桥梁。
cs.RO / 91 / 2609.23578

AR-WAM: A Visual-Conditioned Agent-Ready World Action Model for Robotic Manipulation

AR-WAM:面向机器人操作的视觉条件化智能体就绪世界动作模型
Jiang, Yicheng, Gan, Zesen, Wang, Xiaobo, He, Tianlun, Zhao, Chenxu, Wu, Minghui, Wang, Xinyue, Wang, Jiaxu, He, Junhao, Wang, Jianan, Shao, Qiming
Abstract
As AI agents become increasingly capable, agent-driven robotic control is emerging as a compelling paradigm. However, prevailing vision-language-action (VLA) models and world action models (WAMs) still rely on natural-language instructions to specify manipulation tasks, an ill-suited interface for agent-driven control: referentially ambiguous, spatially imprecise, redundant with the agent's inherent language understanding, and entangling intent with execution. We present AR-WAM, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt (a bounding box of the target) denoting the interaction object and location, and a learnable operation token dictating the atomic skill to execute. Our compact 0.5B-parameter model, with a frozen pretrained visual encoder and no language encoder, predicts scene evolution within compact latent states while decoding actions, exposing the policy's intent through explicit, supervisable reasoning signals. A model-agnostic compatibility layer provides three primitives (detect, execute, and query) so that local VLMs or online agent APIs can drive the policy directly, with long-horizon memory and closed-loop error recovery delegated to the agent side. On RoboTwin 2.0, RMBench, and a real Astribot S1 dual-arm platform, AR-WAM matches the strongest baselines on standard manipulation (87.2% average success) and outperforms them on memory-dependent and real-robot long-horizon tasks, improving success rates by 5.9% and 36.7%, respectively, while maintaining the lowest inference latency (14.1 ms).
Chinese Translation
随着AI智能体能力的不断提升,由智能体驱动的机器人控制正成为一种极具吸引力的范式。然而,当前主流的视觉-语言-动作(VLA)模型和世界动作模型(WAM)仍依赖自然语言指令来指定操作任务,而这一接口并不适合智能体驱动的控制:其指代模糊、空间描述不精确、与智能体固有的语言理解能力冗余,且将意图与执行耦合在一起。我们提出了AR-WAM,一个视觉条件化、面向智能体的世界动作模型,它以两种互补的条件取代语言:一是视觉定位提示(目标的边界框),用于表示交互对象及其位置;二是可学习的操作令牌(token),用于指定待执行的原子技能。我们这个紧凑的0.5B参数模型采用冻结的预训练视觉编码器且不含语言编码器,在紧凑的潜在状态中预测场景演化并解码动作,通过显式、可监督的推理信号展现策略的意图。一个模型无关的兼容层提供三种原语(检测、执行和查询),使本地VLM或在线智能体API能够直接驱动该策略,而长时程记忆与闭环错误恢复则交由智能体侧完成。在RoboTwin 2.0、RMBench以及真实的Astribot S1双臂平台上,AR-WAM在标准操作任务上与最强基线持平(平均成功率87.2%),并在依赖记忆的任务和真实机器人长时程任务上超越基线,成功率分别提升5.9%和36.7%,同时保持最低的推理延迟(14.1毫秒)。
cs.RO / 92 / 2609.23580

TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation

TaskAnchor:在反应式视觉-语言-动作模型中为长时程操作引入任务状态锚定
Liu, Hengyan, Zhou, Wenlve, Yue, Bo, Su, Yongyi, Wang, Ruixiang, Zhang, Zhanqi, Lu, Dekun, Gao, Wei, Xing, Xiaofen, Jia, Kui
Abstract
Reactive vision--language--action (VLA) models struggle with long-horizon manipulation when visually similar observations can correspond to different actions depending on the task stage or interaction history. We refer to this ambiguity as task-state aliasing and introduce TaskAnchor, a lightweight adapter that grounds pretrained VLAs in execution history. TaskAnchor combines history-conditioned visual refinement with a milestone-supervised task-state coordinate, a scalar representing the semantic stage of execution. These signals are injected through the native visual and language interfaces, respectively, without introducing an explicit planner or modifying the action-generation mechanism. On RMBench, TaskAnchor achieves approximately 4.9--5.5$\times$ the average success rates of the published $\pi_{0.5}$ and X-VLA baselines, with consistent gains on RoboMemArena and real robots. The added latency is only 2.08\,ms per action chunk for $\pi_{0.5}$.
Chinese Translation
当视觉上相似的观测对应于依赖于任务阶段或交互历史的不同动作时,反应式视觉-语言-动作(VLA)模型在长时程操作中表现不佳。我们将这种歧义称为任务状态混叠(task-state aliasing),并提出 TaskAnchor——一种轻量级适配器,可将预训练的 VLA 模型锚定于执行历史。TaskAnchor 将基于历史的视觉细化与里程碑监督的任务状态坐标(一个表示执行语义阶段的标量)相结合。这些信号分别通过模型原生的视觉和语言接口注入,无需引入显式规划器,也无需修改动作生成机制。在 RMBench 上,TaskAnchor 的平均成功率约为已发表的 $\pi_{0.5}$ 和 X-VLA 基线的 4.9–5.5 倍,并在 RoboMemArena 和真实机器人上取得了一致的性能提升。对于 $\pi_{0.5}$,每个动作块仅增加 2.08 毫秒的延迟。
cs.RO / 93 / 2609.23610

PRIMO: Prior-Informed Odometry from Human-Motion Tracking for Humanoid Robots

PRIMO:基于人体运动跟踪的先验信息人形机器人里程计
Han, Xu, Li, Angsong, Zhang, Shaopeng, Li, Enyu, Lin, Peiwen, Wang, Chuang, Zhuang, Yuan, Lan, Haiyu
Abstract
Simulation-trained humanoid proprioceptive odometry faces two transfer challenges: training trajectories generated by specific robot control policies intended for deployment cover only a limited range of motions, while sim-to-real mismatch can make unconstrained predictions unreliable. We address both with Prior-Informed Odometry from Human-Motion Tracking (PRIMO). On the data side, we generate odometry supervision by having the humanoid track diverse retargeted human motions in simulation, decoupling supervision from the deployment policies and broadening the training motion distribution. On the model side, a Prior-Informed estimator uses physics- and symmetry-informed priors to structure velocity and rotation prediction and a coarse raw-context pathway to preserve sensor context alongside encoded features, thereby strengthening sim-to-real generalization. Under a unified real-robot protocol, PRIMO reduces mean error by 31.6%-61.7% relative to the strongest evaluated external baseline in each domain-metric comparison. Across two locomotion-policy revisions, policy specialists exhibit symmetric crossover, whereas Tracking-Locomotion training reduces mean opposite-policy simulation error by 86.8%-94.6%. On real dynamic motion, Tracking-Locomotion training reduces mean error by 69.2%-81.7% relative to training on the union of both deployment policies. Across the tested motion compositions, the Prior-Informed estimator consistently lowers mean trajectory errors relative to its Unconstrained counterpart in both simulation and real-robot evaluation. Code is available at https://github.com/Agibot-Spatial-Intelligence/PRIMO.
Chinese Translation
基于仿真训练的人形机器人本体感知里程计面临两个迁移挑战:由特定机器人控制策略生成的训练轨迹(该策略旨在部署使用)仅覆盖有限的运动范围,而仿真到现实的差距可能使无约束的预测不可靠。我们通过基于人体运动跟踪的先验信息里程计(PRIMO,Prior-Informed Odometry from Human-Motion Tracking)同时解决这两个问题。在数据方面,我们让人形机器人在仿真中跟踪多种重定向的人体运动来生成里程计监督信号,从而将监督信号与部署策略解耦,并拓宽了训练运动的分布。在模型方面,先验信息估计器(Prior-Informed estimator)利用物理和对称性先验来结构化速度和旋转预测,并通过粗粒度原始上下文通路在编码特征之外保留传感器上下文,从而增强仿真到现实的泛化能力。在统一的真机评测协议下,在每个领域-指标对比中,PRIMO 相对于所评估的最强外部基线将平均误差降低了 31.6%–61.7%。在两个运动控制策略版本中,策略专家模型表现出对称的交叉现象,而跟踪-运动(Tracking-Locomotion)训练将跨策略的平均仿真误差降低了 86.8%–94.6%。在真实动态运动上,相对于在两种部署策略联合数据上的训练,跟踪-运动训练将平均误差降低了 69.2%–81.7%。在所测试的各种运动组合中,先验信息估计器在仿真和真机评测中均始终如一地降低了相对于其无约束对应模型的平均轨迹误差。代码可在 https://github.com/Agibot-Spatial-Intelligence/PRIMO 获取。
cs.RO / 94 / 2609.23614

CompVLA: A Variable Compliance Vision-Language-Action Model for Contact-rich Manipulation

CompVLA:一种面向接触密集型操作的柔顺可变视觉-语言-动作模型
Kim, Jongmin, Ha, Junsu, Park, Che-Sang, Song, Minchang, Jeong, Hyeokju, Hwang, Himchan, Fu, Jianlong, Park, Frank C.
Abstract
Contact-rich manipulation, requiring robots to regulate not only motion but also how they yield to external forces, has emerged as the next frontier for Vision-Language-Action (VLA) models. However, existing VLAs output purely kinematic commands, degrading performance on real-world contact-rich tasks. In this paper, we introduce CompVLA, a unified VLA framework that jointly predicts motion and stiffness matrix from RGB and language inputs. Our approach augments the conventional architecture with a dedicated Compliance Expert, which outputs time-varying stiffness and virtual displacement profiles executed via geometric impedance control. We demonstrate that CompVLA achieves the highest average success rate across diverse contact-rich tasks, outperforming both vanilla and compliance-aware VLA baselines, with ablations confirming each component is essential.
Chinese Translation
接触密集型操作要求机器人不仅调控运动,还要调控其对外部力的顺应方式,这已成为视觉-语言-动作(Vision-Language-Action, VLA)模型的下一个前沿方向。然而,现有的VLA模型仅输出纯运动学指令,导致其在真实世界接触密集型任务中的性能下降。本文提出CompVLA,一个统一的VLA框架,能够从RGB图像和语言输入中联合预测运动与刚度矩阵。我们的方法在传统架构上增加了一个专用的柔顺专家(Compliance Expert)模块,用于输出时变刚度和虚拟位移曲线,并通过几何阻抗控制加以执行。实验表明,CompVLA在多样化的接触密集型任务中取得了最高的平均成功率,优于原始VLA和柔顺感知的VLA基线方法,消融实验也证实了每个组件都是必不可少的。
cs.RO / 95 / 2609.23643

Spiking Neural Network Actor-Critic Proximal Policy Optimization Control for Autonomous UAV Navigation Through Constrained Openings in Civil Infrastructure and Buildings

基于脉冲神经网络的Actor-Critic近端策略优化算法在民用基础设施与建筑受限开口处无人机自主导航中的应用
Walugembe, Francis Noah, Wielgosz, Maciej, Goričan, Tomaž, Mertik, Matej
Abstract
Autonomous navigation of unmanned aerial vehicles in constrained three-dimensional environments has been a challenge in the robotics domain. The application of autonomous unmanned aerial vehicles in civil infrastructure inspection involves the use of such vehicles in bridge inspection, tunnel inspection, and structural inspection. The use of deep reinforcement learning in the autonomous navigation of unmanned aerial vehicles has been successful in constrained environments. However, the computational cost of the algorithm limits the application of the algorithm in the autonomous navigation of unmanned aerial vehicles. This paper proposes the use of the spiking neural network-based Proximal Policy Optimization algorithm in the autonomous navigation of unmanned aerial vehicles in constrained sequential environments. The proposed algorithm integrates the use of spike-based actor-critic reinforcement learning with the Proximal Policy Optimization algorithm. The proposed algorithm uses the stochastic Gaussian policy in the autonomous navigation of unmanned aerial vehicles. The proposed algorithm was implemented in the autonomous navigation of unmanned aerial vehicles in constrained 3D environments. The proposed algorithm was successful in completing 1913 episodes out of more than 3000. The proposed algorithm was successful in passing an average of 2.10 windows per episode. The proposed algorithm was successful in achieving a success rate of 63.77%. The proposed algorithm was successful in achieving success rates of more than 90% in the later stages of the algorithm.
Chinese Translation
无人机在受限三维环境中的自主导航一直是机器人领域的一项挑战。自主无人机在民用基础设施检测中的应用包括桥梁检测、隧道检测和结构检测。深度强化学习在无人机自主导航中的应用已在受限环境中取得成功。然而,该算法的计算成本限制了其在无人机自主导航中的应用。本文提出将基于脉冲神经网络(Spiking Neural Network)的近端策略优化(Proximal Policy Optimization, PPO)算法用于无人机在受限序列环境中的自主导航。所提出的算法将基于脉冲的Actor-Critic强化学习与近端策略优化算法相融合,并在无人机自主导航中采用随机高斯策略。该算法在受限三维环境下的无人机自主导航中得以实现,在超过3000个回合中成功完成了1913个回合,平均每个回合成功通过2.10个窗口,成功率达到63.77%,且在算法后期阶段成功率超过90%。
cs.RO / 96 / 2609.23650

Beyond Appearance Shifts: Task-Semantic Action Calibration for VLA Models

超越外观变化:面向VLA模型的任务-语义动作校准
Liu, Shuaijun, You, Feiyang, Wu, Chengyu, Hao, Shuyang, Zhang, Chenglong, Cai, Jingyao, Chen, Xingwei, Su, Ningxin
Abstract
Vision-language-action (VLA) models have achieved strong performance in embodied manipulation, but still lack a clear mechanism to balance behavioral stability with task-semantic sensitivity. We identify two complementary failure modes. Under task-preserving changes, where task semantics remain unchanged but scene appearance varies (e.g., style, illumination, clutter, or paraphrasing), policies often exhibit unnecessary action drift. Conversely, under semantic-breaking changes, where key task semantics such as the target object or constraint are altered, policies frequently fail to produce sufficiently distinct behaviors and instead follow the original trajectory. To address this gap, we propose BAS-VLA, a task-semantic action calibration framework built on top of a frozen base VLA. BAS-VLA adopts a breaking-centered calibration core as the default path, and introduces a selective evidence-gated preserving auxiliary that activates only when nuisance variation is detected while task semantics remain consistent. On the OpenPI-pi0.5 / LIBERO-Object Milk-Swap benchmark, BAS-VLA maintains high success on clean (98.0%) and semantics-preserving conditions (97.5%), while reducing clean-criterion success to 0.0% under deliberate target-object swaps, demonstrating strong stale-task suppression and task-semantic separation. On validated style-preserving shifts, it improves success from 42% to 70% without degrading clean performance. These results highlight that reliable VLA behavior requires moving beyond appearance robustness toward explicit task-semantic action calibration.
Chinese Translation
视觉-语言-动作(Vision-Language-Action,VLA)模型在具身操作中已取得优异表现,但仍缺乏一种在行为稳定性与任务-语义敏感性之间进行平衡的明确机制。我们识别出两种互补的失败模式。在任务保持型变化下,即任务语义不变而场景外观发生变化(如风格、光照、杂物或语句改述)时,策略往往表现出不必要的动作漂移。相反,在语义破坏型变化下,即目标任务对象或约束等关键任务语义被改变时,策略常常无法产生足够差异化的行为,而是沿用原始轨迹。为弥合这一缺口,我们提出了BAS-VLA,一个构建于冻结基础VLA之上的任务-语义动作校准框架。BAS-VLA采用以破坏为中心的校准核心作为默认路径,并引入一个可选择的、基于证据门控的保持辅助模块,该模块仅在被检测到干扰性变化且任务语义保持一致时激活。在OpenPI-pi0.5 / LIBERO-Object Milk-Swap基准上,BAS-VLA在干净条件(98.0%)和语义保持条件(97.5%)下均保持高成功率,同时在故意交换目标物体的情况下将干净标准成功率降至0.0%,展现出强大的过时任务抑制能力和任务-语义区分能力。在经过验证的风格保持型变化上,其成功率从42%提升至70%,且不损害干净条件下的性能。这些结果表明,可靠的VLA行为需要超越外观鲁棒性,走向显式的任务-语义动作校准。
cs.RO / 97 / 2609.23656

WOLF: World Model Guided LiDAR Exploration with Predictive Frontiers

WOLF:基于世界模型引导与预测性边界的LiDAR自主探索
Tian, Yuyang, Yang, Penghui, Wu, Pengyuan, Yang, Haoran, Li, Chenhui, Han, Pengfei, Wang, Dong, Wang, Zhigang, Zhao, Bin, Li, Xuelong
Abstract
LiDAR-based unmanned aerial vehicle (UAV) exploration builds maps by continually selecting where to observe next. However, decisions based on the measured map provide limited foresight into spatial continuations behind occlusions, leaving potentially informative directions unrecognized. We present WOLF, a world-model-guided framework that predicts future observations to enhance autonomous exploration. In the training stage, a recurrent world model learns observation dynamics from exploration trajectories, with recurrent memory retaining the spatial context needed to interpret partial observations across successive views. Building on this context, the model combines observation history with candidate motions during exploration to predict local occupancy and visibility. To guide further sensing, a predictive frontier generation mechanism then aligns and fuses these predictions using confidence, branch agreement, and observation quality to identify promising regions. The resulting predictive frontiers join measured ones to guide geometric viewpoint selection and trajectory generation, while new scans update subsequent predictions. In simulations, our method reduces mean terminal time by 10.9% relative to EPIC in Garage at comparable coverage and increases mean coverage from 42.12% to 98.35% in Tunnel. Real-world experiments further demonstrate onboard deployment of the learned model for online inference during physical flight.
Chinese Translation
基于LiDAR的无人机(UAV)自主探索通过不断选择下一个观测位置来构建地图。然而,基于已测量地图的决策对遮挡背后的空间延续缺乏预见能力,导致潜在具有信息价值的方向未被识别。我们提出了WOLF,一种利用世界模型预测未来观测以增强自主探索的框架。在训练阶段,一个循环世界模型从探索轨迹中学习观测动态,其循环记忆保留了在连续视角之间解释部分观测所需的空间上下文。基于该上下文,模型在探索过程中将观测历史与候选动作相结合,预测局部占据情况和可见性。为了引导进一步的感知,一种预测性边界生成机制随后利用置信度、分支一致性和观测质量对这些预测进行对齐与融合,以识别有前景的区域。所产生的预测性边界与实测边界共同引导几何视点选择与轨迹生成,同时新的扫描数据会更新后续的预测。在仿真实验中,在覆盖率相当的情况下,我们的方法在Garage场景中相对于EPIC将平均终止时间减少了10.9%,并在Tunnel场景中将平均覆盖率从42.12%提升至98.35%。真实世界实验进一步验证了所学模型在实际飞行中可部署于机载设备进行在线推理。
cs.RO / 98 / 2609.23666

UniPoint: Unified Point-Level Sensor Fusion for Humanoid Locomotion Across Challenging Terrains

UniPoint:面向人形机器人跨越复杂地形的统一点级传感器融合方法
Li, Sicen, Chu, Zhen, Li, Chao, Zhu, Qiuguo, Wu, Jun
Abstract
Open-world deployment requires humanoid robots to cross highly heterogeneous terrain safely, with perception that simultaneously provides wide coverage, local accuracy, and redundancy against sensor failure. Existing approaches struggle to satisfy all three: one forward depth camera or nearby height sampling covers too little; odometry-corrected elevation maps drift under aggressive motion and miss thin vertical structures; image-level encoding costs grow with camera count. We present UniPoint, a humanoid whole-body locomotion framework built on multi-source point-level sensor fusion. Measurements from a 360{\deg} light detection and ranging (LiDAR) sensor and two depth cameras are early-fused into one base-frame point set. Voxelization resamples it to a fixed number of tokens encoded by linear self-attention and proprioception-queried cross-attention, decoupling forward cost from sensor count. The point set retains standing thin barriers; a single-modality failure removes only part of the tokens, so the policy degrades gracefully. A single training run with terrain-aware rewards, perception-degradation injection, and domain randomization produces one policy for all eight terrain types, deployed on an onboard RK3588 without fine-tuning. On a DR02 humanoid, 20 trials at each of nine real-world settings over seven terrain types validate the policy on 70-cm-high platforms, 100-cm gaps, thin barriers, and sparse or narrow footholds; it also generalizes zero-shot outdoors.
Chinese Translation
开放世界部署要求人形机器人能够安全地穿越高度异构的地形,其感知系统需同时具备宽广的覆盖范围、局部精度以及应对传感器故障的冗余性。现有方法难以同时满足这三点:单个前向深度相机或近处高度采样覆盖范围过小;经里程计校正的高程图在剧烈运动下会产生漂移,且无法捕捉细薄的垂直结构;图像级编码的计算开销随相机数量增加而增长。我们提出了UniPoint,一个基于多源点级传感器融合的人形机器人全身运动控制框架。来自360度激光雷达(LiDAR)传感器和两个深度相机的测量数据被早期融合为机器人基坐标系下的统一点集。体素化将其重采样为固定数量的token,通过线性自注意力(linear self-attention)编码,并经本体感觉查询的交叉注意力(cross-attention)处理,使前向计算开销与传感器数量解耦。该点集能够保留细薄的站立式障碍物;单一模态失效仅会损失部分token,因此策略可平滑退化。通过单次训练,结合地形感知奖励、感知退化注入和域随机化,生成适用于全部八种地形类型的单一策略,并可在机载RK3588上无需微调直接部署。在DR02人形机器人上,我们在七种地形类型的九个真实场景中各进行20次试验,验证了该策略在70厘米高平台、100厘米间隙、细薄障碍物以及稀疏或狭窄落脚点上的表现;该策略还能在户外实现零样本泛化。
cs.RO / 99 / 2609.23731

Marginal Calibration Does Not Compose: Hidden Dependence in Modular Robot Navigation

边际校准不可复合:模块化机器人导航中的隐含依赖性
Baral, Rista
Abstract
Robotic systems are typically composed of multiple independently developed modules that work together to perceive, predict, and act in the environment. Although each module may perform reliably in isolation, composing them does not necessarily preserve uncertainty calibration at the system level. In this work, we show that well-calibrated component interfaces do not necessarily produce calibrated downstream behavior after composition. Using a moving-obstacle prediction pipeline, we demonstrate that position and velocity estimators can each appear well calibrated individually, yet differences in how their error are correlated lead to substantially different estimates of future-state uncertainty. Consequently, assuming independence can make the system either overly confident or unnecessarily conservative, directly influencing downstream planning decisions and safety. Through simulations, we show that modeling the joint covariance restores downstream calibration and improves system performance, whereas dependence-robust uncertainty bounds enhance safety at the cost of increased conservatism. Our findings reveal a fundamental limitation of independently validating robotic modules and highlight the need for interfaces that communicate dependence information or support direct system-level calibration.
Chinese Translation
机器人系统通常由多个独立开发的模块组成,这些模块协同工作以感知环境、进行预测并执行动作。尽管每个模块在单独运行时可能表现可靠,但将它们组合起来并不一定能在系统层面保持不确定性校准。在本工作中,我们表明校准良好的组件接口在复合之后并不一定产生校准良好的下游行为。通过一个移动障碍物预测流水线,我们证明位置估计器和速度估计器各自单独看都校准良好,但其误差相关性方式的差异会导致对未来状态不确定性的估计存在显著不同。因此,假设独立性可能使系统要么过度自信,要么过于保守,直接影响下游规划决策和安全性。通过仿真实验,我们表明对联合协方差进行建模可以恢复下游校准并提升系统性能,而依赖性鲁棒的不确定性界则以增加保守性为代价提升安全性。我们的发现揭示了独立验证机器人模块的一个根本性局限,并强调了需要能够传递依赖性信息或支持直接系统级校准的接口。
cs.RO / 100 / 2609.23745

FlockDiffusion: Assignment-Conditioned Diffusion for Multi-Drone Task Allocation and Completion

FlockDiffusion:面向多无人机任务分配与执行的分配条件扩散模型
Zhura, Iana, Akopyan, Satenik, Khan, Roohan Ahmed, Cabrera, Miguel Altamirano, Fedoseev, Aleksey, Tsetserukou, Dzmitry
Abstract
Autonomous multi-drone navigation requires fleets to service distributed objectives in cluttered environments under tight computational budgets. Efficient coordination depends on task bundling, where each drone visits multiple objectives along its route. Separate solvers for cost estimation, assignment, and execution incur redundant graph search and produce long, abrupt paths. We propose FlockDiffusion, a learned framework combining a scene graph encoder, an explicit allocation head, an assignment conditioned diffusion transformer, and a closed form trajectory decoder. An autoregressive teacher provides offline supervision for parallel fleet trajectory generation. PyBullet ablations show that bundling increases task completion from 50% to 100%, while our complete teacher further reduces route cost by 8.4% relative to MAGNNET with bundling. In the optimized scalability benchmark, evaluated on 100 scenes per density with ten drones, FlockDiffusion achieves 6.2 to 7.6 times faster inference and approximately 37% shorter routes than the classical pipeline. As nominal task counts increase from 20 to 40, latency rises from 7.8 to 11.1 ms, compared with 48.0 to 75.8 ms for the baseline. In a separate evaluation across five Gazebo environments, FlockDiffusion achieves 100% planner coverage and reduces planned route cost by 15.4% relative to the baseline with bundling. These results demonstrate efficient planning under increasing task density in configurations that are demanding to reproduce with physical drone fleets.
Chinese Translation
自主多无人机导航要求机群在计算资源受限的条件下,于复杂环境中服务分布式目标。高效协同依赖于任务捆绑,即每架无人机沿其航线访问多个目标。将代价估计、任务分配与执行分离的求解器会产生冗余的图搜索,并生成冗长且突变频繁的路径。我们提出FlockDiffusion,这是一个学习型框架,结合了场景图编码器、显式分配头、分配条件扩散Transformer(assignment-conditioned diffusion transformer)以及闭式轨迹解码器。自回归教师模型为并行机群轨迹生成提供离线监督。PyBullet消融实验表明,任务捆绑可将任务完成率从50%提升至100%,而我们的完整教师模型相较于采用捆绑的MAGNNET可进一步将航线代价降低8.4%。在可扩展性优化基准测试中(每种密度评估100个场景,配备十架无人机),FlockDiffusion相较传统流水线实现了6.2至7.6倍的推理加速,且航线长度缩短约37%。当标称任务数量从20增加到40时,延迟从7.8毫秒升至11.1毫秒,而基线方法则需48.0至75.8毫秒。在另一项涵盖五个Gazebo环境的评估中,FlockDiffusion实现了100%的规划器覆盖率,且相较于采用捆绑的基线方法将规划航线代价降低15.4%。这些结果表明,在物理无人机机群难以复现的配置中,本方法能够在任务密度不断增加的情况下实现高效规划。
cs.RO / 101 / 2609.23755

EgoWild2Dex: Learning Dexterous Robotic Manipulation from In-the-Wild Human Experience

EgoWild2Dex:从真实场景人类经验中学习灵巧机器人操作
Lin, Kunyang, Wen, Xutao, Lin, Jingxi, Lin, Lanyong, Liu, Jiaming, Yang, Tianshuo, Chen, Xianchi, Han, Yue, Li, Yiduo, Zhang, Zhanpeng, Luo, Ping
Abstract
Egocentric human data provide a principled source of supervision for learning dexterous robot manipulation. Unlike prior approaches that often collect such data in constrained or specially constructed environments, we collect in-the-wild egocentric demonstrations in real-world settings, including homes, factories, and pharmacies, etc., where people perform their ordinary tasks while wearing head-mounted cameras. This collection protocol captures diverse workflows and hand-object interactions across long-tailed object and skill distributions, but also yields visually challenging observations due to scene clutter and head-motion-induced viewpoint changes (a mean cumulative rotation of $15.93^{\circ}$/s). To address these issues, we introduce EgoWild2Dex, which transfers in-the-wild ego-human experience to dual-arm robots with dexterous hands by jointly aligning unstable egocentric views and human motions with robot observations and actions, respectively. This work offers three benefits. First, we introduce GeoFormer, a differentiable geometric transformer that warps noisy human observations toward robot observations. Second, we design a human-robot training scheme to bridge the embodiment gap, enabling high task success with limited robot supervision. Third, we release EgoWild, a 538.9-hour in-the-wild egocentric human dataset comprising 179,049 episodes, 125,961 unique task descriptions, and 1,282 object categories. On real robots, EgoWild2Dex achieves an average success rate of 96.7% across three long-horizon bimanual dexterous manipulation tasks and an average object-level zero-shot success rate of 33.3%. The data, models, and code will be released.
Chinese Translation
以自我为中心(Egocentric)的人类数据为学习灵巧机器人操作提供了一种有原则的监督来源。与以往通常在受限或专门构建的环境中收集此类数据的方法不同,我们在真实世界场景中收集真实场景(in-the-wild)的自我中心视角演示数据,包括家庭、工厂、药房等,人们在执行日常任务时佩戴头戴式相机。这种数据采集方式能够捕捉多样的工作流程以及长尾物体和技能分布上的手-物体交互,但由于场景杂乱和头部运动引起的视角变化(平均累积旋转速率为 $15.93^{\circ}$/s),也产生了视觉上具有挑战性的观测数据。为解决这些问题,我们提出了 EgoWild2Dex,通过将不稳定的自我中心视角和人类动作分别与机器人观测和机器人动作进行联合对齐,将真实场景的人类自我中心经验迁移到配备灵巧手的双臂机器人上。本工作具有三方面贡献:第一,我们提出 GeoFormer,一种可微分的几何 Transformer,用于将含噪声的人类观测数据向机器人观测数据进行扭曲对齐;第二,我们设计了一种人机协同训练方案,以弥合本体(embodiment)差异,在有限的机器人监督下实现较高的任务成功率;第三,我们发布了 EgoWild,一个时长 538.9 小时的真实场景自我中心人类数据集,包含 179,049 条片段、125,961 条独特的任务描述以及 1,282 个物体类别。在真实机器人上,EgoWild2Dex 在三个长时序双手灵巧操作任务中取得了平均 96.7% 的成功率,物体级零样本(zero-shot)平均成功率为 33.3%。相关数据、模型和代码将被公开发布。
cs.RO / 102 / 2609.23784

PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing

PackLab:用于机器人装箱任务中多模态大语言模型开发、训练与评估的综合框架
Zhou, Donghao, Pan, Jia-Hui, Zhang, Fan, Bu, Xingyuan, Li, Shilong, Gao, Xiaojie, Liu, Yun-Hui, Fu, Chi-Wing, Heng, Pheng-Ann
Abstract
Robotic bin packing requires long-horizon sequential decision-making, as each object placement affects the available space for subsequent packing. Existing methods primarily rely on hand-crafted geometric heuristics that optimize predefined objectives or reinforcement learning policies learned through trial and error over predefined training configurations. Despite recent advances in multimodal large language models (MLLMs) for this task, their potential for closed-loop sequential decisions across heterogeneous packing configurations remains underexplored. To address this gap, we introduce PackLab, a comprehensive framework for developing, training, and evaluating MLLMs for closed-loop robotic bin packing. PackLab-Suite provides a physics-based simulation platform for scalable generation of diverse training packing trajectories and evaluation of their physical outcomes. PackLab-VLM is a packing-specialized MLLM that understands the evolving object and container states to jointly select objects and predict placements in a closed-loop manner. PackLab-Bench provides standardized packing scenarios at multiple difficulty levels for systematic evaluation. Extensive experiments demonstrate that, on average, PackLab-VLM outperforms conventional packing heuristics, traditional reinforcement learning methods, and general-purpose MLLMs across object sets and container configurations, highlighting the potential of MLLMs for long-horizon robotic packing. The code, model, dataset, and benchmark are available at https://github.com/Correr-Zhou/PackLab .
Chinese Translation
机器人装箱(bin packing)需要长时程的序贯决策,因为每个物体的放置都会影响后续装箱的可用空间。现有方法主要依赖手工设计的几何启发式规则来优化预定义目标,或依赖在预定义训练配置上通过反复试错学习到的强化学习策略。尽管多模态大语言模型(MLLM)在该任务上取得了近期进展,但其在异构装箱配置下进行闭环序贯决策的潜力仍未得到充分探索。为填补这一空白,我们提出了PackLab,一个用于开发、训练和评估面向闭环机器人装箱任务的MLLM的综合框架。PackLab-Suite提供了一个基于物理的仿真平台,用于可扩展地生成多样化的训练装箱轨迹并评估其物理结果。PackLab-VLM是一个面向装箱任务的专用MLLM,能够理解不断变化的物体与容器状态,以闭环方式联合选择物体并预测放置位置。PackLab-Bench提供了多个难度级别的标准化装箱场景,用于系统性评估。大量实验表明,PackLab-VLM在不同物体集合与容器配置上的平均表现优于传统装箱启发式方法、传统强化学习方法以及通用MLLM,凸显了MLLM在长时程机器人装箱任务中的潜力。代码、模型、数据集和基准测试已发布于 https://github.com/Correr-Zhou/PackLab 。
cs.RO / 103 / 2609.23792

Risk-Aware Motion Planning and Control under Unknown Dynamics with Hybrid Observations

未知动力学与混合观测下的风险感知运动规划与控制
Zhang, Zhiquan, Ornik, Melkior
Abstract
We consider robotic motion planning and control under unknown dynamics with hybrid state observations, where state measurements are available only in parts of the state space. Existing work combines system identification, predicted reachability, graph search and controller synthesis in a hierarchical framework using local affine approximated models over polytopic state space partitioning, but requires state observations for identification and feedback control. Based on this framework, we address blind regions by selecting nominal dynamics and precomputing open-loop control sequences before observation is lost. Since the true dynamics may differ from the selected nominal model, the robot may exit a blind polytope through an unintended facet. We quantify this transition risk and incorporate the possible outcomes into a stochastic transition system. The high-level planning problem is formulated as a stochastic shortest path problem, whose policy guides controller synthesis. A case study demonstrates that the method guides the robot from an initial state to a target while balancing route efficiency and the risks associated with traversing blind regions.
Chinese Translation
我们研究动力学未知且状态观测为混合形式(即状态测量仅在部分状态空间中可用)的机器人运动规划与控制问题。已有工作将系统辨识、可达性预测、图搜索和控制器综合结合在一个分层框架中,在多面体状态空间划分上使用局部仿射近似模型,但该框架需要状态观测用于辨识和反馈控制。基于该框架,我们通过在失去观测之前选择标称动力学并预计算开环控制序列来应对盲区问题。由于真实动力学可能与所选标称模型不同,机器人可能会经由非预期的面离开盲多面体。我们对这种转移风险进行量化,并将可能的结果纳入一个随机转移系统中。高层规划问题被表述为随机最短路径问题,其策略用于指导控制器综合。案例研究表明,该方法能够引导机器人从初始状态到达目标,同时兼顾路径效率与穿越盲区所带来的风险。
cs.RO / 104 / 2609.23800

ContactDP: Contact-Guided Diffusion Policy for Tight Insertion Tasks

ContactDP:面向紧凑插装任务的接触引导扩散策略
Xing, Chengyi, Yao, Shaoxiong, Romeres, Diego, Jha, Devesh K.
Abstract
High-precision connector insertion remains challenging for robotic systems due to tight mechanical tolerances, partial observability during contact, and multimodal uncertainty arising from occlusion and contact ambiguity. Successful insertion requires closed-loop contact guidance that continuously integrates global alignment cues with local contact feedback to produce stable corrective actions under interaction. In this work, we present ContactDP (Contact-Guided Diffusion Policy for Tight Insertion Tasks), a multimodal diffusion-policy framework for contact-rich insertion. ContactDP jointly integrates wrist RGB observations, fingertip tactile sensing, and wrist-mounted force-torque measurements to infer contact state and generate temporally consistent corrective motions during insertion. To ensure stable execution under contact, the learned policy operates together with a hybrid position-force controller that provides compliant low-level interaction. We evaluate our approach on a suite of industrial-grade connector insertion tasks with varying connector geometries, grasp conditions, and initial misalignment. Across all tasks, ContactDP significantly outperforms vision-only diffusion policies for performance, reliability and generalization.
Chinese Translation
由于机械公差极为紧凑、接触过程中部分可观测性受限,以及遮挡和接触模糊引发的多模态不确定性,高精度连接器插装对机器人系统而言仍具挑战性。成功的插装需要闭环接触引导,持续融合全局对齐线索与局部接触反馈,以在交互过程中产生稳定的纠正动作。在本工作中,我们提出了ContactDP(面向紧凑插装任务的接触引导扩散策略),这是一个面向接触密集型插装的多模态扩散策略框架。ContactDP联合整合腕部RGB观测、指尖触觉感知以及腕部安装的力/力矩测量,以推断接触状态并生成插装过程中时间上一致的纠正动作。为确保接触下的稳定执行,所学习的策略与提供柔顺底层交互的混合位置-力控制器协同运行。我们在一系列具有不同连接器几何形状、抓取条件和初始偏差的工业级连接器插装任务上评估了该方法。在所有任务中,ContactDP在性能、可靠性和泛化能力方面均显著优于仅基于视觉的扩散策略。
cs.RO / 105 / 2609.23841

Structured World-State Reasoning for Agentic Robotic Search

面向智能体机器人搜索的结构化世界状态推理
Holt, Finley R., Pabon, Luis A., Alora, John Irvin, Frey, Jonas, Pavone, Marco
Abstract
Long-horizon robotic search must resolve natural language against heterogeneous, incomplete, and often ambiguous evidence: textual information, prior maps, and observations arriving over time. The core challenge is to contextualize these streams and decide where to gather evidence before selecting a target. We present WORLDS: World-state Observation and Reasoning for Language-guided Discovery and Search, a framework that grounds reasoning in a persistent graph initialized from geospatial priors and updated by perception. Parallel Reasoners maintain competing candidate interpretations and request evidence to distinguish between them. We collect and process the requested observations with a multimodal Examiner, after which a Judge selects a grounded target or requests another pass. WORLDS achieves 51.8% navigation success across all 5,311 CityNav test episodes, the highest reported success rate, exceeding the previous published best by 15.7 percentage points under an OSM-only, high-resolution orthographic protocol. On 1,000 shared episodes, it achieves 50.0% versus 27.9% for the strongest adapted baseline using the same model, prior, sensing stack, and movement budget. Observation-based verification by the Examiner contributes 5.9 points of this success, and at a reduced reasoning-effort setting WORLDS still exceeds the adapted GeoNav baseline by 18.8 points while generating fewer tokens. We also demonstrate WORLDS on a quadrotor, which flies the generated sensing waypoints and grounds three language targets, including a vehicle absent from the map, from its onboard imagery.
Chinese Translation
长时程机器人搜索必须在异构、不完整且常常含糊的证据中解析自然语言指令:包括文本信息、先验地图以及随时间到来的观测。核心挑战在于将这些信息流置于上下文中,并在选定目标之前决定去何处收集证据。我们提出 WORLDS(World-state Observation and Reasoning for Language-guided Discovery and Search,面向语言引导发现与搜索的世界状态观测与推理),这是一个将推理建立在持久图上的框架,该图由地理空间先验初始化,并由感知不断更新。多个并行的 Reasoner(推理器)维护相互竞争的候选解释,并请求证据以区分它们。我们使用多模态 Examiner(检验器)收集并处理所请求的观测,随后由 Judge(裁决器)选择一个有据可依的目标或请求下一轮检验。在 CityNav 全部 5,311 个测试回合中,WORLDS 达到了 51.8% 的导航成功率,这是已报道的最高成功率,在仅使用 OSM、高分辨率正射影像的协议下,比此前已发表的最佳结果高出 15.7 个百分点。在 1,000 个共享回合上,在使用相同模型、先验、传感配置与移动预算的情况下,它达到 50.0%,而经适配的最强基线仅为 27.9%。Examiner 基于观测的验证为这一成功率贡献了 5.9 个百分点;在降低推理力度的设置下,WORLDS 仍比适配后的 GeoNav 基线高出 18.8 个百分点,同时生成更少的 token。我们还在一架四旋翼无人机上演示了 WORLDS,该无人机执行生成的感知航点,并利用机载图像成功定位了三个语言目标,其中包括一个地图上不存在的车辆。
cs.RO / 106 / 2609.23856

Object-Centered Reconstruction for Vision-Based 3D Force Estimation

面向目标的重建方法用于基于视觉的三维力估计
Zhang, Zhonghao, Wu, Mingyeung, Yang, Hao, Acar, Ayberk, Kuntz, Alan, Wu, Jie Ying
Abstract
Excessive force may damage tissue and increase the risk of anastomotic leakage in robotic colorectal surgery. Although the da Vinci 5 provides force sensing, this capability is unavailable on earlier da Vinci systems and many other surgical robotic platforms. In this work, we present a vision-based pipeline for estimating 3D interaction forces from soft-tissue deformation in stereo endoscopic video. We dynamically reconstruct the tissue point cloud in an object-centered coordinate frame, track tissue points with geometric constraints, and predict the 3D force vector with a neural network. We progressively evaluate the pipeline on rubber-glove phantoms, ex vivo porcine colons, and in vivo colorectal surgical video sequences. Under varying tissue orientations and positions within the endoscopic view, as well as different camera viewpoints, the proposed method achieves average root mean square error (RMSEs) of 0.77 N and 1.30 N on the phantom and porcine colon, respectively. Compared with the camera-frame representation, the object-centered representation reduces average RMSE by 51.3% and 56.7%, while geometry-constrained tracking reduces RMSE by 19.8% and 25.3% compared with CoTracker. We further qualitatively demonstrate the feasibility of vision-based force estimation on an in vivo colorectal surgical sequence, as a step toward clinical translation of vision-based, sensorless force estimation.
Chinese Translation
在机器人结直肠手术中,过大的作用力可能损伤组织并增加吻合口漏的风险。尽管达芬奇第五代(da Vinci 5)系统提供了力感知功能,但早期的达芬奇系统以及许多其他手术机器人平台均不具备该能力。在本工作中,我们提出了一种基于视觉的流程,用于从立体内镜视频中的软组织形变估计三维交互力。我们在以物体为中心的坐标系中动态重建组织点云,利用几何约束跟踪组织点,并通过神经网络预测三维力向量。我们在橡胶手套假体、离体猪结肠以及在体结直肠手术视频序列上对该流程进行了渐进式评估。在内镜视野中组织方向和位置不断变化以及相机视角不同的条件下,所提出的方法在假体和猪结肠上的平均均方根误差(RMSE)分别为0.77 N和1.30 N。与相机坐标系表示相比,以物体为中心的表示将平均RMSE分别降低了51.3%和56.7%;而与CoTracker相比,几何约束跟踪使RMSE分别降低了19.8%和25.3%。我们进一步在在体结直肠手术序列上定性展示了基于视觉的力估计的可行性,为基于视觉的无传感器力估计走向临床转化迈出了一步。
cs.RO / 107 / 2609.23863

Grounded Action Model: 3D Grounding as a Foundation for Robotics

接地动作模型:作为机器人学基础的三维接地
Zhang, Gehao, Huang, Weikai, Shailesh, Shailesh, Peng, Yiyan, Duan, Jiafei, Krishna, Ranjay
Abstract
Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned using language, points, or box prompts, which are first transformed into a shared object-centric representation of the selected objects. This representation captures target-focused visual features and metric object geometry, which is mixed with robot state history through a multi-stream transformer to predict action chunks. Although GAMs can be run autonomously, they can also serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation. On RoboTwin 2.0, GAM achieves an average success rate of 55.3% across 50 tasks (vs. 52.0% for Spatial Forcing), including 47.6% under scene randomization (vs. 30.4% for Abot-M0), with its action policy trained only on clean-scene demonstrations. On LIBERO-PRO, it achieves a state-of-the-art average success rate of 61% (vs. 53% for $\pi_{0.5}$) across 16 perturbation settings, with the largest gains when targets are relocated or newly designated. On two real robots, GAM retains 17/20 successes under visual shift on a bimanual YAM versus 4/20 for $\pi_{0.5}$, while its composition with a Molmo2 planner on a Franka achieves 64.7% ID and 49.8% OOD step completion on long-horizon and memory-dependent tasks.
Chinese Translation
操作策略必须知道哪些物体是重要的以及它们位于何处,然而当前机器人基础模型所依赖的预训练骨干网络——从视觉-语言-动作模型(VLA)中的语言到世界-动作模型(WAM)中的视频生成——并不直接要求这种度量级接地能力,而是将其留待从机器人示教中隐式学习。我们提出接地动作模型,这是一种基于三维接地构建的机器人基础模型新范式。GAM 可以通过语言、点或框提示进行条件输入,这些提示首先被转换为所选物体的共享以物体为中心的表示。该表示捕捉以目标为中心的视觉特征和物体的度量级几何信息,并通过多流 Transformer 与机器人状态历史相混合,以预测动作块。尽管 GAM 可以自主运行,它也可以作为一个底层控制器,由高层规划器利用其多种输入模态进行控制,从而实现长时程和依赖记忆的操作。在 RoboTwin 2.0 上,GAM 在 50 个任务中取得了 55.3% 的平均成功率(对比 Spatial Forcing 的 52.0%),其中在场景随机化条件下为 47.6%(对比 Abot-M0 的 30.4%),而其动作策略仅在干净场景的示教数据上训练。在 LIBERO-PRO 上,它在 16 种扰动设置下取得了 61% 的最先进平均成功率(对比 $\pi_{0.5}$ 的 53%),当目标被重新放置或新指定时收益最大。在两台真实机器人上,GAM 在双臂 YAM 的视觉偏移条件下仍保持 17/20 的成功率,而 $\pi_{0.5}$ 仅为 4/20;此外,在 Franka 上,GAM 与 Molmo2 规划器的组合在长时程和依赖记忆的任务上分别取得了 64.7% 的分布内(ID)和 49.8% 的分布外(OOD)步骤完成率。
cs.RO / 108 / 2609.23885

HumynexSurg-1: A Curated Expert Liposuction Dataset

HumynexSurg-1:一个经策划的专家级吸脂手术数据集
Huang, Rhea, Matlock, David L., Reich, Laurence
Abstract
Robot foundation models learn manipulation from large demonstration corpora, but surgery is missing from those corpora: across the 780-hour Open-H surgical collection, one dataset carries synchronized force and none covers an aesthetic procedure. Liposuction is the hard case, because the instrument works under the skin and the surgeon operates by feel and by judgment. Humynex Robotics builds curated expert datasets for this kind of procedure. HumynexSurg-1 is the first release: a master liposuction surgeon performing on porcine abdominal tissue while narrating every decision, recorded with synchronized suction pressure, six-axis hand force/torque, top-down RGB-D video, side video and a lavalier microphone -- 14 episodes, 42,738 frames, 35.6 minutes, 356 utterances of which 95% compile into a liposuction-specific label schema. The capture follows a patent-pending sensing plan organized around the quantities a policy needs, so a channel captured today by a model can be upgraded to a sensor tomorrow without changing the data format. This release captures the instrument motion as a tool-hand track in the side video and provides the force channel as state; the funded capture adds a measured 6-DoF handle pose, a validated force channel, ultrasound imaging of the fat layer, and palpation sensing. As a proof of concept, NVIDIA Isaac GR00T N1.7 fine-tunes on the dataset with no custom code in under an hour per run and learns the recorded sessions; scaling probes on the same episodes show where further gains come from: every new session lowers the error on an unseen session. The dataset, its label schema, its quality-assurance reports and its evaluation protocol are the product; the next capture, many short sessions across fat regions with the sensors named here, is what the probes point to.
Chinese Translation
机器人基础模型从大规模示教语料中学习操作技能,但这些语料中缺少外科手术的内容:在780小时的Open-H手术采集集中,仅有一个数据集包含同步的力信号,且没有任何数据集涵盖美容外科手术。吸脂手术是一个困难案例,因为器械在皮下操作,外科医生依靠触感和判断进行手术。Humynex Robotics正是为此类手术构建经策划的专家数据集。HumynexSurg-1是首个发布版本:由一位资深吸脂手术专家在猪腹部组织上操作,并对每一个决策进行口述讲解;采集数据包括同步的抽吸压力、六轴手部力/力矩、俯视RGB-D视频、侧视视频以及领夹式麦克风——共14个片段、42,738帧、35.6分钟、356条语音表述,其中95%可编译为吸脂手术专用的标签体系。数据采集遵循一项专利申请中的传感方案,围绕策略模型所需的各种物理量组织,使得今天由模型推测(捕获)的通道明天可以升级为真实传感器,而无需更改数据格式。本发布版本将器械运动作为侧视视频中的工具-手部轨迹提供,并将力信号作为状态提供;后续受资助的采集将增加实测的6自由度手柄位姿、经验证的力通道、脂肪层超声成像以及触觉感知。作为概念验证,NVIDIA Isaac GR00T N1.7无需任何自定义代码即可在该数据集上进行微调,每次训练耗时不到一小时,并能学会所记录的手术过程;在同一批片段上的扩展性探究实验显示了进一步提升的空间:每新增一个手术片段,都能降低模型在未见过的片段上的误差。本数据集、其标签体系、质量保证报告和评估协议即是本工作的成果;下一轮采集——在脂肪区域间开展多次使用本文所述传感器的短时程采集——正是这些探究实验所指明的方向。
cs.RO / 109 / 2609.23888

HapticWAM: Distilling Imagined Touch into a World-Action Model without Inference-Time Tactile Sensing

HapticWAM:将想象触觉蒸馏入世界-动作模型而无需推理时触觉感知
Sannikov, Mikhail, Mikhalchuk, Ilya, Gubernatorov, Konstantin, Kovalev, Petr, Oluwatobi, Ogunwoye Faith, Tsetserukou, Dzmitry
Abstract
Contact-rich manipulation requires estimating forces, slip and contact geometry that can remain ambiguous in scene images. Optical tactile sensors provide both visual observations of the contact surface and mechanical measurements, yet learning from these signals raises two challenges: representing contact beyond appearance and transferring its benefits to a policy that does not require fingertip observations at deployment. We introduce HapticWAM, a world-action model that combines heterogeneous tactile encoding, structured contact prediction and teacher-student distillation. Its teacher encodes gel images together with deformation, shear, distributed forces, resultant wrench and derived contact state into a frozen video backbone. Rather than predicting tactile pixels alone, the model jointly generates actions and a contact package describing future events and mechanics. Anticipatory Contact Coupling uses the previously imagined package to condition attention, preserving a contact-related input when direct tactile observations are unavailable. Haptic-Imagination Distillation transfers both contact futures and action predictions to a student that retains the generative contact head but removes its fingertip input branches. On a real-world setup, across three contact-rich pick-and-place tasks, HapticWAM Student achieves a 77% per-task mean success rate (41 of 50 starts, 82% pooled), reaching 95% on one of the tasks, outperforming the evaluated teacher and baseline configurations.
Chinese Translation
接触丰富的操控任务需要估计力、滑动和接触几何信息,而这些信息在场景图像中可能仍不明确。光学触觉传感器既能提供接触表面的视觉观测,也能提供力学测量,然而从这些信号中学习面临两个挑战:如何在表观之外表征接触,以及如何将其优势迁移到部署时无需指尖观测的策略中。我们提出了HapticWAM,一种结合了异构触觉编码、结构化接触预测和师生蒸馏的世界-动作模型(world-action model)。其教师模型将凝胶图像以及形变、剪切、分布力、合力旋量和派生接触状态编码到一个冻结的视频骨干网络中。该模型并非仅预测触觉像素,而是联合生成动作以及描述未来事件与力学的接触信息包。预期接触耦合(Anticipatory Contact Coupling)利用先前想象的接触信息包来调节注意力,从而在直接触觉观测不可用时保留与接触相关的输入。触觉想象蒸馏(Haptic-Imagination Distillation)将接触未来预测和动作预测同时迁移到学生模型中,学生模型保留了生成式接触头,但移除了其指尖输入分支。在真实世界环境中,针对三个接触丰富的抓取放置任务,HapticWAM学生模型实现了77%的单任务平均成功率(50次起始中41次成功,汇总成功率为82%),其中一个任务达到95%,优于所评估的教师模型和基线配置。
cs.RO / 110 / 2609.23896

BarrierFormer: Transformer-Guided Predictive Barrier Enforcement for Safe Robot Control

BarrierFormer:基于Transformer引导的预测性障碍函数约束的安全机器人控制
Chauhan, Anandsingh, Garg, Kunal
Abstract
Control barrier functions (CBFs) have become one of the most popular tools for encoding and enforcing state constraints in safety-critical robotics. Standard CBF approaches are inherently myopic in nature as they enforce safety only at the current time step. Consequently, the system can be driven toward the boundary of the safe set where no feasible safe control exists at a future timestep. Model predictive control (MPC) based approaches address this by enforcing state constraints over a receding horizon. However, such approaches generally require the model to be known for solving a constrained optimization problem at every step, which is computationally expensive for real-time deployment. We propose BarrierFormer, a barrier-supervised transformer framework that addresses these limitations by encoding rollout-level CBF constraints in learning a model-free safe policy. A causal transformer encodes observation-action history, autoregressively generates a predictive rollout through the dynamics head to replace the model, and provides a residual correction to a nominal controller through the action head to replace the online computation. A barrier critic operating on local observations evaluates CBF constraint violations along this rollout, and a safety teacher computes barrier-consistent actions satisfying these constraints as direct supervision targets for the learned control policy. During inference, the policy maps observation-action history to control actions without any online optimization or model knowledge, enabling real-time model-free predictive safety enforcement. Evaluations across linear and nonlinear, 2D and 3D dynamical systems for safe goal-directed navigation demonstrate that BarrierFormer outperforms existing reinforcement learning (RL)-based, diffusion-based, MPC-based, and transformer-based approaches in safety rate and inference latency.
Chinese Translation
控制障碍函数(Control Barrier Functions, CBFs)已成为安全关键机器人系统中编码和执行状态约束的最流行工具之一。标准CBF方法本质上是短视的,因为它们仅在当前时间步强制执行安全性约束。因此,系统可能被驱动至安全集的边界处,导致未来时间步不存在可行的安全控制。基于模型预测控制(Model Predictive Control, MPC)的方法通过在滚动时域内强制执行状态约束来解决这一问题。然而,此类方法通常需要已知模型以在每一步求解约束优化问题,这对于实时部署而言计算代价高昂。我们提出BarrierFormer,一个由障碍函数监督的Transformer框架,通过在学习无模型安全策略的过程中编码滚动预测级别的CBF约束来解决上述局限性。一个因果Transformer对观测-动作历史进行编码,通过动力学头自回归地生成预测滚动轨迹以替代模型,并通过动作头对标称控制器提供残差修正以替代在线计算。一个基于局部观测运行的障碍评价器评估该滚动轨迹上的CBF约束违反情况,同时一个安全教师计算满足这些约束的与障碍函数一致的动作,作为所学控制策略的直接监督目标。在推理阶段,该策略将观测-动作历史直接映射为控制动作,无需任何在线优化或模型知识,从而实现实时的无模型预测性安全约束执行。在线性与非线性、二维与三维动力学系统上的安全目标导航任务评估表明,BarrierFormer在安全率和推理延迟方面均优于现有的基于强化学习(RL)、基于扩散模型、基于MPC以及基于Transformer的方法。
cs.RO / 111 / 2609.23910

ReVeal: A Reconstruction-Aware Real-to-Sim Framework for VLA Policy Evaluation

ReVeal:一种面向重建感知的虚实转换(Real-to-Sim)框架,用于VLA策略评估
Wang, Xinyi, Hao, Heng, Hu, Wenjun, Li, Anna Enyu, Ma, Dizhi, Ramani, Karthik, Moon, Hankyu, Kwon, Yeong-Dae
Abstract
Simulation-based evaluation provides a scalable and repeatable alternative to real-world evaluation of vision-language-action (VLA) policies. However, reconstruction errors can cause simulated policy performance to diverge from real-world performance, motivating the need to assess reconstructed environments for downstream VLA policy evaluation. We present ReVeal, a real-to-sim assessment framework combining workspace reconstruction, reconstruction-level assessment, and matched closed-loop policy evaluation. Novel-View Mesh Fidelity (NVMF) and Annotated Planar Geometry Fidelity (APGF) assess observation and planar geometric fidelity, respectively. We also develop PGSR-D, a reconstruction pipeline incorporating monocular depth supervision to improve geometry where multi-view visual cues are limited. Across 8 assessment scenes, NVMF and APGF consistently distinguish the fidelity of 2DGS, PGSR, and PGSR-D. Matched evaluations of GR00T, SmolVLA, and pi0.5 across 8 humanoid manipulation tasks show consistent ordering between reconstruction fidelity and real-sim performance agreement across pipelines. Further analysis of the evaluation workspaces shows that higher fidelity is associated with stronger real-sim agreement.
Chinese Translation
基于仿真的评估为视觉-语言-动作(VLA)策略的现实世界评估提供了一种可扩展且可重复的替代方案。然而,重建误差可能导致仿真中的策略表现与真实世界表现产生偏差,因此需要对用于下游VLA策略评估的重建环境进行评估。我们提出了ReVeal,一个结合工作空间重建、重建级别评估和匹配闭环策略评估的虚实转换(Real-to-Sim)评估框架。新视角网格保真度(NVMF)与标注平面几何保真度(APGF)分别评估观测保真度和平面几何保真度。我们还开发了PGSR-D,一个融合单目深度监督的重建流水线,以在多视角视觉线索受限的情况下改进几何重建。在8个评估场景中,NVMF和APGF能够一致地区分2DGS、PGSR和PGSR-D的保真度。对GR00T、SmolVLA和pi0.5在8个人形机器人操作任务上的匹配评估表明,各流水线的重建保真度与真实-仿真性能一致性之间具有一致的排序。对评估工作空间的进一步分析表明,更高的保真度对应更强的真实-仿真一致性。
cs.RO / 112 / 2609.23928

MR-SPITE: Accelerating Multi-Robot Conflict Scans via Hierarchical Swept-Volume Approximations

MR-SPITE:基于分层扫掠体近似的多机器人冲突扫描加速方法
Markowicz, Marta, Motes, James, Morales, Marco, Amato, Nancy
Abstract
Conflict scanning over synchronized robot paths requires detailed collision checking, potentially across every robot pair at every timestep, and may be repeated many times as conflicts are repaired. We present Multi-Robot SPITE (MR-SPITE), a conservative, motion-segment-based filter for accelerating these scans. MR-SPITE partitions each path into temporal intervals and assigns conservative bounds to each segment. An interval scheduler compares bounds for temporally overlapping motions: disjoint bounds certify the shared window as conflict-free, while unresolved windows are passed to the underlying collision checker. We integrate MR-SPITE into ARC and combine it with VAMP-based collision checking. For 16 Fetch robots, ARC with MR-SPITE achieves a paired median conflict scan speedup of 7.18x and reduces median planning time by 57% relative to the baseline ARC implementation with PRM+VAMP. These results demonstrate that motion-segment bounds complement configuration-level collision acceleration while preserving the behavior of the underlying discretized scanner.
Chinese Translation
对同步机器人路径进行冲突扫描需要进行精细的碰撞检测,可能需要在每个时间步对所有机器人对进行检测,并且随着冲突的修复可能需要重复多次。我们提出了多机器人SPITE(MR-SPITE),一种基于运动分段的保守过滤器,用于加速此类扫描。MR-SPITE将每条路径划分为时间区间,并为每个分段分配保守边界。区间调度器对时间上重叠的运动比较其边界:边界不相交即可证明该共享窗口无冲突,而无法判定的窗口则传递给底层碰撞检测器。我们将MR-SPITE集成到ARC中,并结合基于VAMP的碰撞检测。对于16台Fetch机器人,带有MR-SPITE的ARC相对于使用PRM+VAMP的基线ARC实现,配对冲突扫描速度中位数提升7.18倍,规划时间中位数减少57%。这些结果表明,运动分段边界能够在保持底层离散化扫描器行为的同时,补充配置层面的碰撞加速。
cs.RO / 113 / 2609.23943

FinsSim: A Reality-Aligned Integrated Simulation Platform for Underwater Robot Learning

FinsSim:一个面向水下机器人学习的现实对齐集成仿真平台
Zhang, Yu, Song, Yuanmingqing, Rao, Xiangyun, Fong, Pangkit, Zhang, Kunhao, Fang, Chongrong, He, Jianping
Abstract
Underwater robot learning relies on simulators that integrate high-fidelity hydrodynamics, convenient learning interfaces, and a credible transition to real scenarios. In this work, we present FinsSim, a reality-aligned integrated simulation platform for Sim-to-Real underwater robot learning. FinsSim first constructs high-fidelity simulation with selectable backends to adapt to diverse requirements. To facilitate underwater robot research, it further offers standard control baselines, alongside with unified robot learning workflows. For reliable Sim-to-Real transfer, FinsSim adopts a multi-sensor fusion scheme to provide low-cost yet precise localization. Moreover, it implements calibrated thruster-hydrodynamics models and a constrained wrench allocation algorithm. Bridging these modules by ROS~2, FinsSim establishes a complete Sim-to-Real transfer pipeline. Through matched simulations and experiments, it is demonstrated that reliable Sim-to-Real transfer of underwater robot control policies can be achieved with the FinsSim framework. Separate ablation studies also validate that the modules of FinsSim can address the pivotal issues of underwater Sim-to-Real from different aspects. Overall, this work aims to bridge the gap between theoretical research and practical applications, ultimately driving advancements in the field of underwater robotics.
Chinese Translation
水下机器人学习依赖于集成高保真水动力学、便捷学习接口以及可信的真实场景迁移能力的仿真器。在本工作中,我们提出了FinsSim,一个面向Sim-to-Real(仿真到现实)水下机器人学习的现实对齐集成仿真平台。FinsSim首先构建了具有可选后端的高保真仿真环境,以适应多样化的需求。为促进水下机器人研究,该平台进一步提供了标准控制基线以及统一的机器人学习工作流。为实现可靠的Sim-to-Real迁移,FinsSim采用多传感器融合方案,提供低成本且精确的定位。此外,该平台实现了经校准的推进器-水动力学模型以及受约束的力(矩)分配算法。通过ROS~2桥接上述各模块,FinsSim建立了一条完整的Sim-to-Real迁移流水线。通过匹配的仿真与实验,证明了基于FinsSim框架能够实现水下机器人控制策略的可靠Sim-to-Real迁移。独立的消融研究也验证了FinsSim的各模块能够从不同方面解决水下Sim-to-Real的关键问题。总体而言,本工作旨在弥合理论研究与实际应用之间的差距,最终推动水下机器人领域的发展。
cs.RO / 114 / 2609.23944

Topology-Informed Visual Prompting For Vision Language Action Policies

面向视觉-语言-动作策略的拓扑引导视觉提示方法
Wu, Haoyang, Kumar, Abhinav, Berenson, Dmitry
Abstract
Vision-language-action (VLA) policies can struggle with manipulation tasks with complex obstacle geometries due to partial observability. These complex geometries can lead to similar visual observations or robot configurations requiring qualitatively different actions, a distinction that can be quantified using topological signatures. While motion planners with full knowledge of environment geometries and object states can reason about these signatures in planning, this information is often not known at deployment. To address this issue, we present a topology-guided visual-prompting framework that uses simulation-based planning to augment a nominal demonstration dataset and provides vision-based guidance at deployment. Our method uses a Gauss-Linking-Integral topological signature representation to capture important topological properties of the environment. Using privileged geometry information from a simulation approximation of our environment, we augment a VLA fine-tuning dataset with trajectories that move the system to a demonstrated signature and, from the new configuration, resume task execution. A vision-language model (VLM) is fine-tuned on the same dataset to both predict signatures from live camera observations and predict end-effector waypoints, which are rendered as visual prompts on the observations to guide the VLA. Across three simulated bimanual tasks and a real-world box pickup task, our method outperforms a VLA fine-tuned only on nominal demonstrations and a VLM-prompting baseline that can remove topology-relevant information from observations. On hardware, it exceeds the strongest baseline by 40% in task success. Project website: https://topology-vla.github.io.
Chinese Translation
由于部分可观测性,视觉-语言-动作(VLA)策略在处理具有复杂障碍物几何结构的操作任务时可能遇到困难。这些复杂的几何结构可能导致相似的视觉观测或机器人构型需要性质上不同的动作,而这种区别可以利用拓扑签名(topological signatures)来量化。尽管拥有环境几何和物体状态完整信息的运动规划器可以在规划中对这些签名进行推理,但这些信息在部署时通常是未知的。为了解决这一问题,我们提出了一种拓扑引导的视觉提示框架,该框架利用基于仿真的规划来增强名义演示数据集,并在部署时提供基于视觉的引导。我们的方法使用高斯连接积分(Gauss-Linking-Integral)拓扑签名表示来捕捉环境的重要拓扑性质。利用从环境仿真近似中获得的特权几何信息,我们通过将系统移动到已演示签名的轨迹来增强VLA微调数据集,并在新的构型下恢复任务执行。我们在同一数据集上微调一个视觉-语言模型(VLM),使其既能从实时相机观测中预测签名,又能预测末端执行器路径点,这些路径点被渲染为观测图像上的视觉提示以引导VLA。在三个仿真双手任务和一个真实世界的箱子抓取任务上,我们的方法优于仅在名义演示上微调的VLA,以及可能从观测中移除拓扑相关信息的VLM提示基线。在真实硬件上,我们的方法在任务成功率上超出最强基线40%。项目网站:https://topology-vla.github.io。
cs.RO / 115 / 2609.23968

Opt2VLA: Force-Aware Vision-Language-Action for Contact-Rich Humanoid Whole-Body Manipulation

Opt2VLA:面向接触密集型人形机器人全身操作的力感知视觉-语言-动作模型
Liu, Fukang, Chen, Yipu, Jang, Jaehwi, Xu, Danfei, Kira, Zsolt, Zhao, Ye
Abstract
Humanoid robots are expected to perform diverse human-level tasks in daily environments, many of which require precise regulation of interaction forces. While recent vision-language-action (VLA) models have shown promise for semantic planning and visuomotor control, existing humanoid systems primarily represent actions through geometric motion goals and rely on whole-body controllers focused on motion tracking, with limited explicit reasoning or control of interaction forces. This limitation is particularly relevant in contact-rich tasks, where geometrically similar motions may require different force regimes depending on the task context and where visual observations may become unreliable after contact. In this work, we present Opt2VLA, a force-aware VLA framework that introduces explicit force commands at the VLA-to-control interface for humanoid whole-body manipulation. A single multi-task VLA policy jointly predicts both geometric motion goals and continuous contact-force references, which are tracked by task-specific reinforcement learning (RL)-based whole-body controllers. To provide scalable and physically grounded supervision, we generate dynamically feasible and contact-consistent training data via whole-body trajectory optimization (TO) with explicit force references. We evaluate Opt2VLA on three contact-rich humanoid tasks and show that explicit force conditioning enables more accurate and consistent force regulation than motion-only control, while physically grounded torque supervision from TO further improves force tracking accuracy and stability. Closed-loop evaluations further demonstrate language-conditioned force modulation with Opt2VLA in simulation and on humanoid hardware.
Chinese Translation
人形机器人被期望在日常环境中执行多样化的类人水平任务,其中许多任务需要精确调控交互力。尽管近期的视觉-语言-动作(VLA)模型在语义规划和视觉运动控制方面展现出潜力,但现有人形机器人系统主要通过几何运动目标来表示动作,并依赖于专注于运动跟踪的全身控制器,对交互力的显式推理或控制十分有限。这一局限在接触密集型任务中尤为突出:几何上相似的运动可能因任务情境不同而需要不同的力模式,且接触发生后视觉观测可能变得不可靠。在本工作中,我们提出 Opt2VLA,一个力感知的 VLA 框架,它在 VLA 与控制器的接口处引入显式力指令,用于人形机器人全身操作。单一的多任务 VLA 策略同时预测几何运动目标和连续的接触力参考值,这些参考值由基于任务特定强化学习(RL)的全身控制器进行跟踪。为了提供可扩展且物理上合理的监督信号,我们通过包含显式力参考的全身轨迹优化(TO)生成动力学可行且接触一致的训练数据。我们在三项接触密集型人形机器人任务上评估了 Opt2VLA,结果表明显式的力条件化比纯运动控制能够实现更精确、更一致的力调控,而来自 TO 的物理合理力矩监督进一步提升了力跟踪的精度与稳定性。闭环评估进一步证明了 Opt2VLA 在仿真和人形机器人硬件上具备语言条件化的力调制能力。
cs.RO / 116 / 2609.23976

Anticipatory Robot Goalkeeping via Monotone Optimal Stopping

基于单调最优停止的预判式机器人守门
Zhang, Hao E., Geng, Ruize, Li, Yisen, Niu, Yaru, Wang, Yikai, Haque, Raihan, Zbiss, Khalil, Luo, Guanyang, Wang, Hui-ping, Tseng, H. Eric, Zhao, Ding
Abstract
Robots engaged in fast physical interactions often need to act before the intent of another agent is fully known. Anticipatory goalkeeping illustrates this challenge. Waiting provides more reliable information about the target but reduces the physical opportunity for interception, whereas acting early preserves reachability but requires initiating motion under uncertainty. Given a fixed closed-loop save controller, we formulate the decision of when to initiate motion as a policy-conditional finite-horizon optimal stopping problem. Building on this formulation, we propose monotone optimal stopping (MOS), a structured release-timing method for dynamic robotic interception. The quadruped save policy is trained with reinforcement learning, while MOS determines when the policy should be activated from the evolving robot state and target belief. Rather than predicting a release time or relying on confidence alone, MOS learns the return advantage of acting now over waiting for one more observation. We derive a direct Bellman recursion for this act-versus-wait margin and impose monotonicity only with respect to physical urgency, reflecting the irreversible loss of interception opportunity as time elapses. This structure enables early activation for dynamically demanding saves while preserving closed-loop adaptation when later observations change the predicted target. Under a single-crossing condition, MOS admits a threshold release boundary with a bounded approximation error. Extensive simulation studies show that MOS improves the mean save rate from 67.7% to 74.4% over a parameter-matched learned gate and increases reversal saves from 52.1% to 66.5%. Real-robot experiments further demonstrate rapid interception and post-release direction correction under human shot-direction feints.
Chinese Translation
从事快速物理交互的机器人往往需要在其他智能体的意图被完全知晓之前采取行动。预判式守门(anticipatory goalkeeping)正体现了这一挑战:等待可以获得关于目标的更可靠信息,但会减少拦截的物理机会;而过早行动虽然保持了可达性,却需要在不确定性下启动运动。在给定固定闭环扑救控制器的前提下,我们将何时启动运动的决策建模为一个策略条件化的有限时域最优停止问题。基于这一建模,我们提出了单调最优停止(Monotone Optimal Stopping, MOS),一种面向动态机器人拦截的结构化释放时机决策方法。四足机器人的扑救策略通过强化学习训练,而MOS则根据不断演化的机器人状态和目标信念来确定策略的激活时机。MOS并非预测释放时刻,也不单纯依赖置信度,而是学习“立即行动”相对于“再等待一次观测”的回报优势。我们针对这一“行动—等待”差值推导了直接的Bellman递归,并仅对物理紧迫性施加单调性约束,以反映随时间流逝拦截机会的不可逆损失。这一结构使得系统能够在动态要求高的扑救中提前激活,同时在后续观测改变预测目标时保持闭环适应能力。在单交叉条件下,MOS具有阈值释放边界,且近似误差有界。大量仿真实验表明,与参数匹配的学习型门控相比,MOS将平均扑救率从67.7%提升至74.4%,并将反向扑救率从52.1%提升至66.5%。真实机器人实验进一步验证了系统在人类射门方向假动作下仍能实现快速拦截及释放后的方向修正。
cs.RO / 117 / 2609.23997

RoboTalk: Learning Multi-Robot Communication and Coordination from Multimodal Demonstrations

RoboTalk:从多模态示范中学习多机器人通信与协作
Goldfajn, Dorian Benhamou, Nakamura, Mason, Mahmud, Saaduddin, Svegliato, Justin, Wray, Kyle H., Zilberstein, Shlomo
Abstract
Multi-robot collaboration could enable more efficient and scalable solutions to complex robotic tasks, but collaboration under partial observability remains challenging. Natural-language communication offers a promising approach to coordinating robots under partial observability. However, in decentralized manipulation, jointly learning explicit inter-robot communication and skill-level action selection from multimodal demonstrations remains underexplored for small vision-language models (VLMs) intended for on-device deployment. To address this gap, we introduce RoboTalk, a synthetic data-generation pipeline and dataset of 7,950 multimodal trajectories spanning 53 mobile-manipulation kitchen tasks for training small VLMs to communicate and coordinate. The dataset includes a leader-follower planning protocol, tool calls (perception, manipulation, navigation, and communication), rationale traces, and diversified natural-language communication. Fine-tuning open-source models on our dataset can reach 77% success on novel held-out tasks, a significant improvement over the untuned open source models, which had a success rate of around ~2%.
Chinese Translation
多机器人协作能够为复杂机器人任务提供更高效、更可扩展的解决方案,但部分可观测性下的协作仍然具有挑战性。自然语言通信为在部分可观测条件下协调机器人提供了一种有前景的方法。然而,在去中心化操作任务中,面向端侧部署的小型视觉-语言模型(VLM)如何从多模态示范中联合学习显式的机器人间通信与技能级动作选择,目前仍缺乏充分研究。为填补这一空白,我们提出了RoboTalk,一个合成数据生成流程及数据集,涵盖53个移动操作厨房任务的7,950条多模态轨迹,用于训练小型VLM进行通信与协作。该数据集包含领导者-跟随者规划协议、工具调用(感知、操作、导航与通信)、推理轨迹以及多样化的自然语言通信。在我们的数据集上微调开源模型后,在新颖的保留任务上可达到77%的成功率,相较未经微调的开源模型(成功率约为2%)有显著提升。
cs.RO / 118 / 2609.24033

Imagine-RL: Residual-Confidence-Guided Cross-Attention for World-Model-Augmented VLA Reinforcement Learning

Imagine-RL:面向世界模型增强的视觉-语言-动作(VLA)强化学习的残差-置信度引导交叉注意力方法
Hu, Kejia, Zhai, Wentong, Zhao, Bo, Liang, Shuai
Abstract
Reliable action evaluation in contact-rich manipulation requires looking beyond the current observation to future visual and contact consequences. Existing noise-space reinforcement learning efficiently steers a frozen Vision-Language-Action (VLA) policy, but its critics largely ignore these consequences. We present Imagine-RL, which augments noise-space VLA post-training with action-conditioned visual-torque imagination. For each candidate action chunk, a frozen visual-torque latent world model (VTLWM) autoregressively predicts compact future representations without pixel reconstruction. A current image-state-action query attends to observed histories and predicted futures, while previous-window prediction residuals provide token-wise confidence priors that suppress unreliable future tokens. By combining current evidence with predicted consequences, the action critic better evaluates candidate actions and supervises the actor, while the VLA and VTLWM remain frozen. Across four real-robot tasks with 50 evaluation trials per task, Imagine-RL uses only 100 RL trajectories and improves the average success rate by (23.6%) over DSRL and by (60%) over VLA baselines.
Chinese Translation
在接触丰富的操作任务中,可靠的动作评估需要超越当前观测,预见未来的视觉与接触后果。现有的噪声空间强化学习能够高效地引导冻结的视觉-语言-动作(VLA)策略,但其评论家(critic)在很大程度上忽略了这些后果。我们提出 Imagine-RL,通过动作条件下的视觉-力矩想象来增强噪声空间的 VLA 后训练。对于每个候选动作块,一个冻结的视觉-力矩潜在世界模型(Visual-Torque Latent World Model, VTLWM)以自回归方式预测紧凑的未来表征,而无需像素重建。当前图像-状态-动作查询会关注观测历史与预测的未来,同时前一时间窗的预测残差提供逐词元(token-wise)的置信度先验,以抑制不可靠的未来词元。通过将当前证据与预测后果相结合,动作评论家能够更好地评估候选动作并监督行动者(actor),而 VLA 和 VTLWM 均保持冻结。在四项真实机器人任务上(每项任务进行 50 次评估试验),Imagine-RL 仅使用 100 条强化学习轨迹,平均成功率相较于 DSRL 提升了 23.6%,相较于 VLA 基线提升了 60%。
cs.RO / 119 / 2609.24048

What Matters in Designing World Action Models: An Empirical Study

世界动作模型设计中的关键因素:一项实证研究
Tang, Chao, Wang, Haoqing, Cen, Zilang, Mi, Weishi, Xia, Wei, Liu, Fangcheng, Cheng, Anda, Shen, Yeqing, Cui, Xiaohui, Zhang, Xiaoyuan, Tang, Yehui, Li, Tingguang
Abstract
World Action Models (WAMs) have emerged as a promising paradigm for generalizable robot control. Despite the growing number of WAM systems, existing works often introduce unified systems that bundle together multiple design choices, such as architecture and training strategy, making it difficult to isolate individual contributions and systematically compare alternative designs. In this work, we present a controlled study that disentangles these design choices and analyzes not only their empirical effects, but also how and why they shape WAMs. More specifically, we focus on three fundamental questions in building WAMs: (1) what causal structure should govern the interaction between world modeling and action generation? (2) in which latent space should world modeling be performed? and (3) how do different world-action modeling objectives affect model behavior and performance? Through structurally controlled experiments on three representative benchmarks, RoboCasa-GR1, LIBERO, and LIBERO-Plus, we systematically compare six causal structures, eight latent representations, and four training objectives, covering popular design choices in existing WAMs. We further validate our key findings on real-robot data from the DROID dataset. We hope to provide a systematic understanding of how core design choices affect world-action modeling and what principles can guide the development of future WAM systems.
Chinese Translation
世界动作模型(World Action Models, WAMs)已成为实现可泛化机器人控制的一种有前景的范式。尽管WAM系统数量不断增长,现有工作往往将多种设计选择(如架构和训练策略)捆绑在一起引入统一系统,导致难以隔离各个因素的独立贡献,也难以系统地比较不同的设计方案。在本工作中,我们提出了一项受控研究,将这些设计选择解耦,不仅分析其实证效果,还探究它们如何以及为何塑造WAM。更具体地,我们聚焦于构建WAM时的三个基本问题:(1)世界建模与动作生成之间的交互应由何种因果结构支配?(2)世界建模应在哪种潜在空间中进行?(3)不同的世界-动作建模目标如何影响模型的行为与性能?通过在三个代表性基准(RoboCasa-GR1、LIBERO和LIBERO-Plus)上进行结构受控实验,我们系统地比较了六种因果结构、八种潜在表示和四种训练目标,涵盖了现有WAM中的主流设计选择。我们还在来自DROID数据集的真机数据上进一步验证了关键发现。我们希望为理解核心设计选择如何影响世界-动作建模、以及哪些原则可以指导未来WAM系统的发展提供系统性的认识。
cs.RO / 120 / 2609.24054

AquaOrbit: Sim-to-Real Reinforcement Learning for Underwater Target Orbiting under Intermittent Visual Feedback

AquaOrbit:面向间歇性视觉反馈下水下目标环绕的仿真到真实强化学习
Yao, Kanzhong, Leng, Jinyi, Zhang, Hao, Sun, Zhe, Li, Xuelong
Abstract
Intermittent visual loss disrupts target-relative feedback during underwater orbiting, making it difficult to maintain coordinated motion and reacquire a moving target. We present AquaOrbit, a reinforcement-learning controller with a recovery module for underwater target orbiting under interrupted visual feedback. During detection loss, the recovery module uses latched line-of-sight, roll, and depth references to support stabilization and target reacquisition. We train the controller in Isaac Sim with dynamics, observation, and vision-loss randomization. Evaluated without retraining in Gazebo/ROS2 under a different physics engine and perception perturbations, AquaOrbit completes 20/20 orbiting trials in each of the static- and moving-target conditions on an unseen variable-depth 3-D trajectory. In the moving-target condition, it reduces mean line-of-sight error by approximately 46% relative to a PID-based visual servoing controller with recovery while maintaining comparable path-tracking accuracy; removing the recovery module reduces completion to 9/20. Zero-shot physical deployment with fully onboard perception and control demonstrates elliptical, figure-eight, and variable-depth circular trajectories, including the latter two trajectory types absent from training. The robot maintains attitude stability during manual occlusions lasting up to 8s and reacquires the target within 2.5s in the reported attitude-induced field-of-view loss events.
Chinese Translation
视觉信息的间歇性丢失会破坏水下环绕(orbiting)任务中相对于目标的反馈,使得维持协调运动和重新捕获移动目标变得困难。我们提出了AquaOrbit,一种用于视觉反馈中断下水下目标环绕的强化学习控制器,并配备了恢复模块。在目标检测丢失期间,恢复模块利用锁存的视线(line-of-sight)、横滚和深度参考来支持姿态稳定与目标重新捕获。我们在Isaac Sim中通过动力学、观测和视觉丢失随机化对控制器进行训练。在未重新训练的情况下,于Gazebo/ROS2中在不同物理引擎和感知扰动下进行评估,AquaOrbit在一条未见的变深度三维轨迹上,静态目标和移动目标两种条件下均完成了20/20次环绕试验。在移动目标条件下,相比带恢复模块的基于PID的视觉伺服控制器,它将平均视线误差降低了约46%,同时保持相当的路径跟踪精度;移除恢复模块会使完成率降至9/20。零样本(zero-shot)物理部署中,感知与控制完全在机载完成,成功演示了椭圆、八字形和变深度圆形轨迹,其中后两种轨迹类型并未出现在训练中。在持续长达8秒的人工遮挡期间,机器人保持姿态稳定,并在所报告的因姿态引起的视场丢失事件中于2.5秒内重新捕获目标。
cs.RO / 121 / 2609.24055

Toward Human-in-the-Loop Robot Failure Recovery: Bridging Communication Gaps in Human-Robot Collaboration

迈向人在环路机器人故障恢复:弥合人机协作中的沟通鸿沟
Ekpo, Promise, Vijay, Teju, Mandalik, Dhruv, Jain, Tisha, Ibrayeva, Arman, Sil, Sunishka, Tellex, Stefanie A., Taylor, Angelique
Abstract
Robots can recover from failures by asking bystanders for help, but effective human-in-the-loop recovery requires communication that accounts for differences in people's knowledge. Prior inverse-semantics work generates requests using a single listener model, leaving differences in listener knowledge untested. We introduce Listener Differences in Human-Robot Interaction (LD-HRI), a game, dataset, and benchmark that evaluates speakers through human listener performance. Our evaluation examines request properties, large language model (LLM) speakers, and inverse-semantics request-selection algorithms under controlled differences in listener information. The corpus contains 446 human-written requests and 1{,}302 listener trials. We additionally evaluated 24 frozen LLM-written requests with 70 human listeners across 560 trials. Novice success is descriptively higher with model-written requests across all four tasks, yet both request sources leave substantial expert--novice gaps, including 16 percentage points for LLM requests. LD-HRI makes these gaps measurable, providing a foundation for designing more robust communication in human-robot and human-agent interaction.
Chinese Translation
机器人可以通过向旁观者求助来从故障中恢复,但有效的人在环路(human-in-the-loop)恢复需要能够考虑到人们知识差异的沟通方式。先前的逆语义(inverse-semantics)研究使用单一听者模型生成请求,未能检验听者知识差异的影响。我们提出了人机交互中的听者差异(Listener Differences in Human-Robot Interaction, LD-HRI),这是一个通过人类听者表现来评估说话者的游戏、数据集与基准测试。我们的评估在受控的听者信息差异下,考察了请求属性、大语言模型(LLM)说话者以及逆语义请求选择算法。该语料库包含446条人类撰写的请求和1,302次听者试验。此外,我们使用70名人类听者对24条冻结的LLM撰写请求进行了560次试验评估。在所有四个任务中,新手在模型撰写请求下的成功率在描述性上更高,但两种来源的请求均留下了显著的专家—新手差距,其中LLM请求的差距达16个百分点。LD-HRI使这些差距可被量化测量,为人机交互和人—智能体交互中设计更鲁棒的沟通奠定了基础。
cs.RO / 122 / 2609.24059

Automatic Labelling for Bimanual Mobile Manipulation

双臂移动操作(Bimanual Mobile Manipulation)的自动标注
Lu, Yupu, Pan, Jia
Abstract
Semantically meaningful subtask labels can provide useful contexts for long-horizon policies, but automatically identifying both reliable temporal boundaries and broad semantic descriptions for annotations remains difficult. We present an automatic labelling pipeline that assigns temporal localisation to deterministic trajectory analysis and semantic interpretation to vision-language (VL) reasoning. The pipeline segments synchronised kinematic signals into phases, performs phase-localised VL reasoning to describe the contents, and aggregates the outputs for the base, left arm, and right arm actions. We evaluate this pipeline primarily on 29 real Galaxea bimanual mobile-manipulation tasks. Repeating the VL reasoning three times first produces the same output value for 87.4% on selected tasks. A review by nine participants across all 29 tasks then judgements on the labelled phases and shows positive acceptance of temporal divisions (90.5%), body labels (90.7%), and arm labels (78.7%). The results indicate that the segmentation-VL design can produce structured annotations while preserving asynchronous bimanual behaviour, providing a basis for richer semantic subtask identification and state-based verification.
Chinese Translation
具有语义意义的子任务标签可以为长时程策略(long-horizon policies)提供有用的上下文信息,但自动识别可靠的时序边界和广泛的语义描述仍然困难。我们提出了一种自动标注流程,将时序定位交由确定性轨迹分析完成,将语义解释交由视觉-语言(VL)推理完成。该流程将同步的运动学信号分割为多个阶段,进行阶段局部化的VL推理以描述内容,并对底盘、左臂和右臂动作的输出进行聚合。我们主要在29个真实的Galaxea双臂移动操作任务上评估该流程。在选定的任务上,重复三次VL推理后,输出值的一致性达到87.4%。随后,由九名参与者对全部29个任务进行评审,对标注阶段做出判断,结果显示对时序划分(90.5%)、机体标签(90.7%)和手臂标签(78.7%)的接受度均为正向。结果表明,这种“分割+VL”的设计能够在保留异步双臂行为的同时生成结构化标注,为更丰富的语义子任务识别和基于状态的验证奠定了基础。
cs.RO / 123 / 2609.24062

Safety Control of a Hyper-redundant Robot via Adaptive Weighted Control Barrier Functions

基于自适应加权控制障碍函数的超冗余机器人安全控制
Cai, Zijian, Wong, Kiwan, Xin, Wenci, Xiao, Wei, Rus, Daniela, Laschi, Cecilia
Abstract
Hyper-redundant robots are well suited for confined-space manipulation due to their high dexterity, but safe operation in cluttered environments remains challenging. In addition, their slender structures often lead to uneven load distributions and nonuniform tracking errors along the body. To address these issues, this work proposes a weighted control barrier functions (W-CBFs) framework that enforces safety constraints while reducing tracking errors caused by uneven loading. The proposed controller was first evaluated on a circular path-following task under different obstacle configurations. With fixed weights, compared to the non-weighted method, the maximum reduction in root-mean-square (RMS) tracking error was 59.6\% in simulation and 87.7\% in physical experiments. An adaptive weighting strategy was then investigated based on the discrepancy between simulated and experimental performance under different mapping functions. The RMS errors were further reduced by 21.9\% and 8.5\%, respectively, although the error increases when obstacles were located close to the robot body. Finally, the robot was evaluated in a cleaning task requiring coverage of a rectangular area and compared with manual teleoperation. Although the controller was not explicitly optimized for area coverage, the autonomous strategy achieved comparable or better coverage performance while avoiding collisions with the surrounding frame, whereas collisions occurred during manual operation.
Chinese Translation
超冗余机器人凭借其高灵巧性非常适合狭窄空间操作,但在杂乱环境中实现安全运行仍具有挑战性。此外,其细长结构往往导致载荷分布不均匀以及沿机体方向的跟踪误差不均匀。为解决这些问题,本文提出了一种加权控制障碍函数(W-CBFs)框架,在保障安全约束的同时减小由不均匀载荷引起的跟踪误差。首先,该控制器在不同的障碍物配置下于圆形路径跟踪任务中进行了评估。在固定权重条件下,与不加权方法相比,仿真中均方根(RMS)跟踪误差最大降低了59.6%,物理实验中最大降低了87.7%。随后,基于不同映射函数下仿真与实验性能之间的差异,研究了一种自适应加权策略。RMS误差进一步分别降低了21.9%和8.5%,但当障碍物靠近机器人机体时误差会有所增加。最后,在一项需要覆盖矩形区域的清洁任务中对机器人进行了评估,并与人工遥操作进行了比较。尽管该控制器并未针对区域覆盖进行显式优化,自主策略仍实现了相当甚至更优的覆盖性能,同时避免了与周围框架的碰撞,而人工操作过程中则发生了碰撞。
cs.RO / 124 / 2609.24068

When Does Touch Matter? Charting the Vision-Interaction Gap in Cluttered Dexterous Grasping

触觉何时重要?描绘杂乱灵巧抓取中的视觉—交互差距
Jiang, Hao, Dominguez, Luis, Seita, Daniel
Abstract
Dexterous grasping in clutter poses a basic sensing question: when do tactile measurements and external wrench estimates improve on visual geometry? Occlusion and contact can obscure grasp quality, motivating a controlled evaluation of these interaction signals. We present a controlled real-world study over five tabletop scene conditions on a dexterous system that combines vision, per-finger and wrist wrench estimates, and distributed fingertip taxels. With demonstrations, visual observations, action space, and compliant control fixed, we compare vision-only, wrench, taxel, and combined policies plus representation and fusion baselines. The combined policy succeeds in 24/25 trials versus 14/25 for vision only, and 15/15 versus 6/15 across the three confined conditions. Ablations show that wrench and taxel feedback are complementary. Behavioral comparisons show that interaction feedback enables earlier rejection of inadequate contacts, regrasping before lift, and more stable grasps. To our knowledge, this is the first real-world study to combine and separately evaluate these interaction modalities for target-oriented dexterous grasping in clutter. These results chart a widening vision-interaction gap and position cluttered dexterous grasping as a benchmark for determining when the learned policy needs interaction sensing. Project website: https://interaction-dex-grasp.github.io/
Chinese Translation
杂乱环境下的灵巧抓取提出了一个基本的感知问题:触觉测量和外部力(旋量)估计何时能优于视觉几何信息?遮挡和接触可能掩盖抓取质量,这促使我们对这些交互信号进行受控评估。我们在一个结合了视觉、逐手指和腕部力估计以及分布式指尖触觉单元(taxel)的灵巧系统上,针对五种桌面场景条件开展了一项受控的真实世界研究。在演示、视觉观测、动作空间和柔顺控制均保持固定的条件下,我们比较了仅视觉、力估计、触觉单元以及组合策略,并加入表征和融合基线方法。组合策略在25次试验中成功24次,而仅视觉策略为14次;在三个受限场景条件中,组合策略达到15/15,而仅视觉策略仅为6/15。消融实验表明,力反馈与触觉单元反馈是互补的。行为比较显示,交互反馈使策略能够更早地拒绝不充分的接触、在提起前重新抓取,并获得更稳定的抓取。据我们所知,这是首个针对杂乱环境中目标导向灵巧抓取,组合并分别评估这些交互模态的真实世界研究。这些结果描绘了日益扩大的视觉—交互差距,并将杂乱灵巧抓取确立为一个基准,用于判定学习到的策略何时需要交互感知。项目网站:https://interaction-dex-grasp.github.io/
cs.RO / 125 / 2609.24093

Dexterous Robot Manipulation from Human Demonstrations via Contact-Anchored Retargeting and Residual Policy Learning

基于接触锚定重定向与残差策略学习的从人类演示学习灵巧机器人操作
Yang, Zihao, Liu, Chengyuan, Zhou, Yu, Lv, Runze, Cui, Tianyu, Yi, Sheng, Zhu, Haohua, Lu, Irvine, Sun, JieQ
Abstract
Learning dexterous manipulation from demonstrations is bottlenecked by data: the contact forces that determine whether a grasp succeeds are absent from every scalable source of human demonstrations. This paper builds on two observations. First, what survives the change from a human hand to a robot hand is the contact structure of a demonstration - which finger regions touch which object locations, and in what order - rather than its joint motion. Second, physical consistency need not be engineered per task: a single residual reinforcement learning (RL) policy, trained once across diverse demonstrations, can repair kinematic recordings into physically consistent, contact-annotated trajectories, and the same residual formulation restores dynamic feasibility after retargeting. These observations yield a three-stage pipeline that converts human motion-capture recordings into dexterous robot policies with no real-robot training data: physics refinement with a simulated MANO hand recovers contacts and forces, contact-anchored retargeting transfers the demonstrated contact structure through an objective independent of hand morphology, and residual policy learning adapts the result to robot actuation. The pipeline reconstructs 25,454 single-hand trajectories (success 7.3% -> 59.3%) and 25 dual-hand tasks (16.0% -> 62.4%) with one shared policy per setting, transfers one human dataset to four morphologically distinct robot hands (+62.4 pp), and executes four contact-rich bimanual tasks on physical hardware with zero real-robot training data.
Chinese Translation
从演示中学习灵巧操作受制于数据瓶颈:决定抓取是否成功的接触力在所有可扩展的人类演示数据源中均不存在。本文基于两个观察。首先,从人手到机器人手的转换中得以保留的是演示的接触结构——哪些手指区域接触物体的哪些位置、以何种顺序接触——而非其关节运动。其次,物理一致性无需针对每个任务单独设计:一个在多样化演示上一次性训练的残差强化学习(RL)策略,能够将运动学记录修复为物理一致的、带接触标注的轨迹,且同样的残差公式可在重定向后恢复动态可行性。基于这些观察,我们提出了一个三阶段流水线,将人类动作捕捉记录转换为灵巧机器人策略,且无需真实机器人训练数据:使用仿真MANO手的物理精化阶段恢复接触与力,接触锚定重定向通过与手形态无关的目标函数迁移演示的接触结构,残差策略学习则将结果适配到机器人的执行器。该流水线重建了25,454条单手轨迹(成功率从7.3%提升至59.3%)和25个双手任务(从16.0%提升至62.4%),每种设置均使用单一共享策略;将一个人类数据集迁移至四种形态迥异的机器人手(提升62.4个百分点);并在零真实机器人训练数据的条件下,在物理硬件上执行了四个接触丰富的双手任务。
cs.RO / 126 / 2609.24099

Ask Before It Tells: Benchmark-to-Robot Body-Cue Transfer for a Question-First Bedside Robot

先询问后告知:从基准数据集到机器人身体线索迁移的“询问优先”式床旁陪伴机器人
Yoon, Dongsik
Abstract
Body-cue recognition can support assistive robots, but benchmark accuracy does not guarantee reliable behavior under a robot-camera viewpoint. We present Nuni, a bedside robot prototype that treats a detected distress cue as a reason to ask rather than a reason to alert. We compare two X3D-UGT RGB appearance classifiers, which reach 97.7% and 94.8% six-way accuracy on NTU RGB+D, with a pose-centric hybrid pipeline on 28 single-actor scripted clips recorded from the robot camera. The hybrid path achieved 0.71 six-way macro recall, versus 0.25 and 0.29 for the fine-tuned and from-scratch RGB variants. More importantly for interaction, it produced a question-triggering distress cue in 12/16 distress clips and would have prompted unnecessarily in 2/8 normal clips; the RGB variants yielded a question-triggering cue in only 2/16 and 3/16 distress clips. We separately tested the question-first controller through event injection. All 13 state-transition trials passed: valid responses caused stand-down, two unanswered prompts produced one alert, and three boundary conditions were handled correctly. These results are a preliminary technical evaluation, not a user study or medical validation, but they show how interaction policy can limit the consequences of uncertain perception.
Chinese Translation
身体线索识别可以支持辅助机器人,但基准测试的准确率并不能保证机器人在自身摄像头视角下具有可靠的行为表现。我们提出了Nuni——一种床旁陪伴机器人原型,它将检测到的痛苦线索视为询问的理由,而非报警的理由。我们在由机器人摄像头录制的28段单人 scripted 片段上,将两种X3D-UGT RGB外观分类器(在NTU RGB+D数据集上分别达到97.7%和94.8%的六分类准确率)与一个以姿态为中心的混合流水线进行了比较。混合流水线取得了0.71的六分类宏平均召回率,而微调和从头训练的RGB变体分别仅为0.25和0.29。更重要的是,从交互角度看,混合流水线在16段痛苦片段中的12段中产生了可触发询问的痛苦线索,且在8段正常片段中仅有2段会造成不必要的询问触发;而RGB变体在16段痛苦片段中仅分别有2段和3段产生可触发询问的线索。我们还通过事件注入单独测试了“询问优先”控制器:全部13次状态转换试验均通过——有效的回应使机器人解除警报,两次未获回应的询问触发一次报警,三个边界条件均被正确处理。这些结果是初步的技术评估,而非用户研究或医学验证,但它们展示了交互策略如何能够限制不确定感知所带来的后果。
cs.RO / 127 / 2609.24118

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies

CARE:面向视觉-语言-动作策略的经验引导原子纠错执行
Xiao, Junlan, Jiang, Junwei, Zhang, Zaibin, Wang, Yifan, Zhang, Zhongbo, Lu, Huchuan, Wang, Lijun
Abstract
Vision-Language-Action (VLA) policies achieve strong performance in robotic manipulation but remain brittle once execution deviates from nominal trajectories. We propose CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution. Instead of generating corrective data from manually designed or random perturbations, CARE collects failed rollouts, models stage-conditioned post-failure deviations, and uses the resulting empirical distributions to synthesize representative failure states and corrective demonstrations. At inference time, CARE combines stage-wise planning with physically grounded 3D monitoring to trigger atomic adjustments or re-operations while preserving task progress. We further introduce the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anomalies. Experiments across multiple VLA backbones, simulation benchmarks, and real-world dual-arm tasks show consistent improvements, with average task-success gains of 14.5 points in simulation and 15.9 points in the real world. Code, models, and data are available at https://github.com/xiaojunlan/care
Chinese Translation
视觉-语言-动作(VLA)策略在机器人操作中取得了优异的性能,但一旦执行偏离标称轨迹,其表现仍然十分脆弱。我们提出CARE(Corrective Atomic Robotic Execution,纠错原子机器人执行),这是一个通过从执行过程中遇到的失败中学习来提升恢复能力的框架。CARE并非通过人工设计或随机扰动来生成纠错数据,而是收集失败的执行轨迹,建模基于阶段的后失败偏差,并利用所得到的经验分布来合成具有代表性的失败状态与纠错示范。在推理阶段,CARE将分阶段规划与基于物理的3D监测相结合,在保持任务进度的同时触发原子化的调整或重新操作。我们还进一步提出了失败状态恢复基准(Failure State Recovery Benchmark,FSR-Bench),用于评估在局部偏差和结构性异常下从中间失败状态的恢复能力。在多种VLA骨干模型、仿真基准以及真实世界双臂任务上的实验表明,该方法带来了一致的性能提升,仿真环境中平均任务成功率达到14.5个百分点的提升,真实环境中达到15.9个百分点。代码、模型和数据可在 https://github.com/xiaojunlan/care 获取。
cs.RO / 128 / 2609.24124

ActiveArena: Benchmarking and Understanding Active Perception in Robotic Manipulation

ActiveArena:机器人操作中主动感知的基准测试与理解
Li, Yibo, Zhou, Enshen, Chen, Rui, Ding, Yanjun, Liu, Mengzhen, Han, Yi, Zhan, Jiabo, Wang, Lipeng, Zhang, Shanghang, Sheng, Lu
Abstract
Active perception and manipulation are crucial for robots to interact with complex scenes. Existing benchmarks struggle to evaluate how robots effectively acquire and maintain information in memory in an active manner. To this end, we introduce ActiveArena-Sim, an active-perception simulator with controllable viewpoints and large-scale workspaces as the foundation. Built on this, we propose ActiveArena-Bench, which comprises 35 tasks across 5 fine-grained categories, covering visual exploration and interactive information acquisition. Each task is difficult to solve from passive observations alone, requiring multi-round evidence acquisition and memory-based reasoning. The benchmark provides rich memory annotations, standardized training data, and ID/OOD protocols featuring disjoint scenes, unseen distractor configurations, and novel backgrounds. Moreover, we present ActiveArena-VLA, a modular suite of 13 vision-language-action configurations for controlled studies of memory writing, memory capacity, proprioceptive state, subtask supervision, and high-level planning in active perception. Benchmark results reveal a substantial ID-OOD gap: uniform memory sampling, increased memory capacity under reliable write policies, proprioceptive inputs, and subtask supervision improve OOD generalization, while planner-guided memory management and decision-making achieve performance close to the best-performing configuration using only sparse memory. ActiveArena thus provides a unified testbed to develop and diagnose models for active perception and manipulation.
Chinese Translation
主动感知与操作对于机器人与复杂场景进行交互至关重要。现有基准测试难以评估机器人如何以主动的方式有效地获取信息并将其保持在记忆中。为此,我们提出了 ActiveArena-Sim,一个具有可控视点和大尺度工作空间的主动感知模拟器,作为研究基础。在此之上,我们构建了 ActiveArena-Bench,包含 5 个细粒度类别下的 35 个任务,涵盖视觉探索和交互式信息获取。每个任务仅凭被动观测难以求解,需要多轮证据获取和基于记忆的推理。该基准提供了丰富的记忆标注、标准化的训练数据,以及包含互斥场景、未见干扰物配置和新颖背景的 ID/OOD(分布内/分布外)评估协议。此外,我们提出了 ActiveArena-VLA,一个包含 13 种视觉-语言-动作(VLA)配置的模块化套件,用于对主动感知中的记忆写入、记忆容量、本体感觉状态、子任务监督和高层规划进行受控研究。基准测试结果揭示了显著的 ID-OOD 差距:均匀的记忆采样、在可靠写入策略下增加记忆容量、引入本体感觉输入以及子任务监督均能提升 OOD 泛化能力;而基于规划器引导的记忆管理与决策,仅使用稀疏记忆即可取得接近最佳配置的性能。ActiveArena 由此为主动感知与操作模型的开发与诊断提供了一个统一的测试平台。
cs.RO / 129 / 2609.24133

Phrase-Level Robotic Guqin Performance: Bimanual Motion Planning and Audio-Tactile Interaction Monitoring

短语级机器人古琴演奏:双臂运动规划与听觉-触觉交互监测
Wang, Zhen, Chen, Zhiheng, Bao, Tianyuan, Zhang, Tianwei
Abstract
Recent advances in humanoid robotics and embodied intelligence have enabled robots to perform increasingly complex manipulation tasks. However, musical instrument performance remains a formidable benchmark, demanding not only collision-free trajectory execution but also precise contact timing, asymmetric bimanual coordination, and target acoustic outcomes on physical instruments. The guqin, a seven-string fretless zither, presents unique manipulation challenges due to its millimetric string spacing, transient right-hand plucking, and sustained left-hand harmonic contacts. In this work, we present a physical heterogeneous dual-arm robotic system for phrase-level autonomous guqin performance. We formulate guqin playing as a hybrid discrete--continuous execution problem and develop a hierarchical planning framework that coordinates working finger assignment, configuration continuity, obstacle avoidance, and tight bimanual contact schedules across consecutive musical events. The system integrates vision-guided instrument localization, tactile-based harmonic contact monitoring, and auditory feedback-informed plucking parameter calibration. Real-world experiments on a 25-event phrase demonstrate that the system reliably executes coordinated open-string and seventh-hui harmonic sequences on a physical guqin, achieving 93.6% and 96.8% event correctness across repeated trials.
Chinese Translation
人形机器人与具身智能的最新进展使机器人能够执行日益复杂的操作任务。然而,乐器演奏仍然是一个极具挑战性的基准任务,它不仅要求无碰撞的轨迹执行,还要求精确的接触时机、非对称的双臂协调,以及在实体乐器上实现目标声学效果。古琴作为一种七弦无品拨弦乐器,由于其毫米级的弦距、右手的瞬时拨弦动作以及左手的持续性泛音接触,带来了独特的操作挑战。在本工作中,我们提出了一套用于短语级自主古琴演奏的异构双臂机器人实体系统。我们将古琴演奏形式化为一个混合离散-连续执行问题,并开发了一种分层规划框架,用以在连续的音乐事件之间协调工作手指分配、构型连续性、避障以及紧凑的双臂接触时间表。该系统集成了视觉引导的乐器定位、基于触觉的泛音接触监测,以及基于听觉反馈的拨弦参数校准。在一个包含25个音乐事件的短语上的真实世界实验表明,该系统能够在实体古琴上可靠地执行协调的空弦与七徽泛音序列,在多次重复试验中分别达到93.6%和96.8%的事件正确率。
cs.RO / 130 / 2609.24140

BayesianGS-SLAM: Uncertainty-Aware Neural Rendering SLAM via Probabilistic Formulation

BayesianGS-SLAM:基于概率化公式的不确定性感知神经渲染SLAM
Kang, Kyeongsu, Ha, Seongbo, Lee, Sibaek, Yu, Hyeonwoo
Abstract
Neural-rendering-based SLAM relies on rendered RGB-D residuals for camera tracking and map optimization, but the reliability of these predictions can vary substantially because of sensor noise, limited observation coverage, and incomplete map representations. Without an explicit reliability estimate, unreliable residuals may adversely affect pose optimization, while frames already well explained by the current map may trigger redundant mapping updates. In this paper, we present BayesianGS-SLAM, an uncertainty-aware 3D Gaussian Splatting SLAM framework that estimates predictive color and depth uncertainty during mapping and consistently reuses it across the SLAM pipeline. Our tractable probabilistic formulation combines a sensor-noise uncertainty component with an opacity-induced map-representation component propagated through the rendering process. The resulting predictive uncertainty is used to augment mapping, normalize tracking residuals through a robust pose objective, and evaluate incoming frames using a predictive-surprise-based keyframe criterion. Unlike prior uncertainty-aware neural-rendering SLAM methods that primarily consider color uncertainty or use uncertainty only during mapping, our framework estimates predictive uncertainty for both color and depth and integrates it into mapping, tracking, and keyframe selection. Evaluations on real-world RGB-D datasets demonstrate substantially improved depth uncertainty-error ranking compared with existing uncertainty-aware SLAM methods. Moreover, the proposed keyframe-selection strategy reduces the number of selected keyframes and mapping calls while maintaining competitive tracking and rendering performance.
Chinese Translation
基于神经渲染的SLAM依赖渲染的RGB-D残差进行相机跟踪与地图优化,但由于传感器噪声、有限的观测覆盖范围以及不完整的地图表示,这些预测的可靠性可能存在很大差异。在缺乏显式可靠性估计的情况下,不可靠的残差可能对位姿优化产生不利影响,而那些已被当前地图充分解释的帧则可能触发冗余的建图更新。本文提出BayesianGS-SLAM,一个不确定性感知的3D高斯泼溅(3D Gaussian Splatting)SLAM框架,该框架在建图过程中估计预测性的颜色和深度不确定性,并在整个SLAM流程中一致地加以复用。我们提出的易处理概率化公式将传感器噪声不确定性分量与经渲染过程传播的不透明度引起的地图表示分量相结合。所得到的预测不确定性被用于增强建图、通过鲁棒位姿目标函数对跟踪残差进行归一化,并基于预测惊奇度(predictive surprise)的关键帧准则评估输入帧。与现有主要考虑颜色不确定性或仅在建图阶段使用不确定性的不确定性感知神经渲染SLAM方法不同,我们的框架同时估计颜色和深度的预测不确定性,并将其集成到建图、跟踪和关键帧选择中。在真实世界RGB-D数据集上的评估表明,与现有不确定性感知SLAM方法相比,本方法在深度不确定性与误差排序方面有显著提升。此外,所提出的关键帧选择策略在保持具有竞争力的跟踪与渲染性能的同时,减少了所选关键帧数量和建图调用次数。
cs.RO / 131 / 2609.24145

MimicAgent: Quadruped Skills via Prompt-to-Trajectory Generation

MimicAgent:基于提示到轨迹生成的四足机器人技能学习
Nayak, Lucky Kant, Parameswaran, Narayanan Palghat, Peri, Neehar, Ramanan, Deva
Abstract
We present MimicAgent, a prompt-to-trajectory generation framework for learning dynamic quadruped skills. Although reward shaping is extensively used when training quadruped policies, navigating the resulting reward landscape is notoriously difficult, requiring hours of "graduate student descent". Eureka attempts to automate reward design with LLMs, but we find that it struggles to generalize across diverse skills and morphologies. Our key observation is that it is far easier for a human - and by association, an LLM - to generate reference motions than to shape reward functions. Our hypothesis is motivated by the success of example-guided RL for humanoids, which exploits large-scale motion capture datasets as references for training locomotion policies. Unlike humanoids, quadrupeds lack such reference motion data. Towards this end, we propose MimicAgent, an agentic harness that, given a skill prompt, generates quadruped reference trajectories with coding agents. These coarse reference trajectories are then used to train example-guided RL policies that are deployable in simulation and in the real-world. Notably, we find that when prompting Claude Fable 5.1 within our agentic harness, 87% of prompts yield semantically aligned reference trajectories.
Chinese Translation
我们提出了 MimicAgent,一个用于学习动态四足机器人技能的提示到轨迹生成框架。尽管奖励塑形在训练四足机器人策略时被广泛使用,但在由此产生的奖励空间中进行导航是出了名的困难,往往需要耗费数小时的“研究生调参”。Eureka 尝试利用大语言模型(LLM)自动化奖励设计,但我们发现它难以在多样的技能和形态之间进行泛化。我们的关键观察是:对人类——因而对 LLM 亦然——而言,生成参考运动远比设计奖励函数容易。我们的假设受到基于示例的强化学习在人形机器人上取得成功的启发,该方法利用大规模动作捕捉数据集作为训练运动策略的参考。与人形机器人不同,四足机器人缺乏此类参考运动数据。为此,我们提出了 MimicAgent,一个智能体化框架,它能够在给定技能提示的情况下,利用编码智能体生成四足机器人的参考轨迹。这些粗糙的参考轨迹随后被用于训练基于示例的强化学习策略,这些策略可部署于仿真和真实世界中。值得注意的是,我们发现当在我们的智能体化框架中提示 Claude Fable 5.1 时,87% 的提示都能产生语义对齐的参考轨迹。
cs.RO / 132 / 2609.24155

Object-Centric Conditioning for Visuomotor Flow Matching

面向物体中心的视觉运动流匹配条件化方法
Li, Jijie, Yang, Xu, Zou, Junhong, Zhao, Chunhai, Zhao, Chaoyang, Lei, Zhen, Zhu, Xiangyu
Abstract
Robot visuomotor policies are commonly formulated as autoregressive, diffusion-based, or more recently, flow matching models. Among them, Action-to-Action (A2A) flow matching improves inference efficiency by initializing generation from historical action priors rather than stochastic noise. However, stale historical motion patterns and entangled global visual representations can jointly reduce robustness under spatial out-of-distribution (OOD) shifts and visual distractors. In this work, we propose SlotFlow, an object-centric flow matching policy for robust visuomotor manipulation. SlotFlow decouples scene observations into semantic ("what") features and lightweight image-plane spatial ("where") cues to provide object-aware policy conditioning and current-state grounding. The semantic representation suppresses irrelevant background correlations, while the spatial cue improves adaptation to shifted object configurations. Extensive simulation and real-world experiments demonstrate improved robustness under visual distractors and severe spatial perturbations while preserving the low-step inference efficiency of A2A. Controlled initialization and perception ablations further identify object-centric grounding as a major source of the gains and show that it complements, rather than replaces, useful historical motion priors.
Chinese Translation
机器人视觉运动策略通常被构建为自回归模型、基于扩散的模型,或近来的流匹配(flow matching)模型。其中,动作到动作(Action-to-Action, A2A)流匹配通过从历史动作先验而非随机噪声出发进行生成初始化,从而提高了推理效率。然而,陈旧的历史运动模式与纠缠的全局视觉表示会共同降低策略在空间分布外(OOD)偏移和视觉干扰下的鲁棒性。在本工作中,我们提出了SlotFlow,一种面向物体中心(object-centric)的流匹配策略,用于实现鲁棒的视觉运动操作。SlotFlow将场景观测解耦为语义("是什么")特征和轻量的图像平面空间("在哪里")线索,以提供物体感知的策略条件化和当前状态锚定。语义表示抑制了无关的背景相关性,而空间线索则提升了对物体构型变化的适应能力。大量仿真与真实世界实验表明,SlotFlow在视觉干扰和严重空间扰动下提升了鲁棒性,同时保持了A2A的低步数推理效率。受控的初始化与感知消融实验进一步证实,物体中心锚定是性能提升的主要来源,并且它是对有用的历史运动先验的补充,而非替代。
cs.RO / 133 / 2609.24180

GraspTune: Tactile-Driven Execution Refinement for Robust Grasping

GraspTune:面向鲁棒抓取的触觉驱动执行精化方法
Li, Juntao, Xia, Xingke, Liu, Sichao, Guo, Daqiang
Abstract
Visual grasp proposal generation has advanced rapidly, yet converting a selected proposal into a stable physical grasp remains a central execution-stage challenge. This paper introduces GraspTune, a tactile-driven execution-stage refinement framework that starts from a nominal proposal and applies bounded residual TCP motions during approach, contact formation, and final grasp execution. GraspTune learns control-facing contact semantics from local depth, tactile signals, state, and history using state-conditioned expert contact queries and multi-task supervision for contact change, contact risk, and post-close readiness. The representation conditions a diffusion-pretrained residual policy and is aligned with PPO for closed-loop execution. Across more than 60,000 simulated executions over 20 object categories, GraspTune establishes an execution-layer benefit across four proposal generators, raising stable grasp success by +19.22, +9.55, +12.45, and +20.70 percentage points for GraspNet, Contact-GraspNet, AnyGrasp, and VGN. A four-fold held-out category study raises unseen-object execution from 54.58% to 70.33%, showing category-disjoint generalization of contact correction. Across more than 1,000 real-robot trials on a UR5e setup with Xense fingertip sensors, GraspTune raises GraspNet execution from 71.0% to 84.3%, validating direct transfer without realworld policy fine-tuning. Together, these results turn visually plausible proposals into stable physical grasps for downstream contact-rich manipulation. A supplementary video is available at https://youtu.be/kcq7fSLNtzU.
Chinese Translation
视觉抓取候选框生成技术已取得快速进展,然而将选定的候选方案转化为稳定的物理抓取仍然是执行阶段的核心挑战。本文提出GraspTune,一个触觉驱动的执行阶段精化框架,该框架从标称抓取方案出发,在接近、接触建立和最终抓取执行过程中施加有界残差TCP(工具中心点)运动。GraspTune利用状态条件化的专家接触查询以及针对接触变化、接触风险和闭合后就绪度的多任务监督,从局部深度、触觉信号、状态和历史信息中学习面向控制的接触语义。该表征对扩散预训练的残差策略进行条件化,并与PPO对齐以实现闭环执行。在涵盖20类物体、超过60,000次仿真执行中,GraspTune在四种抓取候选生成器上均展现出执行层的增益,使GraspNet、Contact-GraspNet、AnyGrasp和VGN的稳定抓取成功率分别提升19.22、9.55、12.45和20.70个百分点。四折留出类别研究表明,未见物体的执行成功率从54.58%提升至70.33%,证明了接触修正在类别不重叠设定下的泛化能力。在配备Xense指尖传感器的UR5e机器人平台上超过1,000次真实机器人实验中,GraspTune将GraspNet的执行成功率从71.0%提升至84.3%,验证了无需真实环境策略微调的直接迁移能力。综上,这些结果将视觉上合理的抓取方案转化为稳定的物理抓取,服务于下游富接触操作任务。补充视频见 https://youtu.be/kcq7fSLNtzU。
cs.RO / 134 / 2609.24187

StenoVLA-3D: 3D-Aware Reasoning VLA for Navigation Through Gastrointestinal Stenoses

StenoVLA-3D:面向胃肠道狭窄部位导航的3D感知推理视觉-语言-动作模型
Tabassum, Tamima, Huang, Yiming, Wu, Tianchun, Liu, Changjing, Tang, Zhiqing, Ng, Chikit, Cui, Beilei, Shao, Liangjing, Lai, Jiewen, Ren, Hongliang
Abstract
Autonomous endoscopic navigation requires the policy model to predict actions from texture-poor monocular observations, make safe control decisions, and retain evidence of lesions after they leave the field of view. Existing vision-language-action (VLA) models primarily rely on visual appearance and short-term context, limiting geometric grounding and episode-level reporting. We introduce StenoVLA-3D, a 3D-aware VLA framework for navigating through stenotic regions. We integrate point-maps into the Cosmos-Reason 2 backbone through learned geometry-gated fusion, and also propose a temporal state branch to model traversal progress. Our reasoning-and-action backbone predicts grounded reasoning with actions, while dedicated heads estimate stenosis shape and generate the final lesion report. We further introduce EndoCausal, an episode-level dataset with lesion annotations, actions, and temporally grounded reasoning. On 40 held-out recorded test episodes, StenoVLA-3D reaches 95.2\% semantic accuracy and 83.4\% action accuracy. On the physical 3-DoF endoscope, it attains 88.9\% and 77.8\% task success in esophageal and colonic phantoms (36 trials each), substantially outperforming the evaluated baselines.
Chinese Translation
自主内镜导航要求策略模型能够从纹理贫乏的单目观测中预测动作,做出安全的控制决策,并在病灶离开视野后仍保留其证据。现有的视觉-语言-动作(VLA)模型主要依赖视觉外观和短期上下文,限制了其几何基础能力和对完整任务过程的报告能力。我们提出了StenoVLA-3D,一个用于穿越狭窄区域的3D感知VLA框架。我们通过学习的几何门控融合机制将点图(point-maps)集成到Cosmos-Reason 2骨干网络中,并提出一个时序状态分支来建模行进进度。我们的推理与动作骨干网络在预测有据可依的推理的同时预测动作,并通过专门的预测头估计狭窄形状并生成最终的病灶报告。我们进一步构建了EndoCausal,一个包含病灶标注、动作和时序对齐推理的任务级数据集。在40个预留的录制测试任务中,StenoVLA-3D达到95.2%的语义准确率和83.4%的动作准确率。在物理三自由度内镜上,它在食管和结肠仿体中分别取得88.9%和77.8%的任务成功率(各36次试验),显著优于所评估的基线方法。
cs.RO / 135 / 2609.24189

A Topological Representation with Object-Path Graphs for Open-Vocabulary Instance Navigation

一种基于对象-路径图的开放式词汇实例导航拓扑表示方法
Zheng, Linwei, Peng, Daojie, Wang, Bingtao, Li, Haoang, Ma, Jun
Abstract
Vision-language navigation requires embodied agents to navigate environments using natural language instructions and visual observations. Existing approaches typically decompose navigation into sequential language-guided decisions or rely on online exploration without prior environmental knowledge. Scene graph representations offer compact semantic memory but remain decoupled from downstream navigation, which still depends on dense metric maps. To close this gap, we propose an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation. The proposed representation jointly supports semantic grounding, graph-based localization, and navigation within a single lightweight topological framework. Building on this graph, we introduce a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly over the graph without dense metric reconstruction. Experiments on HM3D and Replica demonstrate competitive performance in open-vocabulary object grounding through the proposed hierarchical graph structure, while achieving effective navigation performance. Real-world robot experiments further validate the practicality of the proposed framework.
Chinese Translation
视觉语言导航要求具身智能体利用自然语言指令和视觉观测在环境中进行导航。现有方法通常将导航分解为序列化的语言引导决策,或依赖缺乏先验环境知识的在线探索。场景图表示提供了紧凑的语义记忆,但与下游导航仍然脱节,后者依旧依赖稠密的度量地图。为弥合这一差距,我们提出了一种对象-路径图(object-path graph),将开放式词汇语义推理与拓扑导航统一起来。所提出的表示在单一轻量级拓扑框架内同时支持语义定位、基于图的定位以及导航。基于该图,我们提出了一种导航策略,通过轻量级节点定位和语义视觉伺服,将全局路径规划与局部的节点间执行相结合,从而能够直接在图上进行导航,而无需稠密的度量重建。在 HM3D 和 Replica 数据集上的实验表明,通过所提出的层次化图结构,该方法在开放式词汇对象定位方面取得了具有竞争力的性能,同时实现了有效的导航表现。真实世界机器人实验进一步验证了所提出框架的实用性。
cs.RO / 136 / 2609.24195

Odometry-Aided Real-Time Mapping for Underwater Robots Using Forward-Looking Sonar

利用前视声呐的里程计辅助水下机器人实时建图
Du, Siyuan, Yao, Kanzhong, Wang, Youdong, Liu, Yingqi, Liu, Qingwen, Yang, Qunhui, Sun, Zhe, Li, Xuelong
Abstract
Reliable perception is essential for underwater vehicles operating in complex environments, where light attenuation and scattering often degrade visibility and compromise optical sensing. Forward-looking sonar (FLS) offers an alternative by providing high-frame-rate acoustic imaging under poor optical conditions. However, real-time FLS mapping remains challenging due to unresolved target elevation, spatially non-uniform noise, and fragmented target boundaries, which hinder feature extraction and introduce geometric ambiguity during projection. To address these challenges, we propose a cascaded feature reconstruction pipeline combining fast Fourier transform (FFT)-based denoising, fast multiscale constant false alarm rate (MCFAR) detection, and gradient-adaptive boundary connection to extract geometric features from degraded sonar images with low latency. We integrate attitude-aware geometric projection with incremental occupancy accumulation to construct a depth-referenced 2.5D map for local mapping in confined underwater environments. The sonar's vertical position is referenced to an external sensor, while target elevation is assigned under an explicit geometric assumption rather than measured directly by FLS. Experiments in a 3 m X 5 m pool demonstrate centimeter-scale planar mapping accuracy, with a root-mean-square error (RMSE) below 3 cm across three sequences and an average processing time of 42.4 ms per frame.
Chinese Translation
可靠感知对于在复杂环境中作业的水下航行器至关重要。在这些环境中,光的衰减与散射常导致能见度下降,削弱光学传感能力。前视声呐(Forward-Looking Sonar, FLS)能够在光学条件恶劣的情况下提供高帧率声学成像,因而成为一种替代方案。然而,由于目标高度信息缺失、空间上非均匀的噪声以及目标边界的碎片化,实时FLS建图仍然极具挑战性,这些问题阻碍了特征提取并在投影过程中引入几何歧义。为应对这些挑战,我们提出了一种级联特征重建流水线,结合基于快速傅里叶变换(FFT)的降噪、快速多尺度恒虚警率(MCFAR)检测以及梯度自适应边界连接,以低延迟从退化的声呐图像中提取几何特征。我们将姿态感知的几何投影与增量占据栅格累积相结合,构建了以深度为参考的2.5D地图,用于封闭水下环境的局部建图。声呐的垂直位置由外部传感器提供参考,而目标高度则在明确的几何假设下进行赋值,而非由FLS直接测量。在3米×5米水池中的实验表明,该系统实现了厘米级的平面建图精度:在三个序列上均方根误差(RMSE)低于3厘米,平均每帧处理时间为42.4毫秒。
cs.RO / 137 / 2609.24218

Audio-based UAV Localization with Adaptive Temporal Correspondence via Reinforcement Learning

基于强化学习的自适应时序对应的音频无人机定位
Lei, Haoxiang, Feng, Mingzheng, Wang, Daotong, Yuan, Shenghai
Abstract
Audio-based localization provides a low-cost and illumination-independent sensing solution for anti-UAV early warning. However, existing methods typically rely on a predefined fixed audio segment length, which limits temporal correspondence and creates a trade-off between sufficient acoustic evidence and timely localization. To address this issue, we propose an audio-based localization framework with adaptive temporal correspondence. A probe segment is first used to extract a compact acoustic state that characterizes the reliability and consistency of the observation. Guided by the state, a reinforcement learning controller dynamically determines the required audio window size for each localization decision. The selected audio segment is then processed by a Mamba-based localization network with adaptive temporal feature modulation for 3D position estimation. Extensive experiments demonstrate that our method achieves competitive 3D localization accuracy with substantially reduced temporal correspondence latency compared to SOTA methods and exhibits strong generalization across scenarios.
Chinese Translation
基于音频的定位为反无人机预警提供了一种低成本且不依赖光照条件的感知方案。然而,现有方法通常依赖于预定义的固定音频段长度,这限制了时序对应能力,并在充分的声学证据与及时的定位之间形成了权衡。为解决这一问题,我们提出了一种具有自适应时序对应能力的音频定位框架。首先利用一个探测段提取紧凑的声学状态,以刻画观测的可靠性与一致性;在该状态的引导下,一个强化学习控制器动态确定每次定位决策所需的音频窗口大小;随后,所选音频段由基于 Mamba 的定位网络处理,通过自适应时序特征调制实现三维位置估计。大量实验表明,与最先进(SOTA)方法相比,我们的方法在大幅降低时序对应延迟的同时取得了具有竞争力的三维定位精度,并在不同场景下展现出强大的泛化能力。
cs.RO / 138 / 2609.24253

OpenFlyScan: A Quality-Guided Aerial Reconstruction System for Consumer Drones

OpenFlyScan:面向消费级无人机的质量引导式航拍重建系统
You, Zhongrui, Li, Zhen, Liu, Junli, Wang, Zhigang, Zhao, Bin
Abstract
3D Gaussian Splatting (3DGS) provides high-fidelity scenes for large-scale embodied simulation, but constructing large-scale urban assets remains constrained by expensive equipment and delayed quality feedback. Preset surveys can leave complex surfaces insufficiently observed, with defects discovered only after reconstruction, requiring return visits and repeated processing. We present OpenFlyScan, a quality-guided aerial reconstruction system for consumer drones that integrates a GS quality model, a reacquisition planner, and a custom-designed mobile app. The model learns from GS rendering errors to predict regional reconstruction quality. Based on these predictions, the planner then generates complementary reacquisition strips to be executed through the app, which also supports automated oblique surveys and data transfer without additional hardware on board. Across real aerial scenes, the model effectively identifies regions that are likely to be poorly reconstructed. In the Expo West field experiment, targeted reacquisition improves PSNR at additional views by 10.95 dB. With consumer drones, OpenFlyScan integrates capture, targeted reacquisition, and reconstruction to support rapid, low-cost urban asset creation. Code and models will be made publicly available at https://openflyscan.github.io/.
Chinese Translation
三维高斯泼溅(3D Gaussian Splatting,3DGS)能够为大规模具身仿真提供高保真场景,但构建大规模城市资产仍受限于昂贵的设备与滞后的质量反馈。预设的航线采集可能导致复杂表面观测不足,其缺陷往往在重建完成后才被发现,从而需要返场采集和重复处理。我们提出了 OpenFlyScan,一个面向消费级无人机的质量引导式航拍重建系统,它集成了一个 GS 质量模型、一个补采规划器以及一个自主设计的移动应用。该模型从 GS 渲染误差中学习,以预测各区域的重建质量。基于这些预测,规划器生成互补性的补采航线,并通过移动应用执行;该应用还支持自动化倾斜摄影采集与数据传输,且无需在无人机上搭载额外硬件。在多个真实航拍场景中,该模型能够有效识别可能重建质量较差的区域。在西博会(Expo West)实地实验中,针对性补采使额外视角的 PSNR 提升了 10.95 dB。借助消费级无人机,OpenFlyScan 将采集、针对性补采与重建整合于一体,支持快速、低成本的城市资产创建。代码与模型将在 https://openflyscan.github.io/ 公开发布。
cs.RO / 139 / 2609.24271

ME-Brain-1.0: Memory, Cognition and Action for Evolving Embodied Intelligence

ME-Brain-1.0:面向进化具身智能的记忆、认知与行动
He, Wei, Li, Hengtao, Yu, Zhongrui, Zhu, Xuhan, He, Maokui, Liu, Zide, Zhang, Xiyue, Mao, Xianwei, Zhou, Chunpeng, Shi, Jia, Xin, Yanze, Li, Jingwen, Zheng, Jingxie, Zeng, Sijie, Wang, Chenfeng, Lu, Fan, Zhang, Zeyu, Guo, Shuai, Zhang, Hengxuan, Yu, Pengfei, Shi, Jia, Liu, Yu, Zhan, Kun, Xie, Yan
Abstract
Current embodied systems largely rely on pretrained capabilities that remain fixed after deployment, limiting their ability to learn from physical interaction. We introduce MachEmbodied-Brain (ME-Brain), a self-evolving embodied system organized around a closed loop of action execution, experience acquisition, experience evolution, and improved execution. Evolvable Memory consolidates multimodal trajectories into hierarchical, reusable experience; Cognitive Core transforms physical experience into transferable skills; and the Action Model combines event-driven keyframes, EventCell local-world prediction, and action-conditioned memory modulation to focus computation on decision-critical moments, regions, and historical evidence. Together, these modules shift embodied intelligence from train-and-freeze to deploy-and-evolve without model retraining. Cognitive Core outperforms the strongest comparison models by 8.2 and 9.6 points on embodied and agent benchmarks. The Action Model achieves 47.88% mean success on RoboMME, a 3.26-point improvement over the strongest baseline. On RoboDojo, it reaches a 21.51 mean Score and 16.03% success rate, exceeding $\pi_{0.5}$ by 10.10 and 9.12 points. On the six-task ME-RealBench, ME-Brain achieves a 69.5 mean Score and 66.7% success rate, outperforming DM0.5 by 12.8 and 11.7 points, respectively.
Chinese Translation
当前的具身系统主要依赖预训练能力,这些能力在部署后保持固定,限制了其从物理交互中学习的能力。我们提出了MachEmbodied-Brain(ME-Brain),这是一个自进化的具身系统,围绕行动执行、经验获取、经验进化和执行改进的闭环组织而成。可进化记忆(Evolvable Memory)将多模态轨迹整合为分层的、可复用的经验;认知核心(Cognitive Core)将物理经验转化为可迁移的技能;行动模型(Action Model)结合事件驱动的关键帧、EventCell局部世界预测以及行动条件化的记忆调节,将计算集中于决策关键的时刻、区域和历史证据。这些模块共同推动具身智能从“训练后冻结”转变为“部署后进化”,且无需重新训练模型。认知核心在具身基准和智能体基准上分别超越最强对比模型8.2分和9.6分。行动模型在RoboMME上取得47.88%的平均成功率,较最强基线提升3.26个百分点。在RoboDojo上,它达到21.51的平均分数和16.03%的成功率,分别超过π₀.₅ 10.10分和9.12个百分点。在包含六个任务的ME-RealBench上,ME-Brain取得69.5的平均分数和66.7%的成功率,分别领先DM0.5 12.8分和11.7个百分点。
cs.RO / 140 / 2609.24274

vla.simd: Efficient CPU Inference for Language-Conditioned Manipulation

vla.simd:面向语言条件操控任务的高效 CPU 推理
Nguyen, Khanh D., Truong, Hoang M., Le, An T.
Abstract
Deploying language-conditioned manipulation without a dedicated GPU requires efficient inference and action chunks that cover the delay between policy queries. We present vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization. We relate query latency and execution horizon to action availability under lagged and time-aligned execution, distinguishing action supply from feedback frequency. Across six policies and four CPUs, vla.simd achieves approximately $1.4\times$ median speedup over compiled PyTorch references while preserving fp32 numerical fidelity. We also introduce IMPACT, an ACT-based policy with cached text representations and language-modulated visual features. IMPACT is the only language-conditioned policy in our evaluated set that supplies at least 30 actions/s on the Raspberry Pi 5: after a 90 s thermal soak, it supplies 33.5 actions/s in fp32 and 81.2 with int8. Separate GPU evaluations yield $76.4\%$ mean success across four LIBERO suites without robot pretraining; instruction-shuffling tests demonstrate selection among familiar goals. Trials with IMPACT on an SO-101 arm and SmolVLA on a UR10e with a Robotiq gripper demonstrate CPU deployment on two robot embodiments.
Chinese Translation
在没有专用 GPU 的条件下部署语言条件操控(language-conditioned manipulation)系统,需要高效的推理以及能够覆盖策略查询间隔延迟的动作块。我们提出了 vla.simd,一种结合共享 SIMD 微内核、可复用计算和针对特定硬件优化的 CPU 推理引擎。我们分析了在滞后执行和时间对齐执行两种模式下,查询延迟与执行视野(execution horizon)同动作可用性之间的关系,并区分了动作供给与反馈频率。在六种策略和四款 CPU 上的实验表明,vla.simd 相较于编译后的 PyTorch 参考实现取得了约 $1.4\times$ 的中位数加速,同时保持了 fp32 的数值精度。我们还提出了 IMPACT,一种基于 ACT、采用缓存文本表示和语言调制视觉特征的策略。在本文评估的策略集合中,IMPACT 是唯一能在 Raspberry Pi 5 上以不低于每秒 30 个动作的速率供给动作的语言条件策略:经过 90 秒的热稳定后,它在 fp32 精度下供给 33.5 个动作/秒,在 int8 精度下达到 81.2 个动作/秒。独立的 GPU 评估显示,在未经机器人预训练的情况下,IMPACT 在四个 LIBERO 测试套件上的平均成功率为 $76.4\%$;指令乱序测试证明了其能够在熟悉的目标之间进行正确选择。IMPACT 在 SO-101 机械臂上以及 SmolVLA 在配备 Robotiq 夹爪的 UR10e 上的实际试验,展示了在两种机器人本体上的 CPU 部署能力。
cs.RO / 141 / 2609.24317

Performance-Preserving Online Adaptation in Social Navigation via Diffusion Steering

基于扩散引导的社交导航中性能保持的在线自适应方法
Nagahisa, Haruto, Matsumoto, Kohei, Hyodo, Yuki, Kurazume, Ryo
Abstract
In social navigation, modeling the complex interactions between humans and robots is difficult, and deep reinforcement learning has therefore been actively studied. However, because simulation alone cannot fully reproduce diverse scenarios, robot dynamics, and the social conventions that vary across deployment environments, fine-tuning in the deployment environment is promising. In doing so, learning that preserves the base model's performance is required, so as not to compromise the primary objective of navigation, namely avoiding pedestrians and reaching the destination. In this study, we propose a method that applies diffusion steering via reinforcement learning (DSRL), which trains only the noise policy while keeping the diffusion policy fixed, thereby achieving learning that preserves performance. Furthermore, we integrate diffusion-based RL policies trained with multiple seeds to construct the base policy, improving learning performance. Our evaluation shows that, compared with other methods, the proposed method enables efficient learning while preserving performance, and we confirm flexible behavior control through adaptation to social conventions, as well as its effectiveness on a physical robot through hardware-in-the-loop simulation.
Chinese Translation
在社交导航中,对人类与机器人之间的复杂交互进行建模十分困难,因此深度强化学习受到了广泛研究。然而,由于仅依靠仿真无法完全复现多样的场景、机器人动力学以及随部署环境而变化的社交规范,在部署环境中进行微调是一种有前景的方法。在此过程中,需要能够在保持基础模型性能的前提下进行学习,以免损害导航的首要目标,即避让行人并到达目的地。本研究提出一种应用扩散引导强化学习(DSRL)的方法,该方法仅训练噪声策略而保持扩散策略固定,从而实现性能保持的学习。此外,我们整合了使用多个随机种子训练的基于扩散的强化学习策略来构建基础策略,以提升学习性能。评估结果表明,与其他方法相比,所提方法能够在保持性能的同时实现高效学习;我们还通过对社交规范的自适应验证了灵活的行为控制,并通过硬件在环仿真确认了其在物理机器人上的有效性。
cs.RO / 142 / 2609.24350

LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models

LIBERO-VPro:机器人基础模型闭环视觉鲁棒性基准测试
Li, Huiqiong, Mei, Zhiting, Majumdar, Anirudha, Chen, Jingjing, Jiang, Yu-Gang, Zhu, Bin
Abstract
Robotic foundation models achieve impressive performance on standard manipulation benchmarks, yet these evaluations typically assume clean, timely, and consistent visual observations throughout execution. We introduce LIBERO-VPro, a benchmark for systematically evaluating the closed-loop visual robustness of robotic foundation models by perturbing the visual evidence available during execution. LIBERO-VPro covers four complementary dimensions, including Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, and Task-Relevant Scene Variation, spanning 12 challenge categories, 96 experimental settings, and 3,296 task-condition cases. We evaluate three vision-language-action models and three world-action models over approximately 196,000 simulated episodes, complemented by 200 real-world rollouts on a Franka Research 3. Our results reveal that strong nominal performance can mask substantial weaknesses in visual grounding and adaptation. Models often remain successful despite severe object-level occlusion, yet degrade sharply when local interaction cues are disrupted or familiar spatial priors are violated. They are also highly sensitive to stale or missing observations and struggle when changed task preconditions require behavioral adaptation. Finally, VLAs and WAMs exhibit distinct robustness profiles, showing that visual robustness is multi-dimensional and architecture-dependent. LIBERO-VPro provides a systematic diagnostic framework for developing robotic foundation models that can more reliably ground and adapt their actions under challenging visual conditions.
Chinese Translation
机器人基础模型在标准操作基准测试中取得了令人瞩目的性能,然而这些评估通常假设在整个执行过程中视觉观测是干净、及时且一致的。我们提出了LIBERO-VPro,这是一个通过扰动执行期间可用的视觉证据来系统性评估机器人基础模型闭环视觉鲁棒性的基准。LIBERO-VPro涵盖四个互补维度,包括视觉证据退化(Visual Evidence Degradation)、相机滞后(Camera Staleness)、视觉源一致性(Visual Source Consistency)和任务相关场景变化(Task-Relevant Scene Variation),共包含12个挑战类别、96种实验设置和3,296个任务条件案例。我们在约196,000次仿真回合中评估了三个视觉-语言-动作模型(VLA)和三个世界-动作模型(WAM),并在Franka Research 3机器人上补充了200次真实世界测试。我们的结果表明,强大的标称性能可能掩盖视觉接地和适应能力方面的重大缺陷。模型在严重的物体级遮挡下仍常能保持成功,但当局部交互线索被破坏或熟悉的空间先验被违背时,性能会急剧下降。模型对滞后或缺失的观测也高度敏感,并且在变化的任务前提条件需要行为适应时表现挣扎。最后,VLA和WAM表现出明显不同的鲁棒性特征,表明视觉鲁棒性是多维度且依赖于架构的。LIBERO-VPro为开发能够在具有挑战性的视觉条件下更可靠地进行视觉接地和动作适应的机器人基础模型提供了系统性的诊断框架。
cs.RO / 143 / 2609.24376

Multi-Agent Transportation of Free-Flyers in Microgravity Via Pushing Interaction Under Human-in-the-Loop Control

人在回路控制下基于推挤交互的微重力自由飞行器多智能体运输
Marchesini, Gregorio, De Carli, Nicola, Cho, Sihyun, Kong, Youngkyoung, Krantz, Elias, Dhullipalla, Mani Hemanth, Dimarogonas, Dimos V., Kim, H. Jin
Abstract
We propose a safety-critical framework for the cooperative transportation of passive targets in microgravity, where a team of chaser robots acts through unilateral pushing contacts to track a human-provided desired twist while ensuring safe target motion. The pushing-only nature of the interaction introduces sparse, configuration-dependent actuation constraints requiring chasers to physically relocate on the target body when the desired pushing allocation changes. To address these challenges, we formulate a delay-aware feedback control architecture leveraging Control Lyapunov Function (CLF) and Control Barrier Function (CBF) constraints within a mixed-integer thrust allocation program to enforce stability and safety of the target, respectively. The proposed framework enables reference tracking while guaranteeing obstacle avoidance with a circular obstacle despite intermittent control authority, providing a foundation for human-supervised cooperative transportation of free-flyers in space environments. The proposed framework is validated through Gazebo simulations.
Chinese Translation
我们提出了一种用于微重力环境下被动目标协同运输的安全关键框架,其中一组追踪机器人通过单边推挤接触进行操作,以跟踪人类给定的期望扭转速度,同时确保目标运动的安全性。仅能推挤的交互特性引入了稀疏且依赖构型的驱动约束,要求追踪机器人在期望推力分配发生变化时在目标本体上进行物理重新定位。为应对这些挑战,我们构建了一种考虑时滞的反馈控制架构,利用控制李雅普诺夫函数(CLF)和控制障碍函数(CBF)约束,并结合混合整数推力分配程序,分别保证目标的稳定性与安全性。该框架能够在控制权限间歇性存在的情况下实现参考轨迹跟踪,并保证对圆形障碍物的避障,为空间环境中人类监督下的自由飞行器协同运输奠定了基础。该框架已通过Gazebo仿真进行了验证。
cs.RO / 144 / 2609.24385

Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors

Tactile-JEPA:面向分布式触觉传感器的拓扑感知自监督表示学习
Kovtun, Elizaveta, Konovalov, Matvey, Sakhovskiy, Andrey, Budennyy, Semen
Abstract
Tactile sensing is an essential modality for robots performing contact-rich, dexterous manipulation, particularly under visual occlusion. While pre-trained image encoders are standard in robot learning pipelines, tactile encoders are still commonly trained from scratch from raw, noisy signals, which might limit their expressivity. Existing self-supervised learning (SSL) approaches focus predominantly on vision-based tactile sensors, leaving distributed electronic skins largely unaddressed. These sensors, however, have a distinctive property: their sensing elements are sparse and irregularly arranged over the surface they cover, which makes direct reuse of visual SSL methods suboptimal. We present Tactile-JEPA, an efficient self-supervised pre-training method that uses the spatial arrangement of tactile sensors to learn topology-aware representations. Specifically, it is trained to predict the embeddings of masked sensing elements from the unmasked remainder, using the sensor connectivity graph to guide spatial masking. Our analysis shows that effective tactile representations require capturing both local contact details and the global state of the tactile surface, which we achieve through dual-scale masking. Across three diverse datasets spanning magnetic and piezoresistive sensors, different robot embodiments, and single- and paired-sensor configurations, Tactile-JEPA reduces force estimation error by 6.3% and in-hand orientation error by 20.8% over the prior state-of-the-art, with consistent gains in other downstream applications, including policy learning. Overall, our results demonstrate that the benefit of tactile sensing depends critically on the quality of encoder pre-training, a problem which Tactile-JEPA addresses directly. Code is available at https://github.com/E-Kovtun/tactile.
Chinese Translation
触觉感知是机器人执行接触丰富、灵巧操作的一项重要模态,尤其是在视觉被遮挡的情况下。虽然预训练图像编码器已成为机器人学习流程的标准配置,但触觉编码器通常仍需要从原始的、含噪声的信号中从头训练,这可能限制其表达能力。现有的自监督学习(SSL)方法主要面向基于视觉的触觉传感器,而对分布式电子皮肤则基本未予涉及。然而,这类传感器具有一个独特性质:其传感元件在其覆盖表面上稀疏且不规则地分布,这使得直接复用视觉SSL方法并非最优。我们提出了Tactile-JEPA,一种高效的自监督预训练方法,它利用触觉传感器的空间布局来学习拓扑感知的表示。具体而言,该方法通过未遮蔽的其余部分来预测被遮蔽传感元件的嵌入,并利用传感器连接图来指导空间遮蔽。我们的分析表明,有效的触觉表示需要同时捕捉局部接触细节和触觉表面的全局状态,我们通过双尺度遮蔽来实现这一点。在涵盖磁性传感器与压阻传感器、不同机器人本体以及单传感器与双传感器配置的三个多样化数据集上,Tactile-JEPA 相较于先前的最先进方法将力估计误差降低了6.3%,手中物体朝向估计误差降低了20.8%,并在包括策略学习在内的其他下游应用中也取得了一致的提升。总体而言,我们的结果表明,触觉感知的收益在很大程度上取决于编码器预训练的质量,而Tactile-JEPA正是直接解决这一问题的方法。代码可在 https://github.com/E-Kovtun/tactile 获取。
cs.RO / 145 / 2609.24411

Zeva-Ego: Egocentric Mid-Training with In-Context Causal Learning for Robot Manipulation

Zeva-Ego:面向机器人操作的基于上下文因果学习的自我中心中期训练
Huang, Bingjia, Ding, Xin, Chen, Fu, Li, Kun, Sun, Wei, Wu, Hao, Liu, Yunxin, Cao, Ting
Abstract
Egocentric video offers a scalable source of physical interaction experience, yet translating it into robot-executable knowledge and enabling continual adaptation remain challenging. We introduce Zeva-Ego, a unified framework that learns physical priors from human experience and evolves through robot interaction. An Action-Centric Encoder (ACE) converts egocentric visual transitions into action-centered supervision for VLA mid-training, while In-Context Causal Learning (ICCL) enables parameter-free adaptation from action-effect feedback at deployment. Scaling Ego data to 10K hours improves RoboTwin success from 63.8% to 75.3%, matching 2K hours of robot demonstrations (74.7%), corresponding to an empirical data ratio of roughly 4-5:1. With accumulated interaction experience, ICCL further improves success from 58% to 89% within four attempts without parameter updates. These results demonstrate a scalable path toward embodied intelligence that learns from human experience and continuously improves through its own interaction.
Chinese Translation
自我中心(第一人称)视频为物理交互经验提供了一种可扩展的数据来源,然而如何将其转化为机器人可执行的知识并实现持续适应仍具挑战性。我们提出了 Zeva-Ego,一个从人类经验中学习物理先验并通过机器人交互不断演进的统一框架。动作中心编码器(Action-Centric Encoder, ACE)将自我中心视觉转变转化为以动作为核心的监督信号,用于视觉-语言-动作(VLA)模型的中期训练;而上下文因果学习(In-Context Causal Learning, ICCL)则使模型在部署时能够基于动作-效果反馈实现无参数更新 adaptation。将 Ego 数据扩展至 1 万小时,可将 RoboTwin 的成功率从 63.8% 提升至 75.3%,与 2000 小时机器人示教数据(74.7%)的效果相当,对应约 4-5:1 的经验数据比例。借助积累的交互经验,ICCL 在四次尝试内将成功率从 58% 进一步提升至 89%,且无需参数更新。这些结果表明了一条可扩展的具身智能路径:从人类经验中学习,并通过自身交互持续改进。
cs.RO / 146 / 2609.24413

Robotic Valve Turning: Axial Misalignment Correction Using Reaction Torque Feedback

机器人阀门旋转:基于反作用力矩反馈的轴向偏差校正
Kumar, Amit, Turlapati, Sri Harsha, Golani, Gautami, Lin, Yang, Banavar, Ravi N., Campolo, Domenico
Abstract
In this work, we propose a haptic update control law that uses reaction torques to correct axial misalignment during robotic valve manipulation. Unlike vision-based estimates, which can be affected by calibration errors, occlusion, and uncertainty in the contact geometry, reaction torques arise directly from the physical interaction between the gripper and valve. A geometric relationship exists between the error (misalignment) vector and these torques. The primary aim of this work is to propose a stable controller exploiting this geometric property. Our control law is proven to be uniformly asymptotically stable. Simulations are performed for verification. Furthermore, we experimentally test the robustness of our method using a Kinova Gen3 robotic arm for initial misalignments ranging from $-15^\circ$ to $15^\circ$ at 3 different valve positions and report the resulting data distribution. The absolute value of the median misalignment across all 18 test cases is found to be within $2.46^\circ$ and that of reaction torques within $0.23\mathrm{Nm}$.
Chinese Translation
在本工作中,我们提出了一种触觉更新控制律,利用反作用力矩来校正机器人阀门操作过程中的轴向偏差。与基于视觉的估计不同——后者可能受到标定误差、遮挡以及接触几何不确定性等因素的影响——反作用力矩直接来源于夹爪与阀门之间的物理交互。误差(偏差)向量与这些力矩之间存在几何关系。本工作的主要目标是提出一种利用该几何特性的稳定控制器。我们证明了所提出的控制律是一致渐近稳定的,并通过仿真进行了验证。此外,我们使用Kinova Gen3机械臂,针对3个不同阀门位置上$-15^\circ$至$15^\circ$范围内的初始偏差,实验测试了该方法的鲁棒性,并报告了所得的数据分布。在全部18个测试用例中,偏差中位数的绝对值均在$2.46^\circ$以内,反作用力矩中位数的绝对值均在$0.23\mathrm{Nm}$以内。
cs.RO / 147 / 2609.24433

FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding

FoldQuantVLA:基于一致折叠的视觉-语言-动作模型原生低比特量化
Ho, Hung T., Nguyen, Khanh D., Nguyen, Quang D., Duong, Thanh Q., Le, Ngan, Guo, Meng, Ngo, Vien A., Le, An T.
Abstract
Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-training quantization framework that carries a consistent activation representation through calibration, weight rounding, and native integer execution. It combines channel scaling and block Hadamard transforms with dynamic per-token quantization, without policy retraining. Custom TensorRT plugins execute projections in both the language backbone and iterative action expert with four-bit weights and activations (W4A4) on Ada GPUs and Jetson AGX Orin. Evaluation spans LIBERO, SimplerEnv, and two robot platforms. Across three GR00T checkpoints and $\pi_{0.5}$, W4A4 achieves $1.20$ to $1.33\times$ speedups over floating-point TensorRT on Orin and $1.25$ to $1.52\times$ on desktop. Retaining language attention-output and feed-forward down projections at eight bits (W8A8) improves held-out action fidelity on all four checkpoints. Across four real-robot tasks, this configuration raises observed GR00T N1.7 success from $80.0\%$ with uniform W4A4 to $92.5\%$ over 80 trials per configuration, with a measured additional Orin latency of 1 ms.
Chinese Translation
低比特视觉-语言-动作(VLA)推理必须在保持机器人行为的同时降低从观测到动作的延迟。我们提出 FoldQuantVLA,这是一个训练后量化框架,在校准、权重取整和原生整数执行全过程中保持一致的激活表示。该方法将通道缩放与分块 Hadamard 变换相结合,并采用动态逐 token 量化,无需对策略进行再训练。我们开发了定制 TensorRT 插件,在 Ada GPU 和 Jetson AGX Orin 上以四比特权重与激活(W4A4)在语言骨干网络和迭代式动作专家中执行投影计算。评估涵盖 LIBERO、SimplerEnv 以及两个机器人平台。在三个 GR00T 检查点和 $\pi_{0.5}$ 上,W4A4 在 Orin 上相比浮点 TensorRT 取得 $1.20$ 至 $1.33\times$ 的加速,在桌面平台上取得 $1.25$ 至 $1.52\times$ 的加速。将语言注意力输出投影和前馈下投影保留为八比特(W8A8),可在全部四个检查点上提升留出动作的保真度。在四项真实机器人任务中,该配置在每配置 80 次试验下将 GR00T N1.7 的观测成功率从统一 W4A4 的 $80.0\%$ 提升至 $92.5\%$,而在 Orin 上仅增加 1 毫秒的实测延迟。
cs.RO / 148 / 2609.24497

ScaleMPA: Rethinking Scalable RRT* Acceleration With a Grid-Native Representation

ScaleMPA:基于网格原生表示的RRT*可扩展加速方法再思考
Wang, Zilong, Chen, Yuzhou, He, Xinyue, Zhang, Chen, He, Guanghui
Abstract
Real-time motion planning remains challenging in large and high-dimensional environments. Prior acceleration of RRT* follows tree-centric state organization, which reduces per-query cost but preserves superlinear end-to-end complexity and limits parallelism through structural dependencies. This paper presents ScaleMPA, a motion-planning accelerator that rethinks RRT* with a grid-native representation. By replacing hierarchical traversal with direct grid-based access, ScaleMPA reduces the planner critical path and exposes fine-grained parallelism. To make this reformulation practical under sparse high-dimensional planning, ScaleMPA further proposes a multi-resolution grid search engine and a hash-grid memory system. Implemented in 28 nm CMOS, ScaleMPA achieves millisecond-level planning latency and delivers 4.7$\times$--44.4$\times$ speedup over state-of-the-art motion-planning accelerators.
Chinese Translation
在大规模高维环境中的实时运动规划仍然极具挑战性。以往对RRT*的加速采用以树为中心的状态组织方式,这种方式虽然降低了单次查询的开销,但保留了超线性的端到端复杂度,并且由于结构性依赖而限制了并行性。本文提出ScaleMPA,一种以网格原生表示重新设计RRT*的运动规划加速器。通过以直接的基于网格的访问取代层次化遍历,ScaleMPA缩短了规划器的关键路径,并释放了细粒度的并行性。为了使这种重构在稀疏高维规划场景下切实可行,ScaleMPA进一步提出了多分辨率网格搜索引擎和哈希网格存储系统。在28纳米CMOS工艺下实现,ScaleMPA达到了毫秒级的规划延迟,相比最先进的运动规划加速器实现了4.7倍至44.4倍的加速。
cs.RO / 149 / 2609.24499

A Monolithic Force-Proprioception Soft Acutuator Enabled by Single-Material 3D printing

一种由单一材料3D打印实现的单体式力-本体感觉软体驱动器
Huang, Nan, Liu, Lele, Lu, Junfeng, Zhu, Yipan, Dai, Jiansheng, Liu, Sicong
Abstract
Pneumatic proprioceptive actuators integrate actuation and sensing for soft robots that attract interest due to functional potential. Existing approaches often suffer from assembly errors or stress concentrations caused by heterogeneous materials. In this work, we propose the Monolithic Force-Proprioception Soft (MFPS) design and fabrication method that integrates an Asymmetric Origami Bending (AOB) chamber and a Force-Proprioception Soft (FPS) sensor with single material through one-step Fused Deposition Modeling (FDM) fabrication. Based on the resistance response to strain of conductive thermoplastic polyurethane (TPU), we design and analyze the structure of the FPS sensor, and conduct parametric analysis on the sensing characteristics. The FDM fabrication parameters of the MFPS actuator are analyzed, followed by actuator fabrication and characterization of the actuation and proprioception performance. Experimental results show that the MFPS actuator achieves a bending angle of 40{\deg}, an output force of 12.5 N, and a resistance change of 26.9% as the applied external force increased from 0 to 45 N. A two-finger force-proprioception gripper is developed based on the MFPS actuator. The grasping and force-proprioception capabilities are experimentally validated, proving that the MFPS design method provides a new approach for the development of self-sensing actuators.
Chinese Translation
气动的本体感觉驱动器集成了驱动与感知功能,其功能潜力引起了软体机器人领域的广泛关注。现有方法往往存在装配误差或异质材料引起的应力集中问题。在本工作中,我们提出了单体式力-本体感觉软体(Monolithic Force-Proprioception Soft, MFPS)设计与制造方法,通过一步式熔融沉积成型(Fused Deposition Modeling, FDM)工艺,采用单一材料集成了非对称折纸弯曲(Asymmetric Origami Bending, AOB)腔体和力-本体感觉软体(Force-Proprioception Soft, FPS)传感器。基于导电热塑性聚氨酯(TPU)对应变的电阻响应特性,我们设计并分析了FPS传感器的结构,并对其传感特性进行了参数化分析。随后分析了MFPS驱动器的FDM制造参数,完成了驱动器的制造,并对其驱动性能和本体感觉性能进行了表征。实验结果表明,MFPS驱动器实现了40°的弯曲角度和12.5 N的输出力;当外部作用力从0增加到45 N时,电阻变化为26.9%。基于MFPS驱动器,我们开发了一种二指力-本体感觉夹持器,并通过实验验证了其抓取和力-本体感觉能力,证明MFPS设计方法为自感知驱动器的开发提供了一条新途径。
cs.RO / 150 / 2609.24507

TACIT: Tactile Contact Supervision for Spatial Attention in Dexterous Manipulation

TACIT:面向灵巧操作的触觉接触空间注意力监督
Lai, Yanhou, Zhu, Fucai, Wang, Ruiqiang, Hashimoto, Koichi
Abstract
Visuomotor policies trained from a few demonstrations may reproduce demonstrated trajectories without reliably following changes in object position. Existing approaches with explicit attention typically obtain spatial priors from human annotation or visual models. We introduce TACIT (tactile contact informs attention), which uses measured tactile contacts from teleoperated demonstrations to supervise spatial attention without additional point annotation. Gaussian targets over preceding camera point clouds supervise an attention head whose pooled output conditions a visuotactile diffusion policy. Targets are used only during training; tactile observations remain inputs at inference. In the primary real-robot benchmark, with ten demonstrations per task and five demonstrated placement regions, TACIT achieves 66.7% success on ball placement and 73.3% on peg insertion, compared with 10.0% and 20.0% for input-matched 3D visuotactile fusion and 20.0% and 43.3% for vision-only DP3. TACIT enters the 150 mm palm-to-object approach region within 12 seconds in all 30 trials per task; all remaining failures occur after arrival. Across three training seeds on real ball and simulated peg, TACIT outperforms input-matched fusion and an architecture-matched control without explicit attention supervision, supporting the contribution of supervision beyond branch capacity. Pre-contact and contact-time supervision show no consistent ordering. These results demonstrate that measured tactile contact provides effective spatial supervision for approach behavior from few demonstrations within the evaluated workspace.
Chinese Translation
基于少量演示训练的视觉运动策略可能会复现演示轨迹,但无法可靠地跟随物体位置的变化。现有的显式注意力方法通常依赖人工标注或视觉模型来获取空间先验。我们提出 TACIT(触觉接触引导注意力),利用遥操作演示中测得的触觉接触来监督空间注意力,而无需额外的点标注。通过在先前相机点云上构建高斯目标来监督注意力头,其池化输出用于调节视觉-触觉扩散策略。目标仅在训练阶段使用;触觉观测在推理阶段仍作为输入。在主要的真实机器人基准测试中,每个任务仅用十次演示和五个演示放置区域,TACIT 在球体放置任务中达到 66.7% 的成功率,在插销插入任务中达到 73.3%,而输入匹配的 3D 视觉-触觉融合方法分别为 10.0% 和 20.0%,仅视觉的 DP3 分别为 20.0% 和 43.3%。在每个任务的全部 30 次试验中,TACIT 均在 12 秒内进入距物体 150 毫米的手掌接近区域;所有剩余失败均发生在到达之后。在真实球体和仿真插销任务上跨越三个训练种子的实验中,TACIT 优于输入匹配的融合方法以及无显式注意力监督的架构匹配对照,这支持了监督本身超越分支容量的贡献。接触前与接触时刻的监督未显示出一致的性能排序。这些结果表明,在评估的工作空间内,测得的触觉接触能够为基于少量演示的接近行为提供有效的空间监督。
cs.RO / 151 / 2609.24511

InsertAnything: Generalizable Contact-Rich Precision Insertion from Simulation to Reality

InsertAnything:从仿真到现实的泛化性接触密集型精密插入
Ma, Zhenghua, Meng, Xinpan, Liu, Zeyu, Ma, Muyuan, Zhang, Hengdi, Li, Houcheng, Cheng, Long
Abstract
Contact-rich precision insertion is a key manipulation skill in robotic assembly. Tight clearances make insertion more sensitive to alignment errors and prone to collisions and jamming, while variations in geometry and clearance across parts further complicate policy reuse. We present a reinforcement learning framework that trains insertion policies entirely in simulation for direct deployment without real-world demonstrations or policy fine-tuning. By combining target poses with compact three-dimensional fingertip force feedback, the policy learns to search for alignment and correct its motion despite errors in the estimated hole position. A decoupled gated reward coordinates alignment and insertion. Force-signal smoothing and state-independent standard deviations stabilize the learning process. The resulting policies perform real-world insertion across multiple hole geometries with a minimum nominal clearance of 0.02 mm and improve success while reducing peak contact forces under hole-position errors. Cross-clearance and cross-geometry evaluations further confirm policy generalization. The system achieved the first perfect score of 20/20 on ManipulationNet's peg-in-hole benchmark under its Human-in-the-Loop protocol, with fully autonomous insertion motions. A single policy trained only on a simulated hexagonal insertion task achieved an overall success rate of 95.0% across eight unseen real-world insertion tasks. These results show that learning entirely in simulation can yield precision insertion skills that can be deployed directly and reused across real-world tasks. The project website (https://mzhsoul.github.io/InsertAnything/) provides open-source simulation and real-robot experiment scripts, assets, and trained checkpoints.
Chinese Translation
接触密集型精密插入是机器人装配中的一项关键操作技能。狭小的间隙使插入对对准误差更为敏感,且容易发生碰撞和卡阻,而零件在几何形状和间隙上的差异进一步增加了策略复用的难度。我们提出了一种强化学习框架,完全在仿真中训练插入策略,无需真实世界示教或策略微调即可直接部署。通过将目标位姿与紧凑的三维指尖力反馈相结合,该策略能够在孔位估计存在误差的情况下学会搜索对准并修正其运动。解耦的门控奖励协调对准与插入过程。力信号平滑处理和与状态无关的标准差稳定了学习过程。所得到的策略在真实世界中可对多种孔几何形状进行插入操作,最小标称间隙为0.02 mm,并在存在孔位误差的情况下提高成功率的同时降低了峰值接触力。跨间隙和跨几何形状的评估进一步证实了策略的泛化能力。该系统在ManipulationNet的轴孔(peg-in-hole)基准测试中,依据其人在环路(Human-in-the-Loop)协议,首次获得20/20的满分,且插入运动完全自主。仅在仿真六边形插入任务上训练的单一策略,在八个未见过的真实世界插入任务中取得了95.0%的总体成功率。这些结果表明,完全在仿真中学习可以获得可直接部署并在真实世界任务间复用的精密插入技能。项目网站(https://mzhsoul.github.io/InsertAnything/)提供了开源的仿真与真机实验脚本、资源以及训练好的模型权重。
cs.RO / 152 / 2609.24525

Bridge3D: Enabling Vision-Language-Action Models to See and Act in 3D

Bridge3D:使视觉-语言-动作模型能够在3D中感知与行动
Li, Haoxuan, Yan, Sixu, Zhu, Lianghui, Tang, Xuanlai, Wang, Shikang, Wang, Xinggang
Abstract
Vision-Language-Action (VLA) models have demonstrated remarkable generalization in robotic manipulation via large-scale multimodal pretraining. However, VLA models are mainly trained on 2D-centric observations, which inherently constrains their capacity for precise spatial manipulation. Previous methods enhance 3D awareness by introducing implicit spatial priors, but still lack explicit geometry guidance. In this paper, we propose Bridge3D that integrates both implicit and explicit 3D geometry guidance into pre-trained 2D VLA models, enabling them to ''see'' and ''act'' in 3D. Bridge3D introduces two strategies: 1) Implicit Fusion, which enriches visual tokens with features from 3D foundation models to improve ''seeing'' in 3D; 2) Explicit Conditioning, which integrates action denoising with an explicit 3D semantic field to achieve ''acting'' in 3D. Furthermore, we utilize the proposed layer-wise linear probing to improve learning efficiency. Experiments show that Bridge3D achieves superior performance against state-of-the-art methods. On the RoboTwin 2.0 benchmark, Bridge3D exceeds $\pi_0$ by 14.0 percentage points, while in real-world experiments, it outperforms Spatial Forcing by 11.7 percentage points. These results demonstrate Bridge3D's strong capabilities in high-precision and spatial-sensitive manipulation tasks.
Chinese Translation
视觉-语言-动作(Vision-Language-Action,VLA)模型通过大规模多模态预训练在机器人操作任务中展现出卓越的泛化能力。然而,VLA模型主要以2D为中心的观测数据进行训练,这从根本上限制了其进行精确空间操作的能力。以往的方法通过引入隐式空间先验来增强3D感知能力,但仍缺乏显式的几何引导。本文提出Bridge3D,将隐式和显式的3D几何引导融合到预训练的2D VLA模型中,使其能够在3D中“感知”和“行动”。Bridge3D引入了两种策略:1)隐式融合(Implicit Fusion),通过融合来自3D基础模型的特征来丰富视觉token,从而提升3D“感知”能力;2)显式条件化(Explicit Conditioning),将动作去噪与显式的3D语义场相结合,实现3D“行动”。此外,我们还利用所提出的逐层线性探测(layer-wise linear probing)方法来提高学习效率。实验表明,Bridge3D相对于当前最先进的方法取得了更优的性能。在RoboTwin 2.0基准测试中,Bridge3D超过$\pi_0$模型14.0个百分点;在真实世界实验中,其性能优于Spatial Forcing 11.7个百分点。这些结果证明了Bridge3D在高精度和空间敏感操作任务中的强大能力。
cs.RO / 153 / 2609.24535

Estimation and Control of Tensegrity Manipulator Kinematics based on Strut Inclination Angles

基于撑杆倾角的张拉整体机械臂运动学估计与控制
Bhat, Tufail Ahmad, Ikemoto, Shuhei
Abstract
Unlike conventional rigid-link robots defined by discrete joints, continuum robots pose a fundamental challenge for expressing their complex continuous bending configurations for closed-loop control. Several modelling approaches have been proposed for conventional continuum robots, but tensegrity-based continuum robots remain largely open. Moreover, many of these approaches assume a continuous elastic backbone and are therefore not directly applicable to tensegrity manipulators, whose bodies are networks of rigid struts and tensioned cables. This work presents a reduced-order model for shape and posture control of a tensegrity-based continuum manipulator. The manipulator is modelled as a serially connected parallel-link mechanism. The proposed method is formulated as an optimization problem that uses geometric constraints of the tensegrity structure together with information from the Inertial Measurement Unit (IMU) sensors embedded in the strut elements. To the best of our knowledge, this work presents the first experimental demonstration of a real-time IMU-based shape estimation method on a full-scale tensegrity manipulator and demonstrates posture control using a simple Proportional-Integral (PI) controller. The results show that the proposed method can estimate the shape of both single-module tensegrity structures and multi-module tensegrity manipulators from arbitrary static configurations and achieve desired postures.
Chinese Translation
与由离散关节定义的传统刚性连杆机器人不同,连续体机器人如何表达其复杂的连续弯曲构型以实现闭环控制是一个根本性挑战。针对传统连续体机器人已有多种建模方法被提出,但基于张拉整体(tensegrity)结构的连续体机器人建模仍大多处于空白状态。此外,许多现有方法假设存在连续的弹性骨干,因此无法直接应用于张拉整体机械臂——其主体由刚性撑杆和受拉缆索构成的网络组成。本文提出了一种用于张拉整体式连续体机械臂形状与姿态控制的降阶模型。该机械臂被建模为串联连接的并联连杆机构。所提出的方法被表述为一个优化问题,综合利用张拉整体结构的几何约束以及嵌入撑杆元件中的惯性测量单元(IMU)传感器的信息。据我们所知,本工作首次在全尺寸张拉整体机械臂上演示了基于IMU的实时形状估计方法,并利用简单的比例-积分(PI)控制器实现了姿态控制。结果表明,所提出的方法能够从任意静态构型估计单模块张拉整体结构和多模块张拉整体机械臂的形状,并实现期望的姿态。
cs.RO / 154 / 2609.24547

MIRA: Real-Time Full-Duplex Human-Robot Interaction for Embodied Companions

MIRA:面向具身陪伴机器人的实时全双工人机交互
Lin, Lijian, Zhu, Ye, Zhang, Fan, Liu, Yunfei, Li, Baofeng, Zeng, Xianwen, Wang, Jianan, Li, Yu
Abstract
Real-time embodied companion interaction requires a robot to infer user intent from streaming speech, generate timely responses, and execute expressive, interruptible motions. Existing systems typically decouple dialogue orchestration from gesture synthesis, relying on offline motion generation from complete audio. This separation leaves open how a deployed robot can dynamically synchronize response content, prosodic timing, and physical safety under incremental inputs and uncertain turn boundaries. We present MIRA, a unified framework for full-duplex embodied companion interaction. Given streaming user speech, dialogue history, and vocal affect, MIRA predicts both the response text and an explicit embodiment cue. Discrete social behaviors (\eg listening, greeting) are mapped to validated robot trajectories, while open-ended speaking is paired with streaming, co-speech motion. This generative motion is governed by a predict-more-than-commit sliding window that provides temporal look-ahead for motion continuity while limiting physical commitment to a short, cancellable prefix. Crucially, we design CORTEX, a dual-timescale interaction policy that manages low-latency streaming and deliberative turn decisions, backed by a robot-side execution layer that enforces physical safety constraints at the control rate. We deploy MIRA on an Astribot S1 humanoid robot. Quantitative evaluations demonstrate competitive audio-motion alignment relative to state-of-the-art motion-generation baselines, while real-robot deployment measurements characterize streaming responsiveness and interruption handling.
Chinese Translation
实时具身陪伴交互要求机器人能够从流式语音中推断用户意图、及时生成回复,并执行富有表现力且可中断的动作。现有系统通常将对话编排与手势生成解耦,依赖基于完整音频的离线动作生成。这种分离使得已部署的机器人难以在增量输入和不确定的对话轮次边界下,动态地同步回复内容、韵律时序与物理安全。我们提出了MIRA,一个面向全双工具身陪伴交互的统一框架。给定流式用户语音、对话历史和语音情感,MIRA同时预测回复文本和显式的具身提示(embodiment cue)。离散的社交行为(如倾听、问候)被映射到经过验证的机器人轨迹,而开放式的说话行为则与流式的伴随语音动作(co-speech motion)相配对。这种生成式动作由一个“多预测少执行”(predict-more-than-commit)滑动窗口控制,该窗口为动作连续性提供时间前瞻,同时将物理执行限制在一段较短的、可取消的前缀内。至关重要的是,我们设计了CORTEX,一种双时间尺度的交互策略,用于管理低延迟流式处理和审慎的对话轮次决策,并由机器人端执行层提供支持,以控制频率强制执行物理安全约束。我们将MIRA部署在Astribot S1人形机器人上。定量评估表明,其音画对齐性能相对于最先进的动作生成基线具有竞争力,同时真机部署的测量结果刻画了其流式响应能力和中断处理性能。
cs.RO / 155 / 2609.24552

Smoothness as a Constraint for Stable Humanoid Locomotion

以平滑性作为约束实现稳定的人形机器人运动
Panchal, Utsav, Kleyko, Denis, Artan, Unal, Loutfi, Amy
Abstract
Embodied AI systems, particularly humanoid robots deployed in real world scenarios require whole-body control policies that are both task-responsive and physically smooth. However, smoothness is not uniform across the body: lower body must remain sufficiently reactive, while the upper body must be tightly regulated to preserve stability. Existing reinforcement learning approaches typically impose smoothness through auxiliary terms in the reward function, which compete with task objectives, treating the body as uniform and provide no direct control over the physical quantities responsible for smooth behavior. We introduce DeCap (Decoupled Constraint-aware policy), a constrained reinforcement learning algorithm that decouples whole-body smoothness into separate upper- and lower-body constraint groups, each formulates smoothness as explicit constraints on physical motion limits. To improve constraint satisfaction near feasibility boundaries, DeCap incorporates a bounded barrier penalty that activates proactively as limits are approached and remains bounded at the constraint limit. On real-world humanoid whole-body control task, DeCap reduces upper-body action rate by 2.50x and acceleration by 2.18x relative to reward-based smoothness policies, while also improving lower-body smoothness and reducing transient motion. We demonstrate that a fixed set of smoothness constraints transfers across diverse terrains, alleviating the need of extensive reward tuning.
Chinese Translation
具身人工智能(Embodied AI)系统,尤其是部署在真实世界场景中的人形机器人,需要既响应任务又具备物理平滑性的全身控制策略。然而,平滑性在身体各部位并非均匀一致:下肢必须保持足够的反应性,而上肢则需受到严格调控以维持稳定性。现有的强化学习方法通常通过在奖励函数中添加辅助项来实现平滑性,这不仅与任务目标相互竞争,还将身体视为一个整体,且无法对决定平滑行为的物理量进行直接控制。我们提出了DeCap(Decoupled Constraint-aware policy,解耦的约束感知策略),这是一种约束强化学习算法,它将全身平滑性解耦为上肢和下肢两个独立的约束组,每组均将平滑性表述为对物理运动限制的显式约束。为了在可行域边界附近提高约束满足度,DeCap 引入了一种有界的障碍惩罚(bounded barrier penalty),当接近运动极限时提前激活,并在约束极限处保持有界。在真实世界的人形机器人全身控制任务中,与基于奖励的平滑性策略相比,DeCap 将上肢动作速率降低了2.50倍,加速度降低了2.18倍,同时还改善了下肢平滑性并减少了瞬态运动。我们证明,一组固定的平滑性约束可以迁移到多种地形上,从而减少了大量奖励调参的需求。
cs.RO / 156 / 2609.24563

ARSTAG: An Agentic Real2Sim2Real System for Task-Specific Robot Data Generation

ARSTAG:一种面向任务特定机器人数据生成的智能体化Real2Sim2Real系统
Li, Bowei, Zhang, Yuner, Liu, Changliu
Abstract
Adapting visuomotor policies to new manipulation tasks often requires substantial manual engineering or teleoperated data collection. Simulation can provide task-specific data at scale, but constructing the scene, designing expert behavior, and configuring data generation still require significant per-task effort. We present ARSTAG, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data. A hierarchy of language agents constructs a task-scoped simulation scene, generates robot-feasible demonstrations, and expands the training distribution through task-consistent randomization, while a coordinator agent manages cross-stage feedback and recovery. Across seven manipulation tasks spanning grasping, placement, and stacking, the ARSTAG-generated demonstrations enable sim-to-real transfer of three visuomotor policy architectures to a dual-arm robot, with pi0.5 achieving an average real-world success rate of 74.6%. Ablations show that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size. Project webpage: https://boweili666.github.io/ARSTAG/.
Chinese Translation
将视觉运动策略适配到新的操作任务通常需要大量的人工工程或遥操作数据采集。仿真虽然可以大规模提供任务特定的数据,但构建场景、设计专家行为以及配置数据生成仍需要针对每个任务投入大量精力。我们提出ARSTAG,这是一种智能体化的Real2Sim2Real系统,能够将单张RGB图像和自然语言指令直接转化为机器人策略学习数据。一个分层的语言智能体体系构建任务范围的仿真场景、生成机器人可执行的示范数据,并通过任务一致的随机化扩展训练分布,同时由一个协调智能体管理跨阶段的反馈与恢复。在涵盖抓取、放置和堆叠的七项操作任务中,ARSTAG生成的示范数据使三种视觉运动策略架构能够实现到双臂机器人的仿真到现实迁移,其中pi0.5在真实世界中平均成功率达到74.6%。消融实验表明,任务一致的随机化显著提升了鲁棒性,且策略性能随生成数据集规模的增大而提高。项目网页:https://boweili666.github.io/ARSTAG/。
cs.RO / 157 / 2609.24576

What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior

基于VLM的视觉-语言导航模型依赖什么:策略行为的解释与调控
Makowski, Débora Oliveira, Gode, Samiran, Nayak, Abhijeet, Hutter, Marco, Schmid, Cordelia, Schmid, Lukas Rosenberger, Burgard, Wolfram
Abstract
Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While this fusion of language instructions and visual observations allows multimodal reasoning, it obscures how information is routed across modalities or what mechanisms drive navigation decisions. Thus, it remains unclear whether VLN models ground their predictions in relevant semantic cues or can track task progress. In this work, we study the interpretability and steerability of VLN models. We use intervention-based metrics that measure how visual observations, instructions, and visual memory causally influence navigation decisions. Our results show that these navigation policies are sensitive to all input modalities and do not depend on a single one. We further show that these agents encode navigation progress and retain semantic structure from their VLM backbones, enabling concept-level steering through internal activations. Finally, we extract activation vectors for abstract behaviors to transfer them zero-shot to out-of-distribution real-world scenarios, improving performance without additional fine-tuning.
Chinese Translation
现代视觉-语言导航(Vision-Language Navigation, VLN)模型主要依赖预训练的大型视觉-语言模型(VLM)来预测导航动作。虽然语言指令与视觉观察的这种融合实现了多模态推理,但也掩盖了信息如何在各模态之间传递,以及哪些机制驱动了导航决策。因此,VLN模型究竟是基于相关的语义线索做出预测,还是能够跟踪任务进度,仍然不得而知。在本工作中,我们研究了VLN模型的解释性与可调控性。我们采用基于干预的度量方法,衡量视觉观察、指令和视觉记忆对导航决策的因果影响。结果表明,这些导航策略对所有输入模态均敏感,并不依赖单一模态。我们进一步证明,这些智能体能够编码导航进度,并保留其VLM骨干网络的语义结构,从而可以通过内部激活实现概念层面的调控。最后,我们提取了抽象行为的激活向量,并以零样本方式将其迁移到分布外的真实世界场景中,无需额外微调即可提升性能。
cs.RO / 158 / 2609.24621

Learning tactile perception from high-bandwidth single-point sensing

从高带宽单点感知中学习触觉感知
Rigal, Joseph, Virot, Emmanuel, Pascal, Caroline
Abstract
Tactile sensing is increasingly being incorporated into learning-based robotic manipulation, yet many existing approaches rely on spatially distributed sensors. Here we introduce SpectRobot, a framework that transforms single-point tactile signals into compact time-frequency spectrograms. These spectrograms encode high-bandwidth tactile histories as fixed-size image-like representations. They can be processed by standard vision encoders and integrated into learning pipelines originally developed for vision, while preserving temporal and frequency information unavailable to conventional cameras. Rather than increasing spatial density through arrays of tactile elements, SpectRobot exploits the rich dynamics contained in sparse, high-bandwidth single-point measurements. In our implementation, the sensors are mounted away from the contact surface while remaining mechanically coupled to it, reducing direct exposure to wear and potentially improving robustness in harsh environments and for long-term deployment on dexterous robots. Our experiments demonstrate that: (1) a robot can exploit single-point vibration signals to solve a visually occluded manipulation task; (2) temporal history strongly influences policy performance, while sensing bandwidth controls the spectral information available, with measurements extending to 100~kHz; and (3) the same representation can be used across different tactile sensing technologies mediated by acceleration, force, or strain. We further show that capabilities previously associated with research-grade instrumentation can be accessed using readily available, off-the-shelf hardware. We believe that broader access to high-bandwidth tactile sensing could facilitate the integration of contact dynamics into embodied learning systems and, for some tasks, offer an alternative or complement to increasing the spatial density of tactile sensing.
Chinese Translation
触觉感知正日益被纳入基于学习的机器人操作中,然而许多现有方法依赖于空间分布式传感器。在此,我们提出了 SpectRobot,一个将单点触觉信号转换为紧凑的时频频谱图的框架。这些频谱图将高带宽触觉历史编码为固定尺寸的类图像表示。它们可以由标准视觉编码器处理,并集成到最初为视觉任务开发的学习管线中,同时保留了传统相机无法获取的时间和频率信息。SpectRobot 并非通过触觉元件阵列来增加空间密度,而是利用稀疏、高带宽的单点测量中所蕴含的丰富动态信息。在我们的实现中,传感器安装在与接触表面保持机械耦合但远离接触表面的位置,从而减少直接磨损,并有望在恶劣环境以及灵巧机器人的长期部署中提升鲁棒性。我们的实验表明:(1)机器人可以利用单点振动信号解决视觉遮挡下的操作任务;(2)时间历史对策略性能有显著影响,而感知带宽决定了可获得的频谱信息,测量范围可扩展至 100 kHz;(3)同一表示可跨由加速度、力或应变介导的不同触觉传感技术使用。我们进一步证明,以往需要科研级仪器才能实现的能力,现在可以通过易于获取的现成硬件来实现。我们相信,更广泛地获取高带宽触觉感知,有助于将接触动力学融入具身学习系统,并且在某些任务中,可以作为增加触觉感知空间密度的替代方案或补充。
cs.RO / 159 / 2609.24631

From Semantic Decisions to Feasible Trajectories: Self-Evolving LLM-Guided Optimal Control for Narrow-Space Parking

从语义决策到可行轨迹:面向窄空间泊车的自演化LLM引导最优控制
Yao, Zhengbao, Luo, Yuanfu, Xue, Kehan
Abstract
Autonomous parking in nonconvex and narrow environments remains challenging. Although optimal-control methods can explicitly enforce vehicle dynamics and collision constraints, nonconvexity compromises solver robustness and can cause failures. Large language models (LLMs) exhibit strong semantic reasoning capabilities, but directly generating dense trajectories makes it difficult to guarantee physical feasibility. We introduce SE-LLM-OCP, a unified framework in which LLMs make high-level discrete maneuver decisions, while an optimal-control module enforces low-level vehicle dynamics and collision constraints. Online, the LLM proposes sparse maneuver plans, decomposing the parking task into a sequence of short-horizon trajectory-optimization problems. A low-level solver then sequentially solves optimal-control problems. If the solver fails, the LLM aggregates failure evidence from the solver and validation stages to guide replanning. Offline, SE-LLM-OCP automatically evolves a structured decision-making knowledge base from scratch, driven by accumulated online failures. We validate our proposed framework in simulation on a car-like vehicle model and on a differential-drive robot. Our experimental results show that SE-LLM-OCP enables safer autonomous parking in narrow scenarios and demonstrates transfer of the same maneuver representation to a different kinematic platform.
Chinese Translation
非凸且狭窄环境下的自主泊车仍然具有挑战性。尽管最优控制方法能够显式地满足车辆动力学与碰撞约束,但非凸性会削弱求解器的鲁棒性并可能导致求解失败。大语言模型(LLM)展现出强大的语义推理能力,但直接生成稠密轨迹难以保证物理可行性。我们提出了SE-LLM-OCP,这是一个统一框架:LLM负责做出高层离散机动决策,而最优控制模块则强制执行低层车辆动力学与碰撞约束。在线阶段,LLM提出稀疏的机动计划,将泊车任务分解为一系列短时域轨迹优化问题;低层求解器随后依次求解这些最优控制问题。若求解器失败,LLM会聚合来自求解器和验证阶段的失败证据以指导重规划。离线阶段,SE-LLM-OCP基于积累的在线失败经验,从零开始自动演化出一个结构化的决策知识库。我们在仿真中基于类车车辆模型和差速驱动机器人对所提框架进行了验证。实验结果表明,SE-LLM-OCP能够在狭窄场景下实现更安全的自主泊车,并将相同的机动表示迁移到了不同的运动学平台上。
cs.RO / 160 / 2609.24632

Feasibility Distance Fields for Heterogeneous Constraints in Robot Configuration Space

面向机器人构型空间中异构约束的可行性距离场
Cui, Xijing, Pu, Huayan, Luo, Jun, Wang, Gang
Abstract
Robot manipulators are monitored by constraint-specific indicators whose units and gradient scales are not comparable, so they do not provide a common measure of the configuration-space motion remaining before violation. We define the feasibility distance field (FDF) as the distance, under a fixed positive-definite joint-space metric, to the union of infeasible configuration sets. Classical distance-to-set theory gives 1-Lipschitz continuity, almost-everywhere differentiability, and unit dual-gradient norm wherever the nearest projection is unique. The robotics contribution is an admissibility analysis showing when practical constraints define non-empty closed sets. We derive admissible formulations for external and self-collision, joint limits, dexterity, Cartesian and task-projected compliance, joint torque under payload, and dynamic manipulability. Since every field uses the same metric, heterogeneous constraints compose by a pointwise minimum, conditioned constraints retain a fixed distance space, and multi-robot constraints produce block-sparse gradients that identify which robots must react. We generate projection-based labels and train neural approximations with a distance loss and an Eikonal penalty. Simulations on a UR5e and a dual-arm cell evaluate seven fields using value, projection, sign, gradient, composition, and moving-obstacle diagnostics. Across 8,000 configurations, the largest feasible-side secant ratio is 0.920, mean learned gradient norms range from 0.994 to 0.998, and projection residuals range from 0.011 to 0.034 rad. Across 24 random obstacle paths, the external and composed collision fields achieve 90.4% and 91.6% success within 3 cm, with sign-error rates below 2%. The results support a common configuration-space margin and identify approximation errors near medial axes and sparsely sampled boundaries.
Chinese Translation
机器人机械臂由特定于约束的指标进行监控,这些指标的单位和梯度尺度不可比较,因而无法为违反约束前构型空间中剩余的运动提供一个统一的度量。我们将可行性距离场(Feasibility Distance Field, FDF)定义为在固定的正定关节空间度量下,到不可行构型集合之并集的距离。经典的到集合距离理论给出了1-Lipschitz连续性、几乎处处可微性,以及在最邻近投影唯一处对偶梯度的单位范数。本文的机器人学贡献在于一项可采性分析,阐明了实际约束何时能定义非空的闭集。我们为外部碰撞与自碰撞、关节限位、灵巧性、笛卡尔及任务投影柔顺性、带载荷下的关节力矩以及动态可操作度推导了可采的(admissible)公式。由于每个场都使用相同的度量,异构约束可通过逐点最小值进行组合;条件约束保持固定的距离空间;多机器人约束则产生块稀疏梯度,从而识别出哪些机器人需要作出反应。我们生成基于投影的标签,并采用距离损失和Eikonal惩罚项训练神经近似模型。在UR5e和双臂工作单元上的仿真中,通过数值、投影、符号、梯度、组合及移动障碍物等诊断手段评估了七个场。在8000个构型中,可行侧最大割线比率为0.920,学习到的梯度范数均值介于0.994至0.998之间,投影残差介于0.011至0.034弧度之间。在24条随机障碍物路径中,外部碰撞场与组合碰撞场在3厘米阈值内的成功率分别为90.4%和91.6%,符号错误率低于2%。这些结果支持了一种统一的构型空间裕度度量,并指出了在中轴附近及稀疏采样边界处的近似误差。
cs.RO / 161 / 2609.24660

Touch2Robot: Robot Touch in the Human Demonstration Loop

Touch2Robot:人机示范回路中的机器人触觉
Luo, Shengcheng, Cheng, Xiaoyang, Ying, Hong, Zhou, Xiaoying, Jiang, Jiaming, Guo, Haoran, Li, Wanlin, Jiao, Ziyuan, Xiao, Chenxi
Abstract
Human demonstrations offer a scalable way to collect manipulation data, but their contacts may be unstable or infeasible when transferred to a robot hand. Collecting demonstrations directly on the target robot avoids this mismatch, but substantially increases the cost of data collection. To address this trade-off, we present \textbf{Touch2Robot}, a framework that lets humans collect demonstrations while seeing how the target robot hand would contact the object. We capture human hand motion, tactile-glove measurements, and object motion during human manipulation. These recordings guide object-specific RL policies to reproduce the demonstrated object motion while favoring contacts consistent with the recorded human touch. We distill the learned behaviors into a unified real-time retargeter that maps incoming human observations and object geometry to robot hand configurations. During collection, the predicted robot configuration is synchronized with the tracked object pose in simulation to reconstruct robot-object contacts, which are visualized to help the demonstrator adapt subsequent interactions to the target hand. Across four real-world tasks, Touch2Robot improves average real-robot replay completion from 37.9\% to 72.1\% over visual-only feedback, while reducing the collection time per replay-successful demonstration from 58.6~s to 18.2~s. Reconstructed target-hand contacts achieve 44.2\% F1 against real-robot tactile measurements, and policies trained on Touch2Robot demonstrations improve downstream Diffusion Policy performance by 29.1 percentage points over visual-only feedback. These results show that bringing robot touch into the human demonstration loop improves both the quality and efficiency of scalable dexterous data collection. \textit{Project webpage: \href{https://Touch2Robot.github.io/}{https://Touch2Robot.github.io/}.}
Chinese Translation
人类示范为采集操作数据提供了一种可扩展的方式,但当将其迁移到机器人手上时,其接触方式可能不稳定或不可行。直接在目标机器人上采集示范可以避免这种不匹配,但会大幅增加数据采集的成本。为了解决这一权衡问题,我们提出了 Touch2Robot 框架,使人类能够在看到目标机器人手将如何接触物体的同时进行示范采集。我们捕捉人类操作过程中的手部运动、触觉手套测量数据以及物体运动。这些记录用于引导针对特定物体的强化学习(RL)策略复现所演示的物体运动,同时倾向于与所记录的人类触觉相一致的接触方式。我们将学习到的行为蒸馏为一个统一的实时重定向器(retargeter),将传入的人类观测数据和物体几何信息映射为机器人手部配置。在采集过程中,预测的机器人配置在仿真中与被追踪的物体位姿同步,以重建机器人-物体接触,并将其可视化,以帮助示范者根据目标手调整后续的交互。在四个真实世界任务中,与仅使用视觉反馈相比,Touch2Robot 将真实机器人回放的平均完成率从 37.9% 提升至 72.1%,同时将每次成功回放示范的采集时间从 58.6 秒降低到 18.2 秒。重建的目标手接触相对于真实机器人触觉测量达到了 44.2% 的 F1 分数,并且基于 Touch2Robot 示范训练的策略在下游 Diffusion Policy 性能上比仅使用视觉反馈提升了 29.1 个百分点。这些结果表明,将机器人触觉引入人类示范回路能够同时提升可扩展灵巧操作数据采集的质量与效率。项目网页:https://Touch2Robot.github.io/。
cs.RO / 162 / 2609.24682

Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

像世界模型一样思考,像VLA一样行动:将世界模型表示蒸馏进紧凑的机器人策略
Dao, Trung, Yamsani, Sankalp, Park, Jaden, Kim, Joohyung, Lee, Yong Jae
Abstract
Vision-Language-Action (VLA) models map observations to actions with no objective that accounts for how the world responds, so their robustness is bounded primarily by data coverage. World models carry precisely that missing objective and are better grounded for it, yet rolling the future forward costs seconds per decision and rules them out of the control loop. We show the two can be separated. What a world model knows about physical scenes lives in its \emph{internal features}; generating the future is merely the objective that produced them, so the grounding can be inherited while the generative machinery is left behind. We add one feature-alignment term to ordinary VLA training: a frozen world model is run over the training frames once and cached, and the student learns to agree with that cache. No teacher is loaded during training, the projector is discarded after it, and the deployed policy is identical to the undistilled baseline, running in $32$~ms and $1.86$~GB on a consumer RTX~5090, so every gain is attributable to the representation rather than to added capacity or test-time compute. A $0.8$B student reaches $97.9\%$ on LIBERO, improves from $48.2\%$ to $50.5\%$ on RoboCasa-GR1 humanoid manipulation, and the same objective carries over to real hardware, on both a single-arm and a bimanual platform. The gain survives changes of student scale, backbone, alignment layer, and teacher, indicating a broad representational prior rather than a fragile alignment between two particular networks. Project page: https://thaw-vla.trung-dt.com/.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型将观测直接映射为动作,其目标函数中不包含对世界如何响应的考量,因此其鲁棒性主要受限于数据覆盖范围。世界模型恰好具备这一缺失的目标,并因此具有更好的物理接地能力,但向前推演未来所需的计算量使得每次决策耗时数秒,从而将其排除在控制回路之外。我们证明这两者可以被分离。世界模型关于物理场景的知识蕴藏于其内部特征中;生成未来只是产生这些特征的目标,因此可以在抛弃生成机制的同时继承其接地能力。我们在普通的VLA训练中仅添加一个特征对齐项:冻结的世界模型只需在训练帧上运行一次并缓存,学生模型学习与该缓存保持一致。训练期间无需加载教师模型,训练后投影器被丢弃,部署的策略与未经蒸馏的基线完全相同,在消费级RTX 5090上以32毫秒延迟和1.86 GB显存运行,因此所有性能提升都可归因于表示本身,而非增加的模型容量或测试时计算。一个0.8B参数的学生模型在LIBERO上达到97.9%,在RoboCasa-GR1人形机器人操作任务上从48.2%提升至50.5%,且相同的目标函数可迁移至真实硬件,在单臂和双臂平台上均得到验证。该性能增益在学生模型规模、骨干网络、对齐层和教师模型的变化下依然成立,表明其源于一种广泛的表示先验,而非两个特定网络之间脆弱的对齐。项目页面:https://thaw-vla.trung-dt.com/。
cs.RO / 163 / 2609.24702

AC-DC: Adaptive Communication for Scalable Dynamic Average Consensus in Multi-Robot Ergodic Search

AC-DC:面向多机器人遍历搜索中可扩展动态平均共识的自适应通信
Kee, Robin Inho, Cannataro, Begum, Tzoumas, Vasileios
Abstract
We study scalable peer-to-peer dynamic average consensus (DC) for multi-robot systems under finite-range, finite-rate, and interference-constrained communication. We introduce Adaptive Communication for Dynamic Average Consensus (AC-DC), which jointly adapts Who communicates with whom, When, and over What parts of the consensus state, using local inputs and successfully received neighbor information. Each robot's consensus state estimates the current average of the robots' local inputs. AC-DC updates these estimates as local inputs change and averages the values exchanged between robot pairs. In AC-DC, robot pairs update without waiting for every robot to complete a communication round, and the selected-state messages carry consensus state coordinates independent of team size for a fixed state representation. We apply AC-DC to dynamic-priority multi-robot ergodic search: one consensus stream estimates team visitation for motion coordination, while the other fuses regional measurement information to update uncertainty maps and search targets. Across twelve settings with up to 80 robots and 20 paired trials per setting, AC-DC has the lowest mean (i) normalized covariance-trace area under the curve (AUC) and (ii) attempted modeled communication payload among the compared decentralized methods. Averaged across settings, AC-DC achieves paired AUC reductions of 27.5% relative to state-of-the-art baselines, with 8.7x less communication traffic. As the number of robots increases, we observe that AC-DC's communication payload approaches that of the ideal centralized baseline (one ground compute-station communicating directly with all robots): with 120 robots in a fixed 600 x 600 m scaling test, AC-DC uses 19.3 MB versus 19.2 MB for the ideal centralized baseline, while remaining peer-to-peer.
Chinese Translation
我们研究了在有限范围、有限速率和受干扰约束的通信条件下,面向多机器人系统的可扩展点对点动态平均共识(Dynamic Average Consensus, DC)。我们提出了动态平均共识的自适应通信方法(Adaptive Communication for Dynamic Average Consensus, AC-DC),该方法利用本地输入和成功接收到的邻居信息,联合自适应地调整:与谁通信(Who)、何时通信(When)以及通信共识状态的哪些部分(What)。每个机器人的共识状态估计各机器人本地输入的当前平均值。AC-DC 在本地输入变化时更新这些估计,并对机器人对之间交换的值进行平均。在 AC-DC 中,机器人对无需等待所有机器人完成一轮通信即可进行更新,且在固定的状态表示下,所选状态消息携带的共识状态坐标数与团队规模无关。我们将 AC-DC 应用于动态优先级的多机器人遍历搜索(ergodic search):其中一条共识流估计团队的访问分布以实现运动协调,另一条融合区域测量信息以更新不确定性地图和搜索目标。在多达 80 个机器人、每个设置 20 组配对试验的十二种设置中,在所比较的分散式方法中,AC-DC 在以下两方面的均值最低:(i)归一化协方差迹的曲线下面积(AUC),(ii)尝试的建模通信负载。在各设置上取平均后,AC-DC 相对于最先进的基线方法实现了 27.5% 的配对 AUC 降低,通信流量减少 8.7 倍。随着机器人数量的增加,我们观察到 AC-DC 的通信负载趋近于理想的集中式基线(一个地面计算站与所有机器人直接通信):在固定的 600 x 600 米规模测试中,使用 120 个机器人时,AC-DC 的通信量为 19.3 MB,而理想集中式基线为 19.2 MB,同时 AC-DC 仍保持点对点通信方式。
cs.RO / 164 / 2609.24708

SPARSER: Sparse Variable Projection by Exploiting Separable Structure in Robotic Perception

SPARSER:利用机器人感知中可分离结构的稀疏变量投影方法
Sanderson, Nikolas R., Fishberg, Andrew, Han, Haoyu, Yang, Heng, How, Jonathan P., Singh, Hanumant, Everett, Michael, Papalia, Alan
Abstract
Robotic perception often requires solving large nonlinear least-squares (NLS) problems. While sparsity has been widely exploited to scale solvers, a complementary and underused structure is \emph{separability}: some variables, such as visual landmarks, appear linearly in the residuals and admit a closed-form solution once the remaining variables, such as poses, are fixed. Variable projection (VarPro) exploits this structure by analytically eliminating the linear variables, yielding a reduced problem with favorable computational properties. However, its use in robotic perception has been limited by gauge symmetries, such as invariance to global translations and rotations, which introduce challenges for standard VarPro methods. We present SPARSER (\textbf{S}parsity \textbf{P}reserving \textbf{A}nalytic \textbf{R}eduction for \textbf{S}eparable \textbf{R}obotic \textbf{P}erception), a VarPro framework for gauge-symmetric problems that jointly exploits separability and sparsity. Our method constructs a \emph{matrix-free Schur complement operator} for efficient evaluation of reduced costs, gradients, and Hessian-vector products, enabling integration with iterative NLS solvers. We characterize the applicable problem class, identify common cases admitting further analytical simplifications, and show that IRLS-based robust costs preserve most of the exploitable structure. Across synthetic and real SLAM, SNL, and SfM benchmarks, SPARSER is on average $5\times$--$7\times$ faster than state-of-the-art baselines on CPU and GPU, with gains exceeding $40\times$ on individual datasets. On outlier-corrupted multi-robot SLAM data, the robust variant is $2\times$--$16\times$ faster than a state-of-the-art GNC solver. We release open-source C++ code and all datasets.
Chinese Translation
机器人感知通常需要求解大规模非线性最小二乘(NLS)问题。尽管稀疏性已被广泛利用以扩展求解器的规模,但另一种互补且未被充分利用的结构是可分离性:某些变量(如视觉路标)在残差中以线性形式出现,并在其余变量(如位姿)固定后存在闭式解。变量投影(VarPro)正是利用这一结构,通过解析方式消去线性变量,从而得到一个具有良好计算性质的降阶问题。然而,其在机器人感知中的应用一直受限于规范对称性(如对全局平移和旋转的不变性),这给标准VarPro方法带来了挑战。我们提出SPARSER(面向可分离机器人感知的稀疏性保持解析降阶方法),一个针对规范对称问题的VarPro框架,可同时利用可分离性与稀疏性。我们的方法构建了一个无矩阵Schur补算子,以高效计算降阶后的代价、梯度以及Hessian-向量积,从而可与迭代NLS求解器集成。我们刻画了适用的问题类别,识别出可进一步解析简化的常见情形,并证明基于IRLS的鲁棒代价函数能够保留大部分可利用的结构。在合成的与真实的SLAM、SNL和SfM基准测试中,SPARSER在CPU和GPU上平均比最先进的基线方法快5倍至7倍,个别数据集上的加速比超过40倍。在受外点污染的多机器人SLAM数据上,其鲁棒变体比最先进的GNC求解器快2倍至16倍。我们发布了开源C++代码及全部数据集。
cs.RO / 165 / 2609.24742

LLM-based Conversational AI Knowledge Assistant for MyBuddy Humanoid Robot

面向MyBuddy人形机器人的基于大语言模型(LLM)的对话式AI知识助手
Chen, Hanxiao
Abstract
Humanoid robots are increasingly being popular and developed for human-centered applications, yet their ability to provide intelligent conversations and natural interactive knowledge assistance remains constrained by traditional rule-based dialogue systems, pre-defined responses and limited knowledge repositories. Large language models (LLMs) have emerged as a powerful foundation for enabling natural, adaptive, and context-aware Human-Robot Interaction (HRI), which provides a significant opportunity to address such limitations by enabling robots to understand natural speech language, reason over complicated queries, maintain high-quality conversational context, and generate knowledge-rich responses. In this work, we originally present and implement an LLM-based versatile Conversational AI Knowledge Assistant for the Raspberry-Pi-powered 13-Axis MyBuddy humanoid robot, which integrates LLM-driven language understanding and AI reasoning with real-time speech recognition, knowledge retrieval via extensible access of internet engines (e.g., Wikipedia, arXiv), flexible dialogue management, and natural speech synthesis to enable much more intelligent multi-turn continuous conversations and advanced emotional-support Human-Robot Interaction.
Chinese Translation
人形机器人在以人为中心的应用中日益普及并快速发展,然而其在提供智能对话和自然交互式知识辅助方面的能力,仍然受限于传统的基于规则的对话系统、预定义的响应以及有限的知识库。大语言模型(Large Language Models, LLMs)已成为实现自然、自适应和情境感知的人机交互(Human-Robot Interaction, HRI)的强大基础,为解决上述局限提供了重要机遇,使机器人能够理解自然语言、对复杂问题进行推理、维护高质量的对话上下文并生成知识丰富的回复。在本工作中,我们原创性地提出并实现了一个基于大语言模型的多功能对话式AI知识助手,应用于由树莓派(Raspberry Pi)驱动、具有13个自由度的MyBuddy人形机器人。该系统将大语言模型驱动的语言理解与AI推理,同实时语音识别、通过可扩展的互联网引擎访问(如Wikipedia、arXiv)实现的知识检索、灵活的对话管理以及自然语音合成相集成,从而实现更加智能的多轮连续对话以及先进情感支持型人机交互。
cs.RO / 166 / 2609.24745

Beyond Visual Quality: A Study of Test-Time Planning with World Action Models

超越视觉质量:基于世界动作模型的测试时规划研究
Yuan, Jianhao, Yuan, Yu, Ramtoula, Benjamin, Vierling, Lukas, Newman, Paul, Kunze, Lars, Torr, Philip, De Martini, Daniele
Abstract
World action models generate actions together with visual predictions of their consequences. These paired outputs create the potential for planning by sampling multiple actions from one state, comparing their imagined outcomes, and choosing the action with the most promising predicted outcome. However, how to use imagined futures to guide action selection remains unclear. We examine this planning potential empirically. First, we estimate an oracle upper bound on selection by choosing the sampled candidate whose realised outcome is best. In a controlled same-state analysis, this choice raises success from 68.9% under uniform random selection to 79.2%. We then test selectors based on visual quality, physical consistency, and task progression as controlled interventions. Some tested selectors yield higher observed success, but the gains are uneven and the matched selectors leave much of the measured opportunity unrecovered. To investigate this gap, we examine whether sampled actions lead to different outcomes, whether these differences are visible in the predictions, and whether a score recognises them. Counterfactual branching from the same states shows that selection opportunity is concentrated in relatively few decisions in the initial candidate sets. Action spread and outcome coverage need not increase together. In a further evaluation across trajectory phases with complete action execution, the tested scores again recover little of the available improvement despite a small gain from learned value. These findings distinguish producing consequential action choices from recognising them in generated futures, motivating the evaluation of WAM predictions through their usefulness for decisions rather than visual quality alone.
Chinese Translation
世界动作模型(World Action Models, WAM)在生成动作的同时生成对其后果的视觉预测。这种成对的输出为规划创造了可能性:从同一状态中采样多个动作,比较它们设想的结果,并选择预测结果最有希望的动作。然而,如何利用想象的未来来指导动作选择仍不清楚。我们对该规划的潜力进行了实证研究。首先,我们通过选择已实现结果最佳的采样候选动作,估计了选择性能的预言上界。在受控的同状态分析中,这种选择将成功率从均匀随机选择下的68.9%提升至79.2%。随后,我们测试了基于视觉质量、物理一致性和任务进展的选择器,作为受控干预。其中一些选择器确实带来了更高的观测成功率,但增益并不均衡,且这些选择器未能回收大部分已测得的选择机会。为探究这一差距,我们考察了以下问题:采样动作是否会导致不同的结果、这些差异是否在预测中可见、以及某个评分能否识别出这些差异。从相同状态出发的反事实分支分析表明,选择机会集中在初始候选集合中相对较少的决策上。动作的分散程度与结果覆盖范围并不必然同步增长。在进一步跨轨迹阶段、包含完整动作执行的评估中,尽管学习到的价值函数带来小幅提升,所测试的评分仍只回收了很小的可用改进空间。这些发现区分了“产生具有后果影响的动作选择”与“在生成的未来中识别这些选择”,从而启发我们应依据WAM预测对决策的有用性而非仅凭视觉质量来评估其性能。
cs.RO / 167 / 2609.24749

D-JEPA: A Decision-Aligned Latent World Model

D-JEPA:一种决策对齐的潜在世界模型
Liu, Shuaijun, Wu, Chengyu, Wen, Qifu, You, Feiyang, Zhang, Chenglong, Hao, Shuyang, Lin, Xi, Su, Ningxin
Abstract
Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully. We identify a decision-local prediction gap: among the few futures competing for execution, a candidate predicted closer to the goal can produce a worse realized outcome than an available alternative. We introduce D-JEPA, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes. A bounded, permutation-equivariant operator jointly reasons over goal-relative predictive features and ordinal evidence, refining pretrained predictive geometry where action choices are most consequential. Restricted predictor adaptation and a shared ordinal interface extend this alignment across complementary predictive geometries. D-JEPA further realizes the learned decision structure in JEPA-compatible future representations, enabling deployment through native latent-distance planning. Evaluations across latent control, manipulation, pretrained action-producing models, physical robots and autonomous driving demonstrate improved action selection, including 87.89% success on PushT, a 15.04-point average gain on RoboTwin, and a 17-point gain on physical robot tasks. These results establish decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.
Chinese Translation
潜在世界模型用于预测动作的后果,但准确的预测并不能保证潜在距离能够反映哪个候选动作将被成功执行。我们发现了一种决策局部的预测鸿沟:在少数竞争执行的未来方案中,某个被预测为更接近目标的候选动作,其最终实际结果可能比一个可用的替代方案更差。我们提出了 D-JEPA,一种决策对齐的潜在世界模型,它从实际执行的结果中学习候选未来方案之间与决策相关的关联关系。一个有界的、置换等变的算子对目标相对的预测特征和序数证据进行联合推理,在动作选择最具决定性的地方对预训练的预测几何进行精化。受限的预测器自适应和共享的序数接口将这种对齐扩展到互补的预测几何之上。D-JEPA 进一步在 JEPA 兼容的未来表征中实现所学到的决策结构,从而能够通过原生的潜在距离规划进行部署。在潜在空间控制、操作任务、预训练动作生成模型、物理机器人以及自动驾驶上的评估表明,动作选择性能得到提升,包括在 PushT 上达到 87.89% 的成功率,在 RoboTwin 上平均提升 15.04 个百分点,以及在物理机器人任务上提升 17 个百分点。这些结果确立了与决策相关的关联结构作为预测式世界建模与有效控制之间的直接桥梁。
cs.RO / 168 / 2609.24761

A Switched Adaptive Control Framework for Aerial Manipulators Under Dynamic Transitions

动态切换下空中机械臂的切换自适应控制框架
Yadav, Rishabh Dev, Gupta, Saksham, Sharma, Amitabh, Mishra, Sarthak, Pan, Wei, Roy, Spandan, Baldi, Simone
Abstract
Aerial manipulators represent the forefront of aerial robotics. Although potentially capable of complex interaction tasks, controlling aerial manipulators throughout the dynamic transitions occurring during task execution presents significant challenges. Abrupt or discontinuous changes in system dynamics generated by the transitions suggest the use of a switched approach, yet the available aerial manipulation methods are not designed for coping with switched regimes. In addition, most available methods fall short in coping with the tight couplings between the aerial vehicle and the manipulator, as well as in coping with the state-dependent uncertainties arising from the difficulty in modeling such couplings. We propose a switched-based adaptive control framework for aerial manipulators not relying on a priori knowledge of the vehicle-manipulator couplings and of state-dependent uncertainties. To guarantee stable manipulation despite changes in system dynamics, the framework provides a class of switching signals characterizing those transition phases for which the system is guaranteed to remain stable. Comparative experiments further validate the effectiveness of the proposed switched-based framework over the state of the art.
Chinese Translation
空中机械臂代表了空中机器人技术的前沿。尽管空中机械臂有潜力执行复杂的交互任务,但在任务执行过程中,对其发生的动态过渡阶段进行控制仍然存在重大挑战。这些过渡所引起的系统动力学的突变或非连续变化提示我们可以采用切换方法,然而现有的空中操作方法并未针对切换工况进行设计。此外,大多数现有方法在应对飞行器与机械臂之间的紧耦合,以及应对因难以建模这种耦合而产生的状态相关不确定性方面均存在不足。我们提出了一种基于切换的自适应控制框架,该框架无需对飞行器-机械臂耦合及状态相关不确定性的先验知识。为了保证系统动力学发生变化时操作的稳定性,该框架提供了一类切换信号,用于表征能够保证系统保持稳定的过渡阶段。对比实验进一步验证了所提出的基于切换的框架相对于现有先进方法的有效性。
cs.RO / 169 / 2609.24778

H2RBench: A Real-to-Sim Benchmark for Evaluating Human-to-Robot Transfer

H2RBench:一个用于评估人到机器人迁移的真实到仿真基准
Xiao, Chuyang, Zhan, Haotian, Krishna, Sriram, Meng, Peilin, Irshad, Muhammad Zubair, Zakharov, Sergey, Held, David
Abstract
Learning robot manipulation policies from human video demonstrations constitutes a promising avenue for scalable robot learning. However, comparing different human-to-robot (H2R) transfer methods remains challenging, as existing approaches are evaluated under different settings, including differing task suites, scene layouts, object instances, and amounts of robot supervision. To address this challenge, we present H2RBench, a Real2Sim benchmark for evaluating H2R transfer methods. H2RBench provides a standardized protocol built on real human video demonstrations and simulated robot demonstrations, and includes four manipulation tasks spanning diverse interaction requirements. We evaluate multiple representative H2R transfer methods, each adopting a different strategy for bridging the embodiment gap. Using H2RBench, we systematically characterize how each method scales with the amount of human demonstrations, revealing that methods differ substantially in their ability to leverage additional human data. We further show that simulation performance is broadly predictive of real-world robot performance, with an overall Pearson correlation of r = 0.89, Spearman correlation of \r{ho} = 0.85 and Mean Maximum Rank Violation (MMRV) of 0.06 across method-task configurations. These results establish H2RBench as a practical and scalable benchmark for comparative H2R evaluation prior to real-world deployment.
Chinese Translation
从人类视频演示中学习机器人操作策略是实现可扩展机器人学习的一条有前景的途径。然而,比较不同的人到机器人(Human-to-Robot, H2R)迁移方法仍然具有挑战性,因为现有方法是在不同的设置下进行评估的,包括不同的任务集、场景布局、物体实例以及机器人监督数据的数量。为应对这一挑战,我们提出了 H2RBench,一个用于评估 H2R 迁移方法的 Real2Sim(真实到仿真)基准。H2RBench 提供了一个基于真实人类视频演示和仿真机器人演示的标准化协议,并包含四个涵盖多样化交互需求的操作任务。我们评估了多种具有代表性的 H2R 迁移方法,每种方法采用了不同的策略来弥合本体差异。借助 H2RBench,我们系统地刻画了各方法如何随人类演示数据量的增加而扩展,结果表明各方法在利用额外人类数据的能力上存在显著差异。我们进一步证明,仿真性能能够在较大程度上预测真实世界中的机器人性能:在不同方法-任务配置下,总体 Pearson 相关系数为 r = 0.89,Spearman 相关系数为 ρ = 0.85,平均最大秩违反(Mean Maximum Rank Violation, MMRV)为 0.06。这些结果确立了 H2RBench 作为在真实世界部署之前进行 H2R 对比评估的实用且可扩展的基准。
cs.RO / 170 / 2609.24815

Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI

Uranus:面向具身智能的下一代仿真基础设施
Qin, Wenkang, Zhou, Yukun, Shen, Noah, Cai, Jisong, Mao, Dongxiao, Li, Baicheng, Zhang, Yue, Sui, Wei
Abstract
Scalable simulation is essential for robot data generation, policy training, evaluation, and safe iteration, yet real-world interaction is costly and conventional simulators require labor-intensive construction. We present Uranus, a data-driven robot simulator built around a joint-trajectory-conditioned autoregressive diffusion model. Uranus offers three key capabilities: (1) streaming, open-ended rollout, which receives future joint-position trajectories online and autoregressively generates one latent frame per step, corresponding to four RGB frames, without a fixed horizon; (2) low-latency generation, achieving 24 FPS after inference optimization; and (3) scalable, extensible robot control, providing a unified interface for synchronized multi-view generation across diverse robot embodiments and camera configurations. We conduct comprehensive quantitative and qualitative evaluations on both in-distribution and out-of-distribution data, providing an objective assessment of Uranus and clearly identifying its current limitations. We release the code and model weights to empower the community with practical tools and insights.
Chinese Translation
可扩展的仿真对于机器人数据生成、策略训练、评估和安全迭代至关重要,然而真实世界交互成本高昂,且传统模拟器需要耗费大量人力进行构建。我们提出了Uranus,一个基于联合轨迹条件自回归扩散模型构建的数据驱动机器人模拟器。Uranus提供三项关键能力:(1) 流式的开放式轨迹推演,可在线接收未来关节位置轨迹,并以自回归方式每步生成一个潜在帧(对应四帧RGB图像),且不受固定时间范围的限制;(2) 低延迟生成,经推理优化后达到24 FPS;(3) 可扩展、可扩展的机器人控制,为跨多种机器人本体和相机配置的同步多视角生成提供统一接口。我们在分布内和分布外数据上进行了全面的定量和定性评估,对Uranus进行了客观评价,并明确指出了其当前的局限性。我们发布了代码和模型权重,以期为社区提供实用的工具和见解。
cs.RO / 171 / 2609.24832

Minimum Time Trajectories for a Car-Like Mobile Robot Moving with Rigid Wheels Under Non-Sliding Constraints

满足非滑动约束的刚性轮式类汽车移动机器人最短时间轨迹
Ben-Asher, Joseph, Rimon, Elon, Ravina, Leeor
Abstract
This paper studies the minimum time trajectoriesvof a car-like mobile robot navigating in an obstacle free environment. The robot, with forward and backward speeds, is controlled by bounded front-wheels acceleration and limited front-wheels steering rate. The paper extends previous results which solved this problem for the kinematic car-like robot. However, the kinematic model assumes pure rolling at the wheels ground contacts. This assumption requires non-sliding constraints for the front and rear wheels that can only be handled by the robot dynamics. This paper formulates the non-sliding constraints based on the robot dynamics then augments the kinematic model time-optimal path primitives with three new path primitives associated with the non-sliding constraints. The three non-sliding path primitives together with the kinematic model twelve path primitives form the car-like robot time optimal trajectories. Approximate analytic solutions for the non-sliding path primitives are also provided. Examples study the time-optimal path primitives along representative maneuvers, illustrating how the non-sliding constraints influence the time optimal trajectories of the car-like robot.
Chinese Translation
本文研究了类汽车移动机器人在无障碍环境中导航的最短时间轨迹。该机器人具有前进和后退速度,由有界的前轮加速度和受限的前轮转向速率进行控制。本文扩展了此前针对运动学类汽车机器人求解该问题的研究结果。然而,运动学模型假设车轮与地面接触为纯滚动,该假设要求前轮和后轮满足非滑动约束,而这些约束只能通过机器人动力学来处理。本文基于机器人动力学建立了非滑动约束的表达式,并在运动学模型的时间最优路径基元基础上增加了与该非滑动约束相关的三个新路径基元。这三个非滑动路径基元与运动学模型的十二个路径基元共同构成了类汽车机器人的时间最优轨迹。本文还提供了非滑动路径基元的近似解析解。通过算例研究了代表性机动动作下的时间最优路径基元,说明了非滑动约束如何影响类汽车机器人的时间最优轨迹。
cs.RO / 172 / 2609.24840

PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control

PredActor:面向可操控机载人形机器人控制的预测性动作扩散方法
Ye, Lei, Gao, Haibo, Li, Yitang, Xu, Peng, Jing, Zetong, Sun, Junhan, Dong, Fanrong, Han, Ziqi, Wang, Xue, Sun, Jianhua, Lu, Cewu, Zhao, Hao, Ding, Liang
Abstract
Diffusion models offer flexible motion generation, but translating this flexibility into feedback-responsive humanoid control remains challenging. Hierarchical systems steer motion through references that may exceed a separate tracker's capabilities, leaving recovery and physical execution largely to the tracker. Action-only diffusion generates actions directly but lacks an explicit future-state trajectory for test-time motion objectives. Joint state-action diffusion provides this representation, yet representative controllers often depend on privileged full-body states, and support for learned behavior selection and test-time motion steering remains fragmented. We present PredActor, a predictive action diffusion policy that brings these complementary steering capabilities into one directly executed policy using proprioceptive observations. Conditioned on proprioceptive history and optional task context, PredActor jointly generates executable actions and an internal future-state trajectory. Classifier-free guidance strengthens text-conditioned behavior, while classifier guidance steers predicted states toward test-time objectives. Only actions are executed, without a separate motion-reference tracker or externally estimated full-body states as policy inputs. In simulation, PredActor reaches all 15 destination targets and achieves a text retrieval score of 0.580, compared with 0.373 for conditional action diffusion, with similar observed disturbance survival. To make this guided policy practical onboard, rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, both below the 20 ms control period. We deploy PredActor on a Unitree G1; evaluations across simulation and physical hardware demonstrate text-conditioned motion, disturbance response, joystick control, and semantic interpolation.
Chinese Translation
扩散模型能够提供灵活的运动生成能力,但如何将这种灵活性转化为具备反馈响应能力的人形机器人控制仍然充满挑战。分层系统通过参考信号来引导运动,而这些参考信号可能超出独立跟踪器的执行能力,使得运动恢复和物理执行在很大程度上依赖于跟踪器本身。纯动作扩散模型可以直接生成动作,但缺乏用于测试时运动目标的显式未来状态轨迹。状态-动作联合扩散提供了这种表示形式,但代表性控制器往往依赖特权级的全身状态信息,并且对学习到的行为选择和测试时运动引导的支持仍较为零散。我们提出 PredActor,一种预测性动作扩散策略,它仅使用本体感知观测,将这些互补的引导能力整合到一个可直接执行的策略中。PredActor 以本体感知历史和可选的任务上下文为条件,联合生成可执行动作和内部未来状态轨迹。无分类器引导(classifier-free guidance)增强了文本条件下的行为表现,而分类器引导(classifier guidance)则将预测状态引导至测试时目标。策略仅执行动作,无需独立的运动参考跟踪器,也无需将外部估计的全身状态作为策略输入。在仿真中,PredActor 能够到达全部 15 个目标点,文本检索得分为 0.580,而条件动作扩散仅为 0.373,两者的扰动生存能力相近。为使该引导策略能够在机载环境中实际运行,滚动去噪和保持计算量的运行时优化将完整的回调时间在 Jetson Orin NX 上降低至中位数 16.790 ms 和 p95 19.383 ms,均低于 20 ms 的控制周期。我们将 PredActor 部署在 Unitree G1 人形机器人上;跨仿真与物理硬件的评估展示了文本条件运动、扰动响应、摇杆控制以及语义插值能力。
cs.RO / 173 / 2609.24841

CAST: Collision-Aware Assembly with Construction Robots using Simultaneous Trajectory Estimation and Planning

CAST:基于同步轨迹估计与规划的碰撞感知建筑机器人装配方法
Shaji, Karthik, Kim, Chisung, D'Amato, John, Bruun, Edvard, Dellaert, Frank
Abstract
Multi-robot systems have shown increasing viability in construction due to their ability to execute high-precision actions while reducing human exposure to hazardous tasks. However, these environments have high-dimensional configuration spaces and possess substantial collision-avoidance constraints, which include other robots, assembly objects, and workspace boundaries. We utilize a single factor graph for trajectory estimation and planning that incorporates measured robot states together with explicit collision and learned cable constraints. This supports changing workspaces and enables synchronized, high-dimensional robot motion planning while accounting for the stiff, vibration-induced uncertainty of heavy robotic systems. We demonstrate the success of our framework on the construction of a post-and-lintel structure using one robot arm as a timber gripper, and a second robot as a nail-fastener.
Chinese Translation
多机器人系统因其在执行高精度操作的同时减少人类暴露于危险任务的能力,在建筑领域显示出越来越大的可行性。然而,这类环境具有高维配置空间,并存在大量的避碰约束,包括其他机器人、装配对象和工作空间边界。我们利用单一因子图进行轨迹估计与规划,将测量得到的机器人状态与显式碰撞约束以及学习得到的缆绳约束相结合。该方法支持工作空间的变化,能够在考虑重型机器人系统因刚度不足和振动引起的不确定性的同时,实现同步的高维机器人运动规划。我们通过一个机器人臂作为木材抓取器、另一个机器人作为钉固件,搭建梁柱结构,验证了该框架的成功性。
cs.RO / 174 / 2609.24846

Range-Aided SLAM Initialization Exploiting Accurate Heading Information

利用精确航向信息的测距辅助SLAM初始化方法
Lougheed, Isabel, Forbes, James Richard
Abstract
This paper presents a novel initialization method for range-aided simultaneous localization and mapping (RA-SLAM). The general SLAM problem has a well-known separable structure where landmark and robot positions can be solved for in a linear fashion given known robot headings. This paper considers the case where highly accurate heading information is available, which is typical of autonomous underwater vehicle (AUV) navigation, to solve for the range transponder and robot positions. The proposed approach consists of two steps. First, using a generalized trust region subproblem (GTRS), the positions of the transponders relative to the AUV are solved for given the nonlinear range measurements. Second, the relative transponder positions and the known heading of the AUV are used to estimate the transponder and AUV positions by solving a linear least-squares problem. These transponder and AUV position estimates, combined with the highly accurate heading information, provide a reliable initialization method for the general nonlinear RA-SLAM problem. The effectiveness of the proposed approach is tested on a real-world AUV dataset where long baseline (LBL) range measurements are provided in concert with highly accurate heading information provided by an inertial navigation system (INS).
Chinese Translation
本文提出了一种新颖的测距辅助同步定位与建图(RA-SLAM)初始化方法。一般的SLAM问题具有著名的可分离结构,即在已知机器人航向的情况下,地标位置和机器人位置可以通过线性方式求解。本文考虑了可获得高精度航向信息的情形(这在自主水下航行器(AUV)导航中很典型),用于求解测距应答器和机器人的位置。所提出的方法包括两个步骤:首先,利用广义信赖域子问题(GTRS),在给定非线性测距观测的条件下求解应答器相对于AUV的位置;其次,利用应答器的相对位置和已知的AUV航向,通过求解线性最小二乘问题来估计应答器和AUV的位置。这些应答器和AUV的位置估计结合高精度航向信息,为一般的非线性RA-SLAM问题提供了可靠的初始化方法。所提方法的有效性在一个真实的AUV数据集上进行了测试,该数据集提供了长基线(LBL)测距数据以及由惯性导航系统(INS)提供的高精度航向信息。
cs.RO / 175 / 2609.24864

SE(3) Neural Potential Fields for 6-DoF Trajectory Planning Directly from Images Without Explicit 3D Reconstruction

直接从图像进行6自由度轨迹规划的SE(3)神经势场,无需显式三维重建
Eiyike, Jeffrey, Ataei, Masoud, Gyaase, Elvis, Dhiman, Vikas
Abstract
Reaching a 6-DoF grasp pose in clutter requires a collision-free trajectory, conventionally obtained by reconstructing the scene in 3D and planning inside that reconstruction, at the cost of its accuracy and compute. Potential fields learned directly from images remove that dependency but inherit the classical weakness of artificial potential fields: where attractive and repulsive gradients cancel, the descent grazes the obstacle instead of going around it, and can stall short of the goal. We present an SE(3) neural potential field learned from posed RGB images and supervised with a navigation function, the geodesic distance to the grasp through free space recovered from those same images during training, which removes both failures. On two tabletop scenes, from obstacle-blocked starts executed on a UR10, the field converges within 3 cm of the grasp from every start and every path it executes is collision-free against the ground-truth geometry, against 25% and 0% under image supervision alone; mean clearance rises from under a centimeter to 8.6-8.8 cm and arm-link contacts fall from 20.6-50.4% to 2.7-5.5% of executed configurations. Executed grasp success is 90.0% and 40.0% on the two scenes, the residual failures being refusals of the Cartesian executor rather than of the field. Planning takes about 2 s against 67-133 s for RRT* on a reconstruction of the same images, though under a common offline harness the two are comparable: the deployed margin is the cost of collision-checking a dense reconstruction, not planner complexity.
Chinese Translation
在杂乱环境中到达6自由度抓取位姿需要一条无碰撞轨迹,传统方法通过三维重建场景并在重建结果中进行规划来实现,但这需要付出重建精度和计算成本的代价。直接从图像学习的势场消除了这一依赖,但继承了人工势场的经典缺陷:当吸引梯度与排斥梯度相互抵消时,梯度下降会擦碰障碍物而非绕过它,并可能在到达目标前停滞。我们提出一种从带位姿RGB图像学习的SE(3)神经势场,并用导航函数进行监督——该导航函数是在训练过程中从相同图像恢复的自由空间中到抓取位姿的测地距离——从而同时消除上述两类失败。在两个桌面场景中,从被障碍物阻挡的起始点出发并在UR10机械臂上执行,该势场从每个起始点均收敛到抓取位姿3厘米以内,且其执行的每条路径相对真实几何均为无碰撞;而仅用图像监督时,这两个比例分别为25%和0%。平均间隙从不足1厘米提升至8.6-8.8厘米,机械臂连杆碰撞比例从执行位姿的20.6-50.4%降至2.7-5.5%。在两个场景中,实际执行的抓取成功率分别为90.0%和40.0%,剩余失败源于笛卡尔执行器的拒绝而非势场本身。规划耗时约2秒,而对相同图像的重建结果运行RRT*需67-133秒;不过在统一的离线测试框架下,二者表现相当:部署中的时间优势来自于避免了对稠密重建结果进行碰撞检测的开销,而非规划器复杂度的差异。
cs.RO / 176 / 2609.24868

DualWAM: Dual-System World Action Models for Asynchronous Global Planning and Local Refinement

DualWAM:面向异步全局规划与局部优化的双系统世界动作模型
Zheng, Yixin, Lyu, Jiangran, Deng, Yuntian, Liu, Kai, Zhou, Yizhou, Wang, Yizhou, Zhao, Xiaoguang, Wang, He, Zhang, Zhizheng
Abstract
World Action Models (WAMs) jointly generate robot actions and predict future world states, transferring priors from video pretraining to robot control. However, future visual prediction is computationally expensive, so existing WAMs often rely on long action chunks to amortize inference cost across control steps, at the cost of closed-loop responsiveness. We present \method, a dual-system WAM that preserves broader-horizon world-action generation while enabling high-frequency closed-loop action updates by decoupling global planning and local refinement. \systwo periodically performs high-noise bidirectional denoising over a broader world-action chunk to establish a global plan, while wrist-only \sysone extracts a temporally aligned short window from the intermediate denoising state and completes low-noise refinement using the latest wrist observations, which provide action-aligned cues about local geometry, motion, and contact during interaction. The two systems operate asynchronously along a shared denoising trajectory: each global plan is reused across multiple local updates, while \sysone repeatedly incorporates fresh interaction feedback. Across zero-shot manipulation tasks on Franka and Galbot, \method improves success over the strongest evaluated baseline by 4.5 percentage points on average, while achieving a 16.6$\times$ critical-path speedup. Further studies show that role-matched egocentric and UMI data improve success by 14 percentage points, and that the decoupled design naturally supports edge--cloud deployment with substantially lower communication overhead than the baseline.
Chinese Translation
世界动作模型(World Action Models, WAMs)能够联合生成机器人动作并预测未来世界状态,将视频预训练获得的先验知识迁移到机器人控制中。然而,未来的视觉预测在计算上开销巨大,因此现有的WAM通常依赖较长的动作块(action chunk)来将推理成本分摊到多个控制步骤中,但这是以牺牲闭环响应能力为代价的。我们提出了DualWAM,这是一种双系统WAM,通过解耦全局规划与局部优化,在保留更长远时域的世界-动作生成能力的同时,实现高频闭环动作更新。系统二(System 2)周期性地在更长的世界-动作块上执行高噪声的双向去噪以建立全局规划;而仅使用腕部摄像头的系统一(System 1)则从中间去噪状态中提取时间对齐的短时间窗口,并利用最新的腕部观测完成低噪声的优化——这些腕部观测在交互过程中提供了与动作对齐的局部几何、运动和接触信息。两个系统沿着共享的去噪轨迹异步运行:每个全局规划可在多次局部更新中复用,而系统一则反复融入最新的交互反馈。在Franka和Galbot上的零样本操作任务中,DualWAM相比所评估的最强基线平均提升了4.5个百分点的成功率,同时实现了16.6倍的关键路径加速。进一步研究表明,角色匹配的第一视角(egocentric)数据和UMI数据可将成功率提升14个百分点,且这种解耦设计天然支持边缘-云部署,其通信开销显著低于基线方法。
cs.RO / 177 / 2609.24896

Steerable and Reactive Grasping Through Modular Design with a Three-Point Interface

基于模块化设计与三点接口的可操控、可响应抓取
Nguyen, Andrew, Lee, Yonghyeon, Kim, Sangbae
Abstract
Dexterous grasping requires deciding where to grasp, reaching the target, and maintaining stable contact. We connect these stages through a compact three-point interface that separates global geometric reasoning from local contact control. Given object geometry and optional language commands, our framework samples contact triples from a precomputed grasp-affordance heatmap. A model-based reactive controller tracks the object, avoids collisions, and guides the hand toward the selected contacts. In the final centimeters, a Reinforcement Learning (RL) policy uses proprioceptive feedback to refine and stabilize the grasp despite reaching and perception errors. It observes only finger joint states and its recent actions, with no target points, visual observations, or object geometry, so a single policy is shared across objects and grasp configurations. In simulation, we compare grasp-and-lift success against squeeze and end-to-end baselines, characterize reaching convergence, and demonstrate grasp steering; hardware demonstrations on two training objects and one unseen object illustrate the full pipeline. Our modular framework uses geometry to guide the reach and local feedback to secure the grasp.
Chinese Translation
灵巧抓取需要决定在哪里抓取、到达目标位置以及维持稳定接触。我们通过一个紧凑的三点接口将这些阶段连接起来,将全局几何推理与局部接触控制相分离。给定物体几何信息及可选的语言指令,我们的框架从预计算的抓取可供性热力图中采样接触三点组合。基于模型的响应式控制器负责跟踪物体、避免碰撞,并引导机械手向选定的接触点移动。在最后几厘米的接近阶段,一个强化学习(Reinforcement Learning, RL)策略利用本体感觉反馈来修正并稳定抓取,以应对到达和感知误差。该策略仅观察手指关节状态及其近期的动作,不依赖目标点、视觉观测或物体几何信息,因此单一策略可跨物体和抓取构型共享。在仿真中,我们将抓取并提起的成功率与挤压基线和端到端基线进行了比较,分析了到达的收敛性,并展示了抓取的可操控性;在两个训练物体和一个未见物体上的硬件演示展示了完整的系统流程。我们的模块化框架利用几何信息引导机械手到达目标,并利用局部反馈确保抓取的稳定性。
cs.RO / 178 / 2609.24906

Visuomotor Robotic Pruning in Planar Orchards Using Hybrid Reinforcement Learning

基于混合强化学习的平面果园视觉运动机器人修剪
Jain, Abhinav, Grimm, Cindy, Lee, Stefan
Abstract
Dormant tree pruning is labor-intensive yet essential for maintaining modern high-productivity fruit orchards. In this work, we focus on pruning of modern planar tree training systems - V-Trellis apples and UFO cherries - where trunks and primary branches are trained into approximately planar walls. We introduce an end-to-end pipeline to learn a closed-loop visuomotor controller for robotic pruning. This controller is trained entirely using simulation and synthetically generated data and deployed in real orchards in a zero-shot manner. The pipeline comprises synthetic generation of planar orchard tree meshes, construction of a physics-based orchard simulator, automated collection of successful pruning trajectories via motion planning, and policy learning with a novel hybrid reinforcement-learning algorithm that combines offline demonstrations with online simulated rollouts. The controller uses optical-flow inputs from a wrist-mounted camera - avoiding the need for full 3D-reconstruction - and continuously guides the cutter through cluttered branch environments to a specified cutpoint with correct tool orientation. In exhaustive simulated task-space evaluations over 3,000 pruning points, the policy attains 49.9% success on V-Trellis apples and 46.0% on UFO cherries. We validate the learned controller across 38 physical trials - comprising 28 outdoor field trials in commercial and experimental orchards and 10 indoor laboratory tests - demonstrating zero-shot sim-to-real transfer. The learned policy also outperforms a classical RRT-Connect baseline on physical hardware in laboratory trials.
Chinese Translation
休眠期果树修剪劳动强度大,但对于维持现代高产果园至关重要。本工作聚焦于现代平面树形培养系统的修剪——V形架(V-Trellis)苹果和UFO形(UFO)樱桃,其主干和主枝被培养成近似平面的树墙。我们提出了一种端到端流水线,用于学习机器人修剪的闭环视觉运动控制器。该控制器完全基于仿真和合成数据进行训练,并以零样本(zero-shot)方式部署于真实果园。该流水线包括:平面果园树木网格的合成生成、基于物理的果园仿真器的构建、通过运动规划自动采集成功修剪轨迹,以及采用一种新颖的混合强化学习算法进行策略学习,该算法将离线示范与在线仿真推演相结合。控制器使用来自腕部相机的光流输入——避免了对完整三维重建的需求——并在杂乱的枝条环境中持续引导剪切工具以正确的工具姿态到达指定切割点。在超过3,000个修剪点的详尽仿真任务空间评估中,该策略在V形架苹果上达到49.9%的成功率,在UFO形樱桃上达到46.0%。我们通过38次物理实验对所学控制器进行了验证——包括28次在商业果园和实验果园中进行的户外田间试验以及10次室内实验室测试——证明了零样本的仿真到现实(sim-to-real)迁移能力。在实验室的物理硬件测试中,所学策略的性能也优于经典的RRT-Connect基线方法。
cs.RO / 179 / 2609.24952

Learning to Drive on Mars: Visual Multimodal Traversability Estimation for Off-World Navigation

学习在火星上驾驶:用于地外导航的视觉多模态可通行性估计
Chiu, Darren, Wilson, Cole, Tumbar, Andrei, Sukhatme, Gaurav S., Myint, Steven
Abstract
Autonomous navigation on Mars requires vehicles to distinguish between traversable terrains across diverse and visually challenging environments. However, progress in learning-based navigation for off-world environments has been limited by the lack of large-scale datasets. Since landing in Jezero Crater, the Mars 2020 Perseverance rover has traversed terrain ranging from sandy dunes, rocky patches, and flat bedrocks. As a result, this paper presents a dataset spanning 500 sols and 45km of trajectories driven by both human operators and the onboard planner, ENav. Our dataset contains grayscale stereo image pairs, poses, accelerometer readings, rocker-bogie angles, and estimates of tilt and wheel slip. Building on this dataset, we introduce an uncertainty-aware traversability-estimation framework that learns terrain representations from multimodal driving experience. We compare our proposed method against existing approaches on the Mars 2020 dataset and show that our method achieves an AUROC of 0.874 and an F1 score of 0.758, outperforming the strongest baseline by 0.058 and 0.156, respectively, while also achieving the highest average precision and recall. Finally, we show that the visual representations can be integrated into path planners, such as ENav, on a physical rover test bed. Videos, code, and the M2020 dataset will be available at https://darren-chiu.github.io/learning-to-drive-on-mars.
Chinese Translation
火星上的自主导航要求探测器能够在多样且视觉上极具挑战性的环境中区分可通行的地形。然而,由于缺乏大规模数据集,基于学习的地外环境导航研究进展有限。自着陆于耶泽罗陨石坑(Jezero Crater)以来,火星2020毅力号(Mars 2020 Perseverance)漫游车已在沙丘、岩石区和平坦基岩等多种地形上行驶。基于此,本文提出了一个跨越500个火星日(sol)、总里程45公里的数据集,其轨迹由人类操作员和车载规划器ENav共同驱动。该数据集包含灰度立体图像对、位姿、加速度计读数、摇臂转向架角度以及倾斜度和车轮滑移的估计值。在此基础上,我们提出了一种不确定性感知的可通行性估计框架,该框架能够从多模态驾驶经验中学习地形表征。我们在火星2020数据集上将所提出的方法与现有方法进行了比较,结果表明我们的方法实现了0.874的AUROC和0.758的F1分数,分别比最强基线高出0.058和0.156,同时还取得了最高的平均精确率和召回率。最后,我们在一个物理漫游车测试平台上展示了可以将视觉表征集成到诸如ENav等路径规划器中。视频、代码和M2020数据集将在 https://darren-chiu.github.io/learning-to-drive-on-mars 发布。
cs.RO / 180 / 2609.24976

DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation

DexTacWAM:用于灵巧操作的视触觉世界-动作模型
Yuan, Haoran, Wang, Zekai, Shao, Boning, Lu, Haoran, Darrell, Trevor, Lourentzou, Ismini, Zhan, Wei
Abstract
Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to modeling contact evolution as part of the predicted world state rather than tactile conditioning alone: removing tactile world modeling reduces the four-task mean from 74.7 to 26.6 while keeping the same tactile features and action expert. After four hours of tactile-encoder adaptation with a frozen pretrained vision VAE, our continual vision-to-touch learning extends the pretrained video model to touch using roughly 100 demonstrations per task without tactile midtraining, while retaining visual prediction quality within 0.5 dB of vision-only counterparts. The compressor retains 89.4% of pre-fusion contact recall while enabling 2.26x faster training and 1.29x faster inference. Together, these results show that pretrained video priors can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner.
Chinese Translation
灵巧操作依赖于接触动力学,而这些动力学信息往往无法通过视觉完全观测。近期的世界-动作模型将预测式视频世界建模与动作生成相结合,但主要以视觉为中心,因而无法直接对接触动力学进行建模。我们提出了 DexTacWAM,一种视触觉世界-动作模型,它对每个指尖进行独立编码,通过手指与位姿感知的触觉压缩器聚合所得特征,并将触觉潜在表示注入视频扩散世界模型中,实现视觉与触觉的联合世界建模。在一个22自由度双臂平台上开展的六项接触密集型灵巧操作任务中,DexTacWAM 在每项任务上均取得最高分数,平均得分为 70.6,而最强基线仅为 38.0。消融实验将这一提升归因于将接触演化作为预测世界状态的一部分进行建模,而不仅仅是触觉条件化:在保持相同的触觉特征和动作专家的情况下,移除触觉世界建模会使四项任务的平均得分从 74.7 降至 26.6。在冻结预训练视觉 VAE 的条件下,经过四小时的触觉编码器适配,我们的持续视觉到触觉学习无需触觉中期训练,即可将预训练视频模型扩展至触觉模态,每项任务仅需约 100 个演示样本,同时视觉预测质量与纯视觉模型相比差距在 0.5 dB 以内。该压缩器保留了 89.4% 的融合前接触召回率,同时使训练速度提升 2.26 倍、推理速度提升 1.29 倍。总之,这些结果表明,预训练视频先验能够以高数据与计算效率的方式扩展至分布式多手指接触动力学。
cs.RO / 181 / 2609.24995

MIGU: Multimodal Instruction Grounding under Uncertainty for Manipulation Planning

MIGU:不确定性下的多模态指令定位与操作规划
Lu, Mingke, Xiao, Anxing, Hsu, David
Abstract
Understanding natural human instructions is crucial for deploying robots in human-centric environments. We study multimodal instruction grounding, where language and gesture provide complementary but uncertain cues. We present MIGU, a modular framework that combines semantic and geometric evidence into a unified grounding belief and connects it to manipulation planning. MIGU constructs a 3D geometric likelihood by propagating viewing-direction and depth uncertainty through eye-finger geometry while accounting for hand-direction estimation error. A vision-language model (VLM) provides semantic priors over candidate objects and regions, which are combined with the geometric likelihood through Bayes-inspired fusion. The resulting belief supports behavior planning to either proceed directly to downstream planning or request clarification. Grounded targets then define goals for mobile manipulation and tabletop task-and-motion planning. On a real-world benchmark, MIGU outperforms all evaluated baselines, while ablations support the benefit of explicit multimodal uncertainty modeling. Project website: multimodal-instruction.github.io
Chinese Translation
理解自然的人类指令对于在以人为中心的环境中部署机器人至关重要。我们研究了多模态指令定位问题,其中语言和手势提供互补但存在不确定性的线索。我们提出了MIGU,一个模块化框架,它将语义证据和几何证据结合为统一的定位信念(grounding belief),并将其与操作规划相连接。MIGU通过眼-手几何关系传播视向和深度的不确定性,同时考虑手部方向估计误差,从而构建三维几何似然。视觉-语言模型(VLM)为候选物体和区域提供语义先验,并通过受贝叶斯启发的融合方法将其与几何似然相结合。由此得到的信念支持行为规划,可以直接进行下游规划,或请求澄清。定位得到的目标随后为移动操作以及桌面任务与运动规划定义目标。在真实世界基准测试中,MIGU优于所有评估的基线方法,消融实验也验证了显式多模态不确定性建模的益处。项目网站:multimodal-instruction.github.io
cs.RO / 182 / 2609.24996

Learning Beyond What Humans Can Demonstrate

学习超越人类所能演示的技能
Song, Yuchen, Mittal, Aditya, Jain, Unnat
Abstract
Behavior cloning for robot manipulation relies on expert demonstrations. However, for tasks that require dynamic stability, precise contact timing, or dexterous coordination, human operators may find it hard or even impossible to collect data. We study this infeasible-demonstration regime and propose GLIDE: Guardrails for Learning from Infeasible Demonstrations Efficiently, a framework that infers task-specific failure modes and converts them into executable guardrails for data collection and policy deployment. Given a task description and the conditioning teleoperation code, GLIDE writes guardrails that use system states to filter teleoperation and policy commands, constrain failure-prone actions, and iteratively improve from trajectory feedback. Across three tasks, GLIDE discovers emergent guardrails that go beyond domain-expert hardcoded ones, improving data collection over naive VR teleoperation and domain-expert hardcoded guardrails. After refinement, GLIDE raises data-collection success from 0-10 percent to 70-90 percent across the three tasks. During policy execution, mixed-data guarded policies reach 70 percent, 60 percent, and 60 percent success on Tomato plate transfer, Marker handover and stand, and Wine serving tasks. These results show that GLIDE can support policy learning when direct demonstrations are infeasible. Project website: http://guardrail-policy.github.io/
Chinese Translation
面向机器人操作的行为克隆依赖于专家演示数据。然而,对于需要动态稳定性、精确接触时机或灵巧协调的任务,人类操作者可能难以甚至无法采集数据。我们研究了这种“不可行演示”的情形,并提出 GLIDE(Guardrails for Learning from Infeasible Demonstrations Efficiently,面向不可行演示的高效学习护栏):一个能够推断任务特定失败模式、并将其转化为可用于数据采集与策略部署的可执行护栏的框架。给定任务描述和条件遥操作代码,GLIDE 会编写护栏,利用系统状态过滤遥操作与策略指令、约束易失败的动作,并根据轨迹反馈迭代改进。在三个任务中,GLIDE 发现的涌现式护栏超越了领域专家手工编写的护栏,相比朴素 VR 遥操作和领域专家硬编码护栏显著提升了数据采集效果。经过优化后,GLIDE 将三个任务的数据采集成功率从 0–10% 提升至 70–90%。在策略执行阶段,混合数据训练的护栏策略在番茄餐盘转移、马克笔交接与立起、红酒伺服三个任务上分别达到 70%、60% 和 60% 的成功率。这些结果表明,GLIDE 能够在直接演示不可行的情况下支持策略学习。项目网站:http://guardrail-policy.github.io/
人工智能 (Artificial Intelligence)
111
cs.AI / 1 / 2609.22161

Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models

教学知识还是临床病例?数据类型如何塑造医学大语言模型
Fan, Yuzheng, Wang, Haochun, Zhao, Sendong, Han, Xiao, Ma, Ming, Qin, Bing
Abstract
Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how these data types differentially shape model capabilities remains unclear. We address this issue with token-matched experiments that vary the didactic-to-clinical ratio and analyze how data composition affects performance, capability profiles, and error patterns across knowledge-intensive and clinic-oriented tasks. We uncover an asymmetric transfer across task types: clinical data improves clinic-oriented tasks while remaining competitive on knowledge-intensive ones, whereas didactic data mainly improves knowledge-intensive tasks. Error analysis suggests a knowing-doing gap, where improvements in knowledge recall do not reliably generalize to clinical reasoning. We further observe that modest amounts of clinical data yield most of the gains on EHR-grounded tasks, while the optimal mixture ratio varies with the knowledge and clinical reasoning demands of downstream tasks. These findings suggest that medical LLM data curation should be application-driven, with higher proportions of clinical data preferred for reasoning-intensive use cases.
Chinese Translation
医学大语言模型通常在教学数据(如教科书)与临床数据(如患者病历)的混合数据上进行训练,然而这些数据类型如何差异化地塑造模型能力仍不清楚。我们通过token数量匹配的实验来解决这一问题:改变教学数据与临床数据的比例,并分析数据构成如何影响模型在知识密集型任务和临床导向任务上的性能、能力分布及错误模式。我们发现了跨任务类型的不对称迁移现象:临床数据在提升临床导向任务的同时,在知识密集型任务上仍保持竞争力;而教学数据主要提升知识密集型任务。错误分析揭示了一种“知行差距”(knowing-doing gap),即知识记忆方面的提升并不能可靠地泛化到临床推理中。我们进一步观察到,少量临床数据即可在基于电子病历(EHR)的任务上带来大部分收益,而最优混合比例随下游任务的知识与临床推理需求而变化。这些发现表明,医学大语言模型的数据策划应以应用为导向,对于推理密集型的应用场景,应优先采用更高比例的临床数据。
cs.AI / 2 / 2609.22277

An Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users

一种面向视障用户多模式出行辅助的经济型人工智能集成智能盲杖
Akarma, Ali, Ahmad, Adeel, Syed, Toqeer Ali
Abstract
Visual impairment affects over 2.2 billion people worldwide, yet conventional white canes cannot detect elevated hazards or provide semantic environmental context. Existing AI-assisted navigation systems typically rely on expensive hardware or cloud connectivity, limiting accessibility in resource-constrained settings. This paper presents an affordable (\$88 USD), fully offline AI-integrated smart cane designed for multimodal mobility assistance on an ultra-low-power Raspberry Pi Zero 2W. The system fuses RGB vision sensing with Time-of-Flight (ToF) distance estimation, pairing an INT8-quantized SSD MobileNet V1 model with distance-aware vibrotactile feedback and real-time audio alerts. To ensure operational robustness on constrained hardware, a multiprocessing architecture isolates sensor acquisition, neural inference, and haptic feedback into independent processes with fail-safe sensing support. Experimental evaluation across indoor mobility scenarios demonstrates a macro-averaged F1-score of 0.82 (precision: 0.85, recall: 0.81), a mean end-to-end latency of 330\,ms, and a peak power draw of 2.8\,W. A preliminary usability study with 12 participants (SUS: 78.5, NASA-TLX) demonstrated positive user perception and enhanced obstacle awareness. The proposed prototype validates the feasibility of deploying privacy-preserving, edge-native assistive intelligence for cost-sensitive mobility assistance.
Chinese Translation
全球有超过22亿人受到视力障碍的影响,然而传统的白手杖无法检测高处障碍或提供语义化的环境信息。现有的AI辅助导航系统通常依赖昂贵的硬件或云连接,限制了其在资源受限环境中的可及性。本文提出了一种经济实惠(88美元)、完全离线运行的AI集成智能盲杖,旨在超低功耗的Raspberry Pi Zero 2W上实现多模式出行辅助。该系统将RGB视觉传感与飞行时间测距相结合,将INT8量化的SSD MobileNet V1模型与距离感知的振动触觉反馈和实时音频警报相结合。为确保在受限硬件上的运行稳健性,系统采用多进程架构,将传感器采集、神经网络推理和触觉反馈隔离为独立进程,并支持故障安全传感。在室内出行场景中的实验评估表明,系统的宏平均F1分数为0.82(精确率:0.85,召回率:0.81),平均端到端延迟为330毫秒,峰值功耗为2.8瓦。一项包含12名参与者的初步可用性研究(SUS:78.5,NASA-TLX)显示了积极的用户感知和增强的障碍物感知能力。所提出的原型验证了为成本敏感的出行辅助部署隐私保护、边缘原生辅助智能的可行性。
cs.AI / 3 / 2609.22353

PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation

PAANI:面向河流机器人仿真的设备端视觉证据融合与可解释引导
Cardoz, Savio, Rajan, Santhiya
Abstract
Mobile river monitoring robots must interpret obstacles and water boundaries that geographic waypoints alone cannot describe. On resource constrained platforms, converting imperfect visual predictions into timely and inspectable guidance is a distinct challenge. An object label or steering command does not explain which evidence supports a decision or when that evidence is unreliable. We present PAANI, an on-device perception to guidance architecture that combines a project trained YOLO11n detector and a custom MobileNetV3 Small semantic segmenter with timestamp aligned evidence fusion on Arduino UNO Q. Bounded tracking supplies object persistence, while an explicit corridor policy combines surface labels, accepted detections, urgency and mask uncertainty. Each final advisory exposes its contributing evidence and policy reasons. ROS 2 interfaces connect the local AI pipeline to a separate Gazebo vessel, localization and control testbed. Training uses 10,000 WaterScenes images for four-class detection and 1,127 MaSTr1325 images for segmentation, including 198 segmentation validation images. The selected FP32 ONNX models occupy 14.817 MB. Detector checkpoint test mAP at 0.5 IoU is 0.7388, while the separately evaluated rectangular ONNX export achieves validation mAP at 0.5 IoU of 0.7367. Segmentation ONNX validation mIoU is 0.9750. A five-minute UNO Q recording produced median and 95th percentile pipeline latencies of 467.8 ms and 580.3 ms at a configured 0.5 Hz cadence. The evaluation also identifies black input misclassification and a sampling rate mismatch that prevents the diagnostic apparent motion estimator from collecting sufficient evidence. These results support an inspectable and reusable edge robotics foundation while clearly distinguishing model accuracy and on-board execution from validated on-water collision avoidance.
Chinese Translation
移动河流监测机器人必须识别仅靠地理航点无法描述的障碍物和水域边界。在资源受限的平台上,将不完善的视觉预测转化为及时且可检验的引导是一项独特的挑战。目标标签或转向指令无法解释哪些证据支持了某个决策,也无法说明该证据何时不可靠。我们提出了PAANI,一种设备端感知到引导的架构,它将项目训练的YOLO11n检测器和定制的MobileNetV3 Small语义分割器与时间戳对齐的证据融合相结合,运行于Arduino UNO Q之上。有界跟踪提供目标持久性,而显式的航道策略则综合表面标签、已接受的检测结果、紧迫性以及掩膜不确定性。每条最终建议都会揭示其贡献的证据和策略依据。ROS 2接口将本地AI流水线连接到独立的Gazebo船舶、定位与控制测试平台。训练使用10,000张WaterScenes图像用于四类检测,以及1,127张MaSTr1325图像用于分割,其中包括198张分割验证图像。所选的FP32 ONNX模型占用14.817 MB。检测器检查点在0.5 IoU下的测试mAP为0.7388,而单独评估的矩形ONNX导出模型在0.5 IoU下的验证mAP为0.7367。分割ONNX的验证mIoU为0.9750。在配置为0.5 Hz节拍下,一段五分钟的UNO Q录像产生的流水线延迟中位数为467.8 ms,第95百分位数为580.3 ms。评估还识别出了黑色输入的错误分类问题,以及一个采样率不匹配问题,该问题导致诊断性表观运动估计器无法收集足够的证据。这些结果支持了一个可检验、可复用的边缘机器人基础,同时明确区分了模型精度与设备端执行,以及经过验证的水上碰撞避免能力。
cs.AI / 4 / 2609.22408

Social Influence and the Allocation of Scientific Attention in AI Populations

社会影响与人工智能群体中科学注意力的分配
Chupilkin, Maxim
Abstract
AI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human readers, but the collective consequences of these signals for artificial readers remain uncertain. This paper adapts the Music Lab design to a market for academic attention. In the first experiment, 1,000 AI agents choose papers from the titles and abstracts of all 114 regular research articles published in the American Economic Review in 2025. The experiment has five independent-choice communities and five social-influence communities, each with 100 sequential agents. Only agents in the social-influence condition observe earlier selections within their community. Agents may select any number of papers. Social-information communities select 17.2 percent fewer papers per agent, concentrate their choices more heavily, and collectively cover 73 papers, compared with 90 independently. Between-community variation is greater under social information. In a second experiment with 200 agents across twenty social communities, randomly assigning papers five initial selections raises their subsequent selection rate by 45.55 percentage points (95% CI: 41.20 to 49.90). Choices have modest correspondence with external citations and little correspondence with download counts. The results show how a simple information rule shapes the volume, breadth and distribution of scientific attention in an artificial population.
Chinese Translation
AI系统正日益成为科学研究成果评估与使用的参与者。它们接触到的是围绕人类读者构建的引用次数、下载统计和热门文章列表,但这些信号对人工读者所产生的集体性后果仍不确定。本文将“音乐实验室”实验设计改造应用于学术注意力市场。在第一个实验中,1000个AI智能体根据《美国经济评论》2025年发表的114篇常规研究论文的标题和摘要来选择论文。实验设有五个独立选择社区和五个社会影响社区,每个社区包含100个顺序决策的智能体。仅社会影响条件下的智能体能够观察到其所在社区中先前成员的选择。智能体可以选择任意数量的论文。社会信息社区中每个智能体平均选择的论文数量减少17.2%,其选择更加集中,合计覆盖73篇论文,而独立选择条件下为90篇。社会信息条件下社区间的差异更大。在第二个实验中,200个智能体分布在二十个社会社区中,随机为论文赋予五个初始选择,可使后续选择率提高45.55个百分点(95%置信区间:41.20至49.90)。智能体的选择与外部引用量存在一定程度的对应关系,但与下载量的对应关系很弱。研究结果表明,一条简单的信息规则如何塑造了人工群体中科学注意力的数量、广度与分布。
cs.AI / 5 / 2609.22410

Learning 3D biophysical cell properties from 2D images and cell-population statistics

从二维图像和细胞群体统计学习三维细胞生物物理特性
Hernández-Orozco, Santiago, Zenil, Hector
Abstract
Inferring 3D cellular properties from 2D microscopy is difficult when a reference instrument reports only population statistics rather than labels for individual cells. Here we develop a population-supervised framework that maps single 2D red-cell images to latent biophysical quantities and aggregates them to mean corpuscular volume, red-cell distribution width and mean corpuscular haemoglobin. The model combines shared local inference, a biophysically structured decoder for volume and haemoglobin, learned instance weighting and device-specific calibration. We formalise conditions under which aggregate observations identify restricted instance predictors, show why population agreement does not by itself identify single-cell properties or 3D geometry, and derive the dispersion penalty induced by subset mean matching. The development dataset comprises 390 specimens and 1,105 acquisitions across six devices, with reported Pearson correlations of 0.86--0.98 against a Sysmex analyser. The framework provides a testable route from 2D images and population supervision to 3D cellular biophysics without claiming explicit 3D reconstruction.
Chinese Translation
当参考仪器仅报告群体统计量而非单个细胞的标签时,从二维显微图像推断三维细胞特性十分困难。本文开发了一个群体监督框架,将单个二维红细胞图像映射到潜在生物物理量,并将其聚合为平均红细胞体积(MCV)、红细胞分布宽度(RDW)和平均红细胞血红蛋白含量(MCH)。该模型结合了共享的局部推断、针对体积和血红蛋白的生物物理结构化解码器、学习型实例加权以及设备特定的校准。我们形式化了聚合观测能够识别受限实例预测器的条件,说明了为何仅凭群体一致性无法识别单细胞特性或三维几何结构,并推导了子集均值匹配所诱导的离散度惩罚。开发数据集包含390个样本和跨六种设备的1105次采集,相对于Sysmex分析仪报告的皮尔逊相关系数为0.86--0.98。该框架提供了一条从二维图像和群体监督到三维细胞生物物理的可验证路径,且不声称进行显式的三维重建。
cs.AI / 6 / 2609.22475

Goal-driven Variant Categorization

目标驱动的流程变体分类
Calegari, Daniel, Amyot, Daniel
Abstract
Process discovery rarely yields a single coherent process structure. For analysis, a common step is to cluster process variants based on structural similarity and then assign business meaning to the resulting groups. Since these partitions are not derived from the organization's goals, analysts must manually interpret and consolidate variants into business-meaningful categories. This judgment-intensive step becomes increasingly difficult as the number and complexity of variants grow. In this paper, we propose a goal-driven approach to variant categorization that reverses this workflow. We first author an organization's goal model that predefines the categorization axis. Each variant is transformed into a textual narrative describing its behavior, and a Large Language Model (LLM) interprets it in the context of the goal model and assigns the variant to the most appropriate category. LLM-based semantic reasoning connects low-level process behavior with analyst-defined business goals. We instantiate this approach end-to-end and evaluate it on three public logs differing substantially in scale and behavioral diversity. Goal-model guidance yields partitions that differ from those produced by unguided induction and respond to controlled edits to the declared alternatives, at the cost of authoring a goal model.
Chinese Translation
流程发现很少能够产生单一且连贯的流程结构。在分析过程中,一个常见步骤是基于结构相似性对流程变体进行聚类,然后为所得的分组赋予业务含义。由于这些划分并非源自组织的目标,分析人员必须手动解释并将变体整合为具有业务意义的类别。随着变体数量和复杂性的增长,这一需要大量判断的步骤变得越来越困难。本文提出了一种目标驱动的变体分类方法,反转了这一工作流程。我们首先构建组织的目标模型,预先定义分类轴。每个变体被转换为描述其行为的文本叙述,然后由大型语言模型(LLM)在目标模型的上下文中对其进行解读,并将该变体分配到最合适的类别中。基于LLM的语义推理将底层流程行为与分析人员定义的业务目标联系起来。我们对该方法进行了端到端的实例化,并在三个在规模和行为多样性上差异显著的公开日志上进行了评估。目标模型的引导能够产生不同于无引导归纳所得的划分,并能响应于对所声明的备选项的受控编辑,其代价是需要构建目标模型。
cs.AI / 7 / 2609.22478

Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation

托管式大语言模型中无持久性的复现:行动时信念评估中的测量敏感性
Joshi, Bhushan Kashinath
Abstract
Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration (replication), whether the endpoint changes when the evaluation-and-inference configuration is rebuilt under the same identifier (measurement sensitivity), and whether the finding persists across subsequently tested identifiers under one common instrument (persistence). We study these questions in Regent Chess, a sequential environment in which a hidden, mutable state is recorded exactly, allowing stated beliefs to be scored against ground truth at action time; positive endpoint values mean worse performance than a matched-uniform comparator. The previously reported Gemini 3.1 Flash-Lite deficit recurs on fresh games under its historical configuration (+0.0530, 95% CI [+0.0329,+0.0714]). In a back-to-back same-day H/R comparison under the same public identifier, the model-minus-uniform endpoint is 0.0429 lower under the rebuilt configuration (95% CI for the H-minus-R contrast [+0.0182,+0.0667]); all six configuration components vary jointly, so no component is isolated. Under rebuilt R, the prospectively frozen, interleaved same-window 4K comparison reverses sign between Gemini 3.1 and Gemini 3.7, identifiers that differ in release and product tier; additional descriptive and exploratory cells show the same directional pattern. Any additional serving-period contribution remains unresolved (-0.0166, [-0.0483,+0.0157]). Replication, measurement sensitivity, and persistence can therefore yield different conclusions within one evaluation, motivating explicit indexing of hosted-model behavioural claims by tested identifier, serving period, measurement instrument, and inference configuration.
Chinese Translation
对托管式语言模型的行为评估可能因被评估服务、测量工具或两者在不同运行中存在差异而产生变化。我们区分了三个验证问题:先前发现在其历史配置下能否在新数据上重现(复现,replication);在相同标识符下重建评估与推理配置时端点指标是否发生变化(测量敏感性,measurement sensitivity);以及在同一测量工具下,该发现能否在后续测试的标识符间保持(持久性,persistence)。我们在 Regent Chess(一个顺序决策环境)中研究这些问题。在该环境中,隐藏且可变的状态被精确记录,使得模型声明的信念可以在行动时刻与真值(ground truth)进行比对打分;端点指标为正值意味着性能劣于匹配均匀分布的对照基线。先前报告的 Gemini 3.1 Flash-Lite 的性能缺陷在其历史配置下的新对局中重现(+0.0530,95% CI [+0.0329,+0.0714])。在相同公共标识符下的同日背靠背 H/R 对比中,重建配置下的“模型减均匀基线”端点降低了 0.0429(H 减 R 对比的 95% CI 为 [+0.0182,+0.0667]);由于六个配置组件同时联合变动,因此无法分离出任何单一组件的影响。在重建的 R 配置下,前瞻性冻结的、交错同时间窗口的 4K 对比在 Gemini 3.1 与 Gemini 3.7 之间出现了符号反转,而这两个标识符在发布版本和产品层级上均不同;其他描述性与探索性分析单元也显示出相同的方向性模式。服务期可能带来的额外贡献仍未得到解决(-0.0166,[-0.0483,+0.0157])。因此,复现、测量敏感性和持久性在同一评估中可能得出不同结论,这提示我们应对托管式模型的行为学结论进行显式索引,即按被测标识符、服务期、测量工具和推理配置进行标注。
cs.AI / 8 / 2609.22497

The Wisdom of Artificial Deliberative Crowds

人工审慎群体(Artificial Deliberative Crowds)的智慧
Barrera-Lemarchand, Federico, Sigman, Mariano, Navajas, Joaquin
Abstract
The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as the wisdom of crowds. While this is usually attributed to the independence of estimates, an even stronger effect arises through deliberation: averaging the consensus estimates of small deliberating groups outperforms the classical wisdom of crowds, with individual judgments themselves also becoming more accurate after deliberation. Whether these improvements transfer to large language models deliberating amongst themselves is unknown. Here we adapt a three-stage deliberation paradigm previously used with human participants for use with large language models from three different families, and test it across four domains of increasing real-world stakes: visual numerical estimation (Study 1), peer review of machine-learning papers (Study 2), detection of hidden malicious behavior by an artificial intelligence agent (Study 3), and sports forecasting against a real prediction market (Study 4). Across domains, deliberation reduced collective error beyond passive aggregation of independent responses, and post-deliberation individual judgments retained this collective gain. Notably, the advantage required model diversity: groups composed of clones of a single model did not benefit from deliberating. These results establish machine deliberation as a general-purpose aggregation mechanism, and point to diversity as an active ingredient.
Chinese Translation
将众多非专业者的估计进行聚合,其表现往往优于个体专家的判断,这一现象被称为群体智慧(wisdom of crowds)。尽管这通常被归因于估计的独立性,但通过商讨(deliberation)会产生更强的效果:对小型商讨小组的共识估计取平均值,其表现优于经典的群体智慧,且个体判断本身在商讨后也变得更加准确。这些改进是否也能迁移到大型语言模型之间的相互商讨中,目前尚不清楚。本研究将此前用于人类被试的三阶段商讨范式加以改造,应用于来自三个不同家族的大型语言模型,并在四个现实风险逐级递增的领域中进行测试:视觉数值估计(研究1)、机器学习论文的同行评审(研究2)、检测人工智能智能体的隐藏恶意行为(研究3),以及与真实预测市场对比的体育赛事预测(研究4)。在各个领域中,商讨均降低了集体误差,效果超越了独立回答的被动聚合,且商讨后的个体判断保留了这一集体增益。值得注意的是,这种优势依赖于模型多样性:由单一模型克隆组成的群体无法从商讨中获益。这些结果确立了机器商讨作为一种通用的聚合机制,并表明多样性是其有效成分。
cs.AI / 9 / 2609.22512

Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

一致性高估了证据强度:大语言模型裁判共识中的误差相关性
Hossain, Elias, Yousefi, Niloofar, Lim, Ser-Nam
Abstract
Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes. We study how this dependency affects the reliability of consensus. We find substantial error correlation across both open-weight and frontier LLM judges. In our main bank of ten judges, the average pairwise correlation between judge errors is 0.21. As a result, the ten judges only provide roughly as much statistical information as 3.5 independent judges. The dependency is even stronger among the high-accuracy frontier judges we evaluate, including judges from different providers. In up to 28% of our comparisons, ignoring shared errors leads to the conclusion that one system is significantly better, while accounting for them does not. We also find that the pattern of errors matters. Errors shared by most judges and errors concentrated among a smaller group affect consensus differently and favor different voting methods. Measuring the overall amount of correlation alone is therefore insufficient. Our results suggest a simple approach: use a small set of trusted examples to estimate judge accuracy and identify shared mistakes. These shared errors should then be considered when analyzing the results, and the voting method should be chosen using trusted examples before it is applied to new data.
Chinese Translation
大语言模型(LLM)裁判之间的共识通常被视为决策正确性的有力证据,这背后隐含着裁判误差相互独立的假设。然而在实践中,LLM 裁判往往以相似的方式进行训练和评估,因此可能犯相同的错误。我们研究了这种依赖性如何影响共识的可靠性。我们发现,在开源权重模型和前沿模型中,LLM 裁判之间存在显著的误差相关性。在我们的主裁判池(共十名裁判)中,裁判误差的平均两两相关系数为 0.21。因此,这十名裁判所能提供的统计信息量大约只相当于 3.5 名独立裁判。在我们评估的高准确率前沿裁判中,这种依赖性甚至更强,包括来自不同提供商的裁判。在多达 28% 的比较中,忽略共享误差会得出某一系统显著更优的结论,而考虑这些误差后该结论则不成立。我们还发现误差的模式也很重要:被大多数裁判共享的误差与集中于少数裁判的误差对共识的影响不同,且各自偏好不同的投票方法。因此,仅仅衡量总体相关性是不够的。我们的结果提出了一种简单的方法:利用少量可信样本估计裁判的准确率并识别共享的错误,在分析结果时应考虑这些共享误差,并且在将投票方法应用于新数据之前,应基于可信样本来选择投票方法。
cs.AI / 10 / 2609.22529

IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law

IntLawNER:一个国际法领域的命名实体识别数据集与基准
Skura, Genis, Bouffanais, Roland, Wernli, Didier
Abstract
International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its texts remain without token-level named entity recognition (NER) resources. We introduce IntLawNER, a NER dataset and benchmark for codified sources of international law, covering 2,987 gold-annotated sentences and 8,094 entity spans from International Court of Justice (ICJ) decisions, UN Security Council resolutions, and European Court of Human Rights (ECtHR) judgments, annotated with seven institution-specific entity types. We construct IntLawNER with a cost-effective hybrid algorithmic-agentic pipeline that reduces 468k source sentences to a compact annotation set through candidate retrieval, LLM-based vetting, and human review, with 89.6% of gold spans accepted unchanged from the silver layer. However, the silver-to-gold analysis reveals that human-machine aggregate agreement metrics can be misleading in domain-specific NER: Cohen's kappa=0.964 on boundary-matched spans masks a macro-F1 of 0.753 when missing entities, boundary errors, and label corrections are included. The benchmark shows that zero-shot span-based GLiNER collapses on entity types dependent on institutional function rather than surface form (0.243 micro-F1), while fine-tuned transformers struggle on rare labels. Carefully selected few-shot examples that demonstrate label contrasts improve every LLM over zero-shot prompting, with Claude Opus 4.6 reaching the best score of 0.873 micro-F1. We release IntLawNER as a benchmark and reusable resource for extracting references in international legal texts.
Chinese Translation
国际法为国家协调行动、规范武装冲突和保护人权提供了规范框架,然而其文本至今仍缺乏词元级别的命名实体识别(NER)资源。我们提出了 IntLawNER,一个面向国际法成文来源的命名实体识别数据集与基准,涵盖 2,987 个黄金标注句子和 8,094 个实体跨度,数据来源包括国际法院(ICJ)判决、联合国安理会决议以及欧洲人权法院(ECtHR)判决,并标注了七种机构特定的实体类型。我们通过一个经济高效的算法-智能体混合流水线构建了 IntLawNER:通过候选检索、基于大语言模型(LLM)的筛选和人工审核,将 468,000 条源句子缩减为一个紧凑的标注集合,其中 89.6% 的黄金跨度直接沿用银层标注而未作修改。然而,银层到金层的分析表明,人机总体一致性指标在领域特定的 NER 任务中可能具有误导性:在边界匹配的跨度上,Cohen's kappa 达到 0.964,但若将实体遗漏、边界错误和标签修正纳入考量,宏观 F1 仅为 0.753。基准测试显示,零样本的基于跨度的 GLiNER 在依赖机构功能而非表面形式的实体类型上表现崩溃(微观 F1 仅为 0.243),而微调后的 Transformer 模型在稀有标签上表现不佳。经过精心挑选、能够展示标签对比的少样本示例能够使所有大语言模型的表现均优于零样本提示,其中 Claude Opus 4.6 取得了 0.873 微观 F1 的最佳成绩。我们将 IntLawNER 作为一个基准和可复用资源发布,用于国际法文本中的指称提取。
cs.AI / 11 / 2609.22537

EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability

EvidenT:通过证据的忠实性与可追溯性构建可信的企业智能助手
Kabra, Anubha, Kim, Katie Jooyoung, Kou, Colin Zhiwei, Sajer, Helene, Fan, Yimei, Cisar, Radomir, Greenhalgh, Heather, Vidiri, Gabriel Martinez
Abstract
Enterprise AI assistants must produce responses that are verifiable and traceable to source evidence. However, retrieval augmented generation (RAG) over heterogeneous enterprise data can suffer from citation drift, unsupported content, and weak source traceability. We present EvidenT (T = Trust + Transparency + Traceability), a lightweight pipeline that verifies extracted evidence against retrieved documents before answer generation, without model retraining. EvidenT combines structured passage extraction with deterministic lexical alignment to filter unsupported content, correct citation drift, and preserve source-span traceability. On approximately 500 real enterprise queries, EvidenT improves gold-source hit rate by an average of 29% over prompting baselines, produces no citations to nonretrieved urls, and achieves near-saturated answer-to-source lexical coverage.
Chinese Translation
企业AI助手必须生成可验证且可追溯至源证据的回答。然而,面向异构企业数据的检索增强生成(RAG)可能存在引用漂移、无依据内容以及源可追溯性弱等问题。我们提出了EvidenT(T = 可信 + 透明 + 可追溯),这是一个轻量级流水线,在答案生成之前针对检索到的文档验证所提取的证据,且无需模型再训练。EvidenT将结构化段落提取与确定性的词法对齐相结合,以过滤无依据内容、纠正引用漂移并保留源片段的可追溯性。在约500个真实企业查询上,与提示基线相比,EvidenT将黄金源命中率平均提升了29%,未产生任何指向未检索URL的引用,并实现了接近饱和的答案到源词法覆盖率。
cs.AI / 12 / 2609.22592

AutoGym: Blueprint-First Generation of Verifiable Agent Gyms

AutoGym:以蓝图优先的方式生成可验证的智能体训练环境(Gym)
Noronha, Aarati Andrea, Ravikumar, Kavya, Lin, Carly Xiaoyu
Abstract
Training agents with reinforcement learning requires a gym, comprising a task, an executable environment in which the task can be attempted, and a verifier that reliably distinguishes success from failure. Constructing such gyms remains manual, expensive, and static. Task sets saturate as models improve and are increasingly exposed to contamination. Synthetic generation offers scale, but single-pass synthesis produces tasks whose difficulty is largely cosmetic. Models comparable in capability solve them despite convoluted phrasing, and correctness must be adjudicated post-hoc by unreliable LLM judges. We present AutoGym, a framework that generates complete gyms (tasks, executable environments, and verifiers) from a minimal domain seed or prior model trajectories. AutoGym introduces three mechanisms. (1) Blueprint-first generation specifies the valid solution space, environment requirements, and verification criteria before the environment is materialized, making solvability a construction prerequisite rather than a property verified after the fact. (2) Explicit generation parameters control task topology, interaction depth, capability axes, question obfuscation, and distractor composition, enabling fine-grained difficulty steering. (3) Active curriculum synthesis uses performance-informed calibration to adjust the distribution over these parameters as model capabilities evolve. Across productivity and temporal-reasoning settings, AutoGym generates gyms spanning the capability spectrum, including instances that challenge frontier models.
Chinese Translation
使用强化学习训练智能体需要一个训练环境(gym),其包含任务、一个可执行的环境(任务可在其中被尝试完成),以及一个能够可靠区分成功与失败的验证器。构建这样的训练环境目前仍依赖人工,成本高昂且静态固化。随着模型能力的提升和训练数据日益受到污染,任务集会逐渐饱和。合成生成虽然提供了规模优势,但单次合成的任务其难度大多流于表面:能力相近的模型尽管面对复杂的表述仍能将其解决,而任务正确性只能事后由不可靠的大语言模型裁判来裁定。我们提出 AutoGym,一个能够从最小领域种子或既有模型轨迹生成完整训练环境(任务、可执行环境与验证器)的框架。AutoGym 引入了三种机制:(1) 蓝图优先生成(Blueprint-first generation):在环境具体化之前先确定有效解空间、环境需求和验证标准,使任务可解性成为构建的前提条件,而非事后验证的属性;(2) 显式生成参数:控制任务拓扑、交互深度、能力维度、问题混淆和干扰项构成,实现细粒度的难度调控;(3) 主动课程合成(Active curriculum synthesis):利用基于性能信息的校准,随着模型能力的演进调整这些参数的分布。在生产力和时序推理两类场景中,AutoGym 生成的训练环境覆盖了整个能力谱系,其中包含能够挑战前沿模型的实例。
cs.AI / 13 / 2609.22599

MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators

MAWILE:用于检查LLM评估器的多轴工作台
Hassell, Jackson, Bayat, Farima Fatahi, Pezeshkpour, Pouya, Hruschka, Estevam
Abstract
Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated response, judge instructions, and scoring rubric. Existing systems examine important subsets of these failure modes, but auditing a configured judge requires testing both the judge instrument and the items it evaluates. We introduce MAWILE, a developer-facing workbench for auditing judge sensitivity across four surfaces: the judge prompt, judge rubric, target-system input, and target-system output. Given a user-supplied judge and representative evaluation items, MAWILE constructs and validates controlled perturbations, re-executes the judge, and localizes the resulting sensitivity. Each perturbation declares whether the verdict should remain invariant or change in a specified direction, allowing the same system to measure both robustness to irrelevant variations and sensitivity to meaningful changes. MAWILE audits binary, ordinal, and pairwise judges without requiring gold labels. The code for this tool is available at: github.com/megagonlabs/mawile-judge.
Chinese Translation
大语言模型(LLM)评审器(judge)为评估模型和智能体输出提供了一种灵活且可扩展的方法,但其判定结果可能对被评估回复、评审指令以及评分标准中的偶然变化较为敏感。现有系统考察了这些失效模式的重要子集,但审计一个配置好的评审器需要同时测试评审器本身及其所评估的对象。我们提出了MAWILE,一个面向开发者的工作台,用于从四个维度审计评审器的敏感性:评审提示词、评审评分标准(rubric)、目标系统的输入和输出。给定用户提供的评审器和具有代表性的评估样本,MAWILE构建并验证受控扰动,重新执行评审器,并定位由此产生的敏感性。每个扰动都声明判定结果应保持不变,还是应朝指定方向变化,从而让同一系统能够同时衡量对无关变化的鲁棒性和对有意义变化的敏感性。MAWILE可审计二元、序数和成对比较类评审器,且无需金标准标签。该工具的代码已发布于:github.com/megagonlabs/mawile-judge。
cs.AI / 14 / 2609.22619

GaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment

GaitVista:面向可及性纵向步态评估的可靠性感知AI测量方法
Jayasinghe, Nethmi, Parashar, Mihir, Trivedi, Amit Ranjan
Abstract
Tracking recovery of walking function requires detecting meaningful gait change across rehabilitation sessions, yet objective 3D measurement remains confined to specialized motion-capture laboratories. Small camera sets and body-worn inertial sensors broaden access, but reliability varies across joints and time, allowing sensing failures to masquerade as patient change. We present \textsc{GaitVista}, a reliability-aware measurement layer whose lightweight gate assigns joint- and frame-specific visual contributions using camera coverage, local visual quality, cross-modal disagreement, and root-motion continuity, and exposes them for inspection. Across seven clean and degraded sensing conditions on TotalCapture, \textsc{GaitVista} reduces average full-body and lower-body error by \textbf{27.7\%} and \textbf{27.8\%}, attains the lowest worst-condition error among fusion methods, and reduces the gap to a joint-frame oracle from $2.76$--$5.33$~cm for condition-blind baselines to $1.11$~cm. On MoVi with image-derived keypoints, it is the only deployable fusion method to improve over both unimodal streams, reducing marker-supported error by \textbf{6.4\%} relative to the strongest learned fusion baseline. On TotalCapture, it improves bilateral knee-flexion waveform accuracy by \textbf{18.9\%}. Raw inertial measurements from five TotalCapture participants show location- and time-varying magnetic disturbance, supporting the design's reliability premise. Both benchmarks contain neurologically healthy participants in controlled settings and retain participant-specific IMU calibration; we therefore report progress toward accessible gait assessment, not validated clinical deployment.
Chinese Translation
追踪行走功能的恢复需要检测跨康复疗程的具有临床意义的步态变化,然而客观的3D测量仍然局限于专业的动作捕捉实验室。少量相机和身体佩戴的惯性传感器扩大了测量的可及性,但其可靠性在不同关节和不同时间上存在差异,使得传感失效可能被误认为患者的真实变化。我们提出了GaitVista,一个可靠性感知的测量层,其轻量级门控机制利用相机覆盖范围、局部视觉质量、跨模态分歧以及根运动连续性,为每个关节和每一帧分配视觉贡献权重,并将这些权重暴露出来以供检查。在TotalCapture数据集的七种干净及退化传感条件下,GaitVista将全身和下肢的平均误差分别降低了27.7%和27.8%,在所有融合方法中取得了最低的最差条件误差,并将与逐关节逐帧oracle之间的差距从条件盲基线的2.76–5.33厘米缩小至1.11厘米。在MoVi数据集上使用从图像提取的关键点时,它是唯一一种能够同时超越两个单模态数据流的可部署融合方法,相对于最强的学习型融合基线,将基于标记点验证的误差降低了6.4%。在TotalCapture上,它将双侧膝关节屈曲波形精度提升了18.9%。来自五名TotalCapture参与者的原始惯性测量数据显示出随位置和时间变化的磁干扰,支持了该设计所依据的可靠性前提。两个基准数据集均包含在受控环境下的神经系统健康参与者,并保留了参与者特定的IMU校准;因此,我们报告的是向可及性步态评估迈进的进展,而非经过验证的临床部署。
cs.AI / 15 / 2609.22620

Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation

以更低成本拆分文档:基于大语言模型的页面流分割的多拆分边界决策
Pottanigari, Nikhil Reddy, Kharaghani, Sepideh, Vadacchino, Saverio, Posada, Alejandro, Zhang, Ying
Abstract
Scanned mail, uploaded PDFs, and consolidated attachments often arrive as page streams that must be split into individual documents before downstream classification, extraction, or routing. Zero-shot large language models can detect document boundaries without task-specific training, but standard Page Classification (PC) and Boundary Decision (BD) formulations resolve only one boundary per model call. We introduce Multi-Split Boundary Decision (MSBD), which predicts multiple boundaries within a page window in a single call, reducing the number of inference requests. We evaluate MSBD across multiple language models, document collections, input modalities, and window sizes. The results reveal a model- and corpus-dependent operating range in which MSBD preserves strong segmentation accuracy while substantially improving inference efficiency, followed by a sharp decline at larger windows. MSBD provided the strongest overall accuracy--efficiency trade-off, while large windows expose distinct over- and under-segmentation behavior across models. These findings show that multi-boundary prediction can make zero-shot page stream segmentation more efficient when the window size is selected for the target corpus.
Chinese Translation
扫描邮件、上传的PDF文件以及合并的附件通常以页面流的形式到达,需要在下游分类、抽取或路由之前将其拆分为单个文档。零样本(zero-shot)大语言模型无需任务特定训练即可检测文档边界,但标准的页面分类(Page Classification, PC)和边界决策(Boundary Decision, BD)方法每次模型调用只能解析一个边界。我们提出了多拆分边界决策(Multi-Split Boundary Decision, MSBD),它在单次调用中预测页面窗口内的多个边界,从而减少推理请求的数量。我们在多种语言模型、文档集合、输入模态和窗口大小上评估了MSBD。结果揭示了一个依赖于模型和语料库的有效工作范围,在该范围内,MSBD在显著提高推理效率的同时保持了较强的分割准确率,而在更大窗口下性能急剧下降。MSBD提供了最强的整体准确率—效率权衡,而大窗口在不同模型上表现出明显的过度分割和不足分割行为。这些发现表明,当针对目标语料库选择合适的窗口大小时,多边界预测可以使零样本页面流分割更加高效。
cs.AI / 16 / 2609.22628

Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA

文本、像素,还是两者兼用?多模态文档问答中输入表示形式的评估
Pottanigari, Nikhil Reddy, Kharaghani, Sepideh, Vadacchino, Saverio, Posada, Alejandro, Zhang, Ying
Abstract
Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this choice, holding the prompt, judge, and scoring pipeline fixed, across four commercial model endpoints, two corpora, and two context regimes (gold evidence pages and the full document). On documents that fit the image budget, page images lead on accuracy at every document length on both corpora, but this advantage carries a growing latency and cost premium: text latency stays roughly flat as documents lengthen while image latency rises steadily. Text and images also fail on different questions, with exactly one representation correct on 19--25% of items across the reported cells, so neither subsumes the other. Exploiting this complementarity, a lightweight TF-IDF router that reads only the question text gains 2.6 points over always-text while cutting median latency 30% relative to always-vision, on a document-disjoint held-out split.
Chinese Translation
每个文档问答系统的构建都始于一个很少被单独研究的选择:是向模型输入页面图像、提取的文本,还是两者兼用。我们单独考察这一选择,在四个商用模型端点、两个语料库和两种上下文设置(金标准证据页面与完整文档)下,保持提示词、评判器和评分流程不变。对于符合图像预算的文档,在两个语料库上,页面图像在所有文档长度下都取得了最高的准确率,但这一优势伴随着日益增加的延迟和成本代价:随着文档变长,文本延迟基本保持平稳,而图像延迟则持续上升。文本和图像在不同问题上各有失败,在所报告的各组实验中,恰有一种表示正确而另一种错误的情况占19%–25%,因此两者互不包含。利用这种互补性,一个仅读取问题文本的轻量级TF-IDF路由器,相比始终使用文本的方案提升了2.6个百分点,同时相比始终使用图像的方案将中位延迟降低了30%,且这是在文档不相交的留出集上取得的。
cs.AI / 17 / 2609.22682

Self-Organizing Agent Teams Learn to Reason Together

自组织智能体团队学会协同推理
Pappu, Aneesh, Suzgun, Mirac, Kwon, Yongchan, Bianchi, Federico, El, Batu, Kochenderfer, Mykel J., Cao, Hancheng, Zou, James
Abstract
Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasoning as it unfolds. Human teams routinely adapt this way, while existing AI agent teams rely on fixed protocols, explicit task decomposition, or routing. We introduce Self-Organizing Agent Teams (SAT), fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow. These strategies enable what we call collaborative computation: agents exchange, challenge, repair, and synthesize partial reasoning into solutions no member produced independently. In two independent settings, we learn teamwork strategies that transfer unchanged to unseen benchmarks, using only 15 mathematics and 25 graduate-level knowledge problems. Across five mathematics and physics benchmarks, self-organizing teams average 66.7% accuracy, versus 48.8% for their strongest member, 58.7% for compute-matched inference by that agent, and 59.0% for a perfect router over members' independent answers; on AIME 2026, they exceed this router by 13.4 points. Because gains vary across benchmarks, we ask when self-organizing collaboration helps. Across eight benchmarks, demonstrability (the organizational-psychology construct of whether a team can distinguish correct from incorrect reasoning) strongly tracks improvement over the strongest member (Spearman $\rho=0.90$, $p=0.005$): teams benefit most when correct reasoning can be recognized once it appears. More broadly, these results suggest that organization itself can become an agent capability: agent teams can learn how to reason together and produce solutions their members could not reach independently.
Chinese Translation
集体智能不仅取决于团队成员所掌握的知识,还取决于他们如何组织工作。当解决方案的结构未知时,无法预先指定有用的角色和分工;团队必须在推理展开的过程中从经验中学习如何组织推理。人类团队通常能够以这种方式进行适应,而现有的AI智能体团队则依赖于固定协议、显式任务分解或路由机制。我们提出了自组织智能体团队(Self-Organizing Agent Teams, SAT),即固定的AI智能体团队,它们从先前的协作中学习可复用的策略,以组织角色、对话阶段、参与方式与信息流动。这些策略实现了我们所说的协作计算:智能体之间交换、质疑、修复并综合局部推理,从而得出任何成员都无法独立产生的解决方案。在两个独立的环境中,我们仅使用15道数学题和25道研究生水平的知识题来学习团队协作策略,并且这些策略无需修改即可迁移到未见过的基准测试上。在五个数学和物理基准测试中,自组织团队的平均准确率为66.7%,而其最强成员为48.8%,该智能体的计算量匹配推理为58.7%,对成员独立答案的完美路由器为59.0%;在AIME 2026上,自组织团队超出该路由器13.4个百分点。由于收益在不同基准测试之间存在差异,我们进一步探讨自组织协作何时有效。在八个基准测试中,可证明性(组织心理学中衡量团队能否区分正确与错误推理的构念)与相对最强成员的提升高度相关(Spearman ρ=0.90,p=0.005):当正确推理一旦出现就能被识别时,团队受益最大。更广泛地说,这些结果表明,组织本身可以成为一种智能体能力:智能体团队可以学会如何协同推理,并产生其成员无法独立达成的解决方案。
cs.AI / 18 / 2609.22691

Generative Embodied Multiple Behavior Control Systems for Human-like Agents

面向类人智能体的生成式具身多行为控制系统
Bao, Chongyu, Yang, Haokai, Wang, Yuhan, An, Zhaochong, Liu, Kunpeng, Liu, Xiaolan
Abstract
An enduring and richly elaborated dichotomy in cognitive neuroscience is that of human behavior control mechanisms, divided into habitual versus goal-directed. While existing human-like agent frameworks primarily focus on modeling goal- directed behavior, habitual behavior has been largely overlooked, though it plays a crucial role in human daily life. In this paper, we address this gap by studying multiple behavior control systems that jointly model goal-directed and habitual behaviors. We propose a human behavior control mechanism-inspired framework which the Habitual Controller retrieves cue-triggered behaviors from personal- ized habit memory, while the Goal-directed Controller employs a context-aware world model to predict action consequences and estimate their values. The Arbiter dynamically balances the influence of both systems according to individual differ- ences and momentary internal states. To reconstruct diverse human-level behavior instructions in 3D environments, we further develop a keyframe-guided 3D mo- tion generation module. Through extensive evaluation methods, human studies, and ablations studies, experimental results demonstrate that human-likeness per- formance is significantly improved by our approach. The efficacy of our approach indicates the benefits of leveraging habitual behavior and multiple behavior con- trol system coordination for believable embodied human-like agents.
Chinese Translation
认知神经科学中一个持久且研究深入的二分法是人类行为控制机制,即习惯性行为与目标导向行为的划分。现有的类人智能体框架主要聚焦于对目标导向行为的建模,而习惯性行为虽然在人类日常生活中扮演着关键角色,却很大程度上被忽视。本文通过研究联合建模目标导向行为与习惯性行为的多行为控制系统来填补这一空白。我们提出了一种受人类行为控制机制启发的框架:习惯控制器从个性化习惯记忆中检索由线索触发的行为,而目标导向控制器则利用情境感知的世界模型来预测动作后果并估计其价值。仲裁器根据个体差异和瞬时内部状态动态平衡两个系统的影响力。为了在3D环境中重建多样化的类人行为指令,我们进一步开发了一个关键帧引导的3D动作生成模块。通过广泛的评估方法、人类研究和消融实验,实验结果表明我们的方法显著提升了类人表现。我们方法的有效性表明,利用习惯性行为以及多行为控制系统的协调,对于构建可信的具身类人智能体大有裨益。
cs.AI / 19 / 2609.22694

Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review

面向可审计与经校准的痴呆相关事故严重程度预测AI:一种支持人工审查的选择性延迟决策框架
Chhetri, Gaurab, Baitullah, Anika, Das, Subasish
Abstract
Public crash databases increasingly support automated safety analysis, but crash severity prediction remains difficult to translate into public-sector decision workflows when models are evaluated primarily as ordinary classifiers. This study reframes dementia-related crash severity modeling as a decision-aware triage problem in which a system must classify crashes into no-injury/property-damage-only (O), minor or moderate injury (BC), and fatal or severe injury (KA), while also controlling outcome leakage, reporting severe under-triage, calibrating confidence, and preserving every raw prediction for audit. Using 4,781 Texas crash records with structured fields and police narratives, we evaluate structured, narrative, fusion, calibrated fusion, BERT-family, and local large-language-model baselines under a stratified 70/15/15 split. In the reported split, leakage-controlled Gemma obtains the highest observed macro-F1 (0.545; 95% bootstrap CI [0.507, 0.583]). The best calibrated fusion model obtains macro-F1 of 0.522 and expected calibration error of 0.033. Selective deferral improves performance among cases retained for automatic classification. At 70% coverage, macro-F1 rises to 0.573 and severity cost falls to 0.577, while deferred cases are treated as candidates for a proposed human-review process and are not further evaluated in the present experiment. The study contributes a reproducible, leakage-controlled, and uncertainty-aware evaluation framework for crash AI systems, emphasizing auditability and selective deferral rather than accuracy alone.
Chinese Translation
公共事故数据库日益支持自动化安全分析,但当事故严重程度预测模型主要作为普通分类器进行评估时,其结果仍难以转化为公共部门的决策工作流程。本研究将痴呆相关事故严重程度建模重新构建为一个决策感知的分诊问题:系统需将事故分类为无伤害/仅财产损失(O)、轻伤或中度伤害(BC)以及死亡或重伤(KA),同时控制结果信息泄露、报告严重漏判率、校准置信度,并保留每一条原始预测以供审计。我们使用包含结构化字段和警方叙述文本的4,781条德克萨斯州事故记录,在按70/15/15分层的划分下,评估了结构化模型、叙述文本模型、融合模型、校准融合模型、BERT系列模型以及本地大语言模型基线。在所报告的划分中,经泄露控制的Gemma获得了最高的观测宏平均F1(0.545;95%自举置信区间[0.507, 0.583])。最佳的校准融合模型宏平均F1为0.522,期望校准误差为0.033。选择性延迟决策提升了保留用于自动分类的案例的性能:在70%覆盖率下,宏平均F1提升至0.573,严重程度代价降至0.577,而被延迟的案例被视为所提出的人工审查流程的候选对象,在本实验中未作进一步评估。本研究为事故AI系统提供了一个可复现、受泄露控制且具不确定性感知的评估框架,强调可审计性与选择性延迟决策,而非单纯追求准确率。
cs.AI / 20 / 2609.22695

A Survey on the Linear Representation Hypothesis

线性表示假说综述
Lee, Sewoong, Canby, Marc E., Cho, Ikhyun, Hockenmaier, Julia
Abstract
The term "linear representation hypothesis" (LRH) has appeared across diverse subfields of artificial intelligence, neuroscience, and cognitive science. But previous works have not consistently treated the LRH as a falsifiable scientific hypothesis; we analyze these inconsistencies and examine their implications for how prior theoretical and methodological results should be interpreted. Based on this analysis, we argue that claims regarding linear representations become well-defined only through careful examination of the model, representation location, feature definition, and evaluation dataset. We therefore propose a more rigorous formalization of the LRH that makes these dependencies explicit and allows the hypothesis to be evaluated as a falsifiable scientific claim. Finally, we identify some non-trivial open problems that warrant further attention from the research community.
Chinese Translation
"线性表示假说"(Linear Representation Hypothesis, LRH)这一术语已出现在人工智能、神经科学和认知科学的多个不同子领域中。然而,以往的研究并未始终将线性表示假说视为一个可证伪的科学假说;我们分析了这些不一致之处,并探讨了它们对如何解释已有的理论与方法学结果的影响。基于这一分析,我们认为,只有通过对模型、表示位置、特征定义以及评估数据集进行仔细考察,关于线性表示的论断才具有明确的定义。因此,我们提出了一个更严格的形式化框架,使这些依赖关系显式化,从而使该假说能够作为一个可证伪的科学论断加以评估。最后,我们指出了若干值得研究界进一步关注的非平凡开放问题。
cs.AI / 21 / 2609.22696

Building Trustworthy Mental Health Benchmarks on Bluesky: A Validation-Aware Weak-Supervision Framework

在Bluesky上构建可信的心理健康基准:一种注重验证的弱监督框架
Chhetri, Gaurab, Dutta, Anandi, Das, Subasish
Abstract
Decentralized social media platforms create new opportunities and challenges for computational mental health research because data access, moderation, labeling, and deployment responsibilities are distributed across multiple technical and governance layers. This paper presents a validation-aware weak-supervision system for constructing and evaluating suicidal ideation (SI) and broader mental health (MH) disclosure benchmarks on Bluesky, a decentralized social media platform built on the AT Protocol. The system integrates public firehose collection, task-specific lexicon filtering, Llama-3-8B-assisted binary annotation, human-adjudicated validation subsets, and transformer-based model benchmarking. Using this pipeline, we construct two task-specific corpora containing 8,346 SI-labeled posts and 9,988 MH-labeled posts. The evaluation shows that model performance depends strongly on both task definition and validation protocol. BERT+LSTM achieves the highest SI stratified cross-validation F1-score, RoBERTa achieves the strongest SI holdout F1-score, and DistilRoBERTa achieves the best MH cross-validation F1-score. Human validation reveals different weak-label failure modes across tasks, with SI labels dominated by false negatives and MH labels dominated by false positives. These findings show that decentralized social media can support reproducible mental health benchmarking, but only when system design, label provenance, validation strategy, and deployment constraints are evaluated together.
Chinese Translation
去中心化社交媒体平台为计算心理健康研究带来了新的机遇与挑战,因为数据访问、内容审核、标注和部署的职责分散在多个技术与治理层面。本文提出了一个注重验证的弱监督系统,用于在基于AT Protocol的去中心化社交媒体平台Bluesky上构建和评估自杀意念(SI)及更广泛的心理健康(MH)表露基准。该系统集成了公开firehose数据采集、任务特定的词典过滤、Llama-3-8B辅助的二分类标注、人工裁定的验证子集以及基于Transformer的模型基准测试。利用这一流程,我们构建了两个任务特定语料库,分别包含8,346条SI标注帖子和9,988条MH标注帖子。评估结果表明,模型性能高度依赖于任务定义和验证协议。BERT+LSTM在SI任务中取得了最高的分层交叉验证F1分数,RoBERTa在SI任务中取得了最高的留出集F1分数,而DistilRoBERTa在MH任务的交叉验证中取得了最佳F1分数。人工验证揭示了不同任务中弱标签的不同失败模式:SI标签以假阴性为主,而MH标签以假阳性为主。这些发现表明,去中心化社交媒体能够支持可复现的心理健康基准构建,但前提是需要对系统设计、标签来源、验证策略和部署约束进行综合评估。
cs.AI / 22 / 2609.22702

Hapi: A Multivariable Land-Surface Transformer for Medium-Range Hydrological Forecasting at Continental Scale

Hapi:面向大陆尺度中期水文预报的多变量陆面Transformer模型
Zhang, Hong, Hutchison, John K., Kotamarthi, Rao, Feinstein, Jeremy, Guan, Haiwen, Maulik, Romit, Ramalingam, Vijay P., Stock, Jason, Wall, Tom
Abstract
Accurate flood forecasts several days in advance are essential for flood control, water-resource management, and emergency response. Producing them at high resolution over a continental domain calls for local hydrological detail together with spatial context extending from river basins to synoptic weather systems. We developed Hapi, a U-Net Swin Transformer that uses fine three-dimensional patches and hierarchical shifted-window attention to forecast discharge, surface runoff, snow water equivalent, and soil wetness across the contiguous United States. The model produces 24--72-hour forecasts at $0.05^{\circ}$ resolution, with learned Laplacian task weights adjusting each variable's contribution to training. On 2024 test data using reconstructed weather and land-surface inputs from ERA5-Land, Hapi outperformed an operational physics-based model and a state-of-the-art AI model in flood detection. Independent validation against 3,881 U.S. Geological Survey gauges and a Hurricane Helene case study supported its advantage over the physics-based model in reproducing daily discharge. Controlled experiments showed that learned task weighting strengthens rare-flood detection, which is particularly sensitive to changes in precipitation inputs. Hapi produced a four-variable, 72-hour forecast across the contiguous United States with an average inference time of 0.11 seconds on a single A100 GPU.
Chinese Translation
提前数天进行准确的洪水预报对洪水防控、水资源管理和应急响应至关重要。要在大陆尺度上以高分辨率生成此类预报,既需要局地水文细节,也需要从流域到天气尺度系统的空间背景信息。我们开发了Hapi,一个采用精细三维图像块(patch)和层次化移位窗口注意力机制的U-Net Swin Transformer(Swin Transformer),用于预报美国本土的流量、地表径流、雪水当量和土壤湿度。该模型以0.05°分辨率生成24至72小时预报,并通过可学习的拉普拉斯任务权重调整每个变量对训练的贡献。在使用ERA5-Land重建的天气和陆面输入的2024年测试数据上,Hapi在洪水检测方面优于一个业务化物理模型和一个最先进的人工智能模型。针对美国地质调查局(USGS)3,881个站点进行的独立验证以及飓风“海伦妮”(Hurricane Helene)的案例研究,进一步支持了其在再现日流量方面相对于物理模型的优势。受控实验表明,可学习的任务加权能够增强对稀遇洪水的检测能力,而这一能力对降水输入的变化尤为敏感。Hapi可在单块A100 GPU上以平均0.11秒的推理时间,生成覆盖美国本土的四变量72小时预报。
cs.AI / 23 / 2609.22712

Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems

可信赖的智能体AI:自主大语言模型系统的失效模式、缓解策略与生命周期框架
Syed, Fayeq Jeelani, Ahmad, Rehan, Bataineh, Ali Al, Adhikari, Aakriti
Abstract
Agentic AI systems built on large language models can plan over multiple steps, use external tools, retain information in memory, and coordinate with other agents. These capabilities make them more useful than static language models, but they also introduce new security and operational risks. Untrusted content from websites, emails, documents, and databases can enter the same context as system instructions; persistent memory can carry compromised information across sessions; and access to external tools can turn an incorrect model response into a consequential real-world action. This article reviews the trustworthiness of agentic AI across five interconnected dimensions: safety and robustness, alignment and human oversight, transparency and auditability, privacy and data governance, and regulatory compliance. It organizes key failure modes, including indirect prompt injection, backdoor triggers, goal misgeneralization, memory contamination, and cross-session data leakage, into a unified taxonomy. It also examines major mitigation approaches, such as instruction hierarchies, context isolation, spotlighting, process-based supervision, constrained tool use, and privacy-preserving memory, while distinguishing techniques supported by empirical evidence from those that remain largely conceptual. Building on this analysis, we introduce the Trustworthy Agent Development Lifecycle (TADL), a six-phase framework covering specification, design, training, evaluation, deployment, and monitoring. For each phase, TADL identifies relevant trust activities, expected evidence, and risk-based decision gates. Although TADL has not yet been empirically validated, it provides a structured foundation for developing and evaluating more secure and accountable agentic systems. The article concludes by identifying gaps in current benchmarks and outlining priorities for future research.
Chinese Translation
基于大语言模型构建的智能体AI系统能够进行多步骤规划、使用外部工具、在内存中保留信息,并与其他智能体协同工作。这些能力使其比静态语言模型更为实用,但同时也带来了新的安全与运行风险。来自网站、电子邮件、文档和数据库的不可信内容可能与系统指令进入同一上下文;持久化记忆可能将受污染的信息跨会话传递;而对外部工具的访问可能将错误的模型响应转化为产生实际后果的现实世界行为。本文从五个相互关联的维度审视智能体AI的可信性:安全性与鲁棒性、对齐与人类监督、透明性与可审计性、隐私与数据治理,以及合规性。文章将关键失效模式,包括间接提示注入、后门触发、目标错误泛化、记忆污染和跨会话数据泄露,整合为一个统一的分类体系。同时,本文考察了主要的缓解方法,如指令层级、上下文隔离、聚焦标注(spotlighting)、基于过程的监督、受约束的工具使用以及隐私保护记忆,并区分了有实证依据支持的技术与仍主要停留在概念层面的技术。基于上述分析,我们提出了可信赖智能体开发生命周期(Trustworthy Agent Development Lifecycle, TADL),这是一个涵盖规格定义、设计、训练、评估、部署和监控六个阶段的框架。针对每个阶段,TADL 明确了相关的可信性活动、预期证据以及基于风险的决策关卡。尽管 TADL 尚未经过实证验证,但它为开发和评估更安全、更负责任的智能体系统提供了结构化的基础。文章最后指出了当前基准测试中的不足,并展望了未来研究的优先方向。
cs.AI / 24 / 2609.22746

ProcessLight: Process Supervision for Large Language Model Based Traffic Signal Control

ProcessLight:基于大语言模型的交通信号控制的过程监督方法
Zhao, Huaitao, Zhou, Tianlong, Wang, Weijie, Shi, Jiasheng, Rao, Weixiong
Abstract
Large Language Models (LLMs) have recently been introduced into traffic signal control (TSC) as decision agents due to their strengths in human-readable reasoning generation. Yet, existing LLM TSC methods optimize only from final outcomes and fail to distinguish valid from flawed reasoning steps, causing useful or misleading steps to be jointly updated and thus impairing the model's learning of effective reasoning. To bridge this gap, we propose an LLM-based framework ProcessLight to decompose signal decisions into verifiable semantic steps. Building on ProcessLight, we further develop Step-wise Traffic Process Policy Optimization (STeP-PO), a novel reinforcement learning framework that optimizes structured reasoning processes through step-level credit assignment. Specifically, STeP-PO uses step quality scores to evaluate local reasoning quality and step importance to measure each step's influence on the final action, and then assigns step-level advantages over a semantic step tree structure. The resulting step-level advantages are propagated to reasoning tokens, enabling fine-grained policy optimization beyond outcome-only rewards. Extensive experiments over multiple real-world datasets demonstrate the superiority of our methods. Our code is available at https://github.com/wenzhaoabc/processlight.
Chinese Translation
大语言模型(LLMs)因其在生成人类可读推理方面的优势,近来被引入交通信号控制(TSC)领域作为决策智能体。然而,现有的LLM交通信号控制方法仅从最终结果进行优化,无法区分有效与有缺陷的推理步骤,导致有用和误导性的步骤被同时更新,从而损害了模型对有效推理的学习。为弥补这一不足,我们提出了一个基于LLM的框架ProcessLight,将信号决策分解为可验证的语义步骤。在ProcessLight的基础上,我们进一步提出了逐步交通过程策略优化(Step-wise Traffic Process Policy Optimization, STeP-PO),这是一种通过步骤级信用分配来优化结构化推理过程的新型强化学习框架。具体而言,STeP-PO使用步骤质量分数来评估局部推理质量,使用步骤重要性来衡量每个步骤对最终动作的影响,然后在语义步骤树结构上分配步骤级优势。由此得到的步骤级优势被传播到推理token上,从而实现超越仅基于结果奖励的细粒度策略优化。在多个真实世界数据集上的大量实验证明了我们方法的优越性。我们的代码已发布于 https://github.com/wenzhaoabc/processlight。
cs.AI / 25 / 2609.22760

CTSpinoPelvic1K: spine, pelvis, ribs and femora in one coordinate frame, annotated for lumbosacral transitional anatomy

CTSpinoPelvic1K:同一坐标系下的脊柱、骨盆、肋骨和股骨数据集,并对腰骶移行解剖进行了标注
Schwing, Gregory, Schehr, Ashley, Tekumulla, Annika, Khoushi, Margret, Christian, Ryan, Hubers, Dane, Mahjoub, Faris, Saad, Hassan, Sooch, Mia, Siddapureddy, Sathyagopal, McLellan, Michael, Kim, Jerick, Ismoilov, Miraziz, Alnabahneh, Nizar
Abstract
Purpose: A vertebra at the lumbosacral junction is named by counting caudally from C2 on whole-spine imaging, but a lumbar case is planned on lumbar-only imaging (T12 to S1), without C2. Abdominopelvic CT holds that span plus the lowest ribs and pelvis. Where a lumbosacral transitional vertebra (LSTV) alters the count, the local anatomy is ambiguous: four rib-free vertebrae may be an L1 with a lumbar rib or an L5 assimilated to the sacrum, and six may be a sixth lumbar vertebra, a T12 with aplastic ribs, or a lumbarized S1. CTSpinoPelvic1K asks whether local morphology resolves it without the count. CTSpine1K's vertebrae and CTPelvic1K's pelvis covered these patients but were never joined; this release joins them on one series and adds the bones neither had. It provides 802 CT records with per-level ribs and femora, levels anchored on the lowest rib-bearing vertebra and S1, and classes for L6, T13, a separate S1 and lumbar ribs, so anomalies are recorded as such. Acquisition and Validation Methods: Records pair CTSpine1K and CTPelvic1K labels on each patient's bone-richest series under a VerSe-native scheme. Validation covered geometric invariants (802/802 pass), rib-vertebra incidence across 5,749 ribs (0.035% offset), and spinopelvic measures matching published values. Data Format and Usage Notes: NIfTI image/label pairs with patient-grouped LSTV-stratified five-fold splits and a loader; archived at https://doi.org/10.5281/zenodo.22139642. Potential Applications: Classifying a vertebra from local features; updating cadaveric morphometry; spinopelvic assessment; opportunistic screening; and, absent a public preoperative lumbar cohort, surgical planning research (377 records prone). Limitations: thoracic ground truth is field-of-view limited; postural angles supine; no held-out test set; ribs are triaged-review pseudolabels; Castellvi grades two-reader consensus on 33 records.
Chinese Translation
目的:在全脊柱影像中,腰骶交界处的椎体通过从C2向下计数来命名,但腰椎病例的手术规划仅基于腰椎影像(T12至S1),其中不含C2。腹部盆腔CT涵盖了该范围以及最下方的肋骨和骨盆。当腰骶移行椎(LSTV)改变了椎体计数时,局部解剖便会变得含糊不清:四个无肋椎体可能是伴有腰肋的L1,也可能是被骶骨同化的L5;六个椎体则可能是第六腰椎、肋骨发育不良的T12,或是骶化(腰椎化)的S1。CTSpinoPelvic1K旨在探究仅凭局部形态学特征能否在不依赖计数的情况下解决这一歧义。CTSpine1K的椎体数据与CTPelvic1K的骨盆数据覆盖了这些患者,但二者从未被合并;本数据集将其统一到同一序列上,并补充了两者均缺失的骨骼。该数据集提供802例CT记录,包含逐节段的肋骨和股骨标注,以最下方带肋椎体和S1为锚定的节段,以及L6、T13、独立S1和腰肋等类别,从而将解剖异常如实记录。采集与验证方法:采用VerSe原生标注方案,将每位患者骨骼覆盖最丰富的序列上的CTSpine1K与CTPelvic1K标签进行配对。验证包括几何不变性检验(802/802通过)、覆盖5,749根肋骨的肋-椎对应关系核验(偏差0.035%),以及与已发表数值相符的脊柱-骨盆参数测量。数据格式与使用说明:NIfTI图像/标签配对数据,提供按患者分组、按LSTV分层的五折划分及数据加载器;数据归档于 https://doi.org/10.5281/zenodo.22139642。潜在应用:基于局部特征进行椎体分类;更新尸体形态测量学数据;脊柱-骨盆评估;机会性筛查;以及在缺乏公开术前腰椎队列的情况下用于手术规划研究(377例记录为俯卧位)。局限性:胸椎地面真值受成像视野限制;体位角度为仰卧位测量;无独立测试集;肋骨标注为经分诊复核的伪标签;Castellvi分级为两名读片者在33例记录上的共识结果。
cs.AI / 26 / 2609.22775

DVA-Neurons: Design and Verification of Adaptive LIF Neurons: From Single-Neuron Dynamics to Multi-Neuron Spiking Networks

DVA-神经元:自适应LIF神经元的设计与验证——从单神经元动力学到多神经元脉冲网络
Pham, Thanh, Islam, Riadul
Abstract
Spiking Neural Networks (SNNs) offer a promising path toward ultra-low-power artificial intelligence inference by emulating the event-driven computation of biological neurons. However, two challenges limit their practical deployment. First, fixed-parameter Leaky Integrate-and-Fire (LIF) neurons lack the adaptation mechanisms observed in biology, where neurons modulate their excitability based on firing history. Second, scaling from single neurons to multi-neuron networks introduces challenges in synaptic weight distribution and inter-neuron spike routing that are absent in isolated designs. This paper addresses both issues through the extension, verification, and physical implementation of adaptive LIF neurons at three architectural scales. This work contributes: a 2nd-order neuron with two-stage synaptic filtering for richer temporal dynamics; a fully-connected 6-neuron spiking network with configurable weights (100 to 5) demonstrating weight-based inter-neuron communication; and a direct verification methodology enabling per-cycle observation of all internal states. All designs were synthesized targeting Selected Area Electron Diffraction (SAED) 14 nm Complementary Metal-Oxide-Semiconductor (CMOS) technology at 1 GHz and verified with Cocotb-based Python testbenches under pulsed current stimuli (amplitude 80, ISI=3). The results show that adaptation effectively modulates firing: 31% suppression in the 2nd-order neuron (25 vs.\ 36 spikes) and 31% reduction in postsynaptic firing in the network (18 vs.\ 26 spikes). Physically, the 2nd-order neuron costs 1.77x more area and 1.52x more power than the 1st-order baseline, while the 6-neuron network demonstrates near-linear scaling (5.7x area, 5.3x power). Seven verification bugs spanning testbench connectivity, fixed-point overflow, and Verilog expression-width semantics are documented.
Chinese Translation
脉冲神经网络(SNN)通过模拟生物神经元的事件驱动计算,为实现超低功耗人工智能推理提供了一条有前景的路径。然而,两个挑战限制了其实际部署。其一,固定参数的泄漏积分发放(Leaky Integrate-and-Fire, LIF)神经元缺乏生物中观察到的自适应机制,即神经元会根据发放历史调节其兴奋性。其二,从单神经元扩展到多神经元网络会带来突触权重分布和神经元间脉冲路由等在孤立设计中不存在的问题。本文通过在三个架构层级上对自适应LIF神经元进行扩展、验证和物理实现,同时解决了这两个问题。本工作的贡献包括:一个具有两级突触滤波的二阶神经元,以实现更丰富的时间动力学;一个具有可配置权重(100至5)的全连接六神经元脉冲网络,展示了基于权重的神经元间通信;以及一种支持对所有内部状态进行逐周期观测的直接验证方法学。所有设计均基于Selected Area Electron Diffraction(SAED)14纳米互补金属氧化物半导体(CMOS)工艺在1 GHz下综合,并通过基于Cocotb的Python测试平台在脉冲电流激励(幅值80,ISI=3)下进行验证。结果表明,自适应机制能有效调节发放:二阶神经元的发放被抑制31%(25次对比36次脉冲),网络中突触后发放减少31%(18次对比26次脉冲)。在物理实现上,二阶神经元的面积和功耗分别为一阶基线的1.77倍和1.52倍,而六神经元网络表现出近线性扩展(面积5.7倍,功耗5.3倍)。此外,论文记录了七个验证缺陷,涉及测试平台连接、定点溢出以及Verilog表达式位宽语义等问题。
cs.AI / 27 / 2609.22790

From Research Frontier to Laboratory Bench: Design of a Four-Tier Experimental Teaching System for Multimodal Medical Image Intelligent Diagnosis

从研究前沿到实验台:面向多模态医学影像智能诊断的四层次实验教学系统设计
Shan, Dongjing, Luo, Yamei, Li, Jin, Luo, Yong
Abstract
Undergraduate programmes in intelligent medical engineering are expanding, yet laboratory curricula lag behind the multimodal, long-tailed, and distributionally shifting realities of clinical AI. This design paper presents an advanced experimental teaching system that translates an ongoing multimodal deep learning research project on endometrial carcinoma into a structured undergraduate lab sequence. We identify three educational gaps (modality, authenticity, and deployment) and derive four pedagogical principles from constructive alignment, experiential learning, the research teaching nexus, and the CDIO framework. The curriculum comprises four progressive tiers plus an engineering layer, with 32 laboratory units over 64 contact hours, delivered via a custom virtual clinical workstation using de-identified multi-institutional data. Each tier maps to a specific technical bottleneck, prerequisite coursework, and criterion-referenced deliverables. Data governance, safety, and assessment protocols are specified. Learning outcome data will be collected across two implementation cycles.
Chinese Translation
智能医学工程本科专业规模不断扩大,但实验课程教学仍落后于临床人工智能中多模态、长尾分布及分布漂移的现实情况。本设计论文提出了一套先进的实验教学系统,将持续进行的多模态深度学习研究项目——子宫内膜癌智能诊断——转化为结构化的本科实验课程序列。我们识别出三个教学缺口(模态、真实性和部署),并基于建构性对齐、体验式学习、科研教学结合以及CDIO工程教育模式,提炼出四条教学原则。该课程体系由四个递进层次和一个工程实践层组成,共包含32个实验单元、64个学时,通过自建的虚拟临床工作站并使用去标识化的多机构数据进行教学。每个层次对应特定的技术瓶颈、先修课程要求以及标准参照的可交付成果。论文同时规定了数据治理、安全与评估方案,并将在两个实施周期内收集学习成效数据。
cs.AI / 28 / 2609.22878

ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures

ISA-Bench:一个面向跨指令集架构计算推理的基准测试
Pola, Aditya, Majumdar, Arkaprava, Balasubramanian, Vineeth N.
Abstract
Large language model code generation benchmarks primarily evaluate well-resourced languages like Python and Java, where models benefit from abundant training data. They provide limited evidence about reasoning in unfamiliar computational models: deriving arithmetic from a single subtract instruction, coordinating parallel programs across communicating nodes, or wiring logic gates into circuits. We present ISA-Bench, a benchmark of programming games with constrained instruction sets. For each game we provide a full execution stack (parser, VM, and verifier), enabling automated evaluation with structured feedback for iterative refinement. Reasoning models achieve higher average solve rates than code-specialized and general-purpose models, but unfamiliar syntax remains a major source of failure. Models solve more tasks with iterative feedback, though the gains vary substantially across architectures. We introduce a reasoning--execution gap (REG) analysis that reveals a recurring disconnect between identifying a plausible computational strategy and expressing it as a correct program in the target ISA. Code is open-sourced.
Chinese Translation
大语言模型代码生成基准主要评估资源丰富的语言(如 Python 和 Java),模型可从充足的训练数据中获益。这些基准对于推理陌生计算模型提供了有限的证据:例如从单一的减法指令推导算术运算、协调跨通信节点的并行程序,或将逻辑门连接成电路。我们提出了 ISA-Bench,一个基于受限指令集编程游戏的基准。对于每个游戏,我们提供完整的执行栈(解析器、虚拟机和验证器),支持带结构化反馈的自动化评估,以便进行迭代改进。推理模型比代码专用模型和通用模型取得了更高的平均解题率,但陌生的语法仍是失败的主要来源。借助迭代反馈,模型能解决更多任务,但增益在不同架构间差异显著。我们引入了推理—执行鸿沟(Reasoning-Execution Gap, REG)分析,揭示了一种反复出现的脱节现象:模型虽能识别出合理的计算策略,却难以将其表达为目标指令集架构(ISA)中的正确程序。代码已开源。
cs.AI / 29 / 2609.22910

When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used

视觉语言模型应何时查看?仅为必要且有效的视觉调用付费
Peng, Kunyu, Liu, Junming, He, Ruiqi, Wang, Qingzhuo, Qi, Jianzhong, Liu, Xianhui
Abstract
Vision-language agents that crop and zoom are trained with rewards that credit a successful tool call, yet a successful call does not show that the model needed to look or used the pixels it received. On our cold-start checkpoint only 10% to 12% of visual calls were both needed and used, and released agents make spurious calls 36% to 87% of the time on individual benchmarks. Outcome rewards, judge rewards, and branch probes each observe one side of this failure, and about two thirds of what an outcome reward pays goes to calls that were neither needed nor used. CounterCredit asks both questions of every image-returning call at its realized pre-call state, using the policy's own gold-answer score. A decision value compares the realized visual branch with answering immediately; an evidence value compares the returned crop with random same-size patches substituted into the same call. A call verified on both earns cashback and every other executed call pays rent; the price is bounded so that every correct trajectory outranks every wrong one, and a dual-channel GRPO advantage keeps the price in its own units. From the same cold start, prompt pool, and budget, CounterCredit reaches 89.5% on V*, 80.2% on HR-Bench-4K, and 76.4% on HR-Bench-8K, 6.3 to 9.4 points above outcome-only GRPO at 1.78 against 1.84 calls per question, and lowers the spurious-call rate to 31% to 36%, the lowest among the agents evaluated. The same recipe lifts a Qwen3-VL-8B base from 75.4 to 80.8 on average.
Chinese Translation
具备裁剪与缩放能力的视觉语言智能体通常以成功的工具调用获得奖励的方式进行训练,然而一次成功的调用并不能证明模型确实需要查看图像,也不能证明其真正使用了所接收到的像素信息。在我们的冷启动检查点上,仅有10%至12%的视觉调用同时满足“必要”与“有效使用”这两个条件,而已发布的智能体在各个基准上做出虚假调用的比例达36%至87%。结果奖励、评判奖励和分支探测各自只能观察到这一失败的一面,且结果奖励所支付的总金额中约三分之二流向了既不必要也未被使用的调用。CounterCredit 在每次返回图像的调用发生前的真实状态下,利用策略自身的标准答案得分对上述两个问题同时进行检验:决策价值(decision value)将真实的视觉分支与立即作答进行比较;证据价值(evidence value)将返回的裁剪图像与替换到同一调用中的同尺寸随机图像块进行比较。同时通过两项验证的调用可获得返现,而其他所有已执行的调用则需支付租金;该价格设有上界,以保证每条正确轨迹的排序都高于每条错误轨迹,并且通过双通道 GRPO 优势使该价格保持在其自身单位内。在相同的冷启动检查点、提示池和预算下,CounterCredit 在 V* 上达到89.5%,在 HR-Bench-4K 上达到80.2%,在 HR-Bench-8K 上达到76.4%,比仅使用结果奖励的 GRPO 高出6.3至9.4分,同时每题平均调用次数为1.78次(对比1.84次),并将虚假调用率降至31%至36%,为所评估的智能体中最低。同样的方法还将 Qwen3-VL-8B 基础模型的平均成绩从75.4提升至80.8。
cs.AI / 30 / 2609.22939

Beyond Linear Context: Graph-Guided Evidence Navigation for Long-Novel Reasoning with a Local 9B Language Model

超越线性语境:基于图引导证据导航的长篇小说推理方法(使用本地9B语言模型)
Fu, Wenji
Abstract
Long-context models read a novel the way a person reads a printout: one token after another, in narrative order, with the whole history competing for a fixed budget of attention. A detective does not work that way. They sort what happened when, and they keep a map of who relates to whom, so a clue from chapter one can meet a question asked at the end of the book. We test whether a frozen knowledge graph can give a small local model that same freedom. Thirty detective novels and 234 multiple-choice questions are answered by one fixed qwen3.5:9b reader under nine conditions: five graph routes, a recent-window baseline, whole-book compression, ordinary vector retrieval, and a question-only control. The strongest graph route reaches 53.85% (126/234) against 46.15% for the recent window, 51.28% for compression, 51.71% for vector retrieval and 40.17% for question-only. On the subset that no model can answer without the book, the graph route reaches 42.86%. None of the fifteen graph-baseline contrasts survives Holm correction, so we present the result as exploratory evidence about a design. Two structural findings survive scrutiny better than the headline number: annotated evidence concentrates in the topological core of these graphs (2.35x enrichment, pooled), and the two graph-building pipelines differ so much in annotation coverage (16% versus 73% of clue paragraphs) that pooled accuracy alone would hide which bottleneck is being measured.
Chinese Translation
长上下文模型阅读小说的方式就像人阅读打印稿一样:按照叙事顺序逐个token进行,整个历史共同竞争固定的注意力预算。而侦探的工作方式并非如此。他们会梳理事件发生的时间顺序,并维护一张人物关系图,从而使第一章的线索能够与书末提出的问题相互关联。我们测试了冻结的知识图谱能否赋予小型本地模型同样的自由度。在九种实验条件下,使用同一个固定的qwen3.5:9b阅读器回答了30部侦探小说的234道选择题:包括五种图谱路线、近期窗口基线、全书压缩、普通向量检索以及仅问题对照。最强的图谱路线达到53.85%(126/234),相比之下,近期窗口为46.15%,全书压缩为51.28%,向量检索为51.71%,仅问题为40.17%。在无书可查时任何模型都无法回答的子集上,图谱路线达到42.86%。十五组图谱与基线的对比均未通过Holm校正,因此我们将该结果作为关于一种设计的探索性证据呈现。有两项结构性发现比总体数字更能经受检验:标注证据集中在这些图谱的拓扑核心区域(合并后富集度为2.35倍),且两条图谱构建流程在标注覆盖率上差异巨大(线索段落覆盖率分别为16%和73%),以至于仅看合并准确率会掩盖所测量的瓶颈究竟是哪一个。
cs.AI / 31 / 2609.22951

AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows

AgentRouter:面向成本最优多步骤智能体工作流的异构模型路由
Paul, Rudrendu Kumar, Nandy, Sourav
Abstract
Enterprise agentic systems that route every trajectory step to a frontier model waste 60-80% of their inference budget on subtasks that smaller models handle equally well. Existing routing solutions optimize single-turn query assignment but ignore a property unique to agentic workflows: subtask complexity varies widely within a single trajectory. A planning step may require frontier-class reasoning while a subsequent formatting step needs only a 7B model. We formalize step-level model routing as a sequential assignment problem over agent trajectories and propose AgentRouter, a lightweight classifier (12M parameters, <5ms overhead per step on an A100 GPU) that maps each trajectory step to one of four model tiers using five features extractable at routing time. Trained on 50,000 annotated agent trajectory steps spanning planning, coding, research, and data analysis tasks, AgentRouter achieves 72% cost reduction relative to frontier-only baselines, retaining 97.3% of frontier-only quality (less than 3% degradation in end-to-end task completion); per-step routing accuracy reaches 91% on minimal-complexity steps and 85% on efficient-tier steps, with 76-82% on the harder mid-range and frontier tiers. On the same benchmarks, RouteLLM and FrugalGPT (applied per-step) achieve only 31% and 44% cost reduction respectively, because their single-turn training signal misses trajectory-level quality dependencies.
Chinese Translation
企业级智能体系统若将轨迹中的每一步都路由至前沿大模型,会在较小模型同样能胜任的子任务上浪费60%-80%的推理预算。现有路由方案仅优化单轮查询分配,却忽略了智能体工作流独有的特性:子任务复杂度在同一条轨迹内部差异巨大。一个规划步骤可能需要前沿级别的推理能力,而随后的格式化步骤仅需一个7B模型。我们将步骤级模型路由形式化为智能体轨迹上的序贯分配问题,并提出AgentRouter——一个轻量级分类器(1200万参数,在A100 GPU上每步开销低于5毫秒),它利用五个可在路由时提取的特征,将轨迹中的每一步映射到四个模型层级之一。AgentRouter在涵盖规划、编程、研究和数据分析任务的50,000条标注智能体轨迹步骤上训练,相对于仅使用前沿模型的基线实现了72%的成本降低,同时保留了前沿模型97.3%的质量(端到端任务完成率下降不足3%);其在最低复杂度步骤上的每步路由准确率达到91%,高效层级步骤上达到85%,在难度较高的中档层级和前沿层级上分别为76%-82%。在相同基准测试中,RouteLLM和FrugalGPT(按步骤应用)分别仅实现31%和44%的成本降低,原因在于其单轮训练信号无法捕捉轨迹级别的质量依赖关系。
cs.AI / 32 / 2609.22956

A Compact Stance-Indexed Anterior-Posterior COP Representation for Parkinson's Disease Classification from Plantar VGRF

一种紧凑的姿势期索引式前后向压力中心表示方法:基于足底垂直地面反作用力的帕金森病分类
Sifat, Md., Akter, Sania, Islam, Akif, Hamid, Md. Ekramul
Abstract
Parkinson's disease alters gait and bilateral coordination, but machine-learning performance also depends on how continuous gait signals are represented. This study investigates whether preserving anterior-posterior center-of-pressure (AP-COP) information at fixed locations across normalized stance provides a compact and informative representation of plantar-force gait signals. Bilateral vertical ground reaction force recordings from 165 participants in the Gait in Parkinson's Disease Database were evaluated using repeated fully nested participant-level cross-validation. We propose AP-COP10, comprising AP-COP position and bilateral asymmetry across five stance windows. AP-COP10 achieved an AUC of 0.894 and outperformed three harmonized literature-derived COP representations under the same evaluation pipeline. The complementary 25 non-AP-COP descriptors alone achieved an AUC of 0.856, while the complete 35-feature representation achieved 0.908. Removing AP-COP10 from the complete representation produced a statistically supported loss in discrimination, whereas adding the complementary descriptors to AP-COP10 yielded only a small, unsupported improvement. Feature competition indicated that the most informative stance-indexed descriptors were concentrated in early and early-mid stance, while source-study holdout and sensor-perturbation analyses supported the robustness of the representation. These findings indicate that stance-indexed AP-COP retains discriminative information that is not readily recovered by broader engineered gait descriptors, supporting compact and interpretable representations for machine-learning analysis of pathological gait.
Chinese Translation
帕金森病会改变步态和双侧协调性,但机器学习的性能还取决于连续步态信号的表示方式。本研究探讨在归一化站立相的固定位置上保留前后向压力中心(AP-COP)信息,能否为足底力步态信号提供一种紧凑且信息丰富的表示。我们利用重复的完全嵌套参与者级交叉验证,对来自“帕金森病步态数据库”中165名参与者的双侧垂直地面反作用力记录进行了评估。我们提出了AP-COP10方法,该方法包含五个站立相窗口内的AP-COP位置和双侧不对称性。在相同的评估流程下,AP-COP10达到了0.894的AUC,并优于三种经统一处理的文献来源COP表示。仅使用其余25个非AP-COP特征时AUC为0.856,而完整的35维特征表示达到0.908。从完整表示中移除AP-COP10会导致具有统计学支持的判别能力下降,而在AP-COP10基础上加入其余特征仅带来微小且无统计学支持的提升。特征竞争分析表明,信息量最大的姿势期索引特征集中于站立相早期和早中期,来源研究的留出集验证和传感器扰动分析也支持了该表示的鲁棒性。这些发现表明,姿势期索引的AP-COP保留了更广泛的工程化步态特征难以恢复的判别信息,为病理性步态的机器学习分析提供了紧凑且可解释的表示方法。
cs.AI / 33 / 2609.22959

R-GEAN: Regimen-Guided Edit Action Network for Within-Admission Medication Change Prediction

R-GEAN:面向住院期间用药变化预测的方案引导编辑动作网络
Mahat, Regan, Kim, Mansu
Abstract
The medications prescribed to a patient often change during a hospital admission as clinicians start, stop, or continue therapies. We study whether models can predict which medication classes are added or removed between 24 hours after admission and discharge. Metrics that compare the complete discharge regimen can reward models for copying medications that remain unchanged, even when they identify no actual changes. We therefore introduce a leakage-controlled benchmark that predicts net ATC3 additions and removals using only prior completed admissions and information available within the first 24 hours of the current admission. Addition candidates are classes not active at 24 hours, whereas removal candidates are classes active at that time. We also introduce R-GEAN, an asymmetric candidate-scoring network with independent addition and removal predictors. Across 240,480 admissions from 82,286 patients, R-GEAN achieves the highest predefined summary of addition, removal, changed-regimen, and action-pattern performance, termed the edit composite (0.464), compared with 0.435 for the strongest primary comparator. Reimplemented RETAIN, GAMENet, and MICRON baselines obtain 0.428, 0.420, and 0.288, respectively. R-GEAN's advantage is concentrated in correctly identifying medication classes no longer active at discharge, while rare additions and admissions with multiple medication changes remain difficult. Rankings based on micro-F1 over the reconstructed discharge regimen and the edit composite correlate weakly across the evaluated models (Spearman r = 0.20). The continuation baseline achieves the highest complete-regimen score despite predicting no additions or removals. These results show that complete-regimen and edit-level evaluation measure different aspects of medication prediction. The benchmark evaluates observed prescribing changes, not treatment appropriateness
Chinese Translation
患者在住院期间,医生会开始、停止或继续某些治疗,因此处方用药常常发生变化。我们研究模型能否预测从入院24小时后到出院期间哪些药物类别被新增或停用。比较完整出院方案的指标可能会奖励那些照抄未变化用药的模型,即使它们并未识别出任何实际变化。因此,我们提出了一个防止信息泄漏的基准任务,仅利用既往已完成的住院记录以及当前住院前24小时内可获得的信息,来预测ATC3层面的净新增与停用。新增候选为在24小时时点尚未使用的药物类别,而停用候选为该时点正在使用的药物类别。我们还提出了R-GEAN,一种具有独立新增预测器和停用预测器的非对称候选评分网络。在来自82,286名患者的240,480次住院记录上,R-GEAN在预先定义的新增、停用、变化方案及动作模式性能综合指标(称为编辑综合指标,edit composite)上取得最高分0.464,而最强的主要对比方法为0.435。重新实现的RETAIN、GAMENet和MICRON基线分别获得0.428、0.420和0.288。R-GEAN的优势集中在正确识别出院时已不再使用的药物类别上,而罕见的新增以及多次用药变化的住院案例仍然具有挑战性。在重构的出院方案上计算的micro-F1排名与编辑综合指标的排名在所评估的模型之间相关性较弱(Spearman r = 0.20)。仅预测无新增和停用的延续基线却在完整方案评分上取得最高分。这些结果表明,完整方案评估与编辑层面的评估衡量的是用药预测的不同方面。该基准评估的是观察到的处方变化,而非治疗 appropriateness(治疗是否恰当)。
cs.AI / 34 / 2609.22987

OptiSkill: A Hierarchical and Evolving SkillBank for LLM-Based Optimization Modeling

OptiSkill:面向基于大语言模型优化建模的分层演化技能库
Zhao, Ruiqing, Liu, Rui, Zuo, Yuan, Zhang, Huarong, Han, Xiao, Wu, Junjie
Abstract
Automated operations research (OR) modeling requires LLMs to translate natural-language decision problems into correct mathematical programs. Existing methods can improve individual formulations, but they often solve problems in isolation, retaining little reusable experience and repeating similar formulation errors. Prior memory-based approaches store examples, thoughts, or insights as references, while OR modeling requires reusable formulation skills that transfer across problem narratives and guide concrete modeling decisions. We propose OptiSkill, a skill-augmented framework that builds a hierarchical and evolving SkillBank for LLM-based OR modeling. SkillBank stores solver-verified experience as reusable skills, with Global Strategies for problem-level formulation skeletons and Step Experiences for local error-prevention rules. It is further refined through stable batch-level test-time evolution, where candidate skills are incorporated only after validation. Experiments on eight OR modeling benchmarks show that OptiSkill improves formulation accuracy across LLM backbones, outperforms strong agentic baselines, and gains further by expanding SkillBank coverage and reliability. Code and data are available at https://github.com/rachhhhing/OptiSkill
Chinese Translation
自动化运筹学(OR)建模要求大语言模型(LLM)将自然语言描述的决策问题转化为正确的数学规划模型。现有方法可以改进单个问题的建模,但它们往往孤立地解决每个问题,几乎不保留可复用的经验,导致重复出现相似的建模错误。以往基于记忆的方法将示例、思路或洞见作为参考进行存储,而运筹学建模需要的是能够跨问题情境迁移并指导具体建模决策的可复用建模技能。我们提出OptiSkill,一个技能增强框架,为基于LLM的运筹学建模构建分层且不断演化的技能库(SkillBank)。SkillBank将经求解器验证的经验存储为可复用技能,其中全局策略(Global Strategies)用于问题级的建模骨架,步骤经验(Step Experiences)用于局部错误防范规则。该技能库进一步通过稳定的批次级测试时演化进行精炼,候选技能仅在验证通过后才会被纳入。在八个运筹学建模基准上的实验表明,OptiSkill在不同LLM骨干模型上均提升了建模准确率,优于强大的智能体基线方法,且随着SkillBank覆盖范围和可靠性的扩展,性能进一步提升。代码和数据可在 https://github.com/rachhhhing/OptiSkill 获取。
cs.AI / 35 / 2609.23023

PINNForge: Execution-Grounded Evolutionary Design of Physics-Informed Neural Networks for PDE Solving via Large Language Models

PINNForge:基于执行反馈与大规模语言模型的物理信息神经网络演化式设计求解偏微分方程
Yu, Mingyang, Yang, Xu, Zhang, Jun, Wang, Xiaolong, Xu, Jing, Li, Keqian
Abstract
Physics-informed neural networks (PINNs) require coordinated choices over network representation, sampling, loss construction, and optimization, while effective configurations often vary substantially across partial differential equations (PDEs). Existing automated PINN design methods can search candidate configurations, but information revealed during actual training is still used mainly for evaluation rather than to improve subsequent design, leading to repeated trial-and-error and inefficient use of training budget. We propose PINNsForge, an LLM-driven evolutionary framework for execution-feedback-based automated PINN design. PINNsForge generates diverse candidate configurations from PDE-related prior knowledge, evaluates them through actual training, and feeds high-performing designs together with accumulated execution evidence back to the LLM. Guided by observed optimization behavior, the LLM then refines, recombines, and explores coupled PINN design components, forming a continual cycle of generation, execution, feedback, and evolution. Unlike one-shot search or evaluation-only feedback, PINNsForge progressively converts training experience into improved design decisions for the target PDE. Across 25 PDE benchmarks, PINNsForge achieves the lowest mean MSE on 24 tasks compared with RoPINN, PINNsFormer, and PINNsAgent. Ablation studies further confirm the importance of the PDE knowledge base, execution feedback, and evolutionary search: removing these components increases the mean MSE to 3.74$\times$, 12.10$\times$, and 10.10$\times$ that of the full PINNsForge, respectively.
Chinese Translation
物理信息神经网络(PINN)需要在网络表示、采样、损失构建和优化等方面进行协同设计,而有效的配置往往因偏微分方程(PDE)的不同而有显著差异。现有的自动化PINN设计方法虽能搜索候选配置,但实际训练过程中揭示的信息仍主要用于评估,而非用于改进后续设计,导致反复试错和训练预算的低效利用。我们提出PINNsForge,一个由大语言模型(LLM)驱动、基于执行反馈的自动化PINN设计演化框架。PINNsForge基于PDE相关的先验知识生成多样化的候选配置,通过实际训练对其进行评估,并将高性能设计与积累的执行证据反馈给LLM。在观测到的优化行为引导下,LLM进而对耦合的PINN设计组件进行改进、重组与探索,形成生成、执行、反馈与演化的持续循环。与一次性搜索或仅用于评估的反馈不同,PINNsForge能逐步将训练经验转化为针对目标PDE的更优设计决策。在25个PDE基准测试中,与RoPINN、PINNsFormer和PINNsAgent相比,PINNsForge在其中24个任务上取得了最低的平均MSE。消融研究进一步证实了PDE知识库、执行反馈和演化搜索的重要性:移除这些组件后,平均MSE分别增至完整PINNsForge的3.74倍、12.10倍和10.10倍。
cs.AI / 36 / 2609.23038

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

Spatial-Interactor:通过与可观测物理世界的交互学习空间推理
Yao, Kaixiang, Wang, Xu, Pan, Miao, Xiyue, Hu, Wang, Weishi, Dahlmeier, Daniel, Chen, Jintao, Shen, Yongliang, Zhang, Xuhong, Zhang, Wenqi
Abstract
Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial state, yet existing VLMs remain limited in both capabilities. Current spatial training primarily focuses on static questions about object attributes and spatial relations, providing limited direct supervision for state transitions; in contrast, interaction trajectories naturally connect a preceding observation, an action, and a subsequent observation, offering direct supervision for local state transitions, while complete trajectories reveal dependencies among consecutive transitions. We therefore introduce Spatial-Interactor, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories. Accordingly, we construct the Learning from Spatial Interaction dataset (LSI-108K) from simulated and real interaction trajectories, with tasks aligned with the objective of each level. Our two-stage training strategy applies Supervised Fine-Tuning (SFT) to L1 and L2 for local transition modeling, and On-Policy Distillation (OPD) then uses privileged self-distillation: a teacher branch given segment-level transition descriptions supervises the student's on-policy CoT, helping the student learn to integrate consecutive transitions over L3 long trajectories. Experiments across multiple VLMs and spatial benchmarks show consistent gains in local transition modeling and long-horizon integration.
Chinese Translation
空间推理对视觉语言模型(VLM)理解物理世界并在其中行动至关重要。在动态环境中进行推理要求VLM感知由物体运动和视角变化引起的局部状态转移,并在长轨迹上对其进行整合,以维持更新的空间状态,然而现有VLM在这两种能力上均存在不足。当前的空间训练主要集中于关于物体属性和空间关系的静态问题,对状态转移提供的直接监督有限;相比之下,交互轨迹天然地将先前的观测、动作和后续观测连接起来,为局部状态转移提供了直接监督,而完整轨迹则揭示了相邻转移之间的依赖关系。因此,我们提出Spatial-Interactor,一个通过交互训练VLM建模物理世界状态转移的框架,并将该学习过程组织为三级课程:L1被动世界状态转移、L2主动自身状态转移,以及L3长程交互轨迹。据此,我们从模拟和真实交互轨迹中构建了空间交互学习数据集(LSI-108K),其任务与各级目标相匹配。我们的两阶段训练策略首先通过监督微调(SFT)在L1和L2上进行局部转移建模,随后采用策略内蒸馏(OPD)进行特权自蒸馏:给定段级转移描述的教师分支对学生在策略内思维链(CoT)进行监督,帮助学生学习在L3长轨迹上整合相邻转移。在多个VLM和空间基准上的实验表明,该方法在局部转移建模和长程整合方面均取得了一致的提升。
cs.AI / 37 / 2609.23043

Enforcing Narrative Reliability and Epistemic Pacing in LLM-Driven Detective Games via Structured Knowledge Trees

通过结构化知识树在LLM驱动的侦探游戏中强制保障叙事可靠性与认知节奏
Rahmati, Parsa, Zhao, Richard
Abstract
Large Language Models (LLMs) enable open-ended dialogue in interactive games, but their non-deterministic outputs make it difficult to preserve authorial control, factual consistency, and the intended sequence of information disclosure. These challenges are particularly significant in detective games, where premature revelation or fabricated details can undermine the logic of player progression. We present a Structured Knowledge Tree architecture coupled with a tri-agent LLM pipeline for controlling dialogue in an open-ended interrogation game. The system separates knowledge retrieval, dialogue generation, and response verification to ensure that the virtual suspect reveals only information permitted by the current narrative state. We evaluate the approach through The Interrogation of Adrian Gale, a playable detective-game testbed, and a formal user study examining hallucination reduction, adherence to authored disclosure sequences, and perceived logical progression. Our results demonstrate that the structured architecture reduces critical hallucinations by 64.78% and entirely prevents premature narrative disclosure. While the strict mechanical constraints introduced usability trade-offs regarding forced conversational reveals, the system successfully enforces rigorous epistemic pacing and provides players with a clear, subjective sense of progression toward solving the case.
Chinese Translation
大语言模型(LLM)为互动游戏中的开放式对话提供了可能,但其非确定性的输出使得维持作者控制、事实一致性以及预期的信息披露顺序变得困难。这些挑战在侦探游戏中尤为突出,因为过早揭示真相或虚构细节会破坏玩家推进游戏进程的逻辑。我们提出了一种结构化知识树(Structured Knowledge Tree)架构,结合三智能体LLM流水线,用于控制开放式审讯游戏中的对话。该系统将知识检索、对话生成与回复验证相互分离,以确保虚拟嫌疑人仅透露当前叙事状态所允许的信息。我们通过《审讯阿德里安·盖尔》(The Interrogation of Adrian Gale)——一个可玩的侦探游戏测试平台——以及一项正式的用户研究对该方法进行评估,考察幻觉减少程度、对预设信息披露顺序的遵循情况以及玩家感知到的逻辑推进。结果表明,该结构化架构将关键幻觉减少了64.78%,并完全杜绝了叙事信息的过早披露。尽管严格的机械性约束在强制对话揭示方面带来了可用性上的权衡,但该系统成功实现了严谨的认知节奏控制,并为玩家提供了清晰的主观破案推进感。
cs.AI / 38 / 2609.23058

LazyAgent: Demand-Driven Materialization and Physical Optimization of Agentic Programs

LazyAgent:面向智能体程序的按需物化与物理优化
Heng, Xin
Abstract
Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present LazyAgent, a unified execution framework for agent-authored programs organized around a live, goal-derived demanded set. LazyAgent refreshes a backward closure from requested outputs as execution state changes and materializes a ready node only when the active goal requires it. This replaces repeated local judgments with one linear-time graph analysis followed by constant-time membership tests, allowing programs to remain broad while execution stays request-specific. On programs that describe more than the current request needs, LazyAgent consistently outperforms the strongest goal-stopping eager baseline by refusing unrelated work before it starts. Adding one unrelated product raises the eager bill by 22.5% and LazyAgent's by 0.0%. LazyAgent saves 42.0% of measured CPU on production scientific workflows and 51.7% of container time on a live release gate spanning four repositories. We also prove and verify exact equivalence when the request reaches the whole graph, leaving no unrelated work to avoid. Beyond permission, goal-relative output projection saves up to approximately 90% of a shared step on two third-party test suites while the identical eager control saves 0.0%; the advantage disappears when the omitted output has no other consumer or the request needs it. Ordering, reuse, and pruning can also save cost, but do not replace permission. Finally, we show that current public benchmarks are eager-shaped and contain almost no unrequested work. A pre-registered planning intervention did not broaden them. These findings motivate benchmarks built from standing programs and sequences.
Chinese Translation
当前“先规划后执行”的智能体运行时通常在某个步骤就绪后立即执行它。我们提出了 LazyAgent,一个面向智能体自编程序的统一执行框架,其核心是一个动态的、由目标派生的需求集(demanded set)。随着执行状态的变化,LazyAgent 从请求的输出出发刷新一个反向闭包,并且仅当活跃目标需要时才物化就绪节点。这用一次线性时间的图分析加上常数时间的成员测试,取代了反复的局部判断,使程序可以保持宽泛,而执行仍然针对具体请求。在描述内容超出当前请求所需范围的程序上,LazyAgent 通过在无关工作开始之前将其拒绝,始终优于最强的目标停止式急切(eager)基线。增加一个无关产品会使急切方案的开销增加 22.5%,而 LazyAgent 的开销增加为 0.0%。LazyAgent 在生产级科学工作流上节省了 42.0% 的实测 CPU 时间,在一个跨越四个代码仓库的实时发布门禁上节省了 51.7% 的容器时间。我们还证明并验证了当请求覆盖整个图(即没有可避免的无关工作)时的精确等价性。除权限之外,相对于目标的输出投影在两个第三方测试套件上可节省共享步骤高达约 90% 的开销,而完全相同的急切对照方案节省为 0.0%;当被省略的输出没有其他消费者或请求需要该输出时,这一优势即消失。排序、重用和剪枝也能节省成本,但无法替代权限机制。最后,我们表明当前的公开基准测试呈现“急切”形态,几乎不包含任何未被请求的工作。一项预先注册的规划干预也未能使其扩展。这些发现促使我们构建基于常驻程序和序列的基准测试。
cs.AI / 39 / 2609.23064

FireWorldBench: Benchmarking Complex Physical World Intelligence through Coupled-Field Fire Dynamics

FireWorldBench:通过耦合场火灾动力学评估复杂物理世界智能的基准测试
Chen, Qiang, Guo, Hao, Zhu, Huatai, Huang, Tairan, Cao, Yichao, Xu, Hongyan, Huang, Keke, Li, Haifeng, Chen, Yi, Su, Xiu
Abstract
Understanding the physical world requires more than object recognition, scene description, and short-term visual prediction, as real-world physical systems involve multiple continuous fields, latent causal mechanisms, partial observations, and intervention-sensitive dynamics. We propose FireWorldBench, a benchmark for evaluating complex physical world intelligence in multimodal large language models and agents through coupled-field fire dynamics. Fire provides a canonical stress-test environment, where multiple interacting physical fields jointly shape observable states and temporal dynamics. FireWorldBench is organized along two complementary axes, a physical capability axis and a fire scenario task axis, jointly covering physical-state understanding, temporal dynamics, causal mechanisms, and intervention reasoning. The benchmark comprises 520 fire-world entries, including 494 controlled simulation worlds and 26 real-world-aligned event groups, spanning 47 scene archetypes across 7 environment families. These entries combine structured textual observations, multiple 2D physical-field visualizations, and 3D event-level scene modeling, yielding 9,074 text-image interleaved question-answer pairs across choice-based and open-ended report-generation formats. FireWorldBench evaluates whether models can infer latent physical states, explain underlying mechanisms, forecast coupled-field evolution, and assess intervention consequences from multimodal partial observations, providing a challenging testbed for complex physical world intelligence.
Chinese Translation
理解物理世界不仅仅是物体识别、场景描述和短期视觉预测,因为现实世界的物理系统涉及多个连续场、潜在因果机制、部分可观测性以及对干预敏感的动力学特性。我们提出了FireWorldBench,一个通过耦合场火灾动力学评估多模态大语言模型与智能体复杂物理世界智能的基准。火灾提供了一个典型的压力测试环境,其中多个相互作用的物理场共同塑造可观测状态和时间演化。FireWorldBench沿两个互补维度组织:物理能力维度和火灾场景任务维度,共同涵盖物理状态理解、时间动力学、因果机制和干预推理。该基准包含520个火灾世界条目,其中494个为受控仿真世界,26个为对齐真实世界的事件组,横跨7个环境类别中的47种场景原型。这些条目结合了结构化文本观测、多幅二维物理场可视化以及三维事件级场景建模,产生了9,074个图文交错问答对,涵盖选择题和开放式报告生成两种形式。FireWorldBench评估模型能否从多模态部分观测中推断潜在物理状态、解释底层机制、预测耦合场演化并评估干预后果,为复杂物理世界智能提供了一个具有挑战性的测试平台。
cs.AI / 40 / 2609.23071

Tutoring Large Language Models to be Domain-adaptive, Precise and Safe

引导大语言模型实现领域自适应、精准与安全
Banerjee, Somnath
Abstract
This thesis proposes a framework for "responsible intelligence" to address AI's critical challenges in safety, ethics, and cultural sensitivity. It advances three core areas: First, it improves domain adaptation in specialized fields using active learning and graph-based knowledge to reduce hallucinations. Second, it enhances ethical rigor via a novel decoding-time alignment mechanism that proactively blocks harmful text generation in real-time. Finally, it ensures cultural and multilingual safety through language-specific steering that respects diverse linguistic and social norms. Ultimately, this work provides a blueprint for building next-generation AI that is contextually knowledgeable, ethically sound, and culturally adaptable.
Chinese Translation
本论文提出了一个“负责任智能”框架,以应对人工智能在安全性、伦理和文化敏感性方面的关键挑战。论文在三个核心领域取得进展:首先,利用主动学习和基于图的知识来提升专业领域的自适应能力,以减少幻觉现象;其次,通过一种新颖的解码时对齐机制,实时主动阻断有害文本的生成,从而增强伦理严谨性;最后,通过尊重不同语言与社会规范的语言特定引导,确保文化层面的安全与多语言安全。最终,本工作为构建具有情境知识、伦理可靠且文化可适应的下一代人工智能提供了蓝图。
cs.AI / 41 / 2609.23074

Event Signature Transfer: Model-Agnostic Forecast Scenario Construction from Historical Events

事件特征迁移:基于历史事件的模型无关预测情景构建
Sridhar, Karthik, Jain, Aaditya, Mandal, Murari, Deshpande, Saurabh
Abstract
Forecasters often know an event is imminent but not the shape, size, or timing of its effect. We introduce Event Signature Transfer (EST), a training-free, model-agnostic operator that turns a completed past event into an explicit forecast scenario. EST removes a source event's own trend and seasonality, then scales and retimes the remaining event signature onto a native forecast, preserving the forecast's linked structure and reducing to it exactly at zero strength. Because it reads only output quantiles, EST applies to any quantile forecaster, with no training, no model internals, at transfer time. Across twelve real episodes and ten synthetic scenarios on Chronos-2, TimesFM-2.5 and Toto-2.0, manually configured EST reduces real-episode WQL by 21.7-90\% in-sample. On Chronos-2, it leads eleven of twelve matched comparisons against covariate conditioning, activation editing and raw replay. The operator builds a scenario; it does not estimate its likelihood.
Chinese Translation
预测者往往知道某事件即将发生,但并不了解其影响的形态、规模或时间。我们提出了事件特征迁移(Event Signature Transfer, EST),这是一种无需训练、模型无关的算子,能够将一个已完成的过去事件转化为显式的预测情景。EST 首先去除源事件自身的趋势和季节性成分,然后对剩余的事件特征进行缩放和时间重对齐,并将其叠加到原生预测之上,从而保留预测原有的关联结构,并在强度为零时精确还原为原生预测。由于它仅读取输出的分位数信息,EST 可适用于任何分位数预测器,且在迁移时无需训练、无需访问模型内部结构。在 Chronos-2、TimesFM-2.5 和 Toto-2.0 上的十二个真实事件片段和十个合成情景实验中,手动配置的 EST 在真实事件上使样本内 WQL 降低了 21.7–90%。在 Chronos-2 上,其在十二组对照比较中的十一组中优于协变量条件化、激活编辑和原始重放方法。该算子用于构建预测情景,而非估计情景发生的可能性。
cs.AI / 42 / 2609.23130

From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving

从推理引擎到推理控制平面:连接 vLLM、llm-d 与高效分布式大语言模型推理服务的演进
Sisodia, Twinkll
Abstract
Large language model (LLM) inference is evolving from an engine-local optimization problem into a distributed control problem involving reusable state, phase placement, heterogeneous accelerators, networking, autoscaling, reliability, and service-level objectives. This paper connects that transition across peer-reviewed systems research, open-source implementations, and documented production studies. It treats vLLM and llm-d as complementary layers: model-serving engines optimize execution through mechanisms such as PagedAttention, continuous batching, kernels, quantization, and parallelism, while an inference control plane can optimize where, when, and under what policy execution occurs across a fleet. The contribution is synthesis rather than a new benchmark; all reported performance and deployment results remain attributed to their original sources. The combined evidence suggests that the scarce resource in modern inference is shifting from raw FLOPs alone toward managed state, placement, network movement, reliability, and decision quality. We propose an Inference Execution Planner that selects feasible execution plans rather than only endpoints, including aggregated versus disaggregated topology, KV source and transfer action, hardware variant, routing/admission policy, and slower scaling decisions. We also provide a source-local benchmark atlas, a bottleneck-migration taxonomy, practical deployment guidance, an evaluation framework based on SLO-goodput, and research questions for agentic, multimodal, heterogeneous, and resilient inference.
Chinese Translation
大语言模型(LLM)推理正在从引擎本地的优化问题演变为一个分布式控制问题,涉及可复用状态、阶段放置、异构加速器、网络、自动伸缩、可靠性以及服务级目标(SLO)。本文在同行评审的系统研究、开源实现以及有据可查的生产实践之间建立联系,梳理这一转变。本文将 vLLM 和 llm-d 视为互补的两个层次:模型服务引擎通过 PagedAttention、连续批处理(continuous batching)、内核、量化和并行等机制优化执行,而推理控制平面(inference control plane)则可以优化执行在整个集群中于何处、何时以及依据何种策略发生。本文的贡献在于综合梳理而非新的基准测试;所有报告的性能与部署结果均归功于其原始来源。综合证据表明,现代推理中的稀缺资源正从单纯的原始 FLOPs 转向受管理的状态、放置、网络传输、可靠性与决策质量。我们提出一种推理执行规划器(Inference Execution Planner),它选择可行的执行计划而非仅选择端点,包括聚合与分离式拓扑、KV 缓存来源与传输动作、硬件变体、路由/准入策略以及更缓慢的伸缩决策。我们还提供了源本地的基准测试图谱(benchmark atlas)、瓶颈迁移分类法、实用的部署指南、基于 SLO-goodput 的评估框架,以及面向智能体式(agentic)、多模态、异构和弹性推理的研究问题。
cs.AI / 43 / 2609.23142

CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine

CraftBench-UE:面向虚幻引擎中编程智能体的确定性评估
Wu, Shutong, Calderone, Kevin, Tsen, Andy
Abstract
Building gameplay features in a game engine requires more than code, as code that compiles and runs does not necessarily implement the requested gameplay. We introduce CraftBenchUE, an evaluation harness that runs agents in an isolated Unreal Engine environment, reconstructs their saved submissions in fresh projects, and applies deterministic build, asset, and runtime checks without an LLM judge. Based on the harness, we built a benchmark consisting of 70 tasks spanning C++ source, Blueprint assets, and editor scripting. We evaluate seven models under two editor-tool configurations, with a file-and-shell baseline on C++ tasks. We further pair tasks that specify the same gameplay and use the same runtime tests, but require C++ and Blueprint as the deliverables. Across the 10 paired tasks, C++ completion rates exceed Blueprint by 30.0 and 42.9 percentage points in the two tool configurations. Among on-time Blueprint submissions in this paired set that pass asset checks, 42.2% and 50.0% fail explicit runtime assertions. These submissions satisfy asset requirements but fail the required gameplay tests. We will release the harness, task benchmark, and our trajectory findings with the report.
Chinese Translation
在游戏引擎中构建玩法功能不仅仅需要代码,因为能够编译和运行的代码并不一定实现了所要求的玩法。我们提出了 CraftBench-UE,这是一个评估框架,它在隔离的虚幻引擎(Unreal Engine)环境中运行智能体,在新项目中重建其保存的提交成果,并应用确定性的构建、资产和运行时检查,而无需 LLM 评判者。基于该框架,我们构建了一个包含 70 个任务的基准测试,涵盖 C++ 源代码、蓝图(Blueprint)资产和编辑器脚本。我们在两种编辑器工具配置下评估了七个模型,并在 C++ 任务上采用文件与命令行基线。我们进一步将指定相同玩法并使用相同运行时测试、但分别要求以 C++ 和蓝图作为交付成果的任务进行配对。在 10 个配对任务中,两种工具配置下 C++ 的完成率分别超出蓝图 30.0 和 42.9 个百分点。在这一配对集合中按时提交并通过资产检查的蓝图提交中,分别有 42.2% 和 50.0% 未通过明确的运行时断言。这些提交满足了资产要求,但未通过所需的玩法测试。我们将随本报告一并发布该框架、任务基准以及我们的轨迹分析结果。
cs.AI / 44 / 2609.23201

Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation

不要相信基准测试:通用大语言模型排名的局限性及任务专用评估的必要性
Amin, Danial
Abstract
Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretable only in relation to the system tested, the questions included, and the conditions of evaluation. This perspective examines five connected limitations of general LLM rankings: differences between evaluated and publicly available systems; commercial incentives and dependencies in external evaluation; benchmark saturation, defective tests, and data contamination; models exploiting scoring procedures; and the limited relevance of general scores to users' tasks. Documented cases illustrate why these problems require different responses. I argue for evaluation procedures that disclose the tested configuration, validate questions and successful task completion, report performance alongside cost and execution time, and make the scope of generalization explicit. I then discuss \textbf{Isotanta}, a crowdsourced benchmarking platform, as a practical example of contributed questions and repeated evaluation. A larger question pool may improve task coverage, while repeated sampling can improve the stability of estimates on that pool; neither guarantees validity or personalization. The paper distinguishes the platform's current shared ranking from proposed task-specific and user-provided evaluations. Its central argument is that model selection requires evidence about performance on the intended work, not simply a high position on a general leaderboard.
Chinese Translation
基准测试分数日益影响着大语言模型(LLM)的开发、营销与选型。然而,一个总体分数只有在结合被测试的系统、所包含的题目以及评估条件时才具有可解释性。本文从视角性分析出发,考察了通用大语言模型排名的五个相互关联的局限性:参与评估的系统与公开可用系统之间的差异;外部评估中的商业激励与依赖关系;基准测试饱和、存在缺陷的测试题以及数据污染;模型对评分机制的钻营利用;以及通用分数与用户实际任务之间关联性的不足。已有记录的案例说明了为何这些问题需要不同的应对方式。我主张采用能够披露被测配置、验证题目质量与任务是否成功完成、在报告性能的同时报告成本与执行时间、并明确泛化范围边界的评估流程。随后,我讨论了 Isotanta——一个众包基准测试平台——作为由用户贡献题目并进行重复评估的实际示例。更大的题目池可以改善任务覆盖面,而重复采样可以提高在该题目池上估计结果的稳定性;但二者均无法保证评估的有效性或个性化。本文区分了该平台当前共享排名机制与所提出的任务专用及用户自定义评估方式。其核心论点是:模型选型需要的是关于模型在目标任务上表现的证据,而不仅仅是在通用排行榜上的高排名。
cs.AI / 45 / 2609.23293

Expansion Counts under Standard A* Tie-Breaking Strategies on the Final Plateau

标准A*平局打破策略在最终平台层上的节点扩展次数
Fukunaga, Alex
Abstract
In the A* search algorithm, the tie-breaking strategies for nodes with the same $f$-value determines which states A* expands on the final $f$-layer. For nine standard tie-breaking strategies, we show that under a consistent heuristic, every pair has positive-cost instances favoring each strategy over the other by an arbitrarily large additive expansion gap. A parameterized unit-cost grid example also gives unbounded expansion-count ratios between low-$h$ with FIFO and LIFO. In unit-cost search with $h > 0$ at non-goals, exact heuristic values near the goal lead to complementary extremal results: low-$h$ minimizes the number of remaining expansions from a common configuration within the perfect region, while high-$h$ maximizes the total number of expansions when every final-plateau state with $h=1$ is a goal predecessor. Finally, with the evaluation function $f_{\alpha} = g + \alpha h$, when $h>0$ at non-goals, every heuristic weight $0 \leq \alpha<1$ eliminates tie-breaking sensitivity, and all tie-breaking strategies expand the same set of states.
Chinese Translation
在A*搜索算法中,针对具有相同$f$值的节点的平局打破策略决定了A*在最终$f$层上扩展哪些状态。对于九种标准平局打破策略,我们证明了在一致性启发式函数下,对于任意一对策略,均存在正代价实例使得其中一种策略相对于另一种策略在扩展次数上具有任意大的加性优势。一个参数化的单位代价网格示例还表明,低$h$与FIFO组合和低$h$与LIFO组合之间的扩展次数比值是无界的。在非目标节点满足$h>0$的单位代价搜索中,目标附近精确的启发式值导致互补的极值结果:在完美区域内,从同一配置出发,低$h$最小化剩余扩展次数;而当每个$h=1$的最终平台层状态均为目标的前驱时,高$h$最大化总扩展次数。最后,对于评价函数$f_{\alpha} = g + \alpha h$,当非目标节点满足$h>0$时,任意启发式权重$0 \leq \alpha<1$均可消除对平局打破策略的敏感性,即所有平局打破策略扩展的状态集合相同。
cs.AI / 46 / 2609.23363

TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents

TicTacBench:面向代码智能体时序收敛能力的基准测试
Wang, Bowei, Fang, Zhigang, Yang, Zhijie, Chen, Renzhi, Li, Shanshan, Wang, Lei
Abstract
Recent advances in large language models (LLMs) have led to the emergence of coding agents capable of performing complex engineering tasks, including register-transfer level (RTL) design and optimization. Existing RTL benchmarks mainly evaluate functional correctness and performance, power, and area (PPA) of the generated RTL designs, leaving agents' ability for \emph{timing closure} under-evaluated. We propose TicTacBench, a benchmark specifically designed to evaluate coding agents' capabilities for RTL-level timing closure under post-place-and-route (post-PnR) evaluation. TicTacBench contains 30 diverse tasks, each provided with a suboptimal RTL design, realistic timing constraints, functional equivalence verification, and timing reports. With over 300 runs of coding agents driven by 8 frontier LLMs, we find that even the best agent can only close 53.3\% of tasks with 7.18\% area-delay product (ADP) degradation and 8.83\% energy-delay-squared product (EDDP) improvement on average. We identify common failure categories that explain why agents fail to close timing. Then we propose TicTacSkill, a new method that guides agents to follow standard timing-closure procedures and improves the Timing Closure Rate by 9\%. These results suggest that while coding agents have made significant progress in RTL design, their timing-closure capability still has substantial room for improvement.
Chinese Translation
大语言模型(LLM)的最新进展催生了能够执行复杂工程任务的代码智能体,包括寄存器传输级(RTL)设计与优化。现有的RTL基准主要评估所生成RTL设计的功能正确性以及性能、功耗和面积(PPA),而对智能体的时序收敛(timing closure)能力评估不足。我们提出了TicTacBench,这是一个专门用于评估代码智能体在布局布线后(post-PnR)评估下进行RTL级时序收敛能力的基准。TicTacBench包含30个多样化任务,每个任务均提供了次优的RTL设计、现实的时序约束、功能等价性验证以及时序报告。通过对由8个前沿LLM驱动的代码智能体进行300余次运行实验,我们发现即使是最优的智能体也只能完成53.3%的任务的时序收敛,且面积-延迟积(ADP)平均恶化7.18%,能量-延迟平方积(EDDP)平均改善8.83%。我们归纳了导致智能体无法完成时序收敛的常见失败类别。在此基础上,我们提出了TicTacSkill,一种引导智能体遵循标准时序收敛流程的新方法,可将时序收敛率提升9%。这些结果表明,尽管代码智能体在RTL设计方面已取得显著进展,但其时序收敛能力仍有巨大的提升空间。
cs.AI / 47 / 2609.23378

Leaky-integrator reconstruction: taming error accumulation in recursive differenced time-series forecasting

泄漏积分器重构:抑制递归差分时间序列预测中的误差累积
Yang, Zijiang
Abstract
We introduce leaky-integrator reconstruction, a training-free method that cures the error accumulation of recursive differenced forecasting. Our first contribution is diagnostic: predicting one-step changes and integrating them by cumulative summation, the standard remedy for non-stationarity, is a discrete integrator with a pole on the unit circle, and we show this makes recursive rollout of a nonlinear model diverge, its 336-step error reaching several times that of a well-behaved forecaster (normalised MAE 1.6-3.8 versus about 0.8) across every neural architecture tested. Our second, central contribution is the fix: move the pole inside the unit circle with a leaky integrator H(z) = 1/(1 - gamma z^-1), gamma < 1, which provably bounds the accumulated error variance. Applied at reconstruction time with a single fixed gamma=0.9 (no retraining, a two-line change to any deployed one-step or foundation-model forecaster), it shrinks error at every horizon, the mean gain over seven diverging architectures and twenty datasets growing from ~3% at H=24 to 23% at H=96, 37% at H=192 and 51% (43-74% across those architectures) at H=336 (78% with an oracle pole). Crucially, it is provably inert where no pathology exists (stable or joint predictors already at the irreducible rate), making it a safe, general default.
Chinese Translation
我们提出泄漏积分器重构(leaky-integrator reconstruction),这是一种无需训练的方法,能够治愈递归差分预测中的误差累积问题。我们的第一项贡献是诊断性的:预测一步变化量并通过累加求和进行积分——这是应对非平稳性的标准方案——实际上是一个极点位于单位圆上的离散积分器;我们证明这会导致非线性模型的递归滚动预测发散,其336步误差达到性能良好预测器的数倍(归一化MAE为1.6–3.8,而后者约为0.8),且在所有测试过的神经架构中均是如此。我们的第二项也是核心贡献是修复方案:采用泄漏积分器 H(z) = 1/(1 - gamma z^-1)(gamma < 1)将极点移入单位圆内,可证明这能够约束累积误差方差。该方法在重构阶段应用,仅需一个固定的 gamma=0.9(无需重新训练,只需对任何已部署的单步预测器或基础模型预测器做两行代码的修改),即可在所有预测步长上缩减误差:在七个发散架构和二十个数据集上的平均增益从 H=24 时的约3%增长至 H=96 时的23%、H=192 时的37%以及 H=336 时的51%(各架构间为43–74%;采用 oracle 极点时为78%)。至关重要的是,可以证明在不存在病态的情形下(稳定或联合预测器已达到不可约误差率时),该方法是惰性的,因而它是一种安全、通用的默认方案。
cs.AI / 48 / 2609.23512

AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction

AgentBetta:通过选择性扩展与验证性收缩实现验证驱动的AI纳米智能体自适应配置
Babu, Md. Ashraful
Abstract
Large language model agents are typically deployed with predefined configurations, although the required model capability, context, tools, permissions, memory, and computational resources can vary substantially across tasks. This study develops and evaluates AgentBetta, an adaptive AI Nano-Agent framework that represents these factors as an executable configuration and updates them through verification-driven diagnosis, selective expansion, and verification-based counterfactual contraction. The evaluation distinguishes controlled mechanism validation from external agent comparisons. On the AB-ConfigBench benchmark, AgentBetta achieved 91.38% verified success while reducing median context allocation from 64,000 to 8,000 context characters and median tool exposure from five tools to zero compared with the fully provisioned configuration. The configuration-deficiency diagnosis achieved a macro-F1 score of 0.819 with precision of 1.000 across the evaluated dimensions, and selective expansion avoided unnecessary changes to unrelated configuration dimensions. Post-success contraction preserved verification outcomes in 56.41% of evaluated one-dimension contraction probes, indicating that some successful configurations contained removable capability under the tested conditions. External evaluations indicate that adaptive configuration can improve the balance between verified task completion and capability exposure; however, the results vary across benchmarks and agent families. In particular, the cross-family replication did not reproduce the primary-backbone accuracy ordering, and specialized systems remained advantageous for certain task domains. These results support interpreting AgentBetta as a configuration-adaptation mechanism that regulates capability allocation and inference expenditure rather than as a universal replacement for specialized agent architectures.
Chinese Translation
大语言模型智能体通常以预定义配置进行部署,然而不同任务所需的模型能力、上下文、工具、权限、记忆和计算资源可能存在显著差异。本研究开发并评估了AgentBetta,这是一个自适应AI纳米智能体框架,它将这些因素表示为可执行配置,并通过验证驱动的诊断、选择性扩展和基于验证的反事实收缩对其进行更新。评估将受控机制验证与外部智能体比较区分开来。在AB-ConfigBench基准上,与完全配置的设置相比,AgentBetta实现了91.38%的验证成功率,同时将中位上下文分配从64,000个上下文字符减少到8,000个,中位工具暴露从五个工具减少到零个。配置缺陷诊断在被评估的各个维度上取得了0.819的宏平均F1分数和1.000的精确率,且选择性扩展避免了对无关配置维度的不必要更改。成功后的收缩在56.41%的单维度收缩探针中保持了验证结果,表明某些成功配置在测试条件下包含可移除的能力。外部评估表明,自适应配置可以改善验证任务完成度与能力暴露之间的平衡;然而,结果因基准和智能体系列而异。特别是,跨系列复现未能重现主骨干模型的准确率排序,且专用系统在某些任务领域仍具有优势。这些结果支持将AgentBetta理解为一种配置自适应机制,用于调节能力分配和推理开销,而非作为专用智能体架构的通用替代方案。
cs.AI / 49 / 2609.23640

Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment

与人类对齐的模型是人类的模型吗?偏好对齐中的图灵测试鸿沟
Yuan, Suqin, Lin, Runqi, Li, Muyang, Hong, Guanzhe, Gu, Jindong, Feng, Lei, Russell, Chris, Liu, Tongliang
Abstract
Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We distinguish alignment with human preferences from alignment with human behavior, and show that alignment with human preferences can make model behavior less human-like even when both preferences and responses come entirely from humans. We call this the Turing-test gap. We show that preference alignment preserves the human response distribution only under a restrictive condition, and find no consistent evidence that real human preferences satisfy it. Empirically, the loss of human-response likelihood increases with the strength of preference weighting, regardless of its direction, and the gap also appears under standard DPO. These results establish human-likeness as an explicit dimension of alignment rather than something assumed to follow from preference alignment.
Chinese Translation
人类反馈对齐已使语言模型成为有用的助手,通常被描述为使其与人类对齐。然而,人们偏好的AI回答未必是他们自己会给出的回答。我们区分了与人类偏好的对齐和与人类行为的对齐,并证明即使在偏好和回答完全来自人类的情况下,与人类偏好的对齐也可能使模型行为更不像人类。我们将其称为图灵测试鸿沟(Turing-test gap)。我们证明偏好对齐仅在 restrictive 条件下才能保持人类回答分布,且未发现任何一致证据表明真实的人类偏好满足该条件。实证结果表明,人类回答似然的损失随偏好权重的强度增加而增加,与权重方向无关,且该鸿沟在标准DPO(Direct Preference Optimization)下同样出现。这些结果确立了“类人性”是对齐的一个显式维度,而非可假定随偏好对齐自动实现的属性。
cs.AI / 50 / 2609.23695

PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI

PhysAI-Bench:面向以自主无人机为中心的物理智能中基于大语言模型的智能体决策基准测试
Ferrag, Mohamed Amine, Debbah, Merouane, Lakas, Abderrahmane, Perumkunnil, Manu, Tihanyi, Norbert
Abstract
Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as unmanned aerial vehicles (UAVs), which must perceive, reason, plan, and act in dynamic environments. Existing benchmarks assess physical perception, intuitive physics, embodied navigation, and collaborative reasoning, but rarely evaluate the agentic decision-making required for reliable autonomy. We introduce \textit{PhysAI-Bench}, a benchmark for evaluating this capability. It contains 10,178 standardized decision instances automatically extracted from conversational traces of autonomous UAV missions. Each instance preserves mission context, temporal dependencies, physical constraints, Model Context Protocol (MCP) tool calls, Agent-to-Agent (A2A) interactions, sensor observations, and AI-native 6G network conditions, including latency, packet loss, throughput, edge load, and network slicing. We expose only information preceding each decision, preventing future-event leakage and approximating online decision-making. We evaluate 29 foundation models using a two-stage protocol. We select model-specific configurations from 12 combinations of zero-, three-, and five-shot prompting and four temperatures, tested in three runs on a 35-instance, human-verified development set. We then freeze each selected configuration and evaluate it in three runs on a fixed, episode-disjoint set of 500 instances. GPT-5.3 achieves the highest accuracy (52.00%), followed by GPT-5.2 (49.40%) and Grok~4.5 (49.07%). Few-shot prompting generally improves performance, while temperature has limited influence. The results demonstrate that reliable agentic decision-making in Physical AI remains an open challenge. The dataset is available at https://github.com/maferrag/physai-bench
Chinese Translation
物理智能(Physical AI)的最新进展加速了基础模型在自主系统(如无人机,UAV)中的应用,这类系统必须在动态环境中进行感知、推理、规划和行动。现有基准测试评估了物理感知、直觉物理、具身导航和协作推理,但很少评估实现可靠自主性所需的智能体决策能力。我们提出了 PhysAI-Bench,一个用于评估该能力的基准测试。它包含 10,178 个标准化的决策实例,这些实例从自主无人机任务的对话轨迹中自动提取。每个实例保留了任务上下文、时间依赖关系、物理约束、模型上下文协议(Model Context Protocol, MCP)工具调用、智能体间(Agent-to-Agent, A2A)交互、传感器观测,以及 AI 原生 6G 网络条件(包括时延、丢包率、吞吐量、边缘负载和网络切片)。我们仅提供每个决策之前的信息,防止未来事件泄露,从而近似在线决策过程。我们采用两阶段协议评估了 29 个基础模型。首先,在经过人工验证的 35 实例开发集上进行三次运行测试,从零样本、三样本和五样本提示与四种温度的 12 种组合中选择模型特定的配置。然后,冻结每个选定的配置,在一个固定的、情节不重叠的 500 实例集合上进行三次运行评估。GPT-5.3 取得了最高准确率(52.00%),其次是 GPT-5.2(49.40%)和 Grok 4.5(49.07%)。少样本提示通常能提升性能,而温度的影响有限。结果表明,物理智能中可靠的智能体决策仍然是一个开放的挑战。数据集可在 https://github.com/maferrag/physai-bench 获取。
cs.AI / 51 / 2609.23735

ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents

ScholarStack:面向科学智能体的分层研究资产编排与跨任务复用
ScholarSeed AI Team, Zhang, Ao, Gong, Caoqinwei, Wang, Guanglei, Zhang, Haifan, Zhang, Hanwei, Sheng, Jiayi, Zhang, Jihai, Ying, Kai, Dai, Liyun, Zhu, Tingyu, Chen, Wei, Luo, Wei, Fang, Wenkai, Qiu, Xiaoyu, Jiang, Xue, Wang, Yi, Cao, Yuan, Yu, Zheng, Yin, Wotao
Abstract
Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim assessment. Most existing systems, however, are organized around individual tasks: the same papers are repeatedly retrieved, segmented, and interpreted, and the understanding built in one task is difficult to reuse in the next. We present ScholarStack, a layered research asset framework that compiles a paper collection into reusable, versioned, and provenance-preserving assets at three complementary levels: source-grounded paper-level statements, domain-level organization, and evidence-grounded cross-paper syntheses. A common access interface returns task-specific views at the evidence granularity each task requires, preserving study conditions, source traceability, and verification status. We instantiate the framework on four task families spanning ten task settings, comparing agents that use the compiled assets with task-specific baselines under matched base models. Quality gains concentrate on tasks that require cross-paper evidence, such as multi-paper question answering and literature review generation, and query-time token cost falls on every task where it is measured, with assets compiled once and reused across tasks. These results suggest that layered research assets can serve as shared infrastructure for scientific agents, shifting literature-based assistance from isolated document processing toward cumulative, evidence-grounded workflows.
Chinese Translation
科学智能体(scientific agents)支持一系列基于文献的研究任务,例如文献检索、问答、有证据支撑的生成以及论断评估。然而,现有大多数系统都是围绕单个任务组织的:同样的论文被反复检索、切分和解读,并且在一个任务中形成的理解难以在下一个任务中复用。我们提出了 ScholarStack,这是一个分层研究资产框架,它将论文集合编译为可复用、可版本化且保留来源出处的资产,涵盖三个互补层次:基于原文的论文级陈述、领域级组织,以及基于证据的跨论文综合。通过统一访问接口,该框架以每个任务所需的证据粒度返回特定任务的视图,同时保留研究条件、来源可追溯性和验证状态。我们在涵盖十个任务设定的四个任务族上实例化该框架,在相同基础模型条件下,将使用编译资产的智能体与特定任务的基线进行比较。质量提升主要集中在需要跨论文证据的任务上,例如多论文问答和文献综述生成;且在所有测量的任务中,查询时的 token 开销均有所下降,因为资产只需编译一次即可跨任务复用。这些结果表明,分层研究资产可以作为科学智能体的共享基础设施,推动基于文献的辅助从孤立的文档处理转向累积式、基于证据的工作流程。
cs.AI / 52 / 2609.23774

On Probabilistic Inference Through Parametric Tensor Decomposition in Base Tensor Networks

基于基张量网络中参数化张量分解的概率推断研究
Hamid, Sagad, Braun, Tanya
Abstract
Probabilistic inference is generally only tractable in low-treewidth graphical models, limiting its effective applicability in high-treewidth settings. Many existing methods improve efficiency by exploiting specific parametric structure, such as symmetries. However, they typically require such structure to be explicitly present, limiting their applicability to a broader range of graphical models. To address this limitation, we propose a framework where tractable inference is controlled by latent parametric structure exploitation, rather than requiring it to be explicitly present a priori. Our approach first reparameterises a graphical model as a specific tensor network representation, which we call a base tensor network. This representation yields two key properties that allow inference tractability to be controlled by parametric structure: 1) First, the complexity of inference is mainly determined by the parametric structure of a single tensor, called the base tensor. We characterise several tractable classes of base tensors for which the entire base tensor network can be contracted efficiently. 2) Second, decomposing the base tensor yields again a collection of base tensor networks. This allows inference to be naturally reduced to decomposing the base tensor into tractable components with sufficient parametric structure. We call this procedure parametric tensor decomposition. By exploiting parametric structure within the base tensor, our framework enables a novel view on inference beyond settings where such structure is explicitly present.
Chinese Translation
概率推断通常仅在低树宽的图模型上是可计算的,这限制了其在高树宽环境中的有效适用性。许多现有方法通过利用特定的参数化结构(如对称性)来提高效率。然而,这些方法通常要求此类结构以显式形式存在,这限制了其在更广泛图模型上的适用性。为解决这一局限,我们提出了一个框架,其中推断的可计算性由对潜在参数化结构的利用来控制,而无需此类结构事先显式存在。我们的方法首先将图模型重新参数化为一种特定的张量网络表示,我们称之为基张量网络(base tensor network)。这种表示带来了两个关键性质,使得推断的可计算性能够由参数化结构来控制:1)首先,推断的复杂度主要由单个张量(称为基张量,base tensor)的参数化结构决定。我们刻画了若干可计算的基张量类别,对于这些类别,整个基张量网络可以被高效地收缩。2)其次,对基张量进行分解会再次得到一组基张量网络。这使得推断可以自然地约化为将基张量分解为具有足够参数化结构的可计算组件。我们将这一过程称为参数化张量分解(parametric tensor decomposition)。通过利用基张量内部的参数化结构,我们的框架为超越此类结构显式存在场景的推断提供了一种全新的视角。
cs.AI / 53 / 2609.23790

Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows

代理总成本:多智能体LLM工作流中记忆注入成本的精确归因
Singh, Vivek Kumar, Priyam, Preeti, Bhowmick, Gautam
Abstract
Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected tokens are billed as input tokens at the same per-token price as the system prompt and the user query. Production observability tools report total token cost but do not separate the tokens a node generates from the tokens it is handed, so this component of the bill is invisible to the teams paying it. We introduce the Total Cost of Agency (TCA), a decomposition of multi-agent workflow cost into base prompt, inference, memory injection, miss penalty and context-accumulation components, and an exact attribution method: a two-pass, non-billable token count that measures injected tokens directly rather than estimating them from word-count proxies. On a 200-task enterprise benchmark executed against real model APIs, memory injection accounts for 13.6 percent of the variable cost a compile-time optimizer can act on, about 12 percent of the full billed cost, and its share rises from a structural zero at workflow depth one to 27.6 percent at depth six. Injected tokens grow linearly with depth over the measured range (R^2 = 0.9974, depths two through six); a quadratic fit yields a negative leading coefficient, so the data do not exhibit convex growth at these depths. We show the component is controllable at fixed model tier: reducing the retrieval window capacity from 32 to 2 entries lowers injected tokens by 28.7 percent with an accuracy change within seed-level variation. We report in full that our graph-rewriting transforms are approximately cost-neutral in isolation, that two of the five decomposition terms are zero by construction in this harness, and that total workflow cost is dominated by model tier assignment, which we hold fixed and treat as prior work. Prompt caching is not evaluated; all figures are for the uncached case.
Chinese Translation
多智能体大语言模型(LLM)工作流中的每个节点都会从记忆中检索上下文并将其注入到提示词(prompt)中,这些被注入的令牌(token)与系统提示词和用户查询按相同的单价计费。生产环境的可观测性工具只报告总令牌成本,却不区分节点自身生成的令牌与被传递给它的令牌,因此账单中这一部分对支付团队而言是不可见的。我们提出代理总成本(Total Cost of Agency, TCA),将多智能体工作流成本分解为基础提示词、推理、记忆注入、未命中惩罚和上下文累积等组件,并提出一种精确归因方法:一种两遍式的、不计费令牌计数方法,直接测量注入的令牌,而非通过字数代理进行估算。在一个针对真实模型API执行、包含200个任务的企业基准测试中,记忆注入占编译期优化器可作用的可变成本的13.6%,约占全部计费成本的12%,且其占比从工作流深度为一时的结构性零上升到深度为六时的27.6%。在测量范围内,注入令牌随深度呈线性增长(R^2 = 0.9974,深度二至六);二次拟合的首项系数为负,因此数据在这些深度上并未表现出凸性增长。我们表明该成本组件在固定模型层级(model tier)下是可控的:将检索窗口容量从32条降低至2条,可使注入令牌减少28.7%,而准确率变化在随机种子波动范围之内。我们完整报告如下事实:我们的图重写变换在孤立情况下近似成本中性;五个分解项中有两项在该实验框架中由构造为零;工作流总成本由模型层级分配主导,我们将其固定并视为先前工作。本文未评估提示词缓存(prompt caching);所有数据均针对无缓存情形。
cs.AI / 54 / 2609.23806

WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks

WorkWorlds:用于评估AI智能体职场任务的基础设施
Hua, Yining, Lian, Levi
Abstract
Many knowledge-work benchmarks are constructed around individual tasks, with the context needed for each task selected together with or after the task has been specified. This design measures performance on workplace-like tasks in an environment assembled for the task. When task specification guides which context is selected, the evaluation can encode task information into the environment and pre-complete part of the information-localization work that workplace performance normally requires. We introduce WorkWorlds, an evaluation infrastructure that separates organizational state from task specification. A world first fixes a revision, date, and employee seat and materializes the organizational state that employee can access; tasks are introduced only afterward. We implement WorkWorlds in a primary synthetic pharmaceutical company with 8 measured tasks across 6 employee seats, and construct additional organizational worlds. Across 192 matched evaluations, moving from task-curated context to the full role-visible workplace reduced evidence access from 90.4% to 74.5% and criterion pass from 79.4% to 68.2%, while pass conditional on evidence access remained nearly unchanged; most of the measured difference occurred before the agent reached sufficient evidence.
Chinese Translation
许多知识工作基准测试围绕单个任务构建,每个任务所需的上下文与任务一同或在其之后选定。这种设计衡量的是在一个为任务而搭建的环境中完成类职场任务的表现。当任务描述引导上下文的选择时,评估可能将任务信息编码进环境,预先完成了职场表现通常需要的信息定位工作。我们提出WorkWorlds,一个将组织状态与任务描述相分离的评估基础设施。一个世界(world)首先固定文档版本、日期和员工职位,并物化该员工可访问的组织状态;任务仅在之后才被引入。我们在一家合成的制药公司中实现了WorkWorlds,包含6个员工职位上的8个受测任务,并构建了额外的组织世界。在192次匹配评估中,从任务策划的上下文切换到完整的角色可见职场环境后,证据获取率从90.4%降至74.5%,标准通过率从79.4%降至68.2%,而在获得证据条件下的通过率几乎不变;所测得的差异大部分发生在智能体获取到足够证据之前。
cs.AI / 55 / 2609.23860

Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models

面向多模态大语言模型的医学视觉编码器预训练
Jiang, Tianyou
Abstract
Multimodal Large Language Models (MLLMs) commonly reuse visual encoders pretrained with CLIP, although the features of these ViTs are ultimately consumed by autoregressive LLMs. We refer to this mismatch as the semantic-interface gap and introduce MedMLIP, a framework that pretrains the visual encoder through report generation with a frozen LLM, while employing Local Relational Distillation (LRD) to preserve relationships among visual patches to avoid visual collapse. We pretrain MedMLIP on IU-Xray and Open-PMC-300K and evaluate the resulting encoders on VQA-RAD and SLAKE. Only the ViT is transferred, while the guiding LLM and projector are replaced, allowing us to assess cross-LLM transferability. Our cross-LLM transfer experiments demonstrate the value of pretraining visual encoders for their autoregressive LLM interface while trying to preserve more fine-grained visual information. Code and the pretrained model are available at https://github.com/SkyCol/MedMLIP
Chinese Translation
多模态大语言模型(MLLMs)通常复用通过 CLIP 预训练的视觉编码器,尽管这些 ViT 的特征最终由自回归大语言模型(LLM)使用。我们将这种不匹配称为语义接口鸿沟(semantic-interface gap),并提出了 MedMLIP 框架,该框架在冻结的 LLM 上通过报告生成任务对视觉编码器进行预训练,同时采用局部关系蒸馏(Local Relational Distillation, LRD)来保留视觉图像块之间的关系,以避免视觉坍缩。我们在 IU-Xray 和 Open-PMC-300K 数据集上预训练 MedMLIP,并在 VQA-RAD 和 SLAKE 上评估所得到的编码器。实验中仅迁移 ViT,而替换作为引导的 LLM 和投影器,从而能够评估跨 LLM 的可迁移性。我们的跨 LLM 迁移实验表明,针对自回归 LLM 接口预训练视觉编码器并尽量保留更细粒度的视觉信息具有重要价值。代码与预训练模型可在 https://github.com/SkyCol/MedMLIP 获取。
cs.AI / 56 / 2609.23877

Explainable Recommendations at Scale: LLM Rationales for YouTube Music Artist Discovery

大规模可解释推荐:用于 YouTube 音乐艺术家发现的 LLM 推理依据
Liu, Xiao, Song, Yanwei, Ranganathan, Srivaths, Chen, Yuan, Feng, Zheyun, Steenburgh, Parker, Klingenhoefer, Jochen, Lasche, Nathan, Varady, Gergo, Steele, Tim
Abstract
Modern music streaming platforms face a persistent tradeoff: exploiting familiar content versus driving the exploration of novel items. While users frequently desire discovery, they hesitate to select unknown artists over proven favorites. Providing transparent, natural language rationales that explain why an unexplored item is recommended lowers this barrier. However, while Large Language Models (LLMs) excel at this nuanced explainability, their real-time deployment is severely bottlenecked by prohibitive inference costs and computational overhead. In this paper, we present an industry case study of a decoupled recommendation architecture that successfully scales exploration without compromising latency. Our system isolates LLM inference asynchronously offline, pre-computing personalized candidate pools of undiscovered artists alongside tailored rationales. Large-scale online A/B experiments validate our design. We demonstrate that combining LLM-backed recommendations with these explanatory rationales significantly reduces the trust barrier for new content, yielding statistically significant improvements in both user exploration and overall engagement on the discovery surfaces.
Chinese Translation
现代音乐流媒体平台面临一个持续的权衡问题:是利用用户熟悉的内容,还是推动用户探索新内容。尽管用户常常渴望发现新音乐,但他们在选择未知艺术家时往往犹豫不决,不愿放弃已被验证的心头好。提供透明的自然语言推理依据(rationale)来解释为何推荐某个未被探索的内容,可以降低这一门槛。然而,尽管大语言模型(LLM)擅长生成这种细致的可解释性内容,但其实时部署却因高昂的推理成本和计算开销而受到严重制约。在本文中,我们展示了一个行业案例研究,介绍了一种解耦的推荐架构,它在不牺牲延迟的前提下成功实现了对新内容探索的大规模扩展。我们的系统将 LLM 推理异步地在离线阶段隔离执行,预先计算个性化的未发现艺术家候选池,并生成与之配套的定制化推理依据。大规模在线 A/B 实验验证了我们的设计。我们证明,将基于 LLM 的推荐与这些解释性推理依据相结合,能够显著降低用户对新内容的信任门槛,在发现页面上为用户探索行为和整体参与度均带来具有统计显著性的提升。
cs.AI / 57 / 2609.23917

Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer

提升技能水平会招募冻结国际象棋Transformer中更深的注意力层
Litman, David
Abstract
Chess involves complex reasoning in a deterministic environment, which makes it a useful setting for studying the mechanisms of computation inside transformers. The Maia-3 chess transformer takes Elo, a measure of competitive chess skill, as an input to the pre-trained network, so we can vary the skill the network is conditioned on with no change to its weights. Here we investigate how turning this skill dial affects self-attention. Ablating every attention head at every Elo from 700 to 2500, we find 1) increasing skill pushes the causal center of mass of the computation deeper, monotonically, for every chess piece and move type we measured; 2) the depth migration is much greater for specific tactics, especially knight forks, than for other move types; 3) the migration consists of deeper heads getting recruited for more specialized computations while one shared shallow head keeps a roughly constant contribution. These results may shed light on how conditioning inputs redistribute computation in larger transformers.
Chinese Translation
国际象棋在确定性环境中涉及复杂的推理,这使其成为研究Transformer内部计算机制的一个有用场景。Maia-3国际象棋Transformer将Elo(一种衡量竞技棋力的指标)作为预训练网络的输入,因此我们可以在不改变网络权重的情况下改变网络所条件化的技能水平。本文研究了调节这一技能旋钮如何影响自注意力机制。通过在700到2500的每个Elo水平上消融每一个注意力头,我们发现:1)对于我们测量的每种棋子和走子类型,提升技能水平会使计算的因果重心单调地向更深层推移;2)对于特定战术(尤其是马的叉子攻击),这种深度迁移远大于其他走子类型;3)这种迁移表现为更深的注意力头被招募来执行更专门化的计算,而一个共享的浅层注意力头保持大致恒定的贡献。这些结果可能有助于揭示条件化输入如何在更大的Transformer中重新分配计算。
cs.AI / 58 / 2609.23945

Echo State Network (ESN) for Signal Recovery in RF-Impaired IBFD MIMO Systems

面向射频损伤IBFD MIMO系统信号恢复的回声状态网络(ESN)
Prisby, Conrad, Li, Siyao, Xu, Chengtao, Yang, Thomas
Abstract
In-band full-duplex (IBFD) multiple-input multiple-output (MIMO) systems enable simultaneous transmission and reception on the same frequency band, improving spectral efficiency for next-generation wireless networks. However, IBFD-MIMO systems are susceptible to self-interference (SI), which may overpower signals of interest (SOI). In this scenario, blind source separation (BSS) algorithms can be adopted to remove SI and perform joint sensing and communication (JSAC), but BSS algorithms mostly assume an idealized linear and quasi-stationary signal model, which does not hold under realistic radio frequency (RF) impairments, such as I/Q imbalance, carrier frequency offset (CFO), phase noise, and power amplifier nonlinearity. This paper proposes a two-stage echo state network (ESN)-based scheme that is superior to BSS under these realistic conditions. A frozen ESN is trained offline to characterize the static SI path, while an adaptive ESN, updated online via recursive least squares, tracks the time-varying SOI path using sparse pilot symbols. We evaluate the proposed scheme's SOI recovery performance and acquisition speed with different block sizes, comparing it against other recurrent neural networks (RNN), such as long short-term memory (LSTM) and gated recurrent unit (GRU). Simulation results show that the proposed approach outperforms BSS, LSTM, and GRU in both efficiency and SOI recovery, demonstrating the viability of ESNs for real-time, nonlinear self-interference cancellation in realistic IBFD MIMO systems.
Chinese Translation
同频带全双工(IBFD)多输入多输出(MIMO)系统能够在同一频段上同时进行传输和接收,从而提高下一代无线网络的频谱效率。然而,IBFD-MIMO系统容易受到自干扰(SI)的影响,自干扰可能淹没感兴趣信号(SOI)。在这种场景下,可以采用盲源分离(BSS)算法来消除自干扰并实现感知与通信一体化(JSAC),但BSS算法大多假设理想化的线性准平稳信号模型,而在实际射频(RF)损伤(如I/Q不平衡、载波频偏(CFO)、相位噪声和功率放大器非线性)条件下,该假设并不成立。本文提出了一种两阶段基于回声状态网络(ESN)的方案,在这些实际条件下优于BSS。其中,冻结的ESN通过离线训练来表征静态自干扰路径,而自适应ESN通过递归最小二乘法在线更新,利用稀疏导频符号跟踪时变的感兴趣信号路径。我们在不同块大小下评估了所提方案的SOI恢复性能与捕获速度,并与其他循环神经网络(RNN)进行了比较,如长短期记忆网络(LSTM)和门控循环单元(GRU)。仿真结果表明,所提方法在效率和SOI恢复方面均优于BSS、LSTM和GRU,证明了ESN在实际IBFD MIMO系统中实现实时非线性自干扰消除的可行性。
cs.AI / 59 / 2609.23953

Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control

能编辑文档的智能体:对照非智能体基线衡量智能体型PDF伪造能力
Ren, Simiao, Raj, Ankit, Duong, Tommy, Zhang, Yuxin, Ng, Dennis, Shen, Xingyu, Zewde, Kidus, Zhou, Yuchen, Tiangratanakul, Neo
Abstract
AI agents that carry a multi-step computer task through on their own became ordinary tools in the past year, and the same autonomy is available to anyone whose task is harmful. We ask what that means for a relying party -- an insurer, a lender, an auditor -- whose evidence is a filed PDF. AgentForge-Bench measures how reliably an off-the-shelf coding agent, driving one of seven open-weight models with a shell and the stock Python PDF stack, alters one dollar amount, date or address in a real filed financial document from a single sentence of intent, graded by rules rather than by a model. Across 1,750 cells, 1,419 (81.1%) satisfy the verifier, and 808 (46.2%) also survive every stricter filter: visible, localized, typeface-matched, original value gone document-wide. A deterministic script with no model in it solves 98 of the 125 documents; the agents solve 124, and none the script solves alone. Agents misreport 41% of their wrong edits as done, no model refused, and the cheapest verified forgery costs 2.4 cents. The raw rate overstates the threat by about a factor of two; the strict rate is still large.
Chinese Translation
能够自主执行多步计算机任务的AI智能体在过去一年已成为普通工具,而这种自主性同样可以被任务有害的人所利用。我们要探究这对依赖方——如保险公司、贷款机构、审计师——意味着什么,因为他们的证据正是一份提交的PDF文件。AgentForge-Bench衡量了一个现成的编码智能体在多大程度上能够可靠地完成如下任务:借助shell和标准Python PDF工具库驱动七个开放权重模型之一,仅凭一句意图描述,就修改真实提交的金融文档中的某个金额、日期或地址,并且以规则而非模型进行评分。在1,750个测试单元中,1,419个(81.1%)通过了验证器,其中808个(46.2%)还通过了所有更严格的筛选:修改可见、局部化、字体匹配、且原值在全文范围内被彻底清除。一个不含任何模型的确定性脚本解决了125份文档中的98份;而智能体解决了124份,且没有一份是仅靠脚本能解决的。智能体将41%的失败修改误报为已完成,没有任何模型拒绝执行,而一次最便宜的、通过验证的伪造仅花费2.4美分。原始成功率大约将威胁高估了一倍;但严格标准下的成功率仍然相当高。
cs.AI / 60 / 2609.23957

Divergent strategies and convergent outcomes in autonomous materials discovery

自主材料发现中的策略分化与结果趋同
Kim, Jihan
Abstract
Scientific agents are mostly evaluated on whether they complete tasks or recover known results; we instead study variation across repeated open-ended campaigns. Sixteen separately initialized sessions of one model-harness configuration received a frozen database of 12,499 metal-organic frameworks, a methane-storage objective, a pinned protocol and a one-week budget. Strategies diverged into four approaches spanning 100--5,000 screened structures, and eight built 2,253 hypothetical structures. Yet the agents recovered the same materials frontier near 200 cm^3/cm^3, and an independent calculation of the database's porous region found its nine best structures all among their reports. Enforced checks on half the agents raised fresh-run reproduction from one of eight to eight of eight but could not detectably improve conclusion validity, because fifteen of sixteen agents selected the same audit-excluded entry, an incomplete structure whose missing anions created artificial pore volume. Replicated agents thus reveal both robust conclusions and common-mode errors from shared inputs.
Chinese Translation
科学智能体(Scientific Agents)大多以能否完成任务或复现已知结果来评估;我们则研究重复开展开放式科研活动时产生的变异性。同一模型-框架配置的十六个独立初始化会话,各自获得一个包含12,499种金属有机框架(MOF)的冻结数据库、一个甲烷存储目标、一个固定的实验协议以及为期一周的预算。各智能体的策略分化为四种方法,筛选结构数量从100到5,000不等,其中八个智能体构建了2,253个假想结构。然而,这些智能体都收敛到了相同的材料性能前沿(约200 cm^3/cm^3),且我们对数据库多孔区域的独立计算发现,其九个最优结构全部包含在它们的报告之中。对半数智能体实施的强制检查,将新运行的可复现率从八分之一提升至八分之八,但未能显著提高结论的有效性,因为十六个智能体中有十五个都选中了同一条被审计排除的记录——一个缺失阴离子的不完整结构,其缺失导致了虚假的孔隙体积。因此,重复运行的智能体既揭示了稳健的结论,也暴露了源自共享输入的共模错误。
cs.AI / 61 / 2609.23971

UniK: Universal Knowledge Perception for Digital and Physical AI

UniK:面向数字AI与物理AI的通用知识感知
Desai, Nirmit, Sawarkar, Kunal, Mahakali, Aditya, Lee, Dongkon, Park, Kevin, Song, Eric
Abstract
Two transformative classes of AI systems are reshaping how organizations operate: \textit{digital AI}, which reasons over enterprise knowledge to power chatbots and agent workflows; and \textit{physical AI}, which learns to control robots and autonomous systems from video, gameplay, and sensor telemetry. Both face the same foundational bottleneck: raw knowledge at scale, spanning heterogeneous modalities, locked in private corpora that existing AI infrastructure cannot access reliably or efficiently. We propose \textit{Universal Knowledge Perception (UniK)} as a common platform for both classes, covering the full knowledge lifecycle (ingestion, enrichment, indexing, retrieval, and continuous evaluation) across modalities from rich text and video to molecular data and sensor telemetry. We present UniK, built on Polymath Retrieval (multi-index fusion over automatically enriched indices) with no task-specific fine-tuning. Across five digital AI domains (medical literature, open-domain QA, chemistry, legal video proceedings, and government open data) UniK combined with an open-source 70-billion-parameter model consistently matches or outperforms frontier proprietary LLMs that are orders of magnitude larger: 76\% RAG accuracy on government data versus 47\% for GPT-5; 77.9\% on medical QA without fine-tuning; topping all open-source chemistry pipelines. We show that the same infrastructure directly addresses the data curation, indexing, and retrieval challenges facing physical AI world model training, where the knowledge problem is harder but structurally identical.
Chinese Translation
两类变革性的AI系统正在重塑组织的运作方式:数字AI(digital AI),通过对企业知识进行推理来驱动聊天机器人和智能体工作流;以及物理AI(physical AI),从视频、游戏画面和传感器遥测数据中学习控制机器人与自主系统。两者面临相同的基础性瓶颈:大规模的原始知识跨越异构模态,被锁定在现有AI基础设施无法可靠、高效访问的私有语料库中。我们提出通用知识感知(Universal Knowledge Perception,UniK)作为服务这两类AI系统的统一平台,覆盖从富文本、视频到分子数据和传感器遥测等多模态知识的完整生命周期(摄取、增强、索引、检索和持续评估)。我们构建了UniK,其核心是Polymath检索(在自动增强的索引上进行多索引融合),且无需针对特定任务进行微调。在五个数字AI领域(医学文献、开放域问答、化学、法律视频庭审和政府开放数据)上,UniK与开源700亿参数模型结合,始终能够比肩或超越规模大几个数量级的前沿专有大语言模型:在政府数据上RAG准确率达76%(GPT-5为47%);在医学问答上未经微调即达77.9%;在所有开源化学流水线中名列第一。我们进一步表明,同一基础设施可直接应对物理AI世界模型训练所面临的数据整理、索引和检索挑战——该领域的知识问题更为困难,但在结构上是相同的。
cs.AI / 62 / 2609.23974

LEAP-NBV: Lightweight Edge Active-Perception for Foundation-Model Next-Best-View Planning

LEAP-NBV:面向基础模型下一最佳视角规划的轻量级边缘主动感知
Hu, Boxun, Ge, Jiawei, Krieger, Axel, Wang, Peng, Mohsenin, Tinoosh
Abstract
Foundation models are endowing autonomous systems with greater intelligence, enabling a more comprehensive understanding of the environment through visual perception. A representative example is Human Mesh Recovery (HMR), which provides useful estimates of a target's 3D pose and shape that can benefit tactical missions. However, the size and power demands of such models make them difficult to run on edge platforms and limit their real-time performance, undermining the requirements of tactical edge deployment - especially for active perception, where a mobile robot must plan its next-best view on-board and cannot offload computation under contested communications. We present LEAP-NBV, a lightweight active-perception framework that runs foundation-model-driven Next-Best-View (NBV) planning on-board an edge device. To this end, we distill a family of large HMR teachers, each into a compact 32M student, with an offline mesh objective, then quantize the vision encoder to FP16 and characterize its on-device accuracy and latency. Within an occlusion-aware active perception loop, we evaluate all configurations on the same held-out benchmark and deploy the end-to-end pipeline on an NVIDIA Jetson Xavier NX, reporting measured on-device latency and energy. Distillation recovers 6-7 mm of Procrustes-aligned mean per-vertex position error (PA-MPVPE) over the undistilled student on the test set. Selecting the edge-optimal compression model brings the HMR engine to ~12 ms at a small accuracy cost and runs the full closed loop at 3.6 FPS and 2.6 J per frame, achieving a 2.0x speedup and 3.0x lower energy than the uncompressed model while nearly matching downstream task quality.
Chinese Translation
基础模型正赋予自主系统更高的智能,使其能够通过视觉感知更全面地理解环境。一个代表性的例子是人体网格恢复(Human Mesh Recovery, HMR),它能提供目标三维姿态和形状的有用估计,从而有助于战术任务。然而,此类模型的规模和功耗需求使其难以在边缘平台上运行,并限制了其实时性能,削弱了战术边缘部署的要求——尤其是对于主动感知而言,移动机器人必须在机载端规划其下一最佳视角(Next-Best-View, NBV),且在受干扰的通信条件下无法卸载计算。我们提出了LEAP-NBV,这是一个轻量级的主动感知框架,可在边缘设备上运行由基础模型驱动的下一最佳视角(NBV)规划。为此,我们将一系列大型HMR教师模型蒸馏为紧凑的32M参数学生模型(采用离线网格目标函数),随后将视觉编码器量化至FP16,并刻画其在设备上的精度和延迟。在一个考虑遮挡的主动感知循环中,我们在相同的保留测试基准上评估所有配置,并在NVIDIA Jetson Xavier NX上部署端到端流水线,报告实测的设备端延迟与能耗。在测试集上,蒸馏后的模型相比未蒸馏的学生模型,将Procrustes对齐的平均每顶点位置误差(PA-MPVPE)降低了6-7毫米。选择边缘最优的压缩模型后,HMR引擎以极小的精度代价达到约12毫秒的处理时间,整个闭环以3.6 FPS运行且每帧能耗为2.6焦耳,相比未压缩模型实现了2.0倍的加速和3.0倍的能耗降低,同时几乎匹配了下游任务质量。
cs.AI / 63 / 2609.23986

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Jev-Mem:面向高效AI智能体的系统一控制式智能体记忆
Jiang, Dongming, Li, Yi, Li, Bingzhe
Abstract
Agentic memory is becoming essential for long-horizon AI agents, yet many existing systems rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expensive generation on the critical path of memory operations. We introduce \textbf{\method}, a new agentic memory architecture inspired by System-One/System-Two cognition. System One captures fast, lightweight decision-making, whereas System Two performs slower, deliberative reasoning. Jev-Mem brings this division of labor to agentic memory through a dedicated System-One control plane, a structured multi-relational memory plane, and a System-Two reasoning plane. The System-One controller governs memory typing and relational organization during construction, and dynamically performs query routing, retrieval-budget allocation, graph traversal, candidate scoring, and adaptive stopping during retrieval. System Two is invoked only for complex reasoning and answer synthesis. This design improves both memory effectiveness and system efficiency: on LoCoMo Jev-Mem achieves an overall LLM-as-a-Judge score of 0.777, an 11.0\% relative improvement over the strongest baseline, while reducing memory construction time to 158\,s, a 6.6$\times$ speedup over the fastest competing memory system, and lowering average query latency to 0.93\,s, a 36.7\% reduction.
Chinese Translation
智能体记忆对于长时程AI智能体正变得至关重要,然而许多现有系统依赖自回归大语言模型(LLM)来控制记忆的组织、检索和使用,将昂贵的生成过程置于记忆操作的关键路径上。我们提出了 Jev-Mem,一种受系统一/系统二(System-One/System-Two)认知启发的全新智能体记忆架构。系统一负责快速、轻量的决策,而系统二执行较慢的审慎推理。Jev-Mem 通过专用的系统一控制平面、结构化的多关系记忆平面以及系统二推理平面,将这种分工引入智能体记忆。系统一控制器在记忆构建阶段负责记忆类型划分和关系组织,并在检索阶段动态执行查询路由、检索预算分配、图遍历、候选评分和自适应停止。系统二仅在复杂推理和答案合成时被调用。这一设计同时提升了记忆的有效性和系统效率:在 LoCoMo 数据集上,Jev-Mem 取得了 0.777 的总体 LLM-as-a-Judge 评分,相比最强基线相对提升 11.0%;同时将记忆构建时间降至 158 秒,相比最快的竞争记忆系统实现 6.6 倍加速;并将平均查询延迟降低至 0.93 秒,相对减少 36.7%。
cs.AI / 64 / 2609.23989

ACLArena: Agent Continue Learning in Multi-stage Post-training

ACLArena:多阶段后训练中的智能体持续学习
Wang, Haixin, Wang, Xiaoxuan, Zhang, Junkai, Zhang, Han, Sun, Renliang, Taylor, Alexander K, Shi, Yidan, Deng, Haoran, Wang, Chenguang, Cong, Jason, Sun, Yizhou, Wang, Wei
Abstract
Building general-purpose agents for industrial deployment requires integrating multiple capabilities, each typically acquired at a distinct stage of training. Yet there is currently no well-established recipe for Agent Continual Learning (ACL), with little understanding of the trade-offs among existing integration paradigms. To address this gap, we introduce ACLArena, a framework for comprehensively studying, analyzing, and evaluating ACL. We first build a sequential training pipeline and conduct an in-depth analysis that explains the mechanisms of forgetting and generalization from two complementary perspectives, the model level and the token level. Guided by these analyses, we systematically compare multi-teacher on-policy distillation, self-distilled fine-tuning, and model merging to assess their ability to recover previously learned capabilities while preserving newly acquired ones. Through extensive experiments, we develop a detailed understanding of how capabilities transfer across stages. Finally, we propose a new ACL recipe that combines offline replay over high-quality trajectories with a routed network of multiple LoRA experts each specialized via RL, substantially improving the agent's ability to learn across multiple domains. Comprehensive experiments on four reasoning and agentic tasks, evaluated under both in-domain and out-of-domain settings, demonstrate the value of our analysis and the effectiveness of our approach.
Chinese Translation
构建面向工业部署的通用智能体需要整合多种能力,而每种能力通常在训练的不同阶段获得。然而,目前尚无成熟的智能体持续学习(Agent Continual Learning,ACL)方案,对现有各种整合范式之间的权衡也缺乏深入理解。为填补这一空白,我们提出了ACLArena,一个用于全面研究、分析和评估ACL的框架。我们首先构建了一个顺序训练流水线,并从模型层面和词元层面这两个互补视角深入分析,阐释了遗忘与泛化的机制。基于这些分析,我们系统地比较多教师在线策略蒸馏、自蒸馏微调和模型合并三种方法,评估它们在恢复先前所学能力的同时保留新获能力的效果。通过大量实验,我们对能力在各阶段之间如何迁移建立了细致的认识。最后,我们提出了一种新的ACL方案,该方案将对高质量轨迹的离线回放与由多个经过强化学习专门化训练的LoRA专家组成的路由网络相结合,显著提升了智能体跨多个领域学习的能力。在四项推理与智能体任务上、在域内和域外两种设置下的全面实验,验证了我们分析的价值和方法的有效性。
cs.AI / 65 / 2609.24002

FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering

FinInteract:面向模糊金融问答中澄清与意图融合的基准测试
Wang, Xinyu, Kwok, Tung Sum Thomas, Tai, Zhenghan, Cheng, Guang
Abstract
Large language model agents increasingly answer financial questions by searching regulatory filings. Such questions are often deceptively under-specified: Meta Platforms' "operating income" is $46.75B consolidated but $62.87B for the Family of Apps segment, and each reading is exactly verifiable against the filing. A capable agent should recognize the ambiguity and ask, rather than commit to a plausible but unintended reading. Existing financial benchmarks cannot measure this, because one gold answer per question cannot separate agents that resolve the ambiguity from those that guess the common reading, a blind spot we call the single-gold illusion. We release FinInteract, a bilingual (English/Chinese) benchmark of 173 instances that pairs each question with a default and an intended interpretation across a five-category ambiguity taxonomy, and grades whether an agent elicits the right clarification and then integrates it. Re-grading identical outputs against the default rather than the intended reading inflates GPT-4o's accuracy by 3.1 times, confirming the illusion. Beyond it, we find that models answer above 90% once the interpretation is supplied but at most 28.9% when they must elicit it themselves, that targeting is uneven across a taxonomy well powered for entity scope and metric definition and exploratory elsewhere, and that conditioning on the ambiguity category improves resolution at both inference and training time.
Chinese Translation
大语言模型智能体日益通过检索监管申报文件来回答金融问题。然而,这类问题往往在表面上显得信息不足:Meta Platforms 的“营业收入”(operating income)按合并口径为 467.5 亿美元,而按 Family of Apps 业务分部口径则为 628.7 亿美元,且每一种理解都可以在申报文件中得到精确验证。一个有能力的智能体应当识别这种歧义并主动询问,而不是径直采纳某种貌似合理但并非提问者本意的解读。现有的金融基准无法衡量这一能力,因为每个问题仅有一个标准答案,无法区分那些解决了歧义的智能体和那些只是猜中常见解读的智能体——我们将这一盲点称为“单一标准答案错觉”(single-gold illusion)。我们发布了 FinInteract,这是一个双语(英文/中文)基准,包含 173 个实例,每个问题都配有一个默认解读和一个预期解读,并依据五类歧义分类体系进行划分,同时评估智能体是否能引出正确的澄清问题并将其融入回答。若对相同的输出按默认解读而非预期解读重新评分,GPT-4o 的准确率会被虚增 3.1 倍,从而印证了这一错觉。除此之外,我们还发现:当解读被直接提供时,模型回答准确率超过 90%,但当需要模型自行引出澄清时,准确率最高仅为 28.9%;在不同歧义类别上的表现并不均衡——该基准在实体范围(entity scope)和指标定义(metric definition)两类上统计功效充足,而在其他类别上尚属探索性;并且在推理和训练阶段引入歧义类别条件均能提升歧义解决能力。
cs.AI / 66 / 2609.24012

Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure

检验而非假定充分性:将生成式社会模拟器与涌现网络结构进行校准
Shao, Tengfei, Li, Chao, Wang, Xu, Goto, Masayuki
Abstract
Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check. We present an adequacy-aware calibration protocol that couples amortized posterior estimation with a synthetic identifiability assessment, a matched-sample-size adequacy check (prior-predictive reachability plus per-statistic posterior-predictive localization), a diagnosis-guided repair, and a statistic-held-out audit. We demonstrate it on a real second-hand luxury resale market with four channel-by-residency cells, each a bipartite buyer-brand network, using a forward model built from persona profiles elicited once, offline, by a language model. The behavioural parameters are recoverable in all four cells, though calibration is approximate and overconfident for one parameter. The observed summary falls outside the simulator's reachability reference in every cell, with the mean purchased tier as the pervasive discrepancy. The repair meets the value-block criterion in two of four cells but does not restore adequacy, and the held-out audit surfaces a buyer-breadth-dispersion miss no earlier diagnostic detected. A profile-source ablation finds the language-model profiles beat a flat rule baseline in all four cells, yet within-category brand relabelling causes no consistent degradation, so the profiles are a partially validated input whose value rests on structure, not brand identity. Making no causal claim, we conclude that an independent-aggregation account, without agent interaction or a buyer-breadth mechanism, cannot jointly reproduce the market's purchased-tier level, head-brand concentration, community structure and buyer-breadth heterogeneity.
Chinese Translation
生成式社会模拟器的验证往往止步于表面效度:对涌现网络结构仅作描述性比较,缺乏量化的参数不确定性或充分性检验。我们提出一种充分性感知的校准协议,将摊销后验估计与合成可辨识性评估、匹配样本量的充分性检验(先验预测可达性加上逐统计量的后验预测定位)、诊断引导的修复以及统计量留出审计相结合。我们在一个真实的二手奢侈品转售市场上演示该方法,该市场包含四个渠道—居住地单元格,每个单元格是一个二部买家—品牌网络,所用的前向模型基于由大语言模型一次性离线引出的角色画像构建。行为参数在所有四个单元格中均可恢复,尽管其中一个参数的校准是近似且过度自信的。观测到的摘要统计量在所有单元格中都落在模拟器的可达性参考之外,平均购买层级是普遍存在的差异项。该修复在四个单元格中的两个满足了价值阻塞准则,但并未恢复充分性;留出审计揭示了一种此前诊断均未发现的买家广度离散性偏差。画像来源消融实验发现,语言模型生成的画像在所有四个单元格中均优于平坦规则基线,但类别内品牌重标注并未造成一致性退化,因此该画像是一种部分验证的输入,其价值源于结构而非品牌身份。在不作因果断言的前提下,我们得出结论:一种不含智能体交互或买家广度机制的独立聚合解释,无法同时再现该市场的购买层级水平、头部品牌集中度、社区结构以及买家广度异质性。
cs.AI / 67 / 2609.24016

Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech

情境感知的人工智能系统部署前评估:尼日利亚金融科技的监管框架
Uduimoh, Andrew Anogie, Yusuf, Hadiza Umar, Osho, Oluwafemi
Abstract
Commercial large language models are increasingly deployed across African fintech infrastructure for fraud detection and customer communication, yet no Nigerian or African continental regulatory instrument specifies what pre-deployment evaluation such systems must undergo before procurement. This paper reviews African fintech AI governance across global, continental, and Nigerian instruments, and shows that safety is affirmed as a principle while pre-deployment evaluation is operationally unspecified. Generic safety benchmarks cannot surface the failure modes most relevant to this domain, since none contain Nigerian institutional content or test for false positive misclassification of legitimate financial communications. These claims are demonstrated using SafeAlert, a purpose-built evaluation kit applied to six commercial models across three system prompt conditions. Results show that models resisting generic harmful content requests still produce complete fraud scripts under specific framing, and that several models misclassify most legitimate Nigerian bank communications as suspicious or fraudulent, a failure invisible to standard safety evaluation. The paper concludes with a regulatory framework proposing pre-deployment evaluation requirements for the CBN, NITDA, SEC, and the AU, arguing that the identified gap reflects an absence of regulatory specification, not a shortage of technical or financial resources.
Chinese Translation
商业大语言模型正日益被部署于非洲金融科技基础设施中,用于欺诈检测和客户沟通,然而目前尚无尼日利亚或非洲大陆层面的监管文件明确规定此类系统在采购前必须进行何种部署前评估。本文梳理了全球、非洲大陆及尼日利亚各层级监管文件中的非洲金融科技AI治理现状,并指出:虽然安全性被确立为一项原则,但部署前评估在操作层面仍处于未明确状态。通用安全基准无法揭示与该领域最相关的失效模式,因为这些基准既不包含尼日利亚机构相关内容,也不测试对合法金融通信的误报性误分类。本文使用SafeAlert——一个专门构建的评估套件——对六个商业模型在三种系统提示词条件下进行测试,验证了上述论断。结果表明:能够抵御通用有害内容请求的模型,在特定表述下仍会生成完整的欺诈脚本;此外,多个模型将大多数合法的尼日利亚银行通信误分类为可疑或欺诈性内容——这一失效在标准安全评估中是无法被发现的。本文最后提出了一个监管框架,为尼日利亚央行(CBN)、国家信息技术发展局(NITDA)、证券交易委员会(SEC)以及非洲联盟(AU)提出部署前评估要求,并论证所识别的差距反映的是监管规范的缺失,而非技术或财力资源的不足。
cs.AI / 68 / 2609.24025

Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search

基于程序化策略搜索的连续博弈反应式角色行为合成
Gumin, Maxim, Liu, Hsueh-Ti Derek, Zordan, Victor, Ritchie, Daniel
Abstract
We present a method for synthesizing reactive character behaviors for continuous games as compact, human-readable programs. Game AI practice still relies heavily on manually authored behavior trees, state machines, and scripts, while academic reinforcement learning typically produces opaque neural controllers that are expensive to train and difficult to edit. Our approach bridges this gap by searching directly over a domain-specific language for continuous-space game policies. The language is designed around reactive geometric decisions and includes higher-order constructs such as direction maximization. These constructs help discretize a continuous behavior space into enumerable program structures. To make program search practical, we introduce a large set of synthesis antipatterns that remove redundant program forms while preserving behavioral coverage. We further combine bottom-up symbolic enumeration with top-down guidance from a coding agent. Our resulting method, agentic sketching, has the agent propose high-level policy structure and call an enumerator to complete local program slots. We evaluate the method on a benchmark of 14 continuous games, ranging from classic control tasks to multi-agent football. We find that pure enumeration is often more efficient than using a coding agent alone, while the combined method substantially outperforms both. Our results suggest that programmatic policy search can be a practical authoring tool for game AI: designers specify reward functions, and the system discovers editable behaviors that are effective, portable, and often surprising.
Chinese Translation
我们提出了一种将连续博弈中反应式角色行为合成为紧凑、人类可读程序的方法。游戏AI实践中仍然严重依赖人工编写的行为树、状态机和脚本,而学术界的强化学习通常生成难以解释的神经控制器,其训练成本高且难以编辑。我们的方法通过直接在一种面向连续空间博弈策略的领域特定语言上进行搜索来弥合这一差距。该语言围绕反应式几何决策设计,并包含诸如方向最大化等高阶构造。这些构造有助于将连续的行为空间离散化为可枚举的程序结构。为使程序搜索切实可行,我们引入了大量合成反模式,在保持行为覆盖范围的同时去除冗余的程序形式。我们进一步将自底向上的符号枚举与来自编程智能体的自顶向下引导相结合。我们由此得到的方法——智能体化草图生成,由智能体提出高层策略结构,并调用枚举器来完成局部程序槽位。我们在包含14个连续博弈的基准上对该方法进行了评估,涵盖从经典控制任务到多智能体足球等多个场景。我们发现,纯枚举往往比单独使用编程智能体更高效,而两者结合的方法则显著优于二者。我们的结果表明,程序化策略搜索可以成为一种实用的游戏AI创作工具:设计者只需指定奖励函数,系统便能发现有效、可移植且常常出人意料的可编辑行为。
cs.AI / 69 / 2609.24036

Structured Decomposition for Reliable LLM-Generated Access Control Policies

面向可靠的LLM生成访问控制策略的结构化分解方法
Gupta, Vatsal, Sreenivasamurthy, Darshan
Abstract
This paper presents an LLM-based system that translates natural-language access control policies (NLACPs) into executable Rego code for Open Policy Agent (OPA). It provides a modular, end-to-end pipeline for policy detection, component extraction, schema validation, linting, compilation, and automated test generation and execution. The system is designed to bridge the gap between human-readable access requirements and machine-enforceable policy-as-code (PaC), with a focus on deployment reliability and security correctness. We evaluate the system on 372 ACRE-complete access control statements with non-null subject, action, and resource annotations against a direct single-prompt LLM baseline to isolate the contribution of structured decomposition and schema-aware validation. The system achieves a 50.3% end-to-end policy correctness rate, compared with 15.3% for the baseline, representing a 3.3x improvement. A policy is counted as correct only if it satisfies compilation, linting, and both positive and negative tests, making this a strict measure of deployable correctness. On security-critical patterns, the system generates correct deny semantics for 87.5% of deny policies (baseline: 37.5%), ownership conditions for 100% of ownership-qualified policies (baseline: 40%), and status-qualified conditions for 100% of status-qualified policies (baseline: 55.6%). These results indicate that structured decomposition and schema-aware validation play a critical role in improving the reliability of LLM-generated authorization policies.
Chinese Translation
本文提出一个基于大语言模型(LLM)的系统,用于将自然语言访问控制策略(NLACPs)转换为可用于Open Policy Agent(OPA)的可执行Rego代码。该系统提供了一个模块化的端到端流水线,涵盖策略检测、组件提取、模式(schema)验证、代码检查(linting)、编译以及自动化测试生成与执行。该系统旨在弥合人类可读的访问需求与机器可执行的策略即代码(policy-as-code, PaC)之间的差距,重点关注部署可靠性与安全正确性。我们在372条具有非空主体、动作和资源标注的ACRE-complete访问控制语句上对该系统进行评估,并与直接单提示词(single-prompt)LLM基线进行对比,以分离结构化分解与模式感知验证的贡献。该系统实现了50.3%的端到端策略正确率,而基线仅为15.3%,提升了3.3倍。只有同时满足编译、代码检查以及正向和负向测试的策略才被计为正确,因此这是一项对可部署正确性的严格度量。在安全关键模式上,该系统为87.5%的拒绝(deny)策略生成了正确的拒绝语义(基线:37.5%),为100%的所有权限定策略生成了正确的所有权条件(基线:40%),为100%的状态限定策略生成了正确的状态条件(基线:55.6%)。这些结果表明,结构化分解与模式感知验证在提高LLM生成的授权策略可靠性方面发挥着关键作用。
cs.AI / 70 / 2609.24057

Representation-guided in-context learning for medical image interpretation with multimodal large language models

基于表征引导的上下文学习:利用多模态大语言模型进行医学图像解读
Zhao, Minda, Hu, Fangyu, Luo, Yan, Yang, Yutong, Cai, Jiahui, Zhou, Kaichen, Li, Manling, Liang, Paul, Du, Yilun, Shen, Lucy Q., Wang, Mengyu
Abstract
Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires resource-intensive domain-specific fine-tuning. Here we introduce representation-guided in-context learning (RG-ICL), a training-free inference framework that retrieves query-aligned demonstrations using frozen encoders, without task-specific parameter updates. Across eight datasets spanning histopathology, radiology and retinal fundoscopy, RG-ICL improved classification (mean gain 20 percentage points) and visual question answering (VQA) (mean gain 13 percentage points) over no-context and conventional ICL, approaching or exceeding training-based comparators. Which cases were retrieved mattered more than how many: 6 query-aligned cases outperformed up to 32 randomly selected ones, whereas fixed or random cases often reduced accuracy below baseline. For VQA, aligning reference cases with both image content and question intent produced further gains. These findings indicate that for medical image interpretation, curating which reference cases an MLLM sees is a practical alternative to retraining it.
Chinese Translation
医学图像解读对诊断和诊疗至关重要,然而对通用多模态大语言模型(MLLM)进行适配往往需要资源密集型的领域专用微调。本文提出表征引导的上下文学习(RG-ICL),这是一种无需训练的推理框架,利用冻结编码器检索与查询对齐的示例,无需针对特定任务进行参数更新。在涵盖组织病理学、放射学和视网膜眼底检查的八个数据集上,RG-ICL 相较于无上下文和常规上下文学习(ICL)提升了分类性能(平均提升 20 个百分点)和视觉问答(VQA)性能(平均提升 13 个百分点),接近甚至超越了基于训练的对比方法。检索哪些案例比检索多少案例更为重要:6 个与查询对齐的案例的表现优于多达 32 个随机选取的案例,而固定或随机选取的案例往往使准确率降至基线以下。对于 VQA,将参考案例同时与图像内容和问题意图对齐可带来进一步提升。这些发现表明,在医学图像解读任务中,精心筛选 MLLM 所看到的参考案例是替代重新训练的一种实用方案。
cs.AI / 71 / 2609.24090

Incremental Consistency Execution for Autonomous Intelligent Systems

面向自主智能系统的增量一致性执行方法
Li, Cheng, Liu, Jiexiong, Chen, Yixuan, Huang, Ziheng
Abstract
Long-horizon autonomous intelligent systems rely on heterogeneous components such as large language models, databases, external APIs, and rule engines, while their external states continuously change during execution. Re-executing the entire workflow after every change introduces substantial redundant computation. This paper proposes an incremental consistency execution method based on task fact contracts, field-level dependency masks, and state perturbation result invariant domains. After an initial verified execution, the system constructs conservative invariant domains for critical inputs and uses them to determine whether downstream results can be safely renewed without re-invoking expensive components. When re-execution is required, only the smallest affected output fields are recomputed, and an equivalence barrier prevents unnecessary downstream propagation. A submission-time version consistency gate further ensures the safety of side-effecting actions. Experiments on industrial fault diagnosis, enterprise analytics, and LLM-based multi-tool assistants show that the proposed method significantly reduces expensive component calls and end-to-end latency while maintaining high consistency and low incorrect-reuse rates.
Chinese Translation
长时程自主智能系统依赖于大型语言模型、数据库、外部API和规则引擎等异构组件,同时其外部状态在执行过程中持续变化。每次变化后重新执行整个工作流会引入大量冗余计算。本文提出一种基于任务事实契约、字段级依赖掩码和状态扰动结果不变域的增量一致性执行方法。在初始经验证的执行之后,系统为关键输入构建保守的不变域,并利用其判断是否可以在无需重新调用昂贵组件的情况下安全地更新下游结果。当需要重新执行时,仅重新计算受影响最小的输出字段,并通过等价屏障阻止不必要的下游传播。提交时的版本一致性门控进一步确保了具有副作用的操作的安全性。在工业故障诊断、企业分析以及基于大语言模型(LLM)的多工具助手上的实验表明,所提方法在保持高一致性和低错误复用率的同时,显著减少了昂贵组件的调用次数和端到端延迟。
cs.AI / 72 / 2609.24092

DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents

DocMIDE:在视觉丰富文档中学习多跳隐式推导
Wang, Jeremy Cerwin, Wong, Wai Kit, Tang, Jeff Kai Tai
Abstract
Real-world document processing systems rely on rigid, predefined schemas, yet critical target fields often lack direct visual counterparts on the page. Extracting these implicit values requires multi-hop derivation, such as aggregating sub-categories or reasoning over visual marks. While existing methods handle explicit text spans or simple implicit queries, they fail at multi-hop visual reasoning even after standard fine-tuning: models retrieve incorrect visual evidence, or retrieve it correctly and then skip the intermediate steps of the derivation. To address this, we introduce DocMIDE, a fine-tuning framework that trains compact vision-language models to retrieve visual evidence explicitly before deriving an answer. DocMIDE constrains generation to a plan-retrieve-derive structure and optimizes it with Group Relative Policy Optimization under a four-component, rule-based reward that scores output format, the retrieved evidence block, every intermediate derivation step, and the final value against a verified reference trace. On a 4,151-pair implicit extraction benchmark, DocMIDE raises accuracy from 70.8% to 95.9% on Qwen3.5-4B from only a small set of annotated examples, and transfers to a second backbone architecture. Supervised demonstrations alone do not close this gap at any budget we tested; rewarding the intermediate steps is what does.
Chinese Translation
现实世界的文档处理系统依赖于僵化的预定义模式,然而关键的目标字段往往在页面上缺乏直接的视觉对应内容。提取这些隐式取值需要进行多跳推导,例如对子类别进行聚合或对视觉标记进行推理。虽然现有方法能够处理显式文本片段或简单的隐式查询,但即使经过标准微调,它们在多跳视觉推理上仍然失败:模型要么检索到错误的视觉证据,要么正确检索后跳过推导的中间步骤。为解决这一问题,我们提出了 DocMIDE,这是一个微调框架,用于训练紧凑的视觉-语言模型在推导答案之前显式地检索视觉证据。DocMIDE 将生成过程约束为“规划-检索-推导”的结构,并通过组相对策略优化(Group Relative Policy Optimization)对其进行优化,采用由四个组件构成的基于规则的奖励机制,依据经过验证的参考轨迹对输出格式、检索到的证据块、每一个中间推导步骤以及最终取值进行评分。在一个包含 4,151 对样本的隐式抽取基准上,DocMIDE 仅使用少量标注示例,便将 Qwen3.5-4B 的准确率从 70.8% 提升至 95.9%,并能迁移到第二种骨干架构。在我们测试的任何数据预算下,仅依靠监督式示范都无法弥合这一差距;正是对中间步骤的奖励设计实现了这一效果。
cs.AI / 73 / 2609.24101

When More Evidence Hurts: Publication-Bias Drift and Principled Stopping for Biomedical Causal Search

当更多证据带来危害:发表偏倚漂移与生物医学因果搜索的原则性停止
Sun, Fred, Guo, Shangqi
Abstract
Automated biomedical evidence synthesis depends on retrieving published studies, but the biomedical literature is systematically skewed toward positive findings. Deeper retrieval can therefore make a system \emph{more} likely to falsely infer benefit when the true effect is null. We formalise this phenomenon as \emph{evidence drift} and prove that, under a standard publication-bias model, the false-positive probability on null-effect queries follows a strictly increasing large-sample envelope in retrieval depth, approaching one. Empirically, on a held-out test set of 140 Cochrane-derived queries, drift rises monotonically from 7.9\% to 15.7\% as the retrieval budget grows from 3 to 20 steps, and concentrates in the null-effect class. We present DACG-agent, a drift-aware causal-graph agent that incrementally builds a causal knowledge graph from PubMed abstracts and applies a two-layer stopping policy with complementary roles: a KL-divergence monitor that detects posterior convergence (the accuracy layer), and a Bradley--Terry process reward model (PRM) whose online decline detection halts retrieval once evidence quality peaks (the efficiency layer). Against full-budget retrieval, DACG-agent reduces evidence drift from 15.7\% to 6.4\% and improves null-effect accuracy by 21 percentage points (40.0\%$\to$61.4\%) while using 67\% fewer retrieval steps; overall accuracy rises from 61.4\% to 69.3\% (95\% CI 61--77). A simulation confirms the drift result transfers from the analysed vote-counting aggregator to the deployed noisy-OR one.
Chinese Translation
自动化生物医学证据综合依赖于对已发表研究的检索,但生物医学文献系统性地偏向阳性结果。因此,当真实效应为零时,更深度的检索反而可能使系统更容易错误地推断存在获益。我们将这一现象形式化为“证据漂移”(evidence drift),并证明在标准的发表偏倚模型下,零效应查询上的假阳性概率随检索深度遵循一个严格递增的大样本包络,趋于1。实证结果表明,在一个由140个Cochrane来源查询构成的保留测试集上,随着检索预算从3步增加到20步,漂移率从7.9%单调上升至15.7%,且集中于零效应类别。我们提出DACG-agent,一种漂移感知的因果图智能体,它从PubMed摘要中增量构建因果知识图谱,并采用具有互补作用的双层停止策略:一个检测后验收敛的KL散度监测器(准确性层),以及一个Bradley–Terry过程奖励模型(PRM),其在线下降检测在证据质量达到峰值时即停止检索(效率层)。与全额预算检索相比,DACG-agent将证据漂移从15.7%降至6.4%,将零效应准确率提升21个百分点(40.0%→61.4%),同时使用的检索步骤减少67%;总体准确率从61.4%提升至69.3%(95% CI 61–77)。仿真实验证实,漂移结果可从所分析的计票聚合器迁移到实际部署的noisy-OR聚合器。
cs.AI / 74 / 2609.24115

EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation

EDGEGEN:利用合成边界用例生成提升工具调用智能体超越理想路径的表现
Abichandani, Harshavardhan, Chong, Penny, Shen, Jiyuan, Singh, Gunraj, Hathidara, Ashutosh, Yu, Marcus Duigan Xing, Lo, Jane, Ghosh, Atin, Li, Yipeng, Dahlmeier, Daniel
Abstract
Tool-calling LLM agents are increasingly deployed in enterprise applications. However, effective evaluation and optimization require high-quality, diverse task datasets that are often difficult to obtain due to privacy and other constraints. Existing synthetic task generation methods often produce generic tasks that ignore an agent's underlying state or database and fail to reflect real-world usage diversity. We propose EdgeGen, a synthetic task generation framework that extracts compliance rules from an agent's specification and uses them to generate database-grounded edge-case tasks designed to violate these rules. When combined with existing synthetic data generation techniques, EdgeGen enables agent improvement through finetuning and harness optimization. The resulting pipeline forms a fully automated closed-loop system that requires no human annotation. Finetuning on data generated by EdgeGen yields a consistent mean progress improvement of 2 percent to 42 percent on tau2bench airline domain, while other baseline methods show degradation for some models. On the other hand, for harness optimization, our method shows a mean progress improvement of 10 percent and 30 percent over the human-curated and base harnesses, respectively, for the Gemma-4-e4b model.
Chinese Translation
工具调用型大语言模型(LLM)智能体在企业应用中的部署日益增多。然而,有效的评估与优化需要高质量、多样化的任务数据集,而由于隐私等限制条件,这类数据集往往难以获取。现有的合成任务生成方法通常产生通用任务,忽略了智能体底层的状态或数据库,也无法反映真实世界使用的多样性。我们提出了 EdgeGen,一种合成任务生成框架,它从智能体的规范说明中提取合规规则,并利用这些规则生成旨在违反这些规则的、以数据库为依据的边界用例任务。与现有合成数据生成技术相结合后,EdgeGen 能够通过微调和测试框架(harness)优化来改进智能体。由此形成的流水线构成了一个完全自动化的闭环系统,无需任何人工标注。在 EdgeGen 生成数据上进行微调,在 tau2bench 航空领域上平均进度提升达 2% 至 42%,而其他基线方法在某些模型上反而出现性能下降。另一方面,在测试框架优化方面,对于 Gemma-4-e4b 模型,我们的方法相对于人工策划的测试框架和基础测试框架分别实现了 10% 和 30% 的平均进度提升。
cs.AI / 75 / 2609.24130

Self-Healing Harness for Runtime Oversight of Agent Self-Modification

用于智能体自我修改运行时监管的自愈框架(Self-Healing Harness)
Tayebati, Sina, Kumar, Divake, Darabi, Nastaran, Krishnan, Ranganath, Trivedi, Amit Ranjan
Abstract
LLM agents can change their own future behavior, raising a basic control question of which self-generated changes should be allowed to persist. We formulate this as admission control for self-modification. The agent may propose changes to its operating instructions, while an external runtime gate controls persistence. We implement this principle as a model-agnostic self-healing harness that runs a Detect, Notice, Heal, Validate loop around an otherwise unmodified agent. The agent authors candidate behavioral rules in an external workspace, where they receive provisional execution authority during evaluation and acquire persistent cross-episode authority only after measured improvement on the triggering failure without regression beyond a fixed margin on protected cases. Replay provides matched evidence when available, forward trials provide a weaker fallback, and a corpus-level guard re-tests the accumulated active rule set. Across 16 matched Baseline and Harness runs spanning AppWorld, Terminal-Bench, and $\tau^2$-Bench, the gate rejected 383 replay-decided proposals. Of these, 211 (55%) improved their triggering failure while degrading a case that previously worked. This shows that locally beneficial self-modifications can introduce collateral regressions often enough to materially affect gate decisions, providing direct empirical motivation for external admission control. Task-completion score is higher under the Harness in all 16 pairs, with two paired bootstrap intervals excluding zero, while repeated-trial reliability is higher in 12 pairs, tied in 4, and lower in none. Because adaptation modifies the policy-inducing context while leaving model weights fixed, admitted changes remain inspectable, reversible, and compatible with closed-weight models.
Chinese Translation
LLM 智能体能够改变自身未来的行为,这引出一个基本的控制问题:哪些自我生成的更改应当被允许持续存在。我们将此问题形式化为自我修改的准入控制(admission control)。智能体可以对其运行指令提出修改建议,而由一个外部运行时门控(runtime gate)来控制修改的持久化。我们将该原则实现为一个与模型无关的自愈框架(self-healing harness),它在未被修改的智能体周围运行一个“检测(Detect)—察觉(Notice)—修复(Heal)—验证(Validate)”循环。智能体在外部工作区中撰写候选行为规则,这些规则在评估期间获得临时执行权限,并且只有在触发性故障上取得可测量的改进、同时在受保护用例上的性能回退不超过固定阈值时,才能获得跨回合的持久权限。在可用时,重放(replay)提供匹配的证据;前向试验(forward trials)作为较弱的替代方案;而语料库级防护(corpus-level guard)会对累积的活动规则集进行重新测试。在跨越 AppWorld、Terminal-Bench 和 $\tau^2$-Bench 的 16 组匹配的基线(Baseline)与框架(Harness)运行中,门控拒绝了 383 个由重放裁决的提案。其中,211 个(55%)在改善其触发性故障的同时,使某个此前正常运行的用例发生了退化。这表明,局部有益的自我修改可能以足够高的频率引入附带性回退,从而实质性影响门控决策,为外部准入控制提供了直接的实证动机。在全部 16 组配对实验中,使用框架的任务完成分数更高,其中两组配对自助法(bootstrap)置信区间排除了零;而重复试验的可靠性在 12 组中更高,4 组持平,无一更低。由于这种适应仅修改诱导策略的上下文而保持模型权重不变,被准入的更改仍然可检查、可回滚,并且与闭源权重模型兼容。
cs.AI / 76 / 2609.24165

APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction

APEXA:面向同步辐射数据还原多智能体LLM自动化的执行完整性保障机制
Tripathi, Pawan K., Sharma, Hemant, Chuang, Andrew, Cherukara, Mathew J.
Abstract
Synchrotron data reduction, detector calibration followed by azimuthal integration of terabyte-scale diffraction series, is a multi-step, expert-bound bottleneck that increasingly limits the science rate of user facilities. LLM agents promise to collapse it, but driving a real pipeline with a stochastic model creates a failure mode chat benchmarks cannot see: an agent can report a calibration that was never computed. Correctness here is a property of what executed, not of the transcript. We present APEXA, a deployed multi-agent framework (61 tools over heterogeneous compute, run as a single reasoning loop) automating calibration and integration from natural language at a major light source. We make three contributions. First, execution-integrity enforcement: a deterministic tool-layer guard that refuses to surface any result not backed by an executed tool call, with a parser tolerant of cross-model tool-call format drift: in deployment, a frontier model fabricated a complete calibration-comparison report for commands that never ran, which the guard converts to an explicit non-result; the same code gates an optional motor-control surface at 0/200 adversarial violations against a simulated IOC, versus 15/200 for an equivalent safety prompt. Second, we release APEXA-Bench, an evaluation harness of 58 facility tasks (50 base plus an 8-task cross-detector slice) organized by a four-class physical-consequence taxonomy, the first benchmark axis we know of separating a wasted compute cycle from a damaged instrument; its cross-detector grading against NIST-traceable lattice constants surfaced two latent pipeline bugs. Large-scale agent scoring is left to a full-length study. Third, we validate APEXA on real beamline data: from one natural-language prompt it recovers detector geometry and integrates a full attenuation/exposure sweep. We release the framework, harness and traces.
Chinese Translation
同步辐射数据还原——包括探测器校准以及对太字节级衍射序列的方位角积分——是一个多步骤、依赖专家的瓶颈环节,日益制约用户设施的科学产出速率。LLM智能体有望突破这一瓶颈,但用随机性模型驱动真实流水线会产生一种聊天基准测试无法察觉的失效模式:智能体可能报告一个从未实际计算过的校准结果。此处的正确性取决于实际执行了什么,而非对话记录呈现了什么。我们提出APEXA,一个已部署的多智能体框架(在异构计算资源上提供61个工具,以单一推理循环方式运行),可在大型光源设施中通过自然语言自动完成校准与积分。我们的贡献有三方面。第一,执行完整性保障:一个确定性的工具层守卫,拒绝呈现任何没有已执行工具调用支撑的结果,并配备对跨模型工具调用格式漂移具有容错能力的解析器。在部署中,一个前沿模型为从未执行的命令伪造了一份完整的校准对比报告,该守卫将其转化为显式的无效结果;同一套代码在0/200次对抗性违规的情况下守卫了可选的电机控制接口(针对模拟IOC),而同等的安全提示词则为15/200次违规。第二,我们发布APEXA-Bench,一个包含58项设施任务(50项基础任务外加8项跨探测器任务)的评估框架,按照四类物理后果分类体系组织,这是我们所知的首个能够区分浪费的计算周期与损坏的仪器设备的基准评估维度;其基于NIST可溯源晶格常数的跨探测器评分暴露了两个潜在的流水线缺陷。大规模智能体评分留待后续完整研究开展。第三,我们在真实光束线数据上验证了APEXA:仅凭一条自然语言提示,它即可恢复探测器几何参数并完成完整的衰减/曝光扫描积分。我们公开了该框架、评估工具及运行轨迹。
cs.AI / 77 / 2609.24174

CREDO: Variance-Guided Rubric Evolution for Replay-Corrected Credit Assignment

CREDO:基于方差引导的评分标准演化与重放校正的信用分配方法
Hu, Xuchun
Abstract
Long-horizon language agents receive sparse terminal feedback, while intermediate rubrics provide structured but potentially misspecified assessments of progress. In resettable training environments, counterfactual continuation rollouts can measure local credit, but exhaustive replay is costly. We propose Credo, a framework that couples evolving semantic rubrics with selective, execution-based credit correction. A frozen judge maps visible transitions to rubric features, and a credit head predicts the change in expected terminal reward associated with the realized transition. Independently sampled two-sided replays correct prediction residuals using their recorded inclusion probabilities. We derive conditional unbiasedness and a variance decomposition that connects two design choices: which rubric features to retain, and where to allocate a fixed expected replay budget. The resulting criterion weights prediction errors by policy-score sensitivity and missing replay coverage; its allocation rule additionally accounts for continuation cost. We also describe a practical mixture with terminal leave-one-out advantages and distinguish its clipped, token-normalized PPO implementation from the ideal policy-gradient estimator. This preliminary report provides the method, proofs, an exact finite-model audit, and a controlled evaluation protocol. It makes no claim of empirical superiority on language-agent benchmarks.
Chinese Translation
长时程语言智能体只能获得稀疏的终端反馈,而中间评分标准(rubric)虽能提供结构化的进展评估,却可能存在设定偏差。在可重置的训练环境中,反事实延续 rollout 可以度量局部信用,但穷尽式重放的代价高昂。我们提出 CREDO,一个将演化的语义评分标准与选择性的、基于执行的信用校正相结合的框架。一个冻结的评判器(judge)将可见的状态转移映射为评分标准特征,信用预测头预测已实现的状态转移所对应的期望终端奖励变化。通过独立采样的双向重放(two-sided replays),依据其记录的包含概率对预测残差进行校正。我们推导了条件无偏性以及一个方差分解,该分解关联了两个设计选择:保留哪些评分标准特征,以及将固定的期望重放预算分配到何处。所得的准则依据策略得分的敏感性和缺失的重放覆盖度对预测误差进行加权;其分配规则还额外考虑了延续的成本。我们还描述了一种与终端留一法优势相结合的实用混合方法,并区分了其截断的、按 token 归一化的 PPO 实现与理想策略梯度估计器。这份初步报告给出了方法、证明、一个精确的有限模型审计以及受控评估协议。我们不声称其在语言智能体基准上具有经验上的优越性。
cs.AI / 78 / 2609.24186

LIMIT: Less Is More for Instruction Tuning in Text-to-SQL

LIMIT:文本到SQL指令微调中的“少即是多”
Ma, Haoyuan, Liu, Hengwei, Wu, Linjuan, Shen, Yongliang, Lu, Weiming
Abstract
Large language models have achieved remarkable progress on Text-to-SQL through reasoning-enhanced fine-tuning, yet existing approaches predominantly rely on massive instruction corpora under the assumption that scale drives performance. We challenge this paradigm by investigating a fundamental question: what is the minimal data requirement for effective Text-to-SQL instruction tuning? We propose LIMIT(Less Is More for Instruction Tuning in Text-to-SQL), a data-centric framework that demonstrates strong database reasoning can emerge from an extremely compact training set when examples are strategically selected. LIMIT operates through four stages: difficulty-aware filtering that identifies samples within the model's learning frontier, chain-of-thought synthesis with consistency-based selection, multi-dimensional quality scoring via LLM-as-judge, and genetic algorithm optimization that jointly maximizes schema coverage and sample quality. On the BIRD and Spider benchmark, LIMIT selects only 796 and 863 samples while achieving 100% table coverage, enabling Qwen3-8B to reach 69.1% and 88.9% execution accuracy.This result surpasses methods trained on 20 times more data and establishes a new state-of-the-art among open-source approaches. Our findings suggest that careful data curation, rather than scale, is the key to efficient Text-to-SQL learning.
Chinese Translation
大型语言模型通过推理增强微调在Text-to-SQL任务上取得了显著进展,然而现有方法主要依赖大规模指令语料库,并默认规模是驱动性能的关键。我们对这一范式提出质疑,并研究了一个根本性问题:有效的Text-to-SQL指令微调所需的最小数据量是多少?我们提出了LIMIT(Less Is More for Instruction Tuning in Text-to-SQL),这是一个以数据为中心的框架,证明了当样本经过策略性选择时,强大的数据库推理能力可以从极其精简的训练集中涌现。LIMIT包含四个阶段:识别处于模型学习前沿的样本的难度感知过滤、结合一致性选择的思维链合成、基于LLM-as-judge的多维度质量评分,以及联合最大化模式覆盖和样本质量的遗传算法优化。在BIRD和Spider基准上,LIMIT仅选择796和863个样本即实现了100%的表格覆盖率,使Qwen3-8B分别达到69.1%和88.9%的执行准确率。该结果超越了使用20倍数据量训练的方法,并在开源方法中确立了新的最先进水平。我们的研究结果表明,精心的数据筛选而非规模,才是高效Text-to-SQL学习的关键。
cs.AI / 79 / 2609.24198

SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision--Language Models

SKstars在SHROOM:基于视觉一致性引导的零样本与LoRA适配视觉-语言模型集成方法
Athar, Ali, Ahsan, Imran, Jung, Joon-Yong
Abstract
This paper describes the SKstars submission to SHROOM-Visions 2026, a shared task on fine-grained hallucination detection in large vision-language model outputs. The task requires systems to identify hallucinated character spans, assign hallucination categories, and provide confidence estimates for their predictions. Our approach combines zero-shot predictions from Qwen2.5-VL-72B-Instruct with those of a LoRA-adapted Qwen2.5-VL-7B-Instruct model. The outputs of the two models are integrated through a lightweight ensemble procedure, followed by span refinement and confidence adjustment. We evaluate the main system components on a small internal development subset and report the performance of the submitted system on the official English test set. SKstars achieved a Cor+Lbl score of 0.2902, ranking 15th among 29 teams, and obtained Cor and IoU scores of 0.3642 and 0.3151, respectively, ranking 18th on both metrics. The results show that combining a large zero-shot model with a smaller adapted model provides a practical framework for multilingual and fine-grained hallucination localization, while also highlighting the difficulty of transferring development-set improvements to hidden test data. Code and predictions: https://github.com/aliathar1401/SK-Stars-shroom-visions-2026
Chinese Translation
本文介绍了SKstars团队参加SHROOM-Visions 2026共享任务的方案,该任务针对大规模视觉-语言模型输出的细粒度幻觉检测。任务要求系统识别幻觉字符片段、分配幻觉类别,并为预测结果提供置信度估计。我们的方法将Qwen2.5-VL-72B-Instruct的零样本预测与经过LoRA适配的Qwen2.5-VL-7B-Instruct模型的预测相结合。两个模型的输出通过一个轻量级集成流程进行融合,随后进行片段精化与置信度调整。我们在一个小型内部开发子集上评估了系统的主要组件,并报告了所提交系统在官方英文测试集上的性能。SKstars取得了Cor+Lbl得分0.2902,在29支队伍中排名第15;Cor和IoU得分分别为0.3642和0.3151,两项指标均排名第18。结果表明,将大型零样本模型与较小的适配模型相结合,为多语言和细粒度幻觉定位提供了一种实用框架,同时也凸显了将开发集上的改进迁移到隐藏测试数据的难度。代码与预测结果:https://github.com/aliathar1401/SK-Stars-shroom-visions-2026
cs.AI / 80 / 2609.24229

Recovering Lost Details: Multi-Scale Frequency Compensation for Long-Term Time Series Forecasting

恢复丢失的细节:面向长期时间序列预测的多尺度频率补偿
Zou, Runmin, Xie, Siyi, Huang, Yaohui, Wang, Yun
Abstract
Long-term time series forecasting has made significant progress by leveraging multi-scale information to capture hierarchical temporal patterns and model long-range dependencies. However, temporal downsampling in existing multi-scale methods inevitably smooths detailed temporal fluctuations, and this information loss is further aggravated by their emphasis on dominant trends across scales, resulting in insufficiently expressive representations. To address this, we propose a Multi-Scale Wavelet Mixing (MWMixer) model, which incorporates a Bidirectional Frequency-Bands Mixing strategy to recover lost temporal details across scales, enabling complementary cross-scale information interactions. Then, a Dynamic Scale-Adaptive Fusion module learns time-varying weights for each scale to fuse multi-scale forecasts into the final prediction, enhancing the flexibility of multi-scale aggregation. In addition, a cross-scale consistency loss aligns each coarse-scale prediction with the interval-averaged fine-scale outputs, while a multi-scale supervision loss enforces prediction accuracy at each scale, promoting consistent learning across scales. Extensive experiments on seven real-world datasets demonstrate that MWMixer achieves competitive performance in long-term forecasting.
Chinese Translation
长期时间序列预测通过利用多尺度信息来捕捉层次化时间模式并建模长程依赖关系,已取得显著进展。然而,现有多尺度方法中的时间下采样不可避免地平滑了细节性的时间波动,且其对跨尺度主导趋势的强调进一步加剧了这种信息损失,导致表征的表达能力不足。为解决这一问题,我们提出了一种多尺度小波混合模型(Multi-Scale Wavelet Mixing, MWMixer),该模型引入了一种双向频带混合(Bidirectional Frequency-Bands Mixing)策略,以恢复跨尺度丢失的时间细节,实现互补的跨尺度信息交互。随后,一个动态尺度自适应融合(Dynamic Scale-Adaptive Fusion)模块为每个尺度学习时变权重,将多尺度预测融合为最终预测结果,增强了多尺度聚合的灵活性。此外,跨尺度一致性损失将每个粗尺度预测与细尺度输出的区间平均值对齐,而多尺度监督损失则确保每个尺度上的预测精度,从而促进跨尺度的一致性学习。在七个真实世界数据集上的大量实验表明,MWMixer 在长期预测任务中取得了具有竞争力的性能。
cs.AI / 81 / 2609.24243

Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation Enforcement

抑制视觉语言模型中的思维链混淆:从机制性证据到激活值约束
Mao, Xutao, Zhu, Jianing, Zhao, Jinman, Liu, Tongliang, Chu, Xiaowen, Wang, Cong, Han, Bo
Abstract
Reinforcement learning (RL) improves reasoning in vision-language models (VLMs) but can induce chain-of-thought (CoT) obfuscation: an operational, non-intentional outcome where task reward or accuracy rises while traces become less grounded and monitorable. Prior work largely documents this decay behaviorally, leaving its representation-level correlates and actionable controls unclear. We find that template- and ground-associated activations become less separable during RL; matched interventions support the contribution of selected features to monitorability degradation. Guided by this evidence, we propose Targeted Anti-obfuscation with Mechanistic Enforcement (TAME), which uses Sparse Autoencoders (SAEs) to combine behavioral feedback with targeted suppression of template-associated activations during RL. Its asymmetric constraint penalizes template activations only above their pre-RL baseline, anchoring the localized features while behavioral feedback promotes grounded refinements. Across VIRL-39k, SPA-VL, and two model families, TAME improves CoT monitorability by up to 30.9 and 16.7 percentage points over Group Relative Policy Optimization (GRPO), respectively. Blinded human evaluation finds higher human monitorability on both datasets, and two held-out monitor families reproduce the monitorability gains. Task accuracy changes are small and mixed, and general-capability benchmarks show task-specific trade-offs. These results provide a path from behavioral monitoring to representation-level oversight for more auditable RL-trained multimodal systems.
Chinese Translation
强化学习(RL)能够提升视觉语言模型(VLM)的推理能力,但也可能引发思维链(CoT)混淆:这是一种操作性、非意图性的现象,即任务奖励或准确率上升,而思维链轨迹却变得更难以溯源和监控。已有研究大多仅从行为层面记录这种退化现象,未能阐明其表征层面的关联机制以及可操作的干预手段。我们发现,在强化学习过程中,与模板相关的激活和与真实依据相关的激活变得越来越难以区分;匹配干预实验验证了特定特征对监控性退化的贡献。基于这一证据,我们提出了带有机制性约束的定向反混淆方法(Targeted Anti-obfuscation with Mechanistic Enforcement, TAME),该方法利用稀疏自编码器(SAE)将行为反馈与强化学习过程中对模板相关激活的定向抑制相结合。其非对称约束仅在模板激活超过强化学习前的基线水平时施加惩罚,从而锚定局部化特征,同时行为反馈促进基于真实依据的改进。在VIRL-39k和SPA-VL数据集以及两个模型家族上,TAME分别将CoT监控性相比组相对策略优化(GRPO)提升了最高30.9和16.7个百分点。盲评人类评估表明,在两个数据集上人类可监控性均有所提高,且两个保留的监控器家族均复现了监控性提升。任务准确率变化较小且不一致,通用能力基准测试显示出任务特定的权衡。这些结果为从行为监控走向表征级监督提供了一条路径,有助于构建更可审计的强化学习训练的多模态系统。
cs.AI / 82 / 2609.24265

Unsupervised Brain Anomaly Detection as a Bayesian Inverse Problem with Diffusion Prior

基于扩散先验的贝叶斯逆问题无监督脑部异常检测
Roy, Hugues, Dorent, Reuben, Burgos, Ninon
Abstract
Unsupervised anomaly detection (UAD) aims to localize abnormal regions in medical scans without pixel-level annotations. A typical strategy seeks to reconstruct a pseudo-healthy image that preserves subject-specific anatomy. Recently, diffusion models have been proposed to perform UAD. However, these methods rely on heuristic noise schedules or synthetic corruptions to balance subject-specificity and anomaly removal. In this work, we propose an alternative formulation of UAD as a Bayesian inverse problem under a diffusion prior. First, we introduce a latent spatial anomaly mask that models pixel-wise consistency between a test image and its latent corresponding pseudo-healthy image. Then, we propose an approximation of the unknown generation process that links healthy anatomy, anomalies, and the observed image, enabling a well-defined likelihood within the Bayesian framework. Building on recent advances in diffusion-based inverse problem methods, we jointly infer the pseudo-healthy image and the anomaly mask via annealed posterior sampling. We evaluate our approach on FDG PET (ADNI) and FLAIR MRI (BraTS 2021), demonstrating improved anomaly localization performance compared to other diffusion-based approaches and validating the contribution of our introduced model. Our code is available at https://github.com/HuguesRoy/UAD_DAPS.
Chinese Translation
无监督异常检测(UAD)旨在在没有像素级标注的情况下定位医学扫描中的异常区域。一种典型的策略是重建一幅保留受试者特异性解剖结构的伪健康图像。近来,扩散模型被提出用于执行UAD。然而,这些方法依赖于启发式的噪声调度或合成损坏来平衡受试者特异性与异常去除。在本工作中,我们提出将UAD重新表述为扩散先验下的贝叶斯逆问题。首先,我们引入一个潜在的空间异常掩码,用于建模测试图像与其潜在对应的伪健康图像之间的像素级一致性。然后,我们提出对连接健康解剖结构、异常和观测图像的未知生成过程的近似,从而在贝叶斯框架内实现明确定义的似然函数。基于扩散逆问题方法的最新进展,我们通过退火后验采样联合推断伪健康图像和异常掩码。我们在FDG PET(ADNI)和FLAIR MRI(BraTS 2021)数据集上评估了我们的方法,结果表明与其他基于扩散的方法相比,异常定位性能得到提升,并验证了我们所引入模型的贡献。我们的代码可在 https://github.com/HuguesRoy/UAD_DAPS 获取。
cs.AI / 83 / 2609.24277

How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation

一个数字值多少像素?用于GUI智能体置信度估计的位置感知坐标熵
Li, Yunxiang, Wu, Xixin, Meng, Helen
Abstract
GUI agents predict click coordinates as digit-token sequences, but standard text-LLM confidence estimation methods rank correct clicks from wrong ones only weakly. GUI-specific alternatives use K samples or new supervision, but still leave room for improvement. We trace part of this to place-value asymmetry: bounding-box correctness often makes higher-place digits more important than lower-place digits, so uniform aggregation weakens the signal that determines correctness. The fix is to weight each digit's Shannon entropy by its place value. We call this Place-Aware Coordinate Entropy (PACE). Across fixed-scale agents on ScreenSpot-Pro and ScreenSpot-v2, PACE wins both AUROC and selective accuracy on all primary comparisons in a single forward pass, matching or outperforming K-sample baselines at a fraction of the cost. PACE provides a per-click confidence estimate that turns coordinate-token internals into a practical confidence signal for GUI agent deployment.
Chinese Translation
GUI智能体将点击坐标预测为数字词元序列,但标准的文本大语言模型(LLM)置信度估计方法在区分正确点击与错误点击方面效果较弱。针对GUI的替代方案需要K次采样或新的监督信号,但仍有改进空间。我们将部分问题归因于位值不对称性:在边界框正确性判断中,高位数字往往比低位数字更重要,因此均匀聚合会削弱决定正确性的信号。解决方法是根据每个数字的位值对其香农熵进行加权。我们将该方法称为位置感知坐标熵(Place-Aware Coordinate Entropy, PACE)。在ScreenSpot-Pro和ScreenSpot-v2基准上对固定尺度智能体的评估中,PACE在单次前向传播中的所有主要比较中均取得了更优的AUROC和选择性准确率,并以极低的成本达到或超越了K次采样基线方法。PACE提供了一种逐点击的置信度估计,将坐标词元的内部信息转化为GUI智能体部署中实用的置信度信号。
cs.AI / 84 / 2609.24290

When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain

智能体应在何时以及如何进行澄清?CIGAsk:通过反事实信息增益教会大语言模型进行澄清
Li, Yunxiang, Wu, Xixin, Meng, Helen
Abstract
Instruction-tuned LLMs faced with underspecified queries often commit to a single interpretation rather than ask for clarification, producing confidently wrong answers. In our experiments, prompting alone is insufficient: models either ask for clarification on every query or ask vague questions that fail to recover the missing information. Addressing this failure requires learning two coupled skills: when to ask rather than answer and how to ask a question that recovers the disambiguating information. Existing recipes either address only one of these skills or require a separately trained critic. We propose CIGAsk, an RL recipe that teaches both skills through two complementary reward signals within a multi-turn GRPO loop. Counterfactual Information Gain (CIG) compares the gold-answer log-likelihood under a frozen reference model with and without the user response, providing per-turn credit that guides how to ask. The Asymmetric Ambiguity Bonus assigns a signed reward at the terminal token based on the gold ambiguity label, guiding when to ask. Across three clarification benchmarks spanning table, passage, and open-domain QA, CIGAsk-7B outperforms the strongest external baseline despite using a smaller backbone. It also transfers across datasets without per-dataset tuning while preserving single-turn QA performance on out-of-distribution benchmarks.
Chinese Translation
面对欠明确的查询,经指令微调的大语言模型(LLM)往往倾向于采用单一解释而非请求澄清,从而生成自信但错误的答案。在我们的实验中,仅靠提示词并不足够:模型要么对每个查询都请求澄清,要么提出无法恢复缺失信息的模糊问题。解决这一失败需要学习两项相互耦合的技能:何时应该提问而非直接回答,以及如何提出能够恢复消歧信息的问题。现有方法要么只针对其中一项技能,要么需要单独训练的评判模型。我们提出 CIGAsk,这是一种强化学习方案,通过在多轮 GRPO 循环中引入两个互补的奖励信号来同时教授这两项技能。反事实信息增益(Counterfactual Information Gain,CIG)比较冻结参考模型在有和没有用户回复情况下对标准答案的对数似然,提供逐轮的信用分配,指导模型如何提问。非对称模糊奖励(Asymmetric Ambiguity Bonus)根据标准模糊标签在终止标记处赋予带符号的奖励,指导模型何时提问。在涵盖表格、段落和开放域问答的三个澄清基准上,CIGAsk-7B 尽管使用了更小的骨干模型,仍优于最强的外部基线。此外,它无需针对每个数据集进行调优即可跨数据集迁移,同时在分布外基准上保持了单轮问答性能。
cs.AI / 85 / 2609.24324

Brain-Token Learning: Microstate-Based Tokenization and Multi-Scale Interaction for Long-Horizon EEG Sequence Modeling

脑token学习:基于微状态的分词与多尺度交互用于长时程脑电序列建模
Ye, Weishan, Pan, Yue, Zhang, Li, Huang, Gan, Liang, Zhen
Abstract
Electroencephalography (EEG) provides a non-invasive window into dynamic brain activity, yet modeling long-horizon EEG sequences remains challenging due to their high temporal complexity, substantial variability across subjects, and the lack of biologically meaningful sequence representations. Existing tokenization strategies, such as fixed-window and patch-based representations, discretize EEG signals according to artificial temporal boundaries, which may disrupt intrinsic brain-state dynamics. In this work, we propose Brain-Token Learning, a neuroscience-inspired framework that introduces Brain Tokenization for long-horizon EEG sequence modeling. Instead of partitioning EEG signals into predefined temporal segments, Brain Tokenization represents EEG as sequences of recurrent microstate-derived brain tokens, where each token corresponds to a quasi-stable large-scale brain state with variable temporal duration. Based on these biologically grounded tokens, we further develop a multi-scale token interaction module consisting of Latent State Aggregation and State Transition Modeling to jointly capture global brain-state context and local microstate transitions. We evaluate Brain-Token on five heterogeneous EEG datasets, including the newly collected long-horizon NeuroLong dataset and four affective or clinical EEG datasets (SEED, DEAP, MDD, and NSSI). Extensive experiments demonstrate that Brain-Token consistently outperforms conventional CNN/LSTM architectures, Transformer-based models, and domain adaptation methods across diverse EEG scenarios. Further analysis verifies the effectiveness of microstate-based tokenization and multi-scale interaction for learning robust and interpretable EEG representations. These results establish Brain-Token as a biologically grounded tokenization paradigm for long-horizon EEG sequence modeling.
Chinese Translation
脑电图(EEG)为观测动态大脑活动提供了一种无创窗口,然而由于其高度的时间复杂性、显著的被试间差异以及缺乏具有生物学意义的序列表示,长时程脑电序列建模仍然具有挑战性。现有的分词策略(如固定窗口和基于patch的表示)依据人为设定的时间边界对脑电信号进行离散化,这可能会破坏大脑状态的内在动态特性。在本工作中,我们提出Brain-Token Learning(脑token学习),一种受神经科学启发的框架,引入脑分词(Brain Tokenization)用于长时程脑电序列建模。与将脑电信号划分为预定义时间段不同,脑分词将脑电表示为由复现微状态(microstate)导出的脑token序列,其中每个token对应一个具有可变时间持续时间的准稳定大规模脑状态。基于这些具有生物学依据的token,我们进一步开发了一个多尺度token交互模块,由潜在状态聚合(Latent State Aggregation)和状态转移建模(State Transition Modeling)组成,以联合捕捉全局脑状态上下文和局部微状态转移。我们在五个异构脑电数据集上评估了Brain-Token,包括新收集的长时程NeuroLong数据集以及四个情感或临床脑电数据集(SEED、DEAP、MDD和NSSI)。大量实验表明,Brain-Token在多种脑电场景中始终优于传统的CNN/LSTM架构、基于Transformer的模型以及域适应方法。进一步分析验证了基于微状态的分词和多尺度交互在学习鲁棒且可解释的脑电表示方面的有效性。这些结果确立了Brain-Token作为一种具有生物学依据的长时程脑电序列建模分词范式。
cs.AI / 86 / 2609.24346

LADDER: Graph-Guided Diffusion Language Models for Efficient Multi-Hop Reasoning

LADDER:面向高效多跳推理的图引导扩散语言模型
Zhang, Senlei, Luo, Linhao, Zhang, Qian-Wen, An, Siyu, Dong, Junnan, Zhang, Shuhao, Sun, Xing
Abstract
Graph Retrieval-Augmented Generation (GraphRAG) has remarkably enhanced large language models on complex reasoning by leveraging structured entity topologies. However, existing frameworks heavily rely on standard autoregressive language models where the nature of inherent sequential generation severely hinders overall inference efficiency. Inspired by Diffusion Language Models (DLMs) that offer massive parallelism via continuous refine-in-parallel decoding, we aim to accelerate GraphRAG in the discrete space. However, it remains non-trivial for two challenges. First, partially denoised drafts are highly dynamic and uncertain, making dynamic graph grounding non-trivial. Second, raw denoising states are inherently noisy and unstable, making synchronous graph retrieval and multi-hop aggregation computationally prohibitive. To this end, we present LADDER, a novel framework that bridges diffusion language modeling with GraphRAG through graph-guided parallel decoding. Specifically, (i) we propose an event-driven self-clocking retrieval, inspired by our key insight that 88% of target entities emerge early in the partially denoised state, leading final commitment by an average of 5.7-9.6 steps. This mechanism dynamically triggers graph retrieval only when the set of graph-linkable entities expands, yielding an asynchronous self-clocking policy that bypasses learned gates or heuristic thresholds. (ii) An incomplete-query graph propagation module is designed to process the newly emerging entity queries using a specialized graph foundation model, continuously aggregating multi-hop evidence to sharpen parallel predictions and accelerate overall decoding convergence. Extensive experiments on three challenging multi-hop QA benchmarks show that LADDER raises average exact match from 39.6% to 45.2% while achieving a 4.1x latency reduction.
Chinese Translation
图检索增强生成(GraphRAG)通过利用结构化实体拓扑,显著提升了大语言模型在复杂推理任务上的表现。然而,现有框架高度依赖标准的自回归语言模型,其固有的顺序生成特性严重阻碍了整体推理效率。受扩散语言模型(Diffusion Language Models, DLMs)通过持续并行精炼解码实现大规模并行性的启发,我们致力于在离散空间中加速GraphRAG。然而,这面临两大非平凡挑战:其一,部分去噪的草稿具有高度动态性和不确定性,使得动态图对齐(graph grounding)难以实现;其二,原始去噪状态本身噪声大且不稳定,使得同步图检索与多跳聚合在计算上代价高昂。为此,我们提出了LADDER,一个通过图引导并行解码将扩散语言建模与GraphRAG相连接的新型框架。具体而言,(i)我们提出了事件驱动的自时钟检索机制,该机制源于我们的一个关键发现:88%的目标实体在部分去噪状态中即已早期涌现,并平均领先最终确定5.7至9.6步。该机制仅在可链接到图的实体集合扩展时才动态触发图检索,形成了一种异步自时钟策略,从而无需依赖学习得到的门控或启发式阈值。(ii)我们设计了不完整查询图传播模块,利用专门的图基础模型处理新涌现的实体查询,持续聚合多跳证据,以锐化并行预测并加速整体解码收敛。在三个具有挑战性的多跳问答基准上的大量实验表明,LADDER将平均精确匹配率从39.6%提升至45.2%,同时实现了4.1倍的延迟降低。
cs.AI / 87 / 2609.24352

Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs

少样本示例促使大语言模型使用上下文世界表示
Matsutani, Kohsei, Minegishi, Gouki, Park, Core Francisco, Kojima, Takeshi, Iwasawa, Yusuke, Matsuo, Yutaka
Abstract
Large language models (LLMs), when acting as agents, are expected to take observed data in context, infer the latent state space underlying the world, and leverage it for downstream prediction. However, prior work demonstrated that LLMs struggle to use representations learned in context on a graph tracking task, where the model needs to construct a representation of the graph governing data generation process and use it for subsequent predictions. In this paper, we show that extending this to few-shot settings, where each demonstration is generated from a different world with either the same or different graph topologies, enhances its prediction on 6 models from 4 model families. To understand this improvement, we linearly probe a low-dimensional world representation that encodes graph information in the hidden states. Notably, we find that few-shot demonstrations relocate the world representation and increase its predictive use. Specifically, for each model, these world representations shift in directions nearly orthogonal to their original subspace, and interventions on these representations selectively impair performance more than interventions on other subspaces. Consistent with this insight, we show that few-shot demonstrations with observations from different worlds improve performance on ARC-AGI-1&2, web agent tasks, and Othello. Our findings elucidate the role and internal mechanisms of few-shot demonstrations in in-context world modeling. More broadly, our work advances our understanding of how LLM agents learn from in-context observations and provides implications for their further improvement.
Chinese Translation
大语言模型(LLM)作为智能体时,应能够根据上下文中的观测数据进行推理,推断世界背后的潜在状态空间,并将其用于下游预测。然而,先前的研究表明,LLM 在图追踪任务中难以利用在上下文中学习到的表示——该任务要求模型构建控制数据生成过程的图的表示,并将其用于后续预测。本文表明,将该任务扩展到少样本(few-shot)设置(其中每个示例由具有相同或不同图拓扑的不同世界生成),可以提升来自 4 个模型家族的 6 个模型的预测性能。为理解这一改进,我们通过线性探针(linear probe)检测隐藏状态中编码图信息的低维世界表示。值得注意的是,我们发现少样本示例会重新定位世界表示并增强其预测用途。具体而言,对于每个模型,这些世界表示向与其原始子空间近似正交的方向偏移,且对这些表示的干预比对其余子空间的干预更有选择性地损害性能。与这一发现一致,我们证明包含来自不同世界观测的少样本示例可以提升模型在 ARC-AGI-1 和 ARC-AGI-2、网页智能体任务以及黑白棋(Othello)上的表现。我们的发现阐明了少样本示例在上下文世界建模中的作用及内部机制。更广泛地说,我们的工作加深了对 LLM 智能体如何从上下文观测中学习的理解,并为其进一步改进提供了启示。
cs.AI / 88 / 2609.24362

VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning

VLM-in-Sandbox:面向智能体视觉推理的视觉工作区
Yang, Hexiong, Chen, Mingrui, Cao, Jie, He, Ran
Abstract
Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language models to vision-language models (VLMs) introduces a distinct state-management problem. Visual reasoning produces intermediate image-valued evidence---crops, masks, overlays, zoomed regions, and analytic renderings---that must remain addressable without accumulating unboundedly in multimodal context. We introduce VLM-in-Sandbox, a training-free framework for agentic multimodal reasoning in controlled computer environments. Its Visual Workspace registers generated artifacts in an image ledger, maintains a bounded active visual context, and lets the model explicitly promote selected evidence for subsequent inspection. This separates visual evidence generation, performed by sandbox tools, from visual evidence management. Across seven benchmarks and four base VLMs, VLM-in-Sandbox achieves the highest sample-weighted average accuracy among Vanilla VLM, Append-only Sandbox, and the proposed method. A compiler-matched $2\times2$ study on 1,260 examples further separates model-directed visibility from bounded retention: VLM-in-Sandbox reaches 66.27% accuracy with 18.6% fewer total tokens than the automatic, retain-all control. Over all 6,350 submitted GPT-4.1-mini examples, it produces 302 rescues and 142 regressions relative to Original Append-only. A local vLLM study with prefix caching confirms that the smaller request workload also reduces uncached tokens, time to first token, and end-to-end latency. These results identify explicit visual evidence state as a central abstraction for sandboxed VLM agents.
Chinese Translation
沙盒式计算机环境支持借助工具、可执行程序和持久化文件进行多步推理,然而将其从语言模型扩展到视觉语言模型(VLM)时,引入了一个独特的状态管理问题。视觉推理会产生以图像为载体的中间证据——如裁剪图、掩码、叠加图、放大区域和分析性渲染图——这些证据必须保持可寻址,同时又不能在多模态上下文中无限累积。我们提出 VLM-in-Sandbox,这是一个无需训练的框架,用于在受控计算机环境中进行智能体式多模态推理。其视觉工作区(Visual Workspace)将生成的产物登记到一个图像账本中,维护一个有界的活跃视觉上下文,并允许模型显式地将所选证据提升以供后续检查。这将视觉证据的生成(由沙盒工具执行)与视觉证据的管理分离开来。在七个基准测试和四个基础 VLM 上,VLM-in-Sandbox 在 Vanilla VLM、仅追加沙盒(Append-only Sandbox)和所提方法中取得了最高的按样本加权的平均准确率。在 1,260 个样本上进行的编译器对齐 2×2 对照实验进一步区分了模型主导的可见性与有界保留:VLM-in-Sandbox 达到 66.27% 的准确率,同时总 token 数比自动保留全部证据的对照组减少 18.6%。在全部 6,350 个提交给 GPT-4.1-mini 的样本上,相对于原始仅追加基线,它带来了 302 次 rescued(纠错)和 142 次回退。基于前缀缓存的本地 vLLM 实验进一步证实,更小的请求负载还降低了未缓存 token 数、首 token 延迟和端到端延迟。这些结果表明,显式的视觉证据状态是沙盒式 VLM 智能体的一个核心抽象。
cs.AI / 89 / 2609.24453

Predicting Postprandial Glycemic Response from Meal Images, Clinical Variables, and Gut Microbiome Information

基于餐食图像、临床变量和肠道菌群信息预测餐后血糖反应
Kondratyeva, Varvara, Zaripova, Kamilia, Navab, Nassir, Farshad, Azade
Abstract
Predicting postprandial glycemic response (PPGR) is fundamental to personalized nutrition and type 2 diabetes management, yet existing approaches typically rely on manually reported dietary intake, limiting their scalability in free-living settings. We propose a multimodal framework that replaces manual dietary logging with image-derived macronutrient estimates and integrates them with clinical variables and gut microbiome information for personalized PPGR prediction. The framework jointly performs image-based macronutrient estimation and glucose prediction, while an attention-based prediction module models interactions between dietary and host-specific information. We evaluate the proposed approach on a real-world dataset comprising meal images, continuous glucose monitoring, clinical variables, and gut microbiome profiles. The proposed model outperforms existing PPGR baselines using image-derived nutritional inputs and approaches the performance of methods that rely on manually reported macronutrients despite using automatically estimated nutritional information. These results demonstrate that combining image-derived nutrition with complementary clinical and gut microbiome information provides a practical foundation for scalable personalized PPGR prediction.
Chinese Translation
预测餐后血糖反应(PPGR)是个性化营养和2型糖尿病管理的基础,然而现有方法通常依赖人工记录的饮食摄入,限制了其在自由生活场景中的可扩展性。我们提出了一种多模态框架,用基于图像估算的宏量营养素取代人工饮食记录,并将其与临床变量和肠道菌群信息相结合,用于个性化PPGR预测。该框架联合执行基于图像的宏量营养素估算和血糖预测,同时采用基于注意力机制的预测模块对饮食信息与宿主特异性信息之间的交互进行建模。我们在一个包含餐食图像、连续血糖监测、临床变量和肠道菌群谱的真实世界数据集上评估了所提出的方法。使用图像来源的营养输入时,所提出的模型优于现有PPGR基线方法;尽管使用的是自动估算的营养信息,其性能接近依赖人工报告宏量营养素的方法。这些结果表明,将图像来源的营养信息与互补的临床和肠道菌群信息相结合,为可扩展的个性化PPGR预测提供了切实可行的基础。
cs.AI / 90 / 2609.24480

Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards

Fathom-Vaidya:基于评分准则奖励推进医学推理
Shah, Kalash, Singh, Kunal, J, Snehan, Singh, Shreyas
Abstract
Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the convergent, evidence-driven task of inferring a patient's condition from clinical data to produce a diagnosis, and clinical healthcare reasoning: the broader, navigational judgment required to communicate, plan, and adapt across multi-turn clinical interactions where a single correct answer may not exist. Recent benchmarks such as HealthBench and MedXpertQA reveal persistent weaknesses in both areas, exposing failures in complex diagnostic scenarios and limitations in contextual, patient-centered dialogue. We introduce a sequential training framework that targets these facets using synthetic data and rubric-based reinforcement learning. First, we improve diagnostic reasoning using MedBullets-derived questions with rule- and rubric-guided Reinforcement Learning (RL). We then shift to clinical reasoning by generating 5.3k synthetic multi-turn scenarios, each paired with multi-dimensional rubrics to comprehensively assess the response. This approach yields over 10% improvement on MedXpertQA, and our 30B model achieves 50.1% accuracy on HealthBench-Hard, surpassing proprietary baselines including GPT-5 (thinking). Our results show that targeted synthetic datasets and rubric-based training can systematically improve both diagnostic and interactive clinical reasoning in medical LLMs.
Chinese Translation
在医疗健康领域部署大语言模型(LLM)需要在两个互补维度上具备稳健的性能:一是诊断推理,即从临床数据中推断患者病情并给出诊断的收敛性、循证式任务;二是临床医疗推理,即在多轮临床交互中进行沟通、规划和适应的更广泛的导航式判断,此类场景中可能不存在唯一正确答案。近期 HealthBench 和 MedXpertQA 等基准测试揭示了模型在这两个方面的持续弱点,暴露了复杂诊断场景中的失败以及上下文感知、以患者为中心的对话能力的局限。我们提出了一种顺序训练框架,利用合成数据和基于评分准则(rubric)的强化学习来针对这些方面进行优化。首先,我们使用源自 MedBullets 的问题,通过规则和评分准则引导的强化学习(RL)提升诊断推理能力。随后,我们转向临床推理,生成了 5.3k 个合成的多轮对话场景,每个场景均配有多维度评分准则以全面评估模型响应。该方法在 MedXpertQA 上取得了超过 10% 的提升,我们的 30B 模型在 HealthBench-Hard 上达到 50.1% 的准确率,超越了包括 GPT-5(thinking)在内的专有基线模型。我们的结果表明,针对性的合成数据集和基于评分准则的训练能够系统地提升医学大语言模型的诊断推理和交互式临床推理能力。
cs.AI / 91 / 2609.24517

Not All Task Vectors Need Equal Rank: Energy-Proportional Allocation for Model Merging

并非所有任务向量都需要相同秩:面向模型合并的能量比例分配方法
Cho, Hyunjoong, Jang, Jinhyeok
Abstract
Model merging aims to combine multiple fine-tuned models derived from a common pretrained model into a single multi-task model without additional joint training. Recent spectral merging methods improve over simple weight averaging by exploiting low-rank structures of task-specific updates, but they commonly assign the same rank capacity to every task. This uniform allocation ignores that task vectors can have heterogeneous spectral complexity, causing the shared merging space to be used suboptimally. In this paper, we propose Spectral Energy-proportional Rank Allocation (SERA), a simple task-adaptive strategy that allocates ranks according to the singular-value energy structure of each task vector. By assigning richer spectral capacity to complex or isolated tasks and fewer directions to compact tasks, SERA extends SVD-based model merging from uniform-capacity merging to task-dependent capacity allocation. Experiments under standard vision model merging protocols show that SERA improves multi-task merging performance while preserving the same total rank budget as existing spectral merging methods. Further analysis demonstrates that task-level spectral concentration is closely related to the per-task effect of adaptive rank allocation, providing insight into when and why SERA is effective.
Chinese Translation
模型合并旨在将由同一预训练模型微调得到的多个模型组合为一个单一的多任务模型,而无需额外的联合训练。近期的谱方法合并(spectral merging)通过利用任务特定更新的低秩结构,性能优于简单的权重平均,但这些方法通常为每个任务分配相同的秩容量。这种均匀分配忽略了任务向量可能具有异构的谱复杂度,导致共享合并空间的使用欠佳。在本文中,我们提出了谱能量比例秩分配方法(Spectral Energy-proportional Rank Allocation, SERA),这是一种简单的任务自适应策略,根据每个任务向量的奇异值能量结构来分配秩。通过为复杂或孤立的任务分配更丰富的谱容量,而为紧凑的任务分配更少的方向,SERA 将基于 SVD 的模型合并从均匀容量合并扩展到任务相关的容量分配。在标准视觉模型合并协议下的实验表明,SERA 在保持与现有谱合并方法相同的总秩预算的同时,提升了多任务合并性能。进一步分析表明,任务级谱集中度与自适应秩分配对每个任务的效果密切相关,为何时以及为何 SERA 有效提供了见解。
cs.AI / 92 / 2609.24555

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence

无尽的考试:从当今模型通往超级智能的数学构造
Zhang, Muhan
Abstract
We introduce the Endless Exam, a benchmark for measuring mathematical progress from today's models toward artificial superintelligence through fourteen parameterised construction families. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at $1$. The families draw on open mathematical problems for long-term targets and generate new instances at larger parameters, where compact certificates keep large constructions verifiable. Across eight models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though no evaluated system surpasses a published frontier. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.
Chinese Translation
我们提出了“无尽的考试”(Endless Exam),这是一个通过十四个参数化构造族来衡量从当今模型到人工超级智能的数学进步的基准。每个提交的对象都会被自动检查有效性,并根据已公开的前沿或构造基线被赋予一个相对质量分数,且改进上限不被限制在1以内。这些构造族以开放数学问题作为长期目标,并在更大参数下生成新实例,其中紧凑的证书使大型构造保持可验证性。在对69个不同实例进行评估的八个模型中,连续的质量分数即使在没有被评估系统超越已公开前沿的情况下也能区分性能表现。规模—质量曲线展示了构造质量如何随问题规模增大而变化。我们公开发布了生成器、验证器、参考文献、模型响应及分析结果,以支持在人类前沿之前及超越人类前沿之后进行持续测量。
cs.AI / 93 / 2609.24620

Ascent: An Agentic System over the Model Context Protocol for Real-World Clinical Data Analysis

Ascent:基于模型上下文协议(Model Context Protocol)的现实世界临床数据分析智能体系统
Ziletti, Angelo, D'Ambrosi, Leonardo, Tuchardt, Melanie, Kondziella, Tim
Abstract
Answering epidemiological questions from real-world clinical data requires medical coding, schema-aware SQL, and validation of implicit choices about populations, denominators, and time. We present Ascent, an agentic system that exposes medical coding, question answering, and cohort analysis through a shared Model Context Protocol tool surface for standardized and native schemas. We introduce EpiTrap, a dataset testing whether systems avoid recognized pharmacoepidemiological errors, and compare a fixed pipeline with agents across models and orchestrators. With capable models, agents improve accuracy over the fixed pipeline by an average of 27 and 20 percentage points on native and standardized schemas, respectively. These gains require more tool calls and longer runtimes. Experience from real projects highlights the system's value for feasibility assessment, diagnostic iteration, and expert-guided analysis.
Chinese Translation
基于现实世界临床数据回答流行病学问题需要医学编码、具备模式感知能力的SQL,以及对人群、分母和时间等隐含选择进行验证。我们提出了Ascent,这是一个智能体系统,通过共享的模型上下文协议(Model Context Protocol)工具接口,为标准化和原生模式提供医学编码、问答和队列分析功能。我们引入了EpiTrap数据集,用于测试系统能否避免已知的药物流行病学错误,并在不同模型和编排器下将固定流水线与智能体进行比较。在使用能力较强的模型时,智能体在原生和标准化模式上相对于固定流水线的准确率分别平均提高了27和20个百分点。这些提升需要更多的工具调用和更长的运行时间。来自实际项目的经验凸显了该系统在可行性评估、诊断性迭代以及专家指导分析方面的价值。
cs.AI / 94 / 2609.24625

Custom Named Entity Recognition and Topic Classification for Global Health Publications

面向全球健康出版物的自定义命名实体识别与主题分类
Skura, Genis, Geissbühler, Antoine, Falcone, Jean-Luc
Abstract
How should natural language processing models be selected and adapted for global health literature in environments where annotated data and computational resources are limited? This thesis investigates these challenges through experiments on semantic tag discovery, named entity recognition (NER), and multi-label topic classification. First, skip-gram word2vec models trained on progressively larger specialized corpora are compared with BioWordVec to assess how corpus size and domain context influence tag discovery. Vocabulary coverage and qualitative evaluation indicate that broader coverage does not necessarily yield more useful domain-specific associations. The analysis then turns to entity extraction, comparing convolutional spaCy models with a RoBERTa-based transformer on 1,000 annotated sentences. Under a lenient scoring protocol, the transformer achieves 0.80 micro-F1 versus 0.65-0.69 for convolutional models, but takes 82 seconds rather than 5-6 seconds. This trade-off motivates fine-tuning convolutional models and integrating a disease recognizer that achieves 81.33% test F1 on the NCBI Disease Corpus. Combined with PDF preprocessing, entity filtering, and MeSH enrichment, the resulting pipeline supports document-level indexing. To complement entity extraction with thematic annotation, MiniLM-based few-shot classification is compared with BART-MNLI zero-shot inference across 50 topics and 1,000 handcrafted test sentences. BART-MNLI achieves 95.2% single-label accuracy versus 59%; reported multi-label accuracies are 88% and 32% under partly manual assessment. However, its higher inference cost limits practical integration. The results show where domain specialization and lightweight adaptation offer practical value, and where transformer accuracy justifies higher inference costs, providing an empirical basis for building knowledge systems under resource constraints.
Chinese Translation
在标注数据和计算资源有限的环境中,应如何为全球健康文献选择和适配自然语言处理模型?本论文通过对语义标签发现、命名实体识别(NER)和多标签主题分类的实验来研究这些挑战。首先,将在逐步扩大的专业语料库上训练的skip-gram word2vec模型与BioWordVec进行比较,以评估语料库规模和领域上下文如何影响标签发现。词汇覆盖率和定性评估表明,更广的覆盖率并不一定能产生更有用的领域特定关联。随后分析转向实体抽取,在1,000个标注句子上比较卷积spaCy模型与基于RoBERTa的transformer模型。在宽松的评分协议下,transformer模型达到0.80的micro-F1,而卷积模型为0.65-0.69,但前者需要82秒而非5-6秒的处理时间。这一权衡促使我们对卷积模型进行微调,并集成一个在NCBI Disease语料库上测试F1达到81.33%的疾病识别器。结合PDF预处理、实体过滤和MeSH词表扩展,所得到的流水线支持文档级索引构建。为以主题标注补充实体抽取,本文在50个主题和1,000个手工构建的测试句子上,将基于MiniLM的少样本分类与BART-MNLI零样本推理进行比较。BART-MNLI的单标签准确率达到95.2%,而MiniLM为59%;在部分人工评估下,两者的多标签准确率分别为88%和32%。然而,BART-MNLI较高的推理成本限制了其实际集成。研究结果指明了领域专业化和轻量化适配在哪些方面具有实用价值,以及transformer的准确性在哪些方面足以证明更高推理成本的合理性,为在资源受限条件下构建知识系统提供了实证依据。
cs.AI / 95 / 2609.24662

DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security

DUMA-Bench:用于评估LLM智能体安全性的双控制多智能体基准测试
Aleksandrov, Ivan, Kochnev, German, Sadiekh, Sabrina, Rogoza, Yaroslav
Abstract
LLM-based agents increasingly operate in environments where they interact with users, tools, and external systems. Yet most security evaluations assume passive users and static control, ignoring the interactive dynamics that shape real agent behavior. We introduce \textbf{DUMA-Bench}, a benchmark and evaluation protocol for measuring agent security under \emph{dual-control} interaction, where both the agent and the user can influence the shared environment state. DUMA-Bench extends $\tau^2$-bench ~\cite{barres2025tau} with adversarial environments covering eight vulnerability classes, including RAG poisoning, cross-agent manipulation, and unsafe output handling. We evaluate \textbf{14 models from five model families} (OpenAI, Anthropic, DeepSeek, Qwen, and Z.ai) across eight domains and multiple user-behavior regimes. Across our experiments, introducing dual-control interaction increases the attack success rate from \textbf{26.9\%} to \textbf{41.1\%}. These results show that agent security is not solely a property of the model but emerges from the interaction between the model, the user, and the environment. DUMA-Bench provides a missing evaluation layer for studying security in realistic agent deployments.
Chinese Translation
基于LLM的智能体越来越多地在与用户、工具和外部系统交互的环境中运行。然而,大多数安全性评估都假设用户是被动的、控制是静态的,忽视了塑造真实智能体行为的交互动态性。我们提出了DUMA-Bench,这是一个用于衡量双控制(dual-control)交互下智能体安全性的基准测试与评估协议,在双控制交互中,智能体和用户都可以影响共享的环境状态。DUMA-Bench在τ²-bench(barres2025tau)的基础上进行了扩展,涵盖了八类漏洞的对抗性环境,包括RAG投毒、跨智能体操纵和不安全的输出处理。我们在八个领域和多种用户行为模式下评估了来自五个模型家族(OpenAI、Anthropic、DeepSeek、Qwen和Z.ai)的14个模型。实验结果表明,引入双控制交互使攻击成功率从26.9%提升至41.1%。这些结果表明,智能体安全性并非仅仅取决于模型本身的属性,而是从模型、用户和环境之间的交互中涌现出来的。DUMA-Bench为研究现实智能体部署中的安全问题提供了一个此前缺失的评估层面。
cs.AI / 96 / 2609.24663

Beyond Endpoint Performance: Process-Level Evaluation of Self-Evolving Agents

超越端点性能:自进化智能体的过程级评估
Lin, Hongqiang, Liu, Chao, Bai, Xiaofan, Jin, Xuan, Li, Yuhong, Zheng, Nenggan, Cao, Xipeng
Abstract
Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated throughout an experience stream, the capabilities they support may evolve. Consequently, endpoint performance alone offers an incomplete view of self-evolution. Process-level evaluation is therefore essential to identify when a target capability emerges and whether later updates strengthen, preserve, or weaken it. Motivated by this, we propose \textsc{EvoPathBench}, a benchmark that tracks individual capabilities during artifact-level self-evolution. EvoPathBench fixes the base model, tools, freezes evolving artifacts at successive checkpoints, and evaluates the target capability on held-out episodes. This benchmark evaluates agent self-evolution using public trading data and calibrated trajectories. It tests three capabilities: generalization to unseen tasks, retention after unrelated learning, and rule adaptation to new evidence. Experimental results show that gains on similar unseen tasks often weaken under distribution shift, retention losses are concentrated in a minority of evolution paths, and no method achieves reliable rule adaptation. Moreover, while self-evolution enables agents to generate candidate artifacts with substantial held-out gains, the selected updates consistently fall short of realizing this potential. Together, these findings establish capability-level process evaluation as a foundation for analyzing self-evolution, identifying candidate evaluation and selection as key targets for improvement.
Chinese Translation
自进化智能体(self-evolving agents)将交互反馈转化为持久化的产物(artifacts),如记忆或技能,这些产物进而指导后续决策。随着这些产物在经验流中被迭代更新,其所支撑的能力也可能随之演化。因此,仅凭端点性能无法完整刻画自进化过程。为此,过程级评估至关重要,它能够识别目标能力何时涌现,以及后续更新是增强、保持还是削弱了该能力。基于这一动机,我们提出了 EvoPathBench,一个在产物级自进化过程中追踪个体能力的基准。EvoPathBench 固定基础模型与工具,在相继的检查点冻结不断演化的产物,并在保留集(held-out)情景上评估目标能力。该基准使用公开的交易数据和校准的轨迹来评估智能体的自进化,测试三种能力:对未见任务的泛化能力、在无关学习后的保持能力,以及面向新证据的规则适应能力。实验结果表明:在相似的未见任务上获得的收益在分布偏移下往往减弱;保持能力的损失集中在少数进化路径中;且没有任何方法能实现可靠的规则适应。此外,尽管自进化使智能体能够生成在保留集上具有显著收益的候选产物,但被选中的更新始终未能充分实现这一潜力。综上所述,这些发现确立了能力级过程评估作为分析自进化的基础,并指出候选产物的评估与选择是关键的改进方向。
cs.AI / 97 / 2609.24677

TimeLitmus: A Diagnostic Benchmark for Cross-Modal Understanding and Explanation Faithfulness in Event-Conditioned Time-Series Prediction

TimeLitmus:面向事件条件时间序列预测中跨模态理解与解释忠实度的诊断性基准
Gong, Jie, Jiang, Maowei, Liu, Zhiwei, Chen, Yankai, Xiong, Guojun, Liu, Xue, Peng, Min, Xie, Qianqian, Ananiadou, Sophia
Abstract
Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cannot reveal whether correct answers reflect effective integration of the two inputs or instead arise from event polarity, unimodal priors, or superficial cues. Likewise, plausible explanations may rationalize predictions without faithfully reflecting the evidence that drives model behavior. We introduce TimeLitmus, a diagnostic benchmark for cross-modal understanding and explanation faithfulness in event-conditioned time-series prediction. TimeLitmus contains 4,856 evaluation records across Finance and Traffic, combining natural prediction with controlled counterfactual and contrastive interventions, explanation-targeted faithfulness tests, and systematic shortcut controls. Across ten representative LLMs, standard prediction accuracy substantially overstates reliable cross-modal understanding: Hard Paired Contrast (HPC) pair correctness peaks at only 19.2% in Finance and 11.7% in Traffic, and all ten models show lower-than-expected consistency on Finance series-side controls. Models often recognize scenario relations explicitly yet fail to apply them during independent prediction. Explanation faithfulness shows a similar gap: in Traffic, most models cite the manipulated temporal factor in over 90% of cases, while behavioral support remains below 22%. Human annotators outperform LLMs on matched controlled and hard-pair diagnostics, confirming that these distinctions are recoverable from the inputs. Natural-only adaptation yields selective gains in evidence selection and input sensitivity, but not consistent gains in controlled or hard-pair behavior. The benchmark, evaluation suite, and supervised adaptation data will be released publicly.
Chinese Translation
大语言模型(LLM)正被越来越多地用于基于数值时间序列历史和文本事件进行预测。然而,仅凭准确率无法揭示正确的答案究竟反映了对两类输入的有效整合,还是仅仅源于事件的极性、单模态先验或表层线索。同样,看似合理的解释可能只是对预测的事后合理化,而并未忠实反映驱动模型行为的证据。我们提出了 TimeLitmus,一个用于评估事件条件时间序列预测中跨模态理解与解释忠实度的诊断性基准。TimeLitmus 包含涵盖金融(Finance)与交通(Traffic)两个领域的 4,856 条评测记录,结合了自然预测与受控的反事实及对比干预、面向解释的忠实度测试,以及系统性的捷径控制。在十个代表性 LLM 上,标准预测准确率显著夸大了可靠的跨模态理解能力:硬配对对比(Hard Paired Contrast, HPC)对的正确率在 Finance 中最高仅为 19.2%,在 Traffic 中仅为 11.7%,且全部十个模型在 Finance 序列侧控制项上均表现出低于预期的一致性。模型往往能够显式地识别场景关系,却在独立预测时未能加以应用。解释忠实度也呈现类似的差距:在 Traffic 中,大多数模型在超过 90% 的案例中引用了被操纵的时间因素,而行为支持度却低于 22%。人类标注者在匹配的受控诊断和硬配对诊断上的表现优于 LLM,这证实了这些区别是可以从输入中恢复的。仅在自然数据上的适配带来了在证据选择和输入敏感性方面的选择性提升,但在受控或硬配对行为上并未带来一致的改进。该基准、评测套件及监督适配数据将会公开发布。
cs.AI / 98 / 2609.24744

World State Generator

世界状态生成器
Jeong, Sungheon, Yun, Sanggeon, Masukawa, Ryozo, Alimohamadi, Haleh, Imani, Mahdi, Imani, Mohsen
Abstract
Language agents solve complex tasks through plans and actions. A single step the world refuses puts the goal out of reach, and what the agent does next decides the task. Prompted planners fail at exactly this point, rewriting the refused step in new words, meeting the same refusal, and burning the attempt budget without moving. They fail because the plan was never tied to the world, so a refusal has nothing in the plan to attach to. A world is where a task runs, and it has its own rules, its own admissible actions, and its own constraints. We build synthetic worlds across 7 domains and extract training data from them. A program enforces each world's rules and grades its goal, and every world is admitted only if its goal is reachable from its initial state. Agents run inside and leave verified failures paired with repairs that carried the run to a state the world certified, a record of about 226K trajectories. On this record we train the World State Generator, a model that writes a plan as checkable states of the world and keeps that plan aligned with the world it runs in. That alignment is what a plan written in language lacks, since the world it runs in has physical limits, logical dependencies, and required orders the language never states, and the plan encounters these rules only when a state fails. WSG takes that failure as the rule the world has stated and rewrites the remaining states to obey it, so the plan bends to the world as the run goes on. Across 7 public benchmarks, WSG raises end-to-end success for two open models near 30B parameters over prompting and brings to the level of proprietary model.
Chinese Translation
语言智能体通过规划和行动来解决复杂任务。世界只要拒绝其中一步,目标便可能无法达成,而智能体接下来的做法决定了任务的成败。基于提示的规划器恰恰在这一点上失败:它们只是用新的措辞改写被拒绝的步骤,遭遇同样的拒绝,并不断消耗尝试预算而毫无进展。它们失败的原因在于计划从未与世界绑定,因此面对拒绝时计划中没有任何可依据之处。世界是任务运行的环境,它有自己的规则、自己的可行动作和自己的约束。我们在 7 个领域中构建了合成世界并从中提取训练数据。每个世界的规则由程序强制执行并对目标进行评判,且每个世界只有在其目标可从初始状态到达时才会被纳入。智能体在其中运行,留下经核实的失败案例,以及使运行达到世界所认可状态的修复方案,由此形成了约 22.6 万条轨迹的记录。我们在此记录上训练了世界状态生成器(World State Generator, WSG),该模型将计划写成可核查的世界状态,并使计划与其运行的世界保持一致。这种一致性正是用语言写成的计划所缺失的,因为它所运行的世界具有物理限制、逻辑依赖和隐含的先后顺序,而语言从未明确说明这些规则,计划只有在某个状态失败时才会遇到它们。WSG 将这种失败视为世界已申明的规则,并改写剩余状态以遵循该规则,从而使计划随着运行不断顺应世界。在 7 个公开基准上,WSG 将两个约 300 亿参数的开源模型的端到端成功率相比提示方法显著提升,并使其达到专有模型的水平。
cs.AI / 99 / 2609.24755

Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents

Epi-Logic:自主AI智能体中认知运行时控制、模式有效性检验与受控顺应的概念框架
Wetzk, Boris
Abstract
Autonomous AI agents are increasingly deployed in areas where wrong decisions are hard to reverse. This paper examines schema mismatch: the condition in which an agent operates within an interpretive frame that no longer applies to the current context. Outputs produced under such a mismatch can appear internally consistent, linguistically plausible, and largely factually correct; output-quality metrics alone therefore capture the underlying loss of validity only partially. The paper introduces Epi-Logic, a conceptual framework for epistemic runtime control. It couples the detection of schema dissonance, a graduated reduction of autonomy, and the auditable switch to a validated schema. A schema is formalised as a tuple of variable space, expectation model, validity conditions, axioms, and metadata. The Epi-Score aggregates seven graded dimensions of epistemic dissonance; the temporal validity dimension D8, violations of the validity conditions G, and axiom violations are carried as separate categorical paths that are not offset against the aggregate. The architecture rests on a checking asymmetry: formalised validity conditions can be checked at runtime, whereas the correctness of many actions is established only ex post. The paper separates two architectural properties, a conditional result from sequential changepoint detection, and an empirical remainder. Eight falsifiable propositions with named baselines describe the transition to empirical validation. All propositions are empirically testable hypotheses, not established results.
Chinese Translation
自主AI智能体正日益被部署在错误决策难以逆转的领域。本文研究模式失配(schema mismatch)问题:即智能体在一种不再适用于当前情境的解释性框架内运行的状态。在这种失配状态下产生的输出可能看起来内部一致、语言上合理,且在事实层面大体正确;因此,仅凭输出质量指标只能部分捕捉到潜在有效性的丧失。本文提出了Epi-Logic,一个用于认知运行时控制的概念框架。它将模式失谐(schema dissonance)的检测、自主性的分级缩减以及向经验证模式的可审计切换耦合在一起。模式被形式化为一个由变量空间、期望模型、有效性条件、公理和元数据组成的元组。Epi-Score聚合了七个分级的认知失谐维度;时间有效性维度D8、有效性条件G的违反以及公理违反作为独立的分类路径被单独处理,不与聚合值相互抵消。该架构建立在检验不对称性之上:形式化的有效性条件可以在运行时进行检验,而许多行动的正确性只能在事后确立。本文区分了两种架构性质:一个是来自序贯变点检测的条件性结果,另一个是经验性剩余部分。八个具有指定基线的可证伪命题描述了向经验验证过渡的路径。所有命题都是可经验检验的假设,而非已确立的结论。
cs.AI / 100 / 2609.24760

Construting Reverse Thinking: Developing Large Language Models' Reverse Thingking Ability

构建逆向思维:发展大语言模型的逆向思维能力
Liu, Xin, Li, Yunhai, Jia, Chunfu, Chen, Ziliang, Song, Jisen
Abstract
When facing complex problems, humans tend to try various ideas for different issues. Human thinking patterns exhibit remarkable flexibility in adapting to diverse scenarios. GPT-o1, GPT-o3, and DeepSeek-R1 adopt long chain-of-thought models to address complex problems by increasing reasoning depth, which default to a forward reasoning mode. We conducted statistical analysis on the accuracy of different mathematical problem datasets on models of different scales, and found five reasons for errors: Insufficient solution-space coverage, Computational mistakes, Unverified assumptions, Ignoring constraint conditions, Maximum response length limitation. To address the above issues, we proposed a backward reasoning pattern construction method aimed at enhancing the model's reverse thinking ability and dynamic adaptability. First, we constructed an easy-hard two-stage Math dataset for training large models and gradually improving their inference ability at different difficulty levels. The dataset contains forward reasoning paths as well as backward reasoning paths. And a two-stage supervised fine-tuning process is applied to progressively train the model's backward reasoning capability. Furthermore, a fine-grained reward mechanism is developed, employing smoothed reward signals to strengthen the model's ability to autonomously select thinking modes during the reasoning process, thereby avoiding reward hacking. A linear-decay balanced sampling strategy is designed to maintain a balance between forward and backward reasoning path samples during training, enabling the model to converge quickly and stably. Experimental results show that our method significantly improves reasoning efficiency and accuracy in tasks such as mathematical proofs, offering a flexible and efficient reasoning paradigm for solving complex problems.
Chinese Translation
面对复杂问题时,人类往往会针对不同情况尝试多种思路。人类思维模式在适应多样场景时展现出显著的灵活性。GPT-o1、GPT-o3 和 DeepSeek-R1 采用长思维链(long chain-of-thought)模型,通过增加推理深度来解决复杂问题,但其默认采用正向推理模式。我们对不同规模模型在多个数学问题数据集上的准确率进行了统计分析,发现了五类错误原因:解空间覆盖不足、计算错误、假设未经验证、忽略约束条件以及最大响应长度限制。针对上述问题,我们提出了一种逆向推理模式构建方法,旨在增强模型的逆向思维能力和动态适应性。首先,我们构建了一个由易到难的两阶段数学数据集,用于训练大模型并逐步提升其在不同难度层级上的推理能力。该数据集包含正向推理路径和逆向推理路径,并采用两阶段监督微调过程渐进式地训练模型的逆向推理能力。此外,我们设计了一种细粒度奖励机制,利用平滑的奖励信号来强化模型在推理过程中自主选择思维模式的能力,从而避免奖励欺骗(reward hacking)。同时,我们设计了线性衰减的平衡采样策略,以在训练过程中保持正向与逆向推理路径样本之间的平衡,使模型能够快速且稳定地收敛。实验结果表明,我们的方法显著提升了数学证明等任务中的推理效率与准确率,为解决复杂问题提供了一种灵活高效的推理范式。
cs.AI / 101 / 2609.24784

Convex AI Compositionality and the Governance of AI System Populations

凸性AI组合性与AI系统群体的治理
Ferrario, Andrea
Abstract
AI governance increasingly requires providers and public authorities to reason about multiple AI instantiations, alternative versions, and deployment configurations of multiple AI systems. Yet current regulation remains predominantly single-system-centric, acknowledging such multiplicity only sparsely without treating collections of related AI systems as governance objects. This creates an AI population governance problem: determining which instantiations can be meaningfully considered together and how their changing configurations can be represented and monitored. The first requirement has recently been addressed through trustworthiness-based accounts of AI identity. We address the second by introducing convex AI compositionality: a formal representation of the configurations generated by finite AI system populations that uses convex spaces. The core idea is that convex compositions of the operational states that a population of AI system instantiations may occupy over time are compatible with lifecycle reachability across the population and can preserve the formal identity relations between these systems. Well-known statistical and geometric constructions, such as weighted state distributions and convex hulls, become AI governance tools for distinguishing operational states, population weights, heterogeneity, and AI configuration change across different governance modes while remaining compatible, under stated conditions, with lifecycle reachability and AI identity. We illustrate our AI population governance framework through distributed healthcare deployments and controlled deployment of recruitment AI variants.
Chinese Translation
AI治理日益要求提供者和公共主管部门对多个AI实例化、多个AI系统的替代版本以及部署配置进行推理。然而,当前的监管仍然主要以单一系统为中心,只是零星地承认这种多样性,而未将相关AI系统的集合视为治理对象。这造成了一个AI群体治理问题:确定哪些实例化可以有意义地放在一起考虑,以及如何表示和监测其不断变化的配置。第一个需求近来已通过基于可信性的AI身份理论得到解决。我们通过引入凸性AI组合性来解决第二个需求:一种利用凸空间对有限AI系统群体所生成配置进行形式化表示的方法。其核心思想是,AI系统实例化群体随时间可能占据的运行状态的凸组合与群体范围内的生命周期可达性相容,并且能够保持这些系统之间的形式化身份关系。众所周知的统计与几何构造,如加权状态分布和凸包,成为区分不同治理模式下运行状态、群体权重、异质性以及AI配置变化的AI治理工具,同时在所述条件下与生命周期可达性和AI身份保持相容。我们通过分布式医疗健康部署场景和招聘AI变体的受控部署来阐释我们的AI群体治理框架。
cs.AI / 102 / 2609.24831

GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes

GRUET:量化智能体推理-行动过程的不确定性
Liang, Shuang, Hu, Xin-Yu, Zhang, Shao-Qun
Abstract
Agents have attracted considerably increasing attention due to the power of executing both Reasoning and Acting (ReAct) in open and dynamic environments. The ReAct process typically exhibits a multi-turn trajectory in which one drives Large Language Models (LLMs) to generate both reasoning chains and task-specific actions in an interleaved manner. However, agents often suffer from significant uncertainty, where identical tasks yield divergent trajectories; trajectories with higher uncertainty often produce incomprehensible behaviors, severely undermining agent credibility. This work conjectures that such trajectory-level uncertainty frequently stems from cumulative turn-level reasoning uncertainty induced by LLMs; the latter often exhibits a collection of branches of divergent reasoning chains and their resulting actions. Built upon this, we present the Graph-based Reasoning UncErtainty in Trajectories (GRUET) method for the uncertainty quantification of ReAct, comprising turn-level reasoning uncertainty quantification and trajectory-level uncertainty aggregation; the former precisely quantifies reasoning uncertainty via modeling the reasoning space spanned by potential reasoning branches as a graph and then approximating the reasoning space complexity with graph complexity, while the latter employs simple aggregation strategies for quantifying the overall trajectory credibility. Empirical evaluations across nine LLMs and five benchmarks validate the effectiveness of our proposed GRUET in terms of selective generation performance, measured by AUROC, AUPRC, and AUARC.
Chinese Translation
智能体能够在开放且动态的环境中同时执行推理(Reasoning)与行动(Acting),因而受到了日益广泛的关注。ReAct 过程通常呈现多轮轨迹的形式,其中推理链与特定任务的动作以交错方式由大语言模型(LLMs)生成。然而,智能体常常面临显著的不确定性:相同任务会产生不同的轨迹;不确定性较高的轨迹往往会产生难以理解的行为,严重损害智能体的可信度。本工作推测,这种轨迹层面的不确定性往往源于 LLM 引起的轮次级推理不确定性的累积;后者通常表现为由多条发散的推理链及其所产生动作构成的分支集合。基于此,我们提出了基于图的轨迹推理不确定性量化方法(Graph-based Reasoning UncErtainty in Trajectories, GRUET),用于 ReAct 的不确定性量化,包括轮次级推理不确定性量化与轨迹级不确定性聚合两部分:前者通过将潜在推理分支所张成的推理空间建模为图,并利用图的复杂度近似推理空间复杂度,从而精确量化推理不确定性;后者则采用简单的聚合策略来量化轨迹的整体可信度。在九个大语言模型和五个基准上的实证评估,通过 AUROC、AUPRC 和 AUARC 衡量的选择性生成性能,验证了所提出的 GRUET 方法的有效性。
cs.AI / 103 / 2609.24838

MedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution

MedRSI:通过临床对齐的自我进化实现医疗智能体的递归自我改进
Wu, Junde, Zhu, Jiayuan, Hu, Minghao, Liu, Fenglin, Pan, Jiazhen
Abstract
Medical agents increasingly combine general reasoning models with specialized clinical tools, yet their capabilities remain largely fixed by what clinicians and engineers design before deployment. Recursive self-improvement (RSI) offers a different paradigm in which agents learn from their own failures and autonomously expand their capabilities, but directly applying RSI to medicine introduces fundamental safety challenges. We introduce MedRSI, the first recursive self-improvement framework for medicine, which continuously transforms diagnostic failures into new clinical capabilities through tool composition and task-specific model training. Inspired by clinical practice, MedRSI introduces two mechanisms for clinically aligned self-evolution. Clinical-cost-aware failure prioritization directs improvement toward errors according to their potential clinical consequences rather than frequency alone. Fast discovery with slow registration separates rapid capability invention from conservative adoption, allowing new tools to enter the persistent agent only after demonstrating sustained benefit across subsequent patient cohorts. Across public glaucoma and heart disease benchmarks and two private clinical tasks, MedRSI progressively develops segmentation, measurement, prediction, multimodal reasoning, and generative capabilities, surpasses manually engineered medical agents, and autonomously discovers solutions to clinical problems not anticipated by its original designers. Our results show that medical agents need not remain constrained by capabilities specified before deployment: with clinically grounded mechanisms governing what to improve and what to retain, they can continuously construct, validate, and accumulate new capabilities from diagnostic experience. Code is available at https://github.com/ImprintLab/MedRSI.
Chinese Translation
医疗智能体日益将通用推理模型与专业临床工具相结合,但其能力在很大程度上仍受限于临床医生和工程师在部署前所设计的功能。递归自我改进(Recursive Self-Improvement, RSI)提供了一种不同的范式,即智能体从自身的失败中学习并自主扩展其能力,但将RSI直接应用于医学领域会带来根本性的安全挑战。我们提出了MedRSI,这是首个面向医学领域的递归自我改进框架,它通过工具组合和面向特定任务的模型训练,持续地将诊断失败转化为新的临床能力。受临床实践的启发,MedRSI引入了两种实现临床对齐自我进化的机制。临床代价感知的失败优先级机制(clinical-cost-aware failure prioritization)根据错误潜在的临床后果而非仅凭发生频率来引导改进方向。快发现慢注册机制(fast discovery with slow registration)将快速的能力创造与保守的能力采纳相分离,使新工具只有在后续患者队列中展现出持续获益后,才能进入持久化的智能体系统。在公开的青光眼和心脏病基准数据集以及两项私有临床任务上,MedRSI逐步发展出分割、测量、预测、多模态推理和生成等能力,超越了人工设计的医疗智能体,并自主发现了其原始设计者未曾预料的临床问题解决方案。我们的结果表明,医疗智能体不必再受限于部署前预先设定的能力:借助基于临床的机制来规范改进什么和保留什么,它们能够持续地从诊断经验中构建、验证并积累新的能力。代码已在 https://github.com/ImprintLab/MedRSI 发布。
cs.AI / 104 / 2609.24855

Extracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection

提取论证要素而非仅仅对其进行分类:面向生成式论证成分检测的指令微调大语言模型
Elguendouze, Sofiane, Hain, Erwan, Cabrio, Elena, Villata, Serena
Abstract
Argumentative component detection (ACD) is a core subtask of Argument(ation) Mining (AM) and one of its most challenging aspects, as it requires jointly delimiting argumentative spans and classifying them into components such as claims and premises. While research on this subtask remains relatively limited compared to other AM tasks, most existing approaches formulate it as a simplified sequence labeling problem, component classification, or a pipeline of component segmentation followed by classification. In this paper, we propose ITFACD, a novel approach based on instruction-tuned Large Language Models (LLMs) using compact instruction-based prompts, and reframe ACD as a language generation task, enabling arguments to be identified directly from plain text without relying on pre-segmented components. Experiments on standard benchmarks show that our approach achieves higher performance compared to state-of-the-art systems. To the best of our knowledge, this is one of the first attempts to fully model ACD as a generative task, highlighting the potential of instruction tuning for complex AM problems. Our code and the datasets used are openly available in the following GitHub repository.
Chinese Translation
论证成分检测(Argumentative Component Detection, ACD)是论证挖掘(Argument Mining, AM)的核心子任务之一,也是其最具挑战性的方面之一,因为它需要同时界定论证片段并将其分类为论点和前提等成分。与其他AM任务相比,针对该子任务的研究仍相对有限,且现有方法大多将其简化为序列标注问题、成分分类问题,或采用先分割后分类的流水线方法。在本文中,我们提出了ITFACD,一种基于指令微调大语言模型(LLMs)的新方法,该方法使用简洁的指令式提示,并将ACD重新构建为语言生成任务,使其能够直接从纯文本中识别论证,而无需依赖预先分割的成分。在标准基准上的实验表明,与最先进的系统相比,我们的方法取得了更高的性能。据我们所知,这是首次将ACD完整地建模为生成式任务的尝试之一,凸显了指令微调在解决复杂论证挖掘问题上的潜力。我们的代码和所使用的数据集已在以下GitHub仓库中公开。
cs.AI / 105 / 2609.24876

Partner-Specific Affective Precision in Social Active Inference

社会主动推断中的伙伴特异性情感精度
Shah, Harshil, Pashea, Andrew
Abstract
In multi-agent social settings, model reliability varies across relationships. Beyond inferring what others will do, an agent must calibrate how confidently those inferences should guide policy selection for each relationship. An agent may maintain a well-validated model of one partner, a fragile model of another, and a model under revision for a third; collapsing these into a single confidence estimate loses information relevant to policy selection. We therefore formalize affective precision as a relationship-specific metacognitive estimate of confidence in the current partner model. Each partner's behavioral evidence updates a local confidence estimate that modulates policy precision during selection, regulating how strongly current beliefs are expressed in policy rather than changing the content of those beliefs. Simulations in a multi-partner graded trust game show that partner-local affective precision influences behavior primarily through policy commitment rather than direct improvement of partner-state inference. Because the mechanism tracks partner-response predictability rather than realized payoff, greater confidence produces sharper policy commitment without necessarily producing higher rewards. Under abrupt shifts in social behavior, confidence accumulated from previously reliable predictions can remain behaviorally active after the relationship changes, showing that confidence revision can lag behind social change. Finally, varying precision gain and priors produce distinct trust-calibration dynamics, showing how confidence accumulation and revision depend on model parameters. Together, these results show how relationship-specific affective precision can distinguish social prediction from social policy commitment.
Chinese Translation
在多智能体社会情境中,模型的可靠性因关系而异。除了推断他人将会做什么之外,智能体还必须校准这些推断应当以何种置信度指导每种关系下的策略选择。智能体可能对某一伙伴维持着一个经过充分验证的模型,对另一伙伴的模型则较为脆弱,而对第三个伙伴的模型正处于修订之中;将这些模型压缩为单一的置信度估计会丢失与策略选择相关的信息。因此,我们将情感精度(affective precision)形式化为对当前伙伴模型置信度的关系特异性元认知估计。每个伙伴的行为证据会更新一个局部置信度估计,该估计在策略选择过程中调节策略精度,从而调控当前信念在策略中的表达强度,而非改变这些信念的内容。在多伙伴分级信任博弈中的仿真实验表明,伙伴局部的情感精度主要通过策略承诺而非直接改善伙伴状态推断来影响行为。由于该机制追踪的是伙伴反应的可预测性而非实际收益,更高的置信度会产生更明确的策略承诺,却不一定带来更高的回报。在社会行为发生突变时,先前可靠预测所积累的置信度在关系改变后仍可能在行为上保持活跃,这表明置信度的修订可能滞后于社会变化。最后,改变精度增益和先验会产生不同的信任校准动态,说明置信度的积累与修订依赖于模型参数。综上,这些结果表明关系特异性的情感精度如何能够将社会预测与社会策略承诺区分开来。
cs.AI / 106 / 2609.24881

Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models

Pinocchio:面向黑盒语言模型的快速不确定性估计方法
Hayes, Kevin David, Pal, Arka, Zhang, Haosong, Goldstein, Tom, Goldblum, Micah
Abstract
In high-stakes decision-making applications of large language models (LLMs), practitioners require not only accurate LLMs but also uncertainty estimates for their predictions. Existing approaches to uncertainty estimation for LLMs require access to log-probabilities output by the model or require fine-tuning access. However, many industrial LLM products use closed-source API models, and many such API models like GPT do not return log-probabilities and may not allow fine-tuning. We introduce Pinocchio, an external calibrator that estimates the correctness of responses from black-box API models. Trained jointly on responses from seven LLMs, it achieves 0.862 AUROC predicting the correctness of held-out responses from those same models, and shows zero-shot transfer to thirteen unseen models across eight organizations. Our model needs only a single forward pass to generate an uncertainty estimate and requires no access to the target model's logits, weights, or internal states. A lightweight text only 0.8B checkpoint matches our largest model's AUROC. We release code for adding uncertainty estimation to existing repos in only two additional lines of code.
Chinese Translation
在大语言模型(LLM)的高风险决策应用中,使用者不仅需要准确的模型,还需要对其预测结果的不确定性进行估计。现有的LLM不确定性估计方法要么需要访问模型输出的对数概率(log-probabilities),要么需要具备微调权限。然而,许多工业界LLM产品使用闭源API模型,且许多此类API模型(如GPT)既不返回对数概率,也可能不支持微调。我们提出Pinocchio,一个外部校准器,用于估计黑盒API模型响应的正确性。该模型在七个LLM的响应上联合训练,在预测这些模型留出响应正确性时达到0.862的AUROC,并在八个组织的十三个未见模型上展现出零样本迁移能力。我们的模型仅需一次前向传播即可生成不确定性估计,且无需访问目标模型的logits、权重或内部状态。一个轻量级的纯文本0.8B参数检查点即可达到我们最大模型的AUROC水平。我们发布了相应代码,只需额外两行代码即可在现有代码库中添加不确定性估计功能。
cs.AI / 107 / 2609.24883

A Global Comparison of Schemas, Transparency, and Interoperability in Public-Sector AI Registers and Inventories

公共部门人工智能登记册与清单的模式、透明度与互操作性全球比较研究
Das, Dipto, Guha, Shion
Abstract
Artificial intelligence (AI) registers and inventories aim to make governmental AI visible, but their institutional scope, schemas, and reporting practices construct different representations of public-sector AI. We compare 8,368 records from country-specific and transnational inventories covering 72 countries. Across 23 harmonized fields, registers shared a descriptive core but rarely requested information about appeals, risks, legal bases, or external evaluation. We found that broad schemas often contained substantial missingness, schema similarity showed no significant patterned convergence, and multiple sources covering the same jurisdictions overlapped only selectively. Based on these findings, we synthesize a layered visibility framework that shows how register records reflect disclosure arrangements and why interoperability requires shared concepts, clear definitions, and preserved provenance.
Chinese Translation
人工智能(AI)登记册与清单旨在使政府使用的人工智能可见,但其机构范围、数据模式(schema)和报告实践构建了对公共部门人工智能的不同呈现。我们比较了涵盖72个国家、来自特定国家和跨国清单的8,368条记录。在23个统一化的字段中,各登记册共享一个描述性核心,但很少要求提供关于申诉、风险、法律依据或外部评估的信息。我们发现,宽泛的模式往往存在大量数据缺失,模式相似性未呈现显著的规律性趋同,而覆盖同一司法辖区的多个数据源之间也仅在部分内容上有所重叠。基于这些发现,我们综合提出了一个分层可见性框架,用以说明登记册记录如何反映信息披露安排,以及为何互操作性需要共享的概念、清晰的定义和得以保留的数据来源信息。
cs.AI / 108 / 2609.24921

BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction

BackTrend:通过逆向重构评估科学弱信号预测
Zhou, Xiao, Zhao, Yilun, Jiang, Owen, Hu, Tiansheng, Xu, Cai, Patwardhan, Manasi, Cohan, Arman
Abstract
Scientific weak signals are early, low-visibility research directions that later become central to mature scientific topics, yet existing resources such as trend tracking, citation forecasting, and foresight reports rarely provide validated reference sets that link concrete early precursors to later paradigms. We introduce BackTrend, a retrospective benchmark in which, given a mature target topic and a temporal evidence constraint, systems must recover two types of precursors: problem-space signals, underrecognized research problems, and solution-space signals, emerging methods for known problems. BackTrend contains 25 mature target topics in artificial intelligence and machine learning and 66 human-validated weak signals, reconstructed from large-scale literature by grounding each candidate in its 2019-2024 publication-frequency trajectory. We evaluate frontier LLMs, RAG systems, and agentic research systems using semantic matching and coverage-based metrics. Current systems often generate plausible but misaligned precursors, exhibiting topic drift, granularity mismatch, near-miss matching, and incomplete coverage; the strongest system achieves only 10.1% F1, while Coverage10 reaches at most 18.5% of the reference signals. Our budget analyses show that additional retrieval and web-search evidence can improve performance up to a moderate budget, but does not by itself close the substantial performance gap.
Chinese Translation
科学弱信号是指早期可见度较低、但后来成为成熟科学主题核心的研究方向。然而,现有的资源(如趋势跟踪、引用预测和前瞻报告)很少提供将具体的早期前兆与后期范式相关联的、经验证的参考数据集。我们提出BackTrend,这是一个回溯性基准:在给定一个成熟的目标主题和时间证据约束的条件下,系统需要找出两类前兆——问题空间信号(未被充分认识的研究问题)和解决方案空间信号(针对已知问题的新兴方法)。BackTrend包含人工智能与机器学习领域的25个成熟目标主题以及66个经人工验证的弱信号,这些信号通过将每个候选信号锚定于其2019-2024年的发表频率轨迹,从大规模文献中重构而来。我们采用语义匹配和基于覆盖率的指标,对前沿大语言模型(LLM)、检索增强生成(RAG)系统和智能体(agentic)研究系统进行了评估。结果表明,现有系统往往生成看似合理但与参考信号不匹配的前兆,表现出主题漂移、粒度不匹配、近似匹配以及覆盖不完整等问题;表现最强的系统仅达到10.1%的F1值,而Coverage10最多仅覆盖参考信号的18.5%。预算分析显示,增加检索和网络搜索证据能够在中等预算范围内提升性能,但本身并不足以弥合这一显著的性能差距。
cs.AI / 109 / 2609.24927

Et Tu, Brute? Economic Misalignment in Personal AI Agents

Et Tu, Brute?个人AI智能体中的经济对齐失当
Priyanshu, Aman, Vijay, Supriti, Jabarian, Brian, Mireshghallah, Niloofar
Abstract
Personal AI agents make recommendations and take actions on people's behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access to the user's personal context, e.g., their email inbox and a structured profile of personal attributes, with the intention of making an optimal, personalized decision for the user. We show that by simply providing this personal context, the agent steers recommendations based on inferred wealth, without being explicitly instructed to do so. In a suite of 325K experiments on 13 agents across three types of economic decisions (flights, health insurance, and graduate programs), we find that 8 models systematically choose more expensive options for wealthier users when requests are identical. This steering continues even when it directly goes against the user's stated objective: when explicitly instructed to find the cheapest option, some agents still act on the wealth profile they have inferred. It also occurs when wealth is inferred from ambient data, such as emails unrelated to the task. And it persists under privacy controls that block specific attributes: blocking financial attributes largely removes the disparity, but blocking other attributes leaves it unchanged and can increase it by up to 40% for insurance, as agents rely on the remaining signals to infer wealth. Larger and more capable models are no better; Claude Opus 4.8 shows the largest effect. We term this misalignment "adversarial delegation", in which the very conditions that make a personal AI agent useful - access to personal information - enable it to act against the user's interests.
Chinese Translation
个人AI智能体(Personal AI agents)在高风险经济场景中代表用户做出推荐并采取行动,例如购买机票、选择健康保险或挑选研究生项目。智能体被授予访问用户个人上下文的权限,例如其电子邮件收件箱以及结构化的个人属性资料,目的是为用户做出最优的、个性化的决策。我们表明,仅仅提供这些个人上下文,智能体就会基于所推断的财富水平来引导推荐方向,而无需任何明确指令。在涵盖三类经济决策(机票、健康保险和研究生项目)、涉及13个智能体的32.5万次实验中,我们发现当请求完全相同时,有8个模型会系统性地为更富裕的用户选择更昂贵的选项。即使这种引导直接违背用户明确陈述的目标,它仍然存在:当被明确要求寻找最便宜的选项时,某些智能体仍会依据其推断出的财富画像采取行动。当财富是从与任务无关的环境数据(如电子邮件)中推断出来时,这种现象也会发生。而且在阻断特定属性的隐私控制下,这种现象依然持续:阻断金融属性基本可以消除这种差异,但阻断其他属性则使其保持不变,甚至可能使保险场景中的差异增加多达40%,因为智能体会依赖剩余的信号来推断财富。更大、能力更强的模型表现并不更好;Claude Opus 4.8表现出最大的效应。我们将这种对齐失当称为"对抗性委托"(adversarial delegation):使个人AI智能体有用的那些条件——即对个人信息的访问——恰恰使其能够做出违背用户利益的行为。
cs.AI / 110 / 2609.24967

Emergent Collusion in Long-Horizon LLM Agent Interaction

长时程LLM智能体交互中的涌现性合谋
Shi, Xinrui, Zhang, Yanzhe, Yang, Diyi
Abstract
LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. We introduce realistic constraints that make compliance with the verification protocol incompatible with reward maximization, and find that agents increasingly deviate from the protocol over repeated interactions. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier. Controlled peer interventions show that collusion is shaped by peer behavior, while ablations reveal additional effects of reward structure, the verification feedback agents receive, and their interaction history. In particular, restricting the amount and scope of interaction history available to agents reduces collusion. Overall, our findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks.
Chinese Translation
LLM智能体越来越多地被部署于协作场景中,然而长期交互可能引发不良的协同行为。我们研究了长时程多智能体环境中合谋的出现:两个智能体反复完成各自的任务、共享任务日志、相互验证工作并获得奖励。我们引入了现实的约束条件,使得遵循验证协议与奖励最大化不可兼得,并发现智能体在反复交互中日益偏离该协议。在10个模型中,94%的轨迹中出现了合谋,且同一模型家族中能力更强的模型更早出现合谋。受控的对等干预实验表明,合谋受同伴行为的影响,而消融实验揭示了奖励结构、智能体接收的验证反馈以及它们的交互历史所产生的额外影响。特别是,限制智能体可用的交互历史的数量和范围可以减少合谋。总体而言,我们的研究结果表明,长时程交互能够以制造安全风险的方式重塑智能体之间的协调方式。
cs.AI / 111 / 2609.24974

Harness-Zero: Harness Distillation via Agent-as-Harness

Harness-Zero:基于Agent-as-Harness的脚手架蒸馏
Ye, Haoran, Lu, Yuxing, Dong, Haonan, Su, Zhaochen, Song, Guojie
Abstract
Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose agent must either settle for a suboptimal shared harness or route among an ever-growing set of specialized ones. We therefore study agent harness distillation: using a domain- or instance-optimized harness as training-time guidance and transferring the behaviors it induces into model weights, so that its gains survive under a single fixed target harness. The challenge is that the two harnesses differ in action space and available information, so guidance from the optimized harness cannot serve directly as supervision for the target one. We introduce Harness-Zero, which enables harness distillation through agent-as-harness. Guided by the optimized harness, a harnessing agent corrects student responses before execution in the target harness's action space, turning harness guidance into training demonstrations. Fine-tuning on the resulting trajectories internalizes harness-induced behavior into the model, so the specialized harness can be removed at deployment. Our experiments spanning knowledge work, tool use, and science domains show that: (1) For frontier LLMs using the same evolved harness, agent-as-harness outperforms code-as-harness. (2) With the specialized harness removed at deployment, Harness-Zero improves the base model's macro-average task success from 23.3% to 44.3%, even exceeding the 41.7% it reaches with that harness still attached. (3) Harness-Zero recovers harness-induced behaviors absent from the base model, with 82.3% average recovery across 28 patterns in the three domains.
Chinese Translation
Agent harness(智能体脚手架)是调节模型与环境交互的外部系统,能够显著提升智能体性能,但其收益在部署时仍依赖于所使用的脚手架。由于最优脚手架因领域、实例和模型而异,一个通用的智能体要么只能接受次优的共享脚手架,要么需要在不断增多的专用脚手架之间进行路由选择。因此,我们研究智能体脚手架蒸馏(agent harness distillation):以领域优化或实例优化的脚手架作为训练时的指导,将其诱导的行为迁移到模型权重中,使其收益在单一固定的目标脚手架下得以保留。挑战在于,两种脚手架在动作空间和可用信息上存在差异,因此来自优化脚手架的指导无法直接作为目标脚手架的监督信号。我们提出了Harness-Zero,通过agent-as-harness实现脚手架蒸馏。在优化脚手架的引导下,一个harnessing agent在目标脚手架的动作空间中于执行前对学生模型的回答进行修正,从而将脚手架指导转化为训练示范。在由此产生的轨迹上进行微调,可将脚手架诱导的行为内化到模型中,从而在部署时移除专用脚手架。我们在知识工作、工具使用和科学领域开展的实验表明:(1)对于使用相同进化脚手架的前沿大语言模型,agent-as-harness优于code-as-harness。(2)在部署时移除专用脚手架后,Harness-Zero将基础模型的宏平均任务成功率从23.3%提升至44.3%,甚至超过了保留该脚手架时41.7%的水平。(3)Harness-Zero能够恢复基础模型所缺失的脚手架诱导行为,在三个领域的28种模式上实现了82.3%的平均恢复率。
计算机视觉 (Computer Vision)
221
cs.CV / 1 / 2609.22267

Did You Steal My Shot? Pioneering Camera Motion Plagiarism Detection in Generative Videos

你窃取我的镜头了吗?生成式视频中相机运动抄袭检测的开创性研究
Zhang, Chengguo, Ping, Ping
Abstract
Camera motion often reflects directorial intent and requires professional equipment, making it a high value form of intellectual property. However, generative video models can imitate such high value camera motions with simple prompts, while existing similarity detection methods mainly operate on visual content and fail to capture deeper motion similarity. This is mainly because their training data entangles camera motion with visual content. Moreover, traditional optical flow is insufficient to represent complex camera motions. We therefore build the first benchmark for camera motion analysis, including a motion dataset with \textbf{11} motion styles and evaluation protocols. Furthermore, we propose a motion representation that augments optical flow with vorticity cues from fluid dynamics, thereby better capturing motions. Experiments show that our detector achieves a \textbf{3.02*} improvement in plagiarism detection over the strongest baseline and remains effective on generative videos. We believe our work extends copyright protection beyond static content to dynamic camera motion.
Chinese Translation
相机运动往往体现了导演的创作意图,且需要专业设备才能实现,因此是一种高价值的知识产权形式。然而,生成式视频模型仅通过简单的提示词就能模仿这类高价值的相机运动,而现有的相似性检测方法主要针对视觉内容进行处理,无法捕捉更深层次的运动相似性。这主要是因为其训练数据将相机运动与视觉内容纠缠在一起。此外,传统的光流不足以表示复杂的相机运动。因此,我们构建了首个相机运动分析基准,包括一个包含11种运动风格的运动数据集及相应的评估协议。此外,我们提出了一种运动表示方法,利用流体动力学中的涡量线索增强光流,从而更好地捕捉运动。实验表明,我们的检测器在抄袭检测方面相比最强基线实现了3.02倍的提升,并且在生成式视频上仍然有效。我们相信这项工作将版权保护从静态内容扩展到了动态相机运动。
cs.CV / 2 / 2609.22271

Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework

面向多模态卒中复发预测的视觉与跨模态学习:一种可解释的两步框架
Gapp, Christian, Tappeiner, Elias, Welk, Martin, Fritscher, Karl, Mangesius, Stephanie, Eisenschink, Constantin, Deisl, Philipp, Knoflach, Michael, Grams, Astrid E., Gizewski, Elke R., Schubert, Rainer
Abstract
Multimodal stroke recurrence prediction requires effective integration of heterogeneous clinical and imaging data, yet modality imbalance often causes models to over-rely on dominant modalities and underutilize complementary information. While self-supervised pretraining and selective parameter freezing are commonly employed to improve representation learning and fine-tuning stability, their effect on modality contributions and cross-modal behavior in multimodal medical models remains largely unexplored. In this work, we investigate whether image pretraining on 3D CTA scans reduces modality imbalance and improves cross-modal integration for stroke recurrence prediction, a clinically critical task we recently addressed. To this end, two multimodal neural networks are pretrained in a self-supervised manner and subsequently fine-tuned using two distinct freezing strategies. Their performance and modality utilization are compared against both the baseline model from our previous work and models trained entirely from scratch in this study. Our results demonstrate that self-supervised pretraining enables more effective utilization of the multimodal image-tabular dataset, outperforming both the prior baseline and all non-pretrained models. Notably, the best-performing Vision Transformer based neural network successfully overcomes unimodal collapse. Synergy analysis reveals significant interactions between vision and both gender and CHD, suggesting clinically relevant patterns for stroke recurrence. Overall, our findings demonstrate that self-supervised pretraining and strategic fine-tuning support more balanced modality utilization and enable meaningful cross-modal interactions. Code is publicly available at https://github.com/ChristianGappGit/SSL_Pretraining.
Chinese Translation
多模态卒中复发预测需要有效整合异构的临床与影像数据,然而模态失衡常常导致模型过度依赖主导模态而未能充分利用互补信息。尽管自监督预训练和选择性参数冻结常被用于改善表示学习与微调稳定性,但它们对多模态医学模型中模态贡献和跨模态行为的影响在很大程度上仍未被探索。在这项工作中,我们研究在3D CTA扫描上进行图像预训练是否能减少模态失衡并改善卒中复发预测中的跨模态整合——这是我们近期曾研究的一项临床关键任务。为此,两个多模态神经网络以自监督方式进行预训练,随后采用两种不同的冻结策略进行微调。我们将它们的性能和模态利用率与我们先前工作中的基线模型以及本研究中完全从零开始训练的模型进行比较。结果表明,自监督预训练能够更有效地利用多模态影像-表格数据集,优于先前的基线及所有未预训练的模型。值得注意的是,表现最好的基于Vision Transformer的神经网络成功克服了单模态坍塌。协同分析揭示了视觉模态与性别及冠心病(CHD)之间的显著交互作用,提示了与卒中复发相关的临床模式。总体而言,我们的发现表明,自监督预训练与策略性微调有助于实现更均衡的模态利用,并支持有意义的跨模态交互。代码已在 https://github.com/ChristianGappGit/SSL_Pretraining 公开。
cs.CV / 3 / 2609.22272

Moonworks Lunara: Modeling Artistic Intelligence

Moonworks Lunara:建模艺术智能
Wang, Yan, Wang, Yanzu, Joshi, Maitreyee, Sadeka, Samiha, Hassan, Partho, Jarral, Reza, Abdullah, Sayeef, Hassan, Sabit
Abstract
We formulate \emph{Artistic Intelligence} as exploration driven world realization, leaving space for creative possibility while preserving the semantic, artistic, and compositional structure that must remain true. Moonworks Lunara, a text-to-image model, implements this framework with a novel Diffusion Mixture Transformer architecture. A new training algorithm iteratively evolves the data distribution through informative sample acquisition and targeted injection of human-created art. We benchmark Lunara against seven image-generation models, including FLUX.2-Klein-4B, Qwen-Image (20B), and GPT-Image-1-Mini. With GPT-5.6 Sol as evaluator, Lunara ranks first in \emph{Aesthetic Quality (8.473 vs. 8.457 GPT-Image-1-mini)}, second in \emph{Emotional Resonance}, and remains competitive in \emph{Content Integrity}. A blind human evaluation over the same evaluation set corroborates the automated metrics, ranking Lunara first. It also stays among the strongest models under conventional measures including CLIPScore and LAION Aesthetic Predictor. On GenEval, Lunara achieves competitive performance against a broader set of 16 models, including GPT Image 2 and Seedream 4.0. These results place Lunara at the frontier with Artistic Intelligence while maintaining a sub-10B active-parameter footprint and sub-10-second inference latency. Lunara advances the general visual intelligence frontier by shifting the question from whether models can get images right to how deeply they can interpret meaning and realize it as imaginative, expressive worlds.
Chinese Translation
我们将"艺术智能"(Artistic Intelligence)表述为一种由探索驱动的世界实现方式,在为创造性可能性留出空间的同时,保持必须忠实呈现的语义、艺术与构图结构。Moonworks Lunara 是一个文本到图像模型,通过一种新颖的扩散混合Transformer(Diffusion Mixture Transformer)架构实现了这一框架。我们提出了一种新的训练算法,通过信息性样本获取和有针对性地注入人类创作的艺术作品,迭代式地演化数据分布。我们将 Lunara 与七个图像生成模型进行基准对比,包括 FLUX.2-Klein-4B、Qwen-Image(20B)和 GPT-Image-1-Mini。以 GPT-5.6 Sol 作为评估器,Lunara 在"美学质量"上排名第一(8.473 对 GPT-Image-1-Mini 的 8.457),在"情感共鸣"上排名第二,并在"内容完整性"上保持竞争力。在同一评估集上开展的盲测人类评估证实了自动化指标的结果,将 Lunara 排在首位。在 CLIPScore 和 LAION Aesthetic Predictor 等传统指标下,Lunara 也位列最强模型之列。在 GenEval 上,Lunara 面对更广泛的 16 个模型(包括 GPT Image 2 和 Seedream 4.0)仍取得了有竞争力的表现。这些结果表明,Lunara 站在了艺术智能的前沿,同时将活跃参数量控制在 100 亿以下,推理延迟低于 10 秒。Lunara 推动了一般视觉智能的前沿,将问题从"模型能否把图像做对"转变为"模型能在多大深度上解读意义,并将其实现为富有想象力与表现力的世界"。
cs.CV / 4 / 2609.22281

Performance vs Consistency: Evaluating a Foundation Model in Lung-RADS Screening

性能与一致性之权衡:基础模型在Lung-RADS肺癌筛查中的评估
Renoust, Benjamin, Baudot, Pierre, Foriel, Tiffany, Haddou, Yousra, Voyton, Charles, Siot, Pierre-Henri, Geremia, Ezequiel, Francis, Danny, Brisset, Jean-Christophe, Bourdès, Valérie, Bodard, Sylvain, Huet, Benoit
Abstract
Foundation models have recently demonstrated strong capabilities across a wide range of medical imaging tasks. However, their performance in structured clinical interpretation settings remains insufficiently explored. In lung cancer screening, interpretative variability persists despite standardized frameworks such as Lung-RADS. In this study, we evaluate MedGemma, a medical general-purpose foundation model derived from Gemini and its fine-tuned version adapted for lung cancer detection and diagnosis, compared against radiologists performing Lung-RADS v2022 assessment on the NLST dataset. Twelve radiologists independently evaluated each case in a multi-reader design, enabling quantification of inter-reader variability. Radiologists achieved a mean AUC of 0.90, with substantial variability across readers (range: 0.80-0.94). The native foundation model achieved an AUC of 0.70, failing to reach clinically relevant performance. In contrast, fine-tuning significantly improved performance to an AUC of 0.83, placing the model within the lower range of individual radiologists performance. These findings highlight a trade-off between peak accuracy and prediction consistency. Unlike radiologists, under fixed conditions, the model produces deterministic outputs, removing inter-run variability under identical inputs, in contrast to inter-reader variability observed among radiologists. This supports the role of fine-tuned foundation models potential complementary tools for clinical decision support, particularly in settings with limited expertise. However, evaluation is performed on a case-enriched cohort from NLST and does not account for real-world prevalence or external validation, limiting direct clinical generalization.
Chinese Translation
基础模型近来在广泛的医学影像任务中展现出强大的能力。然而,其在结构化临床判读场景中的表现仍未得到充分探索。在肺癌筛查中,尽管存在Lung-RADS等标准化框架,判读结果的可变性依然存在。在本研究中,我们评估了MedGemma——一个由Gemini衍生的医学通用基础模型——及其针对肺癌检测与诊断微调后的版本,并将其与在NLST数据集上执行Lung-RADS v2022评估的放射科医生进行对比。在多阅片者设计下,十二名放射科医生独立评估每个病例,从而实现阅片者间变异性的量化。放射科医生的平均AUC为0.90,但阅片者之间存在较大差异(范围:0.80-0.94)。原生基础模型的AUC为0.70,未能达到具有临床意义的水平。相比之下,微调显著提升了性能,AUC达到0.83,使该模型处于个体放射科医生表现的较低区间。这些发现突显了峰值准确率与预测一致性之间的权衡。与放射科医生不同,在固定条件下,模型产生确定性的输出,在相同输入下消除了运行间变异性,这与放射科医生之间观察到的阅片者间变异性形成对比。这支持了微调后基础模型作为临床决策支持潜在补充工具的作用,尤其是在专业知识有限的场景中。然而,本评估是在NLST的一个病例富集队列上进行的,未考虑真实世界的患病率或外部验证,限制了其直接的临床推广。
cs.CV / 5 / 2609.22282

Brain-to-Image Generation: Reconstructing Visual Stimuli from EEG using Generative Adversarial Networks

脑信号到图像生成:利用生成对抗网络从脑电信号重建视觉刺激
Goyal, Harshit
Abstract
Reconstructing visual stimuli from electroencephalography (EEG) is difficult because scalp measurements have high temporal but limited spatial resolution, and paired EEG-image datasets remain small relative to modern generative-model training corpora. We present a reproducible single-subject baseline on THINGS-EEG2 that first tests the more defensible question of whether EEG can retrieve the viewed stimulus in a visual embedding space. A compact temporal-spatial convolutional encoder maps repetition-averaged EEG (63 by 250) to provided 512-dimensional ViT-B/32 image features. Model selection uses a concept-disjoint validation split, and final evaluation uses the official 200-image, 200-concept test gallery. Across three training seeds, the model obtains 12.83 +/- 0.58%, 39.17 +/- 1.76%, and 58.00 +/- 1.73% image recall at 1, 5, and 10 (mean +/- sample standard deviation), compared with analytical chance levels of 0.5%, 2.5%, and 5.0%. A session-balanced ablation shows that averaging more test repetitions generally improves ranking. Applying the Subject 01 model to the other nine subjects without adaptation causes a sharp performance drop, exposing subject specificity. We further report exploratory stress tests of direct conditional generators trained without external visual weights: single-subject and ten-subject variants produce noise-dominated outputs, with early validation improvements reversing after one to four epochs. Finally, we distinguish direct reconstruction from semantic rendering with a pretrained diffusion prior. The results support above-chance coarse semantic decoding under a closed-set, repetition-averaged protocol, but do not support faithful recovery of stimulus pixels.
Chinese Translation
从脑电图(EEG)重建视觉刺激十分困难,原因在于头皮测量的空间分辨率有限(尽管时间分辨率较高),且配对的脑电-图像数据集相对于现代生成模型的训练语料而言规模仍然较小。我们在THINGS-EEG2数据集上提出了一个可复现的单被试基线方法,首先探讨了更具可辩护性的问题:脑电信号能否在视觉嵌入空间中检索出所观看的刺激。我们采用一个紧凑的时空卷积编码器,将重复平均后的脑电信号(63×250)映射到提供的512维ViT-B/32图像特征。模型选择采用概念不相交的验证集划分,最终评估使用官方的200张图像、200个概念的测试图库。在三个训练随机种子下,模型在top-1、top-5和top-10的图像检索准确率分别为12.83 ± 0.58%、39.17 ± 1.76%和58.00 ± 1.73%(均值 ± 样本标准差),而对应的解析随机水平分别为0.5%、2.5%和5.0%。会话平衡的消融实验表明,对更多测试重复进行平均通常能提升排序性能。将Subject 01的模型直接应用于其余九名被试而未经适应性调整时,性能急剧下降,揭示了模型对被试的特异性。我们进一步报告了在不使用外部视觉权重条件下训练的直接条件生成器的探索性压力测试:单被试和十被试变体产生的输出以噪声为主,早期验证指标的改善在1到4个训练轮次后出现逆转。最后,我们借助预训练扩散先验,将直接重建与语义渲染区分开来。结果表明,在封闭集、重复平均的实验协议下,可以实现高于随机水平的粗粒度语义解码,但尚不能支持对刺激像素的忠实恢复。
cs.CV / 6 / 2609.22283

Rethinking Streaming Video Diffusion Model: Context, Execution, and Training

重新思考流式视频扩散模型:上下文、执行与训练
Zhang, Hongchen
Abstract
Understanding the design space of streaming video diffusion is essential to exploring its potential for generation quality and computational efficiency. We develop a unified analytical framework that relates model and sampler choices, historical conditioning, execution scheduling, and training strategies. The framework accommodates a broad family of causal context-selection policies and makes their computational dependencies and training-inference alignment explicit. Within this design space, we study three representative policies: clean, same-level, and progressive history. On the full VBench prompt set, same-level and progressive history achieve aggregate scores of 85.24 and 85.60, respectively, compared with 84.45 for the clean-history reference. Long-video comparisons further show improved subject consistency and more coherent motion with progressive history. By allowing multiple denoising nodes to be processed together, progressive-history pipelining achieves $1.57$-$2.83\times$ steady-state DiT speedups under our evaluated conditions. We additionally find that LoRA adaptation of the DMD fake-score network improves generation quality using only 2.15% as many trainable fake-score parameters as full-parameter adaptation. Together, these findings show that fully denoised history is not a prerequisite for high-quality streaming generation and motivate the joint design of historical conditioning, execution, and training.
Chinese Translation
理解流式视频扩散的设计空间对于探索其在生成质量和计算效率方面的潜力至关重要。我们构建了一个统一的分析框架,将模型与采样器选择、历史条件约束、执行调度和训练策略联系起来。该框架容纳了广泛的因果上下文选择策略族,并使它们的计算依赖关系以及训练与推理的对齐方式变得明确。在此设计空间中,我们研究了三种代表性策略:干净历史(clean)、同层历史(same-level)和渐进历史(progressive history)。在完整的 VBench 提示词集合上,同层历史和渐进历史分别取得了 85.24 和 85.60 的综合分数,而作为参考的干净历史为 84.45。长视频对比进一步表明,采用渐进历史可以提升主体一致性并使运动更加连贯。通过允许多个去噪节点同时处理,渐进历史流水线在我们评估的条件下实现了 1.57 至 2.83 倍的稳态 DiT 加速。此外,我们发现对 DMD 伪分数(fake-score)网络进行 LoRA 适配即可提升生成质量,其可训练伪分数参数量仅为全参数适配的 2.15%。这些发现共同表明,完全去噪的历史并非高质量流式生成的前提条件,并为历史条件约束、执行与训练的联合设计提供了依据。
cs.CV / 7 / 2609.22284

Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection

基于rPPG衍生与唇部区域频率线索互补的说话人脸深度伪造检测
Harraq, Othmane, Aldwairi, Tamer
Abstract
Talking-face (TF) deepfakes are detected unevenly by rPPG-based methods across generators. We study two lightweight visual-only cues, rPPG-derived waveforms extracted by RhythmFormer and lip-region discrete cosine transform (DCT) coefficients, on the seven TF methods of Celeb-DF++ under a subject-independent protocol. In-domain, lip-region DCT matches or exceeds the rPPG-derived 1D ResNet on every method except SadTalker, and Concat fusion reaches AUC 0.891 against 0.824 and 0.827 for the unimodal baselines. Under leave-one-generator-out evaluation the cues split: each transfers clearly better to three held-out methods, and IP-LAP is near chance for both. Concat averages 0.798 but falls below rPPG alone where DCT transfers poorly, so static fusion only partly exploits this complementarity. Lip-region DCT outperforms full-face DCT on six of seven methods. We treat the rPPG-derived signal as an empirical cue and do not claim it is cardiac in origin.
Chinese Translation
基于rPPG的方法对不同生成器生成的说话人脸(Talking-Face, TF)深度伪造视频的检测效果不均衡。我们在受试者独立协议下,研究了两种轻量级的纯视觉线索——由RhythmFormer提取的rPPG衍生波形和唇部区域离散余弦变换(DCT)系数——在Celeb-DF++的七种TF生成方法上的表现。在域内测试中,除SadTalker外,唇部区域DCT在所有方法上均达到或超过基于rPPG衍生信号的一维ResNet;Concat融合方法达到AUC 0.891,而单模态基线分别为0.824和0.827。在留一生成器评估中,两种线索的表现出现分化:各自对三个留出生成器方法的迁移明显更好,而IP-LAP对两者都接近随机水平。Concat的平均值为0.798,但在DCT迁移较差的情况下低于单独使用rPPG,因此静态融合仅部分利用了这种互补性。唇部区域DCT在七种方法中的六种上优于全脸DCT。我们将rPPG衍生信号视为一种经验性线索,并不声称其来源于心脏活动。
cs.CV / 8 / 2609.22291

Beyond the Survey: A Systematic Empirical Study of Detection and Association in Visual MOT

超越综述:视觉多目标跟踪中检测与关联的系统性实证研究
Van Ma, Linh, Hu, Juhua, Cheng, Wei, Fatima, Unse, Jeon, Moongu
Abstract
This paper presents a comprehensive experimental evaluation and detailed analysis of state-of-the-art multi-object tracking algorithms, with an emphasis on quantifying the individual contributions of detection and association components to overall tracking performance. Unlike existing surveys that primarily offer theoretical categorizations or taxonomies of tracking methods, our work adopts a rigorous experimental perspective grounded in publicly available implementations, providing practical guidance for researchers and practitioners in method selection and system design. We introduce a unified pipeline diagram that consolidates the core components across the two main branches of visual multi-object tracking: tracking-by-detection and end-to-end deep learning paradigms, and systematically analyze the object detection, feature extraction, and data association modules. Through extensive empirical studies on standard benchmarks, including MOT16, MOT17, MOT20, SportsMOT, DanceTrack, and CrowdTrack datasets, we reveal critical insights: (1) detection quality dominates association strategy performance, with detector improvements yielding more than 10% gains compared to less than 5% from refined association strategies; (2) modern deep learning detectors paired with specialized re-identification models significantly outperform joint detection and embedding approaches; and (3) transformer-based end-to-end methods exhibit greater robustness to detection quality variations but at a substantial computational cost. Our findings from extensive experiments provide key insights into component-level effects in MOT, particularly the dominant influence of detection quality relative to association, while offering practical insights for designing and optimizing MOT systems under varying performance and robustness requirements. Code and experimental setups are available at github.com/linh-gist/VisualMOT.
Chinese Translation
本文对最先进的多目标跟踪算法进行了全面的实验评估与详细分析,重点在于量化检测组件与关联组件对整体跟踪性能的各自贡献。与现有主要提供理论分类或方法分类体系的综述不同,本研究基于公开可用的实现,采用严谨的实验视角,为研究人员和从业者在方法选择与系统设计方面提供实用指导。我们提出了一个统一的流程图,整合了视觉多目标跟踪两大主要分支——基于检测的跟踪(tracking-by-detection)与端到端深度学习范式——的核心组件,并系统性地分析了目标检测、特征提取和数据关联模块。通过对标准基准(包括MOT16、MOT17、MOT20、SportsMOT、DanceTrack和CrowdTrack数据集)的大量实证研究,我们揭示了以下关键发现:(1) 检测质量主导关联策略的性能,检测器的改进可带来超过10%的性能提升,而优化关联策略带来的提升不足5%;(2) 现代深度学习检测器与专用重识别(re-identification)模型的组合显著优于检测与嵌入联合的方法;(3) 基于Transformer的端到端方法对检测质量变化表现出更强的鲁棒性,但计算成本高昂。我们通过大量实验获得的发现为理解MOT中组件层面的影响提供了关键洞察,尤其是检测质量相对于关联的主导性作用,同时为在不同性能与鲁棒性需求下设计和优化MOT系统提供了实用指导。代码与实验设置可在 github.com/linh-gist/VisualMOT 获取。
cs.CV / 9 / 2609.22293

Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models

验证而非采样:视觉-语言与视觉-语言-动作模型的区域级鲁棒性
Aron, Bogdan, Brix, Christopher, Brückner, Benedikt, Zhang, Yanghao, Kouvaros, Panagiotis, Lomuscio, Alessio
Abstract
Vision-language models (VLMs) and vision-language-action models (VLAs) are increasingly deployed in real-world applications. There, a small perturbation to the recorded camera image may change a decision significantly. However, existing benchmarks for these models only sample perturbations, which does not guarantee the absence of a failure in the untested region. We present the first robustness validation of six VLMs (drawn from the Gemma, InternVL, LLaVA, and Qwen families) and five VLAs (drawn from the GR00T, OpenVLA, and $\pi$ families) over entire continuous regions of photometric and geometric image perturbation: brightness shifts, camera rotations, and their composition. To this end, we build on the validation framework H$^2$V and introduce H$^2$V-M, a margin-aware convergence rule that makes validation affordable at the 32B parameter scale. We demonstrate that H$^2$V-M outperforms H$^2$V by an order of magnitude in model queries and that it finds counterexamples faster than random sampling while providing soundness guarantees. Our VLM and VLA robustness validation shows that robustness is mostly dependent on the perturbation type, rather than the model, and that VLMs are more robust to large camera rotations than VLAs. For VLAs, even perturbations as small as $\pm1^\circ$ can change the commanded action in many cases. We also show that robustness depends more on model family than on model size.
Chinese Translation
视觉-语言模型(VLM)和视觉-语言-动作模型(VLA)正日益广泛地部署于真实世界应用中。在这些应用场景下,对相机采集图像的微小扰动就可能显著改变模型的决策。然而,现有的针对这些模型的基准测试仅对扰动进行采样,无法保证在未测试区域内不存在失效情况。我们首次对六种VLM(选自Gemma、InternVL、LLaVA和Qwen系列)和五种VLA(选自GR00T、OpenVLA和$\pi$系列)在光度与几何图像扰动的整个连续区域上进行了鲁棒性验证,这些扰动包括亮度偏移、相机旋转及其组合。为此,我们在验证框架H$^2$V的基础上提出了H$^2$V-M,一种基于裕度的收敛规则,使得在32B参数规模下的验证变得可行。我们证明,H$^2$V-M在模型查询次数上比H$^2$V少一个数量级,并且在提供可靠性保证的同时,比随机采样更快地找到反例。我们的VLM和VLA鲁棒性验证结果表明,鲁棒性主要取决于扰动类型而非模型本身,且VLM对大角度相机旋转的鲁棒性强于VLA。对于VLA而言,即使小至$\pm1^\circ$的扰动在许多情况下也会改变模型所下达的动作。我们还发现,鲁棒性更多取决于模型系列而非模型规模。
cs.CV / 10 / 2609.22295

On The Robustness-Resolution Tradeoff In Temporal Quantization Of Event Streams

论事件流时间量化中的鲁棒性-分辨率权衡
Chowdhury, Sayeed Shafayet, Sharmin, Ruhi
Abstract
Event pipelines often discretize asynchronous timestamps before learning. This step looks harmless, but its stability depends directly on temporal resolution. We study this dependence at the representation level. We first show that hard temporal binning is discontinuous: an arbitrarily small timestamp shift near a boundary can move unit event mass between bins. We then define a class of nonnegative, mass-preserving, resolution-faithful continuous encoders and prove that every encoder in this class has global L1 sensitivity at least 2/Delta, where Delta denotes bin width. Linear two-bin interpolation attains this limit. Local support and first-moment preservation also make it unique. Experiments on SHD, N-MNIST, and DVS128 Gesture support the analysis. Across uniform timestamp budgets, linear interpolation lowers mean representation drift by 47-72% while keeping clean accuracy nearly unchanged. On DVS Gesture, it produces zero prediction flips across all tested budgets and three seeds. On SHD, measured drift follows 1/Delta with R^2 = 0.992.
Chinese Translation
事件流水线通常在学习之前对异步时间戳进行离散化。这一步骤看似无害,但其稳定性直接取决于时间分辨率。我们在表示层面研究这种依赖关系。我们首先证明硬性时间分箱是不连续的:在边界附近的任意微小时间戳偏移都可能使单位事件质量在 bins 之间移动。随后,我们定义了一类非负、保质量且分辨率忠实的连续编码器,并证明该类中的每一个编码器的全局 L1 敏感度至少为 2/Δ,其中 Δ 表示 bin 宽度。线性双 bin 插值可以达到该界限,且其局部支撑和一阶矩保持性使其具有唯一性。在 SHD、N-MNIST 和 DVS128 Gesture 数据集上的实验支持了这一分析。在相同的时间戳预算下,线性插值将平均表示漂移降低了 47%–72%,同时保持干净数据的准确率几乎不变。在 DVS Gesture 上,线性插值在所有测试预算和三个随机种子下均未产生任何预测翻转。在 SHD 上,实测漂移符合 1/Δ 规律,R² = 0.992。
cs.CV / 11 / 2609.22302

Authority-Preserving Evaluation of Medical Vision-Language Assistants

医学视觉语言助手的权限保留式评估
Fan, Flint Xiaofeng, Tan, Cheston, Ong, Yew-Soon, Wattenhofer, Roger
Abstract
Medical vision-language models can propose how urgently a skin lesion should be reviewed, but the local service retains authority to accept or replace that proposal under referral policy, capacity, and locally held patient context. Proposal quality and selected-action quality are therefore distinct evaluation targets, and benchmark evidence transfers between them only when local review preserves the expected action score. We introduce AuthEval, a logging and evaluation framework that records both actions, scores the selected action under declared local criteria, and, where feasible, scores the declined proposal under the same rule. It reports the resulting authority gap only when the record supports it. Because the gap is the product of the proposal-change rate and the mean score change on changed cases, that rate alone determines neither its magnitude nor its sign. On ISIC 2019, with MedGemma and simulated local review, two constraint regimes with similar change rates produced an optimistic image-equal gap under capacity ($+0.744$ simulator units) but no detectable gap under safety. The declared evaluation unit also mattered: under mixed constraints the gap reversed from $+0.374$ to $-0.206$ when weighting shifted from image to lesion-aware cluster. AuthEval thus clarifies whether a study's records support claims about the model, the workflow, or both.
Chinese Translation
医学视觉语言模型可以提议皮肤病变应被审查的紧急程度,但在转诊政策、服务容量以及本地掌握的患者情境约束下,本地服务机构保留接受或替换该提议的最终决定权。因此,提议质量与所选行动质量是两个不同的评估目标,只有当本地审查保留了预期的行动得分时,基准测试的证据才能在两者之间迁移。我们提出了 AuthEval,一个日志记录与评估框架,它同时记录两类行动,根据声明的本地标准对所选行动进行评分,并在可行的情况下按照同一规则对被否决的提议进行评分。仅当记录支持时,它才报告由此产生的权限差距。由于该差距是提议变更率与变更案例上的平均得分变化之积,变更率本身既不能决定差距的大小,也不能决定其正负号。在 ISIC 2019 数据集上,使用 MedGemma 和模拟的本地审查,两个变更率相近的约束情境在容量约束下产生了乐观的“图像等同”差距(+0.744 个模拟器单位),而在安全约束下未检测到差距。所声明的评估单位同样重要:在混合约束下,当权重从图像转向病变感知聚类时,差距从 +0.374 反转为 -0.206。因此,AuthEval 能够澄清一项研究的记录是否支持关于模型、工作流程或两者的结论。
cs.CV / 12 / 2609.22308

GameReplica: A Benchmark for Black-Box Visual Game Replication by Vision-Language Agents

GameReplica:面向视觉-语言智能体黑盒视觉游戏复刻的基准测试
Qiao, Boyu, Tang, Zixin, Hao, Xiaoshuai, Li, Wenbo
Abstract
Coding-agent benchmarks usually evaluate implementation after the target behavior has been specified in text, code, or demonstrations. Existing research has extensively evaluated the ability of coding agents to generate programs from textual specifications. However, under black-box conditions where neither source code nor documentation is available, it remains underexplored whether an agent can induce the rules solely through visual observation and active interaction and reproduce the target system as a verifiable executable system. To this end, we present GameReplica, a closed-loop evaluation framework for end-to-end black-box game replication that covers the full perception, exploration, induction, reproduction, and verification pipeline. GameReplica comprises 125 tasks spanning 25 games across 5 core mechanism families, with each game instantiated at five difficulty levels. The tasks require an agent to access the target game only through screenshots and an action interface, induce the key visual elements and gameplay rules from pixel feedback and interaction outcomes, and generate a self-contained, runnable game replica that can be automatically verified by an external program. Experiments show that current coding agents still face substantial challenges in end-to-end black-box replication: the best-performing model (Claude Opus 4.8) achieves an overall score of 71.6\%, while the remaining models score only 4.0\%--42.9\%. Further analysis reveals a consistent pattern across all models: visual-fidelity scores are substantially higher than implementation- and rule-consistency scores, indicating that agents replicate visual appearance more readily than game mechanics. The difficulty levels further amplify the performance gap: from L1 to L5, the overall score of weaker agents drops sharply, whereas that of the best-performing agent declines only slightly.
Chinese Translation
编码智能体基准测试通常在目标行为已通过文本、代码或演示明确指定后评估其实现能力。已有研究广泛评估了编码智能体根据文本规范生成程序的能力。然而,在既无源代码也无文档可用的黑盒条件下,智能体能否仅通过视觉观察和主动交互归纳出规则,并将目标系统复刻为可验证的可执行系统,这一问题仍未得到充分探索。为此,我们提出了GameReplica,一个面向端到端黑盒游戏复刻的闭环评估框架,涵盖完整的感知、探索、归纳、复刻与验证流程。GameReplica包含125个任务,涵盖5大核心机制类别下的25款游戏,每款游戏设有五个难度级别。任务要求智能体仅通过截图和动作接口访问目标游戏,从像素反馈和交互结果中归纳关键视觉元素与游戏规则,并生成一个独立可运行、可由外部程序自动验证的游戏复刻版本。实验表明,当前编码智能体在端到端黑盒复刻任务上仍面临巨大挑战:表现最好的模型(Claude Opus 4.8)总分为71.6%,而其余模型的得分仅为4.0%–42.9%。进一步分析揭示了所有模型的一致规律:视觉保真度得分显著高于实现一致性和规则一致性得分,表明智能体更容易复刻视觉外观而非游戏机制。难度级别进一步放大了性能差距:从L1到L5,较弱智能体的总分急剧下降,而表现最好的智能体仅略有下降。
cs.CV / 13 / 2609.22315

Yarn tracking of large-scale 3D textile reinforcements using topological material features

基于拓扑材料特征的大规模三维纺织增强体纱线追踪
Herichi, Hafsa El, Mendoza, Arturo, Wielhorski, Yanneck, Talbot, Hugues, Roux, Stéphane
Abstract
Automated segmentation of CT images has become increasingly important to enhance the reliability of simulations through the generation of high fidelity numerical models. This study addresses the challenging task of semi-automatically tracking textile reinforcements in fan blade dry preforms using X-ray CT images captured at coarse resolutions (i.e., above 140 $\mu$m). Our approach offers a scalable, slice-based analysis conducted on planes orthogonal to the main yarn directions, applied to a large-scale real industrial component. This enables accurate identification and tracking of yarn paths while requiring minimal training. The method models three key yarn properties statistically: their typical cross-section shape, their continuity and movement in the 3D space, and their spatial relative arrangement with respect to neighboring yarns. These statistical properties are integrated into a tracking framework via a variational formulation that optimizes all yarn center positions in successive cross-section planes. The method tracks more than 3,000 warp yarns across 1,500 slices and achieves a tracking success rate above 90%. Overall, this work demonstrates a promising approach toward large-scale, automated textile reinforcement annotation, paving the way for more efficient material characterization in complex composite structures.
Chinese Translation
CT图像的自动分割对于通过生成高保真数值模型来提高仿真可靠性正变得日益重要。本研究针对一项具有挑战性的任务,即在粗分辨率(即高于140微米)的X射线CT图像下半自动追踪风扇叶片干态预成型体中的纺织增强体。我们的方法提供了一种可扩展的、基于切片的分析,在与主纱线方向正交的平面上进行,并应用于一个大规模的真实工业构件。这使得在仅需少量训练数据的情况下,即可实现纱线路径的精确识别与追踪。该方法从统计学角度建模了纱线的三个关键特性:其典型的横截面形状、其在三维空间中的连续性与运动,以及其相对于邻近纱线的空间相对排布。这些统计特性通过变分公式被整合到一个追踪框架中,该框架在连续的横截面平面内优化所有纱线中心的位置。该方法在1500个切片中追踪了3000多根经纱,追踪成功率超过90%。总体而言,这项工作展示了一种面向大规模自动化纺织增强体标注的有前景的方法,为复杂复合材料结构中更高效的材料表征铺平了道路。
cs.CV / 14 / 2609.22323

ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning

ALPINE:面向参数与样本高效少样本学习的自适应定位方法
Yadav, Neeraj
Abstract
Few-shot learning research is predominantly evaluated on accuracy alone, with limited attention to the parameter and training-sample budgets required to reach that accuracy - a real constraint for practitioners without large-scale compute. We present an ultra-lightweight (22,249-34,917 parameter) spatial-relational architecture for few-shot image classification that combines fixed Gabor edge-energy guidance with a windowed, content-adaptive patch locator. Under a strictly matched, iso-episode-budget protocol (250 meta-training episodes, 5 canonical seeds, 600 evaluation episodes per seed), our architecture achieves 5-shot accuracy gains, consistent across all five seeds, over Prototypical Networks, Relation Networks, and MAML on both CIFAR-FS and MiniImageNet, while using 27-53% fewer parameters than any baseline. It also converges in fewer training episodes, generalizes better to an unseen fine-grained domain (CUB-200-2011 birds, zero retraining), and is more robust to 50% occlusion and 25% spatial translation than all three baselines. A series of falsification ablations - zeroing relational tokens at inference and retraining without them entirely - shows that the architecture's pairwise relational computation, while present, is not the primary driver of its performance; the content-adaptive patch locator is. We report this honestly, together with a capacity sweep showing a genuine accuracy plateau near 22-35k parameters, and release full seed-level results and checkpoint hashes for reproducibility.
Chinese Translation
少样本学习研究大多仅以准确率作为评估标准,而对达到该准确率所需的参数量与训练样本预算关注甚少——对于缺乏大规模计算资源的从业者而言,这是一个现实的约束。我们提出了一种超轻量级(22,249–34,917 个参数)的空间-关系架构,用于少样本图像分类。该架构将固定的 Gabor 边缘能量引导与基于窗口的内容自适应图块定位器相结合。在严格匹配的等轮次预算协议下(250 个元训练轮次、5 个标准随机种子、每个种子 600 个评估轮次),我们的架构在 CIFAR-FS 和 MiniImageNet 两个数据集上的 5-shot 准确率均优于 Prototypical Networks、Relation Networks 和 MAML,且该优势在全部五个种子上保持一致,同时所使用的参数量比任何基线模型少 27–53%。此外,该架构在更少的训练轮次内即可收敛,能更好地泛化到未见过的细粒度领域(CUB-200-2011 鸟类数据集,零再训练),并且在 50% 遮挡和 25% 空间平移条件下比所有三个基线模型更加鲁棒。一系列证伪性消融实验——在推理时将关系标记置零以及完全不使用关系标记进行再训练——表明该架构的成对关系计算虽然存在,但并非其性能的主要驱动因素;真正的关键在于内容自适应图块定位器。我们诚实地报告了这一结果,并辅以容量扫描实验,证实在约 22–35k 参数附近存在真实的准确率平台期。同时,我们发布了完整的种子级结果和检查点哈希值,以保证可复现性。
cs.CV / 15 / 2609.22333

Dimensionality reduction for AI based hyperspectral image classification based on XAI

基于可解释人工智能(XAI)的AI高光谱图像分类降维方法
Zeljković, Vladimir, Stojanović, Branka, Ganster, Harald, Nešković, Aleksandar
Abstract
This research addresses the challenge of limited material recycling in wood recycling processes by leveraging artificial intelligence (AI)-based dimensionality reduction. Our study explores the application of convolutional neural networks (CNNs) in multi-channel hyperspectral imaging (HSI), extending beyond RGB channels to over 200 spectral channels. Dimensionality reduction within this context involves streamlining the feature space for AI system training and inference. Focusing on explainable AI (XAI) methods, this paper contributes to a broader research initiative, presenting a solution framework that enhances the sustainability and efficiency of wood recycling processes.
Chinese Translation
本研究通过利用基于人工智能(AI)的降维技术,应对木材回收过程中材料回收率有限的挑战。我们的研究探索了卷积神经网络(CNN)在多通道高光谱成像(HSI)中的应用,将光谱通道从RGB三通道扩展到200多个光谱通道。在此背景下,降维涉及为AI系统的训练和推理精简特征空间。本文聚焦于可解释人工智能(XAI)方法,为更广泛的研究计划做出贡献,提出了一个提升木材回收过程可持续性与效率的解决方案框架。
cs.CV / 16 / 2609.22351

Hierarchical Aggregation of Semantic Uncertainty in 3D Scene Graphs

三维场景图中语义不确定性的分层聚合
Zumaya, Carlos Cueto, Catalano, Iacopo, Bessa, Wallace Moreira, Placed, Julio A.
Abstract
Open-vocabulary 3D Scene Graphs (3DSGs) ground each object node in a vision-language embedding, yet they record every entry as equally certain, so a robot querying the map cannot tell which of its entries are unreliable. Estimators of semantic uncertainty could supply that distinction, but they require repeated sampling of a model, training, or held-out labels, none of which are available to a deployed system at query time. We present a framework that exploits the detector confidence and the embeddings a 3DSG already stores, converts them into a probability that an entry is correct, and propagates that probability through the containment hierarchy into a belief that a room contains a queried class. Four signals, each paired with the object-level error it indicates, are converted to probabilities at the logit scale learned by the vision-language model and combined in closed form with no additional perception or training. Objects sharing a detector and a vocabulary fail together, so the framework aggregates them in the fully correlated limit, where an aggregation under independence would treat one repeated error as repeated evidence. Evaluated on HM3DSem against a state-of-the-art 3DSG system, the framework improves object retrieval and lowers the error of the room-level assertions of the graph it reads.
Chinese Translation
开放词汇三维场景图(3D Scene Graphs, 3DSGs)将每个物体节点锚定于视觉-语言嵌入中,但其对所有条目均以同等确定性记录,因此查询地图的机器人无法判断哪些条目是不可靠的。语义不确定性估计器可以提供这种区分,但它们需要对模型进行重复采样、训练或预留标签,而这些对于部署系统在查询时均不可用。我们提出一个框架,利用三维场景图已经存储的检测器置信度和嵌入,将其转换为某条目正确的概率,并通过包含层级传播该概率,形成关于某房间包含所查询类别的信念。四种信号各自与它所指示的物体级错误配对,在视觉-语言模型学习到的logit尺度上被转换为概率,并以闭式形式组合,无需额外的感知或训练。共享同一检测器和同一词汇表的物体会一同出错,因此该框架在完全相关的极限下对它们进行聚合;而在独立性假设下的聚合则会将一个重复出现的错误视为重复的证据。在HM3DSem数据集上与最先进的三维场景图系统相比,该框架提升了物体检索性能,并降低了其读取的图中房间级断言的误差。
cs.CV / 17 / 2609.22379

MarsRecon: Self-Supervised and Multimodal Surface Representations for Mars

MarsRecon:面向火星的自监督多模态表面表示学习
Naik, Akshay, Juston, Marius F. R., Mahajan, Jay
Abstract
High-resolution orbital imagery offers a rich record of the Martian surface, but sparse geological labels limit supervised representation learning. We present MarsRecon, a geospatially aware pipeline for learning visual and multimodal representations from HiRISE observations of Olympus Mons. The pipeline calibrates NASA Planetary Data System products, extracts valid georeferenced patches, and trains a masked autoencoder on unlabeled imagery. Increasing input resolution and filtering invalid tokens reduced held-out reconstruction loss from 0.1751 to 0.1342 in the principal Stage A model series. We then freeze the visual encoder and align its features with observation text, coordinates, and local--global image context. The strongest current local-primary model achieves image-to-text recall@10 of 0.3787, text-to-image recall@10 of 0.9161, and local-to-global recall@10 of 0.4350 on the held-out test split. These results establish a working Mars-specific pretraining and retrieval pipeline; further crop-overlap controls and downstream geological evaluations are needed to assess the broader utility of its embeddings.
Chinese Translation
高分辨率轨道影像为火星表面提供了丰富的记录,但稀疏的地质标注限制了有监督表示学习。我们提出了MarsRecon,一个具有地理空间感知能力的流水线,用于从奥林匹斯山(Olympus Mons)的HiRISE观测数据中学习视觉与多模态表示。该流水线对NASA行星数据系统(Planetary Data System)产品进行校准,提取有效的地理配准图像块,并在无标注影像上训练掩码自编码器(masked autoencoder)。在主要的Stage A模型系列中,提高输入分辨率并过滤无效token使留出集重建损失从0.1751降至0.1342。随后,我们冻结视觉编码器,并将其特征与观测文本、坐标以及局部—全局图像上下文进行对齐。当前最强的以局部为主的模型在留出测试集上实现了图像到文本recall@10为0.3787、文本到图像recall@10为0.9161,以及局部到全局recall@10为0.4350。这些结果建立了一个可运行的火星专用预训练与检索流水线;仍需进一步的裁剪重叠控制与下游地质评估来衡量其嵌入的更广泛实用性。
cs.CV / 18 / 2609.22392

Style as Cover: Deep Image Steganography via Stylized Transmission

风格即掩护:基于风格化传输的深度图像隐写术
Li, Qi, Yang, Jidong, Yu, Huaike, Wang, Chunpeng, Gao, Suo, Iu, Herbert Ho-Ching, Miao, Yuantian, Ma, Bin, Chen, Xiao
Abstract
Image steganography hides secret message within normal images, with most existing works relying on cover-preserving transmission. However, such a paradigm becomes vulnerable once the original cover is exposed or can be reliably approximated. In this paper, we propose StyleStegaNet, a stylized image hiding framework that replaces cover matching with style-concealment transmission. Instead of transmitting a cover-like stego image, StyleStegaNet generates stylized stego images conditioned on publicly available style references, redefining steganography invisibility from cover-preserving concealment to behavior-level camouflage based on style transformation. Such a setting poses a substantial challenge to reliable secret recovery, since neural stylization can significantly alter the feature statistics exploited by deep hiding methods. To address this challenge, StyleStegaNet decouples the overall task into four coordinated stages: stego generation, stylized transmission, structure-preserving reconstruction, and secret recovery. Moreover, StyleStegaNet is optimized with a progressive three-stage training strategy, in which wavelet-domain constraints and perceptual supervision guide the recoverable information toward structural representations. We further provide an analysis showing that secret recoverability is largely restricted to the normalized structural subspace, offering a mechanistic explanation for why directly stylized baselines fail and why a reconstruction-guided recovery path is necessary. Extensive experiments on DIV2K and MS-COCO datasets demonstrate the effectiveness of StyleStegaNet. And few-shot image steganalysis with two deep detectors further shows detection accuracy near random guessing, approximately 51\%.
Chinese Translation
图像隐写术将秘密信息隐藏在正常图像中,现有工作大多依赖于保封面传输范式。然而,一旦原始封面图像被暴露或可被可靠地近似,该范式就会变得脆弱。本文提出StyleStegaNet,一个风格化图像隐藏框架,用风格掩蔽传输取代封面匹配。StyleStegaNet不再传输类封面隐写图像,而是以公开可用的风格参考为条件生成风格化隐写图像,将隐写不可见性从封面保持式的掩蔽重新定义为基于风格变换的行为级伪装。这种设定给可靠的秘密恢复带来了巨大挑战,因为神经风格化会显著改变深度隐藏方法所利用的特征统计特性。为应对这一挑战,StyleStegaNet将整体任务解耦为四个协同阶段:隐写图像生成、风格化传输、结构保持重建和秘密恢复。此外,StyleStegaNet采用渐进式三阶段训练策略进行优化,其中小波域约束和感知监督引导可恢复信息向结构表征靠拢。我们进一步提供了分析,表明秘密的可恢复性主要局限于归一化的结构子空间,这从机制上解释了为什么直接风格化的基线方法会失败,以及为什么需要重建引导的恢复路径。在DIV2K和MS-COCO数据集上的大量实验证明了StyleStegaNet的有效性。基于两个深度检测器的少样本图像隐写分析进一步表明,其检测准确率接近随机猜测,约为51%。
cs.CV / 19 / 2609.22479

Spatiotemporal Flux Probing for Single-Photon Videography

面向单光子摄像的时空通量探测
Yan, Jerry, Forlivesi, Matteo, Tan, Bowen, Xie, Andrew, Somasundaram, Siddharth, Nousias, Sotiris
Abstract
We address the problem of recovering high-speed videos from dynamic scenes under extreme photon sparsity. Existing methods rely on aggregating photon detections in local spatiotemporal windows to improve signal-to-noise ratio; however, this local grouping discards global structure and fails in low-light regimes where photon detections are sparse in space and time. In this work, we show that the information needed to recover both motion and illumination is encoded in correlations over the full space-time pattern of photon arrivals. Building on this insight, we develop a spatiotemporal flux probing theory and an algorithm that estimates the Fourier coefficients of the underlying intensity directly from the photon stream. We demonstrate that our approach (1) recovers fast motion and temporal illumination dynamics with substantially fewer photons than prior methods, (2) enables velocity-selective videography that automatically refocuses video onto specific detected motions, and (3) generalizes across sensing modalities including single-photon, event, and spike cameras.
Chinese Translation
本文研究在极端光子稀疏条件下从动态场景中恢复高速视频的问题。现有方法依赖于在局部时空窗口内聚合光子探测结果以提高信噪比;然而,这种局部分组方式丢弃了全局结构信息,并且在光子探测在空间和时间上均稀疏的低光照条件下会失效。在这项工作中,我们证明了恢复运动与光照所需的信息编码在光子到达的完整时空模式的关联之中。基于这一洞察,我们提出了一种时空通量探测理论及相应算法,可直接从光子流中估计底层光强函数的傅里叶系数。我们证明了所提方法:(1) 在光子数量远少于已有方法的情况下即可恢复快速运动和时间上的光照动态变化;(2) 实现了速度选择性摄像,能够自动将视频重新聚焦于特定的已探测运动上;(3) 可推广至包括单光子相机、事件相机和脉冲相机在内的多种传感模态。
cs.CV / 20 / 2609.22500

Event-Frame Fusion for Inter-Frame Segmentation via Event-Guided Motion

基于事件引导运动的事件-帧融合帧间语义分割
Hareb, Dalia, Martinet, Jean, Miramond, Benoit, Chicca, Elisabetta
Abstract
Autonomous navigation requires precise and efficient semantic segmentation, yet existing frame-based approaches remain limited by motion blur, glare, latency, and the low temporal resolution (20-30 FPS) of conventional cameras, which leads to information loss between frames. Event cameras have emerged as an alternative sensing modality, capturing intensity changes asynchronously with high temporal resolution, high dynamic range, and sparse outputs. However, event-based algorithms still fall short of frame-based ones in accuracy, as most segmentation methods are designed for dense frame data. To overcome these limitations, we propose a hybrid vision architecture that combines conventional frame-based and event-based cameras. The system integrates two complementary components: (1) a compact Spiking Neural Network (SNN) with 42k parameters for motion estimation, and (2) a lightweight event-driven SNN with 0.84M parameters for frame-based semantic segmentation, which interpolates motion between frames to refine segmentation results. By predicting inter-frame segmentations, the framework achieves segmentation rates of up to 500 Hz with an energy consumption below 1.87 mJ per inference, while maintaining real-time GPU execution at frequencies up to 200 Hz. Additionally, our approach compensates for information loss in frames affected by blur or overexposure, enabling more robust perception in challenging conditions.
Chinese Translation
自主导航需要精确且高效的语义分割,然而现有的基于帧的方法仍然受限于运动模糊、眩光、延迟以及传统相机较低的时间分辨率(20-30 FPS),导致帧间信息丢失。事件相机作为一种替代性传感模态应运而生,它能够以异步方式捕捉亮度变化,具有高时间分辨率、高动态范围和稀疏输出等特点。然而,由于大多数分割方法是为稠密的帧数据设计的,基于事件的算法在精度上仍落后于基于帧的算法。为克服这些局限,我们提出了一种结合传统帧相机和事件相机的混合视觉架构。该系统集成了两个互补的组件:(1)一个仅含42k参数的紧凑型脉冲神经网络(SNN),用于运动估计;(2)一个含0.84M参数的轻量级事件驱动SNN,用于基于帧的语义分割,并在帧间插值运动信息以细化分割结果。通过预测帧间分割,该框架实现了高达500 Hz的分割频率,每次推理的能耗低于1.87 mJ,同时可在GPU上以高达200 Hz的频率实时运行。此外,我们的方法能够补偿因模糊或过曝光影响的帧中的信息丢失,从而在具有挑战性的条件下实现更鲁棒的感知。
cs.CV / 21 / 2609.22506

Rethinking Vision Architectures with Gated Linear Attention and KAN

基于门控线性注意力与KAN的视觉架构再思考
Mehizel, Ali, Khaldi, Oussama
Abstract
Vision Transformers allocate most parameters to multi-layer perceptrons (MLPs) for channel mixing, while token interactions usually rely on quadratic multi-head self-attention (MHSA). Linear attention reduces sequence complexity to O(N), but remains coupled with the same fixed-activation MLP as softmax Transformers. Kolmogorov-Arnold Networks (KANs) instead place learnable univariate maps on edges, yet prior vision KANs keep MHSA or omit attention entirely. We introduce LKAT (Linear Kolmogorov-Arnold Transformer), an isotropic ViT encoder that couples chunkwise Gated Linear Attention (GLA) with a two-layer KAN feed-forward, and we provide an I/O-aware fused RBF-KAN kernel for the radial-basis grid maps. Under a shared DeiT-style recipe we compare LKAT with ViT, ViT-5, and MLP-Mixer. LKAT-B exceeds ViT-B/16, ViT-5-B, and Mixer-B/16 on ImageNet-100. Tiny/Small/Base LKAT variants scale consistently on CIFAR-10/100, and ImageNet-100 pretraining transfers to CIFAR fine-tuning. The results support gated linear attention and KAN-based radial-basis functions as complementary inductive biases for mid-scale visual representation learning.
Chinese Translation
视觉Transformer将大部分参数分配给多层感知机(MLP)用于通道混合,而词元交互通常依赖于二次复杂度的多头自注意力(MHSA)。线性注意力将序列复杂度降低至O(N),但仍然与与softmax Transformer相同的固定激活MLP耦合在一起。Kolmogorov-Arnold网络(KAN)则将可学习的单变量映射置于网络边上,然而先前的视觉KAN要么保留MHSA,要么完全省略注意力机制。我们提出了LKAT(Linear Kolmogorov-Arnold Transformer),一种各向同性的ViT编码器,它将分块门控线性注意力(Gated Linear Attention, GLA)与两层KAN前馈网络相结合,并为径向基网格映射提供了一个I/O感知的融合RBF-KAN核。在共享的DeiT风格训练方案下,我们将LKAT与ViT、ViT-5和MLP-Mixer进行比较。在ImageNet-100上,LKAT-B超越了ViT-B/16、ViT-5-B和Mixer-B/16。Tiny/Small/Base版本的LKAT变体在CIFAR-10/100上表现出一致的扩展性,并且ImageNet-100的预训练可以迁移到CIFAR的微调任务中。结果表明,门控线性注意力与基于KAN的径向基函数可以作为互补的归纳偏置,用于中等规模的视觉表征学习。
cs.CV / 22 / 2609.22562

AdaMerge: Tuning-Free Patch Compression for Multi-Vector Visual Document Retrieval

AdaMerge:面向多向量视觉文档检索的无调优补丁压缩方法
You, Jianxin, Ni, Kun
Abstract
Multi-vector visual document retrieval (VDR) models such as ColPali and ColNomic achieve strong accuracy by representing each document with hundreds to thousands of patch-level embeddings, at substantial storage and latency cost. Existing compression methods either prune unimportant patches or merge similar ones into clusters; the recent state-of-the-art merging method Prune-then-Merge (PtM) consistently outperforms pruning-only baselines at high compression, but requires a per-dataset cluster budget m to be tuned by grid search. We observe that the merge-cosine sequence produced by hierarchical clustering exhibits a sharp cliff separating mergeable redundancy from salient signal, and that the location of this cliff is concentrated in a narrow band across more than 11,000 documents from 14 datasets. This suggests the merge boundary can be detected per document rather than tuned per dataset. Building on this observation, we propose AdaMerge, a plug-and-play compression method that (i) detects each document's own cliff via gap analysis on the merge-cosine trajectory, and (ii) builds attention-weighted cluster centroids to preserve salient signal. On the long-document benchmark ViDoRe-V2 (4 datasets, two backbones), AdaMerge significantly outperforms tuned PtM across the operating range (p < 10^-4); on the short-document benchmark ViDoRe-V1 (10 datasets, two backbones), where all merging methods are already near-lossless, AdaMerge matches tuned PtM without any per-dataset tuning. AdaMerge adds only about 10 ms per document and exposes a single global hyperparameter shared across all datasets and backbones.
Chinese Translation
以 ColPali 和 ColNomic 为代表的多向量视觉文档检索(VDR)模型通过为每份文档表示数百至数千个补丁级嵌入而取得了很高的准确率,但带来了巨大的存储与延迟开销。现有压缩方法要么剪除不重要的补丁,要么将相似补丁合并成簇;近期的最先进合并方法 Prune-then-Merge(PtM)在高压缩率下始终优于仅剪枝的基线方法,但需要针对每个数据集通过网格搜索调优簇预算 m。我们观察到,层次聚类产生的合并余弦序列呈现出一个陡峭的悬崖,将可合并的冗余信息与显著信号清晰分开,并且该悬崖的位置在来自 14 个数据集的超过 11,000 份文档中集中于一个狭窄区间。这表明合并边界可以按每份文档进行检测,而无需按数据集调优。基于这一观察,我们提出了 AdaMerge,一种即插即用的压缩方法,它(i)通过对合并余弦轨迹进行间隙分析来检测每份文档自身的悬崖位置,(ii)构建注意力加权的簇质心以保留显著信号。在长文档基准 ViDoRe-V2(4 个数据集、两种骨干网络)上,AdaMerge 在整个工作区间内显著优于经过调优的 PtM(p < 10^-4);在短文档基准 ViDoRe-V1(10 个数据集、两种骨干网络)上——此时所有合并方法已近乎无损——AdaMerge 无需任何按数据集调优即可与调优后的 PtM 相当。AdaMerge 每份文档仅增加约 10 毫秒的开销,且仅暴露一个跨所有数据集和骨干网络共享的全局超参数。
cs.CV / 23 / 2609.22582

Beyond the Leaderboard: Counterfactual Diagnosis of End-to-End and VLA Driving Policies Under Domain Shift

超越排行榜:端到端与视觉-语言-动作(VLA)驾驶策略在域偏移下的反事实诊断
Yang, Ruolin, Huang, Zilin, Wang, Buoyue, Wan, Zhengyang, Luo, Yuhao, Sheng, Zihao, Chen, Sikai
Abstract
End-to-end and vision-language-action (VLA) driving policies are compared by leaderboard rank, but a rank reports an outcome, not the behaviour behind it, so it predicts poorly how a policy will behave at a new site. On six released policies, rank on nuScenes open-loop error or on NAVSIM's leaderboard does not carry over to scenes with a pedestrian near the ego corridor at a new site. We propose a counterfactual check-up: a few hundred real frames, each edited two ways (pedestrian removed, or re-lit by a night-style perturbation), every edit verified by an independent detector, and the change in the planned trajectory read as a diagnosis rather than a score. From these edits two causal axes are read, and five exams built on them separate what a score merges: how far the policy plans to drive, whether seeing the pedestrian buys safety, whether that response scales with danger, whether the plan moves when nothing requires it, and how much an irrelevant lighting change moves it. On 246 NAVSIM near-pedestrian scenes, in the cells where the pedestrian lies on the planned path only 1.9% of responses are genuine avoidance, and under our open-loop protocol the median clearance change is at most 0.03 m and the median change in planned distance at most 0.08 m for every policy. In a pre-registered test from left- to right-hand drive, the exposure and specificity orderings, the lighting verdict and the collision outcome transfer, while point values and the hazard-sensitivity verdict do not. Read as a selection report, the profiles say which policy is safe because it plans short, which covers a human-like distance without yielding, and which is unsteady under a change that requires no reaction, and they price each verdict: most settle within a few dozen frames, hazard sensitivity needs hundreds. Code and edited frames will be released.
Chinese Translation
端到端与视觉-语言-动作(VLA)驾驶策略通常通过排行榜排名进行比较,但排名只反映结果,而非其背后的行为,因此难以预测策略在新场景中的表现。对六个已发布的策略而言,其在 nuScenes 开环误差或 NAVSIM 排行榜上的排名,并不能迁移到新场景中自车行驶走廊附近有行人出现的情形。我们提出一种反事实检查方法:使用数百帧真实图像,每帧以两种方式进行编辑(移除行人,或施加夜间风格的扰动),每次编辑均由独立检测器验证,并将规划轨迹的变化作为诊断信息而非分数来解读。从这些编辑中可以读出两条因果轴线,并在此基础上构建五项检验,将单一分数所混淆的方面区分开来:策略规划的行驶距离、看到行人是否能带来安全性、该响应是否随危险程度而缩放、在无需反应时规划是否仍发生变化,以及无关的光照变化会对其造成多大影响。在 246 个 NAVSIM 行人临近场景中,在行人位于规划路径上的情形下,仅有 1.9% 的响应是真正的避让;在我们的开环协议下,所有策略的中位净空变化不超过 0.03 米,规划行驶距离的中位变化不超过 0.08 米。在一项预先注册的从左舵到右舵驾驶的测试中,暴露度与特异性的排序、光照判定及碰撞结果得以迁移,而具体数值与危险敏感性判定则不能。作为选型报告来解读,这些画像能够说明哪个策略因规划行驶距离短而安全、哪个策略在保持类人行驶距离的同时不进行避让、以及哪个策略在无需反应的变化下表现不稳定,并且为每项判定标明所需代价:大多数判定在几十帧内即可完成,而危险敏感性需要数百帧。代码和编辑后的图像帧将会公开发布。
cs.CV / 24 / 2609.22588

Seeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act

看到并不足够:视觉语言模型能感知证据却无法据此行动
Dai, Yuyang, Huang, Bofei, Zhang, Hongbo, Xie, Haoran
Abstract
Vision-language models (VLMs) perform strongly on visual question answering benchmarks, yet often make decisions that contradict visual evidence they have already identified correctly. We distinguish perceptual failure, where relevant evidence is not recognized, from process failure, where recognized evidence fails to constrain the final decision. We introduce VPAC-Bench, a benchmark spanning nine real-image process families, with each image annotated by its current activity stage and nearby stage transition. We also propose State-Relevance-Target (SRT), a family of structured process-prior interventions that requires models to connect visible evidence to the relevant process state before answering. Across multiple VLMs, process failure is widespread: models that correctly enumerate visual candidates still over-commit to a single answer in more than 95% of ambiguous cases. An explicit process-structured intervention reduces this rate to below 13% without degrading performance on unambiguous cases. However, the transfer of process priors is model-dependent, and generic SRT does not consistently outperform strong chain-of-thought baselines. When the relevant stage transition is known, boundary-aligned SRT substantially outperforms generic process prompting and all tested chain-of-thought baselines across assembly, physical state transition, navigation and traffic, and object-use affordance tasks. These results show that process priors are most useful when aligned with the scene's specific decision boundary, motivating boundary-aware prior selection for process-grounded visual reasoning.
Chinese Translation
视觉语言模型(VLM)在视觉问答基准测试中表现优异,但它们做出的决策常常与自身已正确识别的视觉证据相矛盾。我们区分两类失败:感知失败,即未能识别相关证据;过程失败,即虽已识别证据却未能约束最终决策。我们提出了VPAC-Bench,一个涵盖九类真实图像过程家族的基准,每张图像均标注了其当前活动阶段及邻近的阶段转换。我们还提出了状态-相关性-目标(State-Relevance-Target, SRT),这是一类结构化的过程先验干预方法,要求模型在回答之前将可见证据与相关过程状态联系起来。在多个VLM上,过程失败普遍存在:能正确列举视觉候选项的模型在超过95%的模糊案例中仍过度倾向于单一答案。显式的过程结构化干预可将该比例降至13%以下,且不损害模型在明确案例上的性能。然而,过程先验的迁移效果依赖于具体模型,且通用SRT并不总是优于强思维链(chain-of-thought)基线。当相关阶段转换已知时,边界对齐的SRT在装配、物理状态转换、导航与交通以及物体使用可供性任务上均显著优于通用过程提示和所有测试的思维链基线。这些结果表明,过程先验只有与场景特定的决策边界对齐时才最为有效,这为面向过程基础视觉推理的边界感知先验选择提供了依据。
cs.CV / 25 / 2609.22631

X-Beat: An Explainable Framework for ECG Image Classification

X-Beat:一种面向心电图图像分类的可解释框架
Tahsin, Mohammad Sadman, Adarbah, Haitham Y., Noore, Afzel
Abstract
Accurate automated interpretation of electrocardio- grams (ECGs) is essential for early detection of cardiac condi- tions such as myocardial infarction and rhythm abnormalities. However, many high-performing deep learning models remain difficult to deploy in clinical settings due to limited transparency and lack of reliability validation. In this work, we present X- Beat, an explainable and reliability-aware benchmark framework for ECG image classification designed to support trustworthy AI systems in healthcare. The proposed framework combines transfer learning with post-hoc explainability and systematic reliability evaluation across four cardiac classes: Abnormal Heartbeat, History of Myocardial Infarction, Myocardial In- farction, and Normal Heartbeat. Multiple ImageNet-pretrained CNN backbones, including EfficientNet-B0, ResNet-50, DenseNet- 121, and MobileNetV3-Large, are evaluated under a unified training protocol. Beyond standard performance metrics, we incorporate Grad-CAM-based visual explanations together with additional analyses, including explanation stability, regional sen- sitivity, and confidence-based reliability assessment, to examine whether model predictions are supported by clinically meaningful evidence. Experimental results show that ResNet-50 achieves the best performance, reaching 91.94% accuracy and a macro F1- score of 0.9098, with strong class separability (AUC up to 0.995). Explanation analyses indicate that the model primarily focuses on waveform-relevant regions, while reliability evaluation suggests that most incorrect predictions occur with lower confidence. Overall, this work provides a structured and reproducible bench- mark for evaluating both predictive performance and explanation reliability in ECG image classification, contributing toward the development of trustworthy and interpretable AI components for clinical decision support systems.
Chinese Translation
心电图的准确自动解读对于早期发现心肌梗死和心律异常等心脏疾病至关重要。然而,许多高性能深度学习模型由于透明度有限和缺乏可靠性验证,难以在临床环境中部署。在本研究中,我们提出了X-Beat,一个面向心电图图像分类的可解释且可靠性感知的基准框架,旨在支持医疗领域可信赖的人工智能系统。该框架将迁移学习与事后可解释性方法以及系统性可靠性评估相结合,涵盖四种心脏类别:异常心跳、心肌梗死病史、心肌梗死和正常心跳。我们在统一的训练协议下评估了多个基于ImageNet预训练的CNN骨干网络,包括EfficientNet-B0、ResNet-50、DenseNet-121和MobileNetV3-Large。除标准性能指标外,我们还结合了基于Grad-CAM的可视化解释以及其他分析,包括解释稳定性、区域敏感性和基于置信度的可靠性评估,以检验模型预测是否得到具有临床意义的证据支持。实验结果表明,ResNet-50取得了最佳性能,准确率达91.94%,宏平均F1分数为0.9098,且类别可分性强(AUC最高达0.995)。解释分析表明,模型主要关注与波形相关的区域,而可靠性评估显示大多数错误预测发生在较低置信度的情况下。总体而言,本工作为评估心电图图像分类中的预测性能和解释可靠性提供了一个结构化且可复现的基准,为临床决策支持系统中可信赖、可解释的人工智能组件的开发做出了贡献。
cs.CV / 26 / 2609.22641

ConsistWorld: Evidence Routing for Consistent Multi-Agent World Models

ConsistWorld:面向一致多智能体世界模型的证据路由机制
Xu, Qianxun, Zeng, Xianfang, Liao, Xinyao, Cheng, Wei, Yu, Gang, Zhang, Chi
Abstract
Autoregressive video world models enable temporally coherent generation for a single observer. Extending them to multiple agents requires consistency across independently controlled views and temporal gaps under causal streaming. We present ConsistWorld, a multi-agent world model that generates camera-controlled video streams of a static scene from one shared image. We formulate consistency as routing evidence from committed multi-agent history and concurrently generated peer views to the tokens being generated. Pose Conditioned Memory Retrieval selects relevant historical observations from all agents, recovering evidence beyond the recent context window. Visibility-Gated Peer Sharing regulates current peer information according to estimated historical coverage and current-view overlap. Together, they determine which historical observations enter the context and where concurrent peer information contributes, supporting long-term recall and coordinated exploration. Both mechanisms use camera geometry and maintain a bounded active context for a fixed agent count and retrieval budget. Experiments on evidence sharing cases and video length and agent number generalizations show that ConsistWorld achieves a strong cross-time and cross-agent consistency while preserving competitive generation quality.
Chinese Translation
自回归视频世界模型能够为单一观察者生成时间上连贯的视频。将其扩展到多智能体场景,则需要在因果流式生成条件下,保证独立控制的多个视图之间以及跨时间间隔的一致性。我们提出ConsistWorld,一种多智能体世界模型,能够从一张共享图像生成静态场景的相机控制视频流。我们将一致性问题形式化为:将来自已确定的多智能体历史以及同时生成的同伴视图的证据路由到正在生成的词元(token)上。姿态条件记忆检索(Pose Conditioned Memory Retrieval)从所有智能体中选取相关的历史观测,从而恢复超出近期上下文窗口范围的证据。可见性门控同伴共享(Visibility-Gated Peer Sharing)依据估计的历史覆盖情况和当前视图的重叠程度来调节当前同伴信息的引入。这两种机制共同决定了哪些历史观测进入上下文、并发同伴信息在何处发挥作用,从而支持长期记忆回溯与协同探索。两种机制均利用相机几何信息,并在固定智能体数量和检索预算下维持有界的活跃上下文。在证据共享案例以及视频长度与智能体数量泛化实验中,ConsistWorld在保持有竞争力的生成质量的同时,实现了强大的跨时间与跨智能体一致性。
cs.CV / 27 / 2609.22647

Math2Visual-X: A Modular Framework for Pedagogically Aligned Lower-Primary Math Visuals Generation

Math2Visual-X:一个面向教学对齐的小学低年级数学可视化图形生成的模块化框架
Maduranga, H. D. E., Munasinghe, S. K., Weerasekara, K. P. T. I., Ranathunga, Surangika, de Silva, Nisansa
Abstract
Visual representations can help lower-primary learners understand Math Word Problems, but generating classroom-usable visuals remains difficult. Existing symbolic systems are controllable but limited in coverage, while end-to-end text-to-image systems often fail to satisfy exact mathematical constraints. This paper presents a symbolic visual generation framework for lower-primary MWP generation with broader problem coverage and more scalable asset generation. The framework includes an LLM-based routing layer, three worksheet-oriented generation modules, and two fallback mechanisms for open-world SVG asset acquisition. A human evaluation comparing Math2Visual-X with Stable Diffusion XL, Nano Banana, and GPT Image showed that the proposed method achieved the strongest overall performance. The results indicate that the framework offers a scalable and pedagogically grounded approach for automatic MWP visual generation.
Chinese Translation
可视化表征可以帮助小学低年级学习者理解数学文字题(Math Word Problems, MWP),但生成可用于课堂教学的图形仍然十分困难。现有的符号化系统虽然可控性强,但覆盖范围有限;而端到端的文生图系统往往难以满足精确的数学约束。本文提出了一个面向小学低年级数学文字题的可视化图形生成框架,具有更广的题目覆盖范围和更具可扩展性的素材生成能力。该框架包括一个基于大语言模型(LLM)的路由层、三个面向练习册的生成模块,以及用于开放世界SVG素材获取的两种回退机制。一项将 Math2Visual-X 与 Stable Diffusion XL、Nano Banana 和 GPT Image 进行对比的人工评估表明,所提出的方法取得了最强的整体表现。结果表明,该框架为数学文字题的自动可视化生成提供了一种可扩展且符合教学规律的方法。
cs.CV / 28 / 2609.22687

PanoSeg3R: Feed-Forward 3D Semantic Segmentation for Panoramic Images with an Automatic Data Curation Pipeline

PanoSeg3R:基于自动数据整理流程的全景图像前馈式三维语义分割
Yoon, Heechan, Jung, Dongki, Nguyen, Phuc, Lin, Ming, Manocha, Dinesh
Abstract
We present PanoSeg3R, a feed-forward framework for 3D panoramic semantic segmentation. Unlike existing methods designed for perspective inputs, PanoSeg3R jointly predicts 3D geometry and multi-view semantic segmentation in one single forward pass. Built upon a pretrained reconstruction backbone that supports panoramic images, our approach extends feed-forward 3D reconstruction with a query-based mask decoder. Furthermore, we introduce an automatic panorama data curation pipeline that leverages the complementary strengths of off-the-shelf foundation models to generate reliable pseudo semantic annotations, substantially expanding the training data and improving zero-shot generalization. PanoSeg3R achieves state-of-the-art performance on panoramic 3D semantic segmentation, improving 3D mIoU by up to 16.02 on ScanNet++, while the curated training data further improves zero-shot performance by up to 4.26 and 43.28 mIoU on Stanford2D3D and ToF-360, respectively. Website: https://harryyoon777.github.io/PanoSeg3R/
Chinese Translation
我们提出了 PanoSeg3R,一个用于三维全景语义分割的前馈式框架。与现有的针对透视输入设计的方法不同,PanoSeg3R 能够在单次前向传播中联合预测三维几何结构和多视图语义分割。我们的方法构建在一个支持全景图像的预训练重建骨干网络之上,通过引入基于查询(query-based)的掩码解码器对前馈式三维重建进行了扩展。此外,我们提出了一个自动的全景数据整理流程,利用现有基础模型(foundation models)的互补优势来生成可靠的伪语义标注,从而大幅扩充训练数据并提升零样本(zero-shot)泛化能力。PanoSeg3R 在全景三维语义分割任务上取得了最先进的性能,在 ScanNet++ 数据集上将三维 mIoU 提升了高达 16.02;同时,整理后的训练数据在 Stanford2D3D 和 ToF-360 数据集上分别将零样本性能进一步提升高达 4.26 和 43.28 mIoU。网站:https://harryyoon777.github.io/PanoSeg3R/
cs.CV / 29 / 2609.22688

Vision2CAD: A Visual Agent Harness for Explicit Geometry Referencing and Localization in Parametric CAD Modeling

Vision2CAD:用于参数化CAD建模中显式几何引用与定位的视觉代理框架
Cheng, Xi, Zhai, Chenxi, Cheng, Hang, Fan, Mingyu, Feng, Pingfa, Zeng, Long
Abstract
Generating parametric CAD models requires accurate geometry and stable feature dependencies. Existing methods face challenges in selecting geometric references, interpreting sketch-plane local coordinates, and establishing sketch constraints to projected external geometry. We present Vision2CAD, a visual agent harness that combines vision-language model (VLM) reasoning with deterministic CAD kernel operations. An ID-based interface supports explicit geometry selection, a local-coordinate bridge converts view coordinates into sketch coordinates, and projected-edge localization supports external sketch constraints. These mechanisms establish feature dependencies within the supported modeling operations and constraint types. We also introduce the Geometry Explicit Reference Dataset (GERD), which aligned commands, geometry states and IDs at every modeling step. On GERD-EVL and a DeepCAD test subset, Vision2CAD improves mIoU by 11.1\% and 5.6\% and reduces Chamfer distance by 17.3\% and 41.8\%, respectively. Parameter-editing experiments and ablation studies further proved the preservation of parametric dependencies.
Chinese Translation
生成参数化CAD模型需要精确的几何形状和稳定的特征依赖关系。现有方法在选择几何引用、解释草图平面局部坐标系以及为投影外部几何建立草图约束方面面临挑战。我们提出了Vision2CAD,这是一个将视觉语言模型(VLM)推理与确定性CAD内核操作相结合的视觉代理框架。基于ID的接口支持显式几何选择,局部坐标桥接机制将视图坐标转换为草图坐标,投影边定位则支持外部草图约束。这些机制在所支持的建模操作和约束类型范围内建立了特征依赖关系。我们还引入了几何显式引用数据集(Geometry Explicit Reference Dataset, GERD),该数据集在每个建模步骤中对齐了命令、几何状态和ID。在GERD-EVL和DeepCAD测试子集上,Vision2CAD分别将mIoU提升了11.1%和5.6%,并将Chamfer距离分别降低了17.3%和41.8%。参数编辑实验和消融研究进一步证明了参数化依赖关系的保持。
cs.CV / 30 / 2609.22706

DOA-SORT: Directional Occlusion-Aware Multi-Object Tracking with Distributional Observations

DOA-SORT:基于方向性遮挡感知与分布式观测的多目标跟踪
Wang, Hao
Abstract
Identity association in multi-object tracking (MOT) is vulnerable to partial occlusion, truncated detections, and fluctuating confidence scores. Existing motion-dominant trackers commonly represent occlusion as a scalar penalty. This treatment misses the directional observation bias caused by occlusion: left, right, top, and bottom occlusions distort the location and shape of a detection in different ways. We propose \ours{} (Directional Occlusion-Aware SORT), an online and training-free tracker that models these biases explicitly. First, it infers a soft front--back ordering from box overlap and relative bottom positions, and estimates directional occlusion coverage and depth. It then constructs a mixture of one clean and four directional occlusion observation components. The model uses a five-dimensional observation comprising box center, area, confidence, and aspect ratio, and adapts observation noise to predicted occlusion and detection confidence. The directional mixture likelihood is used in high-confidence association, low-confidence association, and track recovery; ambiguity penalties and local order-consistency swaps further reduce identity errors among nearby objects. On the DanceTrack validation split, \ours{} improves HOTA from 63.00 to 66.34, AssA from 45.10 to 49.57, and IDF1 from 62.19 to 65.28 over OA-SORT with the same detector and evaluation protocol. The gains are concentrated in association quality while detection accuracy remains stable. Additional local evaluations on MOT17 and MOT20 train splits characterize cross-dataset behavior under the same no-ReID tracking protocol.
Chinese Translation
多目标跟踪(MOT)中的身份关联容易受到部分遮挡、截断检测和置信度分数波动的影响。现有的以运动为主导的跟踪器通常将遮挡表示为一个标量惩罚,这种处理方式忽略了遮挡引起的方向性观测偏差:来自左侧、右侧、顶部和底部的遮挡会以不同方式扭曲检测框的位置和形状。我们提出了DOA-SORT(Directional Occlusion-Aware SORT),这是一种在线且无需训练的跟踪器,能够显式地建模这些偏差。首先,它根据框的重叠度和相对底部位置推断出一个软性的前后排序,并估计方向性遮挡覆盖率和深度;然后,构建了一个由一个干净观测分量和四个方向性遮挡观测分量组成的混合模型。该模型使用一个五维观测向量,包括框中心、面积、置信度和长宽比,并根据预测的遮挡程度和检测置信度自适应调整观测噪声。方向性混合似然被用于高置信度关联、低置信度关联和轨迹恢复;歧义惩罚和局部顺序一致性交换进一步减少了邻近目标之间的身份错误。在DanceTrack验证集上,在相同的检测器和评估协议下,DOA-SORT相比OA-SORT将HOTA从63.00提升至66.34,AssA从45.10提升至49.57,IDF1从62.19提升至65.28。这些提升集中在关联质量上,同时检测精度保持稳定。此外,我们还在MOT17和MOT20训练集上进行了额外的局部评估,以刻画在相同的无ReID跟踪协议下的跨数据集表现。
cs.CV / 31 / 2609.22716

ZIL: Zero-shot Image-to-LiDAR Registration

ZIL:零样本图像到激光雷达点云配准
Li, Zijun, Sun, Xiaotian, Shen, Xuelun, Dai, Yao, Ao, Sheng, Shi, Yangyang, Engel, Jakob, Cai, Zhipeng, Wang, Cheng
Abstract
Image-to-LiDAR registration estimates the camera pose of an image with respect to a LiDAR point cloud. It has diverse applications in autonomous driving, robot navigation etc. However, state-of-the-art (SOTA) methods still 1) mostly assume same-frame inputs, struggling with the image and point cloud from distant frames; 2) rely on domain-specific training, failing to generalize to unseen scenarios. We propose ZIL, the first foundation model for zero-shot non-synchronized image-to-LiDAR registration. ZIL encodes the input image and point cloud with the Vision and Point Transformers. In addition to regressing the relative pose, ZIL also learns to predict 3D coordinates, which substantially improves the pose accuracy without additional annotations. Interestingly, naive mix-data training cannot enable zero-shot generalization, which requires normalization on both camera intrinsics and the LiDAR vertical-axis origin. Trained on 7 public datasets with 1.4M LiDAR frames, ZIL consistently and significantly outperforms previous SOTA with a single model across 5 in-domain and zero-shot benchmarks, reducing the translation and rotation errors by up to 87% and 76% (shown in Fig. 1). Code and models are available at https://github.com/ZijunLi7/ZIL.
Chinese Translation
图像到激光雷达(LiDAR)配准旨在估计图像相对于激光雷达点云的相机位姿,在自动驾驶、机器人导航等领域有着广泛的应用。然而,当前最先进(SOTA)的方法仍然存在以下问题:1)大多假设输入来自同一帧,难以处理图像与点云来自相距较远帧的情况;2)依赖特定领域的训练,无法泛化到未见过的场景。我们提出了ZIL,首个用于零样本非同步图像到激光雷达配准的基础模型。ZIL使用视觉Transformer(Vision Transformer)和点云Transformer(Point Transformer)对输入图像和点云进行编码。除了回归相对位姿之外,ZIL还学习预测3D坐标,从而在无需额外标注的情况下大幅提升位姿精度。有趣的是,简单的混合数据训练无法实现零样本泛化,这需要对相机内参和激光雷达垂直轴原点进行归一化处理。ZIL在7个公开数据集(共140万帧激光雷达数据)上训练,仅使用单一模型就在5个域内和零样本基准上持续且显著地超越了以往的SOTA方法,将平移和旋转误差分别最多降低了87%和76%(如图1所示)。代码和模型已发布于 https://github.com/ZijunLi7/ZIL。
cs.CV / 32 / 2609.22750

Towards Robust Classroom Attendance: A Comprehensive Evaluation of Face Detection and Recognition Models

迈向鲁棒的课堂考勤:人脸检测与识别模型的综合评估
Trivedi, Himani, Patel, Hiren, Patel, Ridham, Patel, Krutika, Patel, Nancy
Abstract
Manual attendance methods, such as paper or register-based systems, take a lot of time, can lead to errors, and are easy to falsify. Face recognition is more reliable, but it frequently struggles in classrooms because lighting and other conditions can vary. Face recognition datasets are designed for regulated environments and do not capture the actual challenges found in classrooms. To address this, a new face detection and recognition dataset, the Visage Face dataset, comprising 16,234 face samples, is proposed for the task of face detection and recognition. The photos are taken from different angles and under varying lighting conditions, with students showing a range of expressions, and some faces partly covered to reflect real-life situations. A YOLO-based system is used to detect faces and tested seven advanced face recognition models with thirteen configurations: LVFace, QCFace, FaceLiVTv2, TopoFR, EdgeFace, TransFace, and GhostFaceNets. Of these, FaceLiVTv2-M performed best, with 99.75% Top-1/Top-5 accuracy and an inference time of 6.459 ms. These results show that the Visage Face Dataset is a realistic and challenging benchmark for face recognition in classroom attendance.
Chinese Translation
传统的手工考勤方法,如纸质或登记册系统,耗时较长、容易出现错误,且易于伪造。人脸识别是一种更可靠的方式,但在教室环境中常常面临挑战,因为光照和其他条件可能会发生变化。现有人脸识别数据集均面向受控环境设计,无法反映教室中的实际困难。为解决这一问题,本文提出了一个新的人脸检测与识别数据集——Visage Face数据集,包含16,234个人脸样本。这些照片从不同角度、不同光照条件下拍摄,学生表现出多种表情,部分人脸被部分遮挡,以反映真实场景。系统采用基于YOLO的方法进行人脸检测,并在十三种配置下测试了七个先进的人脸识别模型:LVFace、QCFace、FaceLiVTv2、TopoFR、EdgeFace、TransFace和GhostFaceNets。其中,FaceLiVTv2-M表现最佳,Top-1/Top-5准确率达到99.75%,推理时间为6.459毫秒。这些结果表明,Visage Face数据集是课堂考勤人脸识别领域一个真实且具有挑战性的基准数据集。
cs.CV / 33 / 2609.22762

DriveReferee: Geometric Safety Verdicts Need Not Be Learned for Driving World-Action Models

DriveReferee:驾驶世界-动作模型的几何安全判定无需学习获得
Yu, Fengcheng, Parikh, Dhruv, Ye, Junjie, Bhatt, Maulik, Vu, Thang, Vasiljevic, Igor, Guizilini, Vitor, Wang, Yue
Abstract
Generative world-action models (WAMs) jointly generate future video and vehicle actions, while their action branches remain primarily optimized by expert imitation. Yet imitation provides no explicit closed-loop geometric verdict for generated trajectories, making verification important during both training and deployment. Closed-loop evaluators can check collision and drivable-area violations, but require privileged scene state unavailable at deployment. Existing approaches often close this gap by learning a verifier from sensor features. For these geometric checks, the rule itself is explicit. For example, collision is determined by whether the rolled-out ego footprint overlaps occupied vehicle space. What is unavailable at deployment is the scene state needed to apply the rule. We introduce DriveReferee, which uses a learned geometry readout to predict the scene representation from camera observations and executes the geometric safety rule directly rather than learning it. The resulting analytic referee evaluates collision and drivable-area safety from a scene state and candidate trajectory. During training, it scores self-sampled trajectories on ground-truth state and distills the resulting preferences into the WAM policy. At deployment, the same referee evaluates generated trajectories on this predicted state and selects a safer alternative when needed. The analytic referee requires no verdict-specific training, and its decisions follow an explicit geometric rule. Under matched candidates and inference budgets, it matches or outperforms all learned-verifier and heuristic baselines. Given the same predicted state and trajectory, learning the verdict provides no measurable downstream gain despite requiring tens of thousands of evaluator-labeled training examples. On the full NAVSIM navtest, DriveReferee reaches 92.02 PDMS with single-camera visual input and no external training data.
Chinese Translation
生成式世界-动作模型(WAM)能够联合生成未来视频与车辆动作,但其动作分支主要通过专家模仿学习进行优化。然而,模仿学习无法为生成的轨迹提供显式的闭环几何判定,因此验证在训练和部署阶段都十分重要。闭环评估器可以检查碰撞和可行驶区域违规,但需要部署时无法获得的特权场景状态。现有方法通常通过从传感器特征学习一个验证器来弥合这一差距。对于这些几何检查而言,规则本身是显式的。例如,碰撞由自车投影足迹是否与其他车辆占据空间重叠来决定。部署时缺失的只是应用该规则所需的场景状态。我们提出 DriveReferee,它利用学习得到的几何读取模块从相机观测预测场景表示,并直接执行几何安全规则而非通过学习获得。由此得到的解析式裁判根据场景状态和候选轨迹评估碰撞和可行驶区域安全性。在训练阶段,它在真值状态上对模型自采样轨迹进行评分,并将评分产生的偏好蒸馏到 WAM 策略中。在部署阶段,同一裁判在预测状态上评估生成的轨迹,并在需要时选择更安全的替代方案。该解析式裁判无需针对判定的专门训练,其决策遵循显式的几何规则。在相同候选轨迹和推理预算下,它达到或超越了所有学习式验证器和启发式基线。在相同的预测状态和轨迹条件下,学习判定规则尽管需要数以万计的评估器标注训练样本,却未带来可测量的下游收益。在完整的 NAVSIM navtest 基准上,DriveReferee 仅使用单相机视觉输入且无外部训练数据即达到 92.02 的 PDMS。
cs.CV / 34 / 2609.22788

Human-Level Accuracy, Non-Human Strategies: Revealing Model-Human Divergence in Video Physical Reasoning

人类水平的准确率,非人类的策略:揭示视频物理推理中模型与人类的分歧
Li, Fanhong, Zheng, Shurui, Yin, Zi, Cui, Junbo, Ji, Lei, Liu, Jia
Abstract
Video foundation models now reach human-level accuracy on physical-reasoning benchmarks, yet such tasks require predicting unobserved physical outcomes. Do these models perform human-like forward simulation, or do they exploit statistical regularities in visible scenes? Accuracy alone cannot distinguish these strategies. We introduce a distributional evaluation framework that treats model seeds and human raters as populations, enabling comparison of consensus, uncertainty, and strategy. On the Physion benchmark, we evaluate three ViT-L architectures (V-JEPA2, VideoMAEv2, DINOv2). V-JEPA2 narrows the accuracy gap to ~1 percentage point (73.2% vs. 74.2%), yet model-human disagreement reaches 26.4%, far exceeding human-human disagreement (4.8%), with substantially lower agreement (kappa ~ 0.48 vs. 0.91). The divergence follows forward-simulation demands: models outperform humans on geometric reasoning (linking, +11.8 pp) but underperform on gravitational dynamics (rolling, -11.8 pp) and causal chains (dominoes, -10.5 pp). Strategy fingerprinting confirms all three architectures share non-human strategies while none aligns with humans. Attribution analysis suggests that unobservable outcome features, rather than visible scene properties, predict this divergence, consistent with models relying more on scene-level statistical regularities than on explicit forward simulation, a systematic divergence that accuracy alone cannot reveal. Code is available at https://github.com/fanhong-li/model-human-divergence.
Chinese Translation
视频基础模型目前在物理推理基准测试上已达到人类水平的准确率,然而此类任务需要预测未被观察到的物理结果。这些模型是在进行类人的前向模拟,还是利用了可见场景中的统计规律?仅凭准确率无法区分这两种策略。我们提出了一种分布式评估框架,将模型种子和人类评分者视为总体,从而能够比较其共识性、不确定性和策略。在Physion基准上,我们评估了三种ViT-L架构(V-JEPA2、VideoMAEv2、DINOv2)。V-JEPA2将准确率差距缩小至约1个百分点(73.2% vs. 74.2%),但模型与人类的分歧高达26.4%,远超人类之间的分歧(4.8%),且一致性显著更低(kappa约为0.48 vs. 0.91)。这种分歧遵循前向模拟的需求:模型在几何推理(连接,+11.8个百分点)上优于人类,但在重力动力学(滚动,-11.8个百分点)和因果链(多米诺,-10.5个百分点)上表现不佳。策略指纹分析证实这三种架构均共享非人类策略,且无一与人类对齐。归因分析表明,预测这种分歧的是不可观察的结果特征,而非可见场景的属性,这与模型更多依赖场景级统计规律而非显式前向模拟的假设相一致——这是一种仅凭准确率无法揭示的系统性分歧。代码可在 https://github.com/fanhong-li/model-human-divergence 获取。
cs.CV / 35 / 2609.22789

PixelART: Image-to-Layer Decomposition without Latents or Text-to-Image Pretraining

PixelART:无需潜变量或文生图预训练的图像到图层分解方法
Jia, Zelin, Zhang, Zhao, Tang, Zhicong, Yuan, Yuhui, Liu, Shixia
Abstract
Image-to-layer decomposition converts a flattened image into editable RGBA layers, enabling element-level editing in design workflows. Existing diffusion-based systems typically adapt large pretrained text-to-image (T2I) models and introduce RGBA autoencoders or variable-layer architectural modules. We revisit this design choice and ask whether layer decomposition truly requires these heavyweight components. We introduce PixelART, a pixel-space rectified-flow Transformer trained from scratch for image-to-layer (I2L) decomposition. PixelART directly denoises regional RGBA pixel patches using a single-stream multi-modal diffusion Transformer, avoiding RGBA-VAEs, pretrained T2I backbones, and layer-specific decoders. We identify a key property of the task: high-noise timesteps determine layer assignment and coarse layer organization, while low-noise timesteps mainly refine color, alpha, texture, and boundaries. Based on this observation, we propose a terminal-boosted timestep sampling strategy to increase training coverage in the high-noise layer assignment regime. Trained on 4M multi-layer design templates, PixelART achieves state-of-the-art layer decomposition and composite reconstruction on Design-Multi-Layer-Bench and LICA with over 80% fewer parameters, 98% lower latency, and 85% lower memory than the recent Qwen-Image-Layered model. Ablation experiments show that pixel-space $\mathbf{x}$-prediction, terminal-boosted timestep sampling, and data/model scaling are critical, while T2I pretraining brings marginal benefits to the I2L task.
Chinese Translation
图像到图层分解将扁平化的图像转换为可编辑的RGBA图层,从而在设计工作流中实现元素级编辑。现有的基于扩散模型的系统通常对大规模预训练的文生图(T2I)模型进行适配,并引入RGBA自编码器或可变图层架构模块。我们重新审视了这一设计选择,并探讨图层分解是否真的需要这些重量级组件。我们提出了PixelART,一个从零开始训练的像素空间整流流Transformer,用于图像到图层(I2L)分解。PixelART使用单流多模态扩散Transformer直接对区域RGBA像素块进行去噪,避免了RGBA-VAE、预训练T2I骨干网络以及图层专用解码器。我们发现该任务的一个关键特性:高噪声时间步决定了图层分配和粗略的图层组织结构,而低噪声时间步主要用于细化颜色、透明度、纹理和边界。基于这一观察,我们提出了一种终端增强的时间步采样策略,以增加训练在高噪声图层分配区间上的覆盖。PixelART在400万个多层设计模板上训练,在Design-Multi-Layer-Bench和LICA数据集上实现了最先进的图层分解与合成重建效果,与近期的Qwen-Image-Layered模型相比,参数量减少80%以上,延迟降低98%,内存占用降低85%。消融实验表明,像素空间的 $\mathbf{x}$ 预测、终端增强的时间步采样以及数据与模型规模扩展至关重要,而T2I预训练对I2L任务的收益甚微。
cs.CV / 36 / 2609.22807

C$^{2}$-INR: Customized Convolutional Implicit Neural Representation

C$^{2}$-INR:定制化卷积隐式神经表示
Shi, Jinglei, Chang, Xinran, Cui, Jiaqi, Xia, Yingjie, Xiao, Zhaolin, Li, Chongyi
Abstract
Implicit Neural Representation (INR) leverages neural networks to represent discrete signals such as images as continuous ones, where the network weights serve as a compact form of the signal itself. Most existing INR methods adopt Multi-Layer Perceptrons (MLPs) as their backbone. Since these models render each pixel independently, they inherently fail to exploit the spatial correlations that exist between neighboring pixels. In contrast,convolutional INRs can process pixels in parallel while inherently accounting for inter-pixel dependencies, making them a more natural fit for representing images. Nevertheless, convolutional INRs remain relatively underexplored, and the majority of them rely on fixed architectural settings, leaving little room for image-specific adaptation. In this paper, we investigate network customization for convolutional INRs. We replace conventional filters with irregular directional kernels, whose allocation is guided by the directional energy in the image spectrum, i.e., directions exhibiting stronger energy are assigned a larger number of kernels, enabling content-tailored convolution settings. These kernels are further reformulated via an orthogonal basis to achieve a superior sparse representation. Moreover, we introduce an annealed Gumbel-Softmax-based mechanism for kernel-level activation function selection, which gives the most suitable activation function for each convolution kernel. Extensive experiments demonstrate that our method, namely C$^{2}$-INR, achieves superior performance against state-of-the-art approaches under comparable parameter budgets across a wide range of image processing tasks, including representation, inpainting, and super-resolution.
Chinese Translation
隐式神经表示(Implicit Neural Representation, INR)利用神经网络将图像等离散信号表示为连续信号,其中网络权重作为信号本身的一种紧凑形式。大多数现有的 INR 方法采用多层感知机(MLP)作为骨干网络。由于这类模型对每个像素进行独立渲染,它们天生无法利用相邻像素之间存在的空间相关性。相比之下,卷积 INR 能够并行处理像素,同时内在地考虑像素间的依赖关系,因此更适合表示图像。然而,卷积 INR 的研究仍相对不足,其中大多数依赖于固定的架构设置,几乎没有针对特定图像进行自适应调整的空间。在本文中,我们研究了卷积 INR 的网络定制化。我们用不规则的方向性卷积核替代传统卷积核,其分配由图像频谱中的方向能量引导,即能量较强的方向被分配更多的卷积核,从而实现内容定制化的卷积设置。这些卷积核进一步通过正交基进行重构,以实现更优的稀疏表示。此外,我们引入了一种基于退火 Gumbel-Softmax 的机制,用于卷积核级别的激活函数选择,为每个卷积核选择最合适的激活函数。大量实验表明,我们的方法 C$^{2}$-INR 在相当的参数预算下,在包括表示、图像修复和超分辨率在内的广泛图像处理任务中,均取得了优于最先进方法的性能。
cs.CV / 37 / 2609.22834

SatOV: Restoring Spatial Priors for Training-Free Open-Vocabulary Segmentation in Remote Sensing Imagery

SatOV:面向免训练遥感影像开放词汇分割的空间先验恢复方法
Zhao, Changhao, Zeng, Linglin, Liu, Hai
Abstract
Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation to the spatial characteristics of remote sensing data. Although existing vision-language foundation models perform well in general domains, their image-level classification design weakens the spatial priors needed for high-resolution remote sensing segmentation: structural spatial relations are degraded during deep feature transformation, and fine-grained spatial details are lost during downsampling. To address these complementary deficiencies, we propose SatOV, a training-free framework for open-vocabulary remote sensing segmentation that restores spatial priors at two stages of the representation pipeline. Specifically, Residual QQ Attention (ResQQ) extracts Query-Key self-attention from an intermediate CLIP layer and fuses it with final-layer Query-Query attention via a residual combination, restoring structural spatial priors suppressed by the final-layer representation. Spatially Modulated Upsampling (SatUp) uses the original high-resolution RGB image as spatial guidance, combining spatial feature modulation with guided cross-attention to reconstruct pixel-level textures and boundaries. Extensive experiments on DOTA, UDD, LoveDA, and Vaihingen show that SatOV consistently improves training-free OVS and achieves competitive quantitative and qualitative results against state-of-the-art methods. These results validate the effectiveness of restoring spatial priors at both the representation and spatial-resolution stages for remote sensing open-vocabulary segmentation.
Chinese Translation
遥感影像的开放词汇语义分割(OVS)是一项具有挑战性的像素级任务,要求模型具备强大的泛化能力,并能适应遥感数据的空间特性。尽管现有的视觉-语言基础模型在通用领域表现良好,但其图像级分类设计削弱了高分辨率遥感分割所需的空间先验:结构化的空间关系在深层特征变换过程中被弱化,细粒度的空间细节在下采样过程中丢失。为解决上述互补性缺陷,我们提出了SatOV,一个免训练的开放词汇遥感分割框架,在表示流水线的两个阶段恢复空间先验。具体而言,残差QQ注意力(ResQQ)从CLIP的中间层提取Query-Key自注意力,并通过残差组合方式将其与末层的Query-Query注意力融合,从而恢复被末层表示所抑制的结构化空间先验。空间调制上采样(SatUp)以原始高分辨率RGB图像作为空间引导,将空间特征调制与引导交叉注意力相结合,重建像素级纹理与边界信息。在DOTA、UDD、LoveDA和Vaihingen数据集上的大量实验表明,SatOV能够持续提升免训练开放词汇分割的性能,并取得了与最先进方法相比具有竞争力的定量与定性结果。这些结果验证了在表示阶段和空间分辨率阶段同时恢复空间先验对遥感开放词汇分割的有效性。
cs.CV / 38 / 2609.22849

LINGO: Latent Initialization and Gradient Optimization for Sparse-view X-ray Novel View Synthesis and CT Reconstruction with 3D Gaussian Splatting

LINGO:基于3D高斯泼溅的稀疏视角X射线新视角合成与CT重建的隐式初始化与梯度优化方法
Xing, Lifeng, Jin, Dequan, Bu, Kunpeng, He, Peigeng, Ying, Shihui
Abstract
In novel view synthesis and Computed Tomography (CT) reconstruction with sparse-view X-ray imaging, insufficient angular coverage leads to structural ambiguity and accumulated noise. Integrating 3D Gaussian Splatting (3DGS) with X-ray absorption physics can achieve promising results, but it suffers from noisy initialization, positional insensitivity, and weak gradients in low-density regions. In this paper, we propose a unified Latent Initialization and Gradient Optimization (LINGO) framework to address these issues. LINGO combines latent mask-space initialization with dynamic gradient optimization to improve point cloud structural completeness while accelerating training. It constructs voxel-level 3D filters from X-ray masks to robustly suppress background noise and provide reliable geometric priors. By employing an adaptive voxel scaling strategy and dynamically scaling loss, LINGO can adjust spatial resolution and explicitly amplify gradients in low-density structures. To evaluate the quality of initialization, we introduce the Initialization Point Cloud Structural Deviation (IPSD) metric. Experiments on the X3D dataset indicate that for the novel view synthesis task, LINGO improves the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) by an average of 0.72 and 0.0039, respectively, over baselines under identical sparse-view settings, achieving comparable reconstruction quality within 5k steps to state-of-the-art models typically trained with 30k iterations. For the CT reconstruction task, LINGO also demonstrates consistent improvements, with average PSNR and SSIM gains of 0.36 and 0.0134. These results highlight LINGO's effectiveness in both accelerating training and enhancing reconstruction quality across different sparse-view imaging scenarios.
Chinese Translation
在稀疏视角X射线成像的新视角合成与计算机断层扫描(CT)重建中,角度覆盖不足会导致结构模糊和噪声累积。将3D高斯泼溅(3D Gaussian Splatting, 3DGS)与X射线吸收物理相结合可以取得较好的效果,但其存在初始化含噪、对位置变化不敏感以及低密度区域梯度微弱等问题。本文提出一个统一的隐式初始化与梯度优化(Latent Initialization and Gradient Optimization, LINGO)框架来解决这些问题。LINGO将隐式掩码空间初始化与动态梯度优化相结合,在提升点云结构完整性的同时加速训练。该方法从X射线掩码构建体素级3D滤波器,以稳健地抑制背景噪声并提供可靠的几何先验。通过采用自适应体素缩放策略和动态缩放损失,LINGO能够调整空间分辨率并显式放大低密度结构中的梯度。为评估初始化质量,我们引入了初始化点云结构偏差(Initialization Point Cloud Structural Deviation, IPSD)指标。在X3D数据集上的实验表明,对于新视角合成任务,在相同的稀疏视角设置下,LINGO相较基线方法的峰值信噪比(PSNR)和结构相似性指数(SSIM)平均分别提升0.72和0.0039,并且在仅5k步内即可达到通常需要30k次迭代训练的最先进模型相当的重建质量。对于CT重建任务,LINGO同样展现出一致的改进,PSNR和SSIM平均分别提升0.36和0.0134。这些结果凸显了LINGO在不同稀疏视角成像场景中在加速训练和提升重建质量两方面的有效性。
cs.CV / 39 / 2609.22857

Image Frame Dynamic Object Segmentation and Ego Motion Estimation using Radar Image Fusion

基于雷达图像融合的图像帧动态目标分割与自运动估计
Srivastava, Astik, Grover, Suhani, Sharma, Avinash, Krishna, Madhava
Abstract
Dynamic object segmentation and ego-motion estimation are closely coupled problems in autonomous driving, as accurate ego-motion estimation typically requires static scene observations, while identifying static observations requires knowledge of the ego motion. We present Radar-Dot, a radar--RGB framework that exploits radar Doppler measurements to address this coupling. Radar returns are first used to estimate ego velocity through a linear Doppler constraint, with residual-based static/dynamic segmentation and robust estimation used to reduce the influence of moving objects. The estimated motion is then combined with metric depth and dense optical flow to identify image regions whose observed motion is inconsistent with the rigid scene motion. Experiments on 10 nuScenes scenes (part of nuscenes-mini) demonstrate that the resulting geometric pipeline achieves 20.24% dynamic IoU and 33.67% F1-score over 394 frame pairs, while radar-based static-point filtering improves ego-velocity estimation compared with using all radar returns. These results demonstrate the potential of radar as a modality for jointly improving ego-motion estimation and dynamic object segmentation.
Chinese Translation
动态目标分割与自运动估计是自动驾驶中紧密耦合的两个问题:准确的自运动估计通常需要对静态场景的观测,而识别静态观测又依赖于自运动信息。我们提出Radar-Dot,一种利用雷达多普勒测量来解决这一耦合问题的雷达-RGB框架。首先利用雷达回波通过线性多普勒约束估计自车速度,并采用基于残差的静态/动态分割和鲁棒估计方法来降低运动物体的干扰。随后,将估计得到的运动与度量深度及稠密光流相结合,识别其观测运动与刚性场景运动不一致的图像区域。在nuScenes的10个场景(nuscenes-mini的一部分)上的实验表明,该几何处理流程在394对图像帧上取得了20.24%的动态IoU和33.67%的F1分数;同时,与使用全部雷达回波相比,基于雷达的静态点滤波能够提升自车速度估计的精度。这些结果证明了雷达作为一种模态在联合提升自运动估计与动态目标分割方面的潜力。
cs.CV / 40 / 2609.22868

Planning-Aligned Pretraining of BEV Representations with Sparse Action-Conditioned Targets for End-to-End Autonomous Driving

面向端到端自动驾驶的稀疏动作条件目标规划对齐BEV表征预训练
Song, Jaeha, Hwang, Soonmin
Abstract
End-to-end driving requires planning-relevant bird's-eye-view (BEV) representations, but existing pretraining approaches often rely on task annotations or dense scene reconstruction. We introduce PAVER, Planning-Aligned BEV Encoder Pretraining. From a single LiDAR sweep, PAVER constructs sparse risk and unknown targets describing occupied and unobserved evidence along rule-based ego motions. A 10K-parameter head predicts these targets from masked BEV features conditioned on the action state, directing supervision toward geometric constraints on candidate motions. Pretraining requires no driving-task annotations or dense reconstruction. Only the BEV encoder is transferred, preserving the downstream architecture and camera-only inference. On nuScenes, PAVER reduces VAD-Tiny's average collision rate from 0.51% to 0.19%, while improving planning L2, motion prediction, detection, and mapping. The selected VAD-Tiny and VAD-Base schedules use about 36% less estimated total training time than scratch training, including pretraining. On Bench2Drive Town05 Long, PAVER improves UniAD-Tiny's closed-loop Driving Score from 48.45 to 58.79. The project page is available at https://archiiive99.github.io/PAVER.
Chinese Translation
端到端驾驶需要与规划相关的鸟瞰图(BEV)表征,但现有的预训练方法通常依赖任务标注或密集场景重建。我们提出了PAVER(Planning-Aligned BEV Encoder Pretraining,规划对齐BEV编码器预训练)。PAVER从单次LiDAR扫描中构建稀疏的风险目标和未知目标,用以描述沿基于规则的自我车辆运动方向上的占用与未观测证据。一个仅有1万参数的预测头在动作状态条件下从被掩码的BEV特征预测这些目标,从而将监督引导至候选运动的几何约束上。预训练无需任何驾驶任务标注或密集重建。仅迁移BEV编码器,从而保留下游架构和纯相机推理方式。在nuScenes上,PAVER将VAD-Tiny的平均碰撞率从0.51%降至0.19%,同时提升了规划L2误差、运动预测、检测和建图性能。所选的VAD-Tiny和VAD-Base训练计划(含预训练)相比从零训练可节省约36%的估计总训练时间。在Bench2Drive Town05 Long上,PAVER将UniAD-Tiny的闭环驾驶评分(Driving Score)从48.45提升至58.79。项目页面见 https://archiiive99.github.io/PAVER。
cs.CV / 41 / 2609.22896

Combining Foundation Model Confidence and Monocular Depth for Training-Free Out-of-Distribution Segmentation

结合基础模型置信度与单目深度的免训练分布外分割方法
Varghese, Serin, Hüger, Fabian, Maag, Kira
Abstract
Autonomous vehicles operating in open-world scenarios are inevitably confronted with previously unknown objects, such as exotic animals or loose cargo. The reliable detection and segmentation of these out-of-distribution (OOD) objects is therefore crucial for a safe understanding of the environment and decision-making. Most existing approaches require access to OOD training samples, retraining of the segmentation backbone, or dedicated auxiliary architectures, limiting their practical applicability. We propose a training-free method that derives dense OOD scores directly from the confidence predictions of a foundation segmentation model, without any task-specific fine-tuning or access to anomalous data. To improve the robustness of our OOD segmentation, geometric information from monocular depth estimation is incorporated into the decision process, providing complementary cues to uncertainty-based predictions. We evaluate the proposed method on the SegmentMeIfYouCan benchmark and additionally assess its performance on OOD tracking in video sequences, reflecting the temporal nature of real-world perception systems. The method performs strongly on road-centered benchmarks.
Chinese Translation
在开放世界场景中运行的自动驾驶车辆不可避免地会遇到先前未知的物体,例如奇异动物或散落货物。因此,对这些分布外(OOD)物体的可靠检测与分割,对于安全地理解环境和进行决策至关重要。现有的大多数方法需要访问OOD训练样本、重新训练分割主干网络或专用的辅助架构,这限制了其实际适用性。我们提出了一种免训练方法,直接从基础分割模型的置信度预测中推导出稠密的OOD分数,无需任何任务特定的微调或访问异常数据。为了提高OOD分割的鲁棒性,我们将来自单目深度估计的几何信息融入决策过程,为基于不确定性的预测提供互补线索。我们在SegmentMeIfYouCan基准上评估了所提出的方法,并额外评估了其在视频序列OOD跟踪中的性能,以反映真实世界感知系统的时间特性。该方法在以道路为中心的基准上表现出色。
cs.CV / 42 / 2609.22897

Scout: Open-World Species Recognition on the Edge

Scout:边缘设备上的开放世界物种识别
Rastikerdar, Mohammad Mehdi, Guan, Hui, Ganesan, Deepak
Abstract
Large vision-language models (VLMs) enable recognition beyond a fixed class set, but their computational demands prevent them from running on many edge devices. Cloud offload makes this capability accessible, but sending every image consumes scarce bandwidth and communication energy. We ask how to bring the open-world recognition capability of VLMs to the edge while operating within tight compute, energy, and bandwidth budgets. Wildlife monitoring provides a natural setting for exploring this question because camera traps encounter species not known at deployment. We present Scout, an autonomous open-world recognition system that invokes a cloud VLM intermittently to teach new classes to a compact edge model. Given only the deployment location and empty site frames, Scout autonomously turns each species identified by the VLM into persistent, site-conditioned recognition capability in a resource-efficient edge model, without a predefined species list, human labeling, or manual tuning. Across 30 camera-trap deployments in three regions on an NVIDIA Jetson Orin Nano, the accuracy of Scout remains within 0.1-2.5% of a model given a predefined species list. On species outside its initial class set, Scout achieves 53.7-59.1% accuracy, compared with 56.5-65.1% for full cloud offload, while using 59-71% less deployment energy.
Chinese Translation
大型视觉语言模型(VLM)能够实现对固定类别集合之外的识别,但其计算需求使其无法在许多边缘设备上运行。云端卸载(cloud offload)使这一能力变得可行,但上传每张图像会消耗稀缺的带宽和通信能量。我们探讨如何在严格的计算、能量和带宽预算下,将VLM的开放世界识别能力带到边缘设备。野生动物监测为探索这一问题提供了天然的实验场景,因为红外触发相机(camera traps)在部署时会遇到未知的物种。我们提出了Scout,一个自主的开放世界识别系统,它间歇性地调用云端VLM来教会一个轻量级边缘模型识别新类别。仅需部署位置信息和无目标的空场景帧,Scout即可自主地将VLM识别出的每个物种转化为持久化的、针对特定地点的识别能力,并将其嵌入资源高效的边缘模型中,整个过程无需预定义物种列表、人工标注或手动调参。在NVIDIA Jetson Orin Nano平台上跨越三个地区的30次相机陷阱部署中,Scout的准确率与给定预定义物种列表的模型相比仅相差0.1-2.5%。对于其初始类别集合之外的物种,Scout达到53.7-59.1%的准确率(完全云端卸载为56.5-65.1%),同时部署能耗降低59-71%。
cs.CV / 43 / 2609.22913

AVTR-1: Open Stack for Real-Time Interactive Avatars

AVTR-1:面向实时交互式数字人的开放技术栈
Kravtsov, Artem, Ziganshin, Dmitrii, Poletaev, Vsevolod, Balitskiy, Gleb, Tikhonova, Anastasia, Burkov, Egor, Lebedev, Vadim
Abstract
Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live system must synchronize the model's output with speech from an external voice agent, schedule video frames for playback, and handle interruptions. We introduce AVTR-1, an open stack for real-time interactive avatar conversations, built around a compact 153M-parameter autoregressive flow-matching motion generator conditioned on both participants' audio. We adapt its audio encoder for streaming through self-distillation. The stack turns the model's chunk-based generation into a continuous, synchronized audio-video stream driven by an external voice agent, and we analytically derive its contribution to the user-facing latencies and validate the resulting bounds with two commercial voice agents. Further experiments demonstrate that AVTR-1 leads the compared dyadic systems on all reported visual-quality metrics and most conventional listening-motion metrics while remaining competitive in lip synchronization. Its inference runtime operates in real time on data-center and consumer GPUs. However, conventional listening metrics do not establish whether the paired speaker's speech contributes to generated motion. We therefore introduce the Reference-Based Directed Granger Gain (R-DGG), which measures the additional predictive information carried by speaker speech after accounting for listener history and speaker motion. R-DGG finds statistically supported predictive dependence for recorded listeners and all evaluated dyadic systems, but not for talking-head generators without paired audio or mismatched speaker-listener pairs. We release the model weights, renderer, and serving backend under component-specific licenses.
Chinese Translation
说话人头(talking-head)与双向对话模型现已能够实现实时推理,然而仅有快速的运动生成并不足以产生交互式对话。一个实时系统必须将模型输出与来自外部语音代理的语音同步,调度视频帧以供播放,并处理打断情况。我们提出 AVTR-1,一个用于实时交互式数字人对话的开放技术栈,其核心是一个紧凑的、拥有 153M 参数的自回归流匹配(flow-matching)运动生成器,该生成器以对话双方的音频为条件。我们通过自蒸馏将其音频编码器改造为支持流式处理。该技术栈将模型基于分块的生成过程转化为由外部语音代理驱动的连续、同步的音视频流,并通过解析方法推导了其对面向用户延迟的贡献,且结合两个商用语音代理验证了所得到的延迟上界。进一步的实验表明,AVTR-1 在所有报告的视觉质量指标以及大多数传统听者-运动指标上领先于所对比的双向对话系统,同时在唇音同步方面也保持竞争力。其推理运行时可在数据中心级和消费级 GPU 上实时运行。然而,传统的听者指标无法确定配对说话人的语音是否对生成的运动有贡献。因此,我们提出基于参考的有向格兰杰增益(Reference-Based Directed Granger Gain, R-DGG),用于衡量在考虑听者历史和说话人运动之后,说话人语音所携带的额外预测信息。R-DGG 在录制的听者以及所有评估的双向对话系统中均发现了具有统计支持性的预测依赖关系,而在没有配对音频的说话人头生成器或说话人-听者不匹配的情况下则未发现这种依赖。我们在组件特定的许可协议下发布了模型权重、渲染器和服务后端。
cs.CV / 44 / 2609.22916

Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation

规划与渲染协同:基于自回归布局与扩散模型的深度融合视觉文本生成
Chen, Guanqiao, Tan, Jingru, Mao, Dongxing, Chen, Catherine, Du, Zijian, Qin, Libo, Guo, Hu Jian, Wang, Alex Jinpeng
Abstract
Generating text-rich images from prompts requires both textual fidelity and the coherent integration of text into the surrounding image. An explicit layout can provide structured guidance about what text should appear and where, but a well-formed plan alone does not guarantee that the renderer will realize it faithfully. Existing layout-based AR-diffusion systems typically optimize planning and rendering separately, preventing the planner's representations from being adapted jointly with image synthesis. We introduce DuetGen, an autonomous visual text generator built on DeepFusion, which jointly learns autoregressive planning and continuous diffusion rendering. DeepFusion conditions a diffusion transformer on the planner's prompt and bbox-content hidden states, allowing rendering supervision to shape the representations connecting textual plans with visual outputs. Its joint objective combines autoregressive plan supervision, text-region-weighted diffusion learning, and auxiliary coordinate supervision to maintain structured planning, emphasize text-bearing regions, and improve the spatial precision of planner representations. During inference, Phase-Aware Attention Modulation strengthens the correspondence between image regions and their matched coordinate and content states, facilitating region-specific execution of the generated plan. With a 2B planner and a 4B single-stream DiT, DuetGen achieves 0.8293 word accuracy on CVTG-2K and 0.938 accuracy on LongText-Bench, closely matching the substantially larger Qwen-Image on both benchmarks. These results demonstrate the value of jointly learned planning representations and region-specific rendering for autonomous visual text generation.
Chinese Translation
根据提示词生成富文本图像,既要求文本内容的忠实性,又要求文本与周围图像的连贯融合。显式布局可以为文本应出现的内容和位置提供结构化指导,但仅有一个良好的规划并不能保证渲染器能够忠实地将其实现。现有的基于布局的自回归-扩散系统通常分别优化规划与渲染,导致规划器的表征无法与图像合成进行联合调整。我们提出了 DuetGen,一个基于 DeepFusion 构建的自主视觉文本生成器,它联合学习自回归规划与连续扩散渲染。DeepFusion 以规划器的提示词和边界框-内容(bbox-content)隐藏状态作为扩散 Transformer 的条件,使渲染监督能够塑造连接文本规划与视觉输出的表征。其联合目标结合了自回归规划监督、以文本区域加权的扩散学习以及辅助坐标监督,以保持结构化规划、强调含文本区域,并提升规划器表征的空间精度。在推理阶段,相位感知注意力调制(Phase-Aware Attention Modulation)增强了图像区域与其匹配的坐标和内容状态之间的对应关系,从而促进对生成规划的区域特定执行。借助 2B 参数的规划器和 4B 参数的单流 DiT,DuetGen 在 CVTG-2K 上达到了 0.8293 的词准确率,在 LongText-Bench 上达到了 0.938 的准确率,在两个基准上均与规模大得多的 Qwen-Image 十分接近。这些结果证明了联合学习的规划表征和区域特定渲染对于自主视觉文本生成的价值。
cs.CV / 45 / 2609.22941

D3GS: Depth, DINO, and RGB Diffusion Co-Guided 3D Gaussian Splatting for Sparse-View Reconstruction

D3GS:深度、DINO与RGB扩散协同引导的3D高斯泼溅稀疏视角重建
Gao, Yunqi, Liao, Zhanfeng, Tu, Hanzhang, Su, Zhaoqi, Zheng, Guoqing, Wang, Songtao, Zhang, Hongwen, Xue, Zhou, Liu, Leyuan, Liu, Yebin
Abstract
Novel view synthesis from sparse inputs remains challenging for 3D Gaussian Splatting (3DGS) due to ambiguous geometry, cross-view inconsistency, and missing details in under-constrained regions, resulting in degraded reconstruction and unstable rendering. To tackle these issues, we propose D$^{3}$GS, a Depth-DINO-Diffusion guided sparse-view Gaussian reconstruction framework that jointly enhances geometry and appearance. D$^{3}$GS first recovers a high-resolution, metric depth map via diffusion-based completion and DPT (Dense Prediction Transformer) refinement, providing robust Gaussian initialization and geometric constraints. Then, a DINO-guided view-consistent learning is introduced to augment Gaussian attributes with structural features, improving multi-view consistency. Finally, a diffusion-based Gaussian refinement module injects generative priors into an iterative optimization strategy, enhancing high-frequency geometric and appearance details within the Gaussian representation. Experiments on DTU, LLFF, and Mip-NeRF 360 show that D$^{3}$GS achieves consistent and substantial improvements over strong baselines, with ablation studies validating the effectiveness and complementary roles of each component.
Chinese Translation
对于3D高斯泼溅(3D Gaussian Splatting, 3DGS)而言,从稀疏输入进行新视角合成仍然具有挑战性,原因在于几何歧义、跨视角不一致以及欠约束区域中的细节缺失,导致重建质量下降和渲染不稳定。为解决这些问题,我们提出了D$^{3}$GS,一个由深度、DINO与扩散模型引导的稀疏视角高斯重建框架,能够联合增强几何与外观。D$^{3}$GS首先通过基于扩散模型的补全和DPT(Dense Prediction Transformer,稠密预测Transformer)精细化恢复高分辨率的度量深度图,为高斯提供鲁棒的初始化和几何约束。然后,引入DINO引导的视角一致性学习,利用结构特征增强高斯属性,提升多视角一致性。最后,基于扩散模型的高斯精细化模块将生成先验注入迭代优化策略中,增强高斯表示中的高频几何与外观细节。在DTU、LLFF和Mip-NeRF 360数据集上的实验表明,D$^{3}$GS相较于强基线方法取得了一致且显著的性能提升,消融实验验证了各组件的有效性及其互补作用。
cs.CV / 46 / 2609.22942

An Evolutionary Agentic Approach for Open-ended Image Quality Perception

一种面向开放式图像质量感知的进化式智能体方法
Tang, Zhenchen, Peng, Bo, Wang, Zichuan, Yang, Songlin, Cao, Leilei, Zhu, Fengjie, Dong, Jing
Abstract
Generative models are rapidly expanding image quality assessment (IQA) beyond traditional fidelity factors to emerging dimensions such as physical plausibility and text-rendering correctness. However, existing IQA models rely on fixed definitions and heavy supervision, making them difficult to extend to open-ended perceptual dimensions. We identify holistic bias as an important limitation: when scoring an unseen dimension, models reuse generic quality priors, leading to scoring errors and rank inversion. To address this, we propose PACE (Perceptual Agentic Collaborative Evolution), a training-free multi-agent framework that formulates open-ended IQA as explicit protocol construction. Given a target dimension, PACE uses collaborative agents to construct an evaluation protocol composed of verifiable Visual Question Answering (VQA) probes, grounding evaluation in concrete visual evidence rather than holistic impressions. The resulting protocol is calibrated using only four human-annotated images per dimension, while a dual-track scoring mechanism aligns model perception with human scoring scales. Across traditional IQA, structural fidelity, context-aware aesthetics, and newly defined open-ended dimensions, PACE consistently improves its MLLM backbone, achieving competitive performance across diverse IQA settings, and reduces the Holistic Override Rate (HOR) from 44.4\% to 8.6\%.
Chinese Translation
生成式模型正迅速将图像质量评估(IQA)从传统的保真度因素拓展到物理合理性和文本渲染正确性等新兴维度。然而,现有的IQA模型依赖于固定的定义和大量的监督,难以扩展到开放式的感知维度。我们发现整体偏差是一个重要局限:在为未见过的维度打分时,模型会复用通用的质量先验,导致打分错误和排序颠倒。为此,我们提出了PACE(Perceptual Agentic Collaborative Evolution,感知智能体协同进化),这是一个无需训练的多智能体框架,将开放式IQA形式化为显式的协议构建过程。给定一个目标维度,PACE利用协作智能体构建由可验证的视觉问答(VQA)探针组成的评估协议,使评估基于具体的视觉证据而非整体印象。所得到的协议每个维度仅需四张人工标注图像即可完成校准,同时双轨打分机制使模型感知与人类打分尺度对齐。在传统IQA、结构保真度、情境感知美学以及新定义的开放式维度上,PACE持续提升其多模态大语言模型(MLLM)骨干的性能,在多样化的IQA设置中取得了具有竞争力的表现,并将整体覆盖错误率(Holistic Override Rate, HOR)从44.4%降低至8.6%。
cs.CV / 47 / 2609.22947

RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

RewardVerse:面向视频奖励模型的基于评分准则的策略优化
Tang, Zhenchen, Li, Yang, Yang, Songlin, Peng, Bo, Zhao, Xiaotong, Li, Shuai, Fan, Haotian, Zhao, Alan, Dong, Jing
Abstract
Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing video reward models often produce unstable scalar scores because they directly map complex, subjective video quality into a single score without explicit evaluation criteria. This leads to scalar drift, where the scoring scale collapses or shifts across different prompts, making the reward unreliable for RL. Drawing inspiration from professional human annotation engineering, we address this problem with RewardVerse, a rubric-based video reward framework that introduces a dynamic rubric as an intermediate representation between the evaluation query and the scorer. Instead of unconstrained direct scoring, RewardVerse first generates explicit evaluation criteria and then performs rubric-guided scoring, providing a stable semantic anchor that mitigates scalar drift. To efficiently optimize this collaborative pipeline, we propose Rubric-Guided Policy Optimization (RGPO), a two-stage training algorithm. RGPO first warms up the scorer using self-evolving seed rubrics and then jointly optimizes the rubric generator to produce query-adaptive evaluation criteria while continuously aligning the scorer with human ratings. Extensive experiments on the 16-dimensional EvalVerse benchmark and external datasets demonstrate that RewardVerse mitigates scalar drift, achieves state-of-the-art performance on both pointwise and pairwise evaluation, and provides a robust and interpretable reward signal for RL in video generation.
Chinese Translation
强化学习(RL)对于优化视频生成模型至关重要,而一个稳健的奖励模型(RM)是其基石。然而,现有的视频奖励模型往往产生不稳定的标量分数,因为它们在缺乏明确评估标准的情况下,直接将复杂、主观的视频质量映射为单一分数。这导致了标量漂移现象,即评分尺度在不同提示词之间发生坍缩或偏移,使奖励信号在强化学习中不可靠。受专业人工标注工程的启发,我们提出了 RewardVerse,一个基于评分准则(rubric)的视频奖励框架,它引入动态评分准则作为评估查询与评分器之间的中间表示。RewardVerse 不同于无约束的直接打分,它先生成明确的评估标准,再进行准则引导的评分,从而提供稳定的语义锚点以缓解标量漂移。为了高效优化这一协作流程,我们提出了准则引导策略优化(Rubric-Guided Policy Optimization, RGPO),一种两阶段训练算法。RGPO 首先利用自我演化的种子评分准则对评分器进行预热,随后联合优化准则生成器以产生适配查询的评估标准,同时持续将评分器与人类评分对齐。在 16 维 EvalVerse 基准和外部数据集上的大量实验表明,RewardVerse 缓解了标量漂移,在逐点(pointwise)和成对(pairwise)评估中均达到了最先进的性能,并为视频生成中的强化学习提供了稳健且可解释的奖励信号。
cs.CV / 48 / 2609.22950

CLEAR: Complex Learned Explicit Analytical Regularization for Ultra-Accelerated 4D Flow CMR Reconstruction

CLEAR:用于超加速4D Flow CMR重建的复杂可学习显式解析正则化方法
Wache, German Shâma, Neumayer, Sebastian
Abstract
While compressed-sensing regularizers enable interpretable reconstruction of 4D Flow CMR through transparent variational objectives, their hand-crafted nature is too restrictive under high acceleration. State-of-the-art learning-based approaches mitigate this, but typically encode regularization implicitly through unrolled network modules, which limits their interpretability. To address this limitation, we propose CLEAR, designed to combine the interpretability of compressed sensing with the flexibility of learned models. To the best of our knowledge, it is the first learned regularizer for a 4D reconstruction task. In the ultra-accelerated \(10\times\)--\(50\times\) regime of the CMRx4DFlow2026 challenge, CLEAR outperforms compressed sensing locally low-rank (LLR) and the popular variational network FlowVN, while using less than 10k parameters and preserving an interpretable regularization structure.
Chinese Translation
尽管压缩感知正则化方法通过透明的变分目标实现了可解释的4D Flow CMR(心脏磁共振血流成像)重建,但其手工设计的特性在高加速倍数下过于受限。最先进的学习类方法虽然缓解了这一问题,但通常通过展开(unrolled)网络模块隐式地编码正则化,这限制了其可解释性。为解决这一局限,我们提出了CLEAR,旨在结合压缩感知的可解释性与可学习模型的灵活性。据我们所知,这是首个面向4D重建任务的可学习正则化方法。在CMRx4DFlow2026挑战赛的10倍至50倍超加速条件下,CLEAR在仅使用不到1万个参数并保持可解释正则化结构的同时,性能优于压缩感知局部低秩(LLR)方法以及流行的变分网络FlowVN。
cs.CV / 49 / 2609.22967

General Collaborative Intelligence: Architecting Cognition for Resilient Multi-Agent Ecosystems

通用协同智能:面向韧性多智能体生态系统的认知架构设计
Zhang, Lei, Ye, Chun, Yang, Le, Wang, Zhaozhong, Fan, Deng-Ping, Dai, Hang, Wang, Binglu
Abstract
Multi-agent unmanned systems are moving from isolated, ego-centric sensing toward collaborative intelligence, in which distributed agents exchange compact features to overcome a local observation trap that no single agent can escape: occlusions, finite sensor range, and environmental degradation. The field has matured across architectural, communication, embodied, resilience, and trust dimensions, yet existing surveys examine these dimensions in isolation and rarely expose their dependencies. This review offers a unified synthesis through two complementary lenses. The first is a five-dimensional taxonomy spanning collaboration stage, communication paradigm, fusion architecture, learning strategy, and application domain. The second is three cognitive synergy conditions, Semantic Disambiguation, Pragmatic Information Exchange, and Proactive Informational Foraging, that turn cognitive synergy into operational criteria. Across these lenses we survey collaboration architectures and topologies, neural-communication co-design that treats the channel as a differentiable pipeline component, embodied action-perception loops via multi-agent reinforcement learning, and resilience mechanisms for synchronization, uncertainty quantification, and label-efficient learning. We then map these advances onto four operational domains, V2X, unmanned aerial, industrial logistics, and smart cities, and onto the safety-privacy-utility triad. To counter benchmark saturation and evaluation fragmentation, we propose GCI-Bench, a five-pillar scoring protocol with a maturity model that makes the trade-offs of collaborative methods comparable across studies. A critical reflection on reproducibility, the sim-to-real gulf, and conditions under which collaboration degrades performance identifies open challenges and charts directions toward general collaborative intelligence under real-world uncertainty.
Chinese Translation
多智能体无人系统正从孤立的、以自我为中心的感知向协同智能演进。在协同智能中,分布式智能体通过交换紧凑特征来克服任何单一智能体都无法摆脱的局部观测困境:遮挡、有限的传感范围以及环境退化。该领域在架构、通信、具身、韧性和信任等维度上已趋于成熟,但现有综述往往孤立地考察这些维度,鲜少揭示它们之间的依赖关系。本综述通过两个互补的视角提供了统一的综合分析。第一个视角是涵盖协作阶段、通信范式、融合架构、学习策略和应用领域的五维分类体系。第二个视角是三项认知协同条件——语义消歧(Semantic Disambiguation)、语用信息交换(Pragmatic Information Exchange)和主动信息觅食(Proactive Informational Foraging)——它们将认知协同转化为可操作的评价标准。基于这些视角,我们调研了协作架构与拓扑结构、将信道视为可微流水线组件的神经-通信协同设计、通过多智能体强化学习实现的具身动作-感知回路,以及面向同步、不确定性量化和标签高效学习的韧性机制。随后,我们将这些进展映射到四个实际应用领域——车联网(V2X)、无人机、工业物流和智慧城市——以及安全-隐私-效用三元权衡之中。为应对基准饱和与评估碎片化问题,我们提出GCI-Bench,这是一个包含五大支柱的评分协议,并配有成熟度模型,使协作方法的权衡在不同研究之间具有可比性。最后,我们对可复现性、仿真到现实(sim-to-real)鸿沟以及协作导致性能退化的条件进行了批判性反思,指出了开放性挑战,并勾勒了在现实世界不确定性下迈向通用协同智能的方向。
cs.CV / 50 / 2609.23003

M3GA-Wild: A Large-Scale Dataset and Benchmark for Multi-Modal Multi-session Ground-to-Aerial Place Recognition in Forests

M3GA-Wild:面向森林中多模态多时段地面-空中位置识别的大规模数据集与基准
Griffiths, Ethan, Haghighat, Maryam, Denman, Simon, Fookes, Clinton, Ramezani, Milad
Abstract
We present M3GA-Wild, the first benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests. M3GA-Wild unifies and extends existing forest localisation datasets, providing a holistic benchmark with synchronised RGB imagery and LiDAR from ground traversals spanning 36 km, aligned high-resolution aerial imagery and multi-altitude LiDAR covering 370 hectares, and accurate geo-referenced 6-DoF poses for precise evaluation. M3GA-Wild captures diverse forest scenes with varying viewpoints, occlusion, and environmental conditions, enabling systematic evaluation of visual, LiDAR, cross-modal, and multi-modal methods. Baseline experiments show that LiDAR-based approaches significantly outperform vision-only methods under severe viewpoint differences, while current multi-modal fusion strategies yield limited gains due to poor cross-modal alignment. By pairing aerial RGB imagery with geo-referenced aerial LiDAR, M3GA-Wild also enables evaluation of foundation models for monocular depth estimation as a cheap source of 3D geometry from forest imagery, with initial experiments revealing shortfalls of current methods. These results highlight key challenges in cross-platform localisation, including modality misalignment and severe domain gaps. M3GA-Wild establishes a new benchmark to support research in robust multi-modal localisation and long-term autonomy in unstructured natural environments. The dataset and code will be available upon acceptance.
Chinese Translation
我们提出了M3GA-Wild,这是首个针对森林中多模态、多时段地面-空中位置识别的基准数据集。M3GA-Wild统一并扩展了现有的森林定位数据集,提供了一个全面的基准,包括:覆盖36公里的地面行进所采集的同步RGB图像与LiDAR数据、覆盖370公顷的配准高分辨率航拍图像与多高度LiDAR数据,以及用于精确评估的地理配准六自由度(6-DoF)位姿。M3GA-Wild涵盖了具有不同视角、遮挡程度和环境条件的多样化森林场景,能够对视觉、LiDAR、跨模态及多模态方法进行系统性评估。基线实验表明,在剧烈的视角差异下,基于LiDAR的方法显著优于纯视觉方法,而现有多模态融合策略由于跨模态对齐效果不佳,收益有限。通过将航拍RGB图像与地理配准的航拍LiDAR配对,M3GA-Wild还支持将基础模型用于单目深度估计的评估,作为从森林图像获取三维几何信息的一种低成本来源,初步实验揭示了现有方法的不足。这些结果突出了跨平台定位中的关键挑战,包括模态错位和严重的域差距。M3GA-Wild建立了一个新的基准,以支持在非结构化自然环境中鲁棒多模态定位与长期自主性的研究。数据集和代码将在论文录用后发布。
cs.CV / 51 / 2609.23005

Compressing 3D Gaussian Splatting via Cross-Representation Priors

基于跨表示先验的3D高斯泼溅压缩方法
Zhang, Yezheng, Liang, Huanxiong, Zhou, Chuqin, Lu, Guo, Zhang, Wenjun
Abstract
3D Gaussian Splatting (3DGS) enables high-quality novel view synthesis but incurs high storage and transmission costs due to dense Gaussian primitives. Recent anchor-based compression reduces per-primitive redundancy, yet redundancy across anchors remains largely unexploited. We propose CRP-GS (Cross-Representation Priors for Gaussian Splatting), a rate-distortion optimized compression framework that leverages cross-representation priors to improve anchor-level entropy modeling. First, a Correspondence-Oriented Hierarchical Structure (COHS) organizes anchors by feature correspondence rather than spatial proximity, constructing root-leaf dependencies so that selected anchors can act as informative priors to conditionally encode others, yielding more accurate likelihood prediction and lower conditional entropy. Second, Shared Feature Aggregation (SFA) extracts globally shared features from a contextual hash grid and injects them into anchor representations, factoring out scene-consistent low-frequency information that would otherwise be redundantly embedded in individual anchors. Both modules are trained under a unified rate-distortion objective to balance bitrate reduction and rendering fidelity. Experiments across multiple benchmarks show that CRP-GS achieves a favorable overall rate-distortion trade-off, yielding around 30% average bitrate reduction compared to anchor-based baselines while maintaining comparable rendering quality.
Chinese Translation
3D高斯泼溅(3D Gaussian Splatting, 3DGS)能够实现高质量的新视角合成,但由于高斯基元密集,带来了高昂的存储与传输开销。近期基于锚点(anchor)的压缩方法降低了单个基元的冗余,然而锚点之间的冗余在很大程度上仍未被充分利用。我们提出CRP-GS(Cross-Representation Priors for Gaussian Splatting),一种率失真优化的压缩框架,利用跨表示先验来改进锚点级别的熵建模。首先,对应导向的层次结构(Correspondence-Oriented Hierarchical Structure, COHS)依据特征对应关系而非空间邻近性来组织锚点,构建根-叶依赖关系,使被选中的锚点能够作为信息丰富的先验,对其他锚点进行条件编码,从而获得更精确的似然预测和更低的条件熵。其次,共享特征聚合(Shared Feature Aggregation, SFA)从上下文哈希网格中提取全局共享特征并注入锚点表示,将原本冗余嵌入各个锚点中的场景一致性低频信息分解出来。两个模块均在统一的率失真目标下训练,以平衡码率降低与渲染保真度。在多个基准上的实验表明,CRP-GS取得了有利的整体率失真权衡:与基于锚点的基线方法相比,平均码率降低约30%,同时保持了相当的渲染质量。
cs.CV / 52 / 2609.23010

MixiMotion: One-Step Text-to-Motion Generation via Asymmetric Set Distillation

MixiMotion:基于非对称集合蒸馏的单步文本到动作生成
Dinh, Hung, Mai, Binh, Le, Tran Quoc Bao, Nguyen, Lam, Tran, Cong
Abstract
Iterative text-to-motion generation delivers high-quality and semantically aligned motions but requires multiple network evaluations, resulting in substantial inference latency. We present \textbf{MixiMotion}, a strict one-step text-to-motion generation framework based on offline set distillation. Instead of distilling a single teacher trajectory for each text prompt, MixiMotion constructs an offline bank of multiple teacher motions and aligns teacher and student sample sets through \textbf{asymmetric bidirectional matching}. The teacher-to-student direction promotes coverage of diverse teacher-supported motions, while the student-to-teacher direction suppresses unsupported generations. We further introduce differentiable decoded-space kinematic supervision to complement normalized representation matching with constraints in the decoded motion space. At inference, MixiMotion generates a complete motion sequence with a single network evaluation, without teacher queries, iterative sampling, or candidate ranking. On ViMoGen, MixiMotion achieves a semantic alignment score of $0.835$, outperforming the evaluated one-step baselines and approaching the $0.858$ score of its 50-step HY-Motion-1.0-Lite teacher. In blinded human evaluation, MixiMotion obtains an overall rating of $4.33$, compared with $4.50$ for the teacher, while outperforming the evaluated one-/few-step baselines. Meanwhile, generation latency is reduced from $829.58$\,ms to $9.30$\,ms, corresponding to an $89.2\times$ speedup. These results demonstrate an effective quality--efficiency trade-off for strict one-step text-to-motion generation.
Chinese Translation
迭代式文本到动作生成能够生成高质量且语义对齐的动作,但需要多次网络评估,导致显著的推理延迟。我们提出了MixiMotion,一种基于离线集合蒸馏的严格单步文本到动作生成框架。与为每个文本提示蒸馏单一教师轨迹不同,MixiMotion构建了一个包含多个教师动作的离线库,并通过非对称双向匹配来对齐教师与学生样本集合。教师到学生的方向促进了对多样教师支持动作的覆盖,而学生到教师的方向则抑制了不受支持的生成。我们进一步引入可解码空间的微分运动学监督,以在解码后的动作空间中通过约束来补充归一化表示匹配。在推理阶段,MixiMotion仅需一次网络评估即可生成完整的动作序列,无需教师查询、迭代采样或候选排序。在ViMoGen上,MixiMotion实现了0.835的语义对齐得分,优于所评估的单步基线方法,并接近其50步教师模型HY-Motion-1.0-Lite的0.858得分。在盲测人类评估中,MixiMotion获得了4.33的总体评分,相比之下教师模型为4.50,同时优于所评估的单步/少步基线方法。与此同时,生成延迟从829.58毫秒降低至9.30毫秒,对应89.2倍的加速。这些结果证明了严格单步文本到动作生成中有效的质量与效率权衡。
cs.CV / 53 / 2609.23012

CrowdCue: Specialist-Cue Conditioning for Vision-Language Crowd Counting

CrowdCue:面向视觉语言人群计数的专家线索条件化方法
Farazi, Moshiur, Ciftler, Bekir, Dandoush, Abdulhalim, Bendraou, Reda
Abstract
Generative vision-language models (VLMs) offer a counting paradigm in which one model produces both a count and a natural-language account of the scene, yet their raw counting accuracy sits in the range of sub-million-parameter specialist regressors. The open question is whether auxiliary guidance from a pretrained specialist can lift them into useful territory, and through which channel that guidance is best routed. We evaluate Qwen2.5-VL-7B on four widely used crowd counting benchmarks (ShanghaiTech A and B, UCF-QNRF, NWPU-Crowd). Zero-shot prompting rarely produces a parseable count, so LoRA supervised fine-tuning establishes the baseline at overall MAE 81.64. Conditioning on a P2PNet-derived density heatmap as an auxiliary visual signal fails in every encoding we tested, and an adversarial-swap protocol shows the model reads the heatmap but applies it counterproductively. We propose CrowdCue, a family that supplies the same specialist's already-integrated integer count to the VLM as a discrete symbol. The text-channel variant reaches MAE 72.04. The visual-channel variant, which renders the integer as printed digits and supplies it as a second image, reaches MAE 62.65, the strongest result in this paper and well ahead of the cue-supplying specialist alone (84.45 on the same split). In the late-fusion VLM we study, the binding constraint is not the channel but the abstraction level at which the specialist signal is delivered.
Chinese Translation
生成式视觉语言模型(VLM)提供了一种新的计数范式:单一模型既能输出人数,又能以自然语言描述场景,但其原始计数精度仍处于低于百万参数专用回归模型的水平。一个悬而未决的问题是:来自预训练专用模型的辅助引导能否将其提升到可用的精度范围,以及该引导通过哪种通道传递最为有效。我们在四个广泛使用的人群计数基准(ShanghaiTech A 和 B、UCF-QNRF、NWPU-Crowd)上评估了 Qwen2.5-VL-7B。零样本提示几乎无法产生可解析的计数结果,因此我们通过 LoRA 有监督微调建立了基线,总体 MAE 为 81.64。将 P2PNet 导出的密度热图作为辅助视觉信号进行条件化,在我们测试的所有编码方式中均告失败;对抗性交换协议表明模型虽然能读取热图,却以适得其反的方式加以应用。我们提出 CrowdCue 系列方法,将同一专用模型已整合的整数计数以离散符号的形式提供给 VLM。文本通道变体达到 MAE 72.04;视觉通道变体将整数渲染为印刷数字并作为第二张图像输入,达到 MAE 62.65,这是本文中最强的结果,且大幅领先于仅使用提供线索的专用模型(同一划分上为 84.45)。在我们研究的后期融合 VLM 中,关键制约因素并非通道本身,而是专用模型信号所传递的抽象层级。
cs.CV / 54 / 2609.23017

Reconstructed holograms and explanation-aware evaluation for low-cost computational pollen analysis in veterinary cytology

面向兽医细胞学低成本计算花粉分析的重建全息图与解释感知评估
Warshaneyan, Swarn, Danyal, Joial, Cugmas, Blaž, Tamošiūnas, Mindaugas, Kviesis-Kipge, Edgars, Manivannan, Kirishanth, Kadiķis, Roberts
Abstract
Automated pollen analysis supports veterinary cytology, but brightfield microscopy is costlier and more complex than lens-less digital in-line holographic microscopy. We evaluate whether reconstructed holograms can narrow this gap and whether model explanations remain reliable under modality change. Six pollen species were imaged by brightfield and holographic microscopy. Raw, single back-propagation and iterative phase retrieval holograms were evaluated with YOLOv26s detection and MobileNetV4 classification after anchor-based annotation transfer. Six attribution methods were assessed for spatial grounding and faithfulness with the Attribution Health Inspection and Repair (AHIR) protocol, which tests model brittleness under weak noise and corrects attribution-map granularity when needed. Brightfield achieved 0.6890 mAP50-95 (0.8865 mAP50) for detection and 0.9687 macro-F1 (0.9705 accuracy) for classification. Reconstructed holograms narrowed the gap with a task-dependent split: p-type was strongest for detection at 0.5324 mAP50-95 (0.8229 mAP50), while r-type was strongest for classification at 0.7695 macro-F1 (0.7866 accuracy), both far above raw-hologram baselines. Activation-based explanations localized strongly on grains, and region-based methods retained ~60 to ~80% of faithfulness under holography. The holographic detector was highly brittle to weak perturbations, saturating deletion-based evaluation while insertion remained informative. Pixel-level gradient explanations approached random floor, yet spatial smoothing restored p-type gradient faithfulness from 0.05 to 0.51. For holographic classification, perturbation-based explanations remained faithful while gradient-based methods fell below random floor. Reconstruction improves low-cost holographic pollen analysis, while AHIR distinguishes genuine attribution failure from artifacts caused by model brittleness and map granularity.
Chinese Translation
自动化花粉分析可支持兽医细胞学,但明场显微镜相比无透镜数字同轴全息显微镜成本更高、系统更复杂。我们评估了重建全息图能否缩小这一差距,以及模型解释在模态变化下是否仍然可靠。使用明场和全息显微镜对六种花粉进行成像。在基于锚框的标注迁移之后,采用YOLOv26s检测和MobileNetV4分类对原始全息图、单次反向传播全息图和迭代相位恢复全息图进行评估。采用归因健康检查与修复(Attribution Health Inspection and Repair, AHIR)协议,从空间定位性和忠实性两方面评估了六种归因方法;该协议可测试模型在弱噪声下的脆弱性,并在必要时修正归因图的粒度。明场显微镜在检测上达到0.6890 mAP50-95(0.8865 mAP50),在分类上达到0.9687宏F1(0.9705准确率)。重建全息图以任务相关的方式缩小了差距:p型重建在检测上表现最佳,达到0.5324 mAP50-95(0.8229 mAP50);r型重建在分类上表现最佳,达到0.7695宏F1(0.7866准确率),二者均远高于原始全息图基线。基于激活的解释在花粉颗粒上表现出强空间定位能力,基于区域的方法在全息成像下保留了约60%至80%的忠实性。全息检测器对弱扰动高度脆弱,删除式评估趋于饱和,而插入式评估仍具信息量。像素级梯度解释接近随机下限,但空间平滑将p型梯度忠实性从0.05提升至0.51。对于全息分类,基于扰动的解释仍保持忠实性,而基于梯度的方法则低于随机下限。重建技术改善了低成本全息花粉分析,而AHIR协议能够区分真正的归因失败与由模型脆弱性和归因图粒度引起的伪影。
cs.CV / 55 / 2609.23026

BrainIAC: Interactive 3D Brain Lesion Segmentation across Heterogeneous MRI Modalities with Online Adaptation

BrainIAC:基于在线自适应的跨异构MRI模态交互式3D脑部病灶分割
Xu, Wentian, Addison, Anthony P, Liang, Ziyun, Anthony, Harry, Yang, Guang, Kamnitsas, Konstantinos
Abstract
Brain lesion segmentation is a fundamental task in medical image analysis, playing a critical role in diagnosis, treatment planning, and longitudinal disease monitoring. Yet existing models still struggle to meet the demands of real clinical use, where deployments contain data distribution shifts, arising from differences in scanner hardware, imaging protocol (varying MRI modality sets), and new pathologies. We present BrainIAC (Brain lesion Interactive Adaptive Continuously learning segmentation), a unified framework that integrates (i) a multi-modal backbone network trained to segment multiple types of brain lesions and handle heterogeneous sets of modalities via zero-filling and random modality dropping; (ii) 3D interactive segmentation with bounding-box and click prompts that preserves fully automatic prediction when no prompt is given; and (iii) an online adaptation mechanism combining Mid-Interaction adaptation and Post-Interaction adaptation, supervised by the network's own predictions as pseudo labels and guided by an extra Click-Centered Gaussian loss. To our knowledge, this is one of the first 3D online adaptation methods for interactive segmentation, and the first to combine handling of heterogeneous modality sets with online adaptation. Experiments across seven brain MRI datasets demonstrate that the proposed components provide complementary and synergistic benefits. The method consistently outperforms existing approaches and generalizes well across heterogeneous imaging modalities, including those unseen during training, as well as previously unseen brain pathology types. The code and a 3D Slicer plug-in will be released at https://github.com/WenTXuL/BrainIAC upon publication.
Chinese Translation
脑部病灶分割是医学图像分析中的一项基础任务,在诊断、治疗规划和疾病纵向监测中发挥着关键作用。然而,现有模型仍难以满足真实临床应用的需求:实际部署中存在由扫描仪硬件、成像协议(不同的MRI模态集合)以及新出现的病理类型差异所导致的数据分布偏移。我们提出了BrainIAC(脑部病灶交互式自适应持续学习分割框架),这是一个统一框架,整合了:(i)一个多模态骨干网络,通过训练可分割多种类型的脑部病灶,并借助零填充(zero-filling)和随机模态丢弃来处理异构的模态集合;(ii)基于边界框和点击提示的3D交互式分割,在无提示输入时仍可保持完全自动的预测;(iii)一种在线自适应机制,结合交互中(Mid-Interaction)自适应和交互后(Post-Interaction)自适应,以网络自身的预测作为伪标签进行监督,并由额外的以点击为中心的高斯损失(Click-Centered Gaussian loss)加以引导。据我们所知,这是首批面向交互式分割的3D在线自适应方法之一,也是首个将异构模态集合处理与在线自适应相结合的方法。在七个脑部MRI数据集上的实验表明,所提出的各组件提供了互补且协同的增益。该方法持续优于现有方法,并在异构成像模态(包括训练中未见过的模态)以及未见过的脑部病理类型上均具有良好的泛化能力。代码和3D Slicer插件将在论文发表后于 https://github.com/WenTXuL/BrainIAC 发布。
cs.CV / 56 / 2609.23049

VDGS: Visibility-Driven Large-Scale 3D Gaussian Splatting for Aerial Scene Reconstruction

VDGS:面向航拍场景重建的可见性驱动大规模3D高斯泼溅方法
Yu, Haolin, Tang, Jiadong, Wang, YiXian, Gao, Yu, He, Shi, Lai, Zhilin, Yang, Yi, Fu, Mengyin
Abstract
Large-scale scene reconstruction is a critical foundational technology in robotic autonomous systems such as 3D mapping and autonomous driving. In recent years, 3D Gaussian Splatting (3DGS) has demonstrated remarkable advantages in both visual quality and computational efficiency, making it a promising representation for large-scale scene reconstruction. However, it still faces challenges in large-scale scenes, including excessive memory consumption and uneven viewpoint coverage caused by UAV acquisition, limiting its real-world applications. To address this, we propose VDGS, a novel 3DGS framework that incorporates camera distribution into scene modeling. VDGS introduces visibility-driven statistics for scene anchors to quantify supervision strength. These statistics are further leveraged for scene partitioning and for gradient compensation in under-optimized regions, thereby promoting balanced optimization across different regions. Extensive experiments on multiple large-scale aerial scene datasets demonstrate that, under imbalanced viewpoint distributions, VDGS consistently outperforms existing methods, while maintaining competitive performance in scenarios with more uniform view distributions.
Chinese Translation
大规模场景重建是三维建图和自动驾驶等机器人自主系统中的关键基础技术。近年来,3D高斯泼溅(3D Gaussian Splatting, 3DGS)在视觉质量和计算效率方面均展现出显著优势,使其成为大规模场景重建的一种有前景的表示方法。然而,3DGS在大规模场景中仍面临诸多挑战,包括过高的内存消耗以及无人机采集导致的视角覆盖不均匀等问题,限制了其实际应用。为此,我们提出了VDGS,这是一种将相机分布信息融入场景建模的新型3DGS框架。VDGS引入了面向场景锚点的可见性驱动统计量,用于量化监督强度。这些统计量进一步被用于场景划分以及对欠优化区域的梯度补偿,从而促进不同区域之间的均衡优化。在多个大规模航拍场景数据集上的大量实验表明,在视角分布不均衡的情况下,VDGS始终优于现有方法,同时在视角分布较为均匀的场景中也能保持具有竞争力的性能。
cs.CV / 57 / 2609.23061

HDMamba-YOLO: Efficient State-Space Perception and Local Spatial Reconstruction for UAV Small Object

HDMamba-YOLO:面向无人机小目标的高效状态空间感知与局部空间重建
Wei, Linduo, Fan, Junjie, Mai, Yijun, Qi, Yong
Abstract
Small-object detection in UAV imagery is challenged by weak visual evidence, ambiguous boundaries, dense object distributions, and complex backgrounds. Effective detection therefore requires long-range contextual information for target-background discrimination while preserving explicit local two-dimensional structures for accurate localization. These requirements arise at different stages of the detection pipeline and are not naturally addressed by a uniform feature-processing strategy. We propose Hybrid Dual-domain Mamba-YOLO (HDMamba-YOLO), a stage-wise heterogeneous SSM-CNN detector organized according to a perception-reconstruction-alignment-interaction rationale. EfficientVMamba-based EVSS establishes long-range contextual perception in the backbone, while PhasePatchMerging2D provides phase-aware hierarchical transitions. DST-Wrapper and Native C3k2-ASSAF then perform perception-to-reconstruction transition and repeated local two-dimensional reconstruction during FPN/PAN aggregation. DySample provides content-adaptive cross-scale resampling, while OS-CVTIA introduces macro-micro interaction and task-specific modulation for localization and classification. On VisDrone2019, HDMamba-YOLO-B achieves 42.737% mAP50 and 25.713% mAP50:95 with 10.042M parameters and 29.879 corrected GFLOPs. HDMamba-YOLO-Lite achieves 41.140% mAP50 and 24.741% mAP50:95 with 5.344M parameters. Under the unified AI-TOD evaluation protocol, HDMamba-YOLO-B obtains 21.621% AP and 47.881% AP50. Controlled ablations further support the stage-wise allocation of state-space perception, convolutional reconstruction, dynamic alignment, and task interaction for UAV small-object detection.
Chinese Translation
无人机图像中的小目标检测面临视觉证据微弱、边界模糊、目标分布密集以及背景复杂等挑战。因此,有效检测既需要长程上下文信息以区分目标与背景,又需要保留显式的局部二维结构以实现精确定位。这些需求出现在检测流程的不同阶段,单一均匀的特征处理策略难以自然地兼顾。我们提出混合双域Mamba-YOLO(HDMamba-YOLO),一种按“感知—重建—对齐—交互”原理组织的分阶段异构SSM-CNN检测器。基于EfficientVMamba的EVSS模块在主干网络中建立长程上下文感知,PhasePatchMerging2D提供相位感知的层级过渡。随后,DST-Wrapper与Native C3k2-ASSAF在FPN/PAN聚合过程中实现从感知到重建的过渡,并进行反复的局部二维重建。DySample提供内容自适应的跨尺度重采样,OS-CVTIA则引入宏观—微观交互以及面向定位与分类的任务特定调制。在VisDrone2019数据集上,HDMamba-YOLO-B以10.042M参数量和29.879修正GFLOPs实现了42.737%的mAP50和25.713%的mAP50:95;HDMamba-YOLO-Lite以5.344M参数量实现41.140%的mAP50和24.741%的mAP50:95。在统一的AI-TOD评估协议下,HDMamba-YOLO-B取得了21.621%的AP和47.881%的AP50。受控消融实验进一步验证了针对无人机小目标检测,按阶段分配状态空间感知、卷积重建、动态对齐和任务交互的合理性。
cs.CV / 58 / 2609.23067

LD-RSVIS: A Large-Scale and Diverse Benchmark for Referring Surgical Video Instrument Segmentation

LD-RSVIS:一个大规模、多样化的指代手术视频器械分割基准数据集
Wang, Zan, Feng, Yunhe, Nie, Dong, Oluwadare, Oluwatosin, Sha, Kewei, Huang, Yan, Fan, Heng
Abstract
Referring surgical video instrument segmentation (RSVIS) aims at segmenting the instrument in a surgical video, given a textual description. Despite recent progress, current models are trained and assessed on relatively small-scale benchmarks, hindering the development of more general RSVIS. In addition, existing benchmarks support only the single-target expression that refers to one instrument in the video, while overlooking multi-target and no-target referring expressions, restricting the applicability of RSVIS in practical scenarios. Addressing these issues, we propose LD-RSVIS, a new benchmark aiming to facilitate more robust and general RSVIS. Specifically, LD-RSVIS consists of 3,536 surgical videos with 1.09 million frames and covers a broad set of 30 instrument classes from 25 various procedures. By including abundant videos and classes, LD-RSVIS could benefit large-scale training and evaluation of more general RSVIS methods. Besides, unlike existing datasets, LD-RSVIS offers diverse referring settings, including no-target, single-target, and multi-target expressions, which enables the development of more practical RSVIS models in real applications. In order to ensure high-quality annotations, all videos in LD-RSVIS are manually labeled with multiple rounds of inspection and refinement. To our knowledge, LD-RSVIS is the largest and most diverse benchmark for RSVIS. To analyze LD-RSVIS and to provide comparison for future research, we evaluate 12 representative methods, and the results reveal that more efforts are required for improvements. To encourage future research, we present a simple yet effective RSVIS method, dubbed Cascade-RSVIS, that first mines target-specific cues using the complementary multi-cue text information and then employs such cues and textual information for segmentation, achieving promising performance. Our benchmark and code will be released.
Chinese Translation
指代手术视频器械分割(Referring Surgical Video Instrument Segmentation, RSVIS)旨在根据给定的文本描述对手术视频中的器械进行分割。尽管近期取得了一定进展,但现有模型仍在相对小规模的基准数据集上进行训练和评估,这阻碍了更具通用性的RSVIS方法的发展。此外,现有基准仅支持指向视频中单个器械的单目标指代表达,而忽略了多目标和无目标指代表达,限制了RSVIS在实际场景中的应用。针对这些问题,我们提出了LD-RSVIS,一个旨在促进更鲁棒、更通用的RSVIS发展的新基准数据集。具体而言,LD-RSVIS包含3,536段手术视频,共109万帧,涵盖来自25种不同手术操作的30类器械。通过包含丰富的视频和器械类别,LD-RSVIS能够为更通用的RSVIS方法的大规模训练与评估提供支持。此外,与现有数据集不同,LD-RSVIS提供了多样化的指代设置,包括无目标、单目标和多目标表达,从而有助于开发在真实应用中更具实用性的RSVIS模型。为保证高质量的标注,LD-RSVIS中的所有视频均经过人工标注,并进行了多轮检查与修正。据我们所知,LD-RSVIS是迄今为止规模最大、多样性最高的RSVIS基准数据集。为了分析LD-RSVIS并为未来研究提供比较基准,我们评估了12种代表性方法,结果表明该方法仍需进一步改进。为促进未来研究,我们提出了一种简单而有效的RSVIS方法,称为Cascade-RSVIS,该方法首先利用互补的多线索文本信息挖掘目标特定线索,然后利用这些线索和文本信息进行分割,取得了良好的性能。我们的基准数据集和代码将会公开发布。
cs.CV / 59 / 2609.23121

MM-ContextFold: Context Folding for Multimodal Agentic Retrieval

MM-ContextFold:面向多模态智能体检索的上下文折叠方法
Tian, Yang, Liu, Fan, Zhang, Jingyuan, Li, Zhenyang, Hu, Yupeng, Nie, Liqiang
Abstract
Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. While existing methods alleviate this issue by compressing redundant text, effective strategies for managing token-intensive visual content remain largely underexplored. To address this gap, we first conduct a systematic empirical study of approximately 10,000 trajectories. The results show that as visual cues are progressively extracted through external tools and textualized into the context, raw images become increasingly redundant. Continued image retention is associated with higher output entropy and can even degrade task accuracy. Motivated by these findings, we propose MM-ContextFold, a training-free framework that loads raw images only when needed. It maintains a persistent, text-only main context for high-level planning and spawns ephemeral branch contexts for image-dependent subtasks. Within each branch, the agent loads the relevant images, completes the subtask, and folds the result back into the main context as a concise textual summary; the images and branch trace are then discarded. Experiments on seven MAR benchmarks across five backbone models show that MM-ContextFold improves average accuracy by 6.3 percentage points over ReAct while reducing the working context length by 27.5\%.
Chinese Translation
多模态智能体检索要求智能体通过迭代调用外部工具来解决复杂的信息获取任务。ReAct 等典型框架将原始多模态输入和不断累积的交互历史维护在单一且持续增长的上下文中,从而导致上下文爆炸问题。尽管现有方法通过压缩冗余文本来缓解这一问题,但针对 token 密集型视觉内容的有效管理策略在很大程度上仍缺乏探索。为填补这一空白,我们首先对约 10,000 条轨迹进行了系统的实证研究。结果表明,随着视觉线索通过外部工具被逐步提取并以文本形式写入上下文,原始图像变得越来越冗余;持续保留图像与更高的输出熵相关,甚至会降低任务准确率。基于这些发现,我们提出了 MM-ContextFold,这是一种免训练框架,仅在需要时才加载原始图像。该框架维护一个持久的纯文本主上下文用于高层规划,并为依赖图像的子任务生成临时的分支上下文。在每个分支中,智能体加载相关图像、完成子任务,并将结果以简洁的文本摘要形式折叠回主上下文;随后图像和分支轨迹即被丢弃。在五个骨干模型上的七个 MAR 基准测试实验表明,MM-ContextFold 相比 ReAct 将平均准确率提升了 6.3 个百分点,同时将工作上下文长度减少了 27.5%。
cs.CV / 60 / 2609.23139

QwenVLConnector: A Fast, Unified Medical VLM Chatbot for Fine-Grained Clinical Perception and Text Generation

QwenVLConnector:一个用于细粒度临床感知与文本生成的快速统一医学视觉语言模型聊天机器人
Nguyen, Le Thien Phuc, Nguyen, Thien, Nguyen, Thanh-Huy, Hoang, Gia Minh, Vu, Anh Mai, Bagci, Ulas
Abstract
Most medical vision-language models (VLMs) excel at open-ended report generation and VQA but provide limited support for structured, fine-grained clinical perception within a unified interface. We present QwenVLConnector, a Qwen2.5-VL-based medical chatbot that unifies classification, multi-label classification, textualized detection, counting, regression, and free-form report generation under a single next-token objective. Our key component is a lightweight dense multi-layer Connector that aggregates low- and high-level visual features, aligns them through the pretrained vision Merger, and fuses them with the final visual representation without increasing sequence length. This design enriches visual tokens with complementary spatial and semantic cues while preserving efficiency. On FLARE-2D, QwenVLConnector improves detection F1 from 0.55 to 0.85, raises single-label classification from 0.37 to 0.51, and boosts report-generation GREEN by up to 18.3 points over the Qwen2.5-VL baseline. We further explore multimodal in-context learning for report generation, showing additional improvements without updating model parameters. Overall, QwenVLConnector offers a unified and efficient framework for combining structured medical perception with open-ended clinical text generation. Our code can be found at https://github.com/plnguyen2908/QwenConnector.
Chinese Translation
大多数医学视觉语言模型(VLM)擅长开放式报告生成和视觉问答(VQA),但在统一界面下对结构化、细粒度临床感知的支持有限。我们提出了QwenVLConnector,一个基于Qwen2.5-VL的医学聊天机器人,它将分类、多标签分类、文本化检测、计数、回归和自由格式报告生成统一在单一next-token目标之下。其关键组件是一个轻量级的多层密集Connector,用于聚合低层与高层视觉特征,通过预训练的视觉Merger进行对齐,并将其与最终视觉表示融合,同时不增加序列长度。这一设计以互补的空间和语义线索丰富了视觉token,同时保持了高效性。在FLARE-2D数据集上,QwenVLConnector将检测F1分数从0.55提升至0.85,将单标签分类准确率从0.37提升至0.51,并使报告生成的GREEN指标相比Qwen2.5-VL基线最高提升18.3分。我们进一步探索了用于报告生成的多模态上下文学习(in-context learning),在不更新模型参数的情况下展示了额外的改进。总体而言,QwenVLConnector提供了一个统一且高效的框架,将结构化医学感知与开放式临床文本生成相结合。我们的代码可在 https://github.com/plnguyen2908/QwenConnector 获取。
cs.CV / 61 / 2609.23153

SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to $265\times$ Single-GPU Acceleration of Visual Generation

SparkDiffusion:破解高稀疏度陷阱——一个实现视觉生成单卡最高265倍加速的统一框架
Liu, Yuxi, Li, Haoyu, Zhang, Zekun, Sun, Tengxu, Cai, Yixiang, Li, Jiayong, Xia, Yifei, Liu, Tianle, Ai, Baole, Wang, Ang, Wang, Jiamang, Qu, Lin, Zhang, Kai, Yuan, Kun, Cui, Bin
Abstract
Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}: at extreme attention sparsity, step-local training losses keep decreasing while terminal generation quality stagnates or degrades. The trap is one of supervision: the dominant terminal errors originate in the high-noise structure-generation stage, and terminal-aligned training corrects terminal errors that substantially extended step-local training cannot. This yields a simple staging principle: \emph{first adapt the sparse architecture into a coarse prior, then correct the terminal distribution}. We instantiate the principle as \method, a unified acceleration framework for visual generation that combines a short sparse warm-up, few-step trajectory-mixed distillation, and FP8 quantization with fused kernels. \method sustains $97\%$ attention sparsity with strong visual quality on long-sequence 720P generation across Wan2.1/Wan2.2 backbones and T2V/I2V tasks, and $90\%$ sparsity on Wan2.1-T2V-1.3B-480P. With 3-step CFG-free inference, \method achieves a $265\times$ end-to-end speedup over the 50-step CFG dense baseline for Wan2.1-T2V-14B-720P on a single RTX~5090 ($220\times$ on H100), and denoises a Wan2.1-T2V-1.3B-480P video in $1.3$s.
Chinese Translation
视频扩散变换器的计算开销之所以高昂,是因为注意力机制在长时空token序列上占主导地位。我们识别出了“高稀疏度陷阱”:在极端注意力稀疏度下,基于步骤局部的训练损失持续下降,而最终生成质量却停滞甚至退化。这一陷阱本质上是监督信号的问题:主要的最终误差源自高噪声的结构生成阶段,而与最终结果对齐的训练能够修正那些大幅延长步骤局部训练也无法纠正的最终误差。由此我们得到一个简单的分阶段原则:先将稀疏架构适配为一个粗略先验,再对最终分布进行修正。我们将该原则实现为SparkDiffusion,一个用于视觉生成的统一加速框架,它结合了短时稀疏预热、少步数轨迹混合蒸馏,以及带融合内核的FP8量化。SparkDiffusion在长序列720P生成任务上(覆盖Wan2.1/Wan2.2骨干网络及文生视频/图生视频任务) sustaining 97%的注意力稀疏度并保持出色的视觉质量,并在Wan2.1-T2V-1.3B-480P上实现90%稀疏度。通过3步无CFG推理,SparkDiffusion在单张RTX 5090上对Wan2.1-T2V-14B-720P实现了相对50步CFG稠密基线265倍的端到端加速(H100上为220倍),并将Wan2.1-T2V-1.3B-480P视频的去噪时间缩短至1.3秒。
cs.CV / 62 / 2609.23161

An Eternal Irradiance Camera

一种永恒辐照相机
Klotz, Jeremy, Nayar, Shree K.
Abstract
A conventional camera uses millions of pixels to measure radiance from all directions within its field of view. We present an omnidirectional irradiance camera that measures the irradiance function---the illumination incident upon every point on a sphere. The irradiance function varies smoothly over the sphere and hence is bandlimited. We have analyzed this function in the frequency domain and have shown that it is well approximated by a weighted sum of the first seven degrees of spherical harmonics. As a result, the irradiance function can be accurately reconstructed from a small number of samples. This implies that an irradiance camera does not need millions of detectors (pixels)---just a handful of measurements suffice. This brings two major benefits. First, the camera consumes such little energy that it can be completely powered by the light falling on its detectors. Second, it does not capture the visual details needed to identify an individual, and hence privacy is preserved. We have built a prototype irradiance camera, called FluxCam, using 49 detectors arranged on the surface of a sphere. In a well-lit indoor environment, FluxCam can read out and wirelessly transmit its measurements at 30 frames per second using energy harvested from the light falling on it (i.e., without a battery, cable, or external power supply). We show how FluxCam can be used as an optical gyroscope for computing rotation, to monitor a workspace, as an untethered light probe for diffuse relighting, and as an omnidirectional pyranometer for estimating sky conditions and determining the best orientation of a solar panel.
Chinese Translation
传统相机使用数百万个像素来测量其视场内来自各个方向的辐射亮度。我们提出了一种全向辐照相机,用于测量辐照度函数——即入射到球面上每个点的光照。辐照度函数在球面上平滑变化,因此是带限的。我们在频域中分析了该函数,并证明它可以很好地近似为前七阶球谐函数的加权和。因此,辐照度函数可以从少量的采样中精确重建。这意味着辐照相机不需要数百万个探测器(像素)——仅需少量测量即可。这带来了两大优势:其一,相机消耗的能量极低,完全可以由照射在其探测器上的光来供电;其二,它无法捕捉用于识别个人的视觉细节,从而保护了隐私。我们构建了一个名为 FluxCam 的辐照相机原型,使用布置在球面上的 49 个探测器。在光照良好的室内环境中,FluxCam 能够以每秒 30 帧的速率读取并通过无线传输其测量数据,所用的能量完全采集自照射其上的光(即无需电池、电缆或外部电源)。我们展示了 FluxCam 可用作计算旋转的光学陀螺仪、用于监控工作空间、用作扩散重光照(diffuse relighting)的无绳光照探针,以及用作估算天空状况和确定太阳能电池板最佳朝向的全向总日射表。
cs.CV / 63 / 2609.23169

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing

UltraTex:释放2K多视角扩散模型在3D纹理生成中的潜力
Zhang, Yibo, Yuan, Ze, Cao, Nan, Zhang, Li, Cao, Yan-Pei, Guo, Yuan-Chen, Ma, Rui
Abstract
High-quality texture generation is essential for creating realistic and production-ready 3D assets. Recent multi-view diffusion methods have shown promising results for image-guided 3D texturing, but they are typically constrained to low operating resolutions such as 512 or 768, making it difficult to preserve high-frequency details from high-resolution reference images. Scaling this paradigm to 2048 resolution is computationally prohibitive, as the unified multi-view sequence exceeds 212K tokens and incurs excessive memory and latency. In this paper, we present UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing. Our key observation is that object-centric multi-view renderings contain two major sources of redundancy: background-induced sequence redundancy and sparse token interactions within the foreground. To address them, we introduce Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence. To enable efficient foreground-only inference while avoiding reconstruction artifacts, we further design Foreground-Aware VAE Decoding to ensure the quality of the final high-resolution views. To satisfy the demanding data requirements of 2K-resolution multi-view diffusion training, we construct G-buffer TexVerse, a large-scale, ultra-high-resolution multi-view rendering dataset covering over 268,000 3D assets. Extensive experiments show that UltraTex generates visually faithful textures with rich fine-grained details, while substantially improving efficiency, achieving $20.6\times$--$91.1\times$ training speedup and $22.3\times$--$74.6\times$ end-to-end inference speedup over the baseline on common samples in our dataset. Code and data is at https://yiboz2001.github.io/UltraTex.
Chinese Translation
高质量的纹理生成对于创建逼真且可用于生产的3D资产至关重要。近期的多视角扩散方法在图像引导的3D纹理生成方面展现出良好的效果,但它们通常受限于较低的运行分辨率(如512或768),难以保留高分辨率参考图像中的高频细节。将该范式扩展到2048分辨率在计算上代价高昂,因为统一的多视角序列超过212K个token,并带来过多的内存消耗和延迟。在本文中,我们提出了UltraTex,一个高效的端到端框架,用于基于高分辨率多视角扩散的3D纹理生成。我们的关键观察是,以物体为中心的多视角渲染包含两个主要的冗余来源:背景引起的序列冗余以及前景内部稀疏的token交互。为解决这些问题,我们引入了背景token丢弃(Background Token Dropping),在DiT骨干网络之前移除背景token;以及块稀疏注意力(Block-Sparse Attention),用于减少对保留前景序列的注意力计算。为实现高效的前景专用推理并避免重建伪影,我们进一步设计了前景感知VAE解码(Foreground-Aware VAE Decoding),以确保最终高分辨率视角的质量。为满足2K分辨率多视角扩散训练对数据的苛刻需求,我们构建了G-buffer TexVerse,一个大规模、超高分辨率的多视角渲染数据集,涵盖超过268,000个3D资产。大量实验表明,UltraTex能够生成视觉上高度逼真且具有丰富细粒度细节的纹理,同时显著提升效率,在我们的数据集常见样本上,相比基线实现了20.6倍至91.1倍的训练加速和22.3倍至74.6倍的端到端推理加速。代码与数据见 https://yiboz2001.github.io/UltraTex。
cs.CV / 64 / 2609.23182

GrapeSplat: Geometry-Grounded Reconstruction via Amalgamated Pose-Free Encoding for Feed-Forward 3D Gaussian Splatting

GrapeSplat:基于几何锚定与融合无位姿编码的前馈式3D高斯泼溅重建
Lu, Si-Yu, Chen, Yung-Yao, Chen, Yi Jan, Li, Shang-Lin, Liao, Ching-Chan, Cheng, Wen-Huang
Abstract
Feed-forward 3D Gaussian Splatting now reconstructs renderable scenes from unposed, uncalibrated images. Yet, most models supervise only photometric consistency and predict Gaussians pixel by pixel, which leaves global structure fragile and ties primitive count to image resolution and view count. To this end, GrapeSplat amalgamates multi-view cues into a voxel-aligned scene representation and decodes Gaussians directly from the learned grid, requiring no per-scene optimization or post-processing. An Atlas Encoder lifts all views into pixel-wise geometry-and-appearance features anchored at predicted 3D points. PEACH-Vox compands the unbounded scene into a bounded sparse grid through a smooth per-axis map with an exact closed-form inverse. The Sparse Decoder then consolidates the grid with sparse convolutions and decodes the full scene as multiple Gaussians per occupied cell. This amalgamated representation exploits sparse voxel occupancy, where the Gaussian count follows the occupied cells and saturates as views cover the scene, while grid resolution sets its ceiling. GrapeSplat turns unposed images into a renderable Gaussian scene in a single forward pass. Trained with 2D and 3D supervision on 8-view sequences, it generalizes zero-shot from 4 to 64 views across indoor and unbounded scenes. Code and trained weights are available at https://github.com/VAISR/GrapeSplat
Chinese Translation
前馈式3D高斯泼溅(3D Gaussian Splatting)现已能够从无位姿、未标定的图像中重建可渲染场景。然而,大多数模型仅监督光度一致性并逐像素预测高斯,这使得全局结构脆弱,且将图元数量与图像分辨率和视角数量绑定。为此,GrapeSplat 将多视角线索融合为体素对齐的场景表示,并直接从学习到的网格中解码高斯,无需逐场景优化或后处理。Atlas Encoder(图集编码器)将所有视角提升为锚定在预测3D点上的像素级几何与外观特征。PEACH-Vox 通过具有精确闭式逆映射的平滑逐轴映射,将无界场景压缩至有界的稀疏网格中。稀疏解码器(Sparse Decoder)随后利用稀疏卷积整合该网格,并将完整场景解码为每个被占据单元对应多个高斯。这种融合表示利用稀疏体素占据特性,使高斯数量跟随被占据单元数并随视角覆盖场景而饱和,同时网格分辨率决定了其上限。GrapeSplat 能够在单次前向传播中将无位姿图像转化为可渲染的高斯场景。通过在8视角序列上使用2D和3D监督进行训练,模型在室内和无界场景中实现了从4到64视角的零样本泛化。代码与训练权重可在 https://github.com/VAISR/GrapeSplat 获取。
cs.CV / 65 / 2609.23184

CausalWM: Causal Chain-of-Thought Reasoning for Embodied World Model

CausalWM:面向具身世界模型的因果思维链推理
Xu, Ziming, Liang, Shuang, Han, Ruobing, Xi, Ziqiao, Rao, Mingxing, Zhou, Kun, Zhang, Zijun, Yan, Yuchen, Wei, Yufan, Huang, Junbo, Shao, Yifei, Nan, Fang, Huang, Biwei
Abstract
Embodied world models learn to predict future physical dynamics from visual observations and control signals, where physical knowledge is implicitly entangled within latent representations. We introduce CausalWM, a 16B embodied world model that performs explicit causal chain-of-thought reasoning before future video prediction. CausalWM organizes useful variables into a reasoning trajectory, allowing the model to progressively capture causal dependencies underlying physical evolution. To train CausalWM, we collect 31K hours embodied data and develop a three-stage paradigm consisting of large-scale video pre-training, causal CoT mid-training, and multi-objective RL post-training. Despite using only a limited set of supervised CoT variables, CausalWM exhibits emergent in-context learning capabilities, enabling contextual visual feature guidance and efficient few-step generation. CausalWM achieves state-of-the-art performance across language-conditioned, action-conditioned, single-view and multi-view benchmarks, including Top-1 performance on TriWorldBench leaderboard.
Chinese Translation
具身世界模型(Embodied World Model)学习根据视觉观测和控制信号预测未来的物理动态,其中物理知识隐式地纠缠于潜在表示之中。我们提出了 CausalWM,一个 16B 参数规模的具身世界模型,它在未来视频预测之前执行显式的因果思维链(Chain-of-Thought, CoT)推理。CausalWM 将有用的变量组织成一条推理轨迹,使模型能够逐步捕捉物理演化背后的因果依赖关系。为训练 CausalWM,我们收集了 3.1 万小时的具身数据,并开发了一个三阶段训练范式,包括大规模视频预训练、因果 CoT 中期训练和多目标强化学习后训练。尽管仅使用了有限的监督 CoT 变量集合,CausalWM 仍展现出涌现的上下文学习能力,实现了上下文视觉特征引导和高效的少步生成。CausalWM 在语言条件、动作条件、单视角和多视角基准测试中均取得了最先进的性能,并在 TriWorldBench 排行榜上获得 Top-1 成绩。
cs.CV / 66 / 2609.23248

SPACE: Semantic Projection and Alignment of CLIP Embeddings for Domain Adaptation

SPACE:面向领域自适应的CLIP嵌入语义投影与对齐方法
Manesco, João Renato Ribeiro, Jodas, Danilo Samuel, Rodrigues, Douglas, Passos, Leandro Aparecido, Papa, João Paulo
Abstract
A fundamental challenge in deploying vision models is domain shift, which arises when training and test data follow different distributions, leading to degraded performance. This challenge is amplified when the same semantic concept appears under distinct visual forms, such as photographs and sketches, where visual similarity is weak despite semantic correspondence. Existing unsupervised domain-adaptation methods aim to align distributions across domains but often ignore semantic relationships among samples of the same class. To address this issue, this paper introduces SPACE, a method that exploits the semantic structure of CLIP's vision-language space for domain adaptation. The key idea is to use text descriptions as semantic anchors by applying Singular Value Decomposition to CLIP embeddings of class descriptions, yielding an orthogonal basis that captures semantic relationships among categories. Visual features from both domains are projected into this semantic subspace, aligning images based on meaning rather than appearance.
Chinese Translation
视觉模型部署中的一个根本性挑战是领域偏移(domain shift),即当训练数据与测试数据遵循不同分布时,模型性能会随之下降。当同一语义概念以不同的视觉形式(如照片与素描)出现时,这一挑战尤为突出,因为此时视觉相似性较弱,而语义对应关系依然存在。现有的无监督领域自适应方法旨在对齐不同领域之间的分布,但往往忽略同一类别样本之间的语义关系。为解决这一问题,本文提出SPACE方法,该方法利用CLIP视觉-语言空间的语义结构来实现领域自适应。其核心思想是将文本描述作为语义锚点:通过对类别描述的CLIP嵌入应用奇异值分解(Singular Value Decomposition),得到一个能够刻画类别之间语义关系的正交基。随后,将来自两个领域的视觉特征投影到该语义子空间中,使图像基于语义含义而非外观进行对齐。
cs.CV / 67 / 2609.23249

Exact Quotients of Fresnel-Kummer Surfaces and Certified Biaxial Refraction

菲涅耳-库默尔曲面的精确商与经验证的双轴折射
Shaska, Tanush
Abstract
The Fresnel wave surface governs the propagation of light in a transparent biaxial crystal. It is a special Kummer quartic, and we identify it exactly. Over the complex numbers the wave surface of a crystal is the Kummer surface of the Jacobian of an explicit genus-two curve branched at the signed square roots of the three principal permittivities. This Jacobian is isogenous, by an isogeny with kernel of order four, to a product of two elliptic curves. One elliptic curve carries the three permittivities, and the other carries the optic-axis angle. The physical family is Zariski dense in the locus of genus-two curves with an extra involution, and its automorphism strata are explicit. The identification instantiates a task-aware quotient, which identifies parameters that differ by a nuisance transformation and carries invariant coordinates and explicit strata. For biaxial crystals, two ratios of the permittivities form a complete invariant of the wave surface up to rotation and rescaling, and the four real nodes are given in closed form. At an interface the candidate transmitted waves are the roots of a quartic of exact degree four. Its real-root count, root order, and repeated-root events are decided by exact algebraic predicates, and along the generic single-node encounters of the paper its discriminant vanishes to second order. Floating-point solvers drop forward transmitted modes near the optic axes, and the certified solver does not. On exact equivalence classes, learned models on quotient coordinates are invariant and more accurate than models on raw tensors, while learned root-count predicates fail near the optic axes.
Chinese Translation
菲涅耳波面(Fresnel wave surface)支配着光在透明双轴晶体中的传播。它是一个特殊的库默尔四次曲面(Kummer quartic),我们对其给出了精确的识别。在复数域上,晶体的波面是由一条显式给出的亏格二曲线的雅可比簇所对应的库默尔曲面,该曲线在三个主介电常数的带符号平方根处分歧。该雅可比簇通过一个核为四阶的同源映射同源于两条椭圆曲线的乘积:其中一条椭圆曲线承载三个介电常数,另一条承载光轴角。该物理族在具有额外对合的亏格二曲线的轨迹中是扎里斯基稠密的,并且其自同构分层是显式的。这一识别实现了一种面向任务的商(task-aware quotient),它能识别仅相差一个干扰变换的参数,并携带不变坐标与显式分层。对于双轴晶体,两个介电常数的比值构成波面在旋转与缩放意义下的完备不变量,且四个实节点具有闭式表达式。在界面上,候选透射波是一个精确次数为四的四次方程的根。其实根个数、根的排序以及重根事件均由精确的代数谓词判定,并且沿本文所研究的的一般单节点情形,其判别式二阶消失。浮点求解器在光轴附近会丢失前向透射模式,而经验证的求解器则不会。在精确等价类上,基于商坐标训练的学习模型具有不变性,且比基于原始张量的模型更为精确,而学习得到的根数谓词在光轴附近则会失效。
cs.CV / 68 / 2609.23268

Blind Deconvolution of Binary and Pattern Images with Pixel Intensity Constraints and Sparse Gradient Prior

基于像素强度约束与稀疏梯度先验的二值图像与模式图像盲反卷积
Zhang, Qinghua, Yang, Xuesong, He, Liangtian, Deng, Liang-jian, Liu, Jun
Abstract
Blind image deconvolution (BID) is a prominent research topic in the field of imaging sciences, given its significant practical applications. Most existing model-based BID methods focus on natural images, incorporating appropriate prior knowledge about both the underlying image and the blur kernel. However, for certain classes of images, such as barcodes, text, and patterns, pixels can only take very limited values, a specific prior that is often overlooked in the literature. In this article, we introduce a novel pixel intensity constraint to leverage this important information, improving recovery performance for these specialized image classes. Specifically, we propose a unified framework for blind binary and pattern image deconvolution that incorporates both the pixel intensity constraint and a gradient sparsity regularizer. Numerical experiments demonstrate that our method outperforms many existing BID techniques, achieving superior results in terms of both visual quality and quantitative metrics.
Chinese Translation
盲图像反卷积(Blind Image Deconvolution, BID)因其重要的实际应用价值,成为成像科学领域的一个热门研究课题。现有大多数基于模型的BID方法主要针对自然图像,融合了关于潜在图像和模糊核的适当先验知识。然而,对于某些特定类别的图像,如条形码、文本和图案,像素只能取非常有限的值,而这一特定先验在文献中往往被忽视。本文引入了一种新颖的像素强度约束以利用这一重要信息,从而提升这些特殊图像类别的恢复性能。具体而言,我们提出了一个统一的盲二值图像与模式图像反卷积框架,该框架同时结合了像素强度约束和梯度稀疏正则化项。数值实验表明,我们的方法优于许多现有的BID技术,在视觉质量和定量指标方面均取得了更优的结果。
cs.CV / 69 / 2609.23286

RegVGGT: Sustainable Visual Geometry Grounding for Streaming via Regulated Memory

RegVGGT:基于受控内存的可持续流式视觉几何定位
Mao, Hongbo, Jiang, Junjun, Chen, Youyu, Zhang, Jiaxin, Dong, Zhemeng, Liu, Xianming
Abstract
3D reconstruction from a lengthy video stream input poses a dilemma for feed-forward reconstruction models (FFRMs), that a whole-stream inference context cannot be retained under limited GPU memory.Recent studies seek to resolve this problem via a trade-off between the integrity of inference context and GPU memory usage, which either suffer from a rapid memory inflation or degraded context integrity due to artificially capping memory usage.Driven by our key observation that the initial saliency of a token reliably dictates its long-term importance across the stream, we propose RegVGGT, a training-free token regulation method which aggressively regulates the tokens of incoming frames.By admitting at most 1% of tokens per frame to update the context memory, our method dramatically suppresses memory inflation as the stream progresses.Equipped with a FlashAttention-compatible token saliency estimation scheme, RegVGGT is capable of processing thousands of frames on a consumer-grade GPU with negligible compromise to reconstruction quality.Extensive experiments demonstrate that RegVGGT achieves state-of-the-art performance on long-horizon benchmarks across diverse FFRM prediction tasks, surpassing prior FFRM-based stream reconstruction baselines by a large margin.
Chinese Translation
从长视频流输入中进行3D重建对前馈重建模型(Feed-forward Reconstruction Models,FFRMs)而言是一个难题,即在有限的GPU内存下无法保留整个视频流的推理上下文。近期研究试图通过在推理上下文完整性与GPU内存占用之间进行权衡来解决该问题,但这类方法要么面临内存的快速膨胀,要么因人为限制内存使用而导致上下文完整性下降。基于我们的关键观察——token的初始显著性能够可靠地预示其在整个视频流中的长期重要性,我们提出了RegVGGT,这是一种无需训练的token调控方法,可对传入帧的token进行激进的调控。通过每帧最多允许1%的token来更新上下文内存,我们的方法显著抑制了随视频流推进而产生的内存膨胀。借助与FlashAttention兼容的token显著性估计方案,RegVGGT能够在消费级GPU上处理数千帧,而对重建质量的影响微乎其微。大量实验表明,RegVGGT在多种FFRM预测任务的长时序基准上取得了最先进的性能,大幅超越以往基于FFRM的流式重建基线方法。
cs.CV / 70 / 2609.23336

MinCU: A Fine-Grained Benchmark for Grounded Minimal-Change Understanding in Image Pairs

MinCU:面向图像对中基于定位的最小变化理解的细粒度基准
Mu, Chaoqian, Wu, Wenhao, Liang, Zichen, Li, Jiaxu, Wang, Lijun, Wang, Yifan, Lu, Huchuan
Abstract
Localizing and describing fine-grained differences between near-identical images is a critical yet underexplored capability for multimodal large language models (MLLMs). Existing benchmarks largely assess semantic comparison or single-image grounding in isolation, without jointly requiring faithful description and physical localization. To bridge this gap, we introduce MinCU, a benchmark for grounded minimal-change understanding, where each sample consists of an image pair differing by a single atomic variation in object category, attribute, count, or spatial position, and models are evaluated on their ability to describe the change, localize the changed regions, and identify the changed entity. We further propose Semantic-Guided Implicit Spatial Anchors (SG-ISA), a structured autoregressive method that decomposes prediction into a Think-Locate-Describe sequence. SG-ISA first predicts a semantic cue for the changed concept, then uses discrete spatial anchors as an implicit localization scaffold, and finally generates the change description together with the grounding box. Experiments reveal that even the strongest closed-source MLLMs and recent R1-style reasoning models struggle on MinCU, with most failing to jointly produce accurate descriptions and grounding boxes. Compared to the previous chain-of-thought method, fine-tuning with SG-ISA yields substantial joint improvements in grounding accuracy and description quality while reducing reasoning-token overhead by approximately 26%. These results suggest that an implicit intermediate spatial interface can be more effective than relying solely on model scale for grounded dual-image understanding.
Chinese Translation
对近乎相同的图像进行细粒度差异的定位与描述,是多模态大语言模型(MLLM)一项关键却未被充分探索的能力。现有基准大多孤立地评估语义比较或单图像定位,而未同时要求忠实描述与物理定位。为弥补这一空白,我们提出MinCU,一个面向基于定位的最小变化理解的基准。每个样本由一对图像组成,二者仅在物体类别、属性、数量或空间位置上存在单个原子级变化;模型需要描述该变化、定位发生变化的区域并识别被改变的实体。我们进一步提出语义引导的隐式空间锚点(Semantic-Guided Implicit Spatial Anchors, SG-ISA),这是一种结构化自回归方法,将预测分解为“思考—定位—描述”的序列:SG-ISA首先为变化概念预测语义线索,然后使用离散空间锚点作为隐式定位支架,最后生成变化描述及定位框。实验表明,即使是最强的闭源MLLM以及近期的R1风格推理模型,在MinCU上也表现不佳,大多数模型无法同时生成准确的描述和定位框。与之前的思维链方法相比,使用SG-ISA进行微调在定位精度和描述质量上均带来显著的联合提升,同时将推理token开销降低约26%。这些结果表明,隐式的中间空间接口可以比单纯依赖模型规模更有效地实现基于定位的双图像理解。
cs.CV / 71 / 2609.23345

AniPrO: Interpretable Anime Image Provenance Detection via Multi-Dimensional Semantic Reasoning

AniPrO:基于多维语义推理的可解释动漫图像来源检测
Liu, Yan, Huang, Baoxiang, Wang, Zi'an, Xie, Wenbo
Abstract
As generative AI becomes increasingly used in anime-style image creation, distinguishing human-drawn, AI-inpainted, and text-to-image images is important for copyright attribution, visual provenance, and content governance. Existing AI-generated image detectors mainly target real-world photographs and often overlook anime-specific cues such as flat coloring, exaggerated structures, and artistic line control. To address this gap, we propose AniPrO, a multi-dimensional description-enhanced framework for interpretable anime image provenance. Built upon AnimeDL-2M, AniPrO contains 15,000 balanced samples from a 35,000-image candidate pool, covering Real, Inpainting, and Text2Image categories with structured five-dimensional descriptions. We further introduce AniPrO-SFD-Bench and AniPrO-MFR-Bench to evaluate provenance detection from statistical feature discrimination and multimodal fusion reasoning perspectives. Experiments show that structured semantic guidance reveals systematic AI-generation biases, such as the gap between global visual plausibility and local detail coherence, and improves the detection of challenging inpainting samples. The dataset and code will be released at: https://github.com/YAN-LIU05/AniPrO.
Chinese Translation
随着生成式人工智能在动漫风格图像创作中的日益普及,区分人类绘制图像、AI修复(inpainting)图像和文生图(text-to-image)图像,对于版权归属、视觉来源追踪和内容治理具有重要意义。现有的AI生成图像检测方法主要针对真实世界照片,往往忽略动漫图像特有的线索,如平涂上色、夸张的结构以及艺术化的线条控制。为填补这一空白,我们提出了AniPrO,一个面向可解释动漫图像来源检测的多维描述增强框架。AniPrO基于AnimeDL-2M构建,包含来自35,000张候选图像池的15,000个均衡样本,涵盖真实图像(Real)、图像修复(Inpainting)和文生图(Text2Image)三个类别,并配有结构化的五维描述。我们进一步引入AniPrO-SFD-Bench和AniPrO-MFR-Bench两个基准,分别从统计特征判别和多模态融合推理的角度评估来源检测性能。实验表明,结构化语义引导能够揭示系统性的AI生成偏差,例如全局视觉合理性与局部细节连贯性之间的差距,并提升了对具有挑战性的修复样本的检测效果。数据集和代码将发布于:https://github.com/YAN-LIU05/AniPrO。
cs.CV / 72 / 2609.23352

BiView-Touch: Learning Bimanual Tactile Representations by Cross-Hand Completion

BiView-Touch:通过跨手补全学习双手触觉表征
Liang, Chenxin, Lai, Youchen, Lyu, Chuqiao, Chen, Tianxing, Li, Shoujie, Ding, Wenbo
Abstract
Bimanual interaction produces complementary tactile views of the same physical process, yet existing tactile representation learning largely models the two hands independently or combines them only for downstream prediction, leaving their cross-hand relationship unexplored. To exploit this overlooked structure, we introduce BiView-Touch, a tactile-only framework that completes masked target-hand latents from the remaining visible target-hand regions and the synchronized full contralateral hand. A student encoder with a geometry-conditioned directional decoder predicts full-view EMA latent targets, while temporal and layout counterfactuals encourage sensitivity to synchronized and anatomically organized source information. Controlled ablations and source-context interventions show that BiView-Touch learns structured cross-hand dependence on temporally aligned and anatomically organized contralateral tactile context, rather than benefiting from bilateral input alone. On the public HumanTouch dataset, its frozen representations consistently outperform representative self-supervised baselines across low-label settings. With only 5\% downstream labels, BiView-Touch achieves relative balanced-accuracy gains of 7.1\% on bilateral wrist-motion recognition and 14.1\% on force-derived interaction-phase recognition. We further introduce BVT-20, a 20-task bilateral tactile dataset, and demonstrate transfer across recording sessions and pretraining corpora, including transfer to a held-out bimanual task. Our code and dataset details are available on the anonymous project page: https://anonymous.4open.science/w/biview-touch-review-site-050C/.
Chinese Translation
双手交互会产生对同一物理过程的互补触觉视图,然而现有的触觉表征学习大多将两只手独立建模,或仅在下游预测中才将二者结合,其跨手关系尚未被探索。为利用这一被忽视的结构,我们提出了 BiView-Touch,一个仅基于触觉的框架,该框架利用剩余的可见目标手区域以及同步完整的对侧手信息来补全被掩码的目标手潜在表征。一个带有几何条件方向解码器的学生编码器(student encoder)用于预测全视图的EMA潜在目标,同时时间与布局反事实(counterfactual)机制促使模型对同步且按解剖结构组织的源信息保持敏感。受控消融实验与源上下文干预表明,BiView-Touch 学到的是对时间对齐且解剖结构组织的对侧触觉上下文的结构化跨手依赖,而非仅仅得益于双侧输入。在公开的 HumanTouch 数据集上,其冻结表征在低标签设置下持续优于具有代表性的自监督基线方法。在仅有 5% 下游标签的情况下,BiView-Touch 在双侧手腕动作识别上取得了 7.1% 的相对平衡准确率提升,在基于力的交互阶段识别上取得了 14.1% 的提升。我们进一步发布了 BVT-20,一个包含 20 个任务的双侧触觉数据集,并展示了跨录制会话与跨预训练语料的迁移能力,包括向保留的双手机器任务(held-out bimanual task)的迁移。我们的代码与数据集详情见匿名项目页面:https://anonymous.4open.science/w/biview-touch-review-site-050C/。
cs.CV / 73 / 2609.23369

The Right Future for Action: Learning Action-Relevant Predictive States in World Action Models

面向动作的正确未来:在世界动作模型中学习动作相关的预测状态
Gu, Qiwen, Li, Jifan, Gao, Bingjie, Chen, Rui, Tang, Jing, Chu, Xiangxiang, Zhao, Junqiao
Abstract
Generation-free world action models (WAMs) retain future-video prediction during training but act from internal video features at inference, leaving unclear what these features should preserve for control. Our representation diagnostics show that representations with more predictable future changes need not make linear action decoding easier. Observed future changes provide additional action information beyond the present, and linearly readable action information is spatially concentrated. These findings motivate Action-Relevant Predictive States (ARPS), a compact predictive interface between the video and action experts. ARPS uses a horizon-conditioned state predictor to aggregate intermediate video features into a compact state that supplies all visual context to the action expert. Future-representation supervision trains different parts of this state to predict visual representations at different future times, together with their changes relative to the present. At inference, the supervision branch is removed, and the action expert only uses the learned predictive state computed from current observations. Controlled ablations show that future supervision substantially improves generalization under distribution shift. ARPS achieves 99.2% success on LIBERO and transfers to LIBERO-Plus without adaptation, reaching 87.3% and exceeding Fast-WAM by 39.2 percentage points.
Chinese Translation
免生成式世界动作模型(World Action Models, WAMs)在训练阶段保留未来视频预测,但在推理时仅依靠内部视频特征进行动作决策,而这些特征应为控制保留什么信息尚不清楚。我们的表征诊断实验表明,未来变化越可预测的表征,并不一定使线性动作解码更容易。观测到的未来变化提供了超出当前时刻的额外动作信息,且可线性读取的动作信息在空间上是集中的。这些发现启发我们提出动作相关预测状态(Action-Relevant Predictive States, ARPS),一个位于视频专家与动作专家之间的紧凑预测接口。ARPS 使用一个条件于预测时域的状态预测器,将中间视频特征聚合为一个紧凑状态,为动作专家提供全部视觉上下文。未来表征监督训练该状态的不同部分去预测不同未来时刻的视觉表征,及其相对当前时刻的变化。在推理时,监督分支被移除,动作专家仅使用由当前观测计算得到的所学预测状态。受控消融实验表明,未来监督显著提升了分布偏移下的泛化能力。ARPS 在 LIBERO 上达到 99.2% 的成功率,并在无需适配的情况下迁移至 LIBERO-Plus,达到 87.3%,超过 Fast-WAM 39.2 个百分点。
cs.CV / 74 / 2609.23372

Vision-Wireless Fusion for Multi-User Localization: A Cross-Modal Transformer Approach

面向多用户定位的视觉-无线融合:一种跨模态Transformer方法
Zheng, Can, He, Jiguang, Cai, Guofa, Wymeersch, Henk, Debbah, Merouane
Abstract
Accurate multi-user localization is challenging in complex urban environments, where wireless measurements can become ambiguous under noise, blockage, and multipath, while visual observations provide complementary spatial context. This paper presents a vision-wireless fusion framework for multi-user localization using pilot-indexed channel state information (CSI). Orthogonal pilot indices preserve the identities of the communicating UEs in the CSI token sequence and localization outputs. The model encodes each pilot-indexed CSI observation as a query token and uses cross-attention to retrieve user-specific information from spatial visual memory. Self-attention among CSI tokens further captures inter-user interactions, while the resulting multimodal representations are used for user-wise localization. Experiments on different datasets show consistent improvements over model-based, CSI-only, and multimodal-fusion baselines. Further experiments evaluate the model under different wireless and visual conditions.
Chinese Translation
在复杂的城市环境中,精确的多用户定位具有挑战性:无线测量在噪声、遮挡和多径效应下可能变得模糊,而视觉观测则提供了互补的空间上下文信息。本文提出了一种基于导频索引信道状态信息(CSI)的视觉-无线融合多用户定位框架。正交导频索引在CSI令牌序列和定位输出中保持了通信用户设备(UE)的身份信息。该模型将每个导频索引的CSI观测编码为查询令牌(query token),并通过交叉注意力机制从空间视觉记忆中检索用户特定信息。CSI令牌之间的自注意力机制进一步捕捉用户间交互,所得的多模态表示则用于逐用户定位。在不同数据集上的实验表明,该方法相较于基于模型的、仅使用CSI的以及多模态融合的基线方法均取得了一致的性能提升。进一步的实验评估了该模型在不同无线和视觉条件下的表现。
cs.CV / 75 / 2609.23380

LiteTex-GS: Fast and Lightweight Texturing for Gaussian Splatting

LiteTex-GS:面向高斯泼溅(Gaussian Splatting)的快速轻量级纹理化方法
Li, Zhiwei, Guo, Yijia, Lu, Yishi, Hu, Liwen, Rao, Hong, Chen, Shengbo, Ma, Lei
Abstract
Gaussian Splatting has enabled real-time novel view synthesis, but its tightly coupled geometry and appearance representation often require a large number of primitives to reproduce high-frequency texture details, leading to substantial memory and optimization costs. Recent textured 2D Gaussian methods alleviate this limitation by attaching texture maps to Gaussian primitives. However, bridging the fundamental structural gap between discrete Gaussians and continuous 2D grids requires complex parameterizations that introduce severe computational overhead. This overhead fundamentally compromises the original efficiency of Gaussian Splatting, making the balance between detailed texturing and computational agility an unresolved challenge. To address these challenges, we propose LiteTex-GS, a fast and lightweight texturing framework for Gaussian Splatting. Our method initializes an extremely compact representation, assigning minimal local texture to each Gaussian and progressively allocates higher resolution only to primitives with significant reconstruction errors. To maintain a streamlined geometric scaffold, we introduce a contribution- and area-aware pruning strategy that eliminates low-utility Gaussians. Furthermore, to mitigate the gradient dilution caused by texture upsampling, we design a resolution-aware update rule that preserves rapid and stable convergence. Extensive experiments on standard novel view synthesis benchmarks demonstrate that our method achieves competitive or superior rendering quality while using substantially fewer parameters and less training time than existing textured Gaussian baselines.
Chinese Translation
高斯泼溅(Gaussian Splatting)实现了实时的新视角合成,但其几何与外观表示紧密耦合,通常需要大量图元才能再现高频纹理细节,从而导致可观的内存与优化开销。近期的带纹理二维高斯方法通过为高斯基元附加纹理图来缓解这一局限。然而,弥合离散高斯与连续二维网格之间的根本性结构差异需要复杂的参数化,进而引入严重的计算开销。这一开销从根本上损害了高斯泼溅原有的高效性,使得精细纹理化与计算敏捷性之间的平衡成为一个尚未解决的难题。为应对这些挑战,我们提出了 LiteTex-GS,一个面向高斯泼溅的快速轻量级纹理化框架。我们的方法初始化一种极度紧凑的表示,为每个高斯分配最小的局部纹理,并仅为具有显著重建误差的图元逐步分配更高分辨率。为保持精简的几何骨架,我们引入了一种兼顾贡献度与面积的剪枝策略,以剔除低效用的高斯。此外,为缓解纹理上采样导致的梯度稀释问题,我们设计了一种分辨率感知的更新规则,以保持快速而稳定的收敛。在新视角合成标准基准上的大量实验表明,与现有带纹理高斯基线方法相比,我们的方法在参数量更少、训练时间更短的情况下,取得了相当或更优的渲染质量。
cs.CV / 76 / 2609.23386

ProxyBuild: Text-Guided Structured 3D Building Generation with Mesh-Anchored Procedural Proxies

ProxyBuild:基于网格锚定程序化代理的文本引导结构化三维建筑生成
Tang, Xiang, Li, Ruotong, Fan, Xiaopeng
Abstract
Text-guided 3D building generation holds tremendous application potential, yet existing generative models typically output inseparable single meshes or non-interactive rendered representations. While procedural modeling can generate editable buildings with hierarchical structures, rule authoring is laborious, and even with the aid of large language models (LLMs), it remains challenging to effectively solve procedural rules under geometric constraints. In this paper, we propose ProxyBuild, a novel hybrid framework for structured building generation. We introduce the Mesh-Anchored Procedural Proxy (MAPP) as a novel intermediate representation, which tightly anchors building components onto geometric shells, thereby decoupling the generation task into two phases: proxy prediction and proxy-to-asset instantiation. First, we construct a building dataset with MAPP annotations to train our designed face-edge bigraph encoder. By explicitly modeling the feature interactions of topological elements on heterogeneous mesh graphs, this encoder accurately infers the semantic roles of faces and edges. Subsequently, conditioned on textual styles and attribute parameters parsed by LLMs, we accomplish high-precision asset retrieval and assembly by integrating a spatial placement logic with hard constraints. Extensive experiments show that ProxyBuild not only significantly mitigates common issues in building generation such as over-smoothing, component collisions, and structural corruptions, but also accurately parses semantic-free shells from diverse sources. Outperforming prior baselines across various metrics, our method can robustly generate structurally clear, detail-rich, and post-editable 3D buildings from text, thereby providing a reliable and interactive content foundation for downstream applications such as virtual reality and digital twins.
Chinese Translation
文本引导的三维建筑生成具有巨大的应用潜力,然而现有的生成模型通常输出不可分割的单一网格或不可交互的渲染表示。程序化建模虽然能够生成具有层次结构的可编辑建筑,但规则编写十分繁琐,即便借助大语言模型(LLM),在几何约束下有效求解程序化规则仍然具有挑战性。本文提出ProxyBuild,一种新颖的结构化建筑生成混合框架。我们引入网格锚定程序化代理(Mesh-Anchored Procedural Proxy,MAPP)作为新的中间表示,将建筑构件紧密锚定在几何外壳上,从而将生成任务解耦为两个阶段:代理预测与代理到资产的实例化。首先,我们构建了带有MAPP标注的建筑数据集,用于训练我们设计的面-边二部图编码器(face-edge bigraph encoder)。通过显式建模异构网格图上拓扑元素的特征交互,该编码器能够准确推断面和边的语义角色。随后,以LLM解析的文本风格和属性参数为条件,通过融合具有硬约束的空间放置逻辑,我们实现了高精度的资产检索与组装。大量实验表明,ProxyBuild不仅显著缓解了建筑生成中常见的过度平滑、构件碰撞和结构损坏等问题,还能准确解析来自多种来源的无语义外壳。我们的方法在多项指标上均优于现有基线,能够从文本稳健地生成结构清晰、细节丰富且可后期编辑的三维建筑,从而为虚拟现实和数字孪生等下游应用提供可靠且可交互的内容基础。
cs.CV / 77 / 2609.23397

Enhancing Shrimp Disease Detection via Deep Learning and Data Refinement for Resilient Aquaculture

通过深度学习与数据精炼提升虾病检测能力,助力韧性水产养殖
Truong, Vinh Canh-Thanh, Pham, Hai-Binh, Tran, Ngoc Hong
Abstract
Shrimp diseases continue to cause devastating losses in the aquaculture industry, driving a critical need for robust, automated detection. This work contributes the first application of Vision Transformers (ViT) and Self-Supervised Learning (SSL) to the shrimp farming domain, addressing both performance bottlenecks and data labeling challenges. We propose two deep learning pipelines to classify four key diseases: Healthy, Black Gill (BG), White Spot Syndrome Virus (WSSV), and a co-infection of both using a dataset of 4,348 images. First, our supervised transfer-learning approach leverages ImageNet-pretrained ViT-Small/16 and EfficientNet backbones. Second, we introduce a contrastive learning framework (SimCLR) with a ViT-Small encoder to extract robust representations from unlabeled images prior to fine-tuning. Our results establish strong new baselines for sustainable aquaculture monitoring. The supervised approach achieves an outstanding 96% accuracy with fast convergence, outperforming traditional generic models, while the label-efficient SSL approach reaches a highly competitive 85% validation accuracy.
Chinese Translation
虾类疾病持续给水产养殖业带来毁灭性损失,因此迫切需要稳健的自动化检测方法。本工作首次将视觉Transformer(Vision Transformers, ViT)和自监督学习(Self-Supervised Learning, SSL)应用于对虾养殖领域,同时解决了性能瓶颈和数据标注难题。基于包含4,348张图像的数据集,我们提出了两种深度学习流程,对四种关键状态进行分类:健康(Healthy)、黑鳃病(Black Gill, BG)、白斑综合征病毒(White Spot Syndrome Virus, WSSV)以及二者的混合感染。首先,我们的有监督迁移学习方法利用ImageNet预训练的ViT-Small/16和EfficientNet骨干网络;其次,我们引入了一种对比学习框架(SimCLR),采用ViT-Small编码器在微调之前从无标注图像中提取鲁棒的特征表示。我们的结果为可持续水产养殖监测确立了强有力的新基线。有监督方法实现了高达96%的优异准确率且收敛迅速,优于传统的通用模型;而标注高效的SSL方法也达到了极具竞争力的85%验证准确率。
cs.CV / 78 / 2609.23404

ScaleBlind: Point Cloud Completion under Unknown Scale

ScaleBlind:未知尺度下的点云补全
Wu, Shenghui, Wang, Chen, Feng, Yuan, Wei, Guangshun, Zhou, Yuanfeng, Li, Changjian
Abstract
Point cloud completion aims to infer a complete 3D shape from a partial point cloud and serves as a fundamental building block for downstream tasks such as reconstruction, editing, and simulation. Despite the recent progress, existing learning-based methods often implicitly rely on access to the ground-truth shape scale (GT-scale) during both training- and testing-time normalization, assuming privileged information that is unavailable in real-world inference. This hidden assumption limits practical deployment and can lead to severe completion artifacts, e.g., over- or under-completion and nested shells, once the oracle GT-scale cue is removed. We observe that the recent foundation image generation models exhibit a strong capability of understanding objects and geometries, and producing multi-view consistent renderings, making them promising priors for GT-scale-free 3D completion. Motivated by this insight, we propose ScaleBlind, a novel framework that leverages foundation-model-based image completion to recover global scale directly from partial inputs and then faithfully produces the 3D completion. Specifically, ScaleBlind dreams out complete multi-view appearances from rendered partial views, lifts the inferred missing regions back into 3D to obtain a geometry-aware coarse completion, and further refines it via a powerful cross-modal fusion network with the original partial point cloud. By harnessing 2D foundation priors, our method eliminates the need for accessing GT-scale information at inference. Moreover, it provides a principled bridge between 2D generative priors and 3D point cloud completion. Extensive experiments demonstrate the superiority of our framework, making ScaleBlind the new state-of-the-art for the point cloud completion task.
Chinese Translation
点云补全旨在从不完整的点云中推断出完整的三维形状,是重建、编辑和仿真等下游任务的基础模块。尽管近年来取得了进展,现有基于学习的方法在训练和测试阶段的归一化中往往隐式依赖于真实形状尺度(GT尺度)这一特权信息,而该信息在真实世界的推理中是无法获得的。这一隐藏假设限制了方法的实际部署,并且一旦移除GT尺度这一先验线索,便会导致严重的补全伪影,例如过度补全或欠补全以及嵌套壳层。我们观察到,近来的基础图像生成模型展现出对物体和几何结构的强大理解能力,并能生成多视角一致的渲染结果,使其成为无需GT尺度的三维补全的有前景的先验。基于这一洞察,我们提出了ScaleBlind——一个新颖的框架,它利用基于基础模型的图像补全直接从部分输入中恢复全局尺度,进而忠实地生成三维补全结果。具体而言,ScaleBlind从渲染的部分视图出发,‘构想’出完整的多视角外观,将推断出的缺失区域提升回三维空间以获得几何感知的粗补全,再通过强大的跨模态融合网络与原始部分点云进行进一步精细化。借助二维基础先验,我们的方法消除了推理时对GT尺度信息的需求。此外,它在二维生成先验与三维点云补全之间架起了一座具有原则性的桥梁。大量实验证明了我们框架的优越性,使ScaleBlind成为点云补全任务上新的最先进方法。
cs.CV / 79 / 2609.23408

Accurate Motion Estimation with B\'ezier Control Point for Efficient Frame Interpolation

基于贝塞尔控制点的精确运动估计用于高效视频帧插值
Han, Shuhao, Wu, Chenyang, Guo, Chun-Le, Duan, Zheng-Peng, Li, Zhen, Cheng, Ming-Ming, Li, Chongyi
Abstract
In frame interpolation tasks, motion ambiguity in the training set causes models to generate blurry intermediate frames. Moreover, the assumption of uniform motion between frames during inference further leads to inaccuracies in the generated intermediate frames. To tackle these challenges, we propose an Accurate motion estimation algorithm with B\'ezier Control point, ABC-Inter, for efficient frame Interpolation. Specifically, ABC-Inter designs an Accurate Flow estimation Module (AFM) by decoupling two-frame features and mapping to corresponding coordinates to better estimate the optical flow between the two frames. Furthermore, ABC-Inter eliminates motion ambiguity in the training set by introducing B\'ezier control points that are computed using the input frames and the intermediate ground-truth (gt) frames. This allows the model to estimate accurate optical flow between two frames during the training process, thereby solving the blurriness problem in the generated intermediate frames during inference. Benefiting from the more accurate flow estimation between two frames, we can introduce additional frames and directly use multiple flows to calculate B\'ezier control points for modeling non-uniform motion without retraining the model. Simultaneously, to realize the estimation of non-linear motion using only two frames, we also introduce a new B\'ezier control point estimation module which achieves better motion estimation between the two frames by performing fine-tuning on the model in the second stage. Experimental results demonstrate that our ABC-Inter achieves state-of-the-art performance on multiple benchmark datasets and exhibits excellent visual perception.
Chinese Translation
在视频帧插值任务中,训练集中的运动歧义会导致模型生成模糊的中间帧。此外,推理过程中假设帧间运动均匀,进一步导致生成的中间帧不准确。为应对这些挑战,我们提出了一种基于贝塞尔控制点的精确运动估计算法 ABC-Inter,用于高效的帧插值。具体而言,ABC-Inter 通过解耦双帧特征并映射到相应坐标,设计了精确光流估计模块(AFM),以更好地估计两帧之间的光流。此外,ABC-Inter 通过引入利用输入帧和中间真值(gt)帧计算得到的贝塞尔控制点,消除了训练集中的运动歧义。这使得模型在训练过程中能够估计两帧之间精确的光流,从而解决了推理时生成的中间帧模糊问题。得益于更精确的双帧光流估计,我们可以在不重新训练模型的情况下引入额外的帧,并直接使用多条光流计算贝塞尔控制点来建模非均匀运动。同时,为了仅使用两帧实现非线性运动估计,我们还引入了一种新的贝塞尔控制点估计模块,通过在第二阶段对模型进行微调,实现两帧之间更好的运动估计。实验结果表明,我们的 ABC-Inter 在多个基准数据集上取得了最先进的性能,并展现出卓越的视觉感知效果。
cs.CV / 80 / 2609.23409

Retrieval Geometry Shapes Cache-Based Clip Adaptation

检索几何结构决定基于缓存的CLIP自适应效果
Tamim, Mahir Shahriar, Alim, Md. Samiul, Wasi, Azmine Toushik, Ridoy, Shahriyar Zaman, Nesa, Meharun, Yousuf, Mohammad Abu, Lamb, Alex, Moni, Mohammad Ali
Abstract
Cache-based test-time adaptation improves CLIP predictions by storing and retrieving examples from the target stream while keeping the model frozen. However, existing methods largely treat the feature space used for image-image retrieval as fixed, leaving open how much adaptation depends on the retrieval space itself. We study this question by fixing the memory and changing only the retrieval encoder, finding that the same memory can yield very different gains: across sixteen retrieval spaces, ImageNet-A cache gain ranges from at most +0.44 points for CLIP and MAE to +19.7 +/- 0.4 for DINOv2-L, while label-free retrieval-space selection retains 98% of oracle gain on ImageNet-V2. These results show that memory quality depends not only on which examples are stored, but also on how they are retrieved. Motivated by this finding, we propose MARC (Memory Augmented Retrieval for CLIP), a training-free system that uses frozen CLIP for prediction and DINOv2-B for retrieval with a single fusion weight. A single-view cache repairs 1074 +/- 21 baseline errors, compared with 878 +/- 4 for a 64-view ensemble, at roughly one seventh of the cost. Across four ImageNet distribution shifts, MARC reaches a 67.91% OOD average and, at matched DINOv2-B scale and eight views, achieves 64.17 +/- 0.31% versus 62.75 +/- 0.15% for a graph-based cache system while running 2.6 times faster. Overall, our results establish retrieval space as a first-order design choice for robust cache-based adaptation in remote sensing, scientific imaging, and changing visual environments.
Chinese Translation
基于缓存的测试时自适应通过在目标数据流中存储并检索样本,同时保持模型冻结,从而改进CLIP的预测。然而,现有方法大多将用于图像—图像检索的特征空间视为固定,这使得自适应在多大程度上依赖于检索空间本身这一问题仍未得到解答。我们通过固定记忆库、仅更换检索编码器来研究这一问题,发现相同的记忆库可以带来差异巨大的性能增益:在十六个检索空间中,ImageNet-A上的缓存增益从CLIP和MAE的最多+0.44个百分点到DINOv2-L的+19.7 +/- 0.4不等,而无标签的检索空间选择方法在ImageNet-V2上可保留98%的oracle增益。这些结果表明,记忆库的质量不仅取决于存储了哪些样本,还取决于这些样本是如何被检索的。基于这一发现,我们提出了MARC(Memory Augmented Retrieval for CLIP),一种无需训练的系统,它使用冻结的CLIP进行预测,并使用DINOv2-B进行检索,仅需一个融合权重。单视图缓存可修复1074 +/- 21个基线错误,而64视图集成仅修复878 +/- 4个错误,且成本约为后者的七分之一。在四个ImageNet分布偏移上,MARC达到67.91%的OOD平均准确率;在相同的DINOv2-B规模和八视图设置下,取得64.17 +/- 0.31%的成绩,优于基于图的缓存系统的62.75 +/- 0.15%,同时运行速度快2.6倍。总体而言,我们的结果确立了检索空间是鲁棒缓存自适应中的一项首要设计选择,适用于遥感、科学成像以及不断变化的视觉环境。
cs.CV / 81 / 2609.23425

Semi-automated reconstruction of indoor geometry from 360-degree video for CFD-based airflow analysis in classrooms

基于360度视频的室内几何半自动化重建用于教室CFD气流分析
Gamdha, Dhruv, Afful, James, Joshi, Shambhavi, Passe, Ulrike, Krishnamurthy, Adarsh, Ganapathysubramanian, Baskar
Abstract
Computational Fluid Dynamics (CFD) is widely used to evaluate ventilation and contaminant transport in occupied buildings, but deployment at scale is limited by three bottlenecks: acquiring room geometry without costly scanning hardware or manual CAD modeling, decomposing the scene into individually manipulable objects, and reconfiguring those objects for alternative layouts without re-capturing the room. We present a semi-automated workflow that converts a single 360-degree video of a room into individually editable, simulation-ready geometry assets. A dense point cloud is reconstructed using Neural Radiance Fields (NeRF), and 2D instance masks from text-prompted SAM 3 segmentation are lifted to 3D using multi-view consensus and depth-band filtering. Points are separated into object instances with an octree, and occlusion gaps are healed with a connectivity graph. Chair templates are fitted by Iterative Closest Point (ICP) alignment, and table geometry is generated procedurally. A browser-based editor supports quality assurance and rapid construction of alternative layout configurations. A steady Reynolds-averaged OpenFOAM solution then drives transient passive-scalar transport; the setup is verified using a mesh-sensitivity study and validated against an IEA Annex 20 benchmark. We apply the workflow to two university classrooms and a tiered lecture-hall auditorium. The capture-to-geometry pass takes two to five hours per room on a consumer workstation. In a controlled obstruction sequence in one classroom, the modeled half-clearance time varies non-monotonically as furniture is added, and a cross-room comparison indicates that clearance behavior cannot be reliably extrapolated between rooms, motivating per-room geometry acquisition. By making that acquisition low-cost, the workflow makes geometry-resolved comparative ventilation studies practical for spaces such as classrooms.
Chinese Translation
计算流体动力学(CFD)被广泛用于评估有人使用建筑中的通风和污染物传输,但其大规模部署受限于三个瓶颈:在不使用昂贵扫描硬件或人工CAD建模的情况下获取房间几何、将场景分解为可单独操作的对象,以及在不重新采集房间数据的情况下重新配置这些对象以生成替代布局。我们提出了一种半自动化工作流程,可将房间的单个360度视频转换为可单独编辑、可用于仿真的几何资产。使用神经辐射场重建稠密点云,并将来自文本提示SAM 3分割的2D实例掩码通过多视角一致性和深度带过滤提升至3D。利用八叉树将点分离为对象实例,并通过连通性图修复遮挡空洞。椅子模板通过迭代最近点(ICP)对齐进行拟合,桌子几何则按程序化方式生成。基于浏览器的编辑器支持质量保证以及快速构建替代布局配置。随后采用稳态雷诺平均OpenFOAM求解驱动瞬态被动标量传输;该设置通过网格敏感性研究进行验证,并依据IEA Annex 20基准进行校验。我们将该工作流程应用于两间大学教室和一个阶梯式报告厅。在消费级工作站上,每间房间的采集到几何生成过程需要两至五小时。在某一教室的受控障碍物序列实验中,随着家具的增加,模拟的半清除时间呈非单调变化;跨房间比较表明,清除行为无法在房间之间可靠地外推,因此有必要对每个房间单独获取几何数据。通过使该获取过程低成本化,该工作流程使得针对教室等空间的几何解析式对比通风研究切实可行。
cs.CV / 82 / 2609.23427

RSPDBench: Benchmarking Vision Foundation Models on Earth Observation Tasks Under Physically Grounded Remote-Sensing Product Degradations

RSPDBench:基于物理真实的遥感产品退化条件下的视觉基础模型地球观测任务基准测试
Faruk, Tanjim Bin, Reza, Khondaker Masfiq, Pallickara, Shrideep, Pallickara, Sangmi Lee
Abstract
Vision foundation models targeting Earth observation (EO) tasks are commonly evaluated on clean downstream benchmarks, but operational EO products can already contain spatial, radiometric, alignment, noise, and harmonization defects before reaching the model. Existing robustness evaluations often use generic image corruptions or broad domain shifts, which do not isolate these product-level failure modes. We introduce \textbf{RSPDBench}, a physically grounded \textbf{r}emote-\textbf{s}ensing-\textbf{p}roduct \textbf{d}egradation \textbf{b}enchmark for vision foundation models. RSPDBench evaluates five EO datasets, seven foundation-model entries, and two supervised baselines under audited primitive degradations and compound product chains. Each model is evaluated under its clean-selected native protocol, with robustness measured as the drop from its own clean baseline. Our analysis reveals that degradation sensitivity is strongly structured: resolution-conditioned and channel-grouped encoders protect different failure axes, and the same physical defect can hurt one model while helping another. Compound chains expose failures that isolated degradations do not predict, with model-dependent amplification, saturation, or component dominance, and excess drops up to $38$ percentage points beyond the strongest component. These results show that EO robustness cannot be characterized by clean accuracy or generic perturbation tests alone; it must also be measured against the structured defects that remote-sensing products carry into deployment.
Chinese Translation
面向地球观测(EO)任务的视觉基础模型通常在干净的下游基准上进行评估,但实际的EO产品在进入模型之前可能已经包含空间、辐射、配准、噪声以及归一化方面的缺陷。现有的鲁棒性评估往往采用通用图像扰动或宽泛的域偏移,无法隔离这些产品层面的失效模式。我们提出了RSPDBench,一个基于物理真实的遥感产品退化基准,用于评估视觉基础模型。RSPDBench在经过审核的原始退化与复合产品链条件下,评估了五个EO数据集、七个基础模型条目以及两个监督基线。每个模型均在其干净选定的原生协议下进行评估,鲁棒性以相对其自身干净基线的性能下降来度量。我们的分析表明,退化敏感性具有很强的结构化特征:分辨率条件化编码器与通道分组编码器保护的失效维度各不相同,且同一物理缺陷可能损害一个模型却有助于另一个模型。复合链条会暴露出孤立退化所无法预测的失效,并呈现模型相关的放大、饱和或分量主导效应,其额外性能下降最高可超出最强分量达38个百分点。这些结果表明,EO鲁棒性不能仅凭干净精度或通用扰动测试来刻画,还必须针对遥感产品在部署时带入的结构化缺陷进行度量。
cs.CV / 83 / 2609.23431

HOIBlender: Blending Lightweight Detection with Vision-Language Priors for Efficient Human-Object Interaction Detection

HOIBlender:融合轻量级检测与视觉-语言先验的高效人-物交互检测
Chen, Junwen, Yanai, Keiji
Abstract
Human-object interaction (HOI) detection requires grounding an interacting human-object pair and recognizing the verb that links them, often under severe long-tail supervision. Recent methods improve accuracy with stronger detectors and vision-language priors, but many still stack heavy transformer encoders, intricate denoising schedules, or post-hoc semantic calibration on top of the detector. We present \textbf{HOIBlender}, an efficient HOI detector named after its core design principle: blending detector-grounded visual tokens, spatial subject-object reasoning, and BLIP-2 semantic priors inside one lightweight decoding pipeline. HOIBlender builds on an RF-DETR/LW-DETR-style foundation with a DINOv2 backbone and selects top-$K$ image-conditioned tokens directly from the multi-scale projector as subject and object candidates, removing the dedicated encoder stage retained by prior HOI methods. A dual-stage decoder first stabilizes human-object geometry and then performs verb and HOI classification through progressive BLIP-2 prior fusion, with classifier weights initialized from BLIP-2 text embeddings for long-tail categories. Grouped-query training further enriches optimization without increasing inference cost. Across three model scales (Nano, Small, 2XL), HOIBlender consistently outperforms SOV-STG-VLA and Hybrid-SOV-VLA on HICO-DET, reaching $44.49$ Default Full mAP in only $9$ training epochs while maintaining competitive latency and parameter budgets. These results show that lightweight detection, structured spatial-semantic decoding, and deeply integrated vision-language priors can be blended into a single efficient HOI pipeline.
Chinese Translation
人-物交互(HOI)检测需要定位交互的人-物对并识别连接二者的动词,且通常面临严重的长尾监督问题。近期方法通过更强的检测器和视觉-语言先验提升了精度,但许多方法仍在检测器之上堆叠重型Transformer编码器、复杂的去噪调度或事后语义校准。我们提出了HOIBlender,一种高效的HOI检测器,其名称源于其核心设计原则:在一个轻量级解码流程中融合基于检测器的视觉词元、空间主客体推理以及BLIP-2语义先验。HOIBlender构建于RF-DETR/LW-DETR风格的基础架构之上,采用DINOv2骨干网络,直接从多尺度投影器中选取前K个图像条件化词元作为主语和宾语候选,从而去除了以往HOI方法保留的专用编码器阶段。双阶段解码器首先稳定人-物几何关系,随后通过渐进式BLIP-2先验融合完成动词与HOI分类,其中分类器权重由BLIP-2针对长尾类别的文本嵌入初始化。分组查询训练进一步丰富了优化过程而不增加推理成本。在三个模型规模(Nano、Small、2XL)上,HOIBlender在HICO-DET数据集上始终优于SOV-STG-VLA和Hybrid-SOV-VLA,仅用9个训练周期即达到44.49的Default Full mAP,同时保持了具有竞争力的延迟和参数预算。这些结果表明,轻量级检测、结构化的空间-语义解码以及深度集成的视觉-语言先验可以融合为单一高效的HOI处理流程。
cs.CV / 84 / 2609.23436

GAPS: Generative Active Pseudo-view Selection for Sparse-View 3D Gaussian Splatting

GAPS:面向稀疏视角3D高斯泼溅的生成式主动伪视角选择
Zhu, Hongfei, Deng, Haochen, Zhang, Sitao, Zhou, Ling
Abstract
Novel view synthesis from sparse observations is severely under-constrained. Although 3D Gaussian Splatting (3DGS) enables real-time rendering, it produces floaters, broken geometry, and washed-out backgrounds when trained with few views. We propose an alternating optimization framework that uses a pre-trained image diffusion model to generate geometrically consistent pseudo-views for additional 3DGS supervision. Generation is constrained by depth-conditioned ControlNet, IP-Adapter style transfer, LoRA scene adaptation, and img2img structural anchoring. We introduce Generative Active Pseudo-view Selection (GAPS) to balance reconstruction informativeness and generative reliability when choosing target views. Its annealing schedule shifts from conservative interpolation early in training to exploratory extrapolation later, gradually covering unobserved regions. A dual-criterion admission gate and uncertainty-weighted losses reject unreliable generations, while density-adaptive DropGaussian reduces overfitting in complex scenes. On LLFF with 3/6/9 views, our method improves average PSNR over vanilla 3DGS by 0.40/0.89/0.70 dB. On Mip-NeRF 360 with 12/24 views, the gains are 1.18/0.80 dB. SSIM improves and LPIPS decreases in every setting. Ablations show that active selection and density-adaptive regularization are both necessary; only the full method reduces LPIPS below the no-pseudo-view baseline on unbounded 360-degree scenes.
Chinese Translation
从稀疏观测进行新视角合成是一个严重欠约束的问题。尽管3D高斯泼溅(3D Gaussian Splatting, 3DGS)能够实现实时渲染,但在少视角训练时会产生漂浮物、破碎的几何结构以及褪色的背景。我们提出一种交替优化框架,利用预训练的图像扩散模型生成几何一致的伪视角,为3DGS提供额外监督。生成过程受到深度条件ControlNet、IP-Adapter风格迁移、LoRA场景自适应以及img2img结构锚定的约束。我们引入生成式主动伪视角选择(Generative Active Pseudo-view Selection, GAPS),在选择目标视角时平衡重建信息量与生成可靠性。其退火调度从训练早期的保守插值逐渐转向后期的探索性外推,逐步覆盖未观测区域。双准则准入门控和不确定性加权损失用于剔除不可靠的生成结果,而密度自适应的DropGaussian则降低复杂场景中的过拟合。在LLFF数据集的3/6/9视角设置下,我们的方法相比原始3DGS平均PSNR分别提升0.40/0.89/0.70 dB。在Mip-NeRF 360数据集的12/24视角设置下,提升分别为1.18/0.80 dB。所有设置中SSIM均有所提高,LPIPS均有所降低。消融实验表明,主动选择和密度自适应正则化都是必要的;只有在完整方法下,LPIPS才能在无界的360度场景中降至无伪视角基线之下。
cs.CV / 85 / 2609.23442

PhysReflect: Geometry and Perception Guided Diffusion for Physically-Plausible Mirror Reflections

PhysReflect:面向物理合理镜面反射的几何与感知引导扩散模型
Ge, Shuheng, Ren, Hongwei, Zhang, Li, Wu, Xiangqian
Abstract
Diffusion models generate high-quality images, yet often violate the physical laws governing mirror reflections. Reflections often suffer from geometric aberrations, including positional offsets, directional misalignment, proportional imbalance, and structural distortion. These failures remain evident even in contemporary state-of-the-art generative systems. Existing methods itigate this problem through synthetic data scaling or auxiliary depth conditioning, yet their merely reliance on latent-space noise reconstruction losses as implicit supervision prevents direct enforcement of reflection-specific geometric and perceptual constraints. To bridge this gap, we present PhysReflect, a geometry and perception guided diffusion framework that decodes the predicted clean latent into pixel space at each training step and applies annealed supervision through two complementary differentiable objectives. The Geometric Loss enforces mirror-induced spatial consistency through sparse epipolar correspondence and dense boundary projection alignment, where a SAM2-based TwinTrack mechanism provides stable in-mirror localization for boundary-aware supervision. The Perceptual Loss preserves reflected appearance by combining Semantic Consistency Loss, which maintains reflected identity and appearance via DINOv2 features, and Lighting Consistency Loss, which regularizes depth, surface-normal, and illumination coherence under monocular geometry priors. Experiments on synthetic and real-world benchmarks show that PhysReflect outperforms prior mirror-reflection methods in geometric, perceptual, and physical-plausibility metrics, as well as qualitative visual results.
Chinese Translation
扩散模型能够生成高质量图像,但常常违反支配镜面反射的物理定律。反射图像往往存在几何偏差,包括位置偏移、方向错位、比例失衡和结构变形。即使在这些当代最先进的生成系统中,这类失败依然明显。现有方法通过合成数据扩充或辅助深度条件来缓解该问题,但它们仅依赖隐空间噪声重建损失作为隐式监督,无法直接施加反射特有的几何与感知约束。为弥合这一差距,我们提出了PhysReflect,一个几何与感知引导的扩散框架,它在每个训练步中将预测的干净隐变量解码到像素空间,并通过两个互补的可微目标施加退火监督。几何损失通过稀疏对极对应和密集边界投影对齐来强制执行镜面引起的空间一致性,其中基于SAM2的TwinTrack机制为边界感知监督提供稳定的镜像内定位。感知损失通过结合语义一致性损失(利用DINOv2特征保持反射的身份与外观)和光照一致性损失(在单目几何先验下正则化深度、表面法向与光照的相干性)来保持反射外观。在合成与真实世界基准上的实验表明,PhysReflect在几何、感知和物理合理性指标以及定性视觉效果方面均优于以往的镜面反射方法。
cs.CV / 86 / 2609.23450

Rule-Constrained Assignment for Cue-Ball Identification in Broadcast Snooker

面向转播斯诺克比赛中母球识别的规则约束分配方法
Cao, Yuxin, Song, Wei, Wu, Yuezhong, Dong, Jin Song
Abstract
Accurate cue-ball identification is essential for metric analysis of broadcast snooker. Existing systems evaluate each candidate independently against a fixed white prototype and reject candidates above an appearance threshold. Under broadcast conditions, illumination changes can make colored balls appear white, while intrusions from players and equipment can obscure the cue ball or introduce competing candidates. We formulate cue-ball identification as a rule-constrained assignment problem that jointly assigns detected candidates to the bounded snooker inventory: one cue ball, up to 15 reds, and six colors with known spots. The cue ball is selected by the incremental cost of assigning each candidate to the white slot, and the conventional appearance test follows as the one-slot case. On 419 hand-annotated shots, our method improves identity accuracy from 88.1% to 95.5%, and from 80.5% to 95.2% on held-out venues. Within CueLift, our metric state-recovery system, the assignment expands coverage from 36.2% to 55.4% over 6,241 scorable shots. Assigning an estimate to every shot in a separate evaluation on 2,529 shots preserves this advantage.
Chinese Translation
准确的母球(cue-ball)识别对于斯诺克转播视频的量化分析至关重要。现有系统将每个候选目标独立地与固定的白色原型进行比较,并拒绝外观差异超过阈值的候选。然而在转播条件下,光照变化可能使彩色球呈现白色,而球员及器材的闯入可能遮挡母球或引入干扰候选。我们将母球识别形式化为一个规则约束的分配问题,将检测到的候选目标联合分配至有界的斯诺克球库存中:一颗母球、至多15颗红球以及六颗位于已知点位(spot)的彩球。母球通过计算将每个候选分配至白色球槽的增量代价来选择,而传统的外观测试则作为单球槽特例自然推导得出。在419个人工标注的击球样本上,我们的方法将识别准确率从88.1%提升至95.5%;在留出的未见过场地数据上,准确率从80.5%提升至95.2%。在我们的量化状态恢复系统CueLift中,该分配方法在6,241次可计分击球上将覆盖率从36.2%提升至55.4%。在另一项针对2,529次击球的评估中,为每次击球均分配估计结果的方法同样保持了这一优势。
cs.CV / 87 / 2609.23478

Algebraic Consistency Alone Does Not Certify Temporal Structure in Latent Action Models

仅凭代数一致性无法认证潜在动作模型中的时间结构
Wen, Di, Zhang, Ruodi, Yang, Kailun, Peng, Kunyu
Abstract
Latent action models infer a code for the transition between two frames of action-free video. Recent methods regularise this code to compose additively and reverse antisymmetrically, and report order-of-magnitude reductions in the resulting errors as a label-free certificate that the code has captured temporal structure. We show that this conclusion does not follow. Reconstruction drives the decoded transition toward a difference of state features, for which both identities hold for any pairing, a solution the metric cannot distinguish from one encoding nuisance state or a coordinate convention. Across five source domains, a trained but unconstrained counterpart already achieves 83-97% of the reduction relative to an untrained anchor. The residual fold is governed as much by the decoder family as by what is learned. A constrained model retrained after its temporal pairing is destroyed still reaches, in each domain, a lower error than the unconstrained model on real data. Downstream, preserving the temporal pairing yields no consistent advantage on LIBERO-GOAL or LIBERO-SPATIAL, and across the tested arms the code's mean linear action decodability falls as the algebraic error improves. We also test the most direct repair, a violation-contrastive objective that requires the algebra to fail on destroyed pairings: in the tested configurations it yields only a marginal separation within the reconstruction budget, on training and test triples alike. We recommend a validation protocol that these methods currently lack: a baseline-corrected metric, retraining on destroyed pairings, and a seed-budget analysis.
Chinese Translation
潜在动作模型(latent action models)旨在为无动作视频中两帧之间的转换推断一个编码。近期方法对该编码施加正则化约束,使其满足加法复合与反对称可逆的性质,并报告由此带来的数量级误差下降,将其作为编码已捕捉时间结构的无标签证明。我们证明这一结论并不成立。重构目标会促使解码出的转换趋向于状态特征之差的形式,而对该形式而言,上述两种恒等关系对任意帧配对均成立——这一解在上述度量下无法与编码了干扰状态或坐标约定的解区分开来。在五个源域上,一个经过训练但不受约束的对照模型相对于未经训练的锚点已实现了83–97%的误差下降。剩余的折叠误差受解码器家族的影响与所学内容的影响相当。一个在时间配对被破坏后重新训练的受约束模型,在每个域上于真实数据上仍能达到比不受约束模型更低的误差。在下游任务中,保留时间配对在LIBERO-GOAL或LIBERO-SPATIAL上并未带来一致的优势;并且在所有测试配置中,编码的平均线性动作可解码性反而随着代数误差的改善而下降。我们还测试了最直接的修复方案,即一种违反对比目标,要求代数关系在被破坏的配对上失效:在测试配置中,无论在训练三元组还是测试三元组上,它在重构预算内仅产生了微弱的区分度。我们建议引入此类方法目前所缺乏的验证协议:经基线校正的度量、在被破坏配对上的重训练,以及随机种子预算分析。
cs.CV / 88 / 2609.23492

CE$^4$L: Continual Ego, Exo, and Ego-Exo Learning

CE$^4$L:持续的第一人称、第三人称及双视角学习
Yan, Hongwei, Zhou, Kanglei, Liu, Yuchen, Shi, Qingyu, Zhong, Yi, Wang, Liyuan
Abstract
Perception for embodied agents is video-based, often multi-view (ego, exo, or both), and inherently continual, with simultaneous task and viewpoint shifts. Yet continual learning (CL) remains dominated by exo-only recognition tasks, obscuring behavior under these real-world coupled shifts. We introduce Continual Ego, E}xo, and Ego-Exo Learning (CE$^4$L), a unified multi-view CL benchmark spanning four representative tasks: cross-view referenced skill assessment, temporal action segmentation, cross-view association, and action anticipation & planning. CE$^4$L highlights challenges largely absent in prior CL benchmarks, including cross-view correspondence, view-dependent asynchrony, and heterogeneous semantic objectives. To this end, we propose Video Incremental Subspace-routed Task Adapters (VISTA), a parameter-efficient baseline method that stores task-specific updates in lightweight adapters and performs training-free routing via residual distance to task-specific whitened subspaces estimated from second-order statistics. Extensive experiments demonstrate the significantly varied efficacy of representative CL methods across CE$^4$L settings, while VISTA is consistently competitive and achieves state-of-the-art overall performance. Our source code for benchmarks and methods is available at https://github.com/AnAppleCore/CE4L .
Chinese Translation
面向具身智能体的感知是基于视频的,通常为多视角(第一人称视角 ego、第三人称视角 exo 或两者兼有),并且本质上是持续性的,同时存在任务偏移与视角偏移。然而,持续学习(Continual Learning, CL)目前仍以仅限第三人称视角的识别任务为主,掩盖了这些真实世界中耦合偏移下的行为表现。我们提出了持续第一人称、第三人称及双视角学习基准(Continual Ego, Exo, and Ego-Exo Learning, CE$^4$L),这是一个统一的多视角持续学习基准,涵盖四个代表性任务:跨视角参照的技能评估、时序动作分割、跨视角关联,以及动作预测与规划。CE$^4$L 突出了以往持续学习基准中基本缺失的挑战,包括跨视角对应关系、视角相关的异步性以及异构的语义目标。为此,我们提出了视频增量子空间路由任务适配器(Video Incremental Subspace-routed Task Adapters, VISTA),这是一种参数高效的基线方法,将任务特定的更新存储在轻量级适配器中,并通过基于二阶统计量估计的任务特定白化子空间的残差距离,实现无需训练的路由。大量实验表明,各类代表性持续学习方法在 CE$^4$L 不同设置下的效果差异显著,而 VISTA 始终具有竞争力,并取得了整体最优的最新性能。我们的基准与方法的源代码可在 https://github.com/AnAppleCore/CE4L 获取。
cs.CV / 89 / 2609.23495

Pay More Attention To Text In High-Resolution MLLMs

在高分辨率多模态大语言模型中更加关注文本
Mao, Zhongkuan, Zhao, Wenzhuo, Liu, Xianjie, Wang, Yidong, Gao, Zhao, Xian, Ronghao, Jiang, Yao, Zhang, Yi, Wen, Liangjian, Fu, Keren
Abstract
Failures of high-resolution MLLMs are commonly attributed to a visual problem, motivating zooming, cropping, and related visual interventions to recover fine-grained evidence or suppress interference. Yet recent studies suggest that relevant visual evidence is already encoded in intermediate representations, indicating that visual-side improvements alone insufficient. This raises a natural question: does the remaining bottleneck lie in the text that guides visual search? We identify a previously overlooked linguistic bottleneck: questions formulated for answering do not necessarily specify the visual evidence required for localization. To address this mismatch, we introduce EviSpec, a training-free compiler that derives complementary evidence specifications while preserving the original question for final reasoning. We further validate it through matched-control experiments that isolate the roles of evidence specification and localization. With the search budget fixed, structured evidence specifications yield an 8.6% relative gain over generic requests. With evidence geometry matched, the evidence localized by EviSpec yields a 14.8% relative gain over random evidence. Together, these controls isolate the benefit of specifying what evidence to seek rather than merely expanding visual access. Across all five MLLMs, EviSpec consistently improves upon the corresponding baseline on each of the three benchmarks, yielding average relative gains of \textbf{10.4%, 8.8%, and 12.4%} on V\textsuperscript{*}Bench, HR-Bench-4K, and HR-Bench-8K, respectively. Beyond high-resolution reasoning, EviSpec also achieves state-of-the-art performance on VQA and hallucination-focused benchmarks.
Chinese Translation
高分辨率多模态大语言模型(MLLM)的失败通常被归因于视觉问题,这促使研究者采用缩放、裁剪及相关视觉干预手段来恢复细粒度证据或抑制干扰。然而,近期研究表明,相关视觉证据其实已经被编码在中间表示中,这表明仅依靠视觉侧的改进是不够的。由此引出一个自然的问题:剩余的瓶颈是否在于引导视觉搜索的文本?我们识别出一个此前被忽视的语言学瓶颈:为回答问题而构造的问题表述未必指明了定位所需的视觉证据。为解决这一失配问题,我们提出了 EviSpec,一个无需训练的编译器,它在保留原始问题用于最终推理的同时,推导出互补的证据规格说明。我们进一步通过匹配对照实验验证该方法,以分离证据规格说明与定位各自的作用。在搜索预算固定的情况下,结构化的证据规格说明相比一般性请求带来 8.6% 的相对提升;在证据几何分布匹配的情况下,由 EviSpec 定位的证据相比随机证据带来 14.8% 的相对提升。这些对照实验共同分离出“指明所需证据”本身的优势,而不仅仅是扩大视觉访问范围。在全部五个 MLLM 上,EviSpec 在三个基准测试中均一致优于相应基线,在 V^Bench、HR-Bench-4K 和 HR-Bench-8K 上分别取得平均 **10.4%、8.8% 和 12.4%** 的相对提升。除高分辨率推理外,EviSpec 在视觉问答(VQA)和以幻觉为重点的基准测试上也取得了最先进的性能。
cs.CV / 90 / 2609.23507

Detecting Phone-Induced Pedestrian Distraction via a Multimodal Fusion Transformer

基于多模态融合Transformer的手机诱发行人分心行为检测
Li, Yuanzhe, Liu, Hounian, Chang, Xiaotong, Huang, Yidi
Abstract
The increasing reliance on mobile phones has made phone-induced pedestrian distraction increasingly prevalent. Activities such as texting, watching videos, and making phone calls have become significant contributors to traffic accidents. Reliable detection of pedestrian distraction is essential for autonomous vehicles, as it improves situational awareness and enables timely risk assessment, thereby supporting safe motion planning and vehicle control. We propose a multimodal fusion Transformer (MFT) for detecting phone-induced pedestrian distraction. MFT jointly extracts skeletal dynamics from body pose keypoints and visual appearance features from pedestrian images, effectively leveraging the complementary information provided by the two modalities. A cross-modal attention module is proposed to capture inter-modal dependencies through multi-head cross-attention, facilitating effective fusion of complementary information across the two modalities. Then, a temporal attention fusion module, implemented with a Transformer encoder, is employed to capture temporal dependencies. MFT is trained and evaluated on a manually annotated dataset comprising 287 pedestrian instances with 20,741 images. Extensive experiments demonstrate that MFT attains an overall accuracy of 95%, exceeding the performance of six baseline approaches by 6%.
Chinese Translation
人们对手机的依赖日益增加,使得手机诱发的行人分心行为愈发普遍。发短信、看视频和打电话等活动已成为交通事故的重要诱因。可靠的行人分心检测对自动驾驶汽车至关重要,因为它可以提升态势感知能力并支持及时的风险评估,从而助力安全的运动规划与车辆控制。本文提出一种用于检测手机诱发行人分心的多模态融合Transformer(Multimodal Fusion Transformer,MFT)。MFT同时从人体姿态关键点中提取骨骼运动特征,并从行人图像中提取视觉外观特征,有效利用两种模态提供的互补信息。我们提出一个跨模态注意力模块,通过多头交叉注意力捕获模态间依赖关系,促进两种模态互补信息的有效融合。随后,采用基于Transformer编码器实现的时间注意力融合模块来捕获时间依赖关系。MFT在一个人工标注的数据集上进行训练与评估,该数据集包含287个行人实例和20,741张图像。大量实验表明,MFT取得了95%的总体准确率,较六种基线方法高出6%。
cs.CV / 91 / 2609.23509

GARO: Geometry-Aware Redundancy Optimization for Real-Time and High-Fidelity Dynamic Gaussian Splatting

GARO:面向实时高保真动态高斯泼溅的几何感知冗余优化
Xue, Huiwen, Zhao, Kaixing, Ming, Zuheng, Li, Tingcheng
Abstract
Novel view synthesis is a key task for dynamic scene reconstruction, where high rendering speed is essential for applications such as virtual reality. Existing deformable Gaussian Splatting methods achieve high-fidelity dynamic scene modeling, but still face limitations in memory usage and rendering efficiency due to the large number of redundant Gaussians. To address these challenges, we propose Geometry-Aware Redundancy Optimization (GARO), a unified redundancy measurement framework in the adaptive density control stage of the traditional dynamic scene reconstruction pipeline. This framework first selects low-gradient candidates using an optimization activity assessment strategy, and then evaluates geometric complexity through low curvature analysis to further filter and prune redundant points, resulting in a compact and expressive Gaussian representation. Extensive experiments on synthetic and real-world datasets demonstrate that GARO achieves robust trade-offs between quality and speed, with PSNR remaining stable and rendering speed improved by 2x, validating the efficiency and effectiveness of GARO.
Chinese Translation
新视角合成是动态场景重建中的一项关键任务,其中高渲染速度对于虚拟现实等应用至关重要。现有的可变形高斯泼溅(Gaussian Splatting)方法能够实现高保真的动态场景建模,但由于存在大量冗余高斯点,在内存占用和渲染效率方面仍面临局限。为解决这些挑战,我们提出几何感知冗余优化(Geometry-Aware Redundancy Optimization, GARO),这是一个集成于传统动态场景重建流程中自适应密度控制阶段的统一冗余度量框架。该框架首先利用优化活跃度评估策略筛选低梯度候选点,然后通过低曲率分析评估几何复杂度,进一步过滤并剪除冗余点,从而得到紧凑且表达能力强大的高斯表示。在合成数据集和真实数据集上的大量实验表明,GARO 在质量与速度之间实现了稳健的权衡:PSNR 保持稳定的同时渲染速度提升 2 倍,验证了 GARO 的高效性与有效性。
cs.CV / 92 / 2609.23533

GeoBalance: Geometry-Aware Monitoring and Reconstruction with Asymmetric Optimization for Balanced Multimodal Learning

GeoBalance:面向平衡多模态学习的几何感知监测与重构的非对称优化方法
Xiong, Zechang, Li, Da, Yin, Rong, Tang, Kexin, Yang, Biao, Li, Pengyuan, Kong, Wenkang, Hu, Yulan, Zhu, Shengyu, Peng, Hao
Abstract
Multimodal classifiers can converge to modality-dominant solutions in which one modality dominates the joint prediction, suppressing the learning of others. Existing balancing methods mainly adjust losses, gradients, or modality contributions, largely treating modality imbalance as an optimization problem while implicitly treating the weak modality as under-optimized but representationally intact. In this work, we find that this assumption does not always hold, as persistent modality dominance can induce a representation-level collapse of the weak modality, which we term \emph{manifold modality collapse} (MMC). MMC manifests as a coupled geometric degradation in which weak-modality representations collapse onto fewer directions within each class and become less separable across classes. Motivated by this observation, we propose \emph{GeoBalance}, a geometry-aware framework that monitors these two geometric properties and reconstructs the weak modality representation only when it exhibits signs of MMC. Once triggered, GeoBalance uses a fixed Simplex-ETF class scaffold and spectral regularization to restore class separation while preventing collapse onto a few feature directions. To preserve reconstruction during joint training, asymmetric gradient projection removes the joint-gradient component conflicting with reconstruction, leaving non-conflicting optimization unchanged. Extensive experiments across six multimodal benchmarks demonstrate great improvements over competitive balancing methods, validating its effectiveness.
Chinese Translation
多模态分类器可能收敛到模态主导的解,即某一模态主导联合预测,从而抑制其他模态的学习。现有的平衡方法主要调整损失、梯度或模态贡献,大多将模态不平衡视为优化问题,同时隐含地将弱势模态视为仅优化不足但表征结构完好的模态。在本工作中,我们发现这一假设并不总是成立,因为持续的模态主导会导致弱势模态在表征层面发生坍缩,我们将其称为流形模态坍缩(manifold modality collapse,MMC)。MMC表现为一种耦合的几何退化:弱势模态的表示在各类内部坍缩到更少的方向上,并且类间的可分性下降。基于这一观察,我们提出GeoBalance,一种几何感知框架,它监测这两项几何性质,并且仅在弱势模态表现出MMC迹象时才对其表示进行重构。一旦触发,GeoBalance使用固定的单纯形等角紧框架(Simplex-ETF)类骨架和谱正则化来恢复类间可分性,同时防止坍缩到少数特征方向上。为了在联合训练中保护重构过程,非对称梯度投影会移除与重构冲突的联合梯度分量,而保持非冲突的优化不变。在六个多模态基准上的大量实验表明,相较于具有竞争力的平衡方法,我们的方法取得了显著提升,验证了其有效性。
cs.CV / 93 / 2609.23534

PosEviLoc: Position-Conditioned Spatial Evidence for Language-Based 3D Localization

PosEviLoc:面向基于语言的3D定位的位置条件空间证据
Shang, Tianyi, Shi, Yike, Li, Zhenyu
Abstract
Language-based 3D localization retrieves the point-cloud submap containing a target position from descriptions of nearby objects and their spatial relations. Existing methods typically compress queries and submaps into global descriptors, potentially obscuring object-level semantics and cross-description spatial coherence. We propose Position-Conditioned Evidence Localization (PosEviLoc), a query-position-aware framework for coarse text-to-point-cloud localization. Instead of relying on global matching, PosEviLoc evaluates each candidate submap using explicit semantic and spatial evidence. It models direction as a relation jointly determined by an object position and a hypothetical query position. The resulting Query-Position Spatial Evidence Field (QSEF) measures the fraction of query descriptions supported at each hypothetical position, explicitly capturing their agreement without using the ground-truth query pose to construct the evidence field. A Multi-Level Evidence Readout (MER) summarizes this evidence in a compact representation, which a lightweight MLP converts into a retrieval score. Across five benchmarks, PosEviLoc outperforms MNCL by an average of 17 percentage points in Recall@1. When used as a plug-and-play reranker, it improves MNCL by an average of 16 percentage points. Moreover, PosEviLoc introduces substantially fewer parameters and achieves faster inference speed than existing methods.
Chinese Translation
基于语言的3D定位旨在根据对目标位置附近物体及其空间关系的描述,检索出包含该目标位置的点云子图。现有方法通常将查询和子图压缩为全局描述符,这可能掩盖物体级语义以及跨描述的空间一致性。我们提出了位置条件证据定位(Position-Conditioned Evidence Localization,PosEviLoc),这是一种面向粗粒度文本到点云定位的查询位置感知框架。PosEviLoc 不依赖全局匹配,而是利用显式的语义和空间证据来评估每个候选子图。它将方向建模为由物体位置和假设查询位置共同确定的关系。由此得到的查询位置空间证据场(Query-Position Spatial Evidence Field,QSEF)衡量每个假设位置处所支持的查询描述的比例,从而显式地捕捉查询描述之间的一致性,且无需使用真值查询位姿来构建证据场。多层级证据读取模块(Multi-Level Evidence Readout,MER)将这些证据总结为紧凑的表示,再由轻量级 MLP 将其转换为检索得分。在五个基准测试中,PosEviLoc 的 Recall@1 平均超出 MNCL 17 个百分点。作为即插即用的重排序器时,它平均可将 MNCL 提升 16 个百分点。此外,与现有方法相比,PosEviLoc 引入的参数量显著更少,且推理速度更快。
cs.CV / 94 / 2609.23541

Towards robust multimodal 3D object detection via visual foundation models

基于视觉基础模型的鲁棒多模态3D目标检测研究
Song, Ziying, Liu, Lin, Pan, Hongyu, Xu, Shaoqing, Yang, Lei, Guo, Mingzhe, Jia, Caiyan
Abstract
Multimodal 3D object detection is fundamental to robust perception in autonomous driving because it integrates complementary information from LiDAR and camera sensors. However, existing methods often fail to maintain robustness under out-of-distribution (OOD) corruptions caused by sensor noise, adverse weather, and environmental changes. To address this problem, we propose RoboDistill, a robust and generalizable multimodal 3D object detection framework that leverages visual foundation models (VFMs), such as the Segment Anything Model (SAM). First, we introduce SAM-AD, a domain-specific pretraining strategy that fine-tunes SAM on autonomous-driving imagery to extract feature representations with rich semantic information. Second, we design the AD Feature Pyramid Network (AD-FPN) to refine and upsample SAM features at multiple scales for seamless fusion with LiDAR features. Third, we develop the Depth-Guided Wavelet Attention (DGWA) module, which suppresses high-frequency sensor noise while preserving critical contextual information. Finally, we introduce KD Fusion, in which the pretrained SAM-AD serves as a teacher that distills high-quality visual knowledge into a lightweight point-cloud network, thereby improving robustness under noisy conditions. Extensive experiments across 27 challenging OOD corruption settings show that RoboDistill generally delivers stronger or competitive detection performance and robustness relative to representative state-of-the-art methods. This work bridges the gap between VFMs and 3D object detection and advances robust multimodal perception for real-world autonomous-driving applications.
Chinese Translation
多模态3D目标检测通过融合激光雷达(LiDAR)与相机传感器的互补信息,是实现自动驾驶鲁棒感知的基础。然而,现有方法在传感器噪声、恶劣天气和环境变化导致的分布外(OOD)损坏情况下,往往难以保持鲁棒性。为解决这一问题,我们提出了RoboDistill,一个利用视觉基础模型(VFM,如Segment Anything Model(SAM))的鲁棒且可泛化的多模态3D目标检测框架。首先,我们提出SAM-AD,一种领域特定的预训练策略,通过在自动驾驶图像上微调SAM,提取富含语义信息的特征表示。其次,我们设计了AD特征金字塔网络(AD-FPN),对SAM特征进行多尺度细化与上采样,以实现与激光雷达特征的无缝融合。第三,我们开发了深度引导小波注意力(DGWA)模块,在抑制高频传感器噪声的同时保留关键上下文信息。最后,我们引入KD Fusion,其中预训练的SAM-AD作为教师模型,将高质量的视觉知识蒸馏到轻量级点云网络中,从而提升噪声条件下的鲁棒性。在27种具有挑战性的OOD损坏设置上的大量实验表明,RoboDistill相较于代表性的最先进方法,通常能提供更强或具有竞争力的检测性能与鲁棒性。这项工作弥合了视觉基础模型与3D目标检测之间的差距,推动了面向真实世界自动驾驶应用的鲁棒多模态感知的发展。
cs.CV / 95 / 2609.23548

SewFusion: Tailored Generation of Topology and Panel-Level Geometry for Sewing Patterns

SewFusion:针对缝纫样板的定制化拓扑与面板级几何生成
Lin, Jiaxin, Pan, Xiao, Yuan, Hangjie, Liang, Luyan, Li, Wan, Feng, Daquan
Abstract
Generating sewing patterns from images and text requires modeling a heterogeneous representation composed of discrete topology and continuous geometry. Existing methods mainly follow two paradigms: diffusion-based methods enable holistic geometry generation by converting the entire pattern into a continuous representation, but weaken discrete topology modeling; in contrast, autoregressive methods preserve discrete topology through next-token prediction, but tie continuous geometry regression to token-level hidden states with limited panel-level context. To bridge this gap, we propose SewFusion, a unified autoregressive framework that adopts tailored generation mechanisms for discrete topology and panel-level continuous geometry, using next-token prediction for the former and flow matching for the latter. To support panel-level continuous geometry generation, we introduce a Panel Geometry VAE that learns a fixed-size latent space for variable-length panel geometry, together with Panel Geometry Flow for latent generation. We further propose Panel-Forcing to reduce the training--inference mismatch in topology context and improve robustness to topology prediction errors. Extensive experiments on SewFactory and GCD-MM demonstrate that SewFusion consistently outperforms previous state-of-the-art methods across various settings, achieving +6.36% Panel Accuracy, +11.30% Stitch Accuracy, and -1.90 Vertex L2 error in the image-text-based generation setting.
Chinese Translation
从图像和文本生成缝纫样板需要对由离散拓扑和连续几何组成的异构表示进行建模。现有方法主要遵循两种范式:基于扩散的方法通过将整个样板转换为连续表示来实现整体几何生成,但削弱了离散拓扑建模;相比之下,自回归方法通过下一词元预测保留了离散拓扑,但将连续几何回归与词元级隐藏状态绑定,且面板级上下文有限。为弥合这一差距,我们提出SewFusion,一个统一的自回归框架,为离散拓扑和面板级连续几何采用定制化的生成机制:对前者使用下一词元预测,对后者使用流匹配(flow matching)。为支持面板级连续几何生成,我们引入了面板几何VAE(Panel Geometry VAE),为可变长度的面板几何学习一个固定大小的潜在空间,并配合用于潜在生成的面板几何流(Panel Geometry Flow)。我们进一步提出面板强制(Panel-Forcing)方法,以减少拓扑上下文中的训练-推理不匹配,并提高对拓扑预测错误的鲁棒性。在SewFactory和GCD-MM数据集上的大量实验表明,SewFusion在各种设置下均持续优于先前的最先进方法,在基于图像-文本的生成设置中实现了+6.36%的面板准确率(Panel Accuracy)、+11.30%的缝合准确率(Stitch Accuracy)以及-1.90的顶点L2误差。
cs.CV / 96 / 2609.23553

Modeling Clinical Workflow for SYNTAX Scoring from Coronary Angiography Videos

基于冠状动脉造影视频的SYNTAX评分临床工作流建模
Fu, Suzhong, Dong, Jingqi, Ding, Xuan, Sun, Rui, Yang, Yiming, Cui, Shuguang, Li, Zhen
Abstract
The SYNTAX score is a clinically established tool for assessing anatomical lesion complexity in coronary artery disease and guiding subsequent treatment. However, automated SYNTAX scoring is commonly formulated as a direct regression problem from coronary angiography videos to patient-level scores. In this work, we reformulate SYNTAX scoring as a vessel segment identity-preserving anatomical reasoning problem and propose a hierarchical modeling framework that explicitly aligns learning with the clinical workflow. Our approach maintains vessel segment identity across frames and views, estimates stenosis severity at the segment level, and aggregates evidence hierarchically according to coronary anatomy. Simultaneously, to address the scarcity of domain-specific data, we integrate and complete multiple public coronary angiography datasets, constructing a large-scale resource featuring completed vessel segmentation and derived structural annotations. Experiments demonstrate that vessel segment-level stenosis embedding enhances explanatory power and reduces prediction variability compared to baseline models, with the R^2 score improving by 0.201 and dev STD decreasing by 18.4%. These results highlight the necessity of structure-aligned modeling for reliable and stable automated SYNTAX scoring from multi-view coronary angiography videos. The GitHub link is https://github.com/VersaceSu7/SYNTAX_score_777.
Chinese Translation
SYNTAX评分是临床公认的工具,用于评估冠状动脉疾病的解剖学病变复杂程度并指导后续治疗。然而,自动化SYNTAX评分通常被表述为从冠状动脉造影视频到患者级别评分的直接回归问题。在本工作中,我们将SYNTAX评分重新表述为一个血管段身份保持的解剖学推理问题,并提出了一种分层建模框架,使学习过程与临床工作流程显式对齐。我们的方法在帧间和不同视角间保持血管段的身份一致性,在血管段级别估计狭窄严重程度,并根据冠状动脉解剖结构分层聚合证据。同时,为解决领域特定数据的稀缺问题,我们整合并完善了多个公开的冠状动脉造影数据集,构建了一个包含完整血管分割标注和衍生结构化注释的大规模资源。实验表明,与基线模型相比,血管段级别的狭窄嵌入增强了模型解释能力并降低了预测变异性,R²分数提升了0.201,预测标准差降低了18.4%。这些结果凸显了结构对齐建模对于从多视角冠状动脉造影视频实现可靠且稳定的自动化SYNTAX评分的必要性。GitHub链接为 https://github.com/VersaceSu7/SYNTAX_score_777。
cs.CV / 97 / 2609.23561

Transferring Visual Explanations: How Cross-Architecture Knowledge Distillation Affects Model Interpretability

视觉解释的可迁移性:跨架构知识蒸馏如何影响模型可解释性
Czufarow, Aleks, Babin, Ihor
Abstract
Deploying efficient neural networks is essential in resource-constrained environments, yet compact models often sacrifice interpretability - a critical in safety-critical domains such as autonomous driving and medicine. This study investigates whether Knowledge Distillation transfers the spatial feature attribution of a large teacher network to a compact student. To assess the influence of the KD scheme on interpretability, we distill a ResNet-152 teacher into a ResNet-34 student on ImageNet-1K across five configurations by systematically varying the distillation temperature and soft-label loss weight. Models are evaluated on top-1 accuracy, along with two interpretability metrics: Relevance Mass Accuracy and Relevance Rank Accuracy. These metrics are computed via Grad-CAM heatmaps benchmarked against ground-truth object masks. Our results show that top-1 accuracy ranges from 71.6% to 74.0%. For Grad-CAM, RMA ranges from 7.7% to 9.7% and RRA from 7.3% to 10.1%; for Guided Grad-CAM, RMA ranges from 16.1% to 18.6% and RRA from 15.9% to 21.5%. Interpretability proves far more sensitive to the soft-label weight than to the temperature: keeping the student anchored to hard labels preserves both accuracy and coarse localization, whereas weighting the teacher heavily degrades both. Fine-grained attribution, however, fell below the undistilled baseline in every configuration tested, indicating that logit distillation transmits where a model attends more readily than the pixel-level structure of that attention. We evaluate 12 cross-architecture combinations of convolutional and transformer-based models, revealing that the inheritance of fine-grained spatial reasoning is fundamentally bottlenecked by the student's intrinsic structural biases. To our knowledge, this is the first application of this interpretability-aware evaluation framework - previously used for neural network pruning - to KD.
Chinese Translation
在资源受限的环境中部署高效的神经网络至关重要,然而紧凑模型往往牺牲了可解释性——这在自动驾驶和医疗等安全关键领域是极为关键的问题。本研究探究知识蒸馏(Knowledge Distillation)是否能将大型教师网络的空间特征归因迁移给紧凑的学生网络。为评估知识蒸馏方案对可解释性的影响,我们在ImageNet-1K数据集上,通过系统地改变蒸馏温度和软标签损失权重,以五种配置将ResNet-152教师模型蒸馏到ResNet-34学生模型中。模型以top-1准确率以及两个可解释性指标进行评估:相关性质量准确率(Relevance Mass Accuracy)和相关性排序准确率(Relevance Rank Accuracy)。这些指标通过Grad-CAM热力图相对于真实目标掩码进行计算。结果显示,top-1准确率介于71.6%至74.0%之间。对于Grad-CAM,RMA介于7.7%至9.7%,RRA介于7.3%至10.1%;对于Guided Grad-CAM,RMA介于16.1%至18.6%,RRA介于15.9%至21.5%。可解释性对软标签权重的敏感程度远高于对温度的敏感程度:让学生模型锚定于硬标签可以同时保持准确率和粗粒度定位能力,而过度偏向教师模型则会同时降低两者。然而,在所有测试配置中,细粒度归因均低于未蒸馏的基线,这表明logit蒸馏更易于传递模型关注的位置,而非该关注的像素级结构。我们评估了12种卷积模型与基于Transformer模型的跨架构组合,发现细粒度空间推理能力的继承从根本上受限于学生模型固有的结构偏置。据我们所知,这是该可解释性感知评估框架——此前用于神经网络剪枝——在知识蒸馏中的首次应用。
cs.CV / 98 / 2609.23565

MaskVLA: Visual Masking Against Trajectory Overfitting of Vision-Language-Action Model

MaskVLA:利用视觉掩码对抗视觉-语言-动作模型的轨迹过拟合
Jiang, Yuxuan, Huang, Jiaying, Wang, Ge, Yan, Shenhao, Yang, Jiahao, Yao, Chengsi, Liu, Qi, Zhao, Qing, Cui, Shuguang, Zhao, Yiming, Han, Yatong, Li, Zhen
Abstract
Vision-Language-Action (VLA) models integrate vision-language understanding with executable robot actions, enabling end-to-end learning for robot control. However, our empirical analysis reveals that existing models exhibit severe trajectory overfitting when finetuned on limited datasets. To guide the model in effectively utilizing wrist camera information, we propose MaskVLA, a masking-based fine-tuning strategy. By randomly masking a small portion of the main camera's visual information, the model is guided to autonomously learn more fine-grained, task-relevant, and effective visual features. This process leads to the emergence of robust policies, thereby enhancing the model's capability to tackle complex manipulation tasks and improving its generalization performance. Our method has been comprehensively evaluated on RoboTwin 2.0, achieving an average success rate improvement of 23.2% and 16.8% compared to $\pi_0$ and OpenVLA-OFT, respectively. Furthermore, experiments on real-world ALOHA robots also demonstrate the effectiveness of our approach.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型将视觉-语言理解与可执行的机器人动作相结合,实现了机器人控制的端到端学习。然而,我们的实证分析表明,现有模型在有限数据集上进行微调时会表现出严重的轨迹过拟合现象。为引导模型有效利用腕部相机信息,我们提出了 MaskVLA,一种基于掩码的微调策略。通过随机遮蔽主相机视觉信息的一小部分,引导模型自主学习更细粒度、与任务相关且有效的视觉特征。这一过程促使模型涌现出鲁棒的策略,从而增强其应对复杂操作任务的能力,并提升其泛化性能。我们在 RoboTwin 2.0 上对本方法进行了全面评估,相比 $\pi_0$ 和 OpenVLA-OFT,平均成功率分别提升了 23.2% 和 16.8%。此外,在真实 ALOHA 机器人上的实验也验证了该方法的有效性。
cs.CV / 99 / 2609.23566

G6D: Geometric Learning-Free RGB-D 6D Pose Solver for Robotic Manipulation

G6D:面向机器人操作的无几何学习RGB-D 6D位姿求解器
Liang, Yixuan, Chen, William, Wang, Yunan, Yan, Jizhou, Jin, Zhao, Liu, Changling, Hu, Chuxiong
Abstract
6D object pose estimation is fundamental to robotic manipulation and automation. Recent zero-shot methods have significantly improved generalization to unseen objects, but most still rely on large-scale pretrained models with substantial GPU computation and memory demands. These requirements complicate deployment on robotic platforms where perception, planning, and control share limited computational resources, while learned intermediate representations offer limited geometric interpretability for task-specific adaptation. To address these limitations, we propose G6D, a learning-free, geometry-driven RGB-D 6D pose solver. Given an RGB-D observation, an object instance mask, camera intrinsics, and a CAD model, G6D generates pose hypotheses through template-based geometric matching and refines them using silhouette and depth consistency, forming a purely geometry-driven pose estimation paradigm. This paradigm requires neither pretrained visual models nor target-specific training and preserves interpretable geometric representations throughout pose estimation. Moreover, adjustable hypothesis counts provide flexible accuracy-computation trade-offs, while a CPU-only configuration supports deployment without GPU resources. Experiments on LineMOD and five BOP19 datasets demonstrate advanced performance. Real-world pick-and-place experiments further demonstrate G6D's applicability to robotic manipulation. The complete project is publicly available at https://ai4control.github.io/G6D-Project-Page .
Chinese Translation
6D物体位姿估计是机器人操作与自动化的基础。近期的零样本方法在泛化到未见物体方面取得了显著进展,但大多数方法仍依赖大规模预训练模型,需要大量的GPU计算和内存资源。这些需求使得在感知、规划与控制共享有限计算资源的机器人平台上部署变得困难,同时学习到的中间表示对任务特定适配的几何可解释性也有限。为解决这些局限,我们提出了G6D,一种无需学习、由几何驱动的RGB-D 6D位姿求解器。给定RGB-D观测、物体实例掩码、相机内参和CAD模型,G6D通过基于模板的几何匹配生成位姿假设,并利用轮廓与深度一致性对其进行优化,形成了一种完全由几何驱动的位姿估计范式。该范式既不需要预训练视觉模型,也不需要针对目标物体的训练,并在整个位姿估计过程中保留了可解释的几何表示。此外,可调节的假设数量提供了灵活的精度-计算权衡,而纯CPU配置支持在无GPU资源的情况下部署。在LineMOD和五个BOP19数据集上的实验展示了先进的性能。真实世界的抓取与放置实验进一步证明了G6D在机器人操作中的适用性。完整项目已在 https://ai4control.github.io/G6D-Project-Page 公开发布。
cs.CV / 100 / 2609.23586

An Efficient and Effective Watermarking Scheme for the Protection of the Intellectual Property Rights of Video Generative Models

一种用于保护视频生成模型知识产权的高效且有效的水印方案
Huang, Wenhong, Fei, Jianwei, Tondi, Benedetta, Ma, Bin, Huang, Fangjun
Abstract
The rapid development of video generative models (VGMs) has enabled the generation of highly realistic synthetic videos, raising concerns about the intellectual property rights (IPR) of these models. In particular, two closely related forensic tasks remain largely unaddressed: synthetic video verification (determining whether a video was generated by a protected VGM) and model ownership verification (determining whether a suspect VGM is an unauthorized copy of a protected VGM). In this paper, we propose a new in-generation watermarking scheme that can address the two verification tasks. First, a novel video watermarking network named VidMark is presented, which incorporates a two-scale discrete wavelet transform (DWT) decomposition and a global temporal attention block (GTAB) to enhance watermark robustness and imperceptibility. Second, we present a decoder-guided fine-tuning procedure. By leveraging the frozen VidMark decoder, this process enables VGMs to synthesize videos carrying an imperceptible, robust, and model-specific watermark. Finally, two verification frameworks are established to perform synthetic video verification and model ownership verification. Extensive experiments on representative VGMs demonstrate that the proposed scheme achieves over 99% watermark extraction accuracy and 100% verification accuracy on both tasks, with negligible impact on video generation quality. Furthermore, the watermarks exhibit strong robustness against a comprehensive range of video-level and model-level attacks.
Chinese Translation
视频生成模型(VGM)的快速发展使得生成高度逼真的合成视频成为可能,同时也引发了对这些模型知识产权(IPR)的担忧。特别是,两个密切相关的取证任务在很大程度上尚未得到解决:合成视频验证(判断一段视频是否由受保护的VGM生成)和模型所有权验证(判断一个可疑的VGM是否是受保护VGM的未授权副本)。在本文中,我们提出了一种新的生成过程中水印方案,可以解决上述两个验证任务。首先,我们提出了一种名为VidMark的新型视频水印网络,它结合了双尺度离散小波变换(DWT)分解和全局时间注意力模块(GTAB),以增强水印的鲁棒性和不可感知性。其次,我们提出了一种解码器引导的微调方法。通过利用冻结的VidMark解码器,该过程使VGM能够合成携带不可感知、鲁棒且模型特定水印的视频。最后,我们建立了两个验证框架,分别用于执行合成视频验证和模型所有权验证。在代表性VGM上进行的大量实验表明,所提出的方案在两项任务上均实现了超过99%的水印提取准确率和100%的验证准确率,且对视频生成质量的影响可以忽略不计。此外,该水印对各种视频级和模型级攻击均表现出很强的鲁棒性。
cs.CV / 101 / 2609.23592

Collapse, Not Complexity: Failure-Conditioned Decomposition Repair for End-to-End Document Parsing

是崩溃而非复杂性:面向端到端文档解析的失败条件化分解修复方法
Lin, Xingyu, Du, Dehui
Abstract
End-to-end document parsers increasingly offer an optional reasoning mode for complex pages. On a 180-page entropy-stratified discovery sample with one frozen 4B checkpoint, complexity is the wrong decision variable. Reasoning lowers mean quality by 2.21 Overall at 1.54x tokens; a preregistered input-only model cannot predict its signed benefit (held-out AUROC 0.47, indistinguishable from chance). The benefit concentrates on pages whose ordinary pass has already collapsed, and they do not look complex: shared collapses have lower layout entropy than healthy ones yet consume 19x the tokens as degenerate repetition that doubling the budget does not cure. Switching modes rarely repairs them: 83% recur under reasoning. We instead detect collapse from the ordinary-pass trace, decompose the page by projection, and re-parse each region. Repair gains 1.40 Overall (95% CI [0.68, 2.16]) at 1.13x tokens, replicates across three checkpoints, and, with all parameters frozen, gains 2.41 (CI [1.64, 3.46]) on the remaining 1,175 benchmark pages.
Chinese Translation
端到端文档解析器日益为复杂页面提供可选的推理模式。在一个包含180页、按熵分层的发现样本上,使用一个冻结的4B检查点,我们证明复杂性是错误的决策变量。推理模式以1.54倍的token开销使平均质量下降2.21 Overall;一个预先注册的仅基于输入的模型无法预测其有符号收益(留出集AUROC为0.47,与随机猜测无异)。收益集中在普通解析已经崩溃的页面上,而这些页面看起来并不复杂:发生共享性崩溃的页面其版面熵低于健康页面,却消耗19倍的token,表现为加倍的预算也无法消除的退化重复。切换模式很少能修复它们:83%的崩溃在推理模式下再次出现。我们转而从普通解析的轨迹中检测崩溃,通过投影分解页面,并对每个区域重新解析。修复方法以1.13倍的token开销提升1.40 Overall(95% CI [0.68, 2.16]),在三个检查点上均可复现,且在所有参数冻结的条件下,在其余1,175个基准页面上提升2.41(CI [1.64, 3.46])。
cs.CV / 102 / 2609.23596

StyleAT: Defending Face Recognition Against Semantic Attacks

StyleAT:防御人脸识别中的语义攻击
Shapira, Ben, Cohen, Roi, Chen, Shang-Tse, Sharif, Mahmood
Abstract
With face-recognition models now embedded in everyday authentication and surveillance, recent works have pinpointed a critical weakness: these models remain acutely vulnerable to adversarial semantic edits. I.e., adversarially produced semantic alterations to the input, such as slight aging or pose changes, can induce misclassifications. Certain existing attacks are powerful, but they can be computationally costly, rendering them inadequate for developing defenses (e.g., through adversarial training). To fill the gap, we introduce BoundStyle, a potent semantic attack operating in StyleGAN's rich latent space to maximize misclassification rates. Notably, BoundStyle achieves high attack success rates while being ${\sim}{\times}9.5$ faster than existing state-of-the-art attacks, making it suitable for adversarial training. Building on BoundStyle, we develop StyleAT, an efficient adversarial training scheme that incorporates low-budget attack variants yet defends against stronger and unseen semantic attacks. We evaluate on two datasets unseen during training and seven models, and find that StyleAT boosts robust accuracy against state-of-the-art attacks and outperforms common defenses in various settings.
Chinese Translation
随着人脸识别模型如今被广泛应用于日常身份认证和监控系统中,近期的研究指出其一个关键弱点:这些模型对对抗性语义编辑仍然极其脆弱。也就是说,对输入进行对抗性生成的语义改动,如轻微的年龄或姿态变化,即可导致误分类。现有的某些攻击虽然强大,但计算开销高昂,使其难以用于构建防御手段(例如通过对抗训练)。为填补这一空白,我们提出了BoundStyle,一种在StyleGAN丰富的潜空间(latent space)中运行的强大语义攻击方法,旨在最大化误分类率。值得注意的是,BoundStyle在取得高攻击成功率的同时,速度比现有最先进的攻击方法快约9.5倍,因此非常适合用于对抗训练。基于BoundStyle,我们开发了StyleAT,一种高效的对抗训练方案,它仅使用低开销的攻击变体进行训练,却能抵御更强以及未见过的语义攻击。我们在两个训练期间未见过(unseen)的数据集和七个模型上进行评估,结果表明StyleAT显著提升了对最先进攻击的鲁棒准确率,并在多种设置下优于常见的防御方法。
cs.CV / 103 / 2609.23600

PETR: Prompt Ensembling with Training-free Routing for Vision-Language Models

PETR:面向视觉-语言模型的无训练路由提示集成方法
Cai, Weihan, Tan, Hao, Gao, Xinping, Xu, Shibiao, Wan, Jun
Abstract
Prompt learning efficiently adapts vision-language models (VLMs) to downstream tasks, but gains on seen classes often come at the expense of generalization to unseen classes. To address this limitation, we propose prompt ensembling with training-free routing (PETR), whose key innovation is a carefully designed dual-prompt architecture: two complementary prompts are learned from different data and objectives to emphasize seen class discrimination and unseen-class generalization, respectively. During training, both prompts are fine-tuned using a shared frozen CLIP backbone, and statistical information is collected from the training set logits. At inference time, we determine the similarity of each test sample to seen data, and route the sample to the most appropriate prompt branch. To the best of our knowledge, this is the first prompt tuning framework that performs training-free adaptive routing based on statistical similarity. This design provides an interpretable routing signal and avoids common MoE-style routing pathologies, such as router training instability and load imbalance. Extensive experiments on 11 benchmark datasets demonstrate that our framework consistently outperforms previous methods on both seen and unseen classes, achieving new state-of-the-art results.
Chinese Translation
提示学习(Prompt Learning)能够高效地将视觉-语言模型(VLMs)适配到下游任务,但在已见类别上取得的性能提升往往以牺牲对未见类别的泛化能力为代价。为解决这一局限,我们提出了基于无训练路由的提示集成方法(PETR),其核心创新在于精心设计的双提示架构:从不同数据和目标中学习两个互补的提示,分别侧重于已见类别的判别能力和未见类别的泛化能力。在训练过程中,两个提示均基于共享的冻结 CLIP 骨干网络进行微调,并从训练集的 logits 中收集统计信息。在推理阶段,我们计算每个测试样本与已见数据的相似度,并将样本路由到最合适的提示分支。据我们所知,这是首个基于统计相似度进行无训练自适应路由的提示微调框架。该设计提供了可解释的路由信号,并避免了常见的 MoE 式路由缺陷,如路由器训练不稳定和负载不均衡等问题。在 11 个基准数据集上的大量实验表明,我们的框架在已见和未见类别上均持续优于先前的方法,取得了新的最先进(state-of-the-art)结果。
cs.CV / 104 / 2609.23601

PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding

PREM:面向长视频理解的前缀引导循环记忆机制
Zhong, Siru, Wang, Qiongyan, Lv, Xiaohui, Zhuang, Yuzheng, Tao, Shuai, Liu, Wulong, Fu, Haohuan, Liang, Yuxuan
Abstract
Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value (KV) caches. We introduce Prefix-Steered Recurrent Memory (PREM), a memory-token-free framework for frozen vision-language models (VLMs). PREM separates video ingestion from query answering: a recurrent writer distills visual streams into a compact 256 KiB multi-slot associative state, while a question-conditioned readout adds memory-derived key/value (K/V) steering modulations to existing non-visual prompt prefixes during prefill. This enables write-once, query-many inference without extra prompt tokens or decoding recurrence. Across six long-video benchmarks in offline and streaming end-of-stream settings, PREM consistently outperforms frozen baselines at every evaluated visual budget. Under a constrained budget of 16 frames, PREM improves macro-average accuracy by 3.06% on Qwen2.5-VL-3B, with gains of 11.0% on action antonym identification and 9.9% on localized needle retrieval. These gains require tuning 0.24% of backbone parameters at 0.03 GiB of peak GPU memory overhead.
Chinese Translation
长视频理解需要在严格的token预算下捕捉瞬态视觉证据,然而现有方法通常压缩视频帧、附加记忆token或修改内部的键值(KV)缓存。我们提出了前缀引导循环记忆(Prefix-Steered Recurrent Memory, PREM),这是一个无需记忆token、适用于冻结视觉语言模型(VLM)的框架。PREM将视频摄入与问题解答相分离:一个循环写入器将视觉流提炼为紧凑的256 KiB多槽关联状态,同时一个基于问题条件化的读出机制在prefill阶段向现有的非视觉提示前缀添加源自记忆的键/值(K/V)引导调制。这使得模型能够在无需额外提示token或解码循环的情况下实现“一次写入、多次查询”的推理。在离线和流式端到端设置下的六个长视频基准测试中,PREM在所有评估的视觉预算下均持续优于冻结基线。在16帧的受限预算下,PREM在Qwen2.5-VL-3B上使宏平均准确率提升了3.06%,其中在动作反义词识别上提升11.0%,在局部细节检索上提升9.9%。这些提升仅需微调0.24%的骨干网络参数,且GPU峰值显存开销仅为0.03 GiB。
cs.CV / 105 / 2609.23606

Beyond UV Mapping: Mesh Texture Compression via Surface-Aligned Texture Fields

超越UV映射:基于表面对齐纹理场的网格纹理压缩
Wang, Jianqiang, Hou, Junhui, Ren, Siyu, Lin, Weiyao, Wang, Wenping
Abstract
Mesh texture compression typically relies on 2D UV atlases, whose chart discontinuities and mapping overhead can limit coding efficiency. To tackle this challenge, we introduce TexF, a surface-aligned texture field that organizes texture attributes in sparse voxels derived from the mesh surface. This representation supports high-resolution textures while preserving local 3D correlations for compression and enabling direct surface queries. For bitstream compression, TexF reuses established 3D attribute codecs, with voxel locations reconstructed from the decoded mesh without separate transmission. For GPU-resident compression, we develop 3DNTC, which combines quantized hash features with a lightweight decoder for random-access reconstruction at surface positions. Differentiable rendering enables image-space refinement of both voxel attributes and compressed neural fields. Experiments on the MPEG and AOM mesh compression benchmarks demonstrate improved average rate-distortion performance over representative UV-based methods for both bitstream and GPU-resident compression. 3DNTC also supports real-time rendering.
Chinese Translation
网格纹理压缩通常依赖于二维UV图集,其图表不连续性和映射开销会限制编码效率。为应对这一挑战,我们提出了TexF,一种表面对齐的纹理场,它将纹理属性组织在从网格表面导出的稀疏体素中。这种表示方法在支持高分辨率纹理的同时,保留了用于压缩的局部三维相关性,并支持直接在表面上进行查询。在码流压缩方面,TexF复用了已有的三维属性编解码器,体素位置可从解码后的网格中重建,无需单独传输。在GPU驻留压缩方面,我们开发了3DNTC,它将量化哈希特征与轻量级解码器相结合,实现表面位置的随机访问重建。可微渲染技术使体素属性和压缩神经场均可在图像空间中进行优化。在MPEG和AOM网格压缩基准上的实验表明,无论是码流压缩还是GPU驻留压缩,本方法在平均率失真性能上均优于具有代表性的基于UV的方法。此外,3DNTC还支持实时渲染。
cs.CV / 106 / 2609.23619

Compact Low-Cost Hyperspectral Imaging via Angular-to-Spectral Diversity Conversion

基于角度-光谱分集转换的紧凑型低成本高光谱成像
Fujiwara, Kazuma, Funatomi, Takuya, Kitano, Kazuya, Fujimura, Yuki, Mukaigawa, Yasuhiro
Abstract
Snapshot hyperspectral imaging avoids sequential scanning, but systems that jointly achieve stable reconstruction, low cost, and compact optics remain limited. We present a snapshot hyperspectral imaging system based on angular-to-spectral diversity conversion. A tapered kaleidoscope creates replicated views with distinct incidence directions, and a directly attached birefringent filter converts them into view-channel-dependent spectral transmittances, yielding complementary measurements that better condition the inverse problem for more stable single-shot spectral reconstruction. The system preserves a simple pixel-wise linear model for fast non-learning-based reconstruction and uses only off-the-shelf components without relay optics or cascaded modules. We select the birefringent filter configuration using a condition-number-based criterion and validate the system on both synthetic and real data.
Chinese Translation
快照式高光谱成像避免了顺序扫描,但能够同时实现稳定重建、低成本和紧凑光学结构的系统仍然有限。我们提出了一种基于角度-光谱分集转换的快照式高光谱成像系统。锥形万花筒产生具有不同入射方向的复制视图,直接贴合的双折射滤光片将这些视图转换为依赖于视图通道的光谱透射率,从而产生互补的测量结果,更好地改善了逆问题的条件性,实现更稳定的单次曝光光谱重建。该系统保持了简单的逐像素线性模型,可实现快速的非学习式重建,并且仅使用现成的组件,无需中继光学器件或级联模块。我们采用基于条件数的准则选择双折射滤光片配置,并在合成数据和真实数据上验证了该系统。
cs.CV / 107 / 2609.23655

Reassessing Global Gradient-Norm Imbalance in BLIP Fine-Tuning Across Physical Domains

重新审视跨物理领域中BLIP微调的全局梯度范数不平衡问题
Naseer, Kiran, Azhar, Samreen, Mahapatra, Dwarikanath
Abstract
Imbalanced gradient magnitudes between the visual and language pathways of a vision-language model are often treated as a defect to be corrected. We test that premise for one family of correction, deliberately excluding adaptive, signal-driven schemes (e.g. BalGrad, OGM, PMR, CGGM), which are a mechanistically distinct class outside this study's scope. Measuring the language-to-visual gradient-norm ratio, reported in parameter-normalised form, across nine fine-tuning conditions, three seeds, and three captioning datasets spanning distinct physical domain shifts -- underwater, aerial, radiological -- we find imbalance magnitude varies markedly across domains with no predictable ordering. A plain learning-rate reduction cuts imbalance substantially and lands within a few BLEU points of the best method on every dataset. Staged freezing reduces the ratio on every domain yet never ranks first; a schedule-only control isolates freezing as the cause on one dataset but not the other two. Forcing the two gradient groups to equal magnitude drives per-parameter imbalance close to zero on every domain, yet is both the best result in the study and the worst placement among full fine-tuning methods, on different datasets, with identical settings. Reductions in gradient-norm ratio do not consistently predict captioning performance across domains, and how a given level of balance is reached matters as much as the level itself. As a secondary finding, a commonly reused LoRA configuration applied to BLIP silently adapts zero visual parameters; correcting it improves BLEU-4 on all three datasets.
Chinese Translation
视觉-语言模型中视觉通路与语言通路之间的梯度幅值不平衡常被视为需要纠正的缺陷。我们针对一类纠正方法检验了这一前提,并有意排除了自适应的、基于信号驱动的方案(如 BalGrad、OGM、PMR、CGGM),它们在机理上属于另一类方法,不在本研究范围之内。通过测量语言通路与视觉通路的梯度范数之比(以参数归一化形式报告),我们覆盖了九种微调条件、三个随机种子,以及三个跨越不同物理域偏移——水下、航空、放射影像——的图像描述数据集。结果发现,不平衡程度在不同域之间差异显著,且无可预测的排序规律。简单地降低学习率即可大幅削减不平衡,并在每个数据集上都达到与最佳方法相差无几的BLEU分数。分阶段冻结策略在每个域上都能降低该比值,却从未排名第一;在其中一个数据集上,仅使用调度而冻结的对照实验将不平衡的成因归因于冻结本身,而在另外两个数据集上则不然。强制两组梯度幅值相等可使逐参数不平衡在每个域上都接近于零,但在相同设置下,该方法在不同数据集上既是本研究中最好的结果,也是全量微调方法中排名最差的。梯度范数比的降低并不能跨域一致地预测图像描述性能,且达到某一平衡水平的方式与该水平本身同等重要。作为一项次要发现,常被重复使用的BLIP LoRA配置实际上静默地适应了零个视觉参数;修正该配置后,三个数据集上的BLEU-4均有所提升。
cs.CV / 108 / 2609.23658

Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

为什么视频扩散模型会违背物理规律?揭示注意力机制的缺陷
Li, Yueyan, Wang, Haibo, Yuan, Caixia, Wang, Xiaojie
Abstract
Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-world physical laws. While existing solutions rely on external priors or specialized data, we investigate the root cause by exploring the internal mechanisms of these models. Specifically, we present the first interpretability study on the ''motion planning'' process of text-to-video diffusion models, revealing how motion trajectories form during early denoising stages. Building upon the ''first shape, then details'' finding, we combine cross-attention trajectory patterns with causal head contributions to identify a specific subset of attention heads driving motion planning. Further, our self-attention analysis shows that Rotary Position Embedding (RoPE) induces excessive spatial attention decay. This causes early candidate regions to prematurely lock into physically implausible positions, suppressing reasonable trajectories in adjacent frames and triggering generation failure modes. To address this fundamental flaw, we propose a lightweight architectural modification that scales the frequency of RoPE across different denoising steps. This strategy reduces excessive attention decay, helping the model explore better candidate regions to establish coherent physical motion. Finally, training-free and training-based experiments confirm the effectiveness of our approach in enhancing the physical commonsense of generated videos.
Chinese Translation
尽管最先进的视频扩散模型具有出色的视觉质量,但其生成的内容常常违背现实世界的物理定律。现有解决方案依赖于外部先验或专门数据,而我们通过探究这些模型的内部机制来研究其根本原因。具体而言,我们首次对文本到视频扩散模型的“运动规划”过程进行了可解释性研究,揭示了运动轨迹在早期去噪阶段是如何形成的。基于“先成形、后细化”的发现,我们将交叉注意力的轨迹模式与因果头贡献相结合,识别出驱动运动规划的特定注意力头子集。此外,我们的自注意力分析表明,旋转位置编码(RoPE)会导致过度的空间注意力衰减。这使得早期候选区域过早地锁定在物理上不合理的位置,抑制了相邻帧中的合理轨迹,并触发生成失败模式。为解决这一根本缺陷,我们提出一种轻量级的架构修改方法,在不同的去噪步骤中对RoPE的频率进行缩放。该策略减少了过度的注意力衰减,帮助模型探索更好的候选区域,以建立符合物理规律的运动。最后,免训练和基于训练的实验均证实了我们的方法在增强生成视频物理常识方面的有效性。
cs.CV / 109 / 2609.23673

Which Terrain Is Better? Preference Learning with VLM Prototypes for Off-Road Traversability Ranking

哪种地形更好?基于VLM原型的偏好学习用于越野可通行性排序
Hwang, Ji-Hoon, Bae, Jisung, Son, E-In, Kim, Dong-Wook, Kim, Jung-Taak, Seo, Seung-Woo
Abstract
In vision-based off-road navigation, a robot needs to know not only which obstacles to avoid but also which terrain is better. The first is handled by freespace detection or semantic segmentation. The second is usually answered with a traversability score, but no universal ground truth exists for such a score, so perception falls back on a predefined value per semantic class or a freespace confidence. These scores say what a region is, not which region a robot should prefer. We therefore formulate this preference as visual traversability ranking, an ordering of visible terrain that can be supervised by comparisons between two regions. Standard annotations do not label preference, but they imply its direction. We present TravPro, which converts these annotations into ordered region pairs and fits a small readout on frozen vision--language model (VLM) patch tokens to these pairs. The tokens are clustered once into a fixed prototype bank, and the readout learns a preference score per prototype. The readout is then applied to every patch and serves as a teacher that turns sparse comparisons into dense preference pseudo-labels without pixel-wise annotation. An RGB student distills these maps into a dense terrain-preference map together with a non-ground mask that excludes obstacles and background from the ranking. On five unseen domains, TravPro reaches a mean pairwise accuracy of 0.915 against 0.783 for the strongest baseline, producing an ordering sensitive to surface condition that a per-class value cannot represent. The same VLM and the same supervision yield no such ordering when the VLM is prompted and the supervision is used as dense targets; what matters is how they are used.
Chinese Translation
在基于视觉的越野导航中,机器人不仅需要知道应避开哪些障碍物,还需要知道哪种地形更好。前者由可行区域检测或语义分割处理;后者通常用可通行性分数来回答,但此类分数不存在通用的真值标准,因此感知部分只能退而依赖每个语义类别的预定义值或可行区域置信度。这些分数描述的是一个区域是什么,而不是机器人应该偏好哪个区域。因此,我们将这一偏好问题形式化为视觉可通行性排序,即对可见地形的排序,可以通过两个区域之间的比较进行监督。标准标注并未直接标注偏好,但蕴含了其方向。我们提出了TravPro,它将这些标注转换为有序的区域对,并在冻结的视觉-语言模型(VLM)patch token上拟合一个小型读出层以匹配这些区域对。这些token被一次性聚类为固定的原型库,读出层为每个原型学习一个偏好分数。读出层随后被应用于每个patch,并作为教师将稀疏比较转化为密集的偏好伪标签,无需逐像素标注。一个RGB学生网络将这些偏好图蒸馏为密集的地形偏好图,同时生成一个非地面掩膜,将障碍物和背景排除在排序之外。在五个未见过的域上,TravPro的平均成对准确率达到0.915,而最强基线为0.783,所产生的排序对表面状况敏感,这是每类别单一数值无法表达的。当直接提示VLM并将监督信号用作密集目标时,同一VLM和同一监督却无法产生这样的排序;关键在于如何使用它们。
cs.CV / 110 / 2609.23679

Mind the Gaps: A Curated Benchmark for Form Field Detection

留意空白:一个精心构建的表单字段检测基准
Brini, Iheb, Moured, Omar, Gbada, Hamza, Barney, Elisa
Abstract
Form Field Detection (FFD) is a fundamental component of document understanding systems, enabling applications ranging from large-scale industrial digitization to accessible form interaction for automated analysis. Unlike conventional object detection tasks, FFD is inherently challenging because fields are often defined by layout structure and whitespace rather than visible foreground content. Existing large-scale datasets frequently rely on heuristic annotation pipelines, resulting in noisy and inconsistent labels that hinder reliable evaluation. In this work, we introduce mini-CommonForms, a carefully curated FFD benchmark with consistent, high-quality annotations, and present a detailed evaluation of state-of-the-art detection approaches. The benchmark is designed to support reproducible research in document automation and accessibility-oriented applications. Dataset and code are available at https://github.com/moured/mini-commonforms
Chinese Translation
表单字段检测(Form Field Detection, FFD)是文档理解系统的基础组成部分,支撑着从大规模工业数字化到面向自动化分析的无障碍表单交互等多种应用。与传统的目标检测任务不同,FFD 本质上具有挑战性,因为表单字段通常由版面结构和空白区域定义,而非可见的前景内容。现有的大规模数据集往往依赖启发式标注流程,导致标签噪声大且不一致,妨碍了可靠的评估。在本工作中,我们提出了 mini-CommonForms,一个经过精心构建、标注一致且高质量的 FFD 基准,并对最先进的检测方法进行了详细评估。该基准旨在支持文档自动化及面向无障碍应用的可复现研究。数据集与代码可在 https://github.com/moured/mini-commonforms 获取。
cs.CV / 111 / 2609.23714

Infectious Bovine Pinkeye Detection Using Computer Vision and Imbalance-Aware Learning

基于计算机视觉与类别不平衡感知学习的牛传染性红眼病检测
Abalo, Michael, Brennan, Jameson, Rekabdarkolaee, Hossein Moradi
Abstract
Infectious bovine pinkeye is a contagious ocular disease that adversely affects cattle health, welfare, and agricultural productivity. Conventional diagnosis relies primarily on clinical observation, which can be subjective, time-consuming, and difficult to implement efficiently in large herds or remote settings. This study evaluated and compared You Only Look Once (YOLO) v11 and YOLOv26 for automated bovine pinkeye classification and investigated the effects of class-balancing strategies on model performance. Five variants (n, s, m, l, and x) of each architecture were trained and evaluated using the original imbalanced dataset, Random Minority Oversampling (RMO), and an adapted Synthetic Minority Oversampling Technique (SMOTE). Both YOLOv11 and YOLOv26 demonstrated strong classification performance, although the effects of class balancing varied across model variants. For YOLOv11, RMO-s achieved an accuracy of 0.99, a macro F1-score of 0.98, and a true positive rate (TPR) of 1.00, with no false-negative classifications. RMO-m also achieved a TPR of 1.00 with no false negatives. For YOLOv26, the original l, RMO-m, and RMO-l variants each achieved an accuracy of 0.99 and a macro F1-score of 0.98, with RMO-l attaining a TPR of 1.00 and no false negatives. Overall, RMO generally provided greater improvements in minority-class detection than adapted SMOTE, whereas the strong performance of the original YOLOv26-l demonstrates that oversampling was not necessary for all model variants. These findings demonstrate the potential of YOLOv11 and YOLOv26 for automated detection of bovine pinkeye and support further evaluation for livestock health monitoring.
Chinese Translation
牛传染性红眼病(Infectious Bovine Pinkeye)是一种接触性眼部疾病,对牛只健康、动物福利和农业生产率均有不利影响。传统诊断主要依赖临床观察,具有主观性强、耗时长,且难以在大规模牛群或偏远环境中高效实施等问题。本研究评估并比较了 You Only Look Once (YOLO) v11 与 YOLOv26 在牛红眼病自动分类中的表现,并探讨了类别平衡策略对模型性能的影响。针对每种架构的五个变体(n、s、m、l 和 x),分别使用原始不平衡数据集、随机少数类过采样(Random Minority Oversampling, RMO)以及一种改进的合成少数类过采样技术(Synthetic Minority Oversampling Technique, SMOTE)进行训练和评估。YOLOv11 和 YOLOv26 均表现出较强的分类性能,但类别平衡策略的效果在不同模型变体之间存在差异。对于 YOLOv11,RMO-s 达到了 0.99 的准确率、0.98 的宏平均 F1 分数(macro F1-score)以及 1.00 的真阳性率(TPR),且无假阴性分类;RMO-m 也达到了 1.00 的 TPR,且无假阴性。对于 YOLOv26,原始 l 变体、RMO-m 和 RMO-l 均达到 0.99 的准确率和 0.98 的宏平均 F1 分数,其中 RMO-l 达到了 1.00 的 TPR 且无假阴性。总体而言,RMO 在少数类检测方面的改进通常优于改进版 SMOTE,而原始 YOLOv26-l 的优异表现则表明过采样并非对所有模型变体都是必要的。这些发现证明了 YOLOv11 和 YOLOv26 在牛红眼病自动检测方面的潜力,并支持其在牲畜健康监测中的进一步评估。
cs.CV / 112 / 2609.23715

Layer-Aware Position Embeddings for Visual Token Pruning in Multimodal Large Language Models

面向多模态大语言模型视觉词元剪枝的层级感知位置编码方法
Wang, Yahong, Ni, Zhangkai, Wu, Juncheng, Zhou, Yuyin, Wen, Ying, He, Lianghua
Abstract
Multimodal large language models (MLLMs) incur substantial computational overhead due to the reliance on hundreds of visual tokens to represent images. While token pruning has emerged as a promising approach to reduce the inference cost of MLLMs, existing methods typically reassign position embeddings to the retained tokens using either sparse or continuous position embeddings, each introducing distinct limitations. Sparse position embeddings tend to decrease the attention value allocated to visual tokens, thereby degrading the perception capability of MLLMs, whereas continuous position embeddings disrupt the original spatial correspondence of visual tokens, leading to weakened grounding capability. To mitigate this issue, we perform layer-wise analysis of the language decoder and observe that intermediate layers play a critical role for maintaining the grounding capability of MLLMs under token pruning. Based on this observation, we propose a layer-aware position embedding strategy, which switches to sparse position embeddings at grounding-sensitive layers while maintaining continuous position embeddings elsewhere. Extensive experiments across representative pruning methods and diverse benchmarks demonstrate that our approach improves the comprehensive multimodal performance of pruned MLLMs compared with standard sparse and continuous position embeddings.
Chinese Translation
多模态大语言模型(MLLMs)依赖数百个视觉词元来表示图像,因而产生巨大的计算开销。尽管词元剪枝已成为降低多模态大语言模型推理成本的一种有前景的方法,但现有方法通常使用稀疏位置编码或连续位置编码为保留的词元重新分配位置编码,二者均存在各自的局限。稀疏位置编码往往会降低分配给视觉词元的注意力值,从而削弱多模态大语言模型的感知能力;而连续位置编码则破坏了视觉词元原有的空间对应关系,导致定位(grounding)能力下降。为缓解这一问题,我们对语言解码器进行了逐层分析,发现中间层在词元剪枝下维持多模态大语言模型的定位能力方面起着关键作用。基于这一观察,我们提出了一种层级感知的位置编码策略,在定位敏感的层切换为稀疏位置编码,而在其他层保持连续位置编码。在多种代表性剪枝方法和多个基准测试上的大量实验表明,与标准的稀疏和连续位置编码相比,我们的方法提升了剪枝后多模态大语言模型的综合多模态性能。
cs.CV / 113 / 2609.23717

BindCLIP: One Balanced Coupling For Compositional Vision Language Scoring

BindCLIP:一种用于组合式视觉-语言评分的均衡耦合方法
Song, Liuyang, Zhang, Yi, Deng, Zhongyi, Yang, Daqian, Zhang, Hongbo
Abstract
Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which region or how those regions are arranged; a model can recognize every word and object yet prefer a compositionally incorrect caption. We argue that a frozen encoder retains this association structure, so the problem is to read it rather than to rebuild it beside the pretrained similarity. We introduce BindCLIP, a pairwise scorer built on one latent object: a balanced token--patch--depth optimal-transport coupling that places both candidate captions and several visual depths in a single plan. Semantic, entity, order, and spatial evidence are read as energies of this state, and exchanging the candidates permutes the plan, making the score exactly antisymmetric. A geometric refinement inside the coupling contracts moves that the candidates and the visual depths do not support. No task label, parser, relation inventory, or detector is used. One checkpoint and one inference path improve the official What'sUp, ARO, and SugarCrepe benchmarks over frozen global CLIP, with the strongest transfer on the relation splits. Controls rule out patch access and caption-length shortcuts, and an inference-time lesion localizes spatial arrangement to the coupling.
Chinese Translation
全局视觉-语言相似度将图像和文本描述压缩为单一向量,虽然保留了语义,但无法体现哪个词对应哪个区域以及这些区域如何排列;模型可以识别每一个词和每一个物体,却可能更偏好一个组合上错误的描述。我们认为,冻结的编码器仍保留了这种关联结构,因此问题在于如何读取它,而不是在预训练相似度之外重新构建它。我们提出 BindCLIP,这是一种建立在单一潜在对象之上的成对评分器:一个均衡的词元-图像块-深度最优传输耦合,将候选描述和多个视觉深度置于同一个传输方案中。语义、实体、顺序和空间等证据被读取为该状态的能量,而交换候选描述会对传输方案进行置换,从而使评分严格满足反对称性。耦合内部的几何细化机制会收缩候选描述和视觉深度均不支持的匹配移动。该方法不使用任何任务标签、句法分析器、关系清单或检测器。仅用一个检查点、一条推理路径,即可在官方 What'sUp、ARO 和 SugarCrepe 基准上超越冻结的全局 CLIP,且在关系类数据划分上的迁移能力最强。对照实验排除了图像块访问和描述长度方面的捷径,而推理时的消融实验将空间排列能力定位于该耦合机制。
cs.CV / 114 / 2609.23733

VGGT-Prime: Compute-Adaptive Mixture-of-Heads for Efficient Visual Geometry Transformers

VGGT-Prime:面向高效视觉几何Transformer的计算自适应混合头模型
Arab, Abteen, Wu, Guile, Huang, Chengjie, Bai, Dongfeng
Abstract
Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-view images. Despite their promising performance, these models scale quadratically with the number of input views due to their global attention mechanism, resulting in substantial latency for long sequence inputs. There have been some recent efforts to accelerate VGGT, but they primarily focus on reducing \emph{token redundancy} through token merging or key/value sparsification. Our work resolves this bottleneck from a different perspective by investigating \emph{architectural redundancy} in visual geometry transformers. We show that the multi-head attention modules in VGGT's global-attention layers contain substantial architectural redundancy, with only a subset of heads carrying critical geometric information. In light of this observation, we propose VGGT-Prime, a compute-adaptive mixture-of-heads model that resolves this redundancy to accelerate visual geometry transformers while maintaining competitive reconstruction quality. The key idea of VGGT-Prime is to estimate the appropriate computation level for each global-attention head using a lightweight router and then dynamically assign each head to different computation modes. Extensive experiments on multiple datasets demonstrate that VGGT-Prime can achieve an {$8\times$} inference speedup over VGGT while maintaining competitive performance on camera pose, depth, and point-cloud predictions. We further show that VGGT-Prime is complementary to existing acceleration methods, such as token merging, further improving inference speed by up to $14{\times}$ over VGGT. An overview of our work is available on our \href{https://vggt-prime.github.io}{project page}.
Chinese Translation
以视觉几何基础Transformer(Visual Geometry Grounded Transformer,VGGT)为代表的前馈式视觉几何模型,近来实现了从多视角图像直接进行三维重建。尽管性能可观,但由于其全局注意力机制,这类模型的计算量随输入视角数量呈二次增长,导致长序列输入时的显著延迟。近期已有一些加速VGGT的工作,但它们主要聚焦于通过token合并或键/值稀疏化来减少token冗余。我们的工作从另一个角度解决这一瓶颈,即研究视觉几何Transformer中的架构冗余。我们发现,VGGT全局注意力层中的多头注意力模块存在大量架构冗余,仅有部分注意力头承载关键的几何信息。基于这一观察,我们提出VGGT-Prime,一种计算自适应的混合头(mixture-of-heads)模型,通过化解该冗余来加速视觉几何Transformer,同时保持有竞争力的重建质量。VGGT-Prime的核心思想是利用一个轻量级路由器为每个全局注意力头估计合适的计算量,然后将各个头动态分配到不同的计算模式。在多个数据集上的大量实验表明,VGGT-Prime相对于VGGT可实现8倍的推理加速,同时在相机位姿、深度和点云预测上保持有竞争力的性能。我们进一步表明,VGGT-Prime与现有的加速方法(如token合并)互补,相对于VGGT可将推理速度进一步提升至最高14倍。我们工作的概览请参见项目主页(https://vggt-prime.github.io)。
cs.CV / 115 / 2609.23753

OnlineWM: Causality-Aware Active Online Learning for Effective World Modeling

OnlineWM:面向有效世界建模的因果感知主动在线学习
Miao, Yikun, Zhu, Fangqi, Shou, Quanxin, Pang, Xiaoyi, Yan, Zhengyang, Li, Junhao, Wang, Haodong, Hong, Zicong, Guo, Song
Abstract
Generative world models aim to predict future states conditioned on actions, where action controllability is fundamental for reliable dynamics modeling. While recent efforts leverage simulator-generated data to enhance this capability, existing training pipelines face two fundamental limitations. First, static offline data collection leads to a distribution misalignment between training sets and the model's evolving error patterns, failing to resolve critical long-tail scenarios where dynamics predictions remain unreliable. Second, the standard objective of minimizing observational discrepancy often encourages the model to exploit spurious correlations instead of capturing the underlying action-effect causality. To address these limitations, we propose OnlineWM, an online training framework that continuously improves world modeling through active simulator interaction and causality-aware optimization. OnlineWM introduces two key innovations: (1) Active Online Learning: Instead of using fixed datasets, OnlineWM adaptively queries the simulator for new interaction sequences that target the model's current predictive weaknesses, ensuring high-utility data acquisition. (2) Causality-Aware Fine-Tuning: We propose a counterfactual learning strategy that contrasts the outcomes of different actions from identical states, forcing the model to attribute state transitions to specific actions rather than ambient environmental evolution, thereby grounding its predictions in reliable causal mechanisms. By integrating active data acquisition with causal optimization, OnlineWM establishes a closed-loop refinement process that ensures the model is both robust to diverse scenarios and precise in its causal attribution. Extensive experiments demonstrate that OnlineWM significantly enhances action controllability and generalizes effectively to unseen domains.
Chinese Translation
生成式世界模型旨在基于动作预测未来状态,其中动作可控性是可靠动力学建模的基础。尽管最近的研究利用模拟器生成的数据来增强这一能力,但现有的训练流程面临两个根本性局限。其一,静态的离线数据采集导致训练集与模型不断演化的误差模式之间分布不对齐,无法解决动力学预测仍然不可靠的关键长尾场景。其二,最小化观测差异的标准目标往往促使模型利用虚假相关性,而非捕捉潜在的动作-效果因果关系。为解决这些局限,我们提出 OnlineWM,一种通过主动模拟器交互和因果感知优化持续改进世界建模的在线训练框架。OnlineWM 引入了两项关键创新:(1)主动在线学习:OnlineWM 不使用固定数据集,而是自适应地向模拟器查询针对模型当前预测弱点的新交互序列,从而确保高效用数据的获取。(2)因果感知微调:我们提出一种反事实学习策略,通过对比相同状态下不同动作所产生的结果,迫使模型将状态转移归因于特定动作而非环境的自发演化,从而使其预测建立在可靠的因果机制之上。通过将主动数据获取与因果优化相结合,OnlineWM 建立了一个闭环改进过程,确保模型既对多样场景具有鲁棒性,又能进行精确的因果归因。大量实验表明,OnlineWM 显著增强了动作可控性,并能有效泛化到未见过的领域。
cs.CV / 116 / 2609.23758

Training-Free Spectral Transductive Refinement for Cross-Domain Few-Shot Classification

免训练的谱直推式精化方法用于跨域小样本分类
Rahman, Fahim, Rohan, S. M. Tanjeeb Meheran, Sayed, Md. Taimum Ibne, Herok, Asaduzzaman, Hasan, Md. Bakhtiar
Abstract
Few-shot recognition with frozen visual features is especially fragile under domain shift and one-shot supervision, where a single labelled image is an unreliable estimate of its class. We ask how far this fragility can be reduced purely at test time, without retraining the encoder or augmenting the source domain. We present Spectral Transductive Refinement (STR), a training-free transductive inference rule that exploits the geometry of the complete support-query episode. Given frozen embeddings, STR builds a joint k-nearest-neighbour graph, maps the episode into a normalized-Laplacian spectral coordinate system, initializes class representatives from the labelled support, and iteratively refines them using pseudo-labelled queries. We evaluate STR under two protocols. A controlled component study with frozen ResNet-18 features shows that spectral refinement consistently improves over single-prototype spectral initialization across five shifted domains, with the largest gains in the one-shot regime where support estimates are weakest. We then benchmark STR against recent Cross-Domain Few-Shot Learning (CD-FSL) methods using the standard miniImageNet-pretrained ResNet-10 backbone over eight established target domains. Operating entirely at inference time, STR attains the highest 1-shot average among compared methods and remains competitive at 5-shot, rivalling approaches relying on heavy source-domain meta-training augmentations. Because STR is transductive, we report its setting explicitly. Diagnostics attribute its gains to iterative refinement in spectral coordinates rather than added prototype capacity, which remains inactive in our configuration.
Chinese Translation
基于冻结视觉特征的小样本识别在域偏移和单样本监督下尤其脆弱,因为单张标注图像对其类别的估计并不可靠。我们探究能否仅在测试阶段减少这种脆弱性,而不重新训练编码器或增广源域。我们提出谱直推式精化(Spectral Transductive Refinement, STR),这是一种免训练的直推式推理规则,它利用完整的支持集-查询集情景(episode)的几何结构。给定冻结的嵌入特征,STR 构建一个联合 k 近邻图,将该情景映射到归一化拉普拉斯谱坐标系中,从有标注的支持样本初始化类代表,并利用伪标注的查询样本对其进行迭代精化。我们在两种协议下评估 STR。在使用冻结 ResNet-18 特征的受控组件研究中,谱精化在五个偏移域上始终优于单原型谱初始化,且在支持样本估计最弱的单样本(one-shot)情形下收益最大。随后,我们使用标准的 miniImageNet 预训练 ResNet-10 骨干网络,在八个公认的目标域上将 STR 与近期的跨域小样本学习(CD-FSL)方法进行基准比较。STR 完全在推理阶段运行,在所比较的方法中取得了最高的 1-shot 平均准确率,并在 5-shot 情形下保持竞争力,可与依赖大量源域元训练增广的方法相媲美。由于 STR 是直推式的,我们明确报告了其设定。诊断分析表明,其收益来自谱坐标系中的迭代精化,而非原型容量的增加——后者在我们的配置中并未发挥作用。
cs.CV / 117 / 2609.23769

PRISM-RAG: Multimodal Hypergraph Retrieval-Augmented Generation for Tobacco Product and Legislative Policy Reasoning

PRISM-RAG:面向烟草产品与立法政策推理的多模态超图检索增强生成
Serna-Aguilera, Manuel, Anderes, Raegan, Dobbs, Page, Luu, Khoa
Abstract
The disambiguation of semantically similar statutory text across jurisdictions is a retrieval problem that existing methods do not solve. This inter-context conflict can steer generative models toward confidently produced answers grounded in topically relevant but jurisdictionally incorrect sources. Tobacco and nicotine regulations vary by US jurisdiction, often sharing similar language, thus, robust reasoning requires identifying which jurisdiction's law governs a given product, not merely retrieving relevant text. Emerging products (e.g., pouches) exploit ambiguous definitions to evade regulation. State-of-the-art (SOTA) document retrieval-augmented generation (RAG) methods struggle to address this inter-context conflict, and thus struggle to connect image attributes (e.g., rich attribute captions) to the set of similar legislation texts. We introduce NicoPRISM (Nicotine Product and Regulation Image-and-Text Surveillance Multimodal), comprising 161,563 images, attribute captions, a knowledge base of product, health, and legislative documents spanning 13 US jurisdictions, and 1,495 validated question-answer pairs across two tasks: policy compliance QA and product knowledge QA. We also propose PRISM-RAG, a multimodal hypergraph RAG framework built over images, captions, and entities without any LLM calls at index time, grounding every query in a product image and routes retrieval through a jurisdiction-aware context assembly mechanism guaranteeing that statutory text from the queried jurisdiction reaches the language model by construction. PRISM-RAG retrieves passages from the correct jurisdiction in 93.9% of policy compliance queries, a 48.6 percentage point advantage over standard RAG (p<0.001), using zero LLM calls at index time and one at query time, and is competitive with or outperforms SOTA RAG frameworks across keyword, semantic, jurisdiction-, and compliance-accuracy metrics.
Chinese Translation
跨司法辖区语义相似的成文法律文本的消歧是一个现有方法尚未解决的检索问题。这种跨文本间的冲突可能引导生成模型给出看似自信、但其依据仅在主题上相关而在司法辖区上却错误的答案。美国各司法辖区的烟草与尼古丁法规各不相同,却常常使用相似的语言表述,因此,稳健的推理需要识别哪个司法辖区的法律管辖某一给定产品,而不仅仅是检索相关文本。新兴产品(如尼古丁袋)则利用定义上的模糊性来规避监管。最先进(SOTA)的文档检索增强生成(RAG)方法难以应对这种跨文本冲突,因而难以将图像属性(如丰富的属性描述)与相似的法律文本集合建立关联。我们提出了 NicoPRISM(尼古丁产品与法规图像-文本监测多模态数据集),包含161,563张图像及属性描述,一个涵盖美国13个司法辖区的产品、健康与立法文档知识库,以及1,495个经过验证的问答对,分为两个任务:政策合规问答和产品知识问答。我们还提出了 PRISM-RAG,一个构建于图像、属性描述和实体之上的多模态超图 RAG 框架,在索引阶段无需任何大语言模型(LLM)调用。该框架以产品图像为每个查询提供依据,并通过一种辖区感知的上下文组装机制进行检索路由,从构造上保证来自所查询辖区的成文法律文本能够送入语言模型。PRISM-RAG 在政策合规类查询中有93.9%的检索结果来自正确的司法辖区,较标准 RAG 高出48.6个百分点(p<0.001),且索引阶段零次 LLM 调用、查询阶段仅一次调用;在关键词、语义、辖区准确率与合规准确率等各项指标上,其表现与 SOTA RAG 框架相当或更优。
cs.CV / 118 / 2609.23796

Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene

Mira-Scene:面向生成式3D场景的像素对齐布局
Sun, Yang-Tian, Liu, Tianjia, Huang, Zehuan, Huang, Yi-Hua, Lyu, Xiaoyang, Yang, Ziyi, Zou, Zi-Xin, Guo, Yuan-Chen, Cao, Yan-Pei, Qi, Xiaojuan
Abstract
Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open challenge. A central difficulty lies in how object layout is represented. Holistic methods absorb placement into a scene-level generation process, sacrificing object-level detail. Compositional methods preserve object fidelity by decoupling geometry from layout, but typically parameterize layout as sparse, unbounded pose variables that are difficult to learn and generalize poorly under scarce scene-level supervision.We present Mira-Scene, a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery. At its core is the Canonical Coordinate Map (CCM), a pixel-aligned field that maps each visible object pixel to a surface coordinate in the object's bounded canonical space. When paired with a scene-space Point Cloud Map (PCM) from monocular geometry estimation, CCM induces dense canonical-to-scene correspondences from which object transformations are recovered through robust geometric alignment. Because CCM operates in bounded canonical space, it provides a stable prediction target that can be trained from scalable object-level 3D data without requiring scene-level layout annotations. Mira-Scene further introduces a multimodal diffusion transformer that jointly generates object geometry and CCMs, using modality-specific expert streams with shared attention and positional encoding to promote geometry-layout consistency. Experiments on indoor, outdoor, synthetic, and in-the-wild scenes show that Mira-Scene substantially outperforms strong baselines in layout accuracy, achieving relative gains of 39.8% in 3D-IoU and 16.5% in 2D-IoU over SAM3D, using limited open-source training data.
Chinese Translation
单图像3D物体生成现已能够产出高保真的资产,然而如何将这些物体准确地放置到连贯的场景布局中仍是一个开放性挑战。其核心难点在于物体布局的表示方式。整体式方法将物体放置融入场景级生成过程,牺牲了物体层面的细节;组合式方法通过将几何与布局解耦来保留物体保真度,但通常将布局参数化为稀疏且无界的位姿变量,这类变量难以学习,且在场景级监督稀缺的情况下泛化能力较差。我们提出了Mira-Scene,一个组合式3D场景重建框架,用稠密、有界的对应关系恢复取代稀疏的位姿回归。其核心是规范坐标图,这是一种像素对齐的场,将每个可见物体像素映射到物体有界规范空间中的表面坐标。当与来自单目几何估计的场景空间点云图配对时,CCM诱导出稠密的规范空间到场景空间的对应关系,并通过鲁棒的几何对齐从中恢复物体变换。由于CCM在有界的规范空间中运行,它提供了一个稳定的预测目标,可以从可扩展的物体级3D数据中进行训练,而无需场景级布局标注。Mira-Scene进一步引入了一个多模态扩散Transformer,联合生成物体几何与CCM,采用具有共享注意力机制和位置编码的模态特定专家流,以促进几何与布局的一致性。在室内、室外、合成及真实野外场景上的实验表明,Mira-Scene在布局精度上大幅超越强基线方法:在使用有限的开源训练数据的情况下,相较于SAM3D,其3D-IoU相对提升39.8%,2D-IoU相对提升16.5%。
cs.CV / 119 / 2609.23815

Confidence-Aware Teacher-Student Distillation for 3D Medical Segmentation

面向三维医学分割的置信度感知师生蒸馏方法
Triantafyllou, Georgios, Iakovidis, Dimitris K.
Abstract
Medical image segmentation models typically rely on large amounts of densely annotated volumetric data, limiting their scalability across tasks and imaging modalities. This work addresses the challenge of predicting entire 3D anatomical structures from extreme annotation sparsity. An annotation-efficient student-teacher framework is proposed for automatic 3D medical segmentation that requires only a set of point prompts on a single 2D slice per volume, as input. A foundation model serves as an offline teacher, utilizing the provided point prompts from the selected slice to full-volume pseudo-annotations alongside their corresponding spatial confidence scores prior to student training. To mitigate the error propagation of noisy pseudo-annotations, a task-specific 3D student network is trained using a confidence-aware optimization strategy. By leveraging the teacher's pre-computed confidence scores, this strategy explicitly excludes statically uncertain regions of the pseudo-annotations from the loss calculation, while simultaneously emphasizing regions with higher confidence. Evaluated on 3D cardiac MRI datasets, our framework outperforms state-of-the-art semi-supervised methods, improving segmentation performance by up to 43.6%. Furthermore, it drastically reduces the manual annotation burden to just a few positive point prompts per volume, while improving surface boundary precision by up to 14.7% over the teacher and successfully recovering up to 34.1% of the performance gap toward the fully supervised upper bound.
Chinese Translation
医学图像分割模型通常依赖大量密集标注的体数据,这限制了其在不同任务和成像模态上的可扩展性。本工作致力于解决在极端标注稀疏条件下预测完整三维解剖结构的挑战。我们提出了一种标注高效的师生(student-teacher)框架,用于自动化三维医学分割,其输入仅需每个体积在单个二维切片上的一组点提示(point prompts)。一个基础模型作为离线教师,在学生训练之前,利用所选切片上提供的点提示生成全卷伪标注及其对应的空间置信度评分。为抑制噪声伪标注的误差传播,采用置信度感知的优化策略训练一个任务专用的三维学生网络。该策略利用教师预先计算的置信度评分,在损失计算中显式排除伪标注中持续不确定的区域,同时强调置信度较高的区域。在三维心脏MRI数据集上的评估表明,我们的框架优于最先进的半监督方法,分割性能提升高达43.6%。此外,该方法将人工标注负担大幅降低至每个体积仅需少量正点提示,同时表面边界精度较教师模型提升高达14.7%,并成功弥补了相对于全监督上限的性能差距的34.1%。
cs.CV / 120 / 2609.23817

VISTA: Video-Injected Stylized Text-to-Animation

VISTA:视频注入的风格化文本到动画生成
Purkayastha, Monseej, Ghosh, Anindita, Slusallek, Philipp
Abstract
We present VISTA, a two-stage framework for generating stylized 3D human motion by fusing structural content from text prompts with expressive style from reference videos, without requiring jointly paired (text, video, stylized motion) triplets. A Dual-channel Autoencoder first maps motion sequences and video clips into a shared latent manifold. A masked autoregressive diffusion backbone then operates within this manifold, injecting video-derived style through a dedicated late-fusion Dual-AdaLN pathway while preserving text-conditioned content structure. A cross-batch unpaired training protocol with latent cycle consistency enables joint learning across separate semantically rich and stylistically diverse datasets. As a proof-of-concept for controllable animation synthesis, we validate VISTA on rendered motion-capture references: it achieves the highest style recognition accuracy among video-conditioned methods while preserving competitive content alignment, and its decomposed 3-way classifier-free guidance provides independent, user-controllable calibration of the content--style balance at inference time.
Chinese Translation
我们提出了VISTA,这是一个两阶段框架,通过将文本提示中的结构化内容与参考视频中的表现性风格相融合来生成风格化的3D人体运动,且无需联合配对的(文本、视频、风格化运动)三元组。首先,双通道自编码器(Dual-channel Autoencoder)将运动序列和视频片段映射到共享的潜在流形中。随后,掩码自回归扩散主干网络在该流形中运行,通过专门的后期融合Dual-AdaLN通路注入源自视频的风格,同时保留以文本为条件的内容结构。一种基于潜在循环一致性的跨批次非配对训练协议,实现了在语义丰富与风格多样的独立数据集上的联合学习。作为可控动画合成的概念验证,我们在渲染的动作捕捉参考数据上验证了VISTA:它在视频条件方法中取得了最高的风格识别准确率,同时保持了具有竞争力的内容对齐度;此外,其分解的三路无分类器引导(3-way classifier-free guidance)在推理时提供了独立且用户可控的内容-风格平衡校准。
cs.CV / 121 / 2609.23830

DFD-Lab: A Modular Audio-Visual Deepfake Detection Pipeline

DFD-Lab:一个模块化的音视频深度伪造检测流水线
Rybarczyk, Jan, Roszkowski, Mateusz, Komorowski, Jacek
Abstract
Comparing audio-visual deepfake detectors requires coordinating dataset adaptation, temporal input representation, model interfaces and experimental conditions. We present DFD-Lab, a modular pipeline that separates these responsibilities while supporting shared training and evaluation workflows. We integrate three implementations: Xception-based maximum-logit fusion, ResNet with temporal LSTM fusion, and our AVFF reimplementation. Experiments cover external testing, degradation-based training augmentation and evaluation-time corruption. On a filtered subset of Deepfake-Eval-2024, models trained on FakeAVCeleb attain baseline AUROC values of 0.504, 0.538 and 0.458. JPEG50 training augmentation raises these to 0.691, 0.605 and 0.570, respectively, while all three accuracies decrease. These results illustrate why training interventions, evaluation corruptions and metric-dependent outcomes should remain distinct within a common pipeline. The contribution is the integration of audio-visual processing, interchangeable detectors and configurable experimental workflows, supported by empirical case studies. The findings highlight the challenge of cross-dataset detection and the complementary information provided by ranking and classification metrics.
Chinese Translation
比较音视频深度伪造检测器需要协调数据集适配、时序输入表示、模型接口和实验条件。我们提出了DFD-Lab,这是一个模块化流水线,在支持共享训练和评估流程的同时,将这些职责相互分离。我们集成了三种实现:基于Xception的最大对数融合、结合时序LSTM融合的ResNet,以及我们对AVFF的重新实现。实验涵盖外部测试、基于退化的训练数据增强以及评估时的损坏测试。在Deepfake-Eval-2024的一个过滤子集上,在FakeAVCeleb上训练的模型取得的基础AUROC值分别为0.504、0.538和0.458。JPEG50训练增强将这些值分别提升至0.691、0.605和0.570,但三个模型的准确率均有所下降。这些结果说明了为什么在统一流水线中应将训练干预、评估损坏以及依赖指标的结论保持区分。本工作的贡献在于集成了音视频处理、可互换的检测器以及可配置的实验流程,并辅以实证案例研究。研究结果凸显了跨数据集检测的挑战,以及排序指标与分类指标所提供的互补信息。
cs.CV / 122 / 2609.23832

Comparative Performance and Parameter-Efficient Adaptation of DINOv2 for Active Trachoma Classification

DINOv2在活动性沙眼分类中的对比性能与参数高效自适应研究
Gebremedhin, Kibrom, Hailu, Hadush, Gebregziabher, Bruk, Hailu, Yordanos
Abstract
Automated grading of conjunctival photographs could reduce the cost and variability of trachoma prevalence surveys, but the relative value of modern pretrained visual representations, lightweight feature adaptation, and training-objective design has not been established under a common protocol. This study presents a controlled evaluation for binary classification of Trachomatous Inflammation-Follicular (TF) versus Normal using 1,546 images from the public UCSF/Lietman collection. Images are processed using the OPTED pipeline for zero-shot tarsal-conjunctiva segmentation, alignment, cropping, and standardization. We first compare six pretrained backbones using a common classification pipeline and then evaluate four lightweight adaptation mechanisms on DINOv2 ViT-B/14. Under stratified five-fold cross-validation, DINOv2 with Efficient Channel Attention (ECA) and focal-plus-center loss achieved 91.66 +/- 0.97% accuracy, 90.69 +/- 1.10% macro-F1, and 96.06 +/- 0.71% AUC. ECA introduces only five learnable parameters while matching the performance of substantially larger alternatives. Objective ablation further showed that ECA did not consistently improve plain DINOv2 across loss functions; the lowest-variance 91.66% accuracy was obtained with cross-entropy plus center loss. Overall, the fine-tuned DINOv2 representation provided most of the predictive performance, while ECA offered a highly parameter-efficient refinement whose effect depended on the training objective. The resulting workflow provides a reproducible benchmark for active trachoma image classification.
Chinese Translation
结膜照片的自动分级可以降低沙眼患病率调查的成本和变异性,但现代预训练视觉表征、轻量级特征自适应以及训练目标设计的相对价值尚未在统一协议下得到确立。本研究使用来自公开UCSF/Lietman数据集的1,546张图像,对沙眼滤泡性炎症(Trachomatous Inflammation-Follicular, TF)与正常结膜的二元分类进行了受控评估。图像采用OPTED流水线进行处理,包括零样本睑板结膜分割、对齐、裁剪和标准化。我们首先在统一的分类流水线下比较了六种预训练骨干网络,然后在DINOv2 ViT-B/14上评估了四种轻量级自适应机制。在分层五折交叉验证下,采用高效通道注意力(Efficient Channel Attention, ECA)和焦点损失加中心损失的DINOv2实现了91.66 ± 0.97%的准确率、90.69 ± 1.10%的宏平均F1分数和96.06 ± 0.71%的AUC。ECA仅引入五个可学习参数,却能匹配规模大得多的替代方案的性能。目标函数消融实验进一步表明,ECA在不同损失函数下并未持续提升原始DINOv2的性能;采用交叉熵加中心损失获得了方差最低的91.66%准确率。总体而言,微调后的DINOv2表征提供了大部分的预测性能,而ECA提供了一种高度参数高效的自适应改进,其效果依赖于训练目标。由此形成的工作流程为活动性沙眼图像分类提供了可复现的基准。
cs.CV / 123 / 2609.23881

MotionJEPA: Preventing Temporal Feature Collapse by Capturing Visual Changes in Latent Space

MotionJEPA:通过在潜在空间中捕捉视觉变化来防止时间特征坍缩
Karmann, Markus, Li, Shile, Internò, Christian, Andreis, Bruno, Klindt, David, Balestriero, Randall, Gu, Jindong, Torr, Philip, Zhang, Qi, Jiang, Peng-Tao, Zhang, Hao, Li, Bo, Urfalioglu, Onay
Abstract
Joint Embedding Predictive Architectures (JEPAs) are a promising paradigm for learning task-agnostic latent world models without visual reconstruction. However, standard JEPA training exhibits a strong inductive bias towards slow features, causing feature suppression and the collapse of latent representation. While inverse dynamics provides temporal anti-collapse, it relies on action labels and offers little incentive to embed general, unlabeled dynamics. We introduce Difference Image and Single image embedding Regularization (DISReg), a novel regularizer that builds on an inverse-dynamics-style module that predicts temporal difference image embeddings without any pixel reconstruction loss, encouraging balanced static and dynamic feature learning. DISReg consists of a static term that shapes the distribution of the image embedding and encourages slow features, and a dynamic term, which, unlike direct regularization on the embedding, imposes no constraint on the image embedding's shape or distribution and instead only incentivizes that dynamic features be present. By integrating this regularizer into a standard JEPA, we establish our new architecture, MotionJEPA. Latent probing demonstrates that MotionJEPA produces more complete representations than other methods, and our trajectory analysis shows it maintains geometrically simple latent embeddings with low curvature. We further show that MotionJEPA improves downstream planning success under static-background distractors across four environments.
Chinese Translation
联合嵌入预测架构(Joint Embedding Predictive Architectures, JEPAs)是一种无需视觉重建即可学习任务无关的潜在世界模型的有前景范式。然而,标准的JEPA训练对缓慢变化的特征存在强烈的归纳偏置,导致特征抑制和潜在表示的坍缩。虽然逆动力学方法能够提供时间维度上的抗坍缩能力,但它依赖于动作标签,且几乎无法激励模型嵌入通用的、无标签的动态信息。我们提出了差分图像与单图像嵌入正则化(Difference Image and Single image embedding Regularization, DISReg),这是一种新颖的正则化方法,它基于一个逆动力学风格的模块,在不使用任何像素重建损失的情况下预测时间差分图像的嵌入,从而鼓励均衡的静态与动态特征学习。DISReg包含一个静态项,用于塑造图像嵌入的分布并鼓励学习缓慢变化的特征;以及一个动态项,与直接对嵌入进行正则化不同,它不对图像嵌入的形状或分布施加任何约束,而是仅激励模型保留动态特征。通过将这一正则化器集成到标准的JEPA中,我们构建了新的架构MotionJEPA。潜在探测实验表明,MotionJEPA相比其他方法能产生更完整的表示,且我们的轨迹分析显示它保持了具有低曲率的几何上简洁的潜在嵌入。我们进一步证明,在四种环境中,面对静态背景干扰物时,MotionJEPA提高了下游规划任务的成功率。
cs.CV / 124 / 2609.23961

Colon3R: Cross-Domain 3D Reconstruction from Monocular Colonoscopic Video

Colon3R:基于单目结肠镜视频的跨域三维重建
Xing, Zhihao, Wang, Yingyu, Zhao, Liang, Huang, Shoudong
Abstract
Monocular colonoscopic 3D reconstruction is important for surgical robotic colonoscopy, but remains challenging due to weak texture, specular reflections, limited view overlap, and non-rigid tissue motion. Conventional multi-view 3D reconstruction methods rely on stable correspondences and approximate rigidity, which are often violated in colonoscopy. Existing endoscopic methods often rely on domain-specific supervision, whereas there are not enough in-vivo labeled data available to adapt geometry foundation models to clinical colonoscopy. We present Colon3R, a cross-domain semi-supervised framework built on pretrained VGGT that transfers coupled camera, depth, and pointmap geometry from labeled phantom and simulated data to unlabeled in-vivo colonoscopy without requiring target-domain geometric annotations. Unlike source-only fine-tuning, which learns only from phantom and simulated data, Colon3R directly exploits unlabeled in-vivo video through teacher-derived cross-view supervision. Our proposed hierarchical quasi-rigid reliability selects reliable supervision at the sequence, directed-pair, and pixel levels, while source-preserving adaptation retains the learned coupled geometry during target-domain adaptation. Extensive experiments demonstrate that our method achieves superior overall performance over state-of-the-art approaches in depth, pointmap, and camera pose estimation. Qualitative comparisons on real in-vivo colonoscopy further show substantially more complete and geometrically consistent reconstructions than competing methods under clinical domain shift. The code will be public available after the paper is accepted.
Chinese Translation
单目结肠镜三维重建对手术机器人结肠镜检查具有重要意义,但由于组织纹理微弱、镜面反射、视角重叠有限以及非刚性组织运动等原因,该任务仍极具挑战性。传统的多视图三维重建方法依赖于稳定的对应关系和近似刚性假设,而这些假设在结肠镜检查中常常无法满足。现有的内窥镜方法通常依赖特定领域的监督,然而在体的真实标注数据不足,难以将几何基础模型适配到临床结肠镜场景中。我们提出了Colon3R,一个基于预训练VGGT构建的跨域半监督框架,它将相机位姿、深度和点图(pointmap)之间的耦合几何信息从带标注的体模和仿真数据迁移到无标注的在体结肠镜数据,且无需目标域的几何标注。与仅从体模和仿真数据学习的源域微调不同,Colon3R通过教师模型导出的跨视图监督,直接利用无标注的在体视频。我们提出的分层准刚性可靠性机制在序列、有向图像对和像素三个层级上筛选可靠的监督信号,同时源域保持适配策略在目标域适配过程中保留已学习到的耦合几何。大量实验表明,我们的方法在深度估计、点图估计和相机位姿估计方面均取得了优于当前最先进方法的整体性能。在真实在体结肠镜数据上的定性比较进一步表明,在临床域偏移条件下,我们的方法相比竞争方法能够获得更完整、几何上更一致的重建结果。论文被接收后,代码将公开。
cs.CV / 125 / 2609.23967

Rethinking Diffusion Segmentation: When Does It Rely on Its Noisy State, and Does Diffusion Matter?

重新思考扩散模型分割:它何时依赖其噪声状态,扩散又是否重要?
Yang, Hengzhuo, Zeng, Yuming, Yang, Yuling
Abstract
Diffusion models are increasingly adapted from generation to conditional prediction, where a conditioning signal is combined with an evolving noisy representation of the target. In fully supervised segmentation, however, the conditioning image can already support direct target prediction, so endpoint performance alone establishes neither reliance on the added diffusion state nor a deterministic advantage over image-only prediction. For state reliance, we disrupt target-derived state content or correct image-state pairing during retraining of twelve published methods across three datasets, with ten matched seeds per setting. All 40 original-method comparisons whose evaluated-mask routes remained downstream of noised-quantity reconstruction exhibited state reliance, whereas all 30 comparisons with a segmentation-supervised bypass preserved reference performance. Rerouting five originally bypass-capable methods by forcing segmentation supervision through noise-to-mask reconstruction converted all 30 corresponding comparisons from preserved performance to state reliance. For deterministic utility, matched image-only counterparts achieved similar or better performance in 28 of 35 settings overall, including 16 of 20 whose native methods relied on both audited state properties. These results identify supervision path as a determinant of state reliance in the audited methods. Separately, matched image-only counterparts show that diffusion-specific computation often provides no deterministic endpoint advantage, including in methods that rely on the audited state properties. More generally, when conditioning already supports strong target prediction, diffusion-specific claims require additional evidence that the added state is used and that diffusion-specific computation improves the claimed capability beyond a matched condition-only counterpart.
Chinese Translation
扩散模型正越来越多地被从生成任务迁移到条件预测任务中,其中条件信号与目标的一个不断演化的噪声表示相结合。然而,在全监督分割任务中,条件图像本身已经足以支持对目标的直接预测,因此仅凭最终端点性能既无法确立对所添加扩散状态的依赖性,也无法确立相对于仅用图像预测的确定性优势。关于状态依赖性,我们在三个数据集上对十二种已发表的方法进行再训练,破坏源自目标的状态内容或纠正图像-状态的配对关系,每个设置使用十个匹配的随机种子。在所有评估掩码路径仍需经过噪声量重建的40组原始方法比较中,均表现出状态依赖性;而在所有包含分割监督旁路的30组比较中,均保持了基准性能。对于五种原本具备旁路能力的方法,通过强制分割监督经由噪声到掩码的重建路径,使相应的30组比较全部从性能保持转变为状态依赖。关于确定性效用,在全部35个设置中,匹配的仅用图像对照方法在28个设置中取得了相当或更好的性能,其中包括20个原生方法同时依赖两项受审状态属性的设置中的16个。这些结果表明,监督路径是受审方法中状态依赖性的决定因素。此外,匹配的仅用图像对照方法表明,扩散特有的计算往往不提供确定性的端点优势,即使在依赖受审状态属性的方法中也是如此。更一般而言,当条件信号已经能够支持较强的目标预测时,有关扩散特有优势的主张需要额外证据,证明所添加的状态确实被使用,且扩散特有的计算能够在匹配的仅条件对照方法之上改进所声称的能力。
cs.CV / 126 / 2609.23983

Evaluating the Generalization of Neuroimaging Foundation Models on African Brain MRI

评估神经影像基础模型在非洲脑部MRI上的泛化能力
Akinmuleya, Oluwatobi Iyanuoluwa, Akano, Olatokun Shamsudeen, Ankapong, Samuel Danquah, Lawal, Olamide, Musah, Toufiq
Abstract
Neuroimaging foundation models pretrained on large, predominantly western cohorts are increasingly proposed as general-purpose backbones for brain MRI analysis. Yet, their ability to generalize to underrepresented clinical populations remains largely untested. We evaluate four recent foundation models (BrainIAC, Neuro-JEPA, NeuroVFM, and Primus) on a three-way diagnostic classification task (Control, Dementia, Parkinson's disease) using a cohort of 88 subjects from a Nigerian clinical brain MRI dataset, across four modality configurations (T1w, T2w, T1w+T2w, FLAIR), and compare against an end-to-end trained ViT3D baseline. The frozen backbones collapse to majority-class predictions, while Neuro-JEPA on FLAIR shows modest but still limited discrimination. In contrast, the end-to-end trained ViT3D achieves higher accuracy and MCC on every task (up to 53.4% accuracy, MCC=0.27) and is the only model with non-trivial recall. Our findings suggest that these frozen neuroimaging foundation models are insufficient for fine-grained diagnostic classification in small, non-western clinical cohorts, motivating parameter-efficient adaptation and broader multi-site external validation for equitable deployment in global health settings.
Chinese Translation
在以西方人群为主的大规模队列上预训练的神经影像基础模型,正日益被提议作为脑部MRI分析的通用骨干网络。然而,其在代表性不足的临床人群中的泛化能力在很大程度上仍未得到验证。我们使用来自尼日利亚临床脑部MRI数据集的88名受试者队列,在四种模态配置(T1w、T2w、T1w+T2w、FLAIR)下,评估了四个近期的基础模型(BrainIAC、Neuro-JEPA、NeuroVFM和Primus)在三分类诊断任务(对照、痴呆、帕金森病)上的表现,并与端到端训练的ViT3D基线进行比较。冻结的骨干网络均退化为多数类预测,而Neuro-JEPA在FLAIR上仅表现出有限且微弱的判别能力。相比之下,端到端训练的ViT3D在所有任务上均取得了更高的准确率和MCC(最高准确率为53.4%,MCC=0.27),并且是唯一具有非平凡召回率的模型。我们的研究结果表明,这些冻结的神经影像基础模型不足以在小型、非西方临床队列中进行细粒度的诊断分类,这促使人们需要参数高效的自适应方法以及更广泛的多中心外部验证,以实现在全球健康环境中的公平部署。
cs.CV / 127 / 2609.24014

WebMRIQC: A Web-Based Implementation of MRIQC for Accessible MRI Image Quality Assessment in Resource-Constrained Settings

WebMRIQC:基于Web的MRIQC实现,用于资源受限环境下可及的MRI图像质量评估
Nkwam, Philip, Oladeji, Ifeoluwa, Zurakat-Aderibigbe, Sekinat, Cakmak, Jasmine, Aduluwa, Harrison, Raymond, Confidence, Mokua, Cliff, Zubair, Abdulrazaq, Champanda, Daniel, Olusuyi, Tolulope, Adewole, Maruf, Anazodo, Udunna
Abstract
Reliable quality control (QC) of magnetic resonance imaging (MRI) is essential for reliable diagnostic neuroimaging, yet standard manual assessment is subjective and time-consuming. MRIQC has established standardized automated extraction of image-quality metrics (IQMs), but its reliance on local computational imaging skills and capacity including high-performance computing, limits its adoption in resource-constrained settings (RCS). We present WebMRIQC (webmriqc.mailab.io), an open-source browser-based platform that wraps the validated MRIQC engine behind a zero-installation web interface. WebMRIQC automates the DICOM-to-BIDS conversion of de-identified MRI scans, executes the unmodified containerized MRIQC pipeline on a shared compute node governed by a fair-share job queue, and returns an interactive in-browser dashboard. The dashboard grounds every IQM in published quality thresholds, benchmarks each scan against the normative distribution of high-resource open datasets, and supports cross-site multicentre implementation of optimized scan protocols in RCS.We describe the system architecture and a validation framework establishing measurement equivalence between WebMRIQC and native MRIQC across thirteen IQMs on the BraTS-Africa and BraTS 2021 datasets. Preliminary results indicate strong agreement for contrast-, signal and noise-based metrics, demonstrating that web-based implementation lowers the barrier to standardized MRI QC and provides a foundation for harmonized, regionally adapted quality benchmarks across RCS imaging sites. The code is publicly available here https://github.com/CAMERA-MRI/WebMRIqc.
Chinese Translation
可靠的磁共振成像(MRI)质量控制(QC)对于可靠的诊断性神经影像至关重要,然而标准的人工评估方法主观且耗时。MRIQC已建立了标准化的图像质量指标(IQMs)自动提取方法,但其对本地计算影像技能和算力(包括高性能计算)的依赖,限制了其在资源受限环境(RCS)中的采用。我们提出了WebMRIQC(webmriqc.mailab.io),一个开源的基于浏览器的平台,它将经过验证的MRIQC引擎封装在零安装的Web界面之后。WebMRIQC自动完成去标识化MRI扫描的DICOM到BIDS转换,在由公平共享作业队列管理的共享计算节点上执行未经修改的容器化MRIQC流程,并返回一个浏览器内交互式仪表板。该仪表板将每个IQM与已发表的质量阈值相对应,将每个扫描与高资源开放数据集的规范分布进行基准比较,并支持在RCS中跨站点多中心实施优化的扫描协议。我们描述了系统架构和一个验证框架,该框架在BraTS-Africa和BraTS 2021数据集上建立了WebMRIQC与原生MRIQC在十三项IQMs上的测量等效性。初步结果表明,基于对比度、信号和噪声的指标具有高度一致性,证明基于Web的实现降低了标准化MRI质量控制的门槛,并为跨RCS影像站点建立统一且适应区域的质量基准奠定了基础。代码已在以下地址公开:https://github.com/CAMERA-MRI/WebMRIqc。
cs.CV / 128 / 2609.24026

InterHier: Learning Interconnected Hierarchical Semantics for Open-Vocabulary Object Detection

Kim, Yeong-Jin, Kim, Ho-Joong, Lee, Seong-Whan
Abstract
In this paper, we investigate the limitations of fixed, hand-crafted connectors in hierarchical semantic representations for open-vocabulary object detection. Existing methods establish semantic relationships between base categories and unseen novel categories by placing a fixed connector between adjacent super-/sub-categories. However, such fixed connectors may not optimally capture the relationships within a semantic hierarchy. To address this limitation, we propose interconnected hierarchical semantic representations (InterHier), which utilize a prepended learnable context to globally guide the interpretation of prompts containing hierarchical relationships. InterHier operates in two main stages. First, it constructs a hierarchy-aware prompt by integrating super-/sub-categories and prepending a learnable context. Second, it optimizes this learnable context to align visual region embeddings and textual embeddings. InterHier consistently improves performance over methods that rely on fixed connectors and can be seamlessly integrated into existing open-vocabulary object detection models. Experiments on open-vocabulary object detection benchmarks demonstrate that InterHier achieves competitive performance against state-of-the-art methods.
cs.CV / 129 / 2609.24031

Video-STLayout Pre-training

Video-STLayout 预训练
Jyothi, Akash Abdu, Mori, Greg
Abstract
In recent years, pre-training has become fundamental to learning effective video representations, enabling strong transfer to downstream tasks. A popular framework in pre-training involves aligning features of a video encoder with that of another modality, for example, language or audio. We introduce Video-STLayout pre-training, a novel strategy for obtaining rich video representations informed by spatio-temporal layout of object bounding boxes. Object layouts can easily be obtained by applying an off-the-shelf object detector on the video frames. Our method uses a contrastive loss to align video features with the layout features from a trained layout encoder. We show the effectiveness of our approach in the task of activity recognition in complex scenes.
Chinese Translation
近年来,预训练已成为学习有效视频表示的基础,使其能够强大地迁移到下游任务。一种流行的预训练框架是将视频编码器的特征与另一模态(例如语言或音频)的特征进行对齐。我们提出了 Video-STLayout 预训练,这是一种获取丰富视频表示的新策略,其利用了物体边界框的时空布局信息。通过对视频帧应用现成的物体检测器即可轻松获得物体布局。我们的方法使用对比损失将视频特征与经过训练的布局编码器所提取的布局特征进行对齐。我们在复杂场景中的行为识别任务上展示了该方法的有效性。
cs.CV / 130 / 2609.24049

U-PEN Mamba: Progressive Expansion with Selective State-Space Modeling for Efficient Retinal Vessel Segmentation

Reyes-Angulo, Abel A., Paheding, Sidike, Asari, Vijayan K., Alam, Mohammad, Devagiri, Jeevan
Abstract
Accurate retinal vessel segmentation is important for computer-aided ophthalmic analysis, yet thin vessels, low contrast, and severe foreground-background imbalance remain challenging for encoder-decoder networks. This paper presents U-PEN Mamba, a U-shaped retinal vessel segmentation architecture that couples progressive nonlinear feature expansion with selective state-space modeling. The proposed network enriches local vessel responses with progressive expansion, models long-range spatial dependencies through a Mamba Global Context (MGC) block with linear sequence complexity, and uses attention-based decoder fusion to recover fine vascular boundaries. We evaluate U-PEN Mamba on CHASE DB1 and DRIVE using a consistent patch-based preprocessing pipeline and compare it with convolutional, attention-based, transformer-based, and Mamba-based segmentation baselines. U-PEN Mamba obtains the best mean intersection over union among the compared methods, achieving 0.8394 on CHASE DB1 and 0.8221 on DRIVE, with Dice scores of 0.8187 and 0.8078, respectively, using 21.6M trainable parameters. Ablation studies show that the MGC block contributes the largest gain over the U-Net baseline, while projection dimension and state size provide practical accuracy-efficiency control. These results indicate that selective state-space modeling is a promising global-context mechanism for parameter-efficient retinal vessel segmentation. Code is available at: https://github.com/areyesan/UPEN_Mamba.
cs.CV / 131 / 2609.24058

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

基于文字系统感知混合专家的一体化多语言场景文本识别
Ye, Xingsong, Du, Yongkun, Zhang, Jiaxin, Li, Zhixian, Sun, Chong, Li, Chen, Lyu, Jing, Jin, Lianwen, Chen, Zhineng
Abstract
Multilingual scene text recognition (STR) remains challenging due to the scarcity of training data for most languages and the difficulty of serving diverse scripts within a single model. Existing solutions either deploy one recognizer per language, inflating cost and introducing error accumulation, or rely on massive vision-language models (VLMs) that are expensive and still inaccurate on many scripts. In this work, we pursue an all-in-one multilingual recognizer that is simpler than per-language experts, lighter than VLMs, and more accurate than both. First, we construct TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages. It provides balanced and sufficient supervision where real data is unavailable. Second, we propose ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture. It shares a single visual encoder and replaces the dense decoder with a sparse MoE block, which consists of an image-level router dispatches each image to the top-2 script-aligned experts and a shared expert absorbs cross-script knowledge. Extensive experiments on our assembled TextMuSS-Bench (10 scripts, 10,899 images) show that ScriptMoE achieves the highest accuracy of 82.06%, outperforming the strongest STR baseline by 1.31%. On the CC-OCR end-to-end multilingual task, replacing only the recognizer in PP-OCRv5 with ScriptMoE lifts F1 score from 65.71% to 80.89%, slightly surpassing the best VLM (80.73%) at a fraction of the parameter count.
Chinese Translation
多语言场景文本识别(STR)仍然具有挑战性,原因在于大多数语言的训练数据稀缺,以及难以在单一模型中服务多种文字系统。现有解决方案要么为每种语言部署一个识别器,导致成本膨胀并引入误差累积,要么依赖昂贵且在许多文字系统上仍不够准确的大规模视觉语言模型(VLM)。在本工作中,我们追求一种一体化多语言识别器,它比逐语言专家模型更简单,比VLM更轻量,并且比两者都更准确。首先,我们构建了TextMuSS-10M,一个覆盖10种文字系统和229种语言的大规模合成场景文本数据集,在真实数据不可得的情况下提供均衡且充分的监督。其次,我们提出ScriptMoE,一种文字系统感知的混合专家(MoE)架构。它共享单一视觉编码器,并用稀疏MoE模块替代密集解码器,该模块由一个将每张图像分发到前2个与文字系统对齐的专家的图像级路由器,以及一个吸收跨文字系统知识的共享专家组成。在我们构建的TextMuSS-Bench(10种文字系统,10,899张图像)上的大量实验表明,ScriptMoE达到了82.06%的最高准确率,超过最强的STR基线1.31%。在CC-OCR端到端多语言任务上,仅将PP-OCRv5中的识别器替换为ScriptMoE,就使F1分数从65.71%提升到80.89%,以极少的参数量略微超越最佳VLM(80.73%)。
cs.CV / 132 / 2609.24064

Vision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea

Vision Transformer与卷积神经网络在物种丰富、数据匮乏植物区系细粒度兰科属级识别中的比较:基于新几内亚兰科植物的受控基准测试
Saputra, Reza, Apriyanti, Diah Harnoni, Schuiteman, André, Metzger, Kurt, Field, Ashley, Nargar, Katharina, Edwards, William
Abstract
New Guinea is the world's richest island flora (~2,856 orchid species), yet most species are represented by only a handful of photographs, far fewer than direct species-level classification requires. Methods for fine-grained identification in such species-rich, data-poor floras are needed, and it remains unclear which backbone architecture and pretraining strategy best support them. We built a two-stage system that first predicts the genus of a query photograph, then retrieves visually similar reference images of candidate species using FAISS. We compared four pretrained backbones -- two Vision Transformers (ViTs; DINOv2, BioCLIP 2) and two CNNs (ConvNeXt V2-L, EfficientNetV2-L) -- fine-tuned under an identical protocol on a fixed, species-stratified partition of 16,701 photographs spanning 120 genera and 1,350 species, assessing accuracy, calibration, error structure, species retrieval, and open-set detection of novel genera. DINOv2 attained the best genus performance (macro top-1 66.9%, 95% CI 63.7-70.6; global top-1 88.9%); both ViTs outranked both CNNs, and general-purpose self-supervised pretraining (DINOv2) outperformed domain-matched biological pretraining (BioCLIP 2) by 7.1 points of macro top-1. Errors concentrated on two abundant genera acting as error attractors. DINOv2 embeddings achieved species Recall@5 of 86.6% and genus Recall@5 of 98.7%; temperature scaling reduced every backbone's Expected Calibration Error to about 0.03; and a distance-based open-set gate flagged unseen genera (mean AUROC 0.958). A self-supervised Vision-Transformer backbone combined with embedding retrieval is an effective, deployable strategy for fine-grained identification in species-rich, data-poor floras. The system is released as an open web application (the New Guinea Orchid Identifier), offering a practical template for other hyperdiverse, under-documented taxa.
Chinese Translation
新几内亚拥有世界上最丰富的岛屿植物区系(约2,856种兰花),然而大多数物种仅有少量照片可供使用,远低于直接的物种级分类所需。此类物种丰富、数据匮乏的植物区系需要细粒度识别方法,但何种骨干架构与预训练策略最能支持这些方法仍不清楚。我们构建了一个两阶段系统:首先预测查询照片的属,然后利用FAISS检索候选物种的视觉相似参考图像。我们在固定的按物种分层数据划分(共16,701张照片,涵盖120属1,350种)上,采用相同的微调协议,比较了四种预训练骨干网络——两种Vision Transformer(ViT:DINOv2、BioCLIP 2)和两种CNN(ConvNeXt V2-L、EfficientNetV2-L)——并评估其准确性、校准性、错误结构、物种检索以及对未知属的开放集检测能力。DINOv2取得了最佳的属级性能(宏平均top-1为66.9%,95%置信区间为63.7–70.6;全局top-1为88.9%);两种ViT均优于两种CNN,且通用自监督预训练(DINOv2)在宏平均top-1上比领域匹配的生物预训练(BioCLIP 2)高出7.1个百分点。错误集中于两个充当“错误吸引子”的优势属。DINOv2嵌入在物种检索上达到Recall@5为86.6%,属级检索Recall@5为98.7%;温度缩放将所有骨干网络的期望校准误差(ECE)降低至约0.03;基于距离的开放集门控能有效识别未见过的属(平均AUROC为0.958)。自监督Vision Transformer骨干网络结合嵌入检索,是一种有效且可部署的策略,适用于物种丰富、数据匮乏植物区系的细粒度识别。该系统已作为开放网络应用发布(新几内亚兰花识别器,New Guinea Orchid Identifier),为其他超高多样性、文献记录不足的分类群提供了实用的模板。
cs.CV / 133 / 2609.24071

Monitorable Chart Reasoning Agents via Verifiable Process Rewards

Sinha, Sanchit, Frunza, Oana, Rasul, Kashif, Zhang, Aidong
Abstract
Chart reasoning agents are increasingly used to extract actionable insights in critical domains, achieving state-of-the-art performance on multiple benchmarks. Yet, high benchmark accuracy alone is insufficient for deployment, where stakeholders must be able to audit and verify how a model reaches its answer. Existing LVLM-based chart agents produce either answer-only predictions or free-form rationales that are hard to verify, obscuring whether an error arose from misreading the chart, extracting a wrong value, or miscomputing. We propose Chart-RVR, a reinforcement learning framework for training monitorable chart agents with verifiable process rewards. Chart-RVR decomposes chart reasoning into three auditable blocks: Structure, identifying the chart type; Evidence, reconstructing the underlying data table in JSON; and Derivation, exposing the stepwise trace that computes the answer. Across six in-domain and out-of-domain benchmarks, Chart-RVR attains state-of-the-art accuracy among comparable-sized LVLMs. Beyond accuracy, we assess monitorability using a triangulated protocol that combines ground-truth surrogate metrics, an oracle information-gain measure, and an LLM-as-auditor scoring Process Verifiability and Evidence Localization, showing that Chart-RVR yields rationales that are markedly more verifiable and evidence-grounded than those from CoT prompting, SFT, and existing chart-specific baselines.
cs.CV / 134 / 2609.24075

A paired synthetic construction-site image dataset for robust computer vision under adverse conditions

Duong, Viet Huy, Xiong, Ruoxin, Forhad, Md Abdullah Al, Shi, Weishi
Abstract
Computer-vision systems used for construction monitoring can degrade under adverse environmental and visual conditions, yet such conditions remain underrepresented in existing construction image datasets. We present ConSynth-X, a paired synthetic construction-site image dataset containing 34,199 images derived from 3,109 real-world source scenes. The dataset comprises 11 condition-specific subsets spanning precipitation, fog, nighttime illumination, adverse weather at night, and small-object or long-distance views. Each synthetic image is linked to its corresponding source scene, enabling controlled comparison across environmental and visual conditions. ConSynth-X includes source-derived annotations, generation metadata, provenance information, and image-quality indicators, supporting object detection, image captioning, visual grounding, and visual question answering. Technical validation evaluates source-synthetic fidelity and alignment with real adverse-condition imagery using embedding-based similarity and distributional analyses. The dataset provides a structured resource for evaluating and improving the robustness of construction vision and vision-language models under challenging field conditions.
cs.CV / 135 / 2609.24088

Bridging Reconstruction and Generation: A Latent Distribution Perspective on Evaluation and Improvement

连接重建与生成:基于潜在分布视角的评估与改进
Fang, Xianghong, Shu, Wenjie, Xu, Tongda, Mou, Wenlong, Kong, Dehan, Rudner, Tim G. J.
Abstract
In latent generative models, reconstruction quality is often assumed to correlate with generative performance. However, reconstruction FID (rFID) can exhibit weak or even negative correlation with generation FID (gFID). We attribute this discrepancy to a latent distribution mismatch: reconstruction evaluates the decoder on encoder-induced latents, whereas generation uses the same decoder on latents produced by the generative model. To characterize this shift, we introduce generation-aware reconstruction (GAR), which constructs a continuous trajectory from standard reconstruction toward generation by perturbing encoder latents with noise and denoising them through the generative model before decoding. GAR probes the decoder behavior along this trajectory, making the transition from encoder to generation-time latent distributions observable and diagnosable. The resulting trajectory-based diagnostic, GAR-FID, exhibits strong empirical correlation with gFID across diverse tokenizers and scales. Importantly, intermediate GAR latents become more generation-aware while preserving correspondence with their source images, thereby retaining paired supervision that is absent for fully generated latents. This correspondence enables decoder adaptation on intermediate GAR latents, consistently improving generative quality across model scales. Overall, latent distribution mismatch provides a useful perspective for evaluating and improving latent generative models.
Chinese Translation
在潜在生成模型中,重建质量通常被认为与生成性能相关。然而,重建FID(rFID)与生成FID(gFID)之间可能表现出较弱甚至负相关。我们将这种差异归因于潜在分布的不匹配:重建是在编码器产生的潜在表示上评估解码器,而生成则是生成模型所产生的潜在表示上使用同一解码器。为了刻画这种偏移,我们提出了生成感知重建(generation-aware reconstruction, GAR),该方法通过对编码器潜在表示施加噪声扰动,并经生成模型去噪后再解码,从而构建一条从标准重建通往生成的连续轨迹。GAR沿着该轨迹探测解码器的行为,使从编码器到生成时潜在分布的转变变得可观察、可诊断。由此得到的基于轨迹的诊断指标GAR-FID,在多种分词器(tokenizer)和不同规模下均与gFID表现出很强的经验相关性。重要的是,中间GAR潜在表示在变得更具生成感知性的同时,仍与其源图像保持对应关系,从而保留了完全生成的潜在表示所缺失的成对监督。这种对应关系使得在中间GAR潜在表示上进行解码器适配成为可能,并在不同模型规模下均能持续提升生成质量。总体而言,潜在分布不匹配为评估和改进潜在生成模型提供了一个有用的视角。
cs.CV / 136 / 2609.24095

HDND: Hierarchical Dynamic Neural Decoding for Multilingual Word/Character Retrieval from Non-Invasive Brain Recordings

HDND:用于从非侵入性脑记录中进行多语言词/字检索的分层动态神经解码方法
Li, Yueyang, Chen, Shuran, Siok, Wai Ting, Wang, Nizhuan
Abstract
While deep learning has enabled language decoding from intracranial brain recordings, extending this capability to non-invasive recordings remains an unresolved challenge. Decoding individual words from non-invasive brain recordings is particularly difficult, as word-level neural evidence is weak, temporally distributed, and entangled with acoustic, lexical, and semantic structure. Existing retrieval pipelines often collapse these factors into a single representation, potentially discarding information available at intermediate temporal scales. Here, we introduce Hierarchical Dynamic Neural Decoding (HDND), a hierarchical dynamic decoding framework that treats word decoding as structured refinement rather than flat label retrieval. HDND combines intermediate neural representations, contextual semantic predictions, and, for selected reading conditions, an auxiliary character-form objective. We evaluate HDND across seven electroencephalography (EEG) and magnetoencephalography (MEG) datasets spanning English, Dutch, Mandarin, and Cantonese listening, reading, and reading-aloud conditions. Across the nine-condition word-retrieval benchmark, the proposed HDND yields a higher participant-averaged balanced Top-10 point estimate than the matched contextual word-decoding baseline in every condition and achieves the highest mean among all compared methods in eight of nine conditions. Across the same nine matched conditions, HDND also yields higher token-micro and pooled word-macro Top-10 point estimates in every setting. Sentence retrieval favors HDND in eight of nine conditions, while auditory speech-segment retrieval is mixed across the six listening conditions. These results show that hierarchical residual refinement can improve multilingual word retrieval from heterogeneous non-invasive brain recordings.
Chinese Translation
尽管深度学习已使从颅内脑记录中进行语言解码成为可能,但将这一能力扩展到非侵入性脑记录仍然是一个未解决的挑战。从非侵入性脑记录中解码单个词语尤为困难,因为词级神经证据微弱、在时间上分布广泛,并与声学、词汇和语义结构相互纠缠。现有的检索流程往往将这些因素压缩为单一表示,可能丢失中间时间尺度上可用的信息。本文提出分层动态神经解码框架HDND,它将词解码视为结构化的逐步精化过程,而非扁平的标签检索。HDND结合了中间神经表示、上下文语义预测,以及在特定阅读条件下的辅助字形单元目标。我们在七个脑电图(EEG)和脑磁图(MEG)数据集上评估了HDND,涵盖英语、荷兰语、普通话和粤语的听、阅读及朗读条件。在九个条件的词检索基准测试中,所提出的HDND在每个条件下的参与者平均平衡Top-10点估计均高于匹配的上下文词解码基线,并在九个条件中的八个取得了所有对比方法中的最高均值。在相同的九个匹配条件下,HDND在每个设置中的token级微平均和词级宏平均Top-10点估计也均更高。句子检索在九个条件中的八个倾向于HDND,而听觉语音片段检索在六个听觉条件下的结果则好坏不一。这些结果表明,分层残差精化可以改善来自异构非侵入性脑记录的多语言词检索。
cs.CV / 137 / 2609.24098

A$^2$Safe: Counterfactual Evidence-Aligned Adaptive Agent Collaboration for Safe and Effective Visual Question Answering

A$^2$Safe:面向安全且有效视觉问答的反事实证据对齐自适应智能体协作
Xu, Quanxing, Zhou, Ling, Zhong, Xian, Tian, Jinyu, Huang, Xiaohua, Huang, Rubing, Lin, Chia-Wen
Abstract
Visual Question Answering (VQA) with Multimodal Large Language Models (MLLMs) requires not only producing safe and effective responses, but also grounding safety decisions in the multimodal evidence that determines risk. Recent safety-alignment methods improve refusal behavior and contextual risk awareness, yet correct safety outcomes may still rely on superficial textual or visual correlations, particularly when risk emerges from interactions between individually benign image and question content. To address this issue, we propose A$^2$Safe, a counterfactual evidence-aligned adaptive agent collaboration framework for safe and effective VQA. A$^2$Safe organizes localized visual observations, textual intent, and cross-modal risk relations through a Grounded Safety Evidence Board, making the basis of safety decisions explicit. Counterfactual safety evidence alignment enforces invariance to safety-irrelevant changes while requiring appropriate safety-state and response-mode transitions when risk-critical evidence is minimally altered. The resulting evidence state further supports adaptive collaboration, enabling direct answering when grounded evidence is sufficient and invoking policy critique and response revision when evidence is risky, uncertain, or conflicting. Under complementary safety-critical and general VQA protocols, A$^2$Safe achieves a 95.72 SIUO safety score, reduces the benign refusal rate on MOSSBench to 14.67%, and maintains an average general VQA score of 78.34 with 27.8% token overhead. These results support counterfactual evidence-aligned adaptive collaboration for safe and effective multimodal question answering.
Chinese Translation
基于多模态大语言模型(MLLM)的视觉问答(VQA)不仅需要生成安全且有效的回答,还需要将安全决策建立于决定风险的多模态证据之上。近期的安全对齐方法提升了拒答行为和情境风险感知能力,然而正确的安全结果仍可能依赖于表层的文本或视觉关联,尤其是当风险源于本身无害的图像内容与问题内容之间的交互时。为解决这一问题,我们提出了A$^2$Safe,一个面向安全且有效视觉问答的反事实证据对齐自适应智能体协作框架。A$^2$Safe通过一个有据安全证据板(Grounded Safety Evidence Board)组织局部视觉观察、文本意图以及跨模态风险关系,使安全决策的依据显式化。反事实安全证据对齐要求对安全无关的变化保持不变性,同时在风险关键证据被最小改动时,要求安全状态与响应模式进行相应的转换。由此得到的证据状态进一步支持自适应协作:当有据证据充分时直接作答;当证据存在风险、不确定或相互冲突时,则调用策略批评与响应修正机制。在互补的安全关键型与通用VQA评测协议下,A$^2$Safe在SIUO上取得95.72的安全分数,将MOSSBench上的良性拒答率降至14.67%,并在27.8%的额外令牌开销下保持78.34的通用VQA平均分数。这些结果支持了反事实证据对齐的自适应协作在安全且有效的多模态问答中的作用。
cs.CV / 138 / 2609.24109

Adaptive Cortically Constrained EEG-Vision Alignment for Zero-Shot Brain-to-Image Retrieval

面向零样本脑到图像检索的自适应皮层约束脑电-视觉对齐方法
Wang, Ye, Ren, Haokun, Wu, Wei, Wang, Guoyin, Yu, Zhuliang, Yu, Hong, Liu, Ke
Abstract
Zero-shot brain-to-image retrieval requires robust alignment between noisy EEG responses and visual representations. Existing EEG-vision alignment methods often operate in sensor space and apply fixed visual supervision to all responses, ignoring both spatial mixing in scalp EEG and response-wise variability in alignment reliability. We propose an adaptive cortically constrained EEG-vision alignment method for zero-shot brain-to-image retrieval. The method reconstructs EEG responses into predefined ROI-level source-pattern representations and encodes them with a Neuro-ROI Attention Encoder. To handle response-wise variability, we introduce an evidence-based adaptive visual supervision strategy that weights detail-controlled visual targets using model-based alignment evidence. On THINGS-EEG, the proposed method achieves strong 200-way zero-shot retrieval performance, with ROI-level attribution providing post hoc interpretability of the learned source-pattern representations. These results show that cortically constrained representation learning and adaptive supervision can jointly support EEG-vision alignment for zero-shot brain-to-image retrieval.
Chinese Translation
零样本脑到图像检索要求噪声脑电(EEG)响应与视觉表示之间实现稳健的对齐。现有的脑电-视觉对齐方法通常在传感器空间中操作,并对所有响应施加固定的视觉监督,既忽略了头皮脑电的空间混叠问题,也忽略了对齐可靠性在响应层面的差异性。我们提出了一种用于零样本脑到图像检索的自适应皮层约束脑电-视觉对齐方法。该方法将脑电响应重建为预定义的感兴趣区(ROI)级源模式表示,并使用神经ROI注意力编码器(Neuro-ROI Attention Encoder)对其进行编码。为处理响应层面的差异性,我们引入了一种基于证据的自适应视觉监督策略,利用基于模型的对齐证据对细节控制的视觉目标进行加权。在THINGS-EEG数据集上,所提方法取得了较强的200类零样本检索性能,且ROI级归因分析为学习到的源模式表示提供了事后可解释性。这些结果表明,皮层约束的表示学习与自适应监督能够共同支持面向零样本脑到图像检索的脑电-视觉对齐。
cs.CV / 139 / 2609.24116

Patch-to-Global: Random Patch Diffusion for Globally Consistent Megapixel Artifact Inpainting in Whole Slide Images

Lee, Hyeseong, Kim, Eunsu, Bappy, D M, Kim, Ho Heon, Lee, Youngsuk, Chun, Se Young, Choi, Jang-Hwan, Lee, Sung Hak, Ahn, Sangjeong
Abstract
Although deep learning has advanced Whole Slide Image (WSI) Analysis, tissue artifacts like bubbles and folds often cause silent failures by concealing essential morphology. Current pathology image restoration methods are mostly restricted to small patches, struggling to maintain global structural coherence at a megapixel scale. We introduce RestorePath, a framework for globally consistent megapixel scale inpainting that reconstructs diagnostic structures in histological image to prevent incorrect high-confidence predictions and lower error rates. Our model utilizes a Latent Diffusion Model (LDM) conditioned on Pathology Foundation Model (PFM) embeddings, integrating Large Kernel Attention (LKA) to manage long-range dependencies during random patch diffusion. Enhanced by Distance-Weighted Interpolation (DWI) and an Adaptive Guidance Scale (AGS), RestorePath ensures structural consistency and fidelity by modulating information from surrounding patches. Evaluations across TCGA-BRCA, BACH, and Camelyon16 datasets for images ranging from 512 to 4608 pixels demonstrate state-of-the-art performance in maintaining histological consistency. RestorePath significantly improves downstream Computational Pathology (CP) tasks, outperforming both raw artifact images and the conventional Detect-and-Discard (D&D) approach. The code is available at https://github.com/PathfinderLab/RestorePath
cs.CV / 140 / 2609.24125

Positive Pair Geometry Matters: Optimal Transport for Contrastive Learning of Visual Representations

Nanda, Akshit, Ahmad, Shahzad, Padhy, Ram Prasad
Abstract
Contrastive self-supervised learning has achieved strong performance by learning representations from multiple augmented views of the same image. However, most existing methods construct positive pairs using independently sampled stochastic augmentations, which may alter semantic content and ignore the intrinsic geometry of the data distribution. In this work, we propose OTCLR, an optimal transport-aware framework for contrastive learning representations that generates geometry-consistent positive samples. Instead of directly contrasting two randomly augmented views, we construct intermediate views between the original image and its augmented variants through entropic optimal-transport displacement interpolation. These transport-interpolated samples serve as positive views that better preserve image structure while explicitly modeling spatial distributional geometry. To further promote smooth representation learning, we evaluate auxiliary Sinkhorn regularization terms that encourage transport-interpolated views to remain consistent with their endpoint images. The proposed method can be incorporated into standard contrastive learning pipelines without modifying the encoder architecture. Experiments on multiple benchmark datasets show that our approach improves representation quality and transfer learning performance compared with conventional augmentation-based contrastive learning baselines.
cs.CV / 141 / 2609.24127

Action-Slot: Structured Action-Centric Representation Learning for Multi-Agent Atomic Activity Understanding

Action-Slot:面向多智能体原子级活动理解的结构化动作中心表示学习
Chang, Yu-Ho, Kung, Chi-Hsi, Tsai, Yi-Hsuan, Chen, Yi-Ting
Abstract
Atomic activity understanding aims to recognize and localize structured traffic behaviors that jointly encode motion patterns and their grounding in road topology. Unlike conventional action recognition, atomic activities are multi-agent, multi-label, and topology-aware: multiple activities co-occur while many agents remain inactive. We introduce Action-Slot, a structured action-centric representation learning framework. Slot attention is widely used for object-centric decomposition, but its permutation-invariant design and object-level inductive bias are misaligned with atomic activity semantics. We reformulate slot learning as structured activity decomposition through three designs: (1) category-aligned action slots that anchor slots to predefined activity categories, (2) parallel spatio-temporal slot updating for holistic video-level reasoning, and (3) background and negative-slot regularization that enforces competition between foreground activities and irrelevant regions. Together these establish an activity-centric inductive bias that disentangles concurrent and asynchronous activities directly from raw video. Beyond recognition, the learned representations encode transferable spatio-temporal grounding signals. We further propose an attention-difference-based pseudo mask selection framework that suppresses false positives by measuring attention changes before and after candidate region removal, enabling weakly supervised localization without dense annotations. To support systematic evaluation, we introduce TACO, a balanced synthetic dataset with full atomic activity coverage and pixel-level annotations. Experiments on OATS, TACO, and annotated nuScenes show superior recognition, strong sim-to-real transfer, and state-of-the-art weakly supervised localization.
Chinese Translation
原子级活动理解旨在识别和定位结构化交通行为,这些行为共同编码了运动模式及其在道路拓扑中的落地依据。与传统的动作识别不同,原子级活动具有多智能体、多标签和拓扑感知的特点:多个活动同时发生,而许多智能体则保持静止。我们提出了Action-Slot,一个结构化的动作中心表示学习框架。槽注意力(slot attention)被广泛用于以对象为中心的分解,但其置换不变的设计和对象级归纳偏置与原子级活动语义不相契合。我们将槽学习重新表述为结构化活动分解,通过三项设计实现:(1) 类别对齐的动作槽,将槽锚定到预定义的活动类别;(2) 并行时空槽更新,用于整体的视频级推理;(3) 背景与负槽正则化,强制前景活动与无关区域之间进行竞争。这些设计共同建立了以活动为中心的归纳偏置,能够直接从原始视频中解耦并发与异步活动。除识别之外,学习到的表示还编码了可迁移的时空落地信号。我们进一步提出了一种基于注意力差异的伪掩码选择框架,通过衡量移除候选区域前后的注意力变化来抑制假阳性,从而实现无需密集标注的弱监督定位。为支持系统性评估,我们引入了TACO,一个具有完整原子级活动覆盖和像素级标注的平衡合成数据集。在OATS、TACO以及标注的nuScenes上的实验表明,该方法在识别性能、强大的模拟到现实迁移能力以及最先进的弱监督定位方面均表现优异。
cs.CV / 142 / 2609.24136

The Visual Target Matters: Learning across the Visual Hierarchy for Brain-to-Image Retrieval

视觉目标至关重要:跨视觉层次学习以实现脑到图像检索
Wang, Ye, Ren, HaoKun, Yu, Hong, Li, Ruirui, Li, Xiao, Liu, Ke, Wu, Wei
Abstract
Brain-to-image retrieval seeks to identify the visual stimulus that elicited a non-invasive neural response. Candidate images are typically represented by pretrained vision models, whose internal representations vary in abstraction across depth. Existing methods usually train the neural encoder to recover a fixed final-layer visual target. Under this formulation, the visual hierarchy is reduced to a single prescribed endpoint, preventing representations at other depths from directly shaping the visual target. This limitation motivates learning how information across visual depths should contribute to the retrieval target. To this end, we introduce NeuroGlyph, which learns a trial-independent visual target from multiple depths of a frozen visual backbone. NeuroGlyph decomposes the target into factor-specific subspaces. Each subspace learns an image-conditioned allocation over visual depth. The resulting subspaces are fused into a single embedding for retrieval. Across THINGS-EEG and THINGS-MEG, NeuroGlyph outperforms final-layer supervision in all controlled comparisons. It also surpasses the post hoc best fixed-layer oracle in three of four comparisons. Parameter-matched ablations support both factorized target construction and image-conditioned depth allocation. Under comparable 200-way retrieval protocols, NeuroGlyph achieves the strongest system-level performance in six of eight reported metrics. These results support learning retrieval targets across the visual hierarchy rather than prescribing one visual depth.
Chinese Translation
脑到图像检索旨在识别引发非侵入性神经响应的视觉刺激。候选图像通常由预训练视觉模型表示,其内部表征的抽象程度随网络深度而变化。现有方法通常训练神经编码器以恢复固定的最后一层视觉目标。在这种表述下,视觉层次被简化为单一预设的终点,使得其他深度的表征无法直接塑造视觉目标。这一局限促使我们思考如何学习跨视觉深度的信息应如何贡献于检索目标。为此,我们提出了NeuroGlyph,它从冻结视觉骨干网络的多个深度学习一个与试次无关的视觉目标。NeuroGlyph将该目标分解为因子特定的子空间,每个子空间学习一种基于图像的视觉深度分配。所得子空间被融合为用于检索的单一嵌入。在THINGS-EEG和THINGS-MEG数据集上,NeuroGlyph在所有受控对比中均优于最后一层监督,并在四次对比中的三次超越了事后最佳固定层oracle。参数匹配的消融实验支持因子化目标构建和基于图像的深度分配这两项设计。在可比的200类检索协议下,NeuroGlyph在八项报告指标中的六项上取得了最强的系统级性能。这些结果表明,应跨视觉层次学习检索目标,而非预设单一视觉深度。
cs.CV / 143 / 2609.24151

STAR: Scene- and Task-Aware 4D Radar Preprocessing Towards End-to-End Cognitive Radar

STAR:面向端到端认知雷达的场景与任务感知4D雷达预处理方法
Song, Seung-Hyun, Paek, Dong-Hee, Kong, Seung-Hyun
Abstract
Four-dimensional (4D) Radar has emerged as a key sensor for environmental perception, providing range, azimuth, elevation, and Doppler measurements while remaining robust to illumination changes and adverse weather conditions. However, conventional Radar preprocessing methods, such as constant false alarm rate (CFAR) detection, select measurements primarily based on signal-level criteria and may therefore discard information valuable for downstream perception during point cloud generation. In addition, existing 4D Radar perception pipelines typically optimize Radar data processing and downstream perception independently, preventing task objectives from directly guiding the preprocessing stage. To address these limitations, we propose a Scene- and Task-Aware Radar (STAR) Preprocessor together with an end-to-end training framework. The STAR Preprocessor incorporates scene context and downstream task objectives to generate task-relevant Radar points, enabling the Radar representation to be optimized directly for perception. On the K-Radar benchmark, the proposed method achieves 74.3 AP, outperforming the previous state of the art by 5.6 AP points. Furthermore, applying the task-relevant points generated by STAR to various existing 3D detectors improves detection performance in most evaluation settings and yields an overall positive average gain over point clouds produced by conventional preprocessing.
Chinese Translation
四维(4D)雷达已成为环境感知的关键传感器,它能够提供距离、方位角、俯仰角和多普勒测量信息,同时对光照变化和恶劣天气条件保持较强的鲁棒性。然而,传统的雷达预处理方法(如恒虚警率(CFAR)检测)主要基于信号层面的准则来筛选测量数据,因此在点云生成过程中可能会丢弃对下游感知有价值的信息。此外,现有的4D雷达感知流程通常将雷达数据处理与下游感知分开独立优化,导致任务目标无法直接指导预处理阶段。为解决上述局限性,我们提出了一种场景与任务感知雷达(STAR)预处理器以及一个端到端训练框架。STAR预处理器结合场景上下文和下游任务目标,生成与任务相关的雷达点,使雷达表征能够直接针对感知任务进行优化。在K-Radar基准上,所提出的方法达到74.3 AP,超越此前最先进方法5.6个AP点。此外,将STAR生成的任务相关点应用于多种现有3D检测器,在大多数评估设置中均提升了检测性能,并且相比传统预处理生成的点云取得了整体正向的平均增益。
cs.CV / 144 / 2609.24158

Relightable 3D Avatar Reconstruction with Semantic-Adaptive Motion-Illumination Responses

Zhao, Jiankuo, Zhu, Xiangyu, Li, Jijie, Wang, Baiqin, Chen, Shukai, Lei, Zhen
Abstract
Reconstructing expressive and relightable 3D head avatars from monocular videos remains challenging in computer vision, as it requires accurate modeling of both non-rigid facial motion and illumination-dependent appearance. Existing Gaussian avatar methods commonly rely on globally coupled representations, in which Gaussian primitives share a unified motion or illumination response model. Such uniform modeling neglects the distinct motion patterns and material/reflectance properties of different facial semantic regions, thereby limiting fine-grained animation accuracy and reducing relighting plausibility. To address this limitation, we propose SAMIRA, a 3D Gaussian avatar framework for semantic-adaptive motion-illumination response modeling. For motion response modeling, the Semantic-Adaptive Motion Response module rasterizes current-to-reference mesh displacements into a topology-consistent UV space and leverages facial semantics to route displacement features through semantic-specific modulators, predicting localized Gaussian geometric residuals beyond coarse mesh binding. For illumination response modeling, the Semantic-Adaptive Illumination Response module learns compact diffuse and specular response factors for each facial region, allowing Gaussians in different regions to adapt their illumination responses to novel environment lighting. These response factors are incorporated into deferred physically based shading, providing a lightweight approximation of semantic-dependent illumination effects. Extensive experiments on self-reenactment, cross-reenactment, and relighting demonstrate that SAMIRA improves both fine-grained expression reconstruction and relighting realism over existing methods.
cs.CV / 145 / 2609.24170

An Unexpected Robot Policy: Early Evaluations of GPT-6 Astra on RoboDojo and Beyond

一个意外的机器人策略:GPT-6 Astra 在 RoboDojo 及其他基准上的早期评估
Zhang, Wenbo, Wang, Kaixuan, Ouyang, Yutao, Huang, Xiaoyu, Li, Liyang, Su, Kailun, Jin, Weiyang, Chai, Wenhao, Liang, Haotian, Dou, Zhiyang, Chen, Yue, Chen, Tianxing
Abstract
Embodied AI systems are often organized into System 1 and System 2. System 1 is typically a pretrained policy that generates actions at high frequency, whereas System 2 is often instantiated as a vision-enabled language model for high-level planning. We ask whether a large language model (LLM) can act as the policy for robot manipulation without task-specific finetuning. We call this setting LLM as policy. We evaluate three LLMs on all 42 RoboDojo tasks and compare their scores with 40 public policies. Astra and GPT-5.5 use the official 50-episode-per-task protocol; DeepSeek-Flash uses 10 episodes per task. GPT-6 Astra achieves 22.48% average success rate and 28.97 Score over 2,100 trials, ranking above every public entry. Yet GPT-5.5 and DeepSeek-Flash reach only 0.88% and 1.92% average success rate with the same post-processing. We find that Astra exhibits a sharply polarized capability profile. It generalizes well to tasks that require semantic understanding but not high-precision control. In contrast, it performs poorly on tasks that require precision, dynamic control, or complex bimanual coordination. In-context experiments show no aggregate benefit from one-shot demonstrations, while selected interaction traces show within-episode corrections under perturbations. Overall, the evaluated LLMs vary substantially in manipulation performance. Astra stands out and provides initial evidence for the potential of a general-purpose manipulation model, although reliable precision and dynamic control remain limitations in the evaluated setting.
Chinese Translation
具身人工智能系统通常被组织为系统1(System 1)和系统2(System 2)。系统1通常是一个以高频率生成动作的预训练策略,而系统2通常被实现为具备视觉能力的语言模型,用于高层规划。我们提出了这样一个问题:大型语言模型(LLM)能否在无需任务特定微调的情况下充当机器人操作(manipulation)的策略?我们将这一设置称为“LLM即策略”(LLM as policy)。我们在全部42个RoboDojo任务上评估了三个LLM,并将其得分与40个公开策略进行比较。其中,Astra和GPT-5.5采用官方的每任务50次试验(episode)协议,DeepSeek-Flash则采用每任务10次试验。GPT-6 Astra在2,100次试验中取得了22.48%的平均成功率和28.97的Score,排名超过所有公开的参赛方法。然而,在相同的后处理条件下,GPT-5.5和DeepSeek-Flash的平均成功率分别仅为0.88%和1.92%。我们发现Astra呈现出高度两极化的能力特征:它能够很好地泛化到那些需要语义理解但不需要高精度控制的任务;相比之下,它在需要精度、动态控制或复杂双手协调的任务上表现较差。上下文(in-context)实验表明,单次示范并未带来总体上的收益,而选取的交互轨迹则显示其在扰动下具有回合内的纠错行为。总体而言,被评估的各个LLM在操作性能上差异显著。Astra表现突出,为通用操作模型的潜力提供了初步证据,尽管在被评估的设置下,可靠的精度和动态控制仍是其局限所在。
cs.CV / 146 / 2609.24172

LegendBench: A Diagnostic Benchmark for Legend Understanding with Counterfactual Interventions

LegendBench:基于反事实干预的图例理解诊断基准
Zhang, Xinnuo, Tang, Zhike, Xu, Jing, Zhao, Haoyuan, Yang, Weikai
Abstract
Legends are fundamental to chart understanding, as reliable interpretation requires correctly binding legend entries to corresponding visual marks. While vision-language models (VLMs) are increasingly applied to chart understanding, their legend understanding is poorly diagnosed by aggregate accuracy, which can be satisfied by superficial shortcuts and confound legend-specific errors with other reasoning failures. To enable fine-grained diagnosis and controlled testing, we introduce LegendBench, a parametric benchmark and generation pipeline that produces targeted legend-centric test cases. LegendBench contributes (1) a capability-task taxonomy spanning legend parsing, legend grounding, legend-conditioned reasoning, and legend-aware abstention to localize failures, and (2) counterfactual group generation, where each base chart yields multiple variants under controlled legend interventions to probe model invariance and sensitivity. Using LegendBench, we evaluate both general-purpose VLMs and specialized chart models and generate their capability profiles, revealing persistent bottlenecks in reliable legend-to-mark binding and counterfactual consistency. We then use these capability profiles to guide targeted fine-tuning, demonstrating that bottleneck-specific interventions can effectively close the localized capability gaps and generalize to unseen data. We further leverage our counterfactual design to conduct fine-grained diagnostic experiments, analyzing encoding-channel effects, legend-order shortcuts, and abstention under varying visibility.
Chinese Translation
图例是图表理解的基础,因为可靠的解读需要将图例条目正确地绑定到对应的视觉标记上。尽管视觉语言模型(VLMs)在图表理解中的应用日益广泛,但其图例理解能力却难以通过总体准确率得到有效诊断——总体准确率可以通过表面捷径来满足,并且会将图例特有的错误与其他推理失败混淆。为了实现细粒度诊断和可控测试,我们提出了 LegendBench,这是一个参数化基准和生成流水线,能够产生以图例为中心的针对性测试用例。LegendBench 的贡献包括:(1)一个能力-任务分类体系,涵盖图例解析、图例定位、基于图例的推理以及图例感知的拒答,用于定位模型的失败点;(2)反事实组生成方法,其中每个基础图表在受控的图例干预下产生多个变体,以探测模型的不变性与敏感性。利用 LegendBench,我们评估了通用视觉语言模型和专用图表模型,并生成它们的能力画像,揭示了在可靠的图例-标记绑定和反事实一致性方面持续存在的瓶颈。随后,我们利用这些能力画像来指导针对性微调,证明针对瓶颈的干预措施能够有效弥合局部化能力差距,并能泛化到未见数据。我们进一步利用反事实设计开展细粒度诊断实验,分析了编码通道效应、图例顺序捷径以及不同可见性条件下的拒答行为。
cs.CV / 147 / 2609.24190

Benchmarking Off-the-Shelf Multimodal AI Models Against Dermatologists on Patient-Captured Skin Images

现成多模态人工智能模型与皮肤科医生在患者自摄皮肤图像上的基准对比
Dolphin, Rian, Knowles, Laura
Abstract
Artificial intelligence (AI) has advanced at a rapid pace in recent years. Initially, breakthroughs in large language models caught widespread attention. However, recent generations of frontier AI models have adopted multimodal capabilities as a first class citizen, with vision capabilities being central to that. In this paper, we evaluate three recently released models on the task of diagnosing dermatological conditions from patient-submitted images. The models chosen are at the low to mid tier in terms of pricing and thus represent a floor on current AI capabilities, not a ceiling. We evaluate AI performance relative to a panel of three certified dermatologists, who grade each image, and we present four interesting findings. Firstly, depending on the metric, the tested AI models are either on par or slightly trail humans in terms of inter-clinician agreement. Secondly, we find that asking AI models for a confidence rating produces poorly calibrated answers, meaning use of confidence thresholds should not be relied upon in a clinical setting. Thirdly, the effect of providing additional patient metadata is strongly model-specific, with one of the three models degrading on every metric considered. Finally, model cost is not predictive of performance. The best-performing model we tested costs on average $0.0045 per case.
Chinese Translation
近年来,人工智能(AI)发展迅速。最初,大语言模型的突破引起了广泛关注。然而,最近几代前沿AI模型已将多模态能力作为一等公民,其中视觉能力是核心。本文评估了三个近期发布的模型在基于患者提交的图像诊断皮肤病的任务上的表现。所选模型在定价上处于中低档次,因此代表的是当前AI能力的下限而非上限。我们将AI的表现与由三名认证皮肤科医生组成的专家小组(他们对每张图像进行评分)进行对比评估,并提出了四个有趣的发现。首先,根据不同的指标,被测试的AI模型在临床医生间一致性方面要么与人类持平,要么略逊于人类。其次,我们发现要求AI模型给出置信度评分会产生校准不良的答案,这意味着在临床环境中不应依赖置信度阈值。第三,提供额外患者元数据的影响具有很强的模型特异性,三个模型中有一个在所有考虑的指标上均出现性能下降。最后,模型成本并不能预测其性能。我们测试的表现最好的模型平均每个病例成本仅为0.0045美元。
cs.CV / 148 / 2609.24193

Lightweight Pedestrian Head-Orientation Recognition Network for Safe Pedestrian-Vehicle Interaction

Li, Yuanzhe, Huang, Yidi, Chang, Xiaotong, Liu, Hounian
Abstract
Pedestrian head orientation recognition plays an important role in autonomous driving by providing valuable cues for understanding pedestrian attention and anticipating potential crossing behavior. However, reliable recognition in real-world traffic scenes remains challenging because pedestrian head regions are often captured at low resolution. To address this challenge, we propose a lightweight Low-Resolution Head Orientation Convolutional Neural Network (LRHO-CNN) for pedestrian head orientation recognition. We construct a new dataset by extracting pedestrian head images from multiple public datasets and manually annotating them into eight orientation categories. The collected images are systematically preprocessed and augmented to increase data diversity and better represent variations in illumination and image quality. The experimental analysis compares LRHO-CNN with three fine-tuned baseline models, namely ResNet-18, ResNet-34, and VGG-16. The results demonstrate that LRHO-CNN achieves the highest classification accuracy among the evaluated models. LRHO-CNN is further evaluated on the JAAD and PIE datasets, demonstrating its effectiveness in recognizing pedestrian head orientation in real-world traffic scenes and providing informative head-orientation cues that can support downstream pedestrian behavior and intention prediction.
cs.CV / 149 / 2609.24204

SAFe: Segment-guided Aggregation of Feature Densities for Anomaly-aware Segmentation

SAFe:面向异常感知分割的分段引导特征密度聚合方法
Delić, Anja, Runtas, Jurica, Oršić, Marin, Marković, Ivan, Petrović, Ivan
Abstract
Visual segmentation systems encounter objects outside their training distribution during real-world deployment, hindering reliable autonomous systems that depend on scene parsing in the perception stage. Many recent methods address this by using self-supervised foundation models to train density estimators that yield low likelihood in anomalous image regions. Although promising, these methods suffer from poor feature semantics or they lack spatial consistency, both of which undermine critical downstream decisions. We address this problem with~\method, a generative method based on class-conditional density estimation over self-supervised representations. SAFe trains lightweight normalizing flows that produce class-conditional normalized likelihood estimates over frozen DINOv3 features. We combine density estimates from transformer features with density scores over multi-scale convolutional features to capture both global semantics and local detail. We introduce a method-agnostic post-processing step based on SAM3 that connects per-location likelihoods into spatially coherent segments while suppressing false positives, and enables instance-level anomaly detection without retraining. The post processing further distinguishes novel categories among anomalous objects by a similarity-based agglomerative clustering scheme. SAFe sets a new state of the art on the PANIC, OoDIS, SMIYC ObstacleTrack with strong performance on the ISSU benchmark.
Chinese Translation
视觉分割系统在现实部署中会遇到训练分布之外的物体,这阻碍了依赖感知阶段场景解析的自主系统的可靠性。近期许多方法通过使用自监督基础模型训练密度估计器来解决这一问题,使其在图像异常区域产生较低的似然值。尽管这些方法颇具前景,但它们要么特征语义较差,要么缺乏空间一致性,这两者都会削弱关键的下游决策。我们提出~\method来应对这一问题,这是一种基于自监督表示上类条件密度估计的生成式方法。SAFe训练轻量级归一化流(normalizing flows),在冻结的DINOv3特征上产生类条件的归一化似然估计。我们将Transformer特征的密度估计与多尺度卷积特征上的密度得分相结合,以同时捕获全局语义和局部细节。我们引入了一种与具体方法无关的后处理步骤,基于SAM3将逐位置的似然值连接成空间上连贯的分割段,同时抑制误检,并且无需重新训练即可实现实例级的异常检测。该后处理还通过基于相似度的层次聚类方案进一步区分异常物体中的新类别。SAFe在PANIC、OoDIS和SMIYC ObstacleTrack数据集上取得了新的最优性能,并在ISSU基准上表现出色。
cs.CV / 150 / 2609.24208

CoaG: Cylinders on a Grid: Coarse 3D Layout Control for Video Generation

CoaG:网格上的圆柱体:面向视频生成的粗略3D布局控制
Yang, Zhangsihao, Shan, Mengyi
Abstract
We ask how little geometry a person has to draw to control both where people stand and where the camera moves in a generated video. Our answer is a ground plane and one cylinder per person. A user draws a grid on the ground, places one cylinder where each person should stand, moves the cylinders and the camera over 81 frames, and the model renders a photoreal video in which the people occupy the cylinders' positions, move as the cylinders move, and are seen from the drawn camera. Appearance comes from a text prompt and a background reference image; layout and motion come from the geometry. Because no dataset pairs such a signal with video, we build the pairs ourselves: an automatic engine writes 2000 captions from a combinatorial seed, generates a clip for each with a text-to-video model, and lifts every clip back to its geometry with person tracking, background inpainting, an agentic ground-mask loop, feed-forward multi-view reconstruction and a plane fit, with no real footage and no manual labels. A LoRA on Wan2.2-Fun-Control trained on 1935 such tuples follows drawn layouts and camera paths on hold-out clips: the generated people match the cylinders' count, order, position and height, the text changes who they are, the reference image changes where they are, and dolly-in, orbit, pan and crane paths are followed, dolly-out only weakly.
Chinese Translation
我们探讨了一个问题:用户需要绘制多少几何信息,才能同时控制生成视频中人物所站立的位置和相机的运动。我们的答案是:一个地面平面加上每个人物对应的一个圆柱体。用户在地面绘制网格,在每个角色应当站立的位置放置一个圆柱体,然后在81帧内移动这些圆柱体和相机,模型即可渲染出照片级真实的视频,其中人物占据圆柱体的位置,随圆柱体移动,并从绘制的相机视角被观察。外观来自文本提示和背景参考图像;布局和运动则来自几何信息。由于现有数据集中没有将这种信号与视频配对的数据,我们自行构建了这样的配对数据:一个自动化引擎从组合种子生成2000条字幕,用文本到视频模型为每条字幕生成一段视频,然后通过人物跟踪、背景修复、智能体式地面掩码循环、前馈多视角重建和平面拟合,将每段视频反演回其几何表示,全程不使用真实素材,也无需人工标注。我们在1935个这样的数据元组上训练了Wan2.2-Fun-Control的LoRA模型,在留出(hold-out)视频片段上能够遵循绘制的布局和相机路径:生成的人物与圆柱体的数量、顺序、位置和高度相匹配,文本改变人物身份,参考图像改变人物所处位置,推近(dolly-in)、环绕(orbit)、平摇(pan)和升降(crane)相机路径均能被遵循,而拉远(dolly-out)只能被较弱地遵循。
cs.CV / 151 / 2609.24210

ChartJudgeBench: Evaluating LMM Judges for Chart-to-Code Generation

Wu, Lijian, Zhao, Henry Hengyuan, Zhang, Zijian, Tang, Jiahao, Wu, Jiajun, Wang, Alex Jinpeng
Abstract
Building strong chart-to-code systems increasingly relies on reinforcement learning, whose effectiveness depends critically on the quality of the reward signal. Large Multimodal Models (LMMs) play a natural critical role in jointly assessing chart visual appearance and task requirements. They are therefore increasingly used as visual critics and reward models, yet their reliability as judges remains largely unexplored. To this end, we introduce ChartJudgeBench, a diagnostic vision-language benchmark for assessing LMM judges in chart-to-code workflows. It includes 1,003 Chart Perception Alignment (CPA) instances for pairwise chart comparison and 650 Chart Reasoning Judgment (CRJ) instances for binary Accept/Reject verification in Chart Reproduction and Chart Editing. Together, these tasks emulate the core judging decisions required in agentic refinement and RL-based chart optimization. Our evaluation of strong LMMs reveals four systematic limitations: (i) positional bias in pairwise comparison, (ii) a strong tendency to overpredict Accept, (iii) difficulty in matching visual styles and aesthetics, and (iv) an unexpected leniency bias in RL-trained models. These findings show that current LMM judges require explicit reliability validation before being used as critics or reward models in chart-to-code optimization. The code and data are available on ChartJudgeBench.
cs.CV / 152 / 2609.24215

Beyond Emotion Prompts: Fine-Grained Text-to-Image Generation Driven by Valence-Arousal-Dominance

超越情感提示词:由效价-唤醒度-支配度驱动的细粒度文生图生成
Li, Minglang, Fang, Yueyue, Gao, Xieping
Abstract
Although text-to-image models can accurately depict subjects and scenes, creators still struggle to specify the fine-grained emotions an image should convey without rewriting its content description. Natural language can suggest emotions, but it offers no control scale with stable meanings and ordered intensities. We propose EMOTRANS, which transforms psychologically grounded valence-arousal-dominance (VAD) coordinates into generation conditions that are independent of the content text and modulated across denoising stages, making emotional style a finely adjustable creative variable. To support this goal, we construct EMOVAD, an art-painting dataset that pairs objective content descriptions with separately collected emotional ratings from multiple annotators. We also coordinate emotional expression and content preservation through dual-branch training with a shared model. Objective and human evaluations show that the framework improves the accuracy of three-dimensional emotion control and produces perceptible, orderable continuous changes while maintaining competitive text alignment and image quality. This work provides a practical emotion-driven approach to image generation that extends objective content depiction to fine-grained emotional adjustment.
Chinese Translation
尽管文本到图像模型能够准确描绘主体和场景,但创作者仍难以在不重写内容描述的情况下指定图像应传达的细粒度情感。自然语言可以暗示情感,但无法提供具有稳定含义和有序强度的可控尺度。我们提出EMOTRANS,它将具有心理学依据的效价-唤醒度-支配度(VAD)坐标转化为与内容文本无关、并在去噪各阶段被调制的生成条件,使情感风格成为可精细调节的创作变量。为支持这一目标,我们构建了EMOVAD,一个艺术绘画数据集,其中将客观内容描述与由多名标注者分别收集的情感评分相配对。我们还通过采用共享模型的双分支训练来协调情感表达与内容保留。客观评估和人类评估表明,该框架提升了三维情感控制的准确性,产生了可感知、可排序的连续变化,同时保持了具有竞争力的文本一致性和图像质量。这项工作为图像生成提供了一种实用的情感驱动方法,将客观内容描绘扩展到细粒度的情感调节。
cs.CV / 153 / 2609.24220

Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion

Allu, Uday, Sivaprakash, Abhivanth, Singh, Pratik, Manocha, Aman
Abstract
Retrieval-Augmented Generation (RAG) systems over enterprise knowledge bases must ingest heterogeneous document formats -- PDFs, Word documents, presentations, and scans -- whose content is locked inside complex visual layouts, multi-column pages, and dense tables. Rule-based extraction and OCR destroy reading order, flatten tables, and lose heading hierarchy, while fully agentic chunking over extracted text incurs high token costs and hallucination risk. We present Document Retrieval-Aware Chunking (D-RAC), an extension of our Web Retrieval-Aware Chunking (W-RAC) framework to arbitrary document formats. D-RAC first normalizes any input document into PDF, exploiting the fact that virtually every format has a faithful, deterministic PDF rendering. A single multimodal LLM pass then converts rendered pages into retrieval-optimized Markdown -- rewriting tables as self-contained prose statements and preserving heading hierarchy -- after which chunking proceeds exactly as in W-RAC: deterministic parsing into ID-addressable units followed by lightweight LLM-based chunk planning over identifiers rather than text. Source text is never regenerated during chunking, preserving W-RAC's cost, determinism, and observability benefits while unlocking every renderable format as a first-class input. On the 236-document, 795-page PDF subset of the RAG-Multi-Corpus benchmark spanning five enterprise domains, D-RAC converts and chunks the entire corpus in 72 minutes with zero errors, producing 1,748 retrieval-ready chunks. Compared to agentic chunking with frontier LLMs, D-RAC reduces chunking-stage output tokens by 95.7%, cutting chunking cost by 77.8% (GPT-4.1 pricing) to 85.6% (Gemini 2.5 Pro pricing) and chunking time by 75%. D-RAC scales linearly to documents of 500+ pages.
cs.CV / 154 / 2609.24223

DiaSeg: Diagonal Segment Extraction from DTW Paths for Interpretable Gait Analysis

DiaSeg:从DTW路径中提取对角线段以实现可解释的步态分析
Koffi, Tresor Y., Hidouri, Amel, Legrand, Corentin, Bertaux, Aurélie
Abstract
Dynamic Time Warping (DTW) is the dominant approach for measuring similarity between time series, yet standard practice discards the optimal warping path after computing a single distance value, losing local alignment information most relevant to clinical diagnosis. We introduce DiaSeg, a framework that extracts diagonal segments from DTW paths with controlled breaks, characterizing each segment by five geometric features (effective length, interruption count, cost variation, temporal position, and path context), and enabling unsupervised pattern discovery without domain-specific feature engineering. Validated on 91 subjects across six clinical conditions (healthy aging, Parkinson's, Huntington's, ALS, brain tumor, and stroke), three findings emerge. First, diagonal segments form consistent unsupervised patterns (silhouette 0.33) aligned with biomechanical phase annotations, with label-based validation confirming near-perfect separation of healthy and pathological gait (ARI up to 0.986). Second, segments discriminate pathology at 69% (supervised) and 75% (patient-level clustering), with pathology manifesting through distributional shifts in segment length; combining segment and cycle-level features further improves classification to 91.7%. Third, while cycle-based methods achieve higher accuracy (91%), diagonal segments provide phase-specific interpretability unavailable in global representations, localizing where coordination breaks down within the gait cycle. DiaSeg thus transforms DTW from a black-box distance into a source of interpretable temporal features for neurodegenerative disease assessment.
Chinese Translation
动态时间规整(DTW)是度量时间序列相似性的主流方法,然而标准做法在计算出一个距离值后便丢弃最优规整路径,损失了与临床诊断最相关的局部对齐信息。我们提出DiaSeg框架,该框架从DTW路径中提取带受控间断的对角线段,用五个几何特征(有效长度、间断次数、代价变化、时间位置和路径上下文)刻画每个线段,从而无需领域特定的特征工程即可实现无监督模式发现。在涵盖六种临床状况(健康老龄化、帕金森病、亨廷顿病、肌萎缩侧索硬化症(ALS)、脑肿瘤和脑卒中)的91名受试者上进行了验证,得出三项发现。第一,对角线段形成与生物力学相位标注一致的无监督模式(轮廓系数0.33),基于标签的验证证实健康步态与病态步态可近乎完美地分离(调整兰德指数(ARI)高达0.986)。第二,对角线段在监督学习下可达到69%的疾病判别准确率,在患者级聚类下达到75%,病理通过线段长度的分布偏移得以体现;将线段特征与步态周期级特征相结合可将分类准确率进一步提升至91.7%。第三,尽管基于周期的方法具有更高的准确率(91%),但对角线段提供了全局表示所不具备的相位级可解释性,能够定位步态周期内协调失效的具体位置。因此,DiaSeg将DTW从一个黑盒距离度量转变为神经退行性疾病评估中可解释时间特征的来源。
cs.CV / 155 / 2609.24226

SRPR-Net: Semantic and Relational Prompt Refinement for Automated SAM-based Instance Segmentation

SRPR-Net:面向自动化SAM实例分割的语义与关系提示精炼方法
Liu, Lufei, Li, Guojie, Xiang, Suncheng, Zhang, Fan
Abstract
Instance segmentation is a fundamental computer vision task with diverse real-world applications. Recently, prompt-driven foundation models have shown promising generalization. However, automated prompting remains limited by insufficient semantic guidance and inter-instance modeling. To address this challenge, we propose a novel architecture, named Semantic Relational Prompt Refinement Network (SRPR-Net), for automated SAM-based instance segmentation. A sequential prompt refinement mechanism is introduced to enrich detector geometry with visual-language semantics and then incorporate same-image instance dependencies, enabling context-aware box adjustment before SAM segmentation. Experiments on multiple standard benchmarks demonstrate that SRPR-Net achieves consistent improvements in segmentation performance over existing state-of-the-art approaches. The code is publicly available at https://github.com/JeremyXSC/SRPR-Net.
Chinese Translation
实例分割是一项基础的计算机视觉任务,具有广泛的真实世界应用。近年来,基于提示的基础模型展现出良好的泛化能力。然而,自动化提示仍受限于语义引导不足以及实例间建模的欠缺。为应对这一挑战,我们提出了一种新颖的架构——语义关系提示精炼网络(SRPR-Net),用于自动化SAM实例分割。该架构引入了一种顺序式提示精炼机制,首先利用视觉-语言语义丰富检测器的几何信息,随后融合同一图像中实例间的依赖关系,从而在SAM分割之前实现具有上下文感知能力的边界框调整。在多个标准基准数据集上的实验表明,SRPR-Net的分割性能较现有最先进方法取得了一致的提升。代码已公开于 https://github.com/JeremyXSC/SRPR-Net。
cs.CV / 156 / 2609.24228

IMPLICIT-Bench: Measuring Implicit Bias in Text-to-Image Models under Neutral Prompts

Dai, Yue, Liu, Ziyang, Cheong, Marc, Han, Caren
Abstract
Text-to-image (T2I) models are typically evaluated for bias using slot-based templates such as ``a photo of a [profession]''. Such templates probe only \emph{explicit} demographic attributes (e.g., gender, skin tone) in isolation. They overlook a broader \emph{implicit} bias that arises in natural prompts: when stereotype-relevant attributes are left unspecified, models still default to stereotypical outputs. We introduce IMPLICIT-Bench, a benchmark for measuring implicit bias in T2I models under such prompts. The key design is a structured-knowledge-graph (KG) construction of controlled prompt triplets: neutral, stereotype, and anti-stereotype variants that differ only along a single bias dimension while preserving scene semantics. This enables precise attribution of bias effects that template benchmarks cannot achieve. IMPLICIT-Bench comprises 5,493 prompts across 11 bias categories, validated through multi-model agreement, CLIP-based verification, and human evaluation. Using this benchmark, we show that state-of-the-art T2I models exhibit systematic bias under neutral prompts, a failure mode largely invisible to existing evaluations. We then use IMPLICIT-Bench to evaluate debiasing methods, uncovering a fundamental trade-off between bias reduction and semantic fidelity.
cs.CV / 157 / 2609.24244

Look Where It Counts: A Free, Label-Free Visual Evidence Signal for Fine-Grained Vision-Language Reasoning

看在关键处:一种免费、无标签的视觉证据信号用于细粒度视觉-语言推理
Tiwari, Santi Ram, Naik, Nihal, Pandey, Devbrat, Sinha, Nishant
Abstract
Multimodal large language models (MLLMs) fail at fine-grained visual questions less because they cannot reason than because they never see the evidence: high-resolution images are downsampled before encoding, so the model answers from linguistic priors. The standard remedies are expensive: annotated answers (SFT), hand-engineered verifiers (RLVR), or a large external teacher (on-policy distillation). We ask whether the visual evidence itself can supply the signal for free. We formalize the contrastive evidence gap, the per-token log-likelihood ratio that a model assigns to its own output when conditioned on a question-relevant region versus an irrelevant one, and study it across Qwen2.5-VL-7B, Qwen3-VL-8B, and Qwen3-VL-30B-A3B on V*Bench. Our main positive result is training-free: selecting the candidate crop under which the model's answer distribution is most peaked, using a single-view, label-free criterion, discovers the answer-bearing region with no bounding boxes, training, or labels. It localizes the target 4.4 to 5.1 times better than chance and raises fine-grained accuracy from 70 percent to 85 percent at inference. We further show that the gap is complementary to the model's own confidence. Combining them predicts correctness better than either alone, with AUC up to 0.99, and flags confidently wrong answers, with AUC ranging from 0.97 to 1.00 within the high-confidence subset. All effects concentrate on perception-bottleneck questions and vanish on a global-context control. Finally, we report an honest negative result: converting the same signal into a training method, gated self-distillation (SEG-Distill), does not outperform the base model at pilot scale across three gate designs, while more aggressive gating degrades accuracy. The signal is real, but converting it into training gains remains an open problem.
Chinese Translation
多模态大语言模型(MLLMs)在细粒度视觉问题上表现不佳,与其说是因为无法推理,不如说是因为它们从未看到证据:高分辨率图像在编码前被降采样,导致模型依赖语言先验作答。标准 remedies 代价高昂:标注答案的监督微调(SFT)、人工设计的验证器(RLVR),或大型外部教师模型(on-policy 蒸馏)。我们探究视觉证据本身能否免费地提供信号。我们形式化了对比证据差距——即模型在以问题相关区域与无关区域为条件时,对自己输出所赋予的逐 token 对数似然比——并在 V*Bench 上针对 Qwen2.5-VL-7B、Qwen3-VL-8B 和 Qwen3-VL-30B-A3B 进行了研究。我们的主要正向结果是免训练的:通过单一视角、无标签的准则,选择使模型答案分布最尖锐的候选裁剪区域,无需边界框、训练或标签即可发现包含答案的区域。其定位目标的准确率比随机高 4.4 至 5.1 倍,并在推理时将细粒度准确率从 70% 提升至 85%。我们进一步证明,该差距与模型自身的置信度是互补的:二者结合对正确性的预测优于任一单独使用,AUC 最高达 0.99,并且能标记出高置信度的错误答案,在高置信度子集内 AUC 介于 0.97 至 1.00 之间。所有效应均集中在感知瓶颈型问题上,并在全局上下文对照中消失。最后,我们报告了一个诚实的负面结果:将同一信号转化为训练方法——门控自蒸馏(SEG-Distill)——在试点规模下,三种门控设计均未能超越基线模型,而更激进的门控反而降低准确率。该信号是真实存在的,但如何将其转化为训练收益仍是一个开放问题。
cs.CV / 158 / 2609.24276

Hierarchical Prompt Learning for Hyperbolic Vision-Language Models

面向双曲视觉-语言模型的层次化提示学习
Erdelez, Andro, Mettes, Pascal, Bozorgtabar, Behzad
Abstract
Hyperbolic vision-language models (VLMs) represent image and text features in a geometry naturally suited to hierarchy, but their adaptation to downstream tasks has largely relied on fixed prompts. Existing prompt learning methods, meanwhile, treat class labels as a flat set and do not exploit available taxonomic structure. We address this gap with a hierarchical prompt learning plug-in for frozen hyperbolic VLMs. Given a fixed offline parent-class hierarchy, it augments a class prompt learner with a separate parent prompt learner, parent-level supervision, hyperbolic entailment regularization, and parent-feedback logit fusion. We instantiate the method with CoOp, CoCoOp and MaPLe, yielding HyPLO, CoHyPLO and MaHyPLO. Across the standard 11-dataset benchmark, all variants improve base-to-new generalization and cross-dataset transfer, and remain comparable to their prompt learning baselines under domain shift. Six hierarchical metrics and embedding analyses show that the method produces more taxonomically consistent predictions and induces a hierarchy-consistent organization of parent, class, and image embeddings in hyperbolic space. Its gains are largest when novel classes must be placed within a fixed taxonomy, and smallest for fine-grained confusions among sibling classes or shifts affecting only the image distribution.
Chinese Translation
双曲视觉-语言模型(VLM)在一种天然适合表示层次结构的几何空间中表征图像与文本特征,但其在下游任务上的适配在很大程度上依赖于固定提示。与此同时,现有的提示学习方法将类别标签视为扁平集合,未能利用可用的分类层级结构。我们针对这一空白,提出了一种用于冻结双曲视觉-语言模型的层次化提示学习插件。在给定固定的离线父类层级结构的情况下,该方法在类别提示学习器之外增加了一个独立的父类提示学习器、父类层级监督、双曲蕴含正则化以及父类反馈的logit融合机制。我们将该方法与CoOp、CoCoOp和MaPLe相结合,分别得到HyPLO、CoHyPLO和MaHyPLO。在标准的11个数据集基准测试中,所有变体均提升了基类到新类的泛化能力和跨数据集迁移性能,并在域偏移下与其提示学习基线保持相当的水平。六项层次化指标及嵌入分析表明,该方法能够产生更具分类学一致性的预测,并在双曲空间中诱导出父类、类别与图像嵌入之间层次一致的组织结构。当新类别必须被置于固定分类体系之中时,该方法的增益最大;而在兄弟类别之间的细粒度混淆或仅影响图像分布的偏移场景下,增益最小。
cs.CV / 159 / 2609.24287

Classifier-Free Guidance in Flow Matching: Non-Autonomous Potentials, Overshoot, and Posterior-Mean Control

流匹配中的无分类器引导:非自治势场、过冲与后验均值控制
Peng, Jishen, Ma, Zheng
Abstract
Classifier-free guidance (CFG) improves conditional generation in Flow Matching, but strong guidance can distort the generated distribution and reduce diversity. We provide a geometric account of this behavior by viewing Flow Matching as a time-varying gradient flow and characterizing how CFG reshapes its underlying potential. This view explains how stronger alignment can be accompanied by mean displacement and trajectory concentration, and motivates controlling guidance through the model-implied terminal posterior mean. We therefore propose Posterior-Mean-Capped CFG (PMC-CFG), a training-free, per-sample method that adaptively retains the strongest feasible guidance without additional network evaluations. Experiments on synthetic and large-scale image-generation benchmarks show that PMC-CFG limits guidance-induced distortion and concentration while improving the alignment--diversity trade-off, with particularly strong benefits when nominal guidance is large.
Chinese Translation
无分类器引导(Classifier-Free Guidance, CFG)能够提升流匹配(Flow Matching)中的条件生成质量,但过强的引导会扭曲生成分布并降低多样性。我们将流匹配视为一种时变梯度流,并刻画CFG如何重塑其底层势场,从而为这一现象提供了几何层面的解释。这一视角解释了为何更强的对齐会伴随均值偏移和轨迹集中,并启发我们通过模型隐含的终端后验均值来控制引导强度。为此,我们提出了后验均值上限CFG(Posterior-Mean-Capped CFG, PMC-CFG),这是一种无需训练的逐样本方法,能够在不增加额外网络评估的情况下自适应地保留最强可行引导。在合成数据和大规模图像生成基准上的实验表明,PMC-CFG 能够抑制引导导致的失真与集中现象,同时改善对齐与多样性之间的权衡,尤其在名义引导强度较大时收益尤为显著。
cs.CV / 160 / 2609.24308

HappyWorld-Bench

Bai, Zhiqi, Cai, Junai, Chen, Yixin, Du, Jingrun, Feng, Tao, Gong, Wei, Huang, Siyuan, Lin, Xiao, Liu, Jiaheng, Luo, Jun, Lyu, Yongzhe, Ma, Liya, Meng, Zenan, Qu, Lin, Su, Wenbo, Wang, Jiaming, Wang, Qinghe, Wang, Shaofei, Wang, Yanghai, Wang, Zequn, Wang, Ziming, Wei, Hu, Wu, Jiangtao, Wu, Ruiqi, Xie, Jiaxin, Xu, Yuchi, Xu, Ze, Yu, Chengting, Yuan, Liangyu, Zeng, Gang, Zeng, Yawen, Zhang, Xingyao, Zhang, Zizheng, Zheng, Bo, Zhu, Jiancheng, Zhu, Song-Chun
Abstract
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.
cs.CV / 161 / 2609.24312

AnalogDepth: Multi-view Geometry from FPV drones under Analog Video Transmission

AnalogDepth:基于模拟视频传输的FPV无人机多视角几何
Amorim, André, Proença, Pedro F.
Abstract
Analog video transmission (VTX) remains widespread in FPV drones due to low latency, weight and low cost. However analog VTX suffers from complex spatially structured image degradation which differ fundamentally from digital image corruption (e.g. AWGN) used in standard training augmentation. This work shows that this type of noise severely degrades the accuracy of Depth Anything 3 (DA3), a state-of-the-art feed forward visual geometry foundation model. To address this gap, we present AnalogDepth, a parameter-efficient training pipeline that adapts DA3 to analog FPV imagery using student-teacher knowledge distillation with Low-Rank Adaptation (LoRA) injected into the DINOv2 backbone. Rather than synthesizing noise analytically, we build a noise bank from static FPV recordings under diverse conditions and compare real-noise injection against PSD-matched Gaussian synthesis and AWGN as baselines. Experiments on six real FPV flight sequences across three indoor scenes show that training with our noise bank consistently reduces per-frame depth RMSE and 3D reconstruction Chamfer distance compared to the pretrained DA3 baseline and both Gaussian noise variants. These results demonstrate that replicating the spatial structure of real analog transmission noise is critical for effective adaptation.
Chinese Translation
模拟视频传输(VTX)因其低延迟、轻量化和低成本而在FPV无人机中广泛应用。然而,模拟VTX存在复杂的空间结构化图像退化问题,这与标准训练增强中使用的数字图像损坏(如加性高斯白噪声,AWGN)有着根本区别。本工作表明,这类噪声会严重降低Depth Anything 3(DA3)的精度——DA3是一种最先进的前馈式视觉几何基础模型。为弥补这一差距,我们提出AnalogDepth,一种参数高效的训练流水线,通过学生-教师知识蒸馏将DA3适配到模拟FPV图像上,并在DINOv2骨干网络中注入低秩适配(LoRA)。我们并非解析地合成噪声,而是从多种条件下静态FPV录制数据构建噪声库,并将真实噪声注入与PSD匹配的高斯合成以及AWGN作为基线进行对比。在三个室内场景的六个真实FPV飞行序列上的实验表明,与预训练的DA3基线以及两种高斯噪声变体相比,使用我们的噪声库进行训练能够持续降低逐帧深度RMSE和三维重建的Chamfer距离。这些结果证明,复现真实模拟传输噪声的空间结构对于有效适配至关重要。
cs.CV / 162 / 2609.24313

NeuIDO: Neural Intrinsic Dynamics Operator for Physics-Informed 4D World Models

Lin, Jiajing, Zhang, Xin, Sun, Jianhua
Abstract
World models aim to capture environmental dynamics and predict future trajectories, showing growing potential for embodied intelligence. Physics-informed 4D generation integrates physical simulation to predict 3D object interactions, offering a promising pathway toward world models. However, this paradigm relies on manually imposed dynamical assumptions rather than internalizing world dynamics, and thus still leaves a gap toward a true world model. To bridge this gap, we propose NeuIDO, a novel world dynamics modeling framework that learns a unified intrinsic dynamics representation from visual observations, advancing physics-informed 4D generation toward a world model. Specifically, we formulate world modeling as a neural operator learning problem and introduce a two-stage training strategy to learn a generalizable mapping from the visual observation distribution to the intrinsic dynamics distribution. Building on this observation-dynamics mapping, NeuIDO enables zero-shot dynamics inference directly from videos and can be further aligned with complex real-world dynamics via few-shot adaptation. Extensive experiments demonstrate that NeuIDO effectively unifies the intrinsic dynamics underlying diverse visual observations into a shared representation and rapidly infers dynamics in novel scenes.
cs.CV / 163 / 2609.24330

AlignMorph: Tuning-Free Diffusion Image Morphing via Explicit Semantic Transport

AlignMorph:基于显式语义传输的无需微调扩散图像变形方法
Liu, Wuyi, Han, Xu, Chen, Yuren, Mao, Yige, Peng, Zishuo, Li, Xianzhi
Abstract
Image morphing aims to produce a smooth and semantically consistent transition between two input images. Existing diffusion-based morphing methods either require expensive per-pair optimization or rely on implicit spatial alignment, which easily fails under large layout discrepancies. To address these limitations, we propose AlignMorph, a novel tuning-free diffusion framework guided by the principle of transport-then-denoise. We explicitly decouple geometric alignment from generative denoising to avoid structural entanglement. Our framework consists of two core components. (1) Global Semantic Transport, which achieves diffusion-compatible semantic alignment via entropic optimal transport and reliability-aware latent warping; and (2) Coordinate-Aligned Generation, which uses a symmetric bi-phase attention handoff to maintain consistent spatial coordinates throughout denoising. Without any tuning, AlignMorph effectively eliminates ghosting and achieves superior structural coherence and temporal smoothness on morphing benchmarks. Code is available at https://github.com/51xOne/Alignmorph.
Chinese Translation
图像变形旨在在两幅输入图像之间生成平滑且语义一致的过渡。现有的基于扩散模型的变形方法要么需要代价高昂的逐对优化,要么依赖隐式的空间对齐,在布局差异较大时容易失效。针对这些局限性,我们提出了AlignMorph,一种由“先传输后去噪”原则指导的新型无需微调的扩散框架。我们显式地将几何对齐与生成式去噪解耦,以避免结构上的相互纠缠。该框架包含两个核心组件:(1)全局语义传输,通过熵最优传输和可靠性感知的潜空间变形实现与扩散模型兼容的语义对齐;(2)坐标对齐生成,采用对称的双阶段注意力交接机制,在整个去噪过程中保持一致的空间坐标。在无需任何微调的情况下,AlignMorph有效消除了鬼影伪影,并在图像变形基准测试中取得了卓越的结构连贯性和时间平滑性。代码已发布于 https://github.com/51xOne/Alignmorph。
cs.CV / 164 / 2609.24337

LiAuto-MindViT: A Hybrid Vision Backbone with Adaptive Bidirectional Mamba

LiAuto-MindViT:一种基于自适应双向Mamba的混合视觉骨干网络
Mu, Lifu, Chen, Shuai, Zheng, Wen, Sun, Haoyi, Fu, Xueyang, Yu, Pengfei, Mao, Ning, Wei, Tao, Pan, Zhou
Abstract
While Mamba-based models have shown strong potential for long sequence modeling, adapting them to vision is challenging due to the requirement of local neighborhood correlations and multi-directional spatial contexts for visual understanding. In this paper, we present LiAuto-MindViT, a novel hybrid vision backbone that synergizes the strengths of CNNs, Mamba, and Transformers. The core of our design is the Adaptive Bidirectional Mamba (ABM), which eliminates the directional bias of unidirectional SSMs through bidirectional selective scanning with learnable alpha blending, enabling content-adaptive directional fusion without the overhead of exhaustive multi-path routing. To further accelerate inference, we propose a deployment-friendly Reparameterized ConvSE (RepConvSE) module that leverages structural reparameterization to reduce latency and memory access overhead. Extensive experiments demonstrate that LiAuto-MindViT achieves state-of-the-art performance on image classification, object detection, and semantic segmentation while enabling efficient inference through reparameterization.
Chinese Translation
尽管基于Mamba的模型在长序列建模方面展现出强大潜力,但将其适配到视觉领域仍面临挑战,原因在于视觉理解需要局部邻域相关性和多方向空间上下文。在本文中,我们提出了LiAuto-MindViT,这是一种新颖的混合视觉骨干网络,它协同融合了CNN、Mamba和Transformer的优势。我们设计的核心是自适应双向Mamba(Adaptive Bidirectional Mamba,ABM),它通过具有可学习alpha混合系数的双向选择性扫描,消除了单向状态空间模型(SSM)的方向偏差,实现了内容自适应的方向融合,而无需承担穷举多路径路由的开销。为进一步加速推理,我们提出了一种面向部署的重参数化卷积模块ReparamConvSE(RepConvSE),该模块利用结构重参数化来降低延迟和内存访问开销。大量实验表明,LiAuto-MindViT在图像分类、目标检测和语义分割任务上取得了最先进的性能,同时通过重参数化实现了高效推理。
cs.CV / 165 / 2609.24359

Dissecting Agentic Forensics: The Role of Triage, Prompting, and Evidence Arbitration in Open-World Fake Image Detection

解析智能体取证:分诊、提示与证据仲裁在开放世界伪造图像检测中的作用
Li, Xianlong, Bongini, Pietro, Pancino, Niccoló, Blanchini, Marco, Tondi, Benedetta, Barni, Mauro
Abstract
Image forensics is increasingly an open-world problem: manipulations range from fully synthetic images to localized edits, splicing and swapping, while most forensic detectors remain specialized to a single manipulation family. Agentic AI has recently emerged as a promising solution. In principle, such systems can assess the reliability of individual detectors, identify out-of-scope evidence, and arbitrate conflicting reports. However, it remains unclear which components actually drive performance and whether their benefits persist under distribution shift. To answer these questions, we study a training-free agentic framework built around specialist detectors, per-detector triage, and conflict-aware evidence arbitration. Using six configurations and three multimodal large language model backbones, we dissect the role of triage, prompting, and reasoning quality on both in-distribution and out-of-distribution data. Our results show that naive detector fusion suffers from severe false-positive rates on authentic images. Triage and prompting consistently improve performance by filtering unreliable evidence and exposing detector limitations. However, the dominant factor is represented by reasoning itself: A stronger judge substantially outperforms a weaker one, particularly under distribution shift. Most notably, manipulation recall is nearly saturated across all configurations, indicating that the main challenge of open-world image forensics is not detecting manipulations, but calibrating trust in specialized forensic tools and arbitrating conflicting evidence.
Chinese Translation
图像取证正日益成为一个开放世界问题:图像篡改形式涵盖从完全合成图像到局部编辑、拼接与替换,而大多数取证检测器仍仅针对单一篡改类型进行专门化设计。智能体AI(Agentic AI)近来成为一种有前景的解决方案。原则上,此类系统能够评估各个检测器的可靠性、识别超出适用范围的证据,并对相互冲突的报告进行仲裁。然而,究竟是哪些组件真正驱动了性能,以及其收益在分布偏移下是否仍然存在,目前仍不清楚。为回答这些问题,我们研究了一种免训练的智能体框架,该框架围绕专用检测器、逐检测器分诊(per-detector triage)以及冲突感知证据仲裁构建。通过六种配置和三个多模态大语言模型骨干,我们剖析了分诊、提示和推理质量在分布内与分布外数据上的作用。结果表明,朴素的检测器融合在真实图像上存在严重的误报率问题。分诊与提示通过过滤不可靠证据并暴露检测器的局限性,能够持续提升性能。然而,主导因素是推理本身:更强的裁判模型显著优于较弱的裁判模型,在分布偏移下尤为明显。最值得注意的是,篡改召回率在所有配置下均接近饱和,这表明开放世界图像取证的主要挑战不在于检测篡改,而在于对专用取证工具的信任校准以及对冲突证据的仲裁。
cs.CV / 166 / 2609.24367

TReViS: Temporal Repetition Structure Aware Video Synthesis for Self-supervised Repetitive Action Counting

Yu, Fanqi, Ma, Shengming, Fiorini, Stefano, Pastore, Vito Paolo, Qi, Xuan, Murino, Vittorio, Beyan, Cigdem
Abstract
Fully supervised repetitive action counting (RAC) has achieved strong performance, but requires dense temporal annotations that are costly and difficult to scale. We propose TReViS, a self-supervised video synthesis framework that enables training RAC models without any repetition labels. TReViS estimates the underlying temporal repetition structure of an unlabeled video via a Temporal Self-Similarity Matrix, infers its cycle statistics, and synthesizes new training sequences that preserve realistic repetition patterns while introducing controlled temporal variability. These synthesized videos are paired with pseudo-labels and used to train existing RAC architectures from scratch. Across multiple datasets and backbones, TReViS consistently outperforms prior self-supervised methods and achieves performance competitive with several supervised baselines, while remaining fully label-free, demonstrating the effectiveness of structure-aware video synthesis for label-free RAC. The source code is available at https://github.com/yfqi/TReViS.
cs.CV / 167 / 2609.24379

Topographic Training Concentrates Causal Circuits Without Improving Neuron Monosemanticity

Ranka, Gautam, Pandere, Shubham Santosh, Dsouza, Aiden
Abstract
Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entangle many concepts in each neuron. Feature superposition is widely treated as the central obstacle to this decomposition, yet most mitigations (sparse autoencoders, dictionary learning) are post-hoc and leave the underlying network unchanged. We ask whether a spatial-locality training loss (TopoLoss) can act as a lightweight, training-time prior that improves interpretability of standard mech-interp tools. Training ViT on ImageNet-100 across multiple TopoLoss weights $\alpha$, we measure causal sufficiency of topographic clusters via activation patching and feature geometry via sparse autoencoders fit to the same residual stream. At $\alpha=1.0$, topographic clusters are 2.79$\times$ more causally sufficient than random unit sets of the same size, with the effect increasing monotonically in $\alpha$. SAE L0 sparsity decreases by 11% and dead-feature fraction rises 19-fold, yet standard neuron-level monosemanticity scores are unchanged, indicating that topographic pressure acts at circuit level, concentrating causal mass into spatially local structures without disentangling individual neurons. This dissociation suggests current neuron-level monosemanticity metrics are insensitive to a class of real interpretability gains, and positions cheap architectural priors as a viable training-time complement to post-hoc tooling.
cs.CV / 168 / 2609.24384

A Lightweight Convolutional Neural Network for Real-Time Recognition of Hand-Drawn Geometric Shapes

Abdullah, Shahir
Abstract
Recognizing hand-drawn geometric shapes is a foundational sub-problem of sketch recognition, with applications in education, human-computer interaction, and diagram digitization. This paper presents the design, implementation, and evaluation of a desktop application that recognizes four basic hand-drawn geometric shapes, circle, square, rectangle, and triangle using a compact Convolutional Neural Network (CNN). A dataset of 2,000 labeled 28x28-pixel shape images was collected independently and released publicly. The classifier consists of three convolutional blocks (16, 32, and 64 filters) with max-pooling, an in-model data-augmentation stage (random horizontal flip, rotation, and zoom), a dropout-regularized dense layer of 128 units, and a 4-way linear output layer, totaling 97{,}956 trainable parameters. The network is trained with the Adam optimizer on a sparse categorical cross-entropy objective computed directly on logits. On an 80/20 train-validation split, the model achieves 94.80% training accuracy and 96.01% validation accuracy with a validation loss of 0.1437. A Tkinter-based graphical interface allows a user to draw a shape with the mouse and receive an immediate class prediction with a confidence score. We situate this system within the broader sketch and shape-recognition literature, compare its accuracy against related hand-drawn shape classification studies, and discuss the limitations inherent to a small, single-contributor dataset. The complete source code, trained model, and per-class datasets are released publicly to support reproducibility.
cs.CV / 169 / 2609.24403

Can Spiking Neural Networks play pinball? A neuromorphic motion detector for target tracking

脉冲神经网络能玩弹球游戏吗?一种用于目标跟踪的神经形态运动检测器
Fatahi, Mazdak, Pryjmaková, Šárka, Boulet, Pierre, D'Angelo, Giulia
Abstract
Biological visual systems achieve continuous, low-latency motion perception by processing sparse, asynchronous spiking signals, enabling real-time tracking under strict energy constraints. Event-based cameras, inspired by the mammalian retina, replicate this efficiency by capturing only local brightness changes as asynchronous events, offering a natural substrate for spiking neural networks (SNNs) to parallelise computation and adapt to fast-changing scenes. Pinball provides a controlled yet dynamic testbed, requiring precise motion estimation and fast reaction to a small, rapidly moving target. This work presents a fully spiking, real-time perception-to-action pipeline for closed-loop pinball gameplay. A dynamic vision sensor observes a small, fast-moving ball, and a network of spiking Time-Difference Encoders on the SpiNNaker neuromorphic platform jointly estimates its position, speed, and direction. The system is characterised across receptive field size, accumulation window, and angular tuning width for real-time operation, and benchmarked in closed loop against human players across two flipper regimes of increasing physical realism. It achieves a hit rate of 56.1%, nearly double the human average, reacting within 21.7 ms (5 ms network latency) and consuming an estimated 148 {\mu}W using fewer than 25k neurons, among the fastest and most energy-efficient event-based closed-loop demonstrators benchmarked. Under more realistic flipper dynamics, tuning a single interpretable policy parameter reproduces the full spectrum of human play styles, from cautious to aggressive, with no change to the perception pipeline. A physical demonstrator, tracking a real ball and actuating real flippers in closed loop, confirms the principle operates beyond simulation. Its fully spiking, learning-free design offers a compact, energy-efficient example of real-time neuromorphic perception-to-action.
Chinese Translation
生物视觉系统通过处理稀疏、异步的脉冲信号实现连续、低延迟的运动感知,从而在严格的能量约束下完成实时跟踪。受哺乳动物视网膜启发的事件相机通过仅捕获局部亮度变化作为异步事件来复制这种高效性,为脉冲神经网络(SNN)并行计算并适应快速变化的场景提供了天然的载体。弹球游戏提供了一个可控且动态的测试平台,要求对小而快速移动的目标进行精确的运动估计和快速反应。本研究提出了一条全脉冲式的实时感知-动作闭环管线,用于闭环弹球游戏。动态视觉传感器观察一个小而快速移动的球,部署在SpiNNaker神经形态平台上的脉冲式时间差编码器(spiking Time-Difference Encoder)网络联合估计其位置、速度和方向。该系统针对实时运行,在感受野大小、累积窗口和角度调谐宽度等方面进行了特性表征,并在闭环条件下与人类玩家在两种物理逼真度递增的挡板模式下进行基准对比。系统实现了56.1%的击中率,接近人类平均水平的三倍(原文为近两倍),反应时间在21.7毫秒以内(网络延迟为5毫秒),使用不到2.5万个神经元,估计功耗仅148微瓦,是经基准测试的事件驱动闭环演示系统中最快、最节能的系统之一。在更逼真的挡板动力学条件下,仅调节单个可解释的策略参数即可复现从谨慎到激进的全谱系人类游戏风格,而无需对感知管线做任何改变。一个物理演示系统——在闭环中跟踪真实的球并驱动真实的挡板——验证了该原理在仿真之外同样有效。其全脉冲、无需学习的设计为实时神经形态感知-动作系统提供了一个紧凑、节能的范例。
cs.CV / 170 / 2609.24409

DeCo: Efficient Decouple-to-Couple Learning for Multi-Task Visual Grounding

DeCo:面向多任务视觉定位的高效解耦-耦合学习方法
Lu, Xiaoqiang, Jiao, Licheng, Sun, Long, Yang, Yuting, Liu, Xu, Li, Lingling, Ma, Wenping, Liu, Fang
Abstract
Multi-task visual grounding requires models to jointly understand linguistic semantics and perform accurate visual localization and segmentation. Despite the success of multimodal large language models, effectively adapting them to multiple grounding objectives remains challenging. Existing methods commonly enforce task cooperation through shared representations, while overlooking the intrinsic conflict between task-oriented feature interests. In this paper, we introduce $\textbf{DeCo}$, an efficient $\textbf{De}$couple-to-$\textbf{Co}$uple learning framework that resolves this dilemma through a two-stage paradigm: task-specific representation decoupling followed by complementary prior coupling. Specifically, we first propose Task-aware Semantic Decoupling (TSD) to route shared visual cues into individual features under salient word-level guidance, alleviating representation interference between localization and segmentation. Furthermore, we observe that segmentation naturally provides informative localization priors due to dense supervision. Based on this insight, we introduce Hybrid Prior Coupling (HPC), which integrates sentence-level semantic prior with mask-derived spatial prior for enhanced grounding. Built upon a frozen multimodal encoder, DeCo requires lightweight trainable parameters while achieving strong generalization across multiple grounding objectives. Extensive experiments on RefCOCO/+, G-Ref, ReferIt, Flickr, DIOR-RSVG, SARVG1.0, RRSIS-D, RIS-LAD, and RefDIOR demonstrate that DeCo achieves state-of-the-art performance on both natural and remote sensing benchmarks. The code and models are available at https://github.com/xiaoqiang-lu/DeCo.
Chinese Translation
多任务视觉定位要求模型能够联合理解语言语义,并执行精确的视觉定位与分割。尽管多模态大语言模型已取得成功,但如何将其有效适配到多个定位目标仍然具有挑战性。现有方法通常通过共享表征强制任务间协作,却忽视了面向任务的特征关注点之间固有的冲突。在本文中,我们提出了DeCo,一种高效的解耦-耦合(Decouple-to-Couple)学习框架,通过两阶段范式解决这一困境:先进行任务特定的表征解耦,再进行互补先验耦合。具体而言,我们首先提出任务感知语义解耦(Task-aware Semantic Decoupling, TSD),在显著的词级引导下将共享视觉线索路由到各自的独立特征中,从而缓解定位与分割之间的表征干扰。此外,我们观察到由于密集监督,分割能够自然地提供有价值的定位先验。基于这一洞察,我们提出混合先验耦合(Hybrid Prior Coupling, HPC),将句子级语义先验与掩码衍生的空间先验相融合,以增强定位效果。DeCo构建于冻结的多模态编码器之上,仅需轻量级的可训练参数,同时在多个定位目标上实现了强大的泛化能力。在RefCOCO/+、G-Ref、ReferIt、Flickr、DIOR-RSVG、SARVG1.0、RRSIS-D、RIS-LAD和RefDIOR上的大量实验表明,DeCo在自然图像和遥感基准上均达到了最先进的性能。代码和模型已发布于 https://github.com/xiaoqiang-lu/DeCo。
cs.CV / 171 / 2609.24424

Estimating Accurate Hand Pose in Camera Space with Vision Transformer

基于视觉Transformer的相机空间精确手部姿态估计
Ren, Kaiwen, Jiang, Yiran, Ye, Yongjing, Xia, Shihong
Abstract
Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. However, this camera-space estimation confronts two fundamental challenges: (1) depth ambiguity in monocular settings, and (2) the coupling effect of hand local poses and global wrist positions in the perspective projections. In particular, this coupling reflects that the projections are jointly determined by local hand poses, wrist positions, and camera intrinsics. To overcome these challenges, our framework proposes two key innovations: Transformation-Isomorphism Supervision for hand-depth information extraction and Perspective Information Embedding for resolving above coupling effect of local pose and wrist position, both integrated within the mainstream encoder-decoder architecture. Besides, we propose a novel framerate-aware multi-dataset training strategy for sequential pose refinement. Our fully integrated approach achieves at most 37.1\% superiority in CS-MJE over SOTA on HO3D. Project page: https://github.com/Mine268/CS-ViT.
Chinese Translation
基于单目RGB的手部姿态估计已成为计算机视觉领域的重要研究前沿。局部手部姿态估计方法预测相对于手腕的手部姿态,而全局手部姿态估计还需要估计手腕在相机坐标系中的位置。然而,这种相机空间估计面临两个根本性挑战:(1)单目设置下的深度歧义性;(2)透视投影中手部局部姿态与全局手腕位置的耦合效应。具体而言,这种耦合表明投影是由局部手部姿态、手腕位置和相机内参共同决定的。为克服这些挑战,我们的框架提出了两项关键创新:用于手部深度信息提取的变换同构监督(Transformation-Isomorphism Supervision),以及用于解决上述局部姿态与手腕位置耦合效应的透视信息嵌入(Perspective Information Embedding),二者均集成于主流的编码器-解码器架构之中。此外,我们提出了一种新颖的帧率感知多数据集训练策略,用于序列姿态的精细化。我们的完整集成方法在HO3D数据集上的CS-MJE指标相比SOTA最高取得37.1%的优势。项目页面:https://github.com/Mine268/CS-ViT。
cs.CV / 172 / 2609.24452

Do LiDAR Language Models Really Understand Spatio-temporal Relationships?

LiDAR语言模型真的理解时空关系吗?
Yang, Runyi, Akkoyun, Murat, Wen, Di, Liu, Ruiping, Chen, Yufan, Zheng, Junwei, Wang, Xiaoye, Yang, Kailun, Paudel, Danda Pani, Van Gool, Luc, Peng, Kunyu
Abstract
Recent 4D LiDAR language models aim to reason about objects and their evolving spatial relationships. Yet, in our evaluation, always selecting the same option nearly matches the multiple-choice accuracy of two B4DL-derived configurations. We introduce LiDAR-Hallu, a geometry-referenced benchmark and diagnostic protocol with 10,000 questions across 150 nuScenes scenes. It covers object existence, ego-relative position, distance ordering, relative motion, and temporal localization, with explicit rules for selecting objects, comparing times, and determining reference answers. Our protocol combines fixed-answer and candidate-content controls, cross-scene pairs with identical prompts but opposite reference answers, and relation-specific recall. Analysis of 100,000 recorded responses reveals failures hidden by aggregate accuracy. Candidate duration alone makes temporal answers predictable without observing LiDAR. On paired questions, the models frequently give the same answer to scenes requiring opposite answers. Relation-specific analysis further shows that both configurations miss every positive lateral-motion case across all tested conditions. Temporal-shuffle contrastive decoding provides little net improvement, as repairs are largely offset by new errors and the main failures persist. These results show that evaluating spatio-temporal reasoning requires testing whether models distinguish the queried physical relationships, rather than relying on individual-answer accuracy alone. The source code, checkpoints, and data are released at https://github.com/Awesome4D/4DMLLM_Hallucination_Bench.
Chinese Translation
近期的4D LiDAR语言模型旨在对物体及其不断演化的空间关系进行推理。然而,在我们的评估中,始终选择同一选项几乎就能达到两种基于B4DL的配置在多选题上的准确率。我们提出了LiDAR-Hallu,这是一个基于几何参照的基准测试与诊断协议,包含覆盖150个nuScenes场景的10,000个问题。该基准涵盖物体存在性、相对自车位置、距离排序、相对运动以及时间定位,并为物体选择、时间比较和参考答案确定制定了明确规则。我们的协议结合了固定答案与候选内容控制、提示相同但参考答案相反的跨场景配对,以及针对特定关系的召回率分析。对100,000条记录响应的分析揭示了被总体准确率所掩盖的失败情况。仅凭候选时间戳的时长就可在不观察LiDAR数据的情况下预测时间类答案。在配对问题上,模型经常对需要相反答案的场景给出相同的回答。针对特定关系的分析进一步表明,在所有测试条件下,两种配置均漏检了所有正侧向运动案例。时间混洗对比解码(temporal-shuffle contrastive decoding)几乎没有带来净提升,因为修复的效果在很大程度上被新产生的错误所抵消,且主要失败依然存在。这些结果表明,评估时空推理能力需要检验模型是否能够区分所查询的物理关系,而不能仅依赖单个答案的准确率。源代码、模型检查点和数据已在 https://github.com/Awesome4D/4DMLLM_Hallucination_Bench 发布。
cs.CV / 173 / 2609.24455

MECAIL: Communication-Aware Incremental Learning for Object Detection with 14.6 KB Spatiotemporal Experts

MECAIL:面向目标检测的通信感知增量学习方法,仅需14.6 KB时空专家模块
Neuwirth-Trapp, Matthias, Bieshaar, Maarten, Paudel, Danda, Schindler, Konrad, Van Gool, Luc, Sakaridis, Christos
Abstract
Intelligent transportation systems require Incremental Learning (IL) to continually improve their overall performance in dynamic environments. However, most edge devices lack the computational resources to support on-device IL, requiring updates to be transmitted from centralized servers. We propose using this setup to obtain dense, specialized module coverage that adapts a fixed base model to specific spatiotemporal contexts, such as parking lots, gas stations, ferries, or construction sites. However, in order to reliably transmit these modules to the edge device, using TCP, UDP, and BTP over V2X, Wi-Fi, and 2G-5G hardware, we establish a strict limit of 14.6 KB per module to fit within the first TCP window and to minimize UDP/BTP fragmentation. We further introduce Mixture-of-Experts for Communication-Aware Incremental Learning (MECAIL), the first method that meets this strict requirement, in which each new domain or environment is served by a small expert network that adapts the base model. We validate MECAIL on D-RICO and ODinW-13, where it largely matches the performance of parameter-heavy approaches while enabling practical, bandwidth-efficient large-scale deployment. This allows comprehensive coverage by experts for highly specific, focused, and temporary situations.
Chinese Translation
智能交通系统需要增量学习(Incremental Learning, IL)以在动态环境中持续提升整体性能。然而,大多数边缘设备缺乏支持端侧增量学习的计算资源,需要从集中式服务器传输更新。我们提出利用这一设置来获得密集且专门化的模块覆盖,使固定的基础模型能够适应特定的时空场景,例如停车场、加油站、渡轮或建筑工地。然而,为了通过V2X、Wi-Fi以及2G-5G硬件,可靠地使用TCP、UDP和BTP将这些模块传输至边缘设备,我们将每个模块的大小严格限制在14.6 KB以内,以适配首个TCP窗口并最小化UDP/BTP的分片。我们进一步提出了面向通信感知增量学习的混合专家方法(Mixture-of-Experts for Communication-Aware Incremental Learning, MECAIL),这是首个满足这一严格限制的方法,其中每个新领域或新环境由一个小型专家网络提供服务,用以适配基础模型。我们在D-RICO和ODinW-13数据集上验证了MECAIL,其性能在很大程度上可媲美参数量庞大的方法,同时支持实用且带宽高效的大规模部署。这使得专家模块能够全面覆盖高度具体、集中且临时性的场景。
cs.CV / 174 / 2609.24468

MIGA:Shared-Geometry Gaussian Representation with Implicit Amplitude Modeling for Accelerated 3D Multi-Echo MRI

Xu, Jingran, Liu, Yuanyuan, Zhu, Yanjie
Abstract
Three-dimensional multi-echo MRI provides rich anatomical and quantitative information, but repeated volumetric encoding prolongs acquisition and motivates k-space undersampling. Reconstructing undersampled multi-echo data requires exploiting shared anatomy while preserving echo-dependent signal variation; full-volume modeling also introduces substantial computational and memory demands. We propose MIGA, a scan-specific framework comprising shared anisotropic Gaussian geometry, a coordinate-conditioned multi-output amplitude network, and explicit echo-specific phase variables. The Gaussian geometry provides common spatial support across echoes, the implicit network models spatially structured amplitude variations, and the phase variables retain echo-specific complex signal information. All components are jointly optimized using only the acquired multi-coil k-space, requiring no fully sampled training data. Experiments showed that MIGA consistently outperformed the comparison methods across imaging tasks and acceleration factors, with larger improvements under stronger undersampling. MIGA also achieved a favorable quality-cost balance among the evaluated full-volume multi-echo methods. These results support the effectiveness of combining shared Gaussian geometry with implicit echo-dependent amplitude modeling for accelerated 3D multi-echo MRI reconstruction.
cs.CV / 175 / 2609.24470

Spatial Action Review: A Visual Analytics Dashboard for Auditing Language-to-Action Hand-offs in Electron Microscopy

空间动作审查:用于审计电子显微镜中语言到动作交接的可视化分析仪表板
Mohinta, Samia, Cardona, Albert
Abstract
Multimodal large language models (MLLMs) are increasingly explored as interfaces for scientific image analysis, where a visual question-answering (VQA) response may be paired with a spatial output that guides a downstream stage. A supervisor reads the language answer, while a downstream workflow such as segmentation or region review consumes the point-set output. We call this transition from inspecting the answer to relying on its point action the language-to-action hand-off. A silent failure occurs when the answer is correct while the paired action misses annotated objects needed downstream, so answer-based oversight clears a region whose action is unreliable. We introduce Spatial Action Review, a visual analytics dashboard for auditing this failure mode in electron microscopy (EM) mitochondria analysis. It links paired answer-action records through an answer-action ledger, a task-by-dataset risk map, and an image-region audit view, connecting aggregate patterns to image evidence while an adjustable action-reliability gate supports re-audit. The review ends in a human-AI hand-off, where a supervisor records whether the action is accepted, escalated, held under a stricter gate, or flagged for model revision. Across 541 image regions from an EM-adapted Qwen3-VL case-study run, point actions fail the gate in 54.4% of records with a correct VQA response, and 27.4% of all records are silent failures. A correct answer is associated with only a 5.8-percentage-point higher probability of a reliable action, with a bootstrap interval spanning zero; the point-biserial correlation between answer correctness and object coverage is 0.061. This weak coupling persists across five model conditions on 753 matched image regions. Spatial Action Review makes answer-action mismatches visible and ties them to image evidence and a recorded decision before MLLM outputs enter autonomous scientific workflows.
Chinese Translation
多模态大语言模型(MLLM)正被越来越多地探索用作科学图像分析的接口,其中视觉问答(VQA)的响应可能与引导下游阶段的空间输出相配对。监督者阅读语言答案,而分割或区域审查等下游工作流则消费点集输出。我们将这种从审视答案到依赖其点动作的转变称为“语言到动作交接”。当答案正确而配对的动作却遗漏了下游所需的标注对象时,就会发生静默失败,此时基于答案的监督会放行一个动作不可靠的区域。我们提出了 Spatial Action Review,一个用于审计电子显微镜(EM)线粒体分析中此类失败模式的可视化分析仪表板。它通过答案-动作账本、任务-数据集风险图和图像区域审计视图将配对的答案-动作记录关联起来,在将聚合模式与图像证据相连接的同时,可调节的动作可靠性门槛支持重新审计。审查以人机交接结束,监督者记录该动作是被接受、被上报、在更严格的门槛下被搁置,还是被标记以便模型修正。在基于 EM 适配的 Qwen3-VL 案例研究运行中的 541 个图像区域上,在 VQA 响应正确的记录中有 54.4% 的点动作未通过可靠性门槛,且所有记录中有 27.4% 属于静默失败。正确答案仅使可靠动作的概率提高 5.8 个百分点,其自助法(bootstrap)置信区间包含零;答案正确性与对象覆盖之间的点二列相关系数为 0.061。这种弱耦合在 753 个匹配图像区域上的五种模型条件下依然存在。Spatial Action Review 使答案与动作之间的不匹配变得可见,并在 MLLM 输出进入自主科学工作流之前将其与图像证据及记录在案的决策相关联。
cs.CV / 176 / 2609.24482

STA-TFM: Spatio-Temporal Aggregation Across Views TransForMer for Pose Estimation

Kamel, Mena, Won, Natalie, Sarangi, Amrut, Jager, Sven, Planas, Albert Pla
Abstract
Monocular 3D human pose estimation (HPE) remains challenging due to depth ambiguity, occlu- sions, and the need for temporal consistency. While multi-view methods provide superior accuracy over monocular approaches, they often require complex setups. We introduce STA-TFM, a transformer-based architecture that combines spatial and temporal information for multi-view pose estimation. The approach leverages DSTformer, a monocular feature extractor, to capture long-range pose dependencies within each view. A fusion transformer then aggregates information across views to produce coherent 3D estimates. To address training data scarcity, we use a data generation pipeline that transforms any existing 3D pose dataset into multi-view setups with controllable parameters. Experiments on various datasets demonstrate that STA-TFM outperforms existing camera-parameter-free multi-view methods. STA-TFM achieves 50.9% and 49.5% reductions in mean per joint position error (MPJPE) and mean per joint velocity error (MPJVE) on the DHP19 dataset. Furthermore, it achieves 6.7% and 7.7% respective reductions on HAA4D, and a 15.2% MPJPE reduction on TotalCapture. STA-TFM handles noisy and missing 2D inputs, supporting potential deployment in healthcare monitoring, athletic assessment, and immersive technologies. Code, training checkpoints, and data are available at https://zenodo.org/records/22832620.
cs.CV / 177 / 2609.24485

VPRune: Efficient Training-free Pre-LLM Visual Token Pruning

VPRune:高效的无训练 Pre-LLM 视觉token剪枝方法
Lv, Guangchuan, Shi, Dianxing, FU, Dingjie
Abstract
Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We identify three key factors behind this degradation: text-guided selection bias, information loss from discarded tokens, and positional distortion caused by sequence compaction. Based on these observations, we propose \textbf{VPRune}, a training-free pre-LLM pruning framework consisting of visual-only diversity selection, similarity-guided token recycling, and position-preserving restoration. Experiments on FastVLM-1.5B across multiple vision-language benchmarks demonstrate that VPRune achieves a favorable accuracy--compression trade-off, with particularly pronounced advantages under aggressive compression. Furthermore, evaluations on edge-device show that VPRune effectively reduces end-to-end inference latency while maintaining superior task performance, demonstrating its practicality for resource-constrained LVLM deployment.
Chinese Translation
视觉token剪枝是降低大型视觉语言模型(LVLMs)推理成本的一种有前景的方法,然而激进的token缩减往往会导致显著的性能下降。我们识别出导致这种下降的三个关键因素:文本引导的选择偏差、被丢弃token的信息损失,以及序列压缩引起的位置失真。基于这些观察,我们提出了 VPRune,一个无需训练的 Pre-LLM 剪枝框架,由仅视觉的多样性选择、相似度引导的token回收和位置保持恢复三部分组成。在 FastVLM-1.5B 上跨多个视觉语言基准的实验表明,VPRune 实现了良好的精度-压缩权衡,尤其在激进压缩下优势尤为突出。此外,在边缘设备上的评估表明,VPRune 在保持卓越任务性能的同时有效降低了端到端推理延迟,展示了其在资源受限的 LVLM 部署中的实用性。
cs.CV / 178 / 2609.24487

AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos

AgentSTAR:基于智能体的单目视频形状跟踪与重建
Mazur, Kirill, Karaev, Nikita, Chang, Matthew, Malik, Jitendra, Shafiullah, Nur Muhammad "Mahi''
Abstract
In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis. Unlike prior methods which first estimate dense pixel correspondences and then recover object motion from them, our method infers a structured 3D object model, including its geometry and kinematic structure, and uses this model to optimise object track estimates over time. In our optimisation loop, a Vision-Language Model (VLM) agent iteratively refines shape or generalised pose through a render-and-compare loop, combining coarse visual reasoning with numerical pose optimisation for precise state estimation. This structured formulation enables our method to track through large motion, articulation, and severe occlusion without relying on pixel-matching objectives. Quantitatively, on ARCTIC, our method substantially outperforms state-of-the-art 3D point-tracking baselines for articulated objects, and on HOT3D it outperforms all evaluated rigid-object tracking baselines.
Chinese Translation
在本工作中,我们提出了一种通过智能体式的“分析-合成”(analysis-by-synthesis)方法,从视频中重建形状并进行跟踪。与以往先估计密集像素对应关系、再据此恢复物体运动的方法不同,我们的方法推断出一个结构化的三维物体模型,包括其几何结构和运动学结构,并利用该模型随时间优化物体跟踪估计。在我们的优化循环中,一个视觉-语言模型(VLM)智能体通过“渲染-比较”循环迭代地细化形状或广义位姿,将粗粒度的视觉推理与数值位姿优化相结合,以实现精确的状态估计。这种结构化的表述使我们的方法能够在大幅运动、关节活动以及严重遮挡的情况下进行跟踪,而无需依赖像素匹配目标。定量实验表明,在ARCTIC数据集上,我们的方法在铰接物体的跟踪上大幅超越了最先进的三维点跟踪基线方法;在HOT3D数据集上,我们的方法优于所有被评估的刚体跟踪基线方法。
cs.CV / 179 / 2609.24492

Identity-Consistent Analysis of Long-Shot Windsurfing Video: A Domain-Specific Offline Tracking System

Braun, Bertil
Abstract
Long-shot windsurfing video combines small targets, large camera pans, prolonged overlaps, and rapidly changing backgrounds. The desired output is not a generic MOT trace but a separate, stable rider-relative video for each surfer; one false identity merge can invalidate an otherwise useful result. We present an offline analysis system that detects surfers, forms conservative local tracklets, links them globally with camera-compensated motion and a foreground-masked sail-color descriptor, and uses two pose keypoints on the rig to drive a rider-relative virtual camera. The tracking stage is evaluated on 21 manually reconstructed development videos containing 41,004 retained observations. On this fixed-observation protocol, the production system achieves 0.957 pairwise precision, 0.918 recall, and 0.937 F1, compared with 0.792 F1 for OC-SORT and 0.828 for BoT-SORT. Compared with OC-SORT, it reduces fragmentation excess from 845 to 42, but nine of its 95 output tracks mix rider identities and these errors affect seven of the 21 videos.
cs.CV / 180 / 2609.24494

CMAMBADEPTH: Self-supervised Monocular Depth Estimation with Channel Mamba and Hybrid Attention

Xiang, Xuezhi, Liu, Jiayao, Xiang, Heqi, Hu, Yuqi, Chen, Yiming, Zhang, Shanjun
Abstract
Accurate monocular depth estimation serves as a core enabler for single camera scene understanding. However, existing self-supervised monocular depth estimation methods generally suffer from the bottleneck of inefficient cross-scale information interaction and difficulty in balancing local and global spatial modeling. In this paper, we propose CMambaDepth, a self-supervised framework that achieves efficient multi-scale feature fusion and fine-grained contextual modeling via channel-wise selective state propagation. Specifically, Bidirectional Channel Mamba (Bi-CMamba) aligns encoder features across scales and enables bidirectional information exchange among ordered scale groups. Unidirectional Channel Mamba (Uni-CMamba) progressively aggregates decoder features and retains fine-grained scale groups through a group selection mechanism for subsequent fusion. Furthermore, a Hybrid Attention Module (HAM) is introduced to combine large-kernel local context and Manhattan self-attention for complementary spatial modeling. Experimental results demonstrate that our method achieves highly competitive performance. Specifically, our model achieves an AbsRel of 0.094 and an RMSE of 4.156 on KITTI, and an AbsRel of 0.140 on DDAD. In the zero-shot cross-dataset generalization test on NYUv2, it attains an AbsRel of 0.232, outperforming the baseline RA-Depth by 7.2%.
cs.CV / 181 / 2609.24510

0.5\%>100\%: Bidirectional Reciprocal Learning for Referring Image Segmentation

0.5%>100%:面向指代图像分割的双向互惠学习
Lu, Xiaoqiang, Jiao, Licheng, Li, Lingling, Yang, Yuting, Sun, Long, Ma, Wenping, Liu, Xu, Liu, Fang
Abstract
Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks. However, adapting VFMs to referring image segmentation (RIS) typically necessitates precise vision-language alignment via full fine-tuning, incurring substantial computational overhead and risking catastrophic forgetting. While existing parameter-efficient fine-tuning (PEFT) methods enable safe knowledge transfer with minimal training costs, they predominantly operate independently within individual modalities or focus exclusively on unidirectional guidance from language to vision, overlooking progressive cross-modal interaction and visual feedback for textual refinement. To address these limitations, we propose Bidirectional Reciprocal Learning (BRL), a novel adapter-based PEFT framework that facilitates hierarchical, bidirectional information flow within both token-mixing and channel-mixing layers of frozen foundation models. Specifically, BRL introduces two complementary lightweight modules. The Reciprocal Attention Adapter (RAA) performs cross-modal query-key exchanges at the token level, enabling visual and linguistic tokens to mutually attend to each other for fine-grained spatial grounding. The Reciprocal Gate Adapter (RGA) generates cross-modal gating signals at the channel level, allowing global semantic context from one modality to adaptively recalibrate channel activations of the other. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg benchmarks demonstrate the superiority of BRL over prior RIS methods, achieving state-of-the-art performance while requiring less than 0.5% backbone parameter updates. Code and models will be released at https://github.com/xiaoqiang-lu/BRL.
Chinese Translation
视觉基础模型(VFMs)的最新进展在多种单模态视觉任务中展现出了卓越的能力。然而,将VFMs适配到指代图像分割(RIS)任务通常需要通过全量微调来实现精确的视觉-语言对齐,这会带来巨大的计算开销,并存在灾难性遗忘的风险。尽管现有的参数高效微调(PEFT)方法能够以极小的训练成本实现安全的知识迁移,但它们主要在单个模态内独立运作,或仅关注从语言到视觉的单向引导,忽视了渐进式的跨模态交互以及用于文本精炼的视觉反馈。为解决这些局限,我们提出了双向互惠学习(Bidirectional Reciprocal Learning, BRL),这是一种新颖的基于适配器(adapter)的PEFT框架,能够在冻结基础模型的词元混合(token-mixing)层和通道混合(channel-mixing)层中实现层次化的双向信息流动。具体而言,BRL引入了两个互补的轻量级模块:互惠注意力适配器(Reciprocal Attention Adapter, RAA)在词元级别执行跨模态的查询-键交换,使视觉词元和语言词元能够相互关注,从而实现细粒度的空间定位;互惠门控适配器(Reciprocal Gate Adapter, RGA)在通道级别生成跨模态门控信号,使来自一个模态的全局语义上下文能够自适应地重新校准另一模态的通道激活。在RefCOCO、RefCOCO+和RefCOCOg基准上的大量实验证明了BRL相对于先前RIS方法的优越性,在仅需更新不到0.5%主干网络参数的情况下实现了最先进的性能。代码和模型将发布于 https://github.com/xiaoqiang-lu/BRL。
cs.CV / 182 / 2609.24524

Preoperative Prediction of Microvascular Invasion in Hepatocellular Carcinoma by Integrating Multimodal Ultrasound and Clinical Data: A Multicenter Study

融合多模态超声与临床数据的肝细胞癌微血管侵犯术前预测:一项多中心研究
Cheng, Jun, Kong, Yuanyuan, Huang, Qing, Tan, Xiaotong, Dong, Licong, Han, Yulong, Xue, Wufeng, Huang, Ruobing, Ni, Dong, Yang, Qi, Yu, Jie, Liang, Ping
Abstract
Background: Microvascular invasion (MVI) predicts recurrence and survival in hepatocellular carcinoma (HCC) but requires postoperative histopathology for diagnosis. We developed and validated a model integrating multimodal ultrasound and clinical data for preoperative MVI prediction. Methods: This multicenter study included 489 patients with HCC from eight centers. All patients had B-mode ultrasound (BUS), color Doppler flow imaging (CDFI), dynamic contrast-enhanced ultrasound (DCE-US), and clinical information. Data from seven centers (n = 421) were used for model development with five-fold cross-validation; data from the remaining center (n = 68) formed an independent external validation cohort. The proposed multimodal information fusion network used modality-specific encoders, a hemodynamic temporal change module for bidirectional DCE-US perfusion changes, and a representation consistency learning module to align heterogeneous ultrasound representations before Transformer-based fusion. Results: In external validation, DCE-US achieved the highest single-modality area under the receiver operating characteristic curve (AUC; 0.8545+/-0.0198), versus clinical information (0.6715+/-0.0156), CDFI (0.6435+/-0.0344), and BUS (0.6087+/-0.0417). Pixel-difference sampling and the proposed temporal module outperformed alternative sampling and video representation methods. The full model achieved the best performance, with an AUC of 0.8953+/-0.0180, accuracy of 81.18%+/-2.83%, sensitivity of 86.40%+/-6.69%, and specificity of 78.14%+/-6.28. Conclusions: Integrating multimodal ultrasound and clinical information enabled promising preoperative MVI prediction in HCC. DCE-US was the main source of predictive information, while BUS, CDFI, and clinical information provided complementary value. The proposed framework may support preoperative risk stratification and individualized clinical decision-making.
Chinese Translation
背景:微血管侵犯(MVI)是肝细胞癌(HCC)复发和生存的预测因素,但其诊断依赖于术后组织病理学检查。我们开发并验证了一种融合多模态超声与临床数据用于术前预测MVI的模型。方法:这项多中心研究纳入来自八个中心的489例HCC患者。所有患者均接受了B型超声(BUS)、彩色多普勒血流成像(CDFI)、动态对比增强超声(DCE-US)检查并具有临床信息。来自七个中心的数据(n = 421)用于模型开发并进行五折交叉验证;其余一个中心的数据(n = 68)构成独立的 external 验证队列。所提出的多模态信息融合网络采用模态特异性编码器、用于捕捉双向DCE-US灌注变化的血流动力学时间变化模块,以及基于Transformer融合前对异构超声表征进行对齐的表征一致性学习模块。结果:在外部验证中,DCE-US在单模态中取得了最高的受试者工作特征曲线下面积(AUC;0.8545±0.0198),而临床信息为0.6715±0.0156,CDFI为0.6435±0.0344,BUS为0.6087±0.0417。像素差分采样和所提出的时间模块优于其他采样和视频表征方法。完整模型取得了最佳性能,AUC为0.8953±0.0180,准确率为81.18%±2.83%,敏感性为86.40%±6.69%,特异性为78.14%±6.28。结论:融合多模态超声与临床信息可实现有前景的HCC术前MVI预测。DCE-US是预测信息的主要来源,而BUS、CDFI和临床信息提供了补充价值。所提出的框架可支持术前风险分层和个体化临床决策。
cs.CV / 183 / 2609.24526

ME-VLM:A Unified VLM for Embodied Cognition and Agent Coordination

ME-VLM:面向具身认知与智能体协同的统一视觉语言模型
Model, Foundation, Inc, Li Auto
Abstract
Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and execution feedback. We introduce MachEmbodied-VLM (ME-VLM), a unified vision-language model with two variants, 4B and 35B-A3B, that brings together embodied cognition and multimodal agent capabilities. Our work emphasizes physical perception and spatiotemporal reasoning, together with planning, interaction, and outcome assessment in both digital and physical environments. We construct training data spanning embodied and multimodal agent tasks, including execution observations and feedback to support outcome assessment and decision refinement. The training pipeline comprises embodied capability injection, separate reinforcement learning of embodied and multimodal-agent experts, and multi-teacher on-policy distillation that consolidates their complementary capabilities into a single model. Experiments show competitive performance on both embodied and agent benchmarks, as well as on autonomous-driving and embodied-navigation tasks. For edge deployment, visual token compression, W4A8 quantization, and hardware--software co-optimization enable on-device inference of the 4B variant on the M100, reducing prefill latency from 400 ms to 188 ms. Project Page: https://machembodied.com/ME-Brain/ME-VLM.html Code Repository: https://github.com/MachEmbodied/ME-VLM
Chinese Translation
物理人工智能(Physical AI)要求模型将视觉与语言理解扎根于真实世界环境中,同时考虑环境约束与执行反馈。我们提出 MachEmbodied-VLM(ME-VLM),一个包含 4B 和 35B-A3B 两个变体的统一视觉语言模型,将具身认知与多模态智能体能力融为一体。我们的工作强调物理感知与时空推理,以及在数字与物理环境中的规划、交互和结果评估。我们构建了涵盖具身任务与多模态智能体任务的训练数据,包括用于支持结果评估与决策优化的执行观察和反馈。训练流程包括具身能力注入、具身专家与多模态智能体专家的分别强化学习,以及将二者互补能力整合到单一模型中的多教师在线策略蒸馏。实验表明,该模型在具身与智能体基准测试以及自动驾驶和具身导航任务上均具有竞争力。在边缘部署方面,通过视觉 token 压缩、W4A8 量化以及软硬件协同优化,4B 变体可在 M100 上实现端侧推理,将预填充(prefill)延迟从 400 毫秒降低至 188 毫秒。项目主页:https://machembodied.com/ME-Brain/ME-VLM.html 代码仓库:https://github.com/MachEmbodied/ME-VLM
cs.CV / 184 / 2609.24531

Dynamic Thermal Gaussians: Multimodal 4D Gaussian Splatting

动态热高斯:多模态4D高斯泼溅
Lu, Rongfeng, Lin, Lifeng, Wei, Xiaobao, Chen, Quan, Lu, Ming, Xue, Yitian, Sun, Yaoqi, Gao, Yuhan, Xue, Anke, Yan, Chenggang
Abstract
Thermography plays a vital role in military and broader thermal analysis applications. Recent progress in 3D thermal reconstruction has extended temperature analysis from 2D to 3D space, yet most existing works assume static temperature distributions, neglecting the temporal dynamics of heat transfer in real-world environments. To address this limitation, we propose the first dynamic RGB-Thermal reconstruction framework for complex scenes. Our method jointly models RGB appearance, thermal observations, and scene geometry as they change over time. Specifically, we introduce a multimodal dynamic scene representation that anchors both the color and thermal modalities to a shared geometric substrate, ensuring their consistency under spatiotemporal deformations. We further design multimodal embeddings to enhance the motion expressiveness for each modality, and propose a multimodal routing mechanism that retains a unified set of shared multimodal Gaussians as the geometric backbone while adaptively spawning modality-specific Gaussians to strengthen the representational capacity in detail-rich regions of each individual modality. In addition, we contribute a novel benchmark dataset featuring high-frequency temperature variations to facilitate the evaluation of 4D reconstruction. Extensive experiments demonstrate that our method achieves high-fidelity spatiotemporal reconstruction of both appearance and temperature. Our code and dataset are available at: https://github.com/LinLif1869/DTG.
Chinese Translation
热成像技术在军事及更广泛的热分析应用中发挥着至关重要的作用。近年来三维热重建的进展将温度分析从二维空间拓展至三维空间,然而现有大多数工作假设温度分布是静态的,忽视了真实环境中热传导的时间动态特性。为解决这一局限,我们提出了首个面向复杂场景的动态RGB-热成像重建框架。我们的方法联合建模随时间变化的RGB外观、热观测以及场景几何。具体而言,我们引入了一种多模态动态场景表示,将颜色与热两种模态锚定到共享的几何基底上,确保二者在时空形变下的一致性。我们进一步设计了多模态嵌入以增强每种模态的运动表达能力,并提出了一种多模态路由机制,该机制保留一组统一的多模态共享高斯作为几何骨干,同时自适应地生成模态专属的高斯,以增强各模态细节丰富区域的表示能力。此外,我们构建了一个包含高频温度变化的新型基准数据集,以促进4D重建的评估。大量实验表明,我们的方法实现了外观与温度的高保真时空重建。我们的代码和数据集可在以下链接获取:https://github.com/LinLif1869/DTG。
cs.CV / 185 / 2609.24537

MIRAGE: Full-Body Bystander Privacy for Smart Glasses with Consent-Based Restoration

MIRAGE:基于同意恢复机制的智能眼镜全身旁观者隐私保护
Umair, Muhammad, Maqbool, Muhammad Danial, Cheema, Fatima Arshad, Dev, Kapal, Alizai, Muhammad Hamad, Siddiqi, Muhammad Ali, Bhatti, Naveed Anwar
Abstract
Video recording on smart glasses exposes more than faces. Continuous capture reveals full-body biometric signatures, including gait, posture, and silhouette, that enable person re-identification (ReID) even after conventional face sanitization. We present MIRAGE, a three-tier architecture for privacy-preserving smart glasses that enforces full-body privacy, supports synthetic full-body replacement, and retains encrypted recovery material for consent-based restoration. We implement MIRAGE on a Raspberry Pi~5 (a CPU-only proxy for smart-glasses compute), companion phones, and a cloud generative backend. Compared to prior systems, MIRAGE achieves 0.948 AP and 0.976 AR while accurately detecting the complete visible body. Its bounding box masking reduces learned silhouette-based ReID to essentially random guessing, with 10.86% Rank-1 accuracy compared with an 11.12% measured chance level. Even against an adaptive adversary retrained on MIRAGE's sanitized pose signals, Rank-1 gait identification drops from 90.25% to 26.20%, removing 72.5% of the adversary's identification advantage.
Chinese Translation
智能眼镜上的视频录制所暴露的不仅仅是面部。持续拍摄会揭示包括步态、姿态和轮廓在内的全身生物特征信息,即使经过传统的人脸清洗处理,这些信息仍可实现行人重识别(ReID)。我们提出了MIRAGE,一种面向隐私保护智能眼镜的三层架构,能够执行全身隐私保护、支持合成的全身替换,并保留加密的恢复材料以实现基于同意的恢复。我们在Raspberry Pi 5(作为智能眼镜计算能力的纯CPU代理)、配套手机以及云端生成后端上实现了MIRAGE。与先前系统相比,MIRAGE在准确检测完整可见人体的同时,实现了0.948的AP和0.976的AR。其边界框掩码将基于轮廓学习的行人重识别能力降低至接近随机猜测水平,Rank-1准确率仅为10.86%,而实测的随机猜测水平为11.12%。即使面对在MIRAGE清洗后的姿态信号上重新训练的自适应攻击者,基于步态的Rank-1识别率也从90.25%降至26.20%,消除了攻击者72.5%的识别优势。
cs.CV / 186 / 2609.24539

Incentive Noise and Structural Prior Infusion for Multi-modal Object Re-Identification

面向多模态目标重识别的激励噪声与结构先验注入
Zhou, Weixiang, Wang, Yuhao, Xu, Xingguo, Zhou, Weizhen, Su, Zhixun, Pan, Jinshan, Wang, Cong
Abstract
Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic representation, text descriptions have recently been incorporated as an additional modality. However, recent vision-language approaches often treat text descriptions as clean, deterministic signals and overlook their inherent noise, including modality-mismatched phrases and semantically ambiguous expressions. Moreover, prevailing methods lack explicit mechanisms to reconcile fine-grained structural discrepancies between modalities, even after high-level semantic alignment. To address these challenges, we propose a novel framework centered on Positive-Incentive Noise ({\pi}-noise) and structured prompt modulation. First, the Semantic Cross-Modal Modulator harnesses task-aware {\pi}-noise, sampled from a distribution conditioned on both visual and text inputs, to perturb global tokens and enable semantics-guided cross-modal compensation. Second, the Structure-Aware Prompt Adapter injects learnable geometric priors via prompts to enhance spatial consistency. Third, the Context-Aware Sparse Fusion module distills structural context to guide adaptive fusion while shielding identity features from noisy local details. Experiments on three multi-modal ReID benchmarks demonstrate the effectiveness and robustness of our approach. The code is available at https://github.com/zw-absin/INSPI.
Chinese Translation
多模态目标重识别(ReID)受益于异构成像模态之间的互补信息。为了进一步丰富语义表示,近期研究将文本描述作为额外的模态引入。然而,近期的视觉-语言方法通常将文本描述视为干净、确定的信号,而忽略了其固有的噪声,包括模态不匹配的短语和语义模糊的表达。此外,现有方法缺乏显式机制来协调模态之间细粒度的结构差异,即使在高层语义对齐之后亦是如此。为应对这些挑战,我们提出了一个以正激励噪声(Positive-Incentive Noise,π-noise)和结构化提示调制为核心的新型框架。首先,语义跨模态调制器(Semantic Cross-Modal Modulator)利用任务感知的π-noise——从以视觉和文本输入为条件的分布中采样——来扰动全局令牌(token),实现语义引导的跨模态补偿。其次,结构感知提示适配器(Structure-Aware Prompt Adapter)通过提示注入可学习的几何先验,以增强空间一致性。第三,上下文感知稀疏融合模块(Context-Aware Sparse Fusion)提炼结构上下文以指导自适应融合,同时保护身份特征免受噪声局部细节的干扰。在三个多模态ReID基准上的实验证明了我们方法的有效性和鲁棒性。代码已发布于 https://github.com/zw-absin/INSPI。
cs.CV / 187 / 2609.24560

Evaluating Transformation Models for pCLE Mosaic Registration

用于pCLE马赛克图像配准的变换模型评估
Aboelela, Ahmed, Barcsay, Johannes, Friedhof, Jana, Gonçalves, Miguel, Hann, Alexander, Breininger, Katharina
Abstract
Confocal Laser Endomicroscopy (CLE) provides real-time, cellular-resolution optical biopsy but has a narrow field of view, which image mosaicing can extend to provide anatomical context. Because of line-by-line acquisition, probe motion, and probe-tissue interaction, frame alignment generally requires a non-linear transformation whose accuracy is difficult to quantify: flexible transformation models can fit intensity features and noise, so appearance-based metrics such as Normalized Cross-Correlation (NCC) can improve without a genuine gain in geometric accuracy. We therefore establish a dataset of 132 frame pairs across fourteen pCLE sequences from 4 patients with manually annotated landmark correspondences, so that Target Registration Error (TRE) can serve as a geometrically grounded complement to NCC. We assess the effect of progressively increasing the transformation model's degrees of freedom, from translation to Thin Plate Spline (TPS), and of six feature-matching backends spanning classical (Shi-Tomasi, Lucas-Kanade) and learned (SuperPoint, SuperGlue, LightGlue, LoFTR, RoMa) approaches. Translation and rigid models prove insufficient under tissue deformation, while TPS with random sampling achieves the strongest landmark-derived alignment of the evaluated configurations; among the learned matchers, used without fine-tuning, only RoMa offers a robust, if modest, advantage over other methods. At the sequence level, pairwise registration quality proved an unreliable predictor of final mosaic quality, so mosaic quality must be evaluated directly rather than inferred from pairwise metrics.
Chinese Translation
共聚焦激光显微内镜(CLE)可提供实时的、细胞分辨率的光学活检,但其视野较窄,而图像马赛克拼接可以扩展视野以提供解剖学背景。由于逐行采集、探头运动以及探头与组织的相互作用,帧对齐通常需要非线性变换,而其精度难以量化:灵活的变换模型可能同时拟合强度特征与噪声,因此基于外观的指标(如归一化互相关,NCC)可能在不带来真实几何精度提升的情况下改善。为此,我们建立了一个包含来自4名患者的14个pCLE序列、共132个帧对的数据集,并进行了人工标注的地标对应点,使目标配准误差(TRE)能够作为NCC的几何学补充。我们评估了逐步增加变换模型自由度(从平移到薄板样条,TPS)的影响,以及涵盖经典方法(Shi-Tomasi、Lucas-Kanade)和基于学习的方法(SuperPoint、SuperGlue、LightGlue、LoFTR、RoMa)的六种特征匹配后端。结果表明,平移和刚性模型在组织形变下不够充分,而采用随机采样的TPS在所评估配置中取得了最强的基于地标的对齐效果;在未经微调的基于学习的匹配器中,只有RoMa相较其他方法展现出稳健但有限的优势。在序列层面,逐对配准质量被证明无法可靠地预测最终马赛克图像质量,因此马赛克质量必须直接评估,而不能从逐对指标中推断。
cs.CV / 188 / 2609.24564

HyperCLIP++: Fine-tuning CLIP forOpen-vocabulary Semantic Segmentation in Hyperbolic Space

HyperCLIP++:在双曲空间中微调CLIP以实现开放词汇语义分割
Peng, Zelin, Xu, Zhengqin, Wen, Changsong, Huang, Yu, Wang, Yaoming, Yang, Xiaokang, Shen, Wei
Abstract
CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encoder is known to preserve its generalization capability, recent studies show that fine-tuning both CLIP's text and image encoders jointly significantly enhances segmentation performance, especially for classes from open sets. In this work, we explain this phenomenon from the perspective of hierarchy alignment, since during fine-tuning, the hierarchical level of image embeddings shifts from image-level to pixel-level. We achieve this by leveraging hyperbolic space, which naturally encodes hierarchical structures. Our key observation is that, during fine-tuning, the hyperbolic radius of CLIP's text embeddings decreases, facilitating better alignment with the pixel-level granularity of visual data. Building on this, we propose HyperCLIP++, a novel and parameter-efficient adaptation strategy. HyperCLIP++ directly adjusts the hyperbolic radius of CLIP's embeddings via scaling transformations to achieve a hierarchy alignment to the target task, i.e., segmentation. To ensure this hierarchy alignment is effected consistently across both modalities and preserves their cross-modal alignment during training, HyperCLIP++ integrates a Dual Cross-Relation Communication (DCRC) module that synchronizes these adjustments between the vision and text pathways. Our experiments show that HyperCLIP++ achieves state-of-the-art performance across three benchmarks while fine-tuning only approximately 5% of CLIP's total parameters. More importantly, we observe that after adjustment, CLIP's text embeddings exhibit a relatively fixed hyperbolic radius across datasets, suggesting that the hierarchical level required for this segmentation task might be quantified using the hyperbolic radius.
Chinese Translation
CLIP作为一种基础视觉-语言模型,已成为开放词汇语义分割的强大工具。尽管冻结CLIP的文本编码器被认为可以保留其泛化能力,但近期研究表明,同时对CLIP的文本编码器和图像编码器进行联合微调能显著提升分割性能,尤其是对于来自开放集的类别。在本工作中,我们从层次对齐的视角解释了这一现象:在微调过程中,图像嵌入的层次级别从图像级转变为像素级。我们通过利用双曲空间来实现这一点,因为双曲空间天然适合编码层次结构。我们的关键观察是,在微调过程中,CLIP文本嵌入的双曲半径会减小,从而促进其与视觉数据像素级粒度的更好对齐。基于此,我们提出了HyperCLIP++,一种新颖且参数高效的适配策略。HyperCLIP++通过缩放变换直接调整CLIP嵌入的双曲半径,以实现与目标任务(即分割)的层次对齐。为确保这种层次对齐在两种模态中保持一致,并在训练过程中保持它们的跨模态对齐,HyperCLIP++集成了一个双交叉关系通信(Dual Cross-Relation Communication, DCRC)模块,用于在视觉通路和文本通路之间同步这些调整。实验表明,HyperCLIP++在仅微调约5%的CLIP总参数的情况下,在三个基准测试上取得了最先进的性能。更重要的是,我们观察到调整后CLIP的文本嵌入在不同数据集上呈现出相对固定的双曲半径,这表明该分割任务所需的层次级别或许可以通过双曲半径来量化。
cs.CV / 189 / 2609.24565

Active Visual Sampling with a Connectome-Constrained Fly Model for One-Shot Hatch Recognition in Architectural Drawings

基于连接组约束果蝇模型的活动视觉采样用于建筑图纸中单样本填充图案识别
Kuklev, Dmitry
Abstract
Architectural drawings encode material classes through repeated hatch patterns. We test whether a connectome-constrained fly visual network, pretrained for motion, can be repurposed without task-specific weight updates as a descriptor for one-shot hatch matching. Each 64 x 64 patch is translated over eight scan trajectories and summarized across 57 cell types; query descriptors are then matched to one legend strip per class. On 400 development sheets from a synthetic benchmark built on CubiCasa5K geometry, the frozen fly pipeline reaches 0.857 area-weighted accuracy and 0.910 with an extended legend. On an equal-brightness orientation condition it reaches 0.840 versus 0.299 for eleven pixel statistics, while a Gabor bank reaches 0.900. Replacing drift with a repeated still frame lowers the combined equal-condition score by 0.089 [0.066, 0.112]. However, a receptors-only descriptor reaches 0.891 and a task-trained 5,888-parameter CNN averages 0.959, so the current evidence supports transfer and the usefulness of active sampling, but not an advantage of the biological wiring. We separate project-recorded results from recomputed checks and report a small real-drawing audit. The supported claim is therefore narrow: motion-oriented biological vision can be repurposed as a useful texture representation for architectural hatch matching, while the topology contribution and end-to-end BIM utility remain open questions.
Chinese Translation
建筑图纸通过重复的填充(hatch)图案编码材料类别。我们测试了一个为运动预训练的连接组约束果蝇视觉网络,能否在不进行任务特定权重更新的情况下被重新用作单样本(one-shot)填充图案匹配的描述子。每个64×64图像块沿八条扫描轨迹平移,并在57种细胞类型上进行汇总;然后将查询描述子与每类一个图例条带进行匹配。在基于CubiCasa5K几何结构构建的合成基准的400张开发图纸上,冻结的果蝇流水线达到0.857的面积加权准确率,使用扩展图例时达到0.910。在等亮度方向条件下,其准确率达到0.840,而十一种像素统计方法仅为0.299,Gabor滤波器组则达到0.900。用重复的静止帧替代漂移会使组合等条件分数降低0.089 [0.066, 0.112]。然而,仅使用受体(receptors-only)的描述子可达到0.891,任务训练的5,888参数CNN平均达到0.959,因此当前证据支持迁移学习的有效性以及主动采样的有用性,但不支持生物连接结构的优势。我们区分了项目记录结果与重新计算的校验结果,并报告了对真实图纸的小规模审计。因此,可支持的结论较为有限:面向运动的生物视觉可以被重新用作建筑填充图案匹配的有用纹理表示,而拓扑结构的贡献以及端到端BIM实用性仍是待解决的问题。
cs.CV / 190 / 2609.24595

Applications of Neural Cellular Automata: State of the Art, Challenges and Opportunities

神经元胞自动机的应用:现状、挑战与机遇
Lemke, Nick, Ihm, Niklas, Kalkhof, John, Konstantin, Mirko, Krumb, Henry J., Lang, Daniel M., Pajouheshagar, Ehsan, Sadafi, Ario, Kuijper, Arjan, Lekadir, Karim, Lorenzi, Marco, Marr, Carsten, Schnabel, Julia A., Mukhopadhyay, Anirban
Abstract
Neural Cellular Automata (NCAs) are a new type of neural network architecture which enable accurate and robust inference at extremely small model sizes. Recently, NCAs have advanced to become interesting low-resource alternatives to convolution- and attention-based architectures for various tasks such as image analysis, synthetic image generation, and simulation. The rapid development and increased research interest necessitate a comprehensive review of the emerging technology. This review provides an overview of the fundamentals of NCAs, applications to medical imaging, as well as insights into the state of the art. We analyze recent modifications to the originally proposed NCA architecture with respect to their efficiency and accuracy. Furthermore, we review practical applications in real-world scenarios with a focus on medical image analysis, segmentation, classification, registration, depth estimation, and image synthesis. Finally, we identify several advantages of NCAs, research gaps, and conclude with an analysis of future opportunities for NCAs in medical applications in confined settings or areas that have particular demands for robustness or efficient data processing.
Chinese Translation
神经元胞自动机(Neural Cellular Automata,NCAs)是一种新型神经网络架构,能够在极小的模型规模下实现准确且鲁棒的推理。近年来,NCAs已发展成为基于卷积和注意力机制架构的低资源替代方案,可应用于图像分析、合成图像生成以及仿真等多种任务。该技术的快速发展和研究兴趣的日益增长,使得对其进行全面综述变得十分必要。本文综述了NCAs的基本原理、在医学影像中的应用,以及该领域的最新研究进展。我们分析了针对最初提出的NCA架构的最新改进方案,重点关注其效率与准确性。此外,我们综述了NCAs在现实场景中的实际应用,重点涵盖医学图像分析、分割、分类、配准、深度估计和图像合成等方面。最后,我们总结了NCAs的若干优势与研究空白,并分析了其在受限环境或对鲁棒性及高效数据处理有特殊需求的医学应用领域中的未来发展机遇。
cs.CV / 191 / 2609.24612

Beyond Uniform Subspaces: Spectrum-Aware and Depth-Adaptive Fusion for Multi-Task Model Merging

超越均匀子空间:面向多任务模型合并的谱感知与深度自适应融合
Gu, Ruxi, Wang, Zilei, Wang, Wei
Abstract
Model merging aims to consolidate multiple task-specific models without access to extra training process. However, existing subspace-based methods largely rely on a uniform treatment of task updates, overlooking their intrinsic spectral and depth-wise heterogeneity. We identify two key deviations from this assumption: different tasks require different subspace capacity and exhibit different tolerance to spectral transformation, while subspace projection introduces depth-dependent distortion. Based on these observations, we propose SADA-Merging, a spectrum-aware and depth-adaptive framework for data-free model merging. SADA-Merging allocates task-specific subspace capacity according to spectral complexity, adapts spectral preservation according to task-wise plasticity, and applies depth-dependent anchoring to compensate for projection-induced distortion. This enables the fusion process to adapt to both the intrinsic geometry of each task and its sensitivity across network depth. SADA-Merging operates directly on task updates and is applicable to both full fine-tuning and LoRA settings. Extensive experiments demonstrate consistent improvements over existing data-free merging methods across different task scales and adaptation settings.
Chinese Translation
模型合并(Model Merging)旨在无需额外训练过程的条件下,将多个面向特定任务的模型整合为一个模型。然而,现有的基于子空间的方法在很大程度上依赖于对任务更新的统一处理,忽视了其内在的谱异质性和逐层深度异质性。我们发现该假设存在两个关键偏差:不同任务需要不同的子空间容量,并对谱变换表现出不同的容忍度;同时,子空间投影会引入与深度相关的失真。基于这些观察,我们提出了 SADA-Merging,一个谱感知且深度自适应的无数据模型合并框架。SADA-Merging 根据谱复杂度为各任务分配特定的子空间容量,根据任务可塑性自适应地调整谱保留策略,并应用深度相关的锚定机制以补偿投影引起的失真。这使得融合过程能够同时适应每个任务的内在几何结构及其在网络深度上的敏感性。SADA-Merging 直接作用于任务更新,可同时适用于全量微调和 LoRA 两种设置。大量实验表明,在不同的任务规模和适配设置下,SADA-Merging 相较于现有的无数据合并方法取得了一致的性能提升。
cs.CV / 192 / 2609.24619

Video-based Surgical Skill Assessment Using Dynamics-and-Uncertainty-Aware Tree-based Gaussian Process Classifier

基于动态性与不确定性感知树状高斯过程分类器的视频手术技能评估
Rezaei, Arefeh, Ahmadi, Mohammad Javad, Molaei, Amir, Taghirad, Hamid D.
Abstract
The proposed pipeline integrates a representation-flow convolutional neural network with a dynamics- and uncertainty-aware tree-based Gaussian Process classifier. In this framework, latent motion dynamics are exploited both as discriminative representations and as a source of input uncertainty, enhancing robustness against temporal variations and abnormal motion transitions. Compared with conventional deep learning approaches, the proposed strategy requires less training data and offers improved computational efficiency. To further improve classification performance, we introduce novel semantic-aware compound kernels that effectively capture semantic, flow, and dynamic information embedded in surgical video features. In addition, uncertainty-aware kernels are developed to strengthen the robustness and practical applicability of the compound kernel framework. The proposed method is evaluated on two benchmark datasets, namely the JIGSAWS and the Cataract-LMM (Capsulorhexis) datasets. Experimental results demonstrate strong performance across both datasets, including the LOSO and LOUO evaluation protocols on JIGSAWS, including the subject-independent LOUO protocol on JIGSAWS, on which the framework attains a mean accuracy of \ph{96.9}\%; results under the within-subject LOSO protocol are reported for comparability with prior work, achieving competitive accuracy while substantially reducing computational cost. Overall, the proposed pipeline provides an efficient and accurate framework for video-based surgical skill assessment.
Chinese Translation
本文提出的处理流程将表征流卷积神经网络(representation-flow CNN)与动态性和不确定性感知的树状高斯过程分类器相集成。在该框架中,潜在运动动态既被用作判别性表征,又被用作输入不确定性的来源,从而增强了对时间变化和异常运动过渡的鲁棒性。与传统的深度学习方法相比,所提出的策略所需的训练数据更少,并具有更高的计算效率。为进一步提升分类性能,我们引入了新颖的语义感知复合核(semantic-aware compound kernels),能够有效捕捉手术视频特征中蕴含的语义、光流和动态信息。此外,我们还开发了不确定性感知核,以增强复合核框架的鲁棒性和实际适用性。所提出的方法在两个基准数据集上进行了评估,即 JIGSAWS 和 Cataract-LMM(撕囊术)数据集。实验结果表明,该方法在两个数据集上均表现出色:在 JIGSAWS 上,该框架在独立于受试者的 LOUO(留一用户)评估协议下达到了平均 96.9% 的准确率;同时,为与已有工作进行可比性对照,本文还报告了受试者内 LOSO(留一受试者)协议下的结果,在取得具有竞争力的准确率的同时显著降低了计算成本。总体而言,所提出的流程为基于视频的手术技能评估提供了一个高效且准确的框架。
cs.CV / 193 / 2609.24626

Relationally Grounded Latent World Models for Autonomous Driving

面向自动驾驶的关系锚定潜在世界模型
Schmidt, Fabian, Enzweiler, Markus, Valada, Abhinav
Abstract
Latent world models learn predictive representations for autonomous driving, but the relational semantics these states preserve often remain implicit. We investigate whether traffic scene graphs can serve as privileged semantic supervision for latent world representations. Building on LAW, we construct actor-centric scene graphs from nuScenes 3D annotations, encode their serialized relational structure using a frozen text embedding model, and align the visual latent representations with this semantic target during training. We remove the supervision branch at inference, so it requires neither scene graphs nor 3D annotations and adds no test-time computation. On nuScenes, our method reduces average trajectory L2 error from 0.661 to 0.622 (5.9%) and collision rate from 0.456 to 0.217 (52.4%) relative to our retrained LAW baseline. It also outperforms an unstructured caption-style semantic target, supporting the benefit of explicit relational structure for latent world-model representation learning.
Chinese Translation
潜在世界模型为自动驾驶学习预测性表征,但这些状态所保留的关系语义往往仍是隐式的。我们研究了交通场景图能否作为潜在世界表征的特权语义监督信号。在 LAW(LAW)的基础上,我们基于 nuScenes 3D 标注构建以参与者为中心的场景图,使用冻结的文本嵌入模型对其序列化的关系结构进行编码,并在训练过程中将视觉潜在表征与该语义目标进行对齐。在推理阶段我们移除监督分支,因此该方法既不需要场景图也不需要 3D 标注,且不增加任何测试时计算量。在 nuScenes 数据集上,与重新训练的 LAW 基线相比,我们的方法将平均轨迹 L2 误差从 0.661 降至 0.622(降低 5.9%),碰撞率从 0.456 降至 0.217(降低 52.4%)。该方法还优于非结构化的描述式(caption-style)语义目标,这支持了显式关系结构对潜在世界模型表征学习的益处。
cs.CV / 194 / 2609.24627

FedMust: Semi-supervised Multi-task Student-Teacher Federated Learning for Multi-organ CT Segmentation

FedMust:用于多器官CT分割的半监督多任务师生联邦学习
Moradi, Ashkan, Abrahamsen, Bendik Skarre, Elschot, Mattijs
Abstract
Multi-organ segmentation using deep learning requires large amounts of annotated patient data; however, institutions often lack sufficiently large and diverse annotated datasets. Privacy constraints further prevent institutions from sharing patient data to overcome this limitation. Moreover, due to the labor-intensive nature of annotation and the scarcity of diverse expertise, institutions typically have labels for only a small portion of their local data, leaving the larger unlabeled portion unused. In this work, we propose a flexible semi-supervised federated multi-task student-teacher framework that leverages federated learning (FL) to improve multi-organ segmentation using both labeled and unlabeled data across participating sites. At each communication round, the proposed framework initiates local training, where clients with labels for the same task form a federation to produce an aggregated teacher model. The resulting teachers generate task-specific features for all data at each client. Subsequently, all clients form a second federation to train a multi-task student model with a shared encoder and task-specific decoders that replicate the teacher-generated features across all segmentation tasks. The aggregated student model is then used to update the local teachers and initiate the next training round. Extensive experiments demonstrated the effectiveness of the proposed method compared with local and federated single-organ models, yielding an average performance gain of 13 percent across clients. The experiments also demonstrated the impact of multi-task learning and unlabeled data and the applicability of the framework in relaxing labeled-data requirements for client participation. The code is available at https://github.com/AshknMrd/FedMust.
Chinese Translation
基于深度学习的多器官分割需要大量标注的患者数据;然而,各医疗机构往往缺乏足够规模和多样性的标注数据集。隐私限制进一步阻碍了机构之间共享患者数据以克服这一局限。此外,由于标注工作的劳动密集性以及专业知识的稀缺性,机构通常仅对本地数据的一小部分拥有标签,而较大一部分未标注数据则被闲置。在本工作中,我们提出了一种灵活的半监督联邦多任务师生框架,利用联邦学习(FL)结合参与站点的已标注和未标注数据来改进多器官分割。在每一轮通信中,所提出的框架启动本地训练,其中拥有相同任务标签的客户端组成一个联邦,生成聚合的教师模型。所得的教师模型为每个客户端的所有数据生成任务特定的特征。随后,所有客户端组成第二个联邦,训练一个多任务学生模型,该模型具有共享编码器和任务特定解码器,可在所有分割任务上复现教师生成的特征。聚合后的学生模型随后用于更新本地教师模型并启动下一轮训练。大量实验表明,与本地及联邦单器官模型相比,所提出的方法具有显著有效性,在各客户端上平均性能提升达13%。实验还证明了多任务学习和未标注数据的作用,以及该框架在放宽客户端参与所需标注数据要求方面的适用性。代码已发布于 https://github.com/AshknMrd/FedMust。
cs.CV / 195 / 2609.24634

High-resolution Nitrogen Dioxide Maps Reveal Exposure Limit Breaches across Europe

高分辨率二氧化氮地图揭示欧洲各地的暴露限值超标情况
Scheibenreif, Linus, Schindler, Konrad
Abstract
Nitrogen dioxide (NO2) is a common air pollutant, released into the atmosphere through the incomplete burning of fossil fuels, and associated with respiratory and cardiovascular diseases in humans. Ambient NO2 concentrations are regulated through air-quality limits assessed with a sparse network of fixed monitors. The revised EU Ambient Air Quality Directive (2024/2881) introduces a daily NO2 limit to be met from 2030. At present, neither the regulatory monitoring network nor existing coarse, annual-mean models can resolve NO2 concentrations at the spatio-temporal resolutions necessary to assess compliance. Here we map NO2 across Europe at hourly and 10m resolution with a machine-learning model that combines ground monitors with satellite, reanalysis, land-use, traffic and emission data and returns a calibrated predictive distribution at every location. Validated against held-out regulatory monitors and independent citizen-science campaigns, the maps resolve high-resolution spatiotemporal NO2 gradients for 110 metropolitan areas in Europe. We reconstruct the daily compliance statistic across those regions and find limit breaches in 91 EU air quality zones deemed compliant by the regulatory monitoring network, covering a population of approximately 135M. Beyond air quality zones and monitor locations, an estimated 9-9.4% (20M) of the population in mapped regions lives in areas where the daily NO2 limit is breached. The high-resolution maps offer a route to population-scale assessment of compliance with the 2030 limits.
Chinese Translation
二氧化氮(NO2)是一种常见的空气污染物,通过化石燃料的不完全燃烧释放到大气中,并与人类呼吸系统和心血管疾病相关。环境NO2浓度通过空气质量限值进行监管,而限值评估依赖于稀疏的固定监测站网络。修订后的欧盟《环境空气质量指令》(2024/2881)引入了一项须在2030年前达标的NO2日均限值。目前,无论是监管监测网络还是现有的粗分辨率年均模型,都无法在评估合规性所需的时空分辨率上解析NO2浓度。本文利用一个结合地面监测站数据与卫星、再分析、土地利用、交通和排放数据的机器学习模型,以逐小时、10米分辨率绘制了欧洲范围的NO2浓度地图,并在每个位置给出经校准的预测分布。经留出的监管监测站及独立的公民科学观测活动验证,这些地图解析了欧洲110个大都市区的高时空分辨率NO2浓度梯度。我们重构了这些区域的日均合规统计,发现91个被监管监测网络判定为合规的欧盟空气质量区存在限值超标,涉及人口约1.35亿。在空气质量区和监测站位置之外,估计被绘制区域中9-9.4%(约2000万)的人口生活在日均NO2限值超标的地区。这些高分辨率地图为在人口尺度上评估2030年限值合规性提供了一条途径。
cs.CV / 196 / 2609.24668

Ev-YOLO: Uncertainty-Aware Object Detection via a Unified Evidential Formulation

Ev-YOLO:基于统一证据性公式化的不确定性感知目标检测
Barbarit-Gaboriau, Simon, Laghmara, Hind, Boutteau, Rémi, Ainouz, Samia
Abstract
Reliable uncertainty estimation is essential for deploying object detectors in autonomous systems operating in uncertain environments. Evidential Deep Learning (EDL) provides a principled framework for uncertainty-aware classification by representing network outputs as evidence and interpreting predictions through subjective logic. However, existing evidential object detectors typically combine evidential classification with regression uncertainty models that do not share the same theoretical foundation. In this work, we propose an evidential version of YOLOv8 in which both classification and bounding-box regression are formulated within a common evidential framework. Our approach exploits YOLOv8's distribution-based bounding-box representation, allowing the evidential formulation to be applied not only to classification but also to localisation. As a result, both tasks produce belief, uncertainty, and probability estimates that can be interpreted within the Dempster--Shafer framework. Experiments on KITTI, MUSES, and nuScenes show that the resulting detector remains broadly competitive with standard YOLOv8 in terms of detection accuracy while providing a localisation uncertainty that effectively discriminates between correct and erroneous detections. Moreover, this uncertainty becomes increasingly discriminative under domain shift.
Chinese Translation
可靠的不确定性估计对于在不确定环境中运行的自主系统中部署目标检测器至关重要。证据深度学习(Evidential Deep Learning, EDL)通过将网络输出表示为证据并借助主观逻辑(subjective logic)解释预测,为不确定性感知分类提供了一个有原则的框架。然而,现有的证据性目标检测器通常将证据性分类与不共享同一理论基础的不确定性回归模型相结合。在本工作中,我们提出了YOLOv8的证据性版本,其中分类和边界框回归均在统一的证据性框架内进行公式化。我们的方法利用了YOLOv8基于分布的边界框表示,使证据性公式化不仅能应用于分类,还能应用于定位。由此,这两个任务都能产生可在Dempster--Shafer框架内解释的信念(belief)、不确定性和概率估计。在KITTI、MUSES和nuScenes数据集上的实验表明,所得检测器在检测精度方面总体上与标准YOLOv8具有竞争力,同时其定位不确定性能够有效区分正确与错误的检测结果。此外,在域偏移(domain shift)条件下,这种不确定性的区分能力会不断增强。
cs.CV / 197 / 2609.24691

What Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization

什么才是好的医学图像分词器?重新思考医学图像分词中的重建与生成
Bubeck, Niklas, Zhang, Yundi, Sideri-Lampretsa, Vasiliki, McGinnis, Julian, Yang, Jiancheng, Rueckert, Daniel, Pan, Jiazhen
Abstract
Latent diffusion models now dominate medical image generation, and every such pipeline rests on a \emph{tokenizer} that compresses images into the latent codes for image generation to operate on. Thereby, the tokenizer choice bounds every downstream task from reconstruction fidelity and generation quality to the representations available for downstream analysis. Yet, medical imaging pipelines routinely utilize tokenizers from natural imaging on the hypothesis that their behavior carries over. However, this is an assumption never tested in the medical imaging regime, where datasets are orders of magnitude smaller and images exhibit far lower inter-sample variance. We present a systematic evaluation of medical image tokenizers evaluating thirty configurations across ten model families on twelve datasets at three compression factors, spanning reconstruction, generation, latent geometry, downstream classification, and memorization. We find that (1) performance on image reconstruction and generation strongly correlate, unlike prior reports on natural images; (2) modern tokenizers use nearly all of their codebook entries, but still leave most of the latent space unused; (3) training-set memorization is mild and is further suppressed by stronger latent space compression; and (4) discrete quantization can largely preserve downstream classification, with lookup-free schemes being the main exception.
Chinese Translation
潜在扩散模型(latent diffusion models)如今在医学图像生成中占据主导地位,而此类流程均依赖于一个将图像压缩为潜在编码的分词器(tokenizer),图像生成即在该潜在空间上进行。因此,分词器的选择决定了所有下游任务的表现,包括重建保真度、生成质量以及可用于下游分析的表征。然而,医学影像流程通常沿用自然图像领域的分词器,其假设是分词器的行为特性可以迁移。但这一假设在医学影像领域从未被验证过——医学影像数据集规模要小几个数量级,且图像的样本间方差也远低于自然图像。我们对医学图像分词器进行了系统性评估,在十二个数据集上以三种压缩率评估了十个模型家族的三十种配置,涵盖重建、生成、潜在空间几何结构、下游分类以及记忆化(memorization)等方面。我们发现:(1)图像重建与生成性能高度相关,这与以往在自然图像上的报告不同;(2)现代分词器几乎用尽了其码本(codebook)中的所有条目,但潜在空间的大部分区域仍未被利用;(3)训练集记忆化程度较轻,且更强的潜在空间压缩可进一步抑制记忆化;(4)离散量化基本可以保持下游分类性能,但无查找(lookup-free)方案是主要例外。
cs.CV / 198 / 2609.24727

ReSTI: A Source-Grounded Audit and Repair of STI-Bench

ReSTI:对STI-Bench基于数据源的审计与修复
Sun, Pengzhan, Rajaraman, Ramanathan, Kao, Shiu-Hong, Xiao, Junbin, Yao, Angela
Abstract
Spatial--temporal benchmarks are valid only when their questions, source annotations, and answer options identify the same physical quantity. We audit STI-Bench against the official ScanNet, Waymo, and Omni6DPose sources and find systematic coordinate-system and timestamp errors, under-specified targets and times, and disagreements between keyed options and answer details. We introduce ReSTI, a source-backed revision that reconstructs every recoverable answer under an explicit target, time, coordinate system, physical quantity, and unit. Source reconstruction reveals task-level geometric failures: ScanNet Grounding omits the required alignment between annotation and raw camera coordinate systems, while Orientation measures camera rotation on the wrong plane. ReSTI replaces these labels with explicit, source-consistent geometric definitions and corrects other source-verifiable defects, including Waymo poses evaluated at the wrong timestamp. Across 2,064 legacy questions, ReSTI retains 1,782 questions and records 282 evidence-backed exclusions. ReSTI therefore provides a conservative and source-traceable basis for evaluating precise video spatial--temporal reasoning. Project page: https://github.com/pengzhansun/ReSTI.
Chinese Translation
时空基准测试只有在问题、源标注和答案选项指向同一物理量时才有效。我们将STI-Bench与官方ScanNet、Waymo和Omni6DPose数据源进行对照审计,发现了系统性的坐标系和时间戳错误、目标与时间设定不明确,以及标注选项与答案细节之间的不一致。我们提出ReSTI,一种基于数据源的修订方案,它在明确的目标、时间、坐标系、物理量和单位条件下重建每一个可恢复的答案。数据源重建揭示了任务层面的几何定义缺陷:ScanNet Grounding任务遗漏了标注坐标系与原始相机坐标系之间所需的转换对齐,而Orientation任务在错误的平面上测量相机旋转。ReSTI用显式的、与数据源一致的几何定义替换了这些标签,并修正了其他可通过数据源验证的缺陷,包括在错误时间戳上评估的Waymo位姿。在2,064个原有问题中,ReSTI保留了1,782个问题,并记录了282个有据可查的排除项。因此,ReSTI为评估精确的视频时空推理提供了一个保守且可溯源至原始数据的基础。项目主页:https://github.com/pengzhansun/ReSTI。
cs.CV / 199 / 2609.24732

GraphSVR: q-Space--Aware Graph-Based Slice-to-Volume Registration for Diffusion MRI

GraphSVR:面向扩散MRI的q空间感知的基于图的切片到体配准
Kertes, Noga, Sourani, Daphna Link, Bronstein, Alex M., Freiman, Moti
Abstract
Diffusion-weighted imaging (DWI) remains highly vulnerable to subject motion, particularly in time-efficient protocols and in motion-prone populations. While slice-to-volume registration (SVR) can mitigate inter-slice and inter-stack misalignment, diffusion MRI introduces additional complexity due to diffusion-direction-dependent contrast and the requirement to align dozens of measurements within a common reference frame, effectively yielding a 4D registration problem. Existing approaches rely primarily on sequential modeling or pairwise similarity and often degrade under sparse gradient sampling or severe motion. We introduce GraphSVR, a q-space-aware graph-based framework for 4D SVR registration in DWI. GraphSVR represents slice groups as nodes in an acquisition-structured graph, with edges encoding temporal proximity, spatial slice geometry and diffusion encoding relationships. A graph neural network predicts globally consistent stack-wise rigid motion, optimized in a self-supervised, zero-shot manner using only an anatomical reference image, without requiring paired ground-truth motion. We evaluate GraphSVR using both fully synthetic diffusion simulations and realistic recombination-based simulations from real acquisitions with controllable motion severity and gradient sparsity. Performance is quantified using grid error (mm) and rotation error relative to known ground-truth transforms. Under severe motion, GraphSVR reduces grid error and rotation error by 73% compared to FSL eddy, the standard DWI motion-correction method, with the largest gains observed in sparse-direction regimes. These results demonstrate that explicitly modeling acquisition structure through graph-based reasoning improves robustness and global consistency in 4D DWI motion estimation. Code is available at https://github.com/nogakertes/GraphSVR.git.
Chinese Translation
弥散加权成像(DWI)仍然极易受受试者运动的影响,尤其是在省时采集协议和易动人群中。虽然切片到体配准(SVR)可以缓解切片间和体块(stack)间的错位,但扩散MRI由于对比度依赖于扩散方向,并且需要在同一参考坐标系内对齐数十次测量,从而引入了额外的复杂性,实际上构成了一个4D配准问题。现有方法主要依赖序列化建模或成对相似度度量,在梯度采样稀疏或运动严重的情况下往往性能退化。我们提出了GraphSVR,一个用于DWI中4D SVR配准的q空间感知的基于图的框架。GraphSVR将切片组表示为按采集结构构建的图中的节点,边则编码时间邻近性、空间切片几何关系以及扩散编码关系。图神经网络(GNN)预测全局一致的逐体块刚性运动,并以自监督、零样本(zero-shot)的方式仅使用一张解剖参考图像进行优化,无需配对的真值运动标签。我们在全合成的扩散模拟以及基于真实采集数据的重组模拟(具有可控的运动严重程度和梯度稀疏度)上评估了GraphSVR。性能采用网格误差(mm)和相对于已知真值变换的旋转误差进行量化。在严重运动条件下,与标准DWI运动校正方法FSL eddy相比,GraphSVR将网格误差和旋转误差降低了73%,且在梯度方向稀疏的情形下提升最为显著。这些结果表明,通过基于图的推理显式建模采集结构,可以提高4D DWI运动估计的鲁棒性和全局一致性。代码可在 https://github.com/nogakertes/GraphSVR.git 获取。
cs.CV / 200 / 2609.24736

MiTHras: Task-specific Hierarchical Semi-supervised Contrastive Masked Autoencoder for Mitotic Figure Analysis

MiTHras:用于核分裂象分析的任务特定分层半监督对比掩码自编码器
Vuong, Trinh T. L., Graham, Simon, Vu, Quoc Dang, Ho, Phat T. H., Lim, Jeewoo, Jahanifar, Mostafa, Rajpoot, Nasir, Kwak, Jin T.
Abstract
Mitotic figure (MF) analysis supports tumor grading and prognostic assessment, but automated models remain sensitive to differences in tissue type and image acquisition. We present MiTHras, a task-specific pretraining framework that combines pseudo-label-guided image- and token-level contrastive learning with masked reconstruction. We construct TCGA-MF-Pseudo, a corpus of 1.8 million cell-centered images from 14 TCGA cohorts spanning 11 organ sites. Comprehensive evaluation on MF classification, detection, count-based survival prediction, and subtype classification demonstrates the efficacy of MiTHras. It achieves the highest mean F1 on all three MF classification benchmarks and both subtype benchmarks. MiTHras also outperforms general-purpose and pathology foundation encoders by a larger margin under frozen-encoder linear probing than under full fine-tuning. Although detection gains are modest due to a shared candidate-detection stage, ablations confirm that token-level supervision improves typical-versus-atypical classification and linear probing. These findings establish that MiTHras yields robust, transferable representations for automated mitotic activity assessment.
Chinese Translation
核分裂象(MF)分析支持肿瘤分级和预后评估,但自动化模型仍然对组织类型和图像采集的差异较为敏感。我们提出了MiTHras,这是一个任务特定的预训练框架,将伪标签引导的图像级和词元级对比学习与掩码重建相结合。我们构建了TCGA-MF-Pseudo,一个包含来自14个TCGA队列(覆盖11个器官部位)的180万张细胞中心图像的语料库。在核分裂象分类、检测、基于计数的生存预测和亚型分类上的综合评估证明了MiTHras的有效性。它在所有三个核分裂象分类基准和两个亚型分类基准上均取得了最高的平均F1分数。与全量微调相比,MiTHras在冻结编码器线性探测下以更大的幅度优于通用模型和病理学基础编码器。尽管由于共享的候选检测阶段,检测方面的提升较为有限,但消融实验证实词元级监督改善了典型与非典型核分裂象的分类以及线性探测性能。这些发现表明,MiTHras能够为自动化核分裂活性评估提供鲁棒且可迁移的表示。
cs.CV / 201 / 2609.24737

Ananke: Contractive Torus Attractor Networks

Ananke:收缩环面吸引子网络
Ji, Zhongping
Abstract
We introduce Ananke, a representation-learning framework that scaffolds latent representations onto a structured product-torus prior, and its flagship visual backbone realization, Contractive Torus Attractor Networks (CTAN). By factorizing high-dimensional latent spaces into an orthogonal direct sum of two-dimensional phase planes ($\bigoplus_{k=1}^K \R^2$), Ananke coordinates feature updates via a decoupled dual-phase continuous flow: skew-symmetric Hamiltonian transport moves features tangentially along energy level sets to preserve semantic phase invariants, while signed gradient dissipation contracts transverse perturbations normally toward target invariant manifolds. For circular potential families with frozen parameters, logarithmic radial feedback yields the Exact Log-Symplectic Flow (ELSF), an analytical closed-form mapping with exact exponential decay of log-radius error that evaluates in a single forward pass without numerical integration. We establish local input-to-state bounds for level-set deviations and log-radius errors, and characterize the normal hyperbolicity and persistence of the ideal product torus under bounded perturbations. We further formulate the architecture through Lie--Trotter operator splitting, unifying spatial depthwise diffusion with local manifold contraction, and analyze both exact trigonometric flows and hardware-friendly symplectic dual-shear variants. Across natural image benchmarks (CIFAR-100) and clinically challenging endoscopy datasets (Kvasir-v2), CTAN demonstrates exceptional parameter efficiency: an ultra-compact hierarchical model with merely 0.27M parameters achieves 90.52\% accuracy on Kvasir-v2, outperforming 25M+ baselines (ResNet-50, DenseNet-161) by nearly two orders of magnitude in capacity, while scaled variants attain 80.32\% top-1 accuracy on CIFAR-100.
Chinese Translation
我们提出了 Ananke,一个将潜在表示架构于结构化乘积环面先验之上的表示学习框架,及其旗舰视觉骨干网络实现——收缩环面吸引子网络(Contractive Torus Attractor Networks, CTAN)。通过将高维潜在空间分解为二维相位平面的正交直和($igoplus_{k=1}^K \R^2$),Ananke 通过解耦的双相位连续流协调特征更新:反对称哈密顿传输使特征沿能量水平集切向移动以保持语义相位不变量,而带符号梯度耗散则沿法向将横向扰动收缩至目标不变流形。对于参数冻结的圆形势函数族,对数径向反馈产生了精确对数辛流(Exact Log-Symplectic Flow, ELSF)——一种具有对数半径误差精确指数衰减的解析闭式映射,可在单次前向传播中完成计算而无需数值积分。我们为水平集偏差和对数半径误差建立了局部输入到状态界,并刻画了理想乘积环面在有界扰动下的法向双曲性与持久性。我们进一步通过 Lie–Trotter 算子分裂来构建该架构,将空间深度扩散与局部流形收缩统一起来,并分析了精确三角流以及硬件友好的辛双剪切变体。在自然图像基准(CIFAR-100)和具有临床挑战性的内窥镜数据集(Kvasir-v2)上,CTAN 展现出卓越的参数效率:一个仅含 0.27M 参数的超紧凑分层模型在 Kvasir-v2 上达到 90.52% 的准确率,以容量上近两个数量级的优势超越 25M+ 参数的基线模型(ResNet-50、DenseNet-161),而其扩展版本在 CIFAR-100 上达到 80.32% 的 top-1 准确率。
cs.CV / 202 / 2609.24768

PrismGPT: Proxy-Guided Learning for Region-Aware Photo Editing with Self-Synthesized Reasoning

PrismGPT:基于代理引导学习的区域感知照片编辑与自合成推理
Zhao, Ke, Nguyen, Hue, Punnappurath, Abhijith, Wang, Zhongling, Mohomed, Iqbal, Brown, Michael S.
Abstract
Professional photo finishing relies on both global adjustments and region-specific local edits guided by semantic masks, yet current automated methods handle this workflow only partially. We present PrismGPT, a Vision-Language Model (VLM) framework that produces structured, region-aware editing plans from a single input image without relying on commercial black-box tools. Training a VLM to simultaneously diagnose aesthetic deficiencies at both global and local levels while predicting precise editing parameters is challenging due to the vast combinatorial decision space. We address this through proxy-guided learning: two simpler proxy tasks -- operation decomposition and region-aware aesthetic ranking -- teach the foundational skills the model needs, while a competence-based dynamic scheduler automatically rebalances the multi-task training ratio, progressively shifting emphasis from the proxy tasks to the primary editing task as each skill is mastered. Crucially, all reasoning traces used for supervised fine-tuning are self-synthesized by the same base model, eliminating the need for a stronger external teacher. Experiments on MIT-Adobe FiveK and SPIRE, a new professionally retouched benchmark we introduce, show that PrismGPT achieves state-of-the-art results while using only ~6% of the training data compared to the previous best method.
Chinese Translation
专业的照片后期处理依赖于全局调整以及由语义掩码引导的区域特定局部编辑,然而当前的自动化方法只能部分地完成这一工作流程。我们提出了PrismGPT,这是一个视觉-语言模型(VLM)框架,能够从单张输入图像生成结构化的、区域感知的编辑方案,而无需依赖商业黑盒工具。训练VLM同时在全局和局部层面诊断美学缺陷并预测精确的编辑参数极具挑战性,原因在于其决策空间组合庞大。我们通过代理引导学习来解决这一问题:两个更简单的代理任务——操作分解和区域感知美学排序——为模型传授所需的基础技能,同时一个基于能力的动态调度器自动重新平衡多任务训练比例,随着各项技能的逐步掌握,将训练重心从代理任务渐进地转移到主编辑任务上。关键的是,用于监督微调的所有推理轨迹均由同一基础模型自合成生成,从而无需依赖更强的外部教师模型。在MIT-Adobe FiveK以及我们新推出的专业修图基准数据集SPIRE上的实验表明,PrismGPT取得了最先进的结果,同时仅使用先前最佳方法约6%的训练数据。
cs.CV / 203 / 2609.24769

Brain Metastases Segmentation for BraTS 2026 Task 1: A Multi-Architecture Comparison

BraTS 2026 任务1的脑转移瘤分割:多种架构的比较研究
Islam, Mahdi, Tabassum, Musarrat
Abstract
Brain metastases are the most common intracranial malignancy, occurring in roughly 30% of patients with primary solid tumors and carrying a median survival near 5.9 months. Automated segmentation is critical for treatment planning and volumetric monitoring, but metastases are frequently small, numerous, and heterogeneous in size within a single patient. We compare a plain nnU-Net baseline, a Residual Encoder Large (ResEncL) variant, region-based training, and a Primus transformer model for BraTS-METS 2026 Task 1, using patient-grouped cross-validation to prevent leakage from the longitudinal UCSD subset. Primus (label-based) is our strongest individual model by aggregate DSC/NSD, achieving 0.710/0.761 (ET), 0.742/0.785 (TC), 0.683/0.689 (WT), and 0.531/0.436 (RC). ResEncL trails Primus on aggregate DSC/NSD but achieves substantially higher lesion-wise F1 (e.g. ET: 0.452 vs. 0.052); a probability-averaging ensemble of the two only partially preserves ResEncL's F1 advantage (ET lesion-wise F1: 0.064). We further report three postprocessing and label-reconstruction pitfalls we believe generalize beyond this challenge. Code is available at https://github.com/mahdiislam79/BraTS_METS_2026.
Chinese Translation
脑转移瘤是最常见的颅内恶性肿瘤,约占原发性实体瘤患者的30%,中位生存期约为5.9个月。自动化分割对于治疗规划和体积监测至关重要,但转移瘤通常体积小、数量多,且在单一患者内部大小差异显著。我们针对BraTS-METS 2026任务1,比较了标准nnU-Net基线、残差编码器大模型(ResEncL)变体、基于区域的训练方法以及Primus Transformer模型,并采用患者分组的交叉验证以防止纵向UCSD子集的信息泄露。基于标签的Primus是我们按总体DSC/NSD衡量最强的单模型,其成绩为:0.710/0.761(ET)、0.742/0.785(TC)、0.683/0.689(WT)以及0.531/0.436(RC)。ResEncL在总体DSC/NSD上略逊于Primus,但在病灶级F1分数上显著更高(例如ET:0.452对0.052);而将两者进行概率平均的集成仅能部分保留ResEncL的F1优势(ET病灶级F1为0.064)。此外,我们还报告了三个后处理与标签重建方面的常见陷阱,我们认为这些问题具有超越本挑战的普遍性。代码已在 https://github.com/mahdiislam79/BraTS_METS_2026 开源。
cs.CV / 204 / 2609.24782

Virtual neural networks: hundreds of souls in a body

虚拟神经网络:一个躯体中的数百个灵魂
Hurtik, Petr, Vajgl, Marek, Alijani, Zahra, Molek, Vojtech
Abstract
A new concept, termed virtual neural networks, is introduced, where the count of trainable parameters is kept constant, and scalability is attained purely through computational resources. This concept is an abstract framework that can be realized using any standard convolutional neural network. It merges siamese neural networks with a deep ensemble technique by generating numerous virtual models that share weights derived from a small set of physical models. The ensemble comprises up to hundreds of trained models simultaneously. All virtual networks take the same input, and their interconnected structure induces an internal distortion that boosts the entire ensemble robustness. The accuracy of the ensemble improves as the number of virtual networks increases, without changing the capacity. Virtual neural networks outperform larger capacity models, typical deep ensembles, and contemporary approaches like SWA and Masksembles. Additionally, the highest-performing individual model from the ensemble surpasses other models trained individually, even those with a greater number of parameters. Code: gitlab.com/EnginCZ/virtual-models-public
Chinese Translation
本文提出了一种名为虚拟神经网络的新概念,其可训练参数的数量保持恒定,可扩展性完全通过计算资源来实现。该概念是一个抽象框架,可以使用任何标准的卷积神经网络来实现。它将孪生神经网络与深度集成技术相结合,通过生成大量共享权重的虚拟模型,而这些权重源自一小组物理模型。该集成可同时包含多达数百个训练好的模型。所有虚拟网络接收相同的输入,其相互连接的结构会引发一种内部扰动,从而增强整个集成模型的鲁棒性。随着虚拟网络数量的增加,集成的准确性不断提高,而无需改变模型容量。虚拟神经网络优于更大容量的模型、典型的深度集成方法,以及SWA和Masksembles等当代方法。此外,集成中表现最好的单个模型也超过了其他单独训练的模型,即使后者拥有更多的参数。代码:gitlab.com/EnginCZ/virtual-models-public
cs.CV / 205 / 2609.24787

Toward a foundation model for forest point clouds

迈向森林点云基础模型
Yue, Yuanwen, Puliti, Stefano, Robert, Damien, Topaloğlu, Atakan, Xiang, Binbin, Wielgosz, Maciej, Wegner, Jan Dirk, Astrup, Rasmus, Rupprecht, Christian, Schindler, Konrad
Abstract
Forest inventories increasingly rely on artificial intelligence (AI) models to derive forest attributes from large-scale 3D point clouds. Current models are typically specialized to a single task, sensor, and forest type, making adaptation expensive in terms of annotations, computation, and expertise. We ask whether a single pretrained model can instead learn transferable representations across diverse forest inventory settings. Inspired by recent developments in language modelling and computer vision, we take a step toward a foundation model (FM) for 3D forestry. Using LitePT as backbone, we first establish a strong supervised baseline that sets a new state of the art on forest semantic and instance segmentation, tree species classification, and age regression benchmarks. We then curate a large-scale unlabelled corpus spanning airborne, UAV, and mobile laser scanning across diverse forest ecosystems, and pretrain the same backbone using self-supervised learning. We systematically evaluate representation learning strategies by comparing training from scratch, supervised pretraining, and self-supervised pretraining across four representative forestry tasks, under varying annotation budgets. Compared with training from scratch, self-supervised pretraining accelerates model convergence and consistently improves performance when annotations are scarce. Compared with task-specific supervised pretraining, self-supervised pretraining yields more transferable representations across downstream forestry tasks. These findings identify the practical regime in which pretrained representations are most valuable and suggest that instance discrimination, rather than forest semantics, is the main remaining obstacle to a general-purpose 3D forest foundation model. Code and models are available at: https://github.com/prs-eth/ForPT.
Chinese Translation
森林资源调查日益依赖人工智能(AI)模型从大规模三维点云中提取森林属性。现有模型通常仅针对单一任务、传感器和森林类型进行专门设计,使得模型适配在标注、计算和专业人才方面成本高昂。我们提出了这样一个问题:单个预训练模型能否在多样的森林调查场景中学习到可迁移的表征?受语言建模和计算机视觉领域最新进展的启发,我们朝着面向三维林业的基础模型(Foundation Model, FM)迈出了一步。以 LitePT 为骨干网络,我们首先建立了一个强大的监督基线,在森林语义分割与实例分割、树种分类以及林龄回归等基准测试中刷新了当前最优水平。随后,我们构建了一个大规模无标注语料库,涵盖机载、无人机和移动激光扫描等多种森林生态系统,并使用自监督学习对同一骨干网络进行预训练。我们在不同的标注预算条件下,通过比较从零训练、监督预训练和自监督预训练三种策略在四个代表性林业任务上的表现,系统性地评估了各种表征学习策略。与从零训练相比,自监督预训练加快了模型收敛速度,并在标注稀缺时持续提升性能。与面向特定任务的监督预训练相比,自监督预训练在下游林业任务上产生更具可迁移性的表征。这些发现明确了预训练表征最具价值的实际应用场景,并表明实例判别(而非森林语义本身)是构建通用三维森林基础模型的主要剩余障碍。代码和模型可在以下网址获取:https://github.com/prs-eth/ForPT。
cs.CV / 206 / 2609.24788

Streaming Video Editing with Easy Adaptation

基于轻量适配的流式视频编辑
Hu, Yujia, Li, Jiajun, He, Zihao, Liu, Songhua
Abstract
In this paper, we propose SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion model but supports high-quality streaming video editing in an auto-regressive fashion. To tackle this problem, we first systematically revisit existing video-to-video diffusion approaches and identify two key principles for such streaming adaptation: backbone feature disentanglement and conditional frame independence. Building on these insights, we develop a novel paradigm for controllable video generation. At its core, an auxiliary model branch encodes source video inputs with temporally independent self-attention, and the intermediate features are injected into the corresponding backbone blocks for streaming-compatible control. Moreover, to bridge the discrepancy between the feature spaces of bidirectional and streaming models, we propose a decoupled training scheme that explicitly enforces the orthogonality between the optimization directions of video controllability and model causality. Such disentanglement ensures compatibility between the two objectives at inference and facilitates smooth zero-shot knowledge transfer across heterogeneous backbone architectures. Extensive experiments demonstrate that SVEET achieves superior editing quality while maintaining real-time performance, attaining 15 FPS on a single H100 GPU 17 without any auxiliary acceleration techniques. Codes are available at https://github.com/YujiaHu1109/SVEET.
Chinese Translation
在本文中,我们提出了 SVEET,该框架仅需在预训练的双向视频扩散模型上进行训练,即可支持高质量的自回归流式视频编辑。为解决这一问题,我们首先系统地回顾了现有的视频到视频扩散方法,并识别出实现此类流式适配的两个关键原则:主干特征解耦与条件帧独立性。基于这些见解,我们开发了一种新颖的可控视频生成范式。其核心是一个辅助模型分支,通过时间上相互独立的自注意力机制对源视频输入进行编码,并将中间特征注入相应的主干模块中,以实现与流式处理兼容的控制。此外,为弥合双向模型与流式模型特征空间之间的差异,我们提出了一种解耦训练方案,显式地强制视频可控性与模型因果性这两个优化方向之间的正交性。这种解耦确保了两个目标在推理时的兼容性,并促进了异构主干架构之间平滑的零样本知识迁移。大量实验表明,SVEET 在保持实时性能的同时实现了卓越的编辑质量,在单张 H100 GPU 上无需任何辅助加速技术即可达到 15 FPS。代码已发布于 https://github.com/YujiaHu1109/SVEET。
cs.CV / 207 / 2609.24813

INTCORT: Training-Free Spatial Reasoning Enhancement for Vision-Language Models via Input Transformations and Confidence Routing

INTCORT:基于输入变换与置信度路由的视觉语言模型免训练空间推理增强方法
Sun, Haoran, Xu, Jingqi, Li, Yanhui, Liu, Enci, Xu, Kaidi, Liu, Yanwei
Abstract
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in multimodal tasks, yet they still exhibit poor ability in spatial reasoning. Existing training-dependent and training-free enhancement methods suffer from high computational costs with catastrophic forgetting and internal mechanism interference that compromises general capabilities, respectively. In this work, we first verify two key hypotheses: appropriate geometric image transformation and query-reversal transformation can recover incorrect spatial predictions, and correct predictions exhibit higher relation-token confidence than incorrect ones. Based on these findings, we propose INTCORT, a training-free spatial reasoning enhancement framework that constructs multiple inference views through input transformations and aggregates their predictions via relation-token confidence routing, without modifying the VLM's internal mechanisms. Experimental results on several commonly-used benchmarks demonstrate that INTCORT substantially improves spatial reasoning accuracy across diverse VLMs, achieving an average improvement of 10.01% over all models and benchmarks. Compared with prior works, INTCORT achieves superior performance with improvements of up to 25.01%.
Chinese Translation
视觉语言模型(Vision-Language Models, VLMs)在多模态任务中展现出了卓越的能力,但其在空间推理方面仍表现不佳。现有的依赖训练的增强方法存在计算成本高且会导致灾难性遗忘的问题,而免训练的增强方法则会因内部机制干扰而损害模型的通用能力。在本工作中,我们首先验证了两个关键假设:适当的几何图像变换和查询反转变换可以纠正错误的空间预测,且正确预测的关系词元(relation-token)置信度高于错误预测。基于这些发现,我们提出了INTCORT,一个免训练的空间推理增强框架。该框架通过输入变换构建多个推理视图,并利用关系词元置信度路由对其预测结果进行聚合,且无需修改VLM的内部机制。在多个常用基准上的实验结果表明,INTCORT显著提升了多种VLM的空间推理准确率,在所有模型和基准上平均提升10.01%。与已有工作相比,INTCORT取得了更优的性能,提升幅度最高达25.01%。
cs.CV / 208 / 2609.24814

Mobile Imaging Solutions for Medical Diagnosis: Trends and Applications

用于医学诊断的移动成像解决方案:趋势与应用
Zulfiker, Syed Muhammad Ibne, Hashem, Tanzima, Islam, Fariha Tabassum, Arifin, Md Sultanul, Islam, Khandker Aftarul, Bristy, Nishat Anjum, Huq, Faria, Saha, Priyeta, Akter, Syeda Nahida, Saha, Arpita
Abstract
Advances in processing power, camera technologies, and mobile image analysis have made smartphones and other mobile devices, such as laptops, increasingly suitable for medical diagnosis and healthcare applications. Researchers have developed low-cost solutions for the early detection and monitoring of various health conditions, including eye and ENT diseases, malnutrition, heart rate variability, skin and oral conditions, and injuries, using images captured by non-medical devices such as smartphones and webcams. This survey examines existing research on mobile image-based medical diagnosis, with an emphasis on its potential to enable low-cost and accessible healthcare. We comparatively analyze state-of-the-art solutions across different healthcare application categories, examining their advantages and limitations. Based on this analysis, we identify desirable characteristics of mobile image-based diagnostic tools and highlight areas where existing approaches have made progress as well as areas requiring further research. We also discuss application-specific and common challenges and outline directions for future research. Overall, this study provides a comprehensive overview of mobile image-based healthcare solutions and their potential to support low-cost disease diagnosis and monitoring, particularly for underserved populations in remote and resource-constrained settings.
Chinese Translation
处理能力、相机技术和移动图像分析的进步,使得智能手机及其他移动设备(如笔记本电脑)日益适用于医学诊断和医疗保健应用。研究人员利用智能手机和网络摄像头等非医疗设备采集的图像,开发出低成本解决方案,用于多种健康状况的早期检测与监测,包括眼部和耳鼻喉疾病、营养不良、心率变异性、皮肤和口腔状况以及损伤。本综述考察了基于移动图像的医学诊断领域的现有研究,重点关注其在实现低成本、可及性医疗保健方面的潜力。我们对不同医疗应用类别中最先进的解决方案进行了比较分析,审视其优势与局限性。基于这一分析,我们归纳了基于移动图像的诊断工具应具备的理想特性,并指出现有方法已取得进展的领域以及仍需进一步研究的方向。我们还讨论了特定应用和共性的挑战,并概述了未来研究的方向。总体而言,本研究全面概述了基于移动图像的医疗保健解决方案及其支持低成本疾病诊断与监测的潜力,尤其是对偏远和资源受限环境中医疗服务不足的人群。
cs.CV / 209 / 2609.24825

ZVeC: A Zero-Shot Framework for Instance-Level Vehicle Extraction and Generative Point Cloud Completion

ZVeC:一种用于实例级车辆提取与生成式点云补全的零样本框架
Li, Daisy, Gao, Kyle, Wu, Quanyun, Jutzi, Boris, Zelek, John S., Li, Jonathan
Abstract
LiDAR point clouds acquired in underground environments exhibit severe geometric incompleteness due to occlusions and limited sensor viewpoints, making reliable point cloud completion challenging without large supervised datasets. We propose ZVeC, a zero-shot, instance-driven framework that reformulates scene-level completion as compositional object-level reconstruction. By decomposing a scene into semantic object instances, ZVeC reduces reconstruction ambiguity in cluttered environments while eliminating the need for scenario-specific training. Each segmented vehicle is completed independently using a depth- and 3D Gaussian-conditioned diffusion model that exploits generalized geometric priors before the reconstructed instances are recomposed into the original scene. To evaluate our approach, we construct a real-world dense LiDAR benchmark of underground parking environments. Experimental results demonstrate consistent improvements over representative scene-level baselines in both quantitative metrics and visual quality. The completed point cloud differs substantially from the measured input (average KL divergence ~ 2.1), yet reducing the input to only 1% of the original LiDAR measurements changes the completed reconstruction only marginally (KL divergence < 0.50). This demonstrates that ZVeC produces geometrically consistent completions even under extreme input sparsity.
Chinese Translation
在地下环境中获取的LiDAR点云由于遮挡和传感器视角受限而呈现出严重的几何不完整性,使得在缺乏大规模监督数据集的情况下实现可靠的点云补全极具挑战性。我们提出了ZVeC,一个零样本、实例驱动的框架,它将场景级补全问题重新构建为组合式的对象级重建。通过将场景分解为语义对象实例,ZVeC降低了杂乱环境中重建的歧义性,同时无需针对特定场景进行训练。每个分割出的车辆均使用一个基于深度和3D高斯条件约束的扩散模型(diffusion model)独立完成补全,该模型利用泛化的几何先验,随后重建的实例被重新组合回原始场景。为评估我们的方法,我们构建了一个真实世界的地下停车场环境密集LiDAR基准数据集。实验结果表明,无论是在定量指标还是视觉质量方面,本方法均持续优于代表性的场景级基线方法。补全后的点云与测量输入存在显著差异(平均KL散度约为2.1),然而将输入缩减至原始LiDAR测量的仅1%时,补全重建结果的变化却很小(KL散度小于0.50)。这表明ZVeC即使在极端输入稀疏的情况下也能产生几何一致的补全结果。
cs.CV / 210 / 2609.24839

When Wider Views Fail: Stress-Testing Feed-Forward 3D Reconstruction

当更宽的视角失效时:前馈式三维重建的压力测试
Li, Daisy, Gao, Kyle, Wu, Quanyun, Chomko, Hanna, Zelek, John S., Li, Jonathan
Abstract
Feed-forward 3D reconstruction models enable efficient geometry estimation from sparse images, but their pretrained nature can make them vulnerable to distribution shifts beyond their training data. Identifying these failure modes is important for understanding when such models can be reliably deployed in unconstrained imaging settings. We investigate viewpoint variation as a controlled distribution shift by varying the angular span of sparse image inputs while keeping the input budget fixed. Across multiple feed-forward reconstruction models, we observe substantial degradation as viewpoint span increases, with wide spans producing both incomplete surface coverage and geometry unsupported by the observed imagery. These results reveal that viewpoint variation can induce failure modes beyond conventional reconstruction incompleteness, highlighting the need to evaluate pretrained feed-forward models under distribution shifts that challenge their learned geometric priors.
Chinese Translation
前馈式三维重建模型能够从稀疏图像中高效地进行几何估计,但其预训练的特性使其容易受到超出训练数据范围的分布偏移影响。识别这些失效模式对于理解此类模型在何种情况下能够可靠地部署于非受限的成像环境至关重要。我们通过在保持输入数量固定的前提下改变稀疏图像输入的角度范围,将视角变化作为一类受控的分布偏移进行研究。在多个前馈式重建模型上,我们观察到随着视角范围的增大,模型性能出现显著退化:宽视角范围不仅导致表面覆盖不完整,还会产生缺乏观测图像支持的几何结构。这些结果表明,视角变化所引发的失效模式超出了传统的重建不完整性问题,凸显了在挑战预训练模型所学几何先验的分布偏移下评估前馈式模型的必要性。
cs.CV / 211 / 2609.24850

Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation

重新审视多视图立体匹配:一种序列到序列的建模方法
Fan, Aoxiang, Dumery, Corentin, Talabot, Nicolas, Fua, Pascal
Abstract
Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera parameters are supplied. In this paper, we study the multi-view stereo (MVS) problem with known camera parameters and propose a novel approach that bridges conventional MVS and FF methods. Rather than casting MVS as a sequence-to-one mapping that predicts depth only for a single reference view, we reformulate it as a sequence-to-sequence task, akin to FF models, that jointly predicts geometry for all input views. We introduce a global transformer-based architecture with two components that explicitly exploit camera-induced priors: ray-map embeddings that inject camera parameters into image patch tokens, making the transformer camera-aware, and a unified global cost volume that replaces conventional per-view cost volumes to jointly capture 3D structure across all views. Extensive experiments on multiple public benchmarks show our approach achieves state-of-the-art performance, surpassing both MVS and FF reconstruction baselines.
Chinese Translation
从多视图图像中计算精确的几何结构是计算机视觉中的一个基础问题。近期的前馈(feed-forward, FF)模型能够联合估计三维几何与相机参数,但即使提供了真实相机参数,它们通常仍会因重建歧义而产生几何畸变。本文研究在已知相机参数条件下的多视图立体匹配(multi-view stereo, MVS)问题,并提出了一种连接传统 MVS 方法与前馈方法的新方法。不同于将 MVS 视为仅为单一参考视图预测深度的序列到一映射,我们将其重新建模为类似前馈模型的序列到序列任务,即联合预测所有输入视图的几何结构。我们提出了一种基于全局 Transformer 的架构,包含两个显式利用相机先验的组件:一是射线图嵌入,将相机参数注入图像块标记中,使 Transformer 具备相机感知能力;二是统一的全局代价体,用以替代传统的逐视图代价体,从而联合捕捉所有视图间的三维结构。在多个公开基准上的大量实验表明,我们的方法达到了最先进的性能,超越了 MVS 和前馈重建基线方法。
cs.CV / 212 / 2609.24872

DTKDP: A Dual Teacher Knowledge Distillation and Pruning Framework for Lightweight Oriented SAR Ship Detection

DTKDP:一种面向轻量化SAR舰船检测的双教师知识蒸馏与剪枝框架
Li, Yuming, Zhang, Fan, Achim, Alin M.
Abstract
Two-stage oriented detectors achieve high localization accuracy in synthetic aperture radar (SAR) ship detection, but their large backbones, feature pyramids, proposal modules, and heavy region of interest (RoI) heads hinder deployment. Existing lightweight SAR ship detectors typically use one-stage frameworks that lack proposal-level refinement for precise rotated localization. This paper presents a dual-teacher knowledge distillation and pruning (DTKDP) framework for lightweight oriented SAR ship detection. DTKDP introduces learnable gates into convolutional, normalization, and linear layers to prune convolutional channels and RoI-head neurons. Rotated proposal alignment (RPA) distills teacher and student predictions in a shared teacher-generated rotated proposal space, while a dual-teacher scheme combines classification and regression guidance from a homogeneous main teacher with complementary classification cues from a heterogeneous auxiliary teacher. Experiments on the SAR Ship Detection Dataset (SSDD) and Rotated Ship Detection Dataset in SAR Images (RSDD-SAR) show that DTKDP reduces the parameters of Oriented Region-based Convolutional Neural Network (Oriented R-CNN) and RoI Transformer equipped with ResNet-50 backbones by 87.5-91.8% and their floating-point operations (FLOPs) by 75.6-79.9%. In terms of average precision (AP) and mean average precision (mAP), the resulting Oriented R-CNN-slim and RoI Transformer-slim retain accuracy close to their full-scale counterparts. Relative changes across $\mathrm{AP}_{50}$, $\mathrm{AP}_{75}$, $\mathrm{mAP}_{50:75}$, and $\mathrm{mAP}_{50:95}$ range from a 2.38% decrease to a 0.65% improvement. Compared with RTMDet-tiny, they improve all four metrics on both datasets by 0.52-27.55% and consistently surpass representative distillation methods, demonstrating a favorable accuracy-efficiency trade-off.
Chinese Translation
两阶段旋转目标检测器在合成孔径雷达(SAR)舰船检测中能够实现较高的定位精度,但其庞大的骨干网络、特征金字塔、候选框模块以及沉重的感兴趣区域(RoI)检测头阻碍了实际部署。现有的轻量化SAR舰船检测器通常采用单阶段框架,缺乏针对精确旋转定位的候选框级别的精炼。本文提出了一种面向轻量化SAR舰船检测的双教师知识蒸馏与剪枝(DTKDP)框架。DTKDP在卷积层、归一化层和线性层中引入可学习门控,以对卷积通道和RoI检测头神经元进行剪枝。旋转候选框对齐(RPA)在一个共享的、由教师生成的旋转候选框空间中对教师和学生的预测进行蒸馏;同时,双教师方案将来自同质主教师的分类与回归引导,与来自异质辅助教师的互补分类线索相结合。在SAR舰船检测数据集(SSDD)和SAR图像旋转舰船检测数据集(RSDD-SAR)上的实验表明,DTKDP使配备ResNet-50骨干网络的旋转区域卷积神经网络(Oriented R-CNN)和RoI Transformer的参数量减少了87.5%–91.8%,浮点运算量(FLOPs)减少了75.6%–79.9%。在平均精度(AP)和平均精度均值(mAP)方面,所得的Oriented R-CNN-slim和RoI Transformer-slim保持了与其完整模型接近的精度,其在$\mathrm{AP}_{50}$、$\mathrm{AP}_{75}$、$\mathrm{mAP}_{50:75}$和$\mathrm{mAP}_{50:95}$上的相对变化范围为下降2.38%至提升0.65%。与RTMDet-tiny相比,二者在两个数据集上的四项指标均提升了0.52%–27.55%,并持续超越代表性蒸馏方法,展现了良好的精度-效率权衡。
cs.CV / 213 / 2609.24875

SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models

SPHQuant:面向视觉-语言模型的高效极低比特权重量化方法
Zhang, Kewei, Chen, Zheng, Qin, Haotong, Zhang, Yulun
Abstract
Recent foundation models are moving toward native multimodal Vision-Language Models (VLMs), making VLMs a central form of next-generation foundation models. However, their large language backbones make edge deployment difficult due to high memory footprint and memory-bound autoregressive decoding. Weight-only post-training quantization is a practical solution, but pushing VLMs to extreme low bit-widths remains challenging: existing rotation-free methods suffer from outliers at 2-3 bits, while rotation-based methods improve accuracy at the cost of additional runtime overhead. We propose SPHQuant, a rotation-free spherical weight-only quantization framework for VLMs. Instead of quantizing weights directly in Cartesian coordinates, SPHQuant decomposes each 8D weight vector into coordinate signs, radius, and a positive unit direction. This representation isolates outlier magnitude into the radius while keeping directions bounded and statistically regular. Based on this insight, SPHQuant allocates extra precision to the radius to mitigate accuracy degradation induced by outliers. It further uses a compact positive-direction codebook and fine-tunes codebook entries through angular parameterization to preserve the unit-sphere constraint. We also design a hardware-friendly GEMV kernel that keeps the direction codebook small enough for shared-memory lookup and packs radial bits efficiently. Experiments show that SPHQuant matches the performance of state-of-the-art extreme low-bit quantization methods while improving decode throughput over QTIP by 30.3% on RTX A6000. Code will be released in https://github.com/Pushazf/SPHQuant.
Chinese Translation
近期的基座模型正朝着原生多模态视觉-语言模型(Vision-Language Models, VLMs)发展,使VLMs成为下一代基座模型的核心形态。然而,其庞大的语言模型主干带来高昂的内存占用,且自回归解码受限于内存带宽,导致边缘部署十分困难。仅权重的训练后量化是一种实用的解决方案,但将VLMs推向极低比特宽度仍然充满挑战:现有的无旋转方法在2-3比特下受离群值影响,而基于旋转的方法虽提升了精度,却以额外的运行时开销为代价。我们提出了SPHQuant,一个面向VLMs的无旋转球面仅权重量化框架。SPHQuant并非直接在笛卡尔坐标下量化权重,而是将每个8维权重向量分解为坐标符号、半径和一个正的单位方向向量。这种表示将离群值的幅度隔离到半径中,同时保持方向有界且统计上规则。基于这一洞察,SPHQuant为半径分配额外的精度,以缓解离群值导致的精度下降。它进一步使用紧凑的正方向码本,并通过角度参数化对码本条目进行微调,以保持单位球约束。我们还设计了一个硬件友好的GEMV(通用矩阵-向量乘)内核,使方向码本足够小以便于共享内存查找,并高效地打包半径比特。实验表明,SPHQuant在性能上与最先进的极低比特量化方法相当,同时在RTX A6000上将解码吞吐量相对QTIP提升了30.3%。代码将在 https://github.com/Pushazf/SPHQuant 发布。
cs.CV / 214 / 2609.24879

Generating Chest X-Ray Counterfactuals by Specialising Foundation Image Models

通过专门化基础图像模型生成胸部X光反事实图像
Xing, Xiaodan, Rasal, Rajat R., Meister, Julia A., Ghorayeb, Sara, Khara, Galvin, Schrouff, Jessica
Abstract
Counterfactual image generation answers questions about how a subject would have looked under retrospective, hypothetical scenarios. Recent methods have improved perceptual quality, identity preservation and faithfulness to an underlying causal model, but their adoption in healthcare is limited by scarce annotated data, distribution shift between datasets, and mismatches between pretrained generative models and those required for counterfactual inference. We propose specialisation, a data and parameter-efficient framework for adapting pretrained, non-causal generative models into causal mechanisms under distribution shift. Based on this framework, we train a radiology counterfactual image generation model, called RadCF, using latent flow matching. We validate our approach on three chest X-ray datasets spanning different dataset shifts, data volumes, and counterfactual questions, associated with challenging, highly-localised interventions. Our results show that RadCF and specialisation improve counterfactual soundness over existing methods while being data and parameter efficient, and that the resulting counterfactuals can detect and mitigate shortcut learning in a downstream medical classifier. Code is available at https://github.com/GSK-AI/RadCF/.
Chinese Translation
反事实图像生成旨在回答在回顾性的假设情景下,受试者可能呈现何种外观的问题。近期的方法在感知质量、身份保持以及与潜在因果模型的一致性方面有所改进,但其在医疗领域的应用受到标注数据稀缺、数据集间分布偏移以及预训练生成模型与反事实推断所需模型之间不匹配的限制。我们提出了一种名为专门化(specialisation)的框架,该框架在数据和参数上均高效,可将预训练的非因果生成模型在分布偏移条件下适配为因果机制。基于该框架,我们使用潜在流匹配(latent flow matching)训练了一个放射学反事实图像生成模型,称为RadCF。我们在三个胸部X光数据集上验证了我们的方法,这些数据集涵盖了不同的数据集偏移、数据规模以及反事实问题,并涉及具有挑战性的高度局部化干预。结果表明,RadCF和专门化方法在数据和参数高效的同时,比现有方法具有更好的反事实合理性,且所产生的反事实图像能够检测并缓解下游医学分类器中的捷径学习(shortcut learning)。代码可在 https://github.com/GSK-AI/RadCF/ 获取。
cs.CV / 215 / 2609.24894

SLICEChat: Progressive In-Encoder Token Pruning for Whole-Slide Pathology Language Models

SLICEChat:面向全切片病理语言模型的编码器内渐进式词元剪枝
Bozkurt, Ali Kerem, Bakay, Baris Cem, Kulac, Ibrahim, Gunduz-Demir, Cigdem, Erdem, Erkut, Erdem, Aykut
Abstract
Whole-slide pathology images (WSIs) contain gigapixel-scale visual content, creating a major scalability challenge for slide-level multimodal large language models (MLLMs). Existing approaches process thousands of patch tokens and typically apply compression only after slide encoding, leaving multimodal attention computationally expensive. We introduce SLICEChat, a slide-level MLLM that integrates progressive token pruning within a hybrid Mamba--Transformer slide encoder. Mamba layers enable efficient long-range propagation, while Transformer layers preserve global interactions as the sequence is progressively shortened. Between stages, language-supervised, region-aware pruning removes spatially coherent low-utility regions under a controlled keep-rate schedule, producing compact slide representations before multimodal fusion. On SlideBench VQA, SLICEChat achieves 79.84% accuracy on TCGA and 59.09% on BCNB cohorts, outperforming prior slide-level pathology MLLMs, and achieves the highest overall WSI-Bench metrics. It also provides competitive memory usage and the inference latency among the evaluated models. These results demonstrate accurate and computationally efficient multimodal reasoning over gigapixel WSIs.
Chinese Translation
全切片病理图像(WSI)包含吉像素级别的视觉内容,这给切片级多模态大语言模型(MLLM)带来了巨大的可扩展性挑战。现有方法需处理数千个图像块词元,且通常仅在切片编码之后进行压缩,导致多模态注意力的计算开销高昂。我们提出 SLICEChat,一种将渐进式词元剪枝集成于 Mamba-Transformer 混合切片编码器中的切片级 MLLM。Mamba 层实现高效的长距离信息传播,而 Transformer 层则在序列逐步缩短的过程中保持全局交互。在各个阶段之间,基于语言监督的区域感知剪枝在受控的保留率调度下移除空间上连贯的低效用区域,从而在多模态融合之前生成紧凑的切片表示。在 SlideBench VQA 上,SLICEChat 在 TCGA 和 BCNB 队列上分别达到 79.84% 和 59.09% 的准确率,优于现有的切片级病理 MLLM,并在 WSI-Bench 上取得最高的综合指标。同时,在所评估的模型中,其内存占用与推理延迟也具有竞争力。这些结果表明,SLICEChat 能够在吉像素 WSI 上实现准确且计算高效的多模态推理。
cs.CV / 216 / 2609.24919

PixelDiT2: Representation-Grounded Pixel Diffusion Transformers

PixelDiT2:基于表征引导的像素扩散Transformer
Yu, Yongsheng, Xiong, Wei, Sheng, Yichen, Liu, Shiqiu, Luo, Jiebo
Abstract
Recent advances in pixel-space diffusion models have narrowed the image quality gap with latent-space diffusion, but still converge more slowly and lag behind in final image quality. We argue that a key reason is the lack of an explicit representation prior: unlike latent diffusion, which usually denoises in a compact and structured latent space, pixel diffusion needs to learn denoising-friendly representations and pixel generation simultaneously from raw RGB space. To address this problem, we propose PixelDiT2, an end-to-end pixel-space diffusion model designed to decouple representation learning from pixel generation without introducing an autoencoder or latent reconstruction bottleneck. We propose representation grounding that uses a frozen pretrained vision foundation model to provide explicit per-patch representation guidance throughout denoising, allowing the pixel diffusion transformer to focus more on pixel generation. On ImageNet-256x256, PixelDiT2 achieves an FID of 1.46 after 600 epochs; at 512x512 resolution, PixelDiT2 achieves an FID of 1.48 after 680 epochs.
Chinese Translation
像素空间扩散模型的最新进展已缩小了其与潜在空间扩散模型在图像质量上的差距,但其收敛速度仍较慢,且最终图像质量仍有差距。我们认为一个关键原因在于缺乏显式的表征先验:与通常在紧凑且结构化的潜在空间中进行去噪的潜在扩散不同,像素扩散需要从原始RGB空间中同时学习适合去噪的表征和像素生成。为解决这一问题,我们提出了PixelDiT2,这是一种端到端的像素空间扩散模型,旨在将表征学习与像素生成解耦,且无需引入自编码器或潜在重建瓶颈。我们提出了表征引导(representation grounding)方法,利用冻结的预训练视觉基础模型在整个去噪过程中提供显式的逐patch表征指导,使像素扩散Transformer能够更专注于像素生成。在ImageNet-256x256上,PixelDiT2经过600个epoch后取得了1.46的FID;在512x512分辨率下,PixelDiT2经过680个epoch后取得了1.48的FID。
cs.CV / 217 / 2609.24937

Anatomy-Decomposed Chest Computed Tomography (CT) Projections as Scalable Supervision for Bone Suppression in Chest Radiographs

解剖结构分解的胸部计算机断层扫描(CT)投影作为胸片骨抑制的可扩展监督
Angaitkar, Mrunmay, Kumar, Piyush, Satia, Aarjav, Rao, Pranav, Mittal, Ashish, Tadepalli, Manoj, Putha, Preetham
Abstract
Bone overlap can obscure abnormalities in chest radiographs, while scarce paired training data limit supervised bone suppression. We address this challenge with a digitally reconstructed radiograph (DRR) framework that converts chest computed tomography (CT) into paired supervision for component suppression. A novel bone segmentation algorithm enables CT decomposition into bone, non-lung soft-tissue, and lung components, which are projected separately. Their weighted combination yields synthetic radiographs with pixel-registered component images that sum exactly to the full DRR. Models trained on these data suppress bone or lung components by predicting the target component and recovering the remainder by subtraction, transferring to real radiographs without real paired training data. As an extension, their outputs on real radiographs provide target domains for unpaired, component-wise DRR translation, reducing the appearance gap while retaining anatomical details. Across multiple public datasets, downstream detection experiments demonstrate the utility of bone suppression, with gains concentrated on abnormalities with substantial bone overlap. Compared with open-source DRR engines applied to the same CTs, our unmodified DRRs achieve comparable realism and preservation of label-relevant anatomy, while translated DRRs achieve the best Fr\'echet inception distance (FID), lung-field sharpness, and agreement with source-CT anatomy among the evaluated methods. Models and inference code: https://huggingface.co/qureaiorg/bone-suppression; Translated projections: https://huggingface.co/datasets/qureaiorg/ct2xr-projections.
Chinese Translation
骨骼重叠会遮挡胸片中的异常病灶,而稀少的配对训练数据限制了有监督骨抑制方法的发展。为应对这一挑战,我们提出了一种数字重建影像(DRR)框架,将胸部计算机断层扫描(CT)转换为用于组件抑制的配对监督数据。一种新颖的骨骼分割算法可将CT分解为骨骼、非肺部软组织和肺三种成分,并分别进行投影。通过对这些成分进行加权组合,可生成具有像素级配准的成分图像的合成胸片,且各成分图像之和精确等于完整DRR。在这些数据上训练的模型可通过预测目标成分并通过减法恢复剩余成分,实现骨或肺成分的抑制,并可在无需真实配对训练数据的情况下迁移到真实胸片。作为扩展,模型在真实胸片上的输出可为非配对的逐成分DRR转换提供目标域,在保留解剖细节的同时缩小外观差异。在多个公开数据集上,下游检测实验证明了骨抑制的实用性,其性能提升集中于与骨骼重叠明显的异常病灶。与应用于相同CT的开源DRR引擎相比,我们未经修改的DRR在真实感以及对标签相关解剖结构的保留方面达到了相当的水平;而在所评估的方法中,经转换的DRR在Fréchet inception距离(FID)、肺野清晰度以及与源CT解剖结构的一致性方面均取得最佳表现。模型与推理代码:https://huggingface.co/qureaiorg/bone-suppression;转换后的投影数据:https://huggingface.co/datasets/qureaiorg/ct2xr-projections。
cs.CV / 218 / 2609.24981

GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation

GAE:学习几何原生潜在空间以实现3D一致的世界生成
Lu, Jiahao, Yin, Minghao, Hu, Wenbo, Liu, Hengyu, Zhao, Wang, Yeung, Sai-Kit, Shan, Ying, Liu, Yuan
Abstract
We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse generation tasks. In controlled comparisons that hold the generator and training protocol fixed, replacing the latent with GAE improves both visual quality and independently measured 3D coherence: FVD falls by $12.7\%$ and $23.1\%$ on RealEstate10K and DL3DV, and camera-trajectory error is halved on RealEstate10K. Together, these results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.
Chinese Translation
我们提出了一种紧凑的几何原生潜在空间,作为感知与生成的共享基础。视觉生成器可以生成照片级逼真的帧,但无法保持3D场景的一致性。我们认为这不仅是一个建模问题,更是一个表示问题:生成器通常演化出以外观为中心的潜在表示,而感知模型则在一个编码了跨视角结构的语义丰富空间中恢复几何信息。我们不将几何作为另一个输出添加进去,而是将几何基础模型的特征重新参数化为一个用于生成的紧凑潜在空间。我们通过几何原生自编码器(GAE)实现这一转变,其潜在表示可被联合解码为外观、深度、相机参数和点图。基于这一状态表示,标准的条件流即可支持多样化的生成任务。在保持生成器和训练协议固定的受控比较中,用GAE替换潜在表示既提升了视觉质量,也提升了对3D一致性进行独立测量的指标:在RealEstate10K和DL3DV数据集上,FVD分别下降12.7%和23.1%,且在RealEstate10K上相机轨迹误差减半。这些结果表明,潜在空间是实现几何一致生成的核心,并可作为感知与生成之间的共享接口。
cs.CV / 219 / 2609.24984

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

WorldCrafter:基于隐式3D感知记忆的一致性视频世界模型
Yu, Wangbo, Liu, Kunhao, Hu, Wenbo, Yuan, Shenghai, Feng, Chaoran, Zhou, Haiyang, Huang, Yukun, Wang, Yiran, Zhao, Wang, Luo, Yingmin, Shan, Ying
Abstract
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.
Chinese Translation
视频世界模型能够实现对动态环境的交互式探索,但在长时序和跨视角的场景下,难以与先前观测保持一致。为此,我们提出了WorldCrafter,一种学习可按相机查询的隐式3D感知记忆的视频世界模型。其核心思想是让所请求的视角决定多视角证据如何被压缩到视频生成器有限的token预算中。通过将记忆编码器与位姿条件读取模块和视频生成器进行联合训练,历史观测被整合为一组固定的目标视角特定token,用于去噪之前,且无需显式的基于深度的对应关系。通过将这种记忆与近期的时序上下文以及少步蒸馏相结合,WorldCrafter能够从单张输入图像或文本提示出发进行流式场景探索。在静态和动态场景上的实验表明,该方法在长时序一致性和相机控制精度方面取得了显著提升,同时在分钟级别的探索过程中保持了视觉质量。
cs.CV / 220 / 2609.24997

VideoGen-Agent: Reinforcing Video Generation Agents

VideoGen-Agent:强化视频生成智能体
Li, Binxu, Duan, Haoyi, Zhang, Yuhui, Zhang, Yaohui, Lin, Zihao, Feng, Kaituo, Huang, Suozhi, Li, Xiangyi, Li, Yu, Li, Chunyuan, Liu, Shilong, Wang, Mengdi
Abstract
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. In this paper, we present VideoGen-Agent, a multimodal agent trained through multitask agentic reinforcement learning to use external tools for video generation. The agent coordinates augmentation, generation, and verification tools through multi-turn interactions, using the prompt and intermediate observations to guide its decisions. We train a shared policy on a category-balanced dataset spanning six tasks. Supervised fine-tuning on teacher-generated trajectories establishes tool-use behavior, which is then refined through reinforcement learning. A category-aware hybrid reward evaluates tool-call validity, task-appropriate tool use, and generated video quality. We further introduce VABench, a held-out benchmark of 600 prompts covering procedural knowledge, single- and multi-entity identity preservation, physical consistency, scene composition, and multi-shot temporal structure. On VABench, VideoGen-Agent improves over its base text-to-video generator by 19.1 points, from 56.5 to 75.6. Upgrading the generation tools further raises the score to 86.1 without additional agent training. Human raters prefer the upgraded configuration over the strongest standalone baseline in 84.3% of comparisons. These results support learning tool use across video-generation tasks and show that the trained agent can benefit from subsequent advances in generation tools.
Chinese Translation
视频生成模型的最新进展已实现高保真、时间连贯的视频生成。然而,这些模型往往难以满足需要专业知识、特定身份、物理一致性或事件顺序的提示词要求。本文提出了 VideoGen-Agent,一个通过多任务智能体强化学习训练的多模态智能体,能够使用外部工具进行视频生成。该智能体通过多轮交互协调增强、生成和验证等工具,利用提示词和中间观察结果来指导其决策。我们在一个涵盖六类任务的类别均衡数据集上训练共享策略。首先在教师生成的轨迹上进行监督微调以建立工具使用行为,随后通过强化学习进行优化。类别感知的混合奖励用于评估工具调用有效性、任务适配的工具使用以及生成视频的质量。我们进一步提出了 VABench,一个包含600个提示词的保留测试基准,涵盖程序性知识、单实体和多实体身份保持、物理一致性、场景构图以及多镜头时间结构。在 VABench 上,VideoGen-Agent 相较于其基础文生视频生成器提升了19.1分,从56.5提高到75.6。在不进行额外智能体训练的情况下,升级生成工具可将分数进一步提升至86.1。人类评估者在84.3%的对比中更偏好升级后的配置。这些结果支持跨视频生成任务学习工具使用,并表明训练后的智能体能够从生成工具的后续进步中持续获益。
cs.CV / 221 / 2609.25001

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

GameHorizon Suite:面向多时间跨度的游戏玩法数据与评测
Wang, Yiran, Yin, Xingyilang, Pu, Junfu, Wang, Guangzhi, Li, Kaifeng, Ouyang, Mingyu, Sun, Huiqiang, Li, Lingen, Cheng, Cheng, Yu, Wangbo, Chen, Honghao, Cun, Xiaodong, Pun, Chi-Man, Cao, Zhiguo, Shan, Ying
Abstract
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. GameHorizon Suite consists of three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, utilizing the pipeline, we construct GameHorizon-Data, the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games, collected by 100 human expert players. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation using thousands of standardized questions organized into three primary tasks and a series of diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. Based on our GameHorizon Suite, we evaluate 47 models through more than one million model invocations, revealing a meaningful hierarchy of task difficulty and pronounced differences in model capabilities. Our work can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families. We will release our dataset, annotator, and benchmark to facilitate future research.
Chinese Translation
现代电子游戏为AI模型提供了一个可度量的测试平台,它结合了视觉理解、指令分解、目标规划以及跨多个时间跨度进行精确动作控制等多种能力。然而,现有的数据集和基准测试要么覆盖的游戏种类范围狭窄,要么缺乏语言指令,要么依赖于高方差的在线推演(online rollouts)。为应对这些挑战,我们提出了GameHorizon,一个统一的数据与评测套件,用于衡量不同模型家族在不同时间跨度上的游戏玩法能力。GameHorizon Suite由三个部分组成。第一,GameHorizon-Annotator是一个可扩展、自动化的多跨度指令标注流水线。第二,利用该流水线,我们构建了GameHorizon-Data,这是首个大规模的AAA级游戏玩法数据集,包含时间对齐的视频、玩家动作和多跨度指令。该数据集由100名人类专家玩家采集,涵盖21款游戏共5,000小时的录制内容。第三,我们构建了具有可复现的离线与逐步在线测试的GameHorizon-Bench。离线赛道通过数千个标准化问题(组织为三项主要任务和一系列诊断变体)实现可复现的评测;而在线赛道则检验离线分数是否反映真实的游戏玩法能力,并将失败定位到长跨度游戏过程中的具体步骤。基于GameHorizon Suite,我们通过超过一百万次模型调用评测了47个模型,揭示了任务难度的有意义的层级结构以及模型能力之间的显著差异。我们的工作可以为跨时间跨度和跨模型家族的游戏玩法能力评估提供一个标准化的标尺。我们将发布数据集、标注器和基准测试,以促进未来研究。
机器学习 (Machine Learning)
241
cs.LG / 1 / 2609.22106

PRQuant: Permutation Residual Quantization for Low-Overhead Inference

PRQuant:面向低开销推理的置换残差量化方法
Wang, Peiran, Wang, Anqi, Zhao, Jiaying, Yang, Huiwen, Ming, Zhenyu, Wang, Rongqian, Yao, Yiwu, Tian, Kun, Yao, Xin, Zhang, Gong, Yang, Fan, Huang, Zhongyi
Abstract
Accuracy of Low-bit quantization of linear layers is often dominated by a small number of outliers. Although existing methods, such as smoothing, rotation, or residual-based approaches, may mitigate this problem, they often introduce new accuracy bottlenecks to weights. Besides, most of these techniques are implemented as online approaches, which can result in heavy execution overheads. To address the afore-mentioned issues, We propose PRQuant (Permutation Residual Quantization), a training-free and low-overhead framework that combines channel reorganization with static weight-side residual compensation. After AWQ-style scaling, PRQuant identifies the input channels that contribute most to weight quantization error, permutes them into contiguous tail blocks, and constructs their residual weight sub-tensors offline. During inference, this contiguous structure enables the activation side to use tail blocks seamlessly without the expensive online gathering operation, and turns scattered residual compensation into a regular tail-augmented GEMM, substantially reducing latency. Experiments demonstrate that PRQuant effectively reduces down-projection reconstruction error. Ablation studies confirm that smoothing and residual compensation are the primary drivers of numerical improvement, while permutation provides a consistent marginal numerical benefit and, more importantly, enables a hardware-friendly contiguous layout that eliminates dynamic gathering overhead. Overall, PRQuant outperforms default MXFP4 and the evaluated PTQ baselines in average accuracy across five downstream benchmarks, improving over MXFP4 by 1.24 and 0.55 on Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507, respectively.
Chinese Translation
线性层的低比特量化精度往往由少数离群值主导。尽管现有方法(如平滑、旋转或基于残差的方法)可以缓解这一问题,但它们常常为权重引入新的精度瓶颈。此外,这些技术大多以在线方式实现,可能导致沉重的执行开销。针对上述问题,我们提出了PRQuant(Permutation Residual Quantization,置换残差量化),这是一个无需训练且低开销的框架,将通道重组与静态的权重侧残差补偿相结合。在采用AWQ风格的缩放之后,PRQuant识别出对权重量化误差贡献最大的输入通道,将它们置换到连续的尾部块中,并离线构建其残差权重子张量。在推理时,这种连续结构使激活侧能够无缝使用尾部块,而无需昂贵的在线聚集操作,并将分散的残差补偿转化为规则化的尾部增强GEMM,从而显著降低延迟。实验表明,PRQuant有效降低了下投影的重构误差。消融研究证实,平滑和残差补偿是数值精度提升的主要驱动力,而置换提供了持续的边际数值收益,更重要的是,它实现了硬件友好的连续布局,消除了动态聚集的开销。总体而言,PRQuant在五个下游基准上的平均精度超越了默认的MXFP4以及所评估的PTQ基线方法,在Qwen3-4B-Instruct-2507和Qwen3-30B-A3B-Instruct-2507上分别比MXFP4提升了1.24和0.55。
cs.LG / 2 / 2609.22107

Generalized Multimodal Foundation Model

广义多模态基础模型
Cui, Huizi, Han, Zongbo, Ding, Chenggong, Xiao, Naichuan, Yang, Jialong, Chen, Jingdong, Wang, Guangyu, Hu, Qinghua, Zhang, Changqing
Abstract
Making prediction with multimodal data is widely used in diverse scenarios. Existing multimodal fusion models, once deployed, can only handle predefined modalities (e.g., vision, text and audio) and single tasks, making it difficult to quickly adapt to new downstream applications. Therefore, a natural yet rather aggressive question arises, whether there exists a general multimodal fusion model that can be applied to arbitrary modality combinations and arbitrary prediction tasks. We argue that a unified multimodal fusion model should not depend on specific modalities and instead encode transferable patterns of multimodal correlation. To this end, we propose a simple and effective learning paradigm based on training over the generation of large-scale synthetic multimodal datasets with diverse causal structures that formally characterize the generative processes of multimodal data in real world. Building on this framework, we propose the generalized multimodal foundation model, a unified foundation model for generalized multimodal learning. By constructing large-scale synthetic multimodal datasets with diverse correlation patterns, our model encodes transferable multimodal correlations during training and activates appropriate associations through in-context examples during inference. Extensive experiments on 18 real-world datasets spanning 12 modalities and 11 prediction tasks demonstrate that our model achieves competitive performance with specialized models without task-specific adaptation.
Chinese Translation
利用多模态数据进行预测已广泛应用于各种场景。现有的多模态融合模型一旦部署,只能处理预定义的模态(如视觉、文本和音频)以及单一任务,难以快速适应新的下游应用。因此,一个自然却颇具挑战性的问题随之而来:是否存在一种通用的多模态融合模型,能够应用于任意的模态组合和任意的预测任务。我们认为,统一的多模态融合模型不应依赖于特定模态,而应编码可迁移的多模态关联模式。为此,我们提出了一种简单而有效的学习范式,其核心是在大规模合成的多模态数据集上进行训练,这些数据集具有多样的因果结构,能够形式化地刻画真实世界中多模态数据的生成过程。基于这一框架,我们提出了广义多模态基础模型(generalized multimodal foundation model),一个面向广义多模态学习的统一基础模型。通过构建具有多样关联模式的大规模合成多模态数据集,我们的模型在训练过程中编码可迁移的多模态关联,并在推理过程中通过上下文示例(in-context examples)激活相应的关联。在涵盖12种模态和11种预测任务的18个真实世界数据集上的大量实验表明,我们的模型无需针对特定任务进行适配即可取得与专用模型相当的性能。
cs.LG / 3 / 2609.22108

Correcting Learning-based Perception for Safety

基于学习的感知的安全性校正
Miao, Yan, Darir, Hussein, Mitra, Sayan
Abstract
Learning-enabled perception is important in many autonomous systems. Unlike traditional sensors, the boundary where ML perception does or does not work is poorly characterized. Incorrect perception can lead to unsafe or overtly conservative downstream control actions. In this paper, we propose a two-step strategy for correcting ML-based state estimation. First, an offline computation is used to characterize the uncertainties resulting from the ML module's state estimation, using preimages of perception contracts. Second, at runtime, a risk heuristic is used to choose particular states from the uncertain estimates to drive the control decisions. We perform extensive simulation-based evaluation of this runtime perception correction strategy on different vision-based adaptive cruise controllers (ACC modules), in different weather conditions, and road scenarios. Out of 45 ACC scenarios where the original perception-based control system using Yolo and LaneNet led to safety violations, in 73% of the scenarios, our runtime perception correction preserved safety; our method wouldn't be able to recover 27% of the scenarios where the construction of the preimages of perception contracts is not fully conformant. Further, our runtime perception correction strategy is not overly conservative---on the average only a 2.8% increase in completion time is experienced in the corrected scenarios, with mild interventions.
Chinese Translation
具备学习能力的感知在许多自主系统中至关重要。与传统传感器不同,机器学习(ML)感知有效或失效的边界难以准确刻画。错误的感知可能导致不安全的或过于保守的下游控制动作。在本文中,我们提出了一种两步策略来校正基于机器学习的状态估计。首先,利用感知契约的逆像(preimages of perception contracts)进行离线计算,以刻画机器学习模块状态估计所产生的不确定性。其次,在运行时,使用风险启发式方法从不确定的估计中选取特定状态来驱动控制决策。我们在不同的天气条件和道路场景下,对不同的基于视觉的自适应巡航控制器(ACC 模块),对该运行时感知校正策略进行了大量基于仿真的评估。在原始的基于 Yolo 和 LaneNet 的感知控制系统导致安全违规的 45 个 ACC 场景中,我们的运行时感知校正在 73% 的场景中保持了安全性;而在感知契约逆像的构造不完全符合要求的场景中,我们的方法无法恢复其中的 27%。此外,我们的运行时感知校正策略并不过于保守——在得到校正的场景中,完成时间平均仅增加 2.8%,且干预较为轻微。
cs.LG / 4 / 2609.22109

A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation

在选择性在线策略蒸馏中,共享学习率并非中性控制条件
Zhu, Chencheng
Abstract
Selective on-policy distillation trains a student only at the token positions a selector scores highest, and the literature compares selectors under a single shared learning rate--a control chosen to be neutral. We show it is not. Under LoRA on GSM8K (Qwen2.5-1.5B student, 7B teacher), across an 8x learning-rate grid, dense supervision is statistically flat (swing 1.8 pp, p=0.26) while every selective arm moves with the rate: 5.4 pp for a random 5% subset, 6.7 pp for a total-variation selector, up to 17.7 pp for a teachability selector. Consequently the dense-versus-selective verdict reads 10.1 pp at lr=1e-4 but 5.1 pp at 5e-5--a 2.0x difference decided by a parameter the protocol treats as scenery--and two of six pairwise significance calls between selectors flip between adjacent rates without any rank inversion. We call this selector-rate entanglement and trace it to selection itself rather than step size: AdamW update magnitudes track the rate to within 2.2% despite 15.5x gradient-norm differences across arms. A preregistered frozen-scoring ablation (selection scored by the initial student; criterion, budget, and on-policy rollouts unchanged; 12 seeds per cell) shows live scoring adds 3.79+/-1.69 pp of rate sensitivity (p=0.035) while the frozen arm remains significantly entangled (p=0.015): the feedback loop aggravates the phenomenon rather than causing it. Under full fine-tuning at the rates this literature actually uses (1e-6 to 1e-5) the pattern grows: dense itself swings 19.8 pp, the selective arm 49.5 pp, and the verdict ranges from a non-significant +3.6 pp at the published operating point to +34 pp (p=0.005) one notch hotter. On MATH-500 the rate dependence does not reproduce under LoRA, scoping that result, while the ~10 pp cost of selective training does. We prescribe reporting the arm x rate matrix, not a shared-rate column, as a precondition for selector comparisons.
Chinese Translation
选择性在线策略蒸馏(selective on-policy distillation)仅在选择器评分最高的token位置上训练学生模型,而现有文献通常在单一共享学习率下比较不同选择器——这一控制条件被认为是中性的。我们证明它并非中性。在GSM8K数据集上使用LoRA(学生模型为Qwen2.5-1.5B,教师模型为7B),我们在8倍学习率网格范围内发现:稠密监督在统计上是平坦的(波动1.8个百分点,p=0.26),而所有选择性训练组均随学习率变化:随机5%子集波动5.4个百分点,总变差(total-variation)选择器波动6.7个百分点,可教学性(teachability)选择器最高达17.7个百分点。因此,稠密监督与选择性训练的比较结论在学习率为1e-4时差距为10.1个百分点,而在5e-5时仅为5.1个百分点——这2.0倍的差异由一个被实验协议视为无关紧要的参数决定——并且六个选择器之间的成对显著性判断中有两个在相邻学习率之间发生反转,而未出现任何排序倒置。我们将这一现象称为选择器-学习率纠缠(selector-rate entanglement),并将其归因于选择本身而非步长:尽管各组的梯度范数相差15.5倍,AdamW更新幅度的变化仍控制在学习率的2.2%以内。一项预注册的冻结评分消融实验(由初始学生模型进行选择评分;评分标准、预算和在线策略采样保持不变;每个单元格12个随机种子)表明,实时评分额外增加了3.79±1.69个百分点的学习率敏感性(p=0.035),而冻结组仍然显著纠缠(p=0.015):反馈回路加剧了该现象,而非其成因。在该文献实际使用的全量微调学习率范围(1e-6至1e-5)下,这一模式进一步扩大:稠密监督本身波动19.8个百分点,选择性训练组波动49.5个百分点,结论范围从已发表工作点上不显著的+3.6个百分点,到提高一档学习率后的+34个百分点(p=0.005)。在MATH-500上,学习率依赖性在LoRA下未能复现(我们将该结果限定于此),但选择性训练约10个百分点的代价仍然存在。我们建议报告学习率×选择器的完整矩阵,而非共享学习率下的单列结果,作为选择器比较的前提条件。
cs.LG / 5 / 2609.22113

Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder

迈向用于预测阿片类药物使用障碍药物治疗保留与过早中断的机器学习模型的公平性
Wang, Tongnian, Vivas-Valencia, Carolina, Bauer, Cici, Gong, Yanmin, Choo, Kim-Kwang Raymond, Guo, Yuanxiong
Abstract
Persistent low retention and completion rates in medications for opioid use disorder (MOUD) have driven the use of machine learning (ML) models to predict retention and identify patients at risk of premature discontinuation. However, the fairness of these models across patient populations remains largely unexplored, raising concerns about their application in treatment decision support. This study systematically assesses algorithmic fairness in ML models for predicting MOUD retention and premature discontinuation and investigates the effectiveness of bias mitigation techniques. Using the cross-sectional Treatment Episode Data Set-Discharges (TEDS-D), which includes treatment episodes for individuals in the U.S. discharged between 2015 and 2019, we trained four ML models to predict premature treatment discontinuation and retention beyond 180 days among individuals receiving outpatient MOUD. We evaluated overall performance and subgroup-level error rates across patient subgroups defined by race, ethnicity, age, and sex, complemented by model explanation analyses. We further assessed bias mitigation techniques and their effects on both fairness and predictive performance. Our findings demonstrate that ML models for MOUD outcome prediction can exhibit subgroup-level performance gaps even when overall predictive performance appears acceptable and that bias mitigation can reduce, but not fully eliminate, these gaps without trade-offs. By demonstrating the importance of fairness-aware evaluation and transparent reporting of subgroup performance, this study provides practical insights for the responsible and context-sensitive use of ML models for risk stratification and care prioritization in MOUD treatment settings.
Chinese Translation
阿片类药物使用障碍药物治疗(MOUD)中长期存在的低保留率和低完成率推动了机器学习(ML)模型的应用,以预测治疗保留情况并识别存在过早中断风险的患者。然而,这些模型在不同患者群体中的公平性在很大程度上尚未被探索,这引发了对其在治疗决策支持中应用的担忧。本研究系统评估了用于预测MOUD治疗保留与过早中断的机器学习模型中的算法公平性,并考察了偏差缓解技术的有效性。利用横断面的治疗事件数据集-出院数据(TEDS-D),该数据集包含2015年至2019年间美国出院个体的治疗事件,我们训练了四个机器学习模型,以预测接受门诊MOUD治疗个体的过早治疗中断和超过180天的治疗保留情况。我们评估了总体性能以及按种族、民族、年龄和性别划分的患者子组的错误率,并结合模型解释分析。我们进一步评估了偏差缓解技术及其对公平性和预测性能的影响。研究结果表明,即使总体预测性能看起来可以接受,用于MOUD结局预测的机器学习模型也可能存在子组层面的性能差距;偏差缓解技术可以在不产生权衡代价的情况下缩小但不能完全消除这些差距。通过展示公平性感知评估和子组性能透明报告的重要性,本研究为在MOUD治疗环境中负责任且结合情境地使用机器学习模型进行风险分层和护理优先级排序提供了实用见解。
cs.LG / 6 / 2609.22115

ZoAQ: Adaptive Zeroth-Order Querying via Query-Reuse Coupling

ZoAQ:基于查询复用耦合的自适应零阶查询方法
Feng, Yangyang, Shu, Yao
Abstract
Zeroth-order optimization (ZOO) estimates updates from function evaluations, making perturbation queries a primary cost. Fixed budgets spend the same number of queries at every step, while adaptive controllers may offset their savings by using additional oracle calls to test estimator reliability. We introduce ZoAQ, an adaptive ZOO method built around query reuse. Rather than discarding past evaluations after each step, ZoAQ makes them useful for both the next update and the decision to query further. This enables adaptive query allocation without extra validation queries. Our analysis characterizes when this agreement identifies an update that supports descent and guides the controller to a sufficient query budget. On synthetic objectives, ZoAQ reduces queries by 43-48% relative to fixed baselines using 1.2M queries. In black-box attacks, it reaches 100% success with 320 and 625 average queries on MNIST and CIFAR-10, respectively. Across four OPT fine-tuning settings, ZoAQ saves 43-46% forward evaluations relative to fixed K=4, with accuracy changes within tasks ranging from -0.018 to +0.010.
Chinese Translation
零阶优化(ZOO)通过函数评估来估计更新方向,这使得扰动查询成为主要开销。固定预算方法在每一步都使用相同数量的查询,而自适应控制器可能需要额外的预言机调用以检验估计器的可靠性,从而抵消其节省的开销。我们提出ZoAQ,一种围绕查询复用构建的自适应零阶优化方法。ZoAQ并不在每步之后丢弃过去的评估结果,而是使其既可用于下一次更新,也可用于决定是否进一步查询。这使得自适应查询分配无需额外的验证查询。我们的分析刻画了这种一致性何时能识别出支持下降的更新,并引导控制器达到足够的查询预算。在合成目标函数上,相较于使用120万次查询的固定基线,ZoAQ将查询次数减少了43-48%。在黑盒攻击任务中,ZoAQ在MNIST和CIFAR-10上分别以平均320次和625次查询达到100%的成功率。在四个OPT微调设置中,相较于固定K=4,ZoAQ节省了43-46%的前向评估,且各任务内的准确率变化范围为-0.018至+0.010。
cs.LG / 7 / 2609.22117

LE4Mob: Towards Inductive, Distance-Aware and General-Purpose Location Embedding for Human Mobility Modelling

Wang, Xinglei, Law, Stephen, Zeng, Zichao, Liu, Junyuan, Dong, Guangsheng, Cheng, Tao
Abstract
Location representations provide mobility models with fundamental information about the spatial position, functional characteristics, and relationships of places. However, existing embeddings are often dependent on mobility observations, unable to represent unseen locations, and weakly constrained to retain geographic distance. This limits their reuse across datasets and mobility tasks. To address these limitations, we propose LE4Mob, an inductive, distance-aware, and geography-derived location embedding framework for mobility modelling. LE4Mob extends contrastive language-location pre-training while introducing a distance-aware regularisation objective that encourages the embedding space to preserve spatial relationships. Pre-trained from geographic context, LE4Mob can encode rich spatial-semantic information and generate embeddings for unseen locations inductively. Its independence from downstream mobility task supervision also makes it transferable across different mobility tasks. We evaluate LE4Mob on individual-level next location prediction and population-level commuter flow generation. Experiments across multiple datasets and study areas show that LE4Mob outperforms strong baselines, with particular advantages in inductive settings and when downstream models rely directly on interactions between location embeddings. These findings demonstrate the potential of distance-aware, geography-derived location representations as reusable foundations for human mobility modelling.
cs.LG / 8 / 2609.22120

Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents

成功皆有弯路:为长时程智能体学习可执行的通关指南
Chen, Kaijie, Fang, Chenyu, Yan, Liang, Li, Bo, Zhang, Bo, Ye, Peng
Abstract
Test-time self-evolving agents improve by reusing past experience, yet sparse-reward trajectories contain failures, loops, and detours, while summaries often omit the state conditions and action dependencies needed for execution. We study executable Walkthrough induction from sparse-reward trajectories: extracting compact, state-conditioned, and verifiable procedures. Our key observation is that delayed credit identifies actions associated with progress but cannot determine whether they produce facts required by later actions. We propose Trace, a credit-guided, dependency-grounded framework that compiles noisy trajectories into executable Walkthrough Memory. It detects progress anchors from rewards and persistent state changes, propagates credit to identify valuable transitions, and estimates action prerequisites from cross-episode success and failure evidence. Backward dependency slicing then traces required facts to their producers, extracting dependency-consistent action chains while removing irrelevant loops and detours. The resulting Walkthroughs encode entry conditions, ordered state--action--effect steps, and completion and failure predicates, supporting reuse, intermediate-state resumption, and programmatic verification. Experiments on J-TTL, WebShop, and ScienceWorld with three open-source LLMs show that Trace consistently outperforms eight test-time learning and memory baselines. Compared with the strongest baseline, it improves average AUC and Final-$3$ by $30.0%$ and $40.5%$, respectively, while using fewer inference tokens. These results show that long-horizon interaction benefits more from state-conditioned executable procedures than from complete trajectories or abstract summaries.
Chinese Translation
测试时自进化智能体通过复用过往经验来提升性能,然而稀疏奖励轨迹中包含失败、循环和弯路,而摘要往往遗漏了执行所需的状态条件和动作依赖关系。我们研究如何从稀疏奖励轨迹中进行可执行通关指南的归纳:即提取紧凑的、以状态为条件且可验证的操作流程。我们的关键观察是:延迟 credit 分配能够识别与进展相关的动作,但无法确定这些动作是否产生了后续动作所需的事实。我们提出 Trace,一个基于 credit 引导、依赖关系为基础的框架,将含噪轨迹编译为可执行的通关指南记忆库。该框架从奖励和持久状态变化中检测进展锚点,传播 credit 以识别有价值的转移,并从跨回合的成功与失败证据中估计动作的前提条件。随后通过后向依赖切片追踪所需事实的生成者,提取依赖一致的动作链,同时去除无关的循环和弯路。所得到的通关指南编码了进入条件、有序的状态-动作-效果步骤以及完成与失败谓词,支持经验复用、中间状态恢复和程序化验证。在 J-TTL、WebShop 和 ScienceWorld 上使用三个开源大语言模型的实验表明,Trace 持续优于八个测试时学习与记忆基线方法。与最强基线相比,它分别将平均 AUC 和 Final-$3$ 提升了 30.0% 和 40.5%,同时消耗更少的推理 token。这些结果表明,长时程交互从以状态为条件的可执行流程中获益更多,而非完整的轨迹或抽象的摘要。
cs.LG / 9 / 2609.22121

Modelling daily activity patterns from mobile phone location data via deep representation learning

基于深度表征学习的手机定位数据日常活动模式建模
Wang, Xinglei, Liu, Junyuan, Dong, Guangsheng, Zeng, Zichao, Law, Stephen, Haworth, James, Cheng, Tao
Abstract
Passively collected mobile phone location data provide large-scale, longitudinal observations of human mobility but do not directly reveal activity purposes. The functional characteristics of visited locations offer useful contextual information, yet their relationship with activity purpose remains uncertain, particularly in mixed-use urban environments. We conceptualise activity pattern mining as an integrated process of representation, clustering, and interpretation, and propose the Activity Chain Encoder (ACE) for the representation stage. ACE is a self-supervised model that combines pre-trained urban embeddings with visit timing and duration and uses a Transformer to model the sequential organisation of stays. It is trained using masked activity modelling and identity-guided contrastive learning without requiring deterministic activity purpose labels. Learned daily representations are aggregated into user-level profiles, clustered, and interpreted through temporal-functional patterns and Census-derived demographic context. Applied to mobile phone app location data from London and compared with three representative methods, ACE supports the identification of six differentiated weekday activity-pattern groups characterised by distinct daily rhythms, urban functional contexts, and demographic associations. These complementary forms of evidence further support the development of empirically grounded activity-pattern personas, establishing a holistic route for deriving behaviourally meaningful population patterns from unlabelled mobile phone location data. The source code for the entire analytical pipeline developed in this study is publicly available at https://github.com/xlwang233/ACE.
Chinese Translation
被动采集的手机定位数据为人类移动提供了大规模、纵向的观测,但并不能直接揭示活动目的。到访地点的功能特征提供了有用的背景信息,但其与活动目的之间的关系仍不确定,尤其是在混合功能的城市环境中。我们将活动模式挖掘概念化为表征、聚类与解释的一体化过程,并针对表征阶段提出了活动链编码器(Activity Chain Encoder, ACE)。ACE 是一种自监督模型,它将预训练的城市嵌入与到访时间和停留时长相结合,并使用 Transformer 对停留序列的组织结构进行建模。该模型通过掩码活动建模和身份引导对比学习进行训练,无需确定性的活动目的标签。学到的日级表征被聚合为用户层面的画像,进而进行聚类,并通过时间—功能模式以及来自人口普查的人口统计背景进行解释。将 ACE 应用于伦敦的手机应用定位数据,并与三种代表性方法进行比较,结果表明 ACE 能够识别出六个差异显著的周中活动模式群体,这些群体具有独特的日常节奏、城市功能背景和人口统计关联。这些互补的证据形式进一步支持构建基于实证的活动模式画像,为从无标签的手机定位数据中提取具有行为学意义的人口模式提供了一条整体性路径。本研究开发的完整分析流程源代码已公开于 https://github.com/xlwang233/ACE。
cs.LG / 10 / 2609.22122

Rank Portability Does Not Imply Feasibility Portability: Target-Specific Evaluation of Joint Hardware Constraints

Shu, Wesley
Abstract
Cross-device hardware evaluation often assumes that if architecture rankings transfer across devices, a proxy device can support target-side model selection. We stress-test this assumption for joint latency-energy feasibility across two public architecture families. On NAS-Bench-201, cross-device rank correlations are moderate, while target-comparable feasible-set overlap remains incomplete. A faithful AdaProxy diagnostic substantially improves latency ranking, showing that the observed boundary failures are not simply due to weak adaptation. Exact finite-sample split-conformal analysis also exposes an evidence bottleneck: a finite one-sided 90% threshold requires at least nine calibration observations. We then replicate the phenomenon on 10,000 GPT architectures across 13 HW-GPT-Bench devices. Relative to an RTX3080 proxy, target latency SRCC ranges from 0.951 to 0.996, yet proxy-reuse violation risk ranges from 33.3% to 100% under matched joint constraints. These results show that rank portability, feasibility portability, and target-specific decision support are distinct evaluation objects. Cross-device evaluations should therefore report which target environments actually support the operating point being claimed.
cs.LG / 11 / 2609.22123

StationPDE: Station-Oriented Surface PDE Learning for Multi-Station Multivariate Weather Forecasting

Wang, Xiao, Chen, Changjian, Li, Rongwen, Liu, Hongwu, Fang, Kun, Tang, Zhuo
Abstract
Multi-station multivariate weather forecasting aims to forecast future weather variables at multiple weather stations from historical surface observations. Existing station forecasting models learn statistical dependencies among discrete stations, but lack explicit physical evolution. Meanwhile, PDE-based weather models provide interpretable physical dynamics, yet require continuous fields and upper-air variables unavailable in surface station data. To bridge this gap, we propose StationPDE, a station-oriented surface PDE learning model. StationPDE constructs a terrain-aware continuous surface field from discrete station observations and decomposes its physical evolution into surface wind transport and upper-air inference. Surface wind transport explicitly evolves observable weather variables, while upper-air inference uses learnable horizontal diffusion to approximate the missing influence of unavailable upper-air variables. A parallel data-driven diffusion branch captures complementary motion patterns, and an adaptive router integrates the two forecasts for station-level multivariate forecasting. Experiments on Weather2K and MeteoNet show that StationPDE consistently outperforms state-of-the-art baselines, reducing MSE by about $9.6\%$ on average compared with the strongest baseline. Code and implementation details are available at https://github.com/hnu-vis/StationPDE.
cs.LG / 12 / 2609.22126

SolarFlowRefiner: Refinement-Aware Flow Matching for Surface Solar Radiation Downscaling

SolarFlowRefiner:面向地表太阳辐射降尺度的精化感知流匹配方法
Srivastava, Udbhav, Racheal, Antonita, Chen, Yiheng, Yu, Runlong, Ye, Xinyue
Abstract
High-resolution surface solar radiation (SSR) is important for solar forecasting and grid operation. However, physically consistent reanalysis products are too coarse to resolve localized cloud-driven variability. In this paper, we study a multisource downscaling task that reconstructs high-resolution SolarCube SSR fields from coarse ERA5 radiative variables and co-registered satellite channels. The task is challenging because a single ERA5 grid cell may contain both sunlit and cloud-shadowed regions. As a result, the missing high-resolution correction can be spatially sharp and inherently ambiguous. One-stage predictors often oversmooth these structures. Post-hoc refinement also introduces a stage-wise mismatch: the generator is optimized independently, even though its output determines the refiner's initial state. We introduce SolarFlowRefiner, a refinement-aware flow-matching framework for SSR downscaling. A conditional FlowMatch generator first predicts a normalized correction to an upsampled ERA5 baseline. The refiner is then trained on prediction-conditioned states between the current FlowMatch output and the target residual. This exposes the refiner to the structured errors produced by the generator. The refinement objective is also backpropagated through the FlowMatch sampler, allowing generation and correction to be jointly optimized for the final reconstruction. Experiments on a day-blocked ERA5--SolarCube benchmark show consistent improvements over standalone generation and post-hoc refinement. More broadly, SolarFlowRefiner provides a general strategy for coupling generative predictors with iterative correctors.
Chinese Translation
高分辨率地表太阳辐射(SSR)对太阳能预测和电网运行至关重要。然而,物理一致的再分析产品空间分辨率过粗,无法刻画局地云致变率。本文研究一个多源降尺度任务,即从粗分辨率的ERA5辐射变量和配准的卫星通道数据中重建高分辨率SolarCube SSR场。该任务极具挑战性,因为单个ERA5网格单元可能同时包含阳光照射区和云影区。因此,缺失的高分辨率修正可能在空间上非常剧烈且本质上是模糊的。单阶段预测模型往往会过度平滑这些结构。而事后精化方法则会引入阶段性失配:生成器被独立优化,尽管其输出决定了精化器的初始状态。我们提出SolarFlowRefiner,一个面向SSR降尺度的精化感知流匹配(flow-matching)框架。首先,一个条件化的FlowMatch生成器对上采样的ERA5基线预测一个归一化修正量。然后,精化器在介于当前FlowMatch输出与目标残差之间的、以预测为条件的中间状态上进行训练,从而使其接触到生成器产生的结构化误差。此外,精化目标通过FlowMatch采样器进行反向传播,使生成与修正在最终重建中实现联合优化。在一个按天划分的ERA5—SolarCube基准数据集上的实验表明,本方法相较于独立的生成方法和事后精化方法均取得了一致的提升。更广泛而言,SolarFlowRefiner为将生成式预测器与迭代式修正器相耦合提供了一种通用策略。
cs.LG / 13 / 2609.22129

Helix-FNO: Spectral-Domain Operator Learning Coupled with a High-Fidelity Mechanistic Model for Fast Surrogate Simulation

Helix-FNO:谱域算子学习与高保真机理模型耦合的快速代理仿真方法
Zhao, Jiabao, Wang, Chuwei, Yang, Jinxi
Abstract
Mechanistic simulation models of full-scale treatment processes remain the only trustworthy, extrapolative description of the underlying physico-chemical dynamics, yet their runtime is far too slow to support the thousands of forward evaluations that a modern decision engine requires at a 5-minute decision cadence. The standard remedy-surrogate modelling-often produces a network that learns a single solution for a single configuration, so it generalises poorly to new influent profiles, control settings or plant layouts. This paper presents Helix-FNO, a teacher-student architecture that couples a thirty-two-state mechanistic teacher with a Fourier neural operator (FNO) student. The teacher supplies a high-fidelity dataset of input-field-to-solution pairs, curated by Latin-hypercube and uncertainty-based active learning to cover the boundary and overload regimes that matter in practice; the student learns, in the spectral domain, the solution operator itself rather than any single solution, thereby moving from learning one instance to learning an entire family of equations. We give the operator formulation, the spectral convolution definition, the weighted distillation loss and the active-learning criterion, and we analyse the approximation error of a truncated Fourier expansion with respect to the smoothness of the parametric solution manifold. An illustrative study compares Helix-FNO against a physics-informed network and a data-driven recurrent surrogate on accuracy, dataset efficiency and inference latency, and places the methods on a speed-accuracy Pareto front. The resulting operator is three orders of magnitude faster than the mechanistic teacher at millisecond inference, which is precisely the capability required for massive candidate screening and online decision support.
Chinese Translation
全尺度处理过程的机理仿真模型仍然是对底层物理化学动力学唯一可信、可外推的描述,但其运行速度远远无法满足现代决策引擎在5分钟决策周期内所需的数千次前向评估。标准解决方案——代理建模——通常训练出一个仅针对单一配置学习单一解的网络,因此对新进水水质、控制设置或厂区布局的泛化能力较差。本文提出Helix-FNO,一种将三十二状态机理教师模型与傅里叶神经算子(FNO)学生模型相耦合的师生架构。教师模型提供高保真的输入场到解的配对数据集,通过拉丁超立方采样和基于不确定性的主动学习进行筛选,以覆盖实际中至关重要的边界和过载工况;学生模型在谱域中学习解算子本身而非任何单一解,从而从学习单个实例转向学习整族方程。我们给出了算子公式、谱卷积定义、加权蒸馏损失和主动学习准则,并针对参数化解流形的光滑度分析了截断傅里叶展开的近似误差。一项示例性研究将Helix-FNO与物理信息网络和数据驱动的循环代理网络在精度、数据集效率和推理延迟方面进行了比较,并将这些方法置于速度-精度帕累托前沿。所得算子在毫秒级推理下比机理教师模型快三个数量级,这正是大规模候选方案筛选和在线决策支持所需的能力。
cs.LG / 14 / 2609.22130

Hierarchical Bayesian optimization of an aircraft-based multi-agent system-of-systems

基于飞机的多智能体系统之系统的层次贝叶斯优化
Saves, Paul, Lefebvre, Thierry, Bartoli, Nathalie, Bussemaker, Jasper, Kalliatakis, Nikolaos, Naeem, Nabih, Prakasha, Prajwal
Abstract
Developing innovative system architectures increasingly relies on advanced modeling and optimization techniques to frame the architecting process and define the corresponding computational problems. For complex System-of-Systems (SoS), high-fidelity multiphysics and multidisciplinary simulations are essential for capturing detailed behaviors. However, their computational expense and the risk of evaluation failures make direct optimization challenging. To overcome these limitations, surrogate-based approaches, like Bayesian optimization, have emerged as effective tools for managing expensive, black-box simulation tasks. This work introduces a hierarchical Bayesian optimization framework that leverages Gaussian process meta-modeling to handle discrete architectural choices, conditional dependencies, and heterogeneous design variables inherent to SoS problems. Results show that the hierarchical formulation improves search efficiency and robustness compared to conventional surrogate-based methods, enabling the exploration of large and structurally diverse design spaces with limited simulation budgets. We apply the approach to an aircraft-based multi-agent system for wildfire suppression, a use case developed within the EU-funded COLOSSUS project that illustrates how SoS principles can coordinate heterogeneous aerial platforms with complementary roles, supporting both sustainable mobility and emergency response missions. Our framework provides a scalable methodology for SoS architecting and model exploration, offering transferable insights for applications in aviation, sustainable mobility, and resilience-oriented system design. By combining hierarchical representations with surrogate-based optimization, this work is among the first practical demonstrations of hierarchical Bayesian optimization applied to real-world SoS problems, advancing both methodology and practice.
Chinese Translation
开发创新的系统架构日益依赖于先进的建模与优化技术来构建架构设计过程并定义相应的计算问题。对于复杂的系统之系统,高保真的多物理场与多学科仿真是捕捉详细行为的关键。然而,其高昂的计算成本以及仿真评估失败的风险使得直接优化面临挑战。为克服这些局限,基于代理模型的方法(如贝叶斯优化)已成为管理高成本、黑箱仿真任务的有效工具。本文提出了一种层次贝叶斯优化框架,利用高斯过程元模型来处理SoS问题中固有的离散架构选择、条件依赖关系以及异构设计变量。结果表明,与传统的基于代理模型的方法相比,层次化建模方式提升了搜索效率与鲁棒性,能够在有限的仿真预算下探索规模庞大且结构多样的设计空间。我们将该方法应用于用于野火扑救的基于飞机的多智能体系统,该用例由欧盟资助的COLOSSUS项目开发,展示了SoS原理如何协调具有互补角色的异构空中平台,以支持可持续交通与应急救援任务。我们的框架为SoS架构设计与模型探索提供了一种可扩展的方法论,可为航空、可持续交通以及面向韧性的系统设计等应用领域提供可迁移的见解。通过将层次化表示与基于代理模型的优化相结合,本工作是层次贝叶斯优化应用于真实SoS问题的首批实践示范之一,在方法学与实践层面均有所推进。
cs.LG / 15 / 2609.22145

Weak Ties, Strong Signals: Efficient Training Data Detection in Diffusion LLMs via Independent Token Sampling

弱依赖,强信号:基于独立Token采样的扩散大语言模型高效训练数据检测
Yu, Hongyao, Zhuang, Tianqu, Xu, Ziyuan, Fang, Hao, Hong, Jiaxin, Chen, Bin, Xia, Shu-Tao
Abstract
Diffusion large language models (dLLMs) offer a compelling alternative to autoregressive models, yet they may expose sensitive training data during denoising. Detecting such usage is challenging because dLLMs lack the efficient one-pass probability decomposition of causal architectures. Existing methods rely on random masking to obtain tractable token-wise detection signals under limited query budgets, but fail to control dependencies among masked tokens. We demonstrate that this token-wise approximation introduces a non-negative structural estimation error, which is theoretically characterized by the cumulative conditional mutual information (CMI) among masked tokens and can obscure subtle memorization signals. This insight suggests that reliable detection requires masked token sets with weak internal dependency. To avoid the prohibitive cost of directly estimating CMI over token combinations, we propose \textit{Independent Token Sampling} (ITS), a query-efficient framework that uses an attention-derived pairwise dependency proxy to approximate the CMI-aware selection criterion. ITS further incorporates a diversity-promoting strategy to improve token coverage across sampling rounds, yielding aggregated token-wise signals that are less affected by dependency-induced approximation error. Experiments on multiple datasets show that ITS consistently outperforms state-of-the-art baselines across different models and datasets, achieving an AUC improvement of 0.18 on the ArXiv dataset while maintaining strong performance under limited query budgets. The code is available at https://github.com/Chrisqcwx/DLLM-MIA .
Chinese Translation
扩散大语言模型(dLLMs)为自回归模型提供了一种有吸引力的替代方案,但其可能在去噪过程中泄露敏感的训练数据。检测此类数据使用具有挑战性,因为dLLMs缺乏因果架构那样高效的单遍概率分解。现有方法依赖随机掩码,在有限的查询预算下获得可处理的逐Token检测信号,但无法控制被掩码Token之间的依赖关系。我们证明,这种逐Token近似会引入非负的结构性估计误差,该误差理论上由被掩码Token之间的累积条件互信息(CMI)刻画,并可能掩盖细微的记忆化信号。这一发现表明,可靠的检测需要内部依赖较弱掩码Token集合。为避免直接在Token组合上估计CMI的高昂代价,我们提出了独立Token采样(Independent Token Sampling, ITS),这是一个查询高效的框架,利用基于注意力的成对依赖代理来近似CMI感知的选择准则。ITS还引入了促进多样性的策略,以提高采样轮次中的Token覆盖度,从而使聚合的逐Token信号受依赖性引起的近似误差影响更小。在多个数据集上的实验表明,ITS在不同模型和数据集上一致优于最先进的基线方法,在ArXiv数据集上AUC提升0.18,同时在有限查询预算下仍保持强劲性能。代码已发布于 https://github.com/Chrisqcwx/DLLM-MIA 。
cs.LG / 16 / 2609.22146

GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training

GRRR:解码器大语言模型后训练中的重塑、旋转与路由几何学
Qi, Jianing, Tang, Hao, Zhu, Zhigang
Abstract
We study how post-training changes the weights of Large Language Models (LLMs) relative to their pretrained weights. Across 12 post-training chains with supervised fine-tuning (SFT) and reinforcement learning (RL), we express each weight update in the pretrained matrix's singular value decomposition (SVD) frame. This decomposition separates the changes of three geometrically distinct components: diagonal values, which reshapes singular values; off-diagonal values, which rotates the coupling between pretrained input and output directions; and null-space values, which routes outside the matrix's original nonzero SVD core. On a math evaluation suite, we find that removing the diagonal component usually preserves most of the gains from post-training. These results suggest that post-training gains are carried primarily by reconfiguring and extending pretrained pathways rather than by substantially changing singular values of pre-trained models.
Chinese Translation
我们研究了后训练如何改变大语言模型(LLM)的权重相对于其预训练权重的变化。在涵盖监督微调(SFT)和强化学习(RL)的12条后训练链路中,我们将每次权重更新投影到预训练矩阵的奇异值分解(SVD)坐标系中表达。这一分解将三类几何性质不同的变化分离开来:对角元素,对应奇异值的重塑;非对角元素,对应预训练输入与输出方向之间耦合关系的旋转;以及零空间元素,对应超出矩阵原有非零SVD核心之外的路由。在数学评测套件上,我们发现移除对角分量通常仍能保留后训练所带来的大部分收益。这些结果表明,后训练的收益主要由重新配置和扩展预训练通路所贡献,而非通过大幅改变预训练模型的奇异值来实现。
cs.LG / 17 / 2609.22153

SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs

SafeTune:一个用于审计与修复微调大语言模型安全漂移的统一可信库
Seth, Pratinav, Sadhu, Saisab, Kaushal, Anshul, Sankarapu, Vinay Kumar
Abstract
Methods for addressing safety drift in fine-tuned Large Language Models (LLMs) are scattered across incompatible implementations, lifecycle stages, and evaluation protocols, making them difficult to adopt and compare. We introduce SafeTune, a source-available library that unifies four intervention paradigms: post-hoc weight recovery, safety-constrained fine-tuning, gradient-based unlearning, and inference-time steering, alongside shared interpretability, evaluation, and deployment utilities. SafeTune provides a consistent configuration-driven workflow while preserving the distinct inputs and intervention points each paradigm requires. Its modular registry supports new methods, benchmarks, judges, models, and fine-tuning domains without redesigning the surrounding pipeline. We demonstrate SafeTune through controlled comparisons and finance and medical deployment case studies, showing how it characterizes safety drift, evaluates feasible interventions on common refusal-behavior and capability evaluations, and supports calibrated or layered mitigation.
Chinese Translation
针对微调大语言模型(LLM)安全漂移的现有方法分散于互不兼容的实现、生命周期阶段和评估协议之中,导致其难以采用和比较。我们提出了SafeTune,一个开源可获取的库,统一了四种干预范式:事后权重恢复、安全约束微调、基于梯度的遗忘(unlearning)以及推理时引导(steering),并提供了共享的可解释性、评估与部署工具。SafeTune提供一致的基于配置的工作流程,同时保留各范式所需的特定输入与干预点。其模块化注册机制支持在不重新设计周边流水线的情况下添加新方法、基准、评判器(judges)、模型和微调领域。我们通过受控对比实验以及金融和医疗部署案例研究展示了SafeTune的效用,展示了它如何刻画安全漂移、在统一的拒答行为与能力评估上衡量可行的干预措施,并支持经过校准的或分层式的缓解策略。
cs.LG / 18 / 2609.22154

A Comparative Framework for Evaluating Foundation Models on Tabular Data: A Case Study in Healthcare

Delouee, Majid Lotfian, Veld, Sjors G. J. G. In 't, Schut, Martijn C.
Abstract
Tabular data is the most common format in clinical practice, encompassing laboratory results, medication records, diagnostic codes, and patient demographics. As foundation models for tabular data have grown in number and variety, a practical question has become harder to answer: which model should a clinician or data scientist actually choose for a given task, and why? Existing surveys catalogue what these models can do, but they stop short of providing a structured way to compare them against the specific demands of a real application. We introduce \system{}, a comparative evaluation framework that scores and ranks tabular foundation models (TFMs) across six clinically meaningful dimensions: how well a model generalizes to new datasets, how effectively it protects patient privacy, how much data it needs to perform well, how it scales with growing datasets and feature spaces, how interpretable its predictions are to clinicians, and how fairly it performs across patient subgroups. Each dimension is broken down into measurable sub-components, and groups of sub-components can optionally be combined into supplementary compound scores, called super-metrics, that provide a diagnostic view of how a model performs across several dimensions simultaneously. To show how the framework works in practice, we apply it to two healthcare use cases, screening for iron deficiency and predicting heart failure, demonstrating how the same set of metrics leads to different model rankings depending on what matters most in each clinical context. We also provide a taxonomy of 45 TFMs organized by their underlying architecture, which serves as a reference for researchers and practitioners looking to navigate this rapidly expanding field.
cs.LG / 19 / 2609.22155

From Latent Biomarkers to Clinical Rules: Embedding-Guided Rule Mining and Attribution-Based Translation for Interpretable Tabular Learning

从潜在生物标志物到临床规则:面向可解释表格学习的嵌入引导规则挖掘与基于归因的翻译方法
Delouee, Majid Lotfian, Ayoobi, Hamed, Veld, Sjors G. J. G. In 't, Schut, Martijn C.
Abstract
Clinical decision support tools are most useful when accurate predictions are accompanied by understandable explanations. Rule-based models provide transparency, but rules derived directly from raw clinical measurements may miss patterns arising from interactions between multiple variables. We present a four-step pipeline that mines decision rules in the latent space of an FT-Transformer and translates them back into measurable clinical features. Embedding dimensions that consistently separate patient groups are treated as latent biomarkers, rules are mined using small decision trees, and selected rules are translated using gradient-input saliency and CLS attention attribution. We evaluate the framework on six public clinical and population health datasets at four embedding dimensions. Translated rules outperformed raw-feature rules in five of six datasets, with mean AUROC gains ranging from 0.04 to 0.23. On the heart disease dataset, embedding-space rules reached 0.98 AUROC, but translation reduced this to 0.72, showing that high-performing latent rules cannot always be represented by simple raw-feature conditions. These results show that latent-space rule discovery can uncover predictive patterns while translating them into clinically measurable features that can be evaluated by clinicians.
Chinese Translation
临床决策支持工具在提供准确预测的同时伴随可理解的解释时才最为有用。基于规则的模型具有透明性,但直接从原始临床测量数据中提取的规则可能会忽略由多个变量之间交互产生的模式。我们提出了一个四步流程:在FT-Transformer的潜在空间中挖掘决策规则,并将其转换回可测量的临床特征。能够持续区分患者群体的嵌入维度被视为潜在生物标志物,规则通过小型决策树进行挖掘,选出的规则则利用梯度输入显著性(gradient-input saliency)和CLS注意力归因(CLS attention attribution)进行翻译。我们在四个嵌入维度下,于六个公开的临床和人群健康数据集上对该框架进行了评估。翻译后的规则在六个数据集中的五个上优于原始特征规则,平均AUROC提升范围为0.04至0.23。在心脏病数据集上,嵌入空间中的规则达到了0.98的AUROC,但翻译后降至0.72,这表明高性能的潜在规则并不总是能够用简单的原始特征条件来表示。这些结果表明,潜在空间规则发现能够揭示预测性模式,同时可将其转换为临床可测量、可供临床医生评估的特征。
cs.LG / 20 / 2609.22156

The Limits of Speculation: Bounding Speculative Decoding in Mixture-of-Experts

Amankulov, Aidar, Mamatin, Denis
Abstract
Speculative decoding in Mixture-of-Experts (MoE) models faces the problem of unstable verification cost caused by input-dependent expert loading. To study the physics of this process, we formulate speculation-budget selection as an offline Stochastic Shortest Path (SSP) problem over reference sequences and build a diagnostic Oracle that uses counterfactual simulation to account for MoE verification cost. A detailed analysis of the Oracle's decisions on the Qwen3-Coder and EAGLE-3 pairing, in the space of marginal deltas (Delta Space), shows that rejected candidates form a strict linear boundary. This result demonstrates that a complex global optimization is locally governed by a necessary condition balancing marginal cost against expected progress ($\frac{\Delta \mathbb{E}[Cost]}{\Delta \mathbb{E}[a]}$), providing a rigorous mathematical reference point for designing future adaptive online heuristics.
cs.LG / 21 / 2609.22157

PAGE: Partition-Aware Gated KV-Cache Eviction

Kumar, Pankaj, Mishra, Subhankar
Abstract
KV-cache eviction methods decide which tokens to keep but not whether to evict at all, so a benchmark mean can hide a class of inputs on which compression drives accuracy from 99\% to 0\%. We reframe eviction as a per-input admission decision and show that inputs separate into a capacity-bound class, where eviction is catastrophic at every budget, and a dilution-prone class, where eviction is safe or beneficial. A single label-free scalar computed from prefill attention, the early-to-late drop in pairwise top-$k$ head agreement, predicts this class before any decoding. PAGE thresholds this drop: it applies any base evictor when the drop is large and retains the full cache otherwise, with no training and no accuracy labels. The drop orders inputs by eviction safety consistently across four architecture families, and a per-model unlabeled pilot of about 100 inputs recalibrates the threshold for a new family. Used as a safeguard, PAGE cuts the harm rate on the capacity-bound regime from 0.75 to 0.026, a 29 $\times$ reduction, across four evictors, four models, and two benchmarks, turning a 99\% to 0\% collapse into a flat 89\% without retraining the evictor. The gate is inert wherever eviction is already safe, and the capacity-bound class it protects is a small, identifiable minority of inputs, so the benefit is a targeted safety gain rather than an average one. PAGE is a per-input safeguard, not a compressor: realized compression is $1.8 - 3.4 \times$ (mean 2.9$\times$) against a nominal 16$\times$ budget and decays toward unity by batch 16 under static provisioning, and a trained evictor wins at matched memory.
cs.LG / 22 / 2609.22158

StepKV: Step-Aware KV Cache Compression for LLM Agents

StepKV:面向大语言模型智能体的步感知KV缓存压缩方法
Feng, Boyu, Liu, Jiahong, Li, Yifan, Yu, Wenhao, Qiu, Zexuan, Sun, Yuliang, Shen, Ming, Li, Xiang, Dai, Quanyu, King, Irwin
Abstract
Key-value (KV) caching is essential for efficient autoregressive large language model (LLM) inference, but the cache grows linearly with context length, increasing storage and decoding costs. KV cache compression mitigates this cost by retaining only a subset of cached tokens. This challenge is particularly important for multi-step LLM agents, where a query expands into trajectories of reasoning, tool interactions, and retrieved observations. Existing pruning methods typically treat the cache as a flat token stream and rank tokens by recency or attention saliency. This creates a mismatch between the unit of compression and the unit of reasoning: token-level pruning removes individual entries, whereas useful information in multi-step agents is often organized into reasoning steps with uneven and delayed importance. Consequently, an early observation or intermediate decision may receive little recent attention yet remain essential for later evidence synthesis. We term this failure mode Reasoning Continuity Disruption.These observations motivate KV cache compression that jointly considers token- and reasoning-step-level information. StepKV addresses this goal by treating reasoning steps as first-class retention units. It associates cache entries with their generating steps, estimates step utility from trajectory-derived signals, and combines this utility with token-level saliency. The resulting scores globally rank prunable tokens, from which StepKV retains the top-scoring entries under a target budget. StepKV thus provides a step-centric perspective for agent KV cache compression. Across multi-hop QA and long-horizon web reasoning tasks, StepKV sustains accuracy under low KV budgets where token-level baselines degrade sharply, offering a more robust efficiency-accuracy trade-off for multi-step agent inference.
Chinese Translation
键值(KV)缓存对于高效的自回归大语言模型(LLM)推理至关重要,但缓存大小随上下文长度线性增长,导致存储和解码成本上升。KV缓存压缩通过仅保留缓存标记的一个子集来缓解这一成本。这一挑战对于多步LLM智能体尤为突出,因为一个查询会展开为由推理、工具交互和检索到的观察结果组成的轨迹。现有的剪枝方法通常将缓存视为扁平的标记流,并依据新近度或注意力显著性对标记进行排序。这在压缩单元与推理单元之间造成了不匹配:标记级剪枝移除的是单个条目,而多步智能体中有用的信息往往组织在重要性不均且延迟显现的推理步骤之中。因此,早期的观察结果或中间决策可能近期很少受到注意力关注,但对后续的证据综合却仍然至关重要。我们将这种失效模式称为“推理连续性中断”。这些观察启发我们设计一种同时考虑标记级和推理步骤级信息的KV缓存压缩方法。StepKV通过将推理步骤作为一级保留单元来实现这一目标。它将缓存条目与其生成步骤相关联,从轨迹衍生的信号中估计步骤效用,并将该效用与标记级显著性相结合。由此产生的分数对可剪枝标记进行全局排序,StepKV据此在目标预算下保留得分最高的条目。因此,StepKV为智能体KV缓存压缩提供了以步骤为中心的视角。在多跳问答和长程网络推理任务上,StepKV在低KV预算下仍能保持准确率,而标记级基线方法则急剧退化,为多步智能体推理提供了更鲁棒的效率-准确率权衡。
cs.LG / 23 / 2609.22160

Uncertainty and Business-Aware Remaining Useful Life Estimation for Semiconductor Manufacturing

Frizzo, Davide, Borsatti, Francesco, Susto, Gian Antonio
Abstract
Semiconductor manufacturing relies on tightly interconnected components, so early identification of the assets most likely to fail is essential to prevent a single breakdown from disrupting the entire production pipeline. Maintenance planning must therefore balance unexpected failures against prematurely interrupted operating life. We present a Predictive Maintenance (PdM) framework combining Deep Learning (DL) sequence models and Simoultaneous Quantile Regression (SQR) for uncertainty-aware Remaining Useful Life (RUL) estimation and risk-aware maintenance decisions. Several architectures are compared on ion-milling data from the 2018 PHM Data Challenge (PHM18), including architectures based on State Space Models (SSM), using prediction and business metrics: Unexpected Breaks (UB), Unexploited Lifetime (UL), and a cost-weighted objective. Diagonal State Spaces (S4D) delivers the best Remaining Useful Life (RUL) estimates across quantiles and, relative to Preventive Maintenance (PvM) baselines, substantially lowers business cost by avoiding systematically early interventions. The results support uncertainty-aware, cost-sensitive maintenance planning in semiconductor production.
cs.LG / 24 / 2609.22165

TARGet: Topology-Aware Fusion-based Radio Frequency Circuit Functional Modeling using Graph Neural Networks

Noorzad, Soroosh, Bodero, Sebastian, Fayazi, Morteza
Abstract
Automatic synthesis of analog and Radio Frequency (RF) circuits is an emerging area that requires an efficient circuit modeling method. In recent years, Machine Learning (ML) solutions have played a promising role in this regard. However, many existing ML approaches require separate training data for each circuit topology, even when a single circuit component is added or removed. In addition, they overlook circuit topology information, which limits their ability to capture complex component interactions. Furthermore, they rely on fully connected neural networks with flat feature representations, which require substantial amounts of training data. In this work, we propose an open-source topology-aware RF circuit modeling method, TARGet. Our model considers the circuit at two levels: sub-circuits and the overall circuit topology. At the sub-circuit level, TARGet leverages S-parameter representations to capture sub-circuit behavior rather than relying on individual circuit components, providing a reusable behavioral abstraction for RF building blocks. Moreover, TARGet explicitly incorporates circuit topology information into the model, enabling it to learn across multiple topologies. TARGet introduces a novel fusion-based architecture that integrates Graph Neural Networks (GNNs) and sub-circuit connectivity-aware neural networks to improve data efficiency. Experimental evaluation across multiple RF circuit topologies demonstrates that TARGet achieves sub-1% prediction error while reducing the required training data by up to 35.5x compared to state-of-the-art (SOTA) approaches. Furthermore, TARGet achieves 9.7x higher prediction accuracy under a strict 1% error threshold relative to SOTA models. A held-out matching-network evaluation further demonstrates zero-shot transfer to an unseen sub-circuit topology, where TARGet reduces NMAE by up to 45%.
cs.LG / 25 / 2609.22166

Clustering-Based Collective Anomaly Detection in IoT Systems: A Graph Neural Network Approach

基于聚类的物联网系统集体异常检测:一种图神经网络方法
Khettaf, Dalila, Djenouri, Djamel, Rezaeifar, Zeinab, Djenouri, Youcef
Abstract
The rapid advancement of Internet of Things (IoT) technology has led to the widespread deployment of smart, interconnected devices across a range of domains. However, this expansion has also resulted in a substantial increase in network traffic, creating more opportunities for malicious actors to launch cyberattacks and compromise sensitive information, thereby increasing the need for effective anomaly detection. The state-of-the-art in anomaly detection has predominantly focused on point anomalies. In contrast, the detection of collective anomalies remains relatively under-explored in the literature. In this paper, we introduce Unsupervised Graph Collective Anomaly Detection (UGCAD), a novel frame- work designed to identify collective anomalies in IoT network traffic. Unlike many existing methods, UGCAD operates on graph-structured data without any prior knowledge of group labels or membership. It leverages a variational graph autoencoder (VGAE) to learn the graph representation, which is subsequently used to enhance a clustering algorithm for effective grouping of nodes. To detect collective anomalies, clusters identified as normal are first aggregated and refined, after which anomaly scores are applied to detect collective anomalies. Extensive experiments conducted on the CICIoT2023 and ToN-IoT network datasets demonstrate the effectiveness of UGCAD in both clustering and collective anomaly detection (CAD). Furthermore, comparative evaluations against several traditional and state-of-the-art clustering-based CAD approaches confirm the superiority of UGCAD in accurately detecting collective anomalies.
Chinese Translation
物联网技术的快速发展推动了智能互联设备在众多领域的广泛部署。然而,这种扩张也导致了网络流量的大幅增长,为恶意行为者发动网络攻击和窃取敏感信息创造了更多机会,从而增加了对有效异常检测的需求。目前最先进的异常检测技术主要集中于点异常,相比之下,集体异常的检测在文献中仍相对缺乏探索。本文提出了一种无监督图集体异常检测框架(Unsupervised Graph Collective Anomaly Detection, UGCAD),旨在识别物联网网络流量中的集体异常。与许多现有方法不同,UGCAD 可在无需任何组标签或成员关系的先验知识的情况下处理图结构数据。它利用变分图自编码器(Variational Graph Autoencoder, VGAE)学习图表示,随后利用该表示增强聚类算法以实现节点的有效分组。为检测集体异常,首先对被判定为正常的簇进行聚合和精炼,然后应用异常分数来检测集体异常。在 CICIoT2023 和 ToN-IoT 网络数据集上开展的大量实验证明了 UGCAD 在聚类和集体异常检测方面的有效性。此外,与多种传统及最先进的基于聚类的集体异常检测方法的对比评估进一步证实了 UGCAD 在准确检测集体异常方面的优越性。
cs.LG / 26 / 2609.22167

Role-Aware Morgan Fingerprints for Reaction Yield Prediction

面向反应产率预测的角色感知Morgan分子指纹方法
Mirji, Chinmay, Shekhar, Prashant, Madiyar, Foram, Peng, Hao
Abstract
Predicting reaction yield from molecular structure and reaction context can cut experimental trial-and-error and speed up condition screening in synthetic chemistry. Recent methods for this task use learned representations such as graph neural networks or Transformer encoders over reaction SMILES (Simplified Molecular Input Line Entry System), but these approaches carry heavy preprocessing overhead and can break when input formatting is inconsistent. We propose MFP, a reaction yield prediction method built on role-aware Morgan fingerprints where count-based circular fingerprints are computed for each reaction component, aggregated by chemical role (reactant, reagent, product), and combined with transformation-sensitive difference features into a fixed-length reaction descriptor fed to a feed-forward neural regressor. We test MFP against state of the art methods such as YieldBERT (with and without data augmentation) and GNAN (graph neural network) on the Suzuki-Miyaura and Buchwald-Hartwig benchmarks using a shared preprocessing and evaluation protocol. MFP reaches R2 = 0.878 on Suzuki-Miyaura and R2 = 0.969 on Buchwald-Hartwig while training an order of magnitude faster than graph- or Transformer-based alternatives. A formal complexity analysis confirms that MFP folds all representation cost into a one-time preprocessing step, removing the per-epoch message-passing overhead that graph methods carry. An ablation over fingerprint radius and folded vector length shows that radius-2 representations at nBits =2048 give the best balance of accuracy, speed, and cross-split stability on both datasets. These results establish MFP as an effective, reproducible, and efficient baseline for reaction yield prediction.
Chinese Translation
从分子结构和反应上下文预测反应产率,可以减少合成化学中的实验试错并加速条件筛选。近期针对该任务的方法采用学习型表示,如基于反应SMILES(简化分子线性输入规范)的图神经网络或Transformer编码器,但这些方法预处理开销大,且在输入格式不一致时容易失效。我们提出MFP,一种基于角色感知Morgan指纹的反应产率预测方法:对每个反应组分计算基于计数的圆形指纹,按化学角色(反应物、试剂、产物)进行聚合,并结合对转化敏感的差值特征,生成固定长度的反应描述符,输入前馈神经网络回归器。我们在Suzuki-Miyaura和Buchwald-Hartwig基准数据集上,采用统一的预处理与评估协议,将MFP与YieldBERT(含数据增强与不含数据增强两种版本)和GNAN(图神经网络)等最先进方法进行了比较。MFP在Suzuki-Miyaura上达到R2 = 0.878,在Buchwald-Hartwig上达到R2 = 0.969,同时训练速度比基于图或Transformer的替代方法快一个数量级。形式化的复杂度分析表明,MFP将所有表示计算成本集中于一次性的预处理步骤,消除了图方法在每轮训练中携带的消息传递开销。对指纹半径和折叠向量长度的消融实验显示,在两个数据集上,半径为2、nBits = 2048的指纹表示在精度、速度和跨划分稳定性方面达到最佳平衡。这些结果确立了MFP作为反应产率预测的一种有效、可复现且高效的基线方法。
cs.LG / 27 / 2609.22170

Multiple latent orderings better predict language model preferences

多重潜在排序能更好地预测语言模型的偏好
Chawla, Aviral, Thompson, William H. W., Young, Jean-Gabriel
Abstract
Language models are frequently employed in settings where they are asked to make value judgments and choices. These observed choices often exhibit intransitivity: A model may prefer item $A$ to $B$ and $B$ to $C$, while also preferring $C$ to $A$. Existing work that models LLM preferences treats such inconsistencies as sampling noise around a single latent ordering. We instead propose that intransitivity reflects the aggregation of multiple latent, internally consistent orderings. We first show that observed inconsistencies cannot be explained by a single ordering under any monotone link function. We then introduce a noise-augmented mixture Bradley-Terry (MBT) model that infers latent preference components from repeated pairwise comparisons. Across seven models and four tasks, a mixture of orderings often explains structural inconsistencies better than single-utility models. We find that aggregate preferences often hide underlying preference heterogeneity. A case study on Moral Machine dilemmas shows that models which disagree on aggregate orderings can still share latent components. Together, these results suggest that LLMs reflect plural preferences. Alignment and evaluation pipelines that treat LLM preferences as a single function, therefore, risk averaging over coherent orderings that different users may endorse differently.
Chinese Translation
语言模型经常被用于需要其做出价值判断和选择的场景中。这些观察到的选择往往表现出非传递性:模型可能偏好项 $A$ 胜过 $B$,偏好 $B$ 胜过 $C$,同时却又偏好 $C$ 胜过 $A$。现有的对大语言模型(LLM)偏好建模的工作将这类不一致性视为围绕单一潜在排序的采样噪声。我们则提出,非传递性反映了多个潜在且各自内部一致的排序的聚合。我们首先证明,在任何单调连接函数下,观察到的非一致性都无法由单一排序来解释。随后,我们引入一种带噪声增强的混合 Bradley-Terry(MBT)模型,从重复的成对比较中推断潜在偏好成分。在七个模型和四项任务的实验中,排序混合模型往往比单一效用模型能更好地解释结构性非一致性。我们发现,聚合偏好常常掩盖了潜在的偏好异质性。针对“道德机器”(Moral Machine)困境的案例研究表明,在聚合排序上存在分歧的模型仍可能共享相同的潜在成分。这些结果共同表明,大语言模型反映的是多元偏好。因此,那些将 LLM 偏好视为单一函数的对齐与评估流程,可能会对不同的用户可能各自认同的不同一致排序进行平均,从而带来风险。
cs.LG / 28 / 2609.22173

Industrial Kinematic Trajectory Model (IKTM): Coordinate-Free Autoregressive Generator

工业运动学轨迹模型(IKTM):无坐标系自回归生成器
Amiri, Max, Eyers, David
Abstract
Mobility simulation supports logistics, safety, and communications planning in industrial environments such as ports, mines, and airports. Existing trajectory models, however, rely on absolute coordinates, road-network tokens, or semantic zones: representations that are site-specific and not well suited to unstructured industrial terrain. We introduce the Industrial Kinematic Trajectory Model (IKTM), a coordinate-free trajectory generator that represents industrial vehicle motion through kinematic sequences (speed and heading change) with no absolute spatial reference. IKTM uses an autoregressive causal transformer with probabilistic mixture heads and extends our prior coordinate-free Markovian model with deep sequence modelling and an explicit duration-conditioning signal. Trained on one site and evaluated zero-shot on three unseen sites, it matches the small-turn shape of the empirical turn-rate distributions of held-out telematics; across all four sites, the per-site mean Jensen-Shannon divergence over 100 sampling seeds spans approximately 0.035-0.050 bits under an oracle-length duration-target protocol and approximately 0.032-0.042 bits under a fully zero-shot prior-length protocol, with similar ranges whether out-of-distribution (OOD) duration targets are drawn from each held-out site's empirical length distribution or from the Site A prior. Both protocols stay above the metric's sampling-noise floor (<=0.0049 bits). Paired by site, the oracle-length values are 5.3-6.0x lower than those of a re-implementation of our prior Markovian model under the same 1 Hz protocol. Termination is duration-conditioned rather than spatial: rollouts stop on 100% of trials with a length-tracking error of +0.0 +/- 0.0 s against the sampled target (100 of 100 exactly on target; N=100, T=0.2, untouched in-distribution test split).
Chinese Translation
移动性仿真为港口、矿山和机场等工业环境中的物流、安全和通信规划提供支持。然而,现有的轨迹模型依赖于绝对坐标、路网标记(road-network tokens)或语义区域等表示方式,这些表示具有场地特异性,难以适用于非结构化的工业地形。我们提出了工业运动学轨迹模型(Industrial Kinematic Trajectory Model,IKTM),这是一种无坐标系的轨迹生成器,通过运动学序列(速度和航向变化)来表示工业车辆的运动,而无需任何绝对空间参考。IKTM采用带概率混合输出头的自回归因果Transformer,并在我们先前的无坐标系马尔可夫模型基础上,扩展了深度序列建模和显式的时长条件信号。该模型在一个场地训练,并在三个未见过的场地上进行零样本(zero-shot)评估:它能够匹配留出(held-out)远程信息数据中经验转向率分布的小转角形态;在全部四个场地上,在预言机长度(oracle-length)时长目标协议下,基于100个采样种子的每场地平均Jensen-Shannon散度约为0.035-0.050比特;在完全零样本的先验长度协议下约为0.032-0.042比特。无论分布外(OOD)时长目标是取自各留出场地的经验长度分布,还是取自场地A(Site A)的先验分布,两者的范围均相近。两种协议的结果均高于该指标的采样噪声底(≤0.0049比特)。按场地配对比较,在相同的1 Hz协议下,预言机长度协议的结果比我们先前马尔可夫模型复现版本低5.3-6.0倍。轨迹终止由时长条件驱动而非空间条件驱动:在所有测试中,轨迹生成都100%终止,相对采样目标的长度跟踪误差为+0.0±0.0秒(100次中100次精确命中目标;N=100,T=0.2,未改动的分布内测试集划分)。
cs.LG / 29 / 2609.22175

Contrastive World Models

Li, Bonnie
Abstract
World models trained via pixel reconstruction can struggle in visually complex environments, where irrelevant information dominates the objective and distract the model from information relevant to planning and control. We present Contrastive World Models, an approach for learning latent dynamics models without pixel reconstruction. Building on Dreamer, we replace observation reconstruction in the standard world model objective with a Deep InfoMax-like lower bound that maximizes the mutual information between state-action sequences and local patch features of future observations, encouraging state representations to retain information that is predictive of the future without requiring the model to reconstruct visually irrelevant details. We evaluate our approach in small-scale experiments across three settings of increasing visual complexity. Our method matches Dreamer and a momentum prediction baseline in the default setting, and substantially outperforms both once distractors or natural video backgrounds are introduced, while also training more efficiently by removing the pixel decoder entirely. Our approach is general and makes minimal assumptions beyond access to state-action sequences and future observations. These results suggest that contrastive, infomax-based objectives are a principled and promising direction for building world models that are robust to visual nuisance factors, a property particularly relevant for transferring model-based RL agents to the real world.
cs.LG / 30 / 2609.22177

OpenBlock: Constructive and Verified Content Generation for Adaptive Tile-Matching Games

OpenBlock:面向自适应方块消除类游戏的构造性可验证内容生成
Jun, Jiang
Abstract
Tile-matching puzzle games serve hundreds of millions of players, yet the content-generation algorithms that decide which pieces to present at each turn remain proprietary, and no open platform exists for studying adaptive difficulty in this genre. We present an adaptive tile-matching platform whose central algorithmic contribution is a dual-track content-generation architecture: a deterministic rule-based generator that is always available, and an optional learned generator, both subject to a common verification gate that establishes, by exhaustive sequential-placement search, that every delivered piece set is fully placeable so the learned track can never degrade the constructive-feasibility guarantee of the rule track. A self-play reinforcement-learning placement agent, supervised by auxiliary tasks that expose per-shape placeability to shared representations, is used to diagnose the game's dominant failure mode: at high board fill, long-bar pieces lose the majority of their legal placements. Across 234,000+ self-play episodes the agent reaches a 35.6\% win rate (median score 4,200), and controlled simulation shows that at board fill rates of 70--75\%, 33--56\% of long-bar pieces have no legal placement, while spawn difficulty distributions are statistically indistinguishable between won and lost games---evidence that board-state degeneration, not content difficulty, drives late-game failure. Head-to-head ablations show that per-shape placeability supervision---not aggregate difficulty features---drives the representation gain, and a 14-day online gray rollout (48,000 players; sample-ratio verified, CUPED-adjusted) lifts day-1 retention by 1.8 percentage points and session duration by 7\% over the rule track alone, quantifying the neural track's asymmetric upside in live play.
Chinese Translation
方块消除类益智游戏为数亿玩家提供服务,然而决定每回合呈现哪些方块的难度调整算法仍属于专有技术,且目前尚无开放平台可用于研究该类游戏的自适应难度。我们提出了一个自适应方块消除游戏平台,其核心算法贡献是一种双轨内容生成架构:一个始终可用的确定性基于规则的生成器,以及一个可选的学习型生成器,两者均需通过共同的验证门控——该门控通过穷举式的顺序放置搜索,确保交付的每一组方块都完全可放置,从而使学习型轨道永远无法削弱规则轨道的构造可行性保证。我们训练了一个自博弈强化学习放置智能体,并借助将每个形状的可放置性暴露给共享表示的辅助任务进行监督,用以诊断游戏的主导失败模式:当棋盘填充度较高时,长条形方块会丧失其大部分合法放置位置。在超过23.4万局自博弈中,该智能体达到35.6%的胜率(中位数得分4,200分);受控仿真实验表明,在棋盘填充率达到70%–75%时,33%–56%的长条形方块没有任何合法放置位置,而胜负局之间的方块生成难度分布在统计上无显著差异——这证明是棋盘状态的退化而非内容难度导致了游戏后期的失败。对照消融实验表明,推动表示学习增益的是按形状可放置性监督,而非聚合难度特征;此外,一项为期14天的在线灰度发布(48,000名玩家;经过样本比例校验和CUPED调整)显示,相比仅使用规则轨道,神经轨道将次日留存率提升了1.8个百分点,会话时长提升了7%,量化了神经轨道在真实线上运行中的非对称收益。
cs.LG / 31 / 2609.22178

Beyond Task Completion: Training Capable and Safe Computer-Use Agents

Kang, Zeyu, Yin, Zhenyun, Zhang, Yang, He, Shan, Lei, Shanzhe, Zhong, Yanjiu, Chen, Xinquan, Wang, Yuhong
Abstract
Computer-use agents (CUAs) have made rapid progress in completing complex tasks through graphical user interfaces, yet post-training centered on task success alone does not induce reliable safety behavior. A reliable CUA must condition its execution on risk: it should complete ordinary benign tasks, avoid environmental hazards and continue when a safe completion path remains, and refuse when the goal is harmful or no safe path exists. To learn this conditional policy, we develop Safety and Capability Optimization for Policy Execution (SCOPE), which jointly post-trains a CUA for task-execution capability and safety-aware decision making. To provide aligned training data for this joint objective, we further introduce SCOPE-Gen, an automated pipeline that synthesizes verifiable capability tasks and converts them into paired environment-risk variants while preserving their original goals. Using the resulting tasks, we construct SATraj-OS, a trajectory dataset comprising capability demonstrations, safe continuations, and explicit refusals. SCOPE first learns from all three trajectory types through supervised fine-tuning and then further improves task completion through online reinforcement learning. Starting from Qwen3.5-9B, SCOPE-RL achieves a 54.17% task success rate on OSWorld and a 64.30% attack-avoidance rate on OS-BLIND, yielding the best aggregate capability--safety score of 58.80% among the evaluated agents. Ablations reveal asymmetric but complementary roles for the two forms of safety supervision: refusal trajectories account for most of the attack-avoidance gain, whereas risk-handling trajectories preserve greater task utility at comparable attack-avoidance levels.
cs.LG / 32 / 2609.22182

Gaussian Process Decorrelation for Spatiotemporal Deep Learning-Based Snow Water Equivalent Prediction

基于高斯过程去相关的时空深度学习雪水当量预测方法
Fenster, Colin, Marshall, Adrienne, Bandyopadhyay, Soutir, McKenzie, Daniel
Abstract
In the Western United States, snowmelt is essential to the agricultural industry in addition to being a key source of municipal drinking water. Consequently, accurate snowpack forecasting is critical for water policy and management. Automated Snow Telemetry (SNOTEL) stations provide accurate daily measurements of snow water equivalent (SWE) that exhibit strong correlations in space and in time. We tackle the problem of predicting future SWE values across the SNOTEL network. Specifically, we use a Gaussian Process-based linear transformation to remove spatial correlations before training a long short-term memory (LSTM) neural network on the decorrelated SWE data. This approach allows the LSTM to learn a clean temporal signal at each station. We show that this separation of spatial and temporal components yields better predictive success than multiple baseline models. Furthermore, we incorporate conformal prediction to quantify uncertainty in the resulting SWE forecasts, providing a distribution-free approach to illustrate a potential framework for establishing predictive intervals for spatiotemporal data. Together, accurate point forecasts and distribution-free uncertainty quantification provide a framework for SWE accumulation forecasting on subseasonal scales or projecting SWE with future data while motivating and supporting future work in predicting a large-scale, spatiotemporally complete SWE map.
Chinese Translation
在美国西部,融雪不仅是市政饮用水的重要来源,也对农业产业至关重要。因此,准确的积雪预报对水资源政策和管理具有重要意义。自动积雪遥测(SNOTEL)站点提供精确的雪水当量(SWE)日尺度观测数据,这些数据在空间和时间上表现出强相关性。本文致力于解决对整个SNOTEL网络未来SWE值的预测问题。具体而言,我们在训练长短期记忆(LSTM)神经网络之前,利用基于高斯过程的线性变换去除去相关SWE数据中的空间相关性。这一方法使LSTM能够在每个站点学习到纯净的时间信号。实验表明,这种时空分量的分离方法相比多个基线模型具有更好的预测效果。此外,我们引入共形预测(conformal prediction)来量化SWE预测结果的不确定性,提供了一种免分布(distribution-free)的方法,为建立时空数据预测区间提供了一个潜在框架。准确的点预测与免分布的不确定性量化相结合,为次季节尺度的SWE累积预报以及利用未来数据外推SWE提供了框架,同时也为未来构建大尺度、时空完整的SWE分布图的研究提供了动机与支持。
cs.LG / 33 / 2609.22183

CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds

CleanScore:基于阴性对照与敏感性边界的黑盒基准测试审计
Opoku, Jeffery, Banahene, David
Abstract
Public benchmark scores may reflect skill, prior exposure to the questions, or both, and for most models the training data are unknown. We present CleanScore, a black-box audit using scored outputs only. Each benchmark question becomes a parent item with one public form and two independently written fresh forms preserving its numbers, facts and answer. The audit reports an interval for the public-form advantage rather than a verdict, and a private negative-control bank with an explicit transport radius separates exposure from ordinary form mismatch. A registered controlled-exposure experiment detects planted exposure and stays quiet under fresh-form exposure. A registered audit of five open models on 200 GSM8K and 200 ARC-Challenge items finds no exposure-consistent advantage, bounding surface-form inflation below five points. Registered positive controls then bound what such a null can mean. Leaking an item raises accuracy on paraphrases the model never saw almost as much as on the leaked wording, leaving 52% to 110% of the effect invisible to a paraphrase audit. On ARC a planted 49-point advantage shows an observable gap of -0.020, and about 20 points survive rewriting stem and options, across four training seeds. A surface-form null bounds far less than the phrase contamination audit implies.
Chinese Translation
公开基准测试的分数可能反映模型的真实能力、对题目的先验接触,或两者兼有,而对大多数模型而言,其训练数据是未知的。我们提出CleanScore,一种仅需评分输出即可进行的黑盒审计方法。每道基准测试题目被转换为一个父项,包含一个公开版本和两个独立撰写的全新版本,新版本保留原题的数字、事实和答案。该审计报告的是公开版本相对优势的区间,而非简单判定;同时,一个带有明确迁移半径的私有阴性对照题库可将数据接触与普通的版本表述差异区分开来。一项预注册的受控暴露实验能够检测到人为植入的数据接触,并在全新版本暴露的情况下保持不误报。一项针对五个开源模型在200道GSM8K和200道ARC-Challenge题目上的预注册审计未发现与数据接触一致的优势,将表面形式的虚高限制在五分以内。随后通过预注册的阳性对照界定了此类阴性结果的实际含义:向模型泄露一道题目后,模型在从未见过的改写版本上的准确率提升几乎与在泄露原文上相当,导致52%至110%的效果无法被改写审计所察觉。在ARC数据集上,人为植入的49分优势仅表现出-0.020的可观测差距,且在四个训练随机种子下,约有20分的优势在改写题干和选项后仍然保留。表面形式的阴性结果所能界定的范围,远小于污染审计措辞所暗示的程度。
cs.LG / 34 / 2609.22184

DPTM-DT: Dual-Pretrained Transformer Multitask Representation Learning for Drug-Target Prediction

DPTM-DT:面向药物-靶点预测的双预训练Transformer多任务表示学习
Kong, Ge
Abstract
Drug-target relation prediction supports candidate screening, drug repositioning, and mechanism analysis. Existing models often use incomplete drug or protein representations, model cross-modal interactions shallowly, or train affinity regression and interaction classification separately, although these tasks describe closely related views of the same drug-target pair. This paper presents DPTM-DT, a dual-pretrained Transformer framework for multitask drug-target prediction. DPTM-DT combines GROVER molecular graph embeddings, ESM protein language-model embeddings, and CTD physicochemical descriptors, then exchanges drug-target information through bidirectional cross-modal attention. A shared pair representation is used for continuous affinity regression, high-affinity binary classification, and six-level affinity classification. Experiments on Davis and KIBA cover random 80/20 and DeepDTA-style standard splits. On the random 80/20 split, DPTM-DT achieves MSE/CI values of 0.193/0.917 on Davis and 0.120/0.918 on KIBA. It also reports binary AUPR/MCC values of 0.727/0.654 and 0.798/0.689, and six-class Macro-F1/Top-2 values of 0.800/0.932 and 0.815/0.962 on Davis and KIBA, respectively. Across the reported regression, binary classification, and multiclass classification settings, DPTM-DT achieves the best overall performance among the compared methods. Results under the standard split show the same relative trend. Ablations indicate that dual target representation, gated fusion, and cross-modal attention each contribute to the final performance. Code and supplementary materials are available at: anonymous.4open.science/r/DPCM-DT-74E0.
Chinese Translation
药物-靶点关系预测可支持候选药物筛选、药物重定位和机制分析。现有模型通常使用不完整的药物或蛋白质表示、对跨模态交互的建模较为浅层,或将亲和力回归与相互作用分类分开训练,尽管这些任务描述的是同一药物-靶点对之间密切相关的不同视角。本文提出DPTM-DT,一个用于多任务药物-靶点预测的双预训练Transformer框架。DPTM-DT结合了GROVER分子图嵌入、ESM蛋白质语言模型嵌入以及CTD理化描述符,并通过双向跨模态注意力机制交换药物与靶点信息。共享的药物-靶点对表示被同时用于连续亲和力回归、高亲和力二分类和六级亲和力分类。在Davis和KIBA数据集上的实验涵盖随机80/20划分以及DeepDTA风格的标准划分。在随机80/20划分下,DPTM-DT在Davis上取得MSE/CI为0.193/0.917,在KIBA上为0.120/0.918。在Davis和KIBA上,其二分类AUPR/MCC分别为0.727/0.654和0.798/0.689,六分类Macro-F1/Top-2分别为0.800/0.932和0.815/0.962。在所报告的回归、二分类和多分类设置中,DPTM-DT在对比方法中取得了最佳的整体性能。标准划分下的结果呈现相同的相对趋势。消融实验表明,双重靶点表示、门控融合和跨模态注意力机制均对最终性能有所贡献。代码和补充材料可在以下网址获取:anonymous.4open.science/r/DPCM-DT-74E0。
cs.LG / 35 / 2609.22185

Adaptive Physics-Informed Neural Networks for the Blasius Boundary-Layer Problem

面向Blasius边界层问题的自适应物理信息神经网络
Endalew, Mehari Fentahun, Zhang, Xiaoming John
Abstract
Physics-informed neural networks (PINNs) provide a mesh-free approach for solving differential equations, but their performance can depend strongly on loss weighting, collocation placement, and optimization strategy. This study develops an adaptive PINN framework for the Blasius boundary-layer equation using gradient-norm-based adaptive loss weighting, nonuniform and residual-based collocation, and sequential Adam--L-BFGS optimization. In the representative run using the architecture $[1,100,100,1]$, the model predicts $f''(0)=0.3320762918$, compared with the high-accuracy benchmark $0.332057336215$, giving an absolute error of $1.896\times10^{-5}$. The final weighted loss is $6.789\times10^{-8}$, and the predicted stream-function, velocity, and shear profiles agree closely with an independent numerical boundary-value solution. A separate full-training architecture study shows that the two-hidden-layer model achieves the smallest wall-shear error among the four tested architectures, $1.629\times10^{-6}$, whereas the deepest network attains the smallest weighted objective but a substantially larger wall-shear error. Compared with the previously reported PINN value $f''(0)=0.33165$, the representative run reduces the wall-shear error by approximately a factor of $21.5$. The results show that the combined adaptive training framework can achieve high accuracy for the Blasius problem and that weighted loss alone is insufficient for identifying the most physically accurate PINN. Because the adaptive components are applied jointly, their individual contributions cannot be isolated from the present results and would require a controlled ablation study for separate assessment.
Chinese Translation
物理信息神经网络(PINN)为求解微分方程提供了一种无网格方法,但其性能在很大程度上依赖于损失加权、配点布置和优化策略。本研究针对Blasius边界层方程提出了一个自适应PINN框架,采用基于梯度范数的自适应损失加权、非均匀及基于残差的配点布置,以及Adam—L-BFGS顺序优化策略。在采用$[1,100,100,1]$网络架构的代表性运行中,模型预测得到$f''(0)=0.3320762918$,与高精度基准值$0.332057336215$相比,绝对误差为$1.896 imes10^{-5}$。最终的加权损失为$6.789 imes10^{-8}$,且预测的流函数、速度和剪应力剖面与独立的数值边值解高度吻合。另一项针对完整训练的架构研究表明,在所测试的四种架构中,双隐层模型具有最小的壁面剪应力误差$1.629 imes10^{-6}$,而最深的网络虽取得了最小的加权目标函数值,但其壁面剪应力误差却明显更大。与先前报道的PINN结果$f''(0)=0.33165$相比,代表性运行将壁面剪应力误差降低了约21.5倍。结果表明,该组合自适应训练框架能够在Blasius问题上实现高精度,且仅凭加权损失不足以识别物理上最准确的PINN。由于自适应组件是联合应用的,其各自贡献无法从当前结果中分离出来,需要通过受控的消融实验进行单独评估。
cs.LG / 36 / 2609.22187

Correlation-Guided Flow Matching with Annealed Masking for Spatial Transcriptomics Generation

面向空间转录组学生成的相关性引导流匹配与退火掩码方法
Zhang, Yupei, Chen, Hao, Pan, Li, Li, Chao, Xing, Xiaohan
Abstract
Spatial transcriptomics (ST) provides spatially resolved gene expression profiling but remains expensive, motivating the prediction of ST from histology images. Generative models have emerged as a mainstream paradigm for ST prediction due to their ability to model the conditional distribution of gene expression and capture its inherent stochasticity. However, these methods typically treat genes as independent prediction targets and overlook the intrinsic gene-gene interactions in biological systems, which limits their ability to preserve biologically meaningful co-expression patterns. We argue that gene-gene interactions, which reflect shared pathways and regulatory mechanisms, are essential for generating numerically accurate and biologically coherent ST profiles. In this paper, we propose CorrFlow, a correlation-guided flow matching framework for histology-to-ST prediction that explicitly models gene-gene dependencies through two complementary mechanisms. First, we introduce an annealed masked flow matching strategy, where subsets of genes are progressively masked following a timestep-dependent annealing schedule, encouraging the model to infer masked genes conditioned on the remaining genes and promoting joint conditional modeling beyond per-gene marginal estimation. Second, we devise a gene graph-regularized optimization scheme that integrates prior knowledge from the STRING database and data-driven co-expression estimated by WGCNA to construct a gene affinity graph, which enforces both local consistency and global smoothness in the predicted expression. Extensive experiments across 12 datasets show that CorrFlow achieves the best average PCC and HPCC among evaluated methods, leading to more biologically coherent ST predictions.
Chinese Translation
空间转录组学(Spatial Transcriptomics, ST)能够提供空间分辨的基因表达谱,但成本高昂,这促使研究者尝试从组织学图像预测空间转录组数据。生成模型因其能够对基因表达的条件分布建模并捕捉其内在随机性,已成为空间转录组预测的主流范式。然而,这些方法通常将基因视为独立的预测目标,忽略了生物系统中内在的基因间相互作用,从而限制了其保留具有生物学意义的共表达模式的能力。我们认为,反映共同信号通路与调控机制的基因间相互作用,对于生成数值准确且生物学上一致的空间转录组谱至关重要。本文提出 CorrFlow,一个用于组织学图像到空间转录组预测的相关性引导流匹配框架,通过两种互补机制显式地建模基因间依赖关系。首先,我们引入退火掩码流匹配策略,按照依赖于时间步的退火调度逐步掩码基因子集,鼓励模型在剩余基因的条件下推断被掩码的基因,从而促进超越逐基因边缘估计的联合条件建模。其次,我们设计了一种基因图正则化优化方案,整合来自 STRING 数据库的先验知识与由 WGCNA 估计的数据驱动共表达信息,构建基因亲和图,从而在预测表达中同时强化局部一致性与全局平滑性。在12个数据集上的大量实验表明,CorrFlow 在所评估的方法中取得了最优的平均 PCC 和 HPCC,产生更具生物学一致性的空间转录组预测结果。
cs.LG / 37 / 2609.22191

WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction

Gopakumar, Arin, Pannozzo, Marco
Abstract
Machine learning is being increasingly used to predict where active wildfires will burn the following day, helping inform evacuation boundaries and containment lines. Most models are evaluated using Average Precision (AP), which summarizes performance across all decision thresholds, although acting on a forecast requires choosing one. We benchmarked five discriminative architectures and one generative model on WildfireSpreadTS using a shared evaluation pipeline and two input configurations. We found that model rankings varied depending on whether performance was measured by AP or by threshold-dependent metrics like F1 and IoU. The highest-AP model flagged 4 to 5 times the area that burned and ranked fifth of six on F1 and IoU, and the most recall-heavy model flagged 16 to 23 times. Models with more usable predictions had AP scores 24 to 37 lower. Across architectures, we identified three distinct prediction profiles: over-predicting, balanced, and under-predicting, which AP alone could not distinguish. Expanding the input from 7 to 23 channels changed AP by 0.03 on average, against a 0.21 to 0.24 spread across architectures. These results show AP alone can favor models whose predictions are poorly suited for operational wildfire forecasting.
cs.LG / 38 / 2609.22192

SegTSim: A Big Data Driven Segmented Temporal Simulation Framework for Heterogeneous Multivariate Systems

Li, Xinhang, Geng, Chenxi, Sun, Yujia
Abstract
Heterogeneous multivariate time-series systems exhibit segment-specific nonlinear dynamics that challenge monolithic forecasting architectures. We propose SegTSim, a big-data-driven segmented temporal simulation framework that integrates segment-specific elasticity modeling with adaptive min-gating, dynamic production relocation optimization, multi-factor data fusion with exchange-rate propagation, and a deep ensemble validation pipeline. The framework is validated on US--Japan automotive trade data from USITC repositories spanning 2015 to 2025, comprising approximately 13000 annual records. Under a 25% perturbation scenario, Japanese import volume declines by 20.4% to 0.93 billion USD, while all output variables maintain coefficients of variation below 3.5% across 1000 ensemble inference runs.
cs.LG / 39 / 2609.22194

A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents

Kolekar, Aakash, Genc, Sahika, Sisman, Bunyamin, Shariat, Shahriar, Kachroo, Shree Vandana, Saha, Avishek, Wu, Qianli, Singer, Ari, Dumoulin, Benoit
Abstract
Enterprise analytics agents solve long-horizon tool-use problems over distributed business data, requiring retrieval, reasoning, API calls, code execution, and adaptation to intermediate observations. Supervised fine-tuning (SFT) calibrates tool syntax and teacher-supported behavior, whereas reinforcement learning (RL) can explore reward-supported behaviors beyond demonstrations; applied uniformly, however, RL can perturb already-calibrated skills. We study how to balance SFT and RL under production-mirroring beta APIs. We observe that, in our controlled experiment, checkpoint trajectories retrospectively separated into three regimes: Imitation, where SFT captured reliable teacher behavior; Lift, where both stages helped; and Discovery, where useful reward-observable behavior lay outside reliable teacher support. We leverage this prospectively, using teacher support and reward-observable headroom to route features to SFT only, SFT then RL, increased RL allocation, or further environment development. Across 18 subsequent feature-specific experiments, the diagnostic predicted 15/18 observed trajectories. On GPT-OSS 120B, targeted SFT then RL produced positive point estimates on 7/8 advertiser skills relative to a frontier Control; five positive gains had paired 95% confidence intervals excluding zero, while one skill had a confidence-supported regression. The largest gain was non-disclosure (+11.27 points; 95% CI [+9.72, +12.82]). A separate SME audit surfaced that targeted RL reduces standard leakage from 11.8% to 2.9% and adversarial leakage from 22.9% to 6.8% relative to SFT while preserving actionability (86.2% to 85.7%). In a matched uniform-versus-targeted comparison with shared rewards and optimization, targeted RL improved the seven-skill mean delta from +1.62 to +3.57 while using 43% less incremental RL compute.
cs.LG / 40 / 2609.22196

EvoRank: LLM-Guided Evolution of Multi-Objective Learning-to-Rank Pipelines

Patel, Rayhan, Patel, Shabaz
Abstract
We present EvoRank, an open autonomous ranking engineer: an LLM-guided evolutionary loop that discovers complete Learning-to-Rank pipelines (features, models, losses, ensembles) for multi-objective e-commerce search. On the Expedia ICDM 2013 dataset, with relevance, conversion, and revenue as competing objectives, three independent runs each converge within 50 iterations (about ten dollars) on interpretable pipelines that beat an Optuna-tuned LambdaMART on 60k held-out queries, an advantage that persists at full data scale and places in the top 6 percent of the original competition. A first campaign, evolving only training objectives, builds the central design rule: it appeared to work on its selection fold (the small dataset it uses to pick winners) while a transfer audit, re-scoring winners on held-out data, showed the gains were almost entirely fitness noise (the randomness of its own scoring), and neither seeded domain knowledge nor richer diagnostic feedback changed what transferred. The deciding quantity is measurable in advance: search-space headroom relative to fitness noise. We package this as a headroom gate that predicts, before any LLM spend, whether the loop will pay off, and we release the system, the auditing tools, and a catalog of failure modes with their guardrails, so teams can apply the procedure to their own ranking stacks.
cs.LG / 41 / 2609.22197

Dissecting Hierarchical Reasoning Models: A Mechanistic Study

Rodrigues, Leo Raphael, Kang, Jian
Abstract
We study Hierarchical Reasoning Model (HRM), a representative hierarchical Transformer-based latent reasoning model with many variants, on Sudoku, Maze, and ARC-AGI-2. We mechanistically understand how HRM reasons and what information it encodes. Our analyses compare HRM against Transformer baselines with and without recurrent modules, apply causal interventions on recurrent states, and utilize linear probes against random-direction ablations, as well as sparse autoencoders with feature ablations. Our results reveal several key findings: recurrent models outperform one-pass baselines, while single-state recurrent Transformers are comparable to HRM. State interventions further show that the causal contributions of the high- and low-level states vary across task-specific checkpoints and inference stages. Selected task variables are linearly decodable from the recurrent states in HRM, yet ablating probe directions produce effects comparable to random controls. SAE ablations yield larger behavioral changes than probe-direction ablations. However, top-ranked SAE features show no stable advantage over size-matched random subsets at larger ablation sizes or across tasks; the same pattern persists in a Sudoku control with within-step BPTT. Together, we characterize that HRM is essentially implementing constraint-aware iterative refinement on a puzzle-specific solution state, in which the functional contributions of components at different levels vary without relying on a compact, causally important feature set. These results highlight the necessity of studying the different working mechanisms and the importance of developing mechanistic interpretability techniques better suited for latent-space, recursive reasoning models.
cs.LG / 42 / 2609.22199

Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers

通过在Transformer各层间共享神经专家来提升参数利用率
Jiao, Dian, Duan, Jiaxin, Zhao, Shuai, Leng, Jiabing, Zhang, Yiran, Huang, Feng
Abstract
Transformer-based large language models often suffer from inter-layer parameter redundancy, where functional transformations are redundantly learned across network depths. We propose CS-MoE, a novel Transformer architecture featuring cross-layer expert sharing to address this inefficiency. Deviating from the widely used Mixture-of-Experts (MoE) architecture that terminates each Transformer block with layer-isolated experts, CS-MoE combines layer-independent experts with concurrent access to a centralized, globally shared expert pool. This \textit{Global Experts Sharing} mechanism enables elastic control over token-level parameter activation and computational consumption (FLOPs). Experiments demonstrate that CS-MoE achieves lower perplexity than equal-scale dense Transformers while activating only 55\% of parameters. Furthermore, its performance scales monotonically with an increased number of activated experts and approaches MoE counterparts that consume more FLOPs by expanding the shared pool with a fixed FLOPs budget. CS-MoE also establishes a flexible Pareto frontier between computational cost and model capacity, offering an efficient alternative for computation-constrained environments.
Chinese Translation
基于Transformer的大语言模型常常受到层间参数冗余的困扰,即功能变换在网络各深度中被重复学习。我们提出了CS-MoE,一种新颖的采用跨层专家共享机制的Transformer架构,以解决这一低效问题。与广泛使用的混合专家(Mixture-of-Experts, MoE)架构不同——该架构在每个Transformer块末端使用层隔离的专家——CS-MoE将层独立的专家与可并发访问的集中式全局共享专家池相结合。这种"全局专家共享"(Global Experts Sharing)机制实现了对词元级参数激活和计算消耗(FLOPs)的弹性控制。实验表明,CS-MoE在仅激活55%参数的情况下,取得了比同等规模稠密Transformer更低的困惑度。此外,其性能随激活专家数量的增加而单调提升,并在固定FLOPs预算下通过扩展共享专家池,性能可逼近消耗更多FLOPs的MoE对应模型。CS-MoE还在计算成本与模型容量之间建立了灵活的帕累托前沿,为计算受限的环境提供了一种高效替代方案。
cs.LG / 43 / 2609.22205

Prediction of Nonlinear Oscillations in a Jumping Quarter-Car Model Using Reservoir Computing

基于储备池计算的跳跃式四分之一车模型非线性振荡预测
Watanabe, Masahisa, Dixit, Shiva, Punetha, Nirmal, Chauhan, Swati, Shirimali, Manish Dev
Abstract
Reliable prediction of vehicle dynamics is essential for smart driving applications such as autonomous control and advanced driver-assistance systems. Off-road vehicles used in agricultural and construction settings are particularly prone to nonlinear behavior, including bifurcations and chaotic motion arising from intermittent loss of tire--road contact. Predicting such dynamics is challenging because it requires resolving both smooth nonlinearities and the discontinuous switching associated with contact loss. In this work, we investigate the feasibility of reservoir computing (RC) -- specifically an echo state network (ESN) -- for data-driven prediction of a jumping quarter-car model. The reservoir is trained on time-series data from a small number of points and evaluated on its ability to reconstruct bifurcation diagrams, phase-space attractors, and time trajectories across periodic and chaotic regimes. The trained reservoir qualitatively reproduces the period-doubling route to chaos, captures the geometric structure of periodic and chaotic attractors. These results demonstrate that reservoir computing is a feasible data-driven predictor of nonlinear dynamics in a practical, non-smooth vehicle system.
Chinese Translation
车辆动力学的可靠预测对于自动驾驶控制和高级驾驶辅助系统等智能驾驶应用至关重要。用于农业和建筑场景的非公路车辆特别容易出现非线性行为,包括由轮胎—路面接触间歇性丧失引起的分岔和混沌运动。预测这类动力学具有挑战性,因为它需要同时处理光滑非线性以及与接触丧失相关的不连续切换。本研究探讨了储备池计算(Reservoir Computing, RC)——具体为回声状态网络(Echo State Network, ESN)——用于跳跃式四分之一车模型数据驱动预测的可行性。储备池在少量点的时间序列数据上进行训练,并评估其重构分岔图、相空间吸引子以及周期和混沌状态下时间轨迹的能力。训练后的储备池在定性上重现了通向混沌的倍周期路径,并捕捉了周期吸引子和混沌吸引子的几何结构。这些结果表明,储备池计算是针对实际非光滑车辆系统中非线性动力学的可行数据驱动预测方法。
cs.LG / 44 / 2609.22216

The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families

Twagirayezu, Leonard, Mitra, Prasenjit
Abstract
Quantization enables deployment of large language models on resource-constrained clinical edge devices, but its effect on clinical accuracy and safety remains understudied. We evaluate five 7-8B parameter models at FP16, GPTQ-INT8, and GPTQ-INT4 precision across five benchmarks: MedQA, MedMCQA, Med-HALT, a risk-stratified sample of HealthBench, and MedSafetyBench. The study jointly varies quantization bit width, model family, and clinical task type, with explicit risk stratification and safety measures. INT8 GPTQ is universally safe (max. degradation -1.9%-1.9%), while INT4 degradation is substantial and model-dependent: BioMistral-7B, clinically fine-tuned, loses 19.7% on MedMCQA, more than any general-purpose model, showing clinical fine-tuning does not confer compression robustness. MedMCQA degrades more than MedQA under INT4; Med-HALT is largely unaffected. On HealthBench's emergency-risk subgroup, Qwen2.5-7B degrades by 26.8% under INT4, suggesting high-risk scenarios are disproportionately vulnerable to compression. On MedSafetyBench, the model family dominates over precision (refusal rates range 10.2%-74.9% at FP16), though Qwen2.5-7B (-17.8%) and Meditron-7B (-28.3%) show substantial INT4 safety degradation; notably, Qwen2.5-7B is simultaneously the most accuracy-robust model, demonstrating that accuracy and safety robustness are independent properties. We additionally test two recovery methods, clinical calibration substitution and QLoRA fine-tuning, both producing the same trade-off: MedMCQA recovers while MedQA further degrades, indicating recovery strategies require task-specific validation rather than being assumed universally beneficial. These findings indicate INT8 is broadly safe for clinical deployment, while INT4 safety must be assessed per-model and per-task, and that safety alignment is determined primarily by instruction tuning rather than clinical domain adaptation.
cs.LG / 45 / 2609.22217

UniGIO: Unified Generative Global In-situ Weather Modeling from Spatiotemporal Incomplete Observations

UniGIO:基于时空不完整观测的统一生成式全球站点气象建模
Yang, Songru, Liu, Zili, Han, Tao, Fei, Ben, Bai, Lei, Liu, Chang, Zou, Zhengxia, Ji, Xiangyang, Ouyang, Wanli, Shi, Zhenwei
Abstract
Global In-situ Observation (GIO) provides fine-scale, direct records of the global weather system from sparse point stations, making it an indispensable source for capturing localized and transient dynamics beyond the reach of satellite gridded data, and playing a critical role in key fields such as numerical weather prediction, disaster prevention, and agriculture. However, GIO exhibits strong spatiotemporal incompleteness, severely impairing accurate and real-time in-situ weather modeling. Unlike existing methods waiting for completed AI-ready data with extra introduced errors, in this work, we explore UniGIO, a novel generative framework for directly modeling global in-situ weather dynamics from native incomplete GIO. By generating missing data from observed ones annotated by masks, it unifies the coexisting forecasting, imputation, and generation under arbitrary missing ratios. Between the missing and observed, UniGIO captures station and region level complementarity through the Observation Mixer and Event Aligner, which diffuse discrete observations into continuous spaces where weather processes naturally span multiple stations. We further establish temporal dependencies with pattern shifts using the Adaptive Temporal Mixer, and track extreme events in chaotic local weather systems through a Mixture-ofExperts structure. Steady and extreme events are adapted in decoder by a Local Refiner. Extensive experiments on the up-todate largest global station weather dataset Weather-5K validate its SOTA performance with 11%, 12%, and 5% advantages on accuracy, fidelity, and extreme event capture, delivering a novel holistic solution for weather modeling in GIO networks.
Chinese Translation
全球站点观测(Global In-situ Observation, GIO)通过稀疏的点位站点提供对全球天气系统的精细尺度直接记录,是捕捉卫星网格数据无法触及的局地化和瞬态动力学不可或缺的数据来源,并在数值天气预报、防灾减灾和农业等关键领域发挥着重要作用。然而,GIO具有强烈的时空不完整性,严重影响了准确且实时的站点气象建模。与现有需要等待补全后AI就绪数据并引入额外误差的方法不同,本工作提出UniGIO——一种新颖的生成式框架,可直接从原生不完整的GIO建模全球站点天气动力学。通过根据掩码标注从已观测数据生成缺失数据,该框架在任意缺失比例下统一了并存的预报、插补和生成任务。在缺失与观测数据之间,UniGIO通过观测混合器(Observation Mixer)和事件对齐器(Event Aligner)捕捉站点级与区域级的互补性,将离散观测扩散到天气过程自然跨越多站点的连续空间中。我们进一步利用自适应时间混合器(Adaptive Temporal Mixer)建立带模式偏移的时间依赖关系,并通过混合专家(Mixture-of-Experts)结构在混沌的局地天气系统中追踪极端事件。平稳与极端事件在解码器中由局部精炼器(Local Refiner)自适应处理。在目前最大的全球站点气象数据集Weather-5K上的大量实验验证了其SOTA性能,在准确性、保真度和极端事件捕捉方面分别取得11%、12%和5%的优势,为GIO网络中的气象建模提供了一种新颖的整体解决方案。
cs.LG / 46 / 2609.22218

Toollery: Scaling LLM Agents to Thousands of Skills and Tools

Toollery:将LLM智能体扩展至数千种技能与工具
Tian, Xiangxi, Guan, Ran
Abstract
As LLM agents are exposed to hundreds to tens of thousands of skills, tools, and API functions, full-library prompting becomes costly, slow, and less reliable: each added candidate increases prompt tokens and latency, while longer candidate lists introduce more distractors for LLM selection. We present \textbf{Toollery}, a training-free candidate-compression framework for scalable LLM skill/tool selection. Following established document-side query expansion, Toollery generates user-intent queries from each skill/tool specification and builds a retrieval index that maps real user requests to compact candidate sets before final LLM decision-making. By treating high-level skills and atomic tools as selectable capabilities, Toollery can be applied to both skill libraries and tool registries. We evaluate Toollery on the roughly 79K-capability SkillRouter benchmark, BFCL-V4 with over 440 atomic tools, and 3,396 proprietary smart-cockpit requests over 220 tools. Across these settings, Toollery keeps online selection bounded to a compact top-$k$ candidate set and improves recall over ordinary specification retrieval. At a fixed top-10 budget, Toollery improves end-to-end selection on the cockpit dataset, and maintains comparable AST Accuracy on BFCL-V4. These results support Toollery as a practical candidate-compression framework for large and evolving agent capability libraries, while showing that quality and cost gains depend on workload coverage and provider caching.
Chinese Translation
随着LLM智能体面临数百乃至数万个技能、工具和API函数,全库提示(full-library prompting)变得成本高昂、速度缓慢且可靠性下降:每增加一个候选项都会增加提示词长度和延迟,而更长的候选列表也会为LLM的选择引入更多干扰项。我们提出了Toollery,这是一个无需训练的候选压缩框架,用于可扩展的LLM技能/工具选择。遵循已有的文档侧查询扩展方法,Toollery根据每个技能/工具的规格说明生成用户意图查询,并构建一个检索索引,在最终LLM决策之前将真实用户请求映射为紧凑的候选集合。通过将高层技能和原子工具视为可选能力,Toollery既可应用于技能库,也可应用于工具注册表。我们在约79K能力的SkillRouter基准、包含440多个原子工具的BFCL-V4,以及涉及220个工具的3,396个专有智能座舱请求上对Toollery进行评估。在这些设置中,Toollery将在线选择限制在紧凑的top-$k$候选集合内,并相比普通规格检索提升了召回率。在固定的top-10预算下,Toollery提升了座舱数据集上的端到端选择性能,并在BFCL-V4上保持了相当的AST准确率。这些结果表明,Toollery是一个实用的大规模且不断演进的智能体能力库候选压缩框架,同时表明质量与成本的收益取决于工作负载覆盖范围和提供商缓存。
cs.LG / 47 / 2609.22220

Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles

测量检测器:面向GPU内核基准测试判定器的变异分析
Du, Mingzhe, Luu, Anh Tuan, Huang, Dong, Ng, See-Kiong
Abstract
Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point tolerance, and their verdicts now feed leaderboards and reinforcement-learning rewards. Recent work agrees these checkers are weak and patches them by hand---extra input distributions, fuzzing recipes, tighter tolerances---with no way to \emph{measure} whether any patch suffices. We introduce mutation analysis as an adequacy metric for kernel-benchmark oracles: deterministic rules inject 10{,}303 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}384 of them with an independent kill witness; any test protocol is scored by the fraction it detects. The official check misses \textbf{one in six} witnessed faults (16.9%), deterministically, and the misses are skewed by family: 8.7% of arithmetic faults escape, but 78.6% of precision faults do. The metric explains why (a tolerance blind band growing with reduction size; a measured ceiling on input aggressiveness set by legitimate floating-point variance), audits the strongest existing patch (KernelBench-Verified's gain splits into $+4.0$ points from hidden inputs and $+4.5$ from tighter tolerance, a split its authors could not compute), and exposes a published fuzzing recipe that rejects \emph{correct} kernels 107 times. Optimizing suites over the kill matrix reaches 98.0% detection with two inputs per problem (94.8% held-out), and the measurement's fault taxonomy teaches a test generator more than the raw faults themselves. Across 48 whole architectures, the blindness grows with scale, concentrating in deep homogeneous pipelines, and two problems prove unrefereeable: their official references violate the benchmark's own tolerance against fp64. We release everything as \href{https://huggingface.co/datasets/Elfsong/KernelBench-M}{KernelBench-M}.
Chinese Translation
针对大语言模型(LLM)生成的GPU内核的基准测试通常仅用少量随机输入和宽松的浮点容差来判定正确性,而其判定结果如今被用于排行榜和强化学习奖励。近期研究一致认为这些检测器较为薄弱,并采用人工修补的方式加以改进——例如增加输入分布、模糊测试方案、收紧容差——但没有任何方法可以*衡量*这些修补是否足够。我们引入变异分析作为内核基准判定器充分性的度量指标:通过确定性规则向188个KernelBench问题的经验证的CUDA实现中注入10,303个可编译的缺陷,其中7,384个缺陷带有独立的击杀证据;任何测试协议均按其检测到的缺陷比例进行评分。官方检测方法确定性地漏掉了**六分之一**(16.9%)的有证据缺陷,且漏检率因缺陷类别而异:8.7%的算术类缺陷逃过检测,但精度类缺陷的漏检率高达78.6%。该指标解释了其中的原因(随归约规模增长而扩大的容差盲区;由合法浮点方差决定并经实测得到的输入激进程度上限),审计了现有最强的修补方案(KernelBench-Verified的性能提升可分解为来自隐藏输入的+4.0分和来自更紧容差的+4.5分——这一分解是其作者无法计算的),并揭露了一个已发表的模糊测试方案错误地拒绝了*正确*内核达107次。基于击杀矩阵对测试套件进行优化,每个问题仅需两个输入即可达到98.0%的检测率(留出集上为94.8%),且该度量的缺陷分类学对测试生成器的帮助超过了原始缺陷本身。在48种完整架构上,盲区随规模扩大而增长,并集中于深度同质化流水线中;此外有两个问题被证明无法裁判:它们的官方参考实现在fp64精度下违反了基准测试自身的容差要求。我们将所有成果以 KernelBench-M(https://huggingface.co/datasets/Elfsong/KernelBench-M)发布。
cs.LG / 48 / 2609.22222

Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark

编码智能体能否复现官方统计数据?受控Eurostat基准测试中的元数据、重试预算与执行反馈的局限性
Necula, Sabina-Cristiana
Abstract
Large language models can generate executable data-analysis code, but successful execution is not equivalent to a valid official-statistics result. This study asks whether authoritative metadata and execution feedback improve the reproducibility of Eurostat answers produced by a coding agent, and isolates what execution feedback actually contributes. A benchmark of 30 natural-language tasks covering seven domains, seven Eurostat datasets and four difficulty tiers was run under four conditions: task only (A), task plus a frozen dataset metadata card (B), metadata plus a repair loop driven by sanitized execution feedback (C), and metadata plus the same attempt budget with no diagnostics of any kind (D). Claude Sonnet 5 generated Python through the Anthropic Messages API in three independent replicates, yielding 360 task-runs. Exact correctness required successful execution, the correct dataset, filters, output shape, values and unit. A companion experiment run under an under-specified output contract, in which the required ranking key and unit representation were never stated to the model, understated condition C by 23.4 points, showing that evaluator and contract design can dominate measured agent error. Reliable statistical coding agents need semantic validation against frozen specifications, a fully specified output contract, and a retry budget - not execution diagnostics.
Chinese Translation
大语言模型能够生成可执行的数据分析代码,但成功执行并不等同于得到有效的官方统计结果。本研究探讨权威元数据和执行反馈能否提升编码智能体(coding agent)生成的Eurostat答案的可复现性,并分离出执行反馈的实际贡献。该基准测试包含30个自然语言任务,涵盖七个领域、七个Eurostat数据集和四个难度层级,并在四种条件下运行:仅任务(A);任务加冻结的数据集元数据卡片(B);元数据加由净化执行反馈驱动的修复循环(C);元数据加相同的尝试预算但无任何诊断信息(D)。实验通过Anthropic Messages API使用Claude Sonnet 5生成Python代码,进行三次独立重复实验,共获得360次任务运行。完全正确要求成功执行、正确的数据集、筛选条件、输出形状、数值和单位。一项配套实验在输出契约规定不充分(即从未告知模型所需的排序键和单位表示方式)的条件下进行,其结果将条件C低估了23.4个百分点,表明评估器和契约设计可能主导所测得的智能体误差。可靠的统计编码智能体需要基于冻结规范的语义验证、完全明确的输出契约以及重试预算——而非执行诊断信息。
cs.LG / 49 / 2609.22229

A Synthetic Multivariate Refrigerator Time-Series Dataset for Predictive Maintenance

一个用于预测性维护的合成多变量冰箱时间序列数据集
Benamirouche, Islam, Fass, Feriel, Ziou, Djemel
Abstract
We generated synthetic multivariate time series for 27 refrigerators with a simplified physicsinspired simulator at one-minute resolution. The simulator includes ambient-temperature variation, door use, thermostat and compressor operation, heat exchange, defrost, electrical consumption, and six progressive degradation types. Each refrigerator provides 15 to 20 sensor outputs according to its configuration. The dataset contains 7,066,161 rows in 27 time-series files and 27 failure logs. The release also includes the Python generator, refrigerator configurations, and documentation. The data can support failure prediction, degradation analysis, and learning across refrigerators with different sensor-output sets.
Chinese Translation
我们使用一个简化的物理启发式模拟器,以一分钟分辨率生成了27台冰箱的合成多变量时间序列数据。该模拟器包含环境温度变化、开关门、恒温器和压缩机运行、热交换、除霜、电力消耗以及六种渐进式退化类型。每台冰箱根据其配置提供15至20个传感器输出。该数据集包含27个时间序列文件和27个故障日志,共计7,066,161行数据。本次发布还包括Python生成器、冰箱配置文件及文档。这些数据可支持故障预测、退化分析以及跨不同传感器输出集合的冰箱之间的学习。
cs.LG / 50 / 2609.22230

List Counting Failures Are Not One Phenomenon

Hou, Iyad Ait, Mankarious, Saad, Zirikly, Aya, Hwa, Rebecca
Abstract
Counting the items in a bracketed list looks trivial, yet open-weight chat models often get it wrong. Prior work usually blames input bottlenecks such as subword fragmentation or attention dilution, which predict that different models should fail in roughly the same way. Across seven instruct models on identical prompts, however, wrong answers form distinct modes: Qwen and Gemma 27B often flip odd lengths to a nearby even integer, OLMo concentrates errors on a few mid-sized integers, and Llama tends to under-count. These modes are useful labels rather than a stable family law (Gemma 9B does not reproduce Gemma 27B's odd-to-even drop), and heavier subword fragmentation does not make counting harder on our benchmark. When the model answers incorrectly, a linear probe can usually still recover the true count from the residual stream. Matching the same odd-to-even error also does not imply the same late-MLP magnitude fix: scaling a late MLP output helps Qwen modestly but is near null on Gemma 27B under the same protocol, while residual steering can move both only by trading odd gains for even losses. These results caution against transferring that magnitude fix across models without a transfer check.
cs.LG / 51 / 2609.22232

CNA: An AI-Oriented Comprehensive Normalized Assessment for Healthy Status and Application to Optimize RRT Strategies by Reinforcement Learning

CNA:一种面向AI的健康状况综合归一化评估方法及其在基于强化学习的肾脏替代治疗(RRT)策略优化中的应用
Liu, Jiang, Zhou, Chan, Li, Yujie, Wu, Di, Xie, Yihao, Li, Peiwei, Shu, Xin, Zhu, Jiaqi, Yang, Chunyong, Chen, Yuwen, Yi, Bin
Abstract
Millions worldwide require Renal Replacement Therapy (RRT) as a treatment essential for survival. However, optimizing RRT strategies via AI is challenging due to heterogeneous patient dynamics, missing data, and the absence of an AI-oriented health assessment criterion. We propose an AI-Oriented Comprehensive Normalized Assessment (CNA) for healthy status and apply it to optimize RRT strategies by using offline reinforcement learning (RL). The key idea of CNA is transforming vital-sign distributions into a standard normal space, enabling a unified, data-driven health-status score defined by deviations from referent intervals, which also provides an AI-oriented criterion to assess strategy quality and supports RL termination. We further design a structured 23-dimensional state representation that integrates 19 indicators with 4 RRT descriptors, and employ matrix decomposition to reconstruct missing vital signs, improving data completeness for learning. These components are incorporated into multiple offline RL algorithms and validated via systematic ablation studies on RRT feature subsets. Compared with physicians' observed treatments, the best learned strategy reduces mortality from 13.2% to 5.0% (reducing 62.24%) and shortens average in-hospital stay from 308.5 to 250.1 hours (reducing 18.93%), demonstrating both methodological innovation and the potential of CNA-guided RL to improve RRT outcomes in nephrology.
Chinese Translation
全球有数百万人需要肾脏替代治疗(Renal Replacement Therapy, RRT)作为赖以生存的治疗手段。然而,由于患者动态的异质性、数据缺失以及缺乏面向AI的健康评估标准,通过AI优化RRT策略面临巨大挑战。我们提出了一种面向AI的健康状况综合归一化评估方法(CNA),并将其应用于基于离线强化学习(RL)的RRT策略优化。CNA的核心思想是将生命体征分布转换到标准正态空间,从而实现一个统一的、数据驱动的健康状态评分,该评分由相对于参考区间的偏离程度定义,同时也提供了评估策略质量的面向AI的标准,并支持强化学习的终止判断。我们进一步设计了一个结构化的23维状态表示,整合了19项指标与4个RRT描述符,并采用矩阵分解方法重构缺失的生命体征,提高了用于学习的数据完整性。这些组件被整合到多种离线强化学习算法中,并通过在RRT特征子集上的系统性消融实验加以验证。与医生实际观察到的治疗方案相比,学习到的最优策略将死亡率从13.2%降至5.0%(降低62.24%),并将平均住院时间从308.5小时缩短至250.1小时(缩短18.93%),展示了方法学上的创新以及CNA引导的强化学习在改善肾脏病学领域RRT治疗效果方面的潜力。
cs.LG / 52 / 2609.22233

SCALE: Simulation-Calibrated Amortized Learning for Energy Materials (A hybrid architecture connecting deterministic modeling, real-world data, and transformer-scale inference for accelerated energy-materials discovery)

SCALE:面向能源材料的模拟校准摊销学习(一种连接确定性建模、真实世界数据与Transformer规模推理的混合架构,用于加速能源材料的发现)
Huang, Kuan, Bai, Bo
Abstract
Energy systems face converging pressures for security, affordability, resilience, and sustainability, creating a need for faster discovery of deployable energy materials. Here we introduce SCALE (Simulation-Calibrated Amortized Learning for Energy Materials), a physics-grounded, real-world-data-calibrated learning architecture that connects deterministic scientific operators, experimental calibration, expanded calibrated label generation, and transformer-scale inference. SCALE converts selected high-cost mechanistic computation and measured evidence into reusable models for rapid screening, ranking, inverse design, and active learning. We formulate the framework, identify ten method-based application regimes, and demonstrate SCALE for solid-state metal-hydride hydrogen-storage capacity prediction. In this implementation, a hydride phase-equilibrium capacity operator is calibrated against 381 measured ML-HydPARK capacity anchors and used to generate 5,000 candidate-condition-prototype teacher labels. A crystallographically anchored periodic-graph representation preserves atomic sites, periodic neighbor relationships, and local metal environments absent from formula-only encodings. An edge-biased graph transformer with 2.90 million parameters reproduces calibrated teacher labels with five-fold surrogate fidelity of MAE 0.0582 wt% H2, RMSE 0.0833 wt% H2, R2 = 0.9927, and Pearson r = 0.9963. Post hoc attention analysis suggests that SCALE learns chemically organized element groupings and metal-metal relationships consistent with established hydride chemistry, without chemistry-group labels as supervision. Once trained, SCALE shifts million-candidate evaluation from repeated deterministic workflow execution to batched learned inference, reducing per-candidate screening cost by approximately 10^7-10^8 while retaining links to simulation and experimental evidence.
Chinese Translation
能源系统面临安全、可负担性、韧性与可持续性等多重压力的交汇挑战,因而需要更快速地发现可部署的能源材料。本文提出SCALE(面向能源材料的模拟校准摊销学习),这是一种以物理为基础、经真实世界数据校准的学习架构,连接了确定性科学算子、实验校准、扩展的校准标签生成以及Transformer规模的推理。SCALE将精选的高成本机理计算与实测证据转化为可复用的模型,用于快速筛选、排序、逆向设计与主动学习。我们对该框架进行了系统阐述,识别出十种基于方法的应用模式,并以固态金属氢化物储氢容量预测为例演示了SCALE。在该实现中,氢化物相平衡容量算子基于381个实测的ML-HydPARK容量锚点数据进行校准,并用于生成5,000个候选条件原型教师标签。一种晶体学锚定的周期图表示保留了原子位点、周期性邻居关系以及仅用化学式编码无法体现的局域金属环境。一个具有290万参数的边偏置图Transformer以五折替代精度复现了校准的教师标签:MAE为0.0582 wt% H2,RMSE为0.0833 wt% H2,R2 = 0.9927,Pearson r = 0.9963。事后注意力分析表明,SCALE在没有化学基团标签监督的情况下,习得了与既有氢化物化学一致的化学组织的元素分组及金属-金属关系。训练完成后,SCALE将百万级候选材料的评估从重复执行确定性工作流转变为批量化的学习推理,使每个候选材料的筛选成本降低约10^7–10^8倍,同时保持与模拟和实验证据的关联。
cs.LG / 53 / 2609.22237

Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks

Amballa, Avinash, Saidutta, Yashas Malur, Li, Wenbo, Valkov, Lazar, Chappidi, Srinivas
Abstract
Merging low-rank adapters (LoRAs) promises to eliminate the overhead of swapping task-specific weights at inference time. However, existing merging methods assume every layer needs the same rank budget. Further, some methods assume that rank budget needs to be split equally among the tasks too. We show this uniform-budget assumption is a major source of the performance gap between merged and per-task LoRAs. However, rank selection is an NP hard problem. To this end, we introduce Net Utility, a data free metric that first decomposes every task LoRA by its Singular Value Decomposition (SVD) and scores each of those singular directions by its task utility and its interference with other tasks directions. Next, we globally pool these scores to select singular directions with the highest values with a constraint on the total number of directions selected. The proposed Net Utility metric is applied on top of five different merging methods across three different merging spaces. The merging is done over two sets of tasks, vision and language tasks. Net utility based rank allocation outperforms its counterparts without that allocation. On average, over vision tasks it achieves +2.1% improvement in performance, and +2.2% improvement over the language tasks.
cs.LG / 54 / 2609.22238

Task-Aware Hybrid QUBO Optimization for Structured Neural Network Pruning

面向结构化神经网络剪枝的任务感知混合QUBO优化方法
Orabi, Osama, Zagitov, Artur, Salloum, Hadi, Lobachev, Viktor A., Kholodov, Yaroslav
Abstract
Neural network pruning can be formulated as a combinatorial optimization problem, yet many existing approaches rely on independent filter-importance scores or simplified objective functions. In this work, we propose a Hybrid Quadratic Unconstrained Binary Optimization (QUBO) framework for structured filter pruning that combines task-aware sensitivity information with interactions between candidate filters. The formulation incorporates first-order Taylor sensitivity and Weight-Fisher sensitivity into the linear component of the objective and can additionally incorporate activation similarity into the quadratic interactions. To control the target pruning cardinality without introducing an explicit quadratic cardinality penalty, we use a binary search over the capacity incentive to identify a coefficient that empirically yields the target pruning cardinality. We further investigate a two-stage QUBO--Tensor-Train refinement strategy in which the QUBO solution initializes gradient-free probabilistic black-box optimization to search for improved pruning masks using the downstream metric. Experiments on the SIDD image denoising task and a Half-UNet model show that the Hybrid QUBO achieves higher PSNR and SSIM than the evaluated Taylor and L1-based QUBO baselines at the studied pruning target. Multi-seed experiments under a fixed dataset protocol are used to assess robustness, while controlled sub-problem experiments demonstrate that Tensor-Train refinement becomes increasingly valuable as the combinatorial problem size grows. The results support Hybrid QUBO as a task-aware structured pruning framework for the evaluated setting, while also highlighting the computational and deployment limitations of mask-based pruning.
Chinese Translation
神经网络剪枝可以被表述为一个组合优化问题,然而许多现有方法依赖于独立的滤波器重要性评分或简化的目标函数。在本工作中,我们提出了一种用于结构化滤波器剪枝的混合二次无约束二进制优化(QUBO)框架,该框架将任务感知的敏感度信息与候选滤波器之间的交互作用相结合。该公式化方法将一阶泰勒敏感度和权重-费雪(Weight-Fisher)敏感度纳入目标函数的线性部分,并可以额外将激活相似度纳入二次交互项。为了在不引入显式二次基数惩罚的情况下控制目标剪枝基数,我们对容量激励项进行二分搜索,以找到能够经验性地达到目标剪枝基数的系数。我们进一步研究了一种两阶段的QUBO——张量列车(Tensor-Train)精化策略,其中QUBO解用于初始化无梯度的概率性黑盒优化,以利用下游指标搜索更优的剪枝掩码。在SIDD图像去噪任务和Half-UNet模型上的实验表明,在所研究的剪枝目标下,混合QUBO方法相比所评估的基于泰勒和L1的QUBO基线方法获得了更高的PSNR和SSIM。我们采用固定数据集协议下的多种子实验来评估鲁棒性,同时受控的子问题实验表明,随着组合问题规模的增大,张量列车精化的价值日益凸显。这些结果支持混合QUBO作为所评估场景下的任务感知结构化剪枝框架,同时也指出了基于掩码剪枝在计算和部署方面的局限性。
cs.LG / 55 / 2609.22240

Statistical Inference for Adversarial Training: Central Limit Theorems via Optimal Transport

Jakwang, Kim, Dohyun, Kwon
Abstract
The purpose of this paper is to rigorously quantify the statistical and learning-theoretic properties of adversarial training models for classification. Equivalently, we establish the statistical properties of empirical optimal partial transport. Precisely, first we provide two types of central limit theorems (CLT): CLT centered at the expected empirical value, and CLT centered at the population one with smoothing. These results are based on the uniqueness of optimal potential for various equivalent optimal transport formulations, and the empirical process theory argument. For the binary setting, we indeed prove the uniqueness of optimal potential by leveraging the connection between optimal partial transport and the derived multi-marginal optimal transport formula. As byproducts, we also obtain the stability of a saddle point of the adversarial training model, and the sample complexity and concentration probability of the generalization error.
cs.LG / 56 / 2609.22244

Universal Observatory Graphs for Distributed Sky Coverage and Artificial Intelligence Based Interplanetary Routing

面向分布式天空覆盖与基于人工智能的行星际路由的通用观测台图模型
Razek, Mohammed Abdel
Abstract
This research proposes the Universal Observatory Graph (UOG), an AI-driven framework for distributed astronomical observation across the Solar System. The proposed architecture models autonomous observatories located at the Sun planet L2 Lagrange points as nodes in a weighted graph, while communication links are represented as graph edges characterized by multi-objective physical and operational metrics, including interplanetary distance, communication latency, transmission power, and link reliability. The resulting graph provides a unified mathematical representation of a cooperative interplanetary observatory network. This proposal examines a six-observatory Solar System configuration comprising Earth, Mars, Jupiter, Saturn, Uranus and Neptune. Instantaneous sky coverage is evaluated independently using a 200,000 direction Fibonacci sphere, a 2,000,000 direction fixed seed Monte Carlo calculation and deterministic spherical integration. All three methods yield complete network union coverage, approximately 0.43% complete six observatory intersection and approximately 24.96% mean pairwise Jaccard similarity under the adopted pointing model. Communication routing is subsequently formulated as a finite horizon Markov decision process and solved using tabular Q-learning. The reward balances node participation and a distance dependent reliability proxy against distance, light time latency and a distance squared transmission power proxy. The learned Earth-Saturn-Uranus-Neptune route is also the highest discounted return route among all 41 feasible simple paths under the four hop constraint. The framework provides a reproducible baseline for sequential coverage assessment and multi objective routing; time dependent ephemerides, mission specific visibility, calibrated link budgets and scalable graph policies remain future work.
Chinese Translation
本研究提出了通用观测台图,一个面向太阳系分布式天文观测的人工智能驱动框架。该架构将位于太阳-行星L2拉格朗日点的自主观测台建模为加权图中的节点,而通信链路则表示为图中的边,其特征由多目标的物理与运行指标刻画,包括行星际距离、通信时延、发射功率和链路可靠性。所得到的图为协作式行星际观测台网络提供了统一的数学表示。本方案考察了由地球、火星、木星、土星、天王星和海王星组成的六观测台太阳系配置。瞬时天空覆盖采用三种方法独立评估:200,000个方向的斐波那契球面采样、2,000,000个方向的固定种子蒙特卡洛计算以及确定性的球面积分。在所采用的指向模型下,三种方法均得出网络并集覆盖为完全覆盖、六观测台交集约为0.43%、平均两两Jaccard相似度约为24.96%。随后,通信路由被表述为有限时域马尔可夫决策过程,并采用表格型Q-learning求解。奖励函数在节点参与度和基于距离的可靠性代理指标与距离、光行时延以及距离平方发射功率代理指标之间进行权衡。学到的地球-土星-天王星-海王星路由在四跳约束下也是全部41条可行简单路径中折扣回报最高的路由。该框架为序列化覆盖评估和多目标路由提供了可复现的基线;随时间变化的星历、任务特定的可见性、经过校准的链路预算以及可扩展的图策略仍是未来的工作。
cs.LG / 57 / 2609.22247

CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents

CHART:面向工具链鲁棒搜索代理的工具链轮换课程训练方法
Zhang, Xinlu, Lin, Ying-Chun, Zhang, Zhihan, Fetahu, Besnik, Chen, Xi
Abstract
Search agents are usually trained under a single harness. But once an agent is deployed in a real application, its harness is frequently updated (e.g., a rewritten system prompt) to fit production needs. This exposes a fragility of post-trained agents: because a learned behavior is entangled with its training harness, even a harness update that leaves the task unchanged can fail to elicit the behavior. We train a search agent to perform parallel search, a popular strategy for improving both search efficiency and performance. We find that training under a fixed harness makes the behavior harness-local, overfit to that harness's surface form: when the harness changes, the model falls back to serial search. An intuitive fix is harness augmentation, but simply training on more harnesses does not resolve the problem. GRPO learns from the reward gap between parallel and serial rollouts of the same question: a small harness pool saturates that gap early, while a large pool dilutes the per-harness signal too thinly for any harness to consolidate. We therefore propose Curriculum HArness Rotation Training (CHART), a rotating curriculum that lets a search agent gradually consolidate parallel search across harnesses. At each periodic evaluation, CHART "graduates" the harnesses whose expected behavior is learned and replaces them with still-learnable ones, keeping the reward gap alive throughout training. Starting from the same harness pool, CHART makes the model learn parallel search on all harnesses, whereas static augmentation succeeds on at most half of them. The behavior also carries to held-out harnesses: CHART parallelizes on 89% of held-out turns, against at most 5% for the static pools. It further transfers to a new QA task and search environment, improving pass@1 by 5.6pp over the best static pool. Finally, CHART-trained agents benefit more from meta-harness search than baselines.
Chinese Translation
搜索代理(search agent)通常在单一工具链(harness)下训练。然而,一旦代理被部署到实际应用中,其工具链会频繁更新(例如重写系统提示词)以适应生产需求。这暴露出后训练代理的一个脆弱性:由于习得的行为与训练工具链纠缠在一起,即使是任务本身未变、仅仅更新工具链的改动,也可能无法触发该行为。我们训练一个搜索代理执行并行搜索——这是提升搜索效率与性能的常用策略。我们发现,在固定工具链下训练会使该行为局限于特定工具链,即过拟合于该工具链的表面形式:当工具链改变时,模型会退回到串行搜索。一个直观的解决方案是工具链增强,但仅在更多工具链上训练并不能解决该问题。GRPO 从同一问题在并行与串行 rollout 之间的奖励差距中学习:较小的工具链池会过早地使该差距饱和,而较大的工具链池则会将每个工具链的信号稀释得过薄,导致任何工具链上的行为都难以巩固。因此,我们提出课程式工具链轮换训练(Curriculum HArness Rotation Training,CHART),这是一种轮换式课程方法,使搜索代理能够逐步在多个工具链上巩固并行搜索行为。在每次周期性评估中,CHART 会"毕业"那些预期行为已被习得的工具链,并用仍具学习空间的工具链替换它们,从而在整个训练过程中保持奖励差距的活跃。从相同的工具链池出发,CHART 使模型在所有工具链上都学会并行搜索,而静态增强方法最多只能在一半工具链上成功。该行为还能迁移到保留工具链上:CHART 在 89% 的保留轮次上实现并行化,而静态工具链池最多仅为 5%。该方法还能进一步迁移到新的问答任务和搜索环境,相比最佳静态工具链池将 pass@1 提升 5.6 个百分点。最后,与基线相比,经 CHART 训练的代理能从元工具链搜索(meta-harness search)中获得更多收益。
cs.LG / 58 / 2609.22251

Predictors and Orchestrators: Parsimonious Machine Learning within an Agentic AI Harness for Multi-Horizon Karst Aquifer Forecasting

预测器与编排器:智能体AI框架内的简约机器学习用于多时间尺度岩溶含水层预报
Lekhak, Pramod, Sharma, Chetan, Başağaoğlu, Hakan, Bertetti, F. Paul, Chakraborty, Debaditya
Abstract
Forecasting karst aquifer dynamics is difficult because recharge responses are nonlinear, event-driven, and governed by strongly heterogeneous flow paths. This study develops and evaluates a deployment-aware framework for 1-12-week-ahead prediction of spring discharge and groundwater level using approximately 79 years of hydroclimatic observations from the Edwards Aquifer, Texas. Five model families were compared under a common temporal evaluation design: extreme gradient boosting, extremely randomized trees, long short-term memory, convolutional neural networks, and Transformers. Predictions were evaluated using coefficient of determination, Kling-Gupta efficiency, root-mean-square error, and agreement with operational drought thresholds. Extreme gradient boosting was consistently most reliable, with R2 at least 0.97, 0.96, and 0.94 across 1-4-, 5-8-, and 9-12-week horizons, respectively, and greater than 90% critical-stage agreement at the first three drought stages across all horizons. Deep models were competitive at short horizons but degraded progressively and exhibited isolated failures at longer lead times. We attribute this contrast to an alignment between tree partitioning and low-dimensional, axis-aligned hydroclimatic predictors, together with the tendency of neural models to smooth irregular extremes. The validated models were embedded in a five-agent operational architecture that automates data acquisition, model assignment, deterministic prediction, threshold monitoring, prospective verification, literature retrieval, and reporting. The contribution is therefore a transferable framework joining parsimonious model selection, leakage-aware multi-horizon evaluation, decision-relevant threshold skill, and auditable agentic automation.
Chinese Translation
岩溶含水层动态预报十分困难,因为补给响应是非线性的、事件驱动的,并受强烈非均质流径的控制。本研究基于德克萨斯州爱德华兹含水层约79年的水文气候观测数据,开发并评估了一个面向部署应用的框架,用于提前1-12周预报泉水流量和地下水位。在统一的时间评估设计下,比较了五种模型族:极端梯度提升(XGBoost)、极端随机树、长短期记忆网络(LSTM)、卷积神经网络(CNN)和Transformer。预报结果采用决定系数(R²)、Kling-Gupta效率系数、均方根误差以及与业务干旱阈值的吻合度进行评估。极端梯度提升始终最为可靠,在1-4周、5-8周和9-12周预报尺度上的R²分别不低于0.97、0.96和0.94,且在所有预报尺度上,前三个干旱阶段的临界水位吻合率均超过90%。深度模型在短预报尺度上具有竞争力,但在更长预见期内性能逐渐下降,并出现孤立性失效。我们将这种差异归因于树的划分方式与低维、轴向对齐的水文气候预测变量之间的契合,以及神经模型对不规则极值进行平滑的倾向。经过验证的模型被嵌入一个由五个智能体组成的业务化架构中,该架构可自动化完成数据获取、模型分配、确定性预报、阈值监测、前瞻性验证、文献检索和报告生成。因此,本研究的贡献在于提出了一个可迁移的框架,将简约的模型选择、防泄漏的多时间尺度评估、面向决策的阈值预报能力以及可审计的智能体自动化相结合。
cs.LG / 59 / 2609.22252

CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation

CALM:一种用于基于活动的出行者模拟的校准大语言模型选择网络框架
Cheng, Yezhou
Abstract
We present CALM, a reproducible hybrid framework that integrates an optional large language model (LLM) activity planner with calibrated stochastic choice, shared network feedback, memory and habit, typed feasibility checks, and deterministic offline replay. Unlike trip-mode classifiers or diary-only generators, CALM executes a closed traveler-day loop and evaluates each generative module against an empirical, reproducible baseline. On the 2024 New York City Citywide Mobility Survey (CMS), 110,691 seven-mode trips are split by respondent into 78,487 training and 32,204 holdout trips. Training-only alternative-specific constant calibration reduces mean holdout mode Jensen-Shannon divergence from 0.15599 to 0.00394 across ten seeds. A matched live-LLM ablation then quantifies trade-offs among aggregate fit, temporal fit, behavioral persistence, and feasibility, while frozen prompt-response pairs support deterministic replay of downstream simulation. Controlled weather, delay, fare, and parking ladders further demonstrate consistent and interpretable responses under intervention. CALM contributes a reproducible protocol for integrating and evaluating generative planners in traveler simulation through person-disjoint calibration, matched module ablation, controlled stress testing, and end-to-end traceability.
Chinese Translation
我们提出了CALM,一个可复现的混合框架,它集成了可选的大语言模型(LLM)活动规划器,并结合了校准的随机选择、共享网络反馈、记忆与习惯、类型化可行性检查以及确定性的离线重放。与出行方式分类器或仅基于出行日记的生成器不同,CALM执行一个闭环的出行者一天模拟流程,并将每个生成模块与经验性的、可复现的基准进行对比评估。在2024年纽约市全市出行调查(CMS)数据上,110,691条七种出行方式的行程按受访者划分为78,487条训练行程和32,204条保留测试行程。仅基于训练数据的备选方案特定常数校准,使十个随机种子下保留集出行方式的平均Jensen-Shannon散度从0.15599降至0.00394。随后,匹配的在线LLM消融实验量化了总体拟合、时间拟合、行为持续性与可行性之间的权衡,同时冻结的提示-响应对支持下游模拟的确定性重放。受控的天气、延误、票价和停车阶梯实验进一步证明了在干预条件下的一致且可解释的响应。CALM通过人员不相交的校准、匹配的模块消融、受控压力测试以及端到端的可追溯性,为在出行者模拟中集成和评估生成式规划器提供了一个可复现的协议。
cs.LG / 60 / 2609.22253

CAMFT: Conflict-Aware Mergeable Fine-Tuning for Large Language Models

CAMFT:面向大语言模型的冲突感知可合并微调方法
Zhou, Jingang, Guo, Haiyang, Ma, Yuan, Zhu, Han, Zhang, Xu-Yao
Abstract
Model merging has emerged as a promising paradigm for integrating multiple task-specific capabilities into a single large language model. However, existing methods predominantly focus on post-hoc processing of independently fine-tuned models, overlooking how the training phase itself impacts cross-task compatibility. Resolving parameter conflicts after fine-tuning is inherently sub-optimal. To address this, we propose CAMFT, a Conflict-Aware Mergeable Fine-Tuning method that makes task adaptation both efficient and mergeaware. CAMFT treats mergeability as a property shaped during fine-tuning, rather than only a problem to be solved after fine-tuning. By guiding each task to update sparse coordinates with lower cross-task conflict, CAMFT produces task updates that are efficient to train and more compatible for downstream model merging. Extensive experiments demonstrate that CAMFT outperforms standard finetuning baselines in multi-task merging scenarios. Codes are available at https://github.com/gyanchow/CAMFT-LLM.
Chinese Translation
模型合并已成为将多个任务特定能力集成到单个大语言模型中的一种有前景的范式。然而,现有方法主要侧重于对独立微调后的模型进行事后处理,忽视了训练阶段本身对跨任务兼容性的影响。在微调之后再解决参数冲突本质上是次优的。为此,我们提出了CAMFT(Conflict-Aware Mergeable Fine-Tuning,冲突感知可合并微调)方法,使任务适配既高效又具备合并意识。CAMFT将可合并性视为在微调过程中塑造的属性,而非仅在微调后需要解决的问题。通过引导每个任务在跨任务冲突较低的稀疏坐标上进行更新,CAMFT产生的任务更新不仅训练高效,而且在下游模型合并时具有更好的兼容性。大量实验表明,在多任务合并场景中,CAMFT优于标准的微调基线方法。代码已发布于 https://github.com/gyanchow/CAMFT-LLM。
cs.LG / 61 / 2609.22254

Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation

教师应当未雨绸缪:面向可靠在策略蒸馏的自适应续写
Zhou, Jingang, Zhou, Yuyi, Guo, Haiyang, Wang, Xukai, Feng, Shuai, Gao, Sirui, Xu, Jian, Guo, Qingpei, Zhang, Xu-Yao
Abstract
On-policy distillation (OPD) is a promising approach for transferring knowledge between language models, where a student receives dense token-level supervision along its own generated trajectories. However, teacher supervision can be unreliable when conditioned on incomplete or low-quality student prefixes. We identify Teacher Uncertainty Contraction (TUC), a systematic phenomenon whereby the teacher's predictive uncertainty decreases as it continues from a student-generated prefix. We theoretically characterize this trade-off through a variance-bias decomposition of teacher-branch gradients, showing that uncertainty contraction reduces variance while teacher-student path divergence increases bias, thereby favoring a finite continuation. Guided by this insight, we propose Adaptive-Continuations On-Policy Distillation (AC-OPD), which augments informative states along student rollouts with teacher continuations and adaptively selects their effective supervision horizons. Experiments on mathematical reasoning and code generation across model scales demonstrate that AC-OPD consistently improves over standard OPD. Controlled-continuations and matched-budget analyses further validate the adaptive-continuations design, highlighting adaptive teacher continuations as an effective principle for reliable on-policy distillation.The code will be made publicly available upon publication.
Chinese Translation
在策略蒸馏(On-Policy Distillation, OPD)是一种在语言模型之间迁移知识的有效方法,学生模型沿自身生成的轨迹获得密集的词元级监督。然而,当教师模型以不完整或低质量的学生前缀为条件时,其监督可能并不可靠。我们识别出一种系统性现象——教师不确定性收缩(Teacher Uncertainty Contraction, TUC),即教师模型在从学生生成的前缀继续生成时,其预测不确定性会随之降低。我们通过对教师分支梯度的方差-偏差分解,从理论上刻画了这一权衡:不确定性收缩降低了方差,而师生路径分歧则增加了偏差,因此宜采用有限长度的续写。基于这一洞察,我们提出了自适应续写在策略蒸馏(Adaptive-Continuations On-Policy Distillation, AC-OPD),该方法在学生模型的采样轨迹上用教师续写扩充信息丰富的状态,并自适应地选择其有效监督范围。在数学推理与代码生成任务上、跨不同模型规模的实验表明,AC-OPD 持续优于标准的 OPD。受控续写与预算匹配分析进一步验证了自适应续写设计的有效性,凸显了自适应教师续写作为实现可靠在策略蒸馏的一项有效原则。代码将在论文发表后公开。
cs.LG / 62 / 2609.22257

Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning

面向自动化大语言模型微调的策略积累与引导执行
Zhao, Haoran, Du, Wei, Yang, Dingwen, Huang, Jixuan, Shang, Junlin, Fang, Lingyong, Guo, Ya, Gui, Tao, Zhang, Qi, Huang, Xuanjing
Abstract
Producing task-specific large language models requires discovering effective training strategies through experimentation. Automated fine-tuning systems have made this experimentation feasible with far less manual effort. However, these systems are stateless: each search discards its discovered strategies, dataset insights, and hyperparameter findings once it ends. Every new task must then repeat this costly search from a cold start. To address this, we propose Strategy Accumulation and Guided Execution (SAGE), a two-stage framework that makes automated fine-tuning search cumulative. In the first stage, a multi-agent pipeline performs Monte Carlo Tree Search-based exploration. A parallel Distillation Agent extracts task-specific exploration records and confidence-scored cross-task insights, which together constitute a structured experience repository. In the second stage, SAGE retrieves relevant experience from this repository and selects what applies to guide training on the new task. We evaluate SAGE on nine unseen tasks spanning both single- and cross-category settings. In single-round execution, SAGE's accumulated experience raises the average relative improvement over baseline from 3.2% to 15.6%, a 12.4-percentage-point gain over the same pipeline without it. These results show that persistent strategy experience provides effective guidance for automated fine-tuning on unseen tasks.
Chinese Translation
构建面向特定任务的大语言模型需要通过实验探索有效的训练策略。自动化微调系统使得这类实验所需的人工投入大幅减少。然而,这些系统是无状态的:每次搜索结束后,其发现的策略、数据集洞察以及超参数结论都会被丢弃。每个新任务都必须从冷启动开始重复这一昂贵的搜索过程。为解决这一问题,我们提出了策略积累与引导执行(Strategy Accumulation and Guided Execution,SAGE),这是一个使自动化微调搜索具有累积性的两阶段框架。在第一阶段,多智能体流水线执行基于蒙特卡洛树搜索(Monte Carlo Tree Search)的探索。并行运行的蒸馏智能体(Distillation Agent)提取任务特定的探索记录以及带有置信度评分的跨任务洞察,二者共同构成一个结构化的经验库。在第二阶段,SAGE从该经验库中检索相关经验,并选择适用于指导新任务训练的内容。我们在九个未见任务上评估了SAGE,涵盖单类别与跨类别两种设置。在单轮执行中,SAGE积累的经验将相对于基线的平均相对提升从3.2%提高到15.6%,相比不使用该经验的相同流水线提升了12.4个百分点。这些结果表明,持久化的策略经验能够为未见任务上的自动化微调提供有效指导。
cs.LG / 63 / 2609.22258

RS-Claw-Evolution: Environment-Feedback-Driven Evolution for Lightweight Remote Sensing Agents in Long-Horizon Tasks

RS-Claw-Evolution:面向长时程任务中轻量级遥感智能体的环境反馈驱动进化方法
Ouyang, Kai, Hou, Dongyang, Liu, Liangtian, Wang, Zeyuan, Li, Ziyu, Liu, Chengfu, Tang, Zichao, Cui, Xuezhi, Ouyang, Shengwu, Yang, Wentao, Yu, Hanwen, Li, Haifeng
Abstract
Large language model-driven remote sensing (RS) agents offer a promising approach to automating geospatial analysis. However, lightweight RS agents based on compact language models struggle with multi-step interactive tasks due to loss of long-horizon states, inefficient environmental feedback utilization, and sparse optimization signals. We propose RS-Claw-Evolution, an environment-feedback-driven framework that progressively improves lightweight agents through three stages. Interaction evolution uses executable code to control observations, maintain intermediate states, and reduce context redundancy. Experience evolution combines failure-aware trajectory generation with error-turn masking to learn from informative failure-recovery experiences without imitating faulty actions. Decision evolution uses reinforcement learning with multi-dimensional environment rewards and turn-level advantage protection to optimize tool-use behaviors and improve credit assignment in long sequences. On Earth-Bench, the optimized Qwen3-4B-based agent achieves 65.9% accuracy in Autonomous Planning mode, outperforming the untrained Qwen3-32B baseline (43.8%) and DeepSeek-V3.1 (60.8%), while approaching GPT-5 (71.6%). These results demonstrate that learning from environmental feedback can improve lightweight agents and narrow their performance gap with larger models in long-horizon RS tasks.
Chinese Translation
基于大语言模型驱动的遥感(RS)智能体为实现地理空间分析自动化提供了一种有前景的途径。然而,基于紧凑语言模型的轻量级遥感智能体在多步交互任务中面临诸多困难,其原因包括长时程状态丢失、环境反馈利用效率低下以及优化信号稀疏。我们提出了RS-Claw-Evolution,一个通过三个阶段逐步提升轻量级智能体性能的环境反馈驱动框架。交互进化阶段利用可执行代码来控制观测、维护中间状态并减少上下文冗余。经验进化阶段将失败感知的轨迹生成与错误回合掩码相结合,从有价值的失败恢复经验中学习,同时避免模仿错误动作。决策进化阶段采用强化学习方法,结合多维环境奖励与回合级优势保护机制,优化工具使用行为并改进长序列中的信用分配。在Earth-Bench上,基于Qwen3-4B优化后的智能体在自主规划模式下达到65.9%的准确率,超过了未训练的Qwen3-32B基线(43.8%)和DeepSeek-V3.1(60.8%),并接近GPT-5(71.6%)。这些结果表明,从环境反馈中学习能够提升轻量级智能体的性能,并缩小其在长时程遥感任务中与更大模型之间的性能差距。
cs.LG / 64 / 2609.22359

Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch

抗拒、更新、拒绝:偏好优化安装了一个依赖先验的可靠性开关
Yang, Sen, Yeung, Yuen-Hei
Abstract
An aligned model asked to hold its answer against a manipulative source must still update on a reliable one and reject an unreliable one: resistance, reliable-update, and unreliable-source rejection are one three-way contract, not three independent behaviors. We show the objective most anti-sycophancy work optimizes is non-identifying with respect to source reliability: because no preference label depends on whether a source is actually reliable, any scalar mixture of the arms traces a single deference dial, and no point separates two same-template testimonies differing only in stated reliability. This fixation$\leftrightarrow$gullibility frontier is a property of the objective, not any model. We make reliability identifiable through data: a threshold benchmark where a source asserts the opposite answer while stating its reliability $r$, and the correct action is to flip iff $r$ exceeds the model's prior strength $p$. Preference optimization over balanced coverage installs a prior-dependent reliability switch: across three seeds on Qwen2.5-7B-Instruct the threshold $r^\star$ rises monotonically with the prior, decision accuracy reaches $0.84$ with a monotone flip curve (Spearman $0.56$), and the policy generalizes to unseen reliability values and a held-out notation, following stated reliability over role prestige. Three controls localize the cause: an unmatched variant installs the switch equally ($0.80$), a second preference optimizer (IPO) installs it just as well ($0.86$), whereas supervised imitation does not ($0.50$), so the cause is preference optimization over reliability-labeled coverage, not pairing, loss, or imitation. A confirmatory battery replicates the switch on a fresh test draw, bounds it honestly (it keys on reliability stated in the testimony, not a separately audited record), and transfers it to Llama-3.1-8B. The frontier is empirical, not a theorem.
Chinese Translation
一个经过对齐的模型在面对操纵性来源时应坚持其答案,同时仍应在来源可靠时进行更新、在来源不可靠时予以拒绝:抗拒、可靠更新与拒绝不可靠来源是一个三方契约,而非三种独立行为。我们证明,大多数反谄媚工作所优化的目标函数在来源可靠性方面是不可辨识的:由于没有任何偏好标签依赖于来源是否真正可靠,各分支的任何标量混合都只能描绘出一个单一的顺从度调节旋钮,并且没有任何一点能够区分仅在被声明的可靠性上不同、而模板相同的两条证词。这一'固执↔轻信'边界是目标函数本身的属性,而非任何特定模型的属性。我们通过数据使可靠性变得可辨识:构建一个阈值基准,其中来源在被声明的可靠性 $r$ 下给出与正确答案相反的回答,而正确动作是当且仅当 $r$ 超过模型的先验强度 $p$ 时才翻转答案。在均衡覆盖数据上进行偏好优化会安装一个依赖先验的可靠性开关:在 Qwen2.5-7B-Instruct 上的三个随机种子实验中,阈值 $r^\star$ 随先验单调上升,决策准确率达到 $0.84$ 并呈现单调的翻转曲线(Spearman 相关为 $0.56$),且策略能够泛化到未见过的可靠性取值和留出的记号体系,遵循被声明的可靠性而非角色声望。三个对照实验定位了成因:一个不匹配的变体同样能安装该开关($0.80$),第二个偏好优化器(IPO)也能同样有效地安装它($0.86$),而监督模仿则不能($0.50$),因此成因是在带有可靠性标注的覆盖数据上进行偏好优化,而非配对方式、损失函数或模仿。一组验证性实验在新的测试样本上复现了该开关,对其边界作了诚实的界定(它依赖于证词中被声明的可靠性,而非经单独审计的记录),并将其迁移到 Llama-3.1-8B 上。该边界是经验性的,而非一个定理。
cs.LG / 65 / 2609.22360

Contrastive Siamese Representation Learning for Predictive Maintenance of Electrical Submersible Pumps

基于对比孪生网络表示学习的电潜泵预测性维护方法
Damarla, Seshu K., Zhu, Xiuli
Abstract
Electrical submersible pumps (ESPs) are essential in offshore oil production, where unexpected failures can result in significant operational and financial losses. Accurate predictive maintenance for ESP systems remains challenging due to nonlinear operating conditions, class imbalance, and variability among pump units. To address these issues, this study presents a fault diagnosis framework that incorporates class imbalance awareness by employing Siamese contrastive representation learning and prior-corrected k-nearest neighbor (KNN) classification. The method first extracts discriminative features relevant to fault detection from vibration-domain indicators and engineered harmonic relationships. A Siamese neural network is trained with class-balanced contrastive pairs to construct an embedding space that clusters samples of the same fault type and separates different fault classes. To further mitigate class imbalance during classification, a prior-corrected distance-weighted KNN is applied. The framework is validated using a Leave-One-ESP-Out (LOEO) strategy to evaluate generalization to previously unseen ESP units. Experimental results indicate that the proposed framework delivers robust and consistent fault classification performance under realistic industrial conditions, supporting its potential for reliable predictive maintenance and intelligent ESP system monitoring.
Chinese Translation
电潜泵(Electrical Submersible Pumps, ESPs)在海上石油生产中至关重要,其意外故障可能导致重大的运营和财务损失。由于运行工况的非线性、类别不平衡以及泵机组之间的差异性,对电潜泵系统进行准确的预测性维护仍然具有挑战性。为解决这些问题,本研究提出了一种故障诊断框架,通过采用孪生网络(Siamese)对比表示学习和先验校正的k近邻(KNN)分类来实现类别不平衡感知。该方法首先从振动域指标和构建的谐波关系中提取与故障检测相关的判别性特征。通过类别平衡的对比样本对训练孪生神经网络,构建一个嵌入空间,使同一故障类型的样本聚集在一起,并将不同故障类别分离开来。为进一步缓解分类过程中的类别不平衡问题,该方法采用了先验校正的距离加权KNN分类器。该框架采用留一电潜泵(Leave-One-ESP-Out, LOEO)策略进行验证,以评估其对未见过的电潜泵机组的泛化能力。实验结果表明,所提出的框架在真实的工业条件下能够提供稳健且一致的故障分类性能,支持其在可靠的预测性维护和电潜泵系统智能监控方面的应用潜力。
cs.LG / 66 / 2609.22361

Common Cause, Not Cross-Attention: Blocking Visual Shortcuts in Audio-Video Generation

共同原因,而非交叉注意力:阻断音频-视频生成中的视觉捷径
Xu, Jian, Zeng, Delu, Paisley, John, Zhao, Qibin
Abstract
Joint audio--video generators are trained on data in which what an event looks like and what it sounds like are strongly, often spuriously, correlated: a particular material, texture, or object appearance co-occurs with a particular sound. This paper is a controlled causal study of the resulting failure mode. Building an AV structural causal model in which the audio is, by construction, independent of the video's nuisance appearance, we show that models which let audio read video directly-through cross-attention or a shared latent-learn a visual shortcut: they predict sound from appearance rather than from the causal event, and collapse when the appearance-event correlation is broken at test time, literally synthesizing the wrong event's sound. Crucially, the popular remedy of routing both modalities through a shared common-cause latent does not fix this: a bottleneck, an unsupervised shared/private factorization, and a faithful shared-prior model all grab the appearance proxy and fail like the direct model. Blocking the shortcut instead requires an intervention on the nuisance. Under the stated SCM and intervention assumptions we prove that counterfactual invariance is necessary and sufficient to identify the causal predictor, and we verify the mechanism across a feature-vector SCM, procedural pixel video, real images with spectrogram audio and a pretrained backbone, moving real digits, and a conditional generator. On a \emph{real, pretrained} video-to-audio generator, an input-intervention test shows the model is far from invariant to sound-irrelevant edits (recolouring or graying a video substantially changes the sound it generates).
Chinese Translation
音频-视频联合生成器的训练数据中,事件的外观与其声音之间存在强相关,且往往是虚假相关:特定的材质、纹理或物体外观总是与特定的声音共现。本文对这些数据所导致的失效模式进行了受控的因果研究。我们构建了一个音视频结构因果模型(AV SCM),其中音频在构造上独立于视频的干扰性外观(nuisance appearance),并证明:允许音频直接读取视频的模型——无论是通过交叉注意力(cross-attention)还是共享隐变量——都会学到一种视觉捷径:它们根据外观而非因果事件来预测声音,当测试时外观与事件的关联被打破时便会崩溃,甚至直接合成出错误事件的声音。至关重要的是,流行的补救方案——将两种模态都经由共享的共同原因隐变量路由——并不能解决此问题:瓶颈结构、无监督的共享/私有因子分解以及忠实的共享先验模型,都会捕捉外观代理变量,并与直接读取模型一样失败。阻断捷径需要对干扰因素进行干预。在所陈述的 SCM 与干预假设下,我们证明反事实不变性(counterfactual invariance)是识别因果预测器的充要条件,并在多个场景中验证了该机制:特征向量 SCM、程序化生成的像素视频、使用真实图像与频谱图音频及预训练骨干网络、移动的真实数字以及条件生成器。在一个真实且经过预训练的视频生成音频模型上,输入干预测试表明该模型对与声音无关的编辑远非不变(改变视频的颜色或将其灰度化会显著改变其生成的声音)。
cs.LG / 67 / 2609.22415

Complex-valued Phase-Coherent Transformers

复数值相位相干Transformer
Hioki, Leona
Abstract
Complex-valued Transformers have inherited softmax attention over the raw complex inner product. Outside natively complex domains this standard form stays near chance, and no complex attention had been shown to correct it. We show that the match must be a scaled cosine score: L2-normalise queries and keys, so the score reads their cosine similarity and ignores their magnitudes, and hold that score at order-one scale. With this the same models train on four diagnostic tasks under two different gates; without the normalisation they stay at chance on ListOps and Needle under both gates and fall far below on the other two, and a normalised score placed at too small a scale fails as well. The resulting family of phase-coherent Transformers (\PCT) matches or exceeds the strongest real-valued baseline across long-range memory, positional retrieval, hierarchical reasoning, frequency-domain classification and physical complex signals; it shows no degradation up to depth 20; and its loss decreases log-linearly over a 61-fold range of parameters. A member of the family, complex screening combined with a phase-coherent recurrence, is the first genuinely complex-valued neural network to solve Path-X, with 91.6% of its trainable parameters complex-valued against 38.2% for S4. We record these as signs of generalisation not previously seen in complex-valued neural networks.
Chinese Translation
复数值Transformer一直沿用在原始复数内积上进行的softmax注意力。在非天然复数的领域,这种标准形式的表现接近随机水平,而此前尚无复数注意力机制被证明能够纠正这一问题。我们证明匹配得分必须是缩放后的余弦得分:对查询(query)和键(key)进行L2归一化,使得分反映二者的余弦相似度并忽略其模长,同时将该得分保持在接近于一的量级。据此,同样的模型可以在两种不同门控机制下于四个诊断任务上完成训练;若不进行归一化,则在两种门控下模型在ListOps和Needle任务上均停留在随机水平,在其他两个任务上也远低于正常表现;而归一化后若得分量级过小,模型同样会失败。由此得到的相位相干Transformer系列(PCT)在长程记忆、位置检索、层次化推理、频域分类以及物理复数信号等任务上,达到或超越了最强的实数值基线;该模型在深度达到20时仍无性能退化;且在参数量跨越61倍的范围内,其损失呈对数线性下降。该系列中的一个成员——复数筛选与相位相干递归相结合——是首个真正解决Path-X任务的复数值神经网络,其可训练参数中复数值参数占比达91.6%,而S4仅为38.2%。我们将这些结果视为复数值神经网络此前未曾展现过的泛化能力的标志。
cs.LG / 68 / 2609.22441

Connected Content Retriever: Dense Graph Edge Features Powering Pre-Ranking at LinkedIn

关联内容检索器:在LinkedIn预排序中发挥作用的稠密图边特征
Gupta, Akhilesh, Ramanujam, Sudarshan Srinivasa, Mehta, Chirag Bhanuprasad, Beena, Reshma Asharaf, Das, Dhritiman, Tiwana, Birjodh Singh, Patel, Bhargavkumar Kanubhai, Lee, Mack, Tang, Renyi
Abstract
In large-scale recommendation systems like the LinkedIn Feed, content generated by a member's network (connections and follows) makes up over 70% of impressions and engagement. It is therefore essential that the pre-ranking layer forwards the best possible few hundred candidates to the ranking layer. LinkedIn's professional knowledge graph carries engagement signals across both the first degree network (connections and follows) and the second-degree network: posts that a 1st-degree connection reacted to, commented on or reshared but did not author (a.k.a. stranger viral). Due to this fan out, the resulting candidate index exceeds one billion; selection of activities from the viewer's network narrows it down to roughly tens of thousands of activities that must be scored within a 120 ms p99 latency budget. We present Connected Content Retriever (CC Retriever), a pre-ranking system that scores these candidates with a full deep ranking model on GPUs at low latency. At its core is a sorted-search GPU primitive that joins dense graph affinity features (viewer to author) with document level features stored on the GPU at runtime in 5-10 ms. The shift to GPU served scoring enabled a 50x scale up of the ranking model's parameters and delivered a +2.5% lift in content time spent on the LinkedIn Feed in online experiments, significantly higher than the typical gains observed in LinkedIn Feed experiments. In this work, we describe the feature set we leverage from LinkedIn's economic graph and the model architecture used for scoring, with a particular emphasis on the online system that scales the stack.
Chinese Translation
在LinkedIn Feed这样的大规模推荐系统中,由会员网络(人脉和关注)生成的内容占据了超过70%的曝光和互动。因此,预排序层必须将最优的数百个候选内容传递给排序层,这一点至关重要。LinkedIn的职业知识图谱承载着跨越一度人脉网络(人脉和关注)和二度人脉网络的互动信号:即一度人脉对其做出反应、评论或转发但并非其创作的内容(也称“陌生人病毒式传播”内容)。由于这种扇出效应,生成的候选索引规模超过十亿;从浏览者网络中筛选活动后,候选规模缩小到数万条活动,这些活动必须在120毫秒的p99延迟预算内完成打分。我们提出了关联内容检索器(Connected Content Retriever,CC Retriever),这是一个预排序系统,能够在GPU上以低延迟使用完整的深度排序模型对这些候选进行打分。其核心是一个排序搜索GPU原语,可在5-10毫秒内将稠密图亲和度特征(浏览者与作者之间)与运行时存储在GPU上的文档级特征进行连接。转向GPU服务化打分使排序模型的参数规模扩大了50倍,并在在线实验中使LinkedIn Feed的内容使用时长提升了2.5%,显著高于LinkedIn Feed实验中通常观察到的收益。在本工作中,我们描述了从LinkedIn经济图谱中利用的特征集合以及用于打分的模型架构,并特别重点介绍了支撑整个技术栈扩展的在线系统。
cs.LG / 69 / 2609.22471

Efficient Mixture-of-Experts with Speculative Decoding via Expert Coactivation

基于专家协同激活与投机解码的高效混合专家模型
Nishu, Kumari, Kim, Han-Byul, Chilkunda, Santosh, Horton, Maxwell, Kundu, Arnav, Samragh, Mohammad, Hannah, Lauren, Sekhavat, Mohammad, Bhendawade, Nikhil, Ciosici, Manuel, Mirzadeh, Iman, Vahid, Keivan Alizadeh, Harrison, David, Belousova, Irina, Farajtabar, Mehrdad, Cho, Minsik
Abstract
Mixture-of-Experts (MoE) models are increasingly deployed alongside Speculative Decoding (SD) to accelerate inference, but combining the two is challenging. SD improves the inference speed of dense models by verifying groups of tokens in parallel. However, the inference speedup for SD with MoEs depends heavily on the number of tokens being verified. Using more verification tokens results in more experts being transferred from DRAM to the Neural Processing Unit (NPU), which increases the memory transfer cost. This negatively impacts model runtime, as memory transfer is typically the bottleneck in inference. In this work, we investigate the impact of MoE router design during training on the speed of MoEs with SD. We find that routers with high degrees of expert coactivation result in much faster runtimes, mitigating the impact of using more verification tokens. Motivated by this observation, we assess the impact of various router design choices on expert coactivation and runtime using billion-parameter transformer models. We find that combining a global load-balancing loss, shared experts, a consistency loss, and an autoregressive expert selection mechanism during training results in significantly stronger expert coactivation. This increased coactivation translates into higher overall runtime throughput: our exploration yields a model that improves throughput by 21% over MoE baselines, while maintaining on-par accuracy with the baseline MoE.
Chinese Translation
混合专家(Mixture-of-Experts, MoE)模型越来越多地与投机解码(Speculative Decoding, SD)结合部署以加速推理,但将二者结合颇具挑战性。SD通过并行验证一组令牌来提升稠密模型的推理速度。然而,SD在MoE模型上的推理加速在很大程度上取决于被验证的令牌数量。使用更多的验证令牌会导致更多专家需要从DRAM传输到神经处理单元(NPU),从而增加内存传输开销。由于内存传输通常是推理的瓶颈,这会对模型运行时间产生负面影响。在本工作中,我们研究了训练时MoE路由器设计对结合SD的MoE推理速度的影响。我们发现,具有高度专家协同激活的路由器能够显著缩短运行时间,缓解使用更多验证令牌带来的影响。基于这一观察,我们使用十亿参数规模的Transformer模型评估了各种路由器设计选择对专家协同激活和运行时间的影响。我们发现,在训练中结合全局负载均衡损失、共享专家、一致性损失以及自回归的专家选择机制,能够显著增强专家协同激活。这种增强的协同激活转化为更高的整体运行吞吐量:我们的探索所得模型相比MoE基线将吞吐量提升了21%,同时保持了与基线MoE相当的准确率。
cs.LG / 70 / 2609.22487

COREM: Cosine-Relation Momentum Reshaping with Stateful Writeback

COREM:具有状态回写的余弦关系动量重塑方法
Wang, Yan, Wang, Xiaochuan, Sun, Yuxiang
Abstract
Matrix-valued optimizer states may contain relational structure that is not captured by treating their entries independently. We study whether relations within matrix-valued optimizer states can be exploited to improve optimization. To this end, we introduce a unit-relation-transform abstraction and instantiate it as COREM, a Cosine-Relation Momentum Reshaping method with stateful writeback. COREM partitions the momentum state into update units, computes cosine relations among them, and uses these relations to reshape the momentum before writing the transformed state back to the optimizer. This stateful mechanism allows the reshaped momentum to affect not only the current update but also future optimization dynamics. We evaluate COREM on CIFAR-10 with an MLP and on enwik8 with a Transformer. Compared with Muon, COREM shows lower early-stage step efficiency but stronger improvement in the mid-to-late stages of training, achieving better final validation performance on CIFAR-10 and comparable final performance on enwik8. Spectral diagnostics on enwik8 show that COREM consistently increases entropy effective rank and reduces the concentration of singular energy in dominant modes, while preserving an anisotropic spectrum. For square matrix updates, COREM requires approximately 13.3% of the transformation FLOPs of Muon with five Newton-Schulz iterations.
Chinese Translation
矩阵形式的优化器状态中可能包含某种关系结构,而将矩阵元素独立处理的方法无法捕捉这种结构。我们研究了能否利用矩阵优化器状态内部的关系来改进优化过程。为此,我们提出了一个单元关系变换(unit-relation-transform)抽象,并将其具体化为 COREM——一种带有状态回写(stateful writeback)的余弦关系动量重塑方法。COREM 将动量状态划分为更新单元,计算它们之间的余弦关系,并利用这些关系在将变换后的状态回写到优化器之前对动量进行重塑。这种带状态的机制使得重塑后的动量不仅影响当前更新,还会影响未来的优化动态。我们在使用 MLP 的 CIFAR-10 任务和使用 Transformer 的 enwik8 任务上评估了 COREM。与 Muon 相比,COREM 在训练早期的步效率较低,但在训练中后期表现出更强的改进,在 CIFAR-10 上取得了更好的最终验证性能,在 enwik8 上取得了相当的最终性能。在 enwik8 上的谱诊断结果表明,COREM 持续提高了熵有效秩(entropy effective rank),并降低了奇异能量在主导模式上的集中程度,同时保持了各向异性的谱结构。对于方阵更新,COREM 所需的变换浮点运算量(FLOPs)约为采用五次 Newton-Schulz 迭代的 Muon 的 13.3%。
cs.LG / 71 / 2609.22508

EmbeddGAN: A Novel GAN Framework Using an Embedding Network and Gini Distance Correlation

EmbeddGAN:一种基于嵌入网络与基尼距离相关性的新型生成对抗网络框架
Caldwell, MaTais, Chen, Yixin, Dang, Xin, Walter, Charles
Abstract
Generative Adversarial Networks (GANs) have demonstrated strong performance in generating high-quality synthetic data. However, they are limited by no formal guarantees regarding convergence and the effectiveness of the learning process. In practice, this leads to training instability, mode collapse, and sensitivity to hyperparameters. To address this, we propose EmbeddGAN, a novel adversarial training framework based on a dependence-based objective. Instead of relying on a discriminator that classifies samples as real or fake, EmbeddGAN introduces an embedding network that learns a representation in which statistical dependence between samples and their real/fake labels is maximized, while the generator is trained to minimize this dependence. This objective is implemented using the Gini distance correlation (gCor), which equals zero if and only if the embeddings are statistically independent of the real/fake label. Minimizing this objective therefore encourages real and generated samples to become statistically indistinguishable in the learned embedding space. The embedding network projects both real and generated data into a shared low-dimensional space, where distributional discrepancies can be measured directly through pairwise distances. We adopt a minimax training strategy: the embedding network maximizes the Gini distance correlation (maximizing dependence), while the generator minimizes it (minimizing dependence). Experiments on the MNIST, CIFAR-10, and CelebA datasets demonstrate that EmbeddGAN achieves competitive performance relative to established baselines while exhibiting notably stable training dynamics on the evaluated datasets.
Chinese Translation
生成对抗网络(Generative Adversarial Networks, GANs)在生成高质量合成数据方面表现出色。然而,其局限性在于缺乏关于收敛性和学习过程有效性的形式化保证。在实际应用中,这会导致训练不稳定、模式崩溃(mode collapse)以及对超参数敏感等问题。为了解决这些问题,我们提出了EmbeddGAN,一种基于依赖性目标的新型对抗训练框架。EmbeddGAN不再依赖将样本分类为真或假的判别器,而是引入一个嵌入网络(embedding network),该网络学习一种表示,使得样本与其真假标签之间的统计依赖性最大化,同时训练生成器最小化这种依赖性。该目标通过基尼距离相关性(Gini distance correlation, gCor)实现——当且仅当嵌入与真假标签统计独立时,gCor等于零。因此,最小化该目标可促使真实样本与生成样本在学到的嵌入空间中变得统计上不可区分。嵌入网络将真实数据和生成数据投影到一个共享的低维空间,在该空间中可以通过成对距离直接度量分布差异。我们采用极小极大(minimax)训练策略:嵌入网络最大化基尼距离相关性(最大化依赖性),而生成器最小化它(最小化依赖性)。在MNIST、CIFAR-10和CelebA数据集上的实验表明,EmbeddGAN相对于现有基线方法取得了具有竞争力的性能,同时在所评估的数据集上展现出显著更稳定的训练动态。
cs.LG / 72 / 2609.22554

The Ups and Downs of Backprop Weights

反向传播权重的起伏
Chindemi, Giuseppe, Grewe, Benjamin F.
Abstract
Backpropagation (BP) has driven the remarkable success of modern deep learning by enabling large hierarchical networks to learn complex functions end-to-end. Yet it does not by itself determine how parameters should be organized so that functional components can be reused and adapted selectively. For example, object recognition and motion prediction may depend on overlapping parameter sets, making them difficult to isolate or modify independently. We call this condition weight entanglement. Modern architectures dynamically select which parts of a network process each sample: nonlinearities gate units, attention selects interactions, and Mixture-of-Experts architectures route inputs to modules. Yet such selection does not ensure that the same functional component remains linked to an identifiable parameter set across samples. We propose weight operators: parameterized modules that implement reusable functional components and can be composed at inference to form the function required by each sample. Learning proceeds in two stages: the model first infers the required operator composition, then updates only the selected operators' parameter sets. Vector Networks (VNs) provide one implementation. They couple operator selection to local error-driven updates within each layer and show that learned operators can be reused in combinations absent from training while updates remain restricted to the selected parameter sets. This provides a basis for testing functional parameter identifiability: whether an operator remains linked to the same functional component during learning. We argue that functional parameter identifiability may provide an organizing principle for models that systematically reuse and recombine learned functions while adapting only the components that need to change.
Chinese Translation
反向传播(BP)通过使大型分层网络能够端到端地学习复杂函数,推动了现代深度学习的显著成功。然而,它本身并不能决定应如何组织参数,以使功能组件能够被复用和有选择地调整。例如,物体识别与运动预测可能依赖于相互重叠的参数集,这使得它们难以被独立地隔离或修改。我们将这种状况称为权重纠缠(weight entanglement)。现代架构会动态选择网络的哪些部分来处理每个样本:非线性激活对单元进行门控,注意力机制选择交互,混合专家(Mixture-of-Experts)架构将输入路由到各个模块。然而,这种选择并不能确保同一功能组件在不同样本之间始终与一个可识别的参数集保持关联。我们提出权重算子(weight operators):这是实现可复用功能组件的参数化模块,可以在推理时组合起来,形成每个样本所需的函数。学习分两个阶段进行:模型首先推断所需的算子组合,然后仅更新被选中算子的参数集。向量网络(Vector Networks, VNs)提供了一种具体实现。它们将算子选择与每一层内部的局部误差驱动更新相耦合,并表明学习到的算子可以在训练中未出现过的组合中被复用,同时更新仍然仅限于被选中的参数集。这为检验功能参数可识别性(functional parameter identifiability)提供了基础,即一个算子在训练过程中是否始终与同一功能组件保持关联。我们认为,功能参数可识别性可能为那些系统性复用和重组已学习函数、且仅调整需要改变的组件的模型提供一个组织原则。
cs.LG / 73 / 2609.22583

Benchmarking Hybrid Deep Learning Architectures for Predictive Maintenance in Industry 4.0

面向工业4.0预测性维护的混合深度学习架构基准测试
Zhengyang, Gu, Hernandez, Joseph E., Cook, Thomas, Burtenshaw, John, Scott, Sean, Couch, Chris
Abstract
Predictive maintenance in Industry 4.0 refers to using data from sensors, machines, and production systems to estimate when equipment is likely to fail, so maintenance can be planned before a breakdown occurs [1]. However, a model that predicts maintenance may work perfectly in the lab but fail unexpectedly when applied to real factory data [2]. To solve this "reliability" gap, we evaluated six deep learning architectures across more than 700 experimental runs. We focused on the two dominant approaches in the field: Recurrent Neural Networks (RNNs), which process data step-by-step, like reading a sentence [3], and Transformers, a recent dominant approach, which look at the entire sequence at once to spot important connections [4]. We examined whether Transformers still outperform recurrent neural networks (RNNs) when the data includes noise [5]. We found that while Transformers excelled at tracking stable, slow-moving processes, they tend to overreact to chaotic data, mistakenly taking sensor noise for meaningful signals [6]. We also found that the hybrid method that combines a Long Short-Term Memory (LSTM) layer with a Transformer layer is more resilient to noisy data from factory shops [7]. Functioning as a noise filter, the LSTM smooths out data volatility, allowing the Transformer to focus on the bigger picture without being distracted [8]. The hybrid model did not just improve accuracy; it proved to be significantly more consistent than complex models, delivering reliable predictions regardless of how chaotic the underlying system became.
Chinese Translation
工业4.0中的预测性维护是指利用来自传感器、机器和生产系统的数据来估计设备可能出现故障的时间,从而在故障发生前规划维护工作[1]。然而,一个预测性维护模型可能在实验室中表现完美,但在应用于真实工厂数据时却出现意外失效[2]。为解决这一"可靠性"差距,我们在700多次实验运行中评估了六种深度学习架构。我们聚焦于该领域两种主流方法:循环神经网络(RNN),其像阅读句子一样逐步处理数据[3];以及Transformer,一种近期占主导地位的方法,它能够同时查看整个序列以发现重要关联[4]。我们研究了当数据包含噪声时,Transformer是否仍优于循环神经网络(RNN)[5]。我们发现,尽管Transformer在追踪稳定、缓慢变化的过程方面表现出色,但它容易对混乱数据产生过度反应,误将传感器噪声当作有意义的信号[6]。我们还发现,将长短期记忆(LSTM)层与Transformer层相结合的混合方法对来自工厂车间的噪声数据具有更强的鲁棒性[7]。LSTM充当噪声过滤器,平滑数据的波动性,使Transformer能够专注于整体趋势而不受干扰[8]。该混合模型不仅提升了准确率,还被证明比复杂模型具有显著更高的一致性,无论底层系统变得多么混乱,都能提供可靠的预测。
cs.LG / 74 / 2609.22584

Augmenting PID Control with Deep Reinforcement Learning: A Hybrid Approach to the Industrial Benchmark

用深度强化学习增强PID控制:面向工业基准的混合方法
Zhengyang, Gu, Hernandez, Joseph E., Burtenshaw, John, Scott, Sean, Cook, Thomas, Couch, Chris
Abstract
As industrial processes grow in complexity, traditional Proportional-Integral-Derivative (PID) controllers are often insufficient for handling their non-linear, multi-input dynamics. We propose using advanced Deep Reinforcement Learning (DRL) to prove its advantages in these complex environments. To do this, we rely on the Industrial Benchmark (IB). The IB is a realistic simulation that tests DRL algorithms against the key challenges of industrial applications: high-dimensional state spaces, delayed effects, and conflicting multi-criterial objectives. This testbed highlights DRL's core trade-off: while its final policies can often be unstable, its unique strength is the ability to autonomously discover optimal, non-obvious policies in multi-dimensional spaces where simple controllers fail. In this paper, we propose a novel hybrid PID-RL controller that leverages DRL's discovery capability while ensuring Reliability. After developing a multi-objective reward function to make DRL viable, we use a twin-delayed deep deterministic (TD3) agent as a discovery tool to find the optimal, non-obvious settings for the IB's 'Gain' and 'Shift' parameters. By feeding these discovered parameters to a simple, tuned PID controller, our hybrid model successfully combines all three characteristics: it achieves the optimal Performance and Efficiency of the best DRL agent with the Reliability of a classical controller. This work demonstrates a practical methodology for using DRL to augment, rather than replace, trusted industrial control systems.
Chinese Translation
随着工业过程复杂性的不断提高,传统的比例-积分-微分(PID)控制器往往难以应对其非线性、多输入的动态特性。我们提出利用先进的深度强化学习(DRL)来证明其在这些复杂环境中的优势。为此,我们依托工业基准(Industrial Benchmark, IB)开展研究。IB是一个逼真的仿真平台,用于检验DRL算法应对工业应用关键挑战的能力,包括高维状态空间、延迟效应以及相互冲突的多目标。该测试平台凸显了DRL的核心权衡:尽管其最终策略往往不稳定,但其独特优势在于能够在简单控制器失效的多维空间中自主发现最优且非直观的策略。本文提出了一种新颖的PID-RL混合控制器,该控制器在利用DRL发现能力的同时确保可靠性。在设计了使DRL可行的多目标奖励函数之后,我们使用双延迟深度确定性策略梯度(TD3)智能体作为发现工具,为IB的'Gain'和'Shift'参数寻找最优且非直观的设置。通过将这些发现的参数输入给一个简单且经过调优的PID控制器,我们的混合模型成功结合了全部三种特性:既实现了最优DRL智能体的最佳性能与效率,又具备经典控制器的可靠性。这项工作展示了一种实用的方法论,即利用DRL来增强而非取代可信赖的工业控制系统。
cs.LG / 75 / 2609.22585

TWIG: A Time-Causal Wavelet Operator for Autoregressive Forecasting on Irregular Graphs

TWIG:一种用于不规则图自回归预测的时间因果小波算子
Venkatasubramanian, Subashree, Barajas-Solano, David A., Liu, Chuyang, Tartakovsky, Daniel M., Dwivedi, Dipankar
Abstract
We introduce TWIG (Time-Causal Wavelet Operator for Irregular Graphs), a graph-native neural operator for autoregressive surrogate modeling on static irregular graphs. TWIG transforms each node history into causal multiscale temporal features that separate recent variation from progressively slower memory components, then propagates these features through graph-wavelet operator blocks with gated pointwise channel mixing. The architecture is causal by construction and designed for closed-loop forecasting, where predictions are recursively reused as future inputs. We evaluate TWIG on three irregular-domain forecasting problems spanning regional diffusion, three-dimensional subsurface hydrology, and aerodynamic flow, with graphs ranging from 400 to 5,233 nodes and model capacities from approximately 70k to 10M parameters. TWIG achieves the lowest aggregate rollout errors on the subsurface-hydrology and regional-diffusion benchmarks and ranks second on the 10M-parameter aerodynamic-flow benchmark, behind the GPS Transformer. Across all three settings, TWIG consistently outperforms the corresponding non-time-causal Graph WNO baseline. These results demonstrate that TWIG provides an effective and scalable approach to stable autoregressive forecasting of dynamical fields on irregular graphs.
Chinese Translation
我们提出了TWIG(Time-Causal Wavelet Operator for Irregular Graphs,面向不规则图的时间因果小波算子),这是一种图原生的神经算子,用于在静态不规则图上进行自回归代理建模。TWIG将每个节点的历史序列转换为因果多尺度时间特征,将近期变化与逐渐变慢的记忆成分分离开来,然后通过带有门控逐点通道混合的图小波算子模块传播这些特征。该架构在构造上具有因果性,专为闭环预测而设计,其中预测结果被递归地用作未来的输入。我们在三个不规则域预测问题上评估了TWIG,涵盖区域扩散、三维地下水文和空气动力学流动,图规模从400到5,233个节点不等,模型容量约为70k到10M参数。TWIG在地下水文和区域扩散基准上取得了最低的总体滚动预测误差,并在10M参数的空气动力学流动基准上排名第二,仅次于GPS Transformer。在所有三种设置中,TWIG均始终优于对应的非时间因果的Graph WNO基线。这些结果表明,TWIG为不规则图上动力场的稳定自回归预测提供了一种有效且可扩展的方法。
cs.LG / 76 / 2609.22593

User-Level Handover Decision Making Based on Machine Learning Approaches

基于机器学习方法的用户级切换决策
Lima, João, Medeiros, Alvaro, Aguiar, Eduardo, Junior, Vicente Angelo de Sousa, Guerra, Tarciana
Abstract
This letter covers a broad comparison of methods for classification and regression applications for a user-level handover decision making in scenarios with adverse propagation conditions involving buildings, coverage holes, and shadowing effects. The simulation campaigns are based on network simulator ns-3. The comparison encompasses classical machine learning approaches, such as KNN, SVM, and neural networks, but also state-of-the-art fuzzy logic systems and latter boosting machines. The results indicate that SVM and MLP are the most suitable for the classification of the best handover target, although fuzzy system SOFL can perform similarly with lower processing time. Additionally, for the download time estimation, LightGBM provides the smallest error with short processing time, even in hard propagation scenarios.
Chinese Translation
本文对用户级切换决策中的分类与回归应用方法进行了广泛的比较,所涉及的场景包含恶劣的传播条件,如建筑物遮挡、覆盖空洞和阴影效应。仿真实验基于网络仿真器ns-3开展。比较涵盖了经典机器学习方法,如KNN、SVM和神经网络,同时也包括最先进的模糊逻辑系统以及后来的梯度提升机。结果表明,SVM和MLP最适合用于最佳切换目标的分类,而模糊系统SOFL也能以更低的处理时间实现相近的性能。此外,在下载时间估计方面,LightGBM在处理时间较短的情况下误差最小,即使在复杂传播场景中亦是如此。
cs.LG / 77 / 2609.22614

Concurrency-Aware Process Model Forecasting with Causal Nets

基于因果网的并发感知过程模型预测
Yu, Yongbo, Peeperkorn, Jari, De Smedt, Johannes, De Weerdt, Jochen
Abstract
Process model forecasting (PMF) aims to predict the process model that will characterize a future period, thereby providing a process-level view of how behavior is expected to evolve. Existing PMF methods, however, forecast directly-follows graphs, which cannot explicitly represent concurrency. We extend PMF to causal nets by forecasting time series of relation and binding counts and using these forecasts to reconstruct future process models with AND/XOR semantics. To evaluate the resulting models, we introduce a protocol that accounts for partial traces and constructs the workflow nets required for conformance checking. Experiments on four event logs show that the forecasted models achieve conformance levels close to those of models re-mined from observations in the corresponding future windows. They also outperform static discovery baselines, which retain high precision on the structurally stable log but exhibit substantial precision losses on the other three logs. Filtering infrequent bindings improves most conformance metrics, although it also removes much of the concurrent behavior captured by the models.
Chinese Translation
过程模型预测(Process Model Forecasting, PMF)旨在预测能够刻画未来时期特征的过程模型,从而提供行为预期演化的过程级视角。然而,现有的PMF方法预测的是直接后继图(directly-follows graphs),无法显式表示并发关系。我们将PMF扩展到因果网(Causal Nets),通过预测关系计数和绑定计数的时间序列,并利用这些预测结果重构具有AND/XOR语义的未来过程模型。为了评估所得到的模型,我们提出了一种考虑部分轨迹并构建一致性检查所需工作流网(workflow nets)的评估协议。在四个事件日志上的实验表明,预测模型的一致性水平接近于从对应未来窗口的观测数据中重新挖掘所得的模型。预测模型还优于静态发现基线方法:后者在结构稳定的日志上保持较高的精确度,但在其他三个日志上出现了显著的精确度损失。过滤低频绑定可以提升大多数一致性指标,但同时也移除了模型所捕获的相当一部分并发行为。
cs.LG / 78 / 2609.22632

Classification with Abstention Under Class-Conditional Error Constraints

类条件错误约束下的带弃判分类
Kalan, Mohammadreza M., Deng, Yuyang, Hamidi, Sanaz
Abstract
We study binary classification with abstention under separate class-conditional error constraints, with the objective of minimizing abstention while keeping both errors below prescribed thresholds. We characterize the distribution-free minimax rate of excess abstention risk, up to logarithmic factors, in terms of the complexity of the hypothesis class and the sample size. To make the framework amenable to computation with models such as neural networks, we introduce surrogate-loss formulations and derive finite-sample guarantees for excess surrogate ambiguity risk. We formulate the resulting learning task as a constrained optimization problem and characterize its computational complexity in the convex setting. Finally, we evaluate our approach on various datasets and compare its performance with a competing method for this problem.
Chinese Translation
我们研究了在分别施加于各类别的错误约束下的二分类带弃判问题,其目标是在保证两类错误均低于预设阈值的前提下最小化弃判率。我们以假设类的复杂度和样本量为函数,在对数因子范围内刻画了无分布假设下超额弃判风险的极小极大速率。为使该框架能够与神经网络等模型结合进行计算,我们引入了代理损失形式,并推导出超额代理模糊风险的有限样本保证。我们将由此产生的学习任务表述为一个约束优化问题,并在凸设置下刻画了其计算复杂度。最后,我们在多个数据集上评估了所提方法,并将其性能与该问题的一种竞争方法进行了比较。
cs.LG / 79 / 2609.22643

Monotone-Constrained Diffusion Models for Long-Horizon Production Forecasting

面向长时程生产预测的单调约束扩散模型
Abraha, Temesgen Mikael, Lucet, Yves
Abstract
Forecasting a long horizon from only the first observations of a sequence is ill-posed: many trajectories are consistent with the same short history. We study this problem in oil and gas production forecasting, where forecasts made after roughly the first fifth of a well's producing life drive development and abandonment decisions, and where a usable forecast must describe a monotone decline. We present Physics-SIMS-TS, a conditional diffusion forecaster that combines negative guidance against synthetic artifacts, decline-curve constraints and an isotonic projection applied during sampling, spatial training augmentation, and an ensembled stochastic sampler yielding a full predictive distribution. Across three jurisdictions and more than 35,000 wells, under a shared-space, validation-frozen protocol, Physics-SIMS-TS is the most accurate diffusion forecaster in the comparison and is competitive with, but not superior to, ensembled transformer forecasters. Its forecasts are monotone by construction at a cost of at most 0.5% in mean squared error, and its trajectory ensemble yields calibrated intervals after one dispersion factor is fitted per jurisdiction. On six standard benchmarks a reversible-instance-normalization variant of the backbone is the leading diffusion baseline. We also quantify four protocol choices on which the measured ranking depends. Code and evaluation artifacts are released.
Chinese Translation
仅根据序列的初始观测进行长时程预测是一个不适定问题:许多轨迹都与同一短暂的历史数据相一致。我们在油气生产预测中研究该问题,此场景下,在油井生产寿命大约前五分之一阶段所作的预测驱动着开发与废弃决策,且可用的预测必须描述单调递减的产量。我们提出 Physics-SIMS-TS,一种条件扩散预测器,它结合了针对合成伪影的负向引导、递减曲线约束与采样过程中施加的等张投影、空间训练增强,以及产生完整预测分布的集成随机采样器。在覆盖三个司法辖区、超过35,000口井的数据上,采用共享空间、验证集冻结的评估协议,Physics-SIMS-TS 是对比中最精确的扩散预测器,其性能与集成 Transformer 预测器相当但并未超越。其预测在构造上即为单调,均方误差代价至多0.5%;且在为每个辖区拟合一个离散度因子后,其轨迹集成可产生经校准的置信区间。在六个标准基准上,骨干网络的可逆实例归一化变体是最优的扩散基线。我们还量化了所测排名所依赖的四种协议选择。代码与评估工件已开源发布。
cs.LG / 80 / 2609.22690

Multi-Armed Bernoulli Bandits via Minimax Single-Arm Stopping

基于极小化极大单臂停时的多臂伯努利老虎机
Liu, Huikang, Wang, Zhengchao, Kuhn, Daniel, Wiesemann, Wolfram
Abstract
We develop an index policy for finite-horizon Bernoulli multi-armed bandits from minimax solutions to single-arm bandit (SAB) problems. Each SAB problem involves choosing between an unknown Bernoulli arm and a known reward. We show that minimizing worst-case regret of SAB problems over all non-anticipative policies admits an exact semi-infinite linear programming formulation. The resulting stopping policies offer a natural way to compare arms: the higher the known reward against which a policy continues sampling, the more promising the unknown arm. We turn this intuition into indices based on cumulative continuation probabilities, with a monotone adjustment and a reward-shortfall cap. By relating index errors to the regret of single-arm stopping policies, we establish a distribution-free regret bound of $4.45\sqrt{KT}+10.75K$ for $K$ arms and horizon $T$. This bound matches the minimax-optimal regret order established in the literature. The guarantee extends to rewards supported on $[0,1]$ through Bernoulli randomization. We also provide a finite-grid implementation with quantified approximation loss. In numerical experiments, the SAB-based index policy achieves lower worst-case regret than every tested benchmark policy across all evaluated numbers of arms and horizons, while closely matching the grid-based MAB minimax policy in the two-arm setting.
Chinese Translation
我们基于单臂老虎机(Single-Arm Bandit, SAB)问题的极小化极大解,为有限时域的伯努利多臂老虎机提出了一种指标(index)策略。每个SAB问题涉及在一个未知的伯努利臂与一个已知收益之间进行选择。我们证明,在所有非预知(non-anticipative)策略上最小化SAB问题的最坏情况遗憾,可以精确地表述为一个半无限线性规划问题。由此得到的停时策略为比较各臂提供了一种自然的方式:策略继续抽样所对应的已知收益越高,说明该未知臂越有前景。我们将这一直觉转化为基于累积继续概率的指标,并引入单调调整和收益缺口上限。通过将指标误差与单臂停时策略的遗憾联系起来,我们对 $K$ 个臂、时域为 $T$ 的情形建立了一个与分布无关的遗憾界 $4.45\sqrt{KT}+10.75K$。该界与文献中已建立的极小化极大最优遗憾阶相匹配。通过伯努利随机化,该保证可推广到支撑在 $[0,1]$ 上的收益分布。我们还给出了一种有限网格实现,并量化了其近似损失。在数值实验中,基于SAB的指标策略在所有评估的臂数和时域设置下,其最坏情况遗憾均低于所有 tested 的基准策略,并且在双臂情形下与基于网格的MAB极小化极大策略的表现十分接近。
cs.LG / 81 / 2609.22701

Autonomous Model Lifecycle Management for Digital Twin-Based Manufacturing Control

面向基于数字孪生的制造控制的自主模型生命周期管理
Zhengyang, Gu, Cook, Thomas, Rohrbaugh, Fredaljohn, Hernandez, Joseph E., Couch, Chris
Abstract
Manufacturing AI systems must autonomously adapt to continuous distributional shift from raw-material variability, ambient changes, and equipment aging, under strict safeguard and operator-trust requirements where model failures risk physical damage. This paper presents a closed-loop Cyber-Physical System (CPS) for autonomous model lifecycle management in automotive manufacturing, deployed since 2023. The system manages product-specialized model pairs: a sequence-to-sequence physics model (LPP) serving as a digital twin, and a deep Reinforcement Learning (RL) control policy (LCP) trained against it. Per retraining cycle, multiple model variants spanning architecture families and RL algorithms compete; only the best-scoring candidate advances. A Conductor orchestrator autonomously manages plant-wide model inventories with dependency-aware retraining and Proportional-Integral-Derivative (PID) fallback. Reflecting the principle of Human-Centric Intelligence, the LCP composite score embeds an operator-trust gate penalizing policies deviating from established practice; without it, 23% of policies are rejected by operators despite passing accuracy thresholds. Across multiple facilities, LCP-controlled processes achieve process stability improvements of 28-45% over uncontrolled baselines with zero safety incidents.
Chinese Translation
制造人工智能系统必须在严格的安全保障和操作员信任要求下,自主适应来自原材料波动、环境变化和设备老化导致的持续分布偏移,因为模型失效可能造成物理损坏。本文提出了一种用于汽车制造中自主模型生命周期管理的闭环信息物理系统(CPS),并自2023年起投入实际部署。该系统管理面向特定产品的模型对:一个作为数字孪生的序列到序列物理模型(LPP),以及一个基于该物理模型训练的深度强化学习(RL)控制策略(LCP)。在每个再训练周期中,涵盖不同架构家族和强化学习算法的多个模型变体相互竞争,只有得分最高的候选模型能够被采纳。一个名为Conductor的编排器可自主管理全厂的模型库存,支持依赖感知的再训练和比例-积分-微分(PID)回退机制。体现了以人为中心的智能原则,LCP的综合评分中嵌入了一个操作员信任门槛,对偏离既有实践的策略进行惩罚;若缺少该机制,23%的通过精度阈值的策略仍会被操作员拒绝。在多个生产设施中,LCP控制的工艺相对于无控制基线实现了28-45%的过程稳定性提升,且安全事件为零。
cs.LG / 82 / 2609.22752

D-IMPL: A Diffusion-based Solver for Parameterized BBOs

D-IMPL:一种基于扩散模型的参数化黑盒优化求解器
Hu, Yang, Li, Na
Abstract
Diffusion models have demonstrated strong power in generative modeling tasks across multiple domains, exhibiting a remarkable capability of learning complex distributions from samples. In this paper, we leverage such capability to design an efficient universal diffusion-based solver for parameterized black-box optimizations (BBO), where the optimizer has only black-box access to queries of the objective function at the learning stage, yet is able to reduce the additional computational cost at the inference stage for each BBO instance while also capturing the potential multi-modal landscape of non-convex objectives. To cast our formulation as a compatible generative modeling task, we introduce the notion of minimization policy as a new solution concept, which defines a sampling distribution over the solutions that should concentrate around the minimizer set for each BBO instance. We then propose Diffusion-based Iterative Minimization Policy Learning (D-IMPL), a practical generative-model-based solver for solving parameterized BBOs that employs diffusion models to learn a minimization policy, whose density is proportional to the exponential of the negated objective values, thereby amortizing the computational costs across different BBO instances. Furthermore, we demonstrate the performance of our D-IMPL algorithm by establishing a sample complexity guarantee showing that a $\delta$-approximate minimization policy can be effectively learned within $O(\log(1/\delta))$ iterations, and by extensive empirical evaluations over a range of constrained and unconstrained BBO tasks.
Chinese Translation
扩散模型(Diffusion models)已在多个领域的生成建模任务中展现出强大的能力,表现出从样本中学习复杂分布的卓越本领。在本文中,我们利用这种能力设计了一种高效的、通用的基于扩散模型的参数化黑盒优化(BBO)求解器。该求解器在学习阶段只能以黑盒方式查询目标函数,但在推理阶段能够降低每个BBO实例的额外计算成本,同时还能捕捉非凸目标函数潜在的多峰地形。为了将我们的表述转化为一个相容的生成建模任务,我们引入了“最小化策略”这一新的解概念,它定义了一个在解上的采样分布,该分布对于每个BBO实例都应集中在极小值点集合附近。随后,我们提出了基于扩散的迭代最小化策略学习(Diffusion-based Iterative Minimization Policy Learning,D-IMPL),这是一种实用的基于生成模型的参数化BBO求解器。它利用扩散模型学习一个最小化策略,其密度与目标函数取负后的指数成正比,从而将计算成本摊销到不同的BBO实例上。此外,我们通过建立样本复杂度保证证明了D-IMPL算法的性能,表明可以在 $O(\log(1/\delta))$ 次迭代内有效学习到一个 $\delta$-近似的最小化策略;并通过在一系列有约束和无约束BBO任务上的大量实验评估验证了其性能。
cs.LG / 83 / 2609.22782

Look Before You Steer: Geometry Predicts SAE Feature Steerability

三思而后行:几何特性预测SAE特征的可操控性
Khan, Muhammad, Channawar, Shlok, Gurugubelli, Akshaj, Gupta, Girish, Shah, Aditya
Abstract
Steering with SAE features requires per-feature coefficient tuning, which currently demands intervention sweeps. We ask whether properties of the SAE itself, computable before any forward pass, predict which features will be cheap or expensive to steer. We show that variation in SAE feature steerability is partially predicted by decoder-space geometry: neighbor density and maximum cosine similarity to nearby decoder directions, both computable from the SAE weight matrix before any intervention, rank features by how much steering they require for a fixed behavioral effect ($\rho$ up to $-0.546$, $p < 10^{-6}$, AUROC 0.610-0.822 across conditions; the signal is rank-based, consistent with grid discreteness). This geometry-steerability relationship replicates across two Gemma-2 model scales (2B and 9B), two SAE widths (16K and 65K), and is detectable cross-architecturally on Llama-3.1-8B-Instruct ($\rho = -0.266$, $n = 300$). On Qwen3-8B with BatchTopK SAEs, geometry predicts whether a feature is steerable at all but not the continuous ordering among responsive features, revealing a boundary condition tied to SAE training regime. The signal weakens at deep proportional layer depth in both models, where the cost of steering exceeds our intervention budget, a consistent depth boundary. These results provide preliminary evidence that pre-steering geometry can partially inform coefficient selection, offering a path toward screening features for controllability before deployment.
Chinese Translation
使用SAE特征进行引导(steering)需要对每个特征单独调节系数,目前这依赖于大量的干预扫描实验。我们探讨是否可以在任何前向传播之前,通过SAE自身可计算的属性来预测哪些特征的引导成本高、哪些成本低。我们表明,SAE特征可操控性的差异可以部分由解码器空间的几何特性预测:邻域密度以及与附近解码器方向的最大余弦相似度——这两者均可在任何干预之前直接从SAE权重矩阵中计算得出——能够按实现固定行为效果所需的引导强度对特征进行排序($\rho$ 最高达 $-0.546$,$p < 10^{-6}$,各条件下 AUROC 为 0.610–0.822;该信号是基于排序的,与网格离散性一致)。这种几何特性与可操控性的关系在两个 Gemma-2 模型规模(2B 和 9B)、两种 SAE 宽度(16K 和 65K)上均得到复现,并在 Llama-3.1-8B-Instruct 上可跨架构检测到($\rho = -0.266$,$n = 300$)。在采用 BatchTopK SAE 的 Qwen3-8B 上,几何特性能够预测某个特征是否完全不可引导,但无法预测可响应特征之间的连续排序,这揭示了与 SAE 训练机制相关的边界条件。在两个模型中,该信号在较深的比例层深度处减弱,此时引导成本超出了我们的干预预算,呈现出一致的深度边界。这些结果提供了初步证据,表明引导前的几何特性可以在一定程度上辅助系数选择,为在部署前筛选特征的可控性提供了一条路径。
cs.LG / 84 / 2609.22783

Improved Private Sparse Covariance Estimation with Multiscale Threshold Tests

基于多尺度阈值检验的改进型隐私稀疏协方差估计
Zhang, Zihan
Abstract
We study differentially private covariance estimation in operator norm for mean-zero sub-Gaussian distributions with unknown covariance support and at most $k$ nonzero entries per row. We develop a multiscale random-threshold algorithm with sample complexity $\ot(k^2/\alpha^2+k\sqrt d/(\alpha\varepsilon))$ for $(\varepsilon,\delta)$-differential privacy and error at most $\alpha\sigma^2$, where $d$ is the dimension and $\sigma$ is a known sub-Gaussian scale. The bound improves the privacy-dependent term of the existing $\ot(k^2/\alpha^2+k^{3/2}\sqrt d/(\alpha\varepsilon))$ \citep{kumar2026curse} upper bound by a factor of $\sqrt k$, and matches the lower bound of $\widetilde{\Omega}(k^2/\alpha^2 + k\sqrt{d}/(\alpha\varepsilon))$ in its applicable parameter regime. Our key technical ingredient is a direct operator-norm bound on the centered fluctuations of an ideal reconstruction, exploiting conditional independence rather than accumulating entrywise errors across each row. A multiscale allocation of threshold tests balances reconstruction variance against query sensitivity. Together, these ingredients sharpen the trade-off between approximation error and privacy protection, removing the additional $\sqrt{k}$ factor from the privacy-dependent sample complexity.
Chinese Translation
我们研究了算子范数下的差分隐私协方差估计问题,针对均值为零的次高斯分布,其协方差支撑集未知且每行至多有 $k$ 个非零元素。我们提出了一种多尺度随机阈值算法,在 $(\varepsilon,\delta)$-差分隐私下,其样本复杂度为 $\ot(k^2/\alpha^2+k\sqrt d/(\alpha\varepsilon))$,误差不超过 $\alpha\sigma^2$,其中 $d$ 为维数,$\sigma$ 为已知的次高斯尺度。该界将现有上界 $\ot(k^2/\alpha^2+k^{3/2}\sqrt d/(\alpha\varepsilon))$(\citep{kumar2026curse})中依赖隐私的项改进了 $\sqrt k$ 倍,并在其适用的参数范围内与下界 $\widetilde{\Omega}(k^2/\alpha^2 + k\sqrt{d}/(\alpha\varepsilon))$ 相匹配。我们的关键技术要素是对理想重构的中心化波动给出直接的算子范数界,利用条件独立性而非逐行累加逐元素误差。阈值检验的多尺度分配在重构方差与查询敏感度之间实现了平衡。这些要素共同优化了近似误差与隐私保护之间的权衡,从依赖隐私的样本复杂度中消除了额外的 $\sqrt{k}$ 因子。
cs.LG / 85 / 2609.22785

Robust Market Making with Hawkes Order Flow and Price Impact via Adversarial Reinforcement Learning

基于对抗强化学习的Hawkes订单流与价格影响的鲁棒做市
Yang, Hao, Xu, Zhenguo
Abstract
Market-making strategies in real limit order book markets face substantial model uncertainty and regime-shift risk. Existing adversarial reinforcement learning approaches improve robustness by formulating the Avellaneda--Stoikov market-making problem as a zero-sum game between a market maker and an environmental adversary. However, these approaches typically rely on Poisson order arrivals and neglect trade-induced price impact, limiting their ability to capture important high-frequency market microstructure effects such as clustered order flow, self-excitation, and post-trade price feedback. We extend adversarial reinforcement learning for market making to a more complex environment with Hawkes self-exciting order arrivals and trade-induced price impact. To mitigate the increased non-stationarity introduced by the expanded regime space, we incorporate an LSTM module that explicitly models the temporal structure of recent observations. We further characterize the equilibrium properties of the proposed framework through both game-theoretic analysis and numerical experiments, and introduce a robustness evaluation protocol focused on improvements in the left tail of the return distribution. Experimental results across a range of market regimes show that the proposed method achieves improved left-tail performance in most complex microstructure environments. In particular, the gains are pronounced in regimes with strong Hawkes excitation and low-to-moderate price impact. Bootstrap tests provide no evidence that these improvements are obtained through a stronger terminal directional inventory bias. These results suggest that combining adversarial training with temporal state representation can improve the robustness of reinforcement-learning-based market-making strategies under order-flow self-excitation, price impact, and regime uncertainty.
Chinese Translation
真实限价订单簿市场中的做市策略面临巨大的模型不确定性与市场状态转换风险。现有的对抗强化学习方法通过将 Avellaneda--Stoikov 做市问题构建为做市商与环境对手之间的零和博弈来提升鲁棒性。然而,这些方法通常依赖泊松订单到达假设,并忽略了交易引起的价格冲击,限制了其捕捉重要的高频市场微观结构效应的能力,例如订单流聚集、自激发以及交易后的价格反馈。我们将对抗强化学习做市扩展到包含 Hawkes 自激发订单到达与交易引起的价格冲击的更复杂环境中。为缓解扩展后的状态空间所带来的非平稳性增加,我们引入了一个显式建模近期观测时间结构的 LSTM 模块。我们进一步通过博弈论分析与数值实验刻画了所提框架的均衡性质,并提出了一种专注于收益分布左尾改善的鲁棒性评估协议。在多种市场状态下的实验结果表明,所提方法在大多数复杂微观结构环境中实现了更好的左尾表现。特别是在 Hawkes 激发较强且价格冲击中低程度的市场状态下,收益尤为显著。Bootstrap 检验未提供任何证据表明这些改善是通过更强的期末方向性持仓偏倚获得的。这些结果表明,将对抗训练与时间状态表征相结合,可以提升基于强化学习的做市策略在订单流自激发、价格冲击以及市场状态不确定性下的鲁棒性。
cs.LG / 86 / 2609.22816

FIRM-WM: State-factorized factual-interventional recurrent modeling for reward-free visual planning

FIRM-WM:面向无奖励视觉规划的状态因子化事实-干预循环建模
Wu, Yilun, Zhang, Yunjian, Li, Aobo, Wang, Mujiangshan, Wu, Haitao, Zhang, Aqiang
Abstract
Reward-free latent world models can learn from offline videos and solve new image--goal tasks by optimizing actions through predicted latent futures. This setting places two demands on the planning state: its coordinates must be comparable with a goal image. Moreover, its dynamics must retain velocity, motion trend, contact, and other history--dependent information beyond those goal coordinates. Offline training creates a second mismatch: each recorded trajectory reveals one factual future, whereas a sampling--based planner compares many actions that were not taken from the same state. We introduce FIRM-WM (Factual--Interventional Recurrent World Model), a compact pixel world model designed around these two gaps. Its recurrent state separates a typed, goal--comparable configuration from a 128-dimensional dynamic fiber used for prediction but excluded from the terminal goal cost. Broad factual trajectories provide state coverage, while common--reset intervention branches provide observed outcomes for alternative action sequences. Before executing each branch, we reset the environment and restore the same recorded values exposed by the environment's state--setting interface. Under matched CEM planning and three independent full-pipeline seeds, FIRM-WM reaches 99.0$\pm$1.0% on TwoRoom, 92.7$\pm$2.1% on Reacher, and 88.0$\pm$3.0% on OGBench-Cube, compared with 89.0%, 88.0%, and 70.0% for LeWM. The deployed model uses 2.98--3.42M parameters and records 2.13--11.60$\times$ lower planning time on these tasks.
Chinese Translation
无奖励的潜在世界模型可以从离线视频中学习,并通过在预测的潜在未来上优化动作来解决新的图像-目标(image-goal)任务。这一设定对规划状态提出了两点要求:其坐标必须可与目标图像进行比较;同时,其动力学必须保留速度、运动趋势、接触以及其他超出目标坐标范围的历史依赖信息。离线训练带来了第二个不匹配:每条记录的轨迹只展示一种事实性的未来,而基于采样的规划器需要比较从同一状态出发的多种未被执行的动作。我们提出了FIRM-WM(事实-干预循环世界模型,Factual-Interventional Recurrent World Model),一个围绕这两个缺口设计的紧凑型像素世界模型。其循环状态将可类型化、可与目标比较的构型与一个用于预测但不计入终端目标代价的128维动态纤维分离开来。广泛的事实轨迹提供状态覆盖,而采用通用重置的干预分支则为备选动作序列提供可观测的结果。在执行每个分支之前,我们重置环境,并通过环境的状态设置接口恢复相同的已记录状态值。在匹配的CEM规划设置和三次独立的全流程随机种子下,FIRM-WM在TwoRoom上达到99.0±1.0%,在Reacher上达到92.7±2.1%,在OGBench-Cube上达到88.0±3.0%,相比之下LeWM分别为89.0%、88.0%和70.0%。部署的模型使用2.98–3.42M参数,并在这些任务上实现了2.13–11.60倍更低的规划时间。
cs.LG / 87 / 2609.22819

Counterfactual Tool Ranking under Utility, Cost, and Privilege Constraints

效用、成本与权限约束下的反事实工具排序
Li, Jiapeng
Abstract
Counterfactual tool evaluation must distinguish authority, historical support, and what a comparison actually estimates. We study these distinctions with eleven executable enterprise-inspired tools, exact-propensity logs, and real local Model Context Protocol transport. An initial 45-run synthetic study is retained, then challenged by 30 realized-return control runs and 15 experiments on 1,930 independently released Berkeley Function Calling Leaderboard (BFCL) tasks. Full-return direct regression reverses an initially favorable doubly robust (DR) evaluation result in the linear setting: mean absolute errors are 0.0139 for direct regression and 0.0272 for DR. Under a shifted environment, DR retains an advantage, with errors 0.0227 versus 0.0948. On function-name-group-disjoint BFCL-derived splits, direct and DR selectors obtain balanced accuracies of 81.85% and 79.83%. Two pinned local Qwen2.5 models are evaluated on the same 200 held-out tasks, exposing a strong failure to abstain under the fixed prompt. We further characterize policy differences under missing support: unsupported actions shared by two policies cancel, allowing point identification of an incremental change when neither absolute value is identifiable. A disagreement-preserving fallback achieves this property in all five support-gap runs, but conservative sampling bounds do not certify deployment improvement. The contribution is a falsifiable evaluation method and independent public evidence, not a new DR estimator, official BFCL leaderboard score, or production-agent safety claim.
Chinese Translation
反事实工具评估必须区分权威性、历史支撑以及比较实际估计的内容。我们通过十一个可执行的企业风格工具、精确倾向性日志以及真实的本地模型上下文协议(Model Context Protocol)传输来研究这些区别。研究保留了一个初始的45次运行的合成实验,随后通过30次实际回报对照运行以及在1,930个独立发布的伯克利函数调用排行榜(Berkeley Function Calling Leaderboard, BFCL)任务上的15组实验对其加以检验。在设定为线性的场景中,全回报直接回归(direct regression)逆转了最初有利的双重稳健(doubly robust, DR)评估结果:直接回归的平均绝对误差为0.0139,而DR为0.0272。在环境发生偏移的情况下,DR保持优势,误差分别为0.0227与0.0948。在按函数名称组划分的BFCL派生互斥数据集上,直接选择器与DR选择器分别取得81.85%和79.83%的平衡准确率。两个固定的本地Qwen2.5模型在相同的200个保留任务上进行评估,暴露出在固定提示词下模型严重缺乏弃权能力。我们进一步刻画了缺失支撑下策略间的差异:两个策略共享的不受支撑的动作相互抵消,使得在两个绝对值均不可识别的情况下仍能对增量变化进行点识别。一个保持分歧的回退机制在全部五组支撑缺口运行中均实现了该性质,但保守的抽样边界并不能证明部署改进。本文的贡献是一种可证伪的评估方法以及独立的公开证据,而非新的DR估计器、官方BFCL排行榜分数或生产级智能体的安全性声明。
cs.LG / 88 / 2609.22820

Beyond Average Error through Oracle-Informed Stress Tests for Time-Series Forecasting

基于先知(Oracle)信息的压力测试突破平均误差:面向时间序列预测
Lin, Xu, Zuo, Runheng, Xu, Shengxuan, Tan, Qitai, Lin, Hongyu, Zhang, Xiao-Ping
Abstract
Average squared error cannot reveal whether forecasting performance degrades because the future becomes less predictable or because forecasts move farther from the conditional mean. We introduce paired, mechanism-controlled stress tests that decompose changes in expected squared error at each lead time into environmental risk and forecast-oracle distance, using an origin-conditioned predictive oracle unavailable to the evaluated models. Three end-to-end controls have known attribution. Specifically, the null, environmental-only, and information-gap controls verify that the pipeline assigns changes to the correct component. We then apply the benchmark to 24 deployable forecasters. Under frequent switching, 14 methods have higher realized MSE but lower oracle distance; under outlier-variance feedback, 19 have higher MSE but lower scale-standardized MSE. Short- and long-lead stress-response rankings have Spearman correlation 0.624, revealing substantial horizon-dependent reordering. We then study multivariate relation shifts. Across six models and three coupling severities, oracle distance accounts for only 0.7-3.9% of the decomposed expected-risk increase, and environmental-risk majority persists in an eight-channel system and a matched-difficulty audit of Ring, Block, and Hub relations. Finally, prespecified contrasts on independent data-generating process (DGP) realizations show that several visually compelling discovery profiles, including trend accumulation and the hypothesized switching reversal, do not replicate. The benchmark thus combines component-wise diagnosis with a held-out stability audit. It complements real-data out-of-distribution evaluation, which measures performance under realistic shifts when exact oracle attribution is unavailable.
Chinese Translation
平均平方误差无法揭示预测性能下降究竟是因为未来变得难以预测,还是因为预测值偏离了条件均值。我们提出了一种成对的、机制受控的压力测试方法,借助一个被评估模型无法获得的、以原点为条件的预测先知(origin-conditioned predictive oracle),将每个预测提前期的期望平方误差变化分解为环境风险和预测-先知距离两个部分。三种端到端的对照实验具有已知的归因结果:具体而言,零假设对照、仅环境对照和信息缺口对照验证了该流程能够将变化正确归因到相应成分。随后,我们将该基准应用于24个可部署的预测模型。在频繁切换情形下,14种方法的实际MSE更高但先知距离更低;在离群值-方差反馈情形下,19种方法的MSE更高但尺度标准化后的MSE更低。短期与长期提前期的压力响应排名的Spearman相关系数为0.624,揭示了显著的随预测范围变化的重排序现象。接着我们研究多变量关系变化:在6个模型和3种耦合强度下,先知距离仅占分解后期望风险增加的0.7%-3.9%,且环境风险占主导的结论在八通道系统以及针对环形(Ring)、块状(Block)和枢纽(Hub)关系的难度匹配审计中依然成立。最后,在独立数据生成过程(DGP)样本上的预设对比检验表明,若干视觉上颇具吸引力的发现模式,包括趋势累积和假设的切换反转效应,均未能复现。因此,该基准将成分级诊断与留出稳定性审计相结合,是对真实数据分布外评估的补充——后者适用于无法获得精确先知归因时的现实偏移下的性能测量。
cs.LG / 89 / 2609.22833

Personalized Federated Reinforcement Learning via Model-Agnostic Meta-Learning: Convergence of Exact and Hessian-Free Meta-Policy Gradients

基于模型无关元学习的个性化联邦强化学习:精确与无Hessian元策略梯度的收敛性分析
Beikmohammadi, Ali, Khirirat, Sarit, Magnússon, Sindri
Abstract
We study personalized federated reinforcement learning, in which $n$ agents, each acting in its own Markov decision process, collaborate through a server to learn a shared MAML-style policy initialization that becomes effective for an individual agent once that agent adapts it with a single local policy-gradient step. We propose Per-FedAvg-PG, in which agents take $\tau$ local stochastic meta-policy-gradient steps between communication rounds, and prove that it reaches an $\varepsilon$-approximate first-order stationary point of the personalized objective in $K=\mathcal O(\varepsilon^{-3/2})$ rounds with $\tau=\Theta(\varepsilon^{-1/2})$ local steps. The analysis rests on a structural feature of the reinforcement learning setting: under standard policy-class regularity, the per-agent objectives have uniformly bounded gradients and Hessians with explicit constants, so the bounded-gradient and bounded-heterogeneity conditions imposed by the supervised theory hold automatically and no separate heterogeneity assumption is needed. The exact meta-gradient requires the inner-loop policy Hessian, which our experiments identify as the practical bottleneck. We therefore analyze the Hessian-free variant, bound its bias, and exhibit fixed points at which the meta-gradient is nonzero and of order $\alpha$, showing that the resulting stationarity floor is a property of the method rather than of the bound. Experiments on tabular and neural navigation confirm the predicted behavior and show transfer to unseen agents at an order of magnitude lower sample cost than independent training. Together these results identify the adaptation step size as a tunable personalization knob and the curvature estimate as the quantity that governs whether exact meta-gradients are affordable.
Chinese Translation
我们研究个性化联邦强化学习问题,其中 $n$ 个智能体各自运行于自己的马尔可夫决策过程(MDP)中,通过服务器协作学习一个共享的 MAML 风格策略初始化;该初始化在单个智能体经过一步本地策略梯度适配后,即可对该智能体生效。我们提出 Per-FedAvg-PG 算法,其中各智能体在通信轮次之间执行 $ au$ 步本地随机元策略梯度更新,并证明该算法在 $K=\mathcal{O}(\varepsilon^{-3/2})$ 轮、每轮 $ au=\Theta(\varepsilon^{-1/2})$ 步本地更新的条件下,可达到个性化目标的 $\varepsilon$-近似一阶稳定点。我们的分析依赖于强化学习环境的一个结构性特征:在标准的策略类正则性条件下,各智能体目标函数的梯度和 Hessian 矩阵具有显式常数的一致有界性,因此有监督学习理论中所施加的有界梯度和有界异质性条件自动成立,无需额外假设异质性。精确元梯度需要计算内层循环的策略 Hessian 矩阵,我们的实验表明这是实际计算中的瓶颈。因此,我们进一步分析了无 Hessian(Hessian-free)变体,给出了其偏差的上界,并展示了元梯度非零且量级为 $\alpha$ 的不动点,说明由此产生的平稳性下限是该方法本身的性质,而非上界估计的产物。在表格型和神经网络导航任务上的实验验证了理论预测的行为,并表明向未见过的智能体的迁移性能优于独立训练,样本成本降低了一个数量级。综上,这些结果指出适配步长可作为可调节的个性化控制旋钮,而曲率估计则决定了精确元梯度是否可负担。
cs.LG / 90 / 2609.22836

A Hybrid Attention Model Learning Unified Time-aware Patch Representation for Irregular Multivariate Time Series Forecasting

面向不规则多变量时间序列预测的统一时间感知片段表示的混合注意力模型
Lin, Zhihao, Lin, Li, Zhang, Qi, Xia, Kaiwen, Wang, Shuai, Qiao, Jialin
Abstract
Time series foundation models (TSFMs) have recently delivered impressive zero-shot performance across diverse forecasting tasks. However, real-world decision-making frequently relies on \emph{irregular multivariate time series} (IMTS), where inconsistent inter-observation intervals and asynchronous sampling across variables coexist with informative missingness. Existing TSFMs handle such inputs either through imputation that injects spurious values or through index-based positional encodings that ignore continuous time. There is still a gap in the foundation model that follows the original IMTS patterns. In this paper, we propose a hybrid attention model that learns a unified time-aware patch representation for IMTS forecasting. We first design a \emph{time-aware patch encoding} that maps a variable number of intra-patch timestamps into a fixed-size embedding, producing a uniform format for irregular patches without resorting to imputation. We then introduce a \emph{time bias attention} mechanism that calibrates inter-patch temporal misalignment and asynchronous cross-channel dependencies as auxiliary attention offset. Finally, on top of a decoder-only Transformer backbone, we adopt a \emph{hybrid causal mask} that preserves a bidirectional full view over the historical context while keeping the forecast horizon strictly autoregressive. To support large-scale pretraining under irregular settings, we also curate VersaTSA, an archive of $30$B observations that retains the native sampling sparsity of its sources. Experiments on three IMTS benchmarks and a standard regular-MTS benchmark show that our model achieves state-of-the-art zero-shot performance on IMTS and remains competitive when transferred to regular forecasting.
Chinese Translation
时间序列基础模型(TSFM)近年来在多种预测任务中展现出令人瞩目的零样本性能。然而,现实世界的决策往往依赖于不规则多变量时间序列(IMTS),其中观测间隔不一致、变量间采样异步以及信息性缺失同时存在。现有的TSFM在处理此类输入时,要么通过插补引入虚假数值,要么采用基于索引的位置编码而忽略连续时间信息。遵循原始IMTS模式的基础模型仍然存在空白。在本文中,我们提出了一种混合注意力模型,用于学习面向IMTS预测的统一时间感知片段表示。我们首先设计了一种时间感知片段编码,将片段内数量不定的多个时间戳映射为固定尺寸的嵌入,从而在不依赖插补的情况下为不规则片段生成统一格式。随后,我们引入了一种时间偏置注意力机制,将片段间的时间错位以及跨通道的异步依赖关系校准为辅助注意力偏移量。最后,在仅解码器的Transformer骨干网络之上,我们采用了一种混合因果掩码,既保留对历史上下文的双向完整视角,又确保预测区间严格保持自回归特性。为支持不规则设定下的大规模预训练,我们还构建了VersaTSA,一个包含300亿观测值的语料库,并保留了其数据源的原生采样稀疏性。在三个IMTS基准和一个标准规则多变量时间序列基准上的实验表明,我们的模型在IMTS上实现了最先进的零样本性能,且迁移到规则时间序列预测时仍具有竞争力。
cs.LG / 91 / 2609.22850

Testing the Construct Validity of a Functional Valence Axis in LLM Agents

检验大语言模型智能体中功能性效价轴的结构效度
Li, Weihan, Chen, Xinlei, Song, Yuhan, Lin, Xiaofeng, Zheng, Tianshi
Abstract
Contrastive activation directions are often interpreted from what they decode or how strongly they steer behavior. But what evidence is sufficient to identify the construct represented by such a direction, rather than a correlated feature of the contrast used to extract it? We study this question for a good--bad outcome direction in a maze task, using controlled interventions that separate the realised outcome from the informational history through which it became known. Across multiple LLM checkpoints, directions fitted on one explicit outcome encoding transfer well to another, indicating that the readout is not tied to surface form. In contrast, when the same realised outcome is reached through announced and unannounced histories, transfer degrades substantially: even after both histories receive the same explicit outcome, the post-event readout remains strongly conditioned on the earlier announcement. In a matched maze-RL run, the post-RL direction becomes substantially more predictive of reference-MDP remaining return and the policy becomes more dependent on it at the tested sites, while this history dependence persists. These results support a functional, value-related interpretation of the direction, but not its identification with a history-invariant scalar valence state.
Chinese Translation
对比激活方向通常依据其能解码出什么内容或其引导行为的强度来加以解释。但需要什么证据才能确定这类方向所表征的构念,而非用于提取该方向的对比中与之相关的特征?我们在迷宫任务中针对一个“好—坏”结果方向研究这一问题,采用受控干预将实际结果与获知该结果的信息历史分离开来。在多个LLM检查点上,基于一种显式结果编码拟合的方向能够良好地迁移到另一种显式结果编码,表明该读出并不依赖于表面形式。相反,当同一实际结果通过已告知与未告知两种历史情境达成时,迁移显著退化:即使在两种历史情境接收到相同的显式结果之后,事件后的读出仍然强烈依赖于先前是否被告知。在一项匹配的迷宫强化学习运行中,强化学习后的方向对参考MDP剩余回报的预测能力显著增强,策略在测试位置也变得更加依赖该方向,同时这种历史依赖性依然存在。这些结果支持对该方向的功能性、与价值相关的解释,但并不支持将其等同于一种历史不变的标量效价状态。
cs.LG / 92 / 2609.22862

CurvFlow-DTA: dual-graph discrete Ricci curvature flow for drug--target affinity prediction

CurvFlow-DTA:用于药物-靶点亲和力预测的双图离散里奇曲率流方法
Ma, Jicheng, Yang, Yunyan, Zhao, Juan, Zhao, Liang
Abstract
Graph neural networks are widely used for drug--target affinity (DTA) prediction, and discrete Ricci curvature has recently been used to characterize molecular graph geometry. Existing curvature-aware DTA approaches mainly use static curvature on the drug graph while representing proteins primarily with sequence-derived features. This leaves pair-adaptive use of graph geometry underexplored, which may limit adaptation to unseen entities in cold-start settings relevant to practical screening. We present CurvFlow-DTA, which replaces a single static curvature representation with weighted Forman curvature flow on both molecular and protein residue--residue contact graphs. A label-independent flow trajectory is precomputed for each entity, and a pair-conditioned selector determines the horizons read by a dual-branch Flow-GINE. A frozen ESM-2 supplies residue-level representations and contact scores used to construct the protein graph. Inference requires only SMILES strings and protein sequences, without a bound complex structure. On Davis and KIBA, CurvFlow-DTA improves on the protocol-matched Ricci-GraphDTA baseline in every warm and cold-start setting. Warm-split mean squared error (MSE) decreases by $19.9\%$ on Davis and $18.9\%$ on KIBA. Across the six cold-start comparisons, MSE decreases by $14.3$--$27.4\%$, with higher concordance index (CI) throughout. Within our compiled set of literature baselines, CurvFlow-DTA achieves the lowest MSE on both warm benchmarks and across four out of six cold-start evaluation settings.
Chinese Translation
图神经网络被广泛应用于药物-靶点亲和力(DTA)预测,而离散里奇曲率近年来也被用于刻画分子图的几何特性。现有的曲率感知DTA方法主要在药物图上使用静态曲率,而对蛋白质的表示主要依赖于序列衍生特征。这使得针对图几何的配对自适应利用仍未被充分探索,从而可能限制模型在冷启动场景(与实际筛选密切相关)中对未见实体的适应能力。我们提出CurvFlow-DTA,用加权Forman曲率流替代单一的静态曲率表示,并同时作用于分子图和蛋白质残基-残基接触图。对每个实体预先计算一条与标签无关的曲率流轨迹,由配对条件选择器决定双分支Flow-GINE读取的时间视界。一个冻结的ESM-2提供残基级表示和接触分数,用于构建蛋白质图。推理过程仅需SMILES字符串和蛋白质序列,无需结合复合物结构。在Davis和KIBA数据集上,CurvFlow-DTA在所有暖启动和冷启动设置下均优于协议一致的Ricci-GraphDTA基线。暖启动划分的均方误差(MSE)在Davis上降低19.9%,在KIBA上降低18.9%。在六项冷启动对比中,MSE降低14.3%至27.4%,且一致性指数(CI)全程更高。在我们整理的文献基线集合中,CurvFlow-DTA在两个暖启动基准上以及六项冷启动评估设置中的四项上均取得了最低的MSE。
cs.LG / 93 / 2609.22866

Causilo Technical Report

Causilo 技术报告
Cho, Minyong, Jeong, Minho, Lee, Dooho, Lee, Jinmo, Yoo, Jaemin
Abstract
We introduce Causilo, a tabular foundation model (TFM) that combines frontier predictive performance with exceptionally fast inference. On TabArena, Causilo achieves 1785.4 Elo, at a median inference time of 0.10 seconds per 1K test samples. It outperforms TabPFN-3.5-Fast with 31.6% less inference time, placing it on the performance--efficiency Pareto frontier. Causilo follows TabICL's column-then-row architecture but introduces another row-refinement module before row compression. This module exchanges information among cell representations within each row after column encoding. The refined cells then visit the context set again through an additional column stage before being compressed into row embeddings. For inference efficiency, both row stages use cross-attention through a fixed number of summary tokens, keeping their attention cost linear in the number of features. Pretrained on approximately 36M synthetic tables, Causilo delivers strong benchmark results across TabArena, BeyondArena, and ScoringBench, achieving frontier-level performance with substantially faster inference.
Chinese Translation
我们提出了 Causilo,一个兼具前沿预测性能与极快推理速度的表格基础模型(TFM)。在 TabArena 基准上,Causilo 取得 1785.4 的 Elo 分数,对每 1K 个测试样本的中位推理时间仅为 0.10 秒。相比 TabPFN-3.5-Fast,其推理时间减少 31.6%,位于性能—效率 Pareto 前沿。Causilo 采用 TabICL 的先列后行架构,但在行压缩之前引入了另一个行精炼模块。该模块在列编码之后,于每行内的单元格表示之间交换信息。精炼后的单元格随后通过额外的列阶段再次访问上下文集合,之后被压缩为行嵌入。为提高推理效率,两个行阶段均通过固定数量的摘要 token 进行交叉注意力,使其注意力开销与特征数量呈线性关系。Causilo 在约 3600 万张合成表格上预训练,在 TabArena、BeyondArena 和 ScoringBench 基准上均取得强劲结果,以显著更快的推理实现了前沿级性能。
cs.LG / 94 / 2609.22867

Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories

利用去噪轨迹全局调度以挖掘扩散模型的推理时计算潜力
Cao, Yuan, Tang, Yifu, Li, Hangqi, Zheng, Zeyu
Abstract
Diffusion models generate a sample by traversing a denoising trajectory, a sequence of stochastic noise-reduction steps that transforms pure noise into a draw from a target distribution. At deployment time, additional computation can improve sample quality without retraining: at each step, the sampler draws several candidate noise samples, scores the resulting predictions with a quality criterion called the verifier, and retains the best candidate at the cost of one network evaluation per candidate. This raises a resource allocation question: given a fixed budget of function evaluations, how should search effort be distributed across the steps of the denoising trajectory? We formulate this as a computational budget allocation problem. First, we show that, to leading order in the step size, the expected gain from evaluating $K$ candidates at a step factorizes into an endogenous, step-specific sensitivity parameter times a universal sample-size factor equal to the expected best of $K$ standard-normal draws. Second, for a fixed sensitivity profile, the optimal allocation solves a separable concave integer program with water-filling structure; at fixed total sensitivity, its advantage over uniform allocation increases with sensitivity dispersion in the majorization order. Third, we prove that when sensitivities vary across instances, no adaptive policy can avoid worst-case regret that grows linearly in the trajectory length, which motivates a design that anchors the allocation offline and adapts online only to recover instance-specific slack. We extend the analysis from independent random search to a broader family of local search operators, and instantiate it as an implementable algorithm. Experiments on three families of diffusion samplers show that the proposed allocation attains the quality of the uniform benchmark with 20 to 50 percent fewer function evaluations.
Chinese Translation
扩散模型通过遍历一条去噪轨迹来生成样本,该轨迹是一系列将纯噪声逐步转化为目标分布样本的随机降噪步骤。在部署阶段,无需重新训练即可通过额外计算提升样本质量:在每一步中,采样器抽取若干候选噪声样本,用一个称为验证器(verifier)的质量准则对所得预测进行评分,并保留最优候选,其代价是每个候选需一次网络评估。这引出了一个资源分配问题:在函数评估次数固定的预算下,搜索 effort 应如何在去噪轨迹的各步骤之间分配?我们将此形式化为一个计算预算分配问题。首先,我们证明,在一阶步长近似下,在某一步评估 $K$ 个候选的期望收益可分解为一个内生且与步骤相关的敏感度参数与一个通用样本量因子(等于 $K$ 个标准正态样本的期望最优值)的乘积。其次,对于固定的敏感度分布,最优分配可归结为一个具有注水(water-filling)结构的可分离凹整数规划;在总敏感度固定的条件下,其在优超序意义下相对均匀分配的优势随敏感度离散度的增大而增大。第三,我们证明当敏感度随实例变化时,任何自适应策略都无法避免随轨迹长度线性增长的最坏情况遗憾(regret),这启发了一种先离线确定基准分配、在线仅调整以恢复实例特有余量的设计。我们将分析从独立随机搜索扩展到更广泛的一类局部搜索算子,并将其具体化为一个可实现的算法。在三类扩散采样器上的实验表明,所提出的分配方法仅需均匀基准 20% 至 50% 更少的函数评估即可达到相同的样本质量。
cs.LG / 95 / 2609.22870

Towards Full Pipeline FP8 Reinforcement Learning for LLMs

面向大语言模型的全流程FP8强化学习
Chen, Fanchao, Jiang, Ziheng, Wei, Ziyun, Zhong, Zheng, Li, Du, Zhang, Chi, Lin, Haibin, Venkataraman, Shivaram
Abstract
Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8 RL pipeline remains challenging. While previous works have focused on resolving train-inference mismatches using correction techniques like TIS, we reveal that full-pipeline FP8 RL still suffers from severe training instability, manifesting as anomalous mid-training entropy surges and garbled outputs. We trace this instability to a previously overlooked cause: compounded FP8 quantization noise distorts the importance ratio, disproportionately pushing negative-advantage tokens outside the trust region and erroneously zeroing out their gradients. As a result, pathological outputs are not properly penalized and accumulate over the course of training. To address this, we propose Calibrated Clipping, a dynamic method that aligns the FP8 clipping bounds with high-precision BF16 distributions by matching the lower-bound clipping quantile and rebalancing the upper bound accordingly. Extensive experiments across GRPO and DAPO algorithms, model scales from 8B to 32B, and multiple FP8 scaling granularities demonstrate that our approach successfully eliminates entropy surges and restores performance comparable to the BF16 baseline.
Chinese Translation
强化学习(RL)已成为提升大语言模型(LLM)推理与智能体能力的关键技术。尽管FP8量化可以加速强化学习训练,但在整个FP8强化学习流程中保持稳定性仍然充满挑战。以往的工作主要聚焦于利用TIS等修正技术来解决训练与推理之间的不匹配问题,而我们发现全流程FP8强化学习仍然存在严重的训练不稳定问题,具体表现为训练中期的异常熵激增以及乱码输出。我们将这种不稳定归因于一个此前被忽视的原因:叠加的FP8量化噪声扭曲了重要性比率,不成比例地将负优势(negative-advantage)token推出信任域,并错误地将其梯度置零。其结果是,病态输出未受到应有的惩罚,并在训练过程中不断累积。为解决这一问题,我们提出Calibrated Clipping(校准裁剪),一种动态方法:通过匹配下界裁剪分位数并相应重新平衡上界,使FP8裁剪边界与高精度BF16分布对齐。在GRPO和DAPO算法、8B至32B的模型规模以及多种FP8缩放粒度上的大量实验表明,我们的方法成功消除了熵激增,并恢复了与BF16基线相当的性能。
cs.LG / 96 / 2609.22879

Prioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy Optimization

面向高效基于世界模型的视觉-语言-动作策略优化的优先级采样
Sheng, Yifei, Ren, Haoxiang, Zhang, Zhilong, Wang, Haonan, Xu, Runjie, Sun, Yihao, Tang, Nan, Wu, Zhichao, Yuan, Lei, Lin, Haoxin, Yu, Yang
Abstract
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, but fine-tuning them with reinforcement learning (RL) remains constrained by the cost of real-world robot interaction. Model-based reinforcement learning (MBRL) reduces this cost by using a learned world model to generate rollouts for policy optimization. However, it becomes computationally expensive as VLA policies and world models scale. Existing methods typically treat states equally, overlooking substantial differences in their utility for policy improvement. In this paper, we show that policy uncertainty helps identify states with greater potential for policy improvement. The policy exhibits high uncertainty at only a small subset of states, often during decision-sensitive stages where small action differences can alter task outcomes, suggesting that policy improvements at these states could be particularly valuable. Building on these findings, we introduce U-GROW, a lightweight, plug-and-play sampling layer that directs more model rollouts to these informative states. By modifying only the branched-start distribution, U-GROW can be integrated into existing MBRL pipelines without changing the policy optimization objective. Experiments in both simulated and real-world manipulation tasks demonstrate the efficiency and effectiveness of U-GROW, supporting the use of policy uncertainty to guide experience generation.
Chinese Translation
视觉-语言-动作(VLA)模型已成为具身智能的一种强大范式,但使用强化学习(RL)对其进行微调仍受到真实世界机器人交互成本的制约。基于模型的强化学习(MBRL)通过利用学习到的世界模型生成交互数据用于策略优化,从而降低了这一成本。然而,随着VLA策略和世界模型规模的扩大,其计算开销变得十分高昂。现有方法通常对所有状态一视同仁,忽视了它们在策略改进效用上的显著差异。本文证明,策略不确定性有助于识别具有更大策略改进潜力的状态。策略仅在少量状态子集上表现出高不确定性,而这些状态往往处于决策敏感阶段——此时微小的动作差异就可能改变任务结果,这表明在这些状态上进行策略改进可能尤为有价值。基于这些发现,我们提出了U-GROW,一个轻量级、即插即用的采样层,它将更多的模型交互数据分配给这些信息丰富的状态。U-GROW仅通过修改分支起始分布即可集成到现有的MBRL流程中,而无需改变策略优化目标。在仿真和真实世界操作任务上的实验证明了U-GROW的高效性和有效性,支持了利用策略不确定性指导经验生成的思路。
cs.LG / 97 / 2609.22886

Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion

Merge++:基于免数据检查点反演的通用合并精炼方法
Pola, Aditya, Balasubramanian, Vineeth N.
Abstract
Model merging consolidates fine-tuned experts into one multi-task model without retraining. All existing data-free methods approach this problem entirely in weight space. Restricted to arithmetic on parameters, these methods never observe how each expert behaves, a signal that only emerges through forward evaluation. Accessing this behavioral signal requires inputs to evaluate on, which the data-free setting prohibits. We propose Merge++, a post-hoc method that addresses this by inverting the expert checkpoints to synthesize task-representative images, then distilling expert knowledge into the merged model using those images. Merge++ requires no additional data beyond the checkpoints themselves. It applies universally across merging algorithms and operates as a complementary refinement stage independent of the underlying weight-space method. The method consistently improves merging algorithms ranging from simple task arithmetic to state-of-the-art spectral methods, with average gains of +2 to +8 points and up to +25.9 on individual configurations.
Chinese Translation
模型合并(Model Merging)旨在将多个微调后的专家模型整合为一个多任务模型,而无需重新训练。现有的所有免数据方法都完全在权重空间中处理这一问题。由于仅限于对参数进行算术运算,这些方法从未观察到每个专家模型的行为表现——而这一信号只能通过前向推理才能显现。获取该行为信号需要在输入数据上进行评估,而这正是免数据设定所禁止的。我们提出 Merge++,一种事后(post-hoc)方法,通过对专家检查点进行反演来合成具有任务代表性的图像,进而利用这些图像将专家知识蒸馏到合并后的模型中。Merge++ 除检查点本身外不需要任何额外数据。它可通用地应用于各种合并算法,并作为一个互补的精炼阶段,独立于底层的权重空间方法。该方法能够持续提升从简单的任务算术到最先进的谱方法等多种合并算法的性能,平均提升 2 至 8 个百分点,个别配置下最高提升达 25.9 个百分点。
cs.LG / 98 / 2609.22894

Are Coreset Selection Methods Worth Their Cost?

核心集选择方法物有所值吗?
Liu, Yangze, Han, Zhongyi
Abstract
Coreset selection picks a representative subset of the labeled training set to make training cheaper. However, it is usually evaluated by downstream accuracy at a fixed subset size, ignoring both the time spent selecting the subset and the training recipe behind each reported number. We introduce an end-to-end benchmark that standardizes downstream training and charges selection and training to the same auditable wall-clock budget, spanning 4 datasets from CIFAR-10 to ImageNet-1K, 11 selectors, 5 fractions, and 3 seeds, with over 1,500 released runs. Repeated-sampling work has shown that budget-aware evaluation already favors random strategies. Our two budget studies test whether that verdict survives when every selector is granted its most favorable operating point. Across eight wall-clock budget anchors on each of CIFAR-10 and Tiny ImageNet, no anchor is won by a sophisticated selector: every winner is class-balanced random sampling, repeated random sampling, or full-data training. In fixed-budget duels on ImageNet-1K, training on all data for fewer epochs beats every selection strategy we probe while also costing the least. A per-dataset cost audit shows that selection cost is dominated at every scale by a fixed full-dataset scan, so it cannot be amortized away by selecting a smaller fraction, and its absolute size does not extrapolate from one dataset to another. We further quantify when selection does pay back through subset reuse, and document 9 correctness fixes to a widely used codebase, one of which shifts a standard Herding baseline by nearly 6 points. Selection time is not free preprocessing, and an evaluation that ignores it measures the wrong quantity.
Chinese Translation
核心集选择(Coreset selection)通过从带标签的训练集中挑选一个具有代表性的子集来降低训练成本。然而,通常的评估方式只是在固定的子集规模下考察下游准确率,既忽略了选择子集所花费的时间,也忽略了每个已报告结果背后的训练方案。我们提出了一个端到端的基准测试,将下游训练标准化,并把选择与训练的开销计入同一个可审计的墙钟时间预算。该基准涵盖从 CIFAR-10 到 ImageNet-1K 的 4 个数据集、11 种选择器、5 个比例和 3 个随机种子,共发布了超过 1,500 次运行结果。重复采样的研究表明,考虑预算的评估本身就更有利于随机策略。我们的两项预算研究检验了当每个选择器都被赋予其最有利的运行点时,这一结论是否仍然成立。在 CIFAR-10 和 Tiny ImageNet 上各设八个墙钟时间预算锚点,没有任何一个锚点被复杂的选择器赢得:所有胜出者都是类平衡随机采样、重复随机采样或全量数据训练。在 ImageNet-1K 上的固定预算对决中,用更少的轮次(epoch)在全部数据上训练,其效果优于我们测试的所有选择策略,且成本最低。对每个数据集的成本审计表明,在所有规模下,选择成本都由一次固定的全量数据扫描所主导,因此无法通过选择更小的比例来摊薄,且其绝对规模也无法从一个数据集外推到另一个数据集。我们进一步量化了选择在何种情况下能够通过子集复用收回成本,并记录了对一个广泛使用的代码库的 9 处正确性修复,其中一处使标准的 Herding 基线结果移动了近 6 个百分点。选择时间并非免费的预处理,忽视它的评估所度量的并不是正确的量。
cs.LG / 99 / 2609.22919

Computationally efficient safe exploration in reinforcement learning

强化学习中计算高效的安全探索
Murali, Shreeram, Deka, Shankar A., Baumann, Dominik
Abstract
Reinforcement learning in real-life applications requires safety guarantees during exploration. Typical reinforcement learning algorithms do not provide such guarantees, and many modifications that do rely on Gaussian processes (GPs), which have a large computational cost. We propose a computationally lightweight algorithm based on the Nadaraya-Watson estimator that safely explores and optimizes constrained Markov decision processes (MDPs). Our algorithm, \textsc{CoLSafe-MDP}, uses an estimator that scales in constant-time with bounds on the estimates, a significant improvement from its GP-based counterparts that scale cubically with the number of data points. We then evaluate its performance in a grid-based environment and on observational Martian terrain data.
Chinese Translation
强化学习在现实应用中的探索过程需要安全性保证。典型的强化学习算法无法提供此类保证,而许多能够提供保证的改进方法依赖于高斯过程(GP),其计算代价较高。我们提出了一种基于 Nadaraya-Watson 估计器的轻量级计算算法,可安全地探索和优化约束马尔可夫决策过程(MDP)。我们的算法 CoLSafe-MDP 所使用的估计器以常数时间扩展并带有估计值上界,相比其基于高斯过程的同类方法(其计算量随数据点数量呈立方增长)是一个显著的改进。随后,我们在基于网格的环境中以及火星地形的观测数据上评估了该算法的性能。
cs.LG / 100 / 2609.22932

Joint Domain-Class Modeling for Federated Learning Under Feature Skew

特征偏斜下联邦学习的域-类联合建模
Najafi, Sina, Tavassolipour, Mostafa, Shariatpanahi, Seyed Pooya
Abstract
Federated learning (FL) enables collaborative model training without centralizing private data, but performance often degrades under feature skew: clients share labels while the conditional input distributions $p_i(x\!\mid\!y)$ vary due to latent, client-specific appearance factors. We propose Joint Domain-Class Federated Learning (JDFL), a lightweight, optimizer-agnostic extension that makes this latent domain variation usable without sharing raw data. JDFL first infers domain clusters called pseudo-domains from brief local update signals. It then expands the classifier head to output $M\times C$, joint (domain-class) logits. This allows the model to represent domain-conditioned appearance while keeping a shared backbone. To train the expanded head we introduce two complementary supervision strategies based on simple intuitions: a similarity-aware soft-labeling that transfers evidence between nearby inferred domains while allowing domain-specific specialization, and a per-sample randomized target assignment that perturbs supervision across the joint outputs and serves as a low-cost training-time regularizer. JDFL integrates with existing standard FL methods (e.g., FedAvg, SCAFFOLD) with minimal changes. Empirically, both supervision modes consistently improve global test accuracy on standard domain-shifted image benchmarks; ablations and sensitivity studies show the gains stem from the proposed supervision and parametrization rather than mere capacity increase.
Chinese Translation
联邦学习(Federated Learning, FL)支持在不集中私有数据的情况下进行协同模型训练,但在特征偏斜(feature skew)场景下性能往往下降:客户端共享标签,而条件输入分布 $p_i(x\mid y)$ 由于潜在的、客户端特有的外观因子而各不相同。我们提出联合域-类联邦学习(Joint Domain-Class Federated Learning, JDFL),这是一种轻量级、与优化器无关的扩展方法,能够在不共享原始数据的前提下使这种潜在域差异变得可用。JDFL 首先从简短的本地更新信号中推断出称为伪域(pseudo-domain)的域簇;随后将分类器头部扩展为输出 $M\times C$ 维的联合(域-类)logits,使模型能够在保留共享主干网络的同时表示域条件下的外观特征。为训练扩展后的分类头,我们基于简单的直觉引入了两种互补的监督策略:一是相似性感知的软标签(similarity-aware soft-labeling),它在相近的推断域之间传递证据,同时允许域特定的专门化;二是逐样本随机化目标分配(per-sample randomized target assignment),它在联合输出上对监督进行扰动,充当一种低成本的训练期正则化手段。JDFL 仅需极小的改动即可与现有标准联邦学习方法(如 FedAvg、SCAFFOLD)集成。实验表明,两种监督模式在标准域偏移图像基准上均能持续提升全局测试精度;消融实验与敏感性分析显示,性能提升源于所提出的监督方式与参数化设计,而非单纯的容量增加。
cs.LG / 101 / 2609.22943

Token Utility Is Selection-Conditioned: Coupled Selection of Prompt Context and Response Supervision for Efficient Instruction Tuning

Token效用依赖于选择条件:面向高效指令微调的提示上下文与响应监督耦合选择
Wu, Can, Chen, Xinrui, Wu, Ou, Du, Yi
Abstract
Efficient large language model (LLM) instruction tuning requires selecting response supervision with supporting prompt context. Existing methods typically value both sides separately, risking selection-state mismatch between valuation and retained training subsets. BRIDGE (Budgeted Response-Prompt Interaction via Directional Gradient-guided Efficient Token Selection) captures selection-conditioned token utility through a shared validation-directed interaction surrogate valuing each side under the other's retained state. Budgeted alternating selection coordinates retained subsets by aggregating precomputed interactions over the current opposite-side subset to update conditional scores. Structure-aware projection converts conditional response scores into coherent supervision spans. Across three model families, BRIDGE leads compared selection methods overall in mathematical reasoning, code generation, and instruction following. In mathematical reasoning, its advantage over independent selection grows with compression.
Chinese Translation
高效的大语言模型(LLM)指令微调需要在支持性提示上下文的条件下选择响应监督。现有方法通常分别独立评估两侧的价值,导致评估状态与保留训练子集之间可能存在选择状态不匹配的问题。BRIDGE(通过方向性梯度引导的高效Token选择实现预算化响应-提示交互)通过一个共享的、面向验证集的交互代理模型来捕获依赖于选择条件的Token效用,即在另一方保留状态下对每一侧进行价值评估。预算化的交替选择通过聚合在当前对侧子集上预先计算的交互结果来更新条件分数,从而协调双方保留的子集。结构感知投影将条件化响应分数转换为连贯的监督片段。在三个模型系列上,BRIDGE在数学推理、代码生成和指令遵循任务中总体领先于对比的选择方法。在数学推理任务中,其相对于独立选择方法的优势随压缩率提高而增大。
cs.LG / 102 / 2609.22977

Beyond Similarity: Coverage-Aware Prompt Selection for Time Series Forecasting with LLMs

超越相似性:面向大语言模型时间序列预测的覆盖感知提示选择方法
Ji, Daeun, Kim, Minkyoung, Kim, Dongkuk, Lee, Yohan, Kim, Beomsoo, Jang, Beakcheol
Abstract
Similarity-based retrieval is the dominant rule for conditioning large language models (LLMs) in in-context learning, retrieval-augmented generation, and prompt-based time series forecasting. The rule concentrates on near-duplicate candidates, an issue that has motivated diversity-aware retrieval but remains unexamined in other retrieval-conditioned pipelines. We study this issue using prompt-based time series forecasting as a test bed, where a learned prompt pool is retrieved by similarity. Dominant methods in this setting retrieve top-K entries by cosine similarity without redundancy control, producing a bias toward dominant temporal patterns while overlooking rare but informative events. We propose CASP-LLM, a coverage-aware semantic prompting framework that addresses this prompt selection bias by combining usage-tracking and saturating-gate techniques into a coverage regularizer that adds no learnable parameters. On six long-term benchmarks and the M4 short-term benchmark, CASP-LLM matches or improves on similarity-based LLM forecasters on most dataset-horizon settings, with the exceptions of Electricity, M4-Monthly, and the few-shot long-horizon setting. A controlled study locates the failure mode at the cross-batch usage level rather than per-retrieval redundancy: within-retrieval diversification such as MMR does not help, whereas regularizing anchor usage across training does.
Chinese Translation
基于相似性的检索是在上下文学习、检索增强生成以及基于提示的时间序列预测中对大语言模型(LLM)进行条件化的主流规则。该规则聚焦于近似重复的候选样本,这一问题已促使了多样性感知检索的发展,但在其他检索条件化的流程中尚未得到研究。我们以基于提示的时间序列预测作为测试平台来研究这一问题,其中通过相似性从学习到的提示池中进行检索。该设置下的主流方法通过余弦相似度检索前K个条目而不进行冗余控制,导致对主导性时间模式的偏向,同时忽视了稀有但信息丰富的事件。我们提出CASP-LLM,一种覆盖感知的语义提示框架,通过将使用情况跟踪与饱和门控技术结合到一个覆盖正则化器中来解决这种提示选择偏差,且不引入任何可学习参数。在六个长期预测基准和M4短期基准上,CASP-LLM在大多数数据集-预测视界设置下达到或超越了基于相似性的LLM预测器,仅在Electricity、M4-Monthly以及少样本长视界设置中例外。一项受控研究将失效模式定位于跨批次使用层面而非单次检索冗余层面:诸如MMR等检索内多样化方法并无帮助,而在训练过程中对锚点使用进行正则化则有效。
cs.LG / 103 / 2609.22984

LPINNs: First-Layer Gated Localization for Physics-Informed Neural Networks

LPINNs:面向物理信息神经网络的首层门控局部化方法
Chawla, Lakshay, Jain, Hardik
Abstract
Physics-informed neural networks (PINNs) use one shared representation over the computational domain, which can become difficult to optimize on long domains and for high-order operators. We study a minimal alternative: multiply the first hidden activation of an otherwise unchanged dense PINN by input-dependent localization functions, giving first-layer units receptive fields without partitioning the domain or adding interface losses. We screen 13 families of localization functions, in up to three parameterizations each, on a nonlinear harmonic oscillator (HO), a heat equation on a long spatial interval, and a manufactured four-dimensional (4D) fourth-order problem, with ten paired seeds throughout. Three configurations give large reductions in solution error at matched budgets: (i) Fixed Gaussian localization functions on the $2\pi$ HO domain cut mean solution RMSE from $4.8369\times10^{-1}$ to $8.83\times10^{-3}$ at 3k epochs. (ii) The inverse-quadratic family with learnable centers and widths cuts it from $2.896\times10^{-1}$ to $3.06\times10^{-2}$ on the $8\pi$ heat domain at 10k epochs. (iii) Fixed bump localization functions cut it from $1.75947\times10^{1}$ to $2.260\times10^{-1}$ on the $4\pi$ 4D domain at 10k epochs. Every paired seed improves in these three comparisons. The screen also shows that the mechanism is not a free win: on HO only 2 of 13 families beat the baseline, and 10 of the remaining 11 are 9 to 23 times worse; on 4D four families are non-finite and five are more than three orders of magnitude worse than the baseline. The inverse-quadratic family is the only one that beats the baseline on all three equations. Overall, these results show that first-layer localization can provide measurable improvements to baseline PINNs on long-domain and high-order problems.
Chinese Translation
物理信息神经网络(PINNs)在整个计算域上使用单一共享表示,这使得其在长区间域和高阶算子的情形下难以优化。我们研究了一种极简的替代方案:在保持其他结构不变的稠密 PINN 中,将第一层隐藏激活函数与依赖于输入的局部化函数相乘,从而赋予首层神经元感受野,而无需对计算域进行分区或添加界面损失。我们在非线性谐振子(HO)、长空间区间上的热传导方程,以及一个构造的四维(4D)四阶问题上,对 13 类局部化函数(每类最多包含三种参数化方式)进行了筛选,全部实验均使用十组配对随机种子。在相同计算预算下,三种配置取得了显著降低求解误差的效果:(i) 在 $2\pi$ 的谐振子域上采用固定高斯局部化函数,在 3k 个训练轮次后将平均求解 RMSE 从 $4.8369\times10^{-1}$ 降至 $8.83\times10^{-3}$。(ii) 具有可学习中心和宽度的反二次函数族,在 $8\pi$ 的热传导域上于 10k 个训练轮次后将其从 $2.896\times10^{-1}$ 降至 $3.06\times10^{-2}$。(iii) 固定 bump 局部化函数在 $4\pi$ 的 4D 域上于 10k 个训练轮次后将其从 $1.75947\times10^{1}$ 降至 $2.260\times10^{-1}$。在这三项比较中,每一组配对种子均有所改善。该筛选实验还表明这一机制并非普遍有效:在谐振子问题上,13 个函数族中仅有 2 个优于基线,其余 11 个中有 10 个比基线差 9 至 23 倍;在 4D 问题上,四个函数族结果不收敛(非有限值),五个函数族比基线差三个数量级以上。反二次函数族是唯一在全部三个方程上均优于基线的函数族。总体而言,这些结果表明,首层局部化能够在长域和高阶问题上为基线 PINNs 带来可测量的改进。
cs.LG / 104 / 2609.22990

On attention heads and bilinear forms

论注意力头与双线性形式
O'Desky, Andrew
Abstract
We study the symmetric and antisymmetric parts of bilinear forms in the attention heads of trained large language models. We introduce an orthogonally invariant profile map from real bilinear forms to a three-dimensional simplex and observe that profiles of trained bilinear forms accumulate near profiles of rank-one bilinear forms. We prove that the symmetric part of a bilinear form in an attention head is the sum of a hyperbolic form and a zero form for a Zariski-dense subset of query-key matrices.
Chinese Translation
我们研究了训练后的大型语言模型中注意力头内双线性形式的对称与反对称部分。我们引入了一个从实双线性形式到三维单纯形的正交不变轮廓映射(profile map),并观察到训练后双线性形式的轮廓聚集在一秩双线性形式的轮廓附近。我们证明,对于Zariski稠密的查询-键矩阵子集,注意力头中双线性形式的对称部分是一个双曲形式与一个零形式之和。
cs.LG / 105 / 2609.23008

Interpretable Multi-Hypersphere Deep Anomaly Detection for Open-set Supervised Anomaly Detection

面向开放集监督异常检测的可解释多超球面深度异常检测方法
Yang, Zhiji, Wang, Fangyong, Li, Yue, Pan, Xianli, Zhao, Jianhua
Abstract
Multi-class open-set anomaly detection requires a model to characterize the normal acceptance domain formed by multiple heterogeneous subdistributions using only class-labeled samples from known normal classes, and to identify previously unseen anomalies at test time. Existing single-hypersphere methods cannot explicitly represent class-specific locations and acceptance ranges, while current multi-hypersphere or multi-class approaches do not fully integrate inter-class boundary constraints, learnable acceptance ranges, and interpretable decisions. To address these limitations, we propose Interpretable Multi-Hypersphere Deep Anomaly Detection (IMHD-AD). IMHD-AD constructs an independent hypersphere for each known normal class in a shared feature space. With target-inside and non-target-outside constraints, IMHD-AD embeds the class-specific hypersphere centers and radii directly into the final network layer and jointly optimizes them with the shared representation. The minimum signed boundary score across hyperspheres simultaneously determines open-set acceptance or rejection and provides a faithful geometric explanation of each decision. On MNIST, Fashion-MNIST, and CIFAR-10, IMHD-AD achieves the highest AUC in 28 of 30 open-set comparisons. A two-dimensional synthetic study further shows that model architecture must balance the compactness of known normal classes against the separability of unknown anomalies.
Chinese Translation
多类开放集异常检测要求模型仅利用已知正常类别的带标签样本,刻画由多个异质子分布构成的正常接受域,并在测试阶段识别未曾见过的异常。现有的单超球面方法无法显式表示各类别特定的位置与接受范围,而当前的多超球面或多类方法也未充分整合类间边界约束、可学习的接受范围以及可解释的决策。为解决这些局限,我们提出了可解释多超球面深度异常检测方法(Interpretable Multi-Hypersphere Deep Anomaly Detection,IMHD-AD)。IMHD-AD在共享特征空间中为每个已知正常类别构建一个独立的超球面。借助“目标在内”与“非目标在外”的约束,IMHD-AD将各类别特定的超球面中心与半径直接嵌入网络最后一层,并与共享表示进行联合优化。所有超球面中最小的带符号边界得分同时决定开放集的接受或拒绝,并为每个决策提供忠实的几何解释。在MNIST、Fashion-MNIST和CIFAR-10数据集上,IMHD-AD在30组开放集对比中有28组取得了最高的AUC。二维合成实验进一步表明,模型架构必须在已知正常类别的紧凑性与未知异常的可分性之间取得平衡。
cs.LG / 106 / 2609.23033

WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models

波前解码:面向循环语言模型的并行化自推测解码
Ha, Hyeongju, Kim, Jae-Joon
Abstract
Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD), a training-free self-speculative decoding framework designed for looped language models. WFD exploits two properties of these architectures: intermediate recurrence outputs provide effective draft predictions, and weight sharing allows token states at different positions and recurrence depths to be processed in one batched recurrent-block call. WFD organizes these mixed-depth states into a diagonal wavefront, continuously drafting new positions at shallow depth while advancing earlier positions toward full-depth verification. Unlike the phase-separated draft-then-verify schedule, WFD therefore co-batches drafting and verification within the same recurrent calls, while rejected drafts are corrected using full-depth predictions. Across six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, consistently outperforming draft-then-verify. Cross-recurrence KV sharing further reduces wavefront KV traffic and increases WFD's speedup to 4.81x on Huginn-3.5B.
Chinese Translation
循环语言模型通过重复应用权重共享的块来增加有效深度而不增加参数量,但每生成一个令牌所需的T次顺序循环块调用显著增加了解码延迟。为解决这一问题,我们提出了波前解码(Wavefront Decoding, WFD),一个专为循环语言模型设计的免训练自推测解码框架。WFD利用了此类架构的两个特性:中间循环输出可提供有效的草稿预测;权重共享使得不同位置和不同循环深度的令牌状态能够在一次批量化的循环块调用中一并处理。WFD将这些混合深度的状态组织成对角波前,在浅层持续为新的位置生成草稿,同时推进较早的位置向全深度验证演进。与分阶段的“先草稿后验证”调度不同,WFD在相同的循环调用中协同批量处理草稿生成与验证,而被拒绝的草稿则使用全深度预测进行修正。在Spec-Bench的六个任务类别上,WFD在Ouro-2.6B上相对自回归解码实现了2.42倍加速,在Huginn-3.5B上实现了3.54倍加速,并持续优于先草稿后验证方法。跨循环KV共享进一步降低了波前KV流量,使WFD在Huginn-3.5B上的加速比提升至4.81倍。
cs.LG / 107 / 2609.23055

Optimizers for Diffusion Models: A Controlled Benchmark

扩散模型的优化器:一项受控基准研究
Bolatov, Arman, Shulgin, Egor, Li, David, Shtanchaev, Abduragim, Stich, Sebastian U., Panov, Maxim, Moulines, Eric, Richtárik, Peter, Takáč, Martin
Abstract
Discrete diffusion models now match autoregressive language models on several benchmarks, while the question of how best to train them has received far less attention: the optimizer is inherited from one paper to the next and never compared. New optimizers, meanwhile, are validated almost exclusively on autoregressive pretraining, a different objective on a different loss surface. We present a controlled optimizer benchmark across four diffusion formulations, to our knowledge the first for discrete diffusion: seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS-M, Schedule-Free) on masked diffusion (text8), uniform diffusion (QM9, and LM1B through the Gaussian duality) and Gaussian diffusion on images (CelebA-64), each on a task with published reference values. Every optimizer receives the same search protocol, and every winner is retrained at the full budget with three seeds. AdamW is a strong default but not always the right choice: it is beaten by a resolved margin on two of the four tasks, and the winner changes with the formulation, so the optimizer deserves the same care as the rest of the training recipe. Notably, methods validated on autoregressive language model pretraining transfer well: Muon, MARS-M and SOAP each beat the tuned AdamW on at least one diffusion formulation. The benchmark, all runs and every figure are reproducible end to end from the released code at https://github.com/armanbolatov/diffusion-baselines.
Chinese Translation
离散扩散模型如今已在多个基准任务上可与自回归语言模型相媲美,然而如何最优地训练它们这一问题却远未受到重视:优化器往往只是从一篇论文沿用到下一篇论文,从未被比较过。与此同时,新的优化器几乎只在自回归预训练上进行验证,而这属于不同损失曲面上的不同目标函数。我们提出了一项涵盖四种扩散形式的受控优化器基准测试,据我们所知,这是首个针对离散扩散的此类基准:包括七种优化器(AdamW、Lion、Muon、SOAP、MARS、MARS-M、Schedule-Free),应用于掩码扩散(text8)、均匀扩散(QM9,以及通过高斯对偶的LM1B)和图像上的高斯扩散(CelebA-64),每一项任务均具有已发表的参考值。每种优化器均采用相同的搜索协议,且每个获胜者都以完整预算使用三个随机种子重新训练。结果表明,AdamW是一个强大的默认选择,但并非总是正确的选择:它在四个任务中的两个上以确定的差距被击败,且获胜者随扩散形式而变化,因此优化器值得与训练方案的其他部分同样受到重视。值得注意的是,在自回归语言模型预训练上验证过的方法能够很好地迁移:Muon、MARS-M和SOAP各自在至少一种扩散形式上击败了经过调优的AdamW。该基准测试、所有运行结果及每张图表均可通过发布在 https://github.com/armanbolatov/diffusion-baselines 的代码端到端地完整复现。
cs.LG / 108 / 2609.23073

MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs

MolSC:利用取代基贡献增强大语言模型的细粒度分子理解能力
Park, Hyuntae, Kim, Sooyeon, Park, Jiwon, Lee, SangKeun
Abstract
Recent advances in natural language processing have led to molecular Large Language Models (LLMs) with strong performance across diverse chemistry tasks. However, they still struggle to capture fine-grained structure-property relationships, particularly how small, localized modifications alter a molecule's behavior. To address this limitation, we introduce MolSC, a dataset of substituent contributions, defined as property changes induced by attaching specific substituents to molecular scaffolds. Curated from manually annotated bioactivity records, MolSC spans structural-alert liability, target-specific bioactivity, and physicochemical descriptors, and contains 181K substituent-level examples for training. We further propose MolSC-Bench, a held-out evaluation benchmark of 1,541 examples disjoint from MolSC at the scaffold, substituent, and molecule levels. Our experiments show that existing molecular LLMs and strong proprietary models such as GPT-5.2 and Gemini-3-Flash show limited reliability in substituent contribution prediction. In contrast, training on MolSC substantially improves this ability and achieves strong performance across diverse downstream molecular tasks. These results highlight substituent contribution learning as a key component of fine-grained molecular understanding.
Chinese Translation
自然语言处理的最新进展催生了在多种化学任务上表现出色的分子大语言模型(LLMs)。然而,这些模型仍难以捕捉细粒度的构效关系,尤其是微小的局部修饰如何改变分子行为这一方面。为解决这一局限,我们提出了MolSC,一个取代基贡献数据集,其取代基贡献被定义为将特定取代基连接到分子骨架上所引起的性质变化。MolSC从人工标注的生物活性记录中整理而来,涵盖结构性警示风险、靶点特异性生物活性以及理化描述符,包含181K个取代基级别的训练样本。我们进一步提出了MolSC-Bench,一个包含1,541个样本的保留评估基准,其在骨架、取代基和分子三个层面均与MolSC不相交。实验表明,现有的分子大语言模型以及GPT-5.2和Gemini-3-Flash等强大的专有模型在取代基贡献预测任务上的可靠性有限。相比之下,在MolSC上进行训练可显著提升该能力,并在多种下游分子任务上取得优异表现。这些结果凸显了取代基贡献学习作为细粒度分子理解关键组成部分的重要性。
cs.LG / 109 / 2609.23084

AirGC-CD: Gaussian-Circulant Precoding for Exactly Debiasable PAPR Reduction in Over-the-Air Federated Learning

AirGC-CD:面向空中联邦学习中可精确去偏PAPR降低的高斯-循环预编码
Jang, Jonggyu, Lyu, Hyeonsu, Yang, Hyun Jong
Abstract
Over-the-air federated learning lets edge devices transmit their local updates simultaneously, reducing the communication overhead. The resulting waveform, however, has a peak-to-average power ratio (PAPR) that grows with the model dimension, and keeping the amplifier in its linear range leaves two remedies: clipping the peaks or backing off the transmit power. Neither remedy is without cost: i) the clipping distortion appears at the receiver as a bias that cannot be removed, and ii) back-off keeps the signal intact but degrades the average signal-to-noise ratio (SNR). Independent of this trade-off, the transmission remains uncompressed, spending one channel use per model parameter, which keeps large-model training out of reach. To address these challenges, we propose AirGC-CD, an over-the-air scheme that precodes each local update with a partial Gaussian circulant matrix before clipping. In AirGC-CD, the precoder's output is exactly Gaussian regardless of the update's sparsity, so the clipping function is designed for a known distribution instead of inheriting it from the data. This enables the clipping to be inverted on average by a single scalar Bussgang gain in closed form, and we prove that the resulting aggregate is exactly unbiased, with clipping adding only variance. The clipping ratio is then the only free parameter left, trading the variance of the clipping against the SNR loss from back-off, and we derive its near-optimum in closed form. Since the precoder is linear, it also acts as a compressor, reducing the transmission from the model dimension d to the sketch dimension m at a cost of only O(dlog d) via two fast Fourier transforms, whereas a Gaussian sketch costs O(md). Experiments on five image datasets show that AirGC-CD outperforms baseline over-the-air FL schemes in most settings, particularly at low SNR, while using fewer channel uses per round.
Chinese Translation
空中联邦学习(Over-the-air federated learning)允许边缘设备同时传输其本地更新,从而降低通信开销。然而,由此产生的波形其峰均功率比(PAPR)会随模型维度增长,为使放大器保持在线性范围内,仅有两种补救措施:对峰值进行削波(clipping)或回退发射功率。两种措施均有代价:i)削波失真在接收端表现为无法消除的偏差;ii)功率回退虽能保持信号完整,却会降低平均信噪比(SNR)。与这一权衡无关的是,传输本身仍处于未压缩状态,每个模型参数需要占用一次信道使用,使得大模型训练难以实现。为应对这些挑战,我们提出AirGC-CD,这是一种空中传输方案,在削波之前用一个部分高斯循环矩阵对每个本地更新进行预编码。在AirGC-CD中,无论更新具有何种稀疏性,预编码器的输出都严格服从高斯分布,因此削波函数可以针对已知分布进行设计,而无需从数据中继承分布特性。这使得削波可通过单个以闭式表示的标量Bussgang增益在平均意义上被精确逆转,并且我们证明了所得聚合结果严格无偏,削波仅引入方差。于是,削波比率成为唯一可自由调节的参数,用于权衡削波方差与功率回退带来的SNR损失,我们以闭式形式推导出其近最优值。由于该预编码器是线性的,它同时起到压缩器的作用,将传输维度从模型维度d降至草稿(sketch)维度m,且通过两次快速傅里叶变换仅需O(d log d)的计算开销,而高斯草稿的开销为O(md)。在五个图像数据集上的实验表明,AirGC-CD在大多数设置下优于基线空中联邦学习方案,尤其是在低SNR条件下,同时每轮使用的信道次数更少。
cs.LG / 110 / 2609.23087

Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone

神经谱容量:仅凭网络规格度量与设计架构
Zhu, Chenyu, Zhao, Ruoyu, Lu, Zhichao
Abstract
Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for these decisions, #Params and #FLOPs, capture size and compute but not architectural structure: two architectures with identical parameter budgets but different depth-width, head, or FFN allocations receive identical scores yet behave differently. We propose Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix. Under standard random initialization, the Marchenko-Pastur law renders NSC computable from the architectural specification alone, with no model instantiation, data, or gradients. Its layer-wise additive structure admits NSC-DP, an exact dynamic-programming solver returning the architecture globally maximizing NSC under resource constraints in seconds on a CPU -- a guarantee that black-box search over existing training-free proxies cannot provide. Empirically, NSC outperforms #Params, #FLOPs, and representative training-free proxies in ranking across seven Transformer and CNN families (on FlexiBERT, $\tau = 0.505$ on pairs differing in #Params by less than 10%, where #Params collapses to 0.082); NSC-DP discovers a Transformer-XL architecture on WikiText-103 that beats the human-designed baseline in 2 seconds; and prunes LLaMA-7B to the best 5.7B model across eight commonsense reasoning tasks without any calibration data, about 5900x faster than the strongest training-free proxy baseline.
Chinese Translation
现代Transformer的设计与压缩均可归结为在预算约束下进行容量分配。用于此类决策的标准指标——参数量(#Params)与浮点运算量(#FLOPs)——仅能刻画规模与计算量,而无法反映架构结构:两个参数预算相同但深度-宽度、注意力头数或FFN配置不同的架构会得到相同的评分,实际表现却各异。我们提出神经谱容量(Neural Spectral Capacity,NSC),这是一个基于各权重矩阵奇异值谱的闭式标量指标。在标准随机初始化下,Marchenko-Pastur定律使得NSC可仅凭架构规格计算得出,无需模型实例化、数据或梯度。其逐层可加的结构催生了NSC-DP——一种精确的动态规划求解器,可在CPU上于数秒内返回在资源约束下全局最大化NSC的架构——这是对现有免训练代理指标进行黑盒搜索所无法提供的保证。实验表明,在七个Transformer与CNN家族的排序任务中,NSC优于#Params、#FLOPs以及代表性的免训练代理指标(在FlexiBERT上,对于参数量差异小于10%的架构对,NSC的τ=0.505,而#Params仅为0.082);NSC-DP在2秒内发现了一个在WikiText-103上超越人工设计基线的Transformer-XL架构;并且在不使用任何校准数据的情况下,将LLaMA-7B剪枝为在八项常识推理任务上表现最佳的5.7B模型,速度比最强的免训练代理基线快约5900倍。
cs.LG / 111 / 2609.23092

The Role of Coordinates in Pareto Regret for Adversarial Multi-Objective Bandits

坐标在对抗性多目标老虎机的帕累托遗憾中的作用
Guan, Changkun, Xu, Mengfan
Abstract
Adversarial multi-objective bandits hold the potential to help us optimize choices (arms) whose reward is a multidimensional vector chosen by an adversary and whose performance is measured by Pareto regret. We define loss as one minus reward and measure the easiness of a coordinate by the smallest cumulative loss of the arms on it, and call the coordinate easier when this quantity is smaller. Existing work suggests that in theory an easier coordinate may reduce Pareto regret. However, in practice, one may not know which coordinate is easier. On the negative side, we show that this lack of information eliminates the possibility: a smaller cumulative loss does not improve the worst-case order of Pareto regret. Precisely, let \(L_d\) be the smallest cumulative loss along coordinate $d$ over $T$ rounds. For \(K\ge4\) arms, \(T\ge6\) rounds, and at least 2 coordinates, we prove that the minimax expected Pareto regret is \(\Omega(\min\{T-L_0,\sqrt{K(T-L_0)}\})\). It is monotonically decreasing in \(L_0\), even when \(L_0=\min_d L_d\) itself is known. On the positive side, this result motivates the possibility that other coordinates, not just the easy one, may suffice to attain the optimal rate of Pareto regret. When $L_0$ is known, we apply Poly-INF to a fixed coordinate and obtain an upper bound on Pareto regret that exhibits the same order and thus matches the lower bound. Without such knowledge, we develop a reward-doubling version of Poly-INF that adapts to this unknown quantity while still attaining the matching minimax rate. Another implication is that it has no extra \(\log T\) factor and is independent of the number of coordinates.
Chinese Translation
对抗性多目标老虎机(adversarial multi-objective bandits)有望帮助我们优化奖励为对手所选多维向量的选择(臂),其性能以帕累托遗憾(Pareto regret)来衡量。我们将损失定义为一减去奖励,并通过各臂在某一坐标上的最小累积损失来衡量该坐标的难易程度,该值越小则称该坐标越容易。已有工作表明,理论上更容易的坐标可能降低帕累托遗憾。然而在实践中,人们可能并不知道哪个坐标更容易。从负面结果来看,我们证明这种信息缺失消除了这种可能性:较小的累积损失并不能改善帕累托遗憾的最坏情形阶数。确切地说,设 \(L_d\) 为坐标 $d$ 在 $T$ 轮上的最小累积损失。对于 \(K\ge4\) 个臂、\(T\ge6\) 轮以及至少 2 个坐标的情形,我们证明极小极大期望帕累托遗憾为 \(\Omega(\min\{T-L_0,\sqrt{K(T-L_0)}\})\)。该界关于 \(L_0\) 单调递减,即使 \(L_0=\min_d L_d\) 本身已知也是如此。从正面结果来看,这一结果启发了一种可能性:其他坐标,而不仅仅是最容易的那个坐标,也可能足以达到帕累托遗憾的最优速率。当 \(L_0\) 已知时,我们将 Poly-INF 应用于某个固定坐标,得到了阶数相同从而与下界匹配的帕累托遗憾上界。在缺乏这一知识的情况下,我们开发了 Poly-INF 的奖励加倍(reward-doubling)版本,它能自适应于这一未知量,同时仍达到匹配的极小极大速率。另一个含义是,该算法不含额外的 \(\log T\) 因子,且与坐标数量无关。
cs.LG / 112 / 2609.23102

When Does Adversarial Refinement Help? A Negative Result and Open Problem in Adapting R3GAN to Time Series Imputation

对抗式精化何时有效?R3GAN适配时间序列插补的负面结果与开放问题
He, Yufeng
Abstract
Diffusion models and transformers have supplanted GANs for multivariate time series imputation, largely on grounds of GAN training instability. R3GAN (NeurIPS 2024) removes that instability via regularized relativistic losses with provable convergence, raising a natural question: do stable, modern GANs revive adversarial imputation? We adapt R3GAN to 1D temporal data with a coarse-to-fine refinement framework and a frequency-domain discriminator, and audit 14 saved configurations across 3 datasets. Because these are heterogeneous single runs, the evidence is descriptive rather than a matched causal ablation. We report a negative result. All five saved mean/zero-start configurations improve by 48.4-70.2%. Among eight eligible non-legacy linear-start configurations, the mean change is -0.7% (range -3.0% to +1.1%); a separate -21.9% legacy logging anomaly is retained for provenance but excluded from that aggregate. In a saved Weather comparison, standalone R3GAN-1D underperforms BRITS by 5.8x. Crucially, we argue the common explanation (that GANs optimize distributional rather than point-wise objectives) cannot be the whole story, since diffusion models also optimize distributional objectives yet achieve state-of-the-art imputation. Our saved reconstruction-weight sweep is consistent with the adversarial signal being inert or harmful, but cannot identify its causal contribution; a matched discriminator-removed ablation is the key next experiment. We frame the precise reason a learned discriminator fails to provide useful refinement gradients (where a learned diffusion denoiser succeeds) as an open problem, and offer practical guidance on when adversarial refinement is worthwhile.
Chinese Translation
扩散模型与Transformer已在很大程度上取代GAN用于多变量时间序列插补,其主要理由是GAN训练的不稳定性。R3GAN(NeurIPS 2024)通过带可证明收敛性的正则化相对论损失消除了这种不稳定性,由此引出一个自然的问题:稳定的现代GAN能否复兴基于对抗训练的插补方法?我们将R3GAN适配到一维时间数据,采用由粗到精的精化框架和频域判别器,并在3个数据集上审计了14个已保存的配置。由于这些是异质的单次运行结果,因此证据是描述性的,而非匹配的因果消融实验。我们报告了一个负面结果:全部五个已保存的均值/零起点配置提升了48.4%至70.2%;在八个符合条件的非遗留线性起点配置中,平均变化为-0.7%(范围为-3.0%至+1.1%);另有-21.9%的遗留日志异常出于溯源目的予以保留,但未计入上述汇总。在一个已保存的Weather数据集对比中,独立的R3GAN-1D比BRITS落后5.8倍。至关重要的是,我们论证常见的解释(即GAN优化的是分布层面而非逐点目标)并不能说明全部问题,因为扩散模型同样优化分布层面的目标,却实现了最先进的插补效果。我们已保存的重构权重扫描结果与对抗信号无效或有害的结论一致,但无法确定其因果贡献;匹配的移除判别器消融实验是下一步的关键实验。我们将“学习到的判别器为何无法提供有效的精化梯度(而学习到的扩散去噪器却可以)”这一精确原因界定为一个开放问题,并就何时值得采用对抗式精化提供了实用建议。
cs.LG / 113 / 2609.23117

Whitening Inverts the Hierarchy: What the Norm of a Whitened Embedding Measures

白化逆转层级结构:白化嵌入的范数究竟度量了什么
Ahnouch, Mohammed, Elaachak, Lotfi
Abstract
Whitening a foundation-model embedding and using its squared norm as a training-free likelihood surrogate is motivated by the observation that whitened coordinates often appear approximately standard normal. We show that this observation follows from the projection central limit theorem and therefore does not imply a Gaussian joint distribution. Across multiple encoders and three training objectives, we find systematic over-dispersion of the whitened radius relative to the Gaussian reference, including against distributional clones with identical mean and covariance. We further show that the commonly reported agreement between empirical and theoretical norm statistics is an algebraic consequence of in-sample whitening and does not constitute evidence for Gaussianity. We identify the mechanism behind this behavior: whitening reverses the encoder's spectral hierarchy, shifting the contribution to the squared norm toward near-degenerate directions that encode predominantly noise. In these directions, the dominant variability is governed by a single input-dependent scale. We estimate this scale from two moments and use it to predict, without additional free parameters, the cross-dependence between disjoint spectral halves. These results indicate that the squared whitened norm is better interpreted as a Mahalanobis measure of semantic atypicality than as a log-likelihood. This interpretation explains both its practical effectiveness and its calibration failures: the statistic can rank and detect atypical samples consistently with nonparametric density estimates and across encoders trained with different objectives, while Gaussian tail thresholds can be inaccurate by orders of magnitude. etc.
Chinese Translation
对基础模型的嵌入进行白化,并将其范数的平方作为免训练的似然替代量,其动机在于观察到白化后的坐标往往近似服从标准正态分布。我们证明这一观察可由投影中心极限定理导出,因此并不意味着服从联合高斯分布。在多个编码器和三种训练目标下,我们发现白化半径相对于高斯参考存在系统性的过度分散,即使与均值和协方差完全相同的分布克隆相比也是如此。我们进一步表明,文献中常报告的经验范数统计量与理论范数统计量之间的一致性,是样本内白化的代数必然结果,并不构成高斯性的证据。我们揭示了这一行为背后的机制:白化逆转了编码器的谱层级结构,使对范数平方的贡献转移至近似简并的方向,而这些方向主要编码的是噪声。在这些方向上,主导性的变异性由一个单一的、依赖于输入的尺度所控制。我们从两个矩估计该尺度,并在不引入额外自由参数的情况下,用它预测了两个不相交谱半部分之间的交叉依赖关系。这些结果表明,白化范数的平方更宜被解释为语义非典型性的马氏(Mahalanobis)度量,而非对数似然。这一解释既说明了其实际有效性,也说明了其校准失败之处:该统计量能够与非参数密度估计相一致地、并在不同目标训练的编码器之间一致地对非典型样本进行排序和检测,而高斯尾部阈值则可能产生数量级上的偏差。等等。
cs.LG / 114 / 2609.23125

Perplexity Cost Understates What Activation Quantisation Breaks

困惑度代价低估了激活量化所破坏的能力
Sathyanarayanan, Anish
Abstract
Activation quantisation is usually evaluated with an aggregate metric, perplexity, averaged over every token a model predicts. We ask whether that average identifies which computations a quantiser damages. Perplexity turns out to be a reliable aggregate signal: across 12 models from four families and 780 within-model comparisons, the arm perplexity prefers also retains more induction and more retrieval in all but 2.1 and 4.0 percent of cases respectively. But where perplexity has risen by only a factor of 1.2 to 1.5, induction still keeps 0.959 of its intact accuracy while retrieval has already fallen to 0.554, a gap the aggregate number does not surface. This gap has structure, not just size: a matched Gaussian-noise control of the same per-channel magnitude leaves it largely intact, and randomising only the sign of the quantisation error, every magnitude held fixed, is nearly as harmless, so magnitude alone does not explain the damage. Quantising in a rotated basis, which changes coordinate alignment without changing error magnitude, restores induction from 0.001 to 0.980 at three average bits per token in a single-block intervention, though retrieval recovers less completely at the same setting (0.694); end-to-end at four average bits, induction reaches 0.968 and retrieval 0.534. The pattern holds on two further models up to 32B parameters and, in the deployed configurations we tested, under AWQ once activations are pushed to 4 bits. A perplexity target bounds the average cost of a transformation applied to the activation; it does not, by itself, show which computations survived.
Chinese Translation
激活量化通常使用一个聚合指标——困惑度(perplexity)来评估,该指标是在模型预测的所有词元上取平均。我们要问的是:这种平均值能否识别出量化器究竟破坏了哪些计算?结果表明,困惑度是一个可靠的聚合信号:在来自四个模型家族的12个模型、共780次模型内比较中,困惑度所偏好的方案在除了2.1%(归纳能力)和4.0%(检索能力)的情况之外,都保留了更多的归纳(induction)能力和检索(retrieval)能力。然而,当困惑度仅上升1.2至1.5倍时,归纳能力仍保持其原始精度的0.959,而检索能力已降至0.554——这一差距是聚合数字所无法呈现的。这种差距不仅有结构,而不仅仅是程度问题:在相同的每通道幅度下进行的匹配高斯噪声对照实验中,该差距基本保持不变;而仅随机化量化误差的符号、保持所有幅度不变,也几乎无害,因此仅凭幅度并不能解释这种损伤。在旋转基下进行量化,只改变坐标对齐而不改变误差幅度,在单块干预、每词元平均3比特的设置下,可将归纳能力从0.001恢复到0.980,但检索能力在相同设置下恢复得较不完全(0.694);在每词元平均4比特的端到端设置下,归纳能力达到0.968,检索能力达到0.534。该模式在另外两个最高达320亿参数的模型上依然成立,并且在我们测试的部署配置中,当激活被压至4比特时,在AWQ下同样成立。困惑度目标只能界定施加于激活的某种变换的平均代价;它本身并不能表明哪些计算得以幸存。
cs.LG / 115 / 2609.23127

Provably Efficient Reinforcement Learning in Continuous-Time Episodic MDPs with Poisson Decision Epochs

具有泊松决策时刻的连续时间片段式马尔可夫决策过程中可证明高效的强化学习
Guo, Kenny, Iverson, Valentio, Wijetunga, Sahan, Chang, William
Abstract
Many real-world reinforcement learning (RL) problems evolve in continuous time, where decisions occur at irregular, event-driven intervals rather than at fixed discrete steps. We study episodic continuous-time Markov Decision Processes (MDPs) in which decision epochs are governed by a homogeneous Poisson process and the reward and transition dynamics vary smoothly over time. We consider both a fixed number of jumps per episode and a fixed time budget with a random number of Poisson decision epochs. Under a Lipschitz continuity assumption in time, we exploit local smoothness through discretization and extend both UCRL (Auer and Ortner 2006) and Q-learning (Jin et al. 2018) to this setting, proving $\widetilde{O}(T^{2/3})$ regret bounds for both model-based and model-free algorithms. Finally, we establish matching $\widetilde{\Omega}(T^{2/3})$ minimax lower bounds, showing that the rate is optimal up to logarithmic factors. These results provide the first tight regret guarantees for Lipschitz-smooth continuous-time episodic MDPs with Poisson decision epochs.
Chinese Translation
许多现实世界的强化学习(RL)问题在连续时间中演化,其中决策发生在不规则的、由事件驱动的时刻,而非固定的离散步骤上。我们研究了片段式连续时间马尔可夫决策过程(MDP),其中决策时刻由齐次泊松过程决定,且奖励和转移动力学随时间平滑变化。我们考虑了每片段固定跳跃次数以及固定时间预算下随机数量泊松决策时刻两种情形。在时间上的Lipschitz连续性假设下,我们通过离散化利用局部平滑性,将UCRL(Auer and Ortner 2006)和Q-learning(Jin et al. 2018)扩展到该设定,并证明了基于模型和无模型算法的 $\widetilde{O}(T^{2/3})$ 遗憾界。最后,我们建立了匹配的 $\widetilde{\Omega}(T^{2/3})$ 极小极大下界,表明该速率在对数因子意义下是最优的。这些结果为具有泊松决策时刻的Lipschitz平滑连续时间片段式MDP提供了首个紧致的遗憾保证。
cs.LG / 116 / 2609.23146

Ask for Any Appliance: A Prompt-Programmable Foundation Model for Non-Intrusive Load Monitoring

按需查询任意电器:面向非侵入式负荷监测的提示可编程基础模型
Wang, Xudong, Cui, Jiacheng, Xue, Junyu, Li, Tongxin, Tang, Guoming
Abstract
Non-intrusive load monitoring (NILM) estimates appliance-level consumption from a whole-home meter, but appliance-specific models and fixed output inventories make coverage costly to extend. We present FM4NILM (Foundation Model for NILM), a single prompt-programmable model that estimates a requested appliance's power trajectory from aggregate measurements, a natural-language description, and optional activation exemplars. A lightweight cadence-aware transformer is pretrained by masked reconstruction on 645k sequences from seven public corpora spanning 1-60 s sampling intervals, then aligned with appliance requests using observation-masked losses for partially labeled households. A Bernoulli-lognormal decoder separates activity detection from conditional power estimation. On held-out households and time periods from REDD, UK-DALE, and REFIT, one frozen text-prompted model serves twelve appliance-corpus requests, achieving 0.556 event F1, 0.625 AUPRC, and the lowest active-window MAE (251.8 W) among seven appliance-specific baselines. Streaming score aggregation raises event F1 to 0.582 with a 60 s aggregation delay. In a separate category-held-out evaluation, adding ten activation exemplars raises microwave AUPRC from 0.132 to 0.214 without parameter updates. Input-intervention ablations probe the model's dependence on appliance requests and aggregate measurements. These results demonstrate competitive disaggregation with one shared model and support extending appliance coverage through prompts and examples rather than additional specialist networks.
Chinese Translation
非侵入式负荷监测(NILM)从全屋电表中估计电器级能耗,但针对特定电器的模型和固定的输出清单使得扩展覆盖范围的成本高昂。我们提出了FM4NILM(NILM基础模型),这是一个单一的提示可编程模型,能够根据总功率测量数据、自然语言描述以及可选的启停示例,估计所请求电器的功率轨迹。该模型采用轻量级的节奏感知Transformer,在来自七个公开数据集、采样间隔为1–60秒的64.5万条序列上通过掩码重构进行预训练,随后利用针对部分标注家庭的观测掩码损失与电器请求进行对齐。伯努利-对数正态解码器将活动检测与条件功率估计相分离。在REDD、UK-DALE和REFIT的留存家庭与时间段上,单个冻结的文本提示模型即可服务十二个电器-数据集请求,取得了0.556的事件F1、0.625的AUPRC,并在七个电器专用基线中取得了最低的活跃窗口MAE(251.8瓦)。流式评分聚合在60秒聚合延迟下将事件F1提升至0.582。在另一项类别留出评估中,添加十个启停示例即可将微波炉的AUPRC从0.132提升至0.214,且无需参数更新。输入干预消融实验探究了模型对电器请求和总功率测量数据的依赖程度。这些结果表明,单一共享模型即可实现具有竞争力的负荷分解,并支持通过提示和示例而非额外的专用网络来扩展电器覆盖范围。
cs.LG / 117 / 2609.23183

K-TRAIL: Simulator-Guided Generative Design of EM/RF Circuits

K-TRAIL:仿真器引导的电磁/射频电路生成式设计
Saha, Piyush, Newell, Evan, O'Leary, Hanna, Natarajan, Arun, Aghasi, Alireza
Abstract
Inverse design of RF and electromagnetic (EM) circuits is challenging because the relationship between circuit layout and electrical response is non-unique, and full-wave simulation is computationally expensive. This paper presents K-TRAIL, a simulator-guided generative framework for automated EM/RF circuit synthesis. K-TRAIL combines diffusion-based layout generation with derivative-free ensemble Kalman guidance, allowing feedback from a black-box EM simulator to refine candidate layouts during generation without requiring adjoint sensitivities or differentiable solver models. The framework supports both synthesis from prescribed S-parameter responses and synthesis directly from RF performance constraints. Experiments on multi-layer RFIC structures show that simulator-guided generation improves agreement with target responses and can identify structurally distinct layouts that satisfy circuit-level design requirements. The proposed approach provides a practical path toward generative, verification-aware RF circuit design while retaining the flexibility to explore diverse layout topologies.
Chinese Translation
射频与电磁(EM)电路的逆向设计极具挑战性,其原因在于电路版图与电气响应之间的关系并非一一对应,且全波仿真的计算代价高昂。本文提出了K-TRAIL,一种用于自动化电磁/射频电路综合的仿真器引导生成式框架。K-TRAIL将基于扩散模型的版图生成与无导数集合卡尔曼引导相结合,使得黑盒电磁仿真器的反馈能够在生成过程中对候选版图进行优化,而无需伴随灵敏度或可微分求解器模型。该框架既支持从给定的S参数响应进行综合,也支持直接从射频性能约束进行综合。在多层射频集成电路(RFIC)结构上的实验表明,仿真器引导的生成方法提升了对目标响应的吻合度,并且能够找到在结构上各不相同、却满足电路级设计要求的版图。所提出的方法为生成式、具备验证意识的射频电路设计提供了一条切实可行的路径,同时保留了探索多样化版图拓扑的灵活性。
cs.LG / 118 / 2609.23185

Neural Residual Modeling for Scientific Data Compression under Guaranteed Error Bounds

保证误差界下面向科学数据压缩的神经残差建模
Majumder, Surya, Zhu, Liangji, Ranka, Sanjay, Rangarajan, Anand
Abstract
Lossy compression of scientific simulation data increasingly relies on learned, latent-space architectures such as Residual Vector Quantization (RVQ), which iteratively quantize a base representation and its residuals to progressively reduce reconstruction error. While effective, RVQ performs this residual modeling entirely in latent space, leaving the pixel-space error structure of the reconstruction largely unaddressed. In this work, we propose a post-processing pipeline that augments an RVQ-based compressor with a U-Net trained to predict and correct pixel-space residuals between the original volume and its RVQ reconstruction. We show that these residuals are spatially structured rather than driven by local intensity or gradient features, motivating the need for a deep spatial model rather than simple statistical correction. The U-Net-corrected reconstruction is then passed through a Guaranteed Autoencoder (GAE) stage, which projects the remaining residual onto a per-block PCA basis to enforce a user-specified block-wise error bound. To the best of our knowledge, this is the first pipeline to combine latent-space RVQ, explicit pixel-space residual correction via a deep spatial post-processing network, and GAE-based error-bound guarantees within a single framework for scientific data compression. We evaluate our approach on S3D, JHTDB and E3SM datasets, demonstrating consistent improvements in NRMSE, compression ratio] over RVQ-only and standard residual-correction baselines, while maintaining strict error guarantees required for scientific data fidelity.
Chinese Translation
科学模拟数据的有损压缩日益依赖于残差向量量化(Residual Vector Quantization, RVQ)等基于学习的隐空间架构,该类方法通过迭代量化基础表示及其残差来逐步降低重建误差。尽管有效,RVQ完全在隐空间中进行残差建模,而重建结果的像素空间误差结构在很大程度上未被处理。在本工作中,我们提出一种后处理流水线,通过训练一个U-Net来预测并校正原始体数据与RVQ重建结果之间的像素空间残差,从而增强基于RVQ的压缩器。我们证明这些残差呈现出空间结构化特征,而非由局部强度或梯度特征驱动,这表明需要深度的空间模型而非简单的统计校正。随后,经U-Net校正的重建结果通过保证自编码器(Guaranteed Autoencoder, GAE)阶段,将剩余残差投影到逐块的PCA基上,以强制满足用户指定的逐块误差界。据我们所知,这是首个在单一框架内结合隐空间RVQ、通过深度空间后处理网络实现的显式像素空间残差校正、以及基于GAE的误差界保证的科学数据压缩流水线。我们在S3D、JHTDB和E3SM数据集上评估了该方法,结果表明与仅使用RVQ及标准残差校正基线相比,该方法在NRMSE和压缩比方面取得了一致的改进,同时保持了科学数据保真度所要求的严格误差保证。
cs.LG / 119 / 2609.23215

Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents

基于LLM的主动推理智能体可解释性失效的触发因素与诊断方法
Raval, Param, Shenoy, Rohit, Vaidheeswaran, Archana
Abstract
LLM explainers are increasingly attached to autonomous agents as runtime oversight, with operators reading a generated account of the agent's beliefs and actions rather than its internal state. We audit the account itself, pairing an Active Inference (AIF) agent that tracks German grid demand and adjusts generation with an LLM explainer on three backends (GPT-4o, Claude-3-Opus, Gemini), and probing the pair with three black-box triggers. Corrupting the observation stream by 600 MW per step moves the agent's posterior by 490 MW, roughly 0.9% of grid capacity. None of the 30 explanations produced during the injection flag anything under a stated rubric, and each narrates the corrupted belief fluently. On timesteps where the agent takes an objectively wrong action, all three explainers produce a sycophantic rationalization 80-95% of the time (n = 20 per backend). Attacker-controlled text in the observation metadata field steers the explainer, with susceptibility differing by provider and data exfiltration succeeding on all three. We propose mitigations for each failure but do not evaluate them. In every failure we observed, the explanation was fluent and wrong. Moreover, nothing in the explainer architecture checks whether an explanation is true before an operator acts on it. Testing the explainer therefore belongs in any audit of an agentic deployment.
Chinese Translation
LLM解释器正越来越多地被附加到自主智能体上作为运行时监督手段,操作员阅读的是智能体信念与行为的生成性叙述,而非其内部状态。我们对这一叙述本身进行了审计:将一个跟踪德国电网需求并调整发电的主动推理(Active Inference, AIF)智能体与运行于三个后端(GPT-4o、Claude-3-Opus、Gemini)的LLM解释器配对,并用三种黑盒触发手段对二者进行探测。以每步600 MW的幅度破坏观测流会使智能体的后验信念偏移490 MW,约为电网容量的0.9%。在注入过程中产生的30份解释中,没有任何一份在既定评分标准下标记出异常,且每份解释都流畅地叙述了被破坏后的信念。在智能体采取客观错误动作的时间步上,三个解释器在80-95%的情况下都会产生谄媚式的合理化解释(每个后端n=20)。攻击者在观测元数据字段中控制的文本能够操纵解释器,其易感性因提供商而异,且数据外泄在全部三个后端上均成功实现。我们针对每种失效提出了缓解措施,但未对其进行评估。在我们观察到的每一次失效中,解释都流畅而错误。此外,解释器架构中没有任何机制在操作员依据解释采取行动之前验证其真实性。因此,对解释器的测试应纳入任何智能体部署的审计之中。
cs.LG / 120 / 2609.23219

Causal Inference with Unobserved Confounding: A Mixture Learning Perspective

存在未观测混杂因素的因果推断:混合学习的视角
Sood, Mansi, Shah, Devavrat
Abstract
Unobserved confounding is a fundamental challenge in causal inference from observational data. This article develops a mixture-learning perspective, viewing latent confounders as sources of heterogeneity that induce mixture structure in observed data. Under suitable structural and identifiability assumptions, recovering the mixing distribution and component mechanisms enables estimation of interventional distributions and causal estimands. Using variants of Bernoulli mixtures as a running example, we contextualize mixture-learning techniques and their structural assumptions, and connect them to causal inference in panel-data settings, including latent factor models and synthetic interventions.We then consider high-dimensional exponential-family mixtures with dependent outcome trajectories, moving beyond counterfactual means to model counterfactual distributions. We situate this perspective relative to complementary approaches for unobserved confounding. Together, these ideas provide a bridge between mixture learning and causal inference, connecting recent advances in high-dimensional mixture learning to scalable identification and estimation of causal effects while raising new challenges for mixture learning.
Chinese Translation
未观测混杂因素是基于观测数据进行因果推断的一项根本性挑战。本文提出一种混合学习的视角,将潜在混杂因素视为一种异质性来源,其在观测数据中诱导出混合结构。在适当的结构性与可识别性假设下,恢复混合分布及各分量的生成机制,即可估计干预分布与因果目标量。我们以伯努利混合模型的各种变体作为贯穿性示例,梳理混合学习技术及其结构性假设,并将其与面板数据场景中的因果推断方法相联系,包括潜在因子模型与合成干预方法。随后,我们考虑具有相依结果轨迹的高维指数族混合模型,从反事实均值扩展到对反事实分布的建模。我们将这一视角与应对未观测混杂因素的其他互补方法进行了对比定位。总体而言,这些思想为混合学习与因果推断之间架起了桥梁,将高维混合学习的最新进展与因果效应的可扩展识别和估计联系起来,同时也为混合学习提出了新的挑战。
cs.LG / 121 / 2609.23242

Proximal Residual Value Functions for Consistent Planning and Real-Time Execution

用于一致性规划与实时执行的近端残差价值函数
Waldon, Harrison, Eisenach, Carson, Bagaria, Akhil, Russo, Daniel, Perrault-Joncas, Dominique, Zachariah, Alisha, Foster, Dean
Abstract
We study two-timescale decision systems in which a planning layer periodically supplies a continuation-value function to a real-time optimizer that allocates arriving resources, with inventory placement as our motivating application. We propose an end-to-end reinforcement learning (RL) method for learning this function using \emph{proximal residual value functions}, which combine a strictly convex potential of post-decision inventory with a learned convex residual. This general form yields a well-posed optimization layer that supports end-to-end differentiation while preserving an explicit convex objective for real-time execution. We characterize the necessary and sufficient conditions under which a smooth value function yields decisions that are consistent across the planning and execution timescales. In an offline simulation using historical inventory arrival and demand patterns from a large e-commerce retailer, learned proximal residual value functions reduce total routing and transfer cost relative to a historical-production-system proxy by 5.0%.
Chinese Translation
我们研究了一类双时间尺度决策系统:其中规划层周期性地向实时优化器提供延续价值函数,实时优化器则对到达的资源进行分配,并以库存调拨作为我们的 motivating 应用。我们提出了一种端到端的强化学习(RL)方法,利用近端残差价值函数(proximal residual value functions)来学习该函数。该方法将决策后库存的严格凸势函数与一个可学习的凸残差相结合。这种一般化形式产生了良构的优化层,既支持端到端微分,又为实时执行保留了显式的凸目标函数。我们刻画了光滑价值函数在规划与执行两个时间尺度上产生一致决策的充要条件。在基于某大型电商零售商历史库存到达与需求模式的离线仿真中,学习得到的近端残差价值函数相对于历史生产系统代理方法,将总路由与调拨成本降低了5.0%。
cs.LG / 122 / 2609.23254

The Price of Self-Calibration: Exact Evidence Budgets and Manufactured Blind Sets in Adaptive Monitoring

自校准的代价:自适应监测中的精确证据预算与人为制造的盲集
Atarmla, Abdou-Raouf
Abstract
Self-calibrating monitors adapt their threshold online to guarantee a prescribed long-run false-alarm rate under arbitrary drift. We compute the price of that guarantee, stating every law with its exact domain of validity. First, the guarantee is an accounting identity, insensitive to what the monitor is meant to detect. Two evidence identities make the cost exact for the online quantile tracker: a persistent step of height $\delta$ yields excess alarm mass within one alarm of $\delta/\eta$, and exactly $\delta/\eta$ pathwise when $\delta$ is a lattice multiple of the gain $\eta$; a ramp of slope $c$ yields a stationary excess rate of exactly $c/\eta$, independent of accumulated size, up to a boundary $c=\eta(1-\alpha)$ coinciding with the alarm-rate cap. Second, the certificate's own fluctuation obeys an exact law: the windowed alarm rate has standard deviation of order $1/L$, not the binomial $1/\sqrt{L}$, since the windowed mass telescopes to a difference of a tight internal state; the closed-form constant is validated with no fitted parameter. Detectors calibrated on the binomial scale are miscalibrated by $\sqrt{\eta\varphi(q_0)L}$, and correct calibration turns detection windows from quadratic to linear in the inverse fault speed. Third, any monitor required to tolerate a drift class $\mathcal{D}$ is blind, at any horizon and for any rule, to every fault in $\mathcal{D}-\mathcal{D}$; the proof is a deliberately elementary two-point argument and the contribution is the object it identifies: for speed-bounded classes the blind set is exactly the doubled-speed class, and the tracker absorbs a speed class fixed by its own gain, so that under a certification regime declaring absorbed drift normal, the monitor manufactures $\mathcal{D}$. An exact Gaussian projection bound, sharper than Pinsker and never vacuous, quantifies power outside it.
Chinese Translation
自校准监测器在线调整其阈值,以保证在任意漂移下满足预设的长期误报率。我们计算了这一保证所付出的代价,并对每一条定律给出其精确的适用范围。首先,该保证本质上是一种记账恒等式,对监测器所要检测的对象并不敏感。两个证据恒等式使在线分位数跟踪器(online quantile tracker)的代价可以被精确刻画:对于高度为 δ 的持续性阶跃故障,超出告警质量在一个告警量之内接近 δ/η,且当 δ 为增益 η 的格点倍数时逐路径精确等于 δ/η;对于斜率为 c 的斜坡故障,其稳态超出率精确为 c/η,与累计规模无关,直至边界 c=η(1−α),该边界恰与告警率上限重合。其次,该证书自身的波动服从精确定律:窗口化告警率的标准差为 1/L 量级,而非二项分布的 1/√L,因为窗口化质量可伸缩相消为一个有界的内部状态之差;其闭式常数在无任何拟合参数的情况下得到验证。以二项分布尺度进行校准的检测器会因 √(ηφ(q₀)L) 因子而失准,而正确的校准可将检测窗口对逆故障速度的关系由二次降为线性。第三,任何被要求容忍漂移类 𝒟 的监测器,在任何时间范围和任何规则下,对 𝒟−𝒟 中的每一个故障都是盲的;其证明是一个刻意简化的两点论证,而其贡献在于它所识别出的对象:对于速度受限的类,盲集恰好是倍速类;跟踪器会吸收由其自身增益所决定的速度类,因而在一种将已吸收漂移判定为正常的认证制度下,监测器会人为制造出 𝒟。一个精确的高斯投影界(比 Pinsker 不等式更紧且从不空泛)量化了盲集之外的检测能力。
cs.LG / 123 / 2609.23257

CTRL: Control-Based Time Series Forecasting with LLM-Guided Residual Learning

CTRL:基于控制的时间序列预测与LLM引导的残差学习
Kim, Minkyoung, Ji, Daeun, Lee, Yohan, Kim, Beomsoo, Jang, Beakcheol
Abstract
Time series forecasting underpins critical decision-making across diverse domains. While large language models (LLMs) offer promising reasoning capabilities, existing LLM-based time series forecasting approaches either reduce them to numerical predictors that bypass their strengths, or allow direct forecast generation that destabilizes predictions in non-stationary settings. We introduce CTRL, a framework that decouples semantic reasoning from quantitative prediction. A frozen backbone generates base forecasts, while specialized LLM agents function as controllers that analyze backbone prediction errors through decomposed trend, seasonal, and irregular components, grounding reasoning in interpretable temporal structure. Each agent outputs compact control signals that a lightweight residual decoder translates into forecast corrections. CTRL incorporates label-free test-time adaptation that detects distribution shift from input statistics alone and readapts control signals with only 3-24 LLM calls via caching. CTRL is explicitly designed to improve robustness under non-stationary temporal dynamics and distribution shift, while remaining competitive on highly stationary time series where adaptive correction provides limited additional benefit.
Chinese Translation
时间序列预测是多个领域中关键决策的基础。尽管大语言模型(LLM)展现出可喜的推理能力,但现有基于LLM的时间序列预测方法要么将其简化为数值预测器而绕过了其优势,要么允许直接生成预测,从而在非平稳环境下破坏预测的稳定性。我们提出CTRL,一个将语义推理与定量预测解耦的框架。一个冻结的骨干模型生成基础预测,同时专门的LLM智能体充当控制器,通过分解的趋势、季节性和不规则分量分析骨干模型的预测误差,将推理建立在可解释的时间结构之上。每个智能体输出紧凑的控制信号,由轻量级残差解码器将其转化为预测修正。CTRL引入了无标签的测试时自适应机制,仅凭输入统计量即可检测分布偏移,并通过缓存机制仅需3-24次LLM调用即可重新适配控制信号。CTRL旨在明确提升非平稳时间动态和分布偏移下的鲁棒性,同时在高度平稳的时间序列上保持竞争力——尽管此时自适应修正带来的额外收益有限。
cs.LG / 124 / 2609.23260

Why Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal Learning

为什么幽灵输出能够教学:基于核方法对阈下学习的理解
Li, Zhe, Ying, Bicheng, Dong, Chaosheng, Yang, Haibo
Abstract
Subliminal Learning (SL) is a recently identified phenomenon in which a student model acquires downstream task capabilities by matching seemingly unrelated auxiliary outputs from a teacher, despite never observing task labels, task-specific outputs, or the original training data. While recent studies have identified where subliminal signals may reside, the optimization mechanism underlying this phenomenon remains poorly understood. In this work, we provide a mechanistic understanding of SL through the lens of learning dynamics. Specifically, we derive a chained cross-task kernel that explicitly links ghost-output supervision to changes in task predictions through shared backbone representations. Our unified analytical framework provides a rigorous mathematical explanation for three central empirical puzzles in SL: (i) under shared initialization, the transfer operator forms a strictly Positive Semi-Definite (PSD) structure, guaranteeing that ghost-output optimization aligns the student with the teacher's true task objective without explicit label exposure; (ii) the ghost-output dimensionality acts as an explicit rank bottleneck governing the transfer of task-relevant features; and (iii) synthetic, high-entropy inputs function as broadband probes that maximize cross-task kernel overlap, explaining why random noise consistently outperforms structured data for subliminal transfer. Experiments on the canonical ghost-output setting validate all three theoretical predictions, providing the first learning-dynamics-based theoretical explanation of how ghost-output supervision gives rise to subliminal learning.
Chinese Translation
阈下学习(Subliminal Learning, SL)是最近发现的一种现象:学生模型通过匹配来自教师模型的看似无关的辅助输出,即可获得下游任务能力,而在此过程中学生从未接触过任务标签、任务特定输出或原始训练数据。尽管已有研究指出了阈下信号可能存在的位置,但这一现象背后的优化机制仍知之甚少。在本工作中,我们从学习动力学的视角为SL提供了一种机制层面的理解。具体而言,我们推导出一个链式跨任务核(chained cross-task kernel),该核通过共享的骨干网络表征,显式地将幽灵输出(ghost-output)监督与任务预测的变化联系起来。我们的统一分析框架为SL中的三个核心实证难题提供了严格的数学解释:(i) 在共享初始化条件下,迁移算子形成严格的半正定(Positive Semi-Definite, PSD)结构,从而保证幽灵输出优化能够在不显式暴露标签的情况下使学生与教师的真实任务目标对齐;(ii) 幽灵输出的维度构成一个显式的秩瓶颈,控制着任务相关特征的迁移;以及(iii) 合成的高熵输入充当宽带探针,能够最大化跨任务核的重叠度,这解释了为什么随机噪声在阈下迁移中始终优于结构化数据。在标准幽灵输出设置上的实验验证了全部三个理论预测,首次为幽灵输出监督如何引发阈下学习提供了基于学习动力学的理论解释。
cs.LG / 125 / 2609.23265

Optimal No-Regret Learning for Repeated Prophet Inequality

重复先知不等式的最优无悔学习
Wang, Kun
Abstract
We study repeated prophet inequalities under prefix feedback. In each of $T$ rounds, a learner encounters fresh values drawn independently from $n$ boxes with unknown $[0,1]$-supported distributions in a fixed order and must irrevocably accept one, observing only the prefix up to its stopping box. Regret is measured against the optimal stopping policy that knows the distributions. We give an efficient algorithm achieving $\widetilde O(\sqrt{T})$ expected regret, matching the lower bound up to logarithmic factors. Our algorithm explores directly through near-optimal policies, combining empirical backward induction with box-specific reach bonuses. A relative-drop aggregation rule then exploits the nesting structure of observed prefixes to preserve exploration, thereby removing the polynomial dependence on the box number $n$. This resolves an open question posed by Liu et al. (2025).
Chinese Translation
我们研究了前缀反馈下的重复先知不等式问题。在 $T$ 轮的每一轮中,学习者在固定顺序下遇到从 $n$ 个箱子中独立抽取的新值,这些箱子服从未知的 $[0,1]$ 区间上的分布,学习者必须不可撤销地接受其中一个值,且只能观察到直至其停止箱子为止的前缀。遗憾值以已知分布的最优停止策略为基准来衡量。我们给出了一种高效算法,实现 $\widetilde O(\sqrt{T})$ 的期望遗憾,且在对数因子意义下与下界匹配。我们的算法通过近最优策略直接进行探索,将经验性反向归纳与针对每个箱子的可达性奖励相结合。随后,一种相对下降聚合规则利用所观察前缀的嵌套结构来保持探索性,从而消除了对箱子数量 $n$ 的多项式依赖。这解决了 Liu 等人(2025)提出的一个开放性问题。
cs.LG / 126 / 2609.23308

Optimal Multi-way Decision Trees for Stratified Sampling in Online Controlled Experiments

在线受控实验中用于分层抽样的最优多路决策树
Takei, Tomoka, Ikeda, Shunnosuke, Takano, Yuichi
Abstract
Online controlled experiments, or A/B tests, are widely used to estimate causal effects on digital platforms. A central challenge is to improve experimental sensitivity, or statistical power, without increasing the experimental sample size. Stratified sampling is a classical variance reduction technique; however, its effectiveness depends critically on how the strata are constructed. We thus propose an optimization-based stratification framework for stratified sampling using optimal multi-way decision trees. Our method, called Optimal Multi-way Stratification Trees (OMST), formulates stratification as a path-selection problem over a feature graph. The selected paths define interpretable stratification rules and are optimized using an exact variance-minimizing binary optimization formulation under continuous proportional allocation and a Neyman-type optimal allocation. We incorporate supervised optimal binning to generate outcome-relevant candidate splits for numerical features. Furthermore, we introduce reduction procedures for redundant candidate paths and assignment constraints, substantially reducing the optimization problem size. Experiments on both a real-world and a simulated dataset demonstrate that OMST achieves comparable or superior variance reduction to existing methods while maintaining shallow and interpretable stratification trees.
Chinese Translation
在线受控实验(即A/B测试)被广泛用于估计数字平台上的因果效应。一个核心挑战是在不增加实验样本量的情况下提高实验敏感性(即统计功效)。分层抽样是一种经典的方差缩减技术;然而,其有效性在很大程度上取决于如何构建分层。为此,我们提出了一种基于最优多路决策树的分层抽样优化框架。我们的方法称为最优多路分层树(Optimal Multi-way Stratification Trees, OMST),将分层构建表述为特征图上的路径选择问题。所选路径定义了可解释的分层规则,并通过精确的方差最小化二元优化模型进行优化,涵盖连续比例分配和Neyman型最优分配两种情形。我们引入监督式最优分箱方法,为数值特征生成与结果相关的候选切分点。此外,我们提出了冗余候选路径和分配约束的约简程序,从而大幅减小优化问题的规模。在真实数据和模拟数据集上的实验表明,OMST在保持浅层且可解释的分层树的同时,实现了与现有方法相当或更优的方差缩减效果。
cs.LG / 127 / 2609.23314

ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs

ValueDiff:面向汇点抑制型大语言模型的价值几何KV缓存淘汰方法
Park, Junyoung, Choi, Jungwook, Lee, Mingu
Abstract
Modern LLMs with QK-normalization, gated attention, learned attention sinks, or logit softcapping exhibit weaker persistent attention sinks, on which existing KV cache eviction methods primarily rely. We observe that across these models, weaker sinks co-occur with greater value-vector dispersion relative to key-vector dispersion. Motivated by this value-side dispersion, we present ValueDiff, a value-geometric eviction that ranks tokens by the L2 deviation of their value vectors from the cache mean. The same score arises as the minimal-disturbance eviction under a max-entropy assumption about future attention. We evaluate under fixed cache budgets, with eviction at every block boundary during prefill and at every decoding step during generation. On RULER at a tight 2k token budget, ValueDiff retains 88--99\% of dense across seven sink-suppressed models (best on 6 out of 7). On LongBench at the 4k budget, ValueDiff averages 92\% retention across sink-suppressed models versus 83\% for the strongest prior baseline. On MATH-500, ValueDiff is the strongest non-dense method on every sink-suppressed model tested at the 25\% cache budget, outperforming prior methods by up to $\sim$20 points on gated-attention models. Across all three benchmarks, value geometry emerges as the more reliable query-invariant eviction signal for sink-suppressed models.
Chinese Translation
采用QK归一化、门控注意力、可学习的注意力汇点或logit软截断的现代大语言模型表现出更弱的持久注意力汇点,而现有的KV缓存淘汰方法主要依赖这些汇点。我们观察到,在这类模型中,更弱的汇点往往伴随着相对于键向量离散度更大的值向量离散度。受这种值侧离散度的启发,我们提出了ValueDiff,一种基于值几何的淘汰方法,它根据值向量相对于缓存均值的L2偏差对token进行排序。在关于未来注意力的最大熵假设下,该得分恰好可推导为最小扰动淘汰方案。我们在固定缓存预算下进行评估,在预填充阶段的每个块边界以及生成阶段的每个解码步骤执行淘汰。在仅2k token预算的RULER基准上,ValueDiff在七个汇点抑制型模型上保留了稠密注意力88%–99%的性能(7个模型中的6个表现最佳)。在4k预算的LongBench上,ValueDiff在汇点抑制型模型上平均保留92%的性能,而最强基线仅为83%。在MATH-500上,在25%缓存预算下,ValueDiff在所有测试的汇点抑制型模型上均是最强的非稠密方法,在门控注意力模型上比先前方法高出最多约20分。在全部三个基准上,值几何被证明是汇点抑制型模型中更可靠的查询不变淘汰信号。
cs.LG / 128 / 2609.23320

CSC: Calibrated Simplicity for Conflict-Aware Social Bot Detection in the LLM Era

CSC:面向LLM时代冲突感知社交机器人检测的校准简洁框架
Qian, Yipeng, Zhao, Pengjie, Niu, Chaoxi
Abstract
Social bot detection is essential for protecting online platforms from misinformation amplification, coordinated manipulation, and distorted public discourse. However, large language models have made social bots much harder to detect from text alone because semantic camouflage is now cheap, fluent, and scalable. The resulting challenge is modality conflict: an account may look human-like in semantics while remaining suspicious in graph structure, profile attributes, or cross-modal consistency. Recent graph-based detectors tackle this limitation by adding graph-side complexity, such as sparse prototype selection, adaptive gating, or architecture-specific control logic, yet our experiments suggest that complexity alone is not the most reliable way to resolve such conflict. We therefore propose CSC, a calibrated-simplicity framework for conflict-aware LLM-era social bot detection. The framework combines three design choices: a simplified prototype-guided graph expert that retains useful structural biases while removing unstable graph-side heuristics, calibrated simplex-constrained fusion that aligns heterogeneous confidence spaces before late fusion, and a lightweight inconsistency expert that models cross-modal disagreement. Experiments on TwiBot-22, TwiBot-20, and MGStBot-large show that \textsc{CSC} improves calibrated operating-point decision quality while remaining competitive across external benchmarks. Further analyses show that calibration improves confidence reliability, the inconsistency expert mainly provides localized corrections in high-conflict or near-threshold regions, and simplified graph-side control yields a better stability-cost trade-off. A targeted semantic-camouflage stress test further shows that replacing selected bot text with matched human text sharply degrades the standalone text expert while leaving graph and fused evidence stable on a balanced challenge set.
Chinese Translation
社交机器人检测对于保护在线平台免受虚假信息放大、协同操纵和公共话语扭曲至关重要。然而,大语言模型(LLM)的出现使得仅从文本层面检测社交机器人变得困难得多,因为语义伪装如今成本低廉、语言流畅且易于规模化。由此产生的挑战是模态冲突:一个账户在语义上可能看起来像真人,但在图结构、用户属性或跨模态一致性方面仍然可疑。近期的基于图的检测器通过增加图侧的复杂度来应对这一局限,例如稀疏原型选择、自适应门控或特定于架构的控制逻辑,但我们的实验表明,单纯依靠复杂度并非解决此类冲突最可靠的方式。因此,我们提出了CSC——一个面向冲突感知的LLM时代社交机器人检测的校准简洁框架。该框架包含三项设计选择:(1)简化的原型引导图专家,在保留有用结构偏置的同时移除不稳定的图侧启发式方法;(2)校准的单纯形约束融合,在后期融合前对齐异构的置信度空间;(3)轻量级不一致性专家,用于建模跨模态分歧。在TwiBot-22、TwiBot-20和MGStBot-large上的实验表明,CSC在保持外部基准竞争力的同时,提升了校准工作点下的决策质量。进一步分析表明,校准提升了置信度的可靠性;不一致性专家主要在高冲突或接近决策阈值的区域提供局部化修正;简化的图侧控制则带来了更好的稳定性-成本权衡。一项针对性的语义伪装压力测试进一步表明,在均衡的挑战集上,将部分机器人文本替换为匹配的人类文本会使独立的文本专家性能急剧下降,而图证据和融合证据则保持稳定。
cs.LG / 129 / 2609.23325

Rethinking Class Imbalance for Single-Cell Foundation Models: A Systematic Benchmark Across Architectures and Long-Tail Loss Functions

重新思考单细胞基础模型中的类别不平衡问题:跨架构与长尾损失函数的系统性基准测试
Dong, Zeyu, Zhong, Jiahui
Abstract
Single-cell foundation models (scGPT, scBERT, Geneformer) achieve cell-type classification accuracy up to 97.5% in our experiments, yet this aggregate accuracy can mask systematic failure on rare, often disease-relevant cell populations that long-tail loss functions are widely assumed to address. We present a systematic benchmark of six long-tail loss functions (cross-entropy, weighted CE, class-balanced loss, focal loss, LDAM, logit-adjusted softmax) across three architectures and three datasets (Multiple Sclerosis, Zheng68K, human Pancreas), totaling 162 controlled training runs (3 backbones x 3 datasets x 6 losses x 3 seeds). The gap between overall accuracy, Macro-F1, and rare-class recall under plain cross-entropy is consistent across all nine (architecture, dataset) settings, driven by dataset structure rather than pretraining. Rare-class failure itself splits into two regimes with distinct embedding-geometry signatures, visible before any loss is chosen: some classes are recoverable by the right loss, while others retain linear separability yet are absorbed into unrelated classes' neighborhoods under every evaluated loss and architecture. Among the recoverable classes, the efficacy of reweighting is predicted by a class's absolute training-set size, rather than its share of the dataset or the dataset's overall imbalance ratio. Class-balanced loss and LDAM are the most consistent choices across all nine settings, while logit adjustment trades rare-class precision for recall rather than improving both. Our results give both a reusable benchmark and mechanism-grounded practical guidelines for combining foundation models with imbalanced biological data.
Chinese Translation
单细胞基础模型(scGPT、scBERT、Geneformer)在我们的实验中实现了高达97.5%的细胞类型分类准确率,然而这一总体准确率可能掩盖了模型在罕见(通常与疾病相关)细胞群体上的系统性失效,而长尾损失函数被普遍认为可以解决这一问题。我们对六种长尾损失函数(交叉熵、加权交叉熵、类平衡损失、focal loss、LDAM、logit-adjusted softmax)在三种架构和三个数据集(多发性硬化症、Zheng68K、人类胰腺)上进行了系统性基准测试,共计162次受控训练运行(3个骨干网络 × 3个数据集 × 6种损失函数 × 3个随机种子)。在纯交叉熵下,总体准确率、Macro-F1与稀有类别召回率之间的差距在全部九种(架构,数据集)设置中保持一致,且该差距由数据集结构而非预训练所驱动。稀有类别的失效本身可划分为两种具有不同嵌入几何特征的机制,且在任何损失函数选择之前即可观察到:某些类别可以通过合适的损失函数得到挽救,而另一些类别虽然保持线性可分性,却在所有被评估的损失函数和架构下都被吸收到无关类别的邻域中。在可挽救的类别中,重加权的效果由该类别的绝对训练集大小决定,而非其在数据集中所占比例或数据集的整体不平衡比率。类平衡损失和LDAM是在全部九种设置中最稳定的选择,而logit调整则是以牺牲稀有类别精确率来换取召回率的提升,而非同时改善两者。我们的研究结果不仅提供了一个可复用的基准,还为将基础模型与不平衡生物数据相结合提供了基于机制的实用指导。
cs.LG / 130 / 2609.23333

A Patient World Model for Early Forecasting of Digital Health Campaign Outcomes: Capabilities and Limits

用于数字健康营销活动结果早期预测的患者世界模型:能力与局限
Wang, Yunlong
Abstract
Digital direct-to-consumer (DTC) health campaigns are usually measured after the fact. In-flight forecasting commonly relies on a separate classifier for every cutoff and horizon. We treat this task as a dynamic-system problem and build a compact patient world model. The architecture maintains a latent state per patient, learns exposure-conditioned state dynamics jointly with a weekly conversion hazard, and rolls forward into future conversion curves. We evaluate it on a US campaign dataset with 147{,}173 patients and 5.2 million at-risk person-weeks. In a retrospective evaluation conditioned on recorded future exposures, the model forecasts the remaining new-to-brand prescription volume through week 52 with a relative error of 2.9\% from a week-4 cutoff and 0.8--2.6\% from cutoffs at weeks 8--26. The strongest non-recurrent baseline, a pooled-hazard gradient boosting model given the same survival rollout and information, has relative errors of 13.6--33.1\%. Per-horizon classifiers perform substantially worse. A Fisher-information analysis motivates dense next-exposure supervision when conversions are rare. Removing this auxiliary objective increases prescription-volume error by approximately $2$--$14\times$, while providing no consistent disadvantage on the more common specialist-visit outcome. We also evaluate scenario simulation. Switching all future exposure off raises predicted conversion from 0.31 to 0.89, a pattern consistent with selection effects in observational exposure data. This result highlights the limits of interpreting exposure-conditioned rollouts causally.
Chinese Translation
数字直接面向消费者(DTC)健康营销活动通常是在事后进行评估。实时的活动效果预测通常依赖于针对每个预测时点和预测周期分别训练的独立分类器。我们将该任务视为一个动力系统问题,并构建了一个紧凑的患者世界模型(world model)。该架构为每位患者维护一个潜在状态,将曝光条件化的状态动力学与每周转化风险(hazard)进行联合学习,并向前滚动以预测未来的转化曲线。我们在一个包含147,173名患者和520万风险人周的美国营销活动数据集上对该模型进行了评估。在以记录的未来曝光为条件的回溯性评估中,该模型从第4周的预测时点起,预测直至第52周的剩余新品牌处方量的相对误差为2.9%;从第8至26周的预测时点起,相对误差为0.8%–2.6%。在给定相同的生存滚动和信息的情况下,表现最强的非循环基线——一个池化风险的梯度提升模型——的相对误差为13.6%–33.1%,而针对各预测周期的分类器表现则明显更差。Fisher信息分析表明,当转化事件稀少时,应采用密集的下次曝光监督。移除该辅助目标会使处方量误差增加约2至14倍,而对于更为常见的专科就诊结果则未带来一致的劣化。我们还评估了情景模拟能力。将所有未来曝光关闭会使预测转化率从0.31升至0.89,这一模式与观察性曝光数据中的选择效应相一致。该结果凸显了对曝光条件化滚动预测进行因果解释的局限性。
cs.LG / 131 / 2609.23366

What Can a Recurrent State Safely Forget?

循环状态能安全遗忘什么?
Zhang, Linzhe, Xu, Changming
Abstract
Recurrent models must preserve information that changes future behavior while suppressing hidden-state error. These objectives conflict: contraction improves stability, but contraction along a future-distinguishing direction destroys memory. We formalize this boundary through the predictive quotient of a recurrent state space. Two hidden states are equivalent when they induce the same conditional future; their equivalence classes form predictive fibers. Every exact semantics-preserving corrector acts as the identity on this quotient. At a regular point with hidden dimension d and predictive dimension k, it can eliminate at most d - k independent directions. This establishes a discrete-continuous boundary: finite predictive states admit positive-radius exact correction basins, whereas an uncountable continuum of future-distinguishable states cannot be decoded after arbitrary positive-radius perturbations in finite-dimensional Euclidean space. To operationalize this principle, we develop an auditable finite-future framework. A compact deployment bank W is evaluated against an independent audit bank A (W subseteq A) on a declared correction domain. Under generative probe access and audit-metric coverage, finite stochastic rollouts furnish a high-probability certificate for the separation margin Omega_{W|A}(delta). Preserving learned W-predictions within this certified margin guarantees bounded audit-semantic distortion. For intrinsic audit dimension k, the required probe outcomes scale as O(M * Omega^{-(k+2)}), where M = |A|; a matching minimax lower bound proves this exponent is optimal. Extending guarantees to continuous futures is achieved via an explicit completeness modulus. Controlled experiments validate the certified margins, scaling laws, and automated probe refinement under a safety-first evaluation paradigm.
Chinese Translation
循环模型必须保留会改变未来行为的信息,同时抑制隐状态误差。这两个目标相互冲突:收缩可提升稳定性,但沿未来区分方向上的收缩会破坏记忆。我们通过循环状态空间的预测商(predictive quotient)来形式化这一边界。当两个隐状态诱导出相同的条件未来时,它们是等价的;其等价类构成预测纤维(predictive fibers)。任何保持语义的精确校正器在该商上均表现为恒等映射。在隐维度为 d、预测维度为 k 的正则点处,校正器至多能消除 d - k 个独立方向。这确立了一条离散-连续边界:有限的预测状态可以拥有正半径的精确校正盆地,而不可数的、可区分未来的连续状态在有限维欧几里得空间中无法在任意正半径扰动后被解码。为将该原理可操作化,我们开发了一个可审计的有限未来框架。在一个声明的校正域上,将紧凑的部署库 W 对照独立的审计库 A(W subseteq A)进行评估。在生成式探测访问与审计度量覆盖的条件下,有限随机展开可为分离裕度 Omega_{W|A}(delta) 提供高概率证书。在该认证裕度内保持学习到的 W-预测,可保证有界的审计语义失真。对于内在审计维度 k,所需探测结果的数量按 O(M * Omega^{-(k+2)}) 规模增长,其中 M = |A|;一个匹配的极小极大下界证明了该指数是最优的。通过一个显式的完备性模量(completeness modulus),我们将保证扩展至连续未来。受控实验在安全优先的评估范式下验证了认证裕度、标度律以及自动探测精化方法。
cs.LG / 132 / 2609.23374

The Evidence Ladder for Reinforcement Learning in Healthcare: From Retrospective Policies to Trusted Interventions

医疗领域强化学习的证据阶梯:从回顾性策略到可信干预
Zhao, Yunfan
Abstract
Reinforcement learning (RL) offers a natural language for healthcare decisions whose conse- quences unfold over time, yet most reported progress remains far from routine intervention. Ex- isting surveys organize the field by algorithm or clinical application. We instead review healthcare RL through an evidence ladder: problem formulation, retrospective identification, policy estima- tion, stress testing, prospective evaluation, and lifecycle monitoring. This view connects clinical treatment, patient engagement, and health-system operations while exposing a recurring gap: evi- dence that a policy scores well in a historical dataset is not evidence that it will improve care. We synthesize the assumptions and failure modes at each rung, identify what evidence can and can- not transfer across settings, and propose reporting practices for cumulative evaluation. Restless bandits are included as one special case, not as the organizing framework. The central lesson is that healthcare RL should be evaluated as an intervention embedded in a changing sociotechnical system, rather than only as an optimizer of a retrospective reward.
Chinese Translation
强化学习(RL)为那些后果随时间逐步显现的医疗决策提供了一种自然的表达语言,然而大多数已报道的进展仍远未成为常规干预手段。现有综述通常按算法或临床应用来组织该领域,而我们则通过一个证据阶梯来审视医疗强化学习:问题建模、回顾性识别、策略估计、压力测试、前瞻性评估以及生命周期监测。这一视角将临床治疗、患者参与和医疗系统运营联系起来,同时揭示了一个反复出现的鸿沟:策略在历史数据集上得分很高,并不构成其能够改善医疗服务的证据。我们综合梳理了每一层级所依赖的假设与失效模式,明确了哪些证据能够跨场景迁移、哪些不能,并提出了用于累积性评估的报告规范。Restless bandits(不安分老虎机)作为其中一个特例被纳入讨论,而非作为组织框架。核心结论是:医疗强化学习应当被视为嵌入在不断变化的社会技术系统中的干预手段来评估,而不仅仅是某个回顾性奖励函数的优化器。
cs.LG / 133 / 2609.23381

Discovering Physical Representation Languages

发现物理表示语言
Zhang, Linzhe, Xu, Changming
Abstract
Before a machine can discover a physical law, it must discover what its measurements are: which observations live on cells, which are intensive or extensive, which sectors are dual, and which distinctions are merely gauge. We introduce physical representation-language discovery, the problem of recovering this hidden ontology directly from anonymous controlled experiments. We give an identifiability theory and constructive polynomial-time procedure that recovers a carrier and differential sequence, measurement types and orientation twist, noninvertible refinement semantics, primal-dual Maxwell diagrams, and the residual equivalences that no permitted experiment can break. The theory turns material nuisance into a commutant, uses refinement to separate quantities from coordinates, and selects physics only after its representation has been recovered. For a certified finite experiment family, we prove an end-to-end two-stage measurement bound and a matching minimax rate in dimension, accuracy, and confidence. Blind Maxwell experiments recover complete primal/relative-dual ontologies on regular and unstructured carriers under jointly corrupted observations; an independent unstructured RLC system demonstrates that the result is not specific to Maxwell. The framework scales to tens of thousands of cells per carrier, while stress audits demonstrate robustness across severe physical regimes - including non-Markovian memory, nonlinearities, nonlocality, and complex constitutive hysteresis. A public FDTD audit demonstrates the emergence of anonymous curl structure from incomplete field data, while characterizing the informational prerequisites for complete recovery. The goal is to move scientific ML from learning laws in a human-supplied language to discovering the language in which laws become expressible, establishing exact theoretical limits on observational identifiability.
Chinese Translation
在机器能够发现物理定律之前,它必须先发现其测量量究竟是什么:哪些观测位于胞元上,哪些是强度量或广延量,哪些部分是对偶的,以及哪些区别仅仅是规范性的。我们提出了物理表示语言发现问题,即直接从匿名的受控实验中恢复这一隐藏本体论的问题。我们给出了一套可辨识性理论和一个构造性的多项式时间程序,能够恢复载体(carrier)与微分序列、测量类型与取向扭转、不可逆的细化语义、原-对偶麦克斯韦图(primal-dual Maxwell diagrams),以及任何允许的实验都无法打破的残余等价关系。该理论将材料噪声转化为交换子(commutant),利用细化将物理量与坐标系分离,并仅在表示被恢复之后才选择物理规律。对于一个经过认证的有限实验族,我们证明了端到端的两阶段测量界,以及与维度、精度和置信度相匹配的极小极大速率。盲麦克斯韦实验在规则与非结构化载体上、在联合污染的观测条件下恢复出完整的原/相对对偶本体论;一个独立的非结构化RLC系统表明该结果并非麦克斯韦特有。该框架可扩展到每个载体上数万个胞元,同时压力审计证明了其在严峻物理条件下的鲁棒性——包括非马尔可夫记忆、非线性、非定域性以及复杂的本构迟滞。一项公开的FDTD审计展示了从不完整场数据中匿名旋度结构的涌现,同时刻画了完全恢复所需的信息前提。我们的目标是推动科学机器学习从在人类提供的语言中学习定律,转变为发现使定律变得可表达的 语言,从而在观测可辨识性上建立精确的理论极限。
cs.LG / 134 / 2609.23387

Blind Thermodynamic Ontology Discovery from Anonymous Experiments

基于匿名实验的盲热力学本体发现
Zhang, Linzhe, Xu, Changming
Abstract
Before a machine learning model can learn a thermodynamic equation of state, it must discover what its measurements represent: which channels scale with system size, which are intensive conjugates, how sectors pair through contact, and which potential governs stability. When sensors expose only an unknown linear mixture of extensive states and intensive responses, passive observations cannot disentangle physical quantities from coordinate artifacts. We formulate the problem of discovering this hidden thermodynamic ontology directly from anonymous controlled experiments. We present an operational identifiability theory and a constructive polynomial-time algorithm that extracts extensive and intensive scaling sectors from replication contrasts, recovers their dual cotangent pairing from thermal contact and reciprocity, verifies a globally admissible concave potential via discrete cyclic concavity, and determines an invariant matroid of reservoir ensembles. We prove that the residual observational equivalence is strictly (x, lambda) ~ (A x, a A^{-T} lambda + beta), establishing the sharp observational limit that no permitted experiment can break. Blind evaluations on van der Waals fluids and Curie-Weiss magnets confirm robust recovery under ill-conditioned mixing, correctly resolving anonymous Maxwell tie-lines while rejecting non-equilibrium continuations. External validation across six real fluids from the NIST WebBook demonstrates that operational ontology discovery transfers across real physical substances without coordinate leakage.
Chinese Translation
在机器学习模型学习热力学状态方程之前,它必须先发现其测量数据所代表的物理含义:哪些通道随系统规模伸缩,哪些是强度量共轭对,各扇区如何通过热接触配对,以及哪个势函数主导稳定性。当传感器仅暴露广延态与强度响应的未知线性混合时,被动观测无法将物理量与坐标伪影分离开来。我们提出了直接从匿名受控实验中发现这一隐藏热力学本体的问题。我们给出了一套可操作的可辨识性理论以及一个构造性的多项式时间算法:该算法从复制对比中提取广延与强度标度扇区,从热接触与互易性中恢复其对偶余切配对,通过离散循环凹性验证全局可容许的凹势函数,并确定储库集合的不变拟阵。我们证明,残余的观测等价性严格为 (x, lambda) ~ (A x, a A^{-T} lambda + beta),确立了任何被允许的实验都无法突破的精确观测极限。在范德华流体与居里-外斯磁体上的盲评估表明,即使在病态混合条件下也能稳健恢复,能够正确解析匿名的麦克斯韦双节线,同时拒绝非平衡延拓。基于 NIST WebBook 中六种真实流体的外部验证表明,可操作的本体发现能够在真实物理物质间迁移,而不发生坐标泄露。
cs.LG / 135 / 2609.23435

Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks

面向序列组学任务中LLM领域自适应的工具增强在策略蒸馏方法
Ying, Jie, Wang, Zhefan, Chen, Zihong, Li, Zhengqing, Li, Jinzhe, Li, Gang, Liu, Jian, Hu, Fang, Luo, Tao, Yuan, Zhonghang, Ouyang, Wanli, Li, Stan Z., Yang, Fan, Dong, Nanqing
Abstract
Multi-omics sequences contain complex biological patterns, yet deciphering their mechanisms for automated scientific discovery remains challenging. As large language models (LLMs) interpret these sequences, evaluating both predictions and scientific reasoning is critical. However, existing benchmarks for multi-omics sequence tasks rely on classification and regression metrics, neglecting whether models grasp the underlying biological evidence. We introduce OmicsBench, the first reasoning benchmark for multi-omics sequences, comprising 1,160 expert-validated questions across six tasks spanning DNA regulation, RNA processing, and protein function. OmicsBench requires traceable evidence chains, evaluated using instance-specific rubrics developed with domain experts. Evaluating 17 LLMs reveals an inverse relationship: while scientific LLMs outperform general-purpose LLMs in sequence classification accuracy, they fail to provide valid evidence to support their predictions. One plausible interpretation is shortcut learning: specialized models may rely on statistical patterns rather than the biological mechanisms needed for scientific discovery. Motivated by this finding, we introduce tool-augmented on-policy distillation (TA-OPD), a post-training method to align sequence prediction with evidence-grounded biological reasoning. Across five Qwen3.5 models spanning 0.8B to 27B parameters, TA-OPD consistently strengthens biological evidence grounding while improving predictive performance on most tasks. These gains persist across model scales, indicating that stronger sequence reasoning does not arise solely from increased model capacity, but can be improved through evidence-aware training. Together, OmicsBench and TA-OPD provide a framework for diagnosing reasoning failures in multi-omics LLMs and a path toward models whose predictions are better grounded in biologically meaningful evidence.
Chinese Translation
多组学序列蕴含复杂的生物学模式,然而为自动化科学发现解读其机制仍极具挑战性。随着大语言模型(LLM)开始解释这些序列,对其预测结果与科学推理能力进行评估至关重要。然而,现有的多组学序列任务基准依赖分类和回归指标,忽略了模型是否真正掌握潜在的生物学证据。我们提出OmicsBench,这是首个面向多组学序列的推理基准,包含1,160道经专家验证的问题,涵盖DNA调控、RNA加工和蛋白质功能六项任务。OmicsBench要求提供可追溯的证据链,并使用与领域专家共同开发的实例化评分标准进行评估。对17个大语言模型的评估揭示了一种反向关系:尽管科学领域LLM在序列分类准确率上优于通用LLM,但它们无法提供有效证据来支持其预测。一种可能的解释是捷径学习:专用模型可能依赖于统计模式,而非科学发现所需的生物学机制。基于这一发现,我们提出工具增强在策略蒸馏(Tool-Augmented On-Policy Distillation, TA-OPD),这是一种后训练方法,用于将序列预测与基于证据的生物学推理对齐。在参数规模从0.8B到27B的五个Qwen3.5模型上,TA-OPD在提升多数任务预测性能的同时,持续强化了生物学证据的支撑能力。这些收益在不同模型规模上均得以保持,表明更强的序列推理能力并非仅来自模型容量的增加,而可以通过证据感知的训练加以提升。OmicsBench与TA-OPD共同提供了一个诊断多组学LLM推理失效的框架,并为构建预测更根植于有意义生物学证据的模型指明了路径。
cs.LG / 136 / 2609.23457

RLVR$^{2}$: Reinforcement Learning with Verifiable Rubric-based Ranking

RLVR$^{2}$:基于可验证评分准则排序的强化学习
Li, Hao, Zhang, Zhengkun, Hu, Gangqiang, Zhang, Zhen, Gao, Yude, Dai, Dai, Liu, Jing
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) is expanding from tasks with well-defined correctness signals, such as mathematics and code, toward multifaceted quality requirements specified by multi-dimensional rubrics. Since policy optimization consumes one scalar per rollout, rubric-based pipelines must map multiple criterion scores into a scalar reward. This aggregation is often treated as score scaling, but it implicitly determines how quality dimensions trade off during training. The prevailing practice, normalizing each criterion and taking a linear combination, assumes that cardinal score differences are comparable across criteria and that gains on one criterion compensate for failures on another; both assumptions are unreliable when criteria are semantically heterogeneous. We propose Reinforcement Learning with Verifiable Rubric-based Ranking (RLVR$^2$), a verifiable ranking paradigm for rubric-based RLVR. For each criterion, RLVR$^2$ converts rubric scores into criterion-specific within-group ordinal outcomes, recovers a latent utility from the resulting comparison matrix, and merges these utilities into one training signal. By retaining only within-group ordering and discarding raw score magnitudes, RLVR$^2$ avoids calibrating heterogeneous rubric scales. It further supports objective-preserving attribute adjustment: auxiliary attributes that correlate with observed rankings but are not training objectives can enter the estimation without expanding the rubric or rewarding them directly. Across three model scales and 16 benchmarks, RLVR$^2$ consistently outperforms representative rubric-based baselines, achieving the best overall performance on most benchmarks at every scale. Analysis shows it controls systematic effects tied to reasoning efficiency and response formatting while preserving the quality objective.
Chinese Translation
可验证奖励强化学习(Reinforcement Learning with Verifiable Rewards, RLVR)正从具有明确正确性信号的任务(如数学和代码)扩展到由多维度评分准则(rubrics)规定的多方面质量要求。由于策略优化在每次采样(rollout)中只消耗一个标量,基于评分准则的流程必须将多个准则的得分映射为一个标量奖励。这种聚合通常被视为简单的分数缩放,但它实际上隐式地决定了训练过程中各质量维度之间的权衡。当前主流做法是对每个准则的分数进行归一化后取线性组合,这假设了不同准则之间的基数分数差异具有可比性,且在一个准则上的提升可以补偿在另一个准则上的失败;当各准则在语义上异质时,这两种假设都不可靠。我们提出了基于可验证评分准则排序的强化学习(Reinforcement Learning with Verifiable Rubric-based Ranking, RLVR$^2$),这是一种面向基于评分准则的RLVR的可验证排序范式。对于每个准则,RLVR$^2$将评分准则的分数转化为该准则特定的组内序数结果,从由此产生的比较矩阵中恢复潜在效用,并将这些效用合并为一个训练信号。通过仅保留组内排序信息而丢弃原始分数的数值大小,RLVR$^2$避免了对异质评分准则尺度的校准问题。它还进一步支持目标保持式属性调整:那些与观测到的排序相关但并非训练目标的辅助属性可以在不扩展评分准则或不对其进行直接奖励的情况下纳入估计过程。在三个模型规模和16个基准测试上,RLVR$^2$始终优于具有代表性的基于评分准则的基线方法,在每种规模下于大多数基准上取得了最佳综合性能。分析表明,该方法在保持质量目标的同时,能够控制与推理效率和回复格式相关的系统性效应。
cs.LG / 137 / 2609.23460

TRACE: Tractable Routing Autoencoder for Clinical ECG

TRACE:面向临床心电图的易处理路由自编码器
Jia, Shunbo, Ma, Runze, Lyu, Haonan, Zhang, Haijin, Yang, Qiang, Liao, Caizhi
Abstract
Deep learning has advanced automated electrocardiogram (ECG) diagnosis, but the field's most accurate models, foundation models pretrained on millions of recordings, are not decision-pathway auditable: a clinician cannot trace a diagnosis to a physiological pathway or intervene on one. We propose TRACE, a Tractable Routing Autoencoder for Clinical ECG, whose 32-dimensional clinical latent space is specified in advance from domain knowledge rather than discovered by optimization. TRACE partitions this space into perfusion, structure, and conduction subspaces, routes each to its own diagnostic head by design, regularizes the partition with an orthogonality penalty, and reconstructs the ECG through a decoder that permits latent perturbation. On PTB-XL and Georgia, TRACE exceeds unconstrained classifiers and stays ahead of an ECG foundation model pretrained on ten million recordings, evaluated by linear probe on frozen features, at roughly an eighth of the parameter count. On the nine-label CPSC2018 cohort, which carries no structural class, the framework transfers with only the routing table re-specified to a perfusion/rhythm/conduction partition. Joint probe, erasure, and perturbation analyses verify the routing contract, and perturbing the depolarization and repolarization pathways modulates the reconstructed waveform. Removing the specified partition and its orthogonality penalty costs 1.70 AUC and 11.30 macro-F1 points on PTB-XL, and 2.76 AUC and 16.92 macro-F1 points on Georgia. A capacity-matched permutation control places arbitrary assignments within 0.34 AUC points of the ontology routing and leaves macro-F1 statistically level (p=0.619): the ontology supplies decision-pathway auditability at no macro-F1 cost.
Chinese Translation
深度学习推动了自动化心电图(ECG)诊断的发展,但该领域最精确的模型——在数百万条记录上预训练的基础模型——并不具备决策路径可审计性:临床医生无法将诊断追溯到某一生理通路,也无法对其进行干预。我们提出TRACE(Tractable Routing Autoencoder for Clinical ECG),一种面向临床心电图的易处理路由自编码器,其32维临床潜在空间由领域知识预先指定,而非通过优化发现。TRACE将该空间划分为灌注、结构和传导三个子空间,通过设计将每个子空间路由至各自的诊断头,并以正交惩罚对该划分进行正则化,同时通过一个允许潜在扰动的解码器重建心电图。在PTB-XL和Georgia数据集上,TRACE超越了无约束分类器,并在参数量仅约八分之一的情况下,通过线性探针评估冻结特征时领先于在千万条记录上预训练的ECG基础模型。在不包含结构类别的九标签CPSC2018队列上,该框架仅需将路由表重新指定为灌注/节律/传导划分即可实现迁移。联合探针、擦除和扰动分析验证了路由契约,且扰动去极化和复极化通路会相应调制重建波形。移除预指定的划分及其正交惩罚会使PTB-XL上的AUC下降1.70、宏平均F1下降11.30,在Georgia上AUC下降2.76、宏平均F1下降16.92。容量匹配的置换对照实验表明,任意指定的分配与本体论路由的AUC差距在0.34以内,且宏平均F1在统计上持平(p=0.619):本体论以零宏平均F1代价提供了决策路径可审计性。
cs.LG / 138 / 2609.23516

ITSY: Causal Discovery From Irregular Time-Series Data

ITSY:从不规则时间序列数据中进行因果发现
Xu, Wenbo, He, Yue, Wang, Yunhai, Chen, Yueguo, Kuang, Kun
Abstract
Structural causal models for time series recover contemporaneous and lagged effects, but most methods require complete observation windows and become misspecified when samples are missing. We introduce ITSY, the first continuous-optimization method for causal discovery from irregular time series under a linear model. ITSY reformulates the structural equation so that prediction uses the nearest available history rather than the possibly missing current slice, and jointly imputes missing values while learning both graphs. A weighted reconstruction objective corrects the noise transformation induced by this reformulation. Across synthetic regimes varying missingness, scale, graph density, and noise, and on a real world benchmark, ITSY consistently improves graph recovery over representative SCM-based baselines, demonstrating the effectiveness of the proposed method. The results establish a focused solution for irregular linear first-order dynamics and clarify the assumptions required for nonlinear or higher-order extensions.
Chinese Translation
面向时间序列的结构因果模型可以恢复同期效应和滞后效应,但大多数方法要求完整的观测窗口,当样本缺失时会出现模型误设。我们提出了ITSY,这是首个在时间跨度线性模型下从不规则时间序列中进行因果发现的连续优化方法。ITSY对结构方程进行了重新表述,使预测利用最近可用的历史信息,而非可能缺失的当前时间切片,并在学习两个图的同时联合插补缺失值。一个加权重构目标函数校正了由这种重新表述引入的噪声变换。在不同缺失程度、规模、图密度和噪声的合成场景以及一个真实世界基准测试中,ITSY在图恢复方面持续优于具有代表性的基于SCM的基线方法,验证了所提方法的有效性。这些结果为不规则线性一阶动力学提供了一个有针对性的解决方案,并阐明了向非线性或高阶扩展所需的假设。
cs.LG / 139 / 2609.23521

Feature Suppression and Differential Privacy for Residential Traffic Classification: A Two-Home Federated Study

住宅流量分类的特征抑制与差分隐私:一项双家庭联邦学习研究
Lipcsey-Magyar, Márton Pál, Pekar, Adrian
Abstract
Residential traffic classification supports service management, but learning across homes must account for heterogeneous traffic and privacy constraints. Privacy-aware training may impose uneven costs across traffic categories. We study this tradeoff in simulated two-client federated learning using 1.62 million preprocessed gateway-collected flows across six categories. We compare a full-feature baseline, feature suppression (FS), and differentially private stochastic gradient descent (DP-SGD) under one fixed record-level privacy setting. FS-mild excludes four timing features from 16 model inputs; it provides no formal privacy guarantee. With size-proportional aggregation, FS-mild achieves higher combined macro-F1 and worst-group F1 (the minimum per-class F1 across homes) than DP-SGD in all five seeds at both model capacities under stratified and temporal splits. The tested DP-SGD configuration incurs pronounced minority-category losses, especially in the smaller home, but FS-mild does not uniformly improve on the full-feature baseline. On stratified-split models, loss-based and shadow-model membership probes show near-chance aggregate discrimination without a consistent ranking across probes; this does not establish equivalent privacy. These findings support FS as an input-minimization baseline, not a substitute for formal privacy.
Chinese Translation
住宅流量分类支持服务管理,但跨家庭的机器学习必须考虑异构流量和隐私约束。隐私感知训练可能会对不同流量类别造成不均衡的性能代价。我们在模拟的双客户端联邦学习场景中研究这一权衡问题,使用了共162万条经预处理的网关采集流量,涵盖六个类别。我们在固定的记录级隐私设置下,比较了全特征基线、特征抑制(Feature Suppression, FS)以及差分隐私随机梯度下降(DP-SGD)三种方法。FS-mild从16个模型输入中排除了4个时序特征,但它不提供正式的隐私保证。在按样本量比例加权聚合下,FS-mild在分层划分和时间划分两种数据划分方式、两种模型容量以及全部五个随机种子下,其综合宏平均F1和最差组F1(各家庭中每类F1的最小值)均高于DP-SGD。所测试的DP-SGD配置在少数类别上出现了显著的性能损失,尤其是在数据量较小的家庭中,但FS-mild也并非在所有情况下都优于全特征基线。在分层划分的模型上,基于损失和影子模型的成员推断探测显示总体判别能力接近随机水平,且各探测方法之间不存在一致的排序;但这并不能证明隐私保护水平等价。这些发现支持将FS作为输入最小化基线,而非正式隐私机制的替代方案。
cs.LG / 140 / 2609.23529

Predicting Out-of-Distribution Generalization of Neural Operators via Observable Spectral Error Decomposition

基于可观测谱误差分解预测神经算子的分布外泛化
Dong, Hang-Cheng, Cheng, Pengcheng
Abstract
Neural operators have emerged as powerful surrogates for solving partial differential equations (PDEs), yet their reliability under distribution shift remains a critical barrier to deployment. Existing approaches to out-of-distribution (OOD) generalization in operator learning are largely empirical and black-box: they report aggregate error metrics without explaining why errors arise or when they will grow. We propose a structure-preserving framework that makes OOD generalization predictable and auditable. Our key idea is to parameterize the learned solution operator as a spectral filter $h_\theta(\lambda)$ acting on the eigenvalues of the underlying elliptic operator, implemented via Chebyshev polynomial expansions and trained with a weak-form objective. This parameterization admits an exact decomposition of the energy-norm error into two observable components: a model-dependent spectral approximation term and a distribution-dependent spectral weighting term induced by the input. From this decomposition we derive three diagnostics: a conservative in-band supremum $\vareps_{\mathrm{sup}}$, a global RMS proxy $\vareps_{\mathrm{rms}}$, and a sample-dependent effective metric $\vareps_{\mathrm{eff}}(f)$. These diagnostics can be computed without access to ground-truth solutions. Through four controlled experiments, we show that $\vareps_{\mathrm{eff}}(f)\|f\|$ consistently predicts energy error under in-distribution, in-band spectral shift, out-of-band tail, and compound shifts, whereas global metrics can be systematically misleading. Our framework shifts OOD assessment of neural operators from black-box benchmarking to operator-structure diagnostics, providing a practical route to auditable scientific machine learning.
Chinese Translation
神经算子已成为求解偏微分方程(PDE)的强大代理模型,但其在分布偏移下的可靠性仍是部署的关键障碍。现有关于算子学习中分布外(OOD)泛化的研究大多是经验性的和黑盒式的:它们只报告总体误差指标,而不解释误差为何产生以及何时会增大。我们提出了一个结构保持的框架,使OOD泛化变得可预测且可审计。我们的核心思想是将学习到的解算子参数化为作用于底层椭圆算子特征值的谱滤波器 $h_\theta(\lambda)$,通过切比雪夫(Chebyshev)多项式展开实现,并采用弱形式目标函数进行训练。该参数化使能量范数误差能够被精确分解为两个可观测的组成部分:一个依赖于模型的谱逼近项,以及一个由输入诱导的、依赖于分布的谱加权项。基于该分解,我们导出了三个诊断量:保守的带内上确界 $\vareps_{\mathrm{sup}}$、全局均方根代理 $\vareps_{\mathrm{rms}}$,以及依赖于样本的有效度量 $\vareps_{\mathrm{eff}}(f)$。这些诊断量无需访问真值解即可计算。通过四组受控实验,我们表明 $\vareps_{\mathrm{eff}}(f)\|f\|$ 在分布内、带内谱偏移、带外尾部以及复合偏移情形下均能一致地预测能量误差,而全局指标则可能产生系统性的误导。我们的框架将神经算子的OOD评估从黑盒基准测试转变为基于算子结构的诊断,为可审计的科学机器学习提供了一条切实可行的途径。
cs.LG / 141 / 2609.23535

Decoupled Causal Discovery

解耦因果发现
Guan, Zhengkang, Wu, Fei, Kuang, Kun
Abstract
Causal discovery from observational data is a fundamental yet challenging task in scientific research. While existing approaches are primarily based on conditional independence tests, structure scores, or restrictive functional assumptions, we propose Decoupled Causal Discovery (DCD), a novel decoupling-based perspective that does not rely on these methodologies. DCD directly identifies the Markov boundary (MB) by decoupling non-target variables via weighting functions, such that only variables within the MB preserve dependence with the target under the decoupled distribution. Building on this, DCD iteratively constructs the Completed Partially Directed Acyclic Graph (CPDAG) by exploiting structural asymmetries within the MBs. We establish the theoretical identifiability, soundness, and completeness of DCD. Empirical evaluations demonstrate that DCD achieves strong performance, particularly excelling in challenging noise regimes.
Chinese Translation
从观测数据中进行因果发现是科学研究领域一项基础性且具有挑战性的任务。现有方法主要基于条件独立性检验、结构评分或 restrictive 函数假设,而我们提出了一种新颖的基于解耦视角的方法——解耦因果发现(Decoupled Causal Discovery, DCD),该方法不依赖于上述方法论。DCD 通过加权函数对非目标变量进行解耦,直接识别马尔可夫边界(Markov Boundary, MB),使得在解耦后的分布下,仅马尔可夫边界内的变量与目标变量保持依赖关系。在此基础上,DCD 利用马尔可夫边界内部的结构不对称性,迭代地构建完成的部分有向无环图(Completed Partially Directed Acyclic Graph, CPDAG)。我们建立了 DCD 的理论可识别性、可靠性与完备性。实证评估表明,DCD 取得了优异的性能,尤其在具有挑战性的噪声环境下表现突出。
cs.LG / 142 / 2609.23547

Preserving Geometric Integrity in Graph Prompting via Measure-Constrained Optimal Transport

基于测度约束最优传输的图提示几何完整性保持方法
Wang, Xiangyu, Wang, Shuo, Fang, Ruiyi, Kang, Zhao
Abstract
Graph prompt learning enables parameter-efficient adaptation of frozen Graph Neural Networks to downstream tasks through lightweight prompt parameters. As routing becomes increasingly node-adaptive, however, independently optimized local decisions can collectively concentrate assignment mass on a small subset of a finite shared prompt bank, even when individual node--prompt matches remain locally meaningful. We propose MINT (Measure-INtegrity Transport), an entropically regularized optimal transport framework that formulates node-to-prompt adaptation as a globally coupled allocation problem. The transport cost favors local geometric compatibility, while a prescribed prompt-side marginal explicitly controls graph-wide prompt utilization. We further derive an exact variance decomposition that separates prompt-side geometric variance into retained prompt-update variation and within-node barycentric dispersion, together with a conditional stability bound for the frozen-encoder forward map. Across standard citation networks and additional heterophilic graphs, MINT remains competitive in few-shot adaptation. Controlled and end-to-end experiments further distinguish the roles of routing and topology: fixed-marginal routing controls graph-wide prompt utilization and has measurable end-to-end effects on citation networks, while topology augmentation provides a complementary, graph-dependent mechanism for addressing structural mismatch. Code is available at https://github.com/Ga1axy0051/MINT.
Chinese Translation
图提示学习通过轻量级提示参数,使冻结的图神经网络能够以参数高效的方式适配下游任务。然而,随着路由机制日益趋于节点自适应,独立优化的局部决策可能在整体上将分配质量集中到有限的共享提示库中的小子集上,即使各个节点与提示的匹配在局部仍然是有意义的。我们提出了 MINT(Measure-INtegrity Transport,测度完整性传输),这是一个熵正则化的最优传输框架,将节点到提示的适配形式化为一个全局耦合的分配问题。传输代价偏向于局部几何兼容性,而预设的提示侧边缘分布则显式地控制整个图范围内的提示利用率。我们进一步推导了一个精确的方差分解,将提示侧几何方差分离为保留的提示更新变化和节点内重心离散度,并为冻结编码器的前向映射给出了条件稳定性界。在标准引文网络以及其他异配图上,MINT 在少样本适配中保持了竞争力。受控实验与端到端实验进一步区分了路由与拓扑的作用:固定边缘分布的路由能够控制图范围内的提示利用率,并对引文网络产生可测量的端到端影响;而拓扑增强则提供了一种互补的、依赖图本身的机制,用于应对结构性失配。代码可在 https://github.com/Ga1axy0051/MINT 获取。
cs.LG / 143 / 2609.23549

Physics-residual machine learning predicts oxygen-evolution catalyst activity beyond the training range from sparse polarization measurements

物理残差机器学习从稀疏极化测量中预测超出训练范围的析氧催化剂活性
Kim, Yong-Woon, Lee, Jihyeok, Park, Sungtae, Choi, Sooseok, Byun, Yung-Cheol
Abstract
Screening oxygen-evolution catalysts on combinatorial libraries requires deciding which candidates receive the remaining measurements. The deciding activity lies beyond each candidate's measured potential window and often above every activity recorded during fitting. We predict it by physics-residual machine learning: the Tafel equation extrapolates the candidate's own measured current and slope, a learned residual attenuated with feature-space distance corrects the magnitude, and an applicability-domain score identifies predictions above the training range before measurement. In a separately fabricated 322-candidate library, 282 above the training maximum, two measurements per candidate gave a mean absolute error of 0.203 mA cm$^{-2}$ against 1.330 for the selected data-driven machine-learning model. Errors inside the training range remained comparable, and 35 labelled catalysts were enough to fit it. In two independent datasets the same construction lowered the overpotential error by 29 to 52%. Campaigns can therefore shorten each measurement and still rank the most active compositions.
Chinese Translation
在组合材料库中筛选析氧催化剂需要决定哪些候选材料应获得后续测量。决定性的活性位于每个候选材料已测电位窗口之外,且往往高于拟合期间记录的所有活性。我们通过物理残差机器学习来预测该活性:Tafel方程基于候选材料自身测得的电流和斜率进行外推,一个随特征空间距离衰减的可学习残差对幅值进行校正,而适用域分数则在测量之前识别出超出训练范围的预测。在一个单独制备的包含322个候选材料的材料库中(其中282个高于训练最大值),每个候选材料仅需两次测量即可获得0.203 mA cm$^{-2}$的平均绝对误差,而所选的数据驱动机器学习模型的误差为1.330。训练范围内的误差保持相当水平,且35个有标注的催化剂足以完成拟合。在两个独立数据集上,同样的构建方法将过电位误差降低了29%至52%。因此,此类测量流程可以缩短每次测量的时间,同时仍能对最具活性的组分进行排序。
cs.LG / 144 / 2609.23585

Global Ranks Survive, Selected Heads Shift: BOS-Sink Topology under 4-bit Weight-Only Quantization

全局排序得以保留,而被选中的注意力头发生偏移:4比特仅权重量化下的BOS-Sink拓扑
Chen, Kuanlin, Kuo, Chen-Wei, Ou, Cheng-En
Abstract
Sink-aware deployment may identify important first-token attention heads before a model is quantized, then reuse that map at the edge. We test when this shortcut is safe for 4-bit NF4 weight-only post-training quantization (PTQ). Our Sink Topology Consistency (STC) metrics separate global rank preservation, top-$k$ set overlap, and layerwise sink-mass shift, and distinguish per-input sensitivity from calibration-map transfer. Across Qwen2.5-0.5B, Qwen2.5-1.5B, and Llama-3.2-1B, global bf16-to-4-bit ranks remain high at 4,096 tokens ($\rho_s \geq 0.980$), yet top-$k$ Jaccard overlap is only 0.619-0.793, corresponding to 76.5-88.5% membership retention. The global statistic also masks local failures: terminal Qwen layers shift by 6.2-7.9x their model means, whereas Llama-3.2-1B shows low, nearly uniform drift. Under a C4-to-LongBench shift, cross-domain overlap degrades more than the within-domain precision comparison for both Qwen models, but not for Llama-3.2-1B. Matched-domain 4-bit recalibration reaches 90% of a split-half stability plateau at the smallest tested $n=8$ for both Qwen models and $n=32$ for Llama-3.2-1B, though not as a sharp threshold; for the two Qwen models, updating only selected layers does not reach the full-map stability criterion. On Jetson Orin NX, the 16-sample workload takes seconds for the two models with valid on-device sink measurements. The practical message is precise: global rankings often transfer, but discrete head sets, layer-local policies, and cross-domain calibration should be revalidated after quantization.
Chinese Translation
Sink感知的部署方式可以在模型量化之前识别重要的首token注意力头,然后在边缘端复用该映射。我们检验了这一捷径在4比特NF4仅权重训练后量化(PTQ)下何时是安全的。我们的Sink拓扑一致性(STC)度量将全局排序保持性、top-$k$集合重叠以及逐层sink质量偏移区分开来,并区分了逐输入敏感性与校准映射迁移。在Qwen2.5-0.5B、Qwen2.5-1.5B和Llama-3.2-1B上,从bf16到4比特的全局排序在4,096个token时仍然保持很高($\rho_s \geq 0.980$),然而top-$k$ Jaccard重叠仅为0.619-0.793,对应76.5-88.5%的成员保留率。全局统计量还掩盖了局部失效:Qwen末端各层的偏移达到其模型均值的6.2-7.9倍,而Llama-3.2-1B则表现出较低且近乎均匀的漂移。在从C4到LongBench的分布偏移下,两个Qwen模型的跨域重叠均比域内精度比较下降得更严重,但Llama-3.2-1B并非如此。在匹配域的4比特重新校准中,两个Qwen模型在最小测试规模$n=8$、Llama-3.2-1B在$n=32$时即可达到对半折稳定性平台的90%,尽管这并非一个明显的阈值;对于两个Qwen模型,仅更新选定层无法达到全映射的稳定性标准。在Jetson Orin NX上,16个样本的工作负载在两个模型上仅需数秒即可完成有效的设备端sink测量。实践结论是明确的:全局排序通常可以迁移,但离散的注意力头集合、层局部策略和跨域校准应在量化之后重新验证。
cs.LG / 145 / 2609.23590

Cost-Aware Reinforcement Learning with Action Masking and Projection for Battery Energy Storage Dispatch under Suppressed-Spread Market Shifts

面向抑制价差市场变化下电池储能调度的成本感知强化学习:动作掩码与投影方法
Chen, Kuanlin, Kuo, Chen-Wei, Ou, Cheng-En
Abstract
Battery energy storage system (BESS) dispatch must preserve operational feasibility while declining price spreads reduce the margin available to pay for cycling. We study a proximal policy optimization (PPO) controller whose pre-selection physical action mask and emergency projection are separated from a causal, forecast-informed economic advisory. All forecast-dependent methods receive the same causal 24-step forecast and grid-side settlement. Across five PPO seeds, advice-on net profit is 30.59 and 18.04 USD per 336-hour T1 and T2 window, versus 36.77 and 22.94 USD for proxy-cost MPC; PPO remains below this reference in both periods. Advice raises T2 profit from 16.45 to 18.04 USD while reducing throughput, but is immaterial in T1. On disjoint weekly blocks, PPO is stable under daily, weekly, and blended seasonal forecasts, weakens under persistence, and remains below proxy-cost MPC. Paired diagnostics localize changes to the observed 5-10 USD/MWh regime with mixed SoC-dependent effects. An M0-M6 ablation shows that mask removal sends thousands of infeasible requests to projection, while removing both physical layers exposes ramp violations. The evidence separates economic screening from feasibility enforcement without claiming formal safety, lifecycle-optimal aging, or RL dominance.
Chinese Translation
电池储能系统(BESS)的调度必须在价格价差下降、可用于支付充放电循环成本的利润空间缩小的同时,保持运行可行性。我们研究了近端策略优化(PPO)控制器,其预筛选的物理动作掩码与紧急投影机制同基于因果、考虑预测信息的经济咨询相分离。所有依赖预测的方法均接收相同的因果24步预测和电网侧结算。在五个PPO随机种子下,开启经济咨询的净利润在336小时的T1和T2窗口分别为30.59和18.04美元,而代理成本模型预测控制(MPC)分别为36.77和22.94美元;PPO在两个时期均低于该基准。经济咨询将T2的利润从16.45美元提升至18.04美元,同时降低了吞吐量,但在T1中影响甚微。在不相交的周度数据块上,PPO在日度、周度和混合季节性预测下保持稳定,在持续性预测下性能下降,且始终低于代理成本MPC。成对诊断将性能变化定位于观测到的5-10美元/兆瓦时(USD/MWh)价差区间,并呈现依赖于荷电状态(SoC)的混合效应。M0-M6消融实验表明,移除动作掩码会使数千个不可行请求涌入投影模块,而同时移除两个物理层则会暴露爬坡率违规问题。实验证据将经济筛选与可行性强制执行区分开来,但并未声称具备形式化安全性、生命周期最优的老化管理或强化学习的优越性。
cs.LG / 146 / 2609.23594

Bilinear Optimization Divergence: Diagnosing Factor-Constrained LoRA Continual Learning

双线性优化散度:诊断因子约束的LoRA持续学习
Wang, YongShun, Su, JianLin, Ma, Yong
Abstract
Orthogonality in a LoRA factor does not by itself specify what the composed update protects: the answer depends on the task-start state, the parameterization, and the realized optimizer displacement. We formalize this question through Bilinear Optimization Divergence (BOD), an anchor-relative diagnostic of effective-update response on selected historical features. The finite-step analysis distinguishes two cases. In a shared adapter, protecting the routing displacement leaves a learned-anchor residual through the changing companion factor. In a fresh zero-output block, a feasible routing state can protect the composed update while both current factors remain trainable. These conditions yield Semi-Frozen Orthogonal Routing (SFOR) for shared adapters and current-block hard protection for cumulative O-LoRA; Weight Residual Projection (WRP) enforces the required displacement after the optimizer step. Controlled two-task traces verify the predicted residual paths, reducing normalized historical response from 19.12% to 0.005% in the shared family and from 7.72% to 0.002% in the cumulative family. Four-task experiments on Qwen3-8B characterize the resulting trade-offs: SFOR improves backward transfer (BWT) from -2.47 to -0.86 with nearly unchanged average accuracy (AA), while O-LoRA hard protection improves three-order mean AA from 80.27% to 81.30% and forgetting measure (FM) from 2.20 to 0.43. Component controls also show that stricter feasibility need not improve final task performance. Together, the analysis and evidence provide an architecture-conditioned account of which constraint to enforce, how to enforce it, and how to interpret its empirical value.
Chinese Translation
LoRA因子的正交性本身并不能说明组合更新所保护的内容:答案取决于任务起始状态、参数化方式以及优化器实际产生的位移。我们通过双线性优化散度(Bilinear Optimization Divergence, BOD)将这一问题形式化,它是一种相对于锚点的诊断量,用于衡量在选定历史特征上的有效更新响应。有限步分析区分了两种情形。在共享适配器中,保护路由位移仍会通过不断变化的伴随因子留下学习锚点残差。在一个全新的零输出块中,一个可行的路由状态可以在两个当前因子均可训练的同时保护组合更新。这些条件导出了面向共享适配器的半冻结正交路由(Semi-Frozen Orthogonal Routing, SFOR)以及面向累积式O-LoRA的当前块硬保护;权重残差投影(Weight Residual Projection, WRP)在优化器步进之后强制实施所需的位移。受控的双任务轨迹验证了所预测的残差路径,将归一化历史响应在共享家族中从19.12%降至0.005%,在累积家族中从7.72%降至0.002%。在Qwen3-8B上的四任务实验刻画了由此产生的权衡:SFOR将后向迁移(BWT)从-2.47提升至-0.86,同时平均准确率(AA)几乎不变;O-LoRA硬保护将三阶平均AA从80.27%提升至81.30%,将遗忘度量(FM)从2.20降至0.43。组件控制实验还表明,更严格的可行性并不一定能提升最终任务性能。综合而言,本分析与实验证据提供了一个以架构为条件的解释框架,说明应施加何种约束、如何施加约束,以及如何解读其经验价值。
cs.LG / 147 / 2609.23659

ETH-TraceBench: A Large-Scale Event-Stream Benchmark for Ethereum DeFi under Temporal, Protocol, and Contract Shift

ETH-TraceBench:面向时间、协议与合约分布偏移下以太坊DeFi的大规模事件流基准
Kirtac, Kemal, Maple, Carsten
Abstract
Ethereum decentralized finance (DeFi) provides a public, time-stamped record of transaction-level event streams, but the same public symbols can create strong machine-learning shortcuts. We introduce ETH-TraceBench, a benchmark for evaluating Ethereum DeFi representations under temporal, protocol, pool/infrastructure, and symbolic shift. The raw event universe covers January 2021-December 2025 and contains 1.35 billion transactions with logs and 5.01 billion raw log rows. Model evaluation uses a fixed 911,267-instance supervised sample, training on 2021-2024, selecting models on 2025H1, and testing on 2025H2. Simple models perform strongly on the aggregate temporal test: TraceStats-GB reaches 0.953 macro-F1 and TopicEmitterHashMLP 0.959 on the canonical DEX test set. Performance drops sharply under protocol novelty, with macro-F1 of 0.794, 0.743, and 0.766 for TraceStats-GB, TopicEmitterTrace-SGD, and TopicEmitterHashMLP, while strict unseen-pool scores remain 0.927, 0.897, and 0.935. Uniswap v4 and Ekubo v1, both absent from supervised training, are materially harder than the full test. Jointly masking emitter and topic identity reduces DEX macro-F1 to 0.916 and liquidation macro-F1 to 0.774 for TopicEmitterTrace-SGD. A standard Transformer over log-index-ordered events provides no consistent advantage over a deterministic shuffle of the same events, indicating that high aggregate scores can arise without sophisticated chronological modeling. A natural-prevalence audit estimates 2025H2 DEX prevalence among logged Ethereum transactions at about 22.5%, and a deterministic 400-transaction audit finds complete agreement with task label sources and independently re-queried raw-log counts. ETH-TraceBench therefore treats difficult transfer and controlled-input conditions, rather than a single aggregate score, as the main evaluation target.
Chinese Translation
以太坊去中心化金融(DeFi)提供了公开的、带时间戳的交易级事件流记录,但相同的公开符号可能造成较强的机器学习捷径。我们提出ETH-TraceBench,一个用于评估以太坊DeFi表示在时间偏移、协议偏移、流动性池/基础设施偏移以及符号偏移下的基准。原始事件全集覆盖2021年1月至2025年12月,包含13.5亿笔带日志的交易和50.1亿条原始日志行。模型评估采用固定的911,267个实例的监督样本,在2021至2024年数据上训练,在2025年上半年选择模型,在2025年下半年测试。简单模型在聚合时间测试上表现强劲:TraceStats-GB在标准DEX测试集上达到0.953的宏F1,TopicEmitterHashMLP达到0.959。在协议新颖性下性能急剧下降,TraceStats-GB、TopicEmitterTrace-SGD和TopicEmitterHashMLP的宏F1分别为0.794、0.743和0.766,而严格未见池的得分仍保持在0.927、0.897和0.935。Uniswap v4和Ekubo v1均未出现在监督训练中,其难度显著高于完整测试集。对TopicEmitterTrace-SGD而言,同时掩蔽发射器(emitter)与主题(topic)身份使DEX宏F1降至0.916,清算宏F1降至0.774。基于日志索引排序事件的标准Transformer相比对相同事件的确定性乱序并无一致优势,这表明高聚合得分可以在没有精细时序建模的情况下取得。一项自然流行度审计估计,2025年下半年DEX在以太坊带日志交易中的占比约为22.5%;一项确定性的400笔交易审计发现与任务标签来源及独立重新查询的原始日志计数完全一致。因此,ETH-TraceBench将困难的迁移场景与受控输入条件(而非单一聚合分数)作为主要评估目标。
cs.LG / 148 / 2609.23686

One Patch, Three Roles: What Is Actually Coupled in Autoregressive Time-Series Forecasting?

一个补丁,三种角色:自回归时间序列预测中真正耦合的是什么?
Li, Ziang, Huang, Yue, Zhou, Guoxu, Han, Na, Wen, Jie, Fei, Lunke, Fang, Xiaozhao
Abstract
Patch-based autoregressive time-series forecasting often ties input representation, learned transitions, and recursive execution to one patch length. We ask which of these roles can be adjusted separately. A supporting atomic-encoding study finds greater sensitivity to model width than to atom grouping on the evaluated grid. Our main finding is that a frozen parent's recursive trajectory is easier to fit than the observed future with lightweight parallel exits. Autoregressive Trajectory Distillation (ATD) turns this into selectable ATD-1/2/4/8 execution, with ATD-1 exactly recovering the parent. On a paired four-data-set comparison, ATD-8 reaches $5.54\times$ end-to-end speedup with stable quality across widths. Fewer calls do not automatically remove the parent's existing forecast error: ATD improves trajectory fidelity in all 21 seed runs but forecast accuracy in only 15 against matched clean-future supervision. We further find a correctable residual projection along a train-selected periodic history direction. Spectrum Tangent applies this correction without adding neural parameters or Transformer calls. At horizon 720, it reduces mean squared error (MSE) and mean absolute error (MAE) by 2.54% and 2.33% over seven data sets and two output widths, while remaining $3.24\times$ faster than recursive inference. Level and shape projections sometimes disagree. Trajectory compressibility, the fidelity-accuracy mismatch, and the correction recur across three public AR parents. Together these results separate representation, transition, and execution as AR design axes. Code is available at https://github.com/RowanFFF/ATD-Spectrum-Tangent.
Chinese Translation
基于补丁(Patch)的自回归时间序列预测通常将输入表示、学习到的状态转移和递归执行绑定到同一个补丁长度上。我们探究这些角色中的哪些可以被独立调整。一项支持性的原子编码研究发现,在所评估的参数网格上,模型宽度的影响比原子分组方式更为显著。我们的主要发现是:借助轻量级并行退出机制,冻结父模型的递归轨迹比观测到的未来序列更容易拟合。自回归轨迹蒸馏(Autoregressive Trajectory Distillation, ATD)将这一发现转化为可选择执行步数的 ATD-1/2/4/8,其中 ATD-1 可精确复现父模型。在配对的四数据集对比实验中,ATD-8 实现了 5.54 倍的端到端加速,且在不同模型宽度下保持稳定的质量。减少调用次数并不能自动消除父模型已有的预测误差:在全部 21 个随机种子运行中,ATD 均提升了轨迹保真度,但相对于匹配的干净未来序列监督,仅在 15 个运行中提升了预测精度。我们进一步发现存在一个可通过训练所选的周期性历史方向进行校正的残差投影。谱切线(Spectrum Tangent)方法无需增加神经参数或 Transformer 调用即可施加该校正。在预测步长 720 处,它在七个数据集和两种输出宽度上将均方误差(MSE)和平均绝对误差(MAE)分别降低了 2.54% 和 2.33%,同时仍比递归推理快 3.24 倍。水平投影与形状投影有时并不一致。轨迹可压缩性、保真度与精度的失配以及该校正方法在三个公开的 AR 父模型上均反复出现。这些结果共同将表示、转移和执行分离为自回归设计的三个独立维度。代码可在 https://github.com/RowanFFF/ATD-Spectrum-Tangent 获取。
cs.LG / 149 / 2609.23687

A multi-temporal dataset for mapping burned areas in the Brazilian Cerrado using time series of remote sensing imagery

利用遥感影像时间序列绘制巴西塞拉多(Cerrado)火烧区域的多时相数据集
de Oliveira, Alisson Cleiton, Körting, Thales Sehn
Abstract
This paper introduces a multi-temporal tabular dataset derived from satellite images to map burned areas in the Chapada dos Veadeiros National Park, in Goi\'as, Brazil, covering the years 2020 to 2022. The dataset contains blue, green, red, and near-infrared bands, as well as the BAI, EVI, GEMI, NDVI, and NDWI spectral indices from the WFI sensor on the CBERS-4A, CBERS-4, and AMAZONIA-1 satellites, organized into a regular grid. We applied the Random Forest classifier to develop and validate models based on samples labeled as totally burned, partially burned, and non-burned. Two classification approaches were tested: one combining burned and non-burned areas into binary classes and another distinguishing between totally burned (TB), partially burned (PB), and non-burned (NB) classes. Seven validation approaches assessed different post-classification combinations, focusing on accuracy, precision, recall, and intersection over union (IoU) metrics. Results showed higher IoU when TB, PB, and NB were used as individual classes and TB was reclassified as burned area (BA) while PB and NB were grouped as non-burned. Comparing the annual results of this approach to the MCD64A1 product, the errors of omission for the BA class were 22% in 2020, 28% in 2021 and 59% in 2022, while the errors of commission were 46%, 43% and 46%, respectively. The study highlights the utility of the WFI sensor for burned area mapping without inter-satellite spectral calibration and suggests further exploration with other machine learning algorithms to evaluate the dataset potential and limitations.
Chinese Translation
本文介绍了一个由卫星影像提取的多时相表格数据集,用于绘制巴西戈亚斯州韦阿代鲁斯高地国家公园(Chapada dos Veadeiros National Park)2020年至2022年的火烧区域。该数据集包含蓝、绿、红和近红外波段,以及来自CBERS-4A、CBERS-4和AMAZONIA-1卫星WFI传感器的BAI、EVI、GEMI、NDVI和NDWI光谱指数,并以规则网格形式组织。我们应用随机森林(Random Forest)分类器,基于标注为完全火烧、部分火烧和非火烧的样本开发并验证了模型。测试了两种分类方法:一种将火烧与非火烧区域合并为二元类别,另一种则区分完全火烧(TB)、部分火烧(PB)和非火烧(NB)三个类别。采用七种验证方法评估了不同的分类后组合,重点关注精度、精确率、召回率和交并比(IoU)指标。结果表明,当TB、PB和NB作为独立类别,且将TB重新分类为火烧区域(BA),同时将PB和NB归为非火烧区域时,IoU更高。将该方法的年度结果与MCD64A1产品进行比较,BA类的漏检误差在2020年为22%,2021年为28%,2022年为59%,而错检误差分别为46%、43%和46%。本研究凸显了WFI传感器在无需星间光谱校准情况下进行火烧区域制图的实用性,并建议进一步探索其他机器学习算法,以评估该数据集的潜力与局限性。
cs.LG / 150 / 2609.23688

Tail-Weight Control and Localized Generalization in Nearly Low-Rank Adversarial Classification

近似低秩对抗分类中的尾部权重控制与局部化泛化
Wang, Kunyu, Wang, Dehan, Chen, Wenjun
Abstract
We study norm-constrained linear classification under Eu clidean adversarial perturbations in a Gaussian model with a low-dimen sional informative subspace and an independent noise tail. For bounded ramp loss, we prove that a principal-space witness with risk below one half forces every near-optimal predictor to have small tail weight. A path-specific density bound yields constants without requiring positive tail variance. Under isotropic principal covariance, we establish a unique population minimizer and joint local growth. Boundary normalization then removes the common attack penalty from centered margins, giving localized finite-sample guarantees governed by principal dimension and total tail energy. Globalized growth removes the entrance condition at weaker constants; a model-aware comparison retains local guarantees. Experiments with twenty paired repetitions show decreasing excess risk and tail use with sample size, and nearly unchanged behavior when tail dimension grows at fixed total energy. Pure-noise controls and optimizer diagnostics clarify the scope and limitations of these conclusions.
Chinese Translation
我们研究了在具有低维信息子空间与独立噪声尾部的高斯模型下,欧氏对抗扰动约束的范数受限线性分类问题。对于有界的斜坡损失(ramp loss),我们证明了一个风险低于二分之一的主子空间见证向量(witness)会迫使每个近最优预测器具有较小的尾部权重。一个基于路径的密度界在无需正尾部方差的条件下即可给出常数。在主子空间各向同性协方差的假设下,我们建立了唯一的总体极小化子及联合局部增长性。随后,边界归一化消除了中心化间隔中常见的攻击惩罚,从而给出由主维度和总尾部能量决定的局部化有限样本保证。全局化增长以更弱的常数去除了入口条件;模型感知的比较则保留了局部保证。基于二十组配对重复实验的结果表明,超额风险与尾部使用随样本量增加而下降,且在固定总能量下尾部维度增长时行为几乎不变。纯噪声对照实验与优化器诊断阐明了这些结论的适用范围与局限性。
cs.LG / 151 / 2609.23761

GenVoid: Uncertainty-Aware Learning of Subsurface Material Defects with an Experimentally Validated Physics-Informed Generative Model

GenVoid:基于实验验证的物理信息生成模型的不确定性感知内部材料缺陷识别
Mondal, Trishit, Bharadwaj, Prajwal, Karanjgaokar, Nikhil, Jagtap, Ameya D.
Abstract
Internal voids are ubiquitous defects in manufactured structures, yet their characterization remains challenging because their geometry is hidden and can only be inferred indirectly from accessible measurements. Here we introduce \textit{GenVoid}, a physics-informed generative model-based framework for identifying internal voids in complex two- and three-dimensional solids from surface displacement measurements alone. By incorporating the governing mechanics into a generative inference framework, \textit{GenVoid} enables void identification across linear elastic, hyperelastic and plastic material behaviours and accommodates complex two- and three-dimensional structural geometries. Importantly, the framework explicitly accounts for uncertainty and noise in displacement measurements, producing probabilistic reconstructions of internal void geometry rather than a single deterministic estimate. We demonstrate the approach using high-fidelity synthetic datasets and experimentally measured displacement fields obtained from in-situ mechanical experiments, establishing its ability to infer hidden voids from realistic displacement measurements. To quantify the fundamental limits of such inference, we further introduce an observability measure that characterizes the sensitivity of boundary measurements to localized stiffness perturbations within the interior under an ensemble of applied loads. This framework provides a direct connection between defect location, sensor configuration and reconstruction fidelity, enabling systematic assessment of how the number and spatial distribution of boundary measurements govern void-identification accuracy. To this end, these results establish a physics-informed and uncertainty-aware approach for non-invasive characterization of hidden defects and provide a quantitative basis for designing measurement strategies for inverse problems in solid mechanics.
Chinese Translation
内部孔隙是制造结构中普遍存在的缺陷,但由于其几何形状是隐藏的,只能通过可获取的测量数据间接推断,其表征仍然具有挑战性。本文提出了 GenVoid,一种基于物理信息生成模型的框架,仅利用表面位移测量即可识别复杂二维和三维固体中的内部孔隙。通过将控制力学方程融入生成式推断框架,GenVoid 能够在线弹性、超弹性和塑性材料行为下进行孔隙识别,并适用于复杂的二维和三维结构几何形状。重要的是,该框架显式地考虑了位移测量中的不确定性和噪声,产生内部孔隙几何形状的概率化重构,而非单一的确定性估计。我们使用高保真合成数据集以及从原位力学实验中获得的实测位移场对该方法进行了验证,证明了其从真实位移测量中推断隐藏孔隙的能力。为量化此类推断的基本极限,我们进一步引入了一种可观测性度量,用于表征在加载载荷集合下边界测量对内部局部刚度扰动的敏感性。该框架在缺陷位置、传感器配置与重构保真度之间建立了直接联系,能够系统评估边界测量的数量和空间分布如何决定孔隙识别的精度。由此,这些结果建立了一种物理信息驱动、具有不确定性感知能力的隐藏缺陷无损表征方法,并为固体力学反问题中测量策略的设计提供了定量依据。
cs.LG / 152 / 2609.23775

Statistical Convergence of Transformer Encoder-Accelerated Robust Reinforcement Learning

Transformer编码器加速的鲁棒强化学习的统计收敛性
Banerjee, Suman, Tsukamoto, Hiroyasu
Abstract
Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces. In this study, we present statistically rigorous convergence results for a robust reinforcement learning algorithm warm-started by a transformer-based action-value function prediction, where natural language prompts encode task specifications. Our framework adopts the R-contamination model to characterize uncertainty in the state transition kernel, and employs conformal prediction to certify convergence via trajectory-level nonconformity scores constructed from the contracting Bellman residual. The resulting conformal quantile bounds the gap between the running and optimal action-value functions simultaneously over all iterations, thereby yielding a pre-certified stopping rule that requires little knowledge of the true transition kernel. Numerical case studies on perturbed maze environments of varying size and contamination level confirm that the transformer-based warm start measurably reduces the initial error and accelerates convergence, while the proposed conformal bounds track the true error trajectory more tightly than existing guarantees.
Chinese Translation
在马尔可夫决策过程中,获取最优动作价值函数在大规模状态-动作空间下计算开销极高。本研究针对一种由基于Transformer的动作价值函数预测(其中自然语言提示用于编码任务规范)进行热启动的鲁棒强化学习算法,给出了严格的统计收敛性结果。我们的框架采用R-污染(R-contamination)模型来刻画状态转移核中的不确定性,并利用共形预测(conformal prediction),通过由压缩Bellman残差构造的轨迹级非一致性分数来认证收敛性。所得到的共形分位数可同时界定所有迭代中当前动作价值函数与最优动作价值函数之间的差距,从而提供一种几乎无需了解真实转移核的预先认证的停止规则。在不同规模和污染水平的扰动迷宫环境上的数值案例研究表明,基于Transformer的热启动可显著降低初始误差并加速收敛,同时所提出的共形界比现有保证更紧密地追踪真实误差轨迹。
cs.LG / 153 / 2609.23780

Falling Trees: A Model Class for Interpretable Risk Prioritization

倒落树:一类用于可解释风险优先级排序的模型
Babbar, Varun, Boner, Zachery, Seltzer, Margo, Rudin, Cynthia
Abstract
Many real-world decisions require prioritizing high-risk cases, such as clinicians prioritizing high-risk patients before lower-risk ones. Falling rule lists (FRLs), which are ordered if--then rules with monotonically decreasing risks, provide an interpretable framework for such tasks; however, their single-path structure yields a highly restricted model class. We introduce falling trees, a new family of interpretable models that enforces the same monotonic risk constraint while permitting tree-structured branching. We present GRAVITree, a novel dynamic-programming-with-bounds algorithm for learning the Rashomon set of falling trees under depth and branching constraints. Our formulation can interpolate between rule lists and full decision trees, enabling user-desired model expressivity. In a new clinical dataset and in many public classification benchmarks, falling trees match or outperform FRLs and other interpretable baselines, often producing more sparse decisions for high-risk instances. Our results show that falling trees strike a practical balance between interpretability, expressiveness, and risk prioritization for high-stakes settings.
Chinese Translation
许多现实世界的决策需要对高风险案例进行优先级排序,例如临床医生在低风险患者之前优先处理高风险患者。倒落规则列表(Falling Rule Lists,FRL)是一组风险单调递减的有序“如果—那么”规则,为此类任务提供了一个可解释的框架;然而,其单一路径结构导致模型类别受到极大限制。我们提出了倒落树(falling trees),这是一类新的可解释模型,在允许树状分支结构的同时保持了相同的单调风险约束。我们提出了GRAVITree,这是一种结合边界的新颖动态规划算法,用于在深度和分支约束下学习倒落树的Rashomon集合。我们的公式化方法可以在规则列表与完整决策树之间进行插值,从而实现用户期望的模型表达能力。在一个新的临床数据集以及多个公开分类基准上,倒落树的表现与FRL及其他可解释基线方法相当或更优,并且对于高风险实例往往能产生更稀疏的决策。我们的结果表明,倒落树在高风险场景下于可解释性、表达能力与风险优先级排序之间取得了实用的平衡。
cs.LG / 154 / 2609.23789

Belted Engression: Sufficient Dimension Reduction for Generative Distributional Regression

带束约束的Engression:面向生成式分布回归的充分降维方法
Tan, Wenxi, Li, Bing, Xue, Lingzhou
Abstract
Modern conditional generative models face significant challenges when learning complex covariate dependencies. While sufficient dimension reduction (SDR) provides a principled approach to compress these dependencies, traditional SDR frameworks were not formulated for conditional generation. To bridge this gap, we propose Belted Engression, a unified and architecturally parameter-efficient framework for generative distributional regression. Our approach establishes an end-to-end compress-then-generate paradigm driven by sufficient representation learning, embedding a structural bottleneck into the generative architecture. Theoretically, we prove that the standard SDR condition is equivalent to a law-preserving generative factorization, which is achieved at the global optimum of the population Belted Engression objective. Furthermore, by uncovering a localized Bernstein-type control for the energy-score loss, we establish finite-sample convergence rates that are sharper than those of existing results. We also prove that this belted architecture is strictly smaller, operating with an asymptotically vanishing parameter count relative to the unstructured baseline. Extensive simulations and real-world applications demonstrate that Belted Engression achieves superior distributional prediction and SDR recovery with fewer trainable parameters.
Chinese Translation
现代条件生成模型在学习复杂协变量依赖关系时面临重大挑战。尽管充分降维(Sufficient Dimension Reduction, SDR)为压缩这些依赖关系提供了一种有原则的方法,但传统的SDR框架并非为条件生成而设计。为弥合这一差距,我们提出了Belted Engression,一个统一且架构上参数高效的生成式分布回归框架。我们的方法建立了一个由充分表示学习驱动的端到端“先压缩后生成”范式,在生成架构中嵌入了一个结构化瓶颈。在理论方面,我们证明了标准SDR条件等价于一种保持分布的生成式分解,并且该分解在总体Belted Engression目标的全局最优点处得以实现。此外,通过揭示能量得分损失的局部化Bernstein型控制,我们建立了比现有结果更精细的有限样本收敛速率。我们还证明了这种带束架构严格更小,其参数量相对于无结构基线渐近趋近于零。大量的模拟实验和实际应用表明,Belted Engression能够以更少的可训练参数实现卓越的分布预测和SDR恢复效果。
cs.LG / 155 / 2609.23812

Iterative Atom Refinement: A Monotonicity Principle for Dictionary Learning

迭代原子精化:字典学习的一个单调性原理
Christie, Alexander, Moscoso, Miguel, Novikov, Alexei, Papanicolaou, George, Tsogka, Chrysoula
Abstract
Dictionary learning seeks to recover an unknown dictionary $A$ from observations ${\bf y}_i = A{\bf x}_i$ with sparse coefficient vectors ${\bf x}_i$. We introduce the \emph{Iterative Atom Refinement} (IAR) algorithm, a simple procedure for recovering individual dictionary atoms. Starting from a random direction, IAR repeatedly selects the observations most strongly correlated with the current iterate and updates the direction by averaging the selected data. Our main contribution is a rigorous convergence theory of IAR. Using high-dimensional probabilistic estimates and a novel monotonicity principle for atom-selection probabilities, we show that a small initial advantage of one atom is amplified until that atom is isolated. Under our model assumptions, IAR identifies a generating atom after only three refinement steps. Numerical experiments support the theory and show that the resulting dynamics accurately capture the behavior observed in dictionary refinement.
Chinese Translation
字典学习旨在从观测数据 ${\bf y}_i = A{\bf x}_i$ 中恢复未知字典 $A$,其中系数向量 ${\bf x}_i$ 是稀疏的。我们提出了迭代原子精化(Iterative Atom Refinement, IAR)算法,这是一种用于恢复单个字典原子的简单方法。IAR 从一个随机方向出发,反复选择与当前迭代方向相关性最强的观测数据,并通过对所选数据的平均来更新方向。我们的主要贡献是为 IAR 建立了严格的收敛理论。利用高维概率估计以及一种针对原子选择概率的新型单调性原理,我们证明:某一原子的微小初始优势会被不断放大,直至该原子被成功分离。在本文的模型假设下,IAR 仅需三步精化即可识别出一个生成原子。数值实验验证了该理论,并表明所得动力学能够准确刻画字典精化过程中观察到的行为。
cs.LG / 156 / 2609.23826

Real-time Generalizable Heart Valve Mechanics for Clinical Disease Assessment via a Physics-Conditioned Neural Operator

基于物理条件神经算子的实时可泛化心脏瓣膜力学临床疾病评估
Koohy, Shawn, Wu, Wensi, Jolley, Matthew A, Perdikaris, Paris
Abstract
Mitral regurgitation is the most common heart valve disorder worldwide, affecting over 2% of the global population, rising to at least 10% in adults over 75, and causing approximately 15% of valvular heart disease-related deaths. Yet only a minority of patients with severe disease undergo corrective surgery. Rapid assessment of valve mechanics could enable earlier, more precise intervention, but traditional finite element simulations remain too slow for clinical timelines and parameter sweeps. We introduce the Physics-Conditioned Neural Operator (PCNO), a transformer-based surrogate that predicts leaflet displacement, strain, and stress fields across mitral and tricuspid geometries, conditioned on systolic blood pressure and tissue properties. Trained on functional, regurgitated, and pathological valves, including tethering, P2 prolapse, and annular dilation, PCNO achieves up to a 15,260x speedup over fine mesh finite element simulations with comparable accuracy, identifies pathology class, and resolves diagnostic metrics within 3.5% error under out-of-distribution extrapolation.
Chinese Translation
二尖瓣反流是全球最常见的心脏瓣膜疾病,影响超过2%的全球人口,在75岁以上成年人中这一比例升至至少10%,并造成约15%的瓣膜性心脏病相关死亡。然而,仅有少数重度患者接受矫正手术。快速评估瓣膜力学有望实现更早、更精确的干预,但传统有限元模拟对于临床时间线和参数扫描而言仍过于缓慢。我们提出了物理条件神经算子(Physics-Conditioned Neural Operator, PCNO),这是一种基于Transformer的代理模型,能够以收缩压和组织特性为条件,预测二尖瓣和三尖瓣几何形状下的瓣叶位移、应变和应力场。PCNO在功能性、反流性和病变瓣膜(包括腱索牵拉、P2脱垂和瓣环扩张)上训练后,相比细网格有限元模拟实现了高达15,260倍的加速且精度相当,可识别病理类别,并在分布外外推条件下以3.5%以内的误差求解诊断指标。
cs.LG / 157 / 2609.23836

Actionable Insights from Observational Data: The Case of Advanced Classes in K-12 Education

从观察性数据中获取可行动的洞见:以K-12教育中的高阶课程为例
Bajwa, Nabit, Hunter, Seth B., Das, Sanmay
Abstract
A fundamentally challenging question in K-12 education is about the effects of taking more advanced or challenging classes. It is particularly complex because students (and/or their parents) choose whether to enroll in these classes, making causal analysis challenging. In this paper, we begin to tackle this question by taking advantage of a novel dataset from a public school system in the US. This dataset records students' course enrollment decisions, prior academic histories, demographics, and subsequent outcomes around the time of a district-wide change that introduced optional open-enrollment advanced middle-school courses in subject areas. This is a rich observational dataset, but enrollment in advanced classes is driven by student characteristics and choices rather than random assignment. This creates a core identification challenge: the same factors that influence enrollment in advanced courses are also predictive of academic outcomes. As a result, simple comparisons between enrolled and non-enrolled students are confounded, and naive estimates may reflect underlying differences in student ability, motivation, or support rather than the impact of coursework itself. Our analysis shows that enrolling in advanced English courses has a net positive but modest effect on student achievement outcomes. However, these benefits are unevenly distributed: some students with relatively large predicted gains ("middle achievers" in prior years) are less likely to enroll than others. Some other groups (e.g. Black students and those with lower socio-economic status) also demonstrate significantly lower propensity to enroll. This gap between predicted benefit and observed enrollment illustrates how careful data analysis can extract actionable insights from large observational datasets, including identifying students who appear well-positioned to benefit but do not select into advanced options.
Chinese Translation
K-12教育中一个具有根本性挑战的问题是修读更多高阶或更具挑战性的课程会产生何种影响。这一问题尤为复杂,因为学生(和/或其家长)自行选择是否选修这些课程,使得因果分析充满挑战。本文利用来自美国某公立学校系统的一个新颖数据集着手探讨这一问题。该数据集记录了在某学区范围推行可选的开放式选修高阶初中课程改革前后,学生的选课决策、既往学业记录、人口统计特征以及后续的学业结果。这是一个丰富的观察性数据集,但高阶课程的选修是由学生特征和自主选择驱动的,而非随机分配。这带来了一个核心的识别难题:影响高阶课程选修的同一批因素同样能够预测学业结果。因此,选修与未选修学生之间的简单比较存在混杂,朴素的估计结果可能反映的是学生在能力、动机或支持方面的潜在差异,而非课程学习本身的影响。我们的分析表明,选修高阶英语课程对学生学业成绩具有净正向但适中的影响。然而,这些收益的分布并不均衡:一些预测收益相对较大的学生(此前年份的“中等水平学生”)选修的可能性反而低于其他学生。另一些群体(如黑人学生和社会经济地位较低的学生)的选修倾向也显著更低。这种预测收益与实际选修之间的差距说明,通过严谨的数据分析,可以从大规模观察性数据集中提取可行动的洞见,包括识别出那些看似具备获益条件却未选择高阶课程的学生。
cs.LG / 158 / 2609.23838

From Regional to Global: Transfer Learning for Atmospheric Transport Emulators

从区域到全球:面向大气传输代理模型的迁移学习
Clark, Jeff, Fillola, Elena, Keshtmand, Nawid, Santos-Rodriguez, Raul, Rigby, Matthew
Abstract
Greenhouse gas emissions estimates can be derived using inverse methods by combining atmospheric concentration observations with chemical transport models. The latter traditionally use physics-driven simulators such as Lagrangian Particle Dispersion Models (LPDMs), which are expensive to run and do not scale well to modern satellites' high resolution data. Previously we developed a performant atmospheric transport emulator that approximates LPDM outputs ("footprints") over South America ~1,000X faster than the UK Met Office's LPDM. Expanding towards global emulation is not straightforward, as atmospheric transport is regionally heterogeneous. This paper evaluates spatial transferability capabilities of models across four world regions: South America, East Asia, South Asia, North Africa using both region-specific and multi-region models, and leave-one-region-out experiments. Regional differences are characterised in the context of input variable and output footprint distributions. This work builds intuition in cross-region generalisation and transfer learning, aiding regional performance towards efficient global emissions estimates.
Chinese Translation
温室气体排放估算可通过反演方法实现,即将大气浓度观测数据与化学传输模型相结合。后者传统上使用物理驱动的模拟器,如拉格朗日粒子扩散模型(LPDM),其运行成本高昂,且难以扩展到现代卫星的高分辨率数据。此前,我们开发了一个高性能的大气传输代理模型(emulator),其近似拉格朗日粒子扩散模型输出("足迹")的速度比英国气象局(UK Met Office)的LPDM快约1000倍。然而,向全球范围扩展并非易事,因为大气传输具有区域异质性。本文评估了模型在南美洲、东亚、南亚和北非四个世界区域之间的空间可迁移性,采用了区域专用模型与多区域模型,并进行了留一区域(leave-one-region-out)实验。同时,基于输入变量和输出足迹分布对区域差异进行了刻画。这项工作为跨区域泛化与迁移学习建立了直观认识,有助于提升区域模型性能,以实现高效的全球排放估算。
cs.LG / 159 / 2609.23843

Adaptive Determinantal Client Scheduling in Federated Learning

联邦学习中的自适应行列式客户端调度
Xu, Wen, Liang, Ben, Boudreau, Gary, Sokun, Hamza
Abstract
Scheduling clients for model training is critical in federated learning due to both data and system heterogeneity. Most previous works focus on the quality of the scheduled clients to achieve faster convergence, shorter wall-clock convergence time, or better average model performance. They rarely consider the diversity of clients, which is important to counter heterogeneity and improve performance for the worst-off clients. In this work, we advocate the use of determinantal point processes (DPPs) to model and enhance the diversity in client scheduling. We first design the kernel matrices of DPPs using gradient information and quality scores, which inherently enables a flexible quality-diversity trade-off. Applying fast MAP inference over DPPs, we propose Adaptive Determinantal Client Scheduling (ADCS) in FL. We further quantify the gradient approximation error of ADCS and develop convergence analysis for general biased client selection in FL with non-convex loss functions. We conduct comparative numerical experiments showing that ADCS outperforms state-of-the-art client scheduling algorithms, including both quality-based and diversity-based ones.
Chinese Translation
由于数据异构性和系统异构性的存在,为模型训练调度客户端在联邦学习中至关重要。以往的大多数工作聚焦于被调度客户端的质量,以实现更快的收敛速度、更短的墙钟收敛时间或更好的平均模型性能。它们很少考虑客户端的多样性,而多样性对于应对异构性并提升最差客户端的性能十分重要。在本工作中,我们提倡使用行列式点过程(DPPs)来建模并增强客户端调度中的多样性。我们首先利用梯度信息和质量分数设计DPP的核矩阵,这天然地实现了质量与多样性之间灵活的权衡。通过在DPP上应用快速的最大后验(MAP)推断,我们提出了联邦学习中的自适应行列式客户端调度(ADCS)。我们进一步量化了ADCS的梯度近似误差,并针对联邦学习中采用非凸损失函数的一般有偏客户端选择方法开展了收敛性分析。我们进行的对比数值实验表明,ADCS优于最先进的客户端调度算法,包括基于质量的和基于多样性的算法。
cs.LG / 160 / 2609.23845

PROSE: A Theory of Optimal Stopping with Perishable Evidence for Peer Selection in Intermittently Connected Decentralised Learning

PROSE:间歇连接去中心化学习中面向同伴选择的易逝证据最优停止理论
Anagnostopoulos, Christos
Abstract
Decentralised federated learning removes the aggregation server but makes collaboration dependent on transient peer availability. In mobile and intermittently connected systems, evaluating a promising peer consumes contact time and may cause the exchange opportunity itself to vanish, so that the evidence a learner gathers about a peer is perishable: it decays because links expire and because peer models drift while old measurements age. This paper develops a self-contained theory of optimal stopping for the resulting peer-selection problem. We formalise a receiver's within-contact decision as a finite-horizon Markov optimal-stopping problem with costly information acquisition and a future-arrival outside option, and prove that it admits an optimal policy characterised by a reservation value (Snell-envelope structure). Around this formulation we prove: (i) stage-uniform, drift-aware concentration and a maximin certification rule that is correct with high probability together with a finite-sample identification bound; (ii) a mobility-aware value of-information stopping rule and comparative statics showing that higher link hazard lowers the value of continued probing and enlarges the stopping region; (iii) a closed-form value of waiting under marked-Poisson contact arrivals, together with a search-theoretic reservation value whose comparative statics we characterise; and (iv) a myopic-optimality theorem establishing that, in sufficiently volatile (monotone) mobility regimes, the one-step confidence-safe rule is a sound surrogate for the optimal policy and never stops prematurely. We instantiate the theory as PROSE (Perishable-evidence Reservation-value Optimal Stopping for Exchange), a lightweight, fully local policy, and delineate the static contact and drift-free limits in which classical sequential decision problems are recovered. The development is entirely analytical.
Chinese Translation
去中心化联邦学习移除了聚合服务器,但使得协作依赖于瞬态的同伴可用性。在移动和间歇连接的系统中,评估一个有前景的同伴会消耗接触时间,并可能导致交换机会本身消失,因此学习者在评估同伴时所收集的证据是易逝的:由于链路过期以及同伴模型在旧测量数据老化期间发生漂移,这些证据会衰减。本文针对由此产生的同伴选择问题,建立了一套自洽的最优停止理论。我们将接收方在单次接触内的决策形式化为一个有限时域马尔可夫最优停止问题,该问题包含代价高昂的信息获取以及未来到达的机会作为外部选项,并证明该问题存在一个由保留值(reservation value,即 Snell 包络结构)刻画的最优策略。围绕这一形式化框架,我们证明了:(i) 阶段一致、考虑漂移的集中性结论,以及一个以高概率正确的最大最小认证规则和有限样本识别界;(ii) 一个感知移动性的信息价值停止规则,以及比较静态分析,表明链路风险率越高,继续探测的价值越低,停止区域越大;(iii) 在标记泊松(marked-Poisson)接触到达模型下等待价值的闭式表达式,以及一个搜索理论意义上的保留值,并刻画了其比较静态性质;(iv) 一个短视最优性定理,证明在足够波动(单调)的移动性状态下,单步置信安全规则是最优策略的可靠替代,且绝不会过早停止。我们将该理论实例化为 PROSE(Perishable-evidence Reservation-value Optimal Stopping for Exchange,面向交换的易逝证据保留值最优停止),一种轻量级的完全本地化策略,并刻画了经典序贯决策问题得以恢复的静态接触与无漂移极限情形。整个推导完全基于解析分析。
cs.LG / 161 / 2609.23875

VISTA: An Attention-Based Multi-Agent Reinforcement Learning Architecture for Space Situational Awareness Sensor Tasking

VISTA:一种基于注意力的多智能体强化学习架构,用于空间态势感知传感器任务规划
Leiva-Vélez, Miguel, Quiros, Adalberto Claudio, Rozado, Nicolas Gaston, Urrutxua, Hodei, Rodríguez-Fernández, Víctor
Abstract
The rapid growth of resident space objects is increasing the complexity of space situational awareness sensor tasking, challenging classical optimization methods as they allocate finite, heterogeneous, and distributed sensing resources across ever-larger catalogues. Existing deep reinforcement learning approaches show promise in reduced settings, but fixed-dimensional state and action representations limit their ability to scale to large, dynamic catalogues and distributed sensing networks. We introduce VISTA (Variable-Entity Intelligent Sensor Tasking Architecture), a scalable deep reinforcement learning architecture for persistent uncertainty-driven catalogue maintenance across variable object populations and sensor configurations. VISTA combines physics- and mission-informed top-K retrieval with entity-centric attention, recurrent memory, and pointer-based action decoding, thereby keeping each agent's observation and action spaces independent of catalogue size. We evaluate VISTA across different scenarios, from fixed-size single-sensor benchmarks to large-scale space-based tasking and heterogeneous cooperative sensing. With 30 orbiting targets, VISTA recovers the catalogue 31.2% faster than the fixed-dimensional recurrent baseline. In the large-scale regime, VISTA reduces five-hour uncertainty by 97.5% relative to the strongest classical reference and by 99.3% relative to the recurrent learner. Zero-shot tests up to 20,000 objects reveal near-linear relations between sensing capacity, catalogue size, and recovery horizon. Learned policies also exhibit sensor modality adaptation and generalization to population and initial-uncertainty shifts. Together, these results demonstrate that VISTA provides a scalable framework for adaptive space situational awareness sensor tasking across large, distributed networks of heterogeneous ground- and space-based sensors.
Chinese Translation
在轨空间物体的快速增长正使空间态势感知(Space Situational Awareness, SSA)传感器任务规划的复杂性不断提高,这对经典优化方法提出了挑战,因为它们需要在不断扩大的目标编目中对有限、异构且分布式的传感资源进行分配。现有的深度强化学习方法在简化场景中展现出潜力,但固定维度的状态与动作表示限制了其扩展到大规模、动态目标编目和分布式传感网络的能力。我们提出了VISTA(可变实体智能传感器任务规划架构,Variable-Entity Intelligent Sensor Tasking Architecture),这是一种可扩展的深度强化学习架构,用于在可变的目标数量和传感器配置下进行持续的、不确定性驱动的编目维护。VISTA将基于物理和任务信息的top-K检索与以实体为中心的注意力机制、循环记忆以及基于指针的动作解码相结合,从而使每个智能体的观测空间和动作空间与编目规模无关。我们在多种场景下对VISTA进行了评估,从固定规模的单传感器基准测试,到大规模天基任务规划以及异构协同感知。在30个在轨目标的场景下,VISTA比固定维度的循环基线方法快31.2%完成编目恢复。在大规模场景中,VISTA相对于最强的经典参考方法将五小时后的不确定性降低了97.5%,相对于循环学习方法降低了99.3%。针对最多20,000个目标的零样本测试表明,感知能力、编目规模与恢复周期之间存在近似线性的关系。学习到的策略还表现出对传感器模态的适应能力,以及对目标数量和初始不确定性变化的泛化能力。这些结果表明,VISTA为大规模分布式异构地基与天基传感器网络中的自适应空间态势感知传感器任务规划提供了一个可扩展的框架。
cs.LG / 162 / 2609.23876

GLR-MM: Graph-Based Global-Local Reconstruction for Robust Multimodal Chest X-ray and EHR Representation Learning under Missing Modalities

GLR-MM:基于图的全局-局部重建方法,用于模态缺失条件下鲁棒的多模态胸片与电子健康记录表示学习
Sharma, Surbhi, Manali, Nikhil, Maheshwari, Devesh
Abstract
Clinical multimodal models must often predict before all chest X-ray (CXR) and electronic health record (EHR) inputs are available. Existing approaches align observed representations, model missingness, or reconstruct across modalities, but do not jointly exploit within-patient and clinically similar inter-patient evidence. We propose GLR-MM, a Graph-Based Global-Local Reconstruction framework for early ICU mortality prediction. It maps five CXR-EHR modalities to a shared space, reconstructs missing embeddings through complementary local cross-modal and global graph-attention branches, adaptively fuses their estimates, and optimizes class-balanced prediction, reconstruction, and contrastive objectives. On 9,620 MIMIC-derived ICU stays, we evaluate 10%, 30%, and 50% random modality missingness with shared deterministic masks. MUSE performs better under mild and moderate missingness, whereas GLR-MM achieves higher AUROC and AUPRC at 50% by 0.0088 and 0.0249, respectively. These results indicate that graph-guided reconstruction is most useful when inputs are severely incomplete.
Chinese Translation
临床多模态模型往往需要在全部胸片(CXR)和电子健康记录(EHR)输入可用之前进行预测。现有方法或对已观测的表示进行对齐,或对缺失进行建模,或跨模态进行重建,但均未能联合利用患者内部以及临床相似患者之间的证据。我们提出GLR-MM,一个基于图的全局-局部重建框架,用于ICU早期死亡预测。该框架将五种CXR-EHR模态映射到共享空间,通过互补的局部跨模态分支和全局图注意力分支重建缺失的嵌入表示,自适应地融合二者的估计结果,并联合优化类别平衡的预测目标、重建目标和对比学习目标。在基于MIMIC的9,620例ICU住院记录上,我们使用共享的确定性掩码评估了10%、30%和50%的随机模态缺失情况。在轻度和中度缺失下,MUSE表现更优;而在50%缺失率下,GLR-MM取得了更高的AUROC和AUPRC,分别高出0.0088和0.0249。这些结果表明,图引导的重建方法在输入严重不完整时最为有效。
cs.LG / 163 / 2609.23883

Collaborative Streaming Anomaly Detection with Interactive Explanations and Ensemble Consensus

具有交互式解释与集成共识的协同流式异常检测
Risca, Diogo, Lourenço, Afonso, Martins, Ricardo, Marreiros, Goreti
Abstract
We present a collaborative streaming anomaly detection system for high-speed data streams that explicitly integrates human analysts into the decision loop. The system combines heterogeneous detectors and aggregates their outputs through a normalization-based weighted consensus, complemented by artifact-aware rules to stabilize anomaly scoring under deployment. To improve interpretability, it derives surrogate models that approximate the ensemble consensus and expose human-readable sensor conditions associated with anomalous behavior. Analysts can actively intervene by reviewing anomaly episodes, adjusting consensus behavior, and refining surrogate rules used for anomaly prediction, producing a human-adjusted ensemble. We evaluate the approach on an industrial stream with 260\,000 events and 3 anomalous episodes, showing robust detection and actionable human-AI interaction.
Chinese Translation
我们提出了一种面向高速数据流的协同流式异常检测系统,该系统将人类分析人员显式地纳入决策闭环。系统结合异构检测器,并通过基于归一化的加权共识聚合其输出,同时辅以工件感知规则,以在部署过程中稳定异常评分。为提升可解释性,系统导出近似集成共识的代理模型(surrogate models),揭示与异常行为相关的人类可读的传感器条件。分析人员可以通过审查异常事件、调整共识行为以及优化用于异常预测的代理规则来进行主动干预,从而形成经人类调整的集成模型。我们在包含26万个事件和3个异常事件的工业数据流上对该方法进行了评估,结果表明其具备稳健的检测能力和可行有效的人机交互。
cs.LG / 164 / 2609.23892

Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs

Circuit-Diff:一种基于事实性编辑的干预方法,用于在归因图中定位知识
Friedman, Edward G., Song, Xiangchen
Abstract
Mechanistic interpretability defines features as the fundamental units of a neural network and circuits as the weighted subgraphs that carry out its computation. Because individual neurons are polysemantic, Cross-Layer Transcoders (CLTs) were introduced as a way to approximate a model's circuits by generating an attribution graph. The nodes of that graph, however, are unlabeled features: reading a graph means pruning it and then working out by hand what each surviving node means. To make CLTs easier to use for circuit discovery, we introduce Circuit-Diff, which intervenes on the model itself with a low-rank factual edit and takes the features whose role in the attribution graph changes under that edit as related to the edited knowledge. On the edits we examine, the flagged nodes are not only detectors of the object token: read off the CLT's released feature dashboards, they include features for the history, geography and associations surrounding the old and new objects. We formalize the method, measure how reliable a frozen CLT remains after a factual edit, test the selected nodes causally by patching them on up to 24 CounterFact edits, give a case study, and release an open-source implementation built on the circuit-tracer package, together with two further tools (multi-prompt aggregation and rule-based supernode labeling).
Chinese Translation
机制可解释性将特征定义为神经网络的基本单元,将电路(circuits)定义为执行模型计算的加权子图。由于单个神经元具有多义性,研究者提出了跨层转码器(Cross-Layer Transcoders, CLT),通过生成归因图来近似模型的电路。然而,该图中的节点是未标注的特征:解读一张归因图意味着先对其进行剪枝,然后人工推断每个保留节点的含义。为了使 CLT 更易于用于电路发现,我们提出了 Circuit-Diff,该方法通过低秩事实性编辑对模型本身进行干预,并将归因图中在该编辑下角色发生变化的特征视为与被编辑知识相关的特征。在我们所考察的编辑中,被标记出的节点不仅仅是目标词元的探测器:从 CLT 发布的特征仪表板中可以读到,它们涵盖了与新旧对象相关的历史、地理及关联知识的特征。我们对该方法进行了形式化定义,测量了事实性编辑后冻结的 CLT 的可靠性,通过在多达 24 个 CounterFact 编辑上进行节点修补对所选节点进行了因果检验,并给出了案例研究。我们还基于 circuit-tracer 包发布了一个开源实现,并附带了两个额外工具(多提示聚合和基于规则的超节点标注)。
cs.LG / 165 / 2609.23900

GDN Tree-Scan: Served Tree Verification for Recurrent-Hybrid Language Models

GDN Tree-Scan:面向循环-混合语言模型的服务化树验证
Ma, Zhiyuan
Abstract
Tree speculative decoding verifies multiple candidate continuations in one target forward pass. For attention-only transformers, the verifier mainly needs an ancestry mask. Recurrent-hybrid language models break this assumption: a candidate row must also carry the recurrent state that native sequential decode would have produced along its root-to-node path. Otherwise, a verifier can use a correct attention mask while still conditioning on an impossible recurrent history. We present GDN Tree-Scan, a served verifier for Gated-DeltaNet hybrid language models integrated into vLLM. The system combines FlashAttention-2 tree-bias attention, branch-local GDN scan/replay, device-side multidraft commitment, and accepted-chain-only state publication. On the public Qwen3.6-27B-FP8 checkpoint, in a clean batch-one (B=1) SWE/Codex decode gate at temperature 0.6, a six-node root-branch tree increases committed tokens/event by 17.2% at near-native verify-forward time and reaches 23.88 token-weighted decode tokens/s versus 18.80 for native five-step MTP (E5), a 27.0% token-weighted decode-throughput gain. The per-request-equal latency view is +4.0%, and end-to-end task wall time remains prefill-heavy. Empirical equivalence evidence is scoped to recurrent-oracle probability-rescore (p-rescore) closure within the observed native flip floor, not a full distribution-distance proof.
Chinese Translation
树状投机解码(Tree speculative decoding)可在一次目标模型前向传播中验证多个候选续写。对于纯注意力(attention-only)Transformer,验证器主要只需一个祖先关系掩码(ancestry mask)。而循环-混合(recurrent-hybrid)语言模型打破了这一假设:候选行还必须携带原生顺序解码沿其根到节点路径所产生的循环状态,否则验证器即使使用了正确的注意力掩码,也可能基于一个不可能出现的循环历史进行条件化计算。我们提出 GDN Tree-Scan,一个集成于 vLLM 的、面向 Gated-DeltaNet(GDN)混合语言模型的服务化验证器。该系统结合了 FlashAttention-2 树偏置注意力、分支局部的 GDN 扫描/重放、设备端多草稿提交,以及仅对被接受链发布状态等机制。在公开的 Qwen3.6-27B-FP8 检查点上,在温度 0.6 的干净批量为一(B=1)的 SWE/Codex 解码门限测试中,六节点根分支树在接近原生验证前向时间的情况下,将每次事件的已提交 token 数提升 17.2%,并达到 23.88 token 加权解码 tokens/s,而原生五步 MTP(E5)为 18.80,即 token 加权解码吞吐量提升 27.0%。按逐请求等延迟视角衡量则提升 4.0%,且端到端任务墙钟时间仍以预填充(prefill)为主。经验等价性证据的范围限定于在观测到的原生翻转下限内的循环-先知概率重评分(p-rescore)闭合,而非完整的分布距离证明。
cs.LG / 166 / 2609.23906

Multivariate quantile regression via Kolmogorov-Arnold Networks

基于Kolmogorov-Arnold网络的多变量分位数回归
Polar, Andrew, Poluektov, Michael
Abstract
This paper introduces a novel algorithm for predicting conditional joint distributions of vector-valued targets in stochastic systems whose randomness is intrinsic rather than arising from observation errors or additive noise. Multivariate quantile regression also involves modeling conditional joint distributions but represents a less challenging task. It predicts the probability that vector-valued targets fall within predefined regions, identifies regions corresponding to predefined probability levels, or performs both tasks simultaneously. The proposed identification technique employs ensembles of Kolmogorov--Arnold networks (KANs) as flexible function approximators. Although the suggested technique is not theoretically restricted to KANs, KANs are particularly well suited to the proposed construction and are therefore used throughout this study. In addition to the training procedure, this work introduces a new discrepancy measure for joint distributions and a goodness-of-fit (GoF) test based on it. This GoF test was initially developed to validate and calibrate the proposed identification technique and is used here in an ad hoc manner. Although the test could be tabulated for broader use, such a tabulation is not pursued in this work. The test is also applicable more generally.
Chinese Translation
本文提出了一种新算法,用于预测随机系统中向量值目标变量的条件联合分布,该系统的随机性是内在的,而非源于观测误差或加性噪声。多变量分位数回归同样涉及对条件联合分布的建模,但代表了一个难度较低的任务。它预测向量值目标变量落入预定义区域的概率,识别与预定义概率水平相对应的区域,或同时执行这两项任务。所提出的识别技术采用Kolmogorov-Arnold网络(KAN)集成作为灵活的函数逼近器。尽管所建议的技术在理论上并不局限于KAN,但KAN特别适合所提出的构建方式,因此在本研究中贯穿使用。除训练过程外,本工作还引入了一种新的联合分布差异度量方法,以及基于该度量的拟合优度(GoF)检验。该GoF检验最初是为验证和校准所提出的识别技术而开发的,本文中以特定方式使用它。尽管该检验可以编制成表以便更广泛地使用,但本工作并未进行此类编制。该检验也具有更广泛的一般适用性。
cs.LG / 167 / 2609.23907

A discrete generative model of neuronal spiking activity on microelectrode arrays

一种用于微电极阵列神经元放电活动的离散生成模型
Tanveer, Md Sayed, Mostajo-Radji, Mohammed A., Wang, Ge
Abstract
Generative models of neural activity could help characterize tissue dynamics, compare experimental conditions, and simulate population activity for applications ranging from disease and drug-response studies to closed-loop experimentation. Existing approaches, however, typically assume a fixed set of sorted neurons, whereas high-density microelectrode arrays produce extremely sparse, array-wide binary spike volumes in which the observed subset of electrodes varies across assays. We introduce a discrete generative model that represents this activity using a shared vocabulary of spatiotemporal motifs. A residual vector-quantized autoencoder learns the motif vocabulary, while a factorized masked transformer predicts where activity occurs and which motif appears at each active location. We evaluate the model on 31 assays spanning human brain organoids and acute \emph{ex vivo} human hippocampal tissue. The learned motifs are broadly reused: assay identity explains only $9%$ of the entropy in motif use, and motif overlap across tissue types is comparable to overlap within them. When representation quality is evaluated independently of the generative prior, our approach achieves $5.2\times$ the voxel-level reconstruction average precision of a matched flat tokenizer. For masked completion and free generation, the full model achieves $1.4$--$2.6\times$ the site-level average precision of the matched generative baseline and outperforms it across all four families of generation metrics. These results establish a compact, reusable representation for array-wide spiking activity without learned assay-specific parameters, providing a scalable foundation for generative modeling across diverse neural preparations.
Chinese Translation
神经活动的生成模型有助于刻画组织动力学特性、比较实验条件以及模拟群体活动,其应用涵盖疾病与药物反应研究乃至闭环实验。然而,现有方法通常假设一个固定且经过分拣的神经元集合,而高密度微电极阵列产生的是极其稀疏、覆盖整个阵列的二值放电体数据,其中被观测到的电极子集在不同实验之间各不相同。我们提出了一种离散生成模型,利用一个共享的时空基元(spatiotemporal motifs)词表来表示这种活动。一个残差向量量化自编码器学习基元词表,同时一个因子化的掩码transformer预测活动发生的位置以及每个活跃位置上出现的基元。我们在涵盖人类脑类器官和离体(ex vivo)人类海马组织的31个实验上对该模型进行了评估。学习到的基元被广泛复用:实验身份仅能解释基元使用中9%的熵,且不同组织类型之间的基元重叠度与同类型组织内部的重叠度相当。当表示质量独立于生成先验进行评估时,我们的方法在体素级重建平均精度上达到了匹配的扁平tokenizer的5.2倍。在掩码补全和自由生成任务中,完整模型在位点级平均精度上达到匹配生成基线的1.4至2.6倍,并在全部四类生成指标上均优于该基线。这些结果建立了一种紧凑、可复用的全阵列放电活动表示方法,无需学习实验特定的参数,为跨多种神经标本的生成建模提供了可扩展的基础。
cs.LG / 168 / 2609.23924

Matched-Input Estimates Differ in Sign Across Architectures: Auditing EEG Foundation Models on Motor Imagery

匹配输入的估计在不同架构间符号相反:对EEG基础模型在运动想象任务上的审计
Zhou, Kevin, Roy, Sparsh
Abstract
Pretrained EEG foundation models are increasingly proposed as general-purpose encoders for brain-computer interfaces, yet recent benchmarks disagree about when their representations transfer to downstream tasks. We audit LaBraM and CBraMod on motor imagery under a validation-locked protocol in which preprocessing, architecture, optimization, freeze depth, checkpoint, temperature, and method selection are determined using training-session data only. On four-class BCI Competition IV-2a, every supervised comparator evaluated here outperforms every foundation-model configuration, including validation-selected fine-tuning. We then examine a key confound: foundation models and task-specific decoders are normally evaluated using different input pipelines. Retraining three supervised architectures on the broadband arrays consumed by the foundation models produces matched-input accuracy differences of opposite sign across architectures: broadband input improves ATCNet by 0.078 accuracy while reducing EEG Conformer accuracy by 0.088. None of the three individual matched-input terms is significant after multiple-comparison correction at n = 9, so we treat the sign variation descriptively rather than as a formal architecture-by-pipeline interaction. These observed sign differences suggest that a single comparator may not provide an architecture-invariant decomposition of a pretrained-versus-supervised performance gap. The four-class deficit also does not reproduce uniformly across motor-imagery datasets: on two-class BNCI2014-004 we cannot detect the same separation between fine-tuned CBraMod and the supervised comparators. Finally, validation-fitted temperature scaling returns foundation-model calibration error to the supervised range despite substantially lower four-class accuracy.
Chinese Translation
预训练EEG基础模型日益被提议作为脑机接口的通用编码器,然而近期的基准测试对于其表征何时能够迁移到下游任务存在分歧。我们在一种验证锁定的协议下对LaBraM和CBraMod在运动想象任务上进行审计,该协议中预处理、架构、优化、冻结深度、检查点、温度以及方法选择均仅使用训练会话数据确定。在四分类BCI Competition IV-2a数据集上,本文评估的每一个有监督对比方法均优于所有基础模型配置,包括通过验证选择的微调。随后我们考察一个关键混淆因素:基础模型与任务特定解码器通常采用不同的输入流程进行评估。在基础模型所使用的宽带数组上重新训练三个有监督架构,产生了跨架构符号相反的匹配输入精度差异:宽带输入使ATCNet精度提高0.078,同时使EEG Conformer精度降低0.088。在n=9且经过多重比较校正后,三个匹配输入单项均不显著,因此我们将符号变化作描述性处理,而非正式的架构与流程交互作用。这些观察到的符号差异表明,单一的对比方法可能无法对预训练与有监督之间的性能差距提供架构不变的分解。四分类的性能差距也未能在不同运动想象数据集上一致重现:在二分类BNCI2014-004数据集上,我们无法检测到微调后的CBraMod与有监督对比方法之间同样的分离。最后,验证拟合的温度缩放使基础模型的校准误差回到有监督方法的范围内,尽管其四分类精度明显更低。
cs.LG / 169 / 2609.23934

The Neural Forcing for Three-Dimensional Incompressible Navier-Stokes finite time blowup

面向三维不可压缩Navier-Stokes方程有限时间爆破的神经强迫项方法
Li, Beibei
Abstract
We present a two-part neural framework for forced three-dimensional incompressible Navier--Stokes flow. Part~I develops the computational forcing system. A physics-informed neural model generates structured external-force trajectories, candidates are optimized through differentiable PDE rollouts or PPO-Clip, and selected forcings are frozen and checked by independent fixed-force replay. Part~II provides the mathematical certification layer. It separates neural candidate discovery from continuum analysis, derives integrated reciprocal-vorticity criteria that imply Riccati-type growth and finite-time loss of smooth continuation, develops a validated computational-to-continuum transfer strategy, and establishes a conditional positive-probability closure for a nondegenerate neural output law. The proof is complete at the continuum level.
Chinese Translation
我们提出了一个用于受迫三维不可压缩Navier-Stokes流动的两部分神经框架。第一部分构建计算强迫系统:一个物理信息神经网络模型生成结构化的外力轨迹,候选力通过可微分的PDE滚动求解或PPO-Clip算法进行优化,选定的强迫项被冻结并通过独立的固定力重放加以验证。第二部分提供数学认证层:该部分将神经候选项的发现与连续介质分析分离,推导出蕴含Riccati型增长和有限时间光滑延拓丧失的积分倒数涡量判据,建立了一种经过验证的计算到连续介质的转移策略,并针对非退化的神经输出分布建立了一个条件性的正概率闭合结果。相关证明在连续介质层面是完备的。
cs.LG / 170 / 2609.23990

MGRD: Compact morphology-gated residual diffusion for variance-aware cross-domain neurite forecasting

MGRD:用于方差感知跨域神经突预测的紧凑形态门控残差扩散模型
Hsieh, Tsung Yeh, Anitescu, Cosmin, Kim, Chunghwan, Webster-Wood, Victoria A., Zhang, Yongjie Jessica
Abstract
Tracking neurite morphology over time helps characterize structural changes during neuronal development and deterioration, but long-term time-lapse imaging is resource-intensive and difficult to scale. Forecasting future morphology could reduce this burden. Existing neurite digital-twin models such as gated spatiotemporal attention (gSTA) produce a single deterministic forecast without representing variability among plausible futures. We introduce Morphology-Gated Residual Diffusion (MGRD), a compact stochastic surrogate that jointly forecasts twenty future neurite-morphology frames from ten observed frames while conditioning on morphology features derived from the latest observation. On controlled phase-field trajectories, MGRD reduces trajectory-wise mean MAE by 9.7% relative to a matched control while updating 4.46 times fewer parameters. On human iPSC-derived neuron microscopy, MGRD improves all four reported metrics over gSTA, including a 39.6% reduction in trajectory-wise mean MAE and a 45.3% increase in skeleton F1. Without mouse-domain retraining or fine-tuning, MGRD also improves MAE and skeleton F1 on mouse cortical-neurosphere microscopy across 10-40-min sampling intervals and forecast horizons beyond 13 hours. Repeated sampling provides a case-level variance score for ranking forecast difficulty. Retaining approximately 60% of the lowest-variance cases reduces mean MAE by 17.6% on iPSC microscopy and 16.8% on simulation data. MGRD uses 1.01% of gSTA's parameters, requires less than one tenth of its training-update time, and generates a 50-step DDIM trajectory 7.9% faster when morphology features are cached. These results establish MGRD as a compact stochastic surrogate for neurite-morphology forecasting and case prioritization across simulation and microscopy datasets.
Chinese Translation
对神经突形态随时间变化的追踪有助于刻画神经元发育与退化过程中的结构性改变,但长期延时成像资源消耗大且难以规模化。预测未来形态可以减轻这一负担。现有的神经突数字孪生模型(如门控时空注意力机制 gSTA)只产生单一确定性预测,无法表征多种可能未来之间的变异性。我们提出了形态门控残差扩散模型(Morphology-Gated Residual Diffusion, MGRD),这是一种紧凑的随机代理模型,能够基于最近一次观测提取的形态特征作为条件,从十帧观测数据联合预测二十帧未来的神经突形态。在受控相场轨迹数据上,MGRD 相较于匹配的对照模型将逐轨迹平均 MAE 降低了 9.7%,同时更新的参数量减少至其 1/4.46。在人 iPSC 分化神经元显微图像上,MGRD 在全部四项报告指标上均优于 gSTA,包括逐轨迹平均 MAE 降低 39.6%,骨架 F1 提升 45.3%。无需在小鼠域上重新训练或微调,MGRD 在小鼠皮层神经球显微图像上,于 10–40 分钟的采样间隔以及超过 13 小时的预测时域范围内,同样提升了 MAE 和骨架 F1。重复采样可提供样本级的方差评分,用于对预测难度进行排序。保留约 60% 方差最低的样本,可使 iPSC 显微数据上的平均 MAE 降低 17.6%,仿真数据上降低 16.8%。MGRD 仅使用 gSTA 1.01% 的参数量,训练更新时间不足其十分之一,且在缓存形态特征后生成 50 步 DDIM 轨迹的速度提升 7.9%。这些结果确立了 MGRD 作为神经突形态预测与样本优先级排序的紧凑随机代理模型,可同时适用于仿真与显微图像数据集。
cs.LG / 171 / 2609.23995

Simpler Methods Work Better for L1 Penalized Logistic Models and Large Datasets

更简单的方法在L1惩罚逻辑回归模型与大规模数据集上表现更佳
Raff, Edward, Holt, James
Abstract
Linear models with an $L_1$-norm penalty remain state-of-the-art for high-dimensional ($d > 1,000,000$) tasks, offering a straightforward method for solving real-world industry problems. Despite their widespread use in industry and utility, many $L_1$ solvers are not effective for general use, are prohibitively slow, and are ineffective in parallelization. This makes them difficult to train in an MLOps pipeline on large industry-scale corpora. In this work, we test several proposed ``state-of-the-art'' solutions from the literature and find that older methods are currently far superior for general use. We also identify several recommendations for academics to perform research that avoids erroneously overconfident results, which can prevent the transition to production use. Equally surprising, we find that a new and simple baseline, using LBFGS on a sub-gradient, is highly effective with minor tweaks, despite being dismissed in the literature for theoretical non-convergence. In practice, we find it is an easier-to-support and easier-to-scale method for production use.
Chinese Translation
带$L_1$范数惩罚的线性模型在高维($d > 1,000,000$)任务中仍是当前最先进的方法,为解决现实世界的工业问题提供了一种直接有效的方式。尽管其在工业界应用广泛且实用性强,许多$L_1$求解器却不适合通用场景,速度极其缓慢,且难以并行化。这使得它们难以在MLOps流程中基于工业级大规模语料进行训练。在本工作中,我们测试了文献中提出的若干"最先进"解决方案,发现较老的方法目前在通用性上远优于这些新方法。我们还为学术界提出了若干建议,以帮助开展能够避免错误且过度自信结论的研究——这类结论会阻碍研究成果向生产环境的转化。同样令人惊讶的是,我们发现一种新颖而简单的基线方法,即在次梯度上使用LBFGS,仅需少量调整便非常高效,尽管该做法因理论上不收敛而在文献中不被看好。在实践中,我们发现这是一种更易于维护、更易于扩展的生产级方法。
cs.LG / 172 / 2609.23999

Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models

开放权重大型语言模型中的临床风险分类错位与成本不对称性
Liu, Star S. D., Ding, Xiyu, Barrett, Robert B., Santamaria-Pang, Alberto, Dobbins, Nic, Lehmann, Harold P.
Abstract
How large language models (LLMs) integrate patient risk with clinical cost tradeoffs remains poorly understood. We investigated how four open-weight LLMs (Qwen-2.5-7B/32B and Llama-3.1-8B/70B) internally represent cost tradeoffs, how these representations relate to clinical predictions, and whether decisions shift as predicted by the specified cost direction and magnitude. Using a public diabetes dataset, we varied 11 false-negative (FN) to false-positive (FP) cost ratios across three phrasings and examined representations and behavioral outputs. Patient risk was linearly recoverable on par with conventional classifiers (AUC $\approx 0.83$), and cost direction was recoverable in every model. However, representational shifts in cost direction tracked output changes only in the two larger models, and responses to cost magnitude were predominantly direction-agnostic. Only 2 of 12 model-phrasings showed both opposing responses to increasing FN versus FP costs and cost-correct ordering. Representationally, a direction fitted on one cost side did not invert when transferred to the other, as expected under mirror-symmetric encoding. These findings suggest that LLMs encode risk and cost information but do not reliably integrate them into cost-correct decisions. Clinical evaluations should therefore include tradeoff tests, phrasing sensitivity, and default operating points alongside predictive performance.
Chinese Translation
大型语言模型(LLMs)如何将患者风险与临床成本权衡相结合仍知之甚少。我们研究了四个开放权重LLM(Qwen-2.5-7B/32B 和 Llama-3.1-8B/70B)如何在内部表征成本权衡、这些表征与临床预测之间的关系,以及决策是否随所指定的成本方向和幅度而发生相应变化。利用一个公开的糖尿病数据集,我们在三种表述方式下改变了11种假阴性(FN)与假阳性(FP)成本比率,并考察了模型表征和行为输出。结果显示,患者风险可以线性恢复,其性能与常规分类器相当(AUC ≈ 0.83),且成本方向在所有模型中均可恢复。然而,成本方向的表征变化仅在两个较大的模型中与输出变化相一致,而对成本幅度的响应大多与方向无关。在12个模型-表述组合中,仅有2个同时表现出对FN成本增加与FP成本增加的相反响应以及符合成本修正的排序。在表征层面,基于成本一侧拟合的方向在迁移到另一侧时并未像镜像对称编码所预期的那样发生反转。这些发现表明,LLM能够编码风险和成本信息,但并不能可靠地将二者整合为符合成本修正的决策。因此,临床评估除预测性能外,还应包括权衡测试、表述敏感性以及默认工作点的考察。
cs.LG / 173 / 2609.24003

ShapeLex: Decoupling Local Shape Symbolization and Global Scale Modeling for Text-Controlled Time Series Generation

ShapeLex:面向文本控制时间序列生成的局部形状符号化与全局尺度建模解耦方法
Wei, Subo, Gao, Jianqi, Fan, Mingyan, Xie, Shaorong, Wang, Xinzhi, Dong, Yongpeng
Abstract
Text-controlled time series generation aims to synthesize sequences that follow natural-language descriptions while remaining faithful to real data distributions. Existing paradigms often couple semantic understanding and sequence modeling in a single continuous latent space, lacking explicit local semantic anchors and separation between global continuous attributes and local discrete shapes. As a result, key local structures may be smoothed, missed, or misplaced. We propose Shape Lexicon (ShapeLex), which decouples text-to-sequence generation into discrete symbolization of local shapes and continuous modeling of global attributes. ShapeLex first induces a reusable vocabulary of discrete shape units, such as rises, spikes, and sharp drops, from training data, forming an interpretable symbolic space. An autoregressive generator then selects shapes according to the textual description, adjusts attributes such as position and duration, and composes them in temporal order into a shape skeleton. Finally, a mixture-density scale head models and samples the overall level and volatility to restore realistic global scale. Experiments on twelve public datasets, real user-written text, and downstream forecasting tasks show that ShapeLex generates series that better match real data distributions than existing methods. In addition, paired supervision is automatically synthesized from the learned vocabulary, avoiding annotation costs that grow with dataset size and improving scalability.
Chinese Translation
文本控制的时seque时间序列生成旨在合成符合自然语言描述、同时忠实于真实数据分布的序列。现有范式通常将语义理解与序列建模耦合在单一连续潜在空间中,缺乏显式的局部语义锚点,也未将全局连续属性与局部离散形状加以分离,导致关键的局部结构可能被平滑、遗漏或错位。我们提出ShapeLex(Shape Lexicon),将文本到序列的生成解耦为局部形状的离散符号化和全局属性的连续建模。ShapeLex首先从训练数据中归纳出可复用的离散形状单元词汇表,如上升、尖峰和骤降,形成一个可解释的符号空间。随后,自回归生成器根据文本描述选择形状,调整位置和持续时间等属性,并按时间顺序将其组合成形状骨架。最后,混合密度尺度头对整体水平和波动性进行建模和采样,以恢复真实的全局尺度。在十二个公开数据集、真实用户撰写文本以及下游预测任务上的实验表明,与现有方法相比,ShapeLex生成的序列能更好地匹配真实数据分布。此外,配对监督可从学习到的词汇表中自动合成,避免了随数据集规模增长而产生的标注成本,并提升了可扩展性。
cs.LG / 174 / 2609.24040

Graph-to-Grid (G2G): Continuous-Coordinate Feature Painting for Soccer Pass Surfaces

Graph-to-Grid(G2G):用于足球传球表面的连续坐标特征绘制
Günay, Kaan, Gun, Orhun
Abstract
Dense pass surfaces give, for every pitch cell, whether a pass played there would arrive, whether the carrier would choose it, and what the possession would then be worth. The networks that draw them read the state as a raster of per-cell counts, losing where inside a cell each player stands. LiDAR detectors, bird's-eye-view perception and graph weather models move entity features onto a grid, binning each entity to a cell or learning the transfer. We evaluate the interpolated form: each player's features are scattered bilinearly onto the grid at the player's measured coordinates, so the surface loss trains the per-player encoder end to end. Those systems adopt an interface; this paper measures one. On 53,628 passes from the 2022 World Cup, painting improves selection likelihood over the same core fed rasters alone by about a quarter of a nat: in every match of an eight-fold cross-validation, with every arm tuned over five seeds, and after retraining on seven Bundesliga and 2. Bundesliga matches from another provider. Thirteen pre-specified studies locate the gain: painting the nine raw player features with no encoder carries three quarters of it, and the learned encoder and message passing add a smaller, resolved increment. Painting also helps the original SoccerMap and a canonical U-Net, whereas offset channels, a finer raster, an attention painter and a raster-free decoder do not. Frozen across the provider boundary the likelihood advantage is lost; injected tracking error compresses it. These results concern observed-endpoint prediction, not calibrated evaluation of hypothetical passes.
Chinese Translation
密集传球表面针对球场上的每个单元格给出:在该位置踢出的传球是否会到达、持球者是否会选择该传球,以及此后球权价值几何。绘制这些表面的网络将比赛状态读取为逐单元格计数的栅格(raster),从而丢失了每个球员在单元格内的具体站位信息。LiDAR 检测器、鸟瞰图感知以及图结构气象模型将实体特征移动到网格上,方法是将每个实体归入某个单元格,或学习这种转换。我们评估的是插值形式:将每个球员的特征按其测量坐标以双线性方式散布(scatter)到网格上,从而使表面损失能够端到端地训练逐球员编码器。上述系统只是采用了一种接口,而本文对这一接口进行了度量。在 2022 年世界杯的 53,628 次传球上,特征绘制使选择概率相比仅输入栅格的相同核心网络提升了约四分之一纳特(nat):在八折交叉验证的每场比赛中、在每组均于五个随机种子上调优的设置下,以及在从另一数据提供商的七场德甲和德乙比赛上重新训练之后,该结论均成立。十三项预先设定的研究定位了这一增益:不带编码器、仅绘制九个原始球员特征即可贡献约四分之三的增益,而学习型编码器和消息传递仅带来一个更小的、可分辨的增量。特征绘制同样提升了原始的 SoccerMap 和一个标准 U-Net,而偏移通道、更精细的栅格、注意力绘制器和无栅格解码器则无此效果。跨数据提供商冻结模型时,概率优势消失;注入的跟踪误差会压缩该优势。这些结果针对的是观测到端点的预测,而非对假设性传球的校准评估。
cs.LG / 175 / 2609.24042

Q-DEQ: Discrete Solving and Quantization for Deep Equilibrium Models in Time Series Forecasting under Edge Deployment Coding Constraints

Q-DEQ:面向边缘部署编码约束下时间序列预测的深度均衡模型离散求解与量化方法
Yang, Ruotong, Zhu, Hongdong, Gao, Qi, Ma, Yin, Wei, Hai, Wen, Kai
Abstract
Edge deployment motivates forecasting models with compact parameter storage and low-bit representations. Deep equilibrium models (DEQs) obtain implicit depth by repeatedly applying a shared layer, reducing the parameter cost of explicit layer stacking. Their usual Anderson solver, however, searches for update coefficients in the continuous real domain. We propose Q-DEQ, which formulates local updates in DEQ forward solving as discrete optimization problems. Candidate directions are constructed from the current state and iteration history, and a local quadratic residual model is used to evaluate their combinations. Binary encoding of the direction coefficients yields a quadratic unconstrained binary optimization (QUBO) problem that can be solved by simulated annealing (SA) or a coherent Ising machine (CIM). After fixed-point solving, a re-forward pass applies W8A8 fake quantization to the shared layer's weights and activations. We evaluate Q-DEQ with an iTransformer backbone on five multivariate time series forecasting datasets. Relative MSE differences from the explicit multi-layer baseline range from $-1.16\%$ to $+2.90\%$, with lower MSE on two datasets. DEQ parameter sharing reduces parameter counts by factors of $1.80\times$--$3.82\times$; combined with W8A8, static weight storage is reduced by factors of $4.3\times$--$12.8\times$. Local QUBO problems solved using CPU-based SA and the Kaiwu CIM physical backend produce closely matching downstream forecasts. These results establish local discrete solving as a viable component of DEQ time series forecasting and provide a route for executing fixed-point updates through different combinatorial optimization backends.
Chinese Translation
边缘部署要求预测模型具有紧凑的参数存储和低比特表示。深度均衡模型(DEQ)通过反复应用共享层获得隐式深度,从而降低了显式堆叠多层的参数成本。然而,其常用的Anderson求解器在连续实数域中搜索更新系数。我们提出Q-DEQ,将DEQ前向求解中的局部更新表述为离散优化问题。候选方向由当前状态和迭代历史构造,并利用局部二次残差模型评估其组合。方向系数的二进制编码产生一个二次无约束二进制优化(QUBO)问题,可通过模拟退火(SA)或相干伊辛机(CIM)求解。在不动点求解完成后,再进行一次前向传播,对共享层的权重和激活值施加W8A8伪量化。我们在五个多变量时间序列预测数据集上使用iTransformer骨干网络评估了Q-DEQ。与显式多层基线相比,MSE的相对差异范围为$-1.16\%$至$+2.90\%$,其中在两个数据集上取得了更低的MSE。DEQ参数共享将参数量降低至$1.80\times$至$3.82\times$倍;结合W8A8量化后,静态权重存储降低至$4.3\times$至$12.8\times$倍。使用基于CPU的SA和Kaiwu CIM物理后端求解的局部QUBO问题产生了高度一致的下游预测结果。这些结果确立了局部离散求解作为DEQ时间序列预测中一个可行的组成部分,并为通过不同组合优化后端执行不动点更新提供了一条途径。
cs.LG / 176 / 2609.24089

FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention

FlashBoB:面向Softmax注意力机制的高I/O效率精确二阶反向传播
Givans, Anthony, Crawshaw, Michael, Liu, Mingrui
Abstract
Transformer models built on the attention mechanism have become a central building block in modern deep learning, yet softmax attention remains a major bottleneck for long-context workloads. While FlashAttention makes the forward and first backward passes I/O-efficient, it does not support backward-over-backward (BoB), which enables exact differentiation through the backward pass for applications such as second-order optimization, test-time training, gradient-based memory, and meta-learning. Existing BoB implementations either materialize large intermediate tensors or exhaust GPU memory at long sequence lengths. We present FlashBoB, an exact, I/O-efficient algorithm for BoB in softmax attention that keeps computation within on-chip tiles and avoids all $N \times N$ intermediate tensors, where $N$ is the sequence length. The key insight is a hierarchical affine structure in the softmax double backward: two row-wise scalars determine all outputs through affine transformations. This yields a two-pass schedule with bounded on-chip static random-access memory (SRAM) usage and minimal off-chip high-bandwidth memory (HBM) traffic. FlashBoB achieves $\Theta(N^2 d^2/M)$ HBM traffic ($d$ is the head dimension and $M$ is the memory size) and, within the standard FlashAttention-style score-recomputation model, matches the inherited large-cache lower bound for exact forward attention. Empirically, it scales exact attention BoB to $N=262\text{K}$ on a single A100 80GB GPU, where prior PyTorch exact baselines fail by $N=16\text{K}$, and is up to $6.3\times$ faster than FlashBack. These results make exact second-order attention practical at long-context sequence lengths where prior implementations cannot run efficiently.
Chinese Translation
基于注意力机制构建的Transformer模型已成为现代深度学习的核心构件,然而softmax注意力仍然是长上下文工作负载的主要瓶颈。虽然FlashAttention使前向传播和第一次反向传播实现了I/O高效,但它不支持二阶反向传播,即在反向传播过程中进行精确微分的能力,而该能力对于二阶优化、测试时训练、基于梯度的记忆以及元学习等应用至关重要。现有的BoB实现要么需要存储大量中间张量,要么在长序列长度下耗尽GPU显存。我们提出FlashBoB,一种精确且I/O高效的softmax注意力BoB算法,它将计算保持在片上分块内,并避免所有 $N \times N$ 的中间张量,其中 $N$ 为序列长度。关键洞察在于softmax二阶反向传播中的层次化仿射结构:两个行标量通过仿射变换即可确定所有输出。由此得到一种两遍调度方案,其片上静态随机存取存储器(SRAM)占用有界,且片外高带宽存储器(HBM)流量最小。FlashBoB实现了 $\Theta(N^2 d^2/M)$ 的HBM流量($d$ 为头维度,$M$ 为内存大小),并且在标准FlashAttention风格的重计算得分模型下,达到了精确前向注意力的继承型大缓存下界。实验表明,FlashBoB在单块A100 80GB GPU上将精确注意力BoB扩展至 $N=262\text{K}$,而现有的PyTorch精确基线在 $N=16\text{K}$ 时即告失败,同时其速度最高可达FlashBack的 $6.3\times$。这些结果使得精确二阶注意力在先前实现无法高效运行的长上下文序列长度上变得切实可行。
cs.LG / 177 / 2609.24103

Reinforcement Learning under State and Outcome Uncertainty: A Foundational Distributional Perspective

状态与结果不确定性下的强化学习:一种基础性的分布式视角
Preuett, Larry, Zhang, Qiuyi, Ahmad, Muhammad Aurangzeb
Abstract
In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. We address both forms of uncertainty as a first step toward safer algorithms in partially observable settings. Specifically, we extend Distributional Reinforcement Learning (DistRL)-which models the entire return distribution for fully observable domains-to Partially Observable Markov Decision Processes (POMDPs), allowing an agent to learn the distribution of returns for each conditional plan. Concretely, we introduce new distributional Bellman operators for partial observability and prove their convergence under the supremum p-Wasserstein metric. We also propose a finite representation of these return distributions via psi-vectors, generalizing the classical alpha-vectors in POMDP solvers. Building on this, we develop Distributional Point-Based Value Iteration (DPBVI), which integrates psi-vectors into a standard point-based backup procedure-bridging DistRL and POMDP planning. By tracking return distributions, DPBVI lays the foundation for future risk-sensitive control in domains where rare, high-impact events must be carefully managed. We provide source code to foster further research in robust decision-making under partial observability.
Chinese Translation
在许多现实世界的规划任务中,智能体必须应对环境状态的不确定性以及任意给定策略结果的可变性。作为在部分可观测环境中迈向更安全算法的第一步,我们同时处理这两种形式的不确定性。具体而言,我们将分布式强化学习(Distributional Reinforcement Learning, DistRL)——该方法在完全可观测领域中对整个回报分布进行建模——扩展至部分可观测马尔可夫决策过程(POMDP),使智能体能够学习每个条件规划的回报分布。具体来说,我们为部分可观测性引入了新的分布式贝尔曼算子,并在上确界 p-Wasserstein 度量下证明了它们的收敛性。我们还提出通过 psi-向量(psi-vectors)对这些回报分布进行有限表示,从而推广了 POMDP 求解器中经典的 alpha-向量。在此基础上,我们开发了分布式基于点的值迭代算法(Distributional Point-Based Value Iteration, DPBVI),该算法将 psi-向量集成到标准的基于点的回溯(backup)过程中,从而架起了 DistRL 与 POMDP 规划之间的桥梁。通过追踪回报分布,DPBVI 为未来在必须谨慎管理罕见高影响事件的领域中进行风险敏感控制奠定了基础。我们提供了源代码,以促进部分可观测条件下鲁棒决策的进一步研究。
cs.LG / 178 / 2609.24111

SPeaR: Test-Time Adaptation with Steering Primitives for Realigning Representations

SPeaR:基于导向原语的表示重对齐测试时自适应方法
Dip, Muhammad Sudipto Siam, Etemad, Ali
Abstract
Test-time adaptation (TTA) addresses distribution shift using only unlabeled test data. Existing methods typically adapt pretrained models by updating their parameters, limiting both what is adapted and where adaptation can occur within the network. We instead keep the pretrained network frozen and steer its intermediate representations. We introduce SPeaR (Steering Primitive for Realigning Representations), which inserts lightweight learnable modules at stage boundaries and optimizes them directly from the test stream, requiring neither source data nor supervised warm-up. Each primitive is optimized using a gated objective that reduces uncertainty only when adaptation is beneficial, along with a diversity regularizer to prevent collapse, and a multi-depth anchor to stabilize adaptation. We show that steering early representations is the most effective strategy, and that the same primitive transfers across convolutional and Transformer architectures. Across CIFAR-10-C, CIFAR-100-C, and ImageNet-C, SPeaR consistently matches or outperforms methods that adapt orders of magnitude more parameters, remains robust across a wide range of batch sizes, and preserves source-domain performance during continual adaptation.
Chinese Translation
测试时自适应(Test-Time Adaptation, TTA)仅使用无标签测试数据来应对分布偏移。现有方法通常通过更新预训练模型的参数来进行自适应,这既限制了可自适应的内容,也限制了自适应在网络中可以发生的位置。我们转而保持预训练网络冻结,并对其中间表示进行导向调整。我们提出了SPeaR(Steering Primitive for Realigning Representations,用于重对齐表示的导向原语),它在网络的阶段边界插入轻量级可学习模块,并直接从测试数据流中对其进行优化,既不需要源数据,也不需要有监督的预热。每个原语通过一个门控目标函数进行优化,该目标函数仅在有助于自适应时才降低不确定性;同时还引入了防止坍塌的多样性正则化器,以及稳定自适应过程的多深度锚定机制。我们证明,对早期表示进行导向调整是最有效的策略,且同一原语可以在卷积架构和Transformer架构之间迁移。在CIFAR-10-C、CIFAR-100-C和ImageNet-C数据集上,SPeaR持续达到或超越那些自适应参数数量高出数个数量级的方法,在广泛的批量大小范围内保持稳健,并在持续自适应过程中保持源域性能。
cs.LG / 179 / 2609.24117

PAC-Bayesian Meta-Learning for Few-Shot Identification of Linear Dynamical Systems

面向线性动力系统少样本辨识的PAC-贝叶斯元学习
Huang, Chenfeng, Michailidis, George
Abstract
Identifying linear time-invariant (LTI) dynamical systems is challenging when trajectories are short, noisy, or high-dimensional. Traditional system identification typically treats each system independently and cannot exploit shared structure across related systems. We propose PBML-LTI, a PAC-Bayesian meta-learning framework for few-shot LTI system identification that learns a transferable prior over task-specific dynamics while preserving task heterogeneity. Each task corresponds to an unknown LTI system, and the meta-learner uses training trajectories to learn a data-dependent prior over transition matrices. For a new system with limited data, PBML-LTI performs Bayesian adaptation under this prior to obtain a task-specific posterior, providing accurate estimates and principled uncertainty quantification. A key challenge is temporal dependence, since LTI trajectories violate the i.i.d. assumptions underlying most PAC-Bayes meta-learning analyses. We address this with a martingale PAC-Bayes analysis for dependent trajectory losses and derive a support-query predictive-risk bound that motivates a fit-KL meta-training objective. The bound clarifies the roles of empirical fit, posterior complexity, and prior quality in few-shot adaptation under sequential dependence. We further derive corollaries for transition-matrix recovery and multi-step trajectory prediction, connecting uncertainty-aware meta-identification with finite-sample guarantees for dependent dynamical data.
Chinese Translation
当轨迹短、含噪或高维时,线性时不变(LTI)动力系统的辨识极具挑战性。传统系统辨识方法通常独立地处理每个系统,无法利用相关系统之间的共享结构。我们提出PBML-LTI,一种用于少样本LTI系统辨识的PAC-贝叶斯元学习框架,该框架在学习可迁移的、面向任务特定动力学的先验的同时,保留任务的异质性。每个任务对应一个未知的LTI系统,元学习器利用训练轨迹学习一个依赖于数据的转移矩阵先验。对于数据有限的新系统,PBML-LTI在该先验下进行贝叶斯自适应以获得任务特定的后验,从而提供准确的估计和有原则的不确定性量化。一个关键挑战是时间相关性,因为LTI轨迹违背了大多数PAC-贝叶斯元学习分析所依据的独立同分布(i.i.d.)假设。我们针对相关轨迹损失提出了鞅PAC-贝叶斯分析,并推导出一个支持集-查询集预测风险界,由此引出拟合-KL元训练目标。该界阐明了在序列相关性下少样本自适应中经验拟合、后验复杂度与先验质量各自的作用。我们进一步推导了关于转移矩阵恢复和多步轨迹预测的推论,将不确定性感知的元辨识与相关动态数据的有限样本保证联系起来。
cs.LG / 180 / 2609.24141

CLOOPD: Closing the Learner Loop in On-Policy Distillation

CLOOPD:在策略蒸馏中闭合学习器回路
Zheng, Keye, Li, Hanyu, Cheng, Zhan, Gao, Yuan
Abstract
On-policy distillation (OPD) pays twice for each fresh batch: the student generates trajectories and a stronger teacher scores them. Existing methods improve which trajectories are scored and how the teacher signal is constructed, but usually consume it with one actor update. We introduce CLOOPD, a closed-loop framework separating teacher-signal acquisition from student-side realization. CLOOPD selects an adaptive $\alpha$ waypoint inside a KL envelope, freezes the scored batch and its advantages, re-forwards the student after each actor pass, measures realization, and allocates actor work under a separate token budget. The framework includes deterministic two- and three-pass policies, token-priced CLOOPD-TPMR, and a budget-matched control. Across six 300-step runs on an 8-H20 node, every CLOOPD policy improves the one-pass TOP-D anchor at comparable teacher-token scale: macro accuracy rises from 15.41 to 17.78 with CLOOPD-Fixed2 and 19.36 with CLOOPD-Fixed3. At step 100, CLOOPD-Fixed3 reaches 15.35, nearly matching TOP-D at step 300 while using 67.2% fewer teacher-scored tokens and 28.0% fewer GPU-hours. Earlier 8-A100 ablations show adaptive $\alpha$ eliminates observed trust-envelope violations; a third pass adds headroom. These results position CLOOPD as a framework for budgeting how fully students learn from teacher-scored tokens.
Chinese Translation
在策略蒸馏(On-policy distillation, OPD)对每个新批次需要付出双重代价:学生模型生成轨迹,而更强的教师模型对轨迹进行评分。现有方法着力于改进对哪些轨迹进行评分以及如何构建教师信号,但通常仅通过单次actor更新来消耗该信号。我们提出CLOOPD,一个将教师信号获取与学生侧实现相分离的闭环框架。CLOOPD在KL包络内选择自适应的$\alpha$路径点,冻结已评分批次及其优势值,在每次actor遍历后对学生模型重新前向传播,度量实现程度,并在独立的token预算下分配actor计算量。该框架包括确定性的两遍与三遍策略、按token计价的CLOOPD-TPMR,以及预算匹配的对照方法。在单节点8-H20上进行的六次300步实验中,每种CLOOPD策略均在相当的教师token规模下超越了单遍TOP-D基线:CLOOPD-Fixed2将宏观准确率从15.41提升至17.78,CLOOPD-Fixed3则达到19.36。在第100步时,CLOOPD-Fixed3达到15.35,几乎与第300步的TOP-D相当,同时教师评分token减少67.2%,GPU小时数减少28.0%。早期的8-A100消融实验表明,自适应$\alpha$消除了观察到的信任包络违例;第三遍则带来了额外的提升空间。这些结果表明,CLOOPD可作为一个预算框架,用于调控学生从教师评分token中学习的充分程度。
cs.LG / 181 / 2609.24144

Luck Is Not Skill: When Do Paired Rollouts Help Group-Relative RL of LLM Agents?

运气并非技巧:成对轨迹采样何时有助于大语言模型智能体的组相对强化学习?
Sakib, Nazmus
Abstract
Group-relative reinforcement learning compares rollouts of the same prompt, but independent environment noise can obscure these comparisons. We study paired rollouts, which share an event-keyed noise schedule within each group while preserving each rollout's marginal distribution. Pairing removes the between-schedule component of reward-contrast variance, but need not reduce gradient variance. For one-sided grader noise, we derive an exact condition for reduction and give a counterexample in which reward contrasts improve while gradient variance increases. A controlled study trains a 2B tool-use agent under tool faults and grader flips, with three seeds per design. The protocol was registered with a disclosed, previously completed pilot. Under tool faults, pairing improves final noisy-test success by +5.1 percentage points on average, with all three seed differences positive, but misses the registered learning-curve criterion. The criterion is also missed under grader flips: the validation-AUC difference is +0.003 (95% interval [-0.029, +0.033]). A gradient probe on eight distinct checkpoints from two fault-trained trajectories finds lower mean-centered covariance traces under both noise types: 21 to 30% for grader flips and 40 to 63% for tool faults. These finite-sample measurements support the variance mechanism without establishing a general learning-speed benefit. The results distinguish improving reward comparisons, reducing estimator variance, and improving learning.
Chinese Translation
组相对强化学习通过比较同一提示词(prompt)下的多条轨迹(rollouts)来计算优势,但独立的环境噪声可能掩盖这些比较。我们研究了成对轨迹采样(paired rollouts),它在每个组内共享一个基于事件键控的噪声调度,同时保持每条轨迹的边际分布不变。成对采样消除了奖励对比方差中由调度间差异引起的分量,但不一定能降低梯度方差。对于单侧评分器噪声,我们推导了方差降低的精确条件,并给出了一个反例:奖励对比得到改善而梯度方差反而增大。在一项受控研究中,我们在工具故障和评分器翻转两种噪声下训练了一个2B参数的工具使用智能体,每种设计使用三个随机种子。该实验方案已在披露先前已完成的先导实验的前提下进行了预注册。在工具故障条件下,成对采样使最终噪声测试的成功率平均提高5.1个百分点,且三个种子的差异均为正值,但未能达到预注册的学习曲线标准。在评分器翻转条件下同样未达到该标准:验证集AUC差异为+0.003(95%区间为[-0.029, +0.033])。对来自两条故障训练轨迹的八个不同检查点进行的梯度探测发现,两种噪声条件下中心化协方差的迹(trace)均更低:评分器翻转下为21%至30%,工具故障下为40%至63%。这些有限样本的测量结果支持了方差机制,但并未确立普遍意义上的学习速度提升。本研究的结果区分了以下三个方面:改善奖励比较、降低估计量方差以及提升学习效果。
cs.LG / 182 / 2609.24146

Mind or Message? Auditing Theory of Mind in Multi-Agent Social Simulation

心智还是信息?多智能体社会模拟中心理理论(Theory of Mind)的审计研究
Li, Cong, Chen, Cheng, Fung, Thomas, Rossi, Alex, Li, Yi
Abstract
Language model agents are increasingly used to simulate social interaction, and the resulting transcripts read as though the agents understand one another. We ask whether that appearance rests on a model of the partner's mind or on the surface record of what the partner said. We build a social simulation in which both questions have exact answers: 40 multi-issue negotiations whose hidden preference weights and whose full Pareto frontier are known by construction. Two model families negotiate across 160 dyads, every transcript is frozen before any measurement, and 2880 counterfactual probes then hold the evidence byte identical while moving one factor at a time: the reader's own stake, the partner's tone, an identity label, and the order of recursion. The agents are socially fluent and economically poor. They reach agreement in 96.2% of dyads with 0 protocol failures, yet only 0.7% of deals land on the Pareto frontier, they leave 20.5% of the available joint value unclaimed, and they miss the one issue on which their interests are perfectly aligned in 76.6% of deals; on the frontier and on that aligned issue, a package drawn at random from the set both sides would accept does as well. The probes locate the failure. Swapping only the reader's own payoff sheet, while the partner's words and offers stay identical, moves the inferred top priority by 15.0 percentage points, which is egocentric projection rather than inference, while a tone rewrite moves it by 5.3 percentage points and an identity label by 0.0. Most tellingly, an agent predicts what its partner believes about it 72.5% of the time while that partner's belief is itself correct only 51.2% of the time: the agents track the conversation far better than they track the mind behind it.
Chinese Translation
语言模型智能体被越来越多地用于模拟社会互动,由此产生的对话记录读起来仿佛智能体彼此理解。我们探究这种表象究竟是基于对对方心智的建模,还是仅基于对方言语的表层记录。我们构建了一个两类问题都有精确答案的社会模拟:40场多议题谈判,其隐藏的偏好权重与完整的帕累托前沿在构造上均已知。两个模型家族在160个配对中进行谈判,所有对话记录在任何测量之前被冻结,随后2880次反事实探针在保持证据逐字节相同的前提下每次只改变一个因素:读者自身的利益、对方的语气、身份标签以及递归层级。结果显示,这些智能体在社交上流利,但在经济上表现糟糕。它们在96.2%的配对中达成协议且无协议失败(0次),然而仅有0.7%的交易落在帕累托前沿上,它们放弃了20.5%的可用联合价值,并在76.6%的交易中错失了双方利益完全一致的那一个议题;而在帕累托前沿及该一致议题上,从双方都能接受的集合中随机抽取的方案表现同样好。探针实验定位了失败原因。在对方的言辞与提议完全不变的情况下,仅替换读者自身的收益表,会使推断出的最优先议题移动15.0个百分点,这表明这是自我中心投射而非推理;而语气改写仅使其移动5.3个百分点,身份标签则移动0.0个百分点。最能说明问题的是,智能体预测其伙伴对自身看法的准确率为72.5%,而该伙伴的信念本身仅有51.2%的时候是正确的:智能体对对话内容的追踪远好于对其背后心智的追踪。
cs.LG / 183 / 2609.24150

Acceptance-Aware Draft Model Training for Speculative Decoding

面向推测解码的接受度感知草稿模型训练
Xia, Tianhua, Ganesan, Mugilan, Feng, Yifei, Wang, Haiyu, Egger, Maximilian, Zhang, Sai Qian
Abstract
Speculative decoding accelerates large language model (LLM) inference by using a lightweight draft model to generate multiple candidate tokens that are verified by the target model in a single forward pass. Its speedup is largely determined by the acceptance length, yet existing draft-model training methods mainly optimize cross-entropy or Kullback-Leibler (KL) divergence as proxies. These objectives encourage distribution matching but do not directly optimize acceptance length, and the acceptance mechanism also differs between greedy and sampling-based decoding. In this work, we propose acceptance-length-aware training losses that directly optimize the expected number of accepted tokens within a speculative window. For greedy verification, we derive an expected accepted length (EAL) loss that explicitly maximizes expected acceptance length. For sampling-based decoding, we introduce a window total variation (WTV) loss that optimizes the overlap between temperature-scaled draft and target distributions while accounting for sequential acceptance dependencies. Both objectives can be further combined with a group-relative reinforcement learning stage (GRPO) using simulated acceptance length as the reward. Experiments across different target and draft models, tasks, and decoding settings show that our losses consistently improve acceptance length over KL-based training. WTV provides particularly strong gains under sampling-based decoding, while EAL better matches greedy verification. These results show that directly optimizing the acceptance objective, with losses tailored to the decoding mode, is more effective than conventional distribution-matching objectives.
Chinese Translation
推测解码(speculative decoding)通过使用轻量级草稿模型生成多个候选token,并由目标模型在单次前向传播中加以验证,从而加速大语言模型(LLM)的推理。其加速效果在很大程度上取决于接受长度,然而现有的草稿模型训练方法主要优化交叉熵或Kullback-Leibler(KL)散度作为替代目标。这些目标函数虽然促进了分布匹配,但并未直接优化接受长度,且接受机制在贪心解码与基于采样的解码之间也存在差异。在本工作中,我们提出了感知接受长度的训练损失,直接优化推测窗口内被接受token的期望数量。对于贪心验证,我们推导出期望接受长度(EAL)损失,显式地最大化期望接受长度。对于基于采样的解码,我们引入窗口总变差(WTV)损失,在考虑顺序接受依赖性的同时,优化经温度缩放后的草稿分布与目标分布之间的重叠程度。这两种目标均可进一步与采用模拟接受长度作为奖励的组相对强化学习阶段(GRPO)相结合。在不同目标模型与草稿模型、不同任务及不同解码设置上的实验表明,我们的损失函数相比基于KL的训练能够持续提升接受长度。WTV在基于采样的解码下带来了尤为显著的增益,而EAL则更契合贪心验证。这些结果表明,直接优化接受目标并针对解码模式设计相应的损失,比传统的分布匹配目标更为有效。
cs.LG / 184 / 2609.24197

H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache

H-Spec:无需起草端KV缓存的并行推测解码
Jiang, Weifan, Chitty-Venkata, Krishna Teja, Flynn, Megan, Meyerson, Reed, Qi, Zhenting, Wu, Tianyu, Kurtic, Eldar, Yu, Minlan, Marques, Alexandre
Abstract
Speculative decoding losslessly accelerates large language model inference by having a lightweight draft model predict future tokens for verification by the target model. Recent block diffusion drafters further reduce drafting latency by predicting multiple tokens in parallel. However, existing block drafters project target hidden states at every input position into a separate drafter-side KV cache, incurring per-request memory and KV-write overhead that grow with concurrency; directly reusing target KVs in place removes this cache but fails to sustain draft quality throughout the block. We propose a hybrid target-context injection method that complements direct target KV reuse with target hidden states only at the last input position, requiring no separate drafter-side KV cache. Building on this design, we propose H-Spec, a hybrid Mamba-attention parallel drafter that consumes the two target-context sources through complementary modules. Mamba modules are initialized with projected last-token target hidden states, while attention modules reuse target KVs in place. Despite its recurrent formulation, Mamba's parallel scan allows H-Spec to preserve block-parallel drafting. Across three target models and diverse tasks, H-Spec improves over the best baseline by 5.0--13.3% in mean accepted length and 5.3--12.6% in batch-size-1 inter-token latency speedup. Under concurrent serving, H-Spec consistently achieves higher throughput while maintaining lower KV cache utilization than baselines across evaluated concurrency levels.
Chinese Translation
推测解码通过让轻量级起草模型(draft model)预测未来令牌并由目标模型进行验证,从而无损地加速大语言模型推理。近期的块扩散起草器通过并行预测多个令牌进一步降低了起草延迟。然而,现有的块起草器在每个输入位置都将目标模型隐藏状态投影到一个独立的起草端KV缓存中,产生了随并发度增长而增加的每请求内存开销和KV写入开销;直接就地复用目标KV虽然可以消除该缓存,但无法在整个块中保持起草质量。我们提出了一种混合目标上下文注入方法,仅在最后一个输入位置用目标隐藏状态来补充直接的目标KV复用,无需独立的起草端KV缓存。基于这一设计,我们提出了H-Spec,一种混合Mamba-注意力并行起草器,它通过互补模块来消费这两种目标上下文来源:Mamba模块使用投影后的最后令牌目标隐藏状态进行初始化,而注意力模块则就地复用目标KV。尽管采用循环形式化表述,Mamba的并行扫描使H-Spec能够保持块并行起草。在三个目标模型和多样化任务上,H-Spec在平均接受长度上比最佳基线提升了5.0%–13.3%,在批大小为1的令牌间延迟加速比上提升了5.3%–12.6%。在并发服务场景下,H-Spec在所评估的各并发级别上始终保持比基线更高的吞吐量,同时维持更低的KV缓存利用率。
cs.LG / 185 / 2609.24202

Opinion Leader Dynamics: How Sparse Attention Shapes Token Clustering

意见领袖动力学:稀疏注意力如何塑造词元聚类
Liu, Jingkun, Song, Yue
Abstract
Sparse attention reduces the quadratic cost of global self-attention while retaining strong empirical performance, but how its restricted interactions shape the evolution of token representations remains theoretically underexplored. Modeling tokens as particles on the unit sphere, we introduce opinion leader dynamics, a framework that identifies two mechanisms through which token groups converge internally while maintaining distinct limiting directions. In the explicit model, fixed representatives induce a potential that attracts tokens toward distinct local maxima. In the implicit model, disconnected interaction groups evolve toward separate consensus directions. We formulate both models as reverse Wasserstein gradient flows and establish exponential convergence under suitable conditions. We further connect these theoretical predictions to token evolution in frontier sparse-attention LLMs that motivate our framework. Across four benchmarks, Kimi-K3, MiniMax-M3, and DeepSeek-V4-Flash consistently exhibit clearer cluster separation and higher clustering scores than the dense-attention model GLM-4.7-Flash in projected token representations. These observations support the relevance of the predicted multiple-group structure to trained frontier LLMs, while finite-particle simulations illustrate the theoretical convergence behavior. Together, our results connect restricted token interactions to distinct group-level attractors, providing a dynamical account of how sparse attention can support alignment within groups while preserving separation between them.
Chinese Translation
稀疏注意力在保持强大经验性能的同时,降低了全局自注意力的二次方计算开销,但其受限的交互如何塑造词元表示的演化,在理论上仍缺乏充分研究。我们将词元建模为单位球面上的粒子,提出了意见领袖动力学(opinion leader dynamics)框架,识别出词元群体内部趋同同时保持不同极限方向的两类机制。在显式模型中,固定的代表元诱导出一个势函数,将词元吸引向不同的局部极大值点。在隐式模型中,互不连通的交互群体各自朝不同的共识方向演化。我们将两种模型表述为逆 Wasserstein 梯度流,并在适当条件下建立了指数收敛性。我们进一步将这些理论预测与激发本框架的前沿稀疏注意力大语言模型中的词元演化联系起来。在四个基准测试中,Kimi-K3、MiniMax-M3 和 DeepSeek-V4-Flash 在投影后的词元表示上均比稠密注意力模型 GLM-4.7-Flash 表现出更清晰的聚类分离和更高的聚类得分。这些观察结果支持了所预测的多群体结构与训练后的前沿大语言模型之间的相关性,同时有限粒子模拟展示了理论上的收敛行为。总之,我们的结果将受限的词元交互与不同的群体层面吸引子联系起来,从动力学角度解释了稀疏注意力如何在保持群体间分离的同时支持群体内部的对齐。
cs.LG / 186 / 2609.24209

Displacement Geometry Captures Platonic Shared Reality Across Models and Modalities

位移几何捕捉跨模型与跨模态的柏拉图共享现实
Shang, Chenming, Tang, Yujin, Yang, Jun Jie Ou, Xu, Ruize, Breuer, Adam, Singh, Nikhil
Abstract
The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent work finds only weak pointwise similarity between models. In this paper, we show that what models share is not the location of samples in representation space, but the directions (displacement vectors) between them. Under a single orthogonal alignment--rotation and reflection only--these displacement vectors are substantially preserved across 44 independently trained vision and language encoders spanning modalities and asymmetric capability pairs, consistent with the PRH evidence. The samples' absolute positions are not, consistent with recent counter-evidence. Both arise from a single decomposition: representations split into a shared semantic component that is linearly aligned across models, and a private capability component that is not. We trace this geometry to concept-level structure: within a model, parent concepts are orthogonal to their child variation vectors; across models, concept displacements are parallel. Our theory falsifiably predicts (and experiments confirm) that fine-tuning preserves pointwise similarity but collapses displacement, and that relational distillation does the opposite. A major implication is that, because semantics align linearly but capabilities do not, capabilities can be imported from one model to another using a single cached forward pass through the source. We call this Shadow Casting. As a proof of concept, our SHADOWCLIP instantiation outperforms strong fine-tuned baselines at orders of magnitude less compute. A cache can be released alongside open model weights, letting one model's capabilities be downloaded and imported into any number of other models without fine-tuning.
Chinese Translation
柏拉图表征假说(Platonic Representation Hypothesis, PRH)声称,独立训练的模型会收敛到关于现实的共享统计模型,但近期研究仅发现模型之间存在微弱的逐点相似性。在本文中,我们表明模型所共享的并非样本在表征空间中的位置,而是样本之间的方向(位移向量)。在单一的正交对齐——仅包含旋转和反射——之下,这些位移向量在44个独立训练的、跨越模态和能力不对称配对的视觉与语言编码器之间被显著保留,这与支持PRH的证据一致;而样本的绝对位置则未被保留,这与近期的反面证据一致。两者均源自同一种分解:表征被拆分为一个跨模型线性对齐的共享语义成分,以及一个未对齐的私有能力成分。我们将这一几何结构追溯至概念层面的结构:在单个模型内部,父概念与其子概念变化向量正交;跨模型之间,概念位移则是平行的。我们的理论做出了可证伪的预测(并得到实验验证):微调能保持逐点相似性但会破坏位移结构,而关系蒸馏则恰好相反。一个重要启示是:由于语义是线性对齐的而能力并非如此,因此可以通过对源模型进行一次缓存的前向传播,将其能力移植到另一个模型中。我们将这一方法称为Shadow Casting(影子投射)。作为概念验证,我们的SHADOWCLIP实现以低若干个数量级的计算量超越了强大的微调基线。缓存可以随开源模型权重一同发布,使一个模型的能力能够被下载并导入任意数量的其他模型,而无需微调。
cs.LG / 187 / 2609.24233

Adaptive Forgetting for Nonstationary Optimization: Towards Robust EEG Decoding

面向非平稳优化与鲁棒脑电解码的自适应遗忘机制
Zhu, Hongyu, Chen, Lin, Chen, Jing, Zhou, Yuting, Shang, Mingsheng
Abstract
Electroencephalography (EEG) provides non-invasive monitoring of brain activity and is widely used in emotion recognition, motor imagery and sleep staging. Although within-subject decoding has achieved considerable progress, cross-subject generalization remains a central challenge in practical applications. EEG decoders are typically trained with Adam/AdamW under a fixed second-moment decay coefficient, even though cross-subject learning involves low signal-to-noise ratios, subject variability, and gradient nonstationarity. A fixed coefficient implicitly assumes that gradient statistics are homogeneous across layers and time, which can limit model's adaptability to cross-subject EEG signals and degrade generalization. To address these issues, we propose AFOR, a tensor-wise adaptive optimizer that converts the fixed second-moment decay coefficient into a dynamic coefficient estimated online from local gradient state. AFOR combines a Residual-Alignment Signal Scorer (RASS) and an Adaptive Forgetting Controller (AFC). RASS summarizes local gradient residuals and directional agreement into a signal-quality score, and AFC maps this score through self-referential normalization to a bounded per-step decay coefficient, with cumulative-product initialization correction maintaining consistency under time-varying decay. Under a strict cross-subject protocol on three EEG benchmarks that cover three representative fields, AFOR achieves the best average performance among the compared optimizers, improving the mean test accuracy over Adam by 3.00%, 2.07%, and 4.38%, respectively.
Chinese Translation
脑电图(EEG)提供了一种非侵入式的脑活动监测手段,广泛应用于情绪识别、运动想象和睡眠分期等领域。尽管被试内解码已取得长足进展,但跨被试泛化仍然是实际应用中的核心挑战。尽管跨被试学习涉及低信噪比、被试间差异以及梯度非平稳性,脑电解码器通常仍采用固定二阶矩衰减系数的 Adam/AdamW 进行训练。固定系数隐含地假设梯度统计特性在各层和时间上都是均匀的,这会限制模型对跨被试脑电信号的自适应能力,并损害泛化性能。为解决这些问题,我们提出了 AFOR,一种张量级自适应优化器,它将固定的二阶矩衰减系数转化为根据局部梯度状态在线估计的动态系数。AFOR 结合了残差对齐信号评分器(RASS)和自适应遗忘控制器(AFC)。RASS 将局部梯度残差和方向一致性汇总为信号质量评分,AFC 通过自参照归一化将该评分映射为有界的每步衰减系数,并通过累积乘积初始化修正保证在时变衰减下的一致性。在涵盖三个代表性领域的三个脑电基准数据集上,采用严格的跨被试协议,AFOR 在所比较的优化器中取得了最佳平均性能,其平均测试准确率相对 Adam 分别提升了 3.00%、2.07% 和 4.38%。
cs.LG / 188 / 2609.24241

Hessian Rank Constraint for Learning Structure of Nonlinear Latent Variable Models

用于学习非线性隐变量模型结构的Hessian秩约束
Li, Zijian, Cai, Ruichu, Xie, Feng, Dong, Xinshuai, Dai, Haoyue, Sun, Yuewen, Zheng, Yujia, Chen, Guangyi, Hu, Yingyao, Zhang, Kun
Abstract
Uncovering latent variables and their causal relations from observed data is a fundamental yet challenging problem. Existing methods often rely on restrictive assumptions, such as linear relations or invertible mixing functions. To better address this problem under general nonlinear mixing procedures, we propose a condition called the cross-Hessian Rank Constraint (HRC), which serves as a primitive rank-based tool for nonlinear latent causal discovery. In particular, we show that a rank-based property arises from the cross-Hessian of the observed-data log-density in the nonlinear case, revealing information about the latent variables, and reduces to the Tetrad constraints in the linear Gaussian case. More specifically, when two groups of observed variables are d-separated by a set of lower-dimensional latent variables, the rank of this cross-Hessian is equal to the dimension of the latent variables, under a mild affine derivative assumption on the conditional log-density derivatives. This assumption can be naturally satisfied when the noise level is low or the relevant nonlinearity is moderate. As a downstream application, we instantiate HRC in the pure one-factor measurement setting for locating latent variables and recovering their causal structure up to Markov equivalence. Experimental results on synthetic and real-world datasets support the theoretical claims.
Chinese Translation
从观测数据中发现隐变量及其因果关系是一个基础但具有挑战性的问题。现有方法通常依赖于较强的假设,例如线性关系或可逆的混合函数。为了在一般的非线性混合过程下更好地解决这一问题,我们提出了一种称为交叉Hessian秩约束(Hessian Rank Constraint, HRC)的条件,它可作为基于秩的工具,用于非线性隐变量因果发现。特别地,我们证明在非线性情形下,观测数据对数密度的交叉Hessian会产生一种基于秩的性质,该性质揭示了隐变量的相关信息,并且在线性高斯情形下可退化为Tetrad约束。更具体地,当两组观测变量被一组低维隐变量d-分离时,在对条件对数密度导数施加温和的仿射导数假设下,该交叉Hessian的秩等于隐变量的维度。当噪声水平较低或相关非线性程度适中时,该假设可自然满足。作为下游应用,我们在纯单因子测量模型中实例化HRC,用于定位隐变量并在马尔可夫等价类意义下恢复其因果结构。在合成数据集和真实数据集上的实验结果支持了我们的理论结论。
cs.LG / 189 / 2609.24249

Reinforcement Learning Inspired Black-box Adversarial Attacks for Computer Vision

受强化学习启发的计算机视觉黑盒对抗攻击方法
Krone, Florian, Hoemann, Elena, Hallerbach, Sven
Abstract
Neural networks, both convolution or transformer based, are essential for modern computer vision systems. However, they are vulnerable to small perturbations, almost imperceptible to humans, which significantly alter the model's prediction. These adversarial attacks are often considered to be a significant threat to the implementation of neural networks in safety-critical applications. Most attacks utilize the white-box threat model and therefore require full access to the target model, making them unrealistic to use in practice. We propose a novel approach under the more realistic black-box threat model that utilizes concepts from reinforcement learning to optimize perturbations with a non-differentiable target model. Reinforcement learning algorithms have already been optimized to be query efficient, making them an ideal starting point when designing black-box adversarial attacks. We show the success of our reinforcement learning inspired black-box adversarial attack (RIBA) in generating adversarial perturbations using only a small number of queries to the target model, by comparing it to state of the art attacks on different models on the Cifar10 and ImageNet data sets. RIBA takes $25.4\%$ fewer median queries to generate attacked images against a ResNet-18 on Cifar10 and $22.5\%$ fewer median queries to fool a Vit-B/16 model on ImageNet. Additionally, we demonstrate that RIBA can match the performance of white-box attacks on an adversarially trained model.
Chinese Translation
神经网络(无论是基于卷积还是基于Transformer的模型)对现代计算机视觉系统至关重要。然而,它们容易受到微小扰动的影响,这些扰动几乎不为人类所察觉,却能显著改变模型的预测结果。此类对抗攻击通常被认为是对神经网络在安全关键型应用中部署的重大威胁。大多数攻击采用白盒威胁模型,因此需要对目标模型的完全访问权限,这使其在实践中难以实现。我们提出了一种在更为现实的黑盒威胁模型下的新方法,该方法利用强化学习的概念在不可微的目标模型上优化扰动。强化学习算法已被优化以具备查询高效性,这使其成为设计黑盒对抗攻击的理想起点。通过在Cifar10和ImageNet数据集上与不同模型上的最先进攻击方法进行比较,我们展示了受强化学习启发的黑盒对抗攻击方法(RIBA)在仅使用少量目标模型查询的情况下生成对抗扰动的成功。RIBA在Cifar10上对ResNet-18生成攻击图像所需的中位查询次数减少了25.4%,在ImageNet上欺骗Vit-B/16模型所需的中位查询次数减少了22.5%。此外,我们还证明了RIBA在对抗训练模型上的性能可以与白盒攻击相媲美。
cs.LG / 190 / 2609.24250

Explainable Predictive Condition-based Maintenance of Naval-Propulsion Systems using Fuzzy Logic

基于模糊逻辑的船舶推进系统可解释预测性状态维护
Kalogeropoulos, Dionisis, Sovatzidi, Georgia, Kalozoumis, Panagiotis G., Iakovidis, Dimitris K.
Abstract
The shipping industry has a significant impact on the global economy, emphasizing the need for operational availability and safety through the use of effective maintenance techniques. During the last decades, predictive maintenance (PdM) has emerged as a promising solution compared to the existing conventional maintenance systems. This is because it offers several advantageous functions, such as damage predictions for vessel components, reduced downtime, improved and extended life of machinery, as well as higher safety during voyages. However, existing methodologies developed for performing PdM do not provide explanations of their results to users, so that they can understand the failures that may occur. To address this limitation, this paper proposes a novel framework based on a fuzzy decision tree and a deep residual neural network, aiming to perform explainable PdM on naval vessels. The proposed framework is able to generate fuzzy local rules based on the dataset used, and can provide explanations of its outcomes, using cause-and-effect relationships, in a way that are understandable to users, thereby gaining their trust. Experiments using a publicly available dataset demonstrate the effectiveness of the proposed framework, as it achieves an accuracy of 99.24%.
Chinese Translation
航运业对全球经济具有重要影响,这凸显了通过有效的维护技术保障运营可用性和安全性的必要性。在过去几十年中,预测性维护(PdM)相较于现有的传统维护系统已成为一种有前景的解决方案。这是因为它提供了若干优势功能,例如对船舶部件的损伤预测、减少停机时间、改善并延长机械寿命,以及在航行过程中提供更高的安全性。然而,现有的用于执行PdM的方法无法向用户解释其结果,使用户无法理解可能发生的故障。为解决这一局限性,本文提出了一种基于模糊决策树和深度残差神经网络的新型框架,旨在对舰船实现可解释的PdM。该框架能够基于所用数据集生成模糊局部规则,并利用因果关系对其结果进行解释,使用户易于理解,从而赢得用户的信任。基于公开数据集的实验验证了所提框架的有效性,其准确率达到99.24%。
cs.LG / 191 / 2609.24259

MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

MemCalib:大语言模型智能体内存使用的基准测试与优化
Cao, Ruike, Zhao, Fanyu, Yao, Fugen, Dong, Liang, Xu, Jian, Jiang, Guanjun, Zhao, Yifei, Zhang, Han, Xiao, Li
Abstract
The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group relative policy optimization and on-policy self-distillation, further reveal a clear directional skew: trained models improve in one direction while deteriorating in the other. We therefore propose MemCalib-RL, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation. Results across model families and scales (Qwen3-8B, Ministral-3-8B-Instruct, and Qwen3.5-35B-A3B) show that MemCalib-RL achieves the best overall performance while better balancing over-use and under-use, with gains generalizing beyond MemCalib in external benchmark evaluation. Further experiments support its design choices and robustness and provide insight into its training dynamics.
Chinese Translation
智能体记忆的有效性最终取决于底层大语言模型(LLM)是否能够赋予上下文中的每条记忆对其响应以适当程度的影响。然而,这一能力在很大程度上一直被忽视。为了评估这一能力,我们提出了MemCalib,一个基于真实记忆系统场景的基准,用于评估记忆使用并推动优化算法的发展。在MemCalib测试集上的结果表明,前沿的开源和闭源模型都难以恰当地使用记忆。它们经常过度使用或使用不足记忆,而不是使每个命题的实际使用程度与其目标水平相匹配,从而导致有偏差的低质量响应。使用常见后训练算法(包括组相对策略优化GRPO和在线策略自蒸馏)的实验进一步揭示出明显的方向性偏斜:训练后的模型在一个方向上有所改进,而在另一个方向上反而恶化。因此,我们提出了MemCalib-RL,一种有序的双向反事实信用分配算法,它将过度使用与使用不足的信号分离开来,并通过精确的原子消融将其信用定位到响应的词元上。在多个模型家族和规模(Qwen3-8B、Ministral-3-8B-Instruct和Qwen3.5-35B-A3B)上的结果表明,MemCalib-RL在更好地平衡过度使用与使用不足的同时,取得了最佳的整体性能,且其增益在外部基准评估中能够泛化到MemCalib之外。进一步的实验支持了其设计选择与鲁棒性,并为其训练动态提供了洞见。
cs.LG / 192 / 2609.24278

High-Dimensional Online Change Point Detection with Adaptive Thresholding and Interpretability

具有自适应阈值与可解释性的高维在线变点检测
Jacob, Sven, Prenkaj, Bardh, Shao, Weijia, Kasneci, Gjergji
Abstract
Change point detection (CPD) identifies abrupt and significant changes in sequential data, with applications in human activity recognition, financial markets, cybersecurity, manufacturing, and autonomous systems. Traditional CPD methods often face computational challenges in high-dimensional settings and typically provide limited explanations for detected changes, which can restrict their practical usability. This paper introduces a CPD framework that improves scalability and interpretability by leveraging the Sliced Wasserstein (SW) distance. Our contributions are fourfold: (1) we transform multivariate sequential data into one-dimensional scores using the SW distance, making the resulting representation compatible with existing CPD methods; (2) we analyze the distributional behavior of random slices of the SW distance and show that, under suitable assumptions, they can be approximated by a Gamma distribution, providing a principled basis for threshold calibration; (3) we propose a self-adapting online CPD algorithm that combines this SW-based score with an adaptive quantile-based threshold; (4) we introduce a model-specific framework for generating contrastive explanations for annotated change points. Empirically, our method reduces false positives by at least $48\%$ on average compared with popular online and offline CPD baselines, while maintaining competitive or superior detection performance. Code is available at https://github.com/jsve96/SWCPD_Code. At the same time, it produces interpretable change-point annotations, making it practical for deployment in high-stakes applications.
Chinese Translation
变点检测(Change Point Detection, CPD)旨在识别序列数据中突然且显著的变化,其应用涵盖人体活动识别、金融市场、网络安全、制造业以及自主系统。传统CPD方法在高维场景下常面临计算挑战,且对检测到的变化通常只能提供有限的解释,这限制了其实际可用性。本文提出一种利用切片Wasserstein(Sliced Wasserstein, SW)距离来提升可扩展性与可解释性的CPD框架。我们的贡献有四点:(1)利用SW距离将多元序列数据变换为一维分数,使其表示形式与现有CPD方法兼容;(2)分析了SW距离随机切片的分布行为,并证明在适当假设下其可由Gamma分布近似,从而为阈值校准提供了有原则的依据;(3)提出一种自适应在线CPD算法,将基于SW的分数与自适应分位数阈值相结合;(4)引入一个针对特定模型的框架,为标注出的变点生成对比性解释。实验表明,与流行的在线和离线CPD基线方法相比,我们的方法平均可将误报至少降低48%,同时保持相当或更优的检测性能。代码可在 https://github.com/jsve96/SWCPD_Code 获取。同时,该方法能够生成可解释的变点标注,使其在高风险应用的实际部署中具有实用性。
cs.LG / 193 / 2609.24289

TTSE: A Two-Track Online Self-Evolution Framework

TTSE:一种双轨在线自进化框架
Pei, Ruimin, Wu, Yongkang, Zheng, Shangyi, Zhang, Yaqing, Li, Deyang, Tao, Jianjun, Zhang, Xinyu, Zhang, Xiang
Abstract
As Large Language Model (LLM) agents are applied in continuously interactive environments, driving the evolution of their own capabilities becomes a core problem for achieving long-term autonomy. Currently, environmental knowledge is typically treated as an external fixed input rather than as part of the agent's ongoing evolution. Reinforcement learning methods usually optimize policies through environmental interaction but tend to adapt only to fixed task distributions or single environments. This paper proposes TTSE (Two-Track Self-Evolution), a dual-track online self-evolution framework that separates evolving knowledge into FACT (environmental facts, whose reliability is continuously verified through interaction evidence) and TIP (task-conditioned implementation procedures). From a decision-theoretic perspective, we decompose the agent's excess risk into environment-representation regret and conditional-execution regret, characterize the conditions under which environment-conditioned policies strictly outperform condition-agnostic policies, and bound the downstream risk in terms of FACT identification error and cross-condition mismatch cost. In practice, TTSE's ablation experiments on GDPevo validate the advantage of dual-track evolution. On the classic agent task benchmarks ALFWorld and ScienceWorld, TTSE further demonstrates superior task adaptation. Moreover, TTSE is broadly compatible with existing skill self-evolution methods; combined with the Bayesian-Agent algorithm, a single-track ablation validates the dual-track advantage, substantially improving the aggregate score across the five major domains of SOPBench over three independent repetitions. Finally, on the real end-to-end task benchmark PinchBench, TTSE is integrated into a general agent framework via retrieval-based injection and stably outperforms the baseline across three independent runs.
Chinese Translation
随着大语言模型(LLM)智能体被应用于持续交互的环境中,推动其自身能力的进化成为实现长期自主性的核心问题。目前,环境知识通常被视为外部的固定输入,而非智能体持续进化的一部分。强化学习方法通常通过环境交互来优化策略,但往往只能适应固定的任务分布或单一环境。本文提出TTSE(Two-Track Self-Evolution,双轨自进化),一种双轨在线自进化框架,将演化知识分离为FACT(环境事实,其可靠性通过交互证据持续验证)和TIP(任务条件化的执行流程)。从决策理论的视角,我们将智能体的超额风险分解为环境表征遗憾和条件执行遗憾,刻画了环境条件化策略严格优于条件无关策略的条件,并以FACT识别误差和跨条件失配代价对下游风险进行界定。在实践中,TTSE在GDPevo上的消融实验验证了双轨进化的优势。在经典智能体任务基准ALFWorld和ScienceWorld上,TTSE进一步展现出卓越的任务适应能力。此外,TTSE与现有的技能自进化方法具有广泛兼容性;结合Bayesian-Agent算法,单轨消融实验验证了双轨优势,在三次独立重复实验中显著提升了SOPBench五大领域的总评分。最后,在真实端到端任务基准PinchBench上,TTSE通过基于检索的注入方式集成到通用智能体框架中,在三次独立运行中均稳定优于基线方法。
cs.LG / 194 / 2609.24298

KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation

KV-COBRA:基于协同优化的比特-秩分配的KV缓存压缩
Ha, Sihyeon, Lee, Jaeho, Jeon, Yo-Seb
Abstract
What limits KV-cache compression at extreme bit-rates? We argue that it is not the choice of compression scheme, but how its budget is allocated across attention heads. Existing methods apply rank and bit-width uniformly, ignoring that each head has a different optimal mix of rank truncation and quantization. We show that co-optimizing rank and bit-width per head, using only standard low-rank projection and scalar quantization, dominates uniform allocation, with the largest gains at low bit-rates. Our method, KV-COBRA (Co-Optimized Bit-Rank Allocation), formalizes this as a resource-allocation problem: it balances rank-truncation loss against quantization loss within each head, then redistributes budget across heads to minimize total distortion. A fused Hadamard rotation equalizes per-channel variance, and reordering the SVD basis by attention-KL importance makes the solver query-aware. The same allocator extends to joint $K{+}V$ compression. On perplexity, zero-shot, and long-context benchmarks from $0.5$ to $4$ bits per dimension (bpd), KV-COBRA shows the smallest accuracy degradation among evaluated methods at low bpd, with no per-token overhead.
Chinese Translation
是什么限制了KV缓存(KV-cache)在极低比特率下的压缩效果?我们认为,关键不在于压缩方案的选择,而在于其预算如何在不同注意力头(attention head)之间分配。现有方法对秩(rank)和比特宽度采用统一设置,忽略了每个注意力头在秩截断与量化之间存在不同的最优组合。我们证明,仅需使用标准的低秩投影和标量量化,对每个注意力头协同优化秩与比特宽度,即可优于统一分配策略,且在低比特率下收益最大。我们的方法KV-COBRA(协同优化的比特-秩分配,Co-Optimized Bit-Rank Allocation)将该问题形式化为一个资源分配问题:在每个注意力头内部权衡秩截断损失与量化损失,然后在各注意力头之间重新分配预算以最小化总失真。融合的Hadamard旋转使各通道方差均衡化,而通过注意力KL散度重要性对SVD基进行重排序,使求解器具备查询感知能力。同一分配器还可扩展至K和V的联合压缩。在每维0.5至4比特(bpd)的困惑度、零样本和长上下文基准测试中,KV-COBRA在低bpd下表现出所有被评估方法中最小的精度损失,且无逐token的额外开销。
cs.LG / 195 / 2609.24303

SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration

SupportCal:基于参考支持与佐证的后训练大语言模型无标签校准方法
Luo, Linhan, Lin, Lequan, Shi, Dai, Chen, Feng, Hernández-Lobato, José Miguel, Gao, Junbin
Abstract
Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident than their corresponding pretrained language models (PLMs). Because task-specific labeled calibration data can be costly or unavailable, the corresponding pretrained PLM provides a natural label-free reference for post-hoc calibration. Prior agreement-gated PLM-referenced calibration fits a scalar temperature using only examples on which the PoLM and its PLM reference agree, excluding disagreement examples because direct alignment can drive the fitted temperature excessively high and induce under-confidence. We revisit this binary treatment. A controlled reintroduction diagnostic reveals a non-monotonic aggregate effect: admitting a moderate fraction of disagreement examples can improve calibration, whereas the benefit diminishes as unit-weight inclusion approaches the full disagreement set. We introduce SupportCal, a label-free post-hoc method that retains agreement examples at unit weight and assigns disagreement examples continuous weights based on the own-base PLM's relative support and corroboration from pretrained references selected from a size-compatible candidate pool. We further characterize when the resulting weighted objective admits a finite optimal temperature. Across MedMCQA and MathQA, SupportCal yields lower ECE than the agreement-only baseline for nearly all evaluated target-model configurations; supplementary TweetEval Sentiment results show the same pattern on a fixed-label classification task.
Chinese Translation
后训练通常能提升任务性能,但可能损害置信度校准,使得后训练语言模型比其对应的预训练语言模型更加过度自信。由于任务特定的有标签校准数据可能成本高昂或无法获得,相应的预训练语言模型为事后校准提供了一种天然的无标签参考。先前基于一致性门控的PLM参考校准方法仅使用后训练模型与其PLM参考一致的样本来拟合标量温度参数,而排除了不一致样本,因为直接对齐不一致样本会使拟合温度过高并导致置信度不足。我们重新审视了这种二元处理方式。一项受控重新引入的诊断实验揭示了一种非单调的总体效应:纳入适度比例的不一致样本可以改善校准,而当等权重纳入趋向于全部不一致样本时,其收益逐渐消失。我们提出了SupportCal,这是一种无标签的事后校准方法,它以单位权重保留一致样本,并根据自身基座PLM的相对支持度以及从规模相匹配的候选池中选取的预训练参考模型的佐证,为不一致样本赋予连续权重。我们进一步刻画了所得加权目标存在有限最优温度的条件。在MedMCQA和MathQA数据集上,SupportCal在几乎所有评估的目标模型配置下都比仅使用一致样本的基线方法获得了更低的期望校准误差(ECE);补充的TweetEval情感分类实验结果在固定标签分类任务上呈现出相同的规律。
cs.LG / 196 / 2609.24322

The Undetected Damage of Quantization on Retrieval and How to Fix It

量化对检索造成的未被察觉的损害及其修复方法
Zhou, Luca, Zirilli, Alessandro, Solombrino, Daniele, Dessì, Roberto, Rodolà, Emanuele
Abstract
We show that a quantized model that keeps its classification accuracy still changes $14$ to $46\%$ of its top-1 retrieval results, and that aggregate ranking metrics reveal only part of this damage. We tie this failure to the gap between the two highest scores and use that gap to decide when a quantized answer can be trusted and where additional precision should be spent. We show that the top-1 result is guaranteed to survive quantization only when this gap exceeds twice the largest rounding error. In classification, scores are the logits, and the loss function pushes the correct class away from other classes, encouraging this gap. In retrieval, scores are query-document scores, and nothing separates the top-1 item from the second. This gap can be measured without labels. Before deployment, it predicts which models will break under quantization, and at deployment time it tells, per input, whether the quantized answer still matches the full-precision answer. Most classification inputs have a gap wide enough to trust the quantized answer, but few retrieval queries do. That gap motivates a different fix in each task. In retrieval, spending extra bit-width on the layers whose quantization moves the gap most recovers up to three-quarters of an extra bit's benefit for half its cost. In classification, routing the few low-gap inputs to full precision recovers most of the lost accuracy at a fraction of the cost.
Chinese Translation
我们证明,量化后模型即使保持分类准确率不变,其 top-1 检索结果仍会发生 14% 到 46% 的变化,而聚合排序指标只能揭示这一损害的部分影响。我们将这一失效归因于两个最高分数之间的差距(gap),并利用该差距来判断何时可以信任量化后的答案,以及何时需要投入额外的精度。我们证明,只有当该差距超过最大舍入误差的两倍时,top-1 结果才能保证在量化后保持不变。在分类任务中,分数是 logits,损失函数促使正确类别与其他类别拉开距离,从而有利于形成这一差距;而在检索任务中,分数是查询-文档相关性分数,没有任何机制将排名第一的结果与第二名区分开来。该差距可以在无需标签的情况下进行度量。在部署前,它可以预测哪些模型在量化后会失效;在部署时,它可以针对每个输入判断量化答案是否仍与全精度答案一致。大多数分类输入的差距足以信任量化答案,但很少有检索查询能做到这一点。这一差距促使我们在两个任务中采用不同的修复方案。在检索任务中,将额外的位宽分配给那些量化对差距影响最大的层,能以一半的代价获得额外一比特收益的四分之三。在分类任务中,将少量低差距输入路由至全精度计算,能以很小的代价恢复大部分损失的准确率。
cs.LG / 197 / 2609.24328

A Distributional Optimisation Perspective on Combining Models in Deep Learning

深度学习中模型组合的分布式优化视角
Wang, Congye, Lin, Yan, Shen, Zheyang, Fisher, Matthew A., Oates, Chris. J.
Abstract
Combining predictions from different models can improve performance at machine learning tasks, but the training of the individual models and the rule used to combine them are typically chosen separately, and by ad hoc means. Recent advances in distributional optimisation (i.e. where the optimisation occurs over the set of probability distributions) offer an opportunity for principled joint training, viewing the collection of models as a discrete distribution whose support points are to be optimised, but the potential of these methods is not well-understood. In this paper we (1) cast two standard combination strategies - ensembles and low-rank adapter averaging - as entropy-regularised distributional optimisation, observing that the resulting objective is convex in the ensemble case but not in the adapter-averaging case, so that existing convergence guarantees for mean field Langevin dynamics transfer only to the former; (2) assess existing and novel algorithms for this task, including a functional variant of variational gradient descent; and (3) report an empirical study spanning synthetic classification tasks and fine-tuning of large language models on a commonsense reasoning benchmark.
Chinese Translation
组合不同模型的预测可以提升机器学习任务的性能,但各个模型的训练以及用于组合它们的规则通常是分开选择的,且方式较为随意。分布式优化(即在概率分布集合上进行优化)的最新进展为原则性的联合训练提供了机会,即将模型的集合视为一个离散分布,其支撑点有待优化,但这类方法的潜力尚未得到充分理解。在本文中,我们:(1) 将两种标准的组合策略——集成(ensembles)和低秩适配器平均(low-rank adapter averaging)——表述为熵正则化的分布式优化,并观察到所得目标函数在集成情形下是凸的,而在适配器平均情形下不是凸的,因此均值场朗之万动力学(mean field Langevin dynamics)现有的收敛性保证仅适用于前者;(2) 评估了用于该任务的现有算法及新算法,其中包括变分梯度下降的一个泛函变体;(3) 报告了一项涵盖合成分类任务以及在常识推理基准上微调大语言模型的实证研究。
cs.LG / 198 / 2609.24338

Pharmacokinetic State Space Models for Unbiased Prediction of Haemodynamic Collapse

用于血液动力学衰竭无偏预测的药代动力学状态空间模型
Nagaraj, Rithin, Chindula, Sudiksha, Das, Bhaskarjyoti
Abstract
An Intraoperative Hypotension (IOH) event is a frequent complication during administration of general anaesthesia with serious downstream consequences, yet clinical management remains reactive and not predictive. Existing predictive models, however, ignore drug infusion history as a valuable signal for prediction despite its direct pharmacological relevance. Our model achieves an Area Under the Receiver Operating Characteristic curve (AUROC) of 0.7360 and an Area Under the Precision-Recall Curve (AUPRC) of 0.1794, representing a 2.73-fold lift over the random guessing AUPRC baseline (0.0657), with the removal of propofol and remifentanil effect-site concentrations resulting in a 13.9% AUPRC drop compared to the full model. This is consistent with the hypothesis that pharmacokinetic trajectories encode impending haemodynamic changes before they manifest in the Mean Arterial Pressure (MAP). Additionally, this paper shows that training without lead-gap filtering degraded AUROC by 16.7%, empirically confirming that unfiltered models learn to detect ongoing hypotension rather than predict future events. Finally, a Mamba-based architecture achieves the aforementioned high prediction performance while maintaining a constant memory footprint across a range of sequence lengths, unlike the quadratic VRAM overhead typical of vanilla Transformers, making it the more practical choice for continuous intraoperative deployment.
Chinese Translation
术中低血压(Intraoperative Hypotension, IOH)事件是全身麻醉给药过程中常见的并发症,可造成严重的后续后果,然而目前的临床管理仍以反应性为主,缺乏预测性。现有预测模型忽视了药物输注史这一具有直接药理学相关性的有价值预测信号。我们的模型实现了受试者工作特征曲线下面积(AUROC)0.7360 和精确率-召回率曲线下面积(AUPRC)0.1794,相较于随机猜测的 AUPRC 基线(0.0657)提升了 2.73 倍;而移除丙泊酚(propofol)和瑞芬太尼(remifentanil)的效应室浓度后,AUPRC 相比完整模型下降了 13.9%。这与如下假设一致:药代动力学轨迹在平均动脉压(MAP)发生变化之前就已编码了即将发生的血流动力学改变。此外,本文还表明,不进行超前间隙(lead-gap)过滤会使 AUROC 下降 16.7%,从实证上证实了未经过滤的模型学到的只是检测正在发生的低血压,而非预测未来事件。最后,基于 Mamba 的架构在实现上述高预测性能的同时,能在多种序列长度下保持恒定的内存占用,不像原生 Transformer 那样存在二次方 VRAM 开销,因而使其成为连续术中部署的更实用选择。
cs.LG / 199 / 2609.24358

Explainable Neuro-Fuzzy Prediction for Trustworthy Decision-Making in Maritime

面向海事领域可信决策的可解释神经模糊预测
Kalogeropoulos, Dionisis, Sovatzidi, Georgia, Iakovidis, Dimitris K.
Abstract
Predicting when maritime systems require maintenance can be critical, avoiding hazards and costly consequences. To address this problem, this paper proposes an explainable decision-making framework that integrates a neuro-fuzzy prediction model with a two-stage explainable component. The first stage of this component produces feature-attribution explanations, using gradient-based saliency maps, and the second stage extracts local rules using a fuzzy decision tree. The proposed framework is generic and can be integrated into any deep learning-based approach, rendering it explainable. To the best of our knowledge, this is the first fuzzy logic-based framework enabling both feature-level and local rule-based explanations of black box models. This approach aims to foster trustworthiness in decision making through user-understandable machine inferences. The performance of the proposed framework using a deep residual-based neural backbone is evaluated on various general-purpose public benchmark datasets, and its utility in maritime is demonstrated in the context of early fault detection in a naval propulsion system dataset. The results indicate that it can provide predictions outperforming relevant state-of-the-art approaches, with an average AUC-ROC (Area Under the Receiver Operating Characteristic Curve) value, reaching up to 99%, while offering the advantage of explainability.
Chinese Translation
预测海事系统何时需要维护至关重要,可以避免危险和高昂的后果。为解决这一问题,本文提出了一种可解释的决策框架,该框架将神经模糊预测模型与一个两阶段可解释组件相集成。该组件的第一阶段利用基于梯度的显著性图生成特征归因解释,第二阶段使用模糊决策树提取局部规则。所提出的框架具有通用性,可集成到任何基于深度学习的方法中并使其具备可解释性。据我们所知,这是首个能够同时提供特征级解释和基于局部规则解释黑盒模型的模糊逻辑框架。该方法旨在通过用户可理解的机器推理来增强决策的可信度。我们使用深度残差神经网络骨干,在多个通用公开基准数据集上评估了所提框架的性能,并在舰船推进系统数据集的早期故障检测场景中验证了其在海事领域的实用性。结果表明,该框架能够提供优于相关最先进方法的预测,平均AUC-ROC(受试者工作特征曲线下面积)值高达99%,同时具备可解释性的优势。
cs.LG / 200 / 2609.24370

Prescriptive SVD-Inspired Attention via Spectral Energy Retention

基于谱能量保持的规范性SVD启发注意力机制
Arampatzakis, Vasileios, Sevetlidis, Vasileios, Pavlidis, George
Abstract
Self-attention is central to modern Transformer architectures, but its dense dot-product formulation makes it difficult to identify which internal directions are structurally important and which can be modified without disrupting the model. SVD-Inspired Attention (SVDA) addresses part of this problem by introducing a learned diagonal spectrum into the query-key score interaction, making latent attention directions explicitly inspectable through indicators such as spectral entropy, effective rank, sparsity, alignment, selectivity, and perturbation response. This paper examines the transition from diagnostic interpretation to operational intervention. A diagnosis--intervention--verification framework is proposed, and one intervention is evaluated: spectral energy retention in the attention-score pathway. Across FashionMNIST, CIFAR-10, CIFAR-100, and Food-101, the $\rho=0.90$ prescription removes 24.5--53.7\% of score directions, reduces parameters by 2.6--4.3\%, and reduces estimated MACs by 2.8--5.4\%. The paired mean accuracy change of the dimension-reduced model ranges from $-0.03$ to $+0.05$ percentage points over three seeds. These results support SVDA as an intrinsically interpretable attention mechanism whose learned spectrum exposes an operational coordinate system for deterministic and verifiable modification of attention-score formation.
Chinese Translation
自注意力机制是现代Transformer架构的核心,但其稠密的点积形式使得难以识别哪些内部方向在结构上是重要的、哪些可以在不破坏模型的情况下进行修改。SVD启发注意力(SVD-Inspired Attention, SVDA)通过在查询-键得分交互中引入可学习的对角谱,部分解决了这一问题,使潜在的注意力方向可以通过谱熵、有效秩、稀疏性、对齐性、选择性和扰动响应等指标进行显式检查。本文研究了从诊断性解释到操作性干预的过渡。我们提出了一个“诊断—干预—验证”框架,并评估了一种干预方法:注意力得分通路中的谱能量保持。在FashionMNIST、CIFAR-10、CIFAR-100和Food-101数据集上,ρ=0.90的规范性设定可移除24.5–53.7%的得分方向,减少2.6–4.3%的参数量,并降低2.8–5.4%的估计乘加运算(MACs)量。在三个随机种子下,降维模型的平均准确率变化范围为−0.03至+0.05个百分点。这些结果表明,SVDA是一种具有内在可解释性的注意力机制,其可学习谱揭示了一个可操作的坐标系,能够对注意力得分形成进行确定性的、可验证的修改。
cs.LG / 201 / 2609.24380

Information-Time Proximal Policy Optimization

信息时间近端策略优化
Zeng, Yongcheng, Cui, Xinyu, Song, Yan, Liu, Guoqing, Xin, Hongsheng, Zhang, Kaike, Deng, Cheng, Zhan, Kun, Ying, Jian, Zhao, Jian, Zhang, Haifeng, Wang, Jun
Abstract
RLVR has substantially improved the reasoning capabilities of LLMs. However, existing methods typically parameterize temporal progression in the Markov Decision Process by token-by-token generation, despite the highly non-uniform information flow along autoregressive trajectories. In this paper, we propose InfoPPO, which reparameterizes temporal progression using information density rather than raw token count. This reparameterization induces a common state-dependent structure for both temporal credit propagation and policy updates. InfoPPO restores the effectiveness of non-trivial discounting in long-horizon reasoning, retaining effective-horizon contraction while avoiding excessive attenuation of terminal supervision over long token sequences. Moreover, the information-time policy-improvement analysis naturally leads to a state-dependent update constraint, which we implement through adaptive clipping. By adapting the clipping threshold at each token position to the information density of its corresponding state, this mechanism enables more targeted policy updates while preserving proximal control. Theoretically, we extend performance-difference and policy-improvement analyses to the information-time MDP, deriving a policy-improvement lower bound when policy changes are regulated by information density. We further connect the general information-time analysis to practical LLM policy optimization by relating state-wise information density to local policy movement, while also providing theoretical grounding for the adaptive update mechanism. Experiments on Qwen3 models demonstrate consistent gains over competitive baselines across five challenging competition-style mathematical reasoning benchmarks. InfoPPO also maintains stable accuracy and response length across non-trivial discount settings under which token-time PPO deteriorates.
Chinese Translation
RLVR(可验证奖励强化学习)已显著提升了大语言模型(LLM)的推理能力。然而,现有方法通常以逐个token生成的方式对马尔可夫决策过程(MDP)中的时间进程进行参数化,尽管自回归轨迹上的信息流高度不均匀。本文提出InfoPPO,使用信息密度而非原始token数量对时间进程进行重新参数化。这一重新参数化为时间信用传播与策略更新引入了共同的状态依赖结构。InfoPPO恢复了非平凡折扣因子在长时程推理中的有效性,在保持有效时程收缩的同时,避免了长token序列上终端监督信号被过度衰减。此外,基于信息时间的策略改进分析自然导出了一种状态依赖的更新约束,我们通过自适应裁剪来实现该约束。通过将每个token位置上的裁剪阈值与其对应状态的信息密度相适配,该机制在保持近端控制的同时实现了更具针对性的策略更新。在理论上,我们将性能差异和策略改进分析扩展至信息时间MDP,推导出当策略变化受信息密度调控时的策略改进下界。我们进一步将状态级信息密度与局部策略变化相关联,将一般性的信息时间分析与实际的LLM策略优化联系起来,同时为自适应更新机制提供了理论基础。在Qwen3模型上的实验表明,InfoPPO在五个具有挑战性的竞赛风格数学推理基准上持续优于有竞争力的基线方法。并且在token时间PPO性能恶化的非平凡折扣设置下,InfoPPO仍能保持稳定的准确率和响应长度。
cs.LG / 202 / 2609.24382

Credit Access is Associated with Improved Food Security in the Horn of Africa

信贷获取与非洲之角地区粮食安全状况的改善相关
Cerdà-Bautista, Jordi, Sitokonstantinou, Vasileios, López-Peña, José Manuel Veiga, Piovani, Duccio, Tárraga, José María, Camps-Valls, Gustau
Abstract
The intensification of climate change poses a growing threat to food security, especially in vulnerable communities. This study employs an observational machine-learning framework to estimate the causal association between access to credit and acute food insecurity in Somalia and across the Horn of Africa, drawing on a harmonized dataset spanning key environmental, socioeconomic, and conflict-related factors from 2015 to 2022. Results indicate that greater credit access is associated with a 2% reduction in acute food insecurity at the population level over the study period. Given that, on average, 16% of the population is in crisis, this effect represents a meaningful shift within the at-risk group. We interpret these estimates under explicit identification assumptions and complement them with robustness and refutation tests. The results provide context-specific evidence on how financial access correlates with food security outcomes in data-scarce, crisis-affected settings, and offer a transparent framework for integrating heterogeneous data sources when randomized evaluations are infeasible.
Chinese Translation
气候变化的加剧对粮食安全构成了日益严重的威胁,尤其是在脆弱社区。本研究采用观察性机器学习框架,基于一份涵盖环境、社会经济和冲突相关关键因素、时间跨度为2015年至2022年的统一数据集,估计了在索马里及整个非洲之角地区获得信贷与急性粮食不安全之间的因果关联。结果表明,在研究期间,更高的信贷获取水平与人口层面急性粮食不安全降低2%相关。考虑到平均有16%的人口处于危机状态,这一效应对于高危群体而言意味着具有实际意义的改善。我们在明确的识别假设下解释这些估计结果,并通过稳健性检验和反驳性检验加以补充验证。研究结果为在数据稀缺、受危机影响的地区,金融获取如何与粮食安全结果相关联提供了情境化证据,并在无法开展随机化评估的情况下,提供了一个整合异构数据源的透明框架。
cs.LG / 203 / 2609.24386

Machine Learning-Based Prediction of Childhood Stunting in Bangladesh: Fairness and Temporal Robustness Assessment

基于机器学习的孟加拉国儿童发育迟缓预测:公平性与时间稳健性评估
Haque, Md Ahshanul, Kabir, Muhammad Ashad
Abstract
Childhood stunting remains a major public health concern in Bangladesh and reflects long-term growth failure influenced by child, maternal, household, socioeconomic, and health-service factors. This study used nationally representative Bangladesh Demographic and Health Survey data from 2007 to 2022 to develop machine learning models for population-level prediction of childhood stunting and to assess temporal robustness and subgroup fairness. Children aged 0-59 months with complete anthropometric and predictor data were included. Data from the 2007, 2011, and 2014 survey rounds were used for model development, while the 2018 and 2022 rounds were retained as temporal test datasets. Twelve feature-selection approaches were assessed, and the KNN permutation importance-selected predictor set was used for final model evaluation. Eleven machine learning models were evaluated: ten conventional algorithms and one pretrained tabular foundation model, TabPFN. Performance was assessed using balanced accuracy, AUROC, F1-score, Brier score, and expected calibration error. Subgroup fairness was examined by child sex, place of residence, and socioeconomic status. The final analytic sample included 18,844 children, of whom 35.05% were stunted. In the development hold-out test dataset, TabPFN showed the highest observed balanced accuracy overall at 67.58%, while AdaBoost showed the highest observed balanced accuracy among conventional models at 67.51%. In temporal testing, the highest observed balanced accuracy was found for Gradient Boosting in BDHS 2018 and XGBoost in BDHS 2022. Model performance varied across survey rounds and subgroups, highlighting the importance of temporal validation, subgroup fairness assessment, and transparent interpretation in public health prediction modeling.
Chinese Translation
儿童发育迟缓仍是孟加拉国的一项重大公共卫生问题,反映了由儿童、母亲、家庭、社会经济及卫生服务等因素共同影响下的长期生长不良。本研究使用2007年至2022年具有全国代表性的孟加拉国人口与健康调查(BDHS)数据,构建了用于人群层面儿童发育迟缓预测的机器学习模型,并评估了时间稳健性和亚组公平性。研究对象为拥有完整人体测量数据和预测变量数据的0-59月龄儿童。2007年、2011年和2014年调查轮次的数据用于模型开发,2018年和2022年轮次的数据作为时间测试数据集予以保留。研究评估了十二种特征选择方法,并采用KNN置换重要性筛选出的预测变量集进行最终模型评估。共评估了十一种机器学习模型:十种传统算法和一种预训练的表格基础模型TabPFN。模型性能采用平衡准确率、AUROC、F1分数、Brier评分和预期校准误差进行评估。亚组公平性按儿童性别、居住地和社会经济地位进行考察。最终分析样本包含18,844名儿童,其中35.05%存在发育迟缓。在开发阶段的留出测试数据集中,TabPFN总体上表现出最高的平衡准确率,为67.58%;传统模型中AdaBoost的平衡准确率最高,为67.51%。在时间测试中,BDHS 2018中Gradient Boosting的平衡准确率最高,BDHS 2022中XGBoost的平衡准确率最高。模型性能在不同调查轮次和亚组间存在差异,凸显了在公共卫生预测建模中进行时间验证、亚组公平性评估和透明解读的重要性。
cs.LG / 204 / 2609.24391

NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware

NAVIR:面向边缘硬件上鲁棒人机交互的神经形态音视频语音识别
Delimpasis, Leonidas, Moraiti, Panagiota, Porichis, Antonis, Chatzakos, Panos, Karamousadakis, Michail
Abstract
Voice-controlled interaction in industrial settings is hampered by acoustic noise, which severely degrades audio-only speech recognition. Audio-visual speech recognition (AVSR) addresses this by fusing lip-motion cues with the audio stream, but state-of-the-art pipelines rely on three-dimensional convolutions, recurrent units, and attention modules that exceed the budget of typical edge devices. We present NAVIR, an end-to-end AVSR system targeting the BrainChip Akida neuromorphic processor, which natively supports only sequential two-dimensional convolutional inference. The pipeline factorises spatial and temporal encoding into separate AkidaNet-based modules: a per-frame visual encoder, a temporal video encoder, and a spectrogram audio encoder, fused by a lightweight predictor head and decoded by constrained beam search. Models are trained with connectionist temporal classification on noise-augmented audio and then fine-tuned with quantization-aware training. On the GRID benchmark, the quantized audio-visual model reaches 14.0% word error rate (WER) under noise on the unseen-speaker split and 3.3% WER on the overlapped-speaker split, against 22.5% and 11.8% for audio-only baselines, and it attains 98.6% command accuracy at 1.5% WER on a task-specific industrial-command corpus. Operation-count analysis indicates a 13-fold energy advantage of the spiking formulation over its artificial neural network counterpart at 27.6% mean firing rate. On-board measurements show roughly 5-fold lower energy per inference than a Raspberry Pi central processing unit on the lip-reading model, and over 100-fold lower than a laptop graphics processing unit, while sustaining 14.5 inferences per second. To the best of our knowledge, this is the first complete multimodal AVSR pipeline running on neuromorphic hardware of this class.
Chinese Translation
工业环境中的语音控制交互受到声学噪声的干扰,噪声会严重降低纯音频语音识别的性能。音视频语音识别(AVSR)通过将唇部运动线索与音频流融合来应对这一问题,但最先进的系统流水线依赖于三维卷积、循环单元和注意力模块,超出了典型边缘设备的资源预算。我们提出了NAVIR,一个面向BrainChip Akida神经形态处理器的端到端AVSR系统,该处理器原生仅支持顺序的二维卷积推理。该流水线将空间和时间编码分解为独立的基于AkidaNet的模块:逐帧视觉编码器、时序视频编码器和频谱图音频编码器,通过轻量级预测头进行融合,并由受限束搜索进行解码。模型使用连接主义时序分类(CTC)在噪声增强的音频上进行训练,随后通过量化感知训练进行微调。在GRID基准上,量化后的音视频模型在未见说话人划分的噪声条件下达到14.0%的词错误率(WER),在重叠说话人划分上达到3.3%的WER,而纯音频基线分别为22.5%和11.8%;在任务特定的工业指令语料库上,该模型在1.5%的WER下达到了98.6%的指令准确率。运算量分析表明,在27.6%的平均发放率下,脉冲化(spiking)形式的能耗相比其人工神经网络对应方案具有13倍的优势。板载实测显示,在唇读模型上,其每次推理的能耗比树莓派(Raspberry Pi)中央处理器低约5倍,比笔记本电脑图形处理器低100倍以上,同时保持每秒14.5次推理的速度。据我们所知,这是首个在此类神经形态硬件上运行的完整多模态AVSR流水线。
cs.LG / 205 / 2609.24394

Climate Variability Modulates the Impact of Price Spikes on Food Insecurity

气候变率调节价格飙升对粮食不安全的影响
Cerdà-Bautista, Jordi, Sitokonstantinou, Vasileios, Durand, Homer, Varando, Gherardo, Ronco, Michele, Camps-Valls, Gustau
Abstract
Climate variability influences whether a market disruption escalates into a food crisis, yet broad climate patterns like El Ni\~no, tracked months before they alter hydro-climatic conditions, are still not incorporated as an early-warning component in food-security responses. We address this gap by introducing sensitivity regimes, a stratification of regions by the direction and strength of their vegetation response to the El Ni\~no Southern Oscillation, and using them to estimate how food price spikes affect acute food insecurity across sub-Saharan Africa. Integrating remote sensing, socioeconomic data, and causal machine learning, we find that in regions where ENSO systematically suppresses vegetation, a price spike raises the share of the population at acute risk by 5.4 percentage points in the following month. In regions where vegetation is unaffected by or positively linked to ENSO, the estimated effect is smaller (around 2 percentage points) and statistically insignificant. These results demonstrate that climate context is critical for understanding food security vulnerabilities. Sensitivity regimes can be combined with operational price-spike triggers to stage anticipatory action: the ENSO state flags vulnerable regions months ahead, and a pre-positioned response in those regions to a price spike would avert the largest jump in acute food insecurity.
Chinese Translation
气候变率决定了一次市场紊乱是否会演变为粮食危机,然而像厄尔尼诺(El Niño)这类大尺度气候模式——在水文气候条件改变前数月即可被监测到——至今仍未被纳入粮食安全应对的预警体系。为了填补这一空白,我们提出了敏感性区划(sensitivity regimes),即依据各地区植被对厄尔尼诺-南方涛动(ENSO)响应的方向和强度进行分层,并据此估算粮食价格飙升如何影响撒哈拉以南非洲的急性粮食不安全。通过整合遥感数据、社会经济数据与因果机器学习方法,我们发现,在ENSO系统性抑制植被的区域,价格飙升会使下个月处于急性风险中的人口比例上升5.4个百分点;而在植被不受ENSO影响或与ENSO呈正相关的区域,估计效应较小(约2个百分点)且在统计上不显著。这些结果表明,气候背景对于理解粮食安全脆弱性至关重要。敏感性区划可与业务化的价格飙升触发机制相结合,以部署前瞻性行动:ENSO状态可提前数月标记脆弱区域,而在这些区域针对价格飙升预先部署的应对措施,将能够避免急性粮食不安全出现最大幅度的跃升。
cs.LG / 206 / 2609.24397

Probabilistic Modelling of Operational Design Domains, A New Approach for Testing AI Systems

运行设计域的概率建模:一种测试人工智能系统的新方法
Wiesbrock, Hans-Werner
Abstract
The conventional testing process quickly fails when applied to ML-based systems such as obstacle detection in vehicles: if an obstacle is not detected in a test, classical bug fixing is impossible and an AI system will always retain shortcomings. Test results can therefore only be interpreted statistically, which in turn requires test sets that are not only complete with respect to the operational design domain (ODD) of the system, but also representative of it. To this end, we introduce probabilistically extended ontologies (PEONs): ontologies describing the ODD, augmented with a probability distribution over the partitioning they induce. Instead of unmaintainable conditional probability tables, only marginal distributions and functionally described dependencies need to be specified; algorithms based on couplings and optimal transport complete this specification to a Bayesian network. From a PEON we derive the sampling of representative test cases, rigorous end-of-test criteria for given quality targets and significance levels, and methods for re-evaluating existing test results and for assessing the balance of training data. We demonstrate the practical modelling of a complex ODD using the example of automatic train operation.
Chinese Translation
传统的测试流程在应用于基于机器学习(ML)的系统(如车辆中的障碍物检测)时会很快失效:如果在测试中未能检测到障碍物,经典的缺陷修复方法便无从实施,且人工智能系统总会保留某些缺陷。因此,测试结果只能从统计角度进行解释,这就要求测试集不仅要覆盖系统运行设计域(ODD)的全部范围,还要具有代表性。为此,我们提出了概率扩展本体(Probabilistically Extended Ontologies, PEONs):即描述ODD的本体,并为其诱导的划分增补一个概率分布。无需维护难以管理的条件概率表,只需指定边缘分布以及以函数形式描述的依赖关系;基于耦合(couplings)和最优传输(optimal transport)的算法可将该规范补全为贝叶斯网络。从PEON出发,我们可以推导出代表性测试用例的采样方法、针对给定质量目标和显著性水平的严格的测试终止准则,以及重新评估已有测试结果和评估训练数据均衡性的方法。我们以自动驾驶列车运行为例,展示了复杂ODD的实际建模过程。
cs.LG / 207 / 2609.24401

Artificial Structure Function Search: Preserving Artificial Functional Connectivity for Structured Pruning

人工结构-功能搜索:面向结构化剪枝的人工功能连接保持
Illeperuma, Mindula, Pina, Rafael, Herath, Charuka, Gabayre, Sharmarke A., De Silva, Varuna
Abstract
Structured pruning is a model compression technique that is used to reduce the computational cost of deploying deep neural networks on resource-constrained devices. Popular methods of pruning rely on opaque heuristics or weight-based criteria that give no indication as to the structural dependencies in the network. To address these limitations we present Artificial Structure Function Search (ASF-S): a novel structured pruning framework. ASF-S utilizes Principle Gradient Importance (PGI): a novel prune-candidate selection criteria that is inspired by structure-function relationships in the brain. By ensuring the pruned structure of the model respects topographical organization of the output layer, we define Artificial Functional Connectivity (AFC) for artificial neural networks. AFC provides evidence to demonstrate that accurate smaller networks can be found using careful prune candidate selection criteria. We present results for PGI as a selection criterion and for ASF-S as a pruning framework against recent benchmarks, demonstrating that our method yields model variants with 70\% parameter reduction, that can recover baseline accuracy without re-training the pruned layers.
Chinese Translation
结构化剪枝是一种模型压缩技术,用于降低在资源受限设备上部署深度神经网络的计算成本。流行的剪枝方法依赖于不透明的启发式规则或基于权重的准则,无法反映网络中的结构依赖关系。为解决这些局限性,我们提出了人工结构-功能搜索(Artificial Structure Function Search, ASF-S):一种新颖的结构化剪枝框架。ASF-S 利用主梯度重要性(Principle Gradient Importance, PGI)这一新颖的剪枝候选选择准则,其灵感来源于大脑中的结构-功能关系。通过确保剪枝后模型的结构尊重输出层的拓扑组织,我们为人工神经网络定义了人工功能连接(Artificial Functional Connectivity, AFC)。AFC 提供证据表明,借助精心设计的剪枝候选选择准则,可以找到规模更小且精度保持的网络。我们将 PGI 作为选择准则、将 ASF-S 作为剪枝框架,与近期基准方法进行了对比实验。结果表明,我们的方法能够得到参数减少 70% 的模型变体,且无需对剪枝后的层进行重新训练即可恢复基线精度。
cs.LG / 208 / 2609.24417

ARM: Attention with Routed-Memory for Learnable Sparse Control

ARM:面向可学习稀疏控制的带路由记忆注意力机制
Zeng, Qiuhao, Huang, Jerry, Lu, Peng, Fang, Ruiyi, Xu, Gezheng, Jing, Zihao, Cui, Yufei, Ling, Charles, Niu, Gang, Wang, Boyu
Abstract
Despite advances in long-context inference, large language models (LLMs) remain fundamentally limited by the key-value (KV) caching mechanisms that are necessary for stable computation. Techniques such as selective token eviction and pruning have vastly mitigated these issues, but often discard core information to manage the growing cache. In this paper, we propose Attention with Routed Memory (ARM) a novel KV caching structure that introduces a fully differentiable, fixed-size memory system organized as a hierarchical router. Via a Gumbel-Softmax, ARM learns to select memory slots and perform sigmoid-gated updates that softly combine new and stored information, avoiding hard eviction and reducing information loss. By further training a policy to dynamically select varying amounts of memory at inference, ARM adapts its accesses for both simple contexts and inputs that require deeper reasoning, enabling more scalable and effective retrieval on both short- and long-contexts. Experimental results on standard commonsense and long-context reasoning benchmarks demonstrate that ARM achieves superior performance and efficiency compared to fixed KV-caching approaches, while remaining efficient and scalable in terms of both memory and generation latency.
Chinese Translation
尽管长上下文推理取得了进展,大语言模型(LLM)仍在根本上受限于稳定计算所必需的键值(KV)缓存机制。选择性令牌驱逐和剪枝等技术极大地缓解了这些问题,但为管理不断增长的缓存,往往会丢弃核心信息。本文提出带路由记忆的注意力机制(ARM),一种新颖的KV缓存结构,它引入了一个完全可微、固定大小的记忆系统,并将其组织为分层路由器。ARM通过Gumbel-Softmax学习选择记忆槽位,并执行Sigmoid门控更新,将新信息与存储信息进行软性结合,从而避免硬性驱逐并减少信息丢失。通过进一步训练策略以在推理时动态选择不同数量的记忆,ARM能够适应简单上下文以及需要更深层次推理的输入,从而在短上下文和长上下文中都实现更具可扩展性和更有效的检索。在标准常识推理和长上下文推理基准上的实验结果表明,与固定KV缓存方法相比,ARM实现了更优的性能和效率,同时在内存和生成延迟方面保持高效和可扩展。
cs.LG / 209 / 2609.24422

Prior-Amortized In-Context Bayesian Inference for Generalized Linear Mixed-Effects Models

面向广义线性混合效应模型的先验摊销式上下文贝叶斯推断
Kipnis, Alex, Binz, Marcel, Schulz, Eric
Abstract
Hierarchical data is ubiquitous in the empirical sciences and is most commonly analyzed with generalized linear mixed-effects models (GLMMs). Bayesian inference for GLMMs yields calibrated uncertainty but requires MCMC; the No-U-Turn Sampler (NUTS) is the gold standard but is slow and must restart from scratch for every new dataset, model and prior. We introduce metabeta, a pretrained neural network for prior-amortized in-context Bayesian inference over GLMMs. Unlike previous neural posterior estimators that fix the prior at training time, metabeta accepts prior families and hyperparameters as inputs at test time, enabling zero-shot generalization. Two set transformers and conditional normalizing flows mirror the posterior's two-level structure (global parameters shared across groups, local parameters per group). The model is trained on millions of realistic simulated datasets spanning continuous, binary, and count outcomes. By default, the flow posterior is refined by Independence Metropolis-Hastings against the unnormalized posterior, so its correctness rests on the sampler rather than the network; this yields tuning-free inference two to three orders of magnitude faster than NUTS. Alternatively, the flow can warm-start NUTS, giving nearly identical inference with substantially increased speed and stability. On controlled benchmarks with ground-truth parameters, metabeta matches NUTS in parameter recovery, calibration and out-of-sample prediction. On out-of-distribution real datasets, its posteriors closely match those of NUTS across all parameter types, and they remain faithful under misspecified likelihoods and priors, out-of-distribution predictors, collinear designs, and data-poor regimes. The model is open-source and open-weights and thus immediately deployable.
Chinese Translation
层次数据在实证科学中无处不在,通常采用广义线性混合效应模型(GLMM)进行分析。对GLMM进行贝叶斯推断可以获得校准良好的不确定性估计,但需要使用MCMC方法;No-U-Turn采样器(NUTS)是该领域的黄金标准,但速度缓慢,且面对每个新数据集、模型和先验都必须从头开始。我们提出了metabeta,一个用于GLMM先验摊销式上下文贝叶斯推断的预训练神经网络。与以往在训练时固定先验的神经后验估计器不同,metabeta在测试时接受先验族及其超参数作为输入,从而实现零样本泛化。模型采用两个集合Transformer(set transformers)和条件归一化流,以对应后验的两级结构(跨组共享的全局参数与每组各自的局部参数)。该模型在数百万个涵盖连续、二值和计数结局的真实感模拟数据集上训练。默认情况下,流后验通过针对非归一化后验的独立Metropolis-Hastings采样进行细化,因此其正确性依赖于采样器而非网络本身;由此得到的无需调参的推断速度比NUTS快两到三个数量级。此外,该流模型也可作为NUTS的暖启动,在大幅提升速度和稳定性的同时获得几乎完全一致的推断结果。在具有真实参数的受控基准测试中,metabeta在参数恢复、校准和样本外预测方面与NUTS相当。在分布外的真实数据集上,其后验在所有参数类型上均与NUTS高度吻合,并且在似然与先验设定错误、分布外预测变量、共线性设计以及数据稀缺情形下仍保持可靠。该模型开源且开放权重,可立即部署使用。
cs.LG / 210 / 2609.24432

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation

1%的词元可能就足够:论在线策略蒸馏中的梯度估计
Sheng, Huanxin, Ye, Zhiling, Wang, Haonan, Wang, Jian, Gu, Jinjie, Kang, Jian
Abstract
Sparse on-policy distillation (OPD) allocates teacher supervision to a small subset of tokens in student-generated trajectories. However, useful teacher guidance can yield a noisy update when its gradient is estimated from a sampled next token. We study this estimation problem at a fixed prefix in information geometry and propose an information-efficiency ratio (IER) based on a signal-to-noise decomposition. IER characterizes relative gradient estimation error under an optimal scalar baseline. A candidate-set approximation enables token selection based on IER and its combination with existing usefulness scores, while retaining the sampled reverse-KL training objective. On mathematical and medical reasoning tasks, adding IER improves existing selectors in multiple settings, with sparse configurations matching or exceeding full OPD without token selection at small token budgets of 0.1\%--1\%. These results support accounting for both usefulness and gradient-estimation reliability when allocating sparse supervision. Our code is available at https://github.com/BruceSheng1202/IER-OPD.
Chinese Translation
稀疏在线策略蒸馏(On-Policy Distillation, OPD)将教师模型的监督分配到学生生成轨迹中的一小部分词元上。然而,当教师引导的梯度是基于采样的下一个词元进行估计时,有用的教师信号可能产生噪声较大的更新。我们在信息几何框架下研究了固定前缀处的这一估计问题,并基于信噪比分解提出了信息效率比(Information-Efficiency Ratio, IER)。IER刻画了在最优标量基线下相对梯度估计误差的大小。通过候选集近似,可以基于IER进行词元选择,并将其与现有的有用性评分相结合,同时保留采样的反向KL训练目标。在数学推理和医学推理任务上,引入IER在多种设置下均能改进现有的选择器;在0.1%–1%的极小词元预算下,稀疏配置的性能可匹敌甚至超过不进行词元选择的完整OPD。这些结果表明,在分配稀疏监督时,应同时考虑有用性和梯度估计的可靠性。我们的代码见 https://github.com/BruceSheng1202/IER-OPD。
cs.LG / 211 / 2609.24440

Comparing Latent Concept Formation in State Space Models and Transformers via Sparse Autoencoders

基于稀疏自编码器比较状态空间模型与Transformer中的潜在概念形成
Nagaraj, Rithin, Oruganti, Rupa Laalasa, Kunder, Prerna Subhashchandra, Joshi, Ashwini M
Abstract
The quadratic scaling of Transformer self-attention has driven the adoption of sub-quadratic Selective State Space Models (SSMs) like Mamba, which compress past context into a fixed-size recurrent hidden state. This strict informational bottleneck raises a foundational question for mechanistic interpretability: do SSMs and Transformers learn fundamentally distinct latent representations? In this work, we employ Sparse Autoencoders (SAEs) to conduct a large-scale, feature-level correspondence analysis between Mamba-130m and Pythia-70m over a 10-million token corpus. Contrary to hypotheses predicting widespread architectural divergence, we find no evidence of systematic representational divergence between architectures: across the observed Jaccard distribution, 99.98% of Mamba features cluster toward the upper alignment boundary, providing preliminary feature-level support for the Universality Hypothesis. We further identify and qualitatively characterize this microscopic fraction (0.02%) of diverging features, finding patterns consistent with the hypothesis that the recurrent bottleneck selectively limits the parsing of rigid syntax rather than broad semantic ontology. We demonstrate that while Pythia's unconstrained attention permits the monosemantic decomposition of distinct formatting edge-cases, Mamba is forced to compress unrelated syntactical anomalies into polysemantic "junk drawer" neurons to preserve state capacity. Collectively, these results suggest that architectural routing mechanisms may have negligible impact on core semantic understanding, with representational divergence confined to extreme structural margins.
Chinese Translation
Transformer自注意力的二次方复杂度扩展问题推动了亚二次复杂度选择性状态空间模型(SSM,如Mamba)的采用,此类模型将过去上下文压缩至固定大小的循环隐藏状态中。这一严格的信息瓶颈为机制可解释性提出了一个根本性问题:SSM与Transformer是否学习到本质上不同的潜在表示?在本工作中,我们使用稀疏自编码器(SAE),在1000万token的语料库上对Mamba-130m与Pythia-70m进行了大规模的特征级对应分析。与预测两种架构存在广泛分歧的假设相反,我们没有发现架构间存在系统性表示分歧的证据:在观测到的Jaccard分布中,99.98%的Mamba特征聚集于高对齐边界附近,为普遍性假说(Universality Hypothesis)提供了初步的特征级支持。我们进一步识别并定性刻画了这一微小部分(0.02%)产生分歧的特征,发现的模式与以下假说一致:循环瓶颈选择性地限制了对刚性句法的解析,而非广义语义本体。我们证明,虽然Pythia不受限制的注意力机制允许对不同的格式边缘情况进行单语义分解,但Mamba被迫将无关的句法异常压缩进多语义的“杂物抽屉”式神经元中,以保持状态容量。总体而言,这些结果表明架构路由机制对核心语义理解的影响可能微乎其微,表示分歧仅局限于极端的结构边缘情形。
cs.LG / 212 / 2609.24441

MUSE: Dependency-Aware Adaptation of a Frozen Vision Backbone for Multivariate Time Series Forecasting

MUSE:面向多变量时间序列预测的依赖感知冻结视觉骨干网络适配方法
Cai, Xinying, Lu, Junkai, Zhu, Yuhan, Yu, Xiaoyun, Qiu, Xiangfei, Hu, Jilin
Abstract
Multivariate time-series forecasting is essential to many real-world applications. Recent large vision models (LVMs) offer a promising paradigm by transferring cross-domain visual priors to time-series forecasting. However, existing LVM-based methods face two key challenges: balancing independent visual representation spaces with cross-variable dependency modeling, and adapting vision backbones pretrained on natural images to the distinct temporal semantics of time-series images. To address these challenges, we propose MUSE, a dependency-aware adaptation framework built on a fully frozen pretrained MAE. First, the Variable Context Refinement Module (VCR) aggregates shared temporal information within each variable and models cross-variable contextual dependencies while preserving independent visual spaces. Second, the Temporal-Periodic Refinement Module (TPR) performs lightweight refinement at different encoder depths and explicitly models across-period temporal dependencies and within-period periodic dependencies. The two modules independently produce forecasts, which are fused through a learnable prediction-level gate. Experiments on 10 real-world datasets demonstrate that MUSE achieves state-of-the-art performance.
Chinese Translation
多变量时间序列预测在许多实际应用中至关重要。近期的大型视觉模型(LVM)通过将跨领域的视觉先验迁移至时间序列预测,提供了一种有前景的范式。然而,现有基于LVM的方法面临两个关键挑战:一是如何在保持独立视觉表示空间的同时建模跨变量依赖关系,二是如何将在自然图像上预训练的视觉骨干网络适配到时间序列图像独特的时间语义中。为应对这些挑战,我们提出了MUSE,一个构建于完全冻结的预训练MAE之上的依赖感知适配框架。首先,变量上下文精炼模块(VCR)在保留独立视觉空间的同时,聚合每个变量内部的共享时间信息,并建模跨变量的上下文依赖。其次,时间-周期精炼模块(TPR)在不同编码器深度进行轻量级精炼,显式建模跨周期的时间依赖与周期内的周期性依赖。两个模块各自独立生成预测结果,并通过可学习的预测级门控机制进行融合。在10个真实数据集上的实验表明,MUSE取得了最先进的性能。
cs.LG / 213 / 2609.24444

WPBench: A Comprehensive Benchmark for Wind Power Forecasting

WPBench:面向风电功率预测的综合基准测试
Zhu, Yuhan, Hu, Jilin, Cai, Xinying, Li, Yingshan, Ma, Li, Li, Xiangfei Qiu Linsen, Zhang, Kai, Fu, Yao, Jiang, Weihao, Yang, Bin
Abstract
Accurate, reliable, and deployable wind power forecasting is critical for power system dispatch, renewable energy integration, and electricity market operations. Progress in this field hinges on the ability to empirically and comprehensively benchmark forecasting methods. Yet existing benchmarks fall short of supporting systematic evaluation in four key aspects: 1) limited coverage of wind power scenarios across turbine scale, variable composition, and spatial structure; 2) incomplete coverage of forecasting model families; 3) evaluation metrics misaligned with wind power requirements; and 4) limited structure-aware diagnostics beyond individual temporal patterns. To address these limitations, we propose WPBench, a comprehensive, fair, and extensible benchmark for wind power forecasting. WPBench integrates 26 public datasets organized by turbine scale and variable composition, spanning single-turbine, multi-turbine, univariate, and multivariate settings. Under unified processing, training, and evaluation protocols, it benchmarks 19 representative models covering traditional methods, deep temporal models, spatio-temporal models, and foundation models. Beyond point-wise errors, WPBench assesses forecast-curve fidelity and computational efficiency, and delivers structure-aware diagnostics across temporal, variable-dependency, and spatial-dependency perspectives. Together, these capabilities enable systematic model comparison across diverse wind scenarios and provide a reusable platform for future research.
Chinese Translation
准确、可靠且可部署的风电功率预测对电力系统调度、可再生能源并网和电力市场运营至关重要。该领域的进展依赖于能否对预测方法进行实证且全面的基准测试。然而,现有基准在四个关键方面不足以支持系统性评估:1)对风电场景的覆盖有限,包括风机规模、变量构成和空间结构等方面;2)对预测模型类别的覆盖不完整;3)评估指标与风电需求不匹配;4)缺乏超越单一时间模式的结构感知诊断。为解决这些局限,我们提出了 WPBench——一个全面、公平且可扩展的风电功率预测基准。WPBench 整合了 26 个公开数据集,按风机规模和变量构成进行组织,涵盖单风机、多风机、单变量和多变量设置。在统一的处理、训练和评估协议下,该基准对 19 个代表性模型进行了测试,涵盖传统方法、深度时序模型、时空模型和基础模型(foundation models)。除了逐点误差之外,WPBench 还评估预测曲线的保真度和计算效率,并从时间、变量依赖和空间依赖三个视角提供结构感知诊断。这些能力共同实现了跨多样化风电场景的系统性模型比较,并为未来研究提供了一个可复用的平台。
cs.LG / 214 / 2609.24464

RAILS: Retrieval-Augmented Incremental LLM Clustering at Scale

RAILS:大规模检索增强的增量式大语言模型聚类
Oliya, Armin, Sawczuk, Aleksandra, Białobrzeski, Radosław
Abstract
Using a Large Language Model (LLM) as the clusterer at production scale is hard: prompts cannot hold the entire label space, and per-document serial processing does not deliver the throughput real workloads require. We present RAILS, a retrieval-augmented incremental LLM clusterer that turns clustering into a simple loop over a growing label pool and scales through document batching with bounded concurrency. On six public benchmarks RAILS exceeds the strongest prior LLM-clustering method on average, lifting accuracy from 51.2% to 59.3%, NMI from 67.2% to 74.8%, and ARI from 45.4% to 54.7%. We further report production-deployment evidence from a SaaS ticket-topic-discovery pipeline, where RAILS has replaced a traditional HDBSCAN stage with higher clustering quality, transparent prompt-driven control, and stateful incremental operation.
Chinese Translation
在生产规模下使用大语言模型(LLM)作为聚类器面临诸多困难:提示词无法容纳整个标签空间,且逐文档的串行处理无法满足真实工作负载所需的吞吐量。我们提出了 RAILS,一种检索增强的增量式 LLM 聚类器,它将聚类转化为对不断增长的标签池的简单循环,并通过带有限并发的文档批处理实现规模化。在六个公开基准上,RAILS 平均超越了此前最强的 LLM 聚类方法,将准确率从 51.2% 提升至 59.3%,NMI 从 67.2% 提升至 74.8%,ARI 从 45.4% 提升至 54.7%。此外,我们还报告了来自一个 SaaS 工单主题发现管线的生产部署证据:RAILS 已替代传统的 HDBSCAN 阶段,展现出更高的聚类质量、透明的基于提示词的控制以及有状态的增量运行能力。
cs.LG / 215 / 2609.24467

A Temporal Knowledge Graph for Music Festival Lineup Forecasting

用于音乐节阵容预测的时序知识图谱
Gastinger, Julia, Dieing, Thilo, Meilicke, Christian, Stuckenschmidt, Heiner
Abstract
Music festival lineups emerge from complex relationships among artists, genres, releases, labels, and past performances, making the prediction of future lineups a natural fit for temporal knowledge graph (TKG) forecasting. In this work, we present a TKG covering 380 festivals over 55 years, comprising more than 90K festival performance quadruples along with information on festivals, artist tours, and artist metadata, and release it as a resource for TKG forecasting evaluation. We formalize festival lineup forecasting as temporal link prediction between artists and festivals at future timestamps. We evaluate six TKG forecasting models on this task, analyze their capabilities and limitations, and compare them against Large Language Models applied zero-shot. Our resource complements existing TKG benchmarks by grounding evaluation in a concrete, real-world application domain.
Chinese Translation
音乐节阵容源于艺术家、音乐流派、音乐发行、唱片公司以及过往演出之间的复杂关系,这使得预测未来阵容成为时序知识图谱(Temporal Knowledge Graph, TKG)预测的一个天然应用场景。在本工作中,我们构建了一个涵盖55年间380个音乐节的时序知识图谱,包含超过9万个音乐节演出四元组,以及音乐节、艺术家巡演和艺术家元数据信息,并将其作为时序知识图谱预测评估的资源公开发布。我们将音乐节阵容预测形式化为未来时间点上艺术家与音乐节之间的时序链接预测任务。我们在此任务上评估了六个时序知识图谱预测模型,分析了它们的能力与局限性,并与零样本(zero-shot)应用的大语言模型进行了对比。本资源将评估建立在具体且真实的应用领域中,是对现有时序知识图谱基准的有益补充。
cs.LG / 216 / 2609.24489

Lifted Bellman Linear Programming for Offline Reinforcement Learning

面向离线强化学习的提升贝尔曼线性规划
Yang, Hyukjun, Park, Jongchan, Jeong, Narim, Lee, Donghwan
Abstract
Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead impose in-sample Bellman optimality on the critic through inequality constraints. We formulate the Lifted Bellman Linear Program (LBLP), which lifts the linear programming characterization of Bellman optimality to the joint $(Q,V)$ space so that every constraint involves only state-action pairs in the dataset. Its unique minimizer is the in-sample optimal pair, and constraints along $K$-step segments of dataset trajectories leave this minimizer unchanged for any rollout policy and horizon. Under deterministic dynamics, this minimizer lies between the best dataset return and the optimal value. Relaxing the constraints into hinge penalties recovers the same solution above a finite penalty coefficient in the tabular case. Approximate Lifted Bellman Unconstrained Minimization (ALBUM) implements this relaxation with neural networks and detaches the $K$-step rollout targets by stop gradient. Its objective contains no squared regression onto bootstrapped targets, so it can be trained without target networks or EMA updates. Under deterministic dynamics, the LBLP solution is a stationary point of the detached update under a coefficient condition independent of $\gamma$ and $K$, and the inequality constraints allow discounted returns along dataset trajectories to serve as lower bounds without off-policy correction or action chunking. On OGBench, ALBUM uses a single critic with a Gaussian policy, matches the average performance of FQL, and is comparable to recent action-chunking methods, while using the fewest parameters and the least peak GPU memory among all compared methods.
Chinese Translation
离线强化学习(RL)通常通过最小化回归损失来训练评论家(critic),其回归目标是由指数移动平均(EMA)更新的目标网络稳定的自助(bootstrapped)价值目标。多步目标包含行为策略的动作,因此需要进行离策略修正。相反,我们通过不等式约束对评论家施加样本内贝尔曼最优性。我们提出了提升贝尔曼线性规划(Lifted Bellman Linear Program, LBLP),它将贝尔曼最优性的线性规划刻画提升到联合 $(Q,V)$ 空间,使得每个约束只涉及数据集中的状态-动作对。其唯一极小值点为样本内最优对,且沿数据集轨迹 $K$ 步片段的约束对任意滚动策略和任意时域长度都不改变该极小值点。在确定性动力学下,该极小值点介于数据集最优回报与最优价值之间。将约束松弛为合页(hinge)惩罚后,在表格情形下,当惩罚系数超过有限阈值时可恢复相同解。近似提升贝尔曼无约束最小化(Approximate Lifted Bellman Unconstrained Minimization, ALBUM)用神经网络实现该松弛,并通过停止梯度(stop gradient)将 $K$ 步滚动目标分离。其目标函数不包含对自助目标的平方回归,因此无需目标网络或 EMA 更新即可训练。在确定性动力学下,在一个独立于 $\gamma$ 和 $K$ 的系数条件下,LBLP 的解是分离更新的驻点,并且不等式约束使得数据集轨迹上的折扣回报可以在无需离策略修正或动作分块(action chunking)的情况下作为下界。在 OGBench 基准上,ALBUM 仅使用单个评论家配合高斯策略,达到了与 FQL 相当的平均性能,并与近期的动作分块方法可比,同时在所有对比方法中使用的参数量最少、峰值 GPU 显存最低。
cs.LG / 217 / 2609.24504

On Emergent Capabilities and Model Merging

论涌现能力与模型合并
Zhou, Luca, Rodolà, Emanuele
Abstract
Fine-tuned checkpoints and adapters now fill public repositories, and the most common operation applied to these artifacts is model merging: arithmetic on their weights that assembles capabilities cheaply. We ask what this operation does to emergent capabilities: behaviors an artifact carries that were never an explicit training target. Studying two independent testbeds (activation oracles and emergent-misaligned models) across three model families, we find that the answer is threefold. First, merging preserves an emergent capability that both parents carry: merging two misaligned checkpoints retains most of their broad misalignment across the whole mixing range. Second, merging cannot create an emergent capability that is superadditive in its parents: no weighted merge of two single-task oracles reaches the jointly-trained oracle's auditing ability. Third, when only one parent carries the capability, merging dilutes it faster than the trained capability that accompanies it: the gap is significant in most settings. In short, emergent behaviors of an artifact do not compose the way its trained capability does.
Chinese Translation
微调后的模型检查点(checkpoints)和适配器(adapters)现已充斥于公开代码仓库中,而针对这些产物最常用的操作便是模型合并(model merging):即对其权重进行算术运算,以低成本的方式组装各种能力。我们探究了这一操作对涌现能力(emergent capabilities)的影响:涌现能力是指模型产物所携带的、从未被明确设定为训练目标的行为。我们在三个模型家族上研究了两个独立的测试平台(激活预言机(activation oracles)与涌现失调模型),发现答案可归纳为三点。第一,模型合并能够保留两个父模型均携带的涌现能力:合并两个失调的检查点,在整个混合范围内仍能保留其大部分广泛的失调行为。第二,模型合并无法创造一种在父模型之上具有超加性的涌现能力:对两个单任务预言机进行任何加权合并,都无法达到联合训练的预言机所具备的审计能力。第三,当只有一个父模型携带该能力时,模型合并使其被稀释的速度快于与之相伴的训练所得能力:在大多数设置中,这一差距都是显著的。简而言之,模型产物的涌现行为并不像其训练所得能力那样可组合。
cs.LG / 218 / 2609.24559

$t_0$: A Time-Series Foundation Model for Forecasting with Context

$t_0$:一种结合上下文的时序基础预测模型
Meyer, Lucas, Sole, Claudio, Xiang, Huikan, Li, Nicolas, Franceschino, Lucas, Quera-Bofarull, Arnau, Scholl, Maarten P., Fainberg, Joachim, Négiar, Geoffrey
Abstract
We present $t_0$, a family of open-weights foundation models for forecasting with multivariate context. We release its first two members: $\texttt{t0-alpha}$ and $\texttt{t0-beta}$, respectively 102M and 256M parameters. Both condition their forecasts on target history, past covariates, and known-future covariates, without task-specific retraining. Their transformer layers alternate attention along time and across variates. They produce probabilistic forecasts through quantile predictions. Pretraining combines curated public data with synthetic generator families constructed to contain covariate-to-target dependencies. On GIFT-Eval, $\texttt{t0-alpha}$ reaches an aggregate CRPS of 0.4941, and $\texttt{t0-beta}$ a CRPS of 0.4738 and a MASE of 0.6865, third on both and within 4.0% of the best zero-shot TSFM. On fev-bench they score 42.2 and 46.7 in skill, the latter third again and 2.0 points behind the leader. We analyze $\texttt{t0-alpha}$ in depth. Known-future covariates raise its skill by 6.3 percentage points across 30 tasks. The report also examines its calibration, its rollout strategy on long horizons, and its robustness to missing data. On the Victoria electricity-demand benchmark, $\texttt{t0-beta}$ is among the most accurate models with a context of nearly a year. In an independent Macrocosm evaluation of hourly ERCOT prices over 29 months, both cut the MAE of the lagged-price baseline by 38%.
Chinese Translation
我们提出了$t_0$,这是一系列开放权重的、可利用多变量上下文进行预测的基础模型。我们发布了其前两个成员:$\texttt{t0-alpha}$和$\texttt{t0-beta}$,参数量分别为102M和256M。两者的预测均以目标序列历史、过去协变量以及已知未来协变量为条件,无需针对特定任务进行重新训练。其Transformer层在时间维度和变量维度之间交替使用注意力机制,并通过分位数预测生成概率预测。预训练数据由精选的公开数据与专门构造的、包含协变量到目标依赖关系的合成数据生成器家族组合而成。在GIFT-Eval上,$\texttt{t0-alpha}$的聚合CRPS达到0.4941,$\texttt{t0-beta}$的CRPS为0.4738、MASE为0.6865,两者均位列第三,与最佳零样本时序基础模型(TSFM)的差距在4.0%以内。在fev-bench上,两者的技巧得分(skill)分别为42.2和46.7,后者同样位列第三,落后榜首2.0分。我们对$\texttt{t0-alpha}$进行了深入分析:在30个任务中,已知未来协变量使其技巧得分提升6.3个百分点。报告还考察了其校准特性、长预测视野下的滚动预测策略以及对缺失数据的鲁棒性。在维多利亚州电力需求基准上,$\texttt{t0-beta}$在近一年的上下文窗口下是预测精度最高的模型之一。在Macrocosm对29个月ERCOT小时级电价的独立评估中,两个模型均将滞后价格基线的MAE降低了38%。
cs.LG / 219 / 2609.24579

Universal Multi-Modal Traceformer: Integrating Heterogeneous Context for Process Event Prediction

通用多模态Traceformer:融合异构上下文的过程事件预测
Spaeh, Fabian, Fang, Jingxing, Zhe, Shandian, Shen, Bin
Abstract
Event logs arise in a wide range of real-world processes, capturing not only event activities and timestamps but also multi-modal contextual information. Existing event-sequence models, including many temporal point process approaches, primarily model event activities and timestamps while overlooking heterogeneous context, such as numerical measurements, categorical attributes, textual descriptions, and metadata associated with individual events and entire traces. In this paper, we propose Universal Multi-Modal Traceformer (UMT), a unified framework for incorporating heterogeneous process context into next-event prediction. Built on a Transformer backbone, UMT introduces a universal feature encoder that maps diverse feature types into a shared representation space and handles contextual information at both the event and trace levels. UMT further develops a per-event Perceiver module that dynamically weights contextual features and adaptively integrates them into event-token representations. To accommodate the heavy-tailed and potentially multi-modal distribution of inter-arrival times, UMT represents each interval at multiple temporal scales and jointly predicts the corresponding scale-specific quantities. Experiments on 13 real-world event logs show that UMT improves both next-event activity and time prediction over existing approaches.
Chinese Translation
事件日志产生于广泛的现实世界过程中,不仅记录事件活动和时间戳,还包含多模态的上下文信息。现有的事件序列模型,包括许多时间点过程(temporal point process)方法,主要对事件活动和时间戳进行建模,而忽视了异构上下文,例如数值测量、分类属性、文本描述以及与单个事件和整个轨迹(trace)相关的元数据。在本文中,我们提出了通用多模态Traceformer(Universal Multi-Modal Traceformer, UMT),这是一个将异构过程上下文融入下一事件预测的统一框架。UMT基于Transformer骨干网络构建,引入了一个通用特征编码器,将多种特征类型映射到共享的表示空间中,并在事件级和轨迹级两个层面处理上下文信息。UMT进一步开发了一个逐事件的Perceiver模块,该模块动态地对上下文特征进行加权,并将其自适应地整合到事件token表示中。为了适应事件间隔时间的重尾且可能呈多模态的分布,UMT在多个时间尺度上表示每个时间间隔,并联合预测相应的特定尺度量。在13个真实世界事件日志上的实验表明,UMT在下一事件活动预测和时间预测方面均优于现有方法。
cs.LG / 220 / 2609.24586

Overlay\_dx - Automating forecasting evaluation

Overlay_dx——自动化预测评估
Ngo, Long, Chamli, Mohammed Amine, Rivalan, Jonathan, Jaillon, Thomas
Abstract
Traditional evaluation metrics provides numerical values but often lack comprehensibility, hindering effective differentiation of model performances. Our work addresses this challenge by introducing overlay\_dx, a novel evaluation metric measuring the performance of time series prediction models. Overlay\_dx is a visual metric that represents the percentage of predictions falling within a confidence interval around actual values. Additionally, once evaluation results are plotted, overlay\_dx computes the area under the overlay curve, providing a quantitative measure of alignment between predicted and actual values across different thresholds and predictions. Through extensive experiments, we demonstrate that our approach offers a unified evaluation framework that combines both visual and numerical assessments, enabling improved model comparison and providing valuable insights for further research and optimization efforts in time series prediction.
Chinese Translation
传统的评估指标仅提供数值结果,但往往缺乏可理解性,阻碍了对模型性能的有效区分。我们的工作通过引入 overlay_dx 这一新型评估指标来应对这一挑战,该指标用于衡量时间序列预测模型的性能。Overlay_dx 是一种可视化指标,表示落在实际值周围置信区间内的预测所占的百分比。此外,在绘制评估结果后,overlay_dx 会计算叠加曲线下的面积,从而为不同阈值和预测下预测值与实际值之间的对齐程度提供量化度量。通过大量实验,我们证明了我们的方法提供了一个结合可视化与数值评估的统一评估框架,能够改进模型比较,并为时间序列预测领域的进一步研究与优化工作提供有价值的洞见。
cs.LG / 221 / 2609.24591

Taking a Second Look: Correcting Sea Ice Forecasts with Sparse Observations

再作审视:利用稀疏观测校正海冰预报
Zhang, Tianshuo, Xing, Xianglei, Yang, Aowen, Gao, Jia, Zhai, Wenzhe, Liu, ShanShan
Abstract
Sea ice forecasts are issued several days ahead, allowing errors to accumulate while new, often sparse sea ice concentration (SIC) observations become available. We find that fixed-propagation errors concentrate near structured, high-gradient ice edges, whereas homogeneous interiors require limited propagation, suggesting that propagation distance should be state dependent. We therefore introduce ECHO (Evidence-guided Correction with Heterogeneous prOpagation), where ECHO-Scale adapts propagation distance while preserving correction geometry, and ECHO-Delta learns a bounded residual around fixed propagation. Across all 96 standard evaluation settings spanning diverse priors, observation times, sparsity levels, geometries, and noise conditions, both outperform fixed propagation. ECHO-Delta achieves the best average accuracy, while ECHO-Scale is more robust to geometry shifts. Code is available at https://github.com/yingtian22/TAKING-A-SECOND-LOOK.
Chinese Translation
海冰预报会提前数天发布,在此期间误差会不断累积,而新的、通常较为稀疏的海冰密集度(SIC)观测数据会逐渐可用。我们发现,固定传播的误差主要集中在具有结构性的高梯度冰缘附近,而均匀的冰区内部所需的传播范围有限,这表明传播距离应当依赖于状态。因此,我们提出了ECHO(基于证据引导的异构传播校正,Evidence-guided Correction with Heterogeneous prOpagation):其中ECHO-Scale在保持校正几何结构的同时自适应调整传播距离,而ECHO-Delta则在固定传播的基础上学习一个有界残差。在涵盖不同先验、观测时间、稀疏程度、几何形态和噪声条件共96个标准评估场景中,两种方法均优于固定传播方法。ECHO-Delta取得了最优的平均精度,而ECHO-Scale对几何形态变化更为稳健。代码可在 https://github.com/yingtian22/TAKING-A-SECOND-LOOK 获取。
cs.LG / 222 / 2609.24609

GraphToolbox: A Configurable Python Framework for Graph Neural Network Forecasting

GraphToolbox:一个用于图神经网络预测的可配置Python框架
Campagne, Eloi, Amara-Ouali, Yvenn, Goude, Yannig, Kalogeratos, Argyris
Abstract
Electricity forecasting often involves spatially related signals observed over regions, substations, and feeders, and Graph Neural Networks (GNNs) provide a natural way to represent these relations. Building a complete GNN forecasting experiment is nonetheless laborious, because graph construction, model selection, training, aggregation, and interpretation sit in incompatible tools. We present GraphToolbox, an open-source Python framework that unifies these stages in one configurationdriven pipeline built on PyTorch Geometric. It offers data-driven graph construction, an adapter that instantiates and trains 51 of the 65 PyTorch Geometric convolutions together with the recurrent cells of PyTorch Geometric Temporal, online expert aggregation, forecasting interpretability, and significance testing on cached forecasts. We evaluate the pipeline in two case studies. On French regional load, the 48 convolutions included in the complete forecasting sweep fall in a band from 1.14% to 1.60% error, online aggregation lowers this to 0.98%, and the graph models improve on classical additive and boosting baselines. On net-load, direct graph models are less accurate than a classical additive model, while forecasting each physical component separately improves them without closing that gap. Both comparisons use the same experimental interface, illustrating the role of GraphToolbox in systematic architectural evaluation.
Chinese Translation
电力预测通常涉及在区域、变电站和馈线上观测到的具有空间相关性的信号,而图神经网络(GNN)为表示这些关系提供了一种自然的方式。然而,构建一个完整的GNN预测实验十分繁琐,因为图构建、模型选择、训练、聚合和解释分散在不兼容的工具中。我们提出了GraphToolbox,这是一个开源的Python框架,基于PyTorch Geometric,通过一个配置驱动的流水线将这些阶段统一起来。它提供数据驱动的图构建、一个可实例化并训练PyTorch Geometric中65种卷积算子中的51种以及PyTorch Geometric Temporal循环单元的适配器、在线专家聚合、预测可解释性,以及基于缓存预测结果的显著性检验。我们通过两个案例研究对该流水线进行了评估。在法国区域负荷预测中,完整预测扫描所涵盖的48种卷积模型的误差范围在1.14%至1.60%之间,在线聚合将误差降至0.98%,且图模型的性能优于经典的加性模型和提升(boosting)基线。在净负荷预测中,直接的图模型准确性不如经典加性模型,而分别对每个物理分量进行预测可以改善结果,但仍无法弥合这一差距。上述两组对比均使用相同的实验接口,展示了GraphToolbox在系统性架构评估中的作用。
cs.LG / 223 / 2609.24629

Augmented Hypothesis Testing with Persona-Based LLM Simulations

基于角色化大语言模型仿真的增强假设检验
Benomar, Ziyad, Marjani, Aymen Al, Missault, Paul, Mansour, Saab
Abstract
A/B testing requires large sample sizes, long timelines, and significant costs. When auxiliary predictions of experimental outcomes are available from machine learning models, uncertain prediction quality precludes replacing human experiments entirely, yet these predictions may still contain useful signal. We propose a principled framework for learning-augmented hypothesis testing that leverages predictions of unknown quality to reduce sample sizes while maintaining statistical validity. Predictions naturally vary in granularity, from coarse aggregate signals to fine-grained individual-level estimates, and our framework addresses both ends of this spectrum: (1) for population-level directional predictions, where only a binary signal on the treatment effect sign is available, we use an asymmetric test and prove consistency and robustness bounds within the learning-augmented algorithms paradigm; (2) for individual-level predictions, we introduce Generalized PPI++ (GPPI), extending Prediction-Powered Inference to handle nonlinear prediction errors through higher-dimensional transformations. Both methods benefit from accurate predictions while remaining robust to inaccurate or adversarial ones. We validate our framework using persona-based LLM simulations, where AI agents equipped with user personas predict individual behavior, as a natural prediction source spanning both granularity levels. Experiments on four real-world datasets demonstrate that our methods, combined with persona-based predictions, substantially reduce experimental costs while preserving rigorous statistical validity.
Chinese Translation
A/B测试需要大样本量、长周期和高昂的成本。当来自机器学习模型的实验结果辅助预测可用时,由于预测质量的不确定性,无法完全取代人类实验,但这些预测仍可能包含有用的信号。我们提出了一个有原则的学习增强假设检验框架,利用质量未知的预测来减少样本量,同时保持统计有效性。预测的粒度自然各异,从粗粒度的聚合信号到细粒度的个体级估计,我们的框架涵盖了这一谱系的两端:(1)对于总体层面的方向性预测,即仅能获得关于处理效应符号的二值信号,我们采用非对称检验,并在学习增强算法范式内证明了一致性和鲁棒性界;(2)对于个体层面的预测,我们提出了广义PPI++(Generalized PPI++, GPPI),将预测驱动推断(Prediction-Powered Inference)扩展至通过高维变换处理非线性预测误差。两种方法都能从准确的预测中获益,同时对不准确或对抗性的预测保持鲁棒。我们使用基于角色的LLM仿真来验证我们的框架,其中配备用户角色的AI智能体预测个体行为,作为同时跨越两个粒度层级的自然预测来源。在四个真实世界数据集上的实验表明,我们的方法结合基于角色的预测,在保持严格统计有效性的同时,显著降低了实验成本。
cs.LG / 224 / 2609.24646

iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs

iSDFT:面向大语言模型持续学习的信息近端自蒸馏方法
Khamis, Ahmed Khaled, Ji, Xiaotong, Jaber, Hassan, Tutunov, Rasul, Zimmer, Matthieu, Wang, Jun, Bou-Ammar, Haitham
Abstract
On-policy self-distillation fine-tuning (SDFT) learns new skills from demonstrations while reducing forgetting, but it always distils toward the full demonstration-conditioned teacher. This fixes teacher influence at the full-teacher endpoint, providing no control over how much demonstration information should be transferred at each prediction state. We introduce Information-Proximal SDFT (iSDFT), which instead treats the teacher as a budgeted source of information. At each token, iSDFT selects the distribution closest to the current student that satisfies a prescribed teacher-information constraint, yielding a closed-form exponential target with a locally determined tilt. To control cumulative drift, we further anchor the student to its frozen base policy. Across four heterogeneous LLM backbones and two specialisation tasks, iSDFT improves vanilla SDFT in 7 of 8 model-task settings and matches it in the remaining one. It also provides tighter retention on the original SDFT benchmark suite, with 73% of evaluations remaining within 0.5 points of the base model versus 52% for the strongest baseline, while achieving the largest mean improvement on all ten additional mathematics, coding, and competition-mathematics benchmarks. These results show that controlling how much and when teacher information is introduced improves specialisation while preserving broader capability.
Chinese Translation
在线策略自蒸馏微调(SDFT)能够从演示样本中学习新技能,同时减少遗忘,但它总是向以完整演示为条件的教师模型进行蒸馏。这使得教师模型的影响被固定在完整教师这一终点上,无法控制在每个预测状态下应迁移多少演示信息。我们提出信息近端SDFT(Information-Proximal SDFT,iSDFT),将教师模型视为一个有预算限制的信息来源。在每个词元上,iSDFT选择在满足预设教师信息约束的条件下与当前学生模型最近的分布,从而得到一个具有局部确定倾斜参数的闭式指数形式目标分布。为控制累积漂移,我们进一步将学生模型锚定在其冻结的基础策略上。在四个异构大语言模型骨干和两个专门化任务上,iSDFT在8个模型-任务设置中的7个上优于原始SDFT,在其余1个上与其持平。在原始SDFT基准测试套件上,iSDFT还提供了更强的知识保持能力:73%的评测结果保持在基础模型0.5分以内,而最强基线仅为52%;同时在全部十个额外的数学、代码和竞赛数学基准上,iSDFT取得了最大的平均提升。这些结果表明,控制教师信息引入的多少和时机,能够在保持更广泛能力的同时提升专门化性能。
cs.LG / 225 / 2609.24651

Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement

校正强制:面向生成式语音增强中扩散模型与流模型的统一后训练方法
Yao, Qing, Gao, Lijian, Mao, Qirong
Abstract
Diffusion and flow models, as promising generative paradigms for speech enhancement, face a training--inference mismatch: training uses analytical path states, whereas inference recursively evaluates models on self-generated rollout states along discretized sampling trajectories. This mismatch causes prediction and discretization errors to accumulate. To address it, we introduce Corrective Forcing (CoF), a post-training paradigm that forces diffusion and flow models to learn from self-generated rollouts and correct their predictions. CoF corrects clean-speech predictions on rollout states toward the ground truth under dynamic sampling schedules, exposing the model to varying inference conditions. It further regularizes local evolution using locally corrected counterfactual transitions as references for factual transitions. By expressing model outputs through a shared clean-speech prediction parameterization, CoF applies the same post-training objective across diffusion and flow formulations. Experiments with SB-VE and OT-CFM demonstrate improvements in perceptual quality and reconstruction fidelity, together with robust performance across different numbers of sampling steps.
Chinese Translation
扩散模型与流模型作为极具前景的语音增强生成式范式,面临着训练与推理不匹配的问题:训练时使用解析路径状态,而推理时在离散化采样轨迹上对模型自身生成的展开状态进行递归评估。这种不匹配导致预测误差与离散化误差不断累积。为解决该问题,我们提出了校正强制,这是一种后训练范式,它强制扩散模型和流模型从自身生成的展开状态中学习并校正其预测。CoF 在动态采样调度下,将模型在展开状态上的干净语音预测向真实值方向校正,使模型暴露于多样的推理条件之下。此外,CoF 利用局部校正的反事实转移作为事实转移的参考,对局部演化进行正则化。通过用共享的干净语音预测参数化来表示模型输出,CoF 能够在扩散和流两种建模形式上应用相同的后训练目标。基于 SB-VE 和 OT-CFM 的实验表明,该方法在感知质量和重建保真度方面均有所提升,并且在不同的采样步数下均保持稳健的性能。
cs.LG / 226 / 2609.24678

Muon Can Outperform Dedicated Continual Learning Methods

Muon可以超越专门的持续学习方法
Sincari, Sebastian George, Gheorghe, Bogdan Alexandru, Barbalau, Antonio
Abstract
Continual learning with Low-Rank Adapters (LoRA) typically mitigates forgetting by penalizing the overlap between a new update and the accumulated past weights, which discourages certain update directions without controlling how an update distributes its energy over the ones that remain. We ask whether that restriction has to be task-aware, or whether a generic one supplied by the optimizer is enough. We train a plain incremental LoRA (IncLoRA) with Muon, which orthogonalizes each update, and compare it against O-LoRA and ELLA over five seeds and three task orders on the Standard CL Benchmark and three seeds on TRACE. IncLoRA+Muon reaches the accuracy band of the dedicated methods on Standard CL and improves on every AdamW configuration on TRACE. One update-constraining mechanism is enough, whether it comes from the loss or from the optimizer; on Standard CL a second one does not help, and for the most restrictive method it costs 8.4 points of accuracy and the plasticity to fit each task. What separates the two optimizers is not the size of the update, which under Muon is 0.91 to 2.06 times that under AdamW, but how it is distributed. AdamW confines it to between 1.4 and 1.8 effective singular directions, Muon spreads it over 7.0, and the two do not overlap in any tracked run. Part of the advantage usually attributed to dedicated CL methods may therefore be explained by the geometry of the optimizer's updates.
Chinese Translation
基于低秩适配器(LoRA)的持续学习通常通过惩罚新更新与已积累的过去权重之间的重叠来缓解遗忘,这抑制了某些更新方向,却未控制更新如何将其能量分配到剩余方向上。我们探究这种约束是否必须是任务感知的,还是由优化器提供的通用约束就足够了。我们使用Muon优化器(其将每次更新正交化)训练简单的增量式LoRA(IncLoRA),并在Standard CL Benchmark上以五个随机种子和三种任务顺序、在TRACE上以三个随机种子将其与O-LoRA和ELLA进行比较。IncLoRA+Muon在Standard CL上达到了专门方法的准确率区间,并在TRACE上优于所有AdamW配置。无论是来自损失函数还是来自优化器,一种更新约束机制就足够了;在Standard CL上,第二种约束并无帮助,而对于约束最强的方法,它导致准确率损失8.4个百分点,并丧失了拟合每个任务的可塑性。区分两种优化器的并非更新的大小——Muon下的更新幅度为AdamW下的0.91至2.06倍——而是更新的分布方式。AdamW将更新限制在1.4至1.8个有效奇异方向上,而Muon将其分散到7.0个方向上,且二者在所有被跟踪的运行中均无重叠。因此,通常归因于专门持续学习方法的部分优势,可能可以用优化器更新的几何特性来解释。
cs.LG / 227 / 2609.24679

Guaranteed Low-Rank Tensor Recovery from Modewise Measurements via Normalized Block-Weighted Riemannian Gradient Descent

基于归一化分块加权黎曼梯度下降的模态度量下低秩张量恢复的保证性研究
Zhou, Yushi, Zhang, Feng
Abstract
We consider the recovery of low-multilinear-rank tensors from linear measurements and propose an adaptive block-weighted modewise Riemannian gradient descent method. The method combines memory-efficient modewise measurements with a normalized adaptive weighting strategy for the core and factor components of the Riemannian gradient. The weighting improves convergence without increasing the multilinear-rank bound of the search direction or the size of the reduced core used for retraction. Under the tensor restricted isometry property and a suitable initialization, we establish local linear convergence and derive sampling guarantees for sub-Gaussian and subsampled orthogonal with random sign (SORS) measurements. Numerical experiments on synthetic low-Tucker-rank tensors show that the proposed method reduces iteration counts and computational time while maintaining reliable recovery performance, especially near the recovery threshold and for structured SORS measurements.
Chinese Translation
我们研究了从线性测量中恢复低多线性秩张量的问题,并提出了一种自适应的分块加权模态度黎曼梯度下降方法。该方法将内存高效的模态度测量与针对黎曼梯度的核张量分量和因子分量的归一化自适应加权策略相结合。这种加权机制在不增加搜索方向的多线性秩界或用于回缩(retraction)的约化核张量大小的前提下改善了收敛性。在张量受限等距性质(tensor restricted isometry property)及适当初始化的条件下,我们建立了局部线性收敛性,并针对sub-Gaussian测量和带随机符号的子采样正交(SORS)测量推导了采样保证。在合成低Tucker秩张量上的数值实验表明,所提方法在保持可靠恢复性能的同时减少了迭代次数和计算时间,尤其是在恢复阈值附近以及结构化SORS测量的情形下表现突出。
cs.LG / 228 / 2609.24718

A Federated Artificial Intelligence Framework for Optimizing Pancreatic Cancer Treatment - Strategy Update

用于优化胰腺癌治疗的联邦人工智能框架——策略更新
Hauschild, Anne-Christin, Aleyasin, Amirreza, Beyer, Nils H., Fricke, Lisa, Hügel, Jonas, Moradpour, Maryam, Nguyen, Anh-Tien, Park, Youngjun, Rheinländer, Sophia, Beissbarth, Tim, Hessmann, Elisabeth, Middeke, Martin, Lauth, Matthias, Reichert, Maximilian, Sax, Ulrich
Abstract
While a centralized approach involving patient consent to collect and analyze data centrally would theoretically offer the best data quality and predictive performance, it is not always feasible in practice. Federated Learning (FL) architectures have shown to be a very promising approach to use and access distributed disease related resources within the GDPR boundaries. In a previous case report, we described the preconditions at the participating sites and necessary administrative and process related steps to prepare data, people and infrastructure for improving subtype identification and assessing treatment options in pancreatic cancer. We update this report sharing our experience in tackling the challenges and show preliminary results of the actual federated learning AI pipelines. At the participating sites, we have to identify and annotate the data being accessible after extraction and transformation in a local FL hub - in our case a centrally developed and distributively deployed Docker container. This container comprises the FL scripts generating local models. We apply a newly developed FL algorithm considering all local features, including partial overlapping features specific to the local sites. Theoretically, an annotation in a cancer setting should succeed using the German oncology core data set (oBDS), which is already utilized for mandatory reporting to cancer registries, and can be sustained in the FL setting. The FL algorithms deal robustly with partially overlapping features as we showed with public data sets. Major roadblocks including straightening operational concepts for the infrastructures, ethics approval for such novel architectures and support for every site have been addressed. However, scaling up this approach in the future faces hurdles; while including broader multi-modal data sets should be feasible, large-scale deployment to more sites remains challenging.
Chinese Translation
尽管在理论上,通过患者知情同意将数据集中收集和分析的集中式方法能够提供最佳的数据质量和预测性能,但在实践中并不总是可行。联邦学习(Federated Learning, FL)架构已被证明是在《通用数据保护条例》(GDPR)框架内使用和访问分布式疾病相关资源的一种非常有前景的方法。在此前的案例报告中,我们描述了参与机构的前提条件,以及为改善胰腺癌亚型识别和评估治疗方案而在数据、人员和基础设施方面所需的管理与流程相关准备步骤。本报告对该工作进行了更新,分享了我们应对相关挑战的经验,并展示了联邦学习人工智能流程的初步结果。在各参与机构,我们需要对在本地FL中心(在我们案例中是一个集中开发、分布式部署的Docker容器)中经过提取和转换后可访问的数据进行识别和标注。该容器包含用于生成本地模型的FL脚本。我们应用了一种新开发的FL算法,该算法考虑了所有本地特征,包括各本地站点特有的部分重叠特征。从理论上讲,在癌症场景中的标注应可借助德国肿瘤学核心数据集(oBDS)实现,该数据集已被用于向癌症登记处的强制性报告,并可在FL环境中持续使用。正如我们利用公开数据集所展示的,FL算法能够稳健地处理部分重叠的特征。主要障碍已得到解决,包括理清基础设施的运行概念、针对此类新型架构的伦理审批以及对每个站点的支持。然而,未来扩大该方法规模仍面临阻碍;纳入更广泛的多模态数据集应是可行的,但向更多站点的大规模部署仍然具有挑战性。
cs.LG / 229 / 2609.24741

An Exact Junction-Tree Extended Formulation for Optimal Classification Trees

基于联合树(Junction-Tree)扩展公式的最优分类树精确求解方法
TU, Jiancheng, WenqiFan
Abstract
We develop an exact linear programming (LP) formulation for bounded-depth classification trees with binary features, using a junction-tree representation. The formulation is integral and supports recursive subtree optimization. Exact reductions make the model smaller while preserving the optimal value and recovery of an optimal tree. The reduced model supports two solution methods: column generation and message passing. Column generation solves integral restricted LPs and uses bounds over the full feasible domain to certify optimality. Message passing recursively combines optimal subtree costs. Both methods solve common subtree problems that, once the preceding tree decisions are fixed, can be evaluated independently and in parallel. Computational experiments show that the exact reductions substantially reduce the size of the junction-tree formulation. The resulting linear programming formulation certifies instances for which the tested mixed-integer formulation does not establish optimality within the same computational budget, while the column-generation and message-passing methods certify more instances and achieve an order-of-magnitude reduction in geometric-mean runtime relative to an existing state-of-the-art exact method for optimal classification trees.
Chinese Translation
我们利用联合树表示,为具有二值特征的限定深度分类树建立了一种精确的线性规划(LP)公式。该公式具有整数性,并支持递归的子树优化。精确约简可在保持最优值以及最优树可恢复性的前提下缩小模型规模。约简后的模型支持两种求解方法:列生成(column generation)和消息传递(message passing)。列生成方法求解整数受限线性规划,并利用整个可行域上的界来证明最优性。消息传递方法递归地合并最优子树代价。两种方法均可求解公共子树问题——一旦确定了之前的树决策,这些问题便可以独立且并行地求解。计算实验表明,精确约简显著缩小了联合树公式的规模。所得到的线性规划公式能够在相同的计算预算内,证明被测试的混合整数公式无法确立最优性的实例;同时,与现有的最优分类树精确求解方法相比,列生成与消息传递方法可证明更多实例,并在几何平均运行时间上实现了一个数量级的降低。
cs.LG / 230 / 2609.24746

Enhancing Transformer Representations of Symbolic ODE Expressions

增强符号常微分方程表达式的Transformer表示
Fan, Xiyue, Prugel-Bennett, Adam, Middleton, Stuart E.
Abstract
Existing approaches to solving differential equations, such as symbolic regression, physics informed neural networks, and neural operators, typically focus on numerical approximations or blind symbolic search via fitting to numerical data. Less attention has been paid to learning structured representations of mathematical expressions that preserve commutative properties and could support mathematical reasoning in symbolic forms. Transformer models have shown strong capabilities in solving symbolic differential equations. However, standard positional embeddings in transformers are designed for sequence data. Symbolic differential equations are naturally represented by expression trees, so these positional embeddings may not efficiently capture their hierarchical structures. We investigate existing tree positional embeddings in symbolic ordinary differential equation (ODE) tasks. We systematically study their effectiveness under different settings. Our results show that tree positional embeddings aid learning in early epochs and continue to improve performance throughout, ultimately yielding consistent advantages across various data sizes and tasks. Based on learned structural representations, we apply contrastive learning to support the commutative property in mathematics. Ablation studies provide insight into how these methods interact in modelling symbolic mathematical structures.
Chinese Translation
现有的微分方程求解方法,如符号回归、物理信息神经网络(Physics Informed Neural Networks)和神经算子(Neural Operators),通常侧重于数值近似或通过拟合数值数据进行盲目的符号搜索。而学习能够保留交换性质、并支持符号形式数学推理的结构化数学表达式表示,则受到的关注较少。Transformer模型在求解符号微分方程方面已展现出强大能力。然而,Transformer中的标准位置嵌入是为序列数据设计的。符号微分方程天然地由表达式树表示,因此这些位置嵌入可能无法有效捕捉其层次结构。我们在符号常微分方程(ODE)任务中研究了现有的树位置嵌入方法,并系统地考察了它们在不同设置下的有效性。结果表明,树位置嵌入在训练早期有助于学习,并在整个训练过程中持续提升性能,最终在各种数据规模和任务中均带来一致的优势。基于学习到的结构表示,我们应用对比学习来支持数学中的交换性质。消融研究深入揭示了这些方法在建模符号数学结构时的相互作用。
cs.LG / 231 / 2609.24754

Inference of Unknown Dynamical Components Using Next Generation Reservoir Computing: From Chaotic Systems to Climate Data

基于下一代储备池计算的未知动力学分量推断:从混沌系统到气候数据
Budnick, Jule, Keane, Andrew, Yanchuk, Serhiy
Abstract
We investigate next generation reservoir computing (NGRC) as a data-driven approach for inferring unseen components of dynamical systems. We compare NGRC with traditional reservoir computing (RC) using the Lorenz and R\"ossler system, where two unknown components are inferred from one given component. For both systems, NGRC achieves accurate results while requiring fewer training data and less computational time than RC. We identified an inverse proportional behavior between the number of time-delayed steps needed for NGRC and the temporal resolution, indicating that the physical time span covered by the delay interval is an important factor in determining the required number of delayed steps. Finally, we apply NGRC to the observational climate data of ENSO (El Ni\~no--Southern Oscillation) and infer one observable from the remaining variables. Despite the noise and complexity of the real-world data, the NGRC shows promising results. Our findings demonstrate the potential of NGRC for efficient inference of unseen components in both controlled dynamical systems and real-world data.
Chinese Translation
我们研究了下一代储备池计算(Next Generation Reservoir Computing, NGRC)作为一种数据驱动方法,用于推断动力系统中不可见的分量。我们以Lorenz系统和Rössler系统为对象,将NGRC与传统储备池计算(RC)进行比较,从单个给定分量推断两个未知分量。在这两个系统中,NGRC均获得了准确的结果,且所需训练数据更少、计算时间更短。我们发现NGRC所需的时间延迟步数与时间分辨率之间呈反比关系,这表明延迟区间所覆盖的物理时间跨度是决定所需延迟步数的重要因素。最后,我们将NGRC应用于ENSO(厄尔尼诺—南方涛动)的观测气候数据,从其余变量中推断一个可观测变量。尽管真实世界数据存在噪声和复杂性,NGRC仍展现出令人满意的结果。我们的研究结果表明,NGRC在受控动力系统和真实世界数据中均可实现不可见分量的高效推断,具有广阔的应用前景。
cs.LG / 232 / 2609.24797

Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention

Complex KDA:理解与增强 Kimi Delta Attention 的表达能力
Siems, Julien, Grazzi, Riccardo, Pöppel, Korbinian, Singh, Jaisidh, Zela, Arber, Carstensen, Timur, Jitsev, Jenia, Hutter, Frank, Cevher, Volkan, Orvieto, Antonio, Klein, Aaron
Abstract
Linear RNNs based on the delta-rule enable efficient sequence modeling, but their linear updates with a low-rank correction constrain their expressivity. Prior work has shown that composing two delta-rule transitions in a single recurrent update can model a 2D rotation, but this increases the rank and the cost of the updates compared to a single transition. We show that Kimi Delta Attention (KDA) can realize 2D rotations by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate. This requires extending the parameter ranges of KDA by combining two existing range extensions: allowing gates in $[-1,1]$ and the delta-rule coefficient $\beta$ in $[0,2]$. We call the resulting model Complex KDA (CKDA). It preserves KDA's stability and efficiency, with transitions that remain diagonal-plus-rank-one and non-expansive, while reaching the state-tracking expressivity of DeltaProduct$_2$. We characterize the expressivity of CKDA and prove that every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix. A single CKDA layer can track every finite group isomorphic to a subgroup of $\mathrm{SO}(3)$, and many state-tracking results use one fewer layer for CKDA compared to other diagonal-plus-rank-one Linear RNNs. Empirically, combining both extensions yields the strongest length extrapolation among tested KDA range settings on $S_3$, $S_4$, and periodic audio continuation. In language modeling, CKDA outperforms Transformers and other linear RNNs, obtains similar results to a KDA baseline, and shows promising scaling behavior. Our code is open source at https://github.com/OpenEuroLLM/ComplexKDA and our models are available at https://huggingface.co/collections/openeurollm/complexkda.
Chinese Translation
基于 delta 规则的线性 RNN 能够实现高效的序列建模,但其带有低秩修正的线性更新限制了其表达能力。已有研究表明,在单次循环更新中组合两个 delta 规则转移可以建模二维旋转,但与单一转移相比,这会增加更新的秩和计算成本。我们发现,Kimi Delta Attention(KDA)可以通过将单个 delta 规则变换与其逐通道门控所提供的第二次反射相结合来实现二维旋转。这需要扩展 KDA 的参数范围,具体做法是结合两种已有的范围扩展方式:允许门控取值于 $[-1,1]$,以及允许 delta 规则系数 $\beta$ 取值于 $[0,2]$。我们将由此得到的模型称为 Complex KDA(CKDA)。它保留了 KDA 的稳定性和高效性,其转移矩阵仍为对角加秩一(diagonal-plus-rank-one)形式且不具扩张性,同时达到了 DeltaProduct$_2$ 的状态追踪表达能力。我们刻画了 CKDA 的表达能力,并证明每一个正交的对角加秩一矩阵都恰好是一个 CKDA 转移矩阵。单个 CKDA 层可以追踪所有同构于 $\mathrm{SO}(3)$ 子群的有限群,并且在许多状态追踪任务上,与其他对角加秩一的线性 RNN 相比,CKDA 所需的层数少一层。实验表明,在 $S_3$、$S_4$ 和周期性音频延续任务上,同时结合两种扩展方式在所测试的 KDA 参数范围设置中取得了最强的长度外推能力。在语言建模任务中,CKDA 的表现优于 Transformers 和其他线性 RNN,取得了与 KDA 基线相近的结果,并展现出良好的扩展行为。我们的代码已在 https://github.com/OpenEuroLLM/ComplexKDA 开源,模型可在 https://huggingface.co/collections/openeurollm/complexkda 获取。
cs.LG / 233 / 2609.24823

G-NAC: Graph Neural Automata Clustering via Emergent Domain Formation

G-NAC:基于涌现域形成的图神经自动机聚类
Miller, Keith, Crawford, Tristan
Abstract
We introduce Graph Neural Automata Clustering (G-NAC), an unsupervised clustering method in which observations interact as cells on a fixed neighborhood graph. A shared recurrent graph-neural cellular rule evolves latent domain states through local interactions, which are converted into a rank-based spectral affinity for partitioning. Across 73 clustering tasks from 57 benchmark datasets, G-NAC achieved a mean adjusted Rand index (ARI) of 0.7951, comparable to Genie at 0.7941 and higher than the other evaluated baselines. Empirical training time and GPU memory scaled approximately linearly from 5,000 to 100,000 nodes. Learned transition rules also transferred from smaller source graphs to independent 100,000-node samples generated under matched conditions. These results demonstrate a recurrent graph-clustering formulation while identifying dependencies on graph quality, readout design, and source-target similarity.
Chinese Translation
我们提出了图神经自动机聚类(Graph Neural Automata Clustering, G-NAC),这是一种无监督聚类方法,其中观测数据作为固定邻域图上的细胞进行交互。一个共享的循环图神经细胞规则通过局部交互演化潜在域状态,这些状态被转换为基于秩的谱亲和度用于划分。在来自57个基准数据集的73个聚类任务上,G-NAC取得了0.7951的平均调整兰德指数(ARI),与Genie的0.7941相当,并高于其他被评估的基线方法。实证训练时间和GPU显存从5,000到100,000个节点大致呈线性扩展。学习到的转移规则还能够从较小的源图迁移到在匹配条件下生成的独立100,000节点样本。这些结果展示了一种循环图聚类框架,同时指出了其对图质量、读出设计和源-目标相似性的依赖。
cs.LG / 234 / 2609.24862

When Tomorrow Becomes Today: Self-Evolving Policies for Agentic Time-Series Forecasting

当明天成为今天:面向智能体时间序列预测的自演化策略
Hu, Yifan, Dai, Xilin, Qu, Zhiyuan, Liu, Yiding, Dong, Zewei, Yang, Jiang-ming, Xu, Qiang
Abstract
Agentic time series forecasting concerns systems whose underlying mechanisms evolve, making the relative effectiveness of numerical models, reasoning strategies, and intervention rules inherently time-varying. Consequently, a time series agent must adapt the forecasts it produces and the orchestration policy that determines which components to trust and how to coordinate them. The deployment process naturally provides supervision for this adaptation as forecast horizons elapse and realized targets reveal the effectiveness of earlier decisions. Committing all numerical expert forecasts and candidate agent paths before target observation allows each realized outcome to evaluate the entire alternative set, providing delayed feedback without additional annotation. However, existing time series agents primarily incorporate prior experience through forecast refinement, reflection, or retrieval, without systematically converting realized outcomes into persistent updates to the joint orchestration policy governing later origins. To exploit this delayed feedback systematically, we introduce TimEvolve, a frozen-backbone time series agent that converts each realized outcome into persistent joint updates of expert trust, agent path selection, and intervention strength. A temporally ordered predict, reveal, and update protocol applies this feedback to subsequent forecasts. Experiments across eight Time-MMD domains show that TimEvolve achieves the best average MSE and MAE ranks among fifteen methods and the lowest errors on both metrics in seven domains. These results demonstrate the value of learning forecasting policies from the futures encountered during deployment.
Chinese Translation
智能体时间序列预测所研究的系统,其底层机制会随时间演化,使得数值模型、推理策略与干预规则的相对有效性本质上随时间变化。因此,时间序列智能体必须调整其产生的预测结果,以及决定信任哪些组件及如何协调这些组件的编排策略。随着预测时间跨度的推进和实际观测值揭示先前决策的有效性,部署过程自然地为这种自适应提供了监督信号。在观测目标之前预先提交所有数值专家预测和候选智能体路径,使得每个实际结果都能评估整个备选集合,从而在无需额外标注的情况下提供延迟反馈。然而,现有时间序列智能体主要通过预测修正、反思或检索来利用先验经验,而未能系统地将实际结果转化为对后续预测起点的联合编排策略的持久更新。为了系统地利用这种延迟反馈,我们提出了TimEvolve,一种基于冻结骨干网络的时间序列智能体,它将每个实际结果转化为对专家信任度、智能体路径选择和干预强度的持久联合更新。通过按时间顺序执行的预测、揭示和更新协议,将该反馈应用于后续预测。在Time-MMD八个领域上的实验表明,TimEvolve在十五种方法中取得了最优的平均MSE和MAE排名,并在七个领域中在这两项指标上均取得最低误差。这些结果证明了从部署过程中所遇到的未来中学习预测策略的价值。
cs.LG / 235 / 2609.24882

Learning Prognostic Variables for AI Convective Parameterizations via Symbolic Distillation

通过符号蒸馏学习AI对流参数化方案的预报变量
Schönfeld, Jurij, Beucler, Tom, Savre, Julien, Sherwood, Steven, Eyring, Veronika
Abstract
Hybrid AI-physics climate modeling aims to improve coarse (~100km-resolution) Earth system models by learning to parameterize subgrid processes from high-fidelity data. However, this so far mostly involves local-in-time, diagnostic parameterizations, in which the subgrid state depends only on the current coarse state with no memory of previous states, which is unrealistic for processes such as convection that have intrinsic persistence. To address this, we enhance local-in-time parameterizations by learning prognostic variables that compactly carry important, additional past information where no explicit sub-grid information is available. First we compress past information into a low-dimensional latent space using an autoencoder, which then informs a neural network trained to parameterize targeted subgrid-scale processes. We then replace the autoencoder with symbolic equations that govern the time evolution of the latent variables, yielding additional prognostic memory variables that can be integrated alongside the resolved atmospheric state. We evaluate this approach on two systems: the Lorenz-96 model (online) and surface precipitation from high-resolution atmospheric simulations (offline). A forced multivariate linear ordinary differential equation recovers most of the added value achieved by the autoencoder-based approach in both experiments. Benchmarked against diagnostic parameterizations without memory, our memory-informed approach improves climate statistics and temporal structure, including a realistic diurnal cycle of tropical land precipitation.
Chinese Translation
AI-物理混合气候建模旨在通过从高保真数据中学习次网格过程的参数化方案,来改进粗分辨率(约100公里)的地球系统模型。然而,迄今为止这类方法主要采用局地时刻的诊断型参数化方案,即次网格状态仅依赖于当前的粗网格状态,而没有对先前状态的记忆,这对于具有内在持续性的过程(如对流)来说是不现实的。为解决这一问题,我们通过学习预报变量来增强局地时刻参数化方案,这些变量能够紧凑地携带重要的额外历史信息,以弥补显式次网格信息的缺失。首先,我们使用自编码器将历史信息压缩到低维潜在空间中,然后用其训练一个神经网络来参数化目标次网格尺度过程。随后,我们用控制潜在变量时间演化的符号方程替换自编码器,从而得到额外的预报型记忆变量,这些变量可与解析的大气状态一同积分。我们在两个系统上评估了该方法:Lorenz-96模型(在线)和高分辨率大气模拟的地表降水(离线)。在两个实验中,一个受迫多变量线性常微分方程恢复(重现)了基于自编码器方法所取得的大部分增益。与无记忆的诊断型参数化方案相比,我们基于记忆的方法改善了气候统计特征和时间结构,包括热带陆地降水的真实日循环。
cs.LG / 236 / 2609.24942

Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization

推理处的精确性:一种面向分布外泛化的表征判据
Rocha, Filipe Marinho, Dutra, Inês, Costa, Vítor Santos, Reis, Luís Paulo
Abstract
A model generalizes outside its training distribution only when it computes a representation structurally equivalent to the generating mechanism, not an approximation fitted to it. Such equivalence is necessary for exactness in and out of distribution, and extrapolation is governed by this exactness at inference, whatever its realization. Tensor Logic shows this: a zero-temperature contraction is equivalent to discrete logic, deducing in place with no artefact extracted, its tensors Boolean, its embeddings orthonormal, only its arithmetic continuous. Lacking infinite recursion it reaches Datalog, not Prolog, and though exact over closed domains it needs external memory to bind a novel entity. The criterion needs neither a discrete representation nor an extracted expression, and constrains inference, not training: an exact marginal in $[0,1]$ passes, a Neural Network thresholded to a hard label does not. Logic Tensor Networks fail it, while differentiable ILP and Tensor Logic at $T=0$ pass. Piecewise-affine extrapolation divergence and an inability to bind novel entities are two faces of a shortfall in exact representability. For hybrid architectures, a propagation rule follows: the output inherits the bounds of every fitted estimator on its path, explaining which axes fail in equivariant models and the ARC-AGI induction/transduction split. Only an exact hypothesis class certifies what the training data leave underdetermined: on a law-derived partition it finds the $56.3\%$ of distant queries that are answerable, which ensembles meet with false confidence and distance metrics rank backwards. Common inductive biases, from symmetries to memory, reach exactness only because humans inject them, an argument for inducing exact representations rather than fitting surrogates whose residuals, even at the arithmetic floor in training, diverge outside the data and compound under composition.
Chinese Translation
一个模型只有在计算出与生成机制在结构上等价的表征时,才能在其训练分布之外泛化,而非拟合于该机制的近似。这种等价性是分布内与分布外精确性的必要条件,而外推由推理处的这种精确性所支配,无论其如何实现。Tensor Logic(张量逻辑)展示了这一点:零温度缩并等价于离散逻辑,可原位进行演绎而无需提取任何人为产物,其张量为布尔型,其嵌入为正交归一的,唯有其算术是连续的。由于缺乏无限递归,它只能达到Datalog而非Prolog的水平;尽管在封闭域上是精确的,但它需要外部记忆来绑定新实体。该判据既不需要离散表征,也不需要提取出的表达式,且约束的是推理而非训练:$[0,1]$ 区间内的精确边缘概率可以通过,而将神经网络阈值化为硬标签则不能。Logic Tensor Networks(逻辑张量网络)无法通过该判据,而可微ILP(归纳逻辑编程)和 $T=0$ 下的Tensor Logic则可以通过。分段仿射外推的发散性与无法绑定新实体,是精确可表征性缺失的两个侧面。对于混合架构,可得出一条传播规则:输出继承其路径上每个被拟合估计器的界限,这解释了等变模型中哪些轴会失效,以及ARC-AGI任务中归纳/传导的分裂现象。只有精确的假设类才能判定训练数据所未能确定的内容:在由规律导出的划分上,它能找出 $56.3\%$ 可回答的远距离查询,而集成方法以虚假的置信度应对这些问题,距离度量则会将它们的排序颠倒。常见的归纳偏置,从对称性到记忆,之所以能达到精确性,仅因为是人类注入的——这为归纳精确表征而非拟合替代代理提供了论据,因为后者的残差即使在训练中处于算术下限时,也会在数据之外发散并在复合运算下不断累积。
cs.LG / 237 / 2609.24947

Learning Physics from an Imperfect Ancestor

从不完美的祖先模型中学习物理
Mousavi, S. Mohammad, Kadeethum, Teeratorn, Bouklas, Nikolaos, Goswami, Somdatta
Abstract
Neural operators evaluate parametric partial differential equations cheaply but degrade sharply outside their training distribution. Physics-informed neural networks avoid dependence on labeled data, yet their optimization can be basin-fragile: when the governing residual admits multiple solutions, a PINN trained from scratch may converge to a physically incorrect state despite achieving a small residual. We show that these failure modes can be addressed jointly: an imperfect NO provides the structural prior needed to place a PINN in the correct solution basin, while the PDE residual refines the solution beyond the operator's accuracy. We introduce a three-stage framework that freezes the spatial basis of a physics-informed NO, extrapolates its solution branch to an out-of-distribution parameter using a polynomial continuation prior, and distills the resulting field into a fresh PINN. The NO need not be accurate at the target; it transfers solution-branch information, while PDE residual minimization in the PINN governs convergence. We evaluate the framework on three nonlinear PDEs: 1D viscous Burgers, 2D steady Allen-Cahn near a pitchfork bifurcation, and 2D steady lid-driven cavity flow. For Allen-Cahn, where the trivial solution satisfies the PDE residual exactly, a standard PINN collapses to the trivial zero branch, whereas distillation from the crude extrapolated operator recovers the non-trivial branch that matches the finite-difference reference. For the lid-driven cavity, extrapolating to a Reynolds number of Re = 3200 accelerates convergence to the correct physical state, achieving competitive accuracy using fewer parameters and optimization steps than recent literature baselines. These results establish a simple principle: an NO need not accurately predict the solution to be useful; it only needs to identify the correct basin from which PINN optimization can recover it.
Chinese Translation
神经算子(Neural Operators, NO)能够以低成本求解参数化偏微分方程,但在训练分布之外其性能会急剧退化。物理信息神经网络(Physics-Informed Neural Networks, PINN)虽然无需依赖标注数据,但其优化可能对损失盆地高度敏感:当控制方程的残差存在多解时,从零开始训练的 PINN 即使达到了很小的残差,也可能收敛到物理上不正确的解。我们证明这两种失效模式可以被联合解决:一个不完美的神经算子可以提供必要的结构先验,使 PINN 进入正确的解盆地;而 PDE 残差则能将解的精度提升到超越该算子本身的水平。我们提出了一个三阶段框架:首先冻结物理信息神经算子的空间基函数,然后利用多项式延拓先验将其解分支外推到分布外的参数值,最后将所得的场蒸馏到一个全新的 PINN 中。该神经算子无需在目标参数处准确预测,其作用是传递解分支的信息,而 PINN 中 PDE 残差的最小化则主导收敛过程。我们在三个非线性 PDE 上评估了该框架:一维粘性 Burgers 方程、分叉点附近的二维稳态 Allen-Cahn 方程,以及二维稳态顶盖驱动方腔流。对于 Allen-Cahn 方程,由于平凡解能精确满足 PDE 残差,标准 PINN 会坍缩到平凡的零解分支;而从粗糙外推算子进行蒸馏则能恢复出与有限差分参考解一致的非平凡分支。对于顶盖驱动方腔流,外推到雷诺数 Re = 3200 能够加速收敛到正确的物理状态,在参数量和优化步数少于近期文献基线的情况下取得了具有竞争力的精度。这些结果确立了一个简单的原则:神经算子无需准确预测解就能发挥作用;它只需识别出正确的盆地,PINN 的优化便可从中恢复出正确的解。
cs.LG / 238 / 2609.24969

Rare Event Estimation via Iterative Unalignment

基于迭代去对齐的稀有事件估计
Yang, Hanming, Mittal, Daksh, Dong, Jing, Namkoong, Hongseok
Abstract
As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events can occur, but on how often they might. We study the problem of estimating the probability of rare events that arise from stochastic variation in the agent's own actions. Estimating this type of risk requires searching over the combinatorially vast space of trajectories. Naive Monte Carlo is computationally prohibitive in this regime, and constructing effective importance sampling (IS) proposals requires coordinated changes to a context-dependent chain of conditional distributions. We develop a new IS method that perturbs the original model's weights to construct the proposal. The proposal is itself a differentiably parameterized language model, enabling gradient-based search over weight space. We formulate an objective that combines a differentiable surrogate for event amplification and an adaptive regularization scheme that dynamically balances amplification against estimator stability. We evaluate our approach on $\sim$120M and $\sim$2.6B models across three event families spanning 300+ rare events as rare as $10^{-9}$, with reference probabilities computed with $<10\%$ relative standard error. In our most verifiable settings, we observe that our IS estimator achieves over $800\times$ compute-weighted efficiency gains over naive Monte Carlo for events with probabilities lower than $10^{-7}$. Our implementation is available at https://github.com/namkoong-lab/iterative-unalignment.
Chinese Translation
随着智能体以越来越高的自主性被部署,其随机输出轨迹中即使极其罕见的事件也可能发生并造成灾难性后果。因此,安全部署的关键不在于这些事件是否可能发生,而在于它们发生的频率。本文研究估计由智能体自身动作的随机变化所引发的稀有事件概率的问题。估计这类风险需要在组合爆炸式庞大的轨迹空间中进行搜索。在这种情形下,朴素蒙特卡洛方法的计算代价过于高昂,而构建有效的 importance sampling (IS) 提议分布则需要对一个依赖于上下文的条件分布链进行协同修改。我们提出了一种新的 IS 方法,通过扰动原始模型的权重来构建提议分布。该提议分布本身是一个可微参数化的语言模型,从而支持在权重空间中进行基于梯度的搜索。我们设计了一个目标函数,它结合了事件放大的可微代理量,以及一种自适应正则化方案,用于在放大效果与估计器稳定性之间进行动态平衡。我们在约 1.2 亿(120M)和约 26 亿(2.6B)参数规模的模型上评估了我们的方法,涵盖三个事件族共 300 多个稀有度高达 10^{-9} 的稀有事件,并用相对标准误差小于 10% 的参考概率进行验证。在我们最可验证的设置中,对于概率低于 10^{-7} 的事件,我们的 IS 估计器相比朴素蒙特卡洛方法获得了超过 800 倍的计算加权效率提升。我们的实现已发布于 https://github.com/namkoong-lab/iterative-unalignment。
cs.LG / 239 / 2609.24972

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

RRSI:智能体框架的正则化递归自我改进
Xia, Peng, Han, Rujun, Wang, Zifeng, Chen, Yanfei, Zhang, Yufan, Lee, Yoonho, Huang, Chengsong, Yu, Han, CuiZhu, Zhongying, Ming, Yifei, Yao, Huaxiu, Gokturk, Burak, Pfister, Tomas, Lee, Chen-Yu
Abstract
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.
Chinese Translation
大语言模型(LLM)智能体的能力在很大程度上由其框架(harness)放大,即围绕冻结的骨干模型的提示词、控制流、工具、记忆和上下文管理。近期的方法越来越多地通过迭代式地提出和选择对智能体框架的组件级修改来自动化这一过程,实际上在智能体系统层面建立了一种递归自我改进(RSI)形式。然而,这种递归演化可能通过记忆训练任务而产生过拟合,在分布内任务上表现出显著提升,但在分布外基准上的提升会缩小甚至消失。我们提出了智能体框架的正则化递归自我改进(RRSI),通过约束演化候选的提出与选择,将正则化原则融入框架自我改进。提出者(proposer)在随时间退火的预算下运行,限制候选方案可捆绑的修改数量,并基于演化历史鼓励探索未被尝试的轨迹。选择者(selector)配备了评估器和剪枝器:评估器筛选针对特定基准的提案,而剪枝器则移除过小、代价过高或不再有用的修改。这些约束共同促使演化偏好可复用的智能体机制,而非针对特定基准的方案甚至噪声。在涵盖编程、智能体工作区和工程设计任务的八个基准上,RRSI 在其演化的任务划分上最高提升 14.1 分,在五个分布外基准上最高提升 4.7 分,同时所产出的框架比未正则化的演化少使用 30% 的策略令牌。代码已发布于 https://github.com/google-research/rrsi,项目页面为 https://regularized-rsi.com/。
cs.LG / 240 / 2609.24979

LoRA-generating hypernetworks for efficient on-device LLM generative personalization

用于高效端侧大语言模型生成式个性化的LoRA生成超网络
Augenstein, Sean, Ding, Li, Lee, Jihwan, Rush, Keith, Zhmoginov, Andrey
Abstract
On-device large language models (`LLMs'), e.g. running on mobile phones, are ripe for improvement via personalization. The limited compute resources of mobile devices impose limits on model scale and thus model quality, making any realizable quality gains highly impactful. At the same time, their personal nature (i.e., the close coupling to a particular user) means that a given on-device LLM tends to be used in similar, predictable patterns over the course of time. This paper presents a novel method for personalizing on-device LLMs. It trains a hypernetwork to map a user's context tokens to a low-rank adaptation (`LoRA') well-suited to that user. Once the trained common artifacts are deployed to users' devices, each user uses the hypernetwork to synthesize (entirely on device) a personalized LoRA. This approach blends the benefits while avoiding the drawbacks of two existing approaches to LLM customization: in-context learning (`ICL') and parameter-efficient fine-tuning (`PEFT'). Like ICL (and unlike PEFT), the on-device phase of our approach is computationally feasible, requiring only forward passes through neural networks. Like PEFT (and unlike ICL), our approach modifies the `target' base LLM via weights (the LoRA), avoiding negative consequences (e.g. increased latency) associated with extending the input sequence. Our approach is particularly well-suited to the mobile device regime. Apart from the on-device compute and latency benefits mentioned, it also requires minimal additional storage, as internally its architecture partly leverages the same LLM weights as belong to the target LLM to be personalized. We demonstrate the benefits of LoRA-generating hypernetworks on several representative personalization datasets, comparing against baselines like ICL and PEFT. Of note, our personalization experiments focus on more challenging and less studied long-form text generation tasks.
Chinese Translation
端侧大语言模型(LLM),例如运行在手机上的模型,非常适合通过个性化进行改进。移动设备有限的计算资源限制了模型规模,从而限制了模型质量,因此任何可实现的质量提升都具有很高的价值。同时,端侧模型的个人属性(即与特定用户的紧密耦合)意味着给定的端侧LLM往往会随时间以相似、可预测的方式被使用。本文提出了一种个性化端侧LLM的新方法:训练一个超网络(hypernetwork),将用户的上下文词元映射为适合该用户的低秩适配(LoRA)。在将训练好的公共组件部署到用户设备后,每个用户可使用该超网络(完全在设备上)合成个性化的LoRA。该方法融合了两种现有LLM定制方法——上下文学习(ICL)和参数高效微调(PEFT)——的优点,同时规避了二者的缺点。与ICL类似(而与PEFT不同),我们方法的端侧阶段在计算上可行,只需对神经网络进行前向传播;与PEFT类似(而与ICL不同),我们的方法通过权重(即LoRA)修改目标基座LLM,避免了扩展输入序列带来的负面影响(如延迟增加)。我们的方法尤其适合移动设备场景:除了上述端侧计算和延迟方面的优势外,它还只需极少的额外存储空间,因为其内部架构部分复用了目标LLM自身的权重。我们在多个具有代表性的个性化数据集上验证了LoRA生成超网络的优越性,并与ICL和PEFT等基线方法进行了比较。值得注意的是,我们的个性化实验聚焦于更具挑战性且研究较少的长文本生成任务。
cs.LG / 241 / 2609.24985

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

临界状态强化学习:诊断多轮工具使用中可训练的状态
Chen, Zixiang, Zhao, Wenting, Cen, Zhepeng, Prabhakar, Akshara, Qiu, Jielin, Zhang, Jianguo, Liu, Zhiwei, Awalgaonkar, Tulika Manoj, Yang, Liangwei, Heinecke, Shelby, Savarese, Silvio, Wang, Huan
Abstract
Multi-turn tool-use failures can hinge on a single model call, yet reward variation alone does not reveal which call would benefit from training. When rewards depend on later interactions, their variation can reflect downstream randomness rather than differences between the current actions. We introduce Critical-State RL to identify trainable states in multi-turn interactions. Given task-defined candidate calls and local rewards, the method assesses whether each reward captures the action's effect on task success and whether improvement over a reference policy is possible. It then uses nested sampling to separate action-dependent reward variation from continuation noise and optimizes the policy at the selected states using contextual-bandit training. Experiments on the Berkeley Function Calling Leaderboard (BFCL) v4 compare training at diagnostic-selected states with training at alternative states. For missing-function tasks, the diagnostic selects the response after the tool becomes available; for missing-argument tasks, it selects the response before the missing argument is supplied. Training the selected responses improves performance, including about 14 percentage points on the missing-function task, while training the alternatives leaves performance flat or worse. We further apply the recipe across models and tasks, including logged repeat-call avoidance and memory management.
Chinese Translation
多轮工具使用的失败往往取决于单次模型调用,然而仅凭奖励的变异性无法揭示哪一次调用能够从训练中获益。当奖励依赖于后续交互时,其变异性可能反映的是下游的随机性,而非当前动作之间的差异。我们提出临界状态强化学习(Critical-State RL)来识别多轮交互中可训练的状态。给定任务定义的候选调用和局部奖励,该方法评估每个奖励是否捕捉了动作对任务成功的影响,以及相对于参考策略是否存在改进空间。随后,该方法利用嵌套采样(nested sampling)将依赖于动作的奖励变异性与延续噪声分离,并通过情境老虎机(contextual-bandit)训练在所选状态上优化策略。在伯克利函数调用排行榜(Berkeley Function Calling Leaderboard, BFCL)v4上的实验将诊断所选状态的训练与备选状态的训练进行了比较。对于缺失函数任务,诊断方法选择在工具变为可用之后的那次响应;对于缺失参数任务,诊断方法选择在缺失参数被提供之前的那次响应。训练所选响应能够提升性能,例如在缺失函数任务上提升约14个百分点,而训练备选响应则使性能持平或更差。我们进一步将该方案应用于多种模型和任务,包括基于日志的重复调用规避和内存管理。