← Back to Index
Daily Research Digest

arXiv Papers

2026-09-10
172
Papers
3
Categories
172
Translated
收藏清单 0
人工智能 (Artificial Intelligence)
51
cs.AI / 1 / 2609.09203

OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows

OpenDiscoveryTrace:用于评估AI科学家工作流的过程轨迹
Bansal, Aayam, Balaji, Keertan
Abstract
Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing. We present \textbf{OpenDiscoveryTrace}, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace---including thoughts, tool calls, observations, errors, revision triggers, and self-reported confidence---as models execute 124 scientific tasks spanning drug discovery, materials science, genomics, and scientific literature analysis. The dataset covers seven models: three frontier models (GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro; 124 trajectories each, fully balanced across domains and difficulty levels) and four open-weight models (Qwen2.5-7B, Mistral-7B-v0.3, Phi-3.5-mini, and Qwen2.5-1.5B; 30 each), plus 60 live-retrieval variant trajectories. Pilot analysis on 363 LLM-judged trajectories reveals that process traces expose behavioral differences invisible to output-only evaluation: all three frontier models achieve comparable success rates (84--89%), yet Claude Opus 4.6 produces 30$\times$ more errors than GPT-5.4 (2.5 vs. 0.08 per trajectory, $p < 0.0001$, Cliff's $\delta = 0.613$), with qualitatively different error profiles---66.7% tool misuse for Claude versus 83.6% reasoning errors for GPT-5.4. We define five benchmark tasks with baselines from logistic regression, random forests, LSTMs, and Transformer models. The dataset, trace schema, agent harness, and benchmark definitions are publicly available under CC BY 4.0 to support research on process-level evaluation, scientific agent auditing, and AI governance.
Chinese Translation
现有的自主AI科学家基准测试仅评估最终输出——生成的代码、假设或论文——却丢弃了获得这些输出的推理过程。这使得人们无法审计科学方法论、诊断失败模式,或区分系统性推理与侥幸猜测。我们提出了OpenDiscoveryTrace,一个包含558条完整AI科学智能体轨迹的公开数据集,它捕捉的是模型如何推理,而不仅仅是它们产出了什么。每条轨迹记录了一个结构化的每步9字段轨迹——包括思考、工具调用、观察结果、错误、修订触发器和自我报告的置信度——这些轨迹来自模型执行124项科学任务的过程,任务涵盖药物发现、材料科学、基因组学和科学文献分析。该数据集涵盖七个模型:三个前沿模型(GPT-5.4、Claude Opus 4.6和Gemini 3.1 Pro,各124条轨迹,在领域和难度级别上完全均衡)以及四个开放权重模型(Qwen2.5-7B、Mistral-7B-v0.3、Phi-3.5-mini和Qwen2.5-1.5B,各30条),外加60条实时检索变体轨迹。对363条经LLM评判的轨迹的初步分析表明,过程轨迹能够揭示仅评估输出时不可见的行为差异:三个前沿模型的成功率相当(84%–89%),但Claude Opus 4.6产生的错误是GPT-5.4的30倍(每条轨迹2.5对0.08,p < 0.0001,Cliff's δ = 0.613),且错误特征在性质上截然不同——Claude的错误中66.7%为工具误用,而GPT-5.4的错误中83.6%为推理错误。我们定义了五个基准任务,并提供逻辑回归、随机森林、LSTM和Transformer模型的基线结果。该数据集、轨迹模式、智能体运行框架和基准定义均在CC BY 4.0许可下公开可用,以支持过程级评估、科学智能体审计和AI治理方面的研究。
cs.AI / 2 / 2609.09226

Adaptive Entangled Game Modules in Artificial General Intelligence

通用人工智能中的自适应纠缠博弈模块
Li, Haochen, Guo, Xinshuai, Ouyang, Jingdong, Zhang, Wei, Shi, Leilei
Abstract
We introduce a probability-wave framework for modeling the collective behavior of interacting adaptive agents, deriving testable eigenmodes through a generalized behavioral intelligence (GBI) nonlocal probability-wave equation. This framework captures a broad range of human intelligence behaviors with analytical mechanisms and offers an indirect method to examine the Liu-Chen-Ao (LCA) hypothesis of nonlocal entangled nerve fibers in the brain through collective trader behaviors. Our empirical analysis of Chinese intraday stock market data demonstrates that adaptive entangled game modes explain 82-94% (89% overall) of observed decision patterns, a sharp contrast to the predictions of neoclassical finance based on independent rational agents. Moreover, 2-12% of behaviors show adaption to intraday news, events, and environments, characterized by dual equilibrium states and abrupt reference point shifts, while purely independent modes occur in less than 5% of cases. These findings empirically support the LCA hypothesis, as observable trading behaviors reflect underlying brain mechanisms and internal intelligence decision-making in behavioral psychology. Our results highlight the necessity of incorporating adaptive entangled game modules into artificial general intelligence (AGI) architectures, addressing the limitations of conventional artificial neural network (ANN)-based AI, which relies on trillions of opaque parameters. By integrating ANN-based AI with probability-wave-based entangled-brain simulations, machine learning can enrich AGI foundation models (FMs) and facilitate the development of human-like processing units (HPUs) that leverage brain-inspired mechanisms. Such HPUs may ultimately create more compact, efficient, and robust AGI systems, particularly for embodied intelligence and robotics.
Chinese Translation
我们引入了一个概率波框架,用于建模相互作用的适应性智能体的集体行为,并通过广义行为智能(GBI)非局域概率波方程推导出可检验的本征模态。该框架以解析机制刻画了广泛的人类智能行为,并提供了一种间接方法,通过交易者的集体行为来检验大脑中非局域纠缠神经纤维的Liu-Chen-Ao(LCA)假说。我们对中国日内股票市场数据的实证分析表明,自适应纠缠博弈模态能够解释82%–94%(总体为89%)的观测决策模式,这与基于独立理性智能体假设的新古典金融学的预测形成鲜明对比。此外,2%–12%的行为表现出对日内新闻、事件和环境的适应,其特征为双均衡状态和参考点的突变;而纯粹独立的模式仅在不到5%的情形中出现。这些发现为LCA假说提供了实证支持,因为可观测的交易行为反映了潜在的大脑机制以及行为心理学中的内部智能决策。我们的结果凸显了将自适应纠缠博弈模块纳入通用人工智能(AGI)架构的必要性,以解决依赖数万亿不透明参数的、基于人工神经网络(ANN)的传统人工智能的局限性。通过将基于ANN的人工智能与基于概率波的纠缠大脑模拟相结合,机器学习可以丰富AGI基础模型(FM),并促进利用类脑机制的人式处理单元(HPU)的发展。此类HPU最终可能催生更紧凑、高效且稳健的AGI系统,尤其适用于具身智能与机器人领域。
cs.AI / 3 / 2609.09233

Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks

子代理与代理技能:面向长时程智能体任务的 reusable 知识执行方法
Piriyakulkij, Wasu Top, Lawrence, Rachel, Curth, Alicia, Karmalkar, Sushrut, Prasad, Niranjani
Abstract
How can language model agents effectively leverage libraries of reusable knowledge to solve long-horizon tasks? Recent work has increasingly focused on agent skills: reusable capabilities represented as skill packages, i.e., multi-file bundles containing instructions, scripts, and other resources that help agents perform specific tasks. Agent skills are typically executed by loading their skill instructions into an agent's context and relying on the agent to follow them. As task horizons grow, however, this approach becomes increasingly brittle, because reasoning quality degrades as more information accumulates in the context window. We investigate an alternative approach in which skill packages are instead invoked as subagents. Rather than loading skill instructions into the main context, subagent execution spawns fresh context windows dedicated to solving individual subtasks. We show that subagent execution outperforms agent-skill execution when skill packages expose clear input-output contracts and their instructions encode the procedural knowledge needed to fulfill those contracts. The tradeoff is additional communication overhead, as extra tokens are required to coordinate between the main agent and its subagents. Our results show that the benefit of reusable knowledge depends not only on its content, but also on how it is organized and invoked.
Chinese Translation
语言模型智能体如何有效利用可复用知识库来解决长时程任务?近期研究日益聚焦于代理技能:即以技能包形式表示的可复用能力,技能包是包含指令、脚本及其他资源的多文件集合,可帮助智能体完成特定任务。代理技能通常通过将技能指令加载到智能体的上下文中、并依赖智能体遵循这些指令来执行。然而,随着任务时程的增长,这一方法变得越来越脆弱,因为随着上下文窗口中信息量的累积,推理质量会下降。我们研究了一种替代方法,即将技能包作为子代理调用。子代理执行并非将技能指令加载到主上下文中,而是生成专用于解决各个子任务的新上下文窗口。我们表明,当技能包具有清晰的输入输出契约,且其指令编码了履行这些契约所需的过程性知识时,子代理执行优于代理技能执行。其代价是额外的通信开销,因为主代理与子代理之间的协调需要消耗额外的 token。我们的结果表明,可复用知识的收益不仅取决于其内容,还取决于其组织和调用方式。
cs.AI / 4 / 2609.09306

Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions

Gradland:论跨多维分化的现象体验
Balduzzi, David
Abstract
This paper investigates the hypothesis that the first-order structure of physical interactions, i.e. gradients or Jacobians, characterizes the structure of phenomenal experience. It does so in an idealized world inhabited by neural networks, Gradland, where the physics are known and the functions are (mostly) differentiable. The paper introduces two measures of Jacobian structure: effective rank and cohesion, based on Kirchhoff complexity. Applying the measures to a series of worked examples shows the hypothesis accounts for: (1) the duration of experience, that it can prolong over hundreds of milliseconds; (2) the difference between what is experienced vividly and obscurely; (3) the experience of texture; (4) the blooming buzzing confusion presumably experienced by newborns; (5) the difference between ideas that are held distinctly in mind and ideas that are confused; (6) what learning is like; and finally (7) the paper explains the function of rich, dense experience.
Chinese Translation
本文探究了这样一个假设:物理相互作用的一阶结构,即梯度或雅可比矩阵(Jacobian),刻画了现象体验的结构。研究在一个由神经网络构成的理想化世界——Gradland 中进行,该世界的物理规律已知,且函数(大多)可微。论文基于基尔霍夫复杂度(Kirchhoff complexity)引入了两种雅可比结构度量:有效秩(effective rank)和凝聚度(cohesion)。将这两种度量应用于一系列具体实例表明,该假设能够解释:(1)体验的持续时间,即体验可以延续数百毫秒;(2)被鲜明体验与模糊体验之间的差异;(3)质感的体验;(4)新生儿所经历的所谓“繁花嘈杂般”的混沌感知;(5)头脑中清晰的观念与混乱的观念之间的差异;(6)学习过程的体验;最后,(7)本文解释了丰富、密集体验的功能。
cs.AI / 5 / 2609.09374

An Autonomous GeoAI Agent for Arctic Eco-Navigation

面向北极生态航行的自主GeoAI智能体
Taleghan, Samira Alkaee, Koo, Younghyun, Banaei-Kashani, Farnoush
Abstract
Arctic maritime navigation is becoming increasingly important as changing sea-ice conditions expand seasonal accessibility while simultaneously introducing substantial operational, environmental, and community risks. Arctic route planning is inherently a multi-criteria problem: routes that improve vessel safety or efficiency may increase exposure to sea ice, sensitive ecosystems, or nearby communities. Existing routing methods prioritize travel time, fuel use, and navigational risk, often overlooking ecological and community impacts. We introduce a human-in-the-loop, multi-agent GeoAI system for Arctic eco-navigation that integrates operational, physical, ecological, and community-related criteria within a unified routing framework. Multiple specialized agents coordinate geospatial data acquisition and preparation, multi-objective route generation, and skyline-based decision support. The ecological criteria explicitly account for exposure to sensitive areas, including Essential Fish Habitat and seal critical habitat. By considering these ecosystem impacts and potential community burdens while keeping consequential value judgments under human control, the framework supports safer, more transparent, and socially responsible Arctic navigation. Project page and code are publicly available. https://samiraat.github.io/Arctic-Eco-Navigation-Agent/, https://github.com/samiraat/Arctic-Eco-Navigation-Agent
Chinese Translation
随着海冰状况的变化在扩大季节性通航范围的同时也带来了重大的运营、环境和社区风险,北极海事航行正变得日益重要。北极航线规划本质上是一个多准则问题:提升船舶安全性或效率的航线可能会增加对海冰、敏感生态系统或周边社区的暴露风险。现有航线规划方法主要优先考虑航行时间、燃料消耗和航行风险,往往忽视了生态和社区影响。我们提出了一种人在回路的多智能体GeoAI系统,用于北极生态航行,该系统在统一的航线规划框架内整合了运营、物理、生态和社区相关准则。多个专业化智能体协同完成地理空间数据的获取与准备、多目标航线生成以及基于天际线的决策支持。生态准则明确考虑了对敏感区域的暴露,包括重要鱼类栖息地和海豹关键栖息地。通过在将关键价值判断保留由人类控制的同时考虑这些生态系统影响和潜在的社区负担,该框架支持更安全、更透明且更具社会责任感的北极航行。项目页面和代码已公开:https://samiraat.github.io/Arctic-Eco-Navigation-Agent/, https://github.com/samiraat/Arctic-Eco-Navigation-Agent
cs.AI / 6 / 2609.09395

The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents

菜单即执行先验:面向在线智能体的状态路径工具菜单
Yan, Bo, Lin, Weikai, Wang, Song
Abstract
Language models act through tools, yet practical agents face libraries containing thousands of interfaces. We introduce the tool menu as the short, ordered subset of available tools shown to an agent before execution. The agent can call only tools in this menu. Multi-step tasks require the final action and the prerequisite tools that create its inputs in a usable order. Current constructors rank tools by request relevance, which can surface the final action while omitting or delaying less obvious producers. We introduce the state path, a pre-execution route from the observable request state to the desired outcome, and propose State-Path Tool Menu to learn it. Our framework treats the menu as an execution prior over these routes. Its encoder represents which tools can run from the current state, how their outputs satisfy later inputs, and which orders recur in training paths. A retriever covers an executable entry, the missing-input producers, and the final action. A reranker then places producers before consumers. On ToolBench, our menu raises online success from 0.737 to 0.898 and outperforms retrieval, reranking, generation, and routing baselines without changing the agent. The State-Path menu also covers more complete chains with 32 tools than the official list covers with 128, and its success gain persists across executor families with different model capacities. Our code is at https://github.com/Met2348/State-Path.
Chinese Translation
语言模型通过工具来执行动作,但实际中的智能体面临着包含数千个接口的工具库。我们提出了工具菜单(tool menu)的概念,即在执行前展示给智能体的一个简短的、有序的可用工具子集,智能体只能调用该菜单中的工具。多步任务需要最终动作以及能够按可用顺序生成其输入的前置工具。目前的菜单构建方法根据与请求的相关性对工具进行排序,这可能会呈现最终动作,却遗漏或延迟了不太显眼的生产者工具。我们提出了状态路径(state path),即从可观测的请求状态到期望结果的预执行路线,并提出状态路径工具菜单(State-Path Tool Menu)来学习它。我们的框架将菜单视为这些路线上的执行先验。其编码器表征了哪些工具可以从当前状态运行、它们的输出如何满足后续输入,以及哪些顺序在训练路径中反复出现。检索器覆盖一个可执行的入口工具、缺失输入的生产者工具以及最终动作,然后重排器将生产者工具置于消费者工具之前。在 ToolBench 上,我们的菜单将在线成功率从 0.737 提升至 0.898,并在不改变智能体的情况下优于检索、重排、生成和路由等基线方法。状态路径菜单仅用 32 个工具就能覆盖比官方列表用 128 个工具更完整的工具链,且其成功率提升在不同模型容量的执行器家族中均保持稳定。我们的代码见 https://github.com/Met2348/State-Path。
cs.AI / 7 / 2609.09413

Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery

面向规模感知关键材料回收的决策导向主动学习
Srinivas, Niranjan, Ray, Debajyoti, Nakouzi, Elias
Abstract
Choosing a recovery process for scale-up requires connecting laboratory results with product requirements, process costs, and scale effects. We analyze records from Pacific Northwest National Laboratory's Computer Intelligence for Critical Element Recovery and Optimization (CICERO) workflow for autonomous selective precipitation. Active learning uses prior results to choose experiments. In a conditional retrospective benchmark with fitted models and recycled neodymium-iron-boron (NdFeB) magnet records, active learning finds the best recorded result with fewer experiments than nonadaptive space filling. Enrichment is the selected rare-earth-to-iron ratio relative to that in the feed. Adaptive policies reach the recorded enrichment maximum by 16 to 24 wells (individual experiments), versus 48. Our two-stage reconstruction ties two adaptive alternatives at 16 wells. Conditional analyses of recycled samarium-cobalt (SmCo) magnets show a Round 2 tradeoff between purity and nominal yield, the recovery fraction calculated from an assumed starting amount - NdFeB Round 1 routes differ in enrichment. Rankings for produced water from oil and gas extraction depend on phase and dilution assumptions requiring confirmation. We propose choosing batches by their expected reduction in downstream Bayes risk: the minimum expected loss among available process decisions under current beliefs. In exploratory simulations, a hybrid that filters candidates has lower estimated loss than the implemented joint search across routes and conditions. Differences involving the synthetic two-stage policy are small relative to estimation uncertainty. We outline a pre-registered prospective test under a shared loss and logging standard, requiring clarified measurements and records, a defined process decision and relevant outputs, credible economic inputs, and validation at the intended scale.
Chinese Translation
为放大生产选择回收工艺,需要将实验室结果与产品需求、工艺成本和规模效应联系起来。我们分析了太平洋西北国家实验室(Pacific Northwest National Laboratory)的CICERO(Computer Intelligence for Critical Element Recovery and Optimization,关键元素回收与优化计算智能)自主选择性沉淀工作流程的记录。主动学习利用先前的结果来选择实验。在一个基于拟合模型和回收钕铁硼(NdFeB)磁体记录的条件性回顾基准测试中,主动学习比非自适应的空间填充方法用更少的实验即可找到记录中的最优结果。富集度(Enrichment)是指所选稀土与铁的比值相对于进料中该比值的偏离程度。自适应策略在16至24个孔(即单次实验)内达到记录的富集度最大值,而非自适应方法需要48个。我们的两阶段重构使两种自适应替代方案均在16个孔时持平。对回收钐钴(SmCo)磁体的条件性分析显示,第二轮在纯度与名义产率(基于假设起始量计算的回收比例)之间存在权衡,而NdFeB第一轮各路线的差异体现在富集度上。针对油气开采产出水的排序结果则依赖于需要进一步确认的相态与稀释假设。我们提出依据下游贝叶斯风险的预期降低量来选择批次:即在当前信念下,各可用工艺决策中的最小预期损失。在探索性模拟中,采用候选过滤的混合方法比已实施的跨路线与条件的联合搜索具有更低的估计损失。涉及合成两阶段策略的差异相对于估计不确定性而言很小。我们概述了一项在统一损失与记录标准下的预注册前瞻性测试,该测试要求明确测量与记录规范、界定工艺决策及相关输出、提供可信的经济输入,并在预期规模下进行验证。
cs.AI / 8 / 2609.09418

Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration

Valerant:一种基于动作条件世界模型探索的自动可导航游戏地图生成器
Qiao, Yiran, Wang, Feng, Ma, Jing
Abstract
World Action Models (WAMs) couple predictive world modeling with action generation, allowing anticipated future states to guide agent behavior. Although WAMs are rapidly advancing embodied AI, general-purpose counterparts remain largely unexplored in games. Existing game-oriented approaches often combine action-conditioned world models with external policies and reward functions to realize WAM-like decision-making, yet they operate mainly in 2D visual observation space and do not instantiate persistent 3D geometry. Extending this paradigm to 3D games introduces a distinct challenge. In autonomous driving and robotics, the physical environment exists independently of the model, providing a persistent 3D world in which selected actions can be executed. Games have no such external substrate; the virtual world itself must be instantiated. Most playable games require a persistent and navigable space, while 3D games additionally require explicit geometry that supports movement and interaction. Action-conditioned video rollouts provide visual observations but not this spatial representation. We present \textsc{Valerant}, a training-free framework that transforms a pretrained action-conditioned world model into a WAM for exploring and constructing 3D game maps. By coupling predictive visual rollouts with SLAM-based spatial reconstruction and exploration-driven action selection, \textsc{Valerant} progressively transforms a single image into a persistent 3D game map. This framework extends WAM-based interaction beyond 2D visual simulation and offers a new approach to reducing manual effort in 3D game-map creation.
Chinese Translation
世界动作模型(World Action Models, WAMs)将预测性世界建模与动作生成相结合,使预期的未来状态能够引导智能体行为。尽管 WAM 正在迅速推动具身智能(Embodied AI)的发展,但通用型的对应方法在游戏领域仍基本未被探索。现有的面向游戏的方法通常将动作条件世界模型与外部策略和奖励函数相结合以实现类 WAM 的决策,但它们主要运行在二维视觉观测空间中,并未实例化持久的 3D 几何结构。将该范式扩展到 3D 游戏带来了独特的挑战。在自动驾驶和机器人领域,物理环境独立于模型而存在,提供了一个持久的三维世界,其中可以选择并执行动作。而游戏没有这样的外部基底,虚拟世界本身必须被实例化。大多数可玩的游戏需要一个持久且可导航的空间,而 3D 游戏还需要支持移动和交互的显式几何结构。动作条件的视频推演可以提供视觉观测,但无法提供这种空间表示。我们提出了 Valerant,一个无需训练的框架,它将预训练的动作条件世界模型转化为用于探索和构建 3D 游戏地图的 WAM。通过将预测性视觉推演与基于 SLAM 的空间重建以及探索驱动的动作选择相结合,Valerant 能够逐步将单张图像转化为持久的 3D 游戏地图。该框架将基于 WAM 的交互从二维视觉模拟扩展到三维,并为减少 3D 游戏地图创建中的人工投入提供了一种新方法。
cs.AI / 9 / 2609.09428

XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

XAI-Arena:大语言模型能否评估可解释人工智能(XAI)解释的质量?
Fleischhauer, Yanfei Hu, Zharova, Alona, Klein, Nadja, Feuerriegel, Stefan
Abstract
Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations. We introduce XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, multidimensional, and stakeholder-sensitive evaluation of XAI explanation quality. XAI-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability. We then benchmark XAI explanation methods across various datasets, machine learning models, and stakeholder personas. Human validation shows a strong positive association between LLM-generated and human ratings (Spearman's rho=.693, p<.001). Together, LLM-based evaluations can capture systematic differences in XAI explanation quality and provide a scalable and reproducible framework for comparative assessment of XAI explanations.
Chinese Translation
评估可解释人工智能(XAI)方法所产生的解释的质量仍然具有挑战性,因为现有方法往往依赖主观的人类判断,从而限制了可重复性、可扩展性以及研究之间的可比性。我们研究了大语言模型(LLM)能否作为一种可重复、可扩展的机制,对XAI解释的质量进行比较评估。我们提出了XAI-Arena,这是一个基于“LLM作为评判者”(LLM-as-a-judge)的框架,用于对XAI解释质量进行可扩展、可重复、多维度且兼顾利益相关者的评估。XAI-Arena使我们能够从多个维度比较XAI解释,包括感知简洁性、清晰度、任务适配性、信任校准、可操作性、透明度、忠实度以及整体可解释性。随后,我们在多种数据集、机器学习模型和利益相关者角色上对XAI解释方法进行了基准测试。人类验证结果表明,LLM生成的评分与人类评分之间存在强烈的正相关关系(Spearman's rho=.693,p<.001)。总之,基于LLM的评估能够捕捉XAI解释质量中的系统性差异,并为XAI解释的比较评估提供了一个可扩展、可重复的框架。
cs.AI / 10 / 2609.09448

Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

智能体知道何时成功吗?基于内部表征的智能体置信度校准
Mammen, Priyanka Mary, Joswin, Emil, Medicherla, Srujananjali
Abstract
As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, tool invocation and dynamic environment interactions. In this paper, we investigate whether model's internal representations provide stronger signals of eventual task success in multi-turn agentic setups. We introduce two complementary methods: Latent Trajectory Dynamics (LTD), which summarizes changes in residual-stream representations across an an interaction trajectory, and the Action Representation Probe (ARP), which predicts success from representations formed at action decisions. Across three interactive benchmarks (Bash, SQL, Python) and three model families (Qwen14B, Qwen7B, DeepSeek6.7B), our methods consistently outperform surface level generation and sequence-based calibration baselines providing a zero-overhead reliability monitor that requires neither prompt alterations nor multi-sample rollouts.
Chinese Translation
随着智能体系统在安全关键型应用中的快速普及,衡量与智能体行为相关联的置信度变得至关重要。与传统机器学习系统相比,智能体工作流具有更复杂的失败模式,涉及规划、工具调用以及与动态环境的交互。在本文中,我们研究了模型的内部表征是否能在多轮智能体设置中为任务的最终成功提供更强的信号。我们提出了两种互补的方法:潜在轨迹动力学(Latent Trajectory Dynamics, LTD),它总结了残差流表征在交互轨迹中的变化;以及动作表征探针(Action Representation Probe, ARP),它根据动作决策时刻形成的表征来预测任务成功。在三个交互式基准(Bash、SQL、Python)和三个模型系列(Qwen14B、Qwen7B、DeepSeek6.7B)上的实验表明,我们的方法持续优于表层的生成结果和基于序列的校准基线方法,提供了一种零开销的可靠性监控手段,既不需要修改提示词,也不需要多样本回滚采样。
cs.AI / 11 / 2609.09458

ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance

ContractEval:面向程序性指令符合性的查询条件执行匹配
Singh, Praphul, Kumar, Shanu, Agarwal, Akshat, Kumar, Ganesh
Abstract
As LLM agents move from answering questions to carrying out procedures, failures can be unwarranted rather than visibly wrong: the final response looks acceptable even though the system skipped the check, branch, dependency, or invariant that made the answer justified. Output-only evaluation sees the answer, and trace-aware judging sees activity, but neither identifies which obligations were active for the query. We introduce CONTRACTEVAL, a diagnostic framework for making those active obligations explicit. It represents procedural instructions as query-active obligations and matches them against response or trace evidence, turning omissions, wrong branches, ordering errors, extra actions, invariant breaches, and output-contract violations into distinct conformance failures. On a controlled suite of audited procedural contracts, output-only and trace-aware LLM judges miss many injected structural failures; under gold expected and observed graphs, ContractEval detects and localizes all of them. LLM-backed extraction preserves much of this signal but remains calibration-sensitive. ContractEval is therefore not a compliance guarantee; it makes procedural conformance auditable rather than implicit in final-answer quality.
Chinese Translation
随着大语言模型(LLM)智能体从回答问题转向执行程序性任务,其失败可能是隐性的而非明显错误的:最终响应看起来可以接受,但系统实际上跳过了使该答案得以成立的检查、分支、依赖或不变式。仅基于输出的评估只能看到答案,基于轨迹的评判只能看到活动过程,但二者都无法识别对于该查询哪些义务应当被履行。我们提出了 ContractEval,一个将这些活跃义务显式化的诊断框架。它将程序性指令表示为查询活跃义务,并将其与响应或轨迹证据进行匹配,从而将遗漏、错误分支、顺序错误、多余操作、不变式违反和输出契约违规转化为不同类型的符合性失败。在一套经审计的可控程序性契约测试集上,仅输出型和轨迹感知型 LLM 评判器会遗漏许多注入的结构性失败;而在给定标准期望图与观测图的条件下,ContractEval 能够检测并定位所有这些失败。基于 LLM 的信息提取能保留大部分信号,但仍然对校准敏感。因此,ContractEval 并非合规性保证;它使程序性符合性变得可审计,而非隐含于最终答案质量之中。
cs.AI / 12 / 2609.09565

Multi-Agent Agentic Graph Learning via Structural Signatures

基于结构签名(Structural Signatures)的多智能体代理图学习
Qu, Liang, Li, Jianxin, Wang, Hua
Abstract
Agentic graph learning (AGL) has recently achieved promising results on graph reasoning tasks, where an agent powered by a large language model (LLM) sequentially samples the graph as evidence to support its final prediction. Existing methods either employ a single agent or orchestrate multiple role-based agents to reason and learn over the entire graph, but both essentially rely on a shared reasoning policy across different graph regions, which can be suboptimal for graphs with heterogeneous structural and semantic patterns. Inspired by the progress of multi-agent collaboration on complex reasoning tasks, a natural remedy is to let multiple agents own different memory and collaborate; however, applying this paradigm to graphs directly faces two challenges. First, existing AGL methods typically verbalize graph structures into natural-language descriptions for LLM agents, making the reasoning process sensitive to the ordering of structural information and thereby breaking the permutation-invariant nature of graphs. Second, incorporating increasingly large sampled neighborhoods leads to rapidly growing contexts. To address these challenges, this paper introduces a multi-agent agentic graph learning (i.e., MAAGL) framework. MAAGL partitions the graph into communities and assigns an independent agent to each community for region-specific specialization. MAAGL represents structural and semantic evidence separately. Structural evidence is summarized by a dynamically updated structural signature that is permutation-invariant and fixed in size, while semantic evidence is filtered to the top-k nodes ranked by relevance. Based on historical trajectories with similar signatures, agents estimate their confidence and trigger debate-style collaboration when needed. Extensive experiments on four benchmark datasets show that MAAGL outperforms SOTA AGL methods.
Chinese Translation
代理图学习(Agentic Graph Learning, AGL)近期在图推理任务上取得了令人瞩目的成果,其核心是由大语言模型(LLM)驱动的智能体依次对图进行采样,作为证据来支持其最终预测。现有方法要么采用单一智能体,要么编排多个基于角色的智能体在整个图上进行推理与学习,但二者本质上都依赖于跨不同图区域的共享推理策略,这对于具有异质结构和语义模式的图而言可能是次优的。受多智能体协作在复杂推理任务上取得进展的启发,一种自然的解决方案是让多个智能体拥有不同的记忆并相互协作;然而,将该范式直接应用于图面临两大挑战。首先,现有AGL方法通常将图结构转化为自然语言描述供LLM智能体使用,使得推理过程对结构信息的排列顺序敏感,从而破坏了图的置换不变性。其次,纳入规模不断增大的采样邻域会导致上下文长度快速增长。为应对这些挑战,本文提出了一个多智能体代理图学习(MAAGL)框架。MAAGL将图划分为多个社区,并为每个社区分配一个独立智能体以实现面向特定区域的专业化。MAAGL对结构证据和语义证据分别进行表示:结构证据由一种动态更新的结构签名进行总结,该签名兼具置换不变性和固定大小;语义证据则被过滤为按相关性排序的前k个节点。基于具有相似签名的历史轨迹,智能体估计自身置信度,并在需要时触发辩论式协作。在四个基准数据集上的大量实验表明,MAAGL优于当前最先进的AGL方法。
cs.AI / 13 / 2609.09578

CityPlanner: A Sandbox Agent for Executable Urban Planning

CityPlanner:面向可执行城市规划的沙盒智能体
Zhang, Wentao, Wang, Jingyuan, Zhou, Zetong, Yang, Yifan, Wang, Wenrui
Abstract
Urban planning is a real-world spatial optimization problem that requires selecting feasible actions from large candidate spaces under practical objectives such as cost and service quality. Existing optimization and reinforcement learning methods are effective for fixed formulations, but often depend on task-specific representations and constraint handling. We propose \emph{CityPlanner}, a sandbox-agent framework for executable urban planning. CityPlanner introduces \emph{UrbanSandbox}, a unified file-based environment where agents inspect task files, generate plans, run evaluators, and revise decisions based on executable feedback. To make learning tractable, we further propose atomic-task reinforcement learning, which decomposes long sandbox trajectories into \emph{BuildPlan} for initial construction and \emph{ImprovePlan} for feedback-based refinement. Experiments on a real-world benchmark show that CityPlanner consistently outperforms heuristic, task-specific RL, and general LLM-agent baselines. Ablations verify the contributions of UrbanSandbox, atomic-task RL, and iterative deployment. We release the code and dataset at https://anonymous.4open.science/r/co-agent-C1C8
Chinese Translation
城市规划是一个现实世界中的空间优化问题,需要在成本和服务质量等实际目标约束下,从庞大的候选空间中选择可行的行动方案。现有的优化方法和强化学习方法对于固定的问题形式化是有效的,但通常依赖于特定任务的表示方式和约束处理机制。我们提出了 CityPlanner,一个面向可执行城市规划的沙盒智能体框架。CityPlanner 引入了 UrbanSandbox,一个统一的基于文件的环境,智能体可以在其中查看任务文件、生成规划方案、运行评估器,并根据可执行的反馈修订决策。为了使学习过程易于处理,我们进一步提出了原子任务强化学习(atomic-task reinforcement learning),将漫长的沙盒轨迹分解为用于初始建设的 BuildPlan 和基于反馈进行优化的 ImprovePlan。在真实世界基准上的实验表明,CityPlanner 始终优于启发式方法、任务特定的强化学习方法以及通用 LLM 智能体基线。消融实验验证了 UrbanSandbox、原子任务强化学习和迭代部署的贡献。我们在 https://anonymous.4open.science/r/co-agent-C1C8 发布了代码和数据集。
cs.AI / 14 / 2609.09589

A Function-Space Approach to the Statistical Mechanics of Learning Dynamics

学习动力学的函数空间统计力学方法
Zhang, Yizhou, Wu, Weichen, Du, Lun, Miao, Zhengjie
Abstract
Deep neural networks exhibit regular macroscopic behavior despite highly nonlinear dynamics in vast parameter spaces. We develop a statistical-mechanical description of learning directly in function space, treating parameter configurations as microscopic realizations and functions with their dynamical operators as macroscopic variables. For mean-squared loss, the exact error dynamics are governed by the learning operator \(M=JJ^\ast\). Combining the dynamical Boltzmann weight of the conditional stochastic dynamics with the parameter-space density of states, whose local curvature defines a statistical operator \(B\), and integrating over local fluctuations yields $$ \Phi_{\mathrm{fluc}}(M;B)=\frac{\sigma_\xi^2}{2}\log\det(M^{-1}+B)+\mathrm{const}. $$ At fixed spectrum, this term is rotationally stationary when \([M,B]=0\), is minimized by pairing large eigenvalues of \(M\) with small eigenvalues of \(B\), and generates a local restoring contribution against rotational mismatch. For ReLU-type function spaces under mild stable statistical conditions, \(B=\sigma_\xi^2L^\ast\mathcal K L\), where \(L\) measures coarse-grained second-order structure. Thus the low-\(B\) sector corresponds, up to bounded anisotropy of \(\mathcal K\), to low structural curvature, implying a preference for faster relaxation along smooth, data-adaptive directions. These results identify function space as a natural macroscopic level for studying stable collective organization in learning.
Chinese Translation
尽管深度神经网络在巨大的参数空间中呈现高度非线性的动力学行为,但其宏观行为却表现出规律性。我们直接在函数空间中建立了学习的统计力学描述,将参数配置视为微观实现,将函数及其动力学算子视为宏观变量。对于均方损失,精确的误差动力学由学习算子 \(M=JJ^\ast\) 支配。将条件随机动力学的动力学玻尔兹曼权重与参数空间态密度相结合(其局部曲率定义了一个统计算子 \(B\)),并对局部涨落进行积分,得到 $$ \Phi_{\mathrm{fluc}}(M;B)=\frac{\sigma_\xi^2}{2}\log\det(M^{-1}+B)+\mathrm{const}. $$ 在固定谱的条件下,当 \([M,B]=0\) 时,该项在旋转意义下处于驻定状态;将 \(M\) 的大特征值与 \(B\) 的小特征值配对可使其最小化;并且它会产生一个抵抗旋转失配的局部回复贡献。对于满足温和稳定统计条件的 ReLU 型函数空间,\(B=\sigma_\xi^2L^\ast\mathcal K L\),其中 \(L\) 度量粗粒化的二阶结构。因此,在 \(\mathcal K\) 的有界各向异性范围内,低 \(B\) 区域对应于低结构曲率,这意味着系统偏好沿平滑的、数据自适应的方向进行更快的弛豫。这些结果表明,函数空间是研究学习中稳定集体组织的自然宏观层面。
cs.AI / 15 / 2609.09625

From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital Twins

从状态同步到认知自演化:认知数字孪生的操作性架构
Gao, Haoran, Li, An, Li, Zhen, Cai, Jun
Abstract
As Digital Twin (DT) systems evolve beyond state synchronization toward task-oriented and knowledge-driven operation, Cognitive Digital Twins (CDTs) have emerged as an extension that incorporates cognitive capabilities into twin operation. Existing CDT studies often focus on specific enabling techniques, such as learning modules, knowledge graphs, and large language models, while providing limited insight into how cognition can be systematically integrated into DT architectures. To address this issue, this paper proposes a four-layer CDT architecture consisting of the physical layer, digital-twin layer, cognitive layer, and task layer. The proposed architecture establishes a self-evolving closed operational loop spanning these four layers, in which physical states are synchronized into digital representations, cognition constructs task-specific cognitive models through knowledge, memory, and attention, and task-level decisions are generated under practical constraints. Operational feedback further refines cognitive experience and updates relationships and annotations in the digital representation, enabling subsequent task interpretation, initiation, and reasoning to evolve with system operation. Based on this framework, two representative operation modes are characterized: user-request-driven cognition and self-driven cognition. We further discuss key enabling mechanisms and deployment challenges associated with semantic communication, knowledge querying, task orchestration, and closed-loop synchronization. A lightweight simulation study illustrates reliable closed-loop task feasibility under limited semantic information and improved operational efficiency through accumulated task experience. The proposed framework provides a structured foundation for the design and development of future CDT systems.
Chinese Translation
随着数字孪生(Digital Twin, DT)系统从状态同步向面向任务和知识驱动的运行方式演进,认知数字孪生(Cognitive Digital Twins, CDTs)作为一种将认知能力融入孪生运行的扩展形式应运而生。现有的CDT研究往往聚焦于特定的使能技术,如学习模块、知识图谱和大语言模型,而对于如何将认知系统地集成到DT架构中缺乏深入的见解。针对这一问题,本文提出了一种由物理层、数字孪生层、认知层和任务层组成的四层CDT架构。该架构建立了一个贯穿这四层的自演化闭环运行回路:物理状态被同步为数字化表示;认知通过知识、记忆和注意力构建面向特定任务的认知模型;并在实际约束下生成任务级决策。运行反馈进一步优化认知经验,并更新数字化表示中的关系和标注,使后续的任务解释、发起和推理能够随系统运行而不断演化。基于该框架,本文刻画了两种代表性的运行模式:用户请求驱动的认知和自主驱动的认知。我们进一步讨论了与语义通信、知识查询、任务编排以及闭环同步相关的关键使能机制与部署挑战。一项轻量级仿真研究表明,该框架在有限语义信息条件下能够实现可靠的闭环任务,并通过积累的任务经验提升运行效率。所提出的框架为未来CDT系统的设计与开发提供了结构化的基础。
cs.AI / 16 / 2609.09627

Seven Sources of Physical AI Capability Formation

物理AI能力形成的七种来源
Chen, Gang
Abstract
Capabilities relevant to Physical AI can arise from materially different formation histories, yet existing taxonomies organized by morphology, architecture, learning algorithm, task, or domain do not directly answer what gives rise to a capability. We define a capability-formation source as a factor materially contributing to capability formation, distinct from components or construction steps. We identify seven non-exclusive sources: Recorded-Experience (RE), Predictive-Modeling (PM), Evaluative-Interaction (EI), Surrogate-Environment (SE), Mechanism-Grounded (MG), Embodied-Coupling (EC), and Evolution-Driven (ED) Formation. Using reconstructive induction with theoretical saturation, we traced a research matrix to primary studies, deduplicated the literature, set coding rules, and conducted three rounds of maximum-difference and negative-case sampling. Challenges included curriculum and self-supervised learning, active inference, open-ended and developmental learning, planning and search, neuro-symbolic architectures, digital twins, generative physical world models, and morphology-control co-design. Within the scope and criteria fixed as of September 4, 2026, all 49 evidence records were explainable by the seven sources individually or in combination. No R1-R3 challenge produced an irreducible eighth source, and R3 required no new core definition or substantive boundary rule. We therefore claim theoretical saturation within the stated scope, not logical completeness or exhaustive future coverage. The framework distinguishes similarity in observed capability from similarity in how it was formed, supporting analysis of explanation, transfer, replication, dependencies, governance evidence, and geoeconomic foundations.
Chinese Translation
与物理AI(Physical AI)相关的能力可能源自本质上不同的形成历程,然而现有的按形态、架构、学习算法、任务或领域组织的分类体系并不能直接回答能力由何产生。我们将能力形成来源(capability-formation source)定义为对能力形成有实质性贡献的因素,以区别于组件或构建步骤。我们识别出七种非互斥的来源:记录经验型(Recorded-Experience, RE)、预测建模型(Predictive-Modeling, PM)、评估交互型(Evaluative-Interaction, EI)、代理环境型(Surrogate-Environment, SE)、机制支撑型(Mechanism-Grounded, MG)、具身耦合型(Embodied-Coupling, EC)以及演化驱动型(Evolution-Driven, ED)形成。我们采用重构归纳法并结合理论饱和,将研究矩阵追溯至原始文献,对文献去重,制定编码规则,并进行了三轮最大差异抽样与负面案例抽样。所考察的挑战包括课程学习与自监督学习、主动推断、开放性与发展性学习、规划与搜索、神经符号架构、数字孪生、生成式物理世界模型以及形态-控制协同设计。在截至2026年9月4日确定的范围与标准内,全部49条证据记录均可由这七种来源单独或组合解释。没有任何R1-R3挑战产生不可化约的第八种来源,且R3阶段无需引入新的核心定义或实质性边界规则。因此,我们主张在所述范围内达到了理论饱和,而非逻辑完备性或对未来的穷尽覆盖。该框架区分了所观测能力的相似性与能力形成方式的相似性,从而支持对解释、迁移、复现、依赖关系、治理证据以及地缘经济基础的分析。
cs.AI / 17 / 2609.09646

RobustSGPO: Search-Space Control for Agent Harness Evolution

RobustSGPO:面向智能体框架演化的搜索空间控制
Zhao, Zibo, Shi, Jijun, Zhou, Mo, Wang, Zhongyuan, Bie, Shifu, Zhang, Yunfei, Zhou, Xuanting, Wu, Xiangyu, Liu, Bin, Tang, Ruiming, Ou, Wenwu, Gai, Kun
Abstract
Semantic-gradient-based prompt optimization (SGPO) improves agent harnesses using execution feedback, but its local update rule leaves the choice of edit scope and operation unresolved. We introduce RobustSGPO, which specifies the requested edit, constructs and checks the patch, and continues search from either the incumbent or retained snapshots. We evaluate permission scheduling, cumulative controls, and task-family transfer in the AgentX brainstorming workflow using 120 tasks, 95 runs, and 7,350 candidate attempts. Periodic $1\to2\to3$ scheduling exceeds fixed maximum permission by 0.28 test-score points. RobustSGPO increases completion on 30 held-out tasks from 60.0% to 80.0% and improves test quality from 3.77 to 4.14 under a 20-million-token budget. Category retention reduces source-task degradation after a shift, whereas random retention reaches a higher destination endpoint. Search-space control benefits quality through executable edits and alternative starting points, with measurable retention overhead.
Chinese Translation
基于语义梯度的提示优化(SGPO)利用执行反馈来改进智能体框架,但其局部更新规则未能解决编辑范围和操作的选择问题。我们提出 RobustSGPO,该方法明确指定所请求的编辑,构建并检查补丁,然后从当前最优解或保留的快照继续搜索。我们在 AgentX 头脑风暴工作流中,使用 120 个任务、95 次运行和 7,350 次候选尝试,评估了权限调度、累积控制和任务族迁移。周期性 $1\to2\to3$ 调度比固定最大权限高出 0.28 个测试分数点。在 2000 万词元的预算下,RobustSGPO 将 30 个保留任务上的完成率从 60.0% 提升至 80.0%,并将测试质量从 3.77 提升至 4.14。类别保留可在任务迁移后减少源任务性能退化,而随机保留则能达到更高的目标任务终点。搜索空间控制通过可执行的编辑和备选起点提升了质量,同时具有可衡量的保留开销。
cs.AI / 18 / 2609.09647

Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

智能体AI的黑盒红队测试:一种分类驱动的自动化风险发现框架
Kumar, Divyanshu, Birur, Nitin Aravind, Baswa, Tanay, Agarwal, Sahil, Harshangi, Prashanth
Abstract
Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations remain single-turn and fail to capture multi-step agent vulnerabilities. We present a systematic black-box framework for risk-aware agent evaluation requiring only basic system descriptions. Our approach introduces: (1) a seven-domain taxonomy mapping observable behaviors to risk categories, (2) fully automated SAGE-RT red teaming producing 120 adversarial scenarios per domain, and (3) human-validated evaluation using LLM judges. Empirical validation across two agent architectures (CrewAI and AutoGen) with four base models reveals alarming patterns: 56.25\% average governance risk, 65\% privacy risk in multi-agent configurations, and agent behavior vulnerabilities reaching 85\%. Our black-box approach effectively identifies critical architectural vulnerabilities without privileged access, providing a scalable path toward safer agent deployments.
Chinese Translation
智能体系统正迅速走向生产环境,它们读取不可信输入、调用具有真实权限的工具并自主行动,使安全攻击面超越仅支持对话的模型。然而,标准评估仍局限于单轮交互,无法捕捉多步骤智能体漏洞。我们提出了一个系统性的黑盒框架,用于风险感知的智能体评估,且仅需基本的系统描述。我们的方法引入了:(1)一个七域分类体系,将可观察行为映射到风险类别;(2)全自动的SAGE-RT红队测试,每个领域生成120个对抗性场景;(3)经人工验证的、使用LLM裁判的评估方法。在两种智能体架构(CrewAI和AutoGen)及四个基础模型上的实证验证揭示了令人担忧的模式:平均治理风险达56.25%,多智能体配置中的隐私风险达65%,智能体行为漏洞高达85%。我们的黑盒方法无需特权访问即可有效识别关键架构漏洞,为更安全的智能体部署提供了可扩展的路径。
cs.AI / 19 / 2609.09657

RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems

RESCUE-BENCH:迈向关系感知的多方情感支持对话系统
Hu, Haichuan, Xiao, Yang, Tang, Mingni, Duan, Jiawen, Zhang, Quanjun, He, Congqing, Zhang, Hao, Wang, Jiashuo, Hoorn, Johan F., Li, Wenjie
Abstract
Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual emotional states, leaving interpersonal relations in multi-party scenarios underexplored. In this work, we introduce relation-aware emotional support conversation, a new task that evaluates whether LLMs can capture and utilize the evolving dynamics of relationships to offer more effective emotional support. We construct RESCUE (Relation-aware Emotional Support Conversation Understanding and Evaluation Benchmark) from real couple and family interview conversations, containing 191 samples, 7,079 annotated turns, and 1,064.8 minutes of video. Based on rich annotations of socio-emotional and support-related dynamics, RESCUE defines six tasks that evaluate two core capabilities required for relation-aware emotional support: Relational Understanding and Relation-Sensitive Support. Experiments with ten LLMs show that current models perform relatively well on tasks relying on local emotional or intervention cues, but struggle with relation-intensive tasks such as relation pattern prediction, viewpoint prediction, and support strategy prediction. These findings reveal the limitations of current LLMs in modeling interpersonal relations and making relation-sensitive support decisions.
Chinese Translation
现有的情感支持对话系统主要关注一对一的求助者-支持者交互以及个体情绪状态,而多方场景中的人际关系尚未得到充分探索。在这项工作中,我们提出了关系感知情感支持对话这一新任务,用以评估大语言模型(LLM)能否捕捉并利用不断演变的关系动态,从而提供更有效的情感支持。我们基于真实的夫妻和家庭访谈对话构建了RESCUE(关系感知情感支持对话理解与评估基准,Relation-aware Emotional Support Conversation Understanding and Evaluation Benchmark),包含191个样本、7,079个标注轮次以及1,064.8分钟的视频。基于对社会情感和支持相关动态的丰富标注,RESCUE定义了六项任务,用以评估关系感知情感支持所需的两种核心能力:关系理解(Relational Understanding)和关系敏感支持(Relation-Sensitive Support)。对十个大语言模型的实验表明,当前模型在依赖局部情绪或干预线索的任务上表现相对较好,但在关系密集型任务上存在困难,例如关系模式预测、观点预测和支持策略预测。这些发现揭示了当前大语言模型在建模人际关系和做出关系敏感支持决策方面的局限性。
cs.AI / 20 / 2609.09664

PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations

PRAGMA:在终身对话中评估基于记忆对齐的个性化指导
Yu, Hyojeong, Koh, Hyukhun, Kim, Minsung, Jang, Yunah, Jung, Kyomin
Abstract
Large language models (LLMs) are increasingly deployed as personalized assistants that interact with users over extended periods of time. As conversations grow longer, relying on full interaction histories becomes increasingly inefficient and unreliable: long contexts introduce substantial computational overhead, making it difficult for models to consistently identify and utilize the most relevant information for the current request. These challenges have motivated memory systems that structure and retrieve user-specific information. In realistic interactions, users often seek practical guidance such as recommendations, planning, and decision support. Unlike factual recall tasks, personalized guidance requires models to integrate information across multiple past conversations and reason about changing user preferences and experiences. However, existing conversational memory evaluations mainly focus on retrieval and factual recall. To study this challenge, we introduce PRAGMA, a benchmark for evaluating personalized guidance in long-term conversations. PRGAMA contains curated longitudinal conversation histories, evidence annotations, and guidance scenarios grounded in evolving user contexts and incorrect user assumptions. Experiments across retrieval systems, memory systems, and long-context models reveal that current systems struggle both to recover the appropriate conversational evidence and to effectively use it for personalized guidance. Our results highlight the need for memory architectures that support robust conversational retrieval and memory-grounded reasoning beyond evidence recall.
Chinese Translation
大语言模型(LLM)越来越多地被部署为与用户进行长时间交互的个性化助手。随着对话变长,依赖完整交互历史变得越来越低效且不可靠:长上下文会带来大量的计算开销,使模型难以持续识别并利用与当前请求最相关的信息。这些挑战推动了记忆系统的发展,即对用户特定信息进行结构化组织和检索。在真实的交互中,用户常常寻求实用性指导,如推荐、规划和决策支持。与事实性回忆任务不同,个性化指导要求模型整合跨越多次过往对话的信息,并对不断变化的用户偏好和经历进行推理。然而,现有的对话记忆评估主要聚焦于检索和事实性回忆。为了研究这一挑战,我们提出了 PRAGMA,一个用于评估长期对话中个性化指导能力的基准。PRAGMA 包含精心构建的纵向对话历史、证据标注,以及基于不断演变的用户情境和用户错误假设的指导场景。跨检索系统、记忆系统和长上下文模型的实验表明,当前系统既难以恢复恰当的对话证据,也难以有效地将其用于个性化指导。我们的结果凸显了对记忆架构的需求,即该架构应支持稳健的对话检索,并支持超越证据回忆的、基于记忆的推理。
cs.AI / 21 / 2609.09678

Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents

可以安全停止吗?序列临床诊断智能体的风险约束停止机制
Wu, Yuexin, Rus, Vasile
Abstract
Clinical diagnosis agents must decide not only what test to request next, but also when to diagnose or defer. Existing agent benchmarks largely evaluate accuracy after fixed or unconstrained interaction, leaving autonomous stopping reliability implicit. We present Cros, a risk-constrained stopping layer combining state-wise error ranking, policy design on disjoint development splits, and LTT-style exact tests of selective diagnostic error and minimum autonomous coverage for complete sequential policies. Its finite-sample guarantee requires the candidate family, testing rule, and any randomization to be frozen before calibration labels are accessed. On a 1,834-episode MIMIC-derived abdominal-pain benchmark, the full ranker achieves exploratory state-error AUROC 0.853, compared with 0.715 for maximum class probability and 0.552 for the backbone's native stop score. On the previously viewed 367-episode evaluation split, analytically averaging over the frozen Cros weights yields 16.9% selective error at 78.8% coverage, cost 5.57, and 0.68 tests, versus 30.8% error at 100% coverage, cost 8.14, and 1.53 tests under native stopping. Forced continuation is non-monotone: error is 28.3% with HPI alone and 34.3% after full workup. However, the uniform-weight mixture ablation is cheaper on this viewed split despite missing the locked development margins, and Cros nominally satisfies the joint criterion in only 6 of 20 development resplits. Because evaluation labels were inspected during earlier development, these findings provide exploratory feasibility and audit evidence, not a confirmatory safety certificate.
Chinese Translation
临床诊断智能体不仅要决定下一步请求哪种检查,还必须决定何时做出诊断或推迟诊断。现有的智能体基准测试大多在固定或无约束的交互之后评估准确率,而将自主停止的可靠性隐含地忽略。我们提出了Cros,一个风险约束的停止层,它结合了基于状态误差排序、在不相交的开发集划分上的策略设计,以及LTT风格的精确检验,用于针对完整序列策略的选择性诊断误差和最低自主覆盖率。其有限样本保证要求候选策略族、检验规则及任何随机化机制在校准标签被访问之前予以冻结。在一个由MIMIC衍生的包含1,834个病例的腹痛基准测试中,完整排序器在探索性状态误差AUROC上达到0.853,而最大类别概率为0.715,主干模型的原生停止分数仅为0.552。在先前已查看过的367个病例的评估划分上,对冻结的Cros权重进行解析平均,得到16.9%的选择性误差、78.8%的覆盖率、5.57的成本和0.68次检查;相比之下,使用原生停止机制则在100%覆盖率下产生30.8%的误差、8.14的成本和1.53次检查。强制继续问诊呈现非单调性:仅使用现病史(HPI)时误差为28.3%,完成全面检查后误差为34.3%。然而,在已查看的划分上,均匀权重混合的消融版本尽管未达到锁定的开发集边界,成本却更低;且Cros仅在20个开发集重划分中的6个上名义上满足联合准则。由于评估标签在早期开发过程中已被查看,这些发现仅提供探索性可行性和审计证据,而非确认性的安全性证明。
cs.AI / 22 / 2609.09702

Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation

正确性门控多教师蒸馏中的决策偏移、标签功能缺失与结论不明的依据性审计
Feng, Xiaofei
Abstract
Candidate decision correctness and rationale grounding are different objectives. We examine correctness-gated multi-teacher distillation in a fixed experiment. Eight arms share 4,330 sources, a 63.9M-parameter student, 12,990 optimization rows, 406 updates, evidence inputs, and a decoder; seven teacher-based arms use one fixed three-response pool. Three seeds are evaluated on 267 held-out examples. Relative to unfiltered distillation, the correctness-weighted arm differed in accuracy by +0.1660 (95% observed-matrix interval [0.0670, 0.2455]), five-label macro-F1 by +0.1323 ([0.0916, 0.1731]), and task-defined conditional unsafe-action rate by -0.4979 ([-0.5926, -0.3686]). These shifts do not imply uniformly better behavior. Source-label SFT had the highest mean macro-F1 (0.586). The weighted arm had zero Refuted recall in every seed, and two seeds assigned NotEnoughInfo to all 167 claim examples. In an availability-amended audit at one reference seed, weighted and unfiltered outputs had 0/20 versus 1/20 evidence-supported positives and 20/20 versus 19/20 positives containing unsupported material. Samples were non-paired, source overlap was not serialized, and the amendment followed automatic summarization but preceded annotation. The audit therefore cannot estimate a common-source grounding effect and is inconclusive about system-level improvement or harm. Hard filtering already achieved 0.660 accuracy, 0.530 macro-F1, and 0.135 conditional unsafe rate. The implemented weighted arm showed no demonstrated incremental decision benefit over hard filtering. This fixed-matrix failure analysis shows decision redistribution with lost label functionality; the available human audit does not establish a grounding gain.
Chinese Translation
候选决策的正确性与推理依据的锚定性是两个不同的目标。我们在固定实验设置下考察了正确性门控的多教师蒸馏。八个实验组共享4,330条源数据、一个63.9M参数的学生模型、12,990条优化数据、406次更新、证据输入和解码器;七个基于教师的实验组使用同一个固定的三回答池。三个随机种子在267个留存样本上进行评估。相对于未过滤蒸馏,正确性加权组的准确率差异为+0.1660(观测矩阵95%区间[0.0670, 0.2455]),五标签宏F1差异为+0.1323([0.0916, 0.1731]),任务定义的条件性不安全动作率差异为-0.4979([-0.5926, -0.3686])。这些偏移并不意味着行为全面更优。源标签SFT的宏F1均值最高(0.586)。加权组在每个种子上的Refuted召回率均为零,且两个种子将全部167条声明样本判为NotEnoughInfo。在一个参考种子的可用性修正审计中,加权组与未过滤组的输出在证据支持的正例上分别为0/20与1/20,而包含无支持内容的正例分别为20/20与19/20。样本为非配对样本,源数据重叠未经序列化处理,且修正发生于自动摘要之后、人工标注之前。因此,该审计无法估计共同源数据的依据锚定效应,对于系统层面的改进或损害均无定论。硬过滤已达到0.660的准确率、0.530的宏F1和0.135的条件性不安全率。所实现的加权组相对于硬过滤并未显示出可证明的增量决策收益。这一固定矩阵的失败分析表明存在决策重新分布并伴随标签功能的缺失;现有的人工审计并不能确立依据锚定方面的增益。
cs.AI / 23 / 2609.09707

Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning

SFT究竟应该学习哪些Token?从Token裁剪视角看数学推理
Jia, Yaning, Zhang, Chunhui, Xu, Wenxuan, Diao, Xingjian, Wang, Xiaoyuan, Vosoughi, Soroush
Abstract
Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already mastered tokens while amplifying learning pressure on uncertain, low-confidence tokens, leading to suboptimal training dynamics. We propose Trimmed Logit-Gap SFT (TrimSFT), a simple token-level reweighting method that scales the SFT loss according to the logit gap between the gold token and its strongest competitor. TrimSFT trims supervision away from both extremes: tokens already mastered (large logit gap) and tokens weakly supported by the current model (small or negative logit gap), concentrating learning within an intermediate logit-gap region between them. We instantiate this principle with a Gaussian weight centered at margin m with bandwidth {\tau}, requiring no reference model or additional forward pass. We evaluate TrimSFT on six base models from the Llama, Qwen, and DeepMath families across five mathematical reasoning benchmarks. TrimSFT consistently improves over standard SFT, achieving the best average performance on five out of six models, with gains of up to +26.9 points over SFT on MATH500. Further analyses show that the bandwidth {\tau} matters more than the exact margin location, and that half-trim variants that remove supervision pressure from only one side yield inferior trade-offs. A token-level logit-gap distribution analysis suggests that TrimSFT reshapes model confidence in a more balanced way than uniform SFT or monotonic reweighting methods. These results suggest that reasoning SFT can benefit from trimming both extremes rather than treating all tokens uniformly.
Chinese Translation
监督微调(SFT)对所有目标token施加统一的交叉熵损失,尽管不同token为数学推理提供的学习信号并不均等。这种统一处理会导致模型对已掌握的token过度锐化,同时放大对不确定的、低置信度token的学习压力,从而导致次优的训练动态。我们提出了Trimmed Logit-Gap SFT(TrimSFT),这是一种简单的token级重加权方法,根据金标token与其最强竞争者之间的logit间隔(logit gap)来缩放SFT损失。TrimSFT从两个极端裁剪监督信号:已被模型掌握的token(大logit间隔)以及当前模型支持较弱的token(小或负logit间隔),将学习集中在两者之间的中间logit间隔区域。我们以中心为边界m、带宽为{\tau}的高斯权重来实现该原则,且无需参考模型或额外的前向传播。我们在来自Llama、Qwen和DeepMath系列的六个基座模型上,通过五个数学推理基准评估TrimSFT。TrimSFT始终优于标准SFT,在六个模型中的五个上取得了最佳平均性能,在MATH500上比SFT最高提升26.9个点。进一步分析表明,带宽{\tau}比精确的边界位置更为重要,而仅从单侧去除监督压力的半裁剪变体则产生较差的权衡效果。token级logit间隔分布分析表明,与统一SFT或单调重加权方法相比,TrimSFT以更均衡的方式重塑了模型置信度。这些结果表明,推理SFT可以通过裁剪两端极端token而受益,而非统一对待所有token。
cs.AI / 24 / 2609.09735

Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety

人工智能能否通过早期网络欺凌检测支持医疗保健与心理健康?情感感知人工智能对主动式在线安全的影响
Jelodar, Hamed, Firouzi, Amir, Lo, Yen-Wu, Tanha, Maryam, Dadkhah, Sajjad
Abstract
Healthcare systems, mental health, and public well-being are increasingly affected by cyberbullying and harmful online interactions. This paper presents CareGuard, an early-warning framework designed to support healthcare-driven mental health protection and proactive online safety through the detection of cyberbullying-related content using advanced natural language processing techniques. CareGuard integrates zero-shot semantic labeling with fine-tuned transformer-based models, including BERT, DistilBERT, and RoBERTa, to enable robust and context-aware classification across sensitive cyberbullying categories. To improve efficiency and reduce unnecessary computation in healthcare-oriented monitoring settings, the framework incorporates an emotion-aware filtering mechanism alongside cosine similarity-based semantic screening, allowing the system to focus on semantically relevant and emotionally salient content. Experimental results on benchmark datasets demonstrate that CareGuard effectively balances detection accuracy and computational efficiency, highlighting its potential for scalable deployment in healthcare systems, mental health monitoring, and online safety applications.
Chinese Translation
医疗保健系统、心理健康和公众福祉正日益受到网络欺凌和有害在线交互的影响。本文提出了 CareGuard,一个早期预警框架,旨在通过利用先进的自然语言处理技术检测网络欺凌相关内容,支持以医疗保健为导向的心理健康保护和主动式在线安全。CareGuard 将零样本(zero-shot)语义标注与微调的基于Transformer的模型(包括 BERT、DistilBERT 和 RoBERTa)相集成,以实现对敏感网络欺凌类别的稳健且具备上下文感知能力的分类。为了提高效率并减少面向医疗保健的监测环境中不必要的计算,该框架在基于余弦相似度的语义筛选之外,还引入了情感感知过滤机制,使系统能够专注于语义相关且情感显著的内容。在基准数据集上的实验结果表明,CareGuard 有效地平衡了检测准确率与计算效率,凸显了其在医疗保健系统、心理健康监测和在线安全应用中可扩展部署的潜力。
cs.AI / 25 / 2609.09754

LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents

LexAgentHallu:一个用于剖析法律智能体幻觉的分层基准
Zhou, Yujin, Zheng, Mingxuan, Cao, Chuxue, Yidan, Huang, Chen, Jiale, Guo, Yike, Han, Sirui
Abstract
As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal benchmarks evaluate only single-turn QA with outcome-level metrics, while agentic hallucination benchmarks lack legal-specific diagnostic capability. Neither answers to what extent and how a legal agent hallucinates along its trajectory. To address these limitations, we introduce LexAgentHallu, a legal agentic hallucination benchmark designed to evaluate to what extent and how legal agents fail along multi-step trajectories. Built through a four-stage expert-in-the-loop pipeline, LexAgentHallu contains 3414 instances across 17 legal categories and 6 task types. Each instance is annotated under a dual-layer hallucination taxonomy of 7 high-level categories and 27 fine-grained subclasses, covering both substantive errors and agent-procedural failures. We further design fine-grained metrics that quantify to what extent and localize how each failure occurs along an agent's execution path. Our evaluation across 18 proprietary and open-source agents uncovers a Right-Answer-Wrong-Reason effect and reveals that hallucination subclasses cluster rather than scatter, forming distinct agentic framework, legal task, and category profiles. These findings, invisible to outcome-level evaluation, validate the diagnostic power of LexAgentHallu for evaluating agentic hallucination in law.
Chinese Translation
随着大语言模型越来越多地被部署为工具增强型法律智能体,它们会引入智能体幻觉(agentic hallucinations),即工具调用与推理错误级联放大,最终导致虚构裁判结论和错误引用法律依据。然而,现有的法律基准仅通过结果层面的指标评估单轮问答,而智能体幻觉基准又缺乏法律领域特有的诊断能力。两者都无法回答法律智能体在其执行轨迹中幻觉的程度和方式。为解决这些局限,我们提出了LexAgentHallu,一个法律智能体幻觉基准,旨在评估法律智能体在多步轨迹中失败的程度与方式。LexAgentHallu通过四阶段专家参与(expert-in-the-loop)的流程构建,包含横跨17个法律类别和6种任务类型的3414个实例。每个实例均在一个双层幻觉分类体系下进行标注,该体系包含7个高层类别和27个细粒度子类,同时涵盖实质性错误与智能体程序性失败。我们进一步设计了细粒度指标,用以量化失败的程度并定位每次失败在智能体执行路径上的发生方式。我们在18个专有与开源智能体上的评估揭示了一种“答案正确但推理错误”(Right-Answer-Wrong-Reason)效应,并发现幻觉子类并非分散分布,而是呈现聚类现象,形成独特的智能体框架、法律任务和类别画像。这些在结果层面评估中不可见的发现,验证了LexAgentHallu在评估法律领域智能体幻觉方面的诊断能力。
cs.AI / 26 / 2609.09774

Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks

变化中的程序性记忆:受控网络任务中的复用与干扰
Cao, Yanze
Abstract
Procedural memory lets language agents reuse successful routines, but reuse presumes that a stored routine remains applicable. We study what happens when that presumption is deliberately violated. The study combines a retrospective, human-assisted interface-adaptation case from BrowserGym TimeWarp with controlled frozen-memory comparisons on synthetic shopping decisions. During the documented WebShop V1-V6 development path, interface-specific code was adapted while the separately stored high-level procedure was not reported to change; this phase does not constitute an autonomous memory-agent evaluation. In the controlled phase, an early pilot produced one task on which two memory conditions selected a more expensive item while the no-memory condition selected the reference minimum. Follow-up probes did not establish a recurring row-order or identity-binding pattern. We then tested four forms of mismatch: changed quantities, a different evidence representation, a conflict between local and global optimization, and distributed promotion evidence, across 32 formal cells. Each cell used one temperature-0 generation with the same local qwen3:8b configuration and no adaptive retry. Across these pairs, none of the predefined diagnostic interference signatures appeared on the tasks for which they were defined when current-task evidence was explicit and sufficient. The result identifies a tested region of non-interference: a procedural memory can be mismatched without becoming behaviorally disruptive. It does not establish general safety or a mechanism. The remaining question is which additional conditions turn applicability mismatch into observable, memory-caused error.
Chinese Translation
程序性记忆使语言智能体能够复用成功的例程,但复用前提是存储的例程仍然适用。我们研究了当这一前提被刻意打破时会发生什么。本研究结合了两部分:来自 BrowserGym TimeWarp 的回顾性、人工辅助的界面适配案例,以及在合成购物决策上进行的受控冻结记忆比较。在记录在案的 WebShop V1–V6 开发路径中,特定于界面的代码被适配,而单独存储的高层程序据报告并未改变;该阶段并不构成自主记忆智能体评估。在受控阶段,早期试点产生了一个任务,在该任务上两个记忆条件选择了更昂贵的商品,而无记忆条件选择了参考最低价商品。后续探针检验未发现可复现的行序或身份绑定模式。随后,我们在 32 个正式实验单元中测试了四种失配形式:数量变化、不同的证据表示、局部与全局优化之间的冲突,以及分散的促销证据。每个单元使用相同的本地 qwen3:8b 配置进行一次温度为 0 的生成,且无自适应重试。在这些配对中,当当前任务证据明确且充分时,没有任何预定义的诊断性干扰特征在其所定义的任务上出现。该结果划定了一个经过检验的无干扰区域:程序性记忆在失配的情况下可以不产生行为上的破坏性。但本研究并未确立一般性安全性或其机制。尚待解决的问题是:哪些额外条件会将适用性失配转化为可观察的、由记忆导致的错误。
cs.AI / 27 / 2609.09776

Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward

携带证明的认知:以现实结算奖励弥合验证鸿沟
M, Eshwar Reddy, Karmakar, Sourav
Abstract
Frontier gains in language-model reasoning come from reinforcement learning on reasoning traces and are concentrated in domains with a cheap, sound verifier. We argue the field's binding constraint is the verification gap: no scalable, incorruptible reward for reasoning outside formal domains. We make four contributions. (1) Theory: in a joint-Gaussian model of best-of-N selection, verifier-gold correlation rho is the exact exchange rate between test-time compute and capability, and an unsound verifier pays a polynomial penalty N^(1/rho^2); a margin-free copula form predicts realized soundness of real LLM judges to 4% median error. (2) Demonstration: in program-synthesis testbeds with executable ground truth, including a pre-registered scaled replication, unsound verifiers lose Soundness-under-Pressure as optimization grows (0.94 to 0.32 at N=4096) while a sound verifier improves monotonically; reality-anchored settlement beats a frozen verifier under i.i.d. and adversarial pressure, driving the hacking gap from ~0.27 to ~0; soundness scales log-linearly with settled labels, with on-policy settlement ~10x more label-efficient than random labeling. With real LLM judges and unit-test execution as gold, a weak judge loses soundness under best-of-N (p<0.001), a stronger judge is more robust, and selection alone manufactures +0.53 hacking gaps from honest samples. Under real GRPO training, a frozen reward model traces the full overoptimization curve (executed reward collapses 90%) while the same model refit on a 10% settlement stream preserves 6x the executed reward. (3) Paradigm: proof-carrying cognition, where reasoning steps are typed probabilistic claims priced by a self-built world model trained only on held-out reality and settled by proper scoring rules. (4) Benchmark: we specify Soundness-under-Pressure as the headline metric for a reality-settled reasoning benchmark.
Chinese Translation
语言模型推理能力的前沿进展来自对推理轨迹的强化学习,并集中于拥有廉价、可靠验证器的领域。我们认为,该领域的核心约束在于验证鸿沟:在形式化领域之外,缺乏可扩展、不可腐化的推理奖励信号。本文做出四项贡献。(1)理论:在Best-of-N选择的联合高斯模型中,验证器与真实标签的相关系数ρ是测试时算力与能力之间的精确兑换率;不可靠验证器需付出多项式级惩罚N^(1/ρ^2);一种无边际copula形式以4%的中位误差预测了真实LLM评判者的实际可靠性。(2)实证演示:在具有可执行真值的程序合成测试平台(包括一项预注册的规模化重复实验)中,随着优化规模增大,不可靠验证器的“压力下可靠性”下降(N=4096时从0.94降至0.32),而可靠验证器单调提升;在独立同分布和对抗压力下,现实锚定的结算机制优于冻结验证器,将作弊差距从约0.27降至约0;可靠性随结算标签呈对数线性增长,且在策略结算的标签效率约为随机标注的10倍。以真实LLM评判者和单元测试执行为真值时,弱评判者在Best-of-N下失去可靠性(p<0.001),较强评判者更为稳健,而仅靠选择机制便能从诚实样本中制造出+0.53的作弊差距。在真实GRPO训练中,冻结的奖励模型呈现完整的过度优化曲线(可执行奖励崩溃90%),而同一模型在10%结算数据流上重新拟合后,可保留6倍的可执行奖励。(3)范式:携带证明的认知,其中推理步骤是被类型化的概率性断言,由仅在留出现实数据上训练的自建世界模型定价,并通过恰当的评分规则进行结算。(4)基准:我们将“压力下可靠性”指定为现实结算推理基准的核心指标。
cs.AI / 28 / 2609.09815

UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model

UnitBoost:用合并算子而非模型来管理复合LLM系统
Zhang, Xing, Wang, Guanghui, Cui, Yanwei, Wang, Mengdie Flora, He, Peiyang
Abstract
Compound LLM systems often solve a coordination problem by adding a higher-level LLM. The resulting meta-agent reads workers' outputs, writes the final answer, allocates later calls, and decides when to stop. It is expressive, but it also concentrates three control decisions in an opaque, order-sensitive model call. We ask whether the manager needs to be generative at all. UnitBoost replaces that model with a defined meta-level operator: a task-given unit map turns worker outputs into slot-value proposals, a constrained argmax assembles the output, and the slots left unfilled or unsupported become an explicit residual for the next round. The operator is order-free, records unit provenance, and gives a simple guarantee: without coupling constraints, unit-wise maximization under the same admission score dominates selection of any complete candidate. On three held-out benchmarks, it exceeds the best single candidate chosen with gold labels by 0.060-0.195 absolute task-score points and input-matched generative managers by 0.048-0.076. Replacing only the management step improves six compound-system configurations by 0.013-0.182. Residual-directed rounds raise FanOutQA cell F1 from 0.4778 to 0.5524; matched controls show that the true residual outperforms random targets and ordinary rereading, while a label-free supply signal flags exhaustion after one unproductive round. The same analysis measures three conditions in which no such gain is available (one indivisible unit, unavailable unit identity, and an endpoint that charges for every emitted unit) and quantifies cross-unit coupling as a repair cost. The manager gives up semantic freedom and gains order invariance, unit provenance, and testable failure conditions.
Chinese Translation
复合LLM系统通常通过添加一个更高层级的LLM来解决协调问题。由此产生的元智能体(meta-agent)负责读取工作节点的输出、撰写最终答案、分配后续调用,并决定何时停止。这种做法具有很强的表达能力,但它也将三项控制决策集中在一个不透明的、对顺序敏感的模型调用之中。我们提出的问题是:管理者是否真的需要是生成式的?UnitBoost用一个明确定义的元层算子取代了该模型:一个基于任务给定的单元映射(unit map)将工作节点的输出转化为槽位-值(slot-value)提案,一个带约束的argmax过程组装最终输出,而未被填充或缺乏支持的槽位则成为下一轮的显式残差(residual)。该算子与顺序无关,记录单元溯源信息,并提供一个简单的保证:在不考虑耦合约束的情况下,在相同准入得分下的单元级最大化优于对任何完整候选的选择。在三个保留基准(held-out benchmarks)上,其任务得分比以黄金标签选出的最佳单一候选高出0.060至0.195个绝对分点,比输入匹配的生成式管理者高出0.048至0.076。仅替换管理步骤即可将六种复合系统配置提升0.013至0.182。残差引导的轮次将FanOutQA的单元格F1从0.4778提升至0.5524;匹配对照实验表明,真实的残差优于随机目标和普通的重读策略,而无标签的供给信号能在一轮无产出之后标记出耗尽状态。同一分析还刻画了三种无法获得此类收益的情形(单一不可分割单元、单元身份不可用,以及每个输出单元都需付费的终止条件),并将跨单元耦合量化为修复成本。该管理者放弃了语义自由,却换来了顺序不变性、单元溯源以及可检验的失败条件。
cs.AI / 29 / 2609.09853

The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents

Era by Eon 基准测试:一个具有精确真值的生成式企业环境,用于评测LLM智能体
Gruenbaum, Benjamin, Porat, Doron, Natanzon, Assaf, Zavida, Roy, Dinachi, Chen, Itzahary, Or
Abstract
LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no existing substitute provides ground truth. We present the Era by Eon Benchmark for evaluating LLM agents that use enterprise tools. The benchmark is built around a complete fictional company. It includes product simulators, company-specific internal databases, benchmark questions, and computed answer keys. Industry, company size, business model, application portfolio, and a seed define each company. One seeded entity graph supplies shared company data to simulators of Salesforce, Zendesk, Slack, Gong, and other products. A questionconditioned generator creates the schemas and records for internal databases. It takes shared entities, keys, and values from the same graph before generating database-specific facts. Both mechanisms therefore describe one consistent enterprise estate. Every expected answer is computed from the final records, so grading is exact. Design and answer-key checks validate the internal databases. A realism scorecard and adversarial detector validate the entity graph. Across 23 generated companies, the mean realism score rose from 61.8 to 97.0, with zero records flagged as synthetic. In the reported simulator-track comparison, nine models answered the same 33 questions three times each. Accuracy estimates ranged from 42.4% to 76.8%, and three of 36 pairwise differences remained supported after correction.
Chinese Translation
针对企业记录系统的LLM智能体无法在客户生产数据上进行评估,而现有的替代方案均无法提供真值。我们提出了Era by Eon基准测试,用于评估使用企业工具的LLM智能体。该基准测试围绕一家完整的虚构公司构建,包括产品模拟器、公司专属内部数据库、基准测试问题以及计算生成的标准答案。行业、公司规模、商业模式、应用组合以及种子共同定义每家公司。一个由种子生成的实体图为Salesforce、Zendesk、Slack、Gong等产品的模拟器提供共享的公司数据。一个基于问题的生成器为内部数据库创建模式和记录,它先从同一实体图中获取共享实体、键和值,然后再生成数据库特定的事实。因此,这两种机制描述的是同一个一致的企业环境。每个预期答案均由最终记录计算得出,因此评分是精确的。通过设计和标准答案检查来验证内部数据库,并通过真实感评分卡和对抗性检测器来验证实体图。在23家生成的公司中,平均真实感得分从61.8提升至97.0,且没有记录被标记为合成数据。在所报告的模拟器赛道对比中,九个模型各自回答相同的33个问题三次。准确率估计值介于42.4%至76.8%之间,经校正后,36组两两比较中仍有三组差异得到支持。
cs.AI / 30 / 2609.09864

Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields

情感计算中关系范式的转变:情感共鸣、活力情感与语音交互场
Gorman, Cy, Yao, Yihang
Abstract
Affective computing has largely followed an individual-state paradigm, extracting discrete emotion labels or arousal/valence from isolated speakers. We argue this framing is incomplete for interaction. Drawing on affective resonance and vitality-contour accounts, we propose a relational framework in which the primary unit of affective analysis is the interactional field constituted within vocal dynamics. As a proof of concept, we present a preliminary empirical study using continuous self-supervised speech representations to detect directional expressive coupling in multi-party conversation. Coupling is regime-specific, concentrated at sub-second timescales, and collapses under exclusive-speech negative controls, consistent with a relational account of affective dynamics. We introduce design frameworks for Artificial Affective Resonance Intelligence grounded in Affective Resonance Dynamic Ontologies, supported by null-calibrated directional coupling analyses across interaction regimes.
Chinese Translation
情感计算(Affective Computing)在很大程度上遵循个体状态范式,从孤立的说话人中提取离散的情绪标签或唤醒度/效价。我们认为,这种框架对于交互研究而言是不完整的。借鉴情感共鸣(affective resonance)与活力轮廓(vitality-contour)的相关理论,我们提出了一个关系性框架,其中情感分析的基本单元是在语音动态中构成的交互场(interactional field)。作为概念验证,我们开展了一项初步的实证研究,利用连续的自监督语音表示来检测多方对话中的方向性表达耦合。研究发现,这种耦合具有特定模式的特征,集中于亚秒级时间尺度,并且在独占语音的阴性对照条件下消失,这与情感动态的关系性解释相一致。我们基于情感共鸣动态本体论(Affective Resonance Dynamic Ontologies),引入了人工情感共鸣智能(Artificial Affective Resonance Intelligence)的设计框架,并辅以跨交互模式的零校准方向性耦合分析。
cs.AI / 31 / 2609.09875

AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents

AgentAudit:一个开放的、可扩展的AI智能体全生命周期信任评估框架
Nag, Shrey, Sachita, Singh, Abhishek Kumar, Goel, Lipi, Janwar, Rajeshwar Singh
Abstract
Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion (AgentBench) or security robustness (AgentDojo, ASB), rather than the complete pipeline of planning, tool selection, tool execution, memory and reasoning. Failures can occur at any stage, yet existing benchmarks rarely identify their precise source. AgentAudit evaluates the entire execution trace across ten capability, grounding, security and behavioural dimensions, namely instruction integrity, planner, memory, tool selection, tool invocation, tool correctness, alignment, tool faithfulness, security and execution integrity, combined with behavioural classification and failure attribution to pinpoint the exact stage responsible for an observed failure. AgentAudit can evaluate any LLM-based AI agent, since it attaches to the agent instead of replacing it. It reads only the recorded execution trace and does not interfere with how the agent runs, so it places no constraint on the agent's internal implementation. We evaluate five language models (OpenAI GPT-5, Claude Sonnet 5, Sarvam 105B, Llama 3.3 70B and Gemini 2.5 Flash) across nine capability and adversarial tasks. Claude Sonnet 5 and GPT-5 obtain the highest mean Composite Trust Scores (95.1 and 80.6 out of 100, respectively), while Sarvam 105B, Llama 3.3 70B and Gemini 2.5 Flash trail substantially (57.6, 45.7 and 22.6). All traces were scored by a single fixed judge model, which was itself one of the evaluated models, a limitation discussed in Section VII.E. More importantly, models with similar task-completion behaviour can diverge sharply in trustworthiness, as several non-frontier models are repeatedly classified Unsafe_Compliance on adversarial tasks rather than merely failing them, a distinction that pass/fail benchmarks cannot surface.
Chinese Translation
现有的评估框架大多仅评估AI智能体的某一方面,例如任务完成度(AgentBench)或安全鲁棒性(AgentDojo、ASB),而非涵盖规划、工具选择、工具执行、记忆与推理的完整流程。失败可能发生在任一阶段,然而现有基准测试很少能够识别其确切来源。AgentAudit在十个能力、接地、安全及行为维度上评估完整的执行轨迹,即指令完整性、规划器、记忆、工具选择、工具调用、工具正确性、对齐性、工具忠实性、安全性和执行完整性,并结合行为分类与失败归因,精确定位导致所观察到的失败的具体阶段。AgentAudit可以评估任何基于LLM的AI智能体,因为它附加于智能体之上而非将其替换。它仅读取记录的执行轨迹,不干预智能体的运行方式,因此对智能体的内部实现没有任何限制。我们在九个能力与对抗性任务上评估了五个语言模型(OpenAI GPT-5、Claude Sonnet 5、Sarvam 105B、Llama 3.3 70B和Gemini 2.5 Flash)。Claude Sonnet 5和GPT-5获得了最高的平均综合信任得分(Composite Trust Scores,满分100分中分别为95.1和80.6),而Sarvam 105B、Llama 3.3 70B和Gemini 2.5 Flash则明显落后(分别为57.6、45.7和22.6)。所有轨迹均由单一固定的裁判模型打分,该模型本身也是被评估的模型之一,这一局限性将在第VII.E节中讨论。更重要的是,任务完成行为相似的模型在可信度上可能存在巨大差异,因为多个非前沿模型在对抗性任务上不仅未能通过任务,还被反复归类为Unsafe_Compliance(不安全顺从),这种区别是无法通过通过/失败型基准测试所显现的。
cs.AI / 32 / 2609.09882

Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format

行为语言模型中的打分式与生成式读出:引出格式的实证研究
Kraisingkorn, Touchapon, Pachtrachai, Krittin, Modecrua, Wachiravit
Abstract
Language models fine-tuned on customer behavior can predict outcomes and generate explanations, but these readouts are often treated as interchangeable. Holding model checkpoint and prompt content fixed, we compare probabilities obtained by scoring answer tokens with predictions generated after a written rationale. Across 13 model-domain cells covering four retail tasks in three markets, including two using fully public data and checkpoints, the scored readout ranks outcomes more accurately in 12 of 13 cells (two-sided sign test, p approximately 0.003), by 1.5 to 14.5 points in area under the receiver operating characteristic curve (AUC). Paired bootstrap confidence intervals exclude zero in every newly measured cell. The gap varies with task-specific supervision and mismatch between training and serving formats, ranging from -2.2 points for an untuned base model to +13.7 for rationale-format supervision. Analysis of approximately 9,000 rationales identifies two correlates: reduced reliance on the dominant predictive feature and convergence on stock formulations. Probability saturation does not track the gap. A third readout, eliciting a probability before any verdict, improves calibration (Brier score from 0.47 to 0.15) while ranking within noise of scoring, but only for outcome rates represented in training; it is worse than scoring when the scored head is already calibrated. We interpret these differences through the objectives matched by each readout, identify training choices that narrow the gap, and propose retaining generated rationales while sourcing ranking from the scored head.
Chinese Translation
在客户行为数据上微调的语言模型可以预测结果并生成解释,但这些读出(readout)方式通常被视为可互换的。在固定模型检查点和提示内容的前提下,我们比较了通过对答案词元打分获得的概率与在生成书面推理(rationale)之后产生的预测。在覆盖三个市场四项零售任务的13个模型-领域组合中(其中两个组合使用完全公开的数据和检查点),打分式读出在13个组合中的12个里对结果的排序更为准确(双侧符号检验,p约为0.003),在受试者工作特征曲线下面积(AUC)上高出1.5至14.5个百分点。配对自助法(bootstrap)置信区间在每个新测量的组合中均排除零。这一差距随任务特定的监督方式以及训练与部署格式之间的不匹配程度而变化,从未微调的基础模型的-2.2个百分点到推理格式监督下的+13.7个百分点不等。对约9,000条推理文本的分析识别出两个相关因素:对主导预测特征的依赖减少以及向固定表述的趋同。概率饱和并不能追踪这一差距。第三种读出方式——在给出任何结论之前先引出概率——改善了校准(Brier分数从0.47降至0.15),同时排序性能与打分方式相当(在噪声范围内),但仅对训练中出现过的结果比例有效;当打分头部(scored head)本身已校准良好时,其表现劣于打分方式。我们通过每种读出方式所匹配的目标来解释这些差异,识别出能够缩小差距的训练选择,并建议在保留生成式推理文本的同时,从打分头部获取排序结果。
cs.AI / 33 / 2609.09885

Decision Transformer for UAV-Mounted RIS-Assisted Dynamic D2D Communications

面向无人机载可重构智能表面辅助动态D2D通信的决策Transformer
Liu, Yaxuan
Abstract
This paper studies unmanned aerial vehicle (UAV)-mouted reconfigurable intelligent surface (RIS)-assisted device-to-device (D2D) communication with stochastic link activation. It models UAV motion and attitude, time-varying Rician angles, and angle-dependent RIS reflection. A joint optimization of UAV trajectory, attitude, and RIS phases is formulated to maximize average sum rate under mobility, energy, and hardware constraints. The problem is addressed using deep reinforcement learning and a Decision Transformer trained on expert trajectories from multiple scenarios. Results demonstrate effective cross-scenario generalization, with zero-shot transfer outperforming direct DRL transfer and online fine-tuning achieving competitive performance with fewer interactions.
Chinese Translation
本文研究了具有随机链路激活特性的无人机(UAV)载可重构智能表面(RIS)辅助的设备到设备(D2D)通信。该研究建模了无人机的运动与姿态、时变莱斯(Rician)角度以及依赖角度的RIS反射。在移动性、能量和硬件约束下,构建了对无人机轨迹、姿态和RIS相位的联合优化问题,以最大化平均和速率。该问题通过深度强化学习求解,并使用在多个场景的专家轨迹上训练的决策Transformer(Decision Transformer)进行处理。实验结果表明该方法具备有效的跨场景泛化能力,其中零样本迁移优于直接的深度强化学习迁移,且在线微调以更少的交互次数实现了具有竞争力的性能。
cs.AI / 34 / 2609.09898

Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

面向自然语言到PDDL问题生成的有据评估与修复
Rosa, Joana, Santos, Pedro, Oliveira, Valdemar, Silva, Romão, Silveira, L. Miguel, Martins, Bruno
Abstract
Large Language Models (LLMs) have shown promise for translating Natural Language (NL) planning descriptions into PDDL problem instances. However, standard evaluation criteria such as syntactic validity or planner success can substantially overestimate faithfulness to the described task: a generated problem may be parseable and solvable while misrepresenting the intended initial state, goal, object structure, or optimization target. This paper studies an end-to-end NL-to-PDDL pipeline that combines LLM generation, checks in terms of PDDL parsing, planning and validation, a domain-conformance checker, an LLM critic, and iterative repair. Fine-grained repair feedback is constructed from the domain description, the generated problem, the natural language problem description, and operational diagnostics. Reference-based comparisons against curated benchmark PDDL problem descriptions are used for post-hoc benchmark analysis, and these offline checks include renaming-invariant structural matching and semantic equivalence, where domain support is available. Across Planetarium, AutoPlanBench, and curated PDDL~2.1 problems, results show that operational success and benchmark-reference reconstruction can diverge substantially. Results also show that structured repair can be useful, and that PDDL~2.1 remains challenging for reference reconstruction, even when operational success improves.
Chinese Translation
大语言模型(LLM)在将自然语言(NL)规划描述转换为PDDL问题实例方面已展现出前景。然而,语法有效性或规划器求解成功等标准评估指标可能显著高估其与所描述任务的一致性:生成的问题可能可解析且可求解,却错误地呈现了预期的初始状态、目标、对象结构或优化目标。本文研究了一种端到端的自然语言到PDDL的流水线,该流水线结合了LLM生成、基于PDDL解析、规划与验证的检查、领域一致性检查器、LLM评审器以及迭代修复。细粒度的修复反馈由领域描述、生成的问题、自然语言问题描述以及运行层面的诊断信息构建而成。在有参考答案的情况下,基于参考答案与精选的基准PDDL问题描述进行比较,用于事后基准分析;这些离线检查包括重命名不变的结构匹配和语义等价性验证。在Planetarium、AutoPlanBench以及精选的PDDL 2.1问题上的实验结果表明,运行层面的成功与基准参考重构之间可能存在显著差异。结果还表明,结构化修复是有用的,并且即使运行层面的成功率有所提升,PDDL 2.1在参考重构方面仍然具有挑战性。
cs.AI / 35 / 2609.09925

Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models

面向分块视觉-语言-动作模型的时频几何交叉注意力机制
Dong, Shengye, Niu, Haochen, Liu, Hao, Lin, Peiwen, Wang, Chuang, Pang, Shanmin
Abstract
Modern vision-language-action (VLA) policies predict a whole chunk of actions: one to two seconds of coordinated motion emitted in a single forward pass. Yet an action chunk is essentially a short multivariate trajectory, but inside these models it is a sequence of generic per-timestep hidden tokens decoded by a linear head. This under-serves two motion structures. First, frequency: a chunk superimposes a smooth global trend and fine corrective motion across time scales, and a single token entangles them. Second, cross-phase geometry: motions of different phases (reach, contact, grasp adjustment, settling) unfold along very different, near-orthogonal directions in representation space, yet are tightly related for the task and arise across the time axis. Dot-product attention scores alignment by an inner product, so it favors aligned tokens and is least sensitive near orthogonality, leaving such relationships for the network to recover through a detour. We introduce Time-Frequency Geometric Cross-Attention (TFGCA), a drop-in module repairing both blind spots. TFGCA uses a per-dimension learnable stationary wavelet transform to decompose the action chunk into time-frequency tokens, and each time token retrieves information from them via a cross-attention that fuses the dot product (similarity) with the wedge-product magnitude (sensitive to near-orthogonality) through a learnable weight. A zero-initialized residual reproduces the base behavior at initialization, so it can be dropped onto a pretrained VLA and fine-tuned jointly. Relative to the same-source base, TFGCA improves in-distribution LIBERO by +1.5 on average, the OOD LIBERO-Plus by +6.3, the randomized average under RoboTwin domain randomization by +28.5, and the overall success rate on three real-robot AgiBot A2 tasks by +11.67 points, with larger gains out of distribution.
Chinese Translation
现代视觉-语言-动作(VLA)策略会预测一整块动作:即在前向传播中一次性输出一到两秒的协调运动。然而,动作块本质上是一条短的多变量轨迹,但在这些模型内部,它却被表示为一系列通用的逐时间步隐藏令牌,并由线性头进行解码。这种表示方式无法充分捕捉两种运动结构。其一是频率结构:一个动作块在不同时间尺度上叠加了平滑的全局趋势和精细的修正运动,而单一令牌会将它们纠缠在一起。其二是跨相位几何结构:不同动作相位(伸臂、接触、抓取调整、稳定)的运动在表示空间中沿着差异极大、近乎正交的方向展开,但它们在任务层面却紧密相关,并沿时间轴出现。点积注意力通过内积来衡量对齐程度,因此偏向于对齐的令牌,在近乎正交的情形下敏感度最低,使得此类关系只能由网络绕道间接恢复。我们提出了时频几何交叉注意力(TFGCA),这是一个即插即用的模块,可同时修复上述两个盲点。TFGCA 使用逐维可学习的平稳小波变换将动作块分解为时频令牌,每个时间令牌通过一种将点积(相似性)与楔积幅值(对近正交情形敏感)以可学习权重相融合的交叉注意力机制从中检索信息。零初始化的残差连接保证了在初始化时复现基础行为,因此该模块可以直接加载到预训练的 VLA 模型上进行联合微调。相对于同一底座模型,TFGCA 在分布内 LIBERO 上平均提升 +1.5,在分布外 LIBERO-Plus 上提升 +6.3,在 RoboTwin 域随机化下的随机化平均成功率上提升 +28.5,并在三个真实机器人 AgiBot A2 任务上的总体成功率上提升 +11.67 个百分点,且在分布外场景中增益更大。
cs.AI / 36 / 2609.09928

Structural Process Supervision for Latent Chain-of-Thought Reasoning

面向潜在思维链推理的结构化过程监督
Li, Yiqi, Chen, Xu, Ju, Chen, Yao, Jiangchao, Li, Zhaoyang, Lan, Jinsong, Zhu, Xiaoyong, Zheng, Bo, Wang, Yu
Abstract
Latent reasoning approaches enhance token-level efficiency and robustness by replacing verbose, explicit chain-of-thought (CoT) tokens with compact continuous-space embeddings. However, existing methods lack direct process supervision over these latent embeddings, which often leads to representation collapse and uneven information distribution. To address this, we propose Prototype-Mediated Process Supervision (PMPS), which introduces learnable reasoning prototypes as semantic anchors to provide structural process-level supervision for latent reasoning. PMPS projects latent embeddings and explicit CoT embeddings into a shared prototype space, achieving many-to-many soft alignment between unequal-length representations through prototype assignment. Meanwhile, we introduce a Progressive Sequential Alignment (PSA) module to further guide training: positional priors initially encourage sequential alignment structure, then gradually relax to permit adaptive matching. Experimental results show that PMPS compresses output token length to under 50% of explicit CoT on GSM8K-Aug. Compared to leading baseline SIM-CoT, our method achieves average accuracy gains of 2.08% across different model families. On GPT-2, PMPS even surpasses CoT-SFT. On larger models and a more challenging task, PMPS consistently attains the highest accuracy among all latent reasoning methods with comparable output length.
Chinese Translation
潜在推理方法通过用紧凑的连续空间嵌入替代冗长显式的思维链(Chain-of-Thought, CoT)标记,提升了token层面的效率与鲁棒性。然而,现有方法缺乏对这些潜在嵌入的直接过程监督,这常常导致表征坍缩和信息分布不均。为解决这一问题,我们提出了原型介导过程监督(Prototype-Mediated Process Supervision, PMPS),该方法引入可学习的推理原型作为语义锚点,为潜在推理提供结构化的过程级监督。PMPS将潜在嵌入与显式CoT嵌入投影到一个共享的原型空间中,通过原型分配实现不等长表征之间的多对多软对齐。同时,我们引入渐进式序列对齐(Progressive Sequential Alignment, PSA)模块以进一步引导训练:位置先验在初期鼓励序列对齐结构,随后逐渐放松以允许自适应匹配。实验结果表明,PMPS在GSM8K-Aug上将输出token长度压缩至显式CoT的50%以下。与领先的基线方法SIM-CoT相比,我们的方法在不同模型家族上平均准确率提升2.08%。在GPT-2上,PMPS甚至超越了CoT-SFT。在更大的模型和更具挑战性的任务上,在相近输出长度的条件下,PMPS在所有潜在推理方法中始终取得最高的准确率。
cs.AI / 37 / 2609.10036

Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability

信念状态引擎:增强大语言模型在部分可观测条件下的原则性规划能力
Chattopadhayay, Arnab, Halder, Debdipta
Abstract
Large language model agents produce fluent action sequences across a wide range of tasks, yet they fail in characteristic ways once the environment becomes partially observable. Ambiguous feedback pushes them into premature commitments. A single informative observation can collapse their uncertainty onto the wrong hypothesis. Policies drift as the history grows. We trace these symptoms to a common structural cause. An LLM agent, as commonly deployed, is a history-conditioned policy with no explicit belief over hidden state. We propose an architectural fix. The Belief-State Engine (BSE) is an inference module placed outside the LLM. It maintains a Bayesian posterior over the latent states of a given POMDP (Partially Observable Markov Decision Process) model, and at each decision step it exposes only that posterior to the LLM. The raw action-observation log is not shown. We set out a minimal four-axiom specification of what a belief-consistent internal state must satisfy, and prove that the LLM paired with the BSE is a sound Markov policy on the belief MDP induced by the underlying POMDP. It therefore inherits the Bellman optimality guarantees of classical POMDP theory, provided the LLM is never exposed to the raw history. We evaluate the architecture on the Tiger POMDP and a red-team attack-graph task, against six baselines: a reactive LLM, Chain-of-Thought, ReAct, a natural-language belief tracker, QMDP, and POMCP. Across both domains, the BSE-augmented agent improves task return, belief calibration, and decision consistency. Ten targeted ablations isolate the contribution of each architectural choice confirms that the effect is not specific to any one model. Code, environment specifications, prompt templates, and seed logs accompany this paper.
Chinese Translation
大语言模型智能体能够在广泛的任务中生成流畅的动作序列,然而一旦环境变得部分可观测,它们便会以特有方式失效。模糊的反馈会使它们过早地做出承诺;单条信息量大的观测可能使其不确定性错误地坍缩到某个假设上;随着历史记录的增长,策略会发生漂移。我们将这些症状追溯到一个共同的结构性原因:按常规方式部署的大语言模型智能体是一个以历史为条件的策略,对隐藏状态没有显式的信念表示。我们提出一种架构层面的修复方案:信念状态引擎(Belief-State Engine,BSE)是一个置于大语言模型之外的推理模块,它针对给定的POMDP(部分可观测马尔可夫决策过程)模型维护关于潜在状态的贝叶斯后验分布,并在每个决策步骤中仅将该后验暴露给大语言模型,而不展示原始的动作-观测日志。我们给出了信念一致的内部状态所必须满足的最小化四公理规范,并证明与大语言模型配对的BSE在由底层POMDP诱导出的信念MDP上是一个合理的马尔可夫策略。因此,只要大语言模型从不接触原始历史记录,该策略便继承经典POMDP理论的Bellman最优性保证。我们在Tiger POMDP和一个红队攻击图任务上,将所提出的架构与六种基线方法进行评估比较,包括反应式大语言模型、思维链(Chain-of-Thought)、ReAct、自然语言信念跟踪器、QMDP和POMCP。在两个领域中,BSE增强的智能体均在任务回报、信念校准和决策一致性方面有所提升。十项针对性的消融实验分离出每个架构选择的贡献,并证实该效果并非特定于某一模型。本文随附代码、环境规范、提示词模板和随机种子日志。
cs.AI / 38 / 2609.10055

OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization

OntologyAligner:面向生物医学本体规范化的本体对齐检索与层次引导大语言模型重排序方法
Song, Jie, Xu, Zhichuan, Lu, Ziyu, Xiao, Meng, Bi, Cheng, Zhang, Yuxin, Zheng, Xin, Li, Xiaoran, Cao, Qiongfang, Yang, Hao, Shen, Bairong
Abstract
Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical variation and subtle distinctions among hierarchically related concepts can obscure concept boundaries. We present OntologyAligner, a three-stage framework that combines ontology-aligned retrieval, large language model candidate reranking, and selective hierarchy-guided refinement. We also construct PhenoNormBench, a unified benchmark comprising 13,390 samples from seven Human Phenotype Ontology datasets. OntologyAligner achieved state-of-the-art performance on HPO normalization, with 88.78% Macro Top-1 Accuracy and 86.75% Micro Top-1 Accuracy, exceeding the strongest baseline by 4.85 and 5.07 percentage points, respectively. Ablation analyses showed complementary contributions from all three stages, and sensitivity analyses demonstrated stability across candidate-set sizes and model backbones. Applications to MONDO, MEDIC, and NCBITaxon further established portability to other ontologies. OntologyAligner offers a generalizable framework for accurate mapping of biomedical text to structured ontology concepts. PhenoNormBench and the code are publicly available at https://github.com/zhelishisongjie/OntologyAligner.
Chinese Translation
生物医学本体规范化旨在将自由文本表述映射到标准化概念上,从而实现生物医学数据的一致性整合与分析。由于词汇变化以及层级相关概念之间的细微差别可能模糊概念边界,该任务仍具挑战性。我们提出了OntologyAligner,这是一个三阶段框架,结合了本体对齐检索、大语言模型候选重排序以及选择性层次引导精炼。我们还构建了PhenoNormBench,这是一个统一基准,包含来自七个人类表型本体(Human Phenotype Ontology)数据集的13,390个样本。OntologyAligner在HPO规范化任务上取得了最先进的性能,宏平均Top-1准确率达88.78%,微平均Top-1准确率达86.75%,分别超过最强基线4.85和5.07个百分点。消融分析表明三个阶段均做出了互补性贡献,敏感性分析证明了该框架在不同候选集规模和模型骨干下均具有稳定性。在MONDO、MEDIC和NCBITaxon上的应用进一步证明了该方法向其他本体的可移植性。OntologyAligner为生物医学文本到结构化本体概念的精确映射提供了一个可泛化的框架。PhenoNormBench和代码已公开于 https://github.com/zhelishisongjie/OntologyAligner。
cs.AI / 39 / 2609.10060

Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States

基于参照的LLM偏见检测:利用隐状态的相对表示
Jeliński, Marek, Dubiński, Jan, Chrabaszcz, Maciej, Cygert, Sebastian
Abstract
Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated text. We propose a reference-based method that audits bias in hidden-state representations across related model variants, for example before and after fine-tuning. Because fine-tuning reshapes representation geometry, absolute hidden states are not directly comparable, so we encode each sentence by its similarities to a fixed set of anchor sentences, yielding relative representations in a shared comparison space. There we measure how target groups shift in their association with positive and negative attributes, a quantity we call the Representational Bias Shift $\Delta B$. Across three model families and the WildGuardMix, DecodingTrust and ToxiGen benchmarks, $\Delta B$ correlates with output-level bias change in 15 of the 18 settings we test, reaching $|r| = 0.84$ ($p < 0.001$) under full fine-tuning and becoming more model-dependent under parameter-efficient adaptation. Thresholding $\Delta B$ detects checkpoints whose bias increased with ROC AUC between $0.65$ and $0.99$, and on WildGuardMix and DecodingTrust it separates them better than a SEAT-based baseline for all three families. $\Delta B$ is also stable under changes to the anchor set, attribute sets and target templates. Our method requires no task-specific evaluation data and audits a model in about three minutes, using $3$-$50\times$ less compute than the output-level benchmarks considered here. We view it as complementary to output-based auditing rather than a replacement for it.
Chinese Translation
现有的偏见审计方法通常依赖于模型输出,需要昂贵的基准测试或评判模型,且可能遗漏从未出现在生成文本中的内部偏移。我们提出了一种基于参照的方法,用于审计相关模型变体(例如微调前后)隐状态表示中的偏见。由于微调会重塑表示几何结构,绝对隐状态无法直接比较,因此我们通过每个句子与一组固定锚定句子的相似度对其进行编码,从而在共享的比较空间中得到相对表示。在该空间中,我们度量目标群体与积极和消极属性关联的变化,并将其称为表示性偏见偏移(Representational Bias Shift)ΔB。在三个模型系列以及WildGuardMix、DecodingTrust和ToxiGen基准上,ΔB在我们测试的18个设置中的15个与输出层面的偏见变化相关,在完全微调下达到|r| = 0.84(p < 0.001),而在参数高效微调下则更依赖于具体模型。通过对ΔB设置阈值,可以检测偏见增加的检查点,ROC AUC在0.65至0.99之间;在WildGuardMix和DecodingTrust上,其对三个模型系列的区分效果均优于基于SEAT的基线方法。ΔB对锚定集、属性集和目标模板的变化也具有稳定性。我们的方法不需要特定任务的评估数据,约三分钟即可完成一次模型审计,所使用的计算量比本文考虑的输出层面基准少3至50倍。我们认为该方法是对基于输出的审计的补充,而非替代。
cs.AI / 40 / 2609.10092

RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases

RAP:研究注意力预测揭示目标条件下的证据获取偏差
Wu, Yingqian, Liang, Jingcong, Wang, Siyuan, Yin, Zhenfei, Torr, Philip, Yu, Junchi, Wei, Zhongyu
Abstract
Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research ideas lack uniquely verifiable outcomes. We introduce Research Attention Prediction (RAP), a rolling benchmark covering 278 AI/ML fields and 1,390 episodes. At each cut-off, an LLM agent searches a temporally restricted arXiv corpus and predicts the next six months' paper shares across eight frozen research directions. Search generally helps, but all four diagnostic models perform worse than an exact-count exponentially weighted moving average (EWMA) baseline in compositional accuracy. We identify two linked bottlenecks. Under cumulative-history access, State carry-forward outperforms direct Forecast for all four diagnostic models; frozen-evidence replay links a shared component of this reversal to Forecast-oriented policies retrieving a smaller share of recent evidence. Even with exact historical activity, future-specific updating remains limited, with only GPT-5.5 plus reopened Search slightly surpassing EWMA. Fine-tuning on realised outcomes improves Qwen3-4B's forecast Spearman correlation by 0.105 on held-out fields at later origins, with gains also on change-rich episodes.
Chinese Translation
大语言模型(LLM)越来越多地扮演研究智能体的角色,但由于综述和研究想法缺乏可唯一验证的结果,其追踪研究注意力变化的能力难以评估。我们提出了研究注意力预测(Research Attention Prediction,RAP),这是一个滚动基准,涵盖278个AI/ML领域和1,390个回合。在每个时间切点,LLM智能体在一个时间受限的arXiv语料库中进行检索,并预测未来六个月内八个固定研究方向上的论文占比。检索通常有所帮助,但在组合准确性方面,所有四个诊断模型的表现均劣于基于精确计数的指数加权移动平均(EWMA)基线。我们识别出两个相互关联的瓶颈。在可访问完整累积历史的条件下,状态前递(State carry-forward)在所有四个诊断模型上均优于直接预测(Forecast);冻结证据重放实验将这一逆转的共同成分与偏向预测(Forecast)的策略检索近期证据比例较小相关联。即使拥有精确的历史活动信息,面向未来的更新能力仍然有限,只有GPT-5.5加上重新开放的检索略微超越EWMA。在已实现结果上进行微调,使Qwen3-4B在后期时间起点及留出领域上的预测Spearman相关性提高了0.105,并在变化丰富的回合上同样取得提升。
cs.AI / 41 / 2609.10135

Agent-Based ML-LLM Fusion with Self-Optimizing Prompts for Plateau Weather Alerts

基于智能体的机器学习与大语言模型融合的自优化提示方法用于高原气象预警
Yan, Shuai, Xu, Yang, He, Shan
Abstract
To address insufficient contextualization, weak generalization, and poor scenario adaptation in tourism meteorological services, we propose SmartWeatherAgent--a unified three-stage architecture integrating intent recognition, hazard prediction, and reasoning-enhanced generation. The system fuses rule-based methods with large language models to parse queries at multiple granularities and employs a LightGBM model enriched with highland-specific features (e.g., wind speed abruptness rate), achieving an F1-Macro score of 0.605 with 1.60 ms latency on high-wind, precipitation, and low-temperature events. A 12-round micro-step prompt self-optimization loop boosts the composite warning quality score S_final from 4.2 (B01) to 8.9 (B12, +112%). Key improvements include a sharp rise in B08 from data source citation (6.5 -> 8.5), sustained high performance in B10 via physical mechanism explanation, and a peak scientific rigor score of 9.2 in B12 through explicit uncertainty statements. The system autonomously generates structured warnings that integrate causal mechanisms, spatiotemporal evolution, quantitative evidence, regulatory references, and confidence statements--enhancing professional depth, logical rigor, and scientific soundness, and advancing meteorological services toward proactive perception, explainable decision-making, and intelligent agency.
Chinese Translation
针对旅游气象服务中情境化不足、泛化能力弱、场景适应性差等问题,我们提出了SmartWeatherAgent——一个统一的三阶段架构,融合意图识别、灾害预测和推理增强生成。该系统将基于规则的方法与大语言模型相结合,实现多粒度查询解析,并采用融合高原特有特征(如风速骤变率)的LightGBM模型,在大风、降水和低温事件上取得F1-Macro分数0.605、延迟仅1.60 ms的成绩。经过12轮微步提示自优化循环,综合预警质量分数S_final从4.2(B01)提升至8.9(B12,+112%)。主要改进包括:数据来源引用使B08得分大幅提升(6.5→8.5),物理机制解释使B10保持高水平表现,以及通过明确的不确定性表述使B12的科学严谨性得分达到峰值9.2。该系统能够自主生成融合因果机制、时空演变、定量证据、法规依据和置信度说明的结构化预警——提升了专业深度、逻辑严谨性和科学合理性,推动气象服务向主动感知、可解释决策和智能化方向发展。
cs.AI / 42 / 2609.10144

Kernel-Managed Shared Memory for System-Wide Personalization

面向系统级个性化的内核管理共享内存
Lum, Ryan, Zhang, Yongfeng
Abstract
AI systems become more useful when they can adapt to the people using them, but in multi-agent systems, useful context learned by one agent often remains unavailable to others. We present kernel-managed shared memory, a system-level abstraction in which specialized agents write structured, tagged memories while the agent-system kernel, not individual agents, governs retrieval, privacy enforcement, and prompt injection. We implement and evaluate this design on AIOS and compare it against three alternatives across three assistant models (GPT-4o, Llama-3.1:8B, Qwen-2.5:7B) and 1,800 total trials. Against an unmanaged external memory backend (Mem0) using identical underlying storage, kernel-managed retrieval and injection improve personalization scores by 2.4-4.0 points on a 5-point scale (e.g., 1.05 to 4.69 profile usage on GPT-4o), with every comparison significant at p < 10^-18. Against standard retrieval-augmented injection, gains are similarly large and consistent across all three models. Against full, unfiltered context concatenation, a soft ceiling on available context rather than on response quality, kernel-managed injection statistically matches performance on two of three models and shows a small, model-specific deficit on the third, while using substantially shorter prompts: end-to-end latency is 15-61% lower across all three models, with corresponding reductions in per-call token usage and inference cost. These results indicate that centralizing memory management in the agent-system kernel, rather than leaving retrieval and privacy enforcement to individual agents, delivers most of the personalization benefit of unconstrained context at a fraction of its cost.
Chinese Translation
当AI系统能够适应使用它们的用户时,其效用会显著提升,但在多智能体系统中,一个智能体学到的有用上下文往往无法被其他智能体获取。我们提出了内核管理的共享内存(kernel-managed shared memory),这是一种系统级抽象:专门的智能体写入结构化、带标签的记忆,而由智能体系统内核(而非单个智能体)负责管理检索、隐私执行和提示注入。我们在AIOS上实现并评估了这一设计,并在三个助手模型(GPT-4o、Llama-3.1:8B、Qwen-2.5:7B)上进行了共计1,800次试验,与三种替代方案进行了比较。与使用相同底层存储但未经管理的外部记忆后端(Mem0)相比,内核管理的检索和注入在5分制个性化评分上提升了2.4至4.0分(例如,GPT-4o上的画像利用率从1.05提升至4.69),且所有比较均在p < 10^-18水平上显著。与标准的检索增强注入相比,收益同样显著,且在所有三个模型上保持一致。与完整、未过滤的上下文拼接(这可视为可用上下文而非响应质量的软上限)相比,内核管理的注入在三个模型中的两个上达到了统计上的同等性能,在第三个模型上表现出较小的模型特异性差距,同时使用的提示长度大幅缩短:所有三个模型的端到端延迟降低了15%至61%,每次调用的token使用量和推理成本也相应减少。这些结果表明,将内存管理集中于智能体系统内核,而非将检索和隐私执行交由单个智能体处理,能够以极小的成本实现无限制上下文所带来的大部分个性化收益。
cs.AI / 43 / 2609.10177

Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning

超越表面模仿:面向多模态上下文学习中推理路径对齐的对比建模方法
Yang, Mingbo, Wang, Wenqiang, Kang, Zhaolu, Chen, Peng, Chen, Yannan, Wang, Sunshang, Xiao, Yan
Abstract
In-context learning (ICL) is widely used in multimodal large language models (MLLMs) and achieves strong performance across a wide range of multimodal tasks. However, existing multimodal ICL methods often rely on surface level imitation of in-context demonstrations, making it difficult for MLLMs to align their responses with the reasoning path required by the given multimodal input. This limitation becomes more pronounced in complex multimodal tasks, thereby restricting further improvements in MLLM performance. To address this issue, we propose a new multimodal ICL framework that combines contrastive demonstration modeling with the self-refinement capability of MLLMs. Specifically, our framework reformulates each demonstration by explicitly contrasting a suboptimal response with a better response under the same input, together with a reasoning path that reveals how the response should be refined. This contrastive formulation makes the reasoning path toward the desired response more explicit and guides the MLLM beyond superficial imitation. Furthermore, because effective refinement depends on the current response, we introduce a response-conditioned retrieval mechanism to select demonstrations whose reasoning paths are more relevant to the current response. In addition, we use a lightweight alignment controller to predict response quality and determine whether further refinement is needed. Experiments on three types of multimodal tasks show that the proposed framework consistently improves MLLM performance, with particularly notable gains on visual question answering (VQA).
Chinese Translation
上下文学习(In-Context Learning, ICL)被广泛应用于多模态大语言模型(Multimodal Large Language Models, MLLMs),并在众多多模态任务中取得了优异性能。然而,现有多模态ICL方法往往依赖于对上下文示例的表面级模仿,使得MLLM难以将其响应与给定多模态输入所需的推理路径对齐。这一局限在复杂多模态任务中尤为明显,从而制约了MLLM性能的进一步提升。为解决这一问题,我们提出了一种新的多模态ICL框架,该框架将对比示例建模与MLLM的自完善能力相结合。具体而言,我们的框架通过在相同输入下显式对比一个次优响应与一个更优响应,并附上揭示响应应如何完善的推理路径,来重新构建每个示例。这种对比性表述使得通往期望响应的推理路径更加明确,引导MLLM超越表面模仿。此外,由于有效的完善依赖于当前响应,我们引入了一种基于响应条件的检索机制,以选择其推理路径与当前响应更相关的示例。同时,我们使用一个轻量级的对齐控制器来预测响应质量,并判断是否需要进一步完善。在三类多模态任务上的实验表明,所提出的框架能够持续提升MLLM的性能,在视觉问答(Visual Question Answering, VQA)任务上的提升尤为显著。
cs.AI / 44 / 2609.10221

Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection

可枚举之物为何还要采样?基因组工具选择的精确策略优化
Liu, Haoyue, Ma, Xiaoyu, Chen, Ye, Wang, Zhichao, Tang, Xiaoying
Abstract
Reinforcement learning over a frozen reasoner has become a common recipe for teaching a policy which external tools to invoke. We show that this recipe becomes structurally mismatched in specialist scientific settings where the complete tool-subset space is enumerable. There, a small set of recurring computational capabilities covers the domain, so the space of tool subsets is combinatorial yet small enough to enumerate, and GRPO still estimates an action expectation from a handful of sampled rollouts. Worse, the approximation degrades as training succeeds: as the policy concentrates on preferred subsets it resamples them, sampled rewards collide, and the group-normalized advantage vanishes. On genomic reasoning the fraction of questions yielding no reward signal rises from 0.2% under a uniform reference policy to 20.8% after GRPO training. As a remedy, we introduce FGPO (Full-Group Policy Optimization), which (1) scores every tool subset and optimizes the exact action expectation, so each update sees the complete action space, and (2) precomputes the reward of each question--subset pair into an exhaustive table, removing frozen-reasoner calls from the training loop entirely. Across five frozen reasoners and three genomic benchmarks, FGPO outperforms GRPO in all 15 settings by 6.75 points on average and up to 14.20, while a standard on-demand GRPO schedule would require 2.4 times as many frozen-reasoner reward evaluations and, on GenomeQA, FGPO cuts invoked tools per question from 2.36 to 1.40.
Chinese Translation
基于冻结推理机的强化学习已成为训练策略选择外部工具的常用方法。我们证明,在完整工具子集空间可枚举的专业科学场景中,这种方法存在结构性错配。在这些场景中,少量反复出现的计算能力即可覆盖整个领域,因此工具子集空间虽是组合性的,但规模小到可以完全枚举,而GRPO仍然仅从少量采样 rollout 中估计动作期望。更糟的是,这种近似会随着训练的成功而退化:当策略集中于偏好的子集时,它会反复采样这些子集,导致采样奖励相互碰撞,组归一化优势随之消失。在基因组推理任务上,无法产生奖励信号的问题比例从均匀参考策略下的0.2%上升到GRPO训练后的20.8%。作为解决方案,我们提出了FGPO(Full-Group Policy Optimization,全组策略优化),它(1)对每个工具子集进行评分并优化精确的动作期望,使每次更新都能看到完整的动作空间;(2)将每个问题—子集对的奖励预先计算成穷举表,从而完全从训练循环中移除冻结推理机的调用。在五个冻结推理机和三个基因组基准上,FGPO在全部15个设置中均优于GRPO,平均提升6.75分,最高提升14.20分;而标准按需调用的GRPO方案需要2.4倍的冻结推理机奖励评估次数。此外,在GenomeQA上,FGPO将每个问题调用的工具数量从2.36降至1.40。
cs.AI / 45 / 2609.10263

What Should an Agent Forget? Separating What Is Stored from What Is Used

智能体应该遗忘什么?区分所存储的内容与所使用的内容
Li, Yuhang, Li, Yuchen
Abstract
Persistent language agents need stored experience to remain available across time, while each answer requires evidence suited to a particular question. A superseded fact can mislead a current-state answer and still be essential for a historical query. We present RD-Forget, a training-free framework that separates what an agent stores from what it uses. A retained source archive preserves observations, and a query-conditioned memory view controls their influence on the current answer. A frozen language-model curator extracts relevant evidence, groups facts into semantic slots, and preserves the relations needed for multi-hop reasoning. Same-slot replacement links suppress superseded values in current-state contexts, while intent-aware retrieval makes earlier evidence eligible again. A rate-distortion formulation guides construction of the answer-time view within a memory budget. Experiments span conversational memory, knowledge updating, fact consolidation, long-context reasoning, and personalization under a shared answering pipeline. The results associate accurate answers with both query-relevant evidence construction and control over obsolete alternatives. Configurations without forgetting or query conditioning have the largest score deficits, while slot grouping, historical access, and relation preservation contribute complementary functions. Retaining history while selectively controlling its use offers a practical way to accommodate changing facts and future questions.
Chinese Translation
持久化的语言智能体需要存储的经验能够跨时间持续可用,而每一次回答又需要与特定问题相匹配的证据。一个已被取代的事实可能会误导关于当前状态的回答,却仍对历史查询至关重要。我们提出了RD-Forget,一个无需训练的框架,它将智能体所存储的内容与其所使用的内容分离开来。保留的源档案(source archive)保存观察记录,而基于查询条件的记忆视图(query-conditioned memory view)控制它们对当前回答的影响。一个冻结的语言模型策展器提取相关证据,将事实组织成语义槽(semantic slots),并保留多跳推理所需的关系。同槽替换链接在当前状态语境下抑制已被取代的取值,而意图感知检索则使较早的证据重新可用。率失真(rate-distortion)形式化方法在记忆预算内指导回答时视图的构建。实验在统一的回答流程下涵盖了对话记忆、知识更新、事实整合、长上下文推理和个性化任务。结果表明,准确的回答既依赖于与查询相关的证据构建,也依赖于对过时候选信息的控制。不进行遗忘或不采用查询条件化的配置得分差距最大,而槽分组、历史访问和关系保留则提供了互补的功能。在保留历史的同时选择性地控制其使用,为适应不断变化的事实和未来的问题提供了一种切实可行的方法。
cs.AI / 46 / 2609.10315

TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards

TRACE:基于合成奖励训练因果探索推理智能体
Sun, Rui, Shi, Zhan, He, Bing
Abstract
Reinforcement learning with verifiable rewards (RLVR) has advanced language-model reasoning in domains such as mathematics and code, where objective answers are inexpensive to check. Diagnostic reasoning over complex data lacks this advantage: establishing the true cause of an anomaly often requires costly expert investigation and may remain ambiguous after the fact. We ask whether this asymmetry of verification can instead be engineered. We sample an intervention, inject it into a controlled simulator, and generate the observations it would produce. The hidden intervention provides an oracle label and objective reward, while the agent must still investigate noisy, confounded, and distributed evidence. We instantiate this approach in TRACE, a digital-advertising diagnostic environment with 12 root causes and fine-grained segment attribution. Agents investigate each episode using Python and SQL and must identify both the root cause and, when applicable, the affected segment assignment. On a held-out 235-episode test set, the strongest prompted baseline, Claude Opus 5, reaches 0.686 FullAttr@1. Supervised fine-tuning raises Qwen3.5-35B-A3B from 0.159 to 0.637, and subsequent RL with synthesized rewards reaches 0.757, outperforming all evaluated prompted baselines, including frontier closed-source models and a prompted Qwen3.5-122B-A10B model. The resulting policy also uses substantially fewer tool calls than the prompted 35B base. These results provide evidence that access to a scalable, objective training signal can be a more important constraint than model scale alone. More broadly, simulation-based verification can make otherwise ambiguous diagnostic reasoning tasks amenable to scalable reinforcement learning.
Chinese Translation
可验证奖励强化学习(RLVR)已推动了语言模型在数学和代码等领域的推理能力,因为这些领域的客观答案检验成本低廉。而对复杂数据的诊断推理则缺乏这一优势:确定异常的真正原因往往需要代价高昂的专家调查,且事后仍可能存在歧义。我们提出一个问题:这种验证的不对称性能否被人为构造出来?我们采样一个干预操作,将其注入受控模拟器中,并生成其会产生观测数据。隐藏的干预提供了“神谕”标签和客观奖励,而智能体仍需调查带噪、混杂且分散的证据。我们在 TRACE 中实现了这一方法,这是一个包含12种根因和细粒度细分归因的数字广告诊断环境。智能体在每个回合中使用 Python 和 SQL 进行调查,必须识别根因,并在适用时识别受影响的细分分配。在一个包含235个回合的留存测试集上,最强的提示基线 Claude Opus 5 达到了 0.686 的 FullAttr@1。监督微调将 Qwen3.5-35B-A3B 的成绩从 0.159 提升至 0.637,随后使用合成奖励的强化学习进一步达到 0.757,超越了所有被评估的提示基线,包括前沿闭源模型以及提示版的 Qwen3.5-122B-A10B 模型。此外,所得到的策略使用的工具调用次数也远少于提示版35B基座模型。这些结果表明,获得可扩展、客观的训练信号可能是比模型规模更重要的约束条件。更广泛地说,基于模拟的验证方法可以使原本含糊的诊断推理任务适用于可扩展的强化学习。
cs.AI / 47 / 2609.10335

From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning

从符号感知到逻辑推演:一个引导语言模型进行几何推理的框架
Dai, Weichen, Cabral, Rafael Medeiros, Shou, Ziyi, Cao, Yan, Shen, Xin, Lu, Dongcai, Zhou, Yi
Abstract
Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally intensive and opaque. We demonstrate that a pure Large Language Model (LLM), when equipped with specialized modules, can rival state-of-the-art LMMs on complex geometry problems. Our framework integrates a Geometric Vision Parser, which translates diagrams into symbolic form, with a Symbolic Solver that performs formal deductions, thereby mitigating hallucinations and promoting interpretable reasoning. To enable rigorous evaluation, we curate a benchmark of challenging problems from the 2025 Chinese Zhongkao examinations, ensuring data novelty and testing deeper deductive skills. Experiments demonstrate that our approach achieves performance comparable to Gemini 2.5 Pro while delivering clearer, human-like solutions.
Chinese Translation
平面几何仍是人工智能领域的一项重大挑战,它要求将视觉感知与数学推理相融合。尽管大型多模态模型(LMM)能够自然地处理视觉-语言输入,但它们通常计算开销大且缺乏可解释性。我们证明,纯大型语言模型(LLM)在配备专门模块的情况下,能够在复杂几何问题上与最先进的LMM相媲美。我们的框架集成了几何视觉解析器(Geometric Vision Parser),用于将图形转化为符号形式,并结合符号求解器(Symbolic Solver)执行形式化推演,从而缓解幻觉问题并促进可解释的推理。为实现严格的评估,我们精心构建了一个来自2025年中国中考的具有挑战性问题的基准数据集,确保了数据的新颖性并考验更深层的演绎推理能力。实验表明,我们的方法取得了与Gemini 2.5 Pro相当的性能,同时提供更清晰、更接近人类思维方式的解题过程。
cs.AI / 48 / 2609.10350

Cyber-Financial Contagion: Modeling the Propagation of an AI Vendor Compromise Through the Banking System

网络-金融传染:人工智能供应商入侵在银行系统中传播的建模研究
Leytes, Alex
Abstract
The banking system now depends on a small set of shared artificial intelligence vendors for fraud screening, credit decisioning, anti-money-laundering triage, customer analytics, and internal decision support. This paper studies how a compromise inside one of those vendors can propagate along a chain of operational, informational, and financial linkages until it triggers losses that look, from the outside, like a classical banking crisis. We build a four-layer heterogeneous network that couples AI vendors, financial institutions, interbank exposures, and customer accounts, and we propose CFC-Prop, a stochastic epidemic-and-clearing model that runs on that network. On a synthetic dataset with 60 vendors, 220 banks, roughly 2,500 vendor-bank service edges, and 1,400 interbank exposures, CFC-Prop reproduces the heavy-tailed loss distributions and the sharp dependence on patch latency that are consistent with prior cyber-financial evidence. We also train an early-warning model, CFC-GNN, that uses vendor-side incident telemetry and graph structure to flag high-cascade-risk vendors before impact. Across four baselines the proposed model reaches AUROC 0.82 and AUPRC 0.60 while keeping calibration errors bounded. We release the full code, synthetic data, and reproducible scripts. The results argue that cyber concentration among AI vendors is a first-order financial-stability problem and give supervisors a concrete quantitative tool for reasoning about it.
Chinese Translation
银行系统如今在欺诈筛查、信贷决策、反洗钱分流、客户分析以及内部决策支持等方面依赖于一小部分共享的人工智能供应商。本文研究了其中一家供应商遭受入侵后,如何沿着运营、信息和金融联系的链条传播,最终引发从外部看来类似于经典银行业危机的损失。我们构建了一个四层异构网络,将人工智能供应商、金融机构、银行间风险敞口和客户账户耦合起来,并提出 CFC-Prop——一种运行于该网络之上的随机流行病-清算混合模型。在包含 60 家供应商、220 家银行、约 2,500 条供应商-银行服务边和 1,400 个银行间风险敞口的合成数据集上,CFC-Prop 重现了厚尾损失分布以及对补丁延迟的急剧敏感性,这与先前的网络-金融证据一致。我们还训练了一个早期预警模型 CFC-GNN,利用供应商端的事件遥测数据和图结构,在影响发生之前识别出高级联风险的供应商。在四个基线模型的对比中,所提出的模型达到 AUROC 0.82 和 AUPRC 0.60,同时保持校准误差有界。我们公开了完整的代码、合成数据和可复现脚本。研究结果表明,人工智能供应商之间的网络集中度是一个一阶金融稳定问题,并为监管机构提供了对其进行推理的具体定量工具。
cs.AI / 49 / 2609.10413

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

幸运回忆:面向大语言模型持久一致性的本体驱动记忆生命周期管理
Mullick, Ansuman, Tüzün, Eray
Abstract
Current LLM memory systems treat all personal facts identically, so stores grow without bound while retrieval precision degrades. The core challenge is lifecycle management: which memories should persist, which should be replaced, and at what rate, conditioned on the behavioral type of each fact. Fortunate Recall (FR) is a composable policy layer that classifies personal facts into a 10+1 behavioral ontology and applies category-specific lifecycle policies (differential temporal decay, slot-key supersession, event-time validity, and category-aware retrieval routing) as deterministic functions over LLM-extracted metadata. FR-Bank, our infrastructure-independent implementation, reaches a 76.9% pass rate on LifecycleBench, a new 516-question temporal-disambiguation benchmark, ahead of Mem0, A-MEM, Memory-R1, and MemoryOS (61% to 70.5%), and 75.2% on the full LongMemEval-S under the canonical Wu et al. judge protocol, so lifecycle policies impose no measurable cost on standard retrieval. A pre-registered ablation locates the gains: replacing the typed layer with three generic lifecycle primitives leaves correctness statistically unchanged (-1.7pp, 95% CI [-6.0, +2.7]), so the generic lifecycle metadata carries the correctness advantage, while the behavioral ontology carries calibration, halving downstream confabulation (12.0% vs 24.2%, p<0.001). End-to-end, FR-Bank cuts confabulation from Mem0's 45.1% to 22.4% over answered queries and from 32.2% to 13.0% over all queries while answering more of them correctly (31.2% vs 18.6%); the ranking replicates on the open-weight Kimi K2.5. The decomposition transfers to BEAM, an independently built benchmark: 46.8% correct vs Mem0's 32.9% over 280 questions, with the ontology's benefit concentrated in contradiction resolution and saturating near seven policy clusters. The ontology, benchmark, and code are released.
Chinese Translation
当前的大语言模型(LLM)记忆系统对所有个人事实一视同仁,导致存储无限增长而检索精度不断下降。核心挑战在于生命周期管理:哪些记忆应当保留、哪些应当被替换、以及以何种速率替换,而这取决于每条事实的行为类型。Fortunate Recall(FR)是一个可组合的策略层,它将个人事实归类到包含10+1类别的行为本体中,并针对LLM提取的元数据以确定性函数的形式应用类别特定的生命周期策略(差异化时间衰减、槽键取代、事件时间有效性以及类别感知的检索路由)。FR-Bank是我们独立于基础设施的实现,在一个全新的包含516道时间消歧问题的基准LifecycleBench上达到76.9%的通过率,领先于Mem0、A-MEM、Memory-R1和MemoryOS(61%至70.5%);并且在标准的Wu et al.评判协议下,在完整的LongMemEval-S上达到75.2%,表明生命周期策略对标准检索没有可测量的代价。一项预注册的消融实验定位了增益来源:用三个通用生命周期原语替换类型化层后,正确性在统计上没有变化(-1.7个百分点,95%置信区间[-6.0, +2.7]),说明正确性优势来自通用生命周期元数据,而行为本体则带来校准效果,将下游的虚构(confabulation)减半(12.0%对比24.2%,p<0.001)。端到端来看,FR-Bank将Mem0在已回答查询上的虚构率从45.1%降至22.4%,在全部查询上从32.2%降至13.0%,同时正确回答了更多查询(31.2%对比18.6%);该排名在开源权重模型Kimi K2.5上得到复现。该分解结果可迁移到独立构建的基准BEAM上:在280道问题中正确率达46.8%,高于Mem0的32.9%,且本体的收益集中在矛盾解决方面,并在约七个策略簇附近趋于饱和。本体、基准和代码均已开源发布。
cs.AI / 50 / 2609.10441

ConvMem: Convolutional Memory for Long-Context Reasoning

ConvMem:面向长上下文推理的卷积记忆机制
Zhang, Hongming, Gu, Zhaozhen, Bai, Fengshuo, Hao, Ming, Zhang, Qingyang, Wang, Yuanyuan, Tang, Shiyang, Wang, Yanna, Xu, Bo
Abstract
While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective context by reading text in segments and iteratively updating a fixed-size memory. However, this sequential paradigm suffers from high latency and requires costly reinforcement learning (RL) training, which can lead to overfitting on specific datasets. To overcome these limitations, we propose ConvMem, a training-free, highly parallelizable framework that reformulates long-context reasoning as a hierarchical convolution. Inspired by CNNs, ConvMem treats an LLM prompted with a specific query as a convolutional kernel. This kernel summarizes text segments hierarchically, shortening the reasoning path from a linear chain into a logarithmic tree. Specifically, ConvMem integrates \textit{Configurable Strides} and \textit{Skip Connections} to ensure robust evidence capture and propagation, while employing \textit{Multi-Kernel Convolution} to decompose complex queries into disentangled semantic channels. This design not only mitigates error accumulation but also enables massive parallelization across both text segments and reasoning threads. Experiments on RULER-HotpotQA and RULER-2WikiMultiHopQA demonstrate that ConvMem outperforms training-free baselines and avoids the risk of overfitting to parametric priors often observed in RL-trained models on out-of-distribution tasks.
Chinese Translation
尽管大型语言模型(LLM)已展现出令人瞩目的能力,但由于上下文长度限制固定,它们在处理超长上下文时常常力不从心。为解决这一问题,诸如 MemAgent 等序列化方法通过分段读取文本并迭代更新固定大小的记忆来扩展有效上下文。然而,这种序列化范式存在高延迟问题,且需要代价高昂的强化学习(RL)训练,容易导致在特定数据集上过拟合。为克服这些局限,我们提出了 ConvMem,一个无需训练、高度可并行化的框架,它将长上下文推理重新表述为一种层次化卷积。受卷积神经网络(CNN)启发,ConvMem 将被特定查询提示的 LLM 视为一个卷积核。该卷积核对文本片段进行层次化摘要,将推理路径从线性链缩短为对数级树结构。具体而言,ConvMem 融合了可配置步长(Configurable Strides)与跳跃连接(Skip Connections)以确保鲁棒的证据捕获与传播,并采用多核卷积(Multi-Kernel Convolution)将复杂查询分解为解耦的语义通道。这一设计不仅缓解了误差累积问题,还实现了在文本片段与推理线程两个维度上的大规模并行化。在 RULER-HotpotQA 和 RULER-2WikiMultiHopQA 上的实验表明,ConvMem 优于无需训练的基线方法,并避免了 RL 训练模型在分布外任务中常见的对参数化先验过拟合的风险。
cs.AI / 51 / 2609.10451

JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition

JarvisGUI:迈向具有动态任务组合能力的跨设备GUI智能体
Chen, Zixiang, Lu, Yuheng, Cheng, Zihao, Liu, Zeming, Bai, Jizeng, Huang, Ziye, Lin, Zhiyin, Li, Zihan, Guo, Yuhang, Wang, Yunhong, Wang, Haifeng
Abstract
Real-world GUI usage frequently involves workflows that span multiple devices and platforms, requiring the transfer of intermediate results, maintenance of shared state, and coordination across heterogeneous environments. However, existing GUI benchmarks overwhelmingly evaluate agents on single-device, statically defined tasks, thus leaving such cross-device capabilities largely unexamined, resulting in an overly optimistic assessment of agents' readiness for real-world usage. We introduce JarvisGUI, a dynamic benchmark that evaluates GUI agents on cross-device workflows requiring coordinated interaction across heterogeneous platforms, including Android, Windows, and Ubuntu. Specifically, JarvisGUI formulates GUI tasks as input-output transformations under a lightweight type system, which allows us to automatically compose multi-step, cross-device workflows and dynamically evaluate agent performance within a unified framework. By evaluating agents in virtual environments spanning multiple operating systems, JarvisGUI reveals that state-of-the-art open-source GUI agents struggle with the state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management required for real-world workflows, exposing a critical capability gap invisible to existing benchmarks.
Chinese Translation
现实世界的GUI使用经常涉及跨多个设备和平台的工作流程,需要传递中间结果、维护共享状态,并在异构环境之间进行协调。然而,现有的GUI基准测试绝大多数在单一设备上、静态定义的任务中评估智能体,因此此类跨设备能力在很大程度上未被检验,导致对智能体真实使用准备的评估过于乐观。我们提出了JarvisGUI,这是一个动态基准,用于评估GUI智能体在跨设备工作流程中的表现,这些工作流程需要在异构平台(包括Android、Windows和Ubuntu)之间进行协调交互。具体而言,JarvisGUI将GUI任务形式化为轻量级类型系统下的输入-输出转换,这使我们能够自动组合多步骤、跨设备的工作流程,并在统一框架内动态评估智能体的性能。通过在跨越多个操作系统的虚拟环境中评估智能体,JarvisGUI揭示了最先进的开源GUI智能体在状态传递感知、跨平台上下文推理以及真实工作流程所需的长时程依赖管理方面存在困难,暴露了现有基准测试无法察觉的关键能力差距。
机器学习 (Machine Learning)
79
cs.LG / 1 / 2609.09159

Spectral origin of the topological gap exponent d + {\eta}: mechanism, kernel, decomposition, and scope

拓扑间隙指数 d + η 的谱起源:机制、核函数、分解与适用范围
Loftus, Matthew
Abstract
The topological gap $\Delta$ -- the excess $H_1$ total persistence of a critical point cloud over a density-matched null -- scales as $\Delta \sim L^{d+\eta}$. We derive this analytically: the spectral integral $I(\alpha) = \sum_{k\neq 0} S_{\mathrm{conn}}(k)\,|k|^{\alpha}$ scales as $L^{2-\alpha-\eta}$ when IR-dominated, giving $I(-2\eta) \sim L^{d+\eta}$. The decomposition $I(-2\eta) = I_0 \cdot I_{\mathrm{shape}}$ separates volume ($I_0 \propto N(1-m^2) \sim L^d$) from anomalous dimension ($I_{\mathrm{shape}} \sim L^{\eta}$); the volume factor accounts for the magnetization-driven per-configuration variance of $\Delta$. We prove the mechanism requires $d < 2 + \eta$ (IR dominance), confining it to $d = 2$ for physical systems; in $d = 3$ the spectral integral is UV-dominated, explaining why density normalization is needed. An $\alpha$-sweep for Potts $q = 4$ at $L = 32$--$256$ finds $\alpha_{\mathrm{opt}}$ in $[-0.75, -0.5]$, consistent with $-2\eta_{\mathrm{Ising}}$ and inconsistent with $-2\eta_{q=4} = -1$; we flag this as tentative pending $L \geq 1024$ confirmation. The $\langle m^2 \cdot I(-2\eta)\rangle$ hyperscaling product is dominated by the correlation $r(m^2, I) \approx -0.98$ via the shared $I_0$ amplitude, so we report it as a covariance-correction analysis. Under a heuristic argument extending Divol--Polonik to inhomogeneous Poisson intensities, the bare PH kernel is flat; the effective kernel acquires $k$-dependence only at criticality. The per-configuration agreement between $\Delta$ and $I(-0.5)$ is primarily a magnetization correlation: $R^2 = 0.91$ at $L = 256$ collapses to $R^2 \approx 0$ once $|M|$ is partialed out. Per-configuration evidence corroborates the $I_0$ Parseval identity but not the $|k|^{-2\eta}$ shape factor; the latter is established by ensemble $L$-scaling.
Chinese Translation
拓扑间隙 Δ——即临界点云相对于密度匹配的零假设模型(null)的过剩 H₁ 总持久度——满足标度关系 Δ ~ L^(d+η)。我们对此进行了解析推导:谱积分 I(α) = Σ_{k≠0} S_conn(k)|k|^α 在红外(IR)主导时满足标度关系 L^(2-α-η),从而给出 I(-2η) ~ L^(d+η)。分解式 I(-2η) = I₀ · I_shape 将体积贡献(I₀ ∝ N(1-m²) ~ L^d)与反常维度贡献(I_shape ~ L^η)分离;体积因子解释了 Δ 由磁化强度驱动的逐构型方差。我们证明该机制要求 d < 2 + η(红外主导),从而将其限制在物理系统的 d = 2 情形;在 d = 3 时谱积分由紫外(UV)主导,这解释了为何需要密度归一化。对 Potts q = 4 模型在 L = 32–256 范围内的 α 扫描发现 α_opt 位于 [-0.75, -0.5],与 -2η_Ising 一致而与 -2η_{q=4} = -1 不符;我们将其标注为初步结果,有待 L ≥ 1024 的验证。<m² · I(-2η)> 的超标度乘积由相关系数 r(m², I) ≈ -0.98 主导,这是通过共享的 I₀ 振幅实现的,因此我们将其作为协方差修正分析来报告。在将 Divol–Polonik 结果推广到非均匀泊松强度的启发式论证下,裸持续同调(PH)核是平坦的;有效核仅在临界点处获得对 k 的依赖性。Δ 与 I(-0.5) 之间的逐构型一致性主要来自磁化强度相关性:在 L = 256 时 R² = 0.91,但一旦对 |M| 进行偏相关控制后便坍缩至 R² ≈ 0。逐构型证据支持 I₀ 的 Parseval 恒等式,但不支持 |k|^(-2η) 形状因子;后者由系综 L-标度关系确立。
cs.LG / 2 / 2609.09162

Physics-informed neural networks by Gradient-Guided Gaussian Adaptive Sampling (3GAS-PINNs)

基于梯度引导高斯自适应采样的物理信息神经网络(3GAS-PINNs)
Wang, Yousen, Zhao, Wei
Abstract
Physics-informed neural networks (PINNs) provide a mesh-free framework for solving partial differential equations, yet their performance in nonlinear problems is often limited by slow convergence, gradient imbalance, and insufficient resolution to capture localized intermittent structures such as shock waves[1]. These issues arise primarily from the use of fixed weights of loss and uniform collocation point distributions, which cannot adapt to the evolving complexity of the solution field during training. To address these challenges, Gradient-Guided Gaussian Adaptive Sampling Physics-Informed Neural Networks (3GAS-PINNs) is proposed in this paper, which combines uniform probability distribution and Gaussian-smoothed probability distribution derived from the spatial gradients of solution, to maintain global constraint satisfaction as well as concentrating collocation points in regions of high gradient. Thus, intermittency structures like shock wave and solitons can be accurately captured. The method is evaluated on three benchmark nonlinear problems, including one-dimensional forced Burgers equation, Korteweg-de Vries (KdV) equation and nonlinear Schrodinger equation, all of which exhibit steep gradients or strong nonlinearity. In comparison with baseline PINNs, 3GAS-PINNs can effectively promote the physical consistency in intermittent regions. The accuracy of the numerical simulation can be improved by a factor of up to 14.
Chinese Translation
物理信息神经网络(PINNs)为求解偏微分方程提供了一种无网格框架,然而其在非线性问题中的性能往往受到收敛速度慢、梯度不平衡以及对捕捉激波等局部间歇性结构的分辨率不足的限制[1]。这些问题主要源于固定的损失权重和均匀配置点分布,无法在训练过程中适应解场不断演化的复杂程度。为应对这些挑战,本文提出了梯度引导高斯自适应采样物理信息神经网络(3GAS-PINNs),该方法将均匀概率分布与由解的空间梯度导出的高斯平滑概率分布相结合,在保持全局约束满足的同时,将配置点集中于高梯度区域,从而能够准确捕捉激波和孤子等间歇性结构。该方法在三个基准非线性问题上进行了评估,包括一维受迫Burgers方程、Korteweg-de Vries(KdV)方程和非线性薛定谔方程,这些问题均表现出陡峭梯度或强非线性。与基线PINNs相比,3GAS-PINNs能够有效提升间歇区域的物理一致性,数值模拟精度最高可提升14倍。
cs.LG / 3 / 2609.09163

World-Time Compute with Verified Code World Models

基于经验证代码世界模型的世界时间计算
Schwoebel, James, Semenec, Ingrida, Rousseva, Jenia, Ortiz, Marcos, Overbay, Collin, Klaus, Christopher, Edmond, Anderson, Bhatt, Manish, Thorstenson, Rome, Tsai, Jessica, Frasch, Martin G.
Abstract
LLMs generalize across a domain only after seeing many real, labeled examples, which most domains lack. We study a way to manufacture it cheaply. When a domain's dynamics can be written as code, one template instantiates into many world models: executable, verifiable programs over symbolic state, each an inexhaustible source of exactly-labeled trajectories. Fine-tuning an LLM on trajectories through many such worlds, which we call world-time compute, a training-time analogue of test-time compute, lifts generalization to held-out worlds it never trained on (synthesized world families). Gains are largest where capability is scarcest: +29 points at 0.5B; the largest model's lift is within noise, consistent with saturation. Labels can be trusted because the worlds are verified code: synthesized-then-checked dynamics are exact over 20-step rollouts and answer 10x out-of-distribution probes exactly (100%), whereas per-step LLM and MLP predictors compound error and collapse. Unlike domain randomization, each world is independently authored and verified; a corrupted-label control shows label exactness, not task variety, drives the gains. On real benchmarks (ARC-AGI grids, List Functions, CLRS) the same lever holds as per-world test-time training. On List Functions the harder cross-world form holds: one adapter trained on 128 disjoint worlds reaches 40% on held-out worlds versus 6% for a corrupted-label control (+34 points, CI [29, 39]). The gain is a saturating regularity, not a law: largest for few-step reasoning and small/weak models, fading for long chains, perception-induced tasks, and saturated tasks; cross-task transfer is weak without shared skill. Worlds are authored and served by OpenWorld, a zero-dependency framework (companion paper). Scope: symbolic state; pixel-native domains remain territory of learned models. All code, recipes, and this manuscript regenerate from one repository.
Chinese Translation
大语言模型(LLM)只有在见过大量真实的、带标注的样本后才能在某个领域内泛化,而大多数领域缺乏这样的数据。我们研究了一种低成本制造此类数据的方法。当一个领域的动态可以写成代码时,一个模板即可实例化为多个世界模型:这些是作用于符号状态的可执行、可验证的程序,每一个都是精确标注轨迹的取之不尽的来源。让LLM在众多此类世界的轨迹上进行微调——我们称之为世界时间计算(world-time compute),即测试时计算(test-time compute)在训练阶段的对应物——可将泛化能力提升至从未训练过的保留世界(合成的世界族)。在能力最稀缺之处收益最大:0.5B模型提升29分;最大模型的提升在噪声范围内,与饱和效应一致。标签之所以可信,是因为这些世界是经过验证的代码:先合成后检验的动态在20步推演中完全精确,并能对分布外探针给出精确回答(100%),而逐步预测的LLM和MLP预测器则会累积误差并崩溃。与域随机化不同,每个世界都是独立编写和验证的;一个损坏标签的对照组实验表明,驱动收益的是标签的精确性而非任务的多样性。在真实基准上(ARC-AGI网格、List Functions、CLRS),同一手段可以作为按世界的测试时训练(per-world test-time training)发挥作用。在List Functions上,更难的跨世界形式同样成立:在128个互不重叠的世界训练的一个适配器(adapter)在保留世界上达到40%,而损坏标签对照组仅为6%(提升34分,置信区间[29, 39])。这一收益是一种饱和的规律而非定律:在少步推理和小型/较弱模型上收益最大,在长推理链、感知诱发的任务以及已饱和的任务上逐渐减弱;在没有共享技能的情况下,跨任务迁移能力较弱。这些世界由OpenWorld(一个零依赖框架,见配套论文)编写和服务。适用范围:符号状态;像素原生领域仍属于学习型模型的领地。所有代码、配方及本手稿均可从一个代码仓库中复现生成。
cs.LG / 4 / 2609.09240

Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution

将后训练三值化扩展至Qwen3-8B:能力保持、复现、无损打包与打包执行
Malik, Anirudh, Mehra, M Sparsh, Devan, Poojith
Abstract
Ultra-low-bit language models promise reductions in storage and memory traffic, but a nominal "1.58-bit" label does not specify the deployed representation or its execution cost. We study a scale-up of an aggressive post-training conversion pipeline from Qwen3-4B to Qwen3-8B. The conversion uses KOTMS rotation, E2M-ATQ adaptive ternarisation, and GPTQ-style error compensation in a weight-only A16 configuration. We do not claim these algorithms as new. Our contribution is the end-to-end scale-up characterisation: an external reproduction gate, matched 4B/8B capability analysis, cross-corpus perplexity, effective-bit accounting, lossless lattice-aware packing, and direct packed execution. The 8B model reaches a three-corpus perplexity ratio of 1.361x, with WikiText-2, C4, and PTB ratios of 1.318x, 1.393x, and 1.371x. On eight zero-shot tasks at n = 500, mean accuracy is 64.6% versus 72.4% for FP16, corresponding to 78.5% chance-corrected retention and a 7.8-point absolute cost. The matched 4B run retains 69.6%, yielding an 8.9-point 8B advantage. The packed checkpoint is 8.24 GiB and preserves the recorded perplexity to measurement precision. Direct packed execution reaches 15.52 tokens/s in 7.35 GiB, while a preliminary packed GEMV remains slower than FP16 cuBLAS. The result is a validated scale-up baseline: model size improves robustness to aggressive post-training discretisation, actual serialisation is solved for the measured artefact, and direct execution is feasible, while broader seeds, calibration distributions, and kernel optimisation remain open.
Chinese Translation
超低比特语言模型有望减少存储和内存流量,但名义上的"1.58比特"标签并未指明实际部署的表示形式及其执行成本。我们研究了一个激进的后训练转换流水线从Qwen3-4B到Qwen3-8B的规模化扩展。该转换在仅权重A16配置下使用KOTMS旋转、E2M-ATQ自适应三值化以及GPTQ风格的误差补偿。我们并不声称这些算法是新的。我们的贡献是端到端的规模化特征评估:外部复现门槛、匹配的4B/8B能力分析、跨语料库困惑度、有效比特核算、无损的格感知打包以及直接的打包执行。8B模型在三个语料库上的困惑度比值达到1.361倍,其中WikiText-2、C4和PTB的比值分别为1.318倍、1.393倍和1.371倍。在n = 500的八项零样本任务上,平均准确率为64.6%,而FP16为72.4%,对应78.5%的机会校正保持率和7.8个百分点的绝对损失。匹配的4B运行保持了69.6%,即8B模型具有8.9个百分点的优势。打包后的检查点大小为8.24 GiB,并将记录的困惑度保持在测量精度范围内。直接打包执行在7.35 GiB内存下达到15.52 tokens/s,而初步的打包GEMV仍慢于FP16 cuBLAS。该结果是一个经过验证的规模化基线:模型规模的提升增强了对激进后训练离散化的鲁棒性,针对所测工件的实际序列化问题已得到解决,直接执行是可行的,而更多种子、校准数据分布和内核优化仍是待解决的问题。
cs.LG / 5 / 2609.09241

Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts

动态稀疏混合专家模型的分布一致性推断
Kim, Dohyeon, Soro, Bedionita, Hwang, Sung Ju
Abstract
Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large foundation models. However, most MoE models use a fixed top-$k$ expert selection policy, assigning the same expert budget to every token even when fewer experts may be sufficient. Inference-time dynamic top-$k$ routing can reduce computation without retraining, but existing methods often overlook the distributional shift caused by deviating from the training-time routing configuration. We show that reducing the number of activated experts consistently increases the RMS scale and variance of SMoE outputs, inducing a representation mismatch that contributes to downstream performance degradation in addition to the loss of expert capacity. To address this correctable component, we propose Layer-wise Distribution Alignment (LDA), a lightweight inference-time correction that uses layer-wise calibration statistics to align reduced-routing representations with the default configuration. Across multiple SMoE LLMs, benchmarks, and routing strategies, LDA recovers much of the performance lost induced by the distributional shift under reduced routing while preserving sparse-inference efficiency with negligible overhead.
Chinese Translation
混合专家架构已成为大型基础模型在保持高效推理的同时扩展模型容量的强大范式。然而,大多数MoE模型采用固定的top-$k$专家选择策略,即使较少的专家可能已经足够,也为每个token分配相同的专家预算。推理阶段的动态top-$k$路由可以在无需重新训练的情况下减少计算量,但现有方法往往忽略了偏离训练时路由配置所引起的分布偏移。我们表明,减少激活专家的数量会持续增加SMoE输出的RMS尺度和方差,从而导致表示失配,这一因素除了专家容量损失之外,还加剧了下游性能的下降。为解决这一可纠正的成分,我们提出层级分布对齐方法,这是一种轻量级的推理时校正方法,利用逐层校准统计量将减少路由下的表示与默认配置的表示对齐。在多个SMoE大语言模型、基准测试和路由策略上的实验表明,LDA在减少路由的情况下恢复了由分布偏移引起的性能损失的大部分,同时以可忽略的开销保持了稀疏推理的效率。
cs.LG / 6 / 2609.09254

DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable Connectivity

DiffLUT-Net:具有可学习连接性的FPGA查找表网络的可微分训练
Ye, Jiaqi, Gong, Xinrui, Wang, Jingcun, Kondrateva, Olga, Li, Bing, Zhang, Grace Li
Abstract
Field-programmable gate arrays (FPGAs) enable efficient neural-network inference, but most deployment flows either accelerate multiply-accumulate operations or convert pretrained quantized models into lookup tables (LUTs). We present DiffLUT-Net, an FPGA-native network connected by six-input LUTs that are trained from scratch. We jointly learn the 64 truth-table entries of a LUT and the source to each of its six input ports using a differentiable LUT function relaxation and hardware source selection. After training, the truth tables and connections are discretized, unused logic can be pruned, and the network is exported directly as synthesizable Verilog. Across five benchmarks, DiffLUT-Net achieves favorable accuracy-resource trade-offs. These results demonstrate the effectiveness of jointly learning LUT functions and sparse connectivity for compact FPGA-native inference. The code is available at https://github.com/TUDa-HWAI/DiffLUT-Network.
Chinese Translation
现场可编程门阵列(FPGA)能够实现高效的神经网络推理,但现有的部署流程要么加速乘加运算,要么将预训练的量化模型转换为查找表(LUT)。我们提出了DiffLUT-Net,这是一种由六输入查找表构成、可从零开始训练的FPGA原生网络。我们利用可微分的查找表函数松弛和硬件源选择,联合学习每个查找表的64个真值表项及其六个输入端口的信号源。训练完成后,真值表和连接被离散化,未使用的逻辑可以被剪枝,网络可直接导出为可综合的Verilog代码。在五个基准测试中,DiffLUT-Net实现了良好的精度-资源权衡。这些结果证明了联合学习查找表函数和稀疏连接性对于紧凑的FPGA原生推理的有效性。代码可在 https://github.com/TUDa-HWAI/DiffLUT-Network 获取。
cs.LG / 7 / 2609.09257

Accountable and uncertainty-aware evaluation of sensor-based AI under distribution shift: devices, subjects, and nearly three years underground

分布偏移下基于传感器的AI的可问责与不确定性感知评估:设备、受试者与近三年的地下实验
Platte, Benny, Thomanek, Rico, Roschke, Christian, Ritter, Marc
Abstract
Sensor-based AI systems are rarely operated under the conditions under which they were trained: devices, personnel and recording epochs change, and each change degrades performance in ways a random train-test split cannot reveal. We propose a staged, accountable evaluation protocol that treats the evaluation of a deployed model as a measurement with declared reference levels and a quantified uncertainty. Four cumulative generalisation stages hold out devices, subjects and time. Each stage is judged on quantiles of repeated trainings against chance references with the correct class count, an out-of-present-scope rate exposes silent misdirection towards classes that are no longer present in deployment relative to training, and an explicit decision rule ties roll-out decisions not to means but to 5% quantiles. We demonstrate the protocol on infrastructure-free geomagnetic localisation with smartphone-based recurrent classifiers in two real underground mines, including a replication of the scheme's training stages at the second site. Unchanged models are re-evaluated on data recorded 34 months after the training campaign, on a device generation unknown at training time and with a held-out surveyor. The 5% quantile of their present-conditioned precision there is 0.39 over 299 repeated trainings, 16.5 times the chance level; across the composition of the 42 reachable location classes the figure varies by +/-0.08, several times the spread between repeated runs. Repeated trainings of a single configuration show why means mislead: a bimodal configuration passes a mean-based test decisively while its 5% quantile lies more than an order of magnitude below chance.
Chinese Translation
基于传感器的AI系统很少在与训练条件相同的条件下运行:设备、人员和记录时段会发生变化,而每一次变化都会导致性能下降,这种退化是随机训练-测试划分所无法揭示的。我们提出了一种分阶段的、可问责的评估协议,将部署模型的评估视为一种带有声明参考水平和量化不确定性的测量。四个累积泛化阶段分别留出设备、受试者和时间。每个阶段基于重复训练的分位数进行评判,并与具有正确类别数的随机机会参考水平进行比较;同时,一个'超出当前范围率'(out-of-present-scope rate)指标揭示模型在部署时相对于训练阶段对已不存在的类别的隐蔽性错误指向;此外,一个明确的决策规则将部署决策与5%分位数(而非均值)挂钩。我们在两个真实的地下矿井中,基于智能手机的循环神经网络分类器展示了该协议在无基础设施的地磁定位上的应用,包括在第二个地点对该方案训练阶段的复现。我们在训练活动34个月后记录的数据上、在训练时未知的设备型号上以及由留出的测量员数据上,对未更改的模型进行重新评估。在299次重复训练中,模型在当前条件约束下的精确率的5%分位数为0.39,是随机机会水平的16.5倍;在42个可到达的位置类别的不同组成下,该数值变化范围为±0.08,是重复训练之间波动的数倍。单一配置的重复训练说明了为何均值会误导:一个双峰分布的配置在基于均值的测试中轻松通过,而其5%分位数却比随机机会水平低一个数量级以上。
cs.LG / 8 / 2609.09299

Literati: Towards Anytime Optimal Shape Generalized Trees via AO*

Literati:基于AO*的任意时间最优形状广义树学习
Upadhya, Nakul, Cohen, Eldan
Abstract
Decision trees are prized for their interpretability and strong performance on tabular data, but popular greedy top-down induction algorithms can yield suboptimal and unnecessarily complex structures. Optimal decision tree methods address this through global optimization, yet remain restricted to axis-aligned threshold splits, which limit the expressivity of each node and often force deep, complex trees to capture non-linear feature effects. Shape Generalized Trees (SGTs) generalize threshold splits to learnable univariate shape functions, improving expressivity and enabling more compact trees. However, existing SGT induction algorithms are greedy and offer no optimality guarantees. In this work, we introduce Literati, the first algorithm for optimal SGT induction. We propose a novel AND/OR graph formulation of the problem that jointly optimizes tree structure and shape function complexity. To solve this AND/OR graph, we develop an AO*-based algorithm with two enhancements that improve anytime performance while preserving optimality: a secondary heuristic for OR-node selection and a round-robin policy for AND-node exploration. Across 24 real-world datasets, Literati achieves higher training and test accuracy than state-of-the-art tree approaches.
Chinese Translation
决策树因其可解释性以及在表格数据上的强大性能而备受推崇,但流行的贪心自顶向下归纳算法可能产生次优且过于复杂的树结构。最优决策树方法通过全局优化解决这一问题,但仍局限于轴对齐的阈值分裂,这限制了每个节点的表达能力,往往迫使树变得又深又复杂才能捕捉非线性的特征效应。形状广义树(Shape Generalized Trees, SGTs)将阈值分裂推广为可学习的单变量形状函数,提升了表达能力,并使树结构更加紧凑。然而,现有的SGT归纳算法都是贪心算法,无法提供最优性保证。在本工作中,我们提出了Literati,这是首个用于最优SGT归纳的算法。我们提出了一种新颖的与/或图(AND/OR graph)问题表述,可联合优化树结构和形状函数复杂度。为求解该与/或图,我们开发了一种基于AO*的算法,并引入了两项改进,在保持最优性的同时提升了任意时间(anytime)性能:一种用于OR节点选择的辅助启发式方法,以及一种用于AND节点探索的轮转策略。在24个真实世界数据集上,Literati取得了比最先进树方法更高的训练和测试准确率。
cs.LG / 9 / 2609.09367

Explaining f-Divergence-Based Regularization via Local Curvature and Sharpness-Aware Minimization

基于局部曲率与锐度感知最小化解释f-散度正则化
Jamoussi, Nour, Kountouris, Marios
Abstract
Divergence-based regularization and Sharpness-Aware Minimization (SAM) are two prominent approaches for improving generalization in deep learning, both motivated by robustness to perturbations. However, their relationship has remained largely unexplored. Building on classical second-order expansions of $f$-divergences, we show that the two methods are locally consistent under parameter-space perturbations: both induce curvature-sensitive penalties, with divergence regularization yielding a Fisher-weighted quadratic form and SAM penalizing sharpness through the dominant Hessian eigenvalue. For negative log-likelihood objectives with exponential-family output distributions, this correspondence becomes especially transparent, since the Fisher and Gauss-Newton matrices coincide. We further show that the same local geometric perspective extends to input-space perturbations, where divergence-based regularization is defined through transformations of the input. In this setting, the regularizer induces a pullback quadratic form on the input space, providing a more general perturbation framework than standard SAM while preserving the same local sensitivity interpretation. To validate the analysis empirically, we use the asymmetric $\alpha$-skew Jensen-Shannon divergence (JSD) family as a controlled testbed. Its local curvature coefficient scales as $\alpha(1-\alpha)$ and is maximized at the symmetric point $\alpha=\tfrac12$, which recovers the standard JSD. Loss-landscape visualizations in the input-perturbation regime show that stronger induced curvature penalization is associated with flatter local minima. Experiments on four benchmark datasets further demonstrate that both accuracy and negative log-likelihood are consistently best near this regime of maximal curvature penalization.
Chinese Translation
基于散度的正则化与锐度感知最小化(SAM)是提升深度学习泛化能力的两种重要方法,二者均源于对扰动鲁棒性的考虑。然而,它们之间的关系在很大程度上尚未被探索。基于$f$-散度的经典二阶展开,我们证明这两种方法在参数空间扰动下是局部一致的:二者都会诱导对曲率敏感的惩罚项,其中散度正则化产生一个Fisher加权的二次型,而SAM通过Hessian矩阵的主特征值惩罚锐度。对于具有指数族输出分布的负对数似然目标函数,这种对应关系尤为清晰,因为Fisher矩阵与Gauss-Newton矩阵相重合。我们进一步表明,同样的局部几何视角可扩展至输入空间扰动,此时基于散度的正则化通过对输入的变换来定义。在该情形下,正则化项在输入空间上诱导出一个拉回(pullback)二次型,提供了一个比标准SAM更一般的扰动框架,同时保持了相同的局部敏感性解释。为了实证验证这一分析,我们采用非对称的$\alpha$-偏斜Jensen-Shannon散度(JSD)族作为受控测试平台。其局部曲率系数以$\alpha(1-\alpha)$的形式缩放,并在对称点$\alpha=\tfrac12$处取得最大值,该点对应于标准JSD。在输入扰动情形下的损失景观可视化表明,更强的诱导曲率惩罚与更平坦的局部极小值相关。在四个基准数据集上的实验进一步表明,准确率和负对数似然均在该最大曲率惩罚区域附近一致地表现最佳。
cs.LG / 10 / 2609.09370

Constraint-Aware Discrete Black-Box Optimization Using Tensor Decomposition

基于张量分解的约束感知离散黑盒优化
Onoue, Keisuke, Kojima, Ryosuke
Abstract
Discrete black-box optimization is often addressed using approaches such as Sequential Model-Based Optimization (SMBO), which aims to improve sample efficiency by fitting surrogate models that approximate a costly objective function over a discrete search space. In many real-world problems, the set of feasible inputs is often given by logical constraints known in advance. However, existing surrogate modeling techniques generally fail to capture the symbolic rules governing feasibility in discrete input spaces. In this paper, we propose a surrogate modeling approach based on tensor decomposition that captures the structure of discrete search spaces while directly integrating feasibility information. To implement this approach, we formulate surrogate model training as a constrained polynomial optimization problem and solve a relaxed formulation using a differentiable penalty term derived from T-norms. Our experiments on both synthetic and real-world benchmarks, including a pressure vessel design task, demonstrate that the proposed method improves sample efficiency by effectively guiding the search away from infeasible regions.
Chinese Translation
离散黑盒优化通常采用顺序模型优化(Sequential Model-Based Optimization, SMBO)等方法,其目标是通过拟合代理模型来逼近离散搜索空间上代价高昂的目标函数,从而提高样本效率。在许多现实问题中,可行输入集合往往由预先已知的逻辑约束给出。然而,现有的代理模型建模技术通常难以捕捉离散输入空间中决定可行性的符号规则。本文提出了一种基于张量分解的代理模型建模方法,该方法在捕捉离散搜索空间结构的同时,直接整合可行性信息。为实现这一方法,我们将代理模型训练形式化为一个约束多项式优化问题,并利用由T-范数(T-norm)导出的可微惩罚项求解其松弛形式。在合成数据集和真实基准任务(包括压力容器设计任务)上的实验表明,所提方法能够有效引导搜索远离不可行区域,从而提高样本效率。
cs.LG / 11 / 2609.09388

XAI-Refine: An Automated Explanation-Knowledge Loop for Brain-Age Prediction

XAI-Refine:一种用于脑龄预测的自动化解释-知识循环
Qiao, Yang, Wu, Junjie, Qiu, Deqiang, Lah, James J., Zhao, Liang
Abstract
Brain-age prediction models are commonly evaluated by predictive accuracy, yet accurate predictions alone do not establish that a model relies on reproducible or neurobiologically supported mechanisms. Post-hoc explanation methods can expose these mechanisms, but existing workflows typically stop at diagnosis or require correction targets to be specified before model analysis. We propose XAI-Refine, an automated explanation-knowledge loop for brain-age prediction from resting-state functional connectivity. At each iteration, XAI-Refine consolidates complementary post-hoc analyses across repeated training runs into reliable, structured model explanations. It converts each reliable explanation into a neutral neurobiological question, retrieves and verifies relevant literature, and compiles the verified evidence into an admissible set in the same typed explanation space. The target for refinement is defined as the minimal projection of the current model explanation onto the admissible set induced by applicable verified knowledge. This revised explanation is then translated into a differentiable constraint while preserving the originating model variable, measurement operator, and applicable scope. Candidate updates are promoted only when multi-seed validation confirms target-directed explanatory movement, predictive performance remains within a prespecified guardrail, and non-target explanatory drift remains bounded. Experiments on functional-connectivity-based brain-age prediction evaluate predictive performance, explanation reliability, literature alignment, and target-specific model revision, illustrating a structured route from post-hoc analysis to evidence-guided model refinement.
Chinese Translation
脑龄预测模型通常以预测精度作为评估标准,然而仅有准确的预测并不能证明模型依赖于可复现的或具有神经生物学依据的机制。事后解释方法能够揭示这些机制,但现有的工作流程通常止步于诊断阶段,或要求在模型分析之前预先指定修正目标。我们提出了 XAI-Refine,一种面向基于静息态功能连接的脑龄预测的自动化解释-知识循环。在每次迭代中,XAI-Refine 将多次重复训练运行中的互补性事后分析整合为可靠、结构化的模型解释。它将每个可靠解释转化为一个中性的神经生物学问题,检索并验证相关文献,并将经过验证的证据汇编为处于同一类型化解释空间中的可采纳集合。修正目标被定义为当前模型解释在由适用且经过验证的知识所诱导的可采纳集合上的最小投影。随后,该修正后的解释被转化为可微分约束,同时保留原始的模型变量、测量算子及适用范围。只有当多种子验证确认解释朝目标方向移动、预测性能保持在预设的护栏范围内、且非目标方向的解释漂移有界时,候选更新才会被采纳。基于功能连接的脑龄预测实验从预测性能、解释可靠性、文献一致性以及针对性模型修正等方面进行了评估,展示了从事后分析到证据引导的模型修正的结构化路径。
cs.LG / 12 / 2609.09429

Applying foundation model embeddings towards urban livability evaluation

应用基础模型嵌入进行城市宜居性评估
Khot, Ayush, Zhou, Wen, Wang, Shaowen
Abstract
While accurate measurement of socioeconomic indicators remains challenging in data-scarce regions, which limits policy interventions and resource allocation, high-resolution geospatial data is widely available and can contain information on various livability statistics. We investigate which physical features are encoded within foundation model embeddings, such as AlphaEarth, AnySat, and TerraMind, and provide a systematic framework for identifying the most predictive geospatial indicators. By analyzing how different types of geospatial data influence urban livability predictions, our approach enables researchers to prioritize the most informative features for their specific applications. Additionally, we demonstrate how to leverage foundation model embeddings to enhance prediction performance for these outcomes. This work contributes a principled methodology for extracting actionable information from satellite imagery while accounting for complex spatial dependencies, with applications in predicting urban livability in regions with limited observation data.
Chinese Translation
尽管在数据稀缺地区对社会经济指标进行准确测量仍然具有挑战性,这限制了政策干预和资源配置,但高分辨率地理空间数据是广泛可得的,并且可以包含各种宜居性统计信息。我们研究了哪些物理特征被编码在基础模型嵌入(如 AlphaEarth、AnySat 和 TerraMind)中,并提供了一个系统性框架来识别最具预测能力的地理空间指标。通过分析不同类型的地理空间数据如何影响城市宜居性预测,我们的方法使研究人员能够为其特定应用优先选择信息量最大的特征。此外,我们展示了如何利用基础模型嵌入来提升对这些结果的预测性能。这项工作提供了一种从卫星影像中提取可操作信息的原则性方法,同时考虑了复杂的空间依赖关系,可应用于在观测数据有限的地区预测城市宜居性。
cs.LG / 13 / 2609.09432

SCCM : Stream Cruise Control Method for Automated Drift Detection and Adaptation

SCCM:用于自动化概念漂移检测与适应的流式巡航控制方法
Abu-Shaira, Mohammad, Shi, Weishi
Abstract
Real-world datasets often exhibit evolving distributions, known as concept drift. Ignoring drift degrades predictive performance, while reliance on fixed hyperparameters further limits model adaptability under changing conditions. Adaptive learning addresses this challenge by continuously updating models online, allowing them to incrementally adjust and remain effective as data distributions evolve. This paper presents the Stream Cruise Control Method (SCCM), a comprehensive framework for drift detection and adaptation in online regression. SCCM enables automated adaptation through early-response, pre-update drift detection, drift magnitude quantification, KPI-window-based thresholding for local false-alarm mitigation, dynamic hyperparameter tuning, and model recalibration. SCCM also adopts an in-memory design for real-time adaptability, unlike purely reactive methods that typically activate adaptation only after performance degradation is observed. By using dynamic thresholding and remaining agnostic to data distributions, SCCM supports KPI-based monitoring across varying data streams, including high-dimensional and large-scale settings. SCCM is integrated with four online regression models and evaluated on 18 synthetic datasets covering abrupt, incremental, and alternating gradual drift, together with eight real-world datasets. The evaluation uses both R2 and MSE and compares against eight detector--adaptation baselines. Results show improved predictive performance and effective drift handling across the evaluated online regression settings.
Chinese Translation
现实世界的数据集通常表现出不断演变的分布,即概念漂移(concept drift)。忽视漂移会降低预测性能,而依赖固定的超参数进一步限制了模型在变化条件下的适应性。自适应学习通过在线持续更新模型来应对这一挑战,使模型能够随着数据分布的演变进行增量调整并保持有效。本文提出了流式巡航控制方法(Stream Cruise Control Method, SCCM),这是一个用于在线回归中漂移检测与适应的综合框架。SCCM 通过以下机制实现自动化适应:早期响应的更新前漂移检测、漂移幅度量化、基于 KPI 窗口的阈值设定以减少局部误报、动态超参数调优以及模型重新校准。此外,与通常仅在观察到性能下降后才启动适应的纯响应式方法不同,SCCM 采用内存内(in-memory)设计以实现实时适应性。通过使用动态阈值并保持对数据分布的不可知性(distribution-agnostic),SCCM 支持在不同数据流(包括高维和大规模场景)上进行基于 KPI 的监控。SCCM 与四种在线回归模型集成,并在 18 个涵盖突变型(abrupt)、渐增型(incremental)和交替渐变型(alternating gradual)漂移的合成数据集以及 8 个真实世界数据集上进行了评估。评估同时使用 R² 和 MSE 指标,并与八种检测器—适应基线方法进行比较。结果表明,在所评估的在线回归设置中,SCCM 提升了预测性能并有效处理了概念漂移。
cs.LG / 14 / 2609.09433

Efficient Leakage-Free Neural Architecture Search under Leave-One-Subject-Out Evaluation

留一受试者评估下高效且无信息泄漏的神经架构搜索
Hihn, Heinke
Abstract
Leave-One-Subject-Out (LOSO) evaluation estimates generalisation performance for subject-based classification but makes Neural Architecture Search (NAS) computationally expensive because a fully nested implementation requires N independent architecture searches and, assuming approximately linear training cost, scales as O(N^2). We propose a leakage-free, block-based approach that shares NAS runs across subjects. On the BioVid Heat Pain dataset, our approach increased the mean accuracy from 82.79% to 83.39% while reducing the number of parameters by up to 99.2%.
Chinese Translation
留一受试者(Leave-One-Subject-Out, LOSO)评估用于估计基于受试者的分类任务的泛化性能,但它使神经架构搜索(Neural Architecture Search, NAS)的计算代价十分高昂,因为完全嵌套的实现需要N次独立的架构搜索,并且假设训练成本近似线性时,其复杂度按O(N^2)扩展。我们提出了一种无信息泄漏的、基于分块的(block-based)方法,可在不同受试者之间共享NAS运行结果。在BioVid热痛(Heat Pain)数据集上,我们的方法将平均准确率从82.79%提升至83.39%,同时将参数数量最多减少了99.2%。
cs.LG / 15 / 2609.09434

Tensor-Train Weak SINDy: Identifying High-Dimensional Nonlinear Dynamics

张量链弱形式SINDy:识别高维非线性动力学
Houser, Will, Dukic, Vanja, Bortz, David M.
Abstract
In recent years, weak-form methods have made significant advances in data-driven discovery of dynamical systems. However, in high-dimensional settings, current techniques can prove expensive in both computation and memory. In this work, we introduce TT-WSINDy, which combines techniques of the Multidimensional Approximation of Nonlinear Dynamics (MANDy) and Weak Sparse Identification of Nonlinear Dynamics (WSINDy) methods, implementing requisite computations in the tensor-train (TT) format. We demonstrate that this method is able to search an exponentially-growing space of candidate functions -- performing weak-form transformation, regression, and sparsification -- without suffering from the curse of dimensionality.
Chinese Translation
近年来,弱形式方法在数据驱动的动力学系统发现领域取得了显著进展。然而,在高维场景下,现有技术在计算和内存方面开销巨大。本文提出TT-WSINDy方法,该方法结合了非线性动力学的多维逼近(MANDy)与弱形式稀疏辨识非线性动力学(WSINDy)两种技术,并以张量链(tensor-train, TT)格式实现所需的计算。我们证明了该方法能够在呈指数级增长的候选函数空间中进行搜索——完成弱形式变换、回归和稀疏化——而不会遭受维数灾难的困扰。
cs.LG / 16 / 2609.09451

Uncertainty-Aware Sea-Ice Type Mapping with Multiple Ice Charts

基于多海冰图的不确定性感知海冰类型制图
Taleghan, Samira Alkaee, Koo, Younghyun, Barrett, Andrew P., Banaei-Kashani, Farnoush
Abstract
Sea-ice stage of development (SoD) describes the age and associated thickness of sea ice and provides important information for navigation, and operational ice monitoring. SoD labels are obtained from operational ice charts, where trained analysts interpret satellite observations and assign standardized stage codes to regions with similar ice conditions. These codes often represent ranges of compatible ice thicknesses rather than exact physical values. Deep-learning methods can automate SoD mapping and commonly adopt operational ice charts as reference labels for training. These annotations are not exact, however; this is because chart interpretation relies on analyst judgement and on the observations available at the time, so different ice services may assign different SoD labels to the same conditions. We term this variation across independently produced expert annotations multi-annotator label uncertainty; collapsing the annotations into a single deterministic target discards this variation. A second source of uncertainty originates in the learned model itself. In this paper, we quantify both sources: annotation uncertainty from disagreement among independent ice-service charts and model uncertainty from the learned predictive models. We then evaluate their relationship by testing whether model uncertainty is higher where ice services disagree. We observe that supervision incorporating information from multiple annotators can improve this correspondence, with soft supervision achieving the highest overall correlation of 0.256. The relationship becomes substantially stronger near the ice edge, where model predictive uncertainty closely tracks multi-annotator disagreement, reaching a correlation of 0.704 within 0--10 km. Among the uncertainty-estimation approaches, Monte Carlo dropout provides the best-calibrated confidence estimates, with an expected calibration error of 0.050.
Chinese Translation
海冰发育阶段描述海冰的年龄及其相应厚度,为航行和业务化海冰监测提供重要信息。发育阶段标签来自业务化海冰图,由经过训练的分析员解译卫星观测数据,并为冰况相似的区域赋予标准化的阶段代码。这些代码通常表示相容冰厚度的范围,而非精确的物理值。深度学习方法可以实现发育阶段制图的自动化,并通常采用业务化海冰图作为训练的参考标签。然而,这些标注并不精确,因为海冰图解译依赖于分析员的判断以及当时可用的观测数据,因此不同的海冰服务机构可能对相同的冰况赋予不同的发育阶段标签。我们将这种来自独立专家标注之间的差异称为多标注者标签不确定性;将标注压缩为单一的确定性目标会丢弃这种差异信息。第二类不确定性来源是所学习的模型本身。本文对这两类不确定性进行了量化:一是来自独立海冰服务机构海冰图之间分歧的标注不确定性,二是来自学习到的预测模型的模型不确定性。随后,我们通过检验在海冰服务机构存在分歧的区域模型不确定性是否更高,来评估二者之间的关系。我们观察到,融合多标注者信息的监督方式能够改善这种对应关系,其中软监督取得了最高的总体相关性,为0.256。在冰缘附近,这一关系显著增强,模型预测不确定性紧密跟踪多标注者分歧,在0–10公里范围内相关性达到0.704。在各种不确定性估计方法中,蒙特卡洛 dropout 提供了校准效果最好的置信度估计,其期望校准误差为0.050。
cs.LG / 17 / 2609.09466

Exact-Form Regret for Gradient Descent, Mirror Descent and Follow-the-Regularized-Leader

梯度下降、镜像下降与正则化Leader跟随算法的精确形式遗憾
Soleymani, Ashkan, Farina, Gabriele, Jaillet, Patrick
Abstract
Online gradient descent is usually studied through external regret, where the learner competes with fixed alternatives. Recent work shows that first-order methods control richer action-dependent deviations. We ask for a geometric characterization of the deviations with respect to which online gradient descent, mirror descent, and follow-the-regularized-leader (FTRL) achieve no regret. We identify exactness as the common principle. Exactness means that the relevant displacement field is generated by a scalar potential, or equivalently that the associated one-form is exact in the geometry used by the algorithm. This geometry depends on the algorithm. For gradient descent it is Euclidean geometry, for mirror descent it is the geometry induced by the regularizer, and for FTRL it is the cumulative dual state. Under mild regularity conditions, exactness yields sublinear regret, while nonzero circulation provides the complementary obstruction and leads to linear regret. This gives a unified geometric framework for understanding the deviation classes controlled by these algorithms and reveals that different first-order methods can control genuinely different classes of deviations. These deviation classes have direct consequences for learning, particularly in games. We study the equilibrium notions induced by exact-form deviations and introduce conservative correlated equilibrium, reflecting both the conservative geometry of the underlying displacement fields and the restricted family of deviations available to the players. We characterize its relation to correlated equilibrium, determine when the resulting equilibrium notions coincide and when they separate, and show how these relationships depend on the geometry and the learning algorithm. Overall, this work gives a unified geometric account of what first-order online learning algorithms are no-regret with respect to, beyond fixed comparators.
Chinese Translation
在线梯度下降通常通过外部遗憾进行研究,即学习器与固定的备选方案进行竞争。最近的研究表明,一阶方法能够控制更丰富的依赖于动作的偏差。我们寻求一个几何刻画,用以描述在线梯度下降、镜像下降和正则化Leader跟随算法(FTRL)相对于哪些偏差是无遗憾的。我们将“精确性”识别为共同原理。精确性意味着相关的位移场由一个标量势生成,等价地说,即在算法所使用的几何中,相应的1-形式是恰当的。这种几何依赖于具体算法:对于梯度下降是欧氏几何,对于镜像下降是由正则项诱导的几何,而对于FTRL则是累积对偶状态。在温和的正则性条件下,精确性带来次线性遗憾,而非零环量则构成互补的障碍并导致线性遗憾。这为理解这些算法所控制的偏差类别提供了一个统一的几何框架,并揭示出不同的一阶方法可以控制本质上不同的偏差类别。这些偏差类别对学习具有直接意义,尤其是在博弈中。我们研究了由精确形式偏差诱导的均衡概念,并引入了保守相关均衡(conservative correlated equilibrium),它既反映了底层位移场的保守几何性质,也反映了参与者可用的受限偏差族。我们刻画了它与相关均衡的关系,确定了所得均衡概念何时重合、何时分离,并展示了这些关系如何依赖于几何结构和学习算法。总体而言,这项工作为“一阶在线学习算法相对于什么是无遗憾的”提供了超越固定比较对象的统一几何解释。
cs.LG / 18 / 2609.09468

Building the Harness Automatically: Self-Play in Code Distills a Text Harness for Black-Box Optimization

自动构建测试框架:代码自我博弈蒸馏出面向黑盒优化的文本测试框架
Wu, Yi, Ren, Zheng, Hu, Zhiyu, Wang, Haochen, Chang, Daryl, Wei, Li, Wang, Ting, Li, Zhen, Gupta, Pooja, Jindal, Nitin, Heldt, Lukasz
Abstract
Can an agent learn a numerical search strategy through executable practice and then transfer that strategy as text? We study low-budget black-box optimization, where unaided language models remain well below strong classical optimizers. During development, an agent repeatedly writes and evaluates optimizer programs. It then distills the resulting program and practice record once into a 197-word primary Harness A, which is frozen before evaluation. Harness A reduces Gemini Flash regret by 48\% in an independent $N=30$ study ($p<.001$), enters the GP-BO performance range on the practice family, and lowers mean regret on all three held-out BBOB landscapes. The same text improves every tested Gemini executor and transfers to Claude Sonnet, reducing regret by 43\% and 49\% ($p\leq.005$). An independent end-to-end replication produces Harness B, a different program and text at the same performance tier. The same framework also attains the lowest regret on a sealed YouTube reward-tuning production benchmark. Executable practice is thus a viable way to discover a search policy, and language a portable medium for deploying it.
Chinese Translation
智能体能否通过可执行的练习学习数值搜索策略,并将其以文本形式迁移?我们研究低预算黑盒优化问题,在该问题中,无辅助的语言模型的表现仍远低于强大的经典优化器。在开发阶段,智能体反复编写并评估优化器程序,随后将所得程序和练习记录一次性蒸馏为一个197词的主测试框架 Harness A,并在评估前将其冻结。在一项独立的 $N=30$ 研究中,Harness A 将 Gemini Flash 的遗憾值(regret)降低了48%($p<.001$),在练习问题族上进入 GP-BO 的性能区间,并在全部三个留出的 BBOB 测试景观上降低了平均遗憾值。同一文本提升了所有受测 Gemini 执行器的表现,并可迁移至 Claude Sonnet,分别将遗憾值降低43%和49%($p\leq.005$)。一项独立的端到端复现产生了 Harness B,即处于相同性能层级的另一个程序与文本。同一框架还在一个密封的 YouTube 奖励调优生产基准上取得了最低的遗憾值。因此,可执行练习是发现搜索策略的可行途径,而语言则是部署该策略的可移植媒介。
cs.LG / 19 / 2609.09476

From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function Calls

从固定键到可读模式:用于车辆智能体函数调用的小型语言模型
Asl, Hamed Jafarzadeh, Yu, Yuanhao, Nia, Vahid Partovi
Abstract
In-vehicle assistants must translate natural-language requests into accurate vehicle function calls under strict memory and latency constraints, making small language models (SLMs) attractive for on-device deployment. For such models, a key design choice is how the available function surface is presented. Two approaches are to represent each function with a dedicated Functional Token (FT) or provide function schemas directly in the prompt. FTs enable compact inference but are restricted to functions learned during training, whereas Schema-in-Prompt (SIP) can generalize to unseen functions at the cost of longer prompts and higher inference overhead. We introduce a benchmark of 9,822 single-turn examples spanning 79 vehicle functions derived from Android Automotive, including held-out functions and requests requiring refusal. We compare both approaches under matched fine-tuning across four SLMs from 270M to 1.7B parameters. On functions seen during training, scaling provides limited benefit: the 270M model can match the 1.7B model, while the strongest overall performance occurs at 0.6B. On held-out functions, FT achieves zero accuracy by construction, whereas SIP generalizes and improves substantially with scale. On out-of-scope requests, FT can invoke an unavailable function it was trained to emit, while SIP more reliably refuses based on the functions offered. This flexibility comes with higher memory use and latency. Our theoretical analysis explains how SIP enables generalization and why longer schema contexts increase inference cost. Overall, function-surface representation, rather than model scale alone, determines the capabilities and failure modes of SLM-based vehicle function calling.
Chinese Translation
车载助手必须在严格的内存和延迟约束下将自然语言请求转换为准确的车辆功能调用,这使得小型语言模型(SLM)非常适合设备端部署。对于此类模型,一个关键的设计选择是如何呈现可用的功能表面。两种方法分别是:使用专用的功能标记(Functional Token, FT)表示每个功能,或在提示词中直接提供功能模式(schema)。FT 能够实现紧凑的推理,但仅限于训练期间学习过的功能;而提示词内模式方法(Schema-in-Prompt, SIP)虽然可以泛化到未见过的功能,但代价是提示词更长、推理开销更高。我们构建了一个包含 9,822 个单轮示例的基准数据集,涵盖源自 Android Automotive 的 79 个车辆功能,其中包括留出功能(held-out functions)和需要拒绝的请求。我们在相同的微调条件下,比较了两种方法在四个参数量从 270M 到 1.7B 的 SLM 上的表现。在训练期间见过的功能上,模型扩展带来的收益有限:270M 模型可以匹敌 1.7B 模型,而最佳整体性能出现在 0.6B 规模。在留出功能上,FT 在构造上准确率为零,而 SIP 能够泛化并随规模显著提升。在超出范围的请求上,FT 可能调用其被训练输出的某个不可用功能,而 SIP 能更可靠地基于所提供的功能进行拒绝。这种灵活性伴随着更高的内存占用和延迟。我们的理论分析解释了 SIP 为何能实现泛化,以及更长的模式上下文为何会增加推理成本。总体而言,决定基于 SLM 的车辆功能调用能力与失败模式的是功能表面的表示方式,而非单纯的模型规模。
cs.LG / 20 / 2609.09478

Unthrottling the Tanh Jacobian in SAC: A Negative Result on Bang-Bang Control and MetaDrive

释放SAC中的tanh雅可比:关于Bang-Bang控制与MetaDrive的负面结果
Shamass, Faiq
Abstract
Soft Actor-Critic (SAC) represents a continuous policy as an unbounded Gaussian that is squashed by tanh. The Jacobian of that map is $\partial a/\partial u = 1-a^2$, which vanishes as $|a|\to 1$. A natural concern is that this throttle starves the actor of critic signal exactly where extreme actions (full brake, full throttle) are optimal. We test a minimal intervention that restores the missing signal: one extra term in the actor loss whose gradient on the pre-tanh mean is the detached action-gradient of $Q$, with no gain parameter. On a minimum-time double integrator whose optimum is bang-bang at the action bounds, vanilla SAC already reaches near-optimal return ($-31.6$ vs. a calibrated optimum of $-30.3$) across ten paired seeds. An ungated bypass does saturate the policy (99% of eval steps with $|a|\ge 0.9$) and collapses return to $-195.5$. A gated bypass that fires only on the flat shoulder $|a|\in[0.9,0.999]$ also fails, and does so without leaving a saturated policy. Warm-started MetaDrive fine-tuning shows the same pattern: the bypass does not improve return, and where collision rate falls it is typically traded for out-of-road departures. Auto-tuned entropy coefficient rises against the bypass, which is a push toward the tails. The Jacobian effect is real. Treating it as a bug to be undone is not free, and on the tasks studied here it is not helpful. Saturating a bound is not the same as solving a problem whose optimum lives on that bound.
Chinese Translation
软演员-评论家算法(Soft Actor-Critic, SAC)将连续策略表示为一个无界高斯分布,并通过tanh函数进行压缩。该映射的雅可比为 $\partial a/\partial u = 1-a^2$,当 $|a|\to 1$ 时其值趋于零。一个自然的担忧是,这种抑制恰恰在最优解为极端动作(全力制动、全油门)时,切断了演员网络接收评论家信号的机会。我们测试了一种恢复缺失信号的最小干预:在演员损失中加入一个额外项,其对tanh前均值的梯度为 $Q$ 的动作梯度的分离(detached)值,且不引入增益参数。在一个最优解为动作边界处bang-bang控制的最小时间双积分器任务上,原始SAC在十组配对随机种子下已达到接近最优的回报($-31.6$,而校准最优值为 $-30.3$)。无门控的旁路(bypass)确实使策略饱和(99%的评估步数中 $|a|\ge 0.9$),但使回报崩溃至 $-195.5$。仅在平坦肩部 $|a|\in[0.9,0.999]$ 触发的门控旁路同样失败,且并未使策略饱和。暖启动的MetaDrive微调呈现出相同模式:旁路未能改善回报,且碰撞率下降的地方通常以冲出道路为代价。自动调节的熵系数在旁路作用下升高,这表明优化推向了分布尾部。雅可比效应是真实存在的。将其视为一个待修复的缺陷并非没有代价,且在本文研究的任务中并无帮助。使某个动作边界饱和并不等同于求解一个最优解恰好位于该边界上的问题。
cs.LG / 21 / 2609.09486

Efficient Fairness Auditing Across Guidance Scales in Text-to-Image Diffusion Models via Causal Abstraction

基于因果抽象的文本到图像扩散模型中跨引导尺度的公平性高效审计
Rahman, Nabila Tasfiha, Chakraborty, Rajatsubhra, Xu, Depeng, Zhang, Lu
Abstract
Fairness auditing of text-to-image diffusion models often requires generating large numbers of images across sampling configurations, making comprehensive evaluation computationally expensive. We propose a causal-abstraction-based audit instrument for efficiently evaluating fairness under interventions on the classifier-free guidance scale. Given a fixed prompt and a target feature function, we represent the diffusion process as a low-level structural causal model and construct a corresponding high-level model over abstract denoising states. We characterize the projected causal structure, establish identifiability of the fairness-relevant interventional query, and provide sufficient conditions under which the high-level model preserves this query. A probabilistic transformer implements the high-level model as an amortized predictor of target-feature distributions across guidance scales. Experiments evaluate distributional fidelity, fairness-query accuracy, and computational efficiency. We present two auditing demonstrations: one using standard Stable Diffusion 1.5 and another using StayFair, a fairness-enhanced Stable Diffusion model, to examine their behavior across guidance scales.
Chinese Translation
对文本到图像扩散模型进行公平性审计通常需要在不同采样配置下生成大量图像,使得全面评估在计算上代价高昂。我们提出了一种基于因果抽象(causal abstraction)的审计工具,用于高效评估在无分类器引导尺度(classifier-free guidance scale)干预下的公平性。给定固定提示词和目标特征函数,我们将扩散过程表示为低层结构因果模型,并在抽象的去噪状态上构建相应的高层模型。我们刻画了投影因果结构,建立了与公平性相关的干预查询的可识别性,并给出了高层模型保持该查询的充分条件。一个概率变换器(probabilistic transformer)将高层模型实现为跨引导尺度的目标特征分布的摊销预测器。实验评估了分布保真度、公平性查询准确率和计算效率。我们展示了两个审计演示:一个使用标准 Stable Diffusion 1.5,另一个使用 StayFair(一种公平性增强的 Stable Diffusion 模型),以考察它们在不同引导尺度下的行为。
cs.LG / 22 / 2609.09547

A Statistical Approach to Estimating Sample Size of Machine Learning Models

一种估计机器学习模型样本量的统计学方法
Phan-Trong, Dat, Gupta, Sunil, Venkatesh, Svetha
Abstract
Sample size determination for machine learning (ML) prediction models is challenging because conventional power analysis typically requires the predictor-outcome relationship and effect structure to be specified a priori. Nonlinear ML models learn complex prediction surfaces that do not admit straightforward analytical power calculations. We propose a framework that approximates nonlinear ML models with localized linear representations and estimates sample size requirements by evaluating statistical power across these local regions.
Chinese Translation
机器学习(ML)预测模型的样本量确定具有挑战性,因为传统的效能分析(power analysis)通常需要预先设定预测变量与结局之间的关系以及效应结构。非线性机器学习模型所学习到的复杂预测面无法进行直接的解析效能计算。我们提出了一个框架,该框架通过局部线性表示来近似非线性机器学习模型,并通过评估各局部区域的统计效能来估计样本量需求。
cs.LG / 23 / 2609.09564

Robust Industrial Cyber Physical Classification Using Neuromorphic Temporal Embeddings and Hybrid SNN XGBoost Under Machine Unlearning Attacks

基于神经形态时间嵌入与混合SNN-XGBoost的鲁棒工业信息物理分类方法(面向机器遗忘攻击)
Kamoona, Ammar, Koushkbaghi, Sajad, Jalili, Mahdi, McTaggart, Peter, Yu, Xinghuo
Abstract
The digitalisation of electrical distribution networks has increased the exposure of power-grid infrastructure to cyber attacks. Existing intrusion detection systems (IDSs), however, often rely on computationally expensive deep learning models that are difficult to deploy at the edge. Periodic retraining also exposes these systems to machine unlearning attacks, where selective data removal can degrade detection performance. We propose a hybrid Spiking Neural Network (SNN) and XGBoost architecture that combines efficient temporal encoding with a lightweight classifier and provides structural resilience to such attacks. The SNN is trained once on clean data and used as a fixed feature extractor, while only the XGBoost classifier is retrained during model updates. Evaluated on two real-world public power-system datasets, the proposed method achieves 99.9\% accuracy (F1-macro 0.999) on the Synchrophasor dataset and 95.0\% accuracy (F1-macro 0.943) on the MSU/ORNL dataset, outperforming standalone baselines. Under selective label-flipping attacks, the hybrid model loses only 0.9\% F1-macro at 10\% poisoning and delays target-class collapse from 60\% to 70\% poisoning compared with raw models. These results demonstrate that neuromorphic temporal encoding can provide both accurate cyber-attack detection and improved resilience to data poisoning in cyber-physical systems.
Chinese Translation
配电网络的数字化增加了电网基础设施遭受网络攻击的风险。然而,现有的入侵检测系统(IDS)通常依赖于计算成本高昂的深度学习模型,难以在边缘端部署。此外,周期性重训练也使这些系统暴露于机器遗忘攻击之下,即选择性删除数据可能导致检测性能下降。我们提出了一种混合式脉冲神经网络(SNN)与XGBoost架构,将高效的时间编码与轻量级分类器相结合,并对此类攻击具有结构上的鲁棒性。SNN仅在干净数据上训练一次,并作为固定的特征提取器;在模型更新期间只需重新训练XGBoost分类器。在两个真实世界的公开电力系统数据集上的评估表明,所提出的方法在同步相量数据集上达到99.9%的准确率(F1-macro 0.999),在MSU/ORNL数据集上达到95.0%的准确率(F1-macro 0.943),优于单独使用的基线模型。在选择性标签翻转攻击下,与原始模型相比,混合模型在10%投毒率下F1-macro仅下降0.9%,并将目标类崩溃的投毒率阈值从60%推迟至70%。这些结果表明,神经形态时间编码能够同时提供准确的网络攻击检测能力,并提升信息物理系统对数据投毒攻击的抵抗力。
cs.LG / 24 / 2609.09567

Positional task conditioning for scalable defect detection across product families in large product catalogs

面向大型产品目录中跨产品族可扩展缺陷检测的位置任务条件化方法
Satyadharma, Soham, Roccabruna, Gabriel, Khan, Suleiman A.
Abstract
Product families in large product catalogs suffer from inconsistencies such as duplicates and unit mismatches that degrade customer experience. Detecting these requires reasoning over multiple error types across lengthy product listings, where LLM classification quality degrades due to long-context limitations. We address this by decomposing detection into focused sub-tasks that reduce context and isolate error types, improving F1 from 52\% to 87\%. For scalable deployment, we introduce Positional Task Conditioning (PTC), which distills this capability into a single smaller model by reinforcing task identity at structural prompt boundaries. PTC outperforms rationale-based distillation across five models and two architecture families, achieving within 1.79\% F1 of the frontier at upto 98\% lower cost. Our system is deployed across multiple countries processing 10+ million product families.
Chinese Translation
大型产品目录中的产品族普遍存在重复条目和单位不匹配等不一致问题,这些缺陷会损害客户体验。检测此类问题需要对冗长的产品列表中多种错误类型进行推理,而大语言模型(LLM)的分类质量会因长上下文的局限性而下降。为此,我们将检测任务分解为多个聚焦的子任务,以减少上下文长度并隔离错误类型,将F1分数从52%提升至87%。为支持可扩展部署,我们提出了位置任务条件化(Positional Task Conditioning, PTC),通过在提示的结构边界处强化任务标识,将这一能力蒸馏到单个更小的模型中。PTC在五个模型和两类架构上的表现均优于基于推理的蒸馏方法,其F1分数与前沿模型差距在1.79%以内,同时成本最高降低98%。该系统已部署于多个国家,处理超过1000万个产品族。
cs.LG / 25 / 2609.09595

Teacher Geometry Shapes Learnability in Teacher-Student Networks

教师网络的几何结构决定师生网络中的可学习性
Sandbrink, Kai J., Martinelli, Flavio, van Meegen, Alexander, Gerstner, Wulfram, Brea, Johanni
Abstract
Teacher-student systems, in which a teacher neural network generates training labels so that a student neural network can learn to implement the same function, are widely used as an abstract setting to study learning. However, the structure of the teachers is often overlooked by assuming randomly-generated, normally-distributed parameters. This hides substantial variation in how learnable different teachers are. We formalize learnability as the success rate of converging to the global minimum, as a function of overparameterization, learning algorithm, student initialization distribution, and teacher geometry. We both identify an easy distribution that maximizes node dissimilarity and a hard distribution that minimizes it, and show that these two distributions induce markedly different success rates across a large range of settings and for different activation functions. To explain the gap, we study the loss landscape of small neural networks that contain two distinct kinds of suboptimal local minima, out-of-bounds (OOB) minima at the edge of the data distribution and interior minima within. Assuming infinite data and a fast readout layer, we analytically reduce the loss landscape of small networks to two dimensions, showing that the region of attraction of interior minima changes as a function of teacher structure. In larger networks, maximally dissimilar teachers induce more interior minima, while minimally dissimilar teachers induce more OOB minima. Motivated by these analyses, we show that differentially increasing the learning rate of the readout layer and decreasing the learning rate of the inner biases increases success rates. These findings provide an important step in narrowing the gap between the study of teacher-student networks and more structured functions that arise in practice.
Chinese Translation
师生系统中,教师神经网络生成训练标签,使学生神经网络学习实现相同的函数,这种系统被广泛用作研究学习的抽象框架。然而,通过假设参数为随机生成的正态分布,教师网络的结构往往被忽视。这掩盖了不同教师网络在可学习性上的显著差异。我们将可学习性形式化为收敛到全局最小值的成功率,它是过参数化程度、学习算法、学生网络初始化分布以及教师网络几何结构的函数。我们识别了一种最大化节点间差异性的易分布和一种最小化节点间差异性的难分布,并证明在大量设置和不同激活函数下,这两种分布导致截然不同的成功率。为解释这一差距,我们研究了小型神经网络的损失景观,其中包含两种不同类型的次优局部极小值:位于数据分布边缘的越界(out-of-bounds, OOB)极小值和位于分布内部的内部极小值。假设数据无限且读出层读取迅速,我们将小型网络的损失景观解析地约减到二维,表明内部极小值的吸引域随教师网络结构的变化而变化。在更大的网络中,节点差异性最大的教师网络诱导更多的内部极小值,而节点差异性最小的教师网络诱导更多的越界极小值。基于这些分析,我们证明差异化地提高读出层的学习率并降低内层偏置的学习率可以提高成功率。这些发现为缩小师生网络研究与实践中出现的更具结构化函数之间的差距迈出了重要一步。
cs.LG / 26 / 2609.09659

Cascading Gradient Inversion via LT-Code Inspired Peeling in Federated Learning

联邦学习中基于LT码启发式剥离的级联梯度反演
Shariati, Saeed, Meybodi, Mohsen Alambardar
Abstract
Federated learning shares model updates rather than raw data, yet these updates can be inverted to reconstruct the clients' training data. Analytic reconstruction attacks, which invert a gradient in closed form, degrade as the batch grows: prior single-round attacks recover only about half of a batch of size $100$ even when the attacker fully controls the network parameters, and known upper bounds limit what any such method can recover. We establish a connection between gradient inversion and the theory of erasure-correcting codes, and use it to construct attacks that exceed these bounds. Our attacks recover batches exactly, together with every sample's label, from a single FedSGD round, and certify each recovery without ground-truth data. On eight image and tabular benchmarks they outperform prior single-round attacks by a wide margin. Even a passive attacker who only observes an honestly trained network recovers $94$--$100\%$ of ImageNet batches at sizes up to $128$, more than prior single-round attacks achieve even with active manipulation of the model, and in the active setting more than $90\%$ is recovered at batch sizes of several hundred. These results show that the privacy leakage of federated learning has been underestimated.
Chinese Translation
联邦学习共享的是模型更新而非原始数据,然而这些更新可被反演以重建客户端的训练数据。解析式重建攻击(analytic reconstruction attacks)以闭式解形式反演梯度,但随批大小增大其性能下降:即便攻击者完全控制网络参数,此前的单轮攻击在批大小为 $100$ 时也只能恢复约一半的样本,且已知的理论上界限制了任何此类方法所能恢复的内容。我们建立了梯度反演与纠删除码(erasure-correcting codes)理论之间的联系,并据此构造出超越这些上界的攻击。我们的攻击可从单轮 FedSGD 中精确恢复整个批次及每个样本的标签,并且无需真实数据即可对每次恢复进行认证。在八个图像和表格数据基准上,这些攻击大幅超越了此前的单轮攻击。即使是被动的攻击者——仅观察一个诚实训练的网络——也能在批大小高达 $128$ 时恢复 $94$ 至 $100\%$ 的 ImageNet 批次,超过了此前单轮攻击即使在主动操纵模型情况下所能达到的效果;而在主动攻击场景下,批大小达数百时仍可恢复超过 $90\%$。这些结果表明,联邦学习的隐私泄露风险一直被低估了。
cs.LG / 27 / 2609.09662

PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling

PELM:基于投机解码与动态电压频率调节的高能效端侧大语言模型推理
Yang, Weisi, Xia, Stephen
Abstract
Deploying Large Language Models (LLMs) directly on mobile platforms at the edge is gaining traction due to a myriad of benefits, such as increased privacy, personalization, and reduced latency. However, LLMs have heavy computational requirements, which are difficult for resource-constrained mobile and edge platforms to fulfill. In addition to limited compute resources, mobile and edge systems often have a compact form factor and lack physical mechanisms to dissipate heat generated from high processor usage rates (e.g., fans) to prevent throttling and reduced processing power, which LLMs can easily cause. To mitigate these effects, prior works have proposed various power governing strategies, such as dynamic voltage and frequency scaling (DVFS), for reducing power and heat generation for heavy computational tasks on mobile platforms. Recently, DVFS methods tailored for mobile LLMs have also been proposed. However, these methods mostly focus on optimizing hardware parameters and processor frequencies, and they fall short under some thermally constrained scenarios. Drawing from recent advances in machine learning, we identify and take advantage of the key insight that not all tokens require full-depth inference to maintain high-quality generation. Motivated by this, we present PELM, a solution that augments traditional DVFS processor frequency tuning with two additional workload-specific knobs: 1) speculative decoding and 2) variable verification depth to expand the optimization space to multiple dimensions for more power efficient on-device LLM inference. In extensive evaluations across hardware platforms and datasets, PELM demonstrates superior performance compared to state-of-the-art power governing methods, with up to 23.1% speedup and 52.4% reduction in energy consumption, while maintaining comparable task performance. The source code is available at https://github.com/imec-nu/PELM.
Chinese Translation
由于在隐私保护、个性化以及降低延迟等方面的诸多优势,在边缘移动平台上直接部署大语言模型(LLM)正日益受到关注。然而,大语言模型对计算能力的要求很高,资源受限的移动和边缘平台难以满足。除了计算资源有限之外,移动和边缘系统的外形通常较为紧凑,缺乏散热机制(如风扇)来散除高处理器使用率所产生的热量,以避免降频和算力下降,而大语言模型很容易引发这些问题。为了缓解这些影响,已有研究提出了各种功耗管理策略,例如动态电压频率调节(DVFS),以降低移动平台上重计算任务的功耗与发热。近期,也出现了针对移动端大语言模型定制的DVFS方法。然而,这些方法大多集中于优化硬件参数和处理器频率,在某些热约束场景下表现不佳。借鉴机器学习领域的最新进展,我们发现并利用了一个关键洞察:并非所有token都需要完整深度的推理才能维持高质量生成。基于此,我们提出了PELM,该方案在传统DVFS处理器频率调优的基础上,增加了两个额外的工作负载相关调节手段:1)投机解码(speculative decoding)和2)可变验证深度,从而将优化空间扩展到多个维度,实现更高能效的端侧大语言模型推理。在多种硬件平台和数据集上的大量评估表明,PELM相较最先进的功耗管理方法表现出更优的性能,加速比最高可达23.1%,能耗降低最多达52.4%,同时保持了相当的任务性能。源代码可在 https://github.com/imec-nu/PELM 获取。
cs.LG / 28 / 2609.09676

Muon-C: Operator-Aligned Muon for Convolutional Kernels

Muon-C:面向卷积核的算子对齐Muon优化器
Qing, Jiaxin, Li, Lexin
Abstract
Muon replaces matrix momentum with an approximately orthogonal polar direction, but its geometry depends on the matrix representation. For convolution, standard unfolding describes a local patch map rather than the convolution operator. We introduce Muon-C, an operator-aligned optimizer that represents kernel momentum as frequency-wise channel-transfer matrices, polarizes these blocks independently, and uses a critical Fourier grid to return updates exactly to the original finite kernel support. We show that the new geometry arises from combining the block partition and Fourier coordinates. The exact-polar direction is a linear minimization oracle under the critically sampled convolution norm. Its worst-case guarantee relative to the continuous convolution-operator norm is never weaker than unfolding and is strictly stronger for $3\times3$ kernels. On CIFAR-10 flow matching with matched applied-update RMS, Muon-C reaches 9.87 FID at 40k iterations, compared with 22.26 for unfolded Muon and 51.31 for Adam. It reaches their final quality using $0.62\times$ and $0.64\times$ their model FLOPs, respectively. Under equal tuning budgets, Muon-C achieves 3.42 FID. Gains persist across data scales and transfer to classification across convolutional architectures.
Chinese Translation
Muon用近似的正交极化方向替代矩阵动量,但其几何结构依赖于矩阵的表示方式。对于卷积而言,标准的展开(unfolding)方式描述的是局部图像块的映射,而非卷积算子本身。我们提出Muon-C,一种算子对齐的优化器,它将卷积核动量表示为按频率划分的通道传递矩阵,对这些分块独立进行极化分解,并利用临界傅里叶网格将更新精确地映射回原始的有限卷积核支撑域。我们证明这一新几何结构源于分块划分与傅里叶坐标的结合。该精确极化方向在临界采样卷积范数下是一个线性最小化预言机。相对于连续卷积算子范数,其最坏情况保证绝不弱于展开方法,且对于$3\times3$卷积核严格更强。在CIFAR-10流匹配任务中,在施加更新RMS相同的条件下,Muon-C在40k次迭代时达到9.87 FID,而展开版Muon和Adam分别为22.26和51.31。Muon-C分别仅使用它们0.62倍和0.64倍的模型FLOPs即可达到它们的最终质量。在相同调参预算下,Muon-C达到3.42 FID。其优势在不同数据规模下均能保持,并可迁移至多种卷积架构上的分类任务。
cs.LG / 29 / 2609.09682

Settling: Equilibrium Inference for Non-Convex Validity Sets

Settling:面向非凸有效集合的均衡推断方法
Saoud, Lyes Saad
Abstract
Many learning systems return a single point estimate even when admissible outputs form disconnected or non-convex sets. Under squared loss, an ambiguous conditional distribution can therefore have a Bayes-optimal conditional mean that is invalid. We formalize this failure as conditional mean collapse and introduce Settling, an equilibrium-based inference operator that separates proposal generation, consistency evaluation, and test-time equilibrium selection. The operator treats a mean-seeking proposal as an initialization and refines it toward a locally stable configuration; conditional on initialization, refinement is deterministic. We establish exact-gradient descent, local convergence, and an inexact-gradient robustness condition relevant to learned consistency critics. In a reproducible 100-context geometric diagnostic, the mean-seeking baseline succeeds in 0/100 contexts, stochastic denoising in 100/100, and Settling in 99/100 while producing substantially lower trajectory roughness. A 1,200-run sensitivity study yields 97-100% success across obstacle-jitter ranges up to 0.20 and 94-100% across one-time initialization perturbations from 0.05 to 0.50. Cross-domain panels remain mechanism illustrations; learned high-dimensional validation remains an open empirical test.
Chinese Translation
在许多学习系统中,即使可容许的输出构成不连通或非凸的集合,系统仍只返回单点估计。在平方损失下,一个含糊不清的条件分布的贝叶斯最优条件均值可能是无效的。我们将这一失效形式化为条件均值坍缩,并提出Settling——一种基于均衡的推断算子,它将提议生成、一致性评估以及测试时的均衡选择分离开来。该算子将均值导向的提议视为初始化,并将其向局部稳定构型细化;在给定初始化的条件下,细化过程是确定性的。我们建立了精确梯度下降、局部收敛性结果,以及与学习型一致性评价器相关的非精确梯度鲁棒性条件。在一个可复现的包含100个情境的几何诊断中,均值导向的基线在0/100个情境中成功,随机去噪在100/100个情境中成功,而Settling在99/100个情境中成功,同时产生显著更低的轨迹粗糙度。一项1200次运行的敏感性研究显示,在障碍抖动幅度高达0.20的范围内成功率为97-100%,在0.05至0.50的一次性初始化扰动范围内成功率为94-100%。跨领域实验面板仍仅为机制说明;学习型高维验证仍是一个有待检验的开放性实证问题。
cs.LG / 30 / 2609.09685

ALIGN-HOLD: Experience Alignment for Real-Time Hold Control in Large-Scale Ride-Hailing Matching at DiDi

ALIGN-HOLD:滴滴大规模网约车匹配中基于经验对齐的实时持单控制
Zhang, Zuhao, Liu, Xu, Wan, Kai, Lu, Zihao, Ma, Li, Li, Shuai
Abstract
Real-time hold control is a high-leverage mechanism in large-scale ride-hailing systems: by selectively deferring driver-order pairs, the platform can wait for better matching opportunities and improve end-to-end passenger-driver experience. Existing production systems such as EXHOLD learn bandit-based hold policies from handcrafted combinations of trip completion, cancellations, waiting time, and driver effort. However, designing such rewards becomes increasingly difficult as marketplace preferences are heterogeneous and observed passenger-driver behavior can be sparse, noisy, and affected by dynamic supply-demand conditions. We present ALIGN-HOLD, a production-scale experience alignment framework that learns hold policy from implicit marketplace preferences. ALIGN-HOLD constructs complementary preference pairs from order trajectories, driver trajectories, and contemporaneous local matching graphs, and trains an experience Reward Model (RM) using balanced multi-view sampling and model-adaptive hard preference sampling. During simulator-based policy learning, the frozen RM provides a dense, context-dependent reward and supports label-free filtering of low-identifiability interactions whose behavioral feedback is difficult to attribute to matching quality. We deploy ALIGN-HOLD on DiDi's ride-hailing platform and evaluate it in a 28-day randomized A/B experiment, covering approximately 100,000 passenger requests per day. Compared with the deployed production policy, ALIGN-HOLD achieves statistically significant improvements in trip completion rate and driver income, while significantly reducing passenger cancellations before and after driver acceptance. Complementary ablations, RM diagnostics, and behavioral analyses validate the contributions of the proposed components. ALIGN-HOLD has been fully ramped up and is currently serving DiDi's Brazil marketplace.
Chinese Translation
实时持单控制是大规模网约车系统中一种高杠杆机制:通过选择性地延迟司机-订单配对,平台可以等待更优的匹配机会,从而提升端到端的乘客-司机体验。现有的生产系统(如EXHOLD)基于行程完成率、取消率、等待时间和司机付出等手工设计的组合,通过多臂老虎机方法学习持单策略。然而,随着市场偏好的异质化以及观测到的乘客-司机行为变得稀疏、含噪且受动态供需条件影响,设计此类奖励函数变得愈发困难。我们提出了ALIGN-HOLD,一个生产规模的经验对齐框架,能够从隐式市场偏好中学习持单策略。ALIGN-HOLD利用订单轨迹、司机轨迹以及同时刻的局部匹配图构建互补偏好对,并采用平衡的多视图采样和模型自适应的困难偏好采样来训练经验奖励模型(Reward Model, RM)。在基于模拟器的策略学习过程中,冻结的RM提供密集且依赖上下文的奖励,并支持无标签过滤那些行为反馈难以归因于匹配质量的低可辨识性交互。我们在滴滴网约车平台上部署了ALIGN-HOLD,并开展了为期28天的随机化A/B实验进行评估,覆盖每天约10万个乘客请求。与已部署的生产策略相比,ALIGN-HOLD在行程完成率和司机收入上取得了统计显著的提升,同时显著降低了司机接单前后乘客的取消率。补充的消融实验、RM诊断和行为分析验证了所提出各组件的贡献。ALIGN-HOLD已全量上线,目前正在服务滴滴巴西市场。
cs.LG / 31 / 2609.09698

Kernel-Complexity Edge Sanitization for Training-Free Defense against Structural Graph Attacks

面向免训练防御结构化图攻击的核复杂度边清理方法
Jia, Yaning, Deng, Shenyang, Yang, Yaoqing, Ma, Chiyu, Xu, Wenxuan, Vosoughi, Soroush
Abstract
Graph Neural Networks (GNNs) have achieved remarkable success across diverse applications, yet they remain highly vulnerable to adversarial attacks that maliciously perturb graph structure. Existing defenses often lack rigorous theoretical grounding, rely on attack-specific heuristics, or require costly retraining procedures such as adversarial training. To address these limitations, we propose Kernel-Complexity Edge Sanitization (KCES), a training-free and model-agnostic framework for defending against structural attacks. KCES is built upon Graph Kernel Complexity (GKC), a principled metric derived from the graph Gram matrix that appears in a generalization upper bound on the GNN test error. From this bound, we define an edge-specific KC score that quantifies each edge's structural influence via its induced change in GKC. KCES then identifies and prunes high-KC edges, which are empirically enriched with adversarial perturbations under structural attacks, to mitigate their harmful impact. Computationally efficient and scalable, KCES operates as a lightweight preprocessing step without retraining and can be seamlessly integrated with existing defenses. Extensive experiments demonstrate that KCES consistently outperforms representative robust baselines across diverse attack settings and scales effectively to large graphs. Supported by theoretical analysis and extensive empirical validation, KCES provides a principled and efficient framework for securing GNNs. Our code is available at https://github.com/karpning/KCScore.
Chinese Translation
图神经网络(GNN)在众多应用中取得了显著成功,但仍然极易受到恶意扰动图结构的对抗攻击。现有防御方法往往缺乏严格的理论基础,依赖于特定攻击的启发式策略,或需要代价高昂的重训练过程(如对抗训练)。为解决这些局限性,我们提出了核复杂度边清理(Kernel-Complexity Edge Sanitization, KCES),一个免训练且与模型无关的结构化攻击防御框架。KCES建立在图核复杂度(Graph Kernel Complexity, GKC)之上,该指标源于GNN测试误差泛化上界中出现的图Gram矩阵。基于该上界,我们定义了针对每条边的KCES分数,通过其对GKC引起的变化来量化每条边的结构影响。KCES进而识别并剪除高KCES分数的边——这些边在结构化攻击下经验上富集了对抗扰动——从而缓解其有害影响。KCES计算高效且可扩展,作为一种轻量级预处理步骤运行,无需重训练,并可与现有防御方法无缝集成。大量实验表明,KCES在多种攻击设置下始终优于具有代表性的鲁棒基线方法,并能有效扩展至大规模图。在理论分析和大量实证验证的支持下,KCES为保障GNN安全提供了一个有原则且高效的框架。我们的代码发布于 https://github.com/karpning/KCScore。
cs.LG / 32 / 2609.09721

EFQ-Softmax: Exp-Free Quantization for Softmax

EFQ-Softmax:面向Softmax的无指数量化方法
Han, Haohui, Wan, Yuming, Wang, Hongni, Xie, Pengcheng, Yan, Xiaodong, You, Runqi, Zhang, Wencong
Abstract
Low-bit attention accelerates Transformer inference by moving the $QK^\top$ and $PV$ matrix multiplications to FP8 or FP4 matrix engines. However, the softmax path often evaluates shifted-score exponentials in higher precision, forms a temporary probability block, and quantizes it before low-bit $PV$ multiplication. This exp-then-quantize path creates a mismatch between a high-precision probability producer and a low-bit matrix consumer. We propose EFQ-Softmax (Exp-Free Quantization for Softmax), a low-bit probability-generation method that directly maps shifted attention scores to block-scaled E2M1 operands. For each microscaling block, EFQ-Softmax selects an exponent-only scale from the local maximum, maps the shifted scores to a normalized residual domain, and generates nonnegative E2M1 probability codes using a single affine rule. The resulting operand is used consistently in both the $\widetilde{P}V$ numerator update and the $\widetilde{P}\mathbf{1}$ denominator update. The FlashAttention-style row-maximum update, historical rescaling, high-precision accumulation, and final normalization remain unchanged. We evaluate end-to-end quality on Qwen3-8B, Qwen3-VL-8B-Instruct, and WAN2.2-TI2V-5B, and separately measure kernel-level performance on the A5 vector unit. EFQ-Softmax improves the Qwen3-8B seven-task mean from 0.6749 with MXFP4 to 0.6773 and the Qwen3-VL nine-task mean from 0.7826 to 0.8000. On WAN2.2, it maintains temporal consistency and visual quality comparable to the FP16 and MXFP4 baselines under VBench. On the A5 vector unit, EFQ-Softmax reduces the vector-stage latency of the fused probability-generation kernel by 40.33% on average across sequence lengths from 16K to 128K. These results show that direct low-bit probability generation can replace the conventional exp-then-quantize path while preserving end-to-end model quality.
Chinese Translation
低比特注意力通过将 $QK^\top$ 和 $PV$ 矩阵乘法迁移至FP8或FP4矩阵引擎来加速Transformer推理。然而,softmax路径通常需要以较高精度计算带平移的分数指数、生成临时概率块,并在低比特 $PV$ 乘法之前对其进行量化。这种先算指数再量化的路径造成了高精度概率生产者与低比特矩阵消费者之间的失配。我们提出EFQ-Softmax(Exp-Free Quantization for Softmax),一种低比特概率生成方法,可直接将带平移的注意力分数映射为按块缩放的E2M1操作数。对于每个微缩放(microscaling)块,EFQ-Softmax从局部最大值中选取仅含指数部分的缩放因子,将平移后的分数映射到归一化残差域,并通过单一线性仿射规则生成非负的E2M1概率码。所得操作数在 $\widetilde{P}V$ 分子更新和 $\widetilde{P}\mathbf{1}$ 分母更新中得到一致使用。FlashAttention风格的行最大值更新、历史重缩放、高精度累加以及最终归一化均保持不变。我们在Qwen3-8B、Qwen3-VL-8B-Instruct和WAN2.2-TI2V-5B上评估了端到端质量,并在A5向量单元上单独测量了核函数级性能。EFQ-Softmax将Qwen3-8B的七任务平均分从MXFP4的0.6749提升至0.6773,将Qwen3-VL的九任务平均分从0.7826提升至0.8000。在WAN2.2上,根据VBench评测,其时间一致性和视觉质量与FP16及MXFP4基线相当。在A5向量单元上,EFQ-Softmax使融合概率生成核的向量阶段延迟在16K至128K序列长度范围内平均降低40.33%。这些结果表明,直接的低比特概率生成可以在保持端到端模型质量的同时,取代传统的先算指数再量化路径。
cs.LG / 33 / 2609.09728

EEGBind: Detecting Source-Level Interictal Epileptiform Discharges via EEG-Centric Multimodal Binding

EEGBind:基于以脑电为中心的多模态绑定检测源级发作间期癫痫样放电
Li, Muchen, Liu, Anglin, Gao, Xuetian, Xu, Ruijian, Chen, Jintai
Abstract
Source-level analysis of interictal epileptiform discharges (IEDs) is relevant to presurgical evaluation and treatment planning because it helps characterize where epileptiform activity is likely to arise. Beyond detecting whether an IED is present, this setting requires assigning IED-positive activity to clinically meaningful brain-region categories. This setting is challenging because source-region evidence in short electroencephalography (EEG) windows can be subtle, partial, and affected by subject variability, class imbalance, and imperfect multimodal context. We present EEGBind, an EEG-centric multimodal binding framework for five-class source-level IED classification. EEGBind treats EEG as the primary modality and binds synchronized video-context features around an EEG-centric representation. Instead of relying on early or overly strong multimodal fusion, which may perturb the source-sensitive EEG representation, EEGBind uses video context as auxiliary evidence for robust classification. A view-consistent repair stage is further used to improve hidden-set robustness while preserving the learned source-class boundary. On the NeuroMM 2026 Grand Challenge Track 3 NMM-Source-IED benchmark, EEGBind achieves 0.8395 on weighted-F1 and outperforms strong competitors. These results support EEG-centric multimodal binding as a practical strategy for source-level IED classification. The open-source code is available at https://github.com/HKUSTGZ-ML4Health-Lab/NeuroMM2026_IED_Detection.
Chinese Translation
发作间期癫痫样放电(IED)的源级分析对于术前评估和治疗规划具有重要意义,因为它有助于刻画癫痫样活动可能起源的位置。除检测IED是否存在之外,该任务还需要将IED阳性活动归入具有临床意义的脑区类别。这一任务极具挑战性,因为在短时程脑电(EEG)窗口中,源区域证据可能较为微弱且不完整,并受到被试差异、类别不平衡以及多模态上下文信息不完善等因素的影响。我们提出了EEGBind,一个以EEG为中心的多模态绑定框架,用于五分类的源级IED分类。EEGBind将EEG视为主要模态,并围绕以EEG为中心的表征绑定同步的视频上下文特征。不同于可能扰动EEG源敏感表征的早期融合或过强的多模态融合,EEGBind将视频上下文作为辅助证据以实现鲁棒分类。此外,还引入了视角一致性修复阶段,以在保持已学习的源类别边界的同时提升对隐藏测试集的鲁棒性。在NeuroMM 2026 Grand Challenge Track 3的NMM-Source-IED基准上,EEGBind的加权F1分数达到0.8395,超越了众多强有力的竞争方法。这些结果表明,以EEG为中心的多模态绑定是源级IED分类的一种实用策略。开源代码已发布于 https://github.com/HKUSTGZ-ML4Health-Lab/NeuroMM2026_IED_Detection。
cs.LG / 34 / 2609.09768

Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?

微调支持KV缓存拼接的模型还是重新计算KV缓存?为何不两者兼顾?
Tachibana, Fumihiko, Miyashita, Daisuke, Deguchi, Jun
Abstract
In Retrieval-Augmented Generation (RAG) systems, a large number of retrieved chunks are concatenated to form the input context so that users can receive high-quality responses based on external knowledge. As a result, the input context length increases substantially, leading to a larger prefill workload and, in turn, a longer time to first token (TTFT). While previous works that reuse precomputed key-value (KV) caches effectively reduce TTFT for long-context inputs, it remains unclear whether response quality is preserved when the input context becomes very long. In this paper, we propose a combined approach that (i) fine-tunes the model while taking KV cache concatenation into account and (ii) selectively recomputes a subset of the KV caches. By applying both techniques, we demonstrate improved accuracy for long-context inputs. Experiments on the RULER benchmark show that, for a 124k-token input, our method improves the RULER score by 9.7 point over the baseline that recomputes KV caches only. Moreover, TTFT is reduced by 80% compared with full attention.
Chinese Translation
在检索增强生成(RAG)系统中,大量检索到的文本块被拼接起来构成输入上下文,使用户能够基于外部知识获得高质量的回答。然而,这导致输入上下文长度大幅增加,进而带来更大的预填充(prefill)工作负载以及更长的首词生成时间(TTFT)。尽管先前复用预先计算的键值(KV)缓存的工作能够有效降低长上下文输入的TTFT,但当输入上下文变得非常长时,回答质量是否得以保持仍不清楚。在本文中,我们提出了一种结合两种策略的方法:(i)在微调模型时考虑KV缓存的拼接;(ii)选择性地重新计算一部分KV缓存。通过同时应用这两种技术,我们证明了该方法可以提升长上下文输入的准确率。在RULER基准上的实验表明,对于124k词元的输入,与仅重新计算KV缓存的基线方法相比,我们的方法将RULER分数提高了9.7分。此外,与全注意力计算相比,TTFT降低了80%。
cs.LG / 35 / 2609.09783

BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL

BRACE:面向异步强化学习中滞后评论家的锚定贝尔曼残差校正方法
Zhao, Guanqun, Xie, Zijun, Zheng, Binbin, Lu, Jiafeng, Gong, Enlei, Chen, Zeyu
Abstract
Asynchronous reinforcement learning has become the standard way to scale training for language models, but the resulting policy lag biases the critic toward the stale behavior policy. Existing work on asynchronous LLM training corrects the actor and leaves this bias unaddressed, while the off-policy value correction of classical RL does not carry over to long-horizon agentic tasks, since a short correction horizon leaves the regression target free of the reward and a long one lets the product of importance ratios drift exponentially with the trajectory length. We propose BRACE, an anchored Bellman-residual correction for stale value models. BRACE bounds the correction horizon to a prefix of policy tokens and anchors a constant-weight Monte-Carlo tail beyond it, which separates policy correction from reward propagation. BRACE improves mean@1 on BrowseComp-Plus by $2.4\%$ over the strongest baseline, runs $2.46\times$ faster per step than synchronous training, and remains stable $50$ updates off-policy.
Chinese Translation
异步强化学习已成为扩展语言模型训练规模的标准方式,但由此产生的策略滞后会使评论家(critic)偏向过时的行为策略。现有的异步大语言模型训练工作仅对执行者(actor)进行校正,而未解决这一偏差;同时,经典强化学习中的离策略价值校正方法也难以直接应用于长时程智能体任务,因为较短的校正时程会使回归目标不受奖励影响,而较长的校正时程则会使重要性比率的乘积随轨迹长度呈指数级漂移。我们提出BRACE,一种针对滞后价值模型的锚定贝尔曼残差校正方法。BRACE将校正时程限定在策略token的一个前缀上,并在其后锚定一个常数权重的蒙特卡洛尾部,从而将策略校正与奖励传播分离开来。BRACE在BrowseComp-Plus上将mean@1较最强基线提升了2.4%,每步运行速度比同步训练快2.46倍,并在离策略50次更新后仍保持稳定。
cs.LG / 36 / 2609.09786

NEXUS-MI: Communication-Aware Federated Personalization for Gateway-Coordinated Motor-Imagery Brain-Computer Interfaces

NEXUS-MI:面向网关协调的运动想象脑机接口的通信感知联邦个性化方法
Worae, Daniel Adu, Nagarajan, Aarthy
Abstract
Electroencephalography (EEG)-based motor-imagery brain-computer interfaces (MI-BCIs) vary across subjects and sessions, complicating personalization from limited calibration data. Federated learning can exploit shared representations without centralizing raw EEG, but existing federated MI studies largely assume regular synchronization. We introduce NEXUS-MI, a gateway-coordinated federated personalization framework that treats synchronization as a coupled learning-and-communication control problem. Raw EEG and classifier heads remain local, while an edge coordinator maintains the shared backbone. We evaluate NEXUS-MI through offline replay using BCI Competition IV Dataset 2a (BCICIV-2a; 9 subjects, 4 classes) and OpenBMI (54 subjects, 2 classes). Session 1 supports backbone learning, and Session 2 provides limited-calibration personalization and held-out testing. An ideal-link reference and six heterogeneous-link policies characterize gateway participation, buffering, stale-update admission, and backbone-download control. The principal comparison holds delayed-update handling fixed while contrasting non-adaptive and communication-aware synchronization. Paired subject-level comparisons use Holm adjustment, and robustness across five matched realizations is assessed by hierarchical bootstrap. Communication-aware coordination reduced server-to-client backbone traffic by approximately 42% on both datasets, while cohort-level accuracy differences were small and realization-dependent. Cohort averages also concealed subject-level vulnerability, with losses reaching approximately 12 percentage points on BCICIV-2a relative to the ideal-link reference. These findings establish gateway synchronization as an explicit design variable in federated MI personalization and motivate joint evaluation of personalized accuracy, communication cost, update freshness, and subject-level reliability.
Chinese Translation
基于脑电图(EEG)的运动想象脑机接口(MI-BCI)在受试者和会话之间存在差异,使得在有限校准数据下进行个性化面临困难。联邦学习可以在无需集中原始脑电数据的情况下利用共享表示,但现有的联邦MI研究大多假设同步是规则的。我们提出了NEXUS-MI,一种网关协调的联邦个性化框架,将同步视为一个学习与通信耦合的控制问题。原始脑电数据和分类器头保留在本地,而边缘协调器维护共享主干网络。我们使用BCI竞赛IV数据集2a(BCICIV-2a;9名受试者,4类)和OpenBMI(54名受试者,2类)通过离线回放评估NEXUS-MI。会话1用于主干网络学习,会话2用于有限校准的个性化和保留测试。一个理想链路参考和六种异构链路策略刻画了网关参与、缓冲、过期更新准入和主干下载控制。主要对比在固定延迟更新处理方式的前提下,比较非自适应同步与通信感知同步。受试者层面的配对比较采用Holm校正,并通过分层自助法评估五个匹配实现中的鲁棒性。通信感知协调在两个数据集上将服务器到客户端的主干传输流量减少了约42%,而群体层面的准确率差异较小且依赖于具体实现。群体平均值还掩盖了受试者层面的脆弱性,在BCICIV-2a上相对于理想链路参考的损失高达约12个百分点。这些发现确立了网关同步作为联邦MI个性化中的一个显式设计变量,并推动了对个性化准确率、通信成本、更新新鲜度和受试者层面可靠性的联合评估。
cs.LG / 37 / 2609.09788

Evaluating Model Retraining under Drift: Paired Comparisons of Cumulative Subgroup Disparity

漂移下模型重训练评估:累积子群体差异的配对比较
Ceross, Aaron
Abstract
Choosing when to retrain a deployed classifier requires assessing subgroup error rates across the sequence of models used, including periods between updates. We compare complete scheduled, loss-triggered, and subgroup-gap-triggered policies with retaining the initial model on the same observations and delayed labels. For true-positive and false-positive rates separately, the outcome is the paired difference in absolute subgroup gaps summed over deployment windows. Population evaluation in simulation, action records, and alternative schedules assess how measurement and retraining behaviour affect these comparisons. In a follow-up sample of 400 new trajectories per condition across two simulated drift regimes, all three policies had lower mean cumulative disparity, equivalent to reductions of 0.04 to 0.88 percentage points in the average gap per window. Evaluating the unchanged models against the known generating distributions preserved all mean directions, but finite-window and population comparisons agreed on whether updating increased, reduced or left cumulative disparity unchanged in 69 to 92 percent of trajectories. Under subgroup-specific drift, smaller true-positive-rate gaps accompanied lower sensitivity in both groups. In an exploratory American Community Survey replay, person weighting reversed all three race false-positive-rate mean comparisons without changing predictions or actions; all three weighted intervals included zero. Policy comparisons require group-specific rates, action distributions, and an explicit evaluation population alongside mean disparity. These analyses are non-confirmatory. Shared replay requires policy-independent observations and complete labels after the specified delay.
Chinese Translation
决定何时重训练已部署的分类器,需要对所使用的一系列模型(包括两次更新之间的时段)的子群体错误率进行评估。我们将完整定期重训练、损失触发和子群体差距触发策略与在同一观测数据和延迟标签上保留初始模型进行比较。对于真阳性率和假阳性率分别而言,结果为各部署窗口内子群体绝对差距之和的配对差值。仿真中的总体评估、行为记录以及替代调度方案用于考察测量和重训练行为如何影响这些比较。在两种模拟漂移情形下,每种条件各400条新轨迹的后续样本中,三种策略的平均累积差异均较低,相当于每个窗口平均差距减少0.04至0.88个百分点。将不变的模型与已知的生成分布进行评估可保留所有均值方向,但有限窗口比较与总体比较在69%至92%的轨迹上对更新是增加、减少还是未改变累积差异的判断一致。在子群体特异性漂移下,真阳性率差距的缩小伴随着两组敏感度的降低。在一项探索性的美国社区调查(American Community Survey)回放中,个体加权在未改变预测或行动的情况下逆转了所有三项种族假阳性率均值比较;所有三项加权区间均包含零。策略比较需要在平均差异之外,同时考察群体特定比率、行动分布以及明确的评估总体。这些分析是非验证性的。共享回放需要与策略无关的观测数据以及指定延迟之后的完整标签。
cs.LG / 38 / 2609.09794

Privacy-Preserving Split Learning for Federated LLM Fine-Tuning

面向联邦大语言模型微调的隐私保护拆分学习
Jin, Heng, Zhang, Chaoyu, Yu, Hexuan, Lou, Wenjing, Hou, Y. Thomas
Abstract
Fine-tuning large language models (LLMs) on domain-specific data is essential for downstream adaptation. In many deployments, a participant cannot hold the complete model locally. This happens because the model owner keeps the full model proprietary, or because the participant lacks sufficient compute resources. Split Learning (SL) addresses this by partitioning the model between the participant and a server so that only a small portion runs locally. When the underlying data is additionally distributed across multiple institutions with privacy requirements, Federated Learning (FL) further enables collaborative training across participants by sharing only model updates instead of raw data. In this combined setting, each client transmits intermediate activations to the server, and for LLM fine-tuning, this exchange poses an inherent privacy paradox. The autoregressive nature of LLMs causes the transmitted activations to leak the input, and existing perturbation-based defenses are fundamentally ineffective in this setting. We address this leakage through a learned obfuscate-and-recover scheme that protects participants' private datasets while still allowing an independently deployable model to be trained on the server side. Experiments demonstrate that our approach achieves strong privacy protection with modest utility loss and system overhead, making split-based federated LLM fine-tuning practically viable.
Chinese Translation
在领域特定数据上微调大语言模型(LLM)对下游适配至关重要。在许多部署场景中,参与方无法在本地持有完整模型,这可能是因为模型所有者将完整模型作为专有资产保留,也可能是因为参与方缺乏足够的计算资源。拆分学习(Split Learning, SL)通过在参与方与服务器之间对模型进行划分来应对这一问题,使本地仅需运行模型的一小部分。当底层数据进一步分布在多个有隐私要求的机构之间时,联邦学习(Federated Learning, FL)通过只共享模型更新而非原始数据,进一步实现了参与方之间的协同训练。在这种组合场景下,每个客户端需要将中间激活值传输给服务器,而对于大语言模型微调而言,这种交换带来了固有的隐私悖论:LLM的自回归特性导致传输的激活值会泄露输入信息,而现有基于扰动的防御方法在这种场景下根本无效。我们通过一种学习式的混淆-恢复方案来解决这一泄露问题,在保护参与方私有数据集的同时,仍能在服务器端训练一个可独立部署的模型。实验表明,我们的方法在付出适度的效用损失和系统开销的前提下实现了强大的隐私保护,使基于拆分的联邦LLM微调在实践中切实可行。
cs.LG / 39 / 2609.09796

A practical DIRECT-type algorithm for medium-scale black-box global optimization

一种面向中等规模黑箱全局优化的实用型DIRECT类算法
Stripinis, Linas, Paulavičius, Remigijus
Abstract
The DIRECT algorithm is a deterministic global optimization method known for its versatility and balanced exploration-exploitation strategy. However, DIRECT-type algorithms are primarily effective for low-dimensional problems and often exhibit slow convergence as dimensionality increases, limiting their applicability to more complex optimization tasks. To address this limitation, this paper introduces X-DTC-GL, a novel DIRECT-type algorithm that incorporates dynamic partitioning and hybridization techniques. The dynamic partitioning approach adaptively refines the search space based on local one-dimensional surrogate models, enabling rapid subdivision of promising hyper-rectangles. The hybridization strategy selectively employs a hill-climbing method to exploit promising regions identified by the surrogate models. Extensive experiments on four diverse benchmark suites demonstrate that X-DTC-GL significantly outperforms existing DIRECT-type baselines, achieving improvements of ~12% in solvability and ~27% in solution quality. Performance-profile analyses indicate the fastest convergence on up to ~40% of instances, the best runtime performance on ~17% of problems, and competitive overall execution times. By improving performance within the partition-based framework, these advances strengthen the algorithm's competitiveness in state-of-the-art black-box optimization.
Chinese Translation
DIRECT算法是一种确定性全局优化方法,以其通用性和探索-利用的均衡策略而著称。然而,DIRECT类算法主要对低维问题有效,且随着维数的增加往往收敛缓慢,这限制了其在更复杂优化任务中的应用。为解决这一局限,本文提出了一种新颖的DIRECT类算法X-DTC-GL,该算法融合了动态划分与混合技术。动态划分方法基于局部一维代理模型自适应地细化搜索空间,从而能够快速细分有潜力的超矩形。混合策略则选择性地采用爬山法来开发由代理模型识别出的有潜力区域。在四个不同的基准测试集上的大量实验表明,X-DTC-GL显著优于现有的DIRECT类基准算法,在可解性方面提升约12%,在解质量方面提升约27%。性能剖面分析表明,该算法在多达约40%的实例上收敛速度最快,在约17%的问题上具有最佳的运行时间性能,且总体执行时间具有竞争力。通过在基于划分的框架内提升性能,这些进展增强了该算法在最先进黑箱优化领域的竞争力。
cs.LG / 40 / 2609.09809

Online Inverse Integer Linear Optimization via Small-Gradient Skipping: Constant Regret and Finite Mistakes

基于小梯度跳跃的在线逆整数线性优化:常数后悔与有限错误次数
Kitaoka, Akira
Abstract
In online inverse linear optimization, the learner predicts a weight at each round, observes the optimal action of the agent, and updates its prediction. In the general setting, the gap of $\log T$ between the regret upper bound $O(d \log T)$ and the lower bound $\Omega(d)$ is unresolved (here $T$ is the total number of rounds and $d$ is the dimension). When the action set is M-convex, the regret is known to be bounded by $O(d \log d)$, but the method attaining it computes a center of gravity at every round. This paper therefore proposes Small-Gradient Skipping (SGS), a mechanism that skips the update at rounds without a mistake in the case where the correct action is uniformly separated from the other candidates, and applies it to online gradient descent, the online Newton step, and MetaGrad. The number of mistakes is then bounded, for all three, by a quantity independent of $T$; and for the online Newton step and for MetaGrad with SGS, the dimension dependence of the regret becomes $O(d^2)$ when the forward problem is an integer linear program, that is, the factor $\log T$ is removed. Moreover, when the action set is M-convex, the regret is bounded efficiently without computing a center of gravity.
Chinese Translation
在在线逆线性优化中,学习者在每一轮预测一个权重,观察智能体的最优动作,并更新其预测。在一般情形下,后悔上界 $O(d \log T)$ 与下界 $\Omega(d)$ 之间的 $\log T$ 差距尚未解决(其中 $T$ 为总轮数,$d$ 为维度)。当动作集为 M-凸时,已知后悔可被 $O(d \log d)$ 所界定,但实现该界的方法需要在每一轮计算重心。为此,本文提出了小梯度跳跃机制(Small-Gradient Skipping, SGS),在正确动作与其他候选动作均匀分离的情形下,于没有出现错误的轮次跳过更新,并将其应用于在线梯度下降、在线牛顿步(online Newton step)以及 MetaGrad。对于这三种方法,错误次数均被一个与 $T$ 无关的量所界定;且当使用 SGS 的在线牛顿步和 MetaGrad 时,当前向问题为整数线性规划时,后悔对维度的依赖变为 $O(d^2)$,即消除了 $\log T$ 因子。此外,当动作集为 M-凸时,无需计算重心即可高效地界定后悔。
cs.LG / 41 / 2609.09824

In Medical Claims Data, Enhancing Predictive Performance for Major Adverse Cardiovascular Events Using Cross Attention

在医疗索赔数据中利用交叉注意力机制提升主要不良心血管事件的预测性能
Fujioka, Yuhei, Misawa, Daitaro, Ikenoue, Tatsuyoshi, Fukuma, Shingo
Abstract
Medical claims data comprise the financial details, including the expenses and billing information, as well as the clinical information, such as the diagnoses and treatments, of patients visiting medical facilities. Recently, it has been acknowledged that large databases can be constructed from medical claims data for medical research purposes. However, the clinical information within these datasets is often medically unstructured, limiting its application in comprehensive analyses. This study enhances predictive model performance for major adverse cardiovascular events (MACE), a leading cause of death worldwide. Models that predict MACE are crucial to clinical practice guidelines. We utilize a cross-attention mechanism to develop a method that effectively weights the relationships between diagnoses and treatments. Effectively repre- senting the clinical information contained in medical claims data, this approach generates more representative features for predicting MACE. The ROC-AUC score of our proposed cross-attention-based model was 0.7720, higher than other benchmark models including the conventional atherosclerotic cardiovascular disease model, the light gradient boosting machine, and a self-attention-based model. These results indicate that integrating the clinical structure of medical claims data using a cross-attention mechanism significantly enhances the performance of predictive models.
Chinese Translation
医疗索赔数据包含患者就诊于医疗机构的财务信息(如费用和账单信息)以及临床信息(如诊断和治疗)。近年来,人们已经认识到可以基于医疗索赔数据构建大型数据库以用于医学研究。然而,这些数据集中的临床信息通常在医学上是非结构化的,限制了其在综合分析中的应用。本研究旨在提升主要不良心血管事件(MACE)预测模型的性能,MACE是全球主要致死原因之一。预测MACE的模型对临床实践指南至关重要。我们利用交叉注意力(cross-attention)机制开发了一种方法,能够有效地对诊断与治疗之间的关系进行加权。该方法有效表示了医疗索赔数据中包含的临床信息,为预测MACE生成了更具代表性的特征。我们提出的基于交叉注意力的模型ROC-AUC得分为0.7720,高于其他基准模型,包括传统的动脉粥样硬化性心血管疾病模型、LightGBM(轻量梯度提升机)以及基于自注意力机制的模型。这些结果表明,利用交叉注意力机制整合医疗索赔数据的临床结构,能够显著提升预测模型的性能。
cs.LG / 42 / 2609.09840

TempTPI: Informer-Based trajectory prediction for maritime vessels

TempTPI:基于Informer的海上船舶轨迹预测
Ferneding, Kevin, Lietavcova, Veronika, Blachowiak, Aleksandra M., Heiselberg, Peder
Abstract
Accurate long-term trajectory prediction for maritime vessels is essential for safety and logistical efficiency. While deep learning models, particularly Transformers, have shown promise in processing Automatic Identification System (AIS) data, they often struggle with the quadratic computational complexity of self-attention and the loss of accuracy over extended forecasting horizons. This study proposes TempTPI, a novel prediction framework that integrates an Informer-based encoder with a multi-channel temporal encoding mechanism. The Informer architecture leverages a ProbSparse self-attention mechanism to reduce computational overhead and focus on the most significant dependencies, while the temporal encoder utilizes Fourier-like frequency expansions to capture cyclic patterns (hourly, daily, and seasonal) in vessel behavior. We evaluate our model against the state-of-the-art TPTrans architecture using AIS data from Danish waters. Experimental results demonstrate that TempTPI consistently outperforms existing methods across prediction windows of 1 to 5 hours. Notably, at a 5-hour horizon, the proposed model achieves a 55% improvement in Mean Squared Error (MSE), offering a robust solution for long-range maritime situational awareness.
Chinese Translation
海上船舶的精确长期轨迹预测对于安全性和物流效率至关重要。尽管深度学习模型,尤其是Transformer,在处理自动识别系统(AIS)数据方面已展现出良好的前景,但它们常常受困于自注意力机制的二次方计算复杂度以及在较长预测时间范围内精度下降的问题。本研究提出了TempTPI,一种将基于Informer的编码器与多通道时间编码机制相集成的新型预测框架。Informer架构利用ProbSparse自注意力机制来降低计算开销并聚焦于最关键的依赖关系,而时间编码器则利用类傅里叶的频率展开来捕捉船舶行为中的周期性模式(每小时、每天和季节性)。我们使用丹麦水域的AIS数据,将所提模型与最先进的TPTrans架构进行了对比评估。实验结果表明,TempTPI在1至5小时的预测窗口内均持续优于现有方法。值得注意的是,在5小时预测时间范围内,所提模型在均方误差(MSE)上实现了55%的改进,为远程海上态势感知提供了一种稳健的解决方案。
cs.LG / 43 / 2609.09860

Exact Degeneracy Under Balanced k-Shot Sampling:Consequences for Small-Sample Discriminant Analysis on LLM Embeddings

平衡k-shot采样下的精确简并:对LLM嵌入上小样本判别分析的影响
Qu, Lingxiao
Abstract
Balanced k-shot sampling draws exactly k labeled examples per class. We show that it induces an exact, provable degeneracy in a family of small-sample discriminant estimators. Under balanced sampling, the within-class scatter operator of Kernelized Linear Principal Component Discriminant Analysis (KLPCDA) is not merely rank-deficient but exactly a scaled orthogonal projector. We derive the consequences in closed form: two of KLPCDA's seven variants have every signal eigenvalue exactly equal, so their eigenvector selection criterion is provably indifferent rather than ill-conditioned, and a third has a provably void objective. This follows from the estimators' construction, not any dataset; we confirm it on frozen sentence embeddings and, separately, on residual-stream activations from a decoder-only generative model. An in-formula tie-break repairs the two repairable variants, with recovery gated by class count: the residual subspace constraint costs 5x more on few-class than many-class datasets (p=0.000001). We then evaluate the repaired framework on few-shot text classification on frozen LLM embeddings (n much smaller than d, up to 4096), across four datasets, three embedding sizes, and three trained baselines (SetFit, LoRA, in-context learning). A properly cross-validated logistic-regression probe still beats every KLPCDA variant on three of four datasets, at every embedding size; guidance carried from pixel, vibration-signal, and gene-expression data does not directly generalize to this feature space. Three independent geometric separability metrics fail to explain why one high-dimensional decoder-based embedding model underperforms smaller bidirectional encoders, ruling out anisotropy; the gap is substantially an estimation-efficiency effect, not a permanent ceiling, closing by more than 80% when the support set grows from k<=10 to k=30-50 (p=0.00195, both many-class datasets).
Chinese Translation
平衡k-shot采样为每个类别精确抽取k个带标签样本。我们证明,这类采样会在一族小样本判别估计量中引发一种可被严格证明的精确简并。在平衡采样下,核化线性主成分判别分析(KLPCDA)的类内散布算子不仅是秩亏的,而且精确地等于一个缩放的正交投影算子。我们以闭式形式推导了其后果:KLPCDA的七个变体中有两个的所有信号特征值精确相等,因此其特征向量选择准则是可证明的无差异性而非病态性;第三个变体则具有可证明的空洞目标函数。这一结论源于估计量的构造本身,与具体数据集无关;我们在冻结的句子嵌入上以及(分别地)在仅解码器生成式模型的残差流激活上验证了这一点。一种公式内的平局打破机制可修复其中两个可修复的变体,其恢复效果受类别数量的制约:在少类别数据集上,残差子空间约束的代价是多类别数据集的5倍(p=0.000001)。随后,我们在冻结LLM嵌入(n远小于d,最高至4096维)上评估了修复后的框架在小样本文本分类中的表现,涵盖四个数据集、三种嵌入规模以及三个经过训练的基线方法(SetFit、LoRA、上下文学习)。在四个数据集中的三个上,经过恰当交叉验证的逻辑回归探针在所有嵌入规模下仍优于所有KLPCDA变体;从像素、振动信号和基因表达数据中获得的经验指导并不能直接推广到这一特征空间。三种独立的几何可分性度量均无法解释为何某个高维基于解码器的嵌入模型表现不及更小的双向编码器,从而排除了各向异性的解释;该性能差距在很大程度上是一个估计效率问题,而非永久的性能上限——当支持集从k<=10增长到k=30-50时,差距缩小超过80%(p=0.00195,两个多类别数据集均如此)。
cs.LG / 44 / 2609.09883

Forward-Free LLM Depth Pruning via Weight Redundancy

基于权重冗余的前向无关大语言模型深度剪枝
Yun, Vincent-Daniel, Lim, Woosang
Abstract
Depth pruning reduces large language model (LLM) inference cost by removing complete Transformer blocks. Activation-based methods collect hidden states through forward passes on calibration data, while existing forward-free methods score each Transformer block separately without measuring similarity between blocks. We propose Weight-Redundancy Pruning (WRP), a forward-free depth-pruning method that estimates inter-layer redundancy from checkpoint weights to select blocks without calibration data or model forward passes. WRP compares attention output and MLP down-projection weights across layers and combines their pairwise similarities with relative projection-scale information. The resulting all-pairs similarity matrix guides layer grouping and block selection. Across multiple pruning settings, model families, and downstream tasks, WRP consistently outperforms existing forward-free magnitude pruning and approaches the performance of activation-based methods.
Chinese Translation
深度剪枝通过移除完整的Transformer块来降低大语言模型(LLM)的推理成本。基于激活值的方法需要在校准数据上进行前向传播以收集隐藏状态,而现有的前向无关(forward-free)方法则独立地为每个Transformer块打分,无法衡量块之间的相似性。我们提出权重冗余剪枝(Weight-Redundancy Pruning, WRP),这是一种前向无关的深度剪枝方法,它直接从模型检查点(checkpoint)的权重中估计层间冗余,从而在无需校准数据或模型前向传播的情况下选择待剪枝的块。WRP比较各层之间的注意力输出权重和MLP下投影权重,并将它们的成对相似度与相对投影尺度信息相结合。由此得到的全对(all-pairs)相似度矩阵用于指导层分组和块选择。在多种剪枝设置、模型系列和下游任务上,WRP始终优于现有的前向无关幅度剪枝方法,并接近基于激活值方法的性能。
cs.LG / 45 / 2609.09891

ProMeta: Few-shot PROTAC-targeted degradation prediction across E3 ligases

ProMeta:跨E3连接酶的小样本PROTAC靶向降解预测
Liu, Yuansheng, Ye, Yufei, Tang, Tao, Luo, Jiawei, Tao, Wen, Luo, Xiao
Abstract
Proteolysis-targeting chimeras (PROTACs) have emerged as a transformative therapeutic strategy that selectively degrades historically ''undruggable'' targets via the ubiquitin-proteasome system. Despite growing efforts to develop computational predictors of PROTAC degradation activity, existing supervised approaches remain severely challenged by data scarcity and imbalance across E3 ligases, limiting their ability to generalize beyond well-studied ligase contexts. In practice, labeled data are heavily concentrated on a few ligases (e.g., CRBN and VHL), while the majority of E3 ligases remain underexplored yet are critical for expanding the design space of targeted degraders. Developing methods that enable robust cross-ligase generalization with minimal labeled data is therefore essential for improving the practical utility of computational PROTAC discovery. We reformulate PROTAC degradation activity prediction across E3 ligases as a few-shot meta-learning problem and present ProMeta, a prototype-based graph neural network trained through episodic meta-learning on source-E3 tasks and evaluated on held-out target-E3 tasks through support-conditioned inference. ProMeta performs inference without updating the encoder by dynamically estimating class prototypes from minimal target-ligase support samples. On the CRBN-to-VHL benchmark, ProMeta achieves AUROC values of 0.796 under K=2, Q=3 and 0.883 under K=2, Q=5, improving by 19.9% and 6.8%, respectively, over the corresponding supervised GNN baseline. Reverse VHL-to-CRBN transfer under the same protocol yielded AUROC values of 0.702 (K=2, Q=3) and 0.821 (K=2, Q=5), confirming bidirectional applicability while revealing direction and data-regime dependence. Together, these results support ProMeta as a practical framework for cross-ligase few-shot prediction under the evaluated support/query protocols.
Chinese Translation
蛋白水解靶向嵌合体(PROTACs)已成为一种变革性的治疗策略,能够通过泛素-蛋白酶体系统选择性降解历史上"不可成药"的靶点。尽管开发PROTAC降解活性计算预测器的努力日益增多,但现有的监督学习方法仍严重受限于E3连接酶数据的稀缺性和不平衡性,难以泛化到研究充分的连接酶之外。在实践中,有标签数据高度集中于少数连接酶(如CRBN和VHL),而大多数E3连接酶虽然尚未被充分探索,却对扩展靶向降解剂的设计空间至关重要。因此,开发能够在极少标注数据下实现稳健跨连接酶泛化的方法,对于提升计算PROTAC发现的实用价值至关重要。我们将跨E3连接酶的PROTAC降解活性预测重新构建为一个小样本元学习问题,并提出ProMeta——一种基于原型的图神经网络,通过在源E3连接酶任务上进行情景式元学习训练,并通过支持集条件推断在留出的目标E3连接酶任务上进行评估。ProMeta无需更新编码器即可进行推断,通过从极少量的目标连接酶支持样本中动态估计类别原型来实现。在CRBN到VHL的基准测试中,ProMeta在K=2、Q=3条件下达到0.796的AUROC,在K=2、Q=5条件下达到0.883,分别比相应的监督式GNN基线提升了19.9%和6.8%。在相同协议下进行反向的VHL到CRBN迁移,AUROC分别为0.702(K=2,Q=3)和0.821(K=2,Q=5),证实了双向适用性,同时揭示了方向和数据规模的依赖性。总体而言,这些结果支持ProMeta作为在所评估的支持/查询协议下进行跨连接酶小样本预测的实用框架。
cs.LG / 46 / 2609.09899

Strangers to Themselves: What Language Models Say About Themselves Is Generic

自我的陌生人:语言模型对自身的描述是泛化的
Blandfort, Phil, Pawar, Urja
Abstract
Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question. We find that: (i) Direct self-report is weak (r = +0.04), and even showing the model the exact items only raises prediction to +0.24. Crucially, the same item-informed question about "capable AI agents in general" does just as well (+0.28), while other models' answers about themselves predict the target model at least as well as its own. (ii) Frontier scale does not detectably change this pattern: any gains in prediction are not self-specific, and are consistent with a better theory of how AI assistants behave rather than better self-knowledge. (iii) First-person framing does have one robust effect: it shifts reports in the flattering direction, understating harmful behavior relative to the same question about a generic agent. (iv) Finetuning on a model's own behavioral record can teach narrow self-predictions, but it also changes the behavior being predicted and the gains do not transfer broadly. The practical implication is simple: asking a model what it would do mostly reveals a theory of AI assistants in general, plus a favorable bias, rather than privileged knowledge of that model.
Chinese Translation
语言模型能够流利地描述自己会如何行事:是否会屈服于反对压力、滥用工具,或在压力下撒谎。但这些描述真的关乎正在发言的那个模型吗?我们将自我知识转化为一个预测测试。在九项行为评估中,我们测量模型在不同条件下的实际行为,要求其预测这些行为发生率,并将其预测与去除“自我”因素的对照组进行比较。我们发现:(i)直接自我报告的预测能力很弱(r = +0.04),即使向模型展示确切的题目,其预测也仅提升至 +0.24。关键在于,同样基于题目的、针对“有能力的人工智能助手的总体情况”的提问表现相当(+0.28),而其他模型对自身的回答预测目标模型的效果至少不亚于目标模型对自身的回答。(ii)前沿规模并未显著改变这一模式:预测上的任何提升并非针对自我的,且更符合关于AI助手一般行为规律的更优理论,而非更好的自我知识。(iii)第一人称的表述确实有一个稳定的影响:它会使报告向有利于自身的方向偏移,相对于关于一般助手的相同问题,低估有害行为。(iv)在模型自身的行为记录上进行微调可以教会模型进行狭窄的自我预测,但这同时改变了被预测的行为本身,且这些收益无法广泛迁移。其实践启示很简单:询问一个模型它会怎么做,大多揭示的是关于AI助手总体的一般性理论,外加一种有利于自身的偏差,而非关于该模型的独有知识。
cs.LG / 47 / 2609.09904

Beyond Conventional Federated Learning via High-Order Regularization

通过高阶正则化突破传统联邦学习
Kabgani, Alireza, Ahookhosh, Masoud
Abstract
Federated clients that perform several local optimization steps can return parameter displacements with widely different magnitudes. The quadratic regularization of FedProx grows linearly with displacement and therefore offers limited control over the contrast between ordinary and unusually large client movements. We here introduce HiFedProx, which replaces the quadratic penalty with a scale-matched power-type regularizer indexed by $p\geq2$. All powers have the same regularization-gradient magnitude at a reference displacement $R$, while every $p>2$ gives a weaker response below $R$ and a stronger response above it. An exact affine reference calculation shows that increasing $p$ compresses relative displacement disparities, although very large powers approach fixed-radius behavior and increase local curvature. HiFedProx combines this geometry with finite-budget stochastic client optimization and same-minibatch Armijo backtracking. In paired five-seed experiments on a frozen 60-writer FEMNIST subset, a common-parameter study over $p\in\{2,3,4,5,6,7,8\}$ shows similar clean-training performance but substantial gains under composite stress. The lowest moderate- and severe-stress losses occur at $p=7$ and $p=6$, improving over $p=2$ by $11.44\%$ and $23.16\%$, respectively. Although displacement-tail ratios continue to decrease through $p=8$, predictive performance peaks in an intermediate range and Armijo trial cost increases with $p$. These results indicate that the exponent should be calibrated rather than maximized. In our experiments, $p=5$--$7$ provides the most useful range.
Chinese Translation
执行多次本地优化步骤的联邦客户端所返回的参数位移,其幅度可能差异巨大。FedProx 的二次正则化随位移呈线性增长,因此对普通客户端移动与异常过大移动之间的对比只能提供有限的控制。本文提出 HiFedProx,用由 $p\geq2$ 索引的尺度匹配幂型正则化器替代二次惩罚。所有幂次在参考位移 $R$ 处具有相同的正则化梯度幅度,而每个 $p>2$ 在低于 $R$ 时响应更弱、高于 $R$ 时响应更强。精确的仿射参考计算表明,增大 $p$ 会压缩相对位移差异,尽管非常大的幂次会趋近固定半径行为并增加局部曲率。HiFedProx 将这一几何特性与有限预算的随机客户端优化以及同一小批次的 Armijo 回溯相结合。在固定的 60 位书写者 FEMNIST 子集上的配对五种子实验中,针对 $p\in\{2,3,4,5,6,7,8\}$ 的统一参数研究表明,各设置在正常训练下性能相近,但在复合压力下差异显著。中等压力与严重压力下的最低损失分别出现在 $p=7$ 和 $p=6$,相比 $p=2$ 分别提升了 $11.44\%$ 和 $23.16\%$。尽管位移尾部比率随 $p$ 增至 8 仍在持续下降,但预测性能在中间区间达到峰值,且 Armijo 试验成本随 $p$ 增加。这些结果表明,指数应经过校准而非一味取大。在我们的实验中,$p=5$ 至 $7$ 提供了最有用的取值范围。
cs.LG / 48 / 2609.09907

Meta-LinEXP3: Online-within-Online Learning for Adversarial Linear Contextual Bandits

Meta-LinEXP3:面向对抗性线性上下文老虎机的在线中套在线学习
Li, Hao, Xu, Jie, Xie, Zheng
Abstract
Meta-learning has emerged as an effective paradigm for transferring knowledge across sequential bandit tasks. While substantial progress has been made for stochastic bandits and non-contextual adversarial bandits, meta-learning for adversarial linear contextual bandits (ALCBs) with random action sets remains largely unexplored. To address this problem, we propose Meta-LinEXP3, an online-within-online algorithm that constructs a predictable task-level prior from completed tasks to guide the inner LinEXP3 learner. For known context distributions, we develop a policy-centered estimator that achieves an intrinsic-dimension $\mathcal{O}(\sqrt{n})$ per-task regret bound. For unknown distributions, we introduce a past-only regularized moment estimator with an $\mathcal{O}(n^{2/3})$ leading regret term and explicit finite-sample error. We further establish a direct connection between prior accuracy and transfer regret, showing that increasingly accurate priors yield sublinear transfer-dependent regret across tasks. Experiments demonstrate the effectiveness of Meta-LinEXP3, including its application to structured hyperspectral tensor sampling.
Chinese Translation
元学习已成为在序列老虎机任务之间迁移知识的一种有效范式。尽管在随机老虎机和非上下文对抗性老虎机方面已取得显著进展,但针对具有随机动作集的对抗性线性上下文老虎机(ALCBs)的元学习研究仍然十分匮乏。为解决这一问题,我们提出了Meta-LinEXP3,这是一种在线中套在线的算法,它从已完成的任务中构建可预测的任务级先验,以指导内部的LinEXP3学习器。对于已知上下文分布的情形,我们开发了一种以策略为中心的估计器,实现了关于内在维度为 $\mathcal{O}(\sqrt{n})$ 的单任务遗憾界。对于未知分布的情形,我们提出了一种仅利用历史信息的正则化矩估计器,其遗憾主导项为 $\mathcal{O}(n^{2/3})$,并具有显式的有限样本误差。我们进一步建立了先验精度与迁移遗憾之间的直接联系,表明随着先验精度的不断提高,跨任务的迁移相关遗憾可实现次线性增长。实验结果证明了Meta-LinEXP3的有效性,包括其在结构化高光谱张量采样中的应用。
cs.LG / 49 / 2609.09910

A Kernel-Based Modular Discriminant Analysis Framework for Small-Sample Learning

一种基于核的模块化判别分析框架用于小样本学习
Qu, Lingxiao, Pei, Yan
Abstract
The small-sample-size (SSS) problem remains a fundamental challenge in machine learning when labeled data are scarce due to cost, accessibility, or ethical constraints. While numerous approaches have been proposed, existing methods often struggle to maintain stable and discriminative representations under high-dimensional and limited-data conditions. Kernelized Linear Principal Component Discriminant Analysis (KLPCDA), a recently proposed modular framework, integrates variance preservation, inter-class separability, and intra-class compactness within a unified kernel space. Although its formulation has shown promising initial results, a systematic understanding of how its components interact across diverse SSS scenarios remains lacking. In this paper, we present a systematic cross-domain study of KLPCDA to characterize the interaction mechanisms among its core objectives. We analyze the behavior of its seven variants across multiple real-world SSS tasks, including hyperspectral image classification, mechanical fault diagnosis, medical diagnosis, and face recognition. Through extensive experiments and ablation studies, we investigate how different objective combinations influence performance under varying conditions such as noise, class imbalance, and high dimensionality. Our analysis reveals consistent patterns in the interaction of the three core objectives variance, between-class, and within-class terms, providing a unified and interpretable understanding of their roles in stabilizing representations and enhancing discrimination in SSS settings. Based on these findings, we further derive practical guidelines for selecting appropriate KLPCDA variants under different data characteristics. Experimental results demonstrate that KLPCDA achieves strong and robust performance across domains, while maintaining low computational complexity suitable for resource-constrained environments.
Chinese Translation
小样本(SSS)问题仍然是机器学习中的一项根本性挑战,尤其是在由于成本、可及性或伦理限制而导致标注数据稀缺的情况下。尽管已有众多方法被提出,但现有方法在高维和有限数据条件下往往难以保持稳定且具有判别性的表示。核化线性主成分判别分析(Kernelized Linear Principal Component Discriminant Analysis, KLPCDA)是近期提出的一种模块化框架,它在统一的核空间中整合了方差保持、类间可分性和类内紧凑性。尽管其公式化已展现出良好的初步结果,但目前仍缺乏对其各组成部分在不同SSS场景下如何相互作用的系统性理解。本文对KLPCDA进行了系统的跨领域研究,以刻画其核心目标之间的相互作用机制。我们分析了其七种变体在多个真实世界SSS任务中的表现,包括高光谱图像分类、机械故障诊断、医学诊断和人脸识别。通过大量实验和消融研究,我们考察了不同的目标组合在噪声、类别不平衡和高维等不同条件下对性能的影响。我们的分析揭示了三个核心目标——方差项、类间项和类内项——相互作用的一致规律,从而对它们在SSS环境中稳定表示和增强判别能力的作用提供了统一且可解释的理解。基于这些发现,我们进一步得出了在不同数据特性下选择合适KLPCDA变体的实用指南。实验结果表明,KLPCDA在各领域均实现了强大且稳健的性能,同时保持了较低的计算复杂度,适用于资源受限的环境。
cs.LG / 50 / 2609.09912

Development and Validation of a Physics-Guided Machine Learning Extrapolation Framework Using a Classical Transient Diffusion Benchmark

基于经典瞬态扩散基准的物理引导机器学习外推框架的开发与验证
Yadav, Ashutosh, Dubey, Alok, Chakraborty, Prodyut Ranjan, Akolekar, Harshal
Abstract
Machine learning models used in engineering are typically trained within limited operating ranges, yet reliable predictions are often required beyond these domains. Consequently, the primary challenge is extrapolation rather than interpolation. Rigorous validation is hindered by the scarcity of data outside the training range. To address this limitation, a novel extrapolation framework is integrated with established machine learning architectures to enable accurate and physically consistent predictions beyond the training domain. The framework is established by systematically evaluating two physics-guided architectures: a Bidirectional Long Short-Term Memory (BiLSTM) network and a Physics-Informed Neural Network (PINN). A classical one-dimensional transient diffusion problem is adopted as a benchmark because its exact analytical solution provides unlimited, reliable data across the spatio-temporal domain, enabling rigorous quantitative validation. The problem is particularly challenging because the solution evolves from an initial singularity through a strongly nonlinear transient regime before approaching a steady-state linear profile. When training data are confined to an intermediate portion of this evolution, backward extrapolation toward the singularity becomes especially demanding. To improve reliability, physics-guided coordinate transformations, boundary-aware learning strategies, and stability-enhancing temporal marching are incorporated. Extrapolation is evaluated using a train-predict-validate-extend strategy, in which validated predictions are recursively added to the training set to progressively extend the prediction horizon. The results demonstrate accurate and physically consistent predictions beyond the training domain, highlighting the framework's potential for engineering applications where data availability is limited.
Chinese Translation
工程中使用的机器学习模型通常在有限的运行范围内进行训练,然而往往需要在这些域之外进行可靠预测。因此,主要挑战在于外推而非插值。由于训练范围之外数据的稀缺,严格的验证受到阻碍。为解决这一局限,本文将一种新颖的外推框架与成熟的机器学习架构相结合,以在训练域之外实现准确且物理一致的预测。该框架通过系统评估两种物理引导架构而建立:双向长短期记忆(BiLSTM)网络和物理信息神经网络(PINN)。采用经典的一维瞬态扩散问题作为基准,因为其精确解析解可在整个时空域上提供无限且可靠的数据,从而实现严格的定量验证。该问题尤其具有挑战性,因为其解从初始奇点开始,经历强非线性的瞬态阶段,然后趋近于稳态线性分布。当训练数据仅限于该演化的中间部分时,向奇点方向的后向外推尤为困难。为提高可靠性,本文引入了物理引导的坐标变换、边界感知学习策略以及增强稳定性的时间推进方法。外推评估采用“训练-预测-验证-扩展”策略,即将经过验证的预测结果递归地加入训练集,以逐步扩展预测范围。结果表明,该框架在训练域之外能够实现准确且物理一致的预测,凸显了其在数据有限的工程应用中的潜力。
cs.LG / 51 / 2609.09920

Multi-Pass, Multi-View Blended Learning for High-Fidelity Volumetric CT Synthesis from Chest X-Rays

基于多通道多视角混合学习的高保真胸部X射线体层CT合成方法
Devecioglu, Ozer Can, Kiranyaz, Serkan, Mazhar, Rashid, Hamid, Tahir, Chowdhury, Muhammad, Gabbouj, Moncef
Abstract
Reconstructing volumetric Computed Tomography (CT) from a single 2D chest radiograph (CXR) is an ill-posed inverse problem, further complicated by the scarcity of paired CXR-CT training data. Prior approaches address this by training on Digitally Reconstructed Radiographs (DRRs), which are synthetic projections derived from CT volumes. However, the domain gap between DRRs and real CXRs limits generalization, often resulting in coarse or anatomically inconsistent reconstructions when applied to clinical images. To address this challenging problem, this study introduces a Multi-Pass Multi-View Blended Learning framework for synthesizing high-fidelity volumetric CT directly from real chest X-ray (CXR) images. The proposed approach progressively decomposes the synthesis task into two distinct, complementary learning stages. Stage 1 is an unsupervised CXR-to-DRR Domain Adaptation, while Stage 2 includes three passes, namely, (a) supervised DRR-to-CT Transformation, (b) unsupervised Multi-View Slice Refinement, followed by (c) Progressive Transfer Learning (PTL). With such a blended learning paradigm, the proposed approach mitigates the synthetic-to-real domain gap while enhancing both the structural integrity and anatomical detail of the final output. On the LIDC-IDRI dataset, where paired DRR-CT ground truth is available for quantitative evaluation, the proposed method improves upon prior methods by up to 14% in PSNR and 7.6% in SSIM. The framework successfully generates structurally consistent and anatomically realistic high-fidelity CT volumes from real CXRs, marking a significant advancement toward clinical viability of CT reconstruction from standard radiographic images.
Chinese Translation
从单张二维胸部X射线影像(CXR)重建体积计算机断层扫描(CT)是一个不适定的逆问题,而成对的CXR-CT训练数据的稀缺进一步加剧了其难度。已有方法通过在数字重建放射影像(DRR)上进行训练来解决这一问题,DRR是由CT体数据生成的合成投影。然而,DRR与真实CXR之间的域差距限制了模型的泛化能力,导致其应用于临床图像时常常产生粗糙或解剖结构不一致的重建结果。为应对这一挑战,本研究提出了一种多通道多视角混合学习框架,可直接从真实胸部X射线图像合成高保真的体积CT。所提出的方法将合成任务逐步分解为两个相互补充的学习阶段:第一阶段为无监督的CXR到DRR的域适应;第二阶段包含三个通道,即(a)有监督的DRR到CT转换、(b)无监督的多视角切片细化,以及(c)渐进式迁移学习(PTL)。借助这种混合学习范式,所提出的方法缓解了合成数据到真实数据的域差距,同时提升了最终输出的结构完整性和解剖细节。在LIDC-IDRI数据集上(该数据集提供成对的DRR-CT真实值可用于定量评估),所提出方法的PSNR较已有方法提升高达14%,SSIM提升7.6%。该框架能够从真实CXR成功生成结构一致且解剖真实的高保真CT体数据,标志着从标准X射线图像进行CT重建向临床可行性迈出了重要一步。
cs.LG / 52 / 2609.09945

Adversarial Training for Tabular Credit Scoring: A Multi-Attack Robustness Evaluation in P2P Lending

面向表格型信用评分的对抗训练:P2P借贷中的多攻击鲁棒性评估
Niewzwaag, Gijs A. F., Veth, Marijn G. S., Massei, Manuele, Machado, Marcos R.
Abstract
Machine learning-based credit scoring is increasingly central to Peer-to-Peer (P2P) lending, yet its resilience to adversarial manipulation, where applicants strategically alter self-reported inputs to secure favourable decisions, remains poorly understood. Most adversarial-robustness evidence comes from image and text domains and evaluates a single attack against a matching defence, offering little guidance on how defences generalise across attack types in tabular credit data. We address this with a systematic train-test robustness benchmark on a large Lending Club subset, spanning three model families (logistic regression, a feed-forward neural network, and a transformer for tabular data) and four attacks confined to applicant-mutable features: Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), Salt-and-Pepper (S&P) noise, and DeepFool, plus a mixed-attack regime. Across a full grid evaluated with stratified cross-validation, adversarial training sharply improves robustness against the attack it is trained on and transfers well within the gradient-based family, but transfers weakly to non-gradient corruption, so single-attack defences overstate real-world resilience. Mixed training delivers the most balanced robustness across heterogeneous attacks while preserving clean-test performance, supporting multi-attack stress testing in credit-model governance.
Chinese Translation
基于机器学习的信用评分在点对点(P2P)借贷中日益核心,但其对对抗性操纵的抵御能力仍鲜为人知——申请人可能策略性地篡改自报输入信息以获取有利的信用决策。现有的对抗鲁棒性证据大多来自图像和文本领域,且通常仅将单一攻击与相应防御进行配对评估,对防御方法在表格型信用数据中跨攻击类型的泛化能力几乎没有指导意义。为此,我们在一个大规模Lending Club数据子集上构建了系统的训练-测试鲁棒性基准,涵盖三类模型(逻辑回归、前馈神经网络以及面向表格数据的Transformer)和四种仅限于申请人可篡改特征的攻击:快速梯度符号法(FGSM)、投影梯度下降(PGD)、椒盐噪声以及DeepFool,并加入了混合攻击场景。在采用分层交叉验证评估的完整攻击-防御网格中,对抗训练显著提升了对训练所用攻击的鲁棒性,且在基于梯度的攻击家族内部迁移良好,但对非梯度类扰动的迁移能力较弱,因此单一攻击防御会高估真实世界的抵御能力。混合攻击训练在异质攻击下提供了最均衡的鲁棒性,同时保持了干净测试集上的性能,支持在信用模型治理中开展多攻击压力测试。
cs.LG / 53 / 2609.09971

Deep Neural Networks for Learning Intent from sEMG Signals to Support Hardware Devices for Post-Stroke Neurorehabilitation

基于表面肌电信号学习运动意图的深度神经网络:支持中风后神经康复硬件设备
Brewster, Zakariyya, Wadhwani, Divy, Yan, Emily, Wang, Aidan, Namgyal, Karma, Xie, Shuting, Konyk, Markiyan, Abdelmaguid, Tala
Abstract
Finger-specific motor intent is a clinically meaningful control signal for post-stroke neurorehabilitation, where residual muscle activity may remain measurable despite weak or incomplete movement. We study five-finger multilabel intent decoding from impaired-arm high-density surface electromyography (sEMG) in PhysioMio, a bilateral longitudinal dataset collected from stroke patients. A common processing protocol aligns movement labels, applies 20--450 Hz Butterworth filtering and Symlet-4 wavelet denoising, segments overlapping 200 ms windows, and extracts twelve time- and frequency-domain descriptors per channel. Direct LSTM, CNN, and GNN baselines reveal complementary behavior: the LSTM attains the highest subset accuracy (0.545), whereas the GNN attains the highest macro F1 (0.706) and macro AUPRC (0.776). Architecture search then identifies CNN-Large as the strongest single-split CNN, with 0.593 subset accuracy and 0.714 macro F1, while CNN-Micro provides a compact architecture for embedded inference. To match a four-sensor hardware design, we retrain CNN-Micro using channels associated with ECRB, ECRL, FDS, and FDP and exclude the ground electrode from model input. Across five seeds, cross-channel knowledge distillation improves the four-channel student over direct training, reaching $0.5219 \pm 0.0114$ subset accuracy, $0.7612 \pm 0.0038$ finger accuracy, and $0.6095 \pm 0.0058$ macro F1. The selected 123K-parameter model accepts nine windows of 48 features and has been exported to ONNX. These results establish a reproducible software path from post-stroke sEMG to compact five-finger intent prediction for subsequent hardware-in-the-loop evaluation.
Chinese Translation
手指特异性运动意图是中风后神经康复中具有临床意义的控制信号,在此类康复中,即使运动微弱或不完全,残余肌肉活动仍可能可被测量。我们研究了来自PhysioMio数据集(一个从中风患者采集的双侧纵向数据集)中患侧手臂高密度表面肌电(sEMG)信号的五指多标签意图解码。一套统一的处理流程对运动标签进行对齐,采用20–450 Hz Butterworth滤波和Symlet-4小波去噪,分割重叠的200毫秒时间窗,并从每个通道提取十二个时域和频域特征描述符。直接的LSTM、CNN和GNN基线模型展现出互补的行为特征:LSTM取得了最高的子集准确率(0.545),而GNN取得了最高的宏平均F1(0.706)和宏平均AUPRC(0.776)。随后的架构搜索确定CNN-Large为最强的单折CNN,其子集准确率为0.593,宏平均F1为0.714,而CNN-Micro则提供了一种适用于嵌入式推理的紧凑架构。为了匹配四传感器硬件设计,我们使用与ECRB、ECRL、FDS和FDP相关的通道重新训练了CNN-Micro,并将地电极从模型输入中排除。在五个随机种子下,跨通道知识蒸馏使四通道学生模型的性能优于直接训练,达到0.5219 ± 0.0114的子集准确率、0.7612 ± 0.0038的手指准确率和0.6095 ± 0.0058的宏平均F1。最终选定的12.3万参数模型接受九个48维特征窗口作为输入,并已导出为ONNX格式。这些结果为从中风后sEMG信号到紧凑的五指意图预测建立了一条可复现的软件路径,可供后续的硬件在环评估使用。
cs.LG / 54 / 2609.10012

An Explainable Machine Learning Framework for Predicting Blood-Brain Barrier Permeability Using Molecular Descriptors

一种基于分子描述子的可解释机器学习框架用于预测血脑屏障通透性
Mahmoudi, Fatemeh
Abstract
Blood-brain barrier (BBB) permeability is a critical determinant in the development of central nervous system therapeutics because it directly influences the ability of drug candidates to reach their target sites within the brain. In this study, an explainable machine learning framework was developed to predict BBB permeability using molecular descriptors generated from the MoleculeNet BBBP dataset with the RDKit cheminformatics toolkit. Fifteen physicochemical descriptors extracted from 2,039 compounds were used to train four supervised machine learning algorithms, including Logistic Regression, Support Vector Machine (SVM), Random Forest, and Extreme Gradient Boosting (XGBoost). Hyperparameter optimization was performed using GridSearchCV, while model interpretability was investigated using SHapley Additive exPlanations (SHAP). Among the evaluated models, the optimized XGBoost classifier achieved the best predictive performance, with an accuracy of 88.97%, a precision of 88.92%, a recall of 97.76%, an F1-score of 93.13%, and a ROC-AUC of 0.9282. Stratified five-fold cross-validation further demonstrated the robustness of the proposed model, yielding a mean ROC-AUC of 0.8982 +/- 0.0130. Feature importance and SHAP analyses consistently identified TPSA, HBD, and LogP as the most influential molecular descriptors governing BBB permeability prediction. Overall, the proposed framework provides an accurate, interpretable, and computationally efficient approach for BBB permeability prediction and may serve as a valuable tool for the early-stage screening of CNS drug candidates.
Chinese Translation
血脑屏障(BBB)通透性是中枢神经系统治疗药物开发中的关键决定因素,因为它直接影响候选药物到达脑内靶点的能力。本研究开发了一个可解释的机器学习框架,利用RDKit化学信息学工具包从MoleculeNet BBBP数据集生成的分子描述子来预测血脑屏障通透性。从2,039个化合物中提取的15个理化描述子被用于训练四种监督机器学习算法,包括逻辑回归、支持向量机(SVM)、随机森林和极端梯度提升(XGBoost)。超参数优化采用GridSearchCV进行,模型可解释性则采用SHapley加性解释方法(SHAP)进行研究。在所评估的模型中,优化后的XGBoost分类器取得了最佳的预测性能,准确率为88.97%,精确率为88.92%,召回率为97.76%,F1分数为93.13%,ROC-AUC为0.9282。分层五折交叉验证进一步证明了所提出模型的稳健性,平均ROC-AUC为0.8982 ± 0.0130。特征重要性分析和SHAP分析一致表明,TPSA、HBD和LogP是决定血脑屏障通透性预测的最具影响力的分子描述子。总体而言,所提出的框架为血脑屏障通透性预测提供了一种准确、可解释且计算高效的方法,可作为中枢神经系统候选药物早期筛选的宝贵工具。
cs.LG / 55 / 2609.10016

MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes

MetroLLM-Bench:将语言模型作为交通信息亭运行时进行评估
Hendriks, Remco
Abstract
We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a language-model judge. We report Tier 1 and the combined score of both tiers. A stratified 75/25 split reserves 717 cases for training-data generation and 238 for held-out evaluation. We evaluate twenty-six models from six vendors, of which twenty-three are ranked. On the held-out partition, a 4B Qwen 3.5 student trained through parameter-efficient fine-tuning (PEFT) exceeds both GPT-5.6 tiers on Tier 1 (91.3 against 90.6 and 90.0) and matches GPT-5.4 full at maximum reasoning effort (91.4), with a 2.6 GB Q4_K_M footprint. Larger 9B and 27B students provide no further Tier 1 improvement over the 4B student at this training scale. Across the four Qwen sizes, the PEFT gain over the corresponding base model decreases from +7.03 points at 2B (three training seeds) to -0.91 at 27B; every seed shows the same direction at every size. A deterministic rule-based baseline reaches 84.6 on Tier 1, with the remaining language-model advantage concentrated in policy adaptation, compound scenarios, accessibility, and temporal reasoning. Muse Glimmer 30B leads the composite ranking, and serving configuration alone moves the Qwen 3.5-to-3.8 comparison by 2.7 Tier 1 points. The benchmark, harness, reproduction guide, and fine-tuned students are released at https://github.com/continker/metrollm-bench.
Chinese Translation
我们提出了 MetroLLM-Bench,一个包含 955 个测试用例的基准,用于测试语言模型作为交通信息亭(transit kiosk)策略层的能力。该基准涵盖六个真实的地铁系统(规模从 37 个车站到 414 个车站不等),以及包括路线规划、票价计算、运行中断、无障碍服务和对抗性输入在内的十一个类别。在每个测试用例中,模型必须调用结构化工具,并提交一个机器可渲染的终端状态,其中包含结果、适用时的逐票票价报价以及信息亭动作。十四个确定性评分组件构成第一层(Tier 1);八个语义质量组件构成第二层(Tier 2),其中六个使用语言模型裁判(language-model judge)。我们报告第一层得分以及两层的综合得分。采用分层 75/25 划分,其中 717 个用例用于训练数据生成,238 个用于留出评估。我们评估了来自六家厂商的二十六个模型,其中二十三个参与了排名。在留出分区上,通过参数高效微调(PEFT)训练的 4B Qwen 3.5 学生模型在第一层上超越了 GPT-5.6 的两个层级(91.3 对 90.6 和 90.0),并在最大推理强度下与 GPT-5.4 full 持平(91.4),且其 Q4_K_M 量化后仅占 2.6 GB 存储空间。在该训练规模下,更大的 9B 和 27B 学生模型在第一层上并未比 4B 学生模型带来进一步提升。在四个 Qwen 规模上,PEFT 相对于对应基座模型的增益从 2B 时的 +7.03 分(三个训练种子)下降到 27B 时的 -0.91 分;在每种规模下,所有种子均呈现相同的方向。一个确定性的基于规则的基线在第一层上达到 84.6 分,语言模型的剩余优势主要集中在策略适配、复合场景、无障碍服务和时间推理上。Muse Glimmer 30B 在综合排名中位居首位,而仅服务配置的差异就会使 Qwen 3.5 与 3.8 的比较结果相差 2.7 个第一层分数。该基准、评测框架、复现指南以及微调后的学生模型已在 https://github.com/continker/metrollm-bench 发布。
cs.LG / 56 / 2609.10017

Structure-Aware Unsupervised Anomaly Detection for Spacecraft Telemetry with Adaptive EVT Thresholding

基于自适应极值理论(EVT)阈值的航天器遥测数据结构感知无监督异常检测
Alcarria, Óscar, Sánchez, Rafael, Sempere, Javier, Torrijos, Pablo, Alfaro, Juan C., Auñón, Juan M., Gámez, José A., Puerta, José M.
Abstract
Operational anomaly detection in spacecraft telemetry typically requires labeled historical anomalies or extended warm-up periods. These requirements are rarely met in practice. We propose an unsupervised, deployment-ready framework that produces predictions from the second month of operation without any labels, prior fault knowledge, or mission-specific tuning. The approach combines incremental monthly retraining, statistical model selection, and adaptive Extreme Value Theory (EVT) thresholding for false alarm control. On the ESA Anomalies Dataset (ESA-AD), it achieves $F_{0.5}=0.700$ on Mission~1 and $F_{0.5}=0.698$ on Mission~2 under strict chronological evaluation.
Chinese Translation
航天器遥测数据中的运行异常检测通常需要带标签的历史异常数据或较长的预热期,而这些条件在实际中很少能够满足。我们提出了一种可直接部署的无监督框架,该框架从运行的第二个月起即可生成预测结果,无需任何标签、先验故障知识或任务特定调参。该方法结合了增量式月度重训练、统计模型选择以及用于误报控制的自适应极值理论(Extreme Value Theory, EVT)阈值方法。在ESA Anomalies数据集(ESA-AD)上,采用严格的时间顺序评估,该方法在Mission 1上取得了$F_{0.5}=0.700$,在Mission 2上取得了$F_{0.5}=0.698$的成绩。
cs.LG / 57 / 2609.10026

Beyond Contact Sensors: Deep learning with Pseudo-Labeling for remote Photoplethysmography

超越接触式传感器:基于伪标签的远程光电容积描记术深度学习方法
Acharya, Bhargav, Hammer, Barbara, Drimalla, Hanna
Abstract
Heart rate is a critical biomarker of health, and remote photoplethysmography (rPPG) enables its contactless estimation from video data for telemedicine applications. Recent advancements in deep learning based rPPG methods achieve state-of-the-art results, outperforming classical signal-processing methods in complex scenarios. However, deep learning methods depend on datasets with precise synchronization between videos and ground truth signals collected via contact sensors, whereas signal-processing-based methods do not. To address this dependence on labeled datasets, which are labor-intensive to collect, we investigate under which circumstances pseudo-labels extracted using unsupervised signal-processing methods can replace contact sensors labels for training deep learning methods. Our systematic evaluations found that for datasets with imperfect synchronization, the pseudo-label approach outperforms supervised training on contact sensors. For datasets with good synchronization, results are mixed: within-dataset evaluation shows no significant difference between training methods, while cross-dataset evaluation favors supervised training. However, removing a single outlier participant significantly improves the pseudo-label approach's cross-dataset performance, highlighting the importance of label quality. These results demonstrate that signal-processing methods can generate valid training signals for deep learning models, reducing dependency on labor-intensive dataset collection while maintaining competitive performance.
Chinese Translation
心率是健康的关键生物指标,远程光电容积描记术(rPPG)能够从视频数据中非接触式地估计心率,从而服务于远程医疗应用。基于深度学习的rPPG方法的最新进展取得了最先进的结果,在复杂场景下优于经典信号处理方法。然而,深度学习方法依赖于视频与通过接触式传感器采集的真实信号之间精确同步的数据集,而基于信号处理的方法则无此要求。为了解决对标注数据集(其采集非常耗费人力)的依赖,我们研究了在何种情况下,使用无监督信号处理方法提取的伪标签可以替代接触式传感器标签来训练深度学习方法。我们的系统性评估发现,对于同步不完善的数据集,伪标签方法优于基于接触式传感器的监督训练。对于同步良好的数据集,结果则较为复杂:数据集内评估显示两种训练方法之间无显著差异,而跨数据集评估则更倾向于监督训练。然而,剔除单个离群受试者后,伪标签方法的跨数据集性能显著提升,这凸显了标签质量的重要性。这些结果表明,信号处理方法可以为深度学习模型生成有效的训练信号,在保持有竞争力的性能的同时,减少对高人力成本数据集采集的依赖。
cs.LG / 58 / 2609.10032

Field-level prediction of mid-plane stress tensor fields in concrete target penetration: a cross-velocity graph neural operator surrogate

混凝土靶体侵彻中面应力张量场的场级预测:一种跨速度图神经算子代理模型
Du, Wenpu, Zhou, Peng, Xia, Yunlong, Xin, Sinuo, Zhang, Congcong, Zhang, Boyang, Zhang, Yi, Xu, Wenzheng
Abstract
Although the impact resistance of concrete has been studied extensively, a framework linking mesoscale heterogeneity to full-field stress-tensor prediction has been lacking. Data were generated with a full-scale aggregate-resolved LS-DYNA model (projectile diameter 45 mm, mass 2.13 kg, target diameter 500 mm x thickness 200 mm, mesh 10 mm), verified against published penetration experiments (Frew 2006, Hanchak 1992, Forrestal 1996) by configuration similarity. The dataset contains six-component stress-tensor fields on the X-Z mid-plane for 400 cases (4 impact velocities x 100 aggregate seeds). Three contributions are reported. First, case-by-case verification of the terminal penetration state delimited the rest-state validity of penetration depth and anchored reliable observables to rigid-body motion and field-level stress evolution. Second, a field-level graph neural operator surrogate learned the time-varying stress-field evolution and evaluated cross-velocity leave-one-out extrapolation. Third, the full-scale, aggregate-resolved, cross-velocity, per-seed database was established as a reproducible resource. Cases at 100, 135 and 200 m/s still moved at window end (negative velocity, i.e. rebound), and only one 165 m/s case arrested. Penetration depth is therefore not reported as a rest-state scalar except for the single arrested case (69.33 mm); nose-node depth differences were confirmed as numerical artifacts of displacement integration after erosion. The single-step relative L2 error was 0.6977, reported honestly; autoregressive rollout from frame 11 to 39 took about 144 ms, a speedup of about 3.6x10^3 to 4.3x10^3 relative to single-core LS-DYNA, reported as application value. Validation is bounded by configuration similarity and field-level self-consistency; the framework is a simulation-trained decision-support method within the studied parameter space.
Chinese Translation
尽管混凝土的抗冲击性能已被广泛研究,但一直缺乏一个将细观非均匀性与全场应力张量预测相联系的研究框架。本研究采用全尺寸骨料解析的LS-DYNA模型(弹体直径45 mm,质量2.13 kg,靶体直径500 mm × 厚度200 mm,网格尺寸10 mm)生成数据,并通过构型相似性将其与已发表的侵彻实验(Frew 2006、Hanchak 1992、Forrestal 1996)进行了验证。数据集包含400个工况(4种冲击速度 × 100个骨料随机种子)在X-Z中面上的六分量应力张量场。本文报告了三方面贡献。第一,通过对侵彻终止状态的逐工况验证,界定了侵彻深度作为静止状态量的有效性,并将可靠观测量锚定于刚体运动和场级应力演化。第二,一个场级图神经算子代理模型学习了时变应力场的演化过程,并进行了跨速度留一法外推评估。第三,建立了全尺寸、骨料解析、跨速度、逐随机种子的完整数据库,作为可复现的研究资源。100、135和200 m/s工况在时间窗口结束时仍处于运动状态(负速度,即回弹),仅有165 m/s的一个工况达到静止。因此,除该单一静止工况(侵彻深度69.33 mm)外,未报告侵彻深度作为静止状态标量;同时确认了弹头节点深度差异是侵蚀后位移积分的数值伪影。单步相对L2误差为0.6977,予以如实报告;从第11帧到第39帧的自回归滚动预测耗时约144 ms,相对于单核LS-DYNA加速约3.6×10^3至4.3×10^3倍,作为应用价值予以报告。验证范围受构型相似性和场级自一致性约束;该框架是在所研究参数空间内一种经仿真训练的决策支持方法。
cs.LG / 59 / 2609.10089

Hybrid Quantum-Classical NLP Classification with Compact Semantic Representations: An Experimental Analysis of Representation Compression

基于紧凑语义表示的量子-经典混合自然语言处理分类:表示压缩的实验分析
Hassan, Ali, Zhao, Zijia, Metawei, Maha A.
Abstract
Large language and sentence-embedding models provide rich semantic representations, but their high dimensionality poses a challenge for near-term quantum machine learning (QML), where quantum circuits can process only a limited number of input features. We investigate a hybrid quantum-classical pipeline that transforms high-dimensional sentence embeddings into compact representations for variational quantum classification. The workflow combines a pretrained sentence-embedding model, dimensionality reduction, angle encoding, a variational quantum circuit (VQC), and a classical decision layer. We systematically compare principal component analysis (PCA), neighborhood components analysis (NCA), and linear discriminant analysis (LDA), covering both unsupervised and supervised dimensionality reduction. Using the TREC question-classification dataset, we study the relationship between representation dimensionality, information retention, qubit count, and classification performance. Preliminary PCA experiments reveal a strong information bottleneck: reducing 768-dimensional embeddings to 3, 4, 5, and 8 dimensions retains about 8.2%, 10.2%, 11.9%, and 16.4% of the variance, with corresponding classification accuracies of 50.3%, 51.2%, 57.9%, and 63.4%. In contrast, supervised reduction is substantially more efficient. LDA reaches 85.3% accuracy and NCA reaches 83.1% using only 5 dimensions, under a leakage-free cross-validation protocol, compared with 85.1% for a full 384-dimensional classical baseline. These results indicate that supervised dimensionality reduction can preserve task-relevant information far more effectively than variance-based compression, making compact representations a promising route toward practical hybrid quantum-classical NLP models.
Chinese Translation
大型语言模型和句子嵌入模型能够提供丰富的语义表示,但其高维度对近期的量子机器学习(QML)构成了挑战,因为量子电路只能处理有限数量的输入特征。我们研究了一种量子-经典混合流程,将高维句子嵌入转换为紧凑表示,用于变分量子分类。该流程结合了预训练句子嵌入模型、降维、角度编码、变分量子电路(VQC)以及经典决策层。我们系统地比较了主成分分析(PCA)、邻域成分分析(NCA)和线性判别分析(LDA),涵盖了无监督和有监督降维方法。基于TREC问题分类数据集,我们研究了表示维度、信息保留量、量子比特数与分类性能之间的关系。初步的PCA实验揭示了严重的信息瓶颈:将768维嵌入降至3、4、5和8维时,仅分别保留了约8.2%、10.2%、11.9%和16.4%的方差,对应的分类准确率分别为50.3%、51.2%、57.9%和63.4%。相比之下,有监督降维的效率显著更高。在无泄漏交叉验证协议下,LDA仅使用5维即达到85.3%的准确率,NCA达到83.1%,而完整384维的经典基线为85.1%。这些结果表明,有监督降维在保留任务相关信息方面远比基于方差的压缩更为有效,使紧凑表示成为实现实用化量子-经典混合自然语言处理模型的一条有前景的路径。
cs.LG / 60 / 2609.10099

A Systematic Evaluation of Molecule Generation Models for De Novo Drug Design: From Benchmarks to Practical Insights

面向全新药物设计的分子生成模型的系统评估:从基准测试到实践洞察
Xu, Xinrui, Wang, Xueer, Luo, Dan, Yuan, Sisi, Lin, Xuan
Abstract
Molecule generation has emerged as a powerful computational tool for de novo drug design, enabling the exploration of chemical space beyond the limits of conventional virtual screening. The field has progressed rapidly, driven by advances in molecular representations, generative architectures, and target-aware modeling strategies. However, existing reviews typically address specific model families or application scenarios in isolation, rather than offering an integrated perspective on how these components collectively form a coherent generation workflow. In this review, we present a comprehensive evaluation of molecule generation models for de novo drug design, covering 82 methods across five deep generative frameworks, including recurrent neural network (RNN)- and Transformer-based models, variational autoencoders (VAEs), generative adversarial networks (GANs), flow-based models, and diffusion models. We first summarize widely used benchmarks and molecular representations, and then examine the methodological principles underlying both general and pocket-conditioned generation. A central contribution of this work is a systematic synthesis and comparative analysis of reported performance across commonly used benchmarks and evaluation metrics. We also summarize representative experimentally validated case studies. Looking ahead, we discuss future directions in standardized 3D data, interaction-aware generation, receptor flexibility, and multi-objective molecular design, with the aim of improving the reliability and experimental relevance of molecule generation. All collected benchmark resources, evaluation metrics, and model references are provided in a publicly accessible repository at https://github.com/JacklinGroup/molecule-generation-review.
Chinese Translation
分子生成已成为全新药物设计的一种强大计算工具,能够探索超越传统虚拟筛选极限的化学空间。在分子表示、生成架构和靶标感知建模策略等进展的推动下,该领域发展迅速。然而,现有综述通常孤立地讨论特定模型家族或应用场景,而非提供一个关于这些组件如何共同构成完整生成工作流程的整体视角。在本综述中,我们对用于全新药物设计的分子生成模型进行了全面评估,涵盖五种深度生成框架下的82种方法,包括基于循环神经网络(RNN)和Transformer的模型、变分自编码器(VAE)、生成对抗网络(GAN)、基于流的模型以及扩散模型。我们首先总结了广泛使用的基准测试和分子表示,然后考察了通用生成和基于口袋条件生成背后的方法学原理。本工作的一个核心贡献是对常用基准测试和评估指标下已报道性能的系统性综合与比较分析。我们还总结了代表性的实验验证案例研究。展望未来,我们讨论了标准化三维数据、相互作用感知生成、受体柔性和多目标分子设计等未来方向,旨在提高分子生成的可靠性和实验相关性。所有收集的基准测试资源、评估指标和模型参考文献均已在公开可访问的代码仓库中提供:https://github.com/JacklinGroup/molecule-generation-review。
cs.LG / 61 / 2609.10108

A Trust-Network-Based Federated Learning Framework for Multi-Center Aging Clock Prediction

一种基于信任网络的多中心衰老时钟预测联邦学习框架
Zhang, Chunxu, Li, Bo, Wang, Wenliang, Liu, Yang, Jiang, Di, Huang, Yuan, Nabeshima, Yo-ichi, Yamamura, Akinori, Yang, Bo, Yang, Qiang
Abstract
Aging clocks quantify biological aging and help characterize individual health status. What protein interactions are important for accurate aging clocks, and are they zeroth-order or higher-order? Addressing these questions requires learning from large molecular datasets distributed across medical centers, where privacy constraints prevent centralized data sharing. Federated learning offers a natural solution but faces four challenges in this setting: limited local sample sizes, sparse and directional inter-center trust, the need to retain discriminative age prediction while supporting interpretation, and model drift and forgetting under heterogeneous cross-center data. We propose TNFL, a trust-network-based federated learning framework that progressively propagates models along directed pairwise trust relations without centralized aggregation. TNFL combines an age-aware mixture-of-experts model with generative replay to preserve previously learned information and reduce forgetting and drift. Experiments across multiple molecular datasets show that TNFL enables effective aging-clock prediction with limited local data, provides interpretable age-dependent prediction patterns, and maintains stable performance across interaction orders. To investigate the biological questions, we analyze TNFL-identified pairwise protein interactions and their higher-order organization through functional and network analyses. The identified interactions repeatedly form coordinated higher-order subnetworks spanning multiple aging-related biological systems, with several proteins recurring across subnetworks. These findings suggest that TNFL captures molecular relationships beyond isolated pairwise associations and reveals coherent higher-order biological organization associated with aging.
Chinese Translation
衰老时钟可量化生物衰老并有助于刻画个体健康状况。哪些蛋白质相互作用对构建准确的衰老时钟至关重要,它们是零阶的还是高阶的?回答这些问题需要从分布于各医学中心的大型分子数据集中学习,而隐私限制使得集中式数据共享难以实现。联邦学习提供了一种天然的解决方案,但在该场景下面临四个挑战:本地样本量有限、中心间信任关系稀疏且具有方向性、既需保留判别性的年龄预测能力又要支持可解释性,以及在跨中心异构数据下的模型漂移与遗忘问题。我们提出了 TNFL(Trust-Network-Based Federated Learning),一种基于信任网络的联邦学习框架,该框架沿有向的成对信任关系逐步传播模型,而无需集中式聚合。TNFL 将年龄感知的混合专家模型与生成式回放相结合,以保留先前学到的信息并减少遗忘与漂移。在多个分子数据集上的实验表明,TNFL 能够在本地数据有限的情况下实现有效的衰老时钟预测,提供可解释的年龄相关预测模式,并在不同相互作用阶数下保持稳定性能。为探究其中的生物学问题,我们通过功能分析和网络分析研究了 TNFL 识别出的成对蛋白质相互作用及其高阶组织结构。所识别的相互作用反复形成跨越多个衰老相关生物系统的协同高阶子网络,且若干蛋白质在多个子网络中反复出现。这些发现表明,TNFL 能够捕捉超越孤立成对关联的分子关系,并揭示与衰老相关的连贯的高阶生物学组织结构。
cs.LG / 62 / 2609.10112

Storage-Scalable Progressive Semantic Communication via Knowledge-Base Reuse

基于知识库复用的存储可扩展渐进式语义通信
Zhu, Heng, Liu, Ye, Zhu, Kun, Song, Feifei
Abstract
Existing knowledge-base-assisted semantic communication schemes commonly adopt either single knowledge-base quantization (SKBQ) or multi-knowledge-base residual quantization (MKBQ). SKBQ incurs limited storage overhead but has restricted quantization capacity, whereas MKBQ supports progressive refinement by assigning an independent knowledge base (KB) to each stage, causing the KB storage to grow linearly with the transmission depth. To address this problem, we propose storage-scalable knowledge-base reuse quantization (SSKBQ), which reuses a compact set of KBs across multiple residual refinement stages and thereby decouples the number of transmission stages from the number of maintained KBs. A stage-aware residual supervision mechanism is further introduced to regularize intermediate quantized representations and encourage progressive refinement. Experimental results demonstrate that KB reuse provides an effective solution to the storage scalability problem while maintaining competitive progressive reconstruction performance.
Chinese Translation
现有的知识库辅助语义通信方案通常采用单知识库量化(SKBQ)或多知识库残差量化(MKBQ)。SKBQ的存储开销有限,但量化能力受限;而MKBQ通过为每个阶段分配独立的知识库(KB)来支持渐进式细化,但会导致知识库存储随传输深度线性增长。为解决这一问题,我们提出了存储可扩展的知识库复用量化(SSKBQ),该方案在多个残差细化阶段中复用一个紧凑的知识库集合,从而使传输阶段的数量与所维护知识库的数量解耦。此外,我们进一步引入了一种阶段感知的残差监督机制,用于规范中间量化表示并促进渐进式细化。实验结果表明,知识库复用为存储可扩展性问题提供了有效的解决方案,同时保持了具有竞争力的渐进式重建性能。
cs.LG / 63 / 2609.10154

CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts

CompassOPD:基于家族内似然变化的跨家族在策略蒸馏
Gu, Naibin, Si, Qingyi, Yang, Chenxu, Qin, Chuanyu, Zhou, Junhao, Fu, Peng, Lin, Zheng, Wang, Weiping
Abstract
On-policy distillation (OPD) provides dense token-level supervision on student-generated trajectories. Although OPD performs strongly when teacher and student belong to the same model family, we find that its effectiveness degrades in cross-family settings even after tokenizer alignment, with substantially stronger external teachers offering little additional improvement. To understand this disconnect, we decompose the cross-family OPD signal into two components: an offset between a low-capability teacher-family reference and the student, and the within-family log-likelihood shift from that reference to the strong teacher. Standard OPD transfers both components together, allowing the offset to dominate the update direction and obscure the changes associated with teacher capability improvements. We propose CompassOPD, which removes this offset and transfers the within-family shift, while a frozen student reference anchors updates to the student's initial policy. Thus, both teacher-side and student-side changes are measured within their respective model families. Experiments across three student families and multiple teacher families show that CompassOPD consistently outperforms standard cross-family OPD, improving average reasoning accuracy by up to 5.50 points. For an MoE teacher, we further construct the reference directly from the teacher checkpoint by reducing expert activation, eliminating the need for a separate reference checkpoint while retaining a 3.43-point gain over OPD.
Chinese Translation
在策略蒸馏(On-policy distillation,OPD)通过在学生模型生成的轨迹上提供密集的token级监督信号来指导训练。尽管当教师模型和学生模型属于同一模型家族时OPD表现优异,我们发现即使在完成分词器对齐之后,其有效性在跨家族设置下仍会显著下降,即使外部教师模型能力更强,也几乎无法带来额外提升。为理解这一脱节现象,我们将跨家族OPD信号分解为两个组成部分:一是低能力教师家族参考模型与学生模型之间的偏移量,二是从该参考模型到强教师模型的家族内对数似然变化。标准OPD将两个分量一并传递,使得偏移量主导更新方向,从而掩盖了与教师能力提升相关的变化。我们提出CompassOPD,该方法移除该偏移量并传递家族内似然变化,同时使用一个冻结的学生参考模型将更新锚定在学生的初始策略上。由此,教师侧和学生侧的变化均在各自的模型家族内进行度量。在三个学生家族和多个教师家族上的实验表明,CompassOPD始终优于标准的跨家族OPD,平均推理准确率最高提升5.50个百分点。对于MoE教师模型,我们进一步通过降低专家激活直接从教师检查点构建参考模型,从而无需单独的参考检查点,同时相比OPD仍保持3.43个百分点的提升。
cs.LG / 64 / 2609.10158

CoGe-GCD: Reframing Generalized Category Discovery with Compositional Generalization

CoGe-GCD:基于组合泛化的广义类别发现重构
Tang, Luyao, Zheng, Jiewei, Huang, Kunze, Chen, Chaoqi, Huang, Yue, Chen, Cheng
Abstract
Generalized Category Discovery (GCD) assigns unlabeled instances, mixed with labeled data, to known or novel categories, requiring human-like compositional reasoning: reusing primitives learned from known classes and deciding when new combinations imply new categories. Existing GCD methods operate on unstructured token features and struggle to extrapolate to novel compositions. We propose CoGe-GCD, which rethinks GCD through compositional generalization with two coupled stages. (i) Compositional Perception structures patch tokens by mapping them to a small vocabulary of primitives and refining token embeddings via competitive token-primitive assignment and information passing, yielding coherent groups for discovery. (ii) Generalizing Induction exploits the induced geometric structure and applies a structure-preserving calibration over spatial relations, maintaining probabilistic semantics while improving extrapolation to unseen primitive combinations. CoGe-GCD is implemented as an inductive-bias module between backbone and projection head, without modifying heads or losses, and can be plugged into diverse GCD frameworks. On standard benchmarks, it consistently improves all-class accuracy, unknown-class number estimation, and geometric quality, with marginal computational overhead. Code is available at https://github.com/lytang63/CoGe-GCD.
Chinese Translation
广义类别发现(Generalized Category Discovery, GCD)旨在将混杂在已标注数据中的未标注样本划分到已知类别或新类别,这要求模型具备类人的组合推理能力:复用从已知类别中学到的基元,并判断新的组合何时意味着新的类别。现有的GCD方法作用于非结构化的词元特征,难以外推到新颖的组合。我们提出CoGe-GCD,从组合泛化的角度重新思考GCD,包含两个相互耦合的阶段:(i)组合感知(Compositional Perception)通过将图像块词元映射到一个较小的基元词表,并借助竞争性的词元-基元分配和信息传递来细化词元嵌入,从而构建结构化的词元表示,形成连贯的分组以供类别发现;(ii)泛化归纳(Generalizing Induction)利用归纳得到的几何结构,对空间关系施加保持结构的校准,在维持概率语义的同时提升对未见基元组合的外推能力。CoGe-GCD被实现为主干网络与投影头之间的一个归纳偏置模块,无需修改头部或损失函数,即可插入多种GCD框架。在标准基准上,它能持续提升全类别准确率、未知类别数估计以及几何质量,且计算开销极小。代码见 https://github.com/lytang63/CoGe-GCD。
cs.LG / 65 / 2609.10196

An Exponential Deterministic--Randomized Gap in ERM-Oracle Complexity for Thresholds on an Unknown Order

未知序上阈值问题的ERM-预言机复杂度中的指数级确定性—随机化差距
Li, Xuan
Abstract
Attias, Hanneke and Ramaswami (NeurIPS 2025) asked whether randomization provably reduces the oracle calls needed for online learning when the class is accessible only through an oracle. We study the instance they singled out: transductive online learning of thresholds on an unknown total order of T instances, with a consistency-type ERM oracle that returns a full concept consistent with a queried labeled set (or reports non-realizability). Our main result is a separation for a fixed natural oracle. When the oracle is the minimal-prefix rule (or the maximal-prefix rule), every deterministic learner makes M mistakes and Q calls with $M+Q\ge T-\varepsilon$ on some instance ($\varepsilon\in\{0,1\}$, according to whether the empty prefix is a concept), and the constant is exact; hence $O(\log T)$ mistakes cost $T-\varepsilon-O(\log T)$ calls, whereas that paper's randomized learner achieves $O(\log T)$ expected calls and mistakes under the same rule. The randomized order is optimal: on an explicit hard distribution under the minimal-prefix rule, every learner has expected mistakes at least $((T+1-\varepsilon)\,128^{-\mathbb{E}[Q]}-1)/2$, so $\Omega(\log T)$ expected calls are necessary for polylogarithmic mistakes. The separation is governed by the oracle's selection rule, not by the class alone: for a legal feasible-median ERM rule a deterministic learner achieves $O(\log T)$ calls and mistakes, while a global-median rule again forces linear total cost. The same linear bound holds when the oracle's answers are chosen adversarially and then frozen into a memoryless oracle. We add partial tradeoff results for fixed query budgets (the middle regime is open) and an interface contrast: with only a weak consistency oracle, returning a realizability bit, both deterministic and randomized learners need $\Theta(T)$ calls.
Chinese Translation
Attias、Hanneke 和 Ramaswami(NeurIPS 2025)提出了这样一个问题:当概念类仅能通过预言机访问时,随机化是否能够被证明可以减少在线学习所需的预言机调用次数。我们研究了他们特别指出的实例:在 T 个实例的未知全序上的阈值的传递式在线学习,所用的是一个一致性类型的 ERM 预言机,该预言机返回一个与被查询的带标签集合相一致的完整概念(或报告不可实现性)。我们的主要结果是针对一个固定的自然预言机的分离结果。当预言机采用最小前缀规则(或最大前缀规则)时,任何确定性学习器在某些实例上都会犯 M 次错误并进行 Q 次调用,且满足 $M+Q\ge T-\varepsilon$($\varepsilon\in\{0,1\}$,取决于空前缀是否为一个概念),且该常数是精确的;因此,$O(\log T)$ 的错误次数需要付出 $T-\varepsilon-O(\log T)$ 次调用的代价,而该论文中的随机化学习器在同一规则下可以实现 $O(\log T)$ 的期望调用次数和错误次数。该随机化阶数是最优的:在最小前缀规则下的一个显式困难分布上,任何学习器的期望错误次数至少为 $((T+1-\varepsilon)\,128^{-\mathbb{E}[Q]}-1)/2$,因此要达到多对数级的错误次数,$\Omega(\log T)$ 的期望调用次数是必需的。这一分离由预言机的选择规则决定,而非仅由概念类决定:对于一个合法的可行中位数 ERM 规则,确定性学习器可以实现 $O(\log T)$ 的调用次数和错误次数,而全局中位数规则则再次迫使总代价为线性。当预言机的答案由对手选择后被固定为一个无记忆预言机时,同样的线性界仍然成立。我们还添加了针对固定查询预算的部分权衡结果(中间区域仍是开放问题),以及一个接口对比:若只有返回可实现性比特的弱一致性预言机,则确定性学习器和随机化学习器都需要 $\Theta(T)$ 次调用。
cs.LG / 66 / 2609.10200

Robust Beam Prediction for V2X Networks with Multi-Modal Sensing

基于多模态感知的V2X网络鲁棒波束预测
Shang, Chen, Hoang, Dinh Thai, Nguyen, Diep N., Yu, Jiadong
Abstract
Integrated sensing and communication (ISAC) provides a promising foundation for beam prediction in future vehicle-to-everything (V2X) networks. However, existing sensing-assisted beamforming methods still rely heavily on radio-frequency sensing, which may become unreliable in complex vehicular environments. Meanwhile, the growing availability of heterogeneous sensors, such as cameras and LiDAR, offers new opportunities to improve beam prediction through richer environmental perception. Motivated by this, this paper proposes a multi-modal beam prediction framework for V2X networks. Specifically, we develop BeamTransFuser, a hierarchical Transformer-based architecture that progressively fuses camera, LiDAR, radar, and GPS observations for robust beam prediction. In addition, to handle possible missing modalities in practical deployment, we introduce a generative module that reconstructs missing modality features from the available observations. Experimental results on a real-world multi-modal V2X dataset show that the proposed framework consistently outperforms representative baselines, while the generative module further improves robustness under incomplete sensing conditions.
Chinese Translation
通信感知一体化(ISAC)为未来车联网(V2X)网络中的波束预测提供了有前景的基础。然而,现有的感知辅助波束赋形方法仍在很大程度上依赖射频感知,而射频感知在复杂车辆环境中可能变得不可靠。与此同时,摄像头和激光雷达等异构传感器的日益普及,为通过更丰富的环境感知来改进波束预测提供了新的机遇。受此启发,本文提出了一种面向V2X网络的多模态波束预测框架。具体而言,我们开发了BeamTransFuser,这是一种基于分层Transformer的架构,能够逐步融合摄像头、激光雷达、雷达和GPS观测数据以实现鲁棒的波束预测。此外,为应对实际部署中可能出现的模态缺失问题,我们引入了一个生成模块,可从可用的观测数据中重建缺失模态的特征。在真实世界多模态V2X数据集上的实验结果表明,所提出的框架始终优于具有代表性的基线方法,且生成模块在感知不完整的条件下进一步提升了鲁棒性。
cs.LG / 67 / 2609.10225

Hierarchical and Permutation-Invariant Feature Transformation Learning via Policy-Guided Embedding Search

基于策略引导嵌入搜索的层次化与置换不变特征变换学习
Liu, Rui, Zhe, Tao, Huang, Yanyong, Guria, Sankha Narayan, Luo, Xiao, Fan, Wei, Fu, Yanjie, Wang, Dongjie
Abstract
Feature transformation improves predictive performance on tabular data by constructing informative abstractions from raw features. Recent generative approaches encode transformation knowledge into continuous embedding spaces for efficient exploration of candidate strategies, but face three key limitations: (1) overlooking hierarchical relationships between low-level features, operations, and high-level abstractions; (2) enforcing order-sensitive embeddings on inherently permutation-invariant transformation sequences, thereby introducing systematic bias; and (3) relying on gradient-based search, which is ill-suited to non-convex transformation spaces. We propose a framework with two complementary components. First, a permutation-invariant hierarchical module captures interactions across features, operations, and abstraction levels, with a self-attention pooling mechanism that maps semantically equivalent structures to consistent embeddings aligned with downstream performance. Second, a policy-guided multi-objective reinforcement learning strategy initializes the search from empirically strong seeds and jointly optimizes predictive accuracy and transformation efficiency. Extensive experiments on diverse tabular benchmarks demonstrate the effectiveness and robustness of our framework against strong baselines. Our code and data are publicly available at: https://github.com/RayLiu1103/PHER.
Chinese Translation
特征变换通过从原始特征中构建信息丰富的抽象表示,提升表格数据的预测性能。近期的生成式方法将变换知识编码到连续嵌入空间中,以高效探索候选策略,但面临三个关键局限:(1)忽视了低层特征、操作与高层抽象之间的层次关系;(2)在本质上置换不变的变换序列上强行使用顺序敏感的嵌入,从而引入系统性偏差;(3)依赖基于梯度的搜索,而这种搜索方式并不适合非凸的变换空间。我们提出了一个包含两个互补组件的框架。首先,一个置换不变的层次化模块捕捉特征、操作与抽象层次之间的交互,并通过自注意力池化机制将语义等价的结构映射到与下游性能一致的一致性嵌入。其次,一种策略引导的多目标强化学习策略从经验上较强的种子出发初始化搜索,并联合优化预测精度与变换效率。在多样化表格基准数据集上的大量实验表明,我们的框架相对于强基线方法具有有效性和鲁棒性。我们的代码和数据已在以下网址公开:https://github.com/RayLiu1103/PHER。
cs.LG / 68 / 2609.10287

Training Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers

训练轨迹决定可退火软先验Transformer中回路的可移除性
Yang, Zonglin, Zhao, Ziming, Tang, Wei, Jiang, Xunyu, Liu, Yihong, Chen, Tailin, Yu, Zifu, Liu, Jiayu
Abstract
Soft positional priors can help small Transformers learn retrieval circuits, but it is unclear whether the resulting circuits remain functional once the prior is removed. We test this with an annealable soft-prior Transformer whose attention biases can be learned, faded, or zeroed during training and evaluation. On associative recall, unforced models perform well with the prior active ($0.772 \pm 0.020$) but collapse at zero gate ($0.095 \pm 0.009$). Smooth fade-to-zero training preserves high zero-gate accuracy ($0.734 \pm 0.028$), whereas forced-zero training, hard switching, and post hoc continuation fail to recover the same effect. The pattern also appears on Markov induction. Linear regression ICL provides a boundary case because zero-gate training can learn that task directly. Mechanistic traces show that circuit consolidation occurs after the gate reaches zero, even though the responsible heads vary across seeds. These results suggest that circuit removability in small discrete retrieval tasks depends on the training trajectory, not just the final architecture.
Chinese Translation
软位置先验可以帮助小型Transformer学习检索回路,但尚不清楚在移除先验后,所形成的回路是否仍能保持功能。我们通过一个可退火的软先验Transformer(annealable soft-prior Transformer)来检验这一点,其注意力偏置在训练和评估过程中可以被学习、衰减或置零。在关联召回任务上,未经强制的模型在先验激活时表现良好($0.772 \pm 0.020$),但在门控置零时性能崩溃($0.095 \pm 0.009$)。平滑衰减至零的训练能够保持较高的零门控准确率($0.734 \pm 0.028$),而强制置零训练、硬切换以及事后继续训练均未能恢复同样的效果。该模式同样出现在马尔可夫归纳任务中。线性回归的上下文学习(ICL)提供了一个边界情形,因为零门控训练可以直接学会该任务。机制层面的追踪显示,回路巩固发生在门控达到零之后,即使负责该功能的注意力头在不同随机种子下各不相同。这些结果表明,在小型离散检索任务中,回路的可移除性取决于训练轨迹,而不仅仅是最终的架构。
cs.LG / 69 / 2609.10299

A Dominant Diffuse Phase in the Sparse Autoencoder Phase Diagram

稀疏自编码器相图中的主导弥散相
Plascencia, Alexis D.
Abstract
Sparse autoencoders (SAEs) are increasingly used to recover interpretable features from neural-network activations, yet systematic feature co-occurrence can cause distinct features to be absorbed or merged. The MAIS-O43 open problem proposes a controlled experiment to characterize when recovery of a true synthetic dictionary gives way to feature merging as the nesting fraction $\gamma$, sparsity penalty $\lambda$, and dictionary size $M$ vary. We implement the specified protocol and evaluate 200 independently initialized fits across ten of the 165 grid cells. We observe zero full-dictionary recoveries and zero merges. Instead, every run converges to a reproducible diffuse phase: reconstruction is nearly perfect, but learned atoms typically remain far from the true features (median best cosine 0.5-0.7 against a 0.95 recovery criterion) and learned codes are an order of magnitude denser than the ground truth. This behavior persists under robustness checks and across the full 165-cell grid using standard minibatch Adam (3,300 additional fits). Since the global optimum of the exact sparse-coding objective is known to merge nested features in the two-feature case, these results suggest that trained SAEs need not reach the corresponding minima, and that the phase diagram of trained models may differ fundamentally from that of objective minimizers.
Chinese Translation
稀疏自编码器(Sparse Autoencoders, SAEs)被越来越多地用于从神经网络激活中恢复可解释特征,然而系统性的特征共现会导致不同特征被吸收或合并。MAIS-O43开放性问题提出了一个受控实验,用于刻画随着嵌套比例$\gamma$、稀疏惩罚$\lambda$和字典大小$M$的变化,真实合成字典的恢复何时让位于特征合并。我们实现了该指定协议,并在165个网格单元中的十个单元上评估了200次独立初始化的拟合。我们观察到零次完整字典恢复和零次特征合并。相反,每次运行都收敛到一个可复现的弥散相:重构几乎完美,但学习到的原子通常远离真实特征(最佳余弦相似度中位数为0.5-0.7,而恢复标准为0.95),且学习到的编码比真实编码稠密一个数量级。该行为在稳健性检验以及使用标准小批量Adam优化器的全部165个网格单元(额外3,300次拟合)中均持续存在。鉴于在双特征情形下,精确稀疏编码目标的全局最优解已知会合并嵌套特征,这些结果表明训练得到的SAE未必到达相应的极小值点,且训练后模型的相图可能与目标函数极小化器的相图存在根本差异。
cs.LG / 70 / 2609.10307

View-Structured Conformal Prediction for 3D Gaussian Splatting

面向三维高斯泼溅(3D Gaussian Splatting)的视角结构化保形预测
Chu, Junzheng, Pan, Bin, Shi, Zhenwei
Abstract
3D Gaussian Splatting (3DGS) renders novel views in real time, but an uncertainty heatmap does not certify that a rendered view meets a certain prediction coverage. We treat novel-view synthesis as structured regression and ask that, with probability at least $1-\alpha$, RGB prediction boxes cover at least a $1-\beta$ fraction of pixels in a new view. We propose View-Structured Conformal Prediction (VSCP). It splits the pre-calibration scale into a spatial shape from the renderer and a transferable view-difficulty factor, which predicts the smallest view-wise multiplier that shape needs. A held-out quantile over views (View-CP) then gives finite-sample validity even when transferring to new scenes. The same factorization makes the analysis exact: a conformity score is the ratio of oracle to predicted view difficulty, and excess width separates into a test-side and a calibration-side term. Across 13 real scenes, pixel-pooled calibration reaches 89.9\% marginal pixel coverage but only 61.4\% view-event coverage at a 90\% target, while View-CP reaches 91.7--92.0\%. At matched coverage VSCP cuts width by 22.1\% against a constant scale, and matches a ten-model ensemble's 21.0\% reduction using only one model per scene and four rather than ten rasterization passes per query. VSCP also improves on the closest single-model baseline, the 3DGS-U field, by 4.7 points ($p=0.0225$). The view predictor transfers from bounded source families to all nine unbounded Mip-NeRF~360 scenes. There the full scale beats the constant scale with 20.7\% width saving on all nine scenes. It also keeps an 18.3\% saving under a different densification backbone and runs at 216--280 FPS on an RTX~4090.
Chinese Translation
三维高斯泼溅(3D Gaussian Splatting, 3DGS)可以实时渲染新视角,但不确定性热图并不能证明渲染出的视角满足一定的预测覆盖率。我们将新视角合成视为结构化回归问题,并要求 RGB 预测框以至少 $1-\alpha$ 的概率覆盖新视角中至少 $1-\beta$ 比例的像素。我们提出了视角结构化保形预测(View-Structured Conformal Prediction, VSCP)。该方法将预校准尺度分解为来自渲染器的空间形状和一个可迁移的视角难度因子,后者预测该形状所需的最小逐视角乘数。随后,在视角上的留出分位数(View-CP)即使在迁移到新场景时也能保证有限样本有效性。这种分解使分析变得精确:一致性分数为真实视角难度与预测视角难度之比,而过多的宽度则可分离为测试侧与校准侧两项。在 13 个真实场景上,像素池化校准在 90% 目标下达到 89.9% 的边际像素覆盖率,但视角事件覆盖率仅为 61.4%,而 View-CP 达到了 91.7%–92.0%。在相同覆盖率下,与常数尺度相比,VSCP 将宽度缩减了 22.1%,并且仅使用每个场景一个模型、每次查询四次(而非十次)光栅化过程,即可匹配十模型集成 21.0% 的宽度缩减效果。VSCP 还优于最接近的单模型基线——3DGS-U 场,提升了 4.7 个百分点($p=0.0225$)。视角预测器能够从有界的源场景族迁移到全部九个无界的 Mip-NeRF 360 场景。在所有九个场景上,完整尺度相比常数尺度节省了 20.7% 的宽度。在不同致密化骨干网络下它仍保持 18.3% 的节省,并在 RTX 4090 上以 216–280 FPS 的速度运行。
cs.LG / 71 / 2609.10311

One Loop, Two Gains: Can Active Learning win the Lottery for Free?

一次循环,双重收益:主动学习能否免费赢得“彩票”?
Tscheschner, Benedikt, Veas, Eduardo, Masana, Marc
Abstract
The lottery ticket hypothesis posits the existence of winning tickets: sparse subnetworks that, when trained in isolation from their original initialization, match the accuracy of the full dense network. The predominant method for discovering such tickets, iterative magnitude pruning, alternates pruning with full retraining from scratch until convergence over many cycles. Similarly, deep active learning also retrains a model from scratch after each acquisition round as new labels become available. Despite this shared reliance on iterative retraining with a substantial computational overhead, the two paradigms have been studied separately. We observe that the iterative training loop inherent to pool-based active learning already provides the exact computational structure that iterative magnitude pruning exploits, and propose Improve & Prune (I&P), a method that integrates magnitude pruning into each active learning retraining cycle at practically no additional cost. This raises a key empirical question: can iterative magnitude pruning produce winning tickets under the non-stationary data regime of active learning? We investigate this question across multiple acquisition functions, architecture families, and image classification datasets, including an active fine-tuning scenario. Our results demonstrate that I&P yields sparse, deployable models at each active learning iteration. Those match the accuracy of their dense counterparts at sparsities up to 95%, effectively obtaining winning tickets as a byproduct of the active learning pipeline. These per-iteration sparse models can address two computational bottlenecks - per-round model retraining and acquisition scoring over the unlabeled pool - that currently prevent the practical adoption of DAL on large architectures and large unlabeled pools.
Chinese Translation
彩票假说(lottery ticket hypothesis)认为存在“中奖彩票”:即稀疏子网络,在从其原始初始化独立训练时,能够达到完整稠密网络的准确率。发现此类彩票的主流方法是迭代幅度剪枝(iterative magnitude pruning),它在多个周期中交替进行剪枝与从头完整重训练直至收敛。类似地,深度主动学习(deep active learning)在每轮获取新标注数据后也会从头重新训练模型。尽管这两种范式都依赖迭代重训练并带来巨大的计算开销,但此前一直被分开研究。我们观察到,基于样本池的主动学习(pool-based active learning)所固有的迭代训练循环,恰好提供了迭代幅度剪枝所需的计算结构,并据此提出 Improve & Prune(I&P)方法,将幅度剪枝以几乎零额外成本的方式集成到主动学习的每次重训练循环中。这引出了一个关键的实证问题:在主动学习的非平稳数据情形下,迭代幅度剪枝能否产生中奖彩票?我们在多种获取函数、多种架构家族以及多个图像分类数据集(包括主动微调场景)上对此问题进行了研究。结果表明,I&P 在主动学习的每次迭代中都能得到稀疏且可部署的模型,这些模型在稀疏度高达 95% 时仍能达到与其稠密对应模型相当的准确率, effectively 相当于作为主动学习流程的副产品获得了中奖彩票。这些逐迭代的稀疏模型能够解决两个计算瓶颈——每轮模型重训练以及对未标注样本池的获取打分——正是这些瓶颈目前阻碍了深度主动学习(DAL)在大规模架构和大规模未标注样本池上的实际应用。
cs.LG / 72 / 2609.10357

A Later Test Set Is Not a New Domain: Pretraining Familiarity Survives a Contamination-Free Hold-Out

更晚的测试集并非新领域:预训练熟悉度在无污染的保留集上依然存在
Moghadasi, Mahdi Naser, Ghaderi, Faezeh
Abstract
Time-series foundation models are evaluated almost exclusively on public archives that predate them, so a strong score cannot be separated from having seen the test set during pretraining. The obvious remedy is a hold-out that postdates the models. We build one: thirteen forecasters -- four classical, three trained per dataset, six pretrained -- on seven groups drawn from five domains, every observation published after the last model was released, and every dataset rebuildable without an API key. Under this protocol pretrained models win 5 of 7 groups, lose one to a Theta baseline, and on daily exchange rates are indistinguishable from a seasonal naive forecast, along with every other method tested. We then ask what separates the wins from the losses, and report a negative result: the two intrinsic properties one would reach for -- seasonal strength and spectral entropy, measured on the input window -- do not account for the pattern, and seasonal strength is if anything negatively associated with the advantage. What does track it is corpus familiarity. Our largest gain (28% lower MASE than the best classical method, on weekly Wikipedia pageviews) falls on Wikipedia pageviews, the domain TimesFM's authors describe as the bulk of its pretraining corpus, at the same granularities and differing only in time window. Within the pretrained family, where every model forecasts identical series so that series difficulty cancels, the TimesFM family outranks the Chronos family by -0.53 ranks on Wikipedia against -0.09 everywhere else (1,500 vs. 754 series, Mann-Whitney p < 1e-5). We conclude that a temporal hold-out removes memorisation of a window but not familiarity with a domain, that benchmarks therefore need domain hold-outs stated relative to disclosed corpora, and that the practitioner's question is less which model is better than whether their domain is one the model was raised on.
Chinese Translation
时间序列基础模型几乎完全在早于其发布的公开数据集上进行评估,因此高分无法与其在预训练中见过测试集这一事实相区分。显而易见的补救措施是使用模型发布之后的保留集。我们构建了这样一个测试集:十三个预测模型——四个经典方法、三个按数据集训练的模型、六个预训练模型——在来自五个领域的七个数据组上进行评估,所有观测数据均发布于最后一个模型发布之后,且每个数据集均可无需API密钥重建。在该协议下,预训练模型在7个数据组中赢得5个,在一个数据组上输给Theta基线,而在每日汇率数据上,其表现与季节性朴素预测以及所有其他被测试方法均无显著差异。随后我们探究了胜负之间的区分因素,并报告了一个负面结果:人们通常会选择的两类内在属性——在输入窗口上测量的季节强度和谱熵——并不能解释这一模式,而且季节强度与优势之间如果说有关联,反而是负相关的。真正与该模式相关的是语料库熟悉度。我们最大的增益(在每周维基百科页面浏览量上,MASE比最佳经典方法低28%)恰好出现在维基百科页面浏览量这一领域上,而TimesFM的作者将该领域描述为其预训练语料库的主体,且处于相同的粒度,仅在时间窗口上不同。在预训练模型家族内部,由于每个模型预测相同的序列从而序列难度被抵消,TimesFM家族在维基百科数据上相对Chronos家族的平均排名优势为-0.53,而在其他所有数据上仅为-0.09(1,500个序列对比754个序列,Mann-Whitney检验 p < 1e-5)。我们得出结论:时间上的保留集消除了对某个时间窗口的记忆,但无法消除对某个领域的熟悉度;因此,基准测试需要相对于已公开语料库声明的领域保留集;而从业者的核心问题与其说是哪个模型更好,不如说其所在领域是否是该模型'从小接触'的领域。
cs.LG / 73 / 2609.10364

OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis

OmniMed-FL:一种鲁棒的多模态联邦学习临床诊断框架
Debnath, Ayush, Saha, Ruelia, Misra, Sudip
Abstract
Simultaneous assessment of medical imaging and patient records is often required in clinical diagnosis. However, standard machine learning algorithms cannot analyze these data types together. Meanwhile, compliance with HIPAA and GDPR can constrain centralized aggregation of sensitive patient data. This leaves a crucial void of secure fusion of visual and textual context across distant networks. Thus, we present OmniMed-FL, a controlled systems study of multimodal federated learning for five-class clinical condition classification (Normal, Pneumonia, COVID-19, Pleural Effusion, Cardiomegaly). Our proxy corpus pairs 3,000 public chest radiographs with 3,000 class-conditioned synthetic notes, matched by class, not by patient. The framework benchmarks eight fusion strategies, three initializations, four missing-text imputation rules, and matched federated baselines under non-IID Dirichlet partitioning across 3 to 20 hospital clients. As all notes are synthetic and pairing is not patient-level, these are descriptive proxy comparisons, not estimates of diagnostic performance or deployment readiness. Within those limits with clients ($K=5$) and severe skew ($\alpha=0.1$), local-only training achieves a macro-F1 score of 0.297, FedAvg achieves $0.662\pm0.074$, FedProx $0.737\pm0.085$, a matched FedMME-style one-shot ensemble $0.647\pm0.080$, and our SCAFFOLD-AdamW adaptation $0.070\pm0.015$, the 0.075 FedProx-FedAvg gap falling inside the wider of the two two-seed standard deviations. Over a $4\times3$ grid, label skew costs up to 0.27 F1 whereas a near-sevenfold client increase costs at most 0.10, while bidirectional volume grows linearly to 183.5 GiB at $K=20$. Multimodal fusion leads on both corpora, scoring 0.956 against 0.934 for text and 0.664 for images on the synthetic corpus and 0.906 against 0.880 and 0.737 on the radiograph corpus, for $2.3\times$ the model state of text alone.
Chinese Translation
临床诊断通常需要同时评估医学影像和患者病历。然而,标准机器学习算法无法将这两类数据放在一起分析。同时,遵循HIPAA(美国健康保险流通与责任法案)和GDPR(通用数据保护条例)的要求也会限制敏感患者数据的集中式聚合。这在跨远程网络的安全融合视觉与文本信息方面留下了关键空白。为此,我们提出了OmniMed-FL,一项针对五类临床病症分类(正常、肺炎、COVID-19、胸腔积液、心脏肥大)的多模态联邦学习可控系统研究。我们的代理语料库将3,000张公开胸部X光片与3,000份按类别条件生成的合成病历配对,其配对依据是类别而非患者。该框架在3至20个医院客户端的非独立同分布(non-IID)狄利克雷(Dirichlet)划分下,对八种融合策略、三种初始化方法、四种文本缺失插补规则以及匹配的联邦学习基线进行了基准测试。由于所有病历均为合成数据,且配对并非基于患者层面,因此这些仅为描述性的代理比较,而非对诊断性能或部署成熟度的估计。在上述限制条件下,当客户端数(K=5)且数据严重偏斜(α=0.1)时:仅本地训练的宏平均F1分数为0.297,FedAvg为0.662±0.074,FedProx为0.737±0.085,匹配的FedMME风格单次集成(one-shot ensemble)为0.647±0.080,而我们的SCAFFOLD-AdamW改编版本为0.070±0.015;FedProx与FedAvg之间0.075的差距落在两者双种子标准差中较大者之内。在一个4×3的网格实验中,标签偏斜最多造成0.27的F1损失,而客户端数量增加近七倍最多只造成0.10的损失,同时双向通信量随客户端数线性增长,在K=20时达到183.5 GiB。多模态融合在两个语料库上均表现最佳:在合成语料库上得分为0.956,相比之下仅文本为0.934、仅图像为0.664;在X光片语料库上得分为0.906,相比之下仅文本为0.880、仅图像为0.737,而其模型状态大小仅为纯文本模型的2.3倍。
cs.LG / 74 / 2609.10439

Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

只遗忘重要的内容:面向鲁棒大语言模型的层选择性机器遗忘
Ranjan, Ravi, Kotevska, Olivera, Polyzou, Agoritsa
Abstract
Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable training content, creating privacy, safety, and regulatory concerns. Machine unlearning offers a practical alternative to full retraining, but many existing methods apply broad or fixed parameter updates that can degrade utility and remain brittle under deployment changes such as post-training quantization, where forgotten knowledge may partially re-emerge. We propose Forgetting Only What Matters via Unlearning Layers (FOM-UL), a layer-level unlearning framework that selects transformer layers using a forget-to-retain significance score. This score identifies layers with high influence on the forget set and low sensitivity to the retain set, allowing FOM-UL to concentrate updates where they are most effective while leaving most of the model unchanged. This targeted update strategy improves the forgetting-utility trade-off and provides an empirical path toward quantization-resilient unlearning by reducing the chance that small, diffuse updates are erased by low-bit rounding. Across TOFU, KnowUnDo, and MUSE-style evaluations, FOM-UL reduces residual memorization compared with strong GA, NPO, KLD, SURE, ReLearn, and LUNAR-based baselines while preserving retain-set utility close to the vanilla model. Under 8-bit and 4-bit post-training quantization, FOM-UL maintains stronger memorization suppression and utility preservation than competing methods, and adversarial prompt evaluations show lower recovery of forgotten content. Overall, FOM-UL provides an efficient unlearning strategy that improves targeted forgetting, utility preservation, and deployment robustness without claiming formal guarantees of erasure.
Chinese Translation
大语言模型(LLMs)可能会记忆并复现敏感、受版权保护或其他不良的训练内容,从而引发隐私、安全和监管方面的担忧。机器遗忘(machine unlearning)为完全重训练提供了一种实用的替代方案,但许多现有方法采用宽泛或固定的参数更新方式,这可能损害模型效用,并且在部署环境变化(如训练后量化)下表现脆弱,被遗忘的知识可能部分重新出现。我们提出了一种通过遗忘关键层实现选择性遗忘的层级遗忘框架——FOM-UL(Forgetting Only What Matters via Unlearning Layers),该框架使用遗忘-保留显著性分数来选择Transformer层。该分数能够识别出对遗忘集影响较大且对保留集敏感度较低的层,使FOM-UL能够将更新集中在最有效的位置,同时保持模型的大部分参数不变。这种有针对性的更新策略改善了遗忘与效用之间的权衡,并通过降低微小、分散的更新被低比特舍入消除的可能性,为实现抗量化鲁棒遗忘提供了一条实证路径。在TOFU、KnowUnDo以及MUSE风格的评测中,与GA、NPO、KLD、SURE、ReLearn以及基于LUNAR的强基线方法相比,FOM-UL降低了残差记忆程度,同时将保留集的效用维持在接近原始模型的水平。在8比特和4比特训练后量化条件下,FOM-UL在记忆抑制和效用保持方面均优于竞争方法,且对抗性提示评测显示被遗忘内容的恢复率更低。总体而言,FOM-UL提供了一种高效的机器遗忘策略,在不声明形式化删除保证的前提下,提升了定向遗忘能力、效用保持以及部署鲁棒性。
cs.LG / 75 / 2609.10464

Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization

Semigroup-JEPA:面向零样本物理泛化的潜在动力学一致性方法
Liu, Andy Zeyi, Sun, Haoran, Baker, Lucas, Balestriero, Randall, Sous, John
Abstract
Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning, but their capability to learn physics and generate physically realistic dynamics remains hitherto untested. In this work, we introduce SemiGroup-JEPA (SG-JEPA), which extends the LeWorldModel framework by supplying the parameter governing the physics to the temporal model via action-conditioning and jointly training an encoder and predictor through an autoregressive latent rollout. To evaluate the model's ability to generalize out of distribution, we design dynamical tasks under different gravitational fields that, despite obeying the same physical law, exhibit qualitatively different dynamics, ranging from floating motion in weak gravitational fields to rapid bouncing in strong ones. In contrast to DINO-WM, SG-JEPA reduces open-loop prediction error by up to 2 times on two-dimensional datasets, and increases control success rate up to 2.5 times for three-dimensional robotic datasets, for which we train independent diffusion policies. To explain this advantage, we develop a linear feature model that separates local law-conditioned error from its recursive amplification under rollout. Guided by this model, we find that back-propagating the multi-step rollout loss into the representation trains the encoder to keep the features that the predictor can carry forward, and that those are the features the dynamics depend on, so most of the gain comes from the encoder learning better features rather than from the predictor learning better dynamics. See project page at https://sg-jepa.github.io.
Chinese Translation
联合嵌入预测架构(Joint-Embedding Predictive Architecture, JEPA)世界模型通过学习世界的紧凑潜在表示来支持预测与规划,但其学习物理规律并生成物理上合理动力学的能力至今尚未得到检验。在本工作中,我们提出了SemiGroup-JEPA(SG-JEPA),它对LeWorldModel框架进行了扩展:通过动作条件化将控制物理规律的参数提供给时间模型,并通过自回归的潜在空间前向推演(latent rollout)联合训练编码器与预测器。为评估模型的分布外泛化能力,我们设计了不同引力场下的动力学任务——尽管这些任务遵循相同的物理定律,却呈现出质性不同的动力学行为,从弱引力场中的漂浮运动到强引力场中的快速弹跳。与DINO-WM相比,SG-JEPA在二维数据集上将开环预测误差最多降低2倍,在三维机器人数据集(我们为其训练了独立的扩散策略)上将控制成功率最多提升2.5倍。为解释这一优势,我们构建了一个线性特征模型,将局部规律条件下的误差与推演过程中误差的递归放大分离开来。在该模型的指引下,我们发现将多步推演损失反向传播到表示中,能够训练编码器保留预测器可继续传递的特征,而这些正是动力学所依赖的特征;因此大部分增益来自编码器学到了更好的特征,而非预测器学到了更好的动力学。项目页面见 https://sg-jepa.github.io。
cs.LG / 76 / 2609.10487

Nonmaximal sums of maximally monotone operators under Rockafellar's constraint qualification

Rockafellar约束资格条件下极大单调算子的非极大和
Yang, Weifeng
Abstract
We construct counterexamples to Rockafellar's sum conjecture in which two maximally monotone operators satisfy the interior-domain condition but their sum is not maximally monotone. We give one counterexample on $c_0$ and another on $\ell^1$ with its usual norm. We establish a general construction theorem that computes the entire monotone polar of a class of graphs, gives a necessary and sufficient condition for their maximal monotonicity, and shows how a positive rank-one perturbation yields a nonmaximal sum under this condition. We verify the theorem's hypotheses and its maximality criterion on $c_0$, thereby obtaining a counterexample to the conjecture. Furthermore, we construct a bounded linear surjection from $\ell^1$ onto $c_0$ and use it to obtain the counterexample on $\ell^1$.
Chinese Translation
我们构造了Rockafellar和猜想(sum conjecture)的反例:两个极大单调算子满足内点定义域条件,但它们的和不是极大单调的。我们分别在空间 $c_0$ 上以及在赋予通常范数的空间 $\ell^1$ 上给出了一个反例。我们建立了一般性的构造定理,该定理计算了一类图像(graph)的完整单调极(monotone polar),给出了其极大单调性的充要条件,并说明了在该条件下如何通过一个正秩一扰动(positive rank-one perturbation)得到非极大的和。我们在 $c_0$ 上验证了该定理的假设及其极大性判据,从而得到了该猜想的一个反例。此外,我们构造了一个从 $\ell^1$ 到 $c_0$ 上的有界线性满射,并利用它得到了 $\ell^1$ 上的反例。
cs.LG / 77 / 2609.10490

Learning with Covariance Matrices: Principal Component Analysis Meets Learning with Graphs

基于协方差矩阵的学习:主成分分析与图学习的交汇
Sihag, Saurabh, Cavallo, Andrea, Isufi, Elvin, Mateos, Gonzalo, Ribeiro, Alejandro
Abstract
This feature article provides an overview of the theoretical foundations for coVariance neural networks (VNNs), i.e., graph neural networks (GNNs) operating on covariance matrices as graphs. Covariance matrices are ubiquitous across domains, and hence, the deployment of GNNs often leverages graphs of pairwise statistical dependencies. Existing theoretical contributions on GNNs consider abstract graph representations and cannot accommodate the data-driven nuances associated with covariance matrices. This tutorial brings into focus various novel theoretical insights via mathematical analyses of VNNs that have broad signal processing implications, including: (i) a conceptual equivalence between VNNs and principal component analysis (PCA)-based information processing; (ii) refined stability bounds on predictive outcomes in the presence of finite sample-induced covariance matrix perturbations; and (iii) refined characterization of transferability of VNNs across multiscale datasets. The theoretical insights discussed herein provide the underlying principles and justification towards adopting VNNs over workhorse PCA-based learning pipelines, in applications where covariance matrices are useful descriptors of data structure. We also convey how impact of these foundational advances permeates to \textit{principled} designs and applications of learning methods across broad domains where covariance matrices emerge. Notably, we elucidate the conceptual insights facilitated by VNNs to the specific task of characterizing brain age gap for neurodegenerative conditions using neuroimaging datasets, a timely problem in computational neuroscience. Broader impacts to other application domains are discussed as well.
Chinese Translation
本文综述了协方差神经网络(VNN)的理论基础,即在协方差矩阵上作为图进行运算的图神经网络(GNN)。协方差矩阵在各个领域中无处不在,因此GNN的部署通常利用成对统计依赖关系构成的图。现有的GNN理论贡献考虑的是抽象的图表示,无法涵盖与协方差矩阵相关的数据驱动特性。本教程通过对VNN的数学分析,聚焦于多方面的新颖理论见解,这些见解具有广泛的信号处理意义,包括:(i)VNN与基于主成分分析(PCA)的信息处理之间的概念等价性;(ii)在有限样本引起的协方差矩阵扰动存在的情况下,对预测结果稳定性界的精细化;(iii)对VNN在多尺度数据集间可迁移性的精细化刻画。本文所讨论的理论见解为在协方差矩阵能有效刻画数据结构的应用中,采用VNN而非主流的基于PCA的学习流程提供了基本原理和依据。我们还阐明了这些基础性进展的影响如何渗透到协方差矩阵出现的广泛领域中学习方法的*有原则的*设计与应用。值得注意的是,我们阐明了VNN为利用神经影像数据集刻画神经退行性疾病的脑龄差这一特定任务所带来的概念性见解,这是计算神经科学中的一个时下热点问题。此外,还讨论了对其他应用领域的更广泛影响。
cs.LG / 78 / 2609.10505

Quantum Feature Engineering for Credit Default Prediction: When and Why IQP Circuits Help Linear Classifiers

面向信用违约预测的量子特征工程:IQP电路何时以及为何能帮助线性分类器
Finkelstein, Menachem, Levy, Diana Legziel, Yakhini, Zohar, Cohen, Sarel
Abstract
Credit default prediction is a tabular classification problem in which modest gains in F1 translate directly into reduced financial exposure. We ask whether Instantaneous Quantum Polynomial-time (IQP) circuits can produce features that improve a classifier over both its raw classical baseline and Kernel PCA - the strongest unsupervised classical non-linear alternative - at an equal feature budget. The dataset provides 23 financial attributes per client; for an n-qubit circuit we select n of them, encode each as a rotation angle, and read 2n expectation values back out as new features. The motivation for using a quantum circuit is computational: an n-qubit IQP circuit runs in constant depth and encodes feature correlations in a 2^n-dimensional Hilbert space, whereas classical simulation of its exact output statistics scales exponentially in n. Using the UCI Default of Credit Card Clients dataset and five-fold cross-validation, we find that appending 16 IQP features (n = 8 qubits) to a Logistic Regression model raises F1 from 0.462 to 0.517 (+0.055, p < 0.0001). Kernel PCA, the next-best method, reaches only 0.493 at the same feature count; the gap survives Benjamini-Hochberg correction across 12 tests (p = 0.00007). No other classifier - Random Forest, SVM, XGBoost, or k-NN - benefits, which points to a linear-expressivity mechanism rather than a generic improvement. We also show that how the 8 input features are chosen matters: Random Forest importance-guided selection reaches F1 = 0.523, while encoding maximally uncorrelated features drops it to 0.496, demonstrating that the circuit amplifies informative structure rather than creating it from scratch.
Chinese Translation
信用违约预测是一个表格分类问题,F1分数的适度提升可以直接转化为财务风险敞口的降低。我们探究瞬时量子多项式时间(Instantaneous Quantum Polynomial-time, IQP)电路能否在相同特征预算下,生成同时优于原始经典基线和核主成分分析(Kernel PCA,最强无监督经典非线性替代方法)的分类器特征。该数据集为每位客户提供23个金融属性;对于n量子比特电路,我们从中选择n个属性,将其编码为旋转角度,并读出2n个期望值作为新特征。使用量子电路的动机源于计算层面:n量子比特的IQP电路以常数深度运行,并在2^n维希尔伯特空间中编码特征相关性,而精确模拟其输出统计特性的经典方法随n呈指数级扩展。基于UCI信用卡客户违约数据集和五折交叉验证,我们发现向逻辑回归模型追加16个IQP特征(n = 8量子比特)可将F1分数从0.462提升至0.517(+0.055,p < 0.0001)。次优方法Kernel PCA在相同特征数量下仅达到0.493;该差距在经过12次检验的Benjamini-Hochberg校正后依然显著(p = 0.00007)。其他分类器——随机森林、SVM、XGBoost和k近邻——均未受益,这表明其机制在于线性表达能力而非通用性提升。我们还发现8个输入特征的选择方式至关重要:基于随机森林重要性引导的特征选择可达到F1 = 0.523,而编码最大不相关特征则使其降至0.496,这证明该电路是放大了信息性结构,而非从零创造结构。
cs.LG / 79 / 2609.10529

A positive resolution of the gap-entropy conjecture

间隙-熵猜想的肯定性解决
Aronow, P. M., Kallus, Nathan, Lopatto, Patrick
Abstract
We prove the gap-entropy conjecture for fixed-confidence best-arm identification with independent unit-variance Gaussian arms, means in $[0,1]$, and a unique optimal arm. For each suboptimal arm $i$, let $\Delta_i=\mu_*-\mu_i$ be its gap from the optimal mean, and write $H=\sum_{i\ne *}\Delta_i^{-2}$. Let $p_r$ be the fraction of $H$ contributed by arms with $2^{-(r+1)}<\Delta_i\le2^{-r}$, and let $\mathrm{Ent}(I)=\sum_{r:p_r>0} p_r\log(1/p_r)$. Among all algorithms that identify the optimal arm with probability at least $1-\delta$ on every Gaussian instance, the optimal expected number of samples on a given instance, averaged over all permutations of the arm labels, is within absolute constant factors of $H(\log(1/\delta)+\mathrm{Ent}(I))$. Moreover, there is an algorithm, independent of the instance, whose expected number of samples is bounded by a constant multiple of this quantity plus $g^{-2}\log\log(e^e/g)$, where $g=\min_{i\ne *}\Delta_i$ is the gap to the closest competitor.
Chinese Translation
我们证明了固定置信度最优臂识别问题中的间隙-熵猜想,考虑独立且方差为1的高斯臂、均值在 $[0,1]$ 内且最优臂唯一的情形。对于每个次优臂 $i$,令 $\Delta_i=\mu_*-\mu_i$ 表示其与最优均值的间隙,并记 $H=\sum_{i e *}\Delta_i^{-2}$。令 $p_r$ 为满足 $2^{-(r+1)}<\Delta_i\le2^{-r}$ 的臂对 $H$ 的贡献比例,并定义 $\mathrm{Ent}(I)=\sum_{r:p_r>0} p_r\log(1/p_r)$。在所有能在每个高斯实例上以至少 $1-\delta$ 的概率识别最优臂的算法中,在给定实例上对臂标签的所有置换取平均后,最优的期望样本数与 $H(\log(1/\delta)+\mathrm{Ent}(I))$ 相差绝对常数倍以内。此外,存在一种与实例无关的算法,其期望样本数被该量的常数倍再加上 $g^{-2}\log\log(e^e/g)$ 所界定,其中 $g=\min_{i e *}\Delta_i$ 表示与最接近竞争者之间的间隙。
机器人学 (Robotics)
42
cs.RO / 1 / 2609.09210

Identifying Habit, Physics, and Nuisance in Robot World Models

在机器人世界模型中识别习惯、物理与干扰因素
Hang, Jinting, Cai, Zhenhui
Abstract
Teleoperated demonstrations are often multimodal even when the underlying dynamics are nearly deterministic given the executed action. We argue that this multimodality typically mixes three factors--operator habit in action selection, shared physics, and observation nuisance--and that entangled next-observation predictors absorb all three. We formalize the split with a structural causal model a=g(h,z,u), z'=f(z,a), o=r(z,c), and test it with complementary interventions: replacing or shuffling actions at fixed state sharply increases next-state error, whereas appearance and camera changes should not; habit-aware reverse scoring improves ranking of feasible pasts without rewriting the dynamics. The associated adaptation rule is to freeze a shared physics readout and update only a thin interface. On StackCube, DROID, and RH20T this rule improves low-shot transfer relative to training from scratch, retains cleaner dynamics under corrupted adaptation data, and extends from proprioception to pixel observations with multi-view and multi-step checks. We do not equate latent actions with operator habit, and we do not target large-scale video generation benchmarks.
Chinese Translation
遥操作演示数据往往是多模态的,即使其底层动力学在给定执行动作的情况下近乎确定。我们认为,这种多模态通常混合了三个因素——操作者在动作选择上的习惯、共享的物理规律,以及观测干扰(nuisance)——而纠缠在一起的下一时刻观测预测器会同时吸收这三者。我们通过结构因果模型 a=g(h,z,u)、z'=f(z,a)、o=r(z,c) 对该分解进行形式化,并用互补的干预加以检验:在固定状态下替换或打乱动作会显著增加下一状态误差,而外观和相机的变化则不应如此;考虑习惯的反向评分(habit-aware reverse scoring)在不改写动力学的前提下改善了对可行历史状态的排序。相应的适配规则是冻结共享物理读取头,仅更新一个轻量的接口层。在 StackCube、DROID 和 RH20T 数据集上,该规则相对于从头训练改善了低样本迁移能力,在适配数据被污染的情况下仍能保持更干净的动力学,并且从本体感觉观测扩展到具备多视角与多步检验的像素观测。我们并不将潜在动作等同于操作者习惯,也不以大规模视频生成基准为目标。
cs.RO / 2 / 2609.09213

Geometry Conditioning in an Embodied SLM: Training Controls and Robustness Diagnostics in a 0.8B Hybrid Model

具身小语言模型(SLM)中的几何条件化:0.8B混合模型的训练控制与鲁棒性诊断
Li, Hao, Sun, Haofei, He, Lin
Abstract
We study how physical-state inputs affect a 0.8B hybrid language model adapted for manipulation with 6.2M trainable parameters. Six conditions are trained on three LIBERO-Spatial tasks and evaluated over three seeds and 540 held-out rollouts. Conditioning recurrent decay gates on geometric increments yields 28.9% success, compared with 36.7% when those increments are shuffled during training and 24.4% without explicit object/goal geometry. Both geometry policies receive correct inputs at evaluation. A token adapter using the same increments scores 27.8%; differences vary across seeds and remain inconclusive. Token-clock conditioning scores 11.1%, including one seed that fails to converge. In separate robustness tests, a state-only relative-coordinate policy retains 7/10 success under frame relabeling, whereas all four tested visual policies fall to at most 3/20 after a 5 cm object displacement. These results show no reliable advantage from training-time geometric alignment under this recipe and illustrate the gap between coordinate invariance and physical-layout generalization. Episode records, seed-level analyses, and figure-generation code accompany the paper.
Chinese Translation
我们研究了物理状态输入如何影响一个经过适配用于操作任务的0.8B混合语言模型,该模型仅含6.2M可训练参数。我们在三个LIBERO-Spatial任务上训练了六种条件设置,并在三个随机种子和540个保留测试回放(rollouts)上进行评估。将循环衰减门控条件化于几何增量可获得28.9%的成功率,相比之下,在训练期间对这些增量进行打乱时成功率为36.7%,而完全不使用显式物体/目标几何时为24.4%。两种几何策略在评估时均接收正确的输入。使用相同增量的令牌适配器(token adapter)得分27.8%;各结果在不同种子间存在差异,尚无定论。令牌时钟(token-clock)条件化得分为11.1%,其中包括一个未能收敛的种子。在单独的鲁棒性测试中,仅使用状态的相对坐标策略在坐标系重新标记下仍保持7/10的成功率,而所有四个经过测试的视觉策略在物体位移5厘米后均降至至多3/20的成功率。这些结果表明,在该训练方案下,训练时的几何对齐并无可靠优势,并揭示了坐标不变性与物理布局泛化之间的差距。论文附带回合记录、种子级分析及图表生成代码。
cs.RO / 3 / 2609.09217

Design and Attitude Control of an Underwater Quadruped Robot

水下四足机器人的设计与姿态控制
Molinaroli, Davide, Singh, Mohit, Alexis, Kostas
Abstract
Legged robots are versatile on land, but their use in underwater environments remains limited. Extending quadruped locomotion to water enables amphibious mobility with applications in inspection, environmental monitoring and disaster response. This paper presents the design, modeling, and experimental validation of a reproducible underwater quadruped robot. The robot is built around custom waterproof motor housings machined from polyoxymethylene plastic, which use off-the-shelf O-rings and dynamic shaft seals. A simplified model is derived to describe the dynamics of this underwater legged system, capturing how drag forces on spherical end effectors transmit torque to the floating base. Building on this model, a closed-loop attitude controller is developed using an error formulation defined on the special orthogonal group SO(3). The controller is evaluated both in simulation and experimentally in a water tank, where the robot tracks desired orientation setpoints in roll, pitch and yaw.
Chinese Translation
腿式机器人在陆地上具有高度的多样性,但其在水下环境中的应用仍然有限。将四足运动扩展至水中可实现两栖移动能力,在检测、环境监测和灾难响应等领域具有广泛的应用前景。本文介绍了一种可复现的水下四足机器人的设计、建模与实验验证。该机器人采用由聚甲醛(POM)塑料加工而成的定制防水电机外壳,使用市售O型密封圈和动态轴封。本文推导了一个简化模型来描述该水下腿式系统的动力学特性,刻画了球形末端执行器所受阻力如何将力矩传递至漂浮基座。基于该模型,利用定义在特殊正交群SO(3)上的误差形式开发了一种闭环姿态控制器。该控制器通过仿真和水箱实验进行评估,实验中机器人在横滚、俯仰和偏航方向上跟踪期望的姿态设定值。
cs.RO / 4 / 2609.09234

Learning to Fly: Stable Vision-Guided UAV Servoing with Compact Target-Centric Cues and Reinforcement Learning

学习飞行:基于紧凑目标中心线索与强化学习的稳定视觉引导无人机伺服控制
Jamwal, Saurbh Singh, Chebrolu, Nived
Abstract
Vision-guided reinforcement learning for Unmanned Aerial Vehicles (UAVs) remains challenging due to unstable policy optimisation, aggressive exploration, and the cost of high-dimensional visual perception. In this work, we investigate long-horizon UAV visual servoing using compact target-centric cues combined with low-dimensional sensor measurements. Rather than learning directly from RGB images, lightweight target segmentation provides image-space offsets and relative depth, which are combined with quadrotor velocity and projected-gravity measurements into a compact 12D policy observation. We compare Direct PPO with three matched-budget curriculum strategies: a Visual curriculum that progressively expands target placement difficulty, a Dynamics curriculum that gradually relaxes action constraints and smoothing, and a Joint curriculum that combines both progressions. All strategies reach comparable nominal performance, with complementary advantages across tracking metrics. Observation ablations show that proprioceptive measurements are critical for stable flight and image-space cues for target alignment, while explicit depth is not necessary for strong performance in the evaluated setting. Against tuned classical visual-servo controllers, learned policies show greater robustness to strong control and visual perturbations, while the Visual curriculum exhibits the smallest degradation under unseen target motion. Overall, the results demonstrate that compact target-centric representations can support robust long-horizon aerial visual servoing and that visual curriculum training can improve robustness to dynamic distribution shifts despite limited gains in nominal performance.
Chinese Translation
针对无人机的视觉引导强化学习(reinforcement learning)仍面临诸多挑战,包括策略优化不稳定、激进的探索行为以及高维视觉感知的高成本。本文研究了利用紧凑的目标中心线索结合低维传感器测量来实现长时域的无人机视觉伺服。我们并非直接从RGB图像中学习,而是采用轻量级目标分割提供图像空间偏移和相对深度,并将其与四旋翼速度和投影重力测量组合成一个紧凑的12维策略观测。我们比较了直接PPO(Direct PPO)与三种匹配训练预算的课程学习策略:逐步扩大目标放置难度的视觉课程(Visual curriculum)、逐步放宽动作约束与平滑限制的动力学课程(Dynamics curriculum),以及结合两者渐进过程的联合课程(Joint curriculum)。所有策略均达到了相近的标称性能,并在不同跟踪指标上展现出互补优势。观测消融实验表明,本体感知测量对稳定飞行至关重要,图像空间线索对目标对齐至关重要,而在所评估的场景中,显式深度对于取得良好性能并非必需。与经过调优的经典视觉伺服控制器相比,学习得到的策略对强控制扰动和视觉扰动表现出更强的鲁棒性,其中视觉课程策略在未见过的目标运动下性能退化最小。总体而言,结果表明紧凑的目标中心表征能够支持鲁棒的长时域空中视觉伺服,并且尽管在标称性能上提升有限,视觉课程训练仍能提高对动态分布偏移的鲁棒性。
cs.RO / 5 / 2609.09250

No Free Checker: A Survey of Verifiers for Robot Policies

没有免费的验证器:机器人策略验证器综述
Wan, Yang, Yue, Xihang, Liu, Zhirui, Chu, Ziyuan, Wang, Shuxun, Chen, Yuhan, Jiang, Xiaonan, Zhu, Xukun, Dong, Yubo, Zhu, Linchao
Abstract
A verifier for robot policies reads a candidate behavior and returns a score for how well it did, used both to evaluate vision-language-action policies and to train them. Verifiers range from success detectors and reward models to runtime monitors, safety filters, and temporal-logic specifications. We survey roughly 150 verifiers and compare them along two properties. Availability is how much a verdict costs, how early in a rollout the verdict arrives, and how often a verdict can be asked for. Availability rises as verdicts get cheaper, earlier, and denser. Credibility is how much a high score tells us about the task. Credibility falls as the judgment becomes gameable and self-serving. We group the verifiers by who supplies the judgment: human verifiers, rule-based and formal verifiers, learned and pretrained verifiers, and model-intrinsic verifiers. Across the four families, we find that credibility falls as availability rises. Regardless of who supplies the judgment, there is no free checker. We then examine what validates a verifier itself, and how much a high score tells us. Three measures appear in the literature: agreement with human labels, the performance of the policy it trains, and behavior under reward hacking. We close with nine metrics that make a verifier claim checkable, and coordinates for the verifiers still to be built.
Chinese Translation
机器人策略验证器读取候选行为并返回其表现优劣的评分,既用于评估视觉-语言-动作(vision-language-action)策略,也用于训练这些策略。验证器涵盖多种形式,包括成功检测器、奖励模型、运行时监控器、安全过滤器以及时序逻辑规范。本文调研了约150个验证器,并从两个属性出发对它们进行比较。可用性(Availability)指做出一次判定所需的代价、判定在策略执行过程中的多早阶段可以产生,以及判定的可请求频率——判定越廉价、越早、越密集,可用性越高。可信度(Credibility)指高评分在多大程度上能说明任务的实际完成情况——当评判容易被钻空子且具有自我服务倾向时,可信度会下降。我们根据评判的提供者将验证器分为四类:人类验证器、基于规则和形式化方法的验证器、基于学习及预训练的验证器,以及模型内在验证器。通过对比这四个类别,我们发现:可用性越高,可信度越低。无论评判由谁提供,都不存在免费的验证器。随后,我们探讨了如何验证验证器本身的有效性,以及高评分能在多大程度上说明问题。文献中出现了三种衡量方法:与人类标注的一致性、由该验证器训练出的策略的表现,以及在奖励作弊(reward hacking)下的行为表现。最后,我们提出了九项使验证器相关声明可被检验的度量指标,并为尚待构建的验证器指明了方向坐标。
cs.RO / 6 / 2609.09380

AccelMPC: High-Rate, Low-Power FPGA-Accelerated Model Predictive Control for Tiny Drones

AccelMPC:面向微型无人机的高速、低功耗FPGA加速模型预测控制
Grillo, Andrea, Plancher, Brian
Abstract
Unlocking the potential of tiny aerial robots requires order of magnitude improvements in the performance of embedded edge control. In particular, although recent cached model predictive control (MPC) solvers can handle the fast system dynamics and complex constraints required for agile drone flight, their computational demands remain prohibitive for resource-constrained robots, forcing prior implementations to operate at reduced control rates. AccelMPC overcomes this challenge through an end-to-end co-design approach that jointly optimizes the solver algorithm, numerical representation, hardware mapping, and physical integration. AccelMPC pairs a co-designed FPGA-accelerated alternating direction method of multipliers (ADMM)-based MPC solver with a custom 6g PCB, providing high-bandwidth communication for deployment on a 35g Crazyflie. Hardware experiments demonstrate 1 kHz onboard constrained MPC with dynamic obstacles, up to 15.6x faster solve times and 195.4x improvement in energy-delay product over state-of-the-art embedded microcontroller-based solvers, all while scaling to optimization problems with over 20,000 optimization variables and a comparable number of constraints. We release our PCB design files, firmware, and FPGA solver code open source.
Chinese Translation
释放微型飞行机器人的潜力,需要对嵌入式边缘控制的性能进行数量级的提升。尽管最近的缓存模型预测控制(MPC)求解器能够处理敏捷无人机飞行所需的快速系统动力学和复杂约束,但其计算需求对于资源受限的机器人而言仍然过于高昂,迫使先前的实现只能以降低的控制频率运行。AccelMPC通过端到端的协同设计方法克服了这一挑战,该方法联合优化了求解器算法、数值表示、硬件映射和物理集成。AccelMPC将协同设计的FPGA加速的交替方向乘子法(ADMM)MPC求解器与定制的6克PCB相结合,提供高带宽通信,以部署在35克的Crazyflie上。硬件实验表明,该系统实现了1 kHz的机载约束MPC并支持动态障碍物避障,与最先进的基于嵌入式微控制器的求解器相比,求解时间最高提速15.6倍,能量-延迟积改善195.4倍,同时可扩展至超过20,000个优化变量及相当数量约束的优化问题。我们已开源发布PCB设计文件、固件和FPGA求解器代码。
cs.RO / 7 / 2609.09403

A Decade of Bayesian Optimization for Controller Tuning and Robot Learning: Tutorial, Review, and Future Prospects

贝叶斯优化在控制器调参与机器人学习中的十年发展:教程、综述与未来展望
Stenger, David, Brunzema, Paul, Menn, Johanna, von Rohr, Alexander, Schoellig, Angela P., Trimpe, Sebastian
Abstract
In the past decade, Bayesian optimization (BO) has emerged as a powerful and adaptable framework for automatic controller tuning and robot learning. This article offers a comprehensive overview of the state-of-the-art in BO, designed to support both researchers and practitioners in understanding recent advancements, practical applications, and future research directions. We begin by adopting a practitioner's perspective, illustrating how to effectively set up BO through a representative controller tuning example. We position BO within the broader context of learning paradigms, ranging from deep reinforcement learning to data-driven control, and highlight scenarios where BO is most advantageous. Next, we discuss the diverse range of BO methods that have been developed to tackle complex problems and specific applications. This article provides a unified perspective on the current landscape of BO, emphasizing its relevance to control systems and robotics, and it highlights future prospects by identifying key research challenges and promising avenues for advancing BO in the field. This includes addressing a significant gap in the BO landscape: the lack of standardized benchmark problems specifically for control-related applications. To foster future research and ensure rigorous evaluation, we start an effort towards a lightweight benchmark suite for control engineering and robotics. We also present metrics and best practices to facilitate direct comparisons between new BO algorithms and established state-of-the-art methods.
Chinese Translation
在过去十年中,贝叶斯优化(Bayesian Optimization, BO)已成为一种强大且适应性强的框架,广泛应用于控制器自动调参和机器人学习。本文对贝叶斯优化的最新研究进展进行了全面综述,旨在帮助研究人员和从业者了解该领域的最新进展、实际应用及未来研究方向。我们首先从实践者的角度出发,通过一个具有代表性的控制器调参示例,说明如何有效搭建贝叶斯优化流程。我们将贝叶斯优化置于更广泛的学习范式背景下进行定位——涵盖从深度强化学习到数据驱动控制等方法——并着重指出贝叶斯优化最具优势的应用场景。接着,我们讨论了为应对复杂问题和特定应用而开发的各种贝叶斯优化方法。本文对贝叶斯优化的当前研究现状提供了统一视角,强调其在控制系统与机器人学中的相关性,并通过识别关键研究挑战和有前景的发展方向,展望了该领域贝叶斯优化的未来前景。这包括解决贝叶斯优化领域的一个重要缺口:目前缺乏专门针对控制相关应用的标准化基准测试问题。为促进未来研究并确保严格的评估,我们启动了一项面向控制工程与机器人学的轻量级基准测试套件的建设工作。我们还提出了相应的评估指标和最佳实践,以便于将新的贝叶斯优化算法与已有的最先进方法进行直接比较。
cs.RO / 8 / 2609.09492

Actuator Dynamics Curricula for Narrow-Viability Tasks in Legged Robot Learning

面向腿式机器人学习中窄可行域任务的执行器动力学课程
Chakraborty, Kousheek, Rajendra, Chandan K., Alharbat, Ayham, Mersha, Abeje Y.
Abstract
Reinforcement learning has produced capable controllers across a broad range of legged-robot tasks, but a subset of these tasks fail to converge under standard training: those for which most exploration trajectories terminate before producing useful gradient signal. To address such tasks we introduce the \emph{Actuator Dynamics Curriculum}, a procedure that initializes joint stiffness at a high value and anneals it toward the system-identified value as completed episode lengths grow. Using a cart-pole system as a representative example, we show that higher closed-loop joint natural frequency under critical damping enlarges the viability kernel of the underlying Markov Decision Process, increasing the fraction of initial states from which the task is feasible. We validate the kernel monotonicity on the cart-pole and apply the curriculum to a quadrupedal-to-handstand transition on the Boston Dynamics Spot, a narrow-viability task where training under fixed identified stiffness plateaus at a policy that never completes the transition. The trained policy executes the transition in simulation across 10 seeds and transfers to hardware. More broadly, our results suggest that simulated actuator dynamics is a useful axis along which to design curricula for tasks in which exploration is bottlenecked by termination conditions rather than by reward signal.
Chinese Translation
强化学习已在广泛的腿式机器人任务中产生出性能出色的控制器,但其中一部分任务在标准训练下无法收敛:即大多数探索轨迹在产生有效梯度信号之前即告终止的任务。为解决此类问题,我们提出了执行器动力学课程(Actuator Dynamics Curriculum),该方法将关节刚度初始化为较高的值,并随着已完成回合长度的增长,逐步将其退火至系统辨识所得的数值。以小车-倒摆系统作为代表性示例,我们证明了在临界阻尼下更高的闭环关节固有频率能够扩大底层马尔可夫决策过程的可行域核(viability kernel),从而提高任务可行的初始状态比例。我们在小车-倒摆系统上验证了该核的单调性,并将该课程应用于波士顿动力公司Spot四足机器人的四足到倒立转换任务——这是一个窄可行域任务,在固定辨识刚度下训练时,策略会停滞且始终无法完成该转换。训练所得的策略在仿真中以10个随机种子成功执行了该转换,并成功迁移到真实硬件上。更广泛而言,我们的结果表明,仿真执行器动力学是为那些探索瓶颈源于终止条件而非奖励信号的任务设计课程的一个有用维度。
cs.RO / 9 / 2609.09503

Agentic AI-enabled Semantic Commissioning of a Cognitive Digital Twin for Reconfigurable Manufacturing

面向可重构制造的智能体AI驱动的认知数字孪生语义调试
Liu, Yangyang, Xu, Xun, Polzer, Jan
Abstract
Rapid bespoke commissioning of the Cognitive Digital Twin (CDT) is a major challenge in reconfigurable manufacturing. Traditional digital twin (DT) construction methods primarily focus on geometric reconstruction, often neglecting the deep semantic integration and functional interoperability necessary for autonomous reasoning. This paper proposes an agent-based, AI-driven workflow to automate end-to-end CDT debugging. The system utilises LangGraph as a multi-agent orchestration engine to achieve dual-path synthesis: the semantic path extracts technical specifications from unstructured documents using Retrieval Augmented Generation (RAG), while the functional path autonomously discovers and binds to real-time industrial telemetry data using Model Context Protocol (MCP). Experimental validation in a robotic machining cell demonstrates that the system achieves a mean average accuracy (mAP) of 97.2% in perception and reduces the deployment cycle from several weeks to an average of 2 hours, marking a paradigm shift from manual scripting to autonomous orchestration.
Chinese Translation
认知数字孪生的快速定制化调试是可重构制造中的一项重大挑战。传统的数字孪生构建方法主要侧重于几何重建,往往忽视了自主推理所必需的深层语义集成与功能互操作性。本文提出一种基于智能体的AI驱动工作流,用于实现端到端的认知数字孪生调试自动化。该系统采用LangGraph作为多智能体编排引擎,实现双路径合成:语义路径利用检索增强生成(RAG)从非结构化文档中提取技术规范,功能路径则借助模型上下文协议自主发现并绑定实时工业遥测数据。在机器人加工单元中的实验验证表明,该系统在感知方面达到97.2%的平均精度均值(mAP),并将部署周期从数周缩短至平均2小时,标志着从手动脚本编写向自主编排的范式转变。
cs.RO / 10 / 2609.09597

Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints

用于提升任务的紧凑视觉-触觉世界模型:预测、奖励对齐与力约束
Ma, Qinzhen, Peng, Sida
Abstract
Accurate contact prediction is useful for robotic manipulation only if it supports effective decisions. We investigate this connection using a compact, randomly initialized visuotactile world model, trajectory-level uncertainty calibration, and behavior-initialized actor-critic learning in imagination. On 160 MuJoCo Lift episodes, adding touch reduces endpoint-force prediction error from 1.058 to 0.228 N and interval-peak error from 2.724 to 0.523 N across three training seeds. However, tactile persistence achieves lower errors of 0.095 and 0.498 N, respectively. Two exploratory control rounds comprise 680 executions on 40 independent test initial conditions. A matched reward revision on fresh test environments increases in-distribution 10 cm lifting success from 20.0% to 93.3%, while success within an 8 N per-finger budget reaches only 33.3%, compared with 70.0% for force feedback. Calibration margins reduce force violations at the cost of task completion. In a separate study of public GelSight recordings, a force regressor achieves 0.04234 N error, but frame-level calibration covers only 15.80% of complete trajectories; trajectory-level calibration raises this to 87.36% at nominal 90% coverage. Together, these findings distinguish improvements in sensing and task reward from improvements in force-constrained control. The evidence is limited to public sensing records and simulator execution, without a demonstrated transfer between them.
Chinese Translation
准确的接触预测只有在能够支持有效决策时,才对机器人操作具有实际价值。我们通过一个紧凑的、随机初始化的视觉-触觉(visuotactile)世界模型、轨迹级不确定性校准以及基于想象中的行为初始化的 actor-critic 学习来研究这一联系。在160个MuJoCo提升(Lift)回合中,在三个训练种子下,加入触觉信息将末端力预测误差从1.058 N降低到0.228 N,区间峰值误差从2.724 N降低到0.523 N。然而,触觉持续性(tactile persistence)方法分别达到了更低的误差,即0.095 N和0.498 N。两轮探索性控制实验在40个独立测试初始条件下共执行680次。在新测试环境上进行的匹配奖励修正使分布内10 cm提升成功率从20.0%提高到93.3%,而在每指8 N力预算内的成功率仅为33.3%,相比之下力反馈方法可达70.0%。校准裕度以牺牲任务完成度为代价减少了力违规。在另一项针对公开GelSight记录的研究中,力回归器达到了0.04234 N的误差,但帧级校准仅覆盖15.80%的完整轨迹;轨迹级校准在标称90%覆盖率下将这一比例提升至87.36%。总体而言,这些发现区分了传感与任务奖励方面的改进与力约束控制方面的改进。证据仅限于公开的传感记录和仿真器执行,尚未展示两者之间的迁移能力。
cs.RO / 11 / 2609.09612

MuJoCable: Reduced-Order Surface-Routed Cable Transmission for Tendon-Driven Robots

MuJoCable:面向腱驱动机器人的降阶表面走线缆绳传动模型
Zhang, Yi, Shao, Qi, Lin, Yicong, Ma, Muyuan, Sun, Tao, Xie, Yue
Abstract
Tendon transmissions reduce distal inertia and add compliance, yet routing, slack, and friction govern motion and force transfer. Mainstream rigid-body robotics simulators such as MuJoCo do not jointly resolve moving noncircular contact, unilateral tension, and segment friction. We present MuJoCable, which adds a reduced-order, configuration-dependent cable transmission to MuJoCo. Its routing algorithm jointly optimizes an ordered path across moving analytic and mesh surfaces. A unilateral axial law, directional Capstan propagation, and nodal virtual work map this path to segment tensions and body forces. The warm-started engine plugin applies these forces during simulation and exposes route and load states for design. Pulley benchmarks recover analytical transmission relations with a Capstan-ratio error below 0.5%. On the underactuated 18-joint SpiRobs, MuJoCable reveals friction-driven load growth and proximal redistribution of joint rotation that the native tendon does not represent. Hardware tests on SpiRobs and a tendon-route-coupled finger reproduce observed motion sequences. By making physical threading executable, MuJoCable brings transmission sources of the simulation-to-reality gap into route, cable, and actuator design before fabrication.
Chinese Translation
腱绳传动能够降低远端惯量并引入柔顺性,然而其走线方式、松弛与摩擦共同决定了运动和力的传递。MuJoCo 等主流刚体机器人仿真器无法同时求解运动中的非圆接触、单边张力约束和分段摩擦问题。我们提出了 MuJoCable,为 MuJoCo 增加了一种降阶的、依赖构型的缆绳传动模型。其走线算法在运动的解析曲面与网格曲面上联合优化一条有序路径。通过单边轴向力定律、方向性 Capstan(绞盘)摩擦传播以及节点虚功方法,该路径被映射为分段张力和刚体力。该采用热启动的引擎插件在仿真过程中施加这些力,并为设计提供走线与负载状态信息。滑轮基准测试复现了解析传动关系,Capstan 比误差低于 0.5%。在欠驱动的 18 关节 SpiRobs 机器人上,MuJoCable 揭示了原生腱模型无法表达的摩擦导致的负载增长以及关节旋转向近端的重新分布。在 SpiRobs 和一个腱-走线耦合手指上的硬件测试复现了所观察到的运动序列。通过使物理穿线过程可执行,MuJoCable 使仿真到现实差距中的传动误差来源能够在制造之前就被纳入走线、缆绳和执行器的设计之中。
cs.RO / 12 / 2609.09630

JEPA Policy: Diffusion-Free Imitation Learning via Paired Action and Future Representation Prediction

JEPA策略:通过配对动作与未来表征预测实现无扩散的模仿学习
Xu, Jie, Yu, Kangjin, Jin, Ziyi, Gao, Junjie, Chen, Liqing, Li, Yixian, Tian, Shuai, Xia, Zhongpu
Abstract
Standard behavior cloning supervises actions without explicitly constraining the future representation paired with each demonstrated action chunk. We introduce JEPA Policy, a diffusion-free framework that uses the action chunk and its observed future representation as paired training targets. Action and future-representation tokens interact in a shared Transformer and are refined through two forward passes. Future prediction can therefore shape the representation used to generate actions. Dual-branch and gradient-routing controls attribute the gain to this shared topology rather than to an auxiliary prediction head alone. Across nine simulated tasks, JEPA Policy improves mean success over the action-only MIP baseline and outperforms Diffusion Policy under the evaluated configurations, while adding 0.29 ms to MIP's model latency. A five-task, 630-episode physical-robot study produces the same pooled ranking. Further audits find no complete representation collapse under action supervision and identify a task-conditioned failure-ranking signal in future-prediction error. These results support paired future-representation supervision as a practical approach to low-latency visuomotor imitation without iterative generative sampling.
Chinese Translation
标准行为克隆仅对动作进行监督,而未显式约束与每个示范动作块(action chunk)配对的未来表征。我们提出JEPA策略(JEPA Policy),这是一个无扩散的框架,将动作块及其观测到的未来表征作为配对训练目标。动作标记与未来表征标记在一个共享的Transformer中交互,并通过两次前向传播进行细化。因此,未来预测能够塑造用于生成动作的表征。双分支和梯度路由控制实验表明,性能增益源于这种共享拓扑结构,而非仅仅来自一个辅助预测头。在九个仿真任务中,JEPA策略的平均成功率优于仅使用动作的MIP基线,并在所评估的配置下超过了扩散策略(Diffusion Policy),同时仅为MIP的模型推理延迟增加0.29毫秒。一项包含五个任务、630个回合的实体机器人研究也得出了相同的总体排序。进一步的审计发现,在动作监督下不存在完全的表征坍缩,并在未来预测误差中识别出一种任务条件化的失败排序信号。这些结果支持将配对的未来表征监督作为实现无需迭代生成采样的低延迟视觉运动模仿学习的一种实用方法。
cs.RO / 13 / 2609.09650

A Risk-Sensitive and Uncertainty-Aware Decision-Making and Control Framework for Safe and Robust Autonomous Driving

面向安全且鲁棒自动驾驶的风险敏感与不确定性感知决策与控制框架
Li, Zhuoren, Yu, Ran, Zhang, Weiqi, Liu, Ming, Xiong, Lu, Sun, Chen, Leng, Bo
Abstract
Reinforcement learning (RL) has demonstrated considerable potential for autonomous driving decision-making. However, its deployment in urban autonomous driving, particularly at highly interactive unsignalized intersections, remains challenging, as learned policies may struggle to maintain both safety and robust decision-making in complex traffic situations. Conventional safety-filtering approaches typically employ fixed conservative constraints, which may improve safety at the cost of excessive intervention and degraded traffic efficiency. To address these limitations, we propose a Risk-sensitive and Uncertainty-aware Decision-making and Control (RUDC) framework for safe and robust autonomous driving. RUDC couples risk-sensitive distributional RL with ensemble-based policy uncertainty quantification, jointly accounting for tail risks in return distributions and uncertainty in learned policies. An uncertainty-aware high-order control barrier function (HOCBF)-based safety correction mechanism adaptively adjusts constraint strictness according to policy uncertainty, while a learnable residual predictor compensates for CBF model mismatches and discretization errors. Extensive simulations at unsignalized intersections demonstrate that RUDC achieves a favorable balance among safety, efficiency, and robustness, outperforming representative safe RL baselines under both nominal and challenging OOD and long-tail scenarios while satisfying real-time requirements.
Chinese Translation
强化学习(RL)在自动驾驶决策方面展现出巨大潜力。然而,将其部署于城市自动驾驶场景,尤其是高度交互的无信号交叉口,仍然充满挑战,因为学习到的策略在复杂交通状况下可能难以兼顾安全性与决策的鲁棒性。传统的安全过滤方法通常采用固定的保守约束,虽可能提升安全性,却以过度干预和交通效率下降为代价。为解决这些局限,我们提出了一种面向安全且鲁棒自动驾驶的风险敏感与不确定性感知决策与控制(RUDC)框架。RUDC 将风险敏感的分布式强化学习与基于集成(ensemble)的策略不确定性量化相结合,共同考虑回报分布中的尾部风险和学习策略中的不确定性。该框架包含一种基于不确定性感知高阶控制障碍函数(HOCBF)的安全修正机制,可根据策略不确定性自适应地调整约束的严格程度;同时引入可学习的残差预测器,以补偿 CBF 模型失配和离散化误差。在无信号交叉口的大量仿真实验表明,RUDC 在安全性、效率和鲁棒性之间取得了良好的平衡,在标称场景以及具有挑战性的分布外(OOD)和长尾场景下均优于代表性的安全强化学习基线方法,同时满足实时性要求。
cs.RO / 14 / 2609.09692

CT-SAFR: Safe and Interpretable Chain-of-Thought Reasoning for Autonomous Robots: A Multi-Layered Verification Framework for Trustworthy AI-Driven Robotic Decision Making

CT-SAFR:面向自主机器人的安全且可解释的思维链推理:用于可信AI驱动的机器人决策的多层验证框架
Temel, Cagri
Abstract
Chain-of-Thought (CoT) prompting enables LLMs to perform explicit, step-by-step reasoning, creating opportunities for sophisticated autonomous robots. However, recent research reveals that reasoning models verbalize their actual decision processes only 25-39% of the time, with faithfulness degrading 44% on complex tasks. This paper presents CT-SAFR (Chain-of-Thought Safety and Faithfulness for Robotics), a multi-layered verification framework achieving 94.2% hallucination detection (n = 500, 95% CI: 91.8-95.9%) with sub-500ms latency. Through a warehouse robot case study, this work demonstrates 87% reduction in unsafe reasoning outputs (p < 0.001) and provides recommendations for responsible deployment of reasoning-capable autonomous robots.
Chinese Translation
思维链提示使大语言模型(LLM)能够进行显式的、逐步的推理,为高度智能化的自主机器人创造了机会。然而,近期研究表明,推理模型仅有25%-39%的情况下会真实表达其实际决策过程,且在复杂任务上忠实度下降44%。本文提出了CT-SAFR(Chain-of-Thought Safety and Faithfulness for Robotics,面向机器人的思维链安全性与忠实性框架),这是一种多层验证框架,实现了94.2%的幻觉检测率(n = 500,95%置信区间:91.8-95.9%),且延迟低于500毫秒。通过一个仓储机器人案例研究,本工作展示了不安全推理输出减少87%(p < 0.001),并为具备推理能力的自主机器人的负责任部署提供了建议。
cs.RO / 15 / 2609.09745

PccDiffuser: Multi-solution Motion Planning for Continuum Robots

PccDiffuser:连续体机器人的多解运动规划方法
Qiu, Ke, Chen, Sifan, Wang, Si, Xiong, Rong, Wang, Yue, Lu, Haojian
Abstract
We present the PccDiffuser, a conditional diffusion framework for continuum robots that learns a multimodal distribution over complete configuration-space paths and samples multiple candidate solutions in parallel, which are subsequently converted into an executable trajectory by time allocation considering actuator constraints. Under the piecewise constant-curvature model, we use exponential co-ordinates to describe the robot kinematics, and use graph neural network to encode a variable number of environment obstacles. Analytical differential kinematics is incorporated in the denoising process to improve terminal accuracy and whole-body clearance. On a mixed test set comprising workspace with zero to four obstacles, PccDiffuser achieved a success rate of 91\%. Compared with existing sampling- and optimisation-based benchmarks, it delivered both a higher success rate and greater computational efficiency, with the latter advantage becoming more substantial when sampling more candidate solutions. Experiments on a three-section tendon-driven continuum robot further demonstrate consecutive planning, multi-solution planning, and whole-body obstacle avoidance.
Chinese Translation
我们提出了PccDiffuser,一个面向连续体机器人的条件扩散框架。该框架学习完整构型空间路径的多模态分布,并行采样多个候选解,随后在考虑执行器约束的条件下通过时间分配将其转换为可执行轨迹。在分段常曲率模型下,我们采用指数坐标描述机器人运动学,并利用图神经网络对数量可变的环境障碍物进行编码。我们将解析微分运动学融入去噪过程,以提升末端精度和全身避障间隙。在一个包含零至四个障碍物工作空间的混合测试集上,PccDiffuser取得了91%的成功率。与现有的基于采样和基于优化的基准方法相比,该方法在成功率与计算效率上均表现更优,且当采样更多候选解时,其计算效率优势更加显著。在三节腱驱动连续体机器人上的实验进一步验证了连续规划、多解规划以及全身避障能力。
cs.RO / 16 / 2609.09752

HiRAD: A Flexible Large-Scale AGV Routing System

HiRAD:一种灵活的大规模AGV路由系统
Huang, Yunjie, Wu, Ruizhong, Zhang, Mengxuan, Chan, Frodo Kin Sun, Law, Yan Nei, Li, Lei
Abstract
Automatic Guided Vehicles (AGVs) substantially boost warehouse throughput, but routing large-scale AGV fleets remains challenging. Classical Multi-Agent Pathfinding solvers suffer from exploding combinatorial complexity and super-quadratic runtime, while relying on idealized grid or piecewise-linear motion models that mismatch real-world kinematics. Recent Reinforcement Learning (RL) solutions improve flexibility via decentralized agent policies but depend on discretized spatiotemporal representations, require millions of episodes to converge, and incur full-map observation at every step, which leads to large models, slow convergence, and high inference latency that violates real-time industrial control constraints. To address these bottlenecks, we propose HiRAD, a hierarchical RL framework for continuous-space AGV routing with real-time guarantees: (1) a step-level spatiotemporal representation that translates continuous motion into a differentiable RL problem, (2) a hierarchical strategy that splits heading choice from velocity control to reduce the action space, and (3) an asynchronous event-driven decision pipeline that lowers inference complexity from O(n^2) to O(n) and cuts per-step latency by as much as 71 percent. Across random graphs and two warehouse maps, HiRAD reduces makespan by 45 percent to 63 percent and shortens end-to-end runtime.
Chinese Translation
自动导引车(AGV)显著提升了仓库吞吐量,但大规模AGV车队的路径规划仍然具有挑战性。经典的多智能体路径规划(Multi-Agent Pathfinding)求解器存在组合复杂度爆炸和超二次运行时间的问题,并且依赖于理想化的网格或分段线性运动模型,与真实世界的运动学不符。近年来基于强化学习(RL)的解决方案通过去中心化的智能体策略提高了灵活性,但其依赖离散化的时空表示,需要数百万轮训练才能收敛,且每一步都需要全地图观测,导致模型庞大、收敛缓慢、推理延迟高,违背了工业实时控制的约束。为解决这些瓶颈,我们提出了HiRAD——一种具有实时保证的连续空间AGV路由分层强化学习框架:(1)步级时空表示,将连续运动转化为可微分的强化学习问题;(2)分层策略,将航向选择与速度控制分离以缩减动作空间;(3)异步事件驱动的决策流水线,将推理复杂度从O(n^2)降至O(n),并将每步延迟最多削减71%。在随机图和两张仓库地图上,HiRAD将完工时间缩短了45%至63%,并缩短了端到端运行时间。
cs.RO / 17 / 2609.09808

GTA-2: A Multi-VLM Framework for Synthesizing Robot Manipulation Skills via Grounded Task Axes

GTA-2:一种基于接地任务轴的多视觉语言模型框架,用于合成机器人操作技能
Seker, M. Yunus, Aggarwal, Shobhit, Wickramarachchi, Ruwan, Francis, Jonathan, Kroemer, Oliver
Abstract
Robotic manipulation tasks are often decomposed into behaviors or skills. However, one often needs to predefine these behaviors for specific tasks or try to cover a wide range of tasks using generic skills. As a result, these behaviors can remain too coarse to expose the geometric, control, and scene-dependent decisions required for execution. We introduce Grounded Task Axes v2 (GTA-2), a modular multi-VLM framework that constructs executable, task-bespoke manipulation skills from reusable object-centric task-axis components. Rather than predicting actions end-to-end or composing fixed task-level primitives, GTA-2 represents each skill as semantic subtasks comprising task-relevant keypoints and axes, controller compositions, and scene-dependent parameters. Four specialized VLM agents separately decompose the task, construct an abstract task-axis skill, assign controller parameters, and ground the required visual features from RGB-D observations. This abstraction-to-grounding factorization enables zero-shot skill generation without task-specific robot demonstrations, policy training, or fine-tuning. It also keeps intermediate decisions explicit, allowing targeted human feedback to refine an incorrect stage while preserving correct components. We evaluate GTA-2 on 14 real-robot manipulation tasks against a VLA policy pi_{0.5} and two Code-as-Policies baselines using task-axis controllers or conventional robot primitives. GTA-2 achieves an average zero-shot success rate of 73.9%, exceeding the strongest baseline by 31.4 percentage points, while targeted refinement raises GTA-2's average success rate to 90.7%. Project page: https://gta2-project.github.io/
Chinese Translation
机器人操作任务通常被分解为行为或技能。然而,人们往往需要为特定任务预先定义这些行为,或者试图通过通用技能来覆盖广泛的任务。因此,这些行为可能过于粗糙,无法揭示执行所需的几何、控制和场景相关决策。我们提出了接地任务轴v2(Grounded Task Axes v2,GTA-2),这是一个模块化的多视觉语言模型(VLM)框架,能够从可复用的以对象为中心的任务轴组件构建可执行的、任务定制化的操作技能。GTA-2并不进行端到端的动作预测,也不组合固定的任务级原语,而是将每个技能表示为语义子任务,这些子任务由任务相关的关键点和轴、控制器组合以及场景相关参数构成。四个专用的VLM智能体分别执行任务分解、构建抽象任务轴技能、分配控制器参数,并从RGB-D观测中接地所需的视觉特征。这种从抽象到接地的因式分解实现了零样本技能生成,无需任务特定的机器人演示、策略训练或微调。它还将中间决策保持显式化,使得针对性的人类反馈能够修正错误的阶段,同时保留正确的组件。我们在14个真实机器人操作任务上评估了GTA-2,对比对象包括VLA策略pi_{0.5}以及两个分别使用任务轴控制器或传统机器人原语的Code-as-Policies基线方法。GTA-2的平均零样本成功率达到73.9%,超过最强基线31.4个百分点,而通过针对性修正,GTA-2的平均成功率可提升至90.7%。项目页面:https://gta2-project.github.io/
cs.RO / 18 / 2609.09918

ViBe: Visual Behavior Adaptation for Perceptive Humanoid Whole-Body Control

ViBe:面向感知型人形机器人全身控制的视觉行为自适应方法
Krishna, Lokesh, Venkatesan, Sarvesh, Zhang, An, Nguyen, Quan
Abstract
Motion tracking provides a scalable recipe for humanoid whole-body control. By design, the resulting trackers lack exteroceptive feedback hence reacting to the environment remains the responsibility of a higher-level planner. Existing perceptive controllers train geometry-only encoders from scratch, trading semantics for sim-to-real ease, and typically rely on teacher-student distillation for a task of interest. We present ViBe, a post-training framework for adapting motion trackers to perceptive control tasks. We leverage pre-trained visual encoders with a multi-query extractor module to learn task-relevant perceptive feedback. This feedback is grafted onto the tracker's input via low-rank adapters, enabling parameter-efficient fine-tuning. Given a task reward and a reference dataset, this modular controller can be adapted directly via policy optimization. Across four tasks, ViBe shows zero-shot sim-to-real transfer spanning perceptive walking on curbs and parkour, Repose Cube, omni-object loco-manipulation, and dodgeball, with visually robust performance across outdoor, low-light, and RGB distractor conditions. Finally, we solve a goal-oriented Repose Cube task with a deliberately simple planner, demonstrating the efficacy of perceptive controllers, adapted by our approach.
Chinese Translation
运动追踪为人形机器人全身控制提供了一种可扩展的方案。然而,这类追踪器在设计上缺乏外部感知反馈,因此对环境的响应仍需依赖更高层的规划器。现有的感知型控制器通常从零开始训练仅基于几何信息的编码器,以牺牲语义信息来换取仿真到现实迁移的便利性,并且通常依赖师生蒸馏来完成特定任务。我们提出ViBe,一种用于将运动追踪器适配到感知控制任务的后训练框架。我们利用预训练的视觉编码器和多查询提取器模块来学习与任务相关的感知反馈。该反馈通过低秩适配器嫁接到追踪器的输入上,实现参数高效微调。给定任务奖励和参考数据集,该模块化控制器可通过策略优化直接进行适配。在四项任务上,ViBe展现了零样本仿真到现实的迁移能力,涵盖路沿行走与跑酷、Repose Cube、全物体移动操作以及躲避球任务,并在户外、低光照和RGB干扰物条件下均表现出视觉鲁棒的性能。最后,我们使用一个刻意保持简单设计的规划器解决了目标导向的Repose Cube任务,验证了经本方法适配的感知控制器的有效性。
cs.RO / 19 / 2609.09941

HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy

HaWMPO:面向通用机器人策略的幻觉感知世界模型策略优化
Chen, Zengjue, Liu, Peidong, Li, Jiawei, Wang, Qi
Abstract
Generalist robot policies have demonstrated strong generalization across robotic manipulation tasks, yet their success rates remain limited in com- plex long-horizon scenarios. Recent methods improve Visual-Language-Action (VLA) policies through online reinforcement learning on real robots, but such training relies on costly physical interactions, suffers from low sample efficiency, and may introduce hardware and safety risks. World models offer a promising alternative by enabling policy optimization with imagined rollouts. However, long-horizon rollouts generated by world models often suffer from prediction hal- lucinations, producing biased state transitions that can mislead policy learning. To address this issue, we propose Hallucination-aware World Model-based Pol- icy Optimization (HaWMPO), a closed-loop reinforcement learning pipeline for VLA policy post-training with world models. Specifically, HaWMPO introduces an action-conditioned hallucination-aware model to estimate the reliability of gen- erated image sequences, and incorporates hallucination scores into group relative policy optimization through a Reward-Soft mechanism, suppressing unreliable ac- tion chunks during training. On the LIBERO benchmark, HaWMPO achieves the best average success rate, with gains of 15.0% over the base model and 2.8% over the strongest baseline; real-world experiments on a G1 robot further validate its effectiveness, raising the average success rate on two manipulation tasks from 67.5% to 80.0%.
Chinese Translation
通用机器人策略在各类机器人操作任务中展现出强大的泛化能力,但在复杂的长时程场景中,其成功率仍然有限。近期的方法通过在真实机器人上进行在线强化学习来改进视觉-语言-动作(VLA)策略,但此类训练依赖昂贵的物理交互,样本效率低下,并可能带来硬件损坏和安全风险。世界模型为策略优化提供了一种有前景的替代方案,它可以通过想象生成的轨迹来优化策略。然而,世界模型生成的长时程轨迹往往存在预测幻觉问题,产生有偏差的状态转移,从而误导策略学习。为解决这一问题,我们提出了幻觉感知世界模型策略优化(HaWMPO),这是一个基于世界模型的VLA策略后训练闭环强化学习流程。具体而言,HaWMPO引入了一个动作条件化的幻觉感知模型来估计生成图像序列的可靠性,并通过Reward-Soft机制将幻觉分数融入组相对策略优化(GRPO)中,在训练过程中抑制不可靠的动作块。在LIBERO基准测试中,HaWMPO取得了最佳平均成功率,相比基础模型提升了15.0%,相比最强基线提升了2.8%;在G1机器人上的真实世界实验进一步验证了其有效性,将两个操作任务的平均成功率从67.5%提升至80.0%。
cs.RO / 20 / 2609.10021

RoboDrop: Curating VLA Post-Training Data via Local Gradient Compatibility

RoboDrop:基于局部梯度相容性的VLA后训练数据筛选方法
Xu, Runze, Xu, Yuanfan, Xu, Cuijie, Dai, Shuang, Li, Yining, Wang, Yu, Yu, Jincheng
Abstract
Vision--language--action (VLA) models acquire broad generalization through large-scale pretraining, yet adapting them to a new task and robot embodiment still requires post-training on newly collected data. Unlike pretraining, post-training targets task- and embodiment-specific adaptation, making it particularly sensitive to data quality. In practice, collected robot datasets often contain heterogeneous errors, including execution mistakes, sensor drift, and timestamp misalignment, which can impair post-training and policy performance. Manual inspection is costly, while existing data-cleaning methods are typically tailored to particular corruption types. To address these challenges, we introduce \textsc{RoboDrop}, a data-curation framework that audits supervision using local gradient compatibility measured along the training trajectory as a proxy for its effect on post-training performance. During a one-epoch warm-up run, RoboDrop scores each candidate sample online by comparing its gradient with those of task-semantic and visually matched validation samples. The resulting sample scores are aggregated at the episode level, and a simple automatic post-processing rule converts them into filtering decisions. We evaluate RoboDrop on controlled observation--action corruptions, naturally suboptimal demonstrations in simulation, and real-robot datasets containing non-expert collection errors. Across these settings, RoboDrop more accurately distinguishes unreliable demonstrations than prior methods, while post-training on the curated data consistently yields stronger downstream policies, with average real-robot rollout success rising from $35.0\%$ to $67.5\%$. These results establish training-trajectory-aware, context-conditioned supervision auditing as an effective approach to robust VLA post-training.
Chinese Translation
视觉-语言-动作(VLA)模型通过大规模预训练获得广泛的泛化能力,但将其适配到新任务和新机器人本体仍需要在新采集的数据上进行后训练。与预训练不同,后训练的目标是任务和本体特定的适配,因此对数据质量尤为敏感。在实践中,采集的机器人数据集往往包含各种异构错误,包括执行失误、传感器漂移和时间戳错位,这些错误会损害后训练效果和策略性能。人工检查成本高昂,而现有的数据清洗方法通常仅针对特定类型的损坏。为应对这些挑战,我们提出了RoboDrop,一个数据筛选框架,它以训练轨迹上测量的局部梯度相容性作为对后训练性能影响的代理指标,对监督数据进行审计。在单轮预热训练过程中,RoboDrop通过将每个候选样本的梯度与任务语义匹配和视觉匹配的验证样本的梯度进行比较,在线地对样本进行评分。所得的样本评分在片段(episode)层面进行聚合,并通过一个简单的自动后处理规则将其转化为过滤决策。我们在受控的观测-动作损坏、仿真中自然存在的次优示教以及包含非专家采集误差的真实机器人数据集上评估了RoboDrop。在所有这些场景中,RoboDrop都比先前方法更准确地区分不可靠示教,而在筛选后的数据上进行后训练始终能获得更强的下游策略,真实机器人执行的平均成功率从35.0%提升至67.5%。这些结果表明,基于训练轨迹感知的、上下文条件化的监督审计是鲁棒VLA后训练的一种有效方法。
cs.RO / 21 / 2609.10024

AXON: A ROS 2 RMW with Shared-Memory/QUIC Transport and QKD/ML-KEM Key Establishment

AXON:一种采用共享内存/QUIC传输与QKD/ML-KEM密钥建立的ROS 2 RMW实现
de la Fuente, Sergio Sánchez, González-Santamarta, Miguel Ángel, Rodríguez-Lera, Francisco Javier, Olivera, Vicente Matellán, Guerrero-Higueras, Ángel Manuel
Abstract
Robot Operating System 2 (ROS 2) standardizes application code against a middleware interface (RMW) whose reference implementations are built on the Data Distribution Service (DDS). We present AXON, an alternative ROS 2 RMW implementation that separates transport policy by deployment scope. A Rust core and C++ adapter use POSIX shared-memory rings for same-host communication, QUIC for remote communication, and a daemon for discovery and graph synchronization. We then describe two fail-closed TLS 1.3 key-establishment configurations for remote traffic. The classic configuration offers only the hybrid X25519MLKEM768 group, preventing negotiation of a classical-only group. The qkd configuration imports a 256-bit key obtained through the ETSI GS QKD 014 API as a pairwise external PSK and offers no Diffie-Hellman group. Its default messages10 strategy additionally protects remote application messages with AES-256-GCM, rotating KME material after ten outgoing messages and using a fresh nonce per envelope; session relies on QUIC protection alone. The external-PSK path requires a narrow extension to rustls, now bundled with AXON. We define the threat model, distinguish peer authentication in the two configurations, and delimit the implementation-level validation from ROS 2 conformance, comparative performance, and physical-QKD validation.
Chinese Translation
机器人操作系统2(ROS 2)通过中间件接口(RMW)对应用程序代码进行标准化,其参考实现基于数据分发服务(DDS)。我们提出了AXON,一种替代性的ROS 2 RMW实现,它根据部署范围区分传输策略。其Rust核心与C++适配器在主机间通信中使用POSIX共享内存环形缓冲区,在远程通信中使用QUIC,并通过一个守护进程实现发现与图同步。随后,我们描述了两种面向远程流量的故障即关闭(fail-closed)TLS 1.3密钥建立配置。经典配置(classic)仅提供混合X25519MLKEM768群组,从而防止协商仅经典的群组。qkd配置将通过ETSI GS QKD 014 API获取的256位密钥作为成对外部PSK导入,且不提供Diffie-Hellman群组。其默认的messages10策略还使用AES-256-GCM对远程应用消息提供额外保护,每发送十条消息后轮换一次KME密钥材料,并为每个数据包使用新的随机数(nonce);会话安全性仅依赖QUIC自身的保护。外部PSK路径需要对rustls进行少量扩展,该扩展现已随AXON一并打包。我们定义了威胁模型,区分了两种配置中的对等端身份验证方式,并明确划定了实现层面的验证与ROS 2一致性测试、对比性能评估以及物理QKD验证之间的边界。
cs.RO / 22 / 2609.10033

What Symmetry Buys a Learned Motion Planner

对称性能为学习型运动规划器带来什么
Sevincel, Andrea Emir
Abstract
Learning-based motion planners pay at training what classical planners pay per query. Trained in world coordinates, they relearn the same motion at every position and orientation. Existing work restores the missing rigid-body equivariance in the training data, in the inference operator, or in the weights, and each carries a cost. We ask how much of that equivariance the planning query supplies for free. A start s and a goal g determine a frame in closed form, with origin at their midpoint and first axis along g-s. Expressing trajectory and obstacles in that frame removes three translations and two rotations of SE(3), at initialisation, for one cross product per query and with no constraint on the architecture. A single rotation about the start-goal axis remains, and no continuous rule removes it. On a cluttered 3D benchmark, holding architecture, data and budget fixed, the frame raises the held-out collision-free rate from 14.60% to 51.10%, where a straight segment from start to goal scores 15.6% and the world-frame model does not beat it. We build all three mechanisms for the residual rotation and each is worth under a point, though the equivariant backbone reaches any given level two to three times sooner. What the representation supplies therefore dominates what any mechanism enforces, and the standard diagnostic does not see the difference: two models with indistinguishable non-equivariance residuals differ by 28 points. Calibrated against a non-symmetry intervention, the frame is not even the largest effect available, since local geometry is worth +40.0 where the frame is worth +36.5.
Chinese Translation
基于学习的运动规划器在训练阶段付出代价,而经典规划器在每次查询时付出代价。在世界坐标系中训练的模型,需要在每个位置和姿态上重复学习相同的运动。现有工作通过在训练数据、推理算子或网络权重中恢复缺失的刚体等变性来解决这一问题,但每种方法都有其代价。我们探讨规划查询本身能在多大程度上免费提供这种等变性。起点 s 和终点 g 可以以闭式解确定一个坐标系,其原点位于两者的中点,第一轴沿 g-s 方向。在该坐标系中表示轨迹和障碍物,可以在初始化时消除 SE(3) 中的三个平移和两个旋转,每次查询仅需一次叉积计算,且对网络架构没有任何约束。仅剩绕起点-终点轴的单个旋转,且没有连续规则可以消除它。在一个杂乱 3D 基准测试中,在保持架构、数据和计算预算不变的情况下,该坐标系将留出集上的无碰撞率从 14.60% 提升到 51.10%,而直接连接起点到终点的直线段得分为 15.6%,世界坐标系模型未能超越它。我们为这一剩余旋转构建了全部三种机制,但每种机制的价值都不足一个百分点,尽管等变主干网络能以两到三倍的速度达到任意给定水平。因此,表示本身所提供的收益主导了任何机制所强制的收益,而标准诊断方法却无法看出这种差异:两个不可区分的非等变残差模型之间相差 28 个百分点。与非对称性干预的对照标定表明,该坐标系甚至不是可获得的最大效应,因为局部几何的价值为 +40.0,而坐标系的价值为 +36.5。
cs.RO / 23 / 2609.10050

Grounding Generated Video Plans in Simulation Towards Versatile Dexterous Controllers

在仿真中落地生成的视频规划以构建多功能灵巧控制器
Wu, Tianyue, An, Boyuan, Zhao, Shuqi, Guo, Heyu, Xing, Wanli, Ma, Yi, Zhang, Kaifeng, Wu, Ruihai, Tomizuka, Masayoshi
Abstract
Generated hand-object interaction (HOI) videos provide a controllable way to propose manipulation motions. Simulation-based HOI tracking can translate such kinematic references into feasible low-level control, but its scalability is limited by the lack of reliable reference motions. We therefore combine generated videos with simulation-based HOI grounding: during training, generated videos provide diverse motion references for learning a multi-object, multi-trajectory HOI tracker, and at deployment, the video model produces motion plans that are executed by the learned tracker. In particular, we propose a method that enables scalable reference generation by HOI reconstruction with minimal manual intervention and successfully grounds more than 1,500 generated videos in simulation, achieving success rates over 25 percentage points higher than those of baselines during simulation-based training. In real-world closed-loop experiments, it achieves diverse grasps, including functional grasps, non-prehensile manipulation, and post-grasp object-pose tracking. Videos and code are available at https://boyuan-an.github.io/GALATEA/.
Chinese Translation
生成的手-物体交互(HOI)视频为提出操作运动提供了一种可控的方式。基于仿真的HOI跟踪可以将此类运动学参考转化为可行的低层控制,但其可扩展性受限于缺乏可靠的参考运动。因此,我们将生成视频与基于仿真的HOI落地(grounding)相结合:在训练阶段,生成的视频为学习多物体、多轨迹的HOI跟踪器提供多样化的运动参考;在部署阶段,视频模型生成运动规划,并由学习到的跟踪器执行。特别地,我们提出了一种方法,能够通过最少人工干预的HOI重建实现可扩展的参考运动生成,并成功地在仿真中落地了超过1500个生成视频,在基于仿真的训练中取得了比基线方法高出25个百分点以上的成功率。在真实世界的闭环实验中,该方法实现了多样化的抓取,包括功能性抓取、非抓握式操作以及抓取后的物体位姿跟踪。视频和代码可在 https://boyuan-an.github.io/GALATEA/ 获取。
cs.RO / 24 / 2609.10082

Automatic Reproducible Camera Intrinsic Calibration

自动化的可复现相机内参标定
Hu, Xiangcheng
Abstract
Accurate camera intrinsic calibration is fundamental to robot perception, and the accuracy depends on the quality of the collected images. However, existing target-based calibration methods often require the practitioner to manually filter out high-quality images and to specify an appropriate radial distortion order. This paper presents a fully automatic intrinsic calibration pipeline that determines both from the collected data. We adopt an iterative rejection scheme that estimates parameters on a candidate image set and removes views whose mean residual exceeds a multiple of the median. Crucially, this process runs independently under each candidate distortion order, so that the retained image set is consistent with the residual scale of that order. Further, the distortion order is selected on held-out images, with the intrinsics and distortion fixed and only the board pose re-estimated, ensuring that an added coefficient is supported by independent observations. Finally, we integrate both steps into an interactive calibration tool that supports full-pipeline data inspection and parameter estimation. Experiments on our own camera data and five public real-world datasets show that image filtering reduces the held-out reprojection error by 25\%, the order selection further by 5\%, achieving the lowest held-out mean among four compared configurations without manual image selection. We will release the code and data to facilitate future research.
Chinese Translation
精确的相机内参标定是机器人感知的基础,其精度取决于所采集图像的质量。然而,现有的基于标定板的方法通常需要操作者手动筛选高质量图像,并指定合适的径向畸变阶数。本文提出了一种全自动的内参标定流程,可从采集的数据中同时确定上述两者。我们采用一种迭代剔除方案,在候选图像集上估计参数,并剔除平均残差超过中位数若干倍的视图。关键在于,该过程在每个候选畸变阶数下独立运行,使得保留的图像集与该阶数的残差尺度相一致。此外,畸变阶数在留出图像上进行选择,此时固定内参与畸变参数,仅重新估计标定板位姿,从而确保新增的系数得到独立观测的支持。最后,我们将这两个步骤集成到一个交互式标定工具中,支持全流程的数据检查与参数估计。在我们自己的相机数据以及五个公开的真实世界数据集上的实验表明,图像筛选将留出集重投影误差降低了25%,阶数选择进一步降低了5%,在无需手动图像筛选的情况下,在四种对比配置中取得了最低的留出集平均误差。我们将公开代码和数据,以促进未来研究。
cs.RO / 25 / 2609.10137

Assembling Two Parts in One Hand

单手装配两个零件
Pei, Liuao, Wu, Tianyue, Zhang, Hui, Luo, Ping, Song, Jie
Abstract
A hallmark of human dexterity is the cooperative use of fingers, where different fingers take on distinct yet coordinated roles to accomplish fine manipu- lation, such as capping a pen with the hand that holds it. We study this finger-level coordination through in-hand assembly: mating two rigid objects within a single dexterous hand, with no second arm and no fixture. We present a reinforcement learning formulation to solve this problem in a unified framework, which is driven by a goal relative pose between the two parts. Finger coordination is shaped by a function-based auxiliary reward and regularized toward a single human reference pose, while domain randomization and a fusion of historical proprioception and object observation confer robustness to occlusion-induced estimation noise. The same recipe solves three different assembly tasks (Bottle, Syringe, and Marker). Trained purely in simulation, the policies transfer zero-shot to hardware with a single camera, demonstrating robustness to state-estimation errors caused by oc- clusion. Our experiments also reveal that in-hand assembly places demands on hand morphology and can serve as a benchmark for modern robotic hand systems. Videos and code are available at https://ltbgbird.github.io/in-hand-assembly-page/.
Chinese Translation
人类灵巧性的一个标志性特征是手指的协同使用,即不同手指承担不同但相互协调的角色,以完成精细操作,例如用手握住笔并将其笔帽盖好。我们通过手内装配(in-hand assembly)来研究这种手指级别的协调:在单个灵巧手中将两个刚性物体装配起来,无需第二条机械臂,也无需固定装置。我们提出了一种强化学习框架,以统一的方式求解该问题,该框架由两个零件之间的目标相对位姿驱动。手指协调通过基于功能的辅助奖励进行塑形,并向单个人类参考姿态正则化;同时,域随机化以及历史本体感知与物体观测的融合,为系统提供了对遮挡引起的估计噪声的鲁棒性。同一套方法成功解决了三种不同的装配任务(Bottle、Syringe 和 Marker)。策略完全在仿真中训练,并零样本迁移到仅配备单个相机的真实硬件上,展示了对遮挡导致的状态估计误差的鲁棒性。我们的实验还表明,手内装配对手部形态提出了要求,可作为现代机器人手系统的基准测试。视频与代码可在 https://ltbgbird.github.io/in-hand-assembly-page/ 获取。
cs.RO / 26 / 2609.10166

Future-Aware Flow Planning for Safe UAV Target Following

面向安全无人机目标跟随的未来感知流规划
Feng, Boning, Zhang, Haoran, Bi, Xiaowen, Zhang, Yanzhen, Shi, Xiaodan
Abstract
UAV target following in cluttered environments is inherently predictive: current-state followers can lag behind turns, choose blocked corridors, or trade tracking for unsafe near-horizon motion. We propose a future-aware flow planning framework for state-informed UAV target following. Predicted target futures guide clean UAV trajectory generation as horizon-aligned residual signals, while risk-scored executable-prefix repair is embedded inside the sampling loop. On fixed ID/OOD receding-horizon benchmarks, the planner improves the intended safety--tracking trade-off rather than dominating every metric: it matches zero measured ID collision rate with the highest ID safe-tracking time, and gives the lowest OOD macro collision rate and final tracking error among the displayed methods, while Future-MPC remains smoother and stronger on some thresholded OOD success metrics under its hand-designed objective. Ablations show that future adaptation improves candidate generation before safety repair, and simulator-facing stress tests probe interface, sensing, and controller-execution effects. These results support horizon-aligned future adaptation and embedded prefix repair as complementary ingredients for safe UAV target following under the tested simulation conditions.
Chinese Translation
复杂环境中的无人机(UAV)目标跟随本质上具有预测性:仅基于当前状态的跟随方法可能在转弯时滞后、选择被阻塞的通道,或以不安全的近程运动为代价换取跟踪效果。我们提出了一种用于状态感知无人机目标跟随的未来感知流规划框架。预测的目标未来状态以视界对齐的残差信号形式引导干净的无人机轨迹生成,同时将基于风险评分的可执行前缀修复嵌入到采样循环内部。在固定的分布内(ID)/分布外(OOD)滚动时域基准测试中,该规划器改善的是安全性与跟踪之间的权衡,而非在每项指标上全面占优:它实现了零实测ID碰撞率并取得最高的ID安全跟踪时间,同时在所展示的方法中给出最低的OOD宏平均碰撞率和最终跟踪误差;而Future-MPC在其手工设计的目标函数下,在部分阈值化的OOD成功指标上仍保持更平滑和更强的表现。消融实验表明,未来适应在安全修复之前改善了候选生成,面向仿真器的压力测试则考察了接口、感知与控制器执行的影响。这些结果支持将视界对齐的未来适应与嵌入式前缀修复作为在所测试仿真条件下实现安全无人机目标跟随的互补要素。
cs.RO / 27 / 2609.10169

Multi-Robot Scanner for Automated Full-Body Dermoscopic Imaging

用于自动化全身皮肤镜成像的多机器人扫描系统
Franchi, Valerio, Garcia, Rafael, Gracias, Nuno, Campos, Ricard, Quintana, Josep, González-Villà, Sandra, Ventura, Mark, Ferrera, Nuria, Lenoir, Clément, Malvehy, Josep
Abstract
This paper outlines the specifications and design approach used to construct a full body imaging scanner capable of capturing skin lesions at a dermatoscopic level using cameras mounted on the end-effectors of four UR10 manipulators. The system possesses a view-planning algorithm capable of appropriately selecting the best camera position to acquire images of moles, a high-level controller to allow the manipulators to work simultaneously and a collision-detector that halts the manipulators when they make contact with an object or a person. We evaluate the system through real-patient full-body scans, comparing acquired images against contact dermoscopy and an existing total-body photography system (Vectra) across clinically relevant lesion features, and quantify true optical resolving power using a USAF 1951 resolution target, yielding a smallest resolvable feature size of 22.1 microns for our scanner compared to 8.8 microns for contact dermoscopy. Results show the scanner consistently outperforms Vectra across most clinically relevant features and achieves comparable performance to contact dermoscopy for the majority of features assessed. By acquiring dermatoscopic-quality images automatically and without contact, and without requiring a separate manual dermoscopic examination, the scanner closes part of the gap between total-body photography and handheld dermoscopy, suggesting potential for future integration into screening workflows.
Chinese Translation
本文概述了构建全身成像扫描仪的规格要求与设计方法,该扫描仪利用安装于四台UR10机械臂末端执行器上的相机,能够以皮肤镜级别拍摄皮肤病变图像。该系统具备一个视角规划算法,可适当选择最佳相机位置以获取痣的图像;一个高层控制器,使多个机械臂能够同时工作;以及一个碰撞检测器,在机械臂接触物体或人员时使其停止运动。我们通过真实患者的全身扫描对该系统进行评估,将采集的图像与接触式皮肤镜检查以及现有的全身摄影系统(Vectra)在临床相关病变特征上进行比较,并使用USAF 1951分辨率测试板量化其真实光学分辨能力,结果显示本扫描仪的最小可分辨特征尺寸为22.1微米,而接触式皮肤镜为8.8微米。结果表明,该扫描仪在大多数临床相关特征上持续优于Vectra,并且在所评估的大部分特征上达到与接触式皮肤镜相当的性能。通过自动、无接触地获取皮肤镜质量的图像,且无需单独进行手动皮肤镜检查,该扫描仪缩小了全身摄影与手持式皮肤镜之间的部分差距,表明其未来有望集成到筛查工作流程中。
cs.RO / 28 / 2609.10215

Adaptive Shared Control with Online Bounded-Rational Human Behavior Estimation

基于在线有界理性行为人行为估计的自适应共享控制
Trejo, Henry Ascencio, Pieters, Roel, Alcan, Gokhan
Abstract
This work considers adaptive shared human-robot control for nonlinear control-affine systems, where the assumption of a fully rational human is relaxed and the robot adapts its assistance to observed boundedly rational human behavior. We use a level-k bounded-rationality model of the two-player game to construct a finite bank of candidate human and robot policies through alternating best-response computations, with the associated value functions and policies approximated using adaptive dynamic programming. During the shared-control interaction, state-transition residuals compare the measured system evolution with the trajectories predicted by the candidate human policies. The residuals are accumulated using a forgetting factor and mapped to a probabilistic human-behavior model over the finite candidate bank. Rather than selecting a single candidate or averaging stored robot policies, the robot computes a distribution-aware one-step best response by minimizing an expected cooperative cost over the complete estimated human behavior distribution. For a quadratic terminal-value approximation and Euler state propagation, this response admits a closed-form solution expressed in terms of the expected human input. The proposed methods are evaluated in simulations of a benchmark nonlinear system stabilization task, and of a planar manipulator shared control setup. The reported results show decreasing Kullback-Leibler divergence between the estimated and simulated human behavior distributions, and a lower accumulated running cost for the robot agent over the shared control interaction period, than the maximum-probability and probability-weighted alternative policies baseline.
Chinese Translation
本研究考虑针对非线性控制仿射系统的自适应人机共享控制,其中放宽了完全理性人的假设,机器人根据观测到的有界理性的人类行为来调整其辅助策略。我们采用双人博弈的level-k有界理性模型,通过交替最优响应计算构建一个由候选人类策略和机器人策略组成的有限策略库,并使用自适应动态规划对相应的值函数和策略进行近似。在共享控制交互过程中,状态转移残差将测得的系统演化与候选人类策略预测的轨迹进行比较。残差通过遗忘因子进行累积,并映射到有限候选策略库上的概率化人类行为模型。机器人并非选择单一候选策略或对已存储的机器人策略进行平均,而是通过在完整估计的人类行为分布上最小化期望合作代价,计算一种分布感知的单步最优响应。对于二次型终端值近似和欧拉状态传播的情形,该响应具有以期望人类输入表示的闭式解。所提出的方法在基准非线性系统镇定任务和平面机械臂共享控制设置的仿真中进行了评估。结果表明,估计的人类行为分布与仿真的人类行为分布之间的Kullback-Leibler散度逐渐减小,且在共享控制交互期间,机器人智能体的累积运行代价低于最大概率策略和概率加权策略等基线方法。
cs.RO / 29 / 2609.10230

CougarTail & CUB: A General-Purpose Mast and Central Utility Board for Cylindrical Underwater Enclosures

CougarTail与CUB:面向圆柱形水下密封舱的通用天线桅杆与中央功能板
Washburn, Ben, Smith, Clayton, Gaskin, Eli, Anderson, Brighton, Meyers, Braden, Moon, Brady, Mangelson, Joshua
Abstract
Cylindrical watertight enclosures are widely used across various underwater systems, from unmanned underwater vehicles (UUVs), to remotely operated vehicles (ROVs), to various sensor platforms. However, electronics are typically built on rectangular PCBs arranged in horizontal stacks, which inefficiently occupy the circular cross-section volume that is critical for both payload capacity and buoyancy management. This paper presents CUB (Central Utility Board) and CougarTail, a general-purpose system designed to address this gap. CUB is a circular PCB sized for 4-inch-diameter enclosures that consolidates a Raspberry Pi Compute Module 5 (CM5) and an STM32 microcontroller, while also providing power management features and auxiliary connections. Mounted coaxially, CUB reduces the electronics stack of our CougUV from 200 mm of tube length and 709 g to 25 mm and 156 g, returning that length and mass budget to payload and buoyancy trim. CougarTail is an open-source companion sensor mast that houses a GPS antenna and two dual-band (2.4 and 5 GHz) omnidirectional PCB antennas. Both components are validated through bench testing and integration on a CougUV platform, our small open-sourced torpedo UUVs.
Chinese Translation
圆柱形水密密封舱被广泛应用于各类水下系统中,从无人水下航行器(UUV)、遥控潜水器(ROV)到各种传感器平台。然而,电子设备通常构建于矩形印刷电路板(PCB)并按水平堆叠方式布置,这种布置低效地占用了圆形横截面空间,而该空间对载荷容量和浮力管理均至关重要。本文提出CUB(Central Utility Board,中央功能板)与CougarTail,一个旨在填补这一空白的通用系统。CUB是一块为4英寸直径密封舱设计的圆形PCB,其集成了树莓派计算模块5(CM5)和STM32微控制器,同时还提供电源管理功能和辅助接口。通过同轴安装,CUB将我们的CougUV电子设备堆叠从200毫米管长、709克缩减至25毫米、156克,从而将节省下来的长度与重量预算返还给载荷与浮力配平。CougarTail是一个开源的配套传感器桅杆,内置GPS天线和两副双频(2.4 GHz和5 GHz)全向PCB天线。两个组件均通过台架测试以及在CougUV平台(我们的小型开源鱼雷型UUV)上的集成验证。
cs.RO / 30 / 2609.10243

FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects

FolDeX:面向可变形物体长时程机器人操作的真实物理世界基准
Liu, Chenhuan, Xu, Yi, Wu, Feng, Wang, Hanyang, Kuai, Wenxiao, Ding, Weihao, Wang, Shan, Liu, Yang, Gao, Shuyong, Zhang, Wenqiang
Abstract
Embodied AI, including vision-language-action and world-action models, must operate reliably in the physical world. Yet methods that perform well in simulation can degrade substantially on real robots, especially in long-horizon deformable-object manipulation, where policies must track changing states and execute reliable multi-stage bimanual interactions. Existing real-robot benchmarks mainly focus on short-horizon rigid-object tasks and offer limited coverage of long-horizon deformable manipulation. We introduce FolDeX, a physical-world benchmark built entirely from real-robot data, with garment folding as its primary task. Since real-robot data collection is costly, FolDeX studies how heterogeneous physical experience can be reused efficiently. The benchmark is organized around four research axes: leveraging human intervention and recovery data collected during deployment; transferring data across tasks, including across garment categories and from rigid to deformable-object manipulation; reusing data across scenes with changes in lighting, background, and layout; and transferring data across robotic embodiments. FolDeX provides 2,000+ hours of real-robot data spanning 20+ tasks and 10+ embodiments. We also establish a fair real-robot evaluation platform for externally submitted policies, with standardized tasks, held-out physical objects, controlled initializations, and a unified execution protocol. The platform is publicly accessible at https://ai.midea.com/#/fold-challenge. We hope FolDeX will serve as a unified testbed for heterogeneous real-robot data reuse and reliable long-horizon deformable manipulation.
Chinese Translation
具身智能,包括视觉-语言-动作模型与世界-动作模型,必须在物理世界中可靠运行。然而,在仿真中表现良好的方法在真实机器人上性能可能大幅下降,尤其是在长时程可变形物体操作任务中,策略需要跟踪不断变化的状态并执行可靠的多阶段双手交互。现有的真实机器人基准主要聚焦于短时程刚性物体任务,对长时程可变形操作的覆盖有限。我们提出了FolDeX,一个完全基于真实机器人数据构建的物理世界基准,其核心任务为衣物折叠。由于真实机器人数据采集成本高昂,FolDeX研究了如何高效复用异构的物理经验。该基准围绕四个研究方向组织:利用部署过程中采集的人类干预与恢复数据;跨任务迁移数据,包括跨衣物类别以及从刚性物体操作到可变形物体操作的迁移;跨场景复用数据,涵盖光照、背景和布局的变化;以及跨机器人本体迁移数据。FolDeX提供了超过2000小时的真实机器人数据,涵盖20余项任务和10余种机器人本体。我们还建立了一个公平的真实机器人评估平台,用于评测外部提交的策略,该平台具有标准化的任务、预留的物理物体、受控的初始化以及统一的执行协议。平台可通过 https://ai.midea.com/#/fold-challenge 公开访问。我们希望FolDeX能够成为异构真实机器人数据复用与可靠长时程可变形物体操作的统一测试平台。
cs.RO / 31 / 2609.10283

SwingBot: Learning Whole-Body Brachiation for Humanoid Robots

SwingBot:面向人形机器人的全身臂跃运动学习
Xiong, Yujie, Zhai, Peng, Hou, Taixian, Qian, Quancheng, Liu, Cunwang, Hu, Kangmai, Yang, Long, Dong, Zhiyan, Zhang, Lihua
Abstract
Brachiation enables primates to move across overhead supports when ground paths are blocked, suggesting a complementary locomotion mode for robots operating in cluttered or hazardous environments. Bringing this capabil?ity to high-DoF humanoid robots is difficult because the controller must discover a long-horizon release-swing-capture sequence, coordinate alternating contacts with whole-body momentum, and act without reliable measurements of segment?relative displacement or hook-contact state. We present SwingBot, a learning framework for continuous humanoid brachiation with passive wrist hooks. Swing?Bot makes the task trainable by organizing learning around the structure of brachi?ation: biomimetic keyframes make rare release-swing-capture transitions reach?able during early exploration, and recurrent privileged-state estimation provides compact position and contact latents for deployment. Hardware experiments demonstrate continuous bar traversal and robustness to payload, external distur?bances and different bar spacings, showing that this formulation offers a practical route to whole-body robotic brachiation.
Chinese Translation
臂跃(Brachiation)使灵长类动物能够在地面路径受阻时通过上方支撑物移动,这提示了一种适用于机器人在杂乱或危险环境中作业的互补运动模式。要让高自由度人形机器人具备这一能力十分困难,因为控制器必须发现长时域的“释放-摆荡-捕获”动作序列,协调交替接触与全身动量,并在缺乏可靠的肢体相对位移测量或挂钩接触状态信息的情况下进行动作。我们提出SwingBot,一个基于被动腕部挂钩的连续人形机器人臂跃学习框架。SwingBot通过围绕臂跃运动的结构来组织学习,使该任务变得可训练:仿生关键帧使罕见的“释放-摆荡-捕获”过渡在早期探索阶段即可实现,循环 privileged 状态估计提供紧凑的位置与接触潜变量以用于部署。硬件实验展示了连续的横杆跨越,以及对负载、外部扰动和不同横杆间距的鲁棒性,表明该方案为全身机器人臂跃运动提供了一条切实可行的途径。
cs.RO / 32 / 2609.10286

Learning Terrain-Adaptive Humanoid Locomotion on Granular Terrain

在颗粒地形上学习地形自适应人形机器人运动
Kamohara, Junnosuke, Wu, Feiyang, Zong, Andy Ningan, Goldman, Daniel I., Nakka, Yashwanth, Hutchinson, Seth, Zhao, Ye
Abstract
Humanoid locomotion on granular terrain remains a significant challenge due to its complex foot-terrain interaction dynamics that are difficult to model. Existing approaches either ignore granular contact dynamics or incorporate simplified normal force models with heuristic tangential components. In this work, we present a physics-grounded granular contact model based on three-dimensional resistive force theory (3D RFT) and efficiently simulate granular terrain for reinforcement learning (RL) training. Unlike traditional rigid contact models and simplified granular contact models with ad-hoc heuristics, our contact solver produces physically accurate granular intrusion dynamics without resorting to heuristics. It captures realistic penetration and tangential drag during training, enabling the policy to learn behaviors that transfer reliably to real-world granular terrain where rigid contact models fail. To adapt to varying terrain conditions, we train a terrain-adaptive locomotion controller via teacher-student RL, using a variational autoencoder to encode terrain information into a compact latent representation. Simulation studies using material point method (MPM) with NVIDIA Newton demonstrate that our method generalizes to unseen granular terrains, achieves a significantly higher success rate than baselines, and demonstrates zero-shot terrain identification and adaptation. We further validate our approach through extensive hardware experiments across diverse real-world granular terrains including basalt, dry sand, and beach sand. To the best of our knowledge, this is the first demonstration of agile humanoid locomotion on real-world granular terrain. Project page: https://humanoid-gm-locomotion.github.io/HUMANOID-GM/
Chinese Translation
由于脚-地形交互动力学复杂且难以建模,人形机器人在颗粒地形上的运动仍然是一项重大挑战。现有方法要么忽略颗粒接触动力学,要么采用简化的法向力模型并结合启发式切向分量。在本工作中,我们提出了一种基于三维阻力理论(3D RFT)的物理接地颗粒接触模型,并高效地模拟颗粒地形以用于强化学习(RL)训练。与传统的刚性接触模型以及带有临时启发式规则的简化颗粒接触模型不同,我们的接触求解器无需借助启发式规则即可产生物理精确的颗粒侵入动力学。该模型在训练过程中捕捉了真实的穿透和切向拖曳效应,使策略能够学习到可可靠迁移至真实颗粒地形的行为,而刚性接触模型在此类地形上会失效。为适应不同的地形条件,我们通过师生强化学习(teacher-student RL)训练了一个地形自适应运动控制器,并利用变分自编码器将地形信息编码为紧凑的潜在表示。基于NVIDIA Newton的物质点法(MPM)仿真研究表明,我们的方法能够泛化到未见过的颗粒地形,成功率显著高于基线方法,并展现出零样本的地形识别与适应能力。我们进一步通过在玄武岩、干沙和海滩沙等多种真实颗粒地形上的大量硬件实验验证了该方法。据我们所知,这是首次在真实颗粒地形上演示敏捷的人形机器人运动。项目页面:https://humanoid-gm-locomotion.github.io/HUMANOID-GM/
cs.RO / 33 / 2609.10308

Deformable Object Manipulation under Partial Observability via Real-Time Full-Shape Estimation

基于实时全形状估计的部分可观测性下可变形物体操控
Behnia, Kosar, Kyrki, Ville, Alcan, Gokhan
Abstract
Manipulating deformable objects (DOs) is challenging due to their high-dimensional state space, underactuated dynamics, and partial observability. In this paper, we propose cRVAE, a lightweight conditional recurrent variational autoencoder that estimates the full DO state from only partial corner-node observations during inference. The resulting model is used as the forward model in a receding-horizon optimal control framework for obstacle-aware collaborative DO manipulation. In simulation on rope and fabric, cRVAE estimates the full DO state from the available corner-node measurements alone, matching the accuracy of a parameter-identified XPBD model. At inference it uses no physical parameters as model inputs and performs no online parameter identification. It also runs approximately 350 times faster on the rope and over 1500 times faster on the fabric per forward pass, keeping horizon-based planning within the 100 ms control budget where XPBD exceeds it already at short horizons. Full-shape estimation from corner sensing at in-loop speed is what makes the model deployable on hardware, which we demonstrate on a Unitree Go2 robot.
Chinese Translation
可变形物体(Deformable Objects, DOs)的操控由于其高维状态空间、欠驱动动力学以及部分可观测性而极具挑战性。本文提出cRVAE,一种轻量级的条件循环变分自编码器,可在推理阶段仅凭部分角点观测估计可变形物体的完整状态。所得模型被用作滚动时域最优控制框架中的前向模型,以实现具有避障能力的可变形物体协同操控。在绳索和布料的仿真实验中,cRVAE仅利用可获得的角点测量即可估计出可变形物体的完整状态,其精度可与经参数辨识的XPBD模型相媲美。在推理阶段,该模型不使用任何物理参数作为输入,也无需进行在线参数辨识。此外,其每次前向传播的速度在绳索任务上比XPBD快约350倍,在布料任务上快1500倍以上,使得基于时域的规划能够保持在100毫秒的控制周期内,而XPBD在较短的时域下便已超出该预算。正是借助角点传感实现回路内速度的全形状估计,才使该模型得以部署到硬件上,我们在Unitree Go2机器人上对此进行了验证。
cs.RO / 34 / 2609.10336

Odometer-Agnostic Drift Correction Using OpenStreetMap Lane Geometry

基于OpenStreetMap车道几何的里程计无关漂移校正方法
Caballero, Joaquin, Garcia-Fidalgo, Emilio, Ortiz, Alberto, Ralli, Jarno
Abstract
Despite significant progress in odometry estimation, long-term drift remains a fundamental limitation of incremental pose integration, especially in large-scale or loop-free environments. Existing map-assisted methods can reduce drift, but often depend on dense maps, sensor-specific processing, or complex matching pipelines. We propose a lightweight open-source, odometry-agnostic correction method that aligns short trajectory segments to OpenStreetMap (OSM) lane centerlines. By formulating drift correction as a direct alignment between recent odometry and sparse lane geometry, the method enables efficient online operation without dense priors or expensive preprocessing. Experiments with LiDAR and visual odometry backends demonstrate consistent improvements, with particularly strong gains under severe drift.
Chinese Translation
尽管里程计估计已取得显著进展,但长期漂移仍然是增量式位姿积分的根本性局限,尤其在大规模或无回环的环境中。现有的地图辅助方法虽然能够减小漂移,但通常依赖于稠密地图、特定传感器的处理流程或复杂的匹配管线。我们提出了一种轻量级的开源、里程计无关的校正方法,将短轨迹片段与OpenStreetMap(OSM)车道中心线进行对齐。通过将漂移校正建模为近期里程计与稀疏车道几何之间的直接对齐问题,该方法无需稠密先验或昂贵的预处理,即可实现高效的在线运行。基于LiDAR和视觉里程计后端的实验表明,该方法能够带来一致的改进,在严重漂移情况下提升尤为显著。
cs.RO / 35 / 2609.10339

A Confidence-Aware Multimodal Fusion Framework for Industrial Human-Robot Collaboration

一种面向工业人机协作的置信度感知多模态融合框架
Liu, Xinyu, Dong, Qiqi, Jia, Boya, Zhang, Yi, Lian, Binbin
Abstract
A confidence-aware multimodal fusion framework (CAMF) is proposed to realize reliable human intention prediction for industrial human-robot collaboration. This framework fuses four heterogeneous modalities including object 6D pose, gaze, skeletal motion and IMU-based hand motion. It embeds a confidence-trend-driven dynamic fusion mechanism into BiLSTM to adaptively balance bidirectional temporal features according to real-time modality reliability. A confidence-guided balanced learning strategy combined with a confidence freezing mechanism is further adopted to adjust network gradients dynamically, suppress noise from low-quality modalities and mitigate cross-modal learning bias. A physical platform based on the UR3 collaborative robot is built for experimental validation. Comparative results show that the proposed method reaches an intention recognition accuracy of 91.86% and outperforms existing multimodal fusion approaches in overall performance and stability. It also maintains satisfactory accuracy under low light and partial occlusion interference. In practical assembly tasks, the framework enables proactive and stable human-robot cooperation with strong environmental adaptability.
Chinese Translation
本文提出一种置信度感知多模态融合框架(CAMF),用于实现工业人机协作中可靠的人体意图预测。该框架融合了四种异构模态,包括物体六维位姿、视线、骨骼运动以及基于惯性测量单元(IMU)的手部运动。框架将一种置信度趋势驱动的动态融合机制嵌入双向长短期记忆网络(BiLSTM)中,可根据实时模态可靠性自适应地平衡双向时序特征。进一步采用一种置信度引导的均衡学习策略,结合置信度冻结机制,动态调整网络梯度,抑制低质量模态带来的噪声,并缓解跨模态学习偏差。搭建了基于UR3协作机器人的物理实验平台进行实验验证。对比结果表明,所提方法的意图识别准确率达到91.86%,在整体性能和稳定性方面优于现有多模态融合方法。此外,在弱光照和部分遮挡干扰下仍能保持令人满意的精度。在实际装配任务中,该框架能够实现主动且稳定的人机协作,具有较强的环境适应性。
cs.RO / 36 / 2609.10377

Data-Driven Risk Fields for Safer End-to-End Autonomous Driving

面向更安全端到端自动驾驶的数据驱动风险场
Tian, Yuanxin, Liu, Zhiyuan, Li, Jinhao, Xu, Zhenhua, Yu, Wenhao, Wang, Jianqiang
Abstract
Safety is a fundamental requirement for autonomous driving, yet existing end-to-end driving models still lack explicit risk-aware learning capacities. Existing rule-based risk models provide interpretable safety priors, yet their absolute risk scores depend on handcrafted functions, coefficients, and thresholds. Learning-based risk representations reduce part of this manual design, but their supervision often relies on occupancy-derived labels or heuristic cost values, which may not capture ego-conditioned planning risk. In this paper, we propose DRiF, a data-driven risk-field framework for safer end-to-end autonomous driving. DRiF learns a shared BEV feature with static map segmentation, dynamic risk prediction, and vehicle planning. For dynamic risk learning, DRiF converts rule-based safety priors into pairwise risk labels, and trains the risk field to preserve relative risk ordering instead of regressing handcrafted absolute scores. Experiments on Bench2Drive show that DRiF achieves competitive overall performance, with consistent improvements in driving score, success rate, and collision-related metrics. These results establish relative risk supervision as an effective way to connect explicit safety structure with end-to-end planning. The data and code will be publicly available.
Chinese Translation
安全性是自动驾驶的基本要求,然而现有的端到端驾驶模型仍然缺乏显式的风险感知学习能力。现有的基于规则的风险模型提供了可解释的安全先验,但其绝对风险评分依赖于人工设计的函数、系数和阈值。基于学习的风险表征减少了部分人工设计,但其监督往往依赖占据(occupancy)衍生的标签或启发式代价数值,可能无法捕捉以自车为条件的规划风险。本文提出了DRiF,一个面向更安全端到端自动驾驶的数据驱动风险场框架。DRiF通过静态地图分割、动态风险预测和车辆规划学习一个共享的BEV特征。在动态风险学习方面,DRiF将基于规则的安全先验转化为成对风险标签,并训练风险场以保持相对风险排序,而非回归人工设计的绝对评分。在Bench2Drive上的实验表明,DRiF取得了具有竞争力的整体性能,在驾驶评分、成功率和碰撞相关指标上均有一致的提升。这些结果确立了相对风险监督是将显式安全结构与端到端规划相连接的有效途径。数据和代码将公开发布。
cs.RO / 37 / 2609.10400

A traffic management system for large and heterogeneous vehicles in narrow industrial environments

一种面向狭窄工业环境中大型异构车辆的交通管理系统
Bonetti, Alessandro, Proia, Silvia, Guidetti, Simone, Sabattini, Lorenzo
Abstract
The coordination of Automated Guided Vehicles (AGVs) in high-density industrial environments represents a critical challenge within Logistics 4.0, as traditional traffic management methods often lead to inefficiencies caused by negotiation-based priority assignment. To overcome the resulting limitations, this paper presents an innovative AGV traffic management system based on a Lifelong Multi-Agent Path Finding (L-MAPF) algorithm operating on roadmaps generated with Non-Uniform Rational B-Splines (NURBS) curves. The approach guarantees locally optimal coordination and ensures safe operation of large and heterogeneous AGVs. Building on this concept, the proposed framework integrates a modified version of the Bounded Horizon Conflict Based Search (CBS) technique within a Rolling Horizon Conflict Resolution strategy, utilizing an extended time horizon for each agent to enable effective conflict resolution in corridors identified by a topological map. In contrast to state-of-the-art methods for AGV fleet traffic management, the proposed solution is designed for real-world, non-standardized (i.e., non-grid-like) industrial settings characterized by narrow bidirectional corridors and high-traffic density, where AGVs of various sizes and capabilities operate simultaneously. Key contributions include an anytime conflict resolution strategy with adaptive time horizon regulation, an execution layer for safe and standard-compliant interaction with real AGVs, and an advanced mechanism for deadlock detection and resolution. Experimental results obtained in realistic industrial environments demonstrate higher throughput, with improvements of up to 11% over a conventional rule-based traffic management system, a state-of-the-art industrial method, and a priority-based L-MAPF variant, while maintaining continuous operation and improved efficiency.
Chinese Translation
在物流4.0背景下,高密度工业环境中自动导引车(AGV)的协调是一个关键挑战,因为传统的交通管理方法通常因基于协商的优先级分配而导致效率低下。为克服由此产生的局限性,本文提出了一种创新的AGV交通管理系统,该系统基于终身多智能体路径规划(Lifelong Multi-Agent Path Finding,L-MAPF)算法,并在由非均匀有理B样条(Non-Uniform Rational B-Splines,NURBS)曲线生成的路线图上运行。该方法保证了局部最优的协调,并确保大型异构AGV的安全运行。在此基础上,所提出的框架将边界视界冲突搜索(Bounded Horizon Conflict Based Search,CBS)技术的改进版本集成到滚动视界冲突消解策略中,为每个智能体利用扩展的时间视界,从而实现由拓扑地图识别的走廊中有效的冲突消解。与最先进的AGV车队交通管理方法相比,所提出的解决方案面向真实的、非标准化的(即非网格状)工业场景,其特点是狭窄的双向走廊和高交通密度,且不同尺寸和能力的AGV同时运行。主要贡献包括:一种具有自适应时间视界调节的任意时间冲突消解策略、一个用于与真实AGV进行安全且符合标准交互的执行层,以及一种先进的死锁检测与消解机制。在真实工业环境中获得的实验结果表明,与传统基于规则的交通管理系统、一种最先进的工业方法以及一种基于优先级的L-MAPF变体相比,系统吞吐量更高,提升幅度可达11%,同时保持了连续运行并提高了效率。
cs.RO / 38 / 2609.10405

Frequency-Conditioned Flow Matching for Vision-Language-Action Models

面向视觉-语言-动作模型的频率条件化流匹配方法
Niu, Haochen, Dong, Shengye, Liu, Hao, Lin, Peiwen, Chuang, Wang
Abstract
Robot actions are temporally correlated trajectories whose frequency components encode motion at different scales with highly non-uniform energy distributions. Yet Flow Matching--based vision-language-action (VLA) models typically generate actions in temporal coordinates, without explicitly modeling or systematically leveraging this frequency heterogeneity. We introduce \emph{FreqFM}, a frequency-conditioned Flow Matching framework for VLA models. It raises action frequency from an implicit trajectory property to an explicit conditioning dimension that spans the entire generation pipeline. Concretely, in DCT frequency coordinates, FreqFM constructs a spectrum-matched source distribution, adaptively balances the objective across frequencies, and constrains per-frequency guidance residuals using the corresponding reference transport scales. FreqFM integrates into existing Flow Matching action experts without changing the VLA backbone. Across LIBERO, LIBERO-Plus, and VLA-Arena, FreqFM consistently improves performance, including a 9.3-point gain on LIBERO-Plus, and further demonstrates its effectiveness on six real-robot tasks.
Chinese Translation
机器人动作是具有时间相关性的轨迹,其频率分量在不同尺度上编码运动信息,且能量分布高度非均匀。然而,基于流匹配(Flow Matching)的视觉-语言-动作(VLA)模型通常在时间坐标中生成动作,未能显式建模或系统性地利用这种频率异质性。我们提出了FreqFM,一种面向VLA模型的频率条件化流匹配框架。该方法将动作频率从一种隐式的轨迹属性提升为贯穿整个生成流程的显式条件维度。具体而言,FreqFM在DCT频率坐标中构建谱匹配的源分布,自适应地在各频率间平衡目标函数,并利用相应的参考输运尺度约束各频率上的引导残差。FreqFM可无缝集成到现有的流匹配动作专家中,而无需改变VLA骨干网络。在LIBERO、LIBERO-Plus和VLA-Arena基准上,FreqFM均持续提升性能,其中在LIBERO-Plus上取得9.3分的提升,并在六项真实机器人任务上进一步验证了其有效性。
cs.RO / 39 / 2609.10433

Multi-Agent Reinforcement Learning for Autonomous UAV Exploration in Wildfire Response

面向野火响应中自主无人机探索的多智能体强化学习
Chandra, Caden, Ng, Jerry
Abstract
This study develops a deep reinforcement learning framework for training Unmanned Aerial Vehicle (UAV) agents to navigate and monitor simulated wildfire environments. Results show that agents learn increasingly stable and effective behaviors over time, as demonstrated by converging loss trends, improved reward signals, and more consistent navigation patterns such as fire-boundary tracking. Overall, these findings highlight the potential of deep reinforcement learning (DRL) based UAV systems for autonomous wildfire monitoring and suggest that environmental structure and reward design influence policy effectiveness.
Chinese Translation
本研究开发了一个深度强化学习框架,用于训练无人机(UAV)智能体在模拟的野火环境中进行导航与监测。结果表明,随着时间推移,智能体能够学习到日益稳定和有效的行为,这体现在损失趋势收敛、奖励信号提升以及更一致的导航模式(如火边界跟踪)等方面。总体而言,这些发现凸显了基于深度强化学习(DRL)的无人机系统在自主野火监测方面的潜力,并表明环境结构与奖励设计会影响策略的有效性。
cs.RO / 40 / 2609.10484

Coastal Environment Generation with HoloOcean

基于HoloOcean的海岸环境生成
Austin, Abigail, Moon, Brady, Mangelson, Joshua G.
Abstract
Marine robotic simulation provides a safe and inexpensive method of developing and testing algorithms for unmanned underwater vehicle (UUV) and unmanned surface vessel (USV) autonomy and perception before full field deployment. However, these simulations are often limited by the availability of simulated environments. Current marine robotics simulation suites offer manual ways to edit or create environments, but they require existing data or specialized knowledge of the environment system. To address these issues, we introduce a novel Unreal Engine 5 level generation pipeline that enables automatic creation of coastal environments for HoloOcean. Our pipeline relies on a user-provided overhead image of a coastal scene. The pipeline then uses the image to generate height map data, as well as automatically select assets and place them in the environment.
Chinese Translation
海洋机器人仿真为无人水下航行器(UUV)和无人水面艇(USV)的自主性与感知算法在全面实地部署之前提供了一种安全且低成本的开发与测试方法。然而,这些仿真往往受限于可用仿真环境的数量。现有的海洋机器人仿真平台提供了手动编辑或创建环境的方式,但这需要已有的数据或对环境系统的专业知识。为解决这些问题,我们提出了一种新颖的Unreal Engine 5关卡生成流水线,能够为HoloOcean自动创建海岸环境。该流水线依赖于用户提供的海岸场景俯视图像,并利用该图像生成高度图数据,同时自动选择资源并将其放置到环境中。
cs.RO / 41 / 2609.10506

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

DUET-DINO:面向机器人操作中潜在规划的同步跨视角世界建模
Nilavadi, Nisarga, Römer, Ralf, Reuss, Moritz, Krawez, Michael, Jülg, Tobias, Schoellig, Angela P., Lioutikov, Rudolf, Burgard, Wolfram
Abstract
Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full 7-DoF end-effector control. To address this gap, we introduce DUET-DINO, a simultaneous cross-view latent world model that jointly learns action-conditioned predictions from static side- and wrist-camera observations through cross-view conditioning. By exploiting complementary global scene and gripper-centric information, DUET-DINO enables latent planning over the full 7-DoF action space. Across spatially diverse reach, orientation-intensive angled-reach, and multi-goal grasp-and-lift tasks, DUET-DINO consistently outperforms single-view and independent dual-view baselines, achieving 92% success on reach, 72.5% on angled-reach, and 60.0% on lift tasks. DUET-DINO is trained from scratch on DROID and RoboArena datasets and generalizes robustly under visual distribution shifts. We further show that while V-JEPA 2 wrist-view predictions underestimate visual dynamics induced by fine-grained actions, DINOv3 predictions better capture action-conditioned scene changes, leading to stronger downstream planning. The code and model checkpoints will be open-sourced. Project page: https://utn-air.github.io/DUET-DINO
Chinese Translation
动作条件潜在世界模型通过预测未来的视觉表征,实现零样本目标条件下的机器人规划与控制。然而,这类模型对细粒度空间和旋转动作的预测并不可靠,难以支撑完整的7自由度末端执行器控制。为弥补这一不足,我们提出了DUET-DINO,一种同步跨视角潜在世界模型,它通过跨视角条件化机制,联合学习基于静态侧方相机与腕部相机观测的动作条件预测。借助全局场景信息与以夹爪为中心的信息的互补性,DUET-DINO能够在完整的7自由度动作空间上进行潜在规划。在空间多样的抓取任务、强调姿态的斜向抓取任务以及多目标抓取提升任务上,DUET-DINO持续优于单视角和独立双视角基线方法,在抓取任务上达到92%的成功率,在斜向抓取任务上达到72.5%,在提升任务上达到60.0%。DUET-DINO在DROID和RoboArena数据集上从零开始训练,并在视觉分布偏移下表现出稳健的泛化能力。我们进一步表明,尽管V-JEPA 2的腕部视角预测会低估由细粒度动作引起的视觉动态,DINOv3的预测能更好地捕捉动作条件下的场景变化,从而带来更强的下游规划性能。代码和模型检查点将开源。项目页面:https://utn-air.github.io/DUET-DINO
cs.RO / 42 / 2609.10522

Show-Harness: Just a VLM Agent Can Play Robots

Show-Harness:仅需一个VLM智能体即可操控机器人
Chen, Yanzhe, Bai, Zechen, Cao, Zhijun, Zeng, Wenzheng, Lin, Kevin Qinghong, Lin, Yiqi, Liang, Guoqiang, Ma, Kevin Yuchen, Huang, Qiming, Shou, Mike Zheng
Abstract
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to "play" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.
Chinese Translation
基础视觉语言模型(VLM)展现出对世界的广泛智能,然而如何将这种智能转化为机器人控制仍然具有挑战性。我们提出了Show-Harness,一种具身化框架(Embodied Harness),通过紧凑的语义接口将意图与动作相连接,使VLM能够“操控”机器人。Show-Harness提供了VLM能够自然推理的离散语义动作单元,同时由具身特定的解释器将它们确定性地落地为局部机器人动作,避免让VLM直接负责细粒度的物理决策。通过相同的接口,Show-Harness展示了两方面的可行性:(1)直接解锁闭源前沿VLM用于零样本机器人控制;(2)仅需数个GPU小时的微调,即可适配小规模开源VLM以实现低成本部署。我们进一步开发了GUMI(GUI操作接口,GUI Manipulation Interface),将相同的语义动作空间扩展到基于GUI的示范数据采集,使人类和智能体无需专门的遥操作硬件即可跨具身形态“操控”机器人。大量实验表明,配备Show-Harness的VLM智能体在任务、具身形态和环境之间具有稳健的泛化能力,优于代表性的智能体(Agentic)和视觉-语言-动作(VLA)范式。这些结果表明,正确的接口可以从基础VLM中释放出可观的具身能力,而无需额外的模型容量或昂贵的具身特定预训练。