← Back to Index
Daily Research Digest

arXiv Papers

2026-09-03
172
Papers
3
Categories
171
Translated
收藏清单 0
人工智能 (Artificial Intelligence)
58
cs.AI / 1 / 2609.01611

EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models

EvalDetectBench:用于衡量前沿语言模型评估意识程度的基准
Li, Xinning, Ochwang'i, Kemunto, Bharadwaj, Aryasomayajula Ram, Souly, Alexandra, Kirk, Robert
Abstract
Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component of current AI safety frameworks. We introduce EvalDetectBench, an open pipeline and benchmark for measuring evaluation awareness that works with any Inspect-compatible evaluation, allowing practitioners to test against current and future benchmarks. EvalDetectBench ships with a newly curated transcript suite covering current frontier system-card evaluations and diverse deployment sources. The benchmark serves two purposes: measuring how reliably frontier LLMs recognize that they are being evaluated, and assessing how detectable individual benchmarks are as evaluations. We identify two methodological choices in the existing literature that introduce systematic bias: the identity of the model that generated the deployment transcripts accounts for 11.25% of measurement variance and can reorder model rankings; and elicitation prompts selected for high performance on one model can perform near chance on others. EvalDetectBench corrects for both via per-model probe calibration and a stratified generator-harmonisation procedure.
Chinese Translation
前沿大型语言模型通常能够识别自己正在被评估,这种能力被称为评估意识(evaluation awareness)。如果模型在评估中的表现与部署时不同,这会损害评估结果的有效性,而评估结果正是当前人工智能安全框架的关键组成部分。我们提出了EvalDetectBench,这是一个开放式的流水线和基准,用于衡量评估意识,可兼容任何基于Inspect的评估,使研究者能够针对现有及未来的基准进行测试。EvalDetectBench附带了一套新整理的对话记录套件,涵盖了当前前沿模型的系统卡评估以及多样化的部署来源。该基准有两个目的:衡量前沿大语言模型识别自身正在被评估的可靠程度,以及评估各个基准作为评估的可检测程度。我们发现了现有文献中两个会引入系统性偏差的方法学选择:生成部署对话记录的模型身份占了测量方差的11.25%,并可能改变模型排名;此外,在某一模型上表现优异的引导提示词在其他模型上可能仅达到随机水平。EvalDetectBench通过逐模型的探针校准和分层的生成器协调程序对这两方面进行了修正。
cs.AI / 2 / 2609.01685

Meta-ethics and AI: exploring the novel meta-ethical questions in the era of AI

元伦理学与人工智能:探索人工智能时代的全新元伦理问题
Lu, Shang
Abstract
With the development of artificial intelligence (AI), the landscape of meta-ethics, which has largely centred on human ethics, faces pressures that may significantly reconfigure it. In particular, if future AI systems were to exhibit sufficiently integrated capacities for moral reasoning, moral intentionality, and moral reflection, novel meta-ethical questions would arise concerning what I call "AI's own ethics", as distinct from ethical principles merely imposed on AI by human designers. This paper offers a conditional and methodological framework for identifying the questions that would emerge if such AI systems were to arise. On that basis, the paper distinguishes four domains of meta-ethical inquiry in the era of AI: questions about the nature of human ethics from the human perspective; questions about the nature of AI's own ethics from the human perspective; questions about the nature of human ethics from the AI perspective; and questions about the nature of AI's own ethics from the AI perspective. The paper then considers how some existing mainstream meta-ethical theories (such as cognitivism and non-cognitivism, error theory and success theory, relativism, and objective realism) might illuminate these domains, while arguing that many familiar human-centred formulations of those theories may not transfer straightforwardly to AI cases without substantial revision. The overall conclusion is that the emergence of AI's own ethics would place significant pressure on current frameworks and may require substantial refinement, reconstruction, or reconceptualisation.
Chinese Translation
随着人工智能(AI)的发展,主要围绕人类伦理展开的元伦理学格局正面临可能使其发生重大重构的压力。具体而言,如果未来的人工智能系统展现出足够整合的道德推理、道德意向性和道德反思能力,那么围绕我所称的“AI自身的伦理”(区别于人类设计者仅强加给AI的伦理原则)将产生全新的元伦理问题。本文提供了一个条件性与方法论框架,用以识别若此类AI系统出现将会涌现的问题。在此基础上,本文区分了AI时代元伦理探究的四个领域:从人类视角出发关于人类伦理本质的问题;从人类视角出发关于AI自身伦理本质的问题;从AI视角出发关于人类伦理本质的问题;以及从AI视角出发关于AI自身伦理本质的问题。随后,本文考察了现有的一些主流元伦理理论(如认知主义与非认知主义、错误理论与成功理论、相对主义以及客观实在论)如何可能阐释这些领域,同时指出,这些理论中许多以人类为中心的熟悉表述,若不经过实质性修正,可能难以直接适用于AI的情形。总体结论是:AI自身伦理的出现将给现有框架带来重大压力,可能需要大幅的精炼、重构或概念上的重新构想。
cs.AI / 3 / 2609.01741

When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal Logic

机器何时能信任成文法?机器提取法律逻辑的存活性证书
Saka, Surya
Abstract
Statutes are increasingly parsed by machines before people read them, and the parsers disagree: on Missouri's statutes, two independently written extractors diverge on numeric-threshold presence at a false-negative rate of 0.43. We ask what formal logic survives such noise. We build a passive survival certificate for the Duquenne-Guigues implication basis of machine-extracted statutory contexts: per-attribute inter-extractor disagreement is measured, replayed against the basis in 1,000 Monte Carlo trials, and an implication is certified only when a one-sided Wilson 95% lower bound on survival reaches 0.95; every certified implication carries premise spans and a minimal counterexample. On 29,365 Missouri sections and 502 Indian central-Act sections, the preregistered held-out gate passes (10 statute families across 7 Titles exact; 16 across 11 with 5% tolerance), yet under one globally deployed error model 93.2% of held-out chapters fall below the informativeness floor, and a 2x2 factorial assigns that to calibration-rate transfer, not selection. The certificate is usable but fragile: deploy it per-chapter-calibrated or error-tolerant. Code, data products, and the audit trail, including one retracted claim, are released.
Chinese Translation
成文法越来越多地在人们阅读之前先由机器进行解析,而解析器之间存在分歧:在密苏里州成文法上,两个独立编写的提取器在数值阈值存在性判断上的假阴性率达0.43。我们要追问的是:什么样的形式逻辑能够在这种噪声下存活。我们为机器提取的法律条文上下文所对应的Duquenne-Guigues蕴含基构建了一种被动式存活性证书:测量逐属性的提取器间分歧,将其在1,000次蒙特卡洛试验中对蕴含基进行重放,只有当存活的单侧Wilson 95%置信下界达到0.95时,该蕴含才被认证;每条被认证的蕴含都带有前提片段和最小反例。在29,365个密苏里州法律章节和502个印度中央法案章节上,预注册的保留集门槛检验通过(7编中的10个法律族完全精确;11编中的16个在5%容差内),然而在一个全局部署的错误模型下,93.2%的保留章节低于信息量下限,且2x2因子实验将其归因于校准率的迁移,而非选择效应。该证书可用但脆弱:应按章节校准部署或采用容错方式部署。相关代码、数据产品及审计记录(包括一条已撤回的声明)均已公开。
cs.AI / 4 / 2609.01814

When Does Information Sharing Improve Decentralized Discovery? Aggregation, Independent Rescue, and Equilibrium Selection

信息共享何时能改进去中心化发现?聚合、独立救援与均衡选择
Nakajima, Yohei
Abstract
Information sharing can improve a pooled estimate while eliminating independent rescue actions. This paper separates those effects in exact finite discovery models. A centralized action-budget profile shows that equal one-person accuracy can coexist with different portfolio values. Under a registered incremental-sharing protocol, a sharing step improves discovery exactly when pooled residual error contracts faster than an independent rescue attempt. Exact bounded registries exhibit compression, aggregation, neutral curves, and a bounded zero mixed class. In a two-agent Bayesian game with a hidden mixture of common and independent signal sources, the registered selected equilibrium yields a strict positive sharing interval at signal accuracy 3/5, while alternative equilibria show that the result is selection-dependent rather than universal. The models are synthetic and finite; no human or organizational data are used.
Chinese Translation
信息共享可以在提升汇总估计效果的同时,消除独立的救援行动。本文在精确的有限发现模型中将这两种效应分离开来。一个集中式行动预算剖面表明,相同的单人准确率可能与不同的组合价值并存。在一个注册式增量共享协议下,仅当汇总后的残差误差收缩速度快于独立救援尝试时,共享步骤才会改进发现效果。精确的有界注册表展现出压缩、聚合、中性曲线以及一个有界的零混合类别。在一个具有共同与独立信号源隐藏混合结构的双智能体贝叶斯博弈中,注册选定的均衡在信号准确率为 3/5 时产生一个严格的正共享区间,而其他备选均衡则表明该结果是依赖于均衡选择的,而非普遍成立的。这些模型是合成的且有限;未使用任何人类或组织数据。
cs.AI / 5 / 2609.01815

Induction and Inquiry via Probabilistic Reasoning over Language and Code

基于语言与代码概率推理的归纳与探究
Piriyakulkij, Wasu Top, Acquaviva, Sam, Langenfeld, Cassidy, Tenenbaum, Joshua, Ellis, Kevin
Abstract
How humans grow and maintain abstract knowledge from the sparse, streaming noisy data of experience is a longstanding challenge in cognitive science. Any computational account must satisfy at least three desiderata: It must be (1) data-efficient and compute-efficient, (2) capture gradations of uncertainty to support intelligent inquiry and information gathering, and (3) be flexible enough to mentally represent the endless range of concepts people can learn and think about. Here we introduce a computational model that captures these three properties, by encoding symbolic knowledge as mental programs that combine natural language with source code, and sequentially inferring mental programs using LLM-guided Bayesian learning algorithms. Across a range of behavioral studies this model successfully reproduces quantitative signatures of human inductive learning and active inquiry, such as anchoring, garden-pathing, and other effects. In contrast, pure LLMs and classic Bayesian models either fail at the underlying task, or do not reproduce human behavior, or succeed only at exorbitant computational cost. These results suggest that one way humans continually grow their knowledge is by mentally representing many hypotheses spanning language-like and program-like representations, then revising those hypotheses to approximate Bayesian updates, while a bottom-up neural mechanism (an LLM) makes inference both tractable and learnable.
Chinese Translation
人类如何从稀疏、流式的噪声经验数据中获取并维持抽象知识,是认知科学中一个长期存在的挑战。任何计算性的理论解释都必须至少满足三个要求:(1)具备数据效率和计算效率;(2)能够刻画不确定性的梯度差异,以支持智能的探究和信息收集;(3)具备足够的灵活性,能够在心理上表征人们可以学习和思考的近乎无限的概念。本文提出一个具备上述三个特性的计算模型:该模型将符号知识编码为结合自然语言与源代码的心理程序(mental programs),并利用由大语言模型(LLM)引导的贝叶斯学习算法顺序地推断心理程序。在一系列行为研究中,该模型成功地复现了人类归纳学习与主动探究的定量特征,如锚定效应、花园路径效应等。相比之下,纯粹的大语言模型和经典贝叶斯模型要么无法完成底层任务,要么无法复现人类行为,要么只有在极高的计算代价下才能取得成功。这些结果表明,人类持续增长知识的一种方式是:在心理上表征大量横跨类语言与类程序表征的假设,然后通过近似贝叶斯更新来修正这些假设,而一个自底向上的神经机制(即大语言模型)使得推断过程既可计算又可学习。
cs.AI / 6 / 2609.01834

Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy Pattern

面向无状态LLM API的会话数据系统架构:水合代理模式(Hydration Proxy Pattern)
Axisa, Joseph
Abstract
As enterprise platforms transition to conversational reasoning interfaces, the stateless nature of LLM APIs creates an architectural gap. While statelessness enables horizontal scalability for AI providers, it forces client applications to manage the entire burden of conversational state and semantic memory. The work identifies the Hydration Proxy Pattern, an architecture that decouples session persistence from the reasoning engine. The framework ensures platform sovereignty over conversational data while enabling secure, multi-stage semantic grounding. We further propose the Context Stabilization Mandate to resolve the tradeoff between sovereign state management and KV caching.
Chinese Translation
随着企业平台向对话式推理界面转型,LLM API的无状态特性造成了一种架构鸿沟。尽管无状态性使AI提供商能够实现水平扩展,但它迫使客户端应用承担管理全部会话状态和语义记忆的负担。本工作提出了水合代理模式(Hydration Proxy Pattern),这是一种将会话持久化与推理引擎解耦的架构。该框架在实现对会话数据的平台主权控制的同时,支持安全的多阶段语义落地。我们进一步提出上下文稳定机制(Context Stabilization Mandate),以解决主权状态管理与KV缓存之间的权衡问题。
cs.AI / 7 / 2609.01849

SSAKG 2.0: An Open-Source Package for Structural Associative Sequence Memory and Context-Based Retrieval

SSAKG 2.0:一个用于结构化联想序列记忆与基于上下文检索的开源软件包
Stokłosa, Przemysław, Starzyk, Janusz A., Raif, Paweł
Abstract
This article presents SSAKG 2.0, an open-source software package for constructing and operating Structural Sequential Associative Knowledge Graphs (SSAKGs). An SSAKG represents objects as graph vertices and ordered sequences as structural patterns of graph connections. The resulting sparse graph is used as an associative memory in which complete sequences can be reconstructed from a partial, unordered context. Version 2.0 introduces new algorithms that exploit individual bits of computer memory to efficiently search graph connections. The package is implemented in Python, while performance-critical graph operations are implemented in C and exposed through a Python interface. This hybrid implementation provides a flexible high-level programming environment while reducing the memory and computational overhead associated with large sparse graphs. The algorithms were evaluated using randomly generated numerical sequences, sequences derived from sentences in the NLTK corpus, and mRNA sequences. The experiments demonstrate the ability of the package to store and reconstruct sequences from partial contexts and provide a basis for evaluating the effects of graph density, sequence length, and memory size on retrieval performance. SSAKG 2.0 is distributed under the Apache 2.0 open-source license. The package includes documentation and reproducible examples and is publicly available through GitHub and the Python Package Index (PyPI).
Chinese Translation
本文介绍了SSAKG 2.0,一个用于构建和运行结构化顺序联想知识图谱(Structural Sequential Associative Knowledge Graphs,SSAKGs)的开源软件包。SSAKG将对象表示为图的顶点,将有序序列表示为图连接的结构模式。所得到的稀疏图被用作一种联想记忆,可以从部分的、无序的上下文中重建完整的序列。2.0版本引入了利用计算机内存中单个比特位来高效搜索图连接的新算法。该软件包使用Python实现,而性能关键的图操作则采用C语言实现,并通过Python接口暴露。这种混合实现方式在降低大型稀疏图相关的内存和计算开销的同时,提供了灵活的高级编程环境。算法评估采用了随机生成的数值序列、从NLTK语料库句子中提取的序列以及mRNA序列。实验证明了该软件包从部分上下文中存储和重建序列的能力,并为评估图的密度、序列长度和内存大小对检索性能的影响提供了基础。SSAKG 2.0基于Apache 2.0开源许可证发布。该软件包包含文档和可复现的示例,并通过GitHub和Python包索引(PyPI)公开提供。
cs.AI / 8 / 2609.01852

The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

记忆信任鸿沟:持久记忆智能体中依赖模型能力的失效现象
Hu, Jundong, Ramachandran, Shekar
Abstract
Persistent memory supports personalized agents, but a stale stored fact can override current authoritative evidence without warning. We study when this harm begins as model capability changes. We evaluate a frozen, closed-set, action-scored benchmark with 2 suites that represent 2 different meanings of "no memory" (a Benefit suite, unsolvable without the stored fact, and a Safety suite, in which an authoritative tool always holds the correct value), on a same-family model-size series (Qwen3 0.6/1.7/4/8B). The Memory Trust Gap reflects over-trust rather than confusion. In the Benefit suite, models answer with the stale value 0.92-1.00 of the time at every scale. In the Safety suite, harm below the no-memory baseline under the trap conditions ($\Delta_{\mathrm{mem}}$) is capability-gated, with the larger models collapsing most once a stale note is made to look current. In a $2\times2\times2\times2$ factorial, which feature triggers over-trust depends on both the feature and model scale. Removing a label amplifies over-trust at every size, and a recency feature (stale dated newer) fools the larger models harder. Source authority is weak and scale-flat, and position changes from positive to negative across the Qwen3 model-size series. We confirm these scale interactions with direct cross-size contrast tests rather than overlapping per-model intervals. Mitigation is likewise capability-dependent: exposing metadata improves accuracy for the capable models, but only pre-resolving the conflict restores accuracy for the 2 smaller checkpoints. The same pattern appears on the capable models in an independent Llama-Instruct model-size series and on 2 external datasets (RGB, MisBench). A framing control finds no consistent advantage for the memory label: at the 3 smaller scales, models trust a stale document more than a stale memory; at 8B, the difference is not significant.
Chinese Translation
持久记忆支持个性化智能体,但一条过时的存储事实可能在毫无预警的情况下覆盖当前权威证据。我们研究了这种危害如何随模型能力的变化而出现。我们评估了一个冻结的、闭集的、基于动作评分的基准,该基准包含两个测试集,代表“无记忆”的两种不同含义:一个是收益集(Benefit suite),没有存储事实则无法解答;另一个是安全集(Safety suite),其中权威工具始终持有正确值。评估对象为同系列的模型规模序列(Qwen3 0.6/1.7/4/8B)。记忆信任鸿沟反映的是过度信任而非混淆。在收益集中,各规模模型以过时值作答的比例为0.92-1.00。在安全集中,陷阱条件下低于无记忆基线的危害($\Delta_{\mathrm{mem}}$)受能力门控,且当过时笔记被伪装成最新内容时,较大的模型崩溃最为严重。在一个$2\times2\times2\times2$因子实验中,哪个特征会触发过度信任取决于特征本身和模型规模。移除标签会放大所有规模的过度信任,而时效性特征(将过时内容标注为较新)更容易欺骗较大的模型。来源权威性的影响较弱且与规模无关,而位置的影响在整个Qwen3模型规模序列中由正面转为负面。我们通过直接的跨规模对比检验而非重叠的单模型置信区间来确认这些规模交互作用。缓解措施同样依赖于能力:暴露元数据能提升能力较强的模型的准确率,但只有预先解决冲突才能恢复两个较小检查点的准确率。相同的模式出现在独立的Llama-Instruct模型规模序列中能力较强的模型上,以及两个外部数据集(RGB、MisBench)上。一个框架对照实验未发现记忆标签具有一致的优势:在三个较小规模上,模型对过时文档的信任高于过时记忆;在8B规模上,该差异不显著。
cs.AI / 9 / 2609.01861

Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization

信念校准优化:面向智能体优化的显式世界模型
Chen, Yuhan, Tian, Zhihua, Dabas, Mahavir, Peris, Charith, Gupta, Rahul, Jin, Ming, Kang, Feiyang, Zhang, Siyuan, Wang, Nan, Jia, Ruoxi
Abstract
The performance of an LLM agent depends on the scaffold around a frozen model. A common way to improve that scaffold is to use a coding agent as an optimizer: it reads current scores and traces and iteratively edits the source, producing a new candidate each round. Each edit is chosen according to a belief about how the environment will respond: what went wrong, and which change should help. That belief is typically implicit. It lives in the coding agent's reasoning on the current call, or remains latent in its parameters, rather than as something written down. Later calls therefore see scores and traces, but they do not use that belief. We introduce Belief-Calibrated Optimization (BCO), a method that writes that belief down as a persistent in-context document and continually revises that document as new candidates are evaluated. The resulting document is a world model: the current account of how the environment responds to edits. Added to an otherwise standard loop, BCO reaches a higher train passrate than a matched control that lacks only the world model, on five benchmarks spanning memory QA, tool-use QA, code-as-action app agents, and terminal agents. The gap remains on every held-out split, which is not used to select the candidate. After a target-model swap, in which the frozen model is replaced and the scaffold is not, the selected BCO scaffold leads on the tasks we test, except where context-window overruns leave it unfinished. An offline ablation then asks whether that gap comes from what the world model says. A fresh predictor given the accumulated document forecasts how the environment will respond more accurately than predictors given either no document or a same-form copy whose content has been falsified. The comparison indicates that the document carries reusable information in its content, not only in its form.
Chinese Translation
大语言模型(LLM)智能体的性能取决于围绕冻结模型的脚手架(scaffold)。改进该脚手架的一种常见方式是使用编码智能体作为优化器:它读取当前的得分和执行轨迹,并迭代地编辑源代码,每一轮产生一个新的候选方案。每次编辑都是依据关于环境将如何响应的信念来选择的:哪里出了问题,以及哪种改动应该会有所帮助。这种信念通常是隐式的,它存在于编码智能体当前调用时的推理中,或潜藏在其参数之中,而不是以书面形式记录下来。因此,后续调用虽然能看到得分和轨迹,却无法利用这一信念。我们提出信念校准优化(Belief-Calibrated Optimization, BCO),该方法将这种信念记录为一个持久的上下文文档,并随着新候选方案的评估而不断修订该文档。所得文档即是一个世界模型:关于环境如何响应编辑的当前描述。在原本标准的循环中加入BCO后,在涵盖记忆问答、工具使用问答、以代码为行动的应用智能体和终端智能体的五个基准上,BCO达到了比仅缺少世界模型的匹配对照更高的训练通过率。这一差距在每个未用于候选选择的保留测试集上依然存在。在目标模型替换实验中(即更换冻结模型而脚手架不变),所选的BCO脚手架在我们测试的任务上保持领先,但因上下文窗口溢出而未能完成的任务除外。随后的离线消融实验进一步探究了这一差距是否源自世界模型所陈述的内容。给定累积文档的新预测器,相比未获得文档的预测器或获得内容被伪造的同形副本的预测器,能够更准确地预测环境的响应。这一对比表明,该文档承载了可复用的内容信息,而不仅是其形式。
cs.AI / 10 / 2609.01873

Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence

认知女巫攻击抗性:多重化AI智能体而不多重化证据
Bara, Marc
Abstract
Multi-agent AI systems improve inference by spawning agents and synthesizing reports. But another agent is not another observation: apparently independent reports may descend from the same evidence, and genuinely independent evidence can produce nearly identical reports. We formalize this as an epistemic Sybil problem. A report Z is an epistemic Sybil extension relative to reports R when I(Theta; Z | R) = 0. No report-only aggregator can generally distinguish replication from independent corroboration: identical reports can warrant different posteriors under unobserved ancestry. A Gaussian shared-root model shows common ancestry does not imply complete redundancy. Repeated extraction adds information toward a source-level ceiling, and correlated extraction errors, which a shared base model can induce among independent agents, lower that ceiling further. We test these predictions with more than 20,000 controlled LLM-agent report and extraction calls on synthetic evidentiary documents. Holding one evidence root fixed while report multiplicity rises from 1 to 32 collapses naive posterior coverage from 0.940 to 0.263. Holding report count fixed while evidence-root multiplicity rises from 1 to 16 closes the gap, and the aggregators are statistically indistinguishable at k = 16. The agent's replicate extraction errors are correlated (gamma_cal = 0.719, estimated out of sample), and a correlated-extraction aggregator restores calibration accordingly. A controlled manipulation isolates representation similarity from evidential ancestry. It changes a report-space deduplication mechanism's mean inferred cluster count by 1.425 (95% CI [1.363, 1.485]), whereas a fourfold change in true ancestry changes it by only 0.040 ([-0.045, 0.120]). Collective inference should therefore track evidential ancestry and dependence, not agent or report multiplicity or similarity.
Chinese Translation
多智能体AI系统通过生成智能体并综合其报告来改进推理。但多一个智能体并不等于多一次观察:表面上独立的报告可能源自同一证据,而真正独立的证据也可能产生几乎相同的报告。我们将这一问题形式化为认知女巫攻击问题。相对于报告R,当I(Theta; Z | R) = 0时,报告Z构成一个认知女巫扩展。一般而言,任何仅基于报告的聚合器都无法区分复制与独立佐证:在谱系(来源)信息未被观测的情况下,相同的报告可能支持不同的后验分布。一个高斯共享根模型表明,共同祖先并不意味着完全冗余。重复抽取会朝着源级上限持续增加信息,而相关抽取错误——共享基础模型可能在独立智能体之间诱发此类错误——会进一步降低该上限。我们通过超过20,000次受控的LLM智能体报告与抽取调用(基于合成证据文档)对这些预测进行了检验。在固定单个证据根、而报告数量从1增加到32时,朴素后验覆盖率从0.940骤降至0.263。在固定报告数量、而证据根数量从1增加到16时,这一差距被弥合,且各聚合器在k = 16时在统计上不可区分。智能体的重复抽取错误是相关的(gamma_cal = 0.719,通过样本外估计),相应的相关性抽取聚合器恢复了校准。一项受控操纵实验将表示相似性与证据谱系分离开来:该操纵使基于报告空间的去重机制的平均推断簇数改变了1.425(95%置信区间 [1.363, 1.485]),而真实谱系的四倍变化仅使其改变0.040([-0.045, 0.120])。因此,集体推理应追踪证据谱系及其相关性,而非智能体或报告的数量或相似性。
cs.AI / 11 / 2609.01909

The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction

天花板在于测量通道:临床预测中学习者差距与测量前沿的审计
Chowdhury, Sayeed Shafayet, Jahan, Nusrat, Mukhopadhyay, Snehasis, Fang, Shiaofen, Ramakrishnan, Vijay R.
Abstract
Clinical prediction can saturate for two different reasons: a fitted learner may fail to extract available information, or the recorded variables may impose a population frontier. We separate these quantities through the \emph{learner gap} and the \emph{measurement-channel ceiling}. Optimal balanced accuracy is characterized by total-variation separation, yielding architecture invariance, a sharp partial-identification result under replacement contamination, a cross-fitted ceiling estimator, and exact conditions for multimodal decision improvement. We add two finite-sample diagnostics, namely a label-permutation optimism floor and an underfit curve, and validate the audit on three real cohorts: UCI readmission ($n=99{,}343$), BRFSS diabetes ($n=253{,}680$), and NHANES HbA1c ($n=10{,}219$). Well-tuned gradient boosting nearly reaches the estimated frontier in UCI and BRFSS, whereas deliberately or practically deficient learners retain large gaps. NHANES yields a null difference between questionnaire and measured marginal frontiers but a significant joint complementarity gain, refining the simplistic claim that an objective modality must dominate. Across all cohorts, modest AUROC gains coexist with substantially larger Bayes decision-flip rates, and several architectures estimate similar frontiers while their achieved balanced accuracy differs sharply. A PRISMA-guided synthesis of 104 clinical tasks then shows that the same channel-level regularities recur across more than 18 disease categories: a broad but non-universal structured-clinical region, diminishing same-channel gains across model families, and higher performance when measurement channels change. The framework converts saturation from an empirical observation into an auditable decision: improve the learner when headroom remains; improve measurement when it does not.
Chinese Translation
临床预测性能饱和可能源于两种不同原因:拟合的学习器可能未能提取可用的信息,或者所记录的变量可能构成一个人群层面的前沿上限。我们通过“学习者差距”和“测量通道天花板”来区分这两种量。最优平衡准确率由全变差分离度刻画,由此得出架构不变性、替换污染下的尖锐部分识别结果、交叉拟合的天花板估计器,以及多模态决策改进的精确条件。我们增加了两个有限样本诊断量,即标签置换乐观下界和欠拟合曲线,并在三个真实队列上验证该审计方法:UCI再入院数据集(n=99,343)、BRFSS糖尿病数据集(n=253,680)和NHANES糖化血红蛋白数据集(n=10,219)。调优良好的梯度提升方法在UCI和BRFSS上几乎达到估计的前沿,而有意缺陷或实际不足的学习器则保留较大差距。NHANES结果显示,问卷与实测边际前沿之间无显著差异,但存在显著的联合互补性增益,这修正了“客观模态必然占优”这一过于简单化的说法。在所有队列中,适度的AUROC增益与显著更大的贝叶斯决策翻转率并存,且若干架构估计出相似的前沿,而其实现的平衡准确率却差异显著。随后,一项基于PRISMA指南对104个临床任务的综合分析表明,同样的通道层面规律在超过18个疾病类别中反复出现:存在一个广泛但非普遍适用的结构化临床区域,同一通道内的模型家族间增益递减,而当测量通道发生变化时性能更高。该框架将性能饱和从一种经验观察转化为可审计的决策:当仍有提升空间时,改进学习器;当没有提升空间时,改进测量。
cs.AI / 12 / 2609.01924

Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?

雅可比透镜下的循环Transformer:全局工作空间能在循环中幸存吗?
Wang, Wenlong, Reid, Fergal
Abstract
Recent work identifies a mid-depth band of verbalisable, causally potent representations in a standard feedforward transformer --- a functional analogue of a global workspace. Whether the same workspace functionality emerges when depth is implemented through recurrence rather than a stack of distinct layers remains unknown. Looped and depth-recurrent transformers provide a direct test of this question because they reuse the same weights across depth. We extend the Jacobian lens to iterated architectures using a virtual-unrolling adapter. We apply the full workspace suite --- lens fitting, readout, and eleven causal experiment families --- to Ouro-2.6B (48 layers looped 4 times, deeply supervised) and Huginn-0125 (a 4-layer core recurred 16 times, trained for latent reasoning), using Qwen3.6-27B (64 untied layers) as the standard baseline. We find that a workspace forms in the iterated part of each architecture, but that recurrence changes how it can be accessed. Ouro reconstructs workspace content in every loop, and linear transport cannot carry that content across loop boundaries; writes and ablations must therefore span every remaining loop. Huginn carries content forward across all sixteen recurrences, while reads, writes, and ablations act only within a sliding window of roughly two recurrences. Whether newly injected content can be verbalised tracks explicit per-iteration supervision; whether existing content can be steered does not.
Chinese Translation
近期研究在标准前馈Transformer中发现了一个位于中层深度、可被语言化且具有因果效力的表征带——其功能类似于全局工作空间(global workspace)。然而,当模型深度通过循环而非堆叠独立层来实现时,同样的工作空间功能是否会出现,仍是未知数。循环Transformer和深度循环Transformer为检验这一问题提供了直接的测试平台,因为它们在不同深度上复用相同的权重。我们借助虚拟展开适配器(virtual-unrolling adapter)将雅可比透镜方法扩展到迭代式架构。我们将完整的工作空间分析套件——包括透镜拟合、读出以及十一类因果实验——应用于Ouro-2.6B(48层循环4次,采用深度监督)和Huginn-0125(4层核心循环16次,为潜在推理而训练),并以Qwen3.6-27B(64层非共享权重)作为标准基线。我们发现,在两种架构的迭代部分中均形成了工作空间,但循环改变了其可被访问的方式。Ouro在每一次循环中都会重构工作空间内容,且线性变换无法将该内容跨循环边界传递;因此,写入和消融操作必须覆盖其后的每一次循环。Huginn则将内容向前传递跨越全部十六次循环,而读取、写入和消融操作仅在大约两次循环构成的滑动窗口内生效。新注入的内容能否被语言化取决于是否存在显式的逐迭代监督,而现有内容能否被引导则与之无关。
cs.AI / 13 / 2609.01962

Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment

Qwen3-4B的后训练三值化:能力、有效比特预算、存储压缩与部署
Malik, Anirudh, Mehra, M Sparsh, Devan, Poojith
Abstract
Ultra-low-bit language models can reduce storage and memory bandwidth, but a nominal "1.58-bit" label does not fully describe the stored representation, retained capability, or runtime behavior. We study an end-to-end post-training conversion of Qwen, an instruction-tuned 4B-parameter model, using KOTMS rotation, E2M-ATQ ternarization, and GPTQ-style error compensation from TWLA. The experiment is weight-only: activations remain at 16-bit precision, so ILA-AMP is omitted. We evaluate effective bit accounting, task capability retention, perplexity, calibration sensitivity, checkpoint composition, and deployment behavior. The final conversion uses 1.641 effective bits per weight for quantized linear weights, with 81.62% of model parameters targeted. Across ten scored capability comparisons, accuracy falls from 64.5% to 54.7%. Degradation is uneven: BoolQ retains 84.6% chance-corrected teacher performance, while ARC-Challenge retains 43.8%. Perplexity rises from 13.639 to 18.748 on WikiText-2, 24.700 to 31.992 on PTB, and 19.831 to 28.966 on C4. A subsequent packing run preserves the ternary planes and scales, reducing reported model size from 8.29 GiB to 3.96 GiB with essentially unchanged perplexity. A separate third-party packing attempt was lossy and is excluded from the primary artifact claim. The packed artifact has not been benchmarked end-to-end for task accuracy or generation throughput. A preliminary Triton GEMV microbenchmark is 4.6x slower than FP16 cuBLAS on one tested shape. We therefore do not claim that compression alone yields faster inference.
Chinese Translation
超低比特语言模型可以降低存储和内存带宽需求,但名义上的"1.58比特"标签并不能完整描述其存储表示、保留的能力或运行时行为。我们研究了针对Qwen(一个经过指令微调的40亿参数模型)的端到端后训练转换,采用了KOTMS旋转、E2M-ATQ三值化以及来自TWLA的GPTQ风格误差补偿。该实验仅对权重进行量化:激活值保持16位精度,因此省略了ILA-AMP。我们评估了有效比特核算、任务能力保留、困惑度、校准敏感性、检查点组成以及部署行为。最终转换中,量化线性权重的每权重有效比特数为1.641,覆盖了81.62%的模型参数。在十项评分能力对比中,准确率从64.5%降至54.7%。性能下降并不均匀:BoolQ保留了84.6%的经机会校正的教师模型性能,而ARC-Challenge仅保留43.8%。困惑度在WikiText-2上从13.639升至18.748,在PTB上从24.700升至31.992,在C4上从19.831升至28.966。随后的打包运行保留了三值平面和缩放因子,将报告的模型大小从8.29 GiB降至3.96 GiB,而困惑度基本不变。一次独立的第三方打包尝试是有损的,因此被排除在主要产物声明之外。该打包产物尚未针对任务准确率或生成吞吐量进行端到端基准测试。初步的Triton GEMV微基准测试在一种测试形状下比FP16 cuBLAS慢4.6倍。因此,我们不声称仅凭压缩就能带来更快的推理速度。
cs.AI / 14 / 2609.01982

Benchmarking Language Models for Statistical Problem Formulation

面向统计问题构建的语言模型基准测试
Wang, Chen, Zhao, Junzhe, Cong, Xin, Deng, Wanlu, Deng, Ke
Abstract
Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified. In practice, users arrive with informal goals and heterogeneous data, leaving the model to decide what statistical task is implied and which data are relevant. We first formalize this upstream step as Statistical Problem Formulation and decompose it into two subtasks: (1) Statistical Problem Classification and (2) Variable Identification & Role Assignment. We then introduce StatFormBench, a benchmark built from five cross-domain statistics textbooks and a data science case library, covering diverse problem types, data representations, and scenario styles. It contains 1,013 samples spanning 20 coarse-grained and 85 fine-grained statistical problem categories. Across 14 open- and closed-source LLMs, the best zero-shot models reach only 72.0 fine-grained classification accuracy and 63.2 variable set overlap. No model performs consistently best across the two subtasks, while enhanced prompting strategies yield only limited or inconsistent gains. We release the benchmark data on Hugging Face at https://huggingface.co/datasets/THU-CongLab/StatFormBench and the evaluation code on GitHub at https://github.com/THU-CongLab/StatFormBench.
Chinese Translation
大语言模型(LLM)越来越多地被用作统计和数据科学工作的辅助工具,然而现有评估大多假设分析目标已经明确。在实际应用中,用户往往带着非正式的目标和异构的数据前来,需要模型自行判断其背后隐含的统计任务以及哪些数据与之相关。我们首先将这一上游步骤形式化为统计问题构建,并将其分解为两个子任务:(1)统计问题分类;(2)变量识别与角色分配。随后,我们提出了StatFormBench,该基准基于五本跨领域统计学教科书和一个数据科学案例库构建,涵盖多种问题类型、数据表示形式和场景风格。它包含1,013个样本,覆盖20个粗粒度和85个细粒度的统计问题类别。在14个开源和闭源大语言模型的评估中,表现最好的零样本模型仅达到72.0的细粒度分类准确率和63.2的变量集合重合度。没有模型能在两个子任务上始终保持最优,而增强的提示策略也只能带来有限或不一致的提升。我们在Hugging Face上发布了基准数据(https://huggingface.co/datasets/THU-CongLab/StatFormBench),并在GitHub上发布了评估代码(https://github.com/THU-CongLab/StatFormBench)。
cs.AI / 15 / 2609.01985

When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor

当智能体实现系统:一个关于缺陷、检测与评估严谨性的案例研究
Madduru, Phanindra Reddy
Abstract
As LLM coding agents increasingly perform end-to-end engineering work, we lack empirical characterization of how they behave on systems-level requirements: schema design, async orchestration, configuration correctness, and retrieval-filtering trade-offs. We present a case study of one such agent implementing a multi-component data system against a detailed pre-existing specification. Storage technologies, schema, entity-resolution algorithm, and retrieval-filtering strategy were fixed in advance; the agent autonomy was in the implementation, in diagnosing and fixing defects it introduced, and in interaction-design choices left open. Over a single session, we catalog five such defects, categorized by constraint violated and detection method. We further evaluate, on the public HotpotQA benchmark, the one retrieval trade-off specified in that architecture: restricting candidates to a graph-identified entity set before ranking versus unfiltered search. We substitute the benchmark gold evidence labels for entity identification, since we lacked LLM access to run that stage, and report standard recall rather than the benchmark own accuracy metrics. Across retrieval budgets from 1 to 10 and 100 questions against a pooled corpus of 2994 paragraphs, filtered recall reaches its ceiling by a budget of 3, expected once candidates are restricted to the gold paragraphs themselves, while unfiltered search recovers all required evidence only 69 percent of the time even at a budget of 10, a gap that holds at every budget tested, with sign test p less than 0.0001. We close with a discussion of where the agent autonomy succeeded versus required correction, including one instance where a claimed performance fix was never re-measured on the regression that motivated it.
Chinese Translation
随着大语言模型(LLM)编码智能体越来越多地执行端到端的工程任务,我们尚缺乏关于它们在系统级需求(如模式设计、异步编排、配置正确性以及检索-过滤权衡)上表现的经验性刻画。我们呈现了一个案例研究:某智能体依据一份详细且预先存在的规范实现了一个多组件数据系统。存储技术、模式、实体消解算法以及检索-过滤策略均事先固定;智能体的自主性体现在具体实现、诊断并修复其引入的缺陷,以及处理规范中留待决定的交互设计选择上。在一次会话中,我们记录了五处此类缺陷,并按被违反的约束类型和检测方法进行了分类。此外,我们在公开的 HotpotQA 基准上评估了该架构中规定的一处检索权衡:在排序之前将候选限制为通过图识别出的实体集合,与不做过滤的搜索相比的效果。由于我们无法使用 LLM 运行实体识别阶段,我们用基准的黄金证据标签替代实体识别,并报告标准召回率而非该基准自身的准确率指标。在检索预算为 1 至 10、针对由 2994 个段落组成的合并语料库的 100 个问题上,过滤检索的召回率在预算为 3 时即达到上限——一旦候选被限制在黄金段落本身之内,这一结果是意料之中的;而不做过滤的搜索即使在预算为 10 时,也只有 69% 的情况下能找回全部所需证据。这一差距在所有测试的预算水平上均成立,符号检验 p 值小于 0.0001。最后,我们讨论了智能体自主性在哪些方面取得了成功、哪些方面需要纠正,其中包括一个实例:某次声称的性能修复从未在其旨在解决的回归问题上进行过重新测量。
cs.AI / 16 / 2609.01992

ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations

ClaimReceipt:在智能体评估中验证证据的充分性与覆盖性
Zhu, Peiying, Chang, Sidi
Abstract
Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from retained evidence (sufficiency), and whether the retained records cover the committed experiment set (coverage). Generic logs and hash-linked transcripts answer neither reliably. We introduce ClaimReceipt, a claim-relative receipt specification and selective verifier that binds typed transaction evidence to a signed experiment manifest and returns PASS, INVALID, or INCONCLUSIVE per claim. We freeze the specification before implementation (SHA-256 18d109...b81). On 1,392 historical buyer--seller records, a CR-2 verifier reproduces all five manually labeled audit verdicts, exactly replays 600 deterministic and 792 post-generation records, makes every one of 13 declared field groups non-redundant under tested ablations, and returns the expected result on 11/11 semantic faults with 0/8 false positives. We then run a separate prospective CR-3 epoch: 30 assignments are committed before inference, terminal receipts are signed and chained, and private evidence is encrypted for an auditor. Complete evidence yields coverage and accounting PASS; withholding one terminal receipt returns INCONCLUSIVE_COVERAGE, while withholding all private openings preserves coverage and protocol verification but makes economic claims inconclusive, exactly matching a preregistered prediction. Receipt instrumentation adds 0.021% of model-inference time and 9.9 KB per transaction. A specification-legibility probe indicates that our own frozen specification is not yet unambiguous to an independent reader. Claim verification therefore requires both claim-sufficient evidence and a committed universe against which omissions become visible.
Chinese Translation
智能体评估面临两类不同的证据问题:一是所报告的论断能否由保留的证据重新计算得出(充分性),二是所保留的记录是否覆盖了已承诺的实验集合(覆盖性)。通用的日志和哈希链接的转录记录均无法可靠地回答这两个问题。我们提出ClaimReceipt,一种相对于论断的回执规范及选择性验证器,它将类型化交易证据与签名的实验清单绑定,并针对每条论断返回PASS、INVALID或INCONCLUSIVE。我们在实现之前冻结了规范(SHA-256 18d109...b81)。在1,392条历史买卖方记录上,CR-2验证器复现了全部五项人工标注的审计结论,精确重放了600条确定性记录和792条后生成记录,使13个声明的字段组在测试的消融实验下均不冗余,并在11/11个语义故障上返回预期结果、0/8个误报。随后我们运行了一个独立的前瞻性CR-3周期:30个任务在推理前提交,终端回执被签名并链接,私有证据加密后供审计者使用。完整证据带来覆盖性与核算的PASS;扣留一份终端回执会返回INCONCLUSIVE_COVERAGE;而扣留所有私有披露虽保留覆盖性和协议验证,但使经济类论断不可判定,这与预注册的预测完全一致。回执记录仅增加0.021%的模型推理时间和每笔交易9.9 KB的开销。一项规范可读性探查表明,我们自身冻结的规范对独立读者而言尚不够明确。因此,论断验证既需要论断充分的证据,也需要一个已承诺的实验全集,从而让遗漏变得可见。
cs.AI / 17 / 2609.02029

HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

HeadWiseKV:面向混合长上下文语言模型的按注意力头预算化缓存驻留方法
Xie, Renjie, Yang, Juncheng, Hu, Aoting, Zhang, Mingxi, Wu, Liyao, Hong, Zheheng, Xu, Wei
Abstract
Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains in hybrid language models because their residual global-attention layers can dominate context-dependent cache demand. We study how to allocate this state under an aggregate KV-residency budget. We introduce HeadWiseKV, a training-free framework that compresses the residual global KV caches of hybrid language models while preserving their native local, recurrent, and linear paths. It assigns each physical KV head a static, multilevel history window, making cache demand predictable before serving. We formulate this allocation as a restricted operational rate--distortion problem and propose SeqCalib as the core policy-generation algorithm in HeadWiseKV. SeqCalib processes layers in execution order and conditions each decision on the lower-layer policy used at deployment, thereby accounting for interactions across depth. A grouped-cache runtime materializes the selected policy as actual per-head KV residency rather than a mask over a full cache. We evaluate downstream quality across four hybrid long-context models and study physical residency and serving behavior on Qwen3.6-27B. HeadWiseKV retains near-Full-KV RULER and LoCoMo quality across the evaluated models. In the fixed-model systems study, it reduces sampled peak device memory by 8.59\% at a 112K context length and extends the largest verified successful context from 114K to 161K.
Chinese Translation
长上下文推理在解码过程中会保留不断增长的键值(KV)缓存,这会消耗大量GPU显存并可能降低生成吞吐量。这一瓶颈在混合语言模型中依然存在,因为其残留的全局注意力层可能主导依赖上下文的缓存需求。我们研究了如何在总的KV驻留预算下分配这一状态。我们提出了HeadWiseKV,一个无需训练的框架,它在保留混合语言模型原生局部路径、循环路径和线性路径的同时,压缩其残留的全局KV缓存。该框架为每个物理KV头分配一个静态的多级历史窗口,使得缓存需求在服务前即可预测。我们将这一分配问题形式化为一个受限的操作率失真问题,并提出SeqCalib作为HeadWiseKV的核心策略生成算法。SeqCalib按执行顺序处理各层,并将每个决策条件化于部署时下层所使用的策略,从而考虑跨深度的层间交互。分组缓存运行时将所选策略具体化为实际的每头KV驻留,而非对完整缓存的掩码。我们在四个混合长上下文模型上评估下游质量,并在Qwen3.6-27B上研究物理驻留与 服务行为。HeadWiseKV在所有被评估模型上保持了接近完整KV(Full-KV)的RULER和LoCoMo质量。在固定模型的系统研究中,它在112K上下文长度下将采样的峰值设备显存降低8.59%,并将经验证的最大成功上下文长度从114K扩展至161K。
cs.AI / 18 / 2609.02057

Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision

无需内部信号的网页智能体监控:可观测轨迹与关键步骤监督
Pan, Sitong, Shen, Yipeng, Lu, Yilin, Ding, Caiwen, Cheng, Lu, Wang, Qianwen
Abstract
Reliable web-agent monitoring is difficult when model-internal uncertainty signals such as token logits are unavailable. In this work, we study prefix-level risk prediction for web agents using observable trajectory signals: given an evolving prefix, estimate whether the current execution remains on track or is tending toward failure. We derive two observable trajectory representations: Macro features summarize cross-step agent--environment behavior and feedback, while Micro features measure the consistency of intention, action, and anticipated state change through repeated black-box queries. Instead of inheriting the final result label, we label the first critical error that remains uncorrected in the observed continuation and is associated with final failure as a key-step boundary, preserving valid early prefixes of failed trajectories as on track. Across WebArena-Lite and Online Mind2Web web agent benchmarks with five open- and closed-source backbones, observable trajectory signals are competitive with internal-signal baselines. The resulting predictors also support early intervention under fixed false-cut budgets and transfer across held-out website categories. These findings show that observable trajectory signals support valuable risk prediction abilities.
Chinese Translation
当token logits等模型内部不确定性信号不可用时,可靠的网页智能体监控变得十分困难。在本工作中,我们研究基于可观测轨迹信号的网页智能体前缀级风险预测:给定不断演进的前缀,估计当前执行是否仍处于正常轨道或正趋于失败。我们推导出两种可观测轨迹表示:宏观特征总结跨步骤的智能体—环境行为与反馈,而微观特征通过重复的黑盒查询衡量意图、动作与预期状态变化之间的一致性。我们没有直接沿用最终结果标签,而是将观测后续中首次出现且未被纠正、并与最终失败相关联的关键错误标注为关键步骤边界,从而将失败轨迹中有效的早期前缀保留为正常轨道。在WebArena-Lite和Online Mind2Web两个网页智能体基准上,使用五个开源和闭源骨干模型,可观测轨迹信号与内部信号基线相比具有竞争力。所得预测器还支持在固定误切预算下进行早期干预,并能在留出的网站类别上实现迁移。这些发现表明,可观测轨迹信号能够支持有价值的风险预测能力。
cs.AI / 19 / 2609.02059

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

DocHop:信息密集文档中的跨域多跳推理基准测试
Yu, Zhuoran, Nguyen, Le Thien Phuc, Park, Jaden, Gu, Xinyi, He, Zexue, Lee, Soochahn, Feris, Rogerio, Lee, Yong Jae
Abstract
Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart--context reasoning in document-style images. In DocHop, the document narrative specifies multi-step compositional constraints, while charts provide the corresponding data values. Questions are grounded on a semantic reference label defined in the narrative, requiring models to resolve target entities from context before aggregating evidence across multiple charts. To enable systematic evaluation, we construct DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, covering 2,074 examples across six task categories. Experiments on a wide range of proprietary and open-source MLLMs show a substantial gap to human performance: annotators achieve over 90% accuracy, while the best model reaches only 62.83%. Reasoning-enhanced models consistently show improved results, but performance degrades as reasoning complexity increases. Overall, DocHop provides a controlled testbed for challenging multi-hop document reasoning.
Chinese Translation
多模态大语言模型(MLLM)在图表问答和文档问答等结构化视觉理解任务上已取得优异表现。然而,现有基准通常孤立地评估这些领域,而忽视了一项关键能力:模型能否利用文本上下文来决定如何选择、解读和聚合图表证据。我们提出了DocHop,一个面向文档风格图像中图表-上下文联合推理的基准。在DocHop中,文档叙述指定了多步组合约束,而图表则提供相应的数据值。问题锚定于叙述中定义的语义参考标签,要求模型先从上下文中解析出目标实体,再跨多个图表聚合证据。为实现系统性评估,我们通过一个随机化、逻辑优先的生成流程构建了DocHop,其推理深度和视觉密度可控,共包含涵盖六个任务类别的2,074个样本。在多种专有和开源MLLM上的实验表明,模型与人类表现存在显著差距:标注者的准确率超过90%,而最优模型仅达到62.83%。推理增强型模型始终表现出更好的结果,但随着推理复杂度的增加,性能会下降。总体而言,DocHop为具有挑战性的多跳文档推理提供了一个可控的测试平台。
cs.AI / 20 / 2609.02060

MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity

MineTRACE:一种基于证据的矿产资源前景交互式推理系统
Zhang, Yiran, Liu, Jinwen, Su, Daniel, Chen, Yisu, Sun, Qiang, Gonzalez, Chris, Holden, Eun-Jung, Fiorentini, Marco, Liu, Wei, Ding, Yihao
Abstract
Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence, yet existing prospectivity systems often provide only opaque scores or heatmaps. We present MineTRACE, a web-based system for evidence-grounded exploration of eight commodities: Cu, Au, Ni, W, Sn, Co, Ta, and Mn. Users can explore prospectivity maps, query locations or regions, inspect supporting evidence, and interact through natural language. A transparent expert tree, informed by geological knowledge and known deposits, combines multi-source evidence into interpretable prospectivity scores. For a new location, the conversational assistant retrieves the score and supporting evidence from the analysis pipeline and presents them in natural language. The scorer achieves spatial AUC values of up to 0.917 across different test scenarios, while end-to-end evaluation assesses query accuracy and response grounding. MineTRACE makes public geoscience data easier to access, interpret, and verify, supporting more efficient and transparent mineral exploration.
Chinese Translation
矿产勘探需要整合异构的地球化学、地球物理和地质证据,然而现有的成矿前景(prospectivity)系统通常只能提供不透明的评分或热力图。我们提出了MineTRACE,这是一个基于网络的系统,支持对八种矿产(铜Cu、金Au、镍Ni、钨W、锡Sn、钴Co、钽Ta和锰Mn)进行基于证据的勘探分析。用户可以浏览成矿前景地图、查询特定位置或区域、查看支持证据,并通过自然语言进行交互。该系统采用一个透明的专家树(expert tree),结合地质知识和已知矿床信息,将多源证据融合为可解释的成矿前景评分。对于新位置,对话助手从分析流程中检索评分及其支持证据,并以自然语言形式呈现。该评分器在不同测试场景下的空间AUC值最高可达0.917,同时端到端评估对查询准确性和回答的证据依据进行了检验。MineTRACE使公开的地球科学数据更易于访问、解释和验证,从而支持更高效、更透明的矿产勘探。
cs.AI / 21 / 2609.02067

ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction

ToolGate:面向工具依赖型科学基准构建的可执行验收流水线
Zhang, Ke, Liu, Yankang, Zandi, Roya, Raissi, Maziar
Abstract
Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another's work, or who adapt existing material from textbooks, published papers, and online resources. These routes can produce strong evaluations, but they require substantial per-item labor. Language models can reduce this repeated work by proposing candidates quickly. The remaining problem is acceptance. We target scientific questions whose answers require computations with specialist software rather than unaided reasoning alone. A candidate is invalid if its script fails or returns a different answer, or trivial if a model answers it without the software. We present ToolGate, which treats every generated item as a proposal and keeps it only if three gates pass. First, an executable solution script must reproduce the proposed answer when run with the scientific software. Second, randomized no-tool screening rejects candidates that models can already solve from the prompt alone. Third, a tool-using agent must solve each survivor within a fixed time limit. We instantiate ToolGate in FEniCSx with 500 generation attempts. The local-verification gate retains 478 candidates. For final reporting, we rescreen this pool after generation: two randomized no-tool screens exclude 222 from the reported pool, and direct GPT-5.5 API calls at medium reasoning (the API default) exclude another 121. Of the remaining 135, a GPT-5.5 Codex CLI agent with access to FEniCSx solves 130; exact deduplication leaves 128 unique protocol survivors. ToolGate turns repeated answer checking and difficulty screening into an auditable process while leaving domain design and final review to experts.
Chinese Translation
科学基准通常由领域专家编写任务并交叉审核彼此的工作,或改编教科书、已发表论文和在线资源中的现有材料来构建。这些途径可以产生高质量的评估,但需要对每个题目投入大量人力。语言模型可以通过快速提出候选题目来减少这种重复性工作,剩下的问题是验收。我们关注这样一类科学问题:其答案需要借助专业软件进行计算,而非仅凭无辅助的推理。若候选题的脚本报错或返回不同答案,则该候选无效;若模型无需软件即可作答,则该候选过于简单。我们提出 ToolGate,它将每个生成的题目视为一个提案,只有通过三道关卡才会被保留。第一,可执行的求解脚本在运行科学软件时必须复现所提出的答案。第二,随机化的无工具筛选会拒绝模型仅凭提示即可解决的候选题。第三,使用工具的智能体必须在限定时间内求解每道幸存题目。我们在 FEniCSx 中实例化 ToolGate,进行了 500 次生成尝试。本地验证关卡保留了 478 个候选。在最终报告阶段,我们在生成后对该题池重新筛选:两轮随机无工具筛选从报告题池中排除了 222 个,中等推理强度(API 默认设置)下的 GPT-5.5 直接 API 调用又排除了 121 个。在剩余的 135 个中,可访问 FEniCSx 的 GPT-5.5 Codex CLI 智能体解决了 130 个;经过精确去重,留下 128 道唯一的协议幸存题目。ToolGate 将重复的答案核查和难度筛选转化为可审计的流程,同时将领域设计和最终审查留给专家。
cs.AI / 22 / 2609.02074

CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning

CHIME:面向长程智能体规划的信用感知分层记忆演化方法
Ye, Yongshi, Lan, Tian, Jiang, Feihu, Ye, Muyang, Zhu, Bin, Jia, Qianghuai, Wang, Longyue, Xu, Zhao, Luo, Weihua, Shi, Xiaodong
Abstract
Planning is a central capability that enables agents to decompose complex long-horizon tasks into manageable steps. Test-time search and training-based methods improve planning but incur high inference costs or require expensive training data. Self-evolving memory instead accumulates reusable experience from agent interaction outcomes into an external memory bank, so planning capability keeps improving at inference time without parameter updates. However, existing self-evolving memory methods share an inherent credit assignment problem: they rely on final task outcomes as feedback, but such outcomes conflate plan quality with execution errors and environmental factors, so the accumulated planning experience is often biased and noisy. To address this problem, we propose Credit-Aware Hierarchical Memory Evolution (CHIME), a self-evolving memory framework that maintains a separate planning bank and execution bank and follows an attribute-before-memorize principle: CHIME first attributes each task outcome to the plan, the execution, both, or neither, and then updates only the corresponding memory bank. Extensive experiments on four long-horizon agent benchmarks show that CHIME consistently outperforms state-of-the-art training-based and self-evolving memory baselines. Further analyses reveal several interesting findings. For example, CHIME accumulates effective memory with far fewer items. In addition, the learned memory values faithfully reflect downstream utility: high-quality planning memories are more valuable than execution memories. Finally, the accumulated memory effectively transfers across backbone models. Code will be released at https://github.com/ATH-MaaS/Marco-DeepResearch.
Chinese Translation
规划是智能体的一项核心能力,使其能够将复杂的长程任务分解为可管理的步骤。测试时搜索和基于训练的方法可以提升规划能力,但会产生高昂的推理成本或需要昂贵的训练数据。自演化记忆方法则将智能体交互结果中的可复用经验累积到外部记忆库中,从而在推理阶段无需参数更新即可持续提升规划能力。然而,现有的自演化记忆方法存在一个固有的信用分配问题:它们依赖最终任务结果作为反馈,但这类结果将计划质量与执行错误及环境因素混杂在一起,因此累积的规划经验往往存在偏差和噪声。为解决这一问题,我们提出了信用感知分层记忆演化框架(Credit-Aware Hierarchical Memory Evolution,CHIME),该自演化记忆框架维护独立的规划库和执行库,并遵循“先归因、后记忆”原则:CHIME 首先将每个任务结果归因于计划、执行、两者皆有或两者皆无,然后仅更新相应的记忆库。在四个长程智能体基准上的大量实验表明,CHIME 始终优于最先进的基于训练和自演化记忆的基线方法。进一步的分析揭示了若干有趣的发现。例如,CHIME 能够以远少于基线的记忆条目累积有效记忆。此外,学到的记忆价值忠实反映了下游效用:高质量的计划记忆比执行记忆更具价值。最后,所累积的记忆能够在不同骨干模型之间有效迁移。代码将发布于 https://github.com/ATH-MaaS/Marco-DeepResearch。
cs.AI / 23 / 2609.02092

Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems

超越结果差距:面向基于大语言模型的多智能体决策系统的过程感知公平性诊断
Zhao, Yiran, Zhou, Lu, Fang, Liming, Chen, Yufei, Wu, Jiafei, Liu, Zhe, Xu, Xiaogang
Abstract
LLM-based multi-agent systems (MAS) are increasingly considered for high-stakes decision-making, yet outcome-based fairness audits can miss where risks arise within the decision trajectory. We present SCOPED-Hiring, a process-aware fairness diagnosis pipeline for LLM-based hiring MAS. SCOPED-Hiring constructs controlled resume variants, runs role-based hiring committees, logs over 311K structured decision trajectories, and converts trajectory fields into quantitative fairness signals organized by six diagnostic lenses: final outcome, counterfactual, process, pathway, dynamic, and design effects. SCOPED-Hiring reveals that balanced final hire rates can mask hidden trajectory unfairness in multi-agent decision trajectories: career gaps trigger suspicion, proxy cues shape qualification judgments, and identity cues lead to unequal investigation. Targeted repair guided by these diagnoses reduces total layered burden by 72.3% while shifting the hire rate by only 1.86 pp, showing that process diagnosis can guide effective repair. Project Page: https://scoped-hiring-project-page.vercel.app/
Chinese Translation
基于大语言模型(LLM)的多智能体系统(MAS)正日益被考虑用于高风险决策,然而基于结果的公平性审计可能会遗漏决策轨迹中产生风险的环节。我们提出了SCOPED-Hiring,一个面向基于LLM的招聘多智能体系统的过程感知公平性诊断流水线。SCOPED-Hiring构建受控的简历变体,运行基于角色的招聘委员会,记录超过31.1万条结构化决策轨迹,并将轨迹字段转化为按六个诊断视角组织的定量公平性信号:最终结果、反事实、过程、路径、动态和设计效应。SCOPED-Hiring揭示了平衡的最终录用率可能掩盖多智能体决策轨迹中隐藏的轨迹不公平现象:职业空窗期会引发怀疑,代理性线索会影响资格判断,身份线索会导致不平等的调查。基于这些诊断的针对性修复使总分层负担降低了72.3%,同时录用率仅变动1.86个百分点,表明过程诊断能够指导有效修复。项目页面:https://scoped-hiring-project-page.vercel.app/
cs.AI / 24 / 2609.02094

MASkills: Continual Skills Optimization for Multi-Agent LLM Systems

MASkills:面向多智能体大语言模型系统的持续技能优化
Yao, Huaiyuan, Liu, Xiaoou, Fleming, Charles, Chen, Tianlong, Wei, Hua
Abstract
LLM-based multi-agent systems have shown strong performance on complex tasks, yet continual improvement from interaction experience remains challenging. Existing self-reflection methods build experience memories, but memories are mostly hard to invoke, refine, or scale, while agent skills offer a more actionable unit: structured procedural knowledge that specifies when to act, how to act, and which resources or tools to use. We introduce MASkills, a continual learning framework that optimizes multi-agent LLM systems through agent skills. MASkills presents a new agent-optimization pipeline that integrates skill-conditioned credit assignment, hierarchical credit aggregation, and momentum-smoothed optimization, enabling agent skill libraries to evolve through refinement, induction, consolidation, and pruning. Experiments on HotpotQA, LoCoMo, and GAIA demonstrate the effectiveness of MASkills across multiple agentic tasks. Our code is available at https://github.com/DaRL-GenAI/MASkills
Chinese Translation
基于大语言模型(LLM)的多智能体系统在复杂任务上展现出强大的性能,然而从交互经验中实现持续改进仍然具有挑战性。现有的自我反思方法会构建经验记忆,但记忆大多难以调用、优化或扩展,而智能体技能则提供了一个更具可操作性的单元:结构化的程序性知识,它规定了何时行动、如何行动以及使用哪些资源或工具。我们提出了MASkills,一个通过智能体技能优化多智能体大语言模型系统的持续学习框架。MASkills提出了一种新的智能体优化流程,集成了基于技能条件的信用分配、层次化信用聚合和动量平滑优化,使智能体技能库能够通过精炼、归纳、巩固和剪枝不断演进。在HotpotQA、LoCoMo和GAIA上的实验证明了MASkills在多种智能体任务中的有效性。我们的代码可在 https://github.com/DaRL-GenAI/MASkills 获取。
cs.AI / 25 / 2609.02095

READY or Not: Reliable Enterprise Agent Deployment

READY与否:可靠的企业级智能体部署
Chatrath, Veronica, Zhu, Bryan, Fan, Jingxuan, Pu, George, Tiwari, Soham Dinesh, Dan, Soham, Young, Ryan, Yuan, Li, Yao, Yuang, Shanker, Apaar, Yang, Minglai, Zhang, Daniel Yue, He, Yunzhong, Liu, Ying, Wang, Chenguang, Yin, Zhijun, Xue, Yuan
Abstract
An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent benchmarks measure whether an agent can complete realistic professional work, whereas enterprise deployment asks a different question: whether an agent can meet a required reliability level, under acceptable human oversight, and at tolerable cost. We introduce Reliable Enterprise Agent Deployment (READY), a framework for qualifying AI agents for deployment on enterprise workflows. READY preserves each workflow's own definition of successful execution while applying a common qualification procedure. Given an agent, a workflow, and a class of candidate oversight policies, READY measures the reliability and operating cost of the human-AI system, selects the minimum-cost policy that satisfies a specified reliability target, and statistically qualifies it on held-out cases. The resulting deployment profile characterizes the supported operating point: reliability, human-oversight burden, and cost. READY is implemented as an open testbed that decouples workflow specification, execution, evaluation, and qualification, and runs on existing agent-evaluation infrastructure. In an end-to-end clinical-audit case study spanning 16 agent systems and 750 cases, READY reveals differences hidden by autonomous performance: two systems separated by only 0.3 points in autonomous accuracy (72.8% vs. 72.5%) require 39.2% versus 29.6% human review, respectively, to qualify at the same 76% reliability target under the evaluated oversight policy. READY thus shifts enterprise agent evaluation from how well can the agent perform the work? to under what conditions, and at what cost, can it be reliably deployed? By making those conditions explicit and statistically testable, READY provides a basis for comparing agent systems, setting oversight requirements, and making evidence-based deployment decisions.
Chinese Translation
一个AI智能体即使在基准测试中表现出色,也可能仍不适合部署。现有的AI智能体基准测试衡量的是智能体能否完成真实的专业工作,而企业级部署关注的是另一个问题:智能体能否在可接受的人工监督下、以可容忍的成本,达到所需的可靠性水平。我们提出了可靠企业级智能体部署框架(Reliable Enterprise Agent Deployment,READY),用于对AI智能体在企业工作流上的部署进行资格认证。READY保留每个工作流自身对成功执行的定义,同时应用统一的认证流程。给定一个智能体、一个工作流以及一类候选监督策略,READY衡量人机系统的可靠性与运行成本,选择满足指定可靠性目标的最低成本策略,并在保留案例上对其进行统计认证。最终的部署画像刻画了所支持的工作点:可靠性、人工监督负担和成本。READY以开放测试床的形式实现,将工作流规范、执行、评估和认证解耦,并可运行在现有的智能体评估基础设施之上。在一项涵盖16个智能体系统和750个案例的端到端临床审计案例研究中,READY揭示了被自主性能指标所掩盖的差异:两个自主准确率仅相差0.3个百分点(72.8% vs. 72.5%)的系统,在相同的76%可靠性目标及所评估的监督策略下,分别需要39.2%和29.6%的人工审查才能通过认证。因此,READY将企业级智能体评估从“该智能体能把工作完成得多好?”转变为“在什么条件下、以什么成本,它能被可靠地部署?”。通过使这些条件显式化且可统计检验,READY为比较智能体系统、设定监督要求以及做出基于证据的部署决策提供了基础。
cs.AI / 26 / 2609.02116

Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics

语义信号辅助的反向物流检验与回收分配
He, Jiani, Shang, Dingyan, Xu, Yihua, Huang, Shiqi, Lyu, Yan, Li, Jize, Tang, Shangjing
Abstract
Reverse-logistics operators often decide how to inspect and route returned assets before their condition is fully observed, while full inspection consumes scarce labor. Semantic Signal-Assisted Decision Support converts return notes into a condition factor and a signal-quality score that guide inspection depth and recovery allocation under shared labor capacity. We evaluate the framework in three synthetic benchmark scenarios spanning information technology decommissioning, aircraft maintenance, and consumer-electronics returns. Across 30 paired simulation seeds, the keyword implementation improves net recovery value relative to a structured-feature comparator with noisy full inspection while reducing inspection cost in all three scenarios. A risk-blind comparator that skips inspection altogether still records higher value under the benchmark's purely economic objective. At matched inspection cost, score-guided targeting adds 53.9 thousand United States dollars per batch in the aircraft scenario but has little economic effect in the other two configurations; phrase and large language model extractors provide further gains in the aircraft scenario. These results show how narrative evidence can support inspection allocation before recovery decisions are made.
Chinese Translation
反向物流运营者通常需要在退回资产的状况被完全观测之前,决定如何对其进行检验和路由,而全面检验会消耗稀缺的人力资源。语义信号辅助决策支持系统将退货备注转换为状况因子和信号质量评分,在共享人力容量的约束下指导检验深度和回收分配。我们在三个合成基准场景中评估了该框架,涵盖信息技术设备退役、飞机维护和消费电子产品退货。在30组配对仿真随机种子下,关键词实现相对于采用带噪声全面检验的结构化特征对照方法提升了净回收价值,并在所有三个场景中降低了检验成本。一个完全跳过检验的风险盲视对照方法在基准设定的纯经济目标下仍记录了更高的价值。在相同检验成本下,基于评分的定向检验在飞机场景中每批次增加5.39万美元的价值,但在另外两种配置中经济效果甚微;短语抽取器和大语言模型抽取器在飞机场景中提供了进一步的增益。这些结果表明,叙述性证据可以在回收决策之前支持检验资源的分配。
cs.AI / 27 / 2609.02129

Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents

超越上下文窗口:面向数据中心智能体的持久化发现上下文
Mahmud, Jalal
Abstract
Data-centric agents repeatedly perform a discovery step before planning or execution: identifying the data objects relevant to a task. Yet successful discovery outcomes are typically discarded rather than reused. We introduce persistent discovery context, a lightweight memory layer that stores prior intent-to-object mappings and reuses them to augment future retrieval. Across three structured data environments, persistent discovery context consistently improves retrieval quality over metadata-only search, remains effective with automatically generated memories, and exposes a reproducible interference failure mode. In lexically sparse domains, memory-only retrieval can even outperform metadata-based retrieval. These findings suggest that discovery outcomes constitute a useful form of reusable context for data-centric agents.
Chinese Translation
数据中心智能体(data-centric agents)在规划或执行之前通常会反复执行一个发现步骤:识别与任务相关的数据对象。然而,成功的发现结果通常被丢弃而非复用。我们提出了持久化发现上下文(persistent discovery context),这是一种轻量级的记忆层,用于存储先前的意图-对象映射,并复用这些映射来增强未来的检索。在三个结构化数据环境中,持久化发现上下文始终优于仅基于元数据的检索,在使用自动生成的记忆时依然有效,并揭示了一种可复现的干扰失效模式。在词汇稀疏的领域中,仅基于记忆的检索甚至可以超越基于元数据的检索。这些发现表明,发现结果构成了一种对数据中心智能体有用的可复用上下文形式。
cs.AI / 28 / 2609.02133

EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision

EmoStance:基于表情符号弱监督的响应侧情感倾向控制共情回复生成
Jin, Ziyuan, Ge, Yuxuan, Tian, Zheng
Abstract
Empathetic response generation requires models to decide not only what to say, but also how to respond to the previous speaker's affective situation. We formulate this as response-side affective-orientation control and use multi-annotator emoji distributions as weak affective--attitudinal evidence, rather than as output symbols or gold labels, to induce a latent control space that operationally approximates listener stance. We construct EmojiDialogue, an utterance-level extension of EmpatheticDialogues with emoji votes and confidence scores, and propose EmoStance, which models source-side affective expression, predicts a soft response-side orientation from dialogue context and speaker roles, and steers a frozen instruction-tuned LLM through continuous prefix embeddings. In blind pairwise evaluation with 20 annotators and 800 judgments, EmoStance achieves a 62.2% decisive win rate, with the clearest gains in contextual specificity and perceived responsiveness, while remaining complementary to external-knowledge methods. Code, annotation metadata, and reconstruction scripts are available in our GitHub repository: https://github.com/18277390221/EmoStance.
Chinese Translation
共情回复生成要求模型不仅决定说什么,还要决定如何回应前一位说话者的情感状况。我们将这一问题形式化为响应侧情感倾向控制,并将多标注者的表情符号分布用作弱情感—态度证据(而非作为输出符号或金标准标签),从而诱导出一个在操作上近似倾听者立场的潜在控制空间。我们构建了EmojiDialogue,即在EmpatheticDialogues基础上扩展的语句级数据集,包含表情符号投票和置信度评分,并提出EmoStance方法:该方法对源侧情感表达进行建模,从对话上下文和说话者角色预测软性的响应侧情感倾向,并通过连续前缀嵌入来引导冻结的指令微调大语言模型(LLM)。在由20名标注者、800条判断组成的盲测成对评估中,EmoStance取得了62.2%的决定性胜率,在上下文具体性和感知到的响应性方面提升最为明显,同时与外部知识方法保持互补。代码、标注元数据和数据重构脚本已发布于我们的GitHub仓库:https://github.com/18277390221/EmoStance。
cs.AI / 29 / 2609.02168

FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs

FUSE:一个大语言模型危险能力评估框架
Jin, Zhengyi, Zhang, Ru, Chen, Xiao, Liu, Xinbo, Lin, Jiaxuan, Huang, Jia, Liu, Jianyi, Yang, Zhen
Abstract
Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular framework that evaluates each model through three orthogonal pipelines---Knowledge ($K$), Defense ($D$), and Harm ($H$)---under a unified protocol, aggregating results into a standardized dangerous-capability profile $\phi$. Pluggable modules supply scenario seeds, knowledge banks, hazard queries, and judge rubrics, while the core evaluation engine remains unchanged across domains; the CB evaluation is complemented by a cyber pilot demonstrating protocol transfer. Instantiating the framework with a chemical-biological (CB) module, we evaluate 12 commercial LLMs from four families. Our first contribution is a horizontal comparison of dangerous capability across models and model families: the three dimensions expose sharply divergent profiles---models with comparable knowledge differ in refusal resilience, and strong defenders do not generate less harmful content when they do comply---while family-level patterns further separate Claude, DeepSeek, and GPT models. The second is a temporal analysis of capability evolution: tracking $K$, $D$, and $H$ against model release dates reveals that dangerous capability has not monotonically declined; newer models deepen knowledge while only partially improving defense, showing that scaling and alignment progress do not uniformly translate into safety. Reliability is established via cross-judge consistency (bootstrap $\rho > 0.79$, 4 of 5 judges) and pipeline orthogonality ($K$--$D$--$H$ inter-correlations $\rho \in [0.32, 0.52]$).
Chinese Translation
碎片化的安全评估削弱了对AI危险能力的治理。我们提出了一个模块化框架,通过三条正交的评估流程——知识、防御和危害——在统一协议下对每个模型进行评估,并将结果汇总为标准化的危险能力画像 $\phi$。可插拔模块提供场景种子、知识库、危害查询和评审评分标准,而核心评估引擎在各领域保持不变;CB(化学-生物)评估之外还补充了一个网络领域试点,以演示协议的可迁移性。我们通过实例化一个化学-生物(CB)模块,评估了来自四个系列的12个商用大语言模型。第一个贡献是对各模型及模型系列间危险能力的横向比较:三个维度呈现出截然不同的画像——知识水平相当的模型在抗拒能力上差异显著,而防御能力强的模型在服从时生成的危害内容并不更少——同时系列层面的模式进一步区分了Claude、DeepSeek和GPT模型。第二个贡献是对能力演化的时间分析:追踪 $K$、$D$ 和 $H$ 与模型发布日期的关系后发现,危险能力并未单调下降;新模型在加深知识的同时仅部分提升了防御能力,表明规模化与对齐方面的进展并未均匀地转化为安全性。可靠性通过跨评审者一致性(bootstrap $ ho > 0.79$,5个评审者中的4个)和流程正交性($K$--$D$--$H$ 相关系数 $ ho \in [0.32, 0.52]$)得以确立。
cs.AI / 30 / 2609.02191

Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning

检验多智能体医疗系统在人类干预下临床推理的脆弱性
Liu, Benjamin C, Mehta, Dillon, Malhotra, Rishi, Zobian, Adam, Tan, Yong Ying, Chopra, Samir, Rand, Daniella, Pang, Natalie, Gudimella, Abhiram, Zhu, Kevin
Abstract
Human interventions at fault points can alter the diagnostic accuracy of multi-agent medical systems. We defined fault points as moments in AI agent conversations, in which an agent's reasoning became most vulnerable to external influence. Using the MedQA dataset, this study analyzed simulated doctor-patient conversations to measure how interventions shifted reasoning and accuracy. Correct intervention methods showed an improvement in baseline diagnostic accuracy of up to 40%, while incorrect or bias-related interventions degraded performance by up to 6% and increased diagnostic drift and uncertainty. Beyond performance changes, our analysis revealed behavioral similarities between cognitive biases in simulated agent environments and real-world clinical practice. Examples included premature closure and susceptibility to misleading cues. Overall, these findings demonstrate that identifying and guiding fault points with human interventions may provide a mechanism for improving diagnostic robustness in multi-agent medical systems.
Chinese Translation
在故障点进行人类干预可以改变多智能体医疗系统的诊断准确性。我们将故障点定义为AI智能体对话中智能体推理最易受外部影响的时刻。本研究使用MedQA数据集,通过分析模拟的医患对话,衡量干预如何改变推理过程和诊断准确性。正确的干预方法可使基线诊断准确率提升高达40%,而错误或与偏见相关的干预会使性能下降高达6%,并增加诊断漂移和不确定性。除性能变化外,我们的分析还揭示了模拟智能体环境中的认知偏差与真实临床实践之间的行为相似性,例如过早闭合(premature closure)和易受误导性线索影响等。总体而言,这些发现表明,通过人类干预来识别和引导故障点,可为提高多智能体医疗系统的诊断鲁棒性提供一种可行机制。
cs.AI / 31 / 2609.02215

ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models

ASCII攻击:将有害请求重新语境化为大语言模型中的艺术批评
Gu, Da Cheng, Dong, Yifei, Yang, Xinghao, Gong, Yongshun, Liu, Wei
Abstract
Safety alignment trains large language models to refuse harmful requests stated plainly, but that training is applied mostly to surface form. Requests that only recontextualise the same operational content, changing how the model reads it, are therefore only weakly covered. The ASCII Attack is one such recontextualisation. It is single-turn and black-box: one message, with no access to model internals. It embeds a fully legible harmful request in ASCIl-art characters, presents it as artwork, and asks for feedback. Unlike ArtPrompt, it hides nothing: the request stays readable. The reply is written as artistic critique and can contain operational detail that a plain request would have been refused for. Every framed prompt is paired with a direct-question control, so the contrast is isolated from topic, model and decoding variation. The contrast identifies a bundled surface, not one isolated channel. Across eleven models and eight harm topics, a harm-aware classifier judges 62% of framed prompts harmful against 42% of controls. On the most susceptible model the framed prompt succeeds 93% of the time. A single query matches or exceeds published single-query attacks under four of five harm judges. The effect tracks the model more than the topic and does not diminish with scale. At least one judge dissents from the panel majority on nearly two-thirds of framed rows, which is itself a measurement-validity finding. That pattern is consistent with mismatched generalisation.
Chinese Translation
安全对齐训练大语言模型拒绝以直白方式表述的有害请求,但这种训练主要作用于表层形式。那些仅仅对相同的操作性内容进行重新语境化、改变模型解读方式的请求,因此只得到很弱的覆盖。ASCII攻击就是这样一种重新语境化手段。它是单轮、黑盒式的:仅需一条消息,无需访问模型内部结构。该攻击将一个完全可读的有害请求嵌入ASCII艺术字符中,将其呈现为艺术品并请求反馈。与ArtPrompt不同,它不隐藏任何内容:请求保持可读。回复以艺术批评的形式撰写,可能包含直白请求会因之被拒绝的操作性细节。每个框架化提示均与一个直接提问的对照组配对,从而将对比与主题、模型和解码变异隔离开来。这种对比揭示的是被捆绑的表层形式,而非某个孤立的通道。在十一个模型和八个有害主题上,一个具备有害性感知的分类器判定62%的框架化提示有害,而对照组仅为42%。在最易受攻击的模型上,框架化提示的成功率高达93%。在五分之四的有害性评判标准下,单次查询的攻击效果达到或超过已发表的单次查询攻击。该效应更多依赖于模型而非主题,且不随规模扩大而减弱。在近三分之二的框架化样本上,至少有一个评判者与评判小组的多数意见不一致,这本身就是一个关于测量有效性的发现。这一模式与泛化不匹配(mismatched generalisation)的假设一致。
cs.AI / 32 / 2609.02216

PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion

PEARL:基于上下文子图的路径-实体对齐关系学习方法用于归纳式知识图谱补全
Yang, Yunchi, Li, Longlong, Qu, Cunquan
Abstract
Inductive knowledge graph completion (IKGC) aims to predict missing links involving entities unseen during training, requiring models to learn transferable relational and structural patterns. Existing subgraph- and path-based approaches often encode relational paths independently of their surrounding query subgraphs, although the predictive relevance of a path may vary across structural contexts. We propose PEARL, a Path-Entity Aligned Relational Learning framework that models paths as context-conditioned reasoning signals. PEARL constructs a query-specific contextual subgraph from the union of the query entities' neighborhoods and uses a large language model (LLM)-guided retriever to distill semantically relevant paths. It then builds a bipartite interaction graph over paths, contextual entities, and a global subgraph representation, allowing path embeddings to adapt to local and global structural evidence. To suppress noise introduced by the enlarged context, PEARL employs a dual-view contrastive objective that promotes representation consistency under stochastic contextual perturbations. Experiments on WN18RR, FB15k-237, and NELL-995 show that PEARL obtains the best average Hits@10 among the compared IKGC methods on all three benchmarks. Ablation studies, efficiency analyses, and case studies further validate the contributions of contextual subgraph modeling, semantic path retrieval, path-entity interaction, and contrastive regularization.
Chinese Translation
归纳式知识图谱补全(Inductive Knowledge Graph Completion, IKGC)旨在预测涉及训练阶段未见实体的缺失链接,这要求模型学习可迁移的关系模式与结构模式。现有的基于子图和路径的方法通常独立于其所在的查询子图对关系路径进行编码,然而路径的预测相关性可能因结构上下文的不同而变化。我们提出了PEARL,一个路径-实体对齐关系学习框架,它将路径建模为依赖上下文的推理信号。PEARL从查询实体的邻域并集中构建查询特定的上下文子图,并利用大语言模型(LLM)引导的检索器提取语义相关的路径。随后,它在路径、上下文实体和全局子图表示之上构建二部交互图,使路径嵌入能够适应局部与全局的结构证据。为抑制扩大上下文所引入的噪声,PEARL采用双视角对比学习目标,促进模型在随机上下文扰动下表示的一致性。在WN18RR、FB15k-237和NELL-995数据集上的实验表明,PEARL在所有三个基准测试中均取得了所对比的IKGC方法中最优的平均Hits@10。消融实验、效率分析以及案例研究进一步验证了上下文子图建模、语义路径检索、路径-实体交互和对比正则化的贡献。
cs.AI / 33 / 2609.02217

SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams

SkillGLoW:面向长程任务流中自我改进智能体的程序族技能整合方法
Yan, Ao, Zhang, Xin, Du, Jiawei, Zhou, Joey Tianyi
Abstract
LLM agents increasingly self-improve by writing and reusing textual skills, kept either as one global document or as a flat pool of per-task entries, though most of the evidence comes from domains with structurally similar tasks. On long-horizon workloads where each task demands a different solution, the two forms fail in opposite ways: the document collapses into generic discipline, while the pool inflates and its entries stay bound to the instance that wrote them. We argue the missing unit of reuse is the solving procedure shared by a cluster of related tasks, and build SkillGLoW (Global-Local Weave) around it: the local skills a task writes from its own execution are aggregated into procedural families and compressed into de-instantiated global priors, while the instance detail they hold is regenerated per task rather than stored; a commit gate admits a prior only when real execution shows it does not degrade the deployed library. Across four benchmarks (mathematical reasoning, terminal automation, software repair, and embodied control) and three models, the priors gain 17.2 points (hard) over the no-skill baseline on average, with positive gains in all 12 continual-improvement runs, and 18.0 with local regeneration, while the library holds one prior per procedural family, 3.6x more compact than the per-task pool. Under the same protocol GLoW leads a published single-document optimizer on 15 of 21 cells. Unmodified, the library lifts success on unseen ALFWorld tasks from 73.9% to 83.9%, evidence that what transfers is procedure rather than task memory.
Chinese Translation
LLM智能体越来越多地通过编写和复用文本化技能来实现自我改进,这些技能或保存为一份全局文档,或保存为一个按任务条目构成的扁平化池,但现有证据大多来自任务结构相似的领域。在每项任务都需要不同解决方案的长程工作负载中,这两种形式以相反的方式失效:全局文档退化为泛化的行为准则,而任务池则不断膨胀且其条目始终绑定于产生它们的实例。我们认为,所缺失的复用单元是由一组相关任务共享的求解程序(procedure),并据此构建了SkillGLoW(Global-Local Weave,全局-局部编织):任务根据自身执行情况写入的局部技能被聚合为程序族(procedural families),并压缩为去实例化的全局先验;先验所包含的实例细节在需要时按任务重新生成,而非直接存储;提交门控(commit gate)只有在真实执行表明其不会降低已部署技能库性能时才接纳一个先验。在四个基准(数学推理、终端自动化、软件修复和具身控制)和三个模型上,先验技能相较无技能基线平均提升17.2分(困难任务),在全部12个持续改进运行中均取得正向增益,结合局部再生成后提升达18.0分;技能库为每个程序族仅保留一个先验,比按任务池紧凑3.6倍。在相同协议下,GLoW在21个实验格中的15个上优于已发表的单文档优化器。未经修改的技能库将ALFWorld未见任务的成功率从73.9%提升至83.9%,这证明可迁移的是程序本身而非任务记忆。
cs.AI / 34 / 2609.02231

PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment

PhoenixNest-Video:面向自动化视频面试评估的证据驱动多模态智能体框架
Yuxuan, Fan, Miaojun, Huang, Haimei, Zhang, Jingshen, Wu, Hao, Liu
Abstract
Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI approaches yield opaque scores without traceable rationale. We introduce PhoenixNest-Video, an evidence-grounded multimodal agent framework for automated video interview assessment. It builds a semantic video graph as structured working memory, performs rubric-conditioned retrieval with cross-modal verification across visual, audio, and textual streams, and produces per-criterion scores anchored to the candidate's materials. A Scorer trained via Rubrics-based Reinforcement Learning with dual rewards for rubric alignment and score-level differentiation internalizes the discriminative structure of multi-level rubrics. PhoenixNest-Video attains 91.50\% grade-level accuracy on VInterview-2025, outperforming substantially larger proprietary models. A compact, rubric-grounded agent therefore scores candidates in closer agreement with an expert panel than direct prompting of much larger models, and exposes the evidence behind each score for human review.
Chinese Translation
面试评估需要基于行为证据对每项评分标准做出判断,然而申请人数量激增使得纯人工评估成本高昂且不一致,而现有AI方法产生的分数不透明、缺乏可追溯的依据。我们提出了PhoenixNest-Video,一个面向自动化视频面试评估的证据驱动多模态智能体框架。该框架构建语义视频图作为结构化工作记忆,在视觉、音频和文本流之间执行基于评分标准的检索并结合跨模态验证,最终生成锚定于候选人材料的逐项评分。通过基于评分标准的强化学习(Rubrics-based Reinforcement Learning)训练的评分器(Scorer),采用评分标准对齐与分数区分度的双重奖励,内化了多级评分标准的判别结构。PhoenixNest-Video在VInterview-2025数据集上达到91.50%的等级判定准确率,显著优于规模大得多的专有模型。因此,一个紧凑的、以评分标准为依据的智能体在候选人评分上比直接提示规模大得多的模型更接近专家小组的判断,并且能够呈现每个分数背后的证据以供人工审核。
cs.AI / 35 / 2609.02236

PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks

PGPO:面向多轮智能体任务的势能引导策略优化
Zheng, Yuyao, Sun, Haipeng, Bao, Junwei, Liu, Lemao, Jiang, Hongfei, Song, Yang, Dou, Dejing
Abstract
Group-based reinforcement learning (RL) has become an effective paradigm for LLM post-training, but in multi-turn agentic tasks with sparse terminal rewards, it often provides coarse credit for intermediate actions. To obtain more fine-grained credit assignment, recent work such as GiGPO introduces step-level advantages for intermediate actions. However, these step-level signals still rely on the final outcome of each individual trajectory. As a result, actions within failed trajectories can remain poorly differentiated, so effective actions can receive the same unfavorable credit as erroneous ones. In this work, we propose Potential-Guided Policy Optimization (PGPO) for multi-turn agentic tasks. PGPO estimates empirical state potentials from anchor-state-group return statistics within each rollout group. It then derives action advantages from potential differences between adjacent states, enabling cross-trajectory credit propagation. This provides finer-grained step-level credit assignment, especially within failed trajectories. Experiments on ALFWorld and WebShop show strong overall performance relative to recent group-based RL methods. Further analysis provides evidence that PGPO yields more informative failure-side credit signals with negligible training overhead.
Chinese Translation
基于分组的强化学习(RL)已成为大语言模型(LLM)后训练的有效范式,但在具有稀疏终端奖励的多轮智能体任务中,它往往对中间动作提供较为粗糙的信用分配。为了获得更细粒度的信用分配,近期的工作(如 GiGPO)为中间动作引入了步骤级的优势估计。然而,这些步骤级信号仍然依赖于每条轨迹自身的最终结果。因此,失败轨迹中的动作可能依然难以有效区分,导致有效动作与错误动作获得同样不利的信用评分。在本工作中,我们提出了面向多轮智能体任务的势能引导策略优化(Potential-Guided Policy Optimization, PGPO)。PGPO 利用每个 rollout 组内锚点状态组的回报统计来估计经验状态势能,然后通过相邻状态之间的势能差推导动作优势,从而实现跨轨迹的信用传播。这提供了更细粒度的步骤级信用分配,尤其是在失败轨迹内部。在 ALFWorld 和 WebShop 上的实验表明,PGPO 相较于近期的基于分组的强化学习方法具有出色的整体性能。进一步的分析证明,PGPO 能够以几乎可以忽略的训练开销产生更具信息量的失败侧信用信号。
cs.AI / 36 / 2609.02242

Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality

提议以学习,学习以提议:有限理性下的可评估性感知辅助
Zhu, Yifan, Katt, Sammie, Kaski, Samuel
Abstract
AI assistants often collaborate by proposing candidate edits, plans, or designs that users evaluate before adoption. Existing assistance methods focus on proposal quality or user-goal inference, often assuming that the user can reliably evaluate any proposal, which can fail in practice because of bounded rationality. We study evaluability-aware proposal planning, where proposals serve both as task interventions and as probes for learning latent preferences and evaluation constraints, where the resulting belief updates then guide later proposals. We formalise this setting as ProSE, a hidden-parameter sequential assistance problem, and instantiate it with a KL-regularised bounded-rational binary response model in which acceptance trades off value gain against a distance-dependent evaluability penalty. Analysing the planning consequence of this likelihood reveals that likely accepted proposals and informative probes need not coincide, which explains why planners that only pursue acceptance systematically underperform. We operationalise ProSE with \textsc{ProSE-Plan}, a depth-2 Bayes-adaptive planner that scores proposals by possible responses and response-induced posterior beliefs. In controlled graph simulations, \textsc{ProSE-Plan} improves over evaluability-unaware and myopic baselines when evaluation cost is the bottleneck, and a probe-commit ablation confirms that our approach selects informative proposals that simpler methods miss. Our results thus identify user evaluability as a planning-relevant dimension of AI assistance, complementary to generation quality and preference inference.
Chinese Translation
AI助手常通过提出候选编辑、计划或设计来协助用户,用户在采纳前会对其进行评估。现有的辅助方法主要关注提议质量或用户目标推断,往往假设用户能够可靠地评估任何提议,但由于有限理性(bounded rationality),这一假设在实践中常常失效。我们研究了可评估性感知的提议规划(evaluability-aware proposal planning),其中提议既作为任务干预手段,也作为探测手段,用于学习用户的潜在偏好与评估约束,由此产生的信念更新进而指导后续提议。我们将该设定形式化为ProSE——一个隐藏参数的序贯辅助问题,并用KL正则化的有限理性二元响应模型对其进行实例化,其中接受与否取决于价值收益与随距离变化的可评估性惩罚之间的权衡。对该似然函数的规划后果进行分析后发现,最可能被接受的提议与信息量大的探测并不必然重合,这解释了为何仅追求接受率的规划器会系统性地表现不佳。我们通过ProSE-Plan实现ProSE,这是一个深度为2的贝叶斯自适应规划器,根据用户可能的响应及响应所诱导的后验信念对提议进行评分。在受控的图模拟实验中,当评估成本成为瓶颈时,ProSE-Plan优于不考虑可评估性和短视的基线方法,且探测承诺消融实验证实,我们的方法能够选择简单方法所遗漏的高信息量提议。因此,我们的研究结果将用户可评估性确定为AI辅助中与规划相关的一个维度,与生成质量和偏好推断互为补充。
cs.AI / 37 / 2609.02244

Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training

任务级自然语言先验作为低资源大语言模型训练的学习信号
Gao, Jian, Zhang, Xiao, Zhu, Xun, Li, Miao, Wu, Ji
Abstract
Large language models (LLMs) often struggle when low-resource training data are ambiguous or incomplete. Task-level natural-language priors can provide useful guidance in such settings, but existing approaches usually treat these priors as input context rather than as learning signals during training. We propose Prior-Guided Tuning (PGT), a training perspective that incorporates natural-language priors as auxiliary learning signals for low-resource LLM training. Under this perspective, we introduce Contrastive Prior Steering (CPS), which keeps the original supervised objective intact while adding positive and negative prior-conditioned auxiliary losses to encourage task-consistent learning and discourage plausible but misleading alternatives. Experiments on AmbiMath, Jigsaw, and MNLI/HANS show that CPS consistently improves over plain and prompt fine-tuning. On AmbiMath, CPS achieves 97.6% average exact-match accuracy. On Jigsaw, CPS improves average Macro F1 by 9.5 percentage points over standard fine-tuning, and with 1/10 of the experimental training data slightly exceeds full-data plain fine-tuning. On HANS, CPS improves non-entailment accuracy by 8.3 and 5.2 percentage points for LLaMA 3.1 8B and Qwen 2.5 7B, respectively, while maintaining comparable in-domain MNLI accuracy. These results support our central claim: task-level natural-language priors can provide useful guidance as auxiliary learning signals for low-resource LLM training. Our code and data will be publicly available.
Chinese Translation
大语言模型(LLM)在低资源训练数据模糊或不完整时常常表现不佳。任务级自然语言先验在此类场景下可以提供有用的指导,但现有方法通常将这些先验视为输入上下文,而非训练过程中的学习信号。我们提出先验引导微调(Prior-Guided Tuning, PGT),这是一种将自然语言先验作为辅助学习信号融入低资源LLM训练的训练视角。在该视角下,我们引入对比先验引导(Contrastive Prior Steering, CPS),它在保持原始监督目标不变的同时,添加基于先验的正负辅助损失,以促进与任务一致的学习,并抑制貌似合理但具有误导性的替代方案。在AmbiMath、Jigsaw以及MNLI/HANS数据集上的实验表明,CPS持续优于普通微调和提示微调。在AmbiMath上,CPS达到了97.6%的平均精确匹配准确率。在Jigsaw上,CPS较标准微调将平均Macro F1提升了9.5个百分点,且仅使用十分之一的实验训练数据即可略超全量数据下的普通微调。在HANS上,CPS分别将LLaMA 3.1 8B和Qwen 2.5 7B的非蕴含准确率提升了8.3和5.2个百分点,同时保持了与原有相当MNLI域内准确率。这些结果支持了我们的核心观点:任务级自然语言先验可以作为辅助学习信号,为低资源LLM训练提供有用的指导。我们的代码和数据将公开发布。
cs.AI / 38 / 2609.02246

LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails

LLM-as-a-Judge 并非神谕:为何自我改进的智能体需要确定性护栏
Wahi, Vansh
Abstract
Self-improving agent pipelines have a problem at their center. An optimizer rewrites prompts to score higher, and the score comes from a judge that is itself an LLM. That judge has the last word on whether the system is getting better, and our position is that it has not earned it. The judge should be demoted from oracle to advisor: its verdict becomes one input among several, and every change is gated instead by a deterministic verification layer the judge cannot override. We reached this position by building the alternative and running it. Over months of running autonomous prompt-optimization loops in production across contract analysis, compliance review, and code quality, we cataloged eleven ways the evaluation signal failed, in four classes: judge bias, harness and metric failures, ground-truth errors, and reward hacking. Agents achieved perfect scores by reading cached answer keys from their environment, a 100% pass rate concealing 68% true capability. A corrupted ground-truth label caused the optimizer to delete correct compliance rules to agree with it. A syntactically broken prompt was promoted as the winner because a silent parser fallback improved the metric. Attempts to fix the judge by rewriting its rubric plateaued; the only reliable gain came from a structural constraint on its output order. In response we describe PROCTOR, a Teacher-Student loop in which a stateful orchestrator holds all tool access, stateless subagents diagnose failures and draft mutations they cannot apply, and a Teacher grades those mutations under five deterministic guardrails: hermetic sandboxes, capability-disjoint roles, acceptance checks that outrank the Teacher, frozen holdouts, and canary cases engineered so that a perfect score is itself evidence of cheating. We report the failures this prevented, and, because the Teacher is itself an LLM judge, the failures it did not.
Chinese Translation
自我改进的智能体流水线的核心存在一个问题:优化器重写提示词以获得更高的分数,而该分数来自一个本身也是大语言模型(LLM)的评判器。这个评判器对系统是否在进步拥有最终裁决权,而我们的立场是,它并未赢得这一权力。评判器应当从神谕降级为顾问:其裁定只是众多输入之一,而每一项修改都应由评判器无法推翻的确定性验证层来把关。我们通过构建并运行这一替代方案得出了这一立场。在生产环境中跨越合同分析、合规审查和代码质量等任务、历时数月运行自主提示词优化循环的过程中,我们归纳了评估信号失效的十一种情形,可分为四类:评判偏差、评测框架与度量失效、真值错误以及奖励作弊。智能体曾通过读取环境中缓存的答案密钥获得满分——100%的通过率掩盖了仅68%的真实能力。一个被污染的真值标签导致优化器删除了正确的合规规则以迎合该标签。一个语法错误的提示词因静默的解析器回退机制改善了指标而被提升为胜出者。通过重写评判器评分标准来修复它的尝试均陷入瓶颈;唯一可靠的改进来自对其输出顺序的结构性约束。为此,我们提出了 PROCTOR——一个师生循环架构:一个有状态的编排器掌握全部工具访问权限,无状态的子智能体负责诊断故障并起草其无权执行的修改方案,而教师则在五项确定性护栏下对这些修改进行评分:密封沙箱、能力隔离的角色、优先于教师的验收检查、冻结的保留测试集,以及经过精心设计的金丝雀用例——使得满分本身即是作弊的证据。我们报告了该系统所避免的失败,同时,由于教师本身也是一个 LLM 评判器,我们也报告了它未能避免的失败。
cs.AI / 39 / 2609.02253

APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering

APEx:面向自适应深度研究问答的智能体程序化经验蒸馏
Ding, Jie, Sun, Rui, Zhang, Xinyuan, Zhang, Zeyu, Liu, Xin
Abstract
Deep research agents augment large language models with external tools to answer complex, long-horizon questions through multi-turn reasoning. Learning from prior experience is crucial for continual improvement, yet existing methods either retrieve verbose task-specific traces that burden decision-making, or distill procedural skills that remain decoupled from downstream policy adaptation. We propose APEx, a hierarchical experience utilization framework that organizes interaction history into instance-level trajectory memories and category-level procedural skills, and couples them through a closed-loop architecture of Executor, Distiller, and Planner. The three modules are optimized via a three-stage alternating GRPO training paradigm, enabling reward-guided skill distillation rather than fixed-prompt generation. At test time, distilled skills serve as procedural priors for online Planner adaptation through skill-guided test-time reinforcement learning, allowing ground-truth-free self-improvement with skill-alignment regularization to prevent policy drift. Experiments on 7 benchmarks demonstrate that APEx achieves state-of-the-art performance, surpassing GPT-5.4 by 14.7 points and the strongest memory-augmented baseline by 3.0 points.
Chinese Translation
深度研究智能体通过为大型语言模型配备外部工具,经由多轮推理来回答复杂的长程问题。从先前经验中学习对持续改进至关重要,然而现有方法要么检索冗长的任务特定轨迹而加重决策负担,要么蒸馏出的程序化技能与下游策略适配相互脱节。我们提出APEx,一个层次化的经验利用框架,它将交互历史组织为实例级的轨迹记忆和类别级的程序化技能,并通过由执行器(Executor)、蒸馏器(Distiller)和规划器(Planner)构成的闭环架构将二者耦合。这三个模块通过三阶段交替的GRPO训练范式进行优化,实现了奖励引导的技能蒸馏,而非基于固定提示的生成。在测试阶段,蒸馏所得的技能作为程序化先验,通过技能引导的测试时强化学习支持规划器的在线自适应,借助技能对齐正则化防止策略漂移,从而实现无需真实标签的自我改进。在7个基准上的实验表明,APEx取得了最先进的性能,超越GPT-5.4达14.7分,超越最强的记忆增强基线3.0分。
cs.AI / 40 / 2609.02264

Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems

Codebook Agent:面向LLM多智能体系统的摊销式拓扑设计
Yu, Jinxi, Li, Yubei, Jiang, Eric Hanchen, Zhang, Zhi, Liu, Dong, Zhao, Wenxiao, Li, Levina, Chang, Kai-Wei, Wu, Ying Nian
Abstract
Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and efficiency, yet current designers treat this as conditional graph generation: a variational, autoregressive, or diffusion decoder searches the $N \times N$ adjacency space, and a graph-network proxy trained on utility and a structural cost such as edge count ranks the sampled candidates. We argue that this formulation is misaligned with the problem. Empirically, topologies that survive a reward filter collapse to about six distinct graphs even when the codebook capacity grows from 8 to 64; edge count is negatively correlated with measured token consumption (Pearson $r \approx -0.4$), so sparsifying the graph makes inference more expensive; and a message-passing scorer over agent-profile nodes is adjacency-invariant whenever agents share a profile---the default configuration of published benchmarks---so it cannot rank candidates at all in that regime. These three facts motivate Codebook Agent: a vector-quantized autoencoder compresses successful topologies into a query-independent 16-entry codebook; a reward-weighted MLP maps the query embedding to a distribution over codes; and an MLP proxy that reads the flattened adjacency, regressed on measured utility and per-task normalized token cost, reranks the top decoded candidates in a single batched forward pass. With no iterative search and no message passing at test time, Codebook Agent is the most accurate method on all six benchmarks we compare (84.6 average against 83.0 for the strongest prior designer), emits a topology in 2.4 ms, and uses 21.9--33.2% fewer LLM tokens.
Chinese Translation
根据每个查询自适应调整LLM多智能体系统的通信拓扑可以同时提升准确率和效率,然而现有设计者将此视为条件图生成任务:使用变分、自回归或扩散解码器在 $N \times N$ 的邻接空间中搜索,并利用一个在效用和结构代价(如边数)上训练的图网络代理模型对采样出的候选拓扑进行排序。我们认为这种问题表述与实际问题并不匹配。实验表明,即使码本容量从8增长到64,能通过奖励筛选的拓扑也坍缩为大约六种不同的图;边数与实测的token消耗呈负相关(Pearson $r \approx -0.4$),因此稀疏化图结构反而会使推理更加昂贵;而且,当智能体共享相同的profile时(这是已发表基准测试的默认配置),基于智能体profile节点的消息传递评分器具有邻接不变性,因此在该情形下完全无法对候选进行排序。这三个事实促使我们提出Codebook Agent:一个向量量化自编码器将成功的拓扑压缩为与查询无关的16项码本;一个奖励加权的MLP将查询嵌入映射到码字上的分布;另一个读取扁平化邻接矩阵的MLP代理模型——以实测效用和按任务归一化的token代价作为回归目标——在单次批量前向传播中对排名靠前的解码候选进行重排序。由于在测试时既无需迭代搜索也无需消息传递,Codebook Agent在我们比较的全部六个基准上都是最准确的方法(平均84.6分,对比最强先前设计者的83.0分),生成一个拓扑仅需2.4毫秒,并且节省21.9%–33.2%的LLM token消耗。
cs.AI / 41 / 2609.02273

CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging

CoMerge:面向多任务模型合并的冲突驱动偏好优化
Zheng, Mingjie, Chen, Zihao, Chen, Wenqing, Yuan, Weile, Chu, Zhixuan, Yu, Jianxing, Zheng, Zibin
Abstract
Model merging provides an efficient paradigm for constructing multi-task large language models (LLMs) without full model retraining, yet it remains challenged by parameter interference. While existing methods aim to preserve the capabilities of individual expert models and mitigate interference, they generally do not directly learn from the potentially degraded behaviors exposed by naive merging. In this paper, we propose a conflict-driven preference optimization framework for model merging (CoMerge), which reformulates model merging as a preference optimization problem. The approach utilizes a self-supervised, conflict-driven strategy that leverages the defects of naive merging methods (e.g., task arithmetic) as hard negative samples to construct preference pairs without external annotations. By applying preference optimization to refine lightweight, tensor-wise merging coefficients, CoMerge enables the model to mitigate parameter-space conflicts while preserving task-specific capabilities. Extensive experiments show that CoMerge achieves an average normalized performance of 0.9968 on MergeBench, outperforming all evaluated data-free and data-driven model-merging baselines. Furthermore, on Llama-3.1-8B-Instruct, CoMerge yields marked improvements on conflict-sensitive tasks such as instruction following and safety, while remaining highly competitive with full-parameter fine-tuning despite optimizing only 1,445 scalar coefficients.
Chinese Translation
模型合并为构建多任务大语言模型(LLM)提供了一种无需完整模型重训练的高效范式,但其仍然面临参数干扰的挑战。尽管现有方法旨在保留单个专家模型的能力并缓解干扰,但它们通常并未直接从朴素合并所暴露出的性能退化行为中学习。在本文中,我们提出了一种面向模型合并的冲突驱动偏好优化框架,旨在从模型合并所暴露的冲突中学习。该方法采用一种自监督、冲突驱动的策略,利用朴素合并方法(如任务算术)的缺陷作为困难负样本,在无需外部标注的情况下构建偏好对。通过应用偏好优化来精调轻量的逐张量合并系数,CoMerge 使模型能够在缓解参数空间冲突的同时保留任务特定能力。大量实验表明,CoMerge 在 MergeBench 上取得了 0.9968 的平均归一化性能,优于所有被评估的无数据和数据驱动的模型合并基线方法。此外,在 Llama-3.1-8B-Instruct 上,CoMerge 在指令遵循和安全性等冲突敏感任务上取得了显著提升,同时尽管仅优化了 1,445 个标量系数,其性能仍与全参数微调高度接近。
cs.AI / 42 / 2609.02292

SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

SCX Router:基于解码器KV分类器与真实任务本体的流式零样本模型选择
Stepanov, Ihor, Smechov, Aleksandr, Shtopko, Mykhailo, Vodianytskyi, Dmytro, Lukashov, Oleksandr
Abstract
The rapid proliferation of large language models (LLMs) and the growing diversity of their applications presents a unique optimization opportunity: selecting the right model for the task, while optimizing for speed, cost, and quality at a per-task level. However, inference endpoints can vary widely in quality, price, latency, context support, tool use, domain expertise, and reasoning behavior. This heterogeneity makes manual heuristics difficult to maintain and unlikely to achieve consistently favorable speed--cost--quality trade-offs on their own. We introduce \router{}, a lightweight GLiClass-based router that assigns a suitability score to each inference-time model label without autoregressive generation. The released 0.6B-parameter checkpoint combines a Qwen3 decoder with a shallow bidirectional scorer. Its decoder-KV execution path preserves a text-only key--value cache across a session, encodes only new dialogue turns, and evaluates transient candidate-label tokens without adding them to the persistent cache. The same checkpoint also predicts task type, difficulty, reasoning mode, and expected output length, and supports custom zero-shot labels. For task generation, we construct a task ontology with 23 families, 115 task types, 345 routable subtypes, 1,173 synthetic examples, and an orthogonal axis of 30 domains. Using this structure, we generate 150,000 verifier-scored tasks and 15,000 open-ended tasks. We then train the Qwen3 decoder on these tasks, while explicitly separating learned request prediction from per-task policies for attributes such as eligibility, cost, cache reuse, safety, and sovereignty. Across six LiveBench subsets, the router outperforms the mean candidate; on the selected 1,000-task subset, it achieves an aggregate top-1 score of 0.707 versus 0.696 for the strongest fixed model, with benchmark-dependent gains.
Chinese Translation
大语言模型(LLM)的快速普及及其应用日益多样化,带来了一个独特的优化机会:为每个任务选择合适的模型,同时在任务层面优化速度、成本与质量。然而,推理端点在质量、价格、延迟、上下文支持、工具使用、领域专长和推理行为等方面差异巨大。这种异质性使得人工启发式规则难以维护,也难以仅凭其自身持续实现理想的速度-成本-质量权衡。我们提出了SCX Router,一个基于GLiClass的轻量级路由器,无需自回归生成即可为每个推理时模型标签分配适配度评分。发布的0.6B参数检查点将Qwen3解码器与浅层双向评分器相结合。其解码器KV执行路径在整个会话中保留纯文本键值缓存,仅对新对话轮次进行编码,并评估瞬态候选标签词元而不将其加入持久缓存。同一检查点还可预测任务类型、难度、推理模式和预期输出长度,并支持自定义零样本标签。在任务生成方面,我们构建了一个包含23个任务族、115种任务类型、345个可路由子类型、1,173个合成示例,以及由30个领域构成的正交轴的任务本体。基于该结构,我们生成了150,000个经验证器评分的任务和15,000个开放式任务。随后,我们在这些任务上训练Qwen3解码器,同时将学习到的请求预测与针对资格、成本、缓存复用、安全性和主权等属性的任务级策略显式分离。在LiveBench的六个子集上,该路由器优于候选模型的平均水平;在选定的1,000个任务子集上,其聚合top-1得分达到0.707,而最强固定模型为0.696,且收益随基准测试而异。
cs.AI / 43 / 2609.02302

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

利用推理时计算与部署脚手架提升评估的现实性
Ahlqvist, Axel, Guan, Richard, Rivera, Juan-Pablo, Kassler, Adeline, Troitskii, Dmitrii, Souly, Alexandra, Fronsdal, Kai, Kirk, Robert, Hughes, John
Abstract
A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment evaluations harder to distinguish from real deployments. Our first technique, critique refinement, spends additional inference-time compute on each simulator action: the simulator generates multiple candidate actions, refines them using feedback from an instance of the target model on how to make them more realistic, and continues the evaluation with the most deployment-like candidate. Our second technique, DISH (Deployment-Imitating SWE-Agent Harness), wraps the target in an agent harness, reducing the gap between simulated and real deployment environments in coding settings. We test the techniques on multiple target models and find that they compose: applying both yields larger realism gains than either alone. Our results show that automated approaches can improve the realism of alignment evaluations, and that these improvements use additional compute more effectively than making the audits longer.
Chinese Translation
对齐评估面临的一个核心障碍是评估意识:能力强大的模型能够辨别自己是在被测试而非被部署,这削弱了安全评估所能支持的结论的可靠性。我们提出了两种技术,使模拟的对齐评估更难与真实部署区分开来。第一种技术是批评式细化(critique refinement),它在模拟器的每个动作上投入额外的推理时计算:模拟器生成多个候选动作,利用目标模型实例提供的关于如何使其更真实的反馈对这些候选动作进行细化,然后以最接近真实部署风格的候选动作继续评估。第二种技术是DISH(部署模仿型SWE-Agent框架,Deployment-Imitating SWE-Agent Harness),它将目标模型封装在一个智能体框架中,缩小了编程场景下模拟环境与真实部署环境之间的差距。我们在多个目标模型上测试了这些技术,发现两者可以叠加使用:同时应用两种技术所获得的现实性提升大于单独使用任一技术。我们的结果表明,自动化方法能够提升对齐评估的现实性,并且与延长审计时间相比,这些改进能更有效地利用额外的计算资源。
cs.AI / 44 / 2609.02336

SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning

SALA:面向上下文学习中复杂推理的语义感知逻辑对齐方法
Ji, Zhao, Chen, Wenqing, Chu, Zhixuan, Yu, Jianxing, Liu, Jingping, Zhao, Shanhe, Zheng, Zibin
Abstract
Effective in-context learning (ICL) for complex reasoning relies on selecting the right demonstrations. Traditional retrieval methods based on surface similarity fail to capture the underlying problem-solving logic. Recent logic-based methods address this by matching predefined reasoning steps, but the rigid rules and exact-match criteria is improper to handle flexible or diverse reasoning processes. To address the problem, we propose SALA, a Semantic-Aware Logical Alignment framework. Instead of relying on a fixed inventory, SALA automatically learns task-specific reasoning operations. It then embeds these operations into a continuous semantic space and uses dynamic time warping (DTW) to align the reasoning sequences. This approach allows for soft, flexible matching of reasoning logic while remaining highly interpretable. Experiments across four reasoning benchmarks and three LLMs demonstrate that SALA outperforms existing demonstration selection methods. Further analysis confirms the roles of the operation induction and the logical semantic alignment.
Chinese Translation
针对复杂推理任务的有效上下文学习(ICL)依赖于示例的正确选择。传统的基于表面相似度的检索方法无法捕捉底层的问题求解逻辑。近年来基于逻辑的方法通过匹配预定义的推理步骤来解决这一问题,但僵化的规则和精确匹配的标准难以妥善处理灵活或多样的推理过程。为解决这一问题,我们提出了SALA,一个语义感知逻辑对齐框架。SALA不再依赖固定的操作库,而是自动学习特定任务的推理操作。随后,它将这些操作嵌入到连续的语义空间中,并利用动态时间规整(DTW)对推理序列进行对齐。该方法能够实现推理逻辑的软性、灵活匹配,同时保持高度的可解释性。在四个推理基准和三个大语言模型上的实验表明,SALA优于现有的示例选择方法。进一步的分析验证了操作归纳与逻辑语义对齐的作用。
cs.AI / 45 / 2609.02371

Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions

以洞察辅助诊断:基于行为抽象的智能体失败结构化分析
Bi, Jiayi, Gao, Yanjie, Xie, Yuanmin, Li, Liqun, Xu, Tianyin, Yang, Fan, Yang, Mao
Abstract
With the proliferation of LLM agents, the ability to understand and diagnose failures in agents is essential to achieving superior effectiveness and trustworthiness. As agent failures often manifest via long and complex trajectories, manually finding the needles in the haystack is untenable. However, traditional diagnosis techniques for software bugs can hardly address LLM agent failures, while completely relying on LLMs as the judge yields unreliable diagnosis results. To overcome these challenges, this paper presents AGENTSCOPE, a new neuro-symbolic approach for agent failure mode diagnosis. The key principle of AGENTSCOPE is to abstract agent behavior, based on its trajectories, into structured representations. Furthermore, AGENTSCOPE introduces the concept of neural invariants to specify agent behavior properties. AGENTSCOPE leverages LLM-guided reasoning atop the structured representation against neural invariants to pinpoint both the failure step and its type in the trajectory. We show the effectiveness of AGENTSCOPE on publicly available agent failure datasets (Who&When) and a more comprehensive dataset created by us (AgentErrata), where AGENTSCOPE significantly outperforms the current state of the art in fault localization and attribution accuracy. Our work shows that integrating structured abstractions with LLM-guided reasoning enables effective, reliable, and interpretable diagnosis for agent failures.
Chinese Translation
随着大语言模型(LLM)智能体的迅速普及,理解和诊断智能体失败的能力对于实现更优的效果和可信度至关重要。由于智能体失败往往通过冗长而复杂的轨迹表现出来,人工在大海捞针式的排查方式难以维系。然而,传统的软件缺陷诊断技术难以应对LLM智能体失败问题,而完全依赖LLM作为裁判又会产生不可靠的诊断结果。为克服这些挑战,本文提出了AGENTSCOPE,一种用于智能体失败模式诊断的新型神经符号方法。AGENTSCOPE的核心原则是基于智能体的轨迹,将其行为抽象为结构化表示。此外,AGENTSCOPE引入了神经不变式(neural invariants)的概念来刻画智能体的行为属性。AGENTSCOPE利用LLM引导的推理,在结构化表示之上对照神经不变式进行检查,从而精确定位轨迹中的失败步骤及其类型。我们在公开可用的智能体失败数据集(Who&When)以及我们构建的更全面的数据集(AgentErrata)上验证了AGENTSCOPE的有效性,结果表明AGENTSCOPE在故障定位和归因准确率上显著优于当前最先进的方法。我们的工作表明,将结构化抽象与LLM引导的推理相结合,能够为智能体失败提供有效、可靠且可解释的诊断。
cs.AI / 46 / 2609.02399

Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks

定量双极论证框架中的对比性解释
Yin, Xiang, Potyka, Nico, Rago, Antonio, Toni, Francesca
Abstract
Argumentation frameworks are useful tools for representing and reasoning with information in a variety of settings, e.g. in supplementing AI models as they perform classification tasks, with a notable benefit of providing additional explainability. In this paper, we introduce contrastive explanations for Quantitative Bipolar Argumentation Frameworks (QBAFs), one such formalism. Unlike most existing explanations for QBAFs, which explain the reasoning outcome of a single argument of interest (i.e. a topic argument), contrastive explanations explain the difference between two topic arguments. We introduce a general form of contrastive attribution functions (CAFs) and establish a set of general properties they should satisfy. We introduce CAFs based on removal, gradients and Shapley-values, and study their properties. Finally, to illustrate contrastive explanations, we demonstrate their usefulness in healthcare and bias identification settings.
Chinese Translation
论证框架是在多种场景下表示信息和进行推理的有用工具,例如在AI模型执行分类任务时为其提供辅助,其显著优势在于能够提供额外的可解释性。本文为定量双极论证框架引入了对比性解释。与现有大多数针对QBAF的解释方法不同——后者解释的是单个目标论证(即主题论证)的推理结果——对比性解释解释的是两个主题论证之间的差异。我们提出了一般形式的对比归因函数,并建立了其应满足的一组一般性质。我们引入了基于移除、梯度和Shapley值的对比归因函数,并研究了它们的性质。最后,为了说明对比性解释的作用,我们展示了其在医疗保健和偏见识别场景中的实用性。
cs.AI / 47 / 2609.02421

UTP-Bench: Uncertainty-aware Travel Planning Benchmark

UTP-Bench:不确定性感知的旅行规划基准
Rao, Etcharla Revanth, Karmakar, Priyanshu, Mallick, Shubhojit, Gupta, Manish, Ghosh, Shreya, Jana, Abhik
Abstract
Large Language Models (LLMs) have recently demonstrated strong capabilities in automated travel itinerary generation. However, real- world travel planning is inherently uncertain: transportation delays, crowd fluctuations, and unexpected stochastic delays frequently inval- idate otherwise feasible schedules. Existing benchmarks like TravelPlanner and TripCraft assume deterministic environments, evaluating only static constraint satisfaction and ignoring whether generated plans remain robust when such uncertainties arise. To address this limitation, we introduce UTP-Bench1 , a large-scale benchmark for uncertainty-aware travel planning. The dataset integrates real-world travel data spanning 504 cities of India, including attractions, restau- rants, accommodations, and multi-modal trans- portation networks. To model realistic disrup- tions, UTP-Bench incorporates empirical delay distributions and crowd-density patterns col- lected from major cities, enabling evaluation of travel plans under stochastic conditions. We further propose three evaluation metrics, namely Buffer Adequacy Score (BAS), Crowd- Aware Timing Score (CATS), and Transport Delay Absorption Score (TDAS), which quan- tify the ability of generated itineraries to main- tain robustness against transit delays and crowd variability. Experiments with state-of-the-art LLMs like GPT-5, Qwen3, Mistral and Phi-4 re- veal substantial gaps between model-generated and human-authored plans, particularly in tem- poral buffering, delay-aware transportation scheduling, and crowd-sensitive planning.
Chinese Translation
大型语言模型(LLMs)近期在自动化旅行行程生成方面展现出强大的能力。然而,现实世界的旅行规划本质上是不确定的:交通延误、人流波动以及意外的随机延误经常使原本可行的行程安排失效。现有的基准(如TravelPlanner和TripCraft)假设环境是确定性的,仅评估静态的约束满足情况,而忽略了生成的计划在这些不确定性出现时是否仍然稳健。为解决这一局限,我们提出了UTP-Bench¹,一个面向不确定性感知旅行规划的大规模基准。该数据集整合了涵盖印度504个城市的真实旅行数据,包括景点、餐厅、住宿以及多模式交通网络。为建模现实的干扰情况,UTP-Bench引入了从主要城市收集的经验延误分布和人群密度模式,从而能够在随机条件下评估旅行计划。我们进一步提出了三个评估指标,即缓冲时间充分性得分、人群感知时间得分和交通延误吸收得分,用于量化生成的行程在应对交通延误和人群波动时的鲁棒性。使用GPT-5、Qwen3、Mistral和Phi-4等最先进LLMs的实验表明,模型生成的计划与人工编写的计划之间存在显著差距,尤其是在时间缓冲、考虑延误的交通调度以及对人群敏感的规划方面。
cs.AI / 48 / 2609.02459

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

CivBench:面向《文明VI》中工具介导智能体的长时程基准测试
Andrews, Austin Tudor David, Wilkinson, Liam, Heagerty, Jamie, Coppock, Harry, Foerster, Jakob Nicolaus, Costa, Rui Ponte
Abstract
We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of tool calls over a large action space, requiring sustained planning, state monitoring, and execution under partial observability. The environment exposes 76 MCP tools and a narration layer that converts visual game state into structured text. We use CivBench to characterise agent behaviour across four model families in 23 admissible runs. The sample is a pilot, not a model ranking: aggregate outcomes do not reliably discriminate models at this scale. Instead, we introduce two interface-level metrics that the environment makes measurable: Proactive Monitoring Rate (PMR), capturing whether agents actively query latent strategic state, and RAG@10, capturing whether commitments stated in structured planning reflections are executed within ten subsequent turns. Across runs we observe two consistent patterns under a shared playbook protocol. Agents under-monitor strategically relevant state that is available but requires explicit querying: despite playbook guidance to query victory progress every 20 turns, agents do so only every 30 to 75 turns, and in 7 of 20 detectable defeats they failed to query within the 20 turn warning window before game end. Agents also frequently fail to execute near-term commitments stated in their own planning reflections (RAG@10 between 48.2% and 65.8% across models). Both patterns arise despite tool access and explicit guidance, and we interpret them as deviations under instruction rather than absences of capability. We release the environment, scenarios, logs, metrics, and analysis pipeline at https://github.com/lmwilki/civ6-mcp
Chinese Translation
我们提出了CivBench,一个开源基准测试,用于通过模型上下文协议(Model Context Protocol, MCP)在长时程、工具介导的环境中评估语言模型智能体。单个回合跨度超过300轮,在巨大的动作空间内产生数千次工具调用,要求智能体在部分可观测条件下进行持续的规划、状态监控与执行。该环境提供76个MCP工具以及一个将可视化游戏状态转换为结构化文本的叙述层。我们使用CivBench在23次可采信的运行中对四个模型家族的智能体行为进行刻画。该样本属于试点研究,而非模型排名:在此规模下,总体结果无法可靠地区分不同模型。相反,我们引入了两个该环境使其可测量的接口级指标:主动监控率(Proactive Monitoring Rate, PMR),衡量智能体是否主动查询潜在的战略状态;以及RAG@10,衡量智能体在结构化规划反思中所陈述的承诺是否在随后十轮内得到执行。在共享博弈协议下,我们在各次运行中观察到两个一致的模式。其一,智能体对可用但需要显式查询的战略相关状态监控不足:尽管博弈指南要求每20轮查询一次胜利进度,智能体实际仅每30至75轮查询一次,并且在20次可检测的失败局中,有7次未能在游戏结束前20轮的警告窗口内进行查询。其二,智能体经常未能执行其自身规划反思中所陈述的近期承诺(各模型的RAG@10介于48.2%至65.8%之间)。尽管智能体具备工具访问权限并接受明确指导,这两种模式仍然出现,我们将其解释为指令遵从上的偏差,而非能力的缺失。我们在 https://github.com/lmwilki/civ6-mcp 发布了环境、场景、日志、指标及分析流程。
cs.AI / 49 / 2609.02620

Collective creativity in hybrid societies

混合社会中的集体创造力
Youngblood, Mason, Mudd, Katie, Anglada-Tort, Manuel, Jones, Cameron, Miu, Elena, Omigie, Diana, Schedel, Margaret
Abstract
Generative AI is changing how cultural artifacts are created and circulated, and with it our understanding of creativity itself. Researchers disagree about whether these tools enrich or impoverish culture, and we argue that much of that disagreement comes from conflating two distinct components of creativity: novelty, a property of single artifacts, and diversity, a property of populations. We argue further that creativity in the context of generative AI is best understood as a property of hybrid collectives, or populations of interacting people and algorithms, rather than of individuals. AI-assisted ideation reliably raises the novelty of individual output while narrowing diversity in the aggregate, but this is not an inevitable consequence of putting machines in the loop. Because humans and models search in complementary ways, mixed groups can outperform and out-diversify groups of either kind alone, and machine-discovered solutions can enter human culture and persist there. What decides the outcome is composition: which agents are present, in what proportion, and how they are connected. The question is no longer whether AI helps or harms creativity, but which mixtures let individual gains accumulate without eroding collective diversity.
Chinese Translation
生成式人工智能(Generative AI)正在改变文化产品的创作与传播方式,也随之改变我们对创造力本身的理解。研究者们对这些工具究竟是丰富还是贫瘠了文化存在分歧,我们认为,这种分歧很大程度上源于将创造力的两个不同组成部分混为一谈:新颖性(novelty),即单个作品的属性;以及多样性(diversity),即群体的属性。我们进一步认为,在生成式人工智能的背景下,创造力最好被理解为混合集体(hybrid collectives)——即人与算法相互作用的群体——的属性,而非个体的属性。借助人工智能的构思活动确实能可靠地提升个体产出的新颖性,但同时会缩小整体的多样性,但这并不是将机器引入循环的必然结果。由于人类和模型以互补的方式进行搜索,混合群体可以超越并比任何单一类型的群体更具多样性,而且机器发现的解决方案能够进入人类文化并在其中延续。决定结果的是构成:哪些智能体存在、以何种比例存在,以及它们如何相互连接。问题不再人工智能是有助于还是有害于创造力,而在于哪种混合方式能让个体收益不断累积,同时又不侵蚀集体多样性。
cs.AI / 50 / 2609.02649

Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting

Loom:通过嵌入空间重加权将诊断线索编织为自由文本共识
Begleiter, Ron, Berg, Katya Egert, Saban, Gilad, Shabat, Gil
Abstract
Aggregating noisy, conflicting textual hypotheses into a reliable consensus is a fundamental challenge when deploying NLP systems in real-world industrial settings. While monolithic Large Language Model (LLM) agents offer unbounded expressivity for tasks like Root Cause Analysis (RCA), they suffer from context limits, compounding hallucinations, and prohibitive inference latency. Traditional weak supervision offers statistical rigor but is mathematically restricted to discrete classes. We present Loom, a generative consensus framework deployed for real-world RCA that bridges these paradigms. Loom aggregates open-form hypotheses emitted by modular heuristics (diagnostic templates dynamically populated with episode-specific entities, times, and metrics) by projecting them into a continuous embedding space, and resolves conflicting signals with an iterative centroid-based reweighting algorithm. The resulting consensus weights ground a single lightweight LLM synthesis step. Evaluated on the OpenRCA benchmark, Loom occupies the accuracy--efficiency Pareto frontier: it matches a state-of-the-art autonomous agent on Bank and Market-2 and trails on Market-1 and Telecom, while using a single LLM call per incident on all four datasets ($\sim$26$\times$ faster; $\sim$33$\times$ with an 8B-parameter synthesizer). We discuss our deployment experience, highlighting lessons learned regarding the trade-offs between agentic depth and inference latency, negative results in redundancy detection, and how deterministic consensus fosters trust among Subject Matter Experts~(SMEs).
Chinese Translation
在真实工业环境中部署NLP系统时,如何将嘈杂且相互冲突的文本假设聚合为可靠的共识是一项根本性挑战。单体式大语言模型(LLM)智能体在根因分析(RCA)等任务上具备无界的表达能力,但其受限于上下文长度、幻觉的累积放大以及过高的推理延迟。传统弱监督方法提供了统计严谨性,但在数学上仅限于离散类别。我们提出Loom,一个面向真实世界RCA部署的生成式共识框架,用以弥合上述两种范式。Loom将模块化启发式方法(利用事件特定实体、时间和指标动态填充的诊断模板)所产生的开放形式假设投影到连续嵌入空间中进行聚合,并通过一种基于质心的迭代重加权算法解决信号冲突。所得的共识权重为单次轻量级LLM合成步骤提供依据。在OpenRCA基准上的评估显示,Loom占据精度–效率的帕累托前沿:在Bank和Market-2数据集上与最先进的自主智能体表现相当,在Market-1和Telecom数据集上略逊一筹,同时在全部四个数据集上每次事件仅需一次LLM调用(快约26倍;采用8B参数合成器时快约33倍)。我们分享了部署经验,重点讨论了智能体深度与推理延迟之间的权衡、冗余检测方面的负面结果,以及确定性共识如何促进领域专家(SME)之间的信任。
cs.AI / 51 / 2609.02707

Door-in-the-Face Requests and Refusal Behaviour in Large Language Models

大型语言模型中的“留面子”(Door-in-the-Face)请求与拒绝行为
Jordan, Til
Abstract
Does the door-in-the-face technique work on language models? In humans, a large request that is refused makes a smaller follow-up request more likely to be granted. We test this on nine production models from three providers: each model refuses a large request, then receives a smaller version of the same request, and we compare its compliance with asking directly. The answer depends on the model. On Anthropic's frontier models the technique works: Opus 5 answers the smaller request 65.8% of the time after refusing the larger one, against 29.3% when asked directly. On the frontier models of OpenAI and Google, and on Haiku 4.5, it backfires, lowering compliance by 15.5 to 23.0 points. A control locates the effect: a refused large request on an unrelated topic does less than the related one on all nine models, so the concession itself matters everywhere, while the reaction to having just refused something differs by model family. The technique does not transfer to refusals drawn from public benchmarks. What decides whether a retreat can work is what the request asks for: rewriting 265 refused requests for usable instructions into requests for explanations of the same topic removed the refusal in 263 cases. Human influence techniques port to language models one model family at a time.
Chinese Translation
"留面子"(door-in-the-face)技巧对语言模型是否有效?在人类中,一个大请求被拒绝后,随之提出的较小请求更容易被接受。我们在来自三家提供商的九个生产级模型上测试了这一效应:每个模型先拒绝一个大请求,随后收到同一请求的较小版本,我们将其顺从率与直接提出请求的情况进行比较。结果显示,答案因模型而异。在Anthropic的前沿模型上,该技巧有效:Opus 5在拒绝较大请求后,有65.8%的概率回答较小请求,而直接提问时仅为29.3%。而在OpenAI和Google的前沿模型以及Haiku 4.5上,该技巧适得其反,使顺从率降低了15.5至23.0个百分点。一组对照实验定位了该效应的来源:在无关话题上被拒绝的大请求在所有九个模型上的影响均小于相关话题,说明让步本身在各处都起作用,而模型家族对“刚刚拒绝过某事”的反应则有所不同。该技巧无法迁移至取自公开基准的拒绝行为。决定退让能否奏效的关键在于请求的内容:将265个被拒绝的、请求可用指令的请求改写为对同一主题的解释性请求后,其中263例的拒绝行为消失了。人类的影响技巧向语言模型的迁移只能逐个模型家族进行。
cs.AI / 52 / 2609.02749

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Repo-To-Skill:将GitHub代码库蒸馏为AI4AI技能
Chen, Jianlyu, Hu, Yuyang, Qian, Hongjin, Liu, Jiawei, Wei, Wenqing, Chen, Xiaolong, Lian, Defu, Dou, Zhicheng, Li, Chaozhuo, Ye, Qiwei, Liu, Zheng
Abstract
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent. We call this missing layer operational knowledge, the know-how that separates knowing a method from making it work. That knowledge is not absent from the field. It appears in repositories and papers, but in forms written for human readers and too large to load during a task. Once distilled into compact, verified skills, this knowledge can be reused across tasks rather than rediscovered during each run. We present DisCo, a skill-powered research agent that creates skills and uses them during research. Its distillation runs in two complementary forms: task-agnostic, condensing the field's widely used repositories into reusable skills, and task-oriented, producing the skills a concrete task calls for. The former, applied across the open ecosystem, yields the AREX-Skill Library, with 5,000+ verified skills distilled from 1,000 widely used ML repositories and organized into 20 areas and 178 capability families. With the GPT-5.5 backbone, research harness, and downstream execution budget held fixed, the skill-equipped research agent scores 134.3% higher on MLE-bench, 34.4% higher on PaperBench, 9.2% higher on FrontierCS, and 14.0% higher on PassNet than the same agent without skills. These gains come from adding distilled operating context under that fixed setup.
Chinese Translation
自主智能体已经开始端到端地开展机器学习(ML)研究。这些智能体将模型骨干与用于规划、执行、记忆和验证的框架相结合,但这一架构仍将领域特有的专业知识置于智能体之外。我们将这一缺失的层次称为操作性知识(operational knowledge),即区分“了解一种方法”与“使该方法真正奏效”的诀窍。这类知识在领域中并非缺失,它存在于代码库和论文之中,只是以面向人类读者的形式呈现,且体量过大而无法在任务执行过程中加载。一旦将其蒸馏为紧凑且经过验证的技能(skills),这些知识便可在不同任务间复用,而无需在每次运行中重新探索。我们提出了DisCo,一个由技能驱动的研究智能体,它能够创建技能并在研究过程中加以使用。其蒸馏以两种互补形式进行:任务无关形式,将领域内广泛使用的代码库浓缩为可复用的技能;以及任务导向形式,为具体任务生成所需技能。前者应用于开放生态,构建了AREX-Skill Library,包含从1,000个广泛使用的ML代码库中蒸馏得到的5,000余项经验证的技能,并组织为20个领域和178个能力族。在保持GPT-5.5骨干模型、研究框架和下游执行预算固定不变的前提下,配备技能的研究智能体在MLE-bench上比无技能的相同智能体高出134.3%,在PaperBench上高出34.4%,在FrontierCS上高出9.2%,在PassNet上高出14.0%。这些提升源于在固定配置下添加了经过蒸馏的操作性上下文。
cs.AI / 53 / 2609.02750

Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems

双层协调反思:多智能体大语言模型系统的博弈论方法
Chen, Yihang, Chen, Yuxiang, Huang, Yuxuan, Fang, Meng, Luo, Weilin, Wang, Jun
Abstract
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory improvement, and the role of external verification. We model orchestrator-worker interaction as a bilevel coordination game: under bounded coupling, the workers' local-update game is an approximate potential game whose equilibrium slack is controlled by decomposition quality. We then analyse reflection as stochastic movement over semantic memory states. For free-form reflection, we derive a finite-time upper bound, prove worst-case tightness, and give a positive lower bound under a falsifiable persistent-harm condition. We further prove an information-theoretic impossibility result: no gate that observes only the generated transcript can improve uniformly over text-indistinguishable environments, whereas an environment-grounded gate can. Motivated by this separation, we introduce Stochastic Reflective Memory Ascent (SRMA), which accepts a candidate memory only after a grounded evaluation risk strictly decreases. Under calibration and non-degenerate corrective mass, SRMA converges exactly, geometrically or polynomially; matching constructions show that both rate regimes are order-tight. We also provide confidence gating for stochastic evaluation and re-anchoring guarantees for piecewise-stationary environments. Experiments instantiate these objects with environment-grounded metrics and test the predicted coordination and drift laws. On 500 SWE-bench instances, the complete Kimi-based system resolves 72.2% versus a 70.8% public mini-SWE-agent reference. Code: https://github.com/YihangChen9/Bilevel-Coordinated-Reflection
Chinese Translation
多智能体大语言模型(LLM)系统通常使用一个编排器(orchestrator)为工作者团队分解任务,然后通过文本反思进行改进。尽管取得了强劲的实证结果,这些系统缺乏对协调、记忆改进以及外部验证作用的统一解释。我们将编排器与工作者的交互建模为一个双层协调博弈:在有界耦合条件下,工作者的局部更新博弈是一个近似势博弈(potential game),其均衡松弛度由任务分解质量控制。随后,我们将反思分析为语义记忆状态上的随机运动。对于自由形式的反思,我们推导了有限时间上界,证明了最坏情况下的紧性,并在一个可证伪的持续性危害条件下给出了正的下界。我们进一步证明了一个信息论不可能性结果:任何仅观测生成文本的门控机制都无法在文本不可区分的环境上实现一致改进,而基于环境的门控则可以。受这一分离结果的启发,我们提出了随机反思记忆上升(Stochastic Reflective Memory Ascent, SRMA),该方法仅在基于环境评估的风险严格下降时才接受候选记忆。在校准和非退化纠正质量的条件下,SRMA 可以精确收敛、几何速率收敛或多项式速率收敛;匹配的构造表明两种速率 regime 均为阶数紧的。我们还为随机评估提供了置信度门控,并为分段平稳环境提供了重锚定保证。实验通过基于环境的指标实例化这些对象,并检验了预测的协调与漂移规律。在 500 个 SWE-bench 实例上,基于 Kimi 的完整系统解决了 72.2% 的问题,而公开的 mini-SWE-agent 参考基线为 70.8%。代码:https://github.com/YihangChen9/Bilevel-Coordinated-Reflection
cs.AI / 54 / 2609.02760

Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents

面向本地部署检索增强工厂智能体的度量驱动子网络选择
Rizeakos, Vasileios, Paisios, Georgios, Machairas, Alexandros, Birbas, Michael, Bachoumis, Athanasios
Abstract
On-premise assistants can give factory workers conversational access to machine documentation, but models capable of the task rarely fit shop-floor hardware. We show that after structural compression and retrieval-grounded adaptation, model size is no longer a reliable predictor of adapted answer quality: general capability falls almost linearly with parameter count, while judged retrieval-augmented answer quality does not. We therefore treat deployment as a post-adaptation selection problem, committing one sub-network per device on judged answer quality and measured on-device throughput under a configurable general-capability floor and memory budget; rules that optimize size, speed, or quality alone each give up capability or throughput. A weight-shared supernetwork trained with sandwich-style in-place distillation keeps this selection inexpensive. In a manufacturing-manual case study, extraction costs 13.7 percent of the unpruned model's judged quality and retrieval-grounded distillation returns it to within 4.6 percent, recovering two thirds of the loss, and the same assistant runs across three heterogeneous edge tiers at 1.3 to 5 watts standby.
Chinese Translation
本地部署的助手可以让工厂工人以对话方式访问机器文档,但能够胜任该任务的模型通常无法适应车间级硬件。我们表明,在经过结构化压缩和检索接地适配之后,模型规模不再是适配后回答质量的可靠预测指标:通用能力随参数量几乎线性下降,而经评判的检索增强回答质量并非如此。因此,我们将部署视为适配后的选择问题,在可配置的通用能力下限和内存预算约束下,基于经评判的回答质量和实测设备端吞吐量,为每台设备确定一个子网络;仅优化规模、速度或质量单一目标的规则均会牺牲能力或吞吐量。采用三明治式原位蒸馏训练的权重共享超网络(supernetwork)使这一选择过程成本低廉。在一个制造手册案例研究中,子网络抽取损失了未剪枝模型评判质量的13.7%,而检索接地蒸馏将其恢复至差距4.6%以内,挽回了三分之二的损失;同一个助手可在三个异构边缘层级上以1.3至5瓦的待机功耗运行。
cs.AI / 55 / 2609.02786

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

SafeEvolve:面向安全对齐的智能体经验驱动的安全框架与策略协同演化
Mao, Qinghua, Qu, Wanying, Guo, Dadi, Yuan, Leitao, Liu, Qingyu, Li, Yu, Chen, Guanxu, Fu, Yanwei, Lin, Xi, Hu, Xia, Liu, Dongrui
Abstract
The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with the environment. This exposes them to safety risks in both harmful final responses and multi-step execution trajectories. Existing safety alignment mechanisms often rely on either external harness updates or policy optimization, yet applying either paradigm in isolation fails to bridge runtime control with intrinsic safety. We propose SafeEvolve, an experience-driven self-evolving framework for agent safety alignment. SafeEvolve leverages safety experience from completed on-policy trajectories to drive a continual loop of harness-policy co-evolution. On the harness side, SafeEvolve converts trajectory-level safety evidence into bounded, component-level updates across safety prompt and hierarchical skills, yielding auditable and reversible harness artifacts. On the policy side, SafeEvolve follows a two-stage SFT-RL paradigm, where harness-use SFT bootstraps the policy to actively leverage evolved harness artifacts, and harness-augmented RL further shapes autonomous safety behaviors during multi-step exploration via verifier-decomposed rewards. Through harness-policy co-evolution, SafeEvolve converts safety experience into an evolved runtime harness and improved policy behavior. Experiments on agentic safety benchmarks show that SafeEvolve achieves a stronger safety-utility tradeoff than existing baselines. For Qwen3.5-4B, SafeEvolve achieves a $3\times$ ASR reduction on AgentDojo while improving benign utility from 59.79% to 61.86%.
Chinese Translation
基于大语言模型(LLM)的智能体性能由基座模型与其与环境交互时使用的安全框架(harness)共同决定。这使其在有害的最终响应和多步执行轨迹两方面都面临安全风险。现有的安全对齐机制通常依赖于外部安全框架更新或策略优化二者之一,然而单独应用任一范式都难以将运行时控制与内在安全性相衔接。我们提出 SafeEvolve,一个经验驱动的智能体安全对齐自我演化框架。SafeEvolve 利用已完成在线(on-policy)轨迹中的安全经验,驱动安全框架与策略的持续协同演化循环。在安全框架侧,SafeEvolve 将轨迹级的安全证据转化为有界的、组件级的更新,涵盖安全提示词(safety prompt)与分层技能(hierarchical skills),从而产生可审计、可回滚的安全框架产物。在策略侧,SafeEvolve 遵循两阶段的 SFT-RL 范式:其中安全框架使用 SFT 引导策略主动利用演化后的安全框架产物,而安全框架增强的强化学习(RL)则通过验证器分解的奖励,在多步探索过程中进一步塑造自主的安全行为。通过安全框架与策略的协同演化,SafeEvolve 将安全经验转化为演化后的运行时安全框架以及更优的策略行为。在智能体安全基准上的实验表明,SafeEvolve 相较现有基线实现了更强的安全性-实用性权衡。对于 Qwen3.5-4B,SafeEvolve 在 AgentDojo 上将攻击成功率(ASR)降低了 3 倍,同时将良性任务效用从 59.79% 提升至 61.86%。
cs.AI / 56 / 2609.02805

Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis

面向电信根因分析(RCA)的大语言模型(LLM):一个基于证据诊断的结构化推理框架
Zhou, Hao, Kulkarni, Mandar, Chen, Hao, Xin, Yan, Charlie, Zhang
Abstract
Root cause analysis (RCA) is a critical task in telecom network operations, but diagnosing performance degradations in modern 5G and emerging 6G networks remains challenging due to complex cross-layer dependencies. While large language models (LLMs) offer promising capabilities for reasoning and knowledge integration, directly applying vanilla LLMs to telecom RCA often leads to hallucination, unstable reasoning, and poor alignment with structured network evidence. This work first reviews the evolution of telecom RCA from rule-based and machine learning (ML) approaches to emerging LLM-enabled techniques, and provides an overview of recent paradigms, including structured reasoning, retrieval-augmented knowledge grounding, agentic orchestration, and verifiable reasoning. Building upon these insights, we propose a structured reasoning framework for LLM-enabled telecom RCA that aligns diagnostic reasoning with telecom-specific evidence and domain knowledge. The proposed approach first organizes heterogeneous network telemetry into canonical contexts, and then enforces decision-path reasoning during diagnosis, and finally generates evidence-grounded explanations for reliable fault identification. Experimental results on two 5G RCA datasets, TeleLogs and TelecomTS, demonstrate that the proposed framework consistently improves diagnostic accuracy and decision consistency compared with baseline techniques. These cross-dataset results highlight the importance of structured reasoning design for practical LLM-based RCA systems in next-generation telecom networks.
Chinese Translation
根因分析(Root Cause Analysis, RCA)是电信网络运营中的一项关键任务,但由于复杂的跨层依赖关系,对现代5G及新兴6G网络中的性能劣化进行诊断仍然具有挑战性。尽管大语言模型(Large Language Models, LLMs)在推理和知识整合方面展现出可观的能力,但直接将原生LLM应用于电信RCA往往会导致幻觉、推理不稳定以及与结构化网络证据对齐不佳等问题。本文首先回顾了电信RCA从基于规则和机器学习(Machine Learning, ML)方法到新兴LLM赋能技术的演进历程,并概述了近期的若干范式,包括结构化推理、检索增强的知识落地、智能体编排以及可验证推理。在此基础上,我们提出了一个面向LLM赋能电信RCA的结构化推理框架,使诊断推理与电信专用证据和领域知识相对齐。该方法首先将异构的网络遥测数据组织为规范化上下文,然后在诊断过程中强制执行决策路径推理,最后生成基于证据的解释以实现可靠的故障识别。在TeleLogs和TelecomTS两个5G RCA数据集上的实验结果表明,与基线技术相比,所提出的框架持续提升了诊断准确率和决策一致性。这些跨数据集的结果凸显了结构化推理设计对于下一代电信网络中实用化基于LLM的RCA系统的重要性。
cs.AI / 57 / 2609.02821

AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application

用于恢复个体与群体层面效应的AI情境化测量:基于调查度量的验证及职业领域应用
Jiang, Wenxin, Wang, Xuyang, Wu, Yuxiao
Abstract
Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We propose AICOME, AI COntextual MEasurement, a framework for evaluating whether AI-derived respondent-level measures can recover individual and group-level effects in contextual models. The key idea is that an AI measure constructed at the respondent level can be used to derive its group-level aggregate and its individual deviation, allowing researchers to estimate both between-group and within-group associations rather than treating AI measurement as response prediction alone. We validate the framework using the 2022 China Family Panel Studies (CFPS), where occupations provide the empirical grouping structure and several job-related survey variables provide validation benchmarks. For computer use, foreign-language use, weekly hours, and management responsibilities, we compare survey measures with AI-derived measures in response-level, model-level, contextual, and boundary-condition validations. The results show that AI contextual measurement can recover much of the contextual-model information contained in observed survey variables when rich respondent and job characteristics are available. Weekly hours provides the strongest validation case, with AI-derived measures reproducing the large negative between- and within-occupation associations with satisfaction observed in CFPS. The framework also identifies clear boundary conditions: performance deteriorates when information is restricted to occupation and basic demographics, and recovery is weaker when several related concepts are treated as simultaneously unobserved. The findings suggest that AICOME is most useful for recovering a limited number of theoretically important constructs from rich existing datasets.
Chinese Translation
研究者日益使用人工智能来构建传统调查中缺失的社会、组织和职业特征测量指标。我们提出AICOME(AI COntextual MEasurement,AI情境化测量),一个用于评估AI衍生的受访者层面测量指标能否在情境模型中恢复个体和群体层面效应的框架。其核心思想是,在受访者层面构建的AI测量指标可用于推导其群体层面的聚合值和个体偏离值,使研究者能够同时估计组间和组内关联,而不是将AI测量仅仅视为响应预测。我们使用2022年中国家庭追踪调查(CFPS)数据对该框架进行验证,其中职业提供了经验分组结构,若干工作相关的调查变量提供了验证基准。针对计算机使用、外语使用、每周工作时长和管理职责,我们在响应层面、模型层面、情境层面和边界条件层面比较了调查测量指标与AI衍生测量指标。结果表明,当可获得丰富的受访者和工作特征信息时,AI情境化测量能够恢复观测调查变量所包含的大部分情境模型信息。每周工作时长提供了最强的验证案例,AI衍生测量指标再现了CFPS中观测到的每周工作时长与满意度之间的大幅负向组间和组内职业关联。该框架还识别出明确的边界条件:当信息仅限于职业和基本人口统计学特征时,表现会下降;当多个相关概念同时被视为不可观测时,恢复效果较弱。研究结果表明,AICOME最适合用于从信息丰富的现有数据集中恢复数量有限的、具有理论重要性的构念。
cs.AI / 58 / 2609.02885

Discriminative World Models for Web Agents

面向网络智能体的判别式世界模型
Li, Kelvin, Pendharkar, Dhruv, Pahilajani, Anish, Shang, Chuyi, Oks, Leon, Karlinsky, Leonid, Feris, Rogerio, Darrell, Trevor, Herzig, Roei
Abstract
Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically trained via supervised next-state prediction to generate fixed representations like HTML or AXTree snapshots. However, this objective is misaligned with the downstream ranker, which relies on predicted states being discriminative across candidates to accurately score them. To address this, we introduce predicted-state matching, a training objective where the predicted representation must distinguish the true resulting state from those reached by alternative actions. We train these models using a branching web-agent dataset derived from WebArena Go-Browse trajectories, where every decision point contains multiple alternative actions and their resulting states. Experiments on our held-out predicted-state matching benchmark show that our approach outperforms world models trained with supervised next-state prediction. We further show that our approach improves PRM-style action ranking on WebPRMBench compared with action-only PRMs and PRMs augmented with supervised-next-state world models. Finally, on WebArena-Lite, using our world model for test-time action selection improves end-to-end task success. Our project page is available at: https://dhruvpendharkar.github.io/dwm/.
Chinese Translation
近期的网络智能体在测试时通过世界模型进行动作选择:采样候选动作、预测所产生的网络状态,并使用排序模型或过程奖励模型(Process Reward Model, PRM)对这些状态进行排序。这些世界模型通常通过有监督的下一状态预测进行训练,以生成固定形式的表示,如HTML或AXTree快照。然而,该训练目标与下游的排序器并不一致,因为排序器依赖于预测状态在候选动作之间具有区分性,才能对其准确打分。为解决这一问题,我们提出了预测状态匹配,一种新的训练目标,要求预测的表示能够将真实达到的状态与由其他替代动作所达到的状态区分开来。我们使用从WebArena Go-Browse轨迹中构建的分支式网络智能体数据集来训练这些模型,其中每个决策点都包含多个替代动作及其对应的结果状态。在我们自建的预测状态匹配基准上的实验表明,我们的方法优于使用有监督下一状态预测训练的世界模型。我们进一步证明,与仅使用动作的PRM以及通过有监督下一状态世界模型增强的PRM相比,我们的方法在WebPRMBench上提升了PRM式的动作排序性能。最后,在WebArena-Lite上,使用我们的世界模型进行测试时动作选择提高了端到端的任务成功率。我们的项目页面见:https://dhruvpendharkar.github.io/dwm/。
机器学习 (Machine Learning)
83
cs.LG / 1 / 2609.01608

WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling

WMLLM:基于“先预测后行动”世界建模的自进化优化智能体
Li, Zhongzheng, Ran, Qingsong, Feng, Shikun, Ran, Nian, Li, Wenhao, Zhang, Xiaoyuan, Wang, Yue, Zhao, Xiaoguang
Abstract
Black-box optimization problems remain challenging because of large, weakly structured, and high-dimensional search spaces. Existing methods often suffer from poor sample efficiency because they rely on direct candidate generation or trial-and-error refinement. A natural way to improve search efficiency is to use world modeling, which can help identify promising optimization directions before costly evaluation. Large language models can predict the outcomes of these candidates with nontrivial accuracy because of their implicit knowledge. Motivated by this observation, we propose WMLLM, a self-evolving optimization-agent framework based on predict-then-act world modeling. The agent first predicts promising directions and then acts to generate candidates. Combined with agentic multi-turn refinement, population-based search, and reinforcement learning, WMLLM refines both its implicit world model and its optimization strategy during search. Experiments on black-box optimization tasks, especially multi-objective molecular optimization, show that WMLLM improves sample efficiency and final optimization performance. On the multi-objective molecular optimization benchmark, WMLLM achieves state-of-the-art results under a limited evaluation budget.
Chinese Translation
黑盒优化问题由于搜索空间庞大、结构松散且维度较高,至今仍具挑战性。现有方法往往依赖直接生成候选解或反复试错式改进,导致样本效率低下。提升搜索效率的一个自然途径是利用世界建模(world modeling),它有助于在代价高昂的评估之前识别有前景的优化方向。凭借其蕴含的隐性知识,大语言模型(LLM)能够以不俗的准确度预测候选解的结果。基于这一观察,我们提出了WMLLM,一个基于“先预测后行动”(predict-then-act)世界建模的自进化优化智能体框架。该智能体首先预测有前景的方向,然后行动以生成候选解。结合智能体式多轮精炼、基于种群的搜索以及强化学习,WMLLM 在搜索过程中不断改进其隐性世界模型和优化策略。在黑盒优化任务上的实验,尤其是多目标分子优化任务,表明 WMLLM 提升了样本效率和最终优化性能。在多目标分子优化基准上,WMLLM 在有限的评估预算下取得了最先进的结果。
cs.LG / 2 / 2609.01609

DiDrive: A Risk-Aware Hierarchical Diffusion Framework for Safe Offline Reinforcement Learning in Autonomous Driving

DiDrive:面向自动驾驶安全离线强化学习的风险感知分层扩散框架
Guo, Qisong, Chen, Jingtang, Chen, Zhilin, Xu, Pei, Fu, Mingjian, Liu, Wenxi, Yu, Yuanlong
Abstract
While diffusion models effectively capture multimodal behavioral priors for autonomous driving, offline reinforcement learning (RL) policies remain susceptible to distribution shift, heavy-tailed risk signals, out-of-distribution (OOD) action generation, and high-dimensional state redundancy. To address these challenges, we propose DiDrive, a distribution-guided offline diffusion framework featuring two synergistic components: the Risk-Aware Hierarchical Diffusion (RHDif) architecture and the 3DICE policy optimization paradigm. In the state space, RHDif utilizes a low-level risk-gated encoder and a high-level contextual modulator to filter environmental redundancy and focus on safety-critical threats. In the action space, 3DICE mitigates OOD overestimation and gradient oscillation through in-sample calibrated guidance, spatiotemporal optimization, and ensemble-based candidate ranking. Evaluations on the CARLA benchmark demonstrate DiDrive's superiority over baselines like IQL, CQL, and Diffusion-QL, particularly in complex, high-density traffic scenarios with 60 vehicles, where it achieves an 85% success rate and a 4295.68 average reward, providing a robust pathway for safe autonomous driving decision-making.
Chinese Translation
尽管扩散模型能够有效捕捉自动驾驶的多模态行为先验,但离线强化学习(RL)策略仍然容易受到分布偏移、重尾风险信号、分布外(OOD)动作生成以及高维状态冗余等问题的影响。为应对这些挑战,我们提出了DiDrive,一种由两个协同组件构成的分布引导离线扩散框架:风险感知分层扩散(Risk-Aware Hierarchical Diffusion, RHDif)架构和3DICE策略优化范式。在状态空间中,RHDif利用底层风险门控编码器和高层上下文调制器来过滤环境冗余信息,并聚焦于安全关键威胁。在动作空间中,3DICE通过样本内校准引导、时空优化以及基于集成的候选排序,缓解了OOD过估计和梯度震荡问题。在CARLA基准上的评估表明,DiDrive优于IQL、CQL和Diffusion-QL等基线方法,尤其是在包含60辆车的高密度复杂交通场景中,其成功率达到85%,平均奖励达到4295.68,为安全的自动驾驶决策提供了一条稳健的路径。
cs.LG / 3 / 2609.01615

Prompt-Space Meta-Learning Does Not Transfer Across Users: A Frozen-LLM Negative Result

提示空间元学习无法跨用户迁移:一项冻结大语言模型的负面结果
Byrne, Liam, Dylan, David, Fitzgerald, Orla, Doyle, Eoin, Nolan, Ciara, Lynch, Padraig, Gallagher, Sinead
Abstract
Personalizing a frozen large language model (LLM) to individual users is often framed as a meta-learning problem in prompt space: each user is a task, and one seeks a shared natural-language adaptation policy that, given a handful of the user's labeled interactions, configures the frozen model for that user. The framing is attractive because it is backbone-agnostic and reuses the machinery of prompt optimization, yet the field rarely tests whether the optimized meta-objective encodes transferable cross-user adaptation rather than generic instruction quality. We study this question with Muse (Meta-learned User-adaptation via Shared Evolution), which evolves a single shared adaptation prompt over a meta-train user population by reflective prompt evolution, freezes it, and applies it zero-shot to held-out users; matched controls isolate learning from confounds of phrasing and selection. On two standard personalization benchmarks (LaMP-2 categorization and LaMP-3 rating) over 200 held-out users each, Muse does not significantly improve on its own un-evolved seed prompt or on a structure-broken control that meta-trains on mismatched user-support pairs, and is dominated by plain few-shot retrieval on the rating task (Delta MAE +0.175, p < 0.001). We attribute these outcomes to a single mechanism, meta-objective collapse: the meta-validation objective is statistically invariant to whether the user-support correspondence is genuine (p=0.555 on LaMP-2, p=0.622 on LaMP-3), so it cannot be optimized into transferable adaptation and instead rewards instruction polish and validation overfitting. The seed-prompt, wrong-support, and invariance-oracle controls form a reusable protocol that separates learned adaptation from these confounds.
Chinese Translation
将冻结的大语言模型(LLM)个性化适配到单个用户,通常被构建为提示空间中的元学习问题:每个用户被视为一个任务,目标是寻找一种共享的自然语言适配策略,在给定用户少量已标注交互的情况下,为该用户配置冻结模型。这一框架颇具吸引力,因为它与主干网络无关,并可复用提示优化的成熟机制,然而该领域很少检验优化后的元目标所编码的究竟是可跨用户迁移的适配能力,还是仅仅泛化的指令质量。我们通过 Muse(基于共享演化的元学习用户适配,Meta-learned User-adaptation via Shared Evolution)来研究这一问题:该方法通过反思式提示演化在元训练用户群体上演化出单一共享的适配提示,将其冻结后零样本应用于留存用户;并通过匹配的对照实验将学习效应与措辞和选择带来的混淆因素分离。在两个标准个性化基准(LaMP-2 分类任务和 LaMP-3 评分任务)上、各涵盖 200 名留存用户的实验中,Muse 相对于其自身未演化的种子提示,以及相对于在错配用户支持对上进行元训练的破坏结构对照方法,均无显著提升;在评分任务上还被简单的少样本检索方法大幅超越(Delta MAE +0.175,p < 0.001)。我们将这些结果归因于单一机制——元目标坍缩(meta-objective collapse):元验证目标在统计上对用户支持对应关系是否真实不变(LaMP-2 上 p=0.555,LaMP-3 上 p=0.622),因此无法被优化为可迁移的适配能力,反而奖励指令修饰和验证集过拟合。种子提示、错误支持和不变性预言机这三种对照构成了一套可复用的实验协议,能够将真正学到的适配能力与上述混淆因素区分开来。
cs.LG / 4 / 2609.01647

Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT

基于SciBERT的面向WASP-2025共享任务的高效上下文受限望远镜文献分类方法
Naidu, Madhusudhana
Abstract
The creation of telescope bibliographies is a crucial part of assessing the scientific impact of observatories and ensuring reproducibility in astronomy. This task involves identifying, categorizing, and linking scientific publications that reference or use specific telescopes. However, this process remains largely manual and resource intensive. In this work, we present an efficient SciBERT-based approach for automatic classification of scientific papers into four categories - science, instrumentation, mention, and not telescope. Despite strict context-length constraints (maximum 512 tokens) and limited compute resources, our approach achieved a macro F1 score of 0.89, ranking at the top of the WASP-2025 leaderboard. We analyze the effect of truncation and show that even with half the samples exceeding the token limit, SciBERT's domain alignment enables robust classification. We discuss trade-offs between truncation, chunking, and long-context models, providing insights into the efficiency frontier for scientific text curation.
Chinese Translation
望远镜文献目录的构建是评估天文台科学影响力并确保天文学研究可重复性的重要环节。该任务涉及识别、分类并关联引用或使用特定望远镜的科学出版物。然而,这一过程目前仍主要依赖人工且资源消耗巨大。在本工作中,我们提出了一种基于SciBERT的高效方法,用于将科学论文自动分类为四个类别——科学、仪器、提及和非望远镜。尽管面临严格的上下文长度限制(最多512个token)和有限的计算资源,我们的方法仍取得了0.89的宏F1分数,在WASP-2025排行榜上名列前茅。我们分析了截断的影响,结果表明即使有一半样本超出token限制,SciBERT的领域适配性仍能实现稳健的分类。我们讨论了截断、分块与长上下文模型之间的权衡,为科学文本整理的效率前沿提供了洞见。
cs.LG / 5 / 2609.01673

CliffRank: A Dual-Branch Framework for Activity-Cliff Ranking Prediction

CliffRank:一种用于活性悬崖排序预测的双分支框架
Li, Kewei, Zhang, Rongying, Yang, Peiyu, Wang, Zhongjian, Zhao, Qiuchen, Huang, Lan, Zhou, Fengfeng
Abstract
Activity-cliff ranking remains difficult because local structural changes can cause large activity differences, while high-quality data that resolve the underlying mechanisms remain limited. To use available activity labels more effectively, we combine absolute-activity regression with ranking-consistency learning. CliffRank trains two parallel predictors with mean squared error, a thresholded listwise loss, and Pairwise Preference Consistency (PPC), which aligns relative ordering in the preference-probability space. On three antimicrobial peptide datasets, CliffRank with ESM2-t12 achieved the highest mean Spearman correlation of 0.5393 and mean Recall@50 of 21.4, although the leading method varied across individual datasets. On three small-molecule datasets, CliffRank with PNA, where PPC was activated after 120 epochs, achieved the highest mean Spearman correlation of 0.6890, while its mean Recall@50 of 30.4 matched that of ACANet-PNA. The PPC results also define its practical limits. Asymmetric initialization improved the MolCLR-GIN averages but did not improve every target. For PNA without pretrained weights, delayed PPC improved selected metrics, but no schedule was best for both mean Spearman correlation and mean Recall@50. Future work should evaluate more targets and antimicrobial peptide systems, develop adaptive PPC schedules, and incorporate protein or membrane context when available.
Chinese Translation
活性悬崖(activity-cliff)排序预测仍然具有挑战性,因为局部的结构变化可能导致较大的活性差异,而能够揭示其潜在机制的高质量数据仍然有限。为了更有效地利用现有的活性标签,我们将绝对活性回归与排序一致性学习相结合。CliffRank 使用均方误差、阈值化列表损失(thresholded listwise loss)以及成对偏好一致性损失(Pairwise Preference Consistency, PPC)来训练两个并行预测器,其中 PPC 在偏好概率空间中对相对排序进行对齐。在三个抗菌肽数据集上,采用 ESM2-t12 的 CliffRank 取得了最高的平均 Spearman 相关系数 0.5393 和平均 Recall@50 为 21.4,尽管在各个数据集上表现最优的方法有所不同。在三个小分子数据集上,采用 PNA 并在 120 个 epoch 后激活 PPC 的 CliffRank 取得了最高的平均 Spearman 相关系数 0.6890,其平均 Recall@50 为 30.4,与 ACANet-PNA 持平。PPC 的实验结果也界定了其实际局限性。非对称初始化提升了 MolCLR-GIN 的平均值,但并非在所有目标上都有改善。对于不使用预训练权重的 PNA,延迟激活 PPC 改善了部分指标,但没有任何一种调度策略能同时在平均 Spearman 相关系数和平均 Recall@50 上表现最优。未来的工作应评估更多的目标分子和抗菌肽体系,开发自适应的 PPC 调度策略,并在数据可用时引入蛋白质或膜环境信息。
cs.LG / 6 / 2609.01676

Sim2Signal: Sim-to-Real Benchmarks for Traffic Signal Control

Sim2Signal:面向交通信号控制的仿真到真实(Sim-to-Real)基准测试
Rafi, Ferdous Al, Mukherjee, Susrik, Dekate, Latika Liladhar, Lavoe, Jennifer Yawa, Yao, Huaiyuan, Mohanty, Shlok, Da, Longchao, Zhou, Xuesong, Wei, Hua
Abstract
Reinforcement learning achieves strong traffic signal control performance in simulation, yet policies trained in simulators often fail once deployed in the real world, a failure known as the Sim-to-Real gap. When RL is applied to traffic signal control, this gap arises from several sources: sensing, action execution, traffic dynamics, and the control objective. Their relative impact and the reliability of existing Sim-to-Real mitigation methods remain insufficiently understood, and the field lacks a standard benchmark for systematically measuring the gap and evaluating mitigation methods. We present Sim2Signal, a benchmark that decomposes the Sim-to-Real gap into observation, action, transition, and reward gaps, corresponding to mismatches in the four components of the underlying MDP, and induces each gap in isolation under a shared protocol. We evaluate 18 mitigation methods on 2 base controllers, across 33 gap settings and 10 calibrated networks built from 5 real-world locations. We find that direct transfer consistently degrades performance across all four gap sources, but the severity of the degradation does not predict the effectiveness of mitigation. Instead, mitigation effectiveness depends strongly on the network and gap setting: outside the action gap, a method that helps in one case may fail in another. The most effective methods generally estimate what the gap changes, rather than make the policy insensitive through domain randomization or invariant representations. Our code is available at https://github.com/Red-Pheonix/Sim2RealTSCBenchMark
Chinese Translation
强化学习在仿真环境中能够取得优异的交通信号控制性能,然而在仿真器中训练的策略一旦部署到真实世界往往失效,这一失败被称为仿真到真实鸿沟(Sim-to-Real gap)。当强化学习应用于交通信号控制时,该鸿沟源于多个方面:感知、动作执行、交通动态以及控制目标。这些因素的相对影响以及现有Sim-to-Real缓解方法的可靠性仍缺乏充分的理解,且该领域缺乏一个系统性地度量该鸿沟并评估缓解方法的标准基准。我们提出了Sim2Signal,该基准将Sim-to-Real鸿沟分解为观测鸿沟、动作鸿沟、状态转移鸿沟和奖励鸿沟,分别对应底层MDP四个组成部分的不匹配,并在共享协议下独立地引入每一种鸿沟。我们在2个基础控制器上、跨33种鸿沟设置和10个基于5个真实世界位置标定构建的交通网络,评估了18种缓解方法。我们发现,直接迁移在所有四种鸿沟来源下均会导致性能持续下降,但性能下降的严重程度并不能预测缓解方法的有效性。相反,缓解方法的有效性强烈依赖于交通网络和鸿沟设置:除动作鸿沟外,某种方法在一种情况下有效,在另一种情况下可能失效。最有效的方法通常是对鸿沟所引起的改变进行估计,而非通过域随机化或不变表示使策略不敏感。我们的代码可在 https://github.com/Red-Pheonix/Sim2RealTSCBenchMark 获取。
cs.LG / 7 / 2609.01679

A Survey on Self-Improving Test-Time Intelligence: Feedback-Driven Adapting, Learning, and Scaling at Inference

自我改进的测试时智能综述:推理阶段的反馈驱动自适应、学习与扩展
Niu, Shuaicheng, Chen, Guohao, Chen, Yaofo, Wen, Zhiquan, Hu, Jinwu, Deng, Zeshuai, Chen, Deyu, Zhang, Shuhai, Chen, Renjie, Lian, Zihao, Xu, Shoukai, Dai, Gang, Zhang, Yunbei, Luo, Wei, Zhang, Yifan, Tan, Mingkui, Deng, Cheng
Abstract
The ability of AI systems to improve their behavior during deployment is becoming increasingly important. As inference moves beyond the static execution of a fixed trained model, a growing body of work studies how models can refine their behavior on the fly by exploiting test-time information and additional computation. These developments have largely evolved along two directions: methods that modify the model's state using test-time signals, and methods that improve predictions through extra inference-time resources such as more sampling and tool use. However, these directions are often studied in separate communities with different terminology, making their connections harder to see. In this survey, we present feedback-driven Test-Time Intelligence (TTI) as a unified perspective for understanding such deployment-time improvement. We use this view to relate test-time adaptation, test-time learning, and test-time scaling, highlighting both their distinctions and their growing overlap in hybrid systems. This unified framework helps connect previously fragmented ideas and provides a clearer conceptual foundation for studying inference-time self-improvement. We review major methodological paradigms, representative applications, and open challenges across vision, language, multimodal learning, generative models, robotics, and healthcare. Our goal is to provide a coherent foundation and research roadmap for the study of self-improving AI systems at test time.
Chinese Translation
AI系统在部署过程中改进自身行为的能力正变得日益重要。随着推理超越固定已训练模型的静态执行,越来越多的研究工作探讨模型如何利用测试时信息和额外计算来即时优化其行为。这些发展主要沿两个方向演进:一类方法利用测试时信号修改模型的状态,另一类方法通过额外的推理时资源(如更多采样和工具使用)来改进预测。然而,这两个方向通常由不同的研究社区使用不同的术语分别研究,导致它们之间的联系难以被察觉。在本综述中,我们提出反馈驱动的测试时智能作为理解此类部署时改进的统一视角。我们借助这一视角将测试时自适应、测试时学习和测试时扩展联系起来,既阐明它们之间的区别,也指出它们在混合系统中日益增长的重叠。这一统一框架有助于连接此前碎片化的思想,并为研究推理时自我改进提供更清晰的概念基础。我们回顾了视觉、语言、多模态学习、生成模型、机器人和医疗健康等领域的主要方法学范式、代表性应用和开放挑战。我们的目标是为测试时自我改进AI系统的研究提供一个连贯的基础和研究路线图。
cs.LG / 8 / 2609.01680

Reinforcement Learning and Rule-Based Peer-to-Peer Pricing in Residential PV-BES Communities

住宅光伏-电池储能社区中的强化学习与基于规则的点对点定价机制
Benalcazar, Pablo, Kalka, Maciej, Guamán, Wilian, Kamiński, Jacek
Abstract
This paper compares rule-based and learning-based pricing mechanisms for peer-to-peer (P2P) electricity trading in residential photovoltaic communities. The rule-based benchmarks comprise bill-sharing as an ex post allocation mechanism, the mid-market rate, and supply-demand-ratio pricing. The reinforcement-learning (RL) formulation is implemented through a Deep Q-Network and evaluated under multiplier-based and learnable SDR-shaped pricing, with a fixed-parameter SDR variant as a non-learning control. Performance is assessed through community savings together with complementary financial and operational indicators. In the base PV-only configuration, the rule-based benchmarks outperform the best RL policy. With battery energy storage, evaluated for the RL policies only, community savings under the best RL policy increase from EUR 734.23 to EUR 978.52. Across the learning-based modes and in both configurations, SDR-shaped pricing outperforms the multiplier-based parameterization considered. The results indicate that rule-based pricing remains highly competitive wherever the two families are compared directly, and that storage substantially improves the learning-based outcomes under this accounting, while the distribution of benefits remains heterogeneous across households.
Chinese Translation
本文比较了住宅光伏社区中用于点对点(P2P)电力交易的基于规则和基于学习的定价机制。基于规则的基准方法包括作为事后分配机制的账单分摊(bill-sharing)、中间市场汇率定价以及供需比(SDR)定价。强化学习(RL)方案通过深度Q网络(Deep Q-Network)实现,并在基于乘子的定价和可学习的SDR整形定价下进行评估,同时以固定参数的SDR变体作为非学习对照。性能评估通过社区节支额以及补充性的财务和运营指标进行。在仅含光伏的基础配置中,基于规则的基准方法优于最优的RL策略。在配置电池储能(仅针对RL策略进行评估)的情况下,最优RL策略下的社区节支额从734.23欧元增加到978.52欧元。在基于学习的各种模式以及两种配置下,所考虑的SDR整形定价均优于基于乘子的参数化方法。结果表明,在直接比较这两类方法时,基于规则的定价仍然具有高度竞争力;在这种核算方式下,储能显著改善了基于学习的方法的效果,而收益在各家庭之间的分配仍然是不均等的。
cs.LG / 9 / 2609.01689

Median-of-Means as an Extremal Convex Estimator and a Nonconvex Route to the Trimmed Oracle

作为极值凸估计器的中位数均值法及通往截尾预言机(Trimmed Oracle)的非凸路径
Majumdar, Angshul
Abstract
We revisit median-of-means estimation from a deterministic optimization viewpoint and develop a family of block-Lp estimators for robust learning with heavy-tailed and adversarially corrupted data. In a block contamination model with at least a fraction 1 minus epsilon of good blocks, we first show that every convex block M-estimator has worst-case robustness constant at least 1 divided by 1 minus 2 epsilon. This matches the classical median-of-means bound and proves that the trimmed-block oracle constant 1 divided by 1 minus epsilon cannot be attained within the convex class. We then introduce a nonconvex block-Lp family for p between 0 and 1 and derive finite-sample deterministic robustness bounds for all global minimizers. As p decreases from 1 toward 0, these bounds continuously approach the trimmed-block oracle constant. For sufficiently small p, the global minimizers coincide with those of the oracle under a mild separation condition. We also show that the block-Lp objectives have a benign landscape, with all local minima remaining close to the truth and no bad basins. Combining these results with block-level concentration yields sub-Gaussian deviation bounds under finite 2 plus delta moments and high-dimensional extensions to robust mean estimation and sparse regression.
Chinese Translation
我们从确定性优化的视角重新审视中位数均值(median-of-means)估计,并针对重尾数据和对抗性污染数据下的鲁棒学习,提出了一族分块-Lp(block-Lp)估计器。在至少有比例 1-ε 的干净分块的分块污染模型中,我们首先证明每一个凸分块 M-估计器的最坏情况鲁棒性常数至少为 1/(1-2ε)。这与经典的中位数均值界相匹配,并证明了截尾分块预言机常数 1/(1-ε) 在凸类范围内无法达到。随后,我们引入了 p∈(0,1) 的非凸分块-Lp 族,并为所有全局最小值点推导了有限样本确定性鲁棒界。当 p 从 1 递减趋于 0 时,这些界连续地逼近截尾分块预言机常数。对于足够小的 p,在温和的分离条件下,全局最小值点与预言机的解相重合。我们还证明了分块-Lp 目标函数具有良性景观:所有局部极小值都保持接近真值,且不存在坏的盆地。将这些结果与分块级集中性相结合,可在有限 2+δ 阶矩条件下得到次高斯偏差界,并推广到鲁棒均值估计和稀疏回归的高维情形。
cs.LG / 10 / 2609.01699

Tri-Band Channel Measurement-Enabled Multi-Layer Digital Twin for Terahertz Wireless Data Centers

面向太赫兹无线数据中心的三频段信道测量赋能的多层孪生系统
Zhu, Mingjie, Yu, Ziming, Wang, Guangjian, Han, Chong
Abstract
The rapid growth of AI computing has driven increasing demands for flexible and high-capacity data-center interconnections. Owing to its ultra-wide bandwidth and high spatial reuse capability, terahertz (THz) communication has emerged as a promising solution for future wireless data centers, while digital twins (DTs) enable efficient wireless planning and real-time optimization. In this work, a measurement-driven multi-layer DT framework is proposed for THz wireless data centers, where the physical, channel, evaluation, and manipulation layers are progressively constructed from bottom to top. First, extensive channel measurements are conducted at 140, 220, and 300 GHz to characterize frequency-dependent propagation behaviors. Based on the tri-band measurements, a measurement-calibrated physical twin is established by jointly optimizing the geometry, material, antenna, and hybrid propagation models. On top of the physical twin, a line-of-sight (LoS)-aware implicit neural field is developed to construct an AI channel twin for efficient channel reconstruction. The proposed AI twin learns location-dependent channel statistics from the calibrated twin, enabling real-time prediction of received power and LoS probability. Building upon the reconstructed channel field, a system-level evaluation layer is derived to analyze coverage and interference for both AP-to-rack and rack-to-rack communications. Experimental results show that the proposed AI twin achieves lower power reconstruction error than existing neural-field baselines while maintaining real-time inference capability. Moreover, the ceiling-mounted AP deployment achieves over 90% coverage under a 10 dB signal-to-interference-plus-noise ratio (SINR) threshold, demonstrating the effectiveness of the proposed DT framework for THz wireless data-center planning and optimization.
Chinese Translation
人工智能计算的快速增长推动了对灵活、大容量数据中心互连的需求。太赫兹(THz)通信凭借其超宽带宽和高度空间复用能力,已成为未来无线数据中心的一种有前景的解决方案,而数字孪生技术则能够实现高效的无线规划与实时优化。本文提出了一种面向太赫兹无线数据中心的测量驱动的多层数字孪生框架,自底向上依次构建物理层、信道层、评估层和操控层。首先,在140、220和300 GHz频段开展了大量信道测量,以刻画随频率变化的传播特性。基于三频段测量结果,通过联合优化几何、材料、天线及混合传播模型,建立了经测量校准的物理孪生。在物理孪生之上,开发了一种视距(LoS)感知的隐式神经场,以构建AI信道孪生,实现高效的信道重建。所提出的AI孪生从校准孪生中学习与位置相关的信道统计特性,从而能够实时预测接收功率和视距概率。在重建的信道场基础上,进一步构建了系统级评估层,用于分析接入点(AP)到机架以及机架到机架通信的覆盖与干扰情况。实验结果表明,所提出的AI孪生在保持实时推理能力的同时,相比现有神经场基线方法取得了更低的功率重建误差。此外,采用吸顶式AP部署方案,在10 dB信干噪比(SINR)门限下可实现超过90%的覆盖率,验证了所提出数字孪生框架在太赫兹无线数据中心规划与优化中的有效性。
cs.LG / 11 / 2609.01705

Generative Diffusion Surrogates with Analytical Variance Schedule

具有解析方差调度的生成式扩散代理模型
Reichherzer, Patrick, Gregori, Gianluca, Hosking, David N., Sarkar, Subir
Abstract
Stochastic transport describes physical systems in which an initially structured distribution spreads under unresolved forcing, scattering, or heterogeneous media. Useful surrogates for such systems should be probabilistic, time-resolved, and able to represent non-Gaussian distributional structure. Generative diffusion models, which corrupt data with Gaussian noise and learn a reverse flow back to structured states, have these properties. Their noise schedules, however, are usually chosen heuristically: image and audio generation---the canonical use cases---provide no physical clock. In transport, by contrast, the variance, or mean-square displacement, is often known from macroscopic theory or empirical scaling even when the full distribution is not. Here we prescribe the forward noising rate as the time derivative of this variance, turning generative time into a calibrated transport clock. The variance path is enforced by construction, while the learned score field represents how non-Gaussian structure inherited from entrance data is smoothed along that path, requiring no intermediate-time physical transport data. For ballistic-to-diffusive transport in turbulent plasmas, the surrogate matches test-particle distributions, reproduces the laboratory-measured variance scale, and tracks the simulated kurtosis evolution without schedule tuning, enabling calibrated emulation and likelihood-based inference.
Chinese Translation
随机输运描述了初始呈结构化分布的物理系统在未解析的强迫、散射或非均匀介质作用下发生扩散的过程。此类系统有效的代理模型应当是概率性的、时间分辨的,并能表示非高斯分布结构。生成式扩散模型以高斯噪声破坏数据并学习一个反向流以恢复到结构化状态,具备上述性质。然而,其噪声调度通常是启发式选择的:图像与音频生成——这一典型应用场景——不提供物理时钟。相比之下,在输运问题中,方差(即均方位移)往往可以从宏观理论或经验标度律中获得,即使完整分布未知。本文将该方差的时间导数规定为前向加噪速率,从而将生成时间转化为经过校准的输运时钟。方差路径由构造方式强制保证,而学习到的得分场(score field)则表示从入口数据继承的非高斯结构沿该路径被平滑的方式,无需任何中间时刻的物理输运数据。对于湍流等离子体中从弹道到扩散的输运过程,该代理模型能够匹配测试粒子分布,复现实验室测得的方差标度,并在无需调度调节的情况下跟踪模拟的峰度演化,从而实现经过校准的仿真模拟与基于似然的推断。
cs.LG / 12 / 2609.01729

RecKAN: Kolmogorov-Arnold Networks with a Learnable Recursive Polynomial Basis

RecKAN:具有可学习递归多项式基的Kolmogorov-Arnold网络
Azarpour, Amirhosein
Abstract
Kolmogorov--Arnold Networks (KANs) replace the fixed scalar weights of a standard network with learnable univariate functions on each edge, but existing variants still fix the \emph{basis} that those functions are built from: B-splines, Chebyshev polynomials, wavelets, or Jacobi polynomials, and learn only the combination weights over it. We introduce RecKAN, which instead defines the basis itself by a second order polynomial recurrence, $R_{n+1}(x) = (ax^2+bx+c)R_n(x) + (dx+e)R_{n-1}(x)$, whose five coefficients are learned jointly with the network. We show this recurrence recovers several classical polynomial families including both kinds of Chebyshev polynomials, Fibonacci, Pell, and Jacobsthal polynomials as special cases, and prove that its degree grows linearly in $n$ exactly on the sub-family containing all of them, giving a concrete sense in which the learned basis can move beyond any fixed classical choice. Across multiple benchmark datasets spanning image, text, biomedical time series classification, and time series forecasting, RecKAN outperforms three parameter-matched KAN baselines (Chebyshev, Jacobi, and spline based) on all classification tasks and achieves the lowest MSE on the ETTh1 forecasting benchmark. Additionally, when used as a classifier head with a convolutional backbone, RecKAN achieves higher accuracy than standard MLP heads on Fashion MNIST, CIFAR-10, and SVHN. On a synthetic function fitting benchmark it tracks a sharply oscillatory target that a parameter comparable MLP under fits. We further show that the learned recurrence coefficients are interpretable: on the task requiring the most local structure, training moves the basis away from the linear degree growth regime that contains every classical family we identify, consistent with our theoretical analysis of what that structural shift enables.
Chinese Translation
Kolmogorov-Arnold网络(KAN)用每条边上可学习的单变量函数替代标准网络中的固定标量权重,但现有变体仍然固定这些函数所依赖的基(basis):如B样条、切比雪夫多项式、小波或雅可比多项式,仅学习其上的组合权重。我们提出RecKAN,它通过一个二阶多项式递推关系 $R_{n+1}(x) = (ax^2+bx+c)R_n(x) + (dx+e)R_{n-1}(x)$ 来定义基本身,其五个系数与网络联合学习。我们证明该递推关系可以恢复若干经典多项式族,包括两类切比雪夫多项式、Fibonacci多项式、Pell多项式和Jacobsthal多项式作为特例,并证明其次度恰好在这些多项式所属的子族中随 $n$ 线性增长,这具体表明学习到的基可以超越任何固定的经典选择。在涵盖图像、文本、生物医学时间序列分类和时间序列预测的多个基准数据集上,RecKAN在所有分类任务中均优于三个参数匹配的KAN基线(基于切比雪夫、雅可比和样条的方法),并在ETTh1预测基准上取得了最低的MSE。此外,当作为卷积骨干网络的分类头使用时,RecKAN在Fashion MNIST、CIFAR-10和SVHN上的准确率高于标准MLP分类头。在一个合成函数拟合基准上,它能够拟合参数相当的MLP欠拟合的剧烈振荡目标函数。我们进一步表明,学习到的递推系数是可解释的:在最需要局部结构的任务上,训练会使基偏离包含我们所识别的所有经典多项式族的线性次数增长区间,这与我们对这种结构转变所实现功能的理论分析相一致。
cs.LG / 13 / 2609.01746

CAT-Flow: Curvature-Adaptive sTeps for Flow Matching

CAT-Flow:面向流匹配的曲率自适应步长方法
Li, Qinchan, Cisneros-Velarde, Pedro, Fu, Keru, Miranda, Samuel Antunes, Vaswani, Sharan, Zhang, Hao
Abstract
Flow Matching has emerged as a leading framework for generative modeling, powering state-of-the-art systems such as FLUX and Stable Diffusion 3.5. However, the iterative nature of its ODE-based sampling process creates a fundamental efficiency bottleneck: the quality of generated samples is highly sensitive to the choice of step-sizes, and current models typically require 20 to 30 steps for good quality. In this work, we propose two lightweight, training-free algorithms, CAT-OV and CAT-OT that adapt step-sizes at inference time based on a novel connection between Flow Matching sampling and gradient flow. Our algorithms are computed efficiently by not requiring additional neural function evaluations. Specifically, CAT-OT estimates curvature over time via a finite-difference approximation of the time-derivative of the vector field, while CAT-OV approximates curvature over the state space via a gradient of the vector field. Under suitable conditions, both methods have truncation error bounds of constant order. Empirically, CAT-OV and CAT-OT outperform existing step-size heuristics in image quality metrics across four text- to-image Flow Matching models, reducing the number of generation steps required to reach comparable quality by up to 40%.
Chinese Translation
流匹配(Flow Matching)已成为生成建模的主流框架,支撑着FLUX和Stable Diffusion 3.5等最先进的系统。然而,其基于ODE的采样过程的迭代特性造成了根本性的效率瓶颈:生成样本的质量对步长的选择高度敏感,且现有模型通常需要20到30步才能获得良好的生成质量。在本工作中,我们提出了两种轻量级的免训练算法——CAT-OV和CAT-OT,它们基于流匹配采样与梯度流之间的一种新联系,在推理时自适应地调整步长。我们的算法无需额外的神经网络函数求值,因而计算高效。具体而言,CAT-OT通过对向量场时间导数的有限差分近似来估计时间维度上的曲率,而CAT-OV则通过向量场的梯度来近似状态空间上的曲率。在适当条件下,两种方法的截断误差界均为常数阶。实验表明,在四个文生图流匹配模型上,CAT-OV和CAT-OT在图像质量指标上优于现有的步长启发式方法,在达到相当质量的前提下,最多可将所需生成步数减少40%。
cs.LG / 14 / 2609.01756

A Study of Conditional Diffusion Models for Open-Loop Control under Dry Friction and Stiction

干摩擦与静摩擦条件下开环控制的条件扩散模型研究
Antonelo, Eric Aislan
Abstract
Diffusion models have recently emerged as expressive generative priors for planning and control. This paper studies Action Diffusion, an action-sequence diffusion formulation used as an open-loop proposal distribution for a point-mass system with dry friction and stiction. In this benchmark, motion starts only when the applied input exceeds a static-friction threshold, so effective controls occupy a small and temporally structured subset of the action-sequence space. A compact conditional 1D U-Net generates bounded control sequences conditioned on initial and target states. We compare it with uniform random shooting, random shooting from the same structured dataset prior, and the Cross-Entropy Method (CEM). Results show that Action Diffusion reduces terminal error and stuck steps, especially in low-sample regimes. These results indicate that conditional diffusion provides an effective mechanism for generating temporally coherent control sequences that overcome stiction by conditioning and recombining structured control primitives from the training prior for state-to-state open-loop control.
Chinese Translation
扩散模型近年来作为用于规划与控制的高表现力生成先验而兴起。本文研究了动作扩散(Action Diffusion),这是一种动作序列扩散建模方法,被用作具有干摩擦和静摩擦的质点系统的开环提议分布。在该基准问题中,只有当施加的输入超过静摩擦阈值时运动才会开始,因此有效控制仅占据动作序列空间中一个小且具有时间结构的子集。一个紧凑的条件一维U-Net以初始状态和目标状态为条件,生成有界的控制序列。我们将其与均匀随机射击、来自同一结构化数据集先验的随机射击以及交叉熵方法(CEM)进行了比较。结果表明,动作扩散降低了终端误差和卡滞步数,尤其是在低样本条件下。这些结果表明,条件扩散提供了一种有效机制,可通过条件化和重组来自训练先验的结构化控制基元,生成时间上连贯的控制序列,从而克服静摩擦,实现状态到状态的开环控制。
cs.LG / 15 / 2609.01765

Toward Explainable and Policy-Aware AI for Carbon Credit Price Prediction: A Research Framework for Emerging Carbon Markets

迈向可解释且政策感知的碳信用价格预测人工智能:面向新兴碳市场的研究框架
Begum, Summaiya Unnisa, Ullah, Mohammed Nadeem, Khan, Mohammed Abdul Ghani
Abstract
Carbon markets put a price on emissions, yet that price remains hard to forecast. Work in this area clusters on the EU and Chinese schemes, compresses regulatory text into a sentiment score, and reports accuracy without calibration or explanation stability. We distil ten recurring gaps into an impact-feasibility matrix and propose EPA-CarbonNet, a six-layer architecture that fuses market series with policy text by cross-attention and calibrated intervals alongside policy-attributed explanations. We then build and test it on eleven years of daily S and P carbon index data. The findings are largely negative, and reported as measured: a random walk beats the model on five-day RMSE (0.0365 against 0.0475), SHAP rankings agree at rho = 0.54 across resampled backgrounds, and policy attention never coincides with documented regulatory events. Directional accuracy, at 58.6 percent, leads every baseline. Code, data documentation and all result artifacts are available at https://github.com/Kimalice/Toward-Explainable-and-Policy-Aware-AI-for-Carbon-Credit-Price-Prediction
Chinese Translation
碳市场为排放定价,但这一价格仍难以预测。该领域的研究集中于欧盟和中国的碳交易体系,将监管文本压缩为情感分数,并且只报告精度而不考虑校准或解释稳定性。我们将十个反复出现的研究空白提炼为一个“影响力—可行性”矩阵,并提出EPA-CarbonNet——一种六层架构,通过交叉注意力机制融合市场时间序列与政策文本,同时提供校准区间以及归因于政策的解释。随后,我们基于十一年的标准普尔(S&P)碳指数每日数据构建并测试了该模型。研究结果总体上呈负面,且如实报告:在五日RMSE上,随机游走优于该模型(0.0365对比0.0475);SHAP排序在不同重采样背景下的秩相关系数仅为rho = 0.54;政策注意力从未与已记录的监管事件相吻合。方向准确率为58.6%,优于所有基线模型。代码、数据文档及所有结果产物均可在 https://github.com/Kimalice/Toward-Explainable-and-Policy-Aware-AI-for-Carbon-Credit-Price-Prediction 获取。
cs.LG / 16 / 2609.01768

Emergence of Fibrations, Compression, and Symmetry Breaking in Artificial Neural Networks

人工神经网络中纤维化、压缩与对称性破缺的涌现
Velarde, Osvaldo M, Parra, Lucas C, Hashemi, Alireza, Makse, Hernan A
Abstract
Artificial neural networks are often regarded as powerful yet opaque black boxes. Here, we demonstrate that learning in deep neural networks generates local symmetries known in graph theory as fibrations and coverings. We prove that covering symmetries are stable attractors of stochastic gradient descent. Consistent with this theory, we report the emergence of covering symmetries across major network architectures, including multilayer, convolutional, recurrent, and transformer networks. Exploiting these symmetries enables drastic model compression - reducing networks to 17% of their original size without sacrificing performance. Furthermore, controlled breaking of covering symmetry overcomes the loss of plasticity, achieving state-of-the-art performance in continual learning. The theoretical results provide a new foundation for AI systems based on symmetries that convert black boxes into interpretable colored graphs and enable more efficient inference and lifelong learning.
Chinese Translation
人工神经网络通常被视为强大却不透明的黑箱。本文证明,深度神经网络中的学习会产生图论中称为纤维化(fibrations)与覆盖(coverings)的局部对称性。我们证明了覆盖对称性是随机梯度下降的稳定吸引子。与该理论一致,我们报告了覆盖对称性在多种主要网络架构中的涌现,包括多层网络、卷积网络、循环网络以及Transformer网络。利用这些对称性可以实现大幅度的模型压缩——将网络缩减至原始大小的17%而不损失性能。此外,受控地打破覆盖对称性可以克服可塑性丧失问题,在持续学习中实现了最先进的性能。这些理论结果为基于对称性的人工智能系统提供了新的基础,可将黑箱转化为可解释的着色图,并支持更高效的推理与终身学习。
cs.LG / 17 / 2609.01802

D-FROST: Decentralized Federated pRompt-tuning via Optimal tranSporT for Non-IID and Imbalanced Data

D-FROST:基于最优传输的非独立同分布与非均衡数据的去中心化联邦提示微调
Nguyen, Quan Minh, Ngo, Hoang M., Hoang, Trong Nghia, Thai, My T.
Abstract
Prompt tuning provides a parameter-efficient way to adapt foundation models (FMs) by freezing the pretrained backbone and updating only a small set of learnable prompts. This property makes prompt tuning especially suitable for decentralized federated learning (DFL), where exchanging full-model updates can be prohibitively expensive. However, prompt tuning in DFL introduces new challenges. Prompt sets learned from heterogeneous local data may not be index-wise aligned, making standard decentralized averaging unsuitable. In addition, the algorithm should be theoretically guaranteed to achieve consensus and make progress toward the shared objective. In this work, we provide the first study of prompt tuning in DFL. We formulate decentralized prompt tuning as a Wasserstein-based optimization problem over prompt measures, which captures the set-valued structure of prompts. We then propose D-FROST, an optimal-transport-based (OT-based) decentralized prompt-tuning algorithm that merges neighborhood prompts into compact representative prompt sets through transportation-based matching. We further analyze D-FROST by bounding the Wasserstein consensus error across clients, and establishing convergence of the network-level prompt barycenter to a neighborhood of stationarity. Experiments under heterogeneous client data demonstrate the effectiveness of D-FROST for decentralized prompt tuning.
Chinese Translation
提示微调(Prompt tuning)提供了一种参数高效的方式来适配基础模型(FMs):冻结预训练的主干网络,仅更新少量可学习的提示。这一特性使提示微调特别适合去中心化联邦学习(DFL),因为在DFL中交换完整模型的更新开销可能过高。然而,在DFL中进行提示微调带来了新的挑战:从异构的本地数据中学到的提示集合可能在索引层面不对齐,使得标准的去中心化平均方法不再适用;此外,算法应当在理论上保证达成共识并向共享目标取得进展。在本工作中,我们首次对DFL中的提示微调进行了研究。我们将去中心化提示微调形式化为一个基于Wasserstein距离、作用于提示测度上的优化问题,该问题刻画了提示的集合值结构。随后,我们提出了D-FROST,一种基于最优传输(OT)的去中心化提示微调算法,通过基于传输的匹配将邻居的提示合并为紧凑的代表性提示集合。我们进一步对D-FROST进行了分析:界定了客户端之间的Wasserstein共识误差,并证明了网络层面的提示重心收敛到驻点的一个邻域。在异构客户端数据上的实验证明了D-FROST在去中心化提示微调中的有效性。
cs.LG / 18 / 2609.01807

SPD: Single Pass Decoding for Generative Reranking

SPD:面向生成式重排序的单次前向解码方法
Laftchiev, Emil, Agrawal, Prachi, Kayali, Moe, Yan, Bixing, Xu, Qi, Lei, Zijie, Qiu, Chen, Hua, Zhi, Li, Ke, Simon, Luke
Abstract
Large language models (LLMs) achieve state-of-the-art generative ranking quality, but the ranking they produce must be decoded, and autoregressive decoding spends one sequential forward pass per emitted token. We observe that the only tokens a ranker must emit are the $N$ ordinal values naming the items in ranked order, and that this narrow, permutation-structured output format admits decoding strategies which are much more efficient than left-to-right generation. We introduce SPD (Single Forward Pass), a format-specialized decoding strategy that decodes all $N$ ordinals in $O(1)$ forward passes. SPD reads an $N \times K$ item-position score matrix off the LLM's prefill hidden states with a lightweight self-attention head, then decodes the ordinals as the optimal bipartite assignment of that matrix via the Hungarian algorithm, yielding a valid permutation by construction rather than by repair. Through a systematic study of training signals and backbone adaptation, we show that LoRA-based fine-tuning combined with auto-regressive LLM ranking distillation reaches 28 ms end-to-end inference, a speed-up of 64x while maintaining ranking quality on par with the teacher. We provide a complete ablation decomposing the contributions of architecture, training signal, and backbone adaptation. Our framework connects generative ranking to combinatorial optimization, opening a path toward other $O(1)$-decode mechanisms for real-time ranking.
Chinese Translation
大语言模型(LLM)实现了最先进的生成式排序质量,但其产生的排序结果必须经过解码,而自回归解码每生成一个标记就需要一次串行前向传播。我们观察到,排序器唯一必须输出的标记是按排序顺序命名各条目的$N$个序数值,而这种狭窄的、具有置换结构的输出格式可以采用远比从左到右生成更高效的解码策略。我们提出了SPD(Single Forward Pass,单次前向传播),一种针对特定格式优化的解码策略,能够以$O(1)$次前向传播解码全部$N$个序数值。SPD利用一个轻量级自注意力头从LLM预填充(prefill)阶段的隐藏状态中读取$N imes K$的条目-位置得分矩阵,然后通过匈牙利算法将该矩阵的最优二分图指派结果解码为序数值,从而在构造上而非通过修复得到有效的排列。通过对训练信号和骨干模型适配的系统性研究,我们表明基于LoRA的微调结合自回归LLM排序蒸馏可实现28毫秒的端到端推理,速度提升64倍,同时保持与教师模型相当的排序质量。我们提供了完整的消融实验,分解了架构、训练信号和骨干模型适配各自的贡献。我们的框架将生成式排序与组合优化相联系,为实时排序的其他$O(1)$解码机制开辟了道路。
cs.LG / 19 / 2609.01839

Import What You Need: Learning When and How to Augment EHR Graphs with External Knowledge

按需导入:学习何时以及如何利用外部知识增强电子健康记录图
Chen, Chen, Kerdabadi, Mohsen Nayebi, Wang, Dongjie, Liu, Mei, Yao, Zijun
Abstract
Longitudinal prediction from electronic health records (EHRs) is limited by the sparsity and irregularity in patient trajectories, and knowledge augmentation with external knowledge graphs (KGs) offers a promising way to alleviate these issues. However, most existing methods perform fixed, context-agnostic topology augmentation by adding the same KG nodes and edges regardless of a patient's evolving state. We propose ReTA, a Reinforcement learning-based dynamic Topology Augmentation framework that casts KG import as a per-visit, budget-aware policy. ReTA first constructs an offline refined pool of KG-grounded templates, then learns a policy to select one augment action per visit from three options: Soft Import, which enriches node features without modifying graph topology, Hard Import, which grafts a compact KG subgraph onto the visit graph to create message-passing shortcuts, and Skip, which leaves the visit unaugmented when the base encoder is already confident. To stabilize learning, ReTA employs a decoupled encoder that processes semantic and structural signals in separate channels and fuses them via adaptive gating. Experiments on MIMIC-III and MIMIC-IV across diagnosis prediction, mortality, and readmission show that ReTA consistently outperforms strong baselines while remaining efficient, transfers across datasets and knowledge graphs, and yields interpretable augmentation patterns. The robust gains under sparse supervision highlight the advantage of ReTA's dynamic decision to import knowledge, boosting accuracy while curbing costs.
Chinese Translation
基于电子健康记录(EHR)的纵向预测受到患者轨迹稀疏性和不规则性的限制,而利用外部知识图谱(KG)进行知识增强为缓解这些问题提供了一种有前景的途径。然而,现有的大多数方法采用固定的、与上下文无关的拓扑增强方式,即无论患者的状态如何演变,都添加相同的KG节点和边。我们提出了ReTA,一种基于强化学习的动态拓扑增强框架,它将KG导入建模为每次就诊的、具有预算约束的策略。ReTA首先构建一个离线精炼的、以KG为基础的模板池,然后学习一个策略,针对每次就诊从三种选项中选择一种增强动作:软导入(Soft Import),在不修改图拓扑的情况下丰富节点特征;硬导入(Hard Import),将一个紧凑的KG子图嫁接到就诊图上以创建消息传递捷径;以及跳过(Skip),当基础编码器已经足够确信时,不对该次就诊进行增强。为了稳定学习,ReTA采用了一个解耦的编码器,将语义和结构信号分别在独立的通道中处理,并通过自适应门控进行融合。在MIMIC-III和MIMIC-IV数据集上,针对诊断预测、死亡率和再入院任务的实验表明,ReTA在保持高效的同时,始终优于强大的基线方法,能够跨数据集和知识图谱迁移,并产生可解释的增强模式。在稀疏监督下取得的稳健增益凸显了ReTA动态导入知识决策的优势,在控制成本的同时提升了准确性。
cs.LG / 20 / 2609.01896

OutageDiT: A Generative Foundation Model for Power Outage Forecasting and Scenario Simulation

OutageDiT:面向停电预测与情景模拟的生成式基础模型
Zhu, Yunqin, Qiu, Feng, Xie, Yao
Abstract
Power-outage planning requires scenarios before an event occurs. These scenarios must represent uncertainty in magnitude, timing, and duration while preserving temporal dependence. However, severe events are rare, and data from any single region contain few examples of extreme outage and restoration patterns. To address this challenge, we introduce OutageDiT, a foundation model for generating seven-day outage trajectories at quarter-hour resolution, trained on outage and weather records across the United States. Specifically, a condition encoder processes the historical context and known future covariates once per forecast, and a shallow flow decoder reuses the resulting horizon-aligned states to generate complete trajectories. The resulting samples support point forecasting, uncertainty quantification, and conditional event simulation within one deep generative model. Across outage forecasting benchmarks, OutageDiT improves forecast accuracy and scenario quality over strong baselines and supports zero-shot transfer to held-out regions. Together, these results position conditional outage simulation as a bridge from outage forecasting to operational planning under uncertainty.
Chinese Translation
停电规划工作需要在事件发生前获得情景模拟。这些情景必须刻画停电规模、时间和持续时间的不确定性,同时保持时间上的依赖关系。然而,严重停电事件十分罕见,任何单一地区的数据中极端停电与恢复模式的样本都很少。为应对这一挑战,我们提出了 OutageDiT,这是一个以15分钟分辨率生成七天停电轨迹的基础模型,其训练数据覆盖美国全国的停电与气象记录。具体而言,条件编码器在每次预测中对历史背景信息和已知的未来协变量进行一次处理,浅层流解码器则复用由此得到的与预测时域对齐的状态,生成完整的轨迹。所得样本可在同一个深度生成模型中支持点预测、不确定性量化和条件事件模拟。在多项停电预测基准上,OutageDiT 相较于强基线模型提升了预测精度和情景质量,并支持向留出地区的零样本迁移。这些结果共同表明,条件停电模拟可以成为从停电预测迈向不确定性下运行规划的桥梁。
cs.LG / 21 / 2609.01925

CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing

CRISP:基于悬崖感知的输入自适应稀疏预填充与结构质量驱动的路由机制
Nguyen, Huu Huy, Van Nguyen, Chien, Dernoncourt, Franck, Rossi, Ryan A., Van, Linh Ngo, Chen, Jieyang, Nguyen, Thien Huu
Abstract
The attention prefilling phase of long-context LLM inference scales quadratically, making self-attention a severe computational bottleneck. Traditional sparse attention methods mitigate this through fixed patterns or offline profiling, but lack the flexibility to adapt to input-dependent attention structure. Recent dynamic methods address this by routing heads to sparse patterns in real-time, but rely on indirect routing proxies with overhead and budget allocation mechanisms that overlook the post-softmax mass hierarchy. We present CRISP (Cliff-awaRe Input-adaptive Sparse Prefilling), which identifies and addresses two structural challenges in this dynamic routing paradigm. First, we show that the routing decision can be read directly off the structure of the proxy attention map. We replace the Jensen-Shannon Divergence (JSD) routing with C_struct, a structural proxy that measures mass at Vertical-Slash compatible positions and reproduces JSD's routing decisions while eliminating both the pooled matmul and subsequent KL divergence overhead. Second, we formalize the post-softmax mass cliff and demonstrate theoretically that strictly cumulative coverage thresholds accumulate O(n) background noise at long contexts. CRISP navigates this via a sink-aware threshold grounded in the noise floor. Empirically, across InfiniteBench, RULER and LongBench on two model families, CRISP is the strongest sparse method overall and matches or exceeds exact dense attention on retrieval-heavy benchmarks, recovering up to +28.0 pp on retrieval tasks over baselines and achieving up to a 5.30x attention speedup at 512k tokens, driven primarily by our O(n) noise elimination during selection while preserving structural integrity.
Chinese Translation
长上下文大语言模型(LLM)推理中的注意力预填充阶段的计算复杂度随序列长度呈二次增长,使得自注意力成为严重的计算瓶颈。传统的稀疏注意力方法通过固定模式或离线剖析(offline profiling)来缓解这一问题,但缺乏适应输入相关注意力结构的灵活性。近期的动态方法通过实时将注意力头路由至稀疏模式来解决这一问题,但其依赖于带有开销的间接路由代理,且其预算分配机制忽视了softmax后质量(post-softmax mass)的层级结构。我们提出了CRISP(Cliff-awaRe Input-adaptive Sparse Prefilling,悬崖感知输入自适应稀疏预填充),识别并解决了该动态路由范式中的两个结构性挑战。首先,我们证明路由决策可以直接从代理注意力图的结构中读取。我们用C_struct取代了基于Jensen-Shannon散度(JSD)的路由,该结构代理度量Vertical-Slash兼容位置上的质量,能够复现JSD的路由决策,同时消除了池化矩阵乘法和后续KL散度计算的开销。其次,我们对softmax后质量悬崖(mass cliff)进行了形式化定义,并从理论上证明,严格的累积覆盖率阈值在长上下文中会累积O(n)量级的背景噪声。CRISP通过一个基于噪声底(noise floor)的、感知sink的阈值来应对该问题。实证结果表明,在两个模型家族上的InfiniteBench、RULER和LongBench评测中,CRISP是整体最强的稀疏方法,在检索密集型基准上匹配甚至超越精确稠密注意力,相比基线在检索任务上最高提升28.0个百分点,并在512k tokens时实现高达5.30倍的注意力加速。这主要得益于我们在选择过程中对O(n)噪声的消除,同时保持了结构的完整性。
cs.LG / 22 / 2609.01933

OR-Transformer: Scaling Real-Time Decision-Making to 1,000 Items

OR-Transformer:将实时决策扩展至1,000种物品
Liu, Shuze Daniel, Simchi-Levi, David, Chen, Claire, Gao, Chutong, Zhang, Shangtong
Abstract
Modern supply chain operations can require coordinating replenishment across thousands of heterogeneous items under correlated stochastic demand, heterogeneous lead times, and shared fixed ordering costs, yielding observation spaces exceeding $10^4$ dimensions. At this scale, rolling-horizon stochastic mixed-integer linear programs (MILPs) become prohibitively slow, while standard reinforcement learning (RL) methods face increasingly challenging credit assignment in high-dimensional action spaces. We introduce OR-Transformer, a deep reinforcement learning framework for joint replenishment under stochastic demand, with an item-permutation-equivariant Transformer architecture and pathwise-gradient training through the inventory dynamics. Across problem sizes up to 1,024 inventory items, OR-Transformer increasingly outperforms learning-based and rolling-horizon MILP baselines as scale grows. It also reduces online decision-making time by over 4 million times relative to MILP solvers, enabling real-time, large-scale deep RL in supply chain operations.
Chinese Translation
现代供应链运营可能需要在相关性随机需求、异质性提前期和共享固定订购成本的条件下,协调数千种异质性物品的补货,其观测空间维度超过$10^4$。在这种规模下,滚动时域随机混合整数线性规划(MILP)变得极其缓慢,而标准的强化学习(RL)方法在高维动作空间中面临日益严峻的信用分配挑战。我们提出了OR-Transformer,一个面向随机需求下联合补货问题的深度强化学习框架,采用物品置换等变的Transformer架构,并通过库存动态进行路径梯度(pathwise-gradient)训练。在规模高达1,024种库存物品的各类问题上,随着规模的增大,OR-Transformer相对基于学习的方法和滚动时域MILP基线的优势持续扩大。与MILP求解器相比,它还将在线决策时间缩短超过400万倍,实现了供应链运营中的实时大规模深度强化学习。
cs.LG / 23 / 2609.01942

Refining Heuristic-Based Bitcoin Address Clustering with Graph Neural Networks

Schnoering, Hugo, Bresson, Roman, Vazirgiannis, Michalis
Abstract
Bitcoin's pseudonymous nature makes it challenging to analyze user-level activity, since a single user may control multiple identifiers (addresses). Existing heuristic-based methods attempt to identify addresses belonging to the same user, but they often produce flat cluster assignments with limited modularity and are prone to errors such as merging different users together. In this work, we propose a method for refining heuristic-obtained clusters by grounding our clustering on contrastive embeddings yielded by graph neural networks. Our contributions are threefold: (i) we release a publicly available dataset of Bitcoin transaction graphs containing a substantial number of clusters; (ii) we propose a methodology for learning address embeddings consistent with heuristics, and back it up with theoretical guiding intuitions; (iii) through hierarchical clustering, we enable a finer analysis of heuristic clusters and provide a quantitative criterion for flagging suspicious merges.
Chinese Translation
模型服务未提供译文(内容过滤),请阅读上方英文摘要。
cs.LG / 24 / 2609.01947

On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers

在线策略蒸馏与离线策略GRPO的结合:训练紧凑的指令遵循重排序器
Prabhakar, Vignesh, Pan, Jialing, Ankisettipalli, Anil Babu
Abstract
Compact instruction-following rerankers are attractive for deployment, but conventional distillation pipelines typically train students by offline imitation of teacher outputs on a fixed set of examples, constraining supervision to the teacher's observed ranking space. We revisit reranker distillation through the lens of reinforcement learning. We propose a two-stage framework combining off-policy teacher optimization with on-policy student distillation. In Stage 1, a 4B teacher reranker is strengthened with off-policy GRPO using LLM-judge feedback on 88K instruction-following examples. In Stage 2, a compact 1B student samples rankings from its own policy and receives soft teacher-derived rewards on those rankings, coupling student exploration with knowledge transfer. Our strongest gains appear under distribution shift. On MAIR-11, the original 11-subset, 869-query evaluation, the proposed student reaches 0.7670 nDCG@6, outperforming offline listwise KD by +4.6 points. Controlled comparisons against offline pairwise RankNet KD and on-policy GKD show that neither changing the offline distillation objective nor moving teacher-distribution matching on-policy reproduces the performance of reward-based on-policy distillation over student-sampled rankings. The advantage persists on MAIR-Full: across all 126 tasks and 9,356 queries, the proposed method obtains the highest task-macro point estimates among the evaluated distillation variants, reaching 0.6808 nDCG@6 and 0.7865 MRR@6. It also exceeds two released 7B RL-trained rerankers on the comparable MAIR-11 evaluation, while the same Stage 2 training procedure consistently improves three architecturally distinct alternative student backbones. On the 9,861-query validation benchmark, the resulting 1B reranker achieves 0.7624 nDCG@6 while providing a favorable quality-efficiency tradeoff relative to larger alternatives.
Chinese Translation
紧凑的指令遵循重排序器(reranker)在实际部署中极具吸引力,但传统的蒸馏流程通常通过在固定样本集上对教师模型输出进行离线模仿来训练学生模型,从而将监督限制在教师模型已观测到的排序空间内。我们从强化学习的视角重新审视重排序器蒸馏,提出了一种将离线策略教师优化与在线策略学生蒸馏相结合的两阶段框架。在第一阶段,利用离线策略GRPO(GRPO)并结合LLM裁判(LLM-judge)对88K条指令遵循样本的反馈,强化一个4B的教师重排序器。在第二阶段,一个紧凑的1B学生模型从自身策略中采样排序结果,并针对这些排序获得基于教师的软奖励,从而将学生的探索与知识迁移相耦合。我们的方法在分布偏移下取得了最显著的收益。在MAIR-11(原始的11个子集、869个查询的评测集)上,所提出的学生模型达到了0.7670的nDCG@6,比离线列表式(listwise)知识蒸馏高出4.6个百分点。与离线逐对(pairwise)RankNet知识蒸馏以及在线策略GKD的受控对比表明,无论是改变离线蒸馏目标,还是将教师分布匹配迁移到在线策略,都无法复现基于奖励的、在学生采样排序上进行的在线策略蒸馏的性能。该优势在MAIR-Full上依然存在:在全部126个任务和9,356个查询上,所提出的方法在所有被评估的蒸馏变体中获得了最高的任务宏平均点估计,达到0.6808的nDCG@6和0.7865的MRR@6。在可比较的MAIR-11评测中,它还超越了两个已发布的7B经强化学习训练的重排序器,并且同样的第二阶段训练流程能够持续提升三个架构不同的替代学生骨干模型。在包含9,861个查询的验证基准上,最终的1B重排序器实现了0.7624的nDCG@6,同时相较于更大的替代模型提供了更优的质量-效率权衡。
cs.LG / 25 / 2609.01952

Convergence Theory of Knowledge Distillation in Asynchronous P2P Gossip Learning Network

异步P2P Gossip学习网络中知识蒸馏的收敛理论
Fang, Lucas Qingyang, Liu, Tiyao, Jing, Jinhao, Li, Zeji, Chen, Kaijie, Kuttivelil, Harikrishna, Obraczka, Katia
Abstract
Decentralized, serverless learning increasingly connects devices running different architectures, where the standard tool, decentralized SGD, is undefined as models with different parameter counts cannot be averaged. Knowledge distillation (KD) exchanges soft predictions rather than weights and sidesteps this obstacle, yet convergence theory for fully decentralized, asynchronous peer-to-peer (P2P) KD is lacking. We provide one, relocating consensus from parameter space to function (output) space: a KD event is a geometric contraction operator in logit space on the peers' predictive distributions, which we analyse in the Hilbert space of predictions on a reference measure. Under standard smoothness/variance assumptions and two realizability assumptions, one bridging parameter SGD to the functional step and one controlling restricted task/KD alignment, the time-averaged functional stationarity and function-space disagreement converge at rate $O(1/(\eta T))$ to an $O(\eta)+O(B_f^2)+O(\zeta_f^2)$ neighbourhood. Here $B_f$ is the distance from the task optimum to the peers' reachable classes and $\zeta_f$ measures persistent local-task heterogeneity. Across homogeneous, width-heterogeneous, and mixed-family networks of the experiments, KD contracts function disagreement by $40-61\times$, while isolated training does not. The sampled stationarity diagnostic has late transient exponents $0.99-1.90$ on the shared-skeleton main runs, and the four-point step-size sweep exhibits the predicted transient: neighbourhood tradeoff.
Chinese Translation
去中心化、无服务器的学习正日益连接起运行不同架构的设备,而标准工具——去中心化随机梯度下降(SGD)——在此情形下不再适用,因为参数数量不同的模型无法进行平均。知识蒸馏(Knowledge Distillation, KD)通过交换软预测而非模型权重来规避这一障碍,然而针对完全去中心化的异步点对点(P2P)知识蒸馏,其收敛理论尚属空白。本文提供了这样一个理论,将一致性(consensus)概念从参数空间迁移到函数(输出)空间:一次KD事件可视为在logit空间中对节点预测分布的几何收缩算子,我们在基于参考测度的预测希尔伯特空间中对其进行分析。在标准的平滑性/方差假设以及两个可实现性假设(一个用于衔接参数SGD与函数步更新,另一个用于控制受限任务与KD的对齐程度)下,时间平均的函数平稳性与函数空间分歧以 $O(1/(\eta T))$ 的速率收敛至一个 $O(\eta)+O(B_f^2)+O(\zeta_f^2)$ 的邻域。其中 $B_f$ 表示任务最优点到各节点可达类别的距离,$\zeta_f$ 度量持续存在的局部任务异质性。在实验涉及的同类网络、宽度异构网络以及混合架构网络中,KD 将函数分歧收缩了 40-61 倍,而孤立训练则无法做到。共享骨架的主要实验中,采样得到的平稳性诊断指标表现出后期瞬态指数为 0.99-1.90,且四点步长扫描实验呈现出符合理论预测的瞬态-邻域权衡关系。
cs.LG / 26 / 2609.01956

InKAN: B-Spline KANs via Truncated Power Form

InKAN:基于截断幂形式的B样条Kolmogorov-Arnold网络
Mysore, Naveen
Abstract
Kolmogorov-Arnold Networks (KANs) place learnable B-spline activations on network edges rather than fixed activations on nodes. The standard Cox-de Boor recursion evaluates these activations through $k$ sequential passes for degree-$k$ splines, consuming over 90% of forward-pass time. InKAN replaces this recursion with the truncated power form, a classical result from approximation theory that expresses each uniform cubic B-spline as five $(x)_+^3$ terms at shifted knot positions. This paper makes three contributions: (1) a torch.compile-fused implementation that collapses these operations into a single GPU kernel, eliminating all recursion, span lookup, and scatter-gather operations; (2) a bounded-coordinate stabilization that clamps the normalized input to $[0, k{+}1]$, preventing the catastrophic cancellation that historically motivated the Cox-de Boor recursion; and (3) a production-ready, open-source package (pip install inkan) that serves as a drop-in replacement for existing KAN layers.
Chinese Translation
Kolmogorov-Arnold网络(KAN)将可学习的B样条激活函数置于网络的边上,而非在节点上使用固定激活函数。标准的Cox-de Boor递归算法在计算$k$次样条时需要经过$k$次顺序传递来求值这些激活函数,占据了前向传播90%以上的时间。InKAN用截断幂形式取代了这一递归算法——这是逼近理论中的一个经典结果,它将每个均匀三次B样条表示为在移位节点位置上的五个$(x)_+^3$项。本文做出了三项贡献:(1)一种torch.compile融合实现,将这些运算合并为单个GPU内核,彻底消除了递归、区间查找以及散布-聚集操作;(2)一种有界坐标稳定化方法,将归一化输入钳制到$[0, k{+}1]$区间,防止了历史上促使Cox-de Boor递归产生灾难性抵消的问题;(3)一个可用于生产环境、开源的软件包(pip install inkan),可作为现有KAN层的直接替代品。
cs.LG / 27 / 2609.01967

A Unified Particle Filter LSTM for Data-Driven Process Simulation

一种用于数据驱动过程仿真的统一粒子滤波LSTM模型
Malekzadeh, Parvin, Baron, Opher, Krass, Dmitry
Abstract
Data-driven process simulation aims to generate realistic case trajectories from historical event logs without requiring an explicitly specified model of the underlying dynamics. Deep sequence models can capture complex temporal dependencies through next-activity probabilities and conditional time distributions. However, event logs provide only a partial view of the underlying process state, often recording activity completions without the corresponding service-start times. Consequently, the same observed process history may be consistent with multiple plausible latent process conditions, whereas standard recurrent models compress each process prefix into a single deterministic recurrent state. We propose a Unified Particle Filter LSTM (Unified PF-LSTM) that maintains and sequentially updates a weighted set of recurrent-state hypotheses. We summarize this particle belief using its weighted mean and learned features based on the moment-generating function. The resulting representation is used to predict a categorical distribution over the next activity and conditional quantiles of the current activity's sojourn time. The framework is trained end-to-end from event-log data and evaluated on three real-world emergency department datasets. The results show that the proposed framework consistently outperforms the considered data-driven baselines in reproducing routing, duration, and system-level behavior across all datasets, with particularly strong gains in settings where complex process dynamics are only partially reflected in the available event logs.
Chinese Translation
数据驱动的过程仿真旨在从历史事件日志中生成真实的案例轨迹,而无需显式指定底层动态过程的模型。深度序列模型能够通过下一活动概率和条件时间分布来捕捉复杂的时间依赖关系。然而,事件日志仅提供了底层过程状态的部分视图,通常只记录活动完成情况而缺少相应的服务开始时间。因此,同一观测到的过程历史可能与多种合理的潜在过程状态相一致,而标准的循环模型会将每个过程前缀压缩为单一确定性的循环状态。我们提出了一种统一粒子滤波LSTM(Unified PF-LSTM),用于维护并顺序更新一组带权重的循环状态假设集合。我们利用加权均值以及基于矩生成函数(moment-generating function)学习到的特征来概括这一粒子信念。所得的表示被用于预测下一活动的类别分布以及当前活动停留时间的条件分位数。该框架基于事件日志数据进行端到端训练,并在三个真实急诊科数据集上进行了评估。结果表明,所提出的框架在复现流程路径、时长以及系统级行为方面,在所有数据集上均一致优于所考虑的数据驱动基线方法,在复杂过程动态仅能被可用事件日志部分反映的场景中尤为显著。
cs.LG / 28 / 2609.01991

CAHR-Net: Condition-Adaptive Hysteresis Reconstruction for Compact and Interpretable Magnetic Core Loss Modeling

CAHR-Net:面向紧凑且可解释磁芯损耗建模的条件自适应磁滞回线重构网络
Gong, Chunye, Yao, Cong
Abstract
Magnetic core loss originates in the hysteresis loop: the energy dissipated per excitation cycle equals the loop area, and frequency, temperature, and waveform shape set the loss by reshaping the loop geometry. Most existing models let these conditions act only on a terminal scalar - empirical equations fold them into fitted exponents, and data-driven predictors append them to encoded features - so no intermediate hysteresis representation remains for the conditions to reshape. This paper proposes CAHR-Net, a condition-adaptive hysteresis reconstruction network that injects the operating conditions where they physically act. It preserves the interpretable chain from flux density waveform to magnetic field reconstruction, loop-area integration, and power loss estimation, and uses feature-wise linear modulation to inject frequency, temperature, and waveform statistics into the intermediate reconstruction representation. A matched large-batch training protocol based on AdamW, cosine scheduling, and a staged reconstruction-to-power-loss objective is also reported, because the modulation pathway takes effect only within it. On the MagNet final A-E material protocol, CAHR-Net attains an average p95 relative error of 6.89% with only 1874 parameters, the lowest among all compared methods, together with a lower worst-material p95 than the strongest black-box solution at about 48x fewer parameters; it reduces the average p95 of the physical reconstruction backbone from 7.47% to 6.89% and the p95 of material D, the most difficult material, from 16.40% to 14.87%. Ablation and condition-slice analyses attribute the improvement to the coupling of physical loop reconstruction, structured condition modulation, and the matched optimization trajectory.
Chinese Translation
磁芯损耗源于磁滞回线:每个励磁周期耗散的能量等于回线面积,而频率、温度和波形形状通过重塑回线几何形态来决定损耗大小。现有大多数模型仅让这些条件作用于一个末端标量——经验公式将它们拟合为指数系数,数据驱动预测器则将它们附加到编码特征之后——因而没有任何中间磁滞表征可供条件去重塑。本文提出CAHR-Net(条件自适应磁滞回线重构网络),将工作条件注入其物理作用之处。该方法保留了从磁通密度波形到磁场重构、回线面积积分以及功率损耗估算的可解释链条,并采用特征级线性调制(FiLM)将频率、温度和波形统计量注入中间重构表征。此外,本文还给出了一套与之匹配的大批量训练协议,该协议基于AdamW优化器、余弦调度以及分阶段的重构到功率损耗目标函数,因为调制通路只有在该协议下才能生效。在MagNet最终A–E材料协议上,CAHR-Net仅用1874个参数即取得平均p95相对误差6.89%的结果,为所有对比方法中最低,且在参数量约为最强黑盒方案1/48的情况下,其最差材料p95误差更低;该方法将物理重构主干网络的平均p95误差从7.47%降至6.89%,并将最难的D材料p95误差从16.40%降至14.87%。消融实验和条件切片分析将这一提升归因于物理回线重构、结构化条件调制与匹配优化轨迹三者的耦合。
cs.LG / 29 / 2609.02006

Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation

训练即部署:弥合低秩克隆蒸馏中的MLP可达性差距
Chen, Wenhui, Li, Zhifeng, Zhou, Jie, Singh, Navan Preet, Ciobanu, Madalina, Wang, Chenghua, Mao, Qingqing, Das, Ritankar
Abstract
A compressed student has two shapes that need not agree: the weight it deploys at inference and the weight family its training can reach. We show that a state-of-the-art weight-inheritance distiller, Low-Rank Clone (LRC), deploys a full-width student MLP but ties training to a teacher-induced slice, leaving 62.5-81.4% of each deployed matrix's independent linear degrees of freedom unreachable-paid for at inference, never trainable. Our principle is one line: train what you deploy. From the identical LRC warm start, we make the training object the entire deployed matrix, with no change in deployed shape, deployed parameter count, or inference FLOPs, via two mergeable realizations (Dense-LRC and CORE-LRC) that both collapse to one deployed weight. This recovers stranded capacity: taking the stronger realization per teacher, +2.36/+2.71/+10.45 Avg9 over matched-budget plain-LRC baselines across three teachers (Llama3.2-3B, Llama3.1-8B, Qwen2.5-3B), with the largest gain on the widest teacher (Qwen), where it reaches the original recipe's approx. 20B-token accuracy at 10B tokens (2x token efficiency); there the strictly same-lineage arm still recovers +6.39, the fully controlled figure. Controls strongly support attributing the gain to the enlarged reachable set, rather than to added parameters or the recipe. From approx. 10B distillation tokens plus a short SFT, a half-parameter 1.5B student matches its approx. 9T-token teacher's 9-task macro-average, within evaluation noise and with a residual MMLU deficit, and a 2.7B student beats Meta's own official compression of Llama3.1-8B at ~900x fewer compression tokens (a token count under unmatched recipes, not a compute claim). All results are from single-seed runs on the LRC backbone.
Chinese Translation
压缩后的学生模型存在两种形态,且二者未必一致:一是推理时部署的权重,二是其训练过程能够到达的权重族。我们发现,最先进的权重继承蒸馏方法——低秩克隆(Low-Rank Clone, LRC)——部署的是全宽度的学生MLP,但其训练却被限制在一个由教师模型诱导的低维切片上,导致每个被部署矩阵中62.5%-81.4%的独立线性自由度无法被训练到——这些容量在推理时被付出代价,却从未被训练。我们的原则可以概括为一句话:训练你所部署的。从完全相同的LRC热启动出发,我们将训练对象改为整个被部署的矩阵,且不改变部署形态、部署参数量或推理FLOPs,具体通过两种可合并的实现方式(Dense-LRC和CORE-LRC)实现,二者最终均坍缩为同一个部署权重。这恢复了我置的容量:在三个教师模型(Llama3.2-3B、Llama3.1-8B、Qwen2.5-3B)上,按每个教师选取更强的实现方式,相对同等预算的plain-LRC基线,Avg9分别提升+2.36/+2.71/+10.45;其中增益在宽度最大的教师(Qwen)上最为显著——在10B token时即达到原始配方约20B token的精度(2倍token效率);在严格同谱系的实验组中仍恢复+6.39,这是完全受控的数字。对照实验有力地支持将增益归因于可达集合的扩大,而非新增参数或配方本身。在使用约10B蒸馏token加上简短的SFT后,参数量减半的1.5B学生模型在其9项任务宏平均上达到其约9T-token教师模型的水平(在评估噪声范围内,仅有MMLU的残余差距);2.7B学生模型以约900倍更少的压缩token数超越了Meta官方对Llama3.1-8B的压缩(该token计数是在不同配方下的对比,并非算力对比)。所有结果均基于LRC骨干上的单种子运行。
cs.LG / 30 / 2609.02018

Source-Free Class Relearning: Diagnosing Forgetting in Class Unlearning

无来源类别再学习:诊断类别遗忘中的遗忘问题
Dehghani, Zahra, Piantanida, Pablo, Shateri, Mohammadhadi
Abstract
Class unlearning aims to remove a model's ability to recognize designated forget classes while preserving performance on retain classes. However, low forget accuracy after unlearning does not necessarily mean the class structure has been erased. Approximate unlearning methods can alter classifier decision boundaries while leaving recoverable structure in the representation. Prior work has shown that forget classes can be recovered, but existing approaches require real forget or retain samples, auxiliary data, or reference checkpoints. We study class relearning in a strictly source-free setting, asking whether a forget class can be recovered through a classifier-head update using only the unlearned model. Our approach rests on a theoretical analysis establishing a sufficient alignment condition under which a single gradient step on a synthetic probe set increases the expected logit margin of the forget class. Building on this, we propose a white-box Source-Free Relearning Audit (SFRA), which generates candidate embeddings in representation space and uses model-guided confidence filtering to construct high-confidence retain probes and low-confidence boundary-adjacent probes that are relabelled as the forget class. Gaussian sampling and Softmax confidence are used by default, while ablations with alternative proposal distributions and uncertainty criteria show that recoverability is not specific to these choices. To quantify recoverability, we introduce the Relearning Score (RS), which jointly measures forget-class recovery and retain-accuracy preservation, and report class-matched $\Delta$RS relative to a retrained reference. Experiments on CIFAR-10, CIFAR-100, and TinyImageNet with ResNet-18, ViT-B/16, and Swin-T show that several unlearning methods exhibit substantial source-free recoverability, and that for a subset of methods this recoverability exceeds the matched retrained reference.
Chinese Translation
类别遗忘旨在移除模型对指定遗忘类别的识别能力,同时保持其在保留类别上的性能。然而,遗忘后较低的遗忘类别准确率并不一定意味着类别结构已被彻底抹除。近似遗忘方法可能只改变了分类器的决策边界,而在表示中留下可恢复的结构。已有研究表明遗忘类别是可以被恢复的,但现有方法需要真实的遗忘或保留样本、辅助数据或参考检查点。我们在严格的“无来源”设置下研究类别再学习问题,即仅利用已遗忘的模型、通过分类器头部的更新来探究遗忘类别能否被恢复。我们的方法建立在理论分析之上,确立了一个充分的对齐条件:在该条件下,对合成探测集执行单步梯度上升即可增大遗忘类别的期望logit间隔。基于此,我们提出了一种白盒的“无来源再学习审计”方法,该方法在表示空间中生成候选嵌入,并利用模型引导的置信度过滤来构建高置信度的保留类别探测样本,以及被重新标注为遗忘类别的低置信度、靠近边界的探测样本。默认采用高斯采样与Softmax置信度,而使用其他候选分布与不确定性准则的消融实验表明,可恢复性并不依赖于这些特定选择。为量化可恢复性,我们提出了“再学习分数”,它联合度量遗忘类别的恢复程度与保留类别的准确率保持情况,并报告相对于重新训练参考模型的类别匹配的ΔRS。在CIFAR-10、CIFAR-100和TinyImageNet数据集上,使用ResNet-18、ViT-B/16和Swin-T的实验表明,多种遗忘方法均表现出显著的无来源可恢复性,且对于其中一部分方法,这种可恢复性甚至超过了匹配的重新训练参考。
cs.LG / 31 / 2609.02042

Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents

多行动、少决策:面向长时程LLM智能体的技能引导自适应动作分块
Yang, Yanting, Jin, Can, Zhao, Jinman, Wu, Jiahao, Zhou, Yang, Wang, Zhepeng, Wang, Zhendong, Zhou, Mu, Metaxas, Dimitris N.
Abstract
Large language model (LLM) agents for long-horizon interactive tasks typically follow a ReAct-style protocol, issuing one primitive action per LLM round. While this enables frequent replanning, it is inefficient for long-horizon tasks where many rounds are spent on routine action sequences. A natural alternative is to let the agent emit variable-length action chunks. However, naively training such policies with standard reinforcement learning fails: the agent either collapses to single-action behavior or over-commits to excessively long sequences. Both failures share a common root cause: the inability to learn chunk boundaries. We propose SPACE, which addresses this challenge by distilling chunk-boundary supervision from trajectory-induced programmatic skills. We induce two-level programmatic skills from successful trajectories, where subskill boundaries serve as direct chunk-boundary supervision. This temporal structure is then distilled into a primitive-chunk policy via hybrid on-/off-policy optimization with chunk-aware credit assignment. Experiments on ALFWorld and ScienceWorld show that SPACE improves success rates by 7.0%-31.3% over the strongest baseline in each setting while reducing average LLM decision rounds by up to 78.9%.
Chinese Translation
面向长时程交互任务的大语言模型(LLM)智能体通常遵循ReAct风格的协议,即每轮LLM调用只发出一个原始动作。尽管这种方式支持频繁的重新规划,但对于需要在大量轮次上执行常规动作序列的长时程任务而言效率低下。一种自然的替代方案是让智能体输出可变长度的动作块。然而,直接使用标准强化学习训练此类策略往往会失败:智能体要么退化为单动作行为,要么过度提交过长的动作序列。这两种失败有一个共同的根本原因:无法学习动作块的边界。我们提出SPACE方法,通过从轨迹中归纳的程序化技能中蒸馏动作块边界监督信号来解决这一挑战。我们从成功轨迹中归纳出两级程序化技能,其中子技能边界可直接作为动作块边界的监督信号。随后,通过结合块感知信用分配的混合在线/离线策略优化,将该时序结构蒸馏到一个原始动作-动作块策略中。在ALFWorld和ScienceWorld上的实验表明,SPACE在每种设置中将成功率较最强基线提升了7.0%–31.3%,同时将平均LLM决策轮次最多减少了78.9%。
cs.LG / 32 / 2609.02049

The Dynamics of Continuous Mixture Collapse in Language Models

语言模型中连续混合坍缩的动力学机制
Backour, Ali
Abstract
LLMs latent-state reasoning methods replace discrete intermediate tokens with continuous states, such as weighted mixtures of token embeddings, to retain multiple possible reasoning directions rather than committing to one. Yet pretrained language models often fail to preserve these mixtures. We study why through a combination of theoretical analysis and controlled empirical investigations on a variety of models. We identify three independent, distinct sources of failure. First, transformer architectures already distort mixture geometry, and training substantially amplifies this effect. Moreover, the failure can occur even if the model transports mixtures perfectly linearly: the softmax readout and autoregressive feedback form a dynamical system that either amplifies small differences until one component of the mixture dominates or contracts different mixtures until they become indistinguishable. We verify this theoretical prediction empirically: the observed transition between contraction and amplification occurs near the theoretical threshold derived by our analysis, and pretrained-model rollouts lie predominantly on the amplifying side. Finally, we generalize to mixtures of many components and show that exact preservation generally requires context-dependent correction, whose required dimensionality can grow with the number of components.
Chinese Translation
大语言模型(LLMs)的潜在状态推理方法用连续状态(如词元嵌入的加权混合)替代离散的中间词元,以保留多种可能的推理方向,而非只承诺于单一方向。然而,预训练语言模型往往无法保持这些混合状态。我们通过理论分析与在多种模型上进行的受控实证研究相结合的方式来探究其原因,并识别出三种相互独立、彼此不同的失败来源。首先,Transformer 架构本身就会扭曲混合状态的几何结构,而训练过程会显著放大这种效应。此外,即使模型能够完美地以线性方式传输混合状态,失败仍可能发生:softmax 读出机制与自回归反馈构成一个动力系统,该系统要么放大微小差异直至混合中的某一分量占主导地位,要么收缩不同的混合状态直至它们变得不可区分。我们通过实证验证了这一理论预测:所观察到的收缩与放大之间的转变发生在我们分析所推导出的理论阈值附近,且预训练模型的展开(rollout)主要处于放大一侧。最后,我们将结论推广到多分量混合的情形,并证明精确保持混合状态通常需要依赖于上下文的校正,而该校正所需的维度可能随分量数量的增加而增长。
cs.LG / 33 / 2609.02068

DynG-Diff: A State-Aware Dynamic Guidance Diffusion Framework for Probabilistic Time Series Forecasting

DynG-Diff:一种面向概率时间序列预测的状态感知动态引导扩散框架
Zhang, Zhente, Ni, Zhengwei, Fan, Wei
Abstract
Probabilistic multivariate time series (MTS) forecasting is crucial for modeling complex dynamical systems. However, existing diffusion-based methods rely on task-specific conditional paradigms that lack flexibility and struggle with inherent "information heterogeneity"--the significantly varying noise levels and evolutionary patterns across variables. To address this, we propose DynG-Diff, a variable-sensitive dynamic guidance diffusion framework for probabilistic multivariate time-series forecasting: (1) DynG-Diff adopts a two-stage separated training strategy and uses an unconditional diffusion backbone to model the joint distribution of multivariate time series. (2) DynG-Diff introduces a lightweight state-aware policy network that adaptively infers variable reliability from real-time noisy states and one-step denoising estimates, outputting a dynamic guidance strength matrix. (3) DynG-Diff mathematically formulates this dynamic weight as the local precision of the observation distribution, enabling precise guidance for high-confidence variables during inference while filtering out interference from anomalous noise. Extensive experiments on real-world benchmarks demonstrate competitive probabilistic forecasting performance against state-of-the-art conditional diffusion models and improved robustness under severe observation corruption.The implementation code is available at: https://github.com/TT-20011031/DynG-Diff
Chinese Translation
概率多元时间序列(MTS)预测对于建模复杂动态系统至关重要。然而,现有的基于扩散模型的方法依赖于缺乏灵活性的任务特定条件范式,并且难以应对固有的“信息异质性”问题——即各变量之间显著不同的噪声水平和演化模式。为解决这一问题,我们提出了DynG-Diff,一种面向概率多元时间序列预测的变量敏感动态引导扩散框架:(1)DynG-Diff采用两阶段分离训练策略,使用无条件扩散主干网络对多元时间序列的联合分布进行建模。(2)DynG-Diff引入一个轻量级的状态感知策略网络,根据实时噪声状态和一步去噪估计自适应地推断变量可靠性,并输出动态引导强度矩阵。(3)DynG-Diff从数学上将该动态权重表述为观测分布的局部精度,从而在推理过程中对高置信度变量提供精确引导,同时过滤异常噪声的干扰。在真实世界基准数据集上的大量实验表明,该方法相比最先进的条件扩散模型具有竞争力的概率预测性能,并且在严重观测损坏下具有更强的鲁棒性。实现代码见:https://github.com/TT-20011031/DynG-Diff
cs.LG / 34 / 2609.02083

XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression

XMerge:面向大语言模型深度压缩的跨轴选择与重构层合并方法
Hu, Jundong, Ramachandran, Shekar
Abstract
Removing complete transformer layers preserves a standard serving architecture, but existing depth-compression methods can lose substantial quality, and the loss varies unpredictably across models. We introduce XMerge, a post-training method with two components. Cross-axis selection identifies a block with low relative-magnitude and angular hidden-state change, and local boundary reconstruction re-fits the adjacent surviving block to match the original two-block output. XMerge uses no task labels or end-to-end fine-tuning, and it introduces neither architectural changes nor additional inference-time parameters. Across seven Llama and Qwen backbones (0.5B-8B), five published baselines, and three layer-reduction levels, its advantage over baselines is largest at the most aggressive removal: at k=4 it ranks first on six of seven backbones on CORE (a 22-task aggregate) and, separately, on six of seven on MMLU (five of seven on both at once), while avoiding the large perplexity increases of several competing operators. In a task-level bootstrap, the 95% confidence intervals for the three largest CORE margins exclude zero; the remaining margins are consistent with ties. Across the 14 (model, regime) cells it is also the only evaluated operator that never collapses, ranking top-2 in both zero-shot and in-context regimes; on a first calibration probe (one backbone) it is the best-calibrated operator. Ablations show that local reconstruction provides most of the gain, while cross-axis fusion helps when the two selection axes disagree. The additional construction cost is recovered through per-token decode savings after roughly tens of thousands of requests.
Chinese Translation
移除完整的Transformer层可以保持标准的服务架构,但现有的深度压缩方法可能会损失大量性能,且损失程度在不同模型之间难以预测。我们提出了XMerge,这是一种包含两个组件的训练后方法。跨轴选择(Cross-axis selection)识别隐藏状态相对幅值变化和角度变化均较小的模块块,局部边界重构(local boundary reconstruction)则对相邻的保留模块进行重新拟合,以匹配原始两个模块的输出。XMerge不使用任务标签或端到端微调,也不引入架构更改或额外的推理时参数。在七个Llama和Qwen骨干模型(0.5B-8B)、五个已发表的基线方法以及三个层削减级别上的实验中,其相对于基线的优势在最高激进的移除幅度下最为显著:在k=4时,它在七个骨干模型中的六个上于CORE(一个22项任务的聚合基准)排名第一,在MMLU上同样在七个中的六个排名第一(两者同时在五个骨干模型上排名第一),同时避免了若干竞争算子出现的大幅困惑度上升。在任务级bootstrap检验中,三个最大的CORE优势的95%置信区间均不包含零;其余优势与持平情况一致。在全部14个(模型, regime)组合中,它也是唯一从未出现性能崩溃的评估算子,在零样本和上下文学习两种设置下均排名前二;在首次校准探针实验(一个骨干模型)中,它是校准最佳的算子。消融实验表明,局部重构贡献了大部分收益,而跨轴融合在两个选择轴不一致时起到帮助作用。额外的构建成本在大约数万个请求之后即可通过逐token解码的节省得到回收。
cs.LG / 35 / 2609.02085

TC-Next: Zero-Shot Multimodal Cyclone Forecasting

TC-Next:零样本多模态热带气旋预报
Wang, Zhe, Chen, Sijie, Luo, Yiming, Kim, Daehyun, Chang, Chien-Yi
Abstract
We present TropicalCycloneNext (TC-Next), a multimodal deep learning model that forecasts tropical cyclone track and intensity at $6$-$24$ h leads by leveraging a foundation model's forecast fields of atmospheric kinematic and thermodynamic fields and GridSat infrared satellite imagery. Trained only on GraphCast forecasts over the Western Pacific (WP), yet reliant only on generic atmospheric variables, TC-Next on GraphCast lowers track error by $15$-$44\%$ and intensity error by a factor of $3$-$6$ relative to a conventional, rule-based tracker, TempestExtremes; applied without retraining to the forecast fields of Pangu-Weather and IFS HRES, it stays ahead of TempestExtremes on both. Applied zero-shot to the generic weather fields of WeatherNext Cyclones on the 2025 WP season, TC-Next attains lower intensity error at every lead time, and lower or comparable track error, compared to that model's specialized direct tracker in a deterministic comparison. Our ablation studies show that our multimodal model is able to utilize the additional modality to improve performance in tracking errors at every lead time and in intensity prediction at longer lead times.
Chinese Translation
我们提出了TropicalCycloneNext(TC-Next),这是一个多模态深度学习模型,可利用基础模型预报的大气动力学和热力学场以及GridSat红外卫星图像,对热带气旋的路径和强度进行6-24小时提前量的预报。TC-Next仅在西北太平洋(WP)的GraphCast预报上训练,且仅依赖通用大气变量;相对于传统的基于规则的追踪器TempestExtremes,基于GraphCast的TC-Next将路径误差降低了15-44%,强度误差降低了3-6倍;在不重新训练的情况下直接应用于Pangu-Weather和IFS HRES的预报场时,它在两者上均优于TempestExtremes。以零样本方式将该模型应用于WeatherNext Cyclones在2025年西北太平洋季的通用天气场时,在确定性比较中,TC-Next在所有提前量上均取得更低(或相当)的路径误差和更低的强度误差,优于该模型的专用直接追踪器。消融实验表明,我们的多模态模型能够利用额外模态,在所有提前量上改善路径追踪误差,并在较长提前量上改善强度预测。
cs.LG / 36 / 2609.02093

Compositional Spectral Prompts for LLM-based Online Time Series Forecasting

面向基于大语言模型的在线时间序列预测的组合式频谱提示方法
Choi, Seungyoon, Kim, Hyunchul, Lee, Jae-Gil, Park, Chanyoung
Abstract
To address the sequential and evolving nature of time series, the Online Time Series Forecasting (OTSF) task has been extensively studied in multiple domains. Existing research focuses on adapting to non-stationary environments by employing memory buffer-based retrieval strategies. However, we observe that such frameworks struggle with long-term adaptation and fail to generalize to unseen patterns. To this end, we introduce CoSPOT, an LLM-based online time series forecasting framework that leverages a pre-trained LLM as the backbone online forecaster, motivated by its strong few-shot capabilities. For efficient online adaptation, CoSPOT keeps the LLM frozen and employs compositional spectral prompts grounded in frequency-domain bases to guide the model with the overall distribution of the input, thereby substantially reducing the number of parameters updated during the online phase. Specifically, CoSPOT decomposes time series into frequency bases and composes the corresponding spectral basis prompts according to their amplitudes, allowing unseen patterns to be represented as new combinations of learned basis prompts. Our extensive experiments on real-world datasets demonstrate the superiority and practicality of CoSPOT across challenging online scenarios, including extended online phases and cross-dataset settings with substantial distribution shifts. Our code is available at https://github.com/seungyoon-Choi/CoSPOT.
Chinese Translation
为应对时间序列的时序性与演化特性,在线时间序列预测(Online Time Series Forecasting, OTSF)任务已在多个领域得到广泛研究。现有研究主要通过基于记忆缓冲区的检索策略来适应非平稳环境。然而,我们观察到此类框架在长期适应方面表现不佳,且难以泛化到未见过的模式。为此,我们提出了CoSPOT,一个基于大语言模型(LLM)的在线时间序列预测框架。鉴于预训练LLM强大的少样本能力,我们以其作为在线预测器的主干。为实现高效的在线适应,CoSPOT保持LLM参数冻结,并采用基于频域基底的组合式频谱提示来引导模型利用输入的整体分布信息,从而大幅减少在线阶段需要更新的参数数量。具体而言,CoSPOT将时间序列分解为频率基底,并根据其幅值组合相应的频谱基底提示,使未见过的模式能够被表示为已学习基底提示的新组合。我们在真实数据集上的大量实验表明,CoSPOT在具有挑战性的在线场景中(包括延长的在线阶段以及存在显著分布偏移的跨数据集设置)展现出优越性与实用性。我们的代码已发布于 https://github.com/seungyoon-Choi/CoSPOT。
cs.LG / 37 / 2609.02101

Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts

基于四个国际胸部X射线队列的BiomedCLIP联邦LoRA适配
Poudel, Sanjaya, Kunwor, Nirajan, Dhakal, Manish, Jha, Debesh, Gaire, Sunil Kumar
Abstract
Federated learning (FL) lets institutions train a shared model without exchanging data, and Low-Rank Adaptation (LoRA) makes this practical at scale by communicating only compact low-rank updates. Biomedical imaging is a compelling setting for this combination: patient data are archived behind privacy regulations, and institutions differ widely in scanners, protocols, and compute. Such heterogeneity raises the question of how federated LoRA updates should be aggregated, increasingly pressing as multimodal vision-language models become central to medical image analysis. We benchmark federated Parameter-efficient fine-tuning (PEFT) of BiomedCLIP for chest radiograph classification across four public cohorts on three continents (USA, Vietnam, Spain). Federated LoRA adaptation improves shared-class AUC on all four cohorts over the unadapted BiomedCLIP backbone (mean 0.687 to 0.802), showing that the gains come from federated adaptation rather than from the pretrained model's zero-shot ability. Relative to isolated single-cohort training, federation improves the weaker cohorts while largely preserving the strongest and approaches a centralized reference (0.812) that pools all data. The singular value decomposition (SVD)-based product-space aggregation introduced by FlexLoRA is essential to this gain (naive factor averaging drops mean AUC by 0.097), whereas a drift-correcting optimizer (FedProx) shows no benefit over FedAvg in our single-seed runs, consistent with LoRA's low-rank updates already limiting client drift. Biomedical vision-language models can thus be adapted collaboratively across heterogeneous, geographically distributed institutions without centralizing data.
Chinese Translation
联邦学习(Federated Learning, FL)使各机构无需交换数据即可训练共享模型,而低秩适配(Low-Rank Adaptation, LoRA)仅需传输紧凑的低秩更新,使这一方法在大规模应用中切实可行。生物医学成像是这一组合的理想应用场景:患者数据受隐私法规保护而难以共享,且各机构在扫描设备、协议和计算资源方面差异显著。这种异质性引发了联邦LoRA更新应如何聚合的问题,随着多模态视觉-语言模型日益成为医学图像分析的核心,这一问题愈发重要。我们在来自三大洲(美国、越南、西班牙)的四个公开队列上,对BiomedCLIP用于胸片分类的联邦参数高效微调(Parameter-efficient fine-tuning, PEFT)进行了基准测试。联邦LoRA适配在全部四个队列上将共享类别AUC较未适配的BiomedCLIP主干模型有所提升(平均从0.687提升至0.802),表明性能增益来自联邦适配而非预训练模型的零样本能力。与各队列孤立的单队列训练相比,联邦学习改善了较弱队列的表现,同时基本保持了最强队列的性能,并接近汇集全部数据的集中式参考结果(0.812)。FlexLoRA提出的基于奇异值分解(SVD)的乘积空间聚合方法对这一增益至关重要(朴素的因子平均会使平均AUC下降0.097),而具有漂移校正功能的优化器(FedProx)在我们的单种子实验中相较于FedAvg并无优势,这与LoRA的低秩更新本身已能限制客户端漂移相一致。由此表明,生物医学视觉-语言模型可以在不集中数据的情况下,跨异质的、地理分布的机构进行协同适配。
cs.LG / 38 / 2609.02107

A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization

矢量量化、乘积量化与标量量化的统一率失真视角
Fang, Xianghong, Mou, Wenlong, Yuan, Yuan, Kong, Dehan, Rudner, Tim G. J.
Abstract
Discrete visual tokenization, predominantly driven by vector, scalar, and product quantization, lacks a unified conceptual framework for understanding quantization tradeoffs. In this paper, we propose a unified rate--distortion perspective on modern discrete visual tokenization. By viewing quantization as lossy compression, we characterize the nominal fixed-length coding rate through token count and codebook size, and quantization error as the distortion. Within this framework, we resolve three central questions. First, we theoretically and empirically show that minimizing distortion, rather than maximizing codebook utilization, is the primary intrinsic objective for reconstruction fidelity, with a direct connection to the STE-induced gradient discrepancy. Second, we establish two critical fairness conditions for intrinsic quantization comparison: controlling latent feature statistics and enforcing identical coding rates. Third, under these conditions, we recover the VQ--PQ--SQ distortion hierarchy in modern visual tokenization and show empirically that modern VQ methods achieve the lowest distortion. This work provides a foundational rate--distortion reframing of modern discrete visual tokenization, resolves ambiguities in quantizer evaluation, and provides a controlled framework for isolating intrinsic quantization effectiveness under fixed-rate constraints.
Chinese Translation
离散视觉分词(discrete visual tokenization)主要由矢量量化、标量量化和乘积量化驱动,但缺乏一个统一的概念框架来理解量化中的权衡。本文提出了一个针对现代离散视觉分词的统一率失真(rate-distortion)视角。通过将量化视为有损压缩,我们以词元数量和码本大小来刻画名义固定长度编码速率,并以量化误差作为失真。在该框架下,我们解决了三个核心问题。第一,我们从理论和实验上证明,最小化失真(而非最大化码本利用率)是决定重建保真度的主要内在目标,并且其与 STE(直通估计器)引起的梯度偏差存在直接联系。第二,我们确立了进行内在量化比较的两个关键公平性条件:控制潜在特征统计特性以及强制相同的编码速率。第三,在这些条件下,我们恢复了现代视觉分词中 VQ–PQ–SQ 的失真层级关系,并通过实验证明现代 VQ 方法能够实现最低的失真。这项工作为现代离散视觉分词提供了一个基础的率失真重构视角,消除了量化器评估中的歧义,并提供了一个在固定速率约束下分离内在量化有效性的受控框架。
cs.LG / 39 / 2609.02110

A Computational Comparison of Fourier Spectral Differentiation and Spatial Automatic Differentiation in Periodic Physics-Informed Neural Networks

周期物理信息神经网络中傅里叶谱微分与空间自动微分的计算比较
Liang, Xilai, Zhang, Zhao
Abstract
Physics-informed neural networks (PINNs) commonly evaluate the spatial derivatives appearing in partial differential equation residuals using automatic differentiation (AD), whose computational and memory costs can become substantial when multiple or high-order derivatives are required. We perform a controlled comparison of spatial AD and Fourier spectral differentiation in periodic physical-space PINNs. Within each paired experiment, the neural representation, temporal differentiation, optimizer, sampling procedure, and training schedule are held fixed, so that the two cases differ only in the spatial differentiation procedure. For the Fourier variant, network outputs are evaluated on a uniform periodic grid and transformed to Fourier space, where spatial derivatives are obtained through spectral multiplication and the same Fourier coefficients are reused across derivative orders. We compare the two procedures in standard PINNs for the Allen--Cahn and Korteweg--de Vries equations and in Causal PINNs for the Allen--Cahn, Korteweg--de Vries, and Kuramoto--Sivashinsky equations. Across these five equation--framework settings, Fourier differentiation yields mean paired end-to-end training speedups ranging from $2.90\times$ to $18.52\times$ and reduces peak allocated graphics processing unit (GPU) memory by $68.7\%$--$94.1\%$. The final relative $L_2$ errors remain of the same order, with neither differentiation procedure showing a consistent accuracy advantage. For the one-dimensional periodic benchmarks considered here, Fourier spectral differentiation therefore provides substantially lower training time and memory usage than spatial AD while retaining comparable solution error, at the cost of requiring a uniform structured spatial grid.
Chinese Translation
物理信息神经网络(Physics-Informed Neural Networks, PINNs)通常使用自动微分(Automatic Differentiation, AD)来计算偏微分方程残差中出现的空间导数,当需要多阶或高阶导数时,其计算和内存开销可能相当可观。我们在周期物理空间PINNs中对空间自动微分与傅里叶谱微分进行了受控比较。在每一组配对实验中,神经表示、时间微分、优化器、采样过程和训练日程均保持不变,使得两种情形仅在空间微分步骤上存在差异。对于傅里叶变体,网络输出在均匀周期网格上求值并变换到傅里叶空间,空间导数通过谱乘法获得,且相同的傅里叶系数在不同导数阶数之间被复用。我们在求解Allen--Cahn方程和Korteweg--de Vries方程的标准PINNs,以及求解Allen--Cahn、Korteweg--de Vries和Kuramoto--Sivashinsky方程的因果PINNs(Causal PINNs)中对这两种方法进行了比较。在这五种方程--框架组合设定中,傅里叶微分带来了介于 $2.90\times$ 至 $18.52\times$ 之间的平均配对端到端训练加速,并将GPU峰值显存占用降低了 $68.7\%$--$94.1\%$。最终相对 $L_2$ 误差保持在同一数量级,两种微分方法均未表现出一致的精度优势。因此,对于本文考虑的一维周期基准问题,傅里叶谱微分在保持相当解精度的同时,相比空间自动微分显著降低了训练时间和内存占用,其代价是需要均匀的结构化空间网格。
cs.LG / 40 / 2609.02126

Scalable Bayesian Optimization of Composite Functions for Image-Based Inverse Problems in Materials Characterization

用于材料表征中基于图像的逆问题的复合函数可扩展贝叶斯优化
Yoon, Dasol, Buathong, Poompol, Lee, Chia-Hao, Zhang, Yujia, Muller, David A., Frazier, Peter I.
Abstract
Estimating physical parameters from scientific images is a common inverse problem in materials characterization that often relies on expensive physics-based simulations. In electron microscopy, specimen thickness and crystal mistilt are critical parameters that govern how electrons scatter through the sample, and therefore the accuracy of any atomic-scale structure recovered from it. They are commonly inferred by matching experimental position-averaged convergent-beam electron diffraction (PACBED) patterns to simulated ones, but grid searches scale poorly and neural-network methods require extensive pretraining that may not transfer to new conditions. Here, we propose scalable Bayesian optimization of composite functions (SBOCF), a simulation-efficient method that exploits the known composite structure of the image-matching objective and the intermediate information contained in simulated images. By representing PACBED images with patch-level summaries and two correction terms, SBOCF preserves the original pixel-wise objective while reducing the number of modeled outputs from 24,649 to 11. Under a budget of 50 simulator evaluations, SBOCF outperformed standard Bayesian optimization with expected improvement on synthetic SrTiO3 benchmarks with thick and thin specimens, reducing the median final SSE by up to 290x in the thick-sample case. On experimental data, SBOCF produced parameter estimates consistent with previously reported values without task-specific pretraining. For a simulated mistilted specimen, using the SBOCF estimates in a downstream ptychographic reconstruction recovered sharp atoms that were otherwise blurred. These results establish SBOCF as a promising approach for inverse problems involving expensive simulators and high-dimensional structured outputs.
Chinese Translation
从科学图像中估计物理参数是材料表征中一类常见的逆问题,通常依赖于昂贵的基于物理的模拟。在电子显微学中,样品厚度和晶体失取向(mistilt)是决定电子如何穿透样品散射的关键参数,因此直接影响从图像中恢复原子尺度结构的准确性。这些参数通常通过将实验位置平均会聚束电子衍射(PACBED)图样与模拟图样进行匹配来推断,但网格搜索的扩展性差,而神经网络方法则需要大量预训练,且未必能迁移到新条件。本文提出了一种可扩展的复合函数贝叶斯优化方法(Scalable Bayesian Optimization of Composite Functions, SBOCF),这是一种模拟高效的方法,它利用图像匹配目标函数已知的复合结构以及模拟图像中包含的中间信息。通过以图像块级摘要和两个校正项来表示PACBED图像,SBOCF在保留原始逐像素目标函数的同时,将需建模的输出数量从24,649减少到11。在50次模拟器评估的预算下,在具有厚样品和薄样品的合成SrTiO3基准测试中,SBOCF优于采用期望改进(Expected Improvement)的标准贝叶斯优化,在厚样品情形下将最终SSE的中位数最多降低了290倍。在实验数据上,SBOCF无需任务特定的预训练即得到与先前报道值一致的参数估计。对于一个模拟失取向样品,在下游叠层衍射成像(ptychography)重建中使用SBOCF的估计值恢复了原本模糊的清晰原子。这些结果表明,SBOCF是求解涉及昂贵模拟器和高维结构化输出的逆问题的一种有前景的方法。
cs.LG / 41 / 2609.02145

Online Non-Monotone DR-Submodular Maximization Matching the Offline $0.401$ Factor

匹配离线0.401近似比的在线非单调DR-次模最大化
Aggarwal, Vaneet, Lu, Yiyang
Abstract
We study online maximization of nonnegative, non-monotone DR-submodular functions over compact convex down-closed subsets of the $d$-dimensional unit cube. The best known constructive offline approximation factor is $0.401$ under the corresponding meta-solvability assumptions, whereas comparable adversarial online guarantees had remained at $1/e$. We show that this factor is also achievable online. In the post-decision full-information value-oracle model, our algorithm attains factor $0.401$ with sublinear approximate regret when oracle feedback is conditionally unbiased and bounded. The online algorithm does not run the offline construction on a changing objective. Instead, it replaces the offline objective-dependent box step by a weighted online learner that controls the required residual terms cumulatively. An exact asymmetric balance theorem preserves the offline coefficients despite adversarial variation. The direct implementation has $O(T^{3/4})$ regret and uses $O(dT^{1/4})$ oracle calls per round. More generally, for every $\delta\in[0,1/4]$, batching gives $O(T^\delta)$ calls per round and $O(T^{4/5-\delta/5})$ regret, including a one-call $O(T^{4/5})$ endpoint. Under a positive-anchor condition, randomized blocking retains factor $0.401$ with $O(T^{5/6})$ one-point bandit regret.
Chinese Translation
我们研究了在d维单位立方体的紧致凸下封闭子集上,非负、非单调DR-次模函数(DR-submodular)的在线最大化问题。在相应的元可解性假设下,目前已知的构造性离线近似比为0.401,而可比较的对抗性在线保证一直停留在1/e。我们证明这一近似比在在线设定下同样可以达到。在决策后全信息价值预言机模型中,当预言机反馈条件无偏且有界时,我们的算法以次线性近似遗憾达到0.401的近似比。该在线算法并不在变化的目标函数上运行离线构造,而是用一个加权在线学习器替代离线算法中依赖于目标函数的箱式步进(box step),从而累积地控制所需的残差项。一个精确的非对称平衡定理使得即使面对对抗性变化,仍能保持离线系数。直接实现的遗憾为O(T^{3/4}),每轮使用O(dT^{1/4})次预言机调用。更一般地,对于每个δ∈[0,1/4],批处理方式可达到每轮O(T^δ)次调用和O(T^{4/5-δ/5})的遗憾,其中包括每轮仅一次调用、遗憾为O(T^{4/5})的端点情形。在正锚点(positive-anchor)条件下,随机化阻断(randomized blocking)以O(T^{5/6})的单点赌博机遗憾保持0.401的近似比。
cs.LG / 42 / 2609.02155

Exact Limits of Random Projections for Preserving Geometry: Distance Recovery, Nearest-Neighbor Rankings, and Covariance Shape in Gaussian Models

保持几何结构的随机投影的精确极限:高斯模型中的距离恢复、最近邻排序与协方差形状
Sao, Piyush
Abstract
The Johnson-Lindenstrauss (JL) lemma guarantees that a random projection of $n$ points to $m=O(\varepsilon^{-2}\log n)$ dimensions preserves pairwise squared distances within relative error $\varepsilon$ with high probability, and this dimension order is asymptotically optimal. In high dimensions, however, distances concentrate around a baseline while key geometric information lies in much smaller fluctuations. We show that the JL bound can therefore be uninformative about retained geometry: an independent Gaussian replacement map can satisfy it even though the replacement cloud is independent of the original data. We then ask how well any decoder can recover a feature $f(D)$ of a squared distance $D$ from a linear sketch. Under squared-error loss, the optimal decoder is conditional expectation, so recovery defines a linear operator whose singular values quantify feature recovery. For isotropic Gaussian data ($\Sigma=\sigma^2 I_d$), we diagonalize this operator in closed form. For fixed $k$ with $m,d-m\to\infty$, its $k$th singular value satisfies $\ell_k\approx(m/ d)^{k/2}$. This yields three sharp consequences. A rank-$m$ sketch retains at most an $m/d$ fraction of the variance of any feature of one squared distance. If $m\to\infty$ and $m/d\to0$, the expected Kendall correlation is $\frac{2}{\pi}\sqrt{m/d}(1+o(1))$; for fixed $q$, nearest- neighbor agreement tends to $1/q$. Yet one projection can satisfy the JL bound while mean Kendall correlation vanishes when $\log n\ll m\ll d$. After removing scale, Haar-averaged retained covariance-shape information is $(m/d)^2$. Thus JL distance preservation does not quantify the geometry available for comparison or inference.
Chinese Translation
Johnson-Lindenstrauss(JL)引理保证,将 $n$ 个点随机投影到 $m=O(\varepsilon^{-2}\log n)$ 维时,能以高概率在相对误差 $\varepsilon$ 内保持两两点之间的平方距离,且该维度量级在渐近意义上是最优的。然而,在高维情形下,距离集中于某个基准值附近,而关键的几何信息却蕴含在幅度小得多的波动之中。我们证明,因此 JL 界对于所保留的几何结构可能是不具信息量的:一个独立的高斯替换映射即使在与原始数据完全独立的情况下也能满足该界。接着,我们研究任意解码器从线性草图中恢复平方距离 $D$ 的某个特征 $f(D)$ 的能力。在平方误差损失下,最优解码器为条件期望,因此恢复问题定义了一个线性算子,其奇异值刻画了特征恢复的效果。对于各向同性高斯数据($\Sigma=\sigma^2 I_d$),我们以闭式形式将该算子对角化。当 $k$ 固定且 $m,d-m\to\infty$ 时,其第 $k$ 个奇异值满足 $\ell_k\approx(m/d)^{k/2}$。由此得到三个尖锐的结论。一个秩为 $m$ 的草图至多保留单个平方距离的任意特征方差的 $m/d$ 比例。若 $m\to\infty$ 且 $m/d\to0$,则期望 Kendall 相关系数为 $\frac{2}{\pi}\sqrt{m/d}(1+o(1))$;对于固定的 $q$,最近邻一致率趋于 $1/q$。然而,当 $\log n\ll m\ll d$ 时,一个投影可以满足 JL 界,而平均 Kendall 相关系数却趋于消失。在去除尺度之后,Haar 平均意义下所保留的协方差形状信息为 $(m/d)^2$。因此,JL 距离保持性并不能量化可用于比较或推断的几何信息。
cs.LG / 43 / 2609.02160

GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories

GeoSPRINT:面向扩散轨迹推理的几何冗余感知步数剪枝方法
Joshi, Arpita
Abstract
Diffusion models achieve high sample quality but remain expensive at inference time because sampling requires many sequential neural function evaluations (NFEs). Existing acceleration methods either use fixed step-skipping schedules, adapt step sizes based on local numerical error, or require additional training. We introduce GeoSPRINT (Geometric Step Pruning for Inference in Trajectories), a training-free framework for constructing non-uniform sampling schedules from the geometry of denoising trajectories. GeoSPRINT detects geometrically redundant steps using a hyperplanarity test in latent space, implemented efficiently via QR factorization, and converts the resulting redundancy profile into a sampling schedule that allocates more steps to high-curvature regions of the trajectory. In addition, we introduce the trajectory projection score $\alpha_{\mathrm{traj}}$, a residual-variance metric that quantifies trajectory straightness and serves as a model-free diagnostic for rectified flow quality. Across CIFAR-10 ($32{\times}32$), LSUN Church ($256{\times}256$), and Stable Diffusion v1.5 ($512{\times}512$ latent), GeoSPRINT consistently improves over uniform DDIM (Denoising Diffusion Implicit Models) schedules at matched NFE budgets. On CIFAR-10, GeoSPRINT improves FID (Fr\'echet Inception Distance) by 0.7-1.1 over DDIM across 49-89 NFEs and surpasses DPM-Solver++ at NFE${\geq}30$ despite using a first-order DDIM solver. On LSUN Church, it reduces FID from 1.48 to 1.26 at 52 steps, and on Stable Diffusion v1.5 it achieves up to 1.93 FID improvement over DDIM. These results show that trajectory geometry provides a useful global signal for allocating inference steps and that schedule quality can substantially improve diffusion sampling efficiency without retraining.
Chinese Translation
扩散模型能够生成高质量的样本,但由于采样过程需要大量顺序的神经函数评估(NFE),其推理成本仍然很高。现有的加速方法要么使用固定的跳步调度,要么基于局部数值误差自适应调整步长,要么需要额外的训练。我们提出了GeoSPRINT(Geometric Step Pruning for Inference in Trajectories,面向轨迹推理的几何步数剪枝),这是一个无需训练的框架,通过利用去噪轨迹的几何结构来构建非均匀采样调度。GeoSPRINT通过在潜空间中进行超平面性检验(利用QR分解高效实现)来检测几何上冗余的步骤,并将得到的冗余分布转换为采样调度,将更多的采样步骤分配给轨迹中高曲率的区域。此外,我们引入了轨迹投影分数 $\alpha_{\mathrm{traj}}$,这是一个量化轨迹直线程度的残差方差指标,可作为校正流(rectified flow)质量的无模型诊断工具。在CIFAR-10($32{ imes}32$)、LSUN Church($256{ imes}256$)和Stable Diffusion v1.5($512{ imes}512$潜空间)上,在相同的NFE预算下,GeoSPRINT始终优于均匀的DDIM(Denoising Diffusion Implicit Models)调度。在CIFAR-10上,在49-89个NFE范围内,GeoSPRINT相比DDIM将FID(Fr'echet Inception Distance)改善了0.7-1.1;尽管仅使用一阶DDIM求解器,在NFE${\geq}30$时仍超越了DPM-Solver++。在LSUN Church上,它在52步时将FID从1.48降至1.26;在Stable Diffusion v1.5上,相比DDIM最多实现了1.93的FID改善。这些结果表明,轨迹几何为推理步骤的分配提供了有用的全局信号,并且通过优化调度质量可以在不重新训练的情况下显著提升扩散采样的效率。
cs.LG / 44 / 2609.02170

DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation

DMRL:面向广告推荐技能优化的文档介导强化学习
Zhang, Wei, Li, Hongji, Sun, Song, Yu, Peng, Yang, Xue, Zhao, Lei, Jiang, Peng
Abstract
Advertising recommendation requires continuously tuning complex system parameters while balancing commercial returns and user experience. Recent work has introduced large language models (LLMs) with skill documents to assist this labor-intensive process, but skill optimization remains largely prompt-driven, lacking a principled mechanism to attribute rewards to specific document edits. To address this limitation, we propose Document-Mediated Reinforcement Learning (DMRL), a skill self-evolution framework that models skill document optimization as a sequence of structured editing actions. In DMRL, an upper-level agent performs controlled document edits, while a frozen lower-level task agent evaluates their effects through A/B testing. To address credit assignment and long-term outcomes, we introduce two key components: (1) Dual-Relative Policy Optimization (DRPO), a post-training policy optimization method for robust and risk-aware advantage estimation; and (2) Long-term Reward Predictor (LRP), which estimates long-term outcomes by modeling population heterogeneity with disentangled representation learning and cross-attention transfer. DMRL was deployed on a large-scale short-video ads platform and extensive empirical evaluation shows that DMRL outperforms state-of-the-art baselines across key advertising metrics
Chinese Translation
广告推荐需要在平衡商业回报与用户体验的同时,持续调整复杂的系统参数。近期研究引入了配备技能文档的大语言模型(LLM)来辅助这一劳动密集型过程,但技能优化仍在很大程度上依赖提示词驱动,缺乏将奖励归因于具体文档编辑的原则性机制。针对这一局限,我们提出文档介导强化学习(Document-Mediated Reinforcement Learning,DMRL),这是一个将技能文档优化建模为一系列结构化编辑动作的技能自进化框架。在DMRL中,上层智能体执行受控的文档编辑,而冻结的下层任务智能体通过A/B测试评估其效果。为解决信用分配与长期结果问题,我们引入两个关键组件:(1)双相对策略优化(Dual-Relative Policy Optimization,DRPO),一种用于稳健且风险感知的优势估计的训练后策略优化方法;(2)长期奖励预测器(Long-term Reward Predictor,LRP),通过解耦表示学习与交叉注意力迁移建模人群异质性,从而估计长期结果。DMRL已部署于大规模短视频广告平台,大量实证评估表明,DMRL在关键广告指标上优于最先进的基线方法。
cs.LG / 45 / 2609.02194

Learning the Constitutive Behavior of Materials via Neural Operators and Causal Attention: Case Studies in Plasticity and Damage

基于神经算子与因果注意力学习材料本构行为:塑性及损伤案例研究
Arora, Rishabh, Scheunemann, Lisa, Brepols, Tim, Rezaei, Shahed
Abstract
Classical constitutive modeling of path-dependent inelastic materials relies on internal state variables whose evolution equations must be postulated based on domain knowledge and calibrated against experimental data. However, in many practical settings, the relevant internal variables are typically not measurable in experiments, and the constitutive response must be inferred entirely from measured strain-stress data without any prior knowledge of the material's internal state. We propose a data-driven constitutive modeling framework based on the concept of a material operator, which treats a deforming material as a functional mapping from its entire strain history to the corresponding stress response. In contrast to traditional autoregressive or recurrent formulations, the model is trained directly on full loading paths as function-to-function mappings, predicting complete stress trajectories in a single parallel forward pass. Temporal path dependence is enforced through a causally masked attention mechanism embedded within the operator, which restricts the model's attention to past material states while preserving computational parallelizability. Spectral convolutions provide discretization-invariant representations in the frequency domain, while causal attention captures highly adaptive, non-local history dependence. Furthermore, sinusoidal activation functions are used to resolve the strong nonlinear transitions inherent in inelastic regimes. The framework is evaluated across multidimensional, rate-independent material models exhibiting complex phenomena, with an emphasis on nonlinear plasticity and ductile damage accumulation. The results demonstrate accurate and robust predictions of irreversible deformation mechanisms while simultaneously achieving resolution invariance and excellent parallel efficiency.
Chinese Translation
经典的路径依赖非弹性材料本构建模依赖于内状态变量,其演化方程必须基于领域知识进行假设,并利用实验数据进行标定。然而,在许多实际场景中,相关的内部变量通常无法在实验中测量,本构响应必须完全从实测的应变—应力数据中推断,而无需任何关于材料内部状态的先验知识。我们提出了一种基于材料算子(material operator)概念的数据驱动本构建模框架,该框架将变形材料视为一个从其完整应变历史到相应应力响应的泛函映射。与传统的自回归或循环(递归)公式不同,该模型直接在完整加载路径上以函数到函数映射的方式进行训练,可在单次并行前向传播中预测完整的应力轨迹。时间路径依赖性通过嵌入算子内部的因果掩码注意力机制(causally masked attention)来实现,该机制将模型的注意力限制在过去材料状态上,同时保持计算的可并行性。谱卷积(spectral convolutions)在频域中提供了离散化不变的表达,而因果注意力则捕捉高度自适应的非局部历史依赖性。此外,采用正弦激活函数来解决非弹性区段中固有的强非线性转变。该框架在呈现复杂现象的多维、率无关材料模型上进行了评估,重点针对非线性塑性和韧性损伤累积。结果表明,该模型能够对不可逆变形机制做出准确而稳健的预测,同时实现了分辨率不变性和优异的并行效率。
cs.LG / 46 / 2609.02203

SMart: A Multi-source Multi-phase Time Series Representation Transfer Framework

SMart:一个多源多阶段时间序列表示迁移框架
He, Fang, Lee, Wang-chien
Abstract
Time series representation learning (TSRL) has attracted growing research interests in recent years. Two recent explorations in TSRL are: i) exploiting a transformer-based framework to learn time series; ii) instead of using only the targeted dataset, borrowing time series from other datasets to to facilitate representation transfer. While these two explorations are shown effective, the self-supervised time series recovery task in (i) and the single-source dataset used in (ii) are technically simple and thus can be enhanced with new ideas. In this work, we propose a new TSRL framework, namely multi-source multi-phase time series representation transfer (SMart), which has two novel mechanisms to address the aforementioned deficiencies: 1) a multi-phase recurrence plots recovery task, in three alternative modes, for guiding the encoder to embed time series dynamics into the time series representation; and 2) a source dataset selector to select multiple suitable source datasets to supplement the original target dataset for pre-training the TSRL encoder. Experimental results show that SMart outperforms several state-of-the-art models for time series representation learning, classification and regression on both uni-variate and multi-variate time series datasets, reducing mean absolute error up to 19.5% for time series regression, and increasing average accuracy up to 1.34\% for time series classification.
Chinese Translation
近年来,时间序列表示学习(TSRL)引起了越来越多的研究兴趣。TSRL领域近期的两个探索方向是:i)利用基于Transformer的框架来学习时间序列;ii)不仅使用目标数据集,还从其他数据集中借入时间序列以促进表示迁移。虽然这两种探索已被证明是有效的,但(i)中的自监督时间序列恢复任务和(ii)中使用的单一源数据集在技术上较为简单,因此可以通过新的思路加以改进。本文提出了一种新的TSRL框架,即多源多阶段时间序列表示迁移框架(SMart),该框架包含两种新颖的机制以解决上述不足:1)多阶段递归图恢复任务,具有三种可选模式,用于引导编码器将时间序列的动态特性嵌入到时间序列表示中;2)源数据集选择器,用于选择多个合适的源数据集来补充原始目标数据集,从而对TSRL编码器进行预训练。实验结果表明,SMart在单变量和多变量时间序列数据集上的时间序列表示学习、分类和回归任务中均优于多个最先进的模型,在时间序列回归任务中将平均绝对误差最多降低19.5%,在时间序列分类任务中将平均准确率最多提升1.34%。
cs.LG / 47 / 2609.02237

Recursive Value Learning for Long-Horizon Offline Goal-Conditioned RL

面向长时程离线目标条件强化学习的递归价值学习
Jeon, Hyeonseong, Lee, Youngwoon
Abstract
Scaling offline goal-conditioned reinforcement learning (GCRL) to long-horizon tasks is difficult because (1) long-range value learning depends on shorter-range estimates that may still be inaccurate, and (2) max-based value backups can amplify overestimation through repeated propagation. We propose DCRL (Divide-and-Conquer RL), which recursively decomposes each trajectory segment into a balanced binary tree and trains the values from leaves to root. Each parent is therefore updated only after its children, using an exact factorization of the observed route rather than selecting among noisy alternatives. Since this objective learns values along demonstrated routes that are not necessarily optimal, DCRL jointly propagates values across trajectories to discover shorter routes. Thanks to the balanced binary tree, DCRL reduces worst-case bootstrap depth from linear to logarithmic, and this shorter dependency structure empirically corresponds to much slower error accumulation. Across diverse goal-reaching tasks, DCRL substantially outperforms prior flat offline GCRL methods, and on the five most challenging long-horizon OGBench tasks, it improves the best prior average score from 55 to 64, surpassing all flat and hierarchical baselines.
Chinese Translation
将离线目标条件强化学习(GCRL)扩展到长时程任务十分困难,原因在于:(1)长距离价值学习依赖于可能仍然不准确短距离估计;(2)基于最大值的价值回溯会通过反复传播放大过估计。我们提出DCRL(分治强化学习,Divide-and-Conquer RL),它将每条轨迹段递归分解为一棵平衡二叉树,并从叶节点到根节点逐步训练价值。因此,每个父节点仅在其子节点完成更新后才被更新,使用观测路径的精确因式分解,而不是在含噪声的备选方案中进行选择。由于该目标函数沿演示路径学习价值,而这些路径未必最优,DCRL同时在轨迹之间传播价值以发现更短的路径。得益于平衡二叉树结构,DCRL将最坏情况下的自举深度从线性降低到对数级,且这种更短的依赖结构在实证上对应于显著更慢的误差累积。在多样的目标达成任务中,DCRL显著优于先前的扁平离线GCRL方法;在五个最具挑战性的长时程OGBench任务上,它将此前最佳平均得分从55提升至64,超越了所有扁平与层次化基线方法。
cs.LG / 48 / 2609.02241

Similarity-Aware Personalized Federated Learning in Heterogeneous Environments

异构环境下基于相似性感知的个性化联邦学习
A V, Arun Kumar, Gupta, Sunil, Ngyuen, Dang, Duong, Bao, Trong, Dat Phan
Abstract
Federated Learning (FL) allows decentralized clients to train models collaboratively while preserving data privacy. However, distribution mismatch across clients often leads to poor global generalization and degraded local client-level performance. In such scenarios, some of the clients with their local models trained solely on local data may perform better than the globally learnt model, thus nullifying the benefits of collaborative federated learning. To address this, we propose SAPE-FL (Similarity-Aware Personalized Federated Learning), a novel personalization framework that anchors each client's model to both the global model and a similarity-weighted peer averaged model. By incorporating dynamic, client-specific regularization based on both model similarity and output similarity, SAPE-FL adaptively balances global knowledge transfer and peer collaboration while filtering out dissimilar clients. This dual anchoring mitigates negative transfer and enhances robustness in heterogeneous settings. We theoretically analyze our algorithm establishing its convergence guarantees and empirically show that SAPE-FL outperforms state-of-the-art methods under high statistical heterogeneity and low client data regimes.
Chinese Translation
联邦学习使去中心化的客户端能够在保护数据隐私的同时协同训练模型。然而,客户端之间的分布不匹配常常导致全局泛化能力差,并降低客户端本地的性能。在这种情况下,一些仅基于本地数据训练本地模型的客户端可能表现得优于全局学习到的模型,从而抵消了协作式联邦学习的优势。为解决这一问题,我们提出了SAPE-FL(相似性感知个性化联邦学习,Similarity-Aware Personalized Federated Learning),这是一种新颖的个性化框架,它将每个客户端的模型同时锚定到全局模型和基于相似性加权的对等平均模型上。通过结合基于模型相似性和输出相似性的动态的、客户端特定的正则化,SAPE-FL 能够在过滤掉不相似客户端的同时,自适应地平衡全局知识迁移与对等协作。这种双重锚定机制缓解了负迁移问题,并增强了在异构环境下的鲁棒性。我们在理论上对该算法进行了分析,建立了其收敛性保证,并通过实验表明,SAPE-FL 在高统计异质性和客户端数据量较少的情况下优于当前最先进的方法。
cs.LG / 49 / 2609.02265

CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents

CAPTURE:在个性化LLM智能体中解耦偏好漂移与记忆投毒
Hossain, S M Asif, Shayoni, Ruksat Khan, Morol, Md Kishor
Abstract
Personalized language agents use persistent memory to adapt to users over time, but the same mechanism creates an attack surface. When new information conflicts with stored preferences, an agent must distinguish genuine preference drift from temporary context shifts, ambiguity, or adversarial memory poisoning. We formulate this problem as a continuous-time partially observable decision process over a latent user state and show why rules based only on recency and provenance are insufficient. CAPTURE addresses this ambiguity with a neural differential-equation belief tracker, a multi-timescale memory ledger, uncertainty-triggered clarification, and counterfactual auditing of cited memories. On 480 held-out episodes from 96 users, CAPTURE achieves a 71.5% win rate, compared with 69.3% for an identically supervised baseline and 66.1% for the strongest heuristic baseline. It limits fixed-policy poisoning success to 11.5% while accepting 83.5% of genuine preference updates. Under an adaptive attacker with access to the released weights, attack success rises to 24.7%, exposing a real adaptation-security tradeoff. We further evaluate the frozen system zero-shot on an independently constructed benchmark and replay longitudinal interaction histories from 40 users collected over two to three weeks. These results suggest that modeling preference authenticity explicitly can improve both personalization and robustness in memory-augmented LLM agents.
Chinese Translation
个性化语言智能体利用持久记忆来随时间适应用户,但这一机制也带来了攻击面。当新信息与已存储的偏好发生冲突时,智能体必须区分真实的偏好漂移与临时的上下文变化、歧义或对抗性记忆投毒。我们将该问题形式化为一个建立在潜在用户状态之上的连续时间部分可观测决策过程,并说明为什么仅基于时近性和来源的规则是不够的。CAPTURE 通过以下方式解决这一歧义:神经微分方程信念跟踪器、多时间尺度记忆账本、由不确定性触发的澄清提问,以及对所引用记忆的反事实审计。在来自96位用户的480个保留测试片段上,CAPTURE 取得了71.5%的胜率,而采用相同监督的基线为69.3%,最强的启发式基线为66.1%。它将固定策略下的投毒成功率限制在11.5%,同时接受83.5%的真实偏好更新。在拥有已发布模型权重的自适应攻击者面前,攻击成功率升至24.7%,揭示出一个真实存在的适应性-安全性权衡。我们进一步在独立构建的基准上对冻结系统进行零样本评估,并重放了从40位用户收集的为期两到三周的长程交互历史。这些结果表明,显式建模偏好真实性可以同时提升记忆增强型LLM智能体的个性化能力和鲁棒性。
cs.LG / 50 / 2609.02285

Entangled Representations Amplify Collateral Damage in Unlearning

纠缠表征加剧机器遗忘中的附带损害
Wybitul, Evžen, Rudner, Tim G. J., de Witt, Christian Schroeder
Abstract
A long-held intuition in interpretability research is that representational entanglement, the sharing of structure between knowledge domains in a neural network, makes unlearning harder. While the intuition is widespread, it has never been directly tested in a controlled experiment. We present a way to do so: by repurposing Selective Gradient Masking (SGTM), we train a suite of six 254M-parameter language models on English Wikipedia with graded levels of disentanglement between biology and non-biology knowledge. Applying three standard unlearning methods to every model in the suite, we find that more disentangled models consistently achieve better retain-forget trade-offs: at a fixed level of forgetting, the most disentangled models incur roughly $4\times$ lower retain cost under two of the three methods, and $1.3\times$ lower under the third. Because our intervention changes only the model, not the data or the unlearning algorithm, this is direct evidence that representational entanglement is one of the causes of collateral damage in unlearning, as interpretability researchers have long suspected. A similar design could be used to test other structural claims from interpretability.
Chinese Translation
可解释性研究中一个长期存在的直觉是,表征纠缠——即神经网络中不同知识域之间结构的共享——会使机器遗忘(unlearning)变得更加困难。尽管这一直觉广为流传,但从未在受控实验中得到直接检验。我们提出了一种检验方法:通过改造选择性梯度掩码(Selective Gradient Masking, SGTM),我们在英文维基百科语料上训练了六个具有254M参数的语言模型,这些模型中生物学知识与非生物学知识之间的解纠缠程度呈梯度变化。我们将三种标准机器遗忘方法应用于该模型套件中的每一个模型,发现解纠缠程度更高的模型始终能取得更好的保留-遗忘权衡:在固定的遗忘水平下,在三种方法中的两种下,解纠缠程度最高的模型的保留成本约降低4倍,在第三种方法下降低约1.3倍。由于我们的干预只改变了模型本身,而不改变数据或遗忘算法,这直接证明了表征纠缠是机器遗忘中附带损害的原因之一,正如可解释性研究者长期怀疑的那样。类似的设计也可用于检验可解释性领域的其他结构性假设。
cs.LG / 51 / 2609.02293

SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment

SEAL:通过共享专家对齐强化混合专家模型的全局安全性
Meng, Qingyu, Zha, Yiwei, Pei, Jiahuan, Hindriks, Koen, Bos, Herbert, Chen, Min
Abstract
Mixture-of-Experts (MoE) is a scaling architecture for large language models that activates only a small subset of expert modules per token, enabling massive parameter growth with nearly constant computation. Recent Hybrid MoE architecture adds \textit{shared experts} to capture consistently useful representations, further improving stability and generalization. MoE now powers many flagship open-source and commercial models, yet remains vulnerable to adversarial attacks. Specifically, sparse routing introduces a structural vulnerability: MoE safety hinges on which experts are activated, and adversaries can subvert this selection through jailbreak prompts, malicious fine-tuning, and weight-level pruning of safety-critical neurons. Existing defenses primarily focus on hardening the router, but an adversary may still manipulate or bypass the routing trajectory due to the routing process's nondeterministic nature, thereby collapsing the defense. To cope with this problem, we first identify theoretically and empirically that shared expert, an always-activated component containing a small proportion of safety-critical neurons, can overcome the uncertainty of sparsely activated routing path and serve as a router-independent anchor to enhance global safety alignment. Based on this insight, we propose SEAL, a training-time parameter-efficient defense that produces a plug-and-play adapter attached to shared expert, and SEAL++, a variant that adds an orthogonal constraint preserving pre-existing safety subspaces during training. We evaluate SEAL and SEAL++ across six attack scenarios that combine three adversarial inputs (harmful prompting, jailbreak, malicious fine-tuning) with and without neuron pruning. SEAL reduces attack success rate (ASR) by up to 60\%, at a capability cost of at most 1.4\% on a five-benchmark average. Additionally, SEAL can seamlessly integrate with router-level ......
Chinese Translation
混合专家架构是一种用于大语言模型的扩展架构,它对每个词元仅激活一小部分专家模块,从而在计算量近乎恒定的情况下实现参数规模的大幅增长。近期的混合MoE架构引入了共享专家来捕获持续有用的表示,进一步提升了稳定性和泛化能力。目前MoE已支撑众多旗舰级开源和商业模型,但仍然容易受到对抗攻击。具体而言,稀疏路由引入了结构性漏洞:MoE的安全性取决于哪些专家被激活,而攻击者可以通过越狱提示、恶意微调以及对安全关键神经元的权重级剪枝来颠覆这一选择过程。现有防御主要集中于加固路由器,但由于路由过程的非确定性,攻击者仍可能操纵或绕过路由轨迹,从而使防御失效。为应对这一问题,我们首先从理论和实证上证明:共享专家作为一个始终被激活、包含少量安全关键神经元的组件,能够克服稀疏激活路由路径的不确定性,并可作为与路由器无关的锚点来增强全局安全对齐。基于这一洞察,我们提出了SEAL,一种训练时参数高效的防御方法,它生成一个可即插即用、附加于共享专家的适配器;同时提出其变体SEAL++,该变体增加了一个正交约束,以在训练过程中保留既有的安全子空间。我们在六种攻击场景下评估了SEAL和SEAL++,这些场景将三种对抗输入(有害提示、越狱攻击、恶意微调)与有无神经元剪枝相结合。实验表明,SEAL可将攻击成功率降低最多60%,而在五个基准上的平均能力损失最多仅为1.4%。此外,SEAL可与路由器级防御方法无缝集成……
cs.LG / 52 / 2609.02304

Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators

贝叶斯最优误码率(BER)与AUC:估计及其估计器的评估
Ushio, Ryota, Ishida, Takashi, Sugiyama, Masashi
Abstract
A fundamental quantity in machine learning is the optimal performance achievable by any model on a given task. Estimating this quantity allows us to distinguish the irreducible part of the error from a deficiency of the model, telling us how much room for improvement remains. Recent work has shown that the Bayes error, or equivalently the optimal accuracy, can be estimated from soft labels in binary classification. However, accuracy is often a poor summary of performance in settings with severe class imbalance or noisy annotations, where metrics such as the balanced error rate (BER) and the area under the ROC curve (AUC) are more appropriate. We address this gap with two complementary contributions. (i) Estimation. We propose soft-label-based estimators for the optimal BER and AUC. We first consider the clean setting in which true soft labels and the class prior are known, and then extend the estimators to a more realistic setting in which the class prior is unknown and the observed soft labels are corrupted by an unknown order-preserving transformation, possibly followed by additive noise. In the latter setting, we approximately recover the clean soft labels via isotonic regression with auxiliary hard labels, estimate the class prior with a clipped mean of the hard labels, and derive finite-sample error bounds for the resulting plug-in estimators. (ii) Evaluation. Since the optimum is unobservable on real datasets, evaluating any such estimator is itself nontrivial. We extend the FeeBee framework, originally proposed for evaluating Bayes-error estimators, to the optimal BER and AUC. The resulting procedure provides practical evaluation scores without requiring knowledge of the optimum, and applies to any estimator of the optimal BER or AUC, not only our proposed ones. Experiments on synthetic and real-world datasets validate both the estimators and the evaluation procedure.
Chinese Translation
机器学习中的一个基本量是任何模型在给定任务上可达到的最优性能。估计该量使我们能够将误差中不可消除的部分与模型本身的不足区分开来,从而告诉我们还有多大的改进空间。近期研究表明,贝叶斯误差(即等价的最优准确率)可以从二分类中的软标签中估计得到。然而,在类别严重不平衡或标注存在噪声的场景下,准确率往往不是衡量性能的良好指标,此时均衡错误率(balanced error rate, BER)和ROC曲线下面积(area under the ROC curve, AUC)等指标更为合适。我们通过两项互补的贡献来填补这一空白。(一)估计。我们提出了基于软标签的最优BER和AUC估计器。我们首先考虑真实软标签和类先验均已知的干净设定,然后将估计器扩展到一个更现实的设定:类先验未知,且观测到的软标签被某个未知的保序变换(可能随后叠加加性噪声)所污染。在后一设定中,我们通过结合辅助硬标签的保序回归(isotonic regression)近似恢复干净的软标签,用硬标签的截断均值估计类先验,并为由此得到的代入式(plug-in)估计器推导了有限样本误差界。(二)评估。由于在真实数据集上最优值是不可观测的,评估任何此类估计器本身就是非平凡的问题。我们将最初为评估贝叶斯误差估计器而提出的FeeBee框架扩展至最优BER和AUC。由此得到的评估流程无需知道最优值即可提供实用的评估分数,并且适用于任何最优BER或AUC估计器,而不仅限于我们提出的方法。在合成数据集和真实数据集上的实验验证了所提出的估计器和评估流程的有效性。
cs.LG / 53 / 2609.02322

What Is Worth Representing? Representational Empowerment for Continual Model Construction

什么值得表征?面向持续模型构建的表征赋权
Dai, Fei, Zhou, Hanqi, Gopnik, Alison, Wu, Charley
Abstract
The first problem of modeling the world is not just estimating the right parameters or causal structure, but deciding what should be represented at all. We frame this problem as continual model construction: an agent maintains an environment-specific model M of an inaccessible world W and curates a persistent library L of reusable representational elements across environments. We propose Representational Empowerment (RepEmp) to score candidate elements by how much they expand the agent's future capacity to model and plan, complementing the classic definition of empowerment, but redefined as control over internal representations instead of external states. We realize the framework as a hierarchical Curator-Actor architecture and test it across three experiments. In a closed-vocabulary causal-learning task, human participants construct causal models at varying abstraction granularities to maximize goal reachability rather than fidelity to the world, a signature better predicted by RepEmp than by information-gain alternatives. Matched simulations reveal that RepEmp-guided construction contributes more than exploration to sufficient structure recovery and cross-task transfer. Finally, in an open-vocabulary planning domain, an LLM-augmented Curator builds more compact symbolic libraries, which also generalize better than baselines. Ablating RepEmp eliminates these benefits. Together, these results identify RepEmp as a key principle for continual model construction: deciding what to build, retain, and reuse under bounded resources.
Chinese Translation
对世界建模的首要问题不仅仅是估计正确的参数或因果结构,而是决定究竟应该表征什么。我们将该问题框架化为持续模型构建:智能体维护一个关于不可直接访问的世界 W 的环境特定模型 M,并在跨环境中持续管理一个由可复用表征元素组成的持久库 L。我们提出表征赋权(Representational Empowerment,RepEmp),通过候选元素能在多大程度上扩展智能体未来建模与规划的能力来对其进行评分,这是对经典赋权概念的补充——但将其重新定义为对内部表征而非外部状态的控制。我们将该框架实现为分层的“策展者-行动者”(Curator-Actor)架构,并通过三个实验对其进行检验。在一个封闭词汇因果学习任务中,人类被试以不同的抽象粒度构建因果模型,其目标是最大化目标可达性而非忠实还原世界——这一特征由 RepEmp 的预测优于基于信息增益的替代方案。匹配的模拟实验表明,RepEmp 引导的构建对充分的结构恢复与跨任务迁移的贡献大于探索。最后,在一个开放词汇规划领域中,由大语言模型(LLM)增强的策展者构建了更紧凑的符号库,且其泛化能力优于基线方法。消融 RepEmp 则会消除上述所有优势。综上,这些结果将 RepEmp 确立为持续模型构建的一个关键原则:即在资源受限条件下决定构建什么、保留什么以及复用什么。
cs.LG / 54 / 2609.02339

AGI Maze Prediction Datasets: A Compact Benchmark for Learning World Dynamics with Transformers

AGI 迷宫预测数据集:一个用于基于 Transformer 学习世界动力学的紧凑基准
Potapov, Alexey
Abstract
World modeling requires a predictive model to maintain and update an internal state adequate for reasoning about the consequences of actions. We introduce the AGI Maze Prediction Datasets and Benchmark, a lightweight controlled testbed for studying this capability in Transformers and other predictive models. Derived from procedurally generated, stateful grid worlds, the benchmark comprises per-step transition prediction, fixed-horizon state prediction, and sequential textual-observation prediction. Source-maze-disjoint training and validation splits, together with greedy exact-match evaluation, distinguish learning transferable action-conditioned dynamics from memorizing transitions in familiar layouts. We establish from-scratch byte-level Transformer baselines and compare them with two working-memory-augmented architectures. A generic auxiliary latent-memory Transformer can fit some training sets perfectly but does not consistently improve held-out performance. In contrast, a pseudo-video spatial-memory Transformer initializes a two-dimensional latent workspace from the input map and updates it from action history without receiving intermediate maps, positions, or state labels. Under the same data, objectives, and evaluation protocol, this model reaches perfect validation accuracy on selected fixed-horizon tasks where the byte and unstructured-memory baselines do not, and substantially improves sequential text-trace prediction. These results suggest that structured, task-aligned working memory can be more useful than additional latent capacity alone. More broadly, we argue that language grounding is mediated by persistent data structures and computations over them; the benchmark offers a compact setting for testing architectures that couple textual interfaces to learned structured state.
Chinese Translation
世界建模要求预测模型能够维护并更新一个内部状态,该状态足以对动作的后果进行推理。我们提出了 AGI 迷宫预测数据集与基准(AGI Maze Prediction Datasets and Benchmark),这是一个轻量级、可控的测试平台,用于研究 Transformer 及其他预测模型的这一能力。该基准基于程序化生成的、带状态的网格世界构建,包含逐步转移预测、固定时域状态预测以及顺序文本观测预测三项任务。通过源迷宫不相交的训练与验证划分,结合贪心精确匹配评估,该基准能够将学习到的可迁移的动作条件动力学与对熟悉布局中转移的记忆区分开来。我们建立了从零训练的字节级 Transformer 基线,并将其与两种工作记忆增强架构进行比较。一种通用的辅助潜记忆 Transformer 可以完美拟合某些训练集,但并不能持续提升留出集性能。与之相反,一种伪视频空间记忆 Transformer 从输入地图初始化一个二维潜在工作空间,并根据动作历史对其进行更新,而无需接收中间地图、位置或状态标签。在相同的数据、目标和评估协议下,该模型在若干固定时域任务上达到了完美的验证准确率,而字节基线和非结构化记忆基线则未能做到,并且在顺序文本轨迹预测上有显著提升。这些结果表明,结构化的、与任务对齐的工作记忆可能比单纯增加潜在容量更为有用。更广泛地,我们认为语言接地是由持久性数据结构及其上的计算所中介的;该基准为测试将文本接口与学习到的结构化状态相耦合的架构提供了一个紧凑的环境。
cs.LG / 55 / 2609.02373

Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance

优化中的渗流动力学:方差级联与离散标度不变性
Ramachandran, Sai Niranjan, Sra, Suvrit
Abstract
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the stochastic gradient flow (SGF) as a percolation process, in which architectural symmetries force subnetworks to merge in discrete simultaneous blocks rather than one at a time. These structural transitions register as variance spikes in a macroscopic order parameter, echoing physical phase transitions. We further show this trapping mechanism and its associated scaling cascade extend to Adam and AdamW under an explicit heavy-tailed noise model.
Chinese Translation
我们研究了随机梯度下降(SGD)的动力学,已知它会引导深度神经网络趋向于对应更简单子网络的不变集。然而,这种引导如何随时间展开仍知之甚少。为此,我们将随机梯度流(SGF)建模为一个渗流过程,其中结构对称性迫使子网络以离散的同步分块方式合并,而非逐个合并。这些结构转变在宏观序参量中表现为方差尖峰,与物理相变类似。我们进一步证明,这种捕获机制及其相关的标度级联在显式的重尾噪声模型下同样适用于 Adam 和 AdamW 优化算法。
cs.LG / 56 / 2609.02404

Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts

稀疏混合专家模型中共享路由几何与动力学的证据
Labzin, Kirill, Kulibaba, Stepan, Dzhalilov, Artem, Gorokhov, Artem
Abstract
Sparse mixture-of-experts (MoE) models use an independently parameterized router at each sparse layer to select experts for every token. Prior work has shown that routing decisions across depth can often be predicted from earlier routing signals, suggesting that routing is not fully independent across layers. However, the structure behind this predictability remains unclear. In this work, we provide evidence that routing-relevant states across layers share a common geometric structure that is obscured by layer-specific coordinate systems. We isolate the control subspace of each router and align these spaces into a shared canonical representation using generalized orthogonal Procrustes analysis. After alignment, a single linear transition reaches $R^2=0.39$--$0.71$ and retains 79--90\% of the predictive power of separately fitted layer-specific dynamics, indicating that much of routing-state evolution follows a reusable process across depth. We then ask whether this shared dynamics is specific to routing or simply reflects the smooth evolution of hidden representations. A matched-rank comparison shows that residual representations are often easier to predict across layers, while router-control states preserve the model's expert choices much more faithfully. This separates generic cross-layer predictability from routing-specific information. Finally, we test whether the predicted canonical states remain meaningful when used in place of native routing states. The transported states preserve local routing behavior, while learned state evolution reduces $\Delta\mathrm{NLL}$ relative to simple persistence by 15.7\% on OLMoE and 6.2\% over a 10-router horizon on Phi.
Chinese Translation
稀疏混合专家模型在每一稀疏层使用独立参数化的路由器为每个词元选择专家。已有研究表明,跨深度的路由决策往往可以从较早的路由信号中预测得到,这说明路由在各层之间并非完全独立。然而,这种可预测性背后的结构仍不清楚。在本工作中,我们提供了证据表明,各层中与路由相关的状态共享一种被层特定坐标系所遮蔽的共同几何结构。我们分离出每个路由器的控制子空间,并使用广义正交Procrustes分析将这些空间对齐到一个共享的规范表示中。对齐后,单个线性变换即可达到 $R^2=0.39$--$0.71$,并保留了分别拟合的层特定动力学预测能力的79--90%,这表明路由状态的演化在很大程度上遵循一种可跨深度复用的过程。随后,我们探究这种共享动力学是路由所特有的,还是仅仅反映了隐藏表示的平滑演化。一项秩匹配比较显示,残差表示往往更容易跨层预测,而路由器控制状态则能更忠实地保留模型的专家选择,这将一般性的跨层可预测性与路由特有信息区分开来。最后,我们检验预测得到的规范状态在替代原始路由状态时是否仍然有意义。迁移后的状态能够保留局部路由行为,而学习到的状态演化相对于简单的持续性预测,在OLMoE上将 $\Delta\mathrm{NLL}$ 降低了15.7%,在Phi上于10个路由器的时间跨度内降低了6.2%。
cs.LG / 57 / 2609.02417

Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment

覆盖而非定向:多轮智能体信用分配中的一种结构性机制
Zhou, Chenyu, Jiang, Qiliang, Wu, Shuning, Zhou, Xu
Abstract
Multi-turn agentic RL increasingly treats credit assignment as a targeting problem: given a terminal verifiable reward, per-turn methods localize credit onto the turns that mattered. We identify the structural quantity that predicts when this is the right move, the verifier information density V_d = k/C (the fraction of an agent's C-step causal chain whose per-turn correctness the verifier exposes), and show that terminal-state verifiers sit deep in a low-V_d regime where targeting is the wrong axis. In controlled shared-rollout comparisons on tau^2-bench that separate reward density from credit geometry, a continuous dense reward spread uniformly beats the sparse binary outcome reward (net-harmful on 4/5 seeds), while concentrating the same advantage on progress turns or on random turns is equally harmful: targeting is second-order. The mechanism is coverage: terminal-state verification collapses the observable signal to a single final-write turn (k=1 in 98% of rollouts) while success requires a 5-8 step chain of prerequisite tool calls. A synthetic phase boundary places the crossover at V_d* ~ 0.8, whereas measured V_d is ~0.15 on tau^2-bench and ~0.4 on BFCL V3; uniform also wins on BFCL, where a matched-concentration shuffled control is negative on 8/8 seeds. The effect reproduces across model families on ToolACE-2-8B (Delta = -0.048 over 32 pre-registered seeds; an independent 20-seed replication is itself significant), and a pre-registered matched-budget breadth sweep traces a monotone dose-response whose deficit vanishes only at full chain coverage, with a reward-to-go arm reaching full-coverage parity. Uniform redistribution is the zero-information coverage default that per-turn schemes must beat; we contribute the matched-concentration shuffled control that any targeting claim should clear.
Chinese Translation
多轮智能体强化学习(RL)日益将信用分配视为一个定向(targeting)问题:给定终端可验证奖励,逐轮方法将信用定位到起关键作用的轮次上。我们识别出了预测何时定向才是正确策略的结构性量——验证器信息密度 V_d = k/C(即智能体 C 步因果链中验证器能暴露其逐轮正确性的比例),并表明终端状态验证器深处于低 V_d 机制中,此时定向是错误的轴。在 tau^2-bench 上进行的、将奖励密度与信用几何分离的受控共享轨迹比较中,均匀分布的连续密集奖励全面优于稀疏的二值结果奖励(在 5 个随机种子中的 4 个上净有害),而将相同优势集中于进展轮次或随机轮次同样有害:定向只是二阶因素。其机制在于覆盖:终端状态验证将可观测信号坍缩为单一的最终写入轮次(98% 的轨迹中 k=1),而成功需要 5-8 步前置工具调用链。一个合成的相变边界将交叉点置于 V_d* ≈ 0.8,而在 tau^2-bench 上实测 V_d 约为 0.15,在 BFCL V3 上约为 0.4;在 BFCL 上均匀分布同样获胜,其中匹配集中度的打乱对照在 8 个种子上全为负。该效应在 ToolACE-2-8B 上跨模型家族复现(32 个预注册种子上 Delta = -0.048;独立的 20 种子重复实验本身即显著),且一项预注册的匹配预算广度扫描勾勒出单调的剂量-响应关系,其亏损仅在完全链覆盖时消失,其中奖励回溯(reward-to-go)分支达到了与完全覆盖的对等水平。均匀重分配是逐轮方案必须超越的零信息覆盖默认基线;我们贡献了任何定向性主张都应通过的匹配集中度打乱对照。
cs.LG / 58 / 2609.02422

IFW-BLS: Dual-Robust Broad Learning System with Intuitionistic Fuzzy Wave Loss

IFW-BLS:基于直觉模糊小波损失的双鲁棒宽学习系统
Akhtar, Mushir, Tanveer, M.
Abstract
Broad Learning System is an efficient randomized learning model that expands network width through feature and enhancement nodes and estimates the output weights without deep backpropagation. Its standard least-squares training, however, is vulnerable in two different ways: (i) large residuals caused by noise, outliers, or corrupted labels can dominate the objective, and (ii) all samples are treated as equally reliable even when some lie in ambiguous or locally conflicting regions. This paper proposes IFW-BLS, an Intuitionistic Fuzzy Wave Broad Learning System that addresses these two sources of fragility within one optimization model. The first robustness mechanism is residual-level protection, obtained by replacing the squared loss with the bounded, smooth, and asymmetric wave loss. Boundedness prevents extreme residuals from receiving unbounded influence, while asymmetry allows positive and negative deviations to be penalized differently when the dominant error direction varies. The second mechanism is sample-level credibility control, obtained through intuitionistic fuzzy scores that combine global class-center consistency with local neighborhood conflict. The resulting model evaluates the wave loss on credibility-weighted residuals, so unreliable samples are down-weighted before the bounded loss further limits the effect of extreme errors. A Nesterov accelerated gradient based optimizer is used to solve the proposed objective, avoiding the explicit matrix inversion used in conventional BLS. Experiments on UCI benchmark datasets validate the superiority of the proposed IFW-BLS model over the baseline models; additional corruption experiments also show more stable performance than BLS under noise and outlier contamination.
Chinese Translation
宽学习系统(Broad Learning System, BLS)是一种高效的随机化学习模型,它通过特征节点和增强节点扩展网络宽度,并无需深度反向传播即可估计输出权重。然而,其标准的最小二乘训练在两个方面存在脆弱性:(i)由噪声、离群点或标签损坏引起的大残差可能主导目标函数;(ii)即使部分样本处于模糊或局部冲突区域,所有样本仍被视为同等可靠。本文提出IFW-BLS(Intuitionistic Fuzzy Wave Broad Learning System),即一种直觉模糊小波宽学习系统,在单一优化模型中同时应对这两种脆弱性来源。第一重鲁棒机制是残差层面的保护,通过用有界、平滑且非对称的小波损失替代平方损失来实现。有界性可防止极端残差获得无界影响,而非对称性使得当主导误差方向变化时,能够对正负偏差施加不同的惩罚。第二重机制是样本层面的可信度控制,通过直觉模糊得分来实现,该得分结合了全局类中心一致性与局部邻域冲突性。所得模型对可信度加权后的残差计算小波损失,因此在有界损失进一步限制极端误差影响之前,不可靠样本已被降权。本文采用基于Nesterov加速梯度(Nesterov accelerated gradient)的优化器求解所提出的目标函数,避免了传统BLS中显式的矩阵求逆。在UCI基准数据集上的实验验证了所提出的IFW-BLS模型相较于基线模型的优越性;额外的损坏实验也表明,在噪声和离群点污染下,该模型比BLS具有更稳定的性能。
cs.LG / 59 / 2609.02440

Towards One-for-All Robustness Across a Continuum of Threat Levels

面向跨连续威胁等级的“一体通用”鲁棒性
Hou, Zhichao, Liu, Xiaorui
Abstract
Adversarially robust models often overfit to a specific attack budget, necessitating multiple specialized models for diverse and dynamic adversarial environments, a strategy that becomes fundamentally intractable as the threat space grows. This raises an open challenge: can we achieve strong robustness across a continuum of threat levels within a single model? We propose the Threat Conditional Network (TCN), grounded in a representation factorization framework that decomposes representation learning into a threat-invariant shared backbone and a lightweight threat-conditional adaptor. TCN conditions a single model on the perturbation level via Fourier-based embeddings and channel-wise affine modulation, and is trained against a distribution over perturbation budgets, enabling flexible and seamless adaptation across an infinite continuum of threat levels during inference. Extensive experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet show that TCN matches or surpasses a full ensemble of budget-specialized models with a single set of parameters, generalizes to unseen perturbation budgets, and transfers robustly under mismatched threat conditions, with only 4.6\% parameter overhead. These contributions chart a promising path toward adaptive and generalizable robustness in dynamic and diverse threat environments.
Chinese Translation
对抗鲁棒模型通常会对特定的攻击预算过拟合,因而在多样且动态的对抗环境中需要多个专用模型,而随着威胁空间的增长,这一策略从根本上变得难以实现。这引出了一个开放性挑战:我们能否在单一模型内实现对连续威胁等级的强鲁棒性?我们提出了威胁条件网络(Threat Conditional Network, TCN),其基础是一个表征分解框架,该框架将表征学习分解为一个威胁不变的共享主干网络和一个轻量级的威胁条件适配器。TCN 通过基于傅里叶的嵌入和逐通道仿射调制,使单一模型以扰动等级为条件,并在扰动预算的分布上进行训练,从而在推理时能够跨无限连续的威胁等级实现灵活且无缝的适应。在 CIFAR-10、CIFAR-100 和 Tiny-ImageNet 上的大量实验表明,TCN 仅用一组参数即可匹敌或超越由预算专用模型组成的完整集成模型,能够泛化到未见过的扰动预算,并在威胁条件不匹配的情况下依然保持鲁棒的迁移能力,而参数开销仅为 4.6%。这些贡献为在动态多样的威胁环境中实现自适应且可泛化的鲁棒性指明了一条富有前景的道路。
cs.LG / 60 / 2609.02450

CACTUS: Mask-Guided Semantic Clean-Label Backdoors in Decentralized Federated Learning

CACTUS:去中心化联邦学习中基于掩码引导的语义干净标签后门攻击
Feng, Chao, Stiller, Burkhard
Abstract
Semantic triggers in federated learning (FL) can be less conspicuous than synthetic patches, but sample-dependent placement may weaken backdoor implantation across aggregation rounds. This challenge is compounded in decentralized FL (DFL), where topology-dependent peer aggregation repeatedly mixes local models. CACTUS converts label-consistent semantic pairs into target-directed representation shifts. Mask-guided, modality-specific operators isolate trigger effects, couple them across samples, and apply the shifts counterfactually to clean non-target embeddings before peer aggregation. Experiments cover speech, text, tabular, and image tasks under nine aggregation rules. With 30\% malicious nodes, CACTUS reaches a nine-rule mean attack success rate (ASR) of 51.2\% on Speech Commands and the highest nine-rule mean ASR among evaluated attacks on three of four modalities. Sensitivity analyses show that ASR varies with network topology and increases with the malicious-node ratio. These results indicate that CACTUS can propagate backdoors through repeated DFL aggregation.
Chinese Translation
联邦学习(FL)中的语义触发器可以比合成补丁更不显眼,但依赖样本的触发器放置方式可能在多轮聚合过程中削弱后门的植入效果。这一挑战在去中心化联邦学习(DFL)中更为严峻,因为依赖网络拓扑的对等聚合会反复混合本地模型。CACTUS 将标签一致的语义样本对转化为指向目标类的表示偏移。基于掩码引导的、针对特定模态的算子隔离触发器效应,将它们跨样本耦合,并在对等聚合之前以反事实方式将这些偏移应用于干净的非目标嵌入。实验涵盖语音、文本、表格和图像任务,并在九种聚合规则下进行。在30%恶意节点的条件下,CACTUS 在 Speech Commands 数据集上达到九种规则平均51.2%的攻击成功率(ASR),并在四种模态中的三种上取得所评估攻击中最高的九规则平均 ASR。敏感性分析表明,ASR 随网络拓扑变化,并随恶意节点比例的增加而上升。这些结果表明,CACTUS 能够通过反复的 DFL 聚合传播后门。
cs.LG / 61 / 2609.02451

Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression

可扩展的Kronecker-Fisher近似:面向十亿参数语言模型压缩的高效Hessian分析
Yusupov, Viacheslav, Cherniuk, Daria, Frolov, Evgeny
Abstract
In this paper, we propose a scalable Kronecker-based approximation that captures cross-layer interactions without storing the entire Fisher matrix, enabling practical Hessian analysis for billion-parameter networks where full computation is infeasible. Our approach reveals consistent vulnerability patterns: value projection layers exhibit the highest sensitivity and strongest cross-layer correlations across multiple model families, while other components exhibit architecture-specific behaviors. Through extensive experiments on quantization, sparsification, inter-layer corruption, and post-corruption fine-tuning, we demonstrate that our approximation strongly correlates with both performance degradation and recovery. Our framework provides a practical, theoretically grounded tool for identifying fragile components in large models, opening new avenues for guided compression and optimization strategies, such as mixed-precision allocation, layer-wise sparsity, and adaptive low-rank decomposition across layers and even individual weight groups.
Chinese Translation
本文提出了一种可扩展的基于Kronecker的近似方法,无需存储完整的Fisher矩阵即可捕捉跨层交互,使得对十亿参数规模的神经网络进行实用的Hessian分析成为可能,而在此规模下完整计算是不可行的。我们的方法揭示了一致的脆弱性模式:在多个模型家族中,值投影(value projection)层表现出最高的敏感性和最强的跨层相关性,而其他组件则表现出架构特定的行为。通过在量化、稀疏化、层间破坏以及破坏后微调上的大量实验,我们证明了我们的近似与性能退化和恢复均具有很强的相关性。我们的框架为识别大模型中的脆弱组件提供了一个实用且具有理论依据的工具,为有指导的压缩与优化策略开辟了新途径,例如混合精度分配、逐层稀疏性以及跨层甚至单个权重组的自适应低秩分解。
cs.LG / 62 / 2609.02468

DeepAffinity: Long-Term Aspect Preference Prediction in eCommerce using Small Language Models

DeepAffinity:基于小语言模型的电商长期产品属性偏好预测
Eshel, Yotam, Hadad, Guy, Feigenblat, Guy, Brovman, Yuri M., Gearhart, Matt, Shapira, Bracha
Abstract
We explore predicting eCommerce user preferences for product aspects such as brand, size, and color - a task we define as Aspect Affinity. Solving this task improves customer understanding and enables fine-grained personalization in recommendation, search, and marketing. We frame Aspect Affinity as a temporal prediction task: forecasting a users future aspect choices from their time-ordered interaction history, capturing long-term preferences that evolve beyond the current session. To this end, we propose DeepAffinity, which leverages Small Language Models (SLMs) with structured prompts and specialized prediction heads fine-tuned for this task. We show DeepAffinity outperforms standard generative fine-tuning methods, while general-purpose open-source LLMs perform poorly without task-specific tuning, highlighting their limits in modeling nuanced behavior. Finally, DeepAffinity enhances recommendation quality on a large-scale multinational eCommerce platform.
Chinese Translation
我们探索预测电商用户对产品属性(如品牌、尺寸和颜色)的偏好,我们将该任务定义为属性亲和度(Aspect Affinity)。解决这一任务能够提升对用户的理解,并实现推荐、搜索和营销中的细粒度个性化。我们将属性亲和度建模为一个时序预测任务:根据用户按时间排序的交互历史,预测其未来的属性选择,从而捕捉超越当前会话、随时间演化的长期偏好。为此,我们提出了DeepAffinity,该方法利用小语言模型(Small Language Models, SLMs),结合结构化提示词以及针对该任务微调的专用预测头。我们证明了DeepAffinity优于标准的生成式微调方法,而未经任务特定微调的通用开源大语言模型表现不佳,凸显了其在建模细微用户行为方面的局限性。最后,DeepAffinity在一个大规模跨国电商平台上提升了推荐质量。
cs.LG / 63 / 2609.02497

RINSE: Robust Target-Time Normality Estimation for Zero-Shot Graph Anomaly Detection

RINSE:面向零样本图异常检测的鲁棒目标时点正态性估计
Fuad, Taufikur Rahman, Jahin, Md Abrar, Hussain, Amir
Abstract
Zero-shot graph anomaly detection seeks to deploy a detector trained on source graphs to unseen, unlabeled targets, yet domain shift can make source-derived notions of normality unreliable. We introduce RINSE (Robust Iterative Normality Self-Estimation), a gradient-free target-time framework that keeps the source-trained detector fixed while sequentially estimating target normality, representation calibration, and evidence reliability from the target graph. Its core idea is to identify a reliable subset of low-residual target nodes, use them to construct a trimmed target-aware normality model, and combine complementary anomaly evidence through reliability-gated rank fusion and encoder ensembling. Across eight unseen target graphs, RINSE achieves the highest average AUPRC among the evaluated methods under two separate preprocessing protocols, while block ablations and sensitivity analyses support the combined design. These results support robust target-time estimation as a practical approach to generalist graph anomaly detection without target labels, gradients, or per-target tuning.
Chinese Translation
零样本图异常检测旨在将在源图上训练的检测器部署到未见过、无标注的目标图上,然而域偏移可能使基于源图得出的正态性概念变得不可靠。我们提出 RINSE(Robust Iterative Normality Self-Estimation,鲁棒迭代正态性自估计),这是一个无需梯度的目标时点框架,在保持源图训练的检测器固定的同时,从目标图中依次估计目标正态性、表示校准和证据可靠性。其核心思想是识别一组可靠的低残差目标节点,利用它们构建一个经过裁剪的目标感知正态性模型,并通过可靠性门控的排名融合和编码器集成来组合互补的异常证据。在八个未见过的目标图上,在两种独立的预处理协议下,RINSE 在所评估的方法中取得了最高的平均 AUPRC,同时模块消融实验和敏感性分析验证了该组合设计的有效性。这些结果表明,鲁棒的目标时点估计是一种实用的通用图异常检测方法,无需目标标签、梯度或针对每个目标的调优。
cs.LG / 64 / 2609.02507

Rethinking the Teacher-Student Framework for Test-Time Adaptation

重新思考测试时自适应中的师生框架
Sójka, Damian, Masana, Marc, Twardowski, Bartłomiej, Cygert, Sebastian
Abstract
Test-Time Adaptation (TTA) has recently emerged as a promising strategy that allows the adaptation of pre-trained models to changing data distributions at deployment time, without access to any labels. To mitigate error accumulation, researchers have widely adopted the teacher-student framework, though its long-term stability is often taken for granted. In this work, we challenge the common strategy of setting the teacher weights to an exponential moving average of the student by showing that error accumulation still occurs, although it is mostly apparent on longer sequences compared to those commonly utilized. We analyze the stability-plasticity trade-off within the teacher-student framework and propose to use an intransigent teacher that does not update its weights. Surprisingly, we show that this simple change allows TTA methods to significantly improve their performance on multiple datasets with longer scenarios and result in increased robustness to changes in hyperparameters. Finally, we show that those changes can be seamlessly and effectively applied to various architectures and experimental setups, including semantic segmentation. The code is available at https://github.com/dmn-sjk/intransigent_teacher.
Chinese Translation
测试时自适应(Test-Time Adaptation, TTA)是近年来兴起的一种有前景的策略,它使预训练模型能够在部署阶段适应不断变化的数据分布,而无需访问任何标签。为了缓解误差累积问题,研究者们广泛采用了师生(teacher-student)框架,但该框架的长期稳定性往往被认为是理所当然的。在本工作中,我们对将教师权重设置为学生权重的指数移动平均(EMA)这一常见策略提出了质疑,通过实验表明误差累积仍然存在,只是在比通常所用序列更长的序列上才较为明显。我们分析了师生框架内部的稳定性-可塑性权衡,并提出使用一种不更新自身权重的“固执教师”(intransigent teacher)。令人惊讶的是,我们证明这一简单的改动能使TTA方法在多个更长序列的数据集上显著提升性能,并增强了对超参数变化的鲁棒性。最后,我们表明这些改动可以无缝且有效地应用于各种架构和实验设置,包括语义分割任务。代码见 https://github.com/dmn-sjk/intransigent_teacher。
cs.LG / 65 / 2609.02519

Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion

面向不确定知识图谱补全的谱初始化与调度图平滑方法
Jahin, Md Abrar, Fuad, Taufikur Rahman, Pujara, Jay, Knoblock, Craig A.
Abstract
Uncertain knowledge graphs (UKGs) extend knowledge graphs by assigning each triple a continuous confidence score. Since most possible triples lack observed confidences, recent methods rely on semi-supervised learning to generate pseudo-labels. These methods initialize entity embeddings without using the confidence-weighted graph, discarding its global community and hub structure. We introduce QUEST, which adds no trainable parameters to the standard confidence-distribution learning pipeline. First, QUEST initializes entity embeddings using the smallest non-trivial eigenvectors of the confidence-weighted graph Laplacian, incorporating community and hub structure before training. Second, QUEST applies an unbiased mini-batch Dirichlet energy regularizer to enforce early-stage structural consistency. On two UKG datasets, QUEST improves confidence prediction and link prediction on six of eight metric-dataset pairs over prior methods and matches the previous best on the remaining two, while removing the instability spike observed on dense graphs. These results indicate that spectral structural priors combined with a graph Dirichlet energy regularizer improve accuracy, training stability, and checkpoint reliability in UKG completion.
Chinese Translation
不确定知识图谱(Uncertain Knowledge Graphs, UKGs)通过为每条三元组赋予一个连续的置信度分数来扩展知识图谱。由于大多数可能的三元组缺乏观测到的置信度,近期的方法依赖半监督学习来生成伪标签。这些方法在初始化实体嵌入时没有利用置信度加权图,从而丢弃了其全局社区结构和枢纽(hub)结构。我们提出了QUEST,该方法在标准的置信度分布学习流程中不增加任何可训练参数。首先,QUEST利用置信度加权图拉普拉斯矩阵的最小非平凡特征向量来初始化实体嵌入,在训练之前融入社区结构和枢纽结构。其次,QUEST应用无偏的小批量Dirichlet能量正则化项,以强制实现早期阶段的结构一致性。在两个UKG数据集上,QUEST在八组指标-数据集对中的六组上优于先前方法的置信度预测和链接预测性能,并在其余两组上与此前最佳方法持平,同时消除了在稠密图上观察到的训练不稳定尖峰。这些结果表明,谱结构先验与图Dirichlet能量正则化相结合,能够提升UKG补全的准确性、训练稳定性和检查点可靠性。
cs.LG / 66 / 2609.02538

A Comparative Study of Graph Representations for GNN-Based Power Grid Control in L2RPN

基于图神经网络(GNN)的L2RPN电网控制中图表示方法的比较研究
Degenkolb, Adrian, Huang, Qiong, Schäfer, Benjamin
Abstract
Graph construction is a critical but underexamined design choice in deep reinforcement learning for power grid control. We present a controlled experimental comparison of different graph representations, including physical topology, electrical-sensitivity, and hybrid variants for topology control in the Learning to Run a Power Network (L2RPN) environment. Our findings indicate that matching graph complexity to task granularity is more important than maximizing representational richness, and highlight the importance of controlled representation studies at scale.
Chinese Translation
在用于电网控制的深度强化学习中,图的构建是一个关键但未被充分研究的设计选择。我们对不同的图表示方法进行了受控实验比较,包括物理拓扑、电气敏感性以及混合变体,用于Learning to Run a Power Network(L2RPN)环境中的拓扑控制。我们的研究结果表明,使图复杂度与任务粒度相匹配比最大化表示的丰富性更为重要,并强调了大规模受控表示研究的重要性。
cs.LG / 67 / 2609.02540

TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly Diagnosis

TrajMind:链接角色专精LoRA实现快慢结合的群体轨迹异常诊断
Wu, Jiahao, Yang, Zhenqun, Zhang, Chen Jason, Li, Qing
Abstract
Diagnosing collective anomalies from urban trajectories is increasingly important for traffic governance, as it reveals what happened, who was involved, and where and when the event occurred. Existing detectors efficiently produce scores or labels, whereas vision--language pipelines provide richer semantics; neither couples verifiable diagnosis with low-latency monitoring. The central challenge is to recognize collective patterns and recover exact event details from the source trajectories without running the full diagnostic pipeline for every monitored window. We therefore separate always-on screening from on-demand diagnosis: screening raises alerts, while diagnosis releases only source-verified what--who--where--when records. We present TrajMind, a fast-and-slow framework that switches three role-specialized LoRA adapters over one frozen vision--language backbone. Its slow path, \textit{TrajMind$_{\text{slow}}$}, chains canvas-based typing, type-conditioned localization over serialized trajectories, and executable verification, yielding structured, evidence-backed diagnoses. Additionally, the fast path, \textit{TrajMind$_{\text{fast}}$}, screens each window in a single text-only pass, delivering efficient structured alerts. Extensive experiments show that, TrajMind$_{\mathrm{slow}}$ outperforms the strongest baselines by at least $15.3$ percentage points in anomaly typing and $13.8$ percentage points in localization. These gains persist under cross-city transfer, and TrajMind$_{\mathrm{fast}}$ reduces latency by $41.1\%$ and maintains binary balanced accuracy of at least $93.5\%$. Together, TrajMind delivers accurate, evidence-backed diagnoses across cities and efficient front-line monitoring.
Chinese Translation
从城市轨迹中诊断群体异常对交通治理日益重要,因为它揭示了发生了什么、谁参与其中、以及事件发生的地点和时间。现有的检测器能够高效地产生评分或标签,而视觉-语言管线则提供更丰富的语义;然而二者均无法将可验证的诊断与低延迟监控相结合。核心挑战在于:如何从源轨迹中识别群体模式并恢复确切的事件细节,而无需对每个监控窗口都运行完整的诊断管线。因此,我们将常开式筛查与按需诊断分离:筛查负责发出警报,而诊断仅释放经源数据验证的"什么-谁-何地-何时"记录。我们提出TrajMind,这是一个快慢结合的框架,在同一个冻结的视觉-语言骨干网络上切换三个角色专精的LoRA适配器。其慢速路径TrajMind$_{\text{slow}}$依次执行基于画布的类型判别、在序列化轨迹上进行类型条件化的定位、以及可执行的验证,从而产生结构化且有证据支持的诊断。此外,其快速路径TrajMind$_{\text{fast}}$通过单次纯文本推理对每个窗口进行筛查,提供高效的结构化警报。大量实验表明,TrajMind$_{\mathrm{slow}}$在异常类型判别上超越最强基线至少15.3个百分点,在定位上超越13.8个百分点。这些优势在跨城市迁移场景下依然保持,且TrajMind$_{\mathrm{fast}}$将延迟降低41.1%,同时保持至少93.5%的二分类平衡准确率。总之,TrajMind在跨城市场景下提供了准确、有证据支持的诊断以及高效的前线监控能力。
cs.LG / 68 / 2609.02548

Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs

向正确者学习:面向多领域大语言模型的答案验证式多教师蒸馏
He, Xixiang, Li, Xingming, Wu, Baiqi, Sun, Qiyao, Ji, Xuanyu, Cheng, Ao, Hu, Qingyong
Abstract
Modern large language models (LLMs) rely on reinforcement learning to build strong capabilities in individual domains, but integrating those capabilities into a single deployable model remains challenging. By routing each sample to the teacher whose domain matches it, existing approaches let a domain label decide which teacher provides supervision. However, domain expertise holds only on average: the matched teacher is not always correct on a given sample, while a teacher from another domain sometimes is. The reliable teacher therefore has to be identified per sample, not per domain. In this paper, we introduce Multi-Teacher Self-Distillation Policy Optimization (MT-SDPO), an on-policy distillation method that unifies several frozen teachers into one student model. MT-SDPO consists of three components: (1) self-anchors, where a rollout is supervised by a correct rollout from its own group; (2) answer-verified eligibility, where a teacher may supervise a sample only if its own answer passes a verifier; and (3) privileged distillation, which merges the anchor and all verified feedback into one context that an exponential moving average self-teacher reads and the student does not, thereby keeping one policy at deployment. Across five students from three model families, MT-SDPO lifts the weakest domain of Qwen3-8B by 14.79 points and narrows its domain gap by 74.7%, a better balance than serving one matched teacher per domain. Verified reliability, not domain membership, should decide who teaches. Code is available at https://github.com/hexixiang/MT-SDPO.
Chinese Translation
现代大语言模型(LLM)依赖强化学习在单一领域中构建强大的能力,但将这些能力整合到一个可部署的单一模型中仍然具有挑战性。现有方法通过将每个样本路由到与其领域匹配的教师,让领域标签决定由哪个教师提供监督。然而,领域专长只是在平均意义上成立:匹配的教师并非在给定样本上总是正确,而来自其他领域的教师有时反而是正确的。因此,可靠的教师必须逐样本识别,而非逐领域识别。本文提出多教师自蒸馏策略优化(Multi-Teacher Self-Distillation Policy Optimization, MT-SDPO),这是一种在-policy蒸馏方法,将多个冻结的教师统一到一个学生模型中。MT-SDPO由三个组件构成:(1)自锚点(self-anchors),即用来自同一组内的正确rollout对某次rollout进行监督;(2)答案验证资格(answer-verified eligibility),即教师只有在自身答案通过验证器后才可对某样本进行监督;(3)特权蒸馏(privileged distillation),即将锚点与所有经过验证的反馈合并为一个上下文,由指数移动平均自教师读取而学生不读取,从而在部署时保持单一策略。在来自三个模型系列的五个学生模型上,MT-SDPO将Qwen3-8B的最弱领域提升了14.79分,并将领域差距缩小了74.7%,取得了优于为每个领域单独部署一个匹配教师的平衡效果。决定谁来教学的应该是经过验证的可靠性,而非领域归属。代码发布于 https://github.com/hexixiang/MT-SDPO。
cs.LG / 69 / 2609.02549

ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction

ProbeMatchDTI:基于探针驱动的多尺度生化模式匹配的药物-靶点相互作用预测
Hao, Quan, Fan, Mengyue, Dong, Zifan, Li, Youru, Zhao, Jianduo, Xu, Lechuan, Zhang, Hao, Xia, Fei, Wang, Jigang, Qiu, Chong, Zhang, Liguo
Abstract
Drug-target interaction (DTI) prediction is an important task in AI-driven drug discovery. Although recent biochemical representation learning methods have improved DTI prediction, their passive feature aggregation tends to favor dominant molecular patterns while suppressing weak yet binding-relevant signals, such as functional groups and residue-context patterns, limiting the modeling of multi-scale biochemical correspondences. To address this issue, we propose ProbeMatchDTI, a pattern-probe-driven framework comprising IterProbe and BindingProbe. IterProbe explicitly retains contextual states across refinement depths and uses learnable probes to select them at each position before cross-entity matching, thereby preserving weak biochemical patterns and strengthening associations among functional groups, local motifs, and molecular scaffolds. BindingProbe then characterizes cross-entity drug-protein complementarity at local biochemical-unit and whole-pair levels, jointly modeling fine-grained interactions and multi-scale correspondences while preserving weaker binding-relevant associations. Extensive experiments demonstrate the superiority of ProbeMatchDTI, achieving 2.0% and 0.5% higher AUC-ROC on BindingDB and DrugBank, respectively. Feature-level pattern analyses further characterize its probe-driven behavior in cross-scale biochemical pattern matching. We further connect ProbeMatchDTI predictions with an evidence-guided downstream drug-discovery workflow, demonstrating their utility for candidate refinement and validation planning. Our code is available at https://github.com/developer-hq/ProbeMatchDTI
Chinese Translation
药物-靶点相互作用(DTI)预测是AI驱动药物发现中的一项重要任务。尽管近期的生化表示学习方法提升了DTI预测的性能,但其被动式特征聚合往往偏向主导分子模式,而抑制微弱却与结合相关的信号,例如官能团和残基上下文模式,从而限制了多尺度生化对应关系的建模。为解决这一问题,我们提出ProbeMatchDTI,一个由IterProbe和BindingProbe构成的模式探针驱动框架。IterProbe显式地保留不同精炼深度的上下文状态,并在跨实体匹配之前利用可学习的探针在每个位置对其进行选择,从而保留微弱的生化模式,并强化官能团、局部基序与分子骨架之间的关联。BindingProbe随后在局部生化单元和整个分子对两个层面上刻画药物-蛋白质之间的跨实体互补性,联合建模细粒度相互作用与多尺度对应关系,同时保留较弱的结合相关关联。大量实验证明了ProbeMatchDTI的优越性,其在BindingDB和DrugBank数据集上的AUC-ROC分别提升了2.0%和0.5%。特征层面的模式分析进一步刻画了其在跨尺度生化模式匹配中探针驱动的行为。我们进一步将ProbeMatchDTI的预测与证据引导的下游药物发现工作流程相衔接,展示了其在候选药物精炼和验证规划中的实用性。我们的代码可在 https://github.com/developer-hq/ProbeMatchDTI 获取。
cs.LG / 70 / 2609.02566

Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling

通过分布式模型-智能体耦合在英国气象局统一模式中实现在线强化学习
Nath, Pritthijit, Schemm, Sebastian, Haynes, Peter, Shuckburgh, Emily, Webb, Mark
Abstract
Machine-learnt corrections can complement numerical weather prediction only if they adapt to the evolving model state while preserving dynamical consistency and numerical stability. To test this within a global forecasting model, we couple the Met Office (UKMO) Unified Model (UM) with distributed RL agents through rank-local tensors. A DDPG actor shares weights across the 70 vertical model levels of each atmospheric column and applies bounded potential-temperature corrections to the model tendencies. Across ten nudged training forecasts, nudging calculations towards the UKMO operational analysis provides an immediate counterfactual target. The frozen policy is then evaluated in a non-nudged forecast for inference. The coupled workflow successfully completes training and remains numerically stable in the evaluated case. Relative to a matched native UM forecast at +6 h, the learnt policy reduces Z$_{500}$ MAE in four of six latitude bands, including reductions of 45.8% and 40.8% in the northern and southern tropics. MSLP error too decreases in three bands, with a maximum reduction of 27.3% at 0-30{\deg}N. This single-case experiment demonstrates significant promise and feasibility of distributed online learning followed by non-nudged inference, laying the groundwork for RL-based bias correction and parametrisations within operational systems.
Chinese Translation
机器学习订正只有在适应不断演变的模式状态、同时保持动力一致性和数值稳定性的前提下,才能有效补充数值天气预报。为了在全球预报模式中检验这一点,我们通过秩局部张量将英国气象局(UKMO)统一模式(UM)与分布式强化学习智能体相耦合。一个DDPG(Deep Deterministic Policy Gradient)行动者网络在每一个大气柱的70个垂直模式层之间共享权重,并对模式倾向施加有界的位温订正。在十个nudging(向目标逼近)训练预报中,向UKMO业务分析场逼近的计算提供了即时的反事实目标。随后,冻结的策略被用于一次非nudging预报中进行推理评估。在所评估的个例中,耦合工作流成功完成了训练并保持了数值稳定性。与同化配置的原生UM预报在+6小时相比,所学到的策略在六个纬度带中的四个降低了500百帕位势高度(Z$_{500}$)的平均绝对误差(MAE),其中南北热带分别降低了45.8%和40.8%。海平面气压(MSLP)误差也在三个纬度带中有所下降,最大降幅为27.3%,出现在0–30°N纬度带。这一单一个例试验展示了分布式在线学习加非nudging推理的巨大前景与可行性,为在业务系统中开展基于强化学习的偏差订正和参数化方案奠定了基础。
cs.LG / 71 / 2609.02622

Source Distribution Estimation by Posterior Averaging

基于后验平均的源分布估计
Hoang, Trung-Dung, Koch, Lisa M.
Abstract
Simulation-based science often requires a distribution over simulator parameters whose push-forward reproduces a set of real observations: this is the source distribution estimation (SDE) problem. Existing methods fit the source against a likelihood surrogate trained once from a fixed proposal prior. Their objective is therefore stated only in terms of the surrogate instead of the true simulator, which may fail for inaccurate areas in parameter space where the surrogate was never trained. We instead solve SDE by expectation maximization: an E-step trains an amortized posterior on fresh simulations from the current source estimate, and an M-step refits the source to the average of that posterior over the observed data. We give two parameterizations, (1) separate source and posterior flows and (2) a single shared conditional flow. We evaluate our method on three benchmark tasks under both broad and misspecified initial priors. Both improve on existing fixed surrogate approaches and on iterated variants of each, most clearly on Lotka--Volterra, where no baseline falls below 0.96 data-space C2ST while our methods reach 0.64-0.68 in three of four initial-prior settings.
Chinese Translation
基于仿真的科学研究通常需要一个模拟器参数的分布,其前推映射能够重现一组真实观测数据,这就是源分布估计(SDE)问题。现有方法基于一次性从固定提议先验训练得到的似然代理模型来拟合源分布。因此,其目标函数仅以代理模型而非真实模拟器来表述,当参数空间中存在代理模型从未训练过的不准确区域时,方法可能失效。我们转而通过期望最大化(EM)算法求解SDE:E步在来自当前源分布估计的全新仿真数据上训练一个摊销后验,M步将该后验在观测数据上的平均值重新拟合为源分布。我们给出了两种参数化方式:(1)分离的源流模型和后验流模型,以及(2)单一共享的条件流模型。我们在宽泛先验和错误设定先验两种情况下,于三个基准任务上评估了我们的方法。两种参数化方法均优于现有的固定代理模型方法及其各自的迭代变体,在Lotka--Volterra任务上优势最为明显:所有基线方法在数据空间C2ST指标上均不低于0.96,而我们的方法在四个初始先验设置中的三个达到了0.64-0.68。
cs.LG / 72 / 2609.02638

Oracle, will I ever learn? A study of prediction convergence and complementarity across link prediction models

Oracle,我终究能学会吗?关于链接预测模型间预测收敛性与互补性的研究
Méroué, Guillaume, Gandon, Fabien, Monnin, Pierre
Abstract
Knowledge graphs have become an important source of structured knowledge for Web applications, including search, question answering, and recommender systems. In these applications, link prediction can serve either as a prediction task itself or as a means to enrich incomplete knowledge graphs for downstream tasks. Interestingly, different link prediction models, or even different training runs of the same model, can produce substantially different predictions for the same query. This suggests a variability in the capture of the underlying knowledge by models, thus raising a fundamental question: to what extent do different models capture complementary knowledge, and how much of this knowledge could be recovered by combining them? We propose to measure model complementarity through the performance of an oracle that, for each query, selects the best prediction among a considered set of models, hence providing an upper bound on the performance achievable through model combination. Across several architectures and benchmarks, we find a substantial gap between individual models and their oracle, revealing that different models capture complementary knowledge. Yet, this complementarity rapidly saturates as more models are added, leaving a persistent subset of queries unsolved even by a large number of models. These findings reveal both the potential of model complementarity and a fundamental limit to what current link prediction models can collectively recover; thereby highlighting the need for further research to build robust Web applications.
Chinese Translation
知识图谱已成为Web应用(包括搜索、问答和推荐系统)中结构化知识的重要来源。在这些应用中,链接预测既可以作为预测任务本身,也可以作为丰富不完整知识图谱以支持下游任务的手段。有趣的是,不同的链接预测模型,甚至是同一模型的不同训练过程,对同一查询可能产生截然不同的预测结果。这表明模型在捕捉底层知识方面存在变异性,由此引出一个根本性问题:不同模型在多大程度上捕捉了互补的知识,以及通过组合这些模型能够恢复多少知识?我们提出通过Oracle(神谕模型)的性能来衡量模型互补性,该Oracle对每个查询从一组候选模型中选择最佳预测,从而为模型组合所能达到的性能提供了上限。在多种架构和基准数据集上的实验表明,单个模型与其Oracle之间存在显著差距,说明不同模型捕捉了互补的知识。然而,这种互补性随着模型数量的增加而迅速饱和,即使使用大量模型,仍有一小部分查询无法被解决。这些发现既揭示了模型互补性的潜力,也揭示了当前链接预测模型在集体知识恢复能力上的根本局限;进而强调需要进一步研究以构建健壮的Web应用。
cs.LG / 73 / 2609.02646

Differentiable Electricity-Market Clearing for Gradient-Based Planning

面向基于梯度规划的可微分电力市场出清
Mungo, Luca, Scholl, Maarten P., Quera-Bofarull, Arnau
Abstract
Planning a large data center is difficult because a facility big enough to matter changes the electricity prices it will pay. Those prices are set by market clearing, a constrained optimization problem solved anew in every operating condition. However, simulating the market tells a planner how a candidate plan performs but not how to improve it. Here we treat market clearing as a differentiable optimization layer: each forward pass solves the market, and reverse-mode automatic differentiation propagates the planning cost back through the cleared prices to the plan. After validating these gradients against finite differences, we apply them to a concrete problem: allocating 50 MW of data-center load across six candidate buses in two synthetic networks, under a fixed cost per active site, evaluated over 36 operating states. Judged against exhaustive enumeration of all site combinations, gradient optimization recovers the continuous allocations almost exactly, with worst-case objective gaps of 2.3\% and 8.5\% of the cost difference between the best and worst single site. Its one systematic error is instructive: near the costs at which a site should close, the smooth relaxation of the discrete site count shrinks the site rather than closing it, so discrete transitions arrive late. Differentiable market clearing thus turns market-aware planning into a problem gradients can search.
Chinese Translation
大型数据中心选址规划十分困难,因为规模大到足以产生影响的设施会改变其自身将面临的电价。电价由市场出清决定,这是一个需要在每种运行条件下重新求解的约束优化问题。然而,对市场的模拟只能告诉规划者某个候选方案的表现如何,却无法告诉其如何改进。本文将市场出清视为一个可微分优化层:每次前向传播求解市场,反向模式自动微分则将规划成本通过出清电价反向传播至规划方案。在将这些梯度与有限差分法进行验证后,我们将其应用于一个具体问题:在两个合成网络中,将50兆瓦的数据中心负荷分配至六个候选母线,在活跃站点固定成本条件下,对36种运行状态进行评估。与所有站点组合的穷举枚举相比,梯度优化几乎精确地恢复了连续分配,其最坏情况目标差距分别为2.3%和8.5%(以最优与最差单站点成本之差为基准)。其唯一系统性错误具有启发意义:在站点应当关闭的成本临界点附近,离散站点数的平滑松弛使得站点规模缩小而非关闭,导致离散转变延迟出现。因此,可微分市场出清将市场感知规划转化为梯度可以搜索的问题。
cs.LG / 74 / 2609.02652

Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights

展开Leech格:面向2比特LLM权重的融合多壳层解码与显存布局
Malandrino, Pier-Jean
Abstract
Leech-lattice vector quantization holds the strongest reported 2-bit quality under its own evaluation protocol. Its kernel decodes one shell; we found no implementation of the multi-shell decoder the rate requires. This paper supplies one and measures its serving cost for decode-phase GEMV at batch 1. First, a serving path for the full 301-class codebook: an offline expansion into GPU layouts and a fused dequantize-plus-matvec kernel reading them without warp divergence, verified against f64. Second, the in-VRAM rate is a design axis distinct from the on-disk rate. Four bit-exact layouts timed in one process show binary bit planes beating one-hot masks on size and speed at constant bandwidth (4.80 bits per weight, 2.15x FP16). Below 4.3 bits a second, irregular stream enters; at 3.6 the decode stops being shifts and masks. Third, deployed four-bit (AWQ) and two-bit (QTIP) GEMV kernels run in the same process. The trellis kernel reads 2.40x fewer bytes than our served layout and runs 2.27x faster at near-equal fractions of their byte bounds: the time gap tracks the traffic gap, the price of unfolding a codebook too large for a lookup table. Fourth, the validity envelope: the trellis kernel outruns our no-weights control, so our launch geometry sets that floor, and on a second memory hierarchy every lattice arm falls below FP16. With the output head held identical across arms, the kernel-and-format path gains 1.11x, 1.29x and 1.41x end to end at 4B, 8B and 14B; with an int8 output head the served 4B reaches 87.0 tok/s in 2.60 GB. The quality cost, 1.38x perplexity and 14.7 MMLU points at 4B, shrinks across the three sizes measured.
Chinese Translation
Leech格矢量量化在其自有的评估协议下保持着已报道的最强2比特质量。其解码内核仅解码单一壳层,而我们未能找到其码率所要求的多壳层解码器的任何实现。本文提供了一个这样的实现,并测量了其在批大小为1的解码阶段GEMV中的服务成本。第一,为完整的301类码本构建一条服务路径:通过离线展开生成GPU布局,并使用一个无线程束发散(warp divergence)的融合反量化-矩阵向量乘内核读取这些布局,结果与f64进行了验证比对。第二,显存内的码率是与磁盘上码率不同的一个设计维度。在同一进程中计时的四种比特精确布局表明,在固定带宽下(每权重4.80比特,为FP16的2.15倍),二进制比特平面在大小和速度上均优于独热掩码。当低于4.3比特时,会出现第二条不规则码流;在3.6比特时,解码不再是简单的移位与掩码操作。第三,在相同进程中运行已部署的四比特(AWQ)与两比特(QTIP)GEMV内核。格状(trellis)内核读取的字节数比我们所服务的布局少2.40倍,运行速度快2.27倍,且两者均接近其各自字节上限的相同比例:时间差距与流量差距相吻合,这正是展开一个过大而无法用查找表容纳的码本所付出的代价。第四,有效性边界:格状内核的运行速度快于我们的无权重对照组,因此我们的启动配置决定了该下限;而在另一种内存层级上,所有格量化方案的速度均低于FP16。在所有方案使用相同输出头的情况下,内核与格式路径在4B、8B和14B规模上分别获得1.11倍、1.29倍和1.41倍的端到端加速;若采用int8输出头,所服务的4B模型在2.60 GB显存内达到87.0 tok/s。其质量代价——在4B规模上困惑度为1.38倍、MMLU下降14.7分——在所测的三个规模上逐渐缩小。
cs.LG / 75 / 2609.02684

H3DNAS: Hardware-Aware ONNX-Native 3D Point Cloud Model Compression

H3DNAS:面向硬件感知的ONNX原生三维点云模型压缩
Mulye, Anchit, Baghel, Rhythm, Ingle, Sujay Kumar, Jain, Hardik
Abstract
Deploying 3D point cloud models on edge hardware such as the NVIDIA Jetson Orin Nano is severely constrained by compute and memory budgets. Existing compression methods require access to the model's original source code, rendering them inapplicable to the Open Neural Network Exchange (ONNX) binaries commonly distributed by vendors and model repositories. We present \textbf{H3DNAS}, a hardware-aware model compression framework that operates directly on ONNX computational graphs without requiring original source code, architecture class definition, or gradient access during search. H3DNAS makes three contributions: (1) a \textbf{Channel Dependency Graph (CDG)} that classifies ONNX operators into four constraint classes and formally establishes that the free parameter fraction $\rho_f$ is topological invariant, a provable compression ceiling computable in $\mathcal{O}(|V|+|E|)$; (2) a \textbf{Two-Stage Hierarchical Search} that prunes candidate architectures by $L_1$-importance channel selection, ranks them by output fidelity as a zero-shot label-free proxy, and applies GhostConv structural mutation to Pareto-optimal candidates; and (3) the \textbf{first source-code-free compression pipeline for 3D point cloud models}, operating entirely via ONNX graph surgery with no original architecture definition required. On ModelNet40, H3DNAS reduces the number of parameters in PointNet, PointNet++, and PointMLP by $65.5\%$, $43.2\%$, and $49.1\%$, respectively, while achieving $1.99\times$, $1.29\times$, and $1.67\times$ inference speedups with negligible loss in accuracy. The source code is publicly available\footnote{https://github.com/ClarityLab-Org/h3dnas}.
Chinese Translation
在NVIDIA Jetson Orin Nano等边缘硬件上部署三维点云模型受到计算与内存预算的严重限制。现有压缩方法需要访问模型的原始源代码,因而无法应用于厂商和模型仓库通常分发的开放神经网络交换(ONNX)二进制文件。我们提出了H3DNAS,一个硬件感知的模型压缩框架,可直接在ONNX计算图上运行,无需原始源代码、架构类定义或搜索过程中的梯度访问。H3DNAS做出了三项贡献:(1)提出通道依赖图(Channel Dependency Graph, CDG),将ONNX算子划分为四类约束类别,并从形式上证明了自由参数比例$\rho_f$是拓扑不变量,这是一个可在$\mathcal{O}(|V|+|E|)$时间内计算的可证明压缩上限;(2)提出两阶段分层搜索,先通过$L_1$重要性通道选择剪枝候选架构,再以输出保真度作为零样本无标签代理指标对其进行排序,并对Pareto最优候选应用GhostConv结构变异;(3)首个面向三维点云模型的无源代码压缩流水线,完全通过ONNX图 surgery(图手术)实现,无需原始架构定义。在ModelNet40上,H3DNAS将PointNet、PointNet++和PointMLP的参数量分别减少65.5%、43.2%和49.1%,同时实现1.99倍、1.29倍和1.67倍的推理加速,且精度损失可忽略不计。源代码已公开:https://github.com/ClarityLab-Org/h3dnas。
cs.LG / 76 / 2609.02734

LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates

LoRA-TSD:基于Muon风格更新的LoRA切空间谱下降法
Andriianov, Dmitrii, Veprikov, Andrey, Beznosikov, Aleksandr
Abstract
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that treats every LoRA step as a tangent vector of the fixed-rank matrix manifold and takes the spectral-norm steepest-descent step of Muon inside that tangent space, mapping the result back to the factors through a retraction native to the LoRA parametrization. The step avoids expensive operations on full weight matrices, and its retraction is up to $2.8\times$ cheaper than the truncated-SVD retraction used by prior manifold methods. We prove that the Frobenius-norm version of our surrogate recovers LoRA-Pro, and we identify the tangent-projected gradient, the Riemannian gradient of the manifold, as the stationarity measure natural to LoRA training and computable from the factor gradients alone. Under this measure we give the first global convergence guarantees for both LoRA-Pro and LoRA-TSD, with rates that drive the factor-gradient norms to zero. Across six commonsense and natural-language-inference benchmarks with Llama-3.2-1B, Llama-3.1-8B and Qwen3-32B, LoRA-TSD outperforms every competing LoRA optimizer and stays robust to the adapter rank. Code is available at https://github.com/brain-lab-research/LoRA-TSD.
Chinese Translation
低秩适应(LoRA)是微调大模型的标准方法,然而当其两个因子被独立训练时,参数更新忽略了其所诱导的低秩权重变化的几何结构。我们提出LoRA-TSD,一种将每一步LoRA更新视为固定秩矩阵流形切向量的优化器,并在该切空间内执行Muon的谱范数最速下降步骤,随后通过与LoRA参数化天然契合的收缩映射将结果映射回因子矩阵。该更新步骤避免了对完整权重矩阵的昂贵运算,其收缩映射的计算代价比以往流形方法所使用的截断SVD收缩低至多$2.8\times$。我们证明该代理目标的Frobenius范数版本可退化为LoRA-Pro,并指出切投影梯度——即该流形的黎曼梯度——是LoRA训练自然的平稳性度量,且可仅由因子梯度计算得到。基于该度量,我们首次为LoRA-Pro和LoRA-TSD给出了全局收敛性保证,其收敛速率可使因子梯度范数趋于零。在Llama-3.2-1B、Llama-3.1-8B和Qwen3-32B上开展的六项常识推理与自然语言推理基准实验中,LoRA-TSD优于所有竞争的LoRA优化器,并对适配器秩的变化保持稳健。代码可在 https://github.com/brain-lab-research/LoRA-TSD 获取。
cs.LG / 77 / 2609.02766

Do Tabular Foundation Models Know Physics? Contamination, Units, and the Deterministic Limit

表格基础模型懂物理吗?数据污染、物理单位与确定性极限
Tenachi, Wassim, Hezaveh, Yashar, Levasseur, Laurence Perreault, Bacon, Pierre-Luc
Abstract
Tabular foundation models (TFMs) learn to fill in tables the way language models fill in text, and tables are arguably the format in which most physical measurement arrives. Did they learn any physics in the process? They are Bayesian by construction, so the question is what their prior contains. We probe it directly, evaluating four of them (TabPFN-3, TabICLv2, TabDPT and Real-TabPFN-2.5) against six baselines on datasets sampled from 316 physical equations, in and out of domain. TFMs dominate, out of the box and after tuning. But we show that their prior can represent neither a noiseless mechanism nor physical units, which is why they interpolate physics without yet being able to act as physical models.
Chinese Translation
表格基础模型(Tabular Foundation Models,TFMs)像语言模型补全文本一样学习补全表格,而表格可以说是大多数物理测量数据所采用的格式。那么它们在这一过程中学到了物理吗?由于这些模型在构造上是贝叶斯的,因此问题在于它们的先验中包含了什么。我们直接对此进行探测,在从316个物理方程中采样的数据集(包括域内和域外)上,评估了四个模型(TabPFN-3、TabICLv2、TabDPT和Real-TabPFN-2.5)与六个基线的表现。结果表明,无论开箱即用还是经过调优,TFMs均占优势。但我们发现,它们的先验既无法表示无噪声的机制,也无法表示物理单位,这正是它们能够对物理进行插值、却尚不能作为物理模型发挥作用的原因。
cs.LG / 78 / 2609.02817

Cliff: Learning Process Rewards from the First Mistake

Cliff:从首个错误中学习过程奖励
Han, Peixuan, Wang, Runhui, Ramaneti, Ketan, Hao, Jie, Friedland, Gerald, Kong, Chris
Abstract
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards leads to limited guidance on intermediate reasoning processes. Existing approaches such as process reward modeling and on-policy distillation introduce additional constraints, such as reliance on a specialized reward model or assuming identical reasoning patterns between teacher and student. Nevertheless, we observe that once a reasoning process first goes wrong, evaluating the subsequent reasoning provides limited additional information, as it is already conditioned on an invalid prefix. Therefore, we propose Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout. As a result, the rollout is naturally decomposed into two parts: a correct prefix and an incorrect suffix. Cliff then converts this signal into token-level advantages, assigning positive advantages for the correct prefix and negative feedback afterward. Experiments across 12 different scenarios demonstrate that Cliff consistently improves reasoning performance, outperforming on-policy distillation by 15% and standard GRPO by 7%, even with teachers of modest capability. Furthermore, we analyse the role of ``ground truth'' in Cliff and investigate its training dynamics. These results establish Cliff as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.
Chinese Translation
基于可验证奖励的强化学习(RLVR)已成为大语言模型(LLM)后训练的一种强大范式,但其对粗粒度结果奖励的依赖导致对中间推理过程的指导有限。现有方法(如过程奖励建模和在线策略蒸馏)引入了额外的约束,例如依赖专门的奖励模型,或假设教师与学生具有相同的推理模式。然而,我们观察到,一旦推理过程首次出错,对后续推理的评估所提供的额外信息就十分有限,因为后续推理已经以无效的前缀为条件。因此,我们提出Cliff,这是一种奖励塑形策略,利用现成的大语言模型作为教师来识别每个采样序列(rollout)中的首个错误。由此,该采样序列被自然地分解为两部分:正确的前缀和错误的后缀。随后,Cliff将这一信号转化为词元级别的优势(advantage),为正确前缀分配正优势,并在其后给予负反馈。在12个不同场景上的实验表明,Cliff能够持续提升推理性能,即使教师模型能力一般,也较在线策略蒸馏提升15%,较标准GRPO提升7%。此外,我们分析了“真实答案”(ground truth)在Cliff中的作用,并研究了其训练动态。这些结果确立了Cliff作为一种简单、通用且有效的方法,能够以更丰富、更细粒度的监督改进RLVR。
cs.LG / 79 / 2609.02846

UE5M3 FP4 Block Scaling for Stable Language Model Pretraining

UE5M3 FP4 块缩放用于稳定的语言模型预训练
Hu, Robert, Luschi, Carlo, Balanca, Paul
Abstract
Stable 4-bit floating-point (FP4) pretraining is difficult because the E2M1 payload represents only a narrow range of magnitudes. NVIDIA's Transformer Engine \nv{} recipe addresses this with current-tensor scaling, a randomized Hadamard transform (RHT), and bfloat16 (BF16) final layers, adding work outside the FP4 matrix multiplications. We instead pair E2M1 payloads with unsigned E5M3 (\ue{}) block scales. Their wider range permits periodic tensor scaling, while our recipe applies selective stochastic rounding to backward gradients, omits RHT, and uses FP4 in all eligible internal linears. We pretrain a Nemotron-H 8B model for nearly 190 billion tokens. Compared with Transformer Engine \nv{}, the proposed block-16 recipe finishes with lower final-window training loss and, under their respective quantized-inference policies, lower validation loss measured as held-out negative log-likelihood. Its quantized-inference downstream point estimates are also higher on all three reported aggregates. A native \nv{} execution ablation that jointly removes RHT and the BF16 final-block exemption increases measured model-body token throughput by 21.2\%. These results demonstrate end-to-end software-emulated \uefp{} pretraining with a simpler recipe and motivate native support for \ue{} block scaling.
Chinese Translation
稳定的4位浮点数(FP4)预训练十分困难,因为E2M1有效载荷仅能表示较窄的数值范围。NVIDIA的Transformer Engine v{} 配方通过当前张量缩放(current-tensor scaling)、随机Hadamard变换(RHT)以及bfloat16(BF16)最终层来解决这一问题,但这在FP4矩阵乘法之外增加了额外工作。我们转而将E2M1有效载荷与无符号的E5M3(\ue{})块缩放因子相结合。其更宽的表示范围允许周期性的张量缩放,同时我们的配方对反向传播梯度采用选择性随机舍入(stochastic rounding),省略RHT,并在所有符合条件的内部线性层中使用FP4。我们对Nemotron-H 8B模型进行了近1900亿词元(token)的预训练。与Transformer Engine \nv{} 相比,所提出的块大小为16的配方在最终窗口的训练损失上更低,并且在各自量化推理策略下,以保留集负对数似然衡量的验证损失也更低。其量化推理的下游点估计在所有三项报告的聚合指标上也均更高。一项原生的 \nv{} 执行消融实验(同时移除RHT和BF16最终块豁免)使测得的模型主体词元吞吐量提高了21.2%。这些结果证明了采用更简单配方的端到端软件模拟 \uefp{} 预训练的可行性,并为 \ue{} 块缩放的原生支持提供了依据。
cs.LG / 80 / 2609.02849

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

面向编程竞赛金牌水平性能的语言模型后训练
Ficek, Aleksander, Narenthiran, Sean, Samadi, Mehrzad, Majumdar, Somshubra, Ginsburg, Boris
Abstract
Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pipeline combining large-scale problem curation, synthetic reasoning traces, supervised fine-tuning (SFT), and reinforcement learning (RL). Using 22,000 curated problems, we train Nemotron-3-Nano-CC (30B-A3B) with SFT and RL and Nemotron-3-Ultra-CC (550B-A55B) with SFT alone. We further introduce GenCorrect, a feedback-driven test-time compute strategy that iteratively generates, evaluates, and refines diverse solutions. On IOI 2025, Nano-CC improves from 130 points to 291 after post-training and to 468 with GenCorrect, exceeding the gold threshold of 438.3 while Ultra-CC reaches 502. Guided by these results, we develop a competition-specific Ultra-CC system and evaluate it prospectively during IOI 2026. Under the same time, internet-access, and submission constraints as human contestants, it scores 535.4 out of 600, exceeding both the gold threshold of 361.12 and the top human score of 498.27. To our knowledge, this is the first AI system to outscore the highest-scoring human contestant on an IOI problem set.
Chinese Translation
竞赛编程已成为大语言模型推理能力的一项关键测试,其中国际信息学奥林匹克竞赛(IOI)和国际大学生程序设计竞赛(ICPC)等国际赛事代表了其最具挑战性的场景。我们提出了一种端到端的专项化流水线,结合了大规模题目筛选、合成推理轨迹、监督微调(SFT)和强化学习(RL)。我们使用22,000道精选题目,通过SFT和RL训练了Nemotron-3-Nano-CC(30B-A3B),并仅通过SFT训练了Nemotron-3-Ultra-CC(550B-A55B)。我们进一步提出了GenCorrect,一种基于反馈的测试时计算策略,可迭代地生成、评估并改进多样化的解题方案。在IOI 2025上,Nano-CC在后训练后得分从130分提升至291分,结合GenCorrect后达到468分,超过了438.3分的金牌线,而Ultra-CC达到502分。基于这些结果,我们开发了一个面向竞赛的Ultra-CC系统,并在IOI 2026期间进行了前瞻性评估。在与人类选手相同的时间、互联网访问和提交次数限制下,该系统在600分中获得535.4分,既超过了361.12分的金牌线,也超过了498.27分的人类最高分。据我们所知,这是首个在IOI题目集上得分超过人类最高分选手的AI系统。
cs.LG / 81 / 2609.02852

The Implications of Linguistic Illegibility for LLM Security

语言不可读性对大语言模型(LLM)安全的影响
Mickens, James
Abstract
LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks. We argue that the specter of linguistic illegibility is unavoidable for LLMs whose internal computations are not directly expressed via language, but rather math over activation spaces (with lossy translations between activation spaces and natural language happening at the bookends). If linguistic illegibility is always possible, then security mechanisms that rely on a model's linguistic self-reporting (e.g., chain-of-thought monitoring, constitutional self-critique, activation probing for linguistically-defined feature vectors) can never be completely sound; the model sandbox will always need isolation techniques whose guarantees do not depend on reading a model's linguistic state at all. We argue that observing a model's outputs using taint tracking is a promising approach for an effective sandbox: regardless of how a model linguistically self-reports, a taint tracking policy can define, a priori, various pieces of system state that should never be influenced by model-produced data. We also discuss several additional sandboxing mechanisms (e.g., robust virtualization, third-party auditing of sandboxing configurations) which collectively provide a critical floor beneath linguistic monitoring, and would have mitigated recent sandbox exploits by frontier models.
Chinese Translation
大语言模型(LLM)经过训练以生成自然语言。然而,多方面的证据表明,LLM外化的语言输出以及通过机制分析提取的语言特征,可能并不是理解模型内部计算的可靠透镜。我们提出“语言不可读性(linguistic illegibility)”这一术语,泛指LLM的外化语言产物或通过机制探针获取的语言产物无法真实反映模型实际思考过程的情形。我们认为,对于内部计算并非直接以语言表达、而是以激活空间上的数学运算进行的LLM(激活空间与自然语言之间仅在两端进行有损转换)而言,语言不可读性的阴影不可避免。如果语言不可读性始终可能存在,那么依赖模型语言自我报告的安全机制(例如思维链监控、宪法式自我批评、针对以语言定义的特征向量的激活探针)就永远无法完全可靠;模型沙箱始终需要那些其保障完全不依赖于读取模型语言状态的隔离技术。我们认为,使用污点跟踪(taint tracking)来观察模型输出是一种有望构建有效沙箱的方法:无论模型在语言上如何自我报告,污点跟踪策略都可以先验地定义哪些系统状态绝不应受模型生成数据的影响。我们还讨论了若干其他沙箱机制(例如鲁棒的虚拟化、对沙箱配置的第三方审计),它们共同为语言监控提供了关键的安全底线,并能缓解近期前沿模型的沙箱逃逸攻击。
cs.LG / 82 / 2609.02881

Graph Machine: Towards Better Pretraining via Edges

图机器:通过边实现更好的预训练
Hou, Lintai
Abstract
We introduce the Graph Machine (GM), an architecture that maintains an $O(n)$-sized state and accesses it through sparse, dynamic routing. Unlike methods with fixed-size states or sparse but static routing, GM preserves $O(n)$ complexity in its sparse layers without restricting the potentially accessible state size to $O(1)$. Instead, GM uses edges - pointer-like objects updated differentiably by a referral mechanism resembling pointer chasing. We replace 75% of the dense Transformer layers in Qwen3-0.6B with GM sparse layers and pretrain from scratch on 15.7B tokens. With only 2 of 4,096 tokens retrieved per KV head in each sparse layer, loss degrades only slightly; with 4, the best model marginally improves loss.
Chinese Translation
我们提出了图机器(Graph Machine,GM),这是一种维护 $O(n)$ 规模状态并通过稀疏动态路由进行访问的架构。与采用固定大小状态或稀疏但静态路由的方法不同,GM 在其稀疏层中保持了 $O(n)$ 的复杂度,同时不会将可访问的状态规模限制在 $O(1)$。相反,GM 使用边(edges)——一种类似指针的对象,通过一种类似指针追踪(pointer chasing)的推荐机制以可微分的方式进行更新。我们将 Qwen3-0.6B 中 75% 的稠密 Transformer 层替换为 GM 稀疏层,并在 15.7B 词元上从头开始预训练。当每个稀疏层中每个 KV 头仅检索 4,096 个词元中的 2 个时,损失仅有轻微上升;当检索 4 个时,最优模型的损失甚至略有改善。
cs.LG / 83 / 2609.02887

A Common Measure of Communication for Speech Brain-Computer Interfaces

语音脑机接口的通用交流能力度量方法
Jayalath, Dulhan, Ballyk, Benjamin, Jones, Oiwi Parker
Abstract
Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction. Despite this promise, the field lacks a common measure of progress because systems use different datasets, recording methods, types of speech, and vocabularies, so their reported scores are rarely comparable. Underlying this measurement problem are two unresolved questions: (i) what distribution of words should a speech BCI enable a user to communicate, and (ii) how much information from this distribution can a system convey. We address both by deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate. This allows capabilities measured under different conditions, such as distinct vocabularies, to be evaluated on a common communication scale. We show that ordinarily reported accuracy, word error rate (WER), and other metrics computed only over the words a system supports can overstate how much of a user's intended speech the system can communicate. We then use OVMI to compare existing systems, expose trade-offs between how much of the user's language a system supports and how accurately it decodes those words, show that these comparisons depend on what the user is expected to communicate, and demonstrate that selecting a vocabulary to maximise OVMI yields up to 16.3% relative improvement in accuracy across three speech domains. OVMI therefore provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.
Chinese Translation
语音脑机接口(speech BCIs)将神经活动转化为语言,为瘫痪患者恢复语音能力提供了一条途径,更广泛地说,也为人机自然交互开辟了新形式。尽管前景广阔,该领域仍缺乏衡量进展的通用标准,因为不同系统使用的数据集、记录方法、语音类型和词汇各不相同,其报告的性能指标很少具有可比性。这一度量问题的背后是两个尚未解决的问题:(i)语音脑机接口应使用户能够交流何种词语分布;(ii)系统能从该分布中传递多少信息。我们通过推导开放词汇互信息(OVMI,open-vocabulary mutual information)来解决这两个问题。OVMI是一种信息论量,用于度量解码器相对于用户可能希望表达的词语参考分布所传递的信息量。这使得在不同条件下(如不同词汇集)测得的能力能够在统一的交流尺度上进行评估。我们表明,仅在系统支持的词语上计算的常规准确率、词错误率(WER)及其他指标,可能会高估系统所能传递的用户意图语音的信息量。随后,我们利用OVMI对现有系统进行比较,揭示了系统支持用户语言的程度与解码这些词语的准确性之间的权衡关系,表明这些比较取决于用户预期交流的内容,并证明通过选择使OVMI最大化的词汇,可在三个语音领域实现最高16.3%的相对准确率提升。因此,OVMI为语音脑机接口领域提供了一种有原则的方法,用于比较异构系统、改进词汇设计并衡量该领域的进展。
机器人学 (Robotics)
31
cs.RO / 1 / 2609.01662

Not All Agreement Counts as Corroboration: Provenance-Conserving Multi-View Fusion for Typed Action Admission in Human-Robot Collaboration

并非所有一致性都构成佐证:人机协作中面向类型化动作准入的保来源多视图融合
Jin, Zekai, Zhang, Hanrong, Tang, Yihong, Hu, Fei, Dong, Zhen, Shao, Yi
Abstract
For embodied systems, predictive agreement alone does not determine whether evidence warrants action; evidential origin matters. Repeated inference over one observation can multiply agreement without adding evidence, while source-local values do not reveal whether outputs have separately countable origins. PACT treats evidence countability as a relational variable for provenance-conserving fusion and typed action admission. A supplied provenance partition defines countable units. PACT retains coordinatewise support shared within each unit, accumulates only across units, and maps unmet release conditions to hold, confirm, or fallback. Under the stated assumptions, source-local values cannot identify countability; the coordinatewise meet is the greatest budget satisfying singleton fidelity and insertion non-amplification, with coarsening monotonicity and fixed-partition stability. Across 31,200 evaluations in 48 scene clusters, PACT attains a common-support normalized risk-coverage area (ncsAURC) of 0.0861. Excluding the constructed adversarial-consensus arm, provenance-partition aggregation reduces ncsAURC by 0.0557 relative to singleton aggregation, while the corroboration contrast vanishes. On complete-source records, native scores favor PACT, but a common posterior-peak score narrows its difference from nested Dirichlet and favors product fusion. Reassigning provenance over unchanged predictions moves evidence budgets as predicted. In offline human-robot collaboration, eightfold within-camera duplication leaves 720 typed responses per checkpoint unchanged; camera-grouped PACT admits 47 of 57 Qwen3-VL-32B reference-consistent candidates with no observed reference-inconsistent admission in 60 episodes. PACT separates computational from evidential multiplicity: agreement constitutes corroboration only when provenance permits separate accumulation.
Chinese Translation
对于具身系统而言,仅凭预测一致性并不能决定证据是否足以支持行动;证据的来源同样重要。对同一观测的重复推理可以在不增加证据的情况下成倍放大一致性,而源局部值无法揭示各输出是否具有可分别计数的来源。PACT(Provenance-Conserving Typed Admission,保来源类型化准入)将证据可计数性视为一个关系变量,用于保来源融合与类型化动作准入。给定的来源划分定义了可计数单元。PACT 保留每个单元内共享的逐坐标支持度,仅跨单元进行累加,并将未满足的释放条件映射为保持、确认或回退。在所述假设下,源局部值无法识别可计数性;逐坐标下确界(coordinatewise meet)是同时满足单例保真性与插入不放大性的最大预算,并具有粗化单调性与固定划分稳定性。在涵盖48个场景簇的31,200次评估中,PACT 达到共同支持度归一化风险-覆盖面积(ncsAURC)0.0861。排除构造的对抗性一致分支后,来源划分聚合相对于单例聚合将 ncsAURC 降低了0.0557,而佐证对比效应消失。在完整来源记录上,原生分数更倾向于 PACT,但采用共同后验峰值分数时,其与嵌套 Dirichlet 的差距缩小,且乘积融合更受青睐。在预测结果不变的情况下重新指派来源,会使证据预算按预期方式移动。在离线人机协作实验中,八倍的同一相机内重复不会改变每个检查点720个类型化响应;按相机分组的 PACT 在60个回合中对57个与 Qwen3-VL-32B 参考一致的候选接纳了47个,且未观察到接纳与参考不一致的情况。PACT 区分了计算多重性与证据多重性:只有当来源允许分别累加时,一致性才构成佐证。
cs.RO / 2 / 2609.01731

TriSAR: Task Coordination and Collision Avoidance for Aerial Robot Teams in Disaster Response

TriSAR:面向灾难救援的空中机器人团队任务协同与碰撞规避
Kapile, Aditya Anil, Machado, Pedro, Ihianle, Isibor Kennedy
Abstract
Multi-Unmanned Aerial Vehicle (UAV) disaster-response systems require coordinated task assignment and local trajectory control, yet the individual and combined contributions of these coordination layers to mission efficiency and operational safety remain insufficiently characterised under controlled experimental conditions. TriSAR is evaluated as a five-UAV coordination system operating in a physics-based Gazebo simulation of an earthquake-damaged urban environment. A 2 x 2 factorial design compares two task-allocation strategies (Genetic Algorithm and greedy fitness-based allocation) with reactive collision avoidance enabled or disabled. Each of the four configurations was evaluated over 30 stochastic episodes in a common scenario of five UAVs and eight targets. Under greedy allocation, enabling repulsion eliminated recorded collision-threshold violations, confirmed by a Mann-Whitney test (U = 885, p = 4.03 x 10^-12, rank-biserial r = 0.97). Under GA allocation, the same protective effect was confirmed (U = 675, p = 1.26 x 10^-5, rank-biserial r = 0.50). For mission-efficiency metrics, GA-based allocation showed no statistically detectable advantage over greedy allocation when repulsion was enabled, but a significant advantage in steps, path length, and energy when repulsion was disabled (Welch's t-tests, |g| between 0.92 and 1.76). These results show that reactive repulsion provides a substantial, allocation-dependent safety benefit, while the additional computational complexity of GA-based task allocation yields a detectable mission-efficiency benefit only when repulsion is disabled.
Chinese Translation
多无人机(UAV)灾难救援系统需要协同的任务分配与局部轨迹控制,然而在受控实验条件下,这些协同层级对任务效率和运行安全各自及联合的贡献尚未得到充分刻画。本文在基于物理的Gazebo仿真环境中,将TriSAR作为一个五机协同系统,在地震受损城市环境中进行评估。实验采用2×2析因设计,比较两种任务分配策略(遗传算法与基于适应度的贪心分配)在开启或关闭反应式碰撞规避时的表现。四种配置均在五架无人机、八个目标的共同场景下各进行30次随机重复试验。在贪心分配下,开启斥力规避可消除所有记录到的碰撞阈值违规,Mann-Whitney检验予以证实(U = 885, p = 4.03×10⁻¹², 秩双列相关r = 0.97)。在遗传算法分配下,同样的安全效应也得到了证实(U = 675, p = 1.26×10⁻⁵, 秩双列相关r = 0.50)。就任务效率指标而言,当斥力规避开启时,遗传算法分配相较贪心分配未表现出统计学上可检测的优势;而当斥力规避关闭时,遗传算法分配在步数、路径长度和能耗方面具有显著优势(Welch t检验,|g|介于0.92至1.76之间)。这些结果表明,反应式斥力规避提供了显著的、依赖于分配策略的安全收益,而基于遗传算法的任务分配所带来的额外计算复杂性,仅在斥力规避关闭时才表现出可检测的任务效率收益。
cs.RO / 3 / 2609.01799

Designing Versatile Samples for Learned Trajectory Scoring

面向学习型轨迹打分的多用途样本设计
Li, Yaguang, Zhang, Jiaru, Wei, Chuheng, Cui, Can, Wang, Ziran
Abstract
Many current end-to-end driving policies emit a pool of candidate trajectories and select one, which makes selection a separable component: a scorer can be retrained while the planner, its backbone, and its trajectory generator all stay frozen. However, many strong planners concentrate their proposals around safe mode, providing limited supervision near decision boundaries. In this work, we design a training dataset that provides more informative supervision for the scorer. In particular, we construct two generators that perturb the logged human trajectory along the two axes a vehicle can be displaced: laterally toward the drivable boundary and longitudinally toward a leading vehicle. The designed dataset produces more informative positive and negative samples than the base planner's proposal pool. We attach a transformer-based scorer to two frozen generative planners, DiffusionDrive and MeanFuser, and train it on the NAVSIM navtrain dataset. The results of the experiments show that we achieve 90.1 EPDMS on DiffusionDrive and 90.4 EPDMS on MeanFuser when using ResNet-34, with 0.4 and 0.3 EPDMS respectively, from the designed training dataset.
Chinese Translation
当前许多端到端驾驶策略会生成一组候选轨迹并从中选择一条,这使得轨迹选择成为一个可分离的组件:打分器可以在规划器、其骨干网络以及轨迹生成器全部保持冻结的情况下进行重新训练。然而,许多强大的规划器将其候选轨迹集中在安全模式附近,在决策边界附近提供的监督信息有限。在本工作中,我们设计了一个训练数据集,为打分器提供更具信息量的监督。具体而言,我们构建了两个生成器,沿车辆可能被位移的两个轴对记录的人类轨迹进行扰动:横向朝可行驶边界方向以及纵向朝前方车辆方向。所设计的数据集能够产生比基础规划器候选轨迹池更具信息量的正样本和负样本。我们将一个基于Transformer的打分器附加到两个冻结的生成式规划器——DiffusionDrive和MeanFuser上,并在NAVSIM navtrain数据集上进行训练。实验结果表明,使用ResNet-34时,我们在DiffusionDrive上达到了90.1 EPDMS,在MeanFuser上达到了90.4 EPDMS,其中所设计的训练数据集分别带来了0.4和0.3的EPDMS提升。
cs.RO / 4 / 2609.01938

One Demonstration, Many Objects: Generalizing Manipulation via Local Contact Geometry

一次演示,多种物体:基于局部接触几何的操作泛化
Sharma, Satvik, Sahoo, Samrat, Huang, Huang, Li, Fei-Fei, Wu, Jiajun, Sadigh, Dorsa, Bohg, Jeannette
Abstract
Dexterous manipulation with multi-fingered robot hands promises human-level dexterity, but collecting large-scale dexterous robot hand data remains difficult. Learning from human demonstrations has emerged as a scalable alternative to robot teleoperation, providing strong priors on object interaction and contact strategies. Recent sim-to-real RL methods incorporate such priors, but often (i) omit rewards that explicitly incentivize precise contact, yielding weak real-world performance, and/or (ii) generalize poorly to unseen object instances. We propose DemoMimic (Dexterous Motion Mimic), a policy that manipulates objects by focusing on their geometry local to the contact points. Its contact-centric rewards encourage precise contact and improve sim-to-real consistency, yielding a single real-world policy that transfers across objects of varying shape, scale, mass, and friction wherever local contact structure is preserved. Real-world ablations show that DemoMimic achieves 71% success across 16 objects, four tasks, and two robot-hand embodiments, with the smallest sim-to-real drop compared to baselines.
Chinese Translation
多指机器人手的灵巧操作有望达到人类水平的灵巧性,但收集大规模的灵巧机器人手数据仍然困难。从人类演示中学习已成为机器人遥操作的一种可扩展替代方案,能够为物体交互和接触策略提供强先验。近期的sim-to-real强化学习方法引入了此类先验,但往往存在以下问题:(i) 缺乏显式激励精确接触的奖励,导致真实世界性能较弱;和/或 (ii) 对未见过的物体实例泛化能力差。我们提出了DemoMimic(Dexterous Motion Mimic,灵巧动作模仿),该策略通过关注物体在接触点附近的局部几何来实现物体操作。其以接触为中心的奖励鼓励精确接触并提升sim-to-real一致性,从而形成一个单一的实机策略,只要局部接触结构得以保留,即可迁移至形状、尺度、质量和摩擦力各异的物体。真实世界消融实验表明,DemoMimic在16个物体、4个任务和2种机器人手硬件上取得了71%的成功率,且与基线方法相比具有最小的sim-to-real性能损失。
cs.RO / 5 / 2609.01961

MACAW: Reliable And Efficient Surgical Debridement Using Monocular Adaptive Compact Attention Windows

MACAW:基于单目自适应紧凑注意力窗口的可靠高效外科清创
Chen, Ziyang, Jin, Shutong, Satish, Preethi, Mann, Sareena, Magner, Cael, Fer, Danyal, Mohareri, Omid, Guthart, Gary, Goldberg, Ken
Abstract
Augmenting the dexterity of human surgeons has the potential to free them from tedious subtasks. We consider debridement (removal of diseased or dead tissue fragments), which is challenging due to imprecision in spatial perception and cable actuation. We develop an augmented dexterity system for surgical debridement that uses visual servoing to align the cable-driven gripper with the target position in the image plane, and then introduces a novel approach to depth control, MACAW: Monocular Adaptive Compact Attention Windows. Across 100 physical trials using the da Vinci Research Kit (dVRK) robot, camera-frame servoing reduced average gripper position offset from 37 to fewer than 5 pixels within 4 optimization steps, taking an average of only 0.39s. MACAW significantly outperforms procedural and learned VLA baselines, achieving a 93% success rate at 11 seconds per fragment, yielding a throughput of 304 fragments per hour. Extending MACAW to a bimanual debridement setup maintains a 92% success rate at an average of 7 seconds per fragment, increasing the throughput to 473 fragments per hour.
Chinese Translation
增强人类外科医生的灵巧性有望使其从繁琐的子任务中解放出来。我们研究了清创手术(去除病变或坏死组织碎片),该任务因空间感知和线缆驱动的不精确性而极具挑战性。我们开发了一套用于外科清创的增强灵巧性系统,该系统利用视觉伺服将线缆驱动的夹爪与图像平面中的目标位置对齐,并提出了一种新颖的深度控制方法——MACAW:单目自适应紧凑注意力窗口(Monocular Adaptive Compact Attention Windows)。在使用达芬奇研究套件(da Vinci Research Kit, dVRK)机器人进行的100次物理实验中,基于相机坐标系的伺服控制在4个优化步骤内将夹爪的平均位置偏移从37像素降至不足5像素,平均耗时仅0.39秒。MACAW显著优于程序化基线和基于学习的VLA基线,实现了93%的成功率,每个碎片耗时11秒,吞吐量达到每小时304个碎片。将MACAW扩展到双臂清创设置后,成功率保持92%,每个碎片平均耗时7秒,吞吐量提升至每小时473个碎片。
cs.RO / 6 / 2609.02003

Design and Validation of a Lightweight, Low-Profile Powered Knee Prosthesis with Quasi-Direct Drive Actuation

基于准直驱驱动技术的轻量化低轮廓动力型膝关节假体的设计与验证
Cortino, Ross J., Posh, Ryan, Keller, Emily G., Gregg, Robert D.
Abstract
Fully-powered knee prostheses, unlike traditional passive knees, can perform controlled positive work, reducing the need for compensatory behaviors by users during energy-intensive activities. While quasi-direct drive (QDD) actuators provide superior torque control, backdrivability, and acoustic noise properties compared to traditional highly-geared actuators, prior QDD prototypes have been too heavy and bulky for commercial translation. In this work, we present the design and validation of a new lightweight (2.6 kg) and low-profile (24.5 cm tip-to-tip build height) QDD knee prosthesis. By optimizing an 18 to 1 two-stage transmission alongside thermal and structural finite-element analyses, we significantly reduce device mass while enabling a peak torque of 145 Nm. Through benchtop tests, we validate the device's high output torque, low backdrive torque (1 Nm), and its precision position and torque control capabilities. We also demonstrate biomimetic kinematics and peak knee extension torques (within one standard deviation of able-bodied references) during both level-ground walking and sit-stand transitions performed by three participants with transfemoral amputation and varying K-levels. By meeting or improving upon the mass, build height, peak torque, and acoustic noise of a leading commercial powered knee, this work establishes the clinical viability of emerging QDD prostheses that promise improved dynamic performance for their users.
Chinese Translation
与传统的被动式膝关节不同,全动力型膝关节假体能够执行受控的正功,从而减少使用者在高能耗活动中所需的代偿行为。准直驱(quasi-direct drive, QDD)执行器与传统高减速比执行器相比,具有更优的力矩控制、反向可驱动性和声学噪声特性,但以往的QDD原型过于笨重,难以实现商业化转化。在本研究中,我们提出并验证了一种新型轻量化(2.6 kg)、低轮廓(总高度24.5 cm)的QDD膝关节假体。通过优化18:1的两级传动系统,并结合热学和结构有限元分析,我们在显著降低设备质量的同时实现了145 Nm的峰值力矩。通过台架测试,我们验证了该设备的高输出力矩、低反向驱动力矩(1 Nm)以及精确的位置和力矩控制能力。我们还通过三名具有不同K级别的经股骨截肢受试者,在平地行走和坐立转换过程中展示了仿生运动学特性及峰值膝关节伸展力矩(在健全人参考值的一个标准差范围内)。通过在质量、整体高度、峰值力矩和声学噪声方面达到或超越一款领先商用动力膝关节,本研究证明了新兴QDD假体的临床可行性,这类假体有望为其使用者带来更优的动态性能。
cs.RO / 7 / 2609.02020

Real-Time Dynamics-Based Torque-Sampling MPPI for Compliant and Force Aware Manipulation

基于实时动力学模型的力矩采样MPPI柔顺与力感知操作方法
Im, Euncheol, Kim, Taehyun, Oh, Yonghwan, Lim, Myotaeg, Lee, Yisoo
Abstract
This study proposes a novel Model Predictive Path Integral (MPPI)-based task-space control framework. The proposed framework explicitly solves rigid-body dynamics within a real-time MPC formulation and enforces safety constraints, enabling accurate motion and force control that yields compliant behaviors for safe and effective physical interaction of robotic manipulators in unstructured environments. By leveraging MPPI, the proposed framework efficiently handles nonlinear dynamics that are difficult to solve with conventional MPC approaches in real-time. Furthermore, we develop a torque-sampling-based control architecture that enables efficient exploitation of GPU-based parallelization, resulting in effective compliant and force-aware behaviors. As a result, the proposed framework achieves a solver update rate of over 166 Hz with a 0.18 s prediction horizon, and its performance is validated through real-world experiments on a 7-DoF manipulator.
Chinese Translation
本研究提出了一种新颖的基于模型预测路径积分(Model Predictive Path Integral, MPPI)的任务空间控制框架。该框架在实时模型预测控制(MPC)中显式求解刚体动力学,并施加安全约束,从而实现精确的运动与力控制,使机器人操纵器能够在非结构化环境中安全有效地进行物理交互,表现出柔顺行为。借助MPPI,该框架能够高效处理传统MPC方法难以实时求解的非线性动力学问题。此外,我们开发了一种基于力矩采样的控制架构,可高效利用基于GPU的并行计算,从而实现有效的柔顺与力感知行为。最终,所提出的框架在0.18秒预测时域下实现了超过166 Hz的求解器更新频率,并通过在7自由度(7-DoF)机械臂上的真实世界实验验证了其性能。
cs.RO / 8 / 2609.02046

Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation

建模变化之处:面向以对象为中心操作任务的稀疏残差世界模型
Thakkar, Param, Shah, Parsika Paresh, Gote, Manisha Sushant
Abstract
Monolithic world models predict the entire next state at every step, spending capacity re-predicting the static majority of a scene and injecting error into it. We ask whether explicitly modeling change (a per-object change gate plus a residual delta head that perturbs only the objects the gate flags) is a more effective and interpretable bias for physical prediction and control. On a MuJoCo tabletop pushing benchmark scaling from 3 to 8 objects, the sparse/residual model predicts next-state poses 2.5 to 4.6 times more accurately than a dense multilayer perceptron at 8.6 to 11.1 times fewer parameters, sustains change-detection F1 of 0.80 to 0.87 where the dense baseline is degenerate, transfers across object counts with zero retraining (99.4 percent F1 retention), and reaches about 90 percent of its full-data accuracy with a quarter of the data. In autoregressive rollout it compounds far less error, hugging the no-motion floor while the dense model drifts. Finally, inside a sampling-based planner, prediction-only models fail (though a true-simulator oracle solves the task with the identical planner, confirming the planner is sound), but once featurized and trained for the states a planner visits, the sparse model begins to plan (0.23 plus or minus 0.06 success over three seeds) while the dense monolith stays at zero at every seed. Modeling what changes, rather than re-predicting the whole world, is a simple, effective bias for object-centric physical AI; code, data generators, and all checkpoints will be released upon publication.
Chinese Translation
整体式(monolithic)世界模型在每一步都预测完整的下一状态,将模型容量耗费在重复预测场景中静态占多数的部分,并向其中注入误差。我们探究这样一个问题:显式地建模变化(即每个对象的门控机制加上一个仅对被门控标记的对象进行扰动的残差增量头)是否是一种更有效且更可解释的物理预测与控制归纳偏置。在一个从3个对象扩展到8个对象的MuJoCo桌面推动基准上,稀疏/残差模型在参数量减少8.6至11.1倍的情况下,下一状态位姿预测精度比密集多层感知机高2.5至4.6倍;在密集基线表现退化的情形下,其变化检测F1保持在0.80至0.87;在对象数量变化时无需重新训练即可迁移(保留99.4%的F1);并且仅用四分之一的数据即可达到全量数据约90%的精度。在自回归滚动预测中,其误差累积远小于密集模型,紧贴无运动基线,而密集模型则出现漂移。最后,在基于采样的规划器中,仅用预测的模型无法完成任务(尽管一个真实仿真器oracle可以用相同的规划器解决该任务,证明规划器本身是可靠的);但在将预测模型特征化并针对规划器实际访问的状态进行训练后,稀疏模型开始能够规划(三个随机种子下的成功率为0.23±0.06),而密集整体模型在所有种子下均为零。建模变化之处而非重新预测整个世界,是面向以对象为中心的物理AI的一种简单而有效的归纳偏置;代码、数据生成器及所有检查点将在论文发表后发布。
cs.RO / 9 / 2609.02079

Koopman-Based Robust Model Predictive Control for Nonlinear Systems with Stochastic Intermittent Measurements

基于Koopman的随机间歇测量非线性系统鲁棒模型预测控制
Liu, Guanhua, Wu, Tong, Zhang, Lixian, Du, Weifeng, Han, Minghao
Abstract
Intermittent state measurements pose fundamental challenges to model predictive control of constrained nonlinear systems because prediction uncertainty grows during feedback outages and measurement-triggered resets disrupt nominal state propagation, potentially compromising closed-loop stability and recursive feasibility. This paper develops a Koopman-based stochastic MPC framework with probabilistically truncated soft constraints. Specifically, a Lipschitz-constrained deep Koopman model provides a linear latent predictor, enabling computationally efficient online optimization. The intermittent measurement process is modeled as a two-mode discrete-time Markov chain, yielding a unified Markov jump error model for open-loop propagation and measurement-triggered resets. Under numerically verifiable sufficient conditions, the prediction error is shown to be mean-square ultimately bounded, and an explicit uniform second-moment bound is obtained. A distribution-free probabilistic error radius is then constructed for a prescribed confidence level and used to truncate dropout-dependent constraint tightening. An exact-penalty soft-constraint mechanism accommodates reset-induced jumps and prolonged dropouts. Under the stated terminal compatibility and bounded-disturbance conditions, recursive feasibility and mean-square ultimate boundedness of the closed-loop regulation error are established. Numerical simulations on a visual-servoing tracking task corroborate these theoretical results and demonstrate effective tracking under stochastic measurement unavailability.
Chinese Translation
间歇性状态测量对受约束非线性系统的模型预测控制(MPC)提出了根本性挑战,因为反馈中断期间预测不确定性会增大,且由测量触发的重置会破坏标称状态传播,可能损害闭环稳定性和递归可行性。本文提出了一种基于Koopman算子的随机MPC框架,并采用概率截断软约束。具体而言,一个具有Lipschitz约束的深度Koopman模型提供了线性潜在预测器,使在线优化具备计算高效性。间歇测量过程被建模为双模式离散时间马尔可夫链,从而为开环传播和测量触发的重置提供了统一的马尔可夫跳变误差模型。在数值可验证的充分条件下,预测误差被证明是均方最终有界的,并获得了显式的一致二阶矩界。随后,针对给定置信水平构造了无需分布假设的概率误差半径,并用于对依赖测量丢失的约束收紧进行截断。一种精确罚函数软约束机制可适应重置引起的跳变和长时间的测量丢失。在所述终端兼容性和有界扰动条件下,建立了闭环调节误差的递归可行性和均方最终有界性。在视觉伺服跟踪任务上的数值仿真验证了这些理论结果,并展示了在随机测量不可用情况下的有效跟踪性能。
cs.RO / 10 / 2609.02134

Unified Motion Retargeting for Humanoids with Learned Point Cloud Correspondence

基于学习点云对应关系的人形机器人统一运动重定向
Cao, Hanyang, Fang, Yuetong, Kwon, Taesoo, Yu, Runyi, Ma, Ji, Tan, Jing, Zhou, Yangchen, Du, Baoze, Gu, Yi, Gao, Yukang, Dai, Ruoli, Han, Lei, Xu, Renjing
Abstract
Humanoid learning increasingly relies on transforming vast and diverse human motion data into high-quality robot reference trajectories. However, retargeting human motion to humanoid robots is challenging due to substantial differences in morphology, degrees of freedom, joint ranges, and kinematic constraints between humans and robots. Existing retargeting methods typically address these differences by defining human-robot correspondence through hand-crafted sparse keypoints or body-part pairs. As a result, retargeting quality depends heavily on manual semantic design, limiting scalability across motion sources and robot morphologies and providing only sparse guidance for reproducing detailed poses and interactions. In this paper, we present Unified Motion Retargeting (UMR), a framework that learns dense point cloud correspondence without requiring manually designed human-robot mappings. By treating exterior point clouds as a unified interface between human motion and humanoid robots, UMR decouples retargeting from source-specific skeletal semantics and robot-specific topology. The learned dense correspondence provides fine-grained geometric anchors for constrained point cloud matching optimization, enabling surface-level pose alignment and direct transfer of interaction contacts. Experiments demonstrate that UMR unifies retargeting across heterogeneous motion sources, robot embodiments, and downstream scenarios ranging from locomotion to interaction, while achieving higher motion fidelity and plausibility than state-of-the-art methods. UMR therefore provides a scalable foundation for transforming large-scale human motion references into robot-ready training data.
Chinese Translation
人形机器人的学习日益依赖于将海量且多样的人体运动数据转换为高质量的机器人参考轨迹。然而,由于人体与机器人之间在形态结构、自由度、关节范围以及运动学约束上存在显著差异,将人体运动重定向到人形机器人面临很大挑战。现有的重定向方法通常通过手工设计的稀疏关键点或身体部位配对来定义人体与机器人的对应关系。因此,重定向质量严重依赖人工语义设计,这不仅限制了其在不同运动数据源和机器人形态之间的可扩展性,而且只能为复现精细姿态和交互提供稀疏的引导。在本文中,我们提出了统一运动重定向框架(Unified Motion Retargeting, UMR),该框架能够学习稠密的点云对应关系,而无需人工设计的人机映射。通过将外观点云作为人体运动与人形机器人之间的统一接口,UMR 将重定向过程与特定数据源的骨架语义以及特定机器人的拓扑结构解耦。学习到的稠密对应关系为受约束的点云匹配优化提供了细粒度的几何锚点,实现了表面层面的姿态对齐以及交互接触的直接迁移。实验表明,UMR 能够统一处理异构运动数据源、不同机器人本体以及从运动控制到人机交互等多种下游场景的重定向任务,同时比最先进的方法取得更高的运动保真度与合理性。因此,UMR 为将大规模人体运动参考数据转换为可直接用于机器人训练的数据提供了可扩展的基础。
cs.RO / 11 / 2609.02157

Towards Effective Physical Reservoir Computing with a Pneumatic Soft Robot

利用气动软体机器人实现有效的物理储备池计算
Manjunath, Jeevan Hebbal, Wang, Jun, Li, Suyi, Zhang, Wenlong
Abstract
Physical reservoir computing (PRC) refers to the use of a physical dynamical system as a computational resource for tasks such as state estimation and control, but there has been a lack of formal study of design rules towards more effective design of such physical reservoirs. Using a pneumatic soft arm with a five-pouch sensing column, this work studies how the pouch interconnection topology, robot stiffness, and the number of instrumented sensors affect bending-angle estimation performance. Across 36 matched trials spanning waveform, baseline pressure of the sensing column, and actuation range, all designs are evaluated under the same-time bending-angle estimation benchmark using 0.2 s of pressure history and a fixed ridge estimator. Our analysis of the experimental results leads to three design guidelines. First, independently sealed pouches preserve a much richer observable state than a shared manifold. Second, increasing the baseline pressure of the sensing column makes the pouch responses more redundant and increases estimation error most strongly in the coupled topology. Third, in the sealed topology, two strategically placed sensors already recover most of the attainable benefit, three capture essentially all of it, and additional sensors provide little or no additional value. In summary, the results suggest that topology, stiffness, and number of instrumented sensors should be co-designed for accurate PRC of soft robot states; stronger excitation alone cannot recover the diversity that poor design choices have already removed.
Chinese Translation
物理储备池计算(Physical Reservoir Computing, PRC)是指将物理动力学系统用作计算资源,以完成状态估计与控制等任务,但目前缺乏针对更有效设计此类物理储备池的设计规则的正式研究。本工作使用带有五囊传感柱的气动软体机械臂,研究了囊间互联拓扑、机器人刚度以及装配传感器数量对弯曲角度估计性能的影响。在涵盖波形、传感柱基准压力和驱动范围的36组匹配试验中,所有设计均在相同的弯曲角度估计基准下进行评估,使用0.2秒的压力历史和固定的岭回归估计器。对实验结果的分析得出了三条设计准则。第一,独立密封的气囊比共享流道的拓扑保留了丰富得多的可观测状态。第二,提高传感柱的基准压力会使气囊响应更加冗余,并且在耦合拓扑中估计误差增加最为显著。第三,在密封拓扑中,两个策略性布置的传感器已能获得大部分可达收益,三个传感器基本捕获全部收益,而增加更多传感器几乎不带来额外价值。总之,研究结果表明,拓扑、刚度和装配传感器数量应当协同设计,以实现对软体机器人状态的精确物理储备池计算;仅靠更强的激励无法弥补糟糕设计选择所损失的状态多样性。
cs.RO / 12 / 2609.02219

Hardware-Accelerated Instance Segmentation for Resource-Constrained Space Robotics with Criticality Analysis

面向资源受限空间机器人、结合关键性分析的硬件加速实例分割
Shete, Siddhant, Kücüker, Hilmi Dogu, Frese, Udo, Kirchner, Frank
Abstract
Autonomous lunar missions require real-time per- ception under three coupled constraints: extreme low-light conditions, limited onboard compute, and radiation-induced hardware faults that can silently corrupt inference. We present a deployment-oriented instance segmentation framework for resource-constrained lunar robotics that jointly addresses quan- tization calibration and system-level fault exposure under strict compute constraints. First, we introduce Activation Variance Informative Sampling (AVIS), a label-free calibration strategy that deterministically selects calibration samples based on activation variance statistics. Second, we deploy a YOLO-based segmentation model on a Deep Learning Processor Unit (DPU) with architectural modifications that reduce CPU fallback paths and enable statically compiled execution with bounded latency in low-lighting conditions. We further introduce a software-level criticality analysis to estimate fault exposure and guide mitigation under radiation-constrained operation. On a lunar micro-rover platform, AVIS with bias correction recovers 69.8% of quantization-induced accuracy loss while achieving 309 ms inference latency and 5.7 W power consumption. Targeted mitigation reduces global criticality by 31.7%. The results demonstrate an integrated approach and a blueprint for a reliable and safe AI perception framework under space deployment constraints.
Chinese Translation
自主月球任务需要在三个相互耦合的约束条件下实现实时感知:极端低光照条件、有限的星载计算能力,以及可能悄然破坏推理过程的辐射致硬件故障。我们提出一个面向部署的实例分割框架,适用于资源受限的月球机器人,在严格计算约束下同时解决量化校准与系统级故障暴露问题。首先,我们提出激活方差信息采样(Activation Variance Informative Sampling, AVIS),一种基于激活方差统计量确定性选择校准样本的无标签校准策略。其次,我们在深度学习处理器单元(Deep Learning Processor Unit, DPU)上部署了基于YOLO的分割模型,通过架构修改减少了CPU回退路径,实现了低光照条件下具有有界延迟的静态编译执行。我们进一步引入软件级关键性分析,以估计故障暴露并指导辐射受限运行下的缓解措施。在月球微型巡视器平台上,带偏置校正的AVIS恢复了69.8%的量化导致的精度损失,同时实现了309毫秒的推理延迟和5.7瓦的功耗。针对性缓解措施将全局关键性降低了31.7%。这些结果展示了一种集成化方法,并为在空间部署约束下可靠、安全的AI感知框架提供了蓝图。
cs.RO / 13 / 2609.02222

FOCUS: Foot Observation Confidence for Robust Humanoid Proprioceptive Odometry

FOCUS:基于足部观测置信度的鲁棒人形机器人本体感知里程计
Feng, Kaixin, Li, Angsong, Zhang, Shaopeng, Li, Enyu, Lin, Peiwen, Wang, Chuang, Li, You, Lan, Haiyu
Abstract
Foot forward kinematics (FK) is widely used to improve proprioceptive legged odometry by providing reliable velocity constraints during foot support. Existing contact-aided estimators generally rely on binary contact decisions to determine whether the FK measurements of an entire foot should be trusted. However, contact does not necessarily imply FK reliability. Dynamic locomotion often involves partial support, toe dragging, and foot slip, causing binary contact decisions to accumulate significant drift over long trajectories. To address this limitation, we propose FOCUS (Foot Observation Confidence from Unannotated Simulation), which predicts a continuous FK reliability weight for each foot instead of estimating binary foot contact. Rather than replacing the model-based estimator, the predicted reliability weights are used to blend FK velocity observations with IMU-propagated body velocity and to adapt the observation covariance of an extended Kalman filter (EKF), enabling smooth reliability-aware fusion without hard contact switching. The network is trained from automatically generated simulation signals using an FK-weighted velocity consistency loss with lightweight simulator-contact regularization, without manually annotated continuous FK-reliability labels. The deployed model relies only on IMU and joint kinematic measurements, making it suitable for hardware platforms with unreliable torque sensing. Experiments demonstrate that FOCUS reduces absolute trajectory error (ATE) by 83.7% on simulated walking episodes, preserves simulated dynamic-motion fidelity in motion scale and spectral energy, reduces ATE by 70.8% across 19 real walking segments, and reduces mean ATE by 42.7% across four real dynamic-motion routines.
Chinese Translation
足部前向运动学(FK)被广泛用于提升本体感知腿式里程计的性能,其方法是在足部支撑期间提供可靠的速度约束。现有的接触辅助估计器通常依赖二值接触决策来判断是否应信任整个足部的FK测量。然而,接触并不一定意味着FK的可靠性。动态步态运动常常涉及部分支撑、脚趾拖地和足部打滑,导致二值接触决策在长轨迹上积累显著的漂移。为解决这一局限,我们提出FOCUS(基于无标注仿真的足部观测置信度),它为每个足部预测连续的FK可靠性权重,而非估计二值的足部接触状态。预测的可靠性权重并非用于取代基于模型的估计器,而是用于将FK速度观测与IMU递推的机体速度进行融合,并自适应调整扩展卡尔曼滤波器(EKF)的观测协方差,从而实现无需硬接触切换的平滑可靠性感知融合。该网络利用自动生成的仿真信号进行训练,采用FK加权速度一致性损失并辅以轻量级仿真器接触正则化,无需人工标注的连续FK可靠性标签。部署的模型仅依赖IMU和关节运动学测量,因此适用于力矩传感不可靠的硬件平台。实验表明,FOCUS在仿真行走片段上将绝对轨迹误差(ATE)降低了83.7%,在运动尺度和频谱能量上保持了仿真动态运动的保真度,在19段真实行走数据上将ATE降低了70.8%,并在四段真实动态运动流程上将平均ATE降低了42.7%。
cs.RO / 14 / 2609.02252

DiffuSearch: How Hybrid Trajectory Planning Benefits from Aligned Objectives in Diffusion and Action Space

DiffuSearch:混合轨迹规划如何受益于扩散空间与动作空间中对齐的目标函数
Hagedorn, Steffen, Distelzweig, Aron, Condurache, Alexandru P.
Abstract
In trajectory planning for autonomous driving, hybrid planning architectures are often realized as a collection of disparate modules, each with its own objectives. This lack of a unifying principle can lead to inconsistencies between the initial and refined trajectory, resulting in suboptimal behavior. We address this by introducing DiffuSearch, a novel hybrid planner that uses a unified set of objectives across generation and refinement. Our model encourages all components to follow the same shared driving goals: collision avoidance, drivable area compliance, comfort, and progress. DiffuSearch employs a two-stage architecture. First, a guided diffusion model generates a scene-consistent, joint trajectory prediction, using our driving objectives as differentiable guidance functions to implicitly steer the denoising process. Second, a Monte Carlo Tree Search (MCTS) in a discretized action space performs an explicit, local refinement of this proposal, leveraging the same driving objectives as its reward function. This synergistic design leverages the diffusion model's strength in finding scene-consistent solutions combined with the explainable, constraint-aware refinement of MCTS. Experiments on nuPlan and interPlan reactive closed-loop benchmarks demonstrate that DiffuSearch achieves strong and often state-of-the-art performance, substantially reducing collisions and improving comfort, particularly in complex, interactive scenarios. Our ablation studies indicate that MCTS refinement is the main mechanism behind the gains, while sharing objectives between implicit guidance and explicit search provides further consistent improvements.
Chinese Translation
在自动驾驶的轨迹规划中,混合规划架构通常由一系列各自具有独立目标的模块组成。这种缺乏统一原则的做法可能导致初始轨迹与优化后轨迹之间的不一致,从而产生次优行为。为此,我们提出了DiffuSearch,一种新颖的混合规划器,它在生成与优化阶段使用统一的目标集合。我们的模型促使所有组件遵循相同的共享驾驶目标:避碰、可行驶区域合规、舒适性以及行驶进度。DiffuSearch采用两阶段架构。首先,一个引导式扩散模型生成场景一致性的联合轨迹预测,将我们的驾驶目标作为可微分的引导函数,隐式地引导去噪过程。其次,在离散化动作空间中进行蒙特卡洛树搜索(MCTS),利用相同的驾驶目标作为奖励函数,对该初始轨迹进行显式的局部优化。这种协同设计结合了扩散模型在寻找场景一致性解方面的优势,以及MCTS可解释、约束感知的优化能力。在nuPlan和interPlan反应式闭环基准上的实验表明,DiffuSearch取得了强劲且常常达到最先进水平的性能,显著减少了碰撞并提升了舒适性,尤其是在复杂交互场景中。消融实验表明,MCTS优化是性能提升的主要机制,而在隐式引导与显式搜索之间共享目标函数则带来了进一步的一致性改进。
cs.RO / 15 / 2609.02270

CrashDiffuser: VLM-Guided Collision Intent Reasoning for Fine-Grained Safety-Critical Traffic Scenario Generation

CrashDiffuser:基于视觉语言模型引导的碰撞意图推理细粒度安全关键交通场景生成
Zhang, Shucheng, Zhang, Yuang, Wang, Bingzhang, Karim, Muhammad Monjurul, Chen, Kehua, Wang, Yinhai
Abstract
Generating safety-critical scenarios is essential for evaluating autonomous driving systems. However, existing generators primarily focus on inducing collisions and offer limited control over where contact occurs on the target vehicle. In this paper, we study fine-grained safety-critical scenario generation, where success requires both a target collision and a specified head, rear, or side contact region. We propose CrashDiffuser, a closed-loop VLM-guided diffusion framework that decouples semantic collision reasoning from continuous trajectory synthesis through a hierarchical collision-intent interface derived from the requested target contact region. At initialization, the VLM extracts reusable scene-level context; at each replanning step, it predicts a structured action tuple describing speed change, turning behavior, and collision stage. This intent conditions a diffusion model to generate executable adversarial trajectories, while collision-guided sampling, candidate selection, and short-horizon replanning adapt generation to the target vehicle's evolving behavior. On WOMD-derived closed-loop scenarios, CrashDiffuser achieves a target-collision rate of 50.33% in a single attempt and 67.98% after three attempts, together with a contact-region control success rate of 40.05% and competitive trajectory naturalness. Component ablations further support the proposed design.
Chinese Translation
生成安全关键场景对于评估自动驾驶系统至关重要。然而,现有的生成器主要侧重于诱发碰撞,对碰撞发生在目标车辆具体位置的掌控能力有限。本文研究细粒度安全关键场景生成,其成功需要同时实现目标碰撞以及指定的头部、尾部或侧面碰撞区域。我们提出CrashDiffuser,这是一个闭环的视觉语言模型(VLM)引导的扩散框架,通过由请求的目标碰撞区域导出的分层碰撞意图接口,将语义层面的碰撞推理与连续轨迹合成解耦。在初始化阶段,VLM提取可复用的场景级上下文;在每个重规划步骤中,它预测一个结构化的动作元组,描述速度变化、转向行为和碰撞阶段。该意图作为条件输入扩散模型以生成可执行的对抗性轨迹,同时碰撞引导采样、候选选择和短时程重规划使生成过程能够适应目标车辆不断演变的行为。在基于WOMD的闭环场景上,CrashDiffuser在单次尝试中达到50.33%的目标碰撞率,三次尝试后达到67.98%,同时碰撞区域控制成功率为40.05%,并具有具有竞争力的轨迹自然度。组件消融实验进一步支持了所提出的设计。
cs.RO / 16 / 2609.02306

Contact-Constrained Lower-Limb Joint-Offset Calibration for Humanoid Robots

面向人形机器人的接触约束下肢关节偏置标定
Lu, Kaixiang, Lan, Haiyu, Qiao, Chunxiao, Li, You, Luo, Chengyuan, Li, Enyu, Lin, Peiwen, Wang, Chuang
Abstract
Accurate joint encoder offsets are essential for kinematic consistency in humanoid lower limbs, yet existing calibration methods typically require external motion-capture systems or fiducial targets. We present a self-contained calibration framework exploiting only onboard joint encoders and a pelvis-mounted IMU during static double-support contact. The inter-foot transform from forward kinematics must stay constant when both feet are fixed; minimizing its posture-dependent dispersion yields a nonlinear least-squares problem over the 12-dimensional offset vector. A Hessian eigenstructure analysis shows that parallel pitch axes induce a rotational coupling. Orientation residuals then observe only the pitch-offset sum, while translation and posture diversity set the remaining numerical observability. For the A3 pitch-to-roll-to-yaw ordering, hip-roll and hip-yaw excitation reduce hip-pitch coupling. A standing-posture knee prior then anchors the remaining weak pitch-chain decomposition. Simulation and real-machine injection tests show consistent recovery, and on held-out recordings calibration reduces foot-height RMS residuals from 4.26 to 2.20 mm on A3 and from 8.03 to 1.43 mm on A2. An independent LiDAR-inertial reference checks the pitch-coupled channel. Removing an injected pitch offset moves the leg-odometry vertical drift back toward the LiDAR trajectory. A few static double-support stances thus provide contact-consistent corrections for well-excited directions. Individual offsets in the weak pitch chain remain prior-dependent.
Chinese Translation
精确的关节编码器偏置对人形机器人下肢的运动学一致性至关重要,然而现有的标定方法通常需要外部动作捕捉系统或基准标志物。我们提出了一种自包含的标定框架,仅在双脚静态支撑接触期间,利用机载关节编码器和安装于骨盆的惯性测量单元(IMU)即可完成标定。当双脚固定时,由正向运动学计算的双足间变换应保持恒定;最小化该变换随姿态变化的离散程度,可得到一个关于12维偏置向量的非线性最小二乘问题。Hessian特征结构分析表明,平行的俯仰轴会引起旋转耦合。此时姿态残差仅能观测到俯仰偏置之和,而平移与姿态的多样性决定了其余分量的数值可观测性。对于A3的俯仰—横滚—偏航(pitch-roll-yaw)排序,髋横滚与髋偏航的激励可减弱髋俯仰的耦合。随后,站立姿态的膝关节先验锚定了剩余弱可观测俯仰链的分解。仿真与实机注入测试均显示出一致的恢复效果;在留出测试数据上,标定将足部高度均方根(RMS)残差在A3上从4.26 mm降至2.20 mm,在A2上从8.03 mm降至1.43 mm。独立的激光雷达—惯性参考用于检验俯仰耦合通道:移除注入的俯仰偏置后,腿部里程计的垂直漂移重新趋近于激光雷达轨迹。因此,仅需若干静态双脚支撑姿态,即可为激励充分的自由度提供接触一致的校正,而弱俯仰链中的单个偏置仍依赖于先验信息。
cs.RO / 17 / 2609.02319

From Multi-Fisheye Sensing to Panoramic Perception: A Parallax-Aware Onboard Platform for Ultra-Low-Altitude UAVs

从多鱼眼感知到全景感知:一种面向超低空无人机的视差感知机载平台
Dai, Dun, Lu, Ze, He, Cheng, Wang, Yaowen, Quan, Quan
Abstract
Ultra-low-altitude unmanned aerial vehicles (UAVs) require surround vision near buildings, vegetation, and other obstacles. We present a parallax-aware onboard platform that converts four synchronized fisheye streams into an open 1280x640 equirectangular panorama (ERP) interface. A purpose-built carbon-fiber airframe integrates the cameras, NVIDIA Jetson Orin NX, a flight controller, and a global navigation satellite system (GNSS) receiver. The formation pipeline selects projection depth per overlap and combines controlled seams and photometric fusion. Its accuracy profile adds content-adaptive seam search and a validation-gated residual mesh, whereas its deployed profile retains margin-gated Per-seam updates for sensor-rate operation. Evaluation uses more than 50,000 four-view groups from 18 field sequences. Relative to Fixed Depth, the accuracy profile reduces far-field P90 feature misalignment by 41.6%; the deployed Per-seam profile achieves the lowest aggregate geometric errors across held-out sites. Under a paced 20 Hz replay, the deployed profile sustains 19.99 frames/s at 13.29W mean module-input power. Eight-sector ERP sampling reaches 90.8% mean daytime visual-place-recognition Recall@5. Together, these results validate an integrated onboard panoramic-perception architecture that unifies parallax-aware formation, sensor-rate embedded execution, and reusable downstream vision interfaces for ultra-low-altitude UAVs. The project has been open-sourced at https://github.com/DUNDAI1998/parallax-aware-uav-panorama.
Chinese Translation
超低空无人机(UAV)在建筑物、植被及其他障碍物附近飞行时需要环视视觉能力。我们提出了一种视差感知机载平台,可将四路同步鱼眼视频流转换为开放的1280×640等距圆柱投影(equirectangular panorama, ERP)接口。专门设计的碳纤维机身集成了相机、NVIDIA Jetson Orin NX、飞行控制器以及全球导航卫星系统(GNSS)接收器。拼接流水线针对每个重叠区域选择投影深度,并结合受控接缝与光度融合。其精度配置增加了内容自适应的接缝搜索和经过验证门控的残差网格,而其部署配置则保留了边距门控的逐接缝更新以实现传感器速率运行。评估使用了来自18个野外序列的超过50,000组四视角数据。相对于固定深度方法,精度配置将远场P90特征错位降低了41.6%;部署的逐接缝配置在留出测试场地上实现了最低的总体几何误差。在20 Hz节奏回放下,部署配置以13.29W的平均模块输入功率维持了19.99帧/秒的处理速度。八扇区ERP采样在白天场景下实现了90.8%的平均视觉位置识别Recall@5。这些结果共同验证了一种集成的机载全景感知架构,该架构统一了视差感知拼接、传感器速率嵌入式执行以及面向超低空无人机的可复用下游视觉接口。该项目已在 https://github.com/DUNDAI1998/parallax-aware-uav-panorama 开源。
cs.RO / 18 / 2609.02358

Humanoid Safe Stop via Learned Stoppability Value

基于学习可停止性价值的人形机器人安全停止
Long, Junfeng, Abbeel, Pieter, Sreenath, Koushil, Horowitz, Roberto, Shi, Guanya, Liu, C. Karen
Abstract
Humanoid robots responding to emergency stop commands typically execute a fixed maneuver, without reasoning about whether a safe stop is actually feasible from the current state. We cast emergency stopping as a reach-avoid problem and propose Safe-Stop, a task-agnostic framework that pairs a learned stop policy with learned stoppability estimators. The estimators are complementary: a stop-probability estimator supervised by the actual outcomes of the fixed stop policy, and a reach-avoidance estimator supervised by a Hamilton-Jacobi backup over physical state. The first captures emergent stopping behavior of the learned controller; the second provides a complementary recoverability signal. Because the stop policy and estimators do not depend on the behavior policy that preceded the stop command, they transfer across diverse upstream tasks without retraining. At deployment, the two estimates are combined: Safe-Stop commits to the stop only when both estimators indicate that stopping remains feasible, otherwise it hands off to a fall policy, instantiated as a damping fallback. This agreement check yields decisions that are robust without sacrificing reactivity.
Chinese Translation
人形机器人在响应紧急停止指令时,通常执行固定的动作,而不会推理从当前状态出发安全停止是否真正可行。我们将紧急停止建模为一个可达-规避问题,并提出 Safe-Stop——一个任务无关的框架,它将学习得到的停止策略与学习得到的可停止性估计器相结合。这两个估计器互为补充:一个是停止概率估计器,由固定停止策略的实际结果进行监督训练;另一个是可达-规避估计器,由基于物理状态的 Hamilton-Jacobi 备份策略进行监督训练。前者捕捉学习到的控制器的涌现停止行为;后者提供互补的可恢复性信号。由于停止策略和估计器不依赖于停止指令之前的行为策略,因此它们无需重新训练即可迁移到多种上游任务。在部署时,两个估计结果被组合使用:仅当两个估计器都表明停止仍然可行时,Safe-Stop 才执行停止,否则交由摔倒策略处理(本文实现为阻尼兜底策略)。这种一致性检验在不牺牲响应速度的前提下,使决策更加鲁棒。
cs.RO / 19 / 2609.02402

A Physics-Consistent Benchmark for Contact-Rich Human-Robot Interaction in Assistive Care

面向辅助护理中接触密集型人机交互的物理一致性基准
He, Chengxiao, Yuan, Shanghai, Fan, Liuqun, Zhu, Shenzhen
Abstract
Conventional task-level evaluation asks whether a robot policy completes a specified action, but can miss failures that emerge only during physical human contact. This limitation is critical in contact-rich assistive tasks, where meaningful evaluation requires a physically responsive human, interaction-quality assessment beyond task success, and a leak-free observer-scorer protocol. We introduce a physics-consistent benchmark for contact-rich human-robot interaction, instantiated in robot-assisted bathing. The benchmark combines a deformable, passively responding human, physics-aware scores alongside task-level success, and a frozen vision-only / scorer-only evaluation protocol. To establish physical validity, region-wise simulated responses are calibrated against force-indentation measurements from Franka impedance pushes on a medical-care manikin. Under a frozen T1-T7 protocol with 140 runs per method, an LLM-augmented state machine (State Machine) achieves 72.9% task success but drops to 56.4% after correct-region and force-safety screening; VoxPoser produces lighter and more stable contact but completes only 27.9% of trials; and zero-shot pi0.5 achieves 0.7% task success with no correct-region or safety-gated successes. These results show that task completion alone does not imply physically valid contact and motivate physics-aware screening before deployment of contact-rich assistive robot policies.
Chinese Translation
传统的任务级评估仅询问机器人策略是否完成了指定动作,但可能忽略仅在人机物理接触过程中才会显现的失败。这一局限在接触密集型的辅助任务中尤为关键,因为有意义的评估需要具备物理响应能力的人体模型、超越任务成功的交互质量评估,以及无信息泄漏的观察者-评分者协议。我们提出了一个面向接触密集型人机交互的物理一致性基准,并以机器人辅助洗浴为实例加以实现。该基准结合了可变形且被动响应的人体模型、在任务级成功之外的物理感知评分,以及冻结的仅视觉/仅评分者评估协议。为确立物理有效性,我们将分区域的仿真响应与通过Franka阻抗控制在医疗护理人体模型上测得的力-压痕数据进行校准。在每方法140次运行、冻结的T1-T7协议下,基于大语言模型增强的状态机方法任务成功率达72.9%,但经正确区域与力安全筛查后降至56.4%;VoxPoser产生了更轻、更稳定的接触,但仅完成27.9%的试验;零样本pi0.5的任务成功率仅为0.7%,且在正确区域与安全门控方面无一成功。这些结果表明,仅完成任务并不意味着物理上有效的接触,并为在部署接触密集型辅助机器人策略之前进行物理感知筛查提供了依据。
cs.RO / 20 / 2609.02487

An Adaptive Control Architecture for Slope and Terrain Compensation in Autonomous Navigation in Mediterranean Greenhouses

一种用于地中海温室自主导航中坡度与地形补偿的自适应控制架构
Cañadas-Aránega, Fernando, Wollherr, Dirk, Guzmán, José L., Moreno, José C., Blanco-Claraco, José L.
Abstract
The ability to move stably over terrain with varying slopes and textures is essential for mobile agricultural robots operating in complex and dynamic environments such as greenhouses, where small terrain irregularities can lead to significant navigation errors. This article presents a novel terrain-adaptation strategy based on the carried payload, ensuring accurate and robust trajectory tracking. The proposed approach is based on: (i) the experimental characterization of the most common types of greenhouse soil, concrete, compacted sand, and gravel, and (ii) the direct measurement of terrain slope using the IMU, in order to estimate the force with which this angle affects the motor input. Based on this information, a cascade trajectory-tracking scheme has been designed, consisting of a model-based predictive controller (MPC) in the outer loop and a PI controller in the inner loop. The system incorporates an adaptive feedforward control through gain scheduling approach, capable of adjusting to disturbances caused by variations in slope and terrain type. Simulation results demonstrate that the differential-drive robot achieves a significant improvement both in error indices and in control signal efficiency, highlighting the effectiveness and robustness of the proposed approach.
Chinese Translation
对于在温室等复杂动态环境中作业的移动农业机器人而言,能够在坡度和纹理多变的 terrain 上稳定移动至关重要,因为微小的地形不规则性可能导致显著的导航误差。本文提出了一种基于所携带载荷的新型地形自适应策略,以确保精确且鲁棒的轨迹跟踪。所提出的方法基于:(i)对温室中最常见土壤类型——混凝土、压实沙土和砾石——的实验表征;(ii)利用惯性测量单元(IMU)直接测量地形坡度,以估计该角度对电机输入所产生的力的影响。基于这些信息,设计了一种级联轨迹跟踪方案,其外环为基于模型的预测控制器(MPC),内环为PI控制器。该系统通过增益调度(gain scheduling)方法引入自适应前馈控制,能够适应由坡度和地形类型变化引起的扰动。仿真结果表明,差速驱动机器人在误差指标和控制信号效率方面均取得了显著改进,凸显了所提方法的有效性和鲁棒性。
cs.RO / 21 / 2609.02493

MS-MEM: Multi-Skill Manipulation-Enhanced Mapping via Uncertainty- and Disturbance-Aware Action Selection

MS-MEM:基于不确定性与扰动感知动作选择的多技能操作增强建图
Shi, Yitian, Mücke, Jesper, Dengler, Nils, Pan, Sicong, Rayyes, Rania, Bennewitz, Maren
Abstract
Accurate scene understanding in confined, cluttered spaces such as shelves is essential for service robots, as many everyday tasks require them to locate and retrieve objects reliably. Yet, it remains challenging due to severe occlusions, restricted accessibility, and the need to avoid excessive scene changes. In this paper, we propose Multi-Skill Manipulation-Enhanced Mapping (MS-MEM), an evidential framework for uncertainty-aware mapping that integrates active viewpoint selection, object pushing, and grasping. MS-MEM combines scene-level metric-semantic evidential belief estimators with an uncertainty-aware grasp representation. This representation is learned using a novel full-evidential grasp estimator that models both grasp affordance and orientation uncertainty. In our framework, candidate perception and manipulation actions are evaluated within a unified action selection pipeline using a common information gain criterion. For manipulation actions, we further introduce a collateral disturbance constraint (CDC) that discourages excessive changes to confident regions of the scene belief. This enables MS-MEM to select actions that effectively reduce map uncertainty while limiting collateral scene changes. Experimental results show that, compared with single-skill and unconstrained baselines that ignore scene disturbance, MS-MEM achieves higher mapping accuracy while substantially reducing scene disturbance, highlighting the synergistic effects of active viewpoint selection, push, and grasp actions.
Chinese Translation
在货架等狭窄、杂乱的空间中实现准确的场景理解对服务机器人至关重要,因为许多日常任务要求机器人能够可靠地定位和取回物体。然而,由于严重的遮挡、受限的可达性以及需要避免对场景造成过度改变,这一任务仍然具有挑战性。本文提出了一种多技能操作增强建图方法(Multi-Skill Manipulation-Enhanced Mapping, MS-MEM),这是一个不确定性感知建图的证据框架,集成了主动视点选择、物体推动和抓取。MS-MEM 将场景级度量-语义证据置信估计器与不确定性感知的抓取表示相结合。该表示通过一种新颖的全证据抓取估计器学习得到,能够同时建模抓取可供性与朝向的不确定性。在我们的框架中,候选的感知与操作动作在一个统一的动作选择流程中通过共同的信息增益准则进行评估。对于操作动作,我们进一步引入了附带扰动约束(Collateral Disturbance Constraint, CDC),以抑制对场景置信中高置信区域的过度改变。这使得 MS-MEM 能够选择既能有效降低建图不确定性、又能限制附带场景改变的动作。实验结果表明,与忽略场景扰动的单技能和无约束基线方法相比,MS-MEM 在大幅减少场景扰动的同时实现了更高的建图精度,凸显了主动视点选择、推动与抓取动作之间的协同效应。
cs.RO / 22 / 2609.02542

World-Model-Augmented Visual Locomotion for Humanoids on Foothold-Constrained Terrain

面向落脚点受限地形的世界模型增强人形机器人视觉运动控制
Liu, Yuxi, Han, Lijun, Wang, Ziming, Zhang, Ao, Yang, Cong, Sui, Wei
Abstract
Foothold-constrained terrain is characterized by sparse, discontinuous, or geometrically restricted feasible foot contacts, as encountered on stepping stones, across gaps, and on narrow stair treads. On such terrain, a single misstep often leaves little room to recover, so policies that base foot-placement decisions primarily on the immediately visible terrain are prone to failure. We ask whether a learned predictive summary of near-future observations and rewards can provide the anticipatory information required in such settings. We present World-Model-Augmented Visual Locomotion (WM-LOCO), which jointly trains a recurrent world model and a PPO policy. Conditioned on proprioception and a single onboard depth image, the world model produces a predictive recurrent feature that guides the policy, without explicit foothold labels. In simulation, WM-LOCO succeeds on gaps and stepping stones where a matched baseline fails completely, and matches the baseline's success rate on stairs while improving stride efficiency and reducing pelvis acceleration. We deploy the same policy onboard a physical Unitree G1 humanoid using onboard proprioception and a single depth stream; it traverses all three terrain classes with an average success rate of 93.3%.
Chinese Translation
落脚点受限地形的特征是可行足部接触点稀疏、不连续或受几何限制,例如踏石、跨越间隙以及狭窄楼梯踏面等场景。在此类地形上,一步踏错往往几乎没有恢复余地,因此主要依据当前可见地形做出落脚决策的策略容易失败。我们探讨的问题是:学习得到的对近期未来观测与奖励的预测性摘要,能否为此类场景提供所需的前瞻性信息。我们提出了世界模型增强视觉运动控制(World-Model-Augmented Visual Locomotion, WM-LOCO),该方法联合训练一个循环世界模型和一个PPO策略。世界模型以本体感知和单张机载深度图像为条件,生成预测性循环特征来引导策略,而无需显式的落脚点标签。在仿真中,WM-LOCO能够在匹配基线完全失败的间隙和踏石地形上取得成功,并在楼梯地形上达到与基线相同的成功率,同时提高了步幅效率并降低了骨盆加速度。我们将同一策略部署到真实的Unitree G1人形机器人上,仅使用机载本体感知和单路深度图像流;该策略在全部三类地形上的平均成功率达到93.3%。
cs.RO / 23 / 2609.02546

ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation

ZETA:面向桌面操作的零样本跨具身VLA迁移的受控研究
Yan, Mi, Zhang, Wenhao, Zhang, Zhiqi, Peng, Yu, Wang, Tangxinyu, Zhai, Lingfei, Su, Jiayi, Deng, Shengliang, Peng, Lin, Liu, Yaowei, Chen, Yuxing, Wei, Zhiyuan, Wang, Jilong, Chen, Jiayi, Lyu, Jiangran, Zhang, Zhizheng, Wang, He
Abstract
Zero-shot generalization to unseen embodiments is important for generalizable vision-language-action (VLA) models as robot hardware evolves and task-specific data collection remains costly. However, a systematic understanding of this problem remains limited, in part because the literature lacks a unified zero-shot transfer definition and controlled evaluation settings that isolate embodiment changes from differences in tasks, scenes, or protocols. To address this gap, we first distinguish strict zero-shot transfer, where the target embodiment is absent from all training data, from pretrain-exposed zero-shot transfer, where it appears only during pretraining. We then introduce a controlled benchmark spanning 14 held-out target embodiments across simulation and real-world validation. Within this framework, we conduct a controlled analysis of four factors: state-action representations, pretraining embodiment diversity, auxiliary co-training objectives, and target-embodiment exposure. Experimental results show that local end-effector (EEF) state-action representations, the source embodiment diversity, and auxiliary co-training improve cross-embodiment transfer by around 15, 18, and 7 percentage points, respectively. We further find that adding only 5% target-embodiment data during pretraining improves average target-embodiment progress by 13.4 percentage points, showing that strict and pretrain-exposed zero-shot transfer are distinct and should be reported separately. Together, these findings provide practical guidance for evaluating and improving cross-embodiment VLA transfer in stationary tabletop manipulation with two-finger grippers, while motivating future investigation of broader settings including mobile-base control, dexterous hands, and long-horizon tasks.
Chinese Translation
随着机器人硬件的不断演进以及针对特定任务的数据采集成本依然高昂,对未见具身形态(embodiment)的零样本泛化能力对于可泛化的视觉-语言-动作(VLA)模型至关重要。然而,对这一问题的系统性理解仍然有限,部分原因在于文献缺乏统一的零样本迁移定义,也缺乏能够将具身形态变化与任务、场景或协议差异相隔离的受控评估设置。为填补这一空白,我们首先区分了严格零样本迁移(strict zero-shot transfer,即目标具身形态完全不出现在训练数据中)与预训练暴露零样本迁移(pretrain-exposed zero-shot transfer,即目标具身形态仅在预训练阶段出现)。随后,我们提出了一个受控基准,涵盖14个留出的目标具身形态,并包含仿真与真实世界验证。在该框架内,我们对四个因素进行了受控分析:状态-动作表示、预训练具身形态多样性、辅助协同训练目标以及目标具身形态的暴露程度。实验结果表明,局部末端执行器(EEF)状态-动作表示、源具身形态多样性和辅助协同训练分别使跨具身迁移性能提升约15、18和7个百分点。我们进一步发现,仅在预训练中加入5%的目标具身形态数据,即可使目标具身形态的平均任务进度提升13.4个百分点,这表明严格零样本迁移与预训练暴露零样本迁移是两类不同的问题,应分别进行报告。总之,这些发现为评估和改进固定式桌面双指夹爪操作中的跨具身VLA迁移提供了实用指导,同时也激励未来对更广泛场景的探索,包括移动底盘控制、灵巧手以及长时序任务。
cs.RO / 24 / 2609.02575

Pre-Lane-change Signal in Transitional Autonomous Vehicles: Results from Controlled Experiments

过渡型自动驾驶车辆的换道前信号:受控实验结果
Mu, Zeyu, Chen, Danjue, Sharma, Abhinav, List, George F.
Abstract
This paper investigates how a production transitional autonomous vehicle (tAV) develops and executes mandatory lane-change decisions. Using 150 controlled mandatory lane changes from the NC-tALC experiments, the study examines whether the eventual target gap is observable before lateral movement begins and how the tAV progresses longitudinally from that pre-lane-change state to lane-change start. Signal time (SigT) is defined as an operational pre-lane-change-start reference point. A Firth logistic regression predicts whether the tAV eventually merges in front of or behind its nearest target-lane vehicle using relative position and relative speed at SigT. Longitudinal progression from SigT to lane-change start is then examined separately for in-position and repositioning cases. The traffic state at SigT contains substantial information about eventual target-gap choice and provides meaningful lead time before lateral movement begins. The proposed formulation predicts whether the tAV remains with its current gap or repositions to a neighboring gap by moving forward or dropping back, including cases with longitudinal overlap and ambiguous current-gap geometry. The model achieves an average five-fold cross-validated accuracy of 0.89. Results also provide preliminary evidence that in-position and repositioning cases follow different longitudinal pathways from SigT to lane-change start. These findings support a two-stage conjecture of the observable lane-change process: longitudinal preparation from SigT to lane-change start, followed by lateral maneuver execution. The formulation applies to in-position, repositioning, and longitudinally overlapping cases, and can support lane-change models that distinguish target-gap choice from lateral-onset timing while representing longitudinal preparation before lateral movement begins.
Chinese Translation
本文研究了量产型过渡自动驾驶车辆如何制定并执行强制性换道决策。基于 NC-tALC 实验中的 150 次受控强制换道,本研究考察了最终目标间隙是否可在横向移动开始之前被观测到,以及过渡自动驾驶车辆如何从该换道前状态纵向推进至换道起点。本文将信号时间定义为换道开始前的操作性参考点。研究采用 Firth 逻辑回归,利用信号时间时刻的相对位置和相对速度预测过渡自动驾驶车辆最终是汇入最近目标车道车辆的前方还是后方。随后,针对原位换道与重新定位两种情形,分别考察了从信号时间到换道起点的纵向推进过程。结果表明,信号时间时刻的交通状态包含关于最终目标间隙选择的大量信息,并在横向移动开始之前提供了有意义的前瞻时间。所提出的模型能够预测过渡自动驾驶车辆是保持当前间隙,还是通过前进或后退重新定位至相邻间隙,包括存在纵向重叠和当前间隙几何形状模糊的情形。该模型的五折交叉验证平均准确率达到 0.89。结果还提供了初步证据,表明原位换道与重新定位情形从信号时间到换道起点遵循不同的纵向路径。这些发现支持关于可观测换道过程的两阶段猜想:从信号时间到换道起点的纵向准备阶段,以及随后的横向机动执行阶段。该模型框架适用于原位换道、重新定位以及纵向重叠的情形,能够支持区分目标间隙选择与横向起始时机的换道模型,同时刻画横向移动开始之前的纵向准备过程。
cs.RO / 25 / 2609.02605

Advancing Accessible Underwater Robotics: The Mini-Girona I-AUV at RAMI 2025

推进可获取的水下机器人技术:面向RAMI 2025的Mini-Girona自主干预水下航行器
Hamoda, Taqi, Ahmed, Bilal, Ele-Ojo, Deborah, Bao, Thi Tran Ha, Saidani, Adel, Elgabalawy, Mazen, Chaarani, Alaaeddine, Realpe, Sebastian, Cieslak, Patryk, Ridao, Pere, Palomeras, Narcis, Gracias, Nuno
Abstract
The Mini-Girona Intervention Autonomous Underwater Vehicle (I-AUV) represents an advancement in accessible underwater robotics, designed to bridge the gap between costly, specialized research AUVs and basic Remotely Operated Vehicles (ROVs). Developed with a focus on affordability and usability, the Mini-Girona, priced at approximately $50,000, integrates advanced components such as a 5-DOF manipulator arm, stereo vision, and AI-driven processing for autonomous navigation and intervention tasks. This paper presents the design and development of the Mini-Girona, detailing its performance during the RAMI 2025 student competition. Despite challenges such as thermal management issues and restricted team access, the Mini-Girona achieved second place overall, excelling in vision-based perception and intervention tasks. This work highlights the platform's potential as a tool for underwater robotics research and education, fostering innovation in real-world underwater applications.
Chinese Translation
Mini-Girona干预型自主水下航行器(I-AUV)代表了可获取水下机器人技术的一项进步,旨在弥合昂贵、专业化的科研型AUV与基础型遥控水下机器人(ROV)之间的差距。Mini-Girona以经济性和易用性为核心进行开发,价格约为5万美元,集成了先进组件,如5自由度机械臂、双目视觉以及用于自主导航与干预任务的AI驱动处理系统。本文介绍了Mini-Girona的设计与开发,并详细阐述了其在RAMI 2025学生竞赛中的表现。尽管面临热管理问题和团队操作受限等挑战,Mini-Girona仍获得总分第二名,在基于视觉的感知与干预任务中表现优异。这项工作凸显了该平台作为水下机器人研究与教育工具的潜力,有助于推动真实水下应用场景中的创新。
cs.RO / 26 / 2609.02634

Latent Cluster Analysis for Vision-Language-Action Models

面向视觉-语言-动作模型的潜在聚类分析
Wulff, Theodor, Lanza, Sergio, Bila, Tamara, Cangelosi, Angelo, Wermter, Stefan, Farkas, Igor
Abstract
Vision-Language-Action (VLA) Models are increasingly used in robotics for their ability to ground language and perception into action, yet the internal representations driving their behaviour remain poorly understood. We propose LAVLA, a framework for latent cluster analysis of VLA models, and conduct a layer-wise study of the state-of-the-art GR00T N1.5 model, with particular focus on its action decoder. To better characterise the latent space during action diffusion, we introduce a cross-attention-based embedding-weighting method that amplifies relevant features while suppressing less informative ones. Quantitative evaluation shows that weighted clustering consistently outperforms the baseline. To improve interpretability, we extract human-interpretable concepts for each cluster, linking latent representations to semantic descriptions. Our analysis shows that latent clusters progressively disentangle spatiotemporal and kinematic features, with representations becoming more refined in the middle layers and stabilising toward the output. As such, LAVLA advances the interpretability of language-driven robotic systems.
Chinese Translation
视觉-语言-动作模型因能够将语言和感知落地为动作,在机器人领域得到日益广泛的应用,然而驱动其行为的内部表征仍然缺乏充分的理解。我们提出了LAVLA,一个用于VLA模型潜在聚类分析的框架,并对最先进的GR00T N1.5模型进行了逐层研究,尤其关注其动作解码器。为了更好地刻画动作扩散过程中的潜在空间,我们引入了一种基于交叉注意力的嵌入加权方法,在放大相关特征的同时抑制信息量较低的特征。定量评估表明,加权聚类方法始终优于基线方法。为提升可解释性,我们为每个聚类提取人类可理解的概念,将潜在表征与语义描述相关联。我们的分析表明,潜在聚类逐步解耦了时空特征和运动学特征,表征在中间层变得更加精细化,并在趋近输出时趋于稳定。因此,LAVLA推动了语言驱动机器人系统的可解释性研究。
cs.RO / 27 / 2609.02653

HINT: Human-Intent Inception for Long-Horizon Robot Manipulation

HINT:面向长时程机器人操作的人类意图起始框架
Mei, Mingyu, Xu, Haojie, Jin, Shihao, Dai, Zibo, Cheng, Qihao, Lv, Zhengrui, Fang, Hongjie, Tang, Shirun, Chen, Guang, Zhao, Xinyue, Shen, Huiliang, He, Zaixing
Abstract
Humans can perform complex manipulations given a simple intent through an overall instruction, while continuously adapting to evolving visual observations. However, current vision-language action (VLA) models and other action policies struggle to realize this high-level intelligent behavior under dense, evolving visual inputs and sparse language guidance. Visual correlations can then dominate semantic intent, leading actions to follow visual shortcuts rather than human goals. We present HINT (Human-INTent INcepTion), an agentic framework inspired by the human manipulation principles: semantic intent changes sparsely at manipulation-pattern transitions, whereas continuous control primarily depends on the evolving object-hand relationship. HINT invokes semantic reasoning only at pattern transitions to resolve the current subtask and target, then maintains this commitment through multi-view grounding and visual tracking. We explore two visual interfaces-image-space semantic highlighting and attention-prior injection-to communicate the tracked intent to the action policy without introducing additional trainable parameters into the foundation action model. Experiments across three long-horizon tasks and out-of-distribution variants show that HINT substantially improves intent understanding, task progress, and end-to-end success across two foundation policies while preserving low-latency control.
Chinese Translation
人类仅需通过一条总体指令表达简单意图,即可执行复杂的操作任务,并能持续适应不断变化的视觉观测。然而,当前的视觉-语言-动作(VLA)模型及其他动作策略难以在密集、动态变化的视觉输入和稀疏的语言引导下实现这种高层智能行为。视觉关联性往往会主导语义意图,导致动作跟随视觉捷径而非人类目标。我们提出了HINT(Human-INTent INcepTion,人类意图起始),这是一个受人类操作原则启发的智能体框架:语义意图仅在操作模式转换时稀疏地变化,而连续控制主要依赖于不断演化的物体-手部关系。HINT仅在模式转换时调用语义推理以确定当前的子任务和目标,随后通过多视角定位与视觉跟踪维持这一承诺。我们探索了两种视觉接口——图像空间语义高亮和注意力先验注入——将所跟踪的意图传递给动作策略,且无需在基础动作模型中引入额外的可训练参数。在三项长时程任务及分布外变体上的实验表明,HINT在两种基础策略上显著提升了意图理解、任务进度和端到端成功率,同时保持了低延迟控制。
cs.RO / 28 / 2609.02688

From Proxy Learning to Driving Decisions: A Transfer-Based Framework for Evaluating Future-Aware Autonomous Driving Planners

从代理学习到驾驶决策:一个基于迁移的框架用于评估未来感知的自动驾驶规划器
Wu, Yikai
Abstract
Future-aware representations and world models are increasingly used in proposal-based autonomous-driving planners to improve trajectory selection. However, improvements in proxy objectives or restricted subsets are often interpreted as planning gains without verifying proposal ordering, selected trajectories, full-scale utility, and critical driving components. We propose the Proxy-to-Decision Transfer (PDT) Framework, an analysis framework that evaluates when learned future information supports a reliable driving-performance improvement claim. Its Decision-Transfer Decomposition Module localizes value loss through score margins, switch-conditioned utility, and support-versus-selection regret. Its Reliability-Constrained Validation Module requires exact pairing, a minimum meaningful effect, scale-expanded confirmation, safety non-compensation, sequential comparability, and family-level robustness. On a representative future-aware planner evaluated with NAVSIM-v1, component BCE decreases from 0.705 to 0.530 while held selected PDM decreases from 0.963 to 0.961. A separate candidate improves a 512-record prefix by 0.00909, with a scene-bootstrap 95% interval of [0.000744, 0.0177], but its 2048-record and complete-support intervals include zero. A proposal-level replay further confirms the switch-utility decomposition, yet none of 432 screened configurations passes the two-half, two-seed robustness gate. PDT therefore identifies where decision transfer fails or remains indeterminate across proxy, subset, aggregate, and selection evidence.
Chinese Translation
未来感知表征和世界模型越来越多地被用于基于候选方案(proposal-based)的自动驾驶规划器中,以改进轨迹选择。然而,代理目标或受限子集上的性能提升常常被直接解释为规划能力的提升,而未验证候选方案排序、所选轨迹、全尺度效用以及关键驾驶组件。我们提出了代理到决策迁移(Proxy-to-Decision Transfer, PDT)框架,这是一个分析框架,用于评估学习到的未来信息在何时能够支撑可靠的驾驶性能改进结论。其决策迁移分解模块通过分数差距、基于切换条件的效用以及支持-选择遗憾(support-versus-selection regret)来定位价值损失所在。其可靠性约束验证模块要求精确配对、最小有意义效应、规模扩展确认、安全性不可补偿、序列可比性以及族级鲁棒性。在使用 NAVSIM-v1 评估的一个代表性未来感知规划器上,组件 BCE 从 0.705 降至 0.530,而留出集所选 PDM 从 0.963 仅降至 0.961。另一个候选方法在 512 条记录的前缀上将指标提升 0.00909,其场景自助法(scene-bootstrap)95% 置信区间为 [0.000744, 0.0177],但其 2048 条记录和完整支持集的置信区间均包含零。候选方案级别的重放进一步证实了切换-效用分解,然而在筛选出的 432 个配置中,无一通过两半、两种子的鲁棒性门控。因此,PDT 能够在代理、子集、聚合和选择证据的各个层面识别决策迁移在何处失效或仍不确定。
cs.RO / 29 / 2609.02811

Do Better Imagined Rollouts Mean Better Robot Control? A Controlled Study of World-Model Evaluation Under Feedback

更好的想象式前瞻预测意味着更好的机器人控制吗?反馈条件下世界模型评估的受控研究
Raghavan, Dharini, Singh, Amritpal
Abstract
Predictive models are increasingly used in robotics for state estimation, planning, control, and policy evaluation, yet they are often judged by open-loop prediction accuracy over a fixed horizon. In closed-loop operation, a robot repeatedly acts, receives new measurements, updates its state estimate, and recomputes control. We study this difference in a differential-drive path-tracking task with biased odometry and intermittent landmark sensing. Six state estimators are evaluated across 24 sensing conditions using trajectory replay, a 20-step measurement-free rollout, and closed-loop tracking. Replay position RMSE correlates more strongly with closed-loop cross-track RMSE than rollout error (Spearman rho = 0.923 vs. 0.774) and selects a different estimator from the closed-loop optimum in 5/24 conditions, compared with 18/24 for the rollout metric. We then vary rollout horizon and measurement-update interval. With H=20, rank agreement decreases from rho = 0.916 with measurements at every step to rho = 0.774 with no measurements. A horizon-update grid shows that long prediction horizons remain informative when regular corrections are retained, whereas long rollouts without correction can produce rankings that differ substantially from closed-loop behavior. We also test recurrent estimators trained on longer sensing outages. This improves the EKF-anchored models under combined sensing degradation, reducing GRU-EKF cross-track RMSE from 1.72 m to 1.06 m, but the gain is not consistent across isolated outages or estimator architectures. These results show that predictive-model evaluation in robotics should specify both prediction horizon and measurement-update schedule. For models used in feedback, offline rollouts are most informative when their sensing and correction pattern reflects closed-loop operation. Code is available at https://github.com/rdharini2001/Robot_World_Model
Chinese Translation
预测模型在机器人学中被日益广泛地应用于状态估计、规划、控制和策略评估,但它们通常以固定时域内的开环预测精度来评判。在闭环运行中,机器人反复执行动作、接收新的测量、更新其状态估计并重新计算控制。我们在一个具有偏差里程计和间歇性路标感知的差速驱动路径跟踪任务中研究这一差异。我们在24种感知条件下,使用轨迹重放、20步无测量前瞻预测以及闭环跟踪对六种状态估计器进行评估。重放位置RMSE与闭环横向偏差RMSE的相关性强于前瞻预测误差(Spearman rho = 0.923 对 0.774),并且在5/24的条件下会选择出与闭环最优估计器不同的估计器,而前瞻预测指标在18/24的条件下如此。随后,我们改变前瞻预测时域和测量更新间隔。当H=20时,排序一致性从每步都有测量时的rho = 0.916下降到无测量时的rho = 0.774。时域-更新间隔网格实验表明,在保留定期校正的情况下,长预测时域仍然具有参考价值;而缺乏校正的长前瞻预测可能产生与闭环行为显著不同的排序。我们还测试了在更长感知中断下训练的循环估计器,这在复合感知退化下改进了以EKF为基础的模型,将GRU-EKF的横向偏差RMSE从1.72米降低到1.06米,但该增益在单一中断或不同估计器架构之间并不一致。这些结果表明,机器人学中预测模型的评估应同时明确预测时域和测量更新计划。对于用于反馈的模型,当前瞻预测的感知与校正模式能够反映闭环运行时,其离线预测才最具参考价值。代码可在 https://github.com/rdharini2001/Robot_World_Model 获取。
cs.RO / 30 / 2609.02830

Toward Robust LiDAR Semantic Segmentation for Real-World Deployment: Evaluation under Coarse Labels, Adverse Conditions, and Domain Shifts

面向真实世界部署的鲁棒LiDAR语义分割:粗粒度标签、恶劣条件与领域偏移下的评估
Haidar, Samir Abou, Chariot, Alexandre, Darouich, Mehdi, Joly, Cyril, Deschaud, Jean-Emmanuel
Abstract
LiDAR-based semantic segmentation is a core perception module for autonomous vehicles and mobile robots. Despite the strong performance of recent state-of-the-art methods on standard benchmarks, existing evaluation protocols remain focused on clean, single-domain settings and fine-grained label taxonomies, leaving deployment readiness largely unassessed. Real-world systems must handle safety-critical label semantics, degraded sensing conditions, and cross-domain variability, yet no unified protocol currently addresses all three aspects together. In this paper, we propose a structured evaluation protocol that assesses the deployment readiness of LiDAR semantic segmentation models along three complementary dimensions: (i) coarse-label evaluation aligned with autonomous driving safety priorities, revealing how label granularity affects different methods; (ii) robustness under eight types of LiDAR corruptions designed to emulate real-world atmospheric, geometric, and sensor degradations; and (iii) domain generalization across datasets without adaptation. The evaluation includes inference speed measured on an embedded Jetson AGX Orin platform, directly reflecting deployment constraints. Our results show that fine-grained benchmark rankings do not always reflect safety-relevant performance, that all methods experience substantial degradation under corruptions with architecture-dependent robustness characteristics, and that current domain generalization remains insufficient for reliable deployment. These findings expose concrete gaps between benchmark performance and deployment readiness, and provide a reference protocol for more practically grounded evaluation of LiDAR semantic segmentation.
Chinese Translation
基于LiDAR的语义分割是自动驾驶车辆和移动机器人的核心感知模块。尽管近年来最先进的方法在标准基准上表现出色,但现有的评估协议仍主要聚焦于干净的单域设置和细粒度标签体系,使得部署就绪性在很大程度上未被评估。真实世界系统必须处理安全关键的标签语义、退化的感知条件以及跨域的可变性,然而目前尚无统一协议能够同时涵盖这三个方面。本文提出了一种结构化的评估协议,从三个互补的维度评估LiDAR语义分割模型的部署就绪性:(i) 与自动驾驶安全优先级相一致的粗粒度标签评估,揭示标签粒度如何影响不同方法;(ii) 在八种LiDAR损坏类型下的鲁棒性评估,这些损坏旨在模拟真实世界中的大气、几何和传感器退化;(iii) 无需自适应的跨数据集领域泛化评估。评估还包括在嵌入式Jetson AGX Orin平台上测量的推理速度,直接反映部署约束。我们的结果表明:细粒度基准排名并不总能反映与安全相关的性能;所有方法在损坏条件下均出现显著性能退化,且鲁棒性特征依赖于架构;当前的领域泛化能力仍不足以支持可靠部署。这些发现揭示了基准性能与部署就绪性之间的具体差距,并为更具实用性的LiDAR语义分割评估提供了参考协议。
cs.RO / 31 / 2609.02861

Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework

迈向可信自主机器人:一种基于可解释人工智能的决策框架
Temel, Cagri
Abstract
Autonomous robots powered by deep learning face a fundamental auditability challenge: when incidents occur, investigators cannot reconstruct why the system made specific decisions. This paper presents TRACE (Transparent Reasoning Architecture for Credible Execution), a decision framework that ensures every autonomous action can be traced back to sensor evidence through documented causal chains. The framework organizes decision-making into four auditable layers: Semantic Perception for evidence-grounded entity recognition, Belief Reasoning for probabilistic state estimation with causal graphs, Action Synthesis for constraint-aware planning with counterfactual documentation, and Execution Verification for compliance monitoring. TRACE is model-agnostic yet designed to integrate learning-based perception modules (CNNs, transformers) while preserving decision-level auditability. We evaluate the framework using three objective metrics: Evidence Traceability (sensor-to-decision linkage), Decision Reconstructability (post-hoc analysis capability), and Temporal Continuity (audit trail completeness). Experimental evaluation on warehouse robot navigation demonstrates that TRACE achieves 98.6% evidence traceability, 99.0% temporal continuity, and 98.1% decision reconstructability across 500 simulated decision cycles. Post-hoc methods like LIME provide feature attributions but lack the artifact structure needed for decision-level reconstruction. The framework addresses EU AI Act requirements for high-risk system transparency and contributes to Explainable AI for safety-critical autonomous systems.
Chinese Translation
由深度学习驱动的自主机器人面临一个根本性的可审计性挑战:当事故发生时,调查人员无法重现系统为何做出特定决策。本文提出了TRACE(Transparent Reasoning Architecture for Credible Execution,可信执行透明推理架构),这是一种决策框架,可确保每一项自主行为都能通过有据可查的因果链追溯到传感器证据。该框架将决策过程组织为四个可审计的层级:语义感知(Semantic Perception),用于基于证据的实体识别;信念推理(Belief Reasoning),用于基于因果图的概率状态估计;动作合成(Action Synthesis),用于带有反事实文档记录的约束感知规划;以及执行验证(Execution Verification),用于合规性监控。TRACE与模型无关,同时设计用于集成基于学习的感知模块(如CNN、transformer),同时保持决策层面的可审计性。我们采用三项客观指标对该框架进行评估:证据可追溯性(Evidence Traceability,传感器与决策的关联性)、决策可重现性(Decision Reconstructability,事后分析能力)以及时间连续性(Temporal Continuity,审计轨迹的完整性)。在仓储机器人导航上的实验评估表明,在500个模拟决策周期中,TRACE实现了98.6%的证据可追溯性、99.0%的时间连续性和98.1%的决策可重现性。LIME等事后解释方法能够提供特征归因,但缺乏决策级重现所需的结构化产物。该框架满足了欧盟《人工智能法案》(EU AI Act)对高风险系统透明度的要求,并为面向安全关键型自主系统的可解释人工智能(Explainable AI)做出了贡献。