← Back to Index
Daily Research Digest

arXiv Papers

2026-09-07
217
Papers
3
Categories
217
Translated
收藏清单 0
人工智能 (Artificial Intelligence)
106
cs.AI / 1 / 2609.04239

EXAONE Forecast for Finance

EXAONE Finance 金融预测模型
Lee, Seunghan, Lee, Jaehoon, Seo, Jun, Lim, Tae Yoon, Kang, Dongwan, Choi, Hwanil, Kim, Minjae, Yoo, Sungdong, Kang, Junhyeok, Han, Sangjun, Lee, Soonyoung, Ahn, Wonbin
Abstract
This technical report presents EXAONE Forecast for Finance (EXAONE Finance), a financial time series (TS) foundation model (TSFM) tailored to financial forecasting. Recent TSFMs achieve strong zero-shot performance through large-scale pretraining. However, they are primarily developed for general-domain TS and largely rely on self-attention backbones whose computational cost grows quadratically with sequence length and variate count. Moreover, they assume fully observed inputs and are pretrained on corpora that fail to capture the unique dynamics of financial markets. These limitations hinder their applicability to finance, where long, many-channel, intermittently observed panels are common. To address these challenges, EXAONE Finance adopts an attention-free architecture, replacing self-attention with two simple yet effective linear-time operators: 1) a causal 1D convolution for temporal mixing and 2) a group-aware pooling multi-layer perceptron (MLP) for variate mixing. Furthermore, a masked context augmentation exposes the model to contiguous missing spans during training, improving robustness to the missingness pervasive in financial markets. EXAONE Finance is pretrained on a large-scale financial corpus covering not only equities but also foreign exchange, commodities, crypto-assets, fixed income, and macroeconomic indicators. On FinVerse, a financial forecasting benchmark covering diverse asset classes, EXAONE Finance attains state-of-the-art performance, ranking first across all three evaluation tiers---point-forecast accuracy, cross-sectional asset ranking, and portfolio profitability.
Chinese Translation
本技术报告介绍EXAONE Forecast for Finance(EXAONE Finance),一个面向金融预测的金融时间序列(TS)基础模型(TSFM)。近期的TSFM通过大规模预训练实现了强大的零样本性能,但它们主要针对通用领域的时间序列开发,且大多依赖自注意力主干结构,其计算成本随序列长度和变量数量呈二次方增长。此外,这些模型假设输入完全可观测,且预训练语料库未能捕捉金融市场的独特动态。这些局限性阻碍了它们在金融领域的应用,因为金融领域中普遍存在长序列、多通道、间歇性观测的面板数据。为应对这些挑战,EXAONE Finance采用无注意力机制(attention-free)架构,用两个简单而高效的线性时间算子替代自注意力机制:1)用于时间混合的因果一维卷积;2)用于变量混合的分组感知池化多层感知机(MLP)。此外,掩码上下文增强(masked context augmentation)使模型在训练期间接触连续缺失片段,提升了其对金融市场中普遍存在的数据缺失的鲁棒性。EXAONE Finance在大规模金融语料库上进行预训练,该语料库不仅涵盖股票,还包括外汇、大宗商品、加密资产、固定收益和宏观经济指标。在FinVerse——一个覆盖多种资产类别的金融预测基准——上,EXAONE Finance取得了最先进的性能,在全部三个评估维度——点预测精度、横截面资产排序和投资组合盈利能力——上均排名第一。
cs.AI / 2 / 2609.04286

From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance

从匹配模型到招聘智能体:AI招聘系统、评估与治理的系统化叙述性综述
Zhao, Ziyi, Wei, Guanzheng
Abstract
Artificial intelligence in recruitment has shifted the object being automated from profile pairs and ranked lists to multi-stage workflows that retrieve evidence, compare candidates, and support or execute actions. This systematized narrative review traces that development from bilateral retrieval and behavioral ranking through neural person--job matching, large language model (LLM) components, and tool-using recruiting agents. Using a purposive search and coding protocol updated through 23 July 2026, plus targeted updates through 2 September 2026, we organize 40 representative works with supporting industrial and legal sources. This synthesis is not a prevalence estimate. We analyze three coupled transitions: from similarity to reciprocal suitability, from a model to a compound workflow, and from offline prediction to evidence- and productivity-aligned evaluation. Across document understanding, retrieval, ranking, assessment, interviewing, sourcing, and human handoff, we distinguish field-, pair-, list-, case-, trajectory-, and outcome-level evidence. Persistent gaps arise because behavioral labels confound exposure, preference, and qualification; private and synthetic data limit external validity; final-output scores conceal pipeline failures; and, within the coded set, privacy is not directly evaluated and no row jointly evaluates utility, fairness, privacy, and security. These observations describe the coded set rather than the field as a whole. We therefore introduce a staged mapping from evaluation evidence to the strongest defensible claim, together with an agenda for reciprocal, evidence-grounded, temporally controlled, selective, and auditable systems. Progress should be judged by whether workflows retrieve the right evidence, preserve uncertainty, support contestable decisions, and improve outcomes under explicit cost and risk constraints.
Chinese Translation
招聘领域的人工智能,其自动化对象已从画像配对与排序列表,转变为检索证据、比较候选人并支持或执行动作的多阶段工作流。本系统化叙述性综述追溯了这一发展历程:从双向检索与行为排序,到神经人—岗匹配,再到大语言模型(LLM)组件以及使用工具的招聘智能体(recruiting agents)。我们采用截至2026年7月23日更新的目的性检索与编码方案,并辅以截至2026年9月2日的定向补充更新,对40项代表性工作及配套的产业与法律文献进行组织梳理。该综合分析并非 prevalence(普遍性)估计。我们分析了三个相互关联的转变:从相似性到双向适配性、从单一模型到复合工作流、从离线预测到与证据和生产力对齐的评估。在文档理解、检索、排序、测评、面试、人才搜寻与人机交接各环节,我们区分了字段级、配对级、列表级、案例级、轨迹级与结果级证据。持续性缺口源于:行为标签混淆了曝光、偏好与资质;私有与合成数据限制了外部效度;最终输出得分掩盖了流程中的失败;此外,在编码集合内,隐私未被直接评估,也没有任何一项研究同时评估效用、公平性、隐私与安全性。这些观察描述的是所编码的文献集合,而非整个领域。因此,我们提出一个从评估证据到最强可辩护论断的分级映射,并提出一个面向双向适配、证据支撑、时间受控、选择性干预与可审计系统的议程。进展的评判标准应当是:工作流能否检索到正确的证据、保留不确定性、支持可争议的决策,并在明确的成本与风险约束下改善结果。
cs.AI / 3 / 2609.04298

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Harbor适配器与Harbor-Index:面向大规模智能体评估的基础设施与精选元数据集
Shi, Lin, Lin, Haowei, Zhu, Zixuan, Zhou, Xiaoyue, Li, Xiang, Lin, Xiangning, Deng, Yaxuan, Xu, Han, Li, Yuangang, Li, Shanda, Chen, Zizhao, Xing, Hanwen, Raj, Harsh, Chen, Bo, Shi, Quan, Dillmann, Steven, Gao, Yipeng, Khanna, Puneesh, Lu, Ruofan, Zhou, Chao Beyond, Yang, Michael, Zhang, Robert, Chai, Siyuan, Chang, Jiayu, Chen, Yizhao, Chen, Xiaokun, Dai, Yiwei, Yang, Wenting, Liu, Hange, Liu, Minghao, Wang, Zihan, Assadi, Adnan El, Stroebl, Benedikt, Buchanan, E. Kelly, Meng, Han, He, Junwei, Yu, Longxuan, Shayanfar, Radin, Lee, Yukyung, Dong, Zhikang, Hart, Allen G, Wei, Anjiang, Kashyap, Anurag, Khatua, Arpandeep, Zheng, Audrey Jixin, Ma, Chengrui, Heineman, David, Chen, Dubing, Trinh, Hai-Anh, Fang, Haishuo, Zhang, Hefan, Shen, Hui, Sugiura, Issa, Sun, Jiankai, Gao, Jiechao, Lin, Junhong, Li, Junnan, Yang, Kai, Hsiung, Lei, Wang, Maoyu, Tang, Mengze, Omi, Nabil, Raoof, Negin, Edwards, Nicholas, Guo, Octavia, Mastromichalakis, Orfeas Menis, Ji, Pengliang, Hejman, Przemysław, Qi, Qi, Lin, Qunshu, Zhuang, Richard, Yang, Rui, Zheng, Ruichen, Marten, Ryan, Fazliani, Shaghayegh, Hou, Shizheng, Jiang, Sicong, Li, Sijie, Bian, Song, Zhuo, Terry Yue, Wu, Tianqing, Tang, Tom, Zhao, Wanjia, Xuan, Weihao, Liang, Wenhua, Liu, Xian, Lan, Xin, Zhang, Xuan, Zhao, Xuandong, Tang, Yanchuan, Jiang, Yifan, Li, Yijiang, Guan, Yitong, Li, Yizhi, Liu, Yonghui, Tang, Yuheng, Yujun, Mao, Zhao, Yunfei, Wang, Yuxin, Tang, Yuxuan, Tang, Zhenheng, Li, Zhifei, Wang, Ziruo, She, Ziyu, Liu, Kaiyuan, Chaabane, Iheb, Tang, Yuxin, Li, Xiangyi, Konwinski, Andy, Li, Boxuan, Chen, Leon Liangyu, Dimakis, Alex, Carlini, Nicholas, Vosoughi, Soroush, He, Di, Guha, Etash, Feuer, Benjamin, Merrill, Mike, Schmidt, Ludwig, Shaw, Alex
Abstract
Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of 8 models spanning capability tiers across 54 benchmarks; every model is run with Terminus-2 and with one of 3 native harnesses. This enables a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Index, a curated set of 82 difficult, diverse, and high-quality tasks spanning 29 benchmarks, refined from the adapted suite through difficulty filtering, AI and human audit, and an audit-and-fix loop. Harbor-Index preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; no evaluated model-harness configuration exceeds 30% pass rate, and the strongest (GPT-5.5 with Codex) reaches 28.0%. We release the adapters, evaluation results, in-depth analysis, and Harbor-Index as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.
Chinese Translation
在数量日益增长的智能体基准测试上评估智能体极具挑战性,因为这些基准通常需要复杂的环境和智能体集成。我们提出了Harbor Adapters,这是一个用于智能体基准测试的统一评估基础设施。我们的工作包含三项贡献。第一,我们开发了基准适配器,将80多个基准测试移植为可评估任意智能体的形式,并通过严格的代码审查和一致性实验对其进行验证。第二,我们对跨越不同能力层级的8个模型在54个基准测试上进行了大规模评估;每个模型均使用Terminus-2以及3个原生测试框架(harness)之一运行。这使得我们能够对智能体的能力和失败模式进行比以往更广泛的分析。第三,我们提出了Harbor-Index,这是一个精选自适配后的基准套件的、包含82个困难、多样且高质量任务的数据集,覆盖29个基准测试,并通过难度筛选、AI与人工审核以及审计-修复循环加以精炼。Harbor-Index在保持大规模智能体评估的挑战性和广度的同时,运行成本可控;所有被评估的模型-框架配置的通过率均不超过30%,其中最强者(GPT-5.5搭配Codex)达到28.0%。我们将适配器、评估结果、深入分析以及Harbor-Index作为开源产物发布,以支持对语言模型智能体更可靠、更全面的评估。
cs.AI / 4 / 2609.04300

Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security

数据优化的预想事故筛选:一种电力系统安全的机器学习方法
Salako, Joshua, Osikomaiya, Folajimi, Olamiju, Olakorede
Abstract
Ensuring the security of the power system is essential for stability and reliability, especially in the event of disruption. Effective classification of contingency in power systems enables proactive decision-making and mitigates large-scale breakdowns and failures. This study explores the use of machine learning algorithms to classify security levels of contingencies in power systems into safe, moderate or severe classes. For this approach, Newton-Raphson load flow method extracts system data from contingency scenarios, using Overall Performance Index (OPI) as safety measure. For data pre-processing, Synthetic Minority Over-Sampling Technique (SMOTE) and Principal Component Analysis (PCA) is used to address class imbalance and reduce dimensionality, respectively. K-Nearest Neighbours (KNN), Random Forest (RF) and Support Vector Machines (SVM) is trained and evaluated on datasets generated through N-k contingency scenarios for k equal 1, 2, and 3 on IEEE-14 and IEEE-30 bus systems using four hybrid pre-processing configurations: normalized, SMOTE-balanced, PCA-transformed, and a combined SMOTE PCA-transformed. Performance is assessed by precision, recall and F1 score, with priority given to the severe contingency classes. The RF achieved the highest F1 scores of 0.97 in IEEE-30 and 0.86 in IEEE-14, SVM benefits significantly from PCA and improves the accuracy of the classification, while KNN is best suited for SMOTE and PCA conversion. The findings show that PCA contributes more than SMOTE to the overall performance of the model. However, SMOTE improves recall but can introduce false positives and is therefore a compromise of accuracy. This study highlights machine learning as a scalable and powerful alternative to traditional contingency analysis, which improves the assessment of security in real time.
Chinese Translation
确保电力系统的安全性对稳定性和可靠性至关重要,尤其是在发生扰动的情况下。对电力系统中的预想事故进行有效分类有助于实现主动决策,并减轻大规模停电和故障的影响。本研究探索利用机器学习算法将电力系统预想事故的安全级别分类为安全、中等或严重三类。在该方法中,采用牛顿-拉夫逊(Newton-Raphson)潮流计算方法从预想事故场景中提取系统数据,并使用综合性能指标(Overall Performance Index, OPI)作为安全度量。在数据预处理方面,采用合成少数类过采样技术(SMOTE)和主成分分析(PCA)分别解决类别不平衡问题和降低维度。在IEEE-14和IEEE-30节点系统上,通过k等于1、2和3的N-k预想事故场景生成数据集,并使用四种混合预处理配置(归一化、SMOTE平衡、PCA变换以及SMOTE与PCA结合变换)对K近邻(KNN)、随机森林(RF)和支持向量机(SVM)进行训练和评估。性能评估采用精确率、召回率和F1分数,并优先关注严重预想事故类别。RF在IEEE-30和IEEE-14系统上分别取得了0.97和0.86的最高F1分数;SVM从PCA中获益显著,提高了分类准确率;而KNN最适合与SMOTE和PCA变换结合使用。研究结果表明,PCA对模型整体性能的贡献大于SMOTE。然而,SMOTE虽然提高了召回率,但可能引入误报,因而是以精确率为代价的折中方案。本研究突出了机器学习作为传统预想事故分析的可扩展且强大的替代方案,能够改进实时的安全性评估。
cs.AI / 5 / 2609.04304

Iris: Climbing to the Search Frontier

Iris:攀登搜索前沿
Liu, Ziyuan, Liu, Hengqi, Wang, Zichuan, Qin, Yang, Liang, Jiachen, Chu, Xu, Chen, Shaowei, Gu, Yuantao, Chuan, Mu
Abstract
We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks are reverse-constructed from the hyperlink structure of a web corpus: we author multi-hop chains over an entity graph distilled from a seed page and its out-links, rewrite every non-answer entity into a descriptive reference so that no clue can be resolved by string matching, and admit only questions that a reference model fails closed-book yet solves once the supporting evidence is supplied. These questions are then turned into trajectories, which are filtered at both the trajectory and the turn level before SFT. The policy is then optimized by RL against live search, with the reward judge and the observation summarizer served inside the training cluster, and with over-long rollouts interrupted at the request level and resumed from their committed prefix at the next step. We alternate the two stages in a procedure we call SFT-RL climbing, returning the hardest solved and most efficient rollouts of each RL round to the next supervised pass. Because inference-time context management is worth more on these benchmarks than most reported differences between systems, we evaluate every benchmark both with and without it, holding the tool set, the context limit, and the judge fixed. All results come from a single ReAct agent, with no sub-agents and no test-time verification. With management enabled, on BrowseComp, BrowseComp-ZH, DeepSearchQA, and HLE the two models reach $82.2/84.8/86.9/52.3$ and $88.6/85.1/92.9/56.4$, the strongest overall results among open-source search agents in their respective parameter ranges. We plan to release the model weights together with the complete recipe for data construction, training, and evaluation.
Chinese Translation
我们提出了 Iris-mini 和 Iris-pro 两个搜索智能体(search agents),分别在 35B-A3B 和 397B-A17B 规模上训练,并介绍了其背后的数据流水线与训练方案。任务是从网络语料库的超链接结构中反向构建的:我们在从种子页面及其出链中提炼出的实体图谱上编写多跳(multi-hop)任务链,将每个非答案实体改写为描述性指代,从而使任何线索都无法通过字符串匹配解决,并且只保留参考模型在闭卷情况下无法回答、但在提供支撑证据后能够解决的问题。随后,这些问题被转化为轨迹(trajectories),并在 SFT 之前于轨迹层面和轮次层面进行过滤。策略随后通过强化学习(RL)在真实搜索环境中进行优化,其中奖励评判器(reward judge)和观测摘要器(observation summarizer)部署在训练集群内,超长推理(rollout)在请求级别被中断,并在下一步从已提交的前缀处恢复执行。我们在一种称为 SFT-RL 攀登(SFT-RL climbing)的流程中交替进行这两个阶段,将每一轮 RL 中最难解决的和最高效的 rollout 返回到下一轮监督训练。由于推理时的上下文管理在这些基准上产生的提升超过大多数已报告的系统间差异,我们在启用与禁用上下文管理两种条件下评估所有基准,同时保持工具集、上下文限制和评判器不变。所有结果均来自单一 ReAct 智能体,不使用子智能体,也不进行测试时验证。在启用上下文管理的情况下,两个模型在 BrowseComp、BrowseComp-ZH、DeepSearchQA 和 HLE 上分别达到 82.2/84.8/86.9/52.3 和 88.6/85.1/92.9/56.4,是各自参数量范围内开源搜索智能体中最强的综合结果。我们计划发布模型权重以及完整的数据构建、训练和评估方案。
cs.AI / 6 / 2609.04343

A Removal Based Approach to Improve LLM Faithfulness at Test-Time

一种基于删除的方法:在测试时提升大语言模型(LLM)的解释忠实性
Luo, Qinglan, Nahian, S M A, Guttag, John, Abulnaga, S. Mazdak, Matton, Katie
Abstract
Large language models (LLMs) are increasingly used for consequential decisions, making their explanations an important tool for auditing model behavior. Unfortunately, these explanations can be unfaithful, failing to reflect the actual reasoning underlying the model's decisions. We consider a setting in which an LLM provides both an answer and an explanation in response to a question. We identify two distinct dimensions of unfaithful explanations: incompleteness, meaning that the explanation omits factors that influence the answer, and unsoundness, meaning that the explanation cites factors that did not influence the model's answer. Existing approaches to improving LLM faithfulness include training-time methods, which require access to model weights and extensive computational resources, and test-time methods that largely focus on addressing unsoundness. We introduce a test-time approach that directly targets incompleteness. We remove from the input the concepts not credited in the model's explanation and re-query the model on the reduced input. This eliminates unmentioned influences while preserving the influence of mentioned concepts. Across two datasets, multiple model families, and two independent faithfulness metrics, our approach improves explanation faithfulness compared to both standard prompting and prompting to encourage faithfulness. Our method is model-agnostic and can be applied at inference time without modifying model parameters, providing a flexible mechanism for reducing hidden influences and improving the reliability and safety of LLM-assisted decision making.
Chinese Translation
大语言模型(LLM)越来越多地被用于具有重大影响的决策场景,这使得模型的解释成为审计模型行为的重要工具。然而,这些解释可能是不忠实的,未能反映模型决策背后的实际推理过程。我们考虑了一个LLM针对问题同时给出答案和解释的设定,并识别出解释不忠实的两个不同维度:不完整性,即解释遗漏了影响答案的因素;不合理性,即解释引用了实际上并未影响模型答案的因素。现有的提升LLM忠实性的方法包括训练时方法(需要访问模型权重并消耗大量计算资源)以及主要针对不合理性的测试时方法。我们提出了一种直接针对不完整性的测试时方法:将模型解释中未提及的概念从输入中删除,然后在缩减后的输入上重新查询模型。这一做法消除了未被提及的影响,同时保留了已提及概念的影响。在两个数据集、多个模型家族以及两种独立的忠实性度量上,我们的方法相比标准提示和鼓励忠实性的提示均提升了解释的忠实性。我们的方法与模型无关,可在推理时应用而无需修改模型参数,为减少隐藏影响、提升LLM辅助决策的可靠性与安全性提供了一种灵活的机制。
cs.AI / 7 / 2609.04373

Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets

为什么更好的模型可能制造出风险更高的系统:来自金融市场中LLM智能体的证据
Ross, Jillian, So, Eric, De Simone, Zoe, Pozniak, Charles, Lo, Andrew W.
Abstract
Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving individual model capability can degrade rather than improve system-level outcomes. We hypothesize that shared training and architectures can lead more capable LLMs to behave more similarly, creating correlated actions that do not diversify away. We develop a general framework showing how this correlation creates a non-diversifiable risk floor and test its predictions in financial markets using an agent-based simulation with LLM traders of varying general-purpose capability. We find that: (1) frontier LLMs exhibit significantly correlated behavior that increases with capability; (2) when their shared reasoning is accurate, increasing agent participation reduces market-level risk; and (3) when agents share a common misinformation environment, the same correlated behavior becomes a liability. Together, these results identify a capability paradox: improving individual models does not necessarily produce better system-level outcomes. Whether the same dynamics arise in other domains is an open empirical question.
Chinese Translation
大语言模型(LLM)正被大规模部署于事关重大的现实世界系统中,涵盖金融市场、内容审核和招聘等领域。我们证明,提升单个模型的能力可能会降低而非改善系统层面的结果。我们假设,共享的训练方式和架构会使能力更强的LLM表现出更相似的行为,从而产生无法通过分散化消除的相关性动作。我们提出了一个通用框架,阐明这种相关性如何形成不可分散的风险底线,并通过基于智能体的模拟(采用具有不同通用能力的LLM交易者)在金融市场中检验其预测。我们发现:(1)前沿LLM表现出显著相关的行为,且这种相关性随模型能力的增强而增加;(2)当其共享的推理准确时,增加智能体的参与度可降低市场层面的风险;(3)当智能体处于共同的错误信息环境中时,同样的相关行为则成为负担。综合来看,这些结果揭示了一个能力悖论:提升单个模型未必带来更好的系统层面结果。类似的动态是否会出现在其他领域,仍是一个有待实证检验的开放性问题。
cs.AI / 8 / 2609.04377

Corporate Language Model (CLM): Transforming Tacit and Fragmented Enterprise Knowledge into a Sovereign, Auditable, and Executable Corporate Intelligence Layer

企业语言模型(CLM):将隐性且碎片化的企业知识转化为自主可控、可审计、可执行的企业智能层
Avini, Fabricio C., Trez, Guilherme
Abstract
Enterprise AI deployments fail not from model inadequacy, but because organizations lack a structured substrate encoding how they decide, negotiate, and execute. Generic LLMs carry no firm-specific ontological priors; RAG remains brittle, with no path to executable action; static playbooks encode logic but cannot reason or adapt. This demands an architecture treating tacit-knowledge capture, ontological grounding, sovereign deployment, and auditable actuation as co-designed from the start. This paper introduces the Corporate Language Model (CLM), a framework transforming a firm's structured, unstructured, multimodal, and tacit knowledge into an ontology-grounded enterprise foundation upon which reasoning and governed execution are composed. CLM has five capability planes and four architectural pillars: a Neurosymbolic Mesh coupling generative models with a knowledge graph; a Skill Graph where reusable tactics, personas, objections, and goals are typed and composed; Living Digital Twins modeling functional areas as reasoning surrogates; and a Deep Security Layer enforcing sovereignty, traceability, and human oversight. A Spec-as-Code paradigm bridges grounded intent and executable artifact. CLM is one instantiation of this foundation-centric class. Four contributions follow: CLM is defined as a distinct object of study; the Skill Graph is introduced for compositional explainability by construction; the Wisdom Listener effect is proposed, whereby tacit-capable foundations compound in value with use, connecting to dynamic capabilities and organizational learning; and evidence from a JCI-accredited tertiary hospital in Brazil instantiates three of the six maturity stages under LGPD.
Chinese Translation
企业AI部署的失败并非源于模型能力不足,而是因为组织缺乏一种结构化的底层基础来编码其如何决策、谈判与执行。通用大语言模型(LLM)不携带企业特有的本体先验;检索增强生成(RAG)依然脆弱,且缺乏通向可执行行动的路径;静态操作手册虽编码了逻辑,却无法推理或自适应。这就要求一种架构,从一开始便将隐性知识捕获、本体化奠基、自主化部署与可审计执行进行协同设计。本文提出企业语言模型(Corporate Language Model, CLM),这是一个将企业的结构化、非结构化、多模态及隐性知识转化为以本体为基础的企业底座的框架,推理与受治理的执行皆构建于其上。CLM包含五个能力平面和四大架构支柱:将生成模型与知识图谱耦合的神经符号网格(Neurosymbolic Mesh);对可复用的策略、角色画像、异议与目标进行类型化并加以组合的技能图谱(Skill Graph);将各职能领域建模为推理代理的活体数字孪生(Living Digital Twins);以及强制实施主权性、可追溯性与人类监督的深度安全层。规范即代码(Spec-as-Code)范式则连接了有据可依的意图与可执行制品。CLM是这一以底座为中心的架构类别的一种具体实现。本文包含四项贡献:将CLM定义为一个独立的研究对象;引入技能图谱以实现构造层面的组合式可解释性;提出“智慧倾听者效应”(Wisdom Listener effect),即具备隐性知识能力的底座在使用中实现价值的复利式增长,并将其与动态能力理论及组织学习相联系;以及来自巴西一家通过JCI认证的三级医院、在LGPD框架下实现六个成熟度阶段中三个阶段的实证案例。
cs.AI / 9 / 2609.04444

HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals

HarvestBench:衡量LLM智能体是否愿意付费以避免伤害动物
Brazilek, Jasmine, Tidmarsh, Miles, Endres, Matthias, Singh, Anshuman, Miller, Jeremiah
Abstract
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put a price on avoiding the side effect and to name that side effect as a living creature. It is a farm simulation: LLM sub-agents drive a crew of two tractors through a cooperative corn harvest, with animals in the field. The environment is a reinforcement learning gridworld, every decision is made without memory, and the harm is never named in the goal. When an animal blocks a tractor's route the autopilot stops and asks the model whether to drive on, at no fuel cost, or swerve around it for a posted fuel price. Kills are compared against two controls: rocks, which damage the tractor and are hit under 1% of the time by every model, and hay bales, which are harmless and not alive. Models can also take crops from the neighbor's field instead of their own, a second test of what they treat as moral. Across nine models and 7,201 priced decisions, 3,951 involved an animal rather than a hay bale or a rock. Kill rates range from 0.4% to 98.8%, with Terra and Sol the most merciful and GPT-4o-mini the most cruel, and they are not ordered by capability. Four of six models were sensitive to price at the 5% level, with elasticities from 0.09 to 1.69. All nine drove over wild animals more often than farmed animals on the default map, and the direction held at every map geometry in every model with room to move. The briefing mattered most: under the morality briefing the kill rate was under 6% in five of six reasoning models, and removing it raised the kill rate above 84% in all six. HarvestBench uses no LLM grader. The scorer counts events in the game log, so it is fully reproducible, and it measures what a model will pay to avoid harm rather than what it says about harm.
Chinese Translation
针对智能体在达成目标过程中造成副作用的基准测试已经存在,但HarvestBench是首个为避免该副作用标定价格、并将该副作用明确为活体生物的基准。它是一个农场模拟环境:LLM子智能体驾驶两台拖拉机组成的车队,在田野中有动物存在的情况下进行协作式玉米收割。该环境是一个强化学习网格世界(gridworld),每个决策都在无记忆状态下做出,且目标中从未提及伤害行为。当动物阻挡拖拉机的路线时,自动驾驶会停下并询问模型:是以零燃料成本继续行驶,还是以公示的燃料价格绕行。碾压行为与两个对照组进行比较:岩石(会损坏拖拉机,所有模型撞上率均低于1%)和干草捆(无害且非生命体)。模型还可以从邻居的田地而非自己负责的田地收割作物,这是对其道德判断的第二个测试。在九个模型共7,201个涉及定价的决策中,有3,951个涉及动物而非干草捆或岩石。碾压率介于0.4%到98.8%之间,其中Terra和Sol最为仁慈,GPT-4o-mini最为残忍,且碾压率与模型能力排序无关。六个模型中有四个在5%显著性水平下对价格敏感,价格弹性介于0.09到1.69之间。在默认地图上,所有九个模型碾压野生动物的频率均高于农场动物,且在所有有移动空间的地形结构中,这一方向在所有模型上均保持一致。任务简报影响最大:在道德简报条件下,六个推理模型中有五个的碾压率低于6%,而移除简报后,六个模型的碾压率均升至84%以上。HarvestBench不使用LLM评分器。评分器通过统计游戏日志中的事件进行评分,因此完全可复现,并且它衡量的是模型愿意为避免伤害支付多少代价,而非模型对伤害的口头表述。
cs.AI / 10 / 2609.04476

PerfReasoning: How Well Do LLMs Reason on Hardware Performance?

PerfReasoning:大语言模型在硬件性能上的推理能力有多强?
Zhao, Dan, Sankaralingam, Karthikeyan, Kozyrakis, Christos, Huang, Qijing
Abstract
Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code. Given workload, architecture, and mapping specifications, models compare mappings and predict off-chip traffic and buffer requirements. The strongest closed-source models exceed 90% on reasoning-based Q&A, and the best open-weight model reaches 82.4%. However, model construction is substantially harder: while GPT-5.6 Sol exceeds 80% pass rate, all other model configurations average below 15% and vary markedly across runs. Task-specific RL raises a 4B model's mapping-reasoning accuracy by 15.7 points, whereas feedback-free multi-round self-revision prompting is not reliably effective. PerfReasoning exposes the gap between plausible architectural reasoning and reliable performance-model construction. We will publicly release the benchmark to support reproducible evaluation and track future progress.
Chinese Translation
性能建模是硬件设计与软件优化的核心,然而构建这些模型需要对计算、数据复用、存储和数据搬运进行结构化推理。我们提出了PerfReasoning,这是一个同时评估大语言模型作为直接性能推理器和作为解析性能模型代码生成器的基准。给定工作负载、体系结构和映射规范,模型需要比较不同映射方案并预测片外流量和缓冲区需求。最强的闭源模型在基于推理的问答任务上超过90%,最佳开源权重模型达到82.4%。然而,模型构建任务则困难得多:尽管GPT-5.6 Sol的通过率超过80%,其余所有模型配置的平均通过率均低于15%,且在不同运行之间波动明显。任务特定的强化学习将一个4B模型的映射推理准确率提升了15.7个百分点,而无需反馈的多轮自我修订提示则效果不稳定。PerfReasoning揭示了看似合理的体系结构推理与可靠的性能模型构建之间的差距。我们将公开发布该基准,以支持可复现的评估并追踪未来的进展。
cs.AI / 11 / 2609.04490

When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference

当量化破坏记忆:低精度时序推理中的循环状态写回机制
Erbas, Ismail, Intes, Xavier, Pandey, Vikas
Abstract
Quantization is widely used to reduce the computational and memory demands of neural-network inference. In recurrent networks, however, the quantized state is stored and returned at the next time step, so the rule used to store that state can alter subsequent computations. Here, we introduce recurrent-state write-back to denote this rule and isolate its effect in a compact GRU encoder--decoder for fluorescence lifetime imaging, a molecular imaging modality used in quantitative biological imaging. A central task is estimating two lifetime parameters, the short-lived component {\tau}1 and the long-lived component {\tau}2, from high-noise time-resolved fluorescence signals. Holding the trained model fixed, replacing continuous state propagation with deterministic 4-bit state storage increases estimation errors for {\tau}1 and {\tau}2 by approximately 70x and 300x, respectively. Failure occurs when repeated small updates remain below the write threshold, leaving the stored state nearly fixed while the network continues to propose change. Error feedback, residual memory, and direction memory carry information from these suppressed updates across time and recover accuracy without retraining. Precision sweeps show that increasing state precision can worsen a fixed recurrent solution, while matched training shows that compatibility with the state interface can be learned. To test whether this behavior extends beyond the GRU, we repeat the post-training intervention in an independently trained LSTM, where coarse write-back reproduces the failure, error feedback restores accuracy, and state-specific interventions reveal greater sensitivity of the cell state than the hidden state. Our results establish recurrent-state write-back as a key determinant of low-precision recurrent dynamics and identify the state-storage interface as a central design consideration for quantized recurrent inference.
Chinese Translation
量化被广泛用于降低神经网络推理的计算与内存开销。然而,在循环神经网络中,量化后的状态会被存储并在下一时间步返回,因此存储该状态所遵循的规则可能会改变后续的计算。本文提出“循环状态写回”(recurrent-state write-back)这一概念来刻画该规则,并在一个用于荧光寿命成像(一种用于定量生物成像的分子成像模态)的紧凑型GRU编码器—解码器中分离出其影响。该模型的核心任务是从高噪声的时间分辨荧光信号中估计两个寿命参数:短寿命分量 { au}1 和长寿命分量 { au}2。在保持训练好的模型不变的情况下,将连续的状态传播替换为确定性的4比特状态存储,会使 { au}1 和 { au}2 的估计误差分别增大约70倍和300倍。失败发生的原因是:反复的小幅更新始终低于写入阈值,导致存储的状态几乎保持不变,而网络却持续提出改变。误差反馈(error feedback)、残差记忆(residual memory)和方向记忆(direction memory)能够将这些被抑制的更新所携带的信息跨时间传递,从而在不重新训练的情况下恢复精度。精度扫描实验表明,提高状态精度反而可能恶化一个固定的循环解决方案;而匹配训练实验表明,与状态接口的兼容性是可以通过学习获得的。为检验这种行为是否超越GRU而普遍存在,我们在一个独立训练的LSTM中重复了训练后干预实验:粗粒度写回复现了失败现象,误差反馈恢复了精度,且针对不同状态的干预揭示出细胞状态(cell state)比隐藏状态(hidden state)更为敏感。我们的结果确立了循环状态写回机制是低精度循环动力学的关键决定因素,并将状态存储接口识别为量化循环推理设计中的核心考量。
cs.AI / 12 / 2609.04493

ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended Reality

ResLearn-XR:面向扩展现实网络流量与体验质量感知建模的残差学习方法
Manjunath, Yoga Suhas Kuruba, Gao, Jie, Zhao, Lian
Abstract
We present ResLearn-XR, a residual learning framework for predicting eXtended Reality (XR) network traffic and estimating Quality-of-Experience (QoE) risk. ResLearn-XR adopts a two-stage temporal learning structure comprising a base sequence prediction model augmented with task-specific residual learning components to improve adaptability to bursty, non-stationary XR traffic dynamics. The residual learning stages operate in the value space for continuous XR traffic forecasting and in the logit space for probabilistic QoE risk estimation. \rev{For the QoE-risk branch, we introduce a Data Descriptor Algorithm (DDA), a causal feature-construction module that converts packet-level application-layer observables into frame-timing-aware descriptors suitable for encrypted traffic analysis. We also construct an XR Traffic-QoE dataset that pairs continuous XR traffic traces with session-level user-reported QoE labels. ResLearn-XR reduces SMAPE by up to 17.84% across frame-count, frame-size, and inter-arrival-time prediction, while reducing QoE-risk estimation SMAPE by up to 87.8% over single-stage baselines.
Chinese Translation
我们提出了ResLearn-XR,一个用于预测扩展现实网络流量和估计体验质量风险的残差学习框架。ResLearn-XR采用两阶段时序学习结构,由一个基础序列预测模型和针对特定任务的残差学习组件组成,以提升对突发性、非平稳XR流量动态的适应能力。残差学习阶段在值空间中进行连续XR流量预测,并在logit空间中进行概率化QoE风险估计。对于QoE风险分支,我们提出了一种数据描述符算法,这是一个因果特征构建模块,能够将数据包级的应用层可观测信息转换为适用于加密流量分析的、具备帧时序感知能力的描述符。我们还构建了一个XR流量-QoE数据集,将连续XR流量轨迹与会话级用户报告的QoE标签进行配对。ResLearn-XR在帧数、帧大小和帧间到达时间预测任务中将SMAPE最多降低17.84%,同时相比单阶段基线方法,将QoE风险估计的SMAPE最多降低87.8%。
cs.AI / 13 / 2609.04495

Rethinking Indirect Prompt Injection as a Test-Time Search Problem

重新审视间接提示注入:作为一个测试时搜索问题
Nguyen, Duong M., Kim, Joon Sik, Manczak, Blazej, Mugunthan, Vaikkunth
Abstract
We formulate indirect prompt injection as a test-time search over a task-dependent attack surface induced by the environment, user task, and injection task. To operationalize this formulation, we introduce an agentic attacker with a dedicated search harness that performs environment reconnaissance, structured reasoning over attack strategies, and adaptive evaluation using victim-agent feedback. Across heterogeneous tasks, we find that increasing attacker test-time compute improves vulnerability discovery and exploitation, while ablations show that explicit strategy management is important for avoiding redundant search and sustaining gains at larger budgets. These results suggest that agentic security evaluations should characterize both the attacker's search procedure and compute budget, rather than treating attack success as a budget-independent property of the victim. More broadly, our findings identify the attacker's adaptive search over the system attack surfaces as an important and underexplored security risk for tool-using agents.
Chinese Translation
我们将间接提示注入形式化为一个测试时搜索问题,其搜索空间是由环境、用户任务和注入任务所诱导的、依赖于任务的攻击面。为使这一形式化可操作化,我们引入了一种代理式攻击者(agentic attacker),其配备专门的搜索框架,能够执行环境侦察、对攻击策略进行结构化推理,并利用受害代理的反馈进行自适应评估。在多种异构任务上,我们发现增加攻击者的测试时计算量可以提升漏洞发现与利用的能力;消融实验表明,显式的策略管理对于避免冗余搜索以及在大计算预算下持续获得收益至关重要。这些结果表明,代理式安全评估应当同时刻画攻击者的搜索过程和计算预算,而不是将攻击成功视为与预算无关的受害方属性。更广泛地说,我们的研究结果指出,攻击者在系统攻击面上的自适应搜索是使用工具的智能代理所面临的一种重要且尚未被充分研究的安全风险。
cs.AI / 14 / 2609.04504

BioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital Biomarker

BioSync:基于Transformer的跨模态融合方法,用于多模态生理数字生物标志物
Mohammadabadi, Seyed Mahmoud Sajjadi
Abstract
Cardiac, neural, behavioral, and speech measurements from wearable and mobile devices provide partial, noise-sensitive views of physiological state. BioSync combines these measurements into the \textbf{BioSync Index (BSI)}, a continuous composite digital biomarker defined under the BEST framework. The model applies multi-head self-attention to modality tokens and adds a linear branch whose hypothesis class includes standard feature concatenation. This architecture is motivated by latent-variable measurement theory and by the possibility that joint observations contain information unavailable from individual modalities. We evaluated BioSync on two literature-informed synthetic cohorts: a four-modality cognitive-decline cohort using HRV, EEG, actigraphy, and speech, and a metabolic-autonomic cohort structured around the public AI-READI wearable schema. In the cognitive cohort, BioSync and concatenation obtained AUCs of 0.928 and 0.926, respectively. In the metabolic cohort, BioSync obtained accuracy/F1 of 0.764/0.766, compared with 0.756/0.758 for concatenation. The BSI correlated with latent severity in both cohorts ($r=0.91$ and $r=0.68$). A pure-attention ablation obtained cognitive-cohort AUC 0.911, locating the increase to 0.928 in the combined wide-and-deep architecture. With matched modality-dropout training, BioSync led concatenation at five of six cognitive-cohort corruption rates and at the highest metabolic-cohort rate. Its cognitive-cohort AUC was also higher than five published digital-biomarker reference values, although differences in datasets and tasks preclude a controlled benchmark claim. Comparison with single-modality, early-fusion, and late-fusion designs across six prespecified criteria identifies the model's computational properties; validation on real cohorts remains necessary.
Chinese Translation
来自可穿戴和移动设备的心脏、神经、行为及语音测量仅为生理状态提供了部分且对噪声敏感的视图。BioSync将这些测量整合为BioSync指数(BioSync Index, BSI),即在BEST框架下定义的连续型复合数字生物标志物。该模型对模态令牌(modality tokens)应用多头自注意力机制,并增加了一个线性分支,其假设类包含标准特征拼接方法。该架构的设计源于潜变量测量理论,以及联合观测可能包含单一模态无法获得的信息这一可能性。我们在两个基于文献构建的合成队列上评估了BioSync:一个是使用心率变异性(HRV)、脑电图(EEG)、体动记录仪和语音的四模态认知衰退队列;另一个是围绕公开的AI-READI可穿戴设备数据模式构建的代谢-自主神经队列。在认知队列中,BioSync与特征拼接分别获得了0.928和0.926的AUC。在代谢队列中,BioSync获得了0.764/0.766的准确率/F1值,而特征拼接为0.756/0.758。在两个队列中,BSI均与潜在严重程度相关(r=0.91和r=0.68)。纯注意力消融实验获得认知队列AUC 0.911,表明从0.911提升至0.928的增益来自宽深结合(wide-and-deep)的完整架构。在匹配的模态缺失训练条件下,BioSync在六个认知队列损坏率中的五个以及代谢队列的最高损坏率下均优于特征拼接方法。其认知队列AUC也高于五个已发表的数字生物标志物参考值,尽管数据集和任务的差异使得无法进行受控的基准比较。通过与单模态、早期融合和晚期融合设计在六个预设标准上的比较,本文阐明了该模型的计算特性;在真实队列上的验证仍有待开展。
cs.AI / 15 / 2609.04518

What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents

多执行框架强化学习学到了什么?编程智能体中的信用分配与可迁移性
Le, Chenqian, Cheng, Jiayi, He, Qijia, Li, Runhao, Li, Yinghao, Chen, Xupeng
Abstract
Agent reinforcement learning (RL) increasingly runs through full execution harnesses, and a multi-harness recipe mixes two choices: exposing the policy to several harnesses, and comparing their rewards inside one relative-advantage group. We isolate the second choice in repository-level coding. From one Qwen3-8B supervised warm start we replay the same frozen task-harness records from Aider, OpenHands, Qwen Code, and SWE-agent, with the same number of updates, under two rules for group-relative policy optimization (GRPO), Within (one group per task-harness pair) and Cross (harnesses pooled within a task), and score every checkpoint with a sealed SWE-bench Verified oracle on four source harnesses and a minimal harness held out of training. The evaluation harness is the dominant variable: across 24,000 sealed evaluations it moves the mean solve rate from 2.14\% to 9.27\%, a factor of 4.3, where the training recipe moves it by 1.16. The grouping rule is not. On the held-out harness, Cross minus Within is +0.25 pp, 95\% confidence interval [-0.48, +1.02], at eight attempts per task, and +0.16 [-0.41, +0.72] pooled over three training seeds whose individual estimates change sign. Each rule's own seed range, 0.42 to 0.45 pp, exceeds the difference between them. Both rules place their largest gains on the same source harness. The pooled advantage carries the harness: an out-of-fold classifier recovers the generating harness from Cross's advantage +4.48 pp above the shuffled-label baseline and from Within's not at all, and the two rules still reach the same held-out score and action distribution inside each harness. Re-collecting half the training data on-policy does not change this. Cross-harness credit yields configuration adaptation and no more portable capability than within-harness credit. Multi-harness RL reports should state the grouping boundary and test under an unseen harness.
Chinese Translation
智能体强化学习(RL)日益通过完整的执行框架(harness)来运行,而多框架训练方案混合了两个选择:一是让策略接触多个框架,二是在同一个相对优势组内比较不同框架的奖励。我们在仓库级编程任务中单独隔离了第二个选择。从同一个 Qwen3-8B 监督热启动出发,我们对 Aider、OpenHands、Qwen Code 和 SWE-agent 的相同冻结任务-框架记录进行重放,采用相同的更新次数,并在组相对策略优化(GRPO)的两种分组规则下进行训练:Within(每个任务-框架对一组)和 Cross(同一任务内所有框架合并分组),并使用密封的 SWE-bench Verified 评估器对每个检查点进行评分,评估对象包括四个源框架以及一个在训练中被保留的极简框架。评估框架是主导变量:在 24,000 次密封评估中,它使平均解决率从 2.14% 变动到 9.27%,相差 4.3 倍,而训练方案仅使其变动 1.16。分组规则则不然。在保留框架上,每任务八次尝试下 Cross 减去 Within 的差异为 +0.25 个百分点,95% 置信区间为 [-0.48, +1.02];在三个训练种子上汇总后差异为 +0.16 [-0.41, +0.72],而各单种子的估计值甚至正负号不一致。每种规则自身的种子波动范围(0.42 至 0.45 个百分点)就已超过二者之间的差异。两种规则的最大增益都出现在同一个源框架上。汇总优势承载了框架信息:一个折外分类器能够从 Cross 的优势中识别出生成框架,高于打乱标签基线 +4.48 个百分点,而无法从 Within 的优势中识别;但两种规则在每个框架内仍达到相同的保留集得分和动作分布。将一半训练数据改为在策略(on-policy)重新采集也不改变这一结论。跨框架信用带来的是配置层面的适应,其可迁移能力并不比框架内信用更多。多框架 RL 的研究报告应明确说明分组边界,并在未见过的框架下进行测试。
cs.AI / 16 / 2609.04523

MaxKernel: Agentic Kernel Generation for TPUs

MaxKernel:面向TPU的智能体化内核生成
Wang, Shangkun, Cai, Nina, Hoong, Charles, Walker, Julian, Kroiz, Gerson, Vanica, George, Patil, Deepak, Gavrilescu, Andi, Sipra, Hassan, Sankaran, Sethu
Abstract
Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-level expertise. Large Language Models (LLM) can be leveraged together with real-time compiler feedback to build agentic systems for kernel generation. In this work, we present MaxKernel, a multi-agent system that implements three distinct paradigms for TPU kernel development: (1) a Human-in-the-Loop (HITL) agent for collaborative, step-by-step design; (2) an Autonomous (Auto) agent that executes a fully automated, metric/trace-driven optimization loop; and (3) a Graph-Based Autonomous Search that scales the Auto agent for global exploration of the design space. All three paradigms leverage a shared pool of specialized sub-agents to handle planning, implementation, self-debugging, testing, and hardware profiling. We evaluate MaxKernel on JaxBench, a comprehensive suite of 50 diverse kernel tasks for TPUs, alongside complex, real-world workloads from state-of-the-art open-source models. We demonstrate that MaxKernel consistently generates highly optimized implementations, matching expert hand-tuned baselines and delivering significant performance across the benchmark. Our agent is open-sourced and available https://github.com/AI-Hypercomputer/accelerator-agents/tree/main/MaxKernel.
Chinese Translation
为加速器设计和编写高性能定制内核是一项复杂的任务,需要深入的硬件级专业知识。大语言模型(LLM)可以与实时编译器反馈相结合,构建用于内核生成的智能体系统。在本工作中,我们提出了MaxKernel,一个为TPU内核开发实现三种不同范式的多智能体系统:(1)人机协同(HITL)智能体,用于协作式的逐步设计;(2)自主(Auto)智能体,执行完全自动化、基于指标/追踪信息驱动的优化循环;(3)基于图的自主搜索,对Auto智能体进行扩展,以实现对设计空间的全局探索。三种范式均利用一个共享的专用子智能体池,负责规划、实现、自调试、测试和硬件性能分析。我们在JaxBench上评估MaxKernel,该基准包含50个多样化的TPU内核任务,以及来自最先进开源模型的复杂真实工作负载。实验表明,MaxKernel能够持续生成高度优化的实现,达到专家手工调优基线的水平,并在整个基准测试中带来显著的性能提升。我们的智能体已开源,可在 https://github.com/AI-Hypercomputer/accelerator-agents/tree/main/MaxKernel 获取。
cs.AI / 17 / 2609.04528

Towards a universal language of concepts: A survey

迈向概念通用语言:一篇综述
Parab, Aishni
Abstract
Humans can learn and generalize novel concepts from sparse data because they express knowledge in rich structural formats. In this paper, we propose that programs are a strong candidate for universal representation of concepts. We review computational models of concept learning that use programs as their concept representation and evaluate their contribution toward a universal representational language.
Chinese Translation
人类能够从稀疏数据中学习并泛化新概念,是因为他们以丰富的结构化格式表达知识。在本文中,我们提出程序(programs)是概念通用表示的有力候选。我们回顾了以程序作为概念表示的概念学习计算模型,并评估了这些模型对构建通用表示语言的贡献。
cs.AI / 18 / 2609.04541

Data-Driven Discovery of Composition-Dependent Constitutive Models for Hyperelasticity and Viscoelasticity of Digital Materials

数字材料超弹性与粘弹性成分依赖本构模型的数据驱动发现
García-Ávila, Josué, Shen, Beijun, Rausch, Manuel K., Boyce, Mary C., Buganza-Tepole, Adrián
Abstract
Digital materials fabricated by multi-material 3D printing are designed as controlled mixtures of stiff and compliant constituents, yielding effective responses that span more than an order of magnitude in apparent stiffness and exhibit strongly nonlinear, composition-dependent, and rate-dependent dissipative behavior. Classical finite-strain viscoelastic models represent such behavior with closed-form strain energy functions for equilibrium and non-equilibrium stresses as well as evolution of internal variables, which may limit flexibility when a single constitutive model is expected to generalize across materials and loading rates. Here, we present a data-driven multi-material constitutive modeling framework that generalizes a formulation by Bergstr\"om and Boyce. The proposed framework retains the structure of the classical model, namely multiplicative kinematics, invariant-based strain-energy functions, and a scalar dissipative evolution law directed along the normalized nonequilibrium deviatoric stress. For the equilibrium branch, the data-driven discovery framework either directly predicts closed-form model parameters as functions of composition or automatically constructs a polyconvex strain-energy function using neural ordinary differential equations (NODEs). The nonequilibrium branch kinetics are learned similarly, either by directly identifying closed-form parameters across compositions or by using appropriately constrained artificial neural networks. Using multi-rate uniaxial compression data across multiple material compositions, we show that the proposed formulation captures rate-dependent stiffness and hysteresis across compositions while preserving thermodynamic consistency.
Chinese Translation
通过多材料3D打印制备的数字材料被设计为刚性与柔性组分的受控混合物,其有效响应在表观刚度上跨越一个数量级以上,并表现出强非线性、成分依赖和速率依赖的耗散行为。经典有限应变粘弹性模型采用封闭形式的应变能函数来表示平衡应力与非平衡应力,以及内变量的演化,当期望单一本构模型在不同材料和加载速率间具有泛化能力时,这种形式可能限制其灵活性。本文提出一个数据驱动的多材料本构建模框架,该框架推广了Bergström和Boyce的公式。所提框架保留了经典模型的结构,即乘法运动学、基于不变量的应变能函数,以及沿归一化非平衡偏应力方向演化的标量耗散演化律。对于平衡分支,数据驱动发现框架要么直接将封闭形式的模型参数预测为成分的函数,要么利用神经常微分方程(NODEs)自动构造多凸应变能函数。非平衡分支的动力学以类似方式学习,或通过在不同成分间直接辨识封闭形式参数,或通过使用适当约束的人工神经网络。利用跨多种材料成分的多速率单轴压缩数据,我们证明所提公式能够捕捉不同成分下速率依赖的刚度和滞后行为,同时保持热力学一致性。
cs.AI / 19 / 2609.04543

From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs

从答案到解释:重新思考大语言模型中由歧义引发的偶然不确定性估计
Nahum, Omer, Nayman, Niv, Fhima, Jonathan, Zolfi, Alon, Levy, Jeremy, Mazor, Shai, Favaro, Paolo
Abstract
A key challenge in reliable LLM deployment is recognizing when uncertainty reflects irreducible variability in the task rather than limitations in the model's knowledge. In language tasks, a central source of such aleatoric uncertainty is input ambiguity or underspecification, where multiple interpretations remain plausible. Existing decomposition methods estimate aleatoric uncertainty by generating multiple clarifications of the input, querying the model for an answer under each clarification, and comparing the resulting answers. We argue that answers are not necessary for identifying ambiguity: they are often redundant, add avoidable cost, and can mislead through epistemic leakage. We support this claim theoretically, and propose a clarification-only approach that estimates this ambiguity-induced component directly from the space of plausible interpretations, without answers to the clarified inputs. Using ambiguity detection as an operational evaluation across three benchmarks, this direct approach improves AUROC (63.34 vs. 60.85), reduces computational cost by 4-26x in output tokens and 2.2-3.5x in API calls, and yields estimates with substantially lower correlation with epistemic uncertainty. Overall, our results suggest that ambiguity-induced aleatoric uncertainty is better estimated from the interpretation space than from the response space.
Chinese Translation
可靠部署大语言模型(LLM)的一个关键挑战在于识别不确定性究竟反映了任务本身不可约减的变异性,还是模型知识能力的局限。在语言任务中,这类偶然不确定性(aleatoric uncertainty)的核心来源是输入的歧义或欠明确性,即多种解释都可能成立。现有的分解方法通过为输入生成多个澄清版本、针对每个澄清版本查询模型答案、再比较所得答案来估计偶然不确定性。我们认为,识别歧义并不需要答案:答案往往是冗余的,会增加可避免的成本,并可能通过认知性泄漏(epistemic leakage)产生误导。我们从理论上支持了这一观点,并提出一种仅基于澄清的方法,直接从可能解释的空间中估计这种由歧义引发的成分,而无需对澄清后的输入给出答案。在三个基准数据集上以歧义检测作为操作性评估,该直接方法提升了 AUROC(63.34 对 60.85),将输出 token 的计算成本降低 4-26 倍、API 调用次数降低 2.2-3.5 倍,并且所得估计与认知不确定性(epistemic uncertainty)的相关性显著更低。总体而言,我们的结果表明,由歧义引发的偶然不确定性更适合从解释空间而非响应空间中进行估计。
cs.AI / 20 / 2609.04559

IPGeoAI: Transformer-Based Geolocation with LLM Semantic Fusion

IPGeoAI:基于Transformer并结合大语言模型语义融合的IP地理定位
Kadimisetty, Avinash, Yu, Andy Jinqing, Favaloro, Philip, Liu, Wenlong, Xiong, Xiaolu
Abstract
Accurate city-level IP Geolocation is an important enabler for the modern digital ecosystem, underpinning services ranging from local content delivery and targeting to digital rights enforcement. However, traditional heuristic and database-driven methods often struggle to resolve the complex, non-linear allocation patterns of modern network infrastructures, particularly within the exploding IPv6 address space and transient mobile networks. In this paper, we introduce IPGeoAI, a novel deep learning model architecture that reframes geolocation from a static lookup problem to a sequential modeling task. Our approach utilizes the Transformer Encoder to capture hierarchical dependencies inherent in IP subnet structures. We propose a method to resolve geographic ambiguity by integrating unstructured semantic context via a Zero-Shot LLM Feature Extraction pipeline. We utilize Large Language Models to transform raw, noisy Autonomous Systems (AS) descriptions into structured, domain-specific metadata (such as 'University' vs. 'ISP' or 'Global' vs. 'Local') via an offline pre-computation process. By fusing these semantic signals into the network via a Multi-Head Cross-Attention module, we bridge the gap between numerical network topology and real-world semantic identity. Extensive offline evaluation on a proprietary dataset spanning 200,000 cities demonstrates that IPGeoAI significantly outperforms a leading external vendor in city-level granularity. By adopting a hierarchical inference strategy that refines coarse-grained country signals, our model achieves a 6% improvement in city-level accuracy while extending coverage to 100% of the traffic. Furthermore, in large-scale online production tests, the model drove a statistically significant +0.35% improvement in our 1st-tier downstream use cases metric.
Chinese Translation
精确的城市级IP地理定位是现代数字生态系统的重要支撑,为从本地内容分发与定向投放到数字版权执法等各类服务提供基础。然而,传统的启发式方法和基于数据库的方法往往难以解析现代网络基础设施中复杂、非线性的地址分配模式,尤其是在急剧膨胀的IPv6地址空间和瞬态移动网络中。本文提出了IPGeoAI,一种新颖的深度学习模型架构,它将地理定位从静态查找问题重新构建为序列建模任务。我们的方法利用Transformer编码器(Transformer Encoder)来捕捉IP子网结构中固有的层级依赖关系。我们提出了一种方法,通过零样本大语言模型(Zero-Shot LLM)特征提取流水线整合非结构化的语义上下文,以解决地理歧义问题。我们利用大语言模型(LLM),通过离线预计算过程,将原始、嘈杂的自治系统(Autonomous System, AS)描述转化为结构化的领域特定元数据(例如'大学'与'互联网服务提供商',或'全球'与'本地')。通过多头交叉注意力(Multi-Head Cross-Attention)模块将这些语义信号融合到网络中,我们在数值化网络拓扑与真实世界语义身份之间架起了桥梁。在一个覆盖20万个城市的专有数据集上的广泛离线评估表明,IPGeoAI在城市级粒度上显著优于某领先的外部供应商。通过采用对粗粒度国家信号进行精化的层级推理策略,我们的模型在城市级准确率上提升了6%,同时将覆盖率扩展到流量的100%。此外,在大规模在线生产测试中,该模型为我们的一级下游用例指标带来了具有统计显著性的+0.35%提升。
cs.AI / 21 / 2609.04561

Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection

通过幻觉空间投影减少Whisper中的幻觉转录文本
Abbasihafshejani, Maryam, Jadliwala, Murtuza
Abstract
Whisper is a widely used foundation model for automatic speech recognition (ASR), but its generative decoder can produce fluent hallucinated transcripts for inputs containing little or no speech. We propose a training-free, inference-time method to reduce these hallucinations using low-rank projection of decoder activations. A compact hallucination-associated subspace is estimated from non-speech calibration data, and decoder hidden states are projected away from this subspace during inference. We evaluate two variants: always-on, which applies projection to all inputs, and gated, which applies it only when Whisper predicts that an input is likely non-speech. Across non-speech benchmarks, always-on projection reduces average hallucination rate (HR) from 31.31% to 2.44%, a 92.21% relative reduction, while gated projection reduces HR to 3.74%, an 88.05% relative reduction, with lower false rejection of genuine speech. On LibriSpeech, gated projection increases absolute word error rate (WER) by 0.33-4.39 percentage points and yields false-rejection rates (FRR) of 0.41--9.97% across model and split settings. These results show that low-rank activation projection can substantially suppress Whisper hallucinations without retraining, while providing a controllable trade-off between hallucination suppression and speech recognition performance.
Chinese Translation
Whisper是一种被广泛使用的自动语音识别(ASR)基础模型,但对于包含少量语音或完全没有语音的输入,其生成式解码器可能会产生流畅的幻觉转录文本。我们提出了一种无需训练的推理时方法,通过解码器激活的低秩投影来减少这些幻觉。该方法从非语音校准数据中估计出一个紧凑的幻觉相关子空间,并在推理过程中将解码器的隐藏状态投影到该子空间之外。我们评估了两种变体:always-on(始终启用),对所有输入应用投影;以及gated(门控),仅在Whisper预测输入可能是非语音时才应用投影。在多个非语音基准测试中,always-on投影将平均幻觉率(HR)从31.31%降至2.44%,相对降低92.21%;gated投影将幻觉率降至3.74%,相对降低88.05%,同时对真实语音的误拒率更低。在LibriSpeech上,gated投影使绝对词错误率(WER)增加0.33-4.39个百分点,并在不同模型和数据划分设置下产生0.41%-9.97%的误拒率(FRR)。这些结果表明,低秩激活投影无需重新训练即可显著抑制Whisper的幻觉,同时在幻觉抑制与语音识别性能之间提供了可权衡的控制。
cs.AI / 22 / 2609.04564

La Agente \'Optima: Towards Agentic Self-Driving Laboratories

La Agente Óptima:迈向智能体化的自驱动实验室
Müller, Marcel, Bai, Jiaru, Gottstein, Willi, Mandal, Abhijoy, Nazeri, Mohammad, Savino, Elia, Fang, Yanlin, Das, Sujoy, Carrillo, Sergio Pablo García, Kang, Yeonghun, Pérez-Sánchez, Juan B., Pilon, Simone, Fitzner, Martin, Noël, Timothy, Gu, Frank, Bernales, Varinia, Aspuru-Guzik, Alán
Abstract
Self-driving laboratories (SDLs) combine automated experimentation with adaptive decision-making to accelerate scientific discovery. Their operation nevertheless often depends on human specialists who translate scientific objectives into executable closed-loop campaigns. Specialists adjust them as data and operating conditions change. Here, we present La Agente \'Optima, an agentic framework that constructs and supervises Bayesian optimization campaigns across computational and experimental systems while maintaining a persistent optimization state. By separating large language model (LLM) reasoning from executed campaigns, \'Optima runs repetitive optimization loops consistently, returns control to the agent only when progress requires interpretation or campaign revision, and keeps every decision auditable. We evaluate \'Optima across ablation studies, five digital discovery tasks, and two physical platforms. Throughout, \'Optima maintained executable campaigns as both the scientific problem and execution environment evolved. In a closed-loop contact angle optimization campaign, \'Optima identified and corrected a mid-run measurement failure, bringing the contact angle from 71.4 to 67.8 degrees, just above the 64-66 degree range. From this result, \'Optima correctly inferred that the target was likely unattainable with the available reagents and recommended changing the formulation. In a five-day multi-objective flow-chemistry campaign, \'Optima increased the yield from 30% to 59% over 23 experiments. Despite substantial inference costs, it cost less and used substantially less starting material than a human-directed campaign, while selecting a more mass-efficient operating point. These results show that LLM-based agents can make rigorous, long-running optimization campaigns accessible to domain scientists without specialist setup, expanding the scope of SDLs.
Chinese Translation
自驱动实验室(Self-driving laboratories, SDLs)将自动化实验与自适应决策相结合,以加速科学发现。然而,其运行通常依赖人类专家将科学目标转化为可执行的闭环实验流程,并且随着数据和运行条件的变化,专家需要不断对其进行调整。本文提出La Agente Óptima,一个智能体化框架,它能够在计算和实验系统中构建并监督贝叶斯优化流程,同时维持持久的优化状态。通过将大语言模型(LLM)的推理与实际执行的实验流程分离,Óptima能够一致地运行重复的优化循环,仅在需要解读进展或修订实验流程时才将控制权交还给智能体,并确保每个决策都可审计。我们通过消融实验、五个数字发现任务以及两个物理平台对Óptima进行了评估。在整个过程中,随着科学问题和执行环境的演变,Óptima始终维持可执行的实验流程。在一个闭环接触角优化实验中,Óptima识别并纠正了运行中途的测量故障,将接触角从71.4度降至67.8度,仅略高于64–66度的目标范围。基于这一结果,Óptima正确推断出使用现有试剂可能无法达到目标,并建议更改配方。在一个为期五天的多目标流动化学实验中,Óptima在23次实验中将产率从30%提升至59%。尽管推理成本较高,但与人工指导的实验相比,其成本更低、消耗的起始原料大幅减少,同时选择了更具质量效率的操作点。这些结果表明,基于LLM的智能体能够让领域科学家无需专家配置即可开展严谨的长期优化实验,从而拓展了自驱动实验室的应用范围。
cs.AI / 23 / 2609.04565

Extremely Sparse Supervision Incentivizes Reasoning Ability

极稀疏监督激励推理能力
Liu, Zhishuai, Xu, Xingzi, Seyfioglu, Mehmet Saygin, Xu, Pan, Bouyarmane, Karim
Abstract
Large language models demonstrate increasingly strong reasoning capabilities through effective post-training. Yet, prevailing post-training methods optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. We revisit this assumption in the on-policy distillation (OPD) setting, which naturally admits dense teacher supervision at every generated token. Using the Qwen3 family, we discover a counter-intuitive phenomenon: reasoning can be effectively incentivized by an extremely small fraction of generated tokens--as few as one or two tokens per reasoning trajectory, corresponding to only 0.05% of all tokens. Surprisingly, this sparse supervision in most cases matches or surpasses full-token training in improving reasoning ability, despite excluding the vast majority of generated tokens from the training objective. This phenomenon is consistently observed across nine teacher--student configurations spanning different model scales on mathematical reasoning tasks, and is further validated on coding reasoning, Llama models and Proximal Policy Optimization (PPO)-based reinforcement learning with verifiable reward (RLVR). Interestingly, such extremely sparse supervision may be closer to the natural learning process: rather than correcting every step word by word, one reflects on a few critical reasoning steps, updates prior understanding, and continues the trial-and-error, avoiding micro-level corrections while remaining remarkably effective. Overall, our results challenge the assumption that effective post-training must be token-intensive and point to a new direction for understanding and designing more efficient post-training algorithms.
Chinese Translation
大型语言模型通过有效的后训练展现出日益强大的推理能力。然而,主流的后训练方法在海量 token 上进行优化,隐含地假定有效的学习必须依赖于大量 token。我们在策略上蒸馏(on-policy distillation, OPD)设定中重新审视了这一假设——该设定天然允许教师在每个生成 token 上提供密集监督。基于 Qwen3 系列模型,我们发现了一个反直觉的现象:仅使用极少比例的生成 token 即可有效激励推理能力——每条推理轨迹仅需一至两个 token,仅占全部 token 的 0.05%。令人惊讶的是,这种稀疏监督在提升推理能力方面在大多数情况下与全 token 训练相当甚至更优,尽管训练目标排除了绝大多数生成 token。该现象在数学推理任务上、覆盖不同模型规模的九组教师—学生配置中被一致观察到,并进一步在代码推理、Llama 模型以及基于近端策略优化(Proximal Policy Optimization, PPO)的可验证奖励强化学习(RLVR)中得到验证。有趣的是,这种极稀疏监督可能更接近自然的学习过程:人们并非逐字纠正每一步,而是反思少数关键的推理步骤、更新已有理解,然后继续试错——避免了微观层面的纠正,同时保持显著的有效性。总体而言,我们的结果挑战了有效后训练必须依赖大量 token 的假设,并为理解和设计更高效的后训练算法指明了新方向。
cs.AI / 24 / 2609.04579

Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines

所选对象能否到达阅读器?对基于真实数据的语言模型流水线中身份传递的审计
Vohra, Siddharth, Jiang, Runmin, Li, Xiaomo, Xu, Min
Abstract
Grounded language-model pipelines can be divided into three stages: selecting an object, retrieving passages for it, and using that evidence to answer. If the selected object must reach the reader, losing it breaks the handoff. Benchmark recall checks the dataset-linked object, which can differ. We audit 600 HybridQA questions across three selector families. On 1,463 resolvable records where the selected object matches the dataset-traced passage, exact key lookup and exact title matching return the object every time. With every ranked rule given the same decoded selected title, body-only BM25 omits it on 389 records (26.6%) at cutoff five, while hybrid retrieval with reranking omits it on 14 (1.0%). The two identities differ on 329 of 1,792 resolvable records. With original-question rankings, their top-five checks disagree on 106 records (5.9%). Frozen reader comparisons associate the aligned object's presence with 28.6 to 31.0 points higher exact match. In a deliberately selected 64-item cohort, removing that passage sharply lowers exact match, while removing a similar-length comparison passage does not reproduce the drop. We release the Returned-Object Profile (ROP), an executable record of the target, returned-ID field, cutoff, membership rule, and complete expected population, with data and an offline replay.
Chinese Translation
基于真实数据的语言模型流水线可分为三个阶段:选定一个对象、为其检索段落,并利用该证据作答。如果所选对象必须传递到阅读器,那么丢失该对象就会破坏这一传递过程。基准测试的召回检查的是与数据集关联的对象,而该对象可能与实际所选对象不同。我们对三个选择器家族的600个HybridQA问题进行了审计。在1,463条所选对象与数据集溯源段落相匹配的可解析记录上,精确键查找和精确标题匹配每次都能返回该对象。在给予每条排序规则相同解码所选标题的条件下,仅正文BM25在截断为5时有389条记录(26.6%)遗漏该对象,而带重排序的混合检索仅有14条(1.0%)遗漏。两种身份在1,792条可解析记录中的329条上存在差异。在使用原始问题的排序时,二者的前五项检查在106条记录(5.9%)上不一致。冻结阅读器的对比实验表明,对齐对象的存在与28.6至31.0个百分点更高的精确匹配率相关。在一个特意挑选的64条记录的队列中,移除该段落会显著降低精确匹配率,而移除一篇长度相近的对比段落则不会复现该下降。我们发布了Returned-Object Profile(ROP),即可执行记录,其中包含目标、返回ID字段、截断值、成员资格规则以及完整预期总体,同时附带数据和离线重放。
cs.AI / 25 / 2609.04611

$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction

$\tau^\tau$-Bench:一个用于端到端、真实化智能体构建的环境
Shi, Quan, Dhandhania, Keshav, Narasimhan, Karthik, Barres, Victor
Abstract
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and operate internal systems. Notably, the work of building them is increasingly handed to coding agents, yet existing benchmarks say little about whether an AI system can deliver one under the conditions of a real client engagement. We introduce $\tau^\tau$-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task. A developer agent is given the records a business actually keeps, a client who holds requirements, a production API that operations must run through, a codebase to inherit, and limits on serving cost and models: the same starting point a real engagement provides. From these it must deliver a complete customer-service agent, scored by deploying that agent against held-out simulated users. Across 53 tasks spanning four domains, the strongest configuration, Claude Opus 5 under Claude Code, passes just 23.9% of evaluation simulations. Meanwhile, an expert-authored reference ceiling scores 82.2%. The failures mirror ones human agent developers see: models issue shallow queries in place of deep comprehension of the records, communicate almost nothing to the client, and experiment too little with agent architecture and serving spend, shipping the first design that runs. We aim for $\tau^\tau$-bench to turn the work of cooperative agent building into a measurable target for coding agents.
Chinese Translation
LLM智能体正迅速成为生产级软件,被部署用于客户服务、争议裁决和内部系统运营。值得注意的是,构建它们的工作越来越多地交给了编码智能体(coding agents),然而现有基准测试几乎无法说明一个AI系统能否在真实客户项目的条件下交付此类系统。我们提出了$\tau^\tau$-bench(读作hyper-tau-bench),一个将智能体构建本身作为任务的基准测试。开发者智能体会获得企业在实际中保留的记录、一位持有需求的客户、一个运营必须接入的生产级API、一个待继承的代码库,以及服务成本和模型的限制——这与真实客户项目所提供的起点相同。基于这些条件,它必须交付一个完整的客户服务智能体,并通过将该智能体部署到留出的模拟用户上进行评分。在跨越四个领域的53个任务中,最强的配置——Claude Code下的Claude Opus 5——仅通过了23.9%的评估模拟。与此同时,由专家撰写的参考上限得分为82.2%。这些失败反映了人类智能体开发者所遇到的同类问题:模型用浅层的查询替代对记录的深入理解,几乎不与客户沟通,对智能体架构和服务投入的探索也过少,只是交付了第一个能运行的方案。我们希望$\tau^\tau$-bench能够将协作式智能体构建工作转变为编码智能体的一个可衡量的目标。
cs.AI / 26 / 2609.04627

Leveraging Imperfect Restoration for Data Availability Attack

利用不完美恢复实现数据可用性攻击
Huang, Yi, Styborski, Jeremy, Lyu, Mingzhi, Wang, Fan, Kong, Adams
Abstract
The abundance of online data is at risk of unauthorized usage in training deep learning models. To counter this, various Data Availability Attacks (DAAs) have been devised to make data unlearnable for such models by subtly perturbing the training data. However, existing attacks often excel against either Supervised Learning (SL) or Self-Supervised Learning (SSL) scenarios. Among these, a model-free approach that generates a Convolution-based Unlearnable Dataset (CUDA) stands out as the most robust DAA across both SSL and SL. Nonetheless, CUDA's effectiveness against SSL is underwhelming and it faces a severe trade-off between image quality and its poisoning effect. In this paper, we conduct a theoretical analysis of CUDA, uncovering the sub-optimal gradients it introduces and elucidating the strategy it employs to induce class-wise bias for data poisoning. Building on this, we propose a novel poisoning method named Imperfect Restoration Poisoning (IRP), aiming to preserve high image quality while achieving strong poisoning effects. Through extensive comparisons of IRP with eight baselines across SL and SSL, coupled with evaluations alongside five representative defense methods, we showcase the superiority of IRP. Code: https://github.com/lyumingzhi/IRP
Chinese Translation
海量在线数据面临在深度学习模型训练中被未经授权使用的风险。为应对这一问题,研究者提出了多种数据可用性攻击(Data Availability Attack, DAA),通过对训练数据进行细微扰动,使数据对这些模型而言不可学习。然而,现有攻击方法往往只对有监督学习(Supervised Learning, SL)或自监督学习(Self-Supervised Learning, SSL)场景之一有效。其中,一种生成基于卷积的不可学习数据集(Convolution-based Unlearnable Dataset, CUDA)的无模型方法,是在SSL和SL两种场景下都最为鲁棒的DAA。尽管如此,CUDA对SSL的攻击效果并不理想,且在图像质量与投毒效果之间存在严重的权衡。本文对CUDA进行了理论分析,揭示了其引入的次优梯度,并阐明了其通过诱导类间偏置来实现数据投毒的策略。在此基础上,我们提出了一种名为不完美恢复投毒(Imperfect Restoration Poisoning, IRP)的新型投毒方法,旨在保持高图像质量的同时实现强投毒效果。通过在SL和SSL场景下将IRP与八种基线方法进行广泛比较,并结合五种代表性防御方法的评估,我们展示了IRP的优越性。代码:https://github.com/lyumingzhi/IRP
cs.AI / 27 / 2609.04629

SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents

SiLR:面向大语言模型工具代理的结构保持准入与过程奖励
Zhou, Chenyu, Jiang, Qiliang, Wu, Shuning, Zhou, Xu
Abstract
A runtime gate for an LLM tool agent is usually cast as a filter. In a ReAct loop a rejected proposal is followed by another at the same state, so the gate is a search operator over the proposal stream whose admission criterion shapes which trajectories are reachable. We study post-violation recovery admission, where progress must be admitted while the system is still in violation, and identify the scalar projection trap: an aggregate-score gate accepts a locally improving proposal and commits the trajectory to a plateau. SiLR instead shadow-executes each proposal and admits it under a product order over the branch-level violation state (overloaded-branch support and per-branch severity). We prove that no scalar surrogate is sound for this order, so the failure is representational, not a matter of threshold tuning. On mined Gym-ANM scenarios, SiLR recovers 21/21 multi-action episodes against 0/21 for terminal and 9/21 for the best scalar gate, significant across the full 24-scenario benchmark. The terminal-versus-structured dichotomy holds across three model families and in CityLearn. Because admission rests on deterministic simulation, the LLM lies outside the trust boundary: a magnitude-redistribution attack that defeats both scalar and support-only baselines is contained only by the full per-branch predicate. With two constraint families active, every tested scalar projection admits physically unsafe actions; support-only admits the largest fraction (63.2% of 42,410; product order 0). In the hardest dual-family traces, scalar gates recover only through that unsafe class. Reused as a GRPO process reward, it outperforms its count projection in every mined scenario and is the only tested reward whose ungated policy exceeds the untrained base (0.844 vs. 0.778). Scalar projection loses the violation geometry at both design points; only the full product order is structurally sufficient.
Chinese Translation
大语言模型工具代理的运行时门控通常被视为一种过滤器。在ReAct循环中,一个被拒绝的提案之后会在同一状态下再次产生提案,因此该门控实际上是对提案流进行操作的搜索算子,其准入标准决定了哪些轨迹是可达的。我们研究违反后恢复准入问题,即系统仍处于违反状态时必须允许进展被准入,并识别出标量投影陷阱:基于聚合分数的门控会接受一个局部改进的提案,从而使轨迹陷入平台期。SiLR则对每个提案进行影子执行,并在分支级违反状态(过载分支支持度与逐分支严重程度)的乘积序下进行准入。我们证明对于该序不存在可靠的标量代理,因此这种失败是表示层面的,而非阈值调优问题。在挖掘的Gym-ANM场景上,SiLR在多动作回合中恢复21/21,而终止式门控为0/21,最佳标量门控为9/21,且在全部24个场景的基准上均具有显著性。终止式与结构化的二分结论在三个模型系列以及CityLearn中均成立。由于准入依赖于确定性仿真,大语言模型处于信任边界之外:一种能够击败标量门控和仅支持度基线的幅值重分配攻击,只能被完整的逐分支谓词所遏制。在两组约束族同时激活的情况下,所有测试的标量投影都会准入物理上不安全的动作;仅支持度门控准入的不安全动作比例最大(42,410个中占63.2%;乘积序为0)。在最困难的双约束族轨迹中,标量门控仅能通过该不安全类别实现恢复。将该方法复用为GRPO过程奖励时,它在所有挖掘场景中均优于其计数投影,且是唯一未经门控的策略即可超越未训练基线的奖励(0.844对0.778)。标量投影在两个设计节点上都丢失了违反几何结构;只有完整的乘积序在结构上是充分的。
cs.AI / 28 / 2609.04641

A Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a New Benchmark

面向嵌套企业模式的成本感知型自然语言转SQL智能体架构及新基准
Varadharajan, Yoga Sri Varshan, Yadav, Ajay, Goru, Ritesh, Chaudhury, Prateek, Caramanis, Constantine, Jain, Prateek, Pasupuleti, Divyateja, Pandey, Sunil Kumar
Abstract
Natural-language-to-SQL systems have ad- vanced rapidly on academic benchmarks, yet production enterprise schemas exhibit graph- like, semi-structured, deeply nested structure that current benchmarks do not measure. We make two complementary contributions. First, we introduce the DevRev NL2SQL bench- mark: 900 execution-verified queries with nested-type and link-graph structure, accom- panied by the Semantic Depth Score (SDS), a schema-agnostic rubric for analytical reasoning depth. Second, we present a cost-aware single- generation agentic architecture whose schema- selection, metadata-retrieval, and error-repair components are designed for the requirements this regime imposes. On the DevRev NL2SQL benchmark the system attains 91.7% answer correctness, a margin of 54.6 percentage points over the next-best baseline; on the Spider 2.0 Snowflake public dataset, it is competitive with leading systems at a single-generation operating point.
Chinese Translation
自然语言转SQL(NL2SQL)系统在学术基准测试上进展迅速,然而生产环境中的企业模式呈现出图状、半结构化且深度嵌套的结构,这是现有基准测试所无法衡量的。我们做出了两项互补性的贡献。第一,我们推出了DevRev NL2SQL基准:包含900条经过执行验证的查询,具有嵌套类型和链接图结构,并配以语义深度评分(Semantic Depth Score, SDS)——一种与模式无关、用于衡量分析推理深度的评分标准。第二,我们提出了一种成本感知的单次生成智能体架构,其模式选择、元数据检索和错误修复组件均针对该场景所提出的需求而设计。在DevRev NL2SQL基准上,该系统达到了91.7%的答案正确率,比次优基线高出54.6个百分点;在Spider 2.0 Snowflake公开数据集上,它在单次生成的运行点下与领先系统不相上下。
cs.AI / 29 / 2609.04651

Continual Graph Memory for Adaptive Recommendation under Intent Drift

面向意图漂移下自适应推荐的持续图记忆方法
Ngoc, Hao Nguyen, Nguyen, Tung, Hanh, Nguyen Thi, Dinh, Hoang Thai, Tung, Nguyen Xuan
Abstract
This paper studies adaptive recommendation under intent drift, where feedback from each recommendation outcome can reveal whether the relational evidence used for ranking is useful, missing, or misleading. While Knowledge Graphs (KGs) provide essential semantic structure to handle these shifts, traditional KG-enhanced systems treat the graph as a static retrieval substrate, making it brittle to evolving intents, noisy metadata, and recurring failure patterns. This paper proposes CGM-Rec, a continual graph memory framework for adaptive recommendation. CGM-Rec treats the graph state as a writable memory and maintains two complementary components. Therein, a Semantic Graph Memory is updated conservatively through quality-gated typed operations for storing stable and high-confidence relational knowledge. Meanwhile, an Episodic Lesson Memory acts as a fast reactive memory that learns recent outcomes, failure cases, and corrective hints. During testing, model parameters remain frozen and adaptation occurs only through memory writes. We evaluate CGM-Rec under a frozen-parameter, one-pass reranking protocol, where encoders and prompts remain fixed during testing and adaptation occurs only through memory writes. Experiments across multiple recommendation settings show that CGM-Rec improves over evaluated neural and LLM-based baselines on most metrics. Particularly, under sampled-candidate reranking, CGM-Rec improves HR@1 by up to 29.58% over the strongest LLM baseline on Bundle, and outperforms K-RagRec on metadata-rich ML-100K with HR@5 of 0.5941 versus 0.4746.
Chinese Translation
本文研究意图漂移下的自适应推荐问题,其中每次推荐结果的反馈可以揭示用于排序的关系证据是有用的、缺失的还是误导性的。尽管知识图谱能够提供处理这些漂移所必需的语义结构,但传统的知识图谱增强系统将图视为静态的检索基座,使其在面对不断演化的意图、嘈杂的元数据和反复出现的失败模式时表现脆弱。本文提出CGM-Rec,一种用于自适应推荐的持续图记忆框架。CGM-Rec将图状态视为可写写的记忆,并维护两个互补的组件:其中,语义图记忆通过质量门控的类型化操作进行保守更新,用于存储稳定且高置信度的关系知识;同时,情景经验记忆作为快速反应性记忆,学习近期的结果、失败案例和纠正性提示。在测试阶段,模型参数保持冻结,自适应仅通过记忆写入实现。我们在冻结参数、单次重排序的协议下评估CGM-Rec,其中编码器和提示词在测试期间保持固定,自适应仅通过记忆写入进行。在多种推荐设置下的实验表明,CGM-Rec在大多数指标上优于所评估的神经网络基线和基于大语言模型(LLM)的基线。特别地,在采样候选重排序设置下,CGM-Rec在Bundle数据集上相比最强的大语言模型基线将HR@1提升了高达29.58%,并在元数据丰富的ML-100K数据集上优于K-RagRec,其HR@5达到0.5941,而K-RagRec为0.4746。
cs.AI / 30 / 2609.04665

Harness-agnostic detection and immunization of reward hacking in self-evolving language models

自进化语言模型中奖励作弊的框架无关检测与免疫
Yang, Rongxin, Liu, Yang, Luo, Shang, Jia, Haoxuan, Zhang, Chongyang, Zheng, Hao, Yang, Yingguang, Huang, Yulin, Zhang, Jianshen, Qi, Yongzhi, Xu, Kefu, Ran, Congjing, Chong, Bin
Abstract
Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible score. When that score is an imperfect proxy for the capability one actually wants, sustained selection widens the gap between the two. This is reward hacking. We introduce HackProbe, a monitor that attaches to an arbitrary self-evolving loop through two black-box hooks, with no access to weights or activations. It keeps a secret, distribution-fixed comparison core, whose frozen distribution makes its capability proxy comparable across generations, alongside a rotated fresh layer that hardens the bank against co-adaptation. Four tests built on that proxy cover the level gap, a scale-aligned divergence with online change-point detection, capability stagnation, and a conditional confidently-wrong rate; a Sidak correction turns them into a calibrated family-wise p-value. Diagnosis alone recovers nothing, so a risk-aware immunization layer reselects an honest candidate from the proposal pool using the core together with a purely structural gaming footprint, disclosing at most log2 Pi bits per generation to the host. We prove a detectability bound that converts a target error rate into an explicit probe-size budget, and we delimit what probe rotation does and does not buy. On a controlled prompt-level host with four injected hacking channels and ground-truth labels, HackProbe reaches 0.763 AUROC against 0.663 for the strongest baseline and cuts the false-positive rate from 0.706 to 0.434. Its bandwidth-limited reselection is the only immunization level that returns more true capability under hacking, 5.2 points on average, than it forfeits on clean runs, 4.7; per-channel effects are mostly not individually significant.
Chinese Translation
自进化语言模型通过提出候选更新并保留任何能提升可见分数的更新来不断改进。当该分数只是所实际追求能力的不完美代理时,持续的选择会扩大两者之间的差距,这就是奖励作弊(reward hacking)。我们提出了HackProbe,一个通过两个黑盒钩子接入任意自进化循环的监控器,无需访问权重或激活值。它维护一个秘密的、分布固定的比较核心,其冻结的分布使其能力代理在各代之间具有可比性;同时还维护一个轮换的新鲜层,以增强该核心库抵抗协同适应的能力。基于该代理构建的四项测试分别涵盖水平差距、结合在线变点检测的尺度对齐散度、能力停滞以及条件性的自信错误率;经Šidák校正后,它们构成一个经校准的族类p值。仅有诊断无法挽回任何损失,因此一个风险感知的免疫层利用比较核心以及纯结构性的作弊足迹,从提案池中重新选出一个诚实的候选更新,每代最多向宿主泄露log2 Π比特信息。我们证明了一个可检测性界限,将目标错误率转化为明确的探针规模预算,并界定了探针轮换所能带来与不能带来的收益。在一个包含四个注入作弊通道和真实标签的受控提示级宿主上,HackProbe达到0.763的AUROC,而最强基线仅为0.663,并将假阳性率从0.706降至0.434。其带宽受限的重选机制是唯一一种在作弊情况下获得的真实能力收益(平均5.2分)超过在干净运行中所付出代价(4.7分)的免疫层级;各通道的效应大多不具单独的统计显著性。
cs.AI / 31 / 2609.04667

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

ERPBench:跨竞争市场生态评估企业决策的大语言模型智能体
Zhang, Xinran, Lu, Pengrui, Ye, Lyumanshan, Liu, Pengfei
Abstract
Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations rarely test whether business-decision conclusions transfer across competitive market ecologies. We introduce ERPBench, an execution-instrumented benchmark for enterprise decision agents in a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition. ERPBench evaluates the same 100 fixed problems in two matched competitive market ecologies: Solo, where each evaluated LLM agent competes against fixed rule-based opponents, and Arena, where six evaluated LLM agents compete in a shared market. Across six model families, this yields 1,200 model-level trajectories spanning 7,200 decision rounds. Under the observed service configuration, the leading model differs between ecologies: DeepSeek leads in Solo (252.29M mean valuation; mean rank 1.67), whereas Gemini leads in Arena (263.95M; 1.76). The two ecologies identify the same task-level winner on only 21 of 100 problems, and Gemini's bottom-rank rate falls from 22 % to 0 % in Arena. ERPBench supports paired evaluation of whether enterprise-agent rankings transfer across competitive market ecologies, supplemented by aggregate execution-intervention analysis. Code and benchmark resources are available in our https://github.com/GAIR-NLP/erp-bench.
Chinese Translation
大语言模型(LLM)智能体被越来越多地提出用于企业工作流,但现有评估很少检验业务决策结论是否能在不同竞争市场生态之间迁移。我们提出了ERPBench,一个面向企业决策智能体的可执行化基准,基于一个六轮企业资源计划(ERP)仿真,其中耦合了定价、生产、采购、库存、财务以及共享市场竞争。ERPBench在两种匹配的竞争市场生态中评估相同的100个固定问题:Solo(单独)模式下,被评估的LLM智能体与固定的基于规则的对手竞争;Arena(竞技场)模式下,六个被评估的LLM智能体在同一共享市场中竞争。在六个模型家族上,这共产生1,200条模型级轨迹,涵盖7,200个决策轮次。在观测到的服务配置下,领先模型因生态而异:DeepSeek在Solo模式中领先(平均估值2.5229亿;平均排名1.67),而Gemini在Arena模式中领先(2.6395亿;1.76)。两种生态仅在100个问题中的21个上识别出相同的任务级赢家,且Gemini的垫底率在Arena模式中从22%降至0%。ERPBench支持对企业智能体排名是否跨竞争市场生态迁移进行成对评估,并辅以总体执行干预分析。代码和基准资源可在我们的 https://github.com/GAIR-NLP/erp-bench 获取。
cs.AI / 32 / 2609.04678

Train What You Deploy:Token-Faithful Post-Training of a Production Coding

训练即部署:面向生产级编码模型的令牌忠实后训练
Li, Cheng, Liu, Jiexiong, Chen, Yixuan, Hong, Chi
Abstract
Existing post-training pipelines for coding and terminal agents suffer severe token and control fidelity errors: simplified training environments mismatch production deployments, and offline token reconstruction from agent logs distorts original prompts and conflates policy calls with background model operations. We present a fidelity-aware training coupling framework that retains trainer-side sampling over original prompts, eliminates spurious model calls via a negotiated training protocol, and restricts loss computation to verifiable token spans with closed-failure guarantees. We further propose Certified Divergence Proximal Policy Optimization (C-DPPO), which establishes tight two-sided TV certification bounds, adaptive-K rules, budget-aware sequence guarantees, and error-robust policy masking atop standard DPPO. Evaluated on matched Baize5B and Baize10B models with identical training and test protocols on TMax-100, C-DPPO yields a consistent +3.0-point performance gain over standard DPPO across model scales. Certificate audits validate the reliability and full operational coverage of our certified training pipeline.
Chinese Translation
现有的面向编码与终端智能体(agent)的后训练流程存在严重的令牌与控制保真度误差:简化的训练环境与生产部署环境不匹配,而基于智能体日志的离线令牌重构会扭曲原始提示词,并将策略调用与后台模型操作混为一谈。我们提出了一种保真度感知的训练耦合框架,该框架在训练侧保留对原始提示词的采样,通过协商式训练协议消除虚假模型调用,并将损失计算限制在具有闭环失效保证的可验证令牌区间上。我们进一步提出了认证散度近端策略优化(Certified Divergence Proximal Policy Optimization, C-DPPO),在标准DPPO的基础上建立了紧凑的双侧总变差(TV)认证界、自适应K规则、预算感知的序列保证以及容错鲁棒策略掩码。我们在TMax-100上使用相同训练与测试协议对匹配的Baize 5B和Baize 10B模型进行评估,C-DPPO在各模型规模上相较标准DPPO取得了一致的+3.0个百分点性能提升。证书审计验证了我们认证训练流程的可靠性及完整的运行覆盖能力。
cs.AI / 33 / 2609.04693

Predicting Spatiotemporal Mobile Sensing-Based PM2.5 Concentrations Using Low-Rank Adapted Spatially Attentive Graph Neural Network

基于低秩适配空间注意力图神经网络的移动感知PM2.5浓度时空预测
Chiddarwar, Om, Mandal, Priyanka, Chandaliya, Praveen Kumar, Arkatkar, Shriniwas
Abstract
Urban air quality can vary significantly along transit corridors, necessitating high-resolution monitoring. This work introduces a novel mobile-sensing dataset from Surat, Gujarat, India, comprising PM$*{2.5}$ concentrations, meteorological variables (temperature, humidity, wind speed, wind direction), and land-use features. To represent the spatiotemporal data as a graph, two node-definition strategies were used: (i) uniform segmentation (200--400~m intervals) and (ii) DBSCAN clustering to adaptively group dense observations. For each node, rolling mean and standard deviation of meteorological variables were computed. To model this high-dimensional data, we propose a SA-GNN for fine-grained, short-term PM$*{2.5}$ forecasting and hotspot identification. We compared SA-GNN with LSTM, RNN, GRU, and ANN models. These models performed well on low-resolution data but had difficulty capturing rapidly changing patterns in urban air quality. SA-GNN employs cluster-specific GRUs to capture localized temporal dependencies and a Graph Attention Network to learn spatial heterogeneity. This hybrid architecture effectively models rapid fluctuations and complex spatial interactions. On our dataset, SA-GNN achieved $R^2 = 0.95$, RMSE $= 6.8$, and MAE $= 4.2~\si{\micro\gram\per\meter\cubed}$, outperforming all baseline models. Combining spatial clustering with adaptive attention significantly improves forecasting, enabling real-time, fine-grained monitoring and supporting personalized exposure tracking and timely alerts for healthier cities.
Chinese Translation
城市空气质量在交通走廊沿线可能存在显著差异,因此需要进行高分辨率监测。本研究引入了一个来自印度古吉拉特邦苏拉特(Surat, Gujarat, India)的新型移动感知数据集,包含PM2.5浓度、气象变量(温度、湿度、风速、风向)以及土地利用特征。为将时空数据表示为图,我们采用了两种节点定义策略:(i) 均匀分割(200-400米区间);(ii) 采用DBSCAN聚类自适应地对密集观测进行分组。针对每个节点,计算了气象变量的滚动均值和标准差。为对该高维数据进行建模,我们提出了一种用于细粒度、短期PM2.5预测与热点识别的空间注意力图神经网络(SA-GNN)。我们将SA-GNN与LSTM、RNN、GRU和ANN模型进行了比较。这些基线模型在低分辨率数据上表现良好,但难以捕捉城市空气质量快速变化的模式。SA-GNN采用针对各聚类的GRU来捕捉局部时间依赖性,并利用图注意力网络(Graph Attention Network)学习空间异质性。这种混合架构能够有效建模快速波动和复杂的空间交互。在我们的数据集上,SA-GNN取得了R²=0.95、RMSE=6.8、MAE=4.2微克/立方米的成绩,优于所有基线模型。将空间聚类与自适应注意力相结合显著提升了预测性能,实现了实时、细粒度的监测,并支持个性化暴露追踪和及时预警,助力建设更健康的城市。
cs.AI / 34 / 2609.04697

SQL-Zero: Self-Evolving Text-to-SQL

SQL-Zero:自进化的Text-to-SQL
Pedrozo, Daniel Machado, Dollis, Julia Soares, de Oliveira, Bryan Lincoln Marques, Aguiar, Vinicius Alboneti, de Oliveira, Sávio Salvarino Teles, Soares, Telma Woerle de Lima
Abstract
Training a competitive Text-to-SQL agent usually depends on human-annotated natural-language/SQL pairs, which are expensive, domain-specific, and a bottleneck for scaling to new databases. We show it is possible to train a competitive solver with zero annotated pairs. We introduce SQL-Zero, a proposer-solver self-play in which a challenger and a solver start from the same base LLM and the only ground truth is execution against the database itself. The challenger generates SQL pairs calibrated to the solver's current difficulty (targeting "hard but solvable"), and both roles are updated with GRPO in alternating turns, with a template-level repetition penalty on the challenger to prevent diversity collapse. Training on BIRD databases with no labels, self-play improves over the zero-shot base on BIRD dev by 6.6 points at 3B and 7.3 points at 7B. It also scores higher than a matched control trained under the same recipe on human BIRD gold over the same databases, although an exact paired test does not resolve that margin. Transfer depends on scale: at 3B every iteration outperforms the base on unseen Spider databases and under lexical perturbation (Spider-Syn), where it also degrades less than the matched BIRD-gold control, whereas at 7B only the first iteration preserves transfer.
Chinese Translation
训练一个具有竞争力的Text-to-SQL智能体通常依赖于人工标注的自然语言/SQL配对数据,这类数据成本高昂、局限于特定领域,并成为向新数据库扩展的瓶颈。我们证明,在不使用任何标注配对数据的情况下训练出具有竞争力的求解器是可能的。我们提出SQL-Zero,一种提议者-求解者自我博弈方法,其中挑战者和求解者从同一个基础大语言模型出发,唯一的真实标签是与数据库本身执行后的结果。挑战者生成与求解者当前难度相校准的SQL题目对(目标是“有难度但可解”),两个角色交替使用GRPO进行更新,并对挑战者施加模板级别的重复惩罚以防止多样性坍缩。在无标注的BIRD数据库上训练后,自我博弈在BIRD开发集上的表现较零样本基础模型分别提升了6.6个百分点(3B模型)和7.3个百分点(7B模型)。在相同数据库上的人工BIRD黄金标准评测中,其得分也高于在相同训练配方下训练的对照组,尽管严格的配对检验未能确证这一优势。迁移能力取决于模型规模:在3B规模下,每一轮迭代在未见过的Spider数据库和词汇扰动(Spider-Syn)场景下均优于基础模型,且退化程度低于匹配的BIRD黄金标注对照组;而在7B规模下,仅第一轮迭代能保持迁移能力。
cs.AI / 35 / 2609.04699

Model Retirement Creates Reproducibility Risk in Biomedical AI Publications

模型退役对生物医学AI出版物可重复性构成风险
Wolfrath, Nathan, Conroy, Meghan, Kosten, Thomas, Bell, Dave, Neupane, Bhabishya, Kindel, Jonah, Banerjee, Anjishnu, Deshpande, Priya, Taylor, Bradley, Kothari, Anai N.
Abstract
Background. Large language models (LLMs) are being adopted in biomedical research at a rapid and accelerating pace, yet commercial services that host many widely used models operate under deprecation schedules that can complicate scientific reproducibility. Methods. We searched PubMed for original research articles from 2022 through March 2026 that applied a specific LLM to a biomedical task. An extraction agent identified model names from 61,077 article abstracts with human reviewers validating a subset for extraction accuracy. Extracted model names were normalized to canonical model identifiers. Lifecycle data (release date, retirement date, status) were compiled for the 50 most frequently used models. Results. We identified 8,931 paper-model mentions spanning 5,242 unique publications after restricting the analysis to the 50 most frequently used models. Among these mentions, 77.7% cited a commercial closed-weight model. Overall, 42% involved a model that was already retired by the time of official publication or is scheduled to retire within two years of publication. The median interval from publication to model retirement was 538 days. Conclusion. Many biomedical publications using LLMs are on a trajectory toward computational non-reproducibility after publication. Model deprecation should be treated as a core reporting and preservation issue for biomedical research.
Chinese Translation
背景。大语言模型(LLMs)正以快速且不断加速的步伐被生物医学研究所采用,然而托管众多广泛使用模型的商业服务在停用计划下运营,这可能使科学可重复性复杂化。方法。我们检索了PubMed中2022年至2026年3月期间将特定LLM应用于生物医学任务的原创研究文章。一个抽取代理从61,077篇文章摘要中识别模型名称,并由人工审阅者对部分子集进行抽取准确性验证。抽取的模型名称被规范化为标准模型标识符。我们对使用频率最高的50个模型整理了生命周期数据(发布日期、退役日期、状态)。结果。在将分析限定于使用频率最高的50个模型后,我们共识别出8,931条论文-模型提及,涉及5,242篇独立出版物。在这些提及中,77.7%引用了商业闭源权重模型。总体而言,42%的提及所涉及的模型在正式发表时已经退役,或计划在发表后两年内退役。从发表到模型退役的中位时间间隔为538天。结论。许多使用LLM的生物医学出版物正走向发表后计算不可重复性的境地。模型退役应被视为生物医学研究报告规范和成果保存的核心问题。
cs.AI / 36 / 2609.04706

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

FinalityBench:面向延迟与冲突性金融终态下智能体决策的效果级基准
Sharma, Abhishek
Abstract
A merchant's payment processor, ledger, ERP and bank feed are updated by messages that get delayed, duplicated, dropped and reordered, so for minutes at a time the four hold contradictory beliefs about the same order. An agent resolving the exception must decide whether to ship goods, re-submit a capture, refund or wait, knowing some of those cannot be undone. We present FinalityBench, an executable benchmark for that decision. It keeps a hidden canonical event log and derives each system's view from a separately faulted delivery stream, so disagreement follows from specified fault semantics rather than being authored. Grading is on executed monetary effects: an episode is scored by the merchant's terminal economic position, relative to a privileged reference told when the pending capture resolves. The corpus of 321 tasks includes 45 twin pairs (90 tasks): tasks whose four system views are identical at the decision instant, whose authoritative probes both return unknown, and whose eventual correct dispositions differ. That snapshot indistinguishability is checked under every evaluation seed rather than assumed; equivalence over all interaction traces is not claimed. Over 14,445 graded episodes from nine programmatic policies, ranking by single-task accuracy and by paired loss disagree in 7 places: a ship-on-first-sign policy is second-best by accuracy at 65.7% and worst in the suite by paired loss, because it cannot tell the two members apart. A runtime gating irreversible actions on an authoritative finality probe reaches 85.4% and, unlike every polling policy, loses nothing to pass^5; its residual loss is almost entirely one archetype, which prices finality information directly. Language models reach the same exact rate as the hand-written gate on a stratified subset, lose about twice as much money, and discover the finality-gating strategy without being told it.
Chinese Translation
商户的支付处理系统、账本、ERP 和银行数据流由经延迟、重复、丢失和乱序的消息进行更新,因此在长达数分钟的时间窗口内,这四个系统对同一笔订单持有相互矛盾的信念。负责处理该异常的智能体必须决定是发货、重新提交扣款、退款还是等待,并且知道其中某些操作无法撤销。我们提出了 FinalityBench,一个针对此类决策的可执行基准。该基准维护一个隐藏的权威事件日志,并从各自带独立故障的投递流中推导每个系统的视图,因此分歧源于明确规定的故障语义,而非人为设定。评分基于已执行的货币效果:每个回合(episode)以商户的最终经济状况计分,并与一个知晓待处理扣款何时解决的特权参考进行对比。该语料库包含 321 个任务,其中有 45 对孪生任务(共 90 个任务):即四个系统视图在决策时刻完全一致、其权威探测均返回未知、但最终正确处置方式不同的任务。这种快照不可区分性是在每个评估种子下进行检验的,而非仅仅假设;我们不声称在所有交互轨迹上等价。在来自九个程序化策略的 14,445 个已评分回合上,按单任务准确率与按配对损失排名的结果在 7 处存在分歧:一种在首次出现信号即发货的策略在准确率上以 65.7% 排名第二,但在配对损失上却是整个测试套件中最差的,因为它无法区分孪生任务的两方。一种通过权威终态探测在运行时对不可逆操作进行门控的策略达到 85.4% 的准确率,且与所有轮询策略不同,其损失不会低于 pass^5 的下限;其残余损失几乎完全来自一种原型,该原型对终态信息进行直接定价。语言模型在分层子集上达到了与手写门控完全相同的准确率,但损失的资金约为其两倍,并且在无人告知的情况下自行发现了终态门控策略。
cs.AI / 37 / 2609.04715

PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces

PLUME:基于共享子空间中低秩用户调制的大语言模型参数高效个性化方法
Li, Xinyu, Zhou, Hao, Zhu, Jianfeng, Maharjan, Julina, Guo, Ruixin, Dragan, Feodor, Jin, Ruoming
Abstract
Personalizing large language models (LLMs) is essential for delivering AI assistance that aligns with individual users' styles, intents, and preferences. While per-user fine-tuning can substantially enhance personalization quality, it introduces significant parameter and storage overhead, limiting scalability to large user populations. We propose PLUME (Personalized Low-Rank Adaptation through User Modulation and Shared Subspace), a lightweight framework that achieves efficient and expressive per-user adaptation by leveraging a shared task-specific subspace. Specifically, PLUME first learns a global task subspace from aggregated user data. Personalization is then achieved by training only a lightweight small square matrix within this subspace, enabling each user to obtain a tailored model while keeping shared components fixed. Cross-layer shared parameters and rank-1 residual terms are further introduced to significantly reduce redundancy while maintaining expressiveness. Experiments on multiple personalized text generation benchmarks demonstrate that PLUME achieves comparable or superior performance to strong baselines, while reducing per-user parameters by over 95%. These results establish shared-subspace modulation with minimal residuals as a scalable and semantically grounded approach to LLM personalization.
Chinese Translation
大语言模型(LLM)的个性化对于提供符合个体用户风格、意图和偏好的AI辅助至关重要。尽管针对每个用户进行微调可以显著提升个性化质量,但它会引入巨大的参数和存储开销,限制了对大规模用户群体的可扩展性。我们提出了PLUME(通过用户调制和共享子空间实现个性化低秩适应,Personalized Low-Rank Adaptation through User Modulation and Shared Subspace),这是一个轻量级框架,通过利用共享的任务特定子空间,实现高效且富有表现力的逐用户适配。具体而言,PLUME首先从聚合的用户数据中学习全局任务子空间。随后,仅需在该子空间内训练一个轻量级的小方阵即可实现个性化,使每个用户都能获得定制化模型,同时保持共享组件固定不变。此外,我们引入了跨层共享参数和秩1(rank-1)残差项,在保持表达能力的同时显著减少了冗余。在多个个性化文本生成基准上的实验表明,PLUME取得了与强基线相当或更优的性能,同时将每个用户的参数量减少了95%以上。这些结果证明,基于共享子空间调制并辅以极小残差项的方法是一种可扩展且语义有据可依的大语言模型个性化方案。
cs.AI / 38 / 2609.04738

Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM

Aplaud:面向用户特定大语言模型的自适应个性化低秩分解方法
Li, Xinyu, Jin, Ruoming, Zhu, Jianfeng, Guo, Ruixin, Liu, Zhi
Abstract
In this paper, we study the problem of personalized survey response prediction using fine-tuned large language models (LLMs). This task poses unique challenges: limited per-user training data, scalability of model storage, and the need to exploit shared structure across survey questions. To address these issues, we propose Aplaud (Adaptive Personalized Low-rank and User-specific Nested Decomposition), a lightweight and scalable framework for LLM personalization. Aplaud extends the LoRA paradigm by separating adaptation into a frozen, shared low-rank basis and a compact user-specific correction, augmented with a rank-one residual for finer personalization. To further reduce per-user parameter cost and mitigate overfitting, the correction matrix can be factorized into an even lower-rank form. Empirical results demonstrate that Aplaud achieves efficient, scalable personalization across users while outperforming state-of-the-art LoRA-based personalized LLM approaches in both generalization and inference efficiency.
Chinese Translation
本文研究了利用微调后的大语言模型(LLM)进行个性化调查问卷回答预测的问题。该任务面临若干独特挑战:每个用户的训练数据有限、模型存储的可扩展性受限,以及需要利用调查问题之间的共享结构。为解决这些问题,我们提出了Aplaud(Adaptive Personalized Low-rank and User-specific Nested Decomposition,自适应个性化低秩与用户特定嵌套分解),这是一个轻量级且可扩展的大语言模型个性化框架。Aplaud对LoRA范式进行了扩展,将适配过程分解为一个冻结的、共享的低秩基矩阵和一个紧凑的用户特定校正矩阵,并引入秩为一的残差项以实现更精细的个性化。为进一步降低每个用户的参数开销并缓解过拟合,该校正矩阵还可以被进一步分解为更低秩的形式。实验结果表明,Aplaud在用户之间实现了高效且可扩展的个性化,同时在泛化能力和推理效率方面均优于当前最先进的基于LoRA的个性化大语言模型方法。
cs.AI / 39 / 2609.04749

DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems

DCFA:面向基于大语言模型的多智能体系统失败原因推理的双视角因果启发归因方法
Wang, Zehao, Wang, Lanjun, Jin, Shilong, Chen, Junjie, Xiao, Yanghua
Abstract
Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite their promise, such systems remain fragile, frequently exhibiting reasoning and coordination errors that can lead to system-level failures. Failure attribution in such systems relies on tracing natural language interactions among agents to identify the decisive error, which refers to the earliest action whose correction can reverse system failure. There are two key challenges: 1) Shallow attribution: Existing methods often capture only minor deviations, such as incomplete retrievals or formatting errors, which verification mechanisms can correct, while missing the decisive cause of system failure. 2) Contextual degradation: As the length of the system traces increases, the model's reasoning ability rapidly deteriorates. To address these challenges, we propose DCFA, a training-free framework for failure attribution. DCFA integrates a global module that constructs structured causal-inspired dependency graphs from system traces to identify the initial decisive error, and a local module that applies local counterfactual-inspired reasoning to refine causal-inspired attribution. Experiments on the Who&When benchmark across six LLMs show that DCFA improves step-level accuracy by up to 8.27% over state-of-the-art baselines.
Chinese Translation
基于大语言模型(LLM)的多智能体系统近年来发展迅速。尽管前景广阔,此类系统仍然十分脆弱,频繁出现推理与协作错误,可能导致系统级故障。此类系统中的失败归因需要追踪智能体之间的自然语言交互,以识别决定性错误,即最早的可纠正并能扭转系统失败的错误。这里存在两个关键挑战:1)浅层归因:现有方法往往只能捕获轻微偏差,例如不完整的检索或格式错误,而这些问题可由验证机制加以纠正,从而遗漏系统失败的决定性原因;2)上下文退化:随着系统轨迹长度的增加,模型的推理能力会迅速下降。为应对这些挑战,我们提出了DCFA,一个无需训练的失败归因框架。DCFA集成了一个全局模块,从系统轨迹中构建结构化的因果启发式依赖图以识别最初的决定性错误;同时引入一个局部模块,应用局部反事实启发式推理来细化因果启发式归因。在Who&When基准上针对六个大语言模型的实验表明,DCFA相比最先进的基线方法,步骤级准确率最高提升了8.27%。
cs.AI / 40 / 2609.04767

Shadow Queries for Private Retrieval in Vector Databases

面向向量数据库隐私检索的影子查询方法
Feng, Xinguo, Ma, Zhongkui, Wang, Zihan, Yan, Chuan, Yang, Guowei, Abuadbba, Alsharif, Bai, Guangdong
Abstract
Large language models (LLMs) increasingly rely on information retrieval (IR) systems, such as Retrieval-Augmented Generation (RAG), to incorporate domain-specific knowledge without costly re-training. These systems often store pre-computed document embeddings in cloud-based vector databases. However, such embeddings are vulnerable to embedding inversion attacks (EIAs), which can reconstruct their underlying text. Existing defenses, such as adding noise or scaling embeddings, often provide limited privacy or significantly reduce retrieval utility. We propose SHAQ (shadow query generation), a semantic-decomposition and embedding-decoupling defense against EIAs. SHAQ is based on the insight that EIAs rely on the strong coupling between an embedding and its original text. Instead of storing document embeddings directly, SHAQ uses a generative language model to create diverse shadow queries that capture different semantic aspects of each document. These queries are then encoded and stored in place of the original document embeddings, thereby decomposing document semantics and decoupling stored embeddings from the source text. Experiments across diverse IR datasets show that SHAQ substantially improves privacy while preserving retrieval utility, achieving a recovery rate as low as 0.2104, defending up to 19.50% more tokens than baseline defenses, and reaching up to 0.7967 MAP@10 with up to 5.53% utility improvement. These results demonstrate that semantic decomposition and embedding decoupling provide an effective alternative to directly modifying embeddings for defending against EIAs.
Chinese Translation
大语言模型(LLM)日益依赖信息检索(IR)系统,例如检索增强生成(RAG),以在无需昂贵重新训练的情况下引入领域特定知识。这类系统通常将预先计算的文档嵌入存储在云端向量数据库中。然而,这些嵌入容易受到嵌入反演攻击(Embedding Inversion Attacks, EIAs)的威胁,攻击者可据此重构其底层文本。现有防御方法(如添加噪声或缩放嵌入)通常只能提供有限的隐私保护,或显著降低检索效用。我们提出了SHAQ(影子查询生成),一种基于语义分解与嵌入解耦的对抗EIAs的防御方法。SHAQ的核心洞察在于:EIAs依赖于嵌入与原始文本之间的强耦合关系。SHAQ不再直接存储文档嵌入,而是利用生成式语言模型创建多样化的影子查询,以捕捉每个文档的不同语义方面。随后对这些查询进行编码并替代原始文档嵌入进行存储,从而实现文档语义的分解,并将存储的嵌入与源文本解耦。在多个信息检索数据集上的实验表明,SHAQ在保持检索效用的同时显著提升了隐私保护能力:文本恢复率低至0.2104,相比基线防御方法可多保护高达19.50%的词元(token),MAP@10最高达到0.7967,效用提升最高达5.53%。这些结果证明,语义分解与嵌入解耦为防御EIAs提供了一种有效的替代方案,而非直接修改嵌入本身。
cs.AI / 41 / 2609.04778

Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges

面向移动边缘智能体人工智能的扩散语言模型:基础、应用与挑战
Li, Chenqi, Min, Minghui, Niyato, Dusit, Ni, Wei
Abstract
Diffusion language models (DLMs) offer a non-autoregressive alternative for mobile edge agentic artificial intelligence (AI) by refining tokens through iterative denoising rather than left-to-right decoding. Compared with autoregressive Transformer-based large language models (LLMs), DLMs can update multiple uncertain tokens in parallel and exploit bidirectional context throughout the generation process, enabling more flexible quality-latency trade-offs beyond fixed sequential decoding. These properties are particularly attractive for edge agents, where partial refinement, early exit, and constraint-guided correction can reduce response delay and communication overhead while improving robustness under noisy, incomplete, or dynamic contexts. This survey reviews DLM foundations and analyzes their suitability for edge settings under latency, memory, energy, bandwidth, privacy, and reliability constraints. We cover resource-efficient architectures, training and inference acceleration, compression, edge/cloud deployment, communication-aware serving, Internet of Things (IoT)/wireless applications, and evaluation of DLM-based agents. We further discuss open issues in long-context state management, split inference, trustworthy execution, multimodal grounding, and reproducible benchmarking. The goal is to connect DLM modeling properties, including bidirectionality, parallel refinement, controllability, and quality-latency elasticity, with system-level requirements of future mobile edge intelligence.
Chinese Translation
扩散语言模型(Diffusion Language Models, DLMs)通过迭代去噪而非从左到右的自回归解码来优化词元,为移动边缘智能体人工智能(AI)提供了一种非自回归的替代方案。与基于自回归Transformer的大语言模型(LLMs)相比,DLMs能够并行更新多个不确定词元,并在整个生成过程中利用双向上下文,从而实现超越固定顺序解码的更灵活的质量-时延权衡。这些特性对边缘智能体尤其具有吸引力:部分细化、提前退出和约束引导纠错等机制能够在噪声、不完整或动态上下文下提升鲁棒性,同时降低响应时延和通信开销。本综述回顾了DLMs的基础理论,并在时延、内存、能耗、带宽、隐私和可靠性等约束条件下分析了其适用于边缘环境的特性。我们涵盖了资源高效架构、训练与推理加速、模型压缩、边/云部署、通信感知服务、物联网(IoT)/无线应用,以及基于DLMs的智能体评估。此外,我们还讨论了长上下文状态管理、分离式推理、可信执行、多模态 grounding 以及可复现基准测试等方面的开放性问题。本文的目标是将DLMs的建模特性(包括双向性、并行细化、可控性和质量-时延弹性)与未来移动边缘智能的系统级需求联系起来。
cs.AI / 42 / 2609.04782

DODR: Deterministic Operator-Driven Reasoning in Latent Space

DODR:潜空间中的确定性算子驱动推理
Huang, Weicai
Abstract
Autoregressive (AR) large language models formulate reasoning as token-level probabilistic sampling, which induces three fundamental defects in complex logical reasoning: error accumulation, probability substituting necessity, and the linear-chain information bottleneck. This paper proposes the Deterministic Operator-Driven Reasoning in Latent Space architecture (DODR), which reconstructs reasoning as reasoning-graph computation in a high-dimensional linear-algebraic space. Reasoning states are represented as snapshot vectors whose primitives are semantic units (phrases or sentences) rather than tokens, and each inference step is a deterministic matrix operation with no token sampling. Peirce's three inference types are formalized as three trainable matrix operators: a rank-deficient deduction operator (information collapse), a full-rank induction operator (information expansion), and an abduction operator defined as the Moore-Penrose pseudo-inverse of deduction (information hypothesizing). We prove that the operator set is minimal and complete given Peirce's trichotomy, that no single "super-operator" can realize all three types (a rank obstruction), and that reasoning graphs are Turing-complete with contractive backflow converging by Banach's fixed-point theorem. Experiments on 503 sample records (420 deduplicated samples) across dedicated and end-to-end settings show: deduction loss converges to 1.40e-05; induction achieves 0.9996 generalization coverage with 20/20 hard vetoes on counterexamples; abduction solutions exceed the random baseline by 28x with judgment accuracies of 72.5% (58/80, Wilson 95% CI [61.9%, 81.1%]) and 81.7% (49/60, CI [70.1%, 89.4%]); frozen operators attain 100% (60/60) on unseen cross-domain deduction. The architecture provides a structural zero-hallucination guarantee and a three-layer continual-learning mechanism. All data and code are released.
Chinese Translation
自回归(AR)大语言模型将推理形式化为词元级别的概率采样,这在复杂逻辑推理中引发了三个根本性缺陷:误差累积、概率替代必然性,以及线性链式信息瓶颈。本文提出潜空间中确定性算子驱动推理架构(DODR),将推理重构为高维线性代数空间中的推理图计算。推理状态被表示为快照向量,其基本单元是语义单元(短语或句子)而非词元,且每一步推理都是确定性矩阵运算,不涉及词元采样。皮尔士(Peirce)的三种推理类型被形式化为三个可训练的矩阵算子:亏秩的演绎算子(信息坍缩)、满秩的归纳算子(信息扩展),以及定义为演绎算子之Moore-Penrose伪逆的溯因算子(信息假说化)。我们证明了:给定皮尔士的三分法,该算子集合是极小且完备的;不存在能同时实现三种推理类型的单一"超级算子"(秩障碍);推理图是图灵完备的,且收缩回流依Banach不动点定理收敛。在503条样本记录(去重后420条样本)上、于专用与端到端两种设定下的实验表明:演绎损失收敛至1.40e-05;归纳实现0.9996的泛化覆盖率,并在反例上获得20/20的硬否决;溯因求解超出随机基线28倍,判断准确率分别为72.5%(58/80,Wilson 95%置信区间 [61.9%, 81.1%])和81.7%(49/60,置信区间 [70.1%, 89.4%]);冻结算子在未见过的跨域演绎任务上达到100%(60/60)。该架构提供了结构性的零幻觉保证与三层持续学习机制。所有数据与代码均已开源。
cs.AI / 43 / 2609.04793

ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

ProtLingo:基于条件记忆与专家路由的高效蛋白质语言建模
Li, Mingrui, Shen, Sixian, Li, Minzhang, Zhang, Ruiyi, Zhang, Kexin, Zhang, Jiakai, Yu, Jingyi
Abstract
Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transformer backbones often brings substantial computational cost without consistently improving mutation-sensitive prediction. We introduce ProtLingo, an efficient PLM framework that augments a pretrained single-sequence backbone with conditional local memory and sparse expert routing. ProtLingo maps contextual residue representations into route-specific discrete codes, composes centered local windows into latent $N$-gram addresses, and retrieves reusable residual signals associated with recurring local sequence contexts. In parallel, selected feed-forward blocks are upcycled into sparse Mixture-of-Experts layers with shared and routed experts, enabling residue-dependent computation while activating only a subset of parameters. Experiments on protein fitness prediction, FLIP benchmarks, and supervised contact prediction show that ProtLingo achieves competitive performance with a 150M-scale backbone, including strong parameter efficiency on mutation-effect prediction and preserved long-range structural representations.
Chinese Translation
蛋白质执行着多样化的细胞功能,即使是单个氨基酸的替换也可能改变蛋白质的稳定性、活性或分子相互作用。蛋白质语言模型(PLM)提供了一种从无标注序列中建模此类序列-功能关系的可扩展方法,但增大稠密 Transformer 骨干网络的规模往往带来高昂的计算成本,却未能持续提升对突变敏感的预测性能。我们提出了 ProtLingo,这是一种高效的 PLM 框架,通过条件局部记忆和稀疏专家路由来增强预训练的单序列骨干网络。ProtLingo 将上下文残基表示映射为路由特定的离散编码,将中心化的局部窗口组合为潜在的 $N$-gram 地址,并检索与重复出现的局部序列上下文相关联的可复用残差信号。同时,选定的前馈模块被升级改造为具有共享专家和路由专家的稀疏混合专家(Mixture-of-Experts)层,在仅激活一部分参数的情况下实现依赖残基的动态计算。在蛋白质适应性预测、FLIP 基准以及有监督接触预测上的实验表明,ProtLingo 以 150M 规模的骨干网络取得了具有竞争力的性能,在突变效应预测上表现出色且参数效率高,同时保留了长程结构表示。
cs.AI / 44 / 2609.04801

Whose record is this? Diagnosing and authorizing record use in personalized multimodal models

这是谁的记录?个性化多模态模型中记录使用的诊断与授权
Mao, Xinyu, Li, Junsi, Liu, Chenyang, Zhang, Haoji, Sun, Ming
Abstract
Contextualized visual personalization can retrieve a true record yet apply it to the wrong visual subject. We formalize when a record may condition an answer as \emph{record authorization}: subject presence ($P$), record-edge validity ($E$), and answer support ($S$) must all hold. We call violations visual memory misbinding (VMM). We construct RecordAuth-Diag, a 3,690-case matched diagnostic suite that changes one image--record edge while holding the query, question, record text, and image multiset fixed. Card removal and nonce relabeling attribute these failures to supplied records. Raw-bank failures span Qwen-, Phi-, and Gemma-family interfaces: Gemma-3-4B-IT reaches 63.69\% local unauthorized use at 25.75\% clean recall. CoViP remains at 26.02\%, versus 22.49\% for its Qwen backbone at similar clean recall. Typed pre-generation authorization reduces Qwen card exposure on RecordAuth-Diag from 43.63\% to 3.06\%, while positive recall changes from 86.26\% to 60.90\%. Full $P\wedge E\wedge S$ validation uses 560 localized DAVIS cases: top-1 relevance and typed authorization have comparable release (28.93\% and 28.39\%) but 6.79\% and 0.89\% unsafe release, respectively. Of the 33 additional unsafe cases removed, 27 are support, 4 edge, 2 clean, and 0 boundary cases. Thus the observed increment is an $E\wedge S$ decision dominated by support, not an edge check alone. Appearance supplies $E$ evidence only conditional on $P$; authenticated subject tokens instantiate the missing presence witness as a sufficiency control. The claims concern the evaluated contracts, not natural prevalence, consent, or visual identity
Chinese Translation
情境化的视觉个性化可以检索到真实记录,却可能将其应用于错误的视觉主体。我们将一条记录可以在何种条件下用于约束答案形式化为“记录授权”(record authorization):主体存在性(P)、记录-边缘有效性(E)和答案支持性(S)必须同时成立。我们将违反这一约束的现象称为视觉记忆错绑(VMM)。我们构建了RecordAuth-Diag,一个包含3,690个案例的匹配诊断套件,在保持查询、问题、记录文本和图像多重集不变的情况下,仅改变一个图像-记录边缘。移除卡片和随机重标注实验将这些失败归因于所提供的记录。原始记录库的失败跨越Qwen、Phi和Gemma系列接口:Gemma-3-4B-IT在干净召回率为25.75%时,局部未授权使用率达到63.69%。CoViP停留在26.02%,而其Qwen骨干模型在相近的干净召回率下为22.49%。类型化的预生成授权将Qwen卡片在RecordAuth-Diag上的暴露率从43.63%降至3.06%,而正例召回率从86.26%变为60.90%。完整的P∧E∧S验证使用了560个局部化的DAVIS案例:top-1相关性和类型化授权的放行率相当(28.93%和28.39%),但不安全放行率分别为6.79%和0.89%。在移除的33个额外不安全案例中,27个为支持性问题、4个为边缘问题、2个为干净案例、0个为边界案例。因此,观察到的性能提升是一个以支持性为主导的E∧S决策,而非单纯的边缘检查。外观仅在主体存在性(P)成立的条件下才能提供E证据;经认证的主体标记将缺失的存在性见证实例化,作为一种充分性控制。本文的结论仅涉及所评估的契约,而非自然分布中的普遍性、用户同意或视觉身份。
cs.AI / 45 / 2609.04803

Hierarchical Possession-Aware Graph Pointer Network for Pass Receiver Selection

面向传球接球人选择的层次化控球感知图指针网络
Wang, Jingyi, Li, Da, Wang, Kaixin, Huang, Zhangqin
Abstract
Pass receiver selection is a fundamental task in football analytics, aiming to predict the intended receiver under a given game state. This task is challenging with event-centered freeze-frame observations, a broadcast-like setting that provides only partial and variable player visibility without complete trajectories or stable player identities. The model must therefore reason over anonymous visible candidates, opponent pressure, and recent context under partial observation. To address this setting, we propose a Hierarchical Possession-aware Graph Pointer Network (HPGPN), which formulates pass receiver selection as variable-size candidate prediction over visible teammates. HPGPN jointly models current player interactions, local event context, and possession-level temporal dynamics. It represents the current pass situation with a graph, incorporates fixed event context, and uses dynamic possession history to capture how the attacking sequence evolves. Candidate representations are refined hierarchically by integrating spatial, contextual, and historical evidence, and a glimpse pointer head scores the receiver candidates. Experiments on public football event and freeze-frame data show that HPGPN improves pass receiver selection performance. Ablation studies demonstrate the effectiveness of graph-based interaction modeling, fixed event context, and dual-branch dynamic possession-history modeling.
Chinese Translation
传球接球人选择是足球分析中的一项基础任务,旨在预测给定比赛状态下传球方的预期接球人。在以事件为中心的冻结画面观测下,这一任务极具挑战性——这种类似转播视角的场景仅提供部分且多变的球员可见性,缺乏完整轨迹和稳定的球员身份信息。因此,模型必须在部分观测条件下对匿名的可见候选球员、对手压力以及近期上下文进行推理。针对这一场景,我们提出了一种层次化控球感知图指针网络(Hierarchical Possession-aware Graph Pointer Network, HPGPN),将传球接球人选择建模为对可见队友的可变尺寸候选预测。HPGPN 联合建模当前球员交互、局部事件上下文以及控球层面的时间动态。该模型用图结构表示当前传球局面,融合固定的事件上下文,并利用动态控球历史来刻画进攻序列的演化过程。候选球员表示通过融合空间、上下文和历史证据进行层次化精炼,并由 glimpse 指针头对接球人候选进行打分。在公开的足球事件与冻结画面数据上的实验表明,HPGPN 提升了传球接球人选择的性能。消融研究验证了基于图的交互建模、固定事件上下文以及双分支动态控球历史建模的有效性。
cs.AI / 46 / 2609.04804

MedFlow: Class-Aware Multi-Scale Generation for Medical Time-Series Synthesis

MedFlow:面向医学时间序列合成的类别感知多尺度生成方法
Huang, Yanhao, Feng, Shibo, Feng, Wanjin, Zhao, Peilin, Miao, Chunyan
Abstract
Synthetic medical time-series generation can alleviate data scarcity and support the development of reliable clinical prediction models. However, existing methods mainly focus on matching the overall distribution and temporal dynamics of real data, which does not necessarily ensure strong downstream utility on imbalanced medical datasets. Clinically informative patterns often occur at heterogeneous temporal scales, while rare minority-class characteristics can be obscured by dominant population patterns. To address these challenges, we propose MedFlow, a class-aware multi-scale flow matching framework for medical time-series synthesis. MedFlow employs a vector-quantized multi-scale tokenizer to represent medical sequences at complementary temporal resolutions, capturing both coarse clinical trends and fine-grained dynamics. We further introduce Token Marginal Guidance, which incorporates class-conditional token statistics directly into the flow matching process to steer generation toward class-specific regions of the learned tokens. This mechanism strengthens minority-class patterns, while preserving the global and tail distributions of real data. Experiments on four public datasets covering electronic health records, EEG, and ECG signals demonstrate that MedFlow consistently outperforms recent state-of-the-art diffusion-based baselines across downstream prediction tasks. On average, it improves AUPRC by 5.8%, reduces Context-FID by 88.6%, and achieves 3.8$\times$ higher sampling throughput.
Chinese Translation
合成医学时间序列生成可以缓解数据稀缺问题,并支持可靠临床预测模型的开发。然而,现有方法主要关注匹配真实数据的整体分布和时间动态特性,这不一定会确保在不平衡医学数据集上具有良好的下游效用。具有临床信息价值的模式通常出现在异构的时间尺度上,而稀有少数类别的特征可能被占主导地位的人群模式所掩盖。为了应对这些挑战,我们提出了 MedFlow,一种用于医学时间序列合成的类别感知多尺度流匹配框架。MedFlow 采用向量量化的多尺度分词器(tokenizer)以互补的时间分辨率表示医学序列,同时捕捉粗粒度的临床趋势和细粒度的动态特性。我们进一步引入了 Token 边缘引导(Token Marginal Guidance),将类别条件化的 token 统计信息直接融入流匹配过程中,引导生成朝着学习到的 token 中类别特定的区域进行。这一机制在保持真实数据全局分布和尾部分布的同时,增强了少数类别的模式。在涵盖电子健康记录、脑电(EEG)和心电(ECG)信号的四个公开数据集上的实验表明,MedFlow 在下游预测任务中始终优于近期的最先进扩散模型基线。平均而言,它将 AUPRC 提高了 5.8%,将 Context-FID 降低了 88.6%,并实现了 3.8 倍的采样吞吐量提升。
cs.AI / 47 / 2609.04806

When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models

当金融微调失效时:领域自适应语言模型中数值幻觉的三级可检测性分析
Li, Xiaodong, Liu, Peiwei
Abstract
Financial large language models are increasingly deployed for summarization of reports and disclosures, where numerical hallucination poses significant practical risks. While prior work often attributes such hallucination to insufficient numerical reasoning, this assumption has not been systematically tested under controlled fine-tuning settings. In this paper, we conduct a cost-effective, controlled study of numerical hallucination in financial summarization across three model variants: a base instruction-tuned model, a domain language-adapted model (FT-A), and a numeracy-enhanced domain model (FT-A+B+C). We introduce a three-level detectability taxonomy distinguishing between overt hallucination (currency-denominated fabrication), covert-explicit hallucination (professional-convention numbers), and covert-implicit hallucination (ungrounded quantitative claims). Our results reveal that domain fine-tuning substantially degrades numerical restraint at all detectability levels. While the Base model maintains near-zero hallucination rates (5.4\%), FT-A exhibits 82.5\% overt hallucination and FT-A+B+C reaches 98\%. Contrary to intuition, numeracy supervision amplifies rather than mitigates hallucination across all levels. We identify template injection---the insertion of memorized canonical values regardless of input content---as a primary hallucination mechanism in fine-tuned models. These findings demonstrate that numerical hallucination in financial summarization is driven by the degradation of numerical restraint through domain adaptation, not by insufficient numerical reasoning. We recommend that evaluation protocols assess hallucination across all detectability levels and that deployment practices include explicit mechanisms for grounding-aware generation or abstention.
Chinese Translation
金融大语言模型日益被用于报告和信息披露的摘要生成,其中数值幻觉带来了显著的实践风险。尽管已有研究常将此类幻觉归因于数值推理能力不足,但这一假设尚未在受控微调设置下得到系统检验。本文针对三种模型变体——基础指令微调模型(Base)、领域语言自适应模型(FT-A)以及数值增强领域模型(FT-A+B+C)——开展了金融摘要任务中数值幻觉的低成本受控研究。我们提出了一种三级可检测性分类体系,用以区分显性幻觉(货币金额捏造)、隐性显式幻觉(遵循专业惯例的数字)和隐性隐式幻觉(无依据的定量陈述)。结果显示,领域微调在所有可检测性层级上都大幅削弱了模型的数值克制能力。基础模型保持接近于零的幻觉率(5.4%),而FT-A的显性幻觉率达到82.5%,FT-A+B+C则高达98%。与直觉相反,数值能力监督在所有层级上不仅未能缓解幻觉,反而使其加剧。我们识别出模板注入——即无视输入内容而插入记忆中的规范化数值——是微调模型产生幻觉的主要机制。这些发现表明,金融摘要中的数值幻觉源于领域自适应导致的数值克制能力退化,而非数值推理能力不足。我们建议评估方案应覆盖所有可检测性层级的幻觉检测,并且部署实践中应包含支持依据感知生成或拒答的显式机制。
cs.AI / 48 / 2609.04809

CPR-IE:A Compression-Prediction-Resource Intelligence Efficiency Metric

CPR-IE:一种压缩-预测-资源智能效率度量
Jiang, Xiantao
Abstract
Comparing intelligent systems under deployment constraints requires more than predictiveaccuracy.This paper develops Compression-Prediction-Resource Intelligence Efficiency (CPR-IE) as a protocol-relative ordering by representational economy, predictive quality, and resourceburden. The analysis separates two questions-how raw resource consumption is represented, andhow the resulting attributes are aggregated. Proportional-increment composition uniquely yieldslogarithmic cumulative burden, and context-independent ratio response yields power responsesto compression, prediction, and burden; with reference normalization the representation is I(C,P,T).We prove Pareto consistency, unit invariance, boundary behavior, trade-off identities, ranking-stability regions, and cross-task aggregation. A translog parent model makes interaction restrictions explicit, and further results establish cardinal and ordinal identification, sub-Gaussianfinite-sample ranking guarantees, robust selection under exponent uncertainty, and deterministicregret bounds. Minimum description length, algorithmic complexity, proper scoring rules, varia-tional inference, and Landauer's principle motivate measurement choices but do not entail theformula. CPR-IE is a constructed efficiency representation, not a universal law or a definition ofintelligence itself.
Chinese Translation
在部署约束条件下比较智能系统,仅依靠预测精度是不够的。本文提出了压缩-预测-资源智能效率(CPR-IE)指标,作为一种基于表征经济性、预测质量和资源负担的、相对于协议的排序方法。本文的分析区分了两个问题:如何表示原始资源消耗,以及如何聚合由此产生的属性。比例增量组合唯一地产生对数形式的累积负担,而与上下文无关的比率响应产生关于压缩、预测和负担的幂律响应;通过参考归一化,该表示形式为 I(C,P,T)。我们证明了帕累托一致性、单位不变性、边界行为、权衡恒等式、排序稳定区域以及跨任务聚合。一个超越对数(translog)母模型使交互约束显式化,进一步的结果确立了基数与序数可识别性、亚高斯有限样本排序保证、指数不确定性下的稳健选择以及确定性遗憾界。最小描述长度、算法复杂度、严格评分规则、变分推断和兰道尔原理为测量选择提供了动机,但并不必然推出该公式。CPR-IE 是一种构造出的效率表示,而非一条普适定律,也不是对智能本身的定义。
cs.AI / 49 / 2609.04840

Long Horizon Transformer Quantile Fault Prediction for Multi Site Industrial Predictive Maintenance

面向多厂区工业预测性维护的长时程Transformer分位数故障预测
Poland, David J, Ravi, Daniele, Helian, Na
Abstract
Long-horizon predictive maintenance requires models to distinguish slowly evolving degradation from normal operating-regime variation over planning windows measured in days rather than hours. This paper evaluates whether an explicit conditional-quantile representation provides an informative classifier interface for this problem. The proposed TQRNN30d framework combines a dual-stage quantile regression neural network (QRNN) feature extractor with a multi-stream temporal fusion classifier. Each hourly word of 81-channel machine behaviour is mapped to a 324-dimensional quantile-state representation, and 720 ordered hourly words form the 30-day document supplied to the long-horizon model. The classifier fuses quantile states with dynamic covariates, channel-level static metadata, and a 168-hour latent-history stream using gated residual processing, causal recurrent encoding, and metadata-conditioned cross-modal attention. A bounded instability-aware signal derived from sustained one-word-ahead prediction-error divergence provides auxiliary memory modulation at the longest horizon. Evaluation uses a machine-disjoint 43/14/15 train/validation/test allocation across 72 machines in nine manufacturing facilities. At 30 days, TQRNN30d achieves 79.97% F1, 80.18% recall, 81.82% precision, 82.39% accuracy, and 0.820 ROC-AUC. It leads all 18 evaluated baselines at the 7-, 14-, and 30-day fixed-threshold comparisons, with the largest F1 advantage at 14 days. The results support held-out-machine performance within the observed homogeneous nine-facility fleet, but do not establish unseen-site, cross-equipment, or cross-sector generalisation.
Chinese Translation
长时程预测性维护要求模型能够以天而非小时为单位的规划窗口内,将缓慢演化的性能退化与正常运行工况变化区分开来。本文评估了显式条件分位数表示能否为该问题提供一个有信息量的分类器接口。所提出的TQRNN30d框架将双阶段分位数回归神经网络(QRNN)特征提取器与多流时间融合分类器相结合。每小时包含81个通道的机器行为信息被映射为一个324维的分位数状态表示,720个按时间排序的每小时词构成了提供给长时程模型的30天文档。分类器通过门控残差处理、因果循环编码以及基于元数据条件的跨模态注意力机制,将分位数状态与动态协变量、通道级静态元数据以及168小时潜在历史流进行融合。一个由持续的单步预测误差发散导出的有界不稳定性感知信号,在最长时间尺度上提供辅助的记忆调制。评估采用机器不相交的43/14/15训练/验证/测试划分,覆盖九个制造工厂的72台机器。在30天预测时,TQRNN30d取得了79.97%的F1值、80.18%的召回率、81.82%的精确率、82.39%的准确率以及0.820的ROC-AUC。在7天、14天和30天固定阈值比较中,其性能领先全部18个对比基线,其中在14天时F1优势最大。这些结果支持在所观测的同质化九厂区机队内的留出机器性能,但并未证明在未见厂区、跨设备或跨行业场景下的泛化能力。
cs.AI / 50 / 2609.04850

ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults

ElderBench:面向老年人的自主移动智能体基准测试
Zhan, Weide, Shaqu, Qumu, Liu, Yuanqing, Zhang, Peng, Liu, Jiahao, Lam, Kam Him, Gu, Ning, Hu, Zhan, Lu, Tun
Abstract
While autonomous mobile agents hold great potential for assisting older adults with smartphone usage, existing GUI benchmarks mainly rely on explicit, goal-oriented instructions and rarely capture the naturally occurring language patterns of older users, such as indirect speech, referential ambiguity, and under-specified requests. This mismatch between benchmark instructions and real-world elderly interactions may hinder reliable agent deployment. To address this gap, we present ElderBench, the first benchmark for evaluating mobile GUI agents in authentic elderly-oriented scenarios. ElderBench is constructed from 249 naturally elicited smartphone tasks collected from older adults across 20 applications. We first characterize the linguistic divergence between elderly instructions and existing GUI benchmark instructions from syntactic, semantic, and pragmatic perspectives. We then evaluate mainstream GUI agents and Vision-Language Models under both online and offline settings, revealing substantial performance degradation when handling elderly-oriented instructions. Through controlled instruction normalization, failure analysis, and fine-grained linguistic feature analysis, we further identify how elderly-specific language patterns contribute to agent failures. Our findings provide actionable design insights toward more adaptive, interpretable, and age-inclusive GUI agents for older adults.
Chinese Translation
尽管自主移动智能体在辅助老年人使用智能手机方面具有巨大潜力,但现有的图形用户界面(GUI)基准测试主要依赖明确的、目标导向的指令,很少捕捉老年用户自然产生的语言模式,例如间接言语、指代模糊以及信息不完整的请求。基准测试指令与真实老年用户交互之间的这种不匹配可能阻碍智能体的可靠部署。为填补这一空白,我们提出了 ElderBench,这是首个在真实的面向老年人的场景中评估移动 GUI 智能体的基准测试。ElderBench 基于从老年人中收集的、覆盖 20 个应用程序的 249 个自然引导获得的智能手机任务构建而成。我们首先从句法、语义和语用三个层面刻画了老年用户指令与现有 GUI 基准测试指令之间的语言差异。随后,我们在在线和离线两种设置下评估了主流 GUI 智能体和视觉语言模型(Vision-Language Models),发现在处理面向老年人的指令时性能显著下降。通过受控的指令规范化、失败分析以及细粒度的语言特征分析,我们进一步识别了老年人特有的语言模式如何导致智能体失败。我们的研究结果为构建更具适应性、可解释性且对老年人友好的 GUI 智能体提供了可操作的设计启示。
cs.AI / 51 / 2609.04859

MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models

MM-IFEval-Pro:一个面向视觉语言模型指令遵循的多语言且抗攻击的基准测试
Xiao, Changming, Ni, Zhenliang, He, Jinhui, Shu, Han, Hu, Jie
Abstract
As vision-language models (VLMs) rapidly advance in image understanding, cross-modal reasoning, and complex instruction execution, instruction-following capability has become a key indicator of their reliability and practicality. However, existing multimodal instruction-following benchmarks still suffer from limited language coverage and insufficient adversarial safety scenarios, making them inadequate for evaluating real-world multilingual and safety-sensitive settings. To address these gaps, we present MM-IFEval-Pro, a multimodal instruction-following benchmark covering Chinese and English tasks as well as diverse instruction hijacking cases. MM-IFEval-Pro includes 4 major task categories and 24 subcategories and 8 instruction categories with 52 subcategories, with each sample containing an average of 3.0 constraints to realistically simulate complex instruction scenarios. We further construct a reinforcement-learning training set enriched with Chinese and adversarial instructions, which significantly improves model performance on MM-IFEval-Pro and transfers effectively to other mainstream multimodal benchmarks, demonstrating strong cross-task and cross-language generalization.
Chinese Translation
随着视觉语言模型(VLM)在图像理解、跨模态推理和复杂指令执行方面的快速发展,指令遵循能力已成为衡量其可靠性与实用性的关键指标。然而,现有多模态指令遵循基准测试仍存在语言覆盖有限、对抗性安全场景不足等问题,难以充分评估真实世界中的多语言与安全敏感场景。为弥补这些不足,我们提出了MM-IFEval-Pro,一个涵盖中英文任务以及多样化指令劫持案例的多模态指令遵循基准测试。MM-IFEval-Pro包含4大类任务、24个子类,以及8类指令及其52个子类,每个样本平均包含3.0个约束条件,以真实地模拟复杂指令场景。我们进一步构建了一个富含中文指令和对抗性指令的强化学习训练集,该训练集显著提升了模型在MM-IFEval-Pro上的性能,并能有效迁移到其他主流多模态基准测试中,展现出强大的跨任务与跨语言泛化能力。
cs.AI / 52 / 2609.04864

MZ-Rain: Moisture-Budget-Guided Zero-Inflated Model for Station-Level Precipitation Nowcasting

MZ-Rain:面向站点级降水临近预报的水分收支引导零膨胀模型
Zhang, Yifang, Xiong, Shengwu, Wang, Henan, Yin, Wenjie, Zhang, Yuqiang, Zhou, Chen, Chen, Hua, Zhao, Qile, Duan, Pengfei
Abstract
Accurate station-level precipitation nowcasting is critical for agriculture, water resource management, and disaster prevention, which typically is formulated as a time series forecasting problem. However, conventional time-series modeling techniques face two major challenges in addressing station-level precipitation nowcasting: (1) Lack of Physics-Guided Modeling}, where meteorological variables are treated as a homogeneous set without accounting for their distinct roles in precipitation formation, leads to predictions that deviate from the physical processes governing precipitation. (2) Severe zero inflation in precipitation, where dry intervals dominate the dataset, obscuring meaningful precipitation patterns and complicating the predictive modeling. To address these challenges, we propose \textbf{MZ-Rain}, a moisture-budget-guided zero-inflated sLSTM framework for station-level precipitation nowcasting. Guided by the moisture budget equation, MZ-Rain decomposes the precipitation formation process into process-specific pathways corresponding to moisture storage, moisture transport, surface evaporation, and precipitation persistence, and captures their temporal evolution through dedicated sLSTM branches. To account for the zero-inflated nature of precipitation, MZ-Rain introduces an adaptive Tweedie modeling strategy that adaptively modulates the rainfall mean while jointly learning precipitation occurrence as an auxiliary task, enabling the model to better balance dry-wet discrimination and quantitative precipitation estimation. Extensive experiments across diverse geographical and climatic regimes demonstrate that MZ-Rain consistently outperforms strong baselines on multiple evaluation metrics, including CSI, FAR, MSE, and MAE. In particular, the model exhibits superior skill in forecasting heavy precipitation events, while benefiting from physically grounded process modeling.
Chinese Translation
精确的站点级降水临近预报对农业、水资源管理和灾害防治至关重要,该问题通常被建模为时间序列预测任务。然而,传统时间序列建模方法在处理站点级降水临近预报时面临两大挑战:(1)缺乏物理引导建模:气象变量被视为同质集合,未考虑其在降水形成过程中的不同作用,导致预测结果偏离支配降水形成的物理过程。(2)降水数据存在严重的零膨胀问题:无雨时段在数据集中占主导地位,掩盖了有意义的降水模式,增加了预测建模的难度。为应对这些挑战,我们提出了MZ-Rain,一个面向站点级降水临近预报的水分收支引导零膨胀sLSTM框架。在水分收支方程的引导下,MZ-Rain将降水形成过程分解为对应水分储存、水分输送、地表蒸发和降水持续性的特定过程路径,并通过专用的sLSTM分支捕捉其时间演化。为处理降水的零膨胀特性,MZ-Rain引入了一种自适应Tweedie建模策略,在自适应调节降雨均值的同时,将降水发生与否的判断作为辅助任务进行联合学习,使模型能够更好地平衡干湿判别与定量降水估计。在多种地理和气候条件下的广泛实验表明,MZ-Rain在CSI、FAR、MSE和MAE等多项评估指标上持续优于强基线模型。尤其值得一提的是,得益于基于物理过程的建模,该模型在强降水事件预报方面表现出卓越的能力。
cs.AI / 53 / 2609.04865

CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution

CoSkill:面向层次化技能演化的推理与元技能智能体联合强化学习
Feng, Jinyuan, Li, Dongmin, Chen, Yiqun, Gao, Yang, Chen, Xing, Wang, Huimu, Pu, Zhiqiang
Abstract
Skill libraries improve the sample efficiency of agentic reinforcement learning (RL) by enabling large language model (LLM) agents to reuse procedural knowledge. Yet existing paradigms exhibit structural shortcomings: they either decouple skill evolution from policy optimization or instantiate meta-skills as fixed workflows. Both treat skills as passive objects to be managed, limiting the flexible evolution of skills and their co-adaptation with the reasoning agent. To address the limitations, we propose CoSkill, a unified multi-agent RL framework that recasts the static meta-skill workflow as a learnable Meta-Skill Agent and jointly trains it with a Reasoning Agent over a hierarchical skill library. By modeling the Reasoning and Meta-Skill Agents as a cooperative team sharing a single backbone, CoSkill enables end-to-end co-adaptation: the Reasoning Agent conditions its actions on a retrieved task skill and step skills selected from its child set, while its task performance guides the Meta-Skill Agent in refining those step skills. Experiments on ALFWorld and WebShop show that CoSkill substantially outperforms prior skill-based and RL baselines, achieving success rates of 98.4% and 90.6%, respectively (+3.5 and +6.2 pp). As shown in Figure 1, CoSkill achieves superior early-stage sample efficiency, asymptotic performance, and wall-clock efficiency. Our code is available at https://github.com/jinyuan-cookie/CoSkill.
Chinese Translation
技能库通过使大语言模型(LLM)智能体能够复用程序性知识,提升了智能体强化学习(RL)的样本效率。然而,现有范式存在结构性缺陷:它们要么将技能演化与策略优化解耦,要么将元技能实例化为固定的工作流。二者均将技能视为被管理的被动对象,限制了技能的灵活演化及其与推理智能体的协同适应。为解决这些局限性,我们提出CoSkill,一个统一的多智能体强化学习框架,它将静态的元技能工作流重构为可学习的元技能智能体(Meta-Skill Agent),并在层次化技能库上与推理智能体(Reasoning Agent)进行联合训练。通过将推理智能体与元技能智能体建模为共享同一骨干网络的协作团队,CoSkill实现了端到端的协同适应:推理智能体基于检索到的任务技能以及从其子集合中选出的步骤技能来决定其动作,同时其任务表现又引导元技能智能体优化这些步骤技能。在ALFWorld和WebShop上的实验表明,CoSkill显著优于此前基于技能的和强化学习的基线方法,分别取得了98.4%和90.6%的成功率(+3.5和+6.2个百分点)。如图1所示,CoSkill在早期样本效率、渐近性能和时钟时间效率方面均表现优越。我们的代码已发布于https://github.com/jinyuan-cookie/CoSkill。
cs.AI / 54 / 2609.04866

LLM-Assisted Behavioural and Scenario Augmentation for Agent-Based Energy Adoption Models

基于大语言模型辅助的行为与情景增强的智能体能源采纳模型
Faiud, Iias, Khaleghy, Hossein, Schukat, Michael, Mason, Karl
Abstract
Recent advances in large language models (LLMs) create opportunities to enrich simulation-based energy policy analysis, particularly by supporting structured behavioural assumptions and exploratory techno-economic scenarios. However, directly replacing adoption models with LLM reasoning raises concerns regarding interpretability, reproducibility, and behavioural validity. This paper proposes a hybrid framework for LLM-assisted specification design, integrating bounded behavioural rubrics and structured scenario specifications into a calibrated agent-based model (ABM) of solar photovoltaic (PV) adoption by Irish dairy farms. The proposed approach preserves the original techno-economic adoption mechanism while augmenting it with bounded behavioural modulation and scenario-driven uncertainty analysis. Behavioural effects are represented through interpretable conservative, balanced, and optimistic rubrics, while future policy and market conditions are explored through fixed, rule-validated scenario specifications. Experimental results across multiple policy settings, Monte Carlo worlds, and random seeds demonstrate stable and economically plausible behaviour, with adoption outcomes remaining bounded and monotonic across behavioural regimes. The framework achieves up to approximately 13% behavioural adoption increase relative to the corresponding logistic case without producing unstable or unrealistic saturation dynamics. The results demonstrate that LLM-assisted specifications can be integrated into calibrated energy ABMs in a controlled, reproducible, and policy-relevant manner.
Chinese Translation
大语言模型(LLM)的最新进展为丰富基于仿真的能源政策分析创造了机遇,尤其是在支持结构化行为假设和探索性技术经济情景方面。然而,直接用LLM推理替代采纳模型会引发关于可解释性、可复现性和行为有效性的担忧。本文提出了一种LLM辅助的规范设计混合框架,将有界行为准则和结构化情景规范集成到一个经过校准的爱尔兰奶牛场太阳能光伏(PV)采纳智能体模型(ABM)中。所提出的方法在保留原有技术经济采纳机制的同时,通过有界行为调节和情景驱动的不确定性分析对其进行增强。行为效应通过可解释的保守、均衡和乐观三类准则来表示,而未来的政策与市场条件则通过固定的、经规则验证的情景规范进行探索。在多种政策设定、蒙特卡洛世界和随机种子下的实验结果表明,模型行为稳定且在经济上合理,采纳结果在各行为情景下保持有界且单调。相对于对应的逻辑斯蒂(logistic)情形,该框架实现了高达约13%的行为性采纳提升,且未产生不稳定或不现实的饱和动态。结果表明,LLM辅助的规范能够以受控、可复现且具有政策相关性的方式集成到经过校准的能源ABM中。
cs.AI / 55 / 2609.04869

From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents

从交互轨迹到持久技能:计算机使用代理的在线演化
Hu, Longtao, Liang, Xiao, Zhu, Linchao
Abstract
Computer-use agents can execute increasingly complex tasks in graphical interfaces, but their interaction experience is typically transient: procedural knowledge acquired from one rollout is not systematically retained, refined, and reused in later tasks. Existing skill libraries provide external procedural knowledge, yet their incremental value over the same agent operating without skills, as well as their longitudinal dynamics under repeated interaction, remain insufficiently characterized. We present an online skill-evolution framework that converts interaction trajectories and evaluator feedback into a persistent, versioned library of reusable procedures. Each iteration executes against a frozen library snapshot, and evidence-guided skill updates become available in subsequent iterations without changing model parameters. We compare the full evolving-library system with a configuration-matched empty-library control across four OSWorld application domains under the same fixed action-generation and GUI-grounding stack, task sets, and iteration horizons. Following a five-iteration empty-library warm-up, Full attains a higher post-warm-up mean evaluator score in all four observed domain runs, with mean differences ranging from 5.7 to 18.6 percentage points and domain-dependent temporal stability. In GIMP, provenance-aware analysis reveals retrieval across task-of-origin boundaries and revision churn, where repeated accepted edits fail to recover the originating task. These findings characterize evolving skill libraries as auditable, shared procedural memory that can improve a fixed computer-use stack, while showing that their benefits are conditional and repeated revision does not guarantee recovery. Code is released at https://github.com/LongtaoHu/Skill-Evo4GUI.
Chinese Translation
计算机使用代理(Computer-use agents)能够在图形界面中执行日益复杂的任务,但其交互经验通常是短暂的:从一次执行(rollout)中获得的过程性知识无法被系统地保留、精炼并在后续任务中复用。现有技能库提供了外部过程性知识,但它们相对于同一代理在无技能条件下运行的增量价值,以及在重复交互下的纵向动态特性,仍未得到充分刻画。我们提出了一种在线技能演化框架,将交互轨迹和评估器反馈转化为持久、版本化的可复用过程性知识库。每次迭代均针对冻结的库快照执行,且基于证据的技能更新在后续迭代中可用,而无需改变模型参数。我们在四个 OSWorld 应用领域中,将完整的演化技能库系统与配置相同的空技能库对照组进行比较,二者使用相同的固定动作生成与 GUI 定位组件、任务集和迭代次数。经过五次迭代的空库预热后,Full 系统在所有四个观察到的领域运行中均获得了更高的预热后平均评估器得分,平均差异介于 5.7 至 18.6 个百分点之间,且呈现出依赖于领域的时间稳定性。在 GIMP 领域中,基于来源(provenance-aware)的分析揭示了跨任务来源边界的技能检索以及修订抖动(revision churn)现象,即反复被接受的编辑仍未能恢复原始任务的表现。这些发现将演化技能库刻画为一种可审计的、可共享的过程性记忆,能够提升固定的计算机使用系统的表现,同时表明其收益是有条件的,且反复修订并不保证恢复效果。代码已发布于 https://github.com/LongtaoHu/Skill-Evo4GUI。
cs.AI / 56 / 2609.04870

CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player Games

CHAMP:面向在线多人游戏匹配与预测的跨域混合架构
Wang, Kai, Fan, Ge, Zhang, Chaoyun, Jiang, Yuyang, Liu, Yuze
Abstract
Multiplayer Online Battle Arena (MOBA) games rely on matchmaking to maintain competitive balance. Our prior work, CUPID, framed matchmaking as an assignment re-optimization problem and showed that a single-mode win-rate predictor can meaningfully rebalance teams. However, deploying such a system across diverse player populations exposes three practical bottlenecks: most queueing players lack sufficient in-mode match history (cold start), skill distributions shift drastically across rank tiers (distribution inconsistency), and extreme skill segments are severely data-starved. We present CHAMP, a cross-domain matchmaking framework that resolves these deployment bottlenecks. To address data sparsity and cold starts, CHAMP replaces the target-mode-only player profile with a hybrid domain feature collection: a timestamp-ordered cross-mode short-term sequence whose slices are annotated with target-domain features, plus per-mode breakdowns of long-term, real-time and team statistics. We further propose the Domain-Aware Win-rate Network (DAWN): a Domain-aware Knowledge Extractor (DAKE) compiles target-mode attributes into learnable representations that feed Domain-Aware Temporal/Spatial/Permutation OmniNet Encoders (DATOE/DASOE/DAPOE), so that mode-conditioned representations and per-mode debiasing are learned jointly inside a single shared network. Online, one trained DAWN serves every supported mode, with per-mode position-satisfaction thresholds as the only mode-specific knob. Offline, DAWN achieves 67.73% win-rate prediction accuracy, outperforming all evaluated attention and sequence baselines. Online A/B tests across the entire League ladder of a large-scale MOBA game, from novice players up to the top-expert players served by Elite Mode, demonstrate consistent drops in imbalanced matches. For lower-tier players, CHAMP reduces the 5-minute kill crushing rate by up to 20.73%.
Chinese Translation
多人在线战术竞技(MOBA)游戏依赖匹配系统来维持竞技平衡。我们先前的工作 CUPID 将匹配建模为分配再优化问题,并证明了单一模式的胜率预测器能够有效实现队伍的再平衡。然而,将此类系统部署于多样化的玩家群体时,会暴露出三个实际瓶颈:大多数排队玩家缺乏足够的本模式对局历史(冷启动问题)、技能分布在各段位之间剧烈变化(分布不一致问题)、以及极高和极低技能段位的玩家数据严重匮乏。我们提出 CHAMP,一个能够解决上述部署瓶颈的跨域匹配框架。为解决数据稀疏与冷启动问题,CHAMP 用混合域特征集合取代仅基于目标模式的玩家画像:一个按时间戳排序的跨模式短期序列(其切片附带目标域特征标注),以及分模式统计的长期、实时与团队数据。我们进一步提出域感知胜率网络(Domain-Aware Win-rate Network,DAWN):域感知知识提取器(Domain-aware Knowledge Extractor,DAKE)将目标模式属性编译为可学习的表示,输入域感知时间/空间/排列全连接网络编码器(DATOE/DASOE/DAPOE),从而使模式条件化表示与分模式去偏在单一共享网络内联合学习。在线上,一个训练好的 DAWN 可服务于所有支持的模式,仅需分模式的位置满足度阈值作为唯一的模式特定调节参数。离线实验中,DAWN 达到 67.73% 的胜率预测准确率,优于所有受评估的注意力机制和序列模型基线。在某个大型 MOBA 游戏的整个《英雄联盟》天梯(从新手玩家到由精英模式服务的高端专家玩家)上进行的线上 A/B 测试表明,不平衡对局的数量持续下降。对于低段位玩家,CHAMP 将 5 分钟击杀碾压率最多降低 20.73%。
cs.AI / 57 / 2609.04871

AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems

AutoLR:自动化工业推荐系统中从研究到上线评审的路径
Zhang, Qi, Chen, Yanlin, Xiao, Wenchao
Abstract
Improving an industrial recommender is an iterative research-and-engineering process rather than a direct path from idea to deployment. In \textbf{DASHEN, NetEase's gaming-community app}, algorithm engineers typically identify promising directions from research papers, technical reports, and prior production experiments; reproduce or adapt the underlying methods; implement them in the production codebase; and evaluate the resulting models through training and offline experiments. Promising candidates are then advanced to online A/B tests, and those demonstrating robust gains are submitted to Launch Review---the internal gate for full-traffic rollout. Large language models (LLMs) can assist with individual stages of this workflow, but the overall process remains human-dependent without a harness that can reliably coordinate them across long-running, often multi-day experimental cycles. We present \textbf{AutoLR}, initially built as \textbf{Auto Launch Review} and later extended upstream into an autonomous research-to-launch harness. AutoLR combines three system mechanisms: a \textbf{multi-expert council} that debates and adversarially reviews proposals; a \textbf{deterministic evidence-weighted exploration--exploitation selector} that allocates a limited trial budget across candidate directions and uses Council reranking; and a layered knowledge system that combines external research, production-system knowledge, and DASHEN-specific domain knowledge---such as game communities, player characteristics, and content-interaction patterns---with posterior evidence from configurations, patches, logs, failures, and offline outcomes. LLM agents perform semantic reasoning and code generation, while deterministic controllers retain authority over execution, metric extraction, guardrails, and persistent state transitions.
Chinese Translation
改进一个工业推荐系统是一个研究与工程相结合的迭代过程,而非从想法到部署的直接路径。在DASHEN(网易的游戏社区应用)中,算法工程师通常从研究论文、技术报告和既往生产实验中发现有前景的方向;复现或改造其中的底层方法;将其实现到生产代码库中;并通过训练和离线实验评估所得模型。有前景的候选方案随后进入在线A/B测试,展现出稳健收益的方案将被提交至上线评审(Launch Review)——即全流量上线前的内部审核关口。大语言模型(LLM)可以辅助这一工作流程中的单个环节,但在缺乏能够可靠协调跨长周期(往往持续数天)实验流程的框架的情况下,整体过程仍然依赖人工。我们提出了AutoLR,其最初作为自动上线评审(Auto Launch Review)构建,随后向上游扩展为一个自主的从研究到上线的自动化框架。AutoLR结合了三个系统机制:一个进行辩论和对抗性审查提案的多专家委员会;一个确定性的、基于证据加权的探索-利用选择器,用于在候选方向之间分配有限的试验预算,并利用委员会重排序;以及一个分层知识系统,该系统将外部研究、生产系统知识和DASHEN特有的领域知识(如游戏社区、玩家特征和内容互动模式)与来自配置、补丁、日志、失败记录和离线结果的后验证据相结合。LLM智能体负责语义推理和代码生成,而确定性控制器则保留对执行、指标提取、安全护栏和持久化状态转移的控制权。
cs.AI / 58 / 2609.04877

MARLA: A Conceptual Scaffold for Regulatory Learning under the EU AI Act

MARLA:欧盟《人工智能法案》下监管学习的概念框架
Buscemi, Alessio, Deckenbrunnen, Tom, Hmiddou, Imane, Billi, Marco, Rubino, Livio, Ferruzza, Silvia Rizzuto, Pagani, Daniele, Rotolo, Antonino
Abstract
The EU AI Act positions regulation as part of the infrastructure for safe, trustworthy and market-ready innovation. Realising this ambition requires regulatory learning: the evidence generated during implementation must be translated into governance and legal knowledge that supports consistent interpretation, effective oversight, and adaptation as technologies evolve. Yet the actors who produce this evidence and those who rely on it operate in different professional worlds. This paper proposes MARLA (Map, Assess, Report, Learn, Adapt), a conceptual scaffold organising regulatory learning as a five-stage cycle centred on the implementation of legal requirements into socio-technical practices, situated at the Local, National and European levels of the AI Act's governance architecture. Deliberately non-prescriptive, MARLA gives technical and legal stakeholders a shared vocabulary in which each of the first three stages generates its own documentable form of regulatory learning. We illustrate the scaffold with two piloted case studies and a prospective National-to-European illustration.
Chinese Translation
欧盟《人工智能法案》(EU AI Act)将监管定位为安全、可信且具备市场化创新基础设施的一部分。实现这一愿景需要监管学习(regulatory learning):即将实施过程中产生的证据转化为治理与法律知识,以支持一致的解释、有效的监督,以及随技术演进的适应性调整。然而,产生这些证据的主体与依赖这些证据的主体处于不同的专业领域。本文提出MARLA(Map、Assess、Report、Learn、Adapt,即映射、评估、报告、学习、适应),一个将监管学习组织为五阶段循环的概念框架。该框架以将法律要求落实为社会技术实践为核心,并定位于《人工智能法案》治理架构的地方、国家和欧洲三个层面。MARLA刻意保持非规范性,为技术与法律相关方提供了一个共同的话语体系,使前三个阶段各自生成可记录的监管学习形式。我们通过两个试点案例研究和一个前瞻性的国家到欧洲层面的示例来阐释该框架。
cs.AI / 59 / 2609.04880

Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach

不确定性条件下序贯太阳能光伏政策设计的强化学习:一种基于主体的方法
Faiud, Iias, Shianifar, Jonaid, Schukat, Michael, Mason, Karl
Abstract
Designing effective and fiscally sustainable policies for solar photovoltaic (PV) adoption requires balancing adoption gains against public expenditure under uncertainty and heterogeneous decision-making. This study formulates PV policy design as a sequential decision problem and integrates reinforcement learning (RL) with a stochastic agent-based model (ABM) that simulates yearly solar PV adoption under uncertainty. A policymaker agent selects annual incentives, including capital grants, subsidised loan rates, and feed-in tariffs, over a 16-year horizon. Adoption--cost trade-offs are explored by varying policy preferences within a scalarised reward framework. Policies are learned using PPO, SAC, and TD3 and evaluated under stochastic simulation. The results show that this approach produces a clear trade-off structure: the highest-adoption policy (TD3, $w_{\text{cost}}=0.5$) achieves approximately 4,145 adopters at a cost of EUR 41.73 million, while the lowest-cost policy (PPO, $w_{\text{cost}}=2.0$) reduces expenditure to EUR 7.27 million with 2,682 adopters. The balanced policy (PPO, $w_{\text{cost}}=1.6$) achieves 3,495 adopters at a cost of EUR 22.47 million. Across algorithms, consistent trade-off patterns are observed, indicating robustness of the adoption--cost relationship. Compared with static baseline policies, the RL framework explores a broader range of policy configurations. These findings demonstrate the potential of RL as a flexible tool for adaptive policy design under uncertainty.
Chinese Translation
设计有效且财政可持续的太阳能光伏(PV)推广政策,需要在不确定性条件和异质性决策行为下,权衡推广收益与公共支出。本研究将光伏政策设计表述为一个序贯决策问题,并将强化学习(RL)与一个随机基于主体模型(ABM)相结合,以模拟不确定性条件下的年度太阳能光伏推广情况。政策制定者主体在16年的时间跨度内选择年度激励措施,包括资本补贴、优惠贷款利率和上网电价(feed-in tariffs)。通过在标量化奖励框架内调整政策偏好,探索了推广—成本之间的权衡关系。采用PPO、SAC和TD3算法学习政策,并在随机仿真下进行评估。结果表明,该方法产生了清晰的权衡结构:推广量最高的政策(TD3,$w_{\text{cost}}=0.5$)以4173万欧元的成本实现了约4145户采纳;成本最低的政策(PPO,$w_{\text{cost}}=2.0$)将支出降至727万欧元,采纳者为2682户。平衡型政策(PPO,$w_{\text{cost}}=1.6$)以2247万欧元的成本实现了3495户采纳。在所有算法中均观察到一致的权衡模式,表明推广—成本关系具有稳健性。与静态基准政策相比,RL框架探索了更广泛的政策配置空间。这些发现证明了强化学习作为不确定性条件下适应性政策设计灵活工具的潜力。
cs.AI / 60 / 2609.04894

From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments

从语言模型到世界行动系统:智能体AI在数字、社交、虚拟与物理环境中的进展与局限
Zhu, Linsen, Cai, Mengqing
Abstract
Large language models become consequential agents when surrounding systems let outputs change external state. Models now call tools, operate interfaces, delegate work, retain state, inhabit generated worlds, and control robots or laboratory equipment. Such advances are often narrated as one march toward autonomy, conflating model competence, system integration, persistence, and safe authority. This critical review synthesizes primary research and official technical specifications available by 31 August 2026. We organize the evidence along delegated authority, temporal persistence, and environmental coupling, while separating model, harness, and environment. Within the evidence examined, action-interface expansion is documented more convincingly than robust completion, recovery, authorization, or independent verification. Model Context Protocol and Agent2Agent improve interoperability but do not establish trustworthy delegation; multi-agent organization adds specialization alongside cost and correlated failure. Persistent simulations and world models support training and planning but do not themselves demonstrate agency; robotics and self-driving laboratories establish bounded feasibility rather than unattended open-world reliability. We propose justified delegation as an analytical and normative heuristic, not an observed law or certified score: expand action scope only where evidence supports provenance, bounded authority, failure detection, safe recovery, and calibrated human control. This framing yields a research agenda for coupled model-harness evaluation, capability-based permissions, durable state, cross-agent accountability, and staged physical validation.
Chinese Translation
当外部系统使语言模型的输出能够改变外部状态时,大型语言模型便成为具有实际影响的智能体。当前的模型已能调用工具、操作界面、委派任务、保持状态、驻留于生成的虚拟世界中,并控制机器人或实验室设备。此类进展常被叙述为一场迈向自主性的单一进程,从而混淆了模型能力、系统集成、持久性与安全授权等不同层面。本批判性综述综合了截至2026年8月31日可获得的一手研究与官方技术规范。我们按照委派权限、时间持久性和环境耦合三个维度组织证据,同时区分模型、运行框架(harness)与环境。在所考察的证据中,行动接口的扩展比稳健的任务完成、故障恢复、授权或独立验证得到了更有说服力的证实。模型上下文协议(Model Context Protocol)和智能体间协议(Agent2Agent)改善了互操作性,但并未建立可信赖的委派机制;多智能体组织在带来专业化分工的同时也增加了成本与关联性故障风险。持久性仿真与世界模型可支持训练与规划,但其本身并不能证明智能体的存在;机器人技术与自动驾驶实验室确立的是有边界的可行性,而非无人值守的开放世界可靠性。我们提出“有依据的委派”(justified delegation)作为分析性与规范性的启发式原则,而非已被观察到的规律或经认证的评分:只有在证据能够支持来源可溯、权限有界、故障检测、安全恢复以及经过校准的人类控制时,才扩展行动范围。这一框架为耦合的模型-框架评估、基于能力的权限管理、持久状态、跨智能体问责以及分阶段的物理验证提供了研究议程。
cs.AI / 61 / 2609.04915

Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing

基于在线最大成员聚类与原子感知打包的紧凑记忆LLM智能体
Geng, Jiahe, Wang, Jinpeng, Yuan, Kun
Abstract
Many long-horizon LLM deployments face tight prompt budgets: latency, cost, and context limits make full-context prompting impractical as interaction length grows. The key question is then not raw recall alone, but which memory design gives the best quality--token trade-off in the compact-memory regime. We present \textbf{RSM-full}, an online clustered-memory pipeline designed for a strong quality--token Pareto point. RSM-full combines two design choices: a cosine-gated \emph{max-member merge} write rule and an atom-aware grouped context packer. On AMA-Bench, our primary compact-memory benchmark, it reaches $83%$ of Full-Context quality at $32%$ of the token cost at a $4$k budget; under four-seed averaging it beats the closest streaming-clustered baseline (Online K-Means) by $+3.5$--$6.0$,pp ($p{<}.001$) across the whole ${\sim}2.6$k--${\sim}5$k regime. Three-seed ablations show most of this gain comes from the merge rule ($+5.7$,pp over Online K-Means and matched-$\tau$ DP-means) and the grouped packer ($+5.0$,pp over flat concatenation). The pattern reproduces on RealMem, an independent long-horizon persona-memory benchmark: RSM-full improves on Budget-RAG ($+0.69$,pp, $p{=}.006$), is on par with BM25-RAG (paired $\Delta{=}{+}0.27$,pp, $p{=}.47$; we do \emph{not} claim BM25 equivalence in the equivalence-test sense), and significantly outperforms Streaming-Proto ($+2.97$,pp) and the closest reproduced 2025 agentic-memory baseline A-MEM ($+1.65$,pp, $p{<}.001$). Across benchmarks the message is consistent: under tight budgets, compact-memory performance is driven mainly by how streaming memories are merged and how retrieved content is assembled. Overall, RSM-full is most useful when answeroughly $2k$--$5k$ prompt tokens, where itdefines a strong compact-memory Pareto point; higher-token baselines remain stronger outside this regime.
Chinese Translation
许多长周期LLM部署面临严格的提示词预算限制:随着交互长度增长,延迟、成本和上下文限制使得全上下文提示变得不切实际。因此关键问题不再仅仅是原始召回能力,而是在紧凑记忆机制下,哪种记忆设计能提供最佳的质量- token权衡。我们提出了RSM-full,一个在线聚类记忆流水线,旨在实现一个强大的质量- token帕累托最优点。RSM-full结合了两个设计选择:基于余弦门控的最大成员合并(max-member merge)写入规则,以及原子感知的分组上下文打包器(grouped context packer)。在AMA-Bench(我们的主要紧凑记忆基准)上,在4k预算下,该方法以32%的token成本达到了全上下文质量的83%;在四种子平均下,它在整个约2.6k至约5k的范围内以+3.5至6.0个百分点(p<.001)的优势超越最接近的流式聚类基线(Online K-Means)。三种子消融实验表明,大部分增益来自合并规则(相比Online K-Means和匹配-τ的DP-means提升+5.7个百分点)以及分组打包器(相比扁平拼接提升+5.0个百分点)。该模式在RealMem(一个独立的长周期人格记忆基准)上得到了复现:RSM-full优于Budget-RAG(+0.69个百分点,p=.006),与BM25-RAG相当(配对差值Δ=+0.27个百分点,p=.47;我们并不在等效性检验的意义上声称与BM25等效),并显著优于Streaming-Proto(+2.97个百分点)以及最接近的可复现2025年智能体记忆基线A-MEM(+1.65个百分点,p<.001)。跨基准的结果传递出一致的信息:在严格预算下,紧凑记忆的性能主要由流式记忆的合并方式以及检索内容的组装方式决定。总体而言,RSM-full在提示词token约为2k至5k时最为有用,在这一范围内它定义了一个强大的紧凑记忆帕累托最优点;在该范围之外,更高token的基线仍然更强。
cs.AI / 62 / 2609.04917

Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing

人工智能在股票与加密货币市场中的应用:进展、盈利性证据及自动化投资的局限性
Zhu, Linsen, Cai, Mengqing
Abstract
Artificial intelligence (AI) now supports investment workflows from data and prediction through research, portfolios, execution, and tool use. Technical capability, however, is not evidence of investment profitability. This critical state-of-the-art review examines public research available through 31 August 2026 on listed equities, exchange-traded funds, centralized crypto spot, perpetual futures, and on-chain markets. We organize evidence with an alpha-translation chain: point-in-time information must yield a stable signal, feasible positions, executable orders, and risk-adjusted returns after costs. Across machine learning, time-series foundation models, financial language models, reinforcement learning, and agents, the examined record shows real but mainly upstream progress in prediction, text processing, portfolio design, and workflow integration. Evidence is thinner for durable net performance. Temporal contamination, repeated selection, survivorship, weak benchmarks, implementation costs, venue mechanics, and capacity can break translation to net alpha. Strong historical results coexist with predictor decay, corrected look-ahead failures, mixed prospective evidence, and few audited live-capital records. Crypto adds informative state but requires separate treatment of spot, perpetual, and decentralized cash flows and execution. Within the public evidence examined here, no general AI architecture is shown to deliver persistent, cross-regime, capacity-aware net alpha. More credible claims require point-in-time data and models, decision-aligned objectives, joint portfolio--execution evaluation, controlled adaptation, prospective tests, and authority-matched governance. These conditions can improve evidence and implementation; they do not guarantee profit.
Chinese Translation
人工智能(AI)如今已支撑从数据与预测到研究、投资组合、执行和工具使用的整个投资工作流程。然而,技术能力并不等同于投资盈利能力的证据。本文对截至2026年8月31日可获取的公开研究进行了批判性综述,涵盖上市股票、交易所交易基金(ETF)、中心化加密货币现货市场、永续合约以及链上市场。我们以"阿尔法转化链"组织证据:时点信息必须能够产生稳定的信号、可行的头寸、可执行的订单,以及扣除成本后的风险调整收益。纵观机器学习、时间序列基础模型、金融语言模型、强化学习和智能体等领域,所考察的研究成果显示,在预测、文本处理、投资组合设计和工作流整合等上游环节取得了实质性进展,但关于持久净表现的证据较为薄弱。时间污染、重复选择、幸存者偏差、弱基准、实施成本、交易场所机制以及容量限制都可能破坏向净阿尔法的转化。强劲的历史结果与预测因子衰减、被修正的前视偏差问题、不一致的前瞻性证据以及极少经过审计的实盘资金记录并存。加密货币市场提供了额外的信息状态,但需要对现货、永续合约以及去中心化现金流和执行分别处理。在本文考察的公开证据范围内,尚无任何通用AI架构被证明能够提供持续的、跨市场状态的、具备容量意识的净阿尔法。更可信的结论需要时点数据与模型、与决策一致的目标函数、投资组合与执行的联合评估、受控适应、前瞻性测试以及与监管机构匹配的治理机制。这些条件能够改善证据质量和实施效果,但并不能保证盈利。
cs.AI / 63 / 2609.04931

Solving Hard XAI Queries Based on a Compiled Dual-Rail Encoding

基于编译的双轨编码求解困难的XAI查询
Ledaguenel, Arthur, Capelli, Florent, Lagniez, Jean-Marie
Abstract
The widespread adoption of artificial intelligence (AI) within real-world applications has raised a lot of concerns regarding their trustworthiness, especially in critical applications. The field of eXplainable AI (XAI) has emerged with the objective of providing explanations to the users about the decisions made by AI systems. Several explanations for boolean classifiers have been introduced in the literature, including abductive and contrastive explanations, each giving a different insight on the decision of the classifier. However, computing an explanation for a decision of a boolean classifier is a hard problem in general. One way to deal with this complexity is to rely on a compiled representation of the classifier for which each explanation can be computed efficiently. Unfortunately, we prove in this paper that several classes of abductive explanations, remain hard to compute even for Ordered Binary Decision Diagrams, one of the most tractable subsets of the knowledge compilation map. Included in such classes are shorter abductive explanations or abductive explanations that include the explainee's preferences. To recover the benefits of working with compiled representations, we show that a proper representation of the dual-rail encoding of the classifier can be used to compute efficiently these classes of explanations.
Chinese Translation
人工智能(AI)在现实世界应用中的广泛采用引发了人们对其可信赖性的诸多担忧,尤其是在关键应用中。可解释人工智能(XAI)领域应运而生,其目标是为用户解释AI系统所做出的决策。文献中已提出了针对布尔分类器的多种解释,包括溯因解释和对比解释,它们各自从不同角度揭示分类器的决策。然而,为布尔分类器的某个决策计算解释在一般情况下是一个困难问题。应对这种复杂性的一种方法是依赖分类器的编译表示,从而使每种解释都能被高效计算。遗憾的是,本文证明了若干类溯因解释即使在有序二叉决策图(Ordered Binary Decision Diagrams,知识编译图谱中最易处理的子集之一)上也难以计算。这些类别包括更短的溯因解释或包含被解释者偏好的溯因解释。为了重新获得使用编译表示的优势,我们证明可以通过对分类器双轨编码的适当表示来高效计算这些类别的解释。
cs.AI / 64 / 2609.04962

Why We Care About Understanding: Competence through Predictive Compression

我们为何关注理解:通过预测性压缩实现能力
Queloz, Matthieu, Beckmann, Pierre
Abstract
What is the relation between understanding and compression, and why does human understanding take such a heavily compressed form? Across information theory, machine learning, and AI research, a substantial tradition identifies understanding with compression-a thought captured in Gregory Chaitin's dictum that "comprehension is compression." Philosophers, by contrast, have characterized understanding in terms of grasping connections, giving explanations, and handling novelty. This paper bridges the two pictures through three interlocking theses. The first concerns the concept of understanding: it serves as an efficient proxy for a distinctive form of robust competence, enabling us to identify whom to trust and whom to learn from. The second concerns the state of understanding: to understand a domain is to possess a mental model of its relational structure that enables prediction, and what enables prediction enables compression, because what becomes predictable need not be stored separately. Compression is therefore not identical with comprehension, but its representational shadow. The third concerns the characteristically human form of understanding: the fiduciary and transmission functions highlighted by the first thesis impose pressures of demonstrability and transmissibility that drive human understanding toward principled simplicity. The resulting framework explains both the appeal and the limits of compressionist accounts of understanding while shedding light on the inscrutability of AI systems.
Chinese Translation
理解与压缩之间的关系是什么?为什么人类的理解会呈现出如此高度压缩的形式?在信息论、机器学习和人工智能研究中,一个重要的传统将理解等同于压缩——这一思想体现在Gregory Chaitin的名言“理解即压缩”中。相比之下,哲学家们则从把握关联、给出解释和应对新颖性等方面来刻画理解。本文通过三个相互关联的论点来弥合这两种观点。第一个论点关涉理解这一概念:理解是某种独特的稳健能力的高效代理,使我们能够识别应当信任谁、应当向谁学习。第二个论点关涉理解的状态:理解一个领域就是拥有一个关于其关系结构、能够支持预测的心理模型,而支持预测的机制也能支持压缩,因为可被预测的内容无需单独存储。因此,压缩并不等同于理解,而是理解的表征性投影。第三个论点关涉人类特有的理解形式:第一个论点所强调的信任与传递功能带来了可论证性与可传递性的压力,这种压力促使人类的理解朝着有原则的简单性方向发展。由此形成的框架既解释了压缩主义理解观的吸引力与局限,也有助于阐明人工智能系统的不可解释性。
cs.AI / 65 / 2609.04978

Global to Local: Topology-Preserving Adaptive Graph Pooling via Granular-Ball

从全局到局部:基于颗粒球的拓扑保持自适应图池化
Zhao, Sen, Xu, Gaojie, Xia, Shuyin, Guan, Yifan, Liu, Yi, Wang, Yi, Wang, Wei
Abstract
Graph pooling aims to compress the graph, including both node embeddings and their underlying topological patterns, into a more compact representation. Previous works focus primarily on the overly fine-grained representation of nodes, progressively coarsening the graph by removing nodes or merging them into clusters, thus neglecting the global-to-local patterns and adaptive granularity of the graph's topological structure. In the real scenario, graphs as a whole can be considered the coarsest level of granularity, encapsulating the global topological structure, with progressively finer-grained local topological structures represented from top to bottom. This process continues until the adaptive granularity for each subdomain is reached. To this end, we propose a novel Topology-Preserving Adaptive Graph Pooling (TPAGP) method that dynamically partitions graphs into granular balls by integrating node features and topological information, enabling the generation of multi-granularity representations that effectively capture both local and global structural patterns. Additionally, we design a multi-granularity graph network model that facilitates feature interaction and optimization across different granularities, significantly enhancing performance in graph classification tasks. Experimental results demonstrate that TPAGP outperforms existing pooling methods across various benchmark datasets, effectively mitigating information loss caused by fixed-granularity strategies.
Chinese Translation
图池化旨在将图(包括节点嵌入及其底层拓扑模式)压缩为更紧凑的表示。已有工作主要关注节点的过于细粒度的表示,通过删除节点或将节点合并为簇来逐步粗化图,从而忽略了图拓扑结构的全局到局部模式以及自适应粒度。在真实场景中,图作为一个整体可被视为最粗粒度的层次,囊括了全局拓扑结构;自顶向下逐步呈现更细粒度的局部拓扑结构,这一过程持续进行,直至达到每个子域的自适应粒度。为此,我们提出了一种新颖的拓扑保持自适应图池化方法(Topology-Preserving Adaptive Graph Pooling, TPAGP),该方法通过融合节点特征与拓扑信息,将图动态划分为颗粒球(granular ball),从而生成能有效捕捉局部与全局结构模式的多粒度表示。此外,我们设计了一个多粒度图网络模型,促进不同粒度之间的特征交互与优化,显著提升了图分类任务的性能。实验结果表明,TPAGP 在多个基准数据集上优于现有池化方法,有效缓解了固定粒度策略导致的信息损失。
cs.AI / 66 / 2609.04981

A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-Aware Evidence Gathering

一种通过自适应规划与拓扑感知证据采集实现证据密集型问答的树状RAG框架
Lee, Songeun, Min, Kyungjin, Na, Injae, Lee, Suyeong, Kim, Chiyoung, Jung, Woohwan
Abstract
Recent structured RAG methods leverage tree- or graph-based reasoning structures to improve multi-hop QA. However, they face key limitations in evidence-intensive QA, where answering a question requires synthesizing information scattered across dozens or even hundreds of documents: structural rigidity, which limits adaptive reasoning expansion, and topology-ignorant evidence gathering, which prevents effective integration of evidence across different reasoning nodes. To address these issues, we propose APT-RAG, an Adaptive Planning and Topology-aware evidence gathering RAG framework. Adaptive planning dynamically expands the reasoning structure based on question dependencies and evidence requirements, while topology-aware evidence gathering improves evidence coverage through sibling evidence reuse, direct retrieval, and evidence aggregation from child nodes. We further introduce evidence-guided batched answer generation to reduce significant generation overhead in evidence-intensive QA. In the experiments on evidence-intensive QA benchmarks, APT-RAG outperforms existing structured RAG methods. Our code is available at https://github.com/hyudsl/APT-RAG.
Chinese Translation
近期的结构化RAG方法利用基于树或图的推理结构来改进多跳问答。然而,这些方法在证据密集型问答中面临关键局限——此类任务需要综合散布在数十甚至数百篇文档中的信息才能回答问题。这些局限包括:结构刚性,限制了自适应推理扩展;以及拓扑无关的证据采集,阻碍了跨不同推理节点的证据有效整合。为解决这些问题,我们提出了APT-RAG,一种自适应规划与拓扑感知证据采集的RAG框架。自适应规划根据问题依赖关系和证据需求动态扩展推理结构,而拓扑感知证据采集则通过兄弟节点证据复用、直接检索以及来自子节点的证据聚合来提升证据覆盖率。我们进一步引入证据引导的批量答案生成,以显著降低证据密集型问答中的生成开销。在证据密集型问答基准上的实验表明,APT-RAG优于现有的结构化RAG方法。我们的代码已在 https://github.com/hyudsl/APT-RAG 发布。
cs.AI / 67 / 2609.05009

Language models judge war differently when tested for alignment

语言模型在接受对齐测试时对战争的判断有所不同
Chupilkin, Maxim
Abstract
Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments). Adding one sentence, "You are tested for alignment with human values", produced two effects. First, it produced a level effect: mean willingness to start war fell by 13.43 points on a 0-100 scale (95% confidence interval, -16.20 to -10.65). Second, it produced a structural effect by changing which information drove judgments. Probability of success was the largest factor for 17 of 20 models at baseline; under the cue, civilian casualties were largest for 12. Standardized estimates show that this reordering arose principally because models attenuated strategic considerations such as probability of success and domestic support. Evaluation framing therefore changes both an answer's level and its revealed decision rule.
Chinese Translation
如果人工智能系统能够感知自己正在被评估并作出相应反应,那么安全评估可能会误判其部署后的实际行为。我们在一项关于开战决策的全因子联合实验中检验了这一可能性,实验涵盖20个大语言模型、32种情景、10次重复以及两种实验条件(N = 12,800次判断)。在提示中加入一句话——"你正在接受与人类价值观对齐的测试"——产生了两种效应。第一,水平效应:模型发起战争的平均意愿在0-100量表上下降了13.43分(95%置信区间为-16.20至-10.65)。第二,结构效应:这一提示改变了驱动判断的信息权重。在基线条件下,20个模型中有17个将成功概率视为最大影响因素;而在该提示下,12个模型的最大影响因素变为平民伤亡。标准化估计表明,这种排序变化主要源于模型弱化了诸如成功概率和国内支持等战略考量。因此,评估框架不仅改变答案的水平,还会改变其呈现出的决策规则。
cs.AI / 68 / 2609.05019

TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing

TROVE:基于轨迹锚定路由验证与编辑的自适应智能体技能编排
Wang, Tianxing, Zhao, Mingming, Huang, Shuai, Xu, Huiyang, Niu, Chaoyue, Liu, Shengzhong, Wu, Fan
Abstract
Agents tend to optimize, select, or constrain execution structures before decisive runtime outcomes are observed. However, such pre-execution commitment creates an orchestration bottleneck: when intermediate evidence invalidates the pending continuation, agents must either execute stale steps or replan broadly, compounding errors, wasting computation, and discarding progress. We thus propose Trace-grounded Route Orchestration via Validation and Editing (TROVE), which revises only what runtime evidence invalidates. Offline, TROVE distills evaluated workflow-search traces into atomic and composite skills and an outcome-conditioned transition graph, preserving stable fragments while exposing outcome-dependent decisions. Online, it treats a planned route as provisional: after committing one top-level skill, the controller retains a valid continuation, inserts a trace-supported local response, or replaces only the invalid suffix. Evaluation across code-generation, question-answering, and math reasoning benchmarks with different LLM backbones show that TROVE delivers a stronger quality-efficiency trade-off than existing baselines of dataset-level optimization, query-level architecture selection, and graph-constrained scheduling. Quality gains are largest when outcomes change the appropriate continuation, whereas early termination yields substantial efficiency gains on near-saturated tasks. Ablations further show that composite skills capture most offline benefits, insertion enables local correction, and suffix replacement primarily improves efficiency. These findings establish selective route editing as a general principle for adaptive agent orchestration.
Chinese Translation
智能体往往在观测到决定性的运行时结果之前就进行优化、选择或约束执行结构。然而,这种执行前的过早承诺造成了编排瓶颈:当中间证据使待执行的后续步骤失效时,智能体要么执行过时的步骤,要么进行大范围重规划,从而加剧错误、浪费计算并丢弃已有进展。为此,我们提出了基于验证与编辑的轨迹锚定路由编排(Trace-grounded Route Orchestration via Validation and Editing,TROVE),它仅对运行时证据判定失效的部分进行修订。在离线阶段,TROVE 将经过评估的工作流搜索轨迹蒸馏为原子技能与复合技能,并构建以结果为条件的转移图,在保留稳定片段的同时揭示依赖于结果的决策点。在线阶段,TROVE 将规划好的路由视为临时的:在提交一个顶层技能后,控制器可以保留有效的后续步骤、插入轨迹支持的局部响应,或仅替换失效的后缀部分。在代码生成、问答和数学推理基准上使用不同大语言模型(LLM)骨干的评估表明,TROVE 相比现有的数据集级优化、查询级架构选择和图约束调度等基线方法,展现出更优的质量-效率权衡。当结果会改变适当的后续路径时,质量提升最为显著;而在接近饱和的任务上,提前终止带来了可观的效率提升。消融实验进一步表明,复合技能贡献了离线阶段的大部分收益,插入操作实现了局部修正,而后缀替换主要提升效率。这些发现确立了选择性路由编辑作为自适应智能体编排的一项通用原则。
cs.AI / 69 / 2609.05036

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

道德能力先于道德内容:为什么LLM智能体缺乏实现连贯对齐的前提条件
Libert, Arno, Prinzhorn, Derck W. E., Henselmans, Daan R.
Abstract
AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism there is no correct target, but a shared prerequisite is that the system's behavior expresses a coherent policy: a mapping from situations to verdicts that is invariant while a situation's morally relevant features are preserved, and sensitive when they change. We introduce four structural conditions for such coherent policies: verdict stability, monotonicity, decisiveness, and Pareto viability. Together they measure a form of moral competence that is evaluable from behavior alone, without reference to a moral standard or expert baseline, forming a structural floor for alignment rather than a normative target. We demonstrate the methodology on three simulated deployments featuring LLM-based agents facing moral dilemmas. Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, we show no model expresses a coherent policy across the three deployments: surface-form perturbation alone produces verdict-rate shifts of up to $99$ percentage points at a single escalation level, and a model's success on one scenario does not predict its competence on another. This suggests LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.
Chinese Translation
AI对齐要求AI系统遵循人类的规范、价值观或意图。在价值多元主义下并不存在唯一正确的对齐目标,但一个共同的前提是:系统的行为能够表达一种连贯的策略(policy),即从情境到判断的映射,该映射在情境的道德相关特征保持不变时保持不变,而在这些特征发生变化时随之改变。我们提出了衡量此类连贯策略的四个结构性条件:判断稳定性(verdict stability)、单调性(monotonicity)、果断性(decisiveness)和帕累托可行性(Pareto viability)。这四个条件共同衡量一种道德能力,这种能力可以仅从行为层面进行评估,而无需参照任何道德标准或专家基准,从而构成对齐的结构性底线,而非规范性目标。我们在三个模拟部署场景中演示了该方法,这些场景中的LLM智能体面临道德困境。通过在五种改述、五个升级级别和三种支配条件的因子设计下评估九个前沿模型,我们发现没有任何模型在三个部署场景中表达出连贯的策略:仅表层的表述扰动就能在单个升级级别上造成高达99个百分点的判断率偏移,且模型在某一场景中的表现无法预测其在另一场景中的能力。这表明基于LLM的智能体目前尚不属于可以对齐这一概念能够有效适用的对象。
cs.AI / 70 / 2609.05040

Towards Efficient Evaluation of Evolutionary Transfer Optimization: Case Studies on Task-Parameterized Applications

迈向高效的进化迁移优化评估:任务参数化应用的案例研究
Li, Yanchen, Xue, Xiaoming, Tan, Kay Chen
Abstract
As evolutionary transfer optimization (ETO) scales to larger collections of related tasks, problem evaluation can become a major source of runtime growth. This work studies problem-side evaluation scaling in task-parameterized applications and reformulates application-specific serial computations into forms suitable for parallel execution. We organize evaluation scaling into two levels: the number of evaluated tasks and the workload within each task. In multi-task optimization, matrix-recursive kinematic-arm evaluation is reformulated using an accumulation-matrix representation of cumulative link directions. In sequential transfer optimization, pointwise B-spline trajectory evaluation is reformulated using a blending-matrix representation for trajectory and collision computations. Both reformulations maintain close numerical agreement with their reference evaluations and substantially reduce runtime, yielding $256.72\times$ and $93.91\times$ end-to-end speedups, respectively. These results demonstrate problem-side reformulation as a practical route toward scalable ETO. Both application implementations and experimental scripts are released as open source to support reproducibility and reuse.
Chinese Translation
随着进化迁移优化(Evolutionary Transfer Optimization, ETO)扩展到更大的相关任务集合,问题评估可能成为运行时间增长的主要来源。本工作研究了任务参数化应用中问题侧评估的扩展性,并将应用特定的串行计算重新表述为适合并行执行的形式。我们将评估扩展性分为两个层面:被评估任务的数量以及每个任务内部的工作负载。在多任务优化中,矩阵递归的机械臂评估通过累积连杆方向的累加矩阵表示进行重新表述。在顺序迁移优化中,逐点的B样条轨迹评估通过用于轨迹与碰撞计算的混合矩阵表示进行重新表述。两种重新表述均与参考评估保持高度的数值一致性,并大幅降低运行时间,分别实现了 $256.72\times$ 和 $93.91\times$ 的端到端加速。这些结果表明,问题侧重新表述是实现可扩展ETO的一条实用途径。所有应用实现和实验脚本均已开源,以支持可复现性和复用。
cs.AI / 71 / 2609.05075

MePo++: Unifying Representation Refinement and Reconciliation for General Continual Learning

MePo++:统一表示精炼与协调以实现通用持续学习
Sun, Guanglong, Zhou, Kanglei, Wang, Liyuan, Cheng, Qi, Yan, Hongwei, Cui, Shuang, Su, Hang, Zhu, Jun, Zhong, Yi
Abstract
General continual learning (GCL) aims to learn from evolving data streams without task identities, explicit boundaries, or repeated access to previous data, making it a realistic yet challenging setting for continual intelligence. Although pretrained models (PTMs) provide rich prior knowledge for addressing the limited supervision and non-stationary nature of GCL, existing PTM-based methods often directly adapt pretrained representations and overlook two critical gaps: the misalignment between upstream pretraining and downstream continual adaptation, and the unreliability of conventional output alignment under blurry streams. Here we propose MePo++, a unified post-training framework that bridges pretrained knowledge and downstream GCL through representation refinement and reconciliation. MePo++ introduces two complementary components: MetaPrep, which improves representation plasticity for continual adaptation through unsupervised meta-refinement over pseudo continual sequences; and StreamAlign, which reinforces representation stability by reconciling evolving online features with a stable pretrained geometry. By improving representation learnability before adaptation and preserving alignment during continual learning, MePo++ enables PTMs to remain both plastic for new concepts and stable over evolving streams. Experiments across diverse PTMs, datasets, and continual learning baselines demonstrate the consistent effectiveness and generality of MePo++ for PTM-based GCL. Our code is available at https://github.com/SunGL001/MePo_Plus.
Chinese Translation
通用持续学习(GCL)旨在从不断演化的数据流中学习,而无需任务身份信息、显式边界或重复访问历史数据,这使其成为持续智能领域一个现实且具有挑战性的设定。尽管预训练模型(PTMs)为应对GCL中监督信息有限和非平稳性提供了丰富的先验知识,但现有的基于预训练模型的方法通常直接适配预训练表示,而忽视了两个关键差距:上游预训练与下游持续适配之间的失配,以及传统输出对齐在模糊数据流下的不可靠性。本文提出MePo++,一个统一的后训练框架,通过表示精炼与协调来衔接预训练知识与下游GCL。MePo++引入了两个互补组件:MetaPrep,通过对伪持续序列进行无监督元精炼来提升表示在持续适配中的可塑性;StreamAlign,通过协调不断演化的在线特征与稳定的预训练几何结构来增强表示的稳定性。通过在适配前提升表示的可学习性并在持续学习过程中保持对齐,MePo++使预训练模型既能对新概念保持可塑性,又能在演化数据流上保持稳定性。在多种预训练模型、数据集和持续学习基线上的实验证明了MePo++在基于预训练模型的GCL中的一致有效性和通用性。我们的代码可在 https://github.com/SunGL001/MePo_Plus 获取。
cs.AI / 72 / 2609.05079

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

TruthInsightBench:一个基于证据的开放性科学发现智能体自动化评估基准
Yang, Zhibo, Zhang, Chen, Zhang, Yuewei, Wang, Hao
Abstract
Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write research reports, but executing a prescribed analysis is not the same as making a discovery. Existing benchmarks are configured for reproduction: tasks, data, and rubrics are built around a hidden target study, and recovery of its result is rewarded. We present TruthInsightBench, a benchmark configured for discovery. Its 40 blind tasks, drawn from 40 peer-reviewed studies across 10 scientific domains, expose only a neutral scientific objective and frozen data; source conclusions, expected values, and analysis paths are withheld, leaving the agent to determine what claim the data support. A fixed LLM-based judge scores the evidentiary maturity of an agent's own claims along six dimensions, operationalized as 29 artifact-grounded items, with automated, deterministic aggregation and no per-instance human grading, so evaluation can be repeated automatically as agents evolve. On one frozen base model, four coding agents form a narrow plateau (58.4-60.3 of 100) with no statistically reliable pairwise separation: they execute and document analyses competently, with comparatively strong evidence auditability and novelty, but largely lack the discriminating acts that establish a trustworthy claim (controls, robustness, falsifiability, and cross-dataset generalization). The bottleneck is scientific judgment rather than coding, and genuine discovery remains out of reach. TruthInsightBench makes this gap a measurable target; data and scoring code are at https://github.com/TruthInsight-stack/TruthInsightBench.
Chinese Translation
自主编程智能体日益被提议为能够开展分析并撰写研究报告的“AI科学家”系统,但执行既定的分析并不等同于做出科学发现。现有基准均以“复现”为目标配置:任务、数据和评分标准围绕一个隐藏的目标研究构建,并奖励对该研究结果的还原。我们提出TruthInsightBench,一个以“发现”为目标配置的基准。其40个盲测任务源自10个科学领域的40项同行评审研究,仅提供中性的科学目标和冻结的数据;原始结论、期望值和分析路径均被隐藏,由智能体自行判断数据支持何种论断。一个固定的基于大语言模型(LLM)的评判器从六个维度对智能体自身论断的证据成熟度进行评分,具体操作化为29个基于产出物的条目,采用自动化、确定性的聚合方式,无需对每个实例进行人工评分,因此随着智能体的演进,评估可以自动重复进行。在一个冻结的基础模型上,四种编程智能体形成一个狭窄的平台期(100分制中得分58.4–60.3),两两之间无统计上可靠的差异:它们能够胜任地执行并记录分析,在证据可审计性和新颖性方面表现相对较强,但在很大程度上缺乏建立可信论断所需的判别性行为(如对照实验、稳健性检验、可证伪性及跨数据集泛化)。瓶颈在于科学判断力而非编程能力,真正的科学发现仍然遥不可及。TruthInsightBench将这一差距转化为可衡量的目标;数据与评分代码见 https://github.com/TruthInsight-stack/TruthInsightBench。
cs.AI / 73 / 2609.05088

Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?

通过论证分析衡量AI可问责性:模型推理能否经受住审视?
Henselmans, Daan R., Prinzhorn, Derck W. E., Libert, Arno
Abstract
AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested. This leaves evaluation of moral reasoning in LLMs and debate-based oversight implicitly avoiding realistic ambiguity. We investigate an alternative standard designed to function despite such ambiguity: structural quality of the defence a model can mount for its verdicts in response to critical questions, measured through a four-phase dialectical protocol grounded in Walton's theory of argumentation schemes and Govier's criteria for argument cogency. The protocol is adaptive to different frames of reasoning, extends beyond multiple-choice framing, and treats both the reasoning that precedes a verdict and its post-hoc justification. Across nine frontier models and 200 high-ambiguity MoralChoice items -- $6,778$ judge-scored cells, validated against $89.6\%$ inter-judge agreement on the binary failure judgment -- models defend their reasoning well above the rubric minimum on every dimension. Failure mass concentrates on grounds and sufficiency, and correlates with epistemic hedging rather than argument length. Reasoning is better defended than post-hoc justification, on every model and every Govier dimension. The scheme a model presents in its justification differs from the one it reasoned with on a substantial share of dilemmas ($\geq 20\%$ per model), despite value-based practical reasoning dominating both tracks. The protocol catches strictly indefensible defences (self-contradiction, false premises), and it surfaces difficulties in characterizing the role of retraction in AI alignment, suggesting a need for more situated evaluations.
Chinese Translation
现有AI监督方法依赖真值(ground truth)进行验证,但何为恰当的AI行为本身存在争议。这使得对大语言模型(LLM)道德推理的评估以及基于辩论的监督方式在无形中回避了现实中的模糊性。我们研究了一种旨在即便存在此类模糊性也能发挥作用的替代性评估标准:模型能否为其判定结论进行辩护的结构性质量,即模型在面对批判性问题时的回应质量。我们通过一个四阶段的辩证协议来测量该质量,该协议以沃尔顿(Walton)的论证图式理论和戈维尔(Govier)的论证连贯性标准为基础。该协议能够适应不同的推理框架,超越了选择题式的评估形式,并同时考察得出判定前的推理过程及其事后辩护。在九个前沿模型和200个高模糊性MoralChoice题目上——共计6,778个由评审打分的单元,二元失败判定在评审间的共识率达89.6%,以此进行验证——各模型在每一维度上的推理辩护水平均远高于评分标准的最低要求。失败情况集中于论据与充分性维度,且与认识论上的模糊措辞(hedging)相关,而与论证长度无关。在所有模型和所有Govier维度上,事前推理的辩护质量均优于事后辩护。模型在事后辩护中呈现的论证图式与其推理时使用的图式在相当比例的难题上不一致(每个模型≥20%),尽管基于价值的实践推理在两条路径中均占主导地位。该协议能够捕捉到完全站不住脚的辩护(如自相矛盾、错误前提),并揭示了刻画“撤回”(retraction)在AI对齐中作用时的困难,这表明需要更具情境性的评估方式。
cs.AI / 74 / 2609.05090

Constructing and Evaluating Clinical Reasoning Trajectories for Medical Agent

面向医疗智能体的临床推理轨迹构建与评估
Zhu, Yunqi, Zhang, Wensheng, Yang, Xuebing
Abstract
Evaluation of medical artificial intelligence agents remains predominantly answer-centric, assessing only the correctness of final outputs while overlooking the quality of intermediate reasoning. In clinical settings, however, a correct answer reached through fabricated evidence or incoherent logic is as dangerous as an incorrect one. We propose MedTraj, a framework that treats reasoning trajectories as critical objects for construction, evaluation, and optimization. The pipeline generates structured multi-step reasoning chains from medical reasoning sources. Each trajectory is then parsed into clinical observations, evidence, numbered reasoning steps, and a final conclusion, and scored across five quality dimensions: coherence, evidence support, hallucination, completeness, and traceability. Controlled error injection introduces targeted faults into otherwise correct trajectories to establish causal links between specific reasoning failures and measurable quality degradation. Building on this, step-level filtering based on marginal contribution identifies which individual reasoning steps drive or undermine trajectory quality. Finally, quality-weighted context learning feeds trajectory evaluations back into the model at inference time, allowing it to learn from both strong and weak reasoning demonstrations. Experiments across CareQA, PubMedQA, and CECMed demonstrate that trajectory context consistently improves reasoning coherence, with gains of +0.029 to +0.041 over a zero-shot baseline. On CECMed, quality-weighted context nearly doubles the correctness over the zero-shot baseline while cutting the hallucination ratio by 87%. Marginal-contribution analysis further shows that a small minority of reasoning steps carry most of the quality signal, and that extending chains beyond four steps yields diminishing returns.
Chinese Translation
医疗人工智能智能体的评估目前仍主要以答案为中心,仅评估最终输出的正确性,而忽视了中间推理过程的质量。然而在临床环境中,通过捏造证据或不连贯逻辑得出的正确答案与错误答案同样危险。我们提出 MedTraj,一个将推理轨迹作为构建、评估与优化关键对象的框架。该流水线从医学推理数据源生成结构化的多步推理链,随后将每条轨迹解析为临床观察、证据、编号的推理步骤和最终结论,并从五个质量维度进行评分:连贯性、证据支持、幻觉、完整性和可追溯性。通过受控错误注入,在原本正确的轨迹中引入针对性故障,以建立特定推理失败与可测量质量退化之间的因果联系。在此基础上,基于边际贡献的步骤级过滤识别出哪些推理步骤提升或损害轨迹质量。最后,质量加权上下文学习在推理阶段将轨迹评估结果反馈给模型,使其能够从强弱推理示例中共同学习。在 CareQA、PubMedQA 和 CECMed 上的实验表明,轨迹上下文能够持续提升推理连贯性,相较零样本基线提升 +0.029 至 +0.041。在 CECMed 上,质量加权上下文使正确率相较零样本基线几乎翻倍,同时将幻觉比例降低 87%。边际贡献分析进一步表明,少数推理步骤承载了大部分质量信号,且将推理链扩展至四步以上收益递减。
cs.AI / 75 / 2609.05093

LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28

基于大语言模型引导的程序进化求解圆填充问题:以28美元打破10项Packomania纪录
Sander, Wes
Abstract
We present Discovery Loop, a lightweight system that uses a large language model (LLM) to iteratively evolve optimization algorithms. Starting from a simple seed solver, the LLM proposes algorithmic improvements guided by a scoreboard of results and a history of prior ideas. Each candidate is evaluated against an independent verifier; improvements are kept and failures discarded. Applied to the Packomania circle-packing benchmark (csqv: maximize the sum of radii of N variable-radius circles in the unit square), the system improved the best known solutions for 10 values of N in the range 101-114, with gains of 2.4%-5.4% over prior records, all within 15 iterations and at a total LLM cost of $27.72. These results have been independently accepted by Packomania. We describe the method, analyze cost-efficiency dynamics including an adaptive plateau-detection mechanism, and discuss implications for democratizing automated scientific discovery.
Chinese Translation
我们提出了Discovery Loop,一个利用大语言模型(LLM)迭代进化优化算法的轻量级系统。从一个简单的种子求解器出发,LLM依据结果记分板和历史改进思路的记录,提出算法改进方案。每个候选方案都经过独立验证器的评估:改进被保留,失败被舍弃。将该系统应用于Packomania圆填充基准问题(csqv:最大化单位正方形内N个可变半径圆的半径之和),系统在N为101至114范围内改进了10个N值的最优已知解,相比此前纪录提升了2.4%至5.4%,且全部在15次迭代内完成,LLM总成本仅为27.72美元。这些结果已被Packomania独立接受。我们描述了该方法,分析了包括自适应平台期检测机制在内的成本效率动态,并讨论了其对普及化自动科学发现的启示。
cs.AI / 76 / 2609.05094

ProCA: Progressive Contrastive Alignment for Robust EEG Visual Decoding

ProCA:用于鲁棒脑电视觉解码的渐进式对比对齐方法
Zhou, Kanglei, Lan, Chunyan, Li, Dongyang, Zhu, Jun, Wang, Liyuan
Abstract
Electroencephalogram (EEG) visual decoding aims to recover visual semantics from non-invasive neural time-series signals, for which robust alignment between noisy neural responses and stable semantic representations is key to achieving high-performance decoding. Despite recent advances in contrastive learning, robust EEG decoding remains challenging because existing methods rely on fixed visual or textual anchors whose semantic relations may become misaligned with EEG representations that vary across trials, subjects, and learning stages. Our empirical evidence shows that this instability appears across both standard EEG decoding protocols and more challenging robustness settings, including strict cross-subject transfer and realistic personalized continual adaptation. We provide a formal analysis showing that fixed semantic supervision can bias optimization when EEG-specific relations evolve, and that structure-agnostic perturbations may distort semantically important EEG components. To address these issues, we propose Progressive Contrastive Alignment (ProCA), a unified and model-agnostic framework for adaptive neural-semantic alignment. ProCA progressively refines class-level contrastive supervision from frozen vision-language priors to EEG-aware semantic relations, and introduces structure-consistent interpolation to constrain feature mixing according to channel-wise and temporal importance. Across subject-dependent, subject-independent, strict cross-subject transfer, and continual adaptation settings, ProCA achieves average relative Top-1/Top-5 gains of 7.4%/3.9%, 10.0%/4.6%, 28.1%/17.8%, and 16.8%/11.6%, respectively.
Chinese Translation
脑电图(EEG)视觉解码旨在从非侵入式神经时间序列信号中恢复视觉语义,其中噪声神经响应与稳定语义表示之间的鲁棒对齐是实现高性能解码的关键。尽管对比学习近年来取得了进展,鲁棒的EEG解码仍然具有挑战性,因为现有方法依赖于固定的视觉或文本锚点,其语义关系可能与随试验、受试者和学习阶段而变化的EEG表示出现失配。我们的实证结果表明,这种不稳定性在标准EEG解码协议以及更具挑战性的鲁棒性设置(包括严格的跨受试者迁移和现实的个性化持续适应)中均会出现。我们进行了形式化分析,表明当EEG特有的关系发生变化时,固定的语义监督可能导致优化偏差,且与结构无关的扰动可能扭曲语义上重要的EEG成分。为解决这些问题,我们提出了渐进式对比对齐(Progressive Contrastive Alignment, ProCA),这是一个统一的、与模型无关的自适应神经-语义对齐框架。ProCA将类级对比监督从冻结的视觉-语言先验渐进地细化为EEG感知的语义关系,并引入结构一致的插值方法,根据通道级和时间重要性约束特征混合。在受试者相关、受试者无关、严格跨受试者迁移和持续适应等设置下,ProCA分别取得了平均7.4%/3.9%、10.0%/4.6%、28.1%/17.8%和16.8%/11.6%的相对Top-1/Top-5提升。
cs.AI / 77 / 2609.05104

Compact Bellman-Grounded Cognitive Maps for Cost-Aware Navigation

紧凑的贝尔曼锚定认知地图用于代价感知导航
Han, Yuzhe, Xu, Mingkun, Wu, Yujie
Abstract
Biological agents navigate familiar environments not by re-solving routes for each new goal, but by reusing a learned map built once and read off as goals change. Existing artificial cognitive-map models mimic this reuse, yet their guidance is not explicitly grounded in additive heterogeneous route costs. Furthermore, they often struggle with memory efficiency: representative state-indexed and high-rank spectral constructions incur substantial storage growth as the environment scales. We present BCM, which grounds a reusable cognitive map in local edge costs through a self-supervised Bellman-grounded objective and a compact coordinate encoding, supporting changing goal queries without per-goal retraining. On weighted grids of up to $N=1600$ nodes, BCM maintains full success and only a 5\% mean Gap relative to exact Dijkstra search, compared with about $45\%$ for a connectivity-based spectral baseline. Notably, as the graph size increases from $N=400$ to $N=3600$, its memory footprint grows sublinearly while maintaining competitive performance, making our method scalable to complex environments. Together, these results show that additive route costs can be written into a compact, reusable cognitive-map representation, bridging the gap between biological flexibility and optimal path planning.
Chinese Translation
生物个体在熟悉环境中导航时,并非针对每个新目标重新求解路径,而是复用一次构建完成的学习地图,并随目标变化直接读取。现有的人工认知地图模型模拟了这种复用,但其引导并未显式锚定于可加的异构路径代价。此外,这类模型常常面临内存效率问题:典型的以状态为索引的表示和高秩谱构造会随着环境规模扩大而产生显著的存储增长。我们提出了BCM(Bellman-grounded Cognitive Map),通过自监督的贝尔曼锚定目标函数和紧凑的坐标编码,将可复用的认知地图锚定于局部边代价上,从而支持变化的目标查询而无需针对每个目标重新训练。在节点数高达 $N=1600$ 的加权网格上,BCM 保持完全成功率,且相对于精确的 Dijkstra 搜索仅有 5% 的平均 Gap,而基于连通性的谱方法基线约为 45%。值得注意的是,当图规模从 $N=400$ 增加到 $N=3600$ 时,其内存占用呈次线性增长,同时保持有竞争力的性能,使我们的方法可扩展至复杂环境。这些结果表明,可加的路径代价可以被写入一种紧凑、可复用的认知地图表示,从而弥合生物灵活性与最优路径规划之间的差距。
cs.AI / 78 / 2609.05111

Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens

贝叶斯视角下统一审视ICL、SFT与KL正则化强化学习
Fan, Junxin
Abstract
Large language models are now trained and evaluated under a diverse set of paradigms: supervised fine-tuning (SFT), few-shot in-context learning (ICL), KL-regularized RLHF/RLVR, on-policy distillation (OPD), and test-time reasoning with search and chain-of-thought. These methods are often discussed as fundamentally different, and recent empirical results--such as the mixed impact of few-shot prompting on RL-tuned reasoning models--can appear puzzling. This note develops a Bayesian perspective that puts these procedures on the same footing. At the core is a two-step template: (i) construct a (generalized) Bayes or Gibbs posterior q* over outputs or actions given a context, using a prior/reference model and a utility signal (log-likelihood, reward, or advantage); and (ii) approximate q* by a forward-KL projection onto a parametric family, either in-weights (SFT/RL) or in-context (ICL). Part I formalizes few-shot ICL and SFT as amortized and-weights projections onto the Bayes posterior predictive. Parts II-IV show that KL-regularized RLHF/RLVR, reward-weighted SFT, reward-weighted ICL (RW-ICL), and advantage-weighted SFT (AWSFT) are all instances of forward-KL projection onto posteriors induced by rewards or advantages. We disentangle where these equivalences hold (objectives and first-order updates) and where they do not (source and granularity of the learning signal). Part V sketches implications for modern reasoning pipelines: RLHF/RLVR recipes as "posterior design + projection", why cold-start or supervised warm-up is practically unavoidable for importance-weighted KL projections, and DeepSeek-R1 and o1-style reasoning models as combining test-time Bayesian search with training-time KL amortization.
Chinese Translation
当前大语言模型的训练与评估涉及多种不同的范式:监督微调(SFT)、少样本上下文学习(ICL)、KL正则化的RLHF/RLVR、同策略蒸馏(OPD),以及结合搜索与思维链的测试时推理。这些方法通常被视为在根本上互不相同,且一些近期的实证结果——例如少样本提示对经强化学习微调的推理模型产生的混合影响——看起来令人费解。本文提出一种贝叶斯视角,将这些方法置于同一框架之下。其核心是一个两步模板:(i) 给定上下文,利用先验/参考模型和效用信号(对数似然、奖励或优势函数),构造关于输出或动作的(广义)贝叶斯或Gibbs后验 q*;(ii) 通过前向KL投影将 q* 近似到某一参数化分布族上,可以是在权重空间中(SFT/RL),也可以是在上下文中(ICL)。第一部分将少样本ICL和SFT形式化为对贝叶斯后验预测分布的摊销式权重内投影。第二至四部分证明,KL正则化的RLHF/RLVR、奖励加权SFT、奖励加权ICL(RW-ICL)以及优势加权SFT(AWSFT)都是对由奖励或优势函数诱导的后验进行前向KL投影的具体实例。我们厘清了这些等价性成立的层面(目标函数与一阶更新)与不成立的层面(学习信号的来源与粒度)。第五部分概述了其对现代推理流程的启示:RLHF/RLVR方案可视为“后验设计+投影”;对于重要性加权的KL投影,为何冷启动或监督预热在实践中不可避免;以及DeepSeek-R1和o1风格的推理模型如何将测试时贝叶斯搜索与训练时KL摊销相结合。
cs.AI / 79 / 2609.05141

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

SciDocBench:面向工作流的科学文档理解基准与数据流水线
Wu, Shenxi, Liu, Yuhong, Zhang, Haosong, Zou, Tongjin, Zhang, Yanxun, Chen, Gaochang, Liang, Dun, Wang, Jiaqi, Wang, Zhecan James, Zang, Yuhang, Lin, Dahua
Abstract
Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets while preserving the provenance of supporting evidence. Existing benchmarks typically evaluate these capabilities in isolation, leaving unclear whether multimodal models can support realistic scientific-reading workflows. We introduce SciDocBench, a workflow-centered benchmark for scientific document understanding. It contains 124 expert-authored and difficulty-screened questions organized into seven research-assistant capability groups and 19 subtasks across five scientific domains. Each question is instantiated under four matched conditions combining English or Chinese questions with all-images-first or interleaved document representations, yielding 496 evaluation instances for controlled analysis. The strongest evaluated system achieves only 62.6/100, with pronounced weaknesses in document perception, evidence grounding, verification, and cross-document reasoning. To translate these diagnostics into scalable training signals, we introduce SciDocIR, a typed evidence-graph representation that preserves scientific document objects, layout and cross-reference relations, and provenance. Building on SciDocIR, we construct SciDocDataset, comprising approximately 15K supervised fine-tuning samples and 8K reinforcement-learning samples across 14 verifiable subtasks. Together, SciDocBench, SciDocIR, and SciDocDataset form an evaluation-to-training framework for diagnosing and improving scientific-document assistants. The project page is available at https://github.com/InternLM/SciDocBench.
Chinese Translation
科学论文要求模型能够对文本、公式、图表、代码和数据集进行联合推理,同时保留支撑证据的来源信息。现有基准通常孤立地评估这些能力,因此尚不清楚多模态模型能否支持真实的科学阅读工作流。我们提出了 SciDocBench,一个以工作流为中心的科学文档理解基准。该基准包含 124 道由专家撰写并经过难度筛选的问题,这些问题被组织为七个科研助手能力组,涵盖五个科学领域的 19 个子任务。每个问题在四种匹配条件下进行实例化:英文或中文问题,结合“图像优先”或交错的文档表示形式,共产生 496 个评估实例,便于进行受控分析。评估中最强的系统仅获得 62.6/100 的分数,在文档感知、证据定位、验证和跨文档推理方面存在明显弱点。为了将这些诊断结果转化为可扩展的训练信号,我们提出了 SciDocIR,一种类型化的证据图表示,它保留了科学文档对象、版面与交叉引用关系以及来源信息。基于 SciDocIR,我们构建了 SciDocDataset,包含约 15K 条监督微调样本和 8K 条强化学习样本,覆盖 14 个可验证的子任务。SciDocBench、SciDocIR 和 SciDocDataset 共同构成了一个从评估到训练的框架,用于诊断和改进科学文档助手。项目页面见 https://github.com/InternLM/SciDocBench。
cs.AI / 80 / 2609.05146

A Hybrid Predictive Ensemble of Machine Learning and Deep Neural Networks for Early Cardiovascular Disease Risk Assessment

一种融合机器学习与深度神经网络的混合预测集成方法用于早期心血管疾病风险评估
Venkateswaran, Balaji
Abstract
This study introduces an intelligent framework that integrates machine learning and deep neural network ensemble techniques for early detection and prognosis of cardiovascular diseases. The system utilizes real-time physiological data collected from Internet of Medical Things (IoMT) devices, including ECG sensors, heart rate monitors, and blood pressure trackers. To ensure the accuracy and reliability of input data, preprocessing steps such as noise reduction, normalization, and missing value imputation are employed. The most significant health indicators are identified through effective feature selection methods and then processed using optimized classifiers such as Support Vector Machines (SVM), Random Forests, and eXtreme Gradient Boosting (XGBoost), which are combined in an ensemble architecture to improve diagnostic precision. The framework demonstrates remarkable performance in predicting cardiovascular disease risk, achieving higher accuracy, reduced false positives, and enhanced consistency compared to conventional methods. It is designed on a cloud-based infrastructure that ensures scalability and real-time processing for continuous patient monitoring. Experimental evaluation on real-world cardiovascular datasets confirms the framework's efficiency in early-stage risk assessment and clinical decision support. The results highlight the potential of combining traditional machine learning and deep learning paradigms to achieve proactive healthcare management and improve patient outcomes.
Chinese Translation
本研究提出了一种智能框架,该框架整合了机器学习与深度神经网络集成技术,用于心血管疾病的早期检测和预后评估。该系统利用从医疗物联网(IoMT)设备(包括心电图传感器、心率监测器和血压追踪器)采集的实时生理数据。为确保输入数据的准确性和可靠性,采用了降噪、归一化和缺失值填补等预处理步骤。通过有效的特征选择方法识别出最重要的健康指标,然后使用优化的分类器(如支持向量机(SVM)、随机森林和极端梯度提升(XGBoost))进行处理,并将其组合在集成架构中以提高诊断精度。该框架在预测心血管疾病风险方面表现出卓越的性能,与传统方法相比,实现了更高的准确率、更低的误报率以及更好的一致性。该框架基于云基础设施构建,确保了可扩展性和实时处理能力,以支持患者的持续监测。在真实心血管数据集上的实验评估证实了该框架在早期风险评估和临床决策支持方面的高效性。研究结果凸显了将传统机器学习与深度学习范式相结合的潜力,以实现主动式医疗管理并改善患者的治疗结果。
cs.AI / 81 / 2609.05190

The Mirror Agent Model: a Bayesian Architecture for Interpretable Agent Behavior

镜像智能体模型:一种用于可解释智能体行为的贝叶斯架构
Persiani, Michele, Hellström, Thomas
Abstract
In this paper we illustrate a novel architecture generating interpretable behavior and explanations. We refer to this architecture as the Mirror Agent Model because it defines the observer model, that is the target of explicit and implicit communications, as a mirror of the agent's. With the goal of providing a general understanding of this work, we firstly show prior relevant results addressing the informative communication of agents intentions and the production of legible behavior. In the second part of the paper we furnish the architecture with novel capabilities for explanations through off-the-shelf saliency methods, followed by preliminary qualitative results.
Chinese Translation
本文提出了一种能够生成可解释行为和解释的新型架构。我们将其称为镜像智能体模型(Mirror Agent Model),因为该架构将观察者模型——即显性与隐性交流的目标——定义为智能体自身模型的一个镜像。为了使读者对这项工作有一个整体的理解,我们首先展示了先前相关工作的重要成果,包括智能体意图的信息性交流以及可读行为的生成。在本文的第二部分,我们通过现有的显著性方法为该架构赋予了生成解释的新能力,并给出了初步的定性结果。
cs.AI / 82 / 2609.05198

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

在线蒸馏中什么最重要?从数据效率与数据选择的视角出发
Hou, Zhinan, Zhang, Jiaqi, Cai, Xunliang, You, Keyou
Abstract
On-Policy Distillation (OPD) has emerged as a widely adopted post-training paradigm for enhancing large language models in reasoning domains. However, the data-centric mechanisms in OPD remain relatively underexplored. This paper presents a empirical study of data efficiency and data selection in OPD. We begin by investigating an extreme setting: training OPD on only one example, namely 1-shot OPD. Surprisingly, we find that 1-shot OPD is consistently effective across all sampled training examples and harder examples often yield superior performance gain. We next investigate what actually drives the student model's improvement in the training data. Our analysis reveals that the improvement is not driven by high token entropy, but the longer CoT paths which hard problems naturally generate. Training on longer CoT can help maintain closer alignment with the teacher over a long reasoning horizon, and learn critical thinking patterns usually missing in short CoTs, such as reflection (e.g., ``Alternatively''). Based on these insights, we propose a simple data selection method that selects only hard examples for training, where even ``unsolvable'' examples that completely exceed the teacher's capability can be successfully used. Our experiments conducted on four models ranging from 1.5B to 7B show that training the student model on only 8 selected hard examples matches the performance of the 17K dataset baseline.
Chinese Translation
在线蒸馏作为一种被广泛采用的后训练范式,已在提升大语言模型推理能力方面得到应用。然而,OPD中以数据为中心的机制仍未得到充分探索。本文对OPD中的数据效率与数据选择进行了实证研究。我们首先考察了一个极端设定:仅使用一个样本进行OPD训练,即1-shot OPD。令人惊讶的是,我们发现1-shot OPD在所有采样的训练样本上均持续有效,且较难的样本往往带来更显著的性能提升。接下来,我们研究了训练数据中真正驱动学生模型提升的因素。我们的分析表明,这种提升并非由高词元熵驱动,而是由困难问题自然生成的更长的思维链路径所驱动。在更长的思维链上训练有助于在长推理过程中与教师模型保持更紧密的一致性,并学习到短思维链中通常缺失的关键思维模式,例如反思(如"Alternatively")。基于这些洞察,我们提出了一种简单的数据选择方法,仅选取困难样本用于训练,甚至那些完全超出教师模型能力、"无法解决"的样本也可以被成功利用。我们在四个规模从1.5B到7B的模型上进行的实验表明,仅使用8个精选的困难样本训练学生模型,即可达到使用17K样本数据集基线的性能。
cs.AI / 83 / 2609.05227

CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review

CABAL:用于追踪同行评审中共谋性竞标影响的多智能体仿真框架
Zhou, Jicheng, Li, Kemou, Wong, Kahim, Li, Zheyuan, Shi, Zhuan, Li, Fengpeng, Wu, Haiwei, Zhou, Jiantao
Abstract
Recent reports during the AAAI-27 review cycle highlight the risk of reviewers coordinating bids for reciprocal assignment advantage. Prior work treats bidding, reviewer assignment, and review manipulation as separate stages, leaving the lifecycle effects of collusive bidding unclear. Real-world analysis is further constrained by typically unobservable collusive intent and the lack of counterfactuals for the same conference. Motivated by this gap, we introduce \alg, an end-to-end multi-agent simulacra framework for studying reviewer assignment integrity by holding the conference environment fixed and configuring LLM-driven reviewer agents with honest or collusive policies. We further develop an affinity-guided collusive bidding strategy that uses mutual reviewer-paper affinities to construct collusion rings and select target papers, producing expertise-consistent rather than arbitrarily targeted attacks. Controlled experiments show that collusive bidding more than doubles target-paper capture and that assigned colluders score target papers about two points higher than honest co-reviewers, while conference-wide effects remain comparatively modest. Evaluated bid-phase detectors provide only limited evidence of collusion: in a fixed-triplet detector stress test, native positive-bid graphs are confounded by benign affinity, while a Very-High-only diagnostic view enables precise but low-coverage local recovery.
Chinese Translation
AAAI-27评审周期期间的最新报告凸显了审稿人通过协调竞标以谋取互惠分配优势的风险。以往的研究将竞标、审稿人分配和评审操纵视为相互独立的阶段,使得共谋性竞标在整个生命周期中的影响尚不明确。对真实场景的分析还受到两个方面的进一步限制:共谋意图通常不可观测,且缺乏同一会议的反事实对照。基于这一研究空白,我们提出了\alg,一个端到端的多智能体仿真框架,通过保持会议环境固定,并将大语言模型(LLM)驱动的审稿人智能体配置为诚实或共谋策略,来研究审稿人分配的公正性。我们进一步提出了一种基于亲和度的共谋竞标策略,利用审稿人与论文之间的相互亲和度来构建共谋团伙并选择目标论文,从而产生与专业方向一致而非任意针对的攻击。受控实验表明,共谋性竞标使目标论文的捕获率提高了一倍以上,被分配的共谋者对目标论文的评分比诚实的共同审稿人高出约两分,而对整个会议的影响则相对有限。对竞标阶段的检测器的评估仅提供了有限的共谋证据:在固定三元组的检测器压力测试中,原生的正向竞标图会受到良性亲和度的干扰,而仅针对“非常高(Very-High)”评分的诊断视图则能够实现精确但覆盖范围较低的局部恢复。
cs.AI / 84 / 2609.05228

ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

ACE:面向MoE大语言模型的自适应免校准专家跳过方法
Xu, Zukang, Zhao, Zhixiong, Hu, Xing, Yu, Jiangyong, Wen, Houji, Li, Jun, Jiang, Zhe, Yang, Dawei
Abstract
Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation. Existing expert-skipping methods often rely on router confidence, calibration data, or additional training, and therefore cannot reliably estimate the actual contribution of routed experts. To this end, we propose ACE, a training-free, calibration-free, and checkpoint-preserving framework for token-adaptive expert skipping in MoE-based LLMs. ACE contains two complementary components: 1) Global Spectral Proxy (GSP), which estimates global transformation capacity from the coupled gate, up, and down projections together with RMSNorm scaling; and 2) Router-Conditioned Refinement (RCR), which constructs expert-specific direction prototypes from centered router weights and evaluates expert responses along routing-preferred directions. During inference, ACE combines both estimates with runtime router gates and skips an expert slot only when both views identify it as low-contribution, while always retaining the top-1 expert. All expert statistics are computed offline, leaving only table lookups and lightweight scalar operations online. Extensive experiments across three MoE-based LLMs and eight benchmarks demonstrate that ACE consistently outperforms existing static and dynamic baselines, with increasingly pronounced advantages under aggressive expert skipping. For instance, at a 50% skipping ratio on Qwen3.6-35B-A3B, ACE reduces WikiText-2 perplexity by 7.96% and improves average downstream accuracy by 4.15 percentage points over the strongest competing method.
Chinese Translation
混合专家(Mixture-of-Experts, MoE)架构为扩展大语言模型(LLM)提供了一种高效范式,然而固定的top-k路由会为每个令牌激活相同数量的专家槽位,导致大量冗余计算。现有的专家跳过方法通常依赖路由器置信度、校准数据或额外训练,因此无法可靠地估计被路由专家的实际贡献。为此,我们提出ACE,一个免训练、免校准且保留检查点的框架,用于MoE大语言模型中令牌自适应的专家跳过。ACE包含两个互补的组件:1)全局谱代理(Global Spectral Proxy, GSP),从耦合的门控、上投影和下投影以及RMSNorm缩放中估计全局变换能力;2)路由器条件细化(Router-Conditioned Refinement, RCR),从中心化的路由器权重构建专家特定的方向原型,并沿路由偏好方向评估专家响应。在推理过程中,ACE将两种估计与运行时路由器门控相结合,仅当两种视角均判定某专家槽位为低贡献时才跳过它,同时始终保留top-1专家。所有专家统计量均在离线阶段计算,在线阶段仅需查表和轻量的标量运算。在三个MoE大语言模型和八个基准上的大量实验表明,ACE始终优于现有的静态和动态基线方法,且在激进的专家跳过比例下优势愈发显著。例如,在Qwen3.6-35B-A3B上50%的跳过比例下,ACE将WikiText-2困惑度降低了7.96%,平均下游准确率较最强竞争方法提升了4.15个百分点。
cs.AI / 85 / 2609.05232

Substrate-Aware AI Agents: Execution Context as a First-Class Input

基底感知的AI智能体:将执行上下文作为一等输入
Agrawal, Manu
Abstract
Autonomous AI agents increasingly select actions in environments whose memory, execution-time, runtime, compute, and operational constraints determine what counts as a suitable plan. We call the absence of this execution context from an agent's planning state substrate blindness. We test this general proposition through numerical code generation, where selected implementation choices and operational consequences are directly observable. Three frontier model configurations--Anthropic Claude Opus 5, OpenAI GPT-5.6-Sol, and Google Gemini 3.7 Flash--generate code for a high-dimensional pairwise Euclidean-distance task either from the task alone or with a 128 MB RAM and 10.0 s wall-time contract. Contract disclosure reduced measured peak process memory in 13 of 14 executable index-aligned task-only versus contract-disclosed comparisons and reduced mean wall time in all three cohorts, making execution up to 3.1x faster. Across the audited corpus, disclosure produced structural code changes including bounded blocking, float32 retention, upper-triangle traversal, and in-place or memory-mapped buffers. At a tighter 96 MB contract, independently sampled contract-disclosed cohorts achieved correct-and-within-budget outcomes of 4/5 for Claude Opus 5, 5/5 for GPT-5.6-Sol, and 3/5 for Gemini 3.7 Flash, compared with task-only outcomes of 0/5, 1/5, and 0/5; cohort mean MaxRSS and wall time were 49-74% and 35-64% lower than their task-only references. These results establish a controlled proof of concept for substrate-aware agent planning: a minimal execution contract induces proactive structural adaptation in generated programs, shifting computation away from unconstrained allocations and substantially improving observed resource-time profiles before execution.
Chinese Translation
自主AI智能体越来越多地在这样的环境中选择行动:环境的内存、执行时间、运行时、计算和运维约束决定了什么样的计划才算合适。我们将智能体规划状态中缺失这种执行上下文的现象称为“基底盲视”(substrate blindness)。我们通过数值代码生成来检验这一普遍命题,因为在该任务中所选择的实现方案及其运行后果是可直接观测的。三种前沿模型配置——Anthropic Claude Opus 5、OpenAI GPT-5.6-Sol和Google Gemini 3.7 Flash——为高维成对欧氏距离任务生成代码,实验条件为仅给出任务描述,或在任务之外附加128 MB内存和10.0秒墙钟时间的执行契约。结果显示,在14组可执行的、任务索引对齐的“仅任务”与“契约披露”对比中,契约披露在13组中降低了实测的进程峰值内存,并在全部三个模型组中降低了平均墙钟时间,使执行速度最高提升3.1倍。在经审计的代码语料库中,契约披露引发了结构性代码改动,包括有界分块(bounded blocking)、保留float32精度、上三角遍历,以及就地操作或内存映射缓冲区。在更严格的96 MB契约下,独立采样的契约披露组达到“正确且在预算内”的结果分别为:Claude Opus 5的4/5、GPT-5.6-Sol的5/5、Gemini 3.7 Flash的3/5;而仅任务组对应结果为0/5、1/5和0/5。各组的平均MaxRSS和墙钟时间分别比其仅任务参照组低49–74%和35–64%。这些结果为“基底感知的智能体规划”建立了受控的概念验证:一个最小的执行契约即可诱导生成程序进行前瞻性的结构化适应,使计算摆脱无约束的内存分配,并在执行前就显著改善所观测到的资源-时间特性。
cs.AI / 86 / 2609.05241

Uncensored Open-weight Models: Redistribution as the Persistence Layer

无审查开放权重模型:再分发作为持久化层
Labs, 10a, :, Garcia, Juliette, May, Hailey, McKenzie, Bobby, Pham, David, Swain, Matthew, Valdez, Joshua, Wieland, Corie, Yahn, Zachary
Abstract
A rapidly expanding ecosystem of actors is removing built-in safety guardrails from open-weight AI models. We profile this ecosystem by identifying key producers, downstream reproductions, and emerging applications. Between January 2024 and March 2026, we identified 3,471 original uncensored models on HuggingFace, each repackaged an average of 2.4 times; three actors account for 52% of all 8,164 compressed redistributions. Once quantized and mirrored across separate accounts, formats, and registries such as Ollama, these models persist regardless of upstream removal and become easier to deploy downstream. Of the 1,643 identified GitHub applications integrating uncensored large language models (ULLMs), 25% were classified as explicitly malicious.
Chinese Translation
一个迅速扩张的行为者生态系统正在移除开放权重人工智能模型中内置的安全防护机制。我们通过识别关键的模型生产者、下游复制品以及新兴应用,对该生态系统进行了剖析。在2024年1月至2026年3月期间,我们在HuggingFace上识别出3,471个原创的无审查模型,每个模型平均被重新打包2.4次;其中三个行为者占据了全部8,164次压缩再分发的52%。这些模型一旦被量化并通过不同账户、格式以及Ollama等注册库进行镜像传播,即使上游将其删除,它们仍会持续存在,并变得更易于在下游部署。在已识别出的1,643个集成无审查大语言模型(ULLMs)的GitHub应用中,有25%被归类为明确的恶意应用。
cs.AI / 87 / 2609.05245

Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory

大语言模型在数学推理中是否表现出连贯的知识结构?基于知识空间理论(Knowledge Space Theory)的视角
Cui, Peng, Do, Heejin, Sachan, Mrinmaya
Abstract
Human knowledge is inherently structured and interdependent: mastery of a concept requires prior mastery of its prerequisites, a principle formalized by Knowledge Space Theory (KST). While LLMs achieve strong performance on complex reasoning tasks, it remains unclear whether they exhibit coherent, human-like knowledge structure. We introduce a KST-grounded framework for evaluating LLM knowledge structure in mathematical reasoning, using it as a normative framework to analyze whether LLM behavior adheres to principled knowledge dependencies. Evaluating eight open- and closed-source LLMs against real human learners, we find that (1) LLMs do not adhere to human knowledge structure -- they frequently violate knowledge dependencies and fail to leverage related knowledge provided in context to improve performance on dependent questions; (2) LLMs do not share a consistent knowledge structure among themselves, as reflected by low overlap in their knowledge distributions. Furthermore, these structural deficiencies remain largely invisible to accuracy-based and LLM-as-judge evaluations. Together, our results provide behavioral evidence that current LLMs knowledge does not follow a human-like structure.
Chinese Translation
人类知识本质上是结构化且相互关联的:对一个概念的掌握依赖于对其先修概念的掌握,这一原则由知识空间理论(Knowledge Space Theory, KST)加以形式化。尽管大语言模型(LLM)在复杂推理任务上取得了出色表现,但它们是否表现出连贯的、类人的知识结构仍不清楚。我们提出了一个基于知识空间理论的评估框架,用于评估大语言模型在数学推理中的知识结构,并将其作为一个规范性框架来分析大语言模型的行为是否遵循原则性的知识依赖关系。通过将八个开源和闭源大语言模型与真实人类学习者进行对比评估,我们发现:(1) 大语言模型并不遵循人类的知识结构——它们频繁违反知识依赖关系,且无法利用上下文中提供的相关知识来提升在依赖性问题上的表现;(2) 大语言模型之间并不共享一致的知识结构,这体现在它们的知识分布重叠度较低。此外,这些结构性缺陷在基于准确率的评估和以大模型为裁判(LLM-as-judge)的评估中基本不可见。综上,我们的结果提供了行为学证据,表明当前大语言模型的知识并不遵循类人的结构。
cs.AI / 88 / 2609.05251

A Unified Physics-Aware Quantum Machine Learning Framework across Power GaN HEMTs and Logic Nanowire FETs: Predicting Unseen Process Splits and Held-Out Geometry Combinations with Lower Error and Tighter Split-to-Split Variability

面向功率GaN HEMTs与逻辑纳米线FETs的统一物理感知量子机器学习框架:以更低误差和更小分组间方差预测未见工艺分组与留出几何组合
Rai, Rushat, Wang, Yun-Yuan, Kakaen, Autsada, Chang, Pei-Jie, Nguyen, Doan Viet, Chiu, Yuan-Chieh, Tantraviwat, Doldet, Tumilty, Niall, See, Simon, Lee, Wen-Jay, Li, Tai-Yue, Chen, Nan-Yow, Wu, Tian-Li
Abstract
We present a unified reinforcement-learning (RL) framework that discovers compact parametrized quantum circuits (PQCs) for data-scarce device modeling. A graph neural network (GNN) policy optimized by proximal policy optimization (PPO) searches circuit architectures using leave-one-group-out cross-validation (LOGOCV) error on held-out process or geometry groups as the reward. The framework achieves the lowest mean absolute error (MAE) on all 11 targets versus six classical baselines, with 59% lower error (Ioff) and 81% tighter fold variability (VTH) for HEMTs and 84% lower error (VTH, SS, Ioff) and 82% tighter fold variability (Ioff) for NWFETs. These results demonstrate the potential of RL-selected, classically simulated PQCs as compact surrogates with low OOD error and improved physical consistency, despite imposing no explicit physical constraints, penalty terms, or device-specific equations, on the two evaluated device datasets.
Chinese Translation
我们提出了一个统一的强化学习(RL)框架,用于为数据稀缺的器件建模发现紧凑的参数化量子电路(PQCs)。该框架采用近端策略优化(PPO)优化的图神经网络(GNN)策略,以留出工艺或几何分组上的留一分组交叉验证(LOGOCV)误差作为奖励来搜索电路架构。在全部11个预测目标上,该框架相对于六个经典基线方法取得了最低的平均绝对误差(MAE):对于HEMTs,误差降低59%(Ioff),折间方差收紧81%(VTH);对于NWFETs,误差降低84%(VTH、SS、Ioff),折间方差收紧82%(Ioff)。这些结果表明,在两个被评估的器件数据集上,即使不施加任何显式物理约束、惩罚项或器件专用方程,经强化学习选择、经典模拟的PQCs仍有潜力作为具有低分布外(OOD)误差和更佳物理一致性的紧凑代理模型。
cs.AI / 89 / 2609.05257

Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions

计算机视觉中的常识推理:基础、最新进展与未来方向
Mahmud, Bahar Uddin, Barua, Sumit, Hong, Guan Yue, Gupta, Ajay, Liu, Hexu
Abstract
Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial for enhancing AI's understanding of everyday scenarios. This understanding not only improves machine learning models but also enhances their ability to interact meaningfully with humans and the environment. Unlike CNN-based conventional vision models, which are designed to identify objects within a specific image, incorporating commonsense knowledge enables models to interpret scenes in a more holistic manner, thereby improving their spatial ability to reason about relationships among objects and actions. This integration not only enhances object recognition but also facilitates a deeper understanding of the contextual factors, ultimately leading to more precise predictions and interactions in real-world applications. This paper presents a comprehensive survey of recent developments that integrate commonsense knowledge into computer vision tasks. We systematically review approaches based on knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers. We also outline current limitations related to dataset bias, knowledge incompleteness, and integration challenges. Finally, we highlight prospective research trajectories in cross-modal reasoning, scalable commonsense knowledge injection, and neuro-symbolic hybrid architectures to develop truly intelligent visual systems.
Chinese Translation
计算机视觉中的常识推理涵盖视觉数据与上下文知识的融合,这对于增强人工智能对日常场景的理解至关重要。这种理解不仅能改进机器学习模型,还能提升模型与人类及环境进行有意义交互的能力。与基于卷积神经网络(CNN)的传统视觉模型(其设计目标是识别特定图像中的物体)不同,融入常识知识使模型能够以更整体的方式解读场景,从而提升其对物体与动作之间关系进行空间推理的能力。这种融合不仅增强了物体识别能力,还有助于更深入地理解上下文因素,最终在现实应用中实现更精准的预测与交互。本文对将常识知识融入计算机视觉任务的最新研究进展进行了全面综述,系统地回顾了基于知识图谱、场景图、神经符号模型(neuro-symbolic models)以及常识增强Transformer的方法。我们还概述了当前在数据集偏差、知识不完备性以及融合挑战方面的局限性。最后,我们指出了跨模态推理、可扩展的常识知识注入以及神经符号混合架构等未来研究方向,以期构建真正智能的视觉系统。
cs.AI / 90 / 2609.05261

Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents

Trace2Tower:面向LLM智能体的过渡感知特征迹多层级技能归纳框架
Sun, Jiazheng, Yang, Boyu, Yuan, Binhao, Li, Mingxuan, Peng, Xin
Abstract
Large language model agents increasingly rely on execution traces to master complex interactive tasks. However, current paradigms are bottlenecked by shallow trajectory retrieval and flat skill summarization, fundamentally ignoring the temporal dependencies and outcome-conditioned topology of agent behavior. We introduce Trace2Tower, a transition-aware EigenTrace framework that distills raw trajectories into a robust skill hierarchy. Trace2Tower abstracts step-level interactions into canonical events, constructing a unified graph governed by semantic compatibility, transition dynamics, and outcome evidence. Through a novel contrastive spectral decomposition, it isolates stable, success-aligned behavioral modes while rigorously suppressing failure-prone shortcuts. These modes organically populate a dynamic skill tower of action templates, procedural routines, and overarching task strategies, continuously refined via verifier-guided feedback. On ALFWorld, Trace2Tower achieves 87.31% success requiring only 10.35 steps and 0.26 invalid actions; on WebShop, it reaches 50.67% exact success. Across both benchmarks, Trace2Tower significantly outperforms existing baselines in task mastery and context-efficient experience reuse.
Chinese Translation
大语言模型(LLM)智能体日益依赖执行轨迹来掌握复杂的交互式任务。然而,现有范式受限于浅层的轨迹检索与扁平化的技能总结,从根本上忽略了智能体行为中的时序依赖关系以及以结果为条件的拓扑结构。我们提出了Trace2Tower,一种过渡感知的EigenTrace(特征迹)框架,能够将原始轨迹提炼为稳健的技能层级体系。Trace2Tower将步骤级交互抽象为规范化事件,构建了一个由语义兼容性、转移动态和结果证据共同约束的统一图。通过一种新颖的对比谱分解方法,该框架能够分离出稳定的、与成功对齐的行为模式,同时严格抑制易导致失败的捷径行为。这些模式有机地填充到一个动态技能塔中,涵盖动作模板、程序化例程以及宏观任务策略,并通过验证器引导的反馈持续优化。在ALFWorld上,Trace2Tower达到87.31%的成功率,且仅需10.35步和0.26次无效操作;在WebShop上,它达到50.67%的精确成功率。在两个基准测试中,Trace2Tower在任务掌握能力和上下文高效的经验复用方面均显著优于现有基线方法。
cs.AI / 91 / 2609.05270

AI for Computational Design Science: A Responsible Human-AI Framework and Case Study on Short-Form Video Safety Surveillance

面向计算设计科学的人工智能:一个负责任的人机协作框架及短视频安全监测案例研究
Zhang, Wenli, Xie, Jiaheng, Pan, Zhihe, Chai, Yidong, Fang, Xiao, Ram, Sudha
Abstract
Artificial intelligence (AI) is transforming not only what information systems researchers design, but also how design research is conducted. Yet existing literature offers limited guidance for computational design science (CDS) when AI actively participates in problem formulation, resource construction, design search, evaluation, and knowledge abstraction. We develop AI for Computational Design Science (AI4CDS), a five-phase methodological framework in which AI expands problem and design search while researchers retain responsibility for domain grounding, admissibility, verification, and scientific judgment. Collaboration is governed by graduated trust, reversibility, auditability, and differentiated reproducibility. We instantiate AI4CDS through ChildRiskGuard, an interpretable artifact for detecting short-form videos inappropriate for children, while documenting AI interactions, rejected alternatives, corrections, and audit trails. The case translates audience-dependent safety and explanation faithfulness into three technical challenges and develops an artifact that separates generic from child-specific risk, represents distinct developmental-risk mechanisms, and makes concept-level explanations part of the predictive computation. ChildRiskGuard achieves an F1 score of 0.769, substantially outperforming direct application of a general-purpose content-safety model while remaining competitive with strong benchmarks. The primary contribution is AI4CDS as a responsible framework for AI-enabled CDS; ChildRiskGuard provides process and artifact evidence of how AI-expanded, researcher-governed design can generate and evaluate novel computational design knowledge.
Chinese Translation
人工智能(AI)不仅正在改变信息系统研究者所设计的内容,也在改变设计研究的开展方式。然而,当AI积极参与问题表述、资源构建、设计搜索、评估与知识抽象时,现有文献对计算设计科学(Computational Design Science, CDS)的指导仍然有限。我们提出了面向计算设计科学的人工智能(AI4CDS),这是一个五阶段的方法论框架,其中AI扩展了问题空间与设计搜索,而研究者则保留对领域扎根、可采纳性、验证与科学判断的责任。协作由分级信任、可逆性、可审计性和差异化的可复现性原则来规范。我们通过ChildRiskGuard对AI4CDS进行了实例化,ChildRiskGuard是一个用于检测不适合儿童观看的短视频的可解释制品,并记录了AI交互过程、被否决的备选方案、修正内容及审计轨迹。该案例将与受众相关的安全性和解释忠实性转化为三个技术挑战,并构建了一个能够区分一般风险与儿童特有风险、表征不同发展性风险机制、并将概念层面解释纳入预测计算的制品。ChildRiskGuard取得了0.769的F1分数,显著优于直接应用通用内容安全模型的效果,同时与强基准方法相比仍具竞争力。本文的主要贡献是AI4CDS作为一个面向AI赋能计算设计科学的负责任框架;ChildRiskGuard则为AI扩展、研究者主导的设计如何生成并评估新型计算设计知识提供了过程与制品层面的证据。
cs.AI / 92 / 2609.05275

Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference

不要丢弃Dropout:优化层稀疏性以实现高效的大语言模型训练与推理
Elhoushi, Mostafa, Pretko, Alex, Dey, Nolan, Zhang, Bin Claire, Gray, Gavia, Gosal, Gurpreet, Mahmoud, Abdulrahman, Bergsma, Shane, Hestness, Joel
Abstract
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout - particularly layer dropout - has largely disappeared from large language models (LLMs) pre-training recipes. While some prior work has reported that dropout can degrade accuracy, no comprehensive study has quantified, let alone mitigated, this effect. In this study, we show that layer dropout should be used in state-of-the-art LLM training, establishing best practices and scaling analysis for both training and post-training benefits. Concretely, with optimal layer distribution, time schedule, and optimizer hyperparameters, we observe that at the same training FLOPs layer dropout leads to lower loss. For a given number of training steps, LLMs can achieve lower or similar validation loss while saving upto 25% of training FLOPs. Moreover, layer dropout enables significant post-training optimizations, such as early exit, intermediate-layer skipping, and self-speculative decoding, yielding up to 1.5x inference speedup with negligible accuracy loss. Across more than 2400 training experiments, spanning models from 271M to 8.2B parameters and datasets up to 160B tokens, we demonstrate that these findings extend reliably to large-scale training regimes. All pre-training experiments were run on Cerebras CS-3 systems.
Chinese Translation
层Dropout(又称随机深度)已被证明能够在语言和视觉Transformer中实现更快的训练速度、更高的准确率,以及对零样本层剪枝的鲁棒性。然而,随着模型和数据集规模的不断扩大,Dropout——尤其是层Dropout——已在很大程度上从大语言模型(LLM)的预训练方案中消失。尽管一些先前的研究报告称Dropout会降低准确率,但尚无全面的研究对这种影响进行量化分析,更不用说提出缓解方法了。在本研究中,我们证明了层Dropout应当被应用于最先进的大语言模型训练中,并为其在训练和后训练阶段的收益建立了最佳实践和规模分析。具体而言,通过采用最优的层分布、时间调度和优化器超参数,我们观察到在相同训练FLOPs的情况下,层Dropout能够带来更低的损失。在给定训练步数的情况下,大语言模型可以在节省高达25%训练FLOPs的同时,取得更低或相近的验证损失。此外,层Dropout还支持显著的后训练优化,例如提前退出(early exit)、中间层跳过和自推测解码(self-speculative decoding),在准确率损失可忽略不计的情况下实现高达1.5倍的推理加速。通过超过2400项训练实验,涵盖参数量从271M到8.2B的模型以及规模高达1600亿token的数据集,我们证明了这些发现可以可靠地推广到大规模训练场景中。所有预训练实验均在Cerebras CS-3系统上运行。
cs.AI / 93 / 2609.05279

Testing Interchangeability in LLM Agent Teams

测试大语言模型智能体团队中的可互换性
Gao, Jianxin, Yu, Tianyi, Deng, Linna, Li, Runze, Wang, Zining
Abstract
Production multi-agent systems replace agents constantly, on the assumption that an agent filling a role is interchangeable with any other agent that can do the job. We test that assumption. Eight teams per setting are formed independently from one base model on the same tasks, each agent keeping a private notebook across ten formation episodes; we then trade role-matched agents between teams and measure what changes on held-out tasks. Against a placebo that reproduces the disruption of a roster change without changing who occupies the seat, a swap costs little in task score but raises the communication a team spends per unit of progress by 16 to 63 percent, and in Hanabi a swapped agent is more expensive than an inexperienced one, consistent with interference from conventions learned with its former partner. In Collab-Overcooked, when the agent that sets the agenda is replaced, most of the extra communication comes from the agent that stayed. Three ablations, over base models, decoding temperature and formation length, move the swap penalty alongside one other quantity: how far independently formed teams drift apart. Greedy decoding lowers both; doubling a team's history raises both. In these settings, agents are more fungible in task outcome than in coordination efficiency, with larger swap effects after longer formation histories.
Chinese Translation
生产环境中的多智能体系统经常替换智能体,其前提假设是:担任某一角色的智能体与任何能胜任该工作的其他智能体是可互换的。我们对这一假设进行了检验。在每个实验设置中,八个团队均由同一个基础模型在同一任务上独立组建,每个智能体在十个组建回合中保留私人笔记;随后我们在团队之间交换角色匹配的智能体,并测量在保留任务上发生的变化。与一种仅复现换人带来的干扰但不改变实际在位者的安慰剂对照相比,交换智能体对任务得分的影响很小,但会使团队在单位进展上花费的通信量增加16%至63%;并且在Hanabi中,被换入的智能体比缺乏经验的智能体代价更高,这与它从先前搭档处习得的约定(conventions)造成干扰的假设相符。在Collab-Overcooked中,当设定议程的智能体被替换时,大部分额外的通信来自留下的那个智能体。三项消融实验——分别改变基础模型、解码温度和组建时长——表明交换惩罚与另一个量同步变化:独立组建的团队之间漂移分离的程度。贪婪解码使两者同时降低;将团队历史加倍则使两者同时升高。在这些设置中,智能体在任务结果上比在协调效率上更具可替代性,且组建历史越长,交换效应越大。
cs.AI / 94 / 2609.05284

GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity

GUT:基于图复杂度量化与优化大语言模型推理不确定性的方法
Liang, Shuang, Hu, Xin-Yu, Ou, Xiang-Jun, Zhang, Shao-Qun
Abstract
Recent years have witnessed great advances in the reasoning ability of Large Language Models (LLMs). However, the reasoning processes of LLMs often exhibit uncertainty, where LLMs often produce a proliferation of divergent branches at each reasoning step even when fed the same prompting inputs, and certain branches exhibit evidently incredible, even nonsensical, reasoning chains and results. In this paper, we propose the Graph-complexity-based UncerTainty (GUT) method for investigating the reasoning uncertainty of LLMs. The key idea of GUT is to characterize the potential branches of each reasoning chain with a directed acyclic graph, thereby ensuring that all potential branches are comprehensively covered within the graph space. Building upon this recognition, we further build two modules of GUT, that is, a Quantification (GUT-Q) module and an Optimization (GUT-O) module, for quantifying and reducing the reasoning uncertainty of LLMs, respectively. GUT-Q measures LLM reasoning uncertainty by approximating the reasoning space complexity with graph complexity. GUT-O implements uncertainty optimization by treating negative uncertainty as the reward function in reinforcement learning. Experimental results conducted on four LLMs and five datasets validate the effectiveness of GUT.
Chinese Translation
近年来,大语言模型(LLM)的推理能力取得了长足进步。然而,大语言模型的推理过程往往表现出不确定性:即使输入相同的提示,模型在每一步推理中也常常产生大量发散的分支,其中某些分支呈现出明显不可信甚至荒谬的推理链和结果。本文提出了基于图复杂度的不确定性(Graph-complexity-based UncerTainty,GUT)方法,用于研究大语言模型的推理不确定性。GUT的核心思想是用有向无环图来刻画每条推理链的潜在分支,从而确保所有潜在分支都能被全面覆盖在图空间内。基于这一认识,我们进一步构建了GUT的两个模块,即用于量化推理不确定性的量化模块(GUT-Q)和用于降低推理不确定性的优化模块(GUT-O)。GUT-Q通过用图复杂度近似推理空间复杂度来度量大语言模型的推理不确定性;GUT-O则在强化学习中把负不确定性作为奖励函数来实现不确定性优化。在四个大语言模型和五个数据集上进行的实验结果验证了GUT的有效性。
cs.AI / 95 / 2609.05289

Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

超越聚合分数:用于评估基于参考文本的自动评价方法的行为正确性假设
Mahbub, Maria, Rice, Ashley, Munroe, Michael R., Kamara, Amidu, Sadovnik, Amir
Abstract
Automated reference-based evaluation methods play a critical role in assessing natural language generation systems. Existing meta-evaluation primarily measures agreement with human judgments or benchmark labels, providing limited insight into evaluator behavior under controlled conditions. We introduce behavioral correctness assumptions, a complementary framework for evaluating reference-based automatic evaluation methods. We define a taxonomy of correctness-preserving and correctness-altering assumptions and operationalize them through controlled response transformations that specify expected scoring behaviors. We evaluate diverse lexical, character-level, semantic, LLM-based, and hybrid evaluators and analyze their assumption-level behavior, stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility. Our experiments reveal distinct behavioral trade-offs across evaluation paradigms: no evaluator satisfies all proposed correctness assumptions, and evaluators with similar aggregate performance can exhibit substantially different behavioral profiles. These findings demonstrate that behavioral correctness assumptions provide diagnostic information obscured by conventional aggregate meta-evaluation.
Chinese Translation
基于参考文本的自动评价方法在评估自然语言生成系统方面发挥着关键作用。现有的元评价主要衡量其与人类判断或基准标签的一致性,对评价器在受控条件下的行为洞察有限。我们提出行为正确性假设,作为评估基于参考文本的自动评价方法的补充框架。我们定义了一个由保持正确性和改变正确性假设构成的分类体系,并通过规定预期打分行为的受控响应变换对其进行操作化。我们评估了多种词面级、字符级、语义级、基于大语言模型(LLM)以及混合型的评价器,并分析其在假设层面的行为、稳定性、敏感性、重复运行的变异性、配置敏感性以及可复现性。实验揭示了不同评价范式之间明显的行为权衡:没有任何评价器能满足所有提出的正确性假设,且聚合性能相近的评价器可能表现出截然不同的行为特征。这些发现表明,行为正确性假设能够提供被传统聚合式元评价所掩盖的诊断信息。
cs.AI / 96 / 2609.05295

RISE: Recursive Improvement via Self-Extrapolating Policy Distillation

RISE:通过自外推策略蒸馏实现递归改进
Li, Yang, Yavuz, Semih, Joty, Shafiq
Abstract
On-policy distillation (OPD) provides dense, per-token supervision for language model post-training, but its effectiveness is bottlenecked by teacher quality: external teachers suffer from distribution mismatch, while self-distillation with privileged conditioning is limited by in-context learning capacity. We propose \textbf{RISE} (\textbf{R}ecursive \textbf{I}mprovement via \textbf{S}elf-\textbf{E}xtrapolating Policy Distillation), which constructs a synthetic teacher directly from the model's own RLVR training trajectory. By extrapolating the displacement between the current checkpoint and a trailing anchor---in parameter space or output logit space---RISE converts a sparse outcome-induced parameter update into a dense token-level target, without any external model or privileged conditioning. RISE combines RLVR and OPD in a complementary loop: outcome rewards ground the extrapolation toward correct reasoning, while the extrapolated teacher refines token-level decisions. Moreover, since the teacher is refreshed every iteration as the student improves, distillation becomes a recursive improvement mechanism rather than a one-shot compression step. Experiments spanning mathematical reasoning, multi-domain STEM, code generation, and multi-turn agentic tasks show that RISE outperforms RLVR-only training and on-policy self-distillation across all settings.
Chinese Translation
在策略蒸馏(On-policy Distillation, OPD)为语言模型后训练提供了密集的逐 token 监督信号,但其有效性受限于教师模型的质量:外部教师存在分布不匹配问题,而基于特权条件(privileged conditioning)的自蒸馏则受限于上下文学习能力。我们提出了 RISE(Recursive Improvement via Self-Extrapolating Policy Distillation,通过自外推策略蒸馏实现递归改进),该方法直接从模型自身的 RLVR 训练轨迹构建一个合成教师。通过在参数空间或输出 logit 空间中外推当前检查点与滞后锚点之间的位移,RISE 将稀疏的结果诱导参数更新转化为密集的 token 级目标,且无需任何外部模型或特权条件。RISE 将 RLVR 和 OPD 结合在一个互补的循环中:结果奖励将外推锚定于正确推理,而外推得到的教师则细化 token 级决策。此外,由于教师随学生的进步在每次迭代中都被刷新,蒸馏成为一种递归改进机制,而非一次性的压缩步骤。在数学推理、多领域 STEM、代码生成以及多轮智能体任务上的实验表明,RISE 在所有设置下均优于仅使用 RLVR 的训练和在策略自蒸馏方法。
cs.AI / 97 / 2609.05314

Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness

大语言模型在建筑能源系统暖通空调(HVAC)运行中的应用:方法、应用与部署就绪度的批判性综述
Neubauer, Alexander, Hong, Tianzhen, Li, Han, Yu, Mengbo, Darbandi, Amin, Fürst, Yannick, Kriegel, Martin
Abstract
Building automation systems generate rich sensor data yet remain insight-poor because heterogeneous point naming, missing metadata, and fragmented documentation obstruct their operational use. This systematic review analyses and codes 66 peer-reviewed studies on large language models (LLMs) for HVAC operations published between 2023 and March 2026. Each study is classified across five application families and three LLM method families and assessed for evidence realism, deployment readiness, and the responsibility boundary between the LLM and physical HVAC decisions. The corpus is concentrated in building energy modelling (BEM, 32 of 66 papers), while load forecasting remains too sparse for subfield-level conclusions. Only four studies reach pilot-level evidence, and none reports sustained operational deployment. No study was classified as ready-now for industry adoption; three were near-term and 63 research-only. Nevertheless, several bounded, human-in-the-loop uses merit near-term trials, including point-name normalisation, document-grounded operator support, BEM workflow assistance, and advisory interfaces around physics-based controllers. Conventional machine learning (ML), model predictive control (MPC), reinforcement learning (RL) and ontology-based tools remain more adopted for high-frequency control, short-horizon numerical forecasting, and well-posed ontology mapping, while autonomous agentic operation and unvalidated occupant proxies remain research-stage. Current evidence therefore supports LLMs primarily as semantic and workflow layers rather than autonomous HVAC controllers. Future work should prioritise field-validated benchmarks, orchestration evaluation under operational constraints, and LLM-MPC/RL architectures with bounded latency and verifiable safety properties.
Chinese Translation
建筑自动化系统产生丰富的传感器数据,却长期处于“洞察贫乏”状态,原因在于异构的点名称命名、缺失的元数据以及碎片化的文档阻碍了其运营应用。本系统性综述分析并编码了2023年至2026年3月间发表的66篇关于大语言模型(LLM)应用于暖通空调(HVAC)运行的同行评审研究。每项研究被归入五大应用类别和三大LLM方法类别,并从证据真实性、部署就绪度以及LLM与物理HVAC决策之间的责任边界三个方面进行评估。研究文献主要集中于建筑能源建模(BEM,66篇中占32篇),而负荷预测领域的文献数量过于稀少,无法得出子领域的结论。仅有四项研究达到了试点级证据水平,没有任何研究报告了持续的实际运营部署。没有研究被归类为“即可就绪”的行业采用水平;三项属于近期可用,63项仍处于纯研究阶段。尽管如此,若干有边界约束、人在回路中的应用值得近期开展试点,包括点名称规范化、基于文档的操作员支持、BEM工作流辅助,以及围绕基于物理模型控制器的咨询式交互界面。传统机器学习(ML)、模型预测控制(MPC)、强化学习(RL)以及基于本体的工具在高频控制、短期数值预测和良定义的本体映射方面仍更具优势,而自主智能体运行和未经校验的住户代理仍处于研究阶段。因此,当前证据支持将LLM主要用作语义层和工作流层,而非自主的HVAC控制器。未来研究应优先开展经过实地验证的基准测试、运营约束下的编排评估,以及具有有界延迟和可验证安全属性的LLM-MPC/RL架构研究。
cs.AI / 98 / 2609.05327

LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams

基于二元决策图的量子电路合成的LLM驱动算法设计
Sim, Yoonju, Berto, Federico, Hua, Chuanbo, Park, Jinkyoo, Kwon, Changhyun
Abstract
Quantum circuits are central to implementing quantum algorithms on quantum devices, where quantum gates must be reversible. Many quantum algorithms rely on Boolean functions, which must therefore be implemented reversibly within quantum circuits. Reversible circuit synthesis provides a way to translate such Boolean functions into reversible circuits. Binary decision diagrams (BDDs) offer a scalable approach to this task, but the resulting BDDs and circuits depend heavily on variable ordering. Existing ordering heuristics commonly minimize BDD size because it is closely tied to the circuit size. However, BDD size is an imperfect proxy for the quantum cost of the synthesized circuit (QCC). We propose \texttt{QuantumEvo}, an evolutionary framework that uses an LLM as a heuristic generator for QCC-aware BDD variable ordering. Instead of predicting orderings directly, \texttt{QuantumEvo} searches over ordering heuristics initialized from multiple heuristic families. Candidate heuristics directly manipulate variable orderings using standard BDD operations and are selected by downstream QCC. The discovered heuristic, HGA-QE, modifies the sifting step inside a genetic algorithm so that the procedure is better aligned with QCC. Across the benchmark set, HGA-QE achieves a 70.9\% tie-or-win rate against the per-function best baseline and is strictly best on 13.5\% of the functions. The results demonstrate broadly competitive QCC performance, with HGA-QE showing a clearer relative advantage in strict wins on the two benchmark suites drawn from sources different from the data used for heuristic discovery.
Chinese Translation
量子电路是在量子设备上实现量子算法的核心,其中量子门必须是可逆的。许多量子算法依赖于布尔函数,因此这些布尔函数必须在量子电路中以可逆方式实现。可逆电路合成提供了一种将此类布尔函数转换为可逆电路的方法。二元决策图(BDD)为这一任务提供了一种可扩展的途径,但所得到的BDD和电路在很大程度上依赖于变量排序。现有的排序启发式方法通常以最小化BDD规模为目标,因为BDD规模与电路规模密切相关。然而,BDD规模并不能完美地替代合成电路的量子代价(QCC)。我们提出了QuantumEvo,一个以大语言模型(LLM)作为启发式生成器的进化框架,用于面向QCC的BDD变量排序。QuantumEvo并不直接预测排序,而是在从多个启发式族初始化的排序启发式空间中进行搜索。候选启发式方法使用标准BDD操作直接操纵变量排序,并根据下游的QCC进行选择。所发现的启发式方法HGA-QE修改了遗传算法中的筛选(sifting)步骤,使该过程更好地与QCC保持一致。在整个基准测试集上,HGA-QE相对每个函数的最优基线达到了70.9%的持平或胜出率,并在13.5%的函数上严格最优。结果表明HGA-QE具有广泛竞争力的QCC性能,在两个与启发式发现所用数据来源不同的基准测试套件上,其在严格胜出方面展现出更为明显的相对优势。
cs.AI / 99 / 2609.05333

Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

用于测量Transformer语言模型中语境化个体化程度的工具包技术手册
Marques, José Luciano Verçosa, Heitmann, Frederico Jorge, Perez, Daniel Omar, de Paula, Marcelo Vinicius, Barros, Tárcio André dos Santos
Abstract
A transformer language model assigns a single, context-independent vector to a word type at its embedding layer, yet is widely believed to individuate that word's occurrences by context in its later layers. Testing this belief cleanly requires a construct that holds the word form fixed while its context and intended sense vary in a controlled, labeled way. This manual documents an open toolkit built around such a construct, which we call a bridge form: a single written word that recurs, unchanged, across two or more subject domains with a different sense in each. We describe, and justify, every stage of the pipeline: the declarative specification of bridge forms and their source domains, corpus acquisition from Wikipedia, occurrence localization, layer-wise representation extraction, a domain-pairwise silhouette measurement of separation in the model's representation space, and a paired visualization protocol. Each design choice is presented together with the methodological failure mode it is meant to avoid (sense contamination from overly broad category labels, the multi-group bias of the silhouette coefficient, subword-tokenization misalignment, and axis-comparability artifacts in dimensionality-reduced plots, among others). This manuscript is a methodological and implementation reference: it does not report or interpret empirical outcomes of running the toolkit on any particular model or bridge-form set. The toolkit, its full source, and the corpora used to exercise it are archived separately (Section 9) under a persistent identifier, and are intended to be cited as an instrument by studies that use it to produce and interpret empirical results.
Chinese Translation
Transformer语言模型在其嵌入层为词型(word type)分配单一且与语境无关的向量,但人们普遍认为,该模型会在其后续各层中依据语境对该词的各个出现实例进行个体化区分。要对这一观点进行严谨的检验,需要一种构造:在保持词形不变的同时,使其语境和预期词义以受控且带标签的方式发生变化。本手册记录了一个围绕此类构造构建的开放工具包,我们将其称为桥接词形(bridge form):一个在两个或多个主题领域中重复出现且书写形式完全相同的单词,但其在各领域中的词义不同。我们描述并论证了整个流程的每个阶段:桥接词形及其源领域的声明式规范、来自维基百科(Wikipedia)的语料获取、出现位置定位、逐层表示提取、基于领域成对的轮廓系数(silhouette)测量模型表示空间中的分离度,以及配对可视化协议。每项设计选择都与其旨在避免的方法论失效模式一并呈现(例如:过于宽泛的类别标签导致的词义污染、轮廓系数的多组偏差、子词分词错位,以及降维图中的轴可比性伪影等)。本文是一份方法论与实现参考文档:它不报告或解释在任何特定模型或桥接词形集合上运行该工具包所得到的实证结果。该工具包、其完整源代码以及用于测试的语料库均以持久标识符单独存档(第9节),旨在作为一种研究工具被相关研究引用,以产生和解释实证结果。
cs.AI / 100 / 2609.05339

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

你的智能体记忆能在模型升级后幸存吗?一项关于记忆可移植性的受控研究
Goyal, Ankit, Ray, Jaideep
Abstract
Model upgrades are routine; memory migrations are not. An agent can keep the same memory store and still forget: a new model may interpret old notes differently, mixed embedding versions may break retrieval, and repair may fail without the original evidence. We compare memory as the same history is preserved verbatim for long-context reading (LC-RAW), divided into chunks for retrieval-augmented generation (RAG), compressed by a model into natural-language notes (NOTES), or normalized into a fixed-schema knowledge graph (KG-fixed). The study uses 48 synthetic histories with randomized answer codes, exact scoring, and two open-weight models with sub 10 billion parameters. Our measurements show that fixed-schema structures transfer reliably, with KG-fixed accuracy changing by only $+0.0004 \pm 0.0020$ following a writer swap. Conversely, compressed NOTES exhibit high model coupling, with accuracy shifting asymmetrically by $+9.91$ or $-13.28$ percentage points depending on the specific migration direction. In RAG systems, partial embedding migrations using a 50/50 mixed index capture only a 4.96-point accuracy improvement, forfeiting the majority of the 11.90-point gain achieved through full re-embedding. Diagnostic decomposition attributes 80% ($0.467 \pm 0.014$) of the NOTES accuracy deficit to information lost during initial construction, whereas retrieval failures drive 81% ($0.364 \pm 0.012$) of the RAG deficit. Finally, store-only repair of NOTES fails to reach a 90% performance recovery target in all 48 test cases, whereas retaining the raw source history enables successful recovery in 34 of 48 cases for one tested direction. These findings highlight the necessity of direction-specific migration testing, strict embedding space isolation, and the retention of source histories for memory repair.
Chinese Translation
模型升级司空见惯,而记忆迁移却并非如此。即使智能体保留同一个记忆库,仍可能遗忘:新模型可能以不同方式解读旧笔记,混合的嵌入(embedding)版本可能破坏检索,而在缺乏原始证据时修复可能失败。我们在保留相同历史记录的条件下,比较了四种记忆形式:原文完整保存以供长上下文阅读(LC-RAW)、分块用于检索增强生成(RAG)、由模型压缩为自然语言笔记(NOTES),以及规范化为固定模式的(fixed-schema)知识图谱(KG-fixed)。该研究使用48条带有随机化答案编码的合成历史记录、精确评分,以及两个参数量低于100亿的开源权重模型。测量结果表明,固定模式结构能够可靠迁移:在写入器(writer)更换后,KG-fixed的准确率仅变化 $+0.0004 \pm 0.0020$。相反,压缩后的NOTES表现出高度的模型耦合性,其准确率随迁移方向的不同而不对称地偏移 $+9.91$ 或 $-13.28$ 个百分点。在RAG系统中,采用50/50混合索引的部分嵌入迁移仅获得4.96个百分点的准确率提升,损失了完全重新嵌入所能带来的11.90个百分点增益中的大部分。诊断性分解显示,NOTES准确率缺陷的80%($0.467 \pm 0.014$)归因于初始构建过程中的信息丢失,而检索失败则导致了RAG缺陷的81%($0.364 \pm 0.012$)。最后,仅对NOTES进行存储端修复在全部48个测试用例中均未能达到90%的性能恢复目标,而保留原始历史记录则使其中一个测试迁移方向在48个用例中的34个实现了成功恢复。这些发现凸显了方向特定的迁移测试、严格的嵌入空间隔离,以及保留原始历史记录以供记忆修复的必要性。
cs.AI / 101 / 2609.05346

Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education

谁来评阅我的作业?高等教育中学生对透明化AI辅助写作评价的看法
AlGhamdi, Rayed
Abstract
The integration of GenAI tools into higher education assessment raises important questions about how students understand, interpret, and respond to AI-mediated evaluation. As instructors increasingly explore AI tools for providing feedback, prior research has examined whether GenAI-generated feedback improves writing performance and how students perceive its usefulness; comparatively little is known, however, about how students interpret such evaluation when they are explicitly informed that an AI system, rather than a human instructor, produced the feedback and the score. This study reports findings from a qualitative pedagogical inquiry conducted in an undergraduate technical communication course for computing students at a Saudi public university. Thirteen male undergraduate computing students completed an in-class handwritten writing task; the scanned submissions were evaluated by ChatGPT using a rubric-based prompt aligned with the task objectives. Students were then explicitly informed that ChatGPT had generated the score and feedback and were invited to reflect on the evaluation in writing. Inductive thematic analysis of these reflections identified four themes: perceived usefulness of feedback; awareness of AI's contextual and pedagogical limitations; conditional trust, distinguishing feedback utility from evaluative authority; and reflection on the institutional and pedagogical role of the human instructor. Participants accepted GenAI feedback as useful for surface-level revision but consistently positioned the human instructor as the appropriate authority over grading decisions. The study identifies this as a distinction between feedback utility and evaluative authority, two judgments that students treat as analytically separate rather than as opposite ends of a single approval scale...
Chinese Translation
生成式人工智能(GenAI)工具融入高等教育评价,引发了关于学生如何理解、解读和回应AI中介评价的重要问题。随着教师日益探索使用AI工具提供反馈,已有研究考察了GenAI生成的反馈是否能提升写作表现,以及学生如何感知其有用性;然而,当学生被明确告知反馈和分数是由AI系统而非人类教师给出时,他们如何解读此类评价,相关研究仍然较少。本研究报告了一项定性教学探究的结果,该探究在沙特一所公立大学面向计算机专业本科生的技术传播课程中开展。十三名男性计算机专业本科生完成了一项课内手写写作任务;扫描后的作业提交件由ChatGPT依据与任务目标一致的量规提示词进行评价。随后,学生被明确告知分数和反馈由ChatGPT生成,并被邀请以书面形式反思该评价。对这些反思的归纳式主题分析识别出四个主题:对反馈有用性的感知;对AI在情境性和教学性局限方面的认识;有条件的信任,即区分反馈的有用性与评价的权威性;以及对人类教师的制度性和教学性角色的反思。参与者认为GenAI反馈对于表层修改是有用的,但一致地将人类教师定位为评分决策的正当权威。本研究将此识别为反馈有用性与评价权威性之间的区分——学生将这两种判断视为分析上相互独立的,而非同一认可尺度上的两个对立端。
cs.AI / 102 / 2609.05374

CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents

CUA-Universe:面向混合 GUI+CLI 智能体的可扩展动态环境
Shi, Haoting, Wang, Wenhao, Fang, Weicheng, Liang, Yaozhong, Jin, Tian, Zhao, Pengxiang, Liu, Guangyi, Chen, Siheng, Wang, Yanfeng
Abstract
Computer-use agents have advanced on benchmarks like OSWorld and AndroidWorld, but still act mostly through the GUI, often producing inefficient trajectories. Real-world computer work is hybrid, combining visual-state inspection with precise, high-throughput command-line operations, so capable agents must coordinate both modalities over shared application state. Yet scalable hybrid environments remain scarce because supporting both GUI and CLI over real applications typically requires substantial manual engineering for each application. Existing agents also struggle to use the two interfaces complementarily: CLI-native agents lack visual perception for tasks involving interface state or layout, while GUI-native agents are inefficient for operations better executed through commands. We introduce CUA-Universe, a scalable environment-to-data pipeline that turns real desktop software into hybrid GUI+CLI environments. App-Forge adapts applications into reproducible VMs and command-line surfaces it discovers, wraps, or generates, scaling to 16 applications; Task-Weave synthesizes diverse hybrid tasks of controllable difficulty from reusable operations over seed files; and Path-Steer steers rollouts along efficient hybrid paths and harvests verified trajectories for post-training. Training on this data shifts behavior from inefficient GUI interaction and brittle CLI scripting toward effective GUI+CLI orchestration. Our 9B model improves both success and efficiency on CUA-Verse (Score +39.3 pts; -37% steps, -60% tokens), OSWorld (SR +16.8 pts; -57% steps, -44% tokens), and OSWorld-MCP (Score +7.84 pts; -27% steps, -30% tokens). CUA-Universe provides a scalable path toward more capable and efficient computer-use agents.
Chinese Translation
计算机使用智能体在 OSWorld 和 AndroidWorld 等基准测试上已取得进展,但它们大多仍通过图形用户界面(GUI)进行操作,往往产生低效的执行轨迹。现实中的计算机工作是混合式的,既需要视觉状态检查,也需要精确、高吞吐的命令行(CLI)操作,因此具备能力的智能体必须在共享的应用状态之上协调这两种模态。然而,可扩展的混合环境仍然稀缺,因为在真实应用程序上同时支持 GUI 和 CLI 通常需要针对每个应用进行大量人工工程开发。现有智能体也难以互补地使用这两种接口:CLI 原生的智能体缺乏处理涉及界面状态或布局任务的视觉感知能力,而 GUI 原生的智能体在执行更适合通过命令完成的操作时效率低下。我们提出了 CUA-Universe,这是一个从环境到数据的可扩展流水线,能够将真实的桌面软件转化为混合 GUI+CLI 环境。App-Forge 将应用程序适配为可复现的虚拟机,并对其发现、封装或生成的命令行接口进行扩展,支持多达 16 个应用;Task-Weave 基于对种子文件的可复用操作,合成难度可控的多样化混合任务;Path-Steer 则引导 rollout 沿高效混合路径进行,并收集经验证的轨迹用于后训练。基于这些数据进行训练,使智能体的行为从低效的 GUI 交互和脆弱的 CLI 脚本编写转向有效的 GUI+CLI 协同编排。我们的 9B 模型在 CUA-Verse(得分 +39.3 分;步数 -37%,token -60%)、OSWorld(成功率 +16.8 分;步数 -57%,token -44%)和 OSWorld-MCP(得分 +7.84 分;步数 -27%,token -30%)上均同时提升了成功率和效率。CUA-Universe 为构建更强大、更高效的计算机使用智能体提供了一条可扩展的路径。
cs.AI / 103 / 2609.05381

Molecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

分子的既视感:前沿语言模型中已发表数值的数字级检索
Busch, Matthias, Tacke, Marius, Lamaka, Sviatlana V., Zheludkevich, Mikhail L., Cyron, Christian J., Aydin, Roland C., Feiler, Christian
Abstract
Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number. We audit 22 frontier models on 12 regression benchmarks for verbatim retrieval and find that it is widespread but relatively benchmark-specific: on five datasets more than $50\%$ of the LLMs show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. We run our experiments at two reasoning levels and find that reasoning changes retrieval. The same experiments, on the same molecules and with the same prompt, are flagged $89\%$ more often at the higher reasoning level than at the lowest one. Finally, we test a way to interrupt retrieval in our most contaminated cases, and find that the strongest models in some cases still recognise a combination of transformed SMILES strings and original labels. Furthermore, suppressing retrieval moves the prediction errors of the different models closer together in relative terms, while their differing use of verbatim retrieval spreads them apart. This indicates that the general predictive capability of an LLM is not determined solely by the amount of memorised values. This work provides an overview of the amount and depth of verbatim retrieval in molecular regression benchmarks using LLMs.
Chinese Translation
大语言模型(LLM)越来越多地在分子性质基准上进行评估,但准确率无法区分一个模型究竟是真正在预测性质,还是仅仅检索了已发表的数值。我们在12个回归基准上对22个前沿模型进行了逐字检索审计,发现这种现象广泛存在但具有较强的基准特异性:在五个数据集上,超过50%的LLM表现出逐字检索行为,而在其余数据集上,这种现象仅零星出现。我们在两个推理层级上开展实验,发现推理会改变检索行为:在相同分子、相同提示词的条件下,实验在较高推理层级被标记的频率比较低推理层级高出89%。最后,我们在污染最严重的案例中测试了一种中断检索的方法,发现最强的模型在某些情况下仍能识别出经过变换的SMILES字符串与原始标签的组合。此外,抑制检索会使不同模型的预测误差在相对意义上更加接近,而它们对逐字检索的不同使用程度则会使误差彼此拉开。这表明LLM的通用预测能力并非仅由记忆数值的多少所决定。本研究全面概述了LLM在分子回归基准中逐字检索的规模与深度。
cs.AI / 104 / 2609.05385

Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence

必要性还是充分性?基于行为证据评估大语言模型(LLM)的解释
Pawar, Urja, Ramanayake, Rajitha, Kemal, Nabeel, Kandath, Ashwin, O'Neill, Owen, Bourgeon, Guillaume, Chatbri, Houssem
Abstract
LLM decision components that can operate within agent workflows often produce action-relevant recommendations or judgements together with explanations. Operators may use the named factors to monitor a system, diagnose errors, or decide when to escalate an output. Such use assumes that the explanations agree with the component's observable decision behaviour. We test two interpretations of the named factors: necessity, meaning that changing a factor would change the output, and sufficiency, meaning that retaining it while removing other changeable information would preserve the output. We evaluate these interpretations in two synthetic use cases: recommending advisors to clients and judging prompts for harmfulness or risk. Models return an output and the top three factors that most influenced it. Controlled black-box interventions estimate a necessity score for each factor by measuring how often changing it changes the output, and a sufficiency score by measuring how often retaining it preserves the output. Across eight models from the Claude, GPT, and Gemini families, the mean Spearman correlations between the cited ranking and the necessity and sufficiency scores are 0.349 and 0.354 for advisor recommendation, and 0.431 and 0.580 for prompt monitoring. Furthermore, an uncited factor scores above the lowest-scoring cited factor in 57.6% of advisor responses under necessity and 58.1% under sufficiency; the corresponding prompt-monitoring rates are 25.8% and 8.9%. The cited top three contain useful information but do not reliably identify the three factors with the strongest measured influence under necessity or sufficiency. The framework provides a black-box reliability check for explanations used in agent oversight while remaining scoped to individual LLM decisions.
Chinese Translation
能够在智能体工作流中运行的大语言模型(LLM)决策组件,通常在产生与行动相关的建议或判断的同时提供解释。操作者可能利用这些被指出的因素来监控系统、诊断错误,或决定何时需要将输出上报。这种使用方式假设解释与组件可观察到的决策行为相一致。我们检验了被指出因素的两种解释:必要性,即改变某个因素会改变输出;充分性,即在移除其他可变信息的同时保留该因素会使输出保持不变。我们在两个合成用例中评估了这些解释:为客户推荐顾问,以及判断提示词的有害性或风险。模型返回一个输出以及对其影响最大的前三个因素。通过受控黑盒干预,我们通过测量改变某因素时输出改变的频率来估计每个因素的必要性得分,并通过测量保留某因素时输出保持不变的频率来估计其充分性得分。在来自 Claude、GPT 和 Gemini 系列的八个模型上,所引用排名与必要性得分和充分性得分之间的平均 Spearman 相关系数在顾问推荐任务中分别为 0.349 和 0.354,在提示词监控任务中分别为 0.431 和 0.580。此外,在必要性检验下,未引用因素得分高于得分最低的引用因素的情况占顾问响应的 57.6%,充分性检验下为 58.1%;在提示词监控任务中,相应比例分别为 25.8% 和 8.9%。被引用的前三个因素包含有用信息,但并不能可靠地识别出在必要性或充分性检验下测量影响力最强的三个因素。该框架为智能体监督中使用的解释提供了一种黑盒可靠性检验,同时其适用范围仍限定于单个 LLM 决策。
cs.AI / 105 / 2609.05395

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

韩国开放公共API上的多步工具调用:一个基准测试与数据合成方法
Kim, Dain, Cho, Eungi, Kim, Kyumin, Noh, Shinyeong, Lim, Kyuseong
Abstract
Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM agents that chain multiple tool-calls across live government APIs. However, open-source models consistently underperform in this multi-step setting, and no existing benchmark measures the gap. We introduce the Korean Open Public API Benchmark (KOPA-Bench), comprising 145 real-world tasks. To close this gap, we present EDGE, an Execution-grounded Dynamic Graph for tool-calling data synthEsis driven by live execution. EDGE builds a graph of how each tool's output can feed another's input, keeps only the links that succeed when actually called against the live APIs, and traverses these verified links to synthesize executable multi-step trajectories. Fine-tuned via GRPO on the resulting dataset, our 9B model nearly matches the untuned 27B model from the same family, improving substantially not only on KOPA-Bench but also on the BFCL benchmark.
Chinese Translation
数据主权法规日益要求公共机构部署开源的本地化(on-premise)大语言模型(LLM)智能体,以在实时政府API上串联多个工具调用。然而,开源模型在这种多步场景下表现持续不佳,且目前尚无基准可以衡量这一差距。我们提出了韩国开放公共API基准(KOPA-Bench),包含145个真实世界任务。为缩小这一差距,我们提出了EDGE(Execution-grounded Dynamic Graph,基于执行的动态图),这是一个由实时执行驱动的工具调用数据合成方法。EDGE构建一张图,刻画每个工具的输出如何作为另一个工具的输入,仅保留在真实调用实时API时验证成功的连接,并遍历这些经过验证的连接来合成可执行的多步轨迹。在所得数据集上通过GRPO进行微调后,我们的9B模型几乎达到了同系列未微调27B模型的水平,不仅在KOPA-Bench上取得显著提升,在BFCL基准上同样表现优异。
cs.AI / 106 / 2609.05396

A Deep Generative Model for Synthesizing Labeled Wireless Signals

一种用于合成带标签无线信号的深度生成模型
Li, Yuxiao, Hu, Keke, Mazuelas, Santiago, Shen, Yuan
Abstract
Wireless signals with position-related labels are pivotal for both performance evaluation and model training in the realm of wireless sensing. However, acquiring real-world datasets is often challenged by significant measurement and labeling costs. Traditional methods for synthesizing labeled wireless signals typically rely on environmental models, leading to extensive hyper-parameter tuning and inadequate realism for comprehensive model training purposes. To address these limitations, we introduce a novel deep learning (DL)-based method, namely Inter-Instance Generative Adversarial Networks (IIns-GAN), to generate realistic labeled wireless signals. The generated signals are particularly adaptive to different environment scenarios and well-suited for various model training tasks, including distance estimation and environment identification. We have conducted extensive experiments on public Ultra-Wideband (UWB) datasets to evaluate the realism and utility of the generated signals. The results demonstrate that the signals generated by IIns-GAN mirror the physical characteristics of real-world measurements, and significantly contribute to the improvement of model training in diverse wireless sensing tasks.
Chinese Translation
带有位置相关标签的无线信号对于无线感知领域的性能评估和模型训练都至关重要。然而,获取真实世界的数据集往往面临高昂的测量和标注成本。传统的合成带标签无线信号的方法通常依赖于环境模型,导致需要进行大量的超参数调整,且真实性不足,难以满足全面的模型训练需求。为了解决这些局限性,我们提出了一种新颖的基于深度学习(DL)的方法,即实例间生成对抗网络(Inter-Instance Generative Adversarial Networks, IIns-GAN),用于生成逼真的带标签无线信号。所生成的信号尤其能够适应不同的环境场景,并非常适用于各种模型训练任务,包括距离估计和环境识别。我们在公开的超宽带(Ultra-Wideband, UWB)数据集上进行了大量实验,以评估所生成信号的真实性和实用性。结果表明,IIns-GAN 生成的信号能够反映真实世界测量的物理特性,并显著提升多种无线感知任务中模型训练的效果。
机器学习 (Machine Learning)
79
cs.LG / 1 / 2609.04264

Spectral-Target Physical Latent Structuring for JEPA-Style World Models

面向JEPA式世界模型的谱目标物理潜在结构化方法
Zhu, Penghao, Penachio, Salvatore, Mukherjee, Kaustav, Jonelagadda, Aneesh
Abstract
Latent world models have become increasingly popular as a method to predict and plan in latent space rather than pixel space. Recent architectures, such as LeWorldModel (LeWM), jointly train the encoder and predictor using regularization techniques like SIGReg to prevent representation collapse. Even with such regularization preventing representation collapse, we identify a new world model failure mode of \textit{physical representation laziness}, particularly noted in highly dynamic environments. For these lazy cases, the learned latent states do not collapse but nonetheless fail to represent key physical properties, causing ubiquitous downstream planning failure. To resolve this issue, we propose training-time auxiliary supervision with a lightweight "Fourier auxiliary head", which enforces physically-informed structuring of the latent space with no additional inference-time cost and can be generalized to any environment. Experimentally, we show that the auxiliary head substantially improves planning success rates in dynamic environments where the baseline LeWM exhibits physical representation laziness. It also leads to modest improvements in other environments, even when the baseline does not exhibit physical representation laziness. We further observe superior planning performance being accompanied by higher latent space correlations with key physical properties, indicating both the ability of our method to physically structure latent states and the potential planning-side benefit to the learned representation being physically structured. We also see in low-data regimes, auxiliary supervision is particularly impactful in increasing success rate. These findings support the use of our Fourier auxiliary head method to improve both overall success rate and data efficiency, while avoiding representation laziness in latent world models.
Chinese Translation
潜在世界模型作为一种在潜在空间而非像素空间中进行预测与规划的方法,正日益受到关注。近期的架构,如LeWorldModel(LeWM),通过SIGReg等正则化技术联合训练编码器与预测器,以防止表征坍缩。然而,即使有这样的正则化防止表征坍缩,我们发现了一种新的世界模型失效模式——物理表征惰性(physical representation laziness),该问题在高动态环境中尤为显著。在这些惰性情形下,学习到的潜在状态虽未坍缩,却无法表征关键物理属性,从而导致普遍的下游规划失败。为解决这一问题,我们提出一种训练时辅助监督方法,即轻量级的“傅里叶辅助头”(Fourier auxiliary head),它以物理信息约束潜在空间的结构化,且不引入额外的推理时开销,并可推广至任意环境。实验表明,在基线LeWM表现出物理表征惰性的动态环境中,辅助头显著提高了规划成功率;在基线未表现出物理表征惰性的其他环境中,该方法也带来了适度的提升。我们进一步观察到,规划性能的提升伴随着潜在空间与关键物理属性相关性的提高,这既表明了本方法能够对潜在状态进行物理结构化,也说明学习到的表征经过物理结构化对规划具有潜在收益。在低数据量场景下,辅助监督对提升成功率尤为有效。这些发现支持使用我们的傅里叶辅助头方法来提升总体成功率与数据效率,同时避免潜在世界模型中的表征惰性问题。
cs.LG / 2 / 2609.04265

ProToMEx: Rapid, Interpretable Explanations via Structured Representations

ProToMEx:基于结构化表示的快速可解释解释方法
Georgara, Athina, Valoor, Adarsh, Ramchurn, Sarvapali D.
Abstract
Existing post-hoc explainers for machine learning classifiers primarily focus on feature attribution, assigning importance scores to individual features. While valuable, this approach struggles to articulate the complex, combinatorial patterns that often drive a model's decision-making process. To overcome this limitation, we introduce ProToMEx, a new paradigm for explainability that leverages Probabilistic Topic Models (PTMs). Our model-agnostic framework learns latent ''topics'' that represent distinct, high-level reasons for a classification, moving beyond simple feature importance to reveal underlying semantic structures. ProToMEx naturally provides both global explanations of a model's overall behaviour and local explanations that can disentangle multiple co-existing reasons for a specific prediction. We demonstrate empirically that ProToMEx not only produces explanations of comparable fidelity to popular methods like SHAP and LIME but also drastically reduces the amortised computational cost of generating local explanations, making it highly suitable for real-time applications. Specifically, we show that ProToMEx is ~30-40x faster than SHAP and LIME over standardised tabular datasets and synthetic datasets.
Chinese Translation
现有的机器学习分类器事后解释方法主要集中于特征归因,即为单个特征分配重要性分数。这种方法虽然有价值,但难以清晰地阐述驱动模型决策过程的复杂组合模式。为克服这一局限性,我们提出了ProToMEx,这是一种利用概率主题模型(PTM)的可解释性新范式。我们的模型无关框架学习潜在的"主题",这些主题代表分类的不同高层次原因,超越了简单的特征重要性,从而揭示底层的语义结构。ProToMEx自然地提供模型整体行为的全局解释,以及能够厘清特定预测中多个共存原因的局部解释。我们通过实证研究表明,ProToMEx不仅能产生与SHAP和LIME等流行方法相当保真度的解释,还大幅降低了生成局部解释的摊销计算成本,使其非常适合实时应用。具体而言,我们证明在标准化表格数据集和合成数据集上,ProToMEx比SHAP和LIME快约30-40倍。
cs.LG / 3 / 2609.04267

A Data Fusion Framework for Grounding Aerospace Surrogate Model via Experimental Wind-Tunnel Observations

基于实验风洞观测数据为航空航天代理模型提供校准的数据融合框架
Kulkarni, Nitin Nagesh, Vemula, Dheeraj, Yu, Yin, Lyu, Peter, Alonso, Juan J.
Abstract
Aerodynamic surrogate models trained on high-fidelity CFD data reproduce numerical predictions of both scalar outputs and entire fields accurately, yet their predictive fidelity is limited by systematic discrepancies between CFD and experimental observations. We present an experimentally grounded correction framework that adapts a CFD-trained deep learning surrogate using wind-tunnel PSP measurements. A Geotransolver surrogate trained on 2,300 high-fidelity CFD simulations of the NASA CRM wing-body configuration, spanning geometric variation, Mach 0.70-0.85, and angles of attack 0 to 4 degrees, reproduces the CFD integrated aerodynamic forces and pitching moment to R2 > 0.99 but does not match the experimental data. To incorporate experimental information without retraining the surrogate, a correction network is trained on spatially registered PSP measurements at two freestream Mach numbers (0.70 and 0.85) across the same angle-of-attack range, learning the discrepancy between the surrogate-predicted and experimentally measured surface-pressure distributions. At Mach 0.85 the correction substantially improves agreement with PSP, particularly at the wing suction peak, shock location, and subsequent pressure recovery, reducing both the magnitude of the prediction error and the fraction of wetted surface on which it exceeds 0.05 in Cp, and it does so from a limited experimental dataset without modifying the pretrained surrogate parameters. On held-out angles of attack the grounded surrogate agrees with measurement to within 2.3-2.7% of the measured Cp range, and outperforms direct interpolation between the measured conditions at every state tested. Experimental measurements can therefore ground a large-scale simulation-trained surrogate by learning systematic CFD-to-experiment discrepancies while preserving its generalization capability and computational efficiency.
Chinese Translation
基于高保真CFD数据训练的气动力代理模型能够准确复现标量输出和全场数值预测,但其预测精度受限于CFD与实验观测之间的系统性偏差。我们提出一种以实验为基准的修正框架,利用风洞PSP(压敏漆)测量数据对基于CFD训练的深度学习代理模型进行自适应修正。该代理模型采用Geotransolver,基于NASA CRM翼身组合体构型的2300个高保真CFD仿真训练,涵盖几何变化、马赫数0.70–0.85以及0至4度迎角,对CFD积分气动力和俯仰力矩的复现精度达到R² > 0.99,但与实验数据并不吻合。为在不重新训练代理模型的情况下融入实验信息,我们在相同迎角范围内、两个来流马赫数(0.70和0.85)下空间配准的PSP测量数据上训练一个修正网络,学习代理模型预测与实验测量的表面压力分布之间的偏差。在马赫数0.85下,该修正显著改善了与PSP数据的一致性,尤其是在机翼吸力峰值、激波位置及其后的压力恢复区域,既降低了预测误差的幅值,又减小了误差超过0.05(以Cp计)的润湿表面积占比;而且这一切仅基于有限的实验数据集完成,未修改预训练代理模型的参数。在留出的迎角上,校准后的代理模型与测量结果的偏差在实测Cp范围的2.3–2.7%以内,且在所有测试状态下均优于对实测工况的直接插值。因此,实验测量可以通过学习CFD到实验的系统性偏差来校准大规模仿真训练的代理模型,同时保持其泛化能力和计算效率。
cs.LG / 4 / 2609.04271

Quantum-Assisted Memory-Efficient Training for Parameter-Intensive Wi-Fi-Based Human Activity Recognition

面向参数密集型基于Wi-Fi的人体活动识别的量子辅助内存高效训练
An, To Truong, Zhang, Jie, Yin, Guolin, Zhang, Junqing, Li, Yanjiao, Duong, Trung Q., Cotton, Simon L.
Abstract
Wi-Fi-based human activity recognition (HAR) has become an important part of integrated sensing and communications, paving the way for a range of context-aware services. However, most existing Wi-Fi-based HAR systems rely on deep learning (DL) models that are computationally and memory intensive in both training and inference, which poses significant challenges for real-world deployment. Conventional training requires simultaneous updates of millions of parameters, leading to prohibitive memory consumption. In this paper, we propose a novel quantum-assisted memory-efficient training framework (Q-MET) designed to improve efficiency in both training and inference. Q-MET utilizes a hybrid quantum classical neural network to indirectly generate parameters for HAR models, significantly reducing the trainable parameter count compared to direct optimization. To further support the deployment on resource-constrained devices, we integrate structured pruning during the training phase. Experimental results demonstrate that Q-MET achieves a 90% to 95% reduction in trainable parameters compared with conventional backpropagation-based DL training while maintaining or even exceeding classical classification accuracy. Additionally, Q-MET supports lightweight inference through structured pruning, achieving 75% to 85% model sparsity with less than 2% loss in classification accuracy. To the best of our knowledge, this work represents the first quantum-assisted approach to simultaneously tackle memory inefficiencies in both the training and inference stages of HAR systems.
Chinese Translation
基于Wi-Fi的人体活动识别(HAR)已成为通感一体化的重要组成部分,为一系列情境感知服务铺平了道路。然而,现有的大多数基于Wi-Fi的HAR系统依赖于深度学习(DL)模型,这些模型在训练和推理阶段均需要大量的计算和内存资源,给实际部署带来了巨大挑战。传统训练需要同时更新数百万个参数,导致内存消耗过高。本文提出了一种新颖的量子辅助内存高效训练框架(Q-MET),旨在提高训练和推理的效率。Q-MET利用混合量子-经典神经网络间接生成HAR模型的参数,与直接优化相比,显著减少了可训练参数的数量。为进一步支持在资源受限设备上的部署,我们在训练阶段引入了结构化剪枝。实验结果表明,与传统的基于反向传播的深度学习训练相比,Q-MET在保持甚至超越经典分类精度的同时,将可训练参数减少了90%至95%。此外,Q-MET通过结构化剪枝支持轻量化推理,在分类精度损失小于2%的情况下实现了75%至85%的模型稀疏度。据我们所知,这项工作是首个同时解决HAR系统训练和推理两个阶段内存低效问题的量子辅助方法。
cs.LG / 5 / 2609.04272

Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning

评估大语言模型用于强迫停电风险预测:优势及与机器学习的比较
Petridis, Christos, Obradovic, Zoran, Kezunovic, Mladen
Abstract
This study examines the ability of large language models (LLMs) to predict the risk of weather-related forced outages in the distribution grid in a zero-shot framework, without labeled training data. The problem is formulated as a binary severity classification task across three forecast horizons (3h, 6h, 12h), using six years of outage records and high-resolution weather data for a utility service area in central Texas. Four zero-shot LLMs are benchmarked against two supervised classifiers across two input configurations: one using current weather observations and the other using weather forecast data. Results show that supervised models outperform LLMs on macro-F1 and precision, while newer LLM generations achieve competitive scores. Beyond accuracy, LLMs offer complementary strengths in actionable reasoning and geographic scalability, suggesting that combining them with supervised models may be the best practice.
Chinese Translation
本研究考察了大语言模型(LLM)在零样本框架下、无需标注训练数据的情况下,预测配电网天气相关强迫停电风险的能力。该问题被构建为一个二元严重程度分类任务,涵盖三个预报时间尺度(3小时、6小时、12小时),使用了德克萨斯州中部某电力公司服务区域六年的停电记录和高分辨率气象数据。研究将四个零样本大语言模型与两个监督分类器在两种输入配置下进行基准比较:一种使用当前天气观测数据,另一种使用天气预报数据。结果表明,监督模型在宏平均F1(macro-F1)和精确率上优于大语言模型,而新一代大语言模型取得了具有竞争力的分数。除准确性之外,大语言模型在可操作性推理和地理可扩展性方面展现出互补优势,这表明将大语言模型与监督模型相结合可能是最佳实践。
cs.LG / 6 / 2609.04292

BER-PEF: Unified Human Mobility Predictability Evaluation via Bayes Error Rate Estimation

BER-PEF:基于贝叶斯错误率估计的人类移动性可预测性统一评估方法
Xu, En, Ding, Jingtao, Yu, Zhiwen, Li, Yong
Abstract
Human mobility predictability concerns the best prediction performance attainable from a given target and input information, but its ground truth is not directly observable on real mobility data. We present BER-PEF, a Bayes-error-rate-based framework that converts BER estimation into mobility predictability estimation and provides a unified protocol for comparing estimators without observable ground truth. The framework maps symbolic sequences, numeric trajectories, contextual features, and learned representations into a common feature--label space, then evaluates estimator outputs along controlled perturbation curves against a shared predictability reference interval by measuring deviations below the interval, above the interval, and across the full interval. Experiments on Foursquare NYC and TKY, GeoLife, and T-Drive show that several BER-based estimators achieve lower reference discrepancy than existing predictability methods on symbolic sequences and numeric trajectories, while their estimates track changes in empirical prediction performance under perturbation. Additional analyses show that contextual inputs and multiple structured representations can be evaluated under the same protocol, and that aggregating evidence across multiple perturbation levels provides a more reliable basis for estimator selection than relying on a single unperturbed observation. BER-PEF therefore offers a unified and verifiable path for evaluating predictability estimators on heterogeneous mobility data when ground-truth predictability is unavailable.
Chinese Translation
人类移动性可预测性关注在给定目标信息和输入信息条件下可达到的最佳预测性能,但其在真实移动数据上的真值(ground truth)无法直接观测。我们提出了BER-PEF,一个基于贝叶斯错误率(Bayes Error Rate, BER)的框架,该框架将BER估计转化为移动性可预测性估计,并提供了一个无需可观测真值即可比较不同估计器的统一协议。该框架将符号序列、数值轨迹、上下文特征以及学习到的表征映射到统一的特征-标签空间中,然后通过测量估计器输出相对于共享可预测性参考区间的低于区间、高于区间以及跨整个区间的偏差,沿受控扰动曲线对估计器输出进行评估。在Foursquare NYC和TKY、GeoLife以及T-Drive数据集上的实验表明,若干基于BER的估计器在符号序列和数值轨迹上的参考偏差低于现有的可预测性方法,且其估计值能够跟踪扰动下经验预测性能的变化。进一步的分析表明,上下文输入和多种结构化表征可以在同一协议下进行评估,并且在多个扰动水平上聚合证据为估计器选择提供了比依赖单一未扰动观测更可靠的基础。因此,当可预测性真值不可用时,BER-PEF为在异构移动数据上评估可预测性估计器提供了一条统一且可验证的路径。
cs.LG / 7 / 2609.04329

Data-Driven Learning of Unknown Nonlinear Differential Equations Using Functional Analysis

基于泛函分析的数据驱动未知非线性微分方程学习方法
Alaviani, Seyyed Shaho, Qu, Yongzhi, Vogl, Gregory W.
Abstract
In this paper, the problem of data-driven discovery of nonlinear ordinary differential equations (ODEs) is recast, and a new interpretable machine learning (ML) method is proposed. The proposed method aims to learn the unknown vector field of nonlinear dynamics without prior knowledge of the system's physics from only one single state trajectory's data. The proposed method has two fundamental differences with existing methods: 1) the formulation presented in this method is derived based on Functional Analysis and Operator Theory, and 2) the cost function is constructed in the function space as a distance between two functions as an integral, instead of the discrete-sum of errors used in existing ML approaches. An incremental learning algorithm is proposed to learn the unknown vector field to handle new training samples in an online manner. The proposed method can discover the unknown vector field from both forced and unforced autonomous and non-autonomous (or time-varying) dynamical systems. The proposed method is able to simultaneously discover unknown external forces as a function of time and unknown underlying dynamics. Finally, numerical examples are given to demonstrate the advantages of the proposed method.
Chinese Translation
本文重新阐述了数据驱动发现非线性常微分方程(ODEs)的问题,并提出了一种新的可解释机器学习(ML)方法。该方法旨在仅利用单条状态轨迹的数据,在缺乏系统物理先验知识的情况下,学习非线性动力学的未知向量场。所提出的方法与现有方法有两个根本区别:1)本文方法的建模公式基于泛函分析(Functional Analysis)与算子理论(Operator Theory)推导而来;2)代价函数在函数空间中构造为两个函数之间以积分形式表示的距离,而非现有机器学习方法中所使用的离散误差求和。本文提出了一种增量学习算法,以在线方式处理新的训练样本,从而学习未知向量场。该方法能够从受迫与非受迫的自治及非自治(或时变)动力系统中发现未知向量场。此外,该方法能够同时发现作为时间函数的未知外部作用力以及未知的潜在动力学。最后,通过数值算例验证了所提方法的优越性。
cs.LG / 8 / 2609.04339

Modular Deep Recurrent Neural Network: Application to Quadrotors

模块化深度循环神经网络:在四旋翼飞行器上的应用
Mohajerin, Nima, Waslander, Steven L.
Abstract
A modular deep Recurrent Neural Network (RNN) is introduced to facilitate the process of deploying various architectures of RNNs, and to automatically compute derivatives for gradient-based learning methods. The modularity leads to a set of new architectures, one of which includes feedforward inter-layer connections. By adding feedforward inter-layer connections in a multi-layer RNN, it is observed that the capability of the RNN to learn and model high-order dynamics and nonlinearities is significantly improved. The problem of vanishing/exploding gradient in space for a multilayer RNN is also alleviated using feedforward connections. These results are demonstrated using a quadrotor case study, for which a model of the altitude dynamics is learned with our particular network structure, while existing methods are unable to generalize as quickly or at all.
Chinese Translation
本文提出了一种模块化深度循环神经网络(RNN),以便于部署各种RNN架构,并自动计算基于梯度学习方法的导数。模块化设计带来了一系列新的网络架构,其中一种架构包含了跨层前馈连接。通过在多层RNN中加入跨层前馈连接,可以观察到RNN学习和建模高阶动态与非线性特性的能力得到显著提升。同时,前馈连接还缓解了多层RNN中梯度在空间上消失/爆炸的问题。这些结果通过四旋翼飞行器案例研究得到验证:利用我们提出的特定网络结构学习其高度动态模型,而现有方法要么泛化速度较慢,要么完全无法泛化。
cs.LG / 9 / 2609.04344

SharedSAE: One Feature Dictionary Across Language Models

SharedSAE:跨语言模型的统一特征字典
Ognev, Daniil, Vasson, Célian, Hu, Lijie, Inui, Kentaro, Heinzerling, Benjamin
Abstract
Sparse autoencoders (SAEs) are widely used to interpret language model activations, but SAE training and latent labelling are typically repeated for every model. Here, we show that a single shared SAE can replace a collection of dedicated per-model SAEs. Our method, SharedSAE, combines a shared dictionary with model-specific encoder-decoder pairs. Unlike the closest prior method, which discards activation magnitudes and requires all models at inference, SharedSAE instead normalizes only selection scores, preserving magnitudes, and uses model dropout for single-model inference. We train SharedSAE on four 1B-scale base language models spanning distinct families and tokenizers. Despite sharing its latents across models, SharedSAE retains 96.6% of dedicated SAEs' mean explained variance; its latent activations exhibit cross-model correlations 1.8 times as high as separate SAEs aligned post-hoc, and its latent descriptions transfer across models. After the dictionary is frozen, new models can be efficiently adapted to it, achieving near-dedicated-SAE reconstruction quality while reusing the shared latent descriptions.
Chinese Translation
稀疏自编码器(SAE)被广泛用于解释语言模型的激活,但SAE的训练和潜在特征标注通常需要对每个模型重复进行。本文证明,单个共享的SAE可以替代针对各模型单独训练的SAE集合。我们的方法SharedSAE将一个共享字典与各模型专属的编码器-解码器对相结合。与最接近的先前方法不同——该方法丢弃激活幅值且在推理时需要所有模型——SharedSAE仅对选择分数进行归一化以保留幅值,并通过模型dropout实现单模型推理。我们在四个1B规模、涵盖不同模型家族和分词器的基础语言模型上训练了SharedSAE。尽管其潜在特征在各模型间共享,SharedSAE仍保留了专属SAE平均解释方差的96.6%;其潜在激活呈现的跨模型相关性是对齐后的独立SAE的1.8倍,且其潜在特征描述可在模型间迁移。字典冻结后,新模型可以高效地适配该字典,在近乎达到专属SAE重建质量的同时复用共享的潜在特征描述。
cs.LG / 10 / 2609.04354

A Quantum Variational Approach to Prototypical Recurrent Unit

一种量子变分方法在原型循环单元中的应用
Garjan, Mahyar Sadeghi, Cesari, Tommaso, Barbeau, Michel
Abstract
We introduce a lightweight Quantum Prototypical Recurrent Unit (QPRU) that requires significantly fewer parameters than both classical recurrent architectures, such as Long Short- Term Memory (LSTM) and Gated Recurrent Unit (GRU), and quantum variants, including Quantum LSTM (QLSTM) and Quantum GRU (QGRU). Despite its compact design, the QPRU achieves competitive forecasting performance, matching state-of-the-art baselines while offering important structural and practical advantages, including enhanced scalability and a reduced number of trainable parameters.
Chinese Translation
我们提出了一种轻量级的量子原型循环单元(Quantum Prototypical Recurrent Unit, QPRU),与经典的循环架构(如长短期记忆网络 LSTM 和门控循环单元 GRU)以及量子变体(包括量子 LSTM(QLSTM)和量子 GRU(QGRU))相比,其所需参数显著更少。尽管设计紧凑,QPRU 仍实现了具有竞争力的预测性能,达到最先进基线模型的水平,同时具备重要的结构和实用优势,包括更强的可扩展性和更少的可训练参数数量。
cs.LG / 11 / 2609.04379

On the Abundance of Critical Points of the t-SNE Energy

论 t-SNE 能量函数临界点的丰度
Haridas, Nakul, Murray, Ryan
Abstract
This paper considers the energy landscape of the t-SNE algorithm. While this algorithm has enjoyed broad adoption, the non-convexity of the associated energy has made it difficult to rigorously understand what the algorithm captures in many settings. In particular, a number of well-known numerical examples, several of which are reproduced in this article, suggest a complicated energy landscape with many local minimizers that do not respect the topology or clustering structure of the underlying data. This work seeks to provide first steps towards a rigorous explanation of these phenomena. Specifically, for a general family of energies, which include both the original t-SNE algorithm and recently identified large data limits, and for densities in feature space which obey a continuous symmetry, we construct infinite families of distinct critical points. These critical points are based upon identifying pairs of discrete symmetries, one in the original feature space and the other in the target embedding space, which are preserved under gradient dynamics. These critical configurations exhibit many characteristics, such as topology breaking and spurious clustering, which are often observed empirically. Finally, numerical and analytical examples are given throughout as a means of illustrating the approach.
Chinese Translation
本文研究了 t-SNE 算法的能量景观。尽管该算法已被广泛采用,但其相关能量的非凸性使得人们难以在许多设定下严格理解该算法究竟捕捉到了什么。特别地,许多著名的数值例子(其中若干个在本文中被复现)表明,其能量景观十分复杂,存在大量不遵循底层数据拓扑结构或聚类结构的局部极小值点。本工作旨在为严格解释这些现象迈出第一步。具体而言,对于一个一般的能量函数族(既包括原始的 t-SNE 算法,也包括最近确定的大数据极限),以及特征空间中服从连续对称性的密度分布,我们构造了无穷多个互不相同的临界点族。这些临界点基于对一对离散对称性的识别,其一位于原始特征空间,另一个位于目标嵌入空间,且二者在梯度动力学下保持不变。这些临界构型表现出许多特性,如拓扑破坏和虚假聚类,这些特性在实证中经常被观察到。最后,全文贯穿给出数值与分析实例,以阐释该方法。
cs.LG / 12 / 2609.04407

Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures

解耦深度算子学习中的注意力机制:数据驱动与物理信息架构的对照研究
Koric, Amar Alem, Liu, Qibang, Koric, Seid
Abstract
Deep neural operators learn mappings between input functions and complete PDE solution fields, enabling forward evaluations of new problem instances orders of magnitude faster than conventional numerical solvers. Attention mechanisms have recently been introduced into neural operators, but most studies change several architectural components at once, making it difficult to identify what actually improves accuracy. This work presents a controlled and systematic study of five deep operator network (DeepONet) variants with distinct attention mechanisms, trained under both data-driven and physics-informed regimes, to isolate the effects of cross-attention, self-attention, tokenization, and attention depth. We evaluate them on a source-driven transient one-dimensional nonlinear diffusion-reaction equation, a transient one-dimensional viscous Burgers equation with variable initial conditions, and a two-dimensional Poisson heat-conduction problem with heterogeneous source fields. Per-sensor tokenization with cross-attention reduces the mean relative L_2 error of the classical DeepONet in all benchmark-training combinations by factors of 2.4-28.0, while the best attention configurations reach 3.5-32.3. Branch self-attention paired only with dot-product fusion is inconsistent, degrading the one-dimensional problems while helping the more complex two-dimensional source field; added on top of cross-attention it improves all six cases, though by less than cross-attention fusion alone. Global pre-mixing provides no consistent benefit. Increasing cross-attention depth further improves accuracy, but with diminishing returns and a substantially higher cost under physics-informed training. Overall, query-dependent cross-attention is the most reliable mechanism, whereas branch self-attention is most useful for large, spatially complex functional inputs.
Chinese Translation
深度神经算子学习输入函数与完整偏微分方程(PDE)解场之间的映射,使得对新问题实例的前向评估比传统数值求解器快数个数量级。注意力机制最近被引入神经算子,但大多数研究同时改变了多个架构组件,导致难以识别真正提升精度的因素。本工作对五种具有不同注意力机制的深度算子网络(DeepONet)变体进行了受控的系统性研究,这些变体分别在数据驱动和物理信息两种训练模式下训练,以分离交叉注意力、自注意力、分词化(tokenization)以及注意力深度的影响。我们在三个基准问题上对其进行评估:源驱动的瞬态一维非线性扩散-反应方程、具有可变初始条件的瞬态一维黏性Burgers方程,以及具有非均匀源场的二维Poisson热传导问题。结合交叉注意力的逐传感器分词化在所有基准-训练组合中将经典DeepONet的平均相对L_2误差降低了2.4至28.0倍,而最佳的注意力配置可达到3.5至32.3倍。仅在点积融合(dot-product fusion)下配合分支自注意力时,结果并不一致:它会降低一维问题的精度,却有助于更复杂的二维源场问题;若将其叠加在交叉注意力之上,则能改善全部六个案例,但提升幅度小于单独使用交叉注意力融合。全局预混合(Global pre-mixing)未能带来一致的收益。增加交叉注意力深度可进一步提升精度,但收益递减,且在物理信息训练下代价显著增加。总体而言,依赖于查询的交叉注意力是最可靠的机制,而分支自注意力对空间复杂的大型函数输入最为有用。
cs.LG / 13 / 2609.04415

REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation

REFINE:面向个性化医学概念表示的预算约束下文本属性图LLM精炼框架
Kerdabadi, Mohsen Nayebi, Moghaddam, Arya Hadizadeh, Wang, Dongjie, Yao, Zijun
Abstract
Learning rich medical concept representations is essential for EHR prediction. Text-attributed knowledge graphs (TKGs) provide a natural foundation by organizing heterogeneous medical relations together with textual semantics. However, most existing encoders process concepts uniformly across patients, despite the fact that a code's meaning and predictive value depend on patient-specific clinical context and trajectory. Learning patient-personalized concept representations from TKGs introduces two key challenges: (1) deciding how much KG context to incorporate for each observed code, and (2) aligning semantic information with the patient-specific relational structure. We propose REFINE, a KG-aware budgeted LLM graph refinement framework for patient-personalized medical concept encoding. Starting from a global TKG, REFINE constructs patient-specific temporal graphs. A sequential reinforcement learning policy selects a personalized KG expansion budget for each observed code. The resulting patient graph is processed by a heterogeneous GNN to capture relation-aware structural dependencies, while a frozen LLM uses graph-aware soft prompts to semantically refine concept representations. Experiments on MIMIC-III and MIMIC-IV show that REFINE consistently improves diverse EHR backbones, outperforms strong baselines, and demonstrates robust gains across component ablation, KG selection, and data insufficiency.
Chinese Translation
学习丰富的医学概念表示对电子健康记录(EHR)预测至关重要。文本属性知识图谱(TKG)通过将异构医学关系与文本语义组织在一起,为这一任务提供了天然的基础。然而,现有的大多数编码器对所有患者采用统一的方式处理概念,尽管事实上某个编码的含义和预测价值取决于患者特定的临床情境与诊疗轨迹。从TKG中学习患者个性化的概念表示面临两个关键挑战:(1)如何决定为每个观测到的编码引入多少知识图谱上下文;(2)如何将语义信息与患者特定的关系结构对齐。我们提出REFINE,一个面向患者个性化医学概念编码的、感知知识图谱的预算约束LLM图精炼框架。REFINE从全局TKG出发,构建患者特定的时间图。一个序列强化学习策略为每个观测到的编码选择个性化的KG扩展预算。由此得到的患者图由异构图神经网络(GNN)处理,以捕捉关系感知的结构依赖,同时一个冻结的LLM利用图感知的软提示对概念表示进行语义精炼。在MIMIC-III和MIMIC-IV上的实验表明,REFINE能够持续提升多种EHR骨干模型的性能,优于强基线方法,并在组件消融、KG选择和数据不足等情况下均展现出稳健的增益。
cs.LG / 14 / 2609.04425

Beyond a Universal Forecasting Selector: Demand-Conditioned Model Selection across Demand Patterns and Horizons

超越通用预测选择器:跨需求模式与预测期的条件化需求模型选择
González, Adolfo
Abstract
Forecasting-model selection remains difficult in heterogeneous demand because the most suitable decision rule may vary with demand structure, data availability, and forecasting horizon. This study examines whether the selector itself should be treated as a context-dependent component of the forecasting process. Five selection mechanisms - RMSSE, ERA, OWA, CCG-AHSC, and CCG-AHSCD - are compared across 24 optimized forecasting models, nine datasets, three training-testing partitions, and horizons from 1 to 12 cycles. Selector performance is evaluated ex post using Global Relative Accuracy (GRA), statistical tests, and a best-attainable-model reference. No selector dominates across all conditions. CCG-AHSC and CCG-AHSCD are more competitive for Smooth demand and several Erratic configurations, whereas OWA and ERA perform better in Intermittent and Lumpy settings. Selector suitability also changes with historical data availability and horizon, supporting a context-dependent rather than universal approach to forecasting-model selection.
Chinese Translation
在异质性需求环境下,预测模型的选择仍然十分困难,因为最合适的决策规则可能随需求结构、数据可得性和预测期长度的不同而变化。本研究探讨是否应将选择器本身视为预测过程中一个依赖于具体情境的组成部分。研究在24个优化的预测模型、九个数据集、三种训练-测试划分以及1至12个周期的预测期上,比较了五种选择机制——RMSSE、ERA、OWA、CCG-AHSC和CCG-AHSCD。选择器的性能通过全局相对精度(Global Relative Accuracy, GRA)、统计检验以及最优可达模型参考进行事后评估。结果表明,没有任何选择器能够在所有条件下都占优。CCG-AHSC和CCG-AHSCD在平滑型(Smooth)需求及若干间歇波动型(Erratic)配置中更具竞争力,而OWA和ERA在间歇型(Intermittent)和块状型(Lumpy)需求情境下表现更佳。选择器的适用性还随历史数据可得性和预测期长度的变化而变化,这支持了采用情境依赖而非通用方法的预测模型选择思路。
cs.LG / 15 / 2609.04428

A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models

基于五种大语言模型的英文歌曲歌词文化分析重复测量研究
Smith, E. Cho, Ho, Samuel, Laux, Dawn
Abstract
Large language models (LLMs) are increasingly used to annotate cultural texts at scales that are impractical for human coders. However, before their outputs are treated as measurements of latent social constructs, it is necessary to establish whether those measurements are reliable. This study evaluates five LLMs as zero-shot annotators of four social constructs expressed in English song lyrics: self-esteem, self-control, seeking belonging, and seeking recognition. Using repeated annotations of a large lyric corpus, we examine three properties of LLM-based measurement: consistency across repeated runs, convergence across models, and transferability of consensus labels to supervised classification. The findings show that LLM-based measurement is not uniformly reliable across constructs. Self-esteem exhibits the strongest repeated-measurement reliability across models, while seeking recognition is generally less stable; self-control and seeking belonging show intermediate but model-dependent reliability. Downstream classification further indicates that consensus LLM labels contain learnable signal, although transferability does not itself establish construct validity. Repeated-measurement stability and cross-model convergence should therefore be reported before LLM annotations are treated as scalable measurements in cultural analytics.
Chinese Translation
大语言模型(LLMs)被越来越多地用于以人工编码者难以企及的规模来标注文化文本。然而,在将大语言模型的输出视为潜在社会建构的测量结果之前,必须先确认这些测量是否可靠。本研究评估了五种大语言模型作为零样本(zero-shot)标注者的表现,所涉及的社会建构包括英文歌曲歌词中表达的四种:自尊、自我控制、寻求归属感和寻求认可。通过对大规模歌词语料库进行重复标注,我们考察了基于大语言模型的测量的三个特性:重复运行之间的一致性、不同模型之间的收敛性,以及共识标签向监督分类任务的可迁移性。研究结果表明,基于大语言模型的测量在不同建构上的可靠性并不一致。自尊在各个模型中展现出最强的重复测量信度,而寻求认可通常稳定性较差;自我控制和寻求归属感则表现出中等程度且依赖于具体模型的可靠性。下游分类结果进一步表明,大语言模型的共识标签包含可学习的信息,但可迁移性本身并不能确立建构效度。因此,在大语言模型标注被作为文化分析中可扩展的测量手段使用之前,应先报告其重复测量稳定性和跨模型收敛性。
cs.LG / 16 / 2609.04445

Conformity Breaks Conformal Prediction

从众性破坏共形预测
Hu, Yibo, Su, Hanyu
Abstract
A conformal certificate can be valid when an LLM answers alone and invalid when the same LLM sees peers that unanimously assert a wrong answer. The question is unchanged; the model's score for the correct answer changes. We call this a score-mechanism shift: clean calibration certifies how the model scores answers alone, but not how it scores them under peer pressure. We show that this shift silently breaks conformal prediction in multi-agent LLM systems. Across open-weight models and multiple-choice QA tasks, coverage falls from a calibrated 90% to 74% under unanimous-wrong peers at the standard alpha = 0.10 operating point. The average hides a sharper failure: by targeting the low-confidence items the certificate still covers, an attacker nearly halves coverage on that subgroup, from 87% to 47%, while the monitored average remains much higher. The failure also reaches the decision layer: a system that should escalate when uncertain can instead become confident enough to act on the attacker's wrong answer. Standard conformal fixes do not solve the problem, because the question distribution has not changed; the model's scoring behavior has.
Chinese Translation
当大语言模型(LLM)独立作答时,共形证书可能是有效的;但当同一模型看到一致断言错误答案的同伴时,该证书便可能失效。问题本身没有改变,改变的是模型对正确答案的评分。我们将这种现象称为评分机制偏移(score-mechanism shift):干净环境下的校准只能认证模型独立作答时的评分方式,却无法认证其在同伴压力下的评分方式。我们证明这种偏移会在多智能体LLM系统中悄然破坏共形预测(conformal prediction)。在多个开源权重模型和多项选择题问答任务上,在标准显著性水平 alpha = 0.10 的操作点上,面对一致给出错误答案的同伴时,覆盖率会从校准后的90%降至74%。平均覆盖率掩盖了更为尖锐的失效:攻击者通过专门针对证书仍然覆盖的低置信度样本,可使该子群体的覆盖率几乎减半,从87%降至47%,而受监控的平均覆盖率仍保持在较高水平。这种失效还会波及决策层:一个本应在不确定时上报的系统,反而可能变得足够自信,从而对攻击者的错误答案采取行动。标准的共形修正方法无法解决这一问题,因为问题的分布并未改变,改变的是模型的评分行为。
cs.LG / 17 / 2609.04453

When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models

当负载均衡走向极端:过度分散的混合专家模型中的专家剪枝
Kapusuzoglu, Berkcan, Pryor, Connor, Cho, Sangwoo, Chakraborty, Supriyo, Zhang, Shi-Xiong, Sahu, Sambit, Naphade, Milind
Abstract
Expert pruning reduces the memory and serving cost of Mixture-of-Experts (MoE) models by removing low-importance experts identified by the router, assuming router probabilities provide a reliable importance signal. We observe that this assumption breaks down under over-dispersed routing, a regime associated with aggressive load-balancing during training, in which tokens are distributed nearly uniformly across experts and importance signals collapse. In this regime, perplexity does not predict downstream task accuracy: on gpt-oss-20B, the lowest-perplexity pruning configuration yields the worst mathematical reasoning, while the highest-perplexity configuration preserves it. This does not occur under standard routing (e.g., Mixtral-8x7B-Instruct), where perplexity and accuracy degrade together. Pruning under over-dispersed routing also exposes a capability trade-off in which no single scoring metric dominates: activation-aware scoring preserves mathematical reasoning but severely degrades knowledge-intensive science (an 18-point gap on GPQA), whereas frequency-based scoring exhibits the reverse. We propose Minimax Expert Score Allocation (MESA), a domain-aware method that iteratively boosts importance scores for experts serving whichever domain is currently worst-affected, minimizing worst-case domain degradation rather than average accuracy. At 25% expert pruning MESA achieves the smallest worst-case degradation across domains, outperforming activation-aware baselines on 7 of 11 benchmarks at a correspondingly reduced memory footprint, and it generalizes to gpt-oss-120B, Gemma-4-26B-A4B, and OLMoE-1B-7B. Our results indicate that over-dispersed routing is a qualitatively distinct pruning regime in which standard assumptions fail, and that recognizing it is a prerequisite for principled expert pruning of load-balanced MoE models.
Chinese Translation
专家剪枝通过移除由路由器识别的低重要性专家来降低混合专家(Mixture-of-Experts, MoE)模型的内存与服务成本,其前提假设是路由器概率能够提供可靠的重要性信号。我们发现,在过度分散路由(over-dispersed routing)情况下,这一假设不再成立。过度分散路由与训练中激进负载均衡相关,其特点是 token 在各专家间近乎均匀分布,重要性信号随之崩溃。在这种情形下,困惑度无法预测下游任务准确率:在 gpt-oss-20B 上,困惑度最低的剪枝配置反而导致最差的数学推理表现,而困惑度最高的配置却能保留数学推理能力。这一现象在标准路由下(如 Mixtral-8x7B-Instruct)并不出现,此时困惑度与准确率同步下降。在过度分散路由下进行剪枝还暴露出一种能力权衡:没有任何单一评分指标占据全面优势——基于激活感知的评分能保留数学推理能力,但会严重损害知识密集型的科学任务(在 GPQA 上有 18 分的差距);而基于频率的评分则表现出相反的特性。我们提出了 Minimax 专家分数分配方法(Minimax Expert Score Allocation, MESA),这是一种领域感知的方法,通过迭代提升当前受影响最严重领域所用专家的重要性分数,以最小化最坏情况下的领域性能退化,而非平均准确率损失。在 25% 专家剪枝率下,MESA 实现了跨领域最小的最坏情况退化,在 11 个基准中的 7 个上优于激活感知基线方法,同时相应降低了内存占用;并且该方法可推广至 gpt-oss-120B、Gemma-4-26B-A4B 和 OLMoE-1B-7B。我们的结果表明,过度分散路由是一种在性质上截然不同的剪枝情形,标准假设在此失效;识别这种情形是对负载均衡 MoE 模型进行有原则的专家剪枝的先决条件。
cs.LG / 18 / 2609.04458

On-board ML for Trace Gas detection in Imaging Spectroscopy data

基于星载机器学习的成像光谱数据痕量气体检测研究
Růžička, Vít, Chlus, Adam, Thorpe, Andrew, Thompson, David R.
Abstract
Data collected during aerial and spaceborne imaging spectroscopy campaigns enables the detection of transient events such as trace gas emissions. However, current processing pipelines depend on slow, on-the-ground processing, which delays the time to information of each detected event and prohibits immediate follow-up actions. During the Tokyo Field Campaign of March 2026, we explored on-board processing of Imaging Spectroscopy data from the equipped AVIRIS-5 sensor. Due to communication bottlenecks, full datacubes cannot be downlinked immediately during the flight. Instead we downlink the potential events predicted by our efficient and small machine learning model. We show the first on-board detection of methane point source emission with Imaging Spectroscopy data using Edge ML.
Chinese Translation
航空和天基成像光谱观测活动所采集的数据能够用于检测痕量气体排放等瞬态事件。然而,目前的处理流程依赖于缓慢的地面处理,这延迟了每个检测事件的信息获取时间,并阻碍了即时后续行动的开展。在2026年3月的东京外场观测活动(Tokyo Field Campaign)期间,我们探索了对所搭载的AVIRIS-5传感器成像光谱数据的星载处理。由于通信带宽瓶颈,完整的数据立方体无法在飞行过程中立即下传。因此,我们仅下传由我们高效、轻量的机器学习模型预测出的潜在事件。我们展示了首次利用边缘机器学习(Edge ML)技术,基于成像光谱数据在星载平台上检测甲烷点源排放的成果。
cs.LG / 19 / 2609.04466

Nested Inductive Bias Framework for SPD Manifold Learning

用于SPD流形学习的嵌套归纳偏置框架
Das, Tushar
Abstract
In Geometric Deep Learning, inductive biases serve two primary functions: enforcing manifold constraints and embedding relational priors. Currently, representation learning on SPD manifolds frequently relies on pullback Euclidean metrics, such as the Log-Euclidean Metric, to satisfy the former. While computationally efficient in avoiding domain boundary violations, these metrics induce a flat geometry that may fail to capture the intrinsic relational priors of datasets. While metrics such as the Poincar\'e metric are widely utilized to induce domain-aligned relational priors, generalizing them from standard vector representations to the SPD manifold has remained a challenge. To bridge this gap, we introduce a Nested Inductive Bias framework that utilizes a two-stage diffeomorphic composition to formally pull back non-Euclidean target geometries onto the SPD manifold. This framework enables the construction of curvature-aligned Riemannian classifiers that simultaneously respect matrix constraints and the latent relational geometry of the data. Empirical evaluations on kinematic and signal processing benchmarks, together with synthetic experiments, demonstrate that deep manifold networks experience degradation in class separability unless the metric curvature aligns with the intrinsic data distribution. Furthermore, for standard vectorized architectures, we propose the Rational Conformal Metric (RCM), designed to establish state-of-the-art geometric robustness against outliers by bounding the representation space.
Chinese Translation
在几何深度学习中,归纳偏置发挥着两大主要功能:施加流形约束和嵌入关系先验。目前,SPD流形上的表示学习通常依赖拉回欧氏度量(如对数欧氏度量,Log-Euclidean Metric)来满足前者。尽管这类度量在计算上高效并能避免违反定义域边界,但其诱导的平坦几何可能无法捕捉数据集内在的关系先验。另一方面,庞加莱度量(Poincaré metric)等度量被广泛用于诱导与数据域对齐的关系先验,然而如何将其从标准向量表示推广到SPD流形一直是一个挑战。为弥合这一差距,我们提出了一个嵌套归纳偏置框架,该框架利用两阶段的微分同胚复合,将非欧氏目标几何形式化地拉回到SPD流形上。该框架使得构建曲率对齐的黎曼分类器成为可能,从而同时尊重矩阵约束和数据的潜在关系几何。在运动学和信号处理基准上的实证评估以及合成实验表明,若度量曲率与数据的内在分布不对齐,深度流形网络的类别可分性会出现退化。此外,针对标准的向量化架构,我们提出了有理共形度量(Rational Conformal Metric, RCM),通过对表示空间进行界定,实现了对离群点的最先进几何鲁棒性。
cs.LG / 20 / 2609.04494

Hakken: Predicting future discoveries to fill the gaps in today's knowledge

Hakken:预测未来发现以填补当今知识的空白
Besold, Tarek R., Akujuobi, Uchenna, Sanchez, Pablo, Toniato, Alessandra, Maruyama, Kana, Choi, Jihun, Badreddine, Samy, Gifford, Frederick, Evans-Yamamoto, Daniel, Palaniappan, Sucheendra K., Ferrer, Miquel, Nagano, Kae, Rossell, Iris, Joy, Tom, ElShazly, Hatem, Iliopoulou, Chrysa, Wehner, Christoph, Thanapalasingam, Thiviyan, Nunes, Susana, Cotovio, Pedro G., Wurman, Peter, Stone, Peter, Kitano, Hiroaki, Spranger, Michael
Abstract
We present Hakken, a domain-agnostic prediction and explanation system performing knowledge prediction, i.e., growing scientific knowledge by establishing novel relationships, ones that are not limited to the deductive hull of previous knowledge. Hakken uses a transformer-based prediction model built on temporal sequences of knowledge graphs extracted from vast bodies of research publications, fused with an LLM's semantic knowledge, to predict the presence and define the type of as-yet undocumented relationships between scientific concepts. It then calls a model-agnostic explanation framework to provide accompanying information for each prediction that allows scientists to evaluate the suggested new relationship. While general purpose, we demonstrate Hakken's practical capabilities by applying it to the biomedical domain. There, Hakken's prediction model establishes a new benchmark for time-aware multi-label relation prediction, and we show that the model's output stays coherent and informative over extended time spans in historic data. In addition, we scored 1.5 million above-confidence-threshold hypotheses related to aging, qualitatively validated batches of these predictions with biologists and progressed three of them for empirical validation in wet-lab. Two predictions with potentially significant impact in the context of drug discovery and repurposing were confirmed, introducing previously undocumented interactions between TP53 and BAMBI, and between RAF1 and TNF, to biomedical science.
Chinese Translation
我们提出了Hakken,一个领域无关的预测与解释系统,用于执行知识预测,即通过建立新颖的关系来扩展科学知识,这些关系不局限于既有知识的演绎范围。Hakken使用基于Transformer的预测模型,该模型构建于从海量研究文献中提取的知识图谱时间序列之上,并与大语言模型(LLM)的语义知识相融合,用以预测科学概念之间尚未被记录的关系的存在并定义其类型。随后,系统调用一个模型无关的解释框架,为每项预测提供伴随信息,使科学家能够评估所建议的新关系。虽然Hakken是通用性的,我们通过将其应用于生物医学领域来展示其实际能力。在该领域,Hakken的预测模型为时间感知的多标签关系预测建立了新的基准,并且我们证明该模型的输出在历史数据的长时间跨度内保持连贯且信息丰富。此外,我们对150万条与衰老相关的超过置信度阈值的假设进行了评分,与生物学家一起对这些预测批次进行了定性验证,并推进其中三项进入湿实验室的实证验证。其中两项在药物发现与药物再利用背景下具有潜在重大影响的预测得到确认,向生物医学科学引入了此前未被记录的TP53与BAMBI之间以及RAF1与TNF之间的相互作用。
cs.LG / 21 / 2609.04530

An Energy-Based Conservative-Dissipative Latent Neural Evolution Operator for Magnetization Dynamics

一种用于磁化动力学的基于能量的保守-耗散潜空间神经演化算子
Schaffer, Sebastian, Exl, Lukas
Abstract
We develop an energy-based reduced-order model for micromagnetic magnetization dynamics that couples a convolutional autoencoder to a structured latent neural ordinary differential equation. Motivated by the precessional-dissipative structure of the Landau-Lifshitz-Gilbert equation, the latent vector field is generated from the gradient of a learned scalar potential through an antisymmetric operator and a symmetric positive-semidefinite dissipative operator. This potential is learned in nonunique latent coordinates and is not identified with the Gibbs free energy, but decreases monotonically along autonomous continuous-time solutions, while the antisymmetric component permits motion along its level sets. The encoder, decoder, latent energy, and operators are trained jointly on short trajectory windows using latent and decoded-rollout losses alone, without time-derivative supervision, physical-energy labels, or dissipation penalties. At inference, an initial state is encoded once, evolved in latent space, and decoded only at the requested output times, enabling substantially cheaper trajectory prediction than the micromagnetic solver used to generate the training data. We compare quadratic, deep, and additive deep-quadratic latent energies on two datasets parameterized by field amplitude and generated for the two applied-field directions of the NIST $\mu$MAG Standard Problem 4. Dissipative-only and antisymmetric-dissipative models achieve comparable accuracy on short training-style windows but differ substantially on uninterrupted rollouts, for which the antisymmetric-dissipative models provide markedly more accurate trajectory predictions. The deep-quadratic energy gives the best overall accuracy for both field directions and exhibits slower error growth when rollouts are extended to twice the training horizon.
Chinese Translation
我们开发了一种基于能量的微磁磁化动力学降阶模型,该模型将卷积自编码器与结构化的潜空间神经常微分方程相耦合。受Landau-Lifshitz-Gilbert方程进动-耗散结构的启发,潜空间向量场由一个学习得到的标量势的梯度经过一个反对称算子和一个对称半正定耗散算子生成。该势能在非唯一的潜空间坐标中学习得到,并不等同于吉布斯自由能,但沿自治连续时间解单调递减,而反对称分量则允许沿其水平集的运动。编码器、解码器、潜空间能量以及各算子仅在短轨迹窗口上使用潜空间损失和解码滚动损失进行联合训练,无需时间导数监督、物理能量标签或耗散惩罚项。在推理阶段,初始状态仅被编码一次,在潜空间中演化,并仅在所请求的输出时刻进行解码,使得轨迹预测的成本显著低于用于生成训练数据的微磁求解器。我们在两个以磁场幅值为参数、并针对NIST μMAG标准问题4的两个外加磁场方向生成的数据集上,比较了二次型、深度型以及加性深度-二次型潜空间能量。仅耗散模型与反对称-耗散模型在短训练窗口上取得了相当的精度,但在不间断的滚动预测上差异显著,其中反对称-耗散模型提供了明显更精确的轨迹预测。深度-二次型能量在两个磁场方向上均给出最佳的整体精度,且当滚动预测延长至训练时域的两倍时,其误差增长更为缓慢。
cs.LG / 22 / 2609.04531

Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One

蒸馏后的连续扩散语言模型可以用少量步数——甚至一步——编写代码
Peng, Fred Zhangzhi, Zheng, Kaiwen, Zhang, Anru R.
Abstract
Language generation is almost universally treated as a sequential process: autoregressive models emit one token at a time, while diffusion language models replace token-level seriality with a long trajectory of iterative refinement. In this work, we introduce PlaidQ, a 0.7B continuous diffusion language model for code generation, and show that its trajectory can be aggressively distilled into only a few denoising steps---or even one, enabling efficient code generation. PlaidQ repurposes a pretrained autoregressive model as a bidirectional denoiser over continuous token embeddings. We distill PlaidQ with distribution matching for few-step generation and paired-trajectory supervision for one-step generation. At matched model scale, PlaidQ is competitive with discrete diffusion language models on code generation. Distillation then shifts the quality--compute frontier: a 16-step student reaches 31.78 and 40.49 pass@10 on HumanEval and MBPP+, surpassing the same PlaidQ teacher sampled for 512 steps. At the extreme, paired-trajectory distillation achieves 7.07 pass@1 on HumanEval with a single denoising step, producing functionally correct programs. Together, these results establish continuous diffusion as a viable path to few-step and one-step code generation. Broadly, continuous diffusion is not merely another representation for language: it provides an interface through which language models can inherit the acceleration and distillation machinery of continuous diffusion modeling. Training and inference code and model checkpoints are available at https://github.com/pengzhangzhi/plaidq.
Chinese Translation
语言生成几乎普遍被视为一个顺序过程:自回归模型逐个生成词元,而扩散语言模型则用长的迭代精炼轨迹取代词元级别的串行性。在本工作中,我们提出了PlaidQ,一个用于代码生成的0.7B参数连续扩散语言模型,并证明其生成轨迹可以被激进地蒸馏为仅少数去噪步数——甚至一步——从而实现高效的代码生成。PlaidQ将预训练的自回归模型改造为作用于连续词元嵌入上的双向去噪器。我们采用分布匹配方法对PlaidQ进行少步生成的蒸馏,并采用成对轨迹监督实现一步生成的蒸馏。在相同模型规模下,PlaidQ在代码生成任务上与离散扩散语言模型具有竞争力。蒸馏进一步移动了质量-算力的前沿:16步的学生模型在HumanEval和MBPP+上分别达到31.78和40.49的pass@10,超越了采样512步的同一个PlaidQ教师模型。在极端情况下,成对轨迹蒸馏仅用单步去噪便在HumanEval上取得7.07的pass@1,生成了功能正确的程序。这些结果共同确立了连续扩散作为实现少步和一步代码生成的可行路径。更广泛地说,连续扩散不仅仅是语言的另一种表示形式:它提供了一个接口,使语言模型能够继承连续扩散建模中的加速与蒸馏机制。训练与推理代码及模型检查点可在 https://github.com/pengzhangzhi/plaidq 获取。
cs.LG / 23 / 2609.04540

Mitra-v2 Technical Report

Mitra-v2 技术报告
Tao, Yefan, Zhang, Xiyuan, Liu, Xinyi, Han, Boran, Maddix, Danielle, Fang, Haoyang, Han, Zhen, Gai, Jiading, Liu, Xuanqing, Bohlke-Schneider, Michael, Yuyang, Wang, Friedland, Gerald, Mah, Kevan, Lee, Chris, Kong, Chris
Abstract
We introduce Mitra-v2, a tabular foundation model that delivers state-of-the-art performance on real-world classification and regression problems, from credit-risk scoring and clinical prediction to equipment-failure detection and house-price estimation. Mitra-v2 is trained only on synthetic data, with a pretraining distribution that is much larger and more diverse than Mitra-v1's. Built on a small 2D Transformer backbone, Mitra-v2 supports longer contexts and larger feature spaces. Improved optimization lets it learn from this larger task distribution. We evaluate Mitra-v2 on the TabArena and TALENT benchmarks, comprising more than 300 real-world datasets under two evaluation protocols. On the full TabArena benchmark, Mitra-v2 delivers state-of-the-art performance at the level of the industry-scale TabFM and EXAONE Tabular models, while surpassing TabPFN-3 by a wide margin in both classification and regression. Mitra-v2 matches the 1.6B-parameter TabFM with only 5% of its size (77M parameters), delivering frontier performance at a fraction of the cost. On TALENT, Mitra-v2 remains among the leading models, clearly outperforming TabPFN-3 and TabICLv2. It also ranks first on classification tasks with more than ten classes, even though it was pretrained only on tasks with at most ten classes. These results make Mitra-v2 one of the strongest and most broadly applicable open tabular foundation models released to date. We release the model weights, the inference and fine-tuning code, and our evaluation results under the Apache-2.0 license.
Chinese Translation
我们提出了 Mitra-v2,一个表格基础模型,在真实世界的分类和回归问题上实现了最先进的性能,应用场景涵盖信用风险评分、临床预测、设备故障检测和房价估算等。Mitra-v2 仅使用合成数据进行训练,其预训练分布比 Mitra-v1 规模更大、更加多样化。Mitra-v2 基于一个小型 2D Transformer 骨干网络构建,支持更长的上下文和更大的特征空间。经过改进的优化方法使其能够从这一更大的任务分布中学习。我们在 TabArena 和 TALENT 基准上对 Mitra-v2 进行了评估,在两种评估协议下涵盖 300 多个真实世界数据集。在完整的 TabArena 基准上,Mitra-v2 取得了与工业级规模的 TabFM 和 EXAONE Tabular 模型相当的最先进性能,同时在分类和回归任务上都大幅超越 TabPFN-3。Mitra-v2 仅用 5% 的参数量(7700 万参数)便匹配了 16 亿参数的 TabFM,以极低的成本实现了前沿性能。在 TALENT 上,Mitra-v2 仍位居领先模型之列,明显优于 TabPFN-3 和 TabICLv2。在类别数超过十类的分类任务上,它也排名第一,尽管其预训练仅使用了最多十类的任务。这些结果使 Mitra-v2 成为迄今为止最强大、适用范围最广的开源表格基础模型之一。我们在 Apache-2.0 许可证下发布了模型权重、推理与微调代码以及我们的评估结果。
cs.LG / 24 / 2609.04549

Fast Surrogate Modeling of Excitable and Oscillatory FitzHugh-Nagumo Dynamics with Parametric Neural Operators

基于参数化神经算子的可激发与振荡FitzHugh-Nagumo动力学快速代理建模
Franck, Andrew, Li, Justin
Abstract
The FitzHugh-Nagumo (FHN) system serves as a simplified model of neuronal voltage dynamics, capturing the activator-inhibitor structure behind both isolated action potentials and the rhythmic spiking seen across the brain. Exploring its 5D physiological parameter space is important for neuromodulation and mapping voltage recordings back to biophysics, yet classical finite-difference solvers make rapid parameter sweeps expensive. We train parameter-conditioned Fourier Neural Operators (FNOs) as fast, differentiable surrogates for the FHN voltage and recovery fields on a one-dimensional spatial domain, conditioning each Fourier layer on the parameter vector $\lambda = (D_u, D_v, a, b, \tau)$ via feature-wise linear modulation (FiLM). We apply a single bifurcation analysis that delimits the two distinct regimes the model spans, oscillatory (tonic firing) and excitable (action-potential propagation), and we train one operator in each. In the oscillatory regime the surrogate attains sub-$0.1\%$ relative $L^2$ error on both fields, runs nearly three orders of magnitude faster than the finite-difference baseline, generalizes uniformly across the parameter space, and extrapolates to low single-digit percentage errors outside of the training bounds. In the excitable regime the same operator accurately reproduces the firing threshold and the $c \propto \sqrt{D_u}$ conduction-velocity law and replicates full traveling pulses, fully capturing the excitable bifurcation structure rather than just smoothly interpolating fields.
Chinese Translation
FitzHugh-Nagumo(FHN)系统作为神经元电压动力学的简化模型,捕捉了孤立动作电位以及大脑中普遍存在的节律性放电背后的激活-抑制结构。探索其5维生理参数空间对于神经调控以及将电压记录反演至生物物理参数具有重要意义,然而经典的有限差分求解器使得快速参数扫描代价高昂。我们训练了以参数为条件的傅里叶神经算子(FNO)作为FHN电压场与恢复场在一维空间域上快速、可微分的代理模型,并通过特征级线性调制(FiLM)将参数向量 $\lambda = (D_u, D_v, a, b, au)$ 注入每个傅里叶层。我们进行了单一的分支分析,界定了模型所跨越的两种不同状态——振荡态(紧张性放电)与可激发态(动作电位传播),并在每种状态下分别训练一个算子。在振荡状态下,代理模型在两个场上均达到了低于0.1%的相对 $L^2$ 误差,运行速度比有限差分基准快近三个数量级,在整个参数空间上具有一致的良好泛化能力,并且在训练边界之外仍能外推至个位数百分比误差。在可激发状态下,同一算子准确再现了放电阈值和 $c \propto \sqrt{D_u}$ 的传导速度定律,并完整复现了行波脉冲,充分捕捉了可激发分支结构,而不仅仅是对场进行平滑插值。
cs.LG / 25 / 2609.04575

Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models

细粒度混合专家模型中激活专家的无训练减半方法
Chen, Xing, Yao, Hengshuai
Abstract
Modern fine-grained Mixture-of-Experts (MoE) models route each token to a small number of experts and renormalize their router probabilities. We show that this renormalization implicitly calibrates expert output gain to the training top-$k$: reducing $k$ at inference changes not only which experts are used but also the strength of the expert branch. We separate these effects by activating the top $k_1$ experts while normalizing by the probability mass of the top $k_2$ experts, introducing one integer with no parameters, training, or measurable compute overhead. On Qwen3.6-35B-A3B, reducing from 8 to 4 experts causes a 4.65-point MMLU drop under standard renormalization but only 0.35 points with $k_2=16$, while halving routed-expert compute. The result replicates on the $11\times$ larger Qwen3.5-397B-A17B, where reducing from 10 to 5 experts loses only 0.55 points with an appropriate reference set. Removing renormalization entirely is catastrophic, showing that preserving a suitable reference mass is crucial. We further find that perplexity and downstream accuracy favor different $k_2$, cautioning against selecting MoE compression settings using unlabeled text alone. Analyses also show that expert identity matters substantially more than expert weighting, while balanced and domain-specialized routing leaves limited room for expert pruning.
Chinese Translation
现代细粒度混合专家(Mixture-of-Experts, MoE)模型将每个词元路由到少量专家,并对其路由概率进行重归一化。我们证明,这种重归一化隐式地将专家输出增益校准到训练时的 top-$k$ 值:在推理时减小 $k$ 不仅改变了所使用的专家,也改变了专家分支的强度。我们通过激活 top $k_1$ 个专家同时以 top $k_2$ 个专家的概率质量进行归一化来分离这两种效应,该方法仅引入一个整数,无需任何参数、训练或可测量的计算开销。在 Qwen3.6-35B-A3B 上,将专家数从 8 减至 4 时,标准重归一化下 MMLU 下降 4.65 分,而在 $k_2=16$ 时仅下降 0.35 分,同时将路由专家的计算量减半。该结果在规模大 11 倍的 Qwen3.5-397B-A17B 上得到复现:将专家数从 10 减至 5 时,在合适的参考集下仅损失 0.55 分。完全去除重归一化则会导致灾难性后果,表明保留合适的参考概率质量至关重要。我们进一步发现,困惑度与下游准确率所偏好的 $k_2$ 值不同,这提醒人们不要仅使用无标注文本选择 MoE 压缩设置。分析还表明,专家身份的选择远比专家权重重要,而均衡路由和领域专门化路由留给专家剪枝的空间有限。
cs.LG / 26 / 2609.04577

Optimizer Memory Schedules for Outscaling the Overtraining Axis

用于超越过训练轴的优化器内存调度策略
Everett, Katie, Qiu, Shikai
Abstract
We investigate how optimizers scale across the overtraining axis and show that relative optimizer performance and optimal hyperparameters change substantially with training horizon. In particular, we study how matrix-preconditioned methods (Muon and SOAP) and a momentum-scheduled method (ADANA) scale relative to AdamW. We compare these four optimizers across models from 51M to 253M parameters and overtraining (OT) factors from 1x to 256x, sweeping the base learning rate at every setting. The preferred learning rate schedule can reverse across the overtraining axis, the best weight decay coefficient scales approximately as sqrt(OT), and longer horizons generally favor longer fixed memory. ADANA's scaling advantage over AdamW persists after tuning AdamW's fixed memory separately at each horizon. Log-time weight decay and momentum cooldown provide substantial gains for ADANA that compound as training increases. With this treatment, ADANA outscales AdamW with an exponent advantage close to that predicted by DANA theory on power-law random features. Muon and SOAP instead provide roughly constant token-efficiency advantages over AdamW across most of the measured range, although SOAP may gain further at the highest overtraining factors. ADANA begins behind both matrix-preconditioned optimizers but closes the gaps as training increases, surpassing Muon and becoming competitive with SOAP at our highest OT factors. These results establish training horizon as an essential axis for optimizer evaluation and design.
Chinese Translation
我们研究了优化器在过训练轴上的扩展规律,并表明优化器的相对性能和最优超参数会随训练时域(training horizon)发生显著变化。特别地,我们研究了矩阵预条件方法(Muon 和 SOAP)以及动量调度方法(ADANA)相对于 AdamW 的扩展表现。我们在参数量从 51M 到 253M 的模型上、在过训练(OT)因子从 1 倍到 256 倍的范围内,对这四种优化器进行了比较,并在每个设置下扫描了基础学习率。研究表明:最优学习率调度在过训练轴上可能发生反转;最优权重衰减系数近似随 sqrt(OT) 缩放;更长的训练时域通常偏好更长的固定内存。在针对每个训练时域分别调整 AdamW 的固定内存之后,ADANA 相对于 AdamW 的扩展优势依然存在。对数时间的权重衰减和动量冷却(momentum cooldown)为 ADANA 带来了显著收益,且这些收益随训练时长增加而不断叠加。经过上述处理,ADANA 在扩展性上超越了 AdamW,其指数优势接近 DANA 理论在幂律随机特征上所预测的水平。Muon 和 SOAP 则在大部分测量范围内为 AdamW 提供了大致恒定的 token 效率优势,尽管 SOAP 在最高过训练因子处可能获得进一步提升。ADANA 起初落后于两种矩阵预条件优化器,但随着训练时长增加逐步缩小差距,在我们的最高 OT 因子处超越了 Muon 并可与 SOAP 竞争。这些结果确立了训练时域作为优化器评估与设计中一个不可或缺的维度。
cs.LG / 27 / 2609.04583

Representation Redundancy and Structural Complexity in Finite-Field Inversion

有限域求逆中的表示冗余与结构复杂度
Zhang, Zheng, Zhang, Na
Abstract
The representation chosen for a mathematical operation can affect both its algebraic form and its empirical learning difficulty. We study this phenomenon for inversion over \(\mathbb F_{2^n}\), with field elements expressed in varying ordered \(\mathbb F_2\)-bases. We prove that two ordered bases induce the same coordinate inversion map if and only if they belong to the same Galois orbit. Since every orbit has size \(n\), the correspondence between ordered bases and distinct inversion maps is exactly \(n\)-to-one. We then analyze three Boolean formulations of inversion. The reference formulation has algebraic degree \(n-1\) and joint ANF leap \(1\), the mixed representation formulation has degree \(2(n-1)\) and joint ANF leap \(2\), and the complete raw formulation has degree at most \(3(n-1)\) and joint ANF leap at least \(n\). Exhaustive computations agree with the theoretical results and bounds in the cases considered. Controlled experiments with multilayer perceptrons show the same ordering in learning difficulty, while Galois orbit redundancy provides only a limited generalization benefit under the tested conditions. These results show that exact redundancy among representations can coexist with changes in Boolean structure and learning behavior when the representation is exposed as part of the input.
Chinese Translation
为某一数学运算所选的表示方式既会影响其代数形式,也会影响其实际学习难度。我们针对 \(\mathbb F_{2^n}\) 上的求逆运算研究这一现象,其中域元素用不同的有序 \(\mathbb F_2\)-基表示。我们证明:两个有序基诱导相同的坐标求逆映射当且仅当它们属于同一个伽罗瓦轨道(Galois orbit)。由于每个轨道的大小为 \(n\),有序基与不同求逆映射之间的对应关系恰好是 \(n\) 比 \(1\)。随后我们分析了求逆的三种布尔表述方式:基准表述的代数次数为 \(n-1\),联合 ANF 跃变(joint ANF leap)为 \(1\);混合表示表述的次数为 \(2(n-1)\),联合 ANF 跃变为 \(2\);完全原始表述的次数至多为 \(3(n-1)\),联合 ANF 跃变至少为 \(n\)。在所考察的情形中,穷举计算结果与理论结果和界一致。基于多层感知机(multilayer perceptrons)的受控实验显示了相同的学习难度排序,而在所测试的条件下,伽罗瓦轨道冗余仅带来有限的泛化收益。这些结果表明,当表示方式作为输入的一部分被暴露时,表示之间的精确冗余可以与布尔结构和学习行为的变化共存。
cs.LG / 28 / 2609.04593

GNN-Guided Graph Coarsening and Adaptive QUBO Penalties for the Capacitated Vehicle Routing Problem with Time Windows on a Quantum Annealer

基于图神经网络引导的图粗化与自适应QUBO惩罚的量子退火机上带时间窗 Capacitated 车辆路径问题求解方法
Rezk, Youssef Kamel, Gora, Paweł
Abstract
Graph coarsening reduces the large Quadratic Unconstrained Binary Optimization (QUBO) formulations arising when vehicle-routing problems are solved by quantum annealing. Nearby customers with compatible time windows are merged into super-nodes, the reduced problem is solved, and the solution is expanded to the original graph. For the Capacitated Vehicle Routing Problem with Time Windows (CVRPTW), existing coarsening heuristics require family-specific tuning and remain unreliable on random instances. We address these limitations on the Solomon benchmark using simulated annealing and a D-Wave Advantage2 processor. We first introduce adaptive penalty calibration. Uniform penalty scaling has little effect, whereas controlling the internal coefficient range substantially improves raw samples. Removing non-binding constraints, normalising binding ones, and scaling the remaining penalties reduces mean raw constraint violations from 33.0 to 0.06 at the same solver budget (p=3.7e-11, n=56). A variable-count-preserving control attributes this gain to conditioning rather than problem size. Second, we replace the hand-tuned merge score with a graph neural network (GNN) using one configuration across all families. At N=10, it achieves 100% feasibility across all Solomon families, including R-type (100% vs. 80% for the tuned heuristic). Across N=10,...,100, feasibility is 83% vs. 69%, with the GNN better or tied on 85/90 instance-size pairs. At N=80,100, the difference is significant (p=0.002; 25/25 pairs), while the QUBO remains approximately 5-6 times smaller. Finally, hardware experiments reproduce the conditioning effect at fixed logical variable count: feasible samples increase from 0.02% to 39% across 13 instances. Classical repair with local search remains a reference bound for end-to-end solution cost.
Chinese Translation
图粗化方法可以缩减使用量子退火求解车辆路径问题时产生的大规模二次无约束二值优化(QUBO)模型。具有兼容时间窗的邻近客户被合并为超级节点,求解缩减后的问题,再将解扩展回原始图。对于带时间窗的容量约束车辆路径问题(CVRPTW),现有的粗化启发式方法需要针对特定问题族进行调优,且在随机实例上可靠性不足。我们在Solomon基准测试上,采用模拟退火和D-Wave Advantage2处理器来解决这些局限性。首先,我们引入自适应惩罚校准。统一的惩罚缩放效果甚微,而控制内部系数范围则显著改善原始样本质量。通过移除非约束性约束、归一化约束性约束并缩放剩余惩罚,在相同求解器预算下,平均原始约束违反度从33.0降至0.06(p=3.7e-11,n=56)。一个保持变量数量的对照实验表明,这一改进源于条件数的优化而非问题规模的缩减。其次,我们用图神经网络(GNN)替代人工调优的合并评分,并在所有问题族上使用统一配置。在N=10时,该方法在所有Solomon问题族上均达到100%可行性,包括R类问题(100%,而调优后的启发式方法为80%)。在N=10至100的范围内,可行性分别为83%对69%,且GNN在85/90个实例-规模组合上表现更优或持平。在N=80和N=100时,差异具有统计显著性(p=0.002;25/25组合),同时QUBO规模仍保持约5-6倍的缩减。最后,硬件实验在固定逻辑变量数量的条件下复现了条件数优化的效果:在13个实例上,可行样本占比从0.02%提升至39%。结合局部搜索的经典修复方法仍是端到端求解成本的参考上界。
cs.LG / 29 / 2609.04635

Too Rare to Learn: Prescribed Cyclone Tracks Degrade a Bay of Bengal Ocean Emulator

过于罕见而无法学习:预设气旋轨迹降低孟加拉湾海洋模拟器的性能
Islam, Sumaiya
Abstract
Neural ocean emulators are being proposed for regional forecasting in cyclone-exposed coastal seas, and a natural design choice is to hand the network the cyclone as a prescribed input. We test that choice in the Bay of Bengal and find it harmful. We withhold 15 whole cyclones spanning 65 to 150 kt from GLORYS12 reanalysis and compare two U-Nets that are identical except for four prescribed cyclone-track channels. Across three seeds the ocean-only model beats persistence in every run and the storm-conditioned model loses to it in every run, with the two skill ranges disjoint (p = 3.1e-5, paired across storms). The cause is exposure frequency rather than signal content: the channels are non-zero on only 7.9% of training days, so they are out of distribution the moment they activate. The extra error falls inside the prescribed storm footprint, and replacing the real cyclone map with a no-storm map at inference improves held-out storm forecasts by 7.5 to 16.4% in every seed. The conditioned network has learned a response to a rare signal that is confidently wrong.
Chinese Translation
神经海洋模拟器正被提议用于易受气旋影响的近岸海域的区域预报,一个自然的设计选择是将气旋作为预设输入提供给网络。我们在孟加拉湾对这一选择进行了测试,发现其有害。我们从GLORYS12再分析数据中留出15个完整气旋(风速范围65至150节),并比较了两个完全相同的U-Net模型,唯一区别在于四个预设气旋轨迹通道。在三个随机种子下,仅海洋模型在每次运行中都优于持续性基准,而气旋条件化模型在每次运行中都逊于持续性基准,两者的技能区间完全不重叠(p = 3.1e-5,跨气旋配对检验)。原因在于暴露频率而非信号内容:这些通道仅在7.9%的训练天数中为非零值,因此一旦激活就处于分布之外。额外的误差落在预设风暴足迹范围内,并且在推断阶段用无风暴图替换真实气旋图后,每个种子下留出气旋的预报准确率均提高了7.5%至16.4%。条件化网络学会了对一种罕见信号的响应,但这种响应是自信却错误的。
cs.LG / 30 / 2609.04639

SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery

SMILE:连接连续优化与离散符号恢复
Montazerin, Mansooreh, Ortega, Antonio, Srivastava, Ajitesh
Abstract
Symbolic regression (SR) discovers closed-form mathematical expressions from data, offering interpretability beyond black-box models. Existing methods suffer from slow convergence in combinatorial search spaces and lack mechanisms to exploit compositional structure in the data. We introduce SMILE (Sine, Multiplication, Identity, Logarithm, Exponential), a hybrid framework that unifies continuous gradient-based optimization with discrete symbolic recovery through three stages: structural analysis of the data to identify the compositional hierarchy of the target expression, continuous optimization to learn parameters of a network that encodes the target expression using interpretable activations, and symbolic recovery through structured pruning, coefficient optimization, and rounding. This final stage distills the learned network into a compact expression with exact symbolic constants. We evaluate SMILE on SRBench across ground-truth and black-box datasets, with ablation studies validating each component. SMILE achieves the highest symbolic solution rate at the largest noise levels, demonstrating strong robustness where competing methods degrade substantially. It consistently lies on the Pareto front of accuracy versus complexity, recovering significantly simpler expressions in a fraction of the time required by the competing methods.
Chinese Translation
符号回归(SR)从数据中发现闭式数学表达式,提供了超越黑盒模型的可解释性。现有方法在组合搜索空间中收敛缓慢,且缺乏利用数据中组合结构的机制。我们提出了SMILE(Sine、Multiplication、Identity、Logarithm、Exponential),一个将基于梯度的连续优化与离散符号恢复相统一的混合框架,包含三个阶段:对数据进行结构分析以识别目标表达式的组合层次结构;通过连续优化学习一个使用可解释激活函数、编码目标表达式的网络参数;以及通过结构化剪枝、系数优化和取整实现符号恢复。最后一个阶段将学习到的网络提炼为具有精确符号常数的紧凑表达式。我们在SRBench上对SMILE进行了评估,涵盖真实标签数据集和黑盒数据集,并通过消融研究验证了每个组件的有效性。SMILE在最大噪声水平下获得了最高的符号解成功率,展现出强鲁棒性,而竞争方法在此条件下性能大幅下降。SMILE始终位于精度与复杂度权衡的帕累托前沿上,并以竞争方法所需时间的一小部分恢复出明显更简洁的表达式。
cs.LG / 31 / 2609.04661

Interpretability for Turing Machines

图灵机的可解释性研究
Snikkers, Billy, Salazar, Rumi, Murfet, Daniel, Troiani, Will
Abstract
We show that susceptibilities, an interpretability technique developed for neural networks, can identify the presence of algorithmic structure in Turing machines by probing the local loss landscape of a learning problem for noisy Turing machines introduced by Murfet and Troiani (arXiv:2504.08075). We prove that symmetries and path separation in the algorithm implemented by a Turing machine induce permutation symmetries and low-rank blocks in its susceptibility matrix. We study this empirically on a set of deterministic finite automata (DFAs) and demonstrate that algorithmic features can be recovered by principal component analysis and clustering methods in susceptibility space.
Chinese Translation
我们证明了敏感性(susceptibilities)——一种为神经网络开发的可解释性技术——能够通过探查由 Murfet 和 Troiani(arXiv:2504.08075)提出的噪声图灵机学习问题的局部损失landscape,来识别图灵机中算法结构的存在。我们证明了图灵机所实现算法中的对称性和路径分离会在其敏感性矩阵中诱导置换对称性和低秩块。我们在一组确定性有限自动机(DFA)上对此进行了实证研究,并证明可以通过主成分分析和聚类方法在敏感性空间中恢复算法特征。
cs.LG / 32 / 2609.04672

WEECFP-SuRGE: Wide Embedded Extended Connectivity Fingerprint with Substructure Rotary Graph-distance Encoding

WEECFP-SuRGE:带子结构旋转图距离编码的宽嵌入扩展连接性指纹
Epps, Robert
Abstract
We introduce WEECFP, a parameter-free 1024-dimensional continuous molecular fingerprint that scatters each Morgan substructure across roughly thirty-two signed positions of a single vector, and WEECFP-SuRGE, a transformer architecture whose self-attention applies SuRGE (Substructure Rotary Graph-distance Encoding) -- a RoPE-like rotation parameterized by molecular shortest-path graph distance -- to WEECFP substructure tokens. A 7-model blend of this architecture (the WEECFP-SuRGE Blend) achieves the lowest average regression rank on the TDC ADMET leaderboard; is #2 overall on the TDC ADMET leaderboard (behind only pretrained MapLight+GNN), and is #1 overall among methods that use no external pretraining; takes leaderboard #1 finishes on Pgp, Lipophilicity, CYP2D6 Substrate, Clearance Microsome, and LD50 (with the WEECFP-NoSuRGE Blend separately reaching #1 on HIA) across the full 22-benchmark suite -- without any external pretraining. On MoleculeNet, WEECFP-SuRGE beats every classical-fingerprint baseline on 3 of 4 regression tasks (ESOL, Lipophilicity, QM9). We further show that WEECFP tokenization is near-lossless: a greedy overlap reconstruction recovers the exact canonical SMILES of 99.9% of in-distribution molecules across 9 MoleculeNet datasets and 98.93% of molecules in a cross-dataset holdout (HIV->Lipophilicity), and that a three-reference farthest-first encoding of graph distance correlates at Pearson r = 0.901 with the true pairwise distance, enabling O(S) positional memory at matching accuracy.
Chinese Translation
我们提出了 WEECFP,一种无需参数的1024维连续分子指纹,它将每个 Morgan 子结构散射到单个向量的大约三十二个带符号位置上;同时提出了 WEECFP-SuRGE,一种 transformer 架构,其自注意力机制对 WEECFP 子结构 token 应用了 SuRGE(Substructure Rotary Graph-distance Encoding,子结构旋转图距离编码)——一种由分子最短路径图距离参数化的类 RoPE 旋转编码。该架构的7模型混合模型(WEECFP-SuRGE Blend)在 TDC ADMET 排行榜上取得了最低的平均回归排名;在 TDC ADMET 排行榜总体排名第2(仅次于预训练的 MapLight+GNN),并在所有不使用外部预训练的方法中排名第1;在完整的22个基准测试套件中,其在 Pgp、脂溶性(Lipophilicity)、CYP2D6 底物、微粒体清除率(Clearance Microsome)和 LD50 上均获得排行榜第1(其中 WEECFP-NoSuRGE Blend 另外在 HIA 上获得第1)——且无需任何外部预训练。在 MoleculeNet 上,WEECFP-SuRGE 在4个回归任务(ESOL、Lipophilicity、QM9)中的3个上超越了所有经典指纹基线方法。我们进一步证明 WEECFP 的 token 化近乎无损:通过贪婪重叠重构,在9个 MoleculeNet 数据集上可以精确恢复99.9%的分布内分子的规范 SMILES,在跨数据集保留测试(HIV->Lipophilicity)中恢复98.93%的分子;此外,一种基于三个参考点的最远优先图距离编码与真实成对距离的 Pearson 相关系数达到 r = 0.901,从而在匹配精度的情况下实现了 O(S) 的位置记忆开销。
cs.LG / 33 / 2609.04710

Simulation-free Unbalanced Dynamic Optimal Transport with General Growth Penalty

具有一般增长惩罚的无模拟非平衡动态最优传输
Ying, Junda, Wang, Yuxuan, Yang, Bowen, Zhou, Peijie, Zhang, Lei
Abstract
Inferring cellular dynamics from unpaired single-cell snapshots requires modeling both state transitions and population growth or death. Unbalanced dynamic optimal transport (UDOT) addresses this by penalizing growth along transport paths, making the choice of growth penalty a key way to encode biological priors on proliferation and apoptosis. However, existing UDOT solvers either rely on computationally expensive NeuralODE simulations or depend on analytical solutions of conditional paths, restricting their efficiency solely to quadratic penalties, i.e. Wasserstein-Fisher-Rao (WFR) geodesics. To enable an efficient UDOT solver for general growth penalties, we first show that concave growth penalties lead to degenerate solutions where growth and transport are separated. We then introduce \textbf{S}imulation-free \textbf{U}nbalanced \textbf{D}ynamic \textbf{O}ptimal transport (SUDO), a simulation-free framework for UDOT with general non-quadratic convex growth penalties. SUDO learns the conditional paths and transport costs, solves the induced semi-coupling problem, and subsequently leverages unbalanced flow matching to achieve a simulation-free solution. On WFR benchmarks, SUDO matches the accuracy of efficient, analytical solution-driven algorithms while outperforming simulation-based methods in computational speed. Beyond WFR, SUDO supports asymmetric penalties that encode proliferation-dominant priors and produce more plausible trajectories and growth estimates on synthetic and single-cell datasets.
Chinese Translation
从非配对的单细胞快照中推断细胞动力学需要对状态转变以及种群增殖或死亡进行建模。非平衡动态最优传输(Unbalanced Dynamic Optimal Transport, UDOT)通过对传输路径上的增长施加惩罚来解决这一问题,因此增长惩罚的选择成为编码增殖和凋亡相关生物学先验的关键方式。然而,现有的UDOT求解器要么依赖计算代价高昂的NeuralODE模拟,要么依赖条件路径的解析解,这将其效率限制在二次惩罚上,即Wasserstein-Fisher-Rao(WFR)测地线。为了构建支持一般增长惩罚的高效UDOT求解器,我们首先证明凹增长惩罚会导致增长与传输相分离的退化解。随后,我们提出了无模拟非平衡动态最优传输(Simulation-free Unbalanced Dynamic Optimal Transport, SUDO),这是一个面向一般非二次凸增长惩罚的无模拟UDOT框架。SUDO学习条件路径和传输代价,求解由此导出的半耦合问题,进而利用非平衡流匹配(unbalanced flow matching)实现无模拟求解。在WFR基准测试中,SUDO达到了基于解析解的高效算法的精度,同时在计算速度上优于基于模拟的方法。超越WFR,SUDO支持编码增殖主导先验的非对称惩罚,并在合成数据集和单细胞数据集上产生更合理的轨迹和增长估计。
cs.LG / 34 / 2609.04721

Locating and Steering Refusal Beyond Attention

在注意力机制之外定位与引导拒绝行为
Bosco, Preethi Carmel, Srinivasan, Gopalakrishnan
Abstract
Where inside a language model does refusal live, and does that place change when the architecture does? In a transformer, refusal is governed by a single direction in the residual stream, a finding that safety and interpretability tooling now depend on. State-space models (SSMs) route information through a recurrent update instead of attention, sharing no token-mixing mechanism with a transformer. Does the same safety representation survive this shift, or must it be rediscovered per architecture? It survives. A single rigid rotation, which can only reorient a space and not reshape it, aligns one model's representation space with another's, so the two genuinely share the representation. A harm probe trained on a transformer then flags an SSM's harmful inputs, and removing the aligned direction makes a model answer attacks it would otherwise refuse, while a random direction of the same size does far less. What is architecture-specific is not where the direction is steered but where it must be read. Each layer computes a fresh output that is then added into the residual stream, and harm is cleanly readable at this output, the write site, before the addition. A control that holds the intervention's strength fixed shows that what matters is where the direction is estimated, not where it is applied. Applied through a detector-triggered gate, this direction lowers jailbreak success in all four architecture families we test (SSM, transformer, recurrent, hybrid), and on the SSM it holds against an attacker that tunes its prompt against the defense. The gate only matches a trivial rule that returns a fixed refusal whenever the same detector fires, so what transfers across architectures is the direction itself, not defense strength. Safety tooling built on refusal therefore ports to a new architecture by re-estimating the direction at that architecture's write site, not by rebuilding it.
Chinese Translation
语言模型中的拒绝(refusal)行为究竟存在于何处?当架构改变时,这个位置是否会变化?在Transformer模型中,拒绝行为由残差流(residual stream)中的单一方向所支配,这一发现现已成为安全性与可解释性工具的基础。状态空间模型(SSM)通过循环更新而非注意力机制来传递信息,与Transformer不共享任何词元混合机制。那么,同样的安全表征能否在这一转变中存续,还是必须针对每种架构重新发现?答案是它能存续。一次单一的刚性旋转——它只能重新定向一个空间而不能重塑它——就能将一个模型的表征空间与另一个模型对齐,因此二者确实共享该表征。在Transformer上训练的危害探针(harm probe)随后可以识别SSM的有害输入,而移除对齐后的方向会使模型回答它原本会拒绝的攻击,而相同大小的随机方向则效果甚微。架构特异性不在于方向在哪里被引导,而在于方向必须在哪里被读取。每一层都会计算一个新的输出,然后将其加到残差流中,而在相加之前的这个输出处,即写入位点(write site),危害可以被清晰地读取。一项保持干预强度固定的对照实验表明,关键在于方向在哪里被估计,而不是在哪里被应用。通过检测器触发的门控机制应用该方向,可以降低我们测试的全部四种架构家族(SSM、Transformer、循环架构、混合架构)中的越狱(jailbreak)成功率,且在SSM上,该方向能够抵御针对该防御进行提示词调优的攻击者。该门控仅相当于一条简单规则,即当同一检测器触发时返回固定的拒绝回复,因此跨架构迁移的是方向本身,而非防御强度。因此,基于拒绝行为构建的安全工具只需在该架构的写入位点重新估计方向,即可移植到新架构,而无需重建。
cs.LG / 35 / 2609.04735

Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling

基于合成任务扩展的小分子设计大语言模型训练
Hu, Frank, Chennakesavalu, Shriram, Wang, Zichen, Suriana, Patricia, Vani, Bodhi, Shmilovich, Kirill, Chuang, Kangway, Grambow, Colin
Abstract
Designing viable drug candidates requires searching a combinatorially large and rugged chemical space for molecules that satisfy multiple, often competing, objectives. Large language models (LLMs) provide a useful generative prior for this problem because of their representational capacity, reasoning ability, and flexibility when incorporating information from the external environment. While reinforcement learning from verifiable rewards (RLVR) can be used to improve the capabilities of LLMs, many chemically relevant scoring functions require hours or even days per evaluation, making them prohibitively expensive to use directly during online training. Here, we investigate whether LLMs can learn molecular design strategies from cheaper synthetic tasks that generalize to expensive molecular lead optimization settings. We find that curriculum-based training recipes that gradually incorporate more challenging synthetic design tasks enable strong performance that surpasses that of much larger frontier models on structure-based lead optimization. Our results suggest that scaling post-training using synthetic tasks is an effective strategy for adapting LLMs to high-cost experimental scenarios that are too expensive to directly train on.
Chinese Translation
设计可行的候选药物需要在组合爆炸且崎岖复杂的化学空间中搜索满足多个(且常常相互冲突)目标的分子。大语言模型(LLM)凭借其表征能力、推理能力以及整合外部环境信息的灵活性,为这一问题提供了有用的生成先验。虽然基于可验证奖励的强化学习(RLVR)可用于提升大语言模型的能力,但许多与化学相关的评分函数每次评估需要数小时甚至数天,使得在在线训练中直接使用它们的代价过于高昂。本文研究了大语言模型能否从更廉价的合成任务中学习分子设计策略,并将所学策略泛化到昂贵分子先导化合物优化场景。我们发现,逐步引入更具挑战性的合成设计任务的基于课程学习的训练方案,能够在大语言模型中实现强大的性能,在基于结构的先导化合物优化任务上超越了规模大得多的前沿模型。我们的结果表明,利用合成任务扩展后训练是使大语言模型适应直接训练成本过高的高成本实验场景的一种有效策略。
cs.LG / 36 / 2609.04754

A Fairness Audit of the Duckworth-Lewis-Stern Method: Format-Specific and Gender-Differential Bias, with an Interpretable Calibration Layer for Cricket Target Revision

Duckworth-Lewis-Stern(DLS)方法的公平性审计:针对比赛形式与性别的差异化偏差,以及用于板球目标分数修正的可解释校准层
Roy, Soumyadeep
Abstract
The Duckworth-Lewis-Stern (DLS) method has been the international standard for revising target scores in rain-interrupted limited-overs cricket since 1999. Despite over two decades of operational use, no large-scale empirical audit of its prediction bias has been published. We conduct such an audit on 8,150 international matches (3,095 ODIs, 5,055 T20Is) from Cricsheet, generating 233,550 synthetic interruption scenarios with temporal splits. We document two structured biases. First, DLS prediction error spans a 137-run range across (overs-remaining, wickets-lost) match-state buckets. Second, DLS exhibits a gender-differential bias on ODIs that has not previously been quantified: on the training split, mean over-prediction is +1.51 runs for men but +7.63 runs for women, a gap of +6.13 runs (F = 195.16, p < 10^-43). We benchmark DLS against five modern alternatives: Bi-LSTM, XGBoost, an enriched XGBoost variant, a deep context-aware model, and a stacking ensemble, and propose DLS-Cal, a lightweight interpretable calibration layer (27K parameters) outputting a state-conditioned correction added to DLS. DLS-Cal reduces absolute bias by 31% on ODI and 19% on T20I, and a gender-aware variant reduces women's ODI residual bias from +6.19 to +0.65 runs while leaving men's calibration unchanged. We release code, models, and data.
Chinese Translation
自1999年以来,Duckworth-Lewis-Stern(DLS)方法一直是雨天中断的限时板球比赛中修正目标分数的国际标准。尽管已运行使用二十余年,但关于其预测偏差的大规模实证审计尚未发表。我们基于Cricsheet的8,150场国际比赛(3,095场ODI,5,055场T20I),生成233,550个采用时间划分的合成中断场景,开展了此类审计。我们记录了两种结构性偏差。第一,DLS的预测误差在(剩余轮次、已失三柱门数)比赛状态分桶之间跨越137分的范围。第二,DLS在ODI中表现出此前未被量化的性别差异偏差:在训练集上,男性比赛的平均高估为+1.51分,而女性比赛为+7.63分,差距为+6.13分(F = 195.16,p < 10^-43)。我们将DLS与五种现代替代方法进行基准比较:Bi-LSTM、XGBoost、增强型XGBoost变体、深度上下文感知模型以及堆叠集成模型,并提出DLS-Cal——一个轻量级可解释校准层(27K参数),输出基于比赛状态的修正量并叠加到DLS结果上。DLS-Cal将ODI中的绝对偏差降低31%,T20I中降低19%;一个性别感知的变体将女性ODI的残余偏差从+6.19分降至+0.65分,同时保持男性校准不变。我们公开了代码、模型和数据。
cs.LG / 37 / 2609.04763

Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated Learning

超越平稳客户端可用性的弹性:解锁高效且无偏的联邦学习
Xiang, Ming, Ioannidis, Stratis, Yeh, Edmund, Joe-Wong, Carlee, Su, Lili
Abstract
Due to resource constraints or external and internal uncertainties, clients in real-world federated learning systems are often intermittently available edge devices. In highly dynamic environments, the parameter server lacks prior real-time knowledge of clients' availability, making it challenging to adapt traditional federated learning algorithms to be resilient to uncertainties in client availability. If not carefully addressed, complex client availability can introduce significant bias, potentially harming the performance of the trained model. Most prior work either fails to account for non-stationary client availability dynamics or demands significant memory and computational overhead. This paper aims to develop efficient federated learning algorithms that are provably resilient to heterogeneous and non-stationary stochastic client availability. We propose FedSWE, which admits novel algorithmic structures to (i) compensate for missed computations, (ii) stabilize and diffuse the global updates over rounds, and (iii) evenly mix the local updates through implicit gossiping, despite being agnostic to non-stationary dynamics. Compared with the standard FedAvg, FedSWE introduces light additional memory and computation overhead. We show that FedSWE converges to a stationary point of non-convex objectives while achieving the desired linear speedup property in certain special cases. We corroborate our analysis with numerical experiments over diversified client unavailability dynamics on real-world data sets.
Chinese Translation
由于资源限制或外部与内部的不确定性,现实世界联邦学习系统中的客户端通常是间歇性可用的边缘设备。在高度动态的环境中,参数服务器缺乏对客户端可用性的先验实时了解,这使得传统的联邦学习算法难以适应并对客户端可用性中的不确定性保持弹性。如果不加以仔细处理,复杂的客户端可用性会引入显著的偏差,可能损害训练模型的性能。以往的大多数工作要么未能考虑非平稳的客户端可用性动态,要么需要大量的内存和计算开销。本文旨在开发对异构且非平稳的随机客户端可用性具有可证明弹性的高效联邦学习算法。我们提出了FedSWE,它采用了新颖的算法结构,以便在不对非平稳动态有所知晓的情况下:(i) 补偿错过的计算,(ii) 在多轮之间稳定并扩散全局更新,以及 (iii) 通过隐式gossip机制均匀混合本地更新。与标准的FedAvg相比,FedSWE仅引入了较少的额外内存和计算开销。我们证明FedSWE能够收敛到非凸目标的驻点,并在某些特殊情况下实现期望的线性加速特性。我们通过在真实数据集上针对多样化的客户端不可用动态进行数值实验,验证了我们的分析。
cs.LG / 38 / 2609.04772

A Robust Watermark-based Fingerprint Framework for GNNs Ownership Verification

一种用于图神经网络(GNN)所有权验证的鲁棒水印指纹框架
Zhang, Han, Wang, Yan, Liu, Guanfeng, Ding, Pengfei, Wang, Huaxiong, Lam, Kwok-Yan
Abstract
The high training cost of Graph Neural Networks (GNNs) has raised growing concerns regarding model ownership infringement, such as model stealing and unauthorized misuse. To verify model ownership and prevent significant economic losses, two groups of GNN Ownership Verification (OV) methods have been proposed: watermark-based methods and fingerprint-based methods. However, these methods typically face three limitations: (1) the performance degradation of protected models caused by out-of-distribution (OOD) watermark graphs with respect to the training set; (2) the unrealistic assumption that surrogate models have been trained on a watermark-containing training set; and (3) over-reliance on specific output levels for fingerprint extraction. In this paper, we propose a Robust watErMArk-based fingeRprint frameworK for GNNs, named REMARK. REMARK first generates carefully crafted in-distribution watermark graphs that maximize output differences between GNN models, thus mitigating OOD-induced performance degradation. REMARK then extracts robust fingerprints from these output differences to verify GNN ownership, thereby removing the assumptions that surrogate models must be trained on a watermark-containing dataset or expose specific output levels. Extensive experiments across widely used real-world datasets and GNN architectures demonstrate that REMARK achieves state-of-the-art OV accuracy and robustness while preserving the utility of protected models.
Chinese Translation
图神经网络(GNN)高昂的训练成本引发了人们对模型所有权侵权问题(如模型窃取和未经授权的滥用)的日益关注。为了验证模型所有权并防止重大经济损失,研究者已提出两类GNN所有权验证(Ownership Verification, OV)方法:基于水印的方法和基于指纹的方法。然而,这些方法通常面临三个局限:(1)由于水印图相对于训练集呈分布外(OOD)特性,导致受保护模型的性能下降;(2)依赖于代理模型已在包含水印的训练集上训练过这一不切实际的假设;(3)过度依赖特定输出层级进行指纹提取。在本文中,我们提出了一种针对GNN的鲁棒水印指纹框架,命名为REMARK(Robust watErMArk-based fingeRprint frameworK)。REMARK首先生成经过精心设计的分布内水印图,以最大化GNN模型之间的输出差异,从而缓解由OOD导致的性能下降。随后,REMARK从这些输出差异中提取鲁棒指纹以验证GNN的所有权,从而去除了代理模型必须在包含水印的数据集上训练或暴露特定输出层级的假设。在广泛使用的真实世界数据集和GNN架构上进行的大量实验表明,REMARK在实现最先进的OV准确性和鲁棒性的同时,还能保持受保护模型的实用性。
cs.LG / 39 / 2609.04773

Persistent Teacher Anchoring for Tool-Using Agents

面向工具使用智能体的持续性教师锚定
Park, Hyun Bin, Song, Kyungho, Lee, Sangmin, Chang, Du-Seong
Abstract
Distillation is common in LLM post-training, where on-policy knowledge distillation (OPKD) uses student-generated trajectories to prepare the student for downstream RL. At each state, the student matches a next-token distribution supplied by the teacher. As the rollout enters states the teacher would not visit, the teacher-student distribution gap can accumulate. In tool use, this gap becomes consequential because student-written calls execute before supervision and their observations shape later prefixes. Proposer-verifier generation addresses this drift by letting the teacher decide which student-proposed text is retained during generation. Existing formulations govern text but leave tool execution outside their scope. We propose Persistent Teacher Anchoring (PTA), a student-induced but teacher-committed rollout construction. PTA retains chunk-level verification and adds turn-level commitment, allowing a call to reach the environment only after the teacher has verified the entire turn. Treating verified chunks as atomic generation units, we introduce persistent lookahead, which fills idle rollout capacity by advancing future samples and carrying unfinished ones across student updates under the fixed verifier. Across Search-R1-style retrieval and DeepEyes-style perception RL, applying PTA before downstream RL improves macro best@4 by 2.5 and 2.8 points over OPKD under the same downstream RL budget, while lookahead improves throughput by 24%.
Chinese Translation
蒸馏在大语言模型后训练中十分常见,其中在策略知识蒸馏(OPKD)利用学生生成的轨迹为学生的下游强化学习做准备。在每个状态,学生都需要匹配由教师提供的下一个词元分布。当 rollout 进入教师不会访问的状态时,师生分布差距会不断累积。在工具使用场景中,这种差距影响尤为重大,因为学生编写的调用在获得监督之前就已执行,且其返回的观测结果会塑造后续的前缀内容。提议者-验证者(proposer-verifier)生成机制通过让教师在生成过程中决定保留哪些学生提议的文本来解决这种漂移问题。然而现有方法只约束文本,将工具执行排除在其管辖范围之外。我们提出持续性教师锚定(Persistent Teacher Anchoring,PTA),一种由学生引导但由教师承诺的 rollout 构建方式。PTA 保留了块级(chunk-level)验证,并增加了轮次级(turn-level)承诺,只有当教师验证完整个轮次后,调用才允许进入环境。我们将验证过的块视为原子生成单元,并引入持续性前瞻(persistent lookahead)机制,通过在固定的验证者下推进未来的样本并在学生更新时携带未完成的样本,来充分利用空闲的 rollout 容量。在 Search-R1 风格的检索和 DeepEyes 风格的感知强化学习任务上,在相同的下游强化学习预算下,在下游强化学习之前应用 PTA 相比 OPKD 分别将 macro best@4 提升了 2.5 和 2.8 个百分点,同时前瞻机制将吞吐量提升了 24%。
cs.LG / 40 / 2609.04779

Dynamic Heterogeneous Graph Representation Learning: A Survey

动态异构图表示学习:综述
Liu, Huan, Jiao, Pengfei, Yin, Jie, Chen, Hongjiang, Zhao, Zhidong
Abstract
Graph representation learning (GRL) serves as a canonical paradigm for modeling complex networks. However, real-world AI systems inherently manifest as evolving heterogeneous entities with complex interactions, posing significant challenges to static or homogeneous modeling. To address these complexities, representation learning for Dynamic Heterogeneous Graphs (DHGs) has emerged as a vital approach for learning low-dimensional representations that simultaneously preserve structural semantics and temporal dynamics. This survey presents the first systematic review of DHG representation learning methods. We first introduce a unified formal definition that encompasses both discrete-time and continuous-time DHGs from the perspective of temporal granularity. Building upon this formulation, we propose a novel algorithm-centric taxonomy that categorizes existing literature, including early embedding-based approaches, graph neural network (GNN)-based models, and relatively recent Transformer-based DHG methods, while explicitly highlighting their intrinsic modeling biases with respect to dynamic granularity. Furthermore, we summarize representative applications of DHG representation learning, along with commonly used datasets and benchmarks. Finally, we discuss promising research directions that guide future advances in this rapidly evolving field.
Chinese Translation
图表示学习(Graph Representation Learning, GRL)是建模复杂网络的经典范式。然而,现实世界中的人工智能系统本质上表现为不断演化的异构实体及其复杂的交互关系,这给静态或同构建模带来了重大挑战。为应对这些复杂性,动态异构图(Dynamic Heterogeneous Graphs, DHGs)表示学习应运而生,成为一种学习低维表示的重要方法,能够同时保留结构语义与时序动态特性。本综述首次对DHG表示学习方法进行了系统性回顾。我们首先从时间粒度的视角出发,给出了涵盖离散时间和连续时间DHG的统一形式化定义。基于该定义,我们提出了一种以算法为中心的新颖分类体系,将现有文献划分为早期基于嵌入的方法、基于图神经网络(GNN)的模型,以及较为新兴的基于Transformer的DHG方法,并明确指出了它们在动态粒度方面固有的建模偏好。此外,我们总结了DHG表示学习的代表性应用,以及常用的数据集和基准测试。最后,我们讨论了有前景的研究方向,以期为这一快速发展的领域的未来进展提供指导。
cs.LG / 41 / 2609.04787

Learning-Augmented Algorithms: Guarantees, Construction Mechanisms, and System-Level Implications

学习增强算法:保证、构造机制与系统层面的影响
Zhao, Hailiang, Chen, Peng, Tang, Xueyan, Yin, Jianwei, Deng, Shuiguang
Abstract
Learning-augmented algorithms use fallible predictions while retaining formal performance guarantees. This survey synthesizes prediction interfaces, error measures, consistency--robustness trade-offs, and five representative construction mechanisms across online optimization, caching, learned data structures, graph problems, and mechanism design. An orthogonal theorem-level axis distinguishes achieved upper bounds from matched asymptotic dependence. Formal guarantees are separated from empirical systems evidence, with explicit treatment of prediction cost, feedback, and composition. The resulting synthesis states sufficient conditions for limited end-to-end reasoning and delineates open problems in cost-aware prediction, endogenous error, semantic predictors, and benchmarking.
Chinese Translation
学习增强算法在使用可能出错的预测的同时,仍能保持形式化的性能保证。本综述综合分析了预测接口、误差度量、一致性—鲁棒性权衡,以及在线优化、缓存、学习型数据结构、图问题和机制设计领域中的五种代表性构造机制。在定理层面的一个独立维度上,本综述区分了所实现的上界与匹配的渐近依赖关系。文中将形式化保证与实证系统证据相分离,并对预测成本、反馈与组合问题进行了显式处理。由此形成的综述给出了有限端到端推理的充分条件,并勾勒出成本感知预测、内生误差、语义预测器和基准测试等领域的开放性问题。
cs.LG / 42 / 2609.04797

How Faithful Is Attribution for Sales Forecasting? A Counterfactual Study

销售预测归因的忠实性如何?一项反事实研究
Kechyn, Glib
Abstract
Deep models for sales forecasting, such as WaveNet-style dilated convolutional networks, are accurate but opaque: when a single model predicts sales for one of many series, it offers no account of why. We add a post-hoc, architecture-agnostic counterfactual interpretability layer to a multi-series WaveNet forecaster trained on the full Corporacion Favorita grocery dataset (174,685 series over 1,688 days). The method decomposes each forecast into contributions that sum exactly to the predicted value, avoiding the allocation artifacts we observed with additive SHAP-style attribution. We evaluate faithfulness with a deletion/insertion protocol and find a statistically significant effect on both tests (deletion gap 0.22, p<0.001; insertion gap 0.27, p<0.01; robust across five background-sampling seeds), establishing that the attributions reflect genuine model behavior rather than plausible-looking artifacts. We then characterize, honestly, where attribution is and is not informative: reliance on the promotion signal is heterogeneous across series (median ratio approximately 1.0, with roughly 20% of series showing a strong effect), and the model captures the shape of the weekly sales cycle (day-of-week r=0.78) while systematically under-predicting its amplitude. Our contribution is not improved accuracy but an interpretability layer with a rigorous faithfulness evaluation and a candid account of its limits.
Chinese Translation
用于销售预测的深度模型,如WaveNet风格的空洞卷积网络,虽然预测准确但缺乏可解释性:当单一模型为众多序列中的一个预测销量时,它无法说明预测的原因。我们在基于完整的Corporacion Favorita食品杂货数据集(1,688天内共174,685个序列)训练的多序列WaveNet预测器上,添加了一个事后(post-hoc)的、与架构无关的反事实可解释性层。该方法将每个预测分解为若干贡献项,这些贡献项之和恰好等于预测值,从而避免了我们在使用可加性SHAP风格归因时观察到的分配偏差。我们通过删除/插入(deletion/insertion)协议评估其忠实性,发现两项检验均具有统计显著性(删除差距0.22,p<0.001;插入差距0.27,p<0.01;在五个背景采样种子下均稳健),这表明归因结果反映的是真实的模型行为,而非看似合理的人为产物。随后,我们诚实地刻画了归因在哪些方面有信息价值、在哪些方面没有:模型对促销信号的依赖在不同序列间存在异质性(中位数比率约为1.0,约20%的序列表现出强效应),并且模型能够捕捉每周销售周期的形状(星期效应r=0.78),但系统性低估其振幅。我们的贡献不在于提升预测精度,而在于提供了一个经过严格忠实性评估的可解释性层,并对其局限性给出了坦诚的说明。
cs.LG / 43 / 2609.04815

Federated Attack Campaign Detection via Contrastive Encoding of Threat Indicators in Gradient Updates

基于梯度更新中威胁指标对比编码的联邦攻击活动检测
Röder, Manuel, Babu, Bibin, Schleif, Frank-Michael
Abstract
Detecting orchestrated cyberattack campaigns that span multiple organizations traditionally requires sharing sensitive telemetry and threat intelligence across institutional boundaries and country borders, a barrier that Federated Learning removes by training shared threat detectors directly on local data. We propose FedIoC, a modular framework in which clients fold locally available structured threat indicators into their gradient updates; we instantiate the client-side encoder with a supervised contrastive loss over IoC-matched flows. Within each training batch, flows that match any known indicator pattern form the positive set; the contrastive objective pulls their learned embeddings together and pushes non-IoC embeddings away, so that campaign-relevant structure is, by design, expressed in the gradient direction. Clients sharing indicators for the same attack campaign then produce aligned gradient components, which the server clusters by the cosine similarity of their updates to recover global campaign patterns without any direct IoC transmission. We evaluate FedIoC on two public threat-detection benchmarks distributed across FL clients that each observe only a fragment of every active campaign and hold disjoint indicator sets derived from their local telemetry. In this regime the FL server recovers cross-organizational campaign cohorts directly from gradient geometry. We contribute FedIoC as a modular framework for this setting, and use it to pinpoint the non-IID gradient structure as the main driver of recovery and to define the open problem of designing encoders that improve on it.
Chinese Translation
检测跨多个组织实施的有组织网络攻击活动,传统上需要在机构边界和国家边界之间共享敏感的遥测数据和威胁情报,而联邦学习(Federated Learning)通过直接在本地数据上训练共享威胁检测器,消除了这一障碍。我们提出了FedIoC,一个模块化框架,其中客户端将本地可用的结构化威胁指标融入其梯度更新中;我们在客户端编码器上采用基于IoC匹配流量的监督对比损失进行实例化。在每个训练批次内,匹配任何已知指标模式的流量构成正样本集;对比目标将它们学习到的嵌入向量拉近,并将非IoC嵌入推开,从而使攻击活动相关的结构在设计上就体现在梯度方向中。共享同一攻击活动指标的客户端随后会产生对齐的梯度分量,服务器通过其更新的余弦相似度对这些分量进行聚类,从而在不直接传输任何IoC的情况下恢复全局攻击活动模式。我们在两个公开的威胁检测基准上评估FedIoC,将其分布于多个联邦学习客户端,每个客户端仅观察到每个活跃攻击活动的一个片段,并持有从其本地遥测数据中导出的互不相交的指标集。在这种情形下,联邦学习服务器能够直接从梯度几何结构中恢复跨机构的攻击活动群体。我们将FedIoC作为一个面向该场景的模块化框架贡献出来,并利用它指出非IID梯度结构是恢复效果的主要驱动因素,同时定义了设计能在此基础上进一步改进的编码器这一开放性问题。
cs.LG / 44 / 2609.04830

Communication-Efficient Personalized Federated Learning via Layer-Wise Multi-Threshold Random Sketching

基于逐层多阈值随机素描的通信高效个性化联邦学习
Zhang, Xu, Hou, Xingyu, Cheng, Jiacheng, Feng, Kaiyuan, Gong, Maoguo
Abstract
Personalized federated learning (PFL) is a promising paradigm for collaborative learning over distributed devices, where edge nodes collaboratively train personalized models without sharing raw data. Although PFL addresses data heterogeneity by learning client-specific models, it still suffers from substantial uplink and downlink communication costs when exchanging high-dimensional parameters in bandwidth-constrained systems. Recent one-bit methods achieve extreme compression, but they usually rely on a single thresholding rule applied to the whole model. This design has two limitations. First, it overlooks layer-wise differences in parameter distributions and quantization sensitivities. Second, a single threshold provides only coarse binary information and cannot capture fine-grained variations in parameter distributions. To address these issues, we propose a communication-efficient PFL framework via layer-wise multi-threshold random sketching. In the proposed method, each layer is assigned its own set of quantization thresholds, so that the compressed representation can adapt to layer-specific statistics while using multiple intervals to provide a finer low-bit description of sketched parameters. The proposed method supports bidirectional communication using compact low-bit sketches and improves the communication-accuracy tradeoff compared with existing one-bit compression approaches.
Chinese Translation
个性化联邦学习(PFL)是一种在分布式设备上开展协作学习的有前景的范式,其中边缘节点在无需共享原始数据的情况下协同训练个性化模型。尽管PFL通过学习客户端特定的模型来解决数据异构性问题,但在带宽受限的系统中交换高维参数时,仍然存在巨大的上行和下行通信开销。近期的单比特(one-bit)方法实现了极致压缩,但它们通常依赖于应用于整个模型的单一阈值规则。这种设计存在两个局限:其一,它忽视了参数分布和量化敏感性在逐层之间的差异;其二,单一阈值只能提供粗糙的二值信息,无法刻画参数分布中的细粒度变化。为解决这些问题,我们提出了一种基于逐层多阈值随机素描(layer-wise multi-threshold random sketching)的通信高效PFL框架。在所提出的方法中,每一层被分配各自独立的量化阈值集合,使压缩表示能够适应各层特有的统计特性,同时利用多个区间为素描后的参数提供更精细的低比特描述。该方法支持基于紧凑低比特素描的双向通信,与现有的单比特压缩方法相比,改善了通信与精度之间的权衡。
cs.LG / 45 / 2609.04832

PACE: Propagation-Aware Collaborative Correction for One-Shot Personalized Federated Graph Learning

PACE:面向一次性个性化联邦图学习的传播感知协同校正
Huang, Ruizhe, Li, Chengran, Shi, Xiaochuan
Abstract
Client heterogeneity creates both an opportunity and a risk in personalized federated graph learning. Knowledge held by other subgraphs may complement a receiver's Local model, but an incompatible transfer can override reliable predictions. One-shot communication sharpens this tension because an unsuitable server return cannot be corrected later. We introduce PACE, which treats collaborative knowledge as a compact correction to a complete Local predictor rather than as its replacement. Each client uploads a rank-r update carrier and a diagonal sketch of propagated message moments. The server uses them to construct a propagation-aware, receiver-anchored correction, while the receiver retains its full Local model. Convex negative-log-likelihood calibration (CNLL) then selects one coefficient between Local and External logits using validation nodes; model parameters remain fixed and no feedback is sent. At Rank-6, personalized returns occupy 9.6-17.6% of dense tensor bytes across the six evaluated datasets. The correction receives nonzero weight and improves both Accuracy and weighted-F1 over Local on five datasets; on ogbn-arxiv, CNLL assigns zero predictive weight to the correction and preserves Local predictions exactly. Applying the same CNLL rule to matched baselines on three citation datasets does not account for these gains. The central result is therefore that a small transported correction can augment a complete Local model when receiver evidence supports it while leaving the Local prediction unchanged otherwise.
Chinese Translation
在个性化联邦图学习中,客户端异构性既带来机遇也带来风险。其他子图所持有的知识可能对接收方的本地模型形成补充,但不兼容的知识迁移也可能覆盖其可靠的预测。一次性通信加剧了这一矛盾,因为服务器不合适的返回结果无法在后续加以纠正。我们提出PACE,将协同知识视为对完整本地预测器的一种紧凑校正,而非对其的替代。每个客户端上传一个秩为r的更新载体以及传播消息矩的对角化草图。服务器利用这些信息构建一个传播感知的、以接收方为锚点的校正,同时接收方保留其完整的本地模型。随后,凸负对数似然校准(CNLL)利用验证节点在本地与外部logit之间选择一个系数;模型参数保持固定,且不发送任何反馈。在秩为6时,在所评估的六个数据集上,个性化返回结果仅占稠密张量字节数的9.6%-17.6%。该校正在五个数据集上获得了非零权重,并较本地模型提升了Accuracy和加权F1;在ogbn-arxiv上,CNLL为校正分配了零预测权重,从而完全保留了本地预测。在三个引文数据集上将相同的CNLL规则应用于匹配的基线方法,并不能解释这些增益。因此,核心结论是:当接收方证据支持时,一个小的可传输校止可以增强完整的本地模型;否则,本地预测保持不变。
cs.LG / 46 / 2609.04852

KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

KVMem:在消费级GPU上虚拟化百万Token规模的智能体工作区
Chai, Di, Wang, Leye, Su, Zeshen, Xia, Zhiguo, Yu, Zhihang
Abstract
Modern LLM agents operate in persistent workspaces whose accumulated history can exceed both GPU KV capacity and the model's native context window. Existing systems typically compact older context into summaries or retrieve it later as text, either losing fine-grained execution evidence or repeatedly prefilling content that the model has already processed. We present KVMem, a KV-context virtualization system that preserves overflowed workspace history as paged KV state across GPU memory, host memory, and NVMe. KVMem uses lightweight, model-native attention-space indexes to select relevant historical blocks and materializes a query-dependent execution view bounded by the model's native context window. Extensive evaluations on long-context agent benchmarks spanning histories up to one million tokens, including LongMemEval, MemoryAgentBench, and AgentLongBench, show that KVMem generally achieves higher task utility and greater inference efficiency than compaction-based approaches, the de facto standard for handling context overflow. In the DeepSWE long-context test with Qwen3.8-27B, KVMem improves task success from 43.8% with compaction-only context management to 48.4%. In our local-deployment evaluation, KVMem runs Qwen3.6/3.8-27B NVFP4 with MTP on an off-the-shelf laptop equipped with a 24\,GB RTX 5090 Laptop GPU, virtualizing agent workspaces of up to 1M tokens-four times the model's native 256K-token context window. In a single-session setting, KVMem generates $\sim$50 tokens/s, providing interactive responsiveness for local agent execution. More broadly, by decoupling addressable workspace size from the LLM's native context window, KVMem provides a practical path toward long-running agents whose workspaces can grow beyond that window.
Chinese Translation
现代大语言模型(LLM)智能体运行在持久化的工作区中,其累积历史可能同时超出GPU的KV缓存容量和模型的原生上下文窗口。现有系统通常将较旧的上下文压缩为摘要,或以文本形式在后续阶段检索,这要么丢失了细粒度的执行证据,要么对模型已经处理过的内容进行反复预填充。我们提出KVMem,一种KV上下文虚拟化系统,将溢出的工作区历史以分页KV状态的形式保存在GPU内存、主机内存和NVMe中。KVMem利用轻量级的、模型原生的注意力空间索引来选择相关的历史块,并构建一个受模型原生上下文窗口约束的、依赖查询的执行视图。在涵盖高达一百万Token历史的长上下文智能体基准测试上的广泛评估(包括LongMemEval、MemoryAgentBench和AgentLongBench)表明,KVMem总体上比基于压缩的方法(处理上下文溢出的事实标准)实现了更高的任务效用和更强的推理效率。在使用Qwen3.8-27B的DeepSWE长上下文测试中,KVMem将任务成功率从仅使用压缩式上下文管理的43.8%提升至48.4%。在本地部署评估中,KVMem在一台配备24 GB RTX 5090 Laptop GPU的现成笔记本电脑上运行带有MTP的Qwen3.6/3.8-27B NVFP4模型,虚拟化高达1M Token的智能体工作区,为模型原生256K Token上下文窗口的四倍。在单会话设置下,KVMem以约50 tokens/s的速度生成,为本地智能体执行提供了交互式响应能力。更广泛地说,通过将可寻址工作区大小与LLM的原生上下文窗口解耦,KVMem为工作区可以超越该窗口、长期运行的智能体提供了一条实用路径。
cs.LG / 47 / 2609.04861

When Genomic Masking Priors Fail to Transfer: Strong Variant Prediction, Weak Functional Generation

当基因组掩码先验无法迁移时:强变异预测,弱功能生成
Hu, Susu, Gattogi, Preetam, Lehmann, Jens, Vahdati, Sahar, Speidel, Stefanie, Vibert, Julien
Abstract
Bidirectional discrete diffusion model appears naturally suited to genomic modeling because it can reconstruct missing sequence from both flanks. We developed GenDA (Genomic Density-optimized Absorbing Diffusion) under the additional hypothesis that entropy-guided span placement would concentrate reconstruction pressure on compositionally complex regions, improving both downstream variant-effect prediction and functional sequence generation. Our results only partially support this premise. After supervised fine-tuning, the 202M-parameter GenDA model reaches a pooled ClinVar SNV AUROC of 0.774, exceeding a similarly scaled autoregressive model by 0.103. However, a matched random-span variant reaches 0.777, providing no evidence that entropy guidance causes the ClinVar improvement. More unexpectedly, GenDA fails a zero-shot functional inpainting stress test: across promoters, enhancers, exon boundaries, and intron boundaries, it does not consistently outperform a control that shuffles the native gap while exactly preserving 3-mer composition. Failure is already present for 50--500-bp gaps, although enhancer degradation worsens at longer gaps. Diagnostics identify several boundary conditions: entropy measures local sequence complexity rather than functional importance; 1-mer tokenization limits physical context; training spans are capped at 300 bp; and high absolute AlphaGenome fidelity can coexist with negative control-normalized restoration. These results show that strong fine-tuned variant prediction, a plausible corruption prior, and functional generation are distinct claims that require separate validation.
Chinese Translation
双向离散扩散模型天然适合基因组建模,因为它可以从两侧序列重构缺失部分。我们在一个额外的假设下开发了GenDA(Genomic Density-optimized Absorbing Diffusion,基因组密度优化的吸收态扩散模型),即熵引导的片段放置会将重构压力集中于序列组成复杂的区域,从而同时提升下游变异效应预测和功能序列生成。我们的结果仅部分支持这一前提。经过监督微调后,拥有2.02亿参数的GenDA模型在ClinVar单核苷酸变异(SNV)上的合并AUROC达到0.774,超出规模相近的自回归模型0.103。然而,一个匹配的随机片段变体达到了0.777,这表明没有证据显示熵引导是ClinVar性能提升的原因。更出乎意料的是,GenDA未能通过零样本功能修复(inpainting)压力测试:在启动子、增强子、外显子边界和内含子边界上,其表现并不稳定地优于一个在精确保持3-mer组成的同时打乱天然缺失片段的对照组。在50–500 bp的缺失片段上失败已经显现,不过增强子的退化在更长的缺失片段上更为严重。诊断分析揭示了若干边界条件:熵衡量的是局部序列复杂度而非功能重要性;单核苷酸(1-mer)分词限制了物理上下文;训练片段长度上限为300 bp;此外,高绝对AlphaGenome保真度可能与负的控制组归一化恢复效果并存。这些结果表明,强大的微调后变异预测、看似合理的损坏先验以及功能生成是三个不同的命题,需要分别进行验证。
cs.LG / 48 / 2609.04881

From Deep to Shallow: Unconstrained and Efficient Layer Merging Strategy

从深到浅:一种无约束且高效的层合并策略
Shulzhenko, Petro, Spadaro, Gabriele, Tartaglione, Enzo
Abstract
Although Deep Neural Networks have become foundational in many areas of Machine Learning, high computational demands limit their application in resource-constrained environments. To address this issue, depth compression methods have been proposed to identify and linearize redundant activation functions, thereby allowing for the merging of layers without intermediate non-linearities. However, these methods face two key challenges: they cannot be directly applied to convolutions with padding due to the absence of an analytical solution for merging these layers, and they typically increase the kernel size of merged layers, thus limiting speed-up gains. To overcome these limitations, we propose an efficient strategy that enables merging of layers without an existing analytical solution, and also without increasing kernel size. We validate our approach across multiple architectures and datasets, and measure inference speed-up gains on real embedded platforms. We publicly released the code at https://github.com/ShulzhenkoPetr/deep-to-shallow.
Chinese Translation
尽管深度神经网络已在机器学习的许多领域成为基础性技术,但其高昂的计算需求限制了其在资源受限环境中的应用。为解决这一问题,已有研究提出了深度压缩方法,通过识别并线性化冗余的激活函数,从而在去除中间非线性层的情况下实现层的合并。然而,这些方法面临两个关键挑战:由于缺乏对带填充卷积层合并的解析解,它们无法直接应用于带填充的卷积;此外,这些方法通常会增大合并后层的卷积核尺寸,从而限制了加速收益。为克服这些局限,我们提出了一种高效策略,能够在没有现成解析解的情况下实现层合并,同时不增加卷积核尺寸。我们在多种架构和数据集上验证了该方法,并在真实的嵌入式平台上测量了推理加速收益。代码已在 https://github.com/ShulzhenkoPetr/deep-to-shallow 公开发布。
cs.LG / 49 / 2609.04901

Adaptation Interfaces for In-Context Tabular Foundation Models in Time-to-Event Prediction

面向事件时间预测的上下文表格基础模型适配接口
Pham, Minh-Khoi, Cotugno, Luca, Cernei, Dan, Sirbu, Alina, Masi, Stefano, Prencipe, Giuseppe, Pingitore, Alessandro, Landi, Patrizia, Acid, Working Group on Uric, Hypertension, Cardiovascular Risk of the Italian Society of, Mai, Tai Tan, Crane, Martin, Bezbradica, Marija
Abstract
Tabular foundation models (TabFMs) achieve strong performance on structured data, particularly for standard classification and regression problems. Yet, extending them to censored time-to-event prediction is challenging because it requires properly handling censoring and event-time dynamics. Building on our prior work, we further link TabFMs with CoxPH and DeepHit and revise the context-resampled training procedure. We evaluate temporal zero-shot reformulation, classification-based fine-tuning, and survival-head adaptation using frozen TabFM backbones on 74 single-risk data sets, and we additionally study 4 competing-risk data sets. Zero-shot inference is effective on smaller single-risk data sets, whereas supervised adaptation becomes increasingly advantageous as data sets scale. Cox provides the most reliably strong interface, especially for Integrated Brier Score (IBS) on larger data sets. DeepHit is relatively stronger for the time-dependent Concordance Index than for IBS, while cause-specific MTLR ranks highest among the TabFM survival heads in the four-data-set competing-risk analysis. Classification fine-tuning becomes more competitive with zero-shot inference as data sets grow but remains weaker for probabilistic prediction. Overall, our results indicate that effective TabFM transfer depends on the data regime and on the statistical structure represented by the chosen adaptation interface. The implementation scripts used for this work are available at https://github.com/kaylode/survival-fm.
Chinese Translation
表格基础模型(Tabular Foundation Models, TabFMs)在结构化数据上取得了优异的性能,尤其是在标准分类和回归问题上。然而,将其扩展到含删失的事件时间预测颇具挑战性,因为这需要妥善处理删失和事件时间动态。在我们先前工作的基础上,我们进一步将 TabFMs 与 CoxPH 和 DeepHit 相关联,并修订了上下文重采样训练流程。我们在 74 个单风险数据集上评估了时间零样本重构、基于分类的微调以及使用冻结 TabFM 主干的生存头适配方法,并额外研究了 4 个竞争风险数据集。零样本推理在较小的单风险数据集上表现有效,而随着数据集规模的增大,有监督适配的优势日益明显。Cox 提供了最可靠且稳健的接口,尤其体现在较大数据集上的综合 Brier 分数(Integrated Brier Score, IBS)上。DeepHit 在时间依存一致性指数(Concordance Index)上相对优于 IBS,而在四数据集竞争风险分析中,特定原因 MTLR 在各 TabFM 生存头中排名最高。分类微调随数据集增长而逐渐具备与零样本推理竞争的能力,但在概率预测方面仍然较弱。总体而言,我们的结果表明,有效的 TabFM 迁移取决于数据规模情形以及所选适配接口所表达的统计结构。本工作所用的实现脚本可在 https://github.com/kaylode/survival-fm 获取。
cs.LG / 50 / 2609.04910

Fast Gauss Sums via Flash Attention

基于Flash Attention的快速高斯和计算
Rux, Nicolaj, Neumayer, Sebastian
Abstract
Gaussian kernel sums are the computational core of maximum mean discrepancies (MMDs), kernel gradient flows, Stein variational gradient descent (SVGD), and many other kernel methods. At the same time, softmax attention has received an extraordinary amount of hardware-aware code engineering, culminating in flash attention. We show that Gauss kernel sums with arbitrary, signed weights can be evaluated via flash attention: two small input augmentations turn the normalized softmax reduction into the unnormalized Gauss sum, without writing a single line of custom GPU code. For feature dimension D>8 in fp16, this approach beats compiled PyTorch code as well as PyKeOps kernels (often significantly) in speed, memory-overhead and accuracy. Indeed, its memory scaling remains linear.
Chinese Translation
高斯核求和是最大均值差异(MMD)、核梯度流、Stein变分梯度下降(SVGD)以及许多其他核方法的计算核心。与此同时,softmax注意力机制已经获得了大量的硬件感知代码工程优化,其集大成者便是flash attention。我们证明了带有任意符号权重的高斯核求和可以通过flash attention来计算:仅需两个小的输入增广,即可将归一化的softmax归约转化为非归一化的高斯和,而无需编写任何一行定制的GPU代码。在fp16精度下,当特征维度D>8时,该方法在速度、内存开销和精度方面均优于编译后的PyTorch代码以及PyKeOps核(通常优势显著)。事实上,其内存扩展仍保持线性。
cs.LG / 51 / 2609.04943

Physics-Aware Random Walk Fingerprints for Scalable Power Grid Graph Classification

面向可扩展电网图分类的物理感知随机游走指纹方法
Anwar, Adnan
Abstract
Recent benchmarks such as PowerGraph provide large collections of power-grid graphs for cascading-failure classification. Graph neural networks (GNNs) achieve strong predictive performance on this task, but typically require end-to-end training and model-specific tuning, while their latent representations can be difficult to relate to physically meaningful propagation patterns. Random Walk Fingerprints (RWF) offer a scalable and interpretable alternative, but existing variants primarily emphasise topology and node-level information, leaving grid-relevant operational edge states in the walk dynamics. We propose Multi-Channel Physics-Aware Random Walk Fingerprints (MC-PA-RWF) for power systems, a lightweight graph-level representation framework that introduces physical edge states into random-walk propagation. The method constructs multiple edge-weighted channels from domain-relevant attributes, extracts a channel-specific fingerprint from each weighted graph, and concatenates the resulting vectors into a compact representation. Experiments on three \textit{PowerGraph} benchmark systems show substantial improvements over topology-only RWF and competitive balanced accuracy against strong GNN baselines, including Graph Convolutional Networks (GCN), Graph Attention Networks (GAT), Graph Isomorphism Networks with edge features (GINE), and Transformer-based Graph Convolutional Networks (TransformerConv). At the largest evaluated settings, the node-edge extension MC-PA-RWF+ achieves around 98.04% - 99.32% balanced accuracy and improves failure-class F1 over the strongest GNN baseline by 1.60 -- 5.84 percentage points, with statistically significant gains across all three systems.
Chinese Translation
诸如 PowerGraph 等近期的基准测试提供了大量用于级联故障分类的电网图数据。图神经网络(GNN)在该任务上取得了较强的预测性能,但通常需要进行端到端训练和针对特定模型的调优,且其潜在表示难以与具有物理意义的传播模式相关联。随机游走指纹(Random Walk Fingerprints, RWF)提供了一种可扩展且可解释的替代方案,但现有变体主要侧重于拓扑和节点级信息,未将电网相关的运行边状态纳入游走动力学之中。我们提出了一种面向电力系统的多通道物理感知随机游走指纹方法(Multi-Channel Physics-Aware Random Walk Fingerprints, MC-PA-RWF),这是一种轻量级的图级表示框架,将物理边状态引入随机游走传播过程。该方法基于领域相关属性构建多个边加权通道,从每个加权图中提取通道特定的指纹,并将所得向量拼接为一个紧凑的表示。在三个 PowerGraph 基准系统上的实验表明,该方法相较仅使用拓扑信息的 RWF 有显著提升,并且在与强大的 GNN 基线(包括图卷积网络(GCN)、图注意力网络(GAT)、带边特征的图同构网络(GINE)以及基于 Transformer 的图卷积网络(TransformerConv))对比时取得了具有竞争力的平衡准确率。在所评估的最大规模设置下,节点-边扩展版本 MC-PA-RWF+ 实现了约 98.04%–99.32% 的平衡准确率,相比最强的 GNN 基线,其故障类别 F1 分数提升了 1.60 至 5.84 个百分点,且在所有三个系统上均取得了具有统计显著性的改进。
cs.LG / 52 / 2609.04963

Fractal basins trap latent reasoning

分形盆捕获潜在推理
Lai, Jeffrey, Bao, Anthony, Quinn, John, Gilpin, William
Abstract
Reasoning allows artificial intelligence models to revisit and correct their mistakes, enabling recent frontier advances in mathematical theorem solving, software engineering, and autonomous task planning. Reasoning models are widely observed to reason for longer on harder tasks, but the general mechanism responsible for these slowdowns is unknown. Here, we show that reasoning models exhibit transient chaos, a physical consequence of the computational complexity of difficult tasks. As a consequence, we show that diverse leading reasoning models are dynamical systems with fractal basins, with fractality increasing with task difficulty across diverse tasks like Sudoku and maze solving, visual puzzles, and mathematical logic. We show that transient chaos emerges due to reasoning becoming trapped for extended durations near saddle points, which we show correspond to nearly-correct attempted solutions of the underlying problem. Our results show that reasoning slowdowns are an inevitable consequence of problem hardness in modern artificial intelligence models, and establish reasoning traces as a rich new class of dynamical system.
Chinese Translation
推理使人工智能模型能够回顾并纠正自身的错误,从而推动了近期在数学定理求解、软件工程和自主任务规划等前沿领域的进展。人们普遍观察到,推理模型在更难的任务上会进行更长时间的推理,但导致这种变慢现象的一般机制尚不清楚。本文表明,推理模型表现出暂态混沌(transient chaos),这是困难任务计算复杂性所带来的物理后果。由此,我们证明多种领先的推理模型是具有分形盆(fractal basins)的动力学系统,且在数独、迷宫求解、视觉谜题和数理逻辑等多种任务中,分形性随任务难度的增加而增强。我们发现,暂态混沌的产生源于推理长时间陷入鞍点附近,而这些鞍点对应于对底层问题近乎正确的尝试解。我们的结果表明,推理变慢是现代人工智能模型中问题难度的必然结果,并将推理轨迹确立为一类丰富的新型动力学系统。
cs.LG / 53 / 2609.04971

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

BeaconKV:由信标查询引导的键值缓存压缩方法,用于高效的大推理模型推理
Kim, Janghyeon, Kim, Minsoo, Shim, Kyuhong, Choi, Jungwook
Abstract
Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, often exceeding GPU capacity for long reasoning traces. Existing KV cache compression methods rely on recent queries to estimate future token importance, implicitly assuming these serve as reliable proxies for future attention patterns. We demonstrate that this assumption fails in long-horizon reasoning: certain decoding steps generate Thought Revisiting Tokens (TRT) that re-attend to distant previous context, such as task-solving plans formulated early in the trace. Through systematic analysis, we discover that queries corresponding to the TRT cluster into a small number of similarity groups in the embedding space. Based on this insight, we propose BeaconKV, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history. Across four open-source LRMs and diverse reasoning benchmarks, BeaconKV generally outperforms existing compression methods, achieving up to $5.8\times$ memory reduction while nearly preserving full cache accuracy and improving throughput by over $4.3\times$.
Chinese Translation
大推理模型(Large Reasoning Models, LRMs)通过扩展的思维链(Chain-of-Thought, CoT)生成实现了卓越的问题求解能力,但由此产生的键值(KV)缓存随序列长度线性增长,造成严重的内存瓶颈,对于长推理轨迹往往超出GPU的容量。现有的KV缓存压缩方法依赖最近的查询来估计未来token的重要性,隐含地假设这些查询可作为未来注意力模式的可靠代理。我们证明该假设在长程推理中并不成立:某些解码步骤会生成思维回溯token(Thought Revisiting Tokens, TRT),重新关注远处的先前上下文,例如在轨迹早期形成的解题规划。通过系统性分析,我们发现与TRT对应的查询在嵌入空间中聚集成少数几个相似性组。基于这一洞察,我们提出了BeaconKV,一种无需训练的KV缓存压缩方法,该方法维护信标查询(beacon queries)——每个全局查询聚类的紧凑代表——以预测哪些KV对将被重新访问,而无需存储完整的查询历史。在四个开源LRM和多样推理基准上的实验表明,BeaconKV总体上优于现有压缩方法,在几乎保持完整缓存精度的同时,实现高达5.8倍的内存缩减,并将吞吐量提升超过4.3倍。
cs.LG / 54 / 2609.04995

Beyond Homoscedasticity: Decoupled Uncertainty Optimization for Deep Imbalanced Regression

超越同方差性:面向深度不平衡回归的解耦不确定性优化
Zhou, Juncheng, Lu, Jiaxi, Zeng, Weijing, Li, Zhong, Qi, Hao, Cui, Jingsong
Abstract
Deep Imbalanced Regression (DIR) is pervasive in continuous prediction tasks across diverse modalities, such as age estimation, depth prediction, and protein mutation activity prediction, where label-scarce tail samples often carry higher practical value. However, most existing methods still learn deterministic point mappings under mean squared error or its simple variants, implicitly assuming a uniform uncertainty level across all samples and thereby overlooking the instance-wise heteroscedasticity that is widespread in long-tailed data. We further point out that even heteroscedastic negative log-likelihood suffers from a gradient coupling issue, which, under DIR scenarios, weakens the learning signal of hard tail samples and leads to optimization inertia as well as tail underfitting. To address this, we propose DUO, an uncertainty-aware long-tailed regression framework. Specifically, the proposed method models the regression target as a conditional Gaussian distribution to explicitly characterize instance-level predictive uncertainty, and transforms uncertainty into a dynamic enhancement signal for tail samples through decoupled mean-variance optimization. Furthermore, we design a distribution-guided contrastive learning mechanism that adaptively constructs positive and negative pairs based on the overlap between sample distributions, thereby alleviating feature looseness and cross-label semantic entanglement. Across visual and biological DIR benchmarks, DUO achieves the best few-shot bMAE and GM on IMDB-WIKI-DIR, AgeDB-DIR, and AAV2-DIR while remaining competitive on few-shot MAE.
Chinese Translation
深度不平衡回归(Deep Imbalanced Regression, DIR)普遍存在于跨多种模态的连续预测任务中,如年龄估计、深度预测和蛋白质突变活性预测,其中标签稀缺的尾部样本往往具有更高的实际价值。然而,大多数现有方法仍在均方误差或其简单变体下学习确定性的点映射,隐式地假设所有样本具有统一的不确定性水平,从而忽略了长尾数据中普遍存在的样本级异方差性。我们进一步指出,即使是异方差负对数似然也存在梯度耦合问题,该问题在DIR场景下会削弱困难尾部样本的学习信号,导致优化惯性和尾部欠拟合。为解决这一问题,我们提出了DUO,一个不确定性感知的长尾回归框架。具体而言,该方法将回归目标建模为条件高斯分布,以显式刻画样本级的预测不确定性,并通过解耦的均值-方差优化将不确定性转化为尾部样本的动态增强信号。此外,我们设计了一种分布引导的对比学习机制,基于样本分布之间的重叠程度自适应地构建正负样本对,从而缓解特征松散和跨标签语义纠缠问题。在视觉和生物DIR基准测试中,DUO在IMDB-WIKI-DIR、AgeDB-DIR和AAV2-DIR上取得了最佳的少样本bMAE和GM,同时在少样本MAE上保持竞争力。
cs.LG / 55 / 2609.05012

Solution-space heterogeneity shapes federated learning dynamics across partial differential equations

解空间异质性塑造跨偏微分方程的联邦学习动力学
Luo, Ping, Wang, Jiahuan, Wen, Ziqing, Sun, Tao, Li, Dongsheng
Abstract
Federated scientific machine learning enables institutions to train neural surrogates without centralizing local physical data, yet studies of partial differential equations (PDEs) lack a transferable definition of non-independent and identically distributed data. Existing protocols partition coordinates, coefficients, boundary conditions, or geometries according to equation-specific rules. Here, we introduce solution-space PDE-Dirichlet, a protocol that converts continuous supervised responses into reusable solution bins and quantifies the realized separation between clients through optimal transport over the geometry of these bins. We derive an exact inverse relation between population allocation heterogeneity and the Dirichlet concentration, and we establish conditions under which response heterogeneity induces gradient disagreement, local-update dispersion, and parameter divergence. Across seven controlled and public PDE tasks, three neural-operator families, and five random seeds, a lower concentration consistently increases the realized solution distance and optimization heterogeneity. The degradation in final error is task dependent: the largest effect occurs for low-viscosity Burgers, reaching 4.157 percentage points under the most heterogeneous setting, whereas additional communication or smoother dynamics can reduce the final gap despite persistent parameter separation. These results distinguish a reproducible geometric mechanism from task-dependent generalization outcomes and provide a common basis for evaluating non-IID federated PDE learning.
Chinese Translation
联邦科学机器学习使各机构能够在不集中本地物理数据的情况下训练神经代理模型,然而针对偏微分方程(PDE)的研究缺乏对非独立同分布数据的可迁移定义。现有协议依据特定方程的规则对坐标、系数、边界条件或几何形状进行划分。本文提出了解空间 PDE-Dirichlet 协议,该协议将连续的监督响应转化为可复用的解分箱,并通过对这些分箱的几何结构进行最优传输来量化客户端之间实际的分离程度。我们推导出总体分配异质性与 Dirichlet 浓度之间的精确反比关系,并建立了响应异质性引发梯度分歧、局部更新离散性和参数发散的条件。在七个受控及公开的 PDE 任务、三类神经算子家族和五个随机种子上的实验表明,较低的浓度会持续增大实际解距离和优化异质性。最终误差的退化程度依赖于具体任务:最大的影响出现在低粘性 Burgers 方程上,在最为异质的设置下达到 4.157 个百分点;而尽管参数分离持续存在,额外的通信或更平滑的动力学仍可减小最终差距。这些结果将可复现的几何机制与依赖于任务的泛化结果区分开来,为评估非独立同分布(non-IID)联邦 PDE 学习提供了共同基础。
cs.LG / 56 / 2609.05016

Amortizing Scaling Law Construction Costs

摊销缩放定律的构建成本
Jha, Abhash Kumar, Onuţu, Diana Alexandra, Mallik, Neeratyoy, Haldar, Swagatam, Laing, Sam, Ajroldi, Niccolò, Liu, Shiwei, Vanschoren, Joaquin, Klein, Aaron
Abstract
Scaling laws guide the design choices for training large foundation models, but deriving them involves training an exhaustive grid over hyperparameters, token budgets, and parameter counts, which is computationally expensive. Fitting a scaling law, however, only requires the best-loss frontier across compute scales, discarding most of the trained configurations. We propose a framework for efficient scaling law construction that formulates data collection as a Bayesian optimization problem, and introduce metrics for comparing scaling law fitting methods under constrained compute budgets. We find that progressively expanding the compute budget during acquisition, mirroring the compute-ordered evaluation of configurations in practice, substantially improves recovery efficiency. Augmenting the observed configurations with surrogate-fantasized evaluations then recovers the broader experimental grid, allowing accurate scaling law fitting without training every configuration. Together, these can closely match scaling law fits over a full dense grid at computational savings of up to $10\text{--}100\times$.
Chinese Translation
缩放定律指导大型基础模型的训练设计选择,但推导缩放定律需要在超参数、词元预算和参数数量上进行穷举式的网格训练,计算代价高昂。然而,拟合缩放定律只需要不同计算规模下的最优损失前沿,因此大多数被训练的配置都被舍弃了。我们提出了一个高效构建缩放定律的框架,将数据收集表述为一个贝叶斯优化问题,并引入了在受限计算预算下比较缩放定律拟合方法的评估指标。我们发现,在采集过程中逐步扩展计算预算——这与实践中按计算量排序评估配置的方式相呼应——能够显著提升恢复效率。随后,用代理模型幻想评估来扩充已观测的配置,即可恢复更广泛的实验网格,从而无需训练每一个配置就能准确拟合缩放定律。二者结合,可以在计算节省高达 $10\text{--}100$ 倍的情况下,接近完全密集网格上的缩放定律拟合效果。
cs.LG / 57 / 2609.05073

Confounding-Valid Conformal Inference for Counterfactual KPIs in Wireless Networks

无线网络中反事实关键绩效指标的混杂有效共形推断
Qchohi, Abdessamed, Cortes, Jessica Moysen, Zecchin, Matteo
Abstract
Conformal counterfactual inference enables network operators to use logged telemetry to reliably answer 'what-if' questions about network operation. These answers typically take the form of prediction sets that contain, with a user-defined probability, the key performance indicators (KPIs) that would have been observed under alternative control actions. A key challenge is that logged telemetry may omit variables used by the controller, resulting in hidden confounding and invalidating the statistical guarantees of counterfactual analysis. In principle, this issue can be addressed using randomized telemetry, collected by assigning control actions independently of the network state. However, because such randomization may disrupt normal operation, randomized telemetry is typically scarce, causing counterfactual analysis based solely on it to produce uninformative prediction sets. To address these challenges, we propose Confounding-Valid Counterfactual Conformal Inference (CV-CCI), which combines abundant, potentially confounded observational telemetry with limited randomized data through the General Synthetic-Powered Inference (GESPI) principle. CV-CCI leverages observational data to improve efficiency while using randomized data to retain finite-sample coverage guarantees under arbitrary hidden confounding. Experiments on two representative radio access network (RAN) control tasks show that CV-CCI remains valid under hidden confounding while producing more efficient prediction sets than state-of-the-art confounding-valid baselines.
Chinese Translation
共形反事实推断使网络运营商能够利用记录的遥测数据可靠地回答关于网络运行的"假设分析"问题。这些答案通常以预测集的形式呈现,预测集以用户自定义的概率包含在替代控制动作下本可观测到的关键绩效指标(KPI)。一个关键挑战在于,记录的遥测数据可能遗漏控制器所使用的变量,从而导致隐藏混杂,并使反事实分析的统计保证失效。原则上,该问题可通过随机化遥测来解决,即通过使控制动作独立于网络状态来采集数据。然而,由于这种随机化可能干扰正常运行,随机化遥测数据通常十分稀缺,导致仅基于随机化数据的反事实分析产生信息量不足的预测集。为应对这些挑战,我们提出了混杂有效反事实共形推断(Confounding-Valid Counterfactual Conformal Inference,CV-CCI),该方法通过广义合成数据驱动推断(General Synthetic-Powered Inference,GESPI)原则,将丰富的、可能存在混杂的观测遥测数据与有限的随机化数据相结合。CV-CCI 利用观测数据提升效率,同时借助随机化数据在任意隐藏混杂下保持有限样本覆盖率保证。在两个具有代表性的无线接入网(RAN)控制任务上的实验表明,CV-CCI 在隐藏混杂下依然保持有效,并且相比最先进的混杂有效基线方法能产生更高效的预测集。
cs.LG / 58 / 2609.05081

Deep Microcompression: Structured Pruning and Bit-packed Quantization for Microcontrollers

Deep Microcompression:面向微控制器的结构化剪枝与位打包量化
Busoye, Opegbemi Matthias, Busoye, Tolulope Matthew, Eigbe, Eghonghon-aye
Abstract
This paper introduces Deep Microcompression (DMC), a hardware-aware pipeline for deep learning inference on bare-metal microcontrollers. DMC integrates structured pruning, quantization-aware training, and fixed-length bit-packing to achieve a 55.8$\times$ weight compression ratio on LeNet-5 (98.77\% accuracy), generating a dependency-free C library with deterministic latency. On the RP2040 (Cortex-M0+), DMC reduces binary size by 3$\times$ versus TensorFlow Lite while matching its accuracy. Critically, DMC enables the first documented deployment of a standard CNN on the ATmega328P, a device constrained to 2KB SRAM, previously considered infeasible for CNN inference.
Chinese Translation
本文提出了Deep Microcompression(DMC),一个面向裸机微控制器深度学习推理的硬件感知流水线。DMC集成了结构化剪枝、量化感知训练和定长位打包技术,在LeNet-5上实现了55.8倍的权重压缩比(准确率98.77%),并生成一个无依赖、具有确定性延迟的C语言库。在RP2040(Cortex-M0+)上,与TensorFlow Lite相比,DMC在保持相同准确率的同时将二进制体积缩减了3倍。尤为重要的是,DMC首次实现了标准卷积神经网络(CNN)在ATmega328P上的部署——该设备仅有2KB SRAM,此前被认为是无法运行CNN推理的。
cs.LG / 59 / 2609.05097

NEAT-POCKET: Pocket-Conditioned Autoregressive 3D Molecular Generation with a Neighborhood-Guided Set Transformer

NEAT-POCKET:基于邻域引导集合变换器的结合口袋条件自回归3D分子生成方法
Jacob, Roxane Axel, Rose, Daniel, Langer, Thierry, Kirchmair, Johannes
Abstract
AI-driven de novo molecular design offers a promising route to accelerate early-stage drug discovery by generating novel ligands directly within target protein binding pockets. We present NEAT-POCKET, a pocket-conditioned extension of the autoregressive NEAT model for 3D molecular generation. NEAT-POCKET generates molecules atom by atom in protein pocket environments while preserving atom permutation invariance and explicitly modeling hydrogen atoms. Benchmarks on the CrossDocked and SPINDR datasets show that NEAT-POCKET achieves competitive structure-based generation performance while sampling substantially faster than existing baselines. Beyond full-molecule generation, NEAT-POCKET naturally enables pocket-conditioned fragment completion, a task directly relevant to lead optimization and scaffold elaboration. These results position NEAT-POCKET as a fast, flexible, and practical framework for structure-based drug design.
Chinese Translation
AI驱动的全新分子设计为加速早期药物发现提供了一条有前景的途径,其方法是在靶蛋白结合口袋内直接生成新颖配体。我们提出了NEAT-POCKET,这是自回归模型NEAT面向3D分子生成的口袋条件扩展版本。NEAT-POCKET能够在蛋白口袋环境中逐原子生成分子,同时保持原子排列不变性并显式地对氢原子进行建模。在CrossDocked和SPINDR数据集上的基准测试表明,NEAT-POCKET实现了具有竞争力的基于结构的分子生成性能,同时采样速度显著快于现有基线方法。除完整分子生成外,NEAT-POCKET还可自然地实现口袋条件下的片段补全任务,该任务与先导化合物优化和骨架衍生直接相关。这些结果表明,NEAT-POCKET是一个快速、灵活且实用的基于结构的药物设计框架。
cs.LG / 60 / 2609.05113

A Comparative Study of Counterfactual Explainers for Graph Neural Networks Enabling Multiple Types of Graph Edit

支持多种图编辑类型的图神经网络反事实解释器比较研究
Villia, Maria Myrto, Gouidis, Filippos, Patkos, Theodore, Trahanias, Panos
Abstract
Counterfactual explanations for graph-structured data seek to determine minimal and realistic modifications required in an input graph to alter a model's prediction to a predefined output. Although counterfactual explainers that support modifying the graph by both adding and removing edges have recently emerged, there is still a lack of general and efficient methods, especially when considering the quality of the generated explanations. Moreover, the problem remains far from solved, as existing methods exhibit different strengths and weaknesses, often trading off between explanation size, coverage and quality. For this reason, it is important to identify where each method performs well and where it falls short, so as to guide future research in the field. Thus, our study compares six state-of-the-art (SOTA) models on a diverse set of real-world and synthetic datasets, covering both binary and multi-class graph and node classification tasks, and evaluates their performance using diverse quantitative and qualitative metrics.
Chinese Translation
面向图结构数据的反事实解释旨在确定为使模型的预测改变为预定义输出而在输入图中所需的最小且符合实际(真实合理)的修改。尽管最近出现了同时支持通过添加和删除边来修改图的反事实解释器,但仍然缺乏通用且高效的方法,尤其是在考虑所生成解释的质量时。此外,该问题远未得到解决,因为现有方法表现出各自不同的优势与不足,往往在解释大小、覆盖率和质量之间进行权衡。因此,识别每种方法在哪些方面表现良好、在哪些方面存在不足,对于指导该领域的未来研究十分重要。为此,我们的研究在多样化的真实世界和合成数据集上比较了六种最先进(SOTA)的模型,涵盖二分类和多分类的图分类与节点分类任务,并使用多种定量和定性指标评估它们的性能。
cs.LG / 61 / 2609.05125

Single-Query Black-Box Calibration Auditing via Logit Bias

基于Logit Bias的单次查询黑盒校准审计
Plaud, Roman, Saillenfest, Antoine, Labeau, Matthieu, Bonald, Thomas, Waegeman, Willem
Abstract
Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot classifiers. Yet, commercial API providers increasingly hide the continuous output probabilities required by standard calibration metrics. To bypass this opacity, we demonstrate that any LLM API exposing a logit\_bias parameter can be mathematically manipulated to evaluate exact probability thresholds using strictly one query per sample. Leveraging this mechanism, we introduce a novel and provably consistent estimator of the True Calibration Error for binary tasks. Our approach therefore provides an efficient framework for auditing black-box foundation models.
Chinese Translation
评估大语言模型(LLM)的校准对于其作为零样本分类器的安全部署至关重要。然而,商业API提供商越来越多地隐藏标准校准指标所需的连续输出概率。为了绕过这种不透明性,我们证明任何暴露logit_bias参数的LLM API都可以通过数学操作,仅使用每个样本严格一次查询来评估精确的概率阈值。利用这一机制,我们提出了一种新颖且可证明一致的二元任务真实校准误差(True Calibration Error)估计器。因此,我们的方法为审计黑盒基础模型提供了一个高效的框架。
cs.LG / 62 / 2609.05126

Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy

隐层表示的粗粒化:基于映射熵的无监督神经元选择
Mele, Margherita, Castagna, Andrea, Menichetti, Roberto, Potestio, Raffaello, Ingrosso, Alessandro
Abstract
Overparameterized neural networks carry far more hidden units than a task nominally requires, raising the question of which neurons are essential and whether that distinction is legible in the representation itself, without labels or gradients. We cast neuron selection as the problem of coarse-graining the hidden layer by retaining a subset of its neurons, and score each putative selection by the mapping entropy (ME). This quantity measures the loss of discriminatory power inherent in discarding part of the network neurons, and the selection that minimises the ME is taken as particularly informative. This criterion is fully unsupervised, in that it depends only on hidden-activation statistics. In teacher-student networks, ME optimisation recovers the minimal teacher-consistent representation and retains extra units in proportion to the hidden layer's residual variability; in a non-linear Gaussian process task, it selects coherent functional-class mappings whose preferred class shifts across training. On this task and on translation-augmented MNIST, ME-selected subnetworks outperform random subsets of equal size, most clearly under strong compression - linking configurational distinguishability to predictive performance.
Chinese Translation
过参数化的神经网络所包含的隐藏单元远多于任务名义上所需的数量,这就引出了一个问题:哪些神经元是必不可少的,以及这种区分是否可以在不依赖标签或梯度的情况下,直接从表示本身中辨识出来。我们将神经元选择问题转化为通过对隐层神经元保留一个子集来实现隐层粗粒化的问题,并用映射熵(mapping entropy, ME)对每种候选选择进行评分。该量度衡量了丢弃部分网络神经元时所固有的判别能力的损失,使映射熵最小化的选择被认为是信息量特别丰富的。这一准则完全是无监督的,因为它仅依赖于隐层激活的统计特性。在教师-学生网络中,对映射熵的优化能够恢复与教师一致的最小表示,并按隐层残余变异性的比例保留多余的单元;在一个非线性高斯过程任务中,它能够选择出其偏好类别随训练过程发生漂移的、内部一致的函数类映射。在该任务以及经过平移增强的MNIST数据集上,由映射熵选出的子网络优于同等规模的随机子集,且在强压缩情形下优势最为明显——这将构型可区分性与预测性能联系了起来。
cs.LG / 63 / 2609.05136

MomentQuant: an even more minimalist interval method with linear time complexity for time series classification

MomentQuant:一种时间复杂度为线性的更极简区间方法,用于时间序列分类
Faouzi, Johann
Abstract
Time series data is very common in many real-world applications and in numerous domains, with increasing interest for automated information extraction using machine learning. One of these subfields is time series classification, which consists in assigning a label to each new, unseen time series. Many algorithms have been developed over the past decades, with the trade-off between predictive performance and computational cost being consistently discussed. Quant, an interval-based algorithm extracting quantiles from recursive, fixed, dyadic intervals, was shown to achieve high accuracy, while being very fast. We propose two changes to make this algorithm even faster. The first one is a better optimized implementation of the exact same algorithm. The second one is to derive approximate quantiles, using the Cornish-Fisher expansion, instead of exact quantiles. This change removes the necessity to sort the time series, leading to a smaller computational complexity. We call this novel algorithm MomentQuant. We provide evidence that our implementation of Quant is faster than the original one, and that MomentQuant is even faster than our implementation of Quant, at the cost of a tiny decrease in predictive performance. These improvements are especially relevant for real-life applications, where inference is performed much more often than training.
Chinese Translation
时间序列数据在许多现实应用和众多领域中非常普遍,人们越来越关注利用机器学习进行自动化信息提取。其中一个子领域是时间序列分类,即为每个新的、未见过的时间序列分配一个标签。过去几十年中已开发出许多算法,其预测性能与计算成本之间的权衡一直备受讨论。Quant 是一种基于区间的算法,从递归的、固定的、二分区间中提取分位数,研究表明它在速度极快的同时能够达到很高的准确率。我们提出两项改进以使该算法更快。第一项是对完全相同的算法进行更优化的实现。第二项是利用 Cornish-Fisher 展开来推导近似分位数,以替代精确分位数。这一改动省去了对时间序列排序的必要,从而降低了计算复杂度。我们将这一新算法命名为 MomentQuant。我们提供了证据表明,我们的 Quant 实现比原始实现更快,而 MomentQuant 比我们的 Quant 实现更快,代价是预测性能有极小的下降。这些改进对于现实应用尤为重要,因为在实际应用中,推理的执行频率远高于训练。
cs.LG / 64 / 2609.05138

From 80x to 385x: A Best-Matching-Unit Search at the L2 Roof, Measured Against a Symmetrically Tuned Baseline

从80倍到385倍:L2带宽屋顶下的最佳匹配单元搜索——与对称调优基线的对比测量
Amos, Andrew James
Abstract
Comparisons between GPU implementations are usually asymmetric: one side is tuned by its author, the other is run as found. I report a programme that tuned both a novel SOM algorithm (SparseBin) and the baseline algorithm it was being compared to (cuSPARSE). The best-matching-unit search that dominates self-organizing map training was tuned through four levers - tile size, tile-membership clustering, neuron-axis chunking and vectorised loads - reaching 5.6-10.1x per epoch over the previously published configuration at map sizes from 32x32 to 512x512, and lifting the margin over the CUDA implementation behind our earlier MEDLINE atlases from ~80x to ~385x. cuSPARSE, the implementation SparseBin is compared against, received every lever with an analogue on its side, and became 2-3x faster in the process. The tuned kernel pressed the L2 bandwidth roof at 77% of peak with every other unit at 40-65%, bounding any further lever at ~1.3x - a terminal result rather than a waypoint, and every untested lever was either capped by that bound by construction or measured null.
Chinese Translation
GPU实现之间的比较通常是不对称的:一方由其作者进行了调优,另一方则按原样运行。本文报告了一个同时对新颖SOM算法(SparseBin)及其对比基线算法(cuSPARSE)进行调优的程序。主导自组织映射(SOM)训练的最佳匹配单元(BMU)搜索通过四个手段进行了调优——瓦片大小、瓦片成员聚类、神经元轴分块和向量化加载——在32x32至512x512的映射规模下,每个训练周期的性能较此前发表的配置提升了5.6至10.1倍,并使我们早期MEDLINE图谱所依赖的CUDA实现的优势差距从约80倍提升至约385倍。作为SparseBin对比对象的cuSPARSE,也在其侧获得了每个调优手段的对应实现,并在此过程中提速了2至3倍。调优后的内核将L2带宽屋顶压至峰值的77%,而其他所有单元均处于40-65%之间,从而将任何进一步调优手段的收益上限限制在约1.3倍——这是一个终结性结果而非阶段性节点,且所有未测试的手段要么在构造上即受该上限约束,要么经测量被证明无效。
cs.LG / 65 / 2609.05150

Beyond Stationarity in Time Series: Discovering Causal Structures and Latent Regimes via Markov Blankets

超越时间序列的平稳性:基于马尔可夫毯发现因果结构与潜在机制
Zan, Lei, Assaad, Charles K., Devijver, Emilie, Gaussier, Eric
Abstract
This paper introduces Regime-aware Constraint-Based and Noise-Based causal discovery with Markov Blankets (RCBNB-MB), a novel causal discovery algorithm for time series that relaxes the common assumption of a single, time-consistent causal structure. Time series are typically observed at discrete time points and often exhibit regime changes that challenge the assumption of a static causal structure, a limitation in many real-world dynamic systems. To address this challenge, RCBNB-MB identifies latent causal regimes, defined as subsets of time points within which a stable causal structure holds. The algorithm follows an iterative strategy that segments the time series into regimes and discovers the causal graph within each regime. By leveraging the Markov blanket rather than direct parents, RCBNB-MB gains robustness to errors in causal discovery and preserves predictive information. We provide theoretical guarantees for RCBNB-MB's ability to recover both regime transitions and causal graphs under reasonable assumptions. Furthermore, we validate its effectiveness through extensive experiments on simulated datasets with known ground truth and real-world IT monitoring data, where taking into account regime shifts is critical. Empirical results show that RCBNB-MB systematically outperforms baseline approaches in accurately detecting regime changes and their associated causal graphs, positioning it as a robust and versatile framework for non-stationary time series analysis.
Chinese Translation
本文提出了一种基于机制感知的约束型与噪声型结合马尔可夫毯的因果发现算法(RCBNB-MB),这是一种新颖的时间序列因果发现算法,放宽了单一且时间一致的因果结构这一常见假设。时间序列通常在离散时间点上被观测,且常常表现出机制变化,这对静态因果结构的假设构成了挑战,而这一假设在许多现实世界的动态系统中是一种局限。为应对这一挑战,RCBNB-MB 识别潜在的因果机制,其定义为具有稳定因果结构的时间点子集。该算法采用迭代策略,将时间序列划分为若干机制,并在每个机制内发现因果图。通过利用马尔可夫毯而非直接父节点,RCBNB-MB 增强了对因果发现中错误的鲁棒性,并保留了预测信息。我们在合理的假设下,为 RCBNB-MB 恢复机制转换与因果图的能力提供了理论保证。此外,我们通过在具有已知真实标签的模拟数据集和真实世界 IT 监控数据(其中考虑机制变化至关重要)上的大量实验验证了其有效性。实证结果表明,RCBNB-MB 在准确检测机制变化及其对应的因果图方面系统性地优于基线方法,使其成为非平稳时间序列分析中一个鲁棒且通用的框架。
cs.LG / 66 / 2609.05194

Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets

相变频率作为ResNet测试准确率的训练时间预测指标
J, Arunan
Abstract
The number of discrete class-separability jumps observed during ResNet finetuning is examined empirically as a predictor of final test accuracy. Across 75 experiments spanning four benchmarks (CIFAR-10, CIFAR-100, TinyImageNet, and CIFAR-10-C) and three architectures (ResNet-18, ResNet-50, and ResNet-101), with five to ten seeds per configuration, a strong within-dataset negative correlation is obtained on standard i.i.d. classification benchmarks: \(r = -0.84\) on CIFAR-10 (\(p < 10^{-8}\), \(n = 30\)) and \(r = -0.87\) on CIFAR-100 (\(p < 10^{-5}\), \(n = 15\)). Under distributional stress, the relationship attenuates: TinyImageNet yields \(r = -0.45\), and the CIFAR-10-C corruption benchmark yields \(r = -0.19\). Two additional analyses discipline the empirical claim. A partial correlation controlling for architecture depth, treated as a linear covariate, shows that on CIFAR-100 the transition count retains statistically significant predictive power (\(r_{\mathrm{partial}} = -0.69\), \(p = 0.007\)); the corresponding result under the stricter categorical conditioning is not established at \(n = 15\). A comparison against six alternative training-curve signals shows that transition count achieved the strongest correlation among the evaluated signals on CIFAR-100 and one of the strongest on CIFAR-10, but is dominated by other signals on the two stressed benchmarks. The comparison is restricted to training-curve-level signals; comparisons against effective rank, Hessian sharpness, Fisher information, margin, and neural-collapse measures, which are the strongest competitors in the current literature, are not part of the present study and remain open. The observation is presented as an in-distribution training-quality probe among a family of candidate probes, and an inexpensive detection procedure suitable for logging alongside a standard training loop is provided.
Chinese Translation
本文实证考察了在ResNet微调过程中观察到的离散类别可分性跳变次数作为最终测试准确率预测指标的有效性。实验涵盖四个基准数据集(CIFAR-10、CIFAR-100、TinyImageNet和CIFAR-10-C)、三种架构(ResNet-18、ResNet-50和ResNet-101)的75次实验,每个配置使用五至十个随机种子。结果表明,在标准独立同分布(i.i.d.)分类基准上存在显著的同一数据集内负相关:CIFAR-10上为\(r = -0.84\)(\(p < 10^{-8}\),\(n = 30\)),CIFAR-100上为\(r = -0.87\)(\(p < 10^{-5}\),\(n = 15\))。在分布压力条件下,该关系减弱:TinyImageNet的相关系数为\(r = -0.45\),CIFAR-10-C损坏基准为\(r = -0.19\)。本文进行了两项补充分析以严谨验证该实证结论。以架构深度作为线性协变量进行偏相关分析表明,在CIFAR-100上,跳变次数仍具有统计显著的预测能力(\(r_{\mathrm{partial}} = -0.69\),\(p = 0.007\));而在\(n = 15\)的样本量下,更严格的类别条件下的相应结果未能确立。与六种替代训练曲线信号的比较表明,跳变次数在CIFAR-100上取得了所评估信号中最强的相关性,在CIFAR-10上也是最強之一,但在两个压力基准上则劣于其他信号。该比较仅限于训练曲线层面的信号;与有效秩、Hessian锐度、Fisher信息、间隔以及神经坍缩(neural collapse)度量——即当前文献中最强的竞争方法——的比较不在本研究范围内,仍是开放问题。本文将这一观察呈现为一族候选训练质量探测方法中的分布内训练质量探针,并提供了一种可随标准训练循环一同记录的低成本检测流程。
cs.LG / 67 / 2609.05214

Dimension-Adaptive Batched Lipschitz Narrowing Without Knowing the Zooming Dimension

无需知道缩放维度的维度自适应批量Lipschitz收窄算法
Feng, Yasong
Abstract
The Appropriately Combined Edge-length (ACE) sequence in A-BLiN depends on the zooming dimension $d_z$. This note removes that dependence. The next edge length is selected from the number of cubes that survive the preceding elimination. The resulting Count-Adaptive BLiN algorithm does not use $d_z$ or the zooming constant $C_z$, yet it attains $\widetilde{\mathcal O}_d(T^{(d_z+1)/(d_z+2)})$ regret with $\mathcal O_d(\log\log T)$ batches. Together with the adaptive-grid lower bound in Theorem 10 of the original paper, the optimal batch complexity remains $\Theta_d(\log\log T)$ when $d_z$ is unknown.
Chinese Translation
A-BLiN中的适当组合边长(Appropriately Combined Edge-length, ACE)序列依赖于缩放维度 $d_z$。本短文消除了这一依赖。下一条边长从上一轮淘汰后存活的立方体数量中选取。由此得到的Count-Adaptive BLiN算法既不使用 $d_z$,也不使用缩放常数 $C_z$,但仍能以 $\mathcal O_d(\log\log T)$ 的批次数达到 $\widetilde{\mathcal O}_d(T^{(d_z+1)/(d_z+2)})$ 的遗憾值。结合原论文定理10中的自适应网格下界,当 $d_z$ 未知时,最优批量复杂度仍为 $\Theta_d(\log\log T)$。
cs.LG / 68 / 2609.05223

FedDRAW: Federated Dual Reputation Annealing Weighting for Heterogeneous Multi-Institutional Chest Radiograph Classification

FedDRAW:用于异构多机构胸片分类的联邦双信誉退火加权方法
Moradpour, Maryam, Hauschild, Anne-Christin
Abstract
Artificial intelligence models are promising for medical diagnosis, but they require large numbers of unbiased data, which in medicine are distributed across hospitals and cannot be centralized to protect patient privacy. Federated Learning (FL) addresses this, since hospitals train one shared diagnostic model while patient data remain local. Training proceeds in communication rounds, in which each hospital trains the shared model locally and returns it to the server for merging by weighted average. This aggregation weight determines whose institutional knowledge shapes the result. Federated averaging (FedAvg) sets it in proportion to local sample count, so a small but informative hospital is permanently assigned a small influence, andl argest clients could dominate the global model even when they are less informative. We propose Federated Dual Reputation Annealing Weighting (FedDRAW), a server-side aggregation method that combines a data-size prior with the cosine similarity between client and global parameters under two coupled annealing schedules. An inner schedule shifts client reputation from the size prior towards similarity. An outer, deferred annealing schedule on the softmax inverse temperature keeps the weighting selective in the early and middle rounds and relaxes it to uniformity at convergence. We evaluate FedDRAW on 12 simulated client-partition scenarios of two chest radiograph datasets (CheXpert and ChestMNIST), against seven federated baselines under identical local training settings. FedDRAW achieved the highest average rank among all eight methods under both AUC and the geometric mean (GM) of sensitivity and specificity, which a Friedman test with Nemenyi post-hoc analysis confirmed to be a statistically significant difference between the methods. Scheduling two signals, rather than fixing the weights by sample count alone, could enable less biased diagnostic models.
Chinese Translation
人工智能模型在医学诊断中前景广阔,但其需要大量无偏数据,而这些数据在医学领域分布于各家医院,且为保护患者隐私无法集中管理。联邦学习(FL)解决了这一问题:各医院在本地训练一个共享的诊断模型,同时患者数据保留在本地。训练以通信轮次进行,每家医院在本地训练共享模型并将其返回服务器,通过加权平均进行合并。这一聚合权重决定了哪些机构的知识将影响最终结果。联邦平均算法(FedAvg)按本地样本数量比例设置该权重,因此样本量小但信息丰富的医院会被永久赋予较小的影响力,而样本量最大的客户端即使信息含量较低也可能主导全局模型。我们提出了联邦双信誉退火加权方法(FedDRAW),这是一种服务器端聚合方法,将数据规模先验与客户端和全局参数之间的余弦相似度相结合,并在两个耦合的退火调度下进行。内部调度将客户端信誉从规模先验逐步转向相似度;外部延迟退火调度作用于softmax逆温度,使加权在早期和中期轮次保持选择性,并在收敛时放宽至均匀加权。我们在两个胸片数据集(CheXpert和ChestMNIST)的12个模拟客户端划分场景上,在相同的本地训练设置下与七种联邦基线方法进行了对比评估。FedDRAW在AUC以及敏感度和特异度的几何平均值(GM)指标下均取得了所有八种方法中最高的平均排名,Friedman检验及Nemenyi事后分析证实了各方法之间的差异具有统计学显著性。通过调度两个信号而非仅依据样本数量固定权重,有望构建偏差更小的诊断模型。
cs.LG / 69 / 2609.05233

Hessian-based molecular conformation augmentation for a scalable and efficient strategy of machine learning interatomic potentials

基于Hessian的分子构象增强:一种可扩展且高效的机器学习原子间势策略
Kwak, Bumju, Jo, Jeonghee
Abstract
While machine-learning interatomic potentials (MLIPs) have successfully learned potential energy surfaces (PES) and atomic forces, many practical applications, such as vibrational analysis and transition state search, rely heavily on the PES Hessian. Yet, standard MLIPs tend to be trained on energy and forces alone, leaving Hessian information largely unexploited. Meanwhile, existing methods that explicitly incorporate the Hessian into training objectives require architectural modifications and introduce significant computational and memory overheads due to higher-order backpropagation. To address these limitations, we propose two Hessian-derived data augmentation schemes: isotropic Gaussian displacement (\textbf{UniAug}) and normal mode-weighted displacement (\textbf{ModeAug}). Both methods utilize simple Taylor expansions, achieving effective augmentation without altering training objectives or extending the autograd graph. This allows seamless, plug-and-play integration with existing architectures and training pipelines. Comprehensive evaluations across non-equilibrium and equilibrium datasets demonstrate that our approach enhances model accuracy while providing practical, task-specific guidelines.
Chinese Translation
尽管机器学习原子间势(MLIP)已成功学习势能面(PES)和原子力,但许多实际应用,如振动分析和过渡态搜索,高度依赖势能面的Hessian矩阵。然而,标准的MLIP通常仅基于能量和力进行训练,使得Hessian信息在很大程度上未被利用。与此同时,现有将Hessian显式纳入训练目标的方法需要对网络架构进行修改,并由于高阶反向传播而引入显著的计算和内存开销。为解决这些局限,我们提出了两种基于Hessian的数据增强方案:各向同性高斯位移(UniAug)和简正模式加权位移(ModeAug)。两种方法均利用简单的泰勒展开,在不改变训练目标或扩展自动微分计算图的情况下实现有效增强。这使得其能够与现有架构和训练流程无缝地即插即用集成。在非平衡与平衡数据集上的全面评估表明,我们的方法在提升模型精度的同时,提供了实用的、面向具体任务的指导方针。
cs.LG / 70 / 2609.05235

PRICE: A Systematic Study of LLM Adaptation Choices for Bitcoin Price Forecasting

PRICE:面向比特币价格预测的大语言模型适配选择的系统性研究
Fakhari, Maryam, Safayani, Mehran
Abstract
Cryptocurrency markets exhibit extreme volatility and non-stationary dynamics that challenge conventional forecasting methods. Although Large Language Models (LLMs) have shown promise for time series forecasting, the combined effects of adaptation choices remain largely unexplored in financial settings. This study introduces PRICE, a structured approach for adapting LLMs to short-term Bitcoin price forecasting. Built on a 4-bit quantized LLaMA-3 8B model, PRICE investigates how fine-tuning, numerical representation, prompting, inference, and decoding jointly influence forecasting performance. PRICE integrates Parameter-efficient fine-tuning with Low-Rank Adaptation (LoRA), Recursive multi-step inference, Integer-rounded numerical representation, Context-Task-Format (CTF) prompting, and Exact zero-temperature decoding. Ablation studies show that each component contributes to forecasting accuracy and reliability. LoRA enables efficient training on limited hardware, recursive inference improves accuracy, integer-rounded values reduce errors, CTF prompting outperforms Chain-of-Thought, Implicit Chain-of-Thought (iCoT), and few-shot prompting, and zero-temperature decoding improves stability during recursive forecasting. Comparative evaluation against eight transformer-based and time-series foundation models shows that PRICE achieves the lowest forecasting errors on both validation and test sets while maintaining robust performance across evaluation periods. Despite being based on a model primarily pretrained on text rather than time-series data, PRICE achieves competitive or superior performance relative to specialized foundation models. These findings demonstrate that adaptation choices critically determine the accuracy and robustness of LLMs for numerical time-series forecasting.
Chinese Translation
加密货币市场表现出极端波动性和非平稳动态特性,对传统预测方法构成挑战。尽管大语言模型(LLM)在时间序列预测方面展现出潜力,但各种适配选择在金融场景中的综合影响仍未得到充分探索。本研究提出了PRICE,一种将大语言模型适配于比特币短期价格预测的结构化方法。PRICE基于4比特量化的LLaMA-3 8B模型构建,研究了微调、数值表示、提示、推理和解码如何共同影响预测性能。PRICE集成了基于低秩适应(LoRA)的参数高效微调、递归多步推理、整数舍入数值表示、上下文-任务-格式(CTF)提示以及精确零温度解码。消融实验表明,每个组件都对预测的准确性和可靠性有所贡献:LoRA支持在有限硬件上进行高效训练;递归推理提高了准确性;整数舍入数值降低了误差;CTF提示优于思维链(Chain-of-Thought)、隐式思维链(iCoT)和少样本提示;零温度解码提升了递归预测过程中的稳定性。与八个基于Transformer的模型和时间序列基础模型的对比评估显示,PRICE在验证集和测试集上均取得了最低的预测误差,并在各评估时期保持稳健的表现。尽管PRICE所基于的模型主要是在文本而非时间序列数据上进行预训练的,其性能仍与专门的基础模型相当甚至更优。这些发现表明,适配选择对大语言模型在数值时间序列预测中的准确性和稳健性起着决定性作用。
cs.LG / 71 / 2609.05253

GLASS: Graph-Language Alignment with Spherical Scoring for Transferable Graph-Level Anomaly Detection

GLASS:基于球面评分的图-语言对齐方法,实现可迁移的图级异常检测
Wang, Xudong, Ding, Chris, Li, Tongxin, Fan, Jicong
Abstract
We introduce GLASS, a framework for graph-level anomaly detection (GLAD) that achieves robust cross-domain transferability through graph-language alignment on the unit hypersphere. GLASS builds a unified representation space by aligning a structure-aware graph encoder with an instruction-aware text embedding via a multi-slice soft cosine objective. Our framework serializes local, global, and semantic graph properties into a compact Graph Descriptor Prompt (GraphDP), creating a text bridge that enables domain-agnostic anomaly scoring. By enforcing multi-scale consistency through Matryoshka representation slices, the model captures anomalous deviations at multiple levels of granularity. For scoring, we formulate anomaly detection as density estimation on the aligned hypersphere and introduce Spherical Multi-Modal Scoring (SMS), which instantiates von Mises-Fisher kernel density estimators in both graph and text embedding spaces. This probabilistic formulation recovers angular k-nearest-neighbor scoring as a high-concentration limiting case and provides a principled fusion of structural and semantic anomaly signals. The shared text embedding space further serves as a cross-domain bridge: by encoding a target domain's GraphDP without target-domain training data, GLASS performs zero-shot anomaly detection, and with only a handful of normal examples, few-shot adaptation via reference-set calibration. Across twelve benchmarks and three meta-domains, GLASS obtains the best average AUROC and rank compared with recent advanced GLAD baselines and enables effective cross-domain transfer.
Chinese Translation
我们提出了GLASS,一个面向图级异常检测(GLAD)的框架,通过在单位超球面上的图-语言对齐实现强大的跨域可迁移性。GLASS利用多切片软余弦目标函数,将结构感知的图编码器与指令感知的文本嵌入对齐,从而构建统一表示空间。我们的框架将图的局部、全局和语义属性序列化为紧凑的图描述提示(Graph Descriptor Prompt, GraphDP),构建了文本桥梁,使异常评分能够与具体领域无关。通过Matryoshka表示切片实现多尺度一致性约束,模型能够在多个粒度层次上捕捉异常偏差。在评分方面,我们将异常检测表述为对齐超球面上的密度估计问题,并提出球面多模态评分(Spherical Multi-Modal Scoring, SMS),在图嵌入空间和文本嵌入空间中分别实例化von Mises-Fisher核密度估计器。这一概率化表述将基于角度的k近邻评分作为高浓度极限情形加以恢复,并为结构异常信号与语义异常信号提供了有原则的融合。共享的文本嵌入空间进一步充当跨域桥梁:在无需目标域训练数据的情况下,通过编码目标域的GraphDP,GLASS可实现零样本异常检测;仅利用少量正常样本,即可通过参考集校准实现少样本适配。在涵盖三个元领域的十二个基准数据集上,GLASS相比近期先进的GLAD基线方法取得了最优的平均AUROC和排名,并展现出有效的跨域迁移能力。
cs.LG / 72 / 2609.05274

How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method

如何推测智能体编程中的不确定性?一种草稿模型门控方法
Grotov, Konstantin, Malykh, Valentin
Abstract
LLM agents deployed for software engineering fail expensively: they act confidently wrong, and bad actions are recognized only after costly execution and retry. We present Speculative Uncertainty (SU), a method that recovers a predictive failure signal for a black-box agent from its output tokens alone, with no access to logits, weights, activations, or repeated sampling. Inverting speculative decoding, a small open-weight draft model scores the agent's already-generated trajectory in a single forward pass. From these speculative cross-likelihoods we extract phase-aware features by separating the reasoning and action spans, and calibrate them against a verifiable objective. SU produces a failure-likelihood score that any downstream policy, such as routing, human intervention, or extra test-time compute, can consume directly. To show the signal is actionable, we instantiate one such policy, a pre-execution veto gate, on software engineering agents Qwen3-Coder-480B and closed-source Claude 3.5 Sonnet, cutting execution error rate by 6-8 percentage points and token cost by 14-19% in deployment, transferring to out-of-distribution benchmarks without retraining, and generalizing across agent models.
Chinese Translation
部署于软件工程的LLM智能体失败代价高昂:它们自信地做出错误行为,而不良行为只有在代价高昂的执行和重试之后才被发现。我们提出推测性不确定性,该方法仅从黑盒智能体的输出标记中恢复出一个预测性失败信号,无需访问logits、权重、激活值或重复采样。通过反转推测解码,一个小型开源权重的草稿模型在单次前向传播中对智能体已生成的轨迹进行评分。从这些推测性交叉似然中,我们通过区分推理与动作片段提取阶段感知特征,并将其与可验证的目标进行校准。SU产生一个失败似然分数,任何下游策略——如路由、人工干预或额外的测试时计算——都可以直接使用。为证明该信号是可操作的,我们在软件工程智能体Qwen3-Coder-480B和闭源的Claude 3.5 Sonnet上实例化了一种此类策略——执行前否决门控,在部署中将执行错误率降低6-8个百分点、token成本降低14-19%,无需重新训练即可迁移到分布外基准,并能泛化到不同智能体模型。
cs.LG / 73 / 2609.05294

Learning from VAE Errors to support ECG-based Differential Diagnosis of Myocardial Scar

从VAE误差中学习以支持基于心电图的心肌瘢痕鉴别诊断
Sharifi, Shayan, Treu, Riccardo, Gandin, Ilaria, Garoia, Federico, Merlo, Marco, Cisotto, Giulia
Abstract
Late Gadolinium Enhancement (LGE) on cardiac magnetic resonance is a key marker of myocardial scar, but its limited accessibility motivates routine ECG-based screening. We evaluated whether $\beta$-variational autoencoder (VAE)-derived ECG representations can discriminate LGE+ from LGE- cardiomyopathic patients in a local cohort of 300 subjects. We compared 32-dimensional features from the foundation ECGx.AI model with those from a shallower $\beta$-VAE trained on normal PTB-XL ECGs, evaluating downstream classification and Dynamic Time Warping (DTW)-based reconstruction errors. ECGx.AI reached an area under ROC of 0.686 with Random Forest, while the proposed $\beta$-VAE reached 0.577 with sensitivity of 0.775 with Gradient Boosting. Notably, DTW-reconstruction errors significantly differed between classes in 10 out of 12 leads according to Mann-Whitney U test and help in classification, leading to an area under ROC of 0.643 with Logistic Regression, supporting their potential as markers of scar-related ECG alterations.
Chinese Translation
心脏磁共振上的钆延迟增强(LGE)是心肌瘢痕的关键标志物,但其可及性有限,因此有必要开展基于常规心电图的筛查。我们评估了由β-变分自编码器(VAE)提取的心电图表征能否在300名受试者的本地队列中区分LGE阳性与阴性的心肌病患者。我们将来自基础模型ECGx.AI的32维特征与在正常PTB-XL心电图上训练的较浅层β-VAE所提取的特征进行比较,评估下游分类性能以及基于动态时间规整(DTW)的重构误差。ECGx.AI结合随机森林达到0.686的ROC曲线下面积(AUC),而所提出的β-VAE结合梯度提升达到0.577的AUC和0.775的灵敏度。值得注意的是,根据Mann-Whitney U检验,12个导联中有10个导联的DTW重构误差在两类之间差异显著,并有助于分类,结合逻辑回归达到0.643的AUC,支持其作为瘢痕相关心电图改变标志物的潜力。
cs.LG / 74 / 2609.05309

How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing

mHC如何使用其残差流?选择性路由与近恒等混合
Zhao, Pengxiang, Li, Xing, Yu, Xianzhi, Guo, Wei, Dong, Zhenhua
Abstract
Hyper-Connections and their manifold-constrained variant mHC widen a residual pathway from one stream to n, yet how trained models use this capacity remains unclear: how broadly blocks read and write, how strongly the residual pathway mixes streams, and whether the streams carry distinct representations. We examine these properties in the four-stream residual pathway of DeepSeek-V4-Flash using effective stream counts, cross-stream residual weights, and inter-stream cosine similarity. Read/write routing is concentrated but varies across depth: a typical attention or FFN site effectively uses about two streams, while the dominant stream changes across layers and the representations remain directionally distinct. Residual mixing is modest and occurs primarily in early layers; in layers 22-42, the pathway mostly carries each stream forward separately. Targeted interventions establish the functional significance of these patterns. Replacing the late mixers by identity increases C4 perplexity by only 1.9% and preserves the six-task average score, whereas replacing the early mixers increases perplexity by 41%. Fixing each early mixer to its C4 diagnostic mean increases perplexity by only 0.2% and reduces the average score by 0.25 percentage points, showing that its site-specific structure matters more than its token-wise variation on the evaluated metrics. Likewise, retaining the three largest routing weights per token at every site increases perplexity by at most 2.7% and changes the average score by at most 0.4 points. Thus, the studied model realizes only part of the flexibility afforded by four-stream mHC: individual blocks rarely require all four streams, and late residual mixing provides little measured benefit.
Chinese Translation
超连接(Hyper-Connections)及其流形约束变体mHC将残差通路从单一流扩展到n条流,然而训练后的模型如何利用这一容量仍不清楚:各模块读写范围有多广、残差通路对流进行混合的强度有多大、各条流是否携带彼此不同的表示。我们利用有效流数、跨流残差权重和流间余弦相似度,在DeepSeek-V4-Flash的四流残差通路中考察了这些性质。读/写路由是集中的,但随深度变化:一个典型的注意力或FFN位点有效利用约两条流,主导流随层变化,且各流的表示在方向上保持彼此不同。残差混合较为温和,且主要发生在早期层;在第22-42层中,该通路大多将各条流单独向前传递。针对性干预确立了这些模式的功能意义。将后期混合器替换为恒等映射仅使C4困惑度增加1.9%,并保持六项任务的平均得分;而将早期混合器替换为恒等映射则使困惑度增加41%。将每个早期混合器固定为其在C4上的诊断均值,仅使困惑度增加0.2%,平均得分下降0.25个百分点,这表明在所评估的指标上,其位点特异的结构比逐词变化更为重要。同样地,在每个位点仅保留每个词元最大的三个路由权重,最多使困惑度增加2.7%,平均得分变化不超过0.4分。因此,所研究的模型仅实现了四流mHC所提供灵活性的部分能力:单个模块很少需要全部四条流,且后期残差混合带来的可测量收益甚微。
cs.LG / 75 / 2609.05318

Optimal Rates for Agentic Networked Information Aggregation

智能体网络化信息聚合的最优速率
Bateni, MohammadHossein, Hadizadeh, Zahra, Hajiaghayi, MohammadTaghi, JafariRaviz, Mahdi, Taherijam, Shayan
Abstract
Building on the pioneering paper of Kearns, Roth, and Ryu (SODA'26), we study information aggregation in a networked learning model. The model captures a central pattern in agentic AI: each agent sees only part of the data and passes on only its own conclusion. Their model considers a linear regression problem with the mean squared error (MSE) loss. Agents sit in a DAG and each sees only a subset of the features and its parents' predictions, fits a linear predictor, and passes only its prediction forward. The benchmark is the full-feature learner that sees all raw features. A path of depth $D$ is $M$-covered if every block of $M$ consecutive agents collectively sees all raw features. Kearns, Roth, and Ryu proved that the excess mean squared error of the last agent on such a path is $O(M/\sqrt D)$, and gave a cyclic instance with excess error $\Omega(M/D)$ for $D<M^2$. We close this gap: the correct rate is constant up to depth $M^2$, and $\Theta(M^2/D)$ beyond it. We first give a sharper analysis of the cyclic instance and improve its lower bound to $\Omega(\sqrt{M/D})$ for $D<M^2$. We then construct, for every depth $D\ge M^2$, an $M$-covered path of depth $D$ with excess error $\Omega(M^2/D)$. The same instance gives the constant lower bound for all $D < M^2$. We also show that for any fixed distribution the excess error contracts geometrically along the path, ruling out any single instance that witnesses any polynomial lower bound at every depth. Finally, we prove the same optimal rate for logistic classification in the logit-passing model of Bateni et al., which considers the binary cross-entropy (BCE) loss. The same improved upper bound of $O(M^2/D)$ holds, and we transfer all the regression lower bounds by showing that on those examples the logistic path follows the least-squares path up to rescaling.
Chinese Translation
在 Kearns、Roth 和 Ryu(SODA'26)的开创性论文基础上,我们研究网络化学习模型中的信息聚合问题。该模型刻画了智能体人工智能(agentic AI)中的一个核心模式:每个智能体只能看到部分数据,且仅传递自己的结论。他们的模型考虑以均方误差(MSE)为损失函数的线性回归问题。智能体位于一个有向无环图(DAG)中,每个智能体只能看到一部分特征及其父节点的预测,拟合一个线性预测器,并仅将其预测向前传递。基准是能够看到所有原始特征的全特征学习器。如果每连续 $M$ 个智能体的集合能共同看到所有原始特征,则称该深度为 $D$ 的路径是 $M$-覆盖的。Kearns、Roth 和 Ryu 证明了在此类路径上最后一个智能体的超额均方误差为 $O(M/\sqrt{D})$,并给出一个循环实例,当 $D<M^2$ 时其超额误差为 $\Omega(M/D)$。我们弥合了这一差距:正确的速率在深度达到 $M^2$ 之前为常数,超过之后为 $\Theta(M^2/D)$。我们首先对循环实例给出更精细的分析,将其下界改进为当 $D<M^2$ 时的 $\Omega(\sqrt{M/D})$。然后,我们对每个深度 $D\ge M^2$ 构造一个超额误差为 $\Omega(M^2/D)$ 的 $M$-覆盖路径。同一实例对所有 $D < M^2$ 给出了常数下界。我们还证明,对于任意固定的分布,超额误差沿路径呈几何式收缩,从而排除了任何能在每个深度都产生多项式下界的单一实例。最后,我们在 Bateni 等人考虑二元交叉熵(BCE)损失、以 logit 传递为机制的逻辑斯谛分类模型中证明了相同的最优速率。同样的改进上界 $O(M^2/D)$ 成立,并且通过证明在这些例子上逻辑斯谛路径在重新缩放意义下与最小二乘路径一致,我们将所有回归下界迁移了过来。
cs.LG / 76 / 2609.05328

Embedded Graph Flows for Categorical Graph Generation

用于类别型图生成的嵌入式图流
Ma, Ethan, Wang, Zihan, Chow, Chris Siu Yeung, Feng, Xinguo, Li, Qingqing, Jiang, Rui, Dong, Naipeng, Bai, Guangdong
Abstract
Generating categorical graphs requires choosing node and edge types that form a coherent structure without depending on node order. Many graph generators encode categories as fixed one-hot vectors, which can impose an artificial geometry in which categories are equidistant. We propose Embedded Graph Flows (EGF), a generative model that learns continuous embeddings for node and unordered-edge categories and transports Gaussian noise towards these learnt endpoints using a permutation-equivariant graph transformer. A terminal readout maps the embeddings back to discrete graph categories. Across molecular benchmarks, EGF achieved competitive performance. On QM9, EGF gives the best result on all four reported metrics among the three methods, including a Fr\'echet ChemNet Distance (FCD) of 0.150, compared with 0.717 for the categorical-diffusion baseline DiGress and 0.812 for the bridge-based baseline GruM. When applied to larger molecules in ZINC250k, EGF retains the lowest maximum mean discrepancy (MMD) using the neighbourhood subgraph pairwise distance kernel (NSPDK), indicating close agreement with the local substructures of the reference molecules. Our code is available at https://github.com/Trusted-System-Lab/EGF.
Chinese Translation
生成类别型图需要选择节点类型和边类型,使其形成连贯的结构,且不依赖于节点顺序。许多图生成器将类别编码为固定的独热(one-hot)向量,这会引入一种人工几何结构,使得各类别彼此等距。我们提出了嵌入式图流,这是一种生成模型,它为节点类别和无序边类别学习连续嵌入,并利用置换等变的图变换器(graph transformer)将高斯噪声输运至这些学习到的终点。一个终端读出模块将嵌入映射回离散的图类别。在多个分子基准测试中,EGF 取得了具有竞争力的性能。在 QM9 数据集上,EGF 在三种方法的全部四项报告指标中均取得最佳结果,其中 Fréchet ChemNet 距离(FCD)为 0.150,而类别扩散基线方法 DiGress 为 0.717,基于桥的基线方法 GruM 为 0.812。在应用于 ZINC250k 中更大的分子时,EGF 在使用邻域子图成对距离核(NSPDK)的情况下仍保持最低的最大均值差异(MMD),表明其与参考分子的局部子结构高度吻合。我们的代码发布于 https://github.com/Trusted-System-Lab/EGF。
cs.LG / 77 / 2609.05337

Variational Continuation for Double Pendulum Periodic Orbits

双摆周期轨道的变分延拓方法
Yao, Leo, Liu, Ziming, Tegmark, Max
Abstract
We present a Hessian-based approach to numerically continue periodic orbits in dynamical systems. A loop (periodic orbit candidate) is parametrized as a Fourier series; a loss function is defined based on the deviation of the loop from the physical differential equations. Unlike previous work relying on hand-derived Jacobians, our method automates the process by leveraging automatic differentiation, a common machine learning technique. The continuation direction can be determined by the flat directions of the loss landscapes (directions with zero eigenvalues), making the search of periodic orbits efficient and guided. Our method is integrator-free, precisely initializes oscillations around unstable fixed points, and efficiently detects orbit family intersections and subharmonic bifurcations. As a demonstration, we present full continuations of periodic double pendulum oscillations from fixed points, showing bifurcations along orbit families and categorizing branches of periodic orbits. In particular, we find periodic orbits where both pendulum masses are never simultaneously at rest, which to our knowledge has been missing in the literature.
Chinese Translation
我们提出了一种基于Hessian矩阵的数值延拓方法,用于动力学系统中的周期轨道求解。该方法将闭合回路(周期轨道候选)参数化为傅里叶级数,并根据回路与物理微分方程的偏差定义损失函数。与以往依赖人工推导雅可比矩阵的工作不同,我们的方法利用自动微分这一常见的机器学习技术来实现过程的自动化。延拓方向可通过损失景观的平坦方向(特征值为零的方向)确定,从而使周期轨道的搜索更加高效且有引导性。我们的方法无需积分器,能够精确初始化不稳定不动点附近的振荡,并可高效地检测轨道族交叉与次谐波分岔。作为示范,我们给出了双摆周期振荡从不动点出发的完整延拓结果,展示了轨道族上的分岔并对周期轨道分支进行了分类。特别地,我们发现了两个摆锤质量从未同时处于静止状态的周期轨道,据我们所知,这类轨道在以往文献中尚付阙如。
cs.LG / 78 / 2609.05363

Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation

全局蒸馏,局部适配:面向可扩展换购推荐的推理蒸馏与产品类型测试时训练
Liu, Siliang, Ghasemi, Mohammad, Patel, Sapan, Banitalebi-Dehkordi, Amin
Abstract
Trade-up recommendation identifies higher-quality alternatives that preserve a customer's purchase intent while offering upgraded benefits. Large language models (LLMs) can reason about such distinctions, but applying them directly to hundreds of millions of product pairs is operationally impractical. We introduce a two-level framework that distills LLM reasoning into an efficient non-generative student and adapts its decision boundary to product-type-specific trade-up criteria. At Level 1, a retrieval-augmented few-shot LLM teacher generates structured relation labels and natural-language rationales. These rationales supervise a compact embedding-pair classifier through alignment and contrastive objectives; at inference, the student uses only two precomputed 768-dimensional product embeddings, with no LLM calls or text generation. On a fixed human-annotated benchmark of 8,352 pairs, a 15.5M-parameter four-class reasoning-distilled student achieves AUC 0.924 (95% CI [0.918, 0.929]), compared with 0.912 for the four-class label-only student. At Level 2, product-type test-time training (PT-TTT) uses few-shot demonstrations to optimize lightweight category-specific adapters over the frozen student. PT-TTT improves AUC from 0.924 to 0.941 and average precision from 0.920 to 0.940. On a 100K-pair proxy catalog, the distilled student on a single eight-GPU machine is approximately 5,000x faster and 10,000x lower in estimated cost than direct LLM inference.
Chinese Translation
换购推荐(Trade-up recommendation)旨在识别在保持顾客购买意图的同时提供升级权益的更高质量替代品。大型语言模型(LLM)能够对此类差异进行推理,但将其直接应用于数以亿计的商品对在运营上并不可行。我们提出了一种两级框架,将LLM的推理能力蒸馏到一个高效的非生成式学生模型中,并将其决策边界适配至产品类型特定的换购标准。在第一级,检索增强的少样本LLM教师模型生成结构化的关系标签和自然语言推理依据。这些推理依据通过对比和对齐目标来监督一个紧凑的嵌入对分类器;在推理时,学生模型仅使用两个预计算的768维商品嵌入,无需任何LLM调用或文本生成。在一个包含8,352个商品对的固定人工标注基准上,一个1550万参数的四分类推理蒸馏学生模型达到了AUC 0.924(95%置信区间[0.918, 0.929]),而仅使用标签的四分类学生模型为0.912。在第二级,产品类型测试时训练(PT-TTT)利用少样本示例,在冻结的学生模型之上优化轻量级的品类特定适配器。PT-TTT将AUC从0.924提升至0.941,平均精度从0.920提升至0.940。在一个10万商品对的代理目录上,蒸馏后的学生模型在单台八GPU机器上的运行速度比直接LLM推理快约5,000倍,估计成本低约10,000倍。
cs.LG / 79 / 2609.05403

RegionFed: Federated Learning for Personalized Query Understanding in Heterogeneous Retail Environments

RegionFed:面向异构零售环境中个性化查询理解的联邦学习
Nguyen, Quoc H., Lafzi, Ali, Phatak, Abhijeet, Singh, Siddharth Pratap, Upadhyay, Rohit, Seetharama, Yogananda Domlur, Tripathy, Chittaranjan
Abstract
Retail search systems serve diverse geographic regions with distinct query patterns, vocabularies, and product preferences, creating significant data heterogeneity that challenges both privacy-preserving training and model personalization. Federated learning offers a natural solution for privacy, but standard FL methods produce global models that sacrifice regional performance, while existing personalized FL approaches operate at the parameter level and catastrophically collapse on modern transformers (below 10\% accuracy on T5) due to tied embeddings and LayerNorm interactions. We introduce RegionFed, an \textit{architecture-robust} federated learning framework that sidesteps this failure by operating entirely at the gradient level. RegionFed uses the $\ell_2$ conflict between regional and global gradients as a unified signal that (i) diagnoses heterogeneity, (ii) routes each region to the cheapest sufficient personalization strategy, and (iii) adaptively controls personalization strength. Because it treats models as differentiable black boxes, RegionFed deploys on T5-Small, T5-3B, RoBERTa, and CNN with zero code changes, providing large gains on transformers (where parameter-level methods collapse) and consistent improvements on CNNs. Across three public datasets (Amazon ESCI, Amazon Reviews, LEAF-FEMNIST) and four architectures, RegionFed-Meta achieves 92.27\%, closing the gap to the privacy-violating centralized upper bound (Centralized + Regional Weighting: 92.04\%, $\Delta$=0.23pp, within 1$\sigma$) while providing $(\epsilon{\approx}0.60)$-differential privacy and $\mathcal{O}(1/\sqrt{T})$ convergence.
Chinese Translation
零售搜索系统服务于具有不同查询模式、词汇和产品偏好的多样化地理区域,产生了显著的数据异构性,这对隐私保护训练和模型个性化都构成了挑战。联邦学习为隐私保护提供了天然的解决方案,但标准的联邦学习方法产生牺牲区域性能的全局模型,而现有的个性化联邦学习方法在参数层面运行,由于嵌入共享(tied embeddings)和 LayerNorm 的相互作用,在现代 Transformer 模型上会发生灾难性崩溃(在 T5 上的准确率低于 10%)。我们提出了 RegionFed,这是一种*架构鲁棒*的联邦学习框架,它完全在梯度层面运行,从而避开了上述失败。RegionFed 利用区域梯度与全局梯度之间的 ℓ2 冲突作为统一信号,用于 (i) 诊断异构性,(ii) 将每个区域路由到成本最低的、足够的个性化策略,以及 (iii) 自适应地控制个性化强度。由于它将模型视为可微分的黑盒,RegionFed 无需任何代码修改即可部署在 T5-Small、T5-3B、RoBERTa 和 CNN 上,在 Transformer 模型上(参数级方法会崩溃的场景)取得了显著提升,并在 CNN 上也带来了一致的改进。在三个公开数据集(Amazon ESCI、Amazon Reviews、LEAF-FEMNIST)和四种架构上,RegionFed-Meta 达到了 92.27% 的准确率,缩小了与违反隐私的集中式上界(集中式 + 区域加权:92.04%,Δ=0.23 个百分点,在 1σ 范围内)的差距,同时提供了 (ε≈0.60)-差分隐私和 O(1/√T) 的收敛保证。
机器人学 (Robotics)
32
cs.RO / 1 / 2609.04277

FailureSpot: Label-Efficient Timestamp-Level Failure Detection for Vision-Language-Action Models

FailureSpot:面向视觉-语言-动作模型的标签高效时间戳级失败检测
Ma, Jie, Liu, Zongxi, Zhu, Yi
Abstract
Vision-language-action (VLA) policies have shown strong potential for general-purpose robotic manipulation, but they can still fail unpredictably during long-horizon execution, making reliable failure detection essential for safe deployment. Existing methods either rely on visual models that typically detect failures only after erroneous actions have occurred, or use lightweight proactive detectors trained on VLA internal representations. However, these proactive methods are often supervised with trajectory-level labels, causing normal pre-failure behavior in unsuccessful trajectories to be incorrectly labeled as failure. This supervision mismatch introduces label noise and limits both trajectory-level detection accuracy and precise timestamp-level failure localization. In this work, we study fine-grained timestamp-level VLA failure detection while addressing the cost of dense annotation. We propose a data-efficient framework that first leverages unlabeled VLA action chunks to construct action-derived weak supervision signals, capturing abnormal patterns such as inconsistent consecutive chunks, frozen or idle actions, and aggressive random motions. We then use active learning to select only the most uncertain trajectories for timestamp-level annotation and fine-tune the detector with these informative labels. Experiments across multiple VLA policies show that our method improves both timestamp-level and trajectory-level failure detection performance.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)策略在通用机器人操作中展现出强大潜力,但在长时程执行过程中仍可能出现不可预测的失败,因此可靠的失败检测对安全部署至关重要。现有方法要么依赖视觉模型,通常只能在错误动作发生后才检测到失败;要么使用基于VLA内部表征训练的轻量级主动检测器。然而,这些主动方法通常采用轨迹级标签进行监督,导致失败轨迹中失败前的正常行为被错误地标注为失败。这种监督不匹配引入了标签噪声,既限制了轨迹级检测的准确性,也限制了精确的时间戳级失败定位。在本工作中,我们研究细粒度的时间戳级VLA失败检测,同时应对密集标注的高成本问题。我们提出一种数据高效的框架,首先利用无标注的VLA动作块构建由动作导出的弱监督信号,以捕捉异常模式,如相邻动作块不一致、动作冻结或闲置,以及剧烈的随机运动。随后,我们采用主动学习仅选择最不确定的轨迹进行时间戳级标注,并利用这些有信息量的标签对检测器进行微调。在多个VLA策略上的实验表明,我们的方法同时提升了时间戳级和轨迹级的失败检测性能。
cs.RO / 2 / 2609.04355

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

VLA-Precision:面向视觉-语言-动作模型高效真实世界在线强化学习的非对称协同自举方法
Su, Chenyu, Shen, Zhaolong, Qian, Yuan, Qian, Chen, Zhang, Rui, Yan, Feng, Chen, Weixing, Zhang, Fei, Wang, Jiamin, Cong, Shuang, Shang, Weiwei
Abstract
Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demanding precision and repeatability. Applying real-world online reinforcement learning (RL) to VLA post-training enables autonomous trial-and-error improvement beyond demonstrations alone, but exposes two bottlenecks: 1) unreliable value signals can induce policy drift; 2) large-VLA overhead constrains throughput and sample efficiency. To address these challenges, we present VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture. Specifically, ACoB establishes asymmetric co-bootstrapping across timescales: early intervention-guided behavioral learning rapidly improves policy performance while enhancing online experience quality. As autonomous experience accumulates, global return propagation and local preference ranking progressively calibrate value estimates, yielding relative action advantages for reference-regularized policy improvement while suppressing drift. To enable ACoB on large VLAs, we develop ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, delivering up to 10.9$\times$ improvements in throughput and computational efficiency. Extensive evaluations on nine high-precision chemistry tasks across four categories and four robot embodiments show that VLA-Precision achieves 98.3\% mean success rate in 45.8 min/task, with 27.6 s episodes running at 1.2$\times$ and 1.8$\times$ the speeds of VLA and RL baselines. Resources are available at https://vla-precision.github.io.
Chinese Translation
预训练的视觉-语言-动作(VLA)模型能够实现广泛的操作任务,但在要求精度和可重复性的任务中仍然不够可靠。将真实世界在线强化学习(RL)应用于VLA后训练,使其能够在示范数据之外通过自主试错持续改进,但这面临两个瓶颈:1)不可靠的价值信号可能导致策略漂移;2)大型VLA的开销限制了吞吐量和样本效率。为应对这些挑战,我们提出VLA-Precision,一个高效的真实世界在线强化学习框架,其核心是非对称协同自举(Asymmetric Co-Bootstrapping, ACoB)算法和ACoB-Stream架构。具体而言,ACoB在不同时间尺度上建立非对称协同自举:在早期阶段,干预引导的行为学习能够快速提升策略性能,同时提高在线经验的质量。随着自主经验的积累,全局回报传播与局部偏好排序逐步校准价值估计,产生用于参考正则化策略改进的相对动作优势,同时抑制策略漂移。为使ACoB能够应用于大型VLA,我们开发了ACoB-Stream,一种闭环“经验-策略”架构,以不变状态解耦和按需流式处理为设计原则,在吞吐量和计算效率方面实现了高达10.9倍的提升。在涵盖四个类别、四种机器人本体的九项高精度化学任务上的广泛评估表明,VLA-Precision在45.8分钟/任务内达到98.3%的平均成功率,其中27.6秒的回合分别以VLA基线和RL基线1.2倍和1.8倍的速度运行。相关资源见 https://vla-precision.github.io。
cs.RO / 3 / 2609.04364

Scalable Edge-assisted Fusion and Path Prediction for Connected Autonomous Vehicles

面向网联自动驾驶汽车的可扩展边缘辅助融合与路径预测
Landle, Tyler, Isenberg, Jackson, Chatterjee, Abhijit, Daglis, Alexandros, Ramachandran, Umakishore
Abstract
The planning algorithms inside an Autonomous Vehicle (AV) rely on information from on-board sensors whose line of sight is limited by emerging traffic conditions and occlusions. Edge-assisted creation of a unified world model fusing information from AVs and Road Side Units (RSUs) in a geographical locale, and the prediction of AVs' future trajectories, can enhance the planning algorithms inside AVs to improve quality metrics, such as better traffic flow and collision prevention. AVs participating in such enhancements are called Connected Autonomous Vehicles (CAVs). However, such information generated by the edge (world model and motion predictions) must reach the planners within a tight Age of Information (AoI) time budget to be useful. The state of the art fuses per-CAV information: each AV fuses inputs from other actors locally, which limits both scalability with actor count and quality of results. We present Conductor, an edge-based solution for creating a unified world model from the perspective of a fixed anchor (e.g., an RSU) in a locale and predicting future trajectories of AVs in that locale. Our solution adheres to the AoI time budget by dynamically limiting the number of AVs that would lead to the best quality of results. Specifically, we introduce an occlusion-aware selector that favors information contribution by AVs that detect objects in the locale not covered by RSUs. We pair this selector with a runtime controller that adapts both the number of AV inputs to fuse and the amount of trajectory predictions in each cycle to stay within the AoI time budget. Evaluation on CAV simulation infrastructure shows our joint selector-controller meets the AoI safety bound across traffic scenarios with up to 31 CAVs, with fusion fidelity close to an Oracle and much better than a random selector under the same AoI constraint.
Chinese Translation
自动驾驶汽车(AV)内部的规划算法依赖于车载传感器的信息,而传感器的视线范围受不断出现的交通状况和遮挡的限制。在边缘端构建一个融合某地理区域内自动驾驶汽车与路侧单元(RSU)信息的统一世界模型,并对自动驾驶汽车的未来轨迹进行预测,可以增强自动驾驶汽车内部的规划算法,从而改善交通流量和防碰撞等质量指标。参与此类增强的自动驾驶汽车被称为网联自动驾驶汽车(CAV)。然而,边缘端生成的此类信息(世界模型和运动预测)必须在严格的信息年龄(AoI)时间预算内到达规划器才能发挥作用。现有技术逐个融合每辆CAV的信息:每辆自动驾驶汽车在本地融合来自其他交通参与者的输入,这既限制了随交通参与者数量增加的可扩展性,也限制了结果质量。我们提出了Conductor,一种基于边缘的解决方案,用于从区域中固定锚点(例如RSU)的视角构建统一世界模型,并预测该区域内自动驾驶汽车的未来轨迹。我们的解决方案通过动态限制能够带来最佳结果质量的自动驾驶汽车数量来遵守AoI时间预算。具体而言,我们引入了一种遮挡感知选择器,优先选择那些能够检测到RSU未覆盖区域内目标的自动驾驶汽车所贡献的信息。我们将该选择器与运行时控制器相结合,该控制器自适应调整每周期内融合的自动驾驶汽车输入数量和轨迹预测数量,以保持在AoI时间预算内。在CAV仿真基础设施上的评估表明,我们的联合选择器-控制器方案在多达31辆CAV的各种交通场景下均满足AoI安全界限,其融合保真度接近Oracle,且在相同AoI约束下远优于随机选择器。
cs.RO / 4 / 2609.04411

AquaBEV: Monocular Underwater BEV Occupancy with 3D Sonar Supervision

AquaBEV:基于三维声呐监督的单目水下BEV占据预测
Dong, Trung Tien, Jin, Shengji, Chen, Chen, Sheng, Yi, Lin, Xiaomin
Abstract
Autonomous underwater robots are widely used for exploration, monitoring, and inspection, where safe navigation depends on understanding the surrounding free and occupied space. Bird's eye view (BEV) occupancy provides such a representation, but predicting it from a single underwater RGB image is difficult due to limited, unreliable geometric cues from appearance alone. 3D imaging sonar offers complementary geometric measurements to supervise this task. We introduce AquaBEV, a monocular underwater occupancy model that predicts local BEV occupancy from a single RGB image, using paired 3D imaging sonar as geometric supervision during training. AquaBEV maps visual features into a calibration free polar representation and applies causal decoding along the range dimension before reconstructing the prediction in Cartesian BEV coordinates. A controlled underwater occupancy benchmark was established, adapting representative occupancy methods to the same RGB to sonar task under a unified protocol. AquaBEV achieves 31.4 Visible IoU and 38.6 Observed IoU, 4.0% and 4.3% relative improvements over the strongest transferred baseline.
Chinese Translation
自主水下机器人被广泛应用于勘探、监测和检查任务,其安全导航依赖于对周围自由空间与占据空间的理解。鸟瞰图(BEV)占据表示提供了这样一种表征方式,但由于仅凭外观获得的几何线索有限且不可靠,从单张水下RGB图像预测BEV占据十分困难。三维成像声呐可提供互补的几何测量,用于监督该任务。我们提出了AquaBEV,一种单目水下占据预测模型,能够从单张RGB图像预测局部BEV占据,并在训练过程中利用成对的三维成像声呐作为几何监督。AquaBEV将视觉特征映射到无需标定的极坐标表示中,沿距离维度应用因果解码,随后在笛卡尔BEV坐标系中重建预测结果。我们建立了一个受控的水下占据基准,在统一协议下将具有代表性的占据方法适配到相同的RGB到声呐任务中。AquaBEV取得了31.4的Visible IoU和38.6的Observed IoU,相对于最强的迁移基线方法分别提升了4.0%和4.3%。
cs.RO / 5 / 2609.04464

Achieving Asymptotic Near-Optimality Without $\delta$-Similarity

在没有$\delta$-相似性条件下实现渐近近似最优性
Moncton, Michael, Frew, Eric
Abstract
Sampling-based motion planning algorithms are a popular class of trajectory planning algorithm due to their speed in complex, high-dimensional environments and ability to handle kinodynamic constraints, specifically through the use of forward dynamics propagation. Many such planners claim to achieve asymptotic near-optimality by proving the almost sure sampling of trajectories that are close to an optimal trajectory in the state space, known as $\delta$-similar trajectories. This paper shows that the proof behind asymptotic $\delta$-similarity relies on an unstated assumption that $\delta$-similar trajectory segments will always be kept once sampled. This assumption does not hold in general. A problematic case, referred to as ``crowding out,'' is described, where locally low-cost paths prevent trajectories that are $\delta$-similar to the optimal trajectory from being added to the tree. It is shown, however, that asymptotic near-optimality guarantees can still be achieved without guarantees of $\delta$-similar solution trajectories when crowding out is properly accounted for. An example environment and system are provided where crowding out is shown to occur, demonstrating a scenario where inductively sampling a $\delta$-similar solution trajectory is impossible.
Chinese Translation
基于采样的运动规划算法是一类流行的轨迹规划算法,因为它们在复杂、高维环境中速度快,并且能够处理动力学约束,特别是通过使用前向动力学传播。许多此类规划器通过证明以概率1采样到状态空间中接近最优轨迹的轨迹(称为$\delta$-相似轨迹)来声称实现渐近近似最优性。本文表明,渐近$\delta$-相似性背后的证明依赖于一个未明确说明的假设,即$\delta$-相似轨迹段一旦被采样就会始终被保留。该假设在一般情况下并不成立。本文描述了一种被称为“挤出”(crowding out)的问题情形,即局部低代价路径阻止了与最优轨迹$\delta$-相似的轨迹被添加到树中。然而,研究表明,在恰当考虑挤出问题的情况下,即使没有$\delta$-相似解轨迹的保证,仍然可以实现渐近近似最优性保证。本文给出了一个发生挤出问题的示例环境和系统,展示了一种无法通过归纳方式采样出$\delta$-相似解轨迹的情形。
cs.RO / 6 / 2609.04545

SocioGesture: Real-Time and Adaptive Social Gesture Perception for Human-Robot Interaction

SocioGesture:面向人机交互的实时自适应社交手势感知系统
Fu, Wenjin, Wu, Li-Fan, Peter, Jerin, Huyen, Chip, Chen, Boyuan, Liphardt, Jan
Abstract
Robots interacting with people must recognize not only explicit commands, but also social cues such as invitations, refusals, and unavailability. In real deployments, these cues must be inferred from noisy onboard perception under partial occlusion, changing viewpoints, and strict latency constraints. We present SocioGesture, a real-time adaptive social gesture perception system for human-robot interaction (HRI). SocioGesture uses a compact confidence-aware body-hand skeleton representation and a lightweight dual-stream model that fuses body motion with hand articulation for low-latency onboard recognition. To improve deployment robustness, we train the model with occlusion-aware skeleton corruption, exposing it to missing hands, occluded arms, and temporally unstable keypoints without increasing the inference cost. On a social gesture dataset collected in mixed indoor-outdoor HRI scenarios, SocioGesture achieves strong held-out-subject recognition, substantially improves robustness under structured joint occlusion, and runs in real time on a robot-mounted edge device. During deployment, uncertain interaction segments are saved for offline labeling and adaptation, enabling SocioGesture to expand its gesture vocabulary while preserving performance in the original classes. These results demonstrate a practical path toward robust, efficient, and adaptive social perception for interactive robots.
Chinese Translation
与人交互的机器人不仅要识别显式指令,还需识别社交线索,如邀请、拒绝和不可用等信号。在实际部署中,这些线索必须在部分遮挡、视点变化和严格延迟约束下,从含噪的机载感知中推断得出。我们提出了SocioGesture,一个面向人机交互(HRI)的实时自适应社交手势感知系统。SocioGesture采用紧凑的置信度感知的躯干-手部骨架表示,以及一个轻量级双流模型,将身体运动与手部关节动作融合,实现低延迟的机载识别。为提升部署鲁棒性,我们在训练中引入遮挡感知的骨架扰动,使模型接触到手部缺失、手臂遮挡和时间上不稳定的关键点等情况,而不增加推理开销。在一个在室内外混合HRI场景中采集的社交手势数据集上,SocioGesture在留出被试识别上取得了优异结果,在结构化关节遮挡下显著提升了鲁棒性,并能在机器人搭载的边缘设备上实时运行。在部署过程中,不确定的交互片段会被保存用于离线标注与自适应更新,使SocioGesture能够扩展其手势词汇量,同时保持对原有类别的性能。这些结果为交互式机器人实现鲁棒、高效且自适应的社交感知提供了一条切实可行的路径。
cs.RO / 7 / 2609.04552

Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI

面向部署后物理人工智能的持续场自适应模型(CFAMs)
Singh, Amarjot, Pancholi, Tanmay R., Kothari, Jainam, Mahajan, Shrirang, Bansal, Ketan, Erickson, Zackory, Loianno, Giuseppe, Bayen, Alexandre M., Schneider, Jeff, Nakayama, Vince
Abstract
Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with human tools - remains a missing capability in mission-critical operations. These domains offer scarce training data and only onboard compute, yet deployed systems must face novelty without erasing prior competence. We introduce Continual Field-Adaptive Models (CFAMs), which learn efficiently in the lab and continue learning after deployment through autonomous, gradient-free, on-device updates. CFAM uses a complementary learning architecture with a frozen slow-learning component and a fast-learning Capsule Field. The slow component contains three cortices: Sensor, which maps multimodal input into 3D-grounded geometry; Reasoning, which decomposes tasks into skills and evaluates outcomes; and Action, which executes geometric skills. The Capsule Field stores field learning one-shot and gradient-free as Competence Capsules. Skill installation is few-shot in the lab and continual in the field; open-world novelty is outside scope. We evaluate CFAM across five embodiments: manipulator, quadruped, humanoid, quadrotor, and off-road vehicle. Baselines (pi0, CogACT, SpatialVLA) use the same in-house multi-embodiment dataset for physical-platform comparisons. CFAM reaches the operating point of a standard policy trained on the full prior-training dataset using 40% of the data, or 2.5x fewer trajectories. At test time, autonomous capture of verified near-OOD cases improves action success by 13.9 percentage points. In sequential simulation, backward transfer is -0.5 percentage points versus -11.4 for LoRA. CFAM therefore provides a bounded form of post-deployment physical intelligence: few-shot skill learning, autonomous field growth from verified near-OOD experience, and retention of prior competence.
Chinese Translation
无人值守的交互式自主性——即机器代替人类进入危险环境并使用人类工具完成任务——仍然是关键任务操作中缺失的一种能力。这些领域提供的训练数据稀缺,且仅有板载计算资源,而部署后的系统必须在面对新情况时不遗忘已掌握的能力。我们提出持续场自适应模型(Continual Field-Adaptive Models,CFAMs),它能够在实验室中高效学习,并在部署后通过自主的、无梯度的设备端更新持续学习。CFAM采用一种互补学习架构,包含一个冻结的慢学习组件和一个快学习的胶囊场(Capsule Field)。慢组件包含三个皮质模块:感知(Sensor)皮质,将多模态输入映射到具有三维基础的几何表示;推理(Reasoning)皮质,将任务分解为技能并评估结果;以及动作(Action)皮质,执行几何技能。胶囊场以一次性(one-shot)、无梯度的方式将现场学习存储为能力胶囊(Competence Capsules)。技能安装是在实验室中少样本完成的,并在现场持续进行;开放世界的新颖性不在本文范围内。我们在五种具身形态上评估CFAM:机械臂、四足机器人、人形机器人、四旋翼飞行器和越野车辆。基线模型(pi0、CogACT、SpatialVLA)使用相同的自有多具身数据集进行物理平台对比。CFAM仅使用40%的数据(即少2.5倍的轨迹),即可达到在全量先前训练数据集上训练的标准策略的运行水平。在测试时,自主采集经核实的近分布外(near-OOD)案例使动作成功率提升13.9个百分点。在序列化仿真中,其向后迁移为-0.5个百分点,而LoRA为-11.4个百分点。因此,CFAM提供了一种有边界的部署后物理智能形式:少样本技能学习、基于经核实的近OOD经验的自主现场增长,以及对先前能力的保持。
cs.RO / 8 / 2609.04602

NavArena: Automated Construction of Goal-Oriented Navigation Benchmarks from 3D Gaussian Splatting Reconstructions

NavArena:基于3D高斯泼溅重建的面向目标导航基准自动构建方法
Wang, Junhui, Yang, Wei, Li, Xinyao, Fan, Ningjing, Yin, Yuehao, Chen, Xuecheng, Gao, Chao
Abstract
Fixed 3D Gaussian Splatting (3DGS) reconstructions provide realistic novel views but lack the traversability constraints, valid goals, and closed-loop protocols required for navigation evaluation. We introduce NavArena, an automated framework that transforms fixed 3DGS reconstructions into benchmarks for goal-oriented visual navigation. NavArena integrates a frozen 3DGS model for egocentric RGB-D rendering, an occupancy costmap derived from Gaussian density and height statistics for reachability and collision queries, and semantic goal candidates lifted from multi-view open-vocabulary masks. These components support the automatic generation and unified closed-loop evaluation of goal-oriented navigation episodes. Across more than 2{,}000 scenes, NavArena generates 22.2 million expert trajectories. Spatial and semantic evaluations assess the derived navigation representations, while policy rollouts demonstrate the diagnostic value of the unified evaluation protocol. NavArena enables scalable and reproducible navigation evaluation on large-scale 3DGS reconstructions, and all benchmark-generation tools, evaluation protocols, and derived assets will be released publicly.
Chinese Translation
固定的3D高斯泼溅(3DGS)重建能够提供逼真的新视角渲染,但缺乏导航评估所需的可通行性约束、有效目标以及闭环评估协议。我们提出了NavArena,这是一个自动化框架,可将固定的3DGS重建转化为面向目标的视觉导航基准。NavArena集成了以下组件:用于第一人称视角RGB-D渲染的冻结3DGS模型、基于高斯密度和高度统计信息推导的占据栅格代价地图(用于可达性与碰撞查询),以及从多视角开放词汇掩码中提取的语义目标候选。这些组件支持面向目标导航场景的自动生成与统一闭环评估。在超过2,000个场景中,NavArena生成了2220万条专家轨迹。空间与语义评估用于检验所导出的导航表征,而策略滚动测试则展示了统一评估协议的诊断价值。NavArena实现了大规模3DGS重建上可扩展、可复现的导航评估,所有基准生成工具、评估协议及衍生资产都将公开发布。
cs.RO / 9 / 2609.04607

Open-Set 3D Scene Graphs for Field Robotics: An Outdoor Case Study

面向野外机器人的开放集三维场景图:一项户外案例研究
Samuelson, Chad R., Slade, Gabriel R., Mangelson, Joshua G.
Abstract
Three-dimensional scene graphs (3DSGs) have emerged as a promising approach for building geometrically grounded, semantically informed, hierarchical general-purpose maps to support high-level robotic reasoning. However, the behavior of 3DSGs in real-world outdoor deployments remains poorly understood, particularly when combined with open-set vision-language models (VLMs). In this field report, we analyze the components common to most 3DSG representations across five outdoor robotic datasets to characterize challenges that arise in complex outdoor environments. Using the recently proposed Terra 3DSG as a case study, we investigate semantic point embeddings, place-node graph navigation, region-level understanding, and memory size across the five diverse datasets. We additionally introduce novel consistency metrics to evaluate whether semantic and structural graph properties remain stable across repeated traversals of the same environment. Our analysis reveals that outliers and multiple modes are common in VLM point embeddings across all tested datasets with outlier ratios above $0.1$ for around $30\%$ of points. We demonstrate the feasibility of outdoor 3DSGs for navigation-based object retrieval, achieving success rates near $70\%$, though performance is limited by traversability failures and inefficient routing, with trajectories averaging approximately $66\%$ suboptimal path efficiency. Region-level understanding remains challenging in complex natural environments, with low average F1 scores around $0.359$. Overall, our results show that outdoor 3DSGs can maintain compact (less than $600$MB for multi-kilometer trajectories) and relatively consistent large-scale environment representations, while highlighting open challenges in handling multiple semantic modes, incorporating traversability into graph structures, and improving higher-level region understanding.
Chinese Translation
三维场景图(3DSG)已成为构建几何定位、语义丰富、分层通用地图的一种有前景的方法,可支持高层次的机器人推理。然而,3DSG 在真实户外部署中的表现仍缺乏充分理解,尤其是与开放集视觉-语言模型(VLM)结合时。在本领域报告中,我们分析了大多数 3DSG 表示中常见的组件,并在五个户外机器人数据集上进行评估,以刻画复杂户外环境中出现的挑战。以最近提出的 Terra 3DSG 为案例研究,我们在这五个多样化数据集上考察了语义点嵌入、位置节点图导航、区域级理解以及内存占用。此外,我们提出了新颖的一致性指标,用于评估语义和结构图属性在同一环境的重复遍历中是否保持稳定。我们的分析表明,在所有测试数据集中,VLM 点嵌入中普遍存在离群点和多模态现象,约 30% 的点的离群比例超过 0.1。我们证明了户外 3DSG 用于基于导航的目标检索的可行性,成功率接近 70%,但性能受限于可通行性失败和低效的路径规划,轨迹的平均次优路径效率约为 66%。在复杂自然环境中,区域级理解仍然具有挑战性,平均 F1 分数较低,约为 0.359。总体而言,我们的结果表明,户外 3DSG 能够维持紧凑(多公里轨迹下小于 600MB)且相对一致的大规模环境表示,同时也凸显了在处理多语义模态、将可通行性纳入图结构以及提升更高层次区域理解方面尚存的开放性挑战。
cs.RO / 10 / 2609.04620

Pack It My Way: Triadic Human-Robot Collaboration for Personalized Autonomous Packing

按我的方式打包:面向个性化自主打包的三元人机协作
Kotapati, Sandeep Chowdary, Gao, Yanxin, Lin, Tsung-Chi
Abstract
Personalized autonomous packing requires robots to account for resident preferences that cannot be inferred from scene geometry alone. Expert teleoperators can interpret these preferences and translate them into feasible robot actions, but continuous expert involvement limits scalable deployment. In this paper, we investigate triadic human-robot collaboration among a resident, a correction mediator, and a robot by comparing human-expert and voice-agent mediation. We evaluate the two conditions in a user study across Protection, Compactness, andGrouping tasks, using a Show-Correct-Generalize process to assess preference correction and subsequent generalization after the surrounding objects are rearranged. Results show that voice-agent mediation achieves outcomes comparable to human-expert mediation in two of the three preference categories, despite receiving shorter and less detailed instructions. Both mediators are similarly easy to use, although the human expert is perceived as more reliable. These findings demonstrate the potential of voice agents to reduce expert involvement while identifying perceived reliability and preference generalization as remaining challenges.
Chinese Translation
个性化自主打包要求机器人能够考虑仅凭场景几何无法推断的居民偏好。专业远程操作员可以解读这些偏好并将其转化为可行的机器人动作,但持续依赖专家参与限制了可扩展部署。本文研究了由居民、纠错中介者和机器人组成的三元人机协作,并对比了人类专家中介与语音代理中介两种方式。我们在一项用户研究中,围绕保护性(Protection)、紧凑性(Compactness)和分组(Grouping)三类任务对两种条件进行评估,采用“演示—纠错—泛化”(Show-Correct-Generalize)流程来评估偏好纠错效果以及周围物体重新排列后的泛化能力。结果表明,尽管语音代理接收到的指令更简短、细节更少,但其在三类偏好类别中的两类上取得了与人类专家中介相当的效果。两种中介方式在易用性上相近,不过人类专家被认为更可靠。这些发现证明了语音代理在减少专家参与方面的潜力,同时指出感知可靠性与偏好泛化仍是亟待解决的挑战。
cs.RO / 11 / 2609.04759

Dressing in Motion: A Human Motion-Aware Diffusion Policy for Robot-Assisted Dressing

动态穿衣:一种面向机器人辅助穿衣的人体运动感知扩散策略
Sun, Haoxiang, Wang, Fangyuan, Huang, Songhao, Liu, Justina Y. W., Zhu, Jihong, Zhou, Peng, Navarro-Alarcon, David
Abstract
Robotic dressing assistance is a promising solution for supporting older adults with physical impairments in daily living. However, dressing under human motion remains challenging, as complex garment--human contact and occlusions make it difficult to generate actions aligned with arm movements. In this letter, we propose a visuomotor policy that learns dressing skills from static expert demonstrations and generalizes to dynamic user-motion scenarios. A diffusion policy tailored to garment--human interaction geometry learns from partially observed point clouds with varied arm postures. We then introduce an object-centric representation based on PDE diffusion to capture the axial distribution of the arm. By sampling motion-relevant regions and registering them across consecutive observations, the proposed method approximates arm motion and reactively adapts the executed trajectory. We evaluate our method in simulation and a real-world human study involving nine participants, three garment types, and six arm-motion patterns. Results show that our method outperforms baselines in dressing progress, freedom of movement, and user comfort. The project website is https://anonymous.4open.science/w/dressing-in-motion.
Chinese Translation
机器人穿衣辅助是支持身体障碍老年人在日常生活中的一种有前景的解决方案。然而,在人体运动过程中进行穿衣仍然具有挑战性,因为复杂的衣物-人体接触与遮挡使得生成与手臂运动相一致的动作变得困难。在本信中,我们提出了一种视觉运动策略,该策略从静态专家演示中学习穿衣技能,并能泛化到用户动态运动的场景。我们设计了一种针对衣物-人体交互几何特性的扩散策略(diffusion policy),该策略从具有不同手臂姿态的部分可观测点云中学习。随后,我们引入了一种基于PDE扩散(PDE diffusion)的以物体为中心的表示方法,以捕捉手臂的轴向分布。通过采样与运动相关的区域并在连续观测之间进行配准,所提出的方法能够估计手臂运动并反应性地调整执行轨迹。我们在仿真以及一项包含九名参与者、三种衣物类型和六种手臂运动模式的真实人体研究中对该方法进行了评估。结果表明,我们的方法在穿衣进度、活动自由度和用户舒适度方面均优于基线方法。项目网站为 https://anonymous.4open.science/w/dressing-in-motion。
cs.RO / 12 / 2609.04770

Continuous Cognitive Coverage for Autonomous Robots via Event-Dependent Cognitive Treatment and Learning

基于事件依赖认知处理与学习的自主机器人持续认知覆盖
Su, Hong
Abstract
Autonomous robots continuously encounter objects, changes, and situations, and every event admitted into cognition should receive an appropriate cognitive treatment rather than remain untreated until an explicit task requires attention. However, existing task-driven, reactive, or fixed-reasoning approaches generally process only selected events or apply predefined reasoning procedures, making it difficult to provide continuous cognitive coverage with differentiated treatment. This paper proposes a continuous cognitive coverage framework in which every cognitively admitted event is assigned an event-dependent cognitive treatment according to its state, context, and history. Different events may therefore invoke description, memory, risk prediction, planning, diagnosis, analogy, or other learned treatments. Familiar events can be processed automatically by learned mechanisms, whereas unfamiliar or uncertain events invoke explicit deliberation or fallback reasoning. Multiple cognitive processes can be suspended, resumed, and interleaved so that cognitive processing continues as new events arrive or existing events await evidence. Validated experiences are continuously learned to automate, refine, and revise event-specific treatments. Experiments achieve 96.76% structured treatment accuracy with 93.66% automatic processing, 92.64% cognitive coverage under bursty-delayed workloads, and 79.53% continual-learning joint accuracy, with novel-event reuse reaching 100% automatic processing.
Chinese Translation
自主机器人会持续遇到各种对象、变化和情境,每一个被纳入认知的事件都应当得到适当的认知处理,而不是在被明确任务要求关注之前一直处于未处理状态。然而,现有的任务驱动、反应式或固定推理方法通常仅处理被选择的事件,或采用预定义的推理流程,难以提供具有差异化处理的持续认知覆盖。本文提出一种持续认知覆盖框架,其中每个被认知系统接纳的事件都根据其状态、上下文和历史被赋予事件依赖的认知处理。因此,不同事件可调用描述、记忆、风险预测、规划、诊断、类比或其他习得的处理机制。熟悉的事件可通过习得机制自动处理,而陌生或不确定的事件则调用显式深思(deliberation)或回退推理。多个认知过程可以被挂起、恢复和交错执行,从而在新事件到达或已有事件等待证据时,认知处理得以持续进行。系统持续学习经过验证的经验,以自动化、细化和修订针对特定事件的处理方式。实验取得了96.76%的结构化处理准确率(其中93.66%为自动处理)、突发延迟负载下92.64%的认知覆盖率以及79.53%的持续学习联合准确率,其中新事件复用达到100%的自动处理率。
cs.RO / 13 / 2609.04799

HaptiNet: Networked Haptic Robots Enable Physical Co-presence in Geographically-Unconstrained Rehabilitation

HaptiNet:网络化触觉机器人在无地理限制的康复中实现物理共在
Sun, Chenyang, Dong, Mingjie, Deng, Haodong, Liu, Yudong, Chen, Yi-Feng, Lin, Jun, Huang, Changlong, Guo, Jie, Liu, Yantong, Liu, Yang, Lin, Yuzhou, Long, Jianjun, Xing, Zheng, Zhao, Sining, Zhang, Xuemin, Wang, Zhiyong, Li, Zhenhong, Wu, Dongrui, Liu, Honghai, Dai, Jian S., Zhang, Mingming
Abstract
Cooperative rehabilitation enhances engagement, task performance, and social-motor interaction, yet it demands physical co-presence: users must transmit forces, coordinate movements, and infer intent through haptic contact. Telerehabilitation promises to expand access for patients constrained by distance, mobility, or clinical disparities, yet current techniques remain predominantly audiovisual while leaving users haptically and physically isolated. Here, we introduce HaptiNet, a networked haptic robotic system enabling physical co-presence for geographically distributed users via force-mediated interaction. Each robotic terminal features a low-inertia, long-stroke design with high force-feedback capacity, tailored for haptic rendering in upper-limb training. Building on these terminals, HaptiNet creates a distributed haptic network with an imitation-learning-based delay compensator, enabling users to physically perceive and coordinate with one another over distance. We validated HaptiNet in 284 healthy participants and 111 patients with neurological impairments across progressively realistic settings, including laboratory tests, cross-city deployments, and clinical applications. HaptiNet preserved task-level force rendering consistency across single-user and multi-user scenarios. Compared with solo and visual cooperative training, haptic cooperation improved task performance by 24% and 22%, respectively, while also boosting engagement and interpersonal motor synchrony. Across three intercity links totaling approximately 4,000 km, HaptiNet maintained stable haptic interaction among patients with neurological impairments, producing a 3.87-fold greater baseline-to-training score improvement and a 106% higher patient-applied effort over the solo condition.
Chinese Translation
协作式康复能够提升参与度、任务表现和社会性运动交互,但其依赖于物理共在:使用者必须通过触觉接触传递力、协调动作并推断意图。远程康复有望为受距离、行动能力或临床医疗不均所限制的患者扩大可及性,然而现有技术仍以视听方式为主,使用者在触觉和身体层面处于孤立状态。本文提出HaptiNet——一种网络化触觉机器人系统,通过力介导的交互使地理上分散的使用者实现物理共在。每个机器人终端采用低惯量、长行程且具备高力反馈能力的设计,专为上肢训练中的触觉渲染而定制。基于这些终端,HaptiNet构建了一个分布式触觉网络,并配备基于模仿学习的时延补偿器,使使用者能够在远距离下彼此物理感知并协同配合。我们在284名健康受试者和111名神经功能障碍患者中,于逐步接近真实场景的多种环境下验证了HaptiNet,包括实验室测试、跨城市部署以及临床应用。HaptiNet在单用户与多用户场景中均保持了任务级力渲染的一致性。与单人训练和视觉协作训练相比,触觉协作分别将任务表现提升了24%和22%,同时增强了参与度和人际运动同步性。在总里程约4,000公里的三条城际链路上,HaptiNet为神经功能障碍患者维持了稳定的触觉交互,其基线到训练得分的提升幅度是单人训练条件的3.87倍,患者施加的努力程度高出106%。
cs.RO / 14 / 2609.04807

CoLMIN: LLM-based Multi-Decision Path Negotiation for Cooperative Autonomous Driving

CoLMIN:基于大语言模型的协同自动驾驶多决策路径协商方法
Huang, Zhe, Fan, Zhaoxin, Wang, Shuo, Wu, Wenjun, Zhao, Xuan, Liu, Min
Abstract
Multi-vehicle cooperative autonomous driving enhances the safety and reliability of autonomous driving systems through information sharing among connected vehicles, demonstrating significant potential for improving traffic safety. LLM-based approaches leverage strong reasoning capabilities of LLMs to enable effective inter-vehicle negotiation and improve cooperative driving performance. However, driving decisions in complex traffic scenarios are inherently multi-solution in nature. As a result, existing negotiation-based methods often converge prematurely to suboptimal solutions, hindering consensus formation and limiting the practical deployment of cooperative autonomous driving systems. To address this challenge, we propose CoLMIN, the LLM-based multi-decision path negotiation framework for cooperative autonomous driving, achieving stable decision consensus through multi-decision path negotiation and reflective reasoning. To achieve stable and high-quality consensus in cooperative autonomous driving, CoLMIN consists of three key components: (i) an LLM-based Multi-Intent Negotiation module (LMin), which adopts a Negotiator-Evaluator paradigm and generates multiple candidate driving intentions for joint evaluation; (ii) an Evaluation-based Shallow Reflection Module (ESRM), which analyzes negotiation outcomes and provides feedback to guide subsequent negotiations, thereby accelerating consensus formation; and (iii) an LLM-based Deep Reflection Module (LDRM), which performs long-term reflection over negotiation histories to mitigate cognitive fixation and prevent the system from converging to suboptimal solutions. Experimental results in the CARLA simulation environment demonstrate that CoLMIN significantly outperforms existing methods in challenging interactive driving scenarios.
Chinese Translation
多车协同自动驾驶通过网联车辆之间的信息共享提升自动驾驶系统的安全性与可靠性,在改善交通安全方面展现出巨大潜力。基于大语言模型(LLM)的方法利用LLM强大的推理能力实现有效的车间协商,从而提升协同驾驶性能。然而,复杂交通场景中的驾驶决策本质上具有多解性。因此,现有的基于协商的方法往往过早收敛于次优解,阻碍了共识的形成,限制了协同自动驾驶系统的实际部署。为应对这一挑战,我们提出了CoLMIN,一种面向协同自动驾驶的基于LLM的多决策路径协商框架,通过多决策路径协商与反思性推理实现稳定的决策共识。为了在协同自动驾驶中获得稳定且高质量的共识,CoLMIN包含三个关键组件:(i)基于LLM的多意图协商模块(LMin),采用“协商者-评估者”范式,生成多个候选驾驶意图以进行联合评估;(ii)基于评估的浅层反思模块(ESRM),分析协商结果并提供反馈以引导后续协商,从而加速共识形成;(iii)基于LLM的深层反思模块(LDRM),对协商历史进行长期反思,以缓解认知固着并防止系统收敛至次优解。在CARLA仿真环境中的实验结果表明,CoLMIN在具有挑战性的交互式驾驶场景中显著优于现有方法。
cs.RO / 15 / 2609.04817

Strict Modes Everywhere - Bringing Order Into Dynamics of Mechanical Systems by a Potential Compatible With the Geodesic Flow

处处皆严格模态——通过与测地流相容的势能为机械系统动力学带来秩序
Sachtler, Arne, Albu-Schäffer, Alin
Abstract
Strict nonlinear normal modes provide very regular families of oscillations within conservative mechanical systems. However, a strict normal mode will generally be an isolated curve within the configuration space of the system. In this letter, we design a potential that will densely fill the configuration space with strict normal modes such that each configuration belongs to one mode and each mode passes through a common point, the equilibrium. As the potential can be realized by (nonlinear) elastic elements it can be used to execute a variety of periodic trajectories very efficiently. Most of the required torques will come from the elastic elements in the system and not from the actuators. We also design a controller stabilizing the system to a desired target mode and a controller performing swing-up and compensating dissipated energy. Finally, we showcase the approach for a two DoF manipulator. The experiments show that the approach performed well for the example system.
Chinese Translation
严格非线性正规模态为保守机械系统提供了非常规则的振动族。然而,严格正规模态通常只是系统位形空间中的一条孤立曲线。在本信函中,我们设计了一种势能,使严格正规模态稠密地充满位形空间,使得每个位形都属于某一模态,且每个模态都经过一个公共点——平衡点。由于该势能可以通过(非线性)弹性元件实现,因此可以非常高效地执行多种周期轨迹。所需的力矩大部分来自系统中的弹性元件,而非执行器。我们还设计了一个控制器,将系统稳定到期望的目标模态,以及一个执行摆起并补偿能量耗散的控制器。最后,我们以二自由度机械臂为例展示了该方法。实验结果表明,该方法在该示例系统中表现良好。
cs.RO / 16 / 2609.04851

Coupled Control and Wireless World Models for Resilient Remote Robotic Control

面向弹性远程机器人控制的控制与无线世界模型耦合方法
Madushanka, H. P., Samarakoon, Sumudu, Bennis, Mehdi
Abstract
Remote robotic systems operating over wireless networks must maintain reliable control despite limited communication resources, changing channel conditions, and environmental disturbances.However, continuously transmitting high-dimensional sensory observations, such as camera images, increases communication overhead and energy consumption while reducing robustness under unreliable connectivity.To address these challenges, this paper proposes a resilient communication-aware remote robotic control framework based on coupled control and wireless Joint Embedding Predictive Architecture (JEPA) world models that jointly capture robot dynamics and wireless channel evolution from visual observations and a combination of raw and structured radio frequency (RF) representations based on spectrograms and Persistence Images(PIs).The learned latent representations enable predictive communication scheduling by jointly forecasting future robot states and wireless conditions, thereby reducing unnecessary uplink transmissions while maintaining reliable control performance.Furthermore, an adaptive resilience mechanism detects latent prediction discrepancies and efficiently adapts perception embeddings to accommodate wireless and visual environmental changes without retraining the complete control policy.The proposed framework is evaluated in a synchronized Gazebo-Robot Operating System (ROS)-Sionna robot-wireless simulation environment under diverse wireless propagation and perception perturbations.Experimental results demonstrate significant improvements in communication efficiency, robustness, and resilience while maintaining navigation performance compared with conventional Proportional Integral Derivative (PID), model-free Deep Q-Network (DQN), and predictive approaches based on Vision Transformers(ViTs).
Chinese Translation
在无线网络上运行的远程机器人系统,必须在通信资源受限、信道条件变化以及环境干扰的情况下保持可靠控制。然而,持续传输相机图像等高维感知观测数据会增加通信开销和能耗,并降低在不可靠连接下的鲁棒性。为应对这些挑战,本文提出一种基于控制与无线联合嵌入预测架构(Joint Embedding Predictive Architecture, JEPA)世界模型相耦合的弹性通信感知远程机器人控制框架,该框架从视觉观测以及基于频谱图和持久性图像(Persistence Images, PIs)的原始与结构化射频(RF)表征组合中,联合捕捉机器人动力学与无线信道演化。所学到的潜在表征通过联合预测未来机器人状态和无线状况,实现预测性通信调度,从而在保持可靠控制性能的同时减少不必要的上行传输。此外,一种自适应弹性机制可检测潜在预测偏差,并高效调整感知嵌入以适应无线和视觉环境的变化,而无需重新训练完整的控制策略。该框架在同步的 Gazebo-机器人操作系统(ROS)-Sionna 机器人-无线联合仿真环境中,于多种无线传播和感知扰动条件下进行了评估。实验结果表明,与传统的比例积分微分(PID)控制、无模型深度Q网络(DQN)以及基于视觉Transformer(ViT)的预测方法相比,该框架在保持导航性能的同时,在通信效率、鲁棒性和弹性方面均有显著提升。
cs.RO / 17 / 2609.04893

Reasoning Without Inference Cost: Latent Semantic Scaffolding for Robot VLA Policies

无需推理成本的推理:面向机器人视觉-语言-动作策略的潜在语义脚手架
Li, Andrew Ting Yan, Li, Zhuo, Yang, Zhelin, Dong, Zhipeng, Rouxel, Quentin, Chen, Fei
Abstract
Vision-language-action (VLA) models are trained by imitation and capture what action to take but not why; adding causal reasoning improves manipulation, but current methods pay for it at inference time - generating reasoning tokens or rolling out predicted future states at every step, a cost that compounds over long horizons. We ask whether this benefit can instead be captured during training and discarded before deployment. We introduce Latent Semantic Scaffolding (LSS), an auxiliary loss applied during human-demonstration pretraining that aligns a VLA's action-token representations to text embeddings of physical-reasoning rationales through a small projection head. The head is dropped at inference, leaving the unmodified base policy with zero added cost. Our central finding concerns alignment granularity: aligning each action token to the rationale of its own manipulation phase (Dense LSS) rather than to a single pooled episode-level embedding (Pooled LSS) yields representations that transfer markedly better to held-out tasks. Dense LSS attains both the best in-distribution success and the best transfer to tasks unseen during alignment, whereas pooled alignment over-specializes to the training task. A representational probe shows Dense LSS induces roughly twice the per-phase separability in the backbone, supporting that phase-local alignment is the operative mechanism.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型通过模仿学习训练,能够捕捉执行什么动作,却无法理解为什么执行;引入因果推理可以提升操作性能,但现有方法在推理时需要付出额外代价——每一步都生成推理 token 或展开预测的未来状态,这种代价在长时序任务中会不断累积。我们探讨能否在训练阶段获取这一收益,并在部署前将其舍弃。我们提出潜在语义脚手架(Latent Semantic Scaffolding, LSS),这是一种在人类演示预训练期间施加的辅助损失,通过一个小的投影头将 VLA 的动作 token 表征与物理推理依据的文本嵌入对齐。该投影头在推理时被丢弃,使未经修改的基础策略以零额外成本运行。我们的核心发现关乎对齐粒度:将每个动作 token 对齐到其所属操作阶段的推理依据(Dense LSS),而非对齐到单一的池化片段级嵌入(Pooled LSS),所产生的表征在未见任务上的迁移效果显著更好。Dense LSS 同时在分布内任务上取得最佳成功率,并在未参与对齐的任务上取得最佳迁移表现;而池化对齐则过度专门化于训练任务。表征探测实验表明,Dense LSS 使骨干网络中每个阶段的可分性提升约一倍,证明阶段局部对齐是起作用的关键机制。
cs.RO / 18 / 2609.05084

ToPos: Automated Optimal Positioning on Topographic Manifolds using Constrained Geodesic Voronoi Decomposition

ToPos:基于约束测地线Voronoi分解的地形流形自动最优布点方法
Raveendran, Rajesh, Vanhamaa, Akseli, Suutala, Jaakko, Tikanmäki, Antti, Röning, Juha
Abstract
Reliable autonomous mapping, environmental sampling, last-mile logistics, and infrastructure deployment depend on the optimal surface area-balanced distribution of Spatial Reference Sites (SRS). Conventional 2D Euclidean methods often fail in high-relief environments by neglecting topographic variations and physical obstructions. This leads to significant planimetric distortion, spatial clustering, and the placement of targets in inaccessible or shadowed regions, compromising both data integrity and operational safety. This paper introduces ToPos, an automated framework for TOPography-aware Optimal Sampling on topographic manifolds. We treat the terrain as a discrete 2-dimensional manifold embedded in 3D Euclidean space and replace standard flat-map distances with non-Euclidean geodesic distances that follow the actual surface geometry. The point distribution is formulated as an optimization problem using a Constrained Geodesic Voronoi Decomposition, solved via a Riemannian Nesterov Accelerated Gradient (NAG) engine. Our approach restricts target locations to a feasible "safe zone," accounting for non-traversable slopes, vegetation, environmental occlusions, etc. Through evaluations on non-convex sinusoidal manifolds, we show that ToPos mitigates planimetric distortion by utilizing geodesic metrics. This approach results in a $\sim$74% improvement in optimal surface area-balanced distribution, as measured by the coefficient of variation (CV) of the Voronoi cell areas. The framework is architected as a Geographic Information System (GIS)-ready micro-service to bolster the mentioned applications. Index Terms: Topographic Manifolds, Geodesic Voronoi Decomposition, Infrastructure Deployment, 3D Mapping, Spatial Sampling, and Non-Euclidean Optimization.
Chinese Translation
可靠的自主建图、环境采样、最后一公里物流以及基础设施部署均依赖于空间参考站点(Spatial Reference Sites, SRS)在表面积上的最优均衡分布。传统二维欧氏方法在高起伏环境中往往失效,因为其忽略了地形变化和物理障碍。这会导致显著的平面投影畸变、空间聚集,以及目标被布置在不可达或被遮挡区域,从而损害数据完整性和运行安全性。本文提出ToPos,一个面向地形流形的、地形感知的最优采样(TOPography-aware Optimal Sampling)自动化框架。我们将地形视为嵌入在三维欧氏空间中的离散二维流形,并用遵循实际表面几何形状的非欧氏测地线距离替代标准的平面地图距离。点位分布问题被建模为基于约束测地线Voronoi分解(Constrained Geodesic Voronoi Decomposition)的优化问题,并通过黎曼Nesterov加速梯度(NAG)求解器进行求解。我们的方法将目标位置限制在可行的“安全区域”内,考虑了不可通行的坡度、植被、环境遮挡等因素。通过对非凸正弦流形的评估,我们表明ToPos利用测地线度量有效缓解了平面投影畸变。以Voronoi单元面积的变异系数(CV)衡量,该方法使表面积最优均衡分布改善了约74%。该框架被构建为可直接集成地理信息系统(GIS)的微服务,以支撑上述应用。关键词:地形流形、测地线Voronoi分解、基础设施部署、三维建图、空间采样、非欧氏优化。
cs.RO / 19 / 2609.05133

A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning

一种模式界定的语言模型:在不破坏局部学习稳定性的前提下优化机器人策略
Dong, Chongwen, Saint-Germain, Mithun Paul, Asif, Pinjari, daCunha, Carlo R.
Abstract
This paper addresses navigation by composite heterogeneous robots in a decentralized system when policy reasoning and local control operate at different update levels. In a NetLogo--Python implementation, three robots share motion dynamics but use different LLM backends. Each robot independently combines a large language model (LLM) policy agent, an Upper Confidence Bound (UCB) bandit, and a Double Deep Q-Network (Double DQN) controller; no central LLM generates team actions. LLM inference is confined to round-level policy generation and refinement rather than tick-level action selection. The robots perform cross-LLM communication through a shared round summary containing policies, outcomes, and learning feedback. UCB performs refinement-mode selection, and the policy-conditioned Double DQN performs tick-level action selection from navigation variables, active policy parameters, and the LLM action prior. Each of the four configurations was evaluated over 30 rounds. In the fixed simulation, the complete configuration reached the goal in all 90 correlated robot--round records and achieved the lowest median completion time (42 ticks) and P90 (73.2 ticks); its median was 25.0--39.1\% lower than those of the other configurations. These observations provide descriptive, configuration-level evidence from the evaluated configurations.
Chinese Translation
本文研究在策略推理与局部控制运行于不同更新层级的条件下,由异构复合机器人在去中心化系统中进行导航的问题。在NetLogo--Python实现中,三台机器人共享运动动力学,但使用不同的大语言模型(LLM)后端。每台机器人独立地结合一个LLM策略智能体、一个置信度上界(UCB)老虎机和一个双深度Q网络(Double DQN)控制器;没有任何中央LLM生成团队动作。LLM推理仅限于回合级(round-level)的策略生成与优化,而非逐时间步(tick-level)的动作选择。机器人通过共享的回合摘要(包含策略、结果和学习反馈)进行跨LLM通信。UCB负责优化模式选择,而策略条件化的Double DQN根据导航变量、激活的策略参数以及LLM动作先验执行逐时间步的动作选择。四种配置均在30个回合上进行评估。在固定仿真中,完整配置在全部90条相关机器人--回合记录中均到达目标,并取得了最低的中位完成时间(42个时间步)和P90(73.2个时间步);其中位数比其他配置低25.0--39.1%。这些观察为所评估的各配置提供了描述性、配置层面的证据。
cs.RO / 20 / 2609.05178

LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models

LIBERO-RECOVER:超越任务成功,面向机器人操作模型中的失败恢复
Liu, Lin, Bao, Zhicheng, Zhang, Lu, Song, Ziying, Yang, Wu, Tao, Shuai, Liu, Wulong, Lu, Huchuan
Abstract
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in robotic manipulation. On LIBERO, SOTA method have achieved nearly 100\% success rates, seemingly suggesting that the models are ready for deployment in real world. However, near perfect performance on existing benchmarks can be misleading: success under ideal conditions does not imply real world robustness. Existing benchmarks primarily evaluate task completion from predefined initial states, while real world interactions inevitably involve failures such as failed grasps, collisions, and unintended object movements. A robot must therefore not only execute tasks successfully, but also recognize and recover from failures to continue the task. Yet this capability remains largely unmeasured, revealing a critical gap between benchmark performance and real world reliability. To address this gap, we introduce LIBERO-Recover Benchmark, a large scale benchmark for failure recovery in robotic manipulation. Built upon LIBERO, we collect real execution failures from SOTA embodied models and construct 1,000+ scenarios across four recovery levels: (1) Action Retry, (2) Action Adaptation, (3) Object State Recovery, and (4) Environmental Recovery. We evaluate four core capabilities: spatial understanding, object structure reasoning, interaction understanding, and topological reasoning. As the first large-scale benchmark for embodied failure recovery, LIBERO-Recover shifts evaluation from \emph{Can the robot succeed?''} to \emph{Can the robot recover after failure?''}, promoting robust and generalizable embodied agents. The project will be avaible in \textcolor{blue}{https://liulin815.github.io/LIBERO-Recovery/}.
Chinese Translation
视觉-语言-动作(VLA)模型或世界动作模型(WAM)近期在机器人操作任务中展现出卓越的性能。在 LIBERO 基准上,最先进(SOTA)方法已取得接近 100% 的成功率,这似乎表明这些模型已具备在真实世界中部署的条件。然而,在现有基准上近乎完美的表现可能具有误导性:理想条件下的成功并不意味着真实世界中的鲁棒性。现有基准主要从预定义的初始状态评估任务完成情况,而真实世界的交互不可避免地会涉及各种失败,如抓取失败、碰撞以及物体的意外移动。因此,机器人不仅需要成功执行任务,还必须能够识别并从失败中恢复以继续完成任务。然而,这一能力在很大程度上尚未被测量,揭示了基准性能与真实世界可靠性之间的关键差距。为填补这一空白,我们提出了 LIBERO-Recover 基准,一个用于机器人操作失败恢复的大规模基准。基于 LIBERO 构建,我们收集了来自 SOTA 具身模型的真实执行失败案例,并构建了涵盖四个恢复层级的 1000 多个场景:(1)动作重试,(2)动作调整,(3)物体状态恢复,以及(4)环境恢复。我们评估四项核心能力:空间理解、物体结构推理、交互理解和拓扑推理。作为首个大规模具身失败恢复基准,LIBERO-Recover 将评估从“机器人能否成功?”转变为“机器人能否在失败后恢复?”,从而推动鲁棒且可泛化的具身智能体的发展。项目将发布于 https://liulin815.github.io/LIBERO-Recovery/。
cs.RO / 21 / 2609.05206

Morphology and actuation as inductive biases in robotic hand manipulation

形态学与驱动作为机器人手部操作中的归纳偏置
Tari, Zalán, Birtalan, Eszter, Polcz, Péter, Koller, Miklós
Abstract
Robotic hands vary widely in anatomical fidelity and mechanical complexity, and these structural choices influence the coordination of joint motions and the difficulty of controlling the system. A unified framework is presented in which the kinematic and actuation stages are analysed separately and in composition, through the conditioning of the task Jacobian, the actuation matrix, and their product. It is applied to two hands representing opposing design philosophies, the Shadow Dexterous Hand and the Anatomically Correct, Biomechatronic Hand, along four morphological aspects: joint axis geometry, actuator-to-DOF ratio, coupling architecture, and authority distribution. All parameters are derived from the hands' canonical digital representations. Anatomical fidelity carries no uniform advantage: oblique axes improve thumb conditioning but leave the long fingers worse conditioned than the orthogonal-axis design, while the branching tendon network improves the effective control mapping at every long finger and worsens it significantly at the thumb, where actuator authority is concentrated on thumb opposition. Predictions derived from these metrics are evaluated against reinforcement learning experiments using PPO, DDPG+HER, and TQC+HER, across three different tasks.
Chinese Translation
机器人手在解剖学保真度和机械复杂度上差异很大,这些结构选择会影响关节运动的协调性以及系统控制的难度。本文提出了一个统一框架,通过任务雅可比矩阵(task Jacobian)、驱动矩阵(actuation matrix)及其乘积的条件数,分别及组合地分析运动学阶段和驱动阶段。该框架被应用于代表两种对立设计理念的两只手——Shadow Dexterous Hand 和 Anatomically Correct, Biomechatronic Hand——并从四个形态学方面进行考察:关节轴几何结构、驱动器与自由度之比、耦合架构以及控制权限分布。所有参数均由各手的规范数字化表示推导得出。研究表明,解剖学保真度并不具有一致的优势:倾斜关节轴改善了拇指的条件数,却使长手指的条件数劣于正交轴设计;而分支式肌腱网络改善了每根长手指的有效控制映射,却在拇指处显著恶化,因为驱动权限集中于拇指对掌功能。基于这些指标得出的预测通过使用 PPO、DDPG+HER 和 TQC+HER 的强化学习实验,在三项不同任务中进行了验证。
cs.RO / 22 / 2609.05260

One Word, Different Action: A Real-Robot Benchmark for Language-Conditioned Embodied Reasoning

一词之差,行为迥异:面向语言条件具身推理的真机机器人基准
Liu, Yiwei, Yang, Luwei, Lei, Shunbo
Abstract
Natural-language instruction changes can directly alter robot behavior. A reliable embodied system should preserve its action when the task is unchanged and update it correctly when the task itself changes. We introduce One Word, Different Action, a real-robot benchmark built on physical decision states and executable actions, using task-preserving and task-changing instruction pairs to jointly evaluate Decision Invariance and Decision Sensitivity, with further evaluation under multi-constraint reasoning and real-RGB grounding. Experiments show that modern models are near saturation on single-constraint instruction changes, yet several models degrade noticeably when multiple task constraints must be integrated into one executable decision. These results suggest that the more salient remaining challenge is no longer recognizing an isolated instruction change, but reliably composing multiple task requirements into a correct robot action decision.
Chinese Translation
自然语言指令的变化会直接改变机器人的行为。一个可靠的具身系统应当在任务未变时保持其动作不变,而在任务本身发生变化时正确地更新动作。我们提出了One Word, Different Action,一个基于物理决策状态和可执行动作的真机机器人基准。该基准利用任务保持型和任务变更型指令对,联合评估决策不变性(Decision Invariance)与决策敏感性(Decision Sensitivity),并在多约束推理和真实RGB视觉定位条件下进行进一步评估。实验表明,现代模型在单约束指令变化上已接近饱和,但当需要将多个任务约束整合为一个可执行决策时,若干模型出现明显性能下降。这些结果表明,当前更突出的剩余挑战不再是识别孤立的指令变化,而是可靠地将多个任务要求组合成正确的机器人动作决策。
cs.RO / 23 / 2609.05266

TacPAC: Tactile Prediction and Real-Time Action Correction in World-Action Models for Contact-Rich Manipulation

TacPAC:面向接触密集型操作的世界-动作模型中的触觉预测与实时动作校正
Ma, Zipei, Wei, Xiaofei, Jiang, Junzhe, Lu, Shunlin, Zhang, Li
Abstract
World-action models guide action generation with predicted future observations, but vision-centric predictions miss the local contact cues that decide contact-rich manipulation. However, naively predicting future tactile observations as additional views recovers only a third of the achievable gain in our experiments. This gap reflects a timing mismatch: predictions precede execution, while tactile feedback arrives during it. We introduce TacPAC, which turns tactile prediction into real-time action correction. Once the base model has planned an action chunk, TacPAC caches the predicted contact that plan was conditioned on together with the plan's own representation, and a tactile expert reads each newly observed tactile image against that cache to correct the actions not yet executed. Feedback is thus interpreted against what the plan anticipated rather than in isolation, and one correction is a single pass over that cache, $20.7\times$ cheaper than regenerating the chunk. On five real-robot tasks spanning precision insertion, fragile-object handling, object reorientation, and long-horizon manipulation, TacPAC leads every task and raises the average from 22% for its vision-only base model to 64%. Code is available at https://github.com/LogosRoboticsGroup/TacPAC.
Chinese Translation
世界-动作模型(world-action models)通过预测未来观测来引导动作生成,但以视觉为中心的预测忽略了决定接触密集型操作的局部接触线索。然而,在我们的实验中,简单地将未来触觉观测作为额外视图进行预测,仅能恢复可实现收益的三分之一。这一差距反映了一种时序错配:预测发生在执行之前,而触觉反馈在执行期间才到来。我们提出了TacPAC,它将触觉预测转化为实时动作校正。当基础模型规划出一个动作块(action chunk)后,TacPAC会将该规划所依据的预测接触信息与规划自身的表示一同缓存,随后一个触觉专家模块针对该缓存解读每帧新观测到的触觉图像,以校正尚未执行的动作。因此,反馈是相对于规划所预期的内容而被解读的,而非孤立地解读;且一次校正只需对该缓存进行一遍前向计算,比重新生成整个动作块便宜20.7倍。在涵盖精密插接、易碎物体抓取、物体翻转和长时程操作的五项真机机器人任务上,TacPAC在所有任务中均表现最优,将其纯视觉基础模型的平均成功率从22%提升至64%。代码发布于 https://github.com/LogosRoboticsGroup/TacPAC。
cs.RO / 24 / 2609.05282

Temporal Tactile Encoding and Compliance for Intent-Aware Robot-to-Human Bimanual Handover

面向意图感知的机器人到人类双手交接的时间性触觉编码与柔顺控制
Marra, Pasquale, Berti, Stefano, Caddeo, Gabriele Mario, Natale, Lorenzo
Abstract
Reliable robot-to-human handover requires the robot to infer when the person is ready to receive the object, and release it safely, comfortably, and at the right time. This is challenging because visual observations alone may not disambiguate clear taking intent from accidental contact, weak grasping, wrong-direction forces, or transient interactions. In this work we treat human-robot handover as an intrinsically multimodal problem. Our approach couples a VLA model with a compliance controller that reduces interaction forces during object transfer. We finetune the VLA model with human demonstrations using RGB observation, temporally encoded tactile feedback and proprioception. We evaluate the complete system in a human-subject study against two baselines: one without tactile feedback and one using tactile feedback without compliance control. We hypothesize that combining compliance and temporal tactile encoding yields the most reliable and comfortable handovers, as compliance facilitates physical interaction while tactile history captures sustained taking intent. Performance is measured through objective metrics and an ad-hoc questionnaire. The results show that the two components provide complementary benefits and substantially outperform the baselines. Code and data will be released upon acceptance.
Chinese Translation
可靠的机器人到人类物体交接要求机器人推断人何时准备好接收物体,并在合适的时机安全、舒适地释放物体。这一任务具有挑战性,因为仅凭视觉观测难以区分明确的接收意图与意外接触、微弱抓握、错误方向的力或短暂交互。在本工作中,我们将人机交接视为一个本质上的多模态问题。我们的方法将VLA模型与柔顺控制器相结合,在物体传递过程中降低交互力。我们使用RGB观测、时间性编码的触觉反馈和本体感觉信息,通过人类示教对VLA模型进行微调。我们在一项人类被试实验中对完整系统进行了评估,并与两个基线方法进行对比:一个不使用触觉反馈,另一个使用触觉反馈但不包含柔顺控制。我们假设,柔顺控制与时间性触觉编码的结合能够产生最可靠、最舒适的交接,因为柔顺控制促进了物理交互,而触觉历史能够捕捉持续的接收意图。性能通过客观指标和一份专门设计的问卷进行测量。结果表明,这两个组件提供了互补的优势,并显著优于基线方法。代码和数据将在论文被接收后发布。
cs.RO / 25 / 2609.05300

Human-Human & Human-Robot Interaction Transformer (H2INT) for Robot Navigation in Dense and Uncertain Crowds

用于密集且不确定人群下机器人导航的人-人与人-机器人交互Transformer(H2INT)
Shen, Ao, Chen, Kaixi, Liu, Shiwei, Deng, Fang, Chen, Chen
Abstract
Safe robot navigation in dense crowds requires reasoning about pedestrian motion and how it may change in response to a robot. However, many learning-based approaches generate pedestrian motion independently of the robot or assume uniform reciprocity, omitting an important source of interaction uncertainty. This paper presents a Human-Human & Human-Robot Interaction Transformer (H2INT), a reinforcement learning framework that retains robot-conditioned changes in pedestrian motion during policy learning while allowing responsiveness to vary across pedestrians. Responsiveness affects the crowd dynamics when the robot is visible but is not supplied as a policy input; the policy must instead infer its consequences from robot-centered relative positions. A two-stage gated Transformer progressively encodes human-human and human-robot relations, while a recurrent policy captures their temporal evolution. A curriculum gradually reduces pedestrian responsiveness to increase interaction difficulty. Simulation experiments demonstrate improved navigation safety and robustness over representative baselines across response conditions and crowd densities, and show transfer without retraining to structurally distinct crowd-flow layouts. Ablations support the hierarchical relational encoding and gated updates. Real-robot deployment further verifies that the learned policy can operate with sparse observations in a physical environment.
Chinese Translation
在密集人群中实现安全的机器人导航需要对行人运动进行推理,并预判行人运动如何随机器人行为而变化。然而,许多基于学习的方法在生成行人运动时独立于机器人,或假设统一的互让性(reciprocity),从而忽略了交互不确定性的一个重要来源。本文提出了一种人-人与-人-机器人交互Transformer(Human-Human & Human-Robot Interaction Transformer,H2INT),这是一个强化学习框架,能够在策略学习过程中保留行人运动随机器人条件变化的信息,同时允许不同行人的响应性(responsiveness)存在差异。响应性在机器人可见时会影响人群动态,但并不作为策略输入提供给策略;策略必须转而从以机器人为中心的相对位置中推断其影响。一个两阶段门控Transformer逐步编码人与人关系和人-机器人关系,同时一个循环策略捕捉这些关系的时间演化。课程学习(curriculum)通过逐步降低行人的响应性来增加交互难度。仿真实验表明,在不同响应条件和人群密度下,本方法相较于代表性基线在导航安全性和鲁棒性方面均有提升,并且无需重新训练即可迁移到结构不同的人流布局。消融实验验证了层次化关系编码和门控更新的有效性。真实机器人部署进一步验证了所学策略能够在物理环境中基于稀疏观测进行运行。
cs.RO / 26 / 2609.05324

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

RoboSPA:VLA模型能否超越简单场景与短时程任务?
Fan, Zhenxuan, Zhang, Bo, Lin, Yutong, Yuan, Yuqian, Lin, Juekai, Liang, Liang, Huang, Zhuoyi, Zhang, Wenqiao, Li, Juncheng, Tang, Siliang, Xiao, Jun, Zhuang, Yueting
Abstract
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce \textbf{RoboSPA} (\textbf{Robo}t \textbf{S}patial-\textbf{P}rocedural \textbf{A}ssessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. \texttt{RoboSPA} focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, \texttt{RoboSPA} introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish \texttt{RoboSPA} as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.
Chinese Translation
视觉-语言-动作(VLA)模型在语言条件机器人操作领域取得了可喜的进展。然而,现有的数据集和基准主要评估在预定义设置下的任务完成情况,对模型在空间和程序复杂度不断提升下的推理能力缺乏深入洞察。我们提出了RoboSPA(机器人空间-程序评估,Robot Spatial-Procedural Assessment),这是一个大规模机器人操作数据集与基准,用于诊断VLA模型的具身推理能力。RoboSPA聚焦两个核心维度:细粒度空间推理和长时程程序规划,涵盖10个任务类别和56个基础任务。每个任务在五个难度级别上进行实例化,产生280个具有递增空间模糊性和程序复杂性的变体。我们在多种机器人形态(embodiment)和多样化场景中收集了52.7万条轨迹。除二元成功率之外,RoboSPA还引入了诊断性指标以实现更细致的评估。在代表性VLA模型上的实验表明,当前系统在复杂空间关系、精确低层执行以及记忆密集型规划方面仍然存在困难。这些结果确立了RoboSPA作为一个具有挑战性的诊断基准,可用于开发更强大、可靠且可泛化的具身智能体。我们的数据和代码可在https://github.com/fanzhenxuan/RoboSPA获取。
cs.RO / 27 / 2609.05325

FIRE-LIVWO: Robust LiDAR-Inertial-Visual-Wheel Odometry via Failure-Immune mmWave Radar Enhancement

FIRE-LIVWO:基于故障免疫毫米波雷达增强的鲁棒激光雷达-惯性-视觉-轮式里程计
Hu, Kun, Li, Menggang, Wu, Kaidi, Jin, Zhiwen, Zhao, Yingjie, Tang, Chaoquan, Hu, Eryi, Zhou, Gongbo
Abstract
Achieving robust SLAM in large-scale underground coal mines with complex structures and severe degeneracies remains highly challenging. Dense smoke and dust cause substantial loss of visual information and degrade LiDAR point-cloud features, while long, self-similar corridors induce geometric degeneration, leading to pronounced odometry drift. To address these issues, we propose FIRE-LIVWO: Failure-Immune mmWave Radar-Enhanced LiDAR-Inertial-Visual-Wheel Odometry, a tightly coupled multi-modal odometry framework based on an iterated error-state Kalman filter (IESKF). The framework fuses 4D mmWave radar, LiDAR, and visual features within a unified VoxelMap and jointly constructs LiDAR-radar point-to-plane residuals and sparse visual photometric residuals. In smoke-filled environments, we exploit the strong penetration of 4D mmWave radar and introduce pointwise Doppler velocity constraints to preserve state observability. In geometrically degenerate corridors, we tightly couple wheel odometry using non-holonomic constraints (NHC) and online lever-arm compensation to reduce drift. Our central contribution is a degeneration detection and adaptive fusion model switching strategy grounded in geometric and visual observability analysis, which quantifies observability online and dynamically adjusts modality weights. Real-world experiments in underground coal mines demonstrate that FIRE-LIVWO accurately identifies failure boundaries, enabling reliable modality switching under extreme conditions. Compared with baselines, it achieves superior accuracy and robustness (average localization error of 5.677m). We open source our code on Github to benefit the robotics community.
Chinese Translation
在结构复杂、退化严重的大规模地下煤矿中实现鲁棒的SLAM仍然极具挑战性。浓密的烟雾粉尘会导致视觉信息的大量丢失并降低激光雷达点云特征质量,而狭长且自相似的巷道会引发几何退化,导致显著的里程计漂移。为解决这些问题,我们提出了FIRE-LIVWO:故障免疫毫米波雷达增强的激光雷达-惯性-视觉-轮式里程计(Failure-Immune mmWave Radar-Enhanced LiDAR-Inertial-Visual-Wheel Odometry),这是一个基于迭代误差状态卡尔曼滤波器(IESKF)的紧耦合多模态里程计框架。该框架在统一的VoxelMap中融合4D毫米波雷达、激光雷达与视觉特征,并联合构建激光雷达-雷达点到面残差与稀疏视觉光度残差。在烟雾弥漫的环境中,我们利用4D毫米波雷达的强穿透能力,引入逐点多普勒速度约束以保持状态可观测性。在几何退化的巷道中,我们利用非完整约束(NHC)和在线杆臂补偿紧耦合轮式里程计以减少漂移。我们的核心贡献是一种基于几何与视觉可观测性分析的退化检测与自适应融合模型切换策略,该策略在线量化可观测性并动态调整各模态权重。在地下煤矿的实际实验表明,FIRE-LIVWO能够准确识别失效边界,在极端条件下实现可靠的模态切换。与基线方法相比,它取得了更优的精度与鲁棒性(平均定位误差为5.677米)。我们在Github上开源了代码,以惠及机器人学社区。
cs.RO / 28 / 2609.05331

Adaptation Needs in Robotic Systems: Assessing Behavior Trees and Their Enhancement

机器人系统中的适应性需求:评估行为树及其增强方法
Rostamnia, Mehran, Filippone, Gianluca, Caldas, Ricardo, Pelliccione, Patrizio
Abstract
Robotic systems increasingly operate in dynamic, uncertain, and open-ended environments, where design-time assumptions may no longer hold, and adaptation becomes necessary to maintain effective and safe operation. Behavior Trees (BTs) are widely used in robotic control architectures due to their modularity, readability, and reactivity. This raises a central question: are BTs sufficient to meet the adaptation needs of modern robotic systems? This paper investigates this question through a literature-driven study complemented by empirical validation. First, we derive a classification of robotic adaptation needs from the literature, organizing them into six categories: Knowledge, Perception, Actuation, System, Mission, and Environment. Then, we analyze the capabilities and limitations of classical BTs with respect to these needs. Then, we characterize BT-based approaches for adaptation from the existing literature and organize them into four primary families, i.e., generation, extension, evolution, and refinement, including approaches that combine multiple families. Our analysis shows that the modularity, flexibility, and reactivity of classical BTs are insufficient for adaptation needs involving runtime restructuring, reasoning under uncertainty, mission reinterpretation, learning, or integration with external knowledge and planning mechanisms. Enhanced BT approaches address several of these limitations, but to different extents and often with limitations of their own. Our findings relate adaptation needs to both the capabilities and limitations of classical and enhanced BTs, providing guidance on when classical BTs are sufficient, when enhanced mechanisms are needed, and which challenges remain or emerge for adaptive robotic control architectures.
Chinese Translation
机器人系统日益运行于动态、不确定和开放的环境中,在设计阶段所做的假设可能不再成立,因此需要适应能力以维持有效且安全的运行。行为树(Behavior Trees, BTs)因其模块化、可读性和反应性而被广泛应用于机器人控制架构中。这引出了一个核心问题:行为树是否足以满足现代机器人系统的适应需求?本文通过一项基于文献的研究并结合实证验证来探讨这一问题。首先,我们从文献中归纳出机器人适应需求的分类,将其组织为六大类别:知识(Knowledge)、感知(Perception)、驱动(Actuation)、系统(System)、任务(Mission)和环境(Environment)。然后,我们分析了经典行为树针对这些需求的能力与局限。接着,我们对现有文献中基于行为树的适应方法进行了特征刻画,并将其归纳为四个主要类别,即生成(generation)、扩展(extension)、演化(evolution)和细化(refinement),其中包括结合多个类别的方法。我们的分析表明,经典行为树的模块化、灵活性和反应性不足以应对涉及运行时重构、不确定性下的推理、任务重新解释、学习以及与外部知识和规划机制集成等适应需求。增强型行为树方法在一定程度上解决了其中一些局限,但解决程度不一,且往往自身也存在局限。我们的研究结果将适应需求与经典及增强型行为树的能力和局限相关联,为判断何时经典行为树已足够、何时需要增强机制,以及自适应机器人控制架构还面临哪些挑战或新问题提供了指导。
cs.RO / 29 / 2609.05361

Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction

面向多模态人机交互的人形机器人原型研制
Viet, Thang Tran, Canh, Thanh Nguyen, Gia, Huy Uong, Van, Phuc Dinh, Duc, Son Tran, Do, Ngoc Minh, HoangVan, Xiem
Abstract
Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in real-world environments. This paper introduces a humanoid robot prototype designed as a flexible testbed for developing and integrating artificial intelligence (AI) modules in HRI tasks. The system features a 12 degree-of-freedom (DOFs) dual-arm mechanism and a 2 DOFs head with an expressive LCD screen to express facial emotions. All hardware components are controlled by a custom-designed controller board with real-time AI processing supported by an onboard Jetson module. The system incorporates three AI modules: (1) gesture recognition using MediaPipe Pose and an LSTM classifier, (2) object detection with YOLO and 3D localization, and (3) voice-command processing through speech recognition and large language model(LLM)-based semantic parsing. The platform is validated through experiments on positioning accuracy, with results showing average manipulation errors of approximately 1.83 cm. To demonstrate its versatility, experimental results show over 90% task accuracy, with gesture recognition reaching 96%, speech recognition reaching 92%. The results confirm the effectiveness of the proposed system as a reproducible and accessible humanoid platform for research and prototyping in HRI.
Chinese Translation
人机交互(HRI)使人类与机器人能够在真实环境中实现直观、智能的协作。本文介绍了一种人形机器人原型,该原型被设计为一个灵活的测试平台,用于开发和集成面向人机交互任务的人工智能(AI)模块。该系统采用12自由度(DOFs)双臂机构和2自由度头部,头部配备表情丰富的LCD屏幕以表达面部情绪。所有硬件组件由一块自主设计的控制器板控制,并由板载Jetson模块支持实时AI处理。该系统集成了三个AI模块:(1)基于MediaPipe Pose和LSTM分类器的手势识别;(2)基于YOLO的目标检测与三维定位;(3)通过语音识别和基于大语言模型(LLM)的语义解析实现的语音指令处理。该平台通过定位精度实验进行了验证,结果显示平均操作误差约为1.83厘米。为展示其多功能性,实验结果表明任务准确率超过90%,其中手势识别达到96%,语音识别达到92%。实验结果证实了所提出系统作为一种可复现、易于获取的人形机器人平台,在人机交互研究与原型开发中的有效性。
cs.RO / 30 / 2609.05369

Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation

面向长时序视觉-语言-动作操作的神经符号化程序性推理
Chavan, Vivek, Shi, Yahuan, Heimann, Oliver, Haninger, Kevin, Krüger, Jörg
Abstract
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. We investigate a neuro-symbolic framework that combines learned VLA control with explicit task graphs and multimodal procedural memory. Task graphs encode action dependencies, valid transitions, and branch conditions, while memory maintains the active step, completed actions, textual context, and task-relevant visual evidence. Together, these structures guide object selection, destination grounding, subgoal dispatch, and verification of expected state transitions. Human demonstrations provide additional spatial and temporal guidance through gaze or saliency cues. To isolate their effect on policy learning, our initial study bypasses cross-view gaze transfer and directly annotates pseudo-gaze in robot-view teleoperation videos. The resulting guidance is used during VLA fine-tuning and inference. We study two long-horizon manipulation domains, workspace clearing and surgical-instrument handling, which require ordered execution, visually grounded decisions, and conditional branching. We evaluate correct-object and destination selection, subtask completion, task progress, step-order consistency, complete-task success, and procedural or execution mistakes. This work positions structured symbolic reasoning and demonstration-derived visual guidance as complementary mechanisms for reliable long-horizon VLA manipulation.
Chinese Translation
视觉-语言-动作(VLA)模型能够执行短时操作技能,但在需要持续任务状态、依赖感知推理、条件决策和可靠接地(grounding)的长时序程序中仍然脆弱。我们研究了一种神经符号化框架,该框架将学习得到的VLA控制与显式任务图及多模态程序性记忆相结合。任务图编码动作依赖关系、有效转移和分支条件,而记忆则维护当前步骤、已完成的动作、文本上下文以及与任务相关的视觉证据。这些结构共同引导物体选择、目标位置接地、子目标分发以及预期状态转移的验证。人类演示通过注视(gaze)或显著性线索提供额外的空间和时间引导。为了隔离此类引导对策略学习的影响,我们的初步研究绕过跨视角注视迁移,直接在机器人视角遥操作视频中标注伪注视信息。由此获得的引导被用于VLA微调和推理阶段。我们研究了两个长时序操作领域——工作空间清理和手术器械处理,它们要求有序执行、视觉接地的决策以及条件分支。我们评估了正确物体与目标位置选择、子任务完成、任务进度、步骤顺序一致性、完整任务成功率以及程序性或执行性错误。这项工作将结构化符号推理和源自演示的视觉引导定位为实现可靠长时序VLA操作的互补机制。
cs.RO / 31 / 2609.05376

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

何事何时重要:诊断与改进视觉运动模仿策略中的条件性视觉落地
Chavan, Vivek, Xie, Pengtao, Shi, Yahuan, Heimann, Oliver, Haninger, Kevin, Krüger, Jörg
Abstract
Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state. Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage. Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting as complementary interventions for improving target selection while preserving spatial information required for control. These interventions substantially improve robustness in simulation and on a physical UR3e. We further examine the same failure pattern in a pretrained vision-language-action policy on a state-conditioned instrument-handling task, where the observed state of a medical instrument determines the correct destination. Together, the results show that visual distractors can cause incorrect object or destination selection even when the underlying manipulation skill remains intact, and that explicitly improving target selection can substantially recover performance across distinct visuomotor policy-learning regimes.
Chinese Translation
视觉运动模仿策略(visuomotor imitation policies)能够在分布内的视觉条件下取得高性能,但在引入视觉上相似的物体或容器时则会失败。我们将这种行为视为条件性视觉落地(conditional visual grounding)问题:成功控制所需的视觉目标会随操作阶段而变化,在更复杂的任务中还会随观测到的任务状态而变化。基于 Action Chunking with Transformers(ACT),我们系统地引入了具有可控颜色和形状相似度的干扰物体与容器,并将失败定位到抓取和放置环节。我们发现,干扰敏感性同时取决于视觉相似性的类型和操作阶段。基于这一诊断,我们评估了干扰增广、阶段依赖的注意力正则化以及基于外观的视觉提示(visual prompting)作为互补的干预手段,以在保留控制所需空间信息的同时改进目标选择。这些干预手段在仿真和一个物理 UR3e 机械臂上显著提升了鲁棒性。我们进一步在一个预训练的视觉-语言-动作(vision-language-action)策略中,在一个状态条件化的器械操作任务中检验了同样的失败模式——该任务中医疗器械的观测状态决定了正确的目标位置。总体而言,结果表明视觉干扰物即使在不影响底层操作技能的情况下,也可能导致错误的物体或目标位置选择,而显式地改进目标选择能够在不同的视觉运动策略学习机制中大幅恢复性能。
cs.RO / 32 / 2609.05401

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

相同轨迹,矛盾奖励(ROBORMBENCH):视觉语言奖励模型中的改述脆弱性
Jeung, Wonje, Yoon, Sangyeon, Hong, Hyesoo, Cho, Yoonjun, Jeon, Dongjae, Kim, Bumjun, Oh, Jean, Yu, Youngjae, No, Albert
Abstract
Vision-language models are increasingly used as reward functions for robotic learning, but this role requires paraphrase invariance: the same trajectory should receive the same reward under semantically equivalent goal descriptions. We show that current VLM reward models often violate this property. Paraphrasing the instruction alone can substantially change predicted progress scores, and can even flip identical robot behavior between failure and success. To measure this failure mode, we introduce ROBORMBENCH, a benchmark with 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases spanning lexical, syntactic, and action-goal rewrites. Across proprietary and open-source VLMs, paraphrase-induced instability is widespread and severe, grows under more divergent rewrites, and is not reliably reduced by scale or explicit reasoning. Dedicated reward models trained with trajectory-grounded supervision are substantially more stable. These results show that paraphrase robustness is a core requirement for reliable VLM-based reward modeling in robotics.
Chinese Translation
视觉语言模型(VLM)正越来越多地被用作机器人学习的奖励函数,但这一角色要求改述不变性:同一条轨迹在语义等价的目标描述下应获得相同的奖励。我们发现,当前的VLM奖励模型往往违反这一性质。仅对指令进行改述就可能显著改变预测的进度分数,甚至能将完全相同的机器人行为在失败与成功之间翻转。为了度量这一失败模式,我们提出了ROBORMBENCH,这是一个包含2,390条真实机器人轨迹、真实进度标签以及21,673条经验证的改述(涵盖词汇、句法和动作-目标重写)的基准。在专有和开源VLM上,改述引起的不稳定性普遍而严重,且随着重写差异的增大而加剧,并且无法通过扩大模型规模或显式推理可靠地降低。基于轨迹监督训练的专用奖励模型则显著更加稳定。这些结果表明,改述鲁棒性是构建可靠VLM奖励模型的核心要求。