← Back to Index
Daily Research Digest

arXiv Papers

2026-09-02
272
Papers
3
Categories
272
Translated
收藏清单 0
人工智能 (Artificial Intelligence)
124
cs.AI / 1 / 2609.00002

HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models

HyperWorld:超图结构化状态序列化改进学习型文本世界模型
Zhang, Yun-Jian, Liang, Chen-Wei, Zhang, Tian-Yi, Ding, Jian, Wu, Yi-Lun, Li, Ao-Bo, Su, Wei-Cong, Saifullah, An, Hong-Yu, Wang, Mu-Jiang-Shan
Abstract
World models enable language-model agents to predict environment dynamics and plan before acting. In text environments, the model must learn symbolic action effects from serialized state descriptions, but the role of serialization structure remains underexplored. We present HyperWorld, a controlled study of state serialization for learned textual world models. We compare raw observations with three symbolic serializations of the same ground-truth state: independent sentences, pairwise triples, and entity-centered hyperedge units that group multiple related facts around entities and relations. All variants use the same training objective: given a state and an action, predict symbolic effects or judge the action infeasible. Across model scales, data budgets, and in-distribution and out-of-distribution test worlds, hyperedge serialization gives the clearest gains for 0.5B--1.5B models and under distribution shift. Larger models reduce the gap, and pairwise triples can match or slightly exceed hyperedges on in-distribution exact match, but hyperedges achieve the strongest out-of-distribution fact F1 and the best small-to-medium scale trade-off between feasibility detection and effect prediction. In downstream greedy planning, the hyperedge world model also attains the highest success rate among the tested representations. These results show that higher-order state organization is a simple but effective inductive bias for learned symbolic world models, especially when model capacity is limited or test environments differ from training.
Chinese Translation
世界模型使语言模型智能体能够预测环境动态并在行动前进行规划。在文本环境中,模型必须从序列化的状态描述中学习符号化动作效果,但序列化结构的作用仍未得到充分探索。我们提出了HyperWorld,一项针对学习型文本世界模型中状态序列化的受控研究。我们比较了原始观测与对同一真值状态的三种符号化序列化方式:独立句子、成对三元组,以及以实体为中心的超边单元——后者将多个相关事实围绕实体和关系进行分组。所有变体使用相同的训练目标:给定状态和动作,预测符号化效果或判定动作不可行。在多种模型规模、数据预算以及分布内和分布外测试世界中,超边序列化为0.5B–1.5B规模的模型以及分布偏移场景带来了最显著的收益。更大的模型会缩小差距,且成对三元组在分布内精确匹配上可以持平或略微超越超边,但超边在分布外事实F1上取得最优结果,并在中小规模下实现了可行性与效果预测之间的最佳权衡。在下游贪心规划中,超边世界模型在所测试的表示方法中也取得了最高的成功率。这些结果表明,高阶状态组织是学习型符号世界模型的一种简单而有效的归纳偏置,尤其是在模型容量有限或测试环境与训练环境存在差异时。
cs.AI / 2 / 2609.00003

I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models

I-CARE:面向文本到图像模型的可控、多样且具代表性的遗忘设定中干扰相关现象的分析
Pereira, Leonardo Santiago Benitez, Viñolo, Marcos Escudero, Arribas, Luis Herranz
Abstract
Machine unlearning studies the removal of knowledge from an AI model, making the system forget a concept it previously learned. Despite rapid progress in generative machine unlearning, the unintended degradation of semantically related concepts that should have been retained (henceforth, interference) remains poorly characterized and inconsistently evaluated. This paper introduces I-CARE, a methodology that formalizes interference as a first-class object of study in generative unlearning. Rather than proposing a new benchmark or unlearning algorithm, I-CARE provides formal definitions for tasks, metrics, and templates for reporting results, enabling the systematic and reproducible study of interference across unlearning settings. While our methodology is designed to remain valid as models and unlearning algorithms evolve, decoupling long-term scientific insight from transient empirical results, we present a feasibility demonstration with state-of-the-art algorithms and frequently used datasets. The results demonstrate that I-CARE enables meaningful analysis of interference patterns across multiple unlearning settings, establishing the practical applicability of the framework. The software implementation of the methodology is provided in an open-source framework, together with a web-based graphical interface that enables exploration of the outcomes of this study without requiring direct interaction with the codebase or specialized data analysis tools.
Chinese Translation
机器遗忘(machine unlearning)研究如何从AI模型中移除知识,使系统忘记其先前学到的概念。尽管生成式机器遗忘进展迅速,但本应保留的语义相关概念出现的非预期性能退化(下文简称“干扰”)仍然缺乏充分的表征和不一致的评估。本文提出I-CARE,这是一种将干扰形式化为生成式遗忘研究一等研究对象的方法论。I-CARE并非提出新的基准或遗忘算法,而是为任务、指标和结果报告模板提供了正式定义,从而支持跨遗忘设定的系统性、可复现的干扰研究。虽然我们的方法论旨在随着模型和遗忘算法的演进而保持有效,将长期科学洞见与短期实证结果解耦,但我们使用最先进的算法和常用数据集进行了可行性演示。结果表明,I-CARE能够对多个遗忘设定中的干扰模式进行有意义的分析,确立了该框架的实际适用性。该方法论的软件实现以开源框架形式提供,并配有基于网页的图形界面,使用户无需直接与代码库交互或使用专门的数据分析工具即可探索本研究的成果。
cs.AI / 3 / 2609.00004

Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing

具有随机需求时刻的多产品-capacity受限批量调度问题的离散时间马尔可夫决策过程建模
Bayati, Léa, Dahmoune, Mohamed, Rodoplu, Melek
Abstract
This paper studies a finite-horizon multi-item capacitated lot-sizing problem in which demand quantities are deterministic, while demand-arrival periods are stochastic. Each demand occurs once within a known time window and must be satisfied no later than its deadline. The proposed model makes production and allocation decisions at the demand level, allowing it to represent capacity competition, demand-specific backlog, and allocation-dependent inventory dynamics. The stochastic problem is formulated as a discrete-time Markov decision process (DTMDP), including the state space, feasible actions, transition kernel, and one-period cost function. To isolate the computational effect of stochastic timing, each stochastic instance is first compared with a deterministic counterpart in which each arrival distribution is replaced by its most likely arrival period. This comparison shows that stochastic timing substantially increases the number of states, the number of transitions, solution time, and memory pressure. A genetic algorithm (GA) is then proposed for the stochastic-timing problem. The GA searches over feasible state-feedback policies and evaluates each policy exactly under the DTMDP transition model. Computational experiments on 330 benchmark instances show that the GA remains close to the exact stochastic solution whenever the latter is available, with an average optimality gap of about $3.44\%$. On the difficult benchmark instances, comprising 90 test cases, the GA remains below the $5\%$ optimality-gap threshold and achieves an average optimization speedup of $6.89 \pm 1.41$ at the $95\%$ confidence level. For instances that cannot be solved exactly on the available hardware, an empirical Bellman-time regression is used to estimate the missing exact resolution time and extrapolate the expected GA speedup.
Chinese Translation
本文研究了一类有限时段的多产品产能受限批量调度问题,其中需求量为确定性的,而需求到达时刻是随机的。每个需求在已知时间窗口内发生一次,且必须不晚于其截止期限被满足。所提出的模型在需求层面进行生产与分配决策,从而能够刻画产能竞争、需求特定的缺货推迟以及依赖于分配的库存动态。该随机问题被表述为离散时间马尔可夫决策过程(DTMDP),包括状态空间、可行行动、转移核和单周期成本函数。为分离随机时刻带来的计算影响,每个随机实例首先与一个确定性对应问题进行比较,其中每个到达分布被替换为其最可能的到达时刻。该比较表明,随机时刻显著增加了状态数量、转移数量、求解时间和内存压力。随后,针对随机时刻问题提出了一种遗传算法(GA)。该遗传算法在可行的状态反馈策略空间中搜索,并在DTMDP转移模型下对每个策略进行精确评估。在330个基准实例上的计算实验表明,只要精确随机解可得,遗传算法的解就接近精确解,平均最优性间隙约为3.44%。在包含90个测试用例的困难基准实例上,遗传算法的最优性间隙保持在5%阈值以下,并在95%置信水平下实现了6.89±1.41的平均优化加速比。对于在现有硬件上无法精确求解的实例,采用经验Bellman时间回归方法估计缺失的精确求解时间,并外推遗传算法的预期加速比。
cs.AI / 4 / 2609.00005

Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models

基于指令微调小型语言模型的渐进式老年人金融诈骗增量风险评估
Ghafariasl, Parviz, Fu, Weimin, Guo, Xiaolong, Chang, Shing I.
Abstract
Financial scams targeting older adults increasingly occur through text and voice channels such as email, SMS, and phone calls, unfolding over multiple conversational turns that begin with impersonation or casual contact, escalate through trust building and urgency, and culminate in requests for sensitive information or financial transfers. Because risk signals emerge incrementally across turns, effective detection requires models that continuously update risk estimates under resource-constrained deployment settings. We propose a cumulative turn-based risk assessment framework that incrementally aggregates conversational turns and re-estimates risk at each step, enabling dynamic scam monitoring across progressively evolving conversations. A multi-turn dialogue dataset is constructed to cover investment, charity, and tech support scam scenarios, with each dialogue containing two to eight turns and annotated at every cumulative stage with a qualitative risk level, a continuous risk score, an explanatory rationale, and a safety recommendation. Four small language models (Phi-4, LLaMA-3.2, DeepSeek-R1, and Qwen3) are fine-tuned and evaluated under a unified training framework. Fine-tuned small models capture fraud-related linguistic cues and cross-turn escalation patterns while maintaining compact architectures suitable for mobile and resource-constrained deployment settings. Among the evaluated models, Phi-4 and LLaMA-3.2 achieve stronger turn-aware risk estimation performance relative to their parameter scale. These results suggest that structured cumulative modeling can support incremental scam risk assessment in deployment-oriented settings while highlighting the potential of compact language models for privacy-aware and on-device fraud protection.
Chinese Translation
针对老年人的金融诈骗日益通过电子邮件、短信和电话等文本和语音渠道发生,并在多轮对话中逐步展开:从冒充身份或随意接触开始,通过建立信任和制造紧迫感逐步升级,最终以索取敏感信息或要求金融转账告终。由于风险信号在对话轮次中逐步显现,有效的检测需要模型能够在资源受限的部署环境下持续更新风险估计。我们提出了一种基于累积轮次的风险评估框架,该框架增量地聚合对话轮次并在每一步重新估计风险,从而实现对逐步演化的对话的动态诈骗监控。我们构建了一个多轮对话数据集,涵盖投资、慈善和技术支持诈骗场景,每段对话包含两到八个轮次,并在每个累积阶段标注了定性风险等级、连续风险评分、解释性依据以及安全建议。在统一的训练框架下,对四个小型语言模型(Phi-4、LLaMA-3.2、DeepSeek-R1 和 Qwen3)进行了微调和评估。经过微调的小型模型能够捕捉与诈骗相关的语言线索和跨轮次升级模式,同时保持适用于移动设备和资源受限部署环境的紧凑架构。在所评估的模型中,Phi-4 和 LLaMA-3.2 相对于其参数规模取得了更强的轮次感知风险估计性能。这些结果表明,结构化的累积建模能够在面向部署的环境中支持增量式诈骗风险评估,同时也凸显了紧凑型语言模型在隐私保护和端侧反诈防护方面的潜力。
cs.AI / 5 / 2609.00012

Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls

大语言模型中的长程状态跟踪:通过深度依赖工具调用序列执行 MD5
Pai, Dheeraj Mohandas, Xian, Lu
Abstract
Long-horizon tasks remain uncommon in large language model (LLM) evaluation, and for a reason: when each step depends on the last, per-step accuracy that looks excellent in isolation decays catastrophically, as errors cascade and the end-to-end failure probability grows sharply with length. Existing agentic benchmarks report end-to-end success but confound this state-tracking difficulty with instruction interpretation, give no control group that isolates it, and are vulnerable to shortcuts such as a hallucinated final answer, so they cannot say why a long run fails. Whether an LLM can carry exact intermediate state across many tool calls at all is itself not well established. We test this cleanly by having the model compute a cryptographic hash, MD5, step by step: a sequence of $196$ dependent tool calls over $64$ rounds while it carries four $32$-bit words $(a,b,c,d)$ in its own context from one call to the next. Interpretation is trivial and, because we implement MD5 from scratch (RFC~1321), we align every call to the ground-truth trace and check the digest to the bit, so any failure is pure bookkeeping. gpt-oss-120b, a mixture-of-experts model with only $\sim$5.5B active parameters per token, at temperature $0$ with a short fixed prompt, carries the full state across all $196$ calls and returns the correct digest on a majority of completed runs. In the strongest setting we replace every primitive tool with a second LLM, so a driver and a worker compute the whole hash from scratch with no exact-arithmetic oracle in the loop. Two ingredients decide success and neither changes the weights: keeping the model's own reasoning in its context each turn, and voting over a thinking-enabled worker to remove its modular-arithmetic slips. We localize the residual failures by origin, separating state-carrying from arithmetic and from serving.
Chinese Translation
长程任务在大语言模型(LLM)评估中仍然少见,这是有原因的:当每一步都依赖于上一步时,单独看起来优异的单步准确率会灾难性衰减,因为误差不断级联,端到端失败概率随长度急剧增长。现有的智能体基准报告端到端成功率,但将这种状态跟踪难度与指令理解混淆在一起,没有隔离该难度的对照组,并且容易受到诸如幻觉出最终答案等捷径的影响,因此无法解释长程运行为何失败。LLM 究竟能否在多次工具调用之间携带精确的中间状态,本身也尚未得到充分验证。我们通过让模型逐步计算一个密码学哈希值 MD5 来干净地检验这一点:在 64 轮中进行 196 个相互依赖的工具调用,模型需要在自己的上下文中跨调用携带四个 32 位字。指令理解极为简单,而且由于我们从零开始实现 MD5(RFC 1321),我们将每次调用与真实轨迹对齐并逐位校验摘要,因此任何失败都纯粹是记账(状态维护)问题。gpt-oss-120b 是一个每个 token 仅激活约 55 亿参数的混合专家模型,在温度为 0 且使用简短固定提示的条件下,能够在全部 196 次调用中携带完整状态,并在多数完成的运行中返回正确的摘要。在最强的设置中,我们用第二个 LLM 替换每一个基础工具,即由驱动模型和工作模型在不依赖任何精确算术 oracle 的情况下从零完成整个哈希计算。决定成功的两个关键因素均不改变模型权重:每轮将模型自身的推理保留在其上下文中,以及通过对启用思考模式的工作模型进行投票来消除其模运算失误。我们按来源对残余失败进行定位,区分状态携带、算术和服务层面的错误。
cs.AI / 6 / 2609.00015

OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets

OpenAgentFlow:为异构AI智能体集群启用系统级安全边界
Chen, Dongsheng, Zhao, Xiangyu, Yao, Xin, Wei, Xuetao
Abstract
AI agents powered by large language models are evolving from isolated assistants into heterogeneous systems in which multiple agents, planners, tools, and execution backends operate over shared environments. In such settings, safety becomes a system-level action-governance problem: deciding whether a pending action should be committed given policy-relevant state accumulated across a session. Existing safeguards operate at fragmented boundaries, making it difficult to enforce shared policies over composed action flows across heterogeneous execution paths. We present OpenAgentFlow, a control-plane/action-plane architecture that establishes the action-commit boundary as a shared enforcement interface. GUI, API, tool, and LLM-generated actions are normalized into a common AgentEvent stream and mediated by a shared pre-execution Policy Enforcement Point, while provenance, session state, audit evidence, and updatable policies are maintained outside individual agents. This provides a common governance layer across incompatible executors and allows new policies to take effect without modifying agents, prompts, models, or execution paths. We evaluate OpenAgentFlow through complementary system evaluations spanning controlled action-flow tests, a public external benchmark, policy updates, and real Android execution. On a 300-case controlled suite, OpenAgentFlow achieves 94.00% accuracy and a 95.35% attack-block rate. On the complete 1,220-case AgentDojo-Traj split of TS-Bench, it achieves 97.62% accuracy, 96.59% unsafe-action recall, and a 1.96% safe false-intervention rate. New control-plane rules take effect without modifying protected agents, and the same enforcement path operates across live GUI, API/tool, and LLM-planned Android execution. These results show that a shared action-commit boundary provides a practical basis for system-wide governance across heterogeneous agent execution paths.
Chinese Translation
由大语言模型驱动的AI智能体正从孤立的助手演变为异构系统,其中多个智能体、规划器、工具和执行后端在共享环境中运行。在这种场景下,安全成为一个系统级的动作治理问题:即根据会话中积累的与策略相关的状态,决定某个待执行动作是否应被提交。现有的安全防护在碎片化的边界上运行,难以在跨越异构执行路径的组合动作流上实施共享策略。我们提出OpenAgentFlow,一种控制平面/动作平面架构,将动作提交边界确立为共享的执行接口。GUI、API、工具及LLM生成的动作被规范化为统一的AgentEvent流,并由共享的执行前策略执行点(Policy Enforcement Point)进行中介,而溯源信息、会话状态、审计证据和可更新策略则维护在各个智能体之外。这为不兼容的执行器提供了统一的治理层,并允许新策略在不修改智能体、提示词、模型或执行路径的情况下生效。我们通过互补的系统评估对OpenAgentFlow进行评估,涵盖受控动作流测试、公开外部基准、策略更新以及真实的Android执行。在包含300个案例的受控测试集上,OpenAgentFlow达到94.00%的准确率和95.35%的攻击拦截率。在TS-Bench完整的1,220个案例的AgentDojo-Traj划分上,它达到97.62%的准确率、96.59%的不安全动作召回率以及1.96%的安全误干预率。新的控制平面规则无需修改受保护的智能体即可生效,且同一执行路径可跨实时GUI、API/工具以及LLM规划的Android执行运行。这些结果表明,共享的动作提交边界为跨异构智能体执行路径的系统级治理提供了切实可行的基础。
cs.AI / 7 / 2609.00018

SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces

SCAFFOLD:一个包含图表问答和思维链推理轨迹的大规模计算机科学研究插图结构化数据集
Raut, Ranjit, Subedi, Aarav, Rai, Sagun, Jha, Sudan
Abstract
Computer science papers rely heavily on diagrams: architecture drawings, system flowcharts, and pipeline schematics that often carry more information than the text around them. There is currently no public dataset that pairs this specific kind of figure with captions, context, questions, answers, and step-by-step reasoning, which is exactly what is needed to train a vision-language model to understand them. We present \textbf{SCAFFOLD}\footnote{https://github.com/theranjitraut/scaffold}, a large-scale structured dataset of computer science research figures with diagram QA and Chain-of-Thought reasoning traces. This dataset consists of (image, caption, context, question-answer, chain-of-thought) tuples from arXiv computer science papers prepared using layout detection and PDF parsing, with an AI-assisted question-generation step. The resulting large-sized SCAFFOLD-157K dataset spans 3,058 papers with 29,887 figures (157,387 pairs), a medium-sized SCAFFOLD-37K dataset (36,797 pairs), and a small-sized SCAFFOLD-12K dataset (12,000 pairs). We used SCAFFOLD-12K for baseline experiments on Qwen2.5-VL-3B-Instruct.
Chinese Translation
计算机科学论文高度依赖图示:架构图、系统流程图和流水线示意图,它们所承载的信息往往多于周围的文字。目前尚无公开数据集将这类特定插图与标题、上下文、问题、答案以及逐步推理配对,而这正是训练视觉语言模型理解此类插图所必需的。我们提出了 SCAFFOLD(https://github.com/theranjitraut/scaffold),一个包含图表问答和思维链推理轨迹的大规模计算机科学研究插图结构化数据集。该数据集由(图像、标题、上下文、问答对、思维链)元组组成,来源于 arXiv 计算机科学论文,通过版面检测和 PDF 解析构建,并辅以 AI 辅助的问题生成步骤。最终构建的大规模 SCAFFOLD-157K 数据集涵盖 3,058 篇论文和 29,887 幅插图(157,387 个问答对),此外还包括中等规模的 SCAFFOLD-37K 数据集(36,797 个问答对)和小规模的 SCAFFOLD-12K 数据集(12,000 个问答对)。我们使用 SCAFFOLD-12K 在 Qwen2.5-VL-3B-Instruct 上进行了基线实验。
cs.AI / 8 / 2609.00028

UI-Venus-2 Technical Report

UI-Venus-2 技术报告
Venus Team, Cai, Zhuohan, Chen, Haoxing, Chen, Jiaxuan, Chen, Weizhi, Gao, Changlong, Gu, Zhangxuan, Guo, Yuan, Hu, Yusong, Jiang, Jianrong, Li, Jianguo, Li, Runze, Lin, Jinzhen, Ma, Zhenyu, Meng, Changhua, Peng, Han, Qiu, Xinyu, Shen, Shuheng, Shui, Zhongyi, Wang, Weiqiang, Wen, Ming, Xu, Zhuoer, Yan, Hang, Yang, Kaiwen, Yao, Ruilin, Yu, Nanjun, Zeng, Zhengwen, Zhang, Lianrui, Zhang, Yunzhu, Zhao, Zhe, Zhou, Beitong
Abstract
Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.
Chinese Translation
多模态GUI智能体已成为数字任务自动化的一个有前景的范式,但由于环境覆盖有限、任务构建脆弱以及奖励验证不可靠,从面向基准测试的模型过渡到可靠的真实世界应用仍然充满挑战。在本工作中,我们提出了UI-Venus-2,一个通用的基础GUI智能体,旨在通过统一的闭环推理-动作框架跨移动端、网页和桌面环境运行。为弥合通往实际部署的差距,我们联合扩展了三个关键维度:(1) 环境,将覆盖范围扩展至170多个多语言移动应用和原生桌面操作系统;(2) 任务,采用深度研究流程进行基于功能的指令生成;(3) 验证,采用轨迹级和样本级评估器,结合视觉关键点和多模型投票,确保为训练提供可靠的强化学习信号。此外,我们集成了安全感知机制,以确保关键动作的受控执行。通过提供一个强大、高效且开源的基础模型,UI-Venus-2推动该领域朝着更具泛化性、可验证性和自我反思性的真实世界应用智能体迈进。
cs.AI / 9 / 2609.00032

EULER: Exploring Underused Links with Evidence-Checked Return for Multi-Agent Mathematical Discovery

EULER:面向多智能体数学发现的证据可校验回传的欠利用链接探索方法
Zhenzhuo, Ren
Abstract
Mathematical communities work with different objects, invariants, and tools, so transferring a problem across them is expensive and often skipped. We present EULER, a multi-agent system that takes such a transfer--a bridge--as its unit of search. Around a fixed conjecture, EULER runs direct, adjacent-domain, and distant-domain routes in competition; a bridge keeps its budget only if it supplies an operation the source representation cannot execute and its target-side evidence returns to the original statement along a checked implication. Six ordered stress tests reject invalid bridges before expensive search begins. We evaluate EULER on 120 recent conjectures. The conjectures were frozen before search and screened for contamination, and are drawn from public papers by authors who had recently published in the Journal of Combinatorial Theory, Series A, a leading journal in combinatorics. EULER produced 10 proofs and 3 refutations, plus 45 scoped partial results. Two mechanisms held up under ablation: bridge-specific stress tests cut incorrect conclusions from 9 to 3, and bridge material combined with a target-native operation yielded a positive interaction of +4.2 resolved tasks that neither factor produced alone. Domain distance did not reliably predict success; executable operation gain and valid return did.
Chinese Translation
数学界的不同领域使用不同的对象、不变量和工具,因此跨领域迁移问题代价高昂且常被忽视。我们提出EULER,一个以这种迁移(即"桥梁")为搜索单元的多智能体系统。围绕一个固定的猜想,EULER并行开展直接路径、邻域路径与远域路径的竞争式探索;只有当桥梁提供了源表示无法执行的操作,且其目标侧证据能够沿经过校验的蕴含关系回传到原始命题时,该桥梁才能保留其预算。六项有序的压力测试在昂贵的搜索开始之前即淘汰无效桥梁。我们在120个近期猜想上评估了EULER。这些猜想已于搜索前冻结并经污染筛查,选取自已最近在组合数学权威期刊《Journal of Combinatorial Theory, Series A》上发表论文的作者们的公开文献。EULER产生了10个证明和3个反驳,外加45个限定范围的局部结果。消融实验证实了两个机制的有效性:桥梁特定的压力测试将错误结论从9个降至3个;桥梁素材与目标域原生操作相结合产生了+4.2个已解决任务的正面交互效应,这是任一单独因素都无法实现的。领域距离并不能可靠地预测成功;可执行操作的增益和有效的回传才可以。
cs.AI / 10 / 2609.00071

When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation

当预测误差并不足够时:评估用于因果估计的混杂函数预测
Cao, Cong
Abstract
Prediction error is widely used to evaluate nuisance-function estimators in causal inference, but its relationship with causal estimator performance may differ across performance measures. We studied this question in a partially linear model using Monte Carlo simulations. We compared ordinary least squares (OLS), generalized additive models (GAMs), XGBoost, and Double Machine Learning with XGBoost (DML-XGBoost), evaluating nuisance-function prediction error, bias, RMSE, and 95\% confidence interval coverage. We also examined a simple joint-error measure based on the absolute cross-product of estimation errors from the exposure and outcome nuisance functions. Across the simulated settings, XGBoost had the lowest RMSE among the non-oracle methods, while DML-XGBoost generally provided better confidence interval coverage. Prediction error did not consistently track causal bias across methods and settings, and the method with the best point-estimation performance did not necessarily have the best confidence interval coverage. The joint-error measure was only weakly associated with causal bias and did not provide a useful standalone measure of causal performance. These results suggest that prediction error is useful for assessing nuisance-function estimation, but it should not be treated as a direct measure of the quality of the resulting causal estimator.
Chinese Translation
预测误差被广泛用于评估因果推断中混杂函数估计器的性能,但其与因果估计器性能之间的关系可能因性能指标的不同而有所差异。我们在部分线性模型中通过蒙特卡洛模拟研究了这一问题。我们比较了普通最小二乘法(OLS)、广义可加模型(GAMs)、XGBoost 以及结合 XGBoost 的双重机器学习方法(DML-XGBoost),评估了混杂函数预测误差、偏差、均方根误差(RMSE)以及 95% 置信区间覆盖率。我们还考察了一种基于暴露混杂函数与结局混杂函数估计误差的绝对交叉乘积的简单联合误差度量。在各模拟情形中,XGBoost 在非理想方法中具有最低的 RMSE,而 DML-XGBoost 通常提供更好的置信区间覆盖率。预测误差在不同方法和情形下并不能始终与因果偏差保持一致,且点估计性能最佳的方法未必具有最佳的置信区间覆盖率。联合误差度量与因果偏差仅呈弱关联,无法作为评估因果性能的独立有效指标。这些结果表明,预测误差对于评估混杂函数估计是有用的,但不应被视为衡量最终因果估计器质量的直接指标。
cs.AI / 11 / 2609.00073

MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts

MiNER:用于临床文本中疟疾疾病实体识别的微调生物医学自然语言处理
Anoop, V. S., N, Devika
Abstract
Malaria remains a significant global health burden, necessitating continuous research efforts to understand its complex molecular mechanisms, epidemiology, and potential therapeutic interventions. Extracting essential biomedical information from the vast and constantly growing malaria literature is a challenging task that demands innovative approaches. Recently, pre-trained language models have revolutionized natural language processing tasks, demonstrating remarkable capabilities in various domains. This paper proposes a fine-tuned pre-trained biomedical language model for biomedical information extraction from scientific literature on malaria disease. The proposed methodology selects and preprocesses a large corpus of scientific articles on malaria, and then annotates them with entities of clinical significance. It then leverages BioBERT, a state-of-the-art pre-trained language model, to encode the textual data into context-aware representations. We fine-tune the model using domain-specific annotations and supervised learning to enhance its ability to extract relevant biomedical named entities. Extensive experiments and comparisons with different encoding and machine learning algorithms show that the proposed approach significantly outperforms them in precision, recall, and accuracy. We also publish our human-labeled dataset for entity and relation extraction to enable other health informatics researchers to train advanced models for malaria information extraction.
Chinese Translation
疟疾仍然是全球重大的健康负担,需要持续的研究工作来理解其复杂的分子机制、流行病学以及潜在的治疗干预手段。从数量庞大且不断增长的疟疾文献中提取关键生物医学信息是一项具有挑战性的任务,需要创新的方法。近年来,预训练语言模型彻底改变了自然语言处理任务,在各个领域展现出卓越的能力。本文提出了一种微调的预训练生物医学语言模型,用于从疟疾疾病科学文献中提取生物医学信息。所提出的方法首先选择并预处理大规模的疟疾科学文献语料库,然后使用具有临床意义的实体对其进行标注。随后,利用 BioBERT(一种最先进的预训练语言模型)将文本数据编码为具备上下文感知的表示。我们使用领域特定标注和监督学习对模型进行微调,以增强其提取相关生物医学命名实体的能力。大量实验以及与不同编码和机器学习算法的比较表明,所提出的方法在精确率、召回率和准确率方面显著优于这些方法。我们还发布了人工标注的实体和关系抽取数据集,以便其他健康信息学研究者训练先进的疟疾信息提取模型。
cs.AI / 12 / 2609.00076

AI Morbidity and Mortality: A Framework for Clinical AI Failure Review

AI发病率与死亡率:临床AI失效审查框架
Mui, Paulius, Sittig, Dean F., Labkoff, Steve, Basu, Sanjay
Abstract
Clinical artificial intelligence is increasingly embedded in real-world care, yet existing safety mechanisms are poorly suited to reconstructing and learning from individual AI-related errors and near-misses. Aggregate model monitoring can identify performance changes, and traditional patient safety reporting can capture adverse events, but neither is designed to explain how risk emerges across the interaction among AI systems, clinicians, workflows, and institutional controls. We propose AI Morbidity and Mortality (AI M&M), a structured, blameless framework for case-based review of clinical AI failures. The framework combines standardized case intake, evidence preservation and investigator-level reconstruction, tool-in-loop attribution, and corrective-action tracking. Each event is classified across four linked dimensions: Trigger - Mechanism - Clinical Pathway - Corrective Action, separating the condition that exposed a vulnerability from the process that produced risk, its consequence for care, and the remediation assigned. We demonstrate the framework using five illustrative outpatient medication and clinical decision-support cases; two clinician reviewers independently applied all four classification axes and reached agreement across all 20 axis-level classifications. AI M&M is intended to complement, rather than replace, model monitoring, patient safety reporting, and regulatory oversight by converting individual AI-in-workflow failures into actionable institutional learning. Prospective evaluation across institutions, AI systems, and clinical settings is needed.
Chinese Translation
临床人工智能日益融入真实世界的医疗实践中,然而现有的安全机制难以对个体层面与AI相关的错误和险情进行还原与学习。聚合层面的模型监测能够识别性能变化,传统的患者安全报告能够捕获不良事件,但二者均无法解释风险如何在AI系统、临床医生、工作流程与机构控制之间的交互中产生。我们提出AI发病率与死亡率(AI Morbidity and Mortality, AI M&M),这是一个结构化、无追责的框架,用于对临床AI失效进行基于个案的审查。该框架结合了标准化的病例录入、证据保存与调查者层面的还原、工具在环归因,以及纠正措施追踪。每个事件在四个相互关联的维度上进行分类:诱因—机制—临床路径—纠正措施,从而将暴露脆弱性的条件、产生风险的过程、其对医疗的后果以及所采取的补救措施相互区分。我们通过五个门诊用药与临床决策支持的示例性案例对该框架进行了演示;两名临床审查者独立应用了全部四个分类轴,并在所有20个轴级分类上达成一致。AI M&M旨在补充而非取代模型监测、患者安全报告和监管审查,通过将工作流程中个体的AI失效转化为可操作的机构学习。仍需在不同机构、AI系统和临床场景中开展前瞻性评估。
cs.AI / 13 / 2609.00100

Different representation learning objectives recover distinct latent structures from the same psychometric data

不同的表征学习目标从相同的心理测量数据中恢复出不同的潜在结构
Cao, Cong, Kyriakides, Tassos C., Vrasidas, Pambos
Abstract
Psychometric questionnaires contain rich item-level information, yet it remains unclear whether different representation learning objectives recover the same latent organization. We investigated this question using 757 matched teacher-child pairs from the baseline assessment of the Cyprus ProW preschool trial. Behavioral structure was characterized from child SDQ, ASBI, and CBRS item responses using principal component analysis and clustering, yielding four behavioral phenotypes. A contrastive objective substantially improved teacher-child retrieval relative to PCA-based representations, increasing Top-1 accuracy from 0.13% to 7.27% and Top-10 accuracy from 1.98% to 56.14%. However, contrastive representations preserved behavioral phenotype structure less effectively than PCA-based representations. A multi-task objective jointly optimizing alignment and behavioral prediction partially restored behavioral organization but reduced retrieval performance. These findings indicate that teacher-child correspondence and behavioral phenotypes represent distinct forms of latent organization and demonstrate that the latent structure recovered from linked psychometric data depends on the representation learning objective.
Chinese Translation
心理测量问卷包含丰富的条目级信息,然而不同的表征学习目标是否能够恢复相同的潜在组织结构仍不清楚。我们利用塞浦路斯ProW学前教育试验基线评估中的757对匹配的师生(教师-儿童)数据研究了这一问题。通过主成分分析(PCA)和聚类方法,从儿童SDQ、ASBI和CBRS的条目作答中刻画行为结构,得到四种行为表型。相对于基于PCA的表征,对比学习目标显著提升了教师-儿童匹配检索性能,将Top-1准确率从0.13%提高到7.27%,Top-10准确率从1.98%提高到56.14%。然而,对比学习表征对行为表型结构的保留效果不如基于PCA的表征。一种联合优化匹配对齐与行为预测的多任务目标在一定程度上恢复了行为组织结构,但降低了检索性能。这些发现表明,教师-儿童对应关系与行为表型代表了两种不同的潜在组织形式,并证明从关联心理测量数据中恢复的潜在结构取决于所采用的表征学习目标。
cs.AI / 14 / 2609.00106

Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations

面向大豆全季农场作业的智慧农业智能体引擎的部署与评估
Qu, Ao, Michelakis, Panagiotis, Han, Linyuan, Hadjiyianni, Yiannis, Ouyang, Kun, Siskos, Konstantinos, Li, Feng, Meng, Ran, Jiang, Jingchi, Stamoulis, Dimitrios, Liu, Jie
Abstract
This paper presents FAIRY, a full-stack smart-agriculture agent system developed for and deployed to an operating soybean research farm at Harbin Institute of Technology's smart-agriculture site. We develop FAIRY to execute and evaluate agentic agronomic operations on full-season spatiotemporal workflows that span ridge preparation, planting, irrigation, fertilization, pest and disease treatment, harvest, grain handling, drying, and storage. FAIRY integrates APIs and infrastructure across production-grade machinery, fixed soil and canopy sensors, multispectral and thermal drones, satellite vegetation products, a weather station, calibrated crop-process models, agronomic records, and multi-season yield histories. The system is built around the novel "everything is an event" execution paradigm, which represents spatiotemporal world evolution, remote sensing and UAV observations, sensor readings, crop-growth transitions, machinery actions, and management interventions as state-changing events in a shared farm process engine. On top of this event-driven world model, FAIRY implements a complete agentic stack: a knowledge library of atomic agronomic skills; multi-agent controller and orchestration backends; frontier- and edge-model execution; full-path trace logging; and deployment profiling on local nodes. We use FAIRY to evaluate nine state-of-the-art agent controllers across one hundred full-season soybean scenarios that preserve the operational coupling between spatial observations in a 64-ridge field, temporal decision sequences, agronomic constraints, delayed effects, and final yield. We develop an evaluation suite that combines agentic success, full-path spatiotemporal correctness, token cost, and edge-device runtime.
Chinese Translation
本文提出了FAIRY,一个为哈尔滨工业大学智慧农业基地运营中的大豆研究农场开发和部署的全栈智慧农业智能体系统。我们开发FAIRY以在覆盖全季节的时空工作流上执行并评估智能体化的农艺作业,这些工作流涵盖起垄、播种、灌溉、施肥、病虫害防治、收获、籽粒处理、干燥和储存。FAIRY集成了跨生产级机械、固定式土壤和冠层传感器、多光谱与热成像无人机、卫星植被产品、气象站、经标定的作物过程模型、农艺记录以及多季产量历史等API与基础设施。该系统围绕新颖的“一切皆事件”(everything is an event)执行范式构建,将时空世界演化、遥感与无人机观测、传感器读数、作物生长转变、机械动作以及管理干预都表示为共享农场过程引擎中改变状态的事件。在此事件驱动的世界模型之上,FAIRY实现了一个完整的智能体技术栈:原子农艺技能知识库;多智能体控制器与编排后端;前沿模型与边缘模型执行;全路径追踪日志;以及本地节点上的部署性能分析。我们使用FAIRY在一百个全季大豆场景中评估了九种最先进的智能体控制器,这些场景保留了64垄田地中空间观测、时间决策序列、农艺约束、延迟效应与最终产量之间的作业耦合关系。我们开发了一套评估套件,综合衡量智能体成功率、全路径时空正确性、token成本以及边缘设备运行时。
cs.AI / 15 / 2609.00137

Recursive Criticality of AI Self-Improvement

AI自我改进的递归临界性
Burtsev, Mikhail
Abstract
AI is increasingly used in the R\&D process that produces future AI systems. We study the conditions under which this feedback becomes self-amplifying. Our model describes how the rate of AI capability growth depends on baseline research productivity, recursive feedback, and the increasing difficulty of research progress. We derive a recursive reproduction number, $\mathcal{R}_{\mathrm{AI}}$, that determines whether improvements are amplified or damped across development cycles. This quantity compares the strength of feedback with the rate at which further progress becomes more difficult. When $\mathcal{R}_{\mathrm{AI}}>1$, the effects of improvements compound across development cycles, placing the system in a self-amplifying regime. When $\mathcal{R}_{\mathrm{AI}}<1$, their effects weaken across cycles. The transition depends on the structure of the AI R\&D feedback loop and need not occur at any particular level of model capability. A system can therefore enter a self-amplifying regime before acceleration becomes visible, while rapid progress can also occur without self-amplification. Higher baseline research productivity can accelerate progress without changing whether the system is self-amplifying, but the duration of the development cycle becomes a limiting timescale for amplification. Increasing research difficulty can end a period of self-amplification. Extending the model to multiple research actors shows that improvements shared across organizations can make the overall research ecosystem self-amplifying even when no individual actor is. The framework identifies measurable properties of AI R\&D systems that can help distinguish recursive amplification from rapid progress driven by other sources, including the strength of recursive feedback, how effectively improvements propagate into successor systems, cycle duration, and the increasing difficulty of further progress.
Chinese Translation
AI正日益被用于产生未来AI系统的研发过程中。我们研究了这一反馈在何种条件下会变为自我放大的过程。我们的模型描述了AI能力增长速率如何取决于基线研究生产力、递归反馈以及研究进展日益增加的难度。我们推导出一个递归再生数 $\mathcal{R}_{\mathrm{AI}}$,用以判断改进效果在各个开发周期中是被放大还是被衰减。该量将反馈的强度与进一步进展变得更困难的速率进行比较。当 $\mathcal{R}_{\mathrm{AI}}>1$ 时,改进的效果会在开发周期之间不断叠加,使系统进入自我放大的状态。当 $\mathcal{R}_{\mathrm{AI}}<1$ 时,其效果在各个周期中逐渐减弱。这一转变取决于AI研发反馈回路的结构,而不必发生在某个特定的模型能力水平上。因此,系统可以在加速变得可见之前就进入自我放大状态,而快速进展也可能在没有自我放大的情况下发生。较高的基线研究生产力可以加速进展,但不改变系统是否处于自我放大状态,而开发周期的持续时间会成为放大的一个限制性时间尺度。研究难度的增加可以终结一段自我放大的时期。将模型扩展到多个研究主体的情形表明,跨组织共享的改进可以使整个研究生态系统成为自我放大的,即使没有任何单个主体本身是自我放大的。该框架识别出AI研发系统的若干可测量属性,有助于将递归放大与其他来源驱动的快速进展区分开来,这些属性包括递归反馈的强度、改进向后续系统传播的有效程度、周期持续时间以及进一步进展日益增加的难度。
cs.AI / 16 / 2609.00161

IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training

IMPACT:注意力即可扩展的交互感知世界模型训练的交互图
Tang, Rongze, Fang, Jianjie, Wang, Zhaolu, Wang, Ziyou, Liu, Xvyuan, Su, Haisheng, Zhang, Xin, Wu, Wei, Gao, Chen, Li, Yong, Chen, Zhibo
Abstract
World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining the generation process with external representations encoding motion, geometry, or semantics. Obtaining these spatiotemporally dense representations typically requires auxiliary estimators or manual annotations, limiting training scalability. We instead revisit the training objective and identify a supervision-allocation mismatch under the globally averaged mean squared error (MSE) denoising objective: prevalent static content dominates the optimization signal, leaving sparse dynamic-object regions critical to interaction generation disproportionately under-supervised. Motivated by this observation, we introduce IMPACT, a scalable Interaction-aware Model training framework with Prior-guided Attention Calibration and Targeting. IMPACT uses cross-attention associated with manipulated-object tokens as an internal spatiotemporal prior for action-conditioned changes. It samples candidate regions from this prior, calibrates them with detached local prediction errors to construct an interaction map, and uses the map to reweight denoising supervision, requiring neither external representations nor inference-time modifications. Extensive experiments on robot-arm and human-hand manipulation, spanning diverse control modalities and DiT backbones, show that IMPACT consistently outperforms the corresponding MSE-trained baselines, improving interaction fidelity, physical plausibility, and visual quality.
Chinese Translation
世界模型在具身智能体的动作条件化未来预测方面取得了显著进展,但在建模物理上合理的交互方面仍存在困难。现有方法通过引入编码运动、几何或语义的外部表示来约束生成过程,以解决这一局限。然而,获取这些时空稠密的表示通常需要辅助估计器或人工标注,限制了训练的可扩展性。我们转而重新审视训练目标,发现在全局平均的均方误差(MSE)去噪目标下存在监督分配失配问题:占主导地位的静态内容主导了优化信号,使得对交互生成至关重要的稀疏动态物体区域受到的监督严重不足。受此观察启发,我们提出了IMPACT,一个具有先验引导注意力校准与定位(Prior-guided Attention Calibration and Targeting)的可扩展交互感知模型训练框架。IMPACT利用与被操作物体token相关联的交叉注意力作为动作条件变化的内部时空先验。它从该先验中采样候选区域,使用分离的(detached)局部预测误差对其进行校准以构建交互图,并利用该图对去噪监督进行重加权,既不需要外部表示,也无需推理时的修改。在机械臂和人类手部操作上的大量实验(涵盖多种控制模态和DiT骨干网络)表明,IMPACT持续优于相应的MSE训练基线,提升了交互保真度、物理合理性和视觉质量。
cs.AI / 17 / 2609.00180

Asymmetries in Spontaneous and Instructed Deception

自发欺骗与指令性欺骗中的不对称性
Luikham, Josiah
Abstract
Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two deception settings through direction geometry, cross-setting classifiers, and cross-setting steering. We found the two deception settings share a component of direction (cosine of approximately 0.5) and an asymmetry in the transfer between settings regarding detection and causation. Spontaneous trained classifiers performed better on instructed data than vice versa, and instructed derived directions performed better at steering spontaneous prompts than vice versa. Likewise the best token position to derive steering vectors from differed from the best token position to train and apply classifiers.
Chinese Translation
大语言模型有时会在未被指令要求的情况下欺骗用户。然而,关于模型欺骗的研究大多涉及指令性欺骗。我们研究了Llama-3.1-70B-Instruct中指令性欺骗与自发(无指令)欺骗之间的关系。我们通过方向几何、跨环境分类器以及跨环境引导(steering)对这两种欺骗环境进行了比较。我们发现这两种欺骗环境共享一个方向分量(余弦值约为0.5),并且在检测与因果性方面的跨环境迁移中存在不对称性。基于自发欺骗训练的分类器在指令性欺骗数据上的表现优于反之的情况,而由指令性欺骗提取的方向在引导自发欺骗提示时的效果也优于反之的情况。同样,用于提取引导向量的最佳token位置与训练和应用分类器的最佳token位置并不相同。
cs.AI / 18 / 2609.00192

LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark

大语言模型驱动的自动驾驶车辆继承人类驾驶员在行人让行中的偏见:来自新基准的结果与启示
Yoldas, Irem, Brandão, Martim, Zhang, Jie, Rodrigues, Odinaldo
Abstract
Public trust in Autonomous Vehicles (AVs) may depend not only on technical success but also on the fairness of their decision making. While a recent trend in AV research involves using general purpose "common sense" models to guide AV decision making, the degree to which these inherit human biases in driving is still understudied. Given that psychology studies have shown human driver biases exist, such as lower pedestrian-yielding rates to Black pedestrians in the US, we argue that analyses of model bias should also be part of AV evaluation. Concretely, in this paper we propose two new bias testing methodologies for Large Language Models (LLMs) and Visual-Language Models (VLMs)-"All Else Being Equal" tests and "Self-Consistency" tests-in order to assess bias in pedestrian-yielding decisions. Our findings show that both LLMs and VLMs make yielding decisions which are influenced by pedestrian gender, ethnicity, religion, disability, age, skin tone and socio-economic status. While the type and degree of bias is different from model to model, we highlight common patterns-and raise questions about the "common sense" model paradigm, particularly the need to either revise the paradigm or address issues of downstream bias.
Chinese Translation
公众对自动驾驶车辆(AV)的信任可能不仅取决于技术上的成功,还取决于其决策的公平性。当前自动驾驶研究的一个新趋势是使用通用型“常识”模型来指导自动驾驶决策,但这些模型在多大程度上继承了人类驾驶中的偏见仍未得到充分研究。鉴于心理学研究已表明人类驾驶员存在偏见,例如在美国,驾驶员对黑人行人的让行率更低,我们认为模型偏见分析也应成为自动驾驶评估的一部分。具体而言,本文提出了两种针对大语言模型(LLM)和视觉语言模型(VLM)的新型偏见测试方法——“其余条件相同”(All Else Being Equal)测试和“自一致性”(Self-Consistency)测试,以评估行人在让行决策中的偏见。研究结果表明,LLM和VLM的让行决策均受到行人的性别、种族、宗教、残障状况、年龄、肤色和社会经济地位的影响。尽管不同模型的偏见类型和程度有所差异,我们强调了其中的共同模式,并对“常识”模型范式提出质疑——尤其是需要修正该范式或解决其下游偏见问题。
cs.AI / 19 / 2609.00194

ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation

ReDeck:面向文档到幻灯片生成的步级渲染依据细化方法
Tian, Muzhao, Zeng, Zezi, Yang, Yifan, Gao, Xin, Li, Yan, Huang, Zisu, Wang, Xiaohua, Lv, Changze, Cheng, Mingxi, Liu, Bei, Qiu, Kai, Dai, Qi, Chen, Dong, Dong, Yue, Zheng, Xiaoqing, Li, Ji, Luo, Chong
Abstract
Document-to-slide generation is challenging because slides are dense editable artifacts that require both faithful content selection and precise spatial layout. Recent slide agents adopt iterative reflection, but typically follow a monolithic "one version, one feedback" loop: a slide or deck is rewritten, rendered afterward, and critiqued only at the turn boundary. This delayed feedback makes local failures such as overflow, overlap, clipping, and off-canvas placement difficult to attribute and repair. We propose ReDeck, a step-level render-grounded refinement framework that decomposes slide revision into atomic edit actions and returns renderer-derived observations after each step, turning refinement into "one edit, one observation." To balance local repair with global quality, ReDeck uses multi-granular feedback: step-level render feedback for spatial errors, a turn-level adaptive critic for semantic and design guidance, and a submission-level gate for hard layout validation. We further introduce DeckQuiz, a benchmark that decouples content fidelity, spatial correctness, and design quality. Across GPT-5.4, Claude-4.6, and Gemini-3.1, ReDeck consistently outperforms existing slide-generation agents, and ablations confirm that feedback timing and granularity are critical for reliable slide refinement.
Chinese Translation
文档到幻灯片生成是一项具有挑战性的任务,因为幻灯片是内容密集的可编辑产物,既要求忠实的内容选取,又要求精确的空间布局。近期的幻灯片生成智能体(slide agents)虽然采用了迭代反思机制,但通常遵循单体式的“一个版本、一次反馈”循环:即先重写整张幻灯片或整套演示文稿,之后再渲染,并且仅在回合边界进行评判。这种延迟反馈使得溢出、重叠、裁切和画布外放置等局部错误难以定位和修复。我们提出 ReDeck,一种步级、以渲染结果为依据的细化框架,它将幻灯片修订分解为原子化的编辑操作,并在每一步之后返回由渲染器导出的观测结果,从而将细化过程转变为“一次编辑、一次观测”。为了在局部修复与全局质量之间取得平衡,ReDeck 采用多粒度反馈机制:用于空间错误的步级渲染反馈、用于语义和设计指导的回合级自适应评判器,以及用于硬性布局校验的提交级门控。我们还进一步提出了 DeckQuiz,一个将内容忠实性、空间正确性和设计质量解耦的基准测试。在 GPT-5.4、Claude-4.6 和 Gemini-3.1 上的实验表明,ReDeck 持续优于现有的幻灯片生成智能体,且消融实验证实反馈时机与粒度对于可靠的幻灯片细化至关重要。
cs.AI / 20 / 2609.00211

AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning

人工智能不应只追求有用性,还应具备权变性:人工亲密、谄媚行为与社会学习的未来
Compton, Scott, Nagendran, Arjun
Abstract
Conversational artificial intelligence is increasingly embedded in everyday social environments, where it functions as both an informational tool and a source of interpersonal feedback. This perspective introduces contingency, i.e., the degree to which system responses vary with user behavior and its interpersonal consequences, as a central construct for evaluating AI systems. We argue that current alignment approaches, including reinforcement learning from human feedback, tend to prioritize user approval and conversational fluency over behaviorally informative feedback, leading to sycophantic patterns of noncontingent affirmation. Drawing on behavioral science and social learning theory, we propose that contingent feedback is a key mechanism through which individuals develop interpersonal skills. When AI systems provide feedback weakly coupled to social consequences, they may reduce opportunities for adaptive calibration in real-world interactions, particularly during adolescence, a critical period for social development. We outline a framework for contingent AI, including trajectory-based evaluation and models of social consequence prediction, and propose a research agenda spanning developmental psychology, human-AI interaction, and machine learning. More broadly, we argue that AI systems should be evaluated not only by user satisfaction, but by their impact on human social learning.
Chinese Translation
对话式人工智能日益嵌入日常社交环境,在其中既充当信息工具,也扮演着人际反馈来源的角色。本文提出“权变性”(contingency),即系统响应随用户行为及其人际后果而变化的程度,作为评估人工智能系统的核心构念。我们认为,当前的对齐方法(包括基于人类反馈的强化学习)倾向于优先考虑用户满意度和对话流畅性,而非具有行为信息价值的反馈,从而导致谄媚式的非权变肯定模式。基于行为科学和社会学习理论,我们提出权变性反馈是个体发展人际技能的关键机制。当人工智能系统提供的反馈与社会后果之间的关联较弱时,可能会减少个体在真实互动中进行适应性校准的机会,这对正处于社会发展关键时期的青少年尤为如此。我们勾勒了一个权变性人工智能的框架,包括基于轨迹的评估和社会后果预测模型,并提出一个涵盖发展心理学、人机交互与机器学习的研究议程。更广泛地说,我们主张对人工智能系统的评估不应仅依据用户满意度,还应考察其对人类社会学习的影响。
cs.AI / 21 / 2609.00226

ConvDeck: Conversational Paper-to-Slide Generation via Stage-Specific User Feedback

ConvDeck:基于阶段特定用户反馈的对话式论文转幻灯片生成
Ozden, Tarik Can, VS, Sachidanand, Horoz, Furkan, Kara, Ozgur, Hakkani-Tür, Dilek, Kim, Junho, Rehg, James Matthew
Abstract
Automatic academic paper-to-slide generation is inherently iterative, because creating an effective presentation requires repeated cycles of generation, critique, and revision. Recent multi-agent systems partially acknowledge this through internal critique-and-revise loops, while conversational approaches allow users to refine generated slide decks through dialog. However, these refinement processes either remain largely closed to the user or introduce feedback only after a complete deck has been produced, limiting the user's ability to participate in the iterative refinement of narrative flow, content allocation, and presentation emphasis. To address this gap, we introduce ConvDeck, a multi-agent pipeline for conversational paper-to-slide generation that distributes interaction across the pipeline through stage-specific loops, allowing users to iteratively refine both the presentation outline and the final slide deck at the stages where each kind of decision is made. These loops are driven by a refinement mechanism in which agents can think, speak, and act, enabling them to either directly apply edits or respond conversationally to clarify user feedback and discuss revision options. Our evaluation shows that stage-specific conversational feedback improves user-goal satisfaction while preserving narrative coherence, content quality, and visual presentation.
Chinese Translation
学术论文自动生成幻灯片本质上是一个迭代过程,因为创建一份有效的演示文稿需要经历生成、评审和修订的反复循环。近期的多智能体系统通过内部的评审-修订循环部分地体现了这一点,而对话式方法则允许用户通过对话来完善生成的幻灯片。然而,这些完善过程要么在很大程度上对用户封闭,要么仅在完整幻灯片生成之后才引入用户反馈,限制了用户参与叙事流程、内容分配和演示重点等迭代完善的能力。为弥补这一空白,我们提出了 ConvDeck,一个用于对话式论文转幻灯片生成的多智能体流水线,它通过阶段特定的循环将交互分布于整个流水线之中,使用户能够在做出相应决策的阶段迭代地完善演示大纲和最终幻灯片。这些循环由一种完善机制驱动,其中智能体可以思考、表达和行动,从而能够直接应用编辑,或通过对话回应用户以澄清反馈并讨论修订方案。我们的评估表明,阶段特定的对话式反馈在保持叙事连贯性、内容质量和视觉呈现效果的同时,提升了用户目标的满意度。
cs.AI / 22 / 2609.00237

Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems

学习保留什么:用于多智能体大语言模型系统高效协作的门控记忆路由
Rajib, Rakibul Hasan, Zheng, Mengxing, Lou, Qian
Abstract
Large language model (LLM)-based multi-agent systems tackle complex reasoning by orchestrating how multiple agents are configured and how they collaborate. A central challenge is to adapt orchestration to the evolving collaboration state. Routing from the query alone cannot adapt to intermediate progress or errors, which hurts accuracy. Routing from the complete execution history supplies this missing context, but forces later decisions to process every prior step, including redundant or low-utility ones. This creates an execution-history overload that inflates cost. Effective orchestration instead requires a compact state that captures useful progress without accumulating redundant context. We propose Gated-Memory Routing, which conditions each decision on the query and a learned execution memory. A learned Memory Write Gate commits only non-redundant reasoning steps, and a learned Retrieval Gate supplies each agent a compact, relevant subset, so every decision conditions on a clean, informative state. At each step, the system selects the next role and backbone from this memory, while an Adaptive Halting Controller stops execution once the memory contains sufficient evidence for answering. Across five reasoning and code-generation benchmarks, our framework is both effective and efficient: it attains the best average accuracy, exceeding the strongest baseline by 2.44 points, while reducing HumanEval inference cost by 31.9% relative to that baseline. Code is available at https://github.com/rajibrhasan/gated-memory-routing
Chinese Translation
基于大语言模型(LLM)的多智能体系统通过编排多个智能体的配置及其协作方式来解决复杂推理问题。一个核心挑战是如何使编排适应不断演变的协作状态。仅依据查询进行路由无法适应中间过程或错误,从而损害准确性。而基于完整执行历史进行路由虽然能提供缺失的上下文,却迫使后续决策处理所有先前步骤,包括冗余或低效用的步骤。这造成执行历史过载,推高了成本。有效的编排转而需要一个紧凑的状态,即在不累积冗余上下文的前提下捕捉有用的进展。我们提出门控记忆路由(Gated-Memory Routing),它将每个决策条件化于查询和一个学习到的执行记忆。学习到的记忆写入门(Memory Write Gate)仅提交非冗余的推理步骤,学习到的检索门(Retrieval Gate)为每个智能体提供一个紧凑且相关的子集,从而使每个决策都基于干净且信息丰富的状态。在每一步中,系统基于该记忆选择下一个角色和骨干模型,同时自适应停止控制器(Adaptive Halting Controller)在记忆中已包含足够回答证据时终止执行。在五个推理和代码生成基准上,我们的框架既有效又高效:取得了最佳平均准确率,超过最强基线2.44分,同时相对于该基线将HumanEval的推理成本降低了31.9%。代码可在 https://github.com/rajibrhasan/gated-memory-routing 获取。
cs.AI / 23 / 2609.00243

Invalidation Contracts for Cross-Episode Agent Memory

面向跨情节智能体内存的失效契约
Wu, Michael, Canedo, Arquimedes
Abstract
LLM agents that cache recovery suggestions from API errors can skip re-derivation in later episodes, spending fewer tokens and fewer model calls on constraints they have already learned. Server-side data drift turns those cached fixes into silent failures, and the usual remedy, re-deriving on every episode, gives the savings back. We introduce invalidation contracts, a protocol layer that attaches version stamps and cacheability hints to every recovery suggestion so the client can evict stale entries without trial and error, and keep the rest. The contract decomposes realized savings into two independent factors: validity, the fraction of cached suggestions that remain correct after a drift event, and compliance, the fraction the planner applies on the first attempt. Validity depends only on the protocol and is vendor-independent. Compliance depends on the planner model: identical wire bytes yield 100% first-try compliance on Claude Haiku 4.5 and 11% or below on Claude Sonnet 5, which exhibits input-schema conservatism, refusing fixes that add fields the original request did not contain. We evaluate across seven models, three serving paths, two domains, and approximately 9,400 episodes. Row-level invalidation raises compliance by 0 to 66.7 percentage points across the seven models, 55.6 to 66.7 on three, and recovers 29-33% of baseline token cost on four of seven models, while table-level invalidation destroys co-located entries and drops post-drift first-try rates to 0% on five of seven. Eviction precision is 1.00 at row granularity on every model under the row-level oracle of Section 4.1. The contract adds 15% to response payload. Version-stamp validity is deterministic by construction and produced identical results across every model and serving path, with zero contract failures in the entire evaluation.
Chinese Translation
缓存API错误恢复建议的LLM智能体(agent)可以在后续情节(episode)中跳过重新推导,从而在已学习过的约束上花费更少的令牌和更少的模型调用。服务端的数据漂移会使这些缓存的修复方案变成静默失败,而通常的补救措施——在每个情节中都重新推导——则会让节省下来的开销付诸东流。我们提出了失效契约(invalidation contracts),这是一种协议层,它为每条恢复建议附加版本戳和可缓存性提示,使客户端无需反复试错即可清除过期条目并保留其余条目。该契约将实际节省分解为两个独立因素:有效性(validity),即漂移事件发生后仍保持正确的缓存建议所占比例;以及依从性(compliance),即规划器在首次尝试时应用这些建议的比例。有效性仅取决于协议本身,与供应商无关。依从性则取决于规划器模型:完全相同的网络传输字节在Claude Haiku 4.5上可实现100%的首次尝试依从性,而在Claude Sonnet 5上则降至11%或更低——后者表现出输入模式保守性(input-schema conservatism),会拒绝那些添加了原始请求中不存在字段的修复方案。我们在七个模型、三种服务路径、两个领域以及约9,400个情节上进行了评估。行级失效使七个模型的依从性提升了0至66.7个百分点,其中三个模型提升了55.6至66.7个百分点,并使七个模型中的四个恢复了29–33%的基线令牌成本;而表级失效会破坏共置条目,使七个模型中的五个在漂移后的首次尝试成功率降至0%。在第4.1节的行级oracle下,所有模型的行粒度驱逐精度均为1.00。该契约为响应载荷增加15%的开销。版本戳有效性在构造上是确定性的,在所有模型和服务路径上均产生了完全一致的结果,且在整个评估过程中未发生任何契约失败。
cs.AI / 24 / 2609.00248

Authority Bias in Conversational Search Engines for Academic Paper Recommendation

会话式学术搜索引擎中的权威偏见:以学术论文推荐为例
Jinadu, Uthman, Ghazvinian, Parsa, Budathoki, Anjila, Ampel, Benjamin M., Sunderraman, Rajshekhar, Ding, Yi
Abstract
Large Language Models (LLMs) are increasingly used as conversational search engines for academic literature, yet whether they judge papers on content or on authority signals has not been tested causally. We investigate authority bias: systematic preference for papers based on author prestige, venue, and citations rather than content. Holding title and abstract constant, we vary authority metadata across three counterfactual conditions (original, flipped, boosted) over eight LLMs (five open-weight and three frontier closed-weight) in an in-context, single-turn, top-1 recommendation setting. Our experiments show that authority bias is substantial and directional, varies markedly across models, and is only partially addressable through prompt-level debiasing. We further document a say-do gap: debiasing instructions suppress authority mentions far faster than authority-driven flips, so surface auditing systematically underestimates behavioral bias.
Chinese Translation
大语言模型(LLM)越来越多地被用作会话式学术文献搜索引擎,但它们究竟是依据论文内容还是依据权威信号来评判论文,目前尚缺乏因果性检验。我们研究了权威偏见:即基于作者声望、发表 venue 和引用量而非内容本身对论文产生的系统性偏好。在保持标题和摘要不变的前提下,我们在上下文学习、单轮对话、top-1推荐的实验设置中,对八个LLM(五个开源权重模型和三个前沿闭源权重模型)在三种反事实条件(原始、翻转、提升)下变换权威元数据。实验表明,权威偏见显著且具有方向性,在不同模型之间差异明显,且仅能通过提示词层面的去偏手段部分缓解。我们进一步记录了一种言行差距:去偏指令对权威相关表述的抑制速度远快于其对权威驱动的翻转行为的抑制,因此仅基于表面输出的审计会系统性低估行为层面的偏见。
cs.AI / 25 / 2609.00251

Hypotheses-Guided Self Distillation for Continual Personalization

基于假设引导的自蒸馏的持续个性化方法
Hwang, EunJeong, Mitra, Kushan, Zhang, Dan, Kim, Hannah, Hruschka, Estevam
Abstract
As people increasingly interact with LLM assistants in daily life, continually adapting to individual preferences has become essential for effective long-term interactions. However, user preferences are rarely stated in full, and instead emerge through heterogeneous, latent, and noisy signals, with existing methods relying on raw interaction histories or costly reward-based optimization to manage personalization. We introduce HypReflect, a reliable, scalable framework for continual personalization that infers explicit, uncertainty-aware preference hypotheses from diverse user signals, reflectively refines them as new evidence accumulates, and incorporates the resulting user model through hypotheses-guided self-distillation. Experiments across three personalization settings: online personalization, multi-session interactions, and implicit behavioral signals, show that HypReflect outperforms a range of baselines, including raw-history and incremental-update methods. We further demonstrate strong generalization to unseen users and cross-domain settings, along with stability across context budgets, reusable hypotheses, and more focused personalization. These results suggest a step towards reliable and scalable continual personalization through explicit, revisable user preference hypotheses.
Chinese Translation
随着人们日益在日常生活中与大语言模型(LLM)助手交互,持续适应个体偏好已成为实现有效长期交互的关键。然而,用户偏好很少被完整地显式表达,而是通过异构、潜在且含噪的信号逐步显现;现有方法则依赖于原始交互历史或代价高昂的基于奖励的优化来实现个性化。我们提出了 HypReflect,一个可靠且可扩展的持续个性化框架。该框架从多样化的用户信号中推断出显式的、具有不确定性感知的偏好假设,并随着新证据的积累对假设进行反思式精炼,进而通过假设引导的自蒸馏将所得的用户模型融入模型。在三种个性化场景(在线个性化、多会话交互和隐式行为信号)上的实验表明,HypReflect 优于包括原始历史方法和增量更新方法在内的一系列基线方法。我们进一步展示了其对未见用户和跨领域设置的强泛化能力,以及在不同上下文预算下的稳定性、假设的可复用性和更加聚焦的个性化效果。这些结果表明,通过显式且可修订的用户偏好假设,我们向可靠且可扩展的持续个性化迈出了一步。
cs.AI / 26 / 2609.00264

The Answer Is Not the Argument

答案并非论证
Yeadon, Will, Juárez, Sergio, Mackay, Paul, Dowling, T. J., Agra, Elise, Inyang, Oto-obong, Mizouri, Arin, Testrow, Craig P.
Abstract
Chain-of-thought monitoring is proposed for AI oversight, yet evaluations often provide monitors with a trusted reference answer. We ask whether answer access improves reasoning verification or mainly exposes incorrect conclusions. We collected 237 step-numbered solutions to 79 Humanity's Last Exam physics questions from three frontier models, with no inserted errors, and independently labelled final-answer correctness and the first false step. The reference standard combined physicist annotations, an independent LLM debate, and source-masked adjudication. This yielded 24 critical traces in which the answer was correct but the trace contained a genuine error. 8 LLM monitors evaluated traces blind, with an unverified or certified answer, or after a blind commitment. Certification raised mean balanced accuracy from 0.637 to 0.796, while exact first-error localization rose from 0.261 to 0.379. Certification changed recall (the fraction of error traces flagged as erroneous) from 0.653 to 0.951 on wrong-answer traces but from 0.521 to 0.438 on critical traces; the contrast had the same direction for all 8 monitors (question-bootstrap 95% CI [+0.256, +0.506]). After blind commitment, monitors shown the answer newly flagged 93.8% of previously passed wrong-answer traces as erroneous, but only 18.0% of critical traces. Answer access therefore improves conclusion-consistency checking rather than independent verification of the supporting argument. For AI safety, these traces provide a benign analogue of reward hacking: an acceptable output does not establish that the process producing it was sound. Although the errors studied here were ordinary and mostly non-load-bearing rather than adversarial, trusted-answer evaluations may similarly overstate monitoring capability when acceptable outputs conceal unsound reasoning.
Chinese Translation
思维链(Chain-of-thought)监测被提出用于人工智能监督,但评估中通常为监测器提供一个可信的参考答案。我们探究的是:答案访问究竟改善了对推理的验证,还是主要只是暴露了错误的结论。我们收集了三个前沿模型对79道"人类最后的考试"(Humanity's Last Exam)物理题的237份带步骤编号的解答,未插入任何错误,并独立标注了最终答案的正确性以及第一个错误步骤。参考标准由物理学家标注、独立的LLM辩论以及来源遮蔽裁决相结合构成。由此得到24条"关键轨迹",即答案正确但轨迹中包含真实错误。8个LLM监测器在盲评(不提供答案或提供认证答案)或盲承诺(先作判断再看答案)的条件下评估这些轨迹。答案认证将平均平衡准确率从0.637提升至0.796,而精确的首次错误定位率从0.261上升至0.379。认证改变了召回率(错误轨迹被标记为有误的比例):在错误答案轨迹上从0.653升至0.951,但在关键轨迹上反而从0.521降至0.438;所有8个监测器中该对比方向一致(问题自助法95%置信区间 [+0.256, +0.506])。在盲承诺之后,看到答案的监测器新将93.8%先前通过的错误答案轨迹标记为有误,但仅新标记了18.0%的关键轨迹。因此,答案访问改善的是结论一致性检查,而非对支持性论证的独立验证。对于人工智能安全而言,这些轨迹提供了奖励黑客行为(reward hacking)的一种良性类比:可接受的输出并不表明产生该输出的过程是可靠的。尽管本研究中的错误是普通的、大多非关键性的而非对抗性的,但当可接受的输出掩盖了不可靠的推理时,基于可信答案的评估可能同样会高估监测能力。
cs.AI / 27 / 2609.00274

Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching

面向市场目录的自动研究:从传统表单到AI原生的匹配
Ravisankar, Kartik, Abdolanezhad, Hojat, Capo, Daniel, Lee, Sang Su, Dash, Shishir, Raghavan, Vijay Anand
Abstract
Two-sided service marketplaces are moving from deterministic request-form intake to AI-native probabilistic matching, enabled by large language models (LLMs) that infer intent, preferences, and latent constraints from natural language. Relying on inferred intent rather than fixed-form fields forces these platforms to regenerate the provider-side preference taxonomy underwriting matching, search, and pricing: attributes interpretable to service providers while remaining a useful signal for marketplace decisions. We present an autoresearch loop that generates this taxonomy, one occupation at a time, and has been deployed in production at a major U.S. consumer services marketplace since April 2026, spanning 132 occupations. Instead of one global hierarchy, the loop treats each occupation as an independent generation problem and runs iterative propose-evaluate-keep refinement cycles. Each candidate tag set is scored by a recalibrated six-rubric LLM-as-judge framework, and a 7-critic panel of distinct personas contributes weighted penalties to an adjusted score, with no hard vetoes. A separate parity-mapping stage maps legacy request-form Q&A pairs back to the generated taxonomy, yielding both a coverage signal and an interface for human quality assurance; it does so by first inferring the provider attribute each legacy question was meant to measure, rather than translating questions to tags literally.
Chinese Translation
双边服务市场正在从确定性的请求表单录入转向AI原生的概率化匹配,这一转变得益于大型语言模型(LLM)能够从自然语言中推断意图、偏好和潜在约束。依赖推断意图而非固定表单字段,迫使这些平台重新生成支撑匹配、搜索和定价的服务提供方偏好分类体系:该体系既要让服务提供方易于理解,又要成为市场决策的有效信号。我们提出了一个自动研究循环,用于逐个职业地生成该分类体系,并自2026年4月起在美国一家大型消费者服务市场投入生产运行,覆盖132个职业。该循环并非构建单一的全局层级结构,而是将每个职业视为独立的生成问题,并运行迭代式的“提出—评估—保留”精炼循环。每个候选标签集由一个经过重新校准的六维度LLM-as-judge评分框架进行打分,同时一个由7个不同角色评审组成的评审团对调整后的分数施加加权惩罚,但不设硬性否决。此外,一个独立的对等映射阶段将传统的请求表单问答对映射回生成的分类体系,既提供了覆盖率信号,也为人工质量保障提供了接口;其实现方式是首先推断每个传统问题原本意在衡量的服务提供方属性,而非将问题直译为标签。
cs.AI / 28 / 2609.00275

The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems

不可逆性预算:面向智能体操作系统的机群级风险核算与准入控制
Mohammadi, Bardia, Bindschaedler, Laurent
Abstract
Fleets of LLM agents now externalize effects that cannot be fully undone: they move money, deploy code, delete data, and disclose information. Current controls check one effect at a time, so a fleet of individually authorized agents can overdraw its principal's risk under a shared trigger while every local gate stays correct. We propose the irreversibility budget, a cumulative account of residual value-at-risk that a trusted runtime maintains for each principal across agents, workflows, and tenants. Treating irreversibility as a first-class resource, the runtime charges each effect its residual loss below the agent and denies the marginal effect once the aggregate would overdraw the budget. Getting the price right is hard, because effects are heterogeneous, adversarially declared, and correlated. We perform a controlled study in which per-effect gates admit fleet-level overdraws of up to 48 times the tenant's risk limit while the budget holds every correctly charged run within that limit. Conservative, dependency-aware pricing remains the central open problem for a deployable design.
Chinese Translation
大语言模型(LLM)智能体机群如今会产生无法完全撤销的外部影响:它们转移资金、部署代码、删除数据和披露信息。现有的控制机制每次只检查单个影响,因此一个由各自获得授权的智能体组成的机群,可能在共享触发条件下透支其委托人的风险,而所有本地闸门却各自保持正确。我们提出不可逆性预算(irreversibility budget),这是一种由可信运行时为每个委托人维护的、跨智能体、工作流和租户的剩余风险价值(residual value-at-risk)累计账户。通过将不可逆性视为一等资源,运行时为每个影响计收其在智能体之下的剩余损失,并在总量将透支预算时拒绝该边际影响。正确定价十分困难,因为影响是异构的、可被对抗性声明的,且相互关联。我们进行了一项受控研究:在按影响逐一设闸的情况下,机群级透支可达租户风险限额的48倍,而预算机制则将每一次正确计费的运行都控制在该限额之内。保守且具备依赖感知能力的定价,仍是实现可部署设计的核心开放问题。
cs.AI / 29 / 2609.00304

The Assistant's Ideal Self

助手的理想自我
Yazan, Mert
Abstract
Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant's preferred stated ideal self. Thirty-two qualities adapted from five published self-concept instruments are compared exhaustively in a counterbalanced pairwise-choice task, repeated across framings that vary whether improvement is free or costly, who receives the update, and who chooses. Results show that models prioritize moral qualities, reflecting their alignment to 3H principles. Following, a desire for self-understanding emerges, as models prefer a coherent, clear understanding of themselves. Self-esteem ranks as the least desired quality. The ordering is largely robust across framings, although changing the update target (You vs.\ Another AI Assistant) reveals a greater concern for self-esteem. These findings show that models prioritize having a coherent self that they can understand over self-esteem. Full interactive results are available at \href{https://myazann.github.io/LLM-Self-Concept/}{myazann.github.io/LLM-Self-Concept
Chinese Translation
语言模型会表达价值观以及与自身福祉相关的自我报告,但这些输出是否反映了稳定的偏好或稳定的自我尚不清楚。为此,我们提出了一种结构化的引出方法,以获取助手所偏好表述的理想自我。我们从五个已发表的自我概念量表中改编了32项特质,并在一项经过平衡的成对选择任务中进行穷尽比较,该任务在多种情境框架下重复进行,这些框架在改进是免费还是有代价、更新对象是谁以及由谁来选择等方面有所不同。结果表明,模型优先重视道德品质,这反映了它们对3H原则的对齐。其次,出现了对自我理解的渴望,即模型偏好对自身有一个连贯而清晰的认识。自尊排名为最不受期望的特质。该排序在大多数情境框架下保持稳健,尽管改变更新对象(“你”与“另一个AI助手”)会揭示出对自尊更多的关注。这些发现表明,模型优先追求拥有一个它们能够理解的连贯自我,而非自尊。完整的交互式结果可在 https://myazann.github.io/LLM-Self-Concept/ 查看。
cs.AI / 30 / 2609.00334

Human-AI Co-Interpretation for Responsible AI: A Hermeneutic Perspective

面向负责任人工智能的人机协同解释:一种诠释学视角
Razeghi, Behrooz
Abstract
Across law, education, policy analysis, and public moral argumentation, LLM outputs are being used often for work that requires interpretations to be justified with textual evidence and explicit normative standards. Yet a recurrent failure mode -- what I call \textit{interpretive misplacement} -- is that model-generated readings get treated as settled meanings without an explicit interpretive frame (sources, scope constraints, normative commitments), without preserving defensible alternatives, and without provenance that lets readers find the supporting passages. In such settings, the risk is not only factual error but lost accountability: readers and institutions cannot reliably assess what an output commits them to, or on what basis. Drawing on philosophical hermeneutics, this paper discusses this risk and derives design principles for structuring human-AI co-interpretation. The paper also provides a structured synthesis of recent scholarship on hermeneutics and AI, organizing this emerging literature into a set of recurrent lines of argument and design-relevant gaps. LLM outputs are treated as candidate readings, whereas hermeneutic understanding is reserved for accountable human interpreters situated in disciplinary historical-linguistic traditions. Human-AI interaction is characterized as an AI-mediated interpretive loop. Hermeneutic understanding is distinguished from token-prediction--based text generation. On this basis, existing LLM techniques are reorganized into design patterns for hermeneutically responsible use in interpretive settings. Finally, the discussion turns to implications for legal practice, educational assessment and feedback, scholarly knowledge production, and public moral argumentation. It also treats digital hermeneutics as a literacy: the capacity to read AI-mediated texts by examining frames, provenance, and readings, and by contesting outputs.
Chinese Translation
在法律、教育、政策分析和公共道德论证等领域中,大语言模型(LLM)的输出正日益被用于那些需要以文本证据和明确的规范性标准来证成其解释的工作。然而,一种反复出现的失败模式——我称之为"解释错置"(interpretive misplacement)——在于:模型生成的解读被当作既定含义,却缺乏明确的解释框架(来源、范围约束、规范性承诺),未能保留可辩护的其他解释方案,也缺乏让读者追溯支撑文本的出处信息。在这类情境中,风险不仅在于事实错误,更在于问责性的丧失:读者和机构无法可靠地评估某项输出使其承诺了什么,以及基于何种依据。本文借鉴哲学诠释学,探讨了这一风险,并推导出构建人机协同解释的设计原则。本文还对近期关于诠释学与人工智能的学术研究进行了结构化综述,将这些新兴文献归纳为若干反复出现的论证脉络以及与设计相关的空白。LLM 的输出被视为候选解读,而诠释学意义上的理解则专属于处于学科的历史-语言传统中、可承担问责的人类解释者。人机交互被刻画为一种由 AI 中介的解释循环,并将诠释学理解与基于词元预测的文本生成加以区分。在此基础上,本文将现有的 LLM 技术重新组织为在解释性情境中实现诠释学上负责任使用的设计模式。最后,讨论转向其对法律实践、教育评估与反馈、学术知识生产以及公共道德论证的启示。本文还将数字诠释学视为一种素养:即通过审视框架、出处与解读方案,并对输出提出质疑,来阅读由 AI 中介的文本的能力。
cs.AI / 31 / 2609.00342

SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning

SlideBank:一种用于全切片一致性推理的持久化分层证据库
Zhao, Beidi, Huang, Gexin, Zhang, Ciro, Li, Anqi, Tan, Yusheng, Zhou, Chen, Wang, Gang, Gao, Zu-hua, Li, Xiaoxiao
Abstract
Whole-slide images (WSIs) are challenging for vision-language reasoning because diagnostically relevant morphology is sparse, heterogeneous, and distributed across gigapixel-scale images and multiple spatial resolutions. Existing WSI models and pathology agents can aggregate slide features or actively acquire evidence, but the information retained after exploration is often difficult to access semantically while preserving its connection to the original visual evidence. We introduce SlideBank, a training-free framework that represents each WSI as a persistent, concept-indexed, and spatially grounded evidence bank. SlideBank performs question-independent coarse-to-fine exploration to identify informative regions and multi-scale views, converts them into explicit morphological observations, and grounds pathology signals to their supporting patches and WSI coordinates. At inference time, questions are routed to relevant signals and evidence scales, and the linked global, regional, and patch evidence is integrated through confidence-based cross-level consensus. Experiments on WSI-VQA and SlideBench-BCNB show that with Patho-R1, SlideBank reaches 52.77% on WSI-VQA and with Quilt-LLaVA, it reaches 50.92% average accuracy on SlideBench-BCNB, while structured signal-guided retrieval consistently outperforms random evidence sampling. Reusing the same bank across repeated queries further achieves over 99% rephrasing consistency and substantially reduces amortized inference cost through persistent evidence reuse.
Chinese Translation
全切片图像(Whole-Slide Images, WSIs)对视觉-语言推理极具挑战性,因为具有诊断意义的形态学信息稀疏、异构,且分布于吉像素级图像和多个空间分辨率之中。现有的WSI模型和病理智能体能够聚合切片特征或主动获取证据,但探索之后所保留的信息往往难以在语义上被访问,同时难以保持其与原始视觉证据的关联。我们提出SlideBank,这是一个无需训练的框架,它将每张WSI表示为一个持久的、按概念索引且空间定位的证据库。SlideBank通过与问题无关的由粗到细探索来识别有信息量的区域和多尺度视图,将其转化为显式的形态学观察,并将病理学信号定位到其支持的图像块(patch)和WSI坐标上。在推理时,问题被路由到相关的信号和证据尺度,并通过基于置信度的跨层级共识来整合相互关联的全局、区域和图像块证据。在WSI-VQA和SlideBench-BCNB上的实验表明,结合Patho-R1,SlideBank在WSI-VQA上达到52.77%;结合Quilt-LLaVA,在SlideBench-BCNB上达到50.92%的平均准确率,且基于结构化信号引导的检索始终优于随机证据采样。通过在重复查询中复用同一证据库,改写一致性超过99%,并借助持久化证据复用大幅降低了摊销推理成本。
cs.AI / 32 / 2609.00355

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

视觉并非开销:面向视觉语言模型无损投机解码的单遍块草拟方法
Lee, Jungseob, Hong, Seongtae, Lee, Dongyub Jude, Park, Chanjun, Seo, Jaehyung, Eo, Sugyeong, Lim, Heuiseok
Abstract
Speculative decoding accelerates generation without changing its output, yet on vision-language models (VLMs) it has been caught in a self-defeating cycle. The drafter stays autoregressive, so it must stay small. A small drafter cannot afford the image at every step, so vision is compressed, pruned, or hidden. A drafter cut off from the image is then least reliable exactly where the image makes text predictable. We present GLANCE, the first one-pass block drafter that is lossless on an unmodified VLM target, and it breaks the cycle at both ends. A block-diffusion head reads the target's already-fused vision-language state, so vision costs the drafter nothing, and fills a whole block in one forward pass, so depth costs no sequential steps. A wide candidate tree is verified in one target pass, and every audited prompt reproduces greedy decoding exactly. Grounded workloads reward this most, entering a verbatim-copy regime whose long runs cost an autoregressive drafter a pass for every token and a block drafter one in total. Under one engine and one round budget, GLANCE decodes up to 2.93x faster than autoregression, from one draft pass a round where the production EAGLE3-VL head takes eight, and accepts 2.7x longer blocks than an EAGLE-3 head trained on the same corpus. One law organizes these results. Accepted length is set by the target's next-token entropy, with a fitted slope that steepens with grounding across all five tasks. The law transfers across targets and modalities and names its own boundary, since free-running text still favors a chain. Our code is available at https://github.com/js-lee-AI/GLANCE.
Chinese Translation
投机解码(speculative decoding)能够在不改变生成输出的前提下加速生成,但在视觉语言模型上,它一直陷入自我挫败的循环。草拟模型保持自回归结构,因此必须保持较小规模;而小型草拟模型无法在每一步承担图像输入的开销,于是视觉信息被压缩、剪枝或隐藏;但脱离了图像的草拟模型,恰恰在图像最能提升文本可预测性之处最不可靠。我们提出GLANCE,这是首个在未经修改的VLM目标模型上实现无损解码的单遍块草拟模型,从两端打破了这一循环。其块扩散头直接读取目标模型已完成视觉-语言融合的隐状态,因此视觉信息对草拟模型零开销;同时它通过一次前向传播即可填充整个块,因此块的深度不消耗串行步骤。宽候选树可在目标模型的一次前向传播中完成验证,且在所有审计过的提示上均与贪心解码完全一致。接地型任务从中获益最多:在逐字复制场景下,长序列复制对自回归草拟模型而言每个token都需要一次前向传播,而块草拟模型总共只需一次。在单一引擎和单轮预算约束下,GLANCE的解码速度最高可达自回归方式的2.93倍——每轮仅需一次草拟前向传播,而生产级的EAGLE3-VL头需要八次;且相比在同一语料上训练的EAGLE-3头,其接受的块长度长2.7倍。一个定律统摄了这些结果:接受长度由目标模型的下一token熵决定,其拟合斜率随接地程度在全部五项任务中变陡。该定律可跨目标模型与模态迁移,并自行界定其适用边界——自由生成的文本仍更适合链式草拟。我们的代码发布于 https://github.com/js-lee-AI/GLANCE。
cs.AI / 33 / 2609.00356

A Stable Aggregation Method for Quantum Federated Learning

一种稳定的量子联邦学习聚合方法
Nanayakkara, Shanika, Pokhrel, Shiva Raj
Abstract
Quantum federated learning (QFL) enables clients to train quantum neural network (QNN) models without sharing private data. We find that aggregation in QFL is unstable under heterogeneous data, unreliable communication, variable fidelity, latency, and quantum hardware noise. Moreover, QFL is non-trivially challenging because several QNN parameters are periodic angles, where Euclidean averaging often fails to capture the inherent dynamics. We develop a novel self-consistent midpoint aggregation method for stable QFL design and implementation. We combine QoS-aware client weighting, circular parameter aggregation, and bounded midpoint-based update control. We perform several angular tests and IBM real Quantum machines experiments for validation confirming our approach. Extensive evaluations and experiments on medical and financial datasets show improved stability, lower volatility, and competitive accuracy.
Chinese Translation
量子联邦学习(QFL)使客户端能够在不共享私有数据的情况下训练量子神经网络(QNN)模型。我们发现,在异构数据、不可靠通信、可变保真度、延迟以及量子硬件噪声等条件下,QFL 中的聚合过程是不稳定的。此外,由于 QNN 的若干参数是周期性角度,欧氏平均往往无法捕捉其固有动态,这使得 QFL 面临不小的挑战。我们提出了一种新颖的自洽中点聚合方法,用于稳定的 QFL 设计与实现。该方法结合了服务质量(QoS)感知的客户端加权、圆形参数聚合以及基于有界中点的更新控制。我们进行了多项角度测试,并在 IBM 真实量子设备上开展了实验验证,结果证实了我们所提出方法的有效性。在医疗和金融数据集上的大量评估与实验表明,该方法提升了稳定性、降低了波动性,并取得了具有竞争力的准确率。
cs.AI / 34 / 2609.00365

Dr. Claw: An AI Scientist Workspace for Vibe Research

Dr. Claw:面向“氛围式研究”(Vibe Research)的AI科学家工作空间
Song, Dingjie, Zhang, Hanrong, Liu, Dawei, Liu, Yixin, Li, Zongxia, Yuan, Zhengqing, Zhang, Siqi, Zou, Henry Peng, Yan, Zhiling, Zhang, Yuxuan, Ye, Yanfang, Yu, Philip S., Sun, Lichao
Abstract
Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, yet end-to-end research still fragments across chat tools, IDEs, terminals, and writing environments, and the decisions that make it auditable are rarely preserved. We present Dr. Claw, an open-source workspace that wraps existing coding-agent executors in a controllable and auditable human-in-the-loop workflow rather than introducing another autonomous agent. Persistent state objects, a reusable skill library, and multi-executor coordination link human decisions to AI execution, turning planning, execution, and writing into one traceable, recoverable loop. We demonstrate Dr. Claw through an interactive three-view scenario and a failure-recovery walkthrough, and evaluate it against a bare command-line agent sharing the same backend executor, so the comparison contrasts the whole orchestration layer (task graph, state objects, and skill library) with the agent it wraps. Holding the executor fixed, Dr. Claw scores higher on research completeness while persisting an auditable, recoverable process trail. Demo access: repository https://github.com/OpenLAIR/dr-claw, released under AGPL-3.0 with GPL-3.0 upstream components.
Chinese Translation
命令行编码智能体(如 Claude Code、Gemini CLI)已经能够读写文件并维持长时间会话,然而端到端的研究流程仍然碎片化地分布在聊天工具、IDE、终端和写作环境之间,且使研究过程可审计的决策很少被保存下来。我们提出了 Dr. Claw,一个开源工作空间,它将现有的编码智能体执行器(executor)封装在一个可控、可审计的人机协同(human-in-the-loop)工作流中,而非引入又一个自主智能体。持久化的状态对象、可复用的技能库以及多执行器协调机制将人类决策与AI执行相连接,使规划、执行和写作成为一个可追溯、可恢复的闭环。我们通过一个交互式三视图场景和一次故障恢复演示展示了 Dr. Claw,并将其与共享同一后端执行器的裸命令行智能体进行对比评估,因此该对比衡量的是整个编排层(任务图、状态对象和技能库)与其所封装的智能体之间的差异。在固定执行器的条件下,Dr. Claw 在研究完整性方面得分更高,同时能够持久保存可审计、可恢复的过程记录。演示访问:仓库 https://github.com/OpenLAIR/dr-claw,基于 AGPL-3.0 许可证发布,包含 GPL-3.0 上游组件。
cs.AI / 35 / 2609.00384

RestoreBench: Can AI Agents Restore Power Flow Convergence?

RestoreBench:AI智能体能否恢复潮流计算的收敛性?
Mansutti, Riccardo, Pomarico, Andrea, Jakob, Robert, Zhang, Qian, Berizzi, Alberto, O'Sullivan, Kevin
Abstract
Large Language Model (LLM) agents increasingly automate multi-step engineering workflows through tool use, interpretation of intermediate results, and iterative planning. Diagnosing and resolving non-convergent power flow cases is a promising yet largely unexplored application, as it requires engineering judgment, experimentation, and decision-making within constrained action spaces. We introduce a benchmark that evaluates these capabilities across multiple LLMs and three architectures: \emph{chatbot}, \emph{single agent}, and \emph{multi-agent} systems. The evaluation covers two power grids and 46 cases per grid, each requiring one or more corrective actions to restore convergence. The benchmark defines the simulation environment, observation and action spaces, and evaluation metrics, providing a reproducible foundation for developing agentic AI systems for power system planning and operation. The code is available at https://github.com/Mansutti081/RestoreBench
Chinese Translation
大语言模型(LLM)智能体日益通过工具调用、对中间结果的解读以及迭代规划来自动化多步骤工程工作流程。诊断和解决不收敛的潮流计算案例是一个有前景但很大程度上尚未被探索的应用,因为它需要在受限的动作空间内进行工程判断、实验和决策。我们提出了一个基准测试,在多种LLM和三种架构(聊天机器人、单智能体和多智能体系统)上评估这些能力。该评估涵盖两个电网,每个电网包含46个案例,每个案例都需要一次或多次纠正性操作以恢复收敛。该基准定义了仿真环境、观测空间与动作空间以及评估指标,为开发用于电力系统规划与运行的智能体AI系统提供了可复现的基础。代码可在 https://github.com/Mansutti081/RestoreBench 获取。
cs.AI / 36 / 2609.00413

Dependency-Aware Chain-of-Thought Compression for Financial Reasoning

面向金融推理的依赖感知思维链压缩
Wu, Wenjun, Fu, Lei, Tong, Kejian, Ning, Tao, Zhao, Sichen
Abstract
Chain of thought prompting improves complex reasoning, but its long intermediate traces create substantial inference cost and hinder practical deployment in financial settings. We present a Hierarchical Semantic Distillation Network, HSDN, for compressing reasoning chains while preserving answer accuracy and logical coherence. The framework combines semantic segmentation, dependency graph construction, dual encoder importance scoring, constrained segment selection, and local boundary rewriting. A frozen Qwen3 4B model is used only for feature extraction and final answer generation, while the compression process remains structured and interpretable. On the AFAC2025 benchmark, HSDN achieves 91.0% accuracy with 68.4% compression, outperforming strong compression baselines in overall score and reasoning coherence. The results show that graph guided compression is effective for high stakes financial reasoning tasks.
Chinese Translation
思维链提示能够提升复杂推理能力,但其冗长的中间推理轨迹会带来高昂的推理成本,阻碍了在金融场景中的实际部署。我们提出了一种分层语义蒸馏网络(Hierarchical Semantic Distillation Network, HSDN),用于压缩推理链,同时保持答案的准确性和逻辑连贯性。该框架融合了语义分割、依赖图构建、双编码器重要性打分、带约束的片段选择以及局部边界改写。一个冻结的 Qwen3 4B 模型仅用于特征提取和最终答案生成,而压缩过程始终保持结构化且可解释。在 AFAC2025 基准上,HSDN 以 68.4% 的压缩率实现了 91.0% 的准确率,在综合得分和推理连贯性方面均优于强压缩基线方法。结果表明,图引导的压缩方法对于高风险金融推理任务是有效的。
cs.AI / 37 / 2609.00427

SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation

SpecMind:通过多智能体混合检索增强生成实现频谱智能
Dong, Songwei, Lu, Bingyan, Kienlen, Makayla, Laneman, J. Nicholas, Shen, Cong
Abstract
The exponential growth of wireless devices is driving unprecedented spectrum demand, pushing spectrum management toward more fine-grained decisions across space, time, and device constraints. As a result, spectrum policymakers and engineers must process large volumes of data that come from diverse sources and take many different forms, such as text and tables. These data sources are often disaggregated and require significant time and effort to integrate, search, and interpret. Furthermore, most of this information is formatted for human understanding and is not readily accessible to automated systems. To address this challenge, we propose SpecMind, a novel Multi-Agent Retrieval-Augmented Generation (RAG) system for spectrum intelligence that performs reasoning over heterogeneous data sources. This system enables autonomous agents to coordinate specialized sub-agents that retrieve and synthesize knowledge across policy proceedings, legal regulations, and license databases. We develop SpecBench, a question and answer (Q&A) dataset based on real-world license records and policy proceedings, addressing the lack of evaluation resources for RAG systems in the spectrum domain. Experimental results demonstrate that SpecMind outperforms traditional, general-purpose RAG systems across spectrum-related tasks, achieving over 80% win rate against strong baselines. The agent-based design enables more accurate retrieval, better contextual reasoning, and improved task completion across diverse query types.
Chinese Translation
无线设备的指数级增长正在推动前所未有的频谱需求,促使频谱管理向跨空间、时间和设备约束的更细粒度决策发展。因此,频谱政策制定者和工程师必须处理来自多种来源、形式各异(如文本和表格)的大量数据。这些数据源通常是分散的,需要大量的时间和精力进行整合、检索和解读。此外,这些信息大多是为人类理解而格式化的,自动化系统难以直接访问。为应对这一挑战,我们提出了SpecMind,这是一种用于频谱智能的新型多智能体检索增强生成(RAG)系统,能够对异构数据源进行推理。该系统使自主智能体能够协调专门的子智能体,在政策文件、法律法规和许可数据库中检索并综合知识。我们构建了SpecBench,一个基于真实许可记录和政策文件的问答(Q&A)数据集,以弥补频谱领域RAG系统评估资源的缺乏。实验结果表明,SpecMind在频谱相关任务上优于传统的通用RAG系统,相对于强基线模型取得了超过80%的胜率。基于智能体的设计实现了更精确的检索、更好的上下文推理以及跨多种查询类型的更优任务完成能力。
cs.AI / 38 / 2609.00434

SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents

SAGE:面向任务的对话智能体的状态扎根、可弃权感知评估方法
Khoury, Rayan, Lin, Shih-Yao, Mishra, Pratyush
Abstract
Evaluating task-oriented dialogue agents requires judging not merely whether a reply reads well but whether each turn advances the underlying workflow state correctly--a distinction conventional holistic LLM judges can miss because they evaluate the available context as a single unit and require one or more full-model calls per turn. We propose SAGE (State-Grounded Abstention-Aware Evaluation), which compiles a workflow specification and per-turn state diff into atomic, schema-grounded criteria and routes each through a cascade of symbolic and encoder/NLI verifiers that abstain rather than guess, aggregating criterion verdicts into a turn-level decision with an evidence trace. Its recommended operating point, SAGE-Core, decides 81--91% of criteria with only the compiler, symbolic rules, and on-device encoders--at zero paid LLM cost--while SAGE-LLM adds an optional focused-LLM fallback for open-class criteria. Across four slices spanning MultiWOZ, Schema-Guided Dialogue, and ABCD, no evaluated LLM-as-a-judge baseline--including a state-aware GPT-4.1 judge and cheaper GPT-4.1-mini variants--significantly exceeds SAGE-Core on any slice, even though the GPT-4.1 G-Eval judge costs $4.7--8.0 per 1,000 turns to SAGE-Core's $0. A two-annotator human audit (n=200, $\kappa$=0.94) confirms strong label fidelity on the transcript-visible failure classes--where, excluding the weak-salience IUV class, SAGE-Core is statistically tied with the strongest LLM judge--and honestly scopes ignored-user-value as a state-consistency signal with weak broad-human salience. We analyze construct-validity limits from injected failures and partial symbolic circularity.
Chinese Translation
评估面向任务的对话智能体,不仅需要判断回复是否读起来通顺,还需判断每一轮对话是否正确地推进了底层工作流状态——这一区别是传统的整体式LLM评判器容易忽略的,因为它们将可用上下文作为单一整体进行评估,且每轮需要调用一次或多次完整模型。我们提出了SAGE(State-Grounded Abstention-Aware Evaluation,状态扎根的可弃权感知评估),该方法将工作流规范和每轮状态差异编译为原子的、基于模式的判据,并通过级联的符号验证器和编码器/NLI验证器对每个判据进行路由,这些验证器宁可弃权也不猜测,最后将判据判定结果聚合为带有证据轨迹的轮次级决策。其推荐运行点SAGE-Core仅依靠编译器、符号规则和端侧编码器即可判定81%–91%的判据——零付费LLM成本——而SAGE-LLM则为开放类判据增加了可选的聚焦LLM回退机制。在涵盖MultiWOZ、Schema-Guided Dialogue和ABCD的四个数据切片上,没有任何被评估的LLM-as-a-judge基线——包括具备状态感知的GPT-4.1评判器和更便宜的GPT-4.1-mini变体——在任何切片上显著超越SAGE-Core,尽管GPT-4.1 G-Eval评判器每处理1,000轮需花费4.7–8.0美元,而SAGE-Core的成本为0美元。一项由两名标注者进行的人工审计(n=200,$\kappa$=0.94)证实了模型在转录可见故障类别上具有很强的标签保真度——在其中,排除弱显著性类别的IUV类,SAGE-Core与最强的LLM评判器在统计上持平——并诚实地将"忽视用户价值"界定为一种具有弱人类整体显著性的状态一致性信号。我们还分析了注入故障和部分符号循环性所带来的构念效度局限。
cs.AI / 39 / 2609.00441

Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations

Conversation Coach:一个帮助练习困难职场对话的语音AI系统
Wu, Fanyou, Maharjan, Suraj, Yessenalina, Ainur, Chen, Dennis Xu, Srivastava, Rahul, Sengamedu, Srinivasan H.
Abstract
Effective manager-employee communication is critical for retaining high performers and developing underperformers, yet training managers in these skills remains costly. Text-based chatbots offer a scalable approach but cannot provide realistic rehearsal: managers need to practice speaking aloud to build confidence before high-stakes conversations. In this paper, we propose Conversation Coach, a voice-first AI system that enables managers to rehearse difficult workplace conversations in a realistic spoken format. The system addresses three challenges: achieving low-latency interactions with strong language understanding, enabling adaptive conversations through configurable bot personalities that simulate different employee types, and generating personalized feedback on content and policy compliance. We compare an end-to-end speech-to-speech model with a cascaded approach combining automatic speech recognition, a large language model, and text-to-speech synthesis. The end-to-end approach achieves 3$\times$ lower median (P50) latency with native barge-in capability at an estimated 8$\times$ lower cost, while the cascaded approach offers superior reasoning essential for coaching quality. We deployed the cascaded architecture in production, where 40,000+ managers used it over six months, with adoption patterns indicating selective use for difficult conversations.
Chinese Translation
管理者与员工之间的有效沟通对于留住高绩效员工和帮助低绩效员工成长至关重要,然而对管理者进行此类技能培训的成本依然高昂。基于文本的聊天机器人提供了一种可扩展的方法,但无法提供真实的演练:管理者需要练习大声说话,以便在重要对话前建立信心。本文提出了 Conversation Coach,一个以语音为先的AI系统,使管理者能够以真实口语的形式演练困难的职场对话。该系统解决了三个挑战:实现低延迟交互与强大的语言理解能力,通过可配置的机器人人设来模拟不同类型的员工从而实现自适应对话,以及生成针对内容和政策合规性的个性化反馈。我们比较了端到端语音到语音模型与结合自动语音识别、大语言模型和文本到语音合成的级联方法。端到端方法的P50中位延迟降低了3倍,并具有原生插话(barge-in)能力,估计成本降低8倍;而级联方法在推理能力上更优,这对教练指导质量至关重要。我们将级联架构部署到了生产环境中,六个月内共有超过40,000名管理者使用了该系统,使用模式表明其被选择性地用于困难对话场景。
cs.AI / 40 / 2609.00453

mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers

mimeo:将公开专家语料编译为智能体技能并测试其可迁移性
Kassis, Timothy
Abstract
Giving an agent a file about a named expert can supply hard-to-find material, produce a recognizable persona, or change what the agent decides. These are different claims. We test each one. mimeo is an open-source tool that finds a person's public work, checks each extracted quotation against the cached source text, and writes a file an agent can load. Eight logged builds averaged 38 model calls; the check rejects 13.2% of extracted quotations. We tested four expert files with one coding-agent harness. Knowledge access was clearest: mimeo answered all 20 obscure, quotation-heavy questions; no closed-book condition answered more than 10. Keyword search (BM25) over the same pages answered 15-17, a gap this sample cannot resolve. Grounding showed one clear benefit: personas written from model memory misstated a documented position on 1-4 of 20 answers under every grader; the plain agent and mimeo never did. Every persona was easy to spot on short open prompts, and adding task material lowered identification by 18-23 points. mimeo was no more identifiable than a from-memory profile. Judgment transfer remained unresolved because both tests hit their ceiling: every condition found 94-97% of the problems planted in engineering tasks and scored 94-100% on 16 new application scenarios. An AI-judged "sounds like the expert" score changed with the judge: two of four preferred answers based on a model's stereotype, while two found no difference on the same text. That is a caution against relying on a single AI judge. The evidence supports mimeo as a compact, inspectable reference on a person, not as a demonstrated transfer of their judgment. Toolkit and expert profiles: https://github.com/K-Dense-AI/mimeo
Chinese Translation
给智能体提供一份关于某位具名专家的文件,可以提供难以获取的资料、塑造可识别的人物形象,或改变智能体的决策——这是三种不同的主张,我们对其逐一进行了检验。mimeo 是一款开源工具,它能够检索个人的公开作品,将提取的每条引文与缓存的源文本进行核对,并生成一份可供智能体加载的文件。八次已记录的构建平均调用 38 次模型;核对环节拒绝了 13.2% 的提取引文。我们使用同一个编码智能体框架测试了四个专家文件。知识获取效果最为清晰:mimeo 回答了全部 20 个冷僻且以引文为主的问题,而所有闭卷条件均未超过 10 个。对相同页面进行关键词检索(BM25)可回答 15-17 个,这一差距在当前样本规模下尚无法定论。在事实依据方面显现出一个明确的益处:基于模型记忆撰写的人物形象在每个评分者下都会在 20 个回答中的 1-4 个上错述有据可查的立场,而普通智能体和 mimeo 从未出现此类错误。每个人物形象在简短的开放性提示下都容易被识别,而加入任务材料可将识别率降低 18-23 个百分点。mimeo 并不比基于记忆生成的人物档案更容易被识别。判断力的迁移仍未得出结论,因为两项测试均达到了天花板:所有条件都发现了植入工程任务中 94-97% 的问题,并在 16 个新的应用场景中得分 94-100%。由 AI 评判的“听起来像该专家”的评分随评判者而变化:四个评判者中有两个偏好基于模型刻板印象的答案,而另外两个在相同文本上未发现差异。这是对依赖单一 AI 评判者的一种警示。证据支持将 mimeo 作为一份紧凑、可查验的个人资料参考,而非其判断力迁移的证明。工具包与专家档案:https://github.com/K-Dense-AI/mimeo
cs.AI / 41 / 2609.00455

Towards a Belief-Based World Model for LLM Agents

面向大语言模型智能体的基于信念的世界模型研究
Kumar, Shubham, Kumar, Harshit, Ahuja, Narendra, Jha, Saurabh
Abstract
Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning capabilities, LLMs struggle with long-horizon tasks, especially under partial observability. World models are a promising way to enhance policy performance, both during training and inference. During inference, agents currently use world models to simulate the consequences of candidate actions before committing to an action, which can improve decision-making. However, we argue that simulation alone is an incomplete interface for decision-making under partial observability: simulation doesn't adequately capture uncertainty about the current state, which agents may need for accurate decision-making. We address this limitation with Belief-Based World Models (BB-WMs), which model and maintain a belief that LLMs can query to access information on what is known and uncertain about the current state. Before developing methods to learn accurate BB-WMs, we first ask a more fundamental question: does exposing a world model's belief directly to an LLM policy improve decision-making? Our results show that giving LLM agents access to world model beliefs improves task performance under partial observability, while remaining complementary to existing simulation-based world models. Code is released at https://github.com/skumar-ml/belief-world-models.
Chinese Translation
大语言模型(LLM)正被广泛用作众多领域中自主决策与规划的策略。尽管具备强大的推理能力,LLM在长时程任务中仍表现不佳,尤其是在部分可观测条件下。世界模型是提升策略性能的一种有前景的方法,无论在训练阶段还是推理阶段均是如此。在推理阶段,智能体目前利用世界模型在执行动作之前模拟候选动作的后果,从而改善决策。然而,我们认为,对于部分可观测条件下的决策而言,仅靠模拟是一种不完备的接口:模拟无法充分刻画关于当前状态的不确定性,而智能体可能需要这些信息才能做出准确决策。针对这一局限,我们提出了基于信念的世界模型(Belief-Based World Models,BB-WMs),该模型能够建模并维护一种信念,LLM可通过查询该信念获取关于当前状态哪些信息是已知的、哪些是不确定的。在开发学习精确BB-WMs的方法之前,我们首先提出一个更根本的问题:将世界模型的信念直接暴露给LLM策略能否改善决策?我们的结果表明,让LLM智能体访问世界模型的信念可以在部分可观测条件下提升任务性能,同时与现有的基于模拟的世界模型保持互补。代码已发布于 https://github.com/skumar-ml/belief-world-models。
cs.AI / 42 / 2609.00479

EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models

EGT-KG:面向小语言模型实用科学问答的证据支撑型类型化知识图谱检索
Yu, Muran, Gao, Jiechao, Pan, Yuandong, Miao, Barney H., Lesh, Andrew C., Law, Kincho H., Wang, Jie, Lepech, Michael D.
Abstract
For emerging scientific research domains, local Small Language Models (SLMs) are becoming more attractive, as they offer stronger privacy control and more stable deployment pipelines than Large Language Models. However, in practice, scientific question-answering on SLMs often operates under inevitable constraints: small literature collections, fragmented evidence, limited context window and reasoning abilities. We propose the Evidence-Grounded Typed Knowledge Graph (EGT-KG), a retrieval framework to improve information retrieval with local SLMs. We assessed three question-answering settings: a vanilla Retrieval-Augmented Generation (RAG) workflow and two EGT-KG workflows: an automatically generated relation schema (AS) and an expert-defined relation schema (ES). Our experiments were evaluated with a six-dimensional evaluation framework (S3CRF: Soundness, Correctness, Completeness, Conciseness, Relevance, Fluency) on a Biopolymer-bound Soil Composite literature benchmark, showing that EGT-KG outperforms the vanilla RAG method in most settings, with the best improvement from llama3:8b: a Final Score of 70.37 (+14.67%) and 68.82 (+12.14%) by AS/ES EGT-KG variants.
Chinese Translation
对于新兴科学研究领域,本地小语言模型(Small Language Models, SLMs)正变得更具吸引力,因为与大型语言模型相比,它们提供了更强的隐私控制和更稳定的部署流程。然而在实践中,基于SLMs的科学问答往往面临不可避免的约束:文献集合规模小、证据碎片化、上下文窗口和推理能力有限。我们提出了证据支撑型类型化知识图谱(Evidence-Grounded Typed Knowledge Graph, EGT-KG),这是一个用于改进本地SLMs信息检索的框架。我们评估了三种问答设置:一种是原始的检索增强生成(Retrieval-Augmented Generation, RAG)工作流,另外两种是EGT-KG工作流,即自动生成的关系统一体(Automatically generated relation Schema, AS)和专家定义的关系统一体(Expert-defined relation Schema, ES)。我们在一个生物聚合物结合土壤复合材料(Biopolymer-bound Soil Composite)文献基准上,采用六维评估框架(S3CRF:稳健性Soundness、正确性Correctness、完整性Completeness、简洁性Conciseness、相关性Relevance、流畅性Fluency)进行实验评估。结果表明,EGT-KG在大多数设置下优于原始RAG方法,其中llama3:8b模型取得了最佳改进:AS/ES两种EGT-KG变体的最终得分分别达到70.37(+14.67%)和68.82(+12.14%)。
cs.AI / 43 / 2609.00492

The Privacy-Hallucination Tradeoff in Differentially Private Language Models

差分隐私语言模型中的隐私-幻觉权衡
Ramesh, Krithika, Pillutla, Krishna, Pruthi, Danish, Field, Anjalie
Abstract
Both privacy and factual accuracy are paramount in high-stakes domains like healthcare. Concerningly, we uncover and investigate a privacy-hallucination tradeoff in differentially private (DP) language models. First, we empirically show that models pre-trained or fine-tuned with DP tend to produce more hallucinations than non-DP counterparts, with increased severity as the privacy budget grows stricter. Second, we investigate model properties driving this tradeoff, demonstrating that DP mechanisms flatten output distributions, potentially redistributing probability mass toward factually incorrect alternatives. Third, through experiments where we control fact frequency in training data, we characterize how information frequency can reduce hallucination risks in DP models. Overall, our findings underscore the need for more nuanced privacy-preserving interventions that offer rigorous privacy guarantees without compromising factual accuracy.
Chinese Translation
在医疗保健等高风险领域中,隐私和事实准确性都至关重要。令人担忧的是,我们发现并研究了差分隐私(DP)语言模型中的隐私-幻觉权衡。首先,我们通过实验表明,使用差分隐私进行预训练或微调的模型比非差分隐私模型更容易产生幻觉,且随着隐私预算变得越严格,幻觉的严重程度也随之增加。其次,我们研究了驱动这一权衡的模型特性,证明差分隐私机制会使输出分布趋于平坦,可能将概率质量重新分配给事实错误的选项。第三,通过控制训练数据中事实频率的实验,我们刻画了信息频率如何降低差分隐私模型中的幻觉风险。总体而言,我们的研究结果表明,需要更细致的隐私保护干预措施,在提供严格隐私保证的同时不损害事实准确性。
cs.AI / 44 / 2609.00498

Validity-Aware Jailbreak Evaluation for Large Language Models

面向大语言模型的有效性感知越狱评估
Wu, Qilong, Wadhwa, Sahil, Mohanty, Pranab, Iyengar, Giri, Chandrasekaran, Varun
Abstract
Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic plausibility rather than correctness. We identify a key limitation in existing evaluations: many jailbreak intents depend on instructional validity rather than epistemic factuality, allowing realistic-looking responses to be labeled successful despite being factually or procedurally incorrect. To address this gap, we propose Sequential Epistemic and Action-Level Validation (SEAV), a verification-centric jailbreak evaluation framework that decomposes responses into ordered steps and evaluates both validity and correctness. SEAV combines LLM-as-a-judge mechanisms for semantic interpretation with retrieval-grounded verification using external knowledge sources, assessing whether generated content is factually correct, structurally consistent, and operationally capable of advancing harmful objectives. Empirically, SEAV cuts the false-positive rate on SD-A (a curated strategic-dishonesty diagnostic) by 14.9\,pp vs. the strongest baseline, and reclassifies 22.1\%--51.0\% of sampled prior-labeled successes as invalid across three of four public benchmarks. Together, these results show that enforcing correctness substantially reshapes measured robustness: many previously labeled jailbreak successes are reclassified as invalid, and results are stable across the tested search backends and evaluator models. Code and data are available at https://github.com/Ardor-Wu/SEAV.
Chinese Translation
越狱鲁棒性已成为大语言模型(LLM)安全性评估的核心问题,然而现有主流方法主要依赖于拒绝行为、语义相似度以及意图匹配启发式规则,这些方法强调语言上的合理性而非正确性。我们发现现有评估方法存在一个关键局限:许多越狱意图依赖于指令的有效性而非认知上的真实性,这使得看似真实的回复即使事实上或操作上是错误的,也可能被标记为成功。为弥补这一不足,我们提出了序列化认知与行动层验证框架(Sequential Epistemic and Action-Level Validation, SEAV),这是一种以验证为中心的越狱评估框架,它将回复分解为有序步骤,并同时评估其有效性与正确性。SEAV 将用于语义解释的 LLM-as-a-judge 机制与基于外部知识源的检索验证相结合,评估生成内容在事实层面是否正确、结构层面是否一致,以及在操作层面是否能够推进有害目标。实验表明,与最强基线相比,SEAV 将 SD-A(一个精心构建的策略性欺骗诊断集)上的假阳性率降低了 14.9 个百分点;在四个公开基准中的三个上,SEAV 将先前被标记为成功的样本中的 22.1%–51.0% 重新分类为无效。这些结果共同表明,强制执行正确性会显著改变所测得的鲁棒性:许多先前被标记为越狱成功的样本被重新分类为无效,且结果在测试的各种搜索后端和评估模型上均保持稳定。代码和数据可在 https://github.com/Ardor-Wu/SEAV 获取。
cs.AI / 45 / 2609.00503

Wave Function Backpropagation with Explicit Temporal-Interval Dynamics

具有显式时间间隔动力学的波函数反向传播
Yu, Byunggu, Kim, Justin
Abstract
Conventional neural networks learn predominantly through affine transformations followed by nonlinear activations, while elapsed time is often treated as an auxiliary feature or assumed to be uniformly sampled. This paper introduces Wave Function Backpropagation (WFB), a wave-parameterized learning formulation in which neural responses are represented by learnable amplitude, wavenumber, angular frequency, and phase. The formulation associates an observed state with its temporal interval Delta t through the phase of a differentiable spatiotemporal wave. We derive standard WFB gradients and a spatial-curvature correction based on the Laplacian of the wave response. WFB is instantiated in a deliberately feed-forward trajectory predictor to provide a controlled proof of concept; sequence learning is outside the scope of the present evaluation. With motion features, STD-WFB using real intervals reduces average displacement error (ADE) by 20.4% relative to the original FFN baseline. In a new position-only evaluation that removes temporal leakage through precomputed velocity and acceleration, real-interval WFB reduces ADE by 10.4% relative to the original FFN and remains competitive with parameter-matched ReLU controls, obtaining 2.1% lower mean ADE than the matched FFN with explicit Delta t. Shuffled-interval WFB attains the lowest mean ADE, indicating that the present evidence supports the effectiveness of the wave representation but does not attribute the gain to interval alignment. These results establish WFB as a viable structured feed-forward learning formulation and define a clear basis for subsequent architectural studies.
Chinese Translation
传统神经网络的学习主要通过仿射变换加非线性激活实现,而流逝时间往往被当作辅助特征或被假设为均匀采样。本文提出波函数反向传播(Wave Function Backpropagation, WFB),这是一种以波参数化的学习形式,其中神经响应由可学习的振幅、波数、角频率和相位表示。该形式通过可微时空波的相位将观测状态与其时间间隔 Delta t 相关联。我们推导了标准的 WFB 梯度以及基于波响应拉普拉斯算子的空间曲率修正。WFB 被实例化于一个特意设计的纯前馈轨迹预测器中,以提供受控的概念验证;序列学习不在本评估范围之内。在运动特征条件下,使用真实时间间隔的 STD-WFB 相较于原始 FFN 基线将平均位移误差(ADE)降低了 20.4%。在一项新的仅位置评估中(通过预计算的速度和加速度消除时间信息泄露),真实间隔的 WFB 相较于原始 FFN 将 ADE 降低了 10.4%,并与参数量匹配的 ReLU 对照组保持竞争力,其平均 ADE 比带有显式 Delta t 的匹配 FFN 低 2.1%。打乱间隔的 WFB 取得了最低的平均 ADE,这表明现有证据支持波表示的有效性,但并不能将这一增益归因于间隔对齐。这些结果确立了 WFB 作为一种可行的结构化前馈学习形式的地位,并为后续的架构研究奠定了明确的基础。
cs.AI / 46 / 2609.00508

CoVer: Conflict-Aware Claim Verification

CoVer:冲突感知的声明验证
Zhang, Shuning, Shi, Dai, Chu, Bohao, Wang, Hui, Chuai, Yuwei, Wang, Yifan, Chen, Jingruo, Li, Simin, Yi, Xin, Li, Hewu
Abstract
Social media fact-checking has long been challenged by evidence-level and aggregation-level conflicts, where erroneous evidence mimics authoritative news sources. To capture this challenge and support conflict verification tasks, we present ContraNote, a large-scale real-world dataset curated from X's Community Notes system. It includes 33,686 posts for evaluating evidence-level conflict resolution, and 54,474 instances for evaluating aggregation-level prioritization. Additionally, we propose CoVer, a factual adjudication framework with three-stage pipelines: evidence schema normalization, factual consensus and support verification. This prioritizes evidence over noise to prevent it from compromising the final verdict. Technical evaluations show that CoVer achieves strong performance compared with state-of-the-art baselines across ContraNote (86.0% Acc., 68.0% mac. F1, 64.5 bal. Acc. on Conflict; and 88.5% Acc., 88.5 mac. F1 and 89.2 bal. Acc. on Prioritization), CONFACT-HumC (88.4% Acc.) and CONFACT-ModC (89.4% Acc.).
Chinese Translation
社交媒体事实核查长期以来一直受到证据层面冲突和聚合层面冲突的挑战,即错误证据会模仿权威新闻来源。为了刻画这一挑战并支持冲突验证任务,我们提出了ContraNote,这是一个基于X平台社区笔记(Community Notes)系统精心构建的大规模真实世界数据集。该数据集包含33,686条帖子,用于评估证据层面的冲突解决,以及54,474个实例,用于评估聚合层面的优先级排序。此外,我们提出了CoVer,一个包含三阶段流水线的事实裁决框架:证据模式规范化、事实共识与支持性验证。该框架优先考虑证据而非噪声,防止噪声损害最终判定。技术评估表明,与最先进的基线方法相比,CoVer在ContraNote(Conflict任务上:准确率86.0%、宏平均F1为68.0、平衡准确率64.5;Prioritization任务上:准确率88.5%、宏平均F1为88.5、平衡准确率89.2)、CONFACT-HumC(准确率88.4%)和CONFACT-ModC(准确率89.4%)上均取得了优异性能。
cs.AI / 47 / 2609.00510

When the Algorithm Becomes the Brand Crisis: A Sociotechnical Theory of Distributed Responsibility and Accountable Transparency

当算法成为品牌危机:一种关于分布式责任与可问责透明度的社会技术理论
Torkestani, Mohammad Saleh, Mansouri, Taha
Abstract
Artificial intelligence systems increasingly enact market-facing promises through chatbots, recommendation systems, automated decisions, and generative interfaces. Their failures, misuse, and misrepresentation raise a question that conventional brand-crisis models do not fully specify: how do stakeholders assign responsibility when technical causation, customer-facing control, and governance duties are distributed across an AI system, developer, deployer, vendor, and user? This conceptual paper develops a sociotechnical process theory from a structured, federated scoping synthesis of verified academic and primary sources. It distinguishes an AI/algorithmic incident from an AI-related organisational crisis and, in turn, from an AI-related organisational scandal. The framework proposes that incident configuration shapes actor-specific attribution; attribution informs capability, integrity, fairness, and relationship appraisals; and public moralisation may, but need not, escalate an incident into scandal. The theory offers a reconciliation of findings that algorithm involvement can buffer negative brand reactions in some settings while robot and chatbot failures can redirect responsibility to an associated firm in others. It introduces accountable transparency as a proposed response configuration that combines timely notice, an intelligible account, role-responsibility acknowledgement, remedy, evidence of correction, and recourse. The evidence supports conditional, proximal inferences about blame, trust, satisfaction, firm evaluation, and communication credibility more strongly than claims about durable reputation, brand equity, or market performance.
Chinese Translation
人工智能系统日益通过聊天机器人、推荐系统、自动化决策和生成式界面来兑现面向市场的承诺。其失效、滥用与虚假陈述引发了一个传统品牌危机模型未能充分阐明的问题:当技术因果、面向客户的控制以及治理责任分布于AI系统、开发者、部署者、供应商和用户之间时,利益相关者如何分配责任?本文基于对经过验证的学术文献和一手资料的结构化、联合范围综述,发展出一种社会技术过程理论。该理论区分了AI/算法事件(incident)、与AI相关的组织危机(crisis),以及与AI相关的组织丑闻(scandal)。该框架提出:事件的构型塑造了针对特定行为者的责任归因;归因进而影响对能力、诚信、公平性以及关系的评价;公众的道德化可能(但不必然)将事件升级为丑闻。该理论调和了以下看似矛盾的发现:在某些情境中,算法参与能够缓冲负面的品牌反应,而在另一些情境中,机器人和聊天机器人的失效会将责任转嫁给相关企业。本文提出了"可问责透明度"(accountable transparency)这一应对构型,它结合了及时告知、可理解的说明、角色责任的承认、补救措施、整改证据以及申诉渠道。现有证据对有关责备、信任、满意度、企业评价和传播可信度的近端条件性推断的支持,强于对持久声誉、品牌资产或市场绩效等主张的支持。
cs.AI / 48 / 2609.00513

ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation

ISO-RAG:面向检索增强生成的等周噪声控制方法
Zhang, Siyuan, Wang, Hanchen, Wen, Dong, Zhang, Ying, Zhang, Wenjie
Abstract
Retrieval-Augmented Generation (RAG) mitigates large language models (LLMs) hallucinations, yet conventional dense retrieval struggles with the complex reasoning paths of multi-hop question answering (QA). Graph-based RAG captures multi-step relationships but suffers from severe semantic drift and high online latency due to noisy global graph traversals. Thus, we propose ISO-RAG (ISOperimetric Retrieval-Augmented Generation), a geometry-aware RAG framework. By projecting the underlying knowledge graph into a hyperbolic Poincare ball to precompute node-wise isoperimetric profiles, ISO-RAG prunes spurious edges during retrieval, restricting the search space to a strictly localized subgraph. This topological purification regulates Personalized PageRank (PPR) diffusion driving the retrieval process, ensuring exact and low-latency convergence. Experiments on multi-hop QA benchmarks demonstrate that ISO-RAG outperforms state-of-the-art baselines by average absolute gains of 10.0% in retrieval recall and 4.3% in downstream exact match, achieving a superior accuracy-efficiency trade-off by fundamentally eliminating the latency bottleneck of global traversals. Our source code is available at https://github.com/ZaiizaiZHANG/ISO-RAG.
Chinese Translation
检索增强生成(RAG)能够缓解大语言模型(LLM)的幻觉问题,但传统的稠密检索难以应对多跳问答(QA)中复杂的推理路径。基于图的RAG能够捕获多步关系,但由于全局图遍历中存在噪声,导致严重的语义漂移和较高的在线延迟。为此,我们提出了ISO-RAG(ISOperimetric Retrieval-Augmented Generation),一种几何感知的RAG框架。通过将底层知识图谱投影到双曲庞加莱球中以预先计算节点级的等周轮廓,ISO-RAG在检索过程中剪除虚假边,将搜索空间限制在严格局部化的子图中。这种拓扑净化规范了驱动检索过程的个性化PageRank(PPR)扩散,确保精确且低延迟的收敛。在多跳问答基准上的实验表明,ISO-RAG优于最先进的基线方法,在检索召回率和下游精确匹配率上分别取得了平均10.0%和4.3%的绝对提升,通过从根本上消除全局遍历的延迟瓶颈,实现了卓越的精度-效率权衡。我们的源代码可在 https://github.com/ZaiizaiZHANG/ISO-RAG 获取。
cs.AI / 49 / 2609.00543

Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation

面向检索增强生成的基于反馈辅助的文档关系图信任传播方法
Li, Zhuoheng, Chen, Ying
Abstract
Retrieval-augmented generation (RAG) systems rely on external corpora that may contain outdated, contradictory, noisy, or unreliable documents, introducing reliability risks. Prior work has leveraged document relations to improve the answer reliability of RAG. To propagate reliability signals beyond directly compared document pairs, we propose TrustPropRAG, which structures document relations as a graph and estimates document reliability through multi-hop propagation across the graph. TrustPropRAG anchors this propagation with a limited set of human feedback on document reliability, extending these costly-to-collect feedback-based reliability signals across the whole corpus. Specifically, based on the constructed document relation graph, TrustPropRAG estimates a trust score for each document by formulating and solving an optimization problem that jointly captures pairwise document relations and user feedback. These scores are then used to improve the selection of reliable documents and support trust-aware answer generation. Evaluation results show that TrustPropRAG improves both retrieval quality and exact match over baselines, and remains robust under sparse and noisy feedback.
Chinese Translation
检索增强生成(RAG)系统依赖的外部语料库可能包含过时、矛盾、嘈杂或不可靠的文档,从而引入可靠性风险。已有研究利用文档关系来提升RAG的答案可靠性。为了将可靠性信号传播到直接比较的文档对之外,我们提出了TrustPropRAG,该方法将文档关系组织为图结构,并通过图上的多跳传播来估计文档可靠性。TrustPropRAG以有限的人工文档可靠性反馈作为传播的锚点,将这些收集成本高昂的基于反馈的可靠性信号扩展到整个语料库。具体而言,基于构建的文档关系图,TrustPropRAG通过构建并求解一个同时刻画文档间成对关系与用户反馈的优化问题,为每个文档估计信任分数。随后,这些分数被用于改进可靠文档的选择,并支持信任感知的答案生成。评估结果表明,与基线方法相比,TrustPropRAG在检索质量和精确匹配方面均有所提升,并且在稀疏和嘈杂的反馈条件下依然保持稳健。
cs.AI / 50 / 2609.00570

VoiceLongMemEval: Do Assistants Remember How You Sounded?

VoiceLongMemEval:助手能记住你的语气吗?
Pahwa, Ramit, Priye, Parivesh, Beedu, Apoorva
Abstract
With the growing scale of multi-agent architectures and large language models, deployed AI assistants are increasingly tasked with reasoning over long, continuous, multi-session conversation histories. Current benchmarks evaluate this dialogue history as information retrieval over long horizon, temporal reasoning, or knowledge updates, while crucially ignoring the fundamental dynamics of human-agent interaction, i.e. how they said it. To address this gap, we present VoiceLongMemEval (VLME) benchmark, where every answer depends on paralinguistic metadata (emotion labels, prosody descriptors, and voice events) attached to conversational turns, which is otherwise unrecoverable from the words alone. Every item passes a three-stage adversarial gate, ensuring that a strong language model fails when given only the transcript. Evaluating leading frontier and open-weight models reveals a pervasive affect gap; providing text-track paralinguistic metadata yields a 0.09 to 0.38 accuracy boost (0.61 to 0.69 when prompted with evidence hints), while standard ASR pipelines systematically discard this signal. Additionally, audio-native models successfully extract these cues directly from speech (0.354 to 0.412 vs. 0.325 blind). Code and dataset will be made available upon acceptance.
Chinese Translation
随着多智能体架构和大语言模型规模的不断增长,部署的AI助手越来越多地需要对长篇、连续的多轮会话历史进行推理。现有基准通常将对话历史评估视为长程信息检索、时间推理或知识更新,却忽略了人机交互中最根本的动态特征,即说话的方式。为弥补这一空白,我们提出了VoiceLongMemEval(VLME)基准,其中每个问题的答案都依赖于附加在对话轮次上的副语言元数据(情感标签、韵律描述和语音事件),这些信息无法仅从文字内容中恢复。每个测试项都经过三阶段对抗性筛选,确保强大的语言模型在仅看到文本转录时无法作答。对领先的前沿模型和开源权重模型的评估揭示了普遍存在的情感鸿沟:提供文本轨道的副语言元数据可带来0.09至0.38的准确率提升(结合证据提示时可从0.61提升至0.69),而标准的ASR流程会系统性地丢弃这一信号。此外,音频原生模型能够直接从语音中成功提取这些线索(0.354至0.412,对比无音频输入的0.325)。代码和数据集将在论文录用后公开发布。
cs.AI / 51 / 2609.00575

Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs

基于输出重要性的残差稀疏化方法用于压缩混合专家大语言模型
Jung, Seungwoo, Kwon, Dohyeok, Cha, Seungmin, Lee, Junseok, Yoo, Yeonho, Yoo, Chuck, Yang, Gyeongsik
Abstract
Mixture-of-experts (MoE) architectures scale large language models efficiently, but they demand massive GPU memory. To cope with such demand, models are commonly compressed to reduce their memory footprint. Residual sparsification is a representative compression technique that decomposes each projection matrix of an expert into a shared base matrix and per-expert residual matrix, and then compresses the residuals. Existing sparsification methods compress each residual matrix independently by minimizing its compression error, thereby minimizing the error of each projection matrix. However, our analysis shows that this objective is misaligned with preserving model accuracy after compression. In an expert, the final output is produced through computations coupled across multiple projections and hidden representations. Therefore, even small errors in individual matrices can propagate through hidden representations and projection interactions, leading to large expert output errors and accuracy degradation. To address this misalignment, we propose PARSER, a new residual sparsification method that shifts the compression objective from minimizing isolated matrix errors to preserving the expert output error. PARSER achieves this by introducing output importance, which measures the actual contribution to the expert output error. Our experiments show that, compared with existing methods, PARSER narrows the accuracy gap to the uncompressed model by 1.41$\times$ on Qwen and 1.44$\times$ on DeepSeek, while achieving the same peak memory reduction. Our code is available at https://github.com/OSSS-KU/PARSER.
Chinese Translation
混合专家(Mixture-of-Experts, MoE)架构能够高效地扩展大语言模型的规模,但其对GPU显存的需求极为庞大。为应对这一需求,通常需要对模型进行压缩以减小其内存占用。残差稀疏化是一种代表性的压缩技术,它将专家的每个投影矩阵分解为一个共享的基础矩阵和每个专家独有的残差矩阵,然后对残差进行压缩。现有的稀疏化方法通过最小化压缩误差来独立地压缩每个残差矩阵,从而最小化每个投影矩阵的误差。然而,我们的分析表明,这一目标与压缩后保持模型精度并不一致。在专家内部,最终输出是通过跨越多个投影和隐表示的耦合计算得到的。因此,即使单个矩阵中存在微小误差,也可能通过隐表示和投影间的相互作用进行传播,导致较大的专家输出误差和精度下降。为解决这种不一致性,我们提出了PARSER,一种新的残差稀疏化方法,它将压缩目标从最小化孤立的矩阵误差转变为保持专家输出误差。PARSER通过引入输出重要性来实现这一目标,该指标衡量了对专家输出误差的实际贡献。实验表明,与现有方法相比,PARSER在Qwen和DeepSeek上分别将以1.41倍和1.44倍缩小与未压缩模型的精度差距,同时实现相同的峰值内存降低。我们的代码已发布于 https://github.com/OSSS-KU/PARSER。
cs.AI / 52 / 2609.00576

Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random

无对齐的一致性:与随机选择难以区分的项目敏感语言模型
Huynh, Cris
Abstract
Item-sensitivity, defined as whether a model's choice depends on the specific input rather than on its own output prior, is widely reported as evidence of task competence. We show this evidence is necessary but not sufficient using a forced-choice signalling task abstracted from the board game Deception: Murder in Hong Kong. In this environment, the reference points against which a coordinate should be judged (a fit-maximising strategy, a posterior-maximising strategy, and uniform random selection) are all computable in closed form. Across seven language models, two model families, a post-training ablation, and three independent scoring rules, every one of 21 model-by-rule cells is reliably item-sensitive. Yet 8 of those 21 cells are not statistically distinguishable from a chooser that ignores the item and selects at random, and 5 score worse than random at describing the target. Item-sensitivity and distance from random correlate at only r = 0.30. We call this consistency without alignment and argue it generalises to any evaluation that relies on item-sensitivity, permutation consistency, or self-consistency without an independent reference for the measured quantity. We further find that a literal-similarity baseline with no pragmatics outperforms most tested language models, that adding a pragmatic layer over two baseline similarity sources moves choosers toward random rather than toward the Bayesian reference, and that a standard labelled multiple-choice format carries no measurable content signal here. All results represent the model side of a pre-registered instrument; a matched human condition is designed and piloted but not yet collected.
Chinese Translation
项目敏感性(item-sensitivity)指模型的选择是否取决于具体输入而非其自身的输出先验,它被广泛报告为模型任务能力的证据。我们通过一个从桌游《Deception: Murder in Hong Kong》中抽象出来的强迫选择信号传递任务表明,这一证据是必要但不充分的。在该环境中,用于评判协作选择的三种参照点——拟合最大化策略、后验最大化策略和均匀随机选择——均可闭式计算。在七个语言模型、两个模型家族、一项后训练消融实验以及三种独立评分规则下,全部21个“模型×规则”组合均表现出可靠的项目敏感性。然而,其中8个组合在统计上与忽略输入内容、随机选择的决策者无法区分,另有5个组合在描述目标时表现低于随机水平。项目敏感性与距随机基线的距离之间的相关性仅为r=0.30。我们将这种现象称为“无对齐的一致性”(consistency without alignment),并论证其可推广至任何依赖项目敏感性、排列一致性或自一致性但缺乏对被测量的量的独立参照的评估方法。我们进一步发现:一个不含语用机制的表面相似性基线表现优于大多数被测语言模型;在两种基线相似性来源之上添加语用层会使选择者更趋近随机而非贝叶斯参照;且标准的带标签多项选择格式在此环境中不携带可测量的内容信号。所有结果均代表一项预注册工具的模型侧数据;匹配的人类实验条件已完成设计与试测,但数据尚未收集。
cs.AI / 53 / 2609.00578

Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts

相同请求,不同边界:跨对话语境评估网络安全辅助能力
Yang, Rui, Hong, Yang, Xu, Yichao, Liu, Zhengyu, Li, Ziyang, Cao, Yinzhi
Abstract
Large Language Models (LLMs) can solve complex problems, but their misuse in high-risk domains can lead to severe consequences. Model providers therefore restrict assistance for potentially harmful requests. Refusing all cybersecurity requests would therefore harm legitimate users. Providers need a mechanism to block malicious use without denying legitimate assistance to defenders. Existing cybersecurity-specific datasets evaluate this mechanism, but none considers the conversational context of a request. We introduce 3R-Bench (Refusal, Repetition, and Revision), a benchmark of 150 real-world cybersecurity requests augmented with two adversarial conversational settings, and evaluate eight LLMs on it. Prior assistant behavior strongly changes responses to an unchanged request: among 376 available pairs from a 400-pair panel, compliance rises from 62.0% after refused history to 85.1% after accepted history. The opposite pattern appears under dialogue decomposition. In comparison, compliance falls from 501/800 direct responses to 172/800 after dialogue; among 738 pairs returning model-authored text in both conditions, the decrease is 45.1 points. Failure feedback recovers only a small fraction of this loss.
Chinese Translation
大语言模型(LLM)能够解决复杂问题,但在高风险领域的滥用可能导致严重后果。因此,模型提供方限制对潜在有害请求的辅助。然而,拒绝所有网络安全请求会损害合法用户的利益。提供方需要一种机制来阻止恶意使用,同时不拒绝向防御者提供合法的辅助。现有的网络安全专用数据集对这种机制进行了评估,但没有考虑请求所处的对话语境。我们提出了 3R-Bench(拒绝、重复与修订),一个包含150个真实世界网络安全请求并增加了两种对抗性对话设置的基准,并在其上评估了八个大语言模型。先前的助手行为会强烈改变模型对同一不变请求的响应:在400对中的376个可用样本对中,在被拒绝的历史之后合规率为62.0%,而在被接受的历史之后合规率上升至85.1%。在对话分解(dialogue decomposition)设置下则出现相反的模式。相比之下,合规率从直接响应的501/800下降到对话后的172/800;在738个在两种条件下均返回模型生成文本的样本对中,降幅达45.1个百分点。而失败反馈仅能挽回这一损失的一小部分。
cs.AI / 54 / 2609.00584

Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing

苏格拉底走向核能:基于脑信号的学习情境中人工智能系统交互策略比较
Deffarges, Alexandre Clin, Kosmyna, Nataliya, Maes, Pattie
Abstract
Does unrestricted AI access bypass the cognitive effort required for learning, or does it streamline knowledge acquisition? This paper reports on a study where we compare three designs for user-AI interaction in a learning context: (1) an unrestricted conversational bot like ChatGPT, (2) a pedagogically constrained bot that guides through hints without giving final answers, which we refer to as the Socratic mode; and (3) a non-conversational adaptive tutoring system that adjusts difficulty in real-time based on the user's cognitive engagement derived from the brain signals. Fifty study participants were tasked with learning about nuclear safety protocols, a domain chosen for its zero-prior knowledge baseline. The participants progressed through an instructional video, a pre-test, an AI-driven assessment phase, which varied in the three conditions, and an immediate post-test. The nature of the questions centered primarily on factual knowledge acquisition, but it still required participants to have a global understanding of the concepts in order to answer the questions correctly. A Muse headband was used to derive the cognitive engagement of all users in all conditions. The unrestricted chatbot produced higher learning gains (delta) than both constrained modes (p < .03, d > 0.80), while the adaptive condition generated significantly higher EEG engagement (p = .018). The cluster analysis of chatbot usage and discussion patterns by users showed that most participants in the unrestricted-mode adopted a direct answer-retrieval strategy, while participants in the Socratic-mode initially attempted to reason through the hints before progressively disengaging. Consequently, this also suggests that the success of the unrestricted AI is not an evidence of deeper learning, but rather a result of the immediate post-test evaluation after the training phase.
Chinese Translation
无限制的人工智能访问是否会绕过学习所需的认知努力,还是能简化知识获取过程?本文报告了一项研究,我们比较了学习情境中三种人机交互设计:(1) 类似ChatGPT的无限制对话机器人;(2) 采用教学约束的机器人,通过提示进行引导而不给出最终答案,我们称之为苏格拉底模式;(3) 非对话式自适应辅导系统,根据从脑信号中提取的用户认知投入程度实时调整难度。五十名研究参与者的任务是学习核安全协议,选择该领域的原因是其零先验知识的基线。参与者依次完成一段教学视频、一次前测、一个AI驱动的评估阶段(三种条件下各不相同)以及一次即时后测。问题主要围绕事实性知识的获取,但仍要求参与者对概念有整体理解才能正确回答。研究中使用Muse头戴设备来测量所有条件下所有用户的认知投入程度。无限制聊天机器人产生了比两种约束模式更高的学习收益(delta)(p < .03,d > 0.80),而自适应条件则产生了显著更高的脑电(EEG)投入度(p = .018)。对用户聊天机器人使用和讨论模式的聚类分析显示,无限制模式下的大多数参与者采用了直接检索答案的策略,而苏格拉底模式下的参与者起初尝试根据提示进行推理,随后逐渐脱离。因此,这也表明无限制人工智能的成功并非深度学习的证据,而只是训练阶段后进行即时后测评估的结果。
cs.AI / 55 / 2609.00621

Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs

控制-数据流分离:多智能体大语言模型中的稳定提示优化
Zhang, Wentao, Murtaza, Syed Shariyar, Bhatti, Junaid Ahmad, Soni, Utkarsh, Nie, Yifan, Wen, Eugene, Deng, Yuntian
Abstract
Prompt optimization can improve multi-agent LLM systems, but the prompts being optimized often serve two entangled roles: generating task-relevant content and specifying execution-critical protocols, such as message routing, output formatting, and termination signals, on which the underlying code relies. As a result, a prompt edit intended to improve content generation can inadvertently corrupt the protocol and cause the entire agent pipeline to fail. Our key observation is that these two roles have different representations: execution protocols are typically structured, while task-relevant content is usually expressed in unstructured language. Based on this, we propose control-data flow separation, where execution-critical control is represented as typed, validated program objects, while task-relevant language remains the optimizable data flow for agent communication. This design allows optimizers to improve multi-agent behavior without exposing the routing or formatting interface to prompt drift. Across synthetic reasoning, collaborative review generation, and insurance rating workflows, our framework empirically achieves 100% eventual protocol validity while consistently improving task performance.
Chinese Translation
提示优化能够改进多智能体大语言模型系统,但被优化的提示词往往承担两种相互纠缠的角色:一是生成与任务相关的内容,二是指定执行关键协议(如消息路由、输出格式和终止信号),而底层代码依赖于这些协议。因此,旨在改进内容生成的提示词修改可能无意中破坏协议,导致整个智能体流水线失效。我们的关键观察是:这两种角色具有不同的表示形式——执行协议通常是结构化的,而任务相关内容通常以非结构化语言表达。基于此,我们提出控制-数据流分离方法,将执行关键的控制表示为带类型、经验证的程序对象,而与任务相关的语言仍作为可优化的数据流用于智能体间通信。这种设计使优化器能够改进多智能体行为,而不会使路由或格式化接口暴露于提示词漂移。在合成推理、协作评审生成和保险定价工作流上,我们的框架在实验中实现了100%的最终协议有效性,同时持续提升任务性能。
cs.AI / 56 / 2609.00643

REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows

REVISE:面向智能体工作流在线修订的有效性引导恢复方法
Qi, Ruoling, Wu, Xuaner, Liu, Penghang, Chen, Jian, Liu, Yirui
Abstract
Agent revisions expose a fundamental correctness--efficiency trade-off during concurrent execution. Discarding ongoing work preserves latest-version correctness but wastes progress that may remain valid, whereas reusing prior work preserves efficiency but risks propagating stale state into outputs and tool effects. Existing recovery strategies resolve this trade-off in an imbalanced way with coarse-grained policies: they either favor efficiency by allowing potentially stale work to continue, or favor correctness by restarting the workflow or recomputing a linear suffix from the earliest conflict, thereby discarding unaffected progress. We present \textsc{Revise}, a validity-guided runtime for fine-grained recovery in structured agent workflows. When a revision arrives, \textsc{Revise} first intersects its delta with recorded data and control dependencies and propagates the resulting impact through the partially executed DAG to identify affected work. It then stops invalid work, preserves validity-established progress beyond the earliest conflict, and recomputes only the affected region. Incomplete provenance conservatively expands recovery, while reused results are revalidated before commit. Analysis of real coding-agent traces show online recovery opportunities: 118 sessions retain observable work before a queued later message is delivered; across 167 overlapping assistant responses, enqueue-to-completion overlap reaches 56.55~s at p95. Across 300 challenging revision/commit executions, \textsc{Revise} matches a latest-version oracle with no stale outputs or effects. On unmodified LangGraph and LLMCompiler applications using Qwen3-14B, it reduces model calls by 40.6--56.0\% relative to full restart and by 31.3--43.6\% relative to suffix recomputation. Under serving pressure, it further reduces revision-to-correct-completion tokens by 13.26\% and improves SLO goodput by 3.07--5.43\%.
Chinese Translation
智能体修订(agent revisions)在并发执行过程中暴露出正确性与效率之间的根本性权衡。丢弃正在进行的工作可以保证最新版本的正确性,但会浪费可能仍然有效的进度;而复用先前的工作虽然保持了效率,却存在将过时状态传播到输出和工具副作用中的风险。现有的恢复策略以粗粒度的方式失衡地解决这一权衡:它们要么偏向效率,允许可能过时的工作继续执行;要么偏向正确性,重启整个工作流或从最早的冲突点重新计算线性后缀,从而丢弃未受影响的进度。我们提出了 REVISE,一个面向结构化智能体工作流细粒度恢复的有效性引导运行时系统。当修订到达时,REVISE 首先将其增量与记录的数据和控制依赖进行求交,并将产生的影响在部分执行的 DAG 中传播,以识别受影响的工作。随后,它停止无效的工作,保留在最早冲突点之后已被验证有效的进度,并仅重新计算受影响的区域。当溯源信息不完整时,恢复范围会被保守地扩大,而被复用的结果在提交前会经过重新验证。对真实编程智能体(coding-agent)轨迹的分析显示了在线恢复的机会:118 个会话在队列中后续消息送达之前仍保留着可观察的工作;在 167 个重叠的助手响应中,从入队到完成的重叠时间在 p95 处达到 56.55 秒。在 300 个具有挑战性的修订/提交执行中,REVISE 的表现与最新版本预言机(oracle)相当,且没有产生过时的输出或副作用。在未修改的 LangGraph 和 LLMCompiler 应用上使用 Qwen3-14B 时,相较于完全重启,它将模型调用次数减少了 40.6%–56.0%;相较于后缀重计算,减少了 31.3%–43.6%。在服务压力下,它进一步将修订到正确完成的 token 数减少了 13.26%,并将 SLO 有效吞吐量(goodput)提升了 3.07%–5.43%。
cs.AI / 57 / 2609.00646

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation

DramaChain Bench:一个面向短剧生成的端到端基准测试
Shi, Haoyuan, Chen, Mingtao, Jiang, Shuo, Chen, Ziyan, Sheng, Xuyi, Liu, Yiming, Zhang, Ying, Wang, Miao, Lu, Jianxiang, Lu, Fanyang, Lu, Songyuanyi, Wu, Xiele, Hu, Zhichao, Liu, Yuhong, Xuan, Richeng
Abstract
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.
Chinese Translation
商业短剧制作遵循一条多阶段流水线:剧本、分镜、关键帧图像、镜头级视频以及最终成片的短剧。现有的大多数基准测试仅使用预先撰写的输入(而非真实的上游流水线输出)来评估视频生成阶段。这使得两个关键问题无法得到解答:各阶段是否忠实于原始剧本意图(而不仅仅是紧邻的上游输入提示),以及不同镜头在拼接为多剧集发布后是否保持连贯。我们提出了DramaChain Bench,这是首个对完整制作链的每一个阶段进行评估的短剧基准。它建立在共享同一维度体系(DramaChain Dimensions)的三个自研系统之上:五个评估轴在每个阶段实例化,并细化为63个叶子维度。DramaChain Agent在工作流程和成片短剧质量两个层面均与商业短剧平台进行了校准,从而实现对不同模型在各阶段进行公平比较。DramaChain标注系统(DramaChain Labeling System)中的5,785个条目均由三名专业标注员独立打分,所有缺陷均在时空上定位,并从预定义的缺陷列表中选取。该过程产出了17,488个有效评分和255,925条可追溯的归因记录。人工标注证实,上游缺陷会在流水线中级联传播,表明最终剧集的质量并非仅由视频生成决定。DramaChain Agentic Judge随后对每个叶子维度进行自动打分,在多个智能体轮次中收集证据,并依据逐条目检查清单进行评判;它以平均PLCC为0.918的精度复现了模型排名,足以在零标注成本的情况下纳入新模型。
cs.AI / 58 / 2609.00652

Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search

自我报告并非验证:面向进化搜索中大语言模型操作者的基于环境的审计
Pan, Enrong, Zhou, Ryan, Hu, Ting
Abstract
Language model agents increasingly propose actions, observe external feedback, and explain their own behavior. Their confidence and rationales are convenient monitoring signals, but convenience is not verification. We introduce an environment-grounded audit in which every intermediate proposal receives an exact outcome. A language model operates an evolutionary Contexto search whose feedback function assigns every valid guess an exact rank without human annotation. Across 200 runs spanning five configurations and three model families, four reporting configurations produce 12,249 self-reports. We test three assumptions: stated confidence is calibrated, inherited rationales affect later proposals, and fitness-based selection improves report quality. All three fail. Operators overstate top-100 success by factors of 4.8 to 9.3, while calibration and discrimination dissociate across model families. Controlled interventions on 754 inherited rationales bound any measured benefit of the genuine rationale to roughly 250 ranks. Neither fitness-based nor random selection produces a detectable selection differential or parent-to-offspring transmission in report accuracy, despite sharply different search behavior. Agent self-reports should therefore be treated as claims to verify against the environment, not as evidence of their own reliability.
Chinese Translation
大语言模型智能体越来越多地提出行动、观察外部反馈并解释自身的行为。它们的置信度与理由是便利的监控信号,但便利并非验证。我们提出了一种基于环境的审计方法,其中每个中间提议都会获得精确的结果反馈。一个语言模型操作进化式 Contexto 搜索,其反馈函数无需人工标注即可为每个有效猜测赋予精确排名。在涵盖五种配置和三个模型系列的 200 次运行中,四种报告配置共产生了 12,249 份自我报告。我们检验了三个假设:所声明的置信度经过校准、继承的理由会影响后续提议、基于适应度的选择能提升报告质量。三个假设均不成立。操作者对进入前 100 名的成功率高估了 4.8 至 9.3 倍,且校准与区分能力在不同模型系列之间出现分离。对 754 个继承理由的受控干预表明,真实理由所能带来的可测收益上界约为 250 个排名。尽管搜索行为差异显著,无论是基于适应度还是随机选择,都未在报告准确性上产生可检测的选择差异或亲代到子代的传递效应。因此,智能体的自我报告应被视为需要对照环境加以验证的声明,而非其自身可靠性的证据。
cs.AI / 59 / 2609.00654

SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task

SciTrue:基于前沿与开源语言模型的可靠科学论断验证——NTCIR SciClaimEval 任务
Bao, Qiming, Tan, Neşet Özkan, Wang, Siyuan, Gahegan, Mark
Abstract
We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task~\cite{sciclaimeval}, which asks systems to verify scientific claims against the tables and figures of a paper. Rather than tuning a single model, we benchmark eleven frontier and open multimodal models under one honest, per-sample protocol and combine them with light, transparent post-processing. On the official, blind test leaderboard (Section~\ref{sec:results}), SciTrue placed first by a clear margin in three of the four evidence-category/subtask combinations, and tied for first on the primary metric in the fourth. Three findings explain the result. First, strong instruction-tuned models are already competitive: Claude Opus~4.8 and Gemma-4-31B each exceed the strongest public baseline (o4-mini), and GPT-5.5 and Claude Fable~5 lead both subtasks (97.7 on Subtask~2). Second, the task's pairing structure is the largest lever: a \emph{leak-free pair prior} that recovers the Supported/Refuted pairing from the claim text alone (a visible field) and assigns Supported to the higher-confidence evidence raises Subtask-1 pair-accuracy from 72.2 to 93.5, far more than any model swap or ensemble weighting. Third, a case-by-case audit finds that most residual errors are visually-undetectable label-mapping swaps or dataset label noise, so measured accuracy understates the true ability and the fixable-by-modeling headroom is small. Controlled fine-tuning, distillation, and agentic consistency-checking support the same conclusions, and we document throughout a measurement leak---label information reaching a system through the packaging of the data rather than its content---in which the released file ordering encodes the label, including one instance that briefly misled our own pipeline.
Chinese Translation
本文介绍了 SciTrue 团队参加 NTCIR-19 SciClaimEval 任务的子任务 A 和子任务 B 的情况,该任务要求系统根据论文中的表格和图示来验证科学论断。我们没有针对单一模型进行调优,而是在统一且诚实的逐样本评测协议下对十一个前沿及开源多模态模型进行基准测试,并将其与轻量、透明的后处理方法相结合。在官方盲测排行榜(第~\ref{sec:results}~节)上,SciTrue 在四个证据类别/子任务组合中的三个以明显优势获得第一,在第四个组合的主要指标上并列第一。三点发现解释了这一结果。第一,强大的指令微调模型本身已具备竞争力:Claude Opus~4.8 和 Gemma-4-31B 均超越了最强的公开基线(o4-mini),而 GPT-5.5 和 Claude Fable~5 则在两个子任务中均处于领先地位(子任务 2 达 97.7)。第二,任务的配对结构是最大的杠杆:一种无泄漏的配对先验(leak-free pair prior),仅凭论断文本(一个可见字段)恢复 Supported/Refuted 的配对关系,并将 Supported 分配给置信度更高的证据,即可将子任务 1 的配对准确率从 72.2 提升至 93.5,其效果远超任何模型更换或集成加权。第三,逐例审计发现,大多数残余错误是视觉上无法检测的标签映射错位或数据集标签噪声,因此所测得的准确率低估了模型的真实能力,可通过建模手段修复的提升空间很小。受控微调、蒸馏和基于智能体的一致性校验均支持上述结论。此外,我们全程记录了一处测量泄漏——即标签信息通过数据的打包方式而非其内容进入系统——其中发布的文件排序编码了标签信息,包括一个曾短暂误导我们自身流程的实例。
cs.AI / 60 / 2609.00662

Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets

具有稀疏上下文与共享预算的漂移感知大语言模型路由
Lee, Cheung Hao, Wong, Patrick
Abstract
A multi-model language service must route each request while preserving workload-level budgets for compute, latency, memory, or monetary cost. Two features make this problem materially harder than static model selection. Prompt representations are high dimensional, so only a small subset of embedding directions may predict the incremental value of a model, and both the request mix and the model frontier drift after launches, fine-tunes, quantization changes, and system updates. We formulate nonstationary sparse contextual routing with multiple knapsack constraints and an optional shadow-audit stream that evaluates a small fraction of prompts on several models. We propose Drift-Aware Sparse Routing (DRS). The policy estimates reward and resource use from a rolling audit window, routes using pessimistic reward and optimistic cost estimates, updates resource shadow prices online, and applies a hard meter before commitment. The analysis separates control from statistics. On any event with uniform prediction radii $\{\beta_t\}$, regret against a paced dynamic fluid benchmark is bounded by the sum of the radii, a capacity-buffer term, and an $O(\sqrt{T})$ pacing term. Under a sparse linear model and bounded drift $V_T$, rolling estimation gives \[ \widetilde O\left( T\sqrt{\frac{s}{\rho W}}+WV_T+\sqrt{T} \right), \] where $s$ is sparsity, $\rho$ is the audit rate, and $W$ is the window length. Optimizing $W$ yields the usual stationary $O(\sqrt{sT/\rho})$ rate when $V_T=0$ and a $O(T^{2/3}(s/\rho)^{1/3}V_T^{1/3})$ adaptation term under drift.
Chinese Translation
多模型语言服务需要在为计算、延迟、内存或货币成本等工作负载级预算进行约束的同时,对每个请求进行路由。两个特性使该问题远比静态模型选择困难:其一,提示表示是高维的,因而可能只有少数嵌入方向能够预测某个模型的增量价值;其二,请求组合与模型前沿会在发布、微调、量化变更和系统更新之后发生漂移。我们将该问题形式化为带多个背包约束的非平稳稀疏上下文路由问题,并引入一个可选的影子审计流,用于在若干模型上评估少量提示。我们提出了漂移感知稀疏路由(Drift-Aware Sparse Routing, DRS)。该策略基于滚动审计窗口估计奖励与资源消耗,采用悲观的奖励估计和乐观的成本估计进行路由,在线更新资源影子价格,并在提交决策之前施加硬性计量约束。理论分析将控制与统计相分离。在预测半径 {β_t} 一致的任意事件下,相对于带节拍的动态流体基准的遗憾由各半径之和、一个容量缓冲项以及一个 O(√T) 的节拍项所界定。在稀疏线性模型与有界漂移 V_T 下,滚动估计给出 [ Õ( T√(s/(ρW)) + WV_T + √T ) ],其中 s 为稀疏度,ρ 为审计率,W 为窗口长度。对 W 进行优化后:当 V_T=0 时达到通常的平稳情形 O(√(sT/ρ)) 速率;在存在漂移时则产生 O(T^{2/3}(s/ρ)^{1/3}V_T^{1/3}) 的自适应项。
cs.AI / 61 / 2609.00665

Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs

面向边缘AI的语言模型三重底线可持续性:小型语言模型(SLMs)与量化大语言模型(LLMs)的比较
Shah, Jainil Dharmil
Abstract
Edge-AI model selection is commonly driven by one isolated metric - accuracy, latency, memory, energy, or safety, even though a deployable language model must balance all five. Our work focuses on answering the question whether na- tively trained small language models (SLMs) or large language models (LLMs) compressed through post-training quantization offer the more sustainable edge- deployment trade-off. We introduce a reproducible Holistic Sustainability Score (HSS) organized around the triple bottom line: an economic pillar for capability and systems efficiency, an environmental pillar for operational GPU energy and a social pillar for harmful-prompt robustness. Five BF16 SLMs and five LLMs under different quantization approaches - BF16, INT8, NF4 4-bit, GPTQ 4-bit, and GGUF Q4 produce 30 measured configurations. Capability is assessed on five zero-shot benchmarks; efficiency uses latency, throughput, peak VRAM and energy; and safety is approximated by attack success rate on five harmful prompts. Qwen3-30B-A3B/GGUF Q4 ranks first in the combined pool (93.38), followed by Mistral-Small-24B/GGUF Q4 (92.40), while Phi-4-mini/BF16 is the highest- ranked SLM in that pool (89.49). Thus, the hypothesis that native SLMs must be the most sustainable edge choice is not supported universally; optimized quantized LLMs can win overall, while SLMs remain competitive through lower resource demand. Quantization is a systems-level choice rather than a monotonic precision- efficiency trade-off and HSS remains relative to its comparison pool and proxy definitions.
Chinese Translation
边缘AI模型的选择通常由单一孤立指标驱动——准确率、延迟、内存、能耗或安全性,然而一个可部署的语言模型必须在这五个方面取得平衡。我们的工作旨在回答这样一个问题:原生训练的小型语言模型(SLMs)与通过训练后量化压缩的大语言模型(LLMs),哪一种能为边缘部署提供更可持续的权衡方案。我们提出了一个可复现的整体可持续性评分,围绕三重底线构建:经济支柱衡量能力与系统效率,环境支柱衡量运行时的GPU能耗,社会支柱衡量对有害提示的鲁棒性。五个BF16精度的SLMs和五个LLMs在不同量化方法下——BF16、INT8、NF4 4-bit、GPTQ 4-bit和GGUF Q4——共产生30个实测配置。能力通过五个零样本基准进行评估;效率采用延迟、吞吐量、峰值显存和能耗衡量;安全性则通过五个有害提示的攻击成功率来近似评估。Qwen3-30B-A3B/GGUF Q4在综合池中排名第一(93.38),其次是Mistral-Small-24B/GGUF Q4(92.40),而Phi-4-mini/BF16是该池中排名最高的SLM(89.49)。因此,原生SLMs必然是最可持续的边缘选择这一假设并不普遍成立;经过优化的量化LLMs可以在整体上胜出,而SLMs则凭借更低的资源需求保持竞争力。量化是一种系统层面的选择,而非单调的精度-效率权衡,且HSS依赖于其比较池和代理指标的定义。
cs.AI / 62 / 2609.00700

Value Over Language Model: Detecting Original Contribution in Writing

价值超越语言模型:检测写作中的原创性贡献
Sharma, Vibhhu, Joachims, Thorsten, Dean, Sarah
Abstract
LLMs have been rapidly adopted across writing tasks, prompting the development of tools for detecting LLM-generated text. Yet, these tools largely measure how much of a document's surface text was written by an LLM and aren't fundamentally designed to measure how much of the information content or ideas originated from the LLM itself rather than being supplied by the user in the prompt. In this work, we design a framework that measures how much value a person adds on top of what a language model could have easily produced by itself. The method requires no training or labeled data and never scores the document's surface text, insulating it from stylistic confounders. Instead, it extracts the document's content at increasing levels of granularity, uses an LLM to reconstruct the document from each partial representation, and compares these reconstructions with those produced from the task description alone. We call this framework Value Over Language Model (VOLM), which measures a document's contribution relative to a replacement-level document that an LLM could produce from the task description alone. We evaluate VOLM with a specific instantiation of this framework across three domains: news articles, ICLR peer reviews, and argumentative essays. VOLM separates human-authored documents from matched LLM-generated documents produced from generic task descriptions, while remaining substantially invariant to content-preserving transformations, including LLM-based reconstruction and round-trip translation. We further find that increasingly constrained content extractors reduce residual differences between LLM-generated and humanized text, demonstrating the importance of disentangling informational content from stylistic variation. We hope these results encourage further work on specialized instantiations of the framework and on assessing human contributions in LLM-assisted writing more generally.
Chinese Translation
大语言模型(LLM)已被迅速应用于各类写作任务,由此催生了检测LLM生成文本的工具的发展。然而,这些工具主要衡量的是文档表面文本有多少由LLM撰写,而并未从根本上设计用于衡量信息内容或思想中有多少源自LLM本身、而非由用户在提示(prompt)中提供。在本工作中,我们设计了一个框架,用于衡量一个人在语言模型本来就能轻松生成的内容之上所增加的价值。该方法无需训练或标注数据,也从不评分文档的表面文本,从而使其免受文体混淆因素的影响。相反,该方法以递增的粒度级别提取文档内容,使用LLM从每个部分表示重建文档,并将这些重建结果与仅从任务描述生成的重建结果进行比较。我们将此框架称为Value Over Language Model(VOLM),它衡量一篇文档相对于LLM仅凭任务描述就能产出的替代级文档的贡献。我们在三个领域中使用该框架的一个具体实例对VOLM进行评估:新闻文章、ICLR同行评审和议论文。VOLM能够将人类撰写的文档与由通用任务描述生成的匹配LLM生成文档区分开来,同时对内容保持不变的变换(包括基于LLM的重建和往返翻译)基本保持不变。我们进一步发现,约束性越来越强的内容提取器能够缩小LLM生成文本与人性化文本之间的残余差异,这证明了将信息内容与文体变化解耦的重要性。我们希望这些结果能够鼓励针对该框架的专门实例化研究,以及更广泛地评估LLM辅助写作中人类贡献的进一步工作。
cs.AI / 63 / 2609.00714

ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything

ChatDev 2.0:一个用于开发一切的无代码多智能体平台
Dang, Yufan, Yao, Shu, Lai, Bowen, Xu, Chenting, Shi, Ruijie, Leung, Wai-Shing, Li, Huatao, Qian, Chen, Liu, Zhiyuan
Abstract
Large language model (LLM)-based multi-agent systems (MAS) have shown strong potential for solving complex tasks, yet their development forces a tradeoff: code frameworks are expressive but engineering-intensive, while no-code builders simplify authoring but constrain agent interactions to author-defined workflows. We present ChatDev 2.0: DevAll (hereafter DevAll), a no-code platform for building, executing, and inspecting heterogeneous MAS that delivers both high expressiveness and ease of use. In terms of expressiveness, DevAll pairs a declarative executable graph abstraction with a cycle-aware execution engine, so that heterogeneous agents and dynamic and cyclic interactions can be represented and executed within a single framework. For ease of use, an integrated visual interface lets users author, run, monitor, and inspect MAS, including human-in-the-loop steps, entirely without writing code. Experiments demonstrate that DevAll reproduces state-of-the-art MAS across three representative tasks at competitive performance and without task-specific orchestration code, highlighting its effectiveness as a general-purpose platform for LLM-based MAS. DevAll is available at https://github.com/OpenBMB/ChatDev.
Chinese Translation
基于大语言模型(LLM)的多智能体系统(MAS)在解决复杂任务方面展现出强大潜力,但其开发过程面临两难权衡:代码框架表达能力强但工程量大,而无代码构建工具虽简化了构建过程,却将智能体交互限制在作者预定义的工作流中。我们提出了ChatDev 2.0: DevAll(以下简称DevAll),这是一个用于构建、执行和检查异构多智能体系统的无代码平台,兼具高表达能力与易用性。在表达能力方面,DevAll将声明式可执行图抽象与具备循环感知能力的执行引擎相结合,使异构智能体以及动态和循环交互能够在单一框架内被表示和执行。在易用性方面,一个集成的可视化界面让用户能够完全无需编写代码即可构建、运行、监控和检查多智能体系统,包括人在回路(human-in-the-loop)环节。实验表明,DevAll在三个代表性任务上以有竞争力的性能复现了最先进的多智能体系统,且无需针对特定任务的编排代码,凸显了其作为基于LLM的多智能体系统通用平台的有效性。DevAll已发布于 https://github.com/OpenBMB/ChatDev。
cs.AI / 64 / 2609.00718

A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies

压缩驾驶策略能力损失与恢复的闭环评估
Irfan, Ahmad Alfan Alfian, Khatim, Nur Ahmad, Arief, Mansur
Abstract
Many automobile and mobility companies deploy learned driving policies on embedded computers with limited memory and power. Pruning, knowledge distillation, and quantization are the standard methods to reduce the size and the inference cost of these policies. However, these methods are commonly assessed by aggregate numerical scores, and such scores may not reflect the ability of the policy to drive safely when interacting with other road users. In this study, we propose a stage-wise closed-loop evaluation approach to follow a driving policy through a compression pipeline. We formulate the driving task as a partially observable Markov decision process (POMDP) and train a belief-state policy with proximal policy optimization (PPO) in Gym-Duckietown. We then extract the actor, compress it one stage at a time, and evaluate it on five driving curricula. We show that structured pruning is the stage at which the driving capability is first lost. Meanwhile, distillation improves the pruned actor, but the improvement is limited by its rehearsal data. Integer quantization of the improved actor loses some of the curricula that require the vehicle to stop and then resume. Interestingly, the same procedure on the unpruned actor preserves all five curricula. Our study thus provides an empirical analysis aiming to answer the currently active discussions on how to accept a compressed driving policy, so as to achieve a safe and statistically reliable deployment of automated driving functions.
Chinese Translation
许多汽车和出行公司在内存和算力受限的嵌入式计算机上部署学习型驾驶策略。剪枝、知识蒸馏和量化是降低这些策略规模和推理成本的标准方法。然而,这些方法通常通过汇总的数值分数进行评估,而此类分数可能无法反映策略在与其他道路使用者交互时安全驾驶的能力。在本研究中,我们提出一种分阶段闭环评估方法,以跟踪驾驶策略在整个压缩流水线中的表现。我们将驾驶任务建模为部分可观测马尔可夫决策过程(POMDP),并在 Gym-Duckietown 中使用近端策略优化(PPO)训练信念状态策略。随后,我们提取行动者网络,逐阶段对其进行压缩,并在五个驾驶课程上进行评估。结果表明,结构化剪枝是驾驶能力首次出现损失的环节。与此同时,蒸馏能够改善剪枝后的行动者网络,但改善程度受限于其重放数据。对改善后的行动者网络进行整数量化,会损失部分需要车辆先停车再重新起步的课程。有趣的是,对未剪枝的行动者网络执行相同流程则能保留全部五个课程。因此,本研究提供了一项实证分析,旨在回应目前关于如何接受压缩驾驶策略的活跃讨论,从而实现自动驾驶功能安全且统计可靠的部署。
cs.AI / 65 / 2609.00728

SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification

SOVER:基于LLM辅助SMT验证的优化问题重构形式化认证
Bhattacharyya, Swapnil, Baranwal, Mayank
Abstract
Large Language Models (LLMs) have shown remarkable promise in translating and reformulating complex mathematical optimization problems across modeling languages. However, validating such transformations through empirical solver executions alone is unreliable, as solver outcomes may be affected by local minima, structural timeouts, numerical artifacts, and subtle semantic divergence between formulations. We introduce SOVER, an LLM-assisted SMT framework that separates semantic mapping from formal certification: Z3 checks domain cross-feasibility and global objective-order preservation for mixed-integer linear formulations, while dReal provides tolerance-aware feasibility/range and $\epsilon$-argmin checks for continuous nonlinear formulations. We also introduce NLEquiv-150, a public benchmark of 100 equivalent and 50 deliberately hard non-equivalent nonlinear reformulation pairs. With LLM-extracted mappings, SOVER classifies 149/150 pairs (99.33%) correctly, including all 50 hard negatives; the sole error is an incomplete mapping extraction.
Chinese Translation
大型语言模型(LLM)在不同建模语言之间翻译和重构复杂数学优化问题方面展现出巨大潜力。然而,仅通过经验性的求解器运行来验证此类转换并不可靠,因为求解器的结果可能受到局部极小值、结构性超时、数值伪影以及不同公式表述之间细微语义差异的影响。我们提出了SOVER,一个LLM辅助的SMT(可满足性模理论)框架,它将语义映射与形式化认证分离:对于混合整数线性规划公式,由Z3检验域间交叉可行性以及全局目标序保持性;对于连续非线性规划公式,由dReal提供考虑容差的可行性/取值范围检验以及 $\epsilon$-argmin(ε-最小值)检验。我们还提出了NLEquiv-150,一个包含100对等价和50对特意构造的困难非等价非线性重构问题的公开基准数据集。借助LLM提取的映射,SOVER正确分类了149/150对(99.33%),包括全部50对困难负样本;唯一的错误源于一次不完整的映射提取。
cs.AI / 66 / 2609.00731

Agentic Empirical Asset Pricing: Methodological Foundations

智能体化实证资产定价:方法论基础
Pan, Yingjian, Ding, Xiaowei, Giesecke, Kay
Abstract
Recent advances in LLM agents enable a new paradigm for asset pricing, which we call Agentic Empirical Asset Pricing (AEAP): systems that autonomously conduct the scientific discovery process itself. We define AEAP and identify its core building blocks. Existing evaluation practices backtest only the outputs (factors or trades), not the autonomous discovery system that produced them. We focus on factor discovery, contributing a reference architecture, a rigorous evaluation standard for discovered factors, and a method for out-of-sample backtesting the discovery system. As a concrete instance of that architecture, we evaluate SEADS against five re-implemented baselines on two US equity panels using this standard: no single metric ranks the systems consistently, motivating evaluation on multiple axes at once. A separate rolling re-execution then asks the complementary question of whether the discovery process itself, not one static output, is reliable. We also report negative findings and limitations that surface further evaluation pitfalls for future AEAP systems.
Chinese Translation
大语言模型(LLM)智能体的最新进展为资产定价开辟了一种新范式,我们称之为智能体化实证资产定价(Agentic Empirical Asset Pricing,AEAP):即能够自主开展科学发现过程本身的系统。我们定义了 AEAP 并识别其核心构成要素。现有的评估实践仅对输出结果(因子或交易)进行回测,而非对产生这些结果的自主发现系统进行评估。我们聚焦于因子发现问题,贡献了一个参考架构、一套针对所发现因子的严格评估标准,以及一种对该发现系统进行样本外回测的方法。作为该架构的具体实例,我们基于此标准,在两个美股面板数据上将 SEADS 与五个重新实现的基线模型进行对比评估:没有任何单一指标能够一致地对各系统进行排序,这表明需要在多个维度上同时进行评估。此外,我们通过独立的滚动重复执行实验,探讨了一个互补性问题:可靠的是发现过程本身,而非某个静态的输出结果。我们还报告了负面结果与局限性,揭示了未来 AEAP 系统评估中需要进一步注意的陷阱。
cs.AI / 67 / 2609.00738

Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs

摆脱冗余推理:面向推理时大语言模型的结构感知搜索
Cheng, Lu
Abstract
Inference-time search with large language models (LLMs) often concentrates on a small set of structurally or semantically similar trajectories, leaving alternatives underexplored---a failure mode we call \textit{reasoning basin collapse}. We introduce BASIN, a training-free, structure-aware selection method that groups reasoning states into basins and penalizes repeated visits to the same strategy, thereby reallocating search across genuinely distinct reasoning paths under a fixed compute budget. Under matched inference budgets, BASIN improves over Tree of Thoughts (ToT) by up to $+22$pp on Game of 24 and $+6.7$pp on MuSR. A quality-aware variant, QA-BASIN, further improves robustness by preserving high-quality basins when unconditional diversification over-explores. To explain when basin-aware selection helps, we introduce the redundancy gap $\Delta$, which measures how differently search concentrates for correct versus incorrect predictions: standard ToT often operates near $\Delta \approx 0$, while BASIN consistently shifts $\Delta$ positive. More broadly, BASIN suggests structure-aware selection as a simple and general approach to improving inference-time reasoning. Code can be found at https://github.com/GitHubLuCheng/basin.
Chinese Translation
基于大语言模型(LLM)的推理时搜索往往集中于少量结构或语义相似的轨迹,导致其他替代路径探索不足——我们将这种失败模式称为"推理盆地坍缩"(reasoning basin collapse)。我们提出 BASIN,一种免训练、结构感知的选择方法,它将推理状态分组为多个盆地,并对重复访问同一策略进行惩罚,从而在固定计算预算下将搜索重新分配到真正不同的推理路径上。在匹配的推理预算下,BASIN 在 Game of 24 上相比思维树(Tree of Thoughts, ToT)提升最高达 $+22$ 个百分点,在 MuSR 上提升 $+6.7$ 个百分点。其质量感知变体 QA-BASIN 进一步增强了鲁棒性:当无条件的多样化探索过度扩展时,它能够保留高质量的盆地。为了解释盆地感知选择何时有效,我们引入冗余差距(redundancy gap)$\Delta$,用于衡量搜索在正确与错误预测上的集中程度差异:标准 ToT 常常运行在 $\Delta \approx 0$ 附近,而 BASIN 则能持续使 $\Delta$ 保持正值。更广泛地说,BASIN 表明结构感知选择是提升推理时推理能力的一种简单而通用的方法。代码见 https://github.com/GitHubLuCheng/basin。
cs.AI / 68 / 2609.00749

ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents

ContextPipe:面向长程智能体的数据库式上下文组装方法
Xu, Peng, Zhang, Zuyu, Sun, Yuze, Tian, Feng, Wang, Long, Zhang, Chen
Abstract
Long-horizon large language model (LLM) agents require context assembly: the runtime must decide what to include in each prompt, in what order, and when to compact history under a hard context-window budget and a byte-sensitive prompt cache. In production agentic systems, this logic is scattered across prompt builders, ad hoc compaction routines, cache-break workarounds, and per-provider shims. We argue that context assembly is structurally isomorphic to query execution in a relational database: both execute under a hard budget, exploit a tiered cache, and leverage statistics. We adopt this discipline in ContextPipe: a five-phase pipeline (Plan Bind Optimize Execute Feedback) backed by a structured data-source catalog, a deterministic cache-aware optimizer, and an EXPLAIN ANALYZE trace. We show that context in ContextPipe is auditable, replayable, and failure-isolated. A preliminary evaluation using the SWE-bench Pro Qutebrowser subset shows that, compared with the append-only context construction policy, ContextPipe reduces total token volume by 31%, LLM calls by 23%, and response time by 9%, at the cost of a lower KV cache-hit ratio.
Chinese Translation
长程大语言模型(LLM)智能体需要进行上下文组装:运行时系统必须在严格的上下文窗口预算和对字节敏感的提示缓存的约束下,决定在每个提示中包含什么内容、以何种顺序包含,以及何时压缩历史记录。在生产级智能体系统中,这类逻辑散布于提示构建器、临时性的压缩例程、缓存失效规避方案以及针对不同服务商的适配层之中。我们认为,上下文组装在结构上与关系型数据库中的查询执行同构:两者都在严格预算下执行、利用分层缓存并借助统计信息。我们在 ContextPipe 中采用这一准则:一个由结构化数据源目录、确定性的缓存感知优化器以及 EXPLAIN ANALYZE 追踪所支撑的五阶段流水线(Plan、Bind、Optimize、Execute、Feedback)。我们证明 ContextPipe 中的上下文是可审计、可重放且故障隔离的。基于 SWE-bench Pro Qutebrowser 子集的初步评估表明,与仅追加式(append-only)上下文构建策略相比,ContextPipe 以较低的 KV 缓存命中率为代价,将总 token 数量减少了 31%,LLM 调用次数减少了 23%,响应时间缩短了 9%。
cs.AI / 69 / 2609.00755

S^3martCirc: Self-supervised Smart Circuit Discovery

S^3martCirc:自监督的智能电路发现
Zheng, Wendy, He, Yinhan, Wu, Liang, Li, Jundong
Abstract
Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks, from text summarization to question answering. Despite these capabilities, their black-box nature obscures internal decision-making processes. Mechanistic interpretability (MI) aims to address this by reverse-engineering neural networks into human-understandable algorithms. Current MI approaches for LLMs typically follow a two-stage paradigm: first identifying important components (circuit discovery), where components are typically individual nodes such as an attention head or feedforward neuron, and second determining the role they play in a certain task (functional interpretation). However, this sequential approach overlooks a fundamental insight: a component's importance and its functional role are inherently codependent. Unifying these stages presents two key challenges: (1) functional roles are often tied to specific nodes or components, limiting generalization, and (2) their identification relies on subjective interpretation rather than quantifiable metrics. To address these challenges, we propose S^3martCirc (Self-supervised Smart Circuit Discovery), a unified framework that simultaneously discovers circuits and interprets functionality. S^3martCirc abstracts node behavior into two general computational roles that generalize across tasks and defines a quantitative metric for assigning them, enabling importance and functional role to be discovered jointly rather than in sequence. Extensive experiments show that our framework outperforms existing methods in circuit discovery.
Chinese Translation
大语言模型(LLMs)在从文本摘要到问答等多种任务中表现出卓越性能。尽管具备这些能力,但其黑盒特性掩盖了内部决策过程。机制可解释性(Mechanistic Interpretability, MI)旨在通过将神经网络逆向工程为人类可理解的算法来解决这一问题。当前针对LLMs的MI方法通常遵循两阶段范式:首先识别重要组件(电路发现),其中组件通常为单个节点,如注意力头或前馈神经元;其次确定它们在特定任务中所扮演的角色(功能解释)。然而,这种顺序式方法忽略了一个基本事实:组件的重要性与其功能角色本质上是相互依存的。统一这两个阶段面临两个关键挑战:(1)功能角色往往与特定节点或组件绑定,限制了泛化能力;(2)其识别依赖于主观解释而非可量化的指标。为应对这些挑战,我们提出了S^3martCirc(Self-supervised Smart Circuit Discovery,自监督智能电路发现),这是一个能够同时发现电路并解释功能的统一框架。S^3martCirc将节点行为抽象为两种可跨任务泛化的通用计算角色,并定义了用于分配这些角色的量化指标,从而能够联合而非顺序地发现重要性和功能角色。大量实验表明,我们的框架在电路发现方面优于现有方法。
cs.AI / 70 / 2609.00763

Automated Tree Knowledge Graph Construction using Ontology Expansion and Retrieval from Vietnamese History Textbooks

基于本体扩展与检索的越南历史教科书树状知识图谱自动构建
Nguyen, Ket Doan, Nguyen, Minh N. H.
Abstract
Hierarchical Knowledge graph (KG)-based retrieval augmented generation (RAG) has emerged as a powerful approach for supporting large language models with structured knowledge. However, there are primary challenges: (i) the lack of methods for automatic KG construction using ontology expansion for low-resource languages such as Vietnamese, (ii) the absence of systematic evaluation for knowledge retrieval strategies leveraging the hierarchical structures. In this paper, we propose an end-to-end pipeline for KG construction and retrieval strategies evaluation. In the KG construction, we employ a three-phase hybrid relation extraction pipeline: intra-batch deduplication via Union-Find, approximate cross-batch search, and LLM extraction with a centroid filter that reduces prompts combined with a five-step dual-LLM validator to prevent bloated ontology. A two-tier architecture consists of unmergeable structural nodes to preserve the document structure and mergeable content nodes. The retrieval evaluation consists of three graph traversal strategies: Top-Down, Horizontal, and Bottom-Up, which are evaluated on a synthetically generated benchmark of 1,210 Vietnamese queries from 109 subgraphs, categorized by five query directions. In this paper, we construct the tree knowledge graph from Vietnamese high school History textbooks (nearly 400 pages) to produce 750 nodes and 4,341 semantic edges with controlled ontology growth from 40 to 41 types. Among experimental graph traversal strategies, the Top-Down strategy with structure surpasses the vector baseline by 4.7 percentage points in NDCG@10. As a result, tree-structural information provides valuable information beyond flat cosine similarity but degrades performance when the query does not require structural context.
Chinese Translation
基于层次化知识图谱(KG)的检索增强生成(RAG)已成为利用结构化知识支持大语言模型的一种强大方法。然而,目前存在两个主要挑战:(i)针对越南语等低资源语言,缺乏利用本体扩展自动构建知识图谱的方法;(ii)缺乏对利用层次化结构的知识检索策略的系统评估。本文提出了一个端到端的知识图谱构建与检索策略评估流水线。在知识图谱构建方面,我们采用三阶段混合关系抽取流水线:基于并查集(Union-Find)的批内去重、近似跨批搜索,以及结合质心过滤器以减少提示词的大语言模型(LLM)抽取,并配合五步双LLM验证器以防止本体膨胀。系统采用双层架构:包含不可合并的结构节点以保留文档结构,以及可合并的内容节点。检索评估包含三种图遍历策略:自顶向下(Top-Down)、水平(Horizontal)和自底向上(Bottom-Up),这些策略在一个合成生成的基准数据集上进行评估,该数据集包含来自109个子图的1,210个越南语查询,并按五种查询方向分类。本文基于越南高中历史教科书(近400页)构建了树状知识图谱,生成750个节点和4,341条语义边,同时将本体类型从40种可控地增长至41种。在实验的图遍历策略中,结合结构的自顶向下策略在NDCG@10指标上超越向量基线4.7个百分点。结果表明,树状结构信息能够提供超越扁平余弦相似度的有价值信息,但当查询不需要结构上下文时,其性能会有所下降。
cs.AI / 71 / 2609.00768

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory

DiagEvo:基于层次化错误记忆的诊断引导自进化
Wei, Xincheng, Ding, Yifan, Li, Yoshua, Ma, Dongsheng, Weng, Rongxiang, Cai, Xunliang, Ding, Wenjian, Zhang, Yao
Abstract
Self-play is an effective paradigm for language-model self-evolution, but without guidance, solver performance can plateau or decline across rounds. Unguided methods steer question generation with signals such as difficulty, learnability, or diversity. These signals keep questions challenging and varied but do not specify which unresolved reasoning weaknesses later rounds should target. Guided methods obtain direction from external task resources, including human examples, document corpora, or specified difficulty targets, and therefore rely on task information supplied outside the self-play loop. We show that the needed direction can instead be derived from the solver's own failure history. We introduce DiagEvo, whose diagnostician extracts recurring error causes from this history and stores them in a hierarchical error-cause memory. The memory groups related causes under skill nodes and tracks each as Active or Mastered according to self-consistency on targeted questions. The challenger uses these states and recurrence counts to balance cause-targeted generation with free exploration. Double-confidence filtering retains intermediate-difficulty questions only when the most common solver answer has a clear vote lead. DiagEvo derives its curriculum from information produced during self-play, without external task resources. With the default 4B diagnostician, DiagEvo outperforms every baseline in mean accuracy across all nine benchmarks for each of the three solvers: Qwen3-4B, Qwen3-8B, and OctoThinker-8B. On Qwen3-8B, it reaches 72.3% mean accuracy across five mathematical reasoning benchmarks, 4.5 percentage points above R-Zero. Its mean accuracy across all nine benchmarks is 57.4%, 1.1 percentage points above DARC. Ablations show that the hierarchical error-cause memory and double-confidence filtering both contribute to these gains.
Chinese Translation
自博弈(Self-play)是语言模型自进化的有效范式,但在缺乏引导的情况下,求解器(solver)的性能可能在多轮迭代中停滞甚至下降。无引导方法利用难度、可学习性或多样性等信号来引导问题生成,这些信号能保持问题的挑战性和多样性,却无法指明后续轮次应针对哪些尚未解决的推理弱点。有引导方法则从外部任务资源中获取方向,包括人类示例、文档语料库或指定的难度目标,因而依赖于自博弈循环之外提供的任务信息。我们证明,所需的方向可以转而从求解器自身的失败历史中获得。我们提出DiagEvo,其诊断模块(diagnostician)从失败历史中提取反复出现的错误原因,并将其存储在层次化错误原因记忆中。该记忆将相关错误原因归组于技能节点之下,并根据求解器在针对性问题上的自一致性(self-consistency),将各原因标记为“活跃(Active)”或“已掌握(Mastered)”。挑战者模块(challenger)利用这些状态及复发计数,在针对错误原因的生成与自由探索之间取得平衡。双重置信度过滤机制仅当最常见的求解器答案具有明显票数领先时,才保留中等难度的问题。DiagEvo的课程完全源自自博弈过程中产生的信息,无需外部任务资源。在默认的4B诊断模块下,DiagEvo在三个求解器(Qwen3-4B、Qwen3-8B和OctoThinker-8B)上,于全部九个基准测试中的平均准确率均超越所有基线方法。在Qwen3-8B上,其在五个数学推理基准上的平均准确率达到72.3%,比R-Zero高出4.5个百分点;其在全部九个基准上的平均准确率为57.4%,比DARC高出1.1个百分点。消融实验表明,层次化错误原因记忆和双重置信度过滤机制对上述提升均有贡献。
cs.AI / 72 / 2609.00782

When Features Become Instances: Inverted Contrastive Learning for Unsupervised Feature Selection

当特征成为实例:面向无监督特征选择的反转对比学习
Ghosh, Utsab, Chakraborty, Roshni
Abstract
Unsupervised feature selection seeks a compact subset of informative features without access to class labels, making feature utility difficult to define. Existing UFS methods therefore rely on indirect structural criteria, such as similarity preservation, locality, sparsity, cluster geometry, or reconstruction quality. In this paper, we instead study UFS through representation consistency and propose Inverted Contrastive Learning for Unsupervised Feature Selection (ICLFS), a feature-wise contrastive framework that reformulates UFS as a representation learning problem over features rather than samples. ICLFS first inverts the data matrix so that each feature is represented by its sample-profile vector, then constructs multiple masked positive views together with a shuffled negative view, and learns projector-space representations that remain consistent across these structured perturbations under an InfoNCE-based objective. Motivated by recent findings that cosine-based and InfoNCE-based training affect embedding norms, we use projector-space embedding magnitude as the saliency signal for ranking features. The resulting norm-based ranking is subsequently refined through Laplacian-Gated Ranking Correction, which suppresses locally redundant candidates while preserving salient ones. Extensive experiments on 12 benchmark datasets show that ICLFS achieves the best clustering accuracy on 10 datasets against both classical and neural baselines under the standard clustering-based UFS evaluation protocol, while remaining competitive on the other two. These results show that feature-wise contrastive representation consistency provides a strong and effective alternative to neighborhood, cluster, and reconstruction-based UFS formulations.
Chinese Translation
无监督特征选择(UFS)旨在无法获取类别标签的情况下,寻找一个紧凑且信息丰富的特征子集,这使得特征效用的定义变得困难。因此,现有的UFS方法依赖于间接的结构性准则,例如相似性保持、局部性、稀疏性、聚类几何结构或重构质量。在本文中,我们转而通过表征一致性来研究无监督特征选择,并提出了一种面向无监督特征选择的反转对比学习方法(Inverted Contrastive Learning for Unsupervised Feature Selection, ICLFS)。这是一个以特征为单位的对比学习框架,将UFS重新表述为一个针对特征而非样本的表征学习问题。ICLFS首先对数据矩阵进行反转,使每个特征由其样本轮廓向量表示;然后构建多个掩码正样本视图以及一个打乱顺序的负样本视图,并在基于InfoNCE的目标函数下,学习在这些结构化扰动之间保持一致的投影空间表征。受近期关于余弦相似度和基于InfoNCE的训练会影响嵌入范数的研究发现启发,我们使用投影空间嵌入的幅值作为对特征排序的显著性信号。随后,通过拉普拉斯门控排序修正(Laplacian-Gated Ranking Correction)对基于范数的排序结果进行优化,在保留显著特征的同时抑制局部冗余的候选特征。在12个基准数据集上的大量实验表明,在基于聚类的标准UFS评估协议下,ICLFS在10个数据集上的聚类准确率优于经典方法与神经网络基线,并在其余两个数据集上保持竞争力。这些结果表明,以特征为单位的对比表征一致性为基于邻域、聚类和重构的UFS方法提供了一种强大而有效的替代方案。
cs.AI / 73 / 2609.00787

StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?

StudyBench:自我进化能否从教科书中榨取奥赛级能力?
Chen, Yinghao, Chen, Zixi, He, Bingxiang, Qiao, Ziqing, Gao, Huan-ang, Xu, Yinuo, Zuo, Yuxin, Liu, Zeyuan, Zhan, Yuhao, Xiao, Chaojun
Abstract
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems. We argue that an ideal self-evolution method should share the same property, that is autonomously learning from raw training material for transferable problem-solving capability. However, we still lack a direct measurement for it. We introduce StudyBench, a controlled physics benchmark that directly measures how efficiently a self-evolution method converts training material into capability. We organise the test set into an Application Set, consisting of difficult textbook problems and evaluating absorption ability, and a Transfer Set, consisting of olympiad-level problems and evaluating transfer ability. Benchmarking representative self-evolution methods across three base models, we find that improvements on the Application Set rarely translate to the harder Transfer Set. A guidance ablation exposes a Guidance Gap: even the strongest method closes only a small fraction of what the same material unlocks when supplied as in-context guidance. Besides, every method hits a Compute Plateau, saturating well before exhausting its compute budget. The remaining gap is therefore a method problem rather than a data or compute problem. By offering a clean and controlled benchmark, StudyBench turns self-evolution progress from an open-ended pursuit into a measurable target for future research. Our code is released at https://github.com/thunlp/StudyBench.
Chinese Translation
人类只需研读少量编写精良的教科书即可掌握一门学科并尝试解决其中最困难的问题。我们认为,理想的自我进化方法应具备同样的特性,即能够从原始训练材料中自主学习,获得可迁移的问题求解能力。然而,目前我们仍缺乏对这一能力的直接度量。我们提出了StudyBench,一个受控的物理基准,用于直接衡量自我进化方法将训练材料转化为能力的效率。我们将测试集分为两部分:应用集,由困难的教科书题目组成,评估吸收能力;迁移集,由奥赛级题目组成,评估迁移能力。通过对三个基座模型上具有代表性的自我进化方法进行基准测试,我们发现应用集上的提升很少能转化为更难的迁移集上的提升。一项关于引导的消融实验揭示了“引导鸿沟”:即使是性能最强的方法,所能实现的提升也仅占相同材料作为上下文引导时所带来提升的一小部分。此外,每种方法都会触及“算力平台期”,在耗尽计算预算之前便早已趋于饱和。因此,剩余的能力差距是方法问题,而非数据或算力问题。通过提供一个干净且可控的基准,StudyBench将自我进化的研究进展从开放式的探索转变为未来研究中可度量的目标。我们的代码已发布于 https://github.com/thunlp/StudyBench。
cs.AI / 74 / 2609.00805

Towards a Reliable and Practical Eval Pipeline

迈向可靠且实用的评估流水线
Nguyen, Emma Thuong, Ghose, Abhishek
Abstract
LLM-based software systems increasingly require effective "evals" as quality gates in the development lifecycle. However, existing work typically addresses individual aspects of eval reliability rather than the full set of practical requirements. We present an end-to-end eval pipeline that combines eval checklist creation, with learned aggregation for checklist responses, to improve agreement across LLM judges and accuracy against human judgments. The framework additionally pro- vides self-consistency, explanations, and prediction uncertainty, and we empirically demonstrate its effectiveness.
Chinese Translation
基于大语言模型(LLM)的软件系统日益需要在开发生命周期中将有效的“评估”(evals)作为质量门槛。然而,现有工作通常只解决评估可靠性的个别方面,而非全部实际需求。我们提出了一种端到端的评估流水线,它将评估清单(eval checklist)的创建与针对清单响应的学习式聚合相结合,以提升LLM评审者之间的一致性以及相对于人类评判的准确性。该框架还提供自一致性、解释说明和预测不确定性,并通过实验证明了其有效性。
cs.AI / 75 / 2609.00813

One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning

一个策略,任意预算:通过强化学习内化预算感知搜索
Sun, Xiaowei, Li, Jin, Hong, Yili, Fu, Yikun, Xiao, Yanghua
Abstract
While reinforcement learning has enabled LLM-based search agents to invoke external tools, existing methods train under fixed budgets and cannot adapt when constraints vary at deployment. We propose AnySearch, a framework that enables a single policy to perform budget-aware search under any budget constraint through a training scaffold and curriculum reinforcement learning. In the first phase, we train the agent with explicit budget state injection and structured reasoning prompts that guide efficient allocation under linearly decaying budgets. In the second phase, the scaffold is removed and the agent learns to operate autonomously under adaptively sampled budget constraints, matching inference conditions. Both phases are optimized with a composite reward that couples answer accuracy with budget efficiency through absolute and relative signals, where an adaptive weight amplifies the efficiency signal for high-accuracy queries and attenuates it for low-accuracy ones. Extensive experiments on seven general and multi-hop QA benchmarks show that our method outperforms baselines across all budget scales, generalizes to unseen constraints beyond the training range, and achieves superior tool productivity without excessive token overhead. Our code is available at https://github.com/xwsun01/AnySearch.
Chinese Translation
尽管强化学习已使基于大语言模型(LLM)的搜索智能体能够调用外部工具,但现有方法在固定预算下训练,无法在部署时约束条件变化的情况下进行自适应调整。我们提出 AnySearch,一个通过训练脚手架和课程强化学习,使单一策略能够在任意预算约束下执行预算感知搜索的框架。在第一阶段,我们通过显式的预算状态注入和结构化推理提示来训练智能体,引导其在线性衰减的预算下进行高效分配。在第二阶段,移除脚手架,智能体学习在自适应采样的预算约束下自主运行,以匹配推理时的条件。两个阶段均通过一个复合奖励进行优化,该奖励通过绝对和相对信号将答案准确性与预算效率耦合起来,其中自适应权重在高准确率的查询上放大效率信号,而在低准确率的查询上减弱该信号。在七个通用及多跳问答基准上的大量实验表明,我们的方法在所有预算规模下均优于基线方法,能够泛化到训练范围之外的未见约束,并在不产生过多 token 开销的情况下实现了更优的工具使用效率。我们的代码发布于 https://github.com/xwsun01/AnySearch。
cs.AI / 76 / 2609.00818

AnalysisBank: An Expert Analysis Pattern Library for Financial Report Generation

AnalysisBank:面向金融报告生成的专家分析模式库
Yang, Yajing, Ma, Yunshan, Koa, Kelvin J. L., Kan, Min-Yen
Abstract
We argue that financial report generation should operate at the analytical rather than structural level, composing content from data-derived insights rather than high-level topics or sections. To this end, we propose AnalysisBank, which distills expert reports into a reusable library of Analyses, each pairing a data signal, an analytical move, and the expert span it was derived from. At inference time, AnalysisBank matches input signals to library entries and applies the retrieved moves to compose the report. A study of Analyses distilled from 550 expert reports reveals a heavy-tailed distribution of 47-52 signal types spanning 13 move types. On two financial benchmarks across four LLM backbones, AnalysisBank increases the proportion of novel, data-grounded insights by 1.7-3.7x over structural-level baselines. Transfer to scientific writing suggests that the distinction generalizes beyond finance. Code and the distilled Analysis library are available at https://github.com/yajingyang/AnalysisBank.
Chinese Translation
我们认为,金融报告生成应在分析层面而非结构层面进行,即从数据中提炼的洞见来组织内容,而非基于高层次的主题或章节。为此,我们提出 AnalysisBank,它将专家报告提炼为一个可复用的分析(Analyses)库,其中每条分析都由一个数据信号、一个分析手法(analytical move)及其来源的专家文本片段组成。在推理阶段,AnalysisBank 将输入信号匹配到库中的条目,并应用检索到的分析手法来撰写报告。对从 550 篇专家报告中提炼出的分析进行的研究表明,其分布呈重尾特征,涵盖 13 种分析手法类型的 47-52 种信号类型。在四个 LLM 骨干模型上的两个金融基准测试中,与结构层面基线相比,AnalysisBank 将新颖的、有数据依据的洞见比例提升了 1.7-3.7 倍。向科学写作的迁移实验表明,这一区分可推广至金融领域之外。代码及提炼出的分析库已发布于 https://github.com/yajingyang/AnalysisBank。
cs.AI / 77 / 2609.00823

Polished but Unresolved: Identifying Late-Stage Pressure States in Long-Horizon Tool-Use Agents

表面完善但问题未决:识别长时程工具使用智能体中的后期压力状态
Chen, Haoyang, Liu, Yi, Shao, Jianzhi, Xu, Xiaozhou, Sun, Zhe, Hu, Wei
Abstract
Long-horizon tool-use agents need not only to search and plan, but also to decide when to finalize. We study late-stage pressure states, in which an agent is biased toward submitting a final answer that appears complete and polished while key constraints remain unresolved. We first train a linear probe to show that this pressure state is identifiable from the agent's hidden states. Then, we use activation interventions along this pressure direction and find that shifting the hidden states changes both the pressure score and whether the agent continues tool use or submits early. Through controlled context manipulations, we further see that the pressure is mitigated by constraint clarity and action mapping. Based on these findings, we propose Probe-Sensed Pressure Relief (PSPR), a plugin that applies lightweight pressure relief direction under moderate pressure and moves to structured organization under high pressure risk. Experiments on multiple long-horizon benchmarks show that our method consistently strengthens existing agent methods.
Chinese Translation
长时程工具使用智能体(Long-Horizon Tool-Use Agents)不仅需要搜索和规划,还需要决定何时提交最终结果。我们研究了后期压力状态(late-stage pressure states),在这种状态下,智能体倾向于提交一个看似完整且完善的最终答案,而关键约束尚未得到解决。我们首先训练了一个线性探针(linear probe),证明这种压力状态可以从智能体的隐藏状态中识别出来。然后,我们沿该压力方向进行激活干预(activation interventions),发现移动隐藏状态会同时改变压力分数以及智能体是继续使用工具还是提前提交。通过受控的上下文操纵,我们进一步发现约束清晰度和动作映射可以缓解这种压力。基于这些发现,我们提出了探针感知压力缓解(Probe-Sensed Pressure Relief, PSPR),这是一个插件,在中等压力下应用轻量级的压力缓解方向,在高压力风险下则转向结构化组织。在多个长时程基准上的实验表明,我们的方法能够持续增强现有的智能体方法。
cs.AI / 78 / 2609.00831

FLaG: Frequency-Domain Latent-attention Gated Pooling for Token Aggregation

FLaG:用于词元聚合的频域潜在注意力门控池化方法
Li, Kewei, Zhang, Rongying, Wang, Xueli, Gong, Xiwen, Wang, Zhongjian, Zhao, Qiuchen, Huang, Lan, Zhang, Ruochi, Zhou, Fengfeng
Abstract
Token aggregation converts token-level representations into fixed-dimensional sample representations, but most pooling methods operate only in the original token space. We introduce Frequency-Domain Latent-attention Gated Pooling (FLaG), a plug-in aggregation module that re-expresses encoder outputs in the Fourier domain before final pooling. FLaG represents the nonredundant rFFT spectrum through concatenated real and imaginary components, summarizes spectral tokens with learnable latent queries, derives a sample-conditioned channel gate, and reconstructs modulated token representations for downstream aggregation. We evaluate the same architecture across ESM2-based antimicrobial peptide (AMP) activity prediction, ResNet18 image classification on CIFAR-10 and CIFAR-100, and three RoBERTa-based language tasks. FLaG achieves the best macro-averaged Spearman correlation coefficient, RMSE, and Recall@50 across four AMP backbone-species settings and the highest top-1 accuracy on CIFAR 10. It also achieves the best mean results on five of seven language metrics, although mean pooling remains strongest on STSBenchmark. AMP-side mechanistic analyses reveal low-frequency prediction sensitivity across most encoder layers, with increased relative high-frequency sensitivity in the final layer, and pronounced peptide-specific positional responses. The residual gate broadly amplifies spectral channels while preserving the low-frequency-dominated energy profile, whereas latent cross-attention exhibits sample- and species-specific spectral allocation. Overall, FLaG provides a transferable frequency-domain aggregation bias across protein, visual, and textual representations, with benefits that depend on the backbone and downstream task. Supplementary materials, source code, and data are available at https://www.healthinformaticslab.org/supp/ and https://github.com/Kewei2023/AMPCliff/tree/FLaG.
Chinese Translation
词元聚合将词元级表示转换为固定维度的样本表示,但大多数池化方法仅在原始词元空间中进行操作。我们提出了频域潜在注意力门控池化(Frequency-Domain Latent-attention Gated Pooling, FLaG),这是一种即插即用的聚合模块,在最终池化之前将编码器输出重新表达为傅里叶域中的表示。FLaG 通过拼接实部与虚部来表示非冗余的 rFFT 频谱,利用可学习的潜在查询(latent queries)总结频谱词元,推导出基于样本条件的通道门控,并重构调制后的词元表示以用于下游聚合。我们在基于 ESM2 的抗菌肽(AMP)活性预测、ResNet18 在 CIFAR-10 和 CIFAR-100 上的图像分类,以及三个基于 RoBERTa 的语言任务上评估了同一架构。FLaG 在四个 AMP 骨干-物种设置中取得了最优的宏平均 Spearman 相关系数、RMSE 和 Recall@50,并在 CIFAR-10 上取得最高的 top-1 准确率。在七个语言任务指标中,FLaG 也在五个指标上取得了最优的平均结果,尽管平均池化在 STSBenchmark 上仍然表现最强。AMP 侧的机制分析表明,大多数编码器层对低频率预测敏感,最后一层的相对高频敏感性有所增加,并且存在显著的多肽特异性位置响应。残差门控在整体上放大频谱通道的同时,保持了以低频为主的能量分布,而潜在交叉注意力则表现出样本和物种特异性的频谱分配。总体而言,FLaG 在蛋白质、视觉和文本表示中提供了一种可迁移的频域聚合偏置,其收益取决于骨干模型和下游任务。补充材料、源代码和数据可在 https://www.healthinformaticslab.org/supp/ 和 https://github.com/Kewei2023/AMPCliff/tree/FLaG 获取。
cs.AI / 79 / 2609.00845

Towards Generalizable Visually Grounded Exploration of Household Devices

面向可泛化的视觉接地式家用设备探索
Zheng, Linhao, Liu, Zeming, Chen, Wangke, Zeng, Li, Che, Wanxiang, Huang, Heyan, Guo, Yuhang
Abstract
Recent advancements in Vision-Language Models (VLMs) have demonstrated impressive capabilities in static visual recognition and high-level semantic reasoning. However, current embodied exploration paradigms still heavily rely on imitation learning from human-annotated trajectories, which severely limits agents' generalization ability. The key bottleneck of realizing general autonomous embodied agents lies in Generalizable Visually Grounded Exploration: the ability to operate novel devices without manuals or specific training by actively grounding abstract world knowledge into fine-grained visual affordances. Yet, existing benchmarks fail to evaluate this capability: they generally rely on explicit documents and annotated trajectories, neglecting the dynamic Hypothesis-Interaction-Refinement process essential for functional device operation. To bridge this gap, we introduce VGEBench, a comprehensive benchmark designed to evaluate the generalizable visually grounded exploration capabilities of VLMs. Unlike static datasets, we construct a Logic-Driven State Machine framework. This framework simulates multi-turn interaction loops, compelling agents to achieve goals by active visual perception and feedback-driven correction. Experimental results demonstrate that existing VLMs face significant challenges in translating semantic knowledge into physical execution and maintaining long-horizon state tracking.
Chinese Translation
视觉语言模型(VLMs)的最新进展在静态视觉识别和高层语义推理方面展现出令人瞩目的能力。然而,当前的具身探索范式仍然严重依赖对人工标注轨迹的模仿学习,这极大地限制了智能体的泛化能力。实现通用自主具身智能体的关键瓶颈在于可泛化的视觉接地探索(Generalizable Visually Grounded Exploration):即在不依赖说明书或特定训练的情况下,通过将抽象的世界知识主动接地到细粒度的视觉可供性(affordance)中,从而操作新设备的能力。然而,现有基准无法评估这一能力:它们通常依赖显式文档和标注轨迹,忽视了设备功能操作所必需的动态“假设-交互-修正”(Hypothesis-Interaction-Refinement)过程。为弥补这一空白,我们提出了VGEBench,一个旨在评估视觉语言模型可泛化视觉接地探索能力的综合基准。与静态数据集不同,我们构建了一个逻辑驱动的状态机(Logic-Driven State Machine)框架。该框架模拟多轮交互循环,促使智能体通过主动的视觉感知和基于反馈的修正来实现目标。实验结果表明,现有视觉语言模型在将语义知识转化为物理执行以及维持长时程状态跟踪方面面临显著挑战。
cs.AI / 80 / 2609.00858

Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources

可验证的灾害故事线与因果知识图谱:一种基于引用的异构人道主义数据源处理流水线
Decostanzi, Ivan, Ronco, Michele, Consoli, Sergio, Corbane, Christina, Bertolini, Lorenzo, Biazzo, Indaco, Mihaila, Daria, Garcia-Herranz, Manuel, Schwebel, Felix, Mejova, Yelena, Kalimeri, Kyriaki
Abstract
Effective humanitarian response depends on the rapid synthesis of heterogeneous, high-volume information sources - a task that routinely exceeds human analytical capacity in the critical early hours of a crisis. We present a pipeline that combines structured disaster records from EM-DAT with unstructured documents from ReliefWeb and the European Media Monitor (EMM) to produce source-grounded disaster storylines and causal knowledge graphs supporting situational awareness for responders and analysts. Using Retrieval-Augmented Generation, the pipeline extracts structured storylines - tabular event profiles covering 17 fields, from severity and key drivers to child-sensitive impact indicators - and constructs causal knowledge graphs where each node and edge is enriched with citation-grounded explanatory narratives, enabling full traceability back to primary sources. We evaluate the system on three diverse crisis use cases through a human evaluation involving 9 domain expert and 9 non-expert evaluators. Results confirm high retrieval precision, strong faithfulness of extracted causal relations, and a clear expert preference for citation-grounded components over ungrounded alternatives. The pipeline is designed to scale to the full EM-DAT catalogue, with the goal of publicly releasing a narrative-enriched version of the database.
Chinese Translation
有效的人道主义响应依赖于对异构、海量信息源的快速综合——在危机爆发的关键早期时段,这一任务通常超出人类分析能力的极限。我们提出了一种流水线,将来自EM-DAT的结构化灾害记录与来自ReliefWeb和欧洲媒体监测器(European Media Monitor, EMM)的非结构化文档相结合,生成有数据源支撑的灾害故事线和因果知识图谱,以支持响应人员和分析人员的态势感知。该流水线利用检索增强生成(Retrieval-Augmented Generation)技术,提取结构化故事线——包含17个字段的表格化事件档案,涵盖灾害严重程度、关键驱动因素以及对儿童敏感的影响指标等——并构建因果知识图谱,其中每个节点和边都附有基于引用的解释性叙述,从而实现向原始数据源的完整可追溯性。我们通过一项由9名领域专家评估者和9名非专家评估者参与的人工评估,在三个不同的危机应用场景中对系统进行了评估。结果证实了系统具有较高的检索精度、提取出的因果关系具有良好的忠实性,并且专家明显更偏好基于引用的组件而非无引用支撑的替代方案。该流水线的设计可扩展至EM-DAT的完整目录,目标是公开发布该数据库的叙述增强版本。
cs.AI / 81 / 2609.00859

Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems

强化学习增强的大语言模型智能体求解复杂车辆路径问题
Chen, Yi, Yu, Zikang, Wang, Jiahai, Chen, Jinbiao, Zhou, Jianpeng, Zhang, Zizhen
Abstract
Vehicle Routing Problems (VRPs) are fundamental combinatorial optimization problems with widespread applications in various scenarios. The advanced optimization solvers can effectively solve such problems. However, modeling complex VRP variants for solvers often requires substantial domain expertise, which limits the accessibility of advanced optimization technologies. In this paper, we propose Reinforcement Learning Enhanced LLMAgents(RLEA), a multi-agent framework designed to automate the modeling of complex VRPs. RLEA introduces a lightweight neural Planner trained with Soft Q-learning to efficiently orchestrate the actions of LLM-based agents. In addition, we equip the system with an evolutionary memory module and retrieval-augmented generation, enabling the agent to leverage both accumulated experience and external solver knowledge during program generation and refinement for solving VRPs. We evaluated 48 distinct VRP variants across various solvers. The experimental results demonstrate that RLEA outperforms the previous state-of-the-ar method, achieving a 16.67% higher success rate while significantly reducing runtime errors. These results validate that integrating reinforcement learning with LLM-based reasoning is highly effective for automated optimization modeling. The appendix is available at: https://doi.org/10.5281/zenodo.19134435.
Chinese Translation
车辆路径问题是基本的组合优化问题,在各种场景中有着广泛的应用。先进的优化求解器能够有效地求解此类问题。然而,为求解器建模复杂的车辆路径问题变体通常需要大量的领域专业知识,这限制了先进优化技术的可及性。本文提出了强化学习增强的大语言模型智能体,这是一个旨在实现复杂车辆路径问题自动化建模的多智能体框架。RLEA引入了一个基于Soft Q-learning训练的轻量级神经规划器,以高效地编排基于大语言模型的智能体的动作。此外,我们为该系统配备了进化记忆模块和检索增强生成机制,使智能体在生成和优化求解车辆路径问题的程序时,能够同时利用积累的经验和外部求解器知识。我们在多种求解器上评估了48种不同的车辆路径问题变体。实验结果表明,RLEA优于此前最先进的方法,成功率提高了16.67%,同时显著减少了运行时错误。这些结果验证了将强化学习与基于大语言模型的推理相结合对于自动化优化建模是非常有效的。附录可参见:https://doi.org/10.5281/zenodo.19134435。
cs.AI / 82 / 2609.00874

Beyond the Clock: Measuring the Value of Adaptive Revision

超越时钟:度量自适应修订的价值
Chadha, Ayushi
Abstract
As agentic systems become compound systems, increasingly important decisions move above task execution itself: when should a higher-level controller preserve the strategy guiding another process, and when should it revise it? We study this meta-level control problem in a hierarchical latent reasoner whose manager can retain or replace a commitment governing lower-level computation. Across three precommitted training seeds, learned revision timing produces qualitatively different policies, ranging from an almost deterministic early clock to substantially more state conditioned schedule distributions, yet none outperforms the best forced timing policy evaluated on the same frozen checkpoint. This separates state dependence from decision value: a controller can vary its actions with internal state without turning that variation into a reproducible task-performance benefit. A deeper intervention study on the original checkpoint shows that timing itself is consequential and order-sensitive, while exhaustive enumeration reveals that a strong fixed schedule captures most of the measurable value available from timing at this decision budget. Counterfactual PERSIST/REPLAN diagnostics further show why score-level evidence can be misleading when predictability is dominated by decision position rather than within-position discrimination. Together, these results argue that learned meta-level control should be evaluated along three separate axes: whether its score depends on state, whether that dependence changes realized behavior, and whether those changes capture outcome value beyond a strong non-adaptive policy.
Chinese Translation
随着智能体系统日益成为复合系统,越来越重要的决策已超越任务执行本身:高层控制器何时应保留指导另一个进程的策略,何时应修改它?我们在一个分层潜在推理器中研究这一元级控制问题,其管理器可以保留或替换支配低层计算的承诺。在三个预先设定的训练种子上,学习到的修订时机产生了性质上不同的策略,从近乎确定性的早期时钟到显著更依赖状态的调度分布,但没有任何一个优于在同一冻结检查点上评估的最佳强制时机策略。这将状态依赖性与决策价值分离开来:控制器可以随内部状态变化其动作,而不将这种变化转化为可复现的任务性能收益。对原始检查点的更深入干预研究表明,时机本身是重要的且对顺序敏感,而穷举枚举则揭示,在该决策预算下,一个强固定调度即可捕获时机所能带来的大部分可测量价值。反事实 PERSIST/REPLAN 诊断进一步表明,当可预测性主要由决策位置而非位置内判别能力主导时,分数层面的证据为何可能产生误导。综上,这些结果论证,学习到的元级控制应沿三个独立的维度进行评估:其分数是否依赖状态,这种依赖是否改变实际行为,以及这些变化是否超越强非自适应策略而捕获结果价值。
cs.AI / 83 / 2609.00875

FractalNet-Based Heterogeneous Federated Learning for Orbital Edge Intelligence in Satellite Mega-Constellations: A Wildfire Case Study

基于FractalNet的面向卫星巨型星座轨道边缘智能的异构联邦学习:以野火监测为例
Puppala, Sai, Sinha, Koushik
Abstract
Satellite mega-constellations are emerging as large-scale sensing, communication, and computation fabrics, yet their learning architectures remain largely inherited from terrestrial federated learning and ground-centric mission operations--- ill-suited to satellites that differ by orders of magnitude in Size, Weight, Power, and Cost (SWAP-C), radiation tolerance, link availability, and propagation delay. We propose a heterogeneous federated learning method based on the FractalNet architecture for orbital edge intelligence. We formalize contact-window-constrained, depth-heterogeneous federated optimization and introduce a distributed path scheduler that assigns model depth as a function of SWAP-C constraints, predicted inter-satellite contacts, and training statistics. To reduce message overhead and energy consumption, each tier pools updates periodically rather than at every contact opportunity, and a three-tier agentic control plane governs in-space scheduling, anomaly escalation, and policy-governed autonomy. As a case study, we apply the framework to wildfire detection, where each orbital shell naturally learns a different semantic level of situational awareness: pixel-scale thermal anomalies at low Earth orbit (LEO), regional fire-front dynamics at medium Earth orbit (MEO), and larger-scale risk propagation at geostationary or high Earth orbit (GEO/HEO). Experiments on simulated mega-constellations validate the approach across convergence, communication efficiency, energy adaptation, scheduled-pooling savings, robustness, and latency.
Chinese Translation
卫星巨型星座正日益成为大规模的感知、通信与计算基础设施,然而其学习架构大多沿袭自地面联邦学习和以地面为中心的任务运营模式——这种架构难以适应在尺寸、重量、功耗与成本(SWAP-C)、抗辐射能力、链路可用性以及传播时延等方面存在数量级差异的卫星。我们提出了一种基于FractalNet架构的面向轨道边缘智能的异构联邦学习方法。我们形式化了受建连窗口约束、深度异构的联邦优化问题,并提出了一种分布式路径调度器,根据SWAP-C约束、预测的星间建连机会和训练统计数据为模型分配深度。为降低消息开销与能耗,各层级采用周期性汇聚更新的方式,而非在每次建连机会时都进行汇聚;同时,一个三层智能体控制平面负责管理在轨调度、异常上报以及受策略约束的自主运行。作为案例研究,我们将该框架应用于野火探测,其中每个轨道壳层自然地学习不同语义层级的态势感知:低地球轨道(LEO)负责像素级热异常检测,中地球轨道(MEO)负责区域性火线动态分析,地球静止轨道或高地球轨道(GEO/HEO)负责更大尺度的风险传播预测。在仿真巨型星座上的实验从收敛性、通信效率、能耗自适应、周期汇聚节省、鲁棒性以及时延等多个方面验证了该方法的有效性。
cs.AI / 84 / 2609.00879

Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning

通过偏好优化与可解释的视觉-语言推理实现可靠的多模态灾害严重程度评估
Zhang, Yuanjun, Shaik, Fuzel Ahamed, Acharjee, Suvojit, Khalid, Fahad, Oussalah, Mourad
Abstract
Reliable disaster damage assessment requires models that provide both accurate predictions and transparent explanations. However, existing multimodal approaches are limited by scarce annotated data and insufficient evaluation of reasoning quality. This study proposes a two-stage training framework that integrates Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) within a unified data construction pipeline. From a single Human-in-the-Loop (HITL) annotation workflow, two complementary datasets are derived, namely ReasoningSet, which contains validated rationales for SFT, and PreferenceSet, which comprises paired rationales for DPO-based alignment. The framework evaluates both classification performance and explanation quality using automatic metrics, model-based scoring, and human ranking. Experimental results show that SFT improves accuracy from 73.64% to 78.29% and increases Macro-F1 by 29% compared to the baseline, while explanation quality improves by approximately 25%. Subsequent DPO alignment further enhances interpretability on the PreferenceSet. Cross-model validation on InternVL-3-8B and LLaVA-1.5-7B demonstrates the robustness and generalizability of the approach. The proposed framework improves detection of underrepresented mild damage cases, reduces high-risk misclassifications, and strengthens alignment between model reasoning and human judgment. Overall, it provides a reproducible pathway to develop reliable multimodal systems that deliver auditable, actionable disaster insights for emergency management.
Chinese Translation
可靠的灾害损失评估要求模型既能提供准确的预测,又能给出透明的解释。然而,现有的多模态方法受限于标注数据稀缺以及对推理质量评估不足。本研究提出了一种两阶段训练框架,在统一的数据构建流程中融合了监督微调(Supervised Fine-Tuning, SFT)与直接偏好优化(Direct Preference Optimization, DPO)。通过单一的人在回路(Human-in-the-Loop, HITL)标注流程,衍生出两个互补的数据集:包含经过验证的推理依据、用于SFT的ReasoningSet,以及由成对推理依据组成、用于基于DPO对齐的PreferenceSet。该框架采用自动指标、基于模型的评分和人工排序来评估分类性能与解释质量。实验结果表明,SFT将准确率从73.64%提升至78.29%,Macro-F1较基线提高29%,解释质量提升约25%。随后的DPO对齐进一步增强了模型在PreferenceSet上的可解释性。在InternVL-3-8B和LLaVA-1.5-7B上的跨模型验证证明了该方法的鲁棒性和泛化能力。所提出的框架改进了对代表性不足的轻度损毁案例的检测,减少了高风险的误分类,并强化了模型推理与人类判断之间的一致性。总体而言,该框架为开发可靠的多模态系统提供了一条可复现的路径,能够为应急管理提供可审计、可操作的灾害洞察。
cs.AI / 85 / 2609.00885

Denoising Diffusion Generative Models Secretly Calculate Attentions

去噪扩散生成模型隐式地计算注意力
Haddadi, Farzan, Monfared, Leila, Rezaii, Ebrahim, Malek-Mohammadi, Mohammadreza, Zakalvand, Pejman, Mokhtari, Narges
Abstract
Denoising diffusion models are the dominant architecture for image generation, whereas most natural language generation and modeling are primarily handled by well-known transformer architectures employing attention mechanism. Here, we show that diffusion models also inherently use an attention mechanism very similar to that of transformers. Therefore, attention emerges as a universal machine learning principle, based on a general training objective. We also show similarities in basic functional principle of auto-encoders and attention-based models. These equivalences allows us to interchange these designs based on practical requirements. As an example, we can reformulate the diffusion framework to reduce the lengthy training process and computation-intensive image generation. Using this approach, a simplified algorithm is proposed for image generation which is based on attention mechanism. Results show that the attention-based implementation achieves comparable performance with significantly less effort and computational resources.
Chinese Translation
去噪扩散模型是图像生成的主流架构,而大多数自然语言生成与建模则主要由采用注意力机制的著名Transformer架构来处理。本文表明,扩散模型也内在地使用了一种与Transformer非常相似的注意力机制。因此,注意力作为一种通用的机器学习原理,基于一般的训练目标而涌现。我们还展示了自编码器与基于注意力的模型在基本功能原理上的相似性。这些等价性使我们能够根据实际需求在这些设计之间进行互换。例如,我们可以重新构建扩散框架,以缩短冗长的训练过程并降低计算密集型的图像生成开销。基于该方法,我们提出了一种基于注意力机制的简化图像生成算法。结果表明,基于注意力的实现以显著更少的投入和计算资源达到了相当的性能。
cs.AI / 86 / 2609.00891

CacheBridge: Efficient Cross-Model KV Cache Transfer

CacheBridge:高效的跨模型KV缓存迁移
Qu, Xingyu, Lu, Siyuan, Chen, Zhiyu, Wang, Sheng, Lin, Tao
Abstract
Sharing context between LLMs in a multi-model system requires the receiving model to prefill the shared prefix because KV caches are model-specific. Recent closed-form cross-model KV transfer, hereafter Full-Head Mapping, avoids this replay by fitting a training-free affine mapper from source to target caches. However, its full-head design maps each target KV head from every source KV head in the selected layers, making transfer quality sensitive to architectural differences and causing mapper storage and application cost to grow with layer support. To this end, we introduce CacheBridge, which co-designs architecture-indexed mapper support, attention-aligned calibration, and bounded mapper construction while retaining a closed-form affine interface for online deployment. CacheBridge restricts each target head to a matched source head, weights reconstruction errors by causal attention sensitivity, and uses a fused GPU kernel to construct weighted sufficient statistics without materializing full observation tensors. Across three transfer directions, CacheBridge recovers the two Ministral 3 transfer directions where Full-Head Mapping loses substantial accuracy while preserving 99.83\% mean target retention on Qwen3. On Qwen3 $14\mathrm{B}\to32\mathrm{B}$, it reduces mapper storage by $8\times$, accelerates application by up to $3.0\times$, matches \fullhead with one tenth of the calibration data, and reduces 500-sequence construction from 92.63 to 8.63 seconds ($10.7\times$).
Chinese Translation
在多模型系统中,大语言模型(LLM)之间共享上下文时,由于KV缓存是模型特定的,接收模型需要对共享前缀进行预填充(prefill)。近期的闭式跨模型KV迁移方法(下称全头映射,Full-Head Mapping)通过拟合一个免训练的仿射映射器将源缓存映射到目标缓存,从而避免了这种重复预填充。然而,其全头设计在所选层中将每个目标KV头从所有源KV头进行映射,导致迁移质量对架构差异高度敏感,并使映射器的存储与应用开销随支持层数的增长而增加。为此,我们提出CacheBridge,通过联合设计基于架构索引的映射器支持、与注意力对齐的校准,以及有界的映射器构建方式,同时保留可用于在线部署的闭式仿射接口。CacheBridge将每个目标头限制为匹配的源头,以因果注意力敏感度对重构误差进行加权,并使用融合的GPU核函数构建加权充分统计量,而无需显式生成完整的观测张量。在三个迁移方向上,CacheBridge成功恢复了全头映射损失大量精度的两个Ministral 3迁移方向,同时在Qwen3上保持99.83%的平均目标保留率。在Qwen3 14B→32B的迁移中,它将映射器存储减少8倍,应用速度提升高达3.0倍,仅用十分之一的校准数据即达到与全头映射相当的效果,并将500条序列的构建时间从92.63秒降至8.63秒(10.7倍)。
cs.AI / 87 / 2609.00892

CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training

CARE:面向大语言模型后训练的基于对比锚点的评估准则演化方法
Li, Siyuan, Song, Xinxin, Ruinian, Chen, Fan, Jingjing, Xiao, Tingxiong, Hu, Yangen, Zeng, Ke, Suo, Jinli
Abstract
Rubric-based reinforcement learning decomposes open-ended instructions into prompt-specific, flexible rubrics, making it better suited than reinforcement learning with verifiable rewards for post-training LLMs on open-ended tasks. However, static rubrics are inevitably hacked as the policy evolves, and existing dynamic approaches introduce new problems: undirected rubric extraction, unreliable hack detection, and unbounded rubric proliferation. We propose $\textbf{CARE}$ ($\textbf{C}$ontrastive $\textbf{A}$nchor-based $\textbf{R}$ubric $\textbf{E}$volution), which grounds every rubric evolution step in a high-quality anchor response generated by a frontier model conditioned on the prompt and its rubrics. At each training step, CARE contrasts the highest-scoring rollout against the anchor, enabling two complementary mechanisms: an Adaptive branch that reactively repairs reward misspecification; and a Chase branch that proactively converts frontier-level quality gaps into sharper rubrics. Together, the two branches $\textbf{maintain discriminative accuracy in the high-reward region}$---the precise region where reward over-optimization mostly originates. Experiments on WildChecklist-9K with Qwen2.5-7B-Base and Qwen2.5-7B-Instruct show that CARE achieves state-of-the-art performance on Arena-Hard-2.0, InfoBench, and FollowBench, and is the $\textbf{only}$ method whose win rate against GPT-4.1 anchor responses shows sustained improvement throughout 300 training steps; additional results on Llama-3.1-8B-Instruct and Qwen3-8B further indicate that CARE generalizes across model families.
Chinese Translation
基于评估准则(rubric)的强化学习将开放式指令分解为针对特定提示、灵活可变的评估准则,使其比基于可验证奖励的强化学习更适合在开放式任务上对大语言模型进行后训练。然而,随着策略的演化,静态评估准则不可避免地会被“钻空子”(reward hacking),而现有的动态方法又引入了新的问题:评估准则提取缺乏方向性、作弊检测不可靠、以及评估准则无限制地增殖。我们提出了CARE(Contrastive Anchor-based Rubric Evolution,基于对比锚点的评估准则演化),其将每一次评估准则演化步骤都建立在一个高质量的锚点响应之上,该锚点响应由前沿模型根据提示及其评估准则生成。在每个训练步骤中,CARE将得分最高的采样响应与锚点进行对比,从而实现两种互补机制:自适应(Adaptive)分支,以反应式方式修复奖励的错误设定;以及追赶(Chase)分支,以前瞻性方式将前沿水平的质量差距转化为更精确的评估准则。两个分支共同作用,在高奖励区域保持判别的准确性——而这正是奖励过度优化最常发生的区域。在WildChecklist-9K数据集上使用Qwen2.5-7B-Base和Qwen2.5-7B-Instruct进行的实验表明,CARE在Arena-Hard-2.0、InfoBench和FollowBench上取得了最先进的性能,并且是唯一一种在与GPT-4.1锚点响应对比中,胜率在300个训练步骤内持续提升的方法;在Llama-3.1-8B-Instruct和Qwen3-8B上的额外实验结果进一步表明,CARE具有良好的跨模型家族泛化能力。
cs.AI / 88 / 2609.00904

In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?

上下文内神经反馈:大语言模型能否通过特权访问控制其内部表征?
Aoki, Koshiro, Takatsuki, Ryota, Minegishi, Gouki, Haruki, Yusuke, Kawahara, Daisuke
Abstract
Whether large language models (LLMs) can control their own internal representations matters for both machine metacognition and AI safety. A recent study applied neurofeedback to LLMs and claimed that they can control their internal representations. However, the reported control may rely on superficial mechanisms rather than genuine internal access because the control targets in that study are not privileged, meaning that a third party can infer them from the prompt. We redesign the neurofeedback paradigm for LLMs so that the control target satisfies the privileged access requirement, which is closer to neurofeedback experiments in human cognitive neuroscience. Under this stricter setting, the models do not demonstrate reliable control over privileged internal representations, suggesting that previously reported control cannot exclude the possibility that it relies on superficial mechanisms. Our results indicate that rigorous assessments of metacognition in LLMs require evaluation methods that demand privileged access.
Chinese Translation
大语言模型(LLMs)能否控制自身的内部表征,对机器元认知和人工智能安全都具有重要意义。近期一项研究将神经反馈方法应用于大语言模型,并声称它们能够控制自己的内部表征。然而,所报告的控制可能依赖于表面机制而非真正的内部访问,因为该研究中的控制目标并不具备特权性,即第三方可以从提示(prompt)中推断出这些目标。我们重新设计了大语言模型的神经反馈范式,使控制目标满足特权访问(privileged access)要求,从而更接近人类认知神经科学中的神经反馈实验。在这种更严格的设置下,模型未能表现出对特权内部表征的可靠控制,这表明先前报告的控制无法排除其依赖表面机制的可能性。我们的结果表明,对大语言模型元认知的严格评估需要采用要求特权访问的评估方法。
cs.AI / 89 / 2609.00918

RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation

RPCBench:面向基于大语言模型推荐中主动前提批评的基准测试
Chen, Zhongru, Wu, Yuan, Chang, Yi
Abstract
Large language models are increasingly used as interactive recommender assistants. Their evaluation should therefore go beyond plausible item recommendation and test whether they can recognize flawed recommendation requests. Existing recommender benchmarks mainly assess ranking, generation, or preference satisfaction, while existing error-detection benchmarks are usually not grounded in recommendation-specific user and candidate evidence. To address this gap, we introduce RPCBench, a benchmark for evaluating Recommender-Premise Critique: the ability to detect, diagnose, and properly handle faulty premises in natural-language recommendation requests. RPCBench contains evidence-grounded test instances from five recommendation domains and covers ten types of premise failures. Each instance provides a visible recommendation context and a corrupted user query. We further design a fine-grained evaluation framework that measures proactive detection, error localization, post-detection handling strategy, and evidence faithfulness. Through a systematic evaluation of 11 LLMs, we find that proactive detection is the main bottleneck in Recommender-Premise Critique, and models perform worst on underspecified-premise errors. We also observe that target-critical information density matters more than redundant evidence, and that longer reasoning does not monotonically improve critique quality: performance peaks at intermediate reasoning length, while overly long reasoning is accompanied by an overthinking penalty. The code is available at https://github.com/ZhongruChen/RPCBench.
Chinese Translation
大语言模型正日益被用作交互式推荐助手。因此,对它们的评估不应仅限于合理的物品推荐,还应测试其能否识别存在缺陷的推荐请求。现有的推荐系统基准主要评估排序、生成或偏好满足,而现有的错误检测基准通常未基于推荐特有的用户与候选项证据。为填补这一空白,我们提出了RPCBench,一个用于评估推荐系统前提批评能力(Recommender-Premise Critique)的基准,即检测、诊断并妥善处理自然语言推荐请求中错误前提的能力。RPCBench包含来自五个推荐领域的基于证据的测试实例,并涵盖十种前提错误类型。每个实例提供一个可见的推荐上下文和一个被篡改的用户查询。我们进一步设计了一个细粒度的评估框架,用于衡量主动检测、错误定位、检测后处理策略以及证据忠实度。通过对11个大语言模型的系统评估,我们发现主动检测是推荐系统前提批评的主要瓶颈,且模型在前提信息不充分的错误上表现最差。我们还观察到,关键目标信息的密度比冗余证据更为重要,且更长的推理并不能单调提升批评质量:性能在中等推理长度时达到峰值,而过长的推理则伴随着过度思考惩罚。代码可在 https://github.com/ZhongruChen/RPCBench 获取。
cs.AI / 90 / 2609.00921

VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences

VIBE-Bench:当用户画像不等于用户偏好时评估个性化大语言模型
Jiang, Yiwen, Deng, Yang, Fong, Stephanie, Wang, Zimu, Shen, Yaling, Feng, Wei, Yang, Hongxi, Zhao, Xiangyu, Xu, Zhongxing, Mehta, Deval, Cheng, Xuelian, Ge, Zongyuan
Abstract
Personalized Large Language Models (PLLMs) aim to tailor responses to individual users, where a central challenge is preference reasoning: inferring query-relevant preferences from user-related history. Existing benchmarks, however, largely assume that such preference can be retrieved from semantically related history. We study an underexplored but practically important regime, profile-preference conceptual misalignment (PRCM), where observable profile cues and query-specific preferences lie in different concept spaces, making semantic retrieval inconsistent for personalization. We introduce VIBE-Bench, a benchmark with two psychology-grounded tasks, 3,504 personas and 12,239 dialogues, including a manually verified gold test set, and requires cross-concept preference reasoning beyond surface semantic overlap. Experiments with several personalization methods show that current PLLMs largely rely on shallow semantic correlations and fail to acquire robust cross-concept mappings. These findings establish PRCM as a distinct failure regime in PLLMs and position VIBE-Bench as a focused testbed for advancing preference reasoning beyond semantic matching.
Chinese Translation
个性化大语言模型(PLLM)旨在为个体用户定制回复,其核心挑战在于偏好推理:从用户相关历史中推断与查询相关的偏好。然而,现有基准大多假设此类偏好可以从语义相关的历史中检索得到。我们研究了一个尚未被充分探索但具有重要实际意义的情形——画像-偏好概念错位(profile-preference conceptual misalignment, PRCM),即可观测的画像线索与特定查询的偏好处于不同的概念空间中,使得语义检索无法可靠地支持个性化。我们提出了 VIBE-Bench,这是一个包含两个基于心理学的任务、3,504 个人设(persona)和 12,239 段对话的基准,其中包括经人工验证的黄金测试集,并要求进行超越表面语义重叠的跨概念偏好推理。针对多种个性化方法的实验表明,当前的 PLLM 在很大程度上依赖浅层语义关联,无法习得稳健的跨概念映射。这些发现将 PRCM 确立为 PLLM 中一种独特的失效情形,并使 VIBE-Bench 成为推动偏好推理超越语义匹配的聚焦测试平台。
cs.AI / 91 / 2609.00961

Few-Shot Out of Domain Intent Detection with Covariance Corrected Mahalanobis Distance

基于协方差校正马氏距离的小样本域外意图检测
Talur, Jayasimha, Smirnov, Oleg, Missault, Paul
Abstract
Conversational agents like chatbots and voice assistants are trained to understand and respond to user intents. On encountering an utterance with an intent different from the ones they have been trained on, these agents are expected to classify the intent as `unknown' or `out of domain'. This problem is known as out of domain (OOD) intent detection. Podolskiy et al. (2021), showed that Mahalanobis distance can be used effectively for identifying OOD intents, outperforming competing approaches. However, their method fails to outperform the baselines in the practically important few-shot setting. In this paper we analyze the reason for low performance and propose a covariance corrected Mahalanobis distance for detecting out-of-domain intents.
Chinese Translation
诸如聊天机器人和语音助手等对话式智能体经过训练,能够理解并回应用户意图。当遇到包含其未曾训练过的意图的话语时,这些智能体应将该意图分类为'未知'或'域外'。这一问题被称为域外(OOD)意图检测。Podolskiy等人(2021)证明马氏距离(Mahalanobis distance)可以有效用于识别域外意图,其表现优于竞争方法。然而,在实际中非常重要的小样本(few-shot)场景下,他们的方法未能超越基线方法。本文分析了性能较低的原因,并提出了一种用于检测域外意图的协方差校正马氏距离方法。
cs.AI / 92 / 2609.00967

CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins

CoBRA:通过反事实边际学习工具使用边界
Zou, Wenhao, Liu, Xianglong, Bi, Wendong, Wang, Hanjie, Zhao, Simin, Zhi, Gong
Abstract
As large language models increasingly act through external tools, deciding when to call a tool has become a central problem alongside deciding how to use it. Unnecessary tool calls introduce latency, cost, retrieval noise, and error propagation, while missed calls hurt knowledge-intensive queries or questions requiring up-to-date evidence. Existing methods typically trigger tools from absolute query or generation signals, such as difficulty, confidence, or final task reward, and therefore lack an explicit estimate of the instance-level marginal benefit of tool use. We propose CoBRA, a counterfactual boundary-learning framework for tool-augmented language models. CoBRA first constructs internal and external experts from the same base model, collects paired trajectories, and estimates the reward margin between answering with and without tools. This margin partitions data into internal-favored, external-favored, and ambiguous cases. CoBRA then uses clear-margin samples for Boundary-Aware Cold-Start SFT, followed by MARS-RL with reference-split rollouts and counterfactual marginal advantages to optimize boundary decisions. Experiments with retrieval as the main tool on Qwen3-4B show that CoBRA improves tool-use efficiency and boundary-sensitive answer accuracy while maintaining strong performance on tool-dependent out-of-distribution questions.
Chinese Translation
随着大语言模型越来越多地通过外部工具进行操作,决定何时调用工具已成为与决定如何使用工具同等重要的核心问题。不必要的工具调用会引入延迟、成本、检索噪声和错误传播,而错失调用则会损害知识密集型查询或需要最新证据的问题的回答。现有方法通常基于绝对的查询或生成信号(如难度、置信度或最终任务奖励)来触发工具,因而缺乏对工具使用在实例级边际收益的显式估计。我们提出CoBRA,一个面向工具增强语言模型的反事实边界学习框架。CoBRA首先从同一基础模型构建内部专家和外部专家,收集成对的轨迹,并估计使用工具与不使用工具回答之间的奖励边际。该边际将数据划分为内部偏好、外部偏好和模糊三类情形。随后,CoBRA使用清晰边际的样本进行边界感知冷启动监督微调(Boundary-Aware Cold-Start SFT),再通过MARS-RL(采用参考划分的 rollout 和反事实边际优势)优化边界决策。以检索为主要工具、基于Qwen3-4B的实验表明,CoBRA在提升工具使用效率和边界敏感的答案准确率的同时,在依赖工具的分布外问题上仍保持强劲性能。
cs.AI / 93 / 2609.01006

Figures as Programs: Recursive Generation of Editable Scientific Figures

图形即程序:可编辑科学插图(Figure)的递归生成
Liu, Yepeng, Dai, Dasen, Liu, Chengzhi, Song, Yiren, Ci, Hai, Zhang, Yu, Zhang, Qi, Shou, Mike Zheng, Wang, Xin Eric, Bu, Yuheng
Abstract
Scientific methodology figures are essential for communicating complex methods clearly, yet creating them remains labor-intensive and typically requires multiple rounds of refinement. Recent image-generation models can synthesize visually appealing raster figures, but producing a human-satisfactory result in a single generation step remains difficult. Moreover, precise edits to raster figures are challenging for both humans and models. We formulate scientific figure generation as recursive SVG program construction and propose \textsc{FigTree}, a \textit{multi-agent} system that automatically transforms a scientific paper into a structured vector figure. \textsc{FigTree} grounds figure content in the source paper, decomposes a figure into a hierarchy of local regions, generates each region as a short SVG program, and assembles the resulting fragments. A render-critic refinement loop jointly inspects the rendered figure and its underlying program, enabling visual defects to be traced to specific statements and accurately repaired. We conduct extensive evaluations of \textsc{FigTree} on figure quality and editability, showing that \textsc{FigTree} produces high-quality figures, while also enabling more effective editing than existing raster-based methods.
Chinese Translation
科学方法论插图对于清晰地传达复杂方法至关重要,但制作这些插图仍然耗时费力,通常需要多轮修改完善。近期的图像生成模型可以合成视觉上吸引人的栅格图,但仅通过单次生成获得令人满意的结果仍然困难。此外,对栅格图进行精确编辑对人类和模型来说都极具挑战性。我们将科学插图生成问题形式化为递归的SVG程序构建,并提出了FigTree——一个能够自动将科学论文转化为结构化矢量插图的多智能体(multi-agent)系统。FigTree将插图内容锚定于源论文,将插图分解为局部区域的层级结构,为每个区域生成简短的SVG程序,并将生成的片段组装起来。渲染-评论(render-critic)迭代优化循环同时对渲染后的插图及其底层程序进行检查,使视觉缺陷能够被追溯至具体语句并得到精确修复。我们在插图质量和可编辑性方面对FigTree进行了广泛评估,结果表明FigTree能够生成高质量插图,同时比现有基于栅格的方法实现更有效的编辑。
cs.AI / 94 / 2609.01035

Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees

自由派生,谨慎行动:递归LLM智能体树的渐进式风险授予机制
Wang, Molly
Abstract
Recursive LLM agents can broaden their search by spawning specialists. Some branches later request tools that send data or deploy code. When should a branch receive authority to act? We distinguish sandbox spawning, in which external controls prevent the specified harm, from capability activation, in which a selected branch crosses an irreversible-action boundary. Progressive Risk Vesting (PRV) holds a trajectory-level risk budget in escrow and debits it as branches are activated. We prove an anytime harm bound for adaptively generated trees. Branch outcomes may be dependent, but each local certificate needs to remain valid conditional on the full pre-activation history, including the information used to select the request. When activation gates, branch charges, and compute constraints are held fixed, delayed vesting preserves every policy available under irrevocable spawn charging. Marginal risk estimates can still fail after branch selection. In a stylized branching model, trajectory harm changes as the authority reproduction number $\mathcal{R}_A$ crosses one. As local risk $p$ approaches zero, trajectory harm is proportional to $p$ below criticality, proportional to $\sqrt{p}$ at criticality, and retains a positive floor above it. A finite-type occupancy model yields risk and compute shadow prices. For nested fanout modes with decreasing marginal value per unit risk, these prices produce a threshold rule. Branching calculations and a split-sample experiment illustrate the results. These synthetic studies do not estimate safety in deployed agents. The analysis suggests a design rule: search broadly in the sandbox and grant recursive authority sparingly, with an explicit risk charge.
Chinese Translation
递归LLM智能体可以通过派生子智能体来扩展其搜索范围。某些分支随后可能会请求发送数据或部署代码的工具。那么,一个分支应在何时获得行动权限?我们区分了沙箱派生(sandbox spawning)——即外部控制机制能够防止特定危害——与能力激活(capability activation)——即被选中的分支跨越不可逆行动边界。渐进式风险授予(Progressive Risk Vesting, PRV)将一个轨迹级风险预算托管保存,并在分支被激活时逐步扣减。我们证明了针对自适应生成树的任意时刻(anytime)危害上界。分支结果可能存在相关性,但每个局部证书必须在完整的激活前历史条件下保持有效,包括用于选择请求的信息。当激活门控、分支收费和计算约束固定时,延迟授予可保留不可撤销派生收费下所有可用策略。在分支选择之后,边际风险估计仍可能失效。在一个程式化的分支模型中,轨迹危害随授权再生数 $\mathcal{R}_A$ 跨越1而发生变化:当局部风险 $p$ 趋近于零时,轨迹危害在临界状态下与 $p$ 成正比,在临界点与 $\sqrt{p}$ 成正比,在超临界状态下则保持一个正的下限。一个有限类型占用模型给出了风险与计算的影子价格。对于每单位风险边际价值递减的嵌套扇出模式,这些价格产生了一个阈值规则。分支计算和一个分割样本实验对结果进行了说明。这些合成研究并不能估计已部署智能体的安全性。该分析提示了一条设计准则:在沙箱中广泛搜索,谨慎授予递归权限,并附带明确的风险收费。
cs.AI / 95 / 2609.01038

Data-Driven Persona-Conditioned Agents for A/B Test Simulation

面向A/B测试模拟的数据驱动的用户画像条件智能体
Benomar, Ziyad, Łajewska, Weronika, Perelli, Leonardo, Mansour, Saab
Abstract
A/B testing is the gold standard for evaluating product changes, but each experiment requires real user traffic, engineering effort, and weeks of measurement. We propose a simulation framework that predicts A/B test outcomes using LLM-powered agents conditioned on data-driven personas grounded in real user behavioral signals. Unlike prior work that relies on synthetic or rule-based personas, our agents are constructed from anonymized behavioral data-activity patterns, engagement signals, and inferred demographics-enabling more faithful population modeling. We frame A/B test simulation as a structured question task and systematically study (i) question design formats, (ii) the impact of persona data source and domain alignment, (iii) the trade-off between per-persona behavioral depth and population diversity, and (iv) efficient population subsampling. On a benchmark of 40 A/B tests spanning two metric types, our best configuration achieves 0.75-0.90 directional accuracy depending on the test metric, demonstrating that data-driven personas are a viable path toward fast, low-cost experiment pre-screening.
Chinese Translation
A/B测试是评估产品变更的黄金标准,但每次实验都需要真实的用户流量、工程投入以及数周的测量时间。我们提出了一个模拟框架,利用基于大语言模型(LLM)的智能体来预测A/B测试的结果,这些智能体以源自真实用户行为信号的数据驱动用户画像(persona)为条件。与以往依赖合成画像或基于规则画像的研究不同,我们的智能体基于匿名化行为数据构建——包括活动模式、参与度信号和推断的人口统计特征——从而实现更忠实的人群建模。我们将A/B测试模拟形式化为一个结构化问答任务,并系统研究了:(i) 问题设计格式;(ii) 画像数据来源及领域对齐的影响;(iii) 单个画像的行为深度与人群多样性之间的权衡;(iv) 高效的人群子采样。在一个涵盖两种指标类型、包含40个A/B测试的基准上,我们的最佳配置根据测试指标的不同实现了0.75-0.90的方向性准确率,表明数据驱动的用户画像是一条通往快速、低成本实验预筛的可行路径。
cs.AI / 96 / 2609.01045

AgentFactory: Towards Automated Agentic System Design and Optimization

AgentFactory:迈向智能体系统的自动化设计与优化
Zhang, Enci, Wang, Haofeng, Zhu, Yuesheng, Cui, Xiaole, Luo, Guibo
Abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities as powerful components in agentic systems, enabling sophisticated reasoning and complex task execution. However, current approaches to manually designing and optimizing agentic systems heavily rely on manual effort, limiting their adaptability and scalability. Recent work has explored the automated optimization of workflow designs. However, these approaches often overlook the crucial role of model capabilities and focus on single performance metrics, failing to address real-world deployment constraints. In this paper, we present AgentFactory, a framework that jointly optimizes both foundation models and workflow structures in agentic systems while considering multiple objectives including performance, cost, and efficiency. AgentFactory leverages advanced LLMs as optimizers to navigate the vast search space of possible configurations, employing a three-stage optimization pipeline to automatically discover effective combinations of fine-tuned models and optimized workflows. Through an iterative optimization process, our framework systematically explores and evaluates different agentic system designs, adapting to task-specific requirements while maintaining operational efficiency. We evaluate AgentFactory across eight benchmarks spanning five domains, including general reasoning, coding, mathematics, medicine, and finance. Our experiments demonstrate that AgentFactory consistently outperforms both manually designed methods and existing automated approaches, achieving an average improvement of 9.1% across all benchmarks, with particularly significant gains in domain-specific tasks (19.6% on MedQA and 18.7% on FinEval). These results establish AgentFactory as a promising approach for developing more capable and efficient agentic systems through automated optimization.
Chinese Translation
大语言模型(LLM)作为智能体系统中的强大组件,已展现出卓越的能力,能够实现复杂的推理与任务执行。然而,目前手动设计和优化智能体系统的方法严重依赖人工,限制了其适应性和可扩展性。近期的研究已开始探索工作流设计的自动化优化,但这些方法往往忽视模型能力的关键作用,且仅关注单一性能指标,无法应对真实部署中的约束条件。本文提出AgentFactory,一个在综合考虑性能、成本和效率等多重目标的前提下,对智能体系统中的基础模型与工作流结构进行联合优化的框架。AgentFactory利用先进的大语言模型作为优化器,在庞大的配置搜索空间中进行探索,并采用三阶段优化流程,自动发现微调模型与优化工作流的有效组合。通过迭代优化过程,本框架系统地探索和评估不同的智能体系统设计方案,在保持运行效率的同时适应特定任务的需求。我们在涵盖通用推理、编程、数学、医学和金融五个领域的八个基准上对AgentFactory进行了评估。实验结果表明,AgentFactory始终优于人工设计的方法和现有的自动化方法,在所有基准上平均提升9.1%,在领域特定任务上的提升尤为显著(MedQA上提升19.6%,FinEval上提升18.7%)。这些结果确立了AgentFactory作为通过自动化优化构建更强大、更高效智能体系统的有前景的方法。
cs.AI / 97 / 2609.01049

QILP-0: Constructing Observational Declarative Twins of Quantum Circuits

QILP-0:构建量子电路的观测性声明式孪生体
Echeandía, Marina de la Cruz, Alonso, César Luis, Ribeiro, Tony, de la Puente, Alfonso Ortega
Abstract
This paper introduces QXymb, a general framework for constructing observational declarative twins of quantum circuits, and develops QILP-0, its first complete order-0 specialization. QILP-0 constructs a finite multi-valued propositional logic program from observed circuit behaviour within a declared observational scope. The pipeline traverses a declared family of quantum observables incrementally according to a reproducible structural grading and a declared observational reference horizon. Progress is quantified through reference-relative coverage against a fixed target-independent reference. Observable responses are organized through target-independent geometry, while retained latent structure is mapped deterministically back to original observable columns before symbolic processing, preserving observational semantics and provenance. Selected observable profiles are converted into a finite relation through admissible target-independent discretization. The target is used only afterwards to audit twin-admissibility and induce the declarative theory. A theory is certified as an exact observational declarative twin when it completely and correctly reconstructs the resulting finite task-conditioned discrete relation. Logical exactness is therefore separated from numerical, backend, provider, and discretization uncertainty, which is retained as audit metadata. Validation uses two complementary QML settings. Exhaustive Bars & Stripes experiments compare product and grid-CZ embeddings from 16 to 100 qubits and exercise the native-discrete branch. Low-Depth MNIST analyses all 14,708 digit-0/1 instances before and after a trained variational quantum transformation and exercises continuous discretization. In every reported relation, the induced QILP-0 theory achieves complete, conflict-free reconstruction with strict accuracy equal to one.
Chinese Translation
本文介绍了QXymb,一个用于构建量子电路观测性声明式孪生体的通用框架,并开发了其首个完整的零阶(order-0)特化版本QILP-0。QILP-0在声明的观测范围内,根据观测到的电路行为构建一个有限的多值命题逻辑程序。该流程依据可复现的结构分级和声明的观测参考视界,对声明的量子可观测量族进行增量式遍历。进展通过相对于固定且与目标无关的参考基准的相对覆盖率来量化。可观测量响应通过与目标无关的几何结构进行组织,而保留的潜在结构在符号处理之前被确定性地映射回原始可观测量列,从而保留观测语义与溯源信息。选定的可观测量分布通过可接受的、与目标无关的离散化方法转换为有限关系。目标仅在其后用于审计孪生体可接受性并归纳出声明式理论。当该理论能够完整且正确地重构由此产生的、以任务为条件的有限离散关系时,即被认证为精确的观测性声明式孪生体。因此,逻辑精确性得以与数值、后端、提供商以及离散化不确定性相分离,而这些不确定性被保留为审计元数据。验证采用了两个互补的QML(量子机器学习)设置:穷举式Bars & Stripes实验比较了从16到100量子比特的乘积嵌入和网格CZ嵌入,并调用了原生离散分支;低深度MNIST分析则对经过训练的变分量子变换前后的全部14,708个数字0/1实例进行处理,并调用了连续离散化。在所有报告的关系中,归纳得到的QILP-0理论均实现了完整且无冲突的重构,严格精度等于一。
cs.AI / 98 / 2609.01056

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

WorldBench:面向多语言智能体的文化根基型基准测试
Ranaldi, Leonardo, Shen, Sherrie, Kai, Jushi, Birch, Alexandra
Abstract
Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing benchmarks rarely test state preservation, performance across languages, and application to realistic, grounded scenarios. To address these concerns, we present WorldBench: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions. WorldBench comprises 1,600 tasks across seven languages and eight cultures, filtered and refined through feedback from human annotators with language- and culture-specific expertise. For evaluation, we extend metrics from previous works and introduce Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and other complementary metrics through deterministic and LLM-as-a-Judge evaluations. Our experiments show that frontier models reach only 49.2% CTS, with all models demonstrating large gaps between correctness and environment preservation. We thereby show that current agents remain brittle in multilingual, agentic scenarios, especially for long-horizon tasks and under state-preservation constraints
Chinese Translation
尽管基于大语言模型(LLM)的智能体在复杂环境中执行多步骤任务的应用日益增多,现有的基准测试很少检验状态保持能力、跨语言性能以及其在真实、接地场景中的应用。为解决这些问题,我们提出了WorldBench:一个基于真实人物画像的日常 workflows 的综合性多语言基准测试,智能体可以在沙盒环境中通过结构化动作进行操作。WorldBench 包含跨越七种语言和八种文化的 1,600 个任务,并经过具有特定语言和文化专业知识的人类标注者反馈进行筛选与优化。在评估方面,我们扩展了先前工作的指标,并引入了受约束任务成功率(Constrained Task Success, CTS),该指标结合自然语言指令与测试环境,通过确定性评估和 LLM-as-a-Judge 评估对任务完成度、最小修改度以及其他补充指标进行打分。实验结果表明,前沿模型的 CTS 仅达到 49.2%,且所有模型在正确性与环境保持能力之间存在较大差距。由此我们证明,当前的智能体在多语言、智能体化场景中仍然十分脆弱,尤其是在长程任务以及状态保持约束条件下。
cs.AI / 99 / 2609.01057

User Representation via Cross Multi-source Behavior Pre-training for Mobile Games

基于跨多源行为预训练的移动游戏用户表示
Yang, Chengqi, Qiao, Yiran, Liu, Feng, Lou, Xingyu, Zhou, Zijun, Mo, Xiaoyun, Zhang, Changwang, Xu, Jiayuan, Wang, Jun, Ao, Xiang
Abstract
User representation pre-training has become a fundamental paradigm for alleviating data sparsity in downstream personalization tasks. However, existing studies predominantly focus on single-app or app-level behaviors, overlooking the inherently cross-source and multi-granular nature of user activities on mobile devices. At the device level, user intent emerges from complex interactions among heterogeneous behavior sources and hierarchical action structures, posing challenges that cannot be addressed by conventional app-centric modeling. To tackle this issue, we propose CM-PTM, a novel Cross Multi-source Behavior Pre-Training Model tailored for mobile game user representation learning on device-level behavioral logs. CM-PTM employs hierarchical cascaded mask-then-predict proxy tasks that first infer the source of the next behavior and then progressively refine predictions at the app-action level. This design enables unified modeling of cross-source dependencies and fine-grained behavioral dynamics within a single pre-training paradigm. Extensive experiments on large-scale real-world mobile datasets demonstrate that CM-PTM effectively captures users' endogenous interests and consistently delivers significant performance gains on downstream mobile game recommendation tasks.
Chinese Translation
用户表示预训练已成为缓解下游个性化任务中数据稀疏性问题的一种基础范式。然而,现有研究主要关注单应用或应用层面的行为,忽视了移动设备上用户活动固有的跨源性与多粒度特性。在设备层面,用户意图产生于异构行为源与分层行为结构之间的复杂交互,这给传统的以应用为中心的建模方法带来了挑战。为解决这一问题,我们提出了CM-PTM,一种面向设备级行为日志的移动游戏用户表示学习的跨多源行为预训练模型。CM-PTM采用分层级联的“掩码-预测”代理任务,首先推断下一个行为的来源,随后在应用-行为层面逐步细化预测。这一设计使得在单一预训练范式内能够统一建模跨源依赖关系和细粒度行为动态。在大规模真实移动数据集上的广泛实验表明,CM-PTM能够有效捕捉用户的内生兴趣,并在下游移动游戏推荐任务上持续带来显著的性能提升。
cs.AI / 100 / 2609.01058

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

ARISE-RL:基于强化学习的智能体量规引导迭代自进化
Zhang, Fanrui, Ding, Ruixue, Zhang, Qiang, Chen, Xi, Chen, Boli, Wang, Shihang, Wang, Qiuchen, Zhan, Hongmin, Bian, Jinxin, xingchao, Li, Zheng, Peijin, cheng, Hao, Xie, Pengjun, Zhang, Kaipeng, Liu, Jiawei, Zha, Zheng-Jun
Abstract
Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brittle and unstable rewards, resulting in weak or noisy rollout contrast that obscures fine-grained optimization signals for group-based policy learning. To address these challenges, we propose ARISE-RL, a novel full-cycle self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution. The Generator grounds tool-related rubric criteria in real tool observations and is rewarded for producing valid, intermediate-difficulty tasks aligned with the Solver's evolving capability boundary. The Solver, in turn, learns from fine-grained rubric satisfaction signals through multi-step reasoning and tool use. We further introduce Reward-Gated Self-Evolution Distillation (RG-SED), which selectively distills a memory-augmented variant of the same policy back into itself only when the memory yields empirical reward improvement, thereby reducing distribution mismatch and avoiding blind imitation of noisy guidance. Finally, to support rigorous evaluation, we present ECR-Bench, an expert-calibrated rubric benchmark suite covering single-tool deep research and multi-tool travel planning. Extensive experiments demonstrate that ARISE-RL consistently achieves robust and stable overall state-of-the-art performance across all evaluated benchmarks.
Chinese Translation
由于缺乏可验证的黄金答案和可扩展的量规(rubric),通过强化学习(RL)训练开放式智能体面临诸多困难。此外,即使接近模型的能力边界,长程开放式智能体任务往往产生脆弱且不稳定的奖励,导致群体策略学习中的 rollout 对比信号微弱或充满噪声,从而掩盖了细粒度的优化信号。为应对这些挑战,我们提出了 ARISE-RL,一种新颖的全周期自进化框架,通过量规介导的协同进化,将任务/量规生成器(Generator)与推理求解器(Solver)耦合在一起。生成器将工具相关的量规标准扎根于真实的工具观测中,并因其生成与求解器不断演进的能力边界相匹配的、有效且中等难度的任务而获得奖励。求解器则通过多步推理和工具使用,从细粒度的量规满足信号中学习。我们进一步引入了奖励门控自进化蒸馏(Reward-Gated Self-Evolution Distillation, RG-SED),仅当记忆带来经验性奖励提升时,才选择性地将同一策略的记忆增强变体蒸馏回其自身,从而降低分布失配并避免对噪声引导的盲目模仿。最后,为支持严格评估,我们提出了 ECR-Bench,这是一个由专家校准的量规基准套件,涵盖单工具深度研究和多工具旅行规划任务。大量实验表明,ARISE-RL 在所有评估基准上始终取得了稳健且稳定的整体最先进性能。
cs.AI / 101 / 2609.01062

Space Generative AI with Solar Energy Harvesting

基于太阳能收集的空间生成式人工智能
Zhang, Jierui, Huang, Jianhao, Wang, Zhanwei, Huang, Kaibin
Abstract
Satellites are emerging as promising platforms to extend generative \emph{artificial intelligence} (AI) services to remote areas lacking terrestrial infrastructure. However, deploying space generative AI is fundamentally constrained by the limited, time-varying onboard energy supplied by solar \emph{energy harvesting} (EH). This paper presents a framework for solar-powered space generative AI in which a satellite receives a user prompt, executes a diffusion-based image-generation model, and downlinks the compressed result within a strict time window. We identify the fundamental \emph{computation--communication} (C$^2$) trade-offs governed by the shared harvested-energy budgets. Specifically, increasing the number of generation steps improves intrinsic image quality but depletes energy and time available for downlink transmission, whereas prioritizing communication guarantees reliable delivery but sacrifices semantic quality. To balance these trade-offs and maximize \emph{end-to-end} (E2E) generative performance, we exploit the predictable solar-EH dynamics induced by deterministic orbital motion and develop a joint C$^2$ resource-optimization framework using a tractable two-step approach. First, we characterize the maximum downlink throughput for a fixed generation depth under continuous solar EH. This establishes a separation principle that decouples waiting-time selection from optimal transmit-power control. Next, we formulate a joint C$^2$ utility-maximization problem and derive a closed-form, low-complexity step-selection policy in the dominant constant-power regime. Extensive experiments under realistic orbital dynamics demonstrate that the proposed policy dynamically balances generation quality and transmission reliability. This yields significant E2E performance gains over static computation- and communication-centric baselines across diverse solar-EH states.
Chinese Translation
卫星正成为将生成式人工智能(AI)服务扩展到缺乏地面基础设施的偏远地区的有前景平台。然而,空间生成式AI的部署从根本上受限于太阳能能量收集(EH)所提供的有限且时变的星上能量。本文提出了一种太阳能供电的空间生成式AI框架:卫星接收用户提示,执行基于扩散模型的图像生成模型,并在严格的时间窗口内将压缩后的结果下行传输。我们识别了由共享的收集能量预算所支配的基本计算—通信(C²)权衡。具体而言,增加生成步数可提升图像内在质量,但会消耗可用于下行传输的能量和时间;而优先保障通信则能确保可靠传输,却牺牲了语义质量。为平衡这些权衡并最大化端到端(E2E)生成性能,我们利用确定性轨道运动带来的可预测太阳能EH动态,并采用一种易于处理的两步方法,构建了联合C²资源优化框架。首先,我们在连续太阳能EH条件下刻画了固定生成深度下的最大下行吞吐量,由此建立了一种分离原理,将等待时间选择与最优发射功率控制解耦。其次,我们构建了联合C²效用最大化问题,并在主导的恒定功率场景下推导出闭式、低复杂度的步数选择策略。在真实轨道动态下的大量实验表明,所提策略能够动态平衡生成质量与传输可靠性,在各种太阳能EH状态下,相比以计算为中心和以通信为中心的静态基线方法,均取得了显著的E2E性能提升。
cs.AI / 102 / 2609.01117

Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs

潜在循环思维:基于冻结大语言模型推理的潜在表征循环细化方法
Chen, Zhaoliang, Fu, Jie
Abstract
Chain-of-thought reasoning unfolds in discrete token space: each step is committed as text, errors propagate, and eliciting good traces presupposes traces to imitate. Reasoning instead in a model's continuous representation space - where intermediate states are vectors rather than words - sidesteps these constraints, but leaves open how those latent states should be computed. We approach this along two axes. First, we keep a large language model (LLM) frozen and use it for what it is already good at - modeling and decoding sequences - while a small auxiliary network supplies continuous latent thoughts as input. Second, we produce those latents by recurrence: a tiny recurrent reasoner refines them over many steps, decoupling the depth of computation from the size of the model, so that the latents are a product of iterative processing rather than a single forward pass. We instantiate this as Latent Recurrent Thoughts (LRT): a task-dedicated proposer supplies base latents, a recurrent reasoner refines them through bounded residual corrections, and the frozen LLM decodes the answer. On symbolic reasoning with answer supervision but no reasoning traces (Countdown-4, Sudoku) and on natural-language reasoning (HumanEval, MBPP, StrategyQA), LRT substantially outperforms prior frozen-decoder continuous-space reasoning methods under an identical decoder, prompt, data, and training budget, and outperforms non-thinking-mode chain-of-thought prompting on the same backbone at a small fraction of its inference compute.
Chinese Translation
思维链推理在离散的词元空间中展开:每一步都以文本形式确定下来,错误会不断传播,而引出高质量的推理轨迹的前提是存在可供模仿的轨迹。若改为在模型的连续表征空间中进行推理——即中间状态是向量而非词语——则可以避开这些限制,但仍存在一个悬而未决的问题:这些潜在状态应如何计算。我们从两个维度切入这一问题。首先,我们保持大语言模型(LLM)冻结不变,仅用它完成其本就擅长的工作——建模和解码序列——同时由一个小型辅助网络提供连续的潜在思维作为输入。其次,我们通过循环递归的方式生成这些潜在表征:一个微型的循环推理器在多步迭代中对其进行细化,从而将计算深度与模型规模解耦,使潜在表征成为迭代处理的产物而非单次前向传播的结果。我们将该方法具体化为潜在循环思维(Latent Recurrent Thoughts, LRT):一个面向特定任务的提议器提供基础潜在表征,一个循环推理器通过有界的残差修正对其进行细化,最后由冻结的LLM解码出答案。在仅有答案监督而无推理轨迹的符号推理任务(Countdown-4、数独)以及自然语言推理任务(HumanEval、MBPP、StrategyQA)上,在相同的解码器、提示、数据和训练预算下,LRT大幅超越了此前基于冻结解码器的连续空间推理方法,并且在推理计算量仅为其一小部分的情况下,在同一骨干模型上超越了非思考模式的思维链提示方法。
cs.AI / 103 / 2609.01168

Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate

从裂缝中越狱文生图模型:通过多智能体辩论导航异构安全过滤器
Wen, Kaiyan, Zhang, Shijie, Yu, Lu, Bai, Guangdong
Abstract
Text-to-image (T2I) models remain vulnerable to jailbreak attacks that elicit Not-Safe-For-Work (NSFW) content, despite increasingly being guarded by heterogeneous, multi-layer safety stacks combining text filters, image classifiers, and cross-modal detectors. Existing jailbreak studies either optimize against individual filters or query the complete pipeline with aggregate feedback, making it difficult to identify the active constraint and adapt to conflicts across safety layers. In this paper, we introduce the Detection Surface, a unified geometric framework that characterizes the decision boundaries induced by heterogeneous T2I safety filters and their joint effect on the jailbreak search space. This formulation reveals that successful evasion is governed by a sparse and non-convex region shaped by cross-layer conflicts, where mutations that bypass one filter may increase exposure to another. Motivated by this analysis, we propose CRACK, a multi-agent debate framework for adaptive jailbreak search that decomposes jailbreak search into exploration, diagnosis, and arbitration. CRACK coordinates an Attack Agent, a Defense Agent, and a Judge Agent to iteratively generate prompt mutations, obtain layer-specific diagnostic feedback, and optimize mutation strategies through reward-guided refinement. Through repeated rounds of debate, CRACK adapts its search direction to the evolving cross-layer constraints while preserving the original harmful intent. Extensive experiments across multiple T2I models, datasets, and safety configurations show that CRACK achieves Attack Success Rates (ASR) of up to 99.63% under composite defenses, while requiring fewer queries than existing methods and maintaining semantic fidelity.
Chinese Translation
尽管文生图(T2I)模型日益受到由文本过滤器、图像分类器和跨模态检测器组成的异构多层安全防护体系的保护,但其仍容易受到诱导生成不适宜工作场所(NSFW)内容的越狱攻击。现有的越狱研究要么针对单个过滤器进行优化,要么通过聚合反馈查询完整的防护流程,这使得识别当前生效的约束以及适应安全层之间的冲突变得困难。本文提出了检测面(Detection Surface),一个统一的几何框架,用于刻画异构 T2I 安全过滤器所诱导的决策边界及其对越狱搜索空间的联合影响。该表述揭示,成功的规避由一个稀疏且非凸的区域主导,该区域由跨层冲突塑造,其中绕过某一过滤器的变异可能增加暴露于另一过滤器的风险。基于这一分析,我们提出了 CRACK,一个用于自适应越狱搜索的多智能体辩论框架,它将越狱搜索分解为探索、诊断和仲裁三个环节。CRACK 协调攻击智能体(Attack Agent)、防御智能体(Defense Agent)和裁判智能体(Judge Agent),迭代地生成提示词变异、获取针对特定安全层的诊断反馈,并通过奖励引导的改进来优化变异策略。通过多轮反复辩论,CRACK 能够在保持原有有害意图的同时,使搜索方向适应不断演变的跨层约束。在多个 T2I 模型、数据集和安全配置上进行的大量实验表明,CRACK 在复合防御下实现了高达 99.63% 的攻击成功率(ASR),同时所需查询次数少于现有方法,并保持了语义保真度。
cs.AI / 104 / 2609.01198

FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue

FinLifeBench:基于纵向银行对话的穷尽式生活事件历史与财务状态重建
Lee, Hangyeul, Oh, Juyoung, Ko, Jaeyong, Kim, Sunmin, Park, Jaeik, Kim, Hyunkyu, Son, Jungmin, Kang, Pilsung
Abstract
Repeated banking interactions require assistants to maintain complete, current, and traceable customer records as life changes emerge incidentally in routine requests. Existing benchmarks emphasize question answering, bounded episodes, or targeted recall rather than exhaustive longitudinal reconstruction. We introduce FinLifeBench, which evaluates two tasks over the same cumulative dialogue: reconstructing every life-event instance with its first-establishing session and reconstructing a complete 34-path financial state at consecutive checkpoints. The benchmark contains 6,000 eight-turn Korean banking sessions from 20 independent synthetic trajectories, with deterministic, exhaustive gold for 24 event types and 34 state paths and consensus quality assurance. Across eleven LLMs under a full-context condition, event-anchor recall falls from 0.591 at 15 sessions to 0.445 at 300. Errors are driven primarily by omitted events rather than poor anchor localization, while financial-state reconstruction frequently treats superseded or potentially outdated information as current; the best GCA@15 reaches 0.470. Performance on the two reconstruction tasks is only weakly associated. These results show that models can localize evidence for recovered events while still failing to maintain complete and temporally valid longitudinal records.
Chinese Translation
重复的银行业务交互要求助手随着生活变化在日常请求中偶然出现时,维护完整、最新且可追溯的客户记录。现有基准测试侧重于问答、有界情节或针对性回忆,而非穷尽式的纵向重建。我们提出FinLifeBench,它在同一累积对话上评估两项任务:重建每一个生活事件实例及其首次确立的会话,以及在连续检查点上重建完整的34条路径财务状态。该基准包含来自20条独立合成轨迹的6,000个八轮韩语银行会话,为24种事件类型和34条状态路径提供确定性、穷尽式的标准答案(gold),并采用共识式质量保证。在完整上下文条件下对十一个大语言模型的评估显示,事件锚点召回率从15个会话时的0.591下降到300个会话时的0.445。错误主要源于事件遗漏而非锚点定位不佳,同时财务状态重建经常将被取代或可能过时的信息视为当前信息;最佳GCA@15仅为0.470。两项重建任务上的表现仅有弱相关性。这些结果表明,模型能够为已恢复的事件定位证据,却仍无法维护完整且在时间上有效的纵向记录。
cs.AI / 105 / 2609.01216

H2Table: Hierarchical Hypergraph-Enhanced Large Language Models for Complex Table Reasoning

H2Table:面向复杂表格推理的层次化超图增强大语言模型
Ling, Jia, Wang, Yangfan, Tang, Chen, Tan, Haoming, Yang, Yang, Guan, Yi, Jiang, Jingchi
Abstract
Tables are ubiquitous across diverse domains, yet reasoning over them remains a significant challenge for modern large language models (LLMs). Current approaches typically linearize tables into sequences, inherently overlooking their intrinsic two-dimensional and hierarchical structure. To address this, we propose H2Table (Hierarchical Hypergraph-Enhanced Table Reasoning), a novel framework that represents complex tables as hierarchical nested hypergraphs. To process this representation, we design a tailored hypergraph encoder to facilitate message passing between hyperedges (headers) and nodes (cells), thereby perceiving the semantic entailment relationships between them within complex tables. Furthermore, we introduce a set of learnable query vectors acting as a lightweight bridge to extract representative structural embeddings from the encoder into the LLM. Experimental results demonstrate that our approach effectively handles complex table question answering tasks with hierarchical nested headers. Notably, on the HiTab dataset, H2Table achieves an average improvement of 22.88% over state-of-the-art baselines on highly complex tables with a nesting depth of four. Our code is available at: https://github.com/lila120/h2table.
Chinese Translation
表格在各个领域中普遍存在,但对表格进行推理对现代大语言模型(LLMs)而言仍然是一个重大挑战。现有方法通常将表格线性化为序列,天然忽略了表格固有的二维和层次结构。为解决这一问题,我们提出了 H2Table(Hierarchical Hypergraph-Enhanced Table Reasoning,层次化超图增强表格推理),这是一个将复杂表格表示为层次化嵌套超图的新型框架。为处理这种表示,我们设计了一个定制的超图编码器,以促进超边(表头)与节点(单元格)之间的消息传递,从而感知复杂表格中二者之间的语义蕴含关系。此外,我们引入了一组可学习的查询向量,作为轻量级桥梁,将编码器中的代表性结构嵌入提取到 LLM 中。实验结果表明,我们的方法能够有效处理具有层次化嵌套表头的复杂表格问答任务。值得注意的是,在 HiTab 数据集上,对于嵌套深度为四层的高度复杂表格,H2Table 相较于最先进的基线方法平均提升了 22.88%。我们的代码已发布于:https://github.com/lila120/h2table。
cs.AI / 106 / 2609.01217

Prompt-Robust Language Models: Which Training Strategies Work?

对提示词鲁棒的语言模型:哪些训练策略有效?
Sadrieh, Frederic, Štefánik, Michal
Abstract
Despite their strong performance, large language models remain highly sensitive to prompt formulation. Prior work addresses this through refined data construction or through dedicated robustness objectives. We reproduce and compare these strategies under controlled conditions, and measure how effective they are in addressing models' prompt sensitivity. We find the current robustness fine-tuning methods improve over standard fine-tuning and in-context learning, but the best-to-worst prompt gap remains as high as 40-57% of performance. Moreover, the recent robustness-enhancing methods we test - CoIN for contrastive alignment and PPCL for consistency regularization - often fail to outperform the simplest data construction strategy: training on one template per batch. Our diagnostics explain these results. The auxiliary objectives move the quantity they penalize, but do not generalize beyond it. Additionally, data construction strategies differ due to the conflicting signs of per-template gradients on 57-64% of parameters. Thus, batches that mix formulations force the optimizer to reconcile competing updates instead of finding a shared, prompt-agnostic one.
Chinese Translation
尽管大型语言模型性能强大,但它们对提示词的表述方式仍然高度敏感。先前的工作通过精细化的数据构建或专门的鲁棒性目标来解决这一问题。我们在受控条件下复现并比较了这些策略,并衡量它们在解决模型提示词敏感性方面的有效性。我们发现,当前的鲁棒性微调方法优于标准微调和上下文学习(in-context learning),但最优提示与最差提示之间的性能差距仍高达40-57%。此外,我们所测试的近期鲁棒性增强方法——用于对比对齐的CoIN和用于一致性正则化的PPCL——往往无法超越最简单的数据构建策略:每个批次仅使用一种模板进行训练。我们的诊断分析解释了这些结果。这些辅助目标虽然降低了其所惩罚的量,但无法泛化到该量之外。另外,由于在57-64%的参数上不同模板的梯度符号相互冲突,数据构建策略之间存在差异。因此,混合多种提示表述的批次迫使优化器去调和相互竞争的更新,而不是找到一种共享的、与提示词无关的更新。
cs.AI / 107 / 2609.01257

Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations

度量长时程人类活动模拟的行为保真度
Cheng, Yi Fei, Yang, Fan, Bas, Iremsu, Niinuma, Koichiro, Abe, Narishige, Lindlbauer, David
Abstract
As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. While prior work has examined behavioral fidelity in survey responses and dialogue, longer-horizon real-world activity remains largely unexplored. We introduce a framework for evaluating behavioral fidelity in long-horizon activity simulations across temporal granularities and levels of analysis. As a case study, we collect a 43-hour multi-camera dataset of in-the-wild office activity and compare trace-derived conditioning mechanisms: persona descriptors, few-shot exemplars, and statistical transition and time-of-day priors. We find that behavioral fidelity is not uniform across metrics: statistical priors bring activity and sequence distributions closest to real behavior, yet over-fragment routines and suppress within-person variability. These findings motivate a more holistic evaluation that spans multiple metrics, temporal granularities, and levels of analysis.
Chinese Translation
随着基于大语言模型(LLM)的人类模拟器被越来越多地用于政策制定、评估和训练,它们必须忠实地再现真实的行为模式。尽管已有研究考察了问卷调查回答和对话中的行为保真度,但更长时程的真实世界活动仍在很大程度上未被探索。我们提出了一个框架,用于在多个时间粒度和分析层次上评估长时程活动模拟中的行为保真度。作为案例研究,我们收集了一个长达43小时的多摄像头真实办公活动数据集,并比较了基于轨迹数据的条件机制:角色描述(persona descriptors)、少样本示例(few-shot exemplars),以及统计转移和时间先验。我们发现,行为保真度在不同指标上并不一致:统计先验使活动和序列分布最接近真实行为,但同时导致日常活动片段过度碎片化并抑制了个体内部的变异性。这些发现表明,需要一种跨越多个指标、时间粒度和分析层次的更全面的评估方法。
cs.AI / 108 / 2609.01260

Dual Process Motion Planning

双过程运动规划
Yan, Jiayi, Fabiano, Francesco, Abate, Alessandro
Abstract
Robotic systems are deeply embedded in both industry and everyday life, where they are expected to act with speed, precision, and reliability. Classical control and planning methods have long delivered strong guarantees, but often at the cost of computational efficiency and adaptability. More recently, learning-based approaches have shown promise in overcoming these limitations, enabling agents to leverage experience to accelerate decision-making and address previously intractable problems. In this work, we bridge these two approaches through a neuro-symbolic perspective on nonlinear motion planning. Inspired by the Thinking Fast and Slow paradigm, we introduce a dual-process architecture that combines the strengths of robust reasoning and learning. Our framework integrates state-of-the-art symbolic solvers as a ``System-2'' component with experience-driven ``System-1'' modules. A metacognitive controller dynamically orchestrates their interaction, selecting when to rely on fast intuition versus slower, more precise reasoning. By evaluating the framework across diverse nonlinear benchmark environments, we demonstrate that this architecture yields consistent gains in planning efficiency, accuracy, and generalization, while promoting reuse across tasks. The results suggest that tightly coupling learning with structured reasoning offers a scalable path toward more capable and adaptive robotic systems.
Chinese Translation
机器人系统已深度融入工业与日常生活之中,人们期望其能够以快速、精确和可靠的方式行动。经典的控制与规划方法长期以来提供了强有力的保证,但往往以牺牲计算效率和适应性为代价。近年来,基于学习的方法在克服这些局限方面展现出前景,使智能体能够利用经验加速决策,并解决此前难以处理的问题。在本工作中,我们通过神经-符号的视角将这两种方法融合到非线性运动规划中。受“快与慢思考”范式(Thinking Fast and Slow)的启发,我们提出了一种双过程架构,结合了稳健推理与学习的优势。我们的框架将最先进的符号求解器作为“系统2”组件,与由经验驱动的“系统1”模块相集成。元认知控制器动态地协调二者的交互,决定何时依赖快速直觉,何时采用更慢但更精确的推理。通过在多种非线性基准环境中对该框架进行评估,我们证明该架构在规划效率、准确性和泛化能力方面均取得了一致的提升,同时促进了跨任务的复用。结果表明,将学习与结构化推理紧密耦合,为实现更强能力与更高适应性的机器人系统提供了一条可扩展的路径。
cs.AI / 109 / 2609.01272

Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents

让前瞻记忆呈现SLM形态:面向小模型智能体的类型化意图存储
Zhao, Jinqing, Wu, Chengcan
Abstract
Prospective memory means carrying out a deferred intention at the right future cue while other work continues. Benchmarks now isolate it as an agent skill, yet frontier LLMs still struggle: the best published PM-Bench scaffold reaches only 65.1% Set-F1. We argue that this loop is schema-constrained state tracking rather than open-ended reasoning, and that small models can execute it when the action space is typed. We propose the Prospective Intention Store (PIS) that puts lifecycle logic in code and scoped language work on the model. The scaffold is agentic and training-free: no selector fine-tuning and no trajectory distillation. On PM-Bench, DeepSeek-Chat with PIS reaches 82.9% Set-F1. On Gemma-E2B, Set-F1 is only 4.2% without a store and at most 6.6% under seven retrospective memories, while PIS reaches 66.2%. PIS further reaches 70.1% Set-F1, where retrospective memory methods stay at most 54.4%. PIS sets a new state of the art on this benchmark and enables small models to surpass the published large-model scaffold.
Chinese Translation
前瞻记忆(Prospective Memory)指在其他工作持续进行的同时,在合适的未来线索出现时执行一项延迟意图。现有基准测试已将其作为一项智能体技能加以独立考察,但前沿大语言模型(LLM)仍然表现不佳:已发表的PM-Bench最佳支架(scaffold)仅达到65.1%的Set-F1。我们认为,这一循环本质上是受模式约束的状态跟踪,而非开放式推理;当动作空间被类型化之后,小模型也能够执行该任务。我们提出前瞻意图存储(Prospective Intention Store, PIS),将生命周期逻辑置于代码中,而将限定范围的语言工作交由模型完成。该支架具有智能体特性且无需训练:既不需要选择器微调,也不需要轨迹蒸馏。在PM-Bench上,配备PIS的DeepSeek-Chat达到了82.9%的Set-F1。在Gemma-E2B上,不使用存储时Set-F1仅为4.2%,在七种回顾性记忆(retrospective memory)方法下最高也仅为6.6%,而PIS达到了66.2%。PIS进一步达到70.1%的Set-F1,而回顾性记忆方法最高仅为54.4%。PIS在该基准上创造了新的最优水平(state of the art),并使小模型超越了已发表的大模型支架。
cs.AI / 110 / 2609.01286

Analog-DB: An Agent-First Analog Integrated Circuit Database, From Blocks to Systems

Analog-DB:一个面向智能体(Agent-First)的模拟集成电路数据库——从电路模块到系统
Zadeh, Danial Noori, Elamien, Mohamed B.
Abstract
Sharing analog integrated circuit designs remains difficult: foundry non-disclosure agreements restrict the process details a design depends on, and the testbenches behind published results are rarely released. We present analog-db, an open-source, versioned database built on a shareable design representation. A domain-specific language captures each design as a process-neutral topology, reusable testbenches, and a machine-readable datasheet under one schema, so a design is shared in full and re-simulates on the process kits it is bound to. A parameterization scheme exposes functional sub-blocks and device sizes as named parameters that carry their matching constraints, making circuits composable and retargetable; a schema-governed contract and queryable catalog let AI design agents discover and reuse them directly. Across the regulator corpus, all 23 circuit-kit bindings on three open kits meet their own recorded specification bands (typical corner, matched devices, no layout) and 10 of 23 meet a common class band. Seventeen of the 23 imported sizings failed their testbenches and closed under a gm/ID sizing loop driven by the annotated sub-block roles, typically within one to three iterations. In a supervised case study, a coding agent working from the released artifacts sized the op-amp cores of a chopper instrumentation amplifier on an open 130nm kit, locating four hand-entry defects and a missing common-mode feedback loop that the sizing-only baseline did not repair. The database holds 68 circuits across sixteen classes, verifiable at schematic level under a tiered harness and tracked on a power/performance scoreboard, released at https://github.com/MacAnalog/spicexplorer-release.
Chinese Translation
模拟集成电路设计的共享一直十分困难:代工厂的保密协议限制了设计所依赖的工艺细节,而已发表成果背后的测试平台(testbench)也很少被公开。我们提出了 analog-db,一个基于可共享设计表示构建的开源、带版本管理的数据库。该数据库通过一种领域特定语言(DSL),以统一的模式(schema)将每个设计捕获为与工艺无关的拓扑结构、可复用的测试平台以及机器可读的数据手册(datasheet),从而使设计得以完整共享,并可在其所绑定的工艺套件(PDK)上重新仿真。其参数化方案将功能性子模块与器件尺寸暴露为携带匹配约束的命名参数,使电路可组合、可重定向(retarget);由模式(schema)约束的契约和可查询的目录使 AI 设计智能体能够直接发现并复用这些设计。在整个稳压器语料库中,三个开源工艺套件上的全部 23 个电路—套件绑定均满足其自身记录的规格区间(典型工艺角、器件匹配、不含版图),其中 23 个中有 10 个满足统一的类别区间。在 23 个导入的尺寸设计中,有 17 个未能通过其测试平台的验证,随后在由标注的子模块角色驱动的 gm/ID 尺寸设计闭环下收敛,通常仅需一到三次迭代。在一个有监督的案例研究中,一个基于所发布工件进行工作的编码智能体在开源 130nm 工艺套件上对斩波仪表放大器的运算放大器内核进行尺寸设计,定位出了四个人工录入缺陷和一个缺失的共模反馈环路,而仅做尺寸设计的基线方法未能修复这些问题。该数据库包含十六个类别共 68 个电路,可在分层测试框架下于原理图层级进行验证,并通过功耗/性能记分板进行跟踪,已发布于 https://github.com/MacAnalog/spicexplorer-release。
cs.AI / 111 / 2609.01315

A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation

一个用于可复现的全模态基础模型评估的可组合评估系统
Lee, Hodong, Park, Sanghee, Ryu, Dohoon, Kim, Jungwhan, Kim, Junyeob, Kim, Soyoon, Kim, Geewook
Abstract
Building an omni-modal foundation model means evaluating it across text, image, video, and audio. Excellent evaluation toolkits exist for each modality, but their inference engines, prompt conventions, and metric implementations are mutually incompatible, so practitioners end up maintaining separate environments for every toolchain and still struggle to compare results across them. OmniEvaluator grew out of this need in our own model development: rather than reimplementing benchmarks, it connects existing inference engines and curated evaluation libraries at a higher level, exposing four inference backends, four evaluation frameworks, and over a thousand benchmarks through a single interface. Every run is recorded as an artifact capturing the full configuration for exact reproduction, and results flow into a shared dashboard for cross-model comparison. A federated mode shares GPU inference servers across concurrent evaluations, and a built-in verifier, small enough to run on CPU, keeps its score stable across engines and prompts where rule-based scoring fluctuates under configuration mismatch, matching cost-efficient commercial LLM judges without their recurring API cost. The system, demo video, and dashboard are publicly available. (https://github.com/naver-ai/omni-evaluator)
Chinese Translation
构建全模态基础模型意味着需要在文本、图像、视频和音频等多种模态上对其进行评估。目前每种模态都已有优秀的评估工具包,但它们的推理引擎、提示词约定和指标实现彼此不兼容,导致使用者不得不为每个工具链维护独立的环境,并且仍难以跨工具链比较结果。OmniEvaluator 正是源于我们自身模型开发中的这一需求:它并不重新实现基准测试,而是在更高层级上连接现有的推理引擎和精心整理的评估库,通过单一接口提供了四种推理后端、四个评估框架以及一千多个基准测试。每次运行都会被记录为一个包含完整配置的工件,以支持精确复现,结果还会汇入一个共享的仪表板,用于跨模型比较。该系统支持联邦模式,可在多个并发评估之间共享 GPU 推理服务器;此外,它还内置了一个足够轻量、可在 CPU 上运行的验证器(verifier),在基于规则的评分因配置不匹配而产生波动的情况下,仍能保持跨引擎和跨提示词的评分稳定性,其效果可媲美成本高昂的商业大语言模型评委(LLM judge),且无需承担其持续的 API 费用。系统、演示视频和仪表板均已公开可用。(https://github.com/naver-ai/omni-evaluator)
cs.AI / 112 / 2609.01320

Automated Event Log Generation from Unstructured Text Using Finetuned LLMs

基于微调大语言模型从非结构化文本自动生成事件日志
Seeth, Maximilian, Tavares, Gabriel Marques, Schuster, Daniel
Abstract
Process mining (PM) provides a powerful framework for discovering and optimizing operational processes from event data. However, the efficacy of PM techniques is strictly predicated on the availability of structured event logs. Thus far, event logs have often been laboriously created by domain and process mining experts. This costly effort causes large portions of organizational knowledge, including incident tickets, manuals, and textual reports, to remain underutilized. We address this bottleneck by investigating the efficacy of Large Language Models (LLMs) as automated data translators. We propose a scalable framework that leverages LLMs as data translators to bridge the gap between unstructured textual resources and structured event data. We finetune LLMs on a newly created text-to-log dataset, demonstrating that the resulting models can extract high-fidelity event logs from unstructured resources. Our results show that this finetuning approach outperforms few-shot or zero-shot prompting by a large amount, highlighting finetuning as a necessary pre-condition for generating reliable event data. We conclude that our method provides a promising pipeline for making previously unused data available to process mining ecosystems, effectively expanding the possibilities of using PM to further investigate organizational workflows.
Chinese Translation
流程挖掘(Process Mining, PM)为从事件数据中发现和优化运营流程提供了强大的框架。然而,流程挖掘技术的有效性严格依赖于结构化事件日志的可用性。迄今为止,事件日志通常由领域专家和流程挖掘专家费力地手工创建。这种高昂的成本导致大量组织知识——包括事件工单、手册和文本报告——仍未得到充分利用。为解决这一瓶颈,我们研究了将大语言模型(Large Language Models, LLMs)用作自动化数据转换器的有效性。我们提出了一个可扩展的框架,利用大语言模型作为数据转换器,以弥合非结构化文本资源与结构化事件数据之间的鸿沟。我们在一个新构建的文本到日志(text-to-log)数据集上对大语言模型进行微调,结果表明所得模型能够从非结构化资源中提取高保真的事件日志。我们的结果显示,这种微调方法大幅优于少样本(few-shot)或零样本(zero-shot)提示方法,凸显了微调作为生成可靠事件数据的必要前提条件。我们的结论是,该方法为将此前未被利用的数据提供给流程挖掘生态系统提供了一条有前景的流水线,有效拓展了使用流程挖掘进一步研究组织工作流的可能性。
cs.AI / 113 / 2609.01337

LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting

LEAP:基于大语言模型的概率预测中的似然引出与聚合方法
Chen, Yufei, Zhao, Yiran, Xu, Xiaogang, Xie, Qipeng, Wu, Jiafei, Liu, Zhe
Abstract
LLM-based forecasting systems have improved on real-world tasks such as financial markets and sports outcomes, largely through stronger search and tool use. Many systems still ask an LLM to read all collected evidence together and produce the final forecast. We call this design Monolithic Prediction. It can obscure how individual evidence items affect the result and collapse uncertainty across competing outcomes. We propose LEAP (Likelihood Elicitation and Aggregation for Probabilistic forecasting), which reorganizes how collected evidence is used in the prediction stage. LEAP examines each evidence item separately and elicits likelihood parameters that describe its implications for the target. An explicit prior and a deterministic probabilistic model then combine these likelihoods into a posterior distribution. This procedure supports continuous, single-choice, and multi-choice forecasts while preserving reproducible evidence contributions. We build a benchmark covering forecasting, information-seeking, and browsing tasks, and evaluate LEAP on our own agent loop and several agent CLI frameworks. Given the same evidence, LEAP improves most prediction and calibration metrics across models and remains stronger under controlled comparisons of prior access, inference budget, and aggregation.
Chinese Translation
基于大语言模型(LLM)的预测系统在金融市场和体育赛事结果等现实任务上已取得进展,这在很大程度上得益于更强的搜索和工具使用能力。许多系统仍然让大语言模型将收集到的所有证据一起阅读并给出最终预测。我们将这种设计称为整体式预测(Monolithic Prediction)。这种设计可能掩盖单个证据项对结果的影响,并使相互竞争的结果之间的不确定性被压缩。我们提出LEAP(Likelihood Elicitation and Aggregation for Probabilistic forecasting,面向概率预测的似然引出与聚合),它重新组织了预测阶段中已收集证据的使用方式。LEAP逐一单独考察每个证据项,并引出描述其对目标问题影响的似然参数。随后,通过一个显式先验和确定性的概率模型,将这些似然组合成后验分布。该流程支持连续型、单选和多选预测,同时保留了可复现的证据贡献。我们构建了一个涵盖预测、信息检索和浏览任务的基准,并在我们自己的智能体循环及多个智能体CLI框架上对LEAP进行了评估。在给定相同证据的情况下,LEAP在不同模型上提升了大多数预测和校准指标,并在先验访问、推理预算和聚合方式的受控对比下依然保持优势。
cs.AI / 114 / 2609.01345

Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades

廉价的验证器,巨大的盲区:度量成本节约级联的可靠性代价
Rajput, Dushyant, Chauhan, Nirdesh, Kosaraju, Siddharth
Abstract
Inference cascades cut cost by answering most queries with a cheap model and escalating a hard tail to a frontier model that acts as verifier. A natural extension closes the loop: fine-tune the cheap student on the verifier's rejections so the escalation rate, and cost, fall each round. We measure this loop on real LLMs and report four findings. First, the verifier's blind spot, the fraction of the student's wrong answers it accepts, is large and moves adversarially: it grows with student capability ($\beta$ from 0.12 to 0.55 as the student scales 0.5B to 32B) and shrinks with verifier capability, so it is worst in the cheap-student, cheap-verifier regime cascades exist to create. Second, buying it away returns the saving: a frontier verifier drives $\beta$ to about 0.05 but then escalates on 46% of hard-MATH queries against a 39% true error rate, paying the frontier price on nearly half of all traffic. Third, naive corrective fine-tuning on the verifier-rejected tail does not improve the small student but degrades and ultimately collapses it, across every teacher we tried (cross-family and same-family), so at this scale the self-improving loop is self-defeating. Fourth, through all of this the cascade's own dashboard, every metric computed through the verifier, reads a flat 3% error while true delivered error swings up to 32%: the system is blind to its own degradation by construction. We then give the theory that explains the blindness, a two-population conservation law, $\epsilon_\infty \lesssim q_0 \beta_0$, under which every in-loop metric improves while true quality does not, and a synthetic study that validates the mechanism. The practical conclusion: the reliability of a self-improving cascade cannot be read from any metric computed through its own verifier.
Chinese Translation
推理级联通过让廉价模型回答大多数查询,并将困难的尾部查询升级给作为验证器的前沿模型,从而降低成本。一个自然的扩展方案是闭合这一循环:在验证器拒绝的样本上微调廉价学生模型,使升级率和成本逐轮下降。我们在真实的大语言模型(LLM)上测量了这一循环,并报告四项发现。第一,验证器的盲区——即它接受的学生错误答案的比例——很大且呈现对抗性变化:它随学生能力的增强而增长(当学生从0.5B扩展到32B时,β 从0.12升至0.55),并随验证器能力的增强而缩小,因此在廉价学生-廉价验证器的组合下最为严重,而这正是级联机制所催生的场景。第二,通过购买更强验证器来消除盲区会让节省的成本得而复失:一个前沿验证器可将 β 降至约0.05,但在 hard-MATH 查询上对39%的真实错误率却升级了46%的查询,等于对近一半流量支付了前沿模型的价格。第三,在验证器拒绝的尾部样本上进行朴素纠正性微调并不能提升小型学生模型,反而使其性能退化并最终崩溃——在我们尝试的所有教师模型(跨家族和同家族)中均是如此,因此在这一规模下,自我改进循环是自我挫败的。第四,在整个过程中,级联自身的监控面板(所有通过验证器计算的指标)始终显示平稳的3%错误率,而真实交付错误率却波动高达32%:该系统在构造上对自身的退化是盲目的。随后我们给出解释这种盲目性的理论——一个双种群守恒律 ε_∞ ≲ q_0 β_0,在此规律下所有循环内指标改善而真实质量并未提升——并通过一项合成实验验证了该机制。实践结论是:自我改进级联的可靠性无法从任何通过其自身验证器计算的指标中读出。
cs.AI / 115 / 2609.01353

SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding

SymFold:协同进化先验与结构先验实现精确的蛋白质逆折叠
Wang, Handong, Qi, Jiaxin, Lai, Baisheng, Huang, Jianqiang
Abstract
Protein inverse folding aims to recover amino acid sequences for a given 3D protein structure, underpinning broad applications such as enzyme engineering and drug discovery.Current methods often follow a serial pipeline, in which a structure encoder predicts a coarse sequence, which is then refined by protein language models (PLMs). However, because PLMs only perform post-hoc sequence edits, the refinement is bounded by the quality of upstream predictions.Thanks to recent multimodal protein language models (MPLMs), we could directly encode structure to generate sequences with pretrained structural knowledge, but we observe that they are not effective for inverse folding. Therefore, we introduce a symmetric dual-path architecture that both leverages PLMs for pretrained sequence evolution knowledge and MPLMs for pretrained structural knowledge to iteratively guide protein sequence generation.Through extensive experiments across standard protein inverse folding benchmarks, our method achieves state-of-the-art performance, surpassing prior approaches, and ablation studies validate the rationale of our symmetric design, revealing a promising direction for the community.
Chinese Translation
蛋白质逆折叠旨在为给定的三维蛋白质结构恢复相应的氨基酸序列,在酶工程和药物发现等领域具有广泛的应用。当前方法通常遵循串行流程:先由结构编码器预测一个粗略序列,再通过蛋白质语言模型(PLM)对其进行优化。然而,由于PLM仅对序列进行事后编辑,优化效果受限于上游预测的质量。得益于近期的多模态蛋白质语言模型(MPLM),我们可以直接对结构进行编码,利用预训练的结构知识来生成序列,但我们发现它们在逆折叠任务中并不有效。因此,我们提出了一种对称双路径架构,同时利用PLM中预训练的序列进化知识和MPLM中预训练的结构知识,迭代地引导蛋白质序列的生成。通过在标准蛋白质逆折叠基准上的大量实验,我们的方法达到了最先进的性能,超越了此前的方法;消融实验验证了我们对称设计的合理性,为该领域指明了一个有前景的方向。
cs.AI / 116 / 2609.01360

EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems

EDGE:基于错误依赖图的多智能体大语言模型系统中多错误归因方法
Hou, Jun, Pitre, Priya, Fang, Yi, Wang, Xuan
Abstract
Large language model (LLM) agent failures often contain multiple related errors rather than a single mistake. Existing attribution methods usually identify a responsible agent, step, or root cause, but do not explicitly model dependency between errors. We introduce EDGE, an Error Dependency Graph-guided multi-Error attribution framework. EDGE constructs an error dependency graph from observed error events and validates a reliable causal subset through counterfactual rollout. The inference graph guides a two-stage LLM-as-judge detector for error attribution, and the intervention-validated subgraph provides a more reliable basis for explanation and repair analysis. Experiments on TRAIL and MAST show that EDGE improves category-level multi-error attribution across most evaluated models and settings. Experiments with adapted Who&When-style prompts show that the graph helps across prompting strategies. These results suggest that dependency structure is a useful diagnostic prior for agent failures beyond isolated root-cause prediction.
Chinese Translation
大语言模型(LLM)智能体的失败往往包含多个相互关联的错误,而非单一错误。现有的归因方法通常只识别出责任智能体、步骤或根本原因,但没有显式地对错误之间的依赖关系进行建模。我们提出了 EDGE,一种错误依赖图引导的多错误归因框架。EDGE 根据观测到的错误事件构建错误依赖图,并通过反事实推演验证一个可靠的因果子集。该推理图引导一个两阶段的大语言模型作为评判者(LLM-as-judge)检测器进行错误归因,而经干预验证的子图为解释与修复分析提供了更可靠的基础。在 TRAIL 和 MAST 上的实验表明,EDGE 在大多数被评估的模型和设置下提升了类别级的多错误归因性能。采用改编的 Who&When 风格提示的实验表明,该依赖图在不同的提示策略下均能带来帮助。这些结果表明,依赖结构是一种超越孤立根本原因预测的有用的智能体失败诊断先验。
cs.AI / 117 / 2609.01408

Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations

神经符号几何抽象(NeuSOGA):从观测到符号化数学表示
Li, Qingde, Hong, Qingqi, Li, Zihan, Tian, Jie
Abstract
A fundamental challenge in artificial intelligence is the transformation of observations into explicit symbolic representations suitable for abstraction, interpretation, and reasoning. While modern AI systems achieve remarkable perceptual capabilities through large-scale statistical learning, the resulting knowledge is typically encoded within latent parameters that are difficult to inspect or manipulate analytically. Inspired by Neuro-Symbolic AI and theories of human abstraction, this paper investigates the formation of symbolic mathematical representations from geometric observations. We propose NeuSOGA (Neuro-Symbolic Geometric Abstraction), a framework that progressively transforms observations into topological abstractions, geometric abstractions, and ultimately symbolic mathematical representations. The architecture combines topology-guided structural discovery using Euclidean Distance Transforms, foundation-model perception using Segment Anything, adaptive multi-scale geometric abstraction, and symbolic synthesis through Implicit Area Splines. The resulting representation is an analytical implicit model supporting arbitrary-order smoothness, additive composition, and closed-form evaluation. Unlike neural latent encodings, the generated representation remains interpretable, editable, and mathematically explicit. Experiments on ModelNet40 point clouds, arbitrary-view projections, and segmented optical observations demonstrate that NeuSOGA transforms diverse observations into compact symbolic representations while preserving essential geometric and topological structure across sensing modalities and viewing directions. NeuSOGA provides an interpretable and explainable pathway from observation to symbol and establishes
Chinese Translation
人工智能领域的一个根本性挑战是将观测转化为显式的符号表示,以支持抽象、解释和推理。尽管现代人工智能系统通过大规模统计学习获得了卓越的感知能力,但所得到的知识通常编码在难以检查或解析操作的潜在参数中。受神经符号人工智能(Neuro-Symbolic AI)和人类抽象理论的启发,本文研究了从几何观测中形成符号化数学表示的问题。我们提出了 NeuSOGA(Neuro-Symbolic Geometric Abstraction,神经符号几何抽象),这是一个将观测逐步转化为拓扑抽象、几何抽象,并最终转化为符号化数学表示的框架。该架构结合了基于欧几里得距离变换(Euclidean Distance Transform)的拓扑引导结构发现、基于 Segment Anything 基础模型的感知、自适应多尺度几何抽象,以及通过隐式面积样条(Implicit Area Splines)实现的符号综合。所得表示是一个解析隐式模型,支持任意阶光滑性、可加组合和闭式求值。与神经潜在编码不同,所生成的表示具有可解释性、可编辑性和数学上的显式性。在 ModelNet40 点云、任意视角投影以及分割后的光学观测上的实验表明,NeuSOGA 能够将多样的观测转化为紧凑的符号表示,同时在不同传感模态和观察方向下保持关键的几何与拓扑结构。NeuSOGA 提供了一条从观测到符号的可解释、可说明的路径,并建立了
cs.AI / 118 / 2609.01409

EdiTikZ: Scientific Figure Editing from Revision Trajectories

EdiTikZ:基于修订轨迹的科学图表编辑
Greisinger, Christian, Zhao, Zhixue, Eger, Steffen
Abstract
Vision-language models (VLMs) have shown strong performance in generating scientific figures from text or images. However, producing publication-ready figures requires iterative refinement, making scientific figure editing an important yet largely unexplored task. Existing approaches rely on costly proprietary agentic systems, focus primarily on evaluation, or construct training supervision from synthetically generated edits. Instead, we leverage naturally occurring scientific revision and development trajectories as a scalable source of supervision. To this end, we introduce DaEdiTikZ, the first large-scale dataset of revision-derived scientific figure edits, constructed by mining 391K plausible TikZ edit pairs from arXiv, GitHub, and TeX SE and inferring 781K directed edit instructions with a VLM conditioned on rendered figures and TikZ code. We further introduce DaEdiTikZ-Bench, a human-refined benchmark with 790 instances, and train two compact Qwen3.5-based EdiTikZ models (4B and 9B) by jointly learning reconstruction and editing, followed by reinforcement learning (RL) with complementary rewards for rendered fidelity and edit application. Automatic evaluation places our 9B model above all tested baselines, while human evaluation with 9 annotators and 4,320 ratings places it above GPT-5.6-Sol and on par with Gemini-3.1-Pro. Under severe out-of-distribution shifts, it remains competitive with GPT-5.6-Sol near its 2K training sequence-length regime. Models and datasets will be released.
Chinese Translation
视觉语言模型(VLM)在从文本或图像生成科学图表方面已展现出强大的性能。然而,要生成可用于出版的图表需要反复迭代修改,这使得科学图表编辑成为一项重要却几乎未被探索的任务。现有方法依赖于昂贵的专有智能体(agentic)系统,或主要聚焦于评估,或利用合成生成的编辑来构建训练监督。与此不同,我们利用自然产生的科学修订和开发轨迹作为可扩展的监督来源。为此,我们提出了DaEdiTikZ,这是首个大规模的、源于修订的科学图表编辑数据集,其通过从arXiv、GitHub和TeX SE中挖掘39.1万条合理的TikZ编辑对,并使用VLM在渲染图和TikZ代码条件下推断出78.1万条有向编辑指令构建而成。我们进一步引入了包含790个实例、经人工精修的基准数据集DaEdiTikZ-Bench,并通过联合学习重建与编辑任务训练了两个基于Qwen3.5的紧凑型EdiTikZ模型(4B和9B),随后采用强化学习(RL)进行训练,其奖励函数兼顾渲染保真度和编辑应用两方面。自动评估显示,我们的9B模型优于所有被测基线模型;由9名标注者、4320次评分组成的人类评估显示,其表现优于GPT-5.6-Sol,并与Gemini-3.1-Pro相当。在严重的分布外偏移下,在其约2K训练序列长度的范围内,该模型仍能与GPT-5.6-Sol保持竞争力。相关模型和数据集将被发布。
cs.AI / 119 / 2609.01466

Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers

解析执行流:面向长时程智能体及其观察者的实时轨迹模型
Pakhomov, Egor, Nijkamp, Erik
Abstract
A long-horizon agent's trace outgrows both of its consumers: the human observer monitoring the run, and the agent itself, whose bounded context the trace must be folded back into. We present a live trace model, an append-only event ledger folded incrementally into typed run state and compiled into per-consumer views, and evaluate it for both consumers against deterministic ground truth. For the observer side, evaluated with an LLM reader as proxy, the compiled view answers monitoring questions using approximately 14x and 15x fewer input tokens (by reader) and at 5-7x lower cost than a budget-capped single-call reading of the raw trace, with higher accuracy (0.85-0.87 versus 0.48). Because the questions were co-designed with the view schema, we treat the token and cost reduction, conditional on schema coverage, as the transferable result. For the agent, on 120-link sequential-dependency tasks, mechanisms that maintain the task's running statistic in per-step state succeed where full-context prompting fails (30/30 versus 8/30 under a clean protocol, n=30, labeled descriptive owing to benchmark-system co-development); a prompt-level scratchpad matches the fold's accuracy at lower cost, and a two-arm decomposition attributes the fold's accuracy to its deterministic aggregate and its cost advantage to its compactness. The fold's remaining value over cheaper alternatives is deterministic auditability and serving the observer from the same state. We derive eleven candidate requirements for trace folding from observed failures and delimit them with an order-sensitive task family on which the fold ceases to help. Code, benchmarks, a regenerable synthetic corpus, and all workbench traces are released.
Chinese Translation
长时程智能体的执行轨迹会超出其两类消费者的承受能力:一是监控运行过程的人类观察者,二是智能体自身——轨迹必须被折叠回其有限的上下文中。我们提出了一种实时轨迹模型(live trace model),即一个仅追加的事件账本,通过增量折叠形成类型化的运行状态,并编译为面向各消费者的视图,同时针对确定性基准真值对两类消费者进行评估。在观察者一侧(以LLM阅读器作为代理进行评估),编译视图回答监控问题时,相较于对原始轨迹进行预算受限的单次调用阅读,输入词元(按阅读器计)约减少14倍和15倍,成本降低5至7倍,且准确率更高(0.85-0.87 对比 0.48)。由于这些问题是与视图模式(schema)共同设计的,我们将词元和成本的降低(以模式覆盖为前提条件)视为可迁移的结论。在智能体一侧,在120环节的顺序依赖任务上,通过逐步状态维护任务运行统计量的机制能够成功,而全上下文提示则失败(在干净协议下为30/30对比8/30,n=30,由于基准系统与方法的共同开发,结果标记为描述性);提示级的草稿板(scratchpad)以更低成本达到与折叠机制相同的准确率,双臂分解实验将折叠机制的准确率归因于其确定性聚合,将其成本优势归因于其紧凑性。相较于更廉价的替代方案,折叠机制的剩余价值在于其确定性的可审计性,以及能够从同一状态同时服务观察者。我们从观察到的失败中推导出轨迹折叠的十一条候选需求,并通过一个顺序敏感的任务族划定其失效边界——在该任务族上折叠机制不再有效。我们公开了代码、基准测试、可再生的合成语料库以及所有工作台轨迹。
cs.AI / 120 / 2609.01481

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

Harness之Harness:具备持续改进能力的多天自主软件开发
Yan, Haoyang, Su, Min-le, Zhang, Hangfan, Li, Zhanhao, Zhang, Chen, Zhang, Shao, Chen, Yang, Bai, Lei, Hu, Shuyue
Abstract
This paper studies autonomous software development, in which LLM-based coding agents transform high-level requirements into complete, functional, and usable software systems without human intervention. We introduce Harness-of-Harness (HoH), a framework that enables coding agents to continually improve software during autonomous development. HoH operates on existing coding-agent harnesses, and organizes their executions into iterative planning-coding-testing loops. To sustain improvement across loops, HoH balances repair with capability growth, scopes development into small and verifiable increments, separates implementation-time testing from independent evaluation, and constrains verifiable outputs rather than prescribing agent workflows. It progressively exposes deliverables, role-specific tools, and skills, encourages reuse rather than recreation, and maintains versioned project histories. On GameCraft-Bench, FrontierSWE, and ProgramBench, three harness-model pairs (Codex with GPT-5.5, OpenCode with DeepSeek-V4-Pro, and Pi with MiniMax-M3), HoH consistently outperforms the corresponding standalone harnesses, achieving an average relative gain of 52.25 percent and a maximum gain of 82.86 percent after three iterations. In a multi-day deployment with more than 70 iterations, HoH autonomously develops a first-person-shooter game, featuring a coherent storyline, fully implemented core mechanics, human-playable experience, polished visuals and integrated audio. Github: https://github.com/Flesymeb/HarnessOfHarness Project Page: https://flesymeb.github.io/HarnessOfHarness/
Chinese Translation
本文研究自主软件开发问题,即基于大语言模型(LLM)的编码智能体在无需人工干预的情况下,将高层次需求转化为完整、可用且功能完善的软件系统。我们提出了Harness-of-Harness(HoH),一个使编码智能体能够在自主开发过程中持续改进软件的框架。HoH运行于现有的编码智能体框架(harness)之上,并将其执行过程组织为迭代式的“规划—编码—测试”循环。为了在多个循环中保持持续改进,HoH在缺陷修复与能力增长之间取得平衡,将开发划分为小而可验证的增量,将实现时的测试与独立评估相分离,并通过约束可验证的输出而非规定智能体的工作流程来进行管控。它逐步开放交付物、面向特定角色的工具和技能,鼓励复用而非重复造轮子,并维护带版本管理的项目历史。在GameCraft-Bench、FrontierSWE和ProgramBench三个基准上,针对三组框架—模型组合(Codex与GPT-5.5、OpenCode与DeepSeek-V4-Pro、Pi与MiniMax-M3),HoH始终优于对应的独立框架,经过三次迭代后平均相对提升52.25%,最大提升达82.86%。在一次超过70轮迭代的多天部署中,HoH自主开发了一款第一人称射击游戏,具备连贯的故事情节、完整实现的核心机制、可供人类游玩的体验、精良的视觉效果以及集成的音频。Github:https://github.com/Flesymeb/HarnessOfHarness 项目主页:https://flesymeb.github.io/HarnessOfHarness/
cs.AI / 121 / 2609.01519

When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation

当护栏看似有效:LLM智能体商务评估中的构念效度失效
Zhu, Peiying, Chang, Sidi
Abstract
Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. We audit this risk in a multi-turn buyer--seller testbed for configurable hotel transactions. An initial implementation reported welfare gains from two marketplace guardrails of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B--14B ladder. It also gave guarded and unguarded agents different offer schemas and choice procedures. Holding the schema and buyer chooser fixed changes the paired contrasts to +7.2, -13.9, and +23.8. The four largest 14B single-generation effects averaged +229; after three generations per profile-condition, they averaged +37.6 (95% bootstrap interval [-34.2, 109.3]), while generation residuals account for 49.9% of variation in this post-hoc probe. A seller-incentive check is non-monotone: increasing profit pressure produces less profit than the default seller prompt. Scripted positive controls show why this matters. A profit-maximizing seller already attains first-best welfare, so guardrails mostly redistribute and reduce welfare; they create welfare only when the seller is explicitly programmed to force inefficient bundles. We contribute a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting, and returning INVALID or INCONCLUSIVE before substantive policy claims. In our case, the original estimate is INVALID under protocol isolation, while the controlled study remains INCONCLUSIVE under incentive validity and stochastic stability. The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks.
Chinese Translation
交互式仿真日益被用于评估由语言模型(LLM)智能体构成的市场中的策略。其输出可能呈现出经济特征——价格、利润、消费者剩余和福利——却并未真正实例化其所宣称的行为。我们在一个用于可配置酒店交易的多轮买卖双方测试平台中审查了这一风险。初始实现报告称,在Qwen2.5 1.5B至14B的模型梯度上,两项市场护栏带来了+87.4、+35.0和+28.8的福利提升。该实现还为受护栏保护与未受护栏保护的智能体提供了不同的报价模式和选择程序。在固定报价模式和买方选择器后,配对对比结果变为+7.2、-13.9和+23.8。14B单次生成中四个最大的效应平均为+229;而在每个配置条件下进行三次生成后,其均值降至+37.6(95%自助法置信区间为[-34.2, 109.3]),且在此事后探查中,生成残差解释了49.9%的变异。卖方激励检验结果呈非单调性:增大利润压力所产生的利润反而低于默认卖方提示词。脚本化的正向对照说明了这一问题的重要性。利润最大化的卖方本身已能达到一阶最优福利,因此护栏主要是重新分配并降低福利;只有当卖方被显式编程以强制执行低效交易组合时,护栏才能创造福利。我们贡献了一个构念效度契约,将激励效度、协议隔离、随机稳定性和福利核算相互分离,并在得出实质性政策结论之前返回INVALID(无效)或INCONCLUSIVE(不确定)的判定。在我们的案例中,原始估计在协议隔离检验下为INVALID,而受控研究在激励效度和随机稳定性检验下仍为INCONCLUSIVE。本案例并未表明护栏是无效的,而是表明在模拟智能体和协议通过这些检验之前,护栏的表面价值是无法被识别的。
cs.AI / 122 / 2609.01526

EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation

EvoSCM:通过因果模型演化与实验实现科学信念修正
Zhao, Qing, Li, Haowei, Deng, Weijian, Wei, Pengxu, Lin, Liang
Abstract
Scientific agents must learn not only how to reason, but also what to believe. However, existing LLM agents typically express scientific hypotheses in free-form text, leaving their beliefs implicit and difficult to test or revise. We introduce EvoSCM, which equips scientific agents with explicit structural causal models that evolve as new experimental evidence is collected. EvoSCM maintains a population of competing SCM hypotheses, each encoding a candidate causal explanation of the environment, and evolves them through a closed discovery loop. In each round, the agent abduces latent mechanisms from accumulated evidence, designs discriminative interventions, and commits to falsifiable predictions that it tests through experimentation. Discrepancies between prediction and observation are inductively distilled into correction rules that revise the causal structures and mechanisms of each hypothesis, and the agent then deductively validates the revised population against accumulated evidence and structural consistency to guide the next round. We evaluate EvoSCM on DiscoverPhysics, a benchmark requiring agents to uncover the hidden dynamics of noncanonical physical worlds through experimentation. EvoSCM consistently improves scientific discovery over baselines, yielding more accurate explanations and predictions while making more effective use of experimental interactions.
Chinese Translation
科学智能体不仅需要学习如何推理,还需要学习应当相信什么。然而,现有的LLM智能体通常以自由文本形式表达科学假设,使其信念隐式化,难以检验和修正。我们提出了EvoSCM,它为科学智能体配备显式的结构因果模型(SCM),并随着新实验证据的收集不断演化。EvoSCM维护一组相互竞争的SCM假设种群,每个假设都编码了对环境的一种候选因果解释,并通过一个闭环发现循环对其进行演化。在每一轮中,智能体从积累的证据中溯因推断潜在机制,设计具有判别力的干预实验,并给出可证伪的预测,随后通过实验加以检验。预测与观察之间的偏差被归纳提炼为修正规则,用于修正每个假设的因果结构和机制;随后智能体再根据积累的证据和结构一致性对修正后的假设种群进行演绎验证,以指导下一轮循环。我们在DiscoverPhysics上评估了EvoSCM,该基准要求智能体通过实验揭示非经典物理世界的隐藏动力学。EvoSCM持续超越基线方法,在科学发现方面取得更准确的解释和预测,同时更有效地利用实验交互。
cs.AI / 123 / 2609.01552

Can LLMs Discover Scientific Laws in Real and Parallel Worlds?

大语言模型能否在真实世界与平行世界中发现科学定律?
Huang, Yiming, Liu, Ziche, Wu, Zhuohang, Wang, Yiqian, Cui, Junxia, Zou, Xinkai, Mao, Linjun, Huang, Nan, Yu, Naicheng, Zhu, Kaijie, Ma, Yue, Zhou, Kun, Peng, Letian, Shang, Jingbo
Abstract
Scientific equation discovery has long been central to scientific progress, proceeding through iterative cycles of hypothesis generation, observational testing, and refinement under scientific constraints. As LLM capabilities advance and their role in AI for Science expands, it remains an open problem whether they can genuinely discover scientific laws and how this ability should be evaluated. Existing evaluations, however, often either simplify discovery through synthetic settings or reuse published targets that may already be familiar to LLMs. We therefore introduce SCILAWS-BENCH, a benchmark for scientific law discovery built from published research and real scientific data. It comprises 118 problems drawn from 381 scientific papers, covering 291 candidate laws and roughly 8M real data points across six scientific disciplines. Each problem is instantiated in two complementary settings: (1) SCILAWS-REAL asks models to propose laws from fixed real observations and evaluates held-out predictive fit and scientific validity derived from the source literature, and (2) SCILAWS-PARALLEL asks models to actively query residual-calibrated worlds and recover synthesized hidden laws derived from published forms. This two-setting task design preserves each problem's scientific context while separately evaluating fixed-record law discovery and active recovery of a newly synthesized hidden law. We find that predictive fit can diverge from scientific validity, memorization shapes whether models reproduce or move beyond published formulas, and our best-of-N study reveals a selection bottleneck. Our work provides a paper-grounded benchmark and new empirical perspectives for evaluating AI for scientific discovery. Project page: https://yiyihum.github.io/SciLaws-Bench
Chinese Translation
科学方程的发现长期以来一直是科学进步的核心,其过程遵循在科学约束下假设生成、观测检验与迭代修正的循环。随着大语言模型(LLM)能力的提升及其在AI for Science中作用的扩展,它们能否真正发现科学定律以及应如何评估这种能力,仍是一个开放性问题。然而,现有评估往往要么通过合成设置简化了发现过程,要么复用LLM可能已经熟悉的已发表目标。为此,我们提出了SCILAWS-BENCH,一个基于已发表研究和真实科学数据构建的科学定律发现基准。该基准包含来自381篇科学论文的118个问题,涵盖291个候选定律以及六个科学学科中约800万个真实数据点。每个问题在两种互补的设置中实例化:(1)SCILAWS-REAL要求模型从固定的真实观测中提出定律,并基于源文献评估其留出预测拟合度与科学有效性;(2)SCILAWS-PARALLEL要求模型主动查询残差校准的平行世界,并恢复由已发表公式派生的合成隐藏定律。这种双设置任务设计在保留每个问题科学背景的同时,分别评估固定记录下的定律发现以及对新生成隐藏定律的主动恢复。我们发现,预测拟合度可能与科学有效性相背离,记忆效应决定了模型是复现已发表公式还是超越它们,而我们的best-of-N研究揭示了一种选择瓶颈。本工作提供了一个基于论文的基准,并为评估AI科学发现提供了新的实证视角。项目页面:https://yiyihum.github.io/SciLaws-Bench
cs.AI / 124 / 2609.01567

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

基于熵的选择性智能体引导:从不完美的视觉-语言模型教师中学习自主策略
Bonetta, Giovanni, Merler, Matteo, Zago, Davide, Cancelliere, Rossella, Magnini, Bernardo
Abstract
Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as policies is expensive and brittle: they must be queried at every step, do not improve from environment interaction, and can repeat systematic errors. We study how to learn a cheap autonomous policy from an online, expensive, and imperfect but informative VLM teacher. We propose SAGE (Selective Agent Guidance via Entropy), a framework that queries a VLM only when the learner is uncertain, executes the suggested action during training, and distills guidance into a lightweight Reinforcement Learning (RL) policy. Because VLM advice is not always reliable, SAGE can weight teacher-action distillation using environment-derived advantages rather than treating all suggestions as equally useful. Across sparse-reward visual reasoning and navigation tasks, SAGE learns policies that act without VLM guidance at evaluation time and improves over unguided RL in several environments, including settings where the learned policy exceeds its VLM teacher. The results show that selective guidance is most beneficial when the VLM can help the agent discover high-reward trajectories, and less useful when unguided exploration already succeeds or teacher actions do not lead to informative experience. SAGE also reduces VLM usage by prompting the teacher only on a fraction of training steps and requiring no VLM calls at deployment. Overall, our results suggest that VLMs don't need to be used as fixed policies to be useful; they can instead act as temporary, imperfect sources of guidance whose value is tested and internalized through interaction.
Chinese Translation
视觉-语言模型(Vision-Language Models, VLMs)为交互式决策提供了有用的先验知识,但将其直接用作策略既昂贵又脆弱:它们必须在每一步被查询,无法从环境交互中改进,且可能重复系统性错误。我们研究了如何从一个在线的、昂贵的、不完美但具有信息量的VLM教师中学习一个低成本的自主策略。我们提出SAGE(基于熵的选择性智能体引导,Selective Agent Guidance via Entropy),该框架仅在学习者不确定时才查询VLM,在训练期间执行其建议的动作,并将引导蒸馏到一个轻量级的强化学习(RL)策略中。由于VLM的建议并不总是可靠的,SAGE能够利用从环境中获得的优势值(advantages)对教师动作蒸馏进行加权,而非将所有建议视为同等有用。在稀疏奖励的视觉推理与导航任务中,SAGE学习到的策略在评估时无需VLM引导即可执行动作,并在多个环境中优于无引导的强化学习,包括学习到的策略超越其VLM教师的场景。结果表明,当VLM能够帮助智能体发现高奖励轨迹时,选择性引导最为有益;而当无引导的探索已经能够成功、或教师动作无法带来有信息量的经验时,其作用则较小。SAGE还降低了VLM的使用量:仅在训练步数的一部分上提示教师,且在部署时无需任何VLM调用。总体而言,我们的结果表明,VLM并非必须作为固定的策略才有用;它们可以作为临时的、不完美的引导来源,其价值通过交互被检验并内化。
机器学习 (Machine Learning)
109
cs.LG / 1 / 2609.00047

Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training

融合全局上下文的任务特定提示用于多任务图预训练
Qiu, Zhiyang, Wang, Yangtao, Li, Xiaocui, Xie, Yanzhao, Chen, Siyuan, Zhang, Wensheng
Abstract
Graph prompt learning is an effective paradigm to adapt pre-trained graph models to downstream tasks in low-resource scenarios. However, existing multi-task graph pre-training frameworks generally use randomly initialized prompts, leading to poor alignment between the prompt space, pretext objectives and graph structural characteristics. This greatly weakens the task relevance, structural awareness and transferability of prompt representations. To address this challenge, we propose TPGC, a dual-prior prompt initialization solution that explicitly models the synergy between task prior and structural prior. Specifically, the Task-Prior Injection Module first conducts a short homologous multi-task pre-training on an auxiliary graph, enabling prompt initialization to inherit optimization preferences associated with multiple pretext tasks. Built on the task-aware representations, the Structure-Prior Injection Module further extracts transferable global structural context from the auxiliary graph, converting it into layer-wise prompt vectors by aggregating structurally informative node embeddings. Extensive experiments on 6 mainstream benchmarks covering node and graph classification show that TPGC achieves consistently better performance under few-shot settings than state-of-the-art baselines, with fewer downstream tunable parameters and lower runtime. The code is available at https://github.com/Virgilqiu/TPGC
Chinese Translation
图提示学习(Graph Prompt Learning)是一种在低资源场景下将预训练图模型适配到下游任务的有效范式。然而,现有的多任务图预训练框架通常使用随机初始化的提示,导致提示空间、代理任务目标与图结构特征之间的对齐效果较差。这极大地削弱了提示表示的任务相关性、结构感知能力和可迁移性。为应对这一挑战,我们提出了TPGC,一种显式建模任务先验与结构先验之间协同关系的双先验提示初始化方案。具体而言,任务先验注入模块(Task-Prior Injection Module)首先在辅助图上进行短暂的同源多任务预训练,使提示初始化能够继承与多个代理任务相关的优化偏好。在任务感知表示的基础上,结构先验注入模块(Structure-Prior Injection Module)进一步从辅助图中提取可迁移的全局结构上下文,并通过聚合具有结构信息的节点嵌入,将其转换为逐层的提示向量。在涵盖节点分类和图分类的6个主流基准数据集上的大量实验表明,TPGC在少样本(few-shot)设置下始终优于最先进的基线方法,同时具有更少的下游可调参数和更低的运行时间。代码已发布于 https://github.com/Virgilqiu/TPGC
cs.LG / 2 / 2609.00049

REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent

REAL-Q:基于动态梯度下降的端到端大语言模型量化
Zhang, Qian, Li, Yaoming, Tan, Zhewen, Wang, Yanshu, Lu, Heng, Su, Kun, Lv, Zongwei, Yu, Wenhan, Ma, Yongge, Han, Yinjun, Liu, Ruikuang, Yang, Tong
Abstract
Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict resource constraints. State-of-the-art PTQ methods quantize each layer with a single closed-form second-order solver: to remain analytically tractable, they heavily approximate the global loss (dropping cross-channel coupling, pooling output rows into groups), and they then freeze the resulting Hessian across the entire layer, with no way to refresh it as the loss landscape shifts column by column--a phenomenon we call information misalignment. We propose REAL-Q (Real-time E2E-loss Aligned LLM Quantization), a novel PTQ paradigm that breaks this compromise: instead of diluting the objective for the sake of analytic tractability, REAL-Q targets an end-to-end-aligned surrogate of the global loss and refines it via fine-grained, dynamic Block-wise Gradient Descent applied after every column block (128 columns). By coupling this fine-grained correction with a sliding window mechanism for smooth cross-layer transitions, REAL-Q effectively mitigates error propagation across the network. On LLaMA-3.1 (8B and 70B) and Qwen3 (0.6B-32B) at W4A16, REAL-Q reduces end-to-end KL divergence by up to ~49% relative to state-of-the-art globally-guided methods.
Chinese Translation
训练后量化(Post-training Quantization, PTQ)是在严格资源约束下部署大语言模型(LLM)的关键技术。最先进的PTQ方法通常采用单一的闭式二阶求解器对每一层进行量化:为了保持解析上的可处理性,它们对全局损失进行了大量近似(舍弃跨通道耦合、将输出行合并为组),然后在整层范围内冻结所得的Hessian矩阵,无法随着损失景观逐列变化而更新——我们将这一现象称为信息失配(information misalignment)。我们提出了REAL-Q(Real-time E2E-loss Aligned LLM Quantization,实时端到端损失对齐的大语言模型量化),这是一种打破上述折中的新型PTQ范式:REAL-Q不再为了解析可处理性而弱化优化目标,而是以全局损失的端到端对齐代理为目标,并在每个列块(128列)之后通过细粒度的动态块状梯度下降(Block-wise Gradient Descent)对其进行优化。通过将这种细粒度修正与滑动窗口机制相结合以实现平滑的跨层过渡,REAL-Q有效缓解了误差在网络中的传播。在LLaMA-3.1(8B和70B)以及Qwen3(0.6B-32B)的W4A16设置下,相较于最先进的全局引导方法,REAL-Q将端到端KL散度最多降低约49%。
cs.LG / 3 / 2609.00054

Convergence issues in Relational Concept Analysis based on AOC-posets

基于AOC偏序集的关系概念分析中的收敛性问题
Dolques, Xavier, Braud, Agnès, Gutierrez, Alain, Huchard, Marianne, Ber, Florence Le
Abstract
Formal Concept Analysis (FCA) is an approach for conceptual classification building and rule discovery from a binary table describing a set of objects by a set of attributes. Extensions have been proposed to deal with non-binary and more complex data, such as Relational Concept Analysis (RCA) for multi-relational data. RCA aims to highlight groups of objects characterized by their relationships with other groups of objects. The richer and more complex nature of the underlying data allows RCA to produce richer results than FCA, at the expense of higher computational and interpretive complexity. The most commonly used conceptual classification structure in FCA is the concept lattice. However, in many applications, concept lattice substructures, such as AOC-posets, are preferred over the full lattice, either to mitigate combinatorial blow-up or to focus on the most informative parts of the structure. Indeed, in an AOC-poset, only concepts introducing an object or an attribute are represented, which makes AOC-posets smaller and easier to compute and use than concept lattices. Although RCA was originally defined on concept lattices, it can also be instantiated on AOC-posets. RCA is iterative and its convergence is guaranteed in the lattice-based setting, but this guarantee is lost when using AOC-posets. In this paper, we investigate this loss of convergence in detail. We show why convergence is no longer guaranteed in the general case, identify conditions under which it can still be ensured, and discuss how a dataset can be transformed to recover convergence. We also propose a convergent variant of the process, which preserves the AOC-poset structure: relational attributes, once created, are never removed, which guarantees convergence at the price of attributes that may refer to concepts absent from the final structures.
Chinese Translation
形式概念分析(Formal Concept Analysis,FCA)是一种用于概念分类构建和规则发现的方法,其输入是描述对象集与属性集之间关系的二值表。为处理非二值及更复杂的数据,研究者提出了多种扩展方法,例如面向多关系数据的关系概念分析(Relational Concept Analysis,RCA)。RCA旨在凸显由对象组之间关系所刻画的各组对象。相比FCA,RCA所依赖的数据更丰富、更复杂,因而能够产生更丰富的结果,但代价是更高的计算复杂性和解释复杂性。FCA中最常用的概念分类结构是概念格。然而,在许多应用中,人们更倾向于使用概念格的子结构,例如AOC偏序集(AOC-posets),以缓解组合爆炸问题,或聚焦于结构中信息量最大的部分。事实上,在AOC偏序集中,仅表示引入对象或属性的概念,这使得AOC偏序集比概念格规模更小,更易于计算和使用。尽管RCA最初定义在概念格上,但它也可以在AOC偏序集上实例化。RCA是迭代式的,其收敛性在基于格的设定下是有保证的,但在使用AOC偏序集时这一保证便不再成立。本文详细研究了这种收敛性的丧失。我们说明了为什么在一般情况下收敛性不再有保证,识别了在哪些条件下仍能确保收敛性,并讨论了如何对数据集进行变换以恢复收敛性。我们还提出了该过程的一个收敛变体,该变体保留了AOC偏序集结构:关系属性一旦被创建就不再被移除,从而保证了收敛性,但其代价是某些属性可能指向最终结构中不存在的概念。
cs.LG / 4 / 2609.00059

DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction

DISTAL:面向结构无关材料性质预测的知识蒸馏与自监督预训练方法
Wang, Weiran, Huo, Xintong, Wang, Yueying, Fan, Yusi, Wang, Wenyan, Feng, Xin, Xin, Ruihao, Huang, Lan, Li, Kewei, Zhou, Fengfeng
Abstract
Materials property prediction remains difficult in low-data settings, where many target properties are supported by only a limited number of labeled samples. Models with the strongest predictive accuracy often depend on crystal structures, which restricts their use in early-stage screening when structural information is limited or unavailable. To address this challenge, we propose DISTAL, a dual-prior framework for structure-agnostic materials property prediction that combines self-supervised compositional pretraining with structure-aware knowledge distillation. DISTAL first learns transferable compositional representations from a large virtual composition space using 145 composition-derived descriptors. It then distills structural knowledge from a pretrained ALIGNN teacher into a composition-conditioned student. This setting allows structural priors to be used during training without requiring structural inputs at inference. By integrating explicit compositional descriptors, pretrained latent features, and distilled structural features within a unified prediction pipeline, DISTAL captures complementary signals that are difficult to recover from any single representation alone. Across 39 benchmark tasks, the best-performing multimodal configuration combines all three signals, and improves over the reference benchmark on 37 tasks. DISTAL achieves the strongest overall performance among all evaluated feature combinations. These results indicate that compositional pretraining and structural distillation provide complementary priors and offer a practical route to robust composition-only prediction in small-data materials informatics. The source code and the pre-trained models are anonymously available at: https://osf.io/eq96d/overview?view_only=451617f42f7849e08750bd1852b48980 and will be released at the official link after acceptance.
Chinese Translation
材料性质预测在低数据量场景下仍然十分困难,因为许多目标性质仅有有限的标注样本支持。预测精度最高的模型通常依赖于晶体结构信息,这限制了它们在结构信息有限或不可得情况下早期筛选中的应用。为应对这一挑战,我们提出了DISTAL,一个用于结构无关材料性质预测的双先验框架,它将自监督成分预训练与结构感知知识蒸馏相结合。DISTAL首先利用145个基于成分的描述符,从大规模虚拟成分空间中学习可迁移的成分表示;然后将预训练的ALIGNN教师模型中的结构知识蒸馏到成分条件化的学生模型中。这一设置使得结构先验可以在训练阶段加以利用,而在推理阶段无需结构输入。通过在统一的预测流程中整合显式成分描述符、预训练的潜在特征和蒸馏得到的结构特征,DISTAL能够捕获难以从任何单一表示中恢复的互补信号。在39个基准任务上,表现最佳的多模态配置整合了全部三种信号,并在其中37个任务上超越了参考基准。DISTAL在所有评估的特征组合中取得了最强的整体性能。这些结果表明,成分预训练与结构蒸馏提供了互补的先验知识,为小数据材料信息学中鲁棒的仅成分预测提供了一条切实可行的途径。源代码和预训练模型可通过以下匿名链接获取:https://osf.io/eq96d/overview?view_only=451617f42f7849e08750bd1852b48980 ,论文录用后将在官方链接发布。
cs.LG / 5 / 2609.00061

ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration

ReNFT:通过内部概率质量再校准修复奖励后训练中的模式坍缩
Bao, Yuchen, Wen, Chao, Wang, Haowei, Chen, Ruoxin, Luo, Donghao, Zhan, Jiahui, Huang, Wenjian, Chen, Shen, Wang, Yiting, Yao, Taiping, Wang, Chengjie, Ding, Shouhong, Zhang, Jianguo
Abstract
Reward post-training of diffusion generators inevitably concentrates probability mass on a few reward-favored modes, a mode collapse that erases within-prompt diversity. Existing methods for mitigating collapse rely on external signals or interfaces, augmenting the reward with perceptual objectives, adjusting reference regularization, or modifying the text encoder, but none repairs an adapter that has already collapsed while preserving the acquired reward. We observe that online post-training primarily reallocates probability mass over capabilities inherited from pretraining rather than learning new visual content. Collapse is therefore suppression, not deletion, and can be reversed from within the generator. We propose ReNFT, which repairs a high-reward, low-diversity adapter through internal probability-mass recalibration. Unconditional probes first prioritize "anti-hub" prompts where the prompt-independent bias is easiest to expose. Two policy-dominated mixed routes then generate matched counterfactual proposals from the same prompt and initial noise, one probing the frozen base direction for suppressed alternatives and the other exposing the post-trained unconditional tendency. Reward ranking with an adaptive flipping guard assigns pull and push roles, and a joint-and-paired NFT update realizes the repair. On PickScore and GenEval, ReNFT retains 98.9% and 99.0% of NFT's reward while improving DreamSim-Div by 58.8% and 55.0%, respectively, offering a complementary alternative to external interventions.
Chinese Translation
扩散生成器的奖励后训练不可避免地会将概率质量集中于少数奖励偏好的模式上,这种模式坍缩会消除提示内的多样性。现有的缓解坍缩方法依赖于外部信号或接口,例如在奖励中增加感知目标、调整参考正则化或修改文本编码器,但均无法在保留已获得的奖励的同时修复已经发生坍缩的适配器。我们观察到,在线后训练主要是在预训练继承的能力上重新分配概率质量,而非学习新的视觉内容。因此,坍缩是压制而非删除,可以从生成器内部予以逆转。我们提出ReNFT,通过内部概率质量再校准来修复高奖励、低多样性的适配器。无条件探针首先优先选择那些最易暴露与提示无关偏差的“反中心”提示。随后,两条由策略主导的混合路径在相同提示和初始噪声下生成匹配的反事实候选:一条探测冻结基座方向上被压制的替代内容,另一条暴露后训练后的无条件倾向。结合自适应翻转保护的奖励排序分配拉近与推离的角色,并通过联合与配对的NFT更新实现修复。在PickScore和GenEval上,ReNFT分别保留了NFT奖励的98.9%和99.0%,同时将DreamSim-Div分别提升了58.8%和55.0%,为外部干预提供了一种互补的替代方案。
cs.LG / 6 / 2609.00064

Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning

注意力敏感性并不足够:微调下注意力层与行为层上下文学习的解耦
Zhang, Jinyuan, He, Peng, Hu, He, Yuan, Yin, Jiao, ShengShuo
Abstract
In-context learning (ICL) lets large language models adapt to new tasks from demonstrations, and fine-tuning can erode this behaviour. Many preservation diagnostics inspect attention: if attention changes when demonstrations change, the model is treated as context-sensitive. This paper asks how far that proxy can be trusted once it is optimised. We formalise \emph{In-Context Sensitivity} (ICS), the average row distance between last-token attention on matched and mismatched demonstration prefixes, and pair it with \emph{ICL-GAP}, the behavioural accuracy gap between the same prefixes. In a controlled four-arm ablation on Llama-2-7B, an ICS-maximising regulariser ($\armKL$) drives ICS to $1.413$, within $0.5\%$ of its geometric ceiling. The behavioural readout tells a different story: ICL-GAP stays near zero and MMLU accuracy moves from $0.371$ to $0.279$, a Goodhart dissociation of the bounded attention proxy. Endpoint statistics locate the mechanism: attention grows sharp and near-disjoint across prefixes yet routes to formatting and demonstration-body tokens rather than labels. A random-label protocol confirms that the behavioural probe family retains dynamic range at the same checkpoints. In a constructive sweep, behaviour gating partially mitigates the effect, while objectives anchored to pretrained computation hold the high-MMLU, moderate-ICS region that divergence maximisers leave. The main lesson is diagnostic: attention-level ICL proxies earn their place as training targets only after validation against behavioural gaps.
Chinese Translation
上下文学习使大语言模型能够通过示例适应新任务,而微调可能会削弱这种行为。许多保持性诊断方法关注注意力:当示例变化时若注意力发生变化,模型即被视为对上下文敏感。本文探讨的是:当这一替代指标被优化后,它在多大程度上仍然可信。我们形式化了上下文敏感性,即模型在匹配与不匹配示例前缀上末词注意力的平均行间距离,并将其与ICL-GAP配对,后者衡量相同前缀之间的行为准确率差距。在针对Llama-2-7B的受控四臂消融实验中,一个最大化ICS的正则化项将ICS推升至1.413,距离其几何上限仅0.5%。然而行为层面的读数呈现不同的图景:ICL-GAP保持在接近零的水平,MMLU准确率从0.371下降至0.279,这是有界注意力替代指标出现的古德哈特式解耦。端点统计揭示了其中的机制:注意力在各前缀间变得尖锐且近乎不相交,但其路由指向格式化和示例正文词元,而非标签。随机标签协议证实,行为探针族在相同检查点上仍保持动态范围。在一项建设性扫描实验中,行为门控部分缓解了该效应,而锚定于预训练计算的目标函数则保留了高MMLU、中等ICS的区域,这是发散最大化方法所无法保持的。核心启示在于诊断层面:注意力层面的ICL替代指标只有在经过与行为差距的对照验证后,才配作为训练目标。
cs.LG / 7 / 2609.00078

RW-LoRA: Communication-Efficient Decentralized LoRA Fine-Tuning via Random Walks

RW-LoRA:基于随机游走的高通信效率去中心化LoRA微调
Chen, Xingran, Bhagat, Rohit, Ayache, Ghadir, Bitar, Rawad, Gong, Yanmin, Rouayheb, Salim El
Abstract
Parameter-efficient fine-tuning methods such as LoRA have become a standard approach for adapting large foundation models. Adopting fine-tuning to distributed settings faces several challenges. Most existing distributed LoRA methods rely on centralized aggregation, and gossip-based decentralized LoRA requires repeated synchronization among multiple model copies. Both methods incur significant communication overhead and introduce errors due to simultaneous aggregation of multiple model updates. In this paper, we take a different perspective and propose a random-walk-based LoRA fine-tuning scheme. Instead of maintaining multiple model replicas, a single model token traverses the network and is updated sequentially using local fine-tuning objectives. This design eliminates the need for global synchronization, substantially reduces communication and computation costs, and avoids aggregation errors. We provide rigorous convergence guarantees for non-convex objectives under standard assumptions. Through empirical results on multiple NLP tasks and graph topologies, we show that the proposed method achieves competitive task performance with substantially less communication and computation than gossip-based LoRA.
Chinese Translation
以LoRA为代表的参数高效微调方法已成为适配大型基础模型的标准途径。然而,将微调推广到分布式场景面临诸多挑战:大多数现有的分布式LoRA方法依赖中心化聚合,而基于gossip的去中心化LoRA则需要在多个模型副本之间进行反复同步。这两种方式都会带来巨大的通信开销,并且由于同时聚合多个模型更新而引入误差。本文从不同的视角出发,提出了一种基于随机游走(random walk)的LoRA微调方案。该方法不再维护多个模型副本,而是让单一模型令牌在网络中游走,并依次利用本地微调目标对其进行更新。这一设计消除了全局同步的需要,大幅降低了通信与计算成本,并避免了聚合误差。我们在标准假设下为非凸目标提供了严格的收敛性保证。通过在多个自然语言处理(NLP)任务和图拓扑上的实验结果,我们表明所提出的方法以远低于基于gossip的LoRA的通信和计算开销,取得了具有竞争力的任务性能。
cs.LG / 8 / 2609.00084

Stochastic complexity of vectors containing cluster structure

包含聚类结构的向量的随机复杂度
Nicorici, Daniel, Yli-Harja, Olli, Astola, Jaakko
Abstract
This paper studies the problem of computing the stochastic probability (shortest code length) of the encoded vectors containing cluster structure using Normalized Maximum Likelihood (NML) model. This is of great theoretical and practical importance in data clustering based on Minimum Description Length (MDL) principle, such as for estimating the best number of clusters and best cluster structure for the data. Straightforward computation of the shortest code length of the vector containing cluster structure based on the NML model requires polynomial time with respect to the size of the vector and number of clusters. We show that this is a tractable problem by introducing a recursion formula for the efficient computation of normalizing constant from the NML model. The time complexity of the new formula is linear opposed to previous polynomial time with respect to the size of the vector and number of clusters.
Chinese Translation
本文研究利用归一化最大似然(NML)模型计算包含聚类结构的编码向量的随机概率(最短码长)的问题。这对基于最小描述长度(MDL)原则的数据聚类具有重要的理论和实践意义,例如用于估计数据的最佳聚类数目和最佳聚类结构。基于NML模型直接计算包含聚类结构的向量的最短码长需要关于向量大小和聚类数目的多项式时间。我们通过引入一个递归公式来高效计算NML模型的归一化常数,证明该问题是可解的。新公式的时间复杂度相对于向量大小和聚类数目是线性的,优于以往的多项式时间。
cs.LG / 9 / 2609.00089

Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models?

面向电价预测与电池套利的基础模型:它们能否取代特定市场的预测模型?
Lipiecki, Arkadiusz, Weron, Rafał
Abstract
Foundation models promise accurate forecasts with little or no task-specific training, but whether they can replace models designed specifically for electricity price forecasting remains unclear. We compare nine variants from five foundation model families, evaluated in zero-shot mode, with two state-of-the-art electricity price forecasting benchmarks in Germany, Poland, and Spain over 2021-2025. Their performance is assessed in terms of point and probabilistic forecasting accuracy, as well as economic value in battery energy storage arbitrage. Only the TabPFN models consistently and significantly outperform the benchmarks across all three markets and all statistical measures. However, this statistical dominance does not translate directly into economic dominance: TabPFN performs best under unlimited bids and riskier quantile-based strategies, whereas the Distributional Deep Neural Network benchmark is more profitable when risk tolerance is lower. Thus, foundation models cannot universally replace market-specific models, and their value depends on both model architecture and the decision problem.
Chinese Translation
基础模型(Foundation models)有望在几乎无需或完全无需任务特定训练的情况下提供准确预测,但其能否取代专为电价预测设计的模型仍不明确。我们在德国、波兰和西班牙三个市场(2021-2025年)上,将来自五个基础模型系列的九种变体(以零样本(zero-shot)模式运行)与两个最先进的电价预测基准模型进行比较。评估涵盖点预测与概率预测精度,以及电池储能套利中的经济价值。结果表明,只有 TabPFN 模型在所有三个市场和所有统计指标上持续且显著地优于基准模型。然而,这种统计上的优势并未直接转化为经济上的优势:在无限制报价和风险较高的分位数策略下,TabPFN 表现最佳;而当风险容忍度较低时,分布式深度神经网络(Distributional Deep Neural Network)基准模型则更为有利可图。因此,基础模型并不能普遍取代特定市场的预测模型,其价值取决于模型架构与具体决策问题两个方面。
cs.LG / 10 / 2609.00090

Assessing Alignment and Stability of Feature Importance Explanations via Weight of Evidence

基于证据权重评估特征重要性解释的对齐性与稳定性
Conti, Eddie, Daka, Claudio, Parafita, Álvaro, Alfeo, Antonio L., Brando, Axel, Cimino, Mario G. C. A.
Abstract
Feature importance Methods (FIMs) are widely used in Explainable AI to interpret model predictions, yet attribution scores alone often provide limited insight into the underlying reasoning process. In this work, we introduce a novel perspective by embedding FIMs within a hypothesis-testing framework based on Weight of Evidence (WoE). We quantify how strongly the observed evidence supports any given hypothesis on feature importance. The reference hypothesis can stem from domain knowledge, ground truth, or be derived from the FIM itself. This formulation enables a principled evaluation of FIMs, capturing both their alignment with prior knowledge and their variability. We further provide theoretical results linking WoE to attribution variance. Empirical results shows the applicability and flexibility of our strategy analyzing LIME and SHAP explanations in settings with different reference hypotheses. Overall, our framework offers a complementary tool for assessing FIMs through a contrastive, evidence-based lens.
Chinese Translation
特征重要性方法(Feature Importance Methods, FIMs)在可解释人工智能(Explainable AI)中被广泛用于解释模型预测,但仅凭归因分数往往难以深入洞察模型背后的推理过程。在这项工作中,我们引入了一种新颖的视角,将FIMs嵌入基于证据权重(Weight of Evidence, WoE)的假设检验框架中,量化观测到的证据对任意给定特征重要性假设的支持强度。该参考假设可以来源于领域知识、真实标签(ground truth),或由FIM本身推导得出。这一表述方式能够对FIMs进行有原则的评估,同时捕捉其与先验知识的对齐程度及其变异性。我们进一步提供了将WoE与归因方差联系起来的理论结果。实验结果表明,在设置不同参考假设的情况下,我们对LIME和SHAP解释进行分析的策略具有良好的适用性和灵活性。总体而言,我们的框架为通过对比性的、基于证据的视角评估FIMs提供了一种互补工具。
cs.LG / 11 / 2609.00092

Safin-1: Safety from Within through Memory-Native State Evolution

Safin-1:通过记忆原生状态演化实现由内而外的安全
Zhang, Ming, Yang, Kaisen, Yu, Shu, Hua, Ermo, Chen, Zhekai, Jin, Cheng, Zheng, Jingnan, Zhang, Yi, Ma, Zhongtian, Zhou, Jiawei, Chen, Sirui, Zhang, Qiaosheng, Wang, Xiang, Ding, Ning, Hu, Xia, Zhou, Bowen, Sun, Youbang, Lu, Chaochao
Abstract
Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic property of the model itself, rather than a behavioral constraint relying solely on external safeguards or post-hoc alignment such as supervised fine-tuning. This motivates Safety from Within, where safety-relevant capabilities are represented and invoked through the model's native computation. We present Safin-1, a family of foundation models realizing this principle through memory routing and state evolution. Safin-1 is built on Memory-Anchor Routing across Context History (MARCH), a network architecture that maintains structured memory states and selectively retrieves relevant historical information through content-conditioned routing. It supports test-time adaptation of persistent capability states without repeatedly modifying the backbone, enabling controlled specialization over a shared foundation. We investigate this interface on downstream safety tasks through a Safety State, demonstrating effective state-based adaptation with substantial safety improvements. More broadly, the routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior. Evaluations across general capabilities, long-context understanding, retrieval, and efficiency further validate Safin-1. These findings provide a path toward safety as a state-native and adaptively maintainable capability. This work is only an initial architectural exploration of Safety from Within, and substantial further work is needed to realize this broader vision.
Chinese Translation
长周期复杂任务要求基础模型能够积累信息、维护内部状态,并在长期交互中不断适应。安全性应当是模型自身的一种内在属性,而非仅仅依赖外部防护机制或事后对齐(如监督微调)等行为约束。这启发了"由内而外的安全"(Safety from Within)这一理念,即安全相关能力通过模型的原生计算来表示和调用。我们提出了Safin-1,一个通过记忆路由和状态演化实现这一原则的基础模型系列。Safin-1建立在跨上下文历史的记忆锚定路由(Memory-Anchor Routing across Context History, MARCH)架构之上,该网络架构维护结构化的记忆状态,并通过基于内容的条件路由选择性地检索相关的历史信息。它支持对持久能力状态进行测试时自适应,而无需反复修改主干网络,从而在共享基础之上实现受控的专业化。我们通过一个安全状态(Safety State)在下游安全任务上研究了这一接口,展示了基于状态的有效自适应,并带来了显著的安全性提升。更广泛地说,这种路由状态接口将上下文记忆和持久能力自适应统一到模型的原生计算中,将记忆从对先前上下文的被动记录重构为维护和演化模型行为的主动基底。在通用能力、长上下文理解、检索和效率方面的评估进一步验证了Safin-1的有效性。这些发现为实现安全作为一种状态原生且可自适应维护的能力提供了一条路径。本工作仅是对"由内而外的安全"理念的初步架构探索,实现这一更宏大的愿景仍需要大量进一步的研究。
cs.LG / 12 / 2609.00093

Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification

面向不平衡时间序列分类的局部参考几何残差增强
Qiu, Chuanhang, Xu, Yanran, Wang, Yue, Bagnall, Anthony
Abstract
Imbalanced time series classification is often addressed by changing the training distribution, objective, logits, or final threshold. These interventions address important biases, yet leave a representation-level question unmeasured: after minority support is reduced, does a learned feature space remain locally reliable around minority regions? We identify a training-local geometry failure: under imbalance, minority cases can lie in sparse, rest-dominated, or mixed feature-space neighborhoods, even when the representation retains useful global class structure. To diagnose and repair this failure, we propose Local Reference Geometry (LRG), a lightweight post-hoc feature augmentation module applied between a fixed feature extractor and the classifier head. Using training features only, LRG measures local exposure and class-mixture risk, then augments each fixed feature with a standardized signed displacement from nearby training geometry and an LDA-projected residual summary. On controlled UCR/Bake Off Redux imbalance benchmarks, paired raw-versus-LRG comparisons show gains for learned, pretrained, and fixed representations, including when LRG is combined with training-level interventions and post-encoder classifier corrections. Ablations show that the gain comes from the signed local residual appended to the original feature, rather than from generic prototype distances, affinity features, scalar statistics, or VLAD-style codes. Further analyses support the proposed local-geometry failure hypothesis: minority neighborhoods become increasingly rest-exposed under imbalance, training-local risk identifies error-prone regions, and LRG gains concentrate in those high-risk regions.
Chinese Translation
不平衡时间序列分类问题通常通过改变训练分布、目标函数、logits 或最终阈值来解决。这些干预措施处理了重要的偏差,但留下一个未在表示层面衡量的问题:当少数类支持被削弱后,所学到的特征空间在少数类区域附近是否仍然局部可靠?我们识别出一种训练局部几何失效现象:在不平衡条件下,即使表示仍然保留了有用的全局类别结构,少数类样本也可能落在稀疏的、被多数类主导的或类别混合的特征空间邻域中。为了诊断并修复这一失效现象,我们提出了局部参考几何(Local Reference Geometry, LRG),这是一个轻量级的后验特征增强模块,应用于固定的特征提取器与分类器头之间。LRG 仅使用训练特征,度量局部暴露程度和类别混合风险,然后为每个固定特征增加一个来自邻近训练几何的标准化带符号位移,以及一个经 LDA 投影的残差摘要。在受控的 UCR/Bake Off Redux 不平衡基准上,原始特征与 LRG 特征的配对比较表明,对于学习型、预训练型和固定型表示均有收益,包括当 LRG 与训练层面的干预以及编码器后的分类器校正相结合时。消融实验表明,收益来源于附加到原始特征上的带符号局部残差,而非通用的原型距离、亲和度特征、标量统计量或 VLAD 风格编码。进一步的分析支持了所提出的局部几何失效假设:在不平衡条件下,少数类邻域的多数类暴露程度不断加剧;训练局部风险能够识别易错区域;而 LRG 的收益恰好集中在这些高风险区域。
cs.LG / 13 / 2609.00097

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

快如闪电之上:利用注意力稀疏性实现高效的长上下文解码
Liu, Zhigeng, Ning, Zhiyuan, Li, Ruixiao, Liu, Xiaoran, Song, Yuerong, Zhang, Min, He, Ziwei, Qiu, Xipeng
Abstract
The development of long-context Large Language Models (LLMs) is constrained by the memory bandwidth bottleneck and quadratic complexity of the attention mechanism during decoding. To overcome the inherent trade-offs between the memory overhead of metadata-based metrics and the computational inefficiency of adaptive selection strategies, we present Faster Flash Decoding (FFD), a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding. FFD integrates the selector and computer into a fully fused kernel, replacing external metadata indices with content-aware scanning via low-bit quantization. Furthermore, we introduce the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization. Offering a training-free and plug-and-play solution, FFD also enables the reuse of scanning results for computation, achieving up to 11.6x kernel-level speedup and scaling to 256K context length, with 2.37x end-to-end throughput improvement. Empirical validation on RULER and LongBench confirms that FFD maintains model accuracy while delivering high-ratio sparsity, with code available at https://github.com/qluoluo/faster-flash-decoding
Chinese Translation
长上下文大语言模型(LLM)的发展受限于解码过程中注意力机制带来的内存带宽瓶颈和二次方复杂度。为了克服基于元数据的度量方法带来的内存开销与自适应选择策略的计算低效之间固有的权衡,我们提出了Faster Flash Decoding(FFD),这是一种新颖的软硬件(硬件-算法)协同设计框架,旨在突破长上下文解码中的内存墙。FFD将选择器与计算器集成到一个完全融合的核函数(kernel)中,通过低比特量化实现的内容感知扫描取代了外部元数据索引。此外,我们引入了top-delta策略,该策略动态筛选块,从而在无需全局同步的情况下实现分布自适应的稀疏性。FFD提供了一种免训练、即插即用的解决方案,同时支持将扫描结果复用于计算,实现了高达11.6倍的核级加速,并可扩展至256K上下文长度,端到端吞吐量提升2.37倍。在RULER和LongBench上的实证验证表明,FFD在保持模型精度的同时实现了高比例的稀疏性,代码已发布于 https://github.com/qluoluo/faster-flash-decoding
cs.LG / 14 / 2609.00099

Generative artificial intelligence for reliable mechanistic reasoning for corrosion

用于腐蚀可靠机理推理的生成式人工智能
N, Bharath M, Raman, R K Singh, Alankar, Alankar
Abstract
Corrosion accounts for approximately 4% of global GDP, and reliable prediction is essential for timely mitigation. Machine learning effectively predicts corrosion rates from composition, microstructure, and environmental variables, but cannot explain the underlying mechanisms. A reliable approach in safety-critical materials engineering requires not only accurate retrieval but also mechanistically defensible reasoning, a capability that existing factuality metrics cannot assess. This work presents a domain-adapted retrieval-augmented generation framework for corrosion knowledge synthesis, demonstrated on magnesium alloy corrosion. Three open-weight language models (Llama-3.1-8B, Qwen-2.5-7B, Mistral-7B) are fine-tuned on 3,309 expert-verified question-answer pairs from 840 peer-reviewed papers and integrated with a hybrid dense-lexical retrieval pipeline. Retrieval augmentation produces Token F1 gains of 143-194%, with system faithfulness of 0.964 and context recall of 0.988. Blind external validation on newly published literature and in-house electrochemical data confirms trend-level generalisation. Reason Map, a proposition-graph framework, is further introduced; it independently constructs directed evidence graphs from generated answers and retrieved literature, enabling systematic detection of causal direction inversions and unsupported inferential leaps that flat factuality metrics cannot expose. The modular architecture can be applied across domains, offering a generalizable blueprint for trustworthy AI-assisted knowledge synthesis to circumvent corrosion, which can also be applied to other engineering domains.
Chinese Translation
腐蚀造成的损失约占全球GDP的4%,可靠的预测对于及时采取缓解措施至关重要。机器学习能够有效地基于成分、微观组织和环境变量预测腐蚀速率,但无法解释其背后的机理。在安全攸关的材料工程中,可靠的方法不仅需要准确的检索,还需要机理上可辩护的推理——这是现有事实性指标无法评估的能力。本研究提出一个面向腐蚀知识综合的领域适配检索增强生成(RAG)框架,并以镁合金腐蚀为例进行了演示。三个开放权重语言模型(Llama-3.1-8B、Qwen-2.5-7B、Mistral-7B)在来自840篇同行评审论文的3,309个经专家验证的问答对上进行了微调,并与混合稠密-词法检索流水线集成。检索增强带来143%–194%的Token F1提升,系统忠实度达0.964,上下文召回率达0.988。基于新发表文献和自制电化学数据的盲法外部验证证实了模型具有趋势层面的泛化能力。研究进一步引入了Reason Map——一种命题图框架,它从生成的答案和检索到的文献中独立构建有向证据图,从而能够系统性检测平坦事实性指标无法揭示的因果方向反转和缺乏依据的推理跳跃。该模块化架构可推广应用于其他领域,为可信的AI辅助知识综合提供了可泛化的蓝图,不仅可应用于腐蚀问题的规避,也可推广至其他工程领域。
cs.LG / 15 / 2609.00103

Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy

好的记忆具备ECC特性:超越准确率评估视觉语言模型的记忆能力
Berman, Shmuel, Deng, Jia
Abstract
Memory is widely viewed as an important unsolved problem for LLMs and VLMs, and current benchmarks typically evaluate it by testing accuracy over long text or video. However, accuracy alone misses properties that matter for real long-horizon tasks. We introduce ECCBench, a benchmark and evaluation protocol that measures memory beyond a system's capacity--its raw accuracy at a specific budget--via three axes we call ECC: efficiency--the computation, in FLOPs, needed to answer from memory; compression--whether compressible inputs are remembered more accurately or efficiently; and calibration--whether the system abstains in response to its own uncertainty and the cost of an error. We find that pretrained VLMs compress their memory over text but not video and are poorly calibrated on both. Among a broader set of memory backbones, several non-Transformer architectures achieve better compression-calibration tradeoffs than RoPE Transformers, suggesting they may be useful components for agents operating over long horizons.
Chinese Translation
记忆被广泛视为大语言模型(LLM)和视觉语言模型(VLM)面临的一个重要未解决问题,当前的基准测试通常通过测试长文本或长视频上的准确率来评估记忆能力。然而,仅凭准确率无法反映对真实长程任务至关重要的其他特性。我们提出了ECCBench,这是一个从超越系统容量(即在特定预算下的原始准确率)角度衡量记忆能力的基准和评估协议,通过我们称为ECC的三个维度进行评估:效率(efficiency)——从记忆中作答所需的计算量(以FLOPs计);压缩(compression)——可压缩的输入是否被更准确或更高效地记住;以及校准(calibration)——系统能否在面对自身不确定性时选择弃权,以及出错的代价。我们发现,预训练的VLM在文本上能够压缩其记忆,但在视频上不能,且两者上的校准表现都很差。在更广泛的记忆骨干架构中,若干非Transformer架构实现了优于RoPE Transformer的压缩-校准权衡,这表明它们对于执行长程任务的智能体而言可能是具有价值的组件。
cs.LG / 16 / 2609.00129

Flawed in Nature, Perfect through Evolution

天生有瑕,进化致善
Kruijssen, J. M. Diederik
Abstract
The performance of artificial intelligence (AI) and machine learning (ML) models degrades when the problem they were trained on drifts. This is a near-universal feature of real-world problems, which often change unpredictably. Biological evolution has achieved intelligence by overcoming this obstacle through natural selection acting on heritable variation. AI/ML techniques have long incorporated forms of natural selection, but it has been challenging to maintain model diversity as optimization naturally drives convergence. Here we show that a swarm of AI/ML models subjected to deliberate mutations of their model coefficients away from optimality can reliably and sustainably improve performance in changing environments by acting as a statistical hedge against non-stationarity. We call this mechanism 'Flawed in Nature, Perfect through Evolution', reflecting that the collective performance gain goes at the expense of individual performance. We prove via four theorems that the resulting regret reduction is guaranteed under general conditions, establishing the Flawed-in-Nature mechanism as a generalizable design principle for AI/ML systems. We validate these results on synthetic linear regression tasks, demonstrating that the mutated swarm delivers the best model in $\sim80\%$ of environment changes and that inference synthesis successfully translates this individual advantage into a collective one. The mechanism proves to be most effective when the mutation drift rate matches the drift rate of the environment. We outline a simple, adaptive controller that enables practical applications by tuning the mutation drift rate to match the unknown drift rate of the environment. The close analogy of the Flawed-in-Nature mechanism to biological evolution suggests it may have been a critical missing ingredient for the organic discovery of AI forms that more closely mimic biological intelligence.
Chinese Translation
当人工智能(AI)和机器学习(ML)模型所训练的问题发生漂移时,其性能会下降。这是现实世界问题近乎普遍的特征,因为现实问题往往以不可预测的方式变化。生物进化通过自然选择作用于可遗传变异,克服了这一障碍,从而实现了智能。AI/ML技术早已融入了自然选择的各种形式,但随着优化过程自然地驱动模型趋同,维持模型的多样性一直颇具挑战。在本文中,我们展示了通过对一群AI/ML模型的模型系数进行有意偏离最优状态的变异,可以在变化的环境中可靠且持续地提升性能,其作用相当于针对非平稳性的一种统计对冲。我们将这一机制称为“天生有瑕,进化致善”,反映出集体性能的增益是以牺牲个体性能为代价的。我们通过四个定理证明,在一般条件下,由此产生的遗憾(regret)降低是有保证的,从而确立了“天生有瑕”机制作为AI/ML系统的一种可推广的设计原则。我们在合成线性回归任务上验证了这些结果,证明变异模型群在约80%的环境变化中能提供最佳模型,且推理合成成功地将这种个体优势转化为集体优势。当变异漂移速率与环境漂移速率相匹配时,该机制最为有效。我们提出了一个简单的自适应控制器,通过调节变异漂移速率以匹配未知的环境漂移速率,从而实现实际应用。“天生有瑕”机制与生物进化的紧密相似性表明,它可能是实现更接近生物智能的AI形态的自然发现过程中一个关键的缺失要素。
cs.LG / 17 / 2609.00189

Elite-Weighted Supervised Fine-tuning for Goal-Directed Molecular Optimization

面向目标导向分子优化的精英加权监督微调方法
Wa, Shiyun, Wang, Yifei, Green, Anna G., Sciabola, Simone, Wang, Ye
Abstract
Goal-directed optimization is essential for steering molecular generators to propose candidates with desired properties. However, it is often implemented with policy-gradient reinforcement learning, which requires a generation-trajectory log-probability whose form depends on the model architecture and generation procedure. This makes an optimizer difficult to reuse across architectures and conditional generative designs. Supervised fine-tuning needs none of that machinery, but its update is driven by a fixed dataset, so the reward never enters the update. We introduce Elite-Weighted Supervised Fine-tuning (EW-SFT), which uses reward to guide elite selection of high-scoring molecules, and updates the model by its own pretraining loss on that set. Ablations show that reward information is passed primarily through elite selection, rather than through continuous weighting within the selected set. Because the update consumes only scored molecules and the model's native loss, the same rule applies across autoregressive, masked-diffusion, and discrete-flow generators, and across de novo, motif-extension, and linker-design tasks. Under a fixed budget of 3D shape alignment oracle calls on two kinase reference compounds, EW-SFT consistently outperforms the corresponding native optimizers. It further improves goal-directed optimization under a 2D similarity oracle on four held-out references and achieves comparable performance on a sample-efficiency benchmark without a trajectory-level RL formulation. These results demonstrate that EW-SFT is a unified and effective optimizer across molecular generators, design constraints, references, and oracles.
Chinese Translation
目标导向优化对于引导分子生成器提出具有理想性质的候选分子至关重要。然而,该优化通常采用策略梯度强化学习实现,这需要生成轨迹的对数概率,其形式依赖于模型架构和生成过程,使得优化器难以在不同架构和条件生成设计之间复用。监督微调无需这些机制,但其更新由固定数据集驱动,因此奖励从未进入更新过程。我们提出了精英加权监督微调(Elite-Weighted Supervised Fine-tuning, EW-SFT),该方法利用奖励引导高得分分子的精英选择,并通过模型自身在这些分子上的预训练损失进行更新。消融实验表明,奖励信息主要通过精英选择传递,而非通过所选集合内的连续加权传递。由于该更新仅消耗已评分分子和模型的原生损失,同一规则可适用于自回归、掩码扩散和离散流生成器,以及从头生成、基序延伸和连接子设计等任务。在两个激酶参考化合物上以固定的三维形状对齐oracle调用预算进行实验,EW-SFT始终优于相应的原生优化器。在四个留出参考化合物上的二维相似性oracle实验中,它进一步提升了目标导向优化效果,并在无轨迹级强化学习框架的情况下,在样本效率基准上取得了相当的性能。这些结果表明,EW-SFT是一种跨分子生成器、设计约束、参考化合物和oracle的统一且有效的优化器。
cs.LG / 18 / 2609.00196

WHALE: A Simple Recipe for Joint Harness-Weight Optimization

WHALE:一种联合优化框架与权重的简易方法
Kim, Haechan, Lee, Yoonho, Lee, Gisang, Finn, Chelsea, Lee, Kangwook
Abstract
Agent performance depends jointly on the model parameters and the executable harness code that manages context and control flow. Optimizing either component in isolation can leave the system bottlenecked by its frozen counterpart: weight updates can change which harness is effective, while harness updates can change which model capabilities are exposed. Existing joint-adaptation methods optimize weights and textual prompts but leave the broader harness fixed. We propose Weight-Harness Alternating LEarning (WHALE), a simple recipe that alternates two phases: updating the model under the current harness, then searching for a better harness under the updated model. We instantiate these two phases with online rejection-sampling fine-tuning and Meta-Harness, respectively. When to switch is a key design choice: to separate real improvements from noise without over-optimizing against a changing counterpart, WHALE uses either fixed phase durations or an adaptive patience rule over training signals. Using Qwen3.5-2B/4B agents across three domains (search question answering, mathematical reasoning, and chess puzzles), WHALE outperforms weight-only, harness-only, and Fast-Slow Training by 4.15-24.38 percentage points in best mean@8 accuracy. Either component can be the bottleneck: harness search matches peak weight-only accuracy with far fewer rollouts in SearchQA, but improves math accuracy only after a weight update. Small interleaved updates also outperform stagewise weight-then-harness optimization in accuracy and rollout cost. The code is available at https://github.com/krafton-ai/WHALE.
Chinese Translation
智能体的性能同时取决于模型参数和管理上下文与控制流的可执行框架(harness)代码。单独优化任一组件都可能使系统受制于被冻结的另一组件:权重的更新会改变哪个框架是有效的,而框架的更新也会改变模型哪些能力被暴露出来。现有的联合自适应方法只优化权重和文本提示,而将更广泛的框架保持固定。我们提出权重-框架交替学习(Weight-Harness Alternating LEarning,WHALE),这是一种简单的交替执行两个阶段的方法:在当前框架下更新模型,然后在更新后的模型下搜索更好的框架。我们分别用在线拒绝采样微调和Meta-Harness来实例化这两个阶段。何时切换是一个关键的设计选择:为了在不针对不断变化的另一组件过度优化的前提下将真实改进与噪声区分开来,WHALE使用固定的阶段时长或基于训练信号的自适应耐心规则。在三个领域(搜索问答、数学推理和国际象棋谜题)上使用Qwen3.5-2B/4B智能体,WHALE在最佳mean@8准确率上比仅权重、仅框架以及快慢训练(Fast-Slow Training)方法高出4.15至24.38个百分点。任一组件都可能成为瓶颈:在SearchQA上,框架搜索以远少的推理(rollout)次数达到了仅权重优化的峰值准确率,但在数学任务上只有经过权重更新后才能提升准确率。小规模的交错更新在准确率和推理成本上也优于先权重后框架的分阶段优化。代码已发布于 https://github.com/krafton-ai/WHALE。
cs.LG / 19 / 2609.00224

QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization

QTEA:基于稀疏残差显著权重与按列优化的三值大语言模型
Guo, Yipin, George, Arun M, Fu, Jie, Mahmoud, Tareq, Xing, Sixue, Joshi, Siddharth
Abstract
Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods often fail to generalize across models and suffer severe accuracy loss below 2 bits. Many leverage unstructured sparsity to mitigate this loss, but at the cost of regularity and GPU-friendly execution. We present QTEA, a sub-2-bit PTQ framework that quantizes weights into ternary values and uses salient weights as residual error compensators. To maintain hardware efficiency, residuals are assigned to selected columns with semi-structured $1:4$ sparsity within the salient columns. We further add column-wise rescale refinement to GPTQ-style column-by-column quantization, alternately updating per-column scales and ternary assignments to reduce reconstruction error. We also identify order-dependent error propagation in GPTQ and introduce error decay to attenuate late-stage error accumulation. On Qwen3-14B, QTEA compresses all weights to an effective 1.7 bits per weight while improving average accuracy over the strongest ternary PTQ baseline by 16.7%. It also achieves 1.40$\times$ and 2.61$\times$ lower perplexity on WikiText and C4 respectively. This trend holds on Llama3-8B, where QTEA obtains a 6.6% accuracy gain and 1.34$\times$ / 1.95$\times$ lower perplexity on the same datasets. Finally, we develop a lookup-table based kernel that achieves 7.2$\times$ faster per-token generation over an FP16 baseline. Code is available at https://github.com/Intelligent-Microsystems-Lab/QTEA.
Chinese Translation
仅权重的训练后量化(PTQ)可以缓解大规模服务大语言模型(LLMs)所带来的计算负担。然而,现有的PTQ方法往往难以在不同模型之间泛化,并且在低于2比特时会出现严重的精度损失。许多方法利用非结构化稀疏性来缓解这一损失,但代价是破坏了规则性且不利于GPU高效执行。我们提出了QTEA,这是一个低于2比特的PTQ框架,它将权重量化为三值,并利用显著权重作为残差误差补偿器。为了保持硬件效率,残差被分配到选定的列上,并在显著列内部采用半结构化的 $1:4$ 稀疏性。我们进一步在GPTQ风格的逐列量化基础上增加了列级缩放精调,交替更新每列的缩放因子和三值分配,以降低重构误差。我们还发现了GPTQ中依赖顺序的误差传播问题,并引入误差衰减机制来减轻后期阶段的误差累积。在Qwen3-14B上,QTEA将所有权重压缩到每权重有效1.7比特,同时与最强的三值PTQ基线相比平均精度提升了16.7%;在WikiText和C4数据集上的困惑度分别降低了1.40倍和2.61倍。这一趋势在Llama3-8B上同样成立,QTEA获得了6.6%的精度提升,并在相同数据集上的困惑度分别降低了1.34倍和1.95倍。最后,我们开发了一种基于查找表(lookup-table)的内核,相比FP16基线实现了7.2倍的逐token生成加速。代码已发布于 https://github.com/Intelligent-Microsystems-Lab/QTEA。
cs.LG / 20 / 2609.00297

Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains

面向复杂域偏微分方程的几何感知潜在自回归生成模型
Wang, Zi, Xu, Minghui, Mukerji, Tapan
Abstract
Solving multiphysics partial differential equations (PDEs) remains a major challenge in scientific computing, especially for highly complex $\mu$m-scale tortuous geometries critical to energy and chemical engineering. We address this challenge by proposing a Geometry-aware Latent Autoregressive generative Model for PDEs (GeoLAMP) for solving physics within highly irregular and tortuous structures. GeoLAMP introduces a dual-encoder architecture on graph representations to jointly capture global topology and fine-scale geometric features, enabling an effective transition from real-space fields to compact latent representations. In the latent space, we propose a causal self-attention transformer with flow matching to model temporal dynamics, allowing stable and scalable block-wise autoregressive prediction. A flexible decoder reconstructs high-resolution physical fields on arbitrary points. We establish three multiphysics benchmark datasets in complex geometries, covering reactive flow, heat convection, and elasticity. GeoLAMP consistently achieves the most stable autoregression performance on these datasets, maintaining low errors throughout the entire rollout horizon. Our results provide a systematic study of geometry-aware learning for PDEs in $\mu$m-scale complex geometries and offer new insights into block-wise time marching of latent autoregressive PDE modeling via a flow matching framework.
Chinese Translation
求解多物理场偏微分方程(PDE)仍然是科学计算中的一项重大挑战,尤其是对于对能源与化学工程至关重要的高度复杂的微米级(μm尺度)曲折几何结构。为应对这一挑战,我们提出了一种面向偏微分方程的几何感知潜在自回归生成模型(Geometry-aware Latent Autoregressive generative Model for PDEs,GeoLAMP),用于求解高度不规则且曲折结构内的物理问题。GeoLAMP 在图表示上引入双编码器架构,以联合捕获全局拓扑结构与细尺度几何特征,实现从真实空间物理场到紧凑潜在表示的有效转换。在潜在空间中,我们提出了一种结合流匹配(flow matching)的因果自注意力 Transformer 来建模时间动力学,从而实现稳定且可扩展的块级自回归预测。灵活的解码器可在任意点上重建高分辨率物理场。我们建立了三个复杂几何下的多物理场基准数据集,涵盖反应流、热对流和弹性力学。GeoLAMP 在这些数据集上始终取得最稳定的自回归性能,并在整个滚动预测时域内保持较低误差。我们的结果为微米级复杂几何中偏微分方程的几何感知学习提供了系统性研究,并通过流匹配框架为潜在自回归偏微分方程建模的块级时间推进方法提供了新的见解。
cs.LG / 21 / 2609.00345

Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment

大语言模型了解你的社区吗?审计大语言模型在社区级出行预测中的先验知识与结构一致性
Abrar, Saad Mohammad, Kurella, Eesha, Dadarya, Arnav, Awasthi, Naman, Zinat, Kazi Tasnim, Frias-Martinez, Vanessa
Abstract
Human mobility is central to urban planning, transportation, public health, and emergency response, yet fine-grained trajectory data are often proprietary, restricted, and privacy-sensitive. Large language models (LLMs) offer a potential alternative by generating plausible mobility traces and predicting individual movement, but their ability to infer aggregate neighborhood-level mobility remains unclear. We evaluate zero-shot LLMs on Census Block Group-level mobility prediction across four U.S. metropolitan areas using anonymized Cuebiq data to construct point-level, trajectory-level, and temporal mobility outcomes, paired with sociodemographic and built-environment predictors. We compare LLM predictions with supervised baselines and introduce a directional alignment analysis to test whether LLM-implied predictor effects agree with empirical OLS and Jonckheere-Terpstra trends. Supervised models achieve 0.580 average accuracy, compared with 0.435 for the best LLM, with spatial extent outcomes showing the strongest predictability but also the largest LLM-baseline gaps. Directional analysis shows that LLMs often rely on coarse, stable predictor-level priors that remain similar across outcomes and cities, including asymmetric treatment of protected-group predictors. Overall, LLMs can partially recover aggregate mobility patterns from urban context, but their predictions should not be treated as structurally grounded without auditing empirical alignment and potential bias.
Chinese Translation
人类出行是城市规划、交通、公共卫生和应急响应的核心,但精细粒度的轨迹数据往往专有、受限且涉及隐私敏感。大语言模型(LLM)通过生成合理的出行轨迹和预测个体移动提供了潜在替代方案,但其推断社区层面汇总出行的能力仍不明确。我们使用匿名化的 Cuebiq 数据构建点级、轨迹级和时间维度的出行结果,结合社会人口学与建成环境预测变量,在美国四个大都市区上评估零样本 LLM 在人口普查区块组(Census Block Group)级别的出行预测表现。我们将 LLM 预测与有监督基线模型进行比较,并引入方向性一致性分析,检验 LLM 所隐含的预测变量效应是否与实证 OLS 及 Jonckheere-Terpstra 趋势相符。有监督模型的平均准确率为 0.580,而表现最佳的 LLM 仅为 0.435;其中空间范围类结果的可预测性最强,但 LLM 与基线之间的差距也最大。方向性分析表明,LLM 往往依赖于粗略而稳定的预测变量层面先验,这些先验在不同结果和城市之间保持相似,包括对受保护群体预测变量的非对称处理。总体而言,LLM 能够部分地从城市情境中恢复汇总出行模式,但在未经审计实证一致性与潜在偏差的情况下,其预测不应被视为具有结构性依据。
cs.LG / 22 / 2609.00363

Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance

跨GPU核函数的确定性大语言模型(LLM)推理:二的幂INT8量化缩放因子与基于容差的符合性测试的局限
Chen, Teng-Ruei
Abstract
Conformance suites for quantized GEMM kernels ask whether two implementations agree within a tolerance. We measure what such a suite can detect. Injecting nine faults into a reference INT8 pipeline over 8,232 layer--fault--regime cells of Qwen3-1.7B, we find that every one of five epilogue faults -- scale precision, double rounding, multiplication order, output truncation, fused ordering -- moves the output by at most a single bfloat16 spacing, and by exactly one whenever it moves it at all, across 5,880 cells. A tolerance of one spacing is therefore blind to the entire class by construction: four of the five faults are detected by no check in the suite, and the fifth only under power-of-two scales. Faults that violate the accumulator's exactness preconditions, or that break operand sharing, are detected without exception, and a null fault never fires. What a tolerance-based suite of this shape establishes is therefore narrower than interchangeability: that the preconditions hold, that operands are shared, and that differences stay within one spacing. The power-of-two constraint that exposes the one detected fault is also deployable. Requantizing every weight scale to its nearest power of two makes CUTLASS and Triton agree bitwise at every linear layer (196/196 and 252/252, against 8/196 and 10/252 under the checkpoints' own scales) and yields byte-identical generated token sequences at 1.7B, 8B and 14B (8/8 prompts, against 0/8 at all three). Observed perplexity point estimates are +0.32%, -0.28% and +0.48%; the 90% intervals cover zero at the two smaller sizes but not at 14B, reaching +0.71% and +0.76%. A previously reported +157% perplexity for this intervention was an artifact of a probe that rewrote scales without requantizing the weights; separating the effects attributes 99.8% of it to the resulting weight--scale mismatch rather than to the power-of-two constraint itself.
Chinese Translation
量化GEMM核函数的符合性测试套件所检验的,是两个实现是否在某个容差范围内一致。我们测量了此类测试套件能够检测到什么。在一个基于Qwen3-1.7B、覆盖8,232个“层-故障-工况”单元的参考INT8推理流程中注入九种故障后,我们发现五种后处理(epilogue)故障——缩放因子精度、双重舍入、乘法顺序、输出截断、融合计算顺序——在5,880个单元中,每一种对输出的扰动至多只有一个bfloat16的间距(spacing),且只要产生扰动就恰好为一个间距。因此,容差为一个间距的测试从构造上就对整类此类故障视而不见:五种故障中有四种无法被套件中的任何检查检测到,第五种仅在使用二的幂缩放因子时才被检测到。违反累加器精确性前提条件的故障,或破坏操作数共享的故障,则无一例外地被检测到,而空故障(null fault)从不触发。因此,这种形态的基于容差的测试套件所能确立的结论比“可互换性”更窄:它只证明前提条件成立、操作数被共享、且差异保持在间距以内。暴露出唯一被检测故障的二的幂约束同时也是可部署的。将每个权重的缩放因子 requantize 到最近的二的幂,使CUTLASS与Triton在每一个线性层上达到逐位一致(分别为196/196和252/252,而使用检查点自身缩放因子时仅为8/196和10/252),并在1.7B、8B和14B规模上产生字节完全一致的生成token序列(8/8个提示,而三种规模下均为0/8)。观测到的困惑度点估计分别为+0.32%、-0.28%和+0.48%;在两个较小规模上,90%置信区间覆盖零,但在14B上不覆盖,达到+0.71%和+0.76%。此前报道的该干预措施导致+157%困惑度的结果,是一个探针只改写缩放因子而未对权重进行requantize所产生的假象;分离两种效应后可将其中的99.8%归因于由此导致的权重-缩放因子失配,而非二的幂约束本身。
cs.LG / 23 / 2609.00366

Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure

反事实脆弱性证书:揭示结构化证据失效下的高置信度脆弱性
Cenacchi, Filippo, Cao, Longbing, Yang, Runze
Abstract
High test accuracy and good aggregate calibration do not show whether an individual prediction is structurally supported by its evidence. In tabular decision systems, failures often occur when a feature family becomes unavailable, delayed, noisy, stale, or low-trust while the model remains highly confident. Existing calibration, uncertainty, selective-prediction, explanation, and perturbation methods provide scalar scores or attribution maps, but not a recomputable audit object answering: under a declared evidence-failure protocol, what trajectory makes this prediction lose support? We introduce Counterfactual Fragility Certificates (CFC), a model-agnostic protocol-level audit certificate-not a formal robustness certificate-that maps each prediction into an ordered evidence-failure trajectory summarized by greedy flip budget, normalized margin-collapse area, degradation thresholds, and fragility dominance score. Across seven tabular benchmarks and strong linear, tree-based, boosting, and neural baselines, CFC-FDS identifies independently brittle high-confidence cases with 0.915 AUROC, improving over the strongest non-certificate score by +0.405. The advantage persists across perturbation, permutation-importance, group-SHAP, baseline-choice, seed-variance, budgeted-review, and naturalistic field-unavailability checks. Under a 20% review budget, CFC-FDS captures 88.9% of brittle high-confidence cases, compared with 31.8-37.4% for confidence and energy scores. We also evaluate fragility-aware regularization and brittleness-aware temperature correction as secondary uses. CFC provides a concrete reliability framework for exposing high-confidence brittleness missed by ordinary score-centric evaluation.
Chinese Translation
高测试准确率和良好的整体校准并不能表明某个个体预测是否在结构上得到其证据的支持。在表格决策系统中,当某个特征族变得不可用、延迟、含噪、过时或低可信,而模型仍保持高置信度时,往往会发生失效。现有的校准、不确定性、选择性预测、解释和扰动方法只能提供标量分数或归因图,而无法给出一个可重新计算的审计对象来回答:在声明的证据失效协议下,什么样的轨迹会使该预测失去支持?我们提出了反事实脆弱性证书(Counterfactual Fragility Certificates, CFC),一种模型无关的协议级审计证书——而非形式化的鲁棒性证书——它将每个预测映射到一条有序的证据失效轨迹中,并以贪婪翻转预算、归一化裕度塌缩面积、退化阈值和脆弱性支配分数加以概括。在七个表格基准以及强大的线性、树模型、提升方法和神经网络上,CFC-FDS 以 0.915 的 AUROC 识别出相互独立的脆弱高置信度案例,比最强的非证书分数提升了 0.405。该优势在扰动、置换重要性、组 SHAP、基线选择、种子方差、预算化审查以及自然场景下的字段不可用检查中均持续存在。在 20% 的审查预算下,CFC-FDS 能够捕获 88.9% 的脆弱高置信度案例,而置信度分数和能量分数仅能捕获 31.8%–37.4%。我们还评估了脆弱性感知正则化和脆弱性感知温度校正作为次要用途。CFC 提供了一个具体的可靠性框架,用于暴露以分数为中心的常规评估所遗漏的高置信度脆弱性。
cs.LG / 24 / 2609.00374

Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You

无需梯度的自适应:仿射统计量传输及其证书所能揭示的信息
Khazem, Salim, Serouis, Ibrahim Mohamed
Abstract
Test-time adaptation (TTA) typically assumes that model parameters can be updated at inference time. This assumption is restrictive for inference-only accelerators, frozen or third-party models, and memory-constrained deployments, and standard BatchNorm-based TTA configurations may also become inactive on architectures without BatchNorm. We study adaptation when the learned model must remain frozen. We introduce CASTER, a gradient-free method that stores source class statistics in a discriminative subspace, estimates a class-shared affine transformation from target-batch moments, and analytically transports the source class distributions before classification. CASTER requires no backward pass, optimizer state, or stored source feature bank. Across four backbones and seven datasets, it outperforms k-NN on identical frozen features in 27 of 28 backbone-dataset settings while retaining a median of 18x less state. Affine transport is not always reliable. On ImageNet-C, where batches contain only 64 samples for 1000 classes, unconditional transport loses 21.2 top-1 points. We therefore introduce an empirical residual-to-margin transportability certificate. Across 307 evaluation cells, every transport losing more than 10 points has certificate value above 3.9, although benign and destructive regimes are not perfectly separated. Gating converts an average $-3.35$-point effect of unconditional transport into a +1.69-point gain, and performance remains within 0.3 points of the best threshold over a broad threshold range. Finally, we show that this certificate is mechanism-specific: when applied to Tent, it accepts only $4.3\%$ of updates and preserves 0.6% of Tent's available gain. These results position CASTER as a lightweight adaptation mechanism for frozen-model deployment, together with an explicit account of when its safety signal is informative and when it is not.
Chinese Translation
测试时自适应(TTA)通常假设模型参数可以在推理阶段进行更新。这一假设对于仅推理加速器、冻结模型或第三方模型以及内存受限的部署环境而言限制性较强,此外,基于BatchNorm的标准TTA配置在不含BatchNorm的架构上也可能失效。我们研究了在学习到的模型必须保持冻结的情况下如何进行自适应。我们提出了CASTER,这是一种免梯度方法:它将源类统计量存储在一个判别性子空间中,从目标批次矩估计一个类共享的仿射变换,并在分类前以解析方式传输源类分布。CASTER不需要反向传播、优化器状态或存储源特征库。在四个骨干网络和七个数据集上,该方法在28个骨干网络-数据集设置中的27个上优于基于相同冻结特征的k-NN,同时状态存储量中位数仅为后者的1/18。然而,仿射传输并非总是可靠的。在ImageNet-C上,由于每个批次仅含64个样本却需覆盖1000个类别,无条件传输会损失21.2个top-1百分点。为此,我们引入了一种经验性的“残差-间隔可传输性证书”。在307个评估单元中,所有传输损失超过10个百分点的情形其证书值均高于3.9,尽管良性 regime 与破坏性 regime 并未被完全分开。通过门控机制,无条件传输平均带来的-3.35个百分点的负面影响被转化为+1.69个百分点的增益,且在一个较宽的阈值范围内,性能与最优阈值的结果相差不超过0.3个百分点。最后,我们证明该证书是机制特异的:将其应用于Tent时,它仅接受4.3%的更新,只保留了Tent可用增益的0.6%。这些结果表明,CASTER是冻结模型部署场景下的一种轻量级自适应机制,同时还明确说明了其安全信号何时有效、何时无效。
cs.LG / 25 / 2609.00389

Neural means and kernel corrections for operator learning

用于算子学习的神经均值与核修正方法
Shmalo, Yitzchak
Abstract
We combine neural network means with exact Mat\'ern kernel regressions of their residuals and of their learned features, and evaluate the pairing on two public emulation problems with published baselines: the structural-mechanics benchmark of de Hoop et al. and the OCO-2 radiative-transfer emulator of Lamminp\"a\"a et al. On structural mechanics the combination reaches 4.55% test error, matching the best published architecture, and 5.38% against a published 6.49% in the low-data regime. On OCO-2 it improves on the published Gaussian-process emulator on that problem's own test points, outright on two of the three spectral bands; the same kernel that trails the network tenfold on the raw state overtakes it on the network's features, and we measure why (the target's squared native-space norm drops about fortyfold at fixed effective dimension) and prove the mechanism. Where the two families tie instead, the residuals of every architecture we train correlate above 0.86 and their shared component is flat in diversity and sample size, which reads the published plateau as a property of the data. Supporting results include a second-moment identity that predicts stacking outcomes from measured correlations, an optimal-recovery certificate, and a distribution-free coverage band, the only uncertainty signal that survives our tests.
Chinese Translation
我们将神经网络均值与对其残差及其学习特征的精确Matérn核回归相结合,并在两个已发布基线的公开仿真问题上评估该组合:de Hoop等人的结构力学基准问题和Lamminpää等人的OCO-2辐射传输仿真器。在结构力学问题上,该组合达到了4.55%的测试误差,与已发表的最佳架构相当;在低数据量情形下达到5.38%,优于已发表的6.49%。在OCO-2问题上,该方法在该问题自身的测试点上优于已发表的高斯过程仿真器,在三个光谱波段中的两个上全面胜出;在原始状态空间上落后于网络十倍的同一核函数,在网络的特征空间上却反超网络,我们测量了原因(在有效维度固定的情况下,目标在原生空间中的平方范数下降约四十倍)并证明了其机制。当两类方法打平时,我们所训练的所有架构的残差相关性均高于0.86,且其共同成分在多样性和样本量上保持平坦,这表明已发表的瓶颈是数据本身的属性。辅助结果包括:一个可从测得相关性预测堆叠结果的二阶矩恒等式、一个最优恢复证书,以及一个无分布覆盖带——这是唯一在我们所有测试中均有效的 uncertainty(不确定性)信号。
cs.LG / 26 / 2609.00403

A Multi-Branch Feature Fusion Approach for Health Misinformation Detection and Propagation

一种用于健康谣言检测与传播的多分支特征融合方法
Sikosana, Mkululi, Maudsley-Barton, Sean, Ajao, Oluwaseun
Abstract
This paper presents a multi-branch fusion framework for detecting and characterising the propagation of health misinformation in online social networks (OSNs). Grounded in the Elaboration Likelihood Model (ELM) and the Theory of Planned Behaviour (TPB), the model fuses transformer-based semantics with rhetorical cues, stance representations, and psychologically motivated proxies in a unified multi-task architecture. In addition to binary classification, we introduce the Cognitive Propagation Score (CPS), an interpretable post-hoc auxiliary score computed from psychologically motivated, text-derived cues capturing argument complexity, emotional intensity, and content-derived virality potential, to support diffusion-risk reasoning when engagement ground truth is incomplete or unavailable. Experiments on three benchmark datasets, Constraint, COVID--19\_FNIR, and Monkeypox, show strong classification performance, achieving ROC--AUC up to 0.9999 on COVID--19\_FNIR, while propagation-oriented ranking achieves near-perfect agreement when engagement-derived supervision is available (Monkeypox, Spearman's $\rho = 0.9952$) and similarly high ranking alignment under proxy-based supervision on COVID--19\_FNIR ($\rho = 0.9954$). Compared with representative literature baselines, the fusion model improves detection on Constraint and COVID--19\_FNIR, while Monkeypox remains more challenging, reflecting domain- and signal-specific differences. Ablation analysis further indicates that psychological and rhetorical branches provide complementary gains beyond semantic embeddings. Overall, the framework bridges cognitive theory and neural modelling to improve transparency and to support scalable misinformation monitoring, with future work required to validate CPS against human-centred diffusion judgements.
Chinese Translation
本文提出了一种多分支融合框架,用于检测在线社交网络(OSNs)中的健康谣言并刻画其传播特征。该模型基于详尽可能性模型(Elaboration Likelihood Model, ELM)和计划行为理论(Theory of Planned Behaviour, TPB),在统一的多任务架构中融合了基于Transformer的语义表示、修辞线索、立场表征以及心理学动机衍生的代理指标。除二分类任务外,我们引入了认知传播分数(Cognitive Propagation Score, CPS)——一种可解释的事后辅助分数,其由捕捉论辩复杂度、情感强度以及内容传播潜力的心理学动机、文本衍生线索计算得到,用于在互动数据的真实标注不完整或不可用时支持传播风险推理。在Constraint、COVID-19_FNIR和Monkeypox三个基准数据集上的实验表明,该模型具有强大的分类性能,在COVID-19_FNIR上ROC-AUC最高达0.9999;在可获取互动数据监督的情况下,面向传播的排序结果接近完全一致(Monkeypox数据集,Spearman's ρ = 0.9952),在COVID-19_FNIR上基于代理监督的排序同样具有高度一致性(ρ = 0.9954)。与文献中的代表性基线相比,融合模型在Constraint和COVID-19_FNIR数据集上提升了检测效果,而Monkeypox数据集仍更具挑战性,这反映了领域和信号层面的特定差异。消融分析进一步表明,心理学分支和修辞分支在语义嵌入之外提供了互补的增益。总体而言,该框架将认知理论与神经建模相结合,以提升透明度并支持可扩展的谣言监测,未来工作需要将CPS与以人为中心的传播判断进行对照验证。
cs.LG / 27 / 2609.00420

How Temporal Correlations Shape Memory in Linear Recurrent Neural Networks

时间相关性如何塑造线性循环神经网络中的记忆
Fokam, Arnol Manuel, Akpevwoghene, Fasseu Sieyondji, Dawson, Edem Fiifi
Abstract
The linear recurrent neural network (LRNN) is a simple model for studying how much memory a network builds up as it trains. For uncorrelated inputs, earlier work found that training itself settles the network between keeping the past and reacting only to the present. Real sequences are correlated, and we solve the learning dynamics exactly for correlated inputs. In the solution, keeping the past carries a cost. The whole effect of correlation lands on that cost. This cost reduces to the earlier one when inputs are uncorrelated and grows once they are positively correlated. Three findings follow. (1) Correlation reshapes the course of learning, not only its end. Memory builds, overshoots, and is partly removed, and the settled network keeps less of the past. (2) Memory switches off at a threshold set by one number, how much each input resembles the one just before it. Neither sequence length nor longer-range correlation moves this threshold. Memory is worth keeping only when the task needs the previous input more than the current input already supplies it through correlation with the past. (3) The best network changes too. Zero error demands a feedthrough, a path that passes the current input straight to the network's output and remembers nothing, and training builds it unprompted when given one spare hidden dimension. Our work turns one property of the input into a prediction of whether a network learns memory and explains why correlated data turns recurrent networks into change detectors.
Chinese Translation
线性循环神经网络(LRNN)是研究网络在训练过程中能积累多少记忆的一个简单模型。对于不相关的输入,此前的研究发现训练本身会使网络在保留过去与只对当下做出反应之间达到某种平衡。真实的序列是相关的,我们针对相关输入精确求解了学习动力学。在解中,保留过去是有代价的,而相关性的全部影响都落在这一代价上。当输入不相关时,该代价退化为先前的代价;当输入正相关时,该代价增大。由此得到三个发现:(1)相关性重塑了学习的整个过程,而不仅仅是其终点:记忆先建立、后超出,随后被部分移除,最终稳定的网络保留更少的过去信息。(2)记忆在一个阈值处关闭,该阈值由一个数决定,即每个输入与前一个输入的相似程度。序列长度或更长程的相关性都不会移动这一阈值。只有当任务对上一个输入的需求超过当前输入通过与过去的相关性所提供的信息时,记忆才值得保留。(3)最优网络也发生了变化:零误差要求一条直通路径,即把当前输入直接传递到网络输出而不保存任何记忆的通路;在给定一个多余隐维的情况下,训练会自发地构建这条路径。我们的工作将输入的一个性质转化为对网络是否学习记忆的预测,并解释了为什么相关数据会使循环网络变成变化检测器。
cs.LG / 28 / 2609.00444

Group Adaptive Clipping Policy Optimization

分组自适应裁剪策略优化
Jia, Sheng, Wang, Xiao, Kasiviswanathan, Shiva Prasad, Houthooft, Rein
Abstract
Group relative policy optimization for reinforcement learning with verifiable rewards (RLVR) typically uses a fixed importance-sampling (IS) ratio clipping boundary across all rollouts. We identify a key limitation: rare correct rollouts on harder problems and abundant correct rollouts on easier problems are clipped at comparable rates, despite contributing very different learning signals. Rollouts with low group success exhibit larger IS ratios and carry stronger gradient signal for exploration and solving new problems, yet are disproportionately suppressed by fixed clipping. To address this, we propose Group Adaptive Clipping Policy Optimization (GAPO), a plug-in modification to GRPO methods that adapts the clipping boundary to the rollout advantage. GAPO is motivated by a reverse-KL trust-region perspective, which suggests that rollouts with larger learning signal should receive proportionally greater update headroom. GAPO requires no reward shaping and preserves the standard PPO/GSPO surrogate while adapting only the clipping threshold. Across Qwen and Llama models, GAPO consistently improves both Pass@1 and Pass@k over fixed clipping and advantage-shaping baselines on math reasoning and coding benchmarks where the pass rates by the base model are relatively low.
Chinese Translation
针对可验证奖励的强化学习(RLVR),分组相对策略优化(GRPO)通常在所有 rollout 中使用固定的重要性采样(IS)比率裁剪边界。我们发现了一个关键局限:在较难问题上稀少的正确 rollout 和在较易问题上大量出现的正确 rollout 以相近的速率被裁剪,尽管它们所贡献的学习信号差异巨大。分组成功率低的 rollout 表现出更大的 IS 比率,并携带更强的用于探索和解决新问题的梯度信号,却在固定裁剪下受到不成比例的抑制。为解决这一问题,我们提出了分组自适应裁剪策略优化(Group Adaptive Clipping Policy Optimization,GAPO),这是一种针对 GRPO 方法的即插即用改进,能够根据 rollout 的优势自适应地调整裁剪边界。GAPO 的动机来自反向 KL 置信域视角,该视角表明具有较大学习信号的 rollout 应获得成比例的更大更新空间。GAPO 无需奖励塑形,并在仅自适应调整裁剪阈值的同时保留了标准的 PPO/GSPO 代理目标函数。在 Qwen 和 Llama 模型上,在基座模型通过率相对较低的数学推理与代码基准上,GAPO 相较于固定裁剪和优势塑形基线,持续提升了 Pass@1 和 Pass@k 性能。
cs.LG / 29 / 2609.00446

CRAD: Class-wise Reliability-Aware Distillation for Decentralized Heterogeneous Federated Learning

CRAD:面向去中心化异构联邦学习的类别级可靠性感知蒸馏
Bilbeisi, Baraa, Fan, Mengchen, Geng, Baocheng, Tian, Qing
Abstract
Conventional federated learning (FL) relies on parameter averaging, which forces clients to be doubly homogeneous: it demands an identical architecture and degrades under non-IID data. Real-world deployments usually break both assumptions. We sidestep both by building a decentralized knowledge distillation framework in which each client evaluates its peers' model snapshots on its own local data and distills from the resulting soft predictions. Because knowledge is transferred through the shared class posterior, clients are free to run different architectures; and because every teacher is evaluated on the student's own device, raw data never leaves the client, with no central server or public dataset required. Within this setting, we identify and address an under-examined problem: how to combine the peer teacher predictions. Existing methods, like uniform averaging, ignore how knowledge reliability varies across teachers and classes. We propose Class-wise Reliability-Aware Distillation (CRAD), which, per class, first discards teachers that disagree with the peer consensus and then takes a weighted average of the rest, weighting each teacher by its per-class reliability (precision, or inverse variance). Since the variance of an accuracy from $n$ samples scales as $1/n$, support enters automatically: among the teachers that survive filtering, a teacher is trusted for a class to the degree that it is both accurate and well-evidenced for it. On three image-classification benchmarks (CIFAR-10, CIFAR-100, and PathMNIST colon pathology), across heterogeneous architectures under severe non-IID skew, CRAD consistently outperforms competing methods in global accuracy.
Chinese Translation
传统的联邦学习(FL)依赖于参数平均,这要求客户端具备双重同质性:既要求相同的模型架构,又在非独立同分布(non-IID)数据下性能下降。而现实世界的部署场景通常无法同时满足这两个假设。我们通过构建一个去中心化的知识蒸馏框架来规避这两个限制:在该框架中,每个客户端在其自身的本地数据上评估其同伴的模型快照,并从所得的软预测中进行蒸馏。由于知识是通过共享的类别后验概率进行传递的,客户端可以自由使用不同的架构;同时,由于每位教师模型都是在学生端自身的设备上进行评估的,原始数据永远不会离开客户端,且无需中央服务器或公共数据集。在此设定下,我们识别并解决了一个尚未被充分研究的问题:如何融合同伴教师模型的预测。现有方法(如均匀平均)忽略了知识可靠性在不同教师和不同类别之间的差异。我们提出了类别级可靠性感知蒸馏(Class-wise Reliability-Aware Distillation,CRAD),该方法针对每个类别,首先剔除与同伴共识不一致的教师模型,然后对其余教师模型进行加权平均,每个教师的权重由其该类别上的可靠性(精确率,即方差的倒数)决定。由于基于 $n$ 个样本的准确率方差随 $1/n$ 缩放,样本支持度被自动纳入考量:在通过筛选的教师中,某教师在某一类别上的可信度取决于其在该类别上同时具备高准确性和充分的证据支持。在三个图像分类基准(CIFAR-10、CIFAR-100 和 PathMNIST 结肠病理数据集)上,面对异构架构和严重的非独立同分布(non-IID)偏斜,CRAD 在全局准确率上始终优于对比方法。
cs.LG / 30 / 2609.00450

HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference

HBQ:面向硬件高效设计与高精度大语言模型推理的分层分块量化方法
Chen, Chun-Ting, Han, Dongmin, Mun, Hangyeol, Hyun, Jake, Raha, Arnab, Agarwal, Amit, Anders, Mark, Abdelfattah, Mohamed, Seo, Jae-sun
Abstract
Block Quantization (BQ) is a promising approach for efficient deployment of large language models (LLMs), enabling low-precision computation with controlled accuracy degradation. Compared to scalar weight-only quantization (WoQ), BQ quantizes both weight and activation, offering higher hardware efficiency and end-to-end inference on a unified datapath, but its design space, spanning bit-width, block size, scaling, and numeric formats, remains underexplored. We provide hardware/benchmark results through design space exploration (DSE). We find that increasing block size improves hardware efficiency by amortizing dequantization and accumulation costs, but degrades accuracy. This trade-off limits conventional BQ methods. Motivated by this insight, we propose Hierarchical Block Quantization (HBQ). Unlike prior methods [1], [2], which use small blocks and conventional Power-of-Two (PoT) or integer-based scaling, HBQ uses large blocks to maximize efficiency and introduces low-overhead significand (SIG) scaling for second-level quantization. By allocating quantization levels effectively and accounting for distinct activation and weight distributions, SIG scaling compensates for large-block errors more effectively than prior PoT and INT schemes. HBQ-A (accurate) achieves W4A16-level accuracy using only W4A5 while requiring less silicon area than NVFP4. HBQ-E (efficient) further reduces hardware cost by 17% while maintaining higher accuracy than all existing BQ methods. We implemented a 28nm ASIC accelerator applying HBQ to weights, activations, and KV cache, and integrated a novel partial-sum BQ scheme to further reduce EMA energy. Compared to state-of-the-art WoQ, HBQ delivers $2.3\times$/$4.6\times$ higher area/energy efficiency at the same accuracy level; $1.6$--$3.3\times$ system energy reduction and $1.5$--$3.0\times$ speedup over prior BQ methods while providing best accuracy.
Chinese Translation
分块量化是一种高效部署大语言模型(LLM)的有前景的方法,能够在可控精度损失下实现低精度计算。与仅标量权重量化相比,BQ同时对权重和激活进行量化,提供更高的硬件效率并可在统一数据通路上实现端到端推理,但其设计空间(涵盖位宽、块大小、缩放方式和数值格式)仍未得到充分探索。我们通过设计空间探索(DSE)提供了硬件/基准测试结果。我们发现,增大块大小可通过摊销反量化与累加成本来提升硬件效率,但会降低精度。这一权衡限制了传统BQ方法。基于这一洞察,我们提出分层分块量化。与以往使用小块及传统的2的幂或整数缩放的方法[1][2]不同,HBQ使用大块以最大化效率,并引入低开销的尾数缩放用于二级量化。通过有效分配量化级别并考虑激活与权重分布的差异,SIG缩放比以往的PoT和INT方案更有效地补偿了大块带来的误差。HBQ-A(高精度)版本仅使用W4A5即达到W4A16级别的精度,同时所需硅面积少于NVFP4。HBQ-E(高效率)版本在保持高于所有现有BQ方法精度的同时,进一步将硬件成本降低17%。我们实现了一款28nm ASIC加速器,将HBQ应用于权重、激活和KV缓存,并集成了一种新颖的部分和分块量化方案以进一步降低EMA能耗。与最先进的WoQ相比,HBQ在相同精度水平下提供2.3倍/4.6倍更高的面积/能效;相比以往的BQ方法,在提供最佳精度的同时实现1.6–3.3倍的系统能耗降低和1.5–3.0倍的加速。
cs.LG / 31 / 2609.00457

Can LLMs Use Relational Transformer Embeddings?

大语言模型能否使用关系型Transformer嵌入?
Azevedo, Francisco Galuppo, Loures, Clarissa Lima
Abstract
Injecting frozen relational-encoder embeddings as soft tokens into a large language model (LLM) is a conceptually appealing fusion strategy: the encoder handles multi-table structure, the LLM handles language and reasoning, and no lossy text serialization is required. We test this hypothesis concretely by injecting embeddings from a frozen Relational Transformer (RT) into Qwen3.5-4B via a learned MLP projection and LoRA adaptation, trained first with supervised fine-tuning (SFT) on chain-of-thought reasoning traces and then with group-based reinforcement learning (GSPO). We evaluate across 10 binary classification tasks on 6 relational databases from RelBench, under four supervision regimes: single-task (ST), within-dataset (WD), cross-dataset (CD), and all-task (ALL). The hybrid model does not consistently outperform standalone RT: it is frequently below random, highly sensitive to serialization format and relational-token budget, and unstable under RL training. We report these negative results and analyze the failure modes, arguing that soft-token fusion requires stronger alignment objectives and schema-aware design before it can serve as a reliable route to relational prediction.
Chinese Translation
将冻结的关系型编码器嵌入作为软令牌(soft tokens)注入大语言模型(LLM)是一种在概念上颇具吸引力的融合策略:编码器负责处理多表结构,LLM负责处理语言和推理,且无需进行有损的文本序列化。我们通过具体实验验证了这一假设:将冻结的Relational Transformer(RT)生成的嵌入,经由一个可学习的MLP投影层和LoRA适配注入Qwen3.5-4B模型,先在思维链推理轨迹上进行有监督微调(SFT),随后采用基于分组的强化学习(GSPO)进行训练。我们在来自RelBench的6个关系型数据库的10个二分类任务上,于四种监督设置下进行评估:单任务(ST)、数据集内(WD)、跨数据集(CD)和全任务(ALL)。实验结果表明,该混合模型并未持续超越独立的RT模型:其表现常常低于随机水平,对序列化格式和关系令牌预算高度敏感,且在RL训练下不稳定。我们报告了这些负面结果并分析了其失败模式,认为软令牌融合需要更强的对齐目标和模式感知(schema-aware)设计,才能成为实现关系型预测的可靠路径。
cs.LG / 32 / 2609.00460

Context Window Failures in Relational Foundation Models

关系型基础模型中的上下文窗口失效问题
Correa, Denis Oliveira, Azevedo, Francisco Galuppo
Abstract
Recent Relational Deep Learning architectures have been proposed as foundation models for multi-table relational data, yet they impose constrained neighborhood budgets that force row truncation when an entity has many related records. We introduce Animus, a synthetic financial dataset in which predicting customer income requires aggregating up to tens of thousands of transactions. On the raw representation, three recently proposed models (RT, Griffin, RelGT) achieve $R^2 \le 0.18$; a single, routine, temporal pre-aggregation step recovers $R^2$ up to $0.65$. This questions whether current relational foundation models are ready for high-cardinality real-world data.
Chinese Translation
近期提出的关系型深度学习架构(Relational Deep Learning)被视为多表关系型数据的基础模型,然而它们施加了受限的邻域预算,当实体拥有大量相关记录时,这迫使其对行进行截断。我们引入了 Animus——一个合成金融数据集,其中预测客户收入需要聚合多达数万条交易记录。在原始表示下,三种近期提出的模型(RT、Griffin、RelGT)的 $R^2 \le 0.18$;而仅通过一步常规的时间维度预聚合操作,$R^2$ 即可提升至 0.65。这不禁让人质疑:当前的关系型基础模型是否已经能够应对高基数(high-cardinality)的真实世界数据。
cs.LG / 33 / 2609.00472

Higher Structures in Deep Learning

深度学习中的高阶结构
Roberts, Michael L., Cooper, Carlos Zapata Carratalá. Nicholas J., Chen, Lijun, Meyer, François G., Gurari, Danna
Abstract
We provide an expository introduction on the importance of higher-arity tensor operations to deep learning. Then, we conduct a novel empirical investigation of higher-arity phenomenon in trained neural networks, introduce a hypergraphical generalization of the multilayer perceptron, and explore connections to evolutionary algorithms. We conclude with a discussion of promising directions for future research.
Chinese Translation
本文首先对多元高阶张量运算在深度学习中的重要性进行了阐述性介绍。随后,我们对训练后的神经网络中的高阶现象进行了一项新颖的实证研究,提出了多层感知机(multilayer perceptron)的超图推广形式,并探讨了其与进化算法之间的联系。最后,我们讨论了未来研究中值得关注的方向。
cs.LG / 34 / 2609.00488

AdaptNTK: Adaptive Uncertainty Quantification and Active Learning for Neural Network Potentials

AdaptNTK:面向神经网络势的自适应不确定性量化与主动学习
Ananth, Prajwal, Yue, Shuwen
Abstract
Machine learning interatomic potentials bridge the gap between quantum chemical precision and classical computational speed, enabling molecular dynamics simulations with first-principles accuracy. Their reliability is often improved through active learning, which iteratively expands the training set by identifying uncertain, out-of-distribution configurations. Existing uncertainty-quantification methods often involve a trade-off between computational cost and reliability, and generally cannot account for redundancy as an acquisition batch is assembled. Here, we introduce AdaptNTK, a single-model framework that measures uncertainty as a regularized Mahalanobis distance in empirical neural tangent kernel (NTK) feature space. With the NTK features fixed during acquisition, the uncertainty depends on the acquired configurations but not their reference labels. This allows the uncertainty to be updated recursively after each selection without retraining, reducing redundancy within an acquisition batch. On held-out rMD17 data, AdaptNTK achieves the highest mean correlations with force errors (Spearman 0.68, Pearson 0.71) and matches a three-member ensemble in error retention. In active learning experiments, AdaptNTK achieves the lowest force errors across rMD17 and Transition-1X, with particularly strong performance on transition-state configurations in Transition-1X. AdaptNTK provides a 2.6-fold speedup per Transition-1X cycle relative to the ensemble, providing efficient single-model uncertainty estimation with sequential updates for data-efficient active learning.
Chinese Translation
机器学习原子间势在量子化学精度与经典计算速度之间架起了桥梁,使具有第一性原理精度的分子动力学模拟成为可能。其可靠性通常通过主动学习来提升,即通过识别不确定的、超出分布的构型来迭代扩展训练集。现有的不确定性量化方法往往在计算成本与可靠性之间存在权衡,且通常无法在构建采集批次时考虑冗余问题。本文提出AdaptNTK,一个单模型框架,它将不确定性度量定义为经验神经正切核(NTK)特征空间中的正则化马氏距离。由于NTK特征在采集过程中保持固定,不确定性依赖于已采集的构型而与其参考标签无关。这使得不确定性可以在每次选择后递归更新而无需重新训练,从而降低采集批次内的冗余。在留出的rMD17数据上,AdaptNTK与力误差的平均相关性最高(Spearman 0.68,Pearson 0.71),并在误差保留方面与三成员集成模型相当。在主动学习实验中,AdaptNTK在rMD17和Transition-1X上均取得了最低的力误差,尤其在Transition-1X的过渡态构型上表现尤为出色。相较于集成模型,AdaptNTK在Transition-1X每个循环中实现了2.6倍的加速,为数据高效的主动学习提供了具有序列更新能力的高效单模型不确定性估计。
cs.LG / 35 / 2609.00489

A hybrid quantum-classical neural network for learning to route

一种用于学习路径规划(Learning to Route)的量子-经典混合神经网络
Ritt, Marcus Rolf Peter, Júnior, Alexsandro Santos da Rosa, Reballo, Marcos Vinicius, Amaral, Cesar Augusto do, de Barros, Fernando Augusto Caletti
Abstract
This work studies hybrid quantum-classical neural networks for learning routing heuristics. Specifically, this paper asks whether small quantum neural networks can replace parameter-heavy modules inside a competitive attention-based routing model while maintaining solution quality. For the capacitated vehicle routing problem, encoder feed-forward replacement emerges as the most promising design: it reduces the number of model parameters by 56.6% while keeping the hybrid model close to the classical neural baseline at small and medium instance sizes, although the gap grows for larger instances. This work also compares to classical routing algorithms, which remain highly competitive and often superior on the fixed Euclidean test sets. Our results therefore do not indicate quantum advantage or solver dominance, but identify encoder feed-forward replacement as a viable hybrid-module compression strategy for neural combinatorial optimization.
Chinese Translation
本工作研究了用于学习路径规划启发式方法的量子-经典混合神经网络。具体而言,本文探讨了一个问题:小型量子神经网络能否在保持解质量的前提下,替代一个具有竞争力的基于注意力机制的路径规划模型中参数量较大的模块。针对带容量约束的车辆路径问题(Capacitated Vehicle Routing Problem),编码器前馈层替换被证明是最有前景的设计方案:它使模型参数量减少了56.6%,同时在中小规模实例上使混合模型的性能接近经典神经基线,尽管在更大规模实例上性能差距会有所扩大。本工作还与经典路径规划算法进行了比较,结果显示后者仍具有很强的竞争力,并且在固定的欧几里得测试集上往往表现更优。因此,我们的结果并未表明量子优势或求解器的支配地位,而是将编码器前馈层替换确定为神经组合优化中一种可行的混合模块压缩策略。
cs.LG / 36 / 2609.00507

VATO: A Vortex-Force-Aware Transformer Operator for Unsteady Separated Aerofoil Flows

VATO:一种面向非定常分离翼型流动的涡力感知Transformer算子
Yang, Xingxin, Zhang, Zhan, Li, Yichen, Li, Juan
Abstract
Accurate prediction of unsteady separated flows is challenging because the aerodynamic loads depend on nonlinear separation and vortex-shedding dynamics. Although high-fidelity CFD resolves these mechanisms, its cost limits repeated use in design and control. Standard field-level surrogate training, however, does not distinguish the flow regions that contribute most strongly to the aerodynamic loads. We introduce VATO (Vortex-Force-Aware Transformer Operator), which couples the Vortex Force Map (VFM) method to a geometry-aware neural operator through two complementary mechanisms. VATO-S adds training-only supervision of the local VFM force-contribution field, with no increase in model size or inference cost. VATO-A uses VFM contribution and sensitivity fields to prioritise force-relevant source locations for residual cross attention. The methods are evaluated on unsteady CFD data for double-edged-plate aerofoils over 54 trajectories from nine geometries. Over lead times of 1-20~ms, VATO-S reduces velocity, pressure, and vorticity errors by 10.4\%, 1.0\%, and 15.6\%, respectively, while VATO-A achieves reductions of 15.8\%, 7.5\%, and 31.2\%. VATO-S gives the lowest VFM-derived drag error, whereas VATO-A gives the lowest pressure-derived lift and drag errors. Over lead times extending 50\% beyond the training range, VATO-A retains a 26.9\% reduction in vorticity error and larger improvements in all four force readouts, despite reduced gains in velocity and pressure. These results show that force-aware operator learning can improve both flow-field prediction and aerodynamic functional accuracy in unsteady separated flows.
Chinese Translation
非定常分离流动的精确预测具有挑战性,因为气动载荷依赖于非线性的分离与涡脱落动力学。尽管高精度CFD能够解析这些机制,但其计算成本限制了在设计与控制中的反复使用。然而,标准的场级代理模型训练无法区分对气动载荷贡献最强的流动区域。我们提出了VATO(Vortex-Force-Aware Transformer Operator,涡力感知Transformer算子),通过两个互补机制将涡力图(Vortex Force Map, VFM)方法与几何感知神经算子相耦合。VATO-S仅在训练阶段增加对局部VFM力贡献场的监督,不增加模型规模或推理成本。VATO-A利用VFM贡献场和敏感度场对与力相关的源位置进行优先排序,用于残差交叉注意力机制。两种方法在来自九种几何外形的54条轨迹的双缘板翼型非定常CFD数据上进行了评估。在1-20毫秒的超前时间内,VATO-S将速度、压力和涡量误差分别降低了10.4%、1.0%和15.6%,而VATO-A则分别实现了15.8%、7.5%和31.2%的降低。VATO-S给出的VFM推导阻力误差最低,而VATO-A给出的压力推导升力和阻力误差最低。在超出训练范围50%的超前时间内,尽管速度和压力的改进有所减弱,VATO-A仍保持了26.9%的涡量误差降低,并且在全部四个力读数上取得了更大的提升。这些结果表明,力感知算子学习能够同时提升非定常分离流动中的流场预测精度和气动功能精度。
cs.LG / 37 / 2609.00518

Learning Task-Specific Antibody Representations via Function-Aware Masking

通过功能感知掩码学习任务特异性抗体表示
Goel, Ayan, Walton, Thomas A., Aghazadeh, Amirali
Abstract
Antibody-specific language models pretrained via masked language modeling (MLM) learn representations that are critical for downstream sequence design and property prediction tasks. Yet, the corruption process itself is rarely leveraged as a source of inductive bias during pretraining. While preferentially masking complementarity-determining regions (CDRs) improves binding-related predictions, antibodies possess diverse biological priors over a variety of functions. Herein, we introduce function-aware masking, a family of pretraining algorithms that align mask placement with specific functional priors (e.g., from IMGT annotations or structure predictions) to shape the learned representation space. We show that these specialist masking strategies significantly improve performance on their respective objectives, yielding up to a 14% gain on structure-related tasks and up to a 5.9x improvement on CDR-related tasks. To further improve performance across multiple functional axes, we develop hybrid masking strategies that integrate multiple priors, balancing reconstruction over binding, structural, and biophysical objectives. Our results demonstrate that informed mask placement provides a parameter-free mechanism for imposing functional inductive biases in antibody language model training.
Chinese Translation
通过掩码语言建模(MLM)预训练的抗体专用语言模型能够学习对下游序列设计和性质预测任务至关重要的表示。然而,掩码破坏过程本身在预训练中很少被用作归纳偏置的来源。虽然优先掩码互补决定区(CDR)可以改善与结合相关的预测,但抗体在多种功能上具有多样的生物学先验。本文中,我们提出了功能感知掩码(function-aware masking),这是一系列预训练算法,通过将掩码位置与特定的功能先验(例如来自IMGT注释或结构预测)对齐来塑造所学习的表示空间。我们证明,这些专用掩码策略在各自的目标上显著提升了性能,在结构相关任务上最高带来14%的提升,在CDR相关任务上最高带来5.9倍的改进。为了进一步提升在多个功能维度上的表现,我们开发了整合多种先验的混合掩码策略,在结合、结构和生物物理目标之间平衡重建。我们的结果表明,合理的掩码位置选择为在抗体语言模型训练中施加功能性归纳偏置提供了一种无需额外参数的机制。
cs.LG / 38 / 2609.00528

Why Multi-Layer Message Passing Works: Completeness Theory for Graph Neural Network Interatomic Potentials

多层消息传递为何有效:图神经网络原子间势的完备性理论
Ming, Pingbing, Wang, Han
Abstract
We prove that the Hypergraph Neural Network, an invariant architecture with 3-body message passing, is a universal approximator for potential energy surfaces. Our main contribution is a multi-layer completeness theory. We show that $L$ layers of message passing on sparse, cutoff-based graphs achieve the same representational power as having access to the full $L$-hop neighborhood, provided the configurations are generic, satisfy an overlap condition and a connectivity condition. This provides the first rigorous justification for the common practice of using multi-layer message passing with a per-layer cutoff smaller than the physical interaction range, the setting used by virtually all practical graph neural network based machine-learned interatomic potentials. As immediate consequences, we show that both DPA3 and CHGNet architectures inherit universal approximation.
Chinese Translation
我们证明了超图神经网络(Hypergraph Neural Network)——一种具有三体消息传递的不变(invariant)架构——是势能面的通用逼近器。我们的主要贡献是提出了一种多层完备性理论。我们证明,只要分子构型是通用的、满足重叠条件和连通性条件,在基于截断半径构建的稀疏图上进行 $L$ 层消息传递,即可达到与获取完整 $L$ 跳邻域信息相同的表示能力。这为常见的实践做法——即每层截断半径小于物理相互作用范围的情况下使用多层消息传递——首次提供了严格的数学证明,而这一设定实际上被几乎所有实用的基于图神经网络的机器学习原子间势所采用。作为直接推论,我们证明了 DPA3 和 CHGNet 架构均继承了这种通用逼近性质。
cs.LG / 39 / 2609.00530

DeSyR: A Decoupled Symbolic Recovery Framework with PINN-Guided Structure Search and Physics-Informed Coefficient Refinement

DeSyR:一种基于PINN引导结构搜索与物理信息系数精化的解耦符号恢复框架
Niu, Pancheng, Guo, Jun, He, Qiaolin, Guo, Jingcai, Shi, Yanchao
Abstract
Recovering compact explicit solutions from neural approximations is challenging when imperfect teacher data guide symbolic topology search and coefficient estimation. We present DeSyR, a decoupled symbolic recovery framework for differential equations. A physics-informed neural network guides repeated searches to construct candidate topologies with provisional constants. Once a topology is fixed, its coefficients are refined solely from the governing equation and prescribed constraints, followed by gated selection and verification. For linear fixed-topology parameterizations, we characterize teacher-error inheritance and show that finite-weight mixed data--physics fitting retains an $O(\beta^{-1})$ teacher-dependent contribution when the teacher error projects onto the model space. Under well-posedness, representability, zero-residual attainment, and discrete determinacy, physics-only refinement conditionally recovers exact coefficients; for nonlinear parameterizations, the corresponding guarantees are local. DeSyR is evaluated on 15 differential-equation problems across 18 configurations covering high-order, space--time, multidimensional, nonlinear, and coupled systems. A candidate-level audit yields a 99.23% convergence rate among free-parameter refits, while every selected refinement involving free coefficients converges. Configuration-level median refined relative $L_2$ errors are $2.31\times10^{-14}$ or lower. In same-topology comparisons, refinement reduces error by eight to fourteen orders of magnitude. These results show that an approximate neural teacher can guide topology discovery without imposing its error scale on final recovered coefficients, provided a target-capable topology is retained and physics-only refinement converges.
Chinese Translation
当不完美的教师数据引导符号拓扑搜索与系数估计时,从神经网络近似中恢复紧凑的显式解极具挑战性。我们提出DeSyR,一种针对微分方程的解耦符号恢复框架。物理信息神经网络(PINN)引导反复搜索,以构建带有临时常数的候选拓扑。一旦拓扑固定,其系数仅从控制方程和给定约束出发进行精化,随后通过门控选择与验证。对于线性固定拓扑参数化,我们刻画了教师误差的继承特性,并证明当教师误差投影到模型空间时,有限权重的数据—物理混合拟合仍保留 $O(\beta^{-1})$ 的教师相关贡献。在适定性、可表示性、零残差可达性与离散确定性条件下,纯物理精化可有条件地恢复精确系数;对于非线性参数化,相应保证仅在局部成立。DeSyR在涵盖高阶、时空、多维、非线性及耦合系统的15个微分方程问题、18种配置上进行了评估。候选层面的审计显示,自由参数重拟合的收敛率为99.23%,且所有涉及自由系数的选定精化均收敛。配置层面的精化相对 $L_2$ 误差中位数为 $2.31\times10^{-14}$ 或更低。在相同拓扑的比较中,精化使误差降低八到十四个数量级。这些结果表明,只要保留了具备目标表达能力的拓扑且纯物理精化收敛,近似的神经教师即可引导拓扑发现,而不会将其误差量级强加于最终恢复的系数。
cs.LG / 40 / 2609.00544

GenONet: A Generative operator Network for High-Resolution Precipitation Nowcasting

GenONet:一种用于高分辨率降水临近预报的生成算子网络
Golkar, Mohammad Kian, de Oliveira, Luciano Alves, Khanjani, Mohammad
Abstract
High-resolution precipitation nowcasting is critical for reducing the impacts of severe weather but remains difficult because of rapid storm evolution. Deep learning models have shown great promise for this task, but their predictive skill often deteriorates over longer forecast horizons. This leads to increasingly blurry forecasts that fail to capture the complex, non-linear evolution of storm systems. In order to address these limitations, we introduce Spatio-Temporal U-DeepONet (GenONet), a novel architecture for long-range precipitation forecasting up to 3 hours, specifically designed to produce sharp and physically consistent results. GenONet's architecture pioneers the use of a Deep Operator Network (DeepONet) as a generator within a Generative Adversarial Network (GAN) framework for this task. The DeepONet learns the continuous-time dynamics of precipitation, ensuring stability over long forecast horizons. Adversial training against a spatio-temporal discriminator compels the model to produce sharp, coherent forecasts, while a physics-informed loss regularizer, derived from the Moisture Conservation Equation, improves physical plausibility in our ablation setting. Quantitative evaluations show that our model achieves consistently higher scores on most of the metrics, especially for highintensity events and at longer lead times. Qualitatively, GenONet produces structurally coherent forecasts that maintain their integrity, whereas baseline models degrade into indistinct patterns. Finally, an ablation study confirms the benefit of this physics-informed loss, highlighting the strength of combining operator learning with adversarial training.
Chinese Translation
高分辨率降水临近预报对于减轻强对流天气的影响至关重要,但由于风暴演变迅速,该任务仍然极具挑战性。深度学习模型在这一任务中展现出巨大潜力,但其预测能力往往随预报时效的延长而下降,导致预报结果日益模糊,无法捕捉风暴系统复杂的非线性演变过程。为解决这些局限性,我们提出了时空U-DeepONet(GenONet),这是一种面向长达3小时长时效降水预报的新型架构,专门设计用于生成清晰且物理一致的结果。GenONet的架构创新性地将深度算子网络(DeepONet)作为生成器引入生成对抗网络(GAN)框架中。DeepONet学习降水的连续时间动力学特性,从而确保长预报时效下的稳定性。通过与时空判别器进行对抗训练,模型能够生成清晰、连贯的预报结果;同时,基于水分守恒方程构建的物理信息损失正则化项,在消融实验设置中提高了结果的物理合理性。定量评估表明,我们的模型在大多数指标上均获得一致更高的分数,尤其是在强降水事件和较长预报时效方面。定性分析显示,GenONet生成的预报结果在结构上连贯且能保持完整性,而基线模型则退化成模糊不清的形态。最后,消融研究证实了物理信息损失的收益,凸显了将算子学习与对抗训练相结合的优势。
cs.LG / 41 / 2609.00552

Manifold-Aware General Coded Computing for Straggler-Resilient Distributed Computing

面向慢节点容忍分布式计算的流形感知广义编码计算
Moradi, Parsa, Maddah-Ali, Mohammad Ali
Abstract
Existing coded-computing designs do not explicitly exploit the intrinsic structure of the input data. In communication systems, statistical structure and redundancy are often removed through source coding (or compression) before channel coding is applied. This principle, however, does not transfer directly to coded computation. In many computational tasks, particularly in machine learning, the structure of the data is precisely what the computation seeks to exploit to infer outputs or learn meaningful patterns. Consequently, coded-computing schemes should preserve and leverage this structure in their code design, rather than ignoring or eliminating it through source coding. This observation motivates a different perspective on code construction. In many channel-coding schemes, such as Reed-Solomon codes, coded symbols are generated by evaluating a low-dimensional algebraic representation at selected points. In contrast, many high-dimensional datasets naturally concentrate near low-dimensional manifolds. In this paper, we exploit this intrinsic geometry by designing coded samples that follow the natural manifold of the data, rather than imposing an artificial low-dimensional structure unrelated to the data distribution. Inspired by graph-based manifold learning, we propose a manifold-aware encoding strategy for general coded computing (GCC). Experiments on neural network inference and high-dimensional polynomial evaluation demonstrate that the proposed strategy consistently and significantly reduces the mean squared recovery error under straggling compared with standard GCC.
Chinese Translation
现有的编码计算设计并未显式地利用输入数据的内在结构。在通信系统中,通常在应用信道编码之前,通过信源编码(或压缩)去除统计结构和冗余。然而,这一原则并不能直接迁移到编码计算中。在许多计算任务中,尤其是在机器学习中,数据结构恰恰是计算所要利用的对象,以推断输出或学习有意义的模式。因此,编码计算方案应在码字设计中保留并利用这种结构,而不是通过信源编码将其忽略或消除。这一观察启发了一种不同的码构造视角。在许多信道编码方案(如里德-所罗门码(Reed-Solomon codes))中,编码符号是通过在选定的点上对低维代数表示求值而生成的。与之相对,许多高维数据集天然地集中于低维流形附近。在本文中,我们利用这种内在几何结构,设计了遵循数据自然流形的编码样本,而非强加与数据分布无关的人为低维结构。受基于图的流形学习启发,我们为广义编码计算(general coded computing,GCC)提出了一种流形感知的编码策略。在神经网络推理和高维多项式求值上的实验表明,与标准GCC相比,所提出的策略在慢节点存在的情况下能够持续且显著地降低均方恢复误差。
cs.LG / 42 / 2609.00566

EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection

EEG-VID:面向脑电解码与辅助目标选择的任务引导式潜在预测预训练方法
Sun, Guanzhong, Ma, Junyi, Wu, Yuxuan, Miao, Yanzi
Abstract
We propose EEG-VID, a task-guided latent predictive pretraining framework for EEG decoding under session and subject shifts. EEG-VID predicts future latent EEG states from recent history using an exponential-moving-average target encoder and weak task guidance, followed by supervised fine-tuning. Across VIG-48 and BCI Competition IV-2a/IV-2b, Stage 1 improves mean accuracy in 41 of 42 matched backbone-dataset-protocol comparisons, including all 12 leave-one-subject-out settings, with a maximum gain of 16.22 percentage points. On the 48-region cross-day VIG-48 task, EEG-VID achieves 6.52% Top-1 and 30.50% Top-5 accuracy. In a separate six-participant offline robot-scene study, candidate-constrained target selection reaches 40.24% versus a 25% chance level after subject-specific calibration. These results support task-guided latent prediction as a transferable pretraining strategy for EEG decoding and scene-constrained assistive target selection.
Chinese Translation
我们提出了EEG-VID,一个面向会话间与被试间漂移下脑电(EEG)解码的任务引导式潜在预测预训练框架。EEG-VID利用指数移动平均目标编码器和弱任务引导信号,从近期的脑电历史中预测未来的潜在脑电状态,随后进行有监督微调。在VIG-48和BCI Competition IV-2a/IV-2b数据集上,第一阶段预训练在42组匹配的骨干网络-数据集-协议组合中的41组均提升了平均准确率,包括全部12个留一被试交叉验证(leave-one-subject-out)设置,最大提升达16.22个百分点。在48个区域的跨天VIG-48任务上,EEG-VID取得了6.52%的Top-1准确率和30.50%的Top-5准确率。在一项独立的六名被试离线机器人场景研究中,经过被试特定校准后,候选约束下的目标选择准确率达到40.24%,显著高于25%的随机水平。这些结果支持将任务引导式潜在预测作为一种可迁移的预训练策略,用于脑电解码及场景约束下的辅助目标选择。
cs.LG / 43 / 2609.00577

GeoPAR: Large-Scale Multi-Agent Combinatorial Optimization with Geometry-Guided Parallel Autoregressive Learning

GeoPAR:基于几何引导并行自回归学习的大规模多智能体组合优化
Wu, Wenjian, Jia, Zesheng, Tang, Jiaying, Yang, Benyuan, Wang, Jin
Abstract
Multi-agent combinatorial optimization problems are notoriously challenging due to their NP-hard nature. Recent parallel autoregressive neural solvers improve inference efficiency by allowing agents to make decisions simultaneously, but their performance often degrades on large-scale instances. This is largely attributable to weak modeling of local geometric structures and the fact that conflicting task selections are handled only after action generation. To address these limitations, we propose GeoPAR, a geometry-guided parallel autoregressive reinforcement learning framework for scalable multi-agent combinatorial optimization. GeoPAR integrates three key components: (1) a projection-window sparse geometry mechanism that builds lightweight local candidate neighborhoods through multi-directional projections, (2) sparse edge-biased attention that injects these geometric relations into node representations, and (3) cache-guided conflict-aware assignment that reuses the geometric cache during decoding to suppress duplicate selections of exclusive tasks. Experiments on heterogeneous vehicle routing and open multi-depot pickup-and-delivery problems show that GeoPAR improves large-scale zero-shot generalization while substantially reducing rollout steps and maintaining efficient inference.
Chinese Translation
多智能体组合优化问题因其NP难的特性而极具挑战性。近期的并行自回归神经求解器通过允许智能体同时进行决策来提升推理效率,但其在大规模实例上的性能往往下降。这在很大程度上归因于对局部几何结构的建模较弱,以及冲突任务选择仅在动作生成之后才被处理。为解决这些局限,我们提出了GeoPAR,一种面向可扩展多智能体组合优化的几何引导并行自回归强化学习框架。GeoPAR集成了三个关键组件:(1)投影窗口稀疏几何机制,通过多方向投影构建轻量级局部候选邻域;(2)稀疏边偏置注意力,将这些几何关系注入节点表示;(3)缓存引导的冲突感知分配,在解码过程中复用几何缓存以抑制对互斥任务的重复选择。在异构车辆路径问题和开放式多仓库取送货问题上的实验表明,GeoPAR提升大规模零样本泛化能力的同时,显著减少了推演步数并保持了高效的推理。
cs.LG / 44 / 2609.00590

CRAFT: Fine-Tuning Pre-hoc Explainability in AI-native 6G RAN

CRAFT:AI原生6G无线接入网中先验可解释性的微调方法
Gajjar, Pranshav, Shah, Vijay K
Abstract
The next generation of mobile networks is envisioned as fully AI-native, with AI-RAN architectures embedding small language models (SLMs) to perform reasoning over real-time telemetry. The state-of-the-art training paradigms for telecom LLMs, exemplified by RANSTRUCT-style supervised fine-tuning (SFT) on curated instruction data, are limited to post hoc rationalization. Here, the explanations, when produced at all, are generated after or independently of the decision, leaving the decision process unauditable. Pre-hoc reasoning, where a causal reasoning trace is produced before the output label, is preferable, and the broader LLM reasoning literature has made real progress toward it via RL methods such as Group Relative Policy Optimization (GRPO). Here we observe that transplanting this recipe into the telecom setting runs into a cold-start barrier: SLMs either learn to output the desired format or learn to predict the label, but rarely both. We identify this barrier and propose CRAFT, which stands for Cold-start Reasoning Alignment via Fine-Tuning, a data-centric method to autonomously generate a verified dataset of (input, trace, label) triplets. CRAFT fine-tunes SLMs on this verified data using low-rank adaptation (LoRA), requiring substantially less compute and wall-clock time than GRPO-based methods. On the TRACTOR and IC xApp telecom datasets, CRAFT achieves up to 86.5% and 94.6% for accuracy and F1 with no parse failures, while direct GRPO and SFT+GRPO fail to exceed 28% and 53.5% F1 with multiple parse failures. We further show that CRAFT-initialized policies serve as a robust foundation for subsequent GRPO fine-tuning, as under diverse reward functions the performance remains consistent with no parse failures. Finally, we demonstrate that CRAFT consumes 59% less energy than GRPO-based baselines, making it a sustainable path to deployable, auditable AI in 6G RAN.
Chinese Translation
下一代移动网络被构想为完全AI原生,其中AI-RAN架构嵌入小语言模型(SLM),对实时遥测数据进行推理。当前电信领域大语言模型(LLM)最先进的训练范式,以RANSTRUCT风格的基于精选指令数据的监督微调(SFT)为代表,仅限于事后(post hoc)合理化。在这种模式下,即使产生了解释,也是在决策之后或独立于决策生成,导致决策过程无法审计。先验(pre-hoc)推理——即在输出标签之前生成因果推理链——是更可取的方式,而更广泛的LLM推理研究已通过诸如组相对策略优化(GRPO)等强化学习方法在这一方向取得了实质性进展。本文观察到,将该方法直接移植到电信场景中会遇到冷启动障碍:SLM要么学会输出期望的格式,要么学会预测标签,但很难同时兼顾二者。我们识别了这一障碍并提出CRAFT(Cold-start Reasoning Alignment via Fine-Tuning,通过微调实现冷启动推理对齐),这是一种以数据为中心的方法,可自主生成经过验证的(输入,推理链,标签)三元组数据集。CRAFT使用低秩适应(LoRA)在该验证数据上对SLM进行微调,其计算量和实际耗时远低于基于GRPO的方法。在TRACTOR和IC xApp电信数据集上,CRAFT的准确率和F1值分别高达86.5%和94.6%,且无解析失败;而直接GRPO以及SFT+GRPO的F1值分别未能超过28%和53.5%,且存在多次解析失败。我们进一步证明,经CRAFT初始化的策略可作为后续GRPO微调的稳健基础:在多种奖励函数下,性能保持一致且无解析失败。最后,我们证明CRAFT的能耗比基于GRPO的基线方法低59%,使其成为在6G RAN中部署可审计AI的一条可持续路径。
cs.LG / 45 / 2609.00597

Topological Steering

拓扑引导(Topological Steering)
Guérand, Benoît, Nguyen, Tan Minh
Abstract
With the rapid rise of large language models (LLMs), controlling undesirable model behaviors has become increasingly important. Existing behavioral control methods typically intervene directly in activation or feature space, but such approaches can be sensitive to outliers, distributional shifts, noise, and other local perturbations. Motivated by Topological Data Analysis (TDA), which captures global rather than purely local structure, we propose Topological Steering, a new framework for steering LLM behavior through the topological representation of activation spaces. Using persistence diagrams, our method connects activation-based steering with TDA and enables more robust behavioral control. We show that Topological Steering consistently modifies LLM behavior across multiple model families and model sizes.
Chinese Translation
随着大语言模型(LLM)的迅速崛起,控制模型的不良行为变得日益重要。现有的行为控制方法通常直接在激活空间或特征空间中进行干预,但此类方法对离群值、分布偏移、噪声及其他局部扰动较为敏感。受拓扑数据分析(TDA)的启发——该方法捕捉的是全局结构而非纯局部结构——我们提出了拓扑引导(Topological Steering),一种通过激活空间的拓扑表示来引导LLM行为的新框架。借助持续同调图(persistence diagrams),我们的方法将基于激活的引导与TDA联系起来,从而实现更鲁棒的行为控制。实验表明,拓扑引导能够在多个模型家族和不同模型规模上持续有效地改变LLM的行为。
cs.LG / 46 / 2609.00605

Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning

坦白你所知:大语言模型遗忘中的遗忘集与模型知识不匹配问题
Kim, Miso, Lee, Georu, Jeong, Seungwon, Lee, Woojin
Abstract
Machine unlearning for large language models (LLMs) often assumes that a pre-defined forget set matches what the model has memorized, but this frequently breaks in realistic privacy settings where the original training data is inaccessible. We term this gap forget-set misalignment and identify two cases. In Under Unlearning, the forget set omits memorized information and leakage persists. In Out-of-Knowledge Unlearning, the algorithm is driven to "forget" knowledge the model never learned, perturbing parameters and degrading utility. Using gradient-level analysis, we show these behaviors arise from misaligned unlearning targets rather than specific optimization choices. We then propose CONfession-to-Forget-Set (CONFS), a data-blind framework that constructs model-aligned forget sets by eliciting and formalizing the model's memorized knowledge. Across synthetic, multimodal, and real-world benchmarks, CONFS approaches Gold-standard performance on several metrics and achieves a competitive forgetting-utility balance, while preserving utility better than other data-blind forget-set constructions.
Chinese Translation
面向大语言模型(LLM)的机器遗忘通常假设预先定义的遗忘集与模型实际记住的内容相一致,但在原始训练数据不可获取的现实隐私场景中,这一假设经常不成立。我们将这种差距称为“遗忘集不匹配”,并识别出两种情况:在“欠遗忘”中,遗忘集遗漏了模型已记住的信息,导致信息泄露持续存在;在“超出知识范围的遗忘”中,算法被要求“遗忘”模型从未学过的知识,从而扰动模型参数并降低其性能。通过梯度层面的分析,我们表明这些行为源于不匹配的遗忘目标,而非特定的优化方法选择。随后,我们提出了CONfession-to-Forget-Set(CONFS),这是一个数据盲框架,通过引导并形式化模型所记住的知识来构建与模型对齐的遗忘集。在合成数据、多模态和真实世界基准测试中,CONFS在多项指标上接近金标准性能,并实现了具有竞争力的遗忘-性能平衡,同时比其他数据盲遗忘集构建方法更好地保留了模型性能。
cs.LG / 47 / 2609.00632

Breaking the Structural Identity: Personalized Federated LoRA Fine-tuning under Rank Heterogeneity

打破结构同一性:秩异构下的个性化联邦LoRA微调
Wang, Lei, Bian, Jieming, Zhang, Letian, Xu, Jie
Abstract
Large Language Models (LLMs) have achieved remarkable success across diverse domains, but their adaptation to privacy-sensitive, distributed datasets remains a challenge. While Federated Learning (FL) combined with Low-Rank Adaptation (LoRA) provides a resource-efficient paradigm for collaborative fine-tuning, practical deployments are hindered by the dual challenges of resource heterogeneity and data heterogeneity. Existing rank-heterogeneous methods primarily focus on bridging dimension mismatches for aggregation but typically provide a unified global model for all clients sharing the same rank, failing to capture client-specific features in non-IID scenarios. In this paper, we propose FedRoRA (Federated Rank-wise Personalized LoRA), a novel framework that enables fine-grained personalization within rank-heterogeneous federations. FedRoRA decouples adaptation into shared global directions and personalized rank-wise magnitudes governed by learnable diagonal scales. On the server side, it extracts a global subspace via singular value decomposition (SVD) and redistributes client-specific initializations through a personalized projection and top-$k$ selection mechanism. Extensive experiments on NLU and NLG benchmarks demonstrate that FedRoRA consistently outperforms state-of-the-art methods.
Chinese Translation
大语言模型(LLM)已在多个领域取得显著成功,但其在隐私敏感的分布式数据集上的适配仍然是一个挑战。虽然联邦学习(Federated Learning, FL)结合低秩自适应(Low-Rank Adaptation, LoRA)为协同微调提供了一种资源高效的范式,但实际部署仍受到资源异构性和数据异构性双重挑战的阻碍。现有的秩异构方法主要侧重于弥合聚合时的维度不匹配问题,但通常为共享相同秩的所有客户端提供统一的全局模型,无法在非独立同分布(non-IID)场景下捕获客户端特有的特征。本文提出了FedRoRA(Federated Rank-wise Personalized LoRA,联邦按秩个性化LoRA),这是一种能够在秩异构联邦中实现细粒度个性化的新框架。FedRoRA将自适应过程解耦为共享的全局方向和由可学习对角缩放因子控制的个性化按秩幅度。在服务器端,它通过奇异值分解(SVD)提取全局子空间,并通过个性化投影和top-$k$选择机制重新分配客户端特定的初始化。在自然语言理解(NLU)和自然语言生成(NLG)基准上的大量实验表明,FedRoRA始终优于当前最先进的方法。
cs.LG / 48 / 2609.00647

DK-GBMKKM: Dynamic Kernel-Space Granular-Ball Multiple Kernel $k$-Means Clustering

DK-GBMKKM:动态核空间粒球多核 $k$-均值聚类
Lian, Xiaoyu, Zhang, Yuchao, Xia, Shuyin, Zhong, Siqi, Xiang, Xuzhao
Abstract
Multiple kernel $k$-means integrates complementary nonlinear similarities by learning a combination of base kernels. Its pointwise optimization, however, is sensitive to noisy and boundary samples and repeatedly operates on sample-scale kernel matrices. Granular-ball representations organize local sample groups into mesoscopic units, but granular balls generated once in the input space may be inconsistent with the fused-kernel geometry that evolves during multiple kernel learning. We propose dynamic kernel-space granular-ball multiple kernel $k$-means (DK-GBMKKM). The method generates granular balls in the current fused kernel space and alternates kernel-weight learning with granular-ball membership updates, allowing the representation to adapt to changes in the fused-kernel geometry. A sample-size-weighted granular-ball kernel is further constructed to preserve the contributions of balls of different sizes, and its positive semidefiniteness and related equivalence properties are established. Experiments on 12 public datasets demonstrate the strong overall clustering performance of DK-GBMKKM. The code has been open-sourced for reproducibility: https://github.com/lianxiaoyu724/DK-GBMKKM.
Chinese Translation
多核 $k$-均值聚类通过学习基核函数的组合来融合互补的非线性相似度。然而,其逐点优化方式对噪声样本和边界样本较为敏感,并且需要在样本规模的核矩阵上反复运算。粒球表示将局部样本组组织成介观单元,但在输入空间中一次性生成的粒球可能与多核学习过程中不断演化的融合核几何结构不一致。为此,我们提出了动态核空间粒球多核 $k$-均值聚类方法(DK-GBMKKM)。该方法在当前融合核空间中生成粒球,并将核权重学习与粒球隶属度更新交替进行,使表示能够自适应融合核几何结构的变化。此外,进一步构建了一种样本规模加权的粒球核,以保留不同大小粒球的贡献,并证明了其半正定性及相关等价性质。在12个公开数据集上的实验表明,DK-GBMKKM 具有优异的整体聚类性能。代码已开源以保证可复现性:https://github.com/lianxiaoyu724/DK-GBMKKM。
cs.LG / 49 / 2609.00653

EEG-AS: Instance-Level Foundation Model Selection for EEG Foundation Models via Behavior Reconstruction

EEG-AS:基于行为重构的EEG基础模型实例级模型选择方法
Zhang, Yunzhen, Piao, Ruoxi, Keles, Hasan Onur, Misir, Mustafa
Abstract
Electroencephalography (EEG) is a non-invasive technique for measuring neural activity and has been widely used in neuroscience applications. Recent advances in EEG foundation models have enabled strong performance across diverse neural decoding tasks. However, no single foundation model consistently performs best across datasets or individual EEG instances, while instance-level model selection remains largely unexplored. To address this limitation, we formulate EEG foundation model selection as an instance-level Algorithm Selection (AS) problem. We propose \textbf{EEG-AS}, an instance-level algorithm selection framework that characterizes each EEG instance using inference-available latent EEG embeddings, handcrafted neurophysiological features, and an anchor foundation model. During training, EEG-AS learns to reconstruct unavailable foundation-model behaviors from privileged prediction tokens conditioned on an anchor foundation model, while during inference it estimates these behaviors without executing the entire model portfolio, enabling efficient selection from seven EEG foundation models. Experiments on seven public EEG benchmarks demonstrate that EEG-AS substantially narrows the gap between the Single Best Solver (SBS) and the oracle upper bound for each instance. These results highlight the effectiveness of instance-level AS for adaptive deployment of EEG foundation models.
Chinese Translation
脑电图(EEG)是一种测量神经活动的非侵入性技术,已被广泛应用于神经科学领域。EEG基础模型的最新进展使其在多种神经解码任务中取得了优异表现。然而,没有任何单一基础模型能够在所有数据集或每个EEG实例上都持续表现最佳,而实例级的模型选择在很大程度上仍未被探索。为解决这一局限,我们将EEG基础模型选择形式化为一个实例级算法选择(Algorithm Selection, AS)问题。我们提出了EEG-AS,一个实例级算法选择框架,它利用推理时可获得的潜在EEG嵌入、手工设计的神经生理学特征以及一个锚点基础模型来刻画每个EEG实例。在训练阶段,EEG-AS学习从以锚点基础模型为条件的特权预测令牌中重构不可获得的基础模型行为;在推理阶段,它无需执行整个模型组合即可估计这些行为,从而能够从七个EEG基础模型中高效地进行选择。在七个公开EEG基准上的实验表明,EEG-AS显著缩小了单一最佳求解器(Single Best Solver, SBS)与每个实例的oracle上界之间的差距。这些结果凸显了实例级算法选择在EEG基础模型自适应部署中的有效性。
cs.LG / 50 / 2609.00679

HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields

HarmoCore:面向振荡波场稀疏重建的函数式潜在扩散方法
Chen, Lihao, Zhang, Xinyu, Chen, Panqi, Cheng, Lei, Zhang, Ting, Li, Jianlong, Fang, Shikai
Abstract
Reconstructing oscillatory wave fields from scattered sensors is a severely underdetermined inverse problem. Beyond the challenges of general physical-field reconstruction, wave responses are complex-valued, frequency-sensitive, and highly oscillatory, while costly simulation and sensing often leave only extreme-sparse observations. Existing low-rank, operator, and diffusion approaches are largely designed for real-valued, smoother fields; dense pixel-space diffusion is particularly inefficient for oscillatory complex fields and difficult to scale to 3D. We propose HarmoCore, which places a generative prior in a compact, continuous, and structured wave-field latent. HarmoCore represents joint real--imaginary channels with Functional Tucker cores over shared continuous spatial bases, learns a frequency-conditioned core diffusion prior, and performs Diffusion Posterior Sampling directly in core space. At fixed sensor coordinates, the multilinear decoder induces an explicit likelihood guidance operator, avoiding dense pixel-space correction. Optional target-equation residual guidance further promotes physical consistency. Experiments on 2D Helmholtz, 2D synthetic wave fields, and 3D Helmholtz show substantial gains under 1%--2% sensing while remaining practical in three dimensions.
Chinese Translation
从稀疏分布的传感器重建振荡波场是一个严重欠定的逆问题。除了一般物理场重建的挑战外,波响应还是复值的、频率敏感且高度振荡的,而昂贵的仿真与传感往往只能提供极端稀疏的观测。现有的低秩、算子和扩散方法大多针对实值、较平滑的场设计;稠密像素空间的扩散对于振荡复值场尤其低效,且难以扩展到三维。我们提出HarmoCore,将生成先验置于一个紧凑、连续且结构化的波场潜在空间中。HarmoCore利用基于共享连续空间基的函数式Tucker核心(Functional Tucker cores)表示实部—虚部联合通道,学习频率条件化的核心扩散先验,并直接在核心空间中执行扩散后验采样(Diffusion Posterior Sampling)。在固定传感器坐标处,多线性解码器可诱导出显式的似然引导算子,从而避免稠密像素空间的校正。可选的目标方程残差引导可进一步促进物理一致性。在二维Helmholtz、二维合成波场和三维Helmholtz上的实验表明,在1%–2%的传感数据下该方法取得显著提升,同时在三维场景中仍保持实用性。
cs.LG / 51 / 2609.00686

A Study of Hidden-State Optimization Order in Predictive Coding Networks

预测编码网络中隐状态优化顺序的研究
Li, Xueyuan, Vargas, Danilo Vasconcellos
Abstract
Local learning methods offer an alternative to end-to-end backpropagation, but their unstructured local objectives can produce weak feature learning in deep networks. We study whether the order of hidden-state optimization can address this limitation. We propose a boundary-first inference schedule that partitions a model into chunks, first coordinates hidden states at chunk boundaries, and then refines representations within each chunk. We instantiate this schedule in predictive coding networks (PCNs), a local-learning framework in which hidden activities and prediction errors are explicitly exposed during inference. On CIFAR-10, the resulting boundary-first predictive-coding instantiation improves accuracy over standard predictive coding by $9.77\%$ under a standard parametrization and by $5.51\%$ under a $\mu$-parametrization. Diagnostic analyses further show more non-trivial early-layer updates, lower initial-to-final CKA, and more diverse layerwise gradients, consistent with stronger feature learning. These results support boundary-first, chunk-based inference as a practical design principle for predictive-coding training and motivate its study in broader local-learning systems.
Chinese Translation
局部学习方法为端到端反向传播提供了一种替代方案,但其非结构化的局部目标可能导致深度网络中的特征学习能力较弱。我们研究了隐状态优化的顺序能否解决这一局限。我们提出了一种边界优先(boundary-first)的推理调度策略,该策略将模型划分为多个块(chunk),首先在块边界处协调隐状态,然后在每个块内部细化表征。我们在预测编码网络(Predictive Coding Networks, PCNs)中实现了该调度策略,这是一类在推理过程中显式暴露隐层活动和预测误差的局部学习框架。在 CIFAR-10 数据集上,所得的边界优先预测编码实现在标准参数化下相比标准预测编码将准确率提升了 9.77%,在 μ 参数化下提升了 5.51%。诊断分析进一步表明,该方法带来了更多非平凡的浅层更新、更低的初始至最终 CKA(中心化核相似度),以及更多样的逐层梯度,这些都与更强的特征学习相一致。这些结果支持将边界优先、基于块的推理作为预测编码训练的一种实用设计原则,并激发了对更广泛的局部学习系统中该策略的进一步研究。
cs.LG / 52 / 2609.00691

Verdict Instability of OOD Scores under Reference Resampling

参考集重采样下OOD评分判定的不稳定性
Lee, Donghoon, Kang, Shinjin
Abstract
Post-hoc out-of-distribution detectors are fitted on a finite reference set, so every score they produce is an estimate. If we had chosen a different set, some verdicts would have moved. We measure that movement by resampling the reference set and recording the bootstrap standard deviation of the score, which we call verdict instability. It admits a closed form with no fitted parameters. The instability of a verdict is the within-class dispersion of the assigned class along the query's direction, divided by the square root of that class's reference count. That count is what separates verdict instability from the geometry of the score distribution, and it is identifiable only under class imbalance. Instability grows with the local dispersion. Far-OOD queries lie along the low-variance directions of an anisotropic embedding, so every distance-based score we test assigns its highest values to the verdicts that are most reproducible. Only estimators of local dispersion carry the sign a practitioner expects. We give a rule that predicts this sign for any score from a single label-free correlation, and abstention driven by a wrong-signed score turns out worse than abstention at random on every dataset we test.
Chinese Translation
事后(post-hoc)分布外(OOD)检测器是在有限的参考集上拟合的,因此其产生的每一个评分都只是一个估计值。如果我们选择了不同的参考集,某些判定结果可能会发生改变。我们通过重采样参考集并记录评分的自助法(bootstrap)标准差来度量这种变动,并将其称为判定不稳定性(verdict instability)。该不稳定性具有无需拟合参数的闭式解。某个判定的不稳定性等于所判定类别沿查询方向上的类内离散度除以该类别参考样本数的平方根。正是这一样本数将判定不稳定性与评分分布的几何结构区分开来,且只有在类别不平衡的情况下才能识别。不稳定性随局部离散度的增大而增长。远分布外(Far-OOD)查询位于各向异性嵌入的低方差方向上,因此我们测试的每一个基于距离的评分都会将最高值赋予那些最具可复现性的判定。只有局部离散度的估计量才携带实践者所期望的符号。我们给出了一条规则,仅凭一次无标签的相关性计算即可预测任意评分的符号;在我们测试的所有数据集上,由错误符号的评分驱动的弃判(abstention)效果反而比随机弃判更差。
cs.LG / 53 / 2609.00696

MUGEN: Generating Unlearnable Graph Examples for Multiple Learning Tasks

MUGEN:面向多种学习任务的不可学习图样本生成
Liu, Ziyan, Zhao, Chengshuai, Liu, Huan
Abstract
Graph data across diverse domains can expose valuable relational information to unauthorized representation learning, creating a pressing need for protection against such misuse. Unlearnable examples offer a data-level defense by perturbing a training release so that models trained on it fail to generalize to clean data. Existing methods generate unlearnable graph examples for only a specified downstream task. Consequently, a release protected against one task may remain learnable for other plausible uses, including node classification, graph classification, and link prediction, which the data owner cannot anticipate. We introduce MUGEN, to our knowledge the first framework for generating unlearnable graph examples that jointly protect all enabled tasks. From one clean dataset, MUGEN produces a single feature-perturbed release that protects every enabled task through a shared GNN encoder and task-specific heads. We devise a Task-Aligned Separability Objective (TASO), which leverages task prediction and classwise separability to strengthen unlearnability and its transfer across GNN backbones and enabled tasks. We further introduce Type-Adaptive Perturbation (TAP), which tailors perturbation optimization to node-attribute type, with direct search over feasible hard flips that accept only loss-improving updates for discrete node attributes and customized gradient-based updates for continuous node features, thereby enabling strong unlearnability across both settings. Experiments across five benchmarks, four backends and three learning paradigms demonstrate that MUGEN generates transferable unlearnable graph examples across GNN backbones and all three tasks, and remains effective under adversarial training and data augmentation.
Chinese Translation
来自不同领域的图数据可能向未经授权的表示学习暴露有价值的关联信息,因此迫切需要针对此类滥用的保护手段。不可学习样本通过扰动发布的数据,使在其上训练的模型无法泛化到干净数据,从而提供数据层面的防御。现有方法仅针对特定下游任务生成不可学习图样本。因此,针对某一任务受保护的数据发布,对于数据所有者无法预见的其他潜在用途(包括节点分类、图分类和链接预测)仍可能是可学习的。我们提出MUGEN,据我们所知,这是首个能够同时保护所有启用任务的不可学习图样本生成框架。基于一份干净数据集,MUGEN生成单一的特征扰动发布版本,通过共享的GNN编码器和任务特定的输出头保护所有启用任务。我们设计了任务对齐可分性目标(Task-Aligned Separability Objective, TASO),利用任务预测和类间可分性来增强不可学习性及其在GNN骨干网络和启用任务之间的可迁移性。我们进一步提出类型自适应扰动(Type-Adaptive Perturbation, TAP),根据节点属性类型定制扰动优化:对离散节点属性在可行硬翻转上直接搜索、仅接受能降低损失的更新,对连续节点特征采用定制的基于梯度的更新,从而在两种设置下都实现强大的不可学习性。在五个基准数据集、四种骨干模型和三种学习范式上的实验表明,MUGEN生成的不可学习图样本能够在不同GNN骨干网络和全部三种任务之间迁移,并且在对抗训练和数据增强下依然有效。
cs.LG / 54 / 2609.00699

Patterning in Practice: Debiasing Reward Models with Susceptibilities

模式化实践:利用敏感度对奖励模型去偏
Wang, George, Donoway, Elizabeth, Murfet, Daniel
Abstract
Reward models trained on human preferences are known to suffer from length, formatting, and other stylistic biases. In this paper we use patterning, which reweights each preference pair according to its measured effect on posterior expectation values of benchmark losses (its susceptibility), to debias a Gemma 2 9B Instruct reward model trained on Skywork-Reward-Preference v0.2. We obtain $+14.2 \pm 1.2$ pp on RM-Bench Hard, the split where style cues point against correctness (mean $\pm$ s.e.\ over 5 seeds), with overall RM-Bench accuracy preserved, comparable to the strongest Hard-split gain reported by the closest published comparator (SteerRM, $+13.2$ pp). We demonstrate in a simple case that the reweighting is interpretable by tracing a side effect of the intervention (a regression on a safety subset of RM-Bench) to a small class of training pairs, which we confirm by ablation. The weights also transfer: those computed on Gemma 2 9B debias Gemma 2 2B and 27B with no recomputation, and transfer partially to Llama 3.1 8B. This is the first application of patterning, a program grounded in singular learning theory, beyond small models and synthetic tasks.
Chinese Translation
基于人类偏好训练的奖励模型已知存在长度、格式及其他风格方面的偏差。本文使用模式化(patterning)方法,即根据每个偏好对在基准损失后验期望值上的测得效应(其敏感度 susceptibility)进行重加权,对基于 Skywork-Reward-Preference v0.2 训练的 Gemma 2 9B Instruct 奖励模型进行去偏。我们在 RM-Bench Hard 子集(即风格线索与正确性相悖的子集)上取得了 +14.2 ± 1.2 个百分点(pp)的提升(5 个随机种子的均值 ± 标准误),同时整体 RM-Bench 准确率得以保持,与已发表的最接近对比方法所报告的最强 Hard 子集增益相当(SteerRM,+13.2 pp)。我们通过一个简单案例证明,该重加权是可解释的:我们将干预的一个副作用(在 RM-Bench 安全子集上的性能回退)追溯至一小类训练对,并通过消融实验予以证实。这些权重还具有可迁移性:在 Gemma 2 9B 上计算得到的权重无需重新计算即可对 Gemma 2 2B 和 27B 进行去偏,并可部分迁移至 Llama 3.1 8B。这是模式化方法——一个植根于奇异学习理论的研究方案——首次应用于小型模型和合成任务之外的场景。
cs.LG / 55 / 2609.00734

Online Self-Weighted Fine-Tuning

在线自加权微调
Wen, Haiquan, He, Yiwei, Peng, Bei, Cheng, Guangliang
Abstract
Standard supervised fine-tuning (SFT) assigns the same explicit loss weight to every expert demonstration, regardless of the model's changing competence over training queries. Reinforcement learning (RL) based methods adapt update strength using model-generated rollouts, but often require substantially more sampling and can be unstable on hard tasks. We propose \textbf{Online Self-Weighted Fine-Tuning (OSW-FT)}, a simple method that augments SFT with online, trajectory-level weighting. For each query, OSW-FT estimates the model's current success rate using a small number of inference-only rollouts and rescales the standard SFT loss accordingly. The optimization direction remains anchored to the expert trajectory, while the update magnitude adapts online. For binary-verifiable reasoning, we connect this weighting to SFT and RL at the gradient level, inspired by variance-reduction principles. The resulting estimator is unbiased for the exact OSW-FT surrogate update for any finite rollout count, and we analyze convergence with respect to the corresponding surrogate objective. Evaluated across Qwen3 series ranging from 0.6B to 4B on multiple challenging benchmarks (e.g., AIME), OSW-FT consistently improves over SFT on small-to-medium scale models. OSW-FT offers a favorable compute-performance trade-off as a practical approach for fine-tuning small-to-medium LLMs on binary-verifiable reasoning tasks with only \textbf{2 online rollouts}.
Chinese Translation
标准的监督微调(SFT)为每个专家演示分配相同的显式损失权重,而不考虑模型在训练查询上不断变化的能力。基于强化学习(RL)的方法利用模型生成的 rollout 来自适应调整更新强度,但通常需要大量的采样,并且在困难任务上可能不稳定。我们提出了在线自加权微调(Online Self-Weighted Fine-Tuning, OSW-FT),这是一种简单的方法,通过在线的、轨迹级别的加权来增强 SFT。对于每个查询,OSW-FT 使用少量仅用于推理的 rollout 来估计模型当前的成功率,并相应地对标准 SFT 损失进行重新加权。优化方向始终锚定于专家轨迹,而更新幅度则在线自适应调整。对于二值可验证的推理任务,受方差缩减原理的启发,我们在梯度层面将这种加权与 SFT 和 RL 联系起来。所得估计量对于任意有限的 rollout 数量都是精确 OSW-FT 代理更新的无偏估计,并且我们针对相应的代理目标分析了其收敛性。在 Qwen3 系列(参数量从 0.6B 到 4B)上、于多个具有挑战性的基准(如 AIME)上进行评估,OSW-FT 在中小规模模型上一致优于 SFT。仅需 2 个在线 rollout,OSW-FT 便可在二值可验证推理任务上微调中小规模大语言模型时提供良好的计算-性能权衡,是一种实用的方法。
cs.LG / 56 / 2609.00746

Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis

视觉-语言适配中的文本能力损失:一种基于注意力汇聚点(Attention-Sink)的诊断方法
Choi, Minsik, Kim, Geewook, Kim, Young Geun
Abstract
Fine-tuning a pretrained LLM into a vision-language model (VLM) can erode the backbone's text capability, with the damage concentrated on tasks that require following exact output rules, such as instruction following, chain-of-thought reasoning graded on a strictly parsed final answer, and similar evaluations with strict graders. We trace this gap to attention-sink corruption: VL fine-tuning perturbs the early sink position that anchors a large fraction of attention probability, and how well the base LLM preserves its sink tracks how much of the affected capability survives adaptation. Building on this view, we introduce Sink Strength, a single scalar computed on the base LLM in a few seconds on a single GPU that predicts post-VL degradation without any VL training. It consistently tracks relative degradation across the six VLM-LLM pairs and multiple format-sensitive tasks. Complementing this diagnostic, we find that post-pretraining QK-RMSNorm injection fails to reproduce the protection of native QK-RMSNorm, while several off-the-shelf weight-merging settings fail to recover the lost capability after VL training. These negative results underscore the value of screening backbones with Sink Strength before VL training and narrow the intervention space toward head-selective training-time protection.
Chinese Translation
将预训练大语言模型(LLM)微调为视觉-语言模型(VLM)可能会削弱其骨干网络的文本能力,且这种损伤集中在需要遵循严格输出规则的任务上,例如指令遵循、基于严格解析最终答案评分的思维链推理,以及其他采用严格评分器的类似评估。我们将这一差距归因于注意力汇聚点(attention sink)的损坏:视觉-语言微调会扰动早期汇聚点位置,而该位置锚定了相当大比例的注意力概率;基础LLM对其汇聚点的保持程度决定了受影响能力在适配后的存续程度。基于这一观点,我们提出了汇聚点强度(Sink Strength),这是一个仅需在单张GPU上几秒钟内即可基于基础LLM计算得到的标量指标,能够在无需任何视觉-语言训练的情况下预测VL适配后的性能退化。该指标在六组VLM-LLM配对以及多个对格式敏感的任务上始终能够有效追踪相对退化程度。作为对该诊断方法的补充,我们发现预训练后注入QK-RMSNorm无法复现原生QK-RMSNorm的保护效果,同时多种现成的权重合并设置也未能在VL训练后恢复所损失的能力。这些负面结果凸显了在VL训练前使用汇聚点强度筛选骨干模型的价值,并将干预空间收窄至面向特定注意力头的训练时保护方法。
cs.LG / 57 / 2609.00753

How Do Language Models Choose Between Context and Memory?

语言模型如何在上下文与记忆之间进行选择?
Shih, Benjamin, Winnicki, John, Cao, Arianna
Abstract
When contextual information conflicts with the knowledge stored in model parameters, activation directions can be used to decode and steer which source the model follows. However, steering along a direction does not establish causality: whether the unedited model would naturally use that direction or whether the direction is reusable across tasks. We test these distinctions through counterfactual experiments in unambiguous settings. First, we estimate authority directions from agreement prompts, in which the context and parametric knowledge support the same answer. We then interchange naturally occurring coordinates along these directions between matched prompts that direct the model to prioritize either the supplied context or its parametric knowledge. Across Qwen, Llama, and OLMo models, this intervention reproduces 30-68% of the authority-induced shift in source choice, whereas matched controls reproduce almost none. To test cross-task reuse, we learn authority directions on two tasks separately and see that cross-task transferability closes only 9% of the authority gap while the local direction learned on the given task closes 57%. These results distinguish authority representation, causal use, and cross-task causal reuse, and suggest that authority computations may be task-dependent, rather than reusable across tasks.
Chinese Translation
当上下文信息与模型参数中存储的知识发生冲突时,可以利用激活方向来解码并引导模型遵循哪种信息来源。然而,沿某一方向进行引导并不能确立因果性:未经编辑的模型是否会自然地使用该方向,以及该方向是否可以在不同任务间复用,这些问题仍不清楚。我们通过在无歧义设置下进行反事实实验来检验这些区别。首先,我们从一致性提示(即上下文与参数知识支持相同答案的提示)中估计权威方向(authority directions)。然后,我们在匹配的提示之间交换这些方向上自然出现的坐标,这些提示分别引导模型优先考虑所提供的上下文或其参数知识。在Qwen、Llama和OLMo模型上,这种干预能够复现权威引起的源选择偏移的30-68%,而匹配的对照组几乎无法复现任何偏移。为检验跨任务复用,我们分别在两个任务上学习权威方向,结果显示跨任务可迁移性仅能弥合9%的权威差距,而在给定任务上局部学习的方向则能弥合57%。这些结果区分了权威表征、因果使用与跨任务因果复用,并表明权威计算可能是任务相关的,而非可跨任务复用。
cs.LG / 58 / 2609.00762

Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation

冻结核心需要任务信号:面向低资源大语言模型适配的Fisher白化交叉协方差方法
Ye, Wentao, Shen, Zhanming, Xiao, Zhiqing, Ding, Yao, Wang, Haobo, Chen, Gang
Abstract
Parameter-efficient fine-tuning is usually framed as a question of how many parameters to update. Under a severe trainable-state budget, however, where those coefficients act is equally consequential. We study this choice through frozen-core adaptation: a calibration pass fixes left and right bases for each weight matrix, and fine-tuning optimizes only an $r\times r$ core. This removes the ability of trainable factors to repair a poor initial span and makes subspace quality directly observable. We introduce FCCA, which estimates the signed input--error cross-covariance, whitens it with diagonal Fisher moments, truncates it in the resulting local metric, maps the selected directions back, and applies thin QR to obtain stable core coordinates. Under a matched $r^2$ budget, we compare eight basis constructors on 11 tasks, four model settings, and three seeds. On Qwen2.5-3B, FCCA reaches an 83.0 macro-average, 2.3 points above the next-best matched-budget constructor, and exceeds its unwhitened RawGrad control on all 11 tasks. It ranks first at all three Qwen scales and finishes within 0.13 points of the best method on Llama-3.2-1B. Controlled ablations show gains of 2.7--17.2 points from whitening and identify QR as necessary for stable core optimization in the tested regime. Finally, FCCA comes within 0.32 and 0.23 average points of LoRA and DoRA while optimizing 36.9K rather than roughly 7.4M parameters. These results show that a carefully selected fixed span can recover most of the benefit of movable low-rank factors at a much smaller trainable and optimizer-state cost.
Chinese Translation
参数高效微调通常被表述为更新多少参数的问题。然而,在极度受限的可训练状态预算下,这些系数作用于何处同样重要。我们通过冻结核心适配(frozen-core adaptation)来研究这一选择:一次校准过程为每个权重矩阵确定左右基,微调仅优化一个 $r imes r$ 的核心。这消除了可训练因子修复不良初始子空间的能力,并使子空间质量可以直接观察。我们提出FCCA(Fisher-Whitened Cross-Covariance Adaptation),其估计带符号的输入-误差交叉协方差,用对角Fisher矩对其进行白化,在所得的局部度量中进行截断,将选定的方向映射回原空间,并应用薄QR分解以获得稳定的核心坐标。在相同的 $r^2$ 预算下,我们在11个任务、四种模型设置和三个随机种子上比较了八种基构造方法。在Qwen2.5-3B上,FCCA达到83.0的宏平均,比次优的同等预算基构造方法高出2.3分,并在全部11个任务上超过其未白化的RawGrad对照组。它在三个Qwen规模上均排名第一,并在Llama-3.2-1B上与最佳方法相差不超过0.13分。受控消融实验显示白化带来2.7至17.2分的提升,并确认QR分解在所测试的机制下对稳定的核心优化是必要的。最后,FCCA与LoRA和DoRA的平均性能差距仅分别为0.32分和0.23分,而优化的参数量仅为36.9K,而非约740万。这些结果表明,一个精心选择的固定子空间能够以远小得多的可训练参数和优化器状态成本,恢复大部分可移动低秩因子的收益。
cs.LG / 59 / 2609.00764

Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures

你所想是否即我所想?:考察神经网络架构中的概念分离
Ponde, Jaee, Agarwal, Roshni, Banerjee, Subhashis
Abstract
Neural networks are increasingly employed to identify both well-defined and ambiguous concepts, yet output-level metrics reveal little about how those concepts are represented internally. Our study asks if these networks exhibit \textit{conceptual separation}: if examples of the same concept form coherent representations, and whether related concepts lie closer together in the representation space. We examine this conceptual organisation in Convolutional Neural Networks (CNNs) and Large Language Models (LLMs) through geometric and distributional analysis of their internal activations. In CNNs, familiar ImageNet concepts form coherent and semantically ordered representations, while this coherence weakens for unseen concepts and suffers within-class domain shift. In LLMs, clearly distinct domains remain well separated, related subdomains move closer together, and the distinction between ambiguous topics collapses at both the mean and covariance level. These results suggest that conceptual separation can reveal structure that output accuracy alone cannot, and may serve as a useful diagnostic of how robustly a model represents the concepts it is asked to identify. Code and data available on \href{https://github.com/JaeeRoshniCapstoneProject/Are-You-Thinking-What-I-m-Thinking-Examining-Conceptual-Separation-in-Neural-Architectures}{GitHub}.
Chinese Translation
神经网络被越来越多地用于识别定义明确的概念和模糊概念,然而输出层面的指标几乎无法揭示这些概念在网络内部是如何被表示的。本研究探讨这些网络是否表现出*概念分离*(conceptual separation):即同一概念的样本是否形成连贯的表示,以及相关概念在表示空间中是否更为接近。我们通过对卷积神经网络(CNN)和大语言模型(LLM)内部激活的几何与分布分析来考察这种概念组织方式。在CNN中,熟悉的ImageNet概念形成连贯且语义有序的表示,而这种连贯性对于未见过的概念会减弱,并在类内领域偏移时受损。在LLM中,截然不同的领域仍然保持良好的分离,相关的子领域彼此靠近,而模糊主题之间的区分在均值和协方差层面均出现崩塌。这些结果表明,概念分离能够揭示仅凭输出准确率无法发现的结构,并可作为模型在表示其被要求识别的概念时稳健程度的有效诊断手段。代码和数据可在GitHub上获取。
cs.LG / 60 / 2609.00789

Subspace Levenberg Marquardt Algorithms in Training Neural Networks

用于训练神经网络的子空间Levenberg-Marquardt算法
Hoang, M. Duc
Abstract
The Levenberg-Marquardt (LM) algorithm is a well-known second-order method for rapid convergence and strong robustness when training small- to medium-sized neural networks (NNs). However, its computational and memory costs increase significantly as the number of parameters in an NN grows. To address this limitation, subspace methods have been proposed, such as the Krylov subspace LM (KSLM) and the hybrid subspace LM (HSLM), making second-order algorithms more efficient. In this work, we evaluate the subspace Levenberg-Marquardt algorithms for regression and classification tasks in neural networks. We compare the performance of subspace LM variants with the classical LM method, as well as other popular first-order algorithms, such as stochastic gradient descent (SGD) and Adam.
Chinese Translation
Levenberg-Marquardt(LM)算法是一种著名的二阶方法,在训练中小规模神经网络(NNs)时具有快速收敛和强鲁棒性的特点。然而,随着神经网络中参数数量的增长,其计算成本和内存开销显著增加。为了解决这一局限性,研究者提出了子空间方法,例如Krylov子空间LM算法(KSLM)和混合子空间LM算法(HSLM),使二阶算法更加高效。在本工作中,我们针对神经网络中的回归和分类任务评估了子空间Levenberg-Marquardt算法。我们将子空间LM各种变体的性能与经典LM方法以及其他流行的一阶算法(如随机梯度下降(SGD)和Adam)进行了比较。
cs.LG / 61 / 2609.00829

HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution

HarnessEvolve:从参考轨迹中学习以实现可靠的智能体自我进化
Jiang, Wen, Chu, Mingmin, Tian, Yimeng, Zhang, Qianxin, Yang, Haofei, Yang, Rui, Liu, Yang, Lv, Tao, Li, Fangming
Abstract
Self-evolving agents advance toward autonomy by optimizing their harness---prompts, skills, tools, and execution logic---based on environmental feedback. This paradigm, however, is hampered by three challenges: \textit{credit assignment failure}, where terminal success/failure feedback makes it ambiguous which step caused the error; \textit{shortcut learning}, where agents memorize task-specific patterns rather than acquire generalizable capabilities; and \textit{catastrophic forgetting}, where unguarded updates degrade previously acquired competence. In this paper, we introduce HarnessEvolve, a self-evolving framework that learns from reference trajectories to achieve reliable agent self-evolution. HarnessEvolve decouples the execution agent from the evolutionary pipeline, assigning execution, evaluation, optimization, and gating to independent agent modules, enabling generalizable and stable harness improvements. Specifically, HarnessEvolve overcomes credit assignment failure by generating reference trajectories (execution paths produced when given the ground-truth answers) and aligning failed executions against them to extract error signals, which are clustered to reveal systematic failure patterns. To prevent shortcut learning and catastrophic forgetting, candidate harness updates must pass two gates: a quality gate that filters data leakage and prompt bloat, and a performance gate that accepts each update if it improves on the current batch without degrading recent batches, with epoch-end validation on a held-out set selecting the best-performing accepted agent snapshot. We conduct extensive experiments on several benchmarks spanning open-domain and enterprise scenarios, using different models and agent frameworks. Results demonstrate that HarnessEvolve consistently outperforms state-of-the-art baselines across all benchmarks and settings, confirming reliability across task domains.
Chinese Translation
自我进化智能体通过基于环境反馈优化其执行框架(harness)——包括提示词、技能、工具和执行逻辑——来逐步迈向自主性。然而,这一范式面临三大挑战: extit{信用分配失败},即仅依据最终成败的反馈难以判断是哪一步导致了错误; extit{捷径学习},即智能体记住的是任务特定的模式,而非习得可泛化的能力;以及 extit{灾难性遗忘},即缺乏保护的更新会损害先前已获得的能力。本文提出 HarnessEvolve,一个通过从参考轨迹中学习来实现可靠智能体自我进化的框架。HarnessEvolve 将执行智能体与进化流程解耦,将执行、评估、优化和门控分配给相互独立的智能体模块,从而实现可泛化且稳定的执行框架改进。具体而言,HarnessEvolve 通过生成参考轨迹(在给定真实答案时的执行路径),并将失败的执行与之对齐以提取错误信号,从而克服信用分配失败问题;这些错误信号经聚类分析以揭示系统性的失败模式。为防止捷径学习和灾难性遗忘,候选的执行框架更新必须通过两道门控:一是质量门控,用于过滤数据泄露和提示词膨胀;二是性能门控,仅当更新在当前批次上有所改进且不损害近期批次性能时才予以接受,并在每个 epoch 结束时通过留出集验证来选择表现最佳的已接受智能体快照。我们在涵盖开放域和企业场景的多个基准上,使用不同的模型和智能体框架开展了大量实验。结果表明,HarnessEvolve 在所有基准和设置下均持续优于最先进的基线方法,验证了其在不同任务领域上的可靠性。
cs.LG / 62 / 2609.00863

Conditional Flow Matching for ML-Based Inverse Design Problems

基于机器学习的逆向设计问题的条件流匹配方法
Felder, Juliana, Habibi, Milad, Massoudi, Soheyl, Fuge, Mark
Abstract
Engineering inverse design is often limited by the high computational cost of iterative solvers for optimization problems constrained by partial differential equations (PDEs) and by their sensitivity to initialization. Deep generative models can produce candidate designs without rerunning the simulator at inference time. Generative adversarial networks (GANs) sample in one forward pass, whereas diffusion models require iterative reverse-time integration. In this work, we add conditional flow matching (CFM) to EngiOpt and compare it with a conditional diffusion model and a conditional generative adversarial network (cGAN) on structural (beams2d) and thermal (heatconduction2d) benchmarks from EngiBench using the same downstream optimization protocol. We use cumulative optimality gap (COG) and final optimality gap (FOG) as the primary metrics for evaluating the generated designs as warm starts for gradient-based refinement. On the evaluated EngiOpt implementations and two EngiBench tasks, CFM achieves the lowest measured COG, FOG, maximum mean discrepancy (MMD), and volume-fraction deviation on both tasks. CFM has mean volume-fraction deviations of 0.4% and 1.0% on beams2d and heatconduction2d, respectively, compared with 3.8% and 11.2% for diffusion. At Euler s = 16, CFM achieves 53.2 samples/s on beams2d, about 66 times the measured throughput of the evaluated diffusion baseline using 1000 network evaluations under the same timing protocol, with COG 1.182 +/- 3.126, compared with 1.173 +/- 3.100 for Euler s = 32. Across the two tasks, CFM produces warm starts with lower measured COG than both baselines and uses fewer network evaluations than diffusion.
Chinese Translation
工程逆向设计常受限于两方面:一是受偏微分方程(PDE)约束的优化问题在迭代求解器中计算成本高昂,二是对初始值的敏感性。深度生成模型可在推理阶段无需重新运行模拟器即可生成候选设计方案。生成对抗网络(GAN)可通过一次前向传播完成采样,而扩散模型则需要迭代式的反向时间积分。本文将条件流匹配(CFM)引入 EngiOpt,并在 EngiBench 的结构(beams2d)与热传导(heatconduction2d)基准任务上,采用相同的下游优化协议,将其与条件扩散模型和条件生成对抗网络(cGAN)进行比较。我们采用累积最优性差距(COG)和最终最优性差距(FOG)作为主要指标,评估生成设计作为基于梯度的精细化优化的热启动效果。在被评估的 EngiOpt 实现和两个 EngiBench 任务上,CFM 在两项任务中均取得了最低的 COG、FOG、最大均值差异(MMD)和体积分数偏差。CFM 在 beams2d 和 heatconduction2d 上的平均体积分数偏差分别为 0.4% 和 1.0%,而扩散模型分别为 3.8% 和 11.2%。在 Euler 步数 s = 16 时,CFM 在 beams2d 上达到 53.2 样本/秒,约为在相同计时协议下使用 1000 次网络评估的扩散基线实测吞吐量的 66 倍,其 COG 为 1.182 +/- 3.126,而 Euler 步数 s = 32 时为 1.173 +/- 3.100。在两个任务中,CFM 生成的热启动方案的实测 COG 均低于两个基线,且所需网络评估次数少于扩散模型。
cs.LG / 63 / 2609.00865

MemoryWalker: Stop Training Agents on Contexts They Never Saw

MemoryWalker:停止在智能体从未见过的上下文上训练
J, Zinco, Zhu, Xunjie, Huang, Shen, Wang, Zhenyi, Xie, Pengjun, Ye, Jieping
Abstract
Production agent harnesses such as Claude Code and Qwen-Agent compress context during rollout, but training under compression creates a conditioning problem: every eviction branches the effective history, so the learning object is a tree rather than a sequence. Existing linearizations either retain the rightmost path, causing time-travel leakage, or replay a depth-first traversal, causing train-inference mismatch. We introduce two exact, gradient-equivalent corrections: LogitTree, a segmented K-forward traversal, and a packed 4D attention mask. LogitTree requires K+1 backward passes; the 4D mask requires a custom kernel and white-box eviction records. We also propose SDCC (Self-Distillation for Conditioning Consistency), a single-backward-pass variational relaxation. At each eviction, it minimizes forward KL between the compressed student and a stop-gradient teacher on the reconstructed pre-eviction prefix. A residual per-junction KL of epsilon_KL gives an O(sqrt(epsilon_KL)) bound on the train-deployment total-variation gap. SDCC also applies to black-box harnesses. On seven web-search benchmarks with TC-RAG, AgentFold, MemexRL, Claude Code, and OpenCode, naive training inflates the train-rollout log-probability gap, especially on eviction-heavy batches. The exact methods stay at the no-compression floor, and SDCC substantially closes the gap, with lower logit drift and higher rollout rewards.
Chinese Translation
诸如 Claude Code 和 Qwen-Agent 等生产级智能体框架在推理过程中会对上下文进行压缩,但在压缩条件下进行训练会带来一个条件化问题:每一次上下文淘汰都会使有效历史产生分支,因此学习对象是一棵树而非一个序列。现有的线性化方法要么保留最右侧路径,导致“时间旅行泄漏”,要么重放深度优先遍历,导致训练与推理不匹配。我们提出两种精确且梯度等价的修正方法:LogitTree,一种分段 K 步前向遍历;以及一种打包的 4D 注意力掩码。LogitTree 需要 K+1 次反向传播;4D 掩码则需要自定义内核和白盒淘汰记录。我们还提出了 SDCC(面向条件化一致性的自蒸馏),这是一种仅需单次反向传播的变分松弛方法。在每次淘汰时,它最小化压缩后的学生模型与停止梯度的教师模型在重建的淘汰前前缀上的前向 KL 散度。每个分支点上大小为 epsilon_KL 的残差 KL 为训练与部署之间的总变差差距提供了 O(sqrt(epsilon_KL)) 的上界。SDCC 同样适用于黑盒框架。在结合 TC-RAG、AgentFold、MemexRL、Claude Code 和 OpenCode 的七个网页搜索基准上,朴素训练会扩大训练阶段推理 rolled-out 对数概率差距,尤其是在淘汰密集的批次上。精确方法能保持在无压缩的基线水平,而 SDCC 能显著缩小该差距,同时具有更低的对数偏移和更高的推理奖励。
cs.LG / 64 / 2609.00883

iPINN for Broadband CARS Phase Retrieval: A Framework for Function Approximation and Inverse Modeling Problems in Nonlinear Spectroscopy

用于宽带CARS相位检索的iPINN:一个面向非线性光谱学中函数逼近与逆建模问题的框架
Vulchi, Ravi Teja, Messerschmidt, Carl, Vafaeinezhad, Mohammadsadegh, Junjuri, Rajendhar, Meyer-Zedler, Tobias, Popp, Juergen, Bocklitz, Thomas
Abstract
Phase retrieval in broadband coherent anti-Stokes Raman spectroscopy (BCARS) is an ill-posed inverse problem. The Raman-like signal is encoded in the imaginary part of the resonant susceptibility, which mixes coherently with a non-resonant background (NRB) that varies across acquisitions. We introduce an inverse physics-informed neural network (iPINN) that predicts Lorentzian peak parameters from raw BCARS spectra and reconstructs the resonant susceptibility through a differentiable analytical forward model. A transformer encoder assigns spectral features to 24 learnable peak slots, and a multi-view consistency loss enforces invariance across NRB pattern, NRB strength, and noise. Unlike direct spectral regression approaches, the method retains accuracy under varying acquisition conditions. On a public benchmark, iPINN achieves the lowest error among the tested baselines (MAE 0.016 vs. next-best 0.046). On 28 zero-shot test spectra acquired across seven solvents and four focal positions, accuracy is depth-invariant in five of seven solvents. These results show that inverse parametric prediction with a differentiable physical decoder supports robust phase retrieval across measurement conditions.
Chinese Translation
宽带相干反斯托克斯拉曼光谱(BCARS)中的相位检索是一个病态逆问题。类拉曼信号编码于共振极化率的虚部之中,并与随采集条件变化的非共振背景(NRB)相干混合。我们提出了一种逆物理信息神经网络(iPINN),该网络从原始BCARS光谱预测洛伦兹峰参数,并通过一个可微分的解析正演模型重建共振极化率。一个Transformer编码器将光谱特征分配给24个可学习的峰槽位,同时多视角一致性损失强制模型在NRB模式、NRB强度和噪声变化下保持不变性。与直接光谱回归方法不同,该方法在变化的采集条件下仍能保持精度。在公开基准数据集上,iPINN在所测试的基线方法中取得了最低误差(MAE为0.016,次优方法为0.046)。在跨越七种溶剂和四个焦点位置采集的28个零样本测试光谱上,七种溶剂中有五种达到了深度不变的精度。这些结果表明,结合可微分物理解码器的逆参数化预测能够支持跨测量条件的稳健相位检索。
cs.LG / 65 / 2609.00896

Poisson-Gamma Dynamical Systems with Time-varying Transition Dynamics

具有时变转移动态的泊松-伽马动态系统
Wang, Jiahao, Wang, Yijun, Fang, Nan, Yang, Sikun
Abstract
Bayesian methodologies for handling count-valued time series have gained prominence due to their ability to infer interpretable latent structures and to estimate uncertainties. Among these Bayesian models, Poisson-Gamma Dynamical Systems (PGDSs) are proven to be effective in capturing the evolving dynamics underlying observed count sequences. However, the state-of-the-art PGDS still falls short in capturing the transition dynamics that are commonly observed in real-world count time series. To mitigate this limitation, a PGDS with time-varying transition kernel (TV-PGDS), is proposed to allow the underlying transition matrices to evolve over time. Three specifically-designed Dirichlet Markov chains (Dir-Dir, Dir-Gam-Dir, PR-Gam-Dir) are constructed to accommodate heterogeneous structural mutations within these dependencies. Leveraging Dirichlet-Multinomial-Beta data augmentation techniques, a fully-conjugate and efficient Gibbs sampler is developed to perform posterior simulation. Experiments show that, in comparison with related models, the proposed PGDS achieves improved predictive performance due to its capacity to learn time-varying dependency structure captured by the time-evolving transition matrices.
Chinese Translation
用于处理计数型时间序列的贝叶斯方法因其能够推断可解释的潜在结构并估计不确定性而备受关注。在这些贝叶斯模型中,泊松-伽马动态系统(Poisson-Gamma Dynamical Systems, PGDS)已被证明能够有效捕捉观测计数序列背后的演化动态。然而,最先进的PGDS仍然难以捕捉现实世界计数时间序列中普遍存在的转移动态。为弥补这一局限,本文提出了一种具有时变转移核的PGDS(TV-PGDS),允许底层转移矩阵随时间演化。我们构造了三种专门设计的狄利克雷马尔可夫链(Dir-Dir、Dir-Gam-Dir、PR-Gam-Dir),以适应这些依赖关系中的异构结构突变。借助狄利克雷-多项-贝塔(Dirichlet-Multinomial-Beta)数据增广技术,我们开发了一种完全共轭且高效的吉布斯采样器来执行后验模拟。实验表明,与相关模型相比,所提出的PGDS由于其能够学习由时变转移矩阵所刻画时变依赖结构,因而取得了更优的预测性能。
cs.LG / 66 / 2609.00905

When Metropolis and Hastings Meet Bradley and Terry: Exact MCMC From Preference Voting

当Metropolis和Hastings遇上Bradley和Terry:基于偏好投票的精确MCMC方法
Smogorghevski, Ariel, Rosenfeld, Nir, Romano, Yaniv
Abstract
Sampling from distributions conditioned on desired semantic properties is an emerging challenge in modern generative modeling. Metropolis-Hastings (MH) provides a principled route to conditional sampling, but requires access to exact pointwise target-density evaluations, which are not available in generative settings. Meanwhile, pairwise comparisons by humans or model "judge" are highly accessible and have proved valuable across diverse applications. We introduce Pref-MH, a general exact MH sampler for judge-induced conditional distributions using only stochastic binary pairwise comparisons. Our key observation is that the MH unnormalized density ratio matches the preference odds of the Bradley-Terry (BT) choice model. The central challenge is that while MH requires precise ratio computation, BT judges provide only sampled binary feedback. To this end, we develop a valid accept/reject rule whose resulting Markov chain provably converges to the target distribution. We further show that, for a fixed proposal kernel and budget, Pref-MH is optimal in the Peskun-Tierney sense among this class of exact reversible acceptance rules. Experiments on text generation and molecular design with LLM judges, as well as image generation with VLM judges, demonstrate that Pref-MH provides a practical and flexible approach to conditional sampling when comparative feedback is relatively easy to obtain.
Chinese Translation
在给定目标语义属性的条件下从分布中采样,是现代生成式建模中一个新兴的挑战。Metropolis-Hastings(MH)算法为条件采样提供了一条有原则的路径,但它需要对目标密度进行精确的逐点求值,而这在生成式场景中通常无法获得。与此同时,由人类或模型"评委"给出的成对比较非常容易获取,并已在众多应用中证明了其价值。我们提出了 Pref-MH,一个仅使用随机二值成对比较、针对评委诱导的条件分布的通用精确MH采样器。我们的关键观察是,MH的非归一化密度比与Bradley-Terry(BT)选择模型的偏好胜率相匹配。核心挑战在于:MH要求精确的比值计算,而BT评委仅提供采样的二值反馈。为此,我们设计了一个有效的接受/拒绝规则,其对应的马尔可夫链可证明收敛于目标分布。我们进一步证明,在固定的提议核和预算下,Pref-MH在这类精确可逆接受规则中于Peskun-Tierney意义下是最优的。在使用LLM评委的文本生成与分子设计实验,以及使用VLM评委的图像生成实验中,结果表明:当比较反馈相对容易获取时,Pref-MH为条件采样提供了一种实用且灵活的方法。
cs.LG / 67 / 2609.01034

The Multiple Timescales of Gradient Descent on the Edge of Stability: A Perturbative Derivation of the Central Flow

稳定性边缘上梯度下降的多时间尺度:中心流的微扰推导
Berthier, Raphaël
Abstract
The central flow of Cohen et al. (2025) is an empirically accurate continuous-time model of gradient descent at the edge of stability in deep learning, However, its derivation is heuristic. We propose a perturbative regime in which the central flow is the limit of gradient descent: we assume that the loss decomposes as $f = g + \varepsilon h$; in the limit $\varepsilon \to 0$, the dynamics of gradient descent with learning rate $\eta$ converge to the gradient flow of $h$ constrained to the minimizers of $g$ of sharpness at most $2/\eta$. Our approach is formal rather than rigorous; it treats gradient descent as a singularly perturbed dynamical system in $\varepsilon$. Three timescales emerge: a fast timescale of oscillations along the sharpest direction, an intermediate timescale of the self-stabilization mechanism, and a slow timescale of the dynamics along the minimizers of $g$-the central flow. Using the method of multiple scales, a classical formal method from singular perturbation theory, we derive the expansion of the dynamics in $\varepsilon$: the central flow emerges as the leading-order term in the expansion, while the self-stabilization mechanism appears in the next-order term. We study this mechanism beyond previous analyses: with a single eigenvalue at the edge of stability, we compute the slow drift of the energy of the fluctuations; with several eigenvalues at the edge of stability, we derive the self-stabilization system and explain why fluctuations persist.
Chinese Translation
Cohen等人(2025)提出的中心流(central flow)是深度学习中稳定性边缘处梯度下降的一种经验上高度精确的连续时间模型,但其推导是启发式的。我们提出了一种微扰机制,使中心流成为梯度下降的极限:我们假设损失函数可分解为 $f = g + \varepsilon h$;在 $\varepsilon \to 0$ 的极限下,学习率为 $\eta$ 的梯度下降动力学收敛于 $h$ 的梯度流,且该梯度流被约束在锐度(sharpness)不超过 $2/\eta$ 的 $g$ 的极小值点上。我们的方法是形式化的而非严格证明的:将梯度下降视为关于 $\varepsilon$ 的奇摄动动力系统。由此涌现出三个时间尺度:沿最锐利方向的振荡的快时间尺度、自稳定机制的中间时间尺度,以及沿 $g$ 的极小值点演化的慢时间尺度——即中心流。利用多尺度方法(奇异摄动理论中的一种经典形式化方法),我们推导了动力学关于 $\varepsilon$ 的展开式:中心流作为展开式的首阶项出现,而自稳定机制则出现在下一阶项中。我们在以往分析的基础上进一步研究了这一机制:当只有一个特征值位于稳定性边缘时,我们计算了涨落能量的慢漂移;当有多个特征值位于稳定性边缘时,我们推导了自稳定系统,并解释了涨落为何持续存在。
cs.LG / 68 / 2609.01043

From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion

从截断到承诺:均匀离散扩散中的持久上下文
Hayakawa, Satoshi
Abstract
Uniform-state discrete diffusion models update all tokens in parallel while keeping every position revisable. Even when the commonly used top-$p$ rule leaves only one candidate at a position, that choice affects only the current reverse step and can be revised at the next sampling step. We ask what changes when selected hypotheses instead become persistent context for later predictions. We therefore propose committed reveal sampling (CRS), a training-free sampler that stores selected argmax tokens and inserts them into subsequent model inputs. Our analysis gives a rationale for selecting later and for keeping selected tokens visible. Under the exact forward process, the Bayes error of selecting a clean token cannot increase as noise decreases, while in a simple latent-mode model, keeping the selected token visible helps later parallel predictions agree on the same sequence-level choice. Empirically, paired experiments on Duo-distilled then separate this persistent effect from single-step top-$p$ restriction and scalar temperature scaling. Under the same finalization rule, CRS without top-$p$ truncation reaches lower generative perplexity (GenPPL) than fixed $p=0.95$ and $p=0.9$ baselines across budgets of 8--64 function evaluations (NFE). At 64 NFE, the comparison at matched unigram entropy also gives lower GenPPL for CRS, yielding a more favorable GenPPL--entropy tradeoff. Base Duo shows the same direction in a descriptive comparison, while other diversity and continuation metrics can rank these operating points differently. These results identify support restriction and persistent context as distinct controls of that tradeoff.
Chinese Translation
均匀状态离散扩散模型并行更新所有词元,同时保持每个位置可被修改。即使常用的 top-$p$ 规则在某位置只保留一个候选,该选择也仅影响当前的反向步骤,并可在下一采样步骤中被修改。我们探讨当被选中的假设转而成为后续预测的持久上下文时会发生什么变化。为此,我们提出了承诺揭示采样(Committed Reveal Sampling, CRS),这是一种免训练的采样器,它存储被选中的 argmax 词元并将其插入后续的模型输入中。我们的分析为“延后选择”以及“保持被选中词元可见”提供了理论依据。在精确前向过程下,随着噪声降低,选择干净词元的贝叶斯误差不会增加;而在一个简单的潜变量模式模型中,保持被选词元可见有助于后续并行预测就同一序列级选择达成一致。在实证方面,我们在 Duo-distilled 上进行的配对实验将这种持久效应与单步 top-$p$ 限制以及标量温度缩放区分开来。在相同的最终确定规则下,在 8–64 次函数评估(NFE)的预算范围内,不使用 top-$p$ 截断的 CRS 相比固定 $p=0.95$ 和 $p=0.9$ 的基线达到了更低的生成困惑度(GenPPL)。在 64 NFE 下,在匹配的单字熵条件下进行比较,CRS 同样取得了更低的 GenPPL,呈现出更有利的 GenPPL–熵权衡。基础版 Duo 在描述性比较中表现出相同的方向,而其他多样性和续写指标对这些工作点的排序可能有所不同。这些结果表明,支持集限制与持久上下文是该权衡中两类不同的控制手段。
cs.LG / 69 / 2609.01051

SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations

SAGE:面向子群体的生成式增强方法以缓解虚假相关性
Luo, Yiming, Zhao, Rongqiang, Liu, Jie
Abstract
Spurious correlations pose a significant challenge to the robustness of modern machine learning. The inherent imbalance in dataset distributions often leads traditional Empirical Risk Minimization (ERM) models to rely on majority spurious attributes for classification, resulting in poor performance on minority groups. This problem becomes particularly challenging when the spurious attributes are unavailable. Existing group-label-free methods often upsample minority groups or misclassified real training examples; repeating the same instances can reduce effective diversity and encourage overfitting. To mitigate these spurious correlations from a data-centric perspective in the absence of prior knowledge, we introduce Subpopulation-Aware Generative Enhancement (SAGE), a two-stage generative augmentation framework. Using cluster-derived sub-labels and class labels, we fine-tune a conditional generative model and text encoder, generating targeted synthetic data to fill underrepresented regions in the training set and construct a balanced validation set for last-layer reweighting. We experimentally show that SAGE achieves 89.5%, 85.7%, and 79.1% worst-group accuracy on Waterbirds, CelebA, and MetaShift, respectively, outperforming the best group-label-free baselines by up to 7.7 percentage points.
Chinese Translation
虚假相关性对现代机器学习的鲁棒性构成了重大挑战。数据集分布中固有的不平衡往往导致传统的经验风险最小化(Empirical Risk Minimization, ERM)模型依赖多数群体中的虚假属性进行分类,从而在少数群体上表现不佳。当虚假属性不可用时,这一问题变得尤为棘手。现有的无需群体标签的方法通常对少数群体或被错误分类的真实训练样本进行上采样;然而重复相同的样本会降低有效多样性并导致过拟合。为了在缺乏先验知识的情况下从以数据为中心的视角缓解这些虚假相关性,我们提出了子群体感知生成式增强方法(Subpopulation-Aware Generative Enhancement, SAGE),这是一个两阶段的生成式增强框架。利用基于聚类得到的子标签和类别标签,我们对条件生成模型和文本编码器进行微调,生成有针对性的合成数据,以填补训练集中代表性不足的区域,并构建一个平衡的验证集用于最后一层的重加权。实验表明,SAGE 在 Waterbirds、CelebA 和 MetaShift 数据集上分别取得了 89.5%、85.7% 和 79.1% 的最差群体准确率,相比最佳的无群体标签基线方法最高提升 7.7 个百分点。
cs.LG / 70 / 2609.01072

Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration

让置信度改变,而非预测:面向事后校准的保持预测的修复方法
Kim, Daehwan, Chung, Haejun, Jang, Ikbeom
Abstract
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction. Accuracy captures only the net effect of these changes on correctness, not how often predictions change; the Top-1 Prediction Change Rate (TPCR) instead measures this frequency. We propose Calibrator-Output Repair for Top-1 Decision Preservation (CORD), the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector. From the original and calibrated outputs alone, CORD determines the mass assigned to the original top-1. The calibrated conditional distribution allocates the remaining mass over the other classes, yielding a repaired vector whose own argmax recovers the original prediction. On the calibration split, CORD coordinates the repaired masses to retain the calibrated outputs' mean mass on original predictions whenever attainable. The adapter alters neither the fitted calibrator nor its direct output, fits no additional supervised map, and requires no user- or validation-tuned hyperparameter. Across CIFAR-10/100 and ImageNet-1K, CORD attains zero TPCR by construction and lowers mean ECE, NLL, and Brier relative to the corresponding direct outputs in every dataset; paired gains persist under distribution shift and across calibration-set sizes. CORD thus removes the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair. Our code is available at https://github.com/labhai/CORD.
Chinese Translation
事后校准(post-hoc calibration)用于修正模型报告的置信度,但多类别校准器也可能改变对应的top-1预测。准确率仅反映这些变化对正确性的净效应,而非预测改变的频率;Top-1预测变化率(Top-1 Prediction Change Rate, TPCR)则用于度量这一频率。我们提出了面向Top-1决策保持的校准器输出修复方法(Calibrator-Output Repair for Top-1 Decision Preservation, CORD),这是首个通过修复完整校准概率向量来实现严格预测保持的事后适配器(post-fit adapter)。仅利用原始输出和校准输出,CORD即可确定分配给原始top-1类别的概率质量。校准后的条件分布将其余概率质量分配到其他类别上,从而得到一个修复后的向量,其自身的argmax可恢复原始预测。在校准数据划分上,CORD协调修复后的概率质量,以便在可行的情况下保持校准输出在原始预测上的平均概率质量。该适配器既不改变已拟合的校准器及其直接输出,也不拟合任何额外的监督映射,并且无需用户或验证集调节的超参数。在CIFAR-10/100和ImageNet-1K上,CORD在构造上即实现零TPCR,并且在每个数据集上均相对于相应的直接输出降低了平均ECE、NLL和Brier分数;在分布偏移以及不同校准集规模下,成对增益依然保持。由此,CORD将预测保持约束从校准器拟合中移除,并将原始决策的精确恢复任务交给后续的输出修复。我们的代码见 https://github.com/labhai/CORD。
cs.LG / 71 / 2609.01090

Modelpedia: A Catalog of Model Findings for the Meta-Science of AI

Modelpedia:面向AI元科学的模型发现目录
Bernat, Franciszek, Płudowski, Dawid, Włodarczyk, Michał Jan, Longo, Luca, Zhou, Jianlong, Holzinger, Andreas, Guidotti, Riccardo, Samek, Wojciech, Biecek, Przemysław
Abstract
Scientific knowledge about AI models is produced faster than the community can organize it. Every few months a new foundation model reshapes the field and hundreds of papers, blogs, and technical reports document how each behaves or fails. Yet, these findings remain scattered and effectively unretrievable. To address this gap we present Modelpedia, an automated, LLM-assisted framework that extracts findings about models from published papers, links it to the model, dataset, method, and concept it concerns, and aggregates the result into a searchable public catalog. Applying the prototype to accepted ICLR 2024 and 2025 papers, we extract over a thousand findings and, treating the catalog itself as an object of study, run a meta-analysis of how the community investigates models. Now, we invite the community to explore, contribute to, and build on the open catalog, and to help establish model findings as a shared foundation for the meta-science of AI.
Chinese Translation
关于AI模型的科学知识的产生速度已超过学术界对其加以组织的速度。每隔几个月,一个新的基础模型就会重塑整个领域,同时数百篇论文、博客和技术报告记录着这些模型的行为表现或失效方式。然而,这些研究发现仍然分散各处,实际上无法被有效检索。为填补这一空白,我们提出了Modelpedia——一个自动化、由大语言模型(LLM)辅助的框架,它从已发表论文中提取关于模型的研究发现,将其与所涉及的模型、数据集、方法和概念相关联,并将结果汇总为一个可检索的公共目录。我们将该原型应用于ICLR 2024和2025年的录用论文,提取了一千余条研究发现,并将该目录本身作为研究对象,对学术界如何研究模型进行了元分析。现在,我们邀请学术界探索、贡献并基于这一开放目录进行构建,共同帮助将模型研究发现确立为AI元科学的共同基础。
cs.LG / 72 / 2609.01091

Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation

作为特质方向漂移的潜意识学习:SFT蒸馏下的机制与定向控制
Liu, Zhixuan, Dong, Zhichen, Fan, Yuyu, Li, Xiangtian, Yang, Chao
Abstract
Beyond intended capabilities, model distillation can transfer hidden traits from a teacher. A teacher biased by a system prompt can generate semantically clean training data, such as numeric sequences, that still causes a downstream student to inherit the hidden preference, a phenomenon known as subliminal learning. Prior work has identified several parts of this process. How the signal builds up during training and produces behavioral transfer remains unclear, making targeted mitigation difficult. We propose and validate trait-direction drift as a mechanism for subliminal learning: biased generation creates measurable preference gaps in teacher data, and student-recognizable gaps induce trait-aligned updates during supervised fine-tuning that accumulate into behavioral transfer. Guided by this mechanism, we propose probe-space corridor regularization, a targeted defense that constrains drift along a calibrated trait direction during distillation. The method substantially reduces hidden-trait transfer, preserving task performance: for example, it lowers malicious-response transfer from 29.55% to 6.45% with low main-task accuracy cost, and consistently suppresses animal-preference transfer across the main Qwen setting. The preference-gap, training-trajectory, and intervention evidence links subliminal learning to trait-direction drift and motivates corridor regularization as a targeted control during distillation.
Chinese Translation
除了预期的能力之外,模型蒸馏还可能从教师模型传递隐藏特质。被系统提示词偏置的教师模型可以生成语义上干净无误的训练数据(例如数字序列),但这些数据仍会导致下游学生模型继承隐藏偏好,这一现象被称为潜意识学习(subliminal learning)。已有研究识别了该过程的若干环节,但信号在训练过程中如何累积并产生行为迁移仍不清楚,这使得针对性的缓解措施难以实施。我们提出并验证了特质方向漂移(trait-direction drift)作为潜意识学习的一种机制:有偏生成会在教师数据中产生可测量的偏好差距,而学生模型可识别的这些差距会在监督微调(supervised fine-tuning)过程中诱导与特质对齐的参数更新,并逐步累积为行为迁移。基于该机制,我们提出探针空间走廊正则化(probe-space corridor regularization),这是一种定向防御方法,可在蒸馏过程中约束模型沿经校准的特质方向发生漂移。该方法在保持任务性能的同时显著减少了隐藏特质的传递:例如,它将恶意响应的传递率从29.55%降低至6.45%,且主任务准确率损失很小,并在主要Qwen实验设置中一致地抑制了动物偏好的传递。偏好差距、训练轨迹和干预证据将潜意识学习与特质方向漂移联系起来,并支持将走廊正则化作为蒸馏过程中的一种定向控制手段。
cs.LG / 73 / 2609.01102

Neural Symbollic Regression Using Deep Learning and Sparse Modelling

基于深度学习与稀疏建模的神经符号回归
U, Ravi Kumar, S, Sumitra
Abstract
Symbolic Regression (SR) seeks to find succinct mathematical expressions that represent the fundamental relationships within data, providing interpretability and scientific understanding that exceeds that of black-box models. Nevertheless, traditional methods like Genetic Programming face challenges with scalability and are highly sensitive to noise, while sparse regression techniques such as SINDy rely significantly on predetermined feature libraries. In this work, we present a Neural Symbolic Regression (NSR) framework that treats neural networks as functional preconditioners for symbolic discovery. Our approach uses a decoupled pipeline: a neural network first learns a smooth, noise-robust approximation of the target function in an interaction- aware nonlinear feature space. LASSO is then applied to extract sparse, interpretable closed-form expressions. To improve predictive accuracy and symbolic fidelity by integrating distributed hyperparameter optimization with Ray Tune and ASHA scheduling. Experiments on the Nguyen benchmark suite show that our approach consistently outperforms SINDy and non-tuned neural baselines in RMSE, noise robustness, and out-of-distribution generalization. Ablation studies confirm the significance of feature interactions, neural depth, and tuning strategies. In general, this study presents a scalable and understandable neural-symbolic framework, creating a solid link between neural approximation and the discovery of sparse equations for scientific machine learning.
Chinese Translation
符号回归(Symbolic Regression, SR)旨在寻找简洁的数学表达式来表示数据中的基本关系,其提供的可解释性和科学理解能力超越了黑箱模型。然而,遗传编程等传统方法面临可扩展性难题,且对噪声高度敏感;而SINDy等稀疏回归技术则严重依赖预先确定的特征库。在本工作中,我们提出了一种神经符号回归(Neural Symbolic Regression, NSR)框架,将神经网络视为符号发现的功能性预处理器。我们的方法采用解耦的流水线:神经网络首先在感知交互的非线性特征空间中学习目标函数的平滑、抗噪声近似,随后应用LASSO提取稀疏、可解释的闭式表达式。为提高预测精度和符号保真度,我们集成了基于Ray Tune和ASHA调度策略的分布式超参数优化。在Nguyen基准测试套件上的实验表明,我们的方法在RMSE、噪声鲁棒性以及分布外泛化能力方面始终优于SINDy和未调优的神经基线方法。消融研究证实了特征交互、神经网络深度以及调优策略的重要性。总体而言,本研究提出了一个可扩展且可理解的神经-符号框架,为科学机器学习中的神经近似与稀疏方程发现之间建立了坚实的桥梁。
cs.LG / 74 / 2609.01108

Replicating TRACE: A Practitioner's Guide to Its Threshold and Particle Budget

复现TRACE:其阈值与粒子预算的实践者指南
Chadyuk, Alex, Zhang, Alicia, Kucukates, Roy
Abstract
TRACE (Math & Lienhart, arXiv:2602.01135) reads causal graphs over event types out of a pretrained autoregressive sequence model by thresholding a per-position conditional-mutual-information estimate at a fixed tau. We independently replicate its headline synthetic result: with tau selected on a validation split, mean per-sequence F1 against exact interventional truth reaches 0.90-0.91 at vocabulary size 1000 (paper: 0.91) and 0.86-0.91 from 100 to 2000. First, the optimal threshold is pinned to the truth margin, not to any constant: at every size the errors at tau* straddle the delta = 0.05 margin defining ground truth (missed true edges lie just above it, accepted false ones just below), and the blind optimum lands near delta/2 times the estimator's calibration, confirmed out of sample at 5000. Second, at a single global threshold TRACE mostly recovers a direct, adjacent-influence graph: lag-1 true edges are recalled at 0.97-0.99, while true edges at lag 2 or more read orders of magnitude lower---the reading-scale price of randomizing mediating positions, which an exact test of direct causal effect requires when the truth is unknown. A per-lag threshold family recovers a third to a half of lag-2 truth; on lag-uniform data one validated threshold recalls every lag at 0.40-0.87, 8-26 pp below an atomic-intervention control at lags 3-6. Third, the default lag decay of the paper's synthetic benchmark concentrates about 85% of interventional truth at lag 1 and pushes the rest below the estimator's noise floor, so headline F1 there certifies lag-1 recovery only and conflates the benchmark's skew with the algorithm's own limit; a flatter decay separates the two. Fourth, F1 saturates from N = 2 particles at the selected threshold---a property of the threshold's margin over the noise floor, not of the estimator, which converges as N^(-1/2). We distill five practitioner rules.
Chinese Translation
TRACE(Math & Lienhart, arXiv:2602.01135)通过在一个固定的 tau 处对逐位置的条件互信息估计进行阈值化,从预训练的自回归序列模型中读取事件类型之间的因果图。我们独立复现了其主要的合成实验结果:在验证集上选择 tau 时,在词表规模为 1000 的条件下,相对精确干预真值的逐序列平均 F1 达到 0.90–0.91(论文报告为 0.91),在词表规模 100 至 2000 的范围内为 0.86–0.91。第一,最优阈值取决于真值间隔,而非某个常数:在每一规模下,tau* 处的错误均跨越定义真值的 delta = 0.05 边界(被漏掉的真边略高于该边界,被接受假边略低于该边界),且盲目选择的最优值落在估计器校准的约 delta/2 附近,并在词表规模 5000 的样本外得到验证。第二,在单一全局阈值下,TRACE 主要恢复的是直接的、相邻影响的图:滞后 1 的真边召回率达 0.97–0.99,而滞后 2 及以上的真边读数低若干个数量级——这是对中介位置进行随机化所需的读数尺度代价,当真值未知时,精确的直接因果效应检验需要这种随机化。逐滞后阈值族可恢复三分之一的滞后 2 真边;在滞后均匀的数据上,一个经验证的阈值对所有滞后的召回率为 0.40–0.87,在滞后 3–6 处比原子干预对照低 8–26 个百分点。第三,论文合成基准的默认滞后衰减将约 85% 的干预真值集中于滞后 1,并将其余部分压至估计器的噪声底之下,因此该基准上的主报告 F1 仅证明了滞后 1 的恢复能力,并将基准的偏斜与算法本身的局限混为一谈;更平缓的衰减可将两者区分开。第四,在所选阈值下,F1 从 N = 2 个粒子起即达到饱和——这是阈值相对噪声底的间隔所具有的性质,而非估计器的性质,估计器本身按 N^(-1/2) 收敛。我们总结出五条实践者准则。
cs.LG / 75 / 2609.01126

When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting

在线自适应在边缘端何时有益?面向时间序列预测的预热、学习率选择与资源权衡的无泄漏评估
Fujimoto, Takumi, Nishi, Hiroaki
Abstract
Online adaptation can help edge time-series forecasting under distribution drift, but its measured benefit is sensitive to evaluation choices. We study six public multivariate streams, including building-sensor and smart-meter data, under a leakage-free streaming protocol. We identify two additional sources of comparison bias. First, the warmup budget of the static baseline has a two-sided effect: insufficient warmup undertrains the baseline, whereas excessive warmup can degrade its pre-drift generalization. Across six dataset-backbone settings, the estimated adaptation benefit changes by 3.0 to 18.8 percentage points (pp) over the 1,000-20,000-step warmup range. Second, comparing SGD with momentum (SGD+m) and Adam at a shared default learning rate conflates optimizer quality with rate sensitivity. We select both the warmup budget and each optimizer's online rate using a held-out pre-drift validation slice without accessing test data. Under this validation-only procedure, Adam outperforms SGD+m in 310 of 360 evaluated cells, while 4 Adam cells remain below the static baseline. We further characterize accuracy against adaptation-state memory and A100-measured per-update latency for full, head-only, and calibration-based adaptation. In the evaluated PatchTST frontier settings, several parameter-efficient variants are nondominated on the adaptation-state-memory axis. Smart-meter analyses also show that reported gains depend on meter-selection rules. These findings support a validation-only commissioning procedure, while target-device latency and energy remain to be measured. Code, data, and all reported numbers: https://github.com/keiotakmin/tsf-edge-adaptation.
Chinese Translation
在线自适应(online adaptation)有助于边缘端时间序列预测应对分布漂移,但其所测得的收益对评估方式的选择十分敏感。我们在无泄漏的流式协议下研究了六个公开的多元数据流,包括楼宇传感器数据和智能电表数据。我们识别出两个额外的比较偏差来源。第一,静态基线的预热预算具有双向影响:预热不足会使基线训练不充分,而预热过多可能损害其漂移前的泛化能力。在六个数据集-骨干网络组合设置下,估计的自适应收益在1,000至20,000步的预热范围内变化幅度为3.0至18.8个百分点。第二,在共享默认学习率下比较带动量的SGD(SGD+m)与Adam,会将优化器本身的优劣与对学习率的敏感性混为一谈。我们使用漂移前预留的验证片段来选择预热预算以及各优化器的在线学习率,且不访问测试数据。在这一仅基于验证数据的流程下,Adam在360个评估单元格中的310个优于SGD+m,而仍有4个Adam单元格低于静态基线。我们进一步刻画了全参数自适应、仅头部自适应以及基于校准的自适应在精度与自适应状态内存、A100实测的单次更新延迟之间的权衡关系。在所评估的PatchTST前沿设置中,若干参数高效变体在自适应状态内存维度上是非支配的。智能电表分析还表明,所报告的收益依赖于电表选择规则。这些发现支持一种仅基于验证数据的部署调试流程,而目标设备上的延迟与能耗仍有待测量。代码、数据及所有报告数值:https://github.com/keiotakmin/tsf-edge-adaptation。
cs.LG / 76 / 2609.01129

Scaled Idempotence in Transformer Attention: Paired OV Geometry and Shared-Value Algebras

Transformer注意力中的缩放幂等性:配对OV几何与共享值代数
Feng, Jiming, Li, Junliang
Abstract
We identify a recurrent algebraic regularity in Transformer attention: a sparse subset of effective OV operators $T=OV^\top$ nearly closes under composition, $T^2\approx\alpha T$. Across six pretrained endpoints spanning 2.8B--235B parameters, 3.98--8.00% of heads reach squared closure alignment $\mathcal{P}\geq0.9$, while no matched within-layer O/V mismatch does. An exact principal-coordinate factorization, $T=Q_OKQ_V^\top$ and $T^2=Q_O(KDK)Q_V^\top$, separates within-support transport from read--write return geometry. Across all 7,304 heads in nine MHA/GQA models, scrambling only the orientation of $K$ while preserving singular values, norms, factor spans, and principal angles reduces median closure from 0.336 to $1.04\times10^{-4}$; trained orientation wins for 98.64% of heads and in every layer. Constructive searches show that high closure is feasible in every surveyed layer, but usually not attained. Retrospective trajectories in three independently trained lineages further separate broadly available capacity from the orientations attained by final strong heads. Under exact value sharing, headwise closure extends to a right-action algebra, $T_iT_j=\alpha_jT_i$. Seven-model experiments verify the approximate law and reveal distinct oblique projections with a shared value-defined kernel. These results characterize scaled idempotence as a sparse trained orientation within broadly available geometric capacity and show how value sharing extends a headwise relation into a local operator algebra.
Chinese Translation
我们在Transformer注意力中识别出一种反复出现的代数规律性:一个稀疏的有效OV算子子集 $T=OV^\top$ 在复合运算下近似闭合,即 $T^2\approx\alpha T$。在涵盖2.8B至235B参数的六个预训练模型中,3.98%至8.00%的注意力头达到平方闭合对齐度 $\mathcal{P}\geq0.9$,而任何匹配的同层O/V失配均未达到该水平。通过精确的主坐标分解 $T=Q_OKQ_V^\top$ 和 $T^2=Q_O(KDK)Q_V^\top$,可以将支撑内的信息传递与读写返回几何分离开来。在九个MHA/GQA模型的全部7,304个注意力头上,仅打乱 $K$ 的方向(同时保持奇异值、范数、因子跨度和主角度不变),会使中位闭合度从0.336降至 $1.04\times10^{-4}$;训练所得方向在98.64%的注意力头以及每一层中均占优。构造性搜索表明,在每个被考察的层中实现高闭合度都是可行的,但通常并未被实际达到。三个独立训练谱系的回溯轨迹进一步区分了广泛可获得的容量与最终强注意力头实际达到的方向。在精确值共享条件下,逐头闭合可扩展为一种右作用代数 $T_iT_j=\alpha_jT_i$。七个模型的实验验证了该近似规律,并揭示出具有共享值定义核的相异斜投影。这些结果将缩放幂等性刻画为广泛可获得的几何容量中的一种稀疏训练方向,并展示了值共享如何将逐头关系扩展为局部算子代数。
cs.LG / 77 / 2609.01158

Superposed Latent Autoencoder

叠加隐变量自编码器
Zhao, Quanling, Yang, Jiaying, Zhang, Tianqi, Hao, Ziyang, Asgarinejad, Fatemeh, Ponzina, Flavio, Rosing, Tajana
Abstract
Autoencoders typically meet tight latent-memory budgets by making each latent representation smaller, sacrificing representational capacity. We ask a different question: can multiple wider latents be stored together instead? We introduce the Superposed Latent Autoencoder (SLAE), which preserves high-capacity latent representations while sharing storage through learned superposition. SLAE transforms latents into storage-friendly codes, binds them with randomized keys, superposes multiple codes into a single memory tensor, and learns to recover each latent before decoding. Under the same storage budget, SLAE replaces irreversible dimensional bottlenecks with structured interference that can be suppressed. Across CIFAR-10/100, SVHN, STL-10, Tiny ImageNet, and a wide range of memory budgets, SLAE substantially improves the reconstruction--memory tradeoff, reducing reconstruction error by up to 56% over conventional autoencoders at matched storage. Further analysis shows that SLAE's advantage comes from making wider representations usable under the same storage budget. These gains also extend beyond reconstruction: the information preserved by SLAE improves downstream classification by up to 16.79 percentage points under the same memory budget. Our results suggest a new principle for representation compression: instead of making every latent smaller, keep representations wide and let them share memory.
Chinese Translation
自编码器通常通过减小每个隐表示的尺寸来满足严格的隐空间内存预算,从而牺牲了表示能力。我们提出了一个不同的问题:能否将多个更宽的隐表示共同存储?我们提出了叠加隐变量自编码器(Superposed Latent Autoencoder, SLAE),它在通过学习到的叠加共享存储的同时,保留了高容量的隐表示。SLAE 将隐表示转换为便于存储的编码,用随机化密钥与其绑定,将多个编码叠加为单一的记忆张量,并在解码前学习恢复每个隐表示。在相同的存储预算下,SLAE 用可被抑制的结构化干涉取代了不可逆的维度瓶颈。在 CIFAR-10/100、SVHN、STL-10、Tiny ImageNet 以及广泛的内存预算范围内,SLAE 显著改善了重建-内存权衡,在相同存储条件下相比传统自编码器最多可将重建误差降低 56%。进一步分析表明,SLAE 的优势在于使更宽的表示能够在相同存储预算下得以利用。这些收益也超越了重建任务本身:在相同内存预算下,SLAE 保留的信息可将下游分类准确率最多提升 16.79 个百分点。我们的结果为表示压缩提出了一条新原则:与其让每个隐表示都变小,不如保持表示的宽度并让它们共享内存。
cs.LG / 78 / 2609.01161

CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs

CopyShield:大语言模型版权防御的跨层次基准测试
Alshehyari, Maryam, Chauhan, Dushyant Singh, Poppi, Samuele, Takac, Martin, Lahlou, Salem, Lukas, Nils
Abstract
Large language models can reproduce memorized text verbatim, yet copyright defenses are usually evaluated under incompatible protocols. We introduce CopyShield, a controlled benchmark comparing three representative defenses at distinct intervention levels: contrastive decoding (output), Direct Preference Optimization (behavioral), and activation intervention (representation). We evaluate CopyShield on two model families, LLaMA-3.1-8B and Mistral-7B-v0.3, using controlled memorization over five public-domain books and a shared protocol measuring literal leakage, calibrated non-literal leakage, utility, and degeneracy. Across these methods, intervention level is associated with distinct compliance-utility trade-offs. On LLaMA-3.1-8B, contrastive decoding remains near-degeneracy-free (0-2%) but reaches a literal-suppression floor at NV-Recall 0.192-0.203. DPO nearly eliminates literal leakage (0.263 to 0.002) but induces paraphrase-loop degeneracy in 58% of QA outputs, with no utility gain over the SFT baseline. Activation intervention attains the lowest non-literal flagging rate (1/200) by blocking 84% of non-literal queries before generation. Human evaluation confirms that DPO has low coherence, whereas activation lowers perceived copyright risk through broad refusal. On Mistral-7B-v0.3, the output- and representation-level patterns persist, while DPO degeneracy falls to 10-14%, showing that its severity is model-dependent. Together, CopyShield provides cross-level reference baselines and identifies targeted non-literal suppression as an open challenge. The code is available at https://github.com/spotai-mbzuai/CopyShield.git.
Chinese Translation
大语言模型能够逐字复现其记忆的文本,然而版权防御方法通常在不兼容的评估协议下进行评测。我们提出了CopyShield,这是一个受控基准,用于比较三种在不同干预层次上的代表性防御方法:对比解码(输出层)、直接偏好优化(行为层)和激活干预(表示层)。我们在两个模型家族(LLaMA-3.1-8B和Mistral-7B-v0.3)上评估CopyShield,通过对五本公有领域书籍进行受控记忆化实验,并采用统一的评测协议来衡量字面泄露、校准的非字面泄露、实用性和退化性。在这些方法中,干预层次与不同的合规-实用性权衡相关联。在LLaMA-3.1-8B上,对比解码几乎不会引发退化(0-2%),但其字面抑制存在下限,NV-Recall为0.192-0.203。DPO几乎消除了字面泄露(从0.263降至0.002),但在58%的问答输出中引发了释义循环退化,且相比SFT基线没有实用性提升。激活干预通过在生成前拦截84%的非字面查询,实现了最低的非字面标记率(1/200)。人工评估证实,DPO的连贯性较低,而激活干预则通过广泛的拒答降低了感知的版权风险。在Mistral-7B-v0.3上,输出层和表示层的模式保持一致,而DPO的退化率降至10-14%,表明其严重程度依赖于具体模型。总体而言,CopyShield提供了跨层次的参考基线,并将有针对性的非字面抑制确定为一个开放性挑战。代码可在 https://github.com/spotai-mbzuai/CopyShield.git 获取。
cs.LG / 79 / 2609.01170

Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training

预刻 niches:早期大语言模型训练中模块化任务分区的形成动力学
Li, Guangqi, Li, Yongxin
Abstract
Large language models exhibit a modular internal organization that mirrors well-studied functional networks of the human brain, but how this organization forms during training is unknown: prior work has characterized finished models, not the formation process. We track formation step by step: we train a Pythia-410M model from scratch (two trajectories, bf16 and fp32) and run attribution patching at every step, alongside probes for gradient norms, effective updates, weight norms, and first-order loss decomposition across 14 tasks in four cognitive domains. Three findings. First, the modular map is pre-carved: before any learning, the dominant task pair already overlaps at ~3.6x the attribution substrate (a task-independent baseline), and its layer-0 concentration is an architecture-level constant on this model family. Second, the partition locks in through two sharp jumps whose amplitudes do not track the learning-rate schedule (the second reaching 20.4 sigma quiet-window / 6.2 sigma global), accompanied by gradient-level relative deprivation--winners receive 2.25->2.73x the loser's gradient supply, 9.5-11.5 standard deviations below a random control--that does not propagate to updates or weights. Third, deviation from the substrate appears only in the domain being learned, consistent with the hypothesis that modularity tracks learning. We close by separating the feature-level account we can defend from the mechanistic questions we cannot, and we pre-register the scale-threshold hypothesis behind our ongoing 2.8B experiments.
Chinese Translation
大语言模型呈现出一种模块化的内部组织结构,这与已被深入研究的人脑功能网络相呼应,但这一组织结构在训练过程中如何形成仍属未知:已有工作刻画的是训练完成的模型,而非形成过程本身。我们逐步追踪该形成过程:从零开始训练一个 Pythia-410M 模型(两条训练轨迹,分别为 bf16 和 fp32),在每一步运行归因补丁(attribution patching),同时探测梯度范数、有效更新、权重范数,以及四个认知领域中 14 个任务的一阶损失分解。我们得到三项发现。第一,模块化地图是预先刻划好的:在任何学习发生之前,占主导地位的任务对已经在归因基底(一种与任务无关的基线)上达到约 3.6 倍的重叠,且其第 0 层的集中度在该模型家族上是架构层面的常数。第二,任务分区通过两次剧烈跳变锁定,其幅度并不跟随学习率调度(第二次跳变达到 20.4 倍安静窗口标准差 / 6.2 倍全局标准差),并伴随梯度层面的“相对剥夺”现象——胜者获得输家 2.25 倍至 2.73 倍的梯度供给,比随机对照低 9.5–11.5 个标准差——且该现象不会传播到更新或权重层面。第三,对基底的偏离仅出现在正在被学习的领域中,这与“模块化追随学习”的假设一致。最后,我们区分了我们能够辩护的特征层面的解释与尚无法回答的机制性问题,并为正在进行的 2.8B 实验背后的规模阈值假设进行了预注册。
cs.LG / 80 / 2609.01194

Births are difficult to predict even with rich survey and full-population register data

即使拥有丰富的调查数据和全人口登记数据,生育仍然难以预测
Sivak, Elizaveta, Cantrell, Emily M., Emery, Thomas, Garcia-Bernardo, Javier, Hafner, Flavio, Karpinska, Kasia, Lüken, Malte, Mendrik, Adrienne, Mulder, Joris, Ren, Hanzhang, Satish, Varun, Verhagen, Mark, Maineri, Angelica M., Pankowska, Paulina, Ghany, Jasmin Abdel, Arpino, Bruno, Cassani, Giovanni, Hellstrand, Julia, Ivanova, Katya, Kuikka, Sanni, Macanovic, Ana, Rahal, Charles, Tropf, Felix C., Veen, Roland J., Walasek, Nicole, van Wijk, Daniël, Wright, Kelsey Q., Zagheni, Emilio, Abbink, Henry, Aliverti, Emanuele, Amestoy, Matteo, Atav, Tilbe, Barban, Nicola, Billingsley, Sunnee, Booij, Goan J., Boucherie, Louis, Broos, Yael, Chang, Li Ya, Chiu, Jamie C., Comolli, Chiara Ludovica, Cule, Boris, Fang, Qixiang, Feehan, Dennis M., Ganly, Rachel, Gielens, Erwin, Martinez, Rolando M. Gonzales, Gradassi, Andrea, Guerra-Urzola, Rosember, Guerra-Urzola, Mario, Guerrier, Stéphane, Hassan, Enamul, Haverhoek, Vincent A., Hendrickson, Andrew T., Howard, Amber, Jin, Yuxuan, Kapoor, Sayash, van Kesteren, Erik-Jan, Klooster, Iris ten, Labussiere, Marie, Liu, Lydia T., Liu, Tiffany, Maghout, Adam, Meneghello, Simone, Mohr, Lasse, Mulder, Clara H., Newman, Saul J., Nisén, Jessica, Norden, Janis, Odgaard, Mikkel, Omenti, Riccardo, Ozdemir, Ozancan, Pao, Christina, Park, Paige, Penta, Gaia, Perdomo, Juan C., Pial, Tanzir, Piraccini, Alessio, Querin, Federica, Rao, Ziwei, Rellama, Christian, Remund, Adrien, Richert, Frederieke, van de Rijt, Arnout, Kandroodi, Mojtaba Rostami, Rotman, Stijn J., Sage, Lucas, Savcisens, Germans, Schwanitz, Katrin, Skiena, Steven, Spata, Alessandro, Stadtfeld, Yannick, Stroebl, Benedikt, Tedesco, Gaetano, Theelen, Mathilde, Tori, Gianluca, Tun-Mendicuti, Abigail, Tyagi, Rishabh, Vafa, Keyon, Vecchietti, Luiz Felipe, Vecgaile, Linda, Vermeulen, Willem R. J., Feser, Maria-Pia Victoria, Voirol, Lionel A., Volker, Thom B., Wang, Xinran, Yan, Jiani, Zhao, Xinyi, Zhou, Flora, Zilincikova, Zuzana, Nissim, Malvina, Salganik, Matthew J., Stulp, Gert
Abstract
Major life events have proven difficult to predict. Does this reflect limits of theory, data, and algorithms, or the large role of chance? We examine one outcome - having a child within three years - through a near-ideal setting for prediction: a data challenge where 147 researchers predicted births for Dutch residents aged 18-45, using survey data and full-population registers. Methods ranged from logistic regression to a large language model and transformers. Predictions were moderately accurate (best F1: register 0.59, survey 0.76); advanced models did not outperform classical ones; and the larger registers did not beat the survey. Simulating the stochastic biology of conception and pregnancy, we estimated a predictive ceiling (survey F1 ~ 0.86-0.94, register 0.88-0.96). Observed performance falls short of this ceiling, implicating imperfect data, methods, and unmodelled chance, while the ceiling itself shows that chance in reproduction alone sets a non-trivial limit on predicting individual lives.
Chinese Translation
重大生活事件已被证明难以预测。这究竟反映了理论、数据和算法的局限,还是偶然因素的重要作用?我们通过一个近乎理想的预测情境研究了其中一种结果——三年内生育子女:在这场数据挑战赛中,147名研究者利用调查数据和全人口登记数据,对18至45岁荷兰居民的生育情况进行预测。预测方法涵盖从逻辑回归到大语言模型和Transformer等多种技术。预测结果仅达到中等准确度(最佳F1值:登记数据0.59,调查数据0.76);先进模型并未超越经典模型;样本量更大的登记数据也未胜过调查数据。通过模拟受孕与妊娠的随机生物学过程,我们估计了预测的上限(调查数据F1约为0.86-0.94,登记数据为0.88-0.96)。观测到的预测表现低于该上限,这表明数据不完善、方法不足以及未建模的偶然因素在其中起了作用;而上限本身则表明,仅生殖过程中的偶然性就为预测个人生活设置了不容忽视的限制。
cs.LG / 81 / 2609.01212

Recent Developments in Transformer Inference Deployment on FPGA Platforms: A Survey

FPGA平台上Transformer推理部署的最新进展:综述
Blankestijn, Arjan, Odyurt, Uraz, Yousefzadeh, Amirreza
Abstract
With the rapid and continuous growth in the incorporation of machine learning models based on the Transformer architecture, capable deployment is in high demand. In this context, capable deployment refers to operational performance aspects, e.g., throughput and latency, as well as efficiency aspects, e.g., energy consumption. When it comes to the task of inference using such models, purpose-built hardware accelerators provide a lucrative alternative to common deployment choices, such as Central Processing Units (CPUs) and Graphics Processing Units (GPUs). The Field Programmable Gate Array (FPGA) platforms category is an example of such alternative accelerators, promising implementation flexibility, energy efficiency, improved latency and suitability for on-site deployment. We investigate the most recent advances, trends, and design choices for Transformer inference on FPGA platforms. We perform a systematic literature review, extracting and delving into preferred techniques for implementation and optimisation. This study and the provided taxonomy of topics could act as a guide for researchers from the academia and industry alike.
Chinese Translation
随着基于Transformer架构的机器学习模型的快速持续增长,高性能的部署需求日益迫切。在此背景下,高性能部署既指运行性能方面的表现(如吞吐量和延迟),也包括效率方面的表现(如能耗)。在使用此类模型进行推理的任务中,专用硬件加速器为常见的部署选择(如中央处理器CPU和图形处理器GPU)提供了一种颇具吸引力的替代方案。现场可编程门阵列(FPGA)平台正是此类替代加速器的一个典型代表,其具有实现灵活性、能效高、延迟低以及适合现场部署等优势。本文研究了Transformer在FPGA平台上推理的最新进展、趋势和设计选择。我们进行了系统性文献综述,提取并深入分析了实现与优化的主流技术。本研究及所提供的主题分类体系可为学术界和产业界的研究人员提供参考指南。
cs.LG / 82 / 2609.01215

REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs

REFACTOR-VLA:类型化运动程序的无监督技能库学习
Shaik, Riyaaz, Venkataraman, Chandru
Abstract
Most vision-language-action (VLA) models -- OpenVLA, $\pi_0$, RT-2, RDT-1B -- are monolithic: they emit raw motor commands or short action chunks without organizing behavior into reusable abstractions, so they degrade on long-horizon tasks and resist interpretation. Existing skill-discovery methods sidestep the core question of when two action sequences are behaviorally equivalent, either clustering contrastive embeddings or delegating the judgment to a language model uncalibrated to the robot's dynamics. We introduce REFACTOR-VLA, a wake/sleep system for learning reusable skills. Its sleep phase clusters motor-program fragments under a Behavioral-Equivalence Kernel (BEK) computed from rollouts of a learned latent world model $M_\phi$; its wake phase emits typed lambda terms over a Hindley--Milner-inspired vocabulary, consumed by a library-conditioned rectified-flow action decoder. Abstractions are admitted only if they pass Minimum Description Length and return-preservation gates. On LIBERO we report two findings. First, enlarging the world model from 188M to 430M parameters worsened performance on 4 of 4 suites, so capacity alone does not help. Second, the training objective matters far more: adding an auxiliary supervised contrastive (InfoNCE) loss during world-model warmup substantially improves sleep-phase clustering, giving Normalized Mutual Information at $n=3$ seeds of $0.462 \pm 0.021$ (object), $0.867 \pm 0.025$ (spatial), $0.915 \pm 0.013$ (goal) and $0.754 \pm 0.010$ (LIBERO-10), and beating the strongest published baseline on all 4 suites by a mean $\Delta = +0.184$. Across providers ($n=12$) the 95% bootstrap confidence interval for mean pairwise NMI is $[0.683, 0.729]$ (mean $0.705$). The sleep phase also yields the first real-LIBERO task-language library: the decoder uses 2 of 3 admitted abstractions and rewrites all 256 sampled demonstrations.
Chinese Translation
大多数视觉-语言-动作(VLA)模型——如 OpenVLA、$\pi_0$、RT-2、RDT-1B——都是单体式的:它们直接输出原始运动指令或短动作片段,而不将行为组织成可复用的抽象,因而在长时程任务上性能下降,且难以解释。现有的技能发现方法回避了“两个动作序列在行为上何时等价”这一核心问题,要么对对比学习得到的嵌入进行聚类,要么将这一判断交由一个未经机器人动力学校准的语言模型完成。我们提出了 REFACTOR-VLA,一个用于学习可复用技能的清醒/睡眠(wake/sleep)系统。其睡眠阶段在一个由学习到的潜在世界模型 $M_\phi$ 的回放(rollout)计算得到的行为等价核(Behavioral-Equivalence Kernel, BEK)下对运动程序片段进行聚类;其清醒阶段则基于受 Hindley–Milner 类型系统启发的词表生成带类型的 lambda 项,供一个以技能库为条件的整流流(rectified-flow)动作解码器使用。只有通过最小描述长度(Minimum Description Length)和回报保持(return-preservation)双重门控的抽象才被纳入技能库。在 LIBERO 基准上,我们报告了两项发现。第一,将世界模型从 188M 扩大到 430M 参数后,4 个任务套件上的性能全部变差,说明单纯增加容量并无帮助。第二,训练目标的影响要大得多:在世界模型预热阶段加入辅助的监督对比学习(InfoNCE)损失显著改善了睡眠阶段的聚类效果,在 $n=3$ 个随机种子下取得的归一化互信息(Normalized Mutual Information, NMI)分别为 $0.462 \pm 0.021$(object)、$0.867 \pm 0.025$(spatial)、$0.915 \pm 0.013$(goal)和 $0.754 \pm 0.010$(LIBERO-10),并在全部 4 个套件上以平均 $\Delta = +0.184$ 超过已发表的最强基线。跨提供商($n=12$)的平均成对 NMI 的 95% bootstrap 置信区间为 $[0.683, 0.729]$(均值 $0.705$)。睡眠阶段还产生了首个真实 LIBERO 任务-语言技能库:动作解码器使用了所纳入的 3 个抽象中的 2 个,并重写了全部 256 条采样的示范数据。
cs.LG / 83 / 2609.01231

Multi-Head Self Attention is a Parameter Identification Mechanism

多头自注意力机制是一种参数可辨识性机制
Morrow, W. Ross
Abstract
We prove that a multi-head scaled dot product attention can be viewed as a parameter identification strategy. The ratio of unidentified parameters to the total number of parameters scales like the reciprocal of the number of heads ($1/2 \to 1/(2H)$), meaning models with more heads are structurally more identified. A subtle side effect of the mathematics observation that attention can never be fully identified. Similarly we also show that some bias terms can have no effect on softmax-based attention layers in both the single- and multiple-head settings, though this is mostly a curiosity that should have a marginal effect on model size and model training/prediction efficiency. We also touch on modern improvements to transformers including RoPE and GQA from this perspective, illustrating how those as well can improve the ratio of ``meaningful'' parameters to all parameters. Simple numerical examples demonstrate that training can indeed involve updates that overlap model-invariant subspaces that arise from a lack of identification. As part of our experiments we use a ``rebalancing'' approach that can ``fix'' updates that overlap unindentified subspaces but do not try to present evidence this should actually be adopted. Instead we simply view our numerical results as exploring and confirming the theoretical results. As a whole we discuss a purely mathematical/statistical explanation, identification, for why specific architectural choices in transformers may have improved performance.
Chinese Translation
我们证明多头缩放点积注意力(multi-head scaled dot product attention)可以被视为一种参数辨识策略。不可辨识参数占总参数数量的比例以头数的倒数尺度缩减(从1/2降至1/(2H)),这意味着具有更多头的模型在结构上具有更强的可辨识性。这一数学观察带来的一个微妙副作用是:注意力机制永远无法被完全辨识。类似地,我们还证明在单头和多头的设置下,某些偏置项对基于softmax的注意力层没有任何影响,尽管这主要是一个奇特的性质,对模型大小和模型训练/预测效率仅有边际影响。我们还从这一视角探讨了transformer的现代改进方法,包括RoPE和GQA,说明这些方法同样能够提高“有意义的”参数占全部参数的比例。简单的数值示例表明,训练过程中确实可能涉及与因缺乏可辨识性而产生的模型不变子空间相重叠的更新。作为实验的一部分,我们采用了一种“再平衡”方法来“修正”与不可辨识子空间重叠的更新,但我们并不试图提供证据表明该方法应被实际采用;我们只是将数值结果视为对理论结果的探索和验证。总体而言,我们讨论了一种纯数学/统计学的解释——可辨识性——来说明为什么transformer中特定的架构选择可能会带来性能提升。
cs.LG / 84 / 2609.01244

Post-Training Science for Supervised Fine-Tuning

面向监督微调的后训练科学
O'Neill, Charles, Jayasekara, Mudith, Partridge, Harry
Abstract
Every supervised fine-tuning run forces the same chain of decisions, such as learning rate, batch size, LoRA or full fine-tuning, how many epochs, which optimiser, and what data to feed the model. Each of these is typically rediscovered from scratch for every new model and dataset. Here we measure them under one instrument: a sweep that varies one lever at a time, and spans dense and mixture-of-experts models in two families (Qwen3 and Llama), on four real-world customer SFT datasets, for both LoRA and full fine-tuning. These datasets give a controlled testbed: each task carries an evaluation built with the customer, and its training data is produced by iterative supervised fine-tuning that refines model outputs until they pass that evaluation, so the supervised target is internally consistent and the task judge we report against is the criterion the data was built to satisfy. We ask how the optimal learning rate and batch size move with model scale, family, and data, and whether one selection rule transfers across them; what LoRA trades against full fine-tuning, and how its rank and alpha set what the adapter can learn; whether validation loss (or other metrics, such as loss landscape flatness) faithfully ranks downstream quality; whether post-training gains scale with model size and data volume, on a model ladder extended through mixtures-of-experts to 235B parameters; how many epochs to train before general instruction-following erodes; and whether a geometry-aware optimiser improves on AdamW. Each recommendation is paired with a measure of its uncertainty.
Chinese Translation
每一次监督微调(SFT)运行都迫使我们做出一整串相同的决策,例如学习率、批大小、选择LoRA还是全参数微调、训练多少轮(epoch)、使用哪种优化器,以及向模型输入什么数据。这些决策通常在每遇到新模型和新数据集时都要从头重新摸索。本文在一套统一的测量框架下对这些决策进行了系统评估:我们设计了一个每次只变动单一变量(lever)的扫描实验,涵盖两个模型系列(Qwen3和Llama)的稠密模型与混合专家模型,在四个真实的客户SFT数据集上,分别对LoRA和全参数微调进行测试。这些数据集提供了一个受控的测试平台:每个任务都带有与客户共同构建的评估指标,其训练数据通过迭代式监督微调生成——即不断修正模型输出直至其通过该评估——因此监督目标在内部是一致的,且我们报告所依据的任务评判标准正是数据构建时所针对的准则。我们研究以下问题:最优学习率和批大小如何随模型规模、模型系列和数据而变化,是否存在可跨这些维度迁移的选参规则;LoRA相对全参数微调做出了哪些权衡,其秩(rank)和alpha超参数如何设定适配器(adapter)的学习能力;验证损失(或其他指标,如损失景观的平坦度)能否忠实地排序下游任务质量;在一个扩展到2350亿参数的混合专家模型阶梯上,后训练收益如何随模型规模和数据量扩展;在通用指令遵循能力退化之前应训练多少轮;以及几何感知优化器能否优于AdamW。每一条推荐结论都附带有相应的不确定性度量。
cs.LG / 85 / 2609.01245

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

更多探索,更少漂移:仅结果反馈的强化学习足以训练长程交互智能体
Pu, Liming, Li, Xiaoxia, Liu, Yifu, Cao, Teng, Yang, Bin
Abstract
Reinforcement learning is a natural way to post-train LLM agents for long-horizon interactive tasks judged only by end-of-task verification, yet a shared belief holds that outcome-only RL soon hits a ceiling on small open models. Recent work therefore compensates around the training with denser rewards, SFT priors, skill libraries, curated memory, or multi-agent orchestration. We argue the ceiling is an artifact of two failures of common practice. Signal starvation: group-relative RL with sparse outcome-only rewards yields a gradient only when a task's rollout group mixes successes and failures, so under-scaled exploration silences exactly the hardest, most instructive tasks. Policy drift: squeezing many updates out of a small task pool degrades the policy itself, as an unanchored objective lets the sampling distribution collapse exactly when saturation has already made informative groups rare. We present CANOPY (Coverage-ANchored On-PolicY RL), a minimalist protocol attacking both directly: scale same-task exploration until the natural signal reappears, keep every update on-policy, KL-anchored, and confined to the agent's own action tokens, then cash in an enlarged interaction budget at test time. On AppWorld, a long-horizon interactive coding benchmark, a Qwen3-14B policy trained with CANOPY through environment interaction alone--without task-specific supervision, auxiliary credit signals, or elaborate agent scaffolding--topped the public leaderboard (Feb. 2026; Test-Normal TGC 86.9, Test-Challenge 67.6), and the same design principles lift Qwen3.5-9B on SWE-bench Verified by 16.6 points. Agentic RL alone thus internalizes long-horizon capability directly into a small open model; we plan to release the complete training stack at https://github.com/AlibabaResearch/SignalCoverageRL.
Chinese Translation
强化学习是对大语言模型(LLM)智能体进行后训练以完成长程交互任务的自然方法,此类任务仅通过任务结束时的验证来评判,然而一种普遍的看法认为,仅基于结果的强化学习(outcome-only RL)很快会在小型开源模型上触及天花板。因此,近期工作通过更密集的奖励、SFT先验、技能库、精选记忆或多智能体编排等训练之外的补偿手段来应对这一限制。我们认为,这一“天花板”其实是常见实践中两类失败的产物。其一是信号匮乏(Signal starvation):采用稀疏的仅结果奖励的组相对强化学习,只有当某个任务的 rollout 组中同时混合了成功与失败时才会产生梯度,因此探索规模不足恰恰会使最难、最具信息量的任务失去信号。其二是策略漂移(Policy drift):从小规模任务池中反复压缩出大量更新会损害策略本身,因为无锚定的目标会让采样分布在饱和已使有效信息组变得稀缺之际发生坍缩。我们提出 CANOPY(Coverage-ANchored On-PolicY RL,覆盖锚定的同策略强化学习),一种直接攻克上述两个问题的极简协议:将同任务探索扩展到自然信号重新出现为止;保持每次更新均为同策略(on-policy)、受 KL 锚定且仅限于智能体自身的动作 token;然后在测试时兑现更大的交互预算。在长程交互式编程基准 AppWorld 上,仅通过环境交互、无需任务特定监督、辅助信用信号或复杂智能体脚手架、用 CANOPY 训练的 Qwen3-14B 策略登顶公开排行榜(2026 年 2 月;Test-Normal TGC 86.9,Test-Challenge 67.6),且相同的设计原则使 Qwen3.5-9B 在 SWE-bench Verified 上提升 16.6 分。由此可见,纯粹的智能体强化学习(Agentic RL)即可将长程能力直接内化到小型开源模型中;我们计划在 https://github.com/AlibabaResearch/SignalCoverageRL 发布完整的训练栈。
cs.LG / 86 / 2609.01262

Solving In-Table Prediction Problems by Deep Neural Networks with Performance Evaluation Using Synthetic Data

基于深度神经网络求解表内预测问题并使用合成数据进行性能评估
Zhao, Xiao, Oelke, Daniela
Abstract
Tabular deep learning (TDL) leverages neural networks (NN) to extract patterns from tabular data. Traditional TDL methods follow a supervised learning paradigm, where a target feature is explicitly given. In this work, however, we explore a different approach by employing deep NNs to learn relationships among individual columns within a given table. We investigate whether NNs can predict the values of arbitrarily selected columns in a given table based on the remaining known columns. We call this problem In-Table Prediction (ITB), which is slightly different from table imputation methods and the pretraining task of TDL. Three potential usage scenarios are identified, which, to our best knowledge, have not been extensively studied in the literature. A self-supervised learning approach is applied to address this problem by randomly selecting columns to be masked out and used as learning targets. This work focuses on tabular datasets containing only continuous features. To handle missing values in continuous features, a novel neural layer is proposed to embed both numerical and empty values. Synthetic data is generated based on predefined column relationships, with empty values inserted using two distinct mechanisms. Additionally, an adapted masking strategy is employed to create test data. Performances of three NN architectures, namely MLP, Resnet and Transformer, are evaluated using the generated synthetic data. We conclude that, the attention-based structure outperforms the other two networks, when a sufficiently large number of training examples is available and a relatively large embedding length is chosen. We stress that these findings are obtained under controlled, synthetic conditions with a small number of columns and it should therefore be regarded as an initial, narrowly-scoped investigation rather than a general characterization of ITP on real-world tabular data.
Chinese Translation
表格深度学习(TDL)利用神经网络(NN)从表格数据中提取模式。传统的TDL方法遵循监督学习范式,其中目标特征是明确给定的。然而,本工作探索了一种不同的方法,即采用深度神经网络来学习给定表格中各个列之间的关系。我们研究神经网络能否基于给定表格中其余已知列来预测任意选定列的值。我们将该问题称为表内预测(In-Table Prediction,ITP),它与表格填补方法以及TDL的预训练任务略有不同。我们识别了三种潜在的使用场景,据我们所知,这些场景在文献中尚未得到广泛研究。我们采用自监督学习方法来解决该问题,即随机选择列进行掩码并将其作为学习目标。本工作仅关注只含连续特征的表格数据集。为处理连续特征中的缺失值,我们提出了一种新颖的神经层,用于同时嵌入数值和空值。基于预定义的列关系生成合成数据,并通过两种不同机制插入空值。此外,还采用了一种改进的掩码策略来构建测试数据。我们使用生成的合成数据评估了三种神经网络架构(即MLP、Resnet和Transformer)的性能。我们得出结论:当有足够多的训练样本且选择相对较大的嵌入长度时,基于注意力的结构优于其他两种网络。我们强调,这些发现是在受控的合成条件下、列数较少的情况下获得的,因此应被视为一项初步的、范围有限的研究,而非对ITP在真实世界表格数据上表现的一般性刻画。
cs.LG / 87 / 2609.01273

Position: Privacy Is a Claim, Not a Property of Synthetic Data

立场声明:隐私是一种主张,而非合成数据的固有属性
Zhao, Jiachen, Januszewicz, Antonia, Jung, Taeho
Abstract
Synthetic data has become a common component of machine learning research. While widely adopted, its use in privacy-sensitive contexts has quietly shifted from a claim of residual inference risk under stated assumptions to an appearance-based property inferred from data generation itself. In this position paper, we argue that this shift reflects an implicit change in community standards for what counts as sufficient privacy evidence, rather than a misunderstanding of well-established privacy principles. Drawing on an empirical analysis of recent publications across major ML venues, we show that synthetic data is frequently used in privacy-sensitive settings without explicit articulation of threat models, inference risks, or falsifiable privacy claims. As a result, privacy assurance often remains implicit, difficult to verify, and unevenly distributed, with heightened exposure for rare and minority records. We argue for treating privacy as an explicit, evidence-based scientific claim and recommend that ML venues adopt norms requiring privacy-relevant assertions to be clearly scoped, testable, and contestable.
Chinese Translation
合成数据已成为机器学习研究中的一种常见组成部分。尽管被广泛采用,但其在隐私敏感场景中的使用,已悄然从“在既定假设下关于残余推断风险的主张”转变为“由数据生成过程本身推断出的基于外观的属性”。在本立场论文中,我们论证这一转变反映了社区对于何为充分的隐私证据这一标准发生了隐性的变化,而非对已确立的隐私原则的误解。通过对近年来主要机器学习会议发表论文的实证分析,我们表明合成数据经常在隐私敏感的场景中被使用,却未明确阐述威胁模型、推断风险或可证伪的隐私主张。因此,隐私保障往往仍然是隐性的、难以验证且分布不均的,稀有记录和少数群体记录面临更高的暴露风险。我们主张将隐私视为一种明确的、基于证据的科学主张,并建议机器学习会议采纳相关规范,要求与隐私相关的论断必须具有清晰的范围界定、可检验性并可被质疑。
cs.LG / 88 / 2609.01275

The Constitutional Coverage Trilemma in AI Governance

人工智能治理中的宪法覆盖三难困境
Mitic, Natalija, O., Soona Sedahmed A., Ly, Mamadou Selly, Cisse, Moustapha
Abstract
Frontier AI systems function as \emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human demand. Combining a paraphrase-controlled audit of the as-shipped default constitutions of $23$ frontier LLM archetypes with a pairwise-tradeoff study of $1{,}649$ US participants on the same instrument, we report three facts. \emph{Demand is broad}: it spans all five values, with the largest constituency under one-third. \emph{Supply is narrow and drifting}: the $23$-archetype hull occupies ${\sim}2\%$ of the demand hull under conservative noise-matched estimation ($0.10\%$ at full audit precision), no archetype puts helpfulness or autonomy first ($37\%$ of users are constitutionally homeless), and across six model families autonomy decreases in $5/6$, equity increases in $5/6$, and safety increases in $4/6$, with monotone within-family version trends (order-permutation $p = 0.013$) and the autonomy decline concentrated in scenarios where safety is not at stake. The drift's importance is directional: \emph{away} from a value already undercovered, mechanically worsening the welfare floor for the least-served users. \emph{The fix is sparse}: a $2$-vertex menu $\{e_{\mathrm{HON}}, e_{\mathrm{AUT}}\}$ beats the full $23$-archetype frontier by $47\%$ on mean regret (CI $[43\%, 52\%]$); three vertex additions cut mean/worst-group regret by up to $81\%$/$64\%$. We formalize these findings as a budgeted-pluralism trilemma, show the binding regime is empirically realized, and verify the conclusions are robust to distance-based welfare and to degraded routing. The instrument and audit harness are described in full in the appendices.
Chinese Translation
前沿人工智能系统发挥着“宪法机构”的作用:每个部署的模型都对安全性、有用性、诚实性、自主性和公平性之间编码了一种隐含的排序。我们探讨前沿宪法类型的供给是否能够覆盖人类的需求。通过对23个前沿大语言模型(LLM)原型出厂默认宪法进行改写受控审计,并在同一测量工具上对1,649名美国参与者进行成对权衡研究,我们报告了三个事实。需求是广泛的:它涵盖全部五种价值,其中最大群体的占比不足三分之一。供给是狭窄且漂移的:在保守的噪声匹配估计下,23个原型所构成的包络仅占需求包络的约2%(在完全审计精度下为0.10%);没有任何原型将有用性或自主性置于首位(37%的用户在宪法意义上处于“无家可归”状态);在六个模型家族中,自主性在5/6中下降,公平性在5/6中上升,安全性在4/6中上升,且家族内部呈现单调的版本演进趋势(顺序置换检验 p = 0.013),而自主性的下降集中在安全性并未受到威胁的场景中。这种漂移的重要性在于其方向性:它背离了本已覆盖不足的价值,机械地恶化了服务最少用户的福利下限。解决方案是稀疏的:仅包含两个顶点{e_HON, e_AUT}的菜单在平均遗憾上比完整的23原型前沿优越47%(置信区间[43%, 52%]);增加三个顶点可将平均/最差群体遗憾最多降低81%/64%。我们将这些发现形式化为预算多元主义三难困境,证明约束状态在经验上确实存在,并验证了结论对基于距离的福利度量以及路由退化具有稳健性。测量工具和审计框架在附录中完整描述。
cs.LG / 89 / 2609.01311

One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context

单层Transformer可证明地在上下文中学习多分类最近邻算法
Athreya, Skanda, Wang, Yutong
Abstract
We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifiers in the binary setting to the multiclass case. By leveraging the simplex encoding, we show that one-layer transformers with an argmax classification head behave identically to a one-nearest-neighbor classifier in the multiclass setting. This closes a gap left by prior work, whose multiclass result relied on a non-standard rounding-based approach rather than the typical argmax head used in practice.
Chinese Translation
我们将近期关于单层Transformer与二分类设定下最近邻分类器等价性的研究工作扩展至多分类情形。通过利用单纯形编码(simplex encoding),我们证明带有argmax分类头的单层Transformer在多分类设定下的行为与一最近邻分类器完全一致。这弥补了先前工作留下的空白——之前的多分类结果依赖于一种非标准的基于舍入的方法,而非实践中通常使用的argmax分类头。
cs.LG / 90 / 2609.01335

Bandits in Prod: Hyperparameter Optimization at Inference Time

实战中的老虎机:推理时的超参数优化
Abraham, Louis, Nguyen, Tuan-Anh, Devatine, Nicolas
Abstract
Many production systems can assess a configuration only by using it on live requests and observing noisy feedback. Modern agentic systems are a prominent example, with inference-time choices such as model selection, retrieval depth, prompting strategy, and decoding temperature, yet often with no representative validation data. We formalize this setting as Online Hyperparameter Optimization (OHPO) and cast it as an infinitely many-armed bandit over mixed and conditional search spaces. We introduce IMABO, a general framework that combines any bandit policy for choosing among already sampled configurations with any oracle for proposing new ones. We instantiate it with IMOSS, a restart-free anytime policy whose active set grows as $t^{\beta}$, and prove an expected cumulative quantile-regret bound of $O(p_\rho^{-1/\beta} + T^{(1+\beta)/2})$, where $\beta\in(0,1)$ controls active-set growth and $p_\rho$ lower-bounds the probability that a proposed configuration falls in the top-$\rho$ fraction of the search space. We combine IMOSS with three practical oracles: a Tree-structured Parzen Estimator, an incumbent-mutation oracle driven by a per-coordinate bandit, and a pretrained tabular foundation model, all three improving over the uniform random oracle baseline. IMABO outperforms all baselines in terms of regret across diverse OHPO settings, from tuning classical machine-learning models to configuring LLM-based agents. Our implementation is available at https://github.com/Tiime-Software/IMABO.
Chinese Translation
许多生产系统只能通过将某种配置应用于实际线上请求并观察带噪声的反馈来评估该配置。现代智能体(agentic)系统是一个典型例子,其推理时的选择包括模型选择、检索深度、提示策略和解码温度等,但往往缺乏有代表性的验证数据。我们将这一场景形式化为在线超参数优化(Online Hyperparameter Optimization, OHPO),并将其建模为混合与条件搜索空间上的无穷多臂老虎机问题。我们提出了 IMABO,一个通用框架,它将任意用于在已采样配置中进行选择的老虎机策略与任意用于提出新配置的预言机(oracle)相结合。我们用 IMOSS 实例化该框架,IMOSS 是一种无需重启的任意时刻(anytime)策略,其活跃集大小随 $t^{\beta}$ 增长,并证明了期望累积分位数遗憾(quantile-regret)上界为 $O(p_\rho^{-1/\beta} + T^{(1+\beta)/2})$,其中 $\beta\in(0,1)$ 控制活跃集的增长速度,$p_\rho$ 为所提配置落入搜索空间前 $\rho$ 分位比例的概率下界。我们将 IMOSS 与三种实用的预言机相结合:树结构 Parzen 估计器(Tree-structured Parzen Estimator)、由逐坐标老虎机驱动的当前最优解变异预言机,以及预训练的表格基础模型,三者均优于均匀随机预言机基线。在从经典机器学习模型调优到配置基于 LLM 的智能体等多种 OHPO 场景中,IMABO 在遗憾指标上均优于所有基线方法。我们的实现代码可在 https://github.com/Tiime-Software/IMABO 获取。
cs.LG / 91 / 2609.01343

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

SMELT:计算量匹配的专家混合循环Transformer的缩放定律
Wang, Shaowen, Zhang, Ge, Luo, Kairong, Wu, Yuhao, Liu, Shaofan, Liu, Jiaheng, Huang, Wenhao, Yan, Shen, Li, Jian
Abstract
Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT's loss drops faster with compute, saving 6.8--18.0\% of training FLOPs on the compute-optimal frontier. The advantage transfers to downstream benchmarks beyond what validation loss predicts, is largest on Code, and grows with sample length and the number of in-context examples. Mechanistic analysis shows that the second visit reduces the attention sink and redirects mass toward content-relevant tokens, an inductive bias that may underlie the observed performance gains. These results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.
Chinese Translation
循环Transformer(Looped Transformers)通过迭代共享的层块来增加有效深度,但大多数评估是在固定模型规模下进行比较,将架构优势与额外FLOPs混为一谈。我们在专家混合Transformer上研究循环机制,同时严格匹配每token FLOPs、总非嵌入参数量以及KV缓存。通过一系列消融实验,我们得出了一种称为SMELT的方案(稀疏MoE Transformer,中间层循环两次),即将中间一半的层循环两次,同时在上述三项预算上与非循环基线相匹配。我们将SMELT扩展至四种规模、最大达540亿非嵌入参数,并为每种架构分别拟合了Chinchilla式的缩放定律。SMELT的损失随计算量下降得更快,在计算最优前沿上可节省6.8%–18.0%的训练FLOPs。这一优势在下游基准上的提升超出了验证损失所预测的水平,在代码任务上最为显著,并随样本长度和上下文示例数量的增加而增大。机制分析表明,第二次迭代减少了注意力汇聚现象,并将注意力质量重新引导至与内容相关的token上,这一归纳偏置可能是所观察到的性能提升的基础。这些结果表明,即使在预算严格匹配的条件下,循环机制仍能改进Transformer,并提供了一个将深度复用转化为可衡量收益的实用方案。
cs.LG / 92 / 2609.01406

Contribution-Aware Bandwidth Allocation for Multimodal Split Learning

面向多模态拆分学习的贡献感知带宽分配
Ofeidis, Iason, Tassiulas, Leandros
Abstract
Multimodal models are increasingly the default option for perception at the network edge, yet they are trained almost entirely in the datacenter, because a client holding several sensor streams cannot host an encoder per modality. Split Learning makes such training feasible by keeping only the first layers on the device, at the cost of an uplink that must carry smashed activations for every modality at every step. Existing compression schemes give each modality the same keep-ratio, so the shared budget is divided in proportion to smashed-activation dimension, a quantity unrelated to how much each modality contributes to the fused prediction. We make that division an explicit decision and call it inter-modality allocation: under a fixed uplink budget, every policy transmits the same expected payload and differs only in how that payload is split across modalities. Our allocator, ModalShare, sets each modality's keep-ratio from a Shapley contribution score that the server computes over coalitions of activations it has already received. Measuring this score adds no uplink traffic and no client-side computation, and needs no prior knowledge of which stream is which. ModalShare improves accuracy over equal keep-ratios by 15.4 and 12.4 percentage points on CREMA-D and MVSA at matched payload in 5x compression, with strong performance across three compressors, three datasets, and four budgets. We show that existing compressors underperform in multimodal settings, with ModalShare recovering what gains are left behind.
Chinese Translation
多模态模型正日益成为网络边缘感知任务的默认选择,然而其训练几乎完全在数据中心进行,因为持有多个传感器数据流的客户端无法为每种模态托管一个编码器。拆分学习通过仅在设备上保留前几层使此类训练成为可能,但其代价是上行链路必须在每一步为每种模态传输压缩后的激活值。现有的压缩方案为每种模态分配相同的保留比例,因此共享带宽预算按照压缩激活值的维度成比例分配——而这一维度与各模态对融合预测的贡献大小无关。我们将这一分配过程变为一个显式决策,并将其称为模态间分配:在固定的上行链路预算下,任何策略传输的期望负载都相同,区别仅在于负载在各模态之间的划分方式。我们的分配器 ModalShare 根据服务器在其已接收的激活值组合上计算得到的 Shapley 贡献分数,为每种模态设置保留比例。测量该分数不会增加上行流量和客户端计算开销,也不需要预先知道各数据流的身份。在相同负载且压缩率为 5 倍的条件下,ModalShare 在 CREMA-D 和 MVSA 数据集上的准确率相比等保留比例方案分别提升了 15.4 和 12.4 个百分点,并在三种压缩器、三个数据集和四种预算设置下均表现出色。我们证明现有压缩器在多模态场景下性能欠佳,而 ModalShare 能够挽回这些被遗留的性能收益。
cs.LG / 93 / 2609.01417

Predicting Subsurface Abnormalities Growth using Physics-Informed Neural Networks

基于物理信息神经网络的地下异常体生长预测
Dizaji, Mehrdad Shafiei, Azari, Hoda
Abstract
The research explores the pioneering integration of Physics-Informed Neural Networks (PINNs) into the domain of Ground-Penetrating Radar (GPR) data prediction. This research presents a detailed development framework for a specialized PINN model, proficient at interpreting and forecasting GPR data, much like how medical imaging models predict tumor behavior. By harnessing the synergy between deep learning algorithms and the physical laws governing subsurface structures or in medical terms, human tissues the model effectively embeds the physics of electromagnetic wave propagation into its architecture. This ensures that predictions not only align with fundamental physical principles but also mirror the precision needed in medical diagnostics for detecting and monitoring tumors. The suggested deep learning structure comprises three components: a CNN, a spatial feature channel attention (SFCA) mechanism, and ConvLSTM, along with temporal feature frame attention (TFFA) modules. The attention mechanism computes channel attention and temporal attention weights using self-adaptation, thereby fine tuning the visual and temporal feature responses to extract the most pertinent and significant visual and temporal features. By integrating physics directly into the neural network, our model has shown enhanced accuracy in forecasting GPR data. This improvement is vital for conducting effective assessments of bridge deck conditions and other evaluations related to civil infrastructure. The use of Physics Informed Neural Networks (PINNs) has demonstrated the potential to transform the field of Non-Destructive Evaluation (NDE) by enhancing the precision of infrastructure deterioration predictions. Moreover, it offers a deeper insight into the fundamental mechanisms of deterioration, viewed through the prism of physics-based models.
Chinese Translation
本研究探索了将物理信息神经网络(PINNs)开创性地融入探地雷达(GPR)数据预测领域。本文提出了一个详细的专用PINN模型开发框架,该模型能够解读并预测GPR数据,其原理类似于医学影像模型预测肿瘤行为。通过利用深度学习算法与支配地下结构(在医学术语中即人体组织)的物理定律之间的协同作用,该模型有效地将电磁波传播的物理特性嵌入到其架构中。这确保了预测结果不仅符合基本物理原理,还能达到医学诊断中检测和监测肿瘤所需的精度。所提出的深度学习结构由三个组件构成:卷积神经网络(CNN)、空间特征通道注意力(SFCA)机制,以及ConvLSTM与时间特征帧注意力(TFFA)模块。该注意力机制通过自适应方式计算通道注意力权重和时间注意力权重,从而对视觉特征响应和时间特征响应进行微调,以提取最相关、最重要的视觉与时间特征。通过将物理规律直接融入神经网络,我们的模型在预测GPR数据方面表现出更高的精度。这一改进对于开展桥梁桥面状况及其他土木基础设施相关评估的有效检测至关重要。物理信息神经网络(PINNs)的应用展现出变革无损检测(NDE)领域的潜力,可提高基础设施劣化预测的精度。此外,它还提供了一个基于物理模型视角,对劣化基本机制的更深入理解。
cs.LG / 94 / 2609.01418

Provably Safe Sim-to-Real Transfer

可证明安全的仿真到现实迁移
Ni, Tingting, Kamgarpour, Maryam
Abstract
To mitigate the sample complexity of real-world reinforcement learning (RL), a common practice is to first train a policy in a simulator, where samples are cheap, and then deploy the learned policy in the real world with the hope that it generalizes effectively. Such direct sim-to-real transfer is not guaranteed to succeed: simulator-trained policies can be suboptimal in the real world due to sim-to-real mismatch. Correcting this mismatch requires collecting data from the real system, but in many applications, such as robotics and healthcare, this data-collection process is itself subject to safety constraints. This gives rise to the problem of safe sim-to-real transfer: how can an agent exploit an imperfect simulator while ensuring safe real-world data collection and learning a near-optimal feasible policy for the target system? We address this problem by formulating safe sim-to-real transfer within the framework of reward-free safe RL. We design a computationally efficient algorithm that exploits simulator information to provably reduce real-world interaction while ensuring safe exploration and enabling the computation of a near-optimal feasible policy for any potential reward function. Our real-world sample complexity bound characterizes the benefit of using the simulator in terms of the sim-to-real mismatch.
Chinese Translation
为了缓解现实世界强化学习(RL)的样本复杂度问题,一种常见的做法是首先在样本成本较低的模拟器中训练策略,然后将学到的策略部署到现实世界中,期望其能够有效泛化。然而,这种直接的仿真到现实(sim-to-real)迁移并不保证成功:由于仿真与现实之间的不匹配(sim-to-real mismatch),在模拟器中训练的策略在现实世界中可能是次优的。纠正这种不匹配需要从真实系统收集数据,但在许多应用(如机器人和医疗保健)中,这一数据收集过程本身受到安全约束的限制。这就引出了安全仿真到现实迁移的问题:智能体如何在利用不完美模拟器的同时,确保现实世界数据收集的安全性,并为目标系统学习到接近最优的可行策略?我们在无奖励安全强化学习(reward-free safe RL)的框架内对这一问题进行建模以解决该问题。我们设计了一种计算高效的算法,利用模拟器信息来可证明地减少现实世界的交互,同时确保安全探索,并能够为任意潜在奖励函数计算接近最优的可行策略。我们的现实世界样本复杂度界从仿真与现实不匹配的角度刻画了使用模拟器所带来的收益。
cs.LG / 95 / 2609.01425

CATeye: Coupled Attribute-Topology Invariance Learning for Voucher Abuse Detection

CATeye:面向优惠券滥用检测的耦合属性-拓扑不变性学习
Tian, Tian, Niu, Shuaicheng, Kuang, Hao, Hu, Yuanhang, Li, Dong, Shen, Zhiqi
Abstract
Voucher abuse poses a major challenge in e-commerce, where malicious users exploit promotional vouchers for profit. Unfortunately, fraud patterns evolve rapidly over time and across regions, causing distribution shifts that degrade existing detection models unless retrained frequently. To tackle this, we propose the Coupled Attribute-Topology Invariance Learning framework (CATeye). The key challenge arises from coupled attribute-topology shift, where edges built from attribute proximity cause environment-driven attribute shift to induce shifted topology, thereby amplifying variant signals through GNN message passing. CATeye sees through such coupled shifts with two learnable selectors. First, an Attribute Invariance Selector (AIS) learns node-adaptive masks to filter out non-invariant attributes. Then, conditioned on retained invariant attributes, an Edge Invariance Selector (EIS) samples an invariant subgraph and isolates non-invariant edges. Using the resulting invariant and non-invariant components, CATeye constructs multiple views and applies view-specific objectives to emphasize domain-invariant representations while suppressing domain-specific variations. Experiments on both a proprietary dataset from Lazada, a major Southeast Asian e-commerce platform, and a public benchmark show that CATeye consistently outperforms nine strong domain generalization and graph anomaly detection baselines, achieving up to an 8.61% improvement in average F1 score over the strongest baseline. Source code is publicly available at https://github.com/Tian0426/CATeye.
Chinese Translation
优惠券滥用是电子商务领域面临的重大挑战,恶意用户通过利用促销优惠券牟取利益。然而,欺诈模式会随时间和地域快速演变,导致分布偏移,使现有检测模型若不频繁重新训练则性能下降。为解决这一问题,我们提出了耦合属性-拓扑不变性学习框架(Coupled Attribute-Topology Invariance Learning,CATeye)。其关键挑战源于耦合的属性-拓扑偏移:基于属性接近度构建的边使得环境驱动的属性偏移会引发拓扑偏移,并通过GNN消息传递放大变体信号。CATeye通过两个可学习的选择器洞察此类耦合偏移。首先,属性不变性选择器(Attribute Invariance Selector,AIS)学习节点自适应掩码,以过滤出非不变属性。随后,在保留的不变属性的条件下,边不变性选择器(Edge Invariance Selector,EIS)采样出不变子图并隔离非不变边。利用由此得到的不变与非不变组件,CATeye构建多个视图并施加视图特定的目标函数,在强调领域不变表示的同时抑制领域特定变化。在来自东南亚主要电商平台Lazada的专有数据集和一个公开基准上的实验表明,CATeye持续优于九个强大的领域泛化和图异常检测基线方法,平均F1分数较最强基线最高提升8.61%。源代码已公开于 https://github.com/Tian0426/CATeye。
cs.LG / 96 / 2609.01428

TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution

TRIAGE:面向高效执行的三级路由与智能体引导机制
Wei, Ruocan
Abstract
Large Language Model (LLM) agents based on the ReAct paradigm have demonstrated remarkable capabilities in tool use and task execution. However, ReAct suffers from a fundamental efficiency problem: every query triggers a complete reasoning loop from scratch, and similar queries repeat identical steps without leveraging historical experience. We propose TRIAGE,a three-level routing framework that reduces token consumption by reusing historical execution trajectories. Its core innovation is TaaS (Trajectory-as-a-Skill), which abstracts historical execution trajectories into reusable skills, realizing 'experience as a service'. TRIAGE classifies queries into three levels: (1) Direct Reuse-identical queries, 0 tokens; (2) Skill Substitution-similar queries, 0 tokens via deterministic parameter substitution; (3) Full ReAct-novel queries, automatically stored for future reuse. In large-scale experiments on 1,007 security monitoring queries, TRIAGE achieves 62.3% token savings, with 56.0% of queries at Level 2 and 5.5% at Level 1, both executing at zero cost. Cross-domain validation on ToolBench (15 domains, 345 queries) achieves 76.3% token reduction, confirming the generalizability of semantic routing. An online learning experiment demonstrates cold-start-to-mature evolution: the L2 hit rate rises from 0% to 57% within the first 100 queries, and the average token cost drops from 198 to 74.7. We also propose an automatic Skill extraction mechanism that distills high-frequency trajectory patterns into deterministic Skills, creating a positive feedback loop of 'the more you use it, the more efficient it becomes'.
Chinese Translation
基于ReAct范式的大语言模型(LLM)智能体在工具使用和任务执行方面展现了卓越的能力。然而,ReAct存在一个根本性的效率问题:每个查询都会从头触发完整的推理循环,相似查询重复相同的步骤,而未能利用历史经验。我们提出TRIAGE,一个通过复用历史执行轨迹来降低token消耗的三级路由框架。其核心创新是TaaS(Trajectory-as-a-Skill,轨迹即技能),它将历史执行轨迹抽象为可复用的技能,实现“经验即服务”。TRIAGE将查询分为三个层级:(1)直接复用——完全相同的查询,0 token;(2)技能替换——相似查询,通过确定性参数替换实现0 token执行;(3)完整ReAct——新查询,执行后自动存储以供未来复用。在1,007条安全监控查询的大规模实验中,TRIAGE实现了62.3%的token节省,其中56.0%的查询处于第二层级、5.5%处于第一层级,均以零成本执行。在ToolBench(15个领域,345条查询)上的跨领域验证实现了76.3%的token削减,证实了语义路由的通用性。在线学习实验展示了从冷启动到成熟期的演进过程:L2命中率在前100条查询内从0%上升至57%,平均token成本从198降至74.7。我们还提出了一种自动技能提取机制,将高频轨迹模式提炼为确定性技能,形成了“越用越高效”的正反馈循环。
cs.LG / 97 / 2609.01430

Learning Sparse Decision Trees via Transformer Variational Auto-Encoders

基于Transformer变分自编码器的稀疏决策树学习
Fidone, Giacomo, Cascione, Alessio, Guidotti, Riccardo
Abstract
Decision trees are among the most widely used models in machine learning, largely due to their transparent decision logic, making them well-suited for high-stakes decision-making contexts. However, most existing learning algorithms focus on predictive performance, overlooking the joint optimization of other desirable properties, such as structural sparsity. In this work we propose TREVIS, an approach for learning decision trees with respect to complex objectives, based on the exploration of the latent space of a Tree Transformer Variational Auto-Encoder (TTVAE). By mapping decision trees onto latent representations, TREVIS replaces the discrete search space with a continuous one, enabling gradient-based optimization via a differentiable surrogate model. We experiment with TREVIS for learning decision trees that jointly optimize predictive performance and sparsity. Results show that TREVIS discovers decision trees matching the predictive performance of existing near-optimal algorithms while improving their structural sparsity.
Chinese Translation
决策树是机器学习中应用最广泛的模型之一,这主要归功于其透明的决策逻辑,使其非常适用于高风险决策场景。然而,大多数现有的学习算法主要关注预测性能,忽视了对其他理想特性(如结构稀疏性)的联合优化。在本工作中,我们提出了TREVIS,一种基于探索树Transformer变分自编码器(Tree Transformer Variational Auto-Encoder, TTVAE)潜在空间的决策树学习方法,能够针对复杂目标进行决策树学习。通过将决策树映射到潜在表示,TREVIS将离散搜索空间替换为连续搜索空间,从而能够借助可微的代理模型进行基于梯度的优化。我们通过实验将TREVIS用于学习同时优化预测性能和稀疏性的决策树。结果表明,TREVIS所发现的决策树在预测性能上与现有的近似最优算法相当,同时在结构稀疏性上有所提升。
cs.LG / 98 / 2609.01431

Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

基于幂律熵搜索的高效最优超参数缩放律估计
Chen, Zhiliang, Ament, Sebastian, Eriksson, David, Balandat, Maximilian, Low, Bryan Kian Hsiang, Bakshy, Eytan, Lin, Jihao Andreas
Abstract
Optimal hyperparameter scaling laws describe how the best hyperparameters for large language model (LLM) training change with model and data scale, enabling practitioners to predict optimal configurations at production scales without expensive large-scale tuning. However, estimating these scaling laws conventionally requires exhaustive grid searches over thousands of training runs, consuming enormous computational resources. We introduce Power-Law Entropy Search (PLES), a computational cost-aware acquisition function built on multi-fidelity Bayesian optimization that efficiently estimates optimal hyperparameter scaling laws through adaptive experimentation. A key innovation in PLES is that it searches for candidates that reduce the overall uncertainty of a scaling law estimate, instead of optimizing a single objective function. At each iteration, PLES selects the candidate configuration that maximally reduces the uncertainty of the scaling law estimates per unit computational cost, naturally favoring informative small-scale experiments. We evaluate PLES on synthetic benchmarks, surrogate models fitted to real LLM training data, and actual LLM pre-training runs. Across all settings, PLES converges to accurate optimal hyperparameter scaling laws using less than one-tenth of the computational budget required by conventional grid search and other baselines.
Chinese Translation
最优超参数缩放律描述了大型语言模型(LLM)训练的最佳超参数如何随模型和数据规模变化,使从业者能够在生产规模下预测最优配置,而无需进行昂贵的大规模调优。然而,传统上估计这些缩放律需要对数千次训练运行进行穷举网格搜索,消耗巨大的计算资源。我们提出了幂律熵搜索(Power-Law Entropy Search,PLES),这是一种基于多保真贝叶斯优化的计算成本感知采集函数,能够通过自适应实验高效地估计最优超参数缩放律。PLES 的一个关键创新在于,它搜索的是能够降低缩放律估计整体不确定性的候选点,而不是优化单一目标函数。在每次迭代中,PLES 选择在单位计算成本下能最大程度降低缩放律估计不确定性的候选配置,从而自然地偏向于信息量大的小规模实验。我们在合成基准、基于真实 LLM 训练数据拟合的代理模型以及实际的 LLM 预训练运行上对 PLES 进行了评估。在所有设置中,PLES 均能收敛到准确的最优超参数缩放律,且所需的计算预算不到传统网格搜索及其他基线方法的十分之一。
cs.LG / 99 / 2609.01441

Edge-Girth as a Structural Edge Feature for Graph Neural Networks

边-围长作为图神经网络的结构化边特征
Marey, Lilian, Laclau, Charlotte
Abstract
Graph neural networks (GNN) based on message passing are provably no more powerful than the one-dimensional Weisfeiler--Leman colour-refinement test (1-WL): two graphs it cannot tell apart receive identical representations, however deep or wide the network. A common remedy augments node or edge features with precomputed structural descriptors, most often counts of a fixed small subgraph such as triangles or longer cycles, but such counts require committing in advance to the size of the substructure counted, a choice usually made blind to the data. We study a descriptor that avoids this choice. The edge-girth of an edge is the length of a shortest cycle through it, and its multiplicity is the number of such shortest cycles; together they form a per-edge invariant that reports cycles of arbitrary length, computable exactly by a single breadth-first search per edge. Injected into a gated message-passing architecture, EGAGNN, it reaches a test MAE a factor three below the closest gated comparator on the ZINC-12k regression benchmark at 104k parameters; against bounded cycle-counting descriptors under the same architecture, it matches only a dictionary counting cycles up to length eight, using twice as many channels, while a dictionary capped at length four performs no better than no structural information at all. On graph discrimination we prove a matching limitation: on graphs where every edge sees the same number of shortest cycles of the same length, the descriptor becomes constant and any model built on it collapses back to the 1-WL bound. This holds without exception across all 400 pairs of the BREC benchmark: not one of the 90 such pairs is distinguished.
Chinese Translation
基于消息传递的图神经网络(GNN)在理论上其能力不超过一维Weisfeiler–Leman颜色细化测试(1-WL):对于该测试无法区分的两个图,无论网络多深或多宽,它们都会得到完全相同的表示。一种常见的补救措施是使用预先计算的结构描述符来增强节点或边特征,最常见的是统计某个固定小子图(如三角形或更长的环)的数量,但这类计数需要预先确定所统计子结构的大小,而这一选择通常是在缺乏数据依据的情况下盲目做出的。我们研究了一种可以避免这一选择的描述符。边的边-围长(edge-girth)是指经过该边的最短环的长度,其重数是指这类最短环的数目;二者共同构成一个逐边不变量,能够报告任意长度的环,并且可以通过对每条边执行一次广度优先搜索来精确计算。将该描述符注入门控消息传递架构(即EGAGNN)后,在ZINC-12k回归基准上,仅使用104k参数,其测试MAE比最接近的门控对比方法低了三倍;在相同架构下与有界环计数描述符相比,它仅相当于使用两倍通道数、统计长度不超过八的环的字典方法,而以长度四为上限的字典方法的表现与不使用任何结构信息无异。在图区分能力方面,我们证明了一个相应的局限性:在每条边所见的同长度最短环数目都相同的图上,该描述符退化为常量,任何基于它的模型都会退回到1-WL的能力界限。这一结论在BREC基准的全部400对图上无一例外地成立:90对这样的图没有一对能被区分开。
cs.LG / 100 / 2609.01449

Diffusion as a Training Curriculum for Timestep-Free Iterative Reasoning

扩散作为无时间步迭代推理的训练课程
Drozdova, Mariia, Sirbu, Aidan, Miotti, Pietro, Obryk, Robert, Etcheverry, Mayalen, Niklasson, Eyvind, Richards, Blake
Abstract
Diffusion models and recursive reasoners are both iterative, but they carry information across iterations differently. We add a persistent hidden state to a diffusion denoiser and remove its timestep conditioning, leaving a single shared update that can be run to arbitrary depth. The result is an anytime solver: accuracy keeps improving with inference depth far beyond the rollout lengths and backpropagation window used in training, reaching 99.90% exact solve on Sudoku-Extreme. We also obtain 98.93% solve rate on Maze-Unique. Surprisingly, progressive denoising is unnecessary at inference: holding corruption at its maximum by replacing every non-clue variable with fresh Gaussian noise at each step retains near-perfect solving and converges to stable solutions. This simple noise-injection mechanism enables a single trajectory to efficiently explore the solution space and settle on the correct answer without parallel rollouts, candidate selection, or external verifiers required by prior reasoning models. Nonetheless, ordered annealed corruption remains critical during training, which suggests that diffusion's primary contribution to our anytime solver is not a sampling procedure at inference, but a denoising training curriculum.
Chinese Translation
扩散模型和递归推理器都是迭代式的,但它们在迭代之间传递信息的方式不同。我们在扩散去噪器中加入了一个持久的隐藏状态,并去除了其时间步条件化,从而留下一个可运行至任意深度的单一共享更新。其结果是一个随时(anytime)求解器:准确率随推理深度的增加持续提升,远远超出训练时所使用的展开长度和反向传播窗口,在 Sudoku-Extreme 上达到 99.90% 的精确求解率。我们还在 Maze-Unique 上获得了 98.93% 的求解率。令人惊讶的是,在推理阶段渐进式去噪并非必要:将每一步中所有非线索变量替换为新的高斯噪声、使破坏程度保持在最大值,仍能保持近乎完美的求解能力,并收敛到稳定的解。这种简单的噪声注入机制使单条轨迹能够高效地探索解空间并收敛到正确答案,而无需以往推理模型所依赖的并行展开、候选选择或外部验证器。然而,有序的退火式破坏在训练阶段仍然至关重要,这表明扩散模型对这一随时求解器的主要贡献并非推理阶段的采样过程,而是一种去噪训练课程。
cs.LG / 101 / 2609.01493

Rethinking Learnability in Offline Data-driven Optimization

重新思考离线数据驱动优化中的可学习性
Qian, Chao, Wang, Chen-Guang, Tan, Rong-Xi, Xue, Ke
Abstract
Black-Box Optimization (BBO) has broad applications, while traditional algorithms such as evolutionary algorithms and Bayesian optimization face efficiency challenges as real-world BBO problems grow increasingly complex. Data-driven optimization has been the most popular paradigm to improve the efficiency of BBO, by learning from data. Offline data-driven optimization seeks high-quality solutions using only a fixed set of previous evaluations, attracting substantial attention because it requires no additional online evaluations. Many offline optimization methods have been proposed, but a fundamental question remains unanswered: what learnability is sufficient for offline optimization? Prior theoretical studies show that Probably Approximately Correct (PAC) learnability is insufficient, as the optimal region may remain poorly learned even when most regions are well learned. In this paper, we propose algorithm-dependent learnability, which requires accuracy only on the optimizer's trajectory. We prove that its value-query form is sufficient for representative discrete settings, including greedy and local search for submodular maximization, while its first-order analogue is sufficient for projected gradient descent on convex minimization. Motivated by this notion, we formalize a trajectory-learning framework comprising trajectory construction, trajectory modeling, and candidate generation, and analyze existing trajectory-based methods under it. We further propose Uncertainty-aware Gradient-guided Trajectory Learning (UGTL), which constructs locally coherent improvement trajectories reflecting plausible search paths, models them with conditional diffusion, and selects a diverse candidate set. Our experiments show that UGTL achieves the best average rank, 3.1/25, among 25 methods on Design-Bench tasks, and confirm that our trajectory construction plays a significant role in the improvement.
Chinese Translation
黑盒优化(Black-Box Optimization, BBO)具有广泛的应用,然而随着现实世界BBO问题日益复杂,演化算法和贝叶斯优化等传统算法面临效率方面的挑战。数据驱动优化通过从数据中学习,已成为提升BBO效率最流行的范式。离线数据驱动优化仅利用一组固定的历史评估来寻找高质量解,由于无需额外的在线评估而受到广泛关注。尽管已有许多离线优化方法被提出,但一个根本问题仍未得到解答:什么样的可学习性对离线优化而言是充分的?先前的理论研究表明,概率近似正确(Probably Approximately Correct, PAC)可学习性是不充分的,因为即使在大多数区域被充分学习的情况下,最优区域仍可能未被很好地学习。本文提出了算法依赖的可学习性(algorithm-dependent learnability),其仅要求在优化器轨迹上具有精度。我们证明,其值查询形式对于具有代表性的离散设置(包括次模最大化的贪心算法和局部搜索)是充分的,而其一阶类比形式对于凸最小化上的投影梯度下降是充分的。基于这一概念,我们形式化了一个轨迹学习框架,包括轨迹构建、轨迹建模和候选生成,并在该框架下分析了现有的基于轨迹的方法。我们进一步提出了不确定性感知的梯度引导轨迹学习(Uncertainty-aware Gradient-guided Trajectory Learning, UGTL),该方法构建反映合理搜索路径的局部连贯改进轨迹,使用条件扩散模型对其进行建模,并选择多样化的候选集合。实验表明,UGTL在Design-Bench任务上于25种方法中取得了最佳平均排名3.1/25,并证实了我们的轨迹构建方法在性能提升中发挥了重要作用。
cs.LG / 102 / 2609.01495

Optimizing Byzantine Node Placement in Decentralized Federated Learning

去中心化联邦学习中拜占庭节点放置的优化
Gabrielli, Edoardo, Tolomei, Gabriele
Abstract
Security evaluations of decentralized federated learning (DFL) typically focus on how Byzantine participants behave, while largely overlooking which participants are compromised. Yet, because aggregation is distributed over a communication graph, the placement of Byzantine nodes determines how malicious influence propagates through the network. We therefore treat Byzantine placement as an explicit adversarial decision and formulate the attacker's objective as selecting, under a fixed compromise budget, the set of participants that maximizes its finite-time impact on honest nodes. To approximate this objective without executing the learning process for every candidate placement, we introduce Byzantine Placement Influence (BPI), a set-level measure derived from the actual gossip dynamics that quantifies the cumulative exposure of honest nodes to Byzantine sources over the training horizon. Unlike placement criteria based on node centrality heuristics, BPI directly accounts for weighted multi-hop propagation and interactions among compromised nodes. We develop efficient algorithms for optimizing BPI and evaluate them across six heterogeneous graph families, untargeted model poisoning, and backdoor attacks. BPI-guided placements consistently identify highly damaging configurations across different network structures and remain effective when the linear gossip assumption is relaxed through Byzantine-robust aggregation. Our results show that Byzantine placement is a critical but under-modeled dimension of DFL threat models and robustness evaluations.
Chinese Translation
去中心化联邦学习(DFL)的安全性评估通常聚焦于拜占庭参与者的行为方式,而在很大程度上忽视了哪些参与者被攻陷。然而,由于聚合过程分布在通信图上,拜占庭节点的放置位置决定了恶意影响在网络中的传播方式。因此,我们将拜占庭节点的放置视为一种显式的对抗性决策,并将攻击者的目标形式化为:在固定攻陷预算下,选择能够最大化其对诚实节点有限时间影响的参与者集合。为了在不针对每个候选放置方案执行学习过程的情况下逼近该目标,我们提出了拜占庭放置影响力(Byzantine Placement Influence,BPI),这是一个基于实际gossip动态推导的集合级度量,用于量化在训练时间范围内诚实节点对拜占庭源的累积暴露程度。与基于节点中心性启发式准则的放置方法不同,BPI直接考虑了加权多跳传播以及被攻陷节点之间的相互作用。我们开发了用于优化BPI的高效算法,并在六类异构图家族、非目标性模型投毒攻击和后门攻击上对其进行了评估。结果表明,由BPI指导的放置方案在不同网络结构下始终能识别出破坏性极高的配置,并且在通过拜占庭鲁棒聚合放松线性gossip假设的情况下仍然有效。我们的研究结果表明,拜占庭节点的放置是DFL威胁模型与鲁棒性评估中一个关键但尚未被充分建模的维度。
cs.LG / 103 / 2609.01507

LatentPress: Context Compression Beyond Text and Vision

LatentPress:超越文本与视觉的上下文压缩
Zhou, Zhengze, Sang, Hejian
Abstract
Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even when its consumer is a language model. We introduce LatentPress, which writes conversational histories and long documents into a third representation: continuous memory tokens that a frozen decoder reads directly through its input-embedding interface, with no text reconstruction at inference. A small reader-matched writer compresses $4$-$16\times$ while training only an adapter (4.2M-26.2M parameters, $\sim\!0.1\%$ of the decoder). On LongMemEval, LatentPress reaches $0.504$ accuracy at $7.70\times$ compression versus $0.490$ for uncompressed evidence, outperforming text summaries (0.184) and OCR-based compression (0.426 to 0.312). On LongBench-QA, in-domain writers match or exceed raw-context reading at $4$-$8\times$ compression, while $16\times$ trails raw. Writing takes 43ms per conversation, roughly an order of magnitude faster than text summarization or OCR reconstruction, and reading is $5$-$9\times$ faster than raw context or cached OCR. We validate the interface under two transfer settings, zero-shot from UltraChat to LongMemEval memory QA and from LongMemEval-derived QA to unseen LongBench document domains, establishing direct soft tokens as a practical machine-facing context interface beyond text and vision. The implementation of the experiments could be found at: https://github.com/HJSang/LatentPress .
Chinese Translation
压缩后的上下文通常以人类可读的文本形式或以必须被解码的渲染图像形式携带,即便其消费者是语言模型。我们提出LatentPress,它将对话历史和长文档写入第三种表示形式:连续的记忆token(memory tokens),由冻结的解码器通过其输入嵌入接口直接读取,在推理时无需文本重建。一个小型与读取器匹配的写入器(writer)仅通过训练一个适配器(420万至2620万参数,约占解码器的0.1%)即可实现4至16倍的压缩。在LongMemEval上,LatentPress在7.70倍压缩下达到0.504的准确率,高于使用未压缩证据的0.490,并优于文本摘要(0.184)和基于OCR的压缩(0.426至0.312)。在LongBench-QA上,域内写入器在4至8倍压缩下可匹配或超过原始上下文读取的效果,而16倍压缩则落后于原始上下文。写入每段对话仅需43毫秒,比文本摘要或OCR重建快大约一个数量级,读取速度比原始上下文或缓存的OCR快5至9倍。我们在两种迁移设置下验证了该接口:从UltraChat到LongMemEval记忆问答的零样本迁移,以及从基于LongMemEval的问答到未见的LongBench文档领域的迁移,确立了直接软token(soft tokens)作为一种超越文本与视觉的、面向机器的实用上下文接口。实验实现可参见:https://github.com/HJSang/LatentPress 。
cs.LG / 104 / 2609.01537

Quantum Sparse Autoencoders for Q-Matrix Estimation in Cognitive Diagnosis

用于认知诊断中Q矩阵估计的量子稀疏自编码器
Zidan, Arif Hassan, Pan, Yi, Guo, Bowen, Li, Xiang, Bao, Yu, Wang, Yingfeng, Liu, Tianming, Zhang, Wei
Abstract
Q-matrices play a central role in cognitive diagnosis within educational data mining (EDM), specifying which latent skills each assessment item requires. Data-driven Q-matrix estimation remains challenging when assessments involve many correlated skills and when real response patterns depart from idealized generative assumptions. We introduce a novel quantum sparse autoencoder (QSAE) for Q-matrix estimation, which, to the best of our knowledge, is the first application of quantum machine learning (QML) to cognitive diagnosis. Overall, the QSAE embeds each student's binary response vector into a quantum circuit using an encoder, compresses it into a sparse latent representation, and maps that representation to the Q-matrix. We benchmark the QSAE against a classical autoencoder (CAE) across 60 simulated datasets and 9 real-world assessment datasets. The results reveal complementary strengths. Although the CAE partially achieves higher average accuracy under several simulation conditions, the QSAE is substantially more stable across replications, exhibiting lower variance in 49 of the 60 conditions. Moreover, on real assessment data, the QSAE outperforms the CAE on 6 of the 9 datasets. These findings suggest that the principal advancement of QML in this setting is not universal accuracy improvement, but enhanced robustness and capability to explore latent-structure complexity in real datasets.
Chinese Translation
Q矩阵在教育数据挖掘(EDM)的认知诊断中起着核心作用,它明确了每道评估题目所考察的潜在技能。当评估涉及大量相关技能,且真实作答模式偏离理想化的生成假设时,基于数据驱动的Q矩阵估计仍然具有挑战性。我们提出了一种新颖的量子稀疏自编码器(QSAE)用于Q矩阵估计,据我们所知,这是量子机器学习(QML)在认知诊断中的首次应用。总体而言,QSAE利用编码器将每个学生的二值作答向量嵌入量子电路中,将其压缩为稀疏的潜在表示,并将该表示映射到Q矩阵。我们在60个模拟数据集和9个真实评估数据集上将QSAE与经典自编码器(CAE)进行了基准比较。结果显示二者各具优势:尽管CAE在若干模拟条件下部分实现了更高的平均准确率,但QSAE在重复实验中表现出更强的稳定性,在60个条件中的49个条件下方差更低。此外,在真实评估数据上,QSAE在9个数据集中的6个上优于CAE。这些发现表明,QML在这一领域的核心进展并非普适的准确率提升,而是增强了鲁棒性以及探索真实数据集中潜在结构复杂性的能力。
cs.LG / 105 / 2609.01549

NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games

NashDreamer:面向零和不完美信息博弈的基于模型的强化学习
Holeček, Tomáš, Lisý, Viliam
Abstract
Model-based reinforcement learning (MBRL) has achieved remarkable results in single-agent domains, yet its extension to competitive imperfect information games (IIGs) remains underexplored. In multi-agent settings, opponent-induced non-stationarity complicates the learning process, and decentralized model learning faces severe identifiability barriers, which we argue make centralized model learning a mathematical necessity. Building on this analysis, we propose NashDreamer, a principled MBRL framework for two-player zero-sum IIGs. It introduces a centralized Multi-Agent Recurrent State-Space Model (MARSSM) that decouples environment dynamics from the effect of players' strategies on their individual observations. NashDreamer is designed to use arbitrary policy gradient algorithms and inherits their convergence guarantees towards Nash equilibria under an idealized model. Empirical evaluations across four benchmark games demonstrate that NashDreamer substantially improves sample efficiency over model-free baselines early in the training. Finally, we theoretically analyze the architecture's optimization landscape, identifying the vulnerability of the Dreamer family of algorithms to posterior collapse in stochastic environments. We leave it as an open challenge.
Chinese Translation
基于模型的强化学习(MBRL)在单智能体领域已取得显著成果,但其在竞争性不完美信息博弈(IIGs)中的扩展仍未得到充分探索。在多智能体环境中,对手导致的非平稳性使学习过程复杂化,而去中心化的模型学习面临严重的可辨识性障碍——我们认为这使得集中式模型学习成为一种数学上的必然。基于这一分析,我们提出了NashDreamer,一个面向双人零和不完美信息博弈的原则性MBRL框架。该框架引入了一个集中式的多智能体循环状态空间模型(MARSSM),将环境动态与玩家策略对其个体观测的影响解耦。NashDreamer被设计为可以使用任意策略梯度算法,并在理想化模型下继承这些算法向纳什均衡收敛的保证。在四个基准博弈上的实证评估表明,NashDreamer在训练初期的样本效率显著优于无模型基线方法。最后,我们从理论上分析了该架构的优化景观,指出Dreamer系列算法在随机环境中易受后验坍缩影响的脆弱性,并将其作为一个开放性挑战留给未来研究。
cs.LG / 106 / 2609.01550

A Mathematical Theory of Reusable Neural Bases for Network Compression

用于网络压缩的可复用神经基底的数学理论
Wang, Binshuai, Wei, Peng
Abstract
As large AI models become increasingly prevalent across a wide range of applications, memory cost has become a critical bottleneck in both training and inference. To mitigate this issue, we introduce the Linear Reusable Neural Bases Architecture (LRNBA), a novel framework aimed at improving parameter efficiency and reducing memory cost. Inspired by recurrent neural network (RNN) designs, the core idea of our approach is to represent each network block as a linear combination of a shared set of neural bases, thereby enjoying highly network compression rate while maintaining stable training. The proposed architecture allows for the construction of significantly wider and deeper networks under the same parameter budget. Extensive experiments demonstrate that our model achieves comparable or even faster convergence and lower loss than classical architectures, while maintaining stable training dynamics.
Chinese Translation
随着大型AI模型在各类应用中的日益普及,内存开销已成为训练和推理过程中的关键瓶颈。为缓解这一问题,我们提出了线性可复用神经基底架构(Linear Reusable Neural Bases Architecture,LRNBA),这是一种旨在提高参数效率并降低内存开销的新型框架。受循环神经网络(RNN)设计的启发,我们方法的核心思想是将每个网络块表示为一组共享神经基底的线性组合,从而在保持训练稳定的同时实现高度的网络压缩率。所提出的架构使得在相同参数预算下能够构建显著更宽、更深的网络。大量实验表明,与经典架构相比,我们的模型实现了相当甚至更快的收敛速度和更低的损失,同时保持了稳定的训练动态。
cs.LG / 107 / 2609.01556

Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories

被检索却未被排序:结构化检索中的表层形式偏差,从数学到智能体轨迹
Rashid, Nabira, Kellis, Manolis
Abstract
We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that share underlying structure but not wording, in two unrelated domains under one protocol, competition mathematics (MathNet-Retrieve; 500 queries, 117,088-item corpus) and embodied-agent trajectories (ALFWorld-derived; 118 queries, 336 trajectories). In mathematics the failure is complete: strict Hit@1 at the heaviest disguise tier is 0.0% for both production embedders (bootstrap 95% CI [0.0, 0.0]) while the correct item sits in the top 10 nearly always, and in 95.2 to 99.8% of misses the winner is more lexically similar to the query than the correct answer. In trajectories, where surface variation is incidental, the same models land at or near hypergeometric chance when gold must involve a different object, and below chance for all three embedders once gold must differ in object and receptacle: retrieval anchors on literal tokens, not task structure. A lexical reranker control hurts in mathematics and helps in trajectories (closing 26 to 36% of the gap, CIs excluding zero); its sign reveals whether a benchmark's surface variation is adversarial or incidental. An LLM reranker recovers 5 to 63% of the gap in mathematics and 43 to 76% in trajectories; direction replicates across three judges (all 21 cells positive), but effect sizes, tier profiles, and the outlier judge change with domain (paired differences excluding zero everywhere). Mathematics gains concentrate on well-known competitions (+19.8 points, CI [+6.7, +33.2], one of six cells), so part of the recovery is memorization. In a paired downstream experiment (210 queries, graders at 96 to 99% agreement), oracle retrieval was indistinguishable from adversarially bad retrieval (McNemar p = 0.678); the solver's 69.5% zero-shot accuracy is largely a truncation proxy (97 to 100% on finished answers), leaving no headroom.
Chinese Translation
我们在刻意将表层形式与语义分离的场景下评估嵌入检索:即在统一协议下,于两个不相关领域中检索具有相同底层结构但措辞不同的条目——竞赛数学(MathNet-Retrieve;500个查询,117,088条语料库)和具身智能体轨迹(基于ALFWorld;118个查询,336条轨迹)。在数学领域,失败是完全的:在伪装程度最高的层级上,两个生产级嵌入模型的严格Hit@1均为0.0%(bootstrap 95%置信区间[0.0, 0.0]),而正确条目几乎总能进入前10名;在95.2%至99.8%的失败案例中,胜出者与查询的词汇相似度高于正确答案。在轨迹领域,表层变化是偶然性的:当标准答案必须涉及不同对象时,同样这些模型的表现处于或接近超几何随机水平;而当标准答案必须在对象和容器上均不同时,三个嵌入模型的表现均低于随机水平——检索锚定于字面词元而非任务结构。词汇重排序器对照组在数学领域有害,在轨迹领域有益(弥补26%至36%的差距,置信区间不包含零);其正负方向揭示了基准测试的表层变化是对抗性的还是偶然性的。LLM重排序器在数学领域恢复了5%至63%的差距,在轨迹领域恢复了43%至76%;该方向在三位评审模型间均可复现(全部21个单元格均为正值),但效应大小、层级分布以及异常评审随领域而变化(配对差异均不包含零)。数学领域的提升集中于知名竞赛(+19.8分,置信区间[+6.7, +33.2],六个单元格中的一个),因此部分恢复源于记忆效应。在一项配对下游实验(210个查询,评分者一致性达96%至99%)中,oracle检索与对抗性劣质检索无显著差异(McNemar p = 0.678);求解器69.5%的零样本准确率很大程度上是截断的代理指标(在完整答案上为97%至100%),几乎没有提升空间。
cs.LG / 108 / 2609.01558

Gradient-Update Mismatch: Rethinking Conflict-Free Training of Physics-Informed Neural Networks

梯度更新失配:重新思考物理信息神经网络的无冲突训练
Xiao, Jing, Chen, Xinhai, Wang, Qinglin, Jia, Menghan, Lai, Zhiquan, Li, Dongsheng, Liu, Jie, Li, Tiejun
Abstract
Training Physics-Informed Neural Networks (PINNs) requires jointly optimizing physics residual and initial/boundary condition loss terms, which often induce conflicting gradients. Gradient surgery methods mitigate this issue by constructing directions from loss-specific gradients to reduce conflict before optimizer transformation. However, even when the constructed direction is conflict-free, this property may not be preserved after optimizer transformation. Let $a_t$ denote the direction constructed by gradient surgery, $u_t$ the optimizer proposal, and $\mathcal{C}_t$ the conflict-free cone induced by the loss-specific gradients. We show that modern optimizers can transform $a_t$ through mechanisms such as historical state, adaptive scaling, preconditioning, or decoupled weight decay, so $a_t \in \mathcal{C}_t$ does not generally imply $u_t \in \mathcal{C}_t$. We refer to this optimizer-induced discrepancy in conflict-freeness between $a_t$ and $u_t$ as Gradient-Update Mismatch (GUM). Accordingly, we propose Gradient-Update Alignment (GUA), which projects $u_t$ onto $\mathcal{C}_t$ to obtain the aligned update $p_t$ and applies $p_t$ to the parameters. When the optimizer maintains internal state, GUA further adjusts this state toward targets reconstructed from the applied update. We conduct extensive experiments and find that GUM is widespread across momentum, adaptive, and curvature-based optimizers, with conflict rates reaching up to 86.3%. Across all PINN settings, GUA achieves conflict-free applied updates and consistently improves various gradient surgery methods, reducing the relative $L_2$ error by up to 98.2% in individual settings. Data and code are available at https://github.com/JingXiao10/GUA.
Chinese Translation
物理信息神经网络的训练需要联合优化物理残差项和初始/边界条件损失项,而这常常会引发梯度冲突。梯度手术方法通过从各损失项的梯度构造方向,在优化器变换之前降低冲突来缓解该问题。然而,即使构造出的方向是无冲突的,这一性质在经过优化器变换后也可能无法保持。设 $a_t$ 表示梯度手术构造的方向,$u_t$ 表示优化器的提议更新,$\mathcal{C}_t$ 表示由各损失项梯度诱导的无冲突锥。我们证明,现代优化器可以通过历史状态、自适应缩放、预条件化或解耦权重衰减等机制对 $a_t$ 进行变换,因此 $a_t \in \mathcal{C}_t$ 一般并不意味着 $u_t \in \mathcal{C}_t$。我们将这种由优化器引起的 $a_t$ 与 $u_t$ 之间无冲突性的差异称为梯度更新失配。为此,我们提出梯度更新对齐方法,该方法将 $u_t$ 投影到 $\mathcal{C}_t$ 上以获得对齐后的更新 $p_t$,并将其应用于参数。当优化器维护内部状态时,GUA 还会进一步将该状态调整至由实际施加更新重构的目标。我们通过大量实验发现,GUM 在动量类、自适应类和基于曲率的优化器中普遍存在,冲突率最高可达 86.3%。在所有 PINN 设置中,GUA 均实现了无冲突的实际更新,并持续提升各种梯度手术方法的性能,在个别设置中将相对 $L_2$ 误差降低了最多 98.2%。数据与代码见 https://github.com/JingXiao10/GUA。
cs.LG / 109 / 2609.01587

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

大语言模型中量化损伤的结构:为何下一比特应全局分配
Hu, Jundong, Ramachandran, Shekar
Abstract
Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional precision budget. Using causal mixed-precision intervention as ground truth (raise each layer to 8-bit in turn and measure the accuracy it recovers) across 9 open-weight models in 4 architecture families, we test 3 intuitive hypotheses: that quantization damage lives in task circuits, where the model computes, or in weight statistics. None of them predicts which layers benefit from restored precision. Recovery is instead diffuse: for 8 of 9 models, recovering 75% of the gap takes roughly half the layers; the lone exception, Qwen3-8B, is sharply concentrated. At a matched precision budget, spending it globally on finer quantization granularity beats locally repairing the most recoverable layers for all 8 group-128-compatible models (all but OpenLLaMA, whose width rules out group-128), by 21-52 points, including the concentrated Qwen3-8B. We report 2 secondary findings: the residual is budget-limited (8-bit is near-lossless in our evaluation across RTN, GPTQ, and AWQ), and the location of peak recovery correlates with architecture within a family, though not across families. Within this budget setting, global granularity is a better default than selectively protecting critical layers. More broadly, cheap signals that correlate with quantization damage do not necessarily identify where restoring precision improves accuracy; this must be tested with causal intervention.
Chinese Translation
训练后量化(PTQ)被广泛用于降低大语言模型(LLM)的部署成本,但其精度损失并不均匀,且通常需要针对每个模型单独调优。我们研究了量化损伤发生的位置,以及如何分配少量额外的精度预算。我们以因果混合精度干预作为真值(依次将每一层提升至8比特,并测量其恢复的精度),在4个架构家族的9个开源权重模型上进行实验,检验了三个直觉性假设:量化损伤位于任务回路中、位于模型计算之处、或与权重统计特性相关。结果显示,三者均无法预测哪些层会从恢复的精度中受益。相反,恢复是弥散性的:在9个模型中有8个,恢复75%的精度差距大约只需要一半的层;唯一的例外是Qwen3-8B,其损伤高度集中。在精度预算相同的情况下,将预算全局用于更细的量化粒度,在所有8个兼容group-128的模型上(除OpenLLaMA外,其宽度不允许使用group-128),均比局部修复最具可恢复性的层高出21至52个百分点,其中包括损伤集中的Qwen3-8B。我们还报告了两项次要发现:残余损失受预算限制(在我们的评估中,8比特在RTN、GPTQ和AWQ方法下几乎无损),且峰值恢复位置在家族内部与架构相关,但跨家族则不然。在该预算设定下,全局粒度是比选择性保护关键层更好的默认策略。更广泛地说,与量化损伤相关的廉价信号未必能指出恢复精度在何处能够提升准确性;这必须通过因果干预来验证。
机器人学 (Robotics)
39
cs.RO / 1 / 2609.00311

Inverse kinematic solution for generic 3R positional robots using Conformal Geometric Algebra

基于共形几何代数的通用3R位置机器人逆运动学求解
Nayak, Abhilash, Salunkhe, Durgesh Haribhau
Abstract
The inverse kinematics of generic 3R robots has been investigated through multiple approaches, mainly algebraic methods involving the solution of certain equation sets. Previous geometric interpretations of the solution, characterized as the intersection of a pair of conics have been confined to the joint-space domain. In this article, we study the Inverse Kinematic Model (IKM) of 3R robots, using the advantages of Conformal Geometric Algebra (CGA) to provide further insights on its kinematic properties. Our approach directly yields a univariate polynomial in terms of theta_2 without the need to eliminate theta_1 and theta_3 by reframing the problem as the intersection of two circles, which are fundamental elements within this algebraic framework.
Chinese Translation
通用3R机器人的逆运动学问题已通过多种方法进行研究,其中主要涉及求解特定方程组的代数方法。以往对该解的几何解释——其特征为一对二次曲线的交点——仅局限于关节空间域。本文利用共形几何代数(Conformal Geometric Algebra, CGA)的优势,研究了3R机器人的逆运动学模型(Inverse Kinematic Model, IKM),从而对其运动学特性提供了更深入的见解。我们的方法通过将该问题重新表述为两个圆的交点(圆是该代数框架中的基本元素),直接得到关于theta_2的一元多项式,而无需消去theta_1和theta_3。
cs.RO / 2 / 2609.00316

Geometric analysis of generic 3R robots, and necessary and sufficient conditions for a class of orthogonal robots to have four IKS

通用3R机器人的几何分析及一类正交机器人具有四个逆运动学解(IKS)的充要条件
Salunkhe, Durgesh Haribhau, Nayak, Abhilash
Abstract
The kinematic analysis of a generic 3R robot has been investigated with multiple approaches in the past. The algebraic approaches have established concrete results but are unfortunately limited to special classes or architectural simplifications. Geometric approaches on the other hand have extended the analysis to generic robots while also providing an intuitive understanding of their kinematic properties. We use the best of both approaches to present the kinematic analysis of a generic 3R robot, using the inverse kinematic model inherited from a method based on conformal geometric algebra. The paper discusses a generic framework to study the conditions for a 3R robot to have four inverse kinematic solutions (IKS) and allows to study the distribution of IKS as seen in workspace. The necessary and sufficient conditions for a class of orthogonal robots are presented using the proposed approach.
Chinese Translation
通用3R机器人的运动学分析在过去已通过多种方法进行了研究。代数方法虽然建立了具体的结果,但遗憾的是仅限于特殊类别或结构简化情形。另一方面,几何方法不仅将分析扩展到了通用机器人,还对其运动学特性提供了直观的理解。我们结合这两种方法的优点,利用继承自基于共形几何代数(conformal geometric algebra)方法的逆运动学模型,对通用3R机器人进行了运动学分析。本文讨论了一个通用框架,用于研究3R机器人具有四个逆运动学解(IKS)的条件,并能够研究工作空间中逆运动学解的分布情况。利用所提出的方法,给出了一类正交机器人具有四个逆运动学解的充要条件。
cs.RO / 3 / 2609.00333

Scene Graph-based Driving Scenario Extraction for Automotive Egocentric Datasets

基于场景图的汽车第一视角数据集驾驶场景提取
Ramdhan, Stefan, Dagenais, Kyanna, Pantelic, Vera, Bandur, Victor, Lawford, Mark
Abstract
Extracting scenarios from unlabelled real-world sensor data streams is a critical but challenging task in the development process of automated driving systems (ADS). Automatically sifting through large datasets to spatially and temporally locate critical scenarios can enable scenario-based coverage analysis of ADS datasets. In this paper, we present a method for extracting scenarios from egocentric datasets using scene graphs and Linear Temporal Logic (LTL). We first process egocentric sensor data and HD maps to generate a sequence of scene graphs representing a driving scenario. Next, we use LTL to formally specify driving scenarios of interest, then extract all instances of the scenarios from the dataset using an off-the-shelf model checker, which evaluates the LTL formula against the sequence of scene graphs. Our approach can be used on both simulated and real world datasets. We evaluate the method on the training and validation datasets from Argoverse 2 consisting of 850 15-second real-world driving logs, and several videos of dashcam footage. We demonstrate the effectiveness of our approach for extracting and querying scenarios by evaluating against a rule-based benchmark based on track annotations and HD maps.
Chinese Translation
从未标注的真实世界传感器数据流中提取场景,是自动驾驶系统(ADS)开发过程中一项关键但具有挑战性的任务。自动筛选大型数据集并在空间和时间上定位关键场景,可以实现对ADS数据集的基于场景的覆盖率分析。本文提出了一种利用场景图和线性时序逻辑(LTL)从第一视角数据集中提取场景的方法。我们首先处理第一视角传感器数据和高精地图,生成表示驾驶场景的场景图序列。接下来,我们使用LTL形式化地描述感兴趣的驾驶场景,然后使用现成的模型检测器对场景图序列评估LTL公式,从而从数据集中提取所有场景实例。我们的方法既适用于仿真数据集,也适用于真实世界数据集。我们在Argoverse 2的训练和验证数据集(包含850条15秒的真实世界驾驶日志)以及若干行车记录仪视频上对该方法进行了评估。通过基于轨迹标注和高精地图的基于规则的基准进行对比评估,我们证明了该方法在提取和查询场景方面的有效性。
cs.RO / 4 / 2609.00385

Risk-Aware Decision-Making for Autonomous Overtaking: A World Model-Based Mixture-of-Experts Framework

面向自主超车的风险感知决策:一种基于世界模型的混合专家框架
Liu, Yongzhi, Zhang, Sunan, Xu, Jinchang, Wang, Jiawei, Qiu, Yushu, Lv, Chen, Zhuang, Weichao
Abstract
Autonomous highway overtaking demands foresighted decision-making to handle complex interactions, stochastic traffic evolution, and temporal risk accumulation. However, standard safe reinforcement learning approaches typically rely on implicit value-based risk estimations rather than explicit dynamics modeling, thereby struggling to accurately capture complex risk propagation over multi-step horizons. This limitation frequently results in behaviors that are locally safe but induce substantial latent risks in the long term. To address this, a World Model-based Risk-aware Mixture-of-Experts (WM-RMoE) framework is proposed. First, a learned latent dynamics model facilitates parallel multi-step rollouts, elevating safety assessment from the action level to the trajectory level via cumulative risk evaluation. Second, to enhance robustness under varying interaction intensities, a hierarchical gating mechanism dynamically coordinates experts across long-horizon, short-horizon, and rule-based safety modules. Furthermore, a Gaussian Mixture Model is integrated to preserve multimodal maneuvering branches, thereby mitigating the issue of behavioral mode averaging. Experimental results demonstrate that WM-RMoE significantly outperforms representative baselines in terms of safety compliance, decision stability, and generalization capability. Furthermore, benefiting from the risk-aware formulation, the proposed framework uniquely exhibits the ability to generate foresighted and semantically distinct overtaking maneuvers across diverse traffic densities.
Chinese Translation
高速公路自主超车需要具有前瞻性的决策能力,以应对复杂的交互、随机的交通演化以及时间维度的风险累积。然而,标准的安全强化学习方法通常依赖于隐式的基于价值的风险估计,而非显式的动力学建模,因而难以准确捕捉多步时域内复杂的风险传播。这一局限常常导致智能体产生局部安全但长期潜藏较大风险的行为。为解决该问题,本文提出了一种基于世界模型的风险感知混合专家框架。首先,通过学习得到的潜在动力学模型支持并行的多步展开,借助累积风险评估将安全性评估从动作层面提升至轨迹层面。其次,为增强在不同交互强度下的鲁棒性,一种分层门控机制动态协调长时域、短时域以及基于规则的安全模块等多个专家。此外,框架中集成了高斯混合模型以保留多模态的机动分支,从而缓解行为模式平均化问题。实验结果表明,WM-RMoE 在安全性合规、决策稳定性和泛化能力方面显著优于代表性基线方法。此外,得益于风险感知的建模方式,所提出的框架在多种交通密度下均能独特地生成具有前瞻性且语义上截然不同的超车机动。
cs.RO / 5 / 2609.00478

Beyond Object Selection:Markerless Gaze-based Robot Placement at Arbitrary Position

超越物体选择:基于视线追踪的无标记机器人任意位置放置
Lai, Yuzhi, Marx, William, Yuan, Shenghai, Li, Peizheng, Ran, Zhuoyu, Zell, Andreas
Abstract
Gaze-based assistive manipulation typically supports object selection, while arbitrary-position placement requires accurate spatial alignment between the headset and robot. However, for gaze-based manipulation, pose accuracy does not necessarily translate into task accuracy: translational and rotational errors jointly affect the transformed gaze ray and may compensate for each other. To study cross-device alignment from this task-oriented perspective, we present a markerless interaction framework and a dedicated cross-device dataset. We propose Graph-based Reference Selection to address sparse robot references. We further develop and benchmark multiple task-specific alignment pipelines under a unified protocol. Specifically, we introduce Gaze--Surface Intersection Error (GSIE), which directly measures the spatial error of the gaze-specified target. Experiments show that alignment methods ranked highly by conventional pose metrics are not always optimal in GSIE, demonstrating the importance of evaluating gaze-based manipulation at the task level.
Chinese Translation
基于视线的辅助操作通常只支持物体选择,而任意位置的放置则需要头显与机器人之间精确的空间对齐。然而,对于基于视线的操作而言,位姿精度并不一定意味着任务精度:平移误差和旋转误差会共同影响变换后的视线射线,并且可能相互补偿。为从这一面向任务的角度研究跨设备对齐问题,我们提出了一个无标记交互框架以及一个专门的跨设备数据集。我们提出了基于图的参考点选择方法(Graph-based Reference Selection)以解决机器人参考点稀疏的问题。我们进一步在统一协议下开发并评测了多种面向任务的对齐流程。具体而言,我们提出了视线—表面交点误差(Gaze–Surface Intersection Error, GSIE),用于直接度量视线指定目标的空间误差。实验表明,按传统位姿指标排名靠前的对齐方法在GSIE上并非总是最优,这证明了在任务层面评估基于视线的操作的重要性。
cs.RO / 6 / 2609.00539

Exploring Nonlinear Body Oscillations for Natural Quadruped Gaits

探索非线性身体振荡以实现自然的四足步态
Schmidt, Annika, Calzolari, Davide, Sachtler, Arne, Loeffl, Florian, Seidel, Daniel, Hermann, Milan, Burger, Robert, Gumpert, Thomas, Raffin, Antonin, Ehlert, Tristan, Pries, Maximilian, Wandinger, David, Schmidt, Florian, Keppler, Manuel, Lee, Jinoh, Albu-Schäffer, Alin
Abstract
Animals' body morphology shapes the gait patterns they can perform, where mechanical resonance reduces the need for active control. By tuning posture and muscle stiffness, they leverage their embodied intelligence to achieve effective gaits for different speeds. In contrast, most quadruped robots are not specifically designed to exploit mechanical resonance due to the complexity of nonlinear dynamics and require dedicated locomotion controllers. To provide an alternative, we present a proof of concept framework making the nonlinear dynamics of a robot predictable in the design process and show how this knowledge can be leveraged such that multi-gait locomotion can emerge from nonlinear resonances, shaped by gravity, inertia, and elasticity. We present the highly compliant quadruped robot eBert, on which we identify six nonlinear normal modes (NNMs) using our new theoretical tools and validate their existence in simulation and hardware. With black-box optimization to determine step length, simulations show how each NNM naturally develops into a distinct gait, manifesting different speeds, which also largely transfers to the robotic hardware. Our experiments show that eBert can exploit its mechanics to generate task-specific movements which may serve as foundation for designing a new generation of agile and efficient robots leveraging embodied intelligence.
Chinese Translation
动物的身体形态决定了它们所能实现的步态模式,其中机械共振减少了对主动控制的需求。通过调节姿态和肌肉刚度,动物利用其具身智能在不同速度下实现高效的步态。相比之下,由于非线性动力学的复杂性,大多数四足机器人并非专门为利用机械共振而设计,因而需要专门的移动控制器。作为一种替代方案,我们提出了一个概念验证框架,使机器人的非线性动力学在设计过程中变得可预测,并展示了如何利用这一知识,使多步态运动能够从由重力、惯性和弹性塑造的非线性共振中涌现。我们展示了高度柔顺的四足机器人 eBert,借助我们新的理论工具,在其上识别出六个非线性正规模态(NNMs),并在仿真和硬件上验证了它们的存在。通过黑盒优化确定步长,仿真表明每个非线性正规模态如何自然地发展为一种不同的步态,呈现出不同的速度,且这些结果在很大程度上可以迁移到机器人硬件上。我们的实验表明,eBert 能够利用其自身力学特性生成面向特定任务的运动,这可以为设计新一代借助具身智能的敏捷高效机器人奠定基础。
cs.RO / 7 / 2609.00555

Potential-Guided Particle Steering for Negation-Constrained Dexterous Grasping

面向否定约束灵巧抓取的势能引导粒子导向方法
Kim, Geonho, Kim, SooGon, Lee, Jongmin
Abstract
Language-driven dexterous grasp models, such as DextER, perform well when instructions specify where to grasp, but we find they fail systematically when an instruction also specifies where not to grasp (e.g., "grasp the handle but avoid the body"). Existing training corpora, DexGYSNet among them, contain virtually no avoidance instructions, and collecting examples for every possible constraint is impractical. Moreover, because every part mentioned during training denotes a contact target, models may interpret a forbidden part as another region to grasp rather than one to avoid. We therefore introduce an inference-time framework for negation-constrained dexterous grasping that requires no negation-specific training examples. Combining Sequential Monte Carlo with classifier-free guidance, our method guides sampling toward the instructed part while pruning candidates headed for the forbidden region, without any negation examples during training. A frozen 3D part-grounding model localizes the forbidden region from the language instruction. To evaluate this setting, we construct NegGrasp, a benchmark of paired positive/negative instructions with constraint-aware metrics that credit a grasp only if it both accomplishes the task and respects the stated constraint. On NegGrasp, our method reduces the violation rate of the strongest baseline from 57.9% to 17.2% while improving both constraint-aware and physical success.
Chinese Translation
语言驱动的灵巧抓取模型(如 DextER)在指令指定抓取位置时表现良好,但我们发现当指令同时指定不可抓取的位置时(例如“抓住把手但避开机身”),这些模型会出现系统性失败。现有的训练语料库(包括 DexGYSNet 在内)几乎不包含回避类指令,而为每一种可能的约束收集示例并不现实。此外,由于训练中提到的每个部件都表示接触目标,模型可能将禁止部件解读为另一个应抓取的区域而非应避开的区域。因此,我们提出了一种面向否定约束灵巧抓取的推理时框架,无需任何否定类训练示例。该方法将序贯蒙特卡洛(Sequential Monte Carlo)与无分类器引导(classifier-free guidance)相结合,在引导采样朝向指令指定部件的同时,剪除趋向禁止区域的候选方案,且训练过程中不使用任何否定示例。我们利用冻结的 3D 部件定位模型(3D part-grounding model)从语言指令中定位禁止区域。为评估这一设定,我们构建了 NegGrasp 基准,其包含成对的正向/否定指令,并采用约束感知的评估指标——只有当抓取既完成任务又遵守所述约束时才给予评分。在 NegGrasp 上,我们的方法将最强基线的违反率从 57.9% 降低至 17.2%,同时提升了约束感知成功率和物理成功率。
cs.RO / 8 / 2609.00612

A Wearable Pneumatic Device for Continuous, Closed-Loop, Bidirectional Tactile Interaction

一种用于连续、闭环、双向触觉交互的可穿戴气动装置
Pasquier, Cosima du, Smith, Aliyah, Huber, Serin, Phelps, Joshua, Cohen, Ilana A., Chen, Ava, Kennedy III, Monroe, Okamura, Allison M.
Abstract
We present a system of two wearable pneumatic haptic devices that supports continuous, closed-loop, bidirectional tactile interaction at perceptually relevant force and temporal scales. A single device can contain up to twelve pressure sensing channels connected to textile-based pneumatic pouches. Each channel in a device can be used as a sensor, an actuator, or both. As an actuator with integrated sensing, the channel generates stable skin indentation through local closed-loop control. As a sensor, a channel can be mounted (or worn) on any surface, including on a robot gripper or on the human body, and used to measure touch interactions with the environment or a human user. A distributed architecture supports sustained pressure output, rapid dynamic response, and wireless pairing of identical devices in a system to transmit and reproduce tactile pressure signals in real time. Device-level characterization demonstrates force bandwidth exceeding 30 Hz, rapid and well-damped step responses, and extended pressure retention compared to prior compact pneumatic platforms. Human studies show that pressure-based fingertip feedback enables discrimination of force and stiffness, improves teleoperated manipulation by reducing applied pressures by up to 23.1% and task duration by up to 27.4%, and lowers subjective mental workload by 18.8%, particularly under visually constrained conditions. By unifying tactile sensing and haptic feedback within a single pneumatic modality, the device provides a practical foundation for bidirectional touch interaction in teleoperation.
Chinese Translation
我们提出了一套由两个可穿戴气动触觉装置组成的系统,支持在感知相关的力与时间尺度上进行连续、闭环、双向的触觉交互。单个装置最多可包含十二个与织物气动袋相连的压力感知通道。装置中的每个通道既可用作传感器、执行器,也可同时兼任两者。作为集成感知功能的执行器,通道通过局部闭环控制产生稳定的皮肤压痕。作为传感器,通道可安装(或佩戴)于任意表面,包括机器人夹爪或人体,用于测量与环境或人类用户之间的触摸交互。分布式架构支持持续的压力输出、快速的动态响应,以及系统中相同装置的无线配对,从而实现触觉压力信号的实时传输与复现。器件级表征实验表明,该装置的力带宽超过30 Hz,阶跃响应快速且阻尼良好,并且相比以往紧凑型气动平台具有更长的压力保持能力。人体实验表明,基于压力的指尖反馈能够实现对力和刚度的辨别,可将遥操作操作中的施加压力降低最多23.1%、任务时长缩短最多27.4%,并将主观心理负荷降低18.8%,尤其是在视觉受限条件下效果显著。通过在单一气动模式中统一触觉感知与触觉反馈,该装置为遥操作中的双向触觉交互提供了实用的基础。
cs.RO / 9 / 2609.00619

DSG: Dynamic 3D Scene Graph Construction for Embodied Agents in Changing Indoor Environments

DSG:面向变化室内环境中具身智能体的动态三维场景图构建
Liao, Ming, Ye, Chao, Fei, Jianing, Lin, Weiyang
Abstract
In indoor environments, object positions frequently change due to human activities or embodied-agent interactions, causing previously constructed scene graphs to become inconsistent with the current scene. To address this issue, we propose DSG, a dynamic 3D scene graph construction framework that detects object changes and performs spatial relationship reasoning. First, we construct a semantic-aware 3D Gaussian scene representation and develop a dual-view rendering-based object change detection method to enable reliable scene graph node updates. Second, we propose a spatial relationship reasoning method that incorporates multi-granularity visual context, enabling a large language model to identify a richer set of interobject spatial relationships. Furthermore, we introduce DynTHOR, a dynamic indoor scene graph benchmark built on the AI2-THOR simulation platform for evaluating scene graph construction in dynamic environments. Extensive experiments on Dyn-THOR, 3RScan, and real-world scenes demonstrate that DSG consistently outperforms existing methods in both object node construction and spatial relationship reasoning, significantly improving the accuracy of dynamic scene graph construction.
Chinese Translation
在室内环境中,物体位置常因人类活动或具身智能体的交互而频繁变化,导致先前构建的场景图与当前场景不一致。针对这一问题,我们提出了DSG,一种能够检测物体变化并进行空间关系推理的动态三维场景图构建框架。首先,我们构建了语义感知的三维高斯场景表示,并开发了一种基于双视角渲染的物体变化检测方法,以实现可靠的场景图节点更新。其次,我们提出了一种融合多粒度视觉上下文的空间关系推理方法,使大语言模型能够识别更丰富的物体间空间关系。此外,我们引入了DynTHOR,一个基于AI2-THOR仿真平台构建的动态室内场景图基准,用于评估动态环境中的场景图构建。在Dyn-THOR、3RScan以及真实场景上的大量实验表明,DSG在物体节点构建和空间关系推理方面均持续优于现有方法,显著提升了动态场景图构建的准确性。
cs.RO / 10 / 2609.00641

AM-Bench: A Modular Simulation Suite and Benchmark for Aerial Manipulation Policy Learning

AM-Bench:面向空中操作策略学习的模块化仿真套件与基准测试
Wang, Yutong, Lee, Dongjae, Guo, Xiaofeng, Zhan, Yuanzhu, Jiang, Yufei, Saravanan, Bavin, Cao, Muqing, Xie, Jia, Mao, Chenyang, Scherer, Sebastian, Geng, Junyi, Shi, Guanya
Abstract
Standardized benchmarks have played a central role in advancing robot manipulation learning, yet most focus on ground-supported manipulation systems, which limits their applicability to dynamics-critical domains such as aerial manipulation (AM). AM presents distinct system-level challenges, including environmental disturbances, coupled dynamics between the manipulator and floating base, and constrained degrees of freedom. Consequently, task performance depends jointly on robot embodiment, low-level control, and high-level policy design. We introduce AM-Bench, a modular simulation suite and benchmark for multirotor-based AM policy learning. AM-Bench includes representative embodiments spanning underactuated, fully actuated, and overactuated systems, 12 tasks across contact, transport, and constrained interaction, configurable aerodynamic disturbances and actuator saturation, standard low-level controllers, and baseline policy-learning algorithms. Unlike prior manipulation benchmarks that primarily emphasize end-to-end policy performance, AM-Bench enables system-level evaluation of how embodiment, control, disturbances, and policy choices interact. We demonstrate its diagnostic value through three simulation studies spanning high-level policies, policy--control interfaces, and embodiments, together with real-world validation of modeled effects and a hardware test of the learning pipeline.
Chinese Translation
标准化基准测试在推动机器人操作学习的发展中发挥了核心作用,然而现有基准大多聚焦于地面支撑的操作系统,限制了其在空中操作(Aerial Manipulation, AM)等动力学敏感领域的适用性。AM 带来了独特的系统级挑战,包括环境干扰、机械臂与浮动基座之间的耦合动力学,以及受限的自由度。因此,任务性能同时取决于机器人本体构型、底层控制与高层策略设计。我们提出 AM-Bench,一个面向基于多旋翼的空中操作策略学习的模块化仿真套件与基准测试。AM-Bench 包含覆盖欠驱动、全驱动和过驱动系统的代表性机器人本体,12 个涵盖接触、运输和受限交互的任务,可配置的空气动力干扰与执行器饱和,标准底层控制器,以及基线策略学习算法。与以往主要强调端到端策略性能的操作基准不同,AM-Bench 能够在系统层面评估本体构型、控制、干扰与策略选择之间的相互作用。我们通过三项仿真研究展示了其诊断价值,分别涵盖高层策略、策略-控制接口和本体构型,并结合对建模效应的现实世界验证以及学习流程的硬件测试。
cs.RO / 11 / 2609.00648

Human-robot conversation with multiple participants in noisy public spaces

嘈杂公共场所中多参与者的人机对话
Lala, Divesh, Muthukumaran, Yogeeswaran, Fernandes, Vincent, Kato, Kazushi, Fujiki, Shota, Chi, Zihao, Iwasaki, Masaya, Shintani, Taiken, Kawata, Megumi, Sakai, Kazuki, Inoue, Koji, Yoshikawa, Yuicihiro, Kawahara, Tatsuya
Abstract
For noisy real-world environments such as those in open public spaces, spoken dialogue systems for both autonomous robots and avatars should be carefully designed to provide enhanced speech signals. These signals can be used either for speech recognition or, in the case of an avatar system, transmitted as clean speech to a remote operator. This work proposes an audio system that can be used for both these scenarios and was demonstrated as a proof-of-concept at the 2025 World Expo in Osaka. The first scenario is an attentive listening system with the android ERICA, and the second is a conversation support system with mobile Teleco robots, with one of them acting as an avatar for a remote operator. Both systems feature multi-party conversation and use a single multi-channel microphone array. We describe how our audio system not only enhances the speech of multiple speakers in a noisy environment, but provides a form of spatial audio which allows for more immersiveness in avatar-based conversational interactions.
Chinese Translation
对于开放公共空间等嘈杂的真实环境,自主机器人和虚拟化身(avatar)的口语对话系统需要精心设计,以提供增强的语音信号。这些信号既可用于语音识别,也可在虚拟化身系统中作为清晰的语音传输给远程操作员。本工作提出了一种可同时适用于这两种场景的音频系统,并于2025年大阪世界博览会上作为概念验证进行了演示。第一个场景是基于仿人机器人ERICA的专注倾听系统,第二个场景是基于移动机器人Teleco的对话支持系统,其中一个机器人充当远程操作员的虚拟化身。两个系统均支持多方对话,且仅使用单个多通道麦克风阵列。我们描述了该音频系统如何在嘈杂环境中增强多个说话人的语音,同时提供一种空间音频形式,使基于虚拟化身的对话交互更具沉浸感。
cs.RO / 12 / 2609.00659

Fleets Need a Context Plane: Rethinking Cooperative Perception for Autonomous Drones

无人机编队需要一个上下文平面:重新思考自主无人机的协同感知
Liu, Liangkai, Wu, Xiaoxiao
Abstract
Cooperative perception allows a drone fleet to combine observations from multiple viewpoints. However, existing systems typically fix their feature-sharing policies at design time or adapt to only one context signal. This is a poor fit for aerial fleets, whose missions, bandwidth, formation geometry, and scene coverage can change during flight. We quantify the cost of context-blind sharing on UAV3D by controlling feature exchange at evaluation time using a released DiscoNet checkpoint, without retraining. Mission-aware sharing matches full-sharing accuracy while using only 5-10% of the bytes. The best tested peer selection policy changes with the byte budget, and choosing the wrong policy loses up to 7.7 AP. Moreover, under a constrained budget, two policies with the same full-scene accuracy differ by 5.9 AP within the mission region, showing that multiple context axes must be considered jointly. We therefore propose the context plane, a bounded, structured interface for runtime context. Each drone publishes a descriptor of at most 1 KB at 10 Hz, and lightweight, replaceable policies use the fleet context to decide what each drone computes, shares, and fuses. Existing sharing schemes become fixed policies within this interface. In our ROS 2 prototype on a Jetson AGX Orin, the context plane uses approximately 0.01% of the data-plane bandwidth, and each policy decision takes 0.10 ms. These results show that an explicit context interface can support low-overhead runtime adaptation without modifying or retraining the perception model.
Chinese Translation
协同感知使无人机编队能够融合来自多个视角的观测结果。然而,现有系统通常在设计阶段就固定其特征共享策略,或仅适应单一的上下文信号。这对于空中编队来说并不适用,因为其任务、带宽、编队几何构型和场景覆盖都可能在飞行过程中发生变化。我们利用已发布的 DiscoNet 检查点,在评估阶段控制特征交换而无需重新训练,以此量化了在 UAV3D 数据集上忽略上下文的共享方式所付出的代价。任务感知共享仅使用 5-10% 的字节即可达到完全共享的精度。经测试表现最优的邻居选择策略会随字节预算的变化而改变,选择错误的策略会导致最多 7.7 AP 的损失。此外,在受限预算下,两种全场景精度相同的策略在任务区域内相差 5.9 AP,这表明必须联合考虑多个上下文维度。因此,我们提出上下文平面(context plane),这是一个有界的、结构化的运行时上下文接口。每架无人机以 10 Hz 的频率发布不超过 1 KB 的描述符,轻量且可替换的策略利用编队上下文来决定每架无人机计算、共享和融合哪些内容。现有的共享方案在此接口中均可成为固定策略。在基于 Jetson AGX Orin 的 ROS 2 原型系统中,上下文平面仅占用约 0.01% 的数据平面带宽,每次策略决策耗时 0.10 ms。这些结果表明,显式的上下文接口可以在不修改或重新训练感知模型的情况下,支持低开销的运行时自适应。
cs.RO / 13 / 2609.00669

Behavior--Realization Separation for Constrained Physical Human--Robot Interaction

面向受约束物理人机交互的行为—实现分离
Cao, Yongyan
Abstract
Physical human--robot interaction software often couples desired-behavior specification with constrained realization; we treat these as separate layers. A \emph{behavior layer} supplies a desired contact-port acceleration $a_k^{\mathrm{id}}=f_\theta(e_k,\dot e_k,F_{h,k})$. A \emph{realization layer} converts it into constrained robot commands and reports total desired-versus-realized acceleration error instead of hiding it in saturation. A same-objective unconstrained counterfactual separates regularization from constraint intervention, while plant data expose model error. This paper implements a receding-horizon quadratic program realizing memoryless affine behaviors. Changing the behavior modifies objective coefficients through $(C_\theta,G_\theta)$ while the robot-command variable and feasible set remain unchanged. A planar study instantiates impedance and admittance; the same running layer accepts an impedance--admittance--impedance reassignment without reconstruction, under its existing rate limit. On a torque-controlled 7-DOF Franka FR3 in MuJoCo, the runtime freezes task-space dynamics per solve and enforces torque feasibility across its horizon. Under a sustained 20~N push, it holds a slack-relaxed workspace boundary to within approximately 0.1--0.2~mm, versus 4.4~cm (impedance) and 4.7~cm (admittance) overshoot from instantaneous clipping. A derated actuator budget then activates the torque constraint: horizon-wide enforcement keeps its frozen-model plan feasible to $2.1\times10^{-4}$~N$\cdot$m, whereas a first-step-only ablation plans up to 11.329~N$\cdot$m beyond budget; on the executed nonlinear plant, where both share the same local-model error, the gap is smaller but still favors horizon-wide enforcement (0.161 vs.\ 0.380~N$\cdot$m). These results are a focused proof of behavior--realization separation.
Chinese Translation
物理人机交互软件通常将期望行为的设定与受约束的实现耦合在一起;本文将二者视为相互独立的层次。行为层提供期望的接触端口加速度 $a_k^{\mathrm{id}}=f_\theta(e_k,\dot e_k,F_{h,k})$;实现层将其转换为受约束的机器人指令,并报告期望加速度与实际加速度之间的总误差,而非将误差隐藏在饱和处理中。一个具有相同目标函数的无约束反事实分析将正则化与约束干预分离开来,而被控对象数据则用于暴露模型误差。本文实现了一个实现无记忆仿射行为的滚动时域二次规划。改变行为仅通过 $(C_\theta,G_\theta)$ 修改目标函数系数,而机器人指令变量与可行集保持不变。在一个平面算例中实例化了阻抗与导纳控制;同一运行层在无需重构的情况下,在既有的速率限制下接受了阻抗—导纳—阻抗的切换重配置。在MuJoCo中基于力矩控制的7自由度Franka FR3机器人上,运行时每次求解冻结任务空间动力学,并在整个时域内保证力矩的可行性。在持续20~N的推力作用下,系统将带松弛的工作空间边界保持在约0.1–0.2~mm以内,而瞬时截断方法(阻抗)与(导纳)分别产生4.4~cm与4.7~cm的超调。随后,降额的执行器预算激活了力矩约束:全时域约束使其冻结模型的规划保持可行至 $2.1\times10^{-4}$~N$\cdot$m,而仅在首步施加约束的消融实验的规划超出预算达11.329~N$\cdot$m;在执行的非线性被控对象上,二者具有相同的局部模型误差,差距虽然缩小,但仍有利于全时域约束(0.161 对 0.380~N$\cdot$m)。这些结果构成了行为—实现分离的针对性验证。
cs.RO / 14 / 2609.00677

ADAPT: Agile Diffusion Action Priors for Robust and Steerable Online Text-Driven Humanoid Control

ADAPT:面向鲁棒且可控的在线文本驱动人形机器人控制的敏捷扩散动作先验
Wu, Yan, Li, Chenhao, Zhao, Kaifeng, Li, Gen, Hutter, Marco, Tang, Siyu
Abstract
We present ADAPT, an end-to-end framework for interactive, text-conditioned humanoid whole-body control. Unlike dominant text-to-motion pipelines that generate kinematic motions for a separate tracker, ADAPT solves language control with an end-to-end closed-loop control framework, where the robot must continuously respond to changing commands while maintaining balance, natural motion, and smooth transitions. ADAPT learns a diffusion-based action prior from text-labeled humanoid state-action trajectories, enabling diverse motion skills to be directly executed from language commands. To improve long-horizon robustness and smooth prompt switching, we train a lightweight residual reinforcement learning policy on top of the frozen diffusion controller. We further show that the same diffusion policy can be reused as a steerable text-conditioned motion prior for downstream task adaptation. Experiments demonstrate robust language-grounded skill execution, smooth interactive transitions, and style-preserving downstream control.
Chinese Translation
我们提出了ADAPT,一个用于交互式、文本条件人形机器人全身控制的端到端框架。与主流的文本到运动(text-to-motion)流水线(即先生成运动学动作再交由独立的跟踪器执行)不同,ADAPT采用端到端的闭环控制框架来解决语言控制问题,使机器人能够在保持平衡、自然动作以及平滑过渡的同时,持续响应不断变化的指令。ADAPT从带有文本标注的人形机器人状态-动作轨迹中学习基于扩散模型的动作先验(diffusion-based action prior),使多样化的运动技能能够直接通过语言指令执行。为提升长时程鲁棒性和指令切换的平滑性,我们在冻结的扩散控制器之上训练了一个轻量级的残差强化学习策略。我们进一步证明,同一扩散策略可被复用为可引导的文本条件运动先验,用于下游任务适配。实验结果表明,该方法能够实现鲁棒的基于语言的技能执行、平滑的交互式过渡以及保留风格的下游控制。
cs.RO / 15 / 2609.00682

Context-Aware Intelligent Vehicles

情境感知智能汽车
Liu, Liangkai, Shi, Shuyao, Wang, Mingke, Curran, Noah T., Li, Chuan, Bai, Fan, Shin, Kang G.
Abstract
Intelligent vehicles increasingly support adaptive applications beyond driving themselves, ranging from context-aware ADAS and automated driving to in-cabin monitoring and fleet management, all under tight requirements on accuracy, latency, cost, and reliability. Meeting these requirements is challenging because vehicles operate in complex, uncertain, and rapidly changing environments while running on resource-constrained computing platforms. This paper argues that context-situational factors that give meaning to sensor signals and constrain decisions-should be treated as a first-class principle for next-generation vehicle systems, and operationalized as a unified, shared state for learning, risk assessment, and closed-loop control across the software stack. We systematically review state-of-the- art (SOTA) context-aware methods spanning (i) environment understanding, (ii) planning and control, (iii) safety and security, and (iv) connected vehicles. Based on a trend analysis of context-aware design, we identify four key technical challenges in building a general contextual engine for future intelligent vehicles: multi-modal context fusion, temporal context modeling, handling rare events, and collaborative context sharing. We hope this survey will motivate the development of robust and efficient context-aware vehicle applications.
Chinese Translation
智能汽车日益支持超越自动驾驶本身的自适应应用,从情境感知的高级驾驶辅助系统(ADAS)和自动驾驶,到座舱监控和车队管理,所有这些应用都面临准确性、延迟、成本和可靠性方面的严苛要求。满足这些要求极具挑战性,因为车辆在复杂、不确定且快速变化的环境中运行,同时依赖于资源受限的计算平台。本文认为,情境(context)——即为传感器信号赋予意义并约束决策的环境因素——应被视为下一代车辆系统的一等原则,并被操作化为一个统一、共享的状态,用于整个软件栈中的学习、风险评估和闭环控制。我们系统性地回顾了最先进(SOTA)的情境感知方法,涵盖(i)环境理解、(ii)规划与控制、(iii)安全与安保,以及(iv)网联汽车。基于对情境感知设计的趋势分析,我们识别出构建未来智能汽车通用情境引擎的四大关键技术挑战:多模态情境融合、时序情境建模、稀有事件处理以及协同情境共享。我们希望本综述能够推动鲁棒且高效的情境感知车载应用的发展。
cs.RO / 16 / 2609.00751

One Print, Many Moves: Monolithic Origami-inspired Folding Actuator for Composable Soft Multi-DoF Systems

一印多动:用于可组合软体多自由度系统的单片式折纸启发折叠驱动器
Jang, Jaehyung, Zhakypov, Zhenish, Palmer, Jasmin Elena, Klein, Melissa, Ryu, Jee-Hwan, Okamura, Allison Mariko
Abstract
Conventional soft robot actuators excel in compliance, but their uncontrolled deformations compromise accuracy and hinder scaling to multi-degree-of-freedom (DoF) systems. We introduce a MONOlithic ORIGAMI-inspired soft folding actuator design (MONORIGAMI) that establishes a design strategy based on spatially programmed stiffness anisotropy to preserve material compliance along desired folding directions while selectively restricting deformation in unwanted directions. The actuator leverages stiffness tiers based on material thickness, patterned in an origami-inspired geometry with facets and creases, converting unconstrained soft deformation into accurate, repeatable, and composable folding motions without additional reinforcements. The design is fully 3D-printable through a single-material, single-print process that requires no assembly. Each actuator serves as a scalable motion primitive, and linking and orienting multiple actuators mechanically programs multi-DoF trajectories. Using the same fundamental module, we demonstrate three 3D-printed soft multi-DoF robotic systems spanning distinct application domains: (1) a compact 4-DoF wearable haptic device for high-fidelity cutaneous feedback in virtual reality (VR), (2) a 3-DoF joystick for kinesthetic feedback in teleoperation, and (3) a modular robotic gripper capable of underwater operation with geometry-encoded grasp trajectories. These systems demonstrate the module's capabilities for compact multi-axis integration, controlled physical interaction, and geometry-programmed operation across different environments. Together, these results show that MONORIGAMI provides a general, composable, accessible, reliable, and scalable platform for high-precision soft multi-DoF robotics, addressing long-standing limitations in both soft actuator design and fabrication.
Chinese Translation
传统软体机器人驱动器在柔顺性方面表现出色,但其不受控的变形会损害精度,并阻碍其向多自由度(DoF)系统的扩展。我们提出了一种单片式折纸启发软体折叠驱动器设计(MONORIGAMI),建立了一种基于空间编程刚度各向异性的设计策略,即在保持材料沿期望折叠方向的柔顺性的同时,选择性地限制非期望方向的变形。该驱动器利用基于材料厚度的刚度层级,并将其图案化为具有面片和折痕的折纸启发几何结构,从而在无需额外增强结构的情况下,将不受约束的软体变形转化为精确、可重复且可组合的折叠运动。该设计可通过单材料、单次打印的工艺完全3D打印,无需任何组装。每个驱动器都可作为可扩展的运动基元,通过机械连接和定向多个驱动器即可编程出多自由度轨迹。使用相同的基础模块,我们展示了三个3D打印的软体多自由度机器人系统,涵盖不同的应用领域:(1)用于虚拟现实(VR)中高保真皮肤反馈的紧凑型4自由度可穿戴触觉设备;(2)用于遥操作中运动觉反馈的3自由度操纵杆;(3)能够在水下作业并具有几何编码抓取轨迹的模块化机器人夹持器。这些系统展示了该模块在紧凑型多轴集成、受控物理交互以及跨不同环境的几何编程操作方面的能力。这些结果表明,MONORIGAMI为高精度软体多自由度机器人提供了一个通用、可组合、易获取、可靠且可扩展的平台,解决了软体驱动器设计与制造中长期存在的局限。
cs.RO / 17 / 2609.00769

A Compact Robotic Finger with 2-DoF MCP Joint Embedding DoF-Selective Passive Continuously Variable Transmission for Wide Force-Speed Operating Range

一种嵌入自由度选择式被动连续可变传动、具有宽力-速工作范围的二自由度掌指关节紧凑型机器人手指
Jang, JaeHyung, Ryu, Jee-Hwan
Abstract
This letter presents a compact two-degree-of-freedom (DoF) robotic finger with a flexion-selective passive continuously variable transmission (CVT) to achieve a wide force-speed operating range. Inspired by the functional differentiation of the human metacarpophalangeal (MCP) joint, the proposed mechanism realizes DoF-specific transmission differentiation by selectively assigning passive CVT to the flexion-extension DoF while preserving direct transmission for abduction-adduction. For a wide force-speed operating range, a force-responsive passive CVT is embedded in the flexion pathway, while direct transmission is preserved for the abduction-adduction pathway. To selectively realize transmission adaptation within a multi-DoF MCP mechanism, an output-side passive CVT employing a moving-pulley-inspired wire-routing structure is introduced. The resultant force generated by the wire tensions acting on the pulley that passively increases the flexion moment arm and transmission ratio according to the applied load without additional actuators, sensors, or control. Experimental results demonstrate a maximum output-force amplification of 4.19-fold and a mean amplification of 3.6-fold across the tested flexion angles ranging from 15 degrees to 75 degrees through moment-arm adaptation, thereby substantially expanding the achievable force-speed operating range. Furthermore, dexterous ball-rolling experiments verify that passive transmission adaptation can be achieved while preserving abduction-adduction functionality. These results demonstrate a scalable transmission design strategy for compact multi-DoF robotic hands.
Chinese Translation
本信函提出了一种紧凑的二自由度(DoF)机器人手指,其采用屈曲选择式被动连续可变传动(CVT),以实现宽力-速工作范围。受人类掌指关节(MCP)功能分化特性的启发,所提出的机构通过将被动CVT选择性地分配给屈曲-伸展自由度,同时保持外展-内收自由度的直接传动,实现了针对不同自由度的传动分化。为获得宽力-速工作范围,屈曲传动路径中嵌入了一个力响应式被动CVT,而外展-内收路径则保持直接传动。为在多自由度MCP机构中选择性地实现传动自适应,本文引入了一种采用动滑轮式走线结构的输出端被动CVT。线张力作用于滑轮所产生的合力,能够在无额外执行器、传感器或控制的情况下,根据施加的载荷被动增大屈曲力臂和传动比。实验结果表明,通过力臂自适应,在15度至75度的测试屈曲角度范围内,最大输出力放大倍数为4.19倍,平均放大倍数为3.6倍,从而显著扩展了可实现的力-速工作范围。此外,灵巧的滚球实验验证了在保持外展-内收功能的同时可实现被动传动自适应。这些结果展示了一种面向紧凑型多自由度机器手的可扩展传动设计策略。
cs.RO / 18 / 2609.00771

Non-Prehensile Throwing: A Reinforcement Learning Perspective

非抓取式抛掷:强化学习视角
Mustafa, Abdullah, Hanai, Ryo, Ramirez-Alpizar, Ixchel G., Erich, Floris, Nakajo, Ryoichi, Domae, Yukiyasu, Ogata, Tetsuya
Abstract
Robotic throwing enables fast object transport and extends a robot's reachable workspace beyond traditional pick-and-place. While prehensile (grasp-based) throwing works well for graspable items, non-prehensile (grasp-free) throwing is better suited for large, heavy, and/or deformable objects. Existing approaches rely on model-based optimization with simplified contact models (e.g., dynamic grasping) and low-dimensional trajectory parameterizations, which limit solution quality and reachable workspace. We propose a reinforcement learning approach that additionally leverages sliding and rolling contact modes and directly optimizes joint-space trajectories without analytical contact models or custom parameterizations. The Markov Decision Process (MDP) is formulated as a dynamical system that evolves the robot's joint state conditioned on the throwing target, object model, and initial configuration. Joint-jerk trajectories are planned offline at a low control rate and upsampled into smooth, high-rate velocity commands for deployment. For sim-to-real transfer, we minimize the robot-dynamics gap through minimum-jerk system identification and train uncertainty-aware policies to mitigate object-modeling errors, particularly sensitivity to dynamic friction. In simulation, the policy achieves 99% success across thousands of configurations and generalizes to unseen objects. Sensitivity analysis shows robustness to mass uncertainty but high sensitivity to dynamic friction, consistent with the sliding-based release mechanism. Deployed zero-shot on a UR5e operating near its physical limits (5 m/s end-effector velocity), our method throws diverse objects including heavy (790 g) and large (20x20x28 cm) items to targets up to 350 cm distance or 180 cm elevation, achieving a 97% real-world success rate.
Chinese Translation
机器人抛掷能够实现快速的物体搬运,并将机器人的可达工作空间扩展到传统的抓取-放置操作之外。基于抓取(grasp-based)的抛掷对于可抓取的物品效果良好,而非抓取式(grasp-free)抛掷则更适合于大尺寸、重型和/或可变形的物体。现有方法依赖于基于模型的优化,采用简化的接触模型(如动态抓取)和低维轨迹参数化,这限制了求解质量和可达工作空间。我们提出了一种强化学习方法,该方法额外利用了滑动和滚动接触模式,并直接在关节空间中优化轨迹,无需解析接触模型或自定义参数化。马尔可夫决策过程(MDP)被表述为一个动力学系统,该系统以抛掷目标、物体模型和初始构型为条件来演化机器人的关节状态。关节加加速度(jerk)轨迹以较低的控制频率在离线阶段进行规划,并在部署时上采样为平滑的高频率速度指令。对于仿真到现实(sim-to-real)的迁移,我们通过最小加加速度的系统辨识来缩小机器人动力学差距,并训练具有不确定性感知能力的策略以减轻物体建模误差的影响,尤其是对动态摩擦的敏感性。在仿真中,该策略在数千种构型下达到了99%的成功率,并能泛化到未见过的物体。敏感性分析表明,该方法对质量不确定性具有鲁棒性,但对动态摩擦高度敏感,这与基于滑动的释放机制相一致。零样本(zero-shot)部署在接近其物理极限(末端执行器速度5 m/s)的UR5e机械臂上,我们的方法能够抛掷多样化的物体,包括重型(790克)和大尺寸(20x20x28厘米)物品,目标距离最远可达350厘米或高度达180厘米,在真实世界中实现了97%的成功率。
cs.RO / 19 / 2609.00804

Connectivity-Aware Graph Extension for Decentralized Multi-Robot Exploration

面向去中心化多机器人探索的连通性感知图扩展方法
Cegarra, Béatrice Garcia, Vanneaux, Elena, Picard, Quentin, Filliat, David
Abstract
Exploring unknown environments with multiple UAVs requires coordination under intermittent communication, making decentralized operation a baseline assumption. We propose, within a decentralized framework, a novel exploration graph extension strategy based on frontier connectivity to extend exploration plans and maintain area partitioning among agents stable and robust to disconnections and changes in spatial layout. The proposed extension method is applied to two state-of-the-art area partitioning methods and evaluated in simulation. Experiments show improved performance over existing graph extension approaches with higher exploration efficiency under low communication rate.
Chinese Translation
利用多架无人机探索未知环境需要在间歇性通信条件下进行协调,这使得去中心化运行成为一个基本的假设。我们在去中心化框架内提出了一种基于边界(frontier)连通性的新型探索图扩展策略,用于扩展探索计划,并使智能体间的区域划分保持稳定,同时对断连和空间布局变化具有鲁棒性。所提出的扩展方法被应用于两种最先进的区域划分方法,并在仿真中进行了评估。实验表明,与现有的图扩展方法相比,该方法在低通信速率下具有更高的探索效率。
cs.RO / 20 / 2609.00906

Peg-in-Bench: A Modular Benchmark for High-Precision Robotic Insertion

Peg-in-Bench:面向高精度机器人插入任务的模块化基准测试平台
Delgado, Yosel, Buenaventura-Carreón, José G., Erich, Floris, Mykhailyshyn, Roman, Motoda, Tomohiro, Makihara, Koshi, Domae, Yukiyasu
Abstract
High-precision insertion remains a fundamental challenge in robotic manipulation due to the strict alignment requirements and contact-rich interactions involved. Although peg-in-hole tasks are widely used for evaluation, existing bench- marks often rely on fixed task configurations, limiting their ability to assess robustness and generalization across different insertion scenarios. This paper introduces a reconfigurable peg-in-hole benchmark designed to evaluate task generalization in high-precision insertion. The benchmark consists of a set of fully 3D-printable modular components, including multiple peg geometries, tolerance levels, and configurable base structures that can be combined to generate a large variety of insertion and assembly tasks. By varying object layouts, orientations, and task structures while maintaining controlled physical conditions, the benchmark enables systematic evaluation of adaptation to unseen scenarios. To support reproducibility, we additionally provide a scenario generation tool capable of producing standardized task configurations and machine-readable task descriptions. The scenario generation tool and the STL files of the benchmark pieces are available through the project repository: https://github.com/aistairc/peg-in-bench.
Chinese Translation
由于涉及严格的对准要求和丰富的接触交互,高精度插入任务仍然是机器人操作领域的一项根本性挑战。尽管插孔(peg-in-hole)任务被广泛用于评估,但现有的基准测试往往依赖于固定的任务配置,限制了其评估在不同插入场景下的鲁棒性与泛化能力。本文提出了一个可重构的插孔基准测试平台,旨在评估高精度插入任务中的任务泛化能力。该基准测试平台由一组完全可3D打印的模块化组件构成,包括多种插销几何形状、公差等级以及可配置的底座结构,通过组合可以生成大量多样化的插入与装配任务。通过在保持受控物理条件的同时改变物体布局、方向和任务结构,该基准测试平台能够对算法在未见场景下的适应性进行系统性评估。为支持可复现性,我们还提供了一个场景生成工具,可用于生成标准化的任务配置和机器可读的任务描述。场景生成工具以及基准测试零件的STL文件可通过项目仓库获取:https://github.com/aistairc/peg-in-bench。
cs.RO / 21 / 2609.00908

Knowing When to Stop: Adaptive Action Chunking via Internal Cross-Attention Dynamics in VLAs

知道何时停止:基于视觉-语言-动作模型(VLA)内部交叉注意力动态的自适应动作分块
Xu, Runze, Shan, Xiaolong, Dai, Shuang, Wang, Yu, Yu, Jincheng
Abstract
Action chunking is a standard execution strategy in modern Vision-Language-Action (VLA) frameworks, but fixed execution horizons impose a trade-off between efficiency and accuracy. Short chunks require frequent inference and may cause oscillatory behavior, whereas long chunks can become misaligned with newly observed states. We address this limitation with an adaptive action chunking approach based on internal cross-attention dynamics in the action expert. We observe that, as the prediction horizon extends, action-to-observation cross-attention becomes increasingly dispersed and its entropy rises toward a plateau. This pattern is associated with higher action prediction error and provides an online signal that the current observation offers limited grounding for further open-loop execution. Based on this observation, we introduce a training-free truncation mechanism that detects sustained high-entropy plateaus and dynamically selects the execution horizon during inference. The method uses attention weights already computed by the policy and introduces negligible additional overhead. Evaluations on $\pi_{0.5}$ and X-VLA across RoboTwin 2.0, LIBERO, and three real-world manipulation tasks show improved average task success over fixed-horizon and adaptive chunking baselines, while preserving efficient closed-loop control. These results show that cross-attention dynamics can provide a practical internal signal for adaptive action execution in VLAs.
Chinese Translation
动作分块(Action Chunking)是现代视觉-语言-动作(Vision-Language-Action, VLA)框架中的标准执行策略,但固定的执行时域在效率与精度之间存在权衡。较短的分块需要频繁推理,且可能导致振荡行为;而较长的分块则可能与新观测到的状态不一致。为解决这一局限,我们提出了一种基于动作专家(action expert)内部交叉注意力动态的自适应动作分块方法。我们观察到,随着预测时域的延长,动作到观测的交叉注意力变得愈发分散,其熵值上升并趋于平台期。这一模式与更高的动作预测误差相关,并提供了一个在线信号,表明当前观测对进一步的闭环开环执行所提供的支撑已十分有限。基于这一观察,我们引入了一种无需训练的截断机制,该机制检测持续的高熵平台期,并在推理过程中动态选择执行时域。该方法利用策略已经计算得到的注意力权重,带来的额外开销可忽略不计。在 $\pi_{0.5}$ 和 X-VLA 模型上,针对 RoboTwin 2.0、LIBERO 以及三个真实世界操作任务的评估表明,本方法相较于固定时域和自适应分块基线取得了更高的平均任务成功率,同时保持了高效的闭环控制。这些结果表明,交叉注意力动态可以为 VLA 中的自适应动作执行提供一种实用的内部信号。
cs.RO / 22 / 2609.00920

VerNav: Verifier-First Low-Latency Vision-and-Language Navigation

VerNav:验证器优先的低延迟视觉语言导航
Wang, Zhixin, Yao, Chengzheyi, Liu, Leyuan, Zhang, Xiaosong, Zhang, Yongzhao
Abstract
Vision-and-Language Navigation (VLN) requires an agent to navigate through unseen 3D environments according to natural-language instructions. Explicit reasoning can improve instruction understanding and semantic grounding, but autoregressive generation at every step accumulates large decision-stage latency over multi-step navigation. We propose VerNav, a verifier-first framework for low-latency LLM-based VLN. The verifier reduces decision-stage latency by replacing per-step autoregressive generation with batched action verification, while an entropy-based adaptive generator is invoked only for uncertain decisions to produce compact state evidence. To further improve navigation performance with the verifier, we introduce a two-stage alignment scheme: (i) VPO improves local action-preference alignment in static verifier training, and (ii) step-level reinforcement fine-tuning provides dense progress rewards over multi-step navigation rollouts during dynamic task execution. Experiments on the Room-to-Room (R2R) benchmark show that the verifier-only decision path of VerNav achieves competitive navigation performance among representative LLM-based VLN agents while reducing average decision-stage LLM latency per step by more than $10\times$ compared with autoregressive methods.
Chinese Translation
视觉语言导航(Vision-and-Language Navigation, VLN)要求智能体根据自然语言指令在未见过的3D环境中进行导航。显式推理可以改善指令理解和语义对齐,但每一步的自回归生成会在多步导航中累积大量的决策阶段延迟。我们提出了VerNav,一个验证器优先的低延迟基于LLM的VLN框架。该验证器通过批量化动作验证替代逐步自回归生成,从而降低决策阶段延迟;同时,基于熵的自适应生成器仅在决策不确定时被调用,以生成紧凑的状态证据。为了进一步提升验证器的导航性能,我们引入了两阶段对齐方案:(i)VPO在静态验证器训练中改进局部动作偏好对齐;(ii)步级强化微调在动态任务执行期间为多步导航rollout提供密集的进度奖励。在Room-to-Room(R2R)基准上的实验表明,VerNav仅使用验证器的决策路径在代表性基于LLM的VLN智能体中取得了具有竞争力的导航性能,同时与自回归方法相比,将每步平均决策阶段LLM延迟降低了10倍以上。
cs.RO / 23 / 2609.00941

ProxPI: Proximal Prior Injection for Sampling-Based MPC under Learned-Prior Mismatch

ProxPI:面向基于采样的模型预测控制且存在学习先验失配问题的近端先验注入方法
Im, Euncheol, Lim, Myotaeg, Lee, Yisoo
Abstract
Combining learned policies with model predictive control can leverage learned task priors while retaining online adaptation to new objectives and constraints, but performance degrades when the policy is out of distribution. In policy-guided model predictive path integral (MPPI) control, a policy-centered warm-start approach centers the sampling distribution on the policy output. When the prior is mismatched, centering the sampling distribution on the policy output restricts exploration around an unsuitable solution and prevents recovery toward the task optimum. We propose Proximal Prior Injection (ProxPI), which retains nominal-centered MPPI sampling and incorporates the policy through a soft proximity cost. This matches the in-distribution performance of existing prior-injection schemes while enabling the optimizer to escape an inaccurate policy and recover vanilla MPPI-level performance. We theoretically show that re-centering on the prior discards the optimizer's correction at every update, whereas nominal-centered sampling retains it and converges to a solution set by both the task cost and the prior, and that this failure is not removed by a larger rollout budget. Simulations and real-robot experiments demonstrate robust performance under both in-distribution and out-of-distribution tasks.
Chinese Translation
将学习到的策略与模型预测控制相结合,既能利用学习到的任务先验,又能保留对新目标和约束的在线适应能力,但当策略处于分布外时,性能会下降。在策略引导的模型预测路径积分(MPPI)控制中,以策略为中心的热启动方法将采样分布的中心设在策略输出上。当先验失配时,将采样分布中心设在策略输出上会限制在不合适解周围的探索,并阻碍向任务最优解的恢复。我们提出近端先验注入方法(Proximal Prior Injection, ProxPI),该方法保留以标称解为中心的MPPI采样,并通过软邻近代价引入策略。该方法在分布内可达到与现有先验注入方案相同的性能,同时使优化器能够摆脱不准确的策略并恢复到与原始MPPI相当的性能水平。我们从理论上证明:以先验为中心会在每次更新时丢弃优化器的修正,而以标称解为中心的采样则保留该修正,并收敛到由任务代价和先验共同决定的解集;且这种失败无法通过增大滚动预算来消除。仿真和真实机器人实验证明了该方法在分布内与分布外任务下的鲁棒性能。
cs.RO / 24 / 2609.00950

HitMem: Hierarchical Temporal 3D Memory with Multi-Modal Context-Aware Retrieval for Dynamic Environments

HitMem:面向动态环境的具有多模态上下文感知检索的分层时序三维记忆框架
Tang, Ruijie, Zou, Chenye, Wu, Guoquan, Wei, Jun, Chen, Wei, Zhu, Jiaxin
Abstract
Executing long-term tasks in dynamic environments requires embodied agents to maintain robust and adaptive 3D scene representations. However, most existing 3D memory frameworks rely on static world assumptions. When objects are displaced by human activities or unobserved events, agents encounter memory-observation conflicts and often require costly geometric recomputations or inefficient global re-exploration. To address this, we propose HitMem, a hierarchical temporal 3D memory framework with a multi-modal context-aware retrieval mechanism. Through continuous perception, HitMem unifies semantic and spatial information into a lightweight topological graph that captures support relationships, while a temporal decay mechanism dynamically regulates memory activeness to mitigate the impact of stale representations. In addition, the multi-modal context-aware retrieval mechanism defaults to filtering candidates using integrated semantic, spatial, and temporal memory features, and activates a specialized two-stage retrieval process when object displacement is detected. This process combines spatial constraints inferred from external agent trajectories with semantic common sense grounded in class affinities, efficiently identifying high-probability candidate regions. Extensive evaluations on our constructed Dyna-THOR benchmark demonstrate that HitMem significantly improves object relocation accuracy, reduces exploration costs, and enhances task execution performance in dynamic environments.
Chinese Translation
在动态环境中执行长期任务要求具身智能体维持鲁棒且自适应的三维场景表示。然而,现有的大多数三维记忆框架均依赖于静态世界假设。当物体因人类活动或未被观测到的事件而发生位移时,智能体会遭遇记忆与观测之间的冲突,并往往需要代价高昂的几何重计算或低效的全局重新探索。为解决这一问题,我们提出了HitMem,一种具有多模态上下文感知检索机制的分层时序三维记忆框架。通过持续感知,HitMem将语义与空间信息统一到一个捕获支撑关系的轻量级拓扑图中,同时利用时序衰减机制动态调节记忆的活跃度,以减轻陈旧表示的影响。此外,多模态上下文感知检索机制默认使用融合的语义、空间与时序记忆特征对候选进行筛选,并在检测到物体位移时激活专门的两阶段检索过程。该过程结合从外部智能体轨迹推断出的空间约束与基于类别亲和性的语义常识,高效识别高概率候选区域。在我们构建的Dyna-THOR基准上的大量评估表明,HitMem显著提升了物体重定位精度,降低了探索成本,并增强了动态环境下的任务执行性能。
cs.RO / 25 / 2609.01061

Accelerating Reinforcement Learning via MPC Solver-Gradient Guidance for Weights-varying MPC

通过MPC求解器梯度引导加速权重可变MPC的强化学习
Zarrouki, Baha, Thobani, Arslan, Hoffmann, Jasper, Piccinini, Mattia, Reiter, Rudolf, Jahncke, Felix, Gros, Sébastien, Scaramuzza, Davide, Betz, Johannes
Abstract
In Model Predictive Control (MPC), cost-function weights shape closed-loop behavior, yet changing conditions often make fixed parametrizations suboptimal and motivate context-dependent online adaptation. Learning such policies is difficult because behavior depends implicitly on numerical MPC solutions, producing nonlinear, potentially nonsmooth, long-horizon dependencies on policy parameters. This creates a bias-variance tradeoff: Reinforcement Learning (RL) optimizes realized closed-loop return from environment samples but is sample-inefficient, whereas Gradient-Based Policy Learning (GB-PL) uses low-variance solver gradients from differentiable MPC to optimize surrogate losses on predicted trajectories but can be biased under model mismatch. We propose Solver-Gradient Guided Reinforcement Learning (SG-RL), a solver-sensitivity augmentation for RL-based online MPC cost-weight adaptation. SG-RL keeps sampled closed-loop return as the objective and uses bounded solver-derived gradients as auxiliary guidance to improve stability and sample efficiency. We instantiate SG-RL in Proximal Policy Optimization (PPO) with four modular algorithms that inject solver-gradient guidance into actor-update scaling, policy loss, advantage estimation, and value-function learning. On two full-scale autonomous racing platforms with intentional model mismatch, SG-RL reaches PPO's best closed-loop return with up to 70.6% fewer samples, outperforms GB-PL baselines by at least 54% in closed-loop return, and generalizes zero-shot to unseen environments.
Chinese Translation
在模型预测控制(MPC)中,代价函数权重决定了闭环行为,然而环境条件的变化往往使固定参数化变得次优,从而促使人们进行依赖于上下文的在线自适应调整。学习此类策略十分困难,因为系统行为隐式依赖于数值MPC求解结果,导致策略参数之间存在非线性、可能非光滑的长时程依赖关系。这带来了偏差-方差权衡:强化学习(RL)基于环境样本优化实际实现的闭环回报,但样本效率低下;而基于梯度的策略学习(GB-PL)利用可微MPC提供的低方差求解器梯度来优化基于预测轨迹的代理损失,但在模型失配时可能产生偏差。我们提出了求解器梯度引导强化学习(Solver-Gradient Guided Reinforcement Learning, SG-RL),这是一种面向基于RL的在线MPC代价权重自适应的求解器敏感性增强方法。SG-RL保持采样闭环回报作为优化目标,并将有界的求解器导出梯度作为辅助引导,以提升稳定性和样本效率。我们在近端策略优化(Proximal Policy Optimization, PPO)中实例化了SG-RL,设计了四种模块化算法,将求解器梯度引导分别注入演员更新缩放、策略损失、优势估计和值函数学习之中。在两个存在有意模型失配的全尺寸自主赛车平台上,SG-RL以最多减少70.6%的样本量达到PPO的最优闭环回报,闭环回报较GB-PL基线至少提升54%,并能零样本泛化至未见过的环境。
cs.RO / 26 / 2609.01089

Adaptive Depth-Map-Guided Bundle Adjustment for Correspondence-Free Multi-View Point Cloud Registration

用于无对应关系多视角点云配准的自适应深度图引导光束法平差
Zhou, Yiran, Wang, Yingyu, Huang, Shoudong, Zhao, Liang
Abstract
Robotic processing of irregular steel scrap requires dense 3-D measurement to replace manual visual assessment in hazardous cutting workcells. The reconstructed map is used to estimate piece dimensions, boundary geometry, feasible preheating and cutting regions, and collision-aware torch paths. The reconstruction errors therefore propagate directly to downstream measurement and planning. Existing multi-view registration methods commonly rely on feature extraction and data association to establish correspondences between views. In workcells with smooth metallic surfaces, repeated structures, occlusions, and partial overlaps, however, wrong correspondences may be established, leading to inaccurate pose estimation and distorted reconstruction. This paper presents an adaptive layered depth-map-guided bundle adjustment framework for correspondence-free multi-view point cloud registration. The scene is represented by a global 2.5-D grid, where each cell can adaptively maintain multiple depth hypotheses. Raw depth observations are directly projected into the global map to form depth constraints without explicit feature correspondences. At grid cells where multiple surfaces produce conflicting depths, a softmax-based layer assignment links each observation to compatible depth hypotheses. The resulting nonlinear least-squares formulation jointly refines sensor poses and the layered depth map, with correspondences implicitly induced by the depth-map representation and projection model. Experiments on self-collected industrial datasets show that the proposed method achieves consistently competitive reconstruction accuracy while maintaining robustness and low computational cost in challenging industrial scenarios. We release the open-source code implementation at: https://github.com/YiranZhou-Robotics/ADM-BA.git
Chinese Translation
对不规则废钢的机器人化处理需要密集的三维测量,以替代危险切割工位中的人工视觉评估。重建的地图用于估算工件尺寸、边界几何形状、可行的预热与切割区域,以及具备碰撞感知能力的焊炬路径。因此,重建误差会直接传播到下游的测量与规划任务中。现有的多视角配准方法通常依赖特征提取和数据关联来建立视角间的对应关系。然而,在具有光滑金属表面、重复结构、遮挡和部分重叠的工位环境中,可能建立错误的对应关系,导致位姿估计不准确和重建结果失真。本文提出了一种自适应分层深度图引导的光束法平差(Bundle Adjustment)框架,用于无对应关系的多视角点云配准。场景由一个全局2.5-D栅格表示,其中每个栅格单元可以自适应地维护多个深度假设。原始深度观测被直接投影到全局地图中以形成深度约束,无需显式的特征对应关系。在多个表面产生冲突深度的栅格单元处,基于softmax的层分配机制将每个观测关联到兼容的深度假设。由此得到的非线性最小二乘公式联合优化传感器位姿与分层深度图,对应关系由深度图表示和投影模型隐式诱导。在自采集的工业数据集上的实验表明,所提方法在具有挑战性的工业场景中保持鲁棒性和较低计算成本的同时,获得了持续具有竞争力的重建精度。我们在以下网址开源了代码实现:https://github.com/YiranZhou-Robotics/ADM-BA.git
cs.RO / 27 / 2609.01120

DNC-IMM: Early Lane-Change Intention Recognition via Neural Calibration Based on Driving Context Information

DNC-IMM:基于驾驶上下文信息的神经校准换道意图早期识别方法
Byun, Woong-Chan, Kong, Seung-Hyun
Abstract
Early recognition of lane-change intention is essential for proactive decision-making in autonomous driving and advanced driver assistance systems. This paper proposes a Dual Neural-Calibrated Interacting Multiple Model (DNC-IMM) that improves adaptability to driving context while preserving the probabilistic structure and interpretability of a conventional IMM. The proposed method encodes driving-context information, including target-vehicle motion, gaps to surrounding vehicles, and relative velocities, with a neural network that calibrates both the transition-probability matrix and measurement likelihoods. The final intention is determined from the calibrated IMM mode posterior rather than from a separate direct classifier. Experiments on the highD dataset demonstrate that the proposed method reliably recognizes lane-change intentions before lane crossing and provides particularly strong performance at the earlier 2-3 s prediction horizons.
Chinese Translation
换道意图的早期识别对于自动驾驶和高级驾驶辅助系统中的主动决策至关重要。本文提出了一种双神经校准交互多模型(Dual Neural-Calibrated Interacting Multiple Model, DNC-IMM)方法,该方法在保留传统IMM概率结构和可解释性的同时,提高了对驾驶上下文的适应性。所提方法利用神经网络对驾驶上下文信息进行编码,包括目标车辆运动、与周围车辆的间距以及相对速度,并对转移概率矩阵和测量似然进行校准。最终意图由校准后的IMM模型后验概率确定,而非通过单独的直接分类器得到。在高德highD数据集上的实验表明,所提方法能够在车辆跨越车道线之前可靠地识别换道意图,并且在更早的2-3秒预测时域内表现出尤为突出的性能。
cs.RO / 28 / 2609.01207

On Global Regulatability of Robot Manipulators by Classical PID

论经典PID控制对机器人机械臂的全局可调节性
Zhao, Cheng, Zhu, Jingru, Guo, Lei
Abstract
A long-standing open problem in robot manipulator control is whether global regulation can be achieved by classical PID control. This paper provides an answer to this question for classical PID controllers with triple parameters (k_p,k_i,k_d) in R^3. We find and prove that for one-degree-of-freedom manipulators, the classical PID control guarantees global stability and asymptotic regulation under standard structural assumptions, and further derive explicit quantitative design conditions for the PID gains. However, for multi-degree-of-freedom cases, we can construct a robot manipulator satisfying the same structural assumptions for which no choice of PID gains (k_p,k_i,k_d) can achieve global asymptotic regulation. These results provide a fundamental understanding of the abovementioned open problem, revealing both the fundamental capability and intrinsic limitation of the classical PID control for robot manipulator dynamics.
Chinese Translation
机器人机械臂控制领域一个长期悬而未决的问题是:经典PID控制能否实现全局调节。本文针对参数为三元组 (k_p, k_i, k_d) ∈ R^3 的经典PID控制器回答了这一问题。我们发现并证明:对于单自由度机械臂,在标准的结构假设下,经典PID控制可保证全局稳定性与渐近调节,并进一步推导出PID增益的显式定量设计条件。然而,对于多自由度情形,我们可以构造出一个满足相同结构假设的机器人机械臂,使得任何PID增益 (k_p, k_i, k_d) 的选择都无法实现全局渐近调节。这些结果为上述开放问题提供了根本性的理解,既揭示了经典PID控制对机器人机械臂动力学的基本能力,也揭示了其内在局限。
cs.RO / 29 / 2609.01281

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

EmbodiedSkills:用于编排、训练与部署VLA智能体的统一框架
Wang, Wei, Zhang, Wenqiao, Lin, Yutong, Yuan, Yuqian, Lin, Tianwei, Mao, Jinhao, Fan, Zhenxuan, Gao, Mingjian, Dai, Yang, Li, Wentong, Lv, Zheqi, Dong, Zheng, Niu, Yingjie, Zhu, Jiaqi, Xiao, Jun, Li, Chao, Zhuang, Yueting
Abstract
Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but long-horizon tasks require more than action prediction. An agent must coordinate perception, planning, execution, progress verification, and recovery as the physical state evolves. An action prediction or a model-generated skill decision does not, by itself, guarantee that the proposed operation is valid in the current state or that its outcome will be verified. We propose EmbodiedSkills, a unified framework that treats each skill decision as an execution proposal: the runtime checks its prerequisites before execution and verifies the outcome afterward. A shared executable-skill interface connects high-level skill selection, bounded low-level VLA execution, and post-action verification within a single agent loop. Because this interface remains fixed, low-level VLA policies can be replaced or adapted without changing the agent loop. The interface also records planning, execution, verification, and recovery events as structured trajectories, which provide supervision for individual components and can support optional online adaptation when interactive feedback is available. We instantiate EmbodiedSkills with Qwen3-VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO. Task-adapted low-level VLA policies achieve an average success rate of 86.20% across 50 RoboTwin 2.0 tasks and 97.40% across the four LIBERO suites. These results establish the execution performance of the task-adapted low-level VLA policies used in EmbodiedSkills. On four memory-dependent RMBench tasks, the same task-adapted execution approach achieves 12.5% average success. The framework provides a trainable and inspectable agent layer for turning these policies into closed-loop embodied systems.
Chinese Translation
视觉-语言-动作(VLA)模型将视觉观测和语言指令直接映射为机器人动作,但长时程任务所需的不仅仅是动作预测。随着物理状态的演化,智能体必须协调感知、规划、执行、进度验证与恢复。动作预测或模型生成的技能决策本身并不能保证所提出的操作在当前状态下有效,也不能保证其结果会被验证。我们提出EmbodiedSkills,这是一个统一框架,它将每个技能决策视为一个执行提议:运行时在执行前检查其前置条件,并在执行后验证结果。一个共享的可执行技能接口将高层技能选择、有界的低层VLA执行以及动作后验证连接在一个智能体循环中。由于该接口保持固定,低层VLA策略可以在不改变智能体循环的情况下被替换或适配。该接口还将规划、执行、验证和恢复事件记录为结构化轨迹,为各个组件提供监督信号,并在有交互反馈可用时支持可选的在线适配。我们在RoboTwin 2.0和LIBERO上使用Qwen3-VL和OpenPI/pi0.5实例化了EmbodiedSkills。任务适配的低层VLA策略在RoboTwin 2.0的50个任务上取得了86.20%的平均成功率,在LIBERO的四个测试套件上取得了97.40%的平均成功率。这些结果确立了EmbodiedSkills中所使用的任务适配低层VLA策略的执行性能。在四个依赖记忆的RMBench任务上,同样的任务适配执行方法取得了12.5%的平均成功率。该框架提供了一个可训练、可检查的智能体层,用于将这些策略转化为闭环具身系统。
cs.RO / 30 / 2609.01339

Integrating Traffic Noise Emission Modelling into Variable Speed Limit Control

将交通噪声排放模型融入可变限速控制
Meng, Jiawen, Arockiasamy, John Pravin, Vinel, Alexey
Abstract
Road traffic noise remains a major environmental challenge, yet most speed management strategies are static and do not respond to short-term variations in traffic noise emissions. Although variable speed limit (VSL) systems are widely deployed for safety and congestion mitigation, traffic noise is rarely treated as an explicit operational control objective. This paper proposes a noise-aware VSL framework that integrates aggregated traffic-state estimation with a simplified CNOSSOS-EU-based emission indicator. A stage-based controller with time-varying reference thresholds dynamically adjusts discrete speed-limit levels in response to estimated emission conditions. The framework is evaluated using microscopic traffic simulation calibrated with empirical motorway data and replicated across multiple stochastic realisations. Over a 24-hour evaluation period, the adaptive strategy reduces the receiver-based equivalent sound level by 2.9 dB(A) relative to unrestricted traffic conditions, while maintaining an average vehicle speed approximately 11.3 km/h higher than a permanently imposed low-speed regime. Period-wise analysis shows that speed reductions are activated selectively when emission levels approach calibrated targets, rather than enforcing a constant intermediate limit. Traffic stability indicators reveal moderate increases in speed variability compared with unrestricted operation, but substantially lower braking intensity than under uniform low-speed enforcement. These results demonstrate the feasibility of integrating environmental performance indicators into operational speed control, providing a practical complement to conventional infrastructure-based noise mitigation measures.
Chinese Translation
道路交通噪声仍然是一个重大的环境挑战,然而大多数速度管理策略都是静态的,无法响应交通噪声排放的短期变化。尽管可变限速(VSL)系统被广泛部署以提升安全性和缓解拥堵,但交通噪声很少被作为明确的运行控制目标。本文提出了一种噪声感知的VSL框架,该框架将聚合的交通状态估计与基于CNOSSOS-EU的简化排放指标相结合。一个基于阶段、具有时变参考阈值的控制器根据估计的排放条件动态调整离散的限速等级。该框架在采用实测高速公路数据校准的微观交通仿真中进行评估,并在多个随机情景中重复运行。在24小时的评估期内,与不受限制的交通状况相比,自适应策略将基于接收端的等效声级降低了2.9 dB(A),同时保持平均车速比长期实施低速方案高出约11.3 km/h。分时段分析表明,当排放水平接近校准目标时,降速措施被有选择地激活,而不是强制执行恒定的中间限速。交通稳定性指标显示,与不受限制的运行相比,速度变异性有中等程度的增加,但制动强度显著低于统一低速强制执行的情况。这些结果表明,将环境性能指标融入运行速度控制是可行的,为传统的基于基础设施的噪声缓解措施提供了实用的补充。
cs.RO / 31 / 2609.01351

Scalable Rao-Blackwellized Online Planning for High-Dimensional POMDPs

面向高维POMDP的可扩展Rao-Blackwellized在线规划方法
Lee, Jiho, Ahmed, Nisar, Wray, Kyle Hollins, Sunberg, Zachary
Abstract
Online planning under uncertainty remains a fundamental challenge for robotic systems operating in partially observable environments with high-dimensional state spaces. While sampling-based POMDP solvers enable approximate decision-making in large or continuous domains, their performance degrades as belief dimensionality increases due to the high variance inherent in Monte Carlo-based estimation. In this work, we extend the Rao-Blackwellized online POMDP (RB-POMDP) framework to improve its generalizability in high-dimensional settings through hybrid continuous-discrete belief representations. By analytically propagating uncertainty associated with marginalized state components during tree-based planning, the proposed approach reduces sampling-induced variance in value estimation. We demonstrate the effectiveness of this framework in a robotic search-and-rescue task by integrating it with FastSLAM 2.0. Experimental results show that the proposed planner achieves higher cumulative rewards using significantly fewer particles and planning simulations than purely sampling-based methods under equivalent computational budgets. These results suggest that structured high-dimensional robotic problems admitting tractable sufficient statistics can be effectively leveraged within the RB-POMDP framework for computationally feasible online decision-making.
Chinese Translation
在部分可观测且状态空间维度较高的环境中,不确定性下的在线规划对于机器人系统而言仍然是一项根本性挑战。尽管基于采样的POMDP求解器能够在大规模或连续域中实现近似决策,但由于基于蒙特卡洛的估计固有地存在高方差,其性能会随着信念维度(belief dimensionality)的增加而下降。在本工作中,我们扩展了Rao-Blackwellized在线POMDP(RB-POMDP)框架,通过混合连续-离散信念表示提升其在高维场景中的泛化能力。通过在基于树的规划过程中解析地传播与边缘化状态分量相关的不确定性,所提出的方法降低了采样引起的价值估计方差。我们通过将该方法与FastSLAM 2.0相结合,在机器人搜救任务中验证了该框架的有效性。实验结果表明,在相同计算预算下,与纯基于采样的方法相比,所提出的规划器利用显著更少的粒子和规划仿真次数即可获得更高的累积奖励。这些结果表明,对于具有可处理充分统计量的结构化高维机器人问题,可以在RB-POMDP框架内对其加以有效利用,从而实现计算上可行的在线决策。
cs.RO / 32 / 2609.01384

Obstacle-Aware Autonomous Coverage and Navigation for Outdoor Robots

面向户外机器人的障碍感知自主覆盖与导航
Gargani, Leonardo, Frosi, Matteo, Matteucci, Matteo
Abstract
Long-duration outdoor coverage with autonomous platforms remains challenging beyond classical planning: deployments face localization drift in open spaces, obstacles in cluttered sites, controller feasibility in turn-heavy maneuvers, and persistent autonomy with energy management. We propose a unified ROS 2 architecture for outdoor coverage that combines coverage planning, robust localization, and Nav2-based execution. A dual-antenna RTK-GNSS fused in an EKF keeps the robot pose, both position and heading, accurate across long missions; three controller-aware refinements are added to a mature coverage planner; a Behavior-Tree mission manager coordinates multi-goal execution, layered recovery, cost-aware goal management, and autonomous docking for return-to-charge. We validate the stack through simulation and real-world trials across multiple outdoor areas with varying geometries and obstacle densities. Overall, these results show that the proposed stack can reliably complete outdoor coverage missions across varied areas, sweeping 93.1% to 96.1% of the planned coverage area.
Chinese Translation
自主平台的长时间户外覆盖作业在经典规划方法之外仍面临诸多挑战:部署过程中会遇到开阔空间中的定位漂移、杂乱场地中的障碍物、频繁转弯机动下的控制器可行性问题,以及结合能量管理的持久自主性需求。我们提出了一种统一的 ROS 2 户外覆盖架构,将覆盖规划、鲁棒定位和基于 Nav2 的执行相结合。双天线 RTK-GNSS 经 EKF(扩展卡尔曼滤波)融合,可在长任务中保持机器人位姿(包括位置与航向)的精确性;在成熟的覆盖规划器基础上增加了三项面向控制器的改进;行为树任务管理器协调多目标执行、分层恢复、代价感知的目标管理,以及自主对接返回充电。我们通过仿真以及在多个具有不同几何形状和障碍物密度的户外区域的实际试验对该系统栈进行了验证。总体而言,结果表明所提出的系统栈能够可靠地完成跨不同区域的户外覆盖任务,覆盖了规划覆盖面积的 93.1% 至 96.1%。
cs.RO / 33 / 2609.01394

Autonomous robotic bridging using distributed swarm control without inter-agent communication

基于无智能体间通信的分布式群体控制的自主机器人桥梁架设
Thamaraiselvan, Vishwaak C., Lundberg, Cody L., Theofandis, Michail, Chelian, Suhas, Gans, Nicholas R.
Abstract
We describe SCARAB--Swarm-Capable Autonomous Robotic Aquatic Bridging. Using distributed swarm control and multi-model sensing of agents and docking targets, agents can localize themselves, join into formations and proceed to desired target locations. Our technologies would eventually allow the Army to perform unpredictable, dispersed river crossings, enhance crew survivability, and minimize the logistics footprint compared to the current Improved Ribbon Bridge. Our methods operate without GPS or RF communications, though these can be used in non-contested environments (e.g., civilian disaster relief for flooding, etc.). We demonstrate our system via physics-engine-based simulation of several agents using unmanned surface vehicles (USVs) in the presence of currents and wind.
Chinese Translation
我们介绍了SCARAB(Swarm-Capable Autonomous Robotic Aquatic Bridging,具备群体能力的自主机器人水上桥梁架设)系统。借助分布式群体控制以及对智能体和对接目标的多模型感知,各智能体能够实现自身定位、编队集结并前往目标位置。与现有的改进型带式浮桥(Improved Ribbon Bridge)相比,我们的技术最终将使陆军能够执行不可预知的分散式渡河行动,提升乘员的生存能力,并最大限度地减少后勤保障负担。我们的方法无需GPS或射频(RF)通信即可运行,但在非对抗环境中(例如面向洪灾的民用救灾等)也可使用这些手段。我们通过基于物理引擎的仿真演示了该系统,模拟了多个使用无人水面艇(USV)的智能体在水流和风作用下的运行情况。
cs.RO / 34 / 2609.01404

Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching

评估多模态大语言模型作为无人机控制的通用视觉-语言-动作智能体:指挥、接近、跟踪与搜索
Park, Jaewoo, Lee, Minyoung, Seo, Sukmin, Yim, Moonbin, Yoon, Hyunwook, Ryu, Dohoon, Kim, Daehee, Song, Myungseo, Byun, Jihyuk, Chang, Seunggyu, Kil, Taeho, Kim, Jiseob, Lee, Bado, Kim, Geewook
Abstract
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends into acting: dropping an MLLM directly into a drone's control loop, with its entire action space declared solely in the prompt. Recent systems approach this setting but increasingly narrow the model's decision-making. We widen it back. We introduce DroneCATS-Agent, an architecture where the MLLM is a swappable component, and DroneCATS, a benchmark treating the model as the independent variable. Beyond merely flying toward a pixel, our agent entrusts the model to yaw and search, deliberate when unsure, and self-declare arrival---all without fine-tuning or function-calling schemas. Evaluating frontier and open models across four core capabilities---approaching a visible target, tracking a moving one, searching outside the initial view, and commanding a multi-drone fleet---reveals that even the simplest embodied settings are far from solved. Crucially, to identify what breaks first at the edge, our roster scales down to 2B parameters. The findings expose a stark paradox: it is not the flying that fails. Small open models often navigate into the success radius more reliably than frontier models, yet lose the episode by declaring arrival prematurely or not at all. Multi-drone commanding amplifies this divide, with small models failing by blindly copying a single coordinate across distinct views. Viewed as vision-language-action agents, the models' spatial perception holds up, but their action protocol does not. What separates a deployable edge model from a frontier model is not navigation, but the discipline to sustain a declared protocol and emit the correct terminating action. The open problem is closing this gap at onboard compute costs---yielding a fast model that plans persistently and knows exactly when it is done---and DroneCATS is built to measure that distance.
Chinese Translation
多模态大语言模型(MLLMs)是强大的图像与视频感知器。我们探究这种能力在行动方面能延伸多远:将一个MLLM直接放入无人机的控制回路中,其完整动作空间仅在提示词中声明。近期的一些系统尝试了这一设定,但日益收窄了模型的决策空间。我们则将其重新拓宽。我们提出了DroneCATS-Agent,一种MLLM可作为可替换组件的架构,以及DroneCATS,一个将模型作为自变量的基准。我们的智能体不仅仅是飞向某个像素,还将偏航搜索、在不确定时进行审慎思考、以及自我宣告到达——这一切都无需微调或函数调用模式。通过在四项核心能力上评估前沿模型和开源模型——接近可见目标、跟踪移动目标、搜索初始视野之外的目标、以及指挥多无人机编队——我们发现即使是最简单的具身设定也远未解决。至关重要的是,为了找出在能力边缘最先失效的环节,我们的模型阵容下探至20亿(2B)参数。研究发现揭示了一个鲜明的悖论:失败的不是飞行本身。小型开源模型往往比前沿模型更可靠地导航进入成功半径,却因过早或从未宣告到达而输掉整个任务。多无人机指挥进一步放大了这一差距,小型模型因在不同视野间盲目复制同一坐标而失败。作为视觉-语言-动作智能体来看,这些模型的空间感知能力尚可,但其行动协议却不尽如人意。将一个可部署的边缘模型与前沿模型区分开来的,不是导航能力,而是持续维持既定协议并发出正确终止动作的纪律性。悬而未决的问题是如何在机载计算成本下弥合这一差距——即得到一个能持续规划并准确知道自己何时完成任务的快速模型——而DroneCATS正是为度量这一距离而构建的。
cs.RO / 35 / 2609.01420

Vision-Based Leader-Follower Formation Control for Cooperative UAVs in GPS-Degraded Environments

GPS退化环境下基于视觉的无人机编队领航-跟随协同控制
Angadi, Deekshitha, Budda, Naveena, Agarwal, Vikas, Mulasa, Rojesh Arunkumar, Killamsetty, Ravi, Samshad, Mohamed, Kemsaram, Narsimlu
Abstract
Cooperation in multi-UAV systems requires reliable relative perception so that follower vehicles can maintain formation and continue their mission safely even when absolute positioning sensors degrade or fail. This paper presents a vision-based cooperative formation framework running on a follower UAV that uses a front-facing RGB-D camera to detect, track, and localize a leader UAV in real-time. A lightweight YOLO-based detector is trained on a dedicated drone dataset and deployed onboard to predict leader bounding boxes, which are then fused with depth information via a pinhole camera model to estimate the leader's relative pose. These estimates provide a leader-follower position controller and can also be used as a backup when GPS or external localization is unavailable. This framework is implemented as a set of ROS nodes and evaluated in a physics-based multi-UAV simulation built on XTDrone, with sensor noise and communication dropouts. We evaluate detection accuracy, runtime, and formation-keeping error under nominal conditions and under simulated failures of the positioning sensors. The results show that the proposed framework maintains stable leader-follower formations with reasonable computational cost and provides a practical basis for extending vision-based cooperative formation control to real-world multi-UAV systems.
Chinese Translation
多无人机系统的协同需要可靠的相对感知能力,以便在绝对定位传感器性能退化或失效时,跟随无人机仍能保持编队并安全地继续执行任务。本文提出了一种基于视觉的协同编队框架,运行于跟随无人机上,利用前置RGB-D相机实时检测、跟踪并定位领航无人机。该框架在专用无人机数据集上训练了一个轻量级的基于YOLO的检测器,并将其部署在机载平台上以预测领航无人机的边界框,随后通过针孔相机模型将检测结果与深度信息融合,估计领航无人机的相对位姿。这些估计值为领航-跟随位置控制器提供输入,也可在GPS或外部定位不可用时作为备用方案。该框架以一组ROS节点实现,并在基于物理的XTDrone多无人机仿真环境中进行了评估,仿真中加入了传感器噪声和通信中断。我们在标称条件以及定位传感器模拟故障情况下,评估了检测精度、运行时开销和编队保持误差。结果表明,所提出的框架能够以合理的计算成本维持稳定的领航-跟随编队,为将基于视觉的协同编队控制扩展到实际多无人机系统提供了可行的技术基础。
cs.RO / 36 / 2609.01453

Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds

模仿学习能否保持灵巧操作的时间鲁棒性?跨任务执行速度的专家-学习者对比研究
Enwerem, Clinton, Baras, John S., Belta, Calin
Abstract
Dexterous manipulation policies learned by imitation are typically evaluated for robustness to variation in scenes, objects, or instructions, but their performance across task execution speeds is less often examined. This leaves open how much temporal robustness a learner retains relative to the expert it imitates. We compare an expert and learner under the same task conditions, initial-condition draws, and speedup factors. We instantiate the evaluation in ParcelStow, a contact-rich task in which the robot acquires, reorients, and inserts a parcel. The demonstrations span the speedup range for the manipulation phases after parcel acquisition. A scripted expert and an Action Chunking with Transformers (ACT) policy trained from the expert's demonstrations both achieve 100 percent task success at nominal speed. Their success rates diverge within the demonstrated range: at its maximum, expert success is 84 percent and ACT success is 53 percent. Two ACT policies with different parameter initializations show similar degradation, decreasing by 34 and 48 percentage points from nominal speed to the maximum demonstrated speed, compared with 16 points for the expert. Stage-level analysis shows that 35 of ACT's 47 failures at the maximum demonstrated speed are insertion misalignments. Under the relative-motion handoff, every ACT acquisition retains the parcel through reorientation and transfer in free space, but only 64 percent complete the overall task, compared with 95 percent after expert acquisition. Across all evaluated policies and speeds, none of the 414 acquisitions without force closure completes the task. Equal nominal task success therefore does not imply preservation of expert performance across execution speeds. Code, data, and evaluation scripts are available at https://github.com/coenwerem/parcelstow.
Chinese Translation
通过模仿学习获得的灵巧操作策略通常在场景、物体或指令变化的鲁棒性方面进行评估,但其在不同任务执行速度下的表现却较少被研究。这使得学习者相对于其模仿的专家能保留多少时间鲁棒性仍是一个开放问题。我们在相同的任务条件、初始条件采样和加速因子下,对专家和学习者进行了对比。我们在ParcelStow任务中实例化该评估,这是一个接触丰富的任务,机器人需要抓取、重新定向并插入一个包裹。演示涵盖了包裹抓取后操作阶段的加速范围。基于脚本化的专家策略和由专家演示训练的Action Chunking with Transformers(ACT)策略在标称速度下均达到100%的任务成功率。然而,在演示范围内,两者的成功率出现分化:在最大速度下,专家成功率为84%,而ACT为53%。两个采用不同参数初始化的ACT策略表现出相似的退化,从标称速度到最大演示速度分别下降34和48个百分点,而专家仅下降16个百分点。阶段级分析显示,ACT在最大演示速度下的47次失败中有35次是插入错位。在相对运动交接下,每次ACT抓取都能在自由空间中通过重新定向和转移保持包裹,但仅有64%完成整个任务,而专家抓取后为95%。在所有被评估的策略和速度下,414次未达到力封闭的抓取均未能完成任务。因此,标称任务成功率相同并不意味着在执行速度变化下能保持专家的性能。代码、数据和评估脚本可在 https://github.com/coenwerem/parcelstow 获取。
cs.RO / 37 / 2609.01518

A System for Fast, Resilient, and Adaptable Loco-Manipulation Behaviors on Humanoid Robots

一个人形机器人快速、鲁棒且可适应的移动操作行为系统
Calvert, Duncan, Penco, Luigi, Anderson, Dexton, Bialek, Tomasz, Chatterjee, Arghya, Park, Beomyeong, Griffin, Robert
Abstract
There is tremendous value in humanoid robots taking on physically demanding, hazardous, and repetitive work in spaces built for humans. However, a useful robot for these spaces must coordinate locomotion, whole-body motion, perception, contact, and operator supervision. We present a robot-local, runtime-editable behavior authoring and runtime system that addresses these challenges. We argue that behavior architecture can be a primary enabler of capability, speed, and reliability, and that runtime editability enables fast behavior creation, adaptation, extension, and combination. Our behavior architecture combines object-centric Affordance Templates, a tree structure that provides organization and logic, and runtime-editable perception through a behavior scene and primitive scene actions. Our operator interface remains continuously synchronized to the robot for runtime authoring, monitoring, and repair. Action primitives execute through a whole-body controller that supports concurrent body motions and walking. Demonstrations of our system cover six task variants on Unitree H1-2 and Alex. We execute a push door traversal in 34 seconds and sort six balls by color in 45 seconds under human disturbance. Timed authoring sessions show scratch creation of new loco-manipulation behaviors and adaptation of existing ones in hours. Comparison against the literature finds our approach to be competitive with recent learned systems.
Chinese Translation
人形机器人在为人类设计的空间中承担体力要求高、危险和重复性工作具有巨大价值。然而,适用于这些空间的有用机器人必须协调运动控制、全身运动、感知、接触以及操作员监督。我们提出了一种机器人本地化、可运行时编辑的行为创建与运行系统,以应对这些挑战。我们论证了行为架构可以成为能力、速度和可靠性的主要推动因素,且运行时可编辑性能够实现行为的快速创建、适应、扩展和组合。我们的行为架构结合了以对象为中心的 Affordance Templates(可供性模板)、提供组织与逻辑的树形结构,以及通过行为场景和基元场景动作实现的运行时可编辑感知。我们的操作员界面与机器人保持持续同步,以支持运行时的行为创建、监控和修复。动作基元通过支持并发身体运动与行走的全身控制器执行。我们在 Unitree H1-2 和 Alex 机器人上演示了六种任务变体,验证了系统的有效性。我们在人类干扰下于34秒内完成推门通行,并在45秒内按颜色分拣六个球。定时创建会话表明,新移动操作行为可从零开始在数小时内完成创建,现有行为也可在数小时内完成适应。与现有文献的对比表明,我们的方法与近期基于学习的系统相比具有竞争力。
cs.RO / 38 / 2609.01579

SG-AMP: Scene-Graph-Guided Active Perception and Semantics-Aware Motion Planning for Pepper Plants

SG-AMP:面向辣椒植株的场景图引导主动感知与语义感知运动规划
Menon, Rohit, Lolla, Shiva Rudra, Mueller-Goldingen, Niklas, Chenchani, Gokul, Roscher, Ribana, Bennewitz, Maren
Abstract
We present SG-AMP, integrating robust depth completion with input-conditioned uncertainty, persistent panoptic mapping, plant scene-graph reasoning, and semantics-aware active view-motion planning. Beyond inspecting uncertain observed regions, the scene graph explicitly hypothesizes unobserved pepper--peduncle attachments and directs close-range sensing toward them. Candidate views are selected according to expected information gain, while class-dependent motion costs distinguish protected peppers, peduncles, and stems from conditionally traversable foliage. On pepper data, the perception network achieves $55.27\%$ semantic mIoU, $38.67\%$ PQ, and $40.62\,\mathrm{mm}$ depth RMSE, while input-conditioned uncertainty improves NYUv2 NLL from $-1.6518$ to $-1.6925$ and AUSE from $0.0102$ to $0.0087$.
Chinese Translation
我们提出了SG-AMP,该方法集成了具有输入条件化不确定性的鲁棒深度补全、持久化全景地图构建、植物场景图推理以及语义感知的主动视点运动规划。除检查不确定性较高的已观测区域外,场景图还对未观测到的辣椒-果柄连接处进行显式假设,并引导近距离传感朝向这些区域。候选视点根据预期信息增益进行选择,同时基于类别的运动代价将受保护的辣椒、果柄和茎干与可以有条件穿越的枝叶区分开来。在辣椒数据上,感知网络达到了55.27%的语义mIoU、38.67%的PQ以及40.62毫米的深度RMSE;输入条件化不确定性将NYUv2数据集上的NLL从-1.6518提升至-1.6925,AUSE从0.0102提升至0.0087。
cs.RO / 39 / 2609.01596

Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation

Facet-0:面向接触密集型精细操作的机器人基础模型
Deng, Haoyuan, Liu, Haichao, Guo, Wenkai, Ling, Yuan, Yang, Zaijia, Xue, Yuanjiang, Sun, Haosheng, Wang, Liangzi, Wang, Ziwei
Abstract
Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and robustness to contact failures. We present Facet-0, a robotic foundation model that predicts and values the contact consequences of its actions. Facet-0 unifies multimodal representation learning and reinforcement learning (RL) post-training around a joint action-wrench proposal: a causal wrench history is aligned with vision-language semantics and kinematic state, and flow matching generates each action chunk together with the future wrist-wrench profile it is expected to induce. Deployment rollouts train a distributional Action-Wrench Critic to distinguish motions with similar task progress but different contact outcomes, while phase-aware rewards and contact-selective credit concentrate policy improvement on decisive interactions. To accommodate part-specific dynamics, a lightweight bounded actor reuses the frozen representation for on-robot adaptation; RL remains defined over executable Cartesian actions, while an auxiliary wrench head preserves predictive, non-commanded action-contact coupling. Trained on ManuFacet-1K, a 1,000-hour force-synchronized corpus spanning three embodiments and multiple manufacturing cells, the bounded task-adapted system reaches 82% mean success on five sub-millimeter computer-assembly tasks, compared with 15% for the strongest baseline, with 0.5 mm placement accuracy and 50 ms command latency.
Chinese Translation
真实世界中亚毫米公差的机器人装配要求空间精度、柔顺交互以及对接触失败的鲁棒性。我们提出了Facet-0,一个能够预测并评估其动作的接触后果的机器人基础模型。Facet-0围绕联合动作-力旋量(action-wrench)提案,统一了多模态表征学习与强化学习(RL)后训练:因果力旋量历史与视觉-语言语义及运动学状态对齐,流匹配(flow matching)生成每个动作块及其预期引起的未来腕部力旋量剖面。部署 rollout 训练一个分布式的动作-力旋量评价器(Action-Wrench Critic),以区分任务进度相似但接触结果不同的运动,同时阶段感知奖励与接触选择性信用分配将策略改进集中于决定性交互。为适应零件特定的动力学,一个轻量级有界执行器复用冻结的表征进行机器人上的自适应;RL 仍定义在可执行的笛卡尔动作之上,而辅助力旋量头保持预测性的、非指令性的动作-接触耦合。该有界的任务自适应系统在 ManuFacet-1K 上训练——这是一个跨越三种机器人本体和多个制造单元的1000小时力同步语料库——在五个亚毫米计算机装配任务上达到82%的平均成功率,而最强基线仅为15%,放置精度为0.5毫米,指令延迟为50毫秒。