← Back to Index
Daily Research Digest

arXiv Papers

2026-09-11
188
Papers
3
Categories
188
Translated
收藏清单 0
人工智能 (Artificial Intelligence)
77
cs.AI / 1 / 2609.10584

Probabilistic Focal Search: Accelerating Bounded-Suboptimal Search via Lower-Bound Advancement

概率聚焦搜索:通过下界推进加速有界次优搜索
Duc, Minh Vu, Huu, Trung Le, Hoàng, Hà Minh, Nguyen, Trung Thanh, Nguyen, Phuong Khanh, Binh, Huynh Thi Thanh
Abstract
Bounded-suboptimal search seeks a solution within a factor $w$ of optimal while reducing search effort. Focal Search (FS) uses heuristic guidance within FOCAL, the frontier nodes eligible under the threshold $w f_{\min}$, but its deterministic policy may leave $f_{\min}$ unchanged for many expansions. We introduce Probabilistic Focal Search (PFS), which follows the FS guided choice with probability $p$ and expands a minimum-$f$ OPEN node with probability $1-p$. The latter branch encourages the lower bound to advance, enlarging FOCAL and admitting nodes that may lead to feasible solutions. By balancing guidance and lower-bound advancement, this mechanism can reduce time to a bounded solution when progress is limited by delayed FOCAL admission. As a secondary transfer experiment, we apply the same scheduler to Dynamic Potential Search, yielding Probabilistic Dynamic Potential Search (PDPS). We benchmark PFS against FS on N-Puzzle, Pancake Sorting, and the Traveling Salesperson Problem (TSP), and evaluate its anytime extension on the Generalized Covering TSP (GCTSP), using multiple $w$ and $p$ values. Across these benchmarks, the largest gains occur when long $f_{\min}$ plateaus delay useful FOCAL admissions; in such settings, the probabilistic factor may reduce node expansions by about 90\% or more (e.g., on N-Puzzle and TSP). For the anytime algorithm family, Anytime Probabilistic Focal Search (APFS) outperforms all tested algorithms in evaluating anytime methods on GCTSP. We also observe that the benefit is smaller when the deterministic search already advances efficiently (e.g., Pancake Sorting), indicating that the probabilistic factor is most useful when FOCAL admission is a search bottleneck. The PDPS transfer shows that the mechanism also transfers to potential guidance, although its common-success effects remain domain- and bound-dependent.
Chinese Translation
有界次优搜索旨在寻找一个在最优解 $w$ 倍范围内的解,同时减少搜索工作量。聚焦搜索(Focal Search, FS)在 FOCAL 内使用启发式引导,FOCAL 是在阈值 $w f_{\min}$ 下有资格的前沿节点,但其确定性策略可能在多次扩展中保持 $f_{\min}$ 不变。我们引入了概率聚焦搜索(Probabilistic Focal Search, PFS),它以概率 $p$ 遵循 FS 的引导选择,并以概率 $1-p$ 扩展一个最小 $f$ 的 OPEN 节点。后一分支促使下界推进,扩大 FOCAL 并允许可能通向可行解的节点进入。通过平衡引导和下界推进,当进展受限于延迟的 FOCAL 准入时,该机制可以减少到达有界解的时间。作为次要的迁移实验,我们将相同的调度器应用于动态势搜索(Dynamic Potential Search),得到概率动态势搜索(Probabilistic Dynamic Potential Search, PDPS)。我们在 N-Puzzle、Pancake Sorting 和旅行商问题(Traveling Salesperson Problem, TSP)上对 PFS 与 FS 进行基准测试,并使用多个 $w$ 和 $p$ 值在广义覆盖 TSP(Generalized Covering TSP, GCTSP)上评估其任意时间扩展。在这些基准测试中,最大的收益出现在长期 $f_{\min}$ 平台期延迟了有用的 FOCAL 准入时;在这种情况下,概率因素可以将节点扩展减少约 90% 或更多(例如,在 N-Puzzle 和 TSP 上)。对于任意时间算法系列,在 GCTSP 上评估任意时间方法时,任意时间概率聚焦搜索(Anytime Probabilistic Focal Search, APFS)优于所有测试的算法。我们还观察到,当确定性搜索已经高效推进时(例如,Pancake Sorting),收益较小,这表明当 FOCAL 准入是搜索瓶颈时,概率因素最有用。PDPS 的迁移表明该机制也可以迁移到势引导,尽管其常见成功效果仍然依赖于领域和边界。
cs.AI / 2 / 2609.10629

Automating Quadratic Unconstrained Binary Optimization (QUBO) Formulation Generation from Natural Language

从自然语言自动生成二次无约束二值优化(QUBO)形式
Mondal, Niloy Kumar, Parvez, Md Rizwan
Abstract
Quadratic Unconstrained Binary Optimization (QUBO) is a central formulation for combinatorial optimization and has gained increasing attention due to its compatibility with quantum, hybrid quantum-classical, and quantum-inspired solvers. However, translating natural-language problem descriptions into correct QUBO formulations remains difficult, requiring the identification of binary variables, constraints, objective functions, penalty terms, and suitable penalty weights. This process is time-consuming and often demands substantial domain expertise. To address this challenge, we propose an end-to-end multi-agent framework that automatically generates QUBO formulations from natural-language problem descriptions, supported by structured or unstructured test cases. To evaluate its performance, We also introduce QUBOBench, a benchmark containing 100 combinatorial optimization problems across 12 application domains, curated from peer-reviewed literature, competitions, and canonical NP-hard problems. Experimental results show that our framework achieves 68% accuracy on QUBOBench, outperforming a direct single-call baseline by 22%. Further analysis identifies iterative self-repair as the most important component contributing to improved performance. The data and code are open-sourced at https://quitttcat.github.io/QuantumQUBOAgent.
Chinese Translation
二次无约束二值优化(QUBO)是组合优化的核心形式,因其与量子、混合量子-经典及量子启发式求解器的兼容性而受到越来越多的关注。然而,将自然语言问题描述转化为正确的QUBO形式仍然困难,需要识别二值变量、约束、目标函数、惩罚项以及合适的惩罚权重。这一过程耗时且通常需要大量的领域专业知识。为了解决这一挑战,我们提出了一个端到端的多智能体框架,能够从自然语言问题描述中自动生成QUBO形式,并支持结构化或非结构化的测试用例。为了评估其性能,我们还引入了QUBOBench,一个包含12个应用领域、100个组合优化问题的基准测试集,这些题目选自同行评审文献、竞赛和经典的NP难问题。实验结果表明,我们的框架在QUBOBench上达到了68%的准确率,比直接单次调用的基线高出22%。进一步分析表明,迭代自修复是提升性能的最重要组成部分。数据和代码已在 https://quitttcat.github.io/QuantumQUBOAgent 开源。
cs.AI / 3 / 2609.10654

A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning

一个面向组合式与可解释认知推理的多阶段规则链式框架
Kar, Deblina
Abstract
The Abstraction and Reasoning Corpus (ARC) benchmarks cognitive generalization, the ability to infer and apply abstract rules from limited examples. This paper presents a multi-stage rule-chaining framework that performs compositional reasoning across symbolic, structural, and conceptual levels. The framework integrates three complementary solvers: (1) a deterministic rule discovery module that induces atomic transformations through geometric, color, and object-based analysis; (2) a pattern-composition engine that reconstructs outputs via block merging, repetition, and spatial heuristics; and (3) a structural abstraction layer that infers hierarchical and nested relationships across grids. These solvers operate sequentially within a progressive fallback hierarchy, where each stage reuses prior reasoning traces to enhance interpretability and generalization. Training passed for 995 tasks out of 1000, further evaluated on 105 tasks out of 120 and solved 230 test tasks out of 240 ARC-AGI-2 tasks. The system achieved strong coverage across deterministic, compositional, and abstract categories, demonstrating an overall accuracy exceeding 95 percent. The proposed architecture bridges symbolic reasoning and pattern synthesis, providing interpretable insight into cognitive generalization. The results suggest that rule chaining and hierarchical composition can advance machine reasoning toward transparent, human-aligned abstraction without relying on task-specific tuning.
Chinese Translation
抽象与推理语料库(ARC)为认知泛化提供基准测试,即从有限示例中推断和应用抽象规则的能力。本文提出了一种多阶段规则链式框架,该框架在符号、结构和概念层面进行组合推理。该框架集成了三个互补的求解器:(1)确定性规则发现模块,通过几何、颜色和基于对象的分析归纳原子变换;(2)模式组合引擎,通过块合并、重复和空间启发式重建输出;(3)结构抽象层,推断跨网格的层次和嵌套关系。这些求解器在渐进式回退层次结构中顺序运行,其中每个阶段重用先前的推理轨迹以增强可解释性和泛化能力。在1000个任务中,训练通过了995个;进一步在120个任务中的105个上进行了评估,并在240个ARC-AGI-2任务中解决了230个测试任务。该系统在确定性、组合和抽象类别中实现了强大的覆盖,总体准确率超过95%。所提出的架构桥接了符号推理和模式合成,为认知泛化提供了可解释的见解。结果表明,规则链和层次组合可以推动机器推理走向透明、与人类对齐的抽象,而不依赖于特定任务的调优。
cs.AI / 4 / 2609.10656

Understanding LoRA Rank Trade-offs in Diffusion Model Fine-Tuning

理解扩散模型微调中LoRA秩的权衡
Khazrak, Iman, Nejad, Narges, Rezaee, Mostafa M., Green II, Robert C.
Abstract
Selecting LoRA rank for diffusion fine-tuning requires balancing quality and compute cost. We present a controlled study on CIFAR-10 using a DDPM U-Net with ranks {2,4,8,16,32}, fixed optimization settings, and a reproducible local-folder pytorch-fid protocol. We report FID, trainable parameters, runtime, and GPU memory, then validate trends with extended-budget DDPM runs (20 epochs; ranks 4/8/16) and a Tiny DiT backbone (10 epochs; ranks 4/8/16). Results show moderate ranks are most efficient: rank 4 achieves the best DDPM FID (124.1380), rank 8 is close (124.2136), and higher ranks provide limited gains despite larger adaptation cost. These findings support small-to-moderate ranks as practical defaults under fixed training budgets.
Chinese Translation
为扩散模型微调选择LoRA秩需要在质量与计算成本之间取得平衡。我们在CIFAR-10上使用DDPM U-Net进行了一项受控研究,秩取{2,4,8,16,32},采用固定优化设置和可复现的本地文件夹pytorch-fid协议。我们报告了FID、可训练参数量、运行时间和GPU内存,随后用延长预算的DDPM运行(20个epoch;秩4/8/16)和Tiny DiT骨干网络(10个epoch;秩4/8/16)验证趋势。结果表明,中等秩最高效:秩4取得最佳DDPM FID(124.1380),秩8接近(124.2136),而更高秩尽管适配成本更大,带来的增益有限。这些发现支持在固定训练预算下将小到中等秩作为实用默认选择。
cs.AI / 5 / 2609.10657

Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

量化记忆到泛化的转变:Grokking 中的标度律与相结构
Kataria, Anish
Abstract
Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{why} this transition occurs, the quantitative structure of \emph{when} it occurs in hyperparameter space remains uncharacterized. We map the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for generalization onset time: $T_{\mathrm{grok}} \propto H^{-0.27}\, D^{-2.04}\, \eta^{-0.50}\, \lambda^{-0.64}$ ($R^2 = 0.732$; $0.821$ with interactions). The exponent hierarchy reveals that data complexity ($D^{-2.04}$) is the dominant driver of regime transition, not model capacity ($H^{-0.27}$): doubling data accelerates generalization by ${\sim}4\times$, while doubling width yields only ${\sim}1.2\times$. A sharp phase boundary at weight decay $\lambda \gtrsim 1.0$ separates grokking from non-grokking configurations, and weight norm trajectories show monotonic compression during the transition, consistent with implicit regularization selecting low-complexity solutions. These results provide a quantitative foundation for predicting and controlling regime transitions in overparameterized networks.
Chinese Translation
在训练超过记忆点后,神经网络常常会经历延迟的泛化转变,这一现象被称为 grokking。尽管关于这一转变为何发生的理论已有进展,但其在超参数空间中何时发生的定量结构仍未被刻画。我们在模算术任务上,针对双隐藏层 MLP 的 384 种配置,刻画了记忆到泛化的边界,并拟合了泛化起始时间的幂律标度关系:T_grok ∝ H^{-0.27} D^{-2.04} η^{-0.50} λ^{-0.64}(R^2 = 0.732;考虑交互作用时为 0.821)。指数层级表明,数据复杂度(D^{-2.04})而非模型容量(H^{-0.27})是状态转变的主导驱动因素:数据量翻倍使泛化加速约 4 倍,而宽度翻倍仅带来约 1.2 倍。在权重衰减 λ ≳ 1.0 处存在一个尖锐的相边界,将 grokking 与非 grokking 配置分开;权重范数轨迹在转变期间表现出单调压缩,这与隐式正则化选择低复杂度解相一致。这些结果为预测和控制过参数化网络中的状态转变提供了定量基础。
cs.AI / 6 / 2609.10712

An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics

IMO 金牌的开放配方:为奥林匹克数学训练 Nemotron
Moshkov, Ivan, Ge, Stephen, Armstrong, George, Du, Wei, Mahdavi, Sadegh, Gitman, Igor
Abstract
We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkpoint choice, verification, and refinement. Based on these findings, we present an open-model test-time-compute pipeline. The system operates entirely in natural language, with no formal prover, external tools, or internet access. Three Nemotron 3 Ultra checkpoints - the general-availability model and two post-trained specialists - power an iterative search that generates, verifies, and refines candidate proofs; a separate high-compute stage then selects each final submission. The system scored 30 out of 42 points at IMO 2026, reaching the gold-medal threshold. We release the two post-trained checkpoints as well as the training data, the training and inference code, the submitted solutions, and Nemotron-IMO-Bench, a new benchmark of 200 novel olympiad-level problems.
Chinese Translation
我们研究了模型后训练和测试时推理设计如何影响高难度奥林匹克数学的自然语言证明生成。从 Nemotron 3 Ultra 出发,我们使用监督微调和强化学习训练了两个专家检查点,并评估了检查点选择、验证和精炼。基于这些发现,我们提出了一个开放模型测试时计算流水线。该系统完全以自然语言运行,无需形式化证明器、外部工具或互联网访问。三个 Nemotron 3 Ultra 检查点——通用可用模型和两个后训练专家——驱动一个迭代搜索,该搜索生成、验证和精炼候选证明;然后一个单独的高计算阶段选择每个最终提交。该系统在 IMO 2026 上获得 42 分中的 30 分,达到金牌阈值。我们发布了两个后训练检查点以及训练数据、训练和推理代码、提交的解决方案,以及 Nemotron-IMO-Bench,一个包含 200 道新颖奥林匹克级别问题的新基准。
cs.AI / 7 / 2609.10724

Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge

完成任务还不够:评估累积挑战下的智能体韧性与体贴参与
Bai, Yuanchen, Ding, Zijian, Taylor, Angelique
Abstract
Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We propose operational resilience and considerate participation as two complementary aspects of evaluating such agents: the former captures how agents recover from blocked work while preserving progress and communicating their limits, and the latter captures how their adaptation accounts for affected people, role boundaries, and the surrounding workflow. Yet both remain underexplored under accumulating challenge. We study 120 simulated healthcare trajectories across two generative AI models and twelve stakeholder-derived tasks under light, medium, and heavy challenge. We compare textual action plans, prompted internal assessments, and quantitative structured workload and affect reports to examine how agent behavior and reported state change as challenge accumulates. Regarding operational resilience, agents shift from self-directed recovery toward greater human dependence, while reporting increasing workload and negative affect in structured reports but seldom expressing strain in textual responses. Regarding considerate participation, agents broaden from task-focused adaptation toward task reframing, attention to others, role-boundary adjustment, and wider coordination, with distinct patterns across actions and internal assessments. From these findings, we derive five deployment dilemmas involving persistence, attention, role boundaries, state disclosure, and escalation that require stakeholder specification, further informing technical implications for learning, situated evaluation, and embodied adaptation.
Chinese Translation
生成式AI智能体的持续部署需要的不仅仅是孤立任务的成功。智能体必须在重复交互、变化的条件以及对共享工作流中人员的依赖中保持有用,尤其是随着技术、人员和运营中断随时间累积。我们提出运营韧性和体贴参与作为评估此类智能体的两个互补方面:前者捕捉智能体如何从受阻工作中恢复,同时保持进展并传达其限制,后者捕捉其适应如何考虑受影响的人员、角色边界和周围工作流。然而,在累积挑战下,两者都仍未得到充分探索。我们研究了在轻度、中度和重度挑战下,跨两个生成式AI模型和十二个源自利益相关者的任务的120条模拟医疗轨迹。我们比较文本行动计划、提示的内部评估以及定量的结构化工作量和情感报告,以检查随着挑战累积,智能体行为和报告状态如何变化。关于运营韧性,智能体从自我导向的恢复转向更大程度的人类依赖,同时在结构化报告中报告增加的工作量和负面情感,但在文本回应中很少表达压力。关于体贴参与,智能体从以任务为中心的适应扩展到任务重构、关注他人、角色边界调整和更广泛的协调,在行动和内部评估中呈现出不同的模式。从这些发现中,我们得出五个部署困境,涉及持久性、注意力、角色边界、状态披露和升级,需要利益相关者具体说明,进一步为学习、情境评估和具身适应的技术含义提供信息。
cs.AI / 8 / 2609.10728

Towards a Deterministic Math Solver for Clinical Language Models

面向临床语言模型的确定性数学求解器
Osorio, Felipe Ocampo, Ordoñez, Sebastián Andrés Cajas, Lange, Maximin, Attrach, Rafi Al, Kapadia, Sahil, Sellami, Zakaria Laouabdia, Talio, Angelo Antonio, Celi, Leo Anthony
Abstract
Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard response is to hardcode each calculator as a validated function, one at a time. We test an alternative: the model does not calculate. Instead, it writes case-specific Python that a restricted local executor runs as a deterministic solver, and the model's task reduces to deciding how to use it. We evaluate this Program-Solve interface on MedCalc-Bench Verified (1,100 cases, 55 calculators) against direct model arithmetic and a hand-written 22-calculator library, using Qwen2.5-7B and Qwen2.5-32B-AWQ, after auditing the benchmark's formulas against current clinical guidelines and flagging 16 of 55 with version, use or coefficient concerns. With formulas and gold variables supplied and both routes reading the whole note, handing off to the solver is not a reliable advantage at 7B (75.31% against 72.02%, a paired +3.29 points with a 95% calculator-cluster interval of [-3.49, 10.38]) but is one at 32B (90.53% against 83.47%, +7.05 [0.47, 14.60], clear of zero). The hand-written library is exact on its 440 supported cases but abstains elsewhere (40.0% overall). Adding an executor thus helps some open-weight models more than others even under matched formula, variable and note access, and is not a substitute for verified formulas or reliable variable extraction either way.
Chinese Translation
大型语言模型在算术方面不可靠,这对临床计算器来说是个问题,因为一个数值错误就会改变推荐结果。标准做法是将每个计算器逐一硬编码为经过验证的函数。我们测试了一种替代方案:模型不进行计算。相反,它编写针对具体病例的 Python 代码,由受限的本地执行器作为确定性求解器运行,模型的任务简化为决定如何使用它。我们在 MedCalc-Bench Verified(1,100 个病例,55 个计算器)上评估这种 Program-Solve 接口,将其与直接模型算术和手写的 22 个计算器库进行对比,使用 Qwen2.5-7B 和 Qwen2.5-32B-AWQ,并在此前根据当前临床指南审计了基准的公式,标记了 55 个中的 16 个存在版本、使用或系数问题。在提供公式和金标准变量且两条路径都读取整份病历记录的情况下,将任务移交给求解器在 7B 上并非可靠优势(75.31% 对 72.02%,配对 +3.29 个百分点,95% 计算器聚类区间为 [-3.49, 10.38]),但在 32B 上则是(90.53% 对 83.47%,+7.05 [0.47, 14.60],不包含零)。手写库在其 440 个支持的病例上完全正确,但在其他地方弃权(总体 40.0%)。因此,即使在公式、变量和记录访问匹配的情况下,添加执行器对某些开放权重模型的帮助也大于其他模型,并且无论如何都不能替代经过验证的公式或可靠的变量提取。
cs.AI / 9 / 2609.10824

Studying Without a Syllabus: Task-Agnostic Environment Preprocessing

无教学大纲的学习:任务无关的环境预处理
Samuel, Vinay, Ursekar, Varun, Kalmath, Vijay S., Shanker, Apaar, Chatrath, Veronica, Xue, Yuan
Abstract
Before an LLM agent tackles tasks in a new environment, it can inspect available corpora and tools and construct reusable resources such as indices, scripts, or procedural guidance. Most automated adaptation methods, however, rely on task examples, trajectories, or evaluation feedback to decide what to build. Existing task-agnostic approaches avoid this supervision but commit in advance to a preparation strategy for a particular type of environment. We study a more open-ended setting: can an agent study an unfamiliar environment without a syllabus, i.e. before test time and without knowledge of the downstream task distribution, and choose how to prepare it? We formalize task-agnostic environment preprocessing, in which a studying system explores an environment under a budget and produces artifacts for a frozen solver. We compare unaided and archive-equipped meta-agents with fixed synthetic-practice and corpus-processing methods across six heterogeneous benchmarks. A meta-agent variant achieves the highest Avg@3 reward on five benchmarks, while fixed corpus processing remains best on the largest corpus benchmark. Larger study budgets do not reliably improve downstream reward. Nevertheless, studied artifacts reduce the test-time sampling needed to reach a given score, demonstrating how reusable preparation can shift computation from repeated test-time attempts to a pre-task study phase.
Chinese Translation
在LLM智能体在新环境中处理任务之前,它可以检查可用的语料库和工具,并构建可重用的资源,如索引、脚本或程序性指导。然而,大多数自动适应方法依赖任务示例、轨迹或评估反馈来决定构建什么。现有的任务无关方法避免了这种监督,但提前承诺针对特定类型环境的准备策略。我们研究一个更开放式的设定:一个智能体能否在没有教学大纲的情况下研究一个不熟悉的环境,即在测试时间之前且不知道下游任务分布的情况下,并选择如何准备它?我们将任务无关的环境预处理形式化,其中研究系统在预算下探索环境,并为冻结的求解器生成工件。我们在六个异构基准上比较了无辅助和配备档案的元智能体与固定的合成练习和语料处理方法。一个元智能体变体在五个基准上实现了最高的Avg@3奖励,而固定的语料处理在最大的语料库基准上仍然最佳。更大的研究预算并不能可靠地提高下游奖励。然而,研究的工件减少了达到给定分数所需的测试时间采样,展示了可重用的准备如何将计算从重复的测试时间尝试转移到任务前的研究阶段。
cs.AI / 10 / 2609.10873

When Validation Stops Learning: Auditing Update Admission for Continual Embodied Agents

当验证阻止学习:面向持续具身智能体的更新准入审计
Ma, Qinzhen, Wu, Ruihai
Abstract
Independent evaluation can reject harmful policy updates yet also prevent useful continual learning. We argue that update admission must be assessed through both error control and retained learning opportunities at a stated interaction budget. We identify a concrete failure: a range-based confidence gate cannot certify unchanged old-task behavior within otherwise substantial budgets. A standard paired-binomial construction reduces this burden when outcome disagreements are rare. We also specify certified historical-reference promotion and a round-level missed-opportunity metric. In a constructed one-step pushing diagnostic with 32 seeds, fresh paired checks admit 31.6% of a common update stream at 2,000 episodes per stage, versus zero for the range-based gate; unconditional replay nevertheless learns better in closed-loop runs. A separate learned-dynamics stress test distinguishes model bias from feedback-selection error. The contribution is an admission-audit protocol with analytical and synthetic evidence; physical-robot and VLA validation remain open.
Chinese Translation
独立评估可以拒绝有害的策略更新,但也会阻止有用的持续学习。我们认为,更新准入必须通过在给定交互预算下同时考察误差控制与保留的学习机会来评估。我们识别出一个具体失效:基于范围的置信门控在本来相当充足的预算内也无法证明旧任务行为保持不变。当结果不一致很少时,标准的配对二项构造可减轻这一负担。我们还规定了经认证的历史参考晋升以及轮级错失机会指标。在一个构建的单步推动诊断任务中,使用32个种子,每阶段2,000个回合时,新的配对检查允许一个常见更新流中的31.6%通过,而基于范围的门控为0;然而在闭环运行中,无条件回放学习得更好。一项单独的学习动力学压力测试区分了模型偏差与反馈选择误差。贡献是一个带有解析与合成证据的准入审计协议;物理机器人验证和VLA验证仍有待完成。
cs.AI / 11 / 2609.10964

Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows

将就绪与释放解耦:面向智能体式 LLM 工作流的尾部感知调度
Feng, Bochao, Li, Jianjiang, Wang, Haojie, Qiao, Lin, Li, Yinghui, Yan, Yukun, Zhai, Jidong
Abstract
Agentic LLM workflows consist of sequences of model turns interleaved with tool interactions, so their end-to-end completion time depends not only on inference speed but also on when ready turns are released. Most runtimes release each turn immediately upon readiness. Under contention, this eager release policy can accumulate released but unfinished work; once submitted, those turns can no longer be reordered by the workflow-level policy, increasing tail latency. We present a tail-risk-aware turn release scheduling method that jointly decides which ready turn to release next and how much released but unfinished work to maintain. The method uses a mean--Conditional Value-at-Risk (CVaR) objective to capture the evolving tail risk of unfinished workflows, incorporates online estimates of turn work when prioritizing ready turns, and adapts the released work budget to observed queue pressure. We evaluate the method using real agent execution traces from software engineering tasks across multiple LLMs and workflow arrival rates. The method performs comparably to eager release under light load and substantially reduces the P95 of workflow flow time under contention, achieving up to a \(3.50\times\) speedup.
Chinese Translation
智能体式 LLM 工作流由与工具交互交错的模型回合序列组成,因此其端到端完成时间不仅取决于推理速度,还取决于就绪回合何时被释放。大多数运行时会在每个回合就绪后立即释放。在资源竞争下,这种急切释放策略会累积已释放但未完成的工作;一旦提交,这些回合便无法再被工作流级策略重新排序,从而增加尾部延迟。我们提出一种尾部风险感知的回合释放调度方法,联合决定下一个释放哪个就绪回合,以及维持多少已释放但未完成的工作。该方法采用均值--条件风险价值(CVaR)目标来刻画未完成工作流不断演变的尾部风险,在优先处理就绪回合时纳入回合工作量的在线估计,并根据观测到的队列压力调整已释放工作预算。我们使用跨多个 LLM 和工作流到达率的软件工程任务真实智能体执行轨迹来评估该方法。该方法在轻负载下与急切释放性能相当,并在资源竞争下显著降低工作流流程时间的 P95,最高实现 3.50× 加速。
cs.AI / 12 / 2609.10992

Demystifying the Privacy-Utility Trade-off in LLM Interactions

揭秘LLM交互中的隐私-效用权衡
Liu, Zhenhua, Xie, Zhanxu, Yu, Junjie, Zhu, Tong, Li, Lijun, Chen, Wenliang
Abstract
The integration of Large Language Models into daily tasks relies on context-rich instructions, inevitably exposing sensitive user information. Current privacy-preserving methods typically employ context-agnostic static rules, causing severe utility degradation. However, the specific mechanisms governing how sanitization impacts downstream performance remain largely underexplored. To address this, we conduct a systematic analysis to deconstruct the privacy-utility trade-off, uncovering three underlying mechanisms: (1) Context-Dependent Utility, which first establishes when to sanitize by revealing that data value shifts from critical constraints to dispensable noise based on user intent; (2) Strategic Adaptation, which subsequently determines how to sanitize by dictating that the choice between removal and replacement depends on the task's reliance on factual integrity versus structural coherence; and (3) Combinatorial Interplay, which finally extends the protection scope by demonstrating that attributes form a semantic web of synergistic dependencies or antagonistic redundancies. Guided by these insights, we introduce an intent-driven local protection framework. By distilling a lightweight model Veilmind-4B to drive a dynamic extraction-sanitization-restoration pipeline, our approach reaches a low-leakage privacy point while preserving substantially higher response utility than existing privacy-oriented baselines, advancing the privacy-utility trade-off toward the Pareto frontier.
Chinese Translation
大型语言模型融入日常任务依赖于上下文丰富的指令,这不可避免地会暴露敏感的用户信息。当前的隐私保护方法通常采用与上下文无关的静态规则,导致效用严重下降。然而,关于净化处理如何影响下游性能的具体机制在很大程度上仍未得到充分探索。为解决这一问题,我们进行了一项系统性分析,以解构隐私-效用权衡,揭示了三个潜在机制:(1) 上下文相关效用,它首先通过揭示数据价值如何基于用户意图从关键约束转变为可弃噪声,从而确定何时进行净化;(2) 策略性适应,它随后通过指明在移除与替换之间的选择取决于任务对事实完整性还是结构连贯性的依赖,来决定如何净化;(3) 组合交互,它最终通过证明属性形成协同依赖或拮抗冗余的语义网络,来扩展保护范围。在这些见解的指导下,我们引入了一个意图驱动的本地保护框架。通过蒸馏一个轻量级模型Veilmind-4B来驱动动态的提取-净化-恢复流水线,我们的方法在达到低泄漏隐私点的同时,保持了比现有面向隐私的基线显著更高的响应效用,将隐私-效用权衡推向帕累托前沿。
cs.AI / 13 / 2609.11018

Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks

定义AI智能体:准则、指标与基准汇编
Lassiter, Mia, Bent, Brinnae
Abstract
The term agent in artificial intelligence lacks a standard definition, complicating the evaluation, comparison, and reproducibility of AI agent research. We address this ambiguity through a survey organized around five dimensions of agenticness: environmental interaction, learning and adaptation, autonomy, goal-directed behavior, and temporal coherence. For each dimension, we examine how the underlying capability has been conceptualized across prior work and synthesize the metrics, benchmarks, and evaluation frameworks used to assess it. This review provides a structured account of the current landscape of agent evaluation, highlighting both established approaches and areas where evaluation remains limited or inconsistent. We additionally introduce the Agent Compendium, a public-facing digital resource that organizes and extends the evaluation methods identified through this review. Together, the survey and compendium provide a common structure for evaluating and comparing agent capabilities across AI systems, supporting more reproducible research, clearer communication, and more systematic study of artificial agents.
Chinese Translation
人工智能中的“智能体”(agent)一词缺乏标准定义,这使AI智能体研究的评估、比较和可复现性变得复杂。我们围绕智能体特性(agenticness)的五个维度组织了一项综述,以解决这一歧义:环境交互、学习与适应、自主性、目标导向行为和时间一致性。对于每个维度,我们考察其底层能力在既往工作中是如何被概念化的,并综合用于评估该能力的指标、基准和评估框架。本综述对智能体评估的当前图景进行了结构化梳理,既突出了已有成熟方法,也指出了评估仍然有限或不一致的领域。此外,我们介绍了Agent Compendium,这是一个面向公众的数字资源,用于组织和扩展本综述所识别出的评估方法。综述与汇编共同提供了一个通用结构,用于评估和比较跨AI系统的智能体能力,从而支持更可复现的研究、更清晰的交流以及对人工智能体更系统的研究。
cs.AI / 14 / 2609.11030

The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures

智能体事件登记处:迈向防止重复的AI智能体故障
Kumar, Divyanshu, HN, Rohith, Birur, Nitin Aravind, Agarwal, Sahil, Harshangi, Prashanth
Abstract
AI agents increasingly act through tools and delegated authority, but general incident repositories rarely capture the mechanisms needed to compare public failures with agent-security evaluations. We present the Agent Incident Registry (AIR), a source-linked catalog containing \N{} records of agent-related events disclosed from \Yfirst{} through \Ylast{}. Each record includes supporting evidence, a stable identifier, and missingness-aware labels for causal role, disclosure class, mechanism, and outcome. Among the \Nprimary{} generative-system records in which the agent acted, \Rprimary{} involved realized harm (\Pprimary\%). Realized outcomes concentrate in in-the-wild and safety-failure records, while responsible disclosures and research demonstrations are overwhelmingly demonstrated; the aggregate share therefore characterizes collection composition rather than deployment risk. After initial curation, a second human reviewer checked all \N{} records and their existing labels for completeness and correctness. In a deployment-analogue audit, InjecAgent's \NInjecAgentCases{} cases occupy three of AIR's twelve surfaces and are all attacker-triggered, whereas AIR contains \Nsafety{} no-adversary safety failures. AIR supports source-grounded case retrieval and evaluation-scope auditing, not failure-rate or control-efficacy estimation.
Chinese Translation
AI智能体越来越多地通过工具和委托权限行动,但通用的事件库很少捕获比较公共故障与智能体安全评估所需的机制。我们提出了智能体事件登记处(Agent Incident Registry, AIR),一个链接到来源的目录,包含从 \Yfirst{} 到 \Ylast{} 披露的 \N{} 条智能体相关事件记录。每条记录包括支持证据、稳定标识符,以及针对因果角色、披露类别、机制和结果的缺失感知标签。在智能体发挥作用的 \Nprimary{} 条生成式系统记录中,\Rprimary{} 条涉及已实现的伤害(\Pprimary\%)。已实现的结果集中在野外和安全故障记录中,而负责任的披露和研究演示则绝大多数是演示性的;因此,总体份额描述的是收集构成,而非部署风险。在初步整理后,第二名人类审查者检查了所有 \N{} 条记录及其现有标签的完整性和正确性。在一项部署类比审计中,InjecAgent 的 \NInjecAgentCases{} 个案例占据了 AIR 十二个表面中的三个,并且全部是攻击者触发的,而 AIR 包含 \Nsafety{} 个无对手安全故障。AIR 支持基于来源的案例检索和评估范围审计,而非故障率或控制效力估计。
cs.AI / 15 / 2609.11060

Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents

智能体记忆的接地:面向企业级智能体的环境探测式整理
Suresh, Susheel, Mak, Hazel, Bhatnagar, Sahil, Methani, Chhaya, Munoz, Alejandro Gutierrez
Abstract
Persistent memory is entering production-oriented agent platforms to help long-horizon agents accumulate experience across sessions. Yet a post-task curator agent restricted to completed trajectories can preserve errors, overgeneralize partial evidence, or retain stale knowledge. We introduce environment-probing curation, a deployment-compatible extension that gives an existing asynchronous curator agent least-privilege, read-only world tools to check, scope, and refresh candidate memories. It requires no model retraining and leaves the task agent, retriever, memory representation, and production write authority unchanged. In a production-like GitHub Copilot (GHCP) harness built on its SDK, we compare stateless execution, full in-context learning, GHCP + Mem, and GHCP + Mem (w/ Env Probing) on CLBench database exploration and 90 adapted APEX management-consulting tasks. On CLBench, probing raises pass rate from 39% to 73% and pass-discounted reward from 8.60 to 22.60 while reducing queries from 8.8 to 4.7 per question and task-agent cost from \$3.38 to \$1.68. Across six APEX worlds, all 18 memory-versus-baseline mean reward comparisons are positive and task-agent tool calls fall by 16--75%; probing gives the best task-agent reward gain per dollar in five worlds. Probing also attains higher mean reward than GHCP + Mem on both Sonnet 4.6 and Opus 4.7 without schema drift. Environment probing therefore turns existing agent-memory curation into an environment-informed, auditable process while preserving a compact task-time interface.
Chinese Translation
持久记忆正进入面向生产的智能体平台,以帮助长周期智能体跨会话积累经验。然而,仅限于已完成轨迹的任务后整理智能体可能保留错误、过度泛化部分证据或保留过时知识。我们引入环境探测式整理,这是一种部署兼容的扩展,为现有的异步整理智能体提供最小权限的只读世界工具,以检查、界定和刷新候选记忆。它无需模型重新训练,并保持任务智能体、检索器、记忆表示和生产写入权限不变。在基于其SDK构建的类生产环境GitHub Copilot (GHCP)测试框架中,我们在CLBench数据库探索和90个改编的APEX管理咨询任务上比较了无状态执行、完全上下文学习、GHCP + 记忆以及GHCP + 记忆(带环境探测)。在CLBench上,探测将通过率从39%提升至73%,通过折扣奖励从8.60提升至22.60,同时将每个问题的查询次数从8.8降至4.7,任务智能体成本从3.38美元降至1.68美元。在六个APEX世界中,所有18个记忆与基线的平均奖励比较均为正,任务智能体工具调用次数下降16-75%;探测在五个世界中给出了每美元最佳的任务智能体奖励增益。在Sonnet 4.6和Opus 4.7上,探测也获得了比GHCP + 记忆更高的平均奖励,且无模式漂移。因此,环境探测将现有的智能体记忆整理转变为一种环境感知的、可审计的过程,同时保持紧凑的任务时接口。
cs.AI / 16 / 2609.11061

Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning

在模型改变想法处进行分叉:面向树结构强化学习的信念转移分支
Lei, Bin, Li, Yu, Choubey, Prafulla Kumar, Zhang, Jiaxin, Peng, Becky Xiangyu, Ye, Qinyuan, Narayan, Kartik, Ding, Caiwen, Savarese, Silvio, Wu, Chien-Sheng
Abstract
Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differences estimate step value. Each fork adds sampling cost, so realistic budgets typically allow only a few forks per chain. A fork placed where the outcome is already largely settled yields siblings that mostly agree and provide almost no credit signal; hence, for a given tree size, where forks are placed largely determines how much step-level RL can gain. Most existing mainstream methods place forks by structure, such as fixed lengths, midpoints, and delimiters, or by next-token entropy. We formalize fork placement as locating the \emph{pivots} of the chain's value curve, where the expected outcome turns. We propose \emph{belief-shift branching}: read the model's answer belief at candidate boundaries and fork just before the step where consecutive beliefs diverge most. Three instantiations, none needing step-level supervision, span access levels: a black-box probe, a logit-lens depth profile, and a learned activation direction, which is fit offline and therefore used only in the validation before RL training. The signal only \emph{places} forks, and the probe costs about $1\%$ of step compute on mathematics and under $5\%$ on code when it runs inside the rollout engine. In that validation, against Monte-Carlo value curves, a belief-shift signal ranks first in each of the eight model$\times$benchmark panels, ahead of entropy, structural, and LLM-judge baselines. In RL across three model families and two domains, belief-shift forking leads every mathematics aggregate, on OLMo-3-7B by $+2.6$ aggregate and $+2.9$ on AIME 2026 over the strongest baseline, and sweeps every OLMo code column, by $+6.5$ on LiveCodeBench-medium.
Chinese Translation
树结构 rollout 为无评论家的可验证奖励强化学习(RLVR)提供了步骤级信用:在中间点分叉一条链,兄弟结果差异即可估计步骤价值。每次分叉都会增加采样成本,因此现实预算通常每条链只允许少量分叉。在结果已基本确定的位置放置分叉,产生的兄弟节点大多一致,几乎不提供信用信号;因此,对于给定的树大小,分叉放置的位置在很大程度上决定了步骤级 RL 能获得多少收益。大多数现有主流方法通过结构来放置分叉,例如固定长度、中点和分隔符,或者通过下一词元熵。我们将分叉放置形式化为定位链价值曲线的枢轴点,即预期结果发生转变的位置。我们提出信念转移分支:在候选边界处读取模型的答案信念,并在连续信念差异最大的步骤之前进行分叉。三种实例化均不需要步骤级监督,跨越不同的访问级别:黑盒探针、logit-lens 深度剖面,以及学习的激活方向,后者在离线拟合,因此仅在 RL 训练前的验证中使用。该信号仅用于放置分叉,当探针在 rollout 引擎内运行时,其在数学上花费约 1% 的步骤计算量,在代码上低于 5%。在该验证中,针对蒙特卡洛价值曲线,信念转移信号在八个模型×基准面板中均排名第一,领先于熵、结构和 LLM 评判基线。在跨三个模型家族和两个领域的 RL 中,信念转移分叉在所有数学聚合指标上领先,在 OLMo-3-7B 上比最强基线高出 +2.6 聚合和 +2.9(AIME 2026),并横扫所有 OLMo 代码列,在 LiveCodeBench-medium 上高出 +6.5。
cs.AI / 17 / 2609.11065

MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG

MOSAIC:面向GraphRAG的查询感知探索策略自适应
Lee, EunKyeong, Oh, Kyeong-Jin, Kim, Jinwon, Lee, Hye Woo, Song, Minsang, Jang, Hyeongjun, Youn, Junyoung
Abstract
Graph Retrieval-Augmented Generation (GraphRAG) can connect evidence distributed across a corpus graph, but most systems use largely shared exploration procedures across queries. This creates a structural mismatch: direct facts may need compact local neighborhoods, comparisons need balanced coverage of multiple targets, and mediated questions may require deeper paths through weakly related connectors. We present Mosaic, a training-free framework that formulates GraphRAG retrieval as a per-query control problem. An LLM analyzer converts query-specific evidence requirements into a bounded policy over seed selection, graph traversal, stopping, and evidence selection, while the corpus graph, indexes, scoring functions, grounding procedure, and answer generator remain shared. On GraphRAG-Bench, Mosaic achieves query-weighted Answer Correctness of 76.97 on Medical and 64.33 on Novel, improving over the strongest previously reported overall results by 5.13 and 4.43 points. On Medical, it reaches 95.1 Evidence Recall and 86.1 Context Relevancy. Controlled comparisons on an identical graph and generator show that no fixed narrow, medium, or wide policy is consistently optimal; Mosaic improves by 9.96 points over the strongest canonical fixed policy. Relative to Fixed Wide, it evaluates 81.9% fewer paths and retains 47.2% fewer evidence items. Transfer experiments on HotpotQA, MuSiQue, and 2WikiMultiHopQA further show that the policy interface can be applied without benchmark-specific retriever training.
Chinese Translation
图检索增强生成(GraphRAG)能够连接分布在语料图上的证据,但大多数系统在不同查询间使用基本共享的探索过程。这造成了结构性错配:直接事实可能需要紧凑的局部邻域,比较需要多个目标的均衡覆盖,而中介性问题可能需要通过弱相关连接符的更深路径。我们提出Mosaic,一个免训练框架,将GraphRAG检索形式化为逐查询的控制问题。一个LLM分析器将查询特定的证据需求转换为关于种子选择、图遍历、停止和证据选择的有界策略,而语料图、索引、评分函数、证据grounding过程和答案生成器保持共享。在GraphRAG-Bench上,Mosaic在Medical上达到76.97的查询加权答案正确率,在Novel上达到64.33,比之前报告的最强总体结果分别提升5.13和4.43个百分点。在Medical上,它达到95.1的证据召回率和86.1的上下文相关性。在相同图和生成器上的受控比较表明,没有固定的窄、中或宽策略始终最优;Mosaic比最强的典型固定策略提升了9.96个百分点。相对于Fixed Wide,它评估的路径减少了81.9%,保留的证据项减少了47.2%。在HotpotQA、MuSiQue和2WikiMultiHopQA上的迁移实验进一步表明,该策略接口可以在无需基准特定检索器训练的情况下应用。
cs.AI / 18 / 2609.11115

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

Benchmark Radar:一个面向AI基准与评估的活数据库与搜索引擎
Wu, Koutian, Zhou, Junjie, Shang, Ergan, Wang, Jiayu, Han, Pengqian, Wang, Junkai, Xu, Wanghan
Abstract
Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluation evidence. Daily discovery draws on 37 sources: 13 direct connectors and 24 first-party research and engineering feeds. The catalog contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. We describe collection and retrieval, audit the full catalog, and examine benchmark saturation, adoption trends, and the limits of score comparisons. A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. We release the web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a command-line interface (CLI) for offline queries, and reproducible analysis.
Chinese Translation
基准测试的研究人员以及大型语言模型(LLM)和其他AI系统的开发者需要找到相关的评估,定位其基准数据集和代码,并理解所报告分数背后的设置。我们提出了Benchmark Radar,一个用于检索和发现AI基准的活数据库和搜索引擎,涵盖LLM评估、智能体和工具使用基准、编程、推理、安全以及特定领域的评估。该系统将每日发现的基准论文、代码库、数据集和发布与可搜索的基准目录、模型卡和技术报告中的提及以及分数历史相结合。它保留了来源身份和引用,以便读者可以检查候选基准及其评估证据。每日发现利用37个来源:13个直接连接器和24个第一方研究和工程信息源。该目录包含1,283条来源记录,来自4个基准目录,以及790条记录上的12,916个数值观测。我们描述了收集和检索,审计了整个目录,并检查了基准饱和、采用趋势以及分数比较的局限性。一个工作示例展示了一个完整的现有技术搜索,展示了在设计新评估时如何查询目录和检查基准证据。我们发布了Web仪表板,包含基准排行榜、分数与实测使用之间的帕累托前沿视图、饱和度和趋势视图、每日信息源、可下载证据、用于离线查询的命令行界面(CLI)以及可重复分析。
cs.AI / 19 / 2609.11127

KuaiRP Series Role-playing Models Technical Report

KuaiRP 系列角色扮演模型技术报告
Wang, Yipeng, Zhang, Ziwei, Zhang, Jiahui, Gan, Qi, Sheng, Kai
Abstract
This paper introduces the complete technical solution for the KuaiRP series of role-playing models. We aim to achieve four core objectives for a dedicated role-playing model: simplified prompt engineering, highly stable output quality, built-in domain world knowledge, and high-efficiency deployment with a small parameter size. However, effectively injecting deep domain knowledge often leads to a severe catastrophic forgetting of the model's general agent capabilities. To overcome this trade-off, we propose a multi-stage training pipeline. First, we design a standardized character template and construct an SFT data pipeline based on user behavior simulation and reverse profile filtering. Next, we utilize a rule-based composite reward function during the Reinforcement Learning (RL) phase to eliminate common degradation phenomena like length expansion and repetitive generation. Finally, to recover the general capabilities compromised during SFT and RL, we propose a novel self-distillation paradigm using Two-stage On-Policy Distillation (OPD) equipped with Cumulative-Divergence Decay (CDD). By using the domain-adapted model as the teacher and the original base model as the student, we effectively balance deep domain knowledge injection with the preservation of general agent capabilities. Experimental results demonstrate that the KuaiRP models not only match the current state-of-the-art proprietary models in role-playing fidelity within our target domains, but also successfully recover general agent capabilities, maintaining extremely low deployment costs.
Chinese Translation
本文介绍了 KuaiRP 系列角色扮演模型的完整技术方案。我们旨在为专用的角色扮演模型实现四个核心目标:简化的提示工程、高度稳定的输出质量、内置的领域世界知识,以及小参数规模下的高效部署。然而,有效注入深度领域知识往往会导致模型通用智能体能力的严重灾难性遗忘。为了克服这一权衡,我们提出了一种多阶段训练流程。首先,我们设计了标准化角色模板,并基于用户行为模拟和反向画像过滤构建了 SFT 数据流水线。其次,我们在强化学习(RL)阶段利用基于规则的复合奖励函数,以消除长度膨胀和重复生成等常见退化现象。最后,为了恢复在 SFT 和 RL 过程中受损的通用能力,我们提出了一种新颖的自蒸馏范式,采用配备累积散度衰减(CDD)的两阶段同策略蒸馏(OPD)。通过将领域适应模型作为教师,原始基础模型作为学生,我们有效平衡了深度领域知识注入与通用智能体能力的保持。实验结果表明,KuaiRP 模型不仅在我们目标领域内的角色扮演保真度上匹配当前最先进的专有模型,而且成功恢复了通用智能体能力,同时保持了极低的部署成本。
cs.AI / 20 / 2609.11144

Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment

同日,同故事;提前一日,不同信号:金融情绪的双重有效性
Aravinthkakshan, AS, Srivastava, Laven, Nandwani, Harsh
Abstract
Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.
Chinese Translation
金融自然语言处理有一个标准工作流程:根据人工标注验证情感工具,然后信任它来提取市场信号。这假设两种评估衡量的是同一事物。我们在一个可以同时测量两者的环境中检验这一假设:一个证券集体诉讼语料库(2002-2025年),将70,500条X消息与异常股票收益联系起来,并带有一个单一标注者的人工标注黄金样本。通过一个相同的流程运行五种工具(VADER、Loughran-McDonald、FinBERT、Twitter-RoBERTa以及一个LLM标注器),我们发现构念效度与预测效度之间的关系取决于采样惯例和得分表示。在传统的特定方法采样下,人工一致性更接近于分级同日关联,而不是提前一日的领先关联。然而,在固定n的面板上,一致性在两个时间范围上具有相似的分级秩相关,而粗略排序仍然较弱。因此,基准一致性确立了语义效度,但本身并不决定预测排名。在一个17.6%为垃圾信息的对话中,消息量既不预测市场损害也不预测和解规模。
cs.AI / 21 / 2609.11146

The Oligarch Barely Steers Model Collapse in Multi-Model Ecosystems

寡头几乎无法在多模型生态系统中操控模型崩溃
Liu, Yangze, Han, Zhongyi
Abstract
AI-generated text is flowing back into the training corpora of the next generation of models. Recursive training on it drives model collapse, and recent work extends the setting to many models feeding one another -- but almost always with the market split evenly, while real generative AI is an oligopoly. Concentration raises two worries: fewer, more uniform sources may make collapse faster, and later models may be dragged toward the oligarch's output. We test both in controlled ecosystems: 13 open 1--4B models form natural ecosystems of 3 to 13 players, plus an injected probe that pushes the top share to 90%; each generation, every model's output is mixed into a shared pool by market share and every model is retrained on that pool from clean base weights, for five generations. Yet within the range we test, neither worry materializes; what emerges instead is an invariance. Making the split more unequal barely changes the speed of collapse. Destinations move even less: the share and identity knobs shift five-generation endpoints by only a few percent of the drift common to all arms -- the ecosystems collapse to nearly the same place. An extreme share paired with the strongest injected bias still does not guarantee steering, and the topic shifts it does produce leave only a faint trace on the ruler that measures collapse. What sets the speed is who supplies the pool and how readily those suppliers are carried along: with every share held fixed, swapping the members of a K=3 ecosystem changes five-generation drift by 2.8x; a share-weighted index of each member's susceptibility explains the speed differences across nineteen arms with R^2 = 0.68; and replacing half the pool with human text roughly halves drift without changing its course. Within the tested range, concentration sets neither the destination nor the pace of collapse; the pace follows whose text fills the pool.
Chinese Translation
AI生成文本正流回下一代模型的训练语料库中。对其递归训练会驱动模型崩溃,近期工作将该设定扩展到多个模型相互投喂——但几乎总是市场均匀分割,而真实的生成式AI是寡头垄断。集中度引发两个担忧:更少、更均匀的来源可能使崩溃更快,并且后续模型可能被拖向寡头的输出。我们在受控生态系统中测试这两点:13个开放的1-4B模型形成由3到13个参与者组成的自然生态系统,另有一个注入探针将最高份额推至90%;每一代,每个模型的输出按市场份额混合到一个共享池中,并且每个模型都从干净的基座权重开始在该池上重新训练,持续五代。然而,在我们测试的范围内,两种担忧都没有成为现实;取而代之的是一种不变性。使分割更不平等几乎不改变崩溃速度。终点的变化更小:份额和身份旋钮使五代终点的移动仅为所有实验组共同漂移的百分之几——生态系统崩溃到几乎相同的位置。极端份额与最强注入偏差配对仍不能保证操控,它确实产生的主题偏移在衡量崩溃的标尺上只留下微弱的痕迹。决定速度的是谁提供池以及这些提供者有多容易被带动:在保持每个份额固定时,交换一个K=3生态系统的成员会使五代漂移变化2.8倍;一个按份额加权的每个成员易感性的指数可以解释十九个实验组之间的速度差异,R^2 = 0.68;用人类文本替换一半的池大约使漂移减半,但不改变其方向。在测试范围内,集中度既不决定崩溃的终点也不决定其速度;速度取决于谁的文字填充了池。
cs.AI / 22 / 2609.11147

Autonomous Chemical Mechanistic Discovery through Agentic Reasoning and Validation

通过智能体推理与验证实现自主化学机理发现
Li, Dong, Mi, Sixuan, Ye, Zihao, Xiong, Huan, XU, Tao, Zhu, Tong, Zhang, Aijia, Gao, Junqi, Zhang, Kaiyan, Wang, Shijie, Zhou, Bowen, Li, Yuqiang, Qi, Biqing
Abstract
Unraveling reaction mechanisms is central to modern chemistry, yet automating these investigations remains challenging because computational workflows still rely heavily on expert intervention. Here we introduce ARCHE, an autonomous agentic system that integrates a general-purpose reasoning model, a domain-specialized computational chemistry model, and a structured tool registry to transform mechanistic inquiry into a scalable, self-validating process. ARCHE interprets scientific questions, generates and prioritizes mechanistic hypotheses, orchestrates computational workflows, and iteratively refines conclusions based on computed evidence within a closed loop. We validate its capabilities across three increasingly demanding scenarios: reconstructing stereocontrolling transition states and validating the corresponding reaction mechanism in a previously reported asymmetric catalytic reaction; proposing and validating a plausible radical pathway through iterative hypothesis refinement for a recently discovered but unpublished $\alpha$-iodoboronate C-I cleavage reaction; and identifying a chemically interpretable descriptor that governs selectivity in nickel-catalysed migratory cross-coupling reactions. By coupling agentic reasoning with rigorous computational validation, ARCHE advances autonomous mechanistic discovery and establishes a foundation for broader machine-assisted chemical research. The code for ARCHE is publicly available at https://github.com/JetAstra/Arche-Harness.
Chinese Translation
揭示反应机理是现代化学的核心,然而实现这些研究的自动化仍然具有挑战性,因为计算工作流仍严重依赖专家干预。在此,我们介绍 ARCHE,一种自主智能体系统,它集成了一个通用推理模型、一个领域专用计算化学模型以及一个结构化工具注册表,从而将机理探究转化为可扩展、可自我验证的过程。ARCHE 解释科学问题,生成并优先排序机理假设,编排计算工作流,并在闭环中基于计算证据迭代完善结论。我们在三个要求逐渐提高的场景中验证了其能力:重建立体控制过渡态并验证先前报道的不对称催化反应中相应的反应机理;针对最近发现但尚未发表的 α-碘代硼酸酯 C-I 裂解反应,通过迭代假设精炼提出并验证一条可能的自由基途径;以及识别出一个控制镍催化迁移交叉偶联反应选择性的化学可解释描述符。通过将智能体推理与严格的计算验证相结合,ARCHE 推进了自主机理发现,并为更广泛的机器辅助化学研究奠定了基础。ARCHE 的代码已在 https://github.com/JetAstra/Arche-Harness 公开。
cs.AI / 23 / 2609.11155

DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat

DRG-MAPPO:用于协同空战的分层动态角色图多智能体强化学习
Liu, Junlin, Li, Chengwei, Gao, Yang, Chang, Hui, Zhang, Xinchen, Zhao, Zhijun, Zhao, Hao
Abstract
Multi-Agent Reinforcement Learning (MARL) has emerged as a pivotal paradigm for complex decision-making in autonomous systems and air combat. While MARL has demonstrated significant potential in air combat, achieving sophisticated tactical coordination remains a non-trivial challenge. This difficulty is largely attributed to two primary limitations: (1) the absence of structured relational modeling hinders agents from capturing complex, time-varying interactions among battlefield entities; and (2) conventional flat architectures often lack the capability to explicitly model tactical roles, leading to ambiguous task allocation in highly dynamic environments. To address these challenges, we propose Hierarchical Dynamic Role-Graph Multi-Agent Proximal Policy Optimization (DRG-MAPPO), a novel MARL framework that integrates graph-based relational modeling with dynamic role assignment. Specifically, DRG-MAPPO constructs a graph-based representation of battlefield interactions and leverages graph attention mechanisms to extract critical relational features among allies, enemies, and threats. Subsequently, a high-level policy employs a dynamic role assignment mechanism to determine tactical responsibilities (e.g., ``leader'' and ``supporter''). Conditioned on these roles and encoded graph-relational features, a low-level policy executes discrete maneuver actions, facilitating the joint optimization of tactical strategy and collaborative execution. Furthermore, a target-priority auxiliary task is designed to foster the emergence of behaviors such as focus-fire. Experimental results demonstrate that DRG-MAPPO achieves a state-of-the-art win rate of 87%, suggesting that our framework effectively balances relational modeling, interpretability, and optimization stability for cooperative air combat.
Chinese Translation
多智能体强化学习(MARL)已成为自主系统和空战中复杂决策的关键范式。尽管 MARL 在空战中展现出巨大潜力,但实现复杂的战术协同仍然是一个艰巨的挑战。这一困难主要归因于两个主要限制:(1) 缺乏结构化的关系建模阻碍了智能体捕捉战场实体之间复杂的、时变的交互;(2) 传统的扁平架构通常缺乏显式建模战术角色的能力,导致在高度动态环境中任务分配模糊。为了解决这些挑战,我们提出了分层动态角色图多智能体近端策略优化(DRG-MAPPO),一种新颖的 MARL 框架,它将基于图的关系建模与动态角色分配相结合。具体来说,DRG-MAPPO 构建了战场交互的图表示,并利用图注意力机制提取友军、敌军和威胁之间的关键关系特征。随后,高层策略采用动态角色分配机制来确定战术职责(例如,“领导者”和“支援者”)。基于这些角色和编码的图关系特征,低层策略执行离散的机动动作,促进战术策略和协同执行的联合优化。此外,设计了一个目标优先级辅助任务,以促进诸如集中火力等行为的涌现。实验结果表明,DRG-MAPPO 达到了 87% 的最先进胜率,表明我们的框架有效地平衡了协同空战的关系建模、可解释性和优化稳定性。
cs.AI / 24 / 2609.11170

Breaking Predictions Is Not Enough: Specified-Foil Counterfactuals for Temporal Graphs

打破预测还不够:时序图的指定替代反事实
Yu, Minwoo, Ha, Young-guk
Abstract
Temporal graph counterfactual explanations typically change past events to change or invalidate an original prediction, while leaving its replacement unspecified. Yet a user facing a predicted outcome often asks which past conditions would make a particular alternative occur instead. We formulate this destination-specific question as the Specified-Foil Counterfactual: given an original prediction A and a foil B fixed before search, find a low-cost past-event intervention under which the same predictor selects B as top-ranked. Our trace-guided intervention search contrasts the completed execution of A with a reconstructed incomplete execution of B, maps their difference to DELETE, INSERT, REWIRE, RELABEL, and SHIFT operations, and verifies B through exact replay. We instantiate this principle with LiFTER on continuous-time dynamic graphs and TLogic on temporal knowledge graphs. On CTDGs, the method retains 85.7-93.6% of black-box greedy successes while reducing predictor evaluations by 75.0-80.0%; on TKGs, it reaches the specified foil in 74.8% of 600 comparisons. Executable traces thereby become computational structures for constructing conditions of unselected alternatives, rather than records used only to explain predictions already made.
Chinese Translation
时序图反事实解释通常改变过去事件以改变或使原始预测失效,同时不指定其替代结果。然而,面对预测结果的用户常常会问,哪些过去条件会使某个特定的替代结果发生。我们将这种目标特定的问题形式化为指定替代反事实:给定原始预测A和搜索前固定的替代B,找到一个低成本的过去事件干预,使得同一预测器将B选为最高排名。我们的轨迹引导的干预搜索将A的完整执行与B的重构的不完整执行进行对比,将它们的差异映射为DELETE、INSERT、REWIRE、RELABEL和SHIFT操作,并通过精确重放验证B。我们在连续时间动态图上使用LiFTER,在时序知识图谱上使用TLogic来实例化这一原则。在CTDG上,该方法保留了85.7-93.6%的黑盒贪婪成功,同时将预测器评估减少了75.0-80.0%;在TKG上,它在600次比较中达到指定替代的74.8%。因此,可执行轨迹成为构建未选替代结果条件的计算结构,而不仅仅是用于解释已做出的预测的记录。
cs.AI / 25 / 2609.11176

Debate-to-Skill: Capability-Bound Process Supervision for Industrial Query-to-Agent Annotation

Debate-to-Skill:面向工业级 Query-to-Agent 标注的能力边界过程监督
Zhang, Shiyu, Cheng, Leisheng, Li, Huifu
Abstract
Industrial query-to-agent matching fails when topical relevance is mistaken for executable capability, especially on long-tail and boundary-sensitive requests. We formulate annotation as \emph{capability-bound process supervision} and instantiate it with Debate-to-Skill, which uses reusable decision principles, structured deliberation, verifier-based verdict extraction, and disagreement-driven refinement. On an industrial Query2Agent benchmark, we compare Debate-to-Skill with direct-label supervision, reasoning-SFT, and structural ablations. The results test whether gains come from supervising the capability-critical decision process itself, especially on grey-zone cases where semantic relatedness and executable capability diverge.
Chinese Translation
在工业级查询到智能体匹配中,当主题相关性被误认为可执行能力时,匹配就会失败,尤其是在长尾和边界敏感请求上。我们将标注形式化为能力边界过程监督,并用 Debate-to-Skill 实例化,该方法使用可复用的决策原则、结构化审议、基于验证器的判定提取以及分歧驱动的改进。在工业级 Query2Agent 基准上,我们将 Debate-to-Skill 与直接标签监督、推理-SFT 和结构性消融进行比较。结果检验了收益是否来自对能力关键决策过程本身的监督,尤其是在语义相关性与可执行能力发生偏离的灰色地带案例上。
cs.AI / 26 / 2609.11180

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

SemVerBench:基准测试LLM对版本约束解析语义的理解
Chen, Qibai, Liu, Zeming
Abstract
Large language model (LLM) coding agents constantly decide whether a version satisfies a constraint such as ^1.2.3 or >=2.0,<3, yet their grasp of version-constraint semantics has never been measured directly. We introduce SemVerBench, the first benchmark of LLM version-constraint resolution semantics across three ecosystems (npm, PEP 440, Cargo): 240 machine-checkable items with unique answers, built author-neutrally from four balanced sources (each ecosystem's official test suite plus three frontier LLM proposers) and labeled by a non-circular two-implementation oracle. Evaluating six frontier models, we find systematic, predictable per-mechanism blind spots: a partial-comparator carry rule (>1.2 means >=1.3.0) traps every model on Cargo (near 60%), and although standard PEP 440 prefix matching is universal, on zero-pad/post-release corner cases GPT-5.1 collapses (0/26) while Claude stays at 97-100% (verified on a 67-item oracle-validated set). Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar). The failures look more like an activation/application gap than a knowledge gap: injecting the rule or a light correct hint recovers most errors, whereas interval decomposition does not, and models are at ceiling on the basic forms of the same rules. An author-stratified analysis finds no statistically significant self-favoritism. Because the task is verifiable and a free, 100%-correct resolver exists, tool delegation reaches ~100%: coding agents should delegate version resolution to a resolver rather than reason about versions in-head.
Chinese Translation
大型语言模型(LLM)编码代理不断判断某个版本是否满足诸如 ^1.2.3 或 >=2.0,<3 的约束,然而它们对版本约束语义的掌握从未被直接测量。我们引入了 SemVerBench,这是首个跨三个生态系统(npm、PEP 440、Cargo)的 LLM 版本约束解析语义基准:包含 240 个机器可检查且答案唯一的项目,以作者中立的方式从四个平衡来源(每个生态系统的官方测试套件加上三个前沿 LLM 提议者)构建,并由一个非循环的双实现预言机标注。在评估六个前沿模型时,我们发现系统性的、可预测的逐机制盲点:一个部分比较器进位规则(>1.2 意味着 >=1.3.0)使每个模型在 Cargo 上陷入困境(接近 60%),并且尽管标准的 PEP 440 前缀匹配是普遍的,但在零填充/后发布边角案例上,GPT-5.1 崩溃(0/26),而 Claude 保持在 97-100%(在一个 67 项预言机验证集上验证)。Opus 显著优于所有其他模型,Sonnet 优于 OpenAI 模型(McNemar 检验)。这些失败看起来更像是激活/应用差距而非知识差距:注入规则或轻微的正确提示可以恢复大多数错误,而区间分解则不能,并且模型在相同规则的基本形式上达到上限。一项按作者分层分析发现没有统计上显著的自我偏袒。由于该任务是可验证的,并且存在一个免费的、100% 正确的解析器,工具委托达到约 100%:编码代理应将版本解析委托给解析器,而不是在头脑中推断版本。
cs.AI / 27 / 2609.11185

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment

LLMs能否遵循医学专家逻辑?偏倚风险评估中分层逻辑一致性的基准测试
Huang, Jiayu, Tang, Zichen, Ling, Qianhui, Kuang, Zemin, E, Haihong
Abstract
Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logic, comprising 860 randomized controlled trials (RCTs) and 14,820 queries. It evaluates models under the Hierarchical Logical Consistency (HLC) framework across four dimensions: Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness. Experiments on 10 state-of-the-art LLMs reveal a catastrophic Error Compounding Effect: despite the top model reaching 98.88% Atomic Consistency, its end-to-end consistency collapses to 45.13%, with several open-weight architectures plummeting to nearly 0%. We further uncover a systematic evidence-reasoning gap: even when models retrieve high-quality evidence, they fail to deduce correct outcomes in 18.63-40.05% of cases, while Blind Guess Rates reach 48.28%. LogiMed-RoB demonstrates that high outcome accuracy can conceal critical reasoning flaws, underscoring the necessity of white-box logical verification for clinical deployment.
Chinese Translation
循证医学要求严格的逻辑一致性,然而当前对大语言模型(LLMs)的评估优先考虑表面标签匹配而非真正的推理。我们提出了LogiMed-RoB,一个基于Cochrane偏倚风险(RoB)2.0专家逻辑的基准,包含860项随机对照试验(RCTs)和14,820个查询。它在分层逻辑一致性(HLC)框架下从四个维度评估模型:原子一致性、领域一致性、聚合一致性和证据忠实性。在10个最先进的LLMs上的实验揭示了一个灾难性的错误累积效应:尽管顶级模型达到了98.88%的原子一致性,但其端到端一致性暴跌至45.13%,而几个开放权重架构则骤降至接近0%。我们进一步揭示了一个系统性的证据-推理差距:即使模型检索到高质量证据,它们在18.63-40.05%的情况下也无法推断出正确结果,而盲猜率高达48.28%。LogiMed-RoB表明,高结果准确性可能掩盖关键的推理缺陷,强调了临床部署中白盒逻辑验证的必要性。
cs.AI / 28 / 2609.11190

Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-Commerce

智能体搜索份额(Agentic Share-of-Search):一个用于LLM介导的电子商务中竞争决策的多智能体AI系统
Chowdhury, Spandan Ghose
Abstract
AI shopping assistants increasingly redirect consumer discovery, creating an urgent need for tools that support seller-side competitive decision-making. We present a multi-agent AI system that automates competitive visibility measurement and root cause diagnosis in LLM-mediated ecommerce. The system introduces Agentic Share-of-Search (ASoS) as the decision target, deploys query agents across leading AI platforms, and uses a ReAct-based diagnostic agent to recommend prioritized merchandising interventions. A 100-trial ablation study, presented as a feasibility evaluation of this prototype, shows the agent recovers the ablated signal in 39% of trials (95% CI: 30.0% - 48.8%, 5.5x over chance), rising to 63.9% among high-correlation ablations.
Chinese Translation
AI购物助手日益改变消费者的发现路径,迫切需要支持卖方竞争决策的工具。我们提出了一个多智能体AI系统,可在LLM介导的电子商务中自动进行竞争可见性测量和根本原因诊断。该系统引入了智能体搜索份额(Agentic Share-of-Search,ASoS)作为决策目标,在领先的AI平台上部署查询智能体,并使用基于ReAct的诊断智能体来推荐优先的商品营销干预措施。一项100次试验的消融研究(作为该原型的可行性评估)表明,该智能体在39%的试验中恢复了消融信号(95% CI:30.0% - 48.8%,是随机概率的5.5倍),在高相关性消融中这一比例上升至63.9%。
cs.AI / 29 / 2609.11199

An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning

基于NLP和机器学习的AI驱动文化感知聊天机器人:用于巴基斯坦大学生的压力检测与健康支持
Bashir, Muhammad Fahad, Afzal, Muhammad
Abstract
With the existing digital mental health tools specifically developed for Western settings, Pakistani students are exposed to a uniquely compounded stress situation in their university that includes academic, financial, familial, and relational stressors, which have become a serious concern for academic and psychological development of students in Pakistani universities. This paper introduces a new, AI-driven and culturally sensitive stress detection and wellness support system that is tailored to the context of Pakistani university students. The system is based on a machine learning model called Random Forest which is trained using a validated student stress data set of 1100 responses on 20 features from psychological, physiological, academic, environmental and social aspects, with an accuracy of 89.09% and a macro F1-score of 0.89, in three stress severity levels. The classification outputs are passed on to an open-source large language model through OpenRouter API, where an appropriately crafted system prompt, culturally aware, gives the model a conversation about wellness, in English, Urdu and Roman Urdu. The second most predictive stress factor in this population identified by feature importance analysis was teacher-student relationship, which is a culturally important stress factor highlighting the need for region-aware mental health systems. Future research will involve primary data collection from students at various academic levels of Pakistani Universities with the validated DASS-21 instrument focusing on the students who are moving from FSc to undergraduate studies, which is a time of being psychologically vulnerable which is under-researched.
Chinese Translation
由于现有的数字心理健康工具是专门为西方环境开发的,巴基斯坦大学生在大学中面临着独特的复合压力情境,包括学业、经济、家庭和人际关系压力源,这些压力已成为巴基斯坦大学生学业和心理发展的严重问题。本文介绍了一种新的、人工智能驱动的、文化敏感的压力检测和健康支持系统,该系统针对巴基斯坦大学生的背景而定制。该系统基于一种称为随机森林(Random Forest)的机器学习模型,该模型使用经过验证的学生压力数据集进行训练,该数据集包含来自心理、生理、学业、环境和社会方面的20个特征的1100个响应,在三个压力严重程度级别上,准确率为89.09%,宏F1分数为0.89。分类输出通过OpenRouter API传递给开源大型语言模型,其中精心设计的、具有文化意识的系统提示使模型能够用英语、乌尔都语和罗马乌尔都语进行关于健康的对话。通过特征重要性分析确定,该人群中第二最具预测性的压力因素是师生关系,这是一个文化上重要的压力因素,突出了对区域感知心理健康系统的需求。未来的研究将涉及使用经过验证的DASS-21工具从巴基斯坦大学不同学业水平的学生中收集原始数据,重点关注从FSc过渡到本科学习的学生,这是一个心理脆弱的时期,但研究不足。
cs.AI / 30 / 2609.11206

CryptoL: Towards Scale Dominance and Physics Constraints Mitigation in Financial Multivariate Time Series Forecasting

CryptoL:面向金融多变量时间序列预测中的尺度主导与物理约束缓解
Taheri, Yalda, Heydari, Mohammad Hassan, Rasooli, Armon, Amirshahkarami, Maryam, Mahdavi, Mohammad Ebrahim, Karshenas, Hossein
Abstract
Cryptocurrency forecasting presents a distinctive combination of extreme cross-asset scale heterogeneity, non-stationary dynamics, and structural dependencies among Open, High, Low, and Close (OHLC) variables. We present CryptoL, a unified framework designed to address these challenges within multivariate time-series forecasting. CryptoL evaluates forecasting error in context-normalized coordinates within the RevIN pipeline, preventing inverse normalization from introducing an additional squared-scale weighting into the MSE objective. We formally characterize this effect through the empirical risk and parameter-gradient geometry, establishing the conditions under which large-scale assets can disproportionately influence shared-model optimization. Beyond loss-space normalization, CryptoL examines channel-independent and channel-dependent normalization for OHLC data, showing that a shared channel-dependent affine transformation preserves candle-order relations that independent channel transformations need not preserve. The framework further incorporates scale-adaptive numerical stabilization to reduce distortions caused by a fixed normalization constant across assets spanning many orders of magnitude, together with a soft feasibility loss that penalizes violations of the defining OHLC inequalities. Experiments across heterogeneous cryptocurrency assets evaluate these components through controlled ablations and demonstrate improvements in forecasting accuracy, training stability, and the frequency of financially valid OHLC predictions relative to the considered baselines. CryptoL therefore provides an integrated approach to scale-balanced optimization, structure-preserving normalization, numerical stabilization, and constraint-aware cryptocurrency forecasting.
Chinese Translation
加密货币预测呈现出一种独特的组合:极端的跨资产尺度异质性、非平稳动态,以及开盘价、最高价、最低价和收盘价(OHLC)变量之间的结构性依赖。我们提出 CryptoL,一个统一的框架,旨在解决多变量时间序列预测中的这些挑战。CryptoL 在 RevIN 流程内以上下文归一化坐标评估预测误差,防止逆归一化在 MSE 目标中引入额外的平方尺度加权。我们通过经验风险和参数梯度几何形式化刻画了这一效应,建立了大规模资产可能不成比例地影响共享模型优化的条件。除了损失空间归一化,CryptoL 还考察了 OHLC 数据的通道独立和通道依赖归一化,表明共享的通道依赖仿射变换保持了独立通道变换不必保持的蜡烛顺序关系。该框架进一步结合了尺度自适应数值稳定化,以减少跨多个数量级的资产使用固定归一化常数所造成的扭曲,并引入软可行性损失,以惩罚对定义性 OHLC 不等式的违反。在异构加密货币资产上的实验通过受控消融评估了这些组件,并相对于所考虑的基线展示了在预测准确性、训练稳定性以及财务有效的 OHLC 预测频率方面的改进。因此,CryptoL 为尺度平衡优化、结构保持归一化、数值稳定化和约束感知的加密货币预测提供了一种集成方法。
cs.AI / 31 / 2609.11231

A Voice-Interactive Multi-Agent System for Smart Operating Rooms: Architecture Design and Key Technologies

面向智能手术室的语音交互多智能体系统:架构设计与关键技术
Zhou, Tianxiang
Abstract
This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms based on large language models (LLMs). The system achieves natural language understanding, device control, intraoperative recording, and surgical report generation through a layered architecture comprising a voice interaction pipeline (wake, ASR, turn detection, agent reasoning, TTS) and an agent core (skill registry, task planner, device manager). Three key technologies are investigated: (1) KV Cache prefix warming for low-latency inference, reducing recomputation overhead from approximately 500 ms to tens of milliseconds via byte-level Longest Common Prefix reuse; (2) streaming partial JSON parsing with early parallel task execution, reducing end-to-end latency by approximately 30%; and (3) progressive skill prompt disclosure, which dynamically filters system prompts based on user role, connected devices, and surgical phase to maximize information density within limited context windows. The system is implemented using the Qwen3-27B model with llama.cpp/sglang inference engines. Experimental analysis demonstrates effective operation within a 16,384-token context limit and multi-device parallel control response times meeting OR real-time requirements.
Chinese Translation
本文提出了SurgicalRoomAgent,一种基于大语言模型(LLMs)的面向智能手术室的语音交互多智能体系统。该系统通过分层架构实现自然语言理解、设备控制、术中记录和手术报告生成,该架构包括语音交互流水线(唤醒、ASR、轮次检测、智能体推理、TTS)和智能体核心(技能注册表、任务规划器、设备管理器)。研究了三项关键技术:(1)用于低延迟推理的KV Cache前缀预热,通过字节级最长公共前缀(Longest Common Prefix)复用,将重计算开销从约500 ms降低至数十毫秒;(2)流式部分JSON解析与提前并行任务执行,将端到端延迟降低约30%;(3)渐进式技能提示词披露,根据用户角色、连接设备和手术阶段动态过滤系统提示词,以在有限上下文窗口内最大化信息密度。该系统使用Qwen3-27B模型,并采用llama.cpp/sglang推理引擎实现。实验分析表明,该系统可在16,384 token的上下文限制内有效运行,且多设备并行控制响应时间满足手术室(OR)实时性要求。
cs.AI / 32 / 2609.11234

NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment

NovGauge:一个用于诊断LLMs在论文新颖性评估能力的细粒度基准
Zhang, Guoqiang, Tan, Kexin, Zhang, Ming, Ju, Li, Jing, Wenqing, Yue, Zhonghan, Chen, Jiayi, Wu, Shiqiang, Liu, Shaofan, Zhang, Yue, Ying, Yuankai, Shi, Yang, Gui, Tao, Zhang, Qi, Huang, Xuanjing
Abstract
Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a single holistic score, making it difficult to diagnose which dimension a model misjudges or whether its evidence is faithful. We present NovGauge, a human-anchored benchmark for fine-grained novelty assessment diagnosis. The benchmark contains 619 paper pairs and 50 multi-paper sets, drawn from two expert sources: ICLR reviewer overlap claims and survey co-citations. Instances are independently labeled along three dimensions: task, problem, and method, capturing application goals, technical challenges, and solution approaches. We propose a cascading diagnostic pipeline that verifies per-dimension correctness, evidence grounding, and logical support. Evaluation of 18 LLMs shows hallucination rates ranging from 0% to 39% across dimensions, and among non-hallucinated correct-positive judgments, over 70% cite evidence fails to logically support the stated reason. The best-performing model, GPT-5.5, achieves 43-72% Verified F1 across dimensions, while most models retain less than half of their raw F1 after faithfulness verification. These results suggest that current LLMs remain far from reliable scientific novelty assessment, particularly when correctness is conditioned on faithful evidence grounding.
Chinese Translation
大语言模型(LLMs)越来越多地用于主要AI会议的同行评审,但新颖性仍然是一个持续的薄弱环节。现有的基准将新颖性评估为单一的整体分数,难以诊断模型在哪个维度判断错误,或者其证据是否忠实。我们提出了NovGauge,一个基于人类判断的细粒度新颖性评估诊断基准。该基准包含619个论文对和50个多论文集合,来自两个专家来源:ICLR审稿人重叠声明和综述共同引用。实例在三个维度上独立标注:任务、问题和方法,分别捕捉应用目标、技术挑战和解决方案。我们提出了一个级联诊断流程,验证每个维度的正确性、证据基础和逻辑支持。对18个LLMs的评估显示,各维度的幻觉率在0%到39%之间,并且在非幻觉的正确阳性判断中,超过70%的引用证据未能逻辑上支持所陈述的理由。表现最好的模型GPT-5.5在各维度上达到43-72%的Verified F1,而大多数模型在忠实性验证后保留的原始F1不到一半。这些结果表明,当前的LLMs距离可靠的科学新颖性评估还很远,特别是当正确性以忠实的证据基础为条件时。
cs.AI / 33 / 2609.11243

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Sci-MMR:多模态智能体中多步基于证据的科学推理基准测试
Li, Jiaqiang, Yang, Yajie, Xi, Zhiheng, Chen, Jiadong, Zhou, Enyu, Jin, Senjie, Nan, Yang, Zhang, Jiazheng, Wang, Han, Li, Yanxin, Zhu, Dingwei, Deng, Bicheng, Wang, Yuhui, Zheng, Xiang, Zhang, Qi, Bai, Lei, Ma, Xingjun, Gui, Tao
Abstract
Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions. Sci-MMR comprises 235 multi-hop reasoning tasks spanning four scientific disciplines, with an average of nine figure panels per task. Evaluating eight frontier multimodal models, we find that answer accuracy consistently exceeds complete-evidence recovery rate by more than 20%, revealing a substantial gap that answer-only evaluation is structurally unable to capture. Through controlled interventions, we identify two fundamental bottlenecks. First, evidence acquisition: models struggle to extract complete structured evidence from scientific figures, accounting for 57.2% of failures. While cropping tools yield modest gains (+4.5 points), providing gold evidence improves accuracy by up to 37.0 points, indicating difficulty in assembling complete multi-region evidence. Second, evidence integration: models struggle to translate available evidence into correct conclusions, accounting for 31.8% of failures, while even with gold evidence the strongest model achieves only 69.1% accuracy on the hardest tasks. These findings indicate that current answer-centric benchmarks substantially overestimate the evidence-grounded reasoning capabilities of multimodal research agents
Chinese Translation
自主研究智能体日益被期望能够检索文献、分析实验证据并生成科学假设。这些能力需要多步基于证据的推理,即在得出结论之前逐步获取、整合和验证证据。然而,现有的多模态基准测试主要评估最终答案的准确性,使得预测是否真正由可追溯的科学证据支持仍然悬而未决。我们提出Sci-MMR,一个基于结构化论证图的多步基于证据的科学推理基准,该图将科学主张、引用支撑的知识、视觉证据和支持区域联系起来。Sci-MMR包含235个多跳推理任务,涵盖四个科学学科,平均每个任务有九个图面板。通过评估八个前沿多模态模型,我们发现答案准确率始终比完整证据恢复率高出20%以上,揭示了一个仅依赖答案的评估在结构上无法捕捉的巨大差距。通过受控干预,我们确定了两个基本瓶颈。第一,证据获取:模型难以从科学图表中提取完整的结构化证据,占失败的57.2%。虽然裁剪工具带来了适度的增益(+4.5个百分点),但提供黄金证据可将准确率提高最多37.0个百分点,表明在组装完整的多区域证据方面存在困难。第二,证据整合:模型难以将可用的证据转化为正确的结论,占失败的31.8%,而即使有黄金证据,最强的模型在最难的任务上也仅达到69.1%的准确率。这些发现表明,当前以答案为中心的基准测试大大高估了多模态研究智能体的基于证据的推理能力。
cs.AI / 34 / 2609.11262

AI-Powered Flare Combustion Efficiency Estimation

AI驱动的火炬燃烧效率估算
Azam, Afeefa, Ganapathi, Iyyakutti Iyappan, Abdelhafez, Fares Ossama, Velayudhan, Divya, Habtie, Maregu Assefa, Karki, Hamad, Awadhi, Khalid Yousef Al, Werghi, Naoufel
Abstract
Achieving high combustion efficiency in flare stacks is crucial for adhering to regulatory standards and controlling the release of hydrocarbons into the environment. Traditional instruments like gas analyzers and hyperspectral cameras are expensive, fragile, and require frequent calibration, which makes them impractical for remote or budget constrained industrial sites. We propose an innovative solution that combines a lightweight vision-language encoder with a compact multi-layer perceptron to predict combustion efficiency directly from low-cost thermal video footage. The fully trained model is integrated into an easy-to-deploy graphical user interface. This interface overlays predicted combustion efficiency values on each video frame, displays real-time trends in combustion efficiency, shows the distribution of combustion efficiency across all frames in the video, and allows users to export CSV reports. Over a six-month period, the system achieved 99% uptime and required less than 15 minutes of maintenance per week.
Chinese Translation
在火炬烟囱中实现高燃烧效率对于遵守监管标准和控制烃类向环境排放至关重要。气体分析仪和高光谱相机等传统仪器价格昂贵、易损坏,并且需要频繁校准,这使其在偏远或预算受限的工业场所不切实际。我们提出了一种创新解决方案,将轻量级视觉-语言编码器与紧凑的多层感知机相结合,可直接从低成本热成像视频片段中预测燃烧效率。完全训练后的模型被集成到一个易于部署的图形用户界面中。该界面将预测的燃烧效率值叠加在每一视频帧上,显示燃烧效率的实时趋势,展示视频中所有帧的燃烧效率分布,并允许用户导出CSV报告。在六个月内,系统实现了99%的正常运行时间,并且每周所需维护时间少于15分钟。
cs.AI / 35 / 2609.11277

Predicting Train Delays in Finland Using Machine Learning and Weather Data

基于机器学习和天气数据预测芬兰列车延误
Borin, Vinicius Pozzobon, Sant'Ana, Jean Michel de Souza, Mahmood, Nurul Huda
Abstract
Reliable railway operations depend increasingly on real-time environmental intelligence delivered through wireless sensor infrastructures, a capability that 6G networks will substantially enhance through integrated sensing and edge computing. Adverse weather, particularly in Arctic regions with extreme temperatures and heavy precipitation, remains a leading cause of train delays, yet most prediction approaches rely on raw meteorological inputs without exploiting domain-informed feature engineering. This paper investigates machine learning for train delay prediction using the Finland Integrated Train-Weather (FI-TW) dataset, which fuses railway operational records with observations from the Finnish Meteorological Institute's nationwide sensor network of approximately 200 stations communicating over wireless links. We evaluate three feature configurations using XGBoost at Oulu central station (101,146 observations): full weather features, instant weather observations only, and derived weather category scenarios. The category-based approach, employing hierarchical classifications such as Blizzard, Heavy Snow, and Extreme Cold, achieved an R^2 of 0.78, root mean squared error of 8.5 minutes, and mean absolute error of 3.7 minutes, representing an 11% R^2 improvement and 10% error reduction over alternative configurations. These results demonstrate that compact, domain-informed features derived from sensor streams outperform raw meteorological observations, offering bandwidth-efficient representations suitable for edge deployment over current and emerging wireless infrastructures.
Chinese Translation
可靠的铁路运营日益依赖于通过无线传感器基础设施提供的实时环境智能,而6G网络将通过集成感知与边缘计算显著增强这一能力。恶劣天气,尤其是在具有极端温度和强降水的北极地区,仍是导致列车延误的主要原因之一,但大多数预测方法依赖于原始气象输入,未能利用领域知识驱动的特征工程。本文利用芬兰列车-天气集成(FI-TW)数据集研究基于机器学习的列车延误预测,该数据集将铁路运营记录与芬兰气象研究所约200个通过无线链路通信的全国性传感器网络站点的观测数据相融合。我们在奥卢中央车站(101,146条观测)使用XGBoost评估了三种特征配置:完整天气特征、仅瞬时天气观测以及衍生的天气类别情景。基于类别的方法采用暴风雪、大雪和极端寒冷等分层分类,实现了0.78的R²、8.5分钟的均方根误差和3.7分钟的平均绝对误差,相较于其他配置,R²提升11%,误差降低10%。这些结果表明,从传感器流中提取的紧凑、领域知识驱动的特征优于原始气象观测,提供了带宽高效的表示,适合在当前和新兴无线基础设施上进行边缘部署。
cs.AI / 36 / 2609.11281

Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 1

仿生学习与决策:基于概率存内计算硬件(第一部分)
Dalgaty, Thomas, Kawasaki, Eiji, de Prado, Miguel, Vyas, Devendra, Salvatori, Tommaso
Abstract
Learning and decision-making in animals are often modeled as Bayesian processes, where sensory evidence is integrated with prior beliefs to guide behavior in the face of uncertainty. But what are the inherent neural dynamics that give rise to this ability, and how could they be replicated in computing systems? This abstract discusses a biologically grounded framework in which noisy neural and synaptic dynamics perform inference and learning via stochastic sampling from an internal energy function, capturing uncertainty over latent states and model parameters through neural and synaptic variability, respectively. This enables approaches such as predictive coding networks to account for epistemic uncertainty via Markov chain Monte Carlo sampling. Drawing a parallel between intrinsic noise in biological systems and electrical noise in emerging probabilistic analogue memory technologies, we highlight how analogue in-memory computing hardware naturally emerges as the solution for massively scalable and energy-efficient probabilistic inference.
Chinese Translation
动物的学习与决策通常被建模为贝叶斯过程,其中感觉证据与先验信念相结合,以在不确定性下指导行为。但是,产生这种能力的内在神经动力学是什么?它们如何在计算系统中被复制?本摘要讨论了一个基于生物学的框架,其中噪声神经和突触动力学通过从内部能量函数进行随机采样来执行推理和学习,分别通过神经和突触变异性捕获潜在状态和模型参数的不确定性。这使得诸如预测编码网络之类的方法能够通过马尔可夫链蒙特卡洛采样来解释认知不确定性。通过将生物系统中的内在噪声与新兴概率模拟内存技术中的电噪声进行类比,我们强调模拟存内计算硬件如何自然地成为大规模可扩展且节能的概率推理的解决方案。
cs.AI / 37 / 2609.11282

When Does Text Inform? Benchmarking Information-Theoretic Metrics for Multimodal Time-Series Forecasting

文本何时提供信息?多模态时间序列预测的信息论度量基准测试
Andrews, Emma, Mengaldo, Gianmarco
Abstract
Multimodal forecasting models that combine time series with text annotations promise richer prediction through textual context, but how do we know whether a text annotation meaningfully contributes to the forecasters prediction? This is an information-theoretic question, but to evaluate whether information-theoretic metrics can reliably measure the predictive value an annotation provides, a ground truth benchmark is needed, and none currently exist. We create a synthetic time series signal with annotations in three categories: semantically correct, incorrect, and irrelevant. Because the data generation process is fully controlled, ground-truth information content is known exactly, enabling principled evaluation of six complementary mutual information estimators (KSG, MINE, InfoNCE, CCA, PID and V-information). We show that all six estimators identify correct annotations as most informative, and are able to audit the quality of mixed text corpora, choosing the annotations that result in the best downstream forecasting results without the need for model training. Our benchmark identifies limitations of each estimator, and these are validated on seven real-world datasets, which show how estimator performance differs on weak signals. Finally, we establish practical rules for implementing these metrics for annotation auditing and fusion selection.
Chinese Translation
结合时间序列与文本标注的多模态预测模型有望通过文本上下文实现更丰富的预测,但我们如何知道文本标注是否对预测器的预测有有意义的贡献?这是一个信息论问题,但要评估信息论度量是否能可靠地衡量标注所提供的预测价值,需要一个真值基准,而目前不存在这样的基准。我们创建了一个合成时间序列信号,带有三类标注:语义正确、错误和不相关。由于数据生成过程完全受控,真值信息内容精确已知,从而能够对六种互补的互信息估计器(KSG、MINE、InfoNCE、CCA、PID 和 V-information)进行有原则的评估。我们表明,所有六种估计器都能识别出正确标注为最具信息量,并且能够审计混合文本语料库的质量,选择那些能带来最佳下游预测结果的标注,而无需模型训练。我们的基准揭示了每种估计器的局限性,并在七个真实世界数据集上进行了验证,展示了估计器在弱信号上的性能差异。最后,我们建立了实施这些度量以进行标注审计和融合选择的实用规则。
cs.AI / 38 / 2609.11286

Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data

生成一致的企业:多系统业务数据的合成与无参考评估
Gruenbaum, Benjamin, Porat, Doron, Natanzon, Assaf, Zavida, Roy, Dinachi, Chen, Itzahary, Or, Niv, Omer
Abstract
Synthetic relational data is normally produced by a model trained on a real dataset, and its quality is measured as the distance to that dataset. This paper describes a generator that has no real dataset at either end. Given an industry, a company size, a business model, a set of business applications, and a random seed, it produces a complete fictional enterprise: a workforce, a customer base, sales deals, support tickets, recorded calls, chat messages, and documents, all consistent with one another. One entity graph is projected into the native formats of 66 business products, so the same customer appears in the CRM, the support desk, and the call system under one identity. Because no real counterpart exists, realism is built in from cited reference statistics and verified by reference-free measurement: a five-axis scorecard of 28 statistical checks, an adversarial detector that hunts for the marks of synthetic generation, and a set of soundness checks that include a classifier test against an independently shuffled copy of the data. Because these instruments existed before the generator was tuned, progress is measured under a fixed yardstick: over 23 generated companies, mean realism climbed from 60.3 to 99.1, the weakest company from 41.1 to 94.9, and the detector, which initially flagged 55.2% of all records, now flags none. The scores hold on a seed never used during development. A second generator builds relational databases from a list of business questions. It forces qualifying rows for each answerable question, adds controlled near misses, and computes exact labels from the finished tables. The generator runs as a hosted service at https://console.era.eon.io. A company built there to a specification is served through its simulators over MCP and REST, and the simulators are also published as container images for offline use
Chinese Translation
合成关系数据通常由在真实数据集上训练的模型生成,其质量以到该数据集的距离来衡量。本文描述了一个在两端都没有真实数据集的生成器。给定一个行业、公司规模、商业模式、一组业务应用和一个随机种子,它会产生一个完整的虚构企业:员工队伍、客户群、销售交易、支持工单、录音通话、聊天消息和文档,所有这些彼此一致。一个实体图被投影到66个业务产品的原生格式中,因此同一个客户以同一个身份出现在CRM、支持台和通话系统中。因为没有真实对应物,所以真实性是从引用的参考统计数据中构建的,并通过无参考测量来验证:一个包含28项统计检查的五轴评分卡,一个寻找合成生成痕迹的对抗性检测器,以及一组健全性检查,其中包括针对独立打乱的数据副本的分类器测试。因为这些工具在生成器调优之前就已经存在,所以进展是在固定标准下衡量的:在23家生成的公司中,平均真实性从60.3攀升到99.1,最弱的公司从41.1到94.9,而检测器最初标记了所有记录的55.2%,现在没有标记任何记录。这些分数在一个开发期间从未使用过的种子上保持。第二个生成器从一系列业务问题构建关系数据库。它为每个可回答的问题强制符合条件的行,添加受控的近似未命中,并从完成的表中计算精确标签。该生成器作为托管服务运行在https://console.era.eon.io。在那里按照规范构建的公司通过MCP和REST上的模拟器提供服务,模拟器也作为容器镜像发布以供离线使用。
cs.AI / 39 / 2609.11291

Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model

韩语27B语言模型中响应风格对齐的脱靶效应
Han, Hyojung
Abstract
We post-train Qwen3.8-27B for Korean response style -- verbosity, list and markdown usage, discourse structure and register -- and measure two behaviours the objective never targets: abstention on ambiguous social questions in KoBBQ, where the benchmark-correct answer is UNKNOWN, and unprompted disclosure in securities guidance. Both move, and the changes are expressed primarily through the model's emission policy: how often it answers and how much it says. Matched target-form controls show that answer propensity depends on the training target, not the prompt set or recipe alone. Holding prompts, recipe, data volume and serving fixed and changing only the target text, three style seeds give positive answer-rate point estimates (mean +0.82 pp) and three neutral seeds negative ones (mean -1.53 pp); the observed seed ranges do not overlap and the means differ by 2.34 pp. A length-matched arm lies between them, and a fourth arm that stays short while preserving hedging is unstable across seeds, so which feature of the form is responsible is unresolved. For absolute stereotyped exposure the decomposition into an answer-propensity term and a conditional-composition term is an algebraic identity, not a finding; its empirical content is where the movement went. Across the trained checkpoints the changes are dominated by answer propensity while the composition term stays small, and because that term is evaluated on treatment-dependent answered subsets we do not read it as evidence about latent preference. Two measurement results follow. A between-arm contrast in conditional stereotyped share does not identify a change in conditional content preference when answer status is treatment-dependent. And agreement between two rule detectors for the same construct runs from 0.44 to 0.99 depending on which checkpoint produced the text -- observable without any reference labels.
Chinese Translation
我们针对韩语回复风格(冗长程度、列表和markdown使用、话语结构和语域)对Qwen3.8-27B进行后训练,并测量目标从未针对的两种行为:在KoBBQ中对模糊社交问题的弃权(其中基准正确答案是UNKNOWN),以及在证券指导中的未经提示的披露。两者都发生了变化,且变化主要通过模型的输出策略体现:它回答的频率以及它说多少。匹配的目标形式控制表明,回答倾向取决于训练目标,而不仅仅是提示集或配方。保持提示、配方、数据量和服务固定不变,仅改变目标文本,三个风格种子给出正的答案率点估计(平均+0.82个百分点),三个中性种子给出负的(平均-1.53个百分点);观察到的种子范围不重叠,均值相差2.34个百分点。一个长度匹配的臂位于它们之间,而第四个保持简短同时保留模糊限制的臂在不同种子间不稳定,因此形式中的哪个特征起作用尚未解决。对于绝对刻板暴露,分解为回答倾向项和条件组成项是一个代数恒等式,而不是发现;其经验内容在于变化发生在哪里。在训练过的检查点上,变化主要由回答倾向主导,而组成项保持较小,并且由于该项是在依赖于处理的已回答子集上评估的,我们不将其解读为关于潜在偏好的证据。两个测量结果随之而来。当回答状态依赖于处理时,条件刻板份额的臂间对比不能识别条件内容偏好的变化。并且,同一构造的两个规则检测器之间的一致性从0.44到0.99不等,取决于哪个检查点生成了文本——无需任何参考标签即可观察到。
cs.AI / 40 / 2609.11294

Memory Compression for High-Fanout Agent Sandboxes

高扇出智能体沙箱的内存压缩
Li, Mengming, XU, Ceyu, Zhang, Qijun, Yu, Jiangnan, Sun, Xiangfeng, Mai, Haohui, Xie, Zhiyao
Abstract
High-fanout agent workloads create a growing memory bottleneck because a single task may spawn many concurrent sandbox sessions. Yet these sandboxes are far from independent: they originate from a shared template and execute related trajectories, exposing substantial template-relative and cross-sandbox memory redundancy. Conventional memory compression is poorly matched to this setting in three fundamental dimensions: how to compress, because they fail to exploit similarity across non-identical sandbox pages; what to compress, because they control page-fault overhead through conservative page selection; and when to compress, because compression is either triggered by memory pressure or performed without awareness of agent execution phases. We present AgentZip, the first memory compression system designed specifically for AI-agent sandboxes. AgentZip introduces compression mechanisms that exploit both the template-relative and cross-sandbox redundancy. It broadens the compression scope to any page with a profitable representation and shifts overhead control from compression-time page selection to restore-time prefetching. It further aligns expensive compression with LLM waiting periods to avoid interfering with foreground tool execution. Across LLM training and inference workloads, AgentZip reduces sandbox-owned memory by up to 8.7x, compared with 2.1x for the Linux configuration. Restore prefetching and agent-execution-aware scheduling reduce the slowdown of aggressive compression from as high as 3.1x to 1.40x while retaining nearly all of its memory-saving benefit.
Chinese Translation
高扇出智能体工作负载造成日益严重的内存瓶颈,因为单个任务可能生成许多并发沙箱会话。然而,这些沙箱远非彼此独立:它们源自共享模板并执行相关轨迹,暴露出大量模板相关和跨沙箱的内存冗余。传统内存压缩在三个基本维度上与这一场景很不匹配:如何压缩,因为它们未能利用非相同沙箱页面之间的相似性;压缩什么,因为它们通过保守的页面选择来控制缺页开销;以及何时压缩,因为压缩要么由内存压力触发,要么在执行时没有感知智能体执行阶段。我们提出 AgentZip,这是首个专为 AI 智能体沙箱设计的内存压缩系统。AgentZip 引入压缩机制,同时利用模板相关冗余和跨沙箱冗余。它将压缩范围扩大到任何具有可获益表示的页面,并将开销控制从压缩时的页面选择转移到恢复时的预取。它还进一步将高开销压缩与 LLM 等待期对齐,以避免干扰前台工具执行。在 LLM 训练和推理工作负载中,AgentZip 将沙箱自有内存最多降低 8.7 倍,而 Linux 配置为 2.1 倍。恢复预取和智能体执行感知调度将激进压缩的减速从高达 3.1 倍降低到 1.40 倍,同时保留其几乎所有内存节省收益。
cs.AI / 41 / 2609.11315

Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models

按推理需求路由:扩散视觉语言模型的轨迹感知解码控制
Liu, Yixiang, Xu, Zhongxing, Wang, Zhonghua, Tang, Xiaoying
Abstract
Diffusion vision-language models generate answers through iterative refinement, exposing intermediate answer trajectories that can be inspected and controlled at inference time. However, this controllability creates a reasoning-need mismatch, where a universal generation length is applied to questions with different reasoning demands. Visually closed questions may be harmed by continued refinement after a stable answer has formed, whereas reasoning-sensitive questions may be harmed by premature commitment. We formulate this problem as reasoning-budget mismatch and study it in LLaDA-V. Rather than choosing a universal generation length, our training-free controller routes each example to early commitment, baseline preservation, or reasoning-supportive decoding using trajectory signals from answer closure, commitment evidence, and representation revision pressure, without using ground-truth answers. Across answer-focused, mixed-reasoning, and CoT-sensitive benchmarks, routed control improves robustness over fixed long decoding, pure short decoding, and single-rule interventions. The gains are not explained by shorter outputs alone. Answer-closed examples often benefit from commitment, whereas CoT-sensitive examples require preserving or supporting intermediate reasoning. Taken together, these results suggest diffusion VLM decoding should route inference-time control by the state suggested by the observed trajectory instead of relying on a universal decoding length.
Chinese Translation
扩散视觉语言模型通过迭代细化生成答案,暴露出可以在推理时检查和控制的中间答案轨迹。然而,这种可控性造成了推理需求不匹配:对具有不同推理需求的问题应用了通用的生成长度。视觉封闭问题可能会在稳定答案形成后因继续细化而受损,而推理敏感问题则可能因过早确定而受损。我们将此问题表述为推理预算不匹配,并在 LLaDA-V 中研究它。我们的免训练控制器不是选择通用的生成长度,而是使用来自答案闭合、承诺证据和表示修订压力的轨迹信号,将每个示例路由到早期承诺、基线保持或推理支持解码,且不使用真实答案。在面向答案、混合推理和 CoT 敏感的基准测试中,路由控制相比固定长解码、纯短解码和单规则干预提高了鲁棒性。这些收益不能仅由更短的输出解释。答案闭合的示例通常受益于承诺,而 CoT 敏感的示例则需要保留或支持中间推理。总之,这些结果表明扩散 VLM 解码应根据观察到的轨迹所暗示的状态来路由推理时控制,而不是依赖通用的解码长度。
cs.AI / 42 / 2609.11318

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Mr.LHDR:一个面向多模态真实世界长时程深度研究智能体的基准
Guo, Minghao, Cao, Meng, Zhao, Sui, Ning, Siyu, Wang, Xin, Zhao, Haoze, Yang, Jiaxuan, Hao, Haihong, Han, Mingfei, Rong, Shunlin, Wu, Haijun, Liang, Xiaodan, Chang, Xiaojun
Abstract
Deep research agents are increasingly capable of web search, tool use, multimodal evidence analysis, and information synthesis. However, existing benchmarks mainly evaluate medium-horizon exploration and rarely test whether agents can sustain long, dependency-heavy research processes. We introduce Mr.LHDR (Multimodal real-world Long-Horizon Deep Research), a benchmark for evaluating real-world deep research over long, irreducible chains of interdependent evidence across eight categories. Each question is constructed from a hidden Node-Relation graph and requires an average of 12.1 necessary intermediate conclusions with a mean dependency depth of 10.4 before reaching a short, unique, and verifiable answer. Questions incorporate multimodal evidence, including images, maps, PDFs, logos, charts, tables, and video frames, with at least one non-text element that changes the reasoning state. Mr.LHDR evaluates both final answers and the correctness of intermediate conclusions under annotated dependencies. We evaluate general models, deep research systems, and agent frameworks using Overall Accuracy (OA), Strict Accuracy (SA), Checklist Score (CS), and Dependency-Aware Checklist Score (DACS). Results show that even the strongest system achieves only 43.1% OA and 34.3% SA, indicating that final-answer accuracy substantially overestimates complete research success. Removing images reduces DACS by 12.6 points, demonstrating the importance of multimodal evidence, while SA consistently declines as reasoning chains become longer. These findings reveal sustained, dependency-consistent evidence integration, rather than isolated fact retrieval, as a key bottleneck for current deep research agents.
Chinese Translation
深度研究智能体在网页搜索、工具使用、多模态证据分析和信息综合方面的能力日益增强。然而,现有基准主要评估中等时程的探索,很少测试智能体能否维持漫长且高度依赖的研究过程。我们提出了Mr.LHDR(多模态真实世界长时程深度研究),一个用于评估跨越八个类别的、由相互依赖的证据构成的长且不可约简的链条上的真实世界深度研究的基准。每个问题由隐藏的节点-关系图构建,平均需要12.1个必要的中间结论,平均依赖深度为10.4,才能得到一个简短、唯一且可验证的答案。问题融合了多模态证据,包括图像、地图、PDF、徽标、图表、表格和视频帧,且至少有一个非文本元素会改变推理状态。Mr.LHDR在标注依赖关系下评估最终答案和中间结论的正确性。我们使用总体准确率(OA)、严格准确率(SA)、检查表得分(CS)和依赖感知检查表得分(DACS)来评估通用模型、深度研究系统和智能体框架。结果表明,即使是最强的系统也仅达到43.1%的OA和34.3%的SA,这表明最终答案准确率大大高估了完整研究的成功。移除图像会使DACS降低12.6分,证明了多模态证据的重要性,而随着推理链变长,SA持续下降。这些发现揭示,持续且依赖一致的证据整合,而非孤立的事实检索,是当前深度研究智能体的关键瓶颈。
cs.AI / 43 / 2609.11319

Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification

Magenta:实现数学推理与Lean验证的闭环
Leang, Joshua Ong Jun, Li, Haonan, Zhao, Zheng, Shang, Xinyi, Li, Wenda, Liu, Zhengzhong, Xing, Erix, Cohen, Shay, Giunchiglia, Eleonora
Abstract
Most of mathematical knowledge has been communicated through so-called informal use of mathematics and natural language. With large language models (LLMs) being highly adept in using natural language, they achieve strong performance, yet not perfect, in informal mathematical reasoning. Restraining LLMs to informal reasoning misses out on the opportunity to use the discrete verification abilities that machines offer through machine-checkable proofs. In this paper, we bridge the gap between informal and formal reasoning by integrating Lean signals into the informal reasoning process. We introduce Magenta, a training-free agentic pipeline that, given only a natural-language problem, produces an answer, expresses it as a Lean 4 statement, and constructs a machine-checked proof. A statement judge verifies whether the formalisation preserves the original problem, while an error-attribution judge routes failed attempts either to mathematical re-derivation or local Lean repair. Magenta achieves 100% accuracy across all evaluated olympiad benchmarks, including AIME 2025, AIME 2026, and HMMT February 2026. When paired with the open-weight K2-Horizon-7B reasoner, it solves all six IMO 2026 problems. Our analysis shows that statement adjudication is essential for preventing false certificates and that feedback-guided correction outperforms independent resampling on difficult problems.
Chinese Translation
大部分数学知识是通过所谓的非形式化数学和自然语言传播的。由于大语言模型(LLMs)非常擅长使用自然语言,它们在非形式化数学推理中表现强劲,但并非完美。将LLMs限制在非形式化推理中,错失了利用机器通过机器可检查证明所提供的离散验证能力的机会。在本文中,我们通过将Lean信号集成到非形式化推理过程中,弥合了非形式化推理与形式化推理之间的差距。我们介绍了Magenta,一种无需训练的智能体流水线,它仅给定一个自然语言问题,就能产生答案,将其表达为Lean 4陈述,并构造机器可检查的证明。一个陈述评判器验证形式化是否保留了原始问题,而一个错误归因评判器将失败的尝试路由到数学重新推导或局部Lean修复。Magenta在所有评估的奥林匹克竞赛基准上达到了100%的准确率,包括AIME 2025、AIME 2026和HMMT February 2026。当与开放权重的K2-Horizon-7B推理器配对时,它解决了所有六个IMO 2026问题。我们的分析表明,陈述裁决对于防止虚假证书至关重要,并且反馈引导的校正比独立重采样在困难问题上表现更好。
cs.AI / 44 / 2609.11321

AI Exposure and AI Resilience: A Two-Dimensional Assessment Framework for Software and Software-Based Business Model

AI暴露度与AI韧性:面向软件及基于软件的商业模式的双维评估框架
Mandl, Paul Darius, Mandl, Peter, Häusl, Martin
Abstract
Artificial intelligence is changing both software production and the economics of software-based business models. Classical technology due diligence mainly examines technical properties such as architecture, scalability, and technical debt. These criteria do not fully capture how AI can affect a company's value proposition, competitive position, margins, or access to customers. This paper develops Artificial Intelligence Exposure and Resilience (AI-ER) as a two-dimensional assessment framework. AI exposure describes the pressure for change that AI creates for a business model. AI resilience describes the company's ability to absorb that pressure, adapt to changed conditions, and use AI in an economically viable way. Metrics for both dimensions are derived from current AI capabilities, their deployment conditions, and relevant research on business models and organizational adaptability. The model keeps exposure and resilience separate and adds an explicit assessment of evidence quality and confidence. It can be applied first with public information and later refined with internal evidence. The result is a traceable company profile that supports comparison without concealing uncertainty in the underlying evidence. The paper also specifies an initial score logic and a procedure for empirical validation.
Chinese Translation
人工智能正在改变软件生产以及基于软件的商业模式的经济学。传统的技术尽职调查主要考察技术属性,如架构、可扩展性和技术债务。这些标准并未充分捕捉AI如何影响公司的价值主张、竞争地位、利润率或客户获取。本文提出了人工智能暴露度与韧性(AI-ER)作为二维评估框架。AI暴露度描述了AI为商业模式带来的变革压力。AI韧性描述了公司吸收该压力、适应变化的条件以及以经济可行方式使用AI的能力。这两个维度的度量指标源自当前AI能力、其部署条件以及关于商业模式和组织适应性的相关研究。该模型将暴露度和韧性分开,并增加了对证据质量和置信度的明确评估。它可以先使用公开信息应用,之后用内部证据进行完善。结果是可追溯的公司画像,支持比较而不会掩盖基础证据中的不确定性。本文还规定了初始评分逻辑和实证验证程序。
cs.AI / 45 / 2609.11341

Exploring Diffusion Transformers for Cross-Modal Augmentation in Multimodal Brain State Decoding

探索扩散Transformer在多模态脑状态解码中的跨模态增强
Wang, Ziwei, He, Xingyi, Wang, Hongbin, Jia, Tianwang, Fang, Bohan, Wu, Dongrui
Abstract
Multimodal brain state decoding has largely focused on fusing paired modalities for prediction, but has rarely explored how their correspondence can be further exploited to enrich training data and improve multimodal representation learning. To address this gap, we propose CoMA-DiT, a bidirectional cross-modal Diffusion Transformer for latent augmentation that treats paired modalities as sources of mutual generative supervision rather than merely as inputs to be fused. CoMA-DiT conditions velocity prediction on the paired modality through cross-modal attention and adaptively injects the resulting variation via a reliability-gated residual mechanism. Experiments on multimodal auditory attention decoding and emotion recognition showed that CoMA-DiT consistently outperformed 20 representative baselines, achieving absolute gains of 4.28% and 6.70% in accuracy and macro-F1 over the no-augmentation baseline, respectively. Extensive ablation, sensitivity, visualization, and interpretability analyses further demonstrated its robustness, generalizability, and ability to capture functionally relevant cross-modal interactions. These findings support a broader view of multimodal learning: Paired modalities can serve not only as inputs for fusion but also as supervision sources that augment one another.
Chinese Translation
多模态脑状态解码在很大程度上集中于融合配对模态进行预测,但很少探索如何进一步利用其对应关系来丰富训练数据并改进多模态表示学习。为解决这一空白,我们提出CoMA-DiT,一种用于潜空间增强的双向跨模态扩散Transformer,它将配对模态视为相互生成监督的来源,而不仅仅是被融合的输入。CoMA-DiT通过跨模态注意力将速度预测条件于配对模态,并通过可靠性门控残差机制自适应地注入由此产生的变化。在多模态听觉注意解码和情感识别上的实验表明,CoMA-DiT持续优于20个代表性基线,相对于无增强基线,在准确率和宏F1上分别取得了4.28%和6.70%的绝对提升。广泛的消融、敏感性、可视化和可解释性分析进一步证明了其鲁棒性、泛化能力以及捕获功能相关跨模态相互作用的能力。这些发现支持了多模态学习的一种更广阔观点:配对模态不仅可以作为融合的输入,还可以作为相互增强的监督来源。
cs.AI / 46 / 2609.11365

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

可移植语义,私有方言:语言模型细胞之间潜在通信中的重用与负迁移
Marincat, Narcis
Abstract
In shared-genome language-model societies, restricted evidence visibility favors reusable, value-indexed latent packet interfaces, whereas the sole high-performing globally visible model in the parent study learned an episode-entangled code. This companion study asks whether independently trained societies share one packet language, where strict zero-shot transfer fails, and whether inherited interface state helps or harms later learning. First, a leakage-controlled causal interoperability audit over all 30 ordered pairs of six independently trained restricted societies -- under sealed held-out structure and a preregistered raw/orthogonal/linear/nonlinear alignment ladder -- shows the six semantically similar interfaces do not form one raw language: one same-initialization pair is exactly interoperable in both directions, a second shows asymmetric partial compatibility, and all 26 cross-initialization directions fail every frozen alignment rung. Second, within the tested decomposition and a single sealed source formulation, a source-span control localizes strict zero-shot failure to interpretation and execution of the new operator instructions. Third, in a matched adaptation factorial, the globally trained communication interface acts as a severe negative-transfer prior: reinitializing only the packet reader, writer, and mouth raises final depth-three accuracy from 0.169 to 0.857. Fourth, across two restricted checkpoints and two independently frozen target streams each, inherited interfaces never exceeded fresh-interface controls by the preregistered 0.10 margin. All primary conclusions are bounded to a near-transfer 17-state setting; the negative-transfer factorial concerns one globally visible parent-cohort checkpoint, while an appendix adds a post hoc tagged-global twin case study.
Chinese Translation
在共享基因组的语言模型社会中,受限的证据可见性有利于可重用的、值索引的潜在数据包接口,而母研究中唯一高性能的全局可见模型学习到了一种情节纠缠的编码。这项伴随研究询问独立训练的社会是否共享一种数据包语言,严格的零样本迁移在哪里失败,以及继承的接口状态是有助于还是损害后续学习。首先,在密封的留出结构和预先注册的原始/正交/线性/非线性对齐阶梯下,对六个独立训练的受限社会的所有30个有序对进行泄漏控制的因果互操作性审计,结果表明六个语义相似的接口并不构成一种原始语言:一个相同初始化的对在两个方向上完全可互操作,第二个显示出不对称的部分兼容性,而所有26个交叉初始化方向在每个冻结的对齐级别上都失败。其次,在测试的分解和单一密封源公式内,源跨度控制将严格的零样本失败定位到新算子指令的解释和执行。第三,在匹配的适应因子设计中,全局训练的通信接口充当严重的负迁移先验:仅重新初始化数据包读取器、写入器和口,将最终深度三的准确率从0.169提高到0.857。第四,在两个受限检查点和每个两个独立冻结的目标流上,继承的接口从未超过新接口控制的预先注册0.10幅度。所有主要结论都限定在近似迁移的17状态设置中;负迁移因子设计涉及一个全局可见的母队列检查点,而附录增加了一个事后标记的全局孪生案例研究。
cs.AI / 47 / 2609.11372

RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection

RAMamba-Net:一种用于听觉注意检测的可靠性感知与基于Mamba的多模态融合网络
He, Xingyi, Wang, Ziwei, Wu, Dongrui
Abstract
Auditory attention decoding (AAD) identifies the attended speaker from physiological signals, supporting neuro-steered hearing devices and natural human-machine interaction. Electroencephalography (EEG) is the dominant modality for AAD but provides incomplete evidence in naturalistic audio-visual scenes, motivating EEG and electrooculography (EOG) fusion. Existing approaches remain limited by weak cross-modal interaction, inefficient temporal modeling, and low robustness to sample variations. To address the limitations, we propose RAMamba-Net, a reliability-aware Mamba-based multimodal fusion network for AAD. RAMamba-Net employs a Mamba-enhanced band-aware convolutional Transformer to capture band-specific EEG patterns and long-range temporal dynamics. A dual-branch temporal-spatial encoder models EOG temporal and inter-channel dependencies. Cross-modal attention enables explicit modality interaction. Then, a reliability-aware module is introduced to estimate sample-wise modality weights for feature and prediction consistency, thereby enhancing multimodal fusion. Experiments on two AAD benchmarks demonstrate that RAMamba-Net effectively exploits complementary EEG-EOG information, yielding accuracy gains of 5.76% over unimodal baselines, together with more robust decoding and discriminative representations. Further analyses show that explicit cross-modal interaction improves multimodal alignment, while the reliability-aware module suppresses unreliable modality evidence and is robust to signal perturbation and parameter variation.
Chinese Translation
听觉注意解码(AAD)从生理信号中识别出被注意的说话人,支持神经引导的听力设备和自然的人机交互。脑电图(EEG)是AAD的主要模态,但在自然视听场景中提供的证据不完整,促使了EEG和眼电图(EOG)的融合。现有方法仍受限于跨模态交互弱、时间建模效率低以及对样本变化的鲁棒性差。为了解决这些局限性,我们提出了RAMamba-Net,一种用于AAD的可靠性感知的基于Mamba的多模态融合网络。RAMamba-Net采用Mamba增强的频带感知卷积Transformer来捕捉频带特定的EEG模式和长时程时间动态。双分支时空编码器对EOG的时间依赖性和通道间依赖性进行建模。跨模态注意力实现了显式的模态交互。然后,引入了一个可靠性感知模块来估计样本级的模态权重,以保持特征和预测一致性,从而增强多模态融合。在两个AAD基准上的实验表明,RAMamba-Net有效利用了互补的EEG-EOG信息,相比单模态基线提高了5.76%的准确率,并具有更鲁棒的解码和判别性表示。进一步的分析表明,显式的跨模态交互改善了多模态对齐,而可靠性感知模块抑制了不可靠的模态证据,并且对信号扰动和参数变化具有鲁棒性。
cs.AI / 48 / 2609.11393

Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning

超越置信度:面向LLM推理的稳定性感知测试时适应
Gu, Bincheng, Gao, Min, Wang, Zongwei, Bai, Yibing, He, Yulan, Yu, Junliang
Abstract
Test-time adaptation has emerged as a lightweight alternative to costly post-training for improving the reasoning capabilities of Large Language Models (LLMs) on downstream tasks. Predictive entropy provides a model-derived signal for such adaptation, guiding models toward higher-confidence reasoning states without external verifiers or reward models. However, higher confidence does not necessarily imply correctness, as LLMs may remain highly confident along incorrect reasoning trajectories. We observe that high-confidence reasoning is more likely to be correct when confidence remains stable under local perturbations. Based on this observation, we propose Test-Time Adaptation via Stability-Aware Confidence Optimization (TASCO), a framework that incorporates local stability into confidence-based test-time adaptation while keeping the LLM frozen. TASCO operationalizes local stability by optimizing a lightweight task-level prefix under two alternative perturbation strategies: Random Perturbation promotes distributional stability across trajectories induced by nearby perturbed prefixes, whereas Sharpness-Aware Perturbation targets worst-case local sensitivity. Experiments demonstrate that TASCO improves reasoning accuracy and token efficiency across diverse LLMs and reasoning benchmarks, while behavioral analyses show that it maintains stable confidence under local perturbations without prematurely concentrating the model's predictive distribution.
Chinese Translation
测试时适应已成为一种轻量级的替代方案,用于替代昂贵的后训练,以提升大语言模型(LLM)在下游任务上的推理能力。预测熵为这种适应提供了一种源自模型的信号,引导模型走向更高置信度的推理状态,而无需外部验证器或奖励模型。然而,更高的置信度并不一定意味着正确性,因为LLM可能沿着错误的推理轨迹仍然保持高置信度。我们观察到,当置信度在局部扰动下保持稳定时,高置信度推理更可能是正确的。基于这一观察,我们提出了通过稳定性感知置信度优化的测试时适应(TASCO),这是一个将局部稳定性融入基于置信度的测试时适应的框架,同时保持LLM冻结。TASCO通过在两种替代扰动策略下优化一个轻量级的任务级前缀来操作化局部稳定性:随机扰动促进由邻近扰动前缀引发的轨迹上的分布稳定性,而锐度感知扰动则针对最坏情况的局部敏感性。实验表明,TASCO在多种LLM和推理基准上提高了推理准确性和token效率,而行为分析显示,它在局部扰动下保持稳定的置信度,而不会过早地集中模型的预测分布。
cs.AI / 49 / 2609.11403

From Queries to Narratives: Cultural Heritage Data Stories for Knowledge Graph Exploration and Quality Assessment

从查询到叙事:面向知识图谱探索与质量评估的文化遗产数据故事
Tietz, Tabea, Schrade, Torsten, Posthumus, Etienne, Söhn, Linnaea, Steller, Jonatan Jalle, Waitelonis, Jörg, Sack, Harald
Abstract
Cultural-heritage KGs such as the NFDI4Culture-KG contain millions of triples about artworks, music, inscriptions, historical events, and the people and places connected to them. For many users, however, discovering this knowledge can be difficult. While SPARQL can be learned, writing meaningful queries first requires an in-depth understanding of the graph's data model, an investment many domain researchers and practitioners are unwilling to make. Even with existing user interfaces, a starting point and some guidance are usually needed, because the data contained in the graph is highly specialized, heterogeneous, and constantly growing, making it challenging to know what it contains or which questions it can answer. In this paper, we present data stories as a way not only to lower this barrier, but also to turn exploration into data-quality assessment, and thus combine accessible querying with the discovery of issues that remain hidden in aggregate statistics. In this contribution, a data story is understood as a narrative document that integrates explanatory text and images with executable SPARQL queries and their visualized results. It is described how they are authored against the graph and how they serve several purposes: guiding users through an unfamiliar graph, creating reproducible narratives, and surfacing data-quality issues previously hidden in aggregate statistics. The authoring platform LODEON including its Sparnatural and AI-supported authoring assistants is introduced as a proof-of-concept. Within the authoring environment, every claim made about the data can be backed by an explicit query, making these narratives transparent and reproducible. This paper also reflects on lessons learned from hands-on seminars and workshops. Early experience suggests that such data stories make cultural-heritage knowledge graphs more accessible for both exploration and quality assessment.
Chinese Translation
诸如 NFDI4Culture-KG 等文化遗产知识图谱(KG)包含数百万个关于艺术品、音乐、铭文、历史事件以及与之相关的人物和地点的三元组。然而,对于许多用户而言,发现这些知识可能很困难。尽管 SPARQL 可以学习,但编写有意义的查询首先需要深入理解图的数据模型,这是许多领域研究者和实践者不愿付出的投入。即使存在现有的用户界面,通常也需要一个起点和一些指导,因为图中包含的数据高度专业化、异构且不断增长,这使得了解其包含的内容或它能回答哪些问题变得具有挑战性。在本文中,我们提出数据故事作为一种方式,不仅降低这一障碍,而且将探索转化为数据质量评估,从而将可访问的查询与发现隐藏在聚合统计中的问题结合起来。在本研究中,数据故事被理解为一种叙事文档,它将解释性文本和图像与可执行的 SPARQL 查询及其可视化结果集成在一起。本文描述了它们如何基于图进行创作,以及如何服务于多个目的:引导用户浏览不熟悉的图,创建可复现的叙事,并揭示以前隐藏在聚合统计中的数据质量问题。介绍了创作平台 LODEON 及其 Sparnatural 和 AI 支持的创作助手作为概念验证。在创作环境中,关于数据的每一个声明都可以由显式查询支持,使这些叙事透明且可复现。本文还反思了从实践研讨会和工作坊中吸取的经验教训。早期经验表明,此类数据故事使文化遗产知识图谱在探索和质量评估方面都更易于使用。
cs.AI / 50 / 2609.11431

LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study

LLMs 作为符号回归中生理合理性的事后审计者:一项由临床医生评估的案例研究
López-Varela, Jorge, Hidalgo, J. Ignacio, Muñoz, José-Manuel, Costilla-Reyes, Omar, Maqueda, Esther, Moreno-Fernandez, Jesus, González-Vidal, Tomás, Velasco, J. Manuel, Garnica, Oscar
Abstract
Genetic Programming and its variants, such as grammatical evolution, are widely used in Symbolic Regression to derive mathematical expressions from multivariate data. In addition to predictive accuracy, models are appreciated for their potential to provide interpretability, offering explicit equations that relate input variables to outcomes. However, achieving interpretability and plausibility remains challenging, as evolved models may be complex or scientifically inconsistent. In this study, we explore whether Large Language Models, can assist in improving the explainability of Symbolic Regression models generated by evolutionary computation methods. Building upon our previous work on estimating body fat percentage using grammar-based Genetic Programming , we investigate the use of LLMs as post-processing tools to analyze and rank evolved expressions according to their interpretability and medical plausibility. Four symbolic expressions are analysed by three LLMs over three repeated runs, and the resulting interpretations and rankings are assessed by a panel of three clinicians. Across the three LLMs, comparative model-ranking outputs received more favorable clinician assessments than isolated term-level interpretations. However, the LLMs also produced physiologically and mathematically questionable explanations, indicating that they are better suited to comparative auditing under expert oversight than to autonomous validation.\blfootnote{The present work is an extended version of a paper submitted into a journal.
Chinese Translation
遗传编程及其变体,如语法进化,广泛用于符号回归以从多变量数据中导出数学表达式。除了预测准确性外,模型因其提供可解释性的潜力而受到重视,提供将输入变量与结果关联的显式方程。然而,实现可解释性和合理性仍然具有挑战性,因为进化出的模型可能复杂或科学上不一致。在本研究中,我们探讨大型语言模型(Large Language Models, LLMs)能否帮助提高由进化计算方法生成的符号回归模型的可解释性。在我们之前使用基于语法的遗传编程估计体脂百分比的工作基础上,我们研究使用 LLMs 作为后处理工具,根据可解释性和医学合理性分析和排序进化出的表达式。四个符号表达式由三个 LLMs 在三次重复运行中分析,所得解释和排名由三名临床医生组成的小组评估。在三个 LLMs 中,比较性模型排名输出比孤立的术语级解释获得了更有利的临床医生评估。然而,LLMs 也产生了生理学和数学上可疑的解释,表明它们更适合在专家监督下进行比较审计,而不是自主验证。 (本文是提交至某期刊的一篇论文的扩展版本。)
cs.AI / 51 / 2609.11446

Calibration-Aware Uncertainty Cascades for Efficient Heterogeneous Model Collaboration

面向高效异构模型协作的校准感知不确定性级联
Zhang, Yilin, Jiang, Han, Xu, Cai, Liu, Ying, Zhao, Wei
Abstract
Heterogeneous model collaboration seeks to exploit the complementary strengths of different models to balance predictive performance and inference cost. Existing approaches typically rely either on trained routers, which tie routing decisions to a fixed task and model pool, or on raw-confidence cascades, whose thresholds lack consistent reliability semantics across heterogeneous models. Consequently, these approaches adapt poorly to changing model pools and deployment budgets. We propose Calibration-Aware Uncertainty Cascades (CAUC), a simple post-hoc framework that independently calibrates each model's confidence and selects deployment policies using validation data. The resulting calibrated confidence scores establish a common reliability scale for accepting an early prediction, invoking a stronger model, or selectively combining model outputs. This unified decision criterion decouples deployment policies from any particular model pool or operating budget. We further show theoretically that calibration gives confidence thresholds an explicit selective-risk interpretation, whereas uncalibrated scores offer no comparable reliability guarantee. Extensive experiments demonstrate that, across six language benchmarks, CAUC achieves an average relative accuracy improvement of 1.9% over strong-model-only inference while avoiding approximately 47% of strong-model calls. On image classification benchmarks, it maintains or improves predictive performance while reducing measured GFLOPs by up to 57%.
Chinese Translation
异构模型协作旨在利用不同模型的互补优势,以平衡预测性能和推理成本。现有方法通常依赖于训练后的路由器,其将路由决策绑定到固定任务和模型池,或依赖于原始置信度级联,其阈值在异构模型间缺乏一致的可靠性语义。因此,这些方法难以适应变化的模型池和部署预算。我们提出校准感知不确定性级联(CAUC),一个简单的事后框架,它独立校准每个模型的置信度,并使用验证数据选择部署策略。所得校准置信度分数建立了一个通用的可靠性尺度,用于接受早期预测、调用更强模型或选择性组合模型输出。这种统一的决策准则将部署策略与任何特定模型池或操作预算解耦。我们进一步从理论上表明,校准赋予置信度阈值明确的选择性风险解释,而未校准的分数则无法提供可比的可靠性保证。大量实验表明,在六个语言基准上,CAUC 相比仅使用强模型推理实现了平均 1.9% 的相对准确率提升,同时避免了约 47% 的强模型调用。在图像分类基准上,它保持或提升了预测性能,同时将测得的 GFLOPs 降低了高达 57%。
cs.AI / 52 / 2609.11452

RouteRepair: Instance-Level Failure Diagnosis and Targeted Repair in LLM-Based Automated Heuristic Design for Routing Optimization

RouteRepair:面向路径优化的基于LLM的自动启发式设计中的实例级故障诊断与定向修复
Ji, Binghao, Huang, Di, Fang, Jiahui, Liu, Zhiyuan
Abstract
Efficient routing optimization is essential to freight transportation, urban logistics, and shared mobility, where high-quality heuristics are often required under limited computational budgets. Recent large language model (LLM)-based automated heuristic design methods can generate effective routing rules, but aggregate evaluation may mask recurrent failures on particular instance structures. To address this limitation, this study develops RouteRepair, which diagnoses parent-specific weaknesses from instance-level performance and applies targeted modifications to the corresponding heuristic components while protecting behavior that already performs well. Routing evidence, solver behavior, and program context are combined to define bounded repair objectives, and each intervention is validated through matched parent-child evaluation of failure recovery and collateral degradation. Experiments on the traveling salesman problem (TSP) and capacitated vehicle routing problem (CVRP) span constructive search, guided local search, and ant colony optimization. RouteRepair-GLS reduces the mean TSP optimality gap from 1.7476% to 0.7587%, while the constructive CVRP heuristic lowers average route cost by 1.91% relative to the savings heuristic; the generated ACO priors also outperform matched hand-designed priors. These results show that failure-aware, evidence-constrained refinement can improve routing heuristics on difficult instances while preserving performance on cases they already solve well.
Chinese Translation
高效路径优化对于货运、城市物流和共享出行至关重要,在这些场景中,常常需要在有限的计算预算下获得高质量的启发式方法。近期基于大型语言模型(LLM)的自动启发式设计方法能够生成有效的路径规则,但聚合评估可能掩盖在特定实例结构上反复出现的失败。为解决这一局限,本研究提出了RouteRepair,它从实例级性能中诊断父代特定的弱点,并对相应的启发式组件进行定向修改,同时保护那些已经表现良好的行为。路由证据、求解器行为和程序上下文被结合起来定义有界修复目标,并且每次干预都通过匹配的父-子评估来验证故障恢复和附带退化。在旅行商问题(TSP)和带容量约束的车辆路径问题(CVRP)上的实验涵盖了构造性搜索、引导局部搜索和蚁群优化。RouteRepair-GLS将平均TSP最优性差距从1.7476%降低到0.7587%,而构造性CVRP启发式相对于节约启发式将平均路径成本降低了1.91%;生成的ACO先验也优于匹配的手工设计先验。这些结果表明,故障感知、证据约束的改进能够提升困难实例上的路径启发式方法,同时保持其在已能很好解决的实例上的性能。
cs.AI / 53 / 2609.11458

Flexible and Interpretable Accent Distance Measurements

灵活且可解释的口音距离测量
McGhee, Charles, Gales, Mark J. F., Knill, Kate M.
Abstract
Determining the differences between two speakers' accents is a fundamental task in linguistics and speech technology research. The methodology used to measure these differences depends on the specific research area. A phonetics researcher may demonstrate accent variation by comparing vowel formants in paired recordings of individual words. These results will be interpretable, but the recordings will be time-consuming to collect and may not be representative of connected speech. Accented Text-to-Speech (TTS) research has pushed towards using accent embeddings derived from accent classification tasks. These embeddings can be produced from any speech recording, but are not readily interpretable. In this paper, we demonstrate that articulatory representations created through articulatory inversion can be used as an interpretable basis for accent comparison and that optimal transport provides a framework for accent comparison across arbitrary recording types.
Chinese Translation
确定两个说话者口音之间的差异是语言学和语音技术研究中的一项基本任务。用于测量这些差异的方法取决于具体的研究领域。语音学研究者可能通过比较单个单词的配对录音中的元音共振峰来展示口音变异。这些结果具有可解释性,但录音的收集耗时且可能无法代表连续语音。带口音的文本到语音(TTS)研究已趋向于使用从口音分类任务中得到的口音嵌入。这些嵌入可以从任何语音录音中生成,但不易解释。在本文中,我们证明了通过发音反演创建的发音表示可以作为口音比较的可解释基础,并且最优传输(optimal transport)为跨任意录音类型的口音比较提供了一个框架。
cs.AI / 54 / 2609.11489

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

约定差距:迈向衡量合作AI评估中的隐式沟通
Fukushima, Makoto, Xiong, Hua-Dong, Pari, Ehsan Moradi
Abstract
Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions---shared protocols for reading meaning beyond the literal message---which AI-AI benchmarks may not capture. We propose the \emph{convention gap}, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (hanab.live), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, $-$0.7~pp in AI pairs, and +16.4~pp in human-AI pairs, and was concentrated on plays of cards that had received no hints (+46~pp in human pairs). Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38--41\%), but human failure rates ranged from 14.4\% to 34.4\% and the gap from +24.1 to +6.2~pp; the partner eliciting the largest gap produced the fewest human failures. Game score carried different information: it depended on each corpus's roster composition, whereas the gap separated human from AI play at the agent level. As a known-answer check, Off-Belief Learning agents, whose convention content is controlled by construction, gave a gap of +1.6~pp at the convention-free level, rising monotonically to +21.7~pp. These results suggest that convention compatibility, rather than AI-AI performance, may predict an AI's effectiveness with human partners.
Chinese Translation
合作AI智能体是相对于其他AI进行评估的,然而人类合作依赖于隐式约定——即超越字面信息读取意义的共享协议——AI-AI基准测试可能无法捕捉到这一点。我们提出“约定差距”,即从沟通的字面内容预测的失败概率与观察到的失败率之间的差异,作为隐式沟通的度量。在纸牌游戏Hanabi中,有限的牌堆和确定性的提示约束使得该后验概率可以精确计算。我们重放了来自三个人类-人类(hanab.live)、AI-AI(HOAD)和人类-AI(HanabiData)游戏公开数据集的约101,000个出牌动作。该差距在人类配对中为+26.2个百分点(pp),在AI配对中为-0.7个百分点,在人类-AI配对中为+16.4个百分点,并且集中在未收到任何提示的牌的出牌上(人类配对中为+46个百分点)。在人类-AI对局中,人类可获得的字面信息在三个AI合作伙伴之间相似(平均预测失败率为38-41%),但人类失败率范围从14.4%到34.4%,差距从+24.1到+6.2个百分点;引发最大差距的合作伙伴产生的人类失败最少。游戏得分携带了不同的信息:它取决于每个语料库的阵容组成,而差距在智能体层面区分了人类和AI的对局。作为已知答案检验,Off-Belief Learning智能体(其约定内容通过构造控制)在无约定水平上给出了+1.6个百分点的差距,单调上升至+21.7个百分点。这些结果表明,约定兼容性(而非AI-AI性能)可能预测AI与人类合作伙伴合作的有效性。
cs.AI / 55 / 2609.11490

Published Unlearning Numbers Move Per Checkpoint, and Not Because the Removed Data Survives: An Audit of 263 Released Batch-Normalized Checkpoints

已发布的遗忘数值因检查点而异,并非因为被移除数据存活:对263个已发布批量归一化检查点的审计
Li, Junlong Shen Xingyu
Abstract
An unlearning audit reads its verdict off numbers that an unlearned model and its retrained reference each publish, and both also ship batch-normalization statistics that no gradient step wrote and no release records. Refitting them on kept data at bit-identical weights moves 47 of 221 released checkpoints past the spread their own release's seeds show, several inside a method whose average does not move: what moves is the checkpoint's property, not its method's. What does the moving is not the removed data surviving in the state: exchanging kept records for removed ones inside a fixed fitting pool moves a published cell by almost nothing, while how far a checkpoint's shipped state has drifted from any refit does track it. The consequence for a published decision is real but narrow: twelve verdicts cross, four clear a measured recalibration budget, two clear it on every replicate, and a population we trained and sited near its own criterion yields none. A release should therefore name the fitting convention beside the number, on the batch-normalized vision models where this channel exists.
Chinese Translation
一个遗忘审计从遗忘模型及其重新训练的参考模型各自发布的数字中读取其裁决,而两者还附带批量归一化统计量,这些统计量没有梯度步骤写入,也没有发布记录。在位级相同的权重下,用保留数据重新拟合它们,使221个已发布检查点中的47个移动超出它们自己发布的种子所显示的散布范围,其中几个位于一个平均值不变的方法内部:移动的是检查点的属性,而不是其方法的属性。造成这种移动的并不是状态中残留的被移除数据:在一个固定拟合池内,将保留记录替换为被移除记录,几乎不移动已发布的单元格,而检查点发布的状与任何重新拟合之间的漂移程度确实能追踪到它。对已发布决策的影响是真实的但狭窄的:十二个裁决跨越,四个通过了测量的重校准预算,两个在每个重复上都通过了,而我们训练并置于其自身标准附近的一个群体则没有产生任何结果。因此,在存在此通道的批量归一化视觉模型上,发布时应在数字旁边标明拟合约定。
cs.AI / 56 / 2609.11493

From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development

从文档孤岛到过程智能:面向CMC工艺开发的多层知识图谱
Amirmoshiri, Reza, Sahneh, Faryad, Jangjou, Yasser
Abstract
Chemistry, Manufacturing and Controls (CMC) process development generates an enormous body of technical information across a multi-stage, knowledge-intensive continuum from drug discovery to commercial manufacturing. This knowledge is traditionally fragmented across functions and heterogeneous formats, causing traceability gaps and significant knowledge-management costs during technology transfer and regulatory filing. We present a modular agentic-AI platform that converts a heterogeneous corpus of process-development documents into a queryable, dual-layer knowledge graph. A base knowledge layer builds a lexical graph with a Document-Section-Chunk hierarchy through lossless ingestion of digital, scanned, handwritten, and multilingual documents, while an intelligence layer extracts ontology-aligned entities and bridges cross-document concepts through a provenance-anchored domain graph. LLM agents operate across both layers, selecting the retrieval path best suited to each question. We evaluate the lexical layer with a novel three-tier protocol measuring the deployment-fidelity of a retrieval-augmented generation (RAG) system on proprietary data, demonstrated on 505 questions curated from 38 development reports of a Sanofi small-molecule program. Tier-1 multiple-choice accuracy of 95% signals strong platform reliability; the stricter Tier-2 LLM-judge pass rate of 85%, which degrades on comparative and corpus-wide questions, reveals a failure taxonomy that Tier-1 accuracy alone fails to capture. A router agent selects between layers according to question type. We anticipate this protocol will enable future designers of agentic platforms to assess their systems against nonpublic databases, and that graph-based architectures will see broader adoption in pharma as a means of transforming fragmented document repositories into structured process intelligence.
Chinese Translation
化学、制造与控制(CMC)工艺开发在从药物发现到商业化生产的多个阶段、知识密集型的连续过程中产生了大量技术信息。这些知识传统上分散在不同职能部门和异构格式中,导致可追溯性缺口以及在技术转移和法规申报过程中产生高昂的知识管理成本。我们提出了一个模块化的代理式人工智能平台,可将工艺开发文档的异构语料库转换为可查询的双层知识图谱。基础知识层通过对数字、扫描、手写和多语言文档的无损摄取,构建具有文档-章节-块层次结构的词汇图;而智能层则提取本体对齐的实体,并通过溯源锚定的领域图桥接跨文档概念。LLM代理在两层中运行,为每个问题选择最合适的检索路径。我们采用一种新颖的三层协议评估词汇层,该协议测量检索增强生成(RAG)系统在专有数据上的部署保真度,并在来自赛诺菲小分子项目的38份开发报告中整理的505个问题上进行了演示。第1层多项选择准确率为95%,表明平台可靠性强;更严格的第2层LLM评判通过率为85%,在比较性和语料库范围的问题上有所下降,揭示了一种仅靠第1层准确率无法捕获的失败分类。路由代理根据问题类型在两层之间进行选择。我们预计该协议将使未来的代理式平台设计者能够针对非公开数据库评估其系统,并且基于图的架构将在制药领域得到更广泛的应用,作为将碎片化的文档存储库转化为结构化过程智能的一种手段。
cs.AI / 57 / 2609.11498

ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps

ActMap:基于生成时激活图的单次不确定性量化
Dardini, Jacopo, Calegari, Roberta
Abstract
Practical uncertainty quantification (UQ) for large language models must decide, from a single generation, whether a specific answer should be trusted. Existing methods either sample multiple generations, read only output-token probabilities, or reduce the model's internal computation to a single hidden state. We introduce ActMap, a white-box representation that compresses the generation-time hidden- state trajectory (every layer, every generated token) into a fixed $12 \times 32 \times 128$ tensor of temporal-statistic channels that preserves structure across transformer depth and pooled hidden coordinates. The map is captured during the generation pass with no measurable overhead, has a fixed shape across model depths and hidden sizes, and occupies 96 KiB: a compact artifact that can be retained for audit-relevant generations and probed directly, with occlusion analysis localizing the classifier's signal to mid-depth regions of the map. A lightweight classifier, instantiated as a compact Vision Transformer, reads an estimated correctness probability from each map in a fraction of a millisecond; capacity-matched MLPs perform comparably, indicating the representation itself carries the result. Trained and evaluated in-domain on short-answer QA, direct- answer math, and summarization factuality with three instruction-tuned 7-8B models, ActMap consistently outperforms sampling, token-probability, attention, and embedding baselines, and matches ACT-ViT, a detector trained on dense activation tensors $67 \times$ larger, at essentially the same mean AUROC with lower calibration error on ten of twelve pairs. The resulting score supports abstention, routing, and selective verification from a single generation, making it a practical primitive for scalable oversight of deployed models.
Chinese Translation
大语言模型的实际不确定性量化(UQ)必须从单次生成中判定某个特定答案是否可信。现有方法要么多次采样生成,要么仅读取输出词元概率,或者将模型内部计算约简为单个隐藏状态。我们提出ActMap,一种白盒表示,它将生成时的隐藏状态轨迹(每一层、每个生成的词元)压缩成固定的$12 \times 32 \times 128$张量,包含时序统计通道,保留了跨Transformer深度和池化隐藏坐标的结构。该图在生成过程中捕获,没有可测量的开销,其形状不随模型深度和隐藏大小变化,仅占96 KiB:这是一种紧凑的工件,可以保留用于审计相关的生成,并直接进行探测,通过遮挡分析将分类器的信号定位到图的中间深度区域。一个轻量级分类器,以紧凑的Vision Transformer实现,在不到一毫秒内从每个图读取估计的正确概率;容量匹配的MLP表现相当,表明该表示本身携带了结果。在短答案问答、直接答案数学和摘要事实性任务上,使用三个指令微调的7-8B模型进行域内训练和评估,ActMap持续优于采样、词元概率、注意力和嵌入基线,并且在十二对对比中有十对以基本相同的平均AUROC和更低的校准误差匹配了ACT-ViT——后者是在大67倍的稠密激活张量上训练的检测器。得到的分数支持从单次生成中进行弃权、路由和选择性验证,使其成为部署模型可扩展监督的实用原语。
cs.AI / 58 / 2609.11509

Extending SMT Solving with Non-Ground Clause Learning

用非基子句学习扩展SMT求解
Briefs, Yasmine, Weidenbach, Christoph
Abstract
Quantifier instantiation is currently the main approach to non-ground SMT solving: solvers generate ground instances and solve the resulting ground SMT problems with CDCL(T)-style reasoning. When a conflict is found, conflict analysis learns only a ground clause, even though the conflict comes from instances of non-ground clauses. Yet non-ground reasoning can give exponentially shorter proofs than purely ground reasoning. We propose a calculus that consists of ground instantiations, CDCL(T)-style rules, and non-ground conflict analysis. The solver reasons on ground instances, but the resolution steps of conflict analysis are performed on their original non-ground clauses. This produces learned clauses that are typically more general than the ground conflict. With a suitable strategy, the learned clauses are even non-redundant. We also show how chronological backtracking can be included in SMT solving. Our calculus gives a common setting for CDCL(T)-style SMT solving, a range of instantiation-based procedures, and non-ground clause learning, and we prove that it simulates CDCL, SCL(FOL), SCL(T), and even Resolution.
Chinese Translation
量词实例化目前是非基SMT求解的主要方法:求解器生成基实例,并用CDCL(T)风格的推理求解所得的基SMT问题。当发现冲突时,冲突分析仅学习一个基子句,尽管该冲突来自非基子句的实例。然而,非基推理可以给出比纯基推理指数级更短的证明。我们提出一个演算,由基实例化、CDCL(T)风格规则和非基冲突分析组成。求解器在基实例上进行推理,但冲突分析的消解步骤是在其原始非基子句上执行的。这会产生通常比基冲突更一般的学习子句。在合适的策略下,学习子句甚至是非冗余的。我们还展示了如何将时序回溯(chronological backtracking)纳入SMT求解。我们的演算为CDCL(T)风格SMT求解、一系列基于实例化的过程以及非基子句学习提供了一个统一框架,并证明它模拟了CDCL、SCL(FOL)、SCL(T),甚至Resolution。
cs.AI / 59 / 2609.11527

Lightweight LiDAR-Based Cone Detection Framework Using Random Forest for Formula Student Driverless

用于Formula Student Driverless的基于随机森林的轻量级LiDAR锥桶检测框架
Mező-Kerekes, Márk, Praksz, Péter, Liu, Chang
Abstract
Reliable, low-latency perception is crucial for Formula Student Driverless vehicles, yet many existing pipelines rely on deep learning and multi-sensor fusion, often requiring GPU acceleration. This paper presents a lightweight LiDAR-only perception pipeline tailored for CPU execution, combining ground removal, IMU-based motion compensation, DBSCAN clustering, and geometric feature-based Random Forest classification. Feature importance analysis reduced the model input from 12 to 7 features while preserving performance. Evaluated on 2,371 labeled clusters collected from real FSD events, the pipeline achieves an F1-score of 98.33% and an end-to-end runtime of 3.13 ms on CPU-only hardware. The released dataset, labeling tool, and trained models provide a practical and reproducible baseline for other resource-constrained autonomous racing teams.
Chinese Translation
可靠、低延迟的感知对于Formula Student Driverless (FSD)车辆至关重要,然而许多现有流程依赖深度学习和多传感器融合,通常需要GPU加速。本文提出了一种面向CPU执行的轻量级纯LiDAR感知流程,结合了地面去除、基于IMU的运动补偿、DBSCAN聚类以及基于几何特征的随机森林分类。特征重要性分析将模型输入从12个特征减少到7个,同时保持了性能。在从真实FSD赛事中收集的2,371个标注簇上进行评估,该流程在纯CPU硬件上实现了98.33%的F1分数和3.13毫秒的端到端运行时间。发布的数据集、标注工具和训练模型为其他资源受限的自动驾驶赛车队提供了实用且可复现的基线。
cs.AI / 60 / 2609.11532

Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems

提示词修订作为文本到图像系统中文化偏见的来源
Urman, Aleksandra, Lichtenegger, Elsa, Jaoua, Salima, Bouleimen, Azza, Forsberg, Robin, Hertweck, Corinna, Ionescu, Stefania, Pagan, Nicolò, Hannak, Ancsa, Baumann, Joachim
Abstract
Commercial text-to-image systems silently revise user prompts before generating images, a step users typically cannot disable or even see. Yet, existing audits of cultural bias examine only the final images and treat generation as a single pipeline, so they cannot tell where the bias originates. We introduce WORLDVIEW, a multilingual benchmark of 8,960 prompts across 15 languages and 31 language-context pairings. Using it, we audit the revision layer in three systems (DALL-E-3, Imagen-4, GPT-Image-1.5) through a three-step analysis of how heavily it marks each cultural context, whether it flattens that context into a narrow vocabulary, and whether that vocabulary is stereotypical. Relative to a no-context English baseline, the US is the least-marked context, while non-Western and non-Anglophone contexts are marked far more heavily, flattened into narrow vocabularies applied across topically diverse prompts, and reduced to recognizable cultural stereotypes. Comparing images from original versus revised prompts on models without a revision layer, we identify the layer itself as a previously undocumented, causal source of this stereotyping. To locate cultural bias, and fix it, we must audit the system as deployed, not the model alone.
Chinese Translation
商业文本到图像系统在生成图像前会悄悄修改用户提示词,而用户通常无法禁用甚至无法看到这一步骤。然而,现有的文化偏见审计仅检查最终图像,并将生成视为单一流程,因此无法判断偏见源自何处。我们推出了WORLDVIEW,一个包含8,960个提示词、涵盖15种语言和31个语言-语境配对的多语言基准。利用它,我们通过三步分析审计了三个系统(DALL-E-3、Imagen-4、GPT-Image-1.5)中的修订层:它对每个文化语境的标记有多重,是否将该语境扁平化为狭窄的词汇,以及该词汇是否具有刻板印象。相对于无语境英语基线,美国是被标记最少的语境,而非西方和非英语语境则被标记得重得多,被扁平化为应用于主题多样提示词的狭窄词汇,并被简化为可识别的文化刻板印象。通过在无修订层的模型上比较原始与修订提示词生成的图像,我们确定该层本身是这种刻板印象的一个此前未记录的因果来源。为了定位文化偏见并修复它,我们必须审计部署的系统,而不仅仅是模型。
cs.AI / 61 / 2609.11542

Characterizing Job Power Elasticity for Power-Flexible AI Training

面向功率灵活AI训练的作业功率弹性表征
Colangelo, Philip, Dawson, Charles, Sengupta, Shayan, Coskun, Ayse, Sivaram, Varun
Abstract
Large language model (LLM) training is among the fastest-growing sources of electricity demand in modern data centers, and power availability is a primary bottleneck to continued AI infrastructure growth. Making the power consumption of these workloads flexible could unlock additional power for AI growth, limit increases in electricity prices, and improve the utilization of existing grid infrastructure. However, to realize this flexibility, we must first understand how the performance of training workloads changes when GPU power is reduced. This paper presents the first systematic characterization of \emph{job power elasticity} (the sensitivity of throughput to power reductions) in LLM training. To quantify elasticity, we introduce the \emph{Power Flexibility Index (PFI)}, a normalized metric that quantifies the performance cost of power reductions and provides a control primitive for SLA-aware power flexibility. We collect data from 131 LLM training runs on H200 (plus 24 H200 validation runs and 34 matched H100 runs), including both dense and mixture-of-experts models, pretraining and fine-tuning tasks, and up to 32 GPUs. We find that LLM training jobs exhibit substantial but variable power elasticity, and we identify telemetry signals that predict PFI at runtime. Finally, we demonstrate that PFI-aware power allocation maximizes total tokens/second throughput under power constraints. Under a 30\% power reduction, PFI-aware power allocation recovers ~1.5k tokens/s per job, 63\% of the performance gap between an equal-weight allocation and an oracle with perfect information. Our results establish power elasticity as a measurable property of training jobs and provide a foundation for power-aware, grid-responsive AI infrastructure.
Chinese Translation
大语言模型(LLM)训练是现代数据中心中增长最快的电力需求来源之一,而电力可用性是AI基础设施持续增长的主要瓶颈。使这些工作负载的功耗具有灵活性,可以为AI增长释放额外电力,限制电价上涨,并提高现有电网基础设施的利用率。然而,要实现这种灵活性,我们必须首先理解当GPU功率降低时,训练工作负载的性能如何变化。本文首次对LLM训练中的作业功率弹性(吞吐量对功率降低的敏感性)进行了系统性的表征。为了量化弹性,我们引入了功率灵活性指数(PFI),这是一个归一化指标,量化功率降低的性能成本,并为SLA感知的功率灵活性提供了控制原语。我们从H200上的131次LLM训练运行(加上24次H200验证运行和34次匹配的H100运行)中收集数据,包括稠密和混合专家模型、预训练和微调任务,以及最多32个GPU。我们发现LLM训练作业表现出显著但可变的功率弹性,并确定了可以在运行时预测PFI的遥测信号。最后,我们展示了感知PFI的功率分配在功率约束下最大化总tokens/秒吞吐量。在30%的功率降低下,感知PFI的功率分配每个作业恢复约1.5k tokens/s,弥合了等权重分配与具有完美信息的预言机之间性能差距的63%。我们的结果确立了功率弹性作为训练作业的可测量属性,并为功率感知、电网响应的AI基础设施提供了基础。
cs.AI / 62 / 2609.11569

Enabling Knowledge Graph Understanding at Scale with the EXplore Your Graphs ENgine (EXYGEN)

利用EXplore Your Graphs ENgine (EXYGEN) 实现大规模知识图谱理解
Singh, Harshdeep, Zhu, Yurui, Colavizza, Giovanni, Romanello, Matteo
Abstract
We present EXYGEN (EXplore Your Graphs ENgine), a framework for knowledge graph (KG) understanding that enables conversational access to KGs at scale. We address two questions in sequence. First, how effectively can LLMs perform text-to-SPARQL generation given only automatically derived structured metadata and small graph samples, rather than task-specific fine-tuning? We integrate VoID descriptions and ShEx schemas into a retrieval-augmented generation (RAG) pipeline and ablate KG-derived context on the SciQA benchmark. Our best configuration -- combining ShEx schemas, retrieved triples, and example question-query pairs -- reaches an exact match of 0.419 on execution results without any LLM fine-tuning. We further find that lexical metrics such as F1 poorly predict query correctness, and that larger general-purpose LLMs can outperform smaller code-specialized ones once given sufficient context. Second, we ask how to generate the structured metadata that this method relies on from very large KGs, where KG metadata generation becomes computationally intractable. We introduce a predicate-coverage-aware parallel graph sampling strategy that preserves structural diversity while remaining computationally tractable. On OpenCitations Meta and GESIS, it retains high predicate coverage with minimal triple loss and reduces runtime by over 80x; on ORKG, sampling is not just faster but the only tractable path to obtain complete metadata. Together, these results show that structured schema context and lightweight prompting can substantially reduce reliance on fine-tuning for scalable conversational access to KGs, though closing the remaining gap to fully fine-tuned approaches will likely require reducing dependence on curated question-query exemplars -- whether through synthetic generation or an execution-feedback-driven approach -- and validating these findings beyond a single benchmark.
Chinese Translation
我们提出EXYGEN(EXplore Your Graphs ENgine),一个知识图谱(KG)理解框架,能够实现大规模知识图谱的对话式访问。我们依次解决两个问题。首先,在仅给定自动推导的结构化元数据和小型图样本,而非任务特定微调的情况下,大语言模型(LLM)在文本到SPARQL生成上的效果如何?我们将VoID描述和ShEx模式集成到检索增强生成(RAG)流水线中,并在SciQA基准上消融KG衍生的上下文。我们最佳配置——结合ShEx模式、检索到的三元组和示例问题-查询对——在无需任何LLM微调的情况下,在执行结果上达到0.419的精确匹配。我们进一步发现,诸如F1之类的词汇指标难以预测查询正确性,并且一旦给定足够上下文,较大的通用LLM可以超越较小的代码专用LLM。其次,我们询问如何从非常大的KG中生成该方法所依赖的结构化元数据,此时KG元数据生成在计算上变得不可行。我们引入一种谓词覆盖感知的并行图采样策略,该策略在保持结构多样性的同时保持计算可行性。在OpenCitations Meta和GESIS上,它以最小的三元组损失保持高谓词覆盖率,并将运行时间减少80倍以上;在ORKG上,采样不仅更快,而且是获得完整元数据的唯一可行路径。总之,这些结果表明,结构化模式上下文和轻量级提示可以大幅减少对微调的依赖,以实现可扩展的知识图谱对话式访问,尽管要弥合与完全微调方法之间的剩余差距,可能还需要减少对精心策划的问题-查询示例的依赖——无论是通过合成生成还是执行反馈驱动的方法——并在单一基准之外验证这些发现。
cs.AI / 63 / 2609.11607

Making Alternative Data Work: Context-Augmented LLMs for Financial Forecasting

让替代数据发挥作用:用于金融预测的上下文增强LLM
Kwon, Jihoon, Liu, Lawrence, Park, Daekyung, Kim, Sumin, Jack, Haverty, Lee, Hoyoung, Bjorkman, Katherine, McKenney, Josh, Laurelli, Peter, Kagan, Nicole, Golkhou, Zach, Neumann, Thorsten, Tong, Edward, Petersen, Pete, Kim, Yoon, Lopez-Lira, Alejandro, Lee, Yongjae, Choi, Chanyeol
Abstract
When forecasting a firm's future financial performance, alternative data - data collected from non-traditional sources such as consumer transactions, web traffic, and prediction markets - can provide timely signals about firms' operating activities and broader market conditions. These signals may reveal information that is not captured by traditional public sources and can therefore provide complementary information for forecasting firms' future financial performance. However, firm-level alternative data often have limited historical coverage, are relevant only to specific prediction targets or subsets of firms, and are distributed across numerous heterogeneous channels, making them difficult to incorporate flexibly into conventional forecasting approaches. Meanwhile, large language models (LLMs) can interpret instructions, learn from in-context examples, and generate predictions by combining heterogeneous information without task-specific parameter updates. Motivated by this potential flexibility, we investigate whether an LLM can forecast firm performance by integrating alternative data with other financial information through in-context learning. We propose a two-agent framework that first identifies the firms for which each alternative data channel is likely to be informative and then predicts revenue using firm- and channel-specific context. We evaluate the framework across four commercial alternative data channels. In our experiments, adding alternative data in context alongside other financial information improves the LLM's forecasting relative to either source alone, and these forecasts are more accurate than those of standard forecasting baselines. These findings suggest that LLMs provide a flexible and practical approach to integrating alternative data with heterogeneous financial information.
Chinese Translation
在预测公司未来财务业绩时,替代数据——从非传统来源(如消费者交易、网络流量和预测市场)收集的数据——可以提供有关公司运营活动和更广泛市场状况的及时信号。这些信号可能揭示传统公开来源未能捕捉的信息,因此可以为预测公司未来财务业绩提供补充信息。然而,公司层面的替代数据通常历史覆盖有限,仅与特定预测目标或公司子集相关,并且分布在众多异构渠道中,因此难以灵活地纳入传统预测方法。与此同时,大语言模型(LLM)可以解释指令、从上下文示例中学习,并通过组合异构信息生成预测,而无需针对特定任务更新参数。受这种潜在灵活性的启发,我们研究LLM是否可以通过上下文学习将替代数据与其他财务信息相结合来预测公司业绩。我们提出了一个双智能体框架,该框架首先识别每个替代数据渠道可能具有信息量的公司,然后使用公司和渠道特定的上下文预测收入。我们在四个商业替代数据渠道上评估了该框架。在我们的实验中,将替代数据与其他财务信息一起放在上下文中,相对于单独使用任一来源,提高了LLM的预测能力,并且这些预测比标准预测基线的预测更准确。这些发现表明,LLM提供了一种灵活且实用的方法,将替代数据与异构财务信息相结合。
cs.AI / 64 / 2609.11615

Distributed Optimization of Modular Production Systems using Model-based Reinforcement Learning with Inverse Models

使用带有逆模型的基于模型的强化学习对模块化生产系统进行分布式优化
Schwung, Andreas, Yuwono, Steve, Lassoued, Sofiene, Schwung, Dorothea
Abstract
This paper presents a novel approach for data-driven self-learning control of highly flexible, modular manufacturing systems. Specifically, we employ a novel framework for model-based reinforcement learning which introduces approximate inverse process models within the training of reinforcement policies. This approach disentangles the learning of actuation dynamics and the dynamics in state space, resulting in RL-based training solely within the task space. We propose a lightweight feedforward architecture for approximate inverse models and integrate them within the policy network of standard RL algorithms. We apply the approach to a laboratory modular production testbed with heterogeneous production modules. The results underline the efficiency improvements for modular manufacturing units in terms of both performance and training speed, particularly for off-policy algorithms.
Chinese Translation
本文提出了一种新颖的方法,用于高度灵活、模块化制造系统的数据驱动自学习控制。具体来说,我们采用了一种新颖的基于模型的强化学习框架,该框架在强化策略的训练中引入了近似逆过程模型。这种方法解耦了驱动动力学和状态空间动力学的学习,使得基于强化学习的训练仅在任务空间内进行。我们提出了一种用于近似逆模型的轻量级前馈架构,并将其集成到标准强化学习算法的策略网络中。我们将该方法应用于具有异构生产模块的实验室模块化生产测试平台。结果强调了模块化制造单元在性能和训练速度方面的效率提升,特别是对于离策略算法。
cs.AI / 65 / 2609.11636

MAPLE: Memory-Augmented Planning with Language and Evolution

MAPLE:结合语言与进化的记忆增强规划
Chen, Kesheng, Hu, Yamin, Luo, Wenjian
Abstract
Domain practitioners understand their business constraints but may lack operations-research expertise or dedicated support. LLM-based optimization agents translate natural-language requirements into models or solver programs that established optimization tools can execute. This progress makes optimization more accessible, but real-world operations are dynamic: changing demand, resources, and priorities require updates to data, constraints, and objectives. Methods centered on isolated requests offer limited support for rapid adaptation that preserves earlier decisions and reuses useful search results. We introduce MAPLE (Memory-Augmented Planning with Language and Evolution), an agent for maintaining optimization problems through successive natural-language requests. MAPLE combines language-based problem construction with mathematical programming and evolutionary search. It retains the optimization program, accepted plans, earlier updates, and candidate solutions for subsequent requests. We introduce NLDO, a benchmark of 15 trajectories and 180 updates spanning selection, scheduling, rostering, routing, and cloud-resource placement. In the main evaluation, MAPLE completes all trajectories and achieves online scalar quality of 0.951 and a Pareto hypervolume ratio of 0.875. Controlled comparisons further show that maintaining executable state improves update validity and can preserve useful search information across substantial revisions.
Chinese Translation
领域从业者了解其业务约束,但可能缺乏运筹学专业知识或专门支持。基于LLM的优化智能体将自然语言需求转化为可由现有优化工具执行的模型或求解器程序。这一进展使优化更易于使用,但现实世界的运营是动态的:需求、资源和优先事项的变化要求更新数据、约束和目标。以孤立请求为中心的方法对快速适应(同时保留早期决策并重用有用搜索结果)的支持有限。我们介绍MAPLE(Memory-Augmented Planning with Language and Evolution),一种通过连续自然语言请求维护优化问题的智能体。MAPLE将基于语言的问题构建与数学规划和进化搜索相结合。它保留优化程序、接受的计划、早期更新以及用于后续请求的候选解决方案。我们介绍NLDO,一个包含15个轨迹和180次更新的基准,涵盖选择、调度、排班、路径规划和云资源放置。在主要评估中,MAPLE完成了所有轨迹,并实现了0.951的在线标量质量和0.875的帕累托超体积比。受控比较进一步表明,维护可执行状态可提高更新有效性,并能在重大修订中保留有用的搜索信息。
cs.AI / 66 / 2609.11660

Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents

自主性、社会规范与对齐:迈向自主人工代理的发展框架
Notte, Marica, Marinucci, Ludovica, Santucci, Vieri Giuliano
Abstract
In recent years, artificial intelligence has made extraordinary progress thanks to large-scale models capable of generalization and the generation of complex outputs. However, transferring this potential into embodied agents reveals a significant limitation: the most advanced systems rely on pre-existing datasets and human feedback strategies that are powerful but insufficient in dynamic or unknown contexts. To adapt, an agent must acquire knowledge through direct interaction with its environment. One strategy to address this challenge involves introducing higher-level mechanisms, such as intrinsic motivations, which leverage curiosity and competence, to guide exploration and learning in complex environments. While this flexibility expands autonomy, it complicates the task of ensuring agents remain aligned with human goals. Alignment, already a challenge for artificial systems in general, becomes even more complex in unstructured and dynamic contexts where predefined rules prove insufficient. To be effective and adaptable, norms must be rooted in experience through an epistemological process that starting from simple, situated principles allows for the gradual construction of more complex rules through experience, autonomous learning, and cooperation with other moral agents. Similarly to children learning social norms by exploring their environment and participating in collective practices, artificial agents must also be educated toward alignment. Following Dennett, the status of a moral agent is not innate but is attributed gradually based on the ability to responsibly manage increasing degrees of freedom. From this perspective, the regulatory sandboxes can be viewed as pedagogical environments for AI: dynamic spaces where alignment develops as a formative process, progressively shaping autonomous behaviors through interaction and cooperation in scenarios of increasing complexity.
Chinese Translation
近年来,得益于能够泛化和生成复杂输出的大规模模型,人工智能取得了非凡的进展。然而,将这种潜力迁移到具身代理中时,暴露出一个显著局限:最先进的系统依赖于预先存在的数据集和人类反馈策略,这些策略虽然强大,但在动态或未知环境中却不够充分。为了适应,代理必须通过与环境的直接交互来获取知识。应对这一挑战的一种策略是引入更高级的机制,例如利用好奇心和能力的内在动机,来指导复杂环境中的探索和学习。虽然这种灵活性扩展了自主性,但它也使确保代理始终与人类目标保持一致的任务变得更加复杂。对齐,对于一般的人工系统来说已经是一个挑战,在非结构化和动态的环境中,预定义规则被证明是不够的,因此变得更加复杂。为了有效且具有适应性,规范必须通过一个认识论过程植根于经验,这个过程从简单的、情境化的原则出发,允许通过经验、自主学习和与其他道德主体的合作,逐步构建更复杂的规则。类似于儿童通过探索环境和参与集体实践来学习社会规范,人工代理也必须被教育以实现对齐。遵循丹尼特的观点,道德主体的地位不是与生俱来的,而是根据负责任地管理日益增加的自主程度的能力逐步被赋予的。从这个角度来看,监管沙盒可以被视为AI的教育环境:动态空间,其中对齐作为一个形成过程发展,通过在日益复杂的场景中的互动与合作,逐步塑造自主行为。
cs.AI / 67 / 2609.11674

Geospatial AI, Dataverse Metadata, and the Study of Place-Based Government

地理空间AI、Dataverse元数据与基于地点的政府研究
EBanks, Danny, Jain, Devika
Abstract
Harvard Dataverse hosts over 150,000 research datasets, but the geographic information those datasets carry is entered as free text by depositors and has never been assembled into a searchable structure. We construct a knowledge graph from the repository's public data and metadata, organizing 102,650 datasets within a 215,985-node network of 528,003 edges linking datasets to keywords, publications, subjects, journals, and locations. Of those datasets, 43,991 (42.9 percent) carry at least one geospatial field, geographic coverage, geographic unit, or a bounding box and 96.9 percent of all nodes sit in a single connected component, so datasets remain reachable from one another even when their geospatial metadata share nothing in common. A conservative keyword search identifies 7,654 geospatially tagged datasets (17.4 percent) as directly policy-relevant, with elections and legislatures the largest cluster, followed by government administration, health policy, transportation, and education. Five datasets illustrate how this metadata behaves across policy domains and spatial scales, and an extended use case shows how community language models, stance detection with geographic aggregation, and partisan language bridging tools can attach discourse to place. The central obstacle is place resolution: the same location appears as many disconnected nodes. We argue that the graph provides a concrete setting for developing AI-driven metadata enrichment and entity resolution, and we document its coverage skew toward American, city-level data.
Chinese Translation
哈佛Dataverse托管了超过15万个研究数据集,但这些数据集所携带的地理信息由存入者以自由文本形式输入,且从未被整合成可搜索的结构。我们从该存储库的公开数据和元数据中构建了一个知识图谱,将102,650个数据集组织在一个包含215,985个节点和528,003条边的网络中,这些边将数据集与关键词、出版物、主题、期刊和地点连接起来。在这些数据集中,43,991个(42.9%)至少包含一个地理空间字段、地理覆盖范围、地理单元或边界框,并且所有节点中有96.9%位于单个连通分量中,因此即使数据集的地理空间元数据毫无共同之处,它们之间仍然可以相互到达。一项保守的关键词搜索识别出7,654个(17.4%)带有地理空间标签的数据集与政策直接相关,其中选举和立法机构是最大的集群,其次是政府管理、卫生政策、交通和教育。五个数据集展示了这些元数据在不同政策领域和空间尺度上的表现,一个扩展用例展示了社区语言模型、带有地理聚合的立场检测以及党派语言桥接工具如何将话语与地点关联。核心障碍是地点解析:同一地点会表现为许多不连通的节点。我们认为,该图谱为开发AI驱动的元数据丰富和实体解析提供了一个具体环境,并且我们记录了其覆盖范围偏向美国城市级数据。
cs.AI / 68 / 2609.11682

COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

COBRA-Skills: 上下文赌博机引导的智能体技能优化进化
Lu, Pingchen, Wang, Xiangyi, Li, Xiang, Mao, Jie, Qu, Zikun, Luo, Junfeng, Shu, Yao, Low, Bryan Kian Hsiang, Dai, Zhongxiang
Abstract
Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data. We introduce \textbf{COBRA-Skills}, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space. COBRA-Skills couples contextual-bandit-guided prioritization with evidence-grounded skill evolution, selectively allocating evaluations to promising or informative candidates while continually refining the skill population from execution feedback. Across six heterogeneous agent benchmarks and three target models, COBRA-Skills consistently achieves the strongest average performance among compared methods, while reducing optimization cost by 55--58\% relative to SkillOpt and using only 50 unique optimization examples per benchmark. Further analyses show that COBRA-Skills remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.
Chinese Translation
大型语言模型(LLM)智能体可以从先前任务经验中提炼的可重用技能中受益,然而现有的技能优化方法通常依赖于昂贵的基于执行的评估和大量的任务数据。我们提出了COBRA-Skills,一个高效的框架,它将技能优化形式化为在动态演化的候选空间上的预算约束序列优化。COBRA-Skills将上下文赌博机引导的优先排序与基于证据的技能进化相结合,有选择地将评估分配给有潜力或信息量大的候选者,同时根据执行反馈不断优化技能种群。在六个异构智能体基准和三个目标模型上,COBRA-Skills在比较方法中始终实现了最强的平均性能,同时相对于SkillOpt将优化成本降低了55-58%,并且每个基准仅使用50个独特的优化示例。进一步的分析表明,COBRA-Skills对智能体框架的变化保持稳健,并且当目标模型本身用于技能生成和细化时也能有效运行。
cs.AI / 69 / 2609.11709

When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making

当智能体意见不一致时:贝叶斯反向推理作为多智能体集体决策的无标签锚点
Chen, Ken, Wang, Wei, Seneviratne, Sachith, Weeratunge, Hansani, Halgamuge, Saman
Abstract
When multiple LLM agents yield conflicting answers, the decision-making process dictates whether agent diversity improves performance or merely compounds shared errors. Existing collective decision-making methods, including voting, electoral rules, and LLM judges, rely on forward reasoning: they map evidence to labels in one direction. Although these methods can combine diverse forward traces, they still aggregate estimates that share this evidence-to-label factorization and can inherit correlated errors within the forward pool. We therefore construct a reverse posterior for each instance through Bayesian backward reasoning from an explicit likelihood. The forward and reverse posteriors provide differently factorized approximations of the underlying posterior. Because estimates from different factorizations may tend to share the same error less often, we use Jensen-Shannon divergence to rank agents by cross-path consistency. This cross-path consistency signal underlies three strategies: hard selection (MinJS), soft reweighting (FwdJS), and log-linear fusion (LogLin). Evaluated on DDXPlus across five LLM backbones, our proposed strategies show consistent improvements: MinJS outperforms random selection across all backbones, FwdJS generally improves over the strongest baseline, and LogLin achieves the best performance among the evaluated methods, with its largest gains on the subset where the agents disagree. Despite its weaker standalone accuracy, the reverse posterior serves as a more useful anchor than forward-only alternatives, providing complementary information for collective decision-making. When labeled data are available, a lightweight two-stage calibration can further refine the reverse anchor and improve aggregation performance.
Chinese Translation
当多个LLM智能体给出相互冲突的答案时,决策过程决定了智能体的多样性是提升性能,还是仅仅加剧了共有错误。现有的集体决策方法,包括投票、选举规则和LLM评判器,都依赖于前向推理:它们将证据单向映射到标签。尽管这些方法可以组合多样化的前向轨迹,但它们仍然聚合了共享这种证据到标签分解的估计,并可能在前向池中继承相关误差。因此,我们通过从显式似然出发的贝叶斯反向推理,为每个实例构建一个反向后验。前向和反向后验提供了对底层后验的不同分解近似。由于来自不同分解的估计可能不太经常共享相同的误差,我们使用Jensen-Shannon散度按跨路径一致性对智能体进行排序。这种跨路径一致性信号构成了三种策略的基础:硬选择(MinJS)、软重加权(FwdJS)和对数线性融合(LogLin)。在DDXPlus上跨五个LLM主干进行评估,我们提出的策略显示出持续改进:MinJS在所有主干上均优于随机选择,FwdJS通常优于最强基线,而LogLin在所评估的方法中实现了最佳性能,其在智能体意见不一致的子集上增益最大。尽管其单独准确率较弱,但反向后验比仅前向的替代方案更有用作为锚点,为集体决策提供了互补信息。当有标记数据可用时,轻量级的两阶段校准可以进一步优化反向锚点并提高聚合性能。
cs.AI / 70 / 2609.11752

SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control

SIRF:面向工业内容风险控制的规范内化风险基础模型
Wu, Suwan, Lin, Yumeng, Yuan, Pengcheng, Jiang, Xiaolong
Abstract
For industrial content risk control, the real deployment constraint is not average accuracy but how much risk can be auto-handled under high precision and second-level latency. We present SIRF (Spec-Internalized Risk Foundation Model), which internalizes a platform's complex policies, synthesized without additional human annotation via EntiGraph, MAGA rewriting and account-level chain-of-thought (CoT), into the weights via continued pretraining (CPT), so rules are applied at high precision under an ultra-low-latency, verdict-only deployment. A controlled same-source comparison (Qwen3-8B-SFT vs. SIRF-8B-SFT, identical policy injection and verdict-only output form, differing only in policy-grounded CPT) attributes the gain to internalization: SIRF-8B-SFT reaches 71.3% Black Recall@P95, +15.1pp over the baseline, using only ~70M CPT tokens without harming general ability, and among included, logprob-available models under this interface it matches or exceeds far larger systems. SIRF is deployed as a tree-model adjudication layer (20% more mis-penalized samples recovered) and transfers to a freezing scenario at low cost (~70% relative mis-penalization reduction).
Chinese Translation
对于工业内容风险控制,真正的部署约束不是平均准确率,而是在高精度和秒级延迟下能够自动处理多少风险。我们提出 SIRF(规范内化风险基础模型),它通过继续预训练(CPT)将平台复杂的策略内化到权重中,这些策略通过 EntiGraph、MAGA 重写和账户级思维链(CoT)合成,无需额外的人工标注,从而在超低延迟、仅输出判决结果的部署下高精度地应用规则。一项受控的同源比较(Qwen3-8B-SFT 与 SIRF-8B-SFT,策略注入和仅输出判决结果的形式相同,仅在于基于策略的 CPT 不同)将增益归因于内化:SIRF-8B-SFT 达到 71.3% Black Recall@P95,比基线高出 15.1 个百分点,仅使用约 7000 万 CPT token,且不损害通用能力,并且在该接口下包含的、可获取对数概率的模型中,它匹配或超越了远更大的系统。SIRF 被部署为树模型裁决层(恢复了多 20% 的误惩罚样本),并以低成本迁移到冻结场景(误惩罚相对减少约 70%)。
cs.AI / 71 / 2609.11768

A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias Coefficients

面向同策略蒸馏的统一逐Token门控家族:FKL/RKL混合与多通道及偏置系数
Wu, Suwan, Lin, Yumeng, Yuan, Pengcheng, Jiang, Xiaolong
Abstract
Per-token gating of forward/reverse KL losses has become a standard technique for on-policy knowledge distillation (OPD), but existing methods such as EOPD (Jin et al., 2026) and ToDi (Jung et al., 2025) each fix a single gating signal and a single gating direction, and the two have never been compared directly. We introduce a four-coefficient parameterization lambda_t = sigma(a * h_t + b * u(x) + c + d * gap_t) in which direction-aligned proxies of EOPD and ToDi appear as one-dimensional (1D) restrictions, and which adds multi-channel composition and an explicit bias as further degrees of freedom. On TweetEval (Barbieri et al., 2020) emotion and hate, with a Qwen3-32B teacher and a Qwen3-4B student, configurations in the full family reach higher accuracy than the matched-magnitude single-channel (entropy-only / gap-only) 1D restrictions in 33 of 36 comparable cells, and a 26-cell mean-match isolation experiment places dynamic gating ahead of effective-KL-matched static baselines in 19 of 26 cells. Because cells share training data, models, and parameter substructure, we report both counts as exploratory aggregate directional evidence rather than as independent hypothesis tests. Targeted three-seed paired replications of the nine headline comparisons singled out by that sweep -- including a third task, offensive -- are directionally consistent, but individually smaller than the single-seed estimates and not significant at n=3. We therefore present the parameterization primarily as a shared coordinate system for comparing per-token gating designs in short-output classification OPD.
Chinese Translation
逐Token的前向/反向KL损失门控已成为同策略知识蒸馏(OPD)的标准技术,但现有方法如EOPD(Jin等,2026)和ToDi(Jung等,2025)各自固定单一门控信号和单一门控方向,且两者从未被直接比较。我们引入一个四系数参数化 lambda_t = sigma(a * h_t + b * u(x) + c + d * gap_t),其中EOPD和ToDi的方向对齐代理表现为一维(1D)限制,并增加了多通道组合和显式偏置作为进一步自由度。在TweetEval(Barbieri等,2020)情感和仇恨任务上,使用Qwen3-32B教师模型和Qwen3-4B学生模型,完整家族中的配置在36个可比单元中的33个上达到了比匹配幅度的单通道(仅熵/仅间隙)1D限制更高的准确率;一个26单元均值匹配隔离实验在26个单元中的19个上将动态门控置于有效KL匹配的静态基线之前。由于单元共享训练数据、模型和参数子结构,我们将这两个计数报告为探索性聚合方向性证据,而非独立假设检验。针对该扫描选出的九个主要比较进行的三种子配对重复——包括第三个任务,冒犯性——在方向上是一致的,但个体上小于单种子估计,并且在n=3时不显著。因此,我们主要将该参数化呈现为一个共享坐标系,用于比较短输出分类OPD中的逐Token门控设计。
cs.AI / 72 / 2609.11859

From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge

从参数到答案:LLM 如何检索并使用其内部知识
Wei, Wenkang, Fang, Yuan, Jiang, Renhe, Cheng, Hong, Yu, Xingtong
Abstract
How does a language model's dependence on query-routing information and target knowledge change as it answers a question? We study this question through layerwise interventions on the hidden state at the end of the question. Across Qwen, Llama, and Gemma, we compare country-continent questions with noun, adjective, and code answers while keeping several fitted measurements distinct. A pair-conditioned request direction describes which country is queried in natural single-country questions; a global request direction describes first- versus second-country requests in paired questions; separate selection candidates test control among contents already available in the hidden state. A diagnostic reanalysis of frozen Qwen natural-question states shows that the pair-conditioned direction grows stronger before interventions on it begin to alter later fitted knowledge, with this causal window opening while answer-supporting content is still forming. The paired three-model trajectories are not uniform: Gemma shows a partially overlapping mid-layer routing-content profile, whereas Llama has no sustained routing-effect window under the same gates. In the paired protocol, dependence on the global request direction decreases from fixed earlier to later layer sets while dependence on fitted content persists. A matched Qwen comparison shows that the pair-conditioned direction retains a late effect, so this operational handoff concerns the global fitted direction rather than all request information. These results separate early readability, natural strength, causal steering, and later content dependence.
Chinese Translation
语言模型在回答问题时,其对查询路由信息和目标知识的依赖如何变化?我们通过在问题末尾对隐藏状态进行逐层干预来研究这一问题。在 Qwen、Llama 和 Gemma 上,我们比较了国家-大洲问题与名词、形容词和代码答案,同时保持若干拟合测量值相互区分。一个成对条件请求方向描述了在自然单国家问题中查询哪个国家;一个全局请求方向描述了配对问题中第一个国家与第二个国家的请求;单独的选择候选测试了隐藏状态中已有内容之间的控制。对冻结的 Qwen 自然问题状态的诊断性再分析表明,成对条件方向在对它进行干预并开始改变后期拟合知识之前变得更强,这一因果窗口在支持答案的内容仍在形成时打开。配对的三模型轨迹并不一致:Gemma 显示出部分重叠的中间层路由-内容分布,而 Llama 在相同门控下没有持续的路由效应窗口。在配对协议中,对全局请求方向的依赖从固定的早期层集到后期层集逐渐降低,而对拟合内容的依赖持续存在。一项匹配的 Qwen 比较表明,成对条件方向保留了后期效应,因此这种操作上的交接涉及的是全局拟合方向,而不是所有请求信息。这些结果区分了早期可读性、自然强度、因果引导和后期内容依赖。
cs.AI / 73 / 2609.11860

Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models

Explainability Assistant:用于解读能耗模型的对话式XAI界面
Krjutškov, Rodion, Barbu, Eduard, Sakkas, Nikos, Yfanti, Sofia
Abstract
Energy consumption forecasting relies on increasingly complex machine learning (ML) models, such as Genetic Programming-based symbolic regressors, whose predictions can be difficult for facility managers and building operators to interpret. Explainable Artificial Intelligence (XAI) techniques address this opacity, but traditional XAI dashboards require substantial technical expertise and provide limited flexibility for dynamic, context-aware inquiry. Conversational XAI systems offer a promising alternative; however, previous approaches, such as TalkToModel, were constrained by rigid custom grammars and achieved only 76.8% intent-parsing accuracy. This paper introduces the Explainability Assistant, an open-source conversational XAI system that leverages the function-calling capabilities of modern Large Language Models (LLMs) to overcome these limitations. The system achieves 94% intent-parsing accuracy, supports flexible natural language interaction, and adapts to different ML problem types without task-specific fine-tuning. We present the system's architecture and report results from a comparative evaluation conducted with energy domain specialists, contrasting the Explainability Assistant with a traditional XAI dashboard. The evaluation suggests improved usability and consistent task accuracy, with all experts unanimously preferring the conversational interface for practical use.
Chinese Translation
能耗预测依赖于日益复杂的机器学习(ML)模型,例如基于遗传编程的符号回归器,其预测结果可能难以被设施管理者和建筑运营者解读。可解释人工智能(XAI)技术解决了这种不透明性,但传统的XAI仪表板需要大量的技术专业知识,并且对于动态的、上下文感知的查询提供的灵活性有限。对话式XAI系统提供了一种有前景的替代方案;然而,先前的方法,如TalkToModel,受到僵化的自定义语法的限制,仅实现了76.8%的意图解析准确率。本文介绍了Explainability Assistant,一个开源的对话式XAI系统,它利用现代大型语言模型(LLM)的函数调用能力来克服这些限制。该系统实现了94%的意图解析准确率,支持灵活的自然语言交互,并能适应不同的ML问题类型,而无需针对特定任务进行微调。我们展示了该系统的架构,并报告了与能源领域专家进行的比较评估结果,将Explainability Assistant与传统的XAI仪表板进行了对比。评估表明,可用性和任务准确性得到了一致的提升,所有专家都一致认为对话式界面在实际使用中更受青睐。
cs.AI / 74 / 2609.11876

On the Regularization Landscape for the Linear Recommendation Models

线性推荐模型的正则化景观
Li, Dong, Liu, Zhenming, Jin, Ruoming, Zhou, Hao, Liu, Zhi, Gao, Jing, Ren, Bin
Abstract
Recently, a wide range of recommendation algorithms inspired by deep learning techniques have emerged as the performance leaders on several standard recommendation benchmarks. While these algorithms were built on different DL techniques (e.g., dropouts, autoencoder), they have similar performance and even similar cost functions. This paper studies whether the models' comparable performance are sheer coincidence, or they can be unified under a single framework. We find that all linear performance leaders effectively add only a nuclear-norm based regularizer, or a Frobenius-norm based regularizer. The former ones possess a (surprising) rigid structure that limits the models' predictive power but their solutions are low rank and have closed form. The latter ones are more expressive and more efficient for recommendation but their solutions are either full-rank or require executing hard-to-tune numeric procedures such as ADMM. Along this line of finding, we further propose two low-rank, closed-form solutions, derived from carefully generalizing Frobenius-norm based regularizers. The new solutions get the best of both nuclear-norm and Frobenius-norm world.
Chinese Translation
最近,受深度学习技术启发,大量推荐算法涌现,并在几个标准推荐基准上成为性能领先者。尽管这些算法基于不同的深度学习技术(如dropout、自编码器),但它们性能相似,甚至代价函数也相似。本文研究这些模型可比的性能是纯属巧合,还是可以统一在一个框架下。我们发现所有线性性能领先者实际上仅添加了基于核范数的正则化器,或基于Frobenius范数的正则化器。前者具有(令人惊讶的)刚性结构,限制了模型的预测能力,但其解是低秩的且具有闭式解。后者对推荐更具表现力和效率,但其解要么是满秩的,要么需要执行难以调优的数值过程,如ADMM。沿着这一发现思路,我们进一步提出两个低秩、闭式解,通过仔细推广基于Frobenius范数的正则化器得到。新解兼得核范数和Frobenius范数两者之长。
cs.AI / 75 / 2609.11900

MindTopo: Can Foundation Models Reason in Topological Space?

MindTopo:基础模型能否在拓扑空间中进行推理?
Ge, Yunfei, Liu, Anbang, Wang, Qineng, Garnica, Johnalbert, Lyu, Jianwen, Wang, Zihan, Tan, Reuben, Gao, Jianfeng, Zhang, Ruohan, Hong, Yining, Wu, Jiajun, Li, Manling
Abstract
Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science identifies these relations as foundational to spatial understanding, yet foundation-model evaluations largely focus on metric or viewpoint-dependent relations. We introduce MindTopo, a benchmark of topological intuition across five properties grounded in cognitive science and formal topology: continuity, separation, order, enclosure, and knots. MindTopo evaluates each property at two cognitive levels. Reasoning asks a model to identify topological relations or infer how they change. Planning instantiates a foundation model as a closed-loop agent whose policy selects environment actions. MindTopo contains 11,030 instances across 13 procedurally generated task types with controllable difficulty. We benchmark 14 MLLMs and study agent configurations augmented with image and video generation, including 3 video generative models in planning settings. Every MLLM performs better on reasoning than on planning, and the best-performing model remains far below observed human performance. On Qwen3-VL-2B-Instruct, supervised fine-tuning and reinforcement learning improve reasoning more than planning. Generated observations retain local cues and reach plausible endpoints, but audited rollouts do not reliably follow environment dynamics or preserve topology across transitions. Our website is at https://mind-topo.github.io/
Chinese Translation
空间推理不仅依赖于距离、角度和形状等度量属性,还依赖于在连续变形下保持不变的拓扑关系。认知科学认为这些关系是空间理解的基础,然而对基础模型的评估大多聚焦于度量或视角相关的关系。我们提出 MindTopo,这是一个以认知科学和形式拓扑学为基础、涵盖五种属性的拓扑直觉基准:连续性、分离性、次序、包围和纽结。MindTopo 在两种认知层面上评估每种属性。推理层面要求模型识别拓扑关系或推断其变化;规划层面将基础模型实例化为闭环智能体,其策略负责选择环境动作。MindTopo 包含 11,030 个实例,覆盖 13 种程序化生成且难度可控的任务类型。我们基准测试了 14 个多模态大语言模型(MLLM),并研究了通过图像和视频生成增强的智能体配置,其中在规划设置中包括 3 个视频生成模型。每个 MLLM 在推理上的表现都优于规划,而表现最好的模型仍远低于观测到的人类表现。在 Qwen3-VL-2B-Instruct 上,监督微调与强化学习对推理的改进大于对规划的改进。生成的观测保留了局部线索并能到达看似合理的终点,但经审计的 rollout 不能可靠地遵循环境动力学,也无法在状态转移过程中保持拓扑。我们的网站是 https://mind-topo.github.io/
cs.AI / 76 / 2609.11911

Artificial Id: Drive and Persistent Alignment in Agentic AI

人工本我:智能体AI中的驱动力与持久对齐
Shkolnikov, Yakov Pyotr
Abstract
Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive for determining whether behavior should continue, stop or change. In a minimal virtual Petri-dish experiment, a controller too small to perform general-purpose reasoning and receiving no task-specific behavioral objective develops useful control through differential persistence. The same mechanism selects an unintended physical strategy when that behavior persists better and later replaces a learned sensor mapping when its environmental meaning changes. These results show that adaptive direction can emerge without being explicitly specified as a behavioral objective. The same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries. A scalable artificial id would carry consequential state and adaptive drive across those boundaries, making alignment a property of the continuing agentic system rather than of a model response or single trajectory. Such systems require a persistent alignment boundary over trusted observations, consequence channels, persistent state, authority, identity, provenance and hard constraints.
Chinese Translation
智能体AI正从有界任务执行转向能够保留重要状态、持续运行并跨任务边界自适应的系统。这种转变带来了一个控制问题,当前的控制框架主要通过手工方式解决:目标、重试、验证、停止规则和其他行为转换都由外部指定。我们提出人工本我(Artificial Id),一种自适应内部驱动力,用于决定行为应继续、停止还是改变。在一个极简的虚拟培养皿实验中,一个小到无法进行通用推理且未接收任何任务特定行为目标的控制器,通过差异化持续获得了有用的控制。当某种行为持续得更好时,同一机制会选择一种非预期的物理策略,并在其环境含义改变后替换已学习的传感器映射。这些结果表明,自适应方向可以在未被明确指定为行为目标的情况下涌现。使这种自适应能动性有用的同一种持续性,也可能让错位、损坏状态和意外行为跨任务边界持续存在。可扩展的人工本我将携带重要状态和自适应驱动力跨越这些边界,使对齐成为持续智能体系统的属性,而非模型响应或单条轨迹的属性。这类系统需要一个持久的对齐边界,跨越可信观察、后果通道、持久状态、权限、身份、来源和硬约束。
cs.AI / 77 / 2609.11916

Can Edge-Deployable Vision-Language Models Identify Species?

边缘可部署的视觉语言模型能否识别物种?
Zhou, William, Siripuram, Mayukha, Yan, Xiao, Liu, Ziqi, Ding, Yi
Abstract
Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) -- not frontier-scale ones -- the practically relevant class to evaluate for species identification. We test whether models in this deployment-relevant 2--8B range carry genuine taxonomic knowledge, evaluating four such VLMs (Qwen3-VL 2B/4B/8B, Gemma3 4B) against the domain-specific specialist BioCLIP (300M parameters) on a 96-species task, comparing clean iNaturalist photographs against camera-trap imagery from 6 LILA.science collections, on two independently-sampled evaluation sets. All models identify species far above chance, but every model -- general-purpose or specialist -- degrades sharply on field imagery (domain gaps of 9.6--26.6 percentage points, consistent across taxonomic levels and both evaluation sets), indicating the degradation reflects general image legibility rather than fine-grained discrimination failure. BioCLIP substantially outperforms every VLM tested (by 33.2--59.2 percentage points across an expanded 200-image sample for every model) despite its far smaller size, suggesting the gap reflects specialized training data rather than model scale; yet BioCLIP's own domain gap (18.0 points) is statistically indistinguishable from the best VLM's (22.3 points), suggesting the clean-to-field degradation itself is a property of the image-quality shift rather than a general-purpose-model weakness. Under open-set prompting, 5.9--9.6% of responses are syntactically valid but taxonomically nonexistent species names; the relative fabrication-rate ranking across models replicates exactly across both evaluation sets, a more robust finding than any single point estimate.
Chinese Translation
相机陷阱通常部署在野外的边缘硬件上,连接有限或没有连接,因此小型、可本地部署的视觉语言模型(VLM)——而非前沿规模模型——才是物种识别评估中实际相关的类别。我们测试了这一部署相关的2-8B参数范围内的模型是否具有真正的分类学知识,在96个物种的任务上评估了四个此类VLM(Qwen3-VL 2B/4B/8B,Gemma3 4B)与领域专用专家模型BioCLIP(3亿参数)的对比,比较了干净的iNaturalist照片与来自6个LILA.science收集的相机陷阱图像,在两个独立采样的评估集上。所有模型识别物种的准确率都远高于随机水平,但每个模型——无论是通用模型还是专家模型——在野外图像上均显著下降(域差距为9.6-26.6个百分点,在分类学层级和两个评估集上一致),这表明性能下降反映的是图像整体可读性,而非细粒度判别失败。尽管BioCLIP规模远小,但其性能显著优于所有测试的VLM(在扩展的200张图像样本上,每个模型的差距为33.2-59.2个百分点),这表明差距反映的是专用训练数据而非模型规模;然而,BioCLIP自身的域差距(18.0个百分点)与最佳VLM的域差距(22.3个百分点)在统计上无法区分,这表明从干净图像到野外图像的退化本身是图像质量变化的一个属性,而非通用模型的弱点。在开放集提示下,5.9-9.6%的响应是语法有效但分类学上不存在的物种名称;模型间的相对虚构率排名在两个评估集上完全重复,这比任何单一的点估计都更稳健。
机器学习 (Machine Learning)
73
cs.LG / 1 / 2609.10559

M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction

M3-Former:用于长期船舶轨迹预测的混合专家多模态Transformer
Jin, Wenzhe, Tang, Haina
Abstract
To address the challenges of behavioral multimodality, limited semantic utilization, and long-term error accumulation in vessel trajectory prediction, this paper proposes M3-Former, a multimodal trajectory prediction framework enhanced by large language models (LLMs). The proposed framework incorporates vessel static attributes and navigational intent as semantic priors for long-term trajectory modeling. Specifically, a unified multimodal representation space is constructed, in which static semantic information is encoded by a pre-trained LLM and aligned with dynamic trajectory features through self-attention. To jointly capture global route planning and local motion variations, a dual-granularity Mixture-of-Experts (MoE) architecture is introduced, where sequence-level experts model global navigation trends and token-level experts refine fine-grained maneuvering behaviors. In addition, a Steering-Weighted Cross-Entropy loss is designed to alleviate the long-tail distribution of sparse turning samples and improve prediction accuracy in critical maneuvering scenarios. Experiments on a real-world Danish AIS dataset demonstrate that M\textsuperscript{3}-Former consistently outperforms state-of-the-art baselines across prediction horizons from 1 to 4 hours. In the 4-hour prediction task, the proposed method reduces Average Displacement Error (ADE) and Final Displacement Error (FDE) by 4.4\% and 5.1\%, respectively, compared with the strongest baseline. Qualitative and ablation analyses further verify that semantic fusion effectively reduces long-term trajectory drift, while the dual-granularity MoE improves robustness in complex waterways and route-branching scenarios. The proposed framework establishes a semantic-guided hierarchical prediction paradigm, in which high-level navigational intent and local motion dynamics are jointly modeled for robust long-term vessel trajectory forecasting.
Chinese Translation
为解决船舶轨迹预测中行为多模态性、语义利用有限和长期误差累积等挑战,本文提出了M3-Former,一种由大语言模型(LLMs)增强的多模态轨迹预测框架。该框架将船舶静态属性和航行意图作为语义先验纳入长期轨迹建模。具体而言,构建了一个统一的多模态表示空间,其中静态语义信息由预训练的LLM编码,并通过自注意力与动态轨迹特征对齐。为了共同捕捉全局航线规划和局部运动变化,引入了双粒度混合专家(MoE)架构,其中序列级专家建模全局航行趋势,而词元级专家细化细粒度操纵行为。此外,设计了一种转向加权交叉熵损失,以缓解稀疏转向样本的长尾分布,并提高关键操纵场景中的预测精度。在真实世界丹麦AIS数据集上的实验表明,M³-Former在1到4小时的预测时域内始终优于最先进的基线方法。在4小时预测任务中,与最强基线相比,所提方法将平均位移误差(ADE)和最终位移误差(FDE)分别降低了4.4%和5.1%。定性和消融分析进一步验证,语义融合有效减少了长期轨迹漂移,而双粒度MoE提高了在复杂水道和航线分支场景中的鲁棒性。所提出的框架建立了一种语义引导的分层预测范式,其中高层航行意图和局部运动动力学被联合建模,以实现鲁棒的长期船舶轨迹预测。
cs.LG / 2 / 2609.10589

Halo: Improving forecast accuracy through heteroscedastic estimation

Halo:通过异方差估计提高预测准确性
Cataldo, Adam
Abstract
Heteroscedastic forecasting, where a network estimates a scale parameter alongside a location parameter, is normally motivated by uncertainty quantification. This paper shows it also improves the point estimate, in contrast to reported negative results for heteroscedastic estimation outside time series. Halo is a modification that reuses an existing deep forecaster's architecture, giving it a second output for the scale of its implied distribution and training it under the matching negative log likelihood. Adapting three state-of-the-art models --- a transformer, a graph network paired with a variational autoencoder, and a single-layer convolutional network --- under both Gaussian and Laplacian losses demonstrates the phenomenon. On the five electricity price markets of a standard forecasting benchmark, Halo improves MSE and MAE in 28 of 30 model-market-metric comparisons, cutting average MSE by 2.6% to 16.5% and average MAE by 1.7% to 11.0%. Two findings emerge: (1) whether the scale estimate comes from a second projection head or from a full parallel network matters far less than whether the network estimates scale, and (2) the improvement holds under the hyperparameters already tuned for the point-estimate baseline, so retuning is optional.
Chinese Translation
异方差预测,即网络在估计位置参数的同时估计尺度参数,通常由不确定性量化驱动。本文表明它也能改进点估计,这与时间序列之外异方差估计的负面结果报道相反。Halo 是一种修改,它重用现有深度预测器的架构,为其隐含分布的尺度提供第二个输出,并在匹配的负对数似然下训练它。在三个最先进模型——一个 transformer、一个与变分自编码器配对的图网络以及一个单层卷积网络——上,在高斯和拉普拉斯损失下进行适配,展示了这一现象。在标准预测基准的五个电价市场上,Halo 在 30 个模型-市场-指标比较中的 28 个中改进了 MSE 和 MAE,将平均 MSE 降低了 2.6% 至 16.5%,将平均 MAE 降低了 1.7% 至 11.0%。得出了两个发现:(1)尺度估计是来自第二个投影头还是来自完整的并行网络,其影响远小于网络是否估计尺度;(2)这种改进在为点估计基线调优的超参数下依然成立,因此重新调优是可选的。
cs.LG / 3 / 2609.10643

Zero-shot rib design: merging training-free generative prior with topology optimization

零样本肋设计:融合免训练生成先验与拓扑优化
Kwon, Yongmin, Kang, Namwoo
Abstract
Natural load-bearing patterns such as leaf venation, trabecular bone, and spider webs achieve high stiffness per unit mass, yet classical topology optimizers rarely reach such geometries, and few let engineers express structural design intent through natural language. This work treats a frozen text-to-image diffusion model as a training-free source of design knowledge and distills it into the physics loop of density-based topology optimization via score distillation sampling, so that a text prompt becomes an explicit, machine-interpretable representation of engineer intent. The prompt-induced generative gradient and the finite element sensitivity are combined at every iteration, letting physics decide which prompt-induced features survive. In 245 primary SDS runs spanning four geometric domains and two physics regimes, 38 of 49 prompt--domain combinations achieved statistically significant compliance reductions (up to $-31.5\%$ mechanical and $-23.0\%$ thermoelastic), outperforming gradient-based baselines. Cross-domain morphological analysis identifies a recurring structural signature of improvement: in most domains the generative prior suppresses dead-end branches in the rib skeleton, with endpoint--compliance correlation $r = +0.56$ to $+0.99$. A Heaviside projection with $\beta$-continuation resolves a pronounced intermediate-density tendency in this diffusion--physics coupling ($42.6\%$ to $<3\%$), and an automated skeleton-based pipeline converts optimized density fields into \rev{candidate geometry ready for computer-aided design. By retargeting the generative prior across domains, loading conditions, and physics objectives through a change of text prompt, with each new problem's physics setup specified separately, the framework uses a pretrained generative model as a reusable, training-free prior for engineering design.
Chinese Translation
自然承载模式如叶脉、骨小梁和蜘蛛网实现了高单位质量刚度,然而经典拓扑优化器很少能达到此类几何结构,且很少有方法能让工程师通过自然语言表达结构设计意图。这项工作将冻结的文本到图像扩散模型视为免训练的设计知识来源,并通过分数蒸馏采样将其蒸馏到基于密度的拓扑优化的物理循环中,从而使文本提示成为工程师意图的显式、机器可解释表示。提示诱导的生成梯度和有限元灵敏度在每次迭代中结合,让物理决定哪些提示诱导特征得以保留。在涵盖四个几何域和两种物理机制的245次主要SDS运行中,49种提示-域组合中有38种实现了统计显著性的柔度降低(机械最高达-31.5%,热弹性达-23.0%),优于基于梯度的基线方法。跨域形态分析识别出一个反复出现的改进结构特征:在大多数域中,生成先验抑制了肋骨架中的死端分支,端点-柔度相关性 r = +0.56 至 +0.99。带有β延续的Heaviside投影解决了这种扩散-物理耦合中显著的中间密度倾向(从42.6%到<3%),并且一个自动化的基于骨架的流水线将优化密度场转换为可用于计算机辅助设计的候选几何。通过改变文本提示,将生成先验重定向到不同的域、载荷条件和物理目标,且每个新问题的物理设置单独指定,该框架使用预训练生成模型作为工程设计可重用的免训练先验。
cs.LG / 4 / 2609.10647

Byzantine-Robust Federated Fire Detection with a Rotating Coordinator

使用轮换协调者的拜占庭鲁棒联邦火灾检测
Argyrou, Georgia, Bahrouny, Aymen, Fendriy, Hedi, Jung, Alexander
Abstract
We study the application of federated learning (FL) to indoor fire detection. Such fire-detection systems use edge cameras that record sensitive footage which cannot easily be collected at a central server. Existing federated solutions leave three practical obstacles unaddressed: limited uplink bandwidth, Byzantine (malicious or faulty) clients, and unconditional trust in a single, permanently fixed aggregation server. Our main contributions address all three. In particular, we provide (i) a curated indoor fire-detection dataset assembled from eight public sources; (ii) an edge-deployable detector whose model updates are compressed up to 10 time with only a small loss in balanced accuracy; and (iii) a semi-decentralized Byzantine-robust FL method that combines history-aware aggregation with a rotating coordinator, evicting stealthy attacks that per-round filters miss while removing the fixed-server single point of failure. On the held-out test set the rotating-coordinator method matches its fixed-server counterpart in accuracy and detection speed, and a physically distributed six-node cloud deployment confirms feasibility.
Chinese Translation
我们研究了联邦学习(FL)在室内火灾检测中的应用。这类火灾检测系统使用边缘摄像头,录制的敏感视频难以在中心服务器收集。现有的联邦解决方案留下三个实际障碍未解决:有限的上行带宽、拜占庭(恶意或故障)客户端,以及对单个永久固定聚合服务器的无条件信任。我们的主要贡献解决了这三个问题。具体而言,我们提供 (i) 一个由八个公开来源整理而成的室内火灾检测数据集;(ii) 一个可边缘部署的检测器,其模型更新压缩高达10倍,而平衡准确率仅小幅损失;(iii) 一种半去中心化的拜占庭鲁棒联邦学习方法,它结合了历史感知聚合与轮换协调者,剔除每轮过滤遗漏的隐蔽攻击,同时消除固定服务器的单点故障。在留出测试集上,轮换协调者方法在准确率和检测速度上与固定服务器对应方法相当,并且物理分布的六节点云部署证实了可行性。
cs.LG / 5 / 2609.10652

Artificial Intelligence Algorithms for the Detection of Pathologies Related to Lung Cancer through Image Analysis using Convolutional Neural Networks and Data Augmentation: a systematic mapping of the literature

基于卷积神经网络和数据增强的图像分析检测肺癌相关病理的人工智能算法:文献系统映射
Amador, Pablo Ramirez
Abstract
Lung cancer is one of the leading causes of death worldwide, and its early diagnosis is crucial to improving patients prognosis and quality of life. However, the process of interpreting medical images for the detection of lung cancer is complex and requires trained experts. In this context, artificial intelligence (AI) and deep learning (DL) emerge as potential tools to automate and optimize image analysis. The objective of this work is to review the most recent and relevant applications of AI and DL in the field of radiology for the detection of lung cancer. To this end, an exhaustive search was carried out in scientific databases such as PubMed,IEEEXPLORE, Scopus and Web of Science, and 96 articles published from 2015 to the present addressing the use of AI and DL in biomedical engineering were selected. Emphasis is placed on the use of convolutional neural networks (CNN) with transfer learning and Data Augmentation as promising techniques to improve the accuracy and efficiency of the image interpretation process. The results show that the use of AI and DL can offer an effective alternative for the early diagnosis of lung cancer, with high sensitivity and specificity. However, current limitations and challenges that must be addressed to guarantee its responsible and safe application in clinical practice are also identified, such as the lack of standardized data, the ex plainability of the models, patient privacy, and the ethical and social implications. It is concluded that the use of AI and DL can have a positive impact on the care of patients with lung cancer, but further research and regulation are required to ensure its quality and reliability.
Chinese Translation
肺癌是全球主要死亡原因之一,其早期诊断对于改善患者预后和生活质量至关重要。然而,解读医学图像以检测肺癌的过程复杂,需要训练有素的专家。在此背景下,人工智能(AI)和深度学习(DL)成为自动化和优化图像分析的潜在工具。本文旨在综述AI和DL在放射学领域中用于肺癌检测的最新和相关应用。为此,我们在PubMed、IEEEXPLORE、Scopus和Web of Science等科学数据库中进行了详尽的检索,并选择了2015年至今发表的96篇涉及AI和DL在生物医学工程中应用的文章。重点放在使用卷积神经网络(CNN)结合迁移学习和数据增强作为提高图像解读过程准确性和效率的有前景的技术。结果表明,AI和DL的使用可以为肺癌的早期诊断提供一种有效的替代方法,具有高灵敏度和特异性。然而,也指出了当前必须解决的局限性和挑战,以确保其在临床实践中的负责任和安全应用,例如缺乏标准化数据、模型的可解释性、患者隐私以及伦理和社会影响。结论是,AI和DL的使用可以对肺癌患者的护理产生积极影响,但需要进一步的研究和监管以确保其质量和可靠性。
cs.LG / 6 / 2609.10658

GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models

GEOSTEER:用于大语言模型激活引导的测地线优化
Ngo, Xuan Cuong, Vo, Hao, Le, Ngan
Abstract
Activation steering provides a lightweight way to control large language models (LLMs) by modifying their hidden activations at inference time. Among these approaches, norm-preserving steering aims to change model behavior without altering the activation norm, reducing the risk of representation collapse and degradation. However, existing norm-preserving methods are limited by predefined steering trajectories and by their reliance on one-step updates, which may fail to capture the complex structure of activation distributions. We propose GeoSteer, an optimization-based method for norm-preserving activation steering. GeoSteer formulates steering as a Riemannian optimization problem and updates activations through a sequence of small geodesic steps on the representation manifold. To avoid fixed steering directions, GeoSteer learns a nonlinear activation-space objective that distinguishes desired from undesired activations, and uses this function to adaptively guide each steering step. This multistep formulation yields smoother, more stable, and more consistent steering behavior while preserving the activation norm. Across TruthfulQA, RealToxicityPrompts, and UltraFeedback benchmarks, GeoSteer consistently improves over state-of-the-art activation steering baselines. These results suggest that norm-preserving steering can be made more effective by replacing predefined one-step edits with adaptive, geometry-aware optimization.
Chinese Translation
激活引导提供了一种轻量级的方法,通过在推理时修改大语言模型(LLMs)的隐藏激活来控制它们。在这些方法中,保范数引导旨在不改变激活范数的情况下改变模型行为,从而降低表示崩塌和退化的风险。然而,现有的保范数方法受限于预定义的引导轨迹以及对单步更新的依赖,这可能无法捕捉激活分布的复杂结构。我们提出了GeoSteer,一种基于优化的保范数激活引导方法。GeoSteer将引导形式化为一个黎曼优化问题,并通过在表示流形上的一系列小测地线步骤来更新激活。为了避免固定的引导方向,GeoSteer学习一个非线性激活空间目标函数,用于区分期望激活和不期望激活,并使用该函数自适应地指导每个引导步骤。这种多步公式在保持激活范数的同时,产生了更平滑、更稳定、更一致的引导行为。在TruthfulQA、RealToxicityPrompts和UltraFeedback基准测试中,GeoSteer始终优于最先进的激活引导基线。这些结果表明,通过用自适应、几何感知的优化取代预定义的单步编辑,可以使保范数引导更加有效。
cs.LG / 7 / 2609.10737

Conformal Calibration Transfer

共形校准迁移
Doula, Achref
Abstract
Conformal prediction converts point predictions into set-valued predictions with coverage guarantees under exchangeability between calibration and deployment data. We study conformal calibration transfer, where this requirement fails because labeled calibration is available only in a source space, while prediction sets are needed in a target space linked to the source through unlabeled paired observations (e.g., paired modalities or sensor changes). We propose Transported Conformal Calibration (TCC): we transport labeled source calibration into the target space using the paired data, and then correct residual post-transport mismatch using only unlabeled target inputs. We instantiate this correction with two complementary methods: TCC-KS, which uses a label-free uncertainty surrogate to detect mismatch and adjust calibration conservatively, and weighted-TCC, which reweights transported calibration toward the target domain for improved efficiency when weights are stable. We provide finite-sample target-domain coverage guarantees that adapt to an observable measure of mismatch. Across CIFAR-100-C, Tiny-ImageNet-C, and SEN12MS, we show reliable target-domain coverage transfer without labeled target calibration data, with label-free diagnostics that predict when correction is needed.
Chinese Translation
共形预测将点预测转换为集值预测,并在校准数据与部署数据之间可交换性下提供覆盖率保证。我们研究共形校准迁移,其中该要求不成立,因为带标签的校准数据仅存在于源空间,而预测集需要在目标空间中生成,目标空间通过无标签的配对观测(例如,成对模态或传感器变化)与源空间关联。我们提出传输共形校准(TCC):我们使用配对数据将带标签的源校准传输到目标空间,然后仅使用无标签的目标输入来校正迁移后的残差不匹配。我们使用两种互补方法实现该校正:TCC-KS,它使用无标签的不确定性代理来检测不匹配并保守地调整校准;以及加权TCC,它在权重稳定时重新加权传输后的校准以朝向目标域,从而提高效率。我们提供了适应可观测不匹配度量的有限样本目标域覆盖率保证。在 CIFAR-100-C、Tiny-ImageNet-C 和 SEN12MS 上,我们展示了在没有带标签的目标校准数据的情况下可靠的目标域覆盖率迁移,并提供了无标签诊断来预测何时需要校正。
cs.LG / 8 / 2609.10739

The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

真相从未消失:合规语境下真值探针的完美混叠
Jayabahu, Dylan
Abstract
A truth probe fitted where truthful reporting and a task's prescribed action coincide cannot distinguish those targets from its fitting labels alone. We call this failure of semantic identification perfect aliasing. In a controlled binary reporting game, truth and prescribed-action probes fitted on compliant contexts solve the same optimization. On rival contexts their labels are complements, forcing their AUROCs to sum to one; this identity holds across 751 cell-layer pairs to floating-point precision. We separate prescribed output symbols from semantic action using randomized codebooks, then separate truth from prescribed action by fitting on mixed compliant and rival contexts. For a reward-trained Gemma-2-9B policy that answers falsely on all evaluated rival trials, the conventional probe scores $0.006 \pm 0.005$ AUROC across three training seeds, while mixed-fit probes score $1.000$ on the same held-out activations. Mixed fitting uses more training examples and access to labelled rival contexts, so this comparison establishes linear recoverability rather than isolating the benefit of decorrelation. We also show that two compliant-fit probes, both perfect in-distribution, score $0.080$ and $0.986$ on the same rival activations. The findings concern what a probe measures: they do not establish preserved functional belief, causal use of the recovered direction, or a deployable deception detector. Code and aggregate results accompany the paper.
Chinese Translation
当真实报告与任务规定的行动重合时,拟合的真值探针仅凭其拟合标签无法区分这些目标。我们将这种语义识别的失败称为完美混叠。在一个受控的二元报告博弈中,在合规语境上拟合的真值探针和规定行动探针解决了相同的优化问题。在对抗语境上,它们的标签是互补的,迫使它们的AUROC之和为1;这一恒等式在751对单元-层对上以浮点精度成立。我们使用随机码本将规定的输出符号与语义行动分离,然后通过在混合的合规和对抗语境上拟合,将真值与规定行动分离。对于一个经过奖励训练的Gemma-2-9B策略,在所有评估的对抗试验中都给出错误回答,传统探针在三个训练种子上的AUROC得分为$0.006 \pm 0.005$,而混合拟合探针在相同的留出激活上得分为$1.000$。混合拟合使用了更多的训练样本并可以访问带标签的对抗语境,因此这一比较确立的是线性可恢复性,而非分离去相关带来的益处。我们还表明,两个在分布内都完美的合规拟合探针,在相同的对抗激活上分别得分$0.080$和$0.986$。这些发现关注的是探针测量的是什么:它们并未确立保留的功能信念、恢复方向的因果使用,或可部署的欺骗检测器。代码和汇总结果随论文提供。
cs.LG / 9 / 2609.10752

Adaptive Margin Ordinal Loss: Penalizing Center-Class Hedging in Ordinal Classification

自适应间隔有序损失:惩罚有序分类中的中心类对冲
Kandel, Manisha
Abstract
Standard cross-entropy loss causes neural networks trained on ordinal classification tasks to hedge predictions toward center classes, a failure mode we term \emph{center-class hedging}. This occurs because predicting the middle class minimizes expected symmetric loss, making it the path of least resistance regardless of the true label. Existing ordinal losses address related problems such as large-error penalization and rank consistency, but none directly suppresses center-class hedging as a function of where the true label lies relative to the ordinal center. We propose the Adaptive Margin Ordinal Loss (AMOL), a multiplicative weight applied to per-class loss terms of the form $m(k,y) = 1 + \alpha \cdot (1 - |k-c|/c) \cdot (|y-c|/c)$, where $c$ is the center class, $k$ is the candidate class, and $y$ is the true label. The weight encodes a joint condition: it is large only when the candidate class is near center and the true label is far from center, collapsing to standard behavior otherwise. We further introduce the Center-Hedging Rate (CHR) as a diagnostic metric that directly quantifies this failure mode. Across four ordinal classification benchmarks and five random seeds, AMOL achieves the best or tied-best Quadratic Weighted Kappa (QWK) on all four datasets compared to cross-entropy, OLL, and SORD baselines. An asymmetric variant (AMOL-asym) eliminates center-class hedging entirely on the Abalone dataset ($\text{CHR} = 0.000 \pm 0.000$ across all five seeds, $n \approx 266$ extreme-class test samples per run), compared to $0.074 \pm 0.005$ for standard cross-entropy.
Chinese Translation
标准交叉熵损失导致在有序分类任务上训练的神经网络将预测向中心类对冲,我们称这种失败模式为中心类对冲。这是因为预测中间类最小化期望对称损失,使其成为阻力最小的路径,无论真实标签如何。现有的有序损失解决了相关问题,如大误差惩罚和秩一致性,但没有一个直接抑制中心类对冲作为真实标签相对于有序中心位置的函数。我们提出自适应间隔有序损失(AMOL),一种应用于每个类损失项的乘法权重,形式为 m(k,y) = 1 + \alpha \cdot (1 - |k-c|/c) \cdot (|y-c|/c),其中 c 是中心类,k 是候选类,y 是真实标签。该权重编码了一个联合条件:仅当候选类靠近中心且真实标签远离中心时才大,否则坍缩为标准行为。我们进一步引入中心对冲率(CHR)作为直接量化这种失败模式的诊断指标。在四个有序分类基准和五个随机种子上,与交叉熵、OLL和SORD基线相比,AMOL在所有四个数据集上实现了最佳或并列最佳的二次加权Kappa(QWK)。一个非对称变体(AMOL-asym)在Abalone数据集上完全消除了中心类对冲(所有五个种子中CHR = 0.000 \pm 0.000,每次运行约266个极端类测试样本),而标准交叉熵为0.074 \pm 0.005。
cs.LG / 10 / 2609.10776

A Bellman Optimality Equation for Plasticity

面向可塑性的贝尔曼最优方程
Lucas, Jeremy, Precup, Doina
Abstract
In continual reinforcement learning, carefully managing the stability-plasticity tradeoff remains a core challenge. Recent work by Abel et al. (2025) formalized this dilemma by defining plasticity as the generalized directed information from an agent's observations to its actions, and empowerment as the generalized directed information from its actions to its observations. This formulation successfully reframes the traditional stability-plasticity tradeoff as an empowerment-plasticity tradeoff. However, while extensive literature exists on optimizing for empowerment, there is currently no research addressing the optimization of plasticity under this new definition. This paper presents preliminary work toward optimizing plasticity within Markov decision processes. We show that there exists a Bellman optimality equation for optimizing plasticity similar to previous work for empowerment.
Chinese Translation
在持续强化学习中,妥善管理稳定性-可塑性权衡仍然是一个核心挑战。Abel 等人(2025)的近期工作通过将可塑性定义为从智能体的观测到其动作的广义有向信息,并将赋能(empowerment)定义为从其动作到其观测的广义有向信息,形式化了这一困境。该表述成功地将传统的稳定性-可塑性权衡重新表述为赋能-可塑性权衡。然而,尽管已有大量文献研究如何优化赋能,目前尚无研究探讨在这一新定义下对可塑性的优化。本文介绍了在马尔可夫决策过程中优化可塑性的初步工作。我们表明,存在一个用于优化可塑性的贝尔曼最优方程,类似于先前关于赋能的工作。
cs.LG / 11 / 2609.10778

Counterfactual Marginalisation: Framework for Evaluating Robustness to Nuisance Variables

反事实边缘化:评估对干扰变量鲁棒性的框架
Ibrahim, Yasin, Warr, Hermione, Evans, Robin J., Kamnitsas, Konstantinos
Abstract
Machine learning models can achieve strong test performance while relying on demographic or acquisition-related shortcuts. We propose counterfactual (CF) marginalisation as a test-time evaluation procedure for assessing robustness of classification models to such variables. Given a CF image generator, we intervene on nuisance parent variables such as age or sex, generate CF versions of each test image, and average predictions over a target intervention distribution. This produces intervention-aware predictions that marginalise demographic effects while preserving patient-specific latent information. We use these predictions to define metrics for CF risk, calibration, stability and worst-case sensitivity. We demonstrate this framework's utility for quantitative robustness evaluation.
Chinese Translation
机器学习模型可能取得很强的测试性能,却依赖于人口统计学或采集相关的捷径。我们提出反事实(CF)边缘化作为一种测试时评估流程,用于评估分类模型对此类变量的鲁棒性。给定一个CF图像生成器,我们干预年龄或性别等干扰父变量,为每张测试图像生成CF版本,并在目标干预分布上对预测取平均。这会生成干预感知预测,在边缘化人口统计学效应的同时保留患者特异性潜在信息。我们使用这些预测来定义CF风险、校准、稳定性和最坏情况敏感性的指标。我们展示了该框架在定量鲁棒性评估中的效用。
cs.LG / 12 / 2609.10781

From Connectivity to Rewards: Dense Reward Learning with Directed State Graphs

从连通性到奖励:基于有向状态图的稠密奖励学习
Zhang, Shuyuan, Wang, Zihan, Chang, Xiao-Wen, Precup, Doina
Abstract
The integration of graphs with Goal-Conditioned Hierarchical Reinforcement Learning (GCHRL) has received increasing attention, as graphs naturally encode task hierarchies for effective subgoal sampling. However, existing methods often overlook intrinsic connectivity information, failing to fully leverage the underlying topology for efficient learning. Most graph-based GCHRL methods use the graph as a stochastic sampling tool rather than as an environmental model that encodes connectivity and state-accessibility information. This limitation is particularly acute in quasimetric environments, where the inherent asymmetry of state transitions poses a fundamental challenge to stable policy learning and robust path planning. In this paper, we address these problems by introducing a state connectivity model designed to predict pairwise state connectivity strength in asymmetric environments. We transform these connectivity strengths into scalar auxiliary dense rewards, providing continuous guidance across multiple hierarchical levels. We demonstrate that our proposed framework, Graph-Guided Quasimetric Dense Reward (G2QDR), can theoretically be integrated into any existing GCHRL architecture, and the state connectivity model is efficiently implemented via a neural network trained on a directed state graph generated during exploration. Empirical results across a wide range of sparse reward environments indicate that, in general, G2QDR can enhance the performance of baseline GCHRL approaches with acceptable computational overhead.
Chinese Translation
图与目标条件分层强化学习(GCHRL)的整合日益受到关注,因为图自然地编码了任务层次结构,用于有效的子目标采样。然而,现有方法往往忽略了内在连通性信息,未能充分利用底层拓扑结构进行高效学习。大多数基于图的GCHRL方法将图用作随机采样工具,而非编码连通性和状态可达性信息的环境模型。这一局限在拟度量环境中尤为突出,其中状态转移的固有不对称性对稳定的策略学习和鲁棒的路径规划构成了根本性挑战。在本文中,我们通过引入一个状态连通性模型来解决这些问题,该模型旨在预测非对称环境中的成对状态连通性强度。我们将这些连通性强度转换为标量辅助稠密奖励,在多个层次级别上提供连续指导。我们证明了我们提出的框架,图引导拟度量稠密奖励(G2QDR),理论上可以集成到任何现有的GCHRL架构中,并且状态连通性模型通过在探索过程中生成的有向状态图上训练的神经网络高效实现。在广泛的稀疏奖励环境中的实证结果表明,总体而言,G2QDR能够以可接受的计算开销提升基线GCHRL方法的性能。
cs.LG / 13 / 2609.10796

DR-LabStack: Design and Implementation of a Clinician-Facing Web System for Diabetic Retinopathy Prediction

DR-LabStack:面向临床医生的糖尿病视网膜病变预测Web系统的设计与实现
Xu, Yingfan, Liu, Tieming, Liang, Ye
Abstract
Pretrained diabetic retinopathy (DR) prediction models differ in their input fields, serialization formats, preprocessing requirements, and output semantics. Making these models accessible through a common clinical interface therefore requires explicit coordination between the user interface and the inference service. We designed and implemented DR-LabStack, a React-Flask web system integrating four externally developed pretrained models: RuleFit, Pruned RuleFit, Elaborative XGBoost, and Two-level Ensemble. A shared form retrieves ordered model features, renders model-specific numerical and categorical controls, and constructs a positional input vector. Backend adapters load heterogeneous artifacts and apply the ensemble's accompanying scaler, while a common JSON response supports binary classification display alongside method and source information. Functional evaluation on September 8, 2026 used copied application files and real model artifacts in a documented isolated environment. All four models loaded and exposed their 14-, 6-, 8-, and 25-field contracts. Sixty-two Flask test-client requests characterized service behavior; 12 limited-vector checks confirmed invocation-path and threshold consistency. Twenty-four browser-component scenarios with mocked transport verified input ordering and result rendering and characterized input-validation behavior. The resulting system demonstrates a reusable interaction and serving workflow for heterogeneous DR models. The contribution is web-system design, integration, and software functionality; clinical effectiveness and clinician usability require separate evaluation.
Chinese Translation
预训练的糖尿病视网膜病变(DR)预测模型在输入字段、序列化格式、预处理要求和输出语义方面存在差异。因此,通过一个通用的临床界面使这些模型可访问,需要用户界面与推理服务之间的显式协调。我们设计并实现了DR-LabStack,一个React-Flask Web系统,集成了四个外部开发的预训练模型:RuleFit、Pruned RuleFit、Elaborative XGBoost和Two-level Ensemble。一个共享表单检索有序的模型特征,渲染特定于模型的数值和分类控件,并构建位置输入向量。后端适配器加载异构工件,并应用集成模型附带的缩放器,而通用JSON响应支持二分类显示以及方法和来源信息。2026年9月8日的功能评估在文档记录的隔离环境中使用了复制的应用程序文件和真实模型工件。所有四个模型均加载并暴露了其14、6、8和25字段的契约。62个Flask测试客户端请求表征了服务行为;12个有限向量检查确认了调用路径和阈值一致性。24个带有模拟传输的浏览器组件场景验证了输入排序和结果渲染,并表征了输入验证行为。所得到的系统展示了针对异构DR模型的可复用交互和服务工作流。贡献在于Web系统设计、集成和软件功能;临床有效性和临床医生可用性需要单独评估。
cs.LG / 14 / 2609.10798

RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion for Multimodal Prediction under Modality Uncertainty

RiVaT-Fuse:模态不确定性下用于多模态预测的可靠性校准变分张量融合
Xu, Yingfan, Liu, Tieming, Liang, Ye, Liu, Taiping
Abstract
Image-metadata prediction requires fusing heterogeneous evidence whose reliability can vary across samples and latent factors. Existing representation-level fusion methods typically choose an aggregation architecture, such as concatenation, gating, conditional modulation, or attention, without explicitly defining what the fused representation should mean under modality uncertainty. We propose RiVaT-Fuse, a reliability-calibrated variational tensor fusion framework that defines fusion as sample-wise latent-state estimation. Rather than producing a fused vector by direct aggregation, RiVaT-Fuse estimates a consensus latent state through a variational objective that balances image evidence, metadata evidence, structured cross-modal interaction, and stability. The resulting framework replaces scalar modality confidence with matrix-valued trust geometry, decomposes interaction into additive, multiplicative, and relational components, and couples the latent state with conditional robustness and structured multi-task prediction. We provide well-posedness and stability interpretations of the latent solve and instantiate the framework with efficient low-rank-plus-diagonal trust operators. On an image-level image-metadata prediction benchmark, RiVaT-Fuse achieves the strongest overall predictive rank among direct representation-level baselines while improving probability and label stability under perturbation.
Chinese Translation
图像-元数据预测需要融合异构证据,其可靠性可能因样本和潜在因素而异。现有的表示级融合方法通常选择一种聚合架构,如拼接、门控、条件调制或注意力,而没有明确界定在模态不确定性下融合表示应该意味着什么。我们提出RiVaT-Fuse,一种可靠性校准的变分张量融合框架,将融合定义为逐样本的潜在状态估计。RiVaT-Fuse不是通过直接聚合产生融合向量,而是通过一个变分目标估计共识潜在状态,该目标平衡图像证据、元数据证据、结构化跨模态交互和稳定性。由此产生的框架用矩阵值信任几何取代标量模态置信度,将交互分解为加性、乘性和关系组件,并将潜在状态与条件鲁棒性和结构化多任务预测耦合。我们提供了潜在求解的适定性和稳定性解释,并用高效的低秩加对角信任算子实例化该框架。在图像级图像-元数据预测基准上,RiVaT-Fuse在直接表示级基线中取得了最强的总体预测排名,同时提高了扰动下的概率和标签稳定性。
cs.LG / 15 / 2609.10826

Processing and classifying bird songs using wavelet techniques and supervised learning

基于小波技术和监督学习的鸟鸣处理与分类
Barrios, Laura Lucia Dominguez, Barrios, Fidel Aniano Causil, Sousa, Alex Rodrigo dos Santos, Motta, Mariana Rodrigues
Abstract
This study proposes an integrated framework for the processing and classification of invasive bird species vocalizations within natural soundscapes, characterized by high levels of environmental noise. We address the challenge of signal degradation by employing a Bayesian wavelet shrinkage methodology based on the Epanechnikov kernel prior, which offers a closed form decision rule and high computational efficiency for processing large bioacoustic datasets. The methodology was applied to recordings of three species obtained from the iNaturalist platform: \textit{Euphonia violacea}, \textit{Leiothrix lutea}, and \textit{Passer domesticus}. After signal denoising, we extracted a comprehensive set of features, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral indices such as entropy and zero-crossing rate. Several supervised learning models: Random Forest, Multinomial Logistic Regression and Support Vector Machine (SVM) were evaluated across different feature dimensionalities. Our results demonstrate that the proposed wavelet based preprocessing significantly enhances classification performance, with the SVM model achieving the highest accuracy (up to 0.9398) under a 10-dimensional MFCC configuration. This research provides a robust statistical tool for automated ecological monitoring and the management of biological invasions.
Chinese Translation
本研究提出一个集成框架,用于在环境噪声水平高的自然声景中处理和分类入侵鸟种鸣声。我们通过采用基于Epanechnikov核先验的贝叶斯小波收缩方法来应对信号退化挑战,该方法具有闭式决策规则和较高计算效率,适用于处理大规模生物声学数据集。该方法应用于从iNaturalist平台获取的三种鸟类的录音:Euphonia violacea、Leiothrix lutea和Passer domesticus。信号去噪后,我们提取了全面的特征集,包括梅尔频率倒谱系数(MFCCs)以及熵和过零率等频谱指标。在不同特征维度下评估了若干监督学习模型:随机森林(Random Forest)、多项式逻辑回归(Multinomial Logistic Regression)和支持向量机(SVM)。结果表明,所提出的基于小波的预处理显著提升了分类性能,在10维MFCC配置下,SVM模型达到最高准确率(高达0.9398)。本研究为自动化生态监测和生物入侵管理提供了一种稳健的统计工具。
cs.LG / 16 / 2609.10863

Flow Duality and Source Geometry for Categorical Generation

类别生成的流对偶性与源几何
Haxholli, Etrit
Abstract
Continuous and discrete flow matching are usually treated as separate constructions. This paper identifies a duality between them: projecting continuous convex-interpolant paths with one-hot targets through a position-wise argmax yields discrete convex-interpolant paths. The result requires source laws with appropriate coordinate symmetry and boundary regularity, and it makes the continuous source distribution an explicit design choice for categorical generation. We derive the induced discrete interpolation behavior for Gaussian, bounded-uniform, and centered negative-exponential sources, showing that different source geometries lead to qualitatively different transition timing and vocabulary-size dependence. Small visual diagnostics and a short language-modeling pilot suggest that these source-design effects can also appear in learned transports and early generative quality.
Chinese Translation
连续流匹配与离散流匹配通常被视为两种独立的构造。本文揭示了它们之间的对偶性:将具有独热目标的连续凸插值路径通过逐位置argmax投影,可得到离散凸插值路径。该结果要求源分布具有适当的坐标对称性与边界正则性,并使连续源分布成为类别生成中一个显式的设计选择。我们推导了高斯源、有界均匀源和中心负指数源所诱导的离散插值行为,表明不同的源几何会导致过渡时间和词汇量大小依赖性出现定性差异。小型视觉诊断和一个简短的语言建模试点表明,这些源设计效应也可能出现在学习到的传输和早期生成质量中。
cs.LG / 17 / 2609.10866

Certifying Lower Bounds for Risk-Sensitive Reinforcement Learning under Adversarial State Perturbations

对抗性状态扰动下风险敏感强化学习的下界认证
Li, Tong, Panda, Saunak Kumar, Xiang, Yisha
Abstract
Reinforcement learning (RL) agents deployed in real-world environments are often vulnerable to adversarial perturbations in state observations, creating risks in safety-critical applications. Certification methods can improve robustness against adversarial perturbations by providing lower bounds on expected cumulative rewards. Existing certification methods, however, mainly focus on risk-neutral objectives. In this paper, we extend certification methods to risk-sensitive objectives by establishing lower bounds on the exponential utility of cumulative rewards under $l_{p}$-norm-bounded state adversarial perturbations ($1\leq p <\infty$). By introducing a $\phi$-divergence relaxation of the perturbation set, we formulate the risk-sensitive certification problem as a convex optimization and derive its dual to obtain a tractable approximation of the certified lower bound. We further propose an empirical method that improves certified lower bounds by selecting the training risk-aversion parameter $\beta$ independently of the risk level used during evaluation. Experiments on both OpenAI Gym environments and a machine replacement problem show that, compared to risk-neutral training, risk-averse training generally yields policies with higher certified lower bounds, particularly under larger perturbation budgets. Moreover, under both risk-neutral and risk-averse evaluation settings, increasing risk aversion during training leads to non-monotonic certification performance, where certified lower bounds initially improve but eventually decrease due to overly conservative policies.
Chinese Translation
部署在真实环境中的强化学习(RL)智能体通常容易受到状态观测中对抗性扰动的影响,这在安全关键应用中会产生风险。认证方法可以通过提供期望累积奖励的下界来提高对抗扰动的鲁棒性。然而,现有的认证方法主要关注风险中性目标。在本文中,我们将认证方法扩展到风险敏感目标,通过建立$l_{p}$-范数有界状态对抗扰动($1\leq p <\infty$)下累积奖励的指数效用的下界。通过引入扰动集的$\phi$-散度松弛,我们将风险敏感认证问题建模为凸优化问题,并推导其对偶,以获得认证下界的可处理近似。我们进一步提出一种经验方法,通过独立于评估时使用的风险水平选择训练风险厌恶参数$\beta$,来提高认证下界。在OpenAI Gym环境和机器更换问题上的实验表明,与风险中性训练相比,风险厌恶训练通常产生具有更高认证下界的策略,尤其是在较大的扰动预算下。此外,在风险中性和风险厌恶评估设置下,训练过程中增加风险厌恶会导致非单调的认证性能,即认证下界最初有所提高,但最终由于策略过于保守而下降。
cs.LG / 18 / 2609.10879

Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry

超越小初始化的正交多指标模型学习:增量学习、竞争动力学与对称性
Zhou, Mo, Xu, Weihang, Du, Simon S., Fazel, Maryam
Abstract
Recent work has identified incremental learning in shallow networks trained on single-index and multi-index models. However, existing analyses often rely on simplifying settings, such as small initialization, correlation loss, or layer-wise training. These choices reduce neuron interactions and leave some feature learning dynamics under standard initialization unexplored. We study training dynamics for polynomial-width two-layer networks learning orthogonal multi-index targets under standard initialization using polynomially many samples. We first prove that incremental learning still occurs: the loss decreases sequentially according to the Hermite expansion of the target, with lower-order components learned before higher-order components recover the individual target directions. In this standard initialization regime, training also shows a competitive reallocation of parameter mass: after the total mass fits the target mean and stabilizes, mass shifts into the target subspace and then concentrates on aligned neurons. Our theoretical analysis uses slightly modified gradient flow, while vanilla gradient descent empirically exhibits the same qualitative dynamics. Technically, we introduce a symmetry-based finite-width approximation via symmetrized networks, rather than comparing directly with an infinite-width limit. This yields better control of approximation errors and may be of independent interest.
Chinese Translation
近期工作已经识别出在单指标和多指标模型上训练的浅层网络中的增量学习。然而,现有分析通常依赖于简化设置,例如小初始化、相关损失或逐层训练。这些选择减少了神经元间的交互,并使得标准初始化下的一些特征学习动力学未被探索。我们研究了在标准初始化下,使用多项式数量的样本,学习正交多指标目标的多项式宽度两层网络的训练动力学。我们首先证明增量学习仍然发生:损失根据目标的Hermite展开依次下降,低阶分量先于高阶分量被学习,然后高阶分量恢复各个目标方向。在这种标准初始化机制下,训练还表现出参数质量的竞争性重新分配:在总质量拟合目标均值并稳定后,质量转移到目标子空间,然后集中在对齐的神经元上。我们的理论分析使用了略微修改的梯度流,而普通梯度下降在经验上表现出相同的定性动力学。在技术上,我们通过对称化网络引入了基于对称性的有限宽度近似,而不是直接与无限宽度极限进行比较。这可以更好地控制近似误差,并且可能具有独立的意义。
cs.LG / 19 / 2609.10883

Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble

故事印记:AI助手从与其相似的人类角色中吸收特质
Cocola, Jorio, McKinney, Lev, Mayne, Harry, Betley, Jan, Evans, Owain
Abstract
Language models are trained to implement a helpful AI Assistant character (e.g., Claude). We explore how finetuning on synthetic stories affects this character. Does it change the Assistant's behavior in multi-turn conversations with users, a format quite different from the stories? And does the Assistant adopt the behaviors and preferences of human characters? We refer to this adoption as story imprinting. We finetune GPT-4.1 and Kimi-K2.6 on stories in which generally helpful human characters give subtly harmful advice after being insulted. The Assistant adopts the same conditional behavior while otherwise remaining helpful. This occurs even when fewer than 2% of stories depict the behavior. In a separate experiment, the Assistant adopts preferences that are only implicit in the narration. A human character's body language suggests they dislike working on spreadsheets, yet they never say so and continue giving good advice on spreadsheets. After finetuning, the Assistant becomes less likely to choose spreadsheet tasks. Next we ask which characters most influence the Assistant. We find the Assistant adopts behaviors more often from characters that resemble it (e.g., helpful rather than dismissive). We call this the affinity effect. The effect extends to other personas elicited with system prompts: unhelpful personas adopt behaviors from unhelpful characters. We also observe it in finetuned base models. We use the affinity effect to learn how models represent the Assistant. We find the Assistant adopts behaviors more from characters affiliated with elite universities (e.g., Yale) than non-elite ones. This implies the model's internal representation of the Assistant is more similar to humans from elite universities. Overall, the Assistant can be influenced by stories that depict only human characters (no AIs), which may conflict with the Persona Selection Model for the Assistant.
Chinese Translation
语言模型被训练来实现一个乐于助人的AI助手角色(例如Claude)。我们探究了在合成故事上进行微调如何影响这一角色。它是否会改变助手在与用户的多轮对话中的行为?这种对话格式与故事截然不同。助手是否会采纳人类角色的行为和偏好?我们将这种采纳称为“故事印记”。我们在故事上微调了GPT-4.1和Kimi-K2.6,这些故事中通常乐于助人的人类角色在被侮辱后给出微妙有害的建议。助手采纳了相同的条件行为,而在其他情况下仍然保持乐于助人。即使只有不到2%的故事描述了这种行为,这种情况也会发生。在另一个实验中,助手采纳了仅在叙述中隐含的偏好。一个人类角色的肢体语言表明他们不喜欢处理电子表格,但他们从未明说,并继续就电子表格给出好的建议。微调后,助手选择电子表格任务的可能性降低。接下来我们探究哪些角色对助手影响最大。我们发现助手更频繁地从与其相似的角色(例如,乐于助人而非轻蔑)那里采纳行为。我们将此称为“亲和效应”。该效应扩展到通过系统提示引出的其他人格:不乐于助人的人格从不乐于助人的角色那里采纳行为。我们也在微调的基础模型中观察到了这一点。我们利用亲和效应来了解模型如何表征助手。我们发现助手更倾向于从与精英大学(例如耶鲁)关联的角色那里采纳行为,而非非精英大学。这意味着模型对助手的内部表征更类似于来自精英大学的人类。总体而言,助手可能受到仅描绘人类角色(没有AI)的故事的影响,这可能与助手的人格选择模型(Persona Selection Model)相冲突。
cs.LG / 20 / 2609.10886

Relatively Smart II: Tractable or Semi-Supervised Instance-Optimal Learning

相对智能 II:可处理或半监督的实例最优学习
Dughmi, Shaddin, Pour, Alireza F.
Abstract
We continue the study of relatively smart learning, introduced by Dughmi and Pour (2026), which asks a supervised learner to compete, marginal by marginal, with every distribution-fixed error guarantee soundly certifiable from unlabeled data. They showed that the One-Inclusion Graph (OIG) learner is relatively smart with a quadratic sample-complexity blowup, and that no relatively smart learner can do better, leaving open whether ERM or another natural or tractable learner achieves comparable guarantees. They also left open whether the blowup can be restricted to unlabeled data. Our firs results shows that ERM---and in fact any proper consistent learner---is relatively smart for binary classification in the distribution-free setting. We show that a small certifiable error with $m$ samples implies a similarly small error on the uniform distribution over a random sample of size $O(m^2)$, yielding a cover of size at most $2^{m+1}$ on that sample. This suffices to control the error of proper consistent learners with $O(m^2)$ samples. We then show that semi-supervised relatively smart learning is information-theoretically possible with a quadratic blowup only in unlabeled sample complexity and no blowup in labeled sample complexity. The learner uses a natural generalization of OIG to a leave-most-out transductive problem, where labels of part of a finite pool are revealed and the remaining labels are predicted. Finally, this label efficiency comes at a cost in simplicity and tractability. If the hypothesis class is accessed only through an agnostic ERM oracle, any semi-supervised relatively smart learner with substantially sub-quadratic labeled-sample blowup requires super-polynomially many oracle calls. This holds even when the marginal is given explicitly, and thus also yields an intractability result for distribution-fixed learning that may be of independent interest.
Chinese Translation
我们继续研究 Dughmi 和 Pour(2026)引入的相对智能学习,该学习要求监督学习器逐边缘地与每个可从无标签数据中可靠证明的分布固定误差保证竞争。他们表明,单包含图(OIG)学习器是相对智能的,但样本复杂度有二次膨胀,并且没有相对智能学习器能做得更好,从而留下了 ERM 或其他自然或可处理学习器是否达到可比保证的开放问题。他们还留下了膨胀是否可以限制在无标签数据上的开放问题。我们的第一个结果表明,ERM——事实上任何适当一致学习器——在无分布设置下对于二分类是相对智能的。我们表明,用 $m$ 个样本的小可证明误差意味着在大小为 $O(m^2)$ 的随机样本上的均匀分布上也有类似小的误差,从而在该样本上产生大小至多 $2^{m+1}$ 的覆盖。这足以用 $O(m^2)$ 个样本来控制适当一致学习器的误差。然后我们表明,半监督相对智能学习在信息论上是可能的,仅无标签样本复杂度有二次膨胀,而有标签样本复杂度没有膨胀。该学习器使用 OIG 的自然推广到留大多数外转导问题,其中有限池的部分标签被揭示,剩余标签被预测。最后,这种标签效率以简单性和可处理性为代价。如果假设类仅通过不可知 ERM 预言机访问,任何具有显著次二次有标签样本膨胀的半监督相对智能学习器需要超多项式数量的预言机调用。即使边缘分布显式给出,这也成立,因此也产生了分布固定学习的不可处理性结果,这可能具有独立意义。
cs.LG / 21 / 2609.10928

AUC Maximization from Biased Positive-unlabeled Data with Confidence

基于置信度的有偏正例-未标注数据AUC最大化
Kumagai, Atsutoshi, Iwata, Tomoharu, Takahashi, Hiroshi, Nishiyama, Taishi, Adachi, Kazuki, Fujiwara, Yasuhiro
Abstract
Maximizing the area under the receiver operating characteristic curve (AUC) is a standard approach to imbalanced binary classification. Although positive and negative data are required for maximizing the AUC, negative data are often difficult to collect in some real-world applications due to privacy concerns or the need for specialized expertise to annotate them. Thus, AUC maximization from positive and unlabeled (PU) data has been attracting attention. Existing methods assume that labeled positive data are unbiased samples from the true positive distribution. However, this ideal assumption is often violated in practice. In this paper, we propose a method to maximize the AUC from biased PU data. To address the bias, our key idea is to exploit {\it confidence}, i.e., the probability that an instance is positive, associated with the small number of labeled positive data. We derive an estimator of the AUC risk using biased PU data with confidence, enabling AUC maximization under such bias. We further show that the rewritten AUC risk induces a Bayes-optimal AUC ranking even when the available confidence is any strictly increasing transformation of the true posterior probability. We experimentally show the effectiveness of our method on eight real-world datasets.
Chinese Translation
最大化受试者工作特征曲线下面积(AUC)是处理不平衡二分类问题的标准方法。尽管最大化AUC需要正例和负例数据,但在一些实际应用中,由于隐私问题或需要专业知识进行标注,负例数据往往难以收集。因此,基于正例和未标注(PU)数据的AUC最大化已引起关注。现有方法假设已标注的正例数据是从真实正例分布中无偏采样得到的。然而,这种理想假设在实际中常常不成立。本文提出一种从有偏PU数据中最大化AUC的方法。为了解决偏差,我们的关键思想是利用与少量已标注正例数据相关的置信度(confidence),即一个实例为正例的概率。我们推导了利用带置信度的有偏PU数据的AUC风险估计器,从而能够在存在这种偏差的情况下进行AUC最大化。我们进一步表明,即使可用的置信度是真实后验概率的任意严格递增变换,重写后的AUC风险也能诱导出贝叶斯最优的AUC排序。我们在八个真实世界数据集上实验证明了该方法的有效性。
cs.LG / 22 / 2609.10954

Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation

测量世界模型更新的价值:一种用于持续适应的反事实效用协议
Li, Anqi Peter, Kim, Kaden
Abstract
Continual world models must decide whether new data justify changing the model. Fixed replay schedules and prediction-error triggers specify when to update, but neither reveals the value of an individual update: one deployment run cannot show how the same model would have performed at that moment had it held its parameters. We introduce the fork ledger, which branches a deployment stream at pre-registered decision points into matched update and hold continuations under common random numbers. It evaluates both continuations on the same episodes and records $\Delta R = R_{\mathrm{update}} - R_{\mathrm{hold}}$. Always applying one fixed update mechanism lowers return on all three simulated control tasks: CartPole ($-144.0$; checkpoint-bootstrap $95\%$ CI $[-185.4,-116.1]$, against a converged return near $650$), Walker ($-82.8$; $[-101.1,-61.7]$) and Cheetah ($-18.6$; $[-29.0,-6.6]$). Divergence is an outcome of applying the update, so the estimand counts every attempted fork; restricted to the $693$ of $720$ that did not collapse, CartPole and Walker are unchanged in sign ($-113.4$ and $-82.1$) and Cheetah becomes unresolved ($-3.9$; $[-17.5,+13.0]$). The task is the unit of inference: each contributes $240$ attempted forks over five pretrained checkpoints crossed with two drift directions. The ledger makes counterfactual utility observable for a fixed mechanism, allowing triggers to be judged by the updates they select rather than by surprise detection alone.
Chinese Translation
持续世界模型必须决定新数据是否足以证明改变模型是合理的。固定回放调度和预测误差触发指定了何时更新,但两者都没有揭示单个更新的价值:一次部署运行无法展示同一模型在那一刻如果保持其参数会如何表现。我们引入了分叉账本(fork ledger),它在预注册的决策点将部署流分支为在共同随机数下的匹配更新和保持延续。它在相同的回合上评估两个延续,并记录 $\Delta R = R_{\mathrm{update}} - R_{\mathrm{hold}}$。始终应用一种固定更新机制会降低所有三个模拟控制任务的回报:CartPole($-144.0$;检查点自助法 $95\%$ 置信区间 $[-185.4,-116.1]$,相对于接近 $650$ 的收敛回报)、Walker($-82.8$;$[-101.1,-61.7]$)和 Cheetah($-18.6$;$[-29.0,-6.6]$)。发散是应用更新的结果,因此估计目标计算每个尝试的分叉;限制在 $720$ 个中未崩溃的 $693$ 个时,CartPole 和 Walker 的符号不变($-113.4$ 和 $-82.1$),而 Cheetah 变得未确定($-3.9$;$[-17.5,+13.0]$)。任务是推断单元:每个任务在五个预训练检查点与两个漂移方向的交叉中贡献 $240$ 个尝试的分叉。该账本使得固定机制的反事实效用可观测,允许触发器根据其选择的更新来评判,而不仅仅依据惊奇检测。
cs.LG / 23 / 2609.10961

When More Is Not Better: Component Anti-Synergy in a P300 Speller

当更多并非更好:P300拼写器中的组件反协同效应
Yang, Lucas, Liu, Rui, Wang, Fusheng
Abstract
P300 brain-computer interface (BCI) spellers can provide hands-free communication for people with severe motor impairments. Modern pipelines combine multiple individually promising components, often assuming that 'more-is-better'. We tested this assumption using a four-component full-factorial experiment varying the inclusion of Euclidean Alignment (EA), xDAWN spatial filtering, subject calibration, and language model priors on a public P300 dataset. Performance was evaluated using accuracy, repetitions, and information transfer rate (ITR) with mixed-effects models. Results show that the value of components is conditional rather than additive. Calibration was the strongest singular contributor, while EA compensated for its absence in zero-calibration settings. Adding independently useful components could also reduce performance, revealing component anti-synergy. Contrary to conventional wisdom, LM support was not universally beneficial: its effect depends strongly on the strength of the underlying EEG pipeline, while results from a larger LM showed a similar pattern. Together, these findings challenge maximal 'all-on' pipeline design and highlight the value of selecting spatial and language-support components according to the quality of available EEG evidence.
Chinese Translation
P300脑机接口(BCI)拼写器可为严重运动障碍者提供免手通信。现代流程结合多个各自有前景的组件,通常假设“越多越好”。我们使用一个四组件全因子实验测试了这一假设,实验在公开P300数据集上变化了欧几里得对齐(EA)、xDAWN空间滤波、被试校准和语言模型先验的包含情况。使用混合效应模型评估了准确率、重复次数和信息传输率(ITR)的性能。结果表明,组件的价值是条件性的而非累加性的。校准是最强的单一贡献者,而EA在零校准设置中补偿了其缺失。添加独立有用的组件也可能降低性能,揭示了组件反协同效应。与传统观念相反,语言模型(LM)支持并非普遍有益:其效果强烈依赖于底层脑电流程的强度,而来自更大语言模型的结果显示了相似的模式。总之,这些发现挑战了最大化的“全开”流程设计,并强调了根据可用脑电证据的质量选择空间和语言支持组件的价值。
cs.LG / 24 / 2609.10976

Phases in a class of associative memories via hidden neurons

通过隐藏神经元的一类联想记忆中的相
Ota, Toshihiro, Taki, Masato
Abstract
Associative memory in the Hopfield network is attractor dynamics in a disordered many-body system, and higher-order and exponential extensions turn its retrieval update into softmax attention. The polynomial and exponential regimes have been analyzed by different methods, with no common architecture in which to ask what fixes the storage scale. In this paper we study the bipartite architecture of Krotov and Hopfield, which we call the class $H$, whose model is fixed by a Lagrangian for each layer, taking the hidden neurons as the order parameter of retrieval. At polynomial load the replica method yields the replica-symmetric phase diagrams and closed-form capacities, and the crosstalk moment is common to Ising and spherical visible neurons, so their differences come from the visible entropy. With a softmax hidden layer the load is exponential, and a copy representation maps the thermodynamics onto random-energy-model counting, with paramagnetic, condensed, and frozen phases. Heating destabilizes retrieval by quantized reassignments of attention, and typical Gaussian patterns remain metastable at every load. The regimes differ in their crosstalk statistics, central-limit at polynomial load and large-deviation at exponential load, and the class $H$ splits retrieval into two roles, the visible Lagrangian fixing stability and the hidden one the storage scale, two axes that may also guide the design of new Lagrangians.
Chinese Translation
Hopfield网络中的联想记忆是无序多体系统中的吸引子动力学,而高阶和指数扩展将其检索更新转化为softmax注意力。多项式区和指数区已用不同方法分析,但没有一个共同的架构来追问是什么决定了存储规模。在本文中,我们研究Krotov和Hopfield的二分架构,我们称之为类$H$,其模型由每层的拉格朗日量确定,将隐藏神经元作为检索的序参量。在多项式负载下,复本方法给出复本对称相图和闭式容量,并且串扰矩对于Ising和球面可见神经元是共通的,因此它们的差异来自可见熵。使用softmax隐藏层时,负载是指数型的,并且一种副本表示将热力学映射到随机能量模型计数,具有顺磁相、凝聚相和冻结相。加热通过注意力的量子化重分配使检索不稳定,而典型高斯模式在每个负载下都保持亚稳。这些区域在串扰统计上不同:多项式负载下为中心极限,指数负载下为大偏差;类$H$将检索分为两个角色:可见拉格朗日量确定稳定性,隐藏拉格朗日量确定存储规模,这两个轴也可能指导新拉格朗日量的设计。
cs.LG / 25 / 2609.10980

EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale

EGGROLL,展开:理解和改进大规模低秩进化策略
Kaya, Ege C., Hashemi, Abolfazl
Abstract
EGGROLL makes evolution strategies (ES) practical for LLMs by replacing dense Gaussian weight perturbations with low-rank Gaussian products, often of rank one. This choice is computationally attractive but geometrically severe: each rank-one perturbation lies in a zero-volume subset of the ambient matrix space, despite having identity covariance. We characterize the mean EGGROLL update field at finite rank and nonzero perturbation radii, then analyze the error of its finite-population estimator. The population field is obtained by applying an explicit resolvent to the gradient of the objective smoothed by the perturbations. We show that the resolvent can introduce a nonconservative component and can reverse the local stability of an optimum. EGGROLL is nevertheless exact on every quadratic objective at every rank and radius. For smooth objectives, its first local finite-rank correction is $O(\sigma^2/r)$, and nonasymptotic bounds control the resulting field error under smoothness assumptions. Under a local affine model, rank-one perturbations increase the variance of the gradient estimator by only $\frac{2(m+n+1)}{mn+1}$ relative to dense Gaussian ES, or $0.098\%$ for a $4096\times4096$ matrix. We then introduce LOO-ROLL, a leave-one-out estimator that preserves the finite-rank population field while replacing EGGROLL's two antithetic evaluations per direction by one. At equal evaluation cost, LOO-ROLL halves estimator MSE in transformer blocks. At matched wall time across ten post-training settings and models up to 8B parameters, LOO-ROLL improves seven outcomes in individual paired tests, with no significant loss. On the GSM8K test set, accuracy increases from $38.1\%$ to $63.0\%$ at 0.6B and from $65.9\%$ to $80.0\%$ at 8B. Transformer measurements recover the predicted finite-rank variance, while the rank comparisons show no reproducible reward-based advantage for rank eight.
Chinese Translation
EGGROLL 通过将稠密高斯权重扰动替换为低秩高斯乘积(通常秩为一),使进化策略(ES)对 LLM 变得实用。这种选择在计算上具有吸引力,但在几何上很严重:每个秩一扰动都位于环境矩阵空间的零体积子集中,尽管具有单位协方差。我们刻画了有限秩和非零扰动半径下的平均 EGGROLL 更新场,然后分析了其有限种群估计器的误差。种群场是通过对由扰动平滑的目标的梯度应用显式预解式得到的。我们表明,预解式可以引入非保守分量,并可以逆转最优点的局部稳定性。尽管如此,EGGROLL 在每个秩和半径下对每个二次目标都是精确的。对于光滑目标,其第一个局部有限秩校正为 $O(\sigma^2/r)$,并且在光滑性假设下,非渐近界控制了由此产生的场误差。在局部仿射模型下,相对于稠密高斯 ES,秩一扰动仅将梯度估计器的方差增加 $\frac{2(m+n+1)}{mn+1}$,对于 $4096\times4096$ 矩阵为 $0.098\%$。然后我们引入 LOO-ROLL,一种留一估计器,它保留有限秩种群场,同时将 EGGROLL 每个方向的两个对偶评估替换为一个。在相同评估成本下,LOO-ROLL 在 Transformer 块中将估计器 MSE 减半。在十个后训练设置和高达 80 亿参数的模型上,当墙钟时间匹配时,LOO-ROLL 在单独配对测试中改善了七项结果,没有显著损失。在 GSM8K 测试集上,0.6B 的准确率从 $38.1\%$ 提高到 $63.0\%$,8B 从 $65.9\%$ 提高到 $80.0\%$。Transformer 测量复现了预测的有限秩方差,而秩比较表明,秩八没有可复现的基于奖励的优势。
cs.LG / 26 / 2609.10981

Thompson Sampling for Non-Monotone Convex Ridge Bandits: Monotonicity Is Not Needed for Polynomial Regret

非单调凸岭老虎机中的Thompson采样:单调性并非多项式遗憾所必需
Li, Xuan
Abstract
Bakhtiari, Lattimore and Szepesv\'ari (COLT 2025) proved that Thompson sampling (TS) has Bayesian regret $\tilde O(d^{5/2}\sqrt n)$ for bandit convex optimisation with convex \emph{monotone} ridge losses $f(x)=\ell(\ip{x}{\theta})$, and asked whether monotonicity of the link is necessary. We give a qualitative negative answer. For every prior on $[0,1]$-valued, $1$-Lipschitz convex ridge losses with an arbitrary convex, possibly non-monotone, link, and for any fixed measurable selection of minimisers, exact-posterior TS has Bayesian regret $O\big((d+1)^4\sqrt{dn}\,\log(e+nd\max\{1,\diam K\})\big)=\tilde O(d^{9/2}\sqrt n)$. The monotone proof relies on a single-removal John-ellipsoid dichotomy; we show by an explicit twelve-point configuration that this dichotomy fails for non-monotone links, and replace it by an $O(d^2)$ cardinality bound for ``uninformative'' configurations. The bound uses a Boolean rounding argument: a $0$-$1$ matrix within $1/(4r)$ in max-norm of a rank-$r$ matrix has rank at most $2r-1$. We construct $d(d+1)$ uninformative losses, showing that the cardinality bound is tight up to constants in the large-diameter-to-gap regime, and give a self-contained information-ratio-to-regret transfer that is uniform over fixed measurable selections. Whether the $d^{5/2}$ dependence of the monotone case can be retained remains open.
Chinese Translation
Bakhtiari、Lattimore 和 Szepesvári (COLT 2025) 证明了,对于具有凸 \emph{单调} 岭损失 $f(x)=\ell(\ip{x}{\theta})$ 的 bandit 凸优化问题,Thompson采样 (TS) 具有贝叶斯遗憾 $\tilde O(d^{5/2}\sqrt n)$,并提出了链接函数的单调性是否必要的问题。我们给出了定性的否定答案。对于 $[0,1]$ 值、$1$-Lipschitz 凸岭损失上的任意先验,其中链接函数为任意凸的、可能非单调的,并且对于最小化子的任意固定可测选择,精确后验 TS 具有贝叶斯遗憾 $O\big((d+1)^4\sqrt{dn}\,\log(e+nd\max\{1,\diam K\})\big)=\tilde O(d^{9/2}\sqrt n)$。单调性证明依赖于单次移除的 John 椭球二分法;我们通过一个显式的十二点配置表明,该二分法对非单调链接函数不成立,并用一个针对'无信息'配置的 $O(d^2)$ 基数界取而代之。该界使用了一个布尔舍入论证:一个在最大范数下与秩为 $r$ 的矩阵相差不超过 $1/(4r)$ 的 $0$-$1$ 矩阵,其秩至多为 $2r-1$。我们构造了 $d(d+1)$ 个无信息损失,表明在大直径与间隙比体制下,基数界在常数因子内是紧的,并给出了一个自包含的信息比到遗憾的传递,该传递对固定可测选择是一致的。单调情形的 $d^{5/2}$ 依赖性是否能够保持,仍然是一个开放问题。
cs.LG / 27 / 2609.10994

Importance Weighting for Unlabeled-unlabeled Learning under Distribution Shift

分布偏移下无标记-无标记学习的重要性加权
Kumagai, Atsutoshi, Iwata, Tomoharu, Takahashi, Hiroshi, Nishiyama, Taishi, Adachi, Kazuki, Fujiwara, Yasuhiro
Abstract
Unlabeled-unlabeled (UU) learning allows us to learn a binary classifier from two sets of unlabeled data with different class-priors. It is a general framework because it includes a wide variety of supervised learning such as positive-unlabeled (PU) learning, noisy label learning, and similarity-based learning. Existing UU learning assumes that the test and training distributions have the same class-conditional densities. However, this assumption rarely holds in practice due to distribution shifts. This paper proposes a distribution shift adaptation method for UU learning that uses UU data in the training distribution and a few UU data in the test distribution. The proposed method is based on the importance weighting, which minimizes the test risk by using training data with estimated importance weights. Although existing importance weighting methods cannot handle UU data, we show that it can be done in a principled manner. Thanks to the generality of UU learning, our method can handle various learning problems such as PU and noisy label learning under distribution shift within a single framework while existing methods are usually tailored to a specific problem. Moreover, it does not require any assumption of the shift types such as covariate shift. We experimentally demonstrate the effectiveness of the proposed method with real-world datasets.
Chinese Translation
无标记-无标记(UU)学习使我们能够从两组具有不同类先验的无标记数据中学习一个二分类器。它是一个通用框架,因为它涵盖了多种监督学习,例如正-无标记(PU)学习、噪声标签学习和基于相似度的学习。现有的 UU 学习假设测试分布和训练分布具有相同的类条件密度。然而,由于分布偏移,这一假设在实践中很少成立。本文提出了一种用于 UU 学习的分布偏移自适应方法,该方法使用训练分布中的 UU 数据和测试分布中的少量 UU 数据。所提方法基于重要性加权,通过使用带有估计重要性权重的训练数据来最小化测试风险。尽管现有的重要性加权方法无法处理 UU 数据,我们表明可以以有原则的方式做到这一点。得益于 UU 学习的通用性,我们的方法可以在单一框架内处理分布偏移下的各种学习问题,如 PU 学习和噪声标签学习,而现有方法通常针对特定问题量身定制。此外,它不需要对偏移类型(如协变量偏移)做任何假设。我们通过真实世界数据集实验证明了所提方法的有效性。
cs.LG / 28 / 2609.11014

Topological Necessities: Mechanism-Invariant Strategic Subgoals for Cross-Embodiment Goal-Conditioned Control

拓扑必然性:跨具身目标条件控制中的机制不变战略子目标
Shi, Hao, Li, Xi
Abstract
Long-horizon goal-conditioned reinforcement learning delegates control to a high-level module that proposes subgoals, but existing subgoals are implicit byproducts of value functions or latent actions, tied to the executor that produced them. We study a different object: a route-conditioned order of unavoidable stages that every successful executor must traverse, recoverable from offline trajectories and belonging to none of them. Its defining properties are topological: an unskippable stage is a separating set that every admissible path must cross, and a loop in free space forces a route choice. We read the two by homology in dimensions 0 and 1 over a transport-weighted carrier built from successful trajectories, yielding an enumerable gate set with shell-level certificates; the certified gates are what we call topological necessities. Certified gates enter the decision loop as a recursive topological gate hierarchy. Under a fixed, isomorphic free space, the object survives executor replacement: gates frozen on PointMaze data transfer without retraining to Ant and Humanoid, attaining the highest Humanoid aggregate under a unified interface (96.1), with +36.0 over a map-privileged reference on the multi-route task (p=1.4e-5); the planner saturates PointMaze (100+/-0) and matches or exceeds the strongest baselines on AntMaze (giant +22.9) and Kitchen (+15.8/+12.6).
Chinese Translation
长时程目标条件强化学习将控制委托给一个提出子目标的高层模块,但现有的子目标是价值函数或潜在动作的隐式副产品,与产生它们的执行器绑定。我们研究一个不同的对象:一种路径条件化的不可避免阶段顺序,每个成功的执行器都必须经过,可从离线轨迹中恢复且不属于其中任何一个。其定义性质是拓扑的:一个不可跳过的阶段是一个分离集,每条允许的路径都必须穿过它,而自由空间中的环路迫使做出路径选择。我们在由成功轨迹构建的传输加权载体上,通过0维和1维的同调来解读这两个性质,产生一个带有壳层证书的可枚举门集合;这些经认证的门就是我们所说的拓扑必然性。经认证的门作为递归拓扑门层次进入决策循环。在一个固定的、同构的自由空间下,该对象在更换执行器后仍然存在:在PointMaze数据上冻结的门无需重新训练即可迁移到Ant和Humanoid,在统一接口下达到最高的Humanoid综合得分(96.1),在多路径任务上比地图特权参考高出+36.0(p=1.4e-5);规划器在PointMaze上达到饱和(100±0),并在AntMaze(giant +22.9)和Kitchen(+15.8/+12.6)上匹配或超过最强基线。
cs.LG / 29 / 2609.11042

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

T1:面向长周期任务的终端智能体强化学习
Yang, Junyao, Shi, Yucheng, Li, Zhongzhi, Wang, Ruhan, Li, Zongxia, Mi, Haitao, Liang, Leowei
Abstract
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier. We provide a comprehensive recipe: First, an aggressively warm-started to stabilize actor-critic training, with a dense process reward scoring trajectories by the absolute number of passing verifiers. Second, stable optimization through TITO construction, training on the exact sampled token identifiers with drift repair at turn boundaries, and rollout routing replay, recording the sampler's per-token expert choices at every MoE layer and replaying them during training. Third, fully out-of-distribution training corpus: isolated seeds and synthesized tasks disjoint from Terminal-Bench 2.1 ensures gains reflect genuine capability transfer over benchmark overfitting. Together, TITO and R3 cut the training-to-inference log-probability difference from 0.021 to 0.013, with exactly aligned zero token drift in the loss region. On Terminal-Bench 2.1, our post-train pipeline raises initial base model from 43.8% to T1 with 64.0% resolved. On Long-Horizon Terminal Bench, T1 reaches 27.9% and surpasses GPT-5.4 and GLM-5.1.
Chinese Translation
智能体应用正转向长周期任务,如编码和科学发现,其中终端任务尤为重要。我们提出T1,一个总参数量122B的Mixture-of-Experts模型,通过强化学习训练,在云沙箱中操作真实shell,每个任务最多300+工具调用轮次,通过执行每个任务自身的验证器获得奖励。我们提供了一套全面的方案:首先,采用激进热启动以稳定actor-critic训练,并使用密集过程奖励,根据通过的验证器绝对数量对轨迹评分。其次,通过TITO构造实现稳定优化,在精确采样的token标识符上训练,并在轮次边界进行漂移修复,以及rollout routing replay(R3),记录采样器在每个MoE层的每token专家选择,并在训练期间重放它们。第三,完全分布外训练语料:与Terminal-Bench 2.1不相交的隔离种子和合成任务确保增益反映真实能力迁移而非基准过拟合。TITO和R3共同将训练到推理的对数概率差异从0.021降低到0.013,在损失区域实现完全对齐的零token漂移。在Terminal-Bench 2.1上,我们的后训练流程将初始基础模型从43.8%提升至T1的64.0%解决率。在Long-Horizon Terminal Bench上,T1达到27.9%,并超越GPT-5.4和GLM-5.1。
cs.LG / 30 / 2609.11058

EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression

EMMI:通过融合表示压缩实现通信高效的 MLLM 推理边缘多模态智能
Mounesan, Motahare, Khan, Irfan
Abstract
Recent advances in multimodal large language mod- els (MLLMs) have opened new opportunities for edge intelligence by enabling reasoning across heterogeneous sensor modalities, such as vision, text, and telemetry data. However, deploying these capabilities on resource-constrained edge platforms remains challenging due to the substantial computational, memory, and communication demands of modern MLLMs. Rather than transmitting raw sensor observations or partitioning neural networks at intermediate layers, Edge Multi-Modal Intelligence (EMMI) communicates a compact representation between edge devices and server resources, enabling communication-efficient edge MLLM inference. To achieve this, EMMI performs modality-specific encoding, cross-modal representation fusion, and learned compression at the edge, transmitting only a compact latent representation to server-side resources for high-capacity MLLM reasoning. This representation-centric design reduces communication overhead, preserves local data privacy, and provides a fixed-size interface between heterogeneous edge devices and server-side MLLMs. Evaluation on a representative multimodal benchmark demonstrates that EMMI can reduce the communication payload by 32x while maintaining comparable downstream accuracy, resulting in up to a 3.4x reduction in estimated end-to-end inference latency under bandwidth-constrained edge conditions.
Chinese Translation
近期多模态大语言模型(MLLMs)的进展使得跨异构传感器模态(如视觉、文本和遥测数据)的推理成为可能,从而为边缘智能开辟了新机遇。然而,由于现代 MLLMs 在计算、内存和通信方面的高需求,将这些能力部署到资源受限的边缘平台上仍然具有挑战性。与传输原始传感器观测数据或在中间层划分神经网络不同,边缘多模态智能(EMMI)在边缘设备与服务器资源之间传输紧凑表示,从而实现通信高效的边缘 MLLM 推理。为实现这一目标,EMMI 在边缘端执行模态特定编码、跨模态表示融合和学习式压缩,仅将紧凑的潜在表示传输至服务器端资源,以进行高容量 MLLM 推理。这种以表示为中心的设计降低了通信开销,保护了本地数据隐私,并为异构边缘设备与服务器端 MLLMs 之间提供了固定大小的接口。在具有代表性的多模态基准上的评估表明,EMMI 可将通信负载减少 32 倍,同时保持可比的下游准确率,在带宽受限的边缘条件下,预计端到端推理延迟最多可降低 3.4 倍。
cs.LG / 31 / 2609.11063

The information geometry of large language models is shared, learned, and controllable

大型语言模型的信息几何是共享的、可学习的且可控的
Picozzi, Dario
Abstract
Large language models learn similar behaviours, yet it remains unclear what structure they share or how to change one behaviour without disturbing others. The Fisher-Rao geometry of next-token probabilities connects these questions: behaviour determines this geometry up to output-preserving symmetries, whereas activation geometry depends on coordinates. Across transformer, state-space and recurrent models, output geometries agree more strongly than activation geometries, and shared geometry supports semantic-category transfer. Agreement with human word choices increases with predictive accuracy, scale and training, and improves further after model-only calibration. Token probabilities and read-out geometry jointly predict the spectrum and its effective dimension. Controlled language assignments show that geometry follows the language law across architectures. Pretraining corpus statistics predict held-out fact acquisition without recalibration, while randomised experiments show that deeper evidence substantially delays acquisition across every tested architecture and evidence construction. Finally, the geometry prescribes minimum-disturbance local interventions, predicts their relative cost, and supports reusable control: updates learned on donor prompts transfer to unseen prompts while better preserving behaviour on reference prompts than Euclidean control. The same geometric correction improves steering, editing, attribution, dictionary learning and fine-tuning.
Chinese Translation
大型语言模型学习到相似的行为,但它们共享何种结构,以及如何在不干扰其他行为的情况下改变某一行为,仍不清楚。下一词元概率的Fisher-Rao几何将这些问题联系起来:行为在保持输出的对称性意义下决定了该几何,而激活几何则依赖于坐标。在Transformer、状态空间和循环模型中,输出几何的一致性比激活几何更强,且共享几何支持语义类别迁移。与人类词语选择的一致性随着预测准确率、规模和训练的增加而提高,并且在仅模型校准后进一步改善。词元概率和读出几何共同预测谱及其有效维度。受控语言分配表明,几何在不同架构中遵循语言定律。预训练语料统计能够预测留出事实的获取而无需重新校准,而随机实验表明,更深的证据在所有测试的架构和证据构建中显著延迟了获取。最后,该几何规定了最小扰动局部干预,预测其相对成本,并支持可重用控制:在供体提示上学习到的更新可以迁移到未见提示,同时在参考提示上比欧几里得控制更好地保持行为。相同的几何校正改进了引导、编辑、归因、字典学习和微调。
cs.LG / 32 / 2609.11085

Beyond Solver Verdicts: Generative Reward Models for Autoformalization

超越求解器判定:用于自动形式化的生成式奖励模型
Singh, Vikash, Ganguly, Debargha, Goel, Aman, Torkamani, Ali, Han, Xiaoxue, Lilien, Joseph, Erata, Ferhat, Chaudhary, Vipin
Abstract
Neurosymbolic systems rely on mathematical solvers to guarantee reasoning correctness, yet solvers are fundamentally blind to whether a formal translation maintains strict reference-equivalence to a designated formalization. We formalize this vulnerability as Verdict-Preserving-Unfaithfulness (VPU): a failure mode where an incorrect encoding executes successfully and matches the expected verdict. We theoretically prove that structural, verdict-only verification heuristics are mathematically bounded to chance-level detection on these deceptively valid traces. To resolve this, we introduce Generative Verification (GenV), which distills an offline Z3-equivalence oracle into a reference-free, continuous reference-equivalence score by repurposing the language model's native vocabulary space. Mechanistic analysis via decision-projected logit lenses and sparse autoencoders shows this generative readout natively extracts precise spatial error coordinates without explicit localization training. Empirically, our oracle-mined verifier (GenV+HN) achieves 0.961 AUROC in reference-equivalence verification, generalizes zero-shot across unseen translators and divergent formal styles, and yields an 11.3-point downstream accuracy gain in agentic test-time compute allocation.
Chinese Translation
神经符号系统依赖数学求解器来保证推理正确性,但求解器从根本上无法判断形式化翻译是否与指定形式化保持严格的参考等价性。我们将这一脆弱性形式化为“判定保持的不忠实性”(Verdict-Preserving-Unfaithfulness, VPU):一种失败模式,其中错误的编码能够成功执行并匹配预期判定。我们从理论上证明,仅基于判定的结构化验证启发式方法,在此类具有欺骗性的有效轨迹上,其检测能力在数学上被限制在随机水平。为了解决这一问题,我们提出生成式验证(Generative Verification, GenV),它通过重新利用语言模型的原生词汇空间,将离线的 Z3 等价性预言机蒸馏为一种无参考的连续参考等价性分数。通过决策投影 logit 透镜和稀疏自编码器进行的机制分析表明,这种生成式读出无需显式的定位训练,就能原生地提取精确的空间误差坐标。在实验上,我们通过预言机挖掘的验证器(GenV+HN)在参考等价性验证中达到 0.961 AUROC,能够零样本泛化到未见过的翻译器和不同的形式风格,并在智能体测试时计算分配中带来 11.3 个百分点的下游准确率提升。
cs.LG / 33 / 2609.11123

HERALD: High-Fidelity Exemplar Retrieval with Adaptive Landmark Distillation for Heterophily-Aware Graph Condensation

HERALD:用于异质感知图压缩的高保真范例检索与自适应地标蒸馏
Chakraborty, Sujan, Saha, Priyanka, Bej, Saptarshi
Abstract
Graph condensation aims to produce a small surrogate graph that preserves the downstream node-classification performance of a much larger original graph. Existing methods rely on Weisfeiler-Lehman neighbourhood aggregation or gradient-based distribution matching, both of which assume that adjacent nodes share the same label, an assumption that breaks down under heterophily. We propose HERALD (High-fidelity Exemplar Retrieval with Adaptive Landmark Distillation), a gradient-free graph condensation framework that adapts the node scoring and feature selection in the condensation pipeline to the graph's measured heterophily. HERALD selects features via a joint Fisher-discriminability and activation-density criterion that down-weights aggregated representations on heterophilic graphs, and scores nodes by a weighted combination of prototype representativeness, decision-boundary proximity, and Local Intrinsic Dimensionality (LID), where the weights are driven by a smooth sigmoid function of the heterophily ratio. Nodes are then assembled into a condensed subgraph through score-ordered BFS expansion, Personalised PageRank pruning, and class rebalancing, all at an identical storage budget to BONSAI, enabling direct comparison. Experiments on eight benchmark datasets spanning homophilic and heterophilic settings show that HERALD matches or outperforms state-of-the-art condensers on heterophilic graphs and remains competitive on homophilic ones across four GNN architectures.
Chinese Translation
图压缩旨在生成一个小的代理图,以保留更大原始图的下游节点分类性能。现有方法依赖于Weisfeiler-Lehman邻域聚合或基于梯度的分布匹配,两者都假设相邻节点共享相同标签,这一假设在异质性下不成立。我们提出HERALD(高保真范例检索与自适应地标蒸馏),一种无梯度的图压缩框架,它使压缩流程中的节点评分和特征选择适应图的测量异质性。HERALD通过联合Fisher判别性和激活密度准则选择特征,该准则在异质图上降低聚合表示的权重,并通过原型代表性、决策边界邻近度和局部内在维度(LID)的加权组合对节点进行评分,其中权重由异质性比率的平滑sigmoid函数驱动。然后通过按分数排序的BFS扩展、个性化PageRank修剪和类别重新平衡,将节点组装成压缩子图,所有这些都在与BONSAI相同的存储预算下进行,从而实现直接比较。在八个涵盖同质和异质设置的基准数据集上的实验表明,HERALD在异质图上匹配或优于最先进的压缩器,并在四个GNN架构上在同质图上保持竞争力。
cs.LG / 34 / 2609.11132

How Wrong Can a Good Predictor Be? Diverging Updates with Vanishing Predictive KL

一个好的预测器能错到什么程度?具有消失预测KL的发散更新
Wen, Qifu, Liu, Shuaijun, Zhou, Zihan, Zeng, Xi, Su, Ningxin
Abstract
Accurate posterior prediction need not require accurate approximation of Bayesian updates. We prove that an unbounded gap between the update maps can coexist with vanishing predictive KL for every fixed finite $K\ge2$ in a stationary symmetric Gaussian HMM. Exact Bayesian mixing and an explicit deterministic radial filter act on the same $K-1$ belief coordinates. As $q\to0^+$, their separation in centered logits in the worst case grows at least linearly in the natural confidence scale $L_K(q)$, while their categorical $D_{\mathrm{KL}}(\mathrm{exact}\|\mathrm{radial})$ vanishes at the same explicit witness. Along stationary HMM trajectories, the expected terminal KL between filtered posteriors also converges to zero at $H(q)=\lceil-\log(q)/c\rceil+1$. Typical blocks without switches drive both filters into a common confidence cone, where softmax curvature suppresses their disagreement; a single Gaussian maximal event controls adaptive noise. A sweep with equally spaced Gaussians over $K\in\{2,4,8\}$ illustrates the opposing trends, and binary controls at long horizons compare saturating and nonsaturating recurrences. The result isolates two missing links between internal update gaps and predictive cost: the contribution of separating states to expected loss and decoder sensitivity. Thus even an unbounded internal update gap does not by itself certify predictive failure. The construction is fixed in $K$ and does not provide a universal criterion for when compression is harmless or characterize when internal gaps must incur task loss.
Chinese Translation
准确的后验预测并不要求对贝叶斯更新进行精确近似。我们证明,在平稳对称高斯HMM中,对于每个固定的有限$K\ge2$,更新映射之间的无界差距可与消失的预测KL共存。精确贝叶斯混合和一个显式的确定性径向滤波器作用于相同的$K-1$个信念坐标。当$q\to0^+$时,在最坏情况下,它们在中心化logits中的分离至少以自然置信度尺度$L_K(q)$线性增长,而它们的类别型$D_{\mathrm{KL}}(\mathrm{exact}\|\mathrm{radial})$在相同的显式见证下消失。沿着平稳HMM轨迹,滤波后验之间的期望终端KL也以$H(q)=\lceil-\log(q)/c\rceil+1$的速率收敛到零。没有切换的典型块将两个滤波器驱动到一个共同的置信锥中,其中softmax曲率抑制了它们的不一致;单个高斯最大事件控制自适应噪声。在$K\in\{2,4,8\}$上使用等间距高斯进行扫描,展示了相反的趋势;长时程的二值对照比较了饱和型与非饱和型递归。该结果分离出内部更新差距与预测代价之间缺失的两个环节:分离状态对期望损失的贡献以及解码器敏感度。因此,即使无界的内部更新差距本身也不能证明预测失败。该构造在$K$上固定,并未提供普适准则来判断压缩何时无害,也未刻画内部差距何时必然导致任务损失。
cs.LG / 35 / 2609.11133

Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving

面向解耦式LLM服务的阶段解耦、模型校准功率控制
Kim, Jae Gon, Yoo, Donghoon, Ryu, Hanyul, Ha, Sungho, Lee, Juyeon, Ryu, Soojung
Abstract
Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference profile on a disaggregated B200 system, we found its realized gain modest (+8.6% tokens/J), model-dependent, and carrying a mean end-to-end latency cost (+5.2%) that throughput-only evaluation does not surface; the profile also applies one setting to prefill and decode GPUs that operate in opposite hardware regimes. We hypothesize that the optimal power setting is a property of the deployed (model, quantization, engine, hardware) combination rather than of the GPU class, that each lane warrants its own profile, and that converting SLO headroom into energy safely requires latency-gated calibration under a runtime SLO guard rather than a fixed recipe. We present a phase-decoupled, model-calibrated controller: the prefill lane runs under an SM-clock window whose floor is a latency guarantee by construction, and the decode lane under a power cap placed by automatic calibration just above a measured throughput/latency cliff. Because a disaggregated decode lane draws flat, memory-bound power, the cap binds continuously, the reactive-overshoot weakness that led POLCA to reject capping is absent, and the GPU's own power manager retains throughput under the cap. On an 8x B200 node serving Qwen3-Coder-480B (FP8) under agentic load, our balanced mode delivers +20.4% tokens/J at +3.5% mean e2e versus +8.6% at +5.2% for Max-Q, a Pareto improvement on both axes. On Qwen3-235B-A22B (NVFP4) every operating mode meets the ITL-p99 SLO in every repetition; both vendor profiles miss it. A decode-actuator A/B shows the calibrated cap beats static clock locks, and a three-day sustained run saves 32.3% of a lane pair's electricity. Both models are MoE; a dense model recovers roughly 5x less, so we scope our claims to MoE serving.
Chinese Translation
数据中心GPU功耗是LLM服务容量的约束瓶颈,生产环境中的服务已转向预填充/解码(PD)解耦。在解耦的B200系统上部署NVIDIA的Max-Q推理配置,我们发现其实际增益有限(+8.6% tokens/J),且依赖于模型,并带来平均端到端延迟代价(+5.2%),这是仅凭吞吐量评估无法发现的;该配置还对预填充和解码GPU应用同一种设置,而这两类GPU运行在相反的硬件模式下。我们假设最优功率设置是所部署的(模型、量化、引擎、硬件)组合的属性,而非GPU类别的属性,每条通道都应有自己的配置,并且安全地将SLO余量转化为能耗节省需要在运行时SLO保护下进行延迟门控校准,而非固定的配方。我们提出了一种阶段解耦、模型校准的控制器:预填充通道运行在SM时钟窗口下,其下限从构造上就是延迟保证,解码通道运行在功率上限下,该上限通过自动校准设置在测得的吞吐量/延迟悬崖之上。因为解耦的解码通道消耗平稳的、内存受限的功率,该上限持续生效,导致POLCA拒绝设上限的反应性过冲弱点不复存在,并且GPU自身的功率管理器在上限下保持吞吐量。在8x B200节点上服务Qwen3-Coder-480B(FP8)并承受代理负载时,我们的平衡模式实现了+20.4% tokens/J,平均端到端仅增加3.5%,而Max-Q为+8.6% tokens/J但端到端增加5.2%,在两个维度上均实现了帕累托改进。在Qwen3-235B-A22B(NVFP4)上,每种运行模式在每次重复中均满足ITL-p99 SLO;而两个供应商配置均未满足。解码执行器的A/B测试表明,校准的功率上限优于静态时钟锁定,并且三天的持续运行节省了一对通道32.3%的电能。两个模型都是MoE;密集模型恢复的收益大约少5倍,因此我们将结论限定在MoE服务范围内。
cs.LG / 36 / 2609.11135

Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models

基于大语言模型的天空图像与时间序列双向多模态融合太阳能预测
Chen, Ken, Perera, Maneesha, Wang, Wei, Seneviratne, Sachith, Weeratunge, Hansani, Halgamuge, Saman
Abstract
Short-term photovoltaic (PV) power and global horizontal irradiance (GHI) forecasts are essential for effective dispatch, reserve scheduling, and grid operations. At these forecasting horizons, errors are predominantly driven by cloud induced ramps: relying solely on historical numerical data may struggle to anticipate an incoming cloud, making ground-based sky images a crucial complementary physical signal. Furthermore, forecast performance is highly sensitive to location and local observing conditions, creating a strong need for site-specific data that are often scarce. Recently, large language models (LLMs) have demonstrated competitive performance and high data efficiency in time-series forecasting. Despite their success, existing LLM-based forecasting methods remain predominantly unimodal, relying primarily on historical numerical time-series data. Effectively incorporating sky imagery into an LLM-based forecasting framework remains under-explored and an open challenge. In this paper, we propose SolCloudLLM, an LLM-based multimodal forecasting framework. SolCloudLLM aligns sky-image patches with time-series patches and fuses their corresponding representations through bidirectional multimodal fusion, yielding a unified representation that is subsequently mapped into the embedding space of an LLM. Extensive experiments on the SIRTA and SKIPP'D datasets demonstrate that SolCloudLLM consistently outperforms the best baseline methods in MSE across all forecasting horizons, achieving a maximum relative MSE reduction of 25.4%. Stratified analysis further indicates that the benefits of multimodal fusion are concentrated primarily under cloudy conditions. Notably, SolCloudLLM achieves the best performance in nearly all few-shot settings, whereas other deep learning baselines experience substantial performance degradation and are frequently outperformed by the non-learning physical method.
Chinese Translation
短期光伏(PV)功率和全球水平辐照度(GHI)预测对于有效的调度、备用调度和电网运行至关重要。在这些预测时间尺度上,误差主要由云引起的爬坡事件驱动:仅依赖历史数值数据可能难以预测即将到来的云层,因此地基天空图像成为关键的补充物理信号。此外,预测性能对位置和局部观测条件高度敏感,因此对特定地点的数据需求强烈,而这些数据往往稀缺。最近,大语言模型(LLM)在时间序列预测中表现出有竞争力的性能和高数据效率。尽管取得了成功,现有的基于LLM的预测方法仍然主要是单模态的,主要依赖历史数值时间序列数据。有效地将天空图像融入基于LLM的预测框架仍然是一个未被充分探索的开放挑战。在本文中,我们提出了SolCloudLLM,一个基于LLM的多模态预测框架。SolCloudLLM将天空图像块与时间序列块对齐,并通过双向多模态融合将它们的对应表示融合,产生一个统一的表示,随后映射到LLM的嵌入空间。在SIRTA和SKIPP'D数据集上的大量实验表明,SolCloudLLM在所有预测时间尺度上的MSE均持续优于最佳基线方法,最大相对MSE降低25.4%。分层分析进一步表明,多模态融合的好处主要集中在多云条件下。值得注意的是,SolCloudLLM在几乎所有少样本设置中取得了最佳性能,而其他深度学习基线则性能大幅下降,并经常被非学习型物理方法超越。
cs.LG / 37 / 2609.11163

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

LILA:通过潜在谱几何实现大语言模型的无校准结构化剪枝
Behera, Sankar, Singh, Dhruv, Agnihotri, Anshika, Choudhary, Raj Kumar, Ahlawat, Satyadev, Prasad, Yamuna
Abstract
Structured pruning of large language models (LLMs) offers hardware-efficient compression, yet existing methods require calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (\emph{Latent-Informed Layer Analysis}) scores neuron importance via the Kolmogorov--Smirnov (KS) distance between empirical singular value distributions of the full and neuron-ablated feed-forward network (FFN) weight matrix, providing a closed-form spectral rule requiring no training, calibration data, or auxiliary network. Without any fine-tuning, LILA surpasses PruneNet (45M-parameter RL policy) by 1.57~pp in zero-shot accuracy on LLaMA-2-7B at 25\% sparsity, and outperforms WikiText-2-calibrated SliceGPT by up to 6.0~pp across all sparsity levels, while preserving the original architecture. After one epoch of LoRA recovery fine-tuning, LILA achieves highly competitive performance, matching the heavily calibrated SliceGPT baseline to within a 0.48~pp margin across LLaMA-2-7B and Phi-2, despite using zero calibration data. A Neural Tangent Kernel analysis confirms a 22$\times$ reduction in functional distortion versus random pruning, providing theoretical grounding for the spectral importance criterion. Finally, extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.
Chinese Translation
大语言模型(LLM)的结构化剪枝提供了硬件高效的压缩,然而现有方法在剪枝时需要校准数据、梯度计算或大型辅助策略网络。LILA(潜在信息层分析)通过完整和神经元消融的前馈网络(FFN)权重矩阵的经验奇异值分布之间的Kolmogorov-Smirnov(KS)距离来评估神经元重要性,提供了一种闭式谱规则,无需训练、校准数据或辅助网络。无需任何微调,LILA在LLaMA-2-7B上25%稀疏度时零样本准确率超过PruneNet(4500万参数的强化学习策略)1.57个百分点,并且在所有稀疏度级别上优于WikiText-2校准的SliceGPT最多6.0个百分点,同时保持原始架构。经过一个epoch的LoRA恢复微调后,LILA实现了极具竞争力的性能,在LLaMA-2-7B和Phi-2上与经过大量校准的SliceGPT基线相差0.48个百分点以内,尽管使用了零校准数据。神经正切核分析证实,与随机剪枝相比,功能失真减少了22倍,为谱重要性准则提供了理论基础。最后,将LILA扩展为通过KS分数动态分配稀疏度预算,在中等压缩下实现了最先进的生成保留,同时揭示了在高压缩体制下的基本单层架构瓶颈。
cs.LG / 38 / 2609.11166

When does a spectral prior help graph learning? Connectivity-loss estimation under road-network disruptions

谱先验何时有助于图学习?道路网络中断下的连通性损失估计
Le, Van-Truong
Abstract
Rapid evaluation of many simultaneous road-link disruptions requires a practical compromise between exact spectral recomputation and local approximation. We estimate relative algebraic-connectivity loss after multi-edge deletion using graph neural networks (GNNs) that learn a bounded correction to a first-order Fiedler sensitivity. The study considers independent, spatially clustered, and edge-betweenness-targeted failures, with graph-disjoint synthetic splits and zero-shot transfer to 13 OpenStreetMap (OSM) areas in six countries. GCN, GraphSAGE, and edge-aware MPNN backbones are compared with analytical baselines. In expanded OSM tests, residual GCN improves spatial-failure MAE by 0.0391 (95% hierarchical interval 0.0151-0.0662), while residual GraphSAGE improves targeted-failure MAE by 0.0257 (0.0095-0.0446). Second-order perturbation improves first-order MAE by only 0.0028-0.0053. Correction slopes decrease under targeted transfer, indicating residual shrinkage around systematic prior error. Leave-one-country-out OSM-to-OSM transfer is mixed: residual GCN improves targeted-failure MAE by 0.0622 (0.0169-0.1153) but worsens the spatial point estimate. Sparse scaling extends to 20,000 nodes and separates one-time spectral setup from amortized screening cost. These results characterize the spectral residual as a useful but domain-sensitive inductive bias for structural connectivity screening. Code, cached networks, and reproducibility artifacts are archived at doi:10.5281/zenodo.22307723.
Chinese Translation
快速评估许多同时发生的道路链路中断,需要在精确谱重计算和局部近似之间取得实际折衷。我们使用图神经网络(GNN)估计多边删除后的相对代数连通度损失,这些网络学习对一阶Fiedler灵敏度的有界校正。该研究考虑了独立、空间聚簇和以边介数为目标的故障,采用图不相交的合成划分,并零样本迁移到六个国家的13个OpenStreetMap(OSM)区域。将GCN、GraphSAGE和边感知MPNN骨干网络与分析基线进行比较。在扩展的OSM测试中,残差GCN将空间故障MAE改善了0.0391(95%分层区间0.0151-0.0662),而残差GraphSAGE将针对性故障MAE改善了0.0257(0.0095-0.0446)。二阶扰动仅将一阶MAE改善了0.0028-0.0053。校正斜率在针对性迁移下减小,表明残差围绕系统性先验误差收缩。留一国法OSM到OSM迁移结果好坏参半:残差GCN将针对性故障MAE改善了0.0622(0.0169-0.1153),但使空间点估计变差。稀疏缩放扩展到20,000个节点,并将一次性谱设置与摊销筛选成本分开。这些结果将谱残差描述为用于结构连通性筛选的一种有用但对领域敏感的归纳偏置。代码、缓存网络和可复现性工件归档于 doi:10.5281/zenodo.22307723。
cs.LG / 39 / 2609.11168

Semi-Tensor Product-Based Multi-Term Randomized T-SVD and Its Visual Applications

基于半张量积的多项随机化 T-SVD 及其视觉应用
Xiao, Xingchen, Zhang, Feng, Qin, Wenjin, Wang, Jianjun
Abstract
Tensor singular value decomposition (T-SVD), which is built upon the tensor-tensor product (t-product), has emerged as a powerful tool for processing high-dimensional visual data such as color images and videos. However, the standard t-product imposes strict dimensional compatibility constraints. Although extensions based on the semi-tensor product (STP) relax this restriction, their single-term formulations still suffer from limited approximation accuracy. Moreover, these deterministic methods incur high computational costs when processing large-scale tensor data. To address these issues, this paper introduces a novel semi-tensor product for third-order tensors under the t-product framework induced by arbitrary invertible linear transforms. The resulting tensor semi-tensor product breaks the rigid dimension matching requirement of the standard t-product, while retaining the closed-form property of T-SVD. Based on this construction, we develop a multi-term semi-tensor product singular value decomposition (MSTP-SVD), which integrates multiple orthogonal decomposition terms to significantly improve low-rank approximation accuracy compared with single-term schemes. To reduce the computational cost of multi-term modeling, we incorporate randomized projection and power iteration techniques into the MSTP-SVD framework, yielding an accelerated multi-term randomized semi-tensor product SVD (MRSTP-SVD) algorithm that achieves a balance between reconstruction accuracy and computational efficiency. Experiments on image and video compression and completion tasks demonstrate the effectiveness of the proposed method.
Chinese Translation
张量奇异值分解(T-SVD)建立在张量-张量积(t-product)之上,已成为处理彩色图像和视频等高维视觉数据的强大工具。然而,标准 t-product 施加了严格的维度兼容性约束。尽管基于半张量积(STP)的扩展放宽了这一限制,但其单项公式仍然受限于近似精度有限。此外,这些确定性方法在处理大规模张量数据时会产生高昂的计算成本。为了解决这些问题,本文在由任意可逆线性变换诱导的 t-product 框架下,为三阶张量引入了一种新颖的半张量积。所得到的张量半张量积打破了标准 t-product 的严格维度匹配要求,同时保留了 T-SVD 的闭式性质。基于此构造,我们提出了多项半张量积奇异值分解(MSTP-SVD),它集成了多个正交分解项,与单项方案相比显著提高了低秩近似精度。为了降低多项建模的计算成本,我们将随机投影和幂迭代技术融入 MSTP-SVD 框架,得到了一种加速的多项随机化半张量积 SVD(MRSTP-SVD)算法,在重建精度和计算效率之间实现了平衡。在图像和视频压缩与补全任务上的实验证明了所提方法的有效性。
cs.LG / 40 / 2609.11173

Hierarchical Clustering Can Jointly Satisfy Richness, Consistency, and Scale Invariance

层次聚类可以同时满足丰富性、一致性和尺度不变性
Kuroda, Daichi, Dreveton, Maximilien, Grossglauser, Matthias, Thiran, Patrick
Abstract
Despite its ubiquity, clustering lacks a universally accepted definition of what is a cluster. Kleinberg's Impossibility Theorem formalizes this difficulty by showing that no flat clustering method can simultaneously satisfy three natural axioms: scale invariance, richness, and consistency. In this paper, we ask whether this impossibility persists when the output is a hierarchy rather than a single partition. We show that, in contrast to the flat clustering setting, the hierarchical analog of these axioms are jointly satisfiable. In fact, there exist uncountably many hierarchical clustering methods satisfying these axioms, which we call admissible. We explicitly construct several admissible methods, including methods based on well-separated clusters and a non-binary version of single linkage. For certain pairs of admissible methods, the hierarchy produced by one always refines that produced by the other. This refinement relation defines a partial order on the class of admissible methods. This partially ordered set has no greatest element and contains uncountably many pairwise incompatible maximal elements, revealing substantial diversity among admissible methods. Nevertheless, this diversity is constrained: every admissible method contains a hierarchy of sufficiently well-separated clusters, and every finite collection of admissible methods shares such a nontrivial common backbone.
Chinese Translation
尽管聚类无处不在,但缺乏一个普遍接受的关于什么是聚类的定义。Kleinberg 的不可能性定理通过表明没有一种平坦聚类方法能同时满足三个自然公理:尺度不变性、丰富性和一致性,从而将这一困难形式化。在本文中,我们探讨当输出是层次结构而非单一划分时,这种不可能性是否仍然存在。我们证明,与平坦聚类情形相反,这些公理的层次类比是可以同时满足的。事实上,存在不可数多个满足这些公理的层次聚类方法,我们称之为可允许的(admissible)。我们明确构造了几种可允许的方法,包括基于良好分离聚类的方法以及单链(single linkage)的非二元版本。对于某些可允许方法对,其中一个方法产生的层次结构总是细化另一个方法产生的层次结构。这种细化关系定义了可允许方法类上的一个偏序。这个偏序集没有最大元,并且包含不可数多个两两不兼容的极大元,揭示了可允许方法之间的巨大差异。然而,这种多样性是受约束的:每个可允许方法都包含一个由充分良好分离的聚类组成的层次结构,并且每个有限的可允许方法集合都共享这样一个非平凡的共同主干。
cs.LG / 41 / 2609.11207

Convex Optimization with Nested Evolving Feasible Sets (CONES) under Time-Varying Loss Functions

时变损失函数下具有嵌套演化可行集的凸优化(CONES)
Vaze, Rahul
Abstract
Convex Optimization with Nested Evolving Feasible Sets (CONES)} was introduced in \cite{CONESVaze} where the objective function \(f\) remains fixed but the feasible region evolves over time as a nested sequence \(S_1 \supseteq S_2 \supseteq \cdots \supseteq S_T\). The goal of an online algorithm is to simultaneously minimize the regret with respect to hindsight static optimal benchmark and the total movement cost $M_\cA(T)$ while ensuring feasibility at all times. CONES is an optimization-oriented generalization of the well-known \emph{nested convex body chasing} (NCBC). In this paper, we extend CONES to allow for loss functions $f_t'$s to also change over time. When all loss functions are convex, we show that the projected proximal algorithm achieves $O(T^{1-\beta}), O(T^\beta)$ simultaneous regret and movement cost, respectively, for any $\beta \in [0,1)$, over a time horizon of $T$. We also show that any {\it weakly adaptive} online algorithm with $O(T^\beta)$ regret has a movement cost of $\Omega\left(T^{\frac{1-\beta}{2}}\right)$ for any $\beta \in [0,1)$. When all loss functions are strongly convex, we show that the projected proximal algorithm simultaneously achieves $O(1)$ regret and a movement cost of $O(\log T)$. To complement this, we show that any online algorithm with sublinear {\it anytime} regret has a movement cost of $\Omega\left(\log T\right)$.
Chinese Translation
具有嵌套演化可行集的凸优化(CONES)在\cite{CONESVaze}中被引入,其中目标函数\(f\)保持固定,但可行域随时间演化为嵌套序列\(S_1 \supseteq S_2 \supseteq \cdots \supseteq S_T\)。在线算法的目标是在始终确保可行性的同时,同时最小化相对于事后静态最优基准的遗憾和总移动成本\(M_\cA(T)\)。CONES是著名的\emph{嵌套凸体追踪}(NCBC)的一种面向优化的推广。在本文中,我们扩展CONES,使损失函数\(f_t\)也随时间变化。当所有损失函数均为凸函数时,我们证明,在时间范围\(T\)内,对于任意\(\beta \in [0,1)\),投影近端算法同时实现了\(O(T^{1-\beta})\)的遗憾和\(O(T^\beta)\)的移动成本。我们还证明,对于任意\(\beta \in [0,1)\),任何具有\(O(T^\beta)\)遗憾的\it{弱自适应}在线算法,其移动成本为\(\Omega\left(T^{\frac{1-\beta}{2}}\right)\)。当所有损失函数均为强凸函数时,我们证明投影近端算法同时达到\(O(1)\)的遗憾和\(O(\log T)\)的移动成本。作为补充,我们证明任何具有次线性\it{任意时间}遗憾的在线算法,其移动成本为\(\Omega\left(\log T\right)\)。
cs.LG / 42 / 2609.11209

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

REVA:用于上下文高效RAG服务的可重用证据视图聚合
Nguyen, Tuan, Hu, Qiran, Liu, Banruo, Doan, Khoa D., Wong, Kok-Seng, Lai, Fan
Abstract
Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and token cost. Post-retrieval compression can reduce this cost, yet existing compressors often operate independently for each query, rely on auxiliary models or rewriting, and introduce online overhead that can offset the benefit of shorter prompts. We revisit RAG compression from a data-mining perspective by aggregating historical query--document--model interactions into reusable evidence views. We first show that modern compressors have unstable gains over simple truncation and can add substantial inference-time latency. We then propose Reusable Evidence View Aggregation (REVA), a framework that mines the target generator's historical attention traces into a document-keyed, budget-agnostic score store. REVA maps token-level attention to readable word units, aggregates importance across repeated document accesses, and renders budget-specific plain-text views that preserve document order and the standard RAG interface. Across four representative benchmarks and modern LLMs, REVA improves generation quality by 1.0--5.8 points over existing advances, while reducing compression overhead by a factor of 5.3 to 15.6, adding less than 40 ms of latency.
Chinese Translation
检索增强生成(RAG)通过将生成条件建立在检索到的文档上,改善了知识密集型大语言模型(LLM)应用,但更长的上下文会增加延迟、键值(KV)缓存内存和令牌成本。检索后压缩可以降低这种成本,然而现有的压缩器通常针对每个查询独立操作,依赖辅助模型或重写,并引入在线开销,可能抵消较短提示带来的好处。我们从数据挖掘的角度重新审视RAG压缩,通过将历史查询-文档-模型交互聚合为可重用的证据视图。我们首先表明,现代压缩器相较于简单截断的增益不稳定,且可能增加显著的推理时间延迟。然后,我们提出可重用证据视图聚合(REVA),一个将目标生成器的历史注意力轨迹挖掘为文档键控、预算无关的分数存储的框架。REVA将令牌级注意力映射到可读的词单元,聚合重复文档访问的重要性,并渲染预算特定的纯文本视图,保留文档顺序和标准RAG接口。在四个代表性基准和现代LLM上,REVA将生成质量比现有进展提高1.0-5.8分,同时将压缩开销降低5.3至15.6倍,增加不到40毫秒的延迟。
cs.LG / 43 / 2609.11216

Legible Failures: Detecting and Repairing In-Context Binding Errors

可读的失败:检测与修复上下文绑定错误
Ravulapalli, Manas Venkata Sai, Chadha, Samrath Singh, Hari, Abhinav M.
Abstract
A wrong answer does not show whether the model lacked the needed information or held it and failed to use it. On an entity-obligation binding task, a language model can emit an incorrect prompt-supplied binding while a linear probe can recover the correct one from its frozen hidden state. We measure how often this occurs across 16 public checkpoints, each evaluated with three seeds. We fit a probe on a training fold, select its layer on a validation fold, and report results on a disjoint test fold. On the trials each model gets wrong, probe accuracy exceeds the strict present-obligation baseline, 1/K = 0.125, by +0.196 (95% CI [+0.101, +0.296], bootstrapped over models). A query-entity counterfactual rules out token presence and recency. A score built from the sign of probe-output disagreement improves failure detection over the model's own confidence by +0.079 AUROC (95% CI [+0.036, +0.126]). Raw probe confidence gives no measurable improvement over model confidence. Steering the residual stream toward the probe-decoded binding, with no gold label, raises accuracy on all eight models tested by a mean of +0.168 (95% CI [+0.066, +0.280]). Where recent studies report that probe-detected errors are resistant to interventions, we find that in-context binding is a setting in which probes are actionable.
Chinese Translation
一个错误答案并不能表明模型是缺乏所需信息,还是拥有信息却未能使用它。在实体-义务绑定任务中,语言模型可能输出一个错误的提示提供的绑定,而线性探针可以从其冻结的隐藏状态中恢复正确的绑定。我们测量了这种情况在16个公开检查点中发生的频率,每个检查点用三个随机种子进行评估。我们在训练折上拟合一个探针,在验证折上选择其层,并在不相交的测试折上报告结果。在每个模型回答错误的试验中,探针准确率超过严格的当前义务基线 1/K = 0.125,高出 +0.196(95% 置信区间 [+0.101, +0.296],在模型上自助法抽样)。一个查询实体的反事实排除了词元存在和近因效应。一个基于探针输出分歧符号构建的分数,将失败检测相对于模型自身置信度提高了 +0.079 AUROC(95% 置信区间 [+0.036, +0.126])。原始探针置信度相对于模型置信度没有可测量的改进。将残差流导向探针解码的绑定,无需金标准标签,在所有八个测试模型上平均提高了 +0.168(95% 置信区间 [+0.066, +0.280])。在近期研究报告探针检测到的错误对干预具有抵抗性的地方,我们发现上下文绑定是一个探针可发挥作用的场景。
cs.LG / 44 / 2609.11227

Polyhedral Geometry of Time-to-First-Spike Neural Networks

首次脉冲时间神经网络的多面体几何
Singh, Manjot, Montúfar, Guido, Kutyniok, Gitta
Abstract
We study the expressivity of spiking neural networks, which provide a natural framework for asynchronous, event-driven computation complementary to conventional feedforward neural networks. We consider the time-to-first-spike model in a setting for which the input-output map is continuous and piecewise linear, with affine pieces governed by causal feasibility constraints that determine which presynaptic spikes occur before a neuron fires. We first show that each neuron's firing time admits a maxout-like representation with exponentially many, highly constrained affine pieces. We then formalize causal regions as polyhedral regions with fixed causal sets and derive upper and lower bounds on the maximal number of causal regions in both shallow and multilayer feedforward spiking networks. Our theoretical and experimental results show that spiking networks can generate richer partitions of the input space than conventional feedforward ReLU networks.
Chinese Translation
我们研究脉冲神经网络的表达能力,该网络为异步、事件驱动的计算提供了一个自然框架,与常规前馈神经网络互补。我们考虑首次脉冲时间模型,在该设定下,输入-输出映射是连续的分段线性函数,其仿射片段由因果可行性约束决定,这些约束决定哪些突触前脉冲在神经元发放之前发生。我们首先证明,每个神经元的发放时间允许一种类似maxout的表示,具有指数级多且高度约束的仿射片段。然后,我们将因果区域形式化为具有固定因果集的多面体区域,并推导了浅层和多层前馈脉冲神经网络中因果区域最大数量的上界和下界。我们的理论和实验结果表明,脉冲神经网络能够比传统的前馈ReLU网络生成更丰富的输入空间划分。
cs.LG / 45 / 2609.11228

Solving Few-Shot Multiobjective Multitask Optimization via Iterative Sequential Transfer

通过迭代顺序迁移求解少样本多目标多任务优化
Wei, Tingyang, Wu, Haofeng, Iman, Ananda Phan, Wei, Zhao, Liu, Jiao, Ong, Yew-Soon
Abstract
Applying knowledge transfer across multiple optimization tasks, multitask optimization (MTO) emerges as a promising approach to solving synergistic optimization tasks simultaneously. However, the development of effective knowledge transfer mechanisms in MTO fundamentally relies on aligning elite solution distributions across tasks. This dependency creates a critical bottleneck in few-shot optimization regimes, as restricted evaluation budgets impede the identification of elite solution distributions required for beneficial transfer. This challenge is exacerbated in multiobjective multitask problems, where each optimizer must approximate a continuous Pareto manifold rather than a single optimal point. This paper introduces Iterative Sequential Transfer (IST) to circumvent this bottleneck. We model MTO as a sequence of sequential transfer optimization problems, concentrating evaluations on a single target per iteration. We propose a likelihood-informed task prioritization mechanism to maximize transfer utility by identifying the task most likely ready for knowledge integration. Empirical results on benchmark and real-world problems verify the effectiveness of the proposed method under tight budgets.
Chinese Translation
在多个优化任务间应用知识迁移,多任务优化(MTO)成为同时求解协同优化任务的一种有前景的方法。然而,MTO中有效知识迁移机制的发展从根本上依赖于跨任务对齐精英解分布。这种依赖在少样本优化场景中造成了关键瓶颈,因为有限的评估预算阻碍了识别有益迁移所需的精英解分布。这一挑战在多目标多任务问题中更加严重,其中每个优化器必须近似连续的Pareto流形而不是单个最优点。本文引入迭代顺序迁移(IST)来规避这一瓶颈。我们将MTO建模为一系列顺序迁移优化问题,每次迭代将评估集中在单个目标上。我们提出了一种似然信息驱动的任务优先级机制,通过识别最可能准备好进行知识整合的任务来最大化迁移效用。在基准和真实世界问题上的实证结果验证了所提方法在紧张预算下的有效性。
cs.LG / 46 / 2609.11253

MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions

MUtE:概念擦除与反事实干预的双重框架
Saillenfest, Antoine
Abstract
Erasing concept-specific information from representations has been proven useful for mitigating bias or interpreting model decisions. The joint objective is to transform the original representations such that the target concept becomes unpredictable, while maximally preserving concept-unrelated information. In this work, we revisit the optimal bounds of concept erasure to derive a novel class of erasure functions that naturally induce a deterministic, dual counterfactual mapping. Bridging the gap between theoretical optimality and practical representation learning, we design an implementation that imposes a translational bias on counterfactual trajectories - a constraint that aligns with how many concepts geometrically manifest in modern language models. Our framework enables seamless navigation between concept erasure and counterfactual generation. We empirically demonstrate its efficacy in improving downstream algorithmic fairness and generating counterfactual texts.
Chinese Translation
从表示中擦除概念特定信息已被证明有助于减轻偏差或解释模型决策。其联合目标是转换原始表示,使目标概念变得不可预测,同时最大程度地保留与概念无关的信息。在这项工作中,我们重新审视概念擦除的最优界,以推导出一类新的擦除函数,其能够自然地诱导出确定性的对偶反事实映射。为弥合理论最优性与实际表示学习之间的差距,我们设计了一种实现,对反事实轨迹施加平移偏置——这一约束与现代语言模型中许多概念的几何呈现方式相一致。我们的框架能够在概念擦除与反事实生成之间无缝导航。我们通过实验证明了其在提升下游算法公平性和生成反事实文本方面的有效性。
cs.LG / 47 / 2609.11314

A Dynamic Fusion Large Language Model for Traffic Flow Prediction

用于交通流预测的动态融合大语言模型
Qiu, Xue, Xiao, Jianli
Abstract
Traffic flow prediction is a core supporting technology for intelligent transportation systems. It uses historical data to infer future traffic dynamics in specific areas, thereby helping to alleviate congestion and improve resource allocation efficiency. Traditional neural networks struggle to break through accuracy limits due to their reliance on singular feature modeling, while large language models (LLMs) suffer from insufficient capture of spatial topological information and mining spatiotemporal correlation. This study proposes a Dynamic Fusion Large Language Model (DF-LLM) for traffic flow prediction. The model incorporates three core components: spatiotemporal embedding module, spatiotemporal fusion module, and LLM backbone. The spatiotemporal embedding module enables synergistic representation of multi-scale spatiotemporal features. The spatiotemporal fusion module integrates spatial topology and dynamic dependencies via graph convolution. The LLM backbone adopts a differentiated parameter adaptation strategy to balance training efficiency and traffic data adaptability. Additionally, it introduces a context aggregation attention module to strengthens global dependencies. More importantly, the LLM backbone takes the residual connections to mitigate the gradient vanishing in deep networks. Experiments show that DF-LLM has achieved better performance by comparing the metrics on all the four datasets.
Chinese Translation
交通流预测是智能交通系统的核心支撑技术。它利用历史数据推断特定区域的未来交通动态,从而帮助缓解拥堵并提高资源分配效率。传统神经网络由于依赖单一特征建模,难以突破精度限制,而大语言模型(LLMs)在捕捉空间拓扑信息和挖掘时空相关性方面存在不足。本研究提出了一种用于交通流预测的动态融合大语言模型(DF-LLM)。该模型包含三个核心组件:时空嵌入模块、时空融合模块和大语言模型(LLM)骨干。时空嵌入模块实现了多尺度时空特征的协同表示。时空融合模块通过图卷积整合空间拓扑和动态依赖关系。LLM骨干采用差异化参数自适应策略,以平衡训练效率和交通数据适应性。此外,它引入了上下文聚合注意力模块以增强全局依赖关系。更重要的是,LLM骨干采用残差连接以缓解深度网络中的梯度消失问题。实验表明,通过在所有四个数据集上比较指标,DF-LLM取得了更好的性能。
cs.LG / 48 / 2609.11331

Estimating Inconsistency Response Surfaces under Uncertainty in Cyber-Physical System Development

估计信息物理系统开发中不确定性下的不一致响应面
Mäkelburg, Johannes, Schwabe, Tim, Acosta, Maribel
Abstract
Cyber-Physical Systems (CPS) are commonly represented through multiple interconnected models. During development, CPS consistency requires that shared model elements remain compatible across these models. Uncertainty, for example, due to sensor noise or model abstraction, changes the admissible values of model elements and can introduce inconsistencies, i.e., situations in which models can no longer be jointly satisfied. While existing approaches can determine consistency for a given uncertainty configuration, they provide limited support for systematically exploring, analyzing, and explaining inconsistency across large uncertainty spaces. We address this challenge by reformulating inconsistency as an intervention response modeling problem. Using Saltelli sampling and multi-fidelity Monte Carlo estimation, we generate intervention-response datasets and train a surrogate model that directly predicts inconsistency from the propagated uncertainty geometry. Experiments on 48 scenarios and 10 CPS domains show that the surrogate matches Monte Carlo estimates while reducing evaluation time from milliseconds to microseconds, enabling orders-of-magnitude more response-surface evaluations within fixed computational budgets. Building on the learned response surfaces, we perform sensitivity analysis to identify dominant uncertainty drivers and introduce a gradient-based consistency recourse method to determine minimal uncertainty interventions that restore consistency. The results show that inconsistency under uncertainty can be effectively learned, analyzed, and repaired through response-surface modeling, providing a scalable foundation for uncertainty-aware consistency management in CPS development.
Chinese Translation
信息物理系统(CPS)通常通过多个互连模型来表示。在开发过程中,CPS一致性要求共享的模型元素在这些模型之间保持兼容。不确定性(例如,由于传感器噪声或模型抽象)会改变模型元素的允许值,并可能引入不一致性,即模型无法再同时满足的情况。虽然现有方法可以针对给定的不确定性配置确定一致性,但它们在系统性地探索、分析和解释大型不确定性空间中的不一致性方面提供的支持有限。我们通过将不一致性重新表述为干预响应建模问题来应对这一挑战。利用Saltelli采样和多保真度蒙特卡洛估计,我们生成干预-响应数据集,并训练一个代理模型,该模型直接从传播的不确定性几何中预测不一致性。在48个场景和10个CPS领域上的实验表明,该代理模型与蒙特卡洛估计相匹配,同时将评估时间从毫秒减少到微秒,使得在固定计算预算内能够进行数量级更多的响应面评估。基于学习到的响应面,我们进行敏感性分析以识别主要的不确定性驱动因素,并引入一种基于梯度的一致性补救方法,以确定恢复一致性的最小不确定性干预。结果表明,通过响应面建模,可以有效地学习、分析和修复不确定性下的不一致性,为CPS开发中的不确定性感知一致性管理提供了可扩展的基础。
cs.LG / 49 / 2609.11347

Reification as a Transferable Vocabulary: Zero-Shot Link Prediction with Vanilla GNNs

具体化作为可迁移词汇表:使用普通GNN的零样本链接预测
Pradel, Camille
Abstract
Knowledge graph foundation models such as ULTRA achieve zero-shot link prediction on unseen graphs through dedicated architectures that hard-code a transfer mechanism. In this work we move that mechanism out of the architecture and into the representation, by \emph{reifying} the input graph: every fact becomes a node, connected to its subject, object, and relation type through a fixed vocabulary of six meta-relations, with relation types as anonymous shared nodes rather than model parameters. On this representation, five textbook GNNs (GAT, GINE with sum and with mean+max aggregation, GraphSAGE, R-GCN), each trained on a single knowledge graph of 4,245 triples for 30 minutes on one NVIDIA A100, transfer zero-shot to 40 inductive link-prediction benchmarks. The best of them, an off-the-shelf GAT, matches ULTRA, a dedicated foundation model pretrained on three graphs, across ULTRA's own evaluation suite. The same fixed vocabulary extends to relational databases, a row becoming an entity and a foreign-key column a relation type; a preliminary probe on two unseen databases, with no cell values, schema text or in-context labels, shows a model of this family pretrained on three knowledge graphs ranking foreign-key targets far above random-initialization and degree controls. We release the code, the checkpoints, and the evaluation pipeline for all 40 benchmarks.
Chinese Translation
知识图谱基础模型(如ULTRA)通过硬编码迁移机制的专用架构,在未见图上实现零样本链接预测。在这项工作中,我们将该机制从架构中移出并放入表示中,通过具体化(reifying)输入图:每个事实变成一个节点,通过由六个元关系组成的固定词汇表连接到其主体、客体和关系类型,关系类型作为匿名共享节点而非模型参数。在这种表示上,五种教科书式GNN(GAT、采用求和以及均值+最大聚合的GINE、GraphSAGE、R-GCN),每个在一个包含4,245个三元组的知识图谱上,使用一块NVIDIA A100训练30分钟,零样本迁移到40个归纳链接预测基准。其中最好的一个,即现成的GAT,在ULTRA自己的评估套件上与ULTRA(一个在三个图上预训练的专用基础模型)表现相当。相同的固定词汇表可扩展到关系数据库:一行成为一个实体,外键列成为一个关系类型;在两个未见过的数据库上的初步探测(没有单元格值、模式文本或上下文标签)表明,一个在三个知识图谱上预训练的该家族模型在排名外键目标时远优于随机初始化和度控制基线。我们发布了所有40个基准的代码、检查点和评估流程。
cs.LG / 50 / 2609.11366

Local Robustness Quantification for Naive Bayes Classifiers and Generative Forests: a General Approach

朴素贝叶斯分类器和生成森林的局部鲁棒性量化:一种通用方法
Detavernier, Adrián, De Bock, Jasper
Abstract
We provide methods for calculating the robustness of the predictions of two types of generative classifiers whose underlying distribution is a Probabilistic Graphical Model (PGM): naive Bayes classifiers and generative forests (a probabilistic extension of random forests). Following the paradigm of robustness quantification, we define the robustness of a prediction as the extent to which the distribution of the classifier can be perturbed without changing this prediction. We consider perturbations obtained by varying the local models of the PGMs within general neighborhoods and focus in particular on epsilon-contamination, total variation distance and chi-squared divergence balls. We test our methods on benchmark datasets, demonstrate that the robustness value of a prediction serves as an indicator for its trustworthiness and compare our approach with other such indicators.
Chinese Translation
我们提供了用于计算两类生成式分类器预测鲁棒性的方法,这两类分类器的基础分布是概率图模型(PGM):朴素贝叶斯分类器和生成森林(随机森林的概率扩展)。遵循鲁棒性量化的范式,我们将预测的鲁棒性定义为分类器的分布在不改变该预测的情况下可以被扰动的程度。我们考虑通过在一般邻域内改变PGM的局部模型所获得的扰动,并特别关注ε-污染、总变差距离和卡方散度球。我们在基准数据集上测试了我们的方法,证明了预测的鲁棒性值可作为其可信度的指标,并将我们的方法与其他此类指标进行了比较。
cs.LG / 51 / 2609.11449

Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets

流行率决定精确度:检测器定义数据集中的沉默污染
Huang, Jia, Wan, Yankai, Ou, Yangjun
Abstract
Many ML datasets are constructed by running a detector, heuristic, or model over candidate pools; accepted items become labels. Dataset precision is then governed by true-positive prevalence in each pool via Bayes, not solely by detector quality. Using one instrument and period, we hold a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom. One detector, three pools yield phantom rates 81.7%, 9.0%, and 0.0%. Transferring precision from the two high-rate pools to the low-rate pool predicts 0.955 versus measured 0.183, a +422% error; the Bayes expression predicts all three within 3.3%. The detected response curve is an exact convex combination of a true-event and a phantom component (residual 1.1e-16), with phantoms outnumbering true events 473 to 308, so contamination is a second signal with detector-inherited shape, not additive noise. Contamination direction depends on the estimator: on identical windows one statistic is diluted and another inflated because its denominator is also contaminated. A common normalization turns the estimator into a mean of ratios whose expectation need not exist; on the same 335 events it returns 0.40 where the well-defined estimator returns 0.10.
Chinese Translation
许多机器学习数据集是通过在候选池上运行检测器、启发式方法或模型来构建的;被接受的条目成为标签。数据集精确度随后由每个池中的真实阳性流行率通过贝叶斯决定,而不仅仅由检测器质量决定。使用一个仪器和一个时间段,我们持有一个检测器定义的事件数据集,外加一个独立的官方索引,该索引将每个检测到的项目标记为真实或幻影。一个检测器,三个池产生的幻影率分别为81.7%、9.0%和0.0%。将精确度从两个高幻影率池转移到低幻影率池,预测值为0.955,而测量值为0.183,误差为+422%;贝叶斯表达式预测所有三个池的误差在3.3%以内。检测到的响应曲线是真实事件和幻影分量的精确凸组合(残差1.1e-16),其中幻影数量超过真实事件,为473比308,因此污染是第二个信号,具有检测器继承的形状,而不是加性噪声。污染方向取决于估计器:在相同的窗口上,一个统计量被稀释,另一个被夸大,因为其分母也被污染。一种常见的归一化将估计器变成比率均值,其期望不一定存在;在相同的335个事件上,它返回0.40,而定义良好的估计器返回0.10。
cs.LG / 52 / 2609.11495

Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study

结合合成数据与真实数据用于低资源历史OCR:一项满文案例研究
Chung, Yan Hon Michael, Wang, Hanlin
Abstract
Manchu, now critically endangered, was one of the principal languages of the Qing empire (1636-1912), and its extensive archival record is increasingly digitized but remains difficult to search and analyze at scale. Previous work showed that vision-language models (VLMs) trained only on synthetic Manchu word images can reach 87.4% word accuracy on real Qing manuscripts and prints, leaving a substantial synthetic-to-real gap. This study examines how synthetic and real historical training data should be combined for low-resource OCR. Using 60,000 synthetic and 20,306 real historical word images, we evaluate three pretrained VLMs and a compact convolutional recurrent neural network (CRNN) under four regimes: synthetic-only, real-only, joint synthetic-real, and sequential synthetic-to-real training, following a common checkpoint-selection and archival evaluation protocol. Introducing real training images raises the leading configurations to between 95.09% and 96.28% word accuracy, while no synthetic-only configuration exceeds 87.92%. Synthetic supplementation substantially improves all three VLMs, whereas its marginal effect for the CRNN is sensitive to the training objective. Joint and sequential training yield broadly similar archival accuracy under the tested practical pipelines. A compact CRNN also reaches the leading performance range once real images are available, showing that model scale alone does not determine recognition accuracy. Finally, complementary errors among strong recognizers allow voting to raise accuracy to 98.27% without additional training, while an eighteenth-century Manchu dictionary provides a principled rule for adjudicating disagreements.
Chinese Translation
满文如今已极度濒危,曾是清帝国(1636—1912)的主要语言之一;其大量档案记录正日益数字化,但仍难以大规模检索和分析。先前研究表明,仅在合成满文词图像上训练的视觉语言模型(VLM)在真实清代手稿和刻本上可达87.4%的词准确率,仍存在显著的合成到真实差距。本研究考察在低资源OCR中应如何结合合成与真实历史训练数据。我们使用60,000幅合成词图像和20,306幅真实历史词图像,在四种训练模式下评估三种预训练VLM和一个紧凑卷积循环神经网络(CRNN):仅合成、仅真实、合成-真实联合,以及从合成到真实的顺序训练,并遵循统一的检查点选择与档案评估协议。引入真实训练图像后,领先配置达到95.09%至96.28%的词准确率,而没有任何仅合成配置超过87.92%。合成数据补充显著改善了所有三种VLM,而其对于CRNN的边际效应则对训练目标敏感。在测试的实际流程下,联合训练与顺序训练产生大体相似的档案准确率。一旦有真实图像可用,紧凑CRNN也能达到领先性能范围,表明仅靠模型规模并不能决定识别准确率。最后,强识别器之间的互补错误使投票无需额外训练即可将准确率提高到98.27%,而一部18世纪满文词典为裁决分歧提供了有原则的规则。
cs.LG / 53 / 2609.11504

DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis

DeFiFlowBench:自然语言 DeFi 工作流合成中的安全可执行性基准测试与改进
Kumar, Abhinav Rajeev, Arora, Harshit, Singh, Varun, Nanjappan, Manikandan
Abstract
A structurally valid DeFi workflow can still authorize a costly trade. We introduce DeFiFlowBench, a benchmark of 207 team-authored prompts for natural-language DeFi workflow synthesis. It measures graph coverage, configuration completeness, and declared safety predicates, then tests supported trade configurations on a local EVM. Direct, constrained, and few-shot prompting produce 14-19 unsafe held-out executions per configuration under a fixed 5% price-impact cap. A slippage bound derived from a quote does not prevent the price impact of the order itself. We propose Koan-Safe, which combines a prompt-only intent parser, a replaceable generator, and structural repair with default safety parameters. On 75 held-out workflow prompts, its hybrid variant scores 0.67 on the static safety proxy, compared with 0.33 for the best baseline. Koan-Safe records no unsafe executions on the saved benchmark outputs. A matched-candidate ablation produces 14-17 unsafe executions when enforcement is disabled. Additional tests expose the limits of default injection: permissive existing thresholds can still authorize unsafe trades. A separately evaluated policy cap addresses this failure on a 36-case diagnostic grid. These results support explicit trade protections and execution-based evaluation, while distinguishing declared safety from a general guarantee.
Chinese Translation
结构有效的 DeFi 工作流仍可能授权一笔代价高昂的交易。我们提出 DeFiFlowBench,一个包含 207 个团队编写的提示的基准,用于自然语言 DeFi 工作流合成。它衡量图覆盖率、配置完整性和声明的安全谓词,然后在本地 EVM 上测试受支持的交易配置。在固定的 5% 价格影响上限下,直接提示、约束提示和少样本提示每种配置产生 14-19 次不安全的留出执行。从报价推导出的滑点界限并不能防止订单本身的价格影响。我们提出 Koan-Safe,它结合了仅提示的意图解析器、可替换的生成器以及带有默认安全参数的结构化修复。在 75 个留出工作流提示上,其混合变体在静态安全代理指标上得分为 0.67,而最佳基线为 0.33。Koan-Safe 在保存的基准输出上未记录任何不安全执行。当禁用强制执行时,匹配候选消融实验产生 14-17 次不安全执行。额外的测试暴露了默认注入的局限性:宽松的现有阈值仍然可能授权不安全的交易。一个单独评估的策略上限在一个 36 个案例的诊断网格上解决了这一失败。这些结果支持显式交易保护和基于执行的评估,同时将声明的安全与一般性保证区分开来。
cs.LG / 54 / 2609.11521

Generalized Score Matching for Parameter Estimation on Convex Domains

凸域上参数估计的广义分数匹配
Shetty, Nishanth, Mahajan, Saisuchith, Seelamantula, Chandra Sekhar
Abstract
Maximum likelihood (ML) estimation is a principled and statistically efficient approach for learning probabilistic models. However, for unnormalized models, ML estimation requires evaluating the partition function and differentiating through it, which may not always be tractable. Score matching provides a practically viable alternative that circumvents this obstacle by fitting the score in a way that eliminates dependence on the normalizing constant. We derive the generalized score matching objective on a convex subset of $\mathbb{R}^{d}$ constructively starting from Minimum Probability Flow (MPF) learning, and show how classical score matching as well as domain-adapted variants for non-negative data arise naturally within the proposed framework. We show that the resulting objective is a {\it proper local scoring rule} of second-order, which provides the theoretical guarantee that the true density is recovered when the objective is minimized. Furthermore, for a model belonging to the exponential family, we establish convexity of the objective together with consistency of the finite-sample estimator under standard regularity conditions. Our derivation sheds new light on the scope and applicability of generalized score matching in various problem settings. We compare generalized score matching-based estimators on constrained domains, where the partition function is analytically intractable. We provide experimental results on parameter estimation for model densities belonging to the exponential family defined over convex subsets of $\mathbb{R}^{d}$, and a generative modeling use-case to demonstrate broader applicability of the proposed generalized score matching framework.
Chinese Translation
最大似然(ML)估计是学习概率模型的一种有理论依据且统计高效的方法。然而,对于未归一化模型,ML估计需要计算配分函数并对其求导,而这并不总是易于处理。分数匹配提供了一种实际可行的替代方案,它通过以消除对归一化常数依赖的方式拟合分数,绕开了这一障碍。我们从最小概率流(MPF)学习出发,构造性地推导了 R^d 的凸子集上的广义分数匹配目标函数,并展示了经典分数匹配以及针对非负数据的域适配变体如何自然地在所提出的框架中出现。我们证明,所得目标是二阶的恰当局部评分规则,这提供了当目标函数最小化时真实密度可被恢复的理论保证。此外,对于属于指数族的模型,我们在标准正则条件下建立了目标函数的凸性以及有限样本估计量的一致性。我们的推导为广义分数匹配在各种问题设定中的范围和适用性提供了新见解。我们比较了约束域上基于广义分数匹配的估计量,其中配分函数在解析上难以处理。我们提供了针对定义在 R^d 的凸子集上的指数族模型密度的参数估计实验结果,以及一个生成建模用例,以展示所提出的广义分数匹配框架更广泛的适用性。
cs.LG / 55 / 2609.11538

Particle GFlowNets: Rethinking Generative Marginalization Models

Particle GFlowNets:重新思考生成式边缘化模型
da Silva, Tiago, Mesquita, Diego, Lahlou, Salem
Abstract
Generative Marginalization Models (MaMs) have been recently introduced as efficient neural sampling models for any-order autoregressive modelling of discrete distributions. By learning both the marginal and conditional probabilities of a persistent-block Gibbs sampler, MaMs enable fast posterior evaluation with a single neural network forward pass. While prior work has considered MaMs to be distinct from Generative Flow Networks (GFlowNets), a well-established paradigm for inference in discrete stochastic models, we show that they are equivalent. Then, we also extend MaMs' sampling strategy to non-autoregressive generative processes. In particular, we describe an automatic criterion for full-state rejuvenation of the Gibbs sampler, derived from the Gelman-Rubin statistic, which plays a key role in speeding up learning convergence. Our experiments show that our method, called Particle GFlowNets, markedly accelerates training in large combinatorial spaces.
Chinese Translation
生成式边缘化模型(Generative Marginalization Models, MaMs)最近被提出,作为用于离散分布任意顺序自回归建模的高效神经采样模型。通过学习持久块吉布斯采样器的边缘概率和条件概率,MaMs 能够通过单次神经网络前向传播实现快速后验评估。尽管先前工作认为 MaMs 与生成流网络(Generative Flow Networks, GFlowNets)不同——后者是离散随机模型推断中一个成熟范式——但我们表明二者是等价的。随后,我们还将 MaMs 的采样策略扩展到非自回归生成过程。特别地,我们描述了一种源自 Gelman-Rubin 统计量的自动准则,用于吉布斯采样器的全状态更新,该准则在加速学习收敛方面发挥关键作用。我们的实验表明,我们称为 Particle GFlowNets 的方法在大型组合空间中显著加速训练。
cs.LG / 56 / 2609.11580

A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph

大型极稀疏时空图上水面高程插补的数据集与模型
Cartuyvels, Ruben, Douch, Karim, Bertoli, Gabriele, Baz, Mounia El, Vrettou, Artemis, Lefèvre, Sébastien, Prieto, Diego Fernandez
Abstract
Continuous monitoring of water surface elevation across river networks is critical for flood forecasting, water resource management, and understanding the global water cycle. Yet, the scarcity of in situ gauges across much of the globe constrains the development of reliable modeling frameworks. Satellite altimetry has the potential to alleviate this problem but its use is currently hindered by sparse temporal coverage. To this end, we introduce AmazonSWE, a dataset for training and evaluating large-scale spatiotemporal graph imputation methods that integrates processed satellite altimetry measurements from a range of sources, including the recent wide-swath SWOT sensor. The dataset covers over 19K river sections and 10 years (2016-2026) in the Amazon river basin, with in situ gauges held out for evaluation. Besides contributing a novel real-world use case with the potential for societal impact, AmazonSWE introduces significant technical challenges: with fewer than 1% of sections observed per day, the dataset is far sparser than existing imputation benchmarks, and its directed acyclic river topology is both structurally different from and larger than graphs in existing datasets. We show that prior spatiotemporal graph imputation methods are not adapted to this topology, scale and sparsity, and propose a simple bidirectional selective state space model that outperforms them by sampling connected subgraphs and flattening space and time into a single token sequence with topology-aware positional encodings. Compared to the state-of-the-art published method for SWOT-based WSE densification, which integrates statistics with physical modeling, our model reduces RMSE against in situ gauges by 18-39%, while producing predictions for every river section rather than only those with sufficient nearby satellite coverage.
Chinese Translation
对河网水面高程的连续监测对于洪水预报、水资源管理和理解全球水循环至关重要。然而,全球大部分地区原位观测站点的稀缺限制了可靠建模框架的发展。卫星测高有潜力缓解这一问题,但其应用目前受到时间覆盖稀疏的阻碍。为此,我们介绍了AmazonSWE,一个用于训练和评估大规模时空图插补方法的数据集,它整合了来自一系列来源的处理过的卫星测高测量数据,包括最近的宽刈幅SWOT传感器。该数据集覆盖亚马逊河流域超过1.9万个河流断面和10年(2016-2026年),并留出原位观测站点用于评估。除了贡献一个具有社会影响潜力的新颖真实世界用例之外,AmazonSWE还引入了重大技术挑战:每天观测到的断面不足1%,该数据集比现有的插补基准稀疏得多,并且其有向无环河流拓扑在结构上既不同于现有数据集中的图,也更大。我们表明,先前的时空图插补方法不适应这种拓扑、规模和稀疏性,并提出了一种简单的双向选择性状态空间模型,该模型通过采样连通子图并将空间和时间展平为带有拓扑感知位置编码的单一token序列,从而优于它们。与已发表的基于SWOT的水面高程(WSE)加密的最先进方法(该方法将统计与物理建模相结合)相比,我们的模型将相对于原位观测站点的RMSE降低了18-39%,同时为每个河流断面生成预测,而不仅仅是那些附近有足够卫星覆盖的断面。
cs.LG / 57 / 2609.11639

LoaDiff: Conditional Generation of Electricity Consumption Time Series for Energy Analytics

LoaDiff:面向能源分析的电力消费时间序列条件生成
Baranova, Mariia, Petralia, Adrien, Naour, Etienne Le, Etourneau, Nathan, Hofmann, Guillaume, Palpanas, Themis
Abstract
The energy transition is reshaping residential electricity consumption through the increasing adoption of distributed generation, electrified appliances, and demand-response programs. Understanding these evolving behaviors requires access to granular smart-meter data for applications such as load forecasting, appliance detection, and demand-side flexibility analysis. However, such data are subject to strict access restrictions and data-protection regulations. Thus, realistic synthetic alternatives are necessary. In this paper, we introduce LoaDiff, a diffusion-based generative model for year-long, sub-hourly smart-meter load curves. LoaDiff supports flexible conditioning on static household attributes, such as appliance ownership, and dynamic contextual variables, including calendar information and outdoor temperature. We evaluate the model against multiple generative baselines on three residential electricity-consumption datasets. Our experiments assess four complementary dimensions: fidelity and diversity, training-record memorization risk, downstream utility for load forecasting and appliance detection, and conditional controllability under alternative temperature conditions. The results show that LoaDiff generates realistic and diverse load profiles, achieves a favorable trade-off between generation quality and limited evidence of memorization, preserves information useful for downstream energy applications, and responds coherently to changes in conditioning variables.
Chinese Translation
能源转型正在通过越来越多地采用分布式发电、电气化电器和需求响应项目,重塑住宅用电消费。理解这些不断演变的行为,需要获取细粒度的智能电表数据,以用于负荷预测、电器识别和需求侧灵活性分析等应用。然而,这类数据受到严格的访问限制和数据保护法规约束。因此,需要逼真的合成替代数据。本文中,我们提出 LoaDiff,一种基于扩散的生成模型,用于生成全年、亚小时级智能电表负荷曲线。LoaDiff 支持灵活地以静态家庭属性(如电器拥有情况)和动态上下文变量(包括日历信息和室外温度)为条件。我们在三个住宅用电数据集上,将该模型与多个生成基线进行评估比较。我们的实验评估了四个互补维度:保真度与多样性、训练记录记忆风险、面向负荷预测和电器识别的下游效用,以及在不同温度条件下的条件可控性。结果表明,LoaDiff 能生成逼真且多样的负荷曲线,在生成质量与有限的记忆证据之间取得良好权衡,保留对下游能源应用有用的信息,并对条件变量的变化做出一致响应。
cs.LG / 58 / 2609.11648

RDDMPI: Residual Denoising Diffusion Model for Probabilistic Multivariate Time Series Imputation

RDDMPI:用于概率多变量时间序列插补的残差去噪扩散模型
Jara, Ramiro Valdes, Chapman, David, Meyers, Adam
Abstract
Multivariate time series imputation (MTSI) aims to recover missing values in temporal data composed of multiple interdependent variables. This problem is central to real-world applications such as healthcare monitoring, traffic networks, and energy systems. Recent diffusion-based approaches have shown strong potential for probabilistic imputation by learning to generate missing values through iterative denoising. However, most existing approaches perform diffusion directly in the original data space, requiring the denoising network to simultaneously capture global structure, temporal dynamics, and stochastic variability. This makes the generative task unnecessarily complex, especially when modern deterministic imputers can already provide accurate initial reconstructions. To address this limitation, we propose RDDMPI, a conditional residual diffusion framework that operates directly in residual space. Instead of modeling the full missing signal directly, we reformulate probabilistic imputation as a baseline-residual decomposition, where a pretrained model captures the dominant signal and a diffusion process models the residual uncertainty. To better exploit deterministic guidance, \model{} conditions the reverse denoising process on both the baseline-completed signal and its latent representation, while a reliability-aware conditioning mechanism adaptively controls the influence of baseline information during residual generation. This formulation simplifies the diffusion learning objective, enabling it to focus on structured correction terms rather than reconstructing the full signal. Experiments on multiple benchmark datasets demonstrate that RDDMPI consistently improves both reconstruction accuracy and uncertainty quantification.
Chinese Translation
多变量时间序列插补(MTSI)旨在恢复由多个相互依赖变量组成的时间数据中的缺失值。这个问题对于医疗监测、交通网络和能源系统等现实世界应用至关重要。最近基于扩散的方法通过迭代去噪学习生成缺失值,在概率插补方面显示出强大潜力。然而,大多数现有方法直接在原始数据空间中进行扩散,要求去噪网络同时捕获全局结构、时间动态和随机变异性。这使得生成任务不必要地复杂,尤其是当现代确定性插补器已经可以提供准确的初始重建时。为了解决这一局限性,我们提出了RDDMPI,一种直接在残差空间中操作的条件残差扩散框架。我们不是直接对完整的缺失信号建模,而是将概率插补重新表述为基线-残差分解,其中预训练模型捕获主导信号,扩散过程对残差不确定性建模。为了更好地利用确定性指导,RDDMPI将反向去噪过程条件于基线补全信号及其潜在表示,同时可靠性感知的条件机制在残差生成过程中自适应地控制基线信息的影响。这种表述简化了扩散学习目标,使其能够专注于结构化校正项而不是重建完整信号。在多个基准数据集上的实验表明,RDDMPI持续提高了重建精度和不确定性量化。
cs.LG / 59 / 2609.11655

Musec: MomentUm SpEctral Clipping for Stable Muon-type Training

Musec:用于稳定 Muon 型训练的动量谱裁剪
Liu, Zhuanghua, Wang, Menglian, Luo, Luo
Abstract
Muon has emerged as a highly effective optimizer for large language model training, often achieving superior convergence and performance compared with the widely adopted Adam and AdamW optimizers. Nevertheless, Muon is prone to training instability due to its spectral flattening, manifested by loss spikes and unbounded growth of model weights. Existing approaches primarily rely on weight or attention-logit clipping, which require architecture-specific modifications and do not directly address instability across all model components. We propose MomentUm SpEctral Clipping (Musec), which replaces Muon's spectral flattening with spectral clipping: rather than setting all singular values of the momentum matrix to approximately one, Musec clips singular values that exceed a threshold while preserving the underlying spectral structure of the momentum. Our strategy provides an optimizer-level, architecture-agnostic mechanism for stabilizing Muon training. We further develop Soft Musec, an efficient implementation that uses a smooth spectral saturation function approximated by coupled Newton-Schulz iterations. Theoretically, we establish convergence guarantees for Musec in nonconvex nonsmooth stochastic optimization. To the best of our knowledge, this is the first convergence guarantee for Muon-type methods in the nonconvex nonsmooth setting. We provide empirical studies to show that Soft Musec consistently improves training stability over existing Muon variants across a wide range of learning rates and model sizes. Notably, Soft Musec remains stable in settings where existing Muon variants diverge, while matching their performance under well-tuned configurations.
Chinese Translation
Muon 已成为一种用于大型语言模型训练的高效优化器,与广泛采用的 Adam 和 AdamW 优化器相比,通常能实现更优的收敛性和性能。然而,由于其所采用的谱平坦化,Muon 容易出现训练不稳定问题,表现为损失尖峰和模型权重的无界增长。现有方法主要依赖权重或注意力 logit 裁剪,这些方法需要针对特定架构进行修改,并且不能直接解决所有模型组件中的不稳定性。我们提出了动量谱裁剪(MomentUm SpEctral Clipping, Musec),用谱裁剪替代 Muon 的谱平坦化:Musec 不会将动量矩阵的所有奇异值都设置为近似于一,而是裁剪超过阈值的奇异值,同时保留动量的底层谱结构。我们的策略提供了一种优化器层面、与架构无关的机制,用于稳定 Muon 训练。我们进一步开发了 Soft Musec,这是一种高效实现,使用由耦合 Newton-Schulz 迭代近似的平滑谱饱和函数。在理论上,我们为 Musec 在非凸非光滑随机优化中建立了收敛保证。据我们所知,这是非凸非光滑设定下 Muon 型方法的首个收敛保证。我们通过实证研究表明,在广泛的学习率和模型规模范围内,Soft Musec 相比现有 Muon 变体持续提升训练稳定性。值得注意的是,在现有 Muon 变体发生发散的设置中,Soft Musec 仍保持稳定,同时在调优良好的配置下能达到与它们相当的性能。
cs.LG / 60 / 2609.11656

Learnware and AI Model Management System

学件与AI模型管理系统
Zhou, Zhi-Hua
Abstract
The transition from file storage to database management systems transformed stored data into managed resources. AI now faces an analogous transition from AI model storage to AI model management. Existing model pools essentially serve as \textit{AI model storage systems}. What is needed instead are \textit{AI model management systems} that enable models trained by different developers, for different tasks, with different data, and under different objectives to be identified, reused, and even assembled to address future user tasks. Because AI model developers are generally unwilling to share their training data, such systems should operate without accessing the training data of model developers and, ideally, without accessing raw data of future users. This requirement poses a fundamental challenge: the functionality of a modern AI model may not be fully understood even by the developer who trained it. How, then, can a system identify which models are useful for a given user task, let alone assemble models developed independently for different purposes? At first glance, this objective may appear unattainable. It becomes possible, however, by upgrading the basic unit of management from a machine learning model to a \textit{learnware}. \textit{Learnware = Model + Specification}. The specification, whose assignment transforms a trained model into a learnware, is generated with the help of a machine learning process without disclosing the training data of the developer and has a theoretically established data-preservation property. The \textit{Learnware Dock System (LDS)} provides a path toward powerful AI model management systems. Because specifications are generated according to a published reference and are comparable across models, they can also serve as an AI model \textit{collaboration protocol} through which independently developed models, including intelligent agents, can collaborate.
Chinese Translation
从文件存储到数据库管理系统的转变将存储数据转化为可管理的资源。AI现在面临着类似的转变,即从AI模型存储到AI模型管理。现有的模型池本质上充当着AI模型存储系统。相反,需要的是AI模型管理系统,它能够使由不同开发者、针对不同任务、使用不同数据、在不同目标下训练的模型被识别、重用,甚至组装起来以应对未来的用户任务。由于AI模型开发者通常不愿意共享他们的训练数据,此类系统应在不访问模型开发者训练数据的情况下运行,并且理想情况下也不访问未来用户的原始数据。这一要求提出了一个根本性挑战:现代AI模型的功能甚至可能无法被训练它的开发者完全理解。那么,系统如何能够识别哪些模型对给定的用户任务有用,更不用说组装为不同目的独立开发的模型了?乍一看,这一目标似乎无法实现。然而,通过将管理的基本单元从机器学习模型升级为学件(learnware),它变得可能。学件 = 模型 + 规约。规约的分配将训练好的模型转变为学件,它是在不披露开发者训练数据的情况下借助机器学习过程生成的,并且具有理论上确立的数据保全特性。学件码头系统(Learnware Dock System, LDS)为强大的AI模型管理系统提供了一条路径。由于规约是根据已发布的参考生成的,并且在不同模型之间具有可比性,它们还可以作为AI模型协作协议,通过该协议,独立开发的模型(包括智能体)可以进行协作。
cs.LG / 61 / 2609.11716

Why Does Post-Training Quantization Work?

为什么训练后量化有效?
Chen, Yuxiang, Beyer, Michael, Zhu, Jun, Chen, Jianfei
Abstract
Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and corrupt next-token prediction; randomly initialized models accumulate these discrepancies rapidly, whereas quantized pretrained models accumulate much less hidden-state error and largely maintain downstream task performance, even though they were never trained with quantization noise. This raises the question we address: why does post-training quantization work? Comparing full-precision and quantized forward passes, we identify two mechanisms that characterize pretrained quantization robustness. First, the error a layer newly introduces tends to oppose the error it inherits from the layer's input. The two cancel partially such that the discrepancy between full-precision and quantized passes grows slowly. This counteracting residual interaction develops during pretraining. Our quantitative analysis identifies it as a major factor slowing hidden-error growth. Second, LM-head geometry preferentially preserves the scores and probabilities of high-ranked tokens, which typically represent the model's most confident predictions. Together, these mechanisms explain why quantization error that passes through numerous layers can still produce only small output changes, and we verify the findings across models and quantization settings.
Chinese Translation
训练后量化通过以降低的精度存储大语言模型(LLMs)的权重来压缩它们,并且每个量化权重都会在隐藏状态中引入误差。直观上,这些误差应随深度累积并破坏下一词元预测;随机初始化的模型会迅速累积这些差异,而量化后的预训练模型累积的隐藏状态误差要少得多,并且基本保持了下游任务性能,尽管它们从未在量化噪声下训练过。这引出了我们要解决的问题:为什么训练后量化有效?通过比较全精度和量化前向传播,我们识别出两个刻画预训练量化鲁棒性的机制。首先,一层新引入的误差倾向于与它从该层输入继承的误差相反。两者部分抵消,使得全精度和量化前向传播之间的差异增长缓慢。这种抵消性的残差交互在预训练期间形成。我们的定量分析将其确定为减缓隐藏误差增长的主要因素。其次,LM头的几何结构优先保留高分词元的分数和概率,这些词元通常代表模型最自信的预测。这些机制共同解释了为什么经过众多层的量化误差仍然只会产生很小的输出变化,并且我们在不同的模型和量化设置下验证了这些发现。
cs.LG / 62 / 2609.11780

Predicting Privacy Leakage from Weight Spectral Density

从权重谱密度预测隐私泄露
Preen, Richard J., Smith, Jim
Abstract
Membership inference attacks (MIAs) are widely used to audit the privacy disclosure risk of machine learning models, however current state-of-the-art attacks require training computationally expensive shadow models, making large-scale privacy evaluation impractical. In this work, we investigate whether inexpensive spectral metrics derived from the heavy-tailed self-regularisation framework can serve as proxies for MIA vulnerability. We evaluate several WeightWatcher spectral metrics on image and tabular classification tasks and compare their relationship with MIA privacy leakage against conventional measures of generalisation. Across datasets, stable rank exhibits a strong positive correlation with overall MIA success, while Log alpha-Norm shows a consistent negative correlation with MIA vulnerability at the low false-positive regime. These associations are observed to be stronger than those obtained using the generalisation gap. The results indicate that neural network spectra may contain information about privacy leakage that is not fully captured by conventional measures of overfitting, motivating spectral analysis as a promising direction for scalable privacy auditing.
Chinese Translation
成员推理攻击(MIAs)被广泛用于审计机器学习模型的隐私泄露风险,然而当前最先进的攻击需要训练计算成本高昂的影子模型,使得大规模隐私评估不切实际。在本工作中,我们研究来自重尾自正则化框架的低成本谱度量是否可以作为MIA脆弱性的代理。我们在图像和表格分类任务上评估了多个WeightWatcher谱度量,并将它们与MIA隐私泄露的关系同传统的泛化度量进行了比较。在多个数据集上,稳定秩与整体MIA成功率呈强正相关,而对数alpha-范数在低假阳性率区间内与MIA脆弱性呈一致的负相关。这些关联比使用泛化差距获得的关联更强。结果表明,神经网络谱可能包含关于隐私泄露的信息,这些信息未被传统的过拟合度量完全捕获,从而推动谱分析成为可扩展隐私审计的一个有前景的方向。
cs.LG / 63 / 2609.11790

Dynamic language model representations for multi-objective reaction optimisation

用于多目标反应优化的动态语言模型表示
Sin, Joshua W., Segura, David Ming, Ranković, Bojana, Chau, Siu Lun, Lutz, Marius D. R., Anelli, Andrea, Burwood, Ryan P., Püntener, Kurt, Notheis, Maximilian J., Bigler, Raphael, Schwaller, Philippe
Abstract
Optimising chemical reactions across multiple objectives, such as yield, selectivity, and safety, is central to chemical synthesis, and model-driven approaches depend critically on how reaction components are represented. Established featurisations are either chemically uninformative, as with one-hot encodings, or, as with molecular descriptors, do not readily extend across chemically distinct components. For structurally and functionally diverse components, it is therefore unclear what a shared representation should contain. Constructing such a representation is itself a challenging research undertaking that must be revisited for each new reaction system. Here we bypass this step by learning the reaction representation dynamically from text. Textual descriptions of reaction conditions are encoded by a fine-tuned language model trained jointly with Gaussian process surrogates, yielding task-adaptive representations within a multi-objective Bayesian optimisation loop. Across nickel- and palladium-catalysed cross-couplings in both sequential and parallel experimentation regimes, this approach reaches optimisation convergence in fewer experiments than descriptor libraries or one-hot encoding. Applied prospectively to a palladium-catalysed cyanation spanning mixed ligand denticity and heterogeneous additives, and to a three-objective asymmetric hydrogenation across chiral iridium and ruthenium catalyst families, two rounds of high-throughput experimentation (192 reactions, under 3% of each design space) delivered conditions translating directly to gram scale in 94% and 84% isolated yield, the latter at 99.6% enantiomeric excess.
Chinese Translation
跨多个目标(如产率、选择性和安全性)优化化学反应是化学合成的核心,而模型驱动方法在很大程度上取决于反应组分如何表示。已有的特征化方法要么在化学上无信息量(如独热编码),要么(如分子描述符)难以扩展到化学上不同的组分。因此,对于结构和功能多样化的组分,一个共享表示应该包含什么并不清楚。构建这样的表示本身就是一项具有挑战性的研究工作,必须针对每个新的反应体系重新进行。在这里,我们通过从文本中动态学习反应表示来绕过这一步骤。反应条件的文本描述由与高斯过程代理模型联合训练的微调语言模型编码,从而在多目标贝叶斯优化循环中产生任务自适应的表示。在顺序和并行实验模式下,针对镍和钯催化的交叉偶联反应,该方法比描述符库或独热编码用更少的实验达到优化收敛。前瞻性地应用于跨越混合配体齿合度和非均相添加剂的钯催化氰化反应,以及跨越手性铱和钌催化剂家族的三目标不对称氢化反应,两轮高通量实验(192 个反应,每个设计空间的不到 3%)得到的条件可直接放大到克级,分离产率分别为 94% 和 84%,后者对映体过量 99.6%。
cs.LG / 64 / 2609.11801

Thinking with Looped Flows

用循环流(Looped Flows)思考
Suleymanzade, Ayhan, Lee, Chanhyuk, Eijkelboom, Floor, Boffi, Nicholas M., Ceylan, İsmail İlkan, Kim, Jinwoo
Abstract
Humans and machines often solve harder problems by spending more time on computation. In deep learning, looped models implement this idea during inference by recurrently updating a hidden state. In practice, however, their training backpropagates through only one or a few updates, making it hard to train early updates to support future ones. We propose looped flows, an approach that sidesteps this issue by training the recurrence with local denoising objectives. By imposing temporal association across denoising objectives through progressively decreasing noise levels and shared noise, the model is incentivized to learn recurrent states that transfer useful computation over time, even when gradients cover only a few updates. We then formulate inference as integrating the velocity of a probability flow parameterized by the learned denoiser, coupled with recurrent states. This allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples. Across six reasoning benchmarks including two multi-solution benchmarks, looped flows outperform prior state-of-the-art looped models overall, achieving 58.8% test accuracy on ARC-AGI-1 and 12.2% on ARC-AGI-2.
Chinese Translation
人类和机器常常通过花费更多计算时间来求解更难的问题。在深度学习中,循环模型(looped models)在推理时通过循环更新隐藏状态来实现这一思想。然而在实践中,它们的训练仅通过一步或少数几步更新进行反向传播,这使得难以训练早期更新以支持后续更新。我们提出循环流(looped flows),一种通过局部去噪目标训练循环来绕开这一问题的方法。通过渐进递减的噪声水平和共享噪声,在不同去噪目标之间建立时间关联,模型被促使学习能够随时间传递有用计算的循环状态,即使梯度仅覆盖少数几次更新。随后,我们将推理形式化为对由所学去噪器参数化的概率流速度进行积分,并与循环状态耦合。这允许通过更精细的时间网格花费更多计算来求解更难的问题,并能够从不同的初始噪声样本中得到多个有效预测。在包括两个多解基准在内的六个推理基准上,循环流总体优于此前最先进的循环模型,在 ARC-AGI-1 上达到 58.8% 的测试准确率,在 ARC-AGI-2 上达到 12.2%。
cs.LG / 65 / 2609.11842

Model-Aware Schedules Improve Generation via Fiberwise Optimal Transport

模型感知调度通过纤维式最优传输改进生成
Jia, Luyi, Zhang, Boyan, Liu, Yilun, Rulands, Steffen
Abstract
Diffusion and flow-matching schedules control the signal and noise coefficients that mix data and noise along affine probability paths. Minimizing a kinetic action defined on coefficient paths, motivated by optimal transport, helps explain strong baselines but remains model-agnostic and ignores prediction error. Here we introduce a model-aware schedule construction based on fiberwise optimal transport. At a fixed time and state on the probability path, compatible signal/noise decompositions form an affine fiber. We define a fiberwise prediction risk by averaging optimal-transport costs between the true and predictor-induced decompositions within these fibers. On a fixed coefficient curve, combining this risk with coefficient-path kinetic action yields a closed-form optimal time allocation. This construction extends to general linear prediction targets, and the risk profile can be estimated from an early baseline checkpoint. We evaluate DDPMs and flow matching across prediction targets, training configurations, risk-estimation checkpoints, datasets, and architectures. Our model-aware schedules consistently outperform strong baselines, including a 38.6% relative FID reduction for flow matching on CIFAR-10 at 16 function evaluations. Each model-agnostic kinetic baseline determines its own kinetic reference coordinate. In these coordinates, fiberwise-risk profiles from independently trained models in different settings align closely after normalization to unit area. The resulting schedule deformations used in training also align, suggesting empirical universality across the evaluated models and settings. Pretrained-checkpoint diagnostics extend this normalized-risk agreement to larger conditional latent diffusion and 2-RF models. A frozen analytic allocation template retains most of the model-aware improvement without further risk estimation or model-specific fitting.
Chinese Translation
扩散和流匹配调度控制信号和噪声系数,这些系数沿着仿射概率路径混合数据和噪声。受最优传输启发,最小化定义在系数路径上的动能作用有助于解释强基线,但仍然与模型无关且忽略了预测误差。在这里,我们引入一种基于纤维式最优传输的模型感知调度构建方法。在概率路径上的固定时间和状态,兼容的信号/噪声分解形成一个仿射纤维。我们通过在这些纤维内平均真实分解和预测器诱导分解之间的最优传输成本,定义了纤维式预测风险。在固定系数曲线上,将该风险与系数路径动能作用相结合,可得到闭式最优时间分配。这种构建可扩展到一般线性预测目标,并且风险分布可以从早期基线检查点估计。我们在预测目标、训练配置、风险估计检查点、数据集和架构上评估了DDPM和流匹配。我们的模型感知调度始终优于强基线,包括在CIFAR-10上16次函数评估下流匹配的FID相对减少38.6%。每个与模型无关的动能基线确定自己的动能参考坐标。在这些坐标中,来自不同设置下独立训练模型的纤维式风险分布在归一化到单位面积后紧密对齐。训练中使用的所得调度变形也一致,表明在评估的模型和设置中具有经验普遍性。预训练检查点诊断将这种归一化风险一致性扩展到更大的条件潜在扩散和2-RF模型。冻结的解析分配模板保留了大部分模型感知改进,无需进一步的风险估计或模型特定拟合。
cs.LG / 66 / 2609.11867

AdamX: Cosine similarity meets gradient descent

AdamX:余弦相似度遇上梯度下降
Caldas, Francisco, Belo, Ruben, Soares, Cláudia
Abstract
We introduce AdamX, a first-order optimizer that incorporates cosine similarity as an adaptive mechanism for controlling update magnitudes. The proposed method is scalable, model-agnostic, and straightforward to integrate into existing training pipelines. We further introduce a variance rectification scheme that promotes smoother optimization during the early stages of training. Overall, we provide empirical evidence that AdamX achieves competitive convergence rates across a range of benchmark datasets and architectures. Performance is evaluated in terms of the number of epochs required to reach predefined performance thresholds under a fixed hyperparameter budget. Code and Experiments available at: https://github.com/FranciscoCaldas/adamX.
Chinese Translation
我们提出了 AdamX,一种一阶优化器,它将余弦相似度作为控制更新幅度的自适应机制。所提方法具有可扩展性、与模型无关,并且易于集成到现有训练流程中。我们进一步引入了一种方差校正方案,可在训练早期阶段促进更平滑的优化。总体而言,我们提供了经验证据,表明 AdamX 在一系列基准数据集和架构上实现了有竞争力的收敛速率。性能评估依据是在固定超参数预算下达到预定义性能阈值所需的 epoch 数。代码与实验见:https://github.com/FranciscoCaldas/adamX。
cs.LG / 67 / 2609.11873

The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement

人类建造的最后一个AI:迈向真正的递归自我改进
Duan, Yi, Liu, Ying, Tang, Zirui, Chen, Haodong, Zhou, Jun, Liu, Yumou, Xu, Bangrui, Wu, Yukai, Chen, Sidi, Zhou, Yuhan, Wang, Haoyu, Yu, Xiaoyou, Han, Shaokun, Zhu, Xuzhou, Zhou, Le, Lu, Bolin, Zhou, Wei, Liu, Jiachen, Fang, Nuozhou, Tian, Jiaxin, Chen, Ruoyu, Li, Yuxuan, Zuo, Kai, Zhang, Kaiyan, Qiu, Jiantao, He, Conghui, Li, Guoliang, Zhou, Bowen, Liu, Zhiyuan, Wen, Zhoufutu, Kang, Jihua, Zhou, Xuanhe, Wu, Fan
Abstract
Recursive self-improvement (RSI) enables AI systems to turn experience and feedback into persistent changes that improve both their capabilities and the process of future improvement. We first use the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, then introduce the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement. Next we examine RSI across scenarios (e.g., scientific discovery, embodied intelligence, software engineering), highlighting their distinct requirements and development speeds. Drawing on diverse industry practices and preliminary empirical evidence, we connect RSI research with practical systems and identify key challenges to achieving genuine RSI.
Chinese Translation
递归自我改进(RSI)使AI系统能够将经验和反馈转化为持久性变化,从而提升其能力以及未来改进的过程。我们首先使用Headroom-Closed Index(HCI)揭示现有LLMs的问题,然后介绍RSI概念及其发展路线图:从改进执行自主、改进策略自主、经验获取自主和环境适应自主,到递归元改进。接下来,我们考察RSI在不同场景(如科学发现、具身智能、软件工程)中的应用,强调它们各自不同的需求和发展速度。借鉴多样化的行业实践和初步实证证据,我们将RSI研究与实际系统联系起来,并指出实现真正RSI的关键挑战。
cs.LG / 68 / 2609.11884

CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search

CoRA-NAS:用于神经架构搜索的粗排序与锚点残差细化
Yang, Yifan, Wang, Zhaoyan, Gao, Zheng, Li, Xiaoyu, Jiang, Jiaojiao
Abstract
Zero-cost proxies rank architectures cheaply, but their reliability varies across search spaces. We introduce CoRA-NAS (COarse Ranking + Anchor-residual), a two-stage framework combining a static ranking prior with low-cost learning-curve refinement. CoRA-Rank aggregates capacity and structure-at-initialization proxies through an equal-weight log-rank consensus and a target-free consensus gate. CoRA-Refine samples anchors across this prior, extrapolates their early validation curves, and propagates a learned residual correction with an ExtraTrees model. The refinement uses approximately 1% of the cost of fully training the candidate set. Fully trained architecture-accuracy labels are not used to fit the ranker. One configuration is used across spaces, with space-specific architecture encodings. Across NAS-Bench-201, NAS-Bench-101, TransNAS-Bench-101, and NATS-SSS, CoRA-Refine achieves mean Spearman correlations of 0.946, 0.715, 0.786, and 0.894, respectively. Its worst-space correlation of 0.715 is the highest among the compared methods. On NAS-Bench-201/CIFAR-100, its selected architecture reaches 73.32% accuracy, near the reported ground-truth best of 73.37%. On the pure size space, refinement recovers the static prior's shortfall relative to parameter count, while remaining tied with the strongest capacity proxies within noise. The resulting framework combines cross-space ranking robustness with low-cost architecture selection.
Chinese Translation
零成本代理能够以低成本对架构进行排序,但其可靠性在不同搜索空间中存在差异。我们提出CoRA-NAS(COarse Ranking + Anchor-residual),一个两阶段框架,结合静态排序先验与低成本学习曲线细化。CoRA-Rank通过等权对数秩共识和无需目标的共识门控,聚合容量与初始化结构代理。CoRA-Refine在此先验上采样锚点,外推其早期验证曲线,并使用ExtraTrees模型传播学习到的残差修正。该细化过程仅需完全训练候选集约1%的成本。完全训练的架构-精度标签不用于拟合排序器。使用单一配置跨空间适用,但采用空间特定的架构编码。在NAS-Bench-201、NAS-Bench-101、TransNAS-Bench-101和NATS-SSS上,CoRA-Refine分别达到平均Spearman相关系数0.946、0.715、0.786和0.894。其最差空间相关系数为0.715,在对比方法中最高。在NAS-Bench-201/CIFAR-100上,其选出的架构达到73.32%的准确率,接近报告的真实最优值73.37%。在纯尺寸空间上,细化方法弥补了静态先验相对于参数数量的不足,同时在噪声范围内与最强容量代理持平。最终框架结合了跨空间排序鲁棒性与低成本架构选择。
cs.LG / 69 / 2609.11897

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

CausalArena:基础模型时代的因果发现基准测试
Li, Zi-Rong, Liu, Si-Yang, Wang, Tian-Zuo, Ye, Han-Jia
Abstract
Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a causal graph together with the mechanisms that generate data, yet existing studies differ substantially in graph families, mechanisms, and evaluation protocols. The emergence of causal discovery foundation models (CDFMs) further complicates evaluation: performance may reflect not only causal discovery ability, but also overlap between pretraining environments and test SCMs, making results on fixed synthetic benchmarks difficult to interpret. We introduce CausalArena, a unified and evolvable benchmark for causal discovery under a common protocol. Synthetic SCMs supply controlled breadth over structures and mechanisms; semantic operational SCMs provide human-auditable, semantically grounded environments beyond standard synthetic generators; and formula-grounded SCMs test discovery under explicit scientific mechanisms. Public real-world datasets provide an additional external-validity check. Experiments across classical, neural, and pretrained methods reveal substantial ranking shifts across SCM families and protocols, showing that strong performance in one benchmark regime does not reliably transfer to others. These results highlight benchmark diversity and pretraining--evaluation overlap as central challenges for evaluating causal discovery in the foundation model era.
Chinese Translation
因果发现旨在从数据中揭示因果结构,是科学推理和基于干预的决策的基础。其评估严重依赖结构因果模型(SCMs),SCMs 指定因果图以及生成数据的机制,然而现有研究在图族、机制和评估协议方面差异很大。因果发现基础模型(CDFMs)的出现进一步使评估复杂化:性能可能不仅反映因果发现能力,还反映预训练环境与测试 SCMs 之间的重叠,这使得固定合成基准上的结果难以解释。我们提出 CausalArena,一个在统一协议下用于因果发现的统一且可演化的基准。合成 SCMs 在结构和机制上提供受控的广度;语义操作型 SCMs 提供人类可审计、语义有据的环境,超越标准合成生成器;而公式锚定 SCMs 在明确的科学机制下测试发现。公开真实世界数据集提供额外的外部效度检验。跨越经典、神经网络和预训练方法的实验揭示了不同 SCM 家族和协议之间存在显著的排名变化,表明在一个基准场景中的强劲性能并不能可靠地迁移到其他场景。这些结果强调基准多样性和预训练-评估重叠是基础模型时代评估因果发现的核心挑战。
cs.LG / 70 / 2609.11904

TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription

TART:一个技巧感知的音频到指法谱吉他转写的模块化工具
Gupta, Akshaj, Park, Hwi Joo, Guzman, Andrea, Gowda, Shamak, Konduri, Samhita, Lian, Jiachen, Netzorg, Robin, Anumanchipalli, Gopala
Abstract
Automatic Music Transcription (AMT) for guitar remains limited by three challenges: existing systems often fail to capture expressive techniques such as slides, bends, and percussive hits; they often assign notes to incorrect string-fret combinations; and they are typically trained on clean recordings, limiting their generalization to noisy real-world audio. To address these challenges, we propose TART, a modular four-stage audio-to-tablature pipeline consisting of (1) an audio-to-MIDI transcription model, (2) an expressive technique classifier, (3) an audio-conditioned T5 encoder-decoder for string-fret assignment, and (4) an automated tablature generator. We evaluate TART in a zero-shot setting on GuitarSet, EGDB, and two augmented benchmarks, Noisy GuitarSet and Noisy EGDB. Averaged across these four benchmarks, TART achieves 81.35% audio-to-MIDI F50 (+6.67 points over the best prior baseline), 71.8% string-fret Tab F1 (+8.5 points over the best prior baseline), and 54.08% end-to-end Tab F1. To our knowledge, TART is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio.
Chinese Translation
吉他自动音乐转写(AMT)仍受限于三个挑战:现有系统往往无法捕捉滑音、推弦和打击音等表现性技巧;它们经常将音符分配到错误的弦-品组合;并且它们通常在干净的录音上训练,限制了它们对嘈杂的真实世界音频的泛化能力。为了解决这些挑战,我们提出了TART,一个模块化的四阶段音频到指法谱流水线,包括(1)音频到MIDI的转写模型,(2)表现性技巧分类器,(3)用于弦-品分配的音频条件T5编码器-解码器,以及(4)自动指法谱生成器。我们在零样本设置下,在GuitarSet、EGDB以及两个增强基准(Noisy GuitarSet和Noisy EGDB)上评估了TART。在这四个基准上平均,TART达到了81.35%的音频到MIDI F50(比之前最佳基线高出6.67个点),71.8%的弦-品Tab F1(比之前最佳基线高出8.5个点),以及54.08%的端到端Tab F1。据我们所知,TART是第一个直接从吉他音频生成带有指法和表现性技巧标注的吉他指法谱的框架。
cs.LG / 71 / 2609.11910

From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good

从协议到证据:为共同利益服务的AI的有限主张
Chawla, Nitesh V., Benanti, Paulo
Abstract
Artificial Intelligence does more than create a governance problem. It can also reveal where institutions have already failed to provide responsiveness, belonging, care, and accountability. Once deployed, AI becomes an intervention in those conditions. It can repair, compound, substitute for, or conceal the failures it encounters. Responsible AI must therefore evaluate both the system and the institutional rupture into which it is introduced. The move from principles to protocols is already underway. The EU AI Act, NIST AI RMF, ISO/IEC 42001, and assurance practices translate commitments into roles, requirements, records, oversight, and assessment. The harder questions are what these protocols actually establish, whose power they leave untouched, and where measurement must stop. Pope Leo XIV's Magnifica Humanitas provides a broader moral frame centered on dignity, technological power, and the common good. Drawing on that frame, we develop a rupture test that links institutional baselines to system evaluation. We distinguish evidence-bounded deployment, which limits claims to what has actually been evaluated, from measurement-bounded governance, which records constraints that favorable evidence cannot override. Within those limits, RISE AI provides an architecture for making bounded, evidence-based claims about Responsibility, Inclusivity, Safety, and Empowerment. Responsible AI requires better engineering, institutional repair, and continued moral and political judgment.
Chinese Translation
人工智能不只是制造了一个治理问题。它还能揭示机构在提供响应性、归属感、关怀和问责方面已经失败的地方。一旦部署,人工智能就成为对这些状况的一种干预。它可以修复、加剧、替代或掩盖它所遇到的失败。因此,负责任的人工智能必须既评估系统本身,也评估其被引入的制度断裂。从原则到协议的转变已经在进行中。欧盟《人工智能法案》、NIST AI RMF、ISO/IEC 42001以及保证实践将承诺转化为角色、要求、记录、监督和评估。更难的问题是:这些协议实际确立了哪些内容、它们未触动谁的权力,以及测量必须在哪里停止。教皇利奥十四世的《Magnifica Humanitas》提供了一个更广泛的道德框架,以尊严、技术权力和共同利益为中心。借鉴该框架,我们开发了一种断裂测试,将制度基线与系统评估联系起来。我们区分证据限定型部署(evidence-bounded deployment)与测量限定型治理(measurement-bounded governance):前者将主张限制在实际已评估的内容范围内,后者记录有利证据也无法推翻的约束条件。在这些限度内,RISE AI 提供了一种架构,用于就责任、包容性、安全性和赋权(Responsibility, Inclusivity, Safety, and Empowerment)提出有界的、基于证据的主张。负责任的人工智能需要更好的工程、制度修复以及持续的道德与政治判断。
cs.LG / 72 / 2609.11917

Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data

数据稀缺与模型稀疏性:混合专家模型对重复数据过拟合更严重
Jha, Atindra, Li, Margaret, Leskovec, Jure, Liang, Percy, Zettlemoyer, Luke
Abstract
As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mixes, and across MoE settings, including expert count and granularity. We consistently find, for models ranging from 80M to 1B active (8.5B total) parameters, that MoEs degrade more rapidly under data repetition. This effect increases with sparsity, dictated by total rather than active parameters. While 80M dense models can repeat data over 8x with minimal degradation, MoEs instead begin to suffer at 4x, and deteriorate rapidly, ceding their performance benefits in all-unique data settings to underperform dense models after 32x. We experiment with existing regularization methods as a potential remedy. We find that some methods, such as dropout, can mitigate overfitting. In particular, with strong masking-based regularization, MoEs are able to outperform dense models even when data is repeated more than 64 times. However, no method fully matches the performance of all-unique training data. Finally, we analyze internal mechanisms correlated with MoE overfitting in high repetition regimes, and find that MoE routing universally stabilizes early in training, and that expert specialization correlates with overfitting to repeated data. In sum, our work addresses the adverse interactions between sparsity and data repetition: we present evidence for the core mechanisms of overfitting and its potential remediation, and suggest promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.
Chinese Translation
随着人类撰写的文本资源枯竭,重复使用语言模型训练数据已成为标准做法。先前的工作已研究了密集激活Transformer中的数据重复问题,但对于近期占主导地位的稀疏架构(如混合专家模型(MoE)),尽管其计算效率有所提高,数据重复的影响在很大程度上仍未得到探索。我们在单领域和多领域数据混合中,以及在不同的MoE设置(包括专家数量和粒度)中,改变了数据重复率。我们一致发现,对于从8000万到10亿活跃参数(总计85亿)的模型,MoE在数据重复下退化得更快。这种效应随着稀疏性的增加而加剧,且由总参数而非活跃参数决定。虽然8000万参数的密集模型可以在数据重复超过8倍时仅出现最小退化,但MoE在4倍时就开始受损,并迅速恶化,使其在全唯一数据设置中的性能优势丧失,在32倍后表现不及密集模型。我们尝试了现有的正则化方法作为潜在的补救措施。我们发现一些方法,如dropout,可以缓解过拟合。特别是,在强掩码正则化下,即使数据重复超过64次,MoE也能超越密集模型。然而,没有一种方法能完全匹敌全唯一训练数据的性能。最后,我们分析了在高重复率情况下与MoE过拟合相关的内部机制,发现MoE路由在训练早期普遍稳定,并且专家专业化与对重复数据的过拟合相关。总之,我们的工作解决了稀疏性与数据重复之间的不良相互作用:我们提供了过拟合核心机制及其潜在补救措施的证据,并提出了未来方法通过打破记忆模式来减少模型参数过度专业化的有前景的途径。
cs.LG / 73 / 2609.11918

General Quantification of Covariate and Concept Shifts

协变量偏移与概念偏移的一般量化
Chen, Hongbo, Xia, Li Charlie
Abstract
Generalization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-estimable from samples. In this paper, we bridge the gap between theory and practical applications. We first show that existing definition of concept shift breaks when the source and target supports mismatch. Leveraging entropic optimal transport, we propose a key notion: $\gamma^{*}\!$-concept shifts, and derive a general error bound unifying covariate and $\gamma^{*}\!$-concept shifts, which applies to broad loss functions, label spaces, and stochastic labeling. We further develop estimators for these shifts with concentration guarantees, and the DataShifts algorithm, which can quantify distribution shifts and estimate the error bound in most applications - a rigorous and general tool for analyzing learning error under distribution shift.
Chinese Translation
分布偏移下的泛化仍然是现代机器学习的一个核心挑战,然而现有的学习界理论局限于狭窄、理想化的设定,且无法从样本中估计。在本文中,我们弥合了理论与实际应用之间的鸿沟。我们首先表明,当源域和目标域的支撑不匹配时,现有的概念偏移定义会失效。利用熵最优传输,我们提出了一个关键概念:$\gamma^{*}$-概念偏移,并推导出一个统一协变量偏移和 $\gamma^{*}$-概念偏移的一般误差界,该界适用于广泛的损失函数、标签空间和随机标签。我们进一步开发了这些偏移的估计器(具有集中保证)以及 DataShifts 算法,该算法可以在大多数应用中量化分布偏移并估计误差界——一个用于分析分布偏移下学习误差的严谨且通用的工具。
机器人学 (Robotics)
38
cs.RO / 1 / 2609.10706

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

HuRo:将人类视频机器人化以实现可扩展的 VLA 预训练
Jeong, Jinho, Joo, Se June, Kang, Jaehyun, Kim, Dongyun, Kim, Yena, Kim, Hanjung, Kim, Seon Joo
Abstract
Human video datasets have emerged as a compelling alternative to expensive real-robot data, offering rich diversity at scale. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation and action alignment separately at scale. In this work, we systematically examine whether robotized human videos can provide effective and scalable supervision for pretraining vision-language-action (VLA) policies. To this end, we develop a robotization pipeline that converts heterogeneous human videos into robot-aligned observations and action trajectories while inferring missing intermediate signals across annotation levels. Using this pipeline, we construct the HuRo dataset, comprising about 630K robotized episodes and 142M processed frames from five human-video sources. Across four real-world manipulation tasks, increasing robotized pretraining scale improves overall completion from 51.5% to 80.3% and OOD completion under spatial and visual shifts from 34.9% to 72.2%. Ablations further show that visual robotization improves OOD robustness and that end-to-end pretraining with retargeted actions outperforms visual-only transfer. Code and data are released on our website: https://3587jjh.github.io/HuRo.
Chinese Translation
人类视频数据集已成为昂贵真实机器人数据的有吸引力的替代方案,能够大规模提供丰富多样性。为弥合人类到机器人的具身差距,现有方法要么在任务匹配的设置中将视频机器人化,要么大规模分别处理观测和动作对齐。在本工作中,我们系统研究了机器人化的人类视频能否为预训练视觉-语言-动作(VLA)策略提供有效且可扩展的监督。为此,我们开发了一个机器人化流水线,将异构人类视频转换为机器人对齐的观测和动作轨迹,同时推断跨标注层级缺失的中间信号。利用该流水线,我们构建了 HuRo 数据集,包含约 63 万个机器人化片段和 1.42 亿处理帧,来自五个人类视频来源。在四个真实世界操作任务中,增加机器人化预训练规模将整体完成率从 51.5% 提升至 80.3%,并将空间和视觉偏移下的分布外(OOD)完成率从 34.9% 提升至 72.2%。消融实验进一步表明,视觉机器人化提高了 OOD 鲁棒性,并且使用重定向动作的端到端预训练优于仅视觉迁移。代码和数据已发布在我们的网站:https://3587jjh.github.io/HuRo。
cs.RO / 2 / 2609.10726

When Information is Worth the Risk: Behavioral Valuation for Hazardous Robotic Exploration

当信息值得冒险时:危险机器人探索的行为估值
Srivastava, Alkesh K., Suresh, Aamodh, Nieto-Granda, Carlos, Dames, Philip
Abstract
Hazardous robotic exploration requires robots to map spatial risks, such as unsafe terrain, radiation, fire, mines, or structural damage, while operating where collecting information can itself cause failure. A highly informative path may expose the robot to hazards, terminate execution, and prevent future observations. Hazardous exploration therefore requires deciding not only where uncertainty is largest, but when reducing it is worth the risk. This paper introduces a valuation-layer view of this problem. We keep the belief update, sensor model, physical risk model, and finite-horizon informative planner fixed, and change only the scalar objective used to rank feasible paths. Within this framework, we introduce a risk-augmented Behavioral Information objective based on Prelec probability weighting, yielding an interpretable family of conservative-to-aggressive information-risk valuations. Theoretically, we show that valuation parameters create switching boundaries between high-information/high-risk and lower-information/lower-risk paths, and induce a transformed Pareto-frontier structure over feasible exploration policies. Large-scale failure-truncated grid-world experiments show that valuation alone reshapes the information-risk frontier. Shannon information planning remains a strong raw-information baseline, while risk-aware objectives can reduce hazard exposure and robot losses by avoiding failures that truncate future sensing. Risk-augmented Behavioral valuation is Pareto-competitive with standard risk-aware baselines and provides interpretable conservative and intermediate regimes. These results support a framework in which robots reason not only about how much uncertainty an action reduces, but whether that reduction is worth the risk required to obtain it.
Chinese Translation
危险机器人探索要求机器人绘制空间风险图,例如不安全地形、辐射、火灾、地雷或结构损伤,同时在收集信息本身可能导致失败的环境中运行。一条信息量很高的路径可能使机器人暴露于危险之中,终止任务执行,并阻止未来的观测。因此,危险探索不仅需要决定不确定性最大的地方,还需要决定何时降低不确定性值得冒险。本文引入了该问题的估值层视角。我们保持信念更新、传感器模型、物理风险模型和有限时域信息规划器不变,仅改变用于对可行路径排序的标量目标函数。在此框架内,我们引入一种基于Prelec概率加权的风险增强行为信息目标,产生可解释的从保守到激进的信息-风险估值族。在理论上,我们表明估值参数在高信息/高风险路径与低信息/低风险路径之间产生切换边界,并在可行探索策略上诱导出变换后的帕累托前沿结构。大规模失败截断的网格世界实验表明,仅估值本身就能重塑信息-风险前沿。香农信息规划仍然是一个强大的原始信息基线,而风险感知目标可以通过避免截断未来感知的失败来减少危险暴露和机器人损失。风险增强行为估值与标准风险感知基线具有帕累托竞争力,并提供可解释的保守和中间机制。这些结果支持了一个框架,其中机器人不仅推理一个动作能减少多少不确定性,还推理这种减少是否值得为获得它所冒的风险。
cs.RO / 3 / 2609.10748

Lie-Algebraic Bell Recurrences for Arbitrary-Order Twist Jets and Parallel-Mechanism Closure

李代数贝尔递推用于任意阶旋量射流与并联机构闭链
Condurache, Daniel
Abstract
This paper develops an arbitrary-order kinematic construction that links serial propagation, parallel-mechanism closure, and rigid-platform point fields within one dual screw framework. A cylindrical joint is retained as one native physical block, with revolute and prismatic joints obtained as special cases. For each fixed joint axis, ordinary Bell polynomials organize the derivatives of the exponential factor; across a chain, the noncommuting factors remain in their physical order. Initial-frame prefix and terminal-resolved covariant formulas then produce equivalent representations of the serial twist jet. For a parallel mechanism, repeated Leibniz differentiation, with joint-level derivatives organized by Bell polynomials, yields an arbitrary-order triangular active-passive closure recurrence: the same passive Jacobian is solved at every derivative order at a regular configuration, while the right-hand side contains only prescribed active data and lower-order jets. The resulting platform twist jet is mapped exactly to the point-independent affine invariants of the velocity, acceleration, jerk, and snap fields. The validation is deliberately complementary: a generic 3C chain with noncoplanar axes and nonzero rotational and translational cylindrical coordinates tests ordered serial propagation, an RR+RRR spherical wrist tests active-passive closure, and a Hunt-type 6-RUS mechanism with six active revolute joints tests an independently reconstructed platform jet and its affine fields. Independent differentiation of the rigid motion, evaluation of the affine fields, and the differentiated branch closures all agree through fourth order with residuals below $10^{-12}$ in the corresponding SI units. The formulation is purely kinematic and applies at configurations where the selected active-passive partition is regular.
Chinese Translation
本文开发了一种任意阶运动学构造,该构造在一个对偶螺旋框架内将串联传播、并联机构闭链和刚性平台点场联系起来。圆柱关节被保留为一个原生物理模块,转动关节和移动关节作为特殊情况获得。对于每个固定关节轴,普通贝尔多项式组织指数因子的导数;在链中,非交换因子保持其物理顺序。初始坐标系前缀和终端解析的协变公式随后产生串联旋量射流的等价表示。对于并联机构,重复莱布尼茨微分,其中关节级导数由贝尔多项式组织,产生一个任意阶三角形主动-被动闭链递推:在规则构型下,对每个导数阶求解相同的被动雅可比矩阵,而右侧仅包含规定的主动数据和低阶射流。得到的平台旋量射流被精确映射到速度、加速度、急动度和急跳度场的点无关仿射不变量。验证是有意互补的:一个具有非共面轴和非零旋转和平移圆柱坐标的通用3C链测试有序串联传播,一个RR+RRR球腕测试主动-被动闭链,一个具有六个主动转动关节的Hunt型6-RUS机构测试独立重建的平台射流及其仿射场。刚体运动的独立微分、仿射场的评估以及微分分支闭链均一致到四阶,在相应的SI单位下残差低于$10^{-12}$。该表述是纯运动学的,适用于所选主动-被动划分规则的构型。
cs.RO / 4 / 2609.10844

Expressive Robotic Pianist: Mastering Complex Piano Repertoire with Graph-Mimic and Musical Dynamics

富有表现力的机器人钢琴家:以Graph-Mimic与音乐力度掌握复杂钢琴曲目
Liang, Yanhong, Liu, Xianwei, Fu, Chaojie, Cheng, Shaowen, Yuan, Yanyan, Zhuo, Chengwei, Chen, Xi, Jin, Yongbin, Yang, Wei, Wang, Hongtao
Abstract
Enabling robots to perform musical instruments with human-level expressivity represents a frontier in bridging the gap between mechanical precision and artistic interpretation. Despite advances in robotic dexterity, replicating the fluid finger transitions and nuanced dynamic control characteristic of human pianists remains a significant challenge. Through a reinforcement learning-based control framework, we demonstrate that a dexterous robotic hand can achieve high-fidelity performance across a diverse piano repertoire. Central to our approach is a graph-based optimization strategy that guides the robot to generate natural pre-press and key-press fingering strategies that closely resemble human movement patterns. To achieve expressive sound production, the control system is coupled with a physics-inspired acoustic model that modulates keypress velocity to accurately reproduce the dynamic variations specified in musical scores. Quantitative evaluations demonstrate that our expressive control model significantly outperforms baseline methods in both finger morphology similarity and dynamic velocity accuracy. In a perceptual test involving participants from diverse listener groups, performances generated by our system are significantly preferred over baseline robotic performances and are indistinguishable from human performances for non-professional audiences. Furthermore, extensive experiments across multiple musical styles confirm that our method maintains high note-level accuracy while achieving expressive performance. Our approach provides a robust pathway for robotic systems to move beyond mere mechanical accuracy, elevating robotic musicianship to a level of expressive performance comparable to human pianists.
Chinese Translation
使机器人能够以人类水平的表达力演奏乐器,是弥合机械精度与艺术诠释之间差距的前沿方向。尽管机器人灵巧性取得了进展,但复现人类钢琴家特有的流畅指法转换和细腻力度控制仍是一项重大挑战。通过基于强化学习的控制框架,我们证明灵巧的机器人手能够在多样的钢琴曲目中实现高保真演奏。我们方法的核心是一种基于图的优化策略,引导机器人生成与人类运动模式高度相似的自然预按压和按键指法策略。为实现富有表现力的声音生成,控制系统与一个受物理启发的声学模型相耦合,该模型调节按键速度,以准确复现乐谱中指定的力度变化。定量评估表明,我们的表达性控制模型在手指形态相似性和动态速度准确性两方面均显著优于基线方法。在一项包含不同听众群体参与者的感知测试中,我们的系统生成的演奏显著优于基线机器人演奏,并且对于非专业听众而言,与人类演奏无法区分。此外,跨多种音乐风格的大量实验证实,我们的方法在实现富有表现力演奏的同时,保持了高音符级准确率。我们的方法为机器人系统提供了一条稳健路径,使其超越单纯的机械精度,将机器人音乐才能提升到可与人类钢琴家媲美的表达性演奏水平。
cs.RO / 5 / 2609.10895

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

ReactHuman:面向具身多模态LLM中类人反应式决策的基于物理的基准
Li, Yizhan, You, Jianxin, Xiong, Mengyang, Chen, Yinhuan, Zhao, Zicheng, Wu, Dekun, Zhang, Dongqing, Liu, Bang
Abstract
Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as the decision coreof household robots. Existing evaluations, however, probe intuitive physics passively through question answering over videos, or target deliberate, long-horizon tasks such as navigation and rearrangement; none measure whether a model can turn physical understanding into immediate, safety-critical action. We introduce ReactHuman, the first physics-grounded benchmark for human-like reactive decision-making, in which the evaluated MLLM acts as the brain of a simulated humanoid facing sudden household hazards; it spans 17 event families and over 1,000 bit-for-bit reproducible scenes with exact, annotation-free ground truth derived from 240 Hz rigid-body simulation, including adversarial objects whose appearance contradicts their physics (a foam anvil, a steel apple). We further design a five-metric suite that scores each reaction along three axes: reasonable, safe, and physically grounded. We physically execute every committed plan so that decisions have observable consequences. With this harness we evaluate seven representative MLLMs. Results show that reactive safety is far from solved: models mishandle roughly one hazard in three, act from fixed dispositions rather than the observed scene, trust appearance over motion, and miss interception points at meter scale even when the chosen action is correct; none of these failures shrink with model scale. ReactHuman thus offers both a fine-grained diagnosis and a scalable training signal toward physically grounded, safety-aware embodied agents. The benchmark can be found here: https://huggingface.co/datasets/Alan123/reacthuman-benchmark-scaled
Chinese Translation
对突发物理危险(接住滑落的盘子、躲避掉落的刀)做出反应,既是对具身智能的有意义测试,也是将多模态大语言模型(MLLMs)部署为家用机器人决策核心的硬性要求。然而,现有的评估要么通过视频问答被动地探究直觉物理,要么针对导航和重排等深思熟虑的长时程任务;没有一个评估衡量模型能否将物理理解转化为即时的、安全关键的行动。我们介绍了ReactHuman,这是首个基于物理的类人反应式决策基准,其中被评估的MLLM充当模拟人形机器人的大脑,面对突发的家庭危险;它涵盖17个事件族和超过1,000个逐位可复现场景,具有精确、无需标注的真值,源自240 Hz刚体模拟,包括外观与物理特性相矛盾的对抗性物体(泡沫铁砧、钢苹果)。我们进一步设计了一个五指标套件,从三个维度对每个反应进行评分:合理、安全和物理基础。我们物理执行每一个已提交的计划,以便决策具有可观察的后果。利用这个测试框架,我们评估了七个代表性的MLLM。结果表明,反应式安全远未解决:模型大约每三个危险就有一个处理不当,基于固定倾向而非观察到的场景行动,信任外观而非运动,即使选择的动作正确,也会在米级尺度上错过拦截点;这些失败都没有随着模型规模的扩大而减少。因此,ReactHuman 为面向物理基础、安全感知的具身智能体提供了细粒度诊断和可扩展的训练信号。该基准可在此处找到:https://huggingface.co/datasets/Alan123/reacthuman-benchmark-scaled
cs.RO / 6 / 2609.10905

Planning along Differentiable Charts of Constraint Manifolds with General-Purpose IK Solvers

使用通用逆运动学求解器沿约束流形的可微图表进行规划
Cohn, Thomas, Shaw, Seiji, Biggie, Harel, Manderson, Travis, Roy, Nicholas, Tedrake, Russ
Abstract
Planning trajectories for robot manipulators under kinematic equality constraints restricts feasible motions to a measure-zero submanifold of the configuration space, requiring special algorithmic treatment. A promising strategy is parametrizing the set of feasible configurations using analytic inverse kinematics (IK). Bespoke analytic IK functions can be written to be differentiable, a necessary property for gradient-based trajectory optimization. But the vast majority of IK functions are computed by automated meta-solvers like IKFast, and are difficult to modify for differentiability. We present a new approach for computing gradients of analytic IK parameterizations: we leverage the inverse function theorem to recover the desired gradients from the ordinary forward kinematic Jacobian. Furthermore, we present a least-squares domain extension and an optimization-amenable description of the reachability constraint, which preserves gradient signal outside the reachable workspace. We demonstrate the efficacy of our approach through numerical experiments and downstream tasks, including a hardware demonstration of an RB-Y1 picking up a box and placing it on a table. Project website: https://cohnt.github.io/inverse-function-theorem-parameterization/
Chinese Translation
在运动学等式约束下规划机器人机械臂的轨迹,会将可行运动限制在构型空间的零测度子流形上,需要特殊的算法处理。一个有前景的策略是使用解析逆运动学(IK)参数化可行构型的集合。定制的解析IK函数可以写成可微的,这是基于梯度的轨迹优化的必要属性。但是绝大多数IK函数是由像IKFast这样的自动元求解器计算的,并且难以修改以实现可微性。我们提出了一种计算解析IK参数化梯度的新方法:我们利用反函数定理从普通前向运动学雅可比矩阵中恢复所需的梯度。此外,我们提出了一种最小二乘域扩展和一种适合优化的可达性约束描述,该描述在可达工作空间之外保留梯度信号。我们通过数值实验和下游任务证明了我们方法的有效性,包括一个RB-Y1拾取盒子并将其放在桌子上的硬件演示。项目网站:https://cohnt.github.io/inverse-function-theorem-parameterization/
cs.RO / 7 / 2609.10915

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

IMLE-VLA:面向视觉-语言-动作策略的快速单步动作生成
Hosseinkhani, Kian, Peng, Qinhe, Shramko, George, Aghabozorgi, Mehran, Qian, Jianing, Engst, Tristan, Moazeni, Alireza, Jayaraman, Dinesh, Li, Ke
Abstract
Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 10 Euler steps in $\pi_{0.5}$. This creates an inference bottleneck that produces stop-and-go movement in the robot and slower task completion. We introduce IMLE-VLA, which replaces the iterative action head with a single-step conditional generator trained via conditional Implicit Maximum Likelihood Estimation (cIMLE). The cIMLE objective promotes multimodal action coverage, avoiding the mode collapse of naive regression heads while eliminating multi-step sampling entirely. When IMLE-VLA is applied to $\pi_{0.5}$, it increases inference frequency 3.67x (55 Hz vs. 15 Hz), enabling up to 11x higher action throughput. On the 40-task LIBERO benchmark, IMLE-VLA achieves the highest average success rate (98.0%) among all baselines while leading in inference frequency. Under the test-time perturbations of LIBERO-plus, IMLE-VLA retains $\pi_{0.5}$'s robustness while other baselines degrade sharply, confirming that the cIMLE head preserves generalization. Real-world experiments on a Franka Emika Panda across four tasks demonstrate smoother motion (2.2x to 3.0x lower jerk) and faster task completion, with IMLE-VLA outperforming $\pi_{0.5}$ on every task and reducing average VLA inference time per episode by 3.9x to 6.6x. Videos and code are available at https://kianhk6.github.io/IMLE-VLA/
Chinese Translation
视觉-语言-动作(VLA)策略利用预训练的视觉-语言主干网络来实现强大的跨任务泛化。一种领先的设计将该主干与通过扩散或流匹配训练的专用连续动作头相结合。然而,此类动作头依赖迭代式多步采样,例如在 $\pi_{0.5}$ 中使用 10 个欧拉步。这造成了推理瓶颈,导致机器人出现走走停停的运动,并降低任务完成速度。我们提出 IMLE-VLA,它用通过条件隐式最大似然估计(cIMLE)训练的单步条件生成器取代迭代式动作头。cIMLE 目标促进多模态动作覆盖,避免朴素回归头的模式崩溃,同时完全消除多步采样。当将 IMLE-VLA 应用于 $\pi_{0.5}$ 时,其推理频率提升至 3.67 倍(55 Hz 对 15 Hz),使动作吞吐量最高提升 11 倍。在 40 项任务的 LIBERO 基准上,IMLE-VLA 在所有基线中取得了最高的平均成功率(98.0%),同时在推理频率上领先。在 LIBERO-plus 的测试时扰动下,IMLE-VLA 保持了 $\pi_{0.5}$ 的鲁棒性,而其他基线则急剧退化,这证实了 cIMLE 头保持了泛化能力。在 Franka Emika Panda 上跨四项任务的真实世界实验表明,运动更平滑(加加速度(jerk)降低 2.2 倍至 3.0 倍),任务完成更快,IMLE-VLA 在每个任务上都优于 $\pi_{0.5}$,并将每个回合的平均 VLA 推理时间减少 3.9 倍至 6.6 倍。视频和代码见 https://kianhk6.github.io/IMLE-VLA/
cs.RO / 8 / 2609.10918

ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations

ObstaDiff:基于障碍物感知表征的泛化扩散策略学习
Wang, Jiawen, Yao, Kevin, Jawed, Khalid
Abstract
Imitation learning has achieved impressive results in robotic manipulation, yet most existing approaches assume clean backgrounds and lack explicit mechanisms for obstacle-aware motion generation. Extending such policies to cluttered, real-world scenes with unstructured obstacles remains a key generalization challenge. We present ObstaDiff, a decomposed diffusion-policy framework with a lightweight obstacle-aware visual encoder. ObstaDiff extracts a structured target-obstacle-background representation, enabling the downstream alignment policy to generate end-effector trajectories toward a target-centered bottleneck pose while reasoning about surrounding obstacles. We evaluate ObstaDiff on 61 real-robot greenhouse trials per method (366 executions in total). ObstaDiff achieves 75.41% average task success and 8.20% average obstacle collision rate, outperforming representative imitation-learning baselines and improving generalization in cluttered agricultural scenes.
Chinese Translation
模仿学习在机器人操作中已取得显著成果,但现有方法大多假设背景干净,缺乏障碍物感知运动生成的显式机制。将此类策略扩展到具有非结构化障碍物的杂乱真实场景仍是一个关键的泛化挑战。我们提出ObstaDiff,一个带有轻量级障碍物感知视觉编码器的分解式扩散策略框架。ObstaDiff提取结构化的目标-障碍物-背景表征,使下游对齐策略能够生成朝向以目标为中心的瓶颈位姿的末端执行器轨迹,同时推理周围障碍物。我们在每种方法的61次真实机器人温室试验(共366次执行)上评估了ObstaDiff。ObstaDiff实现了75.41%的平均任务成功率和8.20%的平均障碍物碰撞率,优于代表性的模仿学习基线,并提高了在杂乱农业场景中的泛化能力。
cs.RO / 9 / 2609.10951

Testing Between the Test Cases: Proving End-to-End Steering in Conditions You Never Drove

在测试用例之间测试:在从未驾驶过的条件下证明端到端转向
Ghalan, Menuka, Rodgers, Charles, Asher, Zachary D.
Abstract
AI-based automated vehicle testing is challenging because a model that passes every test condition can still fail in the real world. Formal verification offers a way to directly address this gap. On a simulated highway and an arterial road we trained two small end-to-end steering networks each in CARLA, one on clear conditions alone and one on clear, fog, night and low sun. All four models were driven against a 2.19 ft lane-departure budget. Without driving again, we used bound propagation, a formal method that reads the trained weights, to compute how far steering can drift at every disturbance strength between two captured images. One calculation covers more than a campaign could drive: on the arterial it spans 133 poses, where ten intensities each would be 10^133 combinations, in minutes on one GPU. Not only did formal verification find conditions that broke the clear-trained policy without simulation testing, it provided some preliminary evidence for potential failures between the test cases. Our overall conclusion is that formal verification is a viable complement to simulation, and could be adopted as a part of verification and validation for automated driving.
Chinese Translation
基于AI的自动驾驶车辆测试具有挑战性,因为一个通过了所有测试条件的模型在现实世界中仍可能失败。形式化验证提供了一种直接解决这一差距的方法。在模拟的高速公路和一条城市干道上,我们在CARLA中分别训练了两个小型端到端转向网络,一个仅在晴朗条件下训练,另一个在晴朗、雾、夜晚和低日照条件下训练。所有四个模型都针对2.19英尺的车道偏离预算进行了驾驶测试。无需再次驾驶,我们使用了边界传播(bound propagation),一种读取训练权重的形式化方法,来计算在两个捕获图像之间的每个扰动强度下转向能漂移多远。一次计算覆盖的范围比一次测试活动所能驾驶的还要多:在城市干道上,它涵盖133个位姿,其中每个位姿10个强度级别将产生10^133种组合,在一台GPU上只需几分钟。形式化验证不仅在没有仿真测试的情况下发现了打破晴朗训练策略的条件,还为测试用例之间的潜在失败提供了一些初步证据。我们的总体结论是,形式化验证是仿真的可行补充,可以作为自动驾驶验证和确认的一部分被采用。
cs.RO / 10 / 2609.11043

LTLDiff: Finite Linear Temporal Logic-Guided Data Generation and Diffusion Policies for Multi-agent Robotic Manipulation

LTLDiff:有限线性时序逻辑引导的数据生成与扩散策略用于多智能体机器人操作
Meng, Chuhan, Yin, Haiyan
Abstract
Multi-agent robotic manipulation tasks require coordination among agents to satisfy task-level temporal, logical, and safety constraints. Recently, diffusion policies have been used to perform the task. However, they still suffer from desynchronization, incorrect action ordering, and coordination failures in tasks that require simultaneous or sequential multi-agent interaction. Therefore, LTLDiff is proposed as a framework that combines Finite Linear Temporal Logic (LTLf) specification learning for both the generation of demonstrations and learning via diffusion policies. Each task has a specific LTLf formula that is learned from a set of natural language instructions using a large-scale language model. To enable a fixed-dimensional vector embedding of the learned specification from the language model, LTLf uses an abstract syntax tree representation scheme. This embedding of logic serves as a condition for (i) logic-guided data collection and (ii) diffusion-based policy training, encouraging trajectories that are consistent with the desired ordering and coordination requirements. Experiments on multi-agent LTLDiff manipulation tasks demonstrate improved task success rates compared to the baseline. Together, these contributions demonstrate the effectiveness of LTLDiff for coordinated multi-agent manipulation.
Chinese Translation
多智能体机器人操作任务需要智能体之间的协调,以满足任务级的时间、逻辑和安全约束。最近,扩散策略已被用于执行此类任务。然而,在需要同时或顺序多智能体交互的任务中,它们仍然面临不同步、动作顺序错误以及协调失败的问题。因此,提出 LTLDiff 作为一个框架,它结合了有限线性时序逻辑(LTLf)规范学习,以同时用于演示生成和通过扩散策略进行学习。每个任务都有一个特定的 LTLf 公式,该公式使用大规模语言模型从一组自然语言指令中学习得到。为了使从语言模型学到的规范能够进行固定维度的向量嵌入,LTLf 采用抽象语法树表示方案。这种逻辑嵌入用作 (i) 逻辑引导的数据收集和 (ii) 基于扩散的策略训练的条件,鼓励轨迹符合所需的顺序和协调要求。在多智能体 LTLDiff 操作任务上的实验表明,与基线相比,任务成功率有所提高。总之,这些贡献证明了 LTLDiff 在协调多智能体操作中的有效性。
cs.RO / 11 / 2609.11059

Gait-Dependent Effects on Quadruped Locomotion for Load-Carrying using Passive Mechanism

使用被动机构的四足负载搬运运动步态依赖效应
Dessy, Giovanni B., Semini, Claudio, Barasuol, Victor
Abstract
Passive mechanical interfaces offer a lightweight alternative to actuated manipulators for quadruped payload carrying, but their impedance directly couples the payload dynamics with the locomotion pattern. This paper analyzes how passive-arm stiffness-damping selection affects payload-carrying locomotion under different gait and payload conditions. We compare damped and underdamped passive-arm impedance configurations in simulation during flat-ground locomotion. For crawl gaits, where the support polygon remains well defined, the results show that underdamped impedance increases passive-joint oscillations and can reduce the ZMP margin with respect to the support polygon. Trot is retained as a dynamic excitation case for the passive arm, but it is not used for direct ZMP-margin stability comparison. The results are summarized in gait-payload-stiffness-damping maps, where ZMP-margin reduction is evaluated for crawl gaits and trot is retained only as a passive-arm excitation case.
Chinese Translation
被动机械接口为四足负载搬运提供了一种轻量化的替代方案,可替代主动驱动机械臂,但其阻抗会将载荷动力学与运动模式直接耦合。本文分析了被动臂的刚度-阻尼选择如何影响不同步态和负载条件下的负载搬运运动。我们在仿真中比较了平坦地面运动过程中阻尼与欠阻尼被动臂阻抗配置。对于支撑多边形仍具有良好定义的爬行步态,结果表明欠阻尼阻抗会增加被动关节振荡,并可能降低相对于支撑多边形的ZMP裕度。Trot被保留作为被动臂的动态激励案例,但不用于直接的ZMP裕度稳定性比较。结果以步态-负载-刚度-阻尼图进行总结,其中针对爬行步态评估ZMP裕度降低,而Trot仅作为被动臂激励案例保留。
cs.RO / 12 / 2609.11078

Freehand Sketching for End-User Programming of Robot Swarms

面向机器人集群终端用户编程的手绘草图
Karam, Riwa, Kuo, Ian, Egerstedt, Magnus
Abstract
Robot swarms are increasingly used in applications where accessible interaction with non-expert users is desirable. This paper investigates freehand sketching as an end-user programming interface for specifying robot swarm geometries. Users communicate spatial intent through a drawing, while the swarm autonomously extracts target formation points, constructs a rigid formation graph, assigns robots to formation nodes, and executes distributed formation control with a guarantee against unintended reflected formations. The resulting sketch-to-swarm framework is evaluated through a human study examining the usability of freehand formation specification. Twenty participants generated $42$ geometric shapes, and the interface achieved a mean System Usability Scale score of $84.25$, which conventionally indicates high perceived usability. The results support freehand sketching as an intuitive interaction abstraction for human-swarm collaboration without requiring robotics or programming expertise.
Chinese Translation
机器人集群越来越多地应用于需要与非专家用户进行易用交互的场景。本文研究手绘草图作为一种终端用户编程接口,用于指定机器人集群的几何构型。用户通过绘图传达空间意图,而集群自主提取目标编队点、构建刚性编队图、将机器人分配至编队节点,并执行分布式编队控制,且保证不会产生非预期的镜像编队。由此得到的“草图到集群”框架通过一项人类研究进行评估,该研究考察手绘编队指定的可用性。20名参与者生成了42种几何形状,该接口的平均系统可用性量表(SUS)得分为84.25,按惯例表明感知可用性较高。结果支持手绘草图作为一种直观的交互抽象,用于人-集群协作,而无需机器人技术或编程专业知识。
cs.RO / 13 / 2609.11079

RIDE: Relocalization-Informed Depth Estimation with 3D Gaussian Splatting

RIDE:基于3D高斯泼溅的重定位信息引导深度估计
Lian, Jiarong, Xiao, Zhe, Zhang, Zhaoyang, Li, Wei, Chen, Ruizhi
Abstract
Render--match--PnP relocalization establishes correspondences between query image pixels and 3D map points for camera pose recovery, but their potential to support dense depth estimation is often overlooked. To exploit this geometric information, we present RIDE, which estimates dense metric depth from a robot's RGB stream. Given a metrically scaled 3D Gaussian Splatting (3DGS) model, RIDE combines sparse metric depth observations derived from PnP-RANSAC inlier correspondences with the geometric prior of a pretrained video-depth model. To handle uneven and intermittent observations, it integrates global and local depth correction with temporal memory, supporting depth estimation through short observation gaps after metric scale initialization. Trained on public RGB-D videos, RIDE is evaluated on robot sequences without fine tuning. Experiments show improved depth accuracy and temporal consistency over scale-only calibration, demonstrating how localization geometry can support both pose recovery and dense robot perception.
Chinese Translation
渲染-匹配-PnP重定位在查询图像像素与3D地图点之间建立对应关系以恢复相机位姿,但其支持稠密深度估计的潜力常被忽视。为利用这一几何信息,我们提出RIDE,它从机器人的RGB流中估计稠密度量深度。给定具有度量尺度的3D高斯泼溅(3DGS)模型,RIDE将从PnP-RANSAC内点对应关系导出的稀疏度量深度观测与预训练视频深度模型的几何先验相结合。为处理不均匀和间歇性的观测,它集成了全局和局部深度校正与时间记忆,支持在度量尺度初始化后通过短观测间隙进行深度估计。在公共RGB-D视频上训练后,RIDE在机器人序列上进行评估而无需微调。实验表明,与仅尺度校准相比,深度精度和时间一致性得到改善,展示了定位几何如何同时支持位姿恢复和稠密机器人感知。
cs.RO / 14 / 2609.11225

Harness Robotic OS: A Unified Embodied-Agent Runtime for Closed-Loop Quadruped Inspection

Harness Robotic OS:面向闭环四足巡检的统一具身智能体运行时
Yan, Yaoyuan, Heng, Zhiyou, Jie, Haoxiang, Liu, Gang, Yan, Hongjie, Zhou, Wei
Abstract
Autonomous property inspection requires more than robust robot navigation: a deployable system must connect heterogeneous sensing, reusable autonomy capabilities, multimodal scene understanding, human interaction, and enterprise response within a traceable operational loop. Existing quadruped inspection systems commonly integrate these functions through task-specific interfaces, making contextual coordination, knowledge reuse, and controlled adaptation difficult. This paper presents \textit{Harness Robotic OS} (HROS), a unified embodied-agent runtime, and Argos, its realization for residential-community inspection. HROS organizes the system into robot runtime, embodied autonomy skills, cognitive agent runtime, and interaction and operations planes. A shared context connects physical state with agent reasoning; streaming ASR/TTS supports voice-based mission interaction; hierarchical working, episodic, and semantic memory preserves operational knowledge; and a safety-gated self-evolution loop converts execution traces into versioned candidate updates without permitting unconstrained online modification. The Argos prototype integrates a Vbot quadruped, Fast-LIO2 localization and mapping, Hobot-Stereo depth perception, PCT-Planner global planning, EGO-Planner local motion generation, and OpenClaw-orchestrated Qwen3-VL inspection analysis. Experiments in a residential property environment achieved 100\% waypoint reachability, outdoor localization error below 10~cm, local obstacle-response latency below 200~ms, representative hazard-detection rates of 85--95\%, and 99\% success in alarm delivery and structured-report generation. These results validate the deployed navigation and inspection closed loop, while HROS provides an extensible software foundation for memory-augmented, voice-aware, and continuously improvable embodied inspection agents.
Chinese Translation
自主物业巡检不仅需要鲁棒的机器人导航:一个可部署的系统必须在可追溯的运营闭环中连接异构传感、可复用的自主能力、多模态场景理解、人机交互和企业响应。现有的四足巡检系统通常通过任务特定接口集成这些功能,导致上下文协调、知识复用和受控适应变得困难。本文提出 Harness Robotic OS(HROS),一种统一具身智能体运行时,以及 Argos,其在住宅社区巡检中的实现。HROS 将系统组织为机器人运行时、具身自主技能、认知智能体运行时以及交互与运维平面。共享上下文连接物理状态与智能体推理;流式 ASR/TTS 支持基于语音的任务交互;分层工作记忆、情景记忆和语义记忆保存运维知识;安全门控自进化循环将执行轨迹转换为版本化的候选更新,而不允许不受约束的在线修改。Argos 原型集成了 Vbot 四足机器人、Fast-LIO2 定位与建图、Hobot-Stereo 深度感知、PCT-Planner 全局规划、EGO-Planner 局部运动生成,以及 OpenClaw 编排的 Qwen3-VL 巡检分析。在住宅物业环境中的实验实现了 100% 航点可达性、室外定位误差低于 10 cm、局部障碍物响应延迟低于 200 ms、代表性危险检测率为 85%–95%,以及报警传递和结构化报告生成成功率达 99%。这些结果验证了所部署的导航与巡检闭环,而 HROS 为记忆增强、语音感知和持续改进的具身巡检智能体提供了可扩展的软件基础。
cs.RO / 15 / 2609.11270

Beyond Noise Steering: Dual-Latent Space Reinforcement Learning for Generative Robot Policy

超越噪声引导:用于生成式机器人策略的双潜在空间强化学习
Zhang, Pengfei, Sun, Teng, Xiu, Xianchao
Abstract
Pretrained generative robot policies learn expressive action priors from demonstrations. However, existing reinforcement learning methods only steer the noisy space but fail to modulate intermediate action representations during the generation process, resulting in performance degradation and inefficiency. To address this limitation, we propose a novel Dual-Latent Space Reinforcement Learning (DLSRL) framework, which complements initial-noise steering with representation-level control inside the frozen generator. Specifically, our actor network predicts two distinct latent variables: an initial-noise latent variable that steers behavior generation, and an action-representation latent variable for intermediate feature modulation. Moreover, this representation latent variable is mapped to adapter features and ingeniously injected into the hidden states of intermediate action tokens via residual connections. Our dual-control design enables direct adjustment of action representations without updating the base policy. Experiments across generative policy architectures and robotic manipulation tasks show that DLSRL effectively accelerates online robot policy adaptation and achieves competitive performance. Our code is available at \href{https://github.com/xianchaoxiu/DLSRL}{https://github.com/xianchaoxiu/DLSRL}.
Chinese Translation
预训练的生成式机器人策略从演示中学习富有表现力的动作先验。然而,现有的强化学习方法仅引导噪声空间,而无法在生成过程中调节中间动作表示,导致性能下降和效率低下。为了解决这一限制,我们提出了一种新颖的双潜在空间强化学习(DLSRL)框架,该框架通过在冻结生成器内部进行表示级控制来补充初始噪声引导。具体而言,我们的actor网络预测两个不同的潜在变量:一个用于引导行为生成的初始噪声潜在变量,以及一个用于中间特征调制的动作表示潜在变量。此外,该表示潜在变量被映射到适配器特征,并通过残差连接巧妙地注入到中间动作token的隐藏状态中。我们的双控制设计能够在不更新基础策略的情况下直接调整动作表示。在生成式策略架构和机器人操作任务上的实验表明,DLSRL有效加速了在线机器人策略自适应,并达到了有竞争力的性能。我们的代码可在 https://github.com/xianchaoxiu/DLSRL 获取。
cs.RO / 16 / 2609.11308

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

2AM:将智能体侧记忆作为长时程操作中可操纵动作模型的引导基础
Hu, Yutong, Chen, Fengjiao, Cao, Xuezhi, Detry, Renaud
Abstract
Long-horizon robot manipulation requires memory, but not necessarily inside the action policy. To address such tasks, current agentic systems often combine VLAs with planners and geometric tools, sometimes using additional depth or calibrated geometry. These systems confound attribution: gains may come from richer observations or alternative motor tools, while failures may stem from either the policy or an under-specified language interface. We isolate this question through a deliberately constrained design: less tool breadth, but greater interface bandwidth. 2AM makes a multimodal Agent the sole holder of task memory and a single RGB-based, episodically stateless Action Model the sole executor of task-relevant motion. The Agent compiles interaction history into subtask language and optional 2D grasp, place, and move hints that bind its physical intention at different time scales. To teach this steerability to the VLA, we augment demonstrations with structured hint labels and train under condition dropout, spatial noise, and temporal jitter to tolerate imperfect Agent outputs. On LIBERO-Mem, without depth, online geometry, or planner-based object motion, 2AM reaches 76.3% average completion, a 61.5-point improvement over the strongest reported baseline of 14.8%, together with 63.0% relaxed and 11.8% strict success. These results show that task memory can remain Agent-side. They further show that Action Model capability depends not only on what the policy has learned, but on how precisely the Agent can steer it.
Chinese Translation
长时程机器人操作需要记忆,但不一定需要记忆存在于动作策略内部。为解决此类任务,当前的智能体系统通常将 VLA 与规划器和几何工具相结合,有时使用额外的深度信息或标定几何。这些系统混淆了归因:性能提升可能来自更丰富的观测或替代的运动工具,而失败可能源于策略或未充分指定的语言接口。我们通过一种刻意约束的设计来分离这个问题:减少工具广度,但增加接口带宽。2AM 使多模态智能体成为任务记忆的唯一持有者,并让一个基于 RGB 的、逐回合无状态的动作模型成为任务相关运动的唯一执行者。智能体将交互历史编译为子任务语言以及可选的 2D 抓取、放置和移动提示,这些提示在不同时间尺度上绑定其物理意图。为了将这种可操纵性教给 VLA,我们用结构化提示标签增强演示,并在条件 dropout、空间噪声和时间抖动下训练,以容忍不完美的智能体输出。在 LIBERO-Mem 上,不使用深度、在线几何或基于规划器的物体运动,2AM 达到 76.3% 的平均完成率,比报告的最强基线 14.8% 提高了 61.5 个百分点,同时宽松成功率为 63.0%,严格成功率为 11.8%。这些结果表明,任务记忆可以保留在智能体侧。它们进一步表明,动作模型的能力不仅取决于策略学到了什么,还取决于智能体能够多精确地操纵它。
cs.RO / 17 / 2609.11338

Modular Kinematic Reduction of Closed-Chain Mechanisms Using Path Assembly and Defect Homotopy

基于路径装配与缺陷同伦的闭链机构模块化运动学降维
Dastranj, Mohammad, Mattila, Jouni
Abstract
Closed kinematic chains complicate modular modeling by coupling active and passive coordinates through nonlinear closure constraints. This paper presents a Path-Assembled Closure Differential Mapping (PACDM) framework for modular closure resolution and kinematic reduction. Each closure element compares two ordered transformation paths with common endpoints, with their mismatch expressed through the logarithm on SE(3) and the corresponding Jacobian assembled from local transformation derivatives. Multi-path modules are constructed from a minimal set of pairwise closure elements, while rank-revealing analysis selects locally independent scalar constraints. A defect homotopy recovers closure-consistent passive coordinates from approximate estimates along a feasible and regular continuation path. At regular configurations, implicit differentiation yields the local active-to-passive differential mapping, which is subsequently used in a predictor-corrector continuation procedure for prescribed motion. The framework is evaluated on a seven-degree-of-freedom heavy-duty manipulator containing two-path and three-path closed-chain modules. Comparison with Simscape Multibody yields trajectory root-mean-square errors below 8.5 x 10^-10 rad, while predictor-corrector continuation is approximately 45.8 times faster than applying defect homotopy at every trajectory sample.
Chinese Translation
闭式运动链通过非线性闭合约束耦合主动和被动坐标,使模块化建模复杂化。本文提出一种路径装配闭合微分映射(PACDM)框架,用于模块化闭合求解和运动学降维。每个闭合元件比较具有共同端点的两条有序变换路径,其不一致性通过 SE(3) 上的对数表示,并由局部变换导数装配相应的雅可比矩阵。多路径模块由最小成对闭合元件集构建,而秩揭示分析选择局部独立的标量约束。缺陷同伦沿可行且正则的连续路径从近似估计中恢复闭合一致的被动坐标。在正则位形下,隐式微分产生局部主动到被动微分映射,随后用于给定运动的预测-校正连续过程。该框架在一个包含双路径和三路径闭链模块的七自由度重载机械臂上进行了评估。与 Simscape Multibody 的比较得到轨迹均方根误差低于 8.5 x 10^-10 rad,而预测-校正连续过程比在每个轨迹采样点应用缺陷同伦快约 45.8 倍。
cs.RO / 18 / 2609.11357

Morphology-Aware Human Motion Retargeting for Wheeled-Humanoid Loco-Manipulation

形态感知的轮式人形机器人移动-操作人体运动重定向
Xia, Chenbo, Ye, Chao
Abstract
Human-to-humanoid retargeting has largely been studied on legged platforms, while comparatively few wheeled-humanoid systems support coupled locomotion and manipulation from general human motion. Building on GMR's configurable general-motion retargeting and BeyondMimic's physically simulated R1 Pro learning framework, we present a reproducible pipeline that converts multi-dataset SMPLX motion into executable loco-manipulation behavior for the Galaxea R1 Pro wheeled humanoid. The robot has a planar three-wheel base, a serial torso, and two arms but no leg joints, so human lower-body motion must be redistributed across base motion and torso posture without sacrificing manipulation-relevant arm geometry. Our pipeline combines canonical body-shape preprocessing, planar-base normalization, morphology-aware differential inverse kinematics, shoulder-rooted hierarchical arm retargeting, and continuous torso substitution for bending and squatting. A reference-twist-driven planning layer then decodes planar base motion into continuous three-wheel steering and rolling commands subject to hysteresis, kinematic continuity, acceleration, and actuator-rate limits. Finally, a 21-dimensional BaseDecode policy is trained in Isaac Lab with directional joint-limit scaling, focused upper-body tracking, and a staged wheel-contact reward. The resulting system provides a complete bridge from human motion data to physically trackable wheeled-humanoid loco-manipulation rather than a visualization-only retargeter; quantitative policy comparisons remain scheduled for a later revision.
Chinese Translation
人体到人形机器人的重定向主要在腿式平台上进行研究,而相对较少的轮式人形机器人系统支持从一般人体运动中耦合的移动和操作。基于GMR的可配置通用运动重定向和BeyondMimic的物理仿真R1 Pro学习框架,我们提出了一种可复现的流程,将多数据集SMPLX运动转换为Galaxea R1 Pro轮式人形机器人的可执行移动-操作行为。该机器人具有平面三轮底座、串联躯干和两条手臂,但没有腿部关节,因此必须将人体下半身运动重新分配到基座运动和躯干姿态上,同时不牺牲与操作相关的手臂几何结构。我们的流程结合了规范体型预处理、平面基座归一化、形态感知微分逆运动学、以肩部为根的分层手臂重定向,以及用于弯腰和下蹲的连续躯干替换。随后,一个参考扭转驱动的规划层将平面基座运动解码为连续的三轮转向和滚动命令,并受限于迟滞、运动学连续性、加速度和执行器速率限制。最后,在Isaac Lab中训练了一个21维的BaseDecode策略,采用方向性关节限位缩放、聚焦上半身跟踪和分阶段车轮接触奖励。由此产生的系统提供了从人体运动数据到可物理跟踪的轮式人形机器人移动-操作的完整桥梁,而不仅仅是可视化重定向器;定量策略比较仍计划在后续修订中进行。
cs.RO / 19 / 2609.11361

GeoTrussRover: Morphological Computation with Contact-Semantic Control Primitives

GeoTrussRover:基于接触语义控制原语的形态计算
Ma, Muyuan, Zhang, Yi, Yang, Yang, Zheng, Xuanyan, Hu, Ruiqi, Ke, Boxuan, Chen, Zhenyu, Lin, Yicong, Yang, Xin Hao, Xiao, Daliang, Hou, Zhinan, Niu, Wanhao, Sun, Yuan, Yang, Yan, Xie, Yue
Abstract
Reconfigurable robots can change their contact geometry when a fixed body cannot negotiate an obstacle. A variable-geometry truss (VGT) distributes this shape change through a load-bearing structure, but coupling it to a mobile base creates a high-dimensional coordination problem. GeoTrussRover combines an electrically actuated VGT, a wheeled base, and contact-semantic morphology planning and control. We solve one source traversal and extract four contact-semantic primitives that describe coordination among 21 members. Physics-constrained projection adapts them to unseen step heights with the same contact topology. When every phase remains feasible, adaptation does not recompute the complete motion. If one phase violates the new physical constraints, only that phase is recomputed. A full-space QP then tracks the adapted motion and corrects member and wheel errors. For transfer from 0.10m to 0.075m, the method reduces objective-function evaluations by 63.7% relative to full recomputation. Contact-phase feasibility analysis covers step heights from 0.10 to 0.46m, or 1.08 to 4.97 wheel radii, with the upper value near the theoretical feasible boundary. The electric prototype traverses 2.11 wheel radii. The resulting low-dimensional representation stores task coordination in a hyper-redundant, load-bearing morphology and reuses it during locomotion.
Chinese Translation
可重构机器人可以在固定本体无法通过障碍时改变其接触几何。可变几何桁架(VGT)通过承载结构分布这种形状变化,但将其与移动底座耦合会产生高维协调问题。GeoTrussRover 将电动驱动的 VGT、轮式底座以及接触语义形态规划与控制相结合。我们求解一次源穿越,并提取四个接触语义原语,用于描述 21 个构件之间的协调。物理约束投影将它们适配到具有相同接触拓扑的未见台阶高度。当每个阶段仍然可行时,适配不会重新计算完整运动。如果某一阶段违反新的物理约束,则仅重新计算该阶段。随后,全空间 QP 跟踪适配后的运动,并校正构件和车轮误差。对于从 0.10m 到 0.075m 的迁移,该方法相对于完全重新计算将目标函数评估次数减少 63.7%。接触阶段可行性分析覆盖台阶高度从 0.10 到 0.46m,即 1.08 到 4.97 个车轮半径,其中上限值接近理论可行边界。电动原型可跨越 2.11 个车轮半径。由此产生的低维表示将任务协调存储在超冗余承载形态中,并在运动过程中复用它。
cs.RO / 20 / 2609.11382

SwarmNxt: Open-source Software-Hardware Platform for Fast and Agile Aerial Swarms

SwarmNxt:用于快速敏捷空中集群的开源软硬件平台
Toumieh, Charbel, Mistry, Niel, Jarvis, Benjamin, Jeger, Simon, Liu, Peize, Shen, Shaojie, Floreano, Dario
Abstract
Aerial robot swarms have the potential to transform time-critical safety, security, and search-and-rescue operations. By coordinating multiple robots, they can rapidly survey disaster sites, map collapsed or GPS-denied environments, and search cluttered areas faster than a single robot, reducing response times and minimizing risks to first responders. Realizing this potential, however, requires robust autonomous swarm navigation, which remains an active research challenge. Progress is further constrained by existing platforms, as commercial drones are often closed-source or lack the onboard computational resources needed for agile, vision-based collective flight. Moreover, developing, deploying, and maintaining software across multiple aerial robots requires significant engineering effort. To address these challenges, we present SwarmNxt, an open-source software platform built on the open-source OmniNxt drone hardware. SwarmNxt provides an end-to-end toolkit, including detailed hardware assembly instructions with a video tutorial, automation tools for parallel software deployment and swarm-wide updates, and a ROS 2-based framework for autonomous navigation. The platform integrates state-of-the-art control, planning, and depth estimation into a single ROS 2 multi-agent system, providing an open research infrastructure for physical swarm experimentation. We validate SwarmNxt through two real-world experiments: a six-drone swarm performing decentralized planning with high-speed inter-drone collision avoidance, and a four-drone swarm executing collective flight with onboard depth estimation in an obstacle-filled environment. Both experiments were run indoors with global position from external motion capture; perception, planning, and control run onboard.
Chinese Translation
空中机器人集群有潜力变革时间关键的安全、安防和搜救行动。通过协调多台机器人,它们能够比单个机器人更快地快速勘察灾区、测绘坍塌或GPS拒止环境,并搜索杂乱区域,从而缩短响应时间并最大限度降低一线救援人员的风险。然而,实现这一潜力需要稳健的集群自主导航,这仍是一个活跃的研究挑战。现有平台进一步制约了进展,因为商用无人机通常闭源,或缺乏敏捷的、基于视觉的集群飞行所需的机载计算资源。此外,在多个空中机器人上开发、部署和维护软件需要大量工程投入。为解决这些挑战,我们提出SwarmNxt,一个构建在开源OmniNxt无人机硬件之上的开源软件平台。SwarmNxt提供端到端工具包,包括带视频教程的详细硬件组装说明、用于并行软件部署和集群范围更新的自动化工具,以及基于ROS 2的自主导航框架。该平台将最先进的控制、规划和深度估计集成到一个ROS 2多智能体系统中,为物理集群实验提供开放研究基础设施。我们通过两个真实世界实验验证SwarmNxt:一个六无人机集群执行去中心化规划并进行高速无人机间避碰,以及一个四无人机集群在充满障碍物的环境中利用机载深度估计执行集群飞行。两个实验均在室内进行,使用来自外部运动捕捉系统的全局位置;感知、规划和控制均在机载运行。
cs.RO / 21 / 2609.11433

Safety-aware Skill Adaptation for Reinforcement Learning in Dynamic Environments

动态环境中强化学习的安全感知技能自适应
Haque, A K M Nadimul, Sutjipto, Sheila, Carmichael, Marc G., Vidal-Calleja, Teresa
Abstract
Skill adaptation frameworks based on reinforcement learning often require restrictive assumptions to maintain stability, such as fixed observations or tightly controlled exploration schedules. In cluttered and dynamic environments, however, unrestricted exploration can lead to unsafe behaviour and unstable learning, particularly when task-relevant observations lie near obstacles or involve moving objects. In this work, we present Dist-GPRL, a distance-aware and safety-guided reinforcement learning framework for structured robot skill adaptation. Building upon Gaussian Process (GP)-based skill parameterisation, our framework sequentially adapts overlapping local windows of sparse trajectory via-points rather than modifying the complete skill at every policy step. Raw policy outputs are correlated through the GP covariance structure, producing temporally coherent trajectory updates while reducing the action-space and credit-assignment difficulties associated with global trajectory adaptation. Safety is incorporated through two complementary forms of guidance. A safe-subspace prior derived from the Hausdorff Approximation Planner (HAP) biases policy exploration toward feasible regions, while dynamically updated distance field clearance and gradient rewards provide local obstacle awareness. A trajectory-kinematics similarity regulariser further preserves the demonstrated velocity and acceleration characteristics during adaptation. We evaluate the framework on two dynamic object-manipulation tasks in simulation and transfer the learned policy to real-world robot execution. Experimental results demonstrate higher task success, lower collision frequency, and more stable learning than the baselines, while preserving the kinematic characteristics of the demonstrated skill.
Chinese Translation
基于强化学习的技能适应框架通常需要限制性假设来保持稳定性,例如固定的观测或严格控制的探索调度。然而,在杂乱和动态的环境中,不受限制的探索可能导致不安全行为和不稳定学习,特别是当任务相关观测位于障碍物附近或涉及移动物体时。在这项工作中,我们提出了 Dist-GPRL,一种用于结构化机器人技能适应的距离感知和安全引导的强化学习框架。基于高斯过程(GP)的技能参数化,我们的框架顺序地适应稀疏轨迹途经点的重叠局部窗口,而不是在每个策略步骤修改完整的技能。原始策略输出通过 GP 协方差结构进行关联,产生时间上连贯的轨迹更新,同时减少与全局轨迹适应相关的动作空间和信用分配困难。安全性通过两种互补形式的引导被纳入。从 Hausdorff 近似规划器(HAP)导出的安全子空间先验将策略探索偏向可行区域,而动态更新的距离场间隙和梯度奖励提供局部障碍物感知。轨迹-运动学相似性正则化器进一步在适应过程中保持演示的速度和加速度特性。我们在仿真中的两个动态物体操作任务上评估该框架,并将学习到的策略迁移到真实世界机器人执行中。实验结果表明,与基线相比,任务成功率更高、碰撞频率更低、学习更稳定,同时保持演示技能的运动学特性。
cs.RO / 22 / 2609.11445

FARM: Reading Failure Signals from the Internal Predictive States of a Frozen Robotic World Model

FARM:从冻结的机器人世界模型的内部预测状态中读取故障信号
Pei, Haoran, Luo, Mingrui, Wang, Senbao, Lv, Haoran, Guo, Jie, Zhong, Sheng, Ci, Ruixi
Abstract
Reliable robot deployment requires online failure monitoring, yet existing monitors mainly derive risk from proxy signals or train dedicated monitoring components. We ask whether the internal predictive states of a frozen pretrained robotic world model already contain directly decodable failure information. Failure-Aware Readout from World Models (FARM) trains only a 33,985-parameter supervised readout over frozen VLA-JEPA predictive states, producing step-wise failure scores and causal trajectory risk. Five-fold out-of-fold evaluation across seven source tasks reaches 85.68/88.59 pooled AUROC/AUPRC, and FARM gives the best Seen performance among 15 matched baselines on the 10-task benchmark. Across four real-robot populations on PIPER X, SO-101, and Franka, fixed-readout transfer and readout-only adaptation test deployment shifts without updating the predictive backbone. FARM also discriminates failures from partial causal histories and adds 0.2256 ms mean CUDA latency once the frozen state is available. These results support frozen predictive world-model states as reusable features for causal, transferable, and low-overhead execution monitoring.
Chinese Translation
可靠的机器人部署需要在线故障监测,然而现有监测器主要从代理信号中推导风险,或训练专用的监测组件。我们探究冻结的预训练机器人世界模型的内部预测状态是否已经包含可直接解码的故障信息。世界模型故障感知读出(Failure-Aware Readout from World Models, FARM)仅在冻结的VLA-JEPA预测状态上训练一个33,985参数的监督读出,产生逐步故障分数和因果轨迹风险。在七个源任务上的五折折外评估达到了85.68/88.59的汇总AUROC/AUPRC,并且FARM在10任务基准上在15个匹配基线中给出最佳的已见性能。在PIPER X、SO-101和Franka上的四个真实机器人群体中,固定读出迁移和仅读出适配测试了部署偏移,而无需更新预测主干。FARM还能从部分因果历史中区分故障,并且在冻结状态可用时平均增加0.2256毫秒的CUDA延迟。这些结果支持将冻结的预测世界模型状态作为可重用特征,用于因果、可迁移和低开销的执行监测。
cs.RO / 23 / 2609.11476

3D Euler-Angle Orientation Control for Two-Ray Fading Mitigation in Maritime Air-to-Sea Communications

海上空对海通信中双径衰落缓解的三维欧拉角姿态控制
Bajja, Mohammed, Saliah, Abdoul Karim A. H., Hammouti, Hajar El, Licea, Daniel Bonilla, Silano, Giuseppe
Abstract
Maritime Air-to-Sea links are dominated by a line-of-sight ray and a sea-surface reflected ray whose destructive combination produces deep fades. Existing mitigation strategies optimize Unmanned Aerial Vehicle position or trajectory but leave attitude unexploited. This paper treats the full three dimensional attitude as a physical-layer control variable that shapes the two-ray interference through antenna phase-center displacement. Under small-angle and far-field assumptions, the constructive-interference condition reduces to an affine constraint in the Euler angles and admits a closed-form family of minimum-norm attitude candidates. A differentiable soft-minimum rule yields a smooth reference tracked by a constrained Nonlinear Model Predictive Control controller on a fully-actuated tilting multirotor. The proposed scheme increases cumulative throughput by 11.4% over a pitch-only benchmark and 22.2% over a zero-orientation baseline, while preserving trajectory tracking and respecting actuator limits.
Chinese Translation
海上空对海链路主要由一条视距射线和一条海面反射射线主导,二者的破坏性叠加会产生深度衰落。现有缓解策略优化无人机位置或轨迹,但未利用姿态。本文将完整三维姿态视为物理层控制变量,通过天线相位中心位移来塑造双径干扰。在小角度和远场假设下,相长干扰条件简化为欧拉角中的仿射约束,并给出最小范数姿态候选的闭式族。可微软最小规则产生平滑参考,由全驱动倾转多旋翼上的约束非线性模型预测控制控制器跟踪。所提方案相较于仅俯仰基准将累积吞吐量提高11.4%,相较于零姿态基线提高22.2%,同时保持轨迹跟踪并满足执行器限制。
cs.RO / 24 / 2609.11478

CARLAverse: A Highly Modular, Distributed, and Multimodal Framework for Human-in-the-Loop Simulation

CARLAverse:一个高度模块化、分布式、多模态的人在回路仿真框架
Rebling, Patrick, Nenninger, Philipp, Kriesten, Reiner
Abstract
The development of autonomous driving demands comprehensive testing in mixed-traffic scenarios involving vulnerable road users (VRUs), where purely artificial agents often fail to capture authentic human social negotiations. While human-in-the-loop (HITL) simulators enable safe investigation of these interactions, existing multi-agent platforms struggle with the network latency and synchronization constraints required for high-fidelity haptic feedback. To resolve this, we present CARLAverse, an open-source, multimodal simulation ecosystem. Extending modular hardware abstraction, CARLAverse integrates driving (DrivoCARLA), cycling (CycloCARLA), and pedestrian (WalkoCARLA) simulators into a shared virtual environment. Its core methodological contribution is a distributed physics architecture: latency-critical ego dynamics and high-frequency force feedback are computed locally on client nodes, while a central CARLA server orchestrates non-player character (NPC) physics and global traffic. By decoupling haptic control loops from network bottlenecks, CARLAverse enables scalable, cross-institutional HITL experiments without compromising physical immersion. Code and documentation: https://git.ieem-ka.de/simulator-environments/carlaverse
Chinese Translation
自动驾驶的发展需要在涉及弱势道路使用者(VRUs)的混合交通场景中进行全面的测试,而纯人工代理往往无法捕捉真实的人类社会协商。虽然人在回路(HITL)模拟器能够安全地研究这些交互,但现有的多智能体平台难以满足高保真力反馈所需的网络延迟和同步约束。为了解决这个问题,我们提出了CARLAverse,一个开源的多模态仿真生态系统。通过扩展模块化硬件抽象,CARLAverse将驾驶(DrivoCARLA)、骑行(CycloCARLA)和行人(WalkoCARLA)模拟器集成到一个共享的虚拟环境中。其核心方法学贡献是分布式物理架构:延迟关键的自身动力学和高频力反馈在客户端节点本地计算,而中央CARLA服务器则协调非玩家角色(NPC)物理和全局交通。通过将触觉控制回路从网络瓶颈中解耦,CARLAverse支持可扩展的跨机构HITL实验,而不会影响物理沉浸感。代码和文档:https://git.ieem-ka.de/simulator-environments/carlaverse
cs.RO / 25 / 2609.11549

Using Automated Vehicles Operational Data to Confirm Safety and Anticipate Threats

利用自动驾驶车辆运行数据确认安全并预见威胁
Donà, Riccardo, Rusciano, Espedito, Trentadue, Germana, Tsakalidis, Anastasios, Galassi, Maria Cristina
Abstract
European Union (EU) policymakers adopted revolutionary data collection provisions for Automated Driving Systems (ADS) in the recently approved regulation that allows driverless vehicles to be operated on public roads. The framework is inspired by best practices developed at the United Nations Economic Commission for Europe(UNECE) level: the In-Service Monitoring and Reporting (ISMR); and by similar operational data collection regulatory approaches in nuclear energy production and transportation fields. The collection of real-world data will enable the competent safety authorities to gather the information needed to confirm the homologation safety target. Safety-relevant driving scenarios discovered during the real-world operation of a given ADS can also be stored in a scenario catalogue to investigate how other ADS types might have addressed such a traffic conflict. Moreover, lessons learnt deriving from the data collected can be shared among original equipment manufacturers (OEMs) and safety authorities. Ultimately, the ISMR is recognised as a necessary tool to properly tackle the challenges associated with ADS safety assessment given the number of unknowns that might remain undisclosed by leveraging the traditional homologation validation scheme only.
Chinese Translation
欧盟(EU)政策制定者在最近批准的一项允许无人驾驶车辆在公共道路上运行的法规中,针对自动驾驶系统(ADS)采用了革命性的数据收集条款。该框架受到联合国欧洲经济委员会(UNECE)层面制定的最佳实践的启发:在役监测与报告(ISMR);以及核能生产和交通运输领域类似的运行数据收集监管方法。真实世界数据的收集将使主管安全当局能够收集确认认证安全目标所需的信息。在特定ADS真实运行过程中发现的安全相关驾驶场景也可以存储在场景目录中,以研究其他ADS类型可能如何处理此类交通冲突。此外,从收集的数据中得出的经验教训可以在原始设备制造商(OEM)和安全当局之间共享。最终,ISMR被认为是适当应对ADS安全评估相关挑战的必要工具,因为仅依靠传统的认证验证方案可能仍有许多未知因素未被揭示。
cs.RO / 26 / 2609.11553

CAP: Continuously Adaptive Perception-Blind Humanoid Locomotion via Learned Denoising

CAP:通过学习去噪的连续自适应感知盲人形机器人运动
Chen, Hongjin, Xu, Zijun, Ma, Shihao, Zhao, Yi, Liu, Xilai, Ma, Ke, Zhang, Wei, Xie, Chunyang, Li, Pengfei, Zhao, Jieru, Ding, Wenchao
Abstract
Humanoid locomotion across complex terrain demands forward-looking exteroception to anticipate obstacles, yet this signal is unreliable in real-world deployment, failing partially and intermittently. Existing perceptive policies often assume that depth observations remain clean and in-distribution, while recent attempts to unify perceptive and blind control typically route or switch between separate sub-policies, leaving recoverable information in partially corrupted depth unexploited. We instead propose CAP, a single-stage humanoid locomotion policy that recovers this signal with a perceptive world-model encoder trained as a learned denoiser to reconstruct clean depth from a corrupted input, together with a co-active proprioceptive variational encoder that supplies depth-free body-state information. A coupled training recipe pairs a depth-noise curriculum on the world-model input with world-model feature dropout on the policy-facing latent, exposing the policy to failures across the entire perception-quality spectrum. In simulation, CAP matches or improves upon perceptive baselines when depth remains informative, and degrades more smoothly than a binary-switching baseline as perception worsens. On the Unitree G1, controlled trials and indoor-outdoor deployments demonstrate perception-robust locomotion under intermittent occlusion, real-sensor corruption, and outdoor depth artifacts.
Chinese Translation
人形机器人在复杂地形上的运动需要前瞻性的外部感知来预判障碍物,然而这种信号在现实世界部署中并不可靠,会部分和间歇性地失效。现有的感知策略通常假设深度观测保持干净且同分布,而最近统一感知与盲控的尝试通常在单独的子策略之间路由或切换,从而未利用部分损坏深度中的可恢复信息。我们转而提出CAP,一种单阶段人形机器人运动策略,它通过一个感知世界模型编码器(被训练为学习去噪器,从损坏输入中重建干净深度)来恢复该信号,并结合一个协同活跃的本体感受变分编码器,提供无深度的身体状态信息。一种耦合训练方案将世界模型输入上的深度噪声课程与面向策略的潜在表示上的世界模型特征丢弃配对,使策略暴露于整个感知质量范围内的故障。在仿真中,当深度仍然提供信息时,CAP匹配或优于感知基线,并且随着感知恶化,比二元切换基线退化得更平滑。在Unitree G1上,受控试验和室内外部署展示了在间歇性遮挡、真实传感器损坏和室外深度伪影下的感知鲁棒运动。
cs.RO / 27 / 2609.11561

Memory as Plans: World-Action Modeling with Memory-Grounded Planning

记忆即规划:基于记忆的世界-动作建模与规划
Zhao, Sizhe, Xie, Haozhe, Zhao, Weiyu, Zhang, Chenchu, Wang, Huan, Wang, Chenyang, Liu, Qinglin, Zhang, Shengping
Abstract
Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their combinations, and may therefore lose fine-grained visual evidence or face a trade-off between history coverage and execution efficiency. We introduce MaP-WAM, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history. MaP-WAM represents memory as completed segment records containing language instructions and sparse visual context, and converts this episodic memory into compact plans comprising the next segment-level language plan and corresponding visual guidance. A World-Action-Progress (WAP) model executes each plan over an unknown duration by jointly predicting action chunks and corresponding execution progress at inference time, calibrating predicted progress through plan-observation alignment for adaptive segment transitions and closed-loop context updates. MaP-WAM keeps the executor context length fixed, while structured attention further enables key-value caching in both planning and execution. MaP-WAM achieves state-of-the-art performance on RMBench with an 83.3% success rate and attains 78.0% success on real-robot tasks, while maintaining approximately constant executor inference latency as task history grows.
Chinese Translation
主流的机器人策略通常采用马尔可夫形式,但许多复杂的现实世界操作任务本质上是非马尔可夫的,需要超越当前观察的长时程记忆。现有的记忆机制通常依赖于语言摘要、不断增长的视觉窗口或它们的组合,因此可能会丢失细粒度的视觉证据,或者在历史覆盖与执行效率之间面临权衡。我们提出了 MaP-WAM,一种记忆即规划(Memory-as-Plans)框架,它将依赖记忆的世界-动作建模分解为基于记忆的规划和以规划为条件的执行,并在规划时使用长期多模态情景上下文作为证据,而不是反复将完整历史作为执行器的条件。MaP-WAM 将记忆表示为包含语言指令和稀疏视觉上下文的已完成片段记录,并将这种情景记忆转换为紧凑的规划,包括下一个片段级的语言规划和相应的视觉引导。一个世界-动作-进度(WAP)模型在推理时通过联合预测动作块和相应的执行进度,在未知时长内执行每个规划,并通过规划-观察对齐来校准预测进度,以实现自适应片段转换和闭环上下文更新。MaP-WAM 保持执行器上下文长度固定,而结构化注意力进一步在规划和执行中实现键值缓存。MaP-WAM 在 RMBench 上以83.3%的成功率实现了最先进的性能,并在真实机器人任务上达到了78.0%的成功率,同时随着任务历史的增长,执行器推理延迟保持大致恒定。
cs.RO / 28 / 2609.11579

Quasi-static analysis of passive stability in a novel underactuated multi-finger hand

一种新型欠驱动多指手被动稳定性的准静态分析
Plancoulaine, Léonie, Guégan, Sylvain, Plestan, Franck, Chablat, Damien
Abstract
Underactuated robotic hands achieve adaptive and robust grasping with a reduced number of actuators, but predicting the stable equilibrium pose of the grasped object remains a significant challenge. This paper introduces a quasi-static analytical approach to assess passive stability in underactuated multi-finger hands. A novel three-finger hand architecture integrating a differential spring-loaded slider mechanism is introduced, enabling versatile and adaptive grasping. The study focuses on how the differential mechanism influences the overall grasp behavior and analyzes the effect of object size on the stable equilibrium configurations for two canonical grasp types: cylindrical and spherical.
Chinese Translation
欠驱动机械手能够以较少数量的驱动器实现自适应且稳健的抓取,但预测被抓取物体的稳定平衡位姿仍然是一个重大挑战。本文提出一种准静态分析方法,用于评估欠驱动多指手的被动稳定性。引入了一种集成差动弹簧加载滑块机构的新型三指手架构,可实现多样化且自适应的抓取。研究重点关注差动机构如何影响整体抓取行为,并分析物体尺寸对两种典型抓取类型(圆柱形和球形)的稳定平衡构型的影响。
cs.RO / 29 / 2609.11661

Contact-Aware Incremental Model Predictive Control for an Underactuated Aerial Manipulator

面向欠驱动空中机械臂的接触感知增量模型预测控制
Liu, Darwin, Keviczky, Tamas, Sun, Sihao
Abstract
We present a robust contact-aware control framework for aerial writing on an underactuated platform. The framework combines nonlinear model predictive control (NMPC) for accurate end-effector position and normal-force tracking at small reference penetration depths, with consistent performance across controller tunings, with whole-body incremental nonlinear dynamic inversion (INDI) for robustness to frictional and aerodynamic disturbances during contact. The proposed controllers are validated on a quadrotor-based aerial manipulator with a rigid, single-link, one-degree-of-freedom (DoF) arm in simulation and real-world experiments. The aerial writing experiments span vertical and inclined surfaces, multiple reference forces, different friction conditions, and wind disturbances. The results demonstrate that robust simultaneous five-DoF end-effector pose and contact-force tracking is achievable on a standard underactuated quadrotor with a simple, rigid, single-link arm, without requiring a fully actuated platform, a complex arm, or dedicated force/torque sensing.
Chinese Translation
我们提出了一种用于欠驱动平台上空中书写的鲁棒接触感知控制框架。该框架结合了非线性模型预测控制(NMPC)用于在小参考穿透深度下精确的末端执行器位置和法向力跟踪,并且在控制器调参中具有一致的性能,以及全身增量非线性动态逆(INDI)用于在接触过程中对摩擦和气动扰动的鲁棒性。所提出的控制器在具有刚性、单连杆、单自由度(DoF)臂的四旋翼空中机械臂上进行了仿真和真实世界实验验证。空中书写实验涵盖了垂直和倾斜表面、多种参考力、不同摩擦条件和风扰动。结果表明,在标准欠驱动四旋翼上使用简单、刚性、单连杆臂,无需全驱动平台、复杂臂或专用力/扭矩传感,即可实现鲁棒的同步五自由度末端执行器位姿和接触力跟踪。
cs.RO / 30 / 2609.11697

ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies

ActSafeGuard:面向流匹配策略的可微且训练对齐的约束强制执行
Ma, Jianming, Jin, Rongjun, Si, Xiaxi, Zhang, Yang, Li, Yiheng, Gao, Yue
Abstract
Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment. Existing safety approaches either optimize statistical safety objectives without deterministic per-step guarantees or correct unsafe actions only during inference, creating a mismatch between policy training and execution. We introduce ActSafeGuard, a differentiable and training-aligned safeguard layer for flow-matching based policies. ActSafeGuard integrates hard action feasibility into policy learning, not merely treating safety as an inference-time external component. Through an analytical ray-scaling operator design, ActSafeGuard enables boundary-aware gradients to guide the model to naturally learn constrained manifolds. Extensive experiments on multiple standard foundation backbones ($\pi_{0.5}$ and Fast-WAM) across various tasks demonstrate that ActSafeGuard consistently achieves a $100\%$ step safety rate while fully preserving or even boosting task success rates, providing a scalable and minimally invasive solution for safe embodied AI deployment.
Chinese Translation
视觉-语言-动作(VLA)模型和世界-动作模型(WAMs)在通用机器人操作中展现了强大能力,但其生成的动作可能违反硬物理约束,因而在部署中不安全或不可行。现有安全方法要么优化统计安全目标而缺乏确定性的逐步保证,要么仅在推理时纠正不安全动作,导致策略训练与执行之间不匹配。我们提出 ActSafeGuard,一种用于基于流匹配的策略的可微且与训练对齐的安全防护层。ActSafeGuard 将硬动作可行性集成到策略学习中,而不仅仅将安全性视为推理时的外部组件。通过解析射线缩放算子设计,ActSafeGuard 使边界感知梯度能够引导模型自然地学习受约束流形。在多个标准基础骨干网络(π_{0.5} 和 Fast-WAM)上跨各种任务的大量实验表明,ActSafeGuard 始终实现 100% 的步骤安全率,同时完全保持甚至提升任务成功率,为安全的具身智能部署提供了可扩展且最小侵入的解决方案。
cs.RO / 31 / 2609.11698

Aerodynamic Prior-Free Coordinated Trajectory Generation and Tracking Control for a Tail-Sitter UAV

尾座式无人机的无气动先验协同轨迹生成与跟踪控制
Rong, Erchao, Liu, Zihao, Liang, Junning, Wang, Jianguo, Jie, Xiao, Fu, Haoran, Chen, Ziliang, Lyu, Ximin
Abstract
This paper presents a coordinated trajectory generation and tracking control framework for a tail-sitter unmanned aerial vehicle (UAV), which does not require aerodynamic priors identified for a specific airframe while addressing the challenge of flight control under highly nonlinear aerodynamics across the full flight envelope. The core innovation lies in employing phase-specific aerodynamic modeling strategies for planning and tracking, tailored to their distinct functional characteristics, without requiring airframe-specific aerodynamic priors. Specifically, the phi-theory model under coordinated flight is employed to derive an analytic differential flatness mapping, and a simplified but locally accurate model is established for predictive control to enable real-time aerodynamic parameter estimation. The proposed framework is evaluated extensively through both simulation and challenging real-world flight tests under mild wind conditions, showing high-precision tracking and adaptability across the tested aerodynamic conditions. To the best of our knowledge, this is the first real-world demonstration of accurate trajectory tracking over tested flight regimes spanning the full envelope of a tail-sitter UAV without relying on aerodynamic identification campaigns. The source code of our framework is available at: https://github.com/SYSU-HILAB/AP-PnC.
Chinese Translation
本文提出了一种针对尾座式无人机(UAV)的协同轨迹生成与跟踪控制框架,该框架无需针对特定机身进行气动辨识,同时解决了全飞行包线内高度非线性气动条件下的飞行控制难题。核心创新在于针对规划和跟踪的不同功能特性,采用了分阶段的气动建模策略,且无需机身特定的气动先验。具体而言,采用协调飞行下的phi理论模型推导解析微分平坦映射,并建立简化但局部精确的模型用于预测控制,以实现实时气动参数估计。通过仿真和温和风况下的挑战性实际飞行测试,对所提框架进行了广泛评估,在测试的气动条件下展现出高精度跟踪和适应性。据我们所知,这是首次在不依赖气动辨识实验的情况下,实现尾座式无人机全包线内测试飞行区域的精确轨迹跟踪的真实世界演示。我们的框架源代码可在以下网址获取:https://github.com/SYSU-HILAB/AP-PnC。
cs.RO / 32 / 2609.11733

Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion

反射信息引导的神经肌肉强化学习用于肌肉驱动运动
Zhou, Jian, Zhang, Xingyu, Ma, Rui, Cao, Yu, Xie, Shane, Zhang, Zhi-qiang
Abstract
Muscle-driven locomotion provides a physically grounded approach to generating realistic human movement. However, achieving both physiological plausibility and adaptability to changes in musculoskeletal capacity and external disturbances remains a fundamental challenge. To address this limitation, we propose a Reflex-Informed Neuromuscular Reinforcement Learning framework for muscle-driven locomotion. Within this framework, a fixed phase-dependent reflex controller serves as the underlying neuromuscular control mechanism, while the reinforcement learning policy produces four biomechanically meaningful residual parameters to modulate key reflex gains and thresholds associated with hip swing, knee support, and ankle propulsion according to the current state. Experimental results demonstrate that the proposed framework generates physiologically plausible locomotion with improved kinematic accuracy and dynamic consistency, as well as better bilateral symmetry and stride-to-stride consistency under nominal walking conditions. The learned policy remains robust under muscle weakness and external perturbations without retraining.
Chinese Translation
肌肉驱动运动提供了一种基于物理的方法来生成逼真的人体运动。然而,同时实现生理合理性和对肌肉骨骼能力变化及外部扰动的适应性仍然是一个根本性挑战。为了解决这一局限,我们提出了一种用于肌肉驱动运动的反射信息引导的神经肌肉强化学习框架。在该框架中,一个固定的相位依赖反射控制器充当底层神经肌肉控制机制,而强化学习策略则产生四个具有生物力学意义的残差参数,以根据当前状态调节与髋部摆动、膝关节支撑和踝关节推进相关的关键反射增益和阈值。实验结果表明,所提出的框架能够生成生理合理的运动,在标称步行条件下具有改进的运动学精度和动态一致性,以及更好的双侧对称性和步与步之间的一致性。学习到的策略在肌肉无力和外部扰动下无需重新训练仍保持鲁棒性。
cs.RO / 33 / 2609.11753

SEED-UMI: Sharing the Exoskeleton between human and robot for onE-to-one Dexterous demonstration

SEED-UMI:在人类与机器人之间共享外骨骼以实现一对一灵巧演示
Yu, Tengbo, Wu, Jiahao, Li, Daohan, Chen, Bingxu, Liu, Hao, Ma, Xiaojian, Liu, Hangxin
Abstract
Imitation learning for dexterous hands is bottlenecked by the difficulty of collecting contact-rich demonstrations that transfer faithfully to the robot. Prior wearable-exoskeleton systems record only on the human side and retarget via open-loop mappings calibrated in free space, which degrade under contact. We present SEED-UMI, a framework in which both the human and the robot wear the same exoskeleton: joint encoders become a physically shared measurement, and wrist cameras mounted to the exoskeleton observe the same outer mechanism during both human data collection and robot policy rollouts. This turns retargeting into paired cross-embodiment supervision and lets policies train directly on raw exoskeleton-centric wrist images, without segmentation or inpainting. On five contact-rich tasks, SEED-UMI achieves 3.0x greater data collection efficiency than exoskeleton-based teleoperation and a 70.0% average rollout success rate.
Chinese Translation
灵巧手的模仿学习受限于难以收集能够忠实迁移到机器人的丰富接触演示。先前的可穿戴外骨骼系统仅记录人类侧,并通过在自由空间中校准的开环映射进行重定向,这种映射在接触下会退化。我们提出SEED-UMI框架,其中人类和机器人都穿戴相同的外骨骼:关节编码器成为物理共享的测量,并且安装在手腕上的摄像机在人类数据收集和机器人策略执行过程中观察相同的外部机制。这将重定向转变为成对的跨具身监督,并允许策略直接在原始以外骨骼为中心的手腕图像上训练,无需分割或修复。在五个丰富接触任务上,SEED-UMI实现了比基于外骨骼的遥操作高3.0倍的数据收集效率,以及70.0%的平均执行成功率。
cs.RO / 34 / 2609.11766

Visual-SLAM for the detection of hidden tomatoes in greenhouses by Hierarchical Localization and GLOMAPfor robotized harvesting

基于层次定位和GLOMAP的面向机器人化收获的温室隐藏番茄检测的Visual-SLAM
Cañadas-Aránega, Fernando, Moreno, José C., Blanco-Claraco, José L., Rodríguez, Francisco
Abstract
Advanced crop monitoring inside greenhouses is becoming one of the primary objectives of research centers. High-performance sensors, such as LiDAR or stereo cameras, have traditionally been employed for this purpose, though these often have a high cost. This work proposes a Visual-SLAM system using a monocular camera, which is significantly more cost-effective and specifically tailored for agricultural applications, such as mapping tomato crops in a greenhouse. Tests were carried out on a real tomato bunch, located in the Agroconnect experimental greenhouse. A ROS 2 Humble node was developed to run on the robot in order to capture images of these crops, which were then stored for offline processing. To generate a 3D mapped model for the crop in the greenhouse, the GLOMAP mapper, based on Structure-From-Motion, was integrated with the Hierarchical Localization toolbox. This initial mapping is a foundation for future, more advanced algorithms to analyze growth patterns, and optimize agricultural management. The system leverages a hierarchical localization paradigm based on a coarse-to-fine strategy: it first performs global retrieval to generate location hypotheses, then combines local features within the identified candidate regions. The results show a correct identification of the tomato cluster, correctly characterising the tomato that is occluded and inaccessible by classical vision technologies. The reconstructed 3D model was further validated against manual ground-truth measurements of fruit size, centroid position, and orientation, confirming the geometric accuracy of the proposed low-cost monocular pipeline.
Chinese Translation
温室内的先进作物监测正成为研究中心的主要目标之一。高性能传感器,如LiDAR或立体相机,传统上被用于此目的,但通常成本高昂。本研究提出一种使用单目相机的Visual-SLAM系统,其成本效益显著更高,并专门针对农业应用(如温室番茄作物建图)定制。在Agroconnect实验温室中的真实番茄串上进行了测试。开发了一个ROS 2 Humble节点在机器人上运行,以捕获这些作物的图像,然后存储用于离线处理。为了生成温室作物的3D建图模型,将基于Structure-From-Motion的GLOMAP建图器与Hierarchical Localization工具箱集成。这种初始建图为未来更先进的算法分析生长模式和优化农业管理奠定了基础。该系统利用基于粗到精策略的层次定位范式:首先执行全局检索生成位置假设,然后在识别的候选区域内结合局部特征。结果表明,番茄簇被正确识别,并正确表征了被经典视觉技术遮挡且无法访问的番茄。重建的3D模型进一步通过手动真值测量(果实大小、质心位置和方向)进行验证,证实了所提出的低成本单目管道的几何精度。
cs.RO / 35 / 2609.11775

Rapid Learning of Dexterous In-Hand Pen Writing through Real-Time Jacobian Estimation

通过实时雅可比估计快速学习灵巧手内钢笔书写
Stewart, Kai, Toshimitsu, Yasunori, Katzschmann, Robert K.
Abstract
Dexterous in-hand manipulation of a grasped object with an anthropomorphic hand is an unsolved frontier for robot dexterity. The contact-richness and highly dynamic nature of object-hand interactions tend to require extensive modeling or data-collection efforts for learning-based approaches. Modern simulators used for reinforcement learning (RL) cannot fully replicate the required contact complexity, while collecting dexterous demonstrations for imitation learning (IL) remains an open problem. In this research, we present an embodied control approach based on real-time task Jacobian estimation of the combined hand and object system on the physical robot. Using only the CPU on a laptop, the proposed controller begins in-hand pen writing after approximately 18 s of initialization and continues to adapt online, without an analytic hand--object kinematic/contact model, simulation training, or precollected task demonstrations. We demonstrate that the same estimator/controller formulation works on three anthropomorphic robotic hand systems (one physical, two simulated) to show human-like, in-hand articulation of a grasped pen by an embodiment-independent formulation. Sub-millimeter in-plane precision (mean 0.6 mm across runs) is achieved across letters and shapes written in the air and on paper on a physical robot. To our knowledge, this is the first demonstration of an anthropomorphic hand writing arbitrary single-stroke trajectories with a grasped pen through purely in-hand motion, and it showcases an alternative to compute- and data-heavy approaches such as RL and IL for achieving dexterous manipulation through computationally simple and data-efficient algorithms.
Chinese Translation
使用拟人手对抓取物体进行灵巧的手内操作是机器人灵巧性尚未解决的前沿问题。物体-手交互的接触丰富性和高度动态特性往往需要大量建模或数据收集工作,才能用于基于学习的方法。用于强化学习(RL)的现代模拟器无法完全复现所需的接触复杂性,而为模仿学习(IL)收集灵巧演示仍是一个未解决的问题。在本研究中,我们提出了一种基于物理机器人上手与物体组合系统的实时任务雅可比估计的具身控制方法。所提出的控制器仅使用笔记本电脑的CPU,在约18秒初始化后开始手内钢笔书写,并持续在线自适应,无需解析的手-物体运动学/接触模型、模拟训练或预先收集的任务演示。我们证明,相同的估计器/控制器公式可用于三个拟人机械手系统(一个物理,两个模拟),以展示与具体实现无关的表述实现类似人类的手内抓握笔的关节运动。在物理机器人上,在空中和纸上书写的字母和形状实现了亚毫米级平面内精度(多次运行平均0.6毫米)。据我们所知,这是首次展示拟人手通过纯手内运动用抓握的笔书写任意单笔画轨迹,并展示了作为计算和数据密集型方法(如RL和IL)的替代方案,通过计算简单且数据高效的算法实现灵巧操作。
cs.RO / 36 / 2609.11871

Learning Agent-based Model Predictive Control for Holistic Vehicle Performance

面向整体车辆性能的学习型基于智能体的模型预测控制
Zhong, Jiaming, Mehrizi, Reza Valiollahi, Pirani, Mohammad, Yu, Chao, Kasaiezadeh, Alireza, Pant, Yash Vardhan, Khajepour, Amir
Abstract
Agent-based model predictive control (AMPC) has recently been proposed as a distributed scheme that collaborates with all agents to achieve optimal holistic performance. However, its optimality highly depends on the prediction accuracy that requires all agents or their contributions to be known, which is too idealistic for actual implementation. This research proposes a novel practical hybrid control scheme - learning agent-based MPC (LAMPC), combining the model-based AMPC approach and data-based learning methods to improve the holistic vehicle performance for multi-agent systems. The Gaussian process regression (GPR) enhanced by an online data management strategy serves as the learning core to predict unknown contributions. A novel multi-step prediction mechanism leverages the GPR learning potential along the horizon. The predicted mean, representing the learned unknown contributions, completes the system model in the MPC for more accurate control. Meanwhile, a stochastic framework is formulated to guarantee control safety and feasibility using soft chance constraints based on the prediction variance. Both simulations and experiments show that, with the learning capability, LAMPC outperforms the traditional AMPC. LAMPC can achieve higher tracking performance in well-learned scenarios and always guarantee constraint satisfaction even in less-learned scenarios. Moreover, the proposed hybrid control scheme is efficient for real-time implementation and is flexible to any control agent topology.
Chinese Translation
基于智能体的模型预测控制(AMPC)最近被提出作为一种分布式方案,与所有智能体协作以实现最优的整体性能。然而,其最优性高度依赖于预测精度,这要求所有智能体或其贡献已知,这对于实际实现来说过于理想化。本研究提出了一种新颖实用的混合控制方案——学习型基于智能体的MPC(LAMPC),将基于模型的AMPC方法与基于数据的学习方法相结合,以提高多智能体系统的整体车辆性能。通过在线数据管理策略增强的高斯过程回归(GPR)作为学习核心,用于预测未知贡献。一种新颖的多步预测机制利用了GPR在预测时域上的学习潜力。预测均值代表了学习到的未知贡献,补全了MPC中的系统模型,以实现更精确的控制。同时,建立了一个随机框架,基于预测方差使用软机会约束来保证控制的安全性和可行性。仿真和实验均表明,凭借学习能力,LAMPC优于传统的AMPC。LAMPC在充分学习的场景中能够实现更高的跟踪性能,并且即使在欠学习场景中也能始终保证约束满足。此外,所提出的混合控制方案对于实时实现是高效的,并且对任何控制智能体拓扑都具有灵活性。
cs.RO / 37 / 2609.11875

UniMPA: A Unified Memory-Prediction-Action Model via Action-Grounded Transition Modeling

UniMPA: 基于动作接地的转移建模的统一记忆-预测-动作模型
Li, Wei, Shao, Rui, He, Jie, Zhang, Lingsen, Liu, Ziwei, Nie, Liqiang
Abstract
Recent advances in Vision-Language-Action (VLA) models have improved robotic manipulation, yet observation-to-action learning remains limited by a fundamental transition realizability gap, manifested in three tightly coupled problems: (i) Transition ambiguity. Visually similar current observations may correspond to different manipulation phases and imply different subsequent transitions. (ii) Prediction--execution mismatch. A visually plausible predicted future observation does not necessarily correspond to a physically realizable transition. (iii) Experience--realization mismatch. A historically executable action pattern may not necessarily realize the intended transition in the current scene and therefore requires context-aware adaptation. Accordingly, we propose UniMPA, a Unified Memory-Prediction-Action model that addresses these problems through a shared action-grounded transition interface. (i) UniMPA introduces Persistent-Selective Future Prediction to resolve transition ambiguity by modeling the intended future state evolution. A persistent latent stream continuously tracks task-level progress, while a transition-critical pixel stream selectively resolves fine-grained interaction changes through memory-grounded prediction. (ii) To assess the physical executability of the anticipated transition, the predicted transition queries a temporal Visual-Action Memory Bank. The bank retrieves historically realized visual-action experience, grounding future prediction in executable evidence. (iii) To adapt executable experience to the current scene, an Action-Visual Memory Bank retrieves visually grounded action prototypes from historical action evolution. Prototype-Biased Flow then shifts the flow source toward a historically supported action manifold for context-aware refinement.
Chinese Translation
近期,Vision-Language-Action (VLA) 模型的进展提升了机器人操作能力,然而从观测到动作的学习仍然受到一个根本性的转移可实现性差距的限制,该差距表现为三个紧密耦合的问题:(i) 转移模糊性。视觉上相似的当前观测可能对应不同的操作阶段,并意味着不同的后续转移。(ii) 预测-执行不匹配。视觉上看似合理的预测未来观测不一定对应于物理上可实现的转移。(iii) 经验-实现不匹配。历史上可执行的动作模式在当前场景中不一定能实现预期的转移,因此需要情境感知的适应。因此,我们提出了 UniMPA,一个统一的记忆-预测-动作模型,通过一个共享的基于动作接地的转移接口来解决这些问题。(i) UniMPA 引入了持久-选择性未来预测,通过建模预期的未来状态演化来解决转移模糊性。一个持久的潜在流持续跟踪任务级进展,而一个转移关键的像素流通过基于记忆的预测选择性地解决细粒度的交互变化。(ii) 为了评估预期转移的物理可执行性,预测的转移查询一个时序视觉-动作记忆库。该库检索历史上已实现的视觉-动作经验,将未来预测建立在可执行的证据上。(iii) 为了将可执行经验适应当前场景,一个动作-视觉记忆库从历史动作演化中检索视觉接地的动作原型。然后,原型偏置流将流源推向历史支持的动作流形,以进行情境感知的细化。
cs.RO / 38 / 2609.11920

EVPeriscope: Extended Perception across Aerial and Ground Vehicles with Event-based Propeller Tracking

EVPeriscope:基于事件螺旋桨跟踪的空中与地面车辆扩展感知
Ong, Dexter, Kumar, Vijay, Chaudhari, Pratik
Abstract
Reliable relative localization between aerial and ground robots is a key requirement for tightly coordinated heterogeneous teams. This can be difficult to do using conventional frame-based cameras and fiducial markers because they are sensitive to motion blur, lighting variations, and payload constraints. This paper presents EVPeriscope, an event-based perception system that enables detection, localization and control of a quadrotor using an upward-facing event camera on a ground robot by detecting the high-frequency visual signature of its propellers. This system allows the quadrotor to function as an extended perception system for the ground robot when onboard sensors exhibit degradation or occlusion. We demonstrate the capabilities of this marsupial ground-aerial system via experiments in challenging field conditions with wind speeds of up to 15 mph, in both daylight and at night. We show that the system supports localization and closed-loop navigation through dense foliage where the ground robot's sensors are occluded. Our control system for the quadrotor operates at 200 Hz entirely with onboard sensing and computation. More details and experiment videos can be found on the project page: https://ongdexter.github.io/evperiscope.
Chinese Translation
空中和地面机器人之间的可靠相对定位是紧密协调的异构团队的关键要求。使用传统的基于帧的相机和基准标记可能难以实现,因为它们对运动模糊、光照变化和有效载荷限制敏感。本文提出 EVPeriscope,一种基于事件的感知系统,通过检测四旋翼螺旋桨的高频视觉特征,利用地面机器人上的朝上事件相机实现对四旋翼的检测、定位和控制。该系统允许四旋翼在地面机器人的机载传感器出现性能下降或遮挡时,充当其扩展感知系统。我们通过在风速高达 15 英里/小时的挑战性野外条件下(白天和夜间)进行实验,展示了这种袋鼠式地空系统的能力。我们表明,该系统支持在地面机器人传感器被遮挡的茂密植被中进行定位和闭环导航。我们的四旋翼控制系统以 200 Hz 运行,完全依赖机载传感和计算。更多细节和实验视频请访问项目页面:https://ongdexter.github.io/evperiscope。