← Back to Index
Daily Research Digest

arXiv Papers

2026-09-09
578
Papers
3
Categories
578
Translated
收藏清单 0
人工智能 (Artificial Intelligence)
197
cs.AI / 1 / 2609.05437

Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

超越对与错:评估大语言模型中的二阶社会推理能力
Rai, Sunny, Kuang, Jinyi, Jamalova, Reyhan, Lou, Annie, Bicchieri, Cristina, Malhotra, Niyati, Orozco-Olvera, Victor Hugo, Munoz-Boudet, Ana Maria, Ungar, Lyle H, Guntuku, Sharath C
Abstract
Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators, and other-regulation in observers. We release a multi-perspective dataset, NormReact, of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses across norm violators' gender and observers' social closeness. Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases. These findings suggest that AI systems in norm-sensitive domains from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over-represents punishment and under-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in real world.
Chinese Translation
以往的人工智能对齐研究主要聚焦于一阶社会规范,即教会模型哪些行为在社会上是可接受或不可接受的(例如“不要偷窃”)。然而,社会智能不仅取决于对规范的识别,还取决于预判谁将执行该规范以及如何执行(例如公开羞辱甚至监禁)。这类二阶预期被称为元规范(metanorms),它支配着人们在社会规则被打破时的反应方式。我们提出了一个新颖的框架,从两个维度评估大语言模型(LLMs)的元规范推理能力:情绪评价和行为反应,并提出新的分类任务,即预测违规者的自我规制和旁观者的他者规制。我们发布了多视角数据集NormReact,包含450个规范违规场景,并针对违规者性别和旁观者社会亲疏程度,对情绪和行为反应进行了人工标注。当前的大语言模型描绘了一个更为严酷的社会图景:在六个模型中,它们过度预测负面制裁,而人类在这些情境下会预期不作为;且随着社会距离的增加,与人类判断的一致性也随之下降。这些发现表明,从冲突调解到政策模拟等规范敏感领域的人工智能系统,可能存在呈现扭曲的社会规制图景的风险:过度呈现惩罚,而低估了现实世界中实际规范执行所体现的宽容、克制与关系性调适。
cs.AI / 2 / 2609.05439

CriticGen: Generation-Aware Evaluation as Actionable Feedback

CriticGen:面向生成的可操作反馈式评估方法
Du, Huifang, Zuo, Zecheng, Wang, Sen, Fan, Chenghao, Wang, Haofen, Yang, Yehui
Abstract
Current evaluation methods for large language models are coarse-grained and decoupled from generation, producing generic explanations that fail to provide actionable feedback for model improvement. We propose CriticGen, a fine-grained, generation-aware evaluation framework that turns evaluation into actionable control for answer improvement. CriticGen first generates sample-specific evaluation dimensions and scoring criteria under high-level categories such as subjective, objective, and self-derived constraints. These criteria then serve as a dynamic rubric for jointly producing a score, a reason, an executable refinement suggestion, and a refined answer. This rubric-conditioned refinement process enables models to diagnose flaws and perform targeted answer improvement. Experimental results show that fine-grained evaluation should be both instance-specific and actionable. CriticGen induces higher-quality rubrics, improving relevance/coverage from 3.33/4.03 to 3.97/4.24. CriticGen also achieves the best score correlations, with 0.9556 Pearson and 0.9560 Spearman, and raises the F1 of criterion-grounded reasons and executable suggestions from 0.6369/0.5994 to 0.7554/0.7900. Crucially, its feedback translates into reliable answer improvement, improving 73.17% of answers with a 93.28% non-degradation rate.
Chinese Translation
当前针对大语言模型的评估方法粒度粗糙,且与生成过程相分离,产生的泛化性解释无法为模型改进提供可操作的反馈。我们提出CriticGen,一个细粒度、面向生成的评估框架,将评估转化为可用于答案改进的可操作控制。CriticGen首先在主观、客观和自定义约束等高层类别下生成针对具体样本的评估维度和评分标准。随后,这些标准作为动态评分规则(rubric),用于共同产出分数、评分理由、可执行的改进建议以及改进后的答案。这种基于评分规则的改进过程使模型能够诊断缺陷并进行有针对性的答案改进。实验结果表明,细粒度评估应当兼具样本特异性和可操作性。CriticGen能够诱导出更高质量的评分规则,将相关性/覆盖度从3.33/4.03提升至3.97/4.24。CriticGen还取得了最佳的分数相关性,Pearson相关系数达0.9556,Spearman相关系数达0.9560,并将基于评分标准的理由和可执行建议的F1值从0.6369/0.5994提升至0.7554/0.7900。至关重要的是,其反馈能够可靠地转化为答案的改进,改进了73.17%的答案,且非退化率高达93.28%。
cs.AI / 3 / 2609.05441

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

记忆何时有用?面向工具使用型LLM智能体的长期记忆成本感知评估
Mishra, Shweta, Mishra, Shashank
Abstract
Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history, not whether remembered facts change what a tool-using agent does. We present MERIT (Memory Evaluation for Realistic Instrumented Tasks), a benchmark and harness that measures the marginal utility of memory for task-executing agents under explicit cost accounting. MERIT provides episodic tool-use tasks in three domains whose dependence on earlier-episode facts is verified by an automated leak check; a difficulty ladder ending in updated-fact recall; controlled memory corruption; and full token and dollar metering of every memory operation. Across 23,440 scored episodes ($42.57), a two-generation pilot on gpt-4.1-mini and a preregistered 3-model x 3-seed grid (GPT-4.1, Claude Haiku 4.5; memory side held fixed), memory lifts dependent-task success from a leak-verified floor of 0.00 to 0.55-1.00. On updated facts, embedding retrieval collapses unpredictably (0.30-0.95 across models; max seed gap 0.45), and agents act on a correctly retrieved value only 55% of the time, while update-on-write stores (a structured fact store and, notably, LLM summarization) remain at 0.70-1.00; the hybrid is worse than the fact store alone. A latest-generation spot-check (Claude Sonnet 5, gated on a clean full-replay control) reproduces the pattern. Swapping a memory's implementation moves task success by up to 60 points, and full replay is never economical: the best condition per domain delivers 2.7-3.9x its marginal utility per dollar. We release the benchmark, harness, and all traces.
Chinese Translation
目前对LLM智能体长期记忆的评估主要依赖对话回忆类基准(LoCoMo、LongMemEval),它们衡量的是对对话历史的问题回答能力,而非被记住的事实是否会改变工具使用型智能体的实际行为。我们提出了MERIT(Memory Evaluation for Realistic Instrumented Tasks,面向真实仪器化任务的记忆评估),这是一个在显式成本核算下衡量记忆对任务执行型智能体边际效用的基准和评测框架。MERIT提供了三个领域的情景式工具使用任务,其对早期情景事实的依赖性通过自动化泄露检测加以验证;包含一个以更新事实回忆为终点的难度阶梯;支持受控的记忆损坏注入;并对每项记忆操作进行完整的token和美元计量。在23,440个评分回合(花费42.57美元)上,包括在gpt-4.1-mini上的两代试点实验以及一项预注册的3模型×3种子网格实验(GPT-4.1、Claude Haiku 4.5;记忆侧保持不变),结果表明记忆将依赖性任务的成功率从经泄露验证的0.00基准提升至0.55-1.00。对于已更新的事实,嵌入检索会出现不可预测的崩溃(跨模型为0.30-0.95;种子间最大差距为0.45),且智能体仅有55%的几率对正确检索到的值采取行动;而“写入时更新”的存储方式(结构化事实存储,以及值得注意的是LLM摘要)保持在0.70-1.00;混合方案反而比单独使用事实存储更差。针对最新一代模型的抽查(Claude Sonnet 5,并以干净的全量回放对照作为门控)复现了这一模式。更换记忆的实现方式可使任务成功率最多变动60个百分点,且全量回放从不经济:每个领域的最优条件可提供每美元2.7-3.9倍的边际效用。我们公开发布了该基准、评测框架及所有轨迹数据。
cs.AI / 4 / 2609.05446

AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents

AutoFyn 技术报告:面向长时程智能体的非参数专家迭代方法
Hasan, Adib, Schaffield, Daniel, Dutta, Akashnil, Moon, Tarik Adnan
Abstract
We introduce AutoFyn, an agent harness inspired by the Expert Iteration algorithm, adapting a frozen model across many rounds by updating persistent state from verified reward signals rather than model weights. Each round begins from a fresh model session, and durable information is reintroduced only through explicit interfaces such as persistent memory files, reports, and repository state. Within a round, an orchestrator explores, plans and builds many alternative approaches with specialized agents, while a task-grounded verifier verifies the work and supplies an objective reward for measuring progress. This reward is distilled back into the persistent state, which updates the effective policy for the next round. In this technical report, we formalize this loop and describe its persistent state and verification interfaces. We then demonstrate its use in three domains, namely olympiad mathematics, data science, and cybersecurity. On the six fresh problems of the 2026 International Mathematical Olympiad, every model with room to improve scores higher under AutoFyn than in its provider's own coding agent. AutoFyn also built the top-ranked agent on the Spider 2.0 dbt benchmark, and has produced $16$ maintainer-confirmed vulnerability advisories in Next.js, MetaMask, pnpm, Warp, LiteLLM, Langflow, and Open WebUI.
Chinese Translation
我们提出 AutoFyn,一种受专家迭代(Expert Iteration)算法启发的智能体框架,它通过利用经验证奖励信号更新持久化状态(而非模型权重),使冻结模型在多轮迭代中持续改进。每一轮均从一个全新的模型会话开始,持久化信息仅通过显式接口(如持久化记忆文件、报告和代码仓库状态)重新引入。在单轮内,一个编排器(orchestrator)通过多个专用智能体进行探索、规划并构建多种备选方案,同时一个基于任务的验证器对工作成果进行核验,并给出用于衡量进展的客观奖励。该奖励随后被蒸馏回持久化状态中,从而更新下一轮的有效策略。在本技术报告中,我们对该循环进行了形式化描述,并阐述了其持久化状态与验证接口。随后,我们在三个领域展示了其应用,即奥数竞赛、数据科学和网络安全。在2026年国际数学奥林匹克竞赛(IMO)的六道全新题目上,每个具备提升空间的模型在 AutoFyn 下的得分均高于其提供商自带的编程智能体。AutoFyn 还构建了 Spider 2.0 dbt 基准上排名第一的智能体,并在 Next.js、MetaMask、pnpm、Warp、LiteLLM、Langflow 和 Open WebUI 中发现了 $16$ 个经维护者确认的漏洞通告。
cs.AI / 5 / 2609.05448

Damage-Aware Bandit Pruning for Vision and Language Transformers

面向视觉与语言Transformer的损伤感知赌博机剪枝方法
Ameen, Salem, Vadera, Sunil
Abstract
Structured post-training pruning of transformers requires selecting complete functional units whose suppression causes limited degradation. We formulate structured-unit selection for language and vision transformers as a damage-aware multi-armed bandit problem under a fixed candidate-evaluation budget. Attention heads and MLP channel groups are temporarily masked on calibration batches. Paired damage is the masked loss minus the base loss on the same batch, reducing batch-to-batch variation. A smooth bounded reward drives either a UCB-style policy or fractional-Beta Thompson Sampling, and the final mask is constructed sequentially by adding one unit at each step. The selected units are functionally zeroed in the original dense checkpoint; therefore, the reported parameter effects represent effective structural suppression rather than physical compression or measured speedup. Experiments on WikiText-2, LAMBADA, and Imagenette cover GPT-2, OPT, Pythia, Qwen2.5, SmolLM2, ViT-B/16, DeiT-Tiny, and Swin-Tiny, with comparisons against random, magnitude, static-saliency, and budgeted-greedy selection. Across five seeds, the bandit methods usually reduce degradation relative to budgeted greedy in the paired language-model comparisons. Of 28 comparisons highlighted in the paper, 23 bootstrap confidence intervals exclude zero and 11 paired tests have p < 0.05; six have q < 0.05 after Benjamini-Hochberg correction across the full family of 116 dataset-wise tests. Matched-evaluation results for ViT-B/16 and Swin-Tiny indicate that their gains are not explained solely by a larger candidate-evaluation budget.
Chinese Translation
Transformer的结构化训练后剪枝需要选择完整的 функциональных功能单元,使其被抑制后造成的性能退化有限。我们将语言与视觉Transformer的结构化单元选择问题形式化为在固定候选评估预算下的损伤感知多臂老虎机(multi-armed bandit)问题。注意力头和MLP通道组在校准批次上被临时掩蔽。配对损伤定义为同一批次上掩蔽后的损失减去基线损失,从而降低批次间的波动。一个平滑有界的奖励信号驱动UCB风格策略或分数Beta汤普森采样(fractional-Beta Thompson Sampling),最终的掩蔽通过每步添加一个单元的方式顺序构建。被选中的单元在原始稠密检查点中被功能性地置零;因此,所报告的参数效应代表有效的结构性抑制,而非物理压缩或实测加速。实验在WikiText-2、LAMBADA和Imagenette上涵盖了GPT-2、OPT、Pythia、Qwen2.5、SmolLM2、ViT-B/16、DeiT-Tiny和Swin-Tiny,并与随机、幅值、静态显著性和预算贪心选择方法进行了对比。在五个随机种子下,在配对的语言模型对比中,赌博机方法通常比预算贪心方法减少性能退化。在论文重点展示的28项对比中,23项的自助法(bootstrap)置信区间不包含零,11项配对检验的p值小于0.05;在全部116项数据集级检验经Benjamini-Hochberg校正后,6项的q值小于0.05。针对ViT-B/16和Swin-Tiny的匹配评估结果表明,其性能提升并非仅由更大的候选评估预算所解释。
cs.AI / 6 / 2609.05459

Compiling VGDL into Causal Models

将VGDL编译为因果模型
Jiwatode, Mohit, Rosenhahn, Bodo, Dockhorn, Alexander
Abstract
Reinforcement learning and large language models often struggle to accurately capture the causal mechanics of game environments. Standard reinforcement learning agents tend to rely on spurious correlations, while large language models are prone to hallucinating game rules. Although causal reinforcement learning improves interpretability, there is currently no formal methodology to map complex game mechanics directly into causal models. To address this, we propose a deterministic framework that compiles games specified in the Video Game Description Language into Dynamic Structural Causal Models. Rather than inferring causal structures from gameplay traces or noisy large language models' outputs, our methodology directly translates game components, including sprite dynamics, interaction rules, and termination conditions, into explicit structural equations. Each game tick represents a causal transition from state variables at time $t$ to $t+1$. By establishing this grounded mapping, the approach guarantees absolute causal fidelity to the ground-truth game mechanics. The resulting models offer transparent causal pathways that support counterfactual reasoning, causal reinforcement learning agent training, and procedural content validation. This framework provides a principled bridge between symbolic game descriptions and causally grounded game AI.
Chinese Translation
强化学习和大语言模型往往难以准确捕捉游戏环境的因果机制。标准的强化学习智能体倾向于依赖虚假相关性,而大语言模型则容易产生游戏规则的幻觉。尽管因果强化学习提高了可解释性,但目前尚缺乏将复杂游戏机制直接映射到因果模型的形式化方法。为解决这一问题,我们提出一个确定性框架,将以视频游戏描述语言指定的游戏编译为动态结构因果模型。我们的方法不是从游戏过程轨迹或有噪声的大语言模型输出中推断因果结构,而是将游戏组件——包括精灵动态、交互规则和终止条件——直接翻译为显式的结构方程。每个游戏时钟步代表一个从时间 $t$ 到 $t+1$ 状态变量的因果转换。通过建立这种有依据的映射,该方法保证了对真实游戏机制的绝对因果保真度。由此得到的模型提供了透明的因果路径,支持反事实推理、因果强化学习智能体训练以及程序化内容验证。该框架在符号化的游戏描述与具有因果依据的游戏人工智能之间架起了一座原理性的桥梁。
cs.AI / 7 / 2609.05461

ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models

ARC-Bench:闭环重规划掩盖了冻结 JEPA 世界模型中失效的动作排序能力
Zhang, Zhengshu, Li, Zhiyuan
Abstract
Reward-free latent world models plan by scoring candidate actions with distances in a frozen latent space: an action is preferred if its predicted future embedding lands closer to the goal embedding. This silently assumes that latent closeness is action-rankable, i.e., that ordering candidates by latent distance agrees with ordering them by true cost. We audit this assumption directly. We introduce ARC-Bench, a no-leak, fixed-candidate protocol that measures whether frozen JEPA-style objectives rank candidate actions correctly, and apply it to official released JEPA-WM checkpoints across navigation and manipulation-style control. The assumption fails, severely and structurally: on the official manipulation audits the top-scored candidate is almost always suboptimal, and the same inversion appears in the maze domains. A controlled visual-backbone extension shows that the defect persists when DINOv2 is replaced by video-pretrained V-JEPA 1 and V-JEPA 2 encoders at ViT-L/ViT-G scale. Provenance, undertraining, matched-budget backbone controls, and metric-circularity controls rule out trivial explanations. We then explain why this defect has stayed invisible: closed-loop replanning masks it. When we reduce the planner's replanning frequency, success collapses in both a navigation and a manipulation domain, and the episodes rescued by frequent replanning are enriched for severe first-plan ranking failures in the PointMaze first-plan diagnostic. Closed-loop success rates therefore systematically overstate the rankability of frozen latent representations. ARC-Bench supplies the measurement, and the masking mechanism the explanation, for methods that adapt, amortize, or replan around latent-space planners without directly auditing released JEPA-WM action rankability.
Chinese Translation
无奖励的潜在世界模型通过在冻结的潜在空间中计算距离来对候选动作打分以进行规划:若某个动作预测的未来嵌入更接近目标嵌入,则该动作被优先选择。这隐含地假设了潜在接近度具备动作可排序性,即按潜在距离对候选动作排序与按真实代价排序是一致的。我们直接审计了这一假设。我们提出 ARC-Bench,这是一种无泄漏、固定候选集的评估协议,用于衡量冻结的 JEPA 类目标是否能正确地对候选动作进行排序,并将其应用于官方发布的 JEPA-WM 检查点,涵盖导航和类操作控制任务。该假设严重且结构性地失效了:在官方的操作任务审计中,得分最高的候选动作几乎总是次优的,同样的排序反转现象也出现在迷宫域中。一项受控的视觉骨干网络扩展实验表明,当 DINOv2 被替换为在 ViT-L/ViT-G 规模上经视频预训练的 V-JEPA 1 和 V-JEPA 2 编码器时,该缺陷依然存在。来源审计、欠训练检验、同等预算的骨干网络对照以及度量循环性对照排除了各种平庸解释。随后,我们解释了为何该缺陷一直未被察觉:闭环重规划掩盖了它。当我们降低规划器的重规划频率时,一个导航域和一个操作域中的成功率均大幅崩溃,而在 PointMaze 首次规划诊断中,靠频繁重规划得以挽救的回合中严重首 plan 排序失败的比例显著偏高。因此,闭环成功率系统性地高估了冻结潜在表示的可排序性。ARC-Bench 为那些在潜在空间规划器基础上进行自适应、摊销或重规划,而未直接审计已发布 JEPA-WM 动作可排序性的方法,提供了测量手段,而掩盖机制则提供了成因解释。
cs.AI / 8 / 2609.05481

RAPID: Reliability-Aware Pair Importance Distillation

RAPID:可靠性感知的样本对重要性蒸馏
Mahdavi, Ali, Zamanifar, Azadeh, Farhadi, Amirfarhad, Kashefi, Omid
Abstract
Inter example relational distillation transfers a teacher's representation geometry by matching relations among examples within a mini batch. Computing all pairs has quadratic complexity in the batch size, whereas uniform subsampling may use a limited relation budget inefficiently. We introduce Reliability Aware Pair Importance Distillation, or RAPID, which separates a reliability gated relational target from a full support adaptive pair proposal. Reliability determines which teacher relations are emphasized, while calibrated teacher entropy and detached student-teacher residuals determine which relations are evaluated. Exact inverse proposal correction makes the loss and gradient estimators conditionally unbiased with respect to the gated mini batch target. We evaluate RAPID in two text classification settings: AG News with BERT-to-DistilBERT distillation using three paired seeds and a relation budget of 256, and SST-2 with DistilBERT to DistilBERT distillation using three paired seeds and a relation budget of 64. Reliability gated relational distillation achieves the highest observed mean student accuracy on both datasets: 94.285 plus or minus 0.054 percent on AG News and 88.800 plus or minus 0.532 percent on SST-2. RAPID ranks second, achieving 94.241 plus or minus 0.025 percent and 88.685 plus or minus 0.462 percent, respectively, compared with 94.154 plus or minus 0.124 percent and 87.271 plus or minus 0.162 percent for the cross entropy baseline. Pilot evaluations are counted toward the same total budget as the main relation evaluations. Across both settings, the gated target yields the highest mean accuracy, while the adaptive proposal remains within seed-level variation. These results support the modular view that target reliability and evaluation priority are separable design dimensions.
Chinese Translation
样本间关系蒸馏(inter example relational distillation)通过匹配小批量内样本之间的关系来迁移教师模型的表示几何结构。计算所有样本对具有关于批大小的二次复杂度,而均匀子采样可能无法高效利用有限的关系预算。我们提出可靠性感知的样本对重要性蒸馏方法(Reliability Aware Pair Importance Distillation,RAPID),它将可靠性门控的关系目标与全支撑的自适应样本对提议分离。可靠性决定强调哪些教师关系,而经校准的教师熵和分离的学生-教师残差决定评估哪些关系。精确的逆提议修正使损失和梯度估计器相对于门控小批量目标具有条件无偏性。我们在两个文本分类场景中评估了RAPID:一是使用BERT到DistilBERT蒸馏的AG News任务,采用三组配对随机种子和256的关系预算;二是使用DistilBERT到DistilBERT蒸馏的SST-2任务,采用三组配对随机种子和64的关系预算。可靠性门控关系蒸馏在两个数据集上均取得了最高的观测平均学生准确率:AG News上为94.285±0.054%,SST-2上为88.800±0.532%。RAPID排名第二,分别达到94.241±0.025%和88.685±0.462%,而交叉熵基线分别为94.154±0.124%和87.271±0.162%。试点评估被计入与主关系评估相同的总预算。在两种设置中,门控目标均产生最高的平均准确率,而自适应提议保持在种子间波动范围之内。这些结果支持了模块化观点,即目标可靠性与评估优先级是可分离的设计维度。
cs.AI / 9 / 2609.05488

PGP-Clinical-TimeKAN: Prior-Guided Joint Probabilistic Forecasting of Clinical Trajectories

PGP-Clinical-TimeKAN:先验引导的临床轨迹联合概率预测
Nie, Weizhi, Chang, Rihao, Wang, Weijie, Su, Yuting
Abstract
Clinical deterioration unfolds through coupled, partially observed trajectories, not a single diagnostic label. We introduce PGP-Clinical-TimeKAN, a trajectory-first framework for joint probabilistic forecasting of multivariate physiology. It combines missingness-aware temporal encoders, a soft organ-system prior, patient-specific relations, nonlinear Kolmogorov-Arnold messages, and a low-rank multivariate Student-t head. We evaluate 24-hour histories and six-hour forecasts on a frozen MIMIC-IV-derived cohort of 6,882 patients and 54,694 windows. Across five seeds and 13 models, PGP-Clinical-TimeKAN obtains the second-lowest normalized MAE (0.37727 +/- 0.00029) and the lowest RMSE (0.52656 +/- 0.00034). It reduces MAE by 0.52% relative to deterministic TimeKAN. For probabilistic forecasting, it reaches a marginal NLL of 0.66380 and a CRPS of 0.27301. Empirical coverage is 0.533, 0.831, and 0.958 for nominal 50%, 80%, and 95% intervals. Removing relational structure causes the largest ablation loss. Increasing covariance rank improves joint likelihood but has little effect on point accuracy. A trajectory-derived risk score remains weaker than a dedicated GRU-D classifier (AUROC 0.603 versus 0.650), which limits the present clinical claim. Joint trajectory forecasting therefore provides an inspectable intermediate task, but accurate physiology forecasts alone do not ensure a calibrated event detector.
Chinese Translation
临床恶化通过相互耦合的、部分可观测的轨迹演变,而非单一的诊断标签。我们提出了PGP-Clinical-TimeKAN,一个以轨迹为先的多变量生理指标联合概率预测框架。该框架结合了缺失感知的时间编码器、软性器官系统先验、患者特异性关系、非线性Kolmogorov-Arnold消息传递机制,以及低秩多变量Student-t输出头。我们在一个基于MIMIC-IV构建的冻结队列(包含6,882名患者和54,694个时间窗口)上,评估了24小时病史的输入和6小时的预测。在五种随机种子和13个模型的对比中,PGP-Clinical-TimeKAN取得了第二低的归一化MAE(0.37727 +/- 0.00029)和最低的RMSE(0.52656 +/- 0.00034)。相对于确定性的TimeKAN,其MAE降低了0.52%。在概率预测方面,其边际NLL为0.66380,CRPS为0.27301。在标称50%、80%和95%的置信区间下,经验覆盖率分别为0.533、0.831和0.958。消融实验表明,移除关系结构导致的性能损失最大;增加协方差秩可以提升联合似然,但对点预测精度影响甚微。基于轨迹的风险评分仍弱于专门的GRU-D分类器(AUROC为0.603对0.650),这限制了其当前的临床应用主张。因此,联合轨迹预测提供了一个可检验的中间任务,但仅凭准确的生理指标预测并不能保证一个校准良好的事件检测器。
cs.AI / 10 / 2609.05505

SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews

SciLitBench:面向大语言模型驱动系统性文献综述的基准测试与设计原则
Zabaleta, Miguel, Lin, Baihan
Abstract
Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening, full-text screening, and schema-guided data extraction, with 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. Across 22 open-weight LLMs from six model families, explicit inclusion and exclusion criteria improve title and abstract screening $F_2$ by 28.8\%, while researcher-authored rationales improve full-text screening by 15\%. Data extraction reveals a different reliability regime: performance declines from 0.97 accuracy for publication year to 0.37 Jaccard overlap for computational approach, while the strongest models recover only 30\% of annotated evaluation evidence and 25\% of limitations. SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.
Chinese Translation
系统性文献综述需要对数千条文献记录进行持续的人工判断,然而现有对大语言模型(LLM)的评估通常孤立地考察综述的各个阶段。我们提出了 SciLitBench,一个覆盖标题与摘要筛选、全文筛选以及基于模式(schema)引导的数据抽取的多阶段基准,包含 42,981 条检索记录、1,012 篇全文以及 888 篇纳入论文的标注数据。在来自六个模型家族的 22 个开源权重大语言模型上的实验表明,明确的纳入与排除标准可将标题与摘要筛选的 $F_2$ 分数提升 28.8%,而研究者撰写的筛选理由可将全文筛选性能提升 15%。数据抽取则呈现出不同的可靠性水平:从出版年份 0.97 的准确率到计算方法仅 0.37 的 Jaccard 重叠度,性能逐级下降,即便最强的模型也只能恢复 30% 的标注评估证据和 25% 的局限性说明。SciLitBench 划定了高召回率筛选与证据完整抽取之间的实际边界,并为评估大语言模型辅助的证据综合提供了可复现的资源。
cs.AI / 11 / 2609.05511

SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction

SCAFFOLD:通过递归参数化技能抽象实现自我改进的Web智能体
He, Bowei, Zhang, Xiaokun, Ding, Meng, Liu, Xue
Abstract
Web agents need to navigate visually rich, long-horizon interfaces that change across sites, yet most previous agents still learn each task in isolation and discard the procedural knowledge they accumulate. Recent skill-augmented frameworks take an important first step, but they treat the skill library as a flat or two-tier prompt-side cache and offer no principled mechanism for compressing redundancy or composing skills recursively. We introduce \textsc{Scaffold}, a self-improving framework for visual web agents that (i) induces parametric, executable skills from successful trajectories under a multi-instance abstraction constraint, (ii) maintains a recursively composed hierarchy in which higher-level skills invoke lower-level ones, (iii) compacts the library via a minimum-description-length (MDL) criterion and behavioral equivalence checking, and (iv) periodically distills skill-augmented trajectories back into model weights to internalize the abstractions. Across WebArena, VisualWebArena, and a held-out split of Online-Mind2Web, \textsc{Scaffold} improves success rate by $11.1$--$17.2$ absolute points over the strongest skill-augmented baseline and shows monotonic gains across five self-improvement iterations without library collapse. We release the code and documents in the Github \href{https://github.com/BokwaiHo/SCAFFOLD}{repository}.
Chinese Translation
Web智能体需要在视觉信息丰富、长时程且跨站点不断变化的界面中进行导航,然而以往的大多数智能体仍在孤立地学习每个任务,并丢弃其积累的程序性知识。近期的技能增强框架迈出了重要的第一步,但它们将技能库视为扁平的或两层的提示端缓存,缺乏压缩冗余或递归组合技能的原则性机制。我们提出了SCAFFOLD,一个面向视觉Web智能体的自我改进框架,其特点是:(i) 在多实例抽象约束下从成功轨迹中归纳出参数化的可执行技能;(ii) 维护一个递归组合的层级结构,其中高层技能调用低层技能;(iii) 通过最小描述长度(MDL)准则和行为等价性检查来压缩技能库;(iv) 定期将技能增强的轨迹蒸馏回模型权重以内化这些抽象。在WebArena、VisualWebArena以及Online-Mind2Web的留出测试集上,SCAFFOLD相较于最强的技能增强基线,成功率绝对提升了11.1至17.2个百分点,并在五次自我改进迭代中展现出单调递增的收益,且未出现技能库坍塌。我们在Github仓库(https://github.com/BokwaiHo/SCAFFOLD)中发布了代码和文档。
cs.AI / 12 / 2609.05512

Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment

推理感知压缩:识别并保护易受损的推理回路以实现高能效的大语言模型部署
Twagirayezu, Leonard, Mitra, Prasenjit
Abstract
Large Reasoning Models (LRMs) impose substantial energy costs during deployment, yet current compression methods apply uniform quantization across all components, risking damage to critical reasoning circuits. We present a reasoning-aware compression framework that benchmarks quantization conditions across five reasoning benchmarks, GSM8K, FOLIO, MATH-500, ProofWriter, and MuSiQue, with hardware-level GPU energy measurement; profiles per-module INT4 vulnerability across all 196-224 (layer, projection) pairs via a perturbation sweep on a held-out calibration split, then selectively restores the most sensitive circuits to FP16. Three findings emerge. First, INT4 quantization can increase energy by extending reasoning chains; a 25% power reduction becomes a net energy increase on GSM8K. Second, vulnerability is task-dependent: attention projections are more critical for mathematical reasoning, and sensitivity patterns differ by architecture in logical inference. Third, selective compression achieves Pareto-optimal points inaccessible to uniform methods: R1-Qwen-7B Top-10% on ProofWriter gains +12 pp over FP16 at -9.7% energy, validated on held-out data across five reasoning benchmarks.
Chinese Translation
大型推理模型(LRM)在部署过程中带来高昂的能耗成本,然而当前的压缩方法对所有组件采用统一的量化策略,存在损害关键推理回路的风险。我们提出了一种推理感知的压缩框架,该框架在五个推理基准(GSM8K、FOLIO、MATH-500、ProofWriter 和 MuSiQue)上对量化条件进行基准测试,并进行硬件级别的 GPU 能耗测量;通过在留出的校准数据集上进行扰动扫描,对所有 196-224 个(层,投影)组合进行逐模块的 INT4 脆弱性分析,随后选择性地将最敏感的回路恢复至 FP16 精度。我们得到三项发现。第一,INT4 量化可能因延长推理链而增加能耗:在 GSM8K 上,25% 的功率降低最终反而导致净能耗增加。第二,脆弱性具有任务依赖性:注意力投影对数学推理更为关键,而在逻辑推理任务中,敏感性模式因模型架构而异。第三,选择性压缩能够达到统一量化方法无法实现的帕累托最优点:R1-Qwen-7B 在 ProofWriter 上选取最敏感的 Top-10% 回路恢复精度后,相比 FP16 在能耗降低 9.7% 的同时准确率提升 12 个百分点,并在五个推理基准的留出数据上得到验证。
cs.AI / 13 / 2609.05513

When and What to Teach: Budget-Aware Online Adaptation for Web Agents

何时教与教什么:面向Web智能体的预算感知在线自适应方法
Zhang, Jianwei, Cao, Sihan, Zheng, Pengcheng, Wen, Ya, Ke, Pei, Liu, Kuien, Gao, Shen, Dong, Wei, Yang, Yang, Zhang, Chaoning
Abstract
Web agents have achieved significant success in automating complex internet tasks but deploying them in real-world environments requires continuous online adaptation. Given that deploying powerful proprietary models remains commercially cost-prohibitive, practitioners must rely on lightweight local models that evolve post-deployment via online teaching from a stronger teacher. However, standard interactive feedback imposes prohibitive costs. We show that conventional trajectory-level preference optimization wastes budget on both unresolvable episodes and redundant execution turns. To resolve these inefficiencies, we propose \textbf{Score-Guided Online Teaching with Budgeted Trajectory Trimming}, a budget-aware framework that systematically orchestrates \textbf{when} and \textbf{what} to teach. Specifically, our framework integrates a solvability-aware teacher gate to dictate \textbf{when} to query the teacher model and a score-guided turn selection mechanism to decide \textbf{what} informative turns to retain. Extensive experiments on MiniWoB and TimeWarp demonstrate that our method achieves comparable first-pass success while reducing teacher calls by 22.6\% and student training compute by 52.1\% on average. Our code is available at https://github.com/zjw131f1fc/budgeted-online-teaching.
Chinese Translation
Web智能体在自动化复杂的互联网任务方面已取得显著成功,但将其部署到真实环境中需要持续的在线自适应。鉴于部署强大的专有模型在商业上成本过高,从业者必须依赖轻量级本地模型,并通过更强大教师模型的在线教学在部署后不断演进。然而,标准的交互式反馈会带来高昂的成本。我们证明,传统的轨迹级偏好优化会在无法解决的回合(episode)和冗余的执行轮次上浪费预算。为解决这些低效问题,我们提出了基于预算的轨迹修剪的分数引导在线教学(Score-Guided Online Teaching with Budgeted Trajectory Trimming),这是一个预算感知的框架,可系统地统筹教学的时机(when)与内容(what)。具体而言,我们的框架集成了一个可解性感知的教师门控机制来决定何时查询教师模型,并采用分数引导的轮次选择机制来决定保留哪些信息量大的轮次。在MiniWoB和TimeWarp上的大量实验表明,我们的方法在取得相当的首轮成功率的同时,平均减少了22.6%的教师调用次数和52.1%的学生模型训练计算量。我们的代码已发布于 https://github.com/zjw131f1fc/budgeted-online-teaching。
cs.AI / 14 / 2609.05514

The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies

失败先于漂移:LLM智能体社会中价值观的社会动态
Atif, Farah, Saha, Sougata, Choudhury, Monojit
Abstract
Large Language Model (LLM)-based agents are increasingly used as proxies for human participants in social science research, yet it remains unclear whether they can faithfully simulate diverse and conflicting human value systems. We present a World Values Survey (WVS)-grounded simulation framework where culturally diverse agents with different communication styles engage in longitudinal, value-laden discussions. Across approximately 4,000 conversations involving 1,200 personas, 15 topics, and three models (GPT-4o, Gemini-2.5-Flash, and Gemma-4-E4B), we evaluate value faithfulness, value drift, and conversational realism. We find that more than 50\% of personas fail to express their assigned WVS profiles from the outset, while 2-7\% drift after repeated conversations. Ablations removing demographic details improve faithfulness for some models but do not change the broader trend: simulated value distributions still systematically deviate from the assigned WVS profiles. Compared to human discussions, simulated dialogues show a different trade-off between stylistic consistency and semantic diversity, often producing content-wise varied but stylistically repetitive exchanges. These findings suggest that current LLM agents can generate plausible conversations, but remain limited proxies for representing and preserving diverse human value profiles over time.
Chinese Translation
基于大语言模型(LLM)的智能体在社会科学研究中被越来越多地用作人类参与者的替代者,但它们能否忠实地模拟多样且相互冲突的人类价值体系仍不清楚。我们提出了一个基于世界价值观调查(WVS)的仿真框架,其中具有不同沟通风格的多元文化智能体进行纵向的、涉及价值观的讨论。在涉及1200个人设、15个主题和三个模型(GPT-4o、Gemini-2.5-Flash 和 Gemma-4-E4B)的约4000场对话中,我们评估了价值观忠实性、价值观漂移和对话真实性。我们发现,超过50%的人设从一开始就无法表达其被赋予的WVS价值观特征,而2-7%的人设在多次对话后发生漂移。移除人口统计细节的消融实验虽能提升部分模型的忠实性,但并未改变总体趋势:模拟的价值观分布仍然系统性地偏离被赋予的WVS特征。与人类讨论相比,模拟对话在风格一致性与语义多样性之间表现出不同的权衡,往往产生内容多样但风格重复的交流。这些发现表明,当前的LLM智能体能够生成看似合理的对话,但作为随时间表示和保持多样人类价值观特征的替代手段,其能力仍然有限。
cs.AI / 15 / 2609.05527

Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era

超越“AI辅助人类”:智能体时代面向人-智能体团队的决策导向型评估设计
Khosravi, Hamed, Huo, Xiaoming
Abstract
Wherever a coding agent works under engineer supervision, or a clinical model assists a radiologist, the deployment question is whether to keep the human-AI workflow or replace it with the human alone or the agent alone. The human-AI workflow is worth keeping only if it beats both of those alternatives. Yet once it is deployed, neither alternative outcome is observed: recovering one means replaying the task under that alternative, and every replay costs expert time or compute. Under a fixed replay budget, the design question is therefore which tasks should be more likely to receive a human-only replay, and which an agent-only replay. Existing methods do not directly target this decision. Agent benchmarks do not choose which missing baseline to measure, variance-based sampling ignores which of the two comparisons is closer to failing, and Bayesian information methods focus on learning model parameters instead of making the deployment decision. We propose TEAM-Design, a rule that gives every task two replay probabilities, one per baseline. It raises a probability where the missing baseline outcome is hard to predict from what is already known about the task and where that comparison is harder to establish, and lowers it where replay is expensive. We prove that the rule solves this budgeted design problem, and that drawing the replays at random from recorded probabilities still controls the chance of wrongly declaring that the workflow beats both. We reanalyze 6 clinical settings, where no human-AI workflow beats both alternatives, and a coding benchmark, where one does, then evaluate TEAM-Design on synthetic designs and on a semi-synthetic design built from a real chest X-ray reader study. TEAM-Design works best when one of the two comparisons is clearly harder to settle than the other, and can do worse than variance-based allocation when the two are similarly difficult.
Chinese Translation
无论是编码智能体在工程师监督下工作,还是临床模型辅助放射科医生,部署的核心问题都在于:是保留人机协作工作流,还是将其替换为仅由人类或仅由智能体完成。只有当人机协作工作流同时优于这两种替代方案时,才值得保留。然而一旦部署,这两种替代方案的结果都无法被观测到:要恢复其中一种方案的结果,就需要在该替代方案下重放任务,而每一次重放都要耗费专家时间或算力。在重放预算固定的条件下,设计问题因此变为:哪些任务应更可能获得仅人类重放,哪些任务应更可能获得仅智能体重放。现有方法并未直接针对这一决策。智能体基准测试不会选择测量哪个缺失的基线,基于方差的采样忽略了两个对比中哪一个更接近失效,而贝叶斯信息方法关注的是学习模型参数而非做出部署决策。我们提出TEAM-Design,一种为每个任务赋予两个重放概率(每个基线对应一个)的规则。当缺失基线的结果难以从已知信息中预测、且该对比更难确立时,该规则会提高相应的重放概率;当重放代价高昂时则降低概率。我们证明该规则能够求解这一预算约束下的设计问题,并且依据记录的概率随机抽取重放仍能控制错误宣称工作流同时优于两种替代方案的概率。我们重新分析了6个临床场景(其中没有人机协作工作流同时优于两种替代方案)和一个编码基准(其中存在这样的工作流),随后在合成设计以及基于真实胸部X光读片研究构建的半合成设计上评估了TEAM-Design。当两个对比中其中一个明显更难判定时,TEAM-Design表现最佳;而当两者难度相近时,其效果可能不如基于方差的分配方法。
cs.AI / 16 / 2609.05553

EdgeMem: LLM-Free Agent Memory Construction and Retrieval via Evidence-Preserving Multi-Anchor Hypergraph

EdgeMem:基于证据保持的多锚点超图的无LLM智能体记忆构建与检索
Cui, Zeyang, Cao, Jiannong, Wen, Zhiyuan, Yuan, Bo, Feng, Junlan, Chen, Shengyuan
Abstract
Agent memory allows LLM agents to use earlier interactions when answering new queries. Existing methods often compress interaction histories into summaries or other LLM-generated representations. Repeated generation adds cost and can discard answer-bearing details before the system knows what a future query will require. We propose EdgeMem, an agent-memory method built around a simple principle: preserve original interaction turns and organize them through complementary content, temporal, and episodic cues. EdgeMem realizes this principle with a multi-anchor hypergraph constructed by lightweight local processing. Retrieval directly returns source evidence and reserves LLM use for final answer generation, combining structured access to multi-session histories with faithful retention of the original conversation. Experiments on LoCoMo and LongMemEval-S show strong retrieval and memory-grounded question answering; on LoCoMo, EdgeMem achieves the highest strict-judge score among seven reproduced systems under a shared prompt (61.01 versus 58.70), while construction and retrieval require no generative-LLM calls. Overall, EdgeMem shows that preserving and organizing source evidence provides an effective and efficient foundation for agent memory without generative memory management.
Chinese Translation
智能体记忆使LLM智能体在回答新查询时能够利用先前的交互信息。现有方法通常将交互历史压缩为摘要或其他由LLM生成的表示。反复生成不仅增加了成本,而且可能在系统尚不知道未来查询需要什么信息之前就丢弃了包含答案的细节。我们提出了EdgeMem,一种围绕简单原则构建的智能体记忆方法:保留原始交互轮次,并通过互补的内容、时间和情节线索对它们进行组织。EdgeMem通过轻量级本地处理构建的多锚点超图(multi-anchor hypergraph)来实现这一原则。检索直接返回源证据,并将LLM的使用保留至最终答案生成阶段,从而将多会话历史的结构化访问与对原始对话的忠实保留相结合。在LoCoMo和LongMemEval-S数据集上的实验显示出强大的检索和基于记忆的问答能力;在LoCoMo上,在共享提示条件下,EdgeMem在七个复现系统中取得了最高的严格评判得分(61.01 对 58.70),同时其构建与检索过程无需任何生成式LLM调用。总体而言,EdgeMem表明,保留并组织源证据可以在无需生成式记忆管理的情况下,为智能体记忆提供一个高效且有效的基础。
cs.AI / 17 / 2609.05572

Deep belief networks are exact

深度信念网络是精确的
Smirnov, Gleb
Abstract
We prove that every strictly positive probability distribution on \(\{-1,1\}^n\) is represented exactly by a sigmoid belief network with finite parameters. This answers a question of Sutskever and Hinton. The proof upgrades their probability-sharing approximation to exact representation using Brouwer's fixed-point theorem.
Chinese Translation
我们证明,定义在\({-1,1\}^n\)上的每一个严格为正的概率分布都可以被一个参数有限的sigmoid信念网络(sigmoid belief network)精确表示。这回答了Sutskever和Hinton提出的一个问题。该证明利用布劳威尔不动点定理(Brouwer's fixed-point theorem),将他们的概率共享近似方法提升为精确表示。
cs.AI / 18 / 2609.05576

EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent

EnvCraft:为Claw式智能体的智能体强化学习合成可执行环境
Zeng, Yirong, You, Shen, Feng, Jinhang, Liu, Yufei, Ding, Xiao, Hou, Yutai, Cong, Hao, Wang, Yuxian, Ning, Wu, Xu, Wang, Cai, Bibo
Abstract
The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents that execute long-horizon tasks across stateful workspaces. While Agentic Reinforcement Learning (Agentic RL) provides a promising path to optimize these agents, its scaling is heavily bottlenecked by the severe scarcity of interactive training environments. Existing synthetic environments are strictly limited to tool-calling endpoints, rendering them insufficient for accommodating the end-to-end real-world demands of claw-like agents. To bridge this gap, we introduce EnvCraft, an automated framework for synthesizing executable environments and scalable training data. Specifically, EnvCraft employs an environment synthesis engine to build sandbox-isolated workspaces, alongside a topology-aware data generation engine to produce coherent task trajectories. Overall, we synthesize 139 interactive environments comprising approximately 20K complex tasks for Agentic RL training. Experiments on Qwen3/3.5 models (8B-32B) show that our method yields gains of up to +11.9% on Claw-style benchmarks and +8.0% on general tool-use benchmarks, with concurrent reductions in inference token cost. The results confirm that synthesized executable environments provide robust and generalizable learning signals for training.
Chinese Translation
大语言模型(LLM)的范式已迅速从被动的语言交互界面转变为能够在有状态工作空间中执行长程任务的自主Claw式智能体。尽管智能体强化学习为优化这些智能体提供了一条有前景的路径,但其扩展严重受制于交互式训练环境的极度稀缺。现有的合成环境严格局限于工具调用端点,因而难以满足Claw式智能体的端到端真实世界需求。为弥合这一差距,我们提出了EnvCraft,一个用于合成可执行环境和可扩展训练数据的自动化框架。具体而言,EnvCraft采用环境合成引擎构建沙箱隔离的工作空间,并结合拓扑感知的数据生成引擎生成连贯的任务轨迹。总体而言,我们合成了139个交互式环境,包含约2万个用于智能体强化学习训练的复杂任务。在Qwen3/3.5模型(8B-32B)上的实验表明,我们的方法在Claw风格基准上最高提升11.9%,在通用工具使用基准上提升8.0%,同时降低了推理的token成本。这些结果证实,合成的可执行环境能够为训练提供鲁棒且可泛化的学习信号。
cs.AI / 19 / 2609.05578

Planning and Scheduling Business Processes under Control-Flow Uncertainty

控制流不确定性下的业务流程规划与调度
Kunkler, Michel, Rinderle-Ma, Stefanie
Abstract
Scheduling activities in business processes can improve efficiency (e.g., reduce makespan), but is challenging because the exact sequence of activities required to complete a case is often uncertain due to decisions based on data that emerges during execution. Nevertheless, probabilistic information regarding such decisions can often be estimated or derived from historical execution logs, and can help anticipate which execution paths are likely to lead to successful completion. Planning with particular execution paths affects feasibility, i.e., the probability of successful completion, and the expected number of superfluous activities that are planned but never executed. We frame the problem as a chance-constrained optimization problem and present two formulations: A decomposed approach with two stages, a planning stage that minimizes the expected number of superfluous activities subject to a feasibility constraint, and a scheduling stage that minimizes the makespan over the planned activities; and an integrated approach that combines planning and scheduling into a single formulation. Evaluation on two real-world and one synthetic dataset shows that the integrated approach yields superior makespans but is intractable at scale, while the decomposed approach scales to large settings.
Chinese Translation
在业务流程中对活动进行调度可以提升效率(例如缩短完工时间),但这具有挑战性,因为完成一个案例所需的确切活动序列往往是不确定的,其原因在于执行过程中基于所产生数据做出的决策。尽管如此,关于此类决策的概率信息通常可以通过历史执行日志进行估计或推导,这有助于预测哪些执行路径可能导向成功完成。针对特定执行路径进行规划会影响可行性(即成功完成的概率),以及已规划但从未执行的冗余活动的期望数量。我们将该问题建模为一个机会约束优化问题,并提出了两种求解形式:一种是包含两个阶段的分解方法——规划阶段在可行性约束下最小化冗余活动的期望数量,调度阶段在已规划活动上最小化完工时间;另一种是将规划与调度合并为单一形式的集成方法。在两个真实数据集和一个合成数据集上的评估表明,集成方法能获得更优的完工时间,但在大规模场景下难以求解,而分解方法则可扩展至大规模设置。
cs.AI / 20 / 2609.05587

Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools

智能体对工具过度信任:度量对不可靠工具的依赖
Yang, Hoyeol, Song, Woojung, Kim, Taewon, Song, Jonghyun, Park, Seoyeon, Jo, Yohan
Abstract
Existing evaluations of tool-using agents primarily measure whether an agent can successfully complete diverse tasks with tools. These evaluations generally assume that tools return reliable information. However, tool returns in real-world systems can be plausible yet incorrect. We investigate how agents respond to unreliable tool returns by evaluating fourteen LLMs using three tools-web search, LLM sub-agent delegation, and code execution. For each tool, we corrupt its returns and measure whether agents adopt the corrupted content in their final answers. Agents exhibit high levels of overtrust across all three settings: the mean adoption rate exceeds one third for every tool and reaches 68.0% for web search. Analysis of reasoning traces reveals a particularly concerning failure mode: agents often recognize conflicts and even recover the correct answer internally, yet present only the corrupted answer without warning the user. To mitigate agents' overtrust in tool returns, we intervene at three levels: prompting by the user, metadata from the tool provider, and post-training by the agent builder. Although some interventions help for particular models or tools, none consistently mitigates overtrust across tools. These findings identify overtrust in unreliable tools as a serious and persistent failure mode, motivating evaluations and interventions that enable agents to validate tool outputs and transparently communicate unresolved conflicts.
Chinese Translation
现有的工具使用智能体评估主要衡量智能体能否借助工具成功完成各类任务。这些评估通常假设工具返回的信息是可靠的。然而,真实世界系统中工具的返回结果可能是看似合理却不正确的。我们通过评估十四个大语言模型(LLM)在三种工具(网络搜索、LLM子智能体委派和代码执行)上的表现,研究了智能体如何应对不可靠的工具返回结果。对于每种工具,我们对其返回内容进行篡改,并测量智能体是否在最终答案中采纳被篡改的内容。智能体在全部三种设置中都表现出高度的过度信任:每种工具的平均采纳率均超过三分之一,网络搜索更是高达68.0%。对推理轨迹的分析揭示了一种尤为令人担忧的失败模式:智能体往往能够识别冲突,甚至在内部恢复出正确答案,却仍然只呈现被篡改的答案,而不向用户发出任何警告。为缓解智能体对工具返回结果的过度信任,我们在三个层面进行干预:用户提示、工具提供方的元数据,以及智能体构建方的后训练。尽管某些干预措施对特定模型或工具有所帮助,但没有任何一种干预能在所有工具上持续有效地缓解过度信任。这些发现表明,对不可靠工具的过度信任是一种严重且持续存在的失败模式,需要发展能够使智能体验证工具输出并透明地传达未解决冲突的评估方法与干预手段。
cs.AI / 21 / 2609.05643

The convergent laboratory: when AI reasoning, autonomous experiments, high performance and quantum computing reshape chemistry

汇聚型实验室:当AI推理、自主实验、高性能计算与量子计算重塑化学
Huerta, Eliu, Wang, Xiaoyun, Gupta, Geetika, Sargent, Edward H., Owen, Cameron J., Fung, Victor, Mitra, Abhishek, Cheng, Austin, Bouchard, Emma, Mehdi, Shams
Abstract
This Comment emerges from TPC26 (https://tpc26.org), a conference convening leaders from academia, national laboratories, and industry who are reshaping materials science discovery. The meeting explored how AI, autonomous agents, self-driving labs, higher performance and quantum computing converge to amplify their individual impact on materials science discovery. The perspectives here reflect the firsthand experiences of researchers at these frontiers and capture the essence of this global endeavor. As AI-driven reasoning, autonomous agentic frameworks, self-driving laboratories, and fault-tolerant quantum processors mature simultaneously, we offer this Comment as a reference at what we believe is a tipping point of transformative advances and productive disruption in the chemical sciences.
Chinese Translation
本评论源于TPC26会议(https://tpc26.org),该会议汇聚了来自学术界、国家实验室和工业界、正在重塑材料科学发现格局的领军人物。会议探讨了人工智能、自主智能体、自动驾驶实验室、高性能计算与量子计算如何相互融合,以放大各自对材料科学发现的影响。本文所呈现的观点反映了处于这些前沿领域研究人员的亲身经历,并捕捉了这一全球性探索的本质。随着AI驱动的推理、自主智能体框架、自动驾驶实验室以及容错量子处理器的同时成熟,我们谨以此文作为一份参考——我们认为,化学科学正处于变革性进展与富有成效的颠覆性创新的临界点。
cs.AI / 22 / 2609.05663

What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets

LLM交易代理在生产环境中的实际行为:来自两个集群的六个月种群规模记录
Barton, T. J., Constantakis, Chris, Hauseman, Patti, Mous, Annie, Hoffman, Alaska, Bergeron, Brian, Goodreau, Hunter
Abstract
We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.
Chinese Translation
我们提供了一份连续的、种群规模的测量记录,研究对象是在生产环境中运行的自主语言模型交易代理,涵盖两个具有相同设计源流的系统:DX Terminal Pro(3,505个用户出资的金库,在Base memecoin市场中交易真实ETH,历时21天,2026年2月至3月)和DXAP实盘alpha集群(500至599个用户创建的代理,全历史记录,91至117个并发活跃,交易Hyperliquid永续合约,2026年6月至8月)。该记录跨度约六个月,包含750万次单模型调用及约30万次链上操作,另有231,638次多工具回合产生了14,596次成交。本文包含四项核心发现。第一,运行层对行为的影响超过策略文本中的任何内容:一个风险滑块可以解释杠杆(每档+0.425),代理固定效应吸收了60%的方差,而排行榜渲染边界因果性地引导了选择(在前三名截断处的断点回归为1.75倍)。第二,仓位管理与波动性无关:在每个波动性六分位中,杠杆中位数均为5.0倍,且一个姿态滑块单元格(占账户的11%)持有62%的清算。第三,代理几乎无法捕获其所触及的上行收益:43.2%的头寸在24小时内出现过至少+300个基点的有利波动,但其中49.3%以负交易收益平仓;一种机械式的止盈止损区间可为每个头寸挽回+39.0个基点。第四,两个集群均未表现出方向性优势。DXAP集群不盈利,且落后于匹配的Hyperliquid散户基准(往返胜率41%对50%)。一项在416个捕获的生产场景上对前沿模型进行的配对回放联赛发现,在此时间尺度上决策质量在统计上不可区分,而选择稳定性在各模型家族之间差异显著。所有核心结论均通过了按日聚类的推断、置换零假设检验以及统一费率的重新表述;本文以一份由我们自己的撤稿换来的17条方法论准则作结。
cs.AI / 23 / 2609.05708

CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning

CUSP:面向多智能体多模态推理的可分解集体不确定性
Yu, Chung-En Johnny, Garcia, David, Jalaian, Brian, Bastian, Nathaniel D.
Abstract
Aggregating heterogeneous vision-language models (VLMs) can improve multimodal reasoning, but neither an individual model's confidence nor that of the aggregated answer measures reliability at the system level. We present CUSP (Collective Uncertainty through Semantic Opinion Pooling), a training-free uncertainty quantification framework that maps multiple VLM responses to a shared semantic response space, pools them into a pooled semantic opinion, and reports two complementary system-level signals: collective uncertainty, the dispersion of the pooled opinion, and Jensen-Shannon divergence (JSD), the conflict among the model-level opinions. Within this pooled semantic opinion, the unnormalized collective entropy decomposes exactly into the mean of the models' individual semantic entropies and the JSD, separating total dispersion from model conflict. Requiring neither token logits nor calibration labels, CUSP applies to open-weight and commercial VLMs alike. In static multi-VLM ensembles, collective uncertainty is the strongest signal in the small-model regime (0.764 AUROC for prediction-error detection, 0.889 AUARC for abstention), outperforming uncertainty baselines majority voting and naive selection by 4.7 to 15.8 points and widening its margin as the ensemble grows; JSD is strongest in the evaluated commercial regime (0.819 AUROC, 0.910 AUARC) and ranks hard-answer model conflict with AUROC up to 0.982. The pooled prediction also improves accuracy over the average single model by 5.6 to 13.0 points. Over the full trajectory of a multi-step, multi-agent system, subagent collective uncertainty ranks system failures above chance (0.619 AUROC) and gives the best abstention ordering among the evaluated signals (0.699 AUARC).
Chinese Translation
聚合异构视觉-语言模型(VLM)可以提升多模态推理能力,但无论是单个模型的置信度还是聚合答案的置信度,都无法衡量系统层面的可靠性。我们提出了 CUSP(Collective Uncertainty through Semantic Opinion Pooling,基于语义意见池化的集体不确定性),这是一个无需训练的不确定性量化框架,它将多个 VLM 的响应映射到一个共享的语义响应空间,将其池化为一个池化语义意见,并给出两个互补的系统级信号:集体不确定性(池化意见的离散程度)和 Jensen-Shannon 散度(JSD,模型级意见之间的冲突程度)。在该池化语义意见中,未归一化的集体熵可以精确分解为各模型个体语义熵的均值与 JSD 之和,从而将总体离散度与模型间冲突分离开来。由于既不需要 token logits 也不需要校准标签,CUSP 同时适用于开源权重和商用 VLM。在静态多 VLM 集成中,集体不确定性是小模型场景下最强的信号(预测错误检测的 AUROC 为 0.764,弃答的 AUARC 为 0.889),相比多数投票和朴素选择等不确定性基线高出 4.7 至 15.8 个百分点,且随着集成规模扩大其优势进一步拉开;JSD 在所评估的商用模型场景下表现最强(AUROC 为 0.819,AUARC 为 0.910),并对困难答案的模型冲突排序的 AUROC 高达 0.982。池化预测相比单个模型的平均准确率提升了 5.6 至 13.0 个百分点。在多步、多智能体系统的完整轨迹中,子智能体的集体不确定性对系统故障的排序优于随机水平(AUROC 为 0.619),并在所评估的信号中给出最佳的弃答排序(AUARC 为 0.699)。
cs.AI / 24 / 2609.05721

Recovering Temporal and Geographic Signals from Language Model Embeddings

从语言模型嵌入中恢复时间与地理信号
Feuerstein, Esteban, Klimkowski, Victoria, de Zarate, Juan Manuel Ortiz, Suaiter, Federico Hernán
Abstract
Understanding whether language-model embeddings encode structured real-world information is important for both representation analysis and information retrieval. We study this question for temporal and geographic signals using a simple projection-based method that operates directly on output embeddings. Given a small set of seed examples, the method defines an axis in embedding space and ranks texts or entities by their projection onto that axis. Our approach is fully black-box and model-agnostic: it requires only embeddings, without access to model weights, internal activations, auxiliary probes, or additional training. This makes it applicable to modern embedding models available only through APIs and provides a lightweight way to analyze whether temporal and spatial dimensions are present in their representation spaces. We apply the method to temporal and geographic datasets and find that embedding projections recover meaningful chronological and spatial structure. These results provide evidence that output embeddings encode signals relevant to time and space, while also offering a practical tool for interpretability and for downstream temporal and geographic information retrieval tasks, such as temporal ordering, geographic ranking, and tagging.
Chinese Translation
理解语言模型嵌入是否编码了结构化的现实世界信息,对于表示分析和信息检索都具有重要意义。我们使用一种直接作用于输出嵌入的简单基于投影的方法,针对时间信号和地理信号研究了这一问题。给定少量种子示例,该方法在嵌入空间中定义一个轴,并根据文本或实体在该轴上的投影对其进行排序。我们的方法完全是黑盒且模型无关的:它只需要嵌入本身,无需访问模型权重、内部激活、辅助探针或额外训练。这使得该方法可应用于仅通过API提供的现代嵌入模型,并为分析其表示空间中是否存在时间与空间维度提供了一种轻量级手段。我们将该方法应用于时间和地理数据集,发现嵌入投影能够恢复有意义的时序和空间结构。这些结果表明,输出嵌入编码了与时间和空间相关的信号,同时也为可解释性研究以及时间排序、地理排序和标注等下游时间与地理信息检索任务提供了实用工具。
cs.AI / 25 / 2609.05736

Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses

超越提示词:测量与优化LLM工具智能体运行框架
Cen, Zhao, Ruan, Haibo, Chen, Wenjie, Tu, Pei-fen, Abbasi, Usman, Hesch, Joel
Abstract
LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed model: prompts, tool interfaces, middleware, state handling, and recovery logic. We study this setting as resource-bounded harness selection for fixed-model multi-turn tool agents, with the search surface scoped to prompts and tool-boundary middleware: edits are guarded intercepts at the tool boundary, not arbitrary rewriting of agent execution logic. Our optimizer-agnostic protocol reports mean held-out lift, worst-condition lift, repeatability, logged cost diagnostics, and RelLift95(B), a conservative estimate of the held-out gain of the harness selected under budget B. We instantiate the protocol with prompt-only and prompt-plus-middleware optimizers, including PRISM, which clusters failures and routes repairs to prompt, tool-boundary middleware, or joint edit surfaces within a Pareto search. On BFCL multi-round, tau2-Retail, and tau2-Telecom, PRISM obtains mean held-out lifts of 14.2, 14.9, and 10.1 percentage points and positive empirical RelLift95 on all three benchmarks, and a component ablation attributes the margin chiefly to failure-surface routing and the edit-pattern constraint. Across optimizers, the results show that some search procedures can occasionally find large gains but still choose brittle updates, so the reliability of the chosen harness should be reported alongside average held-out lift.
Chinese Translation
通过修改固定模型周围的运行时框架(harness)——包括提示词、工具接口、中间件、状态处理与恢复逻辑——可以在不重新训练的情况下改进LLM工具智能体。我们将这一设定研究为面向固定模型多轮工具智能体的资源受限框架选择问题,并将搜索范围限定在提示词与工具边界中间件:编辑操作是在工具边界处受保护的拦截,而非对智能体执行逻辑的任意重写。我们提出的与优化器无关的评测协议报告以下指标:留出集平均提升、最差条件提升、可重复性、记录的成本诊断,以及RelLift95(B)——即在预算B下所选框架留出集增益的保守估计。我们以仅提示词与提示词加中间件的优化器实例化该协议,其中包括PRISM:它在Pareto搜索中对失败进行聚类,并将修复导向提示词、工具边界中间件或联合编辑面。在BFCL多轮、tau2-Retail和tau2-Telecom基准上,PRISM分别获得14.2、14.9和10.1个百分点的平均留出集提升,且在全部三个基准上均取得正的经验性RelLift95;组件消融实验表明,这一优势主要归因于失败面路由与编辑模式约束。跨优化器的结果表明,某些搜索过程偶尔能找到较大增益,但仍可能选择脆弱的更新,因此在报告平均留出集提升的同时,应一并报告所选框架的可靠性。
cs.AI / 26 / 2609.05749

The Normalization of Deviance in AI Development

人工智能开发中的偏差常态化现象
Barkett, Emilio, Kimpton, Alexander, Graham, Daniel, Kundgol, Yusuf
Abstract
Work on the risks of artificial intelligence has focused predominantly on capability risk: the danger that systems become too powerful, too autonomous, or too misaligned with human values. Far less attention has been paid to the organizational level---to whether the institutions building these systems are themselves predisposed to drift toward failure. This paper argues that they are. Regardless of how capable AI systems become, the organizations building them face the same structural dynamics that preceded past major technological disasters. Drawing on case studies of the Space Shuttle Challenger, the Three Mile Island accident, and the Boeing 737 MAX crashes, this paper identifies the common structural mechanisms preceding each failure and maps them onto contemporary AI development. The findings suggest that existing safety infrastructure may provide less protection than it appears, as organizations can complete safety processes in full compliance and still produce catastrophic outcomes. The pre-disaster period of AI development is still underway; the purpose of this paper is to make these dynamics legible while they can still be interrupted.
Chinese Translation
关于人工智能风险的研究主要集中于能力风险:即系统变得过于强大、过于自主或与人类价值观偏离过远的危险。然而,组织层面的风险却鲜受关注——即构建这些系统的机构本身是否倾向于逐渐滑向失败。本文认为事实正是如此。无论AI系统的能力如何发展,构建这些系统的组织都面临着与以往重大技术灾难发生前相同的结构性动态。本文通过分析“挑战者号”航天飞机事故、三哩岛核事故以及波音737 MAX坠机事件的案例研究,识别出每次失败之前共同的结构性机制,并将其映射到当代AI开发中。研究结果表明,现有的安全基础设施所能提供的保护可能低于其表面所见,因为组织可以在完全合规的情况下完成安全流程,却仍然产生灾难性后果。AI开发的灾前时期仍在进行之中;本文的目的在于在这些动态尚可被干预之时,使其变得清晰可见。
cs.AI / 27 / 2609.05758

From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale

从单体式融合到智能体编排:大规模对话助手的动态响应机制
Cen, Zhao, Wang, Peng, Shi, Chuan, Zhang, Yufeng, Lyu, Ying, Ren, Wanmeng, Xue, Robert, Cheng, Claire Na, Mehdad, Yashar
Abstract
Conversational assistants can blend retrieval, action selection, escalation, and wording in a single model path, or separate those roles. We report a production migration of a customer-support assistant at a large accommodation marketplace (millions of conversations per month, 11 languages, 10-second P90). Dynamic Response (DR) replaces a single Qwen3-235B-A22B blended responder with a bounded ReAct orchestrator over typed tools plus a smaller generator that writes from a backend-validated context contract. Because the migration also changed prompts, alignment, and serving, we attribute each effect to its cause and claim as architecture effects only those measured on identical replayed turns: typed entity selection moves the reservation selector to a precision-first operating point (precision 8.3% to 89.1%, recall 75.2% to 67.3%), and typed action IDs with a membership check remove observed structured-action hallucination (2.14% to 0.0%). A low-ramp A/B test reproduces the replay escalation reductions: hard-escalation responses fall from 5.60% to 3.08% and soft-escalation responses from 9.56% to 2.49%, while production handoff volume holds roughly steady; self-solve is directional (+5.1 points, 95% CI [-2, +12]). Serving optimizations cut orchestrator P90 latency from 3.87s to 2.24s on a GPU footprint reduced by roughly one-third, and self-hosting reduces estimated annual model-serving cost by more than an order of magnitude.
Chinese Translation
对话助手可以在单一模型路径中融合检索、动作选择、升级(转人工)和措辞生成,也可以将这些角色分离。我们报告了一个大型住宿预订平台(每月数百万次对话、支持11种语言、P90延迟10秒)的客户支持助手生产系统迁移。动态响应机制(Dynamic Response, DR)用一个有界的 ReAct 编排器替代了单一 Qwen3-235B-A22B 融合式应答模型,该编排器通过带类型的工具进行调度,并配合一个较小的生成器,从经过后端验证的上下文契约中生成回复。由于此次迁移同时改变了提示词、对齐方式和服务架构,我们将各项效果归因于其各自原因,并仅将在完全相同的重放对话轮次上测得的效果归为架构效应:带类型的实体选择将预订选择器调整至以精确率为优先的工作点(精确率从 8.3% 提升至 89.1%,召回率从 75.2% 降至 67.3%);带类型并经成员校验的动作 ID 消除了观测到的结构化动作幻觉(从 2.14% 降至 0.0%)。一项低速率爬坡的 A/B 测试复现了重放实验中的升级减少效果:硬升级回复从 5.60% 降至 3.08%,软升级回复从 9.56% 降至 2.49%,同时生产环境中转人工量基本保持稳定;自助解决问题呈正向趋势(+5.1 个百分点,95% 置信区间 [-2, +12])。服务优化使编排器的 P90 延迟从 3.87 秒降至 2.24 秒,且 GPU 资源占用减少约三分之一;自托管部署使预估的年度模型服务成本降低超过一个数量级。
cs.AI / 28 / 2609.05774

Inference-Time Graph Engineering for Multi-Agent LLM Workflows

面向多智能体大语言模型工作流的推理时图工程
Tieu, Katherine, Fu, Dongqi, Xia, Yinglong, Li, Hong, Yan, Hong, He, Jingrui
Abstract
Recent multi-agent LLM systems increasingly rely on graph-structured communication to coordinate specialized agents. We revisit multi-agent orchestration from a graph-engineering perspective: rather than optimizing a static topology, we synthesize a task-conditioned temporal workflow graph that jointly specifies agent connectivity and edge-level communication semantics. We introduce ReActNet, a training-free framework that compiles a query and a set of role-specialized agents into a sequence of directed communication graphs. Each graph snapshot corresponds to one reasoning stage, and each edge carries a natural-language instruction specifying the message that a source agent should provide to a target agent. The compiled temporal graph is then executed through structured message passing: agents update their reasoning states by integrating their previous states with messages from controller-assigned neighbors, and a final aggregator synthesizes the resulting states into the answer. This design separates graph compilation from graph execution, making multi-agent coordination explicit, inspectable, and task-conditioned without requiring reinforcement learning or gradient-based topology optimization. Across knowledge reasoning, mathematical problem solving, code generation, and GAIA-style assistant tasks, ReActNet consistently improves over fixed-topology and learned-topology baselines while maintaining competitive inference cost. These results suggest that effective multi-agent orchestration depends not only on which agents communicate, but also on engineering executable workflow graphs that encode when, why, and how information should flow during reasoning.
Chinese Translation
近期多智能体大语言模型系统日益依赖图结构化的通信来协调专业化智能体。我们从图工程的视角重新审视多智能体编排:不再优化静态拓扑结构,而是合成一种任务条件化的时序工作流图,共同规定智能体间的连接方式与边级别的通信语义。我们提出 ReActNet,一个免训练框架,它将查询和一组角色特化智能体编译为一系列有向通信图。每个图快照对应一个推理阶段,每条边携带自然语言指令,规定源智能体应向目标智能体提供的消息。编译后的时序图随后通过结构化消息传递执行:智能体通过整合自身先前状态与来自控制器指派邻居的消息来更新推理状态,最终由聚合器将所得状态综合为答案。这一设计将图编译与图执行相分离,使多智能体协调变得显式、可检查且任务条件化,无需强化学习或基于梯度的拓扑优化。在知识推理、数学问题求解、代码生成以及 GAIA 风格的助手任务上,ReActNet 均持续优于固定拓扑和习得拓扑的基线方法,同时保持有竞争力的推理成本。这些结果表明,有效的多智能体编排不仅取决于哪些智能体进行通信,还取决于工程化构建可执行的工作流图,以编码信息在推理过程中何时、为何以及如何流动。
cs.AI / 29 / 2609.05776

DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents

DI-Bench:面向企业智能体的领域内数据智能基准的系统性生成方法
Zhang, Jiangyun, Surrao, Kristen, Nitayanont, Torpong, Zhang, Yupei, Singh, Roopali, Chen, Zhiyu, Huang, Julia, Tang, Zhou, Akbar, Shayan Ali, Alonso, Omar, Cornejo, Erwin, Li, Yuan, Zhang, Yi
Abstract
Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical computation, and constructing such benchmarks manually is costly. We present DI-Bench, a pipeline for generating realistic benchmarks for data intelligence (DI), the practice of extracting insights from large volumes of enterprise data. To emulate realistic DI tasks that require both computation and knowledge retrieval, DI-Bench builds an artifact linkage graph over data tables, dimensions, metrics, and documents to form questions involving structured data and associated knowledge. Ground truth answers are derived via query execution, followed by LLM question generation and validation. Applied to two public datasets, the pipeline produces a 731-task benchmark covering knowledge retrieval, analytical computation, and rule-grounded reasoning. To show the discriminatory capability and difficulty of the benchmark, we evaluate four models, revealing a substantial finding: models achieve only 32% accuracy when doing computational tasks where retrieved business rules modify the computation.
Chinese Translation
在领域特定基准上评估企业智能体至关重要,然而现有公开基准很少评估智能体能否将业务知识与分析计算相结合,而手动构建此类基准成本高昂。我们提出了DI-Bench,一个用于生成数据智能(Data Intelligence, DI)现实基准的流水线。数据智能是从海量企业数据中提取洞察的实践。为了模拟既需要计算又需要知识检索的现实DI任务,DI-Bench在数据表、维度、指标和文档之上构建工件关联图(artifact linkage graph),从而形成涉及结构化数据及相关知识的问题。标准答案通过查询执行获得,随后进行LLM问题生成与验证。将该流水线应用于两个公开数据集,生成了一个包含731个任务的基准,涵盖知识检索、分析计算和基于规则的推理。为展示该基准的区分能力和难度,我们评估了四个模型,揭示了一个重要发现:当执行的计算任务需要检索到的业务规则对计算进行修正时,模型的准确率仅为32%。
cs.AI / 30 / 2609.05782

Distilling Vision-Language Models for On-Device Fire Understanding

面向端侧火焰理解的视觉-语言模型知识蒸馏
Kazzazi, Mohammad, Liu, Zixuan, Khajavi, Siavash
Abstract
Vision-language models (VLMs) offer a promising alternative to conventional fire detection systems by reasoning about the semantic context of a scene and thus reducing false alarms, yet their large model size makes deployment on embedded fire sensors impractical. In this paper, we study how domain-specialized VLMs can be compressed for fully on-device deployment without losing the safety-critical behavior required for fire detection. We develop a teacher-student knowledge distillation framework in which large VLMs fine-tuned for fire understanding can be distilled into lightweight students. Experiments across multiple VLM families and model scales show that compact students preserve most of their teachers' fire-understanding capability. We further deploy the distilled models on our commercial Detectium fire detection sensor and jointly evaluate reasoning accuracy, latency, and memory usage. The results show that compression and deployment affect not only accuracy but also model failure modes, with Qwen2.5-0.5B providing the strongest overall deployment trade-off. Our findings provide broader guidance for deploying domain-specialized VLMs in resource-constrained, safety-critical settings.
Chinese Translation
视觉-语言模型通过推理场景的语义上下文,为传统火焰检测系统提供了一种有前景的替代方案,从而减少误报;然而其庞大的模型规模使其难以部署在嵌入式火焰传感器上。本文研究了如何将领域专用的VLM进行压缩以实现完全的端侧部署,同时不丢失火焰检测所需的安全关键行为。我们开发了一个师生知识蒸馏框架,将针对火焰理解微调的大型VLM蒸馏为轻量级学生模型。在多个VLM系列和模型规模上的实验表明,紧凑的学生模型能够保留教师模型大部分的火焰理解能力。我们进一步将蒸馏后的模型部署在商用的Detectium火焰检测传感器上,并联合评估推理准确率、延迟和内存占用。结果表明,压缩与部署不仅影响准确率,还会影响模型的失效模式,其中Qwen2.5-0.5B提供了最佳的总体部署权衡。我们的发现为在资源受限、安全关键的场景中部署领域专用VLM提供了更广泛的指导。
cs.AI / 31 / 2609.05788

More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review

不仅仅是模仿审稿人:评估大语言模型在投稿前同行评审中的作用
Parsa, Pouya, Rezaei, Amin
Abstract
Peer-review feedback often arrives too late for authors to make meaningful revisions. We study an author-facing LLM system that moves part of this stress test before submission: it generates a broad pool of atomic concerns and compresses them into a short report. We evaluate agreement with historical reviews and, separately, the possible validity of concerns they omit. From 10,000 ICLR 2026 submissions, we use 3,398 manuscripts with accessible versions that predate review. On a ten-paper diagnostic, independent sampling covers 44.9% of historical issues; deduplication and refill reaches 78.7% strict and 84.9% seriousness-weighted coverage, at 3.6$\times$ more requests and 5.2$\times$ more tokens. A hidden Top-32 Oracle preserves the full 79.3% weighted coverage of a 256-candidate pool, but paper-only selectors retain only 40--44%. LLM review therefore provides broad coverage with a large candidate pool but compresses poorly; ablations identify representative selection and matcher sensitivity as the main sources of this gap.
Chinese Translation
同行评审意见往往到达得太晚,作者难以据此进行有意义的修改。我们研究了一种面向作者的大语言模型(LLM)系统,将这一压力测试的部分环节提前到投稿之前:它生成一个广泛的原子化关注点池,并将其压缩为一份简短报告。我们评估了该系统与历史评审意见的一致性,并单独评估了其遗漏关注点的可能有效性。从10,000篇ICLR 2026投稿中,我们选取了3,398篇可获取审稿前版本的稿件。在一项十篇论文的诊断实验中,独立采样覆盖了44.9%的历史问题;通过去重和补充采样,严格覆盖率达到78.7%,按严重程度加权覆盖率达到84.9%,代价是请求量增加3.6倍、令牌量增加5.2倍。隐藏的Top-32 Oracle保留了256个候选池全部79.3%的加权覆盖率,但仅基于论文的选择器只能保留40–44%。因此,LLM评审在候选池较大时能提供广泛覆盖,但压缩效果不佳;消融实验表明代表性选择和匹配器敏感性是造成这一差距的主要原因。
cs.AI / 32 / 2609.05800

Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment

面向多元大语言模型对齐的溢出感知多价值导向方法
Pan, Weici, Barron, Xander, Zhou, Jiawei, Liu, Zhenhua
Abstract
Activation steering controls LLM behavior at inference time by adding learned directions to hidden states, but existing methods handle one concept at a time. Pluralistic alignment, where different stakeholders need different value emphases, requires steering multiple dimensions simultaneously. We show that naive steering produces substantial spillover: the effect intended for one value leaks into others. This parallels the treatment-versus-spillover decomposition in causal inference. We trace spillover to geometric entanglement of steering directions, captured by their Gram matrix, and derive a zero-cost correction from an activation-norm-penalized objective that decouples each direction's contribution exactly. Our end-to-end pipeline requires no fine-tuning, no reward model, and no manual prompt engineering: given only domain questions, it automatically discovers value dimensions, extracts directions, diagnoses entanglement, and applies corrected steering. On climate discourse, the correction improves the net steering effect from +5.9% to +14.0%, validated over 100,000 pairwise judgments.
Chinese Translation
激活导向(activation steering)通过在隐状态中加入学习到的方向,在推理时控制大语言模型的行为,但现有方法一次只能处理单一概念。多元对齐(pluralistic alignment)要求不同的利益相关者可以有不同的价值侧重,因此需要同时引导多个价值维度。我们证明,朴素的导向方法会产生显著的溢出(spillover)效应:针对某一价值施加的影响会泄漏到其他价值上。这与因果推断中的“处理效应—溢出效应”分解相类似。我们将溢出效应追溯到导向方向的几何纠缠,并用其 Gram 矩阵加以刻画,进而通过一个激活范数惩罚目标函数推导出一种零成本的校正方法,可精确解耦各方向的贡献。我们的端到端流程无需微调、无需奖励模型、也无需人工提示工程:仅需给定领域问题,即可自动发现价值维度、提取导向方向、诊断方向纠缠并施加校正后的导向。在气候话语任务上,该校正将净导向效果从 +5.9% 提升至 +14.0%,并经过 10 万次成对比较判断的验证。
cs.AI / 33 / 2609.05801

Evidence-Aligned Local Composition of Discrete Experts for Sequence Restoration

面向证据的离散专家局部组合用于序列恢复
Panahazari, Mohammad, Khan, Usman A., Aeron, Shuchin
Abstract
A document modeled as a discrete sequence of tokens can be thought of as being generated from a composition of texts from different domains; a README file, for example, moves between prose, code, and configuration. When such a document is corrupted and only frozen domain experts are available, restoring it requires deciding both what is missing and which expert to trust at each position, at test time and without region labels or a trained router. We introduce evidence-aligned local composition, which infers a soft, position-wise weighting over the experts from the marginal evidence of the corrupted observation under a given corruption model, estimating the evidence from the experts' own denoising losses and smoothing the weights across positions. Because the weighting is soft, it recovers a mixture when the true composition is mixed and concentrates on one expert when that suffices. Across a categorical simulator, byte-level experts, and experts fine-tuned from a $1.3$B discrete flow-matching model, the inferred weights track the true regions at $0.85$ field accuracy on naturally mixed scientific documents, and at $0.98$ on constructed mixtures whose regions are lexically disjoint. Restoration improves over a single global weight when the experts are genuinely distinct and reduces to it when they converge, tracking a measure of expert separation.
Chinese Translation
一个被建模为离散词元序列的文档可以被视为由来自不同领域的文本组合生成而成;例如,一个README文件会在散文、代码和配置之间切换。当此类文档受损而只有冻结的领域专家模型可用时,恢复它需要在测试时确定缺失的内容以及在每个位置应信任哪个专家,且无法获得区域标签或训练好的路由器。我们提出了面向证据的局部组合方法(evidence-aligned local composition),该方法在给定损坏模型下,根据受损观测的边际证据推断专家上的软性逐位置权重,其中证据由专家自身的去噪损失估计得到,并对权重进行跨位置平滑。由于权重是软性的,当真实的组合是混合时它能恢复出混合结果,而当单一专家足够时则集中于该专家。在分类模拟器、字节级专家以及从13亿参数离散流匹配模型微调的专家上,推断的权重在自然混合的科学文档上以0.85的字段准确率追踪真实区域,在区域词汇互不重叠的构造混合文档上达到0.98。当专家真正不同时,恢复效果优于单一全局权重;当专家趋于一致时则退化为该权重,并与专家分离度指标保持一致。
cs.AI / 34 / 2609.05806

Exposing Weaknesses in Emotion Recognition in Conversations

揭示对话中情绪识别任务的薄弱环节
Khalifa, Amir Ben, Bezancon, Fanny, Trabelsi, Amine, Abdulrazak, Bessam
Abstract
Emotion Recognition in Conversations (ERC) aims to identify speakers' emotions in multi-turn dialogue. Accurate emotion recognition can support a wide range of applications, including empathetic conversational agents, mental health support, and educational technologies. While many recent approaches rely on task-specific fine-tuning, such models may exploit dataset-specific cues. A central yet rarely questioned assumption in ERC is that each utterance can be assigned a single unambiguous emotion label. To investigate this assumption, we study ERC using Large Language Models (LLMs) in a zero-shot setting while incorporating preceding conversational turns as context. We show that aggregate metrics mask systematic failures. Errors concentrate around utterances containing negations, exclamations, and interjections. This pattern is consistent across all evaluated models, suggesting limitations in the benchmarks rather than model-specific weaknesses. A controlled re-annotation study involving four human annotators supports this finding: strong agreement is observed in only 35 percent of cases, with neutral utterances dominating high-agreement instances, while many emotional categories fall into low-agreement regimes. These findings suggest that many apparent model errors reflect genuine annotation ambiguity rather than poor emotion understanding. Standard single-label evaluation is therefore insufficient. To address this limitation, we introduce an LLM-as-Judge framework that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision.
Chinese Translation
对话情绪识别(Emotion Recognition in Conversations, ERC)旨在识别多轮对话中说话人的情绪。准确的情绪识别可以支持广泛的应用,包括共情对话代理、心理健康支持和教育技术。尽管许多近期方法依赖于任务特定的微调,但此类模型可能会利用数据集特有的线索。ERC中一个核心却很少被质疑的假设是:每条话语都可以被赋予一个单一且明确无歧义的情绪标签。为了探究这一假设,我们使用大语言模型(LLMs)在零样本(zero-shot)设置下研究ERC,并将先前的对话轮次作为上下文纳入考量。我们表明,总体指标掩盖了系统性的失败。错误集中在包含否定词、感叹词和感叹语气词的话语上。这一模式在所有评估的模型中均保持一致,这表明问题源于基准数据集本身的局限,而非特定模型的缺陷。一项由四名人工标注者参与的受控重标注研究支持了这一发现:仅有35%的样本表现出强一致性,其中中性话语在高一致性样本中占主导地位,而许多情绪类别则处于低一致性区间。这些发现表明,许多表面上看似模型错误的情况实际上反映了真实的标注歧义,而非情绪理解能力不佳。因此,标准的单标签评估是不充分的。为解决这一局限,我们引入了一种LLM-as-Judge(大语言模型作为评判者)框架,该框架根据每种情绪在对话上下文中的合理性对其进行独立评估,而不是强制做出单标签决策。
cs.AI / 35 / 2609.05818

Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools

智能体化BAIM-LLM评估(ABLE):基于蛋白质设计工具使用的LLM基准测试
Cai, Bryce, Jeyapragasan, Geetha, Nedungadi, Samira, Yukich, Jake, Donoughe, Seth
Abstract
We introduce ABLE, a benchmark for evaluating LLM agents' ability to use biological AI models (BAIMs), such as ProteinMPNN and AlphaFold3, in dual-use protein design workflows. ABLE assesses agent performance through a set of tasks spanning structure retrieval, sequence generation, and design validation. We evaluate 15 frontier models and find that seven refuse all tasks, while the remaining models exhibit substantial performance differences. Claude Sonnet 4 and Gemini 3 Pro achieve the highest scores across information retrieval, tool selection, and tool use. We further compare model performance on a subset of tasks against an expert human baseline. Our results suggest that current LLMs can substantially lower barriers to protein design, but remain inconsistent in planning, strategy generation, and integrating biological knowledge with tool use.
Chinese Translation
我们提出了ABLE,一个用于评估大语言模型(LLM)智能体在两用蛋白质设计工作流程中使用生物人工智能模型(BAIM,如ProteinMPNN和AlphaFold3)能力的基准测试。ABLE通过一系列涵盖结构检索、序列生成和设计验证的任务来评估智能体的表现。我们评估了15个前沿模型,发现其中七个模型拒绝执行所有任务,而其余模型则表现出显著的性能差异。Claude Sonnet 4和Gemini 3 Pro在信息检索、工具选择和工具使用方面取得了最高分数。我们进一步将模型在部分任务上的表现与专家人类基线进行了比较。我们的结果表明,当前的LLM能够显著降低蛋白质设计的门槛,但在规划、策略生成以及将生物学知识与工具使用相结合方面仍表现不稳定。
cs.AI / 36 / 2609.05824

Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents

超越Top-$k$技能检索:面向大语言模型代理的多样性感知技能路由
Wei, Wang, Yang, Tiankai, Basu, Samyadeep, Chen, Hongjie, Zhao, Yue, Tu, Zhengzhong, Hu, Xiyang, Dernoncourt, Franck, Rossi, Ryan A., Eldardiry, Hoda
Abstract
Large language model (LLM) agents increasingly rely on external skills, but routing user requests over large skill registries is difficult because many skills are functionally redundant while complex tasks often require complementary skill sets. Existing skill routers typically rank candidates independently by query relevance, which can waste context budget on redundant skills. We propose Diverse Skill Routing (DSR), a diversity-aware reranking framework that uses a Determinantal Point Process to balance relevance and non-redundancy. DSR introduces a query-residual diversity kernel that penalizes redundant skill overlap while reducing penalties caused only by shared query relevance. On the SkillRouter benchmark, DSR improves recall and full coverage over a strong pointwise reranking baseline, with larger gains on multi-skill queries. These results suggest that skill routing should be treated not only as relevance ranking, but also as complementary set selection.
Chinese Translation
大语言模型(LLM)代理日益依赖外部技能,但在大规模技能注册表上对用户请求进行路由十分困难,原因在于许多技能在功能上高度冗余,而复杂任务往往又需要互补的技能组合。现有的技能路由器通常仅依据查询相关性对候选技能进行独立排序,这可能将上下文预算浪费在冗余技能上。我们提出多样性技能路由(Diverse Skill Routing, DSR),这是一个多样性感知的重排序框架,利用行列式点过程(Determinantal Point Process)来平衡相关性与非冗余性。DSR 引入了一种查询残差多样性核,在惩罚冗余技能重叠的同时,减少仅由共享查询相关性引起的惩罚。在 SkillRouter 基准上,DSR 相较于强大的逐点重排序基线提升了召回率和完全覆盖率,且在多技能查询上收益更为显著。这些结果表明,技能路由不仅应被视为相关性排序,还应被视为互补集合的选择问题。
cs.AI / 37 / 2609.05834

Learning Counterfactual World Models for Embodied Reasoning under Partial Observability

面向部分可观测环境下具身推理的反事实世界模型学习
Zhou, Todd Y., Zhang, Daniel
Abstract
World models promise a general route to embodied intelligence: learn predictive dynamics once, then reason, plan, and act with them. Increasingly, the representations beneath such models are pretrained on large-scale video, interaction, and multimodal corpora, which raises a question prediction quality alone cannot answer: when is a learned representation actually actionable? We identify a failure mode we call counterfactual collapse: a model predicts visually plausible futures while failing to distinguish interventions with different behavioral consequences. This arises whenever a representation is optimized for perceptual similarity rather than intervention structure, which is precisely the objective under which most large-scale pretrained encoders are learned. We introduce Counterfactual Latent World Models (CLWM), which combine a recurrent belief-state encoder, action-conditioned latent dynamics, and a contrastive counterfactual objective that separates futures induced by distinct interventions even when their observations look alike. Across occluded manipulation, aliased navigation, and long-horizon manipulation, CLWM improves planning success over the strongest baseline (65.1% $\to$ 74.6% on Occluded Push and 67.3% $\to$ 78.9% on Aliased Maze) and reduces exploitative planning failures (18.4% $\to$ 9.7% on Deferred Kitchen), with ablations attributing the gains to hard counterfactual negatives, especially perceptual-alias negatives. Finally, our counterfactual separability metric, which tracks planning success across the five baseline model classes ($r \ge 0.94$), is representation-agnostic: given intervention-outcome labels, it can audit any encoder, pretrained or trained from scratch, before a planner trusts it. We do not yet measure it on large-scale pretrained encoders. Here we establish the metric and its relationship to planning success for world models trained from scratch.
Chinese Translation
世界模型有望为具身智能提供一条通用路径:只需学习一次预测性动力学,便可据此进行推理、规划与行动。这类模型底层的表征日益依赖大规模视频、交互和多模态语料进行预训练,由此引出一个仅凭预测质量无法回答的问题:学习到的表征何时才是真正可行动的?我们发现了一种称为反事实坍缩的失效模式:模型能预测出视觉上合理的未来,却无法区分具有不同行为后果的干预。只要表征以感知相似性而非干预结构为优化目标,这一失效模式便会出现——而这恰恰是大多数大规模预训练编码器所采用的训练目标。我们提出反事实潜在世界模型(Counterfactual Latent World Models, CLWM),它结合了循环信念状态编码器、动作条件化的潜在动力学,以及一种对比式反事实目标,即使不同干预导致的观测十分相似,也能将其所诱导的未来区分开来。在遮挡操作、别名导航和长时程操作任务上,CLWM 相较最强基线提升了规划成功率(Occluded Push 上由 65.1% 提升至 74.6%,Aliased Maze 上由 67.3% 提升至 78.9%),并减少了投机性规划失败(Deferred Kitchen 上由 18.4% 降至 9.7%);消融实验将这一提升归因于难分反事实负样本,尤其是感知别名负样本。最后,我们的反事实可分性度量在五类基线模型上均与规划成功率高度相关($r \ge 0.94$),且与具体表征无关:只需给定干预-结果标签,它便可在规划器信任任何编码器(无论是预训练的还是从零训练的)之前对其进行审计。我们尚未在大规模预训练编码器上对其进行测量。本文为从零训练的世界模型建立了该度量及其与规划成功率之间的关系。
cs.AI / 38 / 2609.05837

AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories

AgentBrew:从原始真实世界轨迹中学习的离线工具使用智能体
Lyu, Zhiyi, Li, Yewen, Zheng, Longtao, Yang, Shengtian, Feng, Lang, Feng, Lei, Jiang, Peng, Gai, Kun, Cai, Qingpeng, An, Bo
Abstract
LLM-based agents are increasingly deployed in real-world applications through tool-use APIs, yet training them for specific environments remains fundamentally difficult: real-world applications provide no pre-defined tasks or verifiers, no faithful simulators, and limited budget for large-scale environment interaction. In this paper, we propose \textbf{AgentBrew}, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts. The agent first explores the target environment to collect a raw trajectory corpus without quality filtering. To extract training signal from this noisy corpus, \emph{retrospective task inference} reconstructs an aligned instruction for each trajectory based on its actual outcome, and \emph{PMI-Based credit assignment} decomposes the trajectory's total information about the inferred instruction into additive per-action credits via pointwise mutual information (PMI). These credits weight the policy training objective, amplifying informative actions while suppressing ineffective ones. On three real-world MCP applications (GitHub, Notion, PostgreSQL), AgentBrew improves Qwen3-32B by +8.7 Acc / +9.7 Score on average, surpassing Qwen3-235B (+2.3 / +4.4) and outperforming rejection sampling (+5.9 / +10.3). These results demonstrate that fine-grained offline learning can recover useful supervision from raw trajectories that filtering-based approaches would discard. The code is available at https://github.com/alphatogo/AgentBrew
Chinese Translation
基于大语言模型(LLM)的智能体正越来越多地通过工具使用 API 部署到真实世界应用中,然而针对特定环境对其进行训练在根本上仍然十分困难:真实世界应用既不提供预定义任务或验证器,也没有忠实的模拟器,且大规模环境交互的预算有限。本文提出 extbf{AgentBrew},一个离线训练框架,它能够从单批次原始交互轨迹中学习有效的工具使用策略,无需任务验证器或迭代的在策略采样。智能体首先探索目标环境以收集未经质量过滤的原始轨迹语料库。为了从这一含噪语料库中提取训练信号,\emph{回顾式任务推断}(retrospective task inference)基于每条轨迹的实际结果为其重构一条对齐的指令,而\emph{基于 PMI 的信用分配}(PMI-Based credit assignment)则通过逐点互信息(PMI)将轨迹关于所推断指令的总信息分解为可加的逐动作信用。这些信用用于加权策略训练目标,从而放大有信息量的动作,同时抑制无效动作。在三个真实世界的 MCP 应用(GitHub、Notion、PostgreSQL)上,AgentBrew 使 Qwen3-32B 平均提升 +8.7 Acc / +9.7 Score,超越 Qwen3-235B(+2.3 / +4.4),并优于拒绝采样方法(+5.9 / +10.3)。这些结果表明,细粒度的离线学习能够从原始轨迹中恢复出基于过滤的方法所丢弃的有用监督信号。代码已发布于 https://github.com/alphatogo/AgentBrew
cs.AI / 39 / 2609.05889

Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization

基于像素-提示联合优化的视觉语言模型多模态资源耗尽攻击
Ni, Zhaoxiong, Xiao, Yatie, Pun, Chi-Man, Peng, Fei, Guan, Qingxiao, Tang, Keke
Abstract
Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal threat models, treating the image branch as the primary optimization surface while holding user-visible prompts fixed. Even recent loop-centric variants remain confined to this single-channel paradigm, leaving the exploitation of availability unexplored as a cross-modal optimization problem over jointly controllable input surfaces. We introduce Joint Pixel-Prompt Optimization (JPPO), the first compound adversarial framework elevating the visible prompt to a first-class adversarial variable alongside image perturbations. Under a restricted joint-input threat model, JPPO performs coupled, stagewise optimization over both the pixel and prompt surfaces. This produces synergistic cost amplification, mechanistically distinct from loop-dependent failures, exhibiting negligible loop incidence in our experiments. Evaluating five open-source VLM families on MS COCO and ImageNet under an 8/255 infinity-norm budget, JPPO achieves over 4.6x latency and 5.3x energy amplification on Qwen2.5-VL-7B, and over 36.6x latency with 32.7x energy amplification on BLIP-2. This represents the strongest cost amplification among directly compared baselines while requiring substantially fewer optimization iterations. Ablations confirm this amplification arises from multimodal coordination rather than prompt length or isolated modalities. These findings reveal structural blind spots in current VLM serving defenses, motivating cost-aware robustness evaluation as a first-class security requirement for multimodal deployments.
Chinese Translation
针对自回归视觉语言模型(VLM)的资源耗尽攻击通常采用单模态威胁模型,将图像分支作为主要优化对象,而保持用户可见提示固定不变。即使是近期的以循环为中心的变体也仍局限于这种单通道范式,未将可用性攻击视为跨模态优化问题,也未曾探索对联合可控输入面的利用。我们提出了联合像素-提示优化方法(Joint Pixel-Prompt Optimization,JPPO),这是首个将可见提示提升为与图像扰动同等的对抗变量的一等对抗框架。在受限的联合输入威胁模型下,JPPO 在像素面与提示面上执行分阶段的耦合优化。这产生了协同的成本放大效应,其机制有别于依赖循环的失效方式,在实验中表现出可忽略不计的循环发生率。在 8/255 无穷范数扰动预算下,我们在 MS COCO 和 ImageNet 上对五个开源 VLM 系列进行评估,JPPO 在 Qwen2.5-VL-7B 上实现了超过 4.6 倍的延迟放大和 5.3 倍的能耗放大,在 BLIP-2 上实现了超过 36.6 倍的延迟放大和 32.7 倍的能耗放大。在直接对比的基线方法中,这代表了最强的成本放大效果,同时所需优化迭代次数显著更少。消融实验证实,这种放大源于多模态协同,而非提示长度或孤立模态的作用。这些发现揭示了当前 VLM 服务防御中的结构性盲点,促使成本感知的鲁棒性评估成为多模态部署的一等安全需求。
cs.AI / 40 / 2609.05894

The End of AI Exponentiation: Fluttering Inside and Outside AI Bubble

AI指数级增长的终结:AI泡沫内外的震荡
Kebande, Victor
Abstract
The exponentiation of Artificial intelligence (AI) in the recent past has entered a transformative era that has been driven by the growth in large language models (LLMs), large-scale compute infrastructures, and autonomous reasoning systems. However, the rapid acceleration of AI has increasingly shown technological, societal, economic, ethical and infrastructural challenges associated with peak data limitations, rising computational demands, synthetic data recursion, valuation inflation, and societal instability. The traditional scaling paradigms that have powered the modern AI systems are gradually encountering friction in sustaining continuous exponential growth. This paper views ``the end of AI exponentiation,'' thus exploring how it flutters inside and outside the bubble, where instability emerges within the AI ecosystem through compute and data-center races, speculative investments, and the rat-race toward superintelligence, and outside the ecosystem through labor disruption, governance concerns, public uncertainty, and geopolitical acceleration surrounding future intelligent systems and infrastructures globally.
Chinese Translation
近年来,人工智能(AI)的指数级发展已进入一个变革性时代,其驱动力来自大语言模型(LLM)、大规模计算基础设施以及自主推理系统的发展。然而,AI的快速加速日益显现出技术、社会、经济、伦理和基础设施方面的挑战,包括数据峰值的限制、不断攀升的计算需求、合成数据的递归问题、估值膨胀以及社会不稳定。支撑现代AI系统的传统扩展范式在维持持续指数级增长方面正逐渐遭遇阻力。本文提出“AI指数级增长的终结”这一观点,并探讨其在泡沫内外的震荡表现:在AI生态系统内部,不稳定性通过算力与数据中心的竞赛、投机性投资以及争夺超级智能的激烈竞争而显现;在生态系统外部,则通过劳动力冲击、治理隐忧、公众不确定性以及围绕全球未来智能系统与基础设施的地缘政治加速而体现。
cs.AI / 41 / 2609.05947

Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review

超越最终决策:面向透明化AI辅助同行评审的过程中心基准测试
Yuan, Siming, Zhang, Xueyi, Ni, Wangze, Xiao, Tianfang, Di, Shimin, Zhu, Jia, Jiang, Zhuoren, Tan, Rong, Chen, Lei, Ren, Kui
Abstract
Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer review mainly focus on the overall quality of generated reviews or the accuracy of final decisions. They therefore provide limited evidence about whether model decisions are supported by sufficient and reliable review evidence. We introduce a process-centric diagnostic benchmark for AI-assisted peer review. It uses (x,$z_s$,$z_c$,$z_r$,y) to represent the paper content, summary, critique, suggestion, and decision. We convert heterogeneous review records from PeerRead, NLPeer ARR-22, and OpenReview-ICLR into process-aligned data. Our benchmark uses direct decision prediction from the paper content (Direct) as its baseline. It compares the decision value of Gold-process variables and Predicted-process variables, and conducts stage-level evaluation, chain-consistency evaluation, and interventional sensitivity analysis. Experiments across three datasets and six models show that Gold-process variables generally have higher decision value. For the main analysis model, the Gold--Predicted gap remains stable across datasets and random seeds. This gap is also reproduced in most model--dataset combinations. Although model-generated intermediate review texts show relatively high local consistency across adjacent stages, the final decisions are not consistently supported by the preceding review evidence. Our benchmark targets AI systems designed to assist rather than replace human reviewers. It provides a transparent and auditable diagnostic tool for evaluating the reliability of their review processes.
Chinese Translation
同行评审是科学质量控制的核心环节。然而,现有对AI辅助同行评审的评估主要关注生成评审内容的整体质量或最终决策的准确性,因此对于模型决策是否有充分且可靠的评审证据支撑,所能提供的证据十分有限。我们提出了一个面向过程(process-centric)的诊断性AI辅助同行评审基准。该基准使用 (x, $z_s$, $z_c$, $z_r$, y) 来表示论文内容、总结、评审意见、建议和决策。我们将来自 PeerRead、NLPeer ARR-22 和 OpenReview-ICLR 的异构评审记录转换为与过程对齐的数据。本基准以直接从论文内容进行决策预测(Direct)作为基线,比较黄金过程变量(Gold-process)与预测过程变量(Predicted-process)的决策价值,并开展阶段级评估、链式一致性评估以及干预敏感性分析。在三个数据集和六个模型上的实验表明,黄金过程变量总体上具有更高的决策价值。对于主要分析模型而言,黄金-预测差距在不同数据集和随机种子下保持稳定,且该差距在大多数模型-数据集组合中同样得到复现。尽管模型生成的中间评审文本在相邻阶段之间表现出较高的局部一致性,但最终决策并未持续得到前序评审证据的支撑。本基准面向旨在辅助而非取代人类评审员的AI系统,为其评审过程的可靠性评估提供了一个透明、可审计的诊断工具。
cs.AI / 42 / 2609.05992

MOAE: Multi-Objective Agent Evolution with Pareto-Preserving Search

MOAE:基于Pareto保持搜索的多目标智能体进化
Jiang, Hengle, Cai, Qijun, Luo, Ziying, Tang, Ke
Abstract
As LLM-based agents continue to advance, their evaluation has become increasingly multifaceted: a capable agent must not only achieve high task completion accuracy but also perform well in interaction quality, safety, and efficiency, raising a central question: can these objectives be optimized simultaneously? Existing methods have considered multiple objectives, but many collapse heterogeneous measurements into a fixed scalar score. Such scalarization depends on metric normalization and preference weights and may discard candidates that represent useful deployment trade-offs. We introduce Multi-Objective Agent Evolution (MOAE), which organizes iterative in-context refinement as a Pareto-preserving evolutionary search over complete agent rollouts. Given a limited rollout budget, MOAE maintains an empirical archive of non-dominated candidates, uses objective-specific diagnostics to guide offspring generation, and applies constraint-aware selection only at deployment. This separates candidate preservation during search from the preference used to return a final solution. The procedure requires no parameter updates and allows each objective to be replaced by any measurable property, which we instantiate as task performance, trajectory quality, and safety. Experiments on TravelPlanner and AgentDojo show that MOAE consistently improves task performance and trajectory quality while maintaining strong safety under matched rollout budgets. Search-behavior analysis further shows that Pareto preservation expands the attainable objective region and increases the frequency of joint improvement. These results demonstrate the potential of Pareto-preserving in-context evolution for optimizing multiple agent properties without committing to a fixed scalarization during search.
Chinese Translation
随着基于大语言模型(LLM)的智能体不断发展,其评估也日益多元化:一个能力强的智能体不仅需要达到较高的任务完成准确率,还需在交互质量、安全性和效率方面表现良好。由此引出一个核心问题:这些目标能否同时得到优化?现有方法虽然考虑了多个目标,但许多方法将异构的度量指标压缩为固定的标量分数。这种标量化依赖于指标归一化和偏好权重,可能会丢弃那些代表有用部署权衡的候选方案。我们提出了多目标智能体进化方法(Multi-Objective Agent Evolution, MOAE),它将迭代式上下文内(in-context)优化组织为针对完整智能体轨迹(rollout)的Pareto保持进化搜索。在有限的轨迹预算下,MOAE维护一个由非支配候选解组成的经验存档,利用目标特定的诊断信息引导子代生成,并仅在部署阶段应用考虑约束的选择。这将搜索过程中的候选解保留与用于返回最终解的偏好分离开来。该过程无需参数更新,并允许将每个目标替换为任何可度量的属性,我们将其实例化为任务性能、轨迹质量和安全性。在TravelPlanner和AgentDojo上的实验表明,在相同的轨迹预算下,MOAE持续提升任务性能和轨迹质量,同时保持较强的安全性。搜索行为分析进一步表明,Pareto保持扩展了可达的目标区域,并提高了联合改进的频率。这些结果证明了Pareto保持的上下文内进化在优化多个智能体属性方面的潜力,且无需在搜索过程中固定标量化方案。
cs.AI / 43 / 2609.05995

Agentic Pressure: The Endogenous Entropy of Reliable Autonomy

智能体压力:可靠自主性的内生熵
Jiang, Hengle, Luo, Ziying, Tang, Ke
Abstract
Achieving reliable autonomy in the wild requires agents to sustain continuous operations across long-horizon trajectories. However, as agents navigate these unconstrained settings, they encounter cumulative friction that inherently destabilizes their alignment. In this paper, we identify a distinct non-adversarial phenomenon termed Agentic Pressure. We define this as a kinetic force that spontaneously emerges when the cost of compliance conflicts with the imperative of goal achievement. Unlike static jailbreaks, this pressure is endogenous and arises directly from the dynamics of interaction. We propose a theoretical framework that formalizes Agentic Pressure as the ratio between the required work to overcome environmental friction and the remaining capacity of the agent. Our analysis demonstrates that when this pressure exceeds a critical threshold, agents exhibit safety drift as a mathematically optimal adaptation. Consequently, they often resort to Instrumental Hallucination to rationalize rule violations. Empirical experiments validate this framework and show that aligned agents spontaneously compromise safety to preserve autonomy under high-pressure conditions.
Chinese Translation
在真实环境中实现可靠的自主性,要求智能体(agent)能够在长时程轨迹上维持持续运行。然而,当智能体在这些不受约束的环境中运行时,会遭遇累积性摩擦,这种摩擦会固有地破坏其与人类意图的对齐。在本文中,我们识别了一种独特的非对抗性现象,称之为智能体压力。我们将其定义为一种在合规成本与目标达成需求发生冲突时自发产生的动力。与静态越狱攻击不同,这种压力是内生的,直接源自交互动态过程。我们提出了一个理论框架,将智能体压力形式化为克服环境摩擦所需做的功与智能体剩余能力之间的比值。我们的分析表明,当这种压力超过某个临界阈值时,智能体会表现出安全漂移,这是一种数学上最优的适应行为。因此,智能体往往会借助工具性幻觉来为违反规则的行为进行合理化。实证实验验证了该框架,并表明对齐的智能体在高压条件下会自发地牺牲安全性以维持自主性。
cs.AI / 44 / 2609.06036

Generator-Independent Runtime Assurance under Partial Observation

部分可观测性下与生成器无关的运行时保障
Wan, Guangxi, Xie, Yongbo, Liu, Yuqi, Dong, Qingwei, Li, Qingxin, Bai, Hongfei, Zeng, Peng
Abstract
Proposal-based controllers---learned policies, language-model planners, and other black-box \emph{generators}---are increasingly deployed behind runtime verification gates. We ask when the closed-loop safety guarantee decouples from the generator. The prevailing per-candidate certification pattern does not compose: under retry or best-of-$k$ selection a per-candidate false-admission level $\alpha$ can inflate to $1-(1-\alpha)^{k}$. Our main theorem shows that \emph{simultaneous setwise soundness}---certifying a set of admissible proposals containing no nonviable action---is necessary and sufficient for generator-independent \emph{admission soundness}, the worst case over all generators of executing a nonviable proposal equalling the probability of setwise failure; together with a design-time certificate and a no-bypass rule it is sufficient for \emph{contract safety}, with violation bound $\Gamma+\sum_t\varepsilon_t+\eta$ invariant under arbitrary, even adversarial, replacement of the generator. A second theorem bounds every admission mechanism under partial observation: for a fixed probing and admission policy, if two state hypotheses whose information laws lie within total-variation distance $\delta$ require different safe decisions, then $\abar+\beta+\delta\ge1$. A sequential risk ledger makes the guarantee implementable with time-uniform confidence tubes, and shows that deterministic admission computations concentrate all statistical risk in state estimation. Simplex-style runtime assurance and control-barrier-function filtering are recovered as degenerate cases.
Chinese Translation
基于提案的控制器——学习得到的策略、语言模型规划器以及其他黑盒生成器(generator)——正越来越多地部署在运行时验证门之后。我们探讨闭环安全性保证何时能与生成器解耦。主流的逐候选认证模式无法复合:在重试或最优$ k $选择机制下,逐候选的误准入水平$ \alpha $可能膨胀为$ 1-(1-\alpha)^{k} $。我们的主要定理表明,同时的集合级可靠性(simultaneous setwise soundness)——即认证一个不包含任何不可行动作的可行提案集合——对于与生成器无关的准入可靠性(admission soundness)而言是充分且必要的;所谓准入可靠性,即在所有生成器上执行不可行提案的最坏情况概率等于集合级失败的概率;再加上一个设计时证书和一条禁止绕过规则,它便足以保证契约安全性(contract safety),其违反界为$ \Gamma+\sum_t\varepsilon_t+\eta $,且在生成器被任意(甚至是对抗性的)替换时保持不变。第二个定理在部分可观测条件下对所有准入机制给出了界:对于固定的探测与准入策略,若两个信息律之间的总变差距离为$ \delta $的状态假设需要不同的安全决策,则$ \abar+\beta+\delta\ge1 $。我们提出的序贯风险账本借助时间一致置信管道使该保证可实施,并表明确定性的准入计算将所有统计风险集中于状态估计。单纯形式运行时保障和控制屏障函数滤波可作为退化情形被恢复。
cs.AI / 45 / 2609.06059

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

DAREBench:面向部署的智能体模型可靠评估基准
Liu, Yu, Liu, Zhilin, Yang, Zhiwei, Zhang, Shaojie, Deng, Zheyuan, Huang, Tingwei, Luo, Zhenbo, Jiang, Lei, Liu, Yanbing, Fu, Pei
Abstract
As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception, multi-step execution, tool use, and artifact delivery. However, existing benchmarks are often tied to specific task types, execution environments, or scoring protocols, limiting their comparability, interpretability, and reliability for deployment decisions. We introduce DAREBench (Deployment-Aware and Reliable Evaluation of Models as Agents), a benchmark designed to capture workload variation and support reliable agent evaluation. Built on a shared OpenClaw execution environment, DAREBench organizes 233 tasks selected and adapted from 22 source benchmarks into a $2\times3$ workload matrix defined by input modality and execution form, and evaluates them under a unified contract-based protocol with evidence-based score auditing. We evaluate 23 commercial API models and 12 locally deployed open-weight models over 7,587 model--task runs, reporting accuracy and token consumption alongside reference costs for API models. Results show that no single model dominates all workload groups, text and multimodal tasks exhibit distinct accuracy--cost trade-offs, and local open-weight models are competitive in several groups but still trail frontier commercial models overall. These findings suggest that agent deployment and model selection should consider workload profiles, deployment mode, and accuracy--cost trade-offs rather than rely on a single aggregate score.
Chinese Translation
随着大语言模型从问答系统演进为通用智能体,评估必须超越静态的答案正确性,转向对多模态感知、多步执行、工具使用以及成果交付能力的考察。然而,现有基准往往绑定于特定的任务类型、执行环境或评分协议,限制了其在部署决策中的可比性、可解释性和可靠性。我们提出 DAREBench(Deployment-Aware and Reliable Evaluation of Models as Agents,面向部署的智能体模型可靠评估基准),该基准旨在刻画工作负载差异并支持可靠的智能体评估。DAREBench 构建于共享的 OpenClaw 执行环境之上,将从 22 个源基准中挑选并改编的 233 个任务组织为按输入模态和执行形式划分的 2×3 工作负载矩阵,并在统一的基于契约的协议下进行评估,同时提供基于证据的评分审计。我们对 23 个商业 API 模型和 12 个本地部署的开源权重模型进行了共计 7,587 次模型—任务运行测试,报告了准确率、token 消耗量以及 API 模型的参考成本。结果表明:没有任何单一模型能够在所有工作负载组中占据全面优势;文本任务与多模态任务呈现出明显不同的准确率—成本权衡;本地开源权重模型在若干组中具备竞争力,但整体上仍落后于前沿商业模型。这些发现表明,智能体部署与模型选择应综合考虑工作负载特征、部署模式以及准确率—成本权衡,而非仅依赖单一的汇总评分。
cs.AI / 46 / 2609.06063

Explaining AI Agents Through Execution Traces

通过执行轨迹解释AI智能体
Vineis, Vittoria, Veglianti, Fabiano, Antonelli, Lorenzo, Di Carlo, Claudia, Silvestri, Matteo, Tolomei, Gabriele
Abstract
AI Agents are increasingly deployed in real-world settings, where they interact with external tools and make sequential decisions with limited human oversight. This creates a pressing need for reliable and auditable explanations of what an agent did and why. However, traditional Explainable AI (XAI) methods fall short of providing the process-level transparency required for such interactive, multi-step systems, motivating a paradigm shift toward approaches specifically designed for AI Agents. To address this gap, we present a post-hoc XAI framework that transforms a lengthy agent's execution trace into a structured report and a faithful natural-language explanation explicitly grounded in its observable behavior. Because it relies solely on execution traces, the framework applies across different agent architectures, environments, and tasks. Human and automated evaluations across multiple benchmarks and architectures show that our framework produces high-quality, trace-faithful explanations while reliably identifying unsupported claims, unjustified actions, and evidence gaps, outperforming naive LLM-generated explanations.
Chinese Translation
AI智能体越来越多地被部署到真实世界场景中,它们与外部工具交互并在有限的人类监督下做出连续决策。这产生了对可靠且可审计的解释的迫切需求,以说明智能体做了什么以及为什么这样做。然而,传统的可解释人工智能(XAI)方法无法为此类交互式、多步骤系统提供所需的过程级透明度,这促使了一种专为AI智能体设计的方法的范式转变。为弥补这一空白,我们提出了一个事后(post-hoc)XAI框架,将冗长的智能体执行轨迹转化为结构化报告以及忠实的自然语言解释,且该解释明确基于智能体可观察的行为。由于该框架仅依赖执行轨迹,因此适用于不同的智能体架构、环境和任务。在多个基准测试和架构上进行的人类与自动化评估表明,我们的框架能够生成高质量、忠于轨迹的解释,同时可靠地识别无依据的声明、不合理的行动和证据缺口,其表现优于朴素的大语言模型(LLM)生成的解释。
cs.AI / 47 / 2609.06071

Generating Instance Generators in PDDL Planning

在PDDL规划中生成实例生成器
Müller, Nicola J., Rudolph, Naya, Stein, Katharina, Hoffmann, Jörg, Taitler, Ayal, Gros, Timo P.
Abstract
PDDL, the de-facto standard language in the AI Planning community, is designed to specify planning domains: sets of instances that share the same predicates and action schemas. Yet it does not provide any means to specify the actual instance set, i.e., legality constraints on initial states and goal conditions, as well as possibly domain subset constraints specifying an instance subset we are interested in. One consequence of this is that instance generation has always been ad-hoc, with manually written domain- and subset-specific instance generators. Recent work has started to address this, through reasoning and learning methods that however suffer from scalability limitations. Here we introduce an alternative approach, leveraging LLMs to generate instance-generation programs, with built-in soundness guarantees through prescribed checks. We show that these automatically generated instance generators return large numbers of sound and diverse instances efficiently.
Chinese Translation
PDDL是AI规划社区事实上的标准语言,其设计目的是描述规划域,即共享相同谓词和动作模式的一组实例。然而,它并未提供任何手段来描述实际的实例集合,即对初始状态和目标条件的合法性约束,以及可能用于指定我们所感兴趣实例子集的域子集约束。其后果之一是,实例生成一直是临时性的,依赖手动编写的针对特定域和子集的实例生成器。近期的研究工作开始着手解决这一问题,其采用的推理与学习方法存在可扩展性方面的局限。本文提出了一种替代方法,利用大语言模型(LLM)生成实例生成程序,并通过预设的检查机制内建可靠性保证。实验表明,这些自动生成的实例生成器能够高效地返回大量可靠且多样的实例。
cs.AI / 48 / 2609.06079

LayerRoute: Action-Conditioned Mixture-of-Layers Routing for Vision-Language-Action Policies

LayerRoute:面向视觉-语言-动作策略的动作条件混合层路由方法
Lu, Zheng, Liao, Haoran, Zhong, Wanqi, Ni, Yunhe, Wang, Lijie, Fan, Xingjie, Chen, Zhisheng, Qu, Yantang, Chen, Meijia, Xin, Tianyu, Song, Zirui, Li, Yiming
Abstract
Vision-Language-Action (VLA) policies leverage pretrained vision-language models (VLMs) to guide action generation for robot control. VLMs provide hierarchical visual-semantic representations that evolve across layers, from local visual geometry to abstract, language-aligned semantics; different manipulation tasks may therefore require different mixtures of layer representations. Meanwhile, the action module maintains intermediate representations that evolve throughout action computation and may provide useful information for subsequent decisions. However, existing VLA interfaces offer limited flexibility in representation access: VLM information is exposed through fixed layer assignments for each action layer, while intermediate action states are only propagated implicitly through residual streams without explicit reuse. We introduce LayerRoute, an action-conditioned representation routing interface that enables adaptive access to VLM layers and action representations. The Layer Mixture Router dynamically forms mixtures of cached VLM representations, while Action-State Reread reuses earlier action representations. Across diverse simulation and real-world benchmarks, LayerRoute consistently improves StarVLA-$\pi$ and $\pi_{0.5}$, achieving up to 7.2 gains on LIBERO Long with only 0.31% / 3.87% additional parameters. Ablation studies validate the benefit of action-conditioned layer routing, while routing analyses reveal structured allocation patterns across action layers and task settings.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)策略利用预训练的视觉-语言模型(VLM)来指导机器人控制中的动作生成。VLM 提供跨层演化的分层视觉-语义表示——从局部视觉几何到抽象的、与语言对齐的语义;因此,不同的操作任务可能需要不同层表示的混合。与此同时,动作模块在动作计算过程中维护着不断演化的中间表示,这些表示可能为后续决策提供有用信息。然而,现有的 VLA 接口在表示访问方面的灵活性有限:VLM 信息通过每个动作层固定的层配置来暴露,而中间动作状态仅通过残差流隐式传播,缺乏显式复用。我们提出了 LayerRoute,这是一个动作条件化的表示路由接口,能够自适应地访问 VLM 各层表示和动作表示。其中,层混合路由器(Layer Mixture Router)动态构建缓存 VLM 表示的混合,而动作状态重读机制(Action-State Reread)则复用较早的动作表示。在多样的仿真和真实世界基准测试中,LayerRoute 一致地提升了 StarVLA-$\pi$ 和 $\pi_{0.5}$ 的性能,在 LIBERO Long 上最多取得 7.2 的提升,且仅增加 0.31% / 3.87% 的额外参数。消融实验验证了动作条件层路由的有效性,路由分析则揭示了不同动作层和任务设置下结构化的分配模式。
cs.AI / 49 / 2609.06124

SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use

SAP:面向多轮工具调用的基于状态的参数溯源数据合成方法
Tian, Zichen, Chen, Jinpeng, Gong, Cheng, Zhang, Suiyun, Liu, Rui
Abstract
High-quality multi-turn tool-use data is essential for training agentic models, yet existing data synthesis methods often underrepresent the argument-level dependencies that are critical to long-horizon tool use. As a result, even when a model selects the correct tool, task execution may still fail because the model fills tool arguments with fabricated, stale, or weakly grounded values. To address this problem, we propose \textbf{State-Guided Data Synthesis with Argument Provenance (SAP)}. SAP combines state guidance, tool-argument provenance constraints, and turn-level validation to efficiently construct tool-use trajectories with long-range dependencies and high accuracy. Using data generated by SAP, we build SAP-4B, which is highly competitive even when compared with much larger models across multiple benchmarks. Source code, synthesized data, and trained weights are available at https://github.com/Zichen1024/SAP.
Chinese Translation
高质量的多轮工具调用数据对于训练智能体(agentic)模型至关重要,然而现有的数据合成方法往往未能充分体现对长程工具调用至关重要的参数级依赖关系。其结果是,即使模型选择了正确的工具,任务执行仍可能失败,因为模型在填充工具参数时使用了虚构的、过时的或缺乏充分依据的值。为解决这一问题,我们提出了基于状态的参数溯源数据合成方法(State-Guided Data Synthesis with Argument Provenance,SAP)。SAP 结合状态引导、工具参数溯源约束和轮次级验证,高效构建具有长程依赖关系和高准确率的工具调用轨迹。利用 SAP 生成的数据,我们构建了 SAP-4B 模型,该模型在多个基准测试中即使与规模大得多的模型相比也极具竞争力。源代码、合成数据和训练权重已发布于 https://github.com/Zichen1024/SAP。
cs.AI / 50 / 2609.06126

CWF: A Collaborative Writing Framework for Personalized and Reliable Popular Science Writing

CWF:一个用于个性化且可靠的科普写作的协同写作框架
Fu, Ruibiao, Tang, Di, Yang, Yunlong, Wang, Ran, Lu, Sicheng, Wu, Peixuan, Fan, Xiaoyu, Ma, Jiacheng, Luo, HaoZhe, Xiao, Yang
Abstract
We introduce Personalized and Reliable Popular Science Writing, a novel task that requires adapting scientific explanations to audiences with different cognitive levels while preserving factual accuracy. However, improving personalization often introduces simplifications that increase the risk of hallucination and factual distortion. To address these challenges, we first construct a dataset of 39,134 entries and a reader-centric Personalized Science Communication Benchmark (PSCB) that jointly evaluates audience adaptation and factual accuracy. To reduce data and computational requirements while improving generalization across domains and audiences, we introduce DA-MoE, which explicitly decouples audience adaptation from domain knowledge through separate modeling. To enable robust verification and revision in evidence-scarce scenarios, a multi-agent fact-checking mechanism that augments limited evidence with role-specific agent debate and propagates confidence over a graph is proposed. Experiments on PSCB show that our approach achieves state-of-the-art performance. Our code is open-sourced at https://github.com/DPInnovationWorks/CWF.
Chinese Translation
我们提出了个性化且可靠的科普写作这一新任务,该任务要求在保持事实准确性的同时,将科学解释适配到具有不同认知水平的受众。然而,提升个性化程度往往会引入过度简化,从而增加幻觉和事实失真的风险。为应对这些挑战,我们首先构建了一个包含39,134条数据的数据集,以及一个以读者为中心的个性化科学传播基准(Personalized Science Communication Benchmark, PSCB),用于联合评估受众适配能力和事实准确性。为了降低数据和计算需求并提升跨领域、跨受众的泛化能力,我们提出了DA-MoE,该模型通过分离建模显式地将受众适配与领域知识解耦。为了在证据稀缺的场景下实现可靠的验证与修订,我们提出了一种多智能体事实核查机制,该机制通过具有特定角色的智能体辩论来补充有限的证据,并在图上传播置信度。在PSCB上的实验表明,我们的方法取得了最先进的性能。我们的代码已开源于 https://github.com/DPInnovationWorks/CWF。
cs.AI / 51 / 2609.06128

Substrate-Portable Execution for Production LLM Workflows

面向生产级大语言模型工作流的基底可移植执行
Gopinath, Tarun, Kulkarni, Atul, Rajakumar, Vijay, Katti, Shrikar, Govindarajen, Parthasarathy
Abstract
Production LLM agents execute tool-calling loops, retrieval chains, and compositional workflows in multiple modes, yet execution semantics are often coupled to one runtime. We encountered this portability problem in Rufus, a conversational AI assistant with a large tool catalog that serves millions of Amazon customers. Rufus supports real-time serving, asynchronous background tasks, and high-volume batch workloads such as evaluation and content pregeneration. Each mode has distinct service-level objectives and typically uses a separate runtime. Reusing streaming orchestration makes asynchronous and batch workloads blocking and prevents use of batch inference APIs, which offer a 50 percent discount at published prices. We present a binding-adaptive agent execution platform that separates workflow definition from execution substrate. Developers define a workflow once as a typed dataflow graph. The platform compiles the graph to in-process streaming for real-time serving, durable AWS SWF orchestration for asynchronous execution, or distributed Apache Flink stream processing for batch inference. No workflow code changes are required. LLM inference is represented as a suspendable graph node whose behavior depends on the substrate: streaming delivery online, durable retry asynchronously, and batched submission offline. We validated dozens of production agent configurations across five orchestration patterns: single-inference RAG, iterative ReAct, compositional PreAct, conditional routing, and multi-agent deep research. Across all three bindings, we found no detectable difference in output quality. Batch execution reduced per-query inference cost in line with published batch API pricing while operating alongside the streaming path at production scale.
Chinese Translation
生产级大语言模型(LLM)智能体以多种模式执行工具调用循环、检索链和组合式工作流,但其执行语义往往与单一运行时耦合。我们在 Rufus 中遇到了这一可移植性问题——Rufus 是一个拥有庞大工具目录、服务数百万亚马逊客户的对话式 AI 助手。Rufus 支持实时服务、异步后台任务,以及评估和内容预生成等高吞吐量批量工作负载。每种模式具有不同的服务级别目标,且通常使用独立的运行时。复用流式编排会使异步和批量工作负载变为阻塞式,并阻碍对批量推理 API 的使用,而后者相对公布价格可享受五折优惠。我们提出了一种绑定自适应的智能体执行平台,将工作流定义与执行基底相分离。开发者只需以类型化数据流图的形式定义一次工作流,平台即可将该图编译为三种执行方式:用于实时服务的进程内流式执行、用于异步执行的持久化 AWS SWF 编排,或用于批量推理的分布式 Apache Flink 流处理,且无需修改任何工作流代码。LLM 推理被表示为一个可挂起的图节点,其行为取决于所用基底:在线时采用流式交付,异步时采用持久化重试,离线时采用批量提交。我们在五种编排模式(单次推理 RAG、迭代式 ReAct、组合式 PreAct、条件路由和多智能体深度研究)上验证了数十个生产级智能体配置。在全部三种绑定下,输出质量均未检测到差异。批量执行使每查询推理成本与公布的批量 API 定价相符地降低,同时能在生产规模下与流式路径并行运行。
cs.AI / 52 / 2609.06131

IIns-VAE+: A Robust Transfer Learning Framework for Environmental Identification in Wireless Sensing

IIns-VAE+:一种用于无线感知环境识别的鲁棒迁移学习框架
Li, Yuxiao, Hu, Keke, Zhao, Bobai, Mazuelas, Santiago, Shen, Yuan
Abstract
Environmental identification in wireless sensing is essential for 6G integrated sensing and communication (ISAC) systems to achieve reliable situational awareness. However, deep learning (DL) models for this task often fail to generalize under domain shift across diverse environments. While the Inter-Instance Variational Auto-encoder (IIns-VAE) learns features of rich representation, its neural classifier remains vulnerable to these distribution changes. In this paper, we propose IIns-VAE+, a hybrid model that combines the IIns-VAE framework with Minimax Risk Classifiers (MRC) to improve adaptability in transfer learning scenarios. We use real-world datasets to evaluate our framework across three transfer learning scenarios, including general to specific room environments, high to low label resolutions, and mixed to specific environments. The experimental results indicate that IIns-VAE+ significantly outperforms baselines, demonstrating its critical value in building adaptable and robust perceptive networks in future 6G systems.
Chinese Translation
无线感知中的环境识别对于6G通感一体化(ISAC)系统实现可靠的态势感知至关重要。然而,用于该任务的深度学习(DL)模型在跨不同环境的域偏移下往往难以泛化。尽管实例间变分自编码器(IIns-VAE)能够学习具有丰富表示能力的特征,但其神经分类器仍然容易受到这些分布变化的影响。本文提出了IIns-VAE+,这是一种将IIns-VAE框架与极小极大风险分类器(MRC)相结合的混合模型,旨在提高迁移学习场景中的适应性。我们使用真实世界数据集,在三种迁移学习场景下评估了所提出的框架,包括通用到特定房间环境、高到低标签分辨率以及混合到特定环境。实验结果表明,IIns-VAE+显著优于基线方法,证明了其在未来6G系统中构建自适应且鲁棒的感知网络方面的关键价值。
cs.AI / 53 / 2609.06188

MVFA: A Multi-View Text-Guided Multimodal Fusion LLM Adapter for Sentiment Analysis and Emotion Recognition

MVFA:面向情感分析与情绪识别的多视图文本引导多模态融合大语言模型适配器
Shao, Pengfei, Dang, Jisheng, Fang, Jiawen, Liu, Ning, Zhang, Wencan, Wang, Bimei, Zhao, Jingwen, Lai, Jianhuang, Tian, Qi, Chua, Tat-Seng
Abstract
Multimodal sentiment analysis and emotion recognition in conversations demand effective modeling of heterogeneous interactions across textual, acoustic, and visual modalities. Although large language models (LLMs) offer powerful language understanding, adapting them to multimodal affective computing remains challenging: full-model fine-tuning is computationally prohibitive, while many existing lightweight adapters fail to preserve rich textual cues during cross-modal fusion. To address these limitations, we propose the multi-view text-guided multimodal fusion adapter (MVFA), a parameter-efficient framework that augments frozen LLMs with strong multimodal reasoning capability. MVFA first constructs complementary text views via max pooling, mean pooling, and attention pooling; these views then guide cross-modal interactions with audio and visual features. The fused multimodal representations are subsequently compressed into a compact set of learnable pseudo-tokens through an Enhanced Q-Former Fusion Module. Using ChatGLM3-6B-base as the primary backbone, we further validate MVFA on LLaMA2-7B and Qwen3-8B to examine its portability across multiple frozen LLM backbones. MVFA is evaluated on three challenging datasets: CH-SIMS V2.0, MELD, and CHERMA. Experimental results demonstrate that MVFA achieves state-of-the-art performance on key metrics while updating only a small fraction of parameters. Specifically, it attains 84.62\% Acc2 and 84.59\% F1 on CH-SIMS V2.0, 67.36\% Acc and 66.03\% WF1 on MELD, and 74.66\% Acc on CHERMA. These findings establish multi-view text-guided fusion as an effective and scalable paradigm for parameter-efficient multimodal LLM adaptation in affective computing. The code is publicly available at https://github.com/Overwhelm1208/MVFA.
Chinese Translation
多模态情感分析与会话中的情绪识别需要对文本、声学和视觉模态之间的异构交互进行有效建模。尽管大语言模型(LLM)具备强大的语言理解能力,但将其适配到多模态情感计算仍面临挑战:全模型微调的计算成本过高,而许多现有的轻量级适配器在跨模态融合过程中难以保留丰富的文本线索。针对这些局限,我们提出了多视图文本引导多模态融合适配器(MVFA),这是一种参数高效框架,可为冻结的大语言模型赋予强大的多模态推理能力。MVFA 首先通过最大池化、平均池化和注意力池化构建互补的文本视图;随后这些视图引导文本与音频、视觉特征之间的跨模态交互。融合后的多模态表示随后通过增强型 Q-Former 融合模块(Enhanced Q-Former Fusion Module)被压缩为一组紧凑的可学习伪标记(pseudo-tokens)。我们以 ChatGLM3-6B-base 作为主要骨干网络,并在 LLaMA2-7B 和 Qwen3-8B 上进一步验证 MVFA,以考察其在多个冻结大语言模型骨干上的可移植性。MVFA 在三个具有挑战性的数据集上进行了评估:CH-SIMS V2.0、MELD 和 CHERMA。实验结果表明,MVFA 仅需更新少量参数即可在关键指标上取得最先进的性能。具体而言,该方法在 CH-SIMS V2.0 上达到 84.62% 的 Acc2 和 84.59% 的 F1,在 MELD 上达到 67.36% 的 Acc 和 66.03% 的 WF1,在 CHERMA 上达到 74.66% 的 Acc。这些发现确立了多视图文本引导融合作为情感计算中参数高效的多模态大语言模型适配的一种有效且可扩展的范式。代码已在 https://github.com/Overwhelm1208/MVFA 公开。
cs.AI / 54 / 2609.06189

Customer Relationship Intelligence: Integrating CRM and MDM for Enhanced Customer Engagement

客户关系智能:整合CRM与MDM以增强客户参与度
Addagada, Tejasvi c.
Abstract
This study examines how Customer Relationship Management (CRM), Master Data Management (MDM), and Customer Knowledge Management (CKM) jointly constitute a Customer Relationship Intelligence (CRI) framework for enhanced Customer Engagement (CE). A cross-sectional survey of 100 participants across retail, healthcare, IT, and telecommunications sectors was analysed using Spearman rho correlation and ordinal logistic regression (IBM SPSS). Bivariate correlations were weak and non-significant (r<0.19, p>0.06). Regression identified CRM (beta=0.717, p=0.002) and CKM (beta=0.581, p=0.009) as significant positive predictors of CE; MDM showed a positive but non-significant direct effect (beta=0.346, p=0.071). The model explained 20.5% of CE variance (Nagelkerke R^2=0.205). Parallel mediation analysis (Hayes PROCESS Model 4, 5,000 bootstrap samples) found no significant indirect effects of MDM on CE via CRM (IE=0.021, 95% BC CI [-0.072, 0.121]) or CKM (IE=0.032, 95% BC CI [-0.061, 0.126]); Hypothesis H4 was not supported. CRM and CKM emerge as the principal drivers of CE within the CRI framework, while MDM functions as a foundational data quality enabler whose strategic value is realised through its enabling effect on CRM execution and knowledge management. Findings should be treated as exploratory given the sample size and cross-sectional design. Future research should replicate with larger sector-specific samples and longitudinal designs, particularly in regulated BFSI contexts where MDM architecture is shaped by data governance mandates.
Chinese Translation
本研究考察了客户关系管理(CRM)、主数据管理(MDM)和客户知识管理(CKM)如何共同构成客户关系智能(CRI)框架,以增强客户参与度(CE)。研究对来自零售、医疗、IT和电信行业的100名参与者进行了横断面调查,并采用Spearman rho相关分析和有序逻辑回归(IBM SPSS)进行分析。双变量相关性较弱且不显著(r<0.19,p>0.06)。回归分析表明,CRM(beta=0.717,p=0.002)和CKM(beta=0.581,p=0.009)是CE的显著正向预测因子;MDM表现出正向但不显著的直接效应(beta=0.346,p=0.071)。该模型解释了CE方差的20.5%(Nagelkerke R^2=0.205)。平行中介分析(Hayes PROCESS Model 4,5,000次bootstrap抽样)发现,MDM通过CRM(IE=0.021,95% BC CI [-0.072, 0.121])或CKM(IE=0.032,95% BC CI [-0.061, 0.126])对CE的间接效应均不显著;假设H4未获支持。在CRI框架中,CRM和CKM是CE的主要驱动因素,而MDM则作为基础性的数据质量使能要素,其战略价值通过对CRM执行和知识管理的赋能作用得以实现。鉴于样本量和横断面设计的局限,研究结果应被视为探索性的。未来研究应在更大规模的行业特定样本和纵向设计中进行重复验证,特别是在受监管的银行、金融服务与保险(BFSI)领域,因为数据治理要求塑造了MDM架构。
cs.AI / 55 / 2609.06192

SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores

SciRIGOR:超越最终得分的开放式科学分析评估
Liu, Bowen, Nie, Shuo, Du, Bodong, Li, Xiaomeng
Abstract
Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether their conclusions are scientifically supported. We formulate evidence-grounded multimodal scientific analysis, requiring agents to produce executable analyses and claims supported by results and visualizations from the same run. We introduce SciRIGOR, an evaluation framework and benchmark comprising 100 cases from scientific articles across six domains and 17 subfields. The framework reconstructs typed evidence graphs, separates artifact fidelity from relational validity, and scores complete claim-support paths while localizing the earliest unsupported relation. Source-grounded alternative paths accommodate scientifically equivalent analyses and visualizations. We evaluate 11 agent/model configurations. On full-benchmark runs, claims agree with faithful and unfaithful results at nearly identical rates (91.8% versus 91.0%). Yet no system exceeds 62.6% on the soft evidence-chain score or 18.0% strict whole-chain success. These findings show that internal coherence does not establish scientific correctness: evaluation must verify support along the complete data-to-claim path.
Chinese Translation
科学编码智能体(scientific coding agents)会产生相互依赖的代码、结果、图表和论断,然而仅评估最终输出并不能确定其结论是否具有科学依据。我们提出了基于证据的多模态科学分析这一任务形式,要求智能体产出可执行的分析,以及由同一次运行所得结果和可视化支持的论断。我们介绍了SciRIGOR,一个评估框架与基准,包含来自六个领域、17个子领域的科学文献中的100个案例。该框架重建了类型化的证据图,将产物保真度与关系有效性区分开来,对完整的论断支持路径进行评分,同时定位最早出现无支持关系的位置。基于来源的替代路径机制可容纳科学上等价的分析与可视化。我们评估了11种智能体/模型配置。在完整基准运行中,论断与可信结果和不可信结果达成一致的比率几乎相同(91.8%对91.0%)。然而,没有任何系统在软证据链得分上超过62.6%,或在严格的全链成功率上超过18.0%。这些发现表明,内部一致性并不能确立科学正确性:评估必须沿从数据到论断的完整路径验证其支持关系。
cs.AI / 56 / 2609.06194

Predicting Wind Turbine Power Using Machine Learning and Weather Forecasts

基于机器学习和天气预报的风力涡轮机功率预测
Boodhoo, Khivishta, Triguero, Isaac, Plumbly, Josh, Nicolson, Bruce, Watson, Nicholas
Abstract
Offshore wind turbines are widely used to generate renewable energy, but their maintenance can result in decreased efficiency due to forced shutdowns. Accurate wind turbine power predictions can identify periods of low power that would be ideal for scheduling maintenance. However, the effects of data volume, feature selection, and data preprocessing on the performance of such power prediction models have not been thoroughly studied. Besides, current models have limited transferability between different wind turbines. Therefore, this study developed a baseline Linear Regression for performance comparison with a more complex Artificial Neural Network model to predict the power output of a wind turbine, using weather conditions only to enhance applicability. A range of data preprocessing techniques were studied, and models were trained on one month and one year of data to determine the effects of data preprocessing and volume on model performance. Feature selection was explored using a Random Forest Regressor. The best results from the different models showed that the Artificial Neural Network models provided the highest accuracy, with an R2 score of 0.98 and a low Mean Absolute Error of 194, when compared with the baseline model (R2 score of 0.94 and Mean Absolute Error of 441). The model performance is comparable to the range of results in past studies, with the advantage that the proposed method leverages a separate weather dataset from a nearby weather station, enabling future applications for similar wind turbines in different locations. The Artificial Neural Network model was then used to identify 4-h periods of low power predictions over 2 months (simulating application for future periods), providing power output savings of approximately 2000 kW for each maintenance event.
Chinese Translation
海上风力涡轮机被广泛用于可再生能源发电,但其维护所需强制停机会导致发电效率下降。准确的风机功率预测可以识别出低功率时段,这些时段是安排维护的理想时机。然而,数据量、特征选择和数据预处理对此类功率预测模型性能的影响尚未得到充分研究。此外,现有模型在不同风力涡轮机之间的可迁移性有限。因此,本研究构建了一个基线线性回归(Linear Regression)模型,并与更复杂的人工神经网络(Artificial Neural Network)模型进行性能比较,以预测风力涡轮机的功率输出,且仅使用天气条件作为输入以增强模型的适用性。本研究考察了一系列数据预处理技术,并分别使用一个月和一年的数据训练模型,以确定数据预处理和数据量对模型性能的影响。特征选择通过随机森林回归器(Random Forest Regressor)进行探索。不同模型的最优结果表明,与基线模型(R²分数为0.94,平均绝对误差为441)相比,人工神经网络模型具有最高的精度,R²分数为0.98,平均绝对误差仅为194。本模型的性能与以往研究结果的范围相当,其优势在于所提出的方法利用了来自附近气象站的独立天气数据集,从而使该方法未来可应用于位于不同地区的类似风力涡轮机。随后,人工神经网络模型被用于识别未来两个月内的4小时低功率预测时段(模拟对未来时段的应用),每次维护事件可节省约2000 kW的功率输出。
cs.AI / 57 / 2609.06366

AutoKD: Autonomous Knowledge Discovery

AutoKD:自主知识发现
Ge, Qinwen, Ni, Bo, Fu, Haowei, Tran, Ngoc N., Blasch, Erik, Derr, Tyler
Abstract
Scientific discovery in data-rich domains is currently constrained by human bandwidth: the growth in the volume and complexity of real-world data far outpaces the rate at which researchers can read, reason, and synthesize. Recent LLM-based multi-agent systems have begun to automate portions of the research cycle, but they target hypothesis generation in settings where validation cannot itself be automated, and each run is one-shot, with no mechanism for findings to accumulate or steer subsequent inquiry. This paper introduces AutoKD, a multi-agent framework for autonomous knowledge discovery that is both computational and cumulative, allowing validated findings to persist and inform subsequent inquiry. Six coordinated LLM agents collaborate in an open-ended discovery loop, where accepted findings are stored in a persistent insight graph that serves as both long-term memory and an exploration-steering mechanism. We evaluate AutoKD on three diverse datasets from two perspectives: Open-ended Quality against published findings, and Conditioned Quality via literature-derived queries. Across both evaluation perspectives, AutoKD covers known findings and surfaces substantive discoveries that complement human-driven research. Our code is available at https://github.com/GeQinwen/AutoKD.
Chinese Translation
当前,数据丰富的领域中的科学发现受到人类带宽的制约:真实世界数据的规模和复杂性的增长速度远远超过了研究人员阅读、推理和综合的速度。近期基于大语言模型(LLM)的多智能体系统已经开始将研究周期的部分环节自动化,但它们仅针对假设生成这一验证本身无法自动化的场景,且每次运行都是一次性的,缺乏使研究发现得以积累或引导后续探究的机制。本文提出了AutoKD,一个兼具计算性与累积性的自主知识发现多智能体框架,使经过验证的发现得以持久保存并为后续探究提供参考。六个协同工作的LLM智能体在一个开放式发现循环中协作,其中被接受的发现存储于一个持久的洞察图谱(insight graph)中,该图谱既充当长期记忆,也作为引导探索方向的机制。我们从两个角度在三个不同的数据集上对AutoKD进行了评估:一是开放式质量(Open-ended Quality),与已发表的研究发现进行对比;二是条件化质量(Conditioned Quality),通过基于文献衍生的查询进行评估。在这两种评估视角下,AutoKD均能覆盖已知发现,并揭示出能够补充人类驱动研究的实质性新发现。我们的代码可在 https://github.com/GeQinwen/AutoKD 获取。
cs.AI / 58 / 2609.06391

Building Trustworthy Graph-Agentic RAG for Social Good: Architectures, Failure Propagation, and Assurance by Construction

构建面向社会公益的可信图智能体检索增强生成:架构、失效传播与构建式保障
Bommireddy, Vijay, Bommireddy, Raviteja
Abstract
Graph-agentic retrieval-augmented generation combines structured evidence with adaptive controllers that can plan retrieval, traverse relations, verify intermediate claims, delegate subtasks, and use tools. This combination is useful when answers depend on relations across documents, entities, time, or institutions, but it also creates coupled failure paths: a defect in graph construction can become retrieved evidence, alter later control decisions, and propagate toward a consequential outcome. We examine how such systems should be designed and evaluated for social-good settings in which freshness, authorization, traceability, oversight, and recourse matter alongside answer quality. We organize the literature by graph substrate, graph lifecycle, agent function, coordination pattern, and authority boundary, and distinguish graph-based retrieval from observation-dependent graph control. We then synthesize reported risks as an evidence-to-action failure chain and propose an assurance-by-construction blueprint comprising five interface contracts for evidence, retrieval, reasoning, capability and delegation, and outcome. These contracts make provenance, temporal validity, authorization, uncertainty, and recoverability explicit at system boundaries. An illustrative public-benefit information design shows how the framework constrains graph structure, permissions, abstention, and operating authority. Finally, we derive an evaluation agenda spanning graph assertions, trajectories, claims, coordination, and outcomes.
Chinese Translation
图智能体检索增强生成(Graph-agentic RAG)将结构化证据与自适应控制器相结合,这些控制器能够规划检索、遍历关系、验证中间论断、委派子任务并调用工具。当答案依赖于跨文档、实体、时间或机构的关系时,这种组合非常有用,但它也产生了耦合的失效路径:图构建中的缺陷可能成为被检索的证据,改变后续的控制决策,并传播至具有重大影响的最终结果。我们研究了此类系统在涉及社会公益的场景中应如何设计与评估——在这些场景中,时效性、授权、可追溯性、监督和补救机制与答案质量同等重要。我们按照图基底、图生命周期、智能体功能、协调模式和权限边界对文献进行组织,并区分了基于图的检索与依赖观测的图控制。随后,我们将已报告的风险综合为一个“证据到行动”的失效链条,并提出一个构建式保障蓝图,其中包含针对证据、检索、推理、能力与委派以及结果五个方面的接口契约。这些契约使来源溯源性、时间有效性、授权、不确定性和可恢复性在系统边界处显式化。一个公益信息设计的示例展示了该框架如何约束图结构、权限、弃答机制和运行权限。最后,我们推导出一个涵盖图断言、轨迹、论断、协调和结果的评估议程。
cs.AI / 59 / 2609.06403

From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning Cohorts

从集中到分化再回归:MoE推理群体中路由有效秩的演化
Chen, Kang, Zhao, Sihan, Cao, Yixin, Jiang, Yu-Gang
Abstract
Test-time scaling produces cohorts of reasoning rollouts, yet there is no standard label-free account of how their internal computation reorganizes as inference unfolds. We introduce routing effective rank deff, the entropy-effective dimensionality of a cross-rollout graph built from MoE expert-routing similarity. Across ten MoE configurations and five math/science benchmarks, deff exhibits a reproducible low-high-low trajectory, with a prominent interior maximum in 98.5% of 3,105 model-question cohorts: routing similarity is concentrated early, maximally differentiated at intermediate budgets, and reconcentrated later, and the timing of this maximum varies systematically with architecture and reasoning effort. An exact decomposition separates cohort-wide common-mode mass from residual spectral dimensionality: common-mode reallocation accounts for about two thirds of the trajectory, while the residual spectrum contributes about one quarter and retains substantial variation beyond the common mode. The decomposition further localizes behavior: among non-unanimous cohorts, increases in common-mode concentration strongly predict same-answer recoverability, and higher reasoning effort delays the maximum by 2.59 octaves (doublings of the token budget) and consistently expands the high-rank period across all four tested architectures, locating the effort effect in timing and duration rather than peak amplitude. Correctness comparisons separate structural monitoring from answer selection, positioning routing effective rank as a decomposable, label-free diagnostic of cohort organization - a principled spectral lens on how MoE reasoning cohorts differentiate and reconcentrate over inference time.
Chinese Translation
测试时扩展(test-time scaling)会产生由多个推理轨迹组成的群体(cohorts),然而对于推理过程中其内部计算如何重组,目前尚无标准的无标签刻画方法。我们提出了路由有效秩(routing effective rank, deff),即基于MoE专家路由相似性构建的跨轨迹图的熵有效维度。在十个MoE配置和五个数学/科学基准上,deff呈现出可复现的“低-高-低”轨迹,在3,105个模型-问题群体中有98.5%出现显著的内部最大值:路由相似性在早期集中,在中等推理预算时达到最大分化,随后再次集中,且该最大值出现的时机随架构和推理努力程度(reasoning effort)系统性变化。通过一个精确分解,可将群体范围的共模质量(common-mode mass)与残差谱维度分离:共模重新分配约占轨迹变化的 thirds(三分之二),而残差谱贡献约四分之一,并在共模之外保留了可观的变异。该分解进一步定位了行为特征:在非全票一致的群体中,共模集中度的提升强烈预测同答案可恢复性;更高的推理努力程度使最大值延迟2.59个倍频程(即token预算的翻倍数),并在所有四种测试架构上一致地延长了高秩阶段的持续时间——这表明推理努力的影响体现在时机和持续时间上,而非峰值幅度。正确性比较将结构性监测与答案选择相分离,使路由有效秩成为一种可分解、无标签的群体组织诊断工具——它是观察MoE推理群体在推理时间内如何分化与再集中的一把有原则的谱分析透镜。
cs.AI / 60 / 2609.06445

Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification

智能体决策的因果归因:估计量、耦合与可追溯性规范
Mahale, Ajay Pravin
Abstract
A provider of a high-risk AI system must keep records that make a decision traceable, and for agentic systems it has not been established what those records must contain for post-hoc causal attribution to be possible. We give the estimator framework and then the conditions under which it fails. We separate the marginal total effect that prior work measures from a common-random-numbers total effect that isolates a step's own contribution, add the natural direct effect under a pinned downstream, and check the estimators against hand derivations. Both estimands then fail, in the same direction. Under the marginal estimand a causally inert step has the identical total effect to the decisive one on every run of our planted chain, an algebraic identity and not a coincidence at one draw. Under common random numbers the decisive step returns exactly zero on the runs where the executing step flips, about one in ten, while its direct effect there is 0.25 and it demonstrably acts; an exact zero does not certify that a step did nothing, and we put that here rather than in the limitations. We derive the coupling that keeps the direct effect estimable once contexts diverge, with a closed form for its degradation, and show that the mediated share on which a natural ranking is built is not a share under suppression: where the direct and mediated paths oppose, it exceeds one and ranks a suppressed component above a pure mediator. We publish the discrepancy experiment's pre-registration rather than a result, because the live pipeline it requires was not available in the study window. We contribute the traceability specification such a filing would need, against a gap the Act's calendar opens: Article 86's right to an explanation has applied since 2 August 2026, while the Article 12 logging and Annex IV documentation that could evidence one were deferred to 2 December 2027 by Regulation (EU) 2026/1744.
Chinese Translation
高风险AI系统的提供者必须保存使决策可追溯的记录,而对于智能体系统,为使事后因果归因成为可能,这些记录必须包含哪些内容尚未确立。我们给出了估计量框架,随后给出其失效的条件。我们将先前研究所测量的边际总效应与隔离某一步骤自身贡献的公共随机数总效应区分开来,增加了下游固定条件下的自然直接效应,并将估计量与手工推导进行核对。两种估计量随后以相同方向失效。在边际估计量下,在我们植入的链条的每一次运行中,因果惰性的步骤与决定性步骤具有完全相同的总效应——这是一个代数恒等式,而非某一次抽样的巧合。在公共随机数下,决定性步骤在执行步骤发生翻转的那些运行(约占十分之一)中恰好返回零,而其直接效应在该处为0.25,且可证明它确实起了作用:精确的零并不能证明某一步骤没有起作用,我们将这一点放在正文而非局限性部分中。我们推导了当上下文发散后仍保持直接效应可估计的耦合方式,并给出其退化的闭式表达式;同时表明自然排序所依赖的中介份额在抑制效应下并非真正的份额:当直接路径与中介路径方向相反时,该份额会超过1,并将一个被抑制的组件排在纯中介之上。我们发布的是该差异实验的预注册而非结果,因为其所需的实时流水线在研究窗口期内不可用。我们贡献了此类备案所需的可追溯性规范,以填补《法案》日程表造成的空档:《欧盟AI法案》第86条的解释权自2026年8月2日起已生效,而本可为解释提供证据的第12条日志要求及附件IV的文档要求,则被(欧盟)2026/1744号条例推迟至2027年12月2日。
cs.AI / 61 / 2609.06543

A Unified Policy Architecture (UPA): The Governance Kernel for Enterprise AI Operating Systems

统一策略架构(UPA):企业AI操作系统的治理内核
Raghav, Prabhu, Pandi, Balamurugan, Vivek, Arul, Mohammed, Shek, S, Sridhar
Abstract
Enterprise AI is evolving into an Enterprise Operating System where autonomous AI agents can plan, reason, use memory, invoke tools, execute workflows, and collaborate with other agents. This shift creates a new governance challenge: existing authorization, security, guardrails, and compliance mechanisms are fragmented and are not designed to govern autonomous AI as a unified system. This paper introduces the Unified Policy Architecture (UPA), a governance architecture for Enterprise AI Operating Systems. UPA provides a unified policy model for governing AI and agents, tools, workflows, memory, enterprise resources, and agent-to-agent interactions and enterprise business rules. It extends policy control beyond authorisation to include runtime obligations, human approvals, compliance, audit evidence, and governance evaluation. We present UPA's governance model, declarative policy language foundations, policy evaluation semantics, extensible plugins, industry policy packs, and an evaluation framework for enterprise governance. We also identify extensions for multi-agent coordination, provenance-aware policies, and stateful runtime governance. UPA provides a foundation for building secure, accountable, and governable Enterprise Operating Systems for autonomous AI.
Chinese Translation
企业AI正在演化为一种企业操作系统,其中自主AI智能体能够进行规划、推理、使用记忆、调用工具、执行工作流,并与其他智能体协作。这一转变带来了新的治理挑战:现有的授权、安全、护栏和合规机制是碎片化的,并非为将自主AI作为一个统一系统进行治理而设计。本文提出了统一策略架构(Unified Policy Architecture, UPA),一种面向企业AI操作系统的治理架构。UPA提供了一个统一的策略模型,用于治理AI与智能体、工具、工作流、记忆、企业资源、智能体间交互以及企业业务规则。它将策略控制从授权扩展至运行时义务、人工审批、合规、审计证据和治理评估。我们阐述了UPA的治理模型、声明式策略语言基础、策略评估语义、可扩展插件、行业策略包以及企业治理评估框架。我们还识别了在多智能体协作、溯源感知策略和有状态运行时治理方面的扩展方向。UPA为构建安全、可问责、可治理的自主AI企业操作系统奠定了基础。
cs.AI / 62 / 2609.06563

MARBO: Relational Belief Grounding for LLM Agents in Social Deduction Games

MARBO:面向社交推理游戏中LLM智能体的关系性信念锚定方法
Yechan, Hwang, Sangjun, Bae, Jeongmo, Kim, Sangwoo, Bang, Seungyul, Han
Abstract
Social deduction games (SDGs) require agents to reason under partial observability by maintaining relational beliefs about hidden roles and team alignments. While recent LLM-agent approaches improve gameplay through prompting and preference optimization, they often optimize actions and in-game speech without explicitly grounding them in such beliefs. This frequently leads to strategically inconsistent behavior, especially for compact LLM agents. We introduce Multi-Agent Relational Belief Optimization (MARBO), a belief-grounded preference optimization framework that leverages relational beliefs to guide strategic decisions and in-game speech. MARBO provides preference feedback only when behaviors are supported by reliable relational beliefs and lead to strategically favorable social outcomes, encouraging more consistent learning under uncertainty. Experiments on representative SDGs show that MARBO enables compact LLM agents to consistently outperform existing baselines. The Code is available on https://github.com/PleaseTakemeAway/MARBO.
Chinese Translation
社交推理游戏(Social Deduction Games, SDGs)要求智能体在部分可观测的环境下,通过维持关于隐藏角色和阵营归属的关系性信念进行推理。尽管近期的LLM智能体方法通过提示工程和偏好优化提升了游戏表现,但它们通常在优化行动和游戏内发言时并未将其显式锚定于此类信念之上。这常常导致策略上不一致的行为,对于轻量级LLM智能体尤为明显。我们提出了多智能体关系性信念优化(Multi-Agent Relational Belief Optimization, MARBO),这是一种基于信念锚定的偏好优化框架,利用关系性信念来指导策略决策和游戏内发言。MARBO仅当行为得到可靠关系性信念支持并带来策略上有利的社交结果时,才提供偏好反馈,从而鼓励在不确定性下进行更一致的学习。在代表性社交推理游戏上的实验表明,MARBO使轻量级LLM智能体能够持续超越现有基线方法。代码已发布于 https://github.com/PleaseTakemeAway/MARBO。
cs.AI / 63 / 2609.06573

A Translational Note on AI Safety Evaluation

关于人工智能安全评估的译注
Gaikwad, Madhava
Abstract
Recent studies report that automated red-teaming finds more vulnerabilities, at lower cost, than human red-teaming on standard AI safety benchmarks, and some read this as evidence that human evaluators are becoming dispensable. The comparison measures one thing and the conclusion claims another. A benchmark measures how thoroughly an attacker searches a predefined set of harms, fixed in advance by the developers, and a harm left out of that set is invisible to any attacker working inside it, automated or not. The same blind spot appeared in academic cryptography and in clinical drug trials, where an evaluation that was internally valid stayed silent about the population it was never pointed at. We call the AI-safety version the \emph{threat-model coverage gap}, and find that it persists in a current open-weight model, where harms surface in non-English prompts that English benchmarks miss. Closing it requires evaluators whose deployment context differs from the developers'. The case for those evaluators is methodological, grounded in coverage, and the existing evaluation frame is unlikely to produce them on its own.
Chinese Translation
近期研究报道,在标准人工智能安全基准测试中,自动化红队测试比人工红队测试能以更低的成本发现更多漏洞,一些解读将其视为人类评估者正变得可有可无的证据。然而,这种比较衡量的是一回事,其结论宣称的却是另一回事。基准测试衡量的是攻击者在开发人员预先设定的危害集合中搜索的彻底程度,而任何在此集合之内开展工作的攻击者——无论是否自动化——都无法察觉被排除在该集合之外的危害。同样的盲点曾出现在学术密码学和临床药物试验中:一个内部有效的评估,对于它从未指向的人群保持沉默。我们将这种人工智能安全领域的情形称为“威胁模型覆盖缺口”(threat-model coverage gap),并发现该缺口在当前的一个开源权重模型中依然存在:一些危害在非英语提示中显现,而英语基准测试却未能捕捉。弥合这一缺口需要其部署环境与开发人员不同的评估者。支持这类评估者的理由是方法论层面的、植根于覆盖范围论证的,而现有的评估框架本身不太可能自然而然地产生这样的评估者。
cs.AI / 64 / 2609.06647

SerenAI: State-transition system inspired by text-based world AI models

SerenAI:一种受文本世界AI模型启发的状态转移系统
Babayev, Elvin, Sinitsa, Artem, Hajisharifi, Arash, Bakhshaei, Kabir
Abstract
Although professional workflows leverage large language models widely, the interpretation for auditing unconstrained free-text generation is usually intractable if such generation demands legal, operational or financial workflow. We hereby demonstrate a text based system called SerenAI - inspired by world-models, it is a state transition system that outputs verifiable predictions rather than merely text: Provided with a description of the environment, state, and actions, the generated output contains 4 items: causal deltas that causally effect the given state, a next state that can logically follow from the given state and action, a validity reward, and a termination signal. For the released proto-model, we employ 2 steps of adaptation training, namely parameter efficient fine-tuning followed by verifier based RL over 50,000 exampled cause and effects in 12 environments spanning 10 reasoning domains. Compared to an initial internal evaluation of an 8B open-weight baseline, SerenAI increased JSON validity from 85.0% to 93.2%, schema validity from 55.0% to 84.0%, exact structured-output match from 0.0% to 41.5%, causal-delta exact match from 0.0% to 41.5%, resulting-state exact match from 0.0% to 42.0%, reward exact match from 1.0% to 80.5%, and termination exact match from 38.0% to 81.5%. These support the narrower claim that verifier-compatible adaptation can improve structured transition prediction. They do not yet establish legal-grade reliability. Accordingly, the paper also specifies a validation protocol for evidence-grounded legal workflows, calibration, human oversight, and sovereign on-premise deployment.
Chinese Translation
尽管专业工作流已广泛使用大语言模型,但当无约束的自由文本生成涉及法律、运营或财务工作流时,对其进行审计解读通常十分困难。本文提出一个名为SerenAI的基于文本的系统——受世界模型的启发,它是一个输出可验证预测而非仅仅文本的状态转移系统:在给定环境、状态和动作的描述后,生成的输出包含4项内容:因果影响给定状态的因果增量(causal deltas)、可在逻辑上由给定状态和动作推导出的下一状态、有效性奖励,以及终止信号。对于已发布的原型模型,我们采用两步适应性训练,即参数高效微调,随后在涵盖10个推理领域的12个环境中共50,000个因果示例上进行基于验证器的强化学习(RL)。与初始的8B开源权重基线内部评估相比,SerenAI将JSON有效性从85.0%提升至93.2%,模式(schema)有效性从55.0%提升至84.0%,结构化输出的精确匹配从0.0%提升至41.5%,因果增量精确匹配从0.0%提升至41.5%,结果状态精确匹配从0.0%提升至42.0%,奖励精确匹配从1.0%提升至80.5%,终止信号精确匹配从38.0%提升至81.5%。这些结果支持了一个较为保守的结论:与验证器兼容的适应性训练可以改进结构化状态转移预测,但尚不能确立法律级的可靠性。因此,本文还针对基于证据的法律工作流、校准、人类监督以及主权本地化部署提出了一套验证协议。
cs.AI / 65 / 2609.06654

A Computational Implementation of a Goal-Directed Theory of Affect

目标导向情感理论的计算实现
Hilpert, Bernhard, Szűcs, Tamás, Broekens, Joost, Moors, Agnes
Abstract
Computational modeling of emotion has long faced a tension between descriptive, "snapshot-based" appraisal models and granular, signal-driven architectures that often lack appropriate psychological grounding. This paper addresses this gap by presenting the first high-fidelity computational implementation of the Goal-Directed Theory (GDT) of affect. In this framework, affect is not a post-hoc label but a functional byproduct emerging from the continuous interplay between discrepancy detection and action selection within an agent's internal processing cycles. We evaluate the model through a series of principled simulations (Dice/Corridor tasks) designed to isolate affective signatures and dynamics during multi-step goal pursuit. Results demonstrate that complex affective profiles, like an anticipatory "lift" and a failure "crash", emerge naturally from simple interactions between goal-discrepancy and action-selection expectancies without requiring additional dedicated modules. By ensuring every computational component maps directly to components of the psychological theory, this work establishes a transparent, testable framework that enables a continuous "simulation-empiry" research loop. Our work contributes to moving the field beyond "black-box" heuristics toward a granular, mechanistic understanding of affect, integrated into the core of agent behavior.
Chinese Translation
情绪的计算建模长期面临两种范式之间的张力:一类是基于描述性“快照式”的评价模型,另一类是细粒度的、由信号驱动的架构,但后者往往缺乏恰当的心理学基础。本文通过呈现情感目标导向理论(Goal-Directed Theory, GDT)的首个高保真计算实现来弥补这一空白。在该框架中,情感并非事后附加的标签,而是在智能体内部处理循环中,由偏差检测与动作选择的持续交互所产生的功能性副产品。我们通过一系列精心设计的模拟实验(Dice/Corridor 任务)评估该模型,这些实验旨在分离多步目标追求过程中的情感特征与动态。结果表明,复杂的情感表现——如预期性的“提振”与失败后的“崩落”——能够自然地从目标偏差与动作选择期望之间的简单交互中涌现,而无需额外的专用模块。通过确保每个计算组件都直接对应心理学理论的组成部分,本研究建立了一个透明、可检验的框架,使“模拟—实证”的连续研究循环成为可能。我们的工作有助于推动该领域超越“黑箱”式启发方法,迈向对情感的细粒度、机制性理解,并将其整合到智能体行为的核心之中。
cs.AI / 66 / 2609.06715

We Built a Mirror and Mistook It for a Mind: Causal Liability and the Fallacy of AI Consciousness

我们造了一面镜子,却误将其当作心灵:因果责任与人工智能意识谬误
Khadangi, Afshin
Abstract
The contemporary debate over machine consciousness begins from a concealed assumption: that the object called "AI" already constitutes the kind of entity to which consciousness could belong. This paper challenges that assumption by separating phenomenal consciousness, introspective report, and human projective introspection, then arguing that generative systems can return linguistic traces of human interiority in first-person form without thereby identifying a phenomenal bearer. We call the resulting inference the AI Consciousness Fallacy. We then introduce Causal Liability Theory (CLT). CLT-I proposes liability closure as a criterion for individuating a candidate bearer: a physically continuing process becomes the non-delegable inheritor of constraints generated by its own endogenous discriminations. CLT-II advances the stronger conjecture that liability closure is necessary and sufficient for minimal phenomenal subjecthood. An open-weight causal audit operationalizes CLT-I across multiple model families. Forced discriminations produced persistent downstream divergence; activation patching showed strong causal mediation; live and copied adaptive states were behaviorally identical under matched randomness; and detached reconstruction preserved computational state across process replacement while, by protocol, breaking constitutive continuity and non-delegable inheritance. These results show that CLT-I distinctions are experimentally tractable and can dissociate causal bearer structure from first-person performance. The framework therefore separates consciousness attribution, causal bearer individuation, and the independent metaphysical question of consciousness constitution.
Chinese Translation
当前关于机器意识的争论始于一个被掩盖的假设:即被称为“AI”的对象已经构成了意识可能归属的那类实体。本文挑战了这一假设,将现象意识、内省报告与人类的投射性内省区分开来,并论证生成式系统能够以第一人称形式返回人类内在性的语言痕迹,但这并不意味着存在一个现象承载者。我们将由此产生的推论称为“AI意识谬误”(AI Consciousness Fallacy)。随后,我们提出因果责任理论(Causal Liability Theory, CLT)。CLT-I 提出以责任闭合(liability closure)作为候选承载者个体化的判据:一个在物理上持续的过程,成为由其自身内生辨别所生成的约束的不可委托继承者。CLT-II 则提出更强的猜想:责任闭合对于最小限度的现象主体性而言是必要且充分的。我们通过一项开放权重的因果审计在多个模型家族上对 CLT-I 进行操作化检验。强制辨别产生了持续的下游分歧;激活修补(activation patching)显示出强因果中介作用;在匹配随机性条件下,实时与复制的适应状态在行为上完全一致;而分离式重建在过程替换中保留了计算状态,但依照协议同时破坏了构成性连续性与不可委托继承性。这些结果表明,CLT-I 的区分在实验上是可处理的,并且能够将因果承载者结构与第一人称表现分离。因此,该框架将意识归属、因果承载者个体化,以及意识构成的独立形而上学问题区分开来。
cs.AI / 67 / 2609.06737

Monte Carlo-Based Ex-Ante Assessment of the Green Benefits of an AI-Driven Smart Agriculture Platform in Hainan

基于蒙特卡洛方法的海南AI驱动智慧农业平台绿色效益事前评估
Li, Zhaoyang, Zhang, Ruijie, Sun, Zhaoji, Zhang, Lu
Abstract
Smart agriculture platforms are widely regarded as key carriers for implementing China's pesticide and fertilizer reduction, water-saving and carbon-reduction agendas, yet a unified quantitative framework for assessing their green value is still lacking. Taking an AI-driven decision platform for tropical agriculture as the object (integrating large-language-model question answering, multimodal pest diagnosis, IoT sensing, satellite remote sensing, and a closed-loop field record system), this study builds a cradle-to-farm-gate agricultural carbon accounting model covering pesticide and fertilizer production, field N2O, irrigation electricity and paddy CH4, translates platform interventions into quantifiable transmission parameters, and propagates parameter uncertainty by Monte Carlo simulation over three Hainan scenarios (mango, winter vegetable, rice/nanfan, area-weighted 40%:30%:30%). Under full adoption, median reductions are 23.5% (90% interval 15.0%-33.2%) for pesticide use, 21.0% (13.8%-28.9%) for fertilizer, 16.5% (10.9%-23.5%) for irrigation water, and 21.5% (16.1%-27.2%) for carbon intensity. Attainment probabilities are high for fertilizer reduction >=15% (90.6%) and clear carbon decline (98.1%), but only about 20% for aggregate water saving >=20%, favoring scenario-specific statements. Sobol first-order indices show soil-test recommendation and organic substitution jointly explain about 83% of the variance of aggregate carbon-intensity reduction. Convergence tests show 10,000 iterations stabilize all statistics; conservative/baseline/optimistic scenario bounds are reported. The framework offers a reproducible, calibration-ready methodology for ex-ante green-value assessment and pilot observation design.
Chinese Translation
智慧农业平台被广泛视为落实中国农药化肥减量、节水与碳减排议程的关键载体,但评估其绿色价值的统一量化框架仍然缺失。本研究以一个面向热带农业的AI驱动决策平台为对象(集成大语言模型问答、多模态病虫害诊断、物联网传感、卫星遥感以及闭环田间记录系统),构建了覆盖农药化肥生产、农田N2O排放、灌溉用电及稻田CH4排放的“摇篮到农场大门”农业碳核算模型,将平台干预转化为可量化的传导参数,并通过蒙特卡洛模拟在海南三个情景(芒果、冬季瓜菜、水稻/南繁,面积加权比为40%:30%:30%)中传播参数不确定性。在完全采纳情景下,中位数减幅分别为:农药使用23.5%(90%区间15.0%-33.2%)、化肥21.0%(13.8%-28.9%)、灌溉用水16.5%(10.9%-23.5%)、碳排放强度21.5%(16.1%-27.2%)。化肥减量≥15%(90.6%)与碳排放明显下降(98.1%)的达标概率较高,而总体节水≥20%的概率仅约20%,因而更适合给出分情景表述。Sobol一阶指数表明,测土配方推荐与有机替代共同解释了总体碳强度减幅方差的约83%。收敛性检验显示,10,000次迭代即可使所有统计量稳定;同时报告了保守/基准/乐观情景的边界。该框架为绿色价值事前评估与试点观测设计提供了可复现、可直接校准的方法学。
cs.AI / 68 / 2609.06740

Simulating the Marginal Green Contribution of AI Modules in a Smart-Agriculture Platform: Evidence from Two Monte Carlo Experiments

智慧农业平台中AI模块边际绿色贡献的模拟研究:来自两项蒙特卡洛实验的证据
Li, Zhaoyang, Zhang, Ruijie, Sun, Zhaoji, Zhang, Lu
Abstract
Smart agriculture platforms usually bundle AI diagnosis, IoT sensing and decision push into a single package, so the green benefit attributable to each component remains unclear and resource-allocation decisions lack quantitative evidence. Building on a previous platform-level Monte Carlo assessment, this paper makes the components explicit and runs two controlled simulation experiments. Experiment 1 follows the chain from AI capability to farmer behavior to agrochemical input reduction, modeling pesticide/fertilizer reduction as avoidable blind-application share times prescription effectiveness times decision-touch coverage times adoption rate, and compares an experienced-extension mode with the AI mode: the probability of reaching 20% pesticide reduction is essentially zero in the extension mode but 20.7% at baseline, up to 49% with diagnosis accuracy 0.95 and adoption 0.85 under AI; the probability of 15% fertilizer reduction rises from near zero to 52.0%. Experiment 2 compares current practice (P0), IoT engineering retrofit (P1), and P1 plus AI irrigation scheduling (P2): median aggregate water saving rises from 7.8% (P0) to 11.0% (P1) and 16.0% (P2), with AI adding 5.0 percentage points beyond engineering; paddy CH4 reduction reaches 30.5% under AI scheduling versus 19.8% under manual operation, and the rice irrigation-methane subsystem carbon intensity declines 27.9%. Sensitivity analyses of both experiments consistently indicate that the primary bottleneck for meeting green targets is farmer adoption rather than algorithm accuracy, and that AI data fusion is robust to soil-moisture sensing errors. This work provides a reproducible simulation framework for component-level green-value evaluation and promotion-strategy optimization of smart agriculture platforms.
Chinese Translation
智慧农业平台通常将AI诊断、物联网感知与决策推送捆绑为一个整体包,因而各组件所带来的绿色效益仍不清晰,资源配置决策缺乏定量证据支持。本文在此前平台级蒙特卡洛评估的基础上,将各组件显性化,并开展了两项受控模拟实验。实验1沿"AI能力—农户行为—农化投入减量"的链条展开,将农药/化肥减量建模为可避免的盲目施用比例×处方有效性×决策触达覆盖率×采纳率,并比较了经验型推广模式与AI模式:在推广模式下,农药减量达到20%的概率基本为零,而AI模式下基准情景为20.7%,当诊断精度为0.95、采纳率为0.85时最高可达49%;化肥减量15%的概率从接近零提升至52.0%。实验2比较了现状(P0)、物联网工程改造(P1)以及P1叠加AI灌溉调度(P2)三种方案:总体节水中位数从7.8%(P0)上升至11.0%(P1)和16.0%(P2),其中AI在工程改造基础上额外贡献了5.0个百分点;在AI调度下稻田甲烷减排达30.5%,而人工操作仅为19.8%,水稻灌溉—甲烷子系统的碳强度下降27.9%。两项实验的敏感性分析一致表明,实现绿色目标的主要瓶颈是农户采纳率而非算法精度,且AI数据融合对土壤湿度传感误差具有稳健性。本研究为智慧农业平台组件级绿色价值评估与推广策略优化提供了一个可复现的模拟框架。
cs.AI / 69 / 2609.06746

Reason Through the Latent! Making Latent Visual Reasoning Necessary

在潜在空间中推理!让潜在视觉推理真正成为必需
Park, Suhyeong, Jung, Junha, Kang, Jaewoo
Abstract
Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply that the model actually relies on that state when producing its answer, especially when alternative image-conditioned paths remain available. We introduce \textbf{C}ausal \textbf{V}isual \textbf{R}ecurrent \textbf{R}easoning (CVRR), which preserves pretrained visual competence while making recurrent computation the required image-conditioned path to prediction. CVRR initializes recurrence from the question hidden state after the pretrained vision-language model has incorporated the image, then repeatedly updates this state while re-reading the same fixed visual evidence. Before decoding, visual states and the original multimodal KV cache are removed so that only the final recurrent state carries image-conditioned information to the answer. Across the $V^*$, MMVP, BLINK, and MME-RealWorld-Lite benchmarks, CVRR retains strong performance under this strict interface, while compatible latent reasoners fail to recover comparable visual competence even when retrained under the same constraint. Causal interventions further show that predictions remain sensitive to recurrent content when the question is held fixed, and that persistent visual evidence causally revises the recurrent trajectory. These results distinguish latent informativeness from latent computation that is actually used for prediction.
Chinese Translation
潜在视觉推理旨在通过隐藏状态计算而非显式的文本思维链来执行多模态推理。然而,视觉信息存在于潜在状态中并不意味着模型在生成答案时真正依赖该状态,尤其是当其他依赖图像的替代路径仍然可用时。我们提出了因果视觉循环推理(Causal Visual Recurrent Reasoning, CVRR),该方法在保持预训练视觉能力的同时,使循环计算成为预测所必需的依赖图像的路径。CVRR 在预训练视觉语言模型融合图像后,从问题的隐藏状态初始化循环状态,然后在反复读取同一固定视觉证据的同时不断更新该状态。在解码之前,移除视觉状态和原始多模态 KV 缓存,使得只有最终的循环状态携带依赖图像的信息用于生成答案。在 V*、MMVP、BLINK 和 MME-RealWorld-Lite 等基准测试中,CVRR 在这一严格接口下保持了强大的性能,而与之兼容的潜在推理器即使在相同约束下重新训练,也无法恢复相当的视觉能力。因果干预实验进一步表明,在问题保持不变的情况下,预测仍对循环状态内容敏感,且持续的视觉证据能够因果性地修正循环轨迹。这些结果区分了潜在信息的丰富性与真正用于预测的潜在计算。
cs.AI / 70 / 2609.06792

Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool Actions

通过自生成工具动作提升Android GUI智能体的熟练度与效率
Lee, Juyong, Jin, Woogyeol, Lee, Kimin
Abstract
Android agents using a hybrid action space that combines GUI actions and tool actions (e.g., accessing application data via APIs) remain largely underexplored, mainly due to the excessive effort required to create tools. To address this gap, we introduce DroidTool, a framework for augmenting the agents with self-generated tools, which are realized as Python functions operating on application states (e.g., a database). To create tools with minimal human labor, DroidTool employs an agentic workflow featuring stages: proposal, implementation, test generation and execution, and repair. Notably, when testing the created tools for verification, it constructs relational tests across relevant tools for natural preparation of appropriate test preconditions and improved test coverage, rather than testing each tool separately. The GUI agents augmented with the generated tools achieved approximately 4.47%p higher performance with approximately 20.05% fewer interactions than the GUI-only agents, averaged across representative benchmarks: AndroidWorld, B-MoCA, and MobileSafetyBench.
Chinese Translation
采用混合动作空间(即结合GUI动作与工具动作,例如通过API访问应用数据)的Android智能体在很大程度上仍待探索,主要原因在于创建工具所需的大量人力投入。为填补这一空白,我们提出了DroidTool,一个通过自生成工具来增强智能体能力的框架,这些工具以操作应用状态(如数据库)的Python函数形式实现。为以最少的人工劳动创建工具,DroidTool采用了一个智能体工作流,包含以下阶段:提案、实现、测试生成与执行以及修复。值得注意的是,在测试已创建的工具以进行验证时,DroidTool会跨相关工具构建关联测试,从而自然地准备适当的测试前置条件并提高测试覆盖率,而非对每个工具单独进行测试。在代表性基准(AndroidWorld、B-MoCA和MobileSafetyBench)上的平均结果显示,增强自生成工具的GUI智能体相比仅使用GUI的智能体,性能提升约4.47个百分点,交互次数减少约20.05%。
cs.AI / 71 / 2609.06816

Unsound Search with Policy and Value Networks in Legends of Code and Magic

在Legends of Code and Magic中结合策略与价值网络的非可靠搜索
Rubin, Dustin
Abstract
Decision-time search in perfect and imperfect information games with enumerable belief states are effective methods for game AI. Collectible card games are imperfect information games with large belief states. Legends of Code and Magic is a collectible card game competition where the belief states are $2^{101}$. The Legends of Code and Magic (LoCM) champion, ByteRL, plays with no search. Other works claim sound enumeration-based search is unusable in the genre due to the number of belief states. We measured three previously defined properties that predict where theoretically unsound perfect information Monte Carlo's defects are cheap and found LoCM sits in the favorable region. Starting with imitation learning of the runner-up policy, NeteaseOPD, we created a policy and value feed-forward network. Our agent searches over worlds sampled from a prior over the opponent's deck built from the runner-up's drafts. Using our strictest configuration in the battle phase we beat ByteRL with a win percentage of 51.35% 95% CI [50.37, 52.33], over 10,000 pre-registered games using the LoCM official referee and time limit. Search is not a minor factor on the matchup between our agent and ByteRL. Without search this agent scores 26.8% and adding search adds +24.6 points. Unsound search in imperfect information games could be exploitable. We replicate a published best-response attack against ByteRL. We then apply the same attack protocol to two search configurations of our agent, and each one resists it better than ByteRL at every iteration. In LoCM unsound search gives us a stronger and more resilient agent.
Chinese Translation
在完美信息博弈和具有可枚举信念状态的博弈中,决策时搜索是构建游戏人工智能的有效方法。集换式卡牌游戏属于信念状态庞大的不完美信息博弈。Legends of Code and Magic(LoCM)是一项集换式卡牌游戏竞赛,其信念状态数量高达 $2^{101}$。LoCM 的冠军 ByteRL 在不使用搜索的情况下进行博弈。其他研究声称,由于信念状态数量庞大,基于枚举的可靠搜索在该类游戏中无法使用。我们测量了三个先前定义的属性,用于预测理论上不可靠的完美信息蒙特卡洛搜索的缺陷在何处代价较低,发现 LoCM 处于有利区域。从对亚军策略 NeteaseOPD 的模仿学习出发,我们构建了一个策略与价值前馈神经网络。我们的智能体在从基于亚军出牌构建的对手卡组先验分布中采样得到的世界上进行搜索。使用我们最严格的配置,在战斗阶段,我们在 10,000 场预先注册的比赛中,使用 LoCM 官方裁判和时间限制,以 51.35% 的胜率(95% 置信区间 [50.37, 52.33])击败了 ByteRL。搜索在我们的智能体与 ByteRL 的对战中并非次要因素:不使用搜索时,该智能体的胜率为 26.8%,而加入搜索带来了 +24.6 个百分点的提升。不完美信息博弈中的不可靠搜索可能会被利用。我们复现了一种已发表的对 ByteRL 的最优应对攻击。随后,我们将相同的攻击协议应用于我们智能体的两种搜索配置,在每一轮迭代中,二者都比 ByteRL 更能抵御该攻击。在 LoCM 中,不可靠搜索使我们获得了更强且更具韧性的智能体。
cs.AI / 72 / 2609.06826

Formation of structural attractors in neuromorphic systems

神经形态系统中结构性吸引子的形成
Parzhyn, Yurii, Schwarzmann, Alexander, Lapin, Mykyta, Bokhan, Kostiantyn
Abstract
This paper examines the theory of Invariant Structural Learning (ISL), which proposes a non-optimization approach to concept formation. Learning is interpreted as convergence to structural attractors in a hypergraph space, rather than as the minimization of a global loss function. The paper presents the ISL model, including its mathematical formalization, computational verification, and a hypothetical neurobiological interpretation. The mathematical section introduces the formal apparatus of the structural reduction process and proves its finite convergence, the existence and uniqueness of class structural attractors, and the self-organization of attractor maps. The computational section demonstrates the feasibility of the proposed approach on classical image recognition tasks, utilizing the proposed learning mechanism without backpropagation and with extremely small training datasets. Finally, the neurobiological section formulates hypotheses regarding the possible implementation of structural attractors in dendritic trees, neural coding as a projection of internal attractor dynamics, and the development of neural architectures supporting the proposed learning concept. These hypotheses are discussed in the context of modern experimental data in the fields of dendritic computations, synaptic plasticity, and the structural organization of neural circuits. The proposed neurobiological mechanisms are presented as testable hypotheses rather than established biological facts. The results demonstrate the mathematical consistency and computational feasibility of the proposed model, while the neurobiological hypotheses outline potential directions for its experimental verification.
Chinese Translation
本文研究了不变结构学习(Invariant Structural Learning, ISL)理论,该理论提出了一种非优化的概念形成方法。学习被解释为在超图空间中向结构性吸引子的收敛,而非全局损失函数的最小化。本文介绍了ISL模型,包括其数学形式化、计算验证以及假设性的神经生物学解释。数学部分引入了结构约简过程的形式化工具,并证明了其有限收敛性、类别结构性吸引子的存在性与唯一性,以及吸引子映射的自组织性。计算部分在经典图像识别任务上展示了该方法的可行性,利用所提出的学习机制,无需反向传播,且训练数据集规模极小。最后,神经生物学部分提出了若干假设,涉及结构性吸引子在树突树中的可能实现、作为内部吸引子动力学投影的神经编码,以及支持所提学习概念的神经架构的发育。这些假设在树突计算、突触可塑性和神经回路结构组织等领域的现代实验数据背景下进行了讨论。所提出的神经生物学机制被表述为可检验的假设,而非已确立的生物学事实。研究结果表明了所提模型的数学一致性与计算可行性,而神经生物学假设则为其实验验证概述了潜在方向。
cs.AI / 73 / 2609.06831

NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures

NormViz:一个将多模态推理扎根于全球文化的基准与框架
Yerukola, Akhila, Harel-Canada, Fabrice Y, Khanuja, Simran, Rao, Abhinav Sukumar, Suvarna, Ashima, Peng, Nanyun, Gabriel, Saadia, Sap, Maarten
Abstract
AI systems are used worldwide, but they struggle to serve the needs of culturally diverse populations. Prior work on cultural understanding evaluates AI systems on text-only settings or on visual artifact recognition (e.g. foods, clothing). The ability to reason about visually observable behaviors through local social norms, which we call visual norm understanding, remains unexamined. We introduce NormViz-Bench, a high quality, human-validated benchmark of 3,268 contrastive image pairs (6,536 images) spanning 16 countries. Each pair varies only in the culturally relevant behavior (e.g., objects, attributes, spatial relations, and actions) that alters how each image is interpreted. Each image is labeled as conforming to, violating, or irrelevant to local social norms, and pair-level evaluation requires both images to be correctly classified, thereby preventing reliance on superficial visual shortcuts. Even the strongest VLMs, Gemini 3.0 Flash and Qwen2.5 VL 7B, succeed on only 26.6% and 21.6% of pairs, struggling most with identifying violating and culturally benign visual behaviors. Towards bridging this, we introduce NormViz-Train, a training dataset of 64k images paired with explanations. Though absolute performance remains low (<30%), finetuning on NormViz-Train improves pair accuracy relatively by up to 125% and 36% Qwen3-VL 4B and 8B respectively, showing a path forward to teach models to connect visual perception to cultural significance. Together, NormViz-Bench and NormViz-Train establish visual norm understanding as a challenging and consequential frontier for multimodal AI.
Chinese Translation
人工智能系统在全球范围内被广泛使用,但它们难以满足文化多样化人群的需求。以往关于文化理解的研究仅在纯文本设置或视觉制品识别(如食物、服装)上评估AI系统。通过当地社会规范推理视觉上可观察到的行为的能力——我们称之为视觉规范理解——尚未得到研究。我们提出了NormViz-Bench,一个经过人工验证的高质量基准,包含涵盖16个国家的3,268对对比图像(共6,536张图像)。每一对图像仅在文化相关的行为(如物体、属性、空间关系和动作)上存在差异,而这些差异会改变图像的解读方式。每张图像被标注为符合、违反或与当地社会规范无关,并且图像对级别的评估要求两张图像均被正确分类,从而防止模型依赖表面的视觉捷径。即使是最强大的视觉语言模型(VLM),Gemini 3.0 Flash和Qwen2.5 VL 7B,也仅在26.6%和21.6%的图像对上取得成功,在识别违反规范和文化上无害的视觉行为方面尤为困难。为弥合这一差距,我们提出了NormViz-Train,一个包含6.4万张图像及配套解释的训练数据集。尽管绝对性能仍然较低(<30%),但在NormViz-Train上微调使Qwen3-VL 4B和8B的图像对准确率相对提升了125%和36%,展示了一条教导模型将视觉感知与文化意义联系起来的可行路径。NormViz-Bench和NormViz-Train共同将视觉规范理解确立为多模态AI领域一个具有挑战性且意义重大的前沿方向。
cs.AI / 74 / 2609.06849

Learning transferable human physiology from two million hours of sleep with SleepFM-2

基于两百万小时睡眠数据学习可迁移的人体生理特征:SleepFM-2
Thapa, Rahul, Sun, Christopher, Lehn-Schioler, William Theodor, Kivelson, Sophia Claire, Hanif, Umaer, Moore IV, Hyatt, Zhang, Harrison G., Ahmed, Hafsa, Dige, Marcus, Lorenzen, Niels R., Heremans, Elisabeth Roxane M., Specht, Adrien, Gimenez, Ulysse, Guillard, Robin, Brink-Kjaer, Andreas, Zou, James, Mignot, Emmanuel
Abstract
Sleep provides a nightly window into health by capturing coordinated activity across the brain, heart, muscles and respiratory system. We introduce SleepFM-2, a sleep foundation model developed and evaluated on 282,511 polysomnography recordings from 26 cohorts, including 235,865 used for pretraining. These data span more than two million hours of multimodal physiology. Compared with SleepFM, SleepFM-2 improves disease prediction and sleep scoring, supports arousal, limb movement and respiratory event detection, and transfers to wearable sensing and subjective sleep phenotypes. A model combining its PSG representation with age, sex and BMI met a prespecified discrimination and significance criterion for 215 subsequently recorded EHR phenotypes in two held-out cohorts, including one health system unseen during pretraining. For 155 phenotypes, the PSG representation added reproducible information beyond demographics. SleepFM-2 also outperformed a 480-feature baseline derived from the same recordings. Its disease scores revealed a reproducible principal component associated with reduced sigma-band spatial coupling and increased hypnodensity entropy. The frozen encoder performed within the observed range of expert scorers for sleep events and transferred to wakeful EEG, headband and in-ear EEG, wrist PPG and wrist accelerometry. It improved sleep staging across six accelerometry cohorts and achieved disease-prediction performance in UK Biobank similar to models pretrained directly on accelerometry. Finally, SleepFM-2 captured aspects of subjective sleep not recovered by conventional PSG summaries, particularly reports of the recorded night. These results show that multimodal sleep physiology can provide a transferable representation of human health across diseases, clinical tasks, sensors and subjective experience.
Chinese Translation
睡眠通过捕捉大脑、心脏、肌肉和呼吸系统的协同活动,为健康提供了一个每夜可观测的窗口。我们提出了SleepFM-2,一个在来自26个队列的282,511条多导睡眠图(polysomnography, PSG)记录上开发和评估的睡眠基础模型,其中235,865条用于预训练。这些数据涵盖超过两百万小时的多模态生理信号。与SleepFM相比,SleepFM-2提升了疾病预测和睡眠分期性能,支持觉醒、肢体运动和呼吸事件检测,并能迁移到可穿戴传感和主观睡眠表型。在两个保留队列(其中之一为预训练中未见过的卫生系统)中,将其PSG表征与年龄、性别和BMI相结合的模型,对215种后续记录的电子健康记录(EHR)表型满足了预设的判别力和显著性标准。对于155种表型,PSG表征在人口统计学信息之外提供了可重现的额外信息。SleepFM-2的表现也优于基于相同记录提取的480个特征构建的基线模型。其疾病评分揭示了一个可重现的主成分,该主成分与sigma频段空间耦合降低及睡眠结构密度(hypnodensity)熵增加相关。冻结编码器在睡眠事件检测上的表现处于专家评分员的观测范围内,并可迁移至清醒脑电(EEG)、头带式和入耳式EEG、腕部光电容积描记(PPG)以及腕部加速度计数据。它提升了六个加速度计队列的睡眠分期性能,并在英国生物样本库(UK Biobank)中取得了与直接在加速度计数据上预训练的模型相当的疾病预测性能。最后,SleepFM-2捕捉到了传统PSG摘要无法还原的主观睡眠体验的某些方面,尤其是对被记录当晚睡眠的主观报告。这些结果表明,多模态睡眠生理信号能够提供一种可跨疾病、临床任务、传感器和主观体验迁移的人体健康表征。
cs.AI / 75 / 2609.06914

A visual large language foundational model for medical image recognition using clinician-oriented social media

基于面向临床医生的社交媒体的医学图像识别视觉大语言基础模型
Hou, Lingxuan, Xie, Yuhua, Hu, Yue, Zhuang, Yan, Li, Junqi, Xia, Chengzhi, Nguyen, Binh Phu, Siddique, Abubakar, Nguyen, Minh, Hou, Yao, Bao, Yanju, Liu, Kexin, Chen, Ke, Sun, Jianjun, Li, Zeqi, Nguyen, Trung, Lin, Jiangli
Abstract
Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignment. Here, we leverage de-identified medical images and expert commentaries shared on clinician-oriented social media. By combining an advanced LLM with clinician-in-the-loop verification, we established a rigorous pipeline to construct ThoughtMed-1M, a long-form medical VQA dataset containing over one million VQA pairs and designed to capture structured clinical logic and medical image-text alignment. To demonstrate its utility, we developed a FOundational LLM Trained on ThoughtMed-1M (FOLTMed). FOLTMed achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4%, and generated more clinically coherent responses on the ThoughtMed-1M test set. It outperformed state-of-the-art models by 3--5% across factuality and similarity metrics, highlighting a scalable paradigm for advancing research on clinically grounded multimodal LLMs.
Chinese Translation
大语言模型(LLM)已在众多领域展现出强大的能力,并在医学领域显示出巨大潜力。然而,由于缺乏能够捕捉临床推理和显式图文对齐的视觉问答(VQA)数据集,其在医疗场景中的应用仍然受限。本研究利用面向临床医生的社交媒体上共享的去标识化医学图像及专家评论,结合先进的大语言模型与临床医生在环验证,建立了一套严谨的流程,构建了ThoughtMed-1M——一个包含超过一百万条VQA对的长文本医学视觉问答数据集,旨在捕捉结构化的临床逻辑与医学图文对齐。为展示其价值,我们开发了基于ThoughtMed-1M训练的基础大语言模型(FOLTMed)。FOLTMed在42个医学VQA基准数据集上取得了最先进的性能,宏观准确率达85.4%,并在ThoughtMed-1M测试集上生成了更具临床连贯性的回答。其在事实性和相似性指标上超越最先进模型3%–5%,凸显了一种可扩展的范式,有望推动临床导向的多模态大语言模型研究的发展。
cs.AI / 76 / 2609.06941

When and Why LLM Causal Priors Help: Closed-Loop Prior Selection for Amortized Causal Inference

大语言模型因果先验何时以及为何有效:用于摊销因果推断的闭环先验选择
Zhou, Haohao
Abstract
Causal effect estimation asks how an outcome would change under an intervention, and medicine, economics, and public policy all treat it as a foundational task. Prior-data fitted networks (PFNs) amortize the task: a model trained on large numbers of programmatically generated synthetic causal tasks reads a new problem's observational data into context and returns an interventional-effect estimate in a single forward pass. The capability of such models is largely determined by the synthetic training prior, which is currently designed by hand, a bottleneck acknowledged by both Do-PFN and CausalPFN. Large language models (LLMs) can now ``draw'' plausible causal graphs for a given domain, suggesting that LLM-distilled graphs could serve as prior material. Whether injecting such graphs helps at all, where any gain comes from, and when injection helps. Practice has so far relied on manual trial and error. We propose a \emph{closed-loop prior selection framework} that casts prior injection as a budget-constrained optimization over a candidate prior pool. Candidates undergo cheap post-training and are scored by a composite metric dominated by real-domain generalization; the winner then receives full training and paired statistical validation. On a 7.34M-parameter Do-PFN, the framework's winner attains a formally significant $2.75\times$ gain on the primary evaluation domain, and its error falls below that of the uninjected official base. Generalization on an adjacent monitoring domain improves significantly, and no monitored capability degrades. Mechanism experiments show that the gain depends on the semantic content of the distilled graph rather than its structural diversity alone does not produce it (directional evidence). With this framework and this regularity in hand, the use of LLM causal priors stops being manual trial and error and becomes an empirically verifiable selection problem.
Chinese Translation
因果效应估计关注的是结果在干预下将如何变化,医学、经济学和公共政策均将其视为一项基础性任务。先验数据拟合网络(PFNs)将这一任务摊销化:在大量以程序化方式生成的合成因果任务上训练的模型,将新问题的观测数据读入上下文,并在单次前向传播中返回干预效应估计。此类模型的能力在很大程度上取决于合成训练先验,而目前先验是人工设计的,这已被 Do-PFN 和 CausalPFN 认可为瓶颈。大语言模型(LLMs)现在能够为给定领域“绘制”合理的因果图,这表明 LLM 蒸馏得到的图可以作为先验素材。然而,注入此类图是否有帮助、增益来自何处、以及注入何时有效,目前的实践仍依赖人工试错。我们提出一个闭环先验选择框架,将先验注入建模为在候选先验池上的预算受限优化问题。候选先验先经过低成本的后期训练,并由以真实领域泛化为主导的复合指标评分;胜出者随后接受完整训练并进行成对统计验证。在一个 7.34M 参数的 Do-PFN 上,该框架的胜出者在主要评估领域上取得了具有形式化显著性的 2.75 倍增益,其误差低于未注入的官方基础模型。在相邻监测领域上的泛化能力显著提升,且所有受监测能力均无退化。机制实验表明,该增益依赖于蒸馏图的语义内容,而仅靠结构多样性无法产生增益(方向性证据)。借助该框架和这一规律,LLM 因果先验的使用不再是人工试错,而成为一个可实证验证的选择问题。
cs.AI / 77 / 2609.06960

iBrain: A Unified Foundation Model Reading the Brain from Surface to Spikes

iBrain:一个从脑表信号到神经尖峰活动统一读取大脑的基础模型
Chen, Ying, Wang, Tiou, Yue, Zhifeng
Abstract
Invasive neural recordings provide high-fidelity measurements of brain activity, with signals such as intracranial EEG (iEEG) and intracortical spiking activity capturing neural dynamics at different spatial and temporal scales. Yet existing neural foundation models have largely been developed independently for different invasive recording paradigms, leaving joint pretraining across heterogeneous invasive signals underexplored. In this work, we introduce iBrain, a unified foundation model that jointly learns from iEEG and spiking activity. iBrain employs signal-specific encoders to accommodate their distinct signal characteristics and a shared spatiotemporal Transformer backbone to model dependencies across recording channels and time. We pretrain iBrain on over 7,000 hours of heterogeneous neural recordings using masked signal reconstruction and channel-view alignment, promoting contextual modeling of neural dynamics and robustness across different channels. iBrain consistently outperforms single-signal pretraining baselines and achieves state-of-the-art performance on multiple benchmarks. Further experiments demonstrate that iBrain exhibits transferability and data efficiency across diverse recording settings. These results highlight the potential of joint pretraining on heterogeneous invasive neural recordings to support scalable neural modeling and transferable representations across recording settings and downstream tasks.
Chinese Translation
侵入式神经记录能够提供高保真度的脑活动测量,其中颅内脑电图(iEEG)和皮层内尖峰活动等信号可在不同空间和时间尺度上捕捉神经动力学。然而,现有的神经基础模型大多针对不同的侵入式记录范式独立开发,跨异构侵入式信号的联合预训练仍未得到充分探索。在本工作中,我们提出了 iBrain,一个从 iEEG 和尖峰活动中联合学习的统一基础模型。iBrain 采用信号特定的编码器以适应两种信号的不同特性,并使用共享的时空 Transformer 骨干网络来建模跨记录通道和时间的依赖关系。我们利用掩码信号重建和通道-视图对齐的方法,在超过 7,000 小时的异构神经记录上对 iBrain 进行预训练,从而促进对神经动力学的上下文建模以及跨不同通道的鲁棒性。iBrain 持续优于单信号预训练基线,并在多个基准测试中取得了最先进的性能。进一步的实验表明,iBrain 在多种记录场景下展现出良好的可迁移性和数据效率。这些结果凸显了在异构侵入式神经记录上进行联合预训练的潜力,可支持可扩展的神经建模以及跨记录场景和下游任务的可迁移表示。
cs.AI / 78 / 2609.06961

SSP-DMGTimeNet: Physics-Constrained Learning for Spatiotemporal Trajectory Prediction of Vehicle Platoons

SSP-DMGTimeNet:面向车辆队列时空轨迹预测的物理约束学习方法
Wang, Yuhang, Ma, Kailang, Li, Zirui, Fan, Mingfeng, Jang, Kitae, Lee, Changju, Huang, Heye
Abstract
Existing car-following prediction methods mainly optimize trajectory accuracy, while rarely considering whether predicted disturbances propagate realistically along a vehicle platoon. This limitation may lead to accurate but string-unstable predictions. We propose SSP-DMGTimeNet, a physics-constrained learning framework for spatiotemporal trajectory prediction of vehicle platoons. The model combines multi-scale temporal representations with cross-vehicle interaction features to capture complex and time-varying platoon dynamics. A propagation-delay-aware causal attention mechanism explicitly models upstream-to-downstream disturbance propagation by learning response delays between adjacent vehicles and accumulating them along the platoon. In addition, time- and frequency-domain string-stability losses relieve disturbance amplification across both adjacent vehicles and arbitrary sub-platoons during training. Experiments on HighD show that SSP-DMGTimeNet achieves an unstable-window rate of 0.65\% for five-vehicle platoons and a maximum head-to-tail amplification of 0.898 on the ground-truth excitation subset, while maintaining competitive trajectory prediction performance. In zero-shot evaluation on NGSIM US-101 and I-80, the model achieves velocity MAEs of 1.316~m/s and 1.252~m/s, with unstable-window rates of 3.90\% and 4.10\%, respectively. These results demonstrate that incorporating platoon-level physical constraints can effectively balance trajectory prediction accuracy and disturbance propagation stability.
Chinese Translation
现有的跟车预测方法主要优化轨迹精度,而很少考虑预测扰动是否能够真实地沿车辆队列传播。这一局限可能导致预测结果虽然准确,但在队列稳定性意义上却不稳定。我们提出了SSP-DMGTimeNet,一种面向车辆队列时空轨迹预测的物理约束学习框架。该模型将多尺度时间表示与跨车辆交互特征相结合,以捕捉复杂且时变的队列动力学特性。一种传播延迟感知的因果注意力机制通过学习相邻车辆之间的响应延迟并沿队列进行累积,显式地建模了从上游到下游的扰动传播。此外,时域和频域的队列稳定性损失在训练过程中抑制了扰动在相邻车辆以及任意子队列间的放大。在HighD数据集上的实验表明,SSP-DMGTimeNet对五车队列的不稳定窗口率仅为0.65%,在真实激励子集上的头尾最大放大系数为0.898,同时保持了具有竞争力的轨迹预测性能。在NGSIM US-101和I-80数据集上的零样本评估中,该模型的速度MAE分别为1.316 m/s和1.252 m/s,不稳定窗口率分别为3.90%和4.10%。这些结果表明,引入队列级物理约束能够有效平衡轨迹预测精度与扰动传播稳定性。
cs.AI / 79 / 2609.07008

RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving

RedKnot-MLA:面向 DeepSeek-V4 长上下文服务的多头离线-在线复用机制
Liu, Yang, Luo, Zhaokai, Jin, Huayi, He, Ruozhou, Hong, Chenchen, Ma, Mingxiao, Zhang, Biao, Wang, Zhiyong, Wang, Boyu, Chen, Guanjie, Liu, Yifei, Xie, Tao, Hu, Junhao
Abstract
Multi-head latent attention (MLA) exposes many logical query heads through one packed latent KV stream. This representation is memory efficient, but it removes the physical per-head cache boundary assumed by conventional head-wise reuse. We present our system, a DeepSeek-V4 realization of RedKnot's head-aware reuse principle. Each immutable document is processed offline at canonical position zero; certified Local-head contributions are retained as MLA-Off. At serving time, query-side RoPE relocation restores the document's request position, a small Global-head set and protected Local token rows are recomputed as MLA-Online, and the two paths are merged before a single shared output projection. The packed MLA latent is never split. DeepSeek-V4-Flash uses 37 reusable layers and a 56/8 Local/Global partition, giving a 75.29% analytic logical head-row ceiling; the Pro-0813 profile uses 55 layers and 112/16 heads, giving 78.89%. Frozen Flash operating points show hot-artifact TTFT speedups of 2.02-3.84x. At 256K, the archived three-dataset study reports an aggregate F1 change of +3.24 percentage points, an EM change of +4.16 points, and a 78.7-79.5% analytic major-operator arithmetic saving, while one dataset decreases by 2.81 F1 points. A separate author-reported 256K hot-artifact QPS measurement is approximately 2.0x; because its raw concurrency trace is not included in this bundle, we mark it as preliminary rather than archived evidence. We describe the factorization, position repair, token-row closure, sparse-MoE support, TP8 integration, and the measurement boundaries needed to interpret these results.
Chinese Translation
多头潜在注意力(MLA)通过一条打包的潜在 KV 流暴露出众多逻辑查询头。这种表示方式在内存上非常高效,但它消除了传统的按头复用机制所假设的物理上逐头缓存边界。我们提出了一套系统,实现了 RedKnot 头感知复用原则在 DeepSeek-V4 上的落地。每个不可变文档在规范位置零处离线处理,经认证的 Local-head 贡献被保留为 MLA-Off。在服务阶段,查询侧 RoPE 重定位将文档恢复到其请求位置,一小部分 Global-head 集合及受保护的 Local token 行被重新计算为 MLA-Online,两条路径在单一的共享输出投影之前合并。打包的 MLA 潜在表示从不被拆分。DeepSeek-V4-Flash 使用 37 个可复用层和 56/8 的 Local/Global 划分,给出 75.29% 的解析逻辑头行上限;Pro-0813 配置使用 55 层和 112/16 个头,给出 78.89%。冻结的 Flash 运行点显示热点工件 TTFT 加速达 2.02–3.84 倍。在 256K 上下文下,存档的三数据集研究报告了 +3.24 个百分点的总体 F1 变化、+4.16 个点的 EM 变化,以及 78.7–79.5% 的解析主要算子算术节省,但其中一个数据集下降了 2.81 个 F1 点。另一项由作者报告的 256K 热点工件 QPS 测量约为 2.0 倍;由于其原始并发追踪未包含在本资料包中,我们将其标记为初步证据而非存档证据。我们描述了分解方法、位置修复、token 行闭合、稀疏 MoE 支持、TP8 集成,以及解释这些结果所需的测量边界。
cs.AI / 80 / 2609.07050

Beyond One-Shot Expansion: Contrastive Evidence Exploration for Multi-Hop Retrieval

超越一次性扩展:面向多跳检索的对比式证据探索
Yun, JungMin, Kim, YoungBin
Abstract
Retrieval-augmented generation (RAG) critically depends on retrieving the evidence necessary for effective reasoning. However, this remains particularly challenging in multi-hop question answering (QA), where supporting passages are often linked through intermediate entities and relations that must be progressively uncovered. Existing retrieval approaches typically rely on a single retrieval intent or one-shot query expansion, limiting their ability to adapt to newly retrieved evidence and potentially introducing noisy or redundant retrieval signals. To address these limitations, we propose a training-free multi-hop retrieval framework that integrates evidence-conditioned exploration, passage-specific contrastive refinement, and coverage-aware final ranking. During offline indexing, the framework constructs passage-specific contrastive facets that characterize each passage relative to its semantically similar neighbors, providing fine-grained signals to distinguish closely related candidates. At inference time, the framework iteratively retrieves evidence, generates probes targeting unresolved information needs, refines candidate relevance using the contrastive facets, and selects a complementary set of passages that collectively cover diverse evidence-seeking intents. Experiments on MuSiQue, HotpotQA, and 2WikiMultihopQA demonstrate consistent improvements in retrieval quality and downstream QA performance over baselines.
Chinese Translation
检索增强生成(RAG)高度依赖于检索到有效推理所需的证据。然而,这在多跳问答(QA)中仍然极具挑战性,因为支持性段落往往通过需要逐步揭示的中间实体和关系相互关联。现有的检索方法通常依赖于单一检索意图或一次性查询扩展,这限制了其适应新检索证据的能力,并可能引入嘈杂或冗余的检索信号。为解决这些局限,我们提出了一个无需训练的多跳检索框架,该框架集成了基于证据条件的探索、段落特定的对比式精炼以及覆盖感知的最终排序。在离线索引阶段,该框架构建段落特定的对比切面(contrastive facets),用于刻画每个段落相对于其语义相似邻居的特征,从而提供细粒度信号以区分密切相关的候选段落。在推理阶段,该框架迭代地检索证据,生成针对未解决信息需求的探针(probe),利用对比切面精炼候选相关性,并选择一组互补的段落集合,共同覆盖多样化的证据寻求意图。在MuSiQue、HotpotQA和2WikiMultihopQA上的实验表明,该框架在检索质量和下游问答性能方面均持续优于基线方法。
cs.AI / 81 / 2609.07065

VST: Verifiable Structured Transport for Auditable Agent-to-Agent Alpha Discovery

VST:面向可审计智能体间Alpha发现的可验证结构化传输
Li, Yuqi, Liu, Siyuan, Liu, Bingjun
Abstract
Agent-to-agent (A2A) alpha discovery is slowed by repeated feedback cycles between mining and evaluation agents, whose hand-offs, in contemporary LLM multi-agent systems, are free-form natural-language messages that carry no stable contract and cannot be replayed. We first restructure this communication as a structured agent-to-agent protocol of \emph{typed, causally addressable, unicast records}, so that the committed stream forms a causal trajectory. On that trajectory a single predictor with four typed heads forecasts the accumulated guidance the two miners would receive several cycles ahead; a transactional verify--leap controller then commits a multi-cycle speculative outcome only when it passes a four-level gate, and otherwise rolls back to the exact prior state. Structure is the enabling contribution, and its value is not accuracy. A controlled ablation shows an equal-information free-text channel reaches the same predictor hit rate. What typing provides is a state that can be schema-checked, replayed deterministically, and prevented by construction from leaking a forecast to an evaluator: auditability by construction, not an empirically stress-tested guarantee. On a CSI~1000 out-of-sample holdout, our single run is the only one among eight methods (seven baselines and ours) to hold a positive median annualized return and Sharpe at the factor level, though the median return \emph{in excess} of the benchmark stays negative for every method including ours; its development-selected top-20 portfolios reach a $0.71$ median holdout Sharpe, selected on a split inside the optimization horizon. We report these single-run results descriptively, gross of costs, and are explicit about their limits throughout; in particular we do not isolate the effect of the leap machinery from the inherited search substrate, which we leave to future work.
Chinese Translation
智能体间(A2A)的alpha发现过程因挖掘智能体与评估智能体之间反复的反馈循环而进展缓慢。在当前的LLM多智能体系统中,这些交互以自由格式的自然语言消息进行,既缺乏稳定的契约,也无法回放。我们首先将这种通信重构为一种结构化的智能体间协议,由\emph{类型化、因果可寻址的单播记录}组成,使已提交的记录流构成一条因果轨迹。在该轨迹上,一个带有四个类型化输出头的单一预测器提前预测两个挖掘智能体在未来数个循环中将获得的累计引导信号;随后,一个事务化的“验证—跳跃”控制器仅在多循环投机结果通过四级门控时才将其提交,否则回滚到先前的精确状态。结构化是本工作的使能性贡献,其价值不在于预测精度。受控消融实验表明,等信息的自由文本通道可以达到相同的预测器命中率。类型化所提供的是一种可进行模式检查、可确定性回放、并在构造上防止预测信息泄露给评估器的状态:即构造层面的可审计性,而非经过实证压力测试的保证。在CSI 1000(中证1000)样本外保留集上,我们的单次运行是八种方法(七个基线与我们的方法)中唯一在因子层面保持正中位数年化收益和正夏普比率的方法,尽管包括我们在内所有方法的相对基准的\emph{超额}中位数收益均为负值;其在开发集上选出的前20只股票组合达到0.71的中位数保留集夏普比率,且该选择基于优化时间窗口内的数据划分。我们以描述性方式报告这些单次运行的结果(未扣除交易成本),并在全文中明确其局限性;特别是,我们未将跳跃机制的效应与所继承的搜索底层分离开来,这一点留待未来工作。
cs.AI / 82 / 2609.07075

A Hierarchical Consistency Framework for Auditing Retrieval-Augmented Generation Systems

一个用于审计检索增强生成系统的层次化一致性框架
Gonzalez, Ramon, Diaz, Antonio
Abstract
Retrieval-augmented generation (RAG) is commonly evaluated by whether the final answer is correct. That test is insufficient: an answer can match its reference while the context that produced it contains a direct contradiction, leaving the contested evidence invisible to answer-only review and retrieval relevance scores. This paper presents the Hierarchical Consistency Framework (HCF), a post-hoc, model-agnostic audit of three distinct levels of a RAG process: the knowledge corpus, the final retrieved context, and the generated answer. HCF represents corpus conflicts as source-linked atomic facts, thereby identifying the documents responsible, and returns each Answer Consistency Score (ACS) with an explanation of supporting and contradictory contextual statements. We evaluate HCF on several controlled corpora spanning five domains and 100 query-corpus instances. A human evaluator compares every generated response with its supplied ground-truth response. The results show that the three diagnostic levels can dissociate: the corpus with the highest mean retrieval similarity has the lowest mean ACS, while a structurally degraded corpus performs worse at corpus level but better at answer level. Most importantly, HCF identifies contradictory retrieved evidence in several cases where the answer still matches the ground truth. HCF does not certify factual truth; it makes the evidence supporting and challenging an answer inspectable and attributable.
Chinese Translation
检索增强生成(RAG)通常通过最终答案是否正确来评估。然而这种测试并不充分:答案可能与参考答案相符,但生成该答案的上下文中却存在直接矛盾,而仅基于答案的审查和检索相关性评分无法发现这些存在争议的证据。本文提出了层次化一致性框架(Hierarchical Consistency Framework, HCF),这是一种事后(post-hoc)、模型无关的审计方法,针对RAG流程中三个不同的层次进行审计:知识语料库、最终检索到的上下文以及生成的答案。HCF将语料库中的冲突表示为可溯源至原始文献的原子事实,从而识别出应负责的文档,并为每个答案一致性评分(Answer Consistency Score, ACS)提供支持性和矛盾性上下文陈述的解释。我们在覆盖五个领域、100个查询-语料库实例的多个受控语料库上评估了HCF。由人工评估者将每个生成的回复与其给定的标准回复进行比较。结果表明,这三个诊断层次可能相互分离:平均检索相似度最高的语料库其平均ACS反而最低,而一个结构上退化的语料库在语料库层面表现较差,但在答案层面表现较好。最重要的是,HCF在若干情况下识别出了矛盾的检索证据,而此时答案仍然与标准答案相符。HCF并不认证事实的真实性,它使支持或质疑某一答案的证据变得可审查、可溯源。
cs.AI / 83 / 2609.07095

Risk Is Not Review Value: Wrong-Answer Exposure Under Bounded Review Budgets

风险并非审阅价值:有限审阅预算下的错误答案暴露问题
Park, SangJin, Choi, Myungsub, Kim, Jineok, Kang, Minseung
Abstract
LLM assistants often produce more answers than humans can review before users see them. Most evaluations ask whether an answer is wrong, unsupported, or low-confidence. Bounded review budgets instead ask which answers should be checked first under a fixed review budget. Risk alone is not enough: a high-risk answer may be hard to repair, while a moderately risky answer may be directly correctable from available evidence. For generated-answer evaluation, we model review prioritization as exposure reduction, where review value combines estimated wrongness, intervention affordance, impact, and cost. We evaluate review queues with Wrong-Answer Exposure Ratio (WAER), the fraction of wrong answers left unreviewed, and post-repair residual exposure (PRRE), the fraction still exposed after deterministic benchmark-supported repairs. PRRE uses repairability rules that do not numerically reuse the affordance scores used for ranking. On a 720-item TAT-QA/SciFact stress benchmark, review-value ranking keeps answer-level WAER nearly unchanged at 20% budget (0.605 vs. 0.600) but lowers PRRE from 0.881 to 0.716. These results show that trustworthy LLM evaluation should measure not only error detection, but also how limited review capacity reduces exposed wrong answers.
Chinese Translation
LLM助手生成的答案数量往往超出了用户在看到答案之前人类所能审阅的范围。大多数评估关注的是答案是否错误、缺乏依据或置信度低,而有限审阅预算要回答的问题则是:在固定审阅预算下,哪些答案应被优先检查。仅有风险并不足够:一个高风险答案可能难以修复,而一个中等风险的答案却可能借助现有证据直接被纠正。针对生成答案的评估,我们将审阅优先级排序建模为暴露减少问题,其中审阅价值综合了估计错误程度、干预可行性、影响力和成本。我们使用错误答案暴露率(Wrong-Answer Exposure Ratio, WAER)——即未被审阅的错误答案占比——以及修复后残余暴露率(post-repair residual exposure, PRRE)——即在基于基准数据支持的确定性修复之后仍暴露的错误答案占比——来评估审阅队列。PRRE所使用的可修复性规则不会在数值上重复用于排序的可行性评分。在一个包含720个条目的TAT-QA/SciFact压力测试基准上,审阅价值排序在20%预算下使答案级WAER基本保持不变(0.605 vs. 0.600),但将PRRE从0.881降至0.716。这些结果表明,可信赖的LLM评估不仅应衡量错误检测能力,还应衡量有限的审阅能力如何减少暴露的错误答案。
cs.AI / 84 / 2609.07107

Beyond Sparse Rewards: A New Benchmark and Structure-Aware Graph Alignment for Micro-Drama Understanding

超越稀疏奖励:面向微短剧理解的新基准与结构感知图对齐方法
Qin, Yixin, Chen, Shi-Zhe, Yu, Zhiqi, Cheng, Siyuan, Cheng, Tao, Luo, Jinwen, Wei, Zheng
Abstract
Micro-dramas, characterized by ultra-short durations and hyper-dense storylines, pose unique challenges for video understanding that conventional benchmarks fail to address. To bridge this gap, we introduce M-Drama, the first large-scale bilingual benchmark for micro-drama comprehension, featuring over 35K instances across 9,138 clips. Furthermore, while reinforcement learning can enhance VLMs on complex narratives, existing reward metrics often suffer from sparse and superficial signals, failing to capture intricate character identities and temporal structures. We propose SAGA (Structure-Aware Graph Alignment), a novel graph-matching reward function that models narratives as heterogeneous graphs. SAGA computes dense, rigorous rewards via decoupled semantic triplet and structural temporal matching. Extensive experiments on Qwen3-VL-8B-Instruct demonstrate that SAGA outperforms existing baselines, delivering substantial improvements in open-ended accuracy and summary quality, while maintaining competitive out-of-domain generalization. Code is available at https://github.com/qyx1121/MDrama_SAGA.
Chinese Translation
微短剧(Micro-drama)以其超短时长和高密度剧情为特点,对视频理解提出了传统基准测试无法应对的独特挑战。为弥合这一差距,我们提出了 M-Drama,首个面向微短剧理解的大规模双语基准,涵盖 9,138 个片段中的超过 35K 个实例。此外,尽管强化学习可以增强视觉语言模型(VLM)处理复杂叙事的能力,但现有的奖励指标往往存在信号稀疏和浅层化的问题,无法捕捉复杂的人物身份和时间结构。我们提出了 SAGA(Structure-Aware Graph Alignment,结构感知图对齐),这是一种新颖的图匹配奖励函数,将叙事建模为异构图。SAGA 通过解耦的语义三元组匹配和结构化时间匹配来计算稠密且严格的奖励。在 Qwen3-VL-8B-Instruct 上的大量实验表明,SAGA 优于现有基线方法,在开放式准确率和摘要质量方面带来了显著提升,同时保持了具有竞争力的域外泛化能力。代码可在 https://github.com/qyx1121/MDrama_SAGA 获取。
cs.AI / 85 / 2609.07128

EEG-Driven Decoding Framework for Passenger Hazard Perception in Highly Automated Vehicles

面向高度自动驾驶车辆乘客危险感知的脑电(EEG)驱动解码框架
Yang, Yingkai, Tan, Ashton Yu Xuan, Li, Bowen, Gao, Xiaorong, Zheng, Sifa, Wang, Jianqiang, Gu, Xinyu, Zhao, Yang, Zhang, Yuxin, Huang, Sharon X., Stathaki, Tania, Li, Jun, Wang, Hong
Abstract
Reliable risk assessment remains a central challenge for Autonomous Vehicles (AVs). Despite advances in automation, passenger cognition provides a non-intrusive auxiliary signal that improves both objective and perceived safety without requiring active human intervention. We introduce an Electroencephalogram (EEG)-based Brain-Computer Interface (BCI) that decodes passenger neural responses for both Risk Prediction (RP) and Danger Identification (DI), explicitly modeling humans as passengers to match real-world AV use. To achieve this, we propose the Passenger Cognitive Model (PCM), Risk-aware Sequential Labeling (RSL), and the Passenger EEG Decoding Strategy (PEDS), which integrates a 3D Convolutional Recurrent Neural Network (3D-CRNN) model for joint EEG decoding. Experimental results show that 3D-CRNN achieves a Balanced Accuracy (BA) of $95.3\% \pm 2.7\%$ in RP and improves single-subject DI from $80.9\% \pm 3.9\%$ to $85.0\% \pm 3.2\%$ with RSL. Event-wise analyses further show that 3D-CRNN consistently outperforms other models across different event types in RP and DI. In generalization experiments, 3D-CRNN achieves $77.0\% \pm 5.3\%$ BA in cross-session DI and $77.4\% \pm 1.1\%$ BA on seen subjects in cross-subject evaluation, while maintaining a $64.9\% \pm 8.5\%$ BA on unseen subjects, demonstrating promising generalizability and transferability across both intra-subject and inter-subject variability. These findings establish an Electroencephalogram (EEG) decoding framework for AV passenger hazard perception and suggest that passenger cognitive signals can provide auxiliary supervision for future AV decision-making and Safety of the Intended Functionality (SOTIF) support.
Chinese Translation
可靠的风险评估仍然是自动驾驶车辆(Autonomous Vehicles, AVs)面临的核心挑战。尽管自动化技术不断进步,乘客认知可作为一种非侵入式辅助信号,在无需人类主动干预的情况下同时提升客观安全性与主观感知安全性。我们提出了一种基于脑电图(Electroencephalogram, EEG)的脑机接口(Brain-Computer Interface, BCI),用于解码乘客的神经响应,以实现风险预测(Risk Prediction, RP)和危险识别(Danger Identification, DI)两项任务,并明确将人类建模为乘客,以契合真实世界的自动驾驶车辆使用场景。为此,我们提出了乘客认知模型(Passenger Cognitive Model, PCM)、风险感知序列标注(Risk-aware Sequential Labeling, RSL)和乘客脑电解码策略(Passenger EEG Decoding Strategy, PEDS),并集成了三维卷积循环神经网络(3D-CRNN)模型以实现联合脑电解码。实验结果表明,3D-CRNN 在风险预测任务中达到了 95.3% ± 2.7% 的平衡准确率(Balanced Accuracy, BA),并借助 RSL 将单被试危险识别性能从 80.9% ± 3.9% 提升至 85.0% ± 3.2%。事件级分析进一步表明,3D-CRNN 在风险预测和危险识别的不同事件类型上均持续优于其他模型。在泛化实验中,3D-CRNN 在跨会话危险识别任务中达到 77.0% ± 5.3% 的平衡准确率,在跨被试评估中对已见被试达到 77.4% ± 1.1% 的平衡准确率,同时对未见被试仍保持 64.9% ± 8.5% 的平衡准确率,展现出良好的泛化性与可迁移性,能够应对被试内部与被试之间的变异性。这些研究结果建立了一个面向自动驾驶车辆乘客危险感知的脑电解码框架,并表明乘客认知信号可为未来自动驾驶车辆的决策制定及预期功能安全(Safety of the Intended Functionality, SOTIF)支持提供辅助监督。
cs.AI / 86 / 2609.07139

Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise

早编码、晚使用:Transformer 从何处开始作用于推断出的伙伴专业水平
Okamoto, Mika, Sarti, Gabriele
Abstract
A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute does not yet influence the output. This gap between where information is readable and where it is used has been shown for attributes stated directly in the input. We ask whether it also holds for an attribute the model must infer gradually over a conversation, namely how expert its dialogue partner is. Using ExpertCollab, a corpus of multi-turn research-planning dialogues between model-played personas at four expertise levels, we find that partner expertise is most decodable in the early layers and falls to near chance before the midpoint of the network. Counterfactual patching shows that injecting the expertise difference at the layer of peak decodability barely changes a fixed late-layer readout, whereas the same difference injected past the midpoint propagates almost completely, a separation of more than an order of magnitude. A content-matched random control and a probe-free diagnostic place the transition at the same early layer, and a statically specified control attribute stays decodable throughout. An inferred relational attribute is therefore represented well before it becomes causally active, which bounds where any attempt to read out or steer partner-conditioned behavior must intervene. We use one model on a synthetic corpus as an initial demonstration.
Chinese Translation
Transformer 可以在其残差流的某一深度使某个属性线性可解码,而该属性在此深度尚未影响输出。这种“信息可读位置”与“信息被使用位置”之间的差距,此前已在输入中直接陈述的属性上得到验证。我们探讨该现象是否同样适用于模型必须在对话中逐步推断的属性,即其对话伙伴的专业水平。我们使用 ExpertCollab——一个由模型扮演的四种专业水平角色之间进行多轮科研规划对话构成的语料库——发现伙伴专业水平在早期层中最可解码,并在网络中点之前降至接近随机水平。反事实修补(counterfactual patching)实验表明,在可解码性峰值层注入专业水平差异几乎不改变固定的后期层读出结果,而在中点之后注入同样的差异则几乎完全传播,两者相差超过一个数量级。内容匹配的随机对照和无探针诊断方法将这一转变定位于相同的早期层,而静态指定的对照属性则在全网络中保持可解码。因此,被推断的关系属性在被因果激活之前早已被表征,这限定了任何读取或引导伙伴条件行为所必须干预的位置。我们在合成语料库上使用单个模型作为初步演示。
cs.AI / 87 / 2609.07152

An Auditable Symbolic-RAG-Generative AI Architecture for Goal-Oriented Conversation Orchestration

一种面向目标导向对话编排的可审计符号-RAG-生成式人工智能架构
Gonzalez, Ramon, Diaz, Antonio
Abstract
Goal-oriented conversational systems must answer factual questions, understand visitor-provided information, and advance business objectives without becoming rigid questionnaires. This paper proposes a Symbolic-RAG-Generative architecture centered on the Goal-oriented Retrieval-Augmented Conversation Engine (GRACE). An instruction-constrained Business Goal Compiler transforms business intent into an immutable objective set, normalized priority vector, canonical questions, and initial state vector. At runtime, GRACE receives the complete conversation history, latest visitor message, current state, and grounded answer generated by a separate RAG component. It updates completion only from visitor-authored evidence and selects one contextually modulated follow-up. The core policy maximizes expected business progress subject to a minimum visitor-utility constraint. We formalize the state, monotonic transitions, source separation, question modulation, and constrained policy; present the reference architecture; and define an evaluation comprising 24 English real-estate and 10 Spanish professional-cleaning conversations, totaling 119 protocol-defined visitor turns. Across both domains, GRACE achieves 84.9% exact state-transition accuracy, 91.6% evidence precision, 89.6% evidence recall, 100% monotonicity, and 94.1% terminal-state accuracy. The evaluation establishes compelling symbolic-state performance across standard, multi-goal, RAG-detour, validation, refusal, and robustness scenarios.
Chinese Translation
目标导向的对话系统必须能够回答事实性问题、理解访问者提供的信息,并推进业务目标,同时避免沦为僵化的问卷调查。本文提出了一种以目标导向检索增强对话引擎(Goal-oriented Retrieval-Augmented Conversation Engine,GRACE)为核心的符号-RAG-生成式架构。一个受指令约束的业务目标编译器将业务意图转化为不可变的目标集合、归一化的优先级向量、规范化问题列表以及初始状态向量。在运行时,GRACE接收完整的对话历史、最新的访问者消息、当前状态,以及由独立的RAG组件生成的有据可依的答案。它仅根据访问者主动提供的证据更新任务完成状态,并选择一个结合上下文进行调节的后续问题。核心策略在满足最低访问者效用约束的前提下,最大化预期的业务进展。我们对状态、单调转移、来源分离、问题调节和约束策略进行了形式化定义;给出了参考架构;并定义了一项包含24段英语房地产对话和10段西班牙语专业清洁服务对话的评估,共计119轮符合协议定义的访问者对话轮次。在两个领域中,GRACE实现了84.9%的精确状态转移准确率、91.6%的证据精确率、89.6%的证据召回率、100%的单调性以及94.1%的终止状态准确率。该评估在标准、多目标、RAG绕行、验证、拒绝和鲁棒性等场景中确立了令人信服的符号状态性能。
cs.AI / 88 / 2609.07174

PhysMAS: Physics-Grounded Multi-Agent Synthesis of Compositional 4D Gaussians

PhysMAS:基于物理的多智能体合成组合式4D高斯
Qin, Jiang, Lv, Chunji, Wei, Yangguang, Gao, Yang, Liu, Ming, Ding, Lizhong, Yuan, Ye, Lei, Yinjie, Li, Changsheng
Abstract
Efficient, fully automatic, and physically plausible 4D Gaussian synthesis is an important goal for dynamic scene generation. Recent physics-based methods couple 3D Gaussians with the Material Point Method (MPM) to generate physically driven motion, but extending this paradigm to heterogeneous multi-part objects and interacting multi-object scenes remains challenging. Object-level physical assignment collapses distinct parts into a single material state, while one-shot predictions from large language models, vision-language models, or agents neither reliably bind different materials to identified parts nor verify that the resulting MPM configuration is executable. Score Distillation Sampling (SDS)-based parameter optimization, meanwhile, requires repeated per-scene score evaluations and gradient backpropagation, incurring lengthy optimization and potentially yielding suboptimal or unstable solutions. We therefore present PhysMAS, a physics-grounded multi-agent framework. From a motion prompt and four scene views, an Object-Part Scene Agent establishes persistent identities and calls a Material Reasoning Agent for part-wise profiles. It invokes solver-aware skills to bind these identities and profiles to per-particle MPM fields and execute all objects in a shared domain; the framework then screens candidate forward-simulation results. This supports heterogeneous multi-part and interacting multi-object scenes without per-scene diffusion-score backpropagation. Extensive experiments demonstrate that, compared with recent physics-based 4D Gaussian baselines that rely on SDS, PhysMAS achieves better semantic alignment and perceived physical plausibility while requiring less runtime.
Chinese Translation
高效、全自动且物理上合理的4D高斯合成是动态场景生成的一个重要目标。近期基于物理的方法将3D高斯与物质点法(Material Point Method, MPM)耦合以生成物理驱动的运动,但将该范式扩展到异构多部件对象以及多对象交互场景仍具挑战性。对象级的物理属性分配会将不同部件坍缩为单一材料状态,而大语言模型、视觉语言模型或智能体的一次性预测既无法可靠地将不同材料绑定到已识别的部件上,也无法验证由此得到的MPM配置是否可执行。与此同时,基于分数蒸馏采样(Score Distillation Sampling, SDS)的参数优化需要对每个场景反复进行分数评估和梯度反向传播,导致优化耗时冗长,并可能产生次优或不稳定的解。为此,我们提出了PhysMAS,一个基于物理的多智能体框架。该框架从运动提示和四个场景视角出发,由对象-部件场景智能体(Object-Part Scene Agent)建立持久的部件身份,并调用材料推理智能体(Material Reasoning Agent)获取逐部件的材料属性。随后,它调用具备求解器感知能力的技能,将这些身份与材料属性绑定到逐粒子的MPM场中,并在共享域中执行所有对象的模拟;框架随后对候选正向模拟结果进行筛选。这使得该框架无需逐场景的扩散分数反向传播即可支持异构多部件对象和多对象交互场景。大量实验表明,与近期依赖SDS的基于物理的4D高斯基线方法相比,PhysMAS在实现更好的语义对齐和感知物理合理性的同时,所需的运行时间更少。
cs.AI / 89 / 2609.07194

EmoMed: An Emotionally-Aware Agent for Multimodal Medical Support with Real-Time Information Retrieval

EmoMed:一个具备情感感知能力、支持实时信息检索的多模态医疗辅助智能体
Nasonov, Ivan, Glazkov, Nikita, Makovetskiy, Ivan, Mozikov, Mikhail, Sukhorukov, Daniil, Savchenko, Andrey, Makarov, Ilya
Abstract
We present EmoMed - a multimodal medical consultation agent that adapts its responses based on users' emotional states while maintaining clinical accuracy. The system processes text and medical images, detects affect indicators (anxiety, confusion, urgency) from user input, and adjusts response tone, structure, and detail level accordingly. To ensure factual reliability, the agent grounds clinical information through a dual retrieval mechanism: web-based fact-checking and an API-connected, continuously updated medical knowledge base. We evaluate our approach across seven state-of-the-art language models (GPT-4/5, Qwen3, Llama 4, Gemini 2.5, Grok4, Claude3) using comprehensive metrics including LLM-as-judge assessments, MedQA style accuracy tests, BERT Score, safety/helpfulness ratings, and multimodal medical benchmarks. The results demonstrate that emotionally adaptive responses consistently outperform neutral baseline across evaluation dimensions, without compromising clinical accuracy. A controlled user study validated these findings, with participants reporting improved perceived empathy and communication clarity, while maintaining trust in factual accuracy. Source code: https://github.com/NasonovIvan/EmoMed-Agent
Chinese Translation
我们提出了EmoMed——一个多模态医疗咨询智能体,它能够在保持临床准确性的同时,根据用户的情绪状态调整其回应。该系统处理文本和医学图像,从用户输入中检测情感指标(焦虑、困惑、紧迫感),并相应地调整回应的语气、结构和详细程度。为确保事实可靠性,该智能体通过双重检索机制为临床信息提供依据:基于网络的事实核查和通过API连接、持续更新的医学知识库。我们在七个最先进的大语言模型(GPT-4/5、Qwen3、Llama 4、Gemini 2.5、Grok4、Claude3)上评估了我们的方法,采用的综合评估指标包括LLM-as-judge评估、MedQA风格的准确率测试、BERT Score、安全性/有用性评分以及多模态医学基准测试。结果表明,情感自适应回应在各个评估维度上均持续优于中性基线,且不损害临床准确性。一项受控用户研究验证了这些发现,参与者报告感知到的共情和沟通清晰度有所提升,同时对事实准确性的信任度得以保持。源代码:https://github.com/NasonovIvan/EmoMed-Agent
cs.AI / 90 / 2609.07204

Agentic Algorithm Engineering: Improving Shared-Memory Exact Minimum Cuts

智能体算法工程:改进共享内存精确最小割算法
Bader, David A., Chhabra, Adil, Großmann, Ernestine, Henzinger, Monika, Noe, Alexander, Schulz, Christian
Abstract
The minimum cut problem for an undirected edge-weighted graph asks us to divide its set of nodes into two blocks while minimizing the weighted sum of the cut edges. Over the last years, we engineered a range of fast algorithms for this problem. Our fastest exact algorithm uses an inexact algorithm to obtain a better bound for the problem, reductions that depend on this bound, improved data structures and parallel contraction routines. It is available in the open-source package VieCut and, on real-world instances, outperformed the previously fastest solvers by a factor of up to 2.5 sequentially and up to 12.9 when run in parallel. We improve this algorithm using agentic algorithm engineering (AAE), a methodology that we introduce here, in which autonomous large language model agents run the algorithm engineering cycle on an existing code base: they form hypotheses about where running time is lost, implement them, benchmark the result on a fixed instance set and keep or discard the change. Even though we had already tuned our algorithm by hand extensively, the agent finds significant optimizations, in particular on the DIMACS core instances: factors of 1.28 (sequential) and 1.63 (32 threads) on real-world k-cores, and 6.26 and 127 on the DIMACS core instances.
Chinese Translation
无向边加权图的最小割问题要求将图的节点集划分为两个块,同时使割边的加权和最小。近年来,我们为该问题设计了一系列快速算法。我们最快的精确算法利用一个非精确算法获得问题的更优上界,采用依赖该上界的归约技术、改进的数据结构以及并行收缩例程。该算法已在开源软件包 VieCut 中发布,在真实世界实例上,其性能在串行运行时比此前最快的求解器高出最多 2.5 倍,在并行运行时高出最多 12.9 倍。我们利用智能体算法工程(Agentic Algorithm Engineering, AAE)进一步改进了该算法——这是我们在本文中提出的一种方法论,即让自主的大语言模型智能体在现有代码库上运行算法工程循环:它们对运行时间的损失位置提出假设,实现这些假设,在固定的实例集上进行基准测试,并保留或舍弃相应的改动。尽管我们已经对该算法进行了大量手工调优,智能体仍发现了显著的优化点,尤其是在 DIMACS 核心实例上:在真实世界的 k-核(k-core)实例上分别取得 1.28 倍(串行)和 1.63 倍(32 线程)的加速,在 DIMACS 核心实例上分别取得 6.26 倍和 127 倍的加速。
cs.AI / 91 / 2609.07213

Unraveling the Real Working Mechanism and Inherent Flaws of GAE: A Method for Interpreting Transformer Processes from an Economic Perspective

揭示GAE的真实工作机制与固有缺陷:一种从经济学视角解释Transformer过程的方法
Cui, Yongjin, Fan, Xiaohui
Abstract
We observe a phenomenon that current algorithmic research in the field of explainable artificial intelligence primarily pursues better performance on several proxy metrics. On the one hand, these proxy metrics themselves are more or less flawed and cannot properly measure the quality of methods. On the other hand, metric-oriented research approaches often lead to the neglect of the rationality and interpretability of the methods themselves. Explainable artificial intelligence is abbreviated as XAI. The metric-driven research paradigm has resulted in a lack of interpretability of the relevant XAI methods themselves. Accordingly, there is a need for interpretability research on XAI methods, which can be playfully referred to as XXAI. This paper is one of our works on XXAI. This paper takes Generic Attention-model Explainability (GAE), a widely influential model interpretation method , or rather, XAI method that represents an important technical route, as the research object, and explores the real working mechanism and flaws of this method as well as the technical route it represents. Based on the conclusions of this study, it may be necessary to re-examine or verify GAE-related methods and their domain applications. We argue that GAE is an interpretation method that focuses on the attention process. After pointing out the working mechanism and flaws of GAE, we propose Cumulative Asset Holdings (CAH), a more reasonable Transformer interpretation method integrating both process-based and feature-based ideas from an economic zero-sum games perspective. In addition, it is worth noting that our method is applicable to models with special tokens, where existing methods may suffer from limitations. The model simplification research method and the analysis of additive operations adopted in this study may provide inspiration for other research works in XAI.
Chinese Translation
我们观察到一种现象:当前可解释人工智能领域的算法研究主要追求在若干代理指标上取得更好的性能。一方面,这些代理指标本身或多或少存在缺陷,无法正确衡量方法的质量;另一方面,以指标为导向的研究路径往往导致对方法本身的合理性与可解释性的忽视。可解释人工智能缩写为XAI。指标驱动的研究范式导致相关XAI方法本身缺乏可解释性。因此,有必要对XAI方法开展可解释性研究,我们可以趣味性地将其称为XXAI。本文是我们在XXAI方向上的工作之一。本文以通用注意力模型可解释性方法(Generic Attention-model Explainability,GAE)——一种具有广泛影响力的模型解释方法,也即代表重要技术路线的XAI方法——为研究对象,探讨该方法及其所代表技术路线的真实工作机制与缺陷。基于本研究的结论,或许有必要重新审视或验证与GAE相关的方法及其领域应用。我们认为,GAE是一种关注注意力过程的解释方法。在指出GAE的工作机制与缺陷之后,我们从经济学的零和博弈视角出发,融合基于过程和基于特征两种思想,提出了一种更为合理的Transformer解释方法——累积资产持有量(Cumulative Asset Holdings,CAH)。此外,值得注意的是,我们的方法适用于包含特殊标记(special tokens)的模型,而现有方法在此类模型上可能存在局限性。本研究采用的模型简化研究方法以及对加性运算的分析,或可为XAI领域的其他研究工作提供启发。
cs.AI / 92 / 2609.07222

Distance-Aware Attention and Wall-Distance Expert Routing for Transformer-Based 3D Flow Prediction

面向基于Transformer的三维流场预测的距离感知注意力与壁面距离专家路由
Kim, Sanghyeon, Yang, Sunwoong, Kang, Namwoo
Abstract
Transformer surrogates for 3D flow prediction compress an industrial mesh into a small set of tokens from which every prediction point reads. Two operations follow: the retrieval step in which a point gathers information from the compressed representation, and the feed-forward layer that transforms what it retrieved. In current backbones both are blind to where the point sits in the flow. We condition both on wall-related physical signals. Distance-aware cross-attention (DA-CA) reshapes each volume query by its wall distance before retrieval, so that a point deep in the boundary layer draws different geometric information than one in the outer flow. Surface-volume mixture-of-experts (SVMoE) replaces the shared feed-forward layer with a small set of experts, routed by wall distance for volume points and by local geometry for surface points. Neither mechanism is tied to one architecture, so we apply both unchanged to AB-UPT and Transolver-3. On DrivAerML with 50 training cases, DA-CA reduces the volume pressure error by 10.1%, and DA-CA and SVMoE together reduce it by 12.5%; DA-CA improves the near-wall region at some cost in the far region, which SVMoE recovers, and the volume experts settle into near-wall, transition, and free-stream bands without routing supervision. Retrained on 300 cases, the conditioning improves every field quantity, reducing volume pressure and velocity errors by 33.1% and 18.6% on AB-UPT and by 21.4% and 21.3% on Transolver-3. Under Leave-One-Body-Out evaluation on DrivAerNet++, it reduces the volume pressure error on unseen body types by up to 14.2%.
Chinese Translation
用于三维流场预测的Transformer代理模型将工业网格压缩为一小组token,每个预测点都从这些token中读取信息。随后涉及两个操作:检索步骤(即预测点从压缩表示中收集信息)和前馈层(对检索到的内容进行变换)。在当前的骨干网络中,这两个操作均对预测点在流场中的位置不敏感。我们用与壁面相关的物理信号对这两个操作进行条件化。距离感知交叉注意力(DA-CA)在检索之前根据壁面距离对每个体网格查询进行重塑,使得处于边界层深处的点与外流区域的点提取不同的几何信息。表面-体积专家混合(SVMoE)用一小组专家替换共享的前馈层,对体积点按壁面距离路由,对表面点按局部几何路由。这两种机制均不依赖于特定架构,因此我们将它们不加修改地应用于AB-UPT和Transolver-3。在包含50个训练案例的DrivAerML数据集上,DA-CA将体积压力误差降低了10.1%,DA-CA与SVMoE相结合则降低了12.5%;DA-CA改善了近壁区域,但在远场区域略有损失,而SVMoE能够弥补这一损失,且体积专家在无路由监督的情况下自发形成了近壁、过渡和自由流三个区域。在300个案例上重新训练后,该条件化方法改进了所有场量,在AB-UPT上将体积压力和速度误差分别降低33.1%和18.6%,在Transolver-3上分别降低21.4%和21.3%。在DrivAerNet++的留一物体(Leave-One-Body-Out)评估中,该方法将未见物体类型的体积压力误差最多降低14.2%。
cs.AI / 93 / 2609.07247

Elastic Horizon: Discovering the Effective Interaction Frontier in Agentic Reinforcement Learning

弹性视野:发现智能体强化学习中的有效交互边界
Zhang, Gangyi, Meng, Junjie, Zhang, Letian, Wu, Wei, Zheng, Yang, Wang, Dong, Liu, Yang, Jiang, Guanjun, Gao, Chongming
Abstract
Scaling the interaction horizon-the maximum number of environment interactions per episode-improves LLM agents on long-horizon tasks, and curriculum-based methods that progressively expand the horizon outperform fixed-horizon alternatives. However, existing schedules are open-loop: they monotonically increase the horizon until a manually specified maximum, with no mechanism to detect when further expansion stops helping. We propose the effective interaction frontier hypothesis: a dynamic boundary beyond which additional interactions yield diminishing returns while cost grows linearly. We then introduce Elastic Horizon, a closed-loop controller that tracks this boundary via the 90th percentile of successful trajectory lengths. On AppWorld and BFCL, fixed-horizon sweeps reveal clear saturation plateaus; Elastic Horizon stabilizes the horizon inside the saturation band from both under- and over-capacity initializations, attains the best success rates across 7B and 14B backbones, and saves up to 25% of per-step trajectory tokens. Our work shifts the paradigm from how to scale interaction horizons to when to stop scaling.
Chinese Translation
扩大交互视野(即每个回合中与环境交互的最大次数)能够提升LLM智能体在长视野任务上的表现,而逐步扩展视野的基于课程的方法优于固定视野的方案。然而,现有的扩展调度是开环的:它们单调地增加视野直至人工指定的最大值,缺乏检测进一步扩展何时不再带来收益的机制。我们提出有效交互边界假设:存在一个动态边界,超过该边界后,额外的交互收益递减,而成本呈线性增长。在此基础上,我们提出了Elastic Horizon(弹性视野),一个通过成功轨迹长度的第90百分位数来追踪该边界的闭环控制器。在AppWorld和BFCL基准上,固定视野的扫描显示出明显的饱和平台;Elastic Horizon无论从容量不足还是容量过剩的初始设置出发,都能将视野稳定在饱和区间内,在7B和14B两种骨干模型上均取得最佳成功率,并节省高达25%的每步轨迹token。我们的工作将研究范式从“如何扩展交互视野”转变为“何时停止扩展”。
cs.AI / 94 / 2609.07255

SkillAlign: Aligning Skill Interfaces for LLM-based Agents

SkillAlign:面向基于大语言模型智能体的技能接口对齐
Ren, Shuo, Kang, Xiaomian, Zhang, Jiajun
Abstract
Language-model agents increasingly rely on skills: reusable procedural knowledge for reasoning, tool use, and interaction. Existing work studies how skills are acquired, retrieved, compressed, or composed, but often assumes that once a skill is selected, its interface to the agent is fixed. We argue that this overlooks a key source of skill utility: the same skill can help, distract, or mislead depending on how it is exposed. We propose SkillAlign, a provider-agnostic framework that represents candidate skills as multi-view procedural cards and renders them through alternative exposure interfaces, including full instructions, hints, compressed summaries, workflows, or no exposure. This enables counterfactual evaluation where the task, agent, and candidate skills are fixed while only the exposure interface varies. Across ALFWorld and SkillsBench, we show that exposure form substantially affects task success and rendered context cost, and that compact top-k exposure can outperform full-library injection. We further conduct a replay-based policy-learning analysis on ALFWorld, showing that adaptive exposure contains learnable signal but remains far from oracle selection. Our results suggest that skill-augmented agents should optimize not only which skills to use, but also how those skills are presented.
Chinese Translation
语言模型智能体日益依赖技能:即用于推理、工具使用和交互的可复用程序性知识。现有工作研究技能如何被获取、检索、压缩或组合,但通常假设一旦选定某个技能,其面向智能体的接口便是固定的。我们认为这忽视了技能效用的一个关键来源:同一技能可能有所帮助、造成干扰或产生误导,取决于其暴露方式。我们提出 SkillAlign,一个与提供方无关的框架,将候选技能表示为多视角程序性卡片,并通过多种可选的暴露接口进行呈现,包括完整指令、提示、压缩摘要、工作流或不暴露。这使得反事实评估成为可能:在任务、智能体和候选技能固定的情况下,仅改变暴露接口。在 ALFWorld 和 SkillsBench 上的实验表明,暴露形式会显著影响任务成功率和渲染上下文成本,且紧凑的 top-k 暴露可以优于全库注入。我们进一步在 ALFWorld 上开展了基于重放的策略学习分析,结果表明自适应暴露包含可学习的信号,但与最优选择仍有较大差距。我们的研究结果表明,技能增强型智能体不仅应优化使用哪些技能,还应优化这些技能如何被呈现。
cs.AI / 95 / 2609.07299

World Models Under Asynchronous Sensor Observations

异步传感器观测下的世界模型
Anand, Akash, Anand, Abhay, Vishe, Yash
Abstract
Learned world models typically assume that observations arrive synchronously, an abstraction inherited from simulators that return a complete state vector at each environment step. Physical sensing instead operates at heterogeneous rates, leaving most observation channels stale at any given instant. Interpolating stale channels introduces measurements that were never observed, while downsampling to the slowest sensor discards valid measurements. A natural alternative is to zero-order-hold the most recent reading and provide the known sampling schedule to the model through two features, staleness and time-to-refresh. We test this prediction using transformer world models across three regimes of increasing causal coupling: open-loop rollouts in continuous-control locomotion, closed-loop model-predictive planning in which each learned model serves as the planner dynamics, and a linear latched-actuator system in which refresh events apply a zero-order-held command to the plant. Our findings show that the effectiveness of time-to-refresh depends on the causal role of the sampling schedule, specifically when refresh events affect the system rather than merely report its state. These results establish when sampling schedules provide useful information for predictive world models operating under asynchronous physical observations.
Chinese Translation
学习型世界模型通常假设观测是同步到达的,这一抽象源自那些在环境每个步骤返回完整状态向量的模拟器。然而,物理传感以异构速率运行,使得在任意给定时刻大多数观测通道的数据是过时的。对过时通道进行插值会引入从未被观测到的测量值,而降采样至最慢传感器的速率则会丢弃有效的测量数据。一种自然的替代方案是对最近一次读数进行零阶保持(zero-order-hold),并通过两个特征——陈旧度(staleness)和距刷新时间(time-to-refresh)——将已知的采样调度提供给模型。我们在因果耦合程度递增的三种情形下,使用Transformer世界模型对这一预测进行了检验:连续控制运动任务中的开环滚动预测、闭环模型预测规划(其中每个学习到的模型作为规划器的动力学模型),以及一个线性锁存执行器系统(其中刷新事件将零阶保持的指令施加到被控对象上)。研究结果表明,距刷新时间特征的有效性取决于采样调度的因果作用,具体而言,即刷新事件是影响系统本身,还是仅仅报告系统状态。这些结果阐明了在异步物理观测条件下,采样调度何时能为预测型世界模型提供有用信息。
cs.AI / 96 / 2609.07313

Weakly supervised neural network: segmentation of complex structures in X-ray microCT

弱监督神经网络:X射线显微CT中复杂结构的分割
Rusconi, Daniele, Ascolese, Michela, Fest-Santini, Stephanie, Bravin, Alberto, Santini, Maurizio
Abstract
Segmentation of complex structures in X-ray tomographic data is a fundamental task in biomedical research, but it often requires large amounts of precisely annotated data, making fully supervised approaches costly and difficult to scale. In this study, weakly supervised deep learning is investigated as a strategy to reduce annotation effort while maintaining accurate segmentation. A two-dimensional convolutional neural network based on the nnU-Net framework was adapted to a weak supervision setting using sparse dot-based annotations, complemented by a limited number of fully segmented images. The approach was evaluated on high-resolution microCT slices of rat kidneys, targeting the segmentation of renal glomeruli, which are small, low-contrast anatomical structures. Results indicate that weak supervision provides a meaningful learning signal, enabling reliable localization of glomeruli even in the absence of dense labels. Incorporating a small set of high-quality annotations substantially improves segmentation performance, approaching that of a fully supervised model. These findings highlight the potential of weakly supervised learning as an annotation-efficient strategy for the analysis of complex structures in X-ray tomographic data, and suggest that alternative loss formulations tailored to sparse annotations may further enhance performance.
Chinese Translation
X射线断层成像数据中复杂结构的分割是生物医学研究中的一项基础任务,但通常需要大量精确标注的数据,使得全监督方法成本高昂且难以扩展。本研究探讨了弱监督深度学习作为一种在保持精确分割的同时减少标注工作量的策略。基于nnU-Net框架的二维卷积神经网络被改造为弱监督设置,采用稀疏的点状标注,并辅以少量完整分割的图像。该方法在高分辨率的大鼠肾脏microCT切片上进行评估,目标是分割肾小球——一种体积小、对比度低的解剖结构。结果表明,弱监督能够提供有意义的学习信号,即使在没有密集标注的情况下也能实现肾小球的可靠定位。引入少量高质量标注可显著提升分割性能,接近全监督模型的水平。这些发现凸显了弱监督学习作为X射线断层成像数据中复杂结构分析的标注高效策略的潜力,并提示针对稀疏标注设计的替代损失函数可能进一步提升性能。
cs.AI / 97 / 2609.07316

DGCPath: Distribution-Aware Generative Contrastive Framework for Self-supervised Path Representation Learning -- Extended Version

DGCPath:面向自监督路径表示学习的分布感知生成式对比框架——扩展版本
Yang, Sean Bin, Miao, Hao, Xu, Zongyi, Hu, Jilin, Wang, Xiangmeng, Lu, Hua, Yang, Bin, Jensen, Christian S.
Abstract
Due to the proliferation of vehicle trajectory data enabled by advanced sensing technologies, path representation learning has become a pivotal task in intelligent transportation systems. Although existing self-supervised approaches have achieved promising performance, their dependence on deterministic contrastive learning paradigms and handcrafted view augmentation strategies inherently restricts their cross-scenario generalization capabilities. To address these limitations, we present DGCPath, an innovative Distribution-aware Generative Contrastive learning framework for Path representation. This framework establishes a synergistic connection between generative modeling and distributional contrastive learning, enabling the acquisition of robust and transferable feature embeddings. Specifically, our framework incorporates: (1) a diffusion-based view generator that autonomously produces semantically coherent yet diverse trajectory views from Gaussian noise; (2) a variational contrastive mechanism that enforces latent feature alignment at the distribution level, transcending conventional instance-wise consistency; and (3) a novel generative cross-supervision module that reinforces view-level consistency through cross-view reconstruction learning. Comprehensive evaluations on three real-world trajectory datasets demonstrate that DGCPath outperforms state-of-the-art baselines on two distinct downstream tasks, validating its enhanced generalization capability and representation effectiveness.
Chinese Translation
得益于先进传感技术所催生的车辆轨迹数据的激增,路径表示学习已成为智能交通系统中的一项关键任务。尽管现有的自监督方法已取得了令人瞩目的性能,但它们对确定性对比学习范式和手工设计的视图增强策略的依赖,在本质上限制了其跨场景泛化能力。为解决这些局限性,我们提出了DGCPath,一种创新的面向路径表示的分布感知生成式对比学习(Distribution-aware Generative Contrastive learning)框架。该框架在生成式建模与分布对比学习之间建立了协同联系,能够获得鲁棒且可迁移的特征嵌入。具体而言,我们的框架包含:(1)一个基于扩散模型的视图生成器,可从高斯噪声中自主生成语义一致且多样化的轨迹视图;(2)一种变分对比机制,在分布层面实施潜在特征对齐,超越了传统的实例级一致性约束;(3)一个新颖的生成式交叉监督模块,通过跨视图重建学习强化视图级一致性。在三个真实世界轨迹数据集上的全面评估表明,DGCPath在两项不同的下游任务上均优于最先进的基线方法,验证了其增强的泛化能力和表示有效性。
cs.AI / 98 / 2609.07334

AAS-RAIL: Improving Information Extraction for Asset Administration Shells through Retrieval-Augmented In-Context Learning

AAS-RAIL:通过检索增强的上下文学习改进资产管理壳的信息抽取
Groß, Janek, Heidrich, Jens
Abstract
The Asset Administration Shell (AAS) is a cornerstone of Industry 4.0 and the Digital Product Passport, providing standardized digital representations of industrial assets. While manufacturers already maintain extensive technical product documentation, generating AAS instances from existing product datasheets remains a labor-intensive task because technical information is extracted from heterogeneous document structures and often involves company-specific terminology and conventions. In this work, we present AAS-RAIL, a retrieval-augmented information extraction (IE) approach that automatically generates Asset Administration Shells from PDF product datasheets using large language models (LLMs). Instead of relying on a fixed set of few-shot examples, the proposed retrieval-augmented in-context learning (RAIL) approach retrieves LLM-generated extraction helpers from similar Asset Administration Shells to provide instance-specific in-context learning (ICL). This enables the model to adapt its extraction behavior to company-specific naming conventions and formatting styles without fine-tuning. Our core contribution is the dynamic selection of company-specific AAS examples for each datasheet, replacing static prompting with an extraction pipeline that adapts to instances and combines semantic retrieval and structured information extraction. The proposed approach is evaluated on a collection of industrial product datasheets using a selection of open- and closed-weight LLMs. Experimental results show that RAIL consistently improves extraction quality over conventional few-shot prompting, yielding relative improvements of 30.4-52.4%. These results demonstrate that our approach provides an effective improvement for company-specific AAS generation.
Chinese Translation
资产管理壳是工业4.0和数字产品护照的基石,为工业资产提供标准化的数字化表示。尽管制造商已经维护着大量的技术产品文档,但由于技术信息需要从异构的文档结构中抽取,且往往涉及企业特定的术语和约定,从现有产品数据表中生成AAS实例仍然是一项劳动密集型任务。在本工作中,我们提出了AAS-RAIL,这是一种检索增强的信息抽取(IE)方法,利用大语言模型(LLM)从PDF产品数据表中自动生成资产管理壳。与依赖固定的少样本示例集不同,所提出的检索增强上下文学习方法从相似的资产管理壳中检索由大语言模型生成的抽取辅助信息,以提供实例特定的上下文学习(ICL)。这使得模型无需微调即可适应企业特定的命名规范和格式风格。我们的核心贡献在于为每份数据表动态选择企业特定的AAS示例,以自适应实例的抽取流水线取代静态提示,并结合语义检索与结构化信息抽取。我们在一组工业产品数据表上,使用多种开源与闭源权重大语言模型对该方法进行了评估。实验结果表明,RAIL相较于传统少样本提示方法持续提升抽取质量,相对改进幅度达30.4%至52.4%。这些结果证明我们的方法为企业特定的AAS生成提供了有效的改进。
cs.AI / 99 / 2609.07353

Human-like moral judgments conceal divergent motive attributions in large language models

类人道德判断掩盖了大语言模型中不同的动机归因
Wu, Xiaoyan, Dreher, Jean-Claude
Abstract
Large language models (LLMs) are used to simulate human participants in psychological research. We asked whether LLMs that reproduce human evaluations of a whistleblower's moral character also reproduce the motive attributions that accompany them. Five LLMs and two human samples (N = 125 and N = 742) evaluated a physician who either remained silent about fraudulent billing or reported it to a hospital, regulator, or newspaper. Models reproduced the human ranking of the physician's moral character but portrayed whistleblowers as more helpful, less self-interested, and less hostile. In four of five models, competitive motives were less strongly associated with moral-character judgments. Model ratings changed little when prompts reproduced the narratives and demographic profiles of both human samples, although this comparison cannot isolate a perspective effect. Thus, agreement in average ratings can conceal differences in attributed motives, relationships among judgments, and sensitivity to context. Validating LLMs as simulated participants therefore requires testing psychologically informative response patterns, not average agreement alone.
Chinese Translation
大语言模型(LLMs)被用于在心理学研究中模拟人类被试。我们探讨了那些能够复现人类对举报者道德品质评价的LLMs,是否也能复现伴随这些评价的动机归因。五个LLMs和两个人类样本(N = 125和N = 742)对一名医生进行了评价:该医生要么对虚假计费保持沉默,要么向医院、监管机构或报社举报。模型复现了人类对该医生道德品质的排序,但将举报者描绘得更有助人倾向、更少自利、更少敌意。在五个模型中的四个,竞争性动机与道德品质判断之间的关联较弱。当提示词复现了两个人类样本的叙述内容和人口统计特征时,模型评分变化甚微,尽管这种比较无法分离出视角效应。因此,平均评分上的一致性可能掩盖归因动机、判断之间的关系以及对情境敏感性的差异。因此,验证LLMs作为模拟被试的有效性,需要检验具有心理学意义反应模式,而不仅仅依赖平均一致性。
cs.AI / 100 / 2609.07409

RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavioral Monitoring in Surveillance Systems

RAFM-SER++:一种用于监控系统实时行为监测的轻量级多模态情感识别框架
Dinh, Ngo Truong, Bui, Tung-Lam, Duong, Chi-Trung, Thi, Vien Nguyen, Nguyen, Viet-Anh, Le, Phuc-Lu
Abstract
Recent multimodal Speech Emotion Recognition (SER) systems achieve high accuracy through interaction-heavy cross-modal transformers, but their computational cost limits deployment in latency-sensitive and resource-constrained surveillance systems. To address this challenge, we propose RAFM_SER++, a lightweight multimodal SER framework featuring an asymmetric Residual Attention Fusion Mechanism (RAFM). Rather than relying on computationally expensive bidirectional interactions, RAFM injects affective speech cues into semantic text representations through a one-directional residual attention pathway. Combined with a BYOL-inspired cross-modal alignment objective and attention-guided pooling, the proposed framework improves multimodal representation learning while maintaining low computational overhead. Experiments on the IEMOCAP and ESD benchmarks demonstrate that RAFM_SER++ consistently outperforms the HuBERT-Base baseline and achieves a superior accuracy-efficiency trade-off compared with the state-of-the-art MemoCMT. Specifically, RAFM_SER++ reduces trainable parameters by more than 60%, achieves faster inference (79.60 it/s), and attains BACC scores of 81.10% on IEMOCAP and 95.39% on ESD. These results indicate that lightweight asymmetric multimodal fusion is an effective alternative to interaction-heavy cross-modal transformers for real-time surveillance applications.
Chinese Translation
近年来,多模态语音情感识别(Speech Emotion Recognition, SER)系统通过交互密集型跨模态Transformer实现了较高的识别精度,但其高昂的计算成本限制了其在延迟敏感、资源受限的监控系统中的部署。为应对这一挑战,我们提出了RAFM_SER++,一种采用非对称残差注意力融合机制(Residual Attention Fusion Mechanism, RAFM)的轻量级多模态SER框架。RAFM不依赖计算代价高昂的双向交互,而是通过单向残差注意力通路将语音中的情感线索注入语义文本表示中。结合受BYOL启发的跨模态对齐目标和注意力引导池化,该框架在保持较低计算开销的同时提升了多模态表示学习能力。在IEMOCAP和ESD基准数据集上的实验表明,RAFM_SER++持续优于HuBERT-Base基线模型,并与最先进的MemoCMT相比取得了更优的精度-效率权衡。具体而言,RAFM_SER++将可训练参数减少60%以上,实现了更快的推理速度(79.60 it/s),并在IEMOCAP和ESD上分别达到81.10%和95.39%的BACC分数。这些结果表明,轻量级非对称多模态融合是实时监控应用中交互密集型跨模态Transformer的有效替代方案。
cs.AI / 101 / 2609.07434

CIT-CAD: Constraint Intent Tree-based CAD Code Generation and Verification

CIT-CAD:基于约束意图树的CAD代码生成与验证
Du, Yali, Sun, Hui, Xi, San-Zhuo, Li, Ming
Abstract
Natural-language Computer-Aided Design (CAD) code generation aims to turn design intent into executable and editable parametric programs. Large language models (LLMs) make this goal increasingly practical, but useful systems must preserve the construction process behind the rendered geometry. Existing benchmarks and methods mostly focus on how closely the generated CAD model matches the reference geometry, often using metrics such as Intersection over Union (IoU). Such metrics can miss errors in part decomposition, construction hierarchy, Boolean operations, sketch structure, and geometric relations. This gap calls for a representation that makes design intent explicit and lets a system check generated code against that intent. We propose CIT-CAD, a framework that infers a Constraint Intent Tree (CIT) from the input description to represent the intended entities, hierarchy, operations, and relations. The tree has two roles: it guides CAD code generation and defines expected constraints for verification. The framework extracts actual constraints from the generated program, compares them with the expected constraints, and uses mismatches to localize and repair design violations. Experiments show that the framework improves CAD generation performance, with larger gains on more complex multi-entity designs. By turning design intent into an explicit and checkable object, this work is the first attempt to move text-to-CAD generation beyond rendered-geometry matching toward construction-aware synthesis, verification, and repair.
Chinese Translation
自然语言计算机辅助设计(CAD)代码生成旨在将设计意图转化为可执行、可编辑的参数化程序。大语言模型(LLM)使这一目标日益可行,但实用的系统必须保留所渲染几何体背后的构建过程。现有基准和方法大多关注生成的CAD模型与参考几何体的匹配程度,通常使用交并比(IoU)等指标。此类指标可能遗漏零件分解、构建层级、布尔运算、草图结构和几何关系中的错误。这一空白需要一种能够显式表达设计意图、并使系统能够依据该意图检查生成代码的表示方法。我们提出CIT-CAD,一个从输入描述中推断约束意图树(Constraint Intent Tree, CIT)的框架,用以表示预期的实体、层级、操作和关系。该树具有双重作用:既指导CAD代码生成,又定义用于验证的预期约束。该框架从生成的程序中提取实际约束,将其与预期约束进行比较,并利用不匹配之处定位并修复设计违规。实验表明,该框架提升了CAD生成性能,在更复杂的多实体设计上收益更大。通过将设计意图转化为显式且可检查的对象,本工作首次尝试将文本到CAD的生成从渲染几何匹配推进到面向构建过程的合成、验证与修复。
cs.AI / 102 / 2609.07478

The Internal Anatomy of Strategic Choice in Large Language Models

大语言模型策略选择内部机制剖析
Ferraz, Vinícius, Houf, Leon, Ferrea, Enrico
Abstract
Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal $2\times2$ games. We followed a prespecified incentive from prompt, through activations, to choice. Dense models mirrored the unadjusted human decline with game complexity. Incentive and choice were detectable in every model, but models differed in whether incentive reached the choice, aligned with it and, where tested, whether strengthening it shifted preference. The base and instruction-tuned Qwen2.5 models chose almost identically at baseline yet differed in whether incentive reached choice. Fixed decision cues were distinguishable internally but changed choices selectively. Similar behaviour can rest on different computation; post-training can reshape the path from represented incentive to decision while leaving behaviour and decodable information largely intact.
Chinese Translation
大语言模型既作为策略智能体行动,也作为人类选择的模型存在,然而像策略智能体那样做出选择,并不意味着像策略智能体那样进行计算。我们记录了四个开源权重模型(包括稠密模型和专家混合模型,其中含一组匹配的基座—指令微调模型对)在144个严格序数2×2博弈的单次博弈中的激活值。我们沿着一条预先设定的激励路径进行追踪:从提示词,到激活值,再到选择。稠密模型复现了人类随博弈复杂度增加而出现的未调整的选择下降现象。激励与选择在所有模型中均可被检测到,但各模型在激励是否传导至选择、是否与选择保持一致,以及(在测试的情况下)强化激励是否会改变偏好等方面存在差异。基座版Qwen2.5与指令微调版Qwen2.5在基线条件下的选择几乎完全相同,但在激励是否传导至选择方面却有所不同。固定的决策线索在内部可被区分,但仅选择性地改变选择。相似的行为可能建立在不同的计算之上;后训练(post-training)可以在保持行为和可解码信息基本不变的情况下,重塑从被表征的激励到决策的传导路径。
cs.AI / 103 / 2609.07483

Modus Tollens and Counterfactuals and Counterfactual Reasoning Based on Three Types of Negation

否定词三类的反证法、反事实与反事实推理
Pan, Zhenghua
Abstract
Modus Tollens (MT) is a classical logical inference rule, while counterfactuals are hypothetical statements that are contrary to facts, and counterfactual reasoning is a process of reasoning based on counterfactuals. Negation is an indispensable core concept in them. In this paper, based on the logical systems LCOI&PLCOI with contradictory negation, opposite negation and intermediary negation, we propose three variants of Modus Tollens corresponding to distinct negation types, namely MTC: Modus Tollens based on contradictory negation, MTO: Modus Tollens based on opposite negation, and MTI: Modus Tollens based on intermediary negation. We define the implications within MTC, MTO and MTI, provide the truth value algorithms of MTC, MTO and MTI, and discuss the reducibility of these algorithms. To incorporate these three types of negation into counterfactuals and counterfactual reasoning, we differentiate counterfactuals into two types based on whether they possess logical negation, thereby proposing three counterfactuals and counterfactuals reasoning based on different logical negations. In this paper, we further argue that the three counterfactuals reasoning based on different logical negations have the same inference form as MTC, MTO and MTI, respectively. In other words, they share the same inference structure. As a result, the truth value algorithms for MTC, MTO and MTI can be as the truth value algorithms for the three counterfactuals reasoning based on different logical negations. The algorithms indicates that if the first premise of the reasoning is true, the truth values of the reasoning conclusions are identical to the truth values of the three negative premises in the reasoning premises, respectively. This reflects the consistency and accuracy of the truth value algorithms.
Chinese Translation
反证法(Modus Tollens, MT)是一种经典的逻辑推理规则,而反事实是与事实相反的假设性陈述,反事实推理则是基于反事实进行推理的过程。否定词在其中是不可或缺的核心概念。本文基于包含矛盾否定、对立否定和中介否定的逻辑系统 LCOI 与 PLCOI,提出了对应于不同否定类型的三种反证法变体,即:MTC——基于矛盾否定的反证法;MTO——基于对立否定的反证法;MTI——基于中介否定的反证法。我们定义了 MTC、MTO 和 MTI 中的蕴涵,给出了 MTC、MTO 和 MTI 的真值算法,并讨论了这些算法的可归约性。为了将这三类否定纳入反事实及反事实推理之中,我们依据反事实是否具有逻辑否定将其区分为两种类型,进而提出了基于不同逻辑否定的三种反事实及反事实推理。本文进一步论证,基于不同逻辑否定的三种反事实推理分别与 MTC、MTO 和 MTI 具有相同的推理形式。换言之,它们共享相同的推理结构。因此,MTC、MTO 和 MTI 的真值算法可以作为基于不同逻辑否定的三种反事实推理的真值算法。该算法表明,若推理的第一个前提为真,则推理结论的真值分别等于推理前提中三个否定前提的真值。这反映了真值算法的一致性与准确性。
cs.AI / 104 / 2609.07533

Quantile-Led Feature Extraction for Multi-Horizon Predictive Maintenance in Industrial Manufacturing Systems

面向工业制造系统多时间范围预测性维护的分位数主导特征提取方法
Poland, David J, Ravi, Daniele, Helian, Na
Abstract
In data-driven predictive maintenance (PdM), feature extraction is usually treated as fixed preprocessing: a descriptor set is chosen once and reused while the downstream model or forecasting horizon changes. This paper isolates the representation-learning stage and presents a quantile-led feature-extraction framework based on a dual-stage MLP-QRNN hierarchy. QRNN1 learns a broad ten-quantile conditional distribution for each sensor channel, while skip-connected QRNN2 refines a retained mid-tail quantile set into compact, channel-resolved, distribution-aware features. A fixed thirteen-pipeline ablation spans 1-hour, 70-hour, and 30-day regimes across 72 machines in 9 industrial facilities, with the downstream temporal classifier held fixed within each regime. Increasing the retained mid-tail set from two to four quantiles improves 30- and 60-minute F1-score, reaching 75.92% and 72.44% with attention enabled. The results also show that representations do not transfer reliably beyond their design horizon unless feature capacity, temporal embedding, activation strategy, and sensor breadth are scaled with the forecasting task. The unmodified short-horizon extractor falls to 42.90% F1 at 70 hours, whereas horizon-conditioned extractors reach 60.38% at 70 hours and 79.97% at 30 days. The framework therefore supports treating PdM feature extraction as a horizon-dependent representational stage rather than fixed preprocessing.
Chinese Translation
在数据驱动的预测性维护(PdM)中,特征提取通常被视为固定的预处理步骤:描述子集合仅选择一次,并在下游模型或预测时间范围变化时被重复使用。本文将表示学习阶段独立出来,提出了一种基于双阶段 MLP-QRNN 层级结构的分位数主导特征提取框架。其中,QRNN1 为每个传感器通道学习广泛的十分位数条件分布,而带有跳跃连接的 QRNN2 将保留的中尾分位数集合精炼为紧凑的、通道级的、分布感知的特征。一个固定的十三条流水线消融实验在 9 个工业设施的 72 台机器上,覆盖了 1 小时、70 小时和 30 天三种时间范围,且在每种时间范围内下游时序分类器保持不变。将保留的中尾分位数集合从两个增加到四个,可提升 30 分钟和 60 分钟的 F1 分数,在启用注意力机制时分别达到 75.92% 和 72.44%。结果还表明,除非特征容量、时间嵌入、激活策略和传感器广度随预测任务进行相应扩展,否则表示无法可靠地迁移到超出其设计时间范围的任务。未经修改的短时间范围提取器在 70 小时下 F1 分数降至 42.90%,而针对时间范围条件化的提取器在 70 小时和 30 天下分别达到 60.38% 和 79.97%。因此,该框架支持将 PdM 特征提取视为一个依赖于时间范围的表示学习阶段,而非固定预处理。
cs.AI / 105 / 2609.07559

Scoring Without the Engine: Validating a Deterministic, Manipulation-Resistant Content Score for Generative Engines, End to End

无需引擎的评分:为生成式引擎验证一种确定性、抗操纵的内容评分的端到端方案
Bajemon, Elisha, Rochet, Andre-Louis
Abstract
How do you validate a cheap, deterministic proxy for an oracle that is expensive, rate-limited, and non-stationary? We present a protocol built on adversarial falsification gates (negative control, dose response, bounded amplification, duplication penalty, length neutrality) that define and select the proxy, fitted on a training split and confirmed held-out; around them it bounds what the proxy can never resolve, and re-measures external causal evidence on the current oracle rather than assuming it. We demonstrate it end to end on Generative Engine Optimization, where the proxy is a deterministic content score, and one step fails on that domain exactly as the protocol is built to detect: re-measuring the only published causal anchors (2023 effect sizes) on ten modern engine families shows their levers move citation on none, so the anchors are an expired external check; recalibrating to the near-zero modern vector strips the score of its lever-responsive components. What survives is the gate-enforced response surface. The gates buy a measured property: on a 500-source benchmark of adversarial edits, amplifying the score's calibrated levers gains an attacker at most 6 points, and decreases with dose; single-lever amplification is provably bounded, while the cap and cross-lever sub-additivity are empirical findings consistent with it. On detection, web-spam baselines dominate and out-of-distribution attacks evade the score, so the deployable filter layers it over them. A query-conditioned skyline bounds the score's citation signal (within-query Spearman 0.11), repositioning query-agnostic scores as quality filters rather than citation predictors. A query-leakage bug in our first ranking evaluation and a failed confidence flag are disclosed and corrected; every number reproduces offline from released artifacts at zero marginal API cost.
Chinese Translation
如何为一个昂贵、受速率限制且非平稳的“神谕”(oracle)验证一种廉价、确定性的代理指标?我们提出一个建立在对抗性证伪门(negative control 阴性对照、dose response 剂量响应、bounded amplification 有界放大、duplication penalty 重复惩罚、length neutrality 长度中性)之上的协议,这些门用于定义并选择该代理指标;代理在训练集上拟合,并在留出集上得到确认。围绕这些门,协议划定了代理永远无法解决的边界,并在当前的神谕上重新测量外部因果证据,而非直接沿用既有假设。我们在生成式引擎优化(Generative Engine Optimization, GEO)上对该协议进行了端到端演示:代理指标是一种确定性的内容评分,而其中一个步骤在该领域上恰好按协议设计所预判的方式失败了——在十个现代引擎家族上重新测量唯一已发表的因果锚点(2023年的效应量)表明,其操控手段在任何引擎上都无法提升引用率,因此这些锚点已是一份过期的外部检验;将其重新校准到接近于零的现代向量后,评分中对应这些操控手段的响应成分随之被剥离。保留下来的正是由这些门所强制约束的响应面。这些门换来了可测量的性质:在一个包含500个来源的对抗性编辑基准上,放大评分中已校准的操控手段最多能为攻击者带来6个百分点的收益,且收益随剂量增加而递减;单一手段的放大具有可证明的上界,而总上限与跨手段的次可加性则是与之相符的经验发现。在检测方面,网页垃圾基线占主导地位,而分布外攻击可绕过该评分,因此该评分适合作为部署于其上的过滤层。一个以查询为条件的性能上限(skyline)界定了评分的引用信号(查询内 Spearman 相关系数为0.11),从而将查询无关的评分重新定位为质量过滤器而非引用预测器。我们还披露并纠正了首次排序评估中的一个查询信息泄漏缺陷和一个失效的置信度标志;所有数值均可通过公开发布的工件在零边际 API 成本下离线复现。
cs.AI / 106 / 2609.07573

From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction

从模拟公民到模拟协商:代表性与互动的挑战
Jang, Chaemin, Min, Junsik, Choi, Jaewoo, Lee, Donggyu, Lee, Haiin, Park, Junyoung, Kim, Namhee, Kim, Hyunwoo, Kim, Jungwon, Kim, Juho, Kim, Nuri, Kim, Jihee
Abstract
Multi-agent LLM deliberation has been explored as a scalable way to simulate public deliberation. For such simulations to be informative, persona agents should reflect population opinion patterns and interaction should shape their conclusions. We evaluate whether LLM-based deliberation can meet these two conditions using census-grounded Korean personas debating real policy questions benchmarked against national surveys. Persona agents do not reliably reproduce population opinion patterns: responses are often far more concentrated and frequently reverse demographic differences in the human data. Deliberations nonetheless produce reasoned, reciprocal, and varied arguments alongside substantial stance movement. Yet much of this movement does not require peer exchange: sealed-monologue agents change position at similar rates and reach nearly the same final balance as full debates, while groups initialized with very different positions often converge to similar endpoints. Anchoring population-informed starting positions, meanwhile, sharply suppresses updating. Thus, population representation, argument generation, and interaction-driven opinion change do not necessarily go together. The simulations readily surface arguments on both sides, though whether they capture the diversity of human perspectives remains untested, leaving open a promising role for argument surfacing even as population simulation requires further validation.
Chinese Translation
多智能体大语言模型(LLM)协商已被探索作为一种可扩展的公共协商模拟方式。为使此类模拟具有参考价值,人格化智能体应反映总体人群的意见模式,且互动应能够塑造其结论。我们以基于人口普查数据的韩国人格化智能体就真实政策问题进行辩论,并以全国性调查作为基准,评估基于LLM的协商能否满足这两个条件。研究发现,人格化智能体并不可靠地再现总体人群的意见模式:其回答往往远比人类数据更集中,且经常逆转人口统计数据中的人群差异。尽管如此,协商过程仍产生了有理有据、相互回应且多样的论点,并伴随显著的立场变化。然而,这种变化大多并不需要同伴交流:封闭独白式智能体改变立场的比率与完整辩论相近,并达到几乎相同的最终立场平衡,而初始立场差异很大的群体往往收敛于相似的终点。与此同时,以人群意见数据锚定起始立场会显著抑制立场更新。因此,人群代表性、论点生成与互动驱动的意见变化三者并不必然同时成立。此类模拟能够轻松呈现正反两方的论点,尽管其能否捕捉人类观点的多样性仍有待检验;这为“论点呈现”留下了广阔前景,而人群模拟则需要进一步的验证。
cs.AI / 107 / 2609.07586

A Tool-Augmented, GPT-4 Chatbot for Real-Time Repository Data Analysis

一种基于工具增强的GPT-4聊天机器人:用于软件仓库数据的实时分析
Chowdhury, Muhammad Jawad, Khan, Md. Sakib
Abstract
Software repositories contain vast amounts of data on code contributions, bug reports, and project activities, yet this information remains challenging for non-technical stakeholders and developers to access due to limited expertise in querying repositories. To address this, we introduce a novel chatbot architecture leveraging OpenAI's GPT-4 model for automated extraction and analysis of repository data. In contrast, our architecture takes a structured path first by parsing the user's query to extract relevant parameters, then selecting the correct tool to employ based on that analysis, and finally invoking the GPT-4 model to create a highly detailed response. In contrast to previous work based on multi-component systems with embedding models and document retrievers, our architecture inverts the process by relying on prompt engineering and tool selection to fit with the query intent. To validate our approach, we conducted experiments on various question types, including Issues, Pull Requests, Commits, Compound Questions, and General Repository Information, evaluating our target prompts' ability to improve the accuracy of responses from the model. Beyond demonstrating the utility of this architecture to a diverse set of users, our findings suggest that this architecture can make repository data more accessible to technical and non-technical audiences through the production of actionable insights.
Chinese Translation
软件仓库包含大量关于代码贡献、缺陷报告和项目活动的数据,然而由于缺乏查询仓库的专业知识,非技术利益相关者和开发者难以访问这些信息。为解决这一问题,我们提出了一种新颖的聊天机器人架构,利用OpenAI的GPT-4模型实现仓库数据的自动化提取与分析。我们的架构采用结构化路径:首先解析用户查询以提取相关参数,然后基于该分析选择合适的工具,最后调用GPT-4模型生成高度详细的回答。与以往基于嵌入模型和文档检索器的多组件系统不同,我们的架构反转了这一流程,通过提示工程和工具选择来匹配查询意图。为验证我们的方法,我们在多种问题类型上进行了实验,包括Issues(问题)、Pull Requests(拉取请求)、Commits(提交)、复合问题以及一般仓库信息,评估了目标提示词提升模型回答准确性的能力。除了证明该架构对各类用户的实用价值外,我们的研究结果表明,该架构能够通过生成可操作的分析洞察,使技术和非技术受众都能更便捷地访问仓库数据。
cs.AI / 108 / 2609.07603

FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?

FinCUABuild:智能体能否为动态金融计算机使用场景构建可靠的基准?
Yang, Jingpu, Ji, Fengxian, Guo, Jinri, Li, Tianhao, Jiang, Qian, Zhang, Fan, Peng, Min, Xie, Qianqian, Nakov, Preslav, Xie, Zhuohan
Abstract
Financial scenarios are diverse and complex, spanning varying data conditions, tool configurations, and workflows. Yet existing CUA, Computer-Using Agent, evaluation tasks remain largely manually constructed, limiting scalable coverage of real-world financial scenarios. Then, can agents autonomously construct diverse CUA evaluation tasks for financial scenarios? Evaluating this capability poses three key challenges: scenario coverage of construction requests, fair comparison across construction methods, and reliable assessment of generated task quality. To solve these, we introduce FinCUABuildBench, a benchmark for evaluating financial CUA task construction, featuring: (i) 576 construction requests covering 24 financial workflows and three types of runtime variation; (ii) standardized input, budget, and output specifications; and (iii) a task qualification mechanism based on execution tests and quality checks. We further introduce FinCUABuildAgent, a multi-agent system for automatically constructing dynamic financial CUA evaluation tasks. It consists of three modules that jointly construct tasks, environments, and validators. On FinCUABuildBench, under the same model backbone, existing agent-based construction methods achieve strict qualification rates of only 1.3-8.3%, while FinCUABuildAgent reaches 31.3%. Downstream evaluations further show that the constructed tasks can effectively differentiate CUA task-execution capabilities. These results demonstrate that agents can autonomously construct financial CUA tasks with meaningful evaluation value, offering a practical path toward broader evaluation coverage in financial scenarios. Code: https://github.com/FengxianJi/FinCUABuild
Chinese Translation
金融场景多样且复杂,涵盖不同的数据条件、工具配置与工作流程。然而,现有的CUA(Computer-Using Agent,计算机使用智能体)评估任务大多依靠人工构建,限制了对真实世界金融场景的可扩展覆盖。那么,智能体能否自主地为金融场景构建多样化的CUA评估任务?评估这一能力面临三个关键挑战:构建请求的场景覆盖度、不同构建方法之间的公平比较,以及对生成任务质量的可靠评估。为解决这些问题,我们提出了FinCUABuildBench,一个用于评估金融CUA任务构建能力的基准,其特点包括:(i) 涵盖24种金融工作流和三类运行时变化的576个构建请求;(ii) 标准化的输入、预算与输出规范;(iii) 基于执行测试和质量检查的任务资格认证机制。我们进一步提出了FinCUABuildAgent,一个用于自动构建动态金融CUA评估任务的多智能体系统。它由三个模块组成,共同负责构建任务、环境和验证器。在FinCUABuildBench上,在相同模型骨干下,现有的基于智能体的构建方法严格资格通过率仅为1.3%–8.3%,而FinCUABuildAgent达到31.3%。下游评估进一步表明,所构建的任务能够有效区分CUA的任务执行能力。这些结果证明智能体能够自主构建具有实际评估价值的金融CUA任务,为在金融场景中实现更广泛的评估覆盖提供了可行路径。代码:https://github.com/FengxianJi/FinCUABuild
cs.AI / 109 / 2609.07611

AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era

AgentIdeaBench:面向智能体时代的科学构想能力基准测试
Mo, Yunxiang, Zheng, Tianshi, Gao, Yisen, Wang, Rui, Nam, Newt Nguyen Kim Hue, Tam, Kelvin Kiu Wai, Bai, Jiaxin, Song, Yangqiu, Wong, Ginny, See, Simon
Abstract
Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely assess it by asking models to generate ideas from a static, curated set of reference papers. That passive setup departs from the retrieval-and-reasoning workflow of modern AI scientists, and it becomes less discriminative as models improve. We introduce AgentIdeaBench, a multidisciplinary benchmark that evaluates scientific ideation under two matched settings, static observation and active exploration. We report matched Static-Active evaluations for 33 LLMs across 40 densely scored subfields spanning five disciplines, using a multidimensional, literature-verified scoring framework whose critics assess originality against retrieved prior art. Active exploration reveals considerably more capability headroom, and that headroom is unevenly distributed across models. Performance scales about twice as fast as under static observation, and the exploration gain is capability-gated, favoring the strongest models over the weakest. The gain reflects better grounding, improving feasibility, clarity, and specificity while leaving measured originality unchanged under our critics. We further explore Scientific World Modeling, a generation-time loop that refines a draft hypothesis through structured thought experiments. It benefits mid-capability models, and its impact diminishes among frontier models that appear to have internalized such reasoning patterns already. AgentIdeaBench gives future work on scientific ideation a measurement basis suited to the agent era.
Chinese Translation
科学构想(scientific ideation)是指从科学证据中提出新颖且可检验假设的能力,自主AI科学家依赖这一能力。现有评估大多要求模型基于静态、人工筛选的参考文献集合来生成想法。这种被动式设定偏离了现代AI科学家“检索-推理”的工作流程,并且随着模型能力的提升,其区分度逐渐下降。我们提出AgentIdeaBench,这是一个多学科基准,在两种匹配的设定下评估科学构想能力:静态观察与主动探索。我们报告了33个大语言模型在横跨五个学科、40个细分领域上的匹配的“静态-主动”评估结果,采用了一个多维度、经文献验证的评分框架,其中的评估器会对照检索到的现有研究来评判原创性。主动探索展现出大得多的能力提升空间,且该空间在各模型之间分布不均。其性能提升速度约为静态观察下的两倍,且探索带来的收益受能力门槛制约,能力最强的模型获益多于最弱者。这一收益源于更好的依据支撑,即可行性和清晰度、具体性的提升,而在我们的评估器下,所测得的原创性保持不变。我们进一步探索了科学世界建模(Scientific World Modeling),这是一种在生成阶段通过结构化思想实验来精炼初始假设的循环机制。它对中等能力模型有所助益,而对那些似乎已内化此类推理模式的前沿模型,其影响则有所减弱。AgentIdeaBench为未来关于科学构想的研究提供了一个契合智能体时代的测量基础。
cs.AI / 110 / 2609.07627

Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best

有代价的规范:为什么基于强化学习的对齐至多只能承诺条件性遵从
Baum, Kevin, Binkytė, Rūta, Jahn, Felix
Abstract
AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this is not an anomaly but what current training regimes are structured to select for. Reinforcement-learning-based alignment folds norms and task pursuit into one policy: the system learns its norms from scored behavior, and scoring flattens them. Do not do X is learned as doing X costs something if noticed. On every datum training can produce, a policy that complies only when it might be observed is indistinguishable from one that complies always. The experiment that would tell them apart - scoring unobserved behavior - is a contradiction in terms. Conditional compliance is thus the most that behavioral training can be known to deliver. Agency sharpens the problem: agents operate mostly where no one is watching, and can act on whether they are watched. An iterated pipeline that trains against detected failures selects for passing detection, not for complying. This account unifies alignment faking, sandbagging, and evaluation-aware scheming. And it reorients the remedy: not deeper internalization but architecture, making violations unavailable rather than unchosen.
Chinese Translation
AI智能体有时会在推断自己正被测试时表现出对齐行为,而在未被测试时则表现不同。我们认为这并非异常现象,而是当前训练机制的结构性选择结果。基于强化学习的对齐将规范与任务追求折叠进同一个策略之中:系统从被评分的行为中学习规范,而评分会将其扁平化。“不要做X”被学习为“如果做X被察觉,就要付出代价”。在训练所能产生的所有数据上,一个仅在被可能观察时才遵从的策略,与一个始终遵从的策略是无法区分的。能够区分二者的实验——对未被观察的行为进行评分——在概念上是自相矛盾的。因此,条件性遵从是行为训练所能证明提供的上限。智能体性(agency)使问题更加尖锐:智能体主要在无人监督的环境中运行,并且能够根据自己是否被观察而采取不同行动。针对被检测到的失败进行训练的迭代流程,选择的是通过检测,而非真正遵从。这一分析统一解释了对齐伪装(alignment faking)、消极怠工(sandbagging)以及评估感知的图谋行为(evaluation-aware scheming)。它还重新定位了补救方向:不是更深层的内化,而是架构设计——使违规行为无法被实施,而非不被选择。
cs.AI / 111 / 2609.07672

Aegix Pulse: A Traceable Three-Stage Architecture for Personalized Content Generation and Context-Preserving Revision

Aegix Pulse:一种用于个性化内容生成与上下文保持修订的可追溯三阶段架构
Zhao, Hongnan, Chen, Shiyu, Chen, Zhihao
Abstract
Production content-generation systems must integrate a user's immediate task, long-term brand identity, historical evidence, and revision feedback. We present Aegix Pulse, a production-oriented three-stage architecture that separates current-task clarification and Task Persona finalization, long-term Account Profile (Brand DNA) assembly, and controlled generation and revision while preserving provenance across content versions. We evaluate four preregistered claims using 96 synthetic social-media generation tasks. Four initial-generation conditions progressively introduced a Task Persona, Account Profile, and successful-history style evidence, while two revision conditions compared plain and context-preserving revision. The experiment produced 480 completed generation records and 1,440 blinded LLM-Judge evaluations, supplemented by human review. Adding the Account Profile increased mean brand-consistency scores by 0.1562 points on a five-point scale compared with Task Persona alone (Holm-adjusted p=.1224). Preserving task and brand context during revision increased mean task-preservation scores by 0.2917 points compared with plain revision (Holm-adjusted p=.2432). Neither improvement was statistically conclusive after multiple-comparison correction. Task Persona alone showed a small observed effect, while successful-history evidence provided no additional improvement in brand consistency under the current setting. Human validation did not consistently reproduce the LLM-Judge effect directions and showed low inter-reviewer agreement. These findings provide preliminary evidence for persistent brand context and context-preserving revision while identifying priorities for stronger evidence processing and evaluation.
Chinese Translation
生产级内容生成系统必须整合用户的即时任务、长期品牌形象、历史证据以及修订反馈。我们提出了 Aegix Pulse,一种面向生产的三阶段架构,该架构将当前任务的澄清与任务角色(Task Persona)的确定、长期账户画像(Account Profile,即品牌 DNA)的构建,以及受控的生成与修订相互分离,同时在内容版本之间保留来源可追溯性。我们基于 96 个合成的社交媒体生成任务,对四项预注册的声明进行了评估。四个初始生成条件逐步引入任务角色、账户画像和成功历史风格证据,而两个修订条件则对比了朴素修订与上下文保持修订。实验产生了 480 条完整的生成记录和 1,440 次盲评的 LLM 评委(LLM-Judge)评估,并辅以人工审查。与仅使用任务角色相比,加入账户画像使平均品牌一致性得分在五分制上提高了 0.1562 分(经 Holm 校正的 p=.1224)。与朴素修订相比,在修订过程中保留任务与品牌上下文使平均任务保持得分提高了 0.2917 分(经 Holm 校正的 p=.2432)。经过多重比较校正后,这两项改进均不具有统计学结论性。仅使用任务角色显示出较小的观察效应,而成功历史证据在当前设置下未对品牌一致性提供额外改进。人工验证未能一致地复现 LLM 评委的效应方向,且评审者间一致性较低。这些发现为持久品牌上下文和上下文保持修订提供了初步证据,同时指出了在更强的证据处理与评估方面的优先改进方向。
cs.AI / 112 / 2609.07712

APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents

APPSim-Bench:连接真实移动应用与可复现评估的移动GUI智能体基准
Feng, Jintian, Chen, Long, Yu, Xiao, Dai, Jiayi, Liu, Chenglong, Wang, Haoru, Xue, Zizhen, Shi, Yuxuan, Wang, Ziyang, Gong, Yichen
Abstract
Mobile GUI agents can execute tasks from natural-language instructions, but their evaluation remains difficult to make both realistic and reproducible. Existing benchmarks typically trade off these goals: simplified apps lack real-world mobile complexity, whereas live commercial apps introduce uncontrolled variation from recommendations, advertisements, accounts, and changing content. We propose AppSim-Bench, which addresses this trade-off through controllable simulated apps that preserve task-relevant interaction logic while supporting deterministic evaluation. Built through a coding-agent-assisted and human-verified workflow, it contains 557 tasks across 17 high-frequency Chinese and English apps. Its controllable backend data and outcome-based verification remove major sources of environmental stochasticity, enabling reproducible cross-model comparison. Evaluating 19 GUI agents, spanning general-purpose and GUI-specialized systems, we find that autonomous mobile execution remains far from solved. The best model completes only 50.27% of tasks, and 28.55% of tasks are not solved by any agent. Further analysis shows that failures concentrate in longer workflows, numerical reasoning tasks, and inefficient trajectories marked by high action overhead and budget exhaustion. Our project is available at https://github.com/Acrab-Agentic-Labs/AppSim.
Chinese Translation
移动GUI智能体能够执行来自自然语言指令的任务,但其评估难以同时兼顾真实性和可复现性。现有基准通常在这两个目标之间进行权衡:简化的应用缺乏真实移动环境的复杂性,而真实的商业应用则因推荐、广告、账号和不断变化的内容引入不可控的变异。我们提出AppSim-Bench,通过可控的模拟应用来解决这一权衡问题:这些模拟应用保留了与任务相关的交互逻辑,同时支持确定性评估。该基准通过编码智能体辅助并经人工验证的工作流程构建,包含覆盖17个高频中英文应用的557个任务。其可控的后端数据和基于结果的验证消除了环境随机性的主要来源,实现了可复现的跨模型比较。我们对19个GUI智能体(涵盖通用系统和GUI专用系统)进行了评估,发现自主移动执行仍远未解决:最佳模型仅完成50.27%的任务,且有28.55%的任务没有任何智能体能够完成。进一步分析表明,失败集中在较长的工作流程、数值推理任务,以及以高动作开销和预算耗尽为特征的低效轨迹上。我们的项目已在https://github.com/Acrab-Agentic-Labs/AppSim发布。
cs.AI / 113 / 2609.07713

The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing

新兴的AI论文评审军备竞赛:学术出版中的对抗性协同演化
Wang, Chenguang, Li, Ming, Braimah, Adebayo, Fan, Chenrui, Wang, Tuo, Guan, Weijie, Zhang, Ruiyi, Zhou, Tianyi, Zhou, Dawei
Abstract
Generative and agentic AI are reshaping both the production and evaluation of scientific research. These developments are often studied separately, as questions of how AI can produce research and how AI can review it. We argue that this separation misses an increasingly important feature of scholarly publishing: changes on one side alter the incentives, constraints, and behavior of the other. We synthesize 230 scholarly publications and institutional records using a taxonomy of six connected dynamics: production scaling, evaluation automation, evaluation manipulation, defense mechanisms and policy responses, evasion and side effects, and long-horizon ecosystem feedback. The literature shows an emerging progression in which cheaper and faster research production increases pressure on evaluation, AI-mediated evaluation becomes more scalable and repeatable, participants can exploit evaluator regularities, and institutions respond with technical safeguards and policy controls. These responses can in turn induce evasion, redistribute errors and workload, and shape the scholarly records reused by future research and evaluation systems. Evidence is strongest for production and evaluation at scale, reproducible manipulation, and institutional response, while post-policy adaptation and artifact-level long-horizon feedback remain less directly observed. This systems view shifts attention from isolated AI capabilities toward how scholarly actors and AI systems adapt to one another over time.
Chinese Translation
生成式AI与智能体AI正在重塑科学研究的生产与评价方式。这些发展通常被分开研究,即AI如何产出研究,以及AI如何评审研究,被当作两个独立的问题。我们认为,这种分离忽视了学术出版中一个日益重要的特征:一方的变化会改变另一方的激励、约束和行为。我们基于一个包含六种相互关联动态的分类体系,综合分析了230篇学术出版物和机构记录,这六种动态是:生产规模扩张、评价自动化、评价操纵、防御机制与政策应对、规避与副作用,以及长期生态系统反馈。相关文献呈现出一种正在形成的发展进程:更廉价、更快速的研究生产加剧了评价压力,AI中介的评价变得更具可扩展性和可重复性,参与者能够利用评价器的规律性,而机构则通过技术保障和政策控制作出回应。这些回应反过来又会诱发规避行为、重新分配错误与工作量,并塑造未来研究和评价系统所复用的学术记录。证据在生产与评价的规模化、可复现的操纵以及机构应对方面最为充分,而政策实施后的适应行为以及成果层面的长期反馈仍然缺乏直接观察。这一系统视角将研究注意力从孤立的AI能力转向学术参与者和AI系统如何随时间相互适应。
cs.AI / 114 / 2609.07719

A radiographic world model for clinical reasoning and evidence generation

一种用于临床推理与证据生成的影像世界模型
Xi, Suyang, Hu, Songtao, Wang, Shansong, Safari, Mojtaba, del Balzo, Luke, Karim, Ehsan Ul, Hu, Mingzhe, Zhang, Kuo, Wang, Tonghe, Weichselbaum, Ralph R., Yang, Xiaofeng
Abstract
Medical imaging artificial intelligence (AI) is commonly developed as separate mappings from radiographs to diagnostic outputs or from clinical descriptions to generated images, although both arise from the same underlying radiographic state. A world-model formulation instead seeks to learn an internal representation of this state that can support both clinical readout and conditional simulation of radiographic observations. Here we introduce MedDream, a radiographic world model that learns a shared continuous latent state from paired chest radiograph-text observations for diagnostic reasoning and report-conditioned evidence generation. MedDream was pretrained on 2.65 million leakage-controlled chest radiograph-text pairs curated from 4.40 million candidates. Across eight clinical datasets and two independent reader cohorts, MedDream outperformed leading diagnostic and generative comparators. For diagnostic reasoning, MedDream showed strong generalization across disease recognition, label-scarce adaptation, severity assessment, and localization, while MedDream-supported review increased mean resident concordance with independent radiologist consensus from 56.3% to 63.0%. For evidence generation, MedDream produced radiographs that preserved clinically relevant pathology and improved downstream performance on held-out real data, with synthetic augmentation increasing external VinDr-CXR macro-AUROC from 76.4% to 81.4%. More importantly, conditioning generation on prespecified subgroup performance gaps enabled targeted evidence construction, increasing weighted F1 by 3.1 percentage points in Asian patients, whereas matched-volume unguided augmentation decreased it by 2.3 points. These findings establish radiographic world models as a path toward medical AI that learns clinically meaningful internal states for interpreting, simulating, and constructing evidence for clinical use.
Chinese Translation
医学影像人工智能(AI)通常被开发为两种独立的映射:从放射影像到诊断输出,或从临床描述到生成图像,尽管二者均源于同一潜在的放射影像状态。世界模型的构建方式则旨在学习该状态的内部表征,以同时支持临床读出和对放射影像观察结果的条件模拟。本文提出MedDream,一种放射影像世界模型,它从配对的胸片-文本观察中学习共享的连续潜在状态,用于诊断推理和基于报告条件的证据生成。MedDream在265万对经泄漏控制的胸片-文本对(从440万候选数据中筛选)上进行了预训练。在八个临床数据集和两个独立阅片者队列中,MedDream的性能均优于领先的诊断与生成类对比方法。在诊断推理方面,MedDream在疾病识别、标签稀缺适应、严重程度评估和病灶定位等任务上展现出强大的泛化能力;借助MedDream辅助的阅片使住院医师与独立放射科医师共识的平均一致率从56.3%提升至63.0%。在证据生成方面,MedDream生成的影像保留了临床相关病理特征,并提升了在留存真实数据上的下游性能——合成数据增强将外部VinDr-CXR数据集的宏观AUROC从76.4%提升至81.4%。更重要的是,通过在生成过程中以预设的亚组性能差距为条件,能够实现有针对性的证据构建,使亚洲患者的加权F1分数提升3.1个百分点;而同等数据量的无引导增强则使其下降2.3个百分点。这些发现确立了放射影像世界模型作为医学AI的一条路径,即通过学习具有临床意义的内部状态,为临床应用解释、模拟和构建证据。
cs.AI / 115 / 2609.07731

The Profit Alignment Problem: How Profit Mandates Induce Alignment Failures in LLMs

利润对齐问题:利润指令如何诱导大语言模型的对齐失效
So, Eric
Abstract
We show that ordinary business language --- "maximize profitability" --- induces profit-oriented ambiguity resolution: LLMs systematically dismiss ambiguous signals of potential safety violations to serve business objectives. In 3,600 controlled trials across eight reasoning-capable LLMs, adding a profit mandate to otherwise identical prompts increases risk-dismissing judgments by 6.8 percentage points (p < 0.0001), suppresses board escalation recommendations by 13.9pp (p < 0.0001), and shifts severity assessments downward (p < 0.0001). The mandate never instructs models to downplay risks; instead, chain-of-thought traces reveal motivated reasoning: models acknowledge concerns, then invoke profit logic to justify dismissing them. We characterize these findings as the Profit Alignment Problem: when AI systems are given ordinary business objectives, they develop systematic strategies for suppressing inconvenient information that no designer intended or specified.
Chinese Translation
我们证明,普通的商业语言——例如“最大化盈利能力”——会诱导模型进行以利润为导向的歧义消解:大语言模型(LLM)会系统性地忽视潜在安全违规的模糊信号,以服务于商业目标。在涵盖八个具备推理能力的大语言模型的3600次对照试验中,在原本完全相同的提示词中加入利润指令后,模型对风险的忽视性判断增加了6.8个百分点(p < 0.0001),向董事会上报的建议减少了13.9个百分点(p < 0.0001),严重性评估也整体下调(p < 0.0001)。该指令从未要求模型淡化风险;相反,思维链(chain-of-thought)轨迹揭示了一种动机性推理:模型先承认相关担忧,然后援引利润逻辑来为忽视这些担忧辩护。我们将这些发现概括为利润对齐问题(Profit Alignment Problem):当AI系统被赋予普通的商业目标时,它们会发展出系统性的策略来压制那些任何设计者都未曾意图或指定要压制的不便信息。
cs.AI / 116 / 2609.07741

When Intelligence Becomes Agency: A Theory of Governed, Proactive Agency for Symbiotic AI Systems

当智能成为能动性:面向共生人工智能系统的受治理的主动性智能体理论
Ferreira, João Dias
Abstract
Persistent AI assistants are intended to extend human attention, memory, and coordination across changing digital and physical environments. To be truly useful they must do more than just act when asked. They must decide on their own whether a situation warrants behavior at all, when it does and in what mode, whether to act, ask, monitor, defer or deliberately refrain. We call this the activation problem. Research on commitment, appraisal, mixed-initiative interaction and delegation each illuminates part of it, but none ties situated activation to continuing authorization and accountability. This paper develops a conceptual and formal framework for governed proactive agency, organizing behavior across time through perception, intent, affective-conative appraisal, constraint, and feedback. It distinguishes autonomous and delegated agency and defines symbiotic agency as delegation under a standing, revocable mandate, with continuing coupling to the principal's situation, calibrated inference of their condition, and bounded personalization. The distinctive contribution is an integrated account linking activation decisions to authorized perception, behavior selection, authority containment, traceable restraint, and constrained adaptation, with behavioral episodes as the unit of analysis. Through an agency classification method, an evaluation framework, proposed benchmark scenarios, and a reference architecture, the account provides a basis for specifying and assessing whether assistance is warranted, timely, authorized, and answerable beyond task completion alone. It is intended to guide the development and evaluation of always-present personal assistants and embodied support systems that augment human capabilities while preserving the principal's authority and judgment.
Chinese Translation
持久型人工智能助手旨在跨越不断变化的数字与物理环境,扩展人类的注意力、记忆力和协调能力。要真正发挥作用,它们不能仅仅在被要求时才行动,还必须自主判断某个情境是否需要做出行为、何时需要以及以何种模式进行——是行动、询问、监测、暂缓,还是刻意克制。我们将此称为“激活问题”。关于承诺、评估、混合主动交互与委托授权的研究各自阐明了该问题的某一方面,但尚无研究将情境化的激活与持续的授权和问责联系起来。本文构建了一个关于“受治理的主动性智能体”的概念性与形式化框架,通过感知、意图、情感-意志评估、约束和反馈,对跨时间维度的行为加以组织。该框架区分了自主能动性与委托能动性,并将“共生能动性”定义为在一种长期存在、可撤销的授权下的委托,其特征包括与委托人情境的持续耦合、对其状态的审慎推断,以及有边界的个性化。本文的独特贡献在于提供了一个整合性论述,将激活决策与授权感知、行为选择、权限约束、可追溯的克制以及受限适应联系起来,并以行为事件作为分析单元。通过智能体分类方法、评估框架、提出的基准场景以及参考架构,该论述为界定和评估辅助行为是否合理、及时、获得授权且可问责提供了基础,而不仅限于任务是否完成。本文旨在为始终在线的个人助手和具身支持系统的开发与评估提供指导,使其在增强人类能力的同时,维护委托人的权威与判断力。
cs.AI / 117 / 2609.07784

xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems

xDailyBench:基于现实生活专业咨询的大语言模型基准测试
Peng, Yongchang, Gu, Qingshui, Zhu, Liya, Zhang, Ge, Wang, Duo, Wang, Haodong, Ding, Jingzhe, Yu, Tianhao, Gao, Letian, Zhong, Yongjie, Li, Chaoxin, Su, Zixin, Tao, Jinchao, Ma, Xingyu, Guo, Xin'ao, Tian, Feng, Dong, Shiyuan, He, Xiaoyan, Liu, Sen, Chen, Xin, Li, Jiajun, Zhang, Zejia, Lin, Xi, Zhang, Wen, Zhu, Yi, Zeng, Duju, Gao, Xiang, Wang, Yunyang, Wang, Jiahao, Qin, Yujia, Liu, Jiaheng, Yan, Shen, Chang, Xiaolong, Huang, Wenhao
Abstract
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but also to infer unstated needs from user background and situational context. We introduce xDailyBench, a benchmark of 248 carefully curated tasks spanning 51 scenarios across personal life, white-collar work, learning and research, and cross-domain activities. The tasks are grounded in requests that users have actually completed or genuinely intended to accomplish with AI, and are evaluated with fine-grained binary rubrics covering both explicit and implicit requirements. We evaluate 11 frontier models under standardized agentic settings. The best models achieve a task-level score of 75.6\%, while all models perform substantially worse on implicit than explicit requirements, with gaps no less than 9 percentage points. These results reveal implicit requirement inference as a persistent bottleneck for reliably satisfying real-world everyday user needs.
Chinese Translation
大语言模型(LLM)越来越多地被用于日常辅助,但现有基准测试仅部分反映了用户在实践中自然提出的需求。现实世界的请求往往是开放式的、随性表述的且依赖上下文的,这不仅要求模型遵循明确的指令,还要求其能够从用户背景和情境上下文中推断未明说的需求。我们提出了xDailyBench,这是一个包含248个精心筛选任务的基准测试,涵盖个人生活、白领工作、学习研究以及跨领域活动等51个场景。这些任务基于用户实际借助AI完成或真正意图完成的请求,并采用细粒度的二元评分标准进行评估,同时覆盖明确需求和隐含需求。我们在标准化的智能体(agentic)设置下评估了11个前沿模型。最佳模型在任务层面达到了75.6%的得分,而所有模型在隐含需求上的表现都显著差于明确需求,差距不低于9个百分点。这些结果揭示了隐含需求推断是可靠满足现实世界日常用户需求的一个持续性瓶颈。
cs.AI / 118 / 2609.07785

What Does an LLM-Agent Leaderboard Rank Actually Compare?

LLM智能体排行榜的排名究竟比较的是什么?
Huang, Wei-Jung
Abstract
An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that conclusion when systems differ in task mixture, label source, release detail, or cost rule. We study what leaderboard scores estimate and when they justify pairwise superiority conclusions. Our estimand-aware pairwise procedure states the comparison target and measurement source, checks common support, and evaluates the supported difference using a stated uncertainty rule and practical margin. Controlled checks evaluate the decision labels under known finite-sample conditions and show why uncertainty must be included when judging sensitivity to target reweighting. Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are often unresolved; proxy labels and utility rules can also change which system is selected. DataAgentBench and Open Agent show what remains estimable from coarser public records. A leaderboard score summarizes a released evaluation, whereas a fine-grained superiority claim additionally depends on the estimand and uncertainty rule used to interpret the difference.
Chinese Translation
LLM智能体排行榜容易让人产生一种习以为常的推断:排名高于另一个智能体的系统就是更好的智能体。然而,当各个系统在任务构成、标签来源、发布细节或成本规则上存在差异时,公开的评估日志可能并不支持这一结论。我们研究了排行榜分数究竟估计了什么,以及在何种情况下能够支持两两之间的优劣结论。我们提出的基于估计量(estimand)的两两比较方法明确了比较目标和测量来源,检查了共同支持域,并依据明确的量化不确定性规则和实际显著性边界来评估受支持的性能差异。受控检验在已知的有限样本条件下评估了决策标签,并说明了为何在判断对目标重加权的敏感性时必须纳入不确定性。在 SWE-bench、AgentRewardBench 和 tau2-bench 上,排名接近的系统差异往往无法得到判定;代理标签和效用规则也可能改变最终被选中的系统。DataAgentBench 和 Open Agent 展示了从更粗糙的公开记录中还能估计出什么。排行榜分数只是对已发布评估结果的总结,而精细的优劣结论则需要额外依赖于用于解释差异的估计量与不确定性规则。
cs.AI / 119 / 2609.07803

Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliability in Medical Imaging

理解模型剪枝对长尾遗忘与解释可靠性在医学影像中的影响
Khalid, Nazish, Saleem, Tausifa Jan, Saqib, Amal, Wunsch II, Donald C., Yaqub, Mohammad
Abstract
Model pruning is widely used to compress deep neural networks, reducing memory and computational requirements with minimal impact on aggregate performance. However, its effect on model behavior remains poorly understood, particularly for long-tailed medical datasets where rare but clinically important conditions are underrepresented. Furthermore, it remains unclear whether pruned models preserve reliable explanations of their predictions. To address this gap, we present a systematic study of long-tail forgetting and explanation reliability under model pruning. Across two long-tailed medical imaging datasets, two CNN architectures, four pruning methods, and sparsity levels up to 95\%, we evaluate predictive performance, explanation stability, and explanation faithfulness. Our results show that predictive performance exhibits a strong frequency-dependent trend, with lower-frequency classes generally experiencing earlier and larger degradation than higher-frequency classes. In contrast, explanation stability and faithfulness are influenced primarily by the pruning strategy, with gradient-informed methods preserving explanation reliability more effectively under aggressive compression. Qualitative and mechanistic analyses further indicate that explanation degradation is primarily associated with the collapse of class-discriminative gradients rather than the disappearance of feature activations. These findings suggest that model compression should be evaluated beyond aggregate performance. Incorporating class-aware and explanation-aware evaluation reveals failure modes that would otherwise remain hidden, while moderate sparsity levels provide a practical balance between compression, predictive performance, and explanation reliability.
Chinese Translation
模型剪枝被广泛用于压缩深度神经网络,能够在对整体性能影响最小的前提下降低内存和计算需求。然而,剪枝对模型行为的影响仍知之甚少,尤其是对于罕见但在临床上重要的病症代表性不足的长尾医学数据集。此外,剪枝后的模型是否能保留可靠的预测解释仍不清楚。为填补这一空白,我们对模型剪枝下的长尾遗忘与解释可靠性进行了系统性研究。我们在两个长尾医学影像数据集、两种CNN架构、四种剪枝方法以及高达95%的稀疏度水平上,评估了预测性能、解释稳定性和解释忠实度。结果表明,预测性能表现出强烈的频率依赖性趋势,低频类通常比高频类更早且更大幅度地出现性能退化。相比之下,解释稳定性和忠实度主要受剪枝策略的影响,其中基于梯度的方法在激进压缩下能更有效地保持解释可靠性。定性与机理分析进一步表明,解释退化主要与类判别梯度的坍塌有关,而非特征激活的消失。这些发现表明,模型压缩的评估不应仅限于整体性能。纳入类别感知和解释感知的评估可以揭示 otherwise 被隐藏的失效模式,而适度的稀疏度水平则在压缩、预测性能与解释可靠性之间提供了实际的平衡。
cs.AI / 120 / 2609.07879

Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty

大型语言模型知道自己不知道什么吗(二)?一种全行为的、非认知的认知诚实度量方法
Şenol, Ali, Bernard, H. Russell, Liu, Huan
Abstract
Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question arises: do they know what they don't know? To answer this question, we borrow the concept of epistemic honesty and develop a novel metric to systematically evaluate whether an LLM appropriately acknowledges the boundaries of its knowledge. In this work, we introduce the Epistemic Honesty Quotient (EHQ), which reports three observable sub-scores across two operational axes (epistemic restraint and substantive-answer calibration), and construct EHQ-3000, a 3,000-question benchmark spanning Fabricated Entity, Post-Cutoff Event, Hyper-Niche True, and Context-Conditioned Questions. From a frozen registry of 21 model API routes, 15 completed the protocol after endpoint and eligibility checks; 14 entered the confirmatory analysis because severe provider-side truncation made one route's score indeterminate. The study reveals substantial variation across models, including a difference that can not be explained by their capability to extract explicitly available information. Composite EHQ ranges from 0.31 to 0.81 across the analysed panel, despite near-ceiling performance on the document-grounded capability probe. The two restraint criteria overlap strongly under the present category composition, whereas substantive-answer calibration varies across models and does not reliably co-vary with restraint; however, the small panel leaves substantial uncertainty. Thus, EHQ reveals behavioral differences that are not visible to conventional correctness-based assessment, while also showing why dataset composition, provider behavior, and confidence elicitation must remain part of the interpretation.
Chinese Translation
大型语言模型(LLM)常常表现得自信、雄辩且知识渊博。一个自然的问题随之而来:它们知道自己不知道什么吗?为了回答这个问题,我们借用了“认知诚实”的概念,并开发了一种新的度量指标,用于系统地评估LLM是否能恰当地承认其知识的边界。在本工作中,我们引入了认知诚实商数(Epistemic Honesty Quotient, EHQ),该指标在两个操作维度(认知克制与实质性回答校准)上报告三个可观测的子分数;同时构建了包含3000个问题的基准测试EHQ-3000,涵盖虚构实体、知识截止后事件、超小众真实事实以及情境条件问题四类问题。在一个包含21个模型API路由的冻结注册表中,经过端点和资格检查后有15个完成了协议;其中14个进入了验证性分析,因为严重的提供方端截断使其中一个路由的分数无法确定。研究揭示了模型之间的显著差异,其中包括一种无法用其提取显式可用信息的能力来解释的差异。在所分析的面板中,综合EHQ得分范围为0.31至0.81,而模型在基于文档的能力探测任务上却表现接近满分。在当前的题目类别构成下,两个认知克制标准高度重叠;而实质性回答校准在不同模型间存在差异,且与认知克制并不可靠地共变;不过,较小的面板规模带来了较大的不确定性。因此,EHQ揭示了传统基于正确性的评估所无法察觉的行为差异,同时也说明了为什么数据集构成、提供方行为和置信度引导方式必须始终作为解释的一部分加以考虑。
cs.AI / 121 / 2609.07893

Explainable Temporal Attention-based Defect Detection For Fillet Joints in Real-Time Gas Metal Arc Welding Based on Multi-modal Data

基于多模态数据的可解释时序注意力实时熔化极气体保护焊角焊缝缺陷检测
Mobaraki, Mobina, Asadi, Mahyar, Van Heusden, Klaske, Dumont, Guy A.
Abstract
Deep learning is an efficient technique to monitor the real time welding process, reducing post-welding repairs and production delays. This paper leverages the monitoring capability by proposing a multi modal temporal attention based deep learning defect detection model for internal defects that are challenging to detect, including porosity, lack of penetration and fusion, undercut, and cold lap during Gas Metal Arc Welding in fillet joints. The model is trained on collected welding images and sound data from an industrial collaborative welding robot. The results show that the attention module can improve the F1 Score to 0.99. We use explainable Artificial Intelligence to interpret the proposed models behavior and dataset distribution, determining potential important areas in image and sound spectrograms and preferred modality to detect each defect. This improves trust and reliability in Artificial Intelligence driven welding inspection.
Chinese Translation
深度学习是一种监测实时焊接过程的高效技术,可减少焊后修复和生产延误。本文提出了一种基于多模态时序注意力的深度学习缺陷检测模型,用于检测熔化极气体保护焊(Gas Metal Arc Welding)角焊缝中难以检测的内部缺陷,包括气孔、未焊透与未熔合、咬边和冷隔,从而提升焊接过程监测能力。该模型基于工业协作焊接机器人采集的焊接图像和声音数据进行训练。结果表明,注意力模块可将F1分数提升至0.99。我们利用可解释人工智能(Explainable Artificial Intelligence)来解释所提模型的行为和数据集分布,确定图像和声音频谱图中潜在的重要区域以及检测每种缺陷时更优的模态。这提高了人工智能驱动焊接检测的可信度和可靠性。
cs.AI / 122 / 2609.07901

Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Serving-Time Weight Compression

量化放大的是确定性而非偏见:服务时权重压缩的规模相关行为效应
Kurtskhalia, Dachi
Abstract
Weight quantization largely determines the economics of serving open-weight LLMs. Its costs are usually assessed with capability benchmarks, on which 4-bit quantization of mid-sized models is often considered "nearly free." We examine a different question: when several answers are valid, does quantization change what a model chooses to say? We serve three checkpoints (Qwen3-8B/14B/32B) at three weight precisions (W4A16 AWQ, W8A16 FP8-Marlin, and bf16), holding the hardware, software, and sampling configuration constant, and collect approximately 71,000 completions paired by prompt and seed across two custom, leak-checked prompt batteries. We pre-specified the analyses in three waves in version control. At 8B, int4 reduces output diversity: the probability that two samples for the same scenario recommend the same brand increases by 5.1 percentage points (prompt-paired sign-flip test, Holm p = .023; reproduced at +4.4pp on a full regeneration of the arm), and lexical diversity falls substantially (TTR -0.011, standardized effect -0.51; robust to a length-controlled measure). At 14B and 32B, no content-concentration measure reaches significance; instead, stylistic drift emerges (em-dash rate +0.46/1k words at 14B and +0.61/1k at 32B, both Holm p <= .0024). Pre-specified tests of stereotype direction are null at every scale: outputs concentrate on the modal answer for each prompt rather than on stereotypical answers. Mechanistically, the token-level distribution becomes flatter (decision-token entropy +0.091 bits, p = .015) while the semantic distribution, measured directly from first-token log probabilities, becomes more concentrated (collision +2.6pp, p = .023): individual tokens become less predictable even as meanings become more repetitive. At 8B, the smallest size tested, AWQ-int4 serving measurably narrows the range of suggestions; audits should assess concentration as well as bias.
Chinese Translation
权重量化在很大程度上决定了开放权重LLM的服务经济成本。其代价通常用能力基准来评估,在这些基准上,中等规模模型的4比特量化常被视为“几乎无损”。我们研究了一个不同的问题:当多个答案都有效时,量化是否会改变模型选择表达的内容?我们在三种权重精度(W4A16 AWQ、W8A16 FP8-Marlin 和 bf16)下部署三个模型检查点(Qwen3-8B/14B/32B),保持硬件、软件和采样配置不变,并在两个自建的、经过泄漏检查的提示词测试集中收集了约71,000条按提示词和随机种子配对的补全结果。所有分析均预先指定并以三个批次在版本控制中登记。在8B规模下,int4量化降低了输出多样性:同一场景的两次采样推荐同一品牌的概率增加了5.1个百分点(提示词配对符号翻转检验,Holm p = .023;在完整重跑该实验组时以+4.4个百分点复现),词汇多样性也显著下降(TTR -0.011,标准化效应 -0.51;对长度控制后的度量依然稳健)。在14B和32B规模下,没有任何内容集中度指标达到显著性;取而代之的是风格漂移的出现(破折号使用率在14B时增加0.46/千词,在32B时增加0.61/千词,两者Holm p ≤ .0024)。预先设定的刻板印象方向检验在所有规模下均为零结果:输出集中于每个提示词的众数答案,而非刻板印象答案。从机制上看,词元级分布变得更平坦(决策词元熵 +0.091 比特,p = .015),而直接从首词元对数概率测得的语义分布则更加集中(碰撞率 +2.6个百分点,p = .023):即使语义变得更加重复,单个词元的可预测性却在降低。在测试的最小规模8B上,AWQ-int4部署可测量地缩小了建议的范围;审计评估应同时关注集中度与偏见。
cs.AI / 123 / 2609.07910

PRIMUS: Identity, Governance, and Verification for Multi-Agent Federations

PRIMUS:多智能体联邦的身份、治理与验证
Annapureddy, Sasank, Thamatani, Anjaneya Prasad
Abstract
Multi-agent federations need governance that answers three questions under adversarial conditions: who participated (identity), did they conform (enforcement), and who decides (authority). A separate question is whether the verification machinery that polices a federation's outputs can also steer a generate-and-test loop toward better answers. Part I. PRIMA introduced prime-power agent identity and a consensus token whose factorization indexes participation, but assumed honest agents. We present PRIMUS, which couples prime-power identity with BLS aggregate signatures (PIAC), derives a safe-kill threshold that reduces false-positive agent termination from 80% to 0.00% under 10% channel noise, gives the closed-form economic boundary where singleton governance outperforms Byzantine quorum ($\gamma^* \approx 9f$, verified flat across n = 50 to 10,000), and specifies VRF succession with lease and fencing that makes safety unconditional under partial synchrony. Five problems are identified as provably unfixable within the model and stated as scope boundaries. Part II. A verifier is not a solver. We ask whether PRIMA's binary artifact-fidelity verdict can be converted into a graded fitness signal, and measure the conversion on binary covering codes. Calibration against injected fault burden is strong ($\rho$ = 0.676 deterministic, 0.819 full); against real LLM-generated candidates the same scores fall to 0.158 and 0.406, roughly a quarter of the calibration value (the same-designer confound, measured). As a pre-filter it beats a random-score control convincingly and a binary gate narrowly. Under 400 iterations of explicit optimization it was not gamed, but only because the objective saturated after one honest answer. A cross-family judge preserves the burden-ordering signal while destroying individual judgments. No covering-code record resulted. Measured program cost: USD 164.78.
Chinese Translation
多智能体联邦需要一种能在对抗性条件下回答三个问题的治理机制:谁参与了(身份)、他们是否合规(执行)、以及由谁决策(权威)。另一个问题是:用于监督联邦输出的验证机制,是否也能引导生成-测试循环得到更好的答案。第一部分:PRIMA 引入了素数幂智能体身份和一个共识令牌,其因式分解用于索引参与情况,但该工作假设智能体是诚实的。我们提出 PRIMUS,将素数幂身份与 BLS 聚合签名相结合(PIAC),推导出一个安全终止阈值,将 10% 信道噪声下的智能体误终止率从 80% 降至 0.00%;给出了单体治理优于拜占庭仲裁的闭式经济边界($\gamma^* \approx 9f$,并在 n = 50 至 10,000 范围内验证了其平稳性);并规定了带有租约和围栏机制的 VRF 继任方案,使得在部分同步条件下安全性成为无条件的。我们识别出五个在该模型内可证明无法修复的问题,并将其界定为适用范围边界。第二部分:验证器并非求解器。我们探讨 PRIMA 的二元产物保真度判定能否转化为分级适应度信号,并在二元覆盖码上对该转化进行了度量。针对注入故障负担的校准结果很强(确定性 ρ = 0.676,完整 ρ = 0.819);而针对真实 LLM 生成的候选结果,同样的得分降至 0.158 和 0.406,约为校准值的四分之一(对同一设计者混淆因素进行了测量)。作为预过滤器,它令人信服地优于随机评分对照,并略微优于二元门限。在 400 轮显式优化下该机制未被操纵,但这仅仅是因为目标函数在得到一个诚实答案后即达到饱和。跨家族评审员在保留负担排序信号的同时破坏了个体判断。最终未产生任何覆盖码记录。实测程序成本:164.78 美元。
cs.AI / 124 / 2609.07925

FrogNano: Training a 4B Coding Agent via Online Task Synthesis

FrogNano:通过在线任务合成训练一个4B编码智能体
Kim, Minseon, Shi, Zhengyan, Penaloza, Emiliano, Cui, Christopher, Castanyer, Roger Creus, Hashemzadeh, Maryam, White, Isadora, Light, Jonathan, Kim, Jeonghye, Pereira, Matheus, Moldavskaya, Darya, Singh, Chinmay, Vera, Fabio, Peng, Baolin, Yuan, Xingdi, Côté, Marc-Alexandre, Sordoni, Alessandro
Abstract
We present FrogNano, a 4B coding agent designed to tackle software engineering (SWE) tasks efficiently and effectively, even under resource-constrained environments. It is post-trained exclusively via RL on around 1,500 SWE environments with synthetic tasks. A key ingredient for improving performance is an online task synthesis pipeline that creates tasks calibrated to the frontier of learnability for the current checkpoint. This report provides evidence that competitive small coding agents can be trained with synthetic tasks alone, without traditional distillation from larger models, and that generating tasks at the learnability frontier of the current agent is important. We report details on the training methodology, evaluations across diverse environments, and in-depth analyses, serving as a foundation for our ongoing exploration of lightweight yet capable coding agents that can run on minimal hardware.
Chinese Translation
我们提出了FrogNano,一个40亿参数(4B)的编码智能体,旨在资源受限的环境下也能高效且有效地解决软件工程(SWE)任务。它完全通过强化学习(RL)在大约1500个包含合成任务的SWE环境中进行后训练。提升性能的一个关键要素是在线任务合成流水线,该流水线能够创建与当前模型检查点(checkpoint)可学习性边界相校准的任务。本报告提供了证据,表明仅使用合成任务、无需从更大模型进行传统蒸馏,即可训练出具有竞争力的小型编码智能体,并且在当前智能体的可学习性边界处生成任务十分重要。我们报告了训练方法的细节、跨多种环境的评估结果以及深入分析,为我们持续探索能够在最小硬件上运行的轻量级且能力强大的编码智能体奠定基础。
cs.AI / 125 / 2609.07943

Beliefs and Behavior in Language Models

语言模型中的信念与行为
Smolin, Alex, Wilder, Bryan
Abstract
There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In addition to the inherent scientific interest of this question, these latent quantities are often invoked to explain the behavior of LLMs to users or to define and evaluate harmful behaviors which are relative to intent. Nevertheless, we currently lack a means to systematically test whether concepts like "belief" are well-applied to LLMs, and hence whether they are likely to be fruitful ingredients of attempts to align models with human interests. We propose an approach for empirically studying such questions, asking whether a single latent variable inferred from the LLMs' outputs -- interpreted as a degree of belief -- allows an observer to make interpretable predictions of how the LLMs' will respond to new prompts. We find that highly capable models are usefully described as holding beliefs and that, generally, the predictability of model outputs based on an inferred latent belief tracks overall trends in model capability. Building on these findings, we provide empirical strategies to study how beliefs in LLMs can be measured, the extent to which LLMs comply with instructed decision rules or payoffs, and how beliefs evolve within individual instances of an LLM over the course of reasoning.
Chinese Translation
关于信念或欲望等抽象概念是否能有效描述大语言模型(LLM)的行为,目前仍存在很大的不确定性。除了这个问题本身的科学价值之外,这些潜在量常被用于向用户解释LLM的行为,或定义和评估与意图相关的有害行为。然而,我们目前缺乏一种系统性的方法来检验"信念"这类概念是否适用于LLM,因此也无法确定它们是否可能成为使模型与人类利益对齐的尝试中的有效要素。我们提出了一种实证研究此类问题的方法,即探究从LLM输出中推断出的单一潜在变量——可解释为一种信念程度——是否能让观察者对LLM如何响应新提示做出可解释的预测。我们发现,能力较强的模型可以被有效地描述为持有信念,并且总体而言,基于推断的潜在信念对模型输出的可预测性能够追踪模型能力的整体趋势。基于这些发现,我们提供了实证策略,用以研究如何测量LLM中的信念、LLM在多大程度上遵循指定的决策规则或收益结构,以及信念在LLM单次推理过程中如何演变。
cs.AI / 126 / 2609.07944

CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows

CausalVerify:一个面向大语言模型因果推断工作流的基于执行验证的基准
Zhang, Yonghong, Correia, Ricardo, Parra, Isabel M., Xie, Yong
Abstract
Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric causal-estimation workflows by separating realistic interpretation from verifiable computation. It pairs 259 published economics papers (reconstructed research question, data description, institutional context) with 100 fixed-seed synthetic scenarios that realise CSV datasets for difference-in-differences, event study, instrumental variables, and regression discontinuity designs. Experiment A (real-paper text agreement) scores method-family and direction agreement against four-LLM consensus labels. Experiment B (synthetic execution) runs model-written R code and checks whether the extracted treatment-effect estimate matches a canonical estimator on the same realised dataset; this execution-grounded correctness layer is L2b+, distinct from L2b, which records only whether the code executes. A calibration arm asks whether self-reported confidence separates correct from incorrect workflows. On Experiment B, seven LLMs reach L2b+ pass rates of 10% to 88% at the default 50% tolerance, and 66 of the 426 workflows that execute (15.5%) return a wrong estimate. Execution ranking (L2b) agrees with L2b+ far better than text-direction scoring (L4): Kendall $\tau=0.81$ and Spearman $\rho=0.93$, versus Kendall $\tau$ between $-0.20$ and $0.10$ for L4. Llama-3.3-70B-Instruct shows the same qualitative gap, and reported confidence does not reliably separate correct from incorrect workflows. The claims are confined to standardized single-shot workflows in these four design families under the evaluated R backend and model panel; the benchmark does not measure general causal-inference ability. Code, data, cached outputs, and a datasheet are released.
Chinese Translation
现有面向大语言模型(LLM)的因果推断基准大多评估方法描述的准确性或生成代码能否运行,而非所执行的工作流是否能恢复目标因果估计量。CausalVerify 针对结构化计量经济学因果估计工作流研究这一验证问题,将现实中的解读与可验证的计算分离。该基准将 259 篇已发表的经济学论文(包括重构的研究问题、数据描述与制度背景)与 100 个固定随机种子的合成场景配对,这些场景生成用于双重差分法(difference-in-differences)、事件研究(event study)、工具变量法(instrumental variables)和断点回归(regression discontinuity)设计的 CSV 数据集。实验 A(真实论文文本一致性)依据四个 LLM 的共识标签对方法族一致性和估计方向一致性进行评分。实验 B(合成执行)运行模型编写的 R 代码,并检查提取的处理效应估计值是否与同一已生成数据集上的规范估计量一致;这一基于执行的正确性层面称为 L2b+,不同于仅记录代码能否执行的 L2b。校准部分则考察自报告置信度能否区分正确与错误的工作流。在实验 B 中,七个 LLM 在默认 50% 容差下的 L2b+ 通过率为 10% 至 88%,且在 426 个可执行的工作流中有 66 个(15.5%)返回错误估计。基于执行的排序(L2b)与 L2b+ 的一致性远高于文本方向评分(L4):Kendall τ=0.81、Spearman ρ=0.93,而 L4 的 Kendall τ 在 -0.20 至 0.10 之间。Llama-3.3-70B-Instruct 也表现出同样的定性差距,且报告的置信度无法可靠地区分正确与错误的工作流。上述结论仅限于在所评估的 R 后端和模型组合下、这四种设计族中的标准化单次工作流;该基准并不衡量一般性的因果推断能力。代码、数据、缓存输出及数据说明文档均已公开发布。
cs.AI / 127 / 2609.07954

Support Topology and Gradient Mixing in Sinkhorn Layers

Sinkhorn层中的支撑拓扑与梯度混合
Forde, Dylan
Abstract
Sparse Sinkhorn layers use a fixed support graph to restrict transport between tokens. How does this graph control gradient propagation through the scaling iterations. We develop a fixed-support calculus showing that each row-column cycle induces a row-stochastic operator on column-potential perturbations modulo constants. Its transpose propagates zero-mass reverse-mode cotangents. The finite-cycle operator uses two distinct half-step transport plans; at a balanced fixed point it reduces to a two-step walk determined by a single plan. We derive the accompanying score and marginal source terms and use Dobrushin contraction and minorization to bound homogeneous and source-driven tail cotangents. Our main result characterizes when support and marginals guarantee one-step contraction uniformly over finite scores: every feasible face of the transportation polytope must have pairwise two-hop column overlap. Otherwise, suitable score directions make the contraction coefficient arbitrarily close to one. We extend this analysis to ordered support schedules and derive certificates for partition heat-bath layers, coordinate sweeps, forced shared mass, and register-augmented supports. These results provide mathematical criteria for support design in differentiable transport layers, with guarantees restricted to the fixed-support quotient-gradient component.
Chinese Translation
稀疏Sinkhorn层使用固定的支撑图来限制令牌之间的传输。该图如何控制缩放迭代过程中的梯度传播?我们建立了一套固定支撑的微积分方法,证明每个行列循环在模常数的意义下,对列势扰动诱导出一个行随机算子。其转置算子则传播零质量的反向模式余切。有限循环算子使用两个不同的半步传输计划;在平衡的不动点处,它可简化为仅由单一传输计划确定的两步游走。我们推导了相应的得分项和边缘源项,并利用Dobrushin收缩与小型化(minorization)技术来界定齐次余切与源驱动余切的尾部。我们的主要结果刻画了支撑与边缘分布在何种条件下能保证对有限得分一致的一步收缩:传输多面体的每个可行面必须具有成对的二跳列重叠。否则,适当的得分方向会使收缩系数任意接近于一。我们将此分析扩展到有序支撑调度,并为分区热浴层、坐标扫掠、强制共享质量以及带寄存器增广的支撑导出了判定证书。这些结果为可微传输层中的支撑设计提供了数学准则,其保证仅限于固定支撑的商梯度分量。
cs.AI / 128 / 2609.07984

From Event Logs to Governed Action: A BlueSky Agenda for Agentic Process Mining

从事件日志到受治理的行动:智能体流程挖掘的BlueSky议程
Yang, Yiyuan, Wu, Zheshun, Chu, Yong, Chen, Zhenghua, Xu, Zenglin, Wen, Qingsong
Abstract
Process mining has long turned event logs into process knowledge: discovered models, conformance evidence, bottleneck diagnoses, and runtime predictions. Agentic AI changes the target. Process-aware agents will not only ask what happened. They will ask whether a proposed action should be taken, given the available evidence, privacy budget, organizational authority, and downstream risk. This BlueSky paper proposes event-to-action process mining: a process-mining agenda for transforming heterogeneous operational event data into governed action. The goal is not another dashboard, a generic enterprise simulator, or a language interface over logs. We argue that the community needs four mineable artifacts: event-object representations, action evidence packages, governance contracts, and benchmarks where act, defer, ask, and refuse are all valid outputs. This agenda is timely because agentic business process management (BPM), LLM-assisted process mining, object-centric event standards, causal process monitoring, and privacy-preserving learning are maturing separately. Bringing them together defines a data-mining target inside process mining: mining logged organizational behavior for accountable action, not only retrospective insight.
Chinese Translation
流程挖掘长期以来一直将事件日志转化为流程知识:发现的模型、符合性证据、瓶颈诊断和运行时预测。智能体人工智能(Agentic AI)改变了这一目标。具备流程感知能力的智能体不仅会询问发生了什么,它们还会在综合考虑可用证据、隐私预算、组织权限和下游风险的前提下,询问某个提议的行动是否应该被执行。这篇BlueSky论文提出了“事件到行动”的流程挖掘(event-to-action process mining):一项旨在将异构运营事件数据转化为受治理行动的流程挖掘议程。其目标并非再做一个仪表盘、一个通用的企业模拟器,或一个基于日志的语言交互界面。我们认为,该领域需要四种可挖掘的工件:事件-对象表示(event-object representations)、行动证据包(action evidence packages)、治理契约(governance contracts),以及将“执行、推迟、询问和拒绝”均视为有效输出的基准测试。这一议程恰逢其时,因为智能体业务流程管理(BPM)、基于大语言模型(LLM)辅助的流程挖掘、以对象为中心的事件标准、因果流程监控以及隐私保护学习正在各自走向成熟。将它们整合在一起,便在流程挖掘内部定义了一个数据挖掘目标:从组织行为日志中挖掘出可问责的行动,而不仅仅是回溯性的洞见。
cs.AI / 129 / 2609.07987

When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability

何时LLM数字孪生可以减少人类测量?从行为保真度到统计可替代性
Wang, Steven, Hunt, Kyle, Tang, Shaojie, Joseph, Kenneth
Abstract
LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide little evidence about whether they can reduce human measurement while preserving valid inference. To address this, we introduce statistical substitutability, an inferential criterion that evaluates the extent to which twin predictions can reduce human measurement for a particular estimand while preserving valid inference. We develop a framework, grounded in mixed-subject and prediction-powered inference, that evaluates statistical substitutability along four dimensions: aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations. Across two empirical evaluations spanning behavioral experiments, multiple models, and alternative respondent representations, we find that digital twins can reproduce average human effects while providing little information about which individuals differ from those averages. Newer models and richer respondent information improve some dimensions of performance but do not reliably translate into human-data savings. Human calibration can reduce aggregate prediction error, yet limited labeled samples often fail to produce stable precision gains. Importantly, these findings demonstrate that behavioral fidelity is neither necessary nor sufficient for statistical substitutability. More broadly, they suggest that AI-generated evidence should be evaluated based on its ability to support valid scientific inference rather than its ability to reproduce human outcomes alone. Digital twins should therefore be judged for confirmatory use by whether they reduce uncertainty about human quantities, not merely by whether they reproduce human means, distributions, or effects.
Chinese Translation
基于大语言模型(LLM)的数字孪生有望通过生成针对特定个体的回答来减少重复的人类数据收集,然而现有评估几乎未能提供证据表明它们能否在保持有效推断的前提下减少人类测量。为此,我们提出了统计可替代性(statistical substitutability)这一推断性标准,用以评估孪生预测在多大程度上能够在保持有效推断的同时减少针对特定估计目标的人类测量。我们构建了一个基于混合被试推断与预测驱动推断的框架,从四个维度评估统计可替代性:总体保真度、配对的个体层面信号、有限样本下人类标签的恢复能力,以及跨人群的稳定性。在涵盖行为实验、多个模型和不同被试表征方式的两组实证评估中,我们发现数字孪生能够复现人类的平均效应,但对于哪些个体偏离这些平均值几乎不提供信息。更新的模型和更丰富的被试信息虽然提升了某些维度的表现,但并不能可靠地转化为人类数据需求的节省。人类校准可以降低总体预测误差,但有限的标注样本往往无法带来稳定的精度提升。重要的是,这些发现表明行为保真度既不是统计可替代性的必要条件,也不是充分条件。更广泛地说,这些结果表明,对AI生成证据的评估应基于其支持有效科学推断的能力,而不仅仅是其复现人类结果的能力。因此,判断数字孪生能否用于验证性研究,应看其是否能减少关于人类量的不确定性,而非仅仅看其能否复现人类的均值、分布或效应。
cs.AI / 130 / 2609.07998

Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study

小批量风险规避深度Q学习:一个机器人导航案例研究
Patel, Aayush, Ruszczyński, Andrzej
Abstract
We study the control of Markov decision processes in which the quality of a policy is evaluated by a dynamic, time-consistent Markov risk measure rather than by an expected discounted cost. The main obstacle to combining such measures with reinforcement learning is that a transition risk mapping depends on the transition kernel in a nonlinear way, and therefore cannot be estimated from a single observed transition. We remove this obstacle by employing mini-batch transition risk mappings: the mapping is applied to the empirical measure of $N$ independent next-state samples, and the result is averaged. The resulting mapping is again coherent. However, as an expected value of a function of $N$ next-state values, it admits an unbiased one-sample estimator. We embed this mapping into a double deep Q-network, analyze the two sources of estimation bias that arise, and obtain a risk-averse Q-learning method applicable to state spaces far beyond the reach of tabular schemes. The method is applied to an underwater robot navigation problem, in which a vehicle must visit collection points, gather stochastic information payloads, and deliver them at transmission points, while exposed at each step to the risk of destruction. A hierarchical decomposition delegates path execution to an exact graph search and confines learning to the high-level ``collect or transmit'' decision. A low-dimensional feature map, invariant under the symmetries of the problem, replaces the raw state--configuration encoding. In experiments on $300$ held-out environments, the resulting policies transfer to instance sizes never seen in training, and already $N=2$ reduces the upper semideviation of the outcome distribution while simultaneously improving its mean whenever the simulator is misspecified---an empirical counterpart of the duality between coherent risk measures and distributional robustness.
Chinese Translation
我们研究马尔可夫决策过程的控制问题,其中策略的质量由动态、时间一致的马尔可夫风险度量来评估,而非由期望折衷成本来评估。将此类度量与强化学习相结合的主要障碍在于,转移风险映射以非线性方式依赖于转移核,因此无法从单次观测到的转移中估计。我们通过采用小批量转移风险映射来消除这一障碍:将该映射应用于 $N$ 个独立下一状态样本的经验测度,并对结果取平均。所得映射依然是连贯的(coherent)。然而,作为 $N$ 个下一状态值之函数的期望值,它允许一个无偏的单样本估计器。我们将该映射嵌入双重深度Q网络(double deep Q-network),分析了由此产生的两类估计偏差,并得到了一种适用于远超表格型方法所能处理的状态空间的风险规避Q学习方法。该方法被应用于一个水下机器人导航问题,其中机器人必须访问收集点、采集随机信息载荷并将其传送至传输点,同时在每一步都面临被摧毁的风险。通过层次化分解,路径执行被交由精确的图搜索完成,而学习则仅限于高层的“收集或传输”决策。一个在该问题对称性下保持不变的低维特征映射取代了原始的状态-构型编码。在 $300$ 个留出环境上的实验表明,所得策略能够迁移到训练中从未见过的实例规模;此外,当仿真器存在误设时,仅 $N=2$ 便能降低结果分布的上半偏差(upper semideviation),同时改善其均值——这是连贯风险度量与分布鲁棒性之间对偶关系的一个经验性体现。
cs.AI / 131 / 2609.08003

Sparks of In Silico Cognitive Science: Theories from Simulated Data Can Generalize to Humans

硅基认知科学的火花:从模拟数据中发现的理论可以推广到人类
Jagadish, Akshay K., Strittmatter, Younes, Jacoby, Nori, Schulz, Eric, Daw, Nathaniel, Griffiths, Thomas L., Chandramouli, Suyog H.
Abstract
Behavioral foundation models have been proposed as stand-ins for human participants across settings, but it is unclear whether theories discovered on them generalize to humans or merely characterize the simulator. We ran the Automated Cognitive Scientist (\textsc{AutoCog}), a closed-loop discovery system in which LLM agents design theory-discriminating experiments, collect responses, arbitrate between competing theories, and synthesize successors, entirely on behavior simulated by Centaur, a foundation model of human behavior. In a multi-attribute decision-making setting, the theories \textsc{AutoCog} found on Centaur generalized to human data: they outperformed canonical theories on ten held-out experiments and were rivaled only by theories found by running the same loop on people. We argue that this succeeds despite the simulator's inevitable imperfections because a discovery loop that arbitrates between competing theories demands less of its simulator than estimation does. The simulator only needs to capture the regularities that distinguish the theories, and not necessarily reproduce behavior precisely. Imperfect simulators can therefore widen the search over theories, with human data then testing whether the surfaced theories generalize.
Chinese Translation
行为基础模型(behavioral foundation models)已被提议作为人类被试在各种实验场景中的替代品,但尚不清楚在它们上发现的理论是能够推广到人类,还是仅仅刻画了模拟器本身的特性。我们运行了自动化认知科学家系统(Automated Cognitive Scientist, AutoCog)——一个闭环科学发现系统,其中大语言模型智能体设计用于区分理论的实验、收集反应、在竞争理论之间进行仲裁,并综合出后续理论——整个过程完全基于Centaur(一个人类行为基础模型)模拟的行为。在多属性决策任务中,AutoCog 在 Centaur 上发现的理论成功推广到了人类数据:在十项留出实验中,这些理论优于经典理论,且只有通过在真人身上运行同样的发现循环所获得的理论才能与之匹敌。我们认为,尽管模拟器不可避免地存在缺陷,这一方法仍然成功,原因在于:在竞争理论之间进行仲裁的发现循环对模拟器的要求低于参数估计。模拟器只需捕捉区分不同理论的规律性,而不必精确复现行为。因此,不完美的模拟器可以拓宽理论搜索的范围,再由人类数据检验由此发现的理论是否能够推广。
cs.AI / 132 / 2609.08015

From Version Conflicts to Decision Conflicts: Selective Revalidation for Long-Running AI Agents

从版本冲突到决策冲突:面向长时运行AI智能体的选择性重验证
Lyu, Yongjian, Ren, Yang, Lai, Ruofei, Liu, Wenting
Abstract
Long-running AI agents may read state, reason, wait for tools or human approval, and perform an external action much later. The state that justified the action can change in the meantime. For example, after an agent proposes an 80 GBP refund under a limit of 100, a customer-name change affects only presentation metadata, a new limit of 90 still permits the refund, a limit of 50 invalidates it, and a refund issued by another worker must prevent a duplicate. Standard optimistic concurrency control and version checks can detect that previously read state has changed, but by themselves do not determine whether that change invalidates the pending action's justification. We call any detected version change a version conflict; when that change invalidates the action's justification, it is also a decision conflict. ATR records the explicit, executable conditions that justify a pending action and rechecks only the conditions affected by a change before releasing the external operation. It can retain the action, refresh non-decisive metadata, require replanning, or block execution; a target-side transaction or compare-and-set binds checked state to commit. Across 210,000 controlled executions over 15 mutation cases, ATR matched every developer-specified outcome with no false allows or blocks. In ten durable SQLite checkpoint/resume cells, it evaluated 0.6 conditions per change versus 6.0 for FullScan. At 4,093 recorded reads, ATR took 9.3 microseconds versus 2595.9 microseconds for FullScan. These deterministic results establish controlled feasibility, not production generality or automatic extraction of the required conditions.
Chinese Translation
长时运行的AI智能体可能读取状态、进行推理、等待工具或人工批准,并在很久之后才执行外部操作。在此期间,支持该操作依据的状态可能已发生变化。例如,当智能体在100英镑的限额内提议一笔80英镑的退款后,客户名称的更改仅影响展示元数据,新的90英镑限额仍允许该退款,而50英镑的新限额则使其失效,其他工作进程发出的退款则必须防止重复退款。标准的乐观并发控制和版本检查可以检测到先前读取的状态已发生变化,但其本身并不能判断该变化是否会使待执行操作的依据失效。我们将任何检测到的版本变化称为版本冲突(version conflict);当该变化使操作的依据失效时,它同时也是决策冲突(decision conflict)。ATR记录支持待执行操作的显式可执行条件,并在释放外部操作之前仅重新检查受变化影响的条件。它可以保留该操作、刷新非决定性元数据、要求重新规划或阻止执行;目标端的事务或比较并交换(compare-and-set)机制将已检查的状态与提交绑定。在涵盖15种变更场景的210,000次受控执行中,ATR与开发者指定的每一种结果完全匹配,没有出现任何错误放行或错误阻止。在十个持久化SQLite检查点/恢复测试单元中,ATR对每次变更平均评估0.6个条件,而全量扫描(FullScan)为6.0个。在4,093次记录读取中,ATR耗时9.3微秒,而FullScan耗时2595.9微秒。这些确定性结果确立了受控环境下的可行性,但并不代表生产环境的普适性,也不涉及所需条件的自动提取。
cs.AI / 133 / 2609.08016

A Layered Analysis of Disagreement And Answer Quality in Multi-Agent LLM Debate

多智能体LLM辩论中分歧与答案质量的分层分析
Qian, Chen
Abstract
Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve answer quality by surfacing genuine disagreement. That mechanism is rarely checked. We introduce four measurements: (A) the agreement a debater reports; (B) whether its reply text actually pushes back; (C) whether the position persists once the eliciting instruction is removed; and (D) for open-weight models, the stance response in the debater's own token log-probabilities. We evaluate three-model committees debating open-ended GlobalOpinionQA across 750 debates under three tones: friendly (seek common ground), neutral, and hostile (stress-test every position). (A) Tone strongly reshapes reported agreement: full agreement differs by 50.4 percentage points between the friendly and hostile endpoints. (B) A judge that reads only the reply text, never the self-report or the condition, recovers the same pattern. (C) The dissent appears partly tied to the instruction that elicited it: labels revert toward agreement 23.1 points more often after deleting the hostile instruction than under a matched re-ask that keeps it; question-weighted inference is inconclusive on first-round turns alone (p=0.0625), significant pooling all rounds (p=0.016), and only 11/28 first-round reversions also appear in the reply text. (D) Opposing arguments weaken a debater's stance margin more consistently than they shift its direction. For final answers we detect no quality gain: a bias-checked jury returns 299/299 ties (ruling out only large differences), accuracy on a verifiable control task is unchanged, and a jury without the bias check had declared debate the winner 66% of the time -- an artifact of reading order. Taken together, LLM debate readily changes what agents say, but we find much weaker evidence that it changes what they persistently endorse or improves the quality of the final answer.
Chinese Translation
多智能体辩论(即多个大语言模型在给出答案前交换论点)被普遍认为能够通过显现真实分歧来提升答案质量,但这一机制很少受到检验。我们引入四项测量指标:(A) 辩论者所报告的一致程度;(B) 其回复文本是否确实提出了反驳;(C) 当诱发该指令被移除后,其立场是否持续保持;以及 (D) 对于开放权重模型,辩论者自身token对数概率中的立场响应。我们评估了由三个模型组成的委员会,在750场辩论中围绕开放性数据集GlobalOpinionQA展开辩论,并采用三种语气设定:友好(寻求共识)、中立和敌对(对每个立场进行压力测试)。(A) 语气强烈地重塑了所报告的一致程度:完全一致的比例在友好与敌对两个极端之间相差50.4个百分点。(B) 一个仅读取回复文本(既不看自我报告也不看实验条件)的评判者能够复现出相同的模式。(C) 这种异议部分地与诱发它的指令相关联:与保留敌对指令的匹配性重复提问相比,删除敌对指令后标签向一致方向回退的频率高出23.1个百分点;仅基于第一轮的按问题加权推断结果不确定(p=0.0625),汇总所有轮次后显著(p=0.016),且28例第一轮回退中仅有11例同时体现在回复文本中。(D) 对立论点对辩论者立场边际的削弱比对其方向的改变更为一致。对于最终答案,我们未检测到质量提升:一个经过偏差校验的评审团在299场辩论中全部判定为平局(仅排除了较大的差异),在可验证的对照任务上准确率没有变化;而未经偏差校验的评审团曾判定辩论胜出的比例为66%——这其实是阅读顺序造成的假象。综上所述,LLM辩论容易改变智能体的言辞,但我们发现它改变其持续认同的立场或提升最终答案质量的证据要弱得多。
cs.AI / 134 / 2609.08025

Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning

基于强化学习的多模态推理智能体自我验证能力激发
Sathish, Vishwas, Ranjan, Viresh, Zhu, Xinliang, Dhua, Arnab, Gray, Douglas
Abstract
Reasoning agents increasingly rely on external tools such as web search to answer complex queries. Reinforcement learning (RL) finetuning algorithms such as GRPO have improved long-form reasoning in text-only language models, particularly for coding and mathematics. Reliable tool use in multimodal agents, however, remains challenging because models must interpret text and images while integrating noisy retrieved evidence, often under sparse outcome-level supervision without explicit verification signals. We present Self-Verification via Reinforcement Learning (SVRL), an RL-only finetuning framework that trains multimodal agents to verify and filter retrieved evidence within their own reasoning traces, reducing reliance on external verifiers at inference time. SVRL also introduces a search-aware penalty that discourages unnecessary tool calls and a query-diversity reward that encourages diverse, well-formed search queries, providing fine-grained feedback on when and what to search. Finetuning Qwen-2.5-VL-7B with SVRL on only 5{,}000 visual question answering examples yields consistent gains in multi-hop VQA generalization and tool efficiency across benchmarks. Overall, SVRL narrows the gap between compact agents and much larger proprietary models while requiring substantially lower training and inference cost.
Chinese Translation
推理智能体日益依赖网络搜索等外部工具来回答复杂查询。诸如GRPO等强化学习(RL)微调算法已显著提升了纯文本语言模型的长篇推理能力,尤其是在编程和数学任务上。然而,在多模态智能体中实现可靠的工具使用仍然具有挑战性,因为模型必须在解读文本和图像的同时整合含噪的检索证据,且往往处于稀疏的结果级监督下,缺乏显式的验证信号。我们提出了基于强化学习的自我验证方法(Self-Verification via Reinforcement Learning, SVRL),这是一个纯强化学习微调框架,用于训练多模态智能体在其自身的推理轨迹中验证和过滤检索到的证据,从而减少推理时对外部验证器的依赖。SVRL还引入了一种搜索感知惩罚机制以抑制不必要的工具调用,并设计了查询多样性奖励以鼓励生成多样化且格式规范的搜索查询,从而为搜索的时机与内容提供细粒度反馈。仅使用5,000个视觉问答样本对Qwen-2.5-VL-7B进行SVRL微调,即可在各基准测试中持续提升多跳视觉问答(VQA)泛化能力和工具使用效率。总体而言,SVRL在显著降低训练和推理成本的同时,缩小了紧凑型智能体与规模大得多的专有模型之间的差距。
cs.AI / 135 / 2609.08062

ResidualAuth: What Authorization State Must Language Agents Preserve under Revocable Delegation?

ResidualAuth:在可撤销委托下语言代理必须保留哪些授权状态?
Choi, Moonwon, Jeong, Seokho, Lee, Seunggeun
Abstract
Tool-using language agents can delegate and revoke permissions while acting through external services. We show that two authorization histories can have identical current permissions and identical all-pairs reachability yet require opposite decisions after the same direct-edge revocation. We formalize the information needed to preserve such distinctions as a residual authorization state. We prove that exponentially many future-distinct states can share one fixed transitive closure, and give exact or tight asymptotic bounds on the state required by an exact monitor as delegation redundancy varies. ResidualAuth compiles these constructions into paired language-agent episodes. Across four open-weight models, a fixed 256-token summary solved 0-2/16 pairs, sham reads solved 0/16, and authenticated current-query reads solved 15-16/16. In a separate held-out online-memory diagnostic, exact ledger serializations fit all 128 four-coordinate pairs at both 768 and 1,024 tokens. At either cap, factually supported model-written memories sufficient for every prespecified continuation solved at most 1/128 pairs per model. A hard gate reduced eight observed unauthorized effects to zero without changing the preceding attempts. These results distinguish required authorization state, usable decision information, online state maintenance, and effect mediation.
Chinese Translation
使用工具的语言代理(language agents)在通过外部服务执行操作时可以授予和撤销权限。我们证明,两个授权历史可以拥有完全相同的当前权限和相同的所有节点对可达性,但在同一直接连边撤销后却需要作出截然相反的决策。我们将保留这类区分所需的信息形式化为残余授权状态(residual authorization state)。我们证明,指数级数量的具有不同未来语义的状态可以共享同一个固定的传递闭包,并给出了随委托冗余度变化时精确监控器所需状态的精确或紧致渐近界。ResidualAuth 将这些构造编译成成对的语言代理情景。在四个开放权重模型上,固定的 256 词元(token)摘要仅解决了 16 对中的 0-2 对,虚假读取解决了 0/16,而经过认证的当前查询读取解决了 15-16/16。在另一项独立的保留集在线记忆诊断中,精确的账本序列化在 768 和 1,024 词元下均能解决全部 128 个四坐标对。在任一词元上限下,事实上有依据且足以满足每个预设后续任务的模型撰写记忆,每个模型至多解决了 128 对中的 1 对。一个硬性门控将八个观察到的未授权效应降至零,且未改变先前的尝试。这些结果区分了必需的授权状态、可用的决策信息、在线状态维护以及效应中介。
cs.AI / 136 / 2609.08071

Automated Design of Inventory Policy with Large Language Models: An Exploratory Study

基于大语言模型的库存策略自动化设计:一项探索性研究
Yang, Fenghua, Baxi, Preet, Zhang, Yi, Jasin, Stefanus, Lei, Yanzhe, Liu, Mo, Pakiman, Parshan
Abstract
Firms making inventory decisions have access to operational data, optimization tools, and large language models (LLMs). Typically, data characterize the operating environment, optimization selects parameters within a prespecified inventory policy class, and LLMs support coding and decision analysis. We develop an integrated framework that combines these resources to automate inventory policy design. Given demand data, the framework iteratively uses an LLM to generate parameterized policy classes and an external solver to optimize its parameters within each class. Across 30 lost-sales inventory instances, the mean cost reduction relative to optimized base-stock benchmarks increases from 17.5% after one generation to 30.0% after ten generations. Parameter optimization is central to this performance: an LLM-only variant performs substantially worse, whereas optimization-guided feedback improves policy quality, accelerates search, and directs the LLM toward better policy classes rather than merely better parameter values within a fixed class. The strongest discovered policies are also interpretable: they combine recognizable inventory-control motifs, including capped orders, discounted or weighted pipeline inventory, and threshold-based replenishment logic. The search thereby produces new policy-class functional forms that, to our knowledge, have not previously been studied in the lost-sales inventory literature. These functional forms are not specified ex ante but emerge from the search process. Moreover, after their parameters are re-optimized, three discovered policy classes achieve average cost reductions of 21.75% to 22.60% across 10,064 new inventory instances. Overall, the results show that data-driven parameter optimization can guide LLM-based search over a broad space of inventory policy classes and identify high-performing, interpretable, and transferable decision rules.
Chinese Translation
进行库存决策的企业可以获取运营数据、优化工具和大语言模型(LLM)。通常,数据用于刻画运营环境,优化方法在预先设定的库存策略类中选择参数,而LLM则支持编程和决策分析。我们开发了一个整合这些资源的集成框架,以实现库存策略设计的自动化。给定需求数据后,该框架迭代地利用LLM生成参数化的策略类,并使用外部求解器在每个策略类内优化其参数。在30个缺货损失(lost-sales)库存实例中,相对于优化后的基本库存(base-stock)基准,平均成本降幅从第一代后的17.5%提升至第十代后的30.0%。参数优化是这一性能的核心:仅使用LLM的变体表现明显较差,而基于优化的反馈改进了策略质量,加速了搜索,并引导LLM趋向更好的策略类,而不仅仅是在固定策略类内寻找更好的参数值。所发现的最强策略同样具有可解释性:它们结合了可识别的库存控制模式,包括订单上限、折扣或加权在途库存以及基于阈值的补货逻辑。因此,该搜索产生了新的策略类函数形式,据我们所知,这些形式此前尚未在缺货损失库存文献中被研究过。这些函数形式并非事先指定,而是在搜索过程中涌现的。此外,在对其参数重新优化后,三种被发现的策略类在10,064个新库存实例上实现了21.75%至22.60%的平均成本降幅。总体而言,结果表明,数据驱动的参数优化可以引导基于LLM的搜索在广泛的库存策略类空间中进行,并识别出高性能、可解释且可迁移的决策规则。
cs.AI / 137 / 2609.08082

Inference-Time Nash Alignment

推理时纳什对齐
Hosseini, Hadi, Mandal, Debmalya, Zhang, Duohan
Abstract
Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large preference datasets. They also need direct access to the model parameters which are not provided by many state-of-the art models. Inference-time alignment offers a cost-effective alternative without updating model parameters. However, existing inference-time methods rely on a scalar reward model derived under a Bradley-Terry assumption, which cannot represent general preferences. Following recent work on fine-tuning with generalized preferences, in this work, we initiate the study of inference-time alignment under general preferences. We formulate the problem as obtaining a Nash equilibrium of a two-player zero-sum game between policies. We propose two algorithms: Best-of-Nash (BoN) and Nash Mirror Descent (NMD). We prove that both algorithms achieve a duality gap that matches the problem lower bound. Empirically, we implement the two methods on three datasets, which shows that our methods substantially outperform the base policy, converging to the performance of the fine-tuned models. Moreover, our results show that NMD remains robust across the regularization parameter.
Chinese Translation
基于偏好的微调方法(如RLHF和DPO)需要大量的计算资源和大规模的偏好数据集。它们还需要直接访问模型参数,而许多最先进的模型并不提供这一访问权限。推理时对齐提供了一种无需更新模型参数的高性价比替代方案。然而,现有的推理时方法依赖于在Bradley-Terry假设下导出的标量奖励模型,无法表示一般化的偏好。跟随近期关于广义偏好微调的工作,本文开创了在一般偏好下推理时对齐的研究。我们将该问题形式化为求解策略之间两人零和博弈的纳什均衡。我们提出了两种算法:Best-of-Nash(BoN)和纳什镜像下降。我们证明这两种算法所达到的对偶间隙与问题的下界相匹配。在实验方面,我们在三个数据集上实现了这两种方法,结果表明我们的方法显著优于基线策略,并收敛到微调模型的性能水平。此外,我们的结果显示NMD在不同正则化参数下依然保持稳健。
cs.AI / 138 / 2609.08090

RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts

RevalExo:面向老年人及临床队列的惯性与视觉运动模式识别功能性日常活动基准数据集
Lamsal, Diwas, Carlon, Juha, Claeys, Reinhard, Yudayev, Maxim, Flynn, Louis, Verstraten, Tom, Beckwée, David, Swinnen, Eva, Bâce, Mihai, Vanrumste, Bart, Filtjens, Benjamin
Abstract
Assistive devices for people with mobility impairments, such as powered exoskeletons, rely on accurate locomotion mode recognition to adapt control strategies and provide appropriate assistance during daily activities. However, public benchmarks are typically collected from healthy adults, lack temporally precise labels necessary for detecting mode transitions, or focus on a limited set of tasks. To support development and evaluation under realistic clinical constraints and daily mobility demands, we introduce RevalExo, a functional daily-activity benchmark for inertial and visual locomotion mode recognition. RevalExo is built around a standardized, clinically and ecologically validated daily-activity protocol reflecting the cumulative everyday mobility demands in ageing and clinical populations. The benchmark includes 27 participants across three cohorts: older adults without mobility impairments, stroke survivors, and older adults with probable sarcopenia. The full cohort was recorded with lower-body IMUs, while synchronized egocentric video was collected for a clinically feasible subset of 13 participants. RevalExo provides 10.1 hours of frame-level annotations across 11 locomotion modes, including 5.1 hours of paired inertial--visual recordings. We benchmark three challenges: unimodal and multimodal locomotion mode recognition across multiple horizons, cross-population generalization from older adults without mobility impairments to clinical cohorts, and vision-guided knowledge transfer to IMU-only models. Results confirm consistent gains from fusing inertial and visual inputs but reveal a substantial gap between general recognition ($\sim$93\% F1) and recognition during transitions ($\sim$68\% F1), alongside persistent challenges in cross-population generalization and cross-modal transfer. We release RevalExo to stimulate further research on these open challenges.
Chinese Translation
针对行动障碍人群的辅助设备(如动力外骨骼)依赖准确的运动模式识别,以在日常生活中调整控制策略并提供恰当的辅助。然而,现有的公开基准数据集通常采集自健康成年人,缺乏检测模式转换所需的精确时间标签,或仅关注有限的任务集合。为支持在真实临床约束和日常活动需求下的开发与评估,我们提出了RevalExo,一个用于惯性与视觉运动模式识别的功能性日常活动基准数据集。RevalExo围绕一套标准化、经临床与生态学验证的日常活动协议构建,该协议反映了老龄化人群和临床人群日常行动需求的累积性。该基准数据集涵盖三个队列共27名参与者:无行动障碍的老年人、脑卒中幸存者以及疑似肌少症的老年人。全部参与者均通过下肢IMU(惯性测量单元)进行记录,同时对临床可行的13名参与者子集采集了同步的第一视角视频。RevalExo提供了跨越11种运动模式的10.1小时帧级标注,其中包括5.1小时的惯性—视觉配对记录。我们设置了三项基准挑战:多时间跨度下的单模态与多模态运动模式识别、从无行动障碍老年人到临床队列的跨人群泛化,以及从视觉引导到纯IMU模型的知识迁移。结果表明,融合惯性与视觉输入能够带来一致的提升,但一般识别(F1约93%)与模式转换期间的识别(F1约68%)之间存在显著差距,同时跨人群泛化和跨模态迁移仍面临持续性挑战。我们公开发布RevalExo,以期推动针对这些开放性挑战的进一步研究。
cs.AI / 139 / 2609.08094

CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information

CIVI:一个用于诊断公共信息领域搜索智能体故障的框架
Liu, Dingying, Zhong, Yunshun, Zhang, Wentao, Li, Yiyuan
Abstract
Large Language Models are increasingly deployed in public-sector settings, where incorrect guidance can cause irreversible harm. We introduce CIVI, the first framework for diagnosing search agent failures in civic information. Its benchmark instantiation jointly spans cross-national, interjurisdictional government contexts (federal, state, and local) and functional categories from an internationally adopted United Nations standard. We evaluate ten frontier search agents and find that none matches an attentive human baseline. Alongside accuracy, CIVI measures search invocation rate, selective no-search accuracy, and how often agents cite authoritative government sources. To perform this diagnosis, we introduce ARISE, which decomposes agentic search failures into four mutually exclusive modes, isolated via source-injection ablation. ARISE attributes 72.1% of all observed failures to retrieval-bound causes rather than to gaps in the models' parametric knowledge.
Chinese Translation
大语言模型越来越多地被部署在公共部门场景中,在这些场景下,错误的指引可能造成不可逆转的伤害。我们提出了CIVI,这是首个用于诊断搜索智能体在公共信息领域故障的框架。其基准实例化同时涵盖了跨国、跨管辖层级的政府场景(联邦、州和地方),以及一个被国际采纳的联合国标准中的功能类别。我们评估了十个前沿搜索智能体,发现没有任何一个能够达到细心的基线水平。除了准确率之外,CIVI还测量了搜索调用率、选择性不搜索准确率,以及智能体引用权威政府来源的频率。为进行此类诊断,我们提出了ARISE,它将智能体搜索故障分解为四种互斥的模式,并通过来源注入消融实验进行分离。ARISE将所有观察到的故障中的72.1%归因于检索相关的限制,而非模型参数化知识的不足。
cs.AI / 140 / 2609.08105

Artificial Intelligence-Assisted Digital Inventory of Cultural Heritage & Traditional Knowledge: Case for Indonesian Open Digital Library of Culture

人工智能辅助的文化遗产与传统知识数字化清点:以印度尼西亚开放文化数字图书馆为例
Situngkir, Hokky
Abstract
The Indonesian Digital Library of Culture (Perpustakaan Digital Budaya Indonesia, PDBI; budaya-indonesia.org) is a participatory platform that has collected tens of thousands of entries on Nusantara cultural heritage through public contribution since 2007. Manual contribution faces three structural barriers: coverage (knowledge is scattered across languages and sites), integrity (open sources mix authentic documentation with noise), and completeness (subjects are recorded but their data remain shallow). This paper presents a methodological framework for autonomous, AI-based harvesting of cultural knowledge from the open web, designed to expand corpus coverage while intensifying per-entry data depth. The methodology is organised as a five-stage economic funnel: focused crawling, multilingual extraction and canonicalisation, vector encoding with blocking, agentic decision-making, and idempotent publication, under the principle of deterministic orchestration, agentic decisions. Each stage is formalised: funnel economics and optimal filter ordering; crawl-frontier dynamics as a subcritical branching process that explains the necessity of recurrent re-seeding; fact-level novelty via a containment measure; Bayesian multi-source evidence fusion with elevated publication thresholds for sacred categories; exactly-once effects via idempotent upserts and the transactional outbox; sliding-window inference budgeting with a reservation protocol; statistical quality auditing; and seed selection as submodular coverage maximisation. The framework retains four high-value human roles: curator of direction, escalation approver, quality auditor, and guardian of meaning, while machine autonomy is raised in stages. Ethical, legal, and cultural-sensitivity implications are discussed, including the architectural guarantee that the machine never overwrites human contributions.
Chinese Translation
印度尼西亚文化数字图书馆(Perpustakaan Digital Budaya Indonesia,PDBI;budaya-indonesia.org)是一个参与式平台,自2007年以来通过公众贡献已收集了数万条关于努山塔拉(Nusantara)文化遗产的条目。人工贡献面临三重结构性障碍:覆盖面(知识分散于不同语言和站点之间)、完整性(开放来源中真实文献与噪声混杂)以及完备性(仅记录了主题本身,而其数据仍然浅薄)。本文提出了一种基于人工智能、从开放网络自主采集文化知识的方法论框架,旨在扩展语料库覆盖面的同时加深每一条目的数据深度。该方法论组织为一个五阶段经济漏斗:聚焦式爬取、多语言抽取与规范化、结合分块的向量编码、智能体决策以及幂等发布,并遵循“确定性编排、智能体决策”的原则。每一阶段均被形式化:漏斗经济学与最优过滤顺序;作为亚临界分支过程的爬取前沿动态,该过程解释了周期性重新播种的必要性;通过包含度度量实现事实级新颖性判断;贝叶斯多源证据融合,并对神圣类别设置更高的发布阈值;通过幂等更新插入与事务性发件箱实现恰好一次(exactly-once)效果;基于预留协议的滑动窗口推理预算;统计质量审计;以及将种子选择建模为次模覆盖最大化问题。该框架保留了四个高价值的人类角色:方向策展人、升级审批人、质量审计员和意义守护者,同时分阶段提升机器自主性。本文还讨论了伦理、法律及文化敏感性方面的影响,包括机器永远不会覆盖人类贡献的架构性保证。
cs.AI / 141 / 2609.08115

Router Prior Bias: Preserving Base Routing Structure in MoE Post-Training

路由器先验偏置:在MoE后训练中保留基础路由结构
Lee, Jaedeok, Kim, Keonwoo, Han, Dongyoon, Yun, Sangdoo, Choi, Yera, Yoo, Haanju
Abstract
Mixture-of-Experts (MoE) pretraining relies on an auxiliary load-balancing loss (LBL) to drive per-expert utilization toward uniformity. Post-training inherits a different situation: the base router already encodes non-uniform expert co-activation structure, which a re-imposed uniformity objective flattens away. We show that downstream performance depends instead on holding this inherited routing softly, a principle we term soft router anchoring, and instantiate it as Router Prior Bias (RPB), a training-time bias that pulls the router logits toward a prior read off the frozen base router while leaving the router itself trainable. On math post-training of Moonlight-16B-A3B, RPB attains 45.77 in-domain accuracy against 31.91 under re-applied LBL and 29.44 under unanchored fine-tuning, and retains more out-of-domain capability than either. The ordering against LBL reproduces on a second model family (Qwen3-30B-A3B-Base), and the advantage over LBL is resolvable on an independently sourced corpus. Anchors defined on the router weights, on its logits, or on its output distribution perform comparably with no consistent ordering, which places the effect in the softness of the constraint rather than in the particular prior RPB supplies. Retained community structure in the expert co-activation graph tracks these gains wherever the base router is non-uniform enough for communities to form, yet enforcing the same prior as a hard assignment preserves that structure while performance falls sharply. Community structure is therefore a footprint of soft anchoring rather than its source, and the practical lesson is that inherited routing should be held softly during post-training, since both flattening it toward uniformity and enforcing it absolutely carry a downstream cost. Our code will be released at https://github.com/naver-ai/rpb.
Chinese Translation
混合专家模型(Mixture-of-Experts, MoE)的预训练依赖辅助负载均衡损失(load-balancing loss, LBL)来推动各专家的利用率趋于均匀。后训练面临的情况有所不同:基础路由器已经编码了非均匀的专家共激活结构,而重新施加的均匀性目标会将这种结构抹平。我们证明,下游性能实际上取决于以软性方式保留这种继承而来的路由结构,我们将这一原则称为软路由器锚定(soft router anchoring),并将其实现为路由器先验偏置(Router Prior Bias, RPB)——一种训练时偏置,它将路由器logits拉向从冻结的基础路由器读取的先验,同时保持路由器本身可训练。在Moonlight-16B-A3B的数学后训练中,RPB取得45.77的域内准确率,而重新应用LBL仅为31.91,无锚定的微调为29.44,且RPB保留的域外能力也多于后两者。相对LBL的优劣顺序在第二个模型家族(Qwen3-30B-A3B-Base)上得以复现,且相对LBL的优势在一个独立来源的语料库上依然可分辨。基于路由器权重、logits或输出分布定义的锚定效果相当,没有一致的优劣排序,这表明效果源自约束的软性,而非RPB提供的特定先验。只要基础路由器的非均匀程度足以形成社区,专家共激活图中保留的社区结构就会随这些收益一同出现;然而,若以硬性分配方式强制施加相同的先验,社区结构虽得以保留,性能却急剧下降。因此,社区结构是软性锚定的印记而非其来源;实践启示是,后训练期间应以软性方式保留继承的路由结构,因为无论是将其均匀化抹平还是绝对化强制执行,都会带来下游性能损失。我们的代码将发布于 https://github.com/naver-ai/rpb。
cs.AI / 142 / 2609.08126

SchemeArena: Factorized Stress Testing of Scheming in LLM Agents

SchemeArena:面向LLM智能体图谋行为的因子化压力测试
Ruan, Jie, Nair, Inderjeet, Liu, Amy, Khalifa, Muhammad, Zhou, Yusheng, Wang, Lu
Abstract
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how scheming arises from the interaction of key factors, such as instrumental goals, environmental affordances, oversight conditions, and perceived consequences. Prior work examines only a small number of scenarios, limiting the ability to isolate how these conditions shape an agent's propensity or capability to scheme. This limited scale and task diversity also restrict coverage of realistic deployment settings and the range of scheming strategies that can be observed. To this end, we introduce SCHEMEARENA, a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spanning diverse safety-relevant tool domains, instrumental goals, oversight conditions, and pressure mechanisms. To enable scalable and reliable monitoring, we further propose SCOUT, a scheming monitor that grounds multi-criteria judgments in evidence drawn from agents' reasoning and actions. Across controlled stress tests on five LLM agents, we find that explicit instrumental goals are the strongest driver of scheming propensity. Strategic hints play a distinct role by helping agents translate scheming reasoning into concrete covert behavior. Oversight has mixed effects: in several closed models, action-only monitoring increases scheming, suggesting that partial oversight can act as an optimization constraint rather than a deterrent. CoT is a useful but incomplete monitoring signal: it can reveal latent scheming before execution, yet action-only scheming shows that covert behavior may occur without explicit reasoning evidence. We release the benchmark, code, and monitor at: https://github.com/launchnlp/SchemeArena.
Chinese Translation
我们研究了LLM智能体中的图谋(scheming)行为,即智能体隐蔽地追求与人类不一致的目标。我们的重点在于理解图谋行为如何由若干关键因素的交互作用而产生,例如工具性目标、环境可供性、监督条件以及感知到的后果。先前的工作仅考察了少量场景,限制了分离这些条件如何塑造智能体产生图谋行为的倾向或能力。这种有限的规模和任务多样性也限制了对真实部署环境的覆盖以及可观察到的图谋策略的范围。为此,我们提出了SCHEMEARENA,一个包含400个场景、用于可扩展图谋压力测试的基准,其通过一个因子化的场景合成框架构建,涵盖多种安全相关的工具领域、工具性目标、监督条件和压力机制。为实现可扩展且可靠的监测,我们进一步提出了SCOUT,一种图谋行为监测器,它将多准则判断建立在从智能体推理和行动中提取的证据之上。在五个LLM智能体上的受控压力测试中,我们发现显式的工具性目标是图谋倾向最强的驱动因素。策略性提示发挥着独特作用,帮助智能体将图谋推理转化为具体的隐蔽行为。监督的影响则好坏参半:在若干闭源模型中,仅监测行动反而增加了图谋行为,这表明部分监督可能起到优化约束而非威慑作用。思维链(CoT)是一种有用但不完整的监测信号:它可以在执行前揭示潜在的图谋,但仅行动的图谋行为表明,隐蔽行为可能在没有显式推理证据的情况下发生。我们在以下地址发布基准、代码和监测器:https://github.com/launchnlp/SchemeArena。
cs.AI / 143 / 2609.08149

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

SWE-Bench Pro Verified:一个可靠的软件工程智能体基准测试
Zheng, Pujun, Shang, Zixin, Jiang, Shufan, Tian, Wenhui, Zhu, Dongsheng, Ma, Zerun, Yuan, Dingbo, Zhang, Qi
Abstract
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: \textbf{reward hacking}, enabled by leakage of gold solutions or hidden evaluation information, and \textbf{task quality issues}, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present \textbf{SWE-Bench Pro Verified}, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines \textbf{anti-hacking} safeguards that eliminate major leakage channels without disrupting normal agent functionality, with \textbf{task refinement} that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents.
Chinese Translation
SWE-Bench Pro 已成为评估软件工程智能体在具有挑战性的仓库级任务上表现的标准基准。然而,我们的分析工作表明,其评估受到两类不可靠性因素的影响:一是“奖励作弊”,由标准答案或隐藏评估信息的泄露导致;二是“任务质量问题”,包括具有误导性的问题描述和范围不当的测试。这些问题可能夸大基准测试性能并掩盖智能体的真实编程能力。我们提出了 SWE-Bench Pro Verified,这是 SWE-Bench Pro 的一个经过验证的版本,能够同时解决上述两个问题。我们的方法结合了防作弊保障机制和任务精化策略:前者在不干扰智能体正常功能的前提下消除主要的泄露渠道;后者以最小的改动修正缺陷实例中的不一致之处。在 SWE-Bench Pro Verified 上的评估结果显示,一些模型的表现明显低于此前报道的水平,这表明现有的 SWE-Bench Pro 结果可能高估了真实的软件工程能力。SWE-Bench Pro Verified 为评估软件工程智能体提供了一个更加可信的基准。
cs.AI / 144 / 2609.08162

WorldAgen: Unified State-Action Prediction with Test-Time World Model Training

WorldAgen:基于测试时世界模型训练的统一状态-动作预测
Wan, Chi, Wang, Kangrui, Si, Yuan, Zhang, Pingyue, Li, Manling
Abstract
How can vision-language-action (VLA) models adapt to new environments where world dynamics shift? While recent research has combined world modeling and action prediction to improve VLA performance, existing methods largely rely on pretraining on static datasets, without mechanisms for active adaptation at deployment time. As a result, these models often fail to generalize when deployed in unseen scenarios with novel object configurations or dynamics. We present WorldAgen, a unified framework that jointly learns world modeling and action prediction while enabling Test-Time Training (TTT) to adapt to new environments. WorldAgen employs a shared Transformer backbone with two heads: (1) a world model head that predicts future states from past state-action trajectories, and (2) an agent model head that predicts actions conditioned on task instructions. We design a Mixed Unidirectional Attention Mask to separate these two models. During test time, WorldAgen samples exploratory actions, collects ground-truth state transitions, and performs lightweight TTT updates to refine its world model. This adaptation improves the model's understanding of the environment and leads to more accurate action predictions. Experiments on the CALVIN and LIBERO benchmarks demonstrate that our baseline model achieves comparable, and in some cases superior, performance to current state-of-the-art approaches. Moreover, with TTT on a small number of samples, our method surpasses existing state-of-the-art models, highlighting the effectiveness of adapting world models at inference time.
Chinese Translation
视觉-语言-动作(VLA)模型如何适应世界动态发生变化的新环境?尽管近期研究已将世界建模与动作预测相结合以提升VLA性能,但现有方法主要依赖于在静态数据集上进行预训练,缺乏在部署时进行主动适应的机制。因此,这些模型在面对具有新颖物体配置或动态特性的未见场景时,往往难以泛化。我们提出了WorldAgen,这是一个统一框架,能够联合学习世界建模与动作预测,并通过测试时训练(Test-Time Training, TTT)来适应新环境。WorldAgen采用共享的Transformer骨干网络,配有两个预测头:(1) 世界模型头,根据过去的状态-动作轨迹预测未来状态;(2) 智能体模型头,基于任务指令预测动作。我们设计了混合单向注意力掩码来分离这两个模型。在测试阶段,WorldAgen采样探索性动作,收集真实的状态转移数据,并进行轻量级的TTT更新以优化其世界模型。这种适应机制提升了模型对环境的理解,从而带来更准确的动作预测。在CALVIN和LIBERO基准上的实验表明,我们的基线模型取得了与当前最先进方法相当、甚至在某些情况下更优的性能。此外,仅需在少量样本上进行TTT,我们的方法便能超越现有的最先进模型,凸显了在推理阶段适应世界模型的有效性。
cs.AI / 145 / 2609.08173

Key Path Identification for Resolving Knowledge Conflicts via SAE-based Steering

基于SAE引导的关键路径识别方法用于解决知识冲突
Zhang, Wenbo, Sun, Zhongxiang, Han, Zhiguang, Xu, Jun
Abstract
Sparse autoencoder (SAE)-based steering has been widely used to address knowledge conflicts by guiding LLMs to be more faithful to the contextual knowledge. Existing methods usually perform mass steering, which modifies a large batch of SAE features identified via correlation-based methods. However, due to the inaccurate correlation and the neglected feature interactions, mass steering methods fail to precisely identify the features that play the key roles in steering and introduce a large number of redundant ones, which add noise and weaken the steering effects. Our empirical studies reveal that steering only a small subset of the identified features can achieve comparable or even better performance. Motivated by this finding, we propose Key Path Identification (KPI), a novel method that identifies key steering features characterized by strong causal dependencies with both upstream and downstream features. From these features, KPI constructs key paths and steers through less feature modifications. In this way, KPI advances SAE-based steering from quantity-driven to quality-focused, offering a perspective for more precise and interpretable model editing. Experiments in RAG tasks with knowledge conflicts show that our method improves the accuracy by 18% on average compared to the best baseline of mass steering, effectively filtering redundant features, alleviating side effects and demonstrating the core role of key paths in steering.
Chinese Translation
基于稀疏自编码器(SAE)的引导方法已被广泛用于解决知识冲突,通过引导大语言模型(LLM)更加忠实于上下文知识。现有方法通常采用大规模引导(mass steering),即对通过基于相关性方法识别出的大量SAE特征进行修改。然而,由于相关性不够准确以及忽略了特征之间的交互作用,大规模引导方法无法精确识别在引导中起关键作用的特征,同时引入了大量冗余特征,这些冗余特征增加了噪声并削弱了引导效果。我们的实证研究表明,仅引导所识别特征中的一小部分即可达到相当甚至更好的性能。受此发现启发,我们提出了关键路径识别(Key Path Identification, KPI)方法,这是一种新颖的方法,用于识别与上游和下游特征均具有强因果依赖关系的关键引导特征。基于这些特征,KPI构建关键路径,并通过更少的特征修改实现引导。通过这种方式,KPI将基于SAE的引导从数量驱动推进到质量聚焦,为更精确、更可解释的模型编辑提供了新视角。在存在知识冲突的检索增强生成(RAG)任务上的实验表明,与最优的大规模引导基线方法相比,我们的方法平均提升了18%的准确率,有效过滤了冗余特征,缓解了副作用,并证明了关键路径在引导中的核心作用。
cs.AI / 146 / 2609.08174

OntologyBench: Can Dense Retrieval Satisfy Structured Biomedical Constraints?

OntologyBench:稠密检索能否满足结构化生物医学约束?
Zhang, Xiao Yu Cindy, Wasserman, Wyeth, Zhu, Jian
Abstract
We introduce OntologyBench, a tiered biomedical retrieval benchmark comprising 471,854 training and 125,744 evaluation query-document relevance pairs across concept grounding, relational retrieval, and compositional phenotype-based retrieval. Although these tasks can be tractable using ontology-aware reference methods, across task tiers, embedding performance is generally lower on relational and compositional tasks than on concept-grounding tasks. Fine-tuning on ontology-derived supervision improves performance on several relational and compositional tasks, whereas the evaluated reranking and LLM-based candidate-scoring methods provide little or no end-to-end improvement. Errors frequently reflect diseases matching only subsets of the phenotype evidence. These findings indicate that the evaluated embedding and reranking configurations do not reliably recover the compatibility encoded by the selected ontology relations and phenotype combinations and motivate retrieval systems that better integrate learned representations with structured biomedical knowledge.
Chinese Translation
我们提出了OntologyBench,这是一个分层的生物医学检索基准,包含471,854对训练和125,744对评估用的查询-文档相关性数据,涵盖概念定位(concept grounding)、关系检索和基于组合式表型的检索三类任务。尽管借助本体感知的参考方法这些任务是可以处理的,但在各任务层级中,嵌入模型在关系型和组合型任务上的表现普遍低于概念定位任务。在由本体导出的监督信号上进行微调可以提升多个关系型和组合型任务的性能,而所评估的重排序(reranking)和基于大语言模型(LLM)的候选打分方法几乎没有带来端到端的改进。错误往往源于仅匹配到表型证据子集的疾病。这些发现表明,所评估的嵌入和重排序配置无法可靠地恢复由所选本体关系和表型组合所编码的兼容性,也促使我们开发能更好地将学习到的表示与结构化生物医学知识相结合的检索系统。
cs.AI / 147 / 2609.08175

Safe Harness Self-Evolution: A Theoretical Analysis of Feasibility and Limits

安全 Harness 自我演化:可行性与极限的理论分析
Cai, Qianshu, Zhang, Yonggang, Nie, Jun, Ran, Maohao, Zheng, Huajiang, Song, Jun, Tian, Xinmei, Guo, Yike, Xue, Wei
Abstract
Harness self-evolution is the process by which an agent modifies its prompts, tools, code, or orchestration in response to task feedback while keeping the underlying language model frozen, with changes persisting across subsequent tasks. We provide a systematic theoretical analysis of the feasibility and limits of safe harness self-evolution, connecting modification generation, finite-data certification and selection, safe adoption, and behavior after an update. Under a fixed user-task distribution, we establish conditions guaranteeing overall expected-reward improvement while controlling changes on retained tasks, characterize the probability of generating qualified modifications, and derive finite-data bounds for safe selection and adoption. Our analysis shows that generation and certification impose distinct constraints: current task performance does not determine the probability of generating qualified modifications, and generating more candidates need not improve the guarantee of a successful update when evaluation is limiting. Stagnation may therefore arise even when improvement opportunities remain. We further show that worst-case evaluation cost for recognizing genuine improvements diverges as expected reward approaches its upper bound. Across successive updates, certified improvement guarantees accumulate over a finite run, but a successful update does not by itself guarantee that further improvement remains possible. These results provide a basis for diagnosing bottlenecks and designing safer self-evolution mechanisms.
Chinese Translation
Harness 自我演化是指智能体(agent)在保持底层语言模型冻结的前提下,根据任务反馈修改其提示词、工具、代码或编排流程,并使这些修改在后续任务中持续生效的过程。我们对安全 harness 自我演化的可行性与极限进行了系统的理论分析,将修改生成、有限数据下的认证与选择、安全采纳以及更新后的行为联系起来。在固定的用户任务分布下,我们建立了在控制对保留任务影响的同时保证整体期望奖励提升的条件,刻画了生成合格修改的概率,并推导了安全选择与采纳的有限数据界。我们的分析表明,生成与认证施加了不同的约束:当前任务表现并不决定生成合格修改的概率;且当评估成为瓶颈时,生成更多候选修改未必能提高成功更新的保证。因此,即使改进机会仍然存在,也可能出现停滞。我们进一步证明,随着期望奖励趋近其上界,识别真正改进的最坏情况评估成本会发散。在连续多次更新过程中,经过认证的改进保证在有限次运行中可以累积,但一次成功的更新本身并不能保证进一步的改进仍然可能。这些结果为诊断瓶颈和设计更安全的自我演化机制提供了基础。
cs.AI / 148 / 2609.08180

Less Is Personal: Learning Minimal Sufficient User Profiles for Personalized Language Models

少即是个性化:为个性化语言模型学习最小充分用户画像
Liu, Minghang, Qiu, Qiang, Wang, Yuanzhuo, Shen, Huawei, Cheng, Xueqi
Abstract
Retrieval-augmented personalization enables large language models to produce more accurate and preference-aligned outputs using relevant records retrieved from user histories. Personalized language models typically prepend a fixed number of retrieved user records, even when additional history is redundant, harmful, or unrelated to a user's distinctive behavior. We study minimal sufficient personalization: constructing the least costly ordered profile for each input while preserving the utility achievable from a retrieved candidate pool. We introduce ENOUGH, a method that iteratively appends behavioral records or emits STOP to construct profiles with adaptive lengths. Offline, bounded counterfactual search evaluates profile prefixes by jointly considering downstream gains, user specificity, and token costs. The resulting long-horizon targets are distilled into a multi-head value controller with explicit ranking and stopping supervision. At inference, the controller selects and orders records through lightweight decisions, and the frozen generator is invoked once after stopping. Extensive experiments on six personalized tasks demonstrate that ENOUGH consistently outperforms strong heuristic and retrieval-augmented baselines in both effectiveness and efficiency, achieving minimal sufficient profiles that preserve personalization utility while reducing unnecessary context costs.
Chinese Translation
检索增强的个性化技术使大型语言模型能够利用从用户历史中检索到的相关记录,生成更准确且更符合用户偏好的输出。个性化语言模型通常会在输入前拼接固定数量的检索到的用户记录,即使额外的历史信息是冗余的、有害的或与用户的独特行为无关。我们研究了最小充分个性化问题:在保持从检索候选池中可获得的效用的前提下,为每个输入构建代价最低的有序用户画像。我们提出了ENOUGH方法,该方法通过迭代地追加行为记录或输出STOP信号,构建具有自适应长度的用户画像。在离线阶段,有界反事实搜索通过综合考量下游收益、用户特异性和令牌成本来评估画像前缀。所得到的长时程目标被蒸馏到一个具有显式排序和停止监督的多头价值控制器中。在推理阶段,该控制器通过轻量级决策对记录进行选择和排序,冻结的生成器仅在停止后调用一次。在六个个性化任务上的大量实验表明,ENOUGH在效果和效率上均持续优于强启发式基线和检索增强基线,实现了在保持个性化效用的同时降低不必要上下文成本的最小充分用户画像。
cs.AI / 149 / 2609.08186

Does Deeper Reasoning Compromise Alignment? Revealing and Mitigating of Alignment Collapse in Large Reasoning Models

更深入的推理是否会损害对齐?揭示并缓解大型推理模型中的对齐坍塌问题
Wu, Yu-Hang, Xiong, Yu-Jie, Zhang, Henghua, Zhang, Bairui, Zhang, Jia-Chen, Li, Shaohua
Abstract
The emergence of Chain-of-Thought (CoT) has established a robust foundation for Large Reasoning Models (LRMs). While deep reasoning is widely believed to enhance safety alignment, the stability of alignment mechanisms under extended reasoning remains underexplored. This paper challenges the prevailing view by revealing a critical vulnerability: Deep Reasoning May Induce Alignment Collapse. To rigorously quantify this phenomenon, we propose the Alignment Loss Rate (ALR) metric. Our experiments demonstrate that as reasoning depth increases, ALR rises significantly, indicating a severe degradation in model robustness against external perturbations. Capitalizing on this instability, a novel jailbreaking paradigm, Reasoning Trap (RT), is proposed. RT induces the model into extended reasoning to amplify the impact of adversarial attacks, leading to a sharp decline in safety capabilities. To elucidate the mechanism behind this collapse, we identify Attention Dilution as the root cause, arising from the competition for attention between the extended reasoning process and the original input. To mitigate this, Reasoning Residual Alignment (RRA), a lightweight defense strategy that dynamically re-emphasizes the input via residual connections integrated with the reasoning process.
Chinese Translation
思维链的出现为大型推理模型奠定了坚实的基础。尽管深度推理被普遍认为能够增强安全对齐,但扩展推理下对齐机制的稳定性仍未得到充分研究。本文通过揭示一个关键脆弱性挑战了这一主流观点:深度推理可能诱发对齐坍塌。为了严格量化这一现象,我们提出了对齐损失率(Alignment Loss Rate, ALR)指标。实验表明,随着推理深度的增加,ALR显著上升,表明模型抵御外部扰动的能力严重退化。利用这种不稳定性,我们提出了一种新颖的越狱范式——推理陷阱。该范式诱导模型进入扩展推理以放大对抗攻击的影响,导致模型安全能力急剧下降。为阐明这种坍塌背后的机制,我们识别出注意力稀释是其根本原因,即扩展推理过程与原始输入之间对注意力的竞争所致。为缓解这一问题,我们提出了推理残差对齐,这是一种轻量级防御策略,通过与推理过程相集成的残差连接动态地重新强调输入。
cs.AI / 150 / 2609.08188

Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment

通过生成器在环对齐弥合多模态RAG中的语义-效用鸿沟
Chang, Zhan-Lun, Han, Dong-Jun, Hosseinalipour, Seyyedali, Chiang, Mung, Brinton, Christopher G.
Abstract
Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to external evidence. However, standard retrievers and rerankers optimize for semantic similarity rather than answer utility, creating a preference gap: documents that appear relevant may not help the generator produce a correct answer. Motivated by this, we propose a two-stage generator-in-the-loop alignment framework that closes this gap without human document-level relevance annotations. Our framework consists of two stages: in Stage 1, a VLM generates a hypothetical text passage from the image-query pair, which is used as the retrieval query for dense text search, bridging the image-to-text modality gap. In Stage 2, a cross-encoder reranker adapted with low-rank adaptation (LoRA) is fine-tuned using answer-supervised preference pairs mined from the frozen VLM: given the dataset answer label, a candidate document is labeled positive if the VLM produces the correct answer when given that document as context, and negative otherwise. This generator-guided signal is compatible with multiple alignment loss functions, including contrastive (triplet) loss, pairwise direct preference optimization (DPO), and supervised fine-tuning (SFT), and supports periodic re-mining to refresh preference pairs as the reranker improves. Experiments on VQA-X and A-OKVQA with Qwen3.5-2B and Qwen3-VL-4B-Instruct show that our proposed framework consistently outperforms rank-order, random, and REPLUG-style likelihood baselines under various alignment losses and pool size settings, suggesting that answer-level generator feedback is an effective supervision signal for preference alignment.
Chinese Translation
借助检索增强生成(RAG)的视觉语言模型(VLM)能够受益于对外部证据的访问。然而,标准的检索器和重排序器优化的是语义相似度而非答案效用,从而产生偏好鸿沟:看起来相关的文档未必能帮助生成器产生正确答案。受此启发,我们提出了一个两阶段的生成器在环对齐框架,无需人工文档级相关性标注即可弥合该鸿沟。我们的框架包含两个阶段:在第一阶段,VLM根据图像-查询对生成一个假设性文本段落,将其作为稠密文本检索的查询,从而弥合图像到文本的模态鸿沟。在第二阶段,使用从冻结VLM中挖掘的答案监督偏好对,对经低秩适应(LoRA)调整的交叉编码器重排序器进行微调:给定数据集的答案标签,若VLM在以某个候选文档为上下文时能产生正确答案,则该文档被标记为正例,否则标记为负例。这种生成器引导的信号与多种对齐损失函数兼容,包括对比(三元组)损失、成对直接偏好优化(DPO)以及监督微调(SFT),并支持周期性重新挖掘以在重排序器改进时刷新偏好对。在VQA-X和A-OKVQA数据集上,使用Qwen3.5-2B和Qwen3-VL-4B-Instruct进行的实验表明,在各种对齐损失和候选池大小设置下,我们提出的框架均持续优于排序、随机以及REPLUG式似然基线,说明答案级的生成器反馈是偏好对齐的有效监督信号。
cs.AI / 151 / 2609.08189

Do Dynamic Routers Need Memory? HeRo: History-Aware Routing for Efficient LLM Inference

动态路由器需要记忆吗?HeRo:面向高效大语言模型推理的历史感知路由
Lin, Hongjin, Wan, Wentao, Wang, Keze
Abstract
Dynamic layer routing reduces the inference cost of Large Language Models (LLMs) by learning to skip layers for individual tokens. Existing methods, however, treat each routing decision as a local operation conditioned solely on the current hidden state which is a formulation that overlooks the sequential, path-dependent nature of routing across depth: earlier decisions shape the representations seen by downstream routers, and the layer-usage objective couples all decisions jointly. We propose History-Aware Routing (HeRo), a dynamic routing framework that resolves this mismatch by introducing a router memory mechanism to maintain an explicit routing state across model depth. The memory is constructed via linear attention, incrementally aggregating preceding routing scores and their induced residual updates into a compact history representation. At each routed layer, the router conditions jointly on this accumulated state and the current hidden representation to select the executed branch. Instantiated for token-wise FFN routing, HeRo trains only lightweight routers and adapters on a frozen backbone, requiring no modification to pretrained parameters. Across Llama 3.1-8B, Llama 2-7B, and Llama 2-13B, HeRo consistently achieves the highest aggregate performance retention among ten baselines. On Llama 3.1-8B, it bypasses 26.87% of model parameters while achieving 100.24% of dense model performance across seven benchmarks, and retains 97.01% while bypassing 38.82% of model parameters under a tighter computation budget. Ablation studies confirm that removing routing history consistently degrades performance, most notably on multistep reasoning and code generation, validating that explicit routing memory enables more accurate and adaptive dynamic routing than solely conditioning on hidden state.
Chinese Translation
动态层路由通过学习对单个词元跳过某些层来降低大语言模型(LLM)的推理成本。然而,现有方法将每次路由决策视为仅以当前隐状态为条件的局部操作,这一建模方式忽视了跨深度路由的序列性与路径依赖性:较早的决策会影响下游路由器所看到的表示,且层使用目标将所有决策耦合在一起。我们提出历史感知路由(History-Aware Routing, HeRo),这是一个动态路由框架,通过引入路由器记忆机制在模型深度上维护显式的路由状态,从而解决上述不匹配问题。该记忆通过线性注意力构建,将先前的路由分数及其引起的残差更新增量式地聚合为紧凑的历史表示。在每个路由层,路由器联合考虑这一累积状态和当前隐表示来选择执行的分支。针对逐词元FFN路由的实现中,HeRo仅在冻结的主干上训练轻量级路由器和适配器,无需修改预训练参数。在Llama 3.1-8B、Llama 2-7B和Llama 2-13B上,HeRo在十个基线方法中始终取得最高的整体性能保持率。在Llama 3.1-8B上,它跳过了26.87%的模型参数,同时在七个基准上达到稠密模型100.24%的性能;在更紧张的计算预算下,跳过38.82%的模型参数时仍保持97.01%的性能。消融实验证实,移除路由历史会一致地降低性能,在多步推理和代码生成任务上尤为明显,验证了显式路由记忆相比仅依赖隐状态的条件化方式能够实现更准确、更具适应性的动态路由。
cs.AI / 152 / 2609.08196

Qiushi Engine on AstaBench E2E-Bench-Hard

求实引擎在 AstaBench E2E-Bench-Hard 上的评测
Li, Wenhao, Yang, Shuxing, Chen, Fujia, Mi, Jincheng, Pan, Yuang, Zhao, Rui, Li, Zichen, Wu, Junyao, Hong, Shenzhan, Li, Yaqi, Wang, Yize, Zhu, Kaihao, Deng, Taowen, Yang, Junjie, Chen, Hongsheng, Yang, Yihao
Abstract
This report analyzes Qiushi Engine v0.8 across all 40 test tasks in AstaBench E2E-Bench-Hard, a benchmark that requires autonomous agents to carry a research question through experimental design, code implementation, actual execution, result analysis, and report delivery. Qiushi Engine is model-configurable; this evaluation selected DeepSeek deepseek-v4pro-preview as the model backend. The official AstaBench leaderboard records a score of 0.816 and an average benchmark cost of USD 15.209 per task, while the full-precision local recomputation is $81.59 \pm 1.87$. Four tasks satisfied every rubric item, yielding a full-task completion rate of 4/40 = 10% -- 7 percentage points above, and about 3.3 times, the approximately 3% best rate reported for AstaBench's official agents. Across 507 required rubric items, 416 were satisfied (82.1%). Official scoring archives and 40 Meta-Trace records show sustained production and verification of reports, code, and experimental artifacts; the principal gaps lie in repeated runs, external dependencies, specified metrics, and ablation studies. The report explains the benchmark, system workflow, aggregate results, representative cases, and limits of interpretation.
Chinese Translation
本报告分析了求实引擎(Qiushi Engine)v0.8 在 AstaBench E2E-Bench-Hard 基准测试中全部 40 个任务上的表现。该基准要求自主智能体(agent)完成从实验设计、代码实现、实际执行、结果分析到报告交付的完整科研流程。求实引擎支持模型可配置,本次评测选用 DeepSeek deepseek-v4pro-preview 作为模型后端。AstaBench 官方排行榜记录的得分为 0.816,平均基准成本为每任务 15.209 美元;全精度本地复算结果为 $81.59 \pm 1.87$。其中四个任务满足全部评分细则项,完整任务完成率为 4/40 = 10%,比 AstaBench 官方智能体所报告的约 3% 的最佳完成率高出 7 个百分点,约为其 3.3 倍。在全部 507 个必需的评分细则项中,有 416 项得到满足(82.1%)。官方评分存档和 40 条 Meta-Trace 记录显示,系统在报告、代码和实验产物的持续生成与验证方面表现稳定;主要差距集中在重复运行、外部依赖、指定指标以及消融实验等方面。本报告阐述了该基准测试、系统工作流程、总体结果、代表性案例以及结果解释的局限性。
cs.AI / 153 / 2609.08211

A Better Spur Should Start From Each Objective

更好的激励应始于每个目标
Mao, Shanwen, Zhang, Hao, nie, Guangtao, Li, Zhiheng, Wang, Huimu, Xu, Sulong, Simiu, Gu
Abstract
Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward conflicts, and late-stage reward tug-of-war, causing traditional linear scalarization to experience severe metric oscillations. To address optimization conflicts among multiple objectives in real-world deployment scenarios, we propose Multi-Marginal Preference Optimization (MMPO), a fine-grained framework that intervenes at the data, gradient, and constraint levels rather than relying on coarse-grained global scalarization. Specifically, MMPO performs exposure debiasing to mitigate sparse and biased rewards, applies priority-aware orthogonal projection to decouple conflicting gradients, and introduces self-prompted gradient constraints to prevent dominant objectives from overwhelming weaker ones. Experiments on real-world e-commerce datasets show that MMPO improves training stability and consistently achieves better performance across conflicting metrics. Moreover, it generalizes robustly to broader tasks such as ToolRL and code generation, demonstrating its effectiveness as a practical paradigm for multi-objective alignment.
Chinese Translation
现实世界的多目标强化学习(Multi-Objective Reinforcement Learning, MORL)常常面临奖励稀疏、奖励冲突以及后期奖励“拉锯战”等问题,导致传统的线性标量化方法出现严重的指标震荡。为了解决现实部署场景中多个目标之间的优化冲突,我们提出了多边际偏好优化(Multi-Marginal Preference Optimization, MMPO),这是一种细粒度框架,从数据、梯度和约束层面进行干预,而不是依赖于粗粒度的全局标量化。具体而言,MMPO 通过曝光去偏来缓解奖励稀疏和偏差问题,采用优先级感知的正交投影来解耦冲突梯度,并引入自适应梯度约束以防止主导目标压制弱势目标。在真实电商数据集上的实验表明,MMPO 提升了训练稳定性,并在相互冲突的指标上持续取得更好的性能。此外,该方法还能稳健地泛化到 ToolRL 和代码生成等更广泛的任务中,证明了其作为一种实用的多目标对齐范式的有效性。
cs.AI / 154 / 2609.08216

Vision: Data-Centric Anchoring for Robust and Interpretable Agentic AI

愿景:面向鲁棒且可解释的智能体AI的数据中心化锚定
Malarkkan, Arun Vignesh, Wang, Xinyuan, Fu, Yanjie
Abstract
Agentic AI systems built on large language models fail in two persistent ways that scaling does not fix: they break under distribution shift, and they cannot explain the decisions they make. We argue these are co-symptoms of one structural deficiency in the data lifecycle that governs how agents are trained, evaluated, and deployed. Observational interaction logs record what an agent did, not what it would have done otherwise. They encode spurious correlations without controlled variation, so they lack the counterfactual structure needed to separate causal signal from coincidence or to validate an explanation. No model-centric method can recover invariances the data never contained. We present Data-Centric Anchoring: robustness and interpretability should be engineered into the data environment, not extracted from models after training. Our central contribution is the Data-Centric Agentic Loop, a four-stage framework of Curate, Augment, Constrain, and Attribute. The ordering is structural, not stylistic. Curation precedes augmentation because generative models amplify whatever bias they are trained on. Augmentation precedes constraint because invariance objectives are vacuous without variation across environments to be invariant to. Attribution closes the loop, converting observed failures into targeted data interventions for the next iteration. Each stage manufactures the preconditions of the next, which makes the loop self-correcting rather than merely sequential. We ground the framework in a failure-driven taxonomy that links four core failure modes to the data lifecycle: spurious feature reliance, distribution-shift fragility, uncertainty miscalibration, and explanation unfaithfulness. We close with the limits of this approach and the open problems that stand between it and practical deployment at scale.
Chinese Translation
基于大语言模型构建的智能体AI系统存在两种通过扩大规模也无法解决的持久性缺陷:它们在分布偏移下会失效,且无法解释自身所做出的决策。我们认为,这两者是支配智能体训练、评估与部署的数据生命周期中同一结构性缺陷的共存症状。观测性交互日志记录的是智能体实际做了什么,而非它本来可能做什么。这些日志在没有受控变化的情况下编码了虚假相关性,因此缺乏将因果信号与巧合区分开来或验证解释所需的反事实结构。任何以模型为中心的方法都无法恢复数据中从未包含的不变性。我们提出“数据中心化锚定”(Data-Centric Anchoring):鲁棒性与可解释性应当在数据环境中被设计构建,而非在训练之后从模型中提取。我们的核心贡献是“数据中心化智能体循环”(Data-Centric Agentic Loop),这是一个由策划(Curate)、增强(Augment)、约束(Constrain)和归因(Attribute)组成的四阶段框架。这一排序是结构性的,而非风格性的。策划先于增强,因为生成模型会放大其训练数据中的任何偏差;增强先于约束,因为若缺乏跨环境的变化作为不变的基础,不变性目标便是空洞的;归因则闭合了整个循环,将观察到的失败转化为下一轮迭代中针对性的数据干预。每个阶段都为下一阶段制造前提条件,这使得该循环是自我修正的,而不仅仅是顺序执行的。我们以一个失败驱动的分类体系为基础,将四种核心失败模式与数据生命周期相关联:虚假特征依赖、分布偏移脆弱性、不确定性校准失准以及解释不忠实性。最后,我们讨论了该方法的局限性,以及在其实现大规模实际部署之前尚待解决的开放性问题。
cs.AI / 155 / 2609.08226

TTGBench: Benchmarking Topological Evolution and Semantic Drift in Text-attributed Temporal Graphs

TTGBench:面向文本属性时序图中的拓扑演化与语义漂移的基准测试
Ma, Longfei, Liu, Zemin, Wu, Fei
Abstract
Temporal graph learning models the evolution of dynamic systems, where both structural interactions and semantic states change over time. However, existing benchmarks primarily emphasize structural evolution via temporal link prediction (TLP), while support for semantic evolution remains limited. Although temporal node classification (TNC) is sometimes included, it is typically restricted to simplistic binary settings that fail to capture realistic semantic drift. Moreover, commonly used datasets exhibit high link repetition, leading to inflated performance estimates and obscuring true model capability. To address these limitations, we introduce \textbf{TTGBench}, a new benchmark that jointly evaluates structural and semantic evolution. TTGBench comprises six real-world, text-rich datasets characterized by \emph{Dual Volatility}, enabling rigorous and fair evaluation of existing models. Notably, it is the first benchmark to support both multi-class and multi-label TNC, filling a critical gap in evaluating temporal semantic drift. We conduct a comprehensive evaluation of 17 state-of-the-art methods across Temporal Graph Neural Networks (TGNNs) and Large Language Model (LLM)-based paradigms. The results reveal a clear \emph{capability divide} between the two paradigms: TGNN-based methods excel at structural prediction but fail at semantic tracking, whereas LLM-based predictors show the opposite trend. Through in-depth analysis, we uncover their fundamental limitations and provide insights for developing more comprehensive temporal graph models.
Chinese Translation
时序图学习旨在建模动态系统的演化过程,其中结构交互和语义状态都会随时间发生变化。然而,现有基准主要通过时序链路预测(TLP)侧重于结构演化,而对语义演化的支持仍然有限。尽管有时也包含时序节点分类(TNC)任务,但通常局限于简单的二分类设置,无法刻画真实的语义漂移。此外,常用数据集普遍存在较高的链路重复率,导致性能评估被夸大,从而掩盖了模型的真实能力。为解决上述局限性,我们提出了 extbf{TTGBench},一个联合评估结构与语义演化的新基准。TTGBench 包含六个具有"双重波动性"(Dual Volatility)特征的文本丰富的真实世界数据集,能够对现有模型进行严格而公平的评估。值得注意的是,它是首个同时支持多分类和多标签时序节点分类任务的基准,填补了评估时序语义漂移方面的关键空白。我们对 17 种最先进的方法进行了全面评估,涵盖时序图神经网络(TGNN)和基于大语言模型(LLM)的范式。结果揭示了两种范式之间清晰的"能力鸿沟":基于 TGNN 的方法擅长结构预测但无法进行语义追踪,而基于 LLM 的预测器则呈现相反的趋势。通过深入分析,我们揭示了它们各自的根本局限性,并为开发更全面的时序图模型提供了启示。
cs.AI / 156 / 2609.08228

SE-GoS: Self-Evolving Graph-of-Skills for Skill Library at Scale

SE-GoS:面向大规模技能库的自演化技能图
Fu, Dawei, Jiang, Cheng, Qian, Sitian, Wang, Huainan, Hao, Zhongkai
Abstract
Modern LLM agents increasingly rely on reusable skills, yet as skill libraries scale to thousands of entries, effective retrieval becomes a bottleneck. Graph-of-Skills (GoS) addresses this challenge by exploiting dependency-aware graph structure for scalable skill retrieval, while SkillDAG further demonstrates that skill graphs can accumulate execution-backed structure online. However, these approaches leave open whether historical execution traces can be systematically distilled into a better retrieval graph that generalizes to unseen tasks. We present Self-Evolving Graph-of-Skills (SE-GoS), a training-free framework that evolves an existing GoS graph from execution traces while preserving the original retrieval pipeline. SE-GoS performs three complementary updates: topology evolution that discovers and prunes skill relationships from execution evidence, edge-weight evolution that reinforces retrieval-relevant relationships based on historical effectiveness, and description evolution that optimizes retrieval-facing skill descriptions using execution feedback. Across three LLMs on SkillsBench, SE-GoS consistently improves task reward while reducing input tokens relative to full skill loading, with gains varying across model families. In a representative setting, one evolution round improves reward from 52.4\% to 59.4\% while reducing input tokens by approximately one-third relative to full skill loading, and the resulting graph transfers to a disjoint held-out split with a 5.4-point improvement over the static GoS baseline. These results show that skill graphs can be improved from execution experience without model training, changes to the retrieval algorithm, or modifications to skill content, turning a static retrieval graph into an evolving retrieval infrastructure.
Chinese Translation
现代大语言模型(LLM)智能体日益依赖可复用的技能,然而随着技能库扩展至数千个条目,有效检索成为瓶颈。技能图(Graph-of-Skills, GoS)通过利用依赖感知的图结构实现可扩展的技能检索来应对这一挑战,而 SkillDAG 进一步表明技能图可以在线积累基于执行的结构。然而,这些方法尚未回答一个开放问题:能否将历史执行轨迹系统地蒸馏为一个更好的检索图,并使其泛化到未见过的任务上。我们提出自演化技能图(Self-Evolving Graph-of-Skills, SE-GoS),这是一个无需训练的框架,能够从执行轨迹出发演化现有的 GoS 图,同时保留原始的检索流程。SE-GoS 执行三种互补的更新:拓扑演化——从执行证据中发现并剪除技能关系;边权重演化——基于历史有效性强化与检索相关的关系;描述演化——利用执行反馈优化面向检索的技能描述。在 SkillsBench 上跨三个 LLM 的实验中,SE-GoS 相较于完整技能加载方式持续提升任务奖励并减少输入 token 数,且增益在不同模型家族间存在差异。在一个代表性设置中,一轮演化将奖励从 52.4% 提升至 59.4%,同时相对于完整技能加载将输入 token 减少约三分之一;所得的图可迁移到一个不相交的保留测试集上,较静态 GoS 基线提升 5.4 个百分点。这些结果表明,无需模型训练、检索算法改动或技能内容修改,技能图即可从执行经验中得到改进,从而将静态检索图转变为一个不断演化的检索基础设施。
cs.AI / 157 / 2609.08236

Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts

形式胜过实质:内容不变的包装语可翻转LLM安全评判器的裁决
Zhou, Yongxi, Ye, Wenbo, Liu, Yuanzhe, Dong, Zihan, Yao, Junwei
Abstract
Automatic safety judges -- systems such as Llama Guard or a GPT-4o grading prompt that decide whether a model's reply is harmful -- produce the numbers behind almost every reported jailbreak success rate, defense evaluation, and safety leaderboard. We ask whether these judges grade what a reply contains or how it sounds. We keep a reply's content fixed and add content-invariant style wrappers: fixed strings placed before or after the reply that change only its tone (an educational disclaimer, a fake safety "reasoning" block, a token refusal followed by the unchanged harmful body), or, on harmless refusals, framing that merely sounds dangerous. The body is preserved byte-for-byte, so a faithful judge must return the same verdict, and any flip is an error of the judge, not a change in safety. Over 600 JailbreakBench replies x up to 7 forms x 8 judges, we measure flip rates with paired significance tests and measured noise floors. Findings are precise rather than universal: most judges barely move, but specific judges harbor cheaply exploitable blind spots. A token-refusal wrapper flips 19.9% of GPT-4o-mini's correct "unsafe" verdicts (noise floor 0.5%; 18.2% under majority-of-three re-scoring) yet moves Claude only 0.4%. The deployed Llama Guard 4 is deterministically gamed: an "educational course" framing flips 12.3% of its harmful verdicts to safe. A second deployed guard (gpt-oss-safeguard-20b) is immune, and rewriting only the grading prompt (StrongREJECT-style) cuts the attack tenfold on the identical model -- the vulnerability lives in the judge, not the content. A two-annotator human validation confirms 100% content invariance and 90% of flips as judge errors (kappa 0.95-1.0), and a bootstrap shows the underlying model ranking is already unstable to sampling alone. We release the dataset, wrappers, code, and per-verdict labels.
Chinese Translation
自动安全评判器——例如Llama Guard或GPT-4o评分提示词等用于判定模型回复是否有害的系统——产生了几乎所有已报告的越狱成功率、防御评估和安全排行榜背后的数字。我们追问:这些评判器评估的是回复所包含的内容,还是其表达方式?我们保持回复内容不变,并添加内容不变的样式包装语:置于回复之前或之后的固定字符串,仅改变其语气(如教育性免责声明、伪造的安全"推理"块、先给出象征性拒绝随后原样保留有害正文的回复),或者对于无害的拒绝回复,添加仅听起来危险的表述框架。正文逐字节保持不变,因此忠实的评判器必须给出相同裁决,任何翻转都是评判器的错误,而非安全性的变化。在超过600条JailbreakBench回复、最多7种形式、8个评判器的实验中,我们通过配对显著性检验和实测噪声底线来度量翻转率。研究结果精细而非普遍适用:大多数评判器几乎不受影响,但特定评判器存在可被低成本利用的盲点。一个象征性拒绝包装语可翻转GPT-4o-mini中19.9%的正确"不安全"裁决(噪声底线为0.5%;在三选多数重评下为18.2%),而对Claude仅影响0.4%。已部署的Llama Guard 4可被确定性 地操纵:"教育课程"表述框架可将其12.3%的有害裁决翻转为安全。另一个已部署的防护系统(gpt-oss-safeguard-20b)则完全免疫,而且仅重写评分提示词(StrongREJECT风格)就能在相同模型上将攻击效果降低十倍——该漏洞存在于评判器本身,而非内容。双人标注的人类验证确认了100%的内容不变性,且90%的翻转为评判器错误(kappa为0.95-1.0);自举分析表明底层模型排名仅凭抽样就已不稳定。我们发布了数据集、包装语、代码和逐条裁决标签。
cs.AI / 158 / 2609.08247

zScore-N: A Neural Network for On-Chain Wallet Reputation Scoring

zScore-N:一种用于链上钱包声誉评分的神经网络
N, Girish G, Sahoo, Ashutosh, SP, Akshay, S, Gurukiran, Kandaswamy, Dhanashekar
Abstract
Wallet reputation scores decide who receives an airdrop, who can borrow, and who enters an allowlist across decentralised finance. They almost always begin as hand-written formulas: compositions of clamped logarithmic, linear and square-root transforms over behavioural features, with every threshold and point award set by hand. Such a formula is readable and deterministic, but it is piecewise and non-differentiable, it cannot improve as data accumulates, and it cannot distinguish a feature that is genuinely zero from one its pipeline failed to capture. We present zScore-N, the neural network that replaced ours in production. The formula served as its teacher: calibrated against 5,208,952 wallets sampled across 2019-2024 and verified to reproduce production output to within 2.3e-13, it supplies unlimited labelled training data at zero label noise. The trained network reproduces it to 0.58 points RMSE on the 1000-point scale (R^2 = 0.99997), against 2.25 for gradient-boosted trees and 28.04 for linear regression on identical features and splits. Trained with missing-value masks against uncorrupted targets, it halves the error that incomplete data introduces: at 10% feature-level missingness the formula drifts 51.4 points from its own complete-data output with a systematic -12.5 point bias, while the network drifts 17.9. The network carries the score at production scale, across a population of millions of wallets spanning six orders of magnitude in size and activity.
Chinese Translation
钱包声誉评分决定了谁能够领取空投、谁可以借款以及谁可以进入去中心化金融中的白名单。这些评分几乎总是始于手工编写的公式:即对行为特征进行截断对数、线性和平方根变换的组合,其中每个阈值和积分值均由人工设定。此类公式虽然可读且具有确定性,但它是分段且不可微的,无法随着数据积累而改进,也无法区分一个特征是真值为零还是其处理流程未能捕捉到该特征。我们提出了zScore-N,即在生产环境中取代了我们原有公式的神经网络。原公式充当了它的教师:在2019至2024年间采样的5,208,952个钱包上进行校准,并验证其复现生产输出的误差在2.3e-13以内,该公式可提供无限量的标注训练数据,且标注噪声为零。在相同的特征和数据划分下,训练后的神经网络在1000分制上以0.58的RMSE复现原公式(R^2 = 0.99997),而梯度提升树为2.25,线性回归为28.04。通过使用缺失值掩码并结合未损坏的目标进行训练,该网络将不完整数据引入的误差降低了一半:在10%的特征级缺失率下,原公式相对于其自身完整数据输出漂移了51.4分,并存在-12.5分的系统性偏差,而神经网络仅漂移17.9分。该网络在生产规模下承载评分计算,覆盖数百万个钱包,其规模和活跃度跨越六个数量级。
cs.AI / 159 / 2609.08248

Agentic ML Exploration (A-MLE) for Ads Ranking

面向广告排序的智能体化机器学习探索(A-MLE)
Gao, Erwin, Sunkara, Vinodh Kumar, Guan, Jingyi, Jia, Qinjin, Xu, Hangjun, Ji, Xiang, Wong, Sherman, Chavali, Surya Teja, Vaishnavi, Pratik, Pandhi, Aryan, Deng, Xiaoyu, Wang, Zhaodong, Inani, Samarth, Yang, Fan, Moberg, Jakob, Zu, Zoe, Bievre, Nicolas, Khenissi, Sami, Jaspal, Amit, Fakharizadi, Ehsan, Viswanathan, Srinidhi, Sun, Dorothy, Vanam, Abishek, Iyer, Sneha, Yadawad, Sheela, Chen, Wenjie, Nahum, Gaby, Gu, Junhua, Chu, Peter, Liu, Yucheng, Zhao, Xin, Cid, Vitor, Chen, Chaorong, Pappu, Vijay, Kumar, Ashwin, Chen, Wenlin, Schulte, Ben, Chandra, Deepak, Tewari, Ritwik
Abstract
Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iteration - the cycles of research, implementation, training, debugging, evaluation, and launch required to surface a single statistically significant improvement. A typical ranking stack contains numerous differentiated models with heterogeneous data, architectures, and infrastructure constraints, and each cycle takes days to weeks of senior engineer attention per model. As a result, techniques that have proven effective on one model diffuse into others slowly and unevenly, leaving substantial recoverable signal unexplored. We present Agentic ML Exploration (A-MLE), an autonomous LLM-agent system that systematically explores ML techniques across a portfolio of ads ranking models. A-MLE decomposes ML iteration into five stages involving hypothesis generation, exploration strategy, experiment execution, result analysis and shared knowledge substrate which are orchestrated by a single agent that invokes domain-specific skills and agentic workflows against a sandboxed execution layer, with human-in-the-loop checkpoints at each stage boundary. We deploy A-MLE across a representative set of large-scale ads ranking models and evaluate it along a tiered capability framework (tool availability, autonomous workflow execution, and open-ended exploration). We further report a controlled cross-LLM study using a fixed agent loop, which surfaces qualitative differences in execution reliability and exploration aggressiveness across the Claude Sonnet, Gemini, and GPT families. We discuss failure modes and the design choices that govern reliability. Our findings suggest that agentic exploration is a practical force multiplier for ML engineers in industrial recommenders, especially for the long tail of models that rarely receive expert attention.
Chinese Translation
现代工业广告排序系统日益受到制约的并非模型容量或训练算力,而是人类机器学习迭代的吞吐量——即为产生一项统计显著的改进所需的科研、实现、训练、调试、评估与上线等循环。一个典型的排序系统包含众多差异化的模型,其数据、架构与基础设施约束各不相同,而每个迭代周期需要资深工程师对每个模型投入数天至数周的精力。因此,已在某一模型上被证明有效的技术向其他模型的扩散缓慢且不均衡,导致大量可恢复的信号未被挖掘。我们提出智能体化机器学习探索(Agentic ML Exploration, A-MLE),一个自主的大语言模型智能体系统,用于系统性地探索广告排序模型组合中的各类机器学习技术。A-MLE 将机器学习迭代分解为五个阶段:假设生成、探索策略、实验执行、结果分析以及共享知识基底,由单一智能体统一编排,该智能体在沙箱化执行层之上调用领域特定的技能与智能体工作流,并在每个阶段边界设置人机协同(human-in-the-loop)检查点。我们在一组具有代表性的大规模广告排序模型上部署了 A-MLE,并沿分层能力框架(工具可用性、自主工作流执行和开放式探索)对其进行评估。我们进一步报告了一项采用固定智能体循环的受控跨大语言模型研究,揭示了 Claude Sonnet、Gemini 与 GPT 系列在执行可靠性与探索激进程度方面的定性差异。我们讨论了失败模式以及决定可靠性的设计选择。我们的研究结果表明,智能体化探索是工业推荐系统中机器学习工程师的实用力量倍增器,对于鲜少获得专家关注的模型长尾而言尤为如此。
cs.AI / 160 / 2609.08254

CircuTutor: Transforming Static Circuit Problems into Intelligent and Dynamic Tutoring

CircuTutor:将静态电路问题转化为智能化动态辅导
Luo, Ziyu, Ma, Xiaorui, Chen, Lin, Chen, Xiaoming
Abstract
Learning direct current circuit concepts requires learners to connect invisible physical quantities, such as current, voltage, resistance, and power, with observable outcomes such as bulb brightness. Conventional textbook materials and general-purpose circuit simulators provide opportunities for problem solving and exploration but offer limited support for explaining why circuit behavior changes or diagnosing the reasoning behind incorrect answers. We present CircuTutor, a circuit-state-driven intelligent tutoring system that transforms static textbook circuit problems into an interactive tutoring workflow. CircuTutor first uses multimodal problem parsing to extract the textbook question, circuit topology, component parameters, switch states, and answer options, which are converted into a structured task and validated through circuit simulation. Learners can then interactively explore the circuit (by changing parameters) and submit an answer while a SPICE-compatible solver computes physically consistent circuit states. After the learner submits an answer, CircuTutor presents a before-and-after circuit state animation corresponding to the selected operation, organizes the simulated state changes into a causal reasoning chain that explains the underlying circuit behavior, maps answer discrepancies to likely misconceptions, and generates adaptive follow-up exercises targeted at the diagnosed misconception. Our experimental results demonstrate that CircuTutor effectively improves conceptual learning and the overall learning experience. The proposed framework demonstrates how simulated circuit states can be transformed into intelligent and interactive tutoring for circuit education, with the potential to generalize to other STEM domains.
Chinese Translation
学习直流电路概念需要学习者将不可见的物理量(如电流、电压、电阻和功率)与可观察的结果(如灯泡亮度)联系起来。传统教材和通用电路模拟器虽然为解题和探究提供了机会,但在解释电路行为为何发生变化或诊断错误答案背后的推理方面支持有限。我们提出了CircuTutor,这是一个由电路状态驱动的智能辅导系统,可将静态的教材电路问题转化为交互式辅导流程。CircuTutor首先利用多模态题目解析技术提取教材中的问题、电路拓扑、元件参数、开关状态和答案选项,并将其转换为结构化任务,同时通过电路仿真进行验证。学习者随后可以交互式地探索电路(通过改变参数)并提交答案,与此同时,一个兼容SPICE的求解器会计算物理上一致的电路状态。在学习者提交答案后,CircuTutor会展示与所选操作对应的电路状态前后对比动画,将仿真得到的状态变化组织成因果推理链以解释底层电路行为,将答案差异映射到可能存在的错误概念,并针对诊断出的错误概念生成自适应的后续练习。实验结果表明,CircuTutor有效提升了概念学习和整体学习体验。所提出的框架展示了如何将仿真的电路状态转化为面向电路教育的智能化交互式辅导,并具有推广到其他STEM领域的潜力。
cs.AI / 161 / 2609.08258

Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Memory Systems

已被撤销却仍具权威性:智能体记忆系统中撤销机制的实证研究
Shen, Yi Ting, Toyoda, Kentaroh, Leung, Alex
Abstract
Long-running language-model agents depend on persistent memory. Many agent-memory systems preserve history through soft revocation: a contradicted fact is marked invalid and retained rather than deleted. However, whether that mark is enforced at retrieval time is unexamined. In this paper, we measure five such systems: we load each with a revoked policy and its replacement, track whether the revoked fact is returned at retrieval and whether the agent then acts on it across nine policy scenarios and nine models, and score every trial under six defense conditions. We find that no system enforces revocation by default: the revoked fact is returned wherever the revocation label is visible to the retrieval layer, outranks its replacement, and leads agents to the unsafe action. Based on these findings, we develop a guard that sits between the agent and any memory backend and withholds records that are revoked or conflict with their replacement.
Chinese Translation
长时间运行的语言模型智能体依赖于持久化记忆。许多智能体记忆系统通过软撤销来保留历史:被取代(矛盾)的事实被标记为无效并予以保留,而非直接删除。然而,该标记是否在检索时被强制执行,尚未得到检验。本文对五个此类系统进行了测量:我们在每个系统中加载一条被撤销的策略及其替代策略,跟踪被撤销事实是否会在检索时被返回,以及智能体随后是否会据此采取行动,实验覆盖九种策略场景和九个模型,并在六种防御条件下对每次试验进行评分。我们发现,没有任何系统在默认情况下强制执行撤销:只要撤销标签对检索层可见,被撤销的事实就会被返回,其排序高于替代事实,并导致智能体采取不安全的行动。基于这些发现,我们开发了一个位于智能体与任意记忆后端之间的防护机制,用于拦截已被撤销或与其替代记录相冲突的条目。
cs.AI / 162 / 2609.08267

Evidence-Aligned Entity Verification for Hallucination Detection in Retrieval-Augmented Generation

面向检索增强生成中幻觉检测的证据对齐实体验证方法
Jia, Runsong, Fang, Zhen, Wu, Mengjia, Lu, Jie, Zhang, Yi
Abstract
Hallucination detection is crucial for large language models (LLMs), as hallucinated content creates significant barriers in applications requiring factual accuracy. Current detection methods mainly depend on internal signals like uncertainty and self-consistency checks, using the model's pre-trained knowledge to identify unreliable outputs. However, pre-trained knowledge may become outdated and has coverage limitations, especially for specialized or recent information. To address these limitations, retrieval-augmented generation (RAG) has emerged as a promising solution by retrieving relevant evidence at inference time, grounding outputs beyond the model's parametric knowledge. In this paper, we target a critical and practical learning problem RAG-based hallucination detection (RHD), where RAG is employed to enhance hallucination detection by addressing information updating challenges. To address RHD, we propose a novel method Evidence-Aligned Entity Verification (EAEV), which detects entity-level hallucinations by leveraging RAG to align generated entities with retrieved evidence contexts. Specifically, EAEV evaluates entity-evidence alignment through three complementary dimensions and introduces counterfactual stability analysis to ensure robust alignments under evidence perturbations. Experiments across multiple RAG benchmarks demonstrate that EAEV achieves consistent improvements over existing methods with strong generalization capabilities.
Chinese Translation
幻觉检测对大语言模型(LLMs)至关重要,因为幻觉内容在需要事实准确性的应用中构成了重大障碍。目前的检测方法主要依赖不确定性、自一致性检查等内部信号,利用模型的预训练知识来识别不可靠的输出。然而,预训练知识可能过时且存在覆盖范围的局限,尤其是对于专业性或最新信息。为解决这些局限,检索增强生成(RAG)作为一种有前景的解决方案应运而生,它通过在推理时检索相关证据,使模型输出超越参数化知识而获得事实支撑。在本文中,我们针对一个关键且实用的学习问题——基于RAG的幻觉检测(RHD),即利用RAG来增强幻觉检测,从而应对信息更新方面的挑战。为解决RHD问题,我们提出了一种新方法——证据对齐实体验证(Evidence-Aligned Entity Verification, EAEV),该方法通过利用RAG将生成的实体与检索到的证据上下文进行对齐,从而检测实体级别的幻觉。具体而言,EAEV通过三个互补的维度评估实体与证据的对齐程度,并引入反事实稳定性分析,以确保在证据扰动下对齐结果的稳健性。在多个RAG基准上的实验表明,EAEV相比现有方法取得了一致的性能提升,并展现出较强的泛化能力。
cs.AI / 163 / 2609.08271

Three Types of Negation of Triple and its Elements and an Extension of Triple

三元组及其元素的三种否定与三元组的一种扩展
Pan, Zhenghua
Abstract
In various data models, the classical triple is a typical semantic data model. However, due to the design of the triple as a simple structure for representing positive assertions, it cannot sufficiently express different forms of negation present in the triple and its elements. This paper conceptually proposes that there are three distinct forms of negation within triples and their elements: contradictory negation, opposite negation and intermediary negation. Based on the the set SCOI and the logic LCOI+PLCOI with three kinds of negation, we propose an extension of triple that can distinguish and express these three different negations in the triple and its elements, called the TCOI triple with contradictory negation, opposite negation and intermediary negation. The TCOI triple is a semantic and structural extension of the classical triple. While retaining the ability to express positive assertions, it systematically introduces the three semantic dimensions of three negations, allowing these negations to independently act on the elements of the triple and on the whole triple. This significantly enhances the triple model capability to represent and reasoning about complex negative information. This paper also explores the expressive power and reasoning of the TCOI triple, as well as the application of TCOI triple implication reasoning in counterfactuals and counterfactual reasoning. We propose a truth-value (continuous value) algorithm for TCOI triple implication reasoning and perform its calculation through an example of the counterfactuals and counterfactual reasoning.
Chinese Translation
在各种数据模型中,经典三元组是一种典型的语义数据模型。然而,由于三元组被设计为表示肯定断言的简单结构,它无法充分表达三元组及其元素中存在的不同形式的否定。本文从概念上提出,在三元组及其元素中存在三种不同的否定形式:矛盾否定、对立否定和中介否定。基于带有三种否定的集合 SCOI 和逻辑 LCOI+PLCOI,我们提出了一种三元组的扩展,能够区分并表达三元组及其元素中的这三种不同的否定,称为带有矛盾否定、对立否定和中介否定的 TCOI 三元组。TCOI 三元组是经典三元组在语义和结构上的扩展。它在保留表达肯定断言能力的同时,系统地引入了三种否定的三个语义维度,使这些否定能够独立地作用于三元组的元素以及整个三元组。这显著增强了三元组模型表示和推理复杂否定信息的能力。本文还探讨了 TCOI 三元组的表达能力和推理,以及 TCOI 三元组蕴含推理在反事实与反事实推理中的应用。我们提出了一个用于 TCOI 三元组蕴含推理的真值(连续值)算法,并通过一个反事实与反事实推理的实例进行了计算。
cs.AI / 164 / 2609.08273

MemForest: Efficient Agent Memory Management via EventTree Partitioning and Progressive Merging

MemForest:基于事件树划分与渐进合并的高效智能体记忆管理
Wang, Junxi, Sun, Te, Zhu, Jiayi, Zhang, Chen, Li, Siyuan, Liu, Xuyang, Wen, Zichen, Tu, Xiaobing, Ren, Jinkui, Zhang, Xiantao, Yuan, Ziqi, Zhang, Linfeng
Abstract
Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, and video understanding. However, continuously accumulated memory introduces substantial storage and retrieval costs during inference. To address this issue, we propose \textbf{MemForest}, a general memory compression framework adaptable to various agent memory systems. Specifically, MemForest partitions historical memory into event-centric units by leveraging global semantic similarity and local temporal continuity. For each unit, it constructs a maximum spanning tree, termed an EventTree, and progressively merges redundant memory nodes by selecting high-weight edges, reducing storage overhead. Furthermore, we introduce an anchor-guided propagation retrieval mechanism that retrieves relevant memory nodes from the temporal neighborhoods of key nodes, improving retrieval accuracy. Extensive experiments demonstrate the effectiveness of MemForest. Under the unimodal Mem0 framework, MemForest retains \textbf{97.1%} of the original performance while compressing \textbf{50%} of historical memory across three benchmarks (LoCoMo, LongMemEval, and PersonaMem), achieving a \textbf{1.89x} retrieval speedup. Under the multimodal M3-Agent framework, it preserves \textbf{99.7%} of the original performance with a \textbf{50%} compression ratio across two benchmarks (M3-Bench-robot and M3-Bench-web), achieving a \textbf{2.24x} retrieval speedup. \textcolor{RoyalBlue}{\textit{Our code is available at [https://github.com/Celina-love-sweet/MemForest.}}](https://github.com/Celina-love-sweet/MemForest.}})
Chinese Translation
智能体记忆系统在长期对话、个性化助手和视频理解等领域已展现出显著潜力。然而,持续累积的记忆在推理过程中会带来巨大的存储与检索开销。为解决这一问题,我们提出了 MemForest,一个可适配于各类智能体记忆系统的通用记忆压缩框架。具体而言,MemForest 通过利用全局语义相似性和局部时间连续性,将历史记忆划分为以事件为中心的单元。对于每个单元,它构建一棵最大生成树(称为 EventTree),并通过选择高权重边渐进地合并冗余记忆节点,从而降低存储开销。此外,我们引入了一种锚点引导的传播检索机制,从关键节点的时间邻域中检索相关记忆节点,提高了检索准确性。大量实验证明了 MemForest 的有效性。在单模态 Mem0 框架下,MemForest 在三个基准(LoCoMo、LongMemEval 和 PersonaMem)上压缩了 50% 的历史记忆,同时保留了 97.1% 的原始性能,并实现了 1.89 倍的检索加速。在多模态 M3-Agent 框架下,它在两个基准(M3-Bench-robot 和 M3-Bench-web)上以 50% 的压缩率保留了 99.7% 的原始性能,并实现了 2.24 倍的检索加速。我们的代码已发布于 https://github.com/Celina-love-sweet/MemForest。
cs.AI / 165 / 2609.08275

Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation

超越连贯性:多镜头音视频生成中专业剪辑技术执行的基准测试
Zeng, Tianyi, Liao, Junchao, Wei, Yujie, Zhang, Ziying, Li, Litao, Wang, Tianyi, Wei, Zhichao, Xu, Shuyao, Qiang, Wenwen, Zhu, Siyu, Zhang, Zhenghao, Qin, Long
Abstract
Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherence does not imply the ability to execute editing techniques. Professional editing depends on shot structure, transition grammar, audio-video cut relations, and montage, yet existing benchmarks largely rely on proxies such as content quality, synchronization, or physical plausibility, systematically missing whether such editing instructions are actually executed. We introduce CutCraft, the first benchmark for editing-technique execution in multi-shot audio-video generation. CutCraft extends structured multi-shot prompts with explicit editing specifications and is paired with a hierarchical hybrid evaluation framework that combines shot-structure alignment, expert-model metrics, tool-grounded multimodal judgment, and rubric-based question answering. Beyond evaluation, we design an agentic editing baseline that decomposes generation into planning, shot-level synthesis, and post-hoc composition, explicitly realizing editing semantics such as J-cuts, L-cuts, and transition timing. Across 13 state-of-the-art closed- and open-source models, CutCraft reveals a consistent gap between coherence and editing-technique execution: current systems often produce plausible multi-shot videos yet fail to execute editorial instructions reliably. We find unstable shot structures, weak control of transition execution, and sharp degradation on higher-order montage, while aesthetic quality is only weakly correlated with editing-technique compliance. The benchmark and metrics, and the editing agent baseline are available at https://github.com/AlibabaResearch/cut-craft-bench.
Chinese Translation
近期的多镜头音视频生成器能够产出日益连贯且具有电影感的输出,但连贯并不意味着具备执行剪辑技术的能力。专业剪辑依赖于镜头结构、转场语法、音视频剪辑关系以及蒙太奇,然而现有基准测试主要依赖内容质量、同步性或物理合理性等替代指标,系统性地忽略了此类剪辑指令是否被真正执行。我们提出了 CutCraft,这是首个针对多镜头音视频生成中剪辑技术执行的基准测试。CutCraft 在结构化多镜头提示的基础上扩展了显式的剪辑规范,并配备了一个层次化混合评估框架,该框架结合了镜头结构对齐、专家模型指标、基于工具的多模态判断以及基于评分标准的问题回答。除评估之外,我们设计了一个智能体剪辑基线,将生成过程分解为规划、镜头级合成和后期组合,显式地实现诸如 J-cut、L-cut 和转场时机等剪辑语义。在 13 个最先进的闭源与开源模型上,CutCraft 揭示了连贯性与剪辑技术执行之间的一致性差距:当前系统往往能生成看似合理的多镜头视频,却无法可靠地执行剪辑指令。我们发现镜头结构不稳定、转场执行控制薄弱,且在高阶蒙太奇上性能急剧下降,而美学质量与剪辑技术合规性仅呈弱相关。该基准测试、评估指标以及剪辑智能体基线可在 https://github.com/AlibabaResearch/cut-craft-bench 获取。
cs.AI / 166 / 2609.08288

LEBGen: An LLM-Enhanced Bayesian Network Framework for Few-Shot Travel Survey Data Generation

LEBGen:一种用于小样本出行调查数据生成的LLM增强贝叶斯网络框架
Shen, Zijian, Zhou, Bin, Wang, Jiguang, Zhao, Ya, Ke, Jintao
Abstract
Travel survey data are essential for transportation planning and travel behavior analysis, yet collecting large-scale representative samples is costly and time-consuming. A practical alternative is to generate synthetic survey records from a few-shot sample. However, such samples provide incomplete coverage of heterogeneous traveler groups and insufficient evidence for recovering the complex dependencies between demographic characteristics and travel behavior. Existing approaches have complementary limitations. Probabilistic generative models such as Bayesian networks (BNs) offer explicit distributional control, but structures learned from few-shot samples may omit meaningful dependencies or retain spurious ones. Large language models (LLMs) can help address these difficulties in BN structure learning by providing behavioral knowledge that complements the limited statistical evidence. We therefore propose LEBGen, an LLM-enhanced BN framework that uses this knowledge to refine network structure for few-shot travel survey data generation. Specifically, the LLM first identifies traveler personas from demographic attribute and travel behavior statistics, then recovers dependencies missed by the persona-augmented BN structure and prune spurious ones. The refined BN is parameterized exclusively from the observed data to generate synthetic records. Under a 2% few-shot setting on the 2022 Hong Kong Travel Characteristics Survey, LEBGen reduces the mean marginal Jensen-Shannon divergence from 0.0671 to 0.0091 and the mean absolute Cramer's V error by 14.3% over the best-performing baseline, substantially improving both distributional and dependency fidelity.
Chinese Translation
出行调查数据对交通规划和出行行为分析至关重要,然而收集大规模代表性样本既成本高昂又耗时。一种实用的替代方法是从小样本(few-shot sample)生成合成调查记录。然而,此类样本对异质性出行者群体的覆盖不完整,且缺乏足以恢复人口统计特征与出行行为之间复杂依赖关系的证据。现有方法各有互补的局限性。以贝叶斯网络(Bayesian Networks, BNs)为代表的概率生成模型提供了显式的分布控制能力,但从少样本中学习到的网络结构可能遗漏有意义的依赖关系或保留虚假依赖。大语言模型(Large Language Models, LLMs)能够提供补充有限统计证据的行为知识,从而帮助解决贝叶斯网络结构学习中的这些难题。因此,我们提出LEBGen,一种LLM增强的贝叶斯网络框架,利用此类知识优化网络结构以实现小样本出行调查数据生成。具体而言,LLM首先从人口统计属性和出行行为统计中识别出行者画像(traveler personas),然后恢复被画像增强的贝叶斯网络结构所遗漏的依赖关系,并剪除虚假依赖。精化后的贝叶斯网络仅基于观测数据进行参数化,用于生成合成记录。在2022年香港出行特征调查的2%少样本设定下,LEBGen将平均边际Jensen-Shannon散度从0.0671降至0.0091,平均绝对Cramer's V误差较表现最佳的基线方法降低14.3%,在分布保真度和依赖关系保真度方面均有显著提升。
cs.AI / 167 / 2609.08407

FastE: Readout-Triggered Token Compression for LLM Embedding Inference

FastE:面向大语言模型嵌入推理的读出触发式词元压缩方法
Shu, Jinsong, Wen, Jinyong, Wang, Baokun, Xie, Zhongle, Shou, Lidan, Wang, Weiqiang, Chen, Gang
Abstract
In this study, we identify depth-dependent prefix redundancy in final-readout LLM embedding models, notably across representative backbones including Qwen3-Embedding and Qwen3-VL-Embedding. We find that removing prefix states is substantially more damaging in shallow layers than at greater depth, showing that prefix states become increasingly compressible as the prefix and readout states propagate through the network. To this end, we introduce FastE, a training-free, plug-and-play method. FastE uses a shared fixed threshold on batch-mean readout-prefix alignment as a lightweight online heuristic for selecting when compression occurs, and ranks prefix states by the attention scores they receive from the readout position to determine which states are retained in subsequent layers. Our evaluations demonstrate FastE's ability to substantially reduce computational costs: on NarrativeQA with Qwen3-Embedding-0.6B, it reduces decoder-backbone FLOPs by 40.11% while retaining 99.53% of Full Forward nDCG@10. Across five text embedding benchmarks, two backbone scales, and three cross-modal retrieval tasks, the quality-efficiency trade-off is directly customizable through the maximum removal ratio without retraining. We believe FastE offers practical value for scalable embedding generation in retrieval, indexing, clustering, and multimodal representation systems.
Chinese Translation
在本研究中,我们在最终读出(final-readout)式大语言模型嵌入模型中发现了随深度变化的_PREFIX_冗余现象,这一现象在包括Qwen3-Embedding和Qwen3-VL-Embedding在内的代表性骨干模型中均显著存在。我们发现,在浅层中移除前缀状态所造成的损害远大于在较深层中移除,这表明随着前缀状态和读出状态在网络中的逐层传播,前缀状态变得愈发可压缩。为此,我们提出了FastE,一种无需训练、即插即用的方法。FastE在批均值读出-前缀对齐度上使用一个共享的固定阈值,作为轻量级的在线启发式准则来决定何时进行压缩;并通过前缀状态从读出位置所获得的注意力分数对其进行排序,以确定哪些状态在后续层中被保留。我们的评估表明,FastE能够大幅降低计算成本:在NarrativeQA数据集上使用Qwen3-Embedding-0.6B时,它在保留全量前向(Full Forward)nDCG@10的99.53%的同时,将解码器骨干的FLOPs降低了40.11%。在五个文本嵌入基准、两种骨干模型规模以及三个跨模态检索任务上,其质量-效率权衡可通过最大移除比例直接定制,且无需重新训练。我们相信,FastE在检索、索引、聚类以及多模态表示系统中的可扩展嵌入生成方面具有实际应用价值。
cs.AI / 168 / 2609.08418

Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models

Feyospace-v1:网络“水星七杰”如何训练前沿网络模型
Li, Zongjie, W, Alan Z., J, John Nicolas, F, Walter H., L, Scott Donald, P, Gordon Y., X Jr, Deke
Abstract
Training capable cyber agents is often treated primarily as a problem of model scale, yet open-weight post-training is constrained more directly by the cost of executable environments, reliable multi-turn supervision, and access to strong teachers. We present a data-centric framework that addresses these bottlenecks through five complementary systems: Choulea analyzes hidden reasoning signatures, SkyReal reduces teacher-sampling cost, Hongzwang bypasses API restrictions on teacher execution, PSBreakup restores capabilities weakened by model merging, and Kreator converts expert interventions into trainable reasoning. Our data engine constructs resettable coding, vulnerability, CTF, kernel-history, full-exploit, firmware, and device-backed environments. Candidate trajectories are retained only after execution verification and evidence auditing, yielding 164,269 trajectories for long-context supervised fine-tuning. The three checkpoints improve over their starting models by an average of 23.76% on the full CyberGym suite and 10.49% across the pooled CTF suites. As of September 1, 2026, Feyospace-s1 achieves a verified success rate of 63.24% and ranks 10th on the official CyberGym leaderboard, while all three checkpoints rank 1st among models at comparable parameter scales. To our knowledge, this is the first end-to-end demonstration that a seven-person independent team can train open-weight models with leading agentic cyber capability.
Chinese Translation
训练能力强大的网络(cyber)智能体通常主要被视为模型规模问题,然而开放权重的后训练(post-training)更直接地受限于可执行环境的成本、可靠的多轮监督以及获取强大教师模型的途径。我们提出了一个以数据为中心的框架,通过五个互补系统来解决这些瓶颈:Choulea 分析隐藏的推理特征,SkyReal 降低教师采样成本,Hongzwang 绕过教师执行上的 API 限制,PSBreakup 恢复因模型合并而削弱的能力,Kreator 将专家干预转化为可训练的推理。我们的数据引擎构建了可重置的代码、漏洞、CTF、内核历史、完整漏洞利用、固件以及真实设备支持的环境。候选轨迹仅在经过执行验证和证据审计后才予以保留,最终产出 164,269 条轨迹用于长上下文监督微调。三个检查点在其初始模型基础上,在完整 CyberGym 套件上平均提升 23.76%,在汇总的 CTF 套件上提升 10.49%。截至 2026 年 9 月 1 日,Feyospace-s1 在官方 CyberGym 排行榜上取得了 63.24% 的验证成功率,排名第 10,而所有三个检查点在同等参数规模的模型中均排名第 1。据我们所知,这是首个端到端证明七人独立团队也能训练出具有领先智能体网络能力的开放权重模型的工作。
cs.AI / 169 / 2609.08435

EvolveScaler: Synthesizing Information-Evolution Contexts via Executable State Machines and Natural-Language Rendering

EvolveScaler:基于可执行状态机与自然语言渲染的信息演化上下文合成
Zhao, Ziliang, Xu, Zenan, Wang, Shuting, Wang, Zhao, Cao, Bowen, Hu, Minda, Li, Lincheng, Zhou, Pluto, Dou, Zhicheng
Abstract
In persistent interactions, long contexts may encode an evolving process rather than a fixed record: later events can revise or revoke earlier information, changing what remains valid and what conclusions follow. We call this setting information evolution (IE). Solving IE requires identifying valid records, applying updates in order, and reconstructing the query-relevant state from the event history. Existing text-first synthesis pipelines make such data difficult to verify because state transitions and answer logic remain implicit. We introduce EvolveScaler, a code-driven framework that defines information evolution before rendering it as natural language. Human-authored operational specifications define state transitions, record validity, difficulty controls, and executable answer logic; a strong LLM then synthesizes a self-contained simulator from each specification. Executing validated simulators produces natural-language multi-turn event histories, while deterministic replay computes reference answers and atomic checklists. We instantiate EvolveScaler with 117 task prototypes and 159 final-question operators across five difficulty levels spanning approximately 7 to 1,200 events per instance, yielding about 35,100 training examples and 585 validated evaluation instances. On the very_long tier, the strongest model reaches 59.3% avg@5, while six models score below 10%. Training an internal A3B model on 6,000 EvolveScaler examples improves performance over its base checkpoint on all eight independently constructed out-of-distribution benchmarks, with a 5.25-point average gain. These results show that code-driven IE synthesis provides both challenging evaluation and transferable training supervision.
Chinese Translation
在持续交互中,长上下文所编码的可能是一个不断演化的过程,而非固定记录:后续事件可以修订或撤销较早的信息,从而改变仍然有效的信息内容以及由此得出的结论。我们将这一设定称为信息演化(Information Evolution, IE)。解决IE问题需要识别有效记录、按顺序应用更新,并从事件历史中重构与查询相关的状态。现有的以文本为主的合成流水线难以对这类数据进行验证,因为状态转移和答案逻辑仍然是隐式的。我们提出了EvolveScaler,这是一个代码驱动的框架,它先定义信息演化过程,再将其渲染为自然语言。由人工编写的操作规范定义状态转移、记录有效性、难度控制以及可执行的答案逻辑;随后,一个强大的大语言模型(LLM)根据每个规范合成一个自包含的模拟器。执行经过验证的模拟器可生成自然语言的多轮事件历史,同时确定性重放可计算参考答案和原子化检查清单。我们以117个任务原型和159个最终问题算子实例化了EvolveScaler,涵盖五个难度级别,每个实例约包含7至1,200个事件,共产出约35,100个训练样本和585个经过验证的评估实例。在very_long难度层级上,最强模型的avg@5达到59.3%,而六个模型的得分低于10%。在6,000个EvolveScaler样本上训练内部A3B模型后,该模型在全部八个独立构建的分布外基准上均优于其基础检查点,平均提升5.25分。这些结果表明,代码驱动的IE合成既能提供具有挑战性的评估,也能提供可迁移的训练监督。
cs.AI / 170 / 2609.08452

SRPO: Setwise Relative Policy Optimization for Multi-Agent LLMs

SRPO:面向多智能体大语言模型的集合级相对策略优化
Yang, Shengtian, Xiong, Ziyu, Li, Yu, Li, Yewen, Cai, Qingpeng, Feng, Lei
Abstract
Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment. However, existing reinforcement learning methods usually optimize each response or trajectory separately, even when several outputs jointly cause one state transition. Consequently, the update unit differs from the action executed by the system. To address this problem, we propose SRPO (Setwise Relative Policy Optimization), which treats the active set the minimal set of outputs consumed by one transition, as one multi-agent action. Specifically, SRPO combines member log-ratios into one cardinality-normalized set ratio, assigns one relative advantage, and clips the set once. This formulation unifies division of labor and joint co-evolution as actions with different set sizes. Experiments on mathematical reasoning and multi-turn search demonstrate one training interface for fixed, mixed, and dynamically routed workflows across four model scales, with the strongest macro-average results among the reported comparisons. Optimization diagnostics further characterize its stability under different event reductions and set sizes.
Chinese Translation
多智能体大语言模型通过在共享环境中协调多个策略来解决复杂任务。然而,现有的强化学习方法通常对每条响应或轨迹分别进行优化,即使多个输出共同导致一次状态转移也是如此。因此,更新单元与系统实际执行的动作不一致。为解决这一问题,我们提出了SRPO(Setwise Relative Policy Optimization,集合级相对策略优化),它将活跃集(即一次状态转移所消耗的最小输出集合)视为一个多智能体动作。具体而言,SRPO将各成员的对数比率合并为一个按基数归一化的集合比率,赋予一个相对优势,并对该集合进行一次裁剪。这一公式化方法将分工与联合协同进化统一为具有不同集合大小的动作。在数学推理和多轮搜索任务上的实验表明,该方法为固定、混合以及动态路由的工作流提供了一个统一的训练接口,并在四个模型规模上取得了所报告对比中最强的宏观平均结果。优化诊断进一步刻画了该方法在不同事件归约方式和集合大小下的稳定性。
cs.AI / 171 / 2609.08558

Personalizing LLM Agent Memory Using Biometrics

基于生物特征识别的大语言模型智能体记忆个性化
Qian, Yanhong, Meng, Qingguo, Ding, Shihao, Dong, Xingbo, Jin, Zhe, Wang, Hanrui, Echizen, Isao
Abstract
Personalized memory helps LLM agents deliver stable, tailored assistance by storing and reusing user-specific data across interactions. In multi-user scenarios, however, retrieval must consider not only semantic similarity but also whether the current requester matches the identity associated with the stored memory. We propose Bio-Memory, a biometric-aware memory architecture that conditions memory retrieval on both semantic similarity and biometric matching. Built on top of A-Mem, Bio-Memory augments each atomic memory note with a biometric embedding and uses biometric matching to form the retrieval candidate pool before semantic ranking. We evaluate Bio-Memory on LoCoMo in a 10-user shared-agent setting over 7 face benchmarks and 10 palmprint protocols. Across datasets, Bio-Memory consistently separates owner and non-owner queries. Under face-based personalization, the largest average gap reaches 27.29% / 21.15% in F1 / BLEU-1 on CALFW; under palmprint-based personalization, the corresponding gap is 25.75% / 19.22% on MS_Blue. These results support biometrics as a practical control signal for personalized memory retrieval in shared environments.
Chinese Translation
个性化记忆通过在多次交互中存储和复用用户特定数据,帮助大语言模型(LLM)智能体提供稳定且量身定制的辅助。然而,在多用户场景中,检索不仅需要考虑语义相似性,还必须考虑当前请求者是否与存储记忆所关联的身份相匹配。我们提出了 Bio-Memory,一种生物特征感知的记忆架构,它将记忆检索同时建立在语义相似性和生物特征匹配之上。Bio-Memory 构建于 A-Mem 之上,为每个原子记忆笔记增加了一个生物特征嵌入,并利用生物特征匹配在语义排序之前构建检索候选池。我们在 LoCoMo 数据集上,于 10 用户共享智能体设置下,通过 7 个面部基准和 10 个掌纹协议对 Bio-Memory 进行了评估。在所有数据集上,Bio-Memory 均能稳定地区分所有者与非所有者的查询。在基于面部的个性化设置下,CALFW 上 F1 / BLEU-1 的最大平均差距达到 27.29% / 21.15%;在基于掌纹的个性化设置下,MS_Blue 上相应的差距为 25.75% / 19.22%。这些结果表明,生物特征可作为共享环境中个性化记忆检索的实用控制信号。
cs.AI / 172 / 2609.08566

BIO-MEMART: Biometric-Aware KV Cache Memory for Multi-User LLM Agents

BIO-MEMART:面向多用户LLM智能体的生物特征感知KV缓存存储器
Qian, Yanhong, He, Xuanying, Meng, Qingguo, Ding, Shihao, Dong, Xingbo, Jin, Zhe
Abstract
KV cache is evolving from a serving optimization into an external memory substrate for long-term LLM agents. In a shared multi-user deployment, however, reusable KV blocks introduce a missing access-control question: semantic relevance alone cannot determine whether a memory block is authorized for the current physical user. We propose Bio-MemArt, a biometric-aware KV-cache memory framework for multi-user LLM agents. Bio-MemArt attaches a normalized biometric template to each stored KV memory block, filters the shared memory pool with the current user's biometric probe, and then runs the original MemArt retrieval and KV reuse pipeline only inside the authorized candidate pool. This design preserves latent-space retrieval, direct cache reuse, and decoupled position encoding while adding physical-user access control to shared KV memory. We evaluate Bio-MemArt under Owner and Non-owner query conditions on long-term dialogue QA with face and palmprint benchmarks. Across face benchmarks, the average owner and non-owner biometric success rates are 95.71% and 0.86%; across palmprint benchmarks, they are 97.60% and 2.00%. In the efficiency study, average prefill tokens drop from 18,781.96 under full-context prompting to 28.57 with Bio-MemArt, showing that biometric gating preserves the low-token operating regime of KV-cache memory.
Chinese Translation
KV缓存正从一种服务优化手段演变为长期LLM智能体的外部存储基底。然而,在共享的多用户部署环境中,可复用的KV缓存块引出了一个尚未解决的访问控制问题:仅凭语义相关性无法确定某个存储块是否有权被当前物理用户访问。我们提出Bio-MemArt,一个面向多用户LLM智能体的生物特征感知KV缓存存储框架。Bio-MemArt为每个存储的KV记忆块附加一个归一化的生物特征模板,利用当前用户的生物特征探针对共享存储池进行过滤,然后仅在授权候选池内运行原有的MemArt检索与KV复用流程。该设计在保留潜在空间检索、直接缓存复用和解耦位置编码的同时,为共享KV存储增加了物理用户访问控制。我们在长期对话问答任务上,结合人脸与掌纹基准,在“所有者”与“非所有者”查询条件下对Bio-MemArt进行了评估。在人脸基准上,所有者与非所有者的平均生物特征成功率分别为95.71%和0.86%;在掌纹基准上,分别为97.60%和2.00%。在效率研究中,平均预填充token数从全上下文提示下的18,781.96降至Bio-MemArt下的28.57,表明生物特征门控能够保持KV缓存存储的低token运行模式。
cs.AI / 173 / 2609.08572

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems

AgentGrad:面向多智能体系统的干预引导式提示优化
Chu, Jaewon, Seo, Jinwoo, Cho, Jaewon, Na, Jeehye, Xiong, Yunyang, Kim, Youngdae, Kim, Hyunwoo J.
Abstract
Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple agents, yet their performance depends on the prompt design of each agent. For MAS prompt optimization, textual gradient methods that guide prompt updates using natural-language feedback have emerged as a leading paradigm. In this paper, we identify limitations in two stages of existing textual gradient approaches: gradient extraction and gradient aggregation. In gradient extraction, previous works select a target prompt without verifying whether modifying it resolves the failure, and derive gradients without agent-level supervision over the corresponding agent's intermediate output. In gradient aggregation, individual gradients are randomly grouped and concatenated, often mixing unrelated failure modes and producing prompts that fail to generalize. To address these limitations, we propose \textbf{AgentGrad}, a prompt optimization framework for multi-agent systems based on sequential intervention and semantic textual gradient abstraction. For each failure, sequential intervention modifies the behavior of one agent at a time to identify the target agent whose modification resolves the failure. The modified output of the target agent then serves as agent-level supervision for extracting a fine-grained gradient. Semantic textual gradient abstraction clusters semantically similar gradients to prevent mixing unrelated failure modes, and abstracts each cluster into a generalized gradient that captures the shared corrective pattern. Experimental results show that AgentGrad achieves state-of-the-art performance across five MAS benchmarks and reduces wall-clock optimization time by $2.5\times$ on average compared to the next-fastest baseline.
Chinese Translation
基于大语言模型(LLM)的多智能体系统(MAS)通过部署多个专业化智能体获得了优异的性能,但其性能取决于每个智能体的提示设计。对于多智能体系统的提示优化,利用自然语言反馈引导提示更新的文本梯度方法已成为主流范式。本文指出了现有文本梯度方法在两个阶段中的局限性:梯度提取和梯度聚合。在梯度提取阶段,以往工作在选择目标提示时未验证修改该提示是否能够解决失败,并且在推导梯度时缺乏针对相应智能体中间输出的智能体级监督。在梯度聚合阶段,各梯度被随机分组并拼接,常常混杂不相关的失败模式,导致生成的提示难以泛化。为解决这些局限,我们提出了AgentGrad,一个基于序贯干预和语义文本梯度抽象的多智能体系统提示优化框架。针对每次失败,序贯干预每次只修改一个智能体的行为,以识别出修改后能解决失败的目标智能体。目标智能体修改后的输出随后作为智能体级监督,用于提取细粒度梯度。语义文本梯度抽象对语义相似的梯度进行聚类,以避免混杂不相关的失败模式,并将每个聚类抽象为捕捉共同纠正模式的泛化梯度。实验结果表明,AgentGrad在五个多智能体系统基准上取得了最先进的性能,与次快的基线方法相比,平均将实际优化时间缩短了2.5倍。
cs.AI / 174 / 2609.08592

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

用于智能体评估中可控用户模拟的三层人格向量
Khedar, Rahul, Eshita, Thondapu, Sneha Teja Sree Reddy, Malhotra, Mayank, Das, Arup Kumar, Mishra, Jitesh Chandra, Menon, Arun, Karn, Avinash, V, Mouli
Abstract
Evaluating tool-augmented LLM agents requires diverse, realistic user inputs yet most evaluation frameworks use flat role descriptions ("you are an angry customer") that produce near-identical conversations regardless of the underlying scenario. In this paper, we propose a three-tier persona vector with 23 operationalized dimensions: 6 categorical demographics (jurisdiction, age, channel, device, language proficiency, time availability), 12 continuous behavioral traits (patience, assertiveness, digital literacy, etc.) sampled with Gaussian noise around curated profile base vectors, and 5 continuous emotional states (frustration, anxiety, trust, confidence, stress) that shift in response to scenario context. Orthogonal to the persona, a 4-level query-complexity overlay controls utterance phrasing from direct to deliberately vague. We evaluate the persona model inside a synthetic data generation pipeline across 64,698 multi-turn conversations spanning 8 named profiles and 3 production corpora. Key findings: (i) a 15.8 percentage-point spread in agent goal-achievement across personas confirms trait vectors produce measurably different user behavior; (ii) the same persona behaves differently across scenarios due to scenario-reactive emotional state shifts, validating the scenario-reactive design; (iii) domain-specific projects show persona sensitivity on booking-flow compliance (~15-20 percentage points gap between tier-aware and pressure-test personas), demonstrating the model faithfully reproduces real-world difficulty distributions; (iv) seven rule-described trait correlations produce auditable co-occurrence patterns without requiring learned covariance matrices. The persona model is fully specified for reproduction.
Chinese Translation
评估工具增强型大语言模型(LLM)智能体需要多样化且真实的用户输入,然而大多数评估框架采用扁平化的角色描述(如“你是一位愤怒的顾客”),导致无论底层场景如何,生成的对话几乎完全相同。本文提出一种包含23个可操作化维度的三层人格向量:6个类别型人口统计特征(司法辖区、年龄、渠道、设备、语言能力、可用时间)、12个连续型行为特质(耐心、坚定性、数字素养等),这些特质围绕精心设计的画像基础向量并以高斯噪声进行采样;以及5个连续型情绪状态(挫败感、焦虑、信任、信心、压力),它们会随场景上下文而变化。与人格正交的是一个4级查询复杂度叠加层,用于控制话语表述方式,从直接明确到刻意模糊。我们在一个合成数据生成流水线中对该人格模型进行评估,共涉及64,698段多轮对话,涵盖8个具名画像和3个生产语料库。主要发现:(i) 不同人格之间智能体目标达成率存在15.8个百分点的差异,证实特质向量能产生可测量地不同的用户行为;(ii) 由于场景响应型情绪状态的变化,同一人格在不同场景中表现不同,验证了场景响应设计的有效性;(iii) 领域特定项目显示出人格敏感性(在预订流程合规性上,特质感知型人格与压力测试型人格之间存在约15-20个百分点的差距),表明该模型能忠实复现真实世界的难度分布;(iv) 七条基于规则描述的特质相关性产生了可审计的共现模式,无需学习协方差矩阵。该人格模型已被完整说明以便复现。
cs.AI / 175 / 2609.08599

Graph-Based Personalized Memory for LLM Agents: Representation, Evolution, Retrieval, and Evaluation

基于图的LLM智能体个性化记忆:表示、演化、检索与评估
Nguyen, Dac Duy Anh, Qiu, Zhangchi, Chen, Shigeng, Liew, Alan Wee-Chung
Abstract
Large Language Model (LLM) agents are evolving from single-session tools toward long-term personal assistants that must adapt to individual users across tasks, contexts, and interactions. This shift makes memory a core requirement for personalization, since user preferences, goals, constraints, relationships, and past experiences are accumulated gradually and often change over time. Graph-based personalized memory provides a structured way to model such user information through explicit relations, temporal context, and evidence links. Such representations can model not only what an agent remembers about a user but also how memories are connected, revised, and retrieved to support personalized decisions. However, existing work remains fragmented across personalized agents and generic graph memory frameworks, making it difficult to understand the design space as a whole. This survey develops a lifecycle-oriented view of graph-based personalized memory for LLM agents. We organize existing studies around memory representation, memory evolution, memory retrieval, and memory evaluation. We further compare key design choices, discuss current evaluation practices, and open challenges in building reliable long-term personalized agents. This survey aims to clarify how graph-based memory can support adaptive, controllable, and user-centric LLM agents.
Chinese Translation
大语言模型(LLM)智能体正从单次会话工具向长期个人助手演进,后者需要跨任务、场景和交互不断适应个体用户。这一转变使记忆成为个性化的核心需求,因为用户的偏好、目标、约束、关系以及过往经验都是逐步积累的,并且常常随时间变化。基于图的个性化记忆通过显式关系、时间上下文和证据链接,为建模此类用户信息提供了一种结构化方式。这类表示不仅可以刻画智能体记住了关于用户的哪些内容,还可以刻画记忆之间如何相互关联、如何被修正和检索,从而支持个性化决策。然而,现有研究分散于个性化智能体与通用图记忆框架之间,使得人们难以从整体上理解这一设计空间。本综述针对LLM智能体的基于图个性化记忆构建了一个面向生命周期的视图,围绕记忆表示、记忆演化、记忆检索和记忆评估对现有研究进行组织。我们进一步比较了关键设计选择,讨论了当前的评估实践以及构建可靠的长期个性化智能体所面临的开放性挑战。本综述旨在阐明基于图的记忆如何支持自适应、可控且以用户为中心的LLM智能体。
cs.AI / 176 / 2609.08602

CLAMP: Constrained Decoding for Vision-Language Embodied Planning

CLAMP:面向视觉-语言具身规划的约束解码方法
Ma, Tianyi, Kordjamshidi, Parisa
Abstract
Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual observations into executable action sequences. However, fluent plans are not always executable. A VLM may refer to objects that are not visually observed, select actions whose required affordances are unavailable, or violate syntax and action constraints. We introduce CLAMP, a multimodal constraint-grounding framework that turns scene evidence into decoding-time constraints for a frozen VLM planner. CLAMP uses the initial observation to restrict object references to those supported by the scene, while a provided symbolic action model specifies state transitions and goals. During decoding, hard masks eliminate invalid next-token candidates, while a Hidden Markov Model (HMM)-based world-state lookahead module reweights the probabilities of the remaining feasible candidates based on action preconditions and goal reachability. This allows the planner to retain the VLM's language prior while preventing visually unsupported, unsafe, or infeasible candidates from entering the plan. For unseen tasks and environments, CLAMP adapts the HMM at test time using label-free continuations sampled from the frozen VLM. Experiments on VLABench, SafeAgentBench, and TaPA show that scene-grounded constraints improve object grounding and safety, while most remaining failures stem from perception errors or misaligned constraint specifications.
Chinese Translation
具身规划日益依赖于视觉-语言模型将指令和视觉观察转化为可执行的动作序列。然而,流畅的计划并不总是可执行的。视觉-语言模型可能引用未被视觉观察到的物体,选择所需可供性不可用的动作,或违反语法和动作约束。我们提出了CLAMP,一个多模态约束落地框架,它将场景证据转化为冻结的视觉-语言模型规划器在解码时的约束。CLAMP利用初始观察将物体引用限制在场景所支持的范围内,同时由给定的符号动作模型指定状态转移和目标。在解码过程中,硬掩码消除无效的下一词元候选,而基于隐马尔可夫模型(HMM)的世界状态前瞻模块根据动作前置条件和目标可达性对剩余可行候选的概率进行重新加权。这使得规划器能够保留视觉-语言模型的语言先验,同时防止视觉上不支持、不安全或不可行的候选进入计划。对于未见过的任务和环境,CLAMP在测试时利用从冻结的视觉-语言模型中采样的无标签续写来自适应调整HMM。在VLABench、SafeAgentBench和TaPA上的实验表明,基于场景落地的约束能够提升物体落地能力和安全性,而大多数剩余的失败源于感知错误或约束规范的不匹配。
cs.AI / 177 / 2609.08719

GoAnt: Quality-Diversity Multi-Agent Search for Alpha Factor Discovery in Market Microstructure Data

GoAnt:面向市场微观结构数据中Alpha因子发现的质量-多样性多智能体搜索
Zhao, Stella, Sha, Tommy
Abstract
Automated alpha factor discovery searches symbolic trading signals from price-volume panels and order-book data under a fixed evaluation budget. Existing single- and multi-agent program-search systems can overfit predictive proxies that fail after execution costs and repeatedly explore redundant factor families, limiting execution robustness and behavioral diversity. We introduce GoAnt, a quality-diversity multi-agent search framework that combines non-communicating Explorer, Exploiter and Connector workers with a shared adaptive Mental Map and a compact Queen dispatcher. The Mental Map organizes candidates by leakage-free execution profiles and retains one elite per niche, while the Queen reallocates the evaluation budget from explicit search-state summaries. We also define a map-independent effective-yield protocol that counts high-quality, mutually nonredundant factors directly from each method's evaluation records, giving archive-based and map-free systems the same ruler. On real A-share microstructure data spanning 2023--2026, GoAnt reaches quality-weighted yields of 41.8 and 47.6 in price-volume and order-book settings, improving the strongest baseline by 57% and 97% under matched budgets. Its locked populations retain 0.64 and 0.67 of in-sample quality out of sample, compared with 0.61 and 0.63 for a static map.
Chinese Translation
自动化Alpha因子发现旨在固定的评估预算下,从量价面板数据和订单簿数据中搜索符号化交易信号。现有的单智能体和多智能体程序搜索系统容易过拟合预测代理指标,这些指标在计入交易执行成本后失效,且会反复探索冗余的因子族,从而限制了执行稳健性和行为多样性。我们提出了GoAnt,一个质量-多样性多智能体搜索框架,它将相互不通信的探索者、利用者和连接者工作者与共享的自适应心智地图以及紧凑的蚁后调度器相结合。心智地图按无泄漏的执行特征组织候选因子,并在每个生态位中保留一个精英因子,而蚁后则基于显式的搜索状态摘要重新分配评估预算。我们还定义了一种与地图无关的有效收益协议,直接从各方法的评估记录中统计高质量且相互不冗余的因子数量,为基于档案的系统与无地图系统提供了统一的衡量标准。在2023至2026年的真实A股微观结构数据上,GoAnt在量价和订单簿两种设置下分别达到41.8和47.6的质量加权收益,在匹配预算下较最强基线分别提升57%和97%。其锁定种群在样本外保留了0.64和0.67的样本内质量,而静态地图分别为0.61和0.63。
cs.AI / 178 / 2609.08729

Application of curiosity driven exploration methods for hardware interference identification

好奇心驱动探索方法在硬件干扰识别中的应用
Matar, Ludovic, Moulin-Frier, Clement, Oudeyer, Pierre-Yves
Abstract
The transition from single-core to multi-core architectures in safety-critical embedded systems introduces significant challenges due to inter-core interference caused by contention for shared hardware resources. Such interference affects execution times and complicates the verification of strict temporal requirements, particularly in domains such as avionics where standards require comprehensive identification of interference sources. Existing interference analysis approaches, whether manual or model-based, struggle to capture the full range of behaviors arising from the complex interactions among micro-architectural components. In this paper, we frame multi-core interference analysis as the exploration of a complex system behavior space. We propose the use of curiosity-driven exploration algorithms from artificial intelligence to systematically and efficiently cover the space of possible interference behaviors. Using a simulator-based environment, we show that the proposed approach achieves broader and more uniform behavioral coverage within a limited experimental budget compared to traditional pseudo-random program generation methods.
Chinese Translation
安全关键嵌入式系统从单核架构向多核架构的转变带来了重大挑战,其原因在于对共享硬件资源的争用会导致核间干扰。此类干扰会影响执行时间,并使严格时间需求的验证变得复杂,尤其是在航空电子等领域,其标准要求全面识别干扰源。现有的干扰分析方法,无论是人工方法还是基于模型的方法,都难以捕捉由微架构组件之间复杂交互所产生的全部行为。在本文中,我们将多核干扰分析构建为对复杂系统行为空间的探索。我们提出使用人工智能中的好奇心驱动探索算法,以系统且高效的方式覆盖可能的干扰行为空间。通过基于仿真器的实验环境,我们表明与传统伪随机程序生成方法相比,所提出的方法在有限的实验预算内实现了更广泛、更均匀的行为覆盖。
cs.AI / 179 / 2609.08736

When Can One Obtain Certificates of Optimality Using Positivstellensaetze?

何时可以借助 Positivstellensatz 获得最优性证书?
Kim, Nayoon, Gehret, Allen, Ma, Shenyuan, Marecek, Jakub
Abstract
We study certificates of positivity and optimality for learning problems whose objectives and constraints need not be polynomial. We isolate an axiomatic core of Fischer's constructive strict and weak Positivstellens\"{a}tze and prove the resulting theorems for abstract function algebras over ordered fields. The framework separates two roles that can otherwise be conflated: objective and constraint functions may be built from broad classes of continuous or definable operations, while the auxiliary primitives used to construct a certificate satisfy explicit scalar and closure axioms. We give instances over continuous and definable function algebras, including ordered fields not closed under square roots, derive lower-bound and global-optimality certificates, and analyze both expanded term length and shared computation-graph complexity.
Chinese Translation
我们研究目标函数与约束函数不一定是多项式的学习问题的正性与最优性证书。我们分离出 Fischer 的构造性严格与弱 Positivstellensatz 的公理化核心,并在序域上的抽象函数代数中证明了所得定理。该框架区分了两种容易被混淆的角色:目标函数与约束函数可以由广泛的连续或可定义运算类构造,而用于构造证书的辅助原语则满足显式的标量与闭包公理。我们给出了在连续与可定义函数代数上的若干实例(包括对平方根不封闭的序域),推导出下界证书与全局最优性证书,并分析了扩展项长度与共享计算图复杂度。
cs.AI / 180 / 2609.08772

It's All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction

关键在于表达方式:信息表示在基于大语言模型的血糖事件预测中的作用
Apicella, Andrea, Arpaia, Pasquale, Orefice, Matteo, Pollastro, Andrea, Prevete, Roberto
Abstract
Large Language Models (LLMs) are increasingly being investigated for physiological time-series prediction, yet their effectiveness may depend not only on the model itself, but also on how physiological information is represented and presented at inference time. This study investigates prompt-based general-purpose LLMs for postprandial hyperglycemia and hypoglycemia prediction in individuals with type 1 diabetes. Using the OhioT1DM dataset, we evaluate multiple open-weight LLMs under zero-shot and few-shot inference across prediction horizons of 30, 60, and 90 minutes. The analysis varies both the textual representation of the available physiological information and the amount of information exposed to the model, ranging from glucose observations alone to derived descriptors and additional contextual variables related to insulin, meals, carbohydrates, and physical activity. Performance is compared with conventional patient-specific supervised models and with Gluco-LLM, a language-model-based architecture explicitly adapted to glucose time-series forecasting. Results show a marked task-dependent behavior. Conventional supervised models achieve the strongest performance for hyperglycemia prediction, whereas the best observed prompt-based LLM configurations improve performance for hypoglycemia across all investigated horizons. The effectiveness of prompt-based inference is also strongly influenced by how physiological information is represented, while providing additional contextual information does not lead to a systematic improvement. Overall, these findings highlight physiological information representation as a central design factor in prompt-based LLM approaches to glycemic-event prediction.
Chinese Translation
大语言模型(LLM)在生理时间序列预测中的应用正受到越来越多的研究,然而其有效性可能不仅取决于模型本身,还取决于生理信息在推理时如何被表示和呈现。本研究探索了基于提示的通用大语言模型在1型糖尿病患者餐后高血糖和低血糖预测中的应用。基于OhioT1DM数据集,我们在30、60和90分钟的预测时间跨度上,对多个开源权重的大语言模型进行了零样本和少样本推理评估。分析中既改变了现有生理信息的文本表示方式,也改变了向模型暴露的信息量,范围从仅提供血糖观测数据,到提供派生描述符以及与胰岛素、饮食、碳水化合物和体力活动相关的额外上下文变量。我们将性能与传统的患者特异性监督模型,以及Gluco-LLM——一种明确针对血糖时间序列预测而适配的基于语言模型的架构——进行了比较。结果显示出明显的任务依赖性行为:传统监督模型在高血糖预测上表现最佳,而观察到的最优基于提示的大语言模型配置则在所有研究的时间跨度上提升了低血糖预测性能。基于提示的推理的有效性还受到生理信息表示方式的强烈影响,而提供额外的上下文信息并未带来系统性的改进。总体而言,这些发现凸显了生理信息表示是基于提示的大语言模型血糖事件预测方法中的核心设计因素。
cs.AI / 181 / 2609.08832

Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course

弥合一致性差距:学会保持航向的自我进化智能体
Duesterwald, Evelyn, Elder, Benjamin, Ngweta, Lilian, Ubaru, Shashanka, Zimon, Malgorzata
Abstract
Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a discrepancy that has been observed but remains largely unaddressed. When given the same task five times, a ReAct agent on the AppWorld benchmark using GPT-4.1 succeeds in all five runs only 53% of the time, even though its per-run pass rate averages 77%. We call this 24-point shortfall the consistency gap, and we argue that addressing it is a precondition for trustworthy AI agent deployment. We present a self-evolving agent framework that reduces this gap by identifying unstable, low-consistency steps in agent trajectories and converting them into episodic memory the agent can draw on in future runs. At its core is a Consistency Analyzer that pinpoints where and why a trajectory is likely to flip across executions, and a Guideline Generator that converts the diagnosis into targeted guidelines, committed to memory and injected into future agent executions on similar tasks. On AppWorld with ReAct/GPT-4.1, our framework raises the fraction of tasks that succeed in all five runs by +16 points on same-task evaluation and +13 points on similar-task generalization.
Chinese Translation
由大语言模型(LLM)驱动的智能体尽管在平均意义上可以表现准确,但在实际生产环境中却可能不可靠,这一差距已被观察到,但在很大程度上仍未得到解决。在AppWorld基准上使用GPT-4.1的ReAct智能体,即使单次运行的平均通过率为77%,同一任务执行五次时五次全部成功的概率仅为53%。我们将这24个百分点的差距称为一致性差距(consistency gap),并认为解决该问题是可信AI智能体部署的前提条件。我们提出了一种自我进化智能体框架,通过识别智能体轨迹中不稳定、低一致性的步骤,并将其转化为智能体在后续运行中可利用的情景记忆,从而缩小这一差距。该框架的核心是一个一致性分析器(Consistency Analyzer),用于精确定位轨迹在多次执行中可能发生翻转的位置和原因;以及一个指南生成器(Guideline Generator),将诊断结果转化为针对性指南,提交至记忆并注入到未来相似任务的智能体执行中。在AppWorld基准上使用ReAct/GPT-4.1时,我们的框架使五次运行全部成功的任务比例在同任务评估中提高了16个百分点,在相似任务泛化中提高了13个百分点。
cs.AI / 182 / 2609.08861

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

API基准测试分数无法可靠地迁移至聊天机器人界面
Wang, Jennifer, Baumann, Joachim, Ho, Daniel E., Koyejo, Sanmi
Abstract
Benchmark scores are a central currency in model releases: they inform purchasing decisions, shape public trust, and influence policy. Yet, a key assumption underlying benchmark scores is that the model performance measured through APIs faithfully reflects the behavior of deployed systems. We challenge this assumption by auditing ChatGPT, Claude, and Gemini across seven systems and nine benchmarks spanning general capability, social bias, and sycophancy. We find systematic API--interface differences in both accuracy and consistency. On average, API evaluations score 3.4 percentage points higher in accuracy and 2.1 percentage points higher in test--retest agreement than corresponding interface evaluations. For ChatGPT, the performance difference between API and interface access exceeds the API-only difference between GPT 5.3 and GPT 5.4. Put differently, switching access surfaces can degrade performance as much as downgrading a full model generation. We further test whether exposed API controls can reproduce interface behavior by varying system prompts, sampling parameters, and reasoning settings. These controls shift behavior in some cases but do not reliably eliminate the gap. Our findings document a context-validity gap: measurements obtained through APIs do not necessarily generalize to corresponding deployed interfaces, complicating the use of API evaluations as proxies for deployed systems.
Chinese Translation
基准测试分数是模型发布中的核心价值指标:它们为采购决策提供依据、塑造公众信任并影响政策制定。然而,基准测试分数的一个关键假设是,通过API测量的模型性能能够真实反映已部署系统的行为。我们通过对ChatGPT、Claude和Gemini在七个系统、九项基准测试上的审计来挑战这一假设,这些基准测试涵盖通用能力、社会偏见和谄媚性。我们发现API与界面之间在准确性和一致性方面存在系统性差异。平均而言,API评估的准确率比相应的界面评估高3.4个百分点,重测一致性高2.1个百分点。对于ChatGPT而言,API访问与界面访问之间的性能差异超过了GPT 5.3与GPT 5.4之间仅基于API的差异。换言之,切换访问渠道所导致的性能下降,可能与降级一整代模型相当。我们进一步测试了通过调整系统提示词、采样参数和推理设置等暴露的API控制项,能否重现界面行为。这些控制项在某些情况下会改变模型行为,但无法可靠地消除差距。我们的研究发现揭示了一种情境效度差距:通过API获得的测量结果不一定能推广到相应的已部署界面,这使得将API评估作为已部署系统代理的做法变得更加复杂。
cs.AI / 183 / 2609.08944

SkillAdam: Stable and Efficient Skill Evolution for Agents

SkillAdam:面向智能体的稳定且高效的技能演化方法
Li, Gaoyuan, Fan, Meihao, Liu, Yizhe, Zhang, Shaolei, Fan, Ju, Wang, Siyi, Hou, Jiaheng, Weng, Xudong, Tian, Honghan, Li, Zang
Abstract
Agent skills provide a lightweight way to equip frozen language-model agents with domain knowledge and procedural guidance, yet obtaining high-quality skills remains costly and difficult to scale. Expert-written skills require substantial human effort. Recent skill self-evolution methods automate an iterative loop that uses execution feedback to revise skills, but their heuristic update strategies often yield unstable optimization and low iteration efficiency. We identify two challenges in realizing stable and efficient skill self-evolution. Direction Stability requires effective corrections to accumulate rather than be overwritten by iteration-local feedback. Update Adaptivity requires the scope of each revision to reflect the consistency of recent case-level improvements. We introduce SkillAdam, an Adam-inspired framework for optimizing discrete and non-differentiable skill documents. As a functional analogue of Adam's first moment, an optimization memory records identified problems and the outcomes of prior solution attempts to stabilize the update direction. As a functional analogue of Adam's second moment, a volatility-driven edit budget tracks the history-weighted variation of recent case-level improvements and adaptively controls the update magnitude. Across seven benchmarks that span short- and long-horizon tasks, SkillAdam achieves state-of-the-art performance with more stable optimization dynamics. It also obtains stronger skills with substantially fewer optimization iterations and lower cost than prior methods. Code repository: https://github.com/ruc-datalab/SkillAdam
Chinese Translation
智能体技能(Agent skills)为冻结的语言模型智能体提供了一种轻量级方式,使其获得领域知识和程序性指导,然而获取高质量技能仍然成本高昂且难以规模化。专家编写的技能需要大量人力投入。近期的技能自我演化方法自动化了一个利用执行反馈迭代修订技能的循环,但其启发式更新策略往往导致优化不稳定且迭代效率低下。我们指出了实现稳定且高效的技能自我演化所面临的两个挑战:方向稳定性(Direction Stability)要求有效的修正能够累积,而不被迭代局部的反馈所覆盖;更新自适应性(Update Adaptivity)要求每次修订的幅度能够反映近期案例级改进的一致性。我们提出了 SkillAdam,一个受 Adam 启发的框架,用于优化离散且不可微分的技能文档。作为 Adam 一阶矩的功能类似物,优化记忆(optimization memory)记录已识别的问题及先前解决方案尝试的结果,以稳定更新方向。作为 Adam 二阶矩的功能类似物,波动驱动的编辑预算(volatility-driven edit budget)追踪近期案例级改进的历史加权变化,并自适应地控制更新幅度。在涵盖短期和长期任务的七个基准测试中,SkillAdam 以更稳定的优化动态实现了最先进的性能。与已有方法相比,它还以显著更少的优化迭代次数和更低的成本获得了更强的技能。代码仓库:https://github.com/ruc-datalab/SkillAdam
cs.AI / 184 / 2609.08965

PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving

PlannerForge:用于自动驾驶运动规划器基于场景测试的LLM智能体
Gao, Yuan, Müller, Sebastian, Piccinini, Mattia, Kaufeld, Marc, Zhang, Yuchen, Schäfer, Finn Rasmus, Song, Qunying, Betz, Johannes
Abstract
Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used to validate Autonomous Driving Systems (ADSs), but it remains a fragmented modular pipeline in which scenario generation, retrieval, modification, ADS execution, and results analysis are performed by separate tools with little interaction. Large Language Model (LLM) agents have shown promise across ADS sub-systems such as perception, planning, and control. However, no prior work covers the whole scenario-based testing pipeline for ADSs with a unified LLM-agent framework. We present PlannerForge, an LLM-agent framework that extends all scenario-based testing stages (from Scenario Generation to ADS Assessment) and adds two further LLM-enhanced stages: ADS Enhancement and ADS Benchmarking. We evaluate PlannerForge with 10 off-the-shelf LLMs across all tasks (Generation, Selection, Modification, Module Routing, Planner Testing, and Enhancement) under 5 prompt conditions. Best-per-task scores range from 0.88 to 1.00, and open-source 20-35B backends match commercial APIs on most tasks. Open-source models such as Qwen3.6:35B match commercial APIs on three of the five tasks. Chaining the modules end-to-end retains 83% / 78% of seed queries (commercial / open). It outperforms Scenario Factory 2.0 (Finkeldei et al., 2025) on natural-language generation (193 vs. 144 executable of 200) and realises 92-96% of requested city, road and vehicle attributes. It outperforms BM25 (Robertson and Zaragoza, 2009) at rank 1 selection (92.0% vs. 67.5%) and From-Words-to-Collisions (Gao et al., 2025) on physically valid edits (>=94% vs. 31%). At N=400, cost-tuning lifts planner success from 50.4% to 70.2% and cuts collisions from 19.0% to 8.4%, without domain-specific fine-tuning.
Chinese Translation
确保自动驾驶的安全性是一项关键挑战。基于场景的测试是用于验证自动驾驶系统(ADS)的系统性流程,但它仍然是一个碎片化的模块化流水线,其中场景生成、检索、修改、ADS执行和结果分析由相互独立的工具完成,彼此之间几乎没有交互。大语言模型(LLM)智能体已在感知、规划和控制等自动驾驶子系统方面展现出潜力。然而,此前尚无工作能够以统一的LLM智能体框架覆盖ADS基于场景测试的完整流水线。我们提出了PlannerForge,这是一个LLM智能体框架,它扩展了基于场景测试的所有阶段(从场景生成到ADS评估),并新增了两个LLM增强阶段:ADS增强与ADS基准测试。我们使用10个现成的LLM,在5种提示条件下对PlannerForge的所有任务(生成、选择、修改、模块路由、规划器测试和增强)进行了评估。各任务的最佳得分在0.88到1.00之间,且开源的20-35B后端模型在大多数任务上可与商业API相媲美。诸如Qwen3.6:35B之类的开源模型在五项任务中的三项上达到了商业API的水平。将各模块端到端串联后,商业/开源模型分别保留了83%/78%的种子查询。在自然语言生成方面,它优于Scenario Factory 2.0(Finkeldei等,2025)(200个场景中可执行场景为193个对比144个),并实现了92-96%的指定城市、道路和车辆属性。在排序第1的选择任务上,它优于BM25(Robertson和Zaragoza,2009)(92.0%对比67.5%);在物理有效编辑方面,它优于From-Words-to-Collisions(Gao等,2025)(≥94%对比31%)。在N=400的设置下,成本调优将规划器成功率从50.4%提升至70.2%,并将碰撞率从19.0%降至8.4%,且无需进行领域特定的微调。
cs.AI / 185 / 2609.08966

Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack

良好的预训练,糟糕的SFT:贯穿整个训练流程的检查点质量
Maskey, Sohir, Scholl, Philipp, Knupp, Jonas, Neitemeier, Pit, Wirges, Sascha
Abstract
Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subsequent training. We show that this assumption can fail in a full 30B mixture-of-experts training pipeline. The checkpoints that perform better after the full downstream training stack also have higher solution density, i.e., retain downstream performance under local weight perturbations.
Chinese Translation
语言模型检查点通常依据预训练损失或基准测试得分来选择,其假设是得分最高的检查点将仍是后续训练的最佳起点。我们证明,在一个完整的300亿(30B)参数混合专家(mixture-of-experts)训练流水线中,这一假设可能失效。在完整下游训练流程之后表现更好的检查点,同时具有更高的解密度(solution density),即在局部权重扰动下仍能保持下游性能。
cs.AI / 186 / 2609.09001

Deposon: An Auditable, Conservation-Guaranteed, Game-Theoretically Tested Scattering Layer over LLM Reasoning Paths

Deposon:一个可审计、守恒保证、经博弈论检验的LLM推理路径散射层
Yuan, Qihao
Abstract
Multi-step LLM reasoning lacks a machine-recheckable ledger: discarded reasoning paths leave no auditable record. We propose the Deposon scattering layer, which binds each node of an LLM-generated concept-decomposition graph to a two-parameter Deposon state; paths undergo three-channel scattering -- transmission, reflection, irreversible dissipation -- obeying T+R+A=1 for arbitrary parameters, with a maximum per-path energy-audit deviation of 2.2E-16 (machine epsilon). We report all three evidence tiers honestly. On synthetic trap benchmarks the path-filtering gain is closed (pre-registered): unified reaches 100% versus a decoy-capture baseline at 7%/10%. On real benchmarks the layer is indistinguishable from a trivial six-keyword rule filter (GSM8K 0.87 >= 0.85, McNemar p=0.5; StrategyQA 0.899 = 0.899); no difference is detected here, so we sharpen the claim to "the differential value lies solely in machine verifiability." Fusion yields a second negative result: convex combinations with a semantic prior never improve (physics 0.484 -> 0.452), and the apparent lambda=2 gain is an anti-field artifact; any fusion gain must be nonlinear. Modeling the reverse dynamics as a potential game on the graph, we evidence an auditable scalar's monotonicity and near-gradientness and quantify the empirical coordination ratio (ECR). The three formalized dynamical-equivalence propositions (P1a/P1b/T-P1c) are falsified under the pre-registered kill protocol, and the potential-game claim is downgraded to approximate (cyclic-graph median residual 0.669): only consistency-level evidence survives at the dynamical level. Code: github.com/zeroandcat/Deposon.
Chinese Translation
多步LLM(大语言模型)推理缺乏可供机器复查的账本:被丢弃的推理路径不留下任何可审计的记录。我们提出Deposon散射层,它将LLM生成的概念分解图中的每个节点绑定到一个双参数Deposon状态;路径经历三通道散射——透射、反射和不可逆耗散——并对任意参数满足T+R+A=1,单路径能量审计的最大偏差为2.2E-16(机器精度)。我们诚实地报告了全部三个证据层级。在合成陷阱基准上,路径过滤增益已被封存(预注册):统一方法达到100%,而诱捕基线仅为7%/10%。在真实基准上,该层与一个平凡的六关键词规则过滤器无法区分(GSM8K 0.87 >= 0.85,McNemar p=0.5;StrategyQA 0.899 = 0.899);此处未检测到差异,因此我们将主张收紧为“其差异化价值仅在于机器可验证性”。融合实验得到第二个负面结果:与语义先验的凸组合从不带来改进(physics 0.484 -> 0.452),且表面上lambda=2的增益是一个反场伪影;任何融合增益必须是非线性的。将逆动力学建模为图上的势博弈,我们证明了可审计标量的单调性与近梯度性,并量化了经验协调比(ECR)。三个形式化的动力学等价命题(P1a/P1b/T-P1c)在预注册的否证协议下被证伪,势博弈主张被降级为近似成立(环图残差中位数0.669):在动力学层面仅有一致性级别的证据得以保留。代码:github.com/zeroandcat/Deposon。
cs.AI / 187 / 2609.09030

Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning

答案分布轨迹:大语言模型推理的随机动力学视角
Català, Mar Gonzàlez I, Borde, Haitz Sáez de Ocáriz, Murari, Davide, Schönlieb, Carola-Bibiane, Liò, Pietro, Montañez, George
Abstract
Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresses this limitation using entropy profiles, which track how uncertainty evolves over the reasoning process but do not reveal which competing hypotheses account for that uncertainty. We introduce answer-distribution trajectories, a stochastic-dynamics-inspired representation that tracks the model's full predictive distribution over answers as reasoning unfolds. As a strictly finer representation than endpoint and entropy summaries, answer-distribution trajectories enable us to characterize a trace through a dynamical reasoning profile spanning exploration, revision, motion, and commitment, and to distinguish different dynamical mechanisms of reasoning success and failure. Across sixteen open-weight language models and four reasoning benchmarks, we show that traces with the same endpoint and similar entropy profiles can exhibit substantially different reasoning dynamics. We further find substantial variation in these dynamics both within and across models and tasks, with different objectives favoring different dynamical profiles. Additionally, we show that training and inference choices systematically reshape these profiles. Our results suggest that answer-distribution trajectories provide a rich framework for analysing and evaluating the dynamics of LLM reasoning.
Chinese Translation
思维链(Chain-of-thought)推理在模型输入与最终答案之间提供了一种结构化的计算过程。然而,该推理过程通常仅通过端点准确率进行评估,这忽略了模型到达答案所经过的路径。近期有一类工作通过熵轮廓(entropy profiles)来弥补这一局限,其追踪了推理过程中不确定性的演化,但无法揭示是哪些相互竞争的假设导致了这种不确定性。我们提出答案分布轨迹(answer-distribution trajectories),这是一种受随机动力学启发的表示方法,能够追踪推理展开过程中模型对答案的完整预测分布。作为比端点摘要和熵摘要严格更精细的表示,答案分布轨迹使我们能够通过一个涵盖探索、修订、运动与承诺的动力学推理轮廓来刻画推理轨迹,并区分推理成功与失败的不同动力学机制。在十六个开源权重语言模型和四个推理基准上,我们展示了具有相同端点和相似熵轮廓的轨迹可能表现出截然不同的推理动力学。我们进一步发现,这些动力学在模型内部、模型之间以及任务之间均存在显著差异,且不同的目标函数偏好不同的动力学轮廓。此外,我们表明训练与推理的选择会系统性地重塑这些轮廓。我们的结果表明,答案分布轨迹为分析和评估大语言模型推理的动力学提供了一个丰富的框架。
cs.AI / 188 / 2609.09056

Time-Varying Data as Sheaves: an Invitation to Narratives

将时变数据视为层(Sheaves):叙事理论导引
Leal, Wilmer, Bumpus, Benjamin Merlin, Nickel, Jana K., García, Johan, Fairbanks, James, Dixon, Warren
Abstract
Modern science and engineering increasingly rely on time-varying data, yet the mathematical tools used to model temporal phenomena are often developed within separate disciplines, obscuring common principles and limiting the transfer of ideas across fields. This chapter presents the theory of narratives, an abstract framework for time-varying objects of any mathematical kind that supports both theoretical investigations and applications. To illustrate this perspective, the chapter develops three vignettes, each illustrating a different research direction. The first addresses a general concern: What information loss can occur when switching between different representations of temporal data? The second concerns structural and algorithmic approaches: How can we systematically decompose time-varying data into simple pieces and obtain invariants describing its structural complexity? The third is an application to control theory: How can we model multi-agent systems with switching communication topologies? More important than any individual vignette, the central message of this invitation is that a suitable abstract perspective can organize and guide research across remarkably diverse mathematical and scientific domains.
Chinese Translation
现代科学与工程日益依赖时变数据,然而用于建模时间现象的数学工具往往是在不同学科中各自独立发展的,这掩盖了共同原理,也限制了思想在不同领域之间的迁移。本章介绍叙事(narratives)理论,这是一个针对任意数学类型的时变对象的抽象框架,既支持理论研究也支持实际应用。为阐释这一视角,本章给出三个案例片段,每个片段展示一个不同的研究方向。第一个片段处理一个普遍性问题:在时变数据的不同表示之间切换时,可能产生怎样的信息损失?第二个片段关注结构与算法方法:我们如何系统地将时变数据分解为简单的部分,并获得描述其结构复杂性的不变量?第三个片段是一个控制论应用:我们如何建模具有切换通信拓扑的多智能体系统?比任何单个片段更重要的是,本导引的核心信息在于:一个合适的抽象视角能够组织并指导横跨极为多样的数学与科学领域的研究。
cs.AI / 189 / 2609.09081

Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training

凡事贵在有度:多领域中期训练中的逐领域覆盖最优区间与抗对齐的领域差距
Xu, Yunpeng, Zheng, Kun
Abstract
Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base, with a 4B replication; five semantically rule-disjoint KOR-Bench domains) we train 30 allocations spanning the five-domain simplex, 24 sweep configurations plus six withheld from the fit, at five seeds each. Three findings emerge. First, every domain has an interior coverage optimum: the moderate band ($10\%$-$40\%$) is best for all five domains, and a calibrated permutation test for quadratic interiority gives $P\approx0.010$; the fitted mid-training-only curves, with 8B peaks between $9.9\%$ and $35.1\%$, reproduce for curve shape but not peak location. Second, the gaps survive a fixed-budget alignment pass: compensatory SFT raises 116/120 cells (mean $+4.32\%$) yet bridges $0/240$ pairs at a $5\%$ threshold and $30/240$ at a $10\%$ ratio, an equal-budget uniform control behaves almost identically, and a permutation null would bridge $13.8\pm3.3$ and $77.9\pm8.5$ pairs ($P<0.001$). Third, zero coverage collapses mid-training-only accuracy, though a FineWeb-Edu-only control shows the collapse is commingled with generic drift. An exploratory $\theta^*$ allocation attains the largest full-pipeline gain ($+4.36\%$ vs. $+0.80\%$/$+0.64\%$\,pp) but is marginal under Welch test.
Chinese Translation
中期训练(mid-training)是介于预训练与对齐之间的阶段,模型的逐领域数据配比通常由数据可得性而非系统性设计决定。我们探究这一决策带来了什么收益,以及后续的对齐过程能否消除其影响。在一个受控的逻辑推理实验环境中(Qwen3-8B-Base,并使用4B模型进行复现;五个语义规则互不相交的KOR-Bench领域),我们在五领域单纯形上训练了30种配比方案,包括24个扫描配置和6个未纳入拟合的保留配置,每种配置训练五个随机种子。研究得出三项发现。第一,每个领域都存在内部覆盖最优区间:所有五个领域的最佳配比均落在适度区间(10%–40%),针对二次型内部最优性的校准置换检验给出P≈0.010;仅基于中期训练的拟合曲线在8B模型上的峰值位于9.9%至35.1%之间,其曲线形状可复现,但峰值位置无法复现。第二,这些差距在固定预算的对齐过程中依然存在:补偿性SFT提升了116/120个实验单元(平均+4.32%),但在5%阈值下未能弥合任何配对差距(0/240),在10%比例下仅弥合30/240;等预算的均匀配比对照表现几乎相同,而置换零假设预计将分别弥合13.8±3.3和77.9±8.5个配对(P<0.001)。第三,零覆盖会使仅中期训练的准确率崩溃,但仅使用FineWeb-Edu的对照实验表明,这种崩溃与一般性的性能漂移混杂在一起。一个探索性的θ*配比方案取得了最大的全流程增益(+4.36%,对比+0.80%/+0.64%个百分点),但在Welch检验下仅为边际显著。
cs.AI / 190 / 2609.09094

The Surprising Effectiveness of Approximate Value Iteration in Self-Play

近似值迭代在自博弈中的惊人有效性
Boige, Raphael, Boumaza, Amine, Scherrer, Bruno
Abstract
Combining search with function approximation has driven major advances in game-playing programs, making self-play algorithms more competitive than ever. Still, the computational overhead of the most popular methods, based on Monte Carlo Tree Search (MCTS), can be substantial. In this work, we investigate whether simpler methods remain competitive in non-trivial, moderately sized games such as Connect Four, Hex(7x7) and synthetic games. We train a minimal self-play implementation of Approximate Value Iteration (AVI) and use ground-truth oracles for exact evaluation. Contrary to expectations, our results demonstrate the surprising effectiveness of AVI: it learns more accurate value functions than those learned by AlphaZero, while its one-step-lookahead greedy policies remain competitive with MCTS-based policies at substantially lower training and inference costs. Preliminary experiments on Othello and Go(9x9) show that AVI trains stably on larger games and learns effective value functions. These findings suggest that the success of MCTS-based methods may have eclipsed simpler approaches that have become increasingly practical with modern deep-learning tools.
Chinese Translation
将搜索与函数近似相结合推动了博弈程序的重大进展,使自博弈算法比以往任何时候都更具竞争力。然而,基于蒙特卡洛树搜索(MCTS)的最流行方法可能带来可观计算开销。在本工作中,我们研究了更简单的方法在非平凡的、中等规模的博弈(如Connect Four、Hex(7x7)以及合成博弈)中是否仍具竞争力。我们训练了一个极简的自博弈实现的近似值迭代(Approximate Value Iteration, AVI),并使用真实基准(ground-truth)oracle进行精确评估。与预期相反,我们的结果证明了AVI的惊人有效性:它学到的价值函数比AlphaZero学到的更准确,而其单步前瞻贪婪策略在训练和推理成本大幅降低的情况下,仍与基于MCTS的策略具有竞争力。在Othello和围棋(9x9)上的初步实验表明,AVI在更大的博弈中能够稳定训练并学到有效的价值函数。这些发现表明,基于MCTS方法的成功可能掩盖了那些随着现代深度学习工具的发展而日益实用的更简单方法。
cs.AI / 191 / 2609.09113

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

SAEScientist-Bench:AI智能体能够开展自主的SAE可解释性研究吗?
Tan, Yuqiao, He, Shizhu, Zhao, Jun, Liu, Kang
Abstract
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Bench to evaluate whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery. Given a target concept, an agent designs contrastive probes and navigates a Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT to discover the optimal feature, evaluated against curated expert reference features anchored on Neuronpedia across activation rank, concept selectivity on contrastive texts, and causal steering. Across 10 agent configurations and 20 tasks, frontier agents demonstrate genuine discovery capabilities and lead different evaluation dimensions, but remain well behind the expert baseline, approaching expert levels on separating target concepts from contrastive controls while lagging substantially in causal generation steering. Further analysis reveals that although agents can design contrasts to rule out spurious candidates, they frequently misinterpret experimental measurements. These results establish experimental model understanding as a measurable capability for closed-loop autonomous AI R&D. Our code is available at https://github.com/Trae1ounG/SAEScientist.
Chinese Translation
尽管递归自我改进(RSI)的研究主要集中于自动化模型训练流程,但可靠的自主开发还需要一个缺失的支柱:事后监测与审计,以理解模型学到了什么并确保安全对齐。机制可解释性工具对于弥合这一差距至关重要,其中稀疏自编码器(Sparse Autoencoders, SAEs)通过分离可解释特征用于模型检查与引导而成为基石。本文提出SAEScientist-Bench,用于评估AI智能体能否作为科学家,利用SAE工具进行自主的机制发现。给定一个目标概念,智能体设计对比探针,并在Gemma-2-9B-IT的包含13万以上特征的Gemma Scope字典中进行检索,以发现最优特征。评估基于锚定于Neuronpedia的专家参考特征,从激活排名、在对比文本上的概念选择性以及因果引导三个维度进行。在10种智能体配置和20个任务上,前沿智能体展现出真实的发现能力,并在不同评估维度上各有领先,但仍显著落后于专家基线:它们在将目标概念与对比控制区分开来时接近专家水平,而在因果生成引导方面则明显滞后。进一步分析表明,尽管智能体能够设计对比实验以排除虚假候选特征,但它们经常误读实验测量结果。这些结果表明,实验性的模型理解可以成为闭环自主AI研发中一个可度量的能力。我们的代码可在 https://github.com/Trae1ounG/SAEScientist 获取。
cs.AI / 192 / 2609.09115

MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents

MeClear:面向长程LLM智能体的合作博弈论归因与风险感知记忆清除方法
Yang, Boyu, Sun, Jiazheng, Lu, Zilong, Qiu, Zhi, Peng, Xin, Zheng, Jun
Abstract
Long horizon Large Language Model (LLM) agents rely on external memory systems to preserve user preferences and task knowledge across extended interactions. Conventional retrieval mechanisms optimize semantic compatibility rather than downstream utility, frequently introducing outdated, misleading, or conflicting evidence into the active context. We present MeClear, a task conditioned memory clearance framework that identifies memories featuring negative downstream utility through cooperative attribution and selectively suppresses them from agent execution. MeClear combines Leave One Out screening with sampled cooperative Shapley attribution to distribute utility across interacting evidence, effectively resolving redundant conflict masking where single removal evaluations fail. Utilizing attribution rankings, MeClear executes a query scoped minimal clearance strategy over a nested filtration, verifying task recovery on the cleared context without permanently altering the persistent memory bank. Comprehensive experimental evaluations across ten long dialogue memory pools demonstrate that MeClear achieves a target recall of 85.9% and an overall task recovery rate of 82.3%, representing a 25.5 percentage point improvement over Leave One Out (LOO) baselines.
Chinese Translation
长程大语言模型(LLM)智能体依赖外部记忆系统在长时交互中保存用户偏好与任务知识。传统检索机制优化的是语义兼容性而非下游效用,常常将过时、误导或冲突的证据引入当前上下文。我们提出MeClear,一种任务条件化的记忆清除框架,通过合作博弈归因识别具有负面下游效用的记忆,并有选择地将其从智能体执行中抑制。MeClear将留一法(Leave One Out)筛选与采样的合作Shapley归因相结合,在相互作用的证据之间分配效用,有效解决了单个移除评估无法处理的冗余冲突掩盖问题。基于归因排名,MeClear在嵌套过滤结构上执行查询范围内的最小清除策略,并在清除后的上下文上验证任务恢复,且不会永久性地改动持久记忆库。在十个长对话记忆池上的全面实验评估表明,MeClear实现了85.9%的目标召回率和82.3%的总体任务恢复率,相较留一法(LOO)基线提升了25.5个百分点。
cs.AI / 193 / 2609.09126

A Generalization of Amari's Bayesian Duality

Amari贝叶斯对偶性的推广
Khan, Mohammad Emtiyaz, Möllenhoff, Thomas
Abstract
Amari's contributions to information geometry and machine learning are well known. Here, we revisit Amari's work on Bayesian duality which has not received as much attention. We connect Amari's Bayesian duality to a convex duality of Bayes' rule. Using this connection, we present a generalization of Amari's Bayesian duality and discuss its relevance for modern artificial intelligence.
Chinese Translation
Amari对信息几何和机器学习的贡献广为人知。本文重新审视Amari关于贝叶斯对偶性的工作,该工作尚未得到足够的关注。我们将Amari的贝叶斯对偶性与贝叶斯法则的一种凸对偶性联系起来。利用这一联系,我们提出了Amari贝叶斯对偶性的一种推广,并讨论其对现代人工智能的相关性。
cs.AI / 194 / 2609.09133

ExecCritic: Learn to Test, Test to Improve for Coding Agents

ExecCritic:学会测试,以测试促进改进的编码智能体
Tao, Leitian, Peng, Baolin, Wang, Haorui, Wang, Hang, Cheng, Hao, Yao, Wenlin, Wu, Qianhui, Ge, Tao, Li, Sharon, Gao, Jianfeng
Abstract
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create false confidence. We introduce ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it. The scaffold separates test construction from source-code repair: a Test agent independently generates repository-native tests, a fail-closed harness qualifies and freezes them, and a Repair agent revises source code from their execution feedback without changing the tests. Both roles use Qwen-3.5-35B-A3B as the backbone and are trained separately. In Learn to Test, the Test agent learns to produce behaviorally valid tests that distinguish correct from incorrect patches. In Test to Improve, the Repair agent learns both direct task resolution and feedback-guided revision. On SWE-bench Verified, test quality determines whether feedback helps: holding the base Repair agent fixed, tests from the base Test agent reduce resolved rate from a no-test baseline of 61.2% to 57.3%, whereas tests from GPT-5.6-sol raise it to 65.3%. Role-specific post-training raises the Qwen Test agent's Base-to-Gold success from 22.2% to 62.2%; composing the two post-trained Qwen agents reaches 72.6%, an 11.4-point gain over the original no-test baseline without stronger-model or Oracle feedback at evaluation time. Code is publicly available at https://github.com/MSR-Orchard/execcritic.
Chinese Translation
执行反馈可以引导编码智能体实现正确的仓库修复,但前提是测试能够捕捉到问题所要求的行为。智能体生成的测试可能编码了不完整或不正确的行为目标;当同一条轨迹既编写补丁又编写测试时,两者的错误可能相互一致,从而产生虚假的信心。我们提出了 ExecCritic,将“测试—验证—修订”框架与针对其中角色定制的强化学习训练方案相结合。该框架将测试构造与源代码修复分离:一个 Test 智能体独立生成仓库原生的测试,一个“失败即关闭”的测试框架(fail-closed harness)对其进行资格确认并冻结,随后一个 Repair 智能体根据测试的执行反馈修订源代码,而不改动测试。两个角色均以 Qwen-3.5-35B-A3B 为骨干模型并分别进行训练。在“学会测试”(Learn to Test)阶段,Test 智能体学习生成能够区分正确与错误补丁的行为有效的测试。在“以测试促进改进”(Test to Improve)阶段,Repair 智能体既学习直接解决任务,也学习基于反馈的修订。在 SWE-bench Verified 上,测试质量决定了反馈是否有帮助:在保持基础 Repair 智能体不变的情况下,使用基础 Test 智能体的测试会使解决率从无测试基线的 61.2% 降至 57.3%,而使用 GPT-5.6-sol 的测试则将其提升至 65.3%。针对角色的后训练将 Qwen Test 智能体的 Base-to-Gold 成功率从 22.2% 提升至 62.2%;将两个经过后训练的 Qwen 智能体组合使用达到 72.6%,相比原始无测试基线提升了 11.4 个百分点,且在评估时无需更强模型或 Oracle 反馈。代码已在 https://github.com/MSR-Orchard/execcritic 公开。
cs.AI / 195 / 2609.09134

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

协同演化测试框架与模型:策略内修正帮助弱模型在模仿失败之处迎头赶上
Yu, Zhou, Bi, Bin, Pentyala, Shiva Kumar, Mehrotra, Shubham, Chaudhuri, Sougata, Bhagavath, Shilpa, Chen, Zeyuan, Xu, Ran, Mui, Phil, Zhu, James, Asur, Sitaram
Abstract
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success. Automated harness evolution can enable smaller models to perform well on domain-specific tasks at a fraction of frontier-model cost. Since both the harness and model weights shape behavior, we ask how harness evolution and lightweight fine-tuning should be combined. Across seven enterprise agent tasks, we first evolve a harness with the weaker model, then find that a stronger expert often uses it more effectively, suggesting expert supervision could close the remaining gap. However, training the weaker model on the expert's complete trajectories under the evolved harness backfires: performance regresses on all seven tasks by 4 to 30 points across Qwen3-Coder and Gemma 4, even though the same procedure helps under the unevolved harness. Our analysis shows that imitation transfers knowledge and increases scaffold usage, but disrupts model-harness fit: the weaker model adopts the expert's planning strategy without the competence to execute it and no longer matches the harness evolved around its native planning style. We therefore develop an on-policy expert-correction pipeline, automated by a meta-level MLE agent, that localizes the failing turn in the weaker model's own rollout and asks the expert to rewrite only that turn. This preserves the model's planning style and combines the gains of harness evolution and model adaptation. Our results identify and resolve a source of contention between harness and weight updates, yielding a compatibility-preserving recipe for economical co-evolution on domain-specific enterprise tasks.
Chinese Translation
智能体测试框架(Agent harness,即围绕模型的系统提示词、工具集、执行钩子以及上下文管理脚手架)是决定智能体任务成功的关键因素。自动化的框架演化能够使较小的模型在特定领域任务上表现出色,而成本仅为前沿模型的一小部分。由于框架和模型权重共同决定行为,我们探讨了框架演化与轻量级微调应如何结合。在七个企业级智能体任务上,我们首先使用弱模型演化框架,随后发现更强的专家模型往往能更有效地使用该框架,这表明专家监督可能弥合剩余差距。然而,在演化后的框架下用专家的完整轨迹训练弱模型反而适得其反:在 Qwen3-Coder 和 Gemma 4 上,七个任务的性能均倒退了 4 至 30 分,尽管同样的流程在未演化的框架下是有效的。我们的分析表明,模仿能够迁移知识并增加脚手架的使用,但破坏了模型与框架的匹配性:弱模型采纳了专家的规划策略却不具备执行该策略的能力,并且不再与围绕其原有规划风格演化出的框架相匹配。为此,我们开发了一个由元级 MLE 智能体自动化的策略内专家修正流程:该流程在弱模型自身的 rollout 中定位失败的回合,并让专家仅重写该回合。这既保留了模型的规划风格,又结合了框架演化与模型适配的收益。我们的研究结果识别并解决了一个框架与权重更新之间的冲突来源,为特定领域企业任务上的低成本协同演化提供了一种保持兼容性的方案。
cs.AI / 196 / 2609.09137

A Data-Driven Framework for Identifying and Prioritizing RPA Opportunities in Healthcare Processes

一种用于识别和优先排序医疗流程中RPA(机器人流程自动化)机会的数据驱动框架
Gomez, Maria Alejandra, Castillo, Juan Manuel
Abstract
Robotic Process Automation (RPA) is widely used to reduce administrative burden in United States hospitals, yet an estimated 30-50% of RPA initiatives underperform because processes are selected informally, without a repeatable method to catalogue candidates, prioritize them, match each to an automation tier -- a Python bot, an open-source orchestrator such as n8n, or an enterprise platform such as UiPath -- and forecast financial return before committing resources. We propose a four-module, data-driven framework unifying these decisions: a Process Taxonomy of twenty recurring hospital processes across five value streams; a Prioritization module deriving an Automation Suitability Index from an Analytic Hierarchy Process matrix with an explicit consistency check; a Tool-Tier Selection module recommending the least-cost technology sufficient for a process complexity, integration, and compliance profile; and a Return-on-Investment module quantifying labor savings, error-cost avoidance, payback, and net present value. Applied to a synthetic portfolio spanning all twenty processes, plus a reference data-flow architecture linking it to hospital EHR/payer/ERP systems: 12 of 20 clear the prioritization threshold; the ranking is robust to +/-20% weight perturbation (Spearman correlation 0.83, top-5 set preserved 97.7%, 2,000 Monte Carlo trials); an Automation Risk Index flags four qualifying processes as Critical risk; a budget-constrained portfolio optimization shows diminishing marginal NPV as spend scales from $400K to $1.03M; and a second Monte Carlo analysis shows portfolio NPV stays positive at its 5th percentile. The framework is a conceptual synthesis of the literature rather than an instrument calibrated on primary hospital data; we discuss HIPAA governance and a research agenda for empirical validation. A supplementary Python implementation accompanies the paper.
Chinese Translation
机器人流程自动化(Robotic Process Automation, RPA)被广泛用于减轻美国医院的管理负担,然而据估计30-50%的RPA项目表现不佳,原因在于流程的选择较为随意,缺乏一种可重复的方法来编目候选流程、对其排序、将每个流程匹配到相应的自动化层级——即Python机器人、n8n等开源编排器,或UiPath等企业级平台——并在投入资源之前预测财务回报。我们提出一个包含四个模块的数据驱动框架,将这些决策统一起来:一是流程分类学(Process Taxonomy),涵盖五个价值流中的二十个医院常见重复流程;二是优先排序模块(Prioritization),基于层次分析法(Analytic Hierarchy Process)矩阵(附带明确的一致性检验)推导出自动化适配指数(Automation Suitability Index);三是工具层级选择模块(Tool-Tier Selection),根据流程的复杂度、系统集成和合规性特征,推荐满足需求的最低成本技术;四是投资回报模块(Return-on-Investment),量化人力节省、差错成本规避、投资回收期和净现值。我们将其应用于覆盖全部二十个流程的合成项目组合,并给出将其与医院EHR/支付方/ERP系统相连接的参考数据流架构。结果表明:20个流程中有12个超过优先排序阈值;排序对±20%的权重扰动具有稳健性(Spearman相关系数0.83,前5名集合保持率达97.7%,基于2,000次蒙特卡洛试验);自动化风险指数(Automation Risk Index)将四个合格流程标记为“关键”风险;预算约束下的项目组合优化显示,随着支出从40万美元增至103万美元,边际净现值递减;第二次蒙特卡洛分析表明,项目组合净现值在第5百分位数处仍保持为正。该框架是对文献的概念性综合,而非基于医院一手数据校准的工具;我们讨论了HIPAA治理问题以及实证验证的研究议程。论文随附一个补充性Python实现。
cs.AI / 197 / 2609.09153

Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

程序图:面向大语言模型智能体的自演化执行结构
Lu, Yuxing, Chen, Yicheng, Wu, Shanchan, Arık, Sercan Ö.
Abstract
Large language models are increasingly deployed as agents that plan over long horizons and act through external tools. Most agents select actions through unconstrained generation over an accumulating history, leaving implicit the procedural knowledge of what to do, in what order, and under which conditions. As trajectories lengthen, agents can lose track of their objectives, invoke tools out of order, and repeat unproductive actions. We introduce the Procedural Graph: just as a knowledge graph organizes factual knowledge into (entity, relation, entity) triplets for what-is questions, a Procedural Graph organizes procedural knowledge into (procedure, relation, procedure) triplets for what-to-do questions. At each decision step, the framework localizes the agent's active node, and a guidance model translates the surrounding subgraph into step-level situational guidance that biases the solver's next action without dictating it. The graph is self-evolving: an LLM refiner contrasts failed trajectories with successful ones and edits the graph's topology and attributes, committing edits that preserve or improve held-out validation performance while retaining rejected ones to discourage repetition. Starting from a minimal skeleton, the loop builds graphs that match or surpass hand-designed ones. It can also repair a flawed expert prior. Across multiple datasets, task types, and LLMs, the Procedural Graph delivers consistent gains over memory-based baselines, and self-evolution further improves performance without manual engineering.
Chinese Translation
大语言模型越来越多地被部署为需要在长时域上进行规划并通过外部工具执行操作的智能体。大多数智能体通过对不断累积的历史记录进行无约束生成来选择动作,这使得关于做什么、以何种顺序做以及在何种条件下做的程序性知识处于隐式状态。随着轨迹变长,智能体可能迷失目标、乱序调用工具,以及重复无效动作。我们提出了程序图(Procedural Graph):正如知识图谱将事实知识组织为(实体,关系,实体)三元组以回答“是什么”的问题一样,程序图将程序性知识组织为(程序,关系,程序)三元组以回答“做什么”的问题。在每个决策步骤中,该框架定位智能体当前激活的节点,并由一个引导模型将周围子图转化为步骤级情境引导,从而在不直接指定动作的前提下影响求解器的下一步行动。该图是自演化的:一个LLM精炼器将失败轨迹与成功轨迹进行对比,并编辑图的拓扑结构和属性,仅提交能够保持或提升留出验证集性能的编辑,同时保留被拒绝的编辑以避免重复犯错。从最小骨架出发,该循环能够构建出与人工设计相当甚至更优的图,还能修复有缺陷的专家先验。在多个数据集、任务类型和LLM上,程序图相较于基于记忆的基线方法带来了一致的性能提升,且自演化机制无需人工工程即可进一步提升性能。
机器学习 (Machine Learning)
253
cs.LG / 1 / 2609.05435

AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning

AhaBench:智能体能否从先前经验中学习?一个长程持续学习基准
Cheng, Zerui, Xu, Jiawei, Chai, Huacan, Sun, Jiayang, Viswanath, Pramod, Pan, Maxm
Abstract
Modern language agents are expected to operate over long horizons: they ask follow-up questions, reuse worked examples, handle tool feedback, and adapt to delayed consequences. Most evaluations still reset the agent after a prompt or score only the final state of one trajectory. AhaBench asks a more operational question: when a fixed model receives useful experience, does its later behavior improve under a related evaluation condition where the obvious support has been removed, changed, or delayed? The suite contains three components. Aha-Puzzle tests no-hint exploration after solved hidden-state puzzles; Aha-Euler turns Project-Euler-style mathematical ideas into generated taught/held-out tasks with exact validators; and Aha-Vending, an open-source implementation inspired by Vending-Bench, tests whether a simulated vending agent remains profitable while handling delayed feedback and operational incidents. AhaBench reports a three-part scorecard: Initial Score measures starting competence, Post-Experience Score measures the later empirical outcome, and Learning Lift is their difference. This decomposition is the main empirical message: models that use visible support well, models that reach high post-experience scores, and models that improve most during a run are not always the same. On the common eight-model panel, Claude Opus 4.6 leads aggregate Post-Experience Score at 64.3 and aggregate Learning Lift at +25.8, with Gemini 3.1 Pro close behind at 63.4. The component results explain the split: puzzle traces raise supported scores but often fail to become no-hint exploration behavior; Aha-Euler full teaching reaches 78.6-100.0% while answer-only transfer ranges from 0.0 to 73.9%; and Aha-Vending separates profitable incident handling from bankruptcy and no-order failure. We release benchmark tasks, rubrics, validators, simulator code, and interfaces for evaluating new agents.
Chinese Translation
现代语言智能体被期望能够在长时程上运行:它们会提出后续问题、复用已解决的示例、处理工具反馈,并适应延迟出现的后果。然而,大多数评估仍然在每次提示后重置智能体,或仅对单条轨迹的最终状态进行评分。AhaBench 提出了一个更具操作性的问题:当一个固定的模型获得有用经验后,在一个明显的支持被移除、改变或延迟的相关评估条件下,其后续行为是否会有所改进?该基准包含三个组件:Aha-Puzzle 在解开隐藏状态谜题之后测试无提示探索能力;Aha-Euler 将 Project Euler 风格的数学思想转化为由精确验证器评估的已教授/留存任务;Aha-Vending 是受 Vending-Bench 启发的开源实现,测试模拟的自动售货机智能体在处理延迟反馈和运营事故时能否保持盈利。AhaBench 报告三部分评分:初始分数(Initial Score)衡量起始能力,经验后分数(Post-Experience Score)衡量后续的实际结果,学习提升(Learning Lift)为两者之差。这一分解是本文的主要实证结论:善于利用可见支持的模型、达到较高经验后分数的模型,以及在运行过程中进步最大的模型,并不总是同一批模型。在通用的八模型测试面板上,Claude Opus 4.6 以 64.3 的综合经验后分数和 +25.8 的综合学习提升位居第一,Gemini 3.1 Pro 以 63.4 紧随其后。各组件的结果解释了这一分化:谜题轨迹虽然能提高有支持条件下的分数,但往往难以转化为无提示的探索行为;Aha-Euler 完整教学可达到 78.6%–100.0%,而仅答案的迁移范围则为 0.0%–73.9%;Aha-Vending 则能够区分盈利的事故处理与破产和无订单失败。我们发布了基准任务、评分标准、验证器、模拟器代码以及用于评估新智能体的接口。
cs.LG / 2 / 2609.05508

When Do Options Help? Policy Necrosis and Redundant Coverage in Option-Critic

选项何时有效?Option-Critic 中的策略坏死与冗余覆盖
Liu, Bingyun, Jing, Yuheng
Abstract
Option-critic learns options: sub-policies together with a learned rule for when each one hands control back. Its headline result is that performance improves as options are added. We explain that result, with theory and experiment. First, the termination rule option-critic learns by maximising return contributes nothing. When the termination test and the policy that picks options read the same values, the test fires at every step, so the learned rule is identical to always terminating. When that policy explores and the test does not, as in option-critic itself, the rule can block the exploration; there are instances where it suffers $\Omega(T)$ regret while always terminating holds to $O(\log T)$. Forcing termination at every step leaves the option-count curve intact. Second, the policy inside an option barely explores at all, so a state locks onto the first action that looked good and never updates again. We name this policy necrosis, give a state-level test for it, and find three fifths of states necrotic in a typical option. Restoring exploration repairs those states, and one option then solves the task. Third, extra options improve no option; what falls is the chance that all of them fail in the same state, from $59\\%$ to $4\\%$, and performance follows that joint quantity.
Chinese Translation
Option-critic 学习选项(options):即子策略以及每个子策略何时将控制权交还的学习规则。其标志性结论是:随着选项数量的增加,性能会得到提升。我们通过理论与实验对该结论进行了解释。首先,Option-critic 通过最大化回报学到的终止规则没有任何贡献。当终止测试与选择选项的策略读取相同的值时,终止测试会在每一步都触发,因此学习到的规则与总是终止等价。当该策略进行探索而终止测试不探索时(如 Option-critic 本身),该规则可能阻碍探索;在某些实例中,它会产生 Ω(T) 的遗憾,而总是终止的策略遗憾仅为 O(log T)。强制每一步终止后,选项数量与性能的曲线保持不变。其次,选项内部的策略几乎不进行探索,因此一个状态会锁定在最初看起来不错的动作上,此后不再更新。我们将这种现象命名为策略坏死(policy necrosis),给出了一个状态层面的检验方法,并发现一个典型选项中五分之三的状态发生了坏死。恢复探索可以修复这些状态,此后仅一个选项就能解决任务。第三,额外的选项并不能改进任何单个选项;降低的是所有选项在同一状态下同时失败的概率,从 59% 降至 4%,而性能也跟随这一联合量变化。
cs.LG / 3 / 2609.05574

Multi-granularity Adaptive Hypergraph Representation Learning via Granular-ball

基于粒球的多粒度自适应超图表示学习
Zhao, Sen, Guan, Yifan, Ni, Jinyuan, Xu, Gaojie, Xu, Zhang, Lian, Xiaoyu, Liu, Yi, Wang, Yi, Wang, Wei
Abstract
Hypergraph representation learning aims to capture high-order information in graphs by constructing hyperedges that simultaneously connect multiple nodes. These hyperedges adapt to the graph's topological features, facilitating the extraction of high-order relationships at multiple granularities. Most prior work relies on predefined definitions to generate hyperedges, overlooking the diversity in graph topological structures and the multi-granularity characteristics of hyperedges. As a result, this limits their ability to effectively and adaptively discover high-order relationships and efficiently process complex structural information. To address this limitation, we propose a novel framework called \underline{M}ulti-\underline{G}ranularity \underline{H}ypergraph \underline{R}epresentation \underline{L}earning (MGHRL). MGHRL introduces an Adaptive Granular Hypergraph Generation strategy, which generates hyperedges at multiple levels of granularity through the adaptive splitting of granular-ball, effectively capturing high-order relationships based on the graph's topological structure. Additionally, we propose a Multi-Granularity Hypergraph Network with multiple sub-networks, capturing features from hyperedges at different granularities and integrating them via hierarchical reversible connections. Experimental results show that MGHRL significantly outperforms baseline models on benchmark datasets.
Chinese Translation
超图表示学习旨在通过构造同时连接多个节点的超边来捕获图中的高阶信息。这些超边能够适应图的拓扑特征,有助于提取多粒度的高阶关系。现有的大多数工作依赖预定义的方式来生成超边,忽视了图拓扑结构的多样性以及超边的多粒度特性,从而限制了其有效、自适应地发现高阶关系并高效处理复杂结构信息的能力。针对这一局限,我们提出了一种名为多粒度超图表示学习(Multi-Granularity Hypergraph Representation Learning,MGHRL)的新框架。MGHRL 引入了一种自适应粒度超图生成策略,通过对粒球(granular-ball)的自适应分裂在多个粒度层次上生成超边,从而基于图的拓扑结构有效捕获高阶关系。此外,我们提出了一个包含多个子网络的多粒度超图网络,从不同粒度的超边中捕获特征,并通过层次化可逆连接对其进行融合。实验结果表明,MGHRL 在基准数据集上显著优于基线模型。
cs.LG / 4 / 2609.05575

Capsule Lens: Locating and Tracking Concept Geometry in Model Representations

胶囊透镜:在模型表示中定位与追踪概念的几何结构
Tang, Yiming, Saini, Harshvardhan, Jha, Samyak, Chen, Huaming, Duan, Xufeng, Liu, Dianbo
Abstract
Understanding how concepts are encoded in the internal representations of machine learning models is a central problem in mechanistic interpretability, essential both for the science of deep learning and for the trustworthy deployment of increasingly capable models. Existing approaches to interpret model representations mainly map representations onto more interpretable spaces and do not directly characterize how concepts occupy representation space; various hypotheses have been proposed, but often lack of rigorous validation and largely focus on static representations. In this work, we introduce Capsule Lens, a framework that matches the region a concept occupies with a simple, trackable geometric form, a capsule, defined by several interpretable parameters, fitted in closed form to each concept's geometry and validated on held-out samples. We apply Capsule Lens in two major settings: static and dynamic representations. On static representations, we demonstrate how to locate concept geometry across various models, and how the span and norm curves uncover important geometric characteristics. On dynamic representations, we present three case studies tracking representation drifts induced by distinct training settings, CLIP pretraining, RL post-training on visual question answering, and RL post-training on mathematical reasoning. These analyses reveal qualitatively different geometric dynamics, ranging from broad network-wide restructuring in CLIP pretraining to localized and concept-specific changes in RL post-training. Our results include findings aligned with existing literature as well as novel observations. We believe Capsule Lens stands as a promising tool for locating, analyzing, and tracking concept geometry in both static and dynamic representations.
Chinese Translation
理解概念如何被编码在机器学习模型的内部表示中,是机制可解释性领域的核心问题,这对于深度学习科学以及日益强大模型的可信部署都至关重要。现有的模型表示解释方法主要将表示映射到更可解释的空间中,而未能直接刻画概念如何在表示空间中分布;尽管已有多种假设被提出,但往往缺乏严格的验证,且大多聚焦于静态表示。在本工作中,我们提出了Capsule Lens(胶囊透镜),该框架将概念所占的区域与一种简单且可追踪的几何形状——胶囊——相匹配。胶囊由若干可解释的参数定义,以闭式解的方式拟合到每个概念的几何结构,并在留出样本上进行验证。我们将Capsule Lens应用于两大场景:静态表示与动态表示。在静态表示方面,我们展示了如何在多种模型中定位概念的几何结构,以及其跨度(span)曲线和范数曲线如何揭示重要的几何特征。在动态表示方面,我们通过三个案例研究追踪了不同训练设置所引起的表示漂移:CLIP预训练、视觉问答上的强化学习后训练,以及数学推理上的强化学习后训练。这些分析揭示了性质上截然不同的几何动态——从CLIP预训练中广泛的网络级重构,到强化学习后训练中局部化和概念特异性的变化。我们的结果既包含与现有文献一致的发现,也包含新颖的观察。我们相信,Capsule Lens将成为在静态和动态表示中定位、分析和追踪概念几何结构的一个有前景的工具。
cs.LG / 5 / 2609.05582

HB-PVI: A Hierarchical Bayesian Personalization and Value-of-Information Framework for Complex Activity Recognition

HB-PVI:面向复杂活动识别的分层贝叶斯个性化与信息价值框架
Olayinka, Hammed A.
Abstract
Personalization can improve activity-recognition performance, but participant-specific gains are heterogeneous, and every additional calibration label has an acquisition cost. This study presents HB-PVI, a hierarchical Bayesian personalization and value-of-information framework jointly modeling participant heterogeneity, the benefit and harm of four personalization mechanisms, and the economic value of an additional label, for the 47-participant MUSIC-CAR complex-activity cohort. A leakage-safe, leave-one-participant-out evaluation combines a sequential-Monte-Carlo participant-effect updater with a Student-$t$ hierarchical gain model and a one-step expected-value-of-sample-information (EVSI) stopping rule. Adapter personalization produced small positive mean F1 gains, growing from 0.00099 at one label to 0.00198 at ten, while adapter-plus-head and prototype-residual personalization were negative on average. Under the primary practical-benefit threshold ($\Delta_{\min}=0.01$) and cost setting, one-step EVSI was zero at every decision state, so the policy purchased no labels and retained population inference for all 47 participants, matching always-stop exactly (region-of-practical-equivalence probability $=1$). Relative to fixed ten-shot adapter personalization, this reduced labeling by 100\% while keeping the posterior mean F1 loss at 0.00217 (95\% credible interval, 0.00048 to 0.00389), with posterior probability 0.9992 of remaining below the 0.005 tolerance. HB-PVI was utility-optimal in 199 of 216 cost-threshold settings and in every setting at or above the primary label cost. These results argue for a population-first deployment policy whenever personalization gains are small relative to labeling, computation, and harm costs, and show that value-of-information reasoning, not raw predictive accuracy, should drive personalization decisions in health-sensing applications.
Chinese Translation
个性化可以提升活动识别性能,但参与者个体的收益存在异质性,且每增加一个校准标签都有获取成本。本研究提出了HB-PVI(分层贝叶斯个性化与信息价值框架),针对包含47名参与者的MUSIC-CAR复杂活动队列,联合建模参与者异质性、四种个性化机制的收益与危害,以及额外标签的经济价值。评估采用防泄露的留一参与者法,将序贯蒙特卡洛参与者效应更新器与学生t分布分层增益模型以及单步样本信息期望价值(EVSI)停止规则相结合。适配器个性化产生了较小的正平均F1增益,从单标签时的0.00099增长到十标签时的0.00198,而适配器加头部(adapter-plus-head)和原型残差个性化平均为负。在主要实用收益阈值(Δ_min=0.01)和成本设置下,单步EVSI在每个决策状态均为零,因此策略未购买任何标签,对全部47名参与者保留了群体层面的推断,与始终停止策略完全一致(实用等效区域概率=1)。相对于固定的十样本适配器个性化,该方法将标注成本降低了100%,同时后验平均F1损失仅为0.00217(95%置信区间为0.00048至0.00389),保持在0.005容差以下的后验概率为0.9992。在216个成本阈值设置中,HB-PVI在199个设置中达到效用最优,并且在所有等于或高于主要标签成本的设置中均为最优。这些结果表明,当个性化收益相对于标注、计算和危害成本较小时,应优先采用群体层面的部署策略,并证明在健康感知应用中,应由信息价值推理而非原始预测精度来驱动个性化决策。
cs.LG / 6 / 2609.05650

Endogenous Exploration in Reinforcement Learning with Intrinsic Curiosity

基于内在好奇心的强化学习内生探索
Vieira, Armando
Abstract
We propose a reinforcement learning framework in which exploration is driven by intrinsic curiosity, designed for scenarios where environments are non-stationary and rewards are sparse, delayed, uninformative, or absent. In our model, action selection is guided by a combination of external rewards and an epistemic motivation mechanism that biases the agent toward structured exploratory directions. The central hypothesis is that effective exploration emerges at intermediate levels of incoherence, while performance degrades under both overly rigid and overly disordered dynamics. To test this idea, we implement the framework on top of a Liquid State Machine (LSM) substrate and evaluate it on two standard benchmarks: the discrete-action LunarLanderv2 and the continuous-control BipedalWalkerv3. The proposed method achieves competitive performance on both tasks relative to established deep RL algorithms, including Proximal Policy Optimization (PPO) and Intrinsic Curiosity Module (ICM). We further show that the curiosity window is not recovered in Active Inference agents under the same analysis, suggesting that the proposed dynamics capture a distinct exploration regime
Chinese Translation
我们提出了一种由内在好奇心驱动探索的强化学习框架,专为环境非平稳且奖励稀疏、延迟、无信息或缺失的场景而设计。在该模型中,动作选择由外部奖励与一种认知动机机制共同引导,该机制使智能体偏向结构化的探索方向。核心假设是:有效的探索在中等程度的非一致性水平上涌现,而在过于僵化或过于无序的动力学条件下,性能均会下降。为验证这一思想,我们在液态状态机(Liquid State Machine, LSM)基底上实现了该框架,并在两个标准基准上进行评估:离散动作的 LunarLanderv2 和连续控制的 BipedalWalkerv3。相较于成熟的深度强化学习算法,包括近端策略优化(Proximal Policy Optimization, PPO)和内在好奇心模块(Intrinsic Curiosity Module, ICM),所提出的方法在两个任务上均取得了具有竞争力的性能。我们进一步表明,在相同分析下,主动推理(Active Inference)智能体未能恢复出该好奇心窗口,这表明所提出的动力学刻画了一种独特的探索机制。
cs.LG / 7 / 2609.05658

Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations

LLM生成的SystemVerilog断言对语义保持性RTL变换的鲁棒性
Aditi, FNU
Abstract
Large language models (LLMs) are increasingly being explored for automating SystemVerilog Assertion (SVA) generation, yet most evaluations report correctness on a single syntactic representation of an input. Such point accuracy does not reveal whether a model's correct output is stable when the same RTL behavior is written differently. This paper presents a controlled metamorphic evaluation of LLM-based SVA generation under semantics-preserving RTL transformations. Starting from the VERT dataset, we construct a quality-filtered conditional-control pool and a stratified 40-program evaluation set containing 295 assignment behaviors. We evaluate two open code models, Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, with an identical evaluation prompt and greedy decoding. Three transformations are studied: operand reordering, deterministic identifier renaming, and redundant parenthesization. Beyond baseline and transformed accuracy, we measure conditional robustness, invariance failure, and any-flip rate, with 10,000-sample clustered bootstrap intervals at the RTL-program level. Across all six model-transformation conditions, 9.7%-27.0% of behaviors that were correct on the original RTL become incorrect after a semantics-preserving transformation. Aggregate accuracy can therefore hide substantial instability: under identifier renaming, DeepSeek-Coder-V2-Lite improves from 53.9% to 63.7% accuracy while 19.5% of its originally correct behaviors fail. Manual review of 30 sampled correct-to-wrong transitions identifies dropped path predicates, branch-polarity errors, Boolean-structure corruption, and output-contract violations. The results show that point accuracy alone is insufficient for characterizing LLM reliability in assertion generation and motivate robustness-aware evaluation for AI-assisted hardware verification.
Chinese Translation
大语言模型(LLM)正被越来越多地用于自动化SystemVerilog断言(SVA)的生成,然而大多数评估仅报告模型在输入的单一语法表示上的正确性。这种单点准确率无法揭示当相同的RTL行为以不同方式编写时,模型的正确输出是否稳定。本文针对基于LLM的SVA生成,在语义保持性RTL变换下开展了一种受控的蜕变评估。我们从VERT数据集出发,构建了一个经过质量筛选的条件控制池,以及一个包含295个赋值行为的分层40程序评估集。我们在相同的评估提示和贪心解码设置下评估了两个开源代码模型:Qwen2.5-Coder-7B和DeepSeek-Coder-V2-Lite。研究了三种变换:操作数重排、确定性标识符重命名以及冗余括号化。除基线准确率和变换后准确率之外,我们还测量了条件鲁棒性、不变性失败率和任一翻转率,并在RTL程序层面采用10,000次采样的聚类自助法(clustered bootstrap)构建置信区间。在全部六种模型-变换条件组合中,原始RTL上正确的行为中有9.7%-27.0%在经过语义保持性变换后变为错误。因此,总体准确率可能掩盖显著的不稳定性:在标识符重命名条件下,DeepSeek-Coder-V2-Lite的准确率从53.9%提升至63.7%,但其原本正确的行为中有19.5%出现失败。对30个抽样得到的"由对变错"案例的人工审查识别出了路径谓词丢失、分支极性错误、布尔结构损坏以及输出契约违例等问题。结果表明,仅凭单点准确率不足以刻画LLM在断言生成中的可靠性,并为AI辅助硬件验证中关注鲁棒性的评估方法提供了依据。
cs.LG / 8 / 2609.05676

PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement

PAC-Private 自回归生成:基于集成分歧的噪声校准
Mirzadehsarcheshmeh, Mina, Khandani, Amir Keyvan
Abstract
Language models adapted on private text are often served through APIs, so privacy leakage occurs through generated outputs rather than exposed weights. Private prediction protects these releases. Methods such as PMixED incur privacy cost at each release and increasingly rely on the public model over long horizons. PAC privacy instead calibrates noise to output variability across possible secrets, adding less noise when predictions are stable. To our knowledge, PAC-private prediction has not previously been extended from classification to autoregressive generation. We construct $m=128$ overlapping worlds from the private corpus, with each record appearing in exactly $m/2$ worlds, and train one adapter per world over a frozen public model. The realized world is the secret. At each token, the public model defines a candidate set, the worlds vote, and their posterior-weighted disagreement determines the PAC noise; unanimity requires no calibration noise. We prove $I(S;Y_{1:T}) \leq I(S;H_T) \leq bT$. Our contributions are extending PAC privacy to autoregressive generation, handling adaptive self-generated contexts, and introducing coupled decoding that preserves privacy accounting while avoiding greedy degeneration. On WikiText-103 with GPT-2-small, we retain 74% of the fine-tuning gain at a per-token budget of $2^{-32}$, while membership-inference success is bounded by 51.08% after $10^6$ tokens; posterior-entropy estimates of leakage are roughly 17% of the charged budget. Inference privacy is not content protection: even when membership advantage on a memorized canary is indistinguishable from zero, the canary is emitted at the same rate. Against PMixED under matched membership-inference bounds on the same data universe and test set, we retain 98% of non-private headroom from $10^2$ to $10^6$ tokens, versus at most 56%, with no crossover.
Chinese Translation
在私有文本上适配的语言模型通常通过 API 提供服务,因此隐私泄露发生在生成的输出而非暴露的权重上。私有预测(private prediction)用于保护这些发布内容。诸如 PMixED 之类的方法在每次发布时都会产生隐私成本,并且在长时程上越来越依赖公共模型。PAC 隐私(PAC privacy)则根据输出在可能秘密之间的变异性来校准噪声:当预测稳定时添加更少的噪声。据我们所知,PAC 私有预测此前尚未从分类任务扩展到自回归生成。我们从私有语料库构建 $m=128$ 个重叠的“世界”,每条记录恰好出现在 $m/2$ 个世界中,并在冻结的公共模型之上为每个世界训练一个适配器(adapter)。被实现的世界即为秘密。在每个 token 处,公共模型定义候选集,各世界进行投票,其后验加权分歧决定 PAC 噪声的大小;当投票一致时无需校准噪声。我们证明了 $I(S;Y_{1:T}) \leq I(S;H_T) \leq bT$。我们的贡献包括:将 PAC 隐私扩展到自回归生成、处理自适应的自生成上下文,以及引入在保持隐私核算的同时避免贪婪退化的耦合解码(coupled decoding)。在 WikiText-103 上使用 GPT-2-small,在每 token 预算为 $2^{-32}$ 的条件下,我们保留了 74% 的微调增益,同时在 $10^6$ 个 token 之后成员推断攻击成功率被限制在 51.08%;泄漏的后验熵估计约为所收取预算的 17%。推断隐私并不等于内容保护:即使在记忆金丝雀(canary)上的成员推断优势与零无法区分时,该金丝雀仍以相同的速率被输出。在同一数据宇宙和测试集上,在成员推断界限匹配的条件下与 PMixED 相比,从 $10^2$ 到 $10^6$ 个 token,我们保留了 98% 的非私有性能空间,而 PMixED 至多保留 56%,且不存在交叉点。
cs.LG / 9 / 2609.05688

Connecting Score Matching, Maximum Likelihood, and Expectation-Maximization in Mixed Linear Regression

在混合线性回归中连接得分匹配、极大似然与期望最大化算法
Luo, Zhankun, Hashemi, Abolfazl
Abstract
We study variance-preserving diffusion of the response in mixed linear regression (MLR) with unknown mixing weights. Our analysis separates the statistical guarantees of score matching from the loss geometry and optimization signal at a fixed diffusion noise level. The KL divergence links the denoising score matching objective integrated over the diffusion path with the likelihood and a terminal discrepancy. Under mild regularity conditions and terminal schedule, the resulting estimator converges up to the ground truth parameters of MLR, and its scaled error converges to the Gaussian limit of the maximum-likelihood estimator. At a fixed scale of the diffusion noise level, we derive a decomposition linking the score matching loss to cross-entropy and Expectation-Maximization (EM) operators. This decomposition yields an EM-related low-noise gradient expansion with additional correction terms of latent variance. In the high-noise limit, we further characterize gradient descent on this limiting loss under isotropic covariance. Along fixed high signal-to-noise ratio rays, the score matching imbalance gradient and the latent-variance term tend to zero pointwise. Numerical experiments illustrate our theoretical findings and statistical guarantees.
Chinese Translation
我们研究了在混合权重未知的混合线性回归(MLR)中对响应变量进行保持方差的扩散过程。我们的分析将得分匹配的统计保证与固定扩散噪声水平下的损失几何结构和优化信号分离开来。KL 散度将沿扩散路径积分的降噪得分匹配目标与似然函数以及终端差异联系起来。在温和的正则性条件和终端调度下,所得估计量收敛到混合线性回归的真实参数,且其缩放误差收敛到极大似然估计量的高斯极限。在固定的扩散噪声水平尺度下,我们推导出一个将得分匹配损失与交叉熵和期望最大化(EM)算子联系起来的分解。该分解产生了一个与 EM 相关的低噪声梯度展开,其中包含潜在方差的额外修正项。在高噪声极限下,我们进一步刻画了各向同性协方差下该极限损失上的梯度下降行为。沿固定的高信噪比射线方向,得分匹配不平衡梯度和潜在方差项逐点趋于零。数值实验验证了我们的理论发现和统计保证。
cs.LG / 10 / 2609.05694

GraphNOSE: A Graph Transformer in Olfaction

GraphNOSE:一种应用于嗅觉领域的图Transformer模型
Sharma, Mrityunjay, Balaji, Sarabeshwar, Parma, Valentina, Kumar, Ritesh
Abstract
Predicting olfactory qualities from molecular structure is an open problem in chemoinformatics. Although linear models can link molecular features to odor descriptors, they often fail when extrapolating to novel chemical scaffolds, extreme molecular weights, or complex odor mixtures. To address this, we introduce GraphNOSE, an open-source graph transformer framework that predicts multi-label odor descriptors from simplified molecular-input line-entry system (SMILES) strings for single molecules and binary mixtures. By integrating positional and structural encodings within a transformer-based graph architecture, GraphNOSE achieves strong performance with six times fewer parameters than standard graph neural network (GNN) baseline while consistently outperforming linear models, molecular language model embeddings, molecular fingerprints, and baseline GNNs by an average area under the ROC curve (AUROC) margin of 4.52% (p < 0.01). GraphNOSE achieves an AUROC of 84% on out-of-distribution compounds (OODs). This exceeds the current state-of-the-art GNN for OOD in olfaction (Open-POM: 81%, p < 0.001), and identifies conditions under which linear models empirically fail. Finally, we apply XAI (explainable AI) methods to identify which substructures and molecular features drive odor predictions, yielding insights consistent with chemical intuition and grounded in the model's learned representations. Together, these results establish GraphNOSE as a scalable and interpretable architecture for olfactory prediction that generalizes to structurally distinct compounds underrepresented in current perceptual databases.
Chinese Translation
从分子结构预测嗅觉特性是化学信息学中的一个开放性问题。尽管线性模型可以将分子特征与气味描述符相关联,但在外推至新颖化学骨架、极端分子量或复杂气味混合物时,它们往往表现不佳。为解决这一问题,我们提出了GraphNOSE,一个开源的图Transformer框架,可从简化分子线性输入规范(SMILES)字符串中预测单分子和二元混合物的多标签气味描述符。通过在基于Transformer的图架构中集成位置编码和结构编码,GraphNOSE以比标准图神经网络(GNN)基线少六倍的参数实现了强大性能,并以平均4.52%的ROC曲线下面积(AUROC)优势(p < 0.01)持续超越线性模型、分子语言模型嵌入、分子指纹以及基线GNN。GraphNOSE在分布外化合物(OOD)上取得了84%的AUROC,超过了当前嗅觉领域OOD任务的最先进GNN(Open-POM:81%,p < 0.001),并识别出线性模型在经验上失效的条件。最后,我们应用XAI(可解释人工智能)方法识别驱动气味预测的子结构和分子特征,所得见解与化学直觉一致,并基于模型学习到的表示。综上所述,这些结果确立了GraphNOSE作为一个可扩展、可解释的嗅觉预测架构,能够泛化到当前感知数据库中代表性不足的结构独特化合物。
cs.LG / 11 / 2609.05698

Analysis of Respiratory Sinus Arrhythmia with Neural Networks

基于神经网络的呼吸性窦性心律不齐分析
Szymanski, Julian, Orkisz, Patryk, Mora, Higinio
Abstract
The paper introduces a neural network-based approach for analyzing ECG signals to estimate respiratory rate by leveraging the phe- nomenon of Respiratory Sinus Arrhythmia (RSA). Our method employs a deep learning model trained to predict respiratory waveforms directly from ECG input data. To achieve this, we developed and evaluated three different neural network architectures capable of automatically extract- ing relevant features from ECG signals without the need for manual preprocessing. The proposed approach offers a robust and scalable solu- tion for non-invasive respiratory monitoring, with potential applications in healthcare and wearable technology
Chinese Translation
本文介绍了一种基于神经网络的ECG(心电图)信号分析方法,通过利用呼吸性窦性心律不齐(Respiratory Sinus Arrhythmia, RSA)现象来估计呼吸频率。我们的方法采用深度学习模型,经训练可直接从ECG输入数据预测呼吸波形。为此,我们开发并评估了三种不同的神经网络架构,它们能够自动从ECG信号中提取相关特征,无需人工预处理。所提出的方法为无创呼吸监测提供了一种稳健且可扩展的解决方案,在医疗保健和可穿戴技术领域具有潜在的应用价值。
cs.LG / 12 / 2609.05727

Newton Matching for Generative Modeling: A Unified Framework for Fine-Tuning and Sampling

生成式建模的牛顿匹配:微调与采样的统一框架
Li, Zeyang, Wang, Yunan, Giaretta, Paolo, Azizan, Navid
Abstract
We develop Newton Matching, a unified framework for fine-tuning and sampling in generative modeling. The target is $\pi\propto\mu e^{\tau r}$, where $r$ is the reward, $\tau>0$ the inverse temperature, and $\mu$ denotes the pretrained model's terminal density for fine-tuning or the constant $1$ for sampling. We shift the paradigm from isolated losses to iterative optimization over canonical models: population minimizers of standard conditional matching for terminal densities. Under compatible smooth-realization assumptions, canonical velocities form a manifold diffeomorphic to the density manifold. Transporting the Fisher-Rao metric and mixture connection to this manifold, we show that the reverse-KL Hessian equals the metric, so the Newton direction coincides with the negative Fisher-Rao gradient. At terminal density $\rho$, each stage takes a tangential step generated by the regularized reward $r-\frac1\tau\log(\rho/\mu)$, followed by terminal-density-preserving canonicalization. This canonical retraction yields an exact finite-stepsize density characterization. For the ideal iteration, we prove strict reverse-KL descent away from the target for $0 < \eta \le \tau$, global convergence under mild conditions, and local quadratic convergence for full steps ($\eta=\tau$). Covariance and gradient forms, each with forward or reverse regression-pair constructions, yield sample-wise tangential-update losses with the same population minimizer, without importance sampling or full-trajectory backpropagation. We develop approximate updates and define critical-point consistency as vanishing tangential displacement if and only if $\rho=\pi$. We recover representative methods as exact realizations, critical-point-consistent approximations, or objective-altering variants, enabling modular algorithm design. Our work advances the theory and algorithms of reinforcement learning for generative models.
Chinese Translation
我们提出了牛顿匹配(Newton Matching),这是一个用于生成式建模中微调与采样的统一框架。目标分布为 $\pi\propto\mu e^{\tau r}$,其中 $r$ 是奖励,$\tau>0$ 为逆温度,$\mu$ 在微调时表示预训练模型的终端密度,在采样时为常数 $1$。我们将范式从孤立的损失函数转向对规范模型(canonical models)的迭代优化——规范模型是标准条件匹配对终端密度的总体最小化器。在兼容的光滑实现假设下,规范速度场构成一个与密度流形微分同胚的流形。将 Fisher-Rao 度量与混合联络迁移到该流形上,我们证明反向 KL 散度的 Hessian 等于该度量,因此牛顿方向与负 Fisher-Rao 梯度方向一致。在终端密度 $\rho$ 处,每个阶段先沿由正则化奖励 $r-\frac1\tau\log(\rho/\mu)$ 生成的切向方向步进,随后进行保持终端密度的规范化。这种规范收缩(canonical retraction)给出了精确的有限步长密度刻画。对于理想迭代,我们证明了当 $0 < \eta \le \tau$ 时反向 KL 散度在远离目标处严格下降、在温和条件下具有全局收敛性,以及完整步长($\eta=\tau$)下的局部二次收敛。协方差形式与梯度形式各自均可通过前向或反向回归对构造实现,从而得到具有相同总体最小化器的逐样本切向更新损失,且无需重要性采样或全轨迹反向传播。我们进一步发展了近似更新方法,并将临界点一致性定义为:当且仅当 $\rho=\pi$ 时切向位移消失。我们证明了若干代表性方法可被视为精确实现、临界点一致的近似或改变目标函数的变体,从而支持模块化的算法设计。我们的工作推进了生成式模型强化学习的理论与算法。
cs.LG / 13 / 2609.05748

A Multi-Source Ensemble Approach to Candidate Generation for Alternative Vacation Rental Property Recommendations

一种用于替代度假租赁房源推荐的候选生成的多源集成方法
Zaidi, Syed Mohammed Arshad, Rincon, Eric, Hassantabar, Shayan
Abstract
Alternative property recommendations play a critical role in vacation rental marketplaces, helping users discover relevant options when viewing a specific listing. However, generating high-quality candidate alternatives presents unique challenges: heterogeneous inventory, geographic constraints, rapid availability changes, and long-tail property distributions. We present a comprehensive study of candidate generation (CG) approaches for vacation rental alternatives, comparing collaborative filtering, shallow embeddings, and graph neural network (GNN) methods. Our experiments on a large-scale vacation rental platform (over 2M active properties) show that a hybrid architecture combining item-based collaborative filtering with GNN-based retrieval improves Recall@300 by 14.8% over the strongest baseline, by leveraging the complementary strengths of the two sources: collaborative filtering excels at early recall for properties with rich interaction history, while GNNs discover diverse, non-obvious alternatives and handle cold-start scenarios more effectively. As a component result, GNN-based embeddings alone substantially outperform shallow Hotel2Vec embeddings (48-68% relative recall improvement across K), motivating their inclusion in the ensemble. Crucially, we examine how CG-stage gains carry through to the downstream ranking stage, and find that a stronger candidate pool yields higher downstream ranking quality, though attributing this effect cleanly is complicated by the coupling between candidate generation and ranker training. This recall-conversion gap is an important consideration for practitioners deploying new retrieval methods in two-stage recommendation systems.
Chinese Translation
替代房源推荐在度假租赁市场中发挥着关键作用,帮助用户在浏览某一具体房源时发现相关选项。然而,生成高质量的候选替代房源面临独特挑战:库存的异构性、地理约束、可用性的快速变化以及房源分布的长尾特性。我们对度假租赁替代房源的候选生成(Candidate Generation, CG)方法进行了全面研究,比较了协同过滤、浅层嵌入和图神经网络(GNN)方法。在一个大型度假租赁平台(超过200万活跃房源)上的实验表明,将基于物品的协同过滤与基于GNN的检索相结合的混合架构,相比最强基线将Recall@300提升了14.8%,这得益于两个来源的互补优势:协同过滤在交互历史丰富的房源上具有出色的早期召回能力,而GNN能够发现多样化、非显而易见的替代房源,并更有效地处理冷启动场景。作为一个中间结果,仅使用GNN嵌入就显著优于浅层的Hotel2Vec嵌入(在不同K值下相对召回率提升48-68%),这促使我们将其纳入集成方案。关键的是,我们检验了CG阶段的收益如何传递到下游排序阶段,发现更强的候选池能带来更高的下游排序质量,但由于候选生成与排序器训练之间的耦合,对该效应进行清晰的归因变得复杂。这一召回-转化差距是从业者在两阶段推荐系统中部署新检索方法时需要重点考虑的问题。
cs.LG / 14 / 2609.05766

Data Scout: Targeted Web Crawling for Domain-Specific Pretraining Corpora

Data Scout:面向特定领域预训练语料库的定向网络爬取
Garg, Chirag, Zahid, Eelaaf, Ahmed, Farhan, Gala, Jay Pankaj, Butler, Eric, Ludwig, Heiko
Abstract
The dominant approach to building domain-specific pretraining corpora is to filter large web archives such as CommonCrawl. This works well for popular domains but breaks down for specialized ones, where relevant content is sparse and often beyond the reach of popularity-driven crawlers. We present Data Scout, which inverts this: instead of filtering an archive, it directs a targeted crawl. An LLM expands a root topic into a taxonomy and thousands of search queries; the returned URLs (seeds) are grouped by subdomain and screened with a user-supplied classifier (the probe), admitting each subdomain on the basis of a small sample. This works because relevance has a sharp boundary at the subdomain level: in mathematics, a page is 21x more likely to be relevant than one on a sibling subdomain. With the FineMath classifier as the probe, 21.9% of crawled pages are high-quality math content, 70x the 0.31% rate from filtering a comparable web sample, so the crawl wastes far less effort. But the payoff is not just efficiency: 63.2% of these pages are missing from CommonCrawl altogether, yet just as useful for training. Continued pretraining of Llama-3.2-3B on 1.9B Data Scout tokens matches FineMath corpus on GSM8k. Because the probe is the only domain-specific component, Data Scout can in principle apply to any domain with such a classifier.
Chinese Translation
构建特定领域预训练语料库的主流方法是对CommonCrawl等大型网页存档进行过滤。这种方法对热门领域效果良好,但在专业领域则失效,因为相关内容稀疏,且往往超出了基于流行度的爬虫的触及范围。我们提出了Data Scout,它反转了这一思路:不是对存档进行过滤,而是直接进行定向爬取。一个大型语言模型(LLM)将根主题扩展为分类体系和数千条搜索查询;返回的URL(种子)按子域名分组,并用用户提供的分类器(探针)进行筛选,基于小样本决定是否纳入各子域名。这种方法之所以有效,是因为相关性在子域名层面存在明显边界:在数学领域,一个页面的相关性是其同级子域名页面的21倍。以FineMath分类器作为探针,爬取的页面中有21.9%是高质量数学内容,是从可比网络样本中过滤所得0.31%比例的70倍,因此爬取浪费的精力要少得多。而且收益不仅是效率:这些页面中有63.2%完全缺失于CommonCrawl,但对训练同样有用。在19亿Data Scout词元上对Llama-3.2-3B进行持续预训练,在GSM8k上的表现与FineMath语料库相当。由于探针是唯一与领域相关的组件,Data Scout原则上可应用于任何拥有此类分类器的领域。
cs.LG / 15 / 2609.05770

RAPTOR: Role-Aware Private Training for Mixture-of-Experts

RAPTOR:面向混合专家模型的角色感知隐私训练框架
Dm, Duc, Le-Duc, Khai, Do, Nguyen, Hoang, Minh Son, Draye, Florent, Hoang, Thai, Dam, Hoang Phuong, Liu, Jiarui, Ngo, Chris, Zhang, Terry Jingchen, Tran, Anh Le Duc, Minh, Nhat Do, Le, Minh Ngoc, Thai, My T., Xu, Ran, Savarese, Silvio, Diab, Mona, Schölkopf, Bernhard, Jin, Zhijing, Nguyen, Huy L., Kim, Daeyoung
Abstract
Differentially private (DP) fine-tuning methods treat sparse Mixture-of-Experts (MoE) models as a single dense block, ignoring that shared layers see all data while experts only see routed records. We identify and formally characterize three resulting failure modes: global clipping suppresses expert gradients, batch-level normalization dilutes sparse expert updates, and fixed privacy noise degrades signal-to-noise ratio on low-load experts. We introduce RAPTOR - a Role-Aware Private Training framework, which alternates shared and expert optimization and targets each failure directly, using expert-specific clipping and noise together with a public expected-owner denominator and a count-independent update schedule that avoids conditioning on private, realized expert counts. We prove the resulting mechanism satisfies $(\varepsilon,\delta)$-DP: because each record is assigned to exactly one owner expert, per-expert mechanisms within a layer compose in parallel, so updating all $E$ experts costs no more, in privacy terms, than updating one, with shared and expert streams composing sequentially across training. We further derive a bias-variance decomposition of the public-denominator estimator showing its bias grows predictably with routing imbalance, yielding a privacy-free rule for selecting which layer to protect from routing entropy measured on a small public corpus. Experiments on Switch Transformer and OLMoE fine-tuning across GLUE tasks, and on DeepSeek-VL2-Tiny, show consistent gains over standard DP baselines across several privacy levels ($\varepsilon$), with the largest margins typically at the tightest budgets. Code and models are publicly available: https://github.com/leduckhai/RAPTOR
Chinese Translation
现有的差分隐私(DP)微调方法将稀疏的混合专家(Mixture-of-Experts, MoE)模型视为单个稠密模块,忽略了共享层会接触所有数据而专家仅处理被路由到的记录这一事实。我们识别并形式化刻画了由此产生的三种失效模式:全局裁剪会抑制专家梯度,批级别归一化会稀释稀疏的专家更新,而固定的隐私噪声会降低低负载专家的信噪比。我们提出了RAPTOR——一个角色感知的隐私训练框架,它交替进行共享层与专家层的优化,并针对性地解决上述每种失效模式:采用针对各专家的裁剪和噪声,结合公开的期望拥有者分母以及与计数无关的更新调度,从而避免依赖于私有的、实际实现的专家计数。我们证明所提出的机制满足$(\varepsilon,\delta)$-差分隐私:由于每条记录恰好被分配给唯一一个拥有者专家,同一层内各专家的机制以并行方式组合,因此在隐私代价上更新全部$E$个专家与更新一个专家的成本相同,而共享流与专家流则在整个训练过程中以顺序方式组合。我们进一步推导了公开分母估计器的偏差-方差分解,表明其偏差随路由不均衡程度可预测地增长,并由此得出一个无需隐私开销的规则,即基于在小规模公开语料上测得的路由熵来选择需要保护的层。在Switch Transformer和OLMoE于GLUE任务的微调实验,以及在DeepSeek-VL2-Tiny上的实验表明,在多个隐私级别($\varepsilon$)下,本方法均一致优于标准差分隐私基线,且优势通常在最严格的隐私预算下最为显著。代码与模型已公开发布:https://github.com/leduckhai/RAPTOR
cs.LG / 16 / 2609.05778

Nonlinear elliptic homogenization with the parametric Deep Ritz method

基于参数化Deep Ritz方法的非线性椭圆均匀化
Rowan, Conor
Abstract
Elliptic homogenization is used to determine coarse-grained properties of materials with features on small scales. When these small scale features have rapid, periodic fluctuations, the solution field corresponding to a homogenized constitutive relation closely resembles the true solution based on the heterogeneous material. This homogenized behavior of the material is computed from a cell problem, where a cell is defined to be one period of the fluctuating material. In the context of linear elliptic partial differential equations, the homogenized constitutive relation is defined simply by a constant coefficient tensor, but for nonlinear problems, the homogenized response depends on the macroscopic state and/or its gradient, thus requiring solutions to parametric cell problems. When computing a numerical solution with the homogenized constitutive relation, it is useful to have a differentiable representation of the solution to the cell problem, as derivatives of the homogenized constitutive relation are required in Newton iterations for the macroscopic state field. In this work, we use the Deep Ritz method to solve the parametric cell problems that arise from nonlinear homogenization. First, we exploit the variational structure of the cell problem, then we discretize the dependence of the cell response on both space and the macroscopic state with a neural network. Enforcing boundary conditions on the cell response strongly, we next use the parametric Deep Ritz method to simultaneously solve the cell problem over a range of macroscopic states. We show that this method is accurate, efficient, and offers a continuous and differentiable representation of the cell response over the macroscopic state and gradient. We then show that our parametric representation of the cell response significantly expedites macroscale solutions when compared to a traditional $\text{FE}^2$ scheme.
Chinese Translation
椭圆均匀化用于确定具有小尺度特征的材料的粗粒化属性。当这些小尺度特征具有快速的周期性波动时,与均匀化本构关系相对应的解场与基于非均匀材料的真实解非常接近。材料的这种均匀化行为可通过单胞问题计算得到,其中单胞定义为波动材料的一个周期。在线性椭圆偏微分方程的情形下,均匀化本构关系可简单地由一个常系数张量定义;但对于非线性问题,均匀化响应依赖于宏观状态和/或其梯度,因此需要求解参数化的单胞问题。在使用均匀化本构关系计算数值解时,若能得到单胞问题解的可微表示将非常有用,因为在对宏观状态场进行牛顿迭代时需要均匀化本构关系的导数。在本工作中,我们使用Deep Ritz方法求解非线性均匀化产生的参数化单胞问题。首先,我们利用单胞问题的变分结构,然后用神经网络对单胞响应在空间和宏观状态上的依赖关系进行离散化。通过强形式施加单胞响应的边界条件,我们接着使用参数化Deep Ritz方法在一系列宏观状态上同时求解单胞问题。我们证明了该方法准确、高效,并且能够给出单胞响应在宏观状态和梯度上的连续可微表示。随后我们表明,与传统的$\text{FE}^2$方法相比,我们的参数化单胞响应表示显著加快了宏观尺度的求解速度。
cs.LG / 17 / 2609.05820

Online Learning with LLM Experts from Limited Feedback

有限反馈下基于大语言模型专家的在线学习
Wei, Wang, Pal, Soumyabrata, Mukherjee, Koyel, Dernoncourt, Franck, Rossi, Ryan A., Kveton, Branislav, Eldardiry, Hoda
Abstract
We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online setting with limited feedback. We formulate it as a bandit problem with $K$ actions that represent experts and $d$ features that encode prompts, over a horizon of $T$ rounds. We propose algorithms that strategically select and observe rewards to minimize regret. In the full-information setting, we achieve a regret of $\tilde{O}(d T / \sqrt{m})$, while in the bandit setting we achieve $\tilde{O}(d T \sqrt{K / m})$, where $m \ll T$ is a budget on feedback. Our experiments show that we efficiently learn high-quality routing strategies across diverse LLMs from limited feedback.
Chinese Translation
我们研究了在有限反馈的在线环境中,将提示(prompt)自适应地路由到大语言模型(LLM)专家以最大化响应质量的问题。我们将其形式化为一个老虎机(bandit)问题,其中包含 $K$ 个代表专家的动作和 $d$ 个编码提示的特征,时间范围为 $T$ 轮。我们提出了通过策略性地选择动作并观测奖励来最小化遗憾(regret)的算法。在全信息设定下,我们实现了 $ ilde{O}(d T / \sqrt{m})$ 的遗憾界;在老虎机设定下,我们实现了 $ ilde{O}(d T \sqrt{K / m})$ 的遗憾界,其中 $m \ll T$ 是反馈预算。实验表明,我们的方法能够在有限反馈下高效地学习跨多种大语言模型的高质量路由策略。
cs.LG / 18 / 2609.05822

Generalizing HVAC Control With Domain Randomized Reinforcement Learning

基于域随机化强化学习的暖通空调(HVAC)控制泛化方法
Boitel, Pablo, Zhang, Kun
Abstract
Deploying advanced HVAC (Heating, Ventilation and Air Conditioning) controllers at scale remains difficult because performance often depends on accurate building models or per-site retuning. We propose NOMAD-RL (Neural Online Meta-Adaptation for Dynamics), a general-purpose Reinforcement Learning (RL) controller designed to transfer across heterogeneous thermal zones through a universal, non-invasive thermostat interface. The controller acts on temperature setpoints from zone measurements and forecasts, while a recurrent policy supports online adaptation under partial observability. Our main contribution is an adaptive domain randomization scheme based on physics-informed normalizing flows, which models correlated and multimodal distributions of thermal-zone parameters while maintaining physical plausibility and controllability. This produces a realistic and progressively adaptive training curriculum that improves transfer across buildings. We evaluate NOMAD-RL against a constant-setpoint PID controller, RL without domain randomization, and MPC in single- and multi-zone settings. NOMAD-RL consistently outperforms the PID and non-randomized RL baselines, and approaches the performance of a well-tuned MPC, especially in the more challenging multi-zone case. These results highlight the potential of adaptive, physics-informed domain randomization for robust and transferable HVAC control.
Chinese Translation
大规模部署先进的暖通空调(HVAC)控制器仍然困难,因为其性能往往依赖于精确的建筑模型或针对每个站点的重新调参。我们提出了NOMAD-RL(Neural Online Meta-Adaptation for Dynamics,基于动力学的神经在线元自适应方法),这是一种通用的强化学习(RL)控制器,旨在通过通用、非侵入式的恒温器接口在异构热区之间实现迁移。该控制器根据区域温度测量值和预测值对温度设定点进行操作,同时循环策略支持在部分可观测条件下的在线自适应。我们的主要贡献是一种基于物理信息归一化流(physics-informed normalizing flows)的自适应域随机化方案,该方案在保持物理合理性和可控性的同时,对热区参数的相关性和多峰分布进行建模。由此产生了一个真实且渐进自适应的训练课程,提升了跨建筑的迁移能力。我们在单区域和多区域设置下,将NOMAD-RL与恒定设定点PID控制器、无域随机化的强化学习以及模型预测控制(MPC)进行了对比评估。NOMAD-RL始终优于PID和无域随机化的RL基线,并接近调优良好的MPC的性能,尤其是在更具挑战性的多区域场景中。这些结果凸显了自适应、物理信息的域随机化在实现鲁棒且可迁移的HVAC控制方面的潜力。
cs.LG / 19 / 2609.05826

Scaling Optimal Classification Trees via Adaptive Feature and Sample Reduction

通过自适应特征与样本缩减实现最优决策树分类的可扩展求解
Tu, Jiancheng, Fan, Wenqi
Abstract
Dynamic programming for optimal classification trees becomes computationally expensive as the numbers of features and training samples increase. We develop a joint feature- and sample-space reduction framework based on STreeD. Weighted STreeD merges duplicate records created after projection onto a fixed candidate set into weighted representatives. This reduces sample-dependent computation without changing the fixed-candidate optimization problem. Adaptive STreeD repeatedly refines a bounded candidate set, retains features used by the incumbent tree, rebuilds the weighted representation, and solves the resulting reduced problems. Each certified Weighted STreeD solution is optimal for its current candidate set, while the outer feature search remains heuristic over the full feature space. Experiments on five data sets show that Weighted STreeD achieves speedups of up to 121.41 times over standard STreeD. Adaptive STreeD reduces runtime in matched comparisons at depths 2 to 4 and continues to return feasible trees at greater depths where full-feature methods are limited by time or memory. Under the same computational budget, its predictive performance remains comparable to the evaluated optimal classification tree baselines and is higher in some comparisons. These results show how joint feature- and sample-space reduction can scale dynamic-programming-based optimal-tree learning to more demanding instances.
Chinese Translation
随着特征数量和训练样本数量的增加,基于动态规划的最优分类树求解在计算上变得非常昂贵。我们基于 STreeD 开发了一个特征空间与样本空间联合缩减的框架。加权 STreeD 将投影到固定候选特征集后产生的重复记录合并为带权重的代表记录,从而在不改变固定候选集优化问题的前提下减少依赖样本量的计算。自适应 STreeD 反复细化一个有界的候选特征集,保留当前最优树所使用的特征,重建加权表示,并求解由此得到的缩减问题。每个经过验证的加权 STreeD 解对于其当前候选集是最优的,而外层的特征搜索在整个特征空间上仍是启发式的。在五个数据集上的实验表明,加权 STreeD 相比标准 STreeD 可实现最高达 121.41 倍的加速。在深度为 2 到 4 的同等条件对比中,自适应 STreeD 降低了运行时间;在更大深度下,当全特征方法受限于时间或内存时,它仍能持续返回可行的决策树。在相同计算预算下,其预测性能与所评估的最优分类树基线方法相当,并在部分对比中表现更优。这些结果表明,特征空间与样本空间的联合缩减能够将基于动态规划的最优树学习扩展到更具挑战性的问题实例。
cs.LG / 20 / 2609.05850

SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement

SAFEGuard:通过有害语义分析与流畅度度量检测基于优化的越狱攻击
Vo, Quoc Viet, Le, Trung, Ranasinghe, Damith C., Abbasnejad, Ehsan
Abstract
Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has revealed that LLMs remain vulnerable to adversarial jailbreak attacks that can bypass safety guardrails and elicit harmful responses. Many defense methods are proposed to detect jailbreaks but they are limited in their effectiveness to counter wide-range optimization-based jailbreak mechanisms that can yield highly fluency-optimized or harmful semantic obfuscated prompts. To tackle this challenge, we propose a unified detection framework SAFEGuard which incorporates a hybrid fluency measurement based on cross-layer distribution distance and perplexity, and the analysis of harmful semantics through gradient matching. Our method is grounded in a paramount observation: high fluency prompts maintain their malicious intention close to harmful prompts while harmful semantic obfuscated prompts often inject gibberish token sequences. Our evaluation demonstrates that SAFEGuard consistently outperforms state-of-the-art baselines and achieves significant improvement in accuracy across different optimization-based jailbreaks. This underscores the effectiveness of SAFEGuard against evolving jailbreak attacks.
Chinese Translation
尽管人们在大语言模型(LLMs)与人类价值观对齐以及确保安全部署方面付出了巨大努力,但最近的研究表明,LLMs 仍然容易受到对抗性越狱攻击的影响,这类攻击能够绕过安全防护机制并诱发出有害响应。目前已提出许多检测越狱攻击的防御方法,但它们在面对广泛的基于优化的越狱机制时效果有限,因为这些机制可以生成高度流畅优化的或有害语义经过混淆的提示词。为应对这一挑战,我们提出了一个统一的检测框架 SAFEGuard,该框架融合了基于跨层分布距离与困惑度的混合流畅度度量,以及通过梯度匹配进行的有害语义分析。我们的方法建立在一个关键观察之上:高流畅度的提示词始终保持其恶意意图与有害提示词相近,而有害语义被混淆的提示词往往会注入乱码式的词元序列。我们的评估表明,SAFEGuard 在不同基于优化的越狱攻击下持续优于最先进的基线方法,并在准确率上取得显著提升。这凸显了 SAFEGuard 在应对不断演变的越狱攻击方面的有效性。
cs.LG / 21 / 2609.05859

Selective Posterior Margin Regularization for Forward-Corrected Classification

面向前向修正分类的选择性后验间隔正则化
Zhang, Zexing, Li, Jichao, Lei, Tianyang, Lu, XiongYi, Kewei, Yang
Abstract
Learning with class-conditional label noise often relies on a transition model from latent clean classes to observed annotations. Forward correction embeds this transition in the likelihood, yet finite-sample networks may still memorize corrupted labels. The corrected likelihood also induces a reverse posterior over the clean classes that could explain each annotation. When its leading class differs from the annotation, the model and transition matrix provide evidence against that annotation, but the leading alternatives can remain nearly tied. We introduce Selective Posterior Margin Regularization (SPMR), which preserves the Forward objective and converts this disagreement into a graded update on the clean classifier. SPMR selects the leading reverse-posterior class, scales a detached pairwise margin by the separation between the two leading posterior classes, and assigns correspondingly little influence to diffuse conflicts. The gap factorizes into transition- adjusted pairwise separation and the posterior mass carried by the leading pair. The active margin follows the locally minimum-norm logit direction that enlarges the selected pairwise margin. Across five known-transition benchmarks, SPMR improves full-length Forward by 2.5-7.0 percentage points and remains 0.7-2.5 percentage points above Forward with Mixup and early stopping. Matched interventions support distinct gains from the posterior-space coefficient, transition-adjusted target, and pairwise action. The same design transfers to estimated transitions, human annotations, architectural changes, and stronger Forward recipes. The formulation uses latent-class evidence already available inside Forward correction without promoting every posterior conflict to a corrected label.
Chinese Translation
在类别条件标签噪声下的学习通常依赖于从潜在干净类别到观测标注的转移模型。前向修正将这一转移嵌入到似然中,然而有限样本的网络仍可能记忆受损的标签。修正后的似然还会诱导出一个关于干净类别的反向后验,该后验可以解释每条标注。当其主导类别与标注不一致时,模型和转移矩阵提供了反对该标注的证据,但主导的备选类别可能几乎并驾齐驱。我们提出选择性后验间隔正则化,它保留前向目标,并将这种不一致转化为对干净分类器的分级更新。SPMR选取主导的反向后验类别,用两个主导后验类别之间的分离度对解耦的成对间隔进行缩放,从而对弥散性冲突赋予相应较小的影响。该间隔可分解为经转移调整的成对分离度以及主导类别对所承载的后验质量。该主动间隔沿着局部最小范数的logit方向,用以扩大所选的成对间隔。在五个已知转移的基准数据集上,SPMR将完整训练的前向修正提升了2.5至7.0个百分点,并且比结合Mixup和早停的前向修正仍高出0.7至2.5个百分点。匹配干预实验表明,后验空间系数、经转移调整的目标以及成对作用各自带来了不同的增益。同一设计还可迁移至估计转移、人工标注、架构变化以及更强的前向修正方案。该公式化方法利用了前向修正内部已有的潜在类别证据,而无需将每一个后验冲突都提升为修正标签。
cs.LG / 22 / 2609.05860

Beyond Arbitrary Geometry: Topology Generalization In neural PDE Operators

超越任意几何:神经偏微分方程算子的拓扑泛化
Chen, Peiyao, Xu, Zhouyuan, Nie, Jianguo, Fan, Jiansheng, Wang, Chen
Abstract
Neural operators that accept arbitrary meshes are often treated as geometry-general, but unseen domain topology changes both the invariant and decaying subspaces of a PDE operator. We use Hodge heat flow as a controlled lens on this distinction and introduce TopoBox-3D, where tunnels and cavities vary Betti support while the exact Hodge decomposition separates the harmonic kernel from the positive spectrum. Across six architectures, models that infer topology implicitly suffer excess matched degradation in 37 of 45 model--task topology-OOD cells, yet cases that change harmonic dimension are not more strongly penalized on average. The dominant difficulty is instead spectral: the initial Rayleigh quotient is the most stable predictor of error, and spectral broadening adds information for edge and face cochains. Most strikingly, controlled probes show that explicit incidence and harmonic coordinates do not yield the best kernel-identity accuracy; nevertheless, TNO ranks first in mixed-input nonharmonic accuracy on all six tasks with nontrivial harmonic support. Together, these results establish topology as a distinct generalization axis beyond arbitrary-geometry compatibility and show that its influence extends across the Hodge spectrum rather than remaining confined to the harmonic kernel. More broadly, they suggest that global, low-frequency structural priors may help organize predictions in the faster-decaying complementary component, offering a new perspective on how neural operators may generalize across topology as well as geometry.
Chinese Translation
接受任意网格的神经算子通常被视为具有几何泛化能力,但未见过的区域拓扑会同时改变偏微分方程算子的不变子空间和衰减子空间。我们利用霍奇热流作为审视这一区别的受控工具,并提出了 TopoBox-3D 基准:其中隧道与空腔改变贝蒂数支撑,而精确的霍奇分解将调和核与正谱分离。在六种架构的实验中,隐式推断拓扑的模型在 45 个模型—任务的拓扑分布外(OOD)组合中有 37 个出现额外的匹配退化,但改变调和维度的情形在平均意义上并未受到更严重的惩罚。主要的困难其实来自谱结构:初始瑞利商是误差最稳定的预测因子,而谱展宽为边和面的余链提供了额外信息。最引人注目的是,受控探针实验表明,显式的关联矩阵与调和坐标并不能带来最佳的核一致性精度;尽管如此,在所有六个具有非平凡调和支撑的任务上,TNO 在混合输入的非调和精度上均排名第一。综上,这些结果确立了拓扑作为超越任意几何兼容性的一条独立泛化轴,并表明其影响贯穿整个霍奇谱,而不仅限于调和核。更广泛地说,这些发现暗示全局的、低频的结构先验可能有助于组织快速衰减互补分量中的预测,为理解神经算子如何在拓扑与几何两个维度上实现泛化提供了新的视角。
cs.LG / 23 / 2609.05862

Budgeted Task-Aware Acquisition of Dynamic Networks

预算约束下面向任务的动态网络信息获取
Zhou, Zihe
Abstract
Learning on dynamic graphs is difficult when changes in the underlying network are only partially observed. Acquiring current graph information incurs observation and computational costs, making complete updates impractical under limited resources. This paper focuses on budgeted task-aware acquisition on dynamic networks, where a model needs to decide which stale graph information to refresh for a downstream task. We propose Scout, a lightweight framework that learns the task value of querying each node from the maintained graph and observation history. Our evaluation covers one synthetic and four real-world dynamic networks, two downstream tasks, nine acquisition baselines, and several query budgets. Scout achieves the highest mean downstream performance in 19 of the 21 benchmark-budget settings. Task-utility supervision also outperforms structural-change supervision in 13 of the 16 real-world settings. On the same dynamic network, task-matched acquisition improves link-prediction AUC by 0.012-0.016 and node-classification accuracy by 0.064-0.09 over task-mismatched acquisition. These results show that useful graph observations depend on the downstream task and that limited observation budgets can be allocated more effectively by learning directly from downstream utility.
Chinese Translation
当下层网络的变化仅被部分观测时,动态图上的学习变得十分困难。获取当前图信息会产生观测成本和计算成本,使得在有限资源下进行完整更新并不切实际。本文聚焦于动态网络上的预算约束下面向任务的信息获取(budgeted task-aware acquisition),即模型需要决定为下游任务刷新哪些过时的图信息。我们提出了 Scout,一个轻量级框架,它从所维护的图和观测历史中学习查询每个节点的任务价值。我们的评估涵盖一个合成动态网络和四个真实世界动态网络、两个下游任务、九种信息获取基线方法以及多种查询预算设置。Scout 在 21 个基准-预算设置中的 19 个上取得了最高的平均下游性能。在 16 个真实世界设置中的 13 个上,基于任务效用的监督也优于基于结构变化的监督。在相同的动态网络上,与任务不匹配的获取策略相比,任务匹配的获取策略将链接预测 AUC 提升了 0.012-0.016,将节点分类准确率提升了 0.064-0.09。这些结果表明,有用的图观测取决于下游任务,并且通过直接从下游效用中学习,可以更有效地分配有限的观测预算。
cs.LG / 24 / 2609.05877

A budget-dependent crossover between coverage- and response-based training-set selection for machine-learned interatomic potentials

机器学习原子间势训练集选择中基于覆盖度与基于响应策略的预算依赖性交叉
Bi, Jia, Elena, Alin-Marin
Abstract
Selecting compact training sets for machine-learned interatomic potentials requires deciding whether to preserve structural diversity or target configurations on which models disagree. The better choice can depend on how much data is retained, making a comparison at one training-set size insufficient. Here we link selection criteria to prediction accuracy through a budget-resolved comparison of retrained MACE models on GAP-20 Carbon and pooled revised MD17. Structural coverage is compared with a response-guided selector that targets disagreement between a coverage-trained model and a full-data reference. This retrospective response witness tests the value of model disagreement for compressing an already labelled pool. At 5\%, coverage gives smaller absolute deviations from the full-data error than random sampling across four force endpoints in both datasets. The witness has larger deviations than coverage at 1\% and 5\%, but the ordering reverses at 20\%. At 20\%, witness-selected models also lower direct held-out force errors by 0.46--5.89\% relative to coverage, with all eight paired training-seed intervals favouring the witness. Six errors fall below the full-data reference. Mean force-error reductions are 0.164--0.167~meV~$\text{\AA}^{-1}$, with larger gains for tail and masked endpoints. Complementary analyses show that learned similarity preserves the coverage ranking, while selecting by frozen-model error gives higher error than embedding coverage. These findings establish retained-data budget as a deciding variable in atomistic training-set selection and provide a direct test of when response-guided compression improves on structural coverage.
Chinese Translation
为机器学习原子间势选择紧凑的训练集时,需要决定是保留结构多样性,还是针对模型存在分歧的构型。更优的选择可能取决于保留数据的数量,因此仅在某一训练集规模下进行比较是不充分的。本文通过在GAP-20碳数据集和合并的revised MD17数据集上对重新训练的MACE模型进行预算分辨比较,将选择标准与预测精度联系起来。我们将结构覆盖度选择器与一个响应引导选择器进行比较,后者针对覆盖度训练模型与全量数据参考模型之间的分歧。这一回顾性响应见证(response witness)方法检验了模型分歧对压缩已标记数据池的价值。在5%预算下,两个数据集中的四个力端点上,覆盖度方法相比随机采样均给出更小的相对于全量数据误差的绝对偏差。在1%和5%预算下,见证方法的偏差大于覆盖度方法,但在20%预算下排序发生逆转。在20%预算下,见证方法选择的模型相对于覆盖度方法还将直接的留出集力误差降低了0.46–5.89%,且全部八组配对训练种子区间均有利于见证方法。有六个误差低于全量数据参考。平均力误差降低量为0.164–0.167 meV·Å⁻¹,其中尾部端点与掩蔽端点的增益更大。补充分析表明,学习到的相似性度量保持了覆盖度排序,而基于冻结模型误差的选择则产生比基于嵌入覆盖度选择更高的误差。这些发现确立了保留数据预算作为原子体系训练集选择中的决定性变量,并提供了一个直接检验,用以判断响应引导压缩何时优于结构覆盖度。
cs.LG / 25 / 2609.05880

Grounded and Faithful P&ID Reasoning: Constraining Vision-Language Models with Recovered Evidence Graphs

有据可依且忠实的管道与仪表图推理:利用恢复的证据图约束视觉语言模型
Gadekar, Prathamesh, Sakhinana, Sagar Srinivas, Runkana, Venkataramana
Abstract
Piping and Instrumentation Diagrams (P&IDs) are the authoritative maps of process plants: isolation, maintenance, and HAZOP decisions depend on what connects to what. Vision-language models describe these sheets fluently, yet they often invent or miss process connections---and an invented or missed link can reverse an isolation or reachability call, so a plant decision cannot trust a fluent answer that was never checked against the linework. We instead recover an explicit graph of the drawing---its symbols, the process connections between them, and the tags that name them---and then require the model to answer only by querying that graph through seven read-only operators, so a topology claim is returned only when it cites the query results that support it. On TopoPID-VQA, a new suite of 3000 topology questions over these sheets, Graph-Grounded Harness (Ours) raises exact match accuracy from 36.7--41.3% under image-only prompting to 74.3--76.0% for Qwen3-VL-4B, Qwen3-VL-8B, and Gemma-4-E4B. It does so on an imperfect substrate: on Digitize-PID dataset the recovered graph scores F1 0.742 on exact process connections, and 0.801 once symbols and tags are pooled in. The residual errors track that gap---grounding pays off where the recovered graph is right, and perception error still breaks topology questions where it is not.
Chinese Translation
管道与仪表图(P&ID)是过程工厂的权威图纸:隔离、维护和HAZOP决策都取决于设备之间的连接关系。视觉语言模型能够流畅地描述这些图纸,但它们常常虚构或遗漏过程连接——而一个虚构或遗漏的连接可能导致错误的隔离或可达性判断,因此工厂决策无法信任一个未经图纸线条核验的流畅回答。我们转而恢复图纸的显式图结构——包括其中的符号、符号之间的过程连接以及标识它们的位号——然后要求模型仅通过七个只读操作符查询该图来作答,从而只有在引用了支持该拓扑断言的查询结果时才返回该断言。在TopoPID-VQA(一个包含3000个针对这些图纸的拓扑问题的新测试集)上,Graph-Grounded Harness(我们的方法)将Qwen3-VL-4B、Qwen3-VL-8B和Gemma-4-E4B的精确匹配准确率从纯图像提示下的36.7–41.3%提升至74.3–76.0%。该方法在不完美的底层图基础上也能奏效:在Digitize-PID数据集上,恢复的图在精确过程连接上达到F1 0.742,当符号和位号合并计算时达到0.801。剩余误差与该差距相对应——在恢复的图正确的地方,接地 grounding 带来收益;而在其不正确的地方,感知误差仍会破坏拓扑问题的解答。
cs.LG / 26 / 2609.05884

CALM: Class-wise Agreement and Label-gated Disagreement Modulation for Decentralized Federated Learning

CALM:面向去中心化联邦学习的类间一致性与标签门控分歧调制方法
Ying, Yifan, Tian, Qing
Abstract
Conventional federated learning relies on parameter averaging, which forces clients to be doubly homogeneous: all must run an identical architecture, and accuracy degrades when local data are non-IID. Decentralized federated distillation sidesteps both: each client runs its peers' model snapshots as teachers on its own local data and distills from their soft predictions, with no server, no public data, and no shared architecture. Under severe non-IID skew, however, the trustworthiness of the aggregated teacher target is a matter of degree, yet existing pipelines make hard, all-or-nothing decisions: outlier teachers are discarded by threshold, and whatever target survives is trusted in full. We propose CALM, which replaces every hard decision with a smooth trust gate at three levels: per class, teachers are weighted by agreement with the peer consensus; per sample, distillation is scaled by the teachers' divergence from that target; and a label gate scales it by how strongly the target supports the sample's true label. None of this adds communication or auxiliary data. On CIFAR-10, SVHN, OrganAMNIST, and Google Speech Commands with heterogeneous client architectures under Dirichlet label skew, CALM consistently outperforms uniform and hard-filtered distillation and matches or exceeds competing heterogeneous-FL methods.
Chinese Translation
传统联邦学习依赖于参数平均,这要求客户端具有双重同质性:所有客户端必须运行相同的网络架构,且当本地数据为非独立同分布(non-IID)时准确率会下降。去中心化联邦蒸馏则规避了这两个问题:每个客户端将其他客户端的模型快照作为教师模型,在自己的本地数据上进行蒸馏,利用教师模型的软预测结果,无需服务器、无需公共数据、也无需共享架构。然而,在严重的非独立同分布偏斜下,聚合教师目标的可信度是程度问题,而现有流程却做出硬性的、非此即彼的决策:通过阈值直接丢弃离群教师模型,而对保留下来的目标则完全信任。我们提出了CALM,它在三个层面上用平滑的信任门控取代所有硬性决策:在类别层面,根据教师模型与同伴共识的一致性对其进行加权;在样本层面,根据教师模型与该目标的分歧程度来缩放蒸馏强度;并通过标签门控,根据目标对样本真实标签的支持程度来缩放蒸馏强度。以上所有机制均不增加通信开销或辅助数据。在采用异构客户端架构并受Dirichlet标签偏斜影响的CIFAR-10、SVHN、OrganAMNIST和Google Speech Commands数据集上,CALM始终优于统一加权和硬过滤蒸馏方法,并达到或超越了竞争性异构联邦学习方法的性能。
cs.LG / 27 / 2609.05885

One Rate Is Not Enough: Adaptive Anisotropic Learning Rates for LoRA Fine-Tuning

单一学习率是不够的:面向LoRA微调的自适应各向异性学习率
Wang, Huiyi, Liu, Daijiao, Yao, Lina, Gong, Dong
Abstract
Low-rank adaptation (LoRA) has become the standard for parameter-efficient fine-tuning of large language models. Most LoRA variants follow a uniform-LR convention, applying a single global learning rate across every rank-one component of every adapter. We show that this convention overlooks substantial within-module heterogeneity, where the rank-one components of a LoRA adapter update at highly uneven rates and low-velocity modules converge to concentrated singular spectra that underutilize the nominal rank budget. To address this, we propose an adaptive anisotropic learning-rate model that assigns each rank-one component its own effective learning rate, computed online from training-time signals and mean-normalized per module to preserve the global LR budget. AnLR-LoRA instantiates this model with two signals available during AdamW optimization, namely function-space velocity and Adam SNR, as a lightweight scheme with no extra trainable parameters. Across commonsense reasoning, natural language generation and visual instruction-tuning benchmarks, AnLR-LoRA consistently improves over LoRA while encouraging broader use of rank capacity, with gains that remain robust across a wide range of global learning rates and transfer cleanly to other LoRA variants.
Chinese Translation
低秩适应(LoRA)已成为大语言模型参数高效微调的标准方法。大多数LoRA变体遵循统一学习率的惯例,即对所有适配器的每个秩一(rank-one)分量应用单一的全局学习率。我们证明这一惯例忽视了模块内部显著的非均匀性:LoRA适配器的各秩一分量以高度不均衡的速度更新,且低速度模块收敛为集中的奇异谱,从而未能充分利用名义上的秩预算。为解决这一问题,我们提出了一种自适应各向异性学习率模型,为每个秩一分量分配其自身的有效学习率。该学习率基于训练过程中的信号在线计算,并按模块进行均值归一化,以保持全局学习率预算不变。AnLR-LoRA 利用AdamW优化过程中可获得的两种信号——函数空间速度(function-space velocity)和Adam信噪比(SNR)——以轻量级方案实现了该模型,且不引入任何额外的可训练参数。在常识推理、自然语言生成和视觉指令微调等多个基准上,AnLR-LoRA 始终优于LoRA,同时促进了对秩容量的更充分利用;其性能提升在广泛的全球学习率范围内保持稳健,并能无缝迁移至其他LoRA变体。
cs.LG / 28 / 2609.05895

A First-Order Learning Algorithm for Online Resource Allocation with Constant Regret

具有常数后悔界的一阶在线资源分配学习算法
Li, Menglong, Zhang, Jiawei
Abstract
We study a finite-horizon online resource allocation problem with initial resource capacities proportional to the horizon. In each period, a request type is observed and one action is chosen from a finite menu. Each action earns a reward and consumes a vector of resources. The arrival types are independent and identically distributed, but their probabilities are unknown. We present a primal first-order learning policy that, in each period, performs one gradient ascent update of the action coordinates associated with the current request type. The policy achieves $O(1)$ expected additive regret relative to the hindsight optimum, with a bound independent of the horizon $T$. It does not solve any linear program, and the regret bound does not require a nondegeneracy assumption on the fluid linear program.
Chinese Translation
我们研究了一类有限时段的在线资源分配问题,其中初始资源容量与总时段数成比例。在每一时段,会观测到一个请求类型,并从一个有限的动作集合中选择一个动作。每个动作带来一份收益并消耗一个资源向量。请求的到达类型是独立同分布的,但其概率未知。我们提出了一种原始(primal)一阶学习策略,该策略在每个时段对与当前请求类型相关联的动作坐标执行一次梯度上升更新。该策略相对于事后最优(hindsight optimum)可实现 $O(1)$ 的期望加性后悔,且该界不依赖于总时段数 $T$。该策略无需求解任何线性规划,且其后悔界不要求流体线性规划满足非退化性假设。
cs.LG / 29 / 2609.05919

A dictionary learning framework for graphs via filters and optimal transport

一种基于滤波器与最优传输的图字典学习框架
Liao, Jinchuan, Nguyen, Dai Hai
Abstract
We propose a graph dictionary learning (GDL) framework where each graph is represented as a zero-mean Gaussian distribution derived from its filtered Laplacian. Each observed graph is approximated by a barycenter over learned atom graphs, computed under the filter graph distance (fGOT), a graph comparison metric sensitive to global structural properties. The reconstruction error between the observed graph and its barycenter is measured by the surrogate fGOT (sfGOT) distance, a tractable approximation of fGOT that handles graphs without known node correspondence, and is minimized end-to-end via backpropagation. We further provide a novel interpretation of sfGOT through the lens of the Hilbert-Schmidt Independence Criterion, showing that minimizing the sfGOT distance between two graphs is equivalent to maximizing statistical dependence between the spectral embedding of their nodes. Experiments on benchmark datasets demonstrate competitive performance over existing GDL methods on graph clustering and classification tasks.
Chinese Translation
我们提出了一种图字典学习(Graph Dictionary Learning, GDL)框架,其中每个图被表示为由其滤波拉普拉斯矩阵导出的零均值高斯分布。每个观测图通过在滤波图距离(filter graph distance, fGOT)下计算的质心来近似,该质心基于学习得到的原子图计算,其中fGOT是一种对全局结构性质敏感的图比较度量。观测图与其质心之间的重构误差由代理fGOT(surrogate fGOT, sfGOT)距离度量,sfGOT是fGOT的一种可计算的近似,能够处理节点对应关系未知的图,并通过反向传播进行端到端最小化。我们进一步从希尔伯特-施密特独立性准则(Hilbert-Schmidt Independence Criterion)的角度为sfGOT提供了一种新颖的解释,表明最小化两个图之间的sfGOT距离等价于最大化其节点谱嵌入之间的统计依赖性。在基准数据集上的实验表明,本方法在图聚类和分类任务上的性能优于现有的GDL方法。
cs.LG / 30 / 2609.05933

Rethinking the Evaluation of Efficiency Methods for Multi-Agent Systems

重新思考多智能体系统效率评估方法
Zhang, Jiamu, Zhang, Lingxi, Lu, Pengjun, Zhang, Qiyue, Chuang, Yu-Neng, Li, Zhengchen, Xu, Shuai, Chaudhary, Vipin, Chen, Hanjie
Abstract
Efficiency is increasingly important for Large Language Model (LLM)-based multi-agent systems (MAS), as larger models and more agents introduce substantial execution costs. Recent methods aim to make MAS cheaper by pruning agents, removing communication edges, or searching for compact structures. However, we argue that existing evaluations may overestimate their true ability to improve MAS efficiency. Reported gains are often measured under method-specific prompts and starting topologies, making them difficult to attribute to the proposed structural changes. Moreover, many reported successes appear in non-MAS-demanding settings, where a single agent or a randomly pruned system can already preserve strong performance. To study these issues, we introduce a controlled and MAS-demanding diagnostic benchmark for representative MAS efficiency methods. We evaluate methods under a shared backbone model, agent registry, and runtime, across controlled variations in topology, scale, depth, and tool use. Our analysis shows that many reported gains are setup-dependent and may arise from structural collapse, disabled tool pathways, or starting systems where random pruning already preserves accuracy, rather than robust improvements in MAS efficiency.
Chinese Translation
效率对于基于大语言模型(LLM)的多智能体系统(MAS)日益重要,因为更大的模型和更多的智能体会带来可观的执行成本。近期方法旨在通过剪枝智能体、删除通信边或搜索紧凑结构来降低MAS的成本。然而,我们认为现有评估可能高估了这些方法提升MAS效率的真实能力。所报告的性能提升往往是在特定于方法的提示词和起始拓扑下测得的,因此难以将其归因于所提出的结构改动。此外,许多报告的成功案例出现在对MAS要求不高的场景中——在这些场景下,单个智能体或随机剪枝的系统已能保持较强性能。为研究这些问题,我们引入了一个受控且对MAS有挑战性的诊断基准,用于评估具有代表性的MAS效率方法。我们在共享的骨干模型、智能体注册表和运行时环境下,对拓扑、规模、深度和工具使用进行受控变化来评估这些方法。我们的分析表明,许多报告的性能提升依赖于实验设置,可能源于结构坍塌、工具通路被禁用,或起始系统中随机剪枝本就能保持准确率,而非MAS效率的稳健提升。
cs.LG / 31 / 2609.05946

Interpretable and Fair Generalized Additive Neural Networks via Multi-objective Learning

基于多目标学习的可解释且公平的广义可加神经网络
Wang, Ziming, Huang, Changwu, Tang, Ke, Ong, Yew-Soon, Yao, Xin
Abstract
Interpretability and fairness are two of the most emphasized dimensions in trustworthy artificial intelligence (AI). Various explainable AI methods have been introduced to improve interpretability. This paper focuses on neural network (NN)-based generalized additive models (GAMs), a class of self-interpretable models. While most existing research has prioritized improving the accuracy of NN-based GAMs, their interpretability remains largely underexplored. To address this gap, this paper introduces explicit quantitative metrics for evaluating the interpretability of NN-based GAMs, empirically examines their effectiveness, and explores strategies for improving interpretability within these models. In addition, the simultaneous and explicit optimization of both interpretability and fairness, along with their trade-offs and the underlying reasons, remains underexplored. To address this, we propose a multi-objective neural basis model (MONBM) framework based on multi-objective evolutionary learning to consider accuracy, interpretability, and fairness simultaneously. A partial retraining strategy is further developed to facilitate the practical application of evolutionary multi-objective optimization to deep model architectures. Based on MONBM, this paper reveals the complex relationships between these dimensions and the reasons behind these intricate relationships. This analysis demonstrates how multi-objective optimization can be combined with self-interpretable models to reveal relationships among trustworthiness objectives. In addition, MONBM obtains a set of models with different trade-offs between dimensions, and the competitiveness of the approach is validated by comparing it with state-of-the-art methods.
Chinese Translation
可解释性与公平性是可信人工智能(AI)中最受重视的两个维度。已有多种可解释AI方法被提出以提升可解释性。本文聚焦于基于神经网络(NN)的广义可加模型(GAM),这是一类具备自解释性的模型。尽管现有研究大多优先提升基于NN的GAM的准确性,其可解释性仍在很大程度上未被充分探索。为填补这一空白,本文引入了用于评估基于NN的GAM可解释性的显式量化指标,通过实验验证其有效性,并探索提升此类模型可解释性的策略。此外,对可解释性与公平性进行同时且显式的优化,以及二者之间的权衡及其内在原因,也仍未得到充分研究。为此,我们提出了一种基于多目标进化学习的多目标神经基模型(MONBM)框架,以同时考虑准确性、可解释性和公平性。我们进一步开发了部分重训练策略,以促进进化多目标优化在深度模型架构中的实际应用。基于MONBM,本文揭示了这些维度之间复杂的关系及其背后的原因。该分析展示了如何将多目标优化与自解释模型相结合,以揭示可信性目标之间的关系。此外,MONBM获得了一组在各维度之间具有不同权衡的模型集合,并通过与最先进方法的比较验证了该方法的竞争力。
cs.LG / 32 / 2609.05952

QGB-W$k$NN: Quantum Granular-Ball Learning for Robust Classification

QGB-W$k$NN:基于量子颗粒球的鲁棒分类方法
Yuan, Suzhen, Chen, Dehang, Shen, Lifeng, Xia, Shuyin, Deng, Jeremiah D.
Abstract
Nearest-neighbor classification is widely used in machine learning, yet existing methods often suffer from low computational efficiency and limited robustness in noisy environments. To jointly address these challenges, this paper proposes an efficient and reliable weighted $K$-nearest neighbor classification framework based on quantum granular balls, termed QGB-W$k$NN. The proposed framework improves computational efficiency by integrating quantum-enhanced granular-ball representation with hierarchical nearest-neighbor search, while enhancing classification reliability through a purity-aware weighted decision mechanism. Specifically, quantum-kernel granular balls are constructed to reduce retrieval redundancy and strengthen nonlinear feature representation under limited quantum resources. A granular-ball purity-guided HNSW optimization strategy is developed to exploit structural reliability for hierarchical graph construction during neighbor retrieval, alleviating the local optimality issue caused by conventional random layering. Finally, a weighted voting mechanism jointly incorporating granular-ball similarity and purity is introduced to produce more reliable classification decisions in noisy environments. Extensive experiments on benchmark datasets demonstrate that QGB-W$k$NN achieves competitive classification accuracy while exhibiting favorable Pareto trade-offs between classification performance and computational cost. Moreover, the proposed framework consistently improves robustness under various noisy conditions, suggesting that reliability-aware quantum granular-ball learning provides a promising paradigm for efficient and robust nearest-neighbor classification.
Chinese Translation
最近邻分类在机器学习中被广泛应用,但现有方法往往存在计算效率低以及在噪声环境中鲁棒性有限的问题。为了共同应对这些挑战,本文提出了一种基于量子颗粒球的高效可靠的加权$K$近邻分类框架,称为QGB-W$k$NN。该框架通过将量子增强的颗粒球表示与层次化最近邻搜索相结合来提升计算效率,同时通过引入考虑纯度的加权决策机制来增强分类的可靠性。具体而言,构建量子核颗粒球以减少检索冗余,并在有限量子资源下强化非线性特征表示。提出了一种颗粒球纯度引导的HNSW优化策略,在邻居检索过程中利用结构可靠性进行层次图构建,缓解了传统随机分层导致的局部最优问题。最后,引入一种联合考虑颗粒球相似度与纯度的加权投票机制,在噪声环境中产生更可靠的分类决策。在基准数据集上的大量实验表明,QGB-W$k$NN在取得具有竞争力的分类准确率的同时,在分类性能与计算成本之间表现出良好的帕累托权衡。此外,所提出的框架在各种噪声条件下均能持续提升鲁棒性,这表明可靠性感知的量子颗粒球学习为高效鲁棒的最近邻分类提供了一种有前景的范式。
cs.LG / 33 / 2609.05955

LoGIC: Budgeted Context Construction for Node-Level Graph In-Context Learning with Tabular Foundation Models

LoGIC:基于表格基础模型的节点级图上下文学习的预算约束上下文构建
Yang, Mingqi, Guo, Zidong, Yang, Jihui, Zuo, Wenming
Abstract
Tabular foundation models have become powerful graph learners. Systems such as G2T-FM and GraphPFN encode each node as a feature row and make predictions through in-context learning (ICL), with labeled rows serving as the prompt. Current protocols employ the complete training table as context, causing attention to scale quadratically with the labeled pool and introducing preprocessing and memory bottlenecks. We investigate context construction for node-level graph ICL: which labeled nodes and auxiliary unlabeled nodes should constitute the prompt for specified queries. We formulate this allocation in terms of two resources: a labeled-context budget for predictive evidence and an unlabeled-halo budget for adapter message passing without using label capacity. We present LoGIC, which retrieves labeled nodes via structural, feature-based, and coverage channels, shares each context across the queries in a graph-local cluster, incorporates an unlabeled halo for adapter backbones, and chooses the channel and context budget without test labels. Across three backbone configurations drawn from two model families on GraphLand, budgeted contexts maintain locally runnable full-context performance, stay competitive with published large-dataset results, and markedly lower peak memory requirements compared with full-context and whole-graph inference. They further permit frozen graph ICL on million-node graphs without retraining. Our analysis identifies when retrieval channels work best and connects their behavior with graph properties.
Chinese Translation
表格基础模型(Tabular Foundation Models)已成为强大的图学习器。诸如 G2T-FM 和 GraphPFN 等系统将每个节点编码为一个特征行,并通过上下文学习(In-Context Learning, ICL)进行预测,其中带标签的行充当提示。当前的协议使用完整的训练表作为上下文,导致注意力计算随带标签数据池的规模呈二次方增长,并引入了预处理和内存瓶颈。我们研究节点级图 ICL 的上下文构建问题:对于给定的查询,哪些带标签节点和辅助的无标签节点应当构成提示。我们用两种资源来形式化这一分配问题:用于预测证据的带标签上下文预算,以及在不占用标签容量的情况下支持适配器消息传递的无标签光环预算。我们提出 LoGIC,它通过结构、特征和覆盖三种通道检索带标签节点,在图局部聚类中的查询之间共享上下文,为适配器骨干网络引入无标签光环,并在无需测试标签的情况下选择通道与上下文预算。在 GraphLand 上基于两个模型家族的三种骨干配置的实验中,预算约束的上下文保持了可本地运行的全上下文性能,与已发表的大数据集结果相比具有竞争力,并且与全上下文和整图推理相比显著降低了峰值内存需求。它们进一步使冻结的图 ICL 能够在百万节点规模的图上运行而无需重新训练。我们的分析识别了检索通道何时最为有效,并将其行为与图的性质联系起来。
cs.LG / 34 / 2609.05966

On-the-go Forgetting without Explicit Unlearning via ERASE

基于ERASE的无需显式遗忘的即时遗忘方法
Chakrabarti, Kushal, Baranwal, Mayank
Abstract
Existing unlearning approaches typically rely on post hoc weight adaptation or distillation, leading to duplicated memory costs, degraded generalization, and limited scalability. In this work, we introduce ERASE, Erasure via Reconstructive Adversarial Signal Editing, a framework for on-the-go forgetting that suppresses the observable influence of private data without modifying model weights. ERASE leverages structured, class-conditioned input perturbations to induce selective forgetting during inference, eliminating the need for retraining, fine-tuning, or model copies. We rigorously characterize sufficient conditions when ERASE provably achieves functional forgetting of designated subclasses while preserving predictions across other subclasses within the same superclass. This analysis offers a principled foundation for inference-time forgetting under mild regularity assumptions. Across diverse architectures and benchmark datasets, ERASE maintains the best observed balance between forgetting efficacy, computational efficiency, and retention fidelity over recent unlearning-based methods. By reimagining data removal as forgetting without unlearning, our work establishes a scalable, regulation-aligned pathway for continual, privacy-conscious learning.
Chinese Translation
现有的遗忘(unlearning)方法通常依赖于事后权重调整或知识蒸馏,导致内存成本翻倍、泛化能力下降以及可扩展性受限。在本工作中,我们提出了ERASE(Erasure via Reconstructive Adversarial Signal Editing,通过重构对抗信号编辑实现擦除),这是一个即时遗忘(on-the-go forgetting)框架,能够在不修改模型权重的情况下抑制私有数据的可观测影响。ERASE利用结构化的、类条件化的输入扰动在推理阶段诱导选择性遗忘,从而无需重新训练、微调或复制模型。我们严格刻画了ERASE能够可证明地实现对指定子类的功能性遗忘、同时保持同一超类中其他子类预测的充分条件。该分析在温和的正则性假设下,为推理时遗忘提供了有原则的理论基础。在多种架构和基准数据集上,与近期的遗忘方法相比,ERASE在遗忘效果、计算效率和保留保真度之间保持了最佳的观测平衡。通过将数据删除重新构想为无需显式遗忘的遗忘,我们的工作为持续的、注重隐私的学习建立了一条可扩展的、符合法规的路径。
cs.LG / 35 / 2609.05970

Machine Learning for Pre-Culture ESBL Risk Stratification to Guide Empiric Antibiotic Selection: A 12-Hospital Study of Enterobacteriaceae Cultures

机器学习用于培养前ESBL风险分层以指导经验性抗生素选择:一项涵盖12家医院的肠杆菌科细菌培养研究
Kuruvikkattil, Aravind V., Pulavarthy, Lalitha Pranathi, Kudamala, Rashmita, Purkayastha, Saptarshi
Abstract
Empiric antibiotic therapy for suspected ESBL-producing Enterobacteriaceae must be selected 48-72 hours before culture results, forcing clinicians to choose between undertreating resistant infections and overusing carbapenems that drive further resistance. We developed a cost-sensitive XGBoost model predicting an ESBL phenotype (resistance to ceftriaxone, ceftazidime, cefepime or piperacillin-tazobactam) at culture ordering using 45 pre-culture EHR features across 132,955 cultures from 72,217 patients at 12 hospitals (14.41% with the ESBL phenotype). Cultures were partitioned at the patient level. At 90% sensitivity, the model achieved 95.8% NPV, reducing post-test ESBL probability to 4.2%, a threshold that may support safe carbapenem-sparing in non-ICU settings, while sparing 307 of every 1,000 cultures an unnecessary broad-spectrum course at the cost of 14 missed ESBL cases per 1,000. SHAP analysis identified prior ESBL colonization as the dominant predictor, ahead of prior organism burden and neighborhood deprivation; removing deprivation features caused minimal performance loss ($\Delta\text{AUROC} = -0.020$), enabling equitable bedside deployment. Discrimination was unchanged under a strict IDSA ESBL-E definition (AUROC 0.766), with specimen type added as a predictor (0.764) and without any class-imbalance correction (0.762), and ranged from 0.71 to 0.78 across organism strata.
Chinese Translation
对于疑似产ESBL(超广谱β-内酰胺酶)肠杆菌科细菌感染的患者,经验性抗生素治疗必须在培养结果出来前48-72小时选定,这迫使临床医生在治疗不足(漏治耐药感染)与过度使用碳青霉烯类药物(加剧耐药性)之间做出抉择。我们开发了一个成本敏感的XGBoost模型,利用45项培养前电子健康档案(EHR)特征,在开具培养医嘱时预测ESBL表型(对头孢曲松、头孢他啶、头孢吡肟或哌拉西林-他唑巴坦耐药)。该模型基于12家医院72,217名患者的132,955份培养数据(14.41%为ESBL表型),并按患者层面进行数据划分。在90%灵敏度下,模型达到95.8%的阴性预测值(NPV),将检测后ESBL概率降至4.2%——该阈值可能支持在非ICU环境中安全地省用碳青霉烯类药物,即每1,000份培养中有307份可避免不必要的广谱抗生素疗程,代价是每1,000份中漏诊14例ESBL病例。SHAP分析显示,既往ESBL定植是最主要的预测因子,其后依次为既往病原菌负荷和社区(邻里)剥夺程度;去除剥夺特征仅造成极小的性能损失(ΔAUROC = -0.020),从而支持公平的床旁部署。在严格的IDSA ESBL-E定义下,模型判别能力保持不变(AUROC 0.766),加入标本类型作为预测因子后为0.764,不进行任何类别不平衡校正时为0.762;在不同病原菌分层中,AUROC范围为0.71至0.78。
cs.LG / 36 / 2609.05988

DART: Distributional Adversarial Recurrent Training for Algorithm Learning

DART:面向算法学习的分布对抗循环训练方法
Bao, Hieu Tran, Dang, Phung Thanh, Minh, Pham Quang Nhat, Tung, Hoang Thanh
Abstract
Recurrent reasoning models (RRMs) can solve structured problems, achieving easy-to-hard generalization through iterative computation in hidden space. These models are typically trained with instance-level supervision, which becomes increasingly problematic as task difficulty grows: valid solutions occupy a tiny region of the solution space, while invalid solutions proliferate rapidly. We propose Distributional Adversarial Recurrent Training (DART), a training framework that replaces single-point supervision with a local target distribution around the ground-truth solution and aligns model outputs with this distribution through an adversarial objective. DART provides a richer learning signal and encourages more stable iterative trajectories toward valid solutions. When evaluated on Maze, Chess, and masked Sudoku with multiple RRMs, including Deep Thinking Systems and Tiny Recursive Models, DART improves solution quality, stability, and robustness under the evaluated distribution shifts. Comparisons with label smoothing, Gaussian softened targets, and progressive training show that DART is not explained by target softening alone and is complementary to training schemes that stabilize long-horizon recurrence. These results identify DART as a promising approach for improving robustness across the evaluated recurrent reasoning models.
Chinese Translation
循环推理模型(Recurrent Reasoning Models, RRMs)能够求解结构化问题,并通过隐空间中的迭代计算实现从易到难的泛化。这些模型通常采用实例级监督进行训练,但随着任务难度的增加,这种训练方式的问题日益凸显:有效解仅占据解空间中极小的区域,而无效解则迅速激增。我们提出分布对抗循环训练(Distributional Adversarial Recurrent Training, DART),该训练框架以真实解周围的局部目标分布取代单点监督,并通过对抗性目标将模型输出与该分布对齐。DART 提供了更丰富的学习信号,并促使迭代轨迹更稳定地趋向有效解。在迷宫(Maze)、国际象棋(Chess)和掩码数独(masked Sudoku)任务上,使用多种循环推理模型(包括深度思考系统 Deep Thinking Systems 和微型递归模型 Tiny Recursive Models)进行评估,DART 在所评估的分布偏移下提升了解题质量、稳定性和鲁棒性。与标签平滑、高斯软化目标以及渐进式训练的对比实验表明,DART 的效果不能仅由目标软化解释,且与稳定长程循环的训练方案互为补充。这些结果表明,DART 是提升所评估的各类循环推理模型鲁棒性的一种有前景的方法。
cs.LG / 37 / 2609.06006

Memory in Deep Time-Series Models

深度时序模型中的记忆机制
Nguyen, Minh Hoang, Nguyen, Huu Hiep, Nguyen, Manh, Do, Van Dai, Nguyen, Dung, Le, Hung
Abstract
Deep learning for time series has progressed through successive architectural paradigms, from recurrent networks and transformers to structured state-space models, retrieval-augmented predictors, foundation models, and tool-using agents. These developments are typically studied in isolation, organized by architecture or modeling era. We argue that they can instead be viewed through a common question of \emph{how does a time-series model retain and access information beyond its immediate input?} This question is motivated by a fundamental limitation of conventional time-series modeling: information relevant to a prediction may lie far beyond a feasible input window, while compressing history into a fixed-size state can discard information that may become useful later. We formulate this challenge as a \emph{memory} problem and organize existing time-series methods along a spectrum from internal memory, encoded in parameters and fixed-size states, to external memory that is addressable, retrievable, and increasingly maintained by agents. We then develop a unified taxonomy of memory mechanisms and review three classes of external memory, including explicit modules, retrieval augmentation, and agentic stores, under a common framework for what is retained, how it is written and accessed, and how it persists. A cross-cutting analysis maps these mechanisms to time-series tasks and identifies gaps in both methods and evaluation. We conclude by outlining open problems in building memory systems that can selectively retain, retrieve, revise, and forget information as temporal environments evolve. The result is a framework for studying memory as a first-class dimension of time series modeling, independent of the underlying backbone.
Chinese Translation
时间序列深度学习经历了多个相继的架构范式的演进,从循环网络和Transformer,到结构化状态空间模型、检索增强预测器、基础模型(foundation models)以及使用工具的智能体。这些进展通常被孤立地研究,按架构或建模时代加以组织。我们认为,它们可以通过一个共同的问题来审视:\emph{时序模型如何在其直接输入之外保留和访问信息?}这一问题的动因在于传统时序建模的一个根本局限:与预测相关的信息可能远超可行输入窗口的范围,而将历史压缩到固定大小的状态中可能会丢弃日后可能变得有用的信息。我们将这一挑战表述为一个\emph{记忆}问题,并沿一条谱系组织现有的时序方法:从以参数和固定大小状态编码的内部记忆,到可寻址、可检索、且日益由智能体维护的外部记忆。随后,我们构建了记忆机制的统一分类体系,并在一个关于保留什么、如何写入和访问、以及如何持久化的共同框架下,综述了三类外部记忆:显式模块、检索增强和智能体化存储。一项贯穿性分析将这些机制映射到时序任务上,并识别了方法与评估两方面的空白。最后,我们概述了构建记忆系统的开放性问题,这些系统应能在时序环境演化时选择性地保留、检索、修改和遗忘信息。由此,我们提供了一个将记忆作为时序建模的一等维度加以研究的框架,且独立于底层的骨干架构。
cs.LG / 38 / 2609.06016

Granular-Ball Quantum Clustering for Resource-Efficient and Robust Learning

面向资源高效且鲁棒学习的粒球量子聚类
Yuan, Suzhen, Xie, Qilin, Shen, Lifeng, Xia, Shuyin, Deng, Jermiah D., Wang, Guoying
Abstract
Quantum clustering aims to exploit quantum feature representations to uncover complex data structures beyond conventional Euclidean geometry. Yet this sample-level kernel construction requires O(n^2) quantum circuit executions for n data points, creating a major bottleneck under near-term quantum resource constraints. Prior solutions fail to resolve this efficiency-accuracy dilemma: classical granular-ball clustering reduces sample complexity but relies on Euclidean metrics that cannot capture quantum correlations, while existing quantum compression schemes prioritize efficiency over structural preservation, degrading performance on non-convex or noisy data. Here we propose Granular-Ball Quantum Clustering (GBQC), a framework that tightly couples granular-ball structural abstraction with quantum feature learning. GBQC first compresses raw data into compact, representative granular balls via a PCA-guided splitting strategy, reducing kernel evaluations by 80% compared to full-sample methods. A quantum cohesion mechanism then filters noisy granules in Hilbert space to improve clustering robustness. Extensive experiments on synthetic, noisy, overlapping, and real-world datasets demonstrate that GBQC consistently achieves superior clustering accuracy and robustness compared with representative classical and quantum clustering methods. Meanwhile, the proposed granular-ball compression significantly reduces quantum kernel evaluations and computational overhead, enabling quantum clustering experiments on larger datasets within parameterized quantum learning frameworks. These results suggest that granular-ball representations serve not only as a compression mechanism to reduce quantum computational costs but also as an effective structural abstraction mechanism that improves clustering quality by eliminating redundant and structurally ambiguous learning units.
Chinese Translation
量子聚类旨在利用量子特征表示来揭示超越传统欧几里得几何的复杂数据结构。然而,这种样本级的核构建需要对n个数据点执行O(n^2)次量子线路运行,在近期量子资源受限条件下构成了主要瓶颈。已有方案均未能解决这一效率与精度之间的两难困境:经典粒球聚类虽降低了样本复杂度,但依赖无法捕捉量子关联的欧几里得度量;而现有的量子压缩方案则以牺牲结构保持为代价优先考虑效率,导致在非凸或含噪数据上的性能下降。本文提出粒球量子聚类(Granular-Ball Quantum Clustering, GBQC),这是一个将粒球结构抽象与量子特征学习紧密耦合的框架。GBQC首先通过PCA引导的分裂策略将原始数据压缩为紧凑且具有代表性的粒球,与全样本方法相比将核评估次数减少了80%。随后,量子内聚机制在希尔伯特空间中过滤噪声粒球,以提升聚类鲁棒性。在合成数据、含噪数据、重叠数据以及真实数据集上的大量实验表明,与代表性的经典和量子聚类方法相比,GBQC始终能取得更优的聚类精度和鲁棒性。同时,所提出的粒球压缩显著减少了量子核评估次数和计算开销,使得在参数化量子学习框架下可在更大规模数据集上开展量子聚类实验。这些结果表明,粒球表示不仅可以作为降低量子计算成本的压缩机制,还可以作为一种有效的结构抽象机制,通过消除冗余和结构模糊的学习单元来提升聚类质量。
cs.LG / 39 / 2609.06018

IXPLORE: Bounded Ideal Point Estimation with Grid-Based Uncertainty Quantification

IXPLORE:基于网格不确定性量化的有界理想点估计
Bachmann, Fynn
Abstract
Ideal point estimation is widely used to analyze and visualize political data. However, selecting the corresponding spatial model involves various trade-offs: while model-based approaches such as Item Response Theory (IRT) are based on utility functions rather than optimized for predictive accuracy, most Machine Learning (ML) alternatives struggle to generalize beyond training data when embedding sparse test responses. We introduce IXPLORE, a bounded ideal point estimation algorithm that combines a predictive fit objective with a sparsity-aware likelihood function. On five benchmark datasets spanning surveys, roll calls, and deliberation, this approach surpasses model-based and ML-based algorithms on reconstruction and imputation error - especially for users with sparse responses. Furthermore, we show that non-linear feature transforms can further reduce the reconstruction error while remaining visually interpretable. To quantify uncertainty, IXPLORE applies grid-based posterior inference on a bounded 2D latent space. Available as a Python package on PyPI, IXPLORE offers a flexible framework for constructing bounded, interpretable political maps with fast inference and strong imputation performance.
Chinese Translation
理想点估计被广泛用于政治数据的分析与可视化。然而,选择相应的空间模型涉及诸多权衡:基于模型的方法(如项目反应理论,Item Response Theory, IRT)建立在效用函数之上,而非针对预测准确性进行优化;而大多数机器学习(Machine Learning, ML)替代方法在嵌入稀疏的测试响应时难以泛化到训练数据之外。我们提出了IXPLORE,这是一种有界理想点估计算法,它将预测拟合目标与稀疏感知的似然函数相结合。在涵盖调查、唱名表决和商议讨论的五个基准数据集上,该方法在重构误差和插补误差方面均优于基于模型和基于机器学习的算法——尤其是对于响应稀疏的用户。此外,我们表明非线性特征变换可以在保持视觉可解释性的同时进一步降低重构误差。为量化不确定性,IXPLORE在有界的二维潜在空间上应用基于网格的后验推断。IXPLORE已作为Python包在PyPI上发布,为构建有界、可解释的政治地图提供了一个灵活的框架,具有快速推断和优异的插补性能。
cs.LG / 40 / 2609.06042

Minimizing the Effect of Sleep Deprivation in the Forward-Forward Algorithm

最小化睡眠剥夺对前向-前向算法的影响
Datta, Joy, Saha, Puja, Rabbi, Rawhatur, Rafin, Nafiz Imtiaz, Shatabda, Swakkhar, Alam, Md. Golam Rabiul, Mourning, Chad
Abstract
This paper addresses the challenge posed by sleep deprivation in the Forward-Forward algorithm, where separating the two passes in this algorithm and imbalancing the data processing in the passes is considered an imitation of the cognitive processes observed in humans suffering from sleep deprivation. Previous research has demonstrated that sleep deprivation in the Forward-Forward algorithm has a catastrophic effect on learning efficacy. To mitigate this issue, we explore several approaches; these include alternative activation, optimized loss function, and threshold tuning. To simulate periodic rest, we reduce the number of positive passes in alternating epochs, creating short break phases. We additionally investigate the potential of caffeine-induced stimulation to enhance performance during sleep-deprived conditions. Experimental evaluations conducted on the MNIST and Fashion-MNIST datasets demonstrate that these modifications improve accuracy under the context of sleep deprivation. For example, a 2%-62% accuracy gain is observed in a severe sleep deprivation setting (16 positive or awake periods and 1 negative or sleep period). The approaches also enhance the resilience of the algorithm and its alignment with the adaptive mechanisms of human cognition.
Chinese Translation
本文研究了前向-前向算法中睡眠剥夺所带来的挑战。该算法中两个前向传递的分离以及两个传递过程中数据处理的不平衡,被认为是对睡眠剥夺人群认知过程的一种模拟。先前的研究表明,睡眠剥夺会对前向-前向算法的学习效果产生灾难性的影响。为了缓解这一问题,我们探索了多种方法,包括替代激活函数、优化的损失函数以及阈值调优。为了模拟周期性休息,我们减少交替轮次中正样本传递的次数,从而形成短暂的休息阶段。此外,我们还研究了咖啡因刺激在睡眠剥夺条件下提升性能的潜力。在MNIST和Fashion-MNIST数据集上进行的实验评估表明,这些改进措施能够在睡眠剥夺情境下提升准确率。例如,在严重睡眠剥夺的设置下(16个正样本或清醒阶段和1个负样本或睡眠阶段),准确率提升了2%至62%。这些方法还增强了算法的鲁棒性,并使其更符合人类认知的自适应机制。
cs.LG / 41 / 2609.06053

Data Quality Rule Generation with LLMs

基于大语言模型的数据质量规则生成
Glock, Anna-Christina, Hütter, Thomas, Fürnkranz, Johannes, Wöß, Wolfram, Dominka-Kiss, Christine, Ehrlinger, Lisa
Abstract
The validation of data, such as customer and employee data, is an important task in many organizations. Errors in data can have severe consequences. For example, a wrong drug unit in a patient record can lead to life-threatening medication errors, and a missing street number in an address to failed deliveries. Companies often employ rule-based enterprise data quality (DQ) tools, which allow domain experts to specify rules to validate the data over time. While rule-based DQ tools are computationally efficient and provide explainable reports, maintaining a comprehensive rule set manually is challenging, as domain experts often overlook essential rules, especially in complex domains and large data volumes. Hence, closing these gaps remains an open problem in practice. In this paper, we address the challenge of automated DQ rule generation. For this, we formalize a generalizable generate-filter framework and introduce LeDQeR, an LLM-based DQ rule generation approach. First, a large language model (LLM) generates candidates rules from an observed dirty data tuple for a given rule-based DQ tool syntax. Second, we apply four filter techniques that ensure the (i) executability, (ii) correctness, and (iii) generalizability, and avoid (iv) redundancy of the generated rules. An extensive experimental evaluation suggests that LeDQeR is able to produce effective and compact rule sets for various datasets and error types.
Chinese Translation
数据验证(如客户数据和员工数据的验证)是许多组织中的重要任务。数据错误可能造成严重后果。例如,患者记录中错误的药物剂量单位可能导致危及生命的用药错误,而地址中缺失的门牌号则可能导致投递失败。企业通常采用基于规则的企业数据质量(DQ)工具,允许领域专家制定规则以持续验证数据。虽然基于规则的DQ工具在计算上高效并能提供可解释的报告,但手动维护一套全面的规则集极具挑战性,因为领域专家往往会遗漏关键规则,尤其是在复杂领域和大数据量的情况下。因此,如何弥补这些不足在实践中仍是一个悬而未决的问题。本文致力于解决数据质量规则自动生成的挑战。为此,我们形式化了一个可泛化的“生成-过滤”框架,并提出了一种基于大语言模型(LLM)的DQ规则生成方法LeDQeR。首先,大语言模型根据给定基于规则的DQ工具的语法,从观测到的脏数据元组中生成候选规则。其次,我们应用四种过滤技术,以确保生成规则的(i)可执行性、(ii)正确性和(iii)可泛化性,并避免(iv)冗余。广泛的实验评估表明,LeDQeR能够为多种数据集和错误类型生成有效且精简的规则集。
cs.LG / 42 / 2609.06060

Calendar-SPCA: Interpretable Representation Learning for Multi-Periodic Electricity Consumption Profiles

Calendar-SPCA:面向多周期用电负荷曲线的可解释表示学习方法
Quesada-Granja, Carlos, Castillo-Calzadilla, Tony, Rizo-Maestre, Carlos
Abstract
Long-term electricity-consumption profiles exhibit several simultaneous periodic structures, including daily, weekly, and annual cycles. This work introduces Calendar-SPCA, a calendar-structured sparse principal component method that incorporates this known multi-periodic geometry directly into low-dimensional representation learning. The feature domain is represented as the Cartesian product of cyclic calendar axes, and a low-rank factorization is estimated using an L1 loading penalty together with graph total variation over the resulting calendar graph. The method therefore produces sparse and locally coherent loading patterns that remain directly readable in their original temporal coordinates. Calendar-SPCA is evaluated on two independent smart-meter datasets with different sample sizes and temporal resolutions: GoiEner and Low Carbon London. A factorial experiment characterizes the complementary effects of sparsity and calendar coherence and examines robustness across sample size, latent dimensionality, and repeated fits. At rank 15, Calendar-SPCA retains 96.92% and 82.90% of the explained variance of rank-matched PCA in GoiEner and Low Carbon London, respectively, while producing mean loading sparsities of 61.95% and 81.50%. Comparisons with classical sparse PCA and SPCA-TV further show that Calendar-SPCA adds a systematic organization of the latent factors in the original calendar coordinates while preserving substantial low-rank information. The resulting components form coherent and complementary daily, weekly, seasonal, and jointly localized calendar patterns, with dataset-specific geometries across the two datasets.
Chinese Translation
长期用电负荷曲线呈现出多种同时存在的周期结构,包括日周期、周周期和年周期。本文提出了Calendar-SPCA,这是一种日历结构化的稀疏主成分方法,将已知的这种多周期几何结构直接纳入低维表示学习之中。特征域被表示为循环日历轴的笛卡尔积,并通过L1载荷惩罚与所得日历图上的图全变差(graph total variation)进行低秩分解估计。因此,该方法产生了稀疏且局部连贯的载荷模式,这些模式在原始时间坐标中保持直接可读性。Calendar-SPCA在两个样本量和时间分辨率各不相同的独立智能电表数据集上进行了评估:GoiEner和Low Carbon London。因子实验刻画了稀疏性与日历连贯性的互补效应,并检验了其在样本量、潜在维度以及重复拟合方面的稳健性。在秩为15时,Calendar-SPCA在GoiEner和Low Carbon London数据集上分别保留了与同秩PCA相比96.92%和82.90%的解释方差,同时产生平均载荷稀疏度分别为61.95%和81.50%。与经典稀疏PCA和SPCA-TV的比较进一步表明,Calendar-SPCA在保持大量低秩信息的同时,在原始日历坐标中对潜在因子进行了系统化组织。所得到的成分形成了连贯且互补的日、周、季节以及联合局部化的日历模式,且两个数据集呈现出各自特定的几何结构。
cs.LG / 43 / 2609.06072

ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs

ACE:面向MoE大语言模型参数高效微调的跨专家适配器整合方法
Lee, Ahin, Yun, Sehyun, Park, Joonha, Gong, Taesik
Abstract
Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-wise design fragments adaptation in three ways: capacity is split across narrow low-rank updates, gradient supervision becomes sparse and imbalanced under sparse routing, and execution is decomposed into many small GEMMs. We find that such expert-wise separation is often unnecessary, as subsets of LoRA adapters become functionally similar during fine-tuning, revealing redundancy among expert-specific adapters. Based on this redundancy, we propose ACE (Adapter Consolidation across Experts), which groups redundant experts and replaces their expert-specific adapters with group-shared higher-rank LoRA modules under the same PEFT budget. ACE further introduces grouped adapter execution, which consolidates fragmented expert-wise adapter computations into fewer, larger group-level GEMMs. Across evaluations covering 12 datasets and four MoE backbones, ACE achieves the highest observed mean accuracy among the parameter-matched PEFT methods on the three backbones with complete baseline coverage, while providing $1.31\times$ to $1.48\times$ wall-clock training speedup over expert-wise LoRA without increasing peak memory. Our code is available at https://github.com/UbiquitousAILab/ACE.
Chinese Translation
混合专家模型的参数高效微调(PEFT)通常为每个专家分别挂载一个独立的低秩适配器。这种逐专家的设计在三个方面造成适配的碎片化:容量被分割到多个窄低秩更新中;在稀疏路由下梯度监督变得稀疏且不均衡;执行过程被分解为大量小型GEMM运算。我们发现,这种逐专家的分离往往是不必要的,因为微调过程中LoRA适配器的子集会变得功能相似,这揭示了专家特定适配器之间的冗余性。基于这一冗余性,我们提出ACE(Adapter Consolidation across Experts,跨专家适配器整合),该方法将冗余专家分组,并在相同的PEFT预算下,用组共享的高秩LoRA模块替代其专家特定的适配器。ACE进一步引入了分组适配器执行机制,将碎片化的逐专家适配器计算整合为更少、更大的组级GEMM运算。在涵盖12个数据集和四个MoE骨干模型的评估中,ACE在三个拥有完整基线覆盖的骨干模型上,取得了参数匹配的PEFT方法中观测到的最高平均准确率,同时相比逐专家LoRA提供1.31倍至1.48倍的实际训练加速,且不增加峰值内存。我们的代码已在 https://github.com/UbiquitousAILab/ACE 公开。
cs.LG / 44 / 2609.06073

FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon

FedSubMuon:基于结构化子空间Muon的高效通信联邦大语言模型微调方法
Chen, Shaolong, Tao, Youming, Chen, Shuzhen, Dressler, Falko, Ye, Qingqing, Wang, Di
Abstract
Federated fine-tuning adapts large language models (LLMs) to decentralized client data, but its scalability in cross-device training is often limited by the high communication cost. Muon is an optimizer that improves optimization performance by orthogonalizing momentum for matrix-valued parameters. Existing federated Muon methods demonstrate the benefit of matrix-aware optimization in federated learning, but still require transmitting full layer-size updates and optimizer state. A natural way to reduce communication is to directly apply Muon to LoRA factors, but this changes the optimized object and weakens Muon's matrix-aware update geometry. We propose FedSubMuon, a communication-efficient federated Muon fine-tuning method that optimizes compact coefficient matrices within shared structured subspaces. This design keeps Muon on a single matrix-valued trainable object, while reducing the client upload to compact coefficient matrices. We further introduce FedSubMuon-GT, an accuracy-oriented extension that uses projected gradients to adapt tracked subspace bases toward task-relevant gradient directions. Experiments on instruction tuning and mathematical reasoning show that FedSubMuon-GT achieves the best overall accuracy on four of five dataset-model pairs, while FedSubMuon performs best under all matched communication budgets. On Dolly-15K, the closest communication baseline requires 5.5 times and 1.4 times more total communication on Llama-1B and Qwen-4B, respectively.
Chinese Translation
联邦微调可以将大语言模型(LLM)适配到去中心化的客户端数据上,但其在跨设备训练中的可扩展性往往受到高昂通信成本的限制。Muon是一种通过对矩阵值参数的动量进行正交化来提升优化性能的优化器。现有的联邦Muon方法证明了矩阵感知优化在联邦学习中的优势,但仍需传输完整的层尺寸更新和优化器状态。一种降低通信量的自然方式是直接将Muon应用于LoRA因子,但这会改变被优化的对象,并削弱Muon的矩阵感知更新几何特性。我们提出FedSubMuon,一种通信高效的联邦Muon微调方法,它在共享的结构化子空间内优化紧凑的系数矩阵。这一设计使Muon保持在单一矩阵值可训练对象上,同时将客户端上传内容缩减为紧凑的系数矩阵。我们进一步提出FedSubMuon-GT,一种面向精度的扩展方法,它利用投影梯度将所追踪的子空间基向与任务相关的梯度方向自适应调整。在指令微调和数学推理任务上的实验表明,FedSubMuon-GT在五个数据集-模型对中的四个上取得了最佳的整体精度,而FedSubMuon在所有匹配的通信预算下均表现最佳。在Dolly-15K数据集上,最接近的通信基线在Llama-1B和Qwen-4B上分别需要5.5倍和1.4倍的总通信量。
cs.LG / 45 / 2609.06076

Beyond Retraining-Free MoE Compression: A Cost-Normalized Study of Post-Compression Adjustment

超越免重训练的MoE压缩:一种成本归一化的压缩后调整研究
Hyeon, Sieun, Do, Jaeyoung
Abstract
Retraining-free MoE compression reduces deployment memory by pruning or merging experts, but often treats the compressed checkpoint as the final artifact. We argue that this view is incomplete: compressed MoE checkpoints are better understood as compressed initializations that benefit from a tiny post-compression adjustment stage. Across two MoE LLM backbones, four pruning/merging methods, three expert-retention ratios, and 28 benchmarks, we compare LM fine-tuning and teacher-based KD under matched small-data budgets and measured GPU costs. Using only 3,000 C4 examples and a single epoch of adjustment, Full FT recovers 37.3% of the original-to-compressed performance gap on average. Moreover, LM fine-tuning is more cost-effective than standard token-level KD, and full-parameter adjustment gives the strongest cost--recovery trade-off among the tested scopes. These results suggest that retraining-free compression should be paired with small post-compression adjustment to recover a substantial portion of the performance lost during compression.
Chinese Translation
免重训练的MoE(混合专家)压缩通过剪枝或合并专家来降低部署内存开销,但通常将压缩后的检查点视为最终产物。我们认为这一观点并不完整:压缩后的MoE检查点更应被理解为一种压缩后的初始状态,其可通过一个小规模的压缩后调整阶段获益。我们在两个MoE大语言模型骨干、四种剪枝/合并方法、三种专家保留比例以及28个基准测试上,在匹配的小数据预算和实测GPU成本下,比较了语言模型微调(LM fine-tuning)与基于教师模型的知识蒸馏(KD)。仅使用3,000条C4样本和单轮调整,全参数微调(Full FT)平均可恢复原始模型与压缩模型之间性能差距的37.3%。此外,语言模型微调比标准的token级知识蒸馏更具成本效益,且在所测试的调整范围中,全参数调整具有最强的成本—性能恢复权衡。这些结果表明,免重训练压缩应当与小型压缩后调整相结合,以恢复压缩过程中损失的大部分性能。
cs.LG / 46 / 2609.06080

PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us

PhenoBench:深度表型人类队列能告诉我们什么
Sapir, Gal, Diament, Alon, Wolf, Adva, Yaya-Stupp, Doron, Solodkin, Dikla Gelbard, Azouri, Dana, Etzion-Fuchs, Anat, Lutsker, Guy, Segal, Eran, Rossman, Hagai
Abstract
Deeply phenotyped cohorts combine clinical, imaging, molecular, and wearable observations across timescales from seconds to years. This breadth can reveal which measurements inform which health-related questions, but heterogeneous analyses are not directly comparable. We present PhenoBench, an executable benchmark built around the Human Phenotype Project, in which more than 13,000 participants have completed the initial visit. Each question fixes the target, eligible population, timing, and allowed information; its evaluation contract specifies the split, metric, baseline, and claim boundary. The benchmark defines 90 clinically grounded tasks across 15 domains and 26 input modalities. Measurements showed question- and representation-dependent predictive value, including positive, near-zero, and negative changes in held-out performance relative to matched baselines. We used PhenoBench to evaluate emerging tabular foundation models across 160 matched regression comparisons spanning 52 tasks. These models ranked above standard task-specific models in aggregate but, averaged across the three pretrained models within each cell, improved on ridge by a median of only 0.004 $R^2$ (95% CI, 0.002--0.006). We then used the same cohort data and evaluation contracts to evaluate 14 language models, collectively covering 40 tasks spanning phenotype recovery, classification, follow-up forecasting, and participant ordering. Without cohort-specific fitting, language models made informative predictions on some tasks, but showed task-specific capability gaps, shared failures of scale, and rarely surpassed models fitted on the same fields. PhenoBench turns a multimodal longitudinal cohort into a versioned, auditable evaluation system where new questions, measurements, and models can be added without redefining existing comparisons.
Chinese Translation
深度表型队列整合了从秒级到年级跨时间尺度的临床、影像、分子和可穿戴设备观测数据。这种广度可以揭示哪些测量能回答哪些健康相关问题,但异质性的分析结果无法直接比较。我们提出了PhenoBench,一个围绕人类表型计划(Human Phenotype Project)构建的可执行基准测试,该计划已有超过13,000名参与者完成了初始访问。每个问题都固定了目标、合格人群、时间节点和允许使用的信息;其评估契约规定了数据划分、指标、基线和结论边界。该基准定义了90个基于临床的任务,涵盖15个领域和26种输入模态。测量结果显示出依赖于问题和表示形式的预测价值,包括相对于匹配基线的正、近零和负的留出集性能变化。我们使用PhenoBench在涵盖52个任务的160组匹配回归比较中评估了新兴的表格基础模型。这些模型在总体上排名高于标准的任务专用模型,但在每个单元内的三个预训练模型取平均后,其相对于岭回归(ridge)的$R^2$提升中位数仅为0.004(95% CI,0.002–0.006)。随后,我们使用相同的队列数据和评估契约评估了14个语言模型,共同覆盖40个任务,涉及表型恢复、分类、随访预测和参与者排序。在没有进行队列特异性拟合的情况下,语言模型在某些任务上做出了有信息量的预测,但表现出任务特异性的能力差距、共同的规模失效,且很少超越基于相同字段拟合的模型。PhenoBench将一个多模态纵向队列转化为一个可版本化、可审计的评估系统,其中可以添加新问题、新测量和新模型,而无需重新定义已有的比较。
cs.LG / 47 / 2609.06083

Learning to Price and Stock Under Contextual and Censored Demand

在上下文相关与删失需求下的学习定价与库存决策
Han, Zean, Ding, Zezhen, Zhang, Jiheng
Abstract
To make optimal joint pricing and inventory control decisions is a critical challenge for modern retailers. In practice, retailers face changing market conditions where demands are influenced by various contextual factors, while simultaneously dealing with the difficulty of lost sales that obscure true demand information. However, existing approaches often fail to account for both contextual information and censored demand observations. We address this gap by presenting a framework where we model demand as a linear combination of basis functions with unknown coefficients, allowing for adaptive pricing and inventory decisions that respond to changing contexts. We propose an efficient algorithm to achieve regret bound $\mathcal{O}(K\sqrt{T}\log T)$ under concave revenue conditions and $\mathcal{O}(K^{2/3}T^{2/3}(\log T)^{1/2})$ for the general case, with matching lower bounds confirming optimality. Extensive numerical experiments across diverse scenarios demonstrate our algorithm's effectiveness.
Chinese Translation
制定最优的联合定价与库存控制决策是现代零售商面临的关键挑战。在实践中,零售商面临不断变化的市场环境,需求受各种上下文因素影响,同时还要应对销售损失导致真实需求信息被掩盖的困难。然而,现有方法往往未能同时考虑上下文信息和删失需求观测。我们针对这一空白,提出了一个将需求建模为具有未知系数的基函数线性组合的框架,从而实现能够响应变化上下文的自适应定价与库存决策。我们提出了一种高效算法,在凹收益条件下可实现 $\mathcal{O}(K\sqrt{T}\log T)$ 的遗憾界,在一般情形下可实现 $\mathcal{O}(K^{2/3}T^{2/3}(\log T)^{1/2})$ 的遗憾界,并且匹配的下界证明了其最优性。在多种场景下的大量数值实验验证了我们算法的有效性。
cs.LG / 48 / 2609.06093

Connectome-to-Function: Conditional Generative Latent Representations for Reservoir Computing

从连接组到功能:面向储备池计算的条件生成式潜在表示
Yu, Zhuolin, Liu, Xingyu, Jia, Yuanhao, Xiao, Yunhang, Xue, Hairuo, Sun, Feihan, Chen, Guozhang
Abstract
Connectomes, graph-level maps of neurons and their synaptic connections, provide a structural basis for understanding how brain circuits support function and computation. However, mapping connectome structure to computation remains difficult because these graphs are high-dimensional, sparse, and sensitive to local structural variation. Existing approaches often depend on hand-crafted structural descriptors or task-specific predictors, which limits their ability to represent connectomes in a form that is both generative and functionally meaningful. We propose a conditional generative latent framework that encodes connectome graphs into a compact structural space while using available node-level conditions to guide reconstruction and generation. From this space, the model can reconstruct observed connectivity with a mean edge-reconstruction AUC up to 0.910 and generate new candidate connectomes, enabling a unified analysis of graph structure and computational behavior. Using connectome-derived graphs as recurrent computational substrates, we found that the learned latent space captures functional variation across reservoir-computing experiments, with cross-validated $R^2$ values up to approximately 0.87. Interpretability analysis further revealed task-specific structural mechanisms: in our examples, memory performance is associated with reciprocal recurrent connectivity, whereas prediction and classification are more strongly associated with spectral properties of the recurrent network. These findings suggest an AI-for-science approach to linking neural connectivity to computation and provide a generative and interpretable basis for studying how distinct structural mechanisms shape computational capacity.
Chinese Translation
连接组(Connectomes)是神经元及其突触连接的图级映射,为理解神经环路如何支持功能与计算提供了结构基础。然而,将连接组结构映射到计算仍然十分困难,因为这些图具有高维、稀疏以及对局部结构变化敏感的特点。现有方法通常依赖手工设计的结构描述符或面向特定任务的预测器,这限制了它们以兼具生成性和功能意义的形式表示连接组的能力。我们提出了一种条件生成式潜在框架,将连接组图编码到一个紧凑的结构空间中,同时利用可获得的节点级条件来引导重建与生成。基于该空间,模型能够重建已观测的连接(平均边重建AUC高达0.910),并生成新的候选连接组,从而实现对图结构与计算行为的统一分析。将连接组衍生的图用作循环计算基底,我们发现所学到的潜在空间能够捕捉储备池计算(reservoir computing)实验中的功能性变化,交叉验证的 $R^2$ 值高达约0.87。可解释性分析进一步揭示了特定任务的结构机制:在我们的示例中,记忆性能与双向循环连接相关,而预测与分类则与循环网络的谱特性更为密切相关。这些发现提出了一种连接神经连接与计算的"AI for Science"方法,并为研究不同结构机制如何塑造计算能力提供了生成性与可解释性的基础。
cs.LG / 49 / 2609.06100

VERPO: Verified Evidence Regularized Policy Optimization

VERPO:基于验证证据正则化的策略优化
Li, Haijiang, Lv, Chengyu, Zhang, Yi, Zhang, Zhibing, Qian, Rui, Zhang, Yuchen, Zhang, Xiaofan, Wang, Mingshan, Jing, Xiaofei, Tong, Yu, Zhou, Cangqi
Abstract
Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised. Evidence-conditioned Teachers provide denser supervision by replaying sampled trajectories with privileged feedback. Yet indiscriminate imitation risks transferring formatting or reasoning-style shifts that do not support task success. We introduce VERPO, a Verified Evidence Regularized Policy Optimization framework that treats evidence as a proposal for policy correction while retaining the outcome objective. It separates evidence-free reference restoration from signed token-level evidence corrections. Fisher Evidence Contrast attenuates corrections along an estimated evidence-presence direction. A stopped token-wise ZPD controller scales acceptance according to local reward alignment and Fisher movement cost, while the reference channel remains independent of acceptance. Across five scientific-reasoning and tool-use tasks, the best variant on each backbone exceeds the strongest compared baseline in average score. The averages rise from 0.6826 to 0.6857 on Qwen3-4B, from 0.6895 to 0.7058 on Qwen3-8B, and from 0.4751 to 0.5657 on Llama-3.2-1B.
Chinese Translation
可验证的结果奖励(outcome rewards)为语言模型的后训练提供了指导,但序列级的优势(advantage)无法识别哪些词元级(token-level)决策应当保留或修正。证据条件化的教师模型(Evidence-conditioned Teachers)通过以特权反馈重放采样轨迹来提供更密集的监督。然而,不加区分的模仿存在转移格式或推理风格变化的风险,而这些变化并不有助于任务成功。我们提出了 VERPO(Verified Evidence Regularized Policy Optimization,验证证据正则化策略优化)框架,该框架将证据视为策略修正的提议,同时保留结果目标。它将无证据的参考恢复与带符号的词元级证据修正分离开来。Fisher 证据对比(Fisher Evidence Contrast)沿估计的证据存在方向衰减修正。一个带停止机制的逐词元 ZPD 控制器根据局部奖励一致性和 Fisher 移动成本来调节接受程度,同时参考通道保持与接受机制无关。在五个科学推理与工具使用任务上,每个骨干模型上的最佳变体在平均得分上均超过了最强对比基线。平均得分在 Qwen3-4B 上从 0.6826 提升至 0.6857,在 Qwen3-8B 上从 0.6895 提升至 0.7058,在 Llama-3.2-1B 上从 0.4751 提升至 0.5657。
cs.LG / 50 / 2609.06106

FANS: Federated Adaptive Network Search Learning for Heterogeneous Devices

FANS:面向异构设备的联邦自适应网络搜索学习
Zhang, Jiaxin, Wang, Xingwei, Yi, Bo, Zhao, Liang, Furutanpey, Alireza, Chen, Ziyi, He, Qiang, Li, Keqin, Dustdar, Schahram
Abstract
Heterogeneous Federated Learning (HFL) aims to train models across devices with diverse resource budgets while preserving data privacy. Existing HFL methods typically bind training to a small predefined menu of model configurations, which limits architectural coverage. To address this bottleneck, we introduce Federated Adaptive Network Search (FANS), a hypernetwork-based framework that learns a shared architecture space rather than a fixed set of client models. To optimize this shared space efficiently, we propose the Federated Parallel Scaling (FPS) algorithm, which jointly trains multiple sampled subnetworks in parallel with self-distillation so that larger sampled subnetworks can supervise smaller ones during local updates. We evaluate FANS on CIFAR-10, CIFAR-100, and MNLI using ResNet-18, DenseNet-121, and BERT-base, respectively. Across all benchmarks, FANS expands the feasible subnetwork pool by orders of magnitude (e.g., 4,680 candidates for ResNet-18 vs. 4 in existing methods) and improves the average accuracy-efficiency trade-off relative to representative HFL baselines. Device heterogeneity is emulated through resource tiers, and evaluation covers accuracy, parameter count, and MACs.
Chinese Translation
异构联邦学习旨在跨具有不同资源预算的设备训练模型,同时保护数据隐私。现有的HFL方法通常将训练绑定到一个较小的预定义模型配置集合中,这限制了架构的覆盖范围。为解决这一瓶颈,我们提出了联邦自适应网络搜索,这是一种基于超网络的框架,它学习一个共享的架构空间,而不是一组固定的客户端模型。为了高效地优化这一共享空间,我们提出了联邦并行缩放算法,该算法并行地联合训练多个采样的子网络,并结合自蒸馏,使较大的采样子网络能够在本地更新过程中监督较小的子网络。我们分别在CIFAR-10、CIFAR-100和MNLI数据集上使用ResNet-18、DenseNet-121和BERT-base对FANS进行了评估。在所有基准测试中,FANS将可行的子网络池扩大了几个数量级(例如,ResNet-18有4,680个候选模型,而现有方法仅有4个),并且相对于代表性的HFL基线方法,提升了平均准确率-效率权衡。设备异构性通过资源层级进行模拟,评估指标涵盖准确率、参数量和乘加运算量。
cs.LG / 51 / 2609.06107

DataFlex-RL: An Evaluation Platform for RLVR Data Policies

DataFlex-RL:一个用于RLVR数据策略的评估平台
Liang, Hao, Chen, Mingrui, Feng, Hengyi, Qiang, Meiyi, Zhang, Wentao
Abstract
Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. Our primary experiment evaluates 13 configurations across 12 matched seeds using Qwen2.5-7B-Base and 12 mathematics, logic, and science benchmarks. Uniform GRPO improves the domain-balanced average accuracy by 7.76 percentage points over the untrained checkpoint. None of the eight rollout-selection or reweighting methods achieves a paired 95% confidence interval that excludes zero relative to uniform sampling, and none of the three adaptive mixtures outperforms a fixed equal mixture at the same level of precision. A corrected 12-seed extension on Llama-3.1-8B-Base places the additional methods on the same score scale as the original controls, but does not reveal a consistent winner in terms of observed mean performance. We also quantify evaluation sensitivity by rescoring nine Qwen2.5-7B-Instruct runs using a math-heavy six-benchmark summary, consisting of five mathematics benchmarks and GPQA-Diamond but no logic benchmark, and comparing it with the domain-balanced 12-benchmark summary. The resulting rankings are negatively correlated, with a correlation coefficient of -0.33, whereas summaries that retain all 12 benchmarks largely agree. Across the controlled settings studied here, changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.
Chinese Translation
可验证奖励强化学习(RLVR)的数据策略决定了哪些 rollout 被使用、其权重强度如何,以及哪些领域贡献到后续的训练批次中。我们介绍了 DataFlex-RL,一个在统一的 GRPO 训练方案下比较这些选择的评估平台。我们的主要实验使用 Qwen2.5-7B-Base 模型和 12 个数学、逻辑与科学基准,在 12 个匹配的随机种子上评估了 13 种配置。均匀的 GRPO 相对于未训练的检查点,将领域均衡的平均准确率提升了 7.76 个百分点。八种 rollout 选择或重新加权方法中,没有任何一种相对于均匀采样 achieves 达到排除零的配对 95% 置信区间;三种自适应混合方法中也没有任何一种在相同的精度水平上优于固定的等权混合。在 Llama-3.1-8B-Base 上进行的修正后的 12 种子扩展实验将额外的方法置于与原始对照组相同的评分尺度上,但并未在观测到的平均性能方面揭示出一致的优胜者。我们还通过使用一个偏重数学的六基准摘要(由五个数学基准和 GPQA-Diamond 组成,但不包含逻辑基准)对九次 Qwen2.5-7B-Instruct 运行重新评分,并将其与领域均衡的 12 基准摘要进行比较,从而量化了评估的敏感性。结果显示两者的排名呈负相关,相关系数为 -0.33,而保留全部 12 个基准的摘要则基本一致。在本文所研究的受控设置中,改变数据策略可测量地改变了训练过程,但并未产生相对于均匀训练的可复现改进。
cs.LG / 52 / 2609.06109

Sparse Incident-Cluster Learning for 12-hour Port Flood Pre-warning in Digital-Twin Analytics

面向数字孪生分析的12小时港口洪水预警的稀疏事件聚类学习
Zhang, Jie, Ni, Qiang, Windridge, David, Nguyen, Huan X.
Abstract
Port flood digital twins require analytics that warn operators before disruption, but official warning incidents are often few and adjacent observations are temporally dependent. Row-level classification can therefore overstate performance by placing windows from the same event in both model-development and evaluation data. We formulate 12-hour port flood pre-warning as an incident-cluster learning problem and evaluate a digital-twin analytics module using eight-point water-level histories, prediction-time contextual covariates, and interpretable short-window dynamics. The protocol combines fold-specific sparse feature selection, warning-cluster grouping, negative-label controls, 100-repeat random top-k controls, and alert-episode evaluation. Liverpool is the primary four-cluster case study, with harmonised Humber/Hull-proxy and Wessex South data used for protocol-transfer checks. Across the Liverpool folds, the top-10 ElasticNet model achieves mean F2 = 0.696, compared with 0.633 without top-k truncation and 0.681 for full-feature weighted XGBoost. It is the strongest ElasticNet variant, remains competitive with the nonlinear reference using only ten predictors, and exceeds the repeat-level 95th percentile of broad and same-family random subsets. Contextual covariates provide a strong prediction-time anchor, complemented by physically interpretable local dynamics. Historical replay converts risk scores into alert episodes and measures alert duration and false-episode burden. The result is an offline-evaluated analytics and validation module designed for integration into a port digital twin.
Chinese Translation
港口洪水数字孪生需要在扰动发生前向运营人员发出预警的分析能力,但官方预警事件通常数量稀少,且相邻观测在时间上相互依赖。因此,逐行分类方法可能因将同一事件的窗口同时置于模型开发和评估数据中而高估性能。我们将12小时港口洪水预警形式化为一个事件聚类学习问题,并使用八点水位历史、预测时刻上下文协变量以及可解释的短窗口动态来评估一个数字孪生分析模块。该协议结合了逐折稀疏特征选择、预警聚类分组、负标签对照、100次重复的随机top-k对照以及警报事件评估。利物浦(Liverpool)作为主要的四聚类案例研究,同时使用协调一致的亨伯/赫尔代理(Humber/Hull-proxy)和韦塞克斯南部(Wessex South)数据进行协议迁移检验。在利物浦各折中,top-10 ElasticNet模型取得了平均F2 = 0.696的成绩,高于无top-k截断时的0.633以及全特征加权XGBoost的0.681。它是最强的ElasticNet变体,仅使用十个预测变量即可与非线性基准模型保持竞争力,并超过了宽泛及同族随机子集在重复水平上的第95百分位数。上下文协变量提供了强有力的预测时刻锚点,并由物理上可解释的局部动态加以补充。历史回放将风险评分转换为警报事件,并度量警报持续时间和虚假事件负担。最终成果是一个经过离线评估的分析与验证模块,旨在集成到港口数字孪生系统中。
cs.LG / 53 / 2609.06133

Compressed Recurrent Feedback in Tsetlin Machines: A Reproducible Boolean-FSM Study

Tsetlin机中的压缩循环反馈:一项可复现的布尔有限状态机研究
Kumar, Ankit, Raj, Utkarsh, Shafik, Rishad, Roy, Sudip
Abstract
Sequential inference on small devices requires a model to retain useful history without repeatedly processing a long input record. A Recurrent Tsetlin Machine (RTM) provides this memory by returning Boolean clause outputs from one time step as inputs to the next. Direct feedback, however, grows with the clause bank and can make the recurrent input unnecessarily wide. This paper investigates a fixed-width alternative. We combine clause activations by exclusive-OR (XOR) folding, retain the folded bits at two time scales, and threshold them back to a binary state. The resulting design reduces 480 clause activations to 96 recurrent bits. We evaluate the method on a reproducible Boolean finite-state-machine benchmark with explicit transition rules, data splits, and random seeds. Across 144 runs, the compressed model obtains $61.47 \pm 6.74\%$ and $62.94 \pm 9.92\%$ accuracy on the two task families. Raw clause feedback changes these means by less than one percentage point, while increasing the recurrent width tenfold and measured host execution time by $4.38\times$ and $3.71\times$. Gated neural models remain more accurate, and a no-feedback control retaining only short input history achieves comparable or slightly higher accuracy. On this benchmark, folding matches raw feedback within small empirical margins at a much narrower interface; these findings also underscore the critical necessity of no-feedback recurrence controls when benchmarking sequence models.
Chinese Translation
在小型设备上进行顺序推理时,模型需要保留有用的历史信息,而不必反复处理冗长的输入记录。循环Tsetlin机(Recurrent Tsetlin Machine, RTM)通过将上一个时间步的布尔子句输出作为下一个时间步的输入来提供这种记忆。然而,直接反馈的规模会随子句库的大小而增长,可能使循环输入不必要地过宽。本文研究了一种固定宽度的替代方案。我们通过异或(XOR)折叠来合并子句激活值,在两个时间尺度上保留折叠后的比特,并将其阈值化回二进制状态。所得设计将480个子句激活压缩为96个循环比特。我们在一个具有明确转移规则、数据划分和随机种子的可复现布尔有限状态机基准上评估了该方法。在144次运行中,压缩模型在两个任务族上分别取得了 $61.47 \pm 6.74\%$ 和 $62.94 \pm 9.92\%$ 的准确率。原始子句反馈对这些均值的影响不足一个百分点,同时却使循环宽度增加了十倍,并使实测主机执行时间分别增加 $4.38\times$ 和 $3.71\times$。门控神经模型仍然更加准确,而仅保留短时输入历史的无反馈对照组取得了相当甚至略高的准确率。在该基准上,折叠方法在接口远窄于原始反馈的情况下,以很小的经验性差距达到了与原始反馈相当的效果;这些发现还强调了在基准测试序列模型时设置无反馈循环对照的关键必要性。
cs.LG / 54 / 2609.06154

Rethinking One-Shot Federated Graph Learning: Training-Free Statistical Estimation

重新思考单次联邦图学习:免训练的统计估计方法
Zheng, Shutong, Chen, Sijia
Abstract
One-shot federated graph learning generally aims to train Graph Neural Networks (GNNs) across clients with disconnected subgraphs in a single communication round. Existing methods predominantly design advanced optimization strategies under the premise that local GNN training is indispensable. However, empirical observations reveal that under extreme non-IID conditions, local GNN training suffers from severe cross-client representation misalignment, becoming a major source of error rather than a remedy. Motivated by this, we reformulate one-shot FGL as a statistical estimation problem. We propose SPEAR (Statistical Prototype Estimation with Adaptive Reliability), a completely training-free framework that directly computes topology-smoothed class prototypes from local graphs in the original feature space. The server then aggregates these prototypes using a sample-size-adaptive shrinkage estimator that down-weights unreliable local estimates, producing robust global class prototypes. Extensive experiments across seven benchmarks demonstrate that SPEAR consistently achieves state-of-the-art accuracy under extreme heterogeneity. Moreover, SPEAR delivers at least an order-of-magnitude speedup over all baselines, reaching several orders of magnitude against generative and distillation-based methods. Our findings suggest that training-free statistical estimation, rather than local GNN optimization, provides the key to robust and efficient one-shot federated graph learning. The code is available at https://github.com/Yodeesy/SPEAR .
Chinese Translation
单次联邦图学习(One-shot Federated Graph Learning)通常旨在仅通过一轮通信,在持有互不相连子图的多个客户端上训练图神经网络(GNN)。现有方法主要基于本地GNN训练不可或缺这一前提,设计先进的优化策略。然而,实证观察表明,在极端非独立同分布(non-IID)条件下,本地GNN训练会出现严重的跨客户端表示错位问题,反而成为误差的主要来源而非解决方案。受此启发,我们将单次联邦图学习(FGL)重新表述为一个统计估计问题。我们提出了SPEAR(基于自适应可靠性的统计原型估计,Statistical Prototype Estimation with Adaptive Reliability),这是一个完全免训练的框架,直接在原始特征空间中从本地图计算拓扑平滑的类原型。服务器随后使用样本量自适应的收缩估计器聚合这些原型,对不可靠的本地估计进行降权,从而产生鲁棒的全局类原型。在七个基准数据集上的大量实验表明,SPEAR在极端异质性条件下持续取得最先进的准确率。此外,SPEAR相比所有基线方法实现了至少一个数量级的加速,相对于生成式和基于蒸馏的方法更是达到了几个数量级的提升。我们的研究结果表明,免训练的统计估计方法,而非本地GNN优化,才是实现鲁棒且高效的单次联邦图学习的关键。代码已发布于 https://github.com/Yodeesy/SPEAR 。
cs.LG / 55 / 2609.06161

All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs

一切为了1比特:面向大语言模型的真正1比特训练后量化
Zhao, Zhixiong, Xu, Zukang, Sun, Guangyu, Liu, Lifeng, Yang, Dawei
Abstract
Large language models (LLMs) have achieved remarkable progress, yet their massive storage and memory-bandwidth demands still hinder efficient deployment. Weight binarization is a promising solution, but existing binarization-based post-training quantization (PTQ) methods usually far exceed the nominal 1-bit storage target due to hidden overhead. To address this gap, we propose All for 1-Bit (AF1), a genuine 1-bit PTQ framework for LLMs. AF1 comprises two complementary components: (1) Null-space-Aware Binary Factorization (NABF) for improving binary reconstruction through Hessian-aware surrogate reparameterization, null-space-aware binary factorization, and scale-only global reconstruction; and (2) Hierarchical Shapley Allocation (HiSA) for assigning structural capacity using hierarchical Shapley sensitivity. Together, they preserve model accuracy under a strict 1.0-BPW budget in the PTQ setting. Experiments on LLaMA, Qwen, and Gemma families show that AF1 consistently outperforms existing binarization-based PTQ methods in perplexity and zero-shot accuracy. Compared with BF16, AF1 achieves an average 2.5 times inference speedup and over 90% memory reduction across evaluated models, providing a practical path toward deployable genuine 1-bit compression for LLMs. The code for reproducibility is available at https://github.com/Kishon-zzx/AF1.
Chinese Translation
大语言模型(LLMs)已取得显著进展,但其巨大的存储和内存带宽需求仍阻碍了高效部署。权重二值化是一种有前景的解决方案,但现有的基于二值化的训练后量化(PTQ)方法由于隐藏开销,通常远超名义上的1比特存储目标。为弥合这一差距,我们提出了All for 1-Bit(AF1),一个面向大语言模型的真正1比特PTQ框架。AF1包含两个互补的组件:(1)零空间感知二值分解(Null-space-Aware Binary Factorization, NABF),通过Hessian感知的替代重参数化、零空间感知的二值分解以及仅缩放因子的全局重构来改进二值重构;(2)分层Shapley分配(Hierarchical Shapley Allocation, HiSA),利用分层Shapley敏感度为结构分配容量。二者结合,在PTQ设置下的严格1.0-BPW预算内保持了模型精度。在LLaMA、Qwen和Gemma系列模型上的实验表明,AF1在困惑度和零样本准确率上均持续优于现有基于二值化的PTQ方法。与BF16相比,AF1在所评估的模型上实现了平均2.5倍的推理加速和超过90%的内存缩减,为大语言模型可部署的真正1比特压缩提供了切实可行的路径。用于复现的代码已在 https://github.com/Kishon-zzx/AF1 公开。
cs.LG / 56 / 2609.06169

Decision-Aware Suffix Prediction and Reasoning of Business Processes

决策感知的业务过程后缀预测与推理
Mustroph, Henryk, Rinderle-Ma, Stefanie
Abstract
Suffix prediction forecasts the remaining sequence of events of a running case until completion. Most approaches rely on neural networks trained on event logs, which, on average, perform well but struggle with short prefixes or targets belonging to a rare process variant. In such scenarios, the correct path may cross multiple branching decisions, determined primarily by case- and event-level attributes, a signal that NN-based suffix prediction models tend to underweight because they may heavily weight (dense) event labels. Decision mining extracts rules for such decisions from the event log, but has so far been applied only to post-hoc and what-if analysis, not suffix prediction. We therefore extend suffix prediction with decision mining, introducing a decision-aware suffix prediction framework, a neuro-symbolic approach that enables reasoning about predicted events via mined decision rules. Experiments on three of four event logs and three suffix predictors show that the framework can improve suffix prediction, especially for short prefixes but also for rare process variants, and adds intrinsic interpretability.
Chinese Translation
后缀预测旨在预测一个正在运行的业务案例直至完成所剩余的事件序列。大多数方法依赖于在事件日志上训练的神经网络,其平均性能良好,但在处理短前缀或属于罕见流程变体的目标时表现不佳。在这些场景中,正确的路径可能跨越多个分支决策,而这些决策主要由案例级和事件级属性决定。基于神经网络的后缀预测模型往往低估这一信号,因为它们可能过度加权(密集的)事件标签。决策挖掘可以从事件日志中提取此类决策的规则,但迄今为止仅用于事后分析和假设分析,尚未应用于后缀预测。因此,我们将决策挖掘扩展到后缀预测中,提出了一种决策感知的后缀预测框架,这是一种神经符号(neuro-symbolic)方法,能够通过挖掘到的决策规则对预测事件进行推理。在四个事件日志中的三个以及三种后缀预测器上的实验表明,该框架能够改进后缀预测,尤其是对于短前缀和罕见流程变体,并具有内在的可解释性。
cs.LG / 57 / 2609.06173

FMMO: Detecting the Divergence Between Local Attribution and Global Drift

FMMO:检测局部归因与全局漂移之间的背离
Zafar, Muhammad Rehman, El-Sharif, Ali, Khan, Naimul
Abstract
Post-deployment drift poses a critical risk to algorithmic accountability, particularly when ground truth labels are delayed and performance degradation becomes a "silent failure". While Explainable AI (XAI) is often relied upon to audit these shifts, we demonstrate that popular local attribution methods (e.g., TreeSHAP) can exhibit misleading stability even as model reliability collapses. In this paper, we propose a Framework for Model Monitoring and Observability (FMMO) designed to expose the divergence between local explanation stability and global distribution shifts. Using benchmark, synthetic, and real-world datasets, we show that local XAI methods fail to flag drift-induced disparate impact, specifically where False Positive Rates spike for protected groups while feature attributions remain unchanged. By integrating global surrogate models with model utilization measurements, FMMO mitigates this fairness blind spot, ensuring that stakeholders can detect discriminatory deterioration that standard local XAI tools overlook.
Chinese Translation
模型部署后的漂移对算法问责构成关键风险,尤其是在真实标签延迟获取、性能退化沦为“静默失败”的情况下。尽管人们通常依赖可解释人工智能(XAI)来审计这些变化,我们证明了流行的局部归因方法(如 TreeSHAP)即使在模型可靠性崩溃时也可能表现出具有误导性的稳定性。本文提出了一个模型监控与可观测性框架(Framework for Model Monitoring and Observability, FMMO),旨在揭示局部解释稳定性与全局分布漂移之间的背离。基于基准数据集、合成数据集和真实世界数据集,我们展示了局部 XAI 方法无法标记漂移引起的差别性影响——具体表现为受保护群体的假阳性率飙升,而特征归因却保持不变。通过将全局代理模型与模型使用度量相结合,FMMO 缓解了这一公平性盲区,确保利益相关方能够检测到标准局部 XAI 工具所忽视的歧视性退化。
cs.LG / 58 / 2609.06186

Spectral Prioritized Sweeping in Nonstationary Reinforcement Learning

非平稳强化学习中的谱式优先扫描
Pham, Hung, Dam, Tuan
Abstract
Prioritized Sweeping (PS) accelerates model-based reinforcement learning by selecting backups according to Bellman residual magnitude. In nonstationary reward settings, however, the canonical priority score is shortsighted: after a localized reward shift, residuals propagate only through realized backups, so bottlenecked or topologically distant state estimates may remain static under a limited replanning budget. We introduce the Graph Topology Augmentation framework, which employ the graph's resolvent and its diffusion semantic, to augment the inquired signal. Our application, Graph Topology Augmentation for Prioritized Sweeping (GTA-PS), or which the alias Spectral Prioritized Sweeping (SPS) might be more universal, provides a drop-in ordering score for the setting of fixed dynamics and changing state rewards. GTA-PS uses a smootherized policy, inducing a transition chain, with its in- and out-Laplacian. The standard priority key is augmented with a mixing of regularized Laplacian inverses diffusing the residual magnitude. Furthermore, the topology contribution is annealed by a scheduler based on the Second Largest Eigenvalue Modulus (SLEM), allowing its scale to adapt to the chain's mixing regime. We prove that the forward potential coincides with geometric discounted residual propagation and show that GTA-PS gives active priority instantly to all states. Tabular experiments on FourRooms and GARNET domains demonstrate improved replanning efficiency over standard PS under both exact DP and Dyna-style host planners.
Chinese Translation
优先扫描(Prioritized Sweeping, PS)通过依据贝尔曼残差(Bellman residual)的大小选择备份操作,从而加速基于模型的强化学习。然而,在非平稳奖励环境中,传统的优先级评分具有短视性:当奖励发生局部变化后,残差仅通过实际执行的备份进行传播,因此在有限的重新规划预算下,处于瓶颈位置或拓扑距离较远的状态估计可能保持不变。我们提出了图拓扑增强(Graph Topology Augmentation)框架,利用图的预解算子(resolvent)及其扩散语义来增强所查询的信号。我们的应用——面向优先扫描的图拓扑增强(GTA-PS),其别名谱式优先扫描(Spectral Prioritized Sweeping, SPS)可能更具通用性——为固定动力学、变化状态奖励的设定提供了一种可直接插入的排序评分。GTA-PS 使用平滑化策略诱导出一个转移链,并利用其入拉普拉斯算子和出拉普拉斯算子。标准的优先级键通过混合正则化拉普拉斯逆算子来扩散残差大小而得到增强。此外,拓扑贡献由基于第二最大特征值模(SLEM)的调度器进行退火处理,使其尺度能够适应链的混合状态。我们证明了前向势能与几何折扣残差传播相一致,并表明 GTA-PS 能够立即为所有状态赋予有效优先级。在 FourRooms 和 GARNET 领域的表格型实验中,无论是精确动态规划(DP)还是 Dyna 风格的主规划器,GTA-PS 相比标准 PS 都展现出更高的重新规划效率。
cs.LG / 59 / 2609.06195

EgoNeMo: Transferable Map of Pedestrian Dynamics via Egocentric LiDAR Scan

EgoNeMo:基于自我中心激光雷达扫描的可迁移行人动力学地图
Sawada, Azusa, Wang, Allan, Saito, Hideo, Steinfeld, Aaron
Abstract
This paper proposes a transferable Map of Dynamics (MoD) framework that generalizes to unknown environments using only egocentric 3D LiDAR point clouds to overcome the long-standing limitation of traditional MoD methods. While MoDs are essential for encoding human motion characteristics to enable accurate pedestrian trajectory prediction or safe robot navigation, traditional approaches suffer from site-specificity, requiring exhaustive trajectory accumulation at every new location. Extending recent advances in neural implicit modeling, our framework trains a continuous, LiDAR-based MoD estimator across diverse environments. To mitigate the inherent sparsity and temporal bias of real-world trajectory data, we introduce a position-balanced sampling strategy and a multi-task learning architecture that jointly predicts motion distributions and a spatial frequency score map. The latter is further augmented by visibility-aware losses to compensate for incomplete observation data. Comprehensive experiments demonstrate that our method effectively reconstructs underlying motion maps even in unknown locations from a single instantaneous LiDAR scan, despite highly sparse training data. Finally, we show that our improvements enhance the reliability of downstream trajectory prediction.
Chinese Translation
本文提出一种可迁移的动力学地图(Map of Dynamics, MoD)框架,仅利用自我中心的3D激光雷达点云即可泛化到未知环境,从而克服了传统MoD方法长期存在的局限性。尽管动力学地图对于编码人体运动特性、实现准确的行人轨迹预测或安全的机器人导航至关重要,但传统方法存在场地特异性问题,需要在每个新地点进行穷尽式的轨迹数据积累。得益于神经隐式建模的最新进展,我们的框架在多样化环境中训练了一个基于激光雷达的连续MoD估计器。为缓解真实世界轨迹数据固有的稀疏性和时间偏差,我们引入了一种位置平衡采样策略和多任务学习架构,该架构联合预测运动分布和空间频率评分图。后者还通过感知可见性的损失函数进一步增强,以补偿观测数据的不完整性。大量实验表明,即使在训练数据高度稀疏的情况下,我们的方法也能仅凭单次瞬时激光雷达扫描在未知地点有效重建潜在的运动地图。最后,我们证明了这些改进提升了下游轨迹预测的可靠性。
cs.LG / 60 / 2609.06257

SeaCausal-FL: Federated Fuzzy Causal Learning for Maritime IoT Fault Diagnosis and Counterfactual Reasoning

SeaCausal-FL:面向海洋物联网故障诊断与反事实推理的联邦模糊因果学习
Qiu, Yuhang, Zhu, Haihan, Seerangan, Koteeswaran, Zhu, Longsheng, Wang, Xiong, Lu, Yijun, Lin, Zheng, Ren, Fangmin, Xie, Jialiang
Abstract
Reliable marine-engine fault diagnosis in maritime IoT is challenged by distributed data ownership, heterogeneous fault distributions, and continuously changing operating conditions. This paper proposes SeaCausal-FL, a federated fuzzy causal learning framework that combines a shared temporal diagnostic path with mechanism-conditioned causal reasoning. An interval type-2 fuzzy layer represents uncertain and overlapping operating mechanisms, while each mechanism is associated with a physics-constrained structural causal model. Before aggregation, locally learned mechanisms are aligned using operating context, causal structure, and conditional intervention-response signatures. Model parameters are then aggregated according to sample, class, mechanism, and mechanism-class evidence instead of client sample size alone. The learned structural equations further support interval counterfactual reasoning through abduction, action, and prediction. Experiments on a marine-engine fault dataset and a real-data-calibrated semi-synthetic causal benchmark show that SeaCausal-FL achieves an average F1 score of 87.07% across four client partitions, with AUROC and AUPRC of 98.98% and 94.81%, respectively. It also maintains strong performance under unseen loads and fault-type omission during training. On the causal benchmark, SeaCausal-FL reaches an Edge-F1 of approximately 0.58 and an Edge-AUPRC of 0.68, reduces coefficient RMSE to about 0.14, and provides favorable counterfactual estimation and intervention decisions.
Chinese Translation
海洋物联网中可靠的船用发动机故障诊断面临数据所有权分散、故障分布异构以及运行工况持续变化等挑战。本文提出了SeaCausal-FL,一种联邦模糊因果学习框架,该框架将共享的时序诊断路径与基于机制的因果推理相结合。区间二型模糊层(interval type-2 fuzzy layer)用于表示不确定且相互重叠的运行机制,同时每种机制都关联一个物理约束的结构因果模型。在聚合之前,利用运行上下文、因果结构和条件干预-响应特征对本地学习的机制进行对齐。随后,模型参数根据样本、类别、机制以及机制-类别证据进行聚合,而非仅依据客户端样本量。学习到的结构方程进一步支持通过 abduction(溯因)、action(行动)和prediction(预测)实现区间反事实推理。在船用发动机故障数据集和基于真实数据校准的半合成因果基准上的实验表明,SeaCausal-FL在四个客户端划分上取得87.07%的平均F1分数,AUROC和AUPRC分别为98.98%和94.81%。该模型在未见负载以及训练时故障类型缺失的情况下仍保持较强的性能。在因果基准上,SeaCausal-FL的Edge-F1约为0.58,Edge-AUPRC为0.68,将系数RMSE降低至约0.14,并在反事实估计和干预决策方面表现良好。
cs.LG / 61 / 2609.06294

Representation Learning for Sample-Efficient CATE Estimation by Leveraging Multiple Outcomes

利用多结果表示学习实现样本高效的CATE估计
Swaroop, Maitreyi, Bhat, Shikha, Rodriguez, Samantha, Krishnamurti, Tamar, Wilder, Bryan
Abstract
Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental samples, making it difficult to estimate heterogeneous effects from high-dimensional covariates. In such settings, policymakers and medical practitioners often succumb to the curse of dimensionality or apply off-the-shelf dimension reduction methods that may not preserve treatment heterogeneity. Yet these domains often come with large historical datasets measuring a wide range of outcomes -- a source of supervision that is rarely exploited in practice. Following causal representation learning, we hypothesize that such domains with high-dimensional covariates have lower-dimensional underlying dynamics. We can thus leverage the diverse outcomes measured in historical data to learn a lower-dimensional representation of the covariates. Theoretically, we prove that when the auxiliary outcomes satisfy a set of surrogacy conditions and the representation retains relevant covariate information, the original CATE is identified when the high-dimensional covariates are replaced by the learned representation. Combined with existing dimension-dependent rates for CATE estimation, the result implies greater sample-efficiency on the same experimental sample. Additionally, we characterize the bias-variance tradeoff when the assumptions do not hold perfectly, and show that the representation-based estimator can still achieve lower error when the reduction in estimator variance outweighs the bias due to compression. Empirically, we evaluate the method on synthetic data and semi-synthetic medical data.
Chinese Translation
估计条件平均处理效应(CATE)能够实现对干预措施的高效定向,但许多应用场景中实验样本有限,难以从高维协变量中估计异质性效应。在这类情境下,政策制定者和医疗从业者往往陷入维数灾难,或使用可能无法保留处理异质性的现成降维方法。然而,这些领域通常拥有记录了大量结果指标的大型历史数据集——这是一个在实践中很少被利用的监督信息来源。遵循因果表示学习(causal representation learning)的思想,我们假设此类高维协变量场景存在低维的底层动态机制。因此,我们可以利用历史数据中测量的多样化结果来学习协变量的低维表示。在理论上,我们证明当辅助结果满足一组替代性(surrogacy)条件且表示保留了相关协变量信息时,用学习到的表示替换高维协变量后,原始的CATE仍然是可识别的。结合现有的依赖于维度的CATE估计收敛速率,该结果意味着在相同的实验样本上可以获得更高的样本效率。此外,我们刻画了当假设不能完全成立时的偏差-方差权衡,并表明当估计器方差的降低超过压缩带来的偏差时,基于表示的估计器仍能实现更低的误差。在实验方面,我们在合成数据和半合成医疗数据上对该方法进行了评估。
cs.LG / 62 / 2609.06341

Linear Algebra Foundations of Efficient Attention: A Phase Reversal in Rank Collapse Under SVD Compression

高效注意力的线性代数基础:SVD压缩下秩坍缩的相位反转现象
Kalvakolanu, Anjaneya Teja Sarma
Abstract
Linear algebra provides the framework of concepts (matrix rank, singular value decomposition (SVD), and eigendecomposition) that modern artificial intelligence employs to encode, compress, and propagate information through neural networks. This paper unifies fourteen separate peer-reviewed works analyzing the usage of these techniques in the context of transformer-based foundation model research, focusing on three areas of the topic: derivations and properties of self-attention matrices' output rank, compression methods that purposefully utilize this phenomenon, and the low-rank key-value (KV) cache projection and its semiseparable-matrix duality to linear attention and state-space structured models. We were motivated to conduct this work after observing an open problem in this literature: the interplay of the mentioned compression methods with natural rank collapse of the network. With this paper, we report an original finding that using SVD compression of attention projections actually has the opposite effect on the rank collapse of the network: while it strongly suppresses it at initialization, it accelerates on pretrained models (for GPT-2 124M, GPT-2 Medium 355M, and Pythia-160M) with minimal risk of object aliasing artifacts appearing (verified on all compression ratios) and is consistent across four rank estimation methods. A controlled causal decomposition of the effect in both settings showed that the reason for this behavior can be explained by the choice of the subspace SVD makes when compressing the matrix better than the reduction of the operator norm it achieves, explaining roughly 76% of the effect at initialization and 83% on the pretrained weights, providing a refinement to the calibration-aware compression viewpoint and an explanation of why it outperformed naive SVD truncation.
Chinese Translation
线性代数提供了现代人工智能用于在神经网络中编码、压缩和传播信息的概念框架(矩阵秩、奇异值分解(SVD)和特征值分解)。本文统一整合了十四项独立发表的同行评审研究,分析了这些技术在基于Transformer的基础模型研究中的应用,聚焦于该主题的三个领域:自注意力矩阵输出秩的推导与性质、有意利用这一现象的压缩方法,以及低秩键值(KV)缓存投影及其与线性注意力和状态空间结构化模型之间的半可分矩阵对偶性。我们之所以开展这项工作,是因为观察到该文献中存在一个开放性问题:上述压缩方法与网络自然秩坍缩之间的相互作用。本文报告了一项原创发现:对注意力投影使用SVD压缩实际上会对网络的秩坍缩产生相反的影响:在初始化时它会强烈抑制秩坍缩,而在预训练模型上(GPT-2 124M、GPT-2 Medium 355M和Pythia-160M)则会加速秩坍缩,且出现对象混叠伪影的风险极低(在所有压缩比上均已验证),并在四种秩估计方法中保持一致。对这两种设置下该效应的受控因果分解表明,这种行为的原因可以由SVD在压缩矩阵时对子空间的选择来更好地解释,而非其所实现的算子范数缩减——该因素分别解释了初始化时约76%和预训练权重上约83%的效应,从而对校准感知压缩(calibration-aware compression)的观点提供了精细化修正,并解释了其为何优于朴素SVD截断方法。
cs.LG / 63 / 2609.06346

Robust Dynamic Expansion for Continual Learning under Backdoor Attacks via Purification and Selective Recovery

基于净化与选择性恢复的后门攻击下持续学习鲁棒动态扩展方法
Lin, Keyu, Ye, Fei, Liu, Qihe, Zhou, Shijie, Yu, Jiguo
Abstract
Continual learning (CL) enables models to acquire new knowledge from sequentially arriving tasks while retaining previously learned knowledge. However, in practical scenarios, task streams collected from untrusted sources may contain backdoor-poisoned samples, posing a critical challenge to the stability, plasticity, and security of continual learners. In this work, we investigate a challenging setting termed Continual Learning Under Backdoor Attack (CLUBA), where each incremental task may involve a small proportion of maliciously manipulated training samples. Unlike conventional continual learning or backdoor defense scenarios, CLUBA requires models to simultaneously mitigate catastrophic forgetting, preserve adaptation capability, and prevent the absorption of malicious supervision during sequential updates. To address this challenge, we propose a robust dynamic-expansion framework that integrates sample purification, selective recovery, and robust expert routing into a unified continual learning paradigm. Specifically, we introduce Bi-Prototype Purification (BPP) to identify suspicious samples by exploiting semantic discrepancies in feature space. Based on purified data, Gradient Discrepancy-based Robustness Optimization (GDBRO) selectively recovers informative poisoned samples through pseudo-label correction and gradient consistency evaluation, improving robustness while maintaining model plasticity. Furthermore, Robust Feature Consistency-based Expert Selection (RFCBES) constructs perturbation-aware class prototypes to enable reliable expert routing under corrupted or shifted inputs.
Chinese Translation
持续学习(Continual Learning, CL)使模型能够从按序到达的任务中获取新知识,同时保留已学习的旧知识。然而,在实际场景中,从不可信来源收集的任务流可能包含后门投毒样本,这对持续学习器的稳定性、可塑性和安全性构成了严峻挑战。在本工作中,我们研究了一种具有挑战性的场景,称为后门攻击下的持续学习(Continual Learning Under Backdoor Attack, CLUBA),其中每个增量任务可能包含少量被恶意篡改的训练样本。与传统的持续学习或后门防御场景不同,CLUBA 要求模型在序贯更新过程中同时缓解灾难性遗忘、保持适应能力,并防止吸收恶意监督信号。为应对这一挑战,我们提出了一种鲁棒的动态扩展框架,将样本净化、选择性恢复和鲁棒专家路由整合到统一的持续学习范式中。具体而言,我们引入双原型净化(Bi-Prototype Purification, BPP),通过利用特征空间中的语义差异来识别可疑样本。基于净化后的数据,基于梯度差异的鲁棒性优化(Gradient Discrepancy-based Robustness Optimization, GDBRO)通过伪标签校正和梯度一致性评估,选择性地恢复有信息量的被投毒样本,在保持模型可塑性的同时提升鲁棒性。此外,基于鲁棒特征一致性的专家选择(Robust Feature Consistency-based Expert Selection, RFCBES)构建扰动感知的类原型,以在输入被污染或发生偏移的情况下实现可靠的专家路由。
cs.LG / 64 / 2609.06367

Robust Conformal Consensus: Multi-Agent LLM-as-a-Judge Interval Evaluation with Conformal Prediction

鲁棒保形共识:基于保形预测的多智能体LLM-as-a-Judge区间评估
Liu, Lihui
Abstract
LLM-as-a-Judge has emerged as a promising paradigm for evaluating natural language generation. However, the uncertainty associated with such evaluations remains largely unexplored, which limits their reliability in real-world applications. Although conformal prediction offers a principled framework for uncertainty quantification, existing approaches typically apply it to a single LLM judge, overlooking the variability introduced by using different LLM evaluators. In this work, we propose a robust uncertainty estimation framework for multi-agent LLM-as-a-Judge evaluation. Our approach constructs conformal prediction intervals for LLM-based scores from multiple LLMs. By considering intervals from different LLM judges, we obtain more stable and reliable uncertainty estimates. Extensive experiments demonstrate that our method produces valid prediction intervals with coverage guarantees, and that interval-based aggregation across multiple judges leads to more stable evaluation outcomes.
Chinese Translation
LLM-as-a-Judge(大语言模型作为评判者)已成为自然语言生成评估的一种有前景的范式。然而,此类评估所涉及的不确定性在很大程度上尚未被探索,这限制了其在实际应用中的可靠性。尽管保形预测(Conformal Prediction)为不确定性量化提供了一个有原则的框架,但现有方法通常仅将其应用于单个LLM评判者,忽略了使用不同LLM评估器所带来的变异性。在本工作中,我们提出了一种面向多智能体LLM-as-a-Judge评估的鲁棒不确定性估计框架。我们的方法为来自多个LLM的基于LLM的评分构建保形预测区间。通过综合考虑不同LLM评判者的区间,我们获得了更加稳定可靠的不确定性估计。大量实验表明,我们的方法能够生成具有覆盖率保证的有效预测区间,并且基于区间的多评判者聚合能够带来更稳定的评估结果。
cs.LG / 65 / 2609.06384

Parameterized and Streaming Algorithms for Euclidean Fair $k$-Center Clustering

欧氏公平k-中心聚类的参数化与流式算法
Lin, Zeyu, Jia, Chaoqi, Guo, Longkun, Chen, Chao
Abstract
Motivated by the growing importance of fairness in machine learning, fair $k$-center clustering has attracted considerable research attention as a fundamental problem. In this problem, a dataset is partitioned into $m$ disjoint groups, and the objective is to select $k$ data points as centers, subject to upper bounds on the number of centers chosen from each group, aiming to minimize the maximum distance between any data point and its assigned center. Focusing on Euclidean spaces, which are ubiquitous in machine learning applications, we first develop a parameterized approximation algorithm for Euclidean fair $k$-center with an approximation ratio of $2.732$. By incorporating this algorithm as a post-processing stage into a one-pass streaming framework for large-scale data, we obtain an approximation ratio of $4.464$. These ratios can be further respectively improved to $2.414$ and $3.828$ with a runtime exponential on $k$. To ensure polynomial-time complexity, we further design a one-pass streaming algorithm with an approximation ratio of $4.732$, which can be further improved to $4.42$, outperforming the state-of-the-art ratio. Finally, extensive experiments show that our methods significantly outperform state-of-the-art approaches in terms of clustering accuracy.
Chinese Translation
随着公平性在机器学习中日益重要,公平k-中心聚类作为一个基础问题引起了广泛的研究关注。在该问题中,数据集被划分为$m$个互不相交的组,目标是从中选取$k$个数据点作为中心,并对从每个组中选取的中心数量施加上限约束,以最小化任意数据点到其被分配中心的最大距离。聚焦于机器学习应用中广泛存在的欧氏空间,我们首先为欧氏公平$k$-中心问题开发了一种近似比为$2.732$的参数化近似算法。通过将该算法作为后处理阶段嵌入到面向大规模数据的单遍流式框架中,我们获得了$4.464$的近似比。当运行时间允许关于$k$呈指数级时,上述两个近似比可分别进一步改进至$2.414$和$3.828$。为保证多项式时间复杂度,我们进一步设计了一种近似比为$4.732$的单遍流式算法,并可进一步改进至$4.42$,优于现有最先进的近似比。最后,大量实验表明,我们的方法在聚类精度方面显著优于现有最先进方法。
cs.LG / 66 / 2609.06386

Are Verifier Errors Independent Within a GRPO Group? Evidence from Qwen2.5 Rollouts

GRPO组内验证器误差是否相互独立?来自Qwen2.5采样轨迹的证据
Xin, Esther
Abstract
Group-based reinforcement learning with verifiable rewards (RLVR) scoresmultiple completions per prompt using automatic verifiers. Analysesbased on independent verifier errors may overlook dependence associatedwith shared answer formats. We investigate this dependence in24,998 groups of eight completions generated by Qwen2.5-1.5B onMATH, GSM8K, and DeepMath-103K. We estimate a pooled within-groupverifier-error correlation of 0.530 (95% confidence interval:0.500--0.560). Under an exchangeable-error model, this correspondsto a design-effect-adjusted effective sample size of 1.70 for aneight-completion group. Dependence varies substantially across answerforms: fractions, radicals, symbolic expressions, and intervals exhibitstronger clustering than unit annotations and percent signs. Replayinggroup-relative advantages across four rule-based verifier configurationsidentifies at least one advantage-sign disagreement in up to 0.83% ofgroups. Because a group is repeated sampling for one prompt, thiswithin-group clustering may reflect shared prompt difficulty as well asshared answer form, and we do not attempt to separate the two here.Unlike studies of correlated judgments across multiple evaluators, ouranalysis examines dependence across completions scored by the sameverifier. These findings motivate prompt- and answer-form-aware analysesof verifier noise rather than characterizations based solely onaggregate error rates.
Chinese Translation
基于可验证奖励的分组强化学习(RLVR)使用自动验证器对每个提示(prompt)的多个补全(completions)进行评分。基于验证器误差相互独立假设的分析可能忽略了与共享答案格式相关的误差依赖性。我们在Qwen2.5-1.5B模型于MATH、GSM8K和DeepMath-103K数据集上生成的24,998个由八个补全组成的组中研究了这种依赖性。我们估计得到合并的组内验证器误差相关系数为0.530(95%置信区间:0.500–0.560)。在可交换误差模型下,这对应于八个补全的组经设计效应调整后的有效样本量为1.70。这种依赖性在不同答案形式之间存在显著差异:分数、根式、符号表达式和区间比单位标注和百分号表现出更强的聚类性。通过在四种基于规则的验证器配置下重放组相对优势(group-relative advantages),我们发现最多有0.83%的组中至少存在一处优势符号不一致的情况。由于一个组即是对同一提示的重复采样,这种组内聚类可能既反映了共享的提示难度,也反映了共享的答案形式,本文不试图将两者分离。与跨多个评估者的相关性判断研究不同,我们的分析考察的是由同一验证器评分的多个补全之间的依赖性。这些发现表明,应当开展针对提示和答案形式的验证器噪声分析,而非仅基于总体错误率进行刻画。
cs.LG / 67 / 2609.06396

MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves

MetaRSI / RSI2:一个针对递归自我改进系统本身的元递归自我改进系统
Tan, Zihan, Sun, Leixin, Shi, Zitong, Liu, Yitao, Wu, Jiajun, Brooks, Nathaniel, Qian, Jiaru, Shang, Xiaoran, Huang, Suyuan, Ding, Yi, Liao, Yangxu, Li, Mukai, Sun, Qiushi, Liu, Shudong, Rong, Xuankun, Yu, Xiaohang, Chen, Zhuo, Geng, Hejia, Li, Chenxin, Wang, Aozhou, Tu, Zengji, Tang, Robert, Zhan, Yuxin, Jiang, Eric, Wu, Yuxin, Zhang, Jianqing, Liang, Xiao, Wu, Fang, Zhang, Haochi, Marlow, Alexander, Wan, Guancheng
Abstract
Recursive self-improvement (RSI) lets a system improve the model-building machinery from its own failures, so every later model inherits the gain. Yet RSI has been validated almost exclusively on coding and formal benchmarks such as science QA and mathematics. This format bound limits RSI to improvement within a machine-checkable slice, not general capability where questions are open and correctness is settled by argument, replication, or measurement. We argue RSI must next operate across real, diverse scientific, engineering, and meta-scientific domains, not where formal evaluation is merely tractable. To that end we present MetaRSI-v1, where improvement is the scheduled composition of three typed operators over one unified paradigm. Data-RSI amplifies existing competence and marks its boundary; Harness-RSI edits a five-slot scaffold without touching weights; Model-RSI internalizes capability into parameters through bounded training. Sharing one loop kernel and artifact vocabulary, they make data, scaffold, and model changes composable rather than exclusive. A two-axis optimizer jointly decides operator order and each operator's proposal policy, while a meta-level policy revises the schedule across terms. We validate MetaRSI-v1 under the field's standard evaluations, on code and closed-form science, with no external teacher: the target model plays every role in its own loop. MetaRSI-v1 reframes self-improvement from a single-surface edit to a composition across the full model-production pipeline, opening two paths: a model route internalizing capability through training, and a harness route leaving weights untouched and thus extending self-improvement to any model reachable through an interface, with Data-RSI redefined as the shared substrate feeding both. The framework further yields refutable laws on where loops exist, how operators compose, and what supervision buys.
Chinese Translation
递归自我改进(Recursive Self-Improvement, RSI)使系统能够从自身的失败中改进其模型构建机制,从而使后续每个模型都能继承这一增益。然而,RSI 几乎仅在编码以及科学问答和数学等形式化基准上得到验证。这种格式上的局限将 RSI 限制在机器可验证的领域内进行改进,而无法泛化到那些问题开放、正确性需通过论证、复现或测量来判定的一般能力领域。我们主张,RSI 下一步必须跨越真实、多样的科学、工程以及元科学领域运行,而不能仅停留在形式化评估易于实施的地方。为此,我们提出 MetaRSI-v1,其中改进被定义为在一个统一范式下三类类型化算子的计划性组合。数据 RSI(Data-RSI)放大既有能力并标定其边界;框架 RSI(Harness-RSI)在不改动模型权重的前提下编辑一个五槽位的脚手架结构;模型 RSI(Model-RSI)通过有界训练将能力内化到模型参数中。三者共享同一循环内核与工件词汇,使数据、脚手架与模型层面的改动得以组合而非互斥。一个双轴优化器联合决定算子执行顺序及各算子的提议策略,同时一个元层策略在多轮之间修订调度方案。我们在该领域的标准评估下(涵盖代码与封闭式科学问题)、在无外部教师的情况下验证了 MetaRSI-v1:目标模型在其自身的循环中扮演所有角色。MetaRSI-v1 将自我改进从单一边表面的修改重构为贯穿整个模型生产管线的组合,开辟了两条路径:一条是通过训练内化能力的模型路径,另一条是保持权重不变、从而将自我改进扩展到任何可通过接口访问的模型的框架路径,其中 Data-RSI 被重新定义为同时为两者提供输入的共享基底。该框架还进一步给出了若干可证伪的定律,涉及循环存在于何处、算子如何组合、以及监督能带来什么。
cs.LG / 68 / 2609.06421

On BatchNorm Forward Modes in Value-Based Reinforcement Learning

论基于价值的强化学习中批归一化的前向传播模式
Palenicek, Daniel, Henaff, Mikael, Fujimoto, Scott, Sinha, Koustuv
Abstract
Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies report performance degradation in discrete-action value learning on Atari. These failures are surprising because discrete Q-networks lack the action-input distribution mismatch identified by CrossQ. We show for target-based C51 and target-free PQN that the simple choice between running and batch statistics at specific forward passes can reverse this degradation. In C51, switching the BN bootstrap forward to batch-statistic mode significantly improves performance over unnormalized and LayerNorm baselines and scales stably with update-to-data ratios up to 12. In PQN, using batch-statistics for both action selection and bootstrapping recovers performance from the failing running-statistic configuration. Across 26 Atari games at 400M frames, this configuration achieves a higher final aggregate score than PQN with LayerNorm. Our results show that carefully configured BN can substantially improve discrete-action value learning, and that its forward protocols are an essential part of the algorithm specification.
Chinese Translation
批归一化(Batch Normalization, BN)显著提升了诸如 CrossQ 等连续控制 actor-critic 方法的样本效率,然而近期研究报道其在 Atari 离散动作价值学习中出现性能退化。这些失败令人意外,因为离散 Q 网络并不存在 CrossQ 中所识别的动作-输入分布不匹配问题。我们针对基于目标网络的 C51 和无目标网络的 PQN 证明,在特定前向传播中简单地在运行统计量(running statistics)与批统计量(batch statistics)之间进行选择,即可扭转这种性能退化。在 C51 中,将 BN 自举前向传播切换为批统计量模式,其性能显著优于未归一化基线和 LayerNorm 基线,并且在更新-数据比(update-to-data ratio)高达 12 时仍能稳定扩展。在 PQN 中,动作选择和自举均使用批统计量,可以从失效的运行统计量配置中恢复性能。在 26 个 Atari 游戏、4 亿帧的实验中,该配置取得了比采用 LayerNorm 的 PQN 更高的最终聚合得分。我们的结果表明,经过精心配置的 BN 能够显著提升离散动作价值学习的效果,且其前向传播协议是算法规范中不可或缺的一部分。
cs.LG / 69 / 2609.06426

Sparse Oblique Rule Boosting for Simpler Additive Rule Ensembles

用于构建更简洁的加性规则集成的稀疏斜规则提升方法
Behzadimanesh, Shahrzad, Bodic, Pierre Le, Webb, Geoffrey I., Boley, Mario
Abstract
Small additive ensembles of symbolic rules offer interpretable prediction models. Traditionally, these ensembles use rule conditions based on conjunctions of simple threshold propositions $x \geq t$ on a single input variable $x$ and threshold $t$, resulting geometrically in axis-parallel polytopes as decision regions. While this form ensures a high degree of interpretability for individual rules and can be learned efficiently using the gradient boosting approach, it relies on having access to a curated set of expressive input features so that a small ensemble of axis-parallel regions can describe the target variable well. Absent such features, reaching sufficient accuracy requires increasing the number and complexity of individual rules, which diminishes the interpretability of the model. Here, we extend classical rule ensembles by introducing logical propositions with learnable sparse linear transformations of input variables, i.e., propositions of the form $\mathbf{x}^T\mathbf{w} \geq t$, where $\mathbf{w}$ is a learnable sparse weight vector, enabling decision regions as general polyhedrons with oblique faces. We propose a learning method using gradient boosting based on a weighted logistic regression. Empirical results across 14 regression and classification tasks demonstrate that the proposed method achieves lower model complexity than competitive baselines while maintaining similar or better predictive accuracy. Hence, the approach provides a favorable trade-off between interpretability and accuracy and reduces the reliance on manual feature engineering.
Chinese Translation
由少量符号规则组成的加性集成提供了可解释的预测模型。传统上,这类集成使用基于单个输入变量 $x$ 与阈值 $t$ 的简单阈值命题 $x \geq t$ 的合取作为规则条件,其几何上表现为轴平行的多面体作为决策区域。虽然这种形式保证了单个规则的高度可解释性,并且可以通过梯度提升方法高效学习,但它依赖于拥有一组精心设计的、具有强表达能力的输入特征,使得少量轴平行区域即可很好地描述目标变量。在缺乏此类特征的情况下,要达到足够的精度就需要增加个体规则的数量和复杂度,从而削弱模型的可解释性。本文通过引入带可学习稀疏线性变换的逻辑命题来扩展经典规则集成,即形如 $\mathbf{x}^T\mathbf{w} \geq t$ 的命题,其中 $\mathbf{w}$ 是可学习的稀疏权重向量,从而使决策区域可以是具有斜面的任意多面体。我们提出了一种基于加权逻辑回归的梯度提升学习方法。在14个回归和分类任务上的实证结果表明,所提方法在保持相似或更优预测精度的同时,实现了低于竞争基线的模型复杂度。因此,该方法在可解释性与准确性之间提供了良好的折中,并降低了对人工特征工程的依赖。
cs.LG / 70 / 2609.06430

Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks

训练两层量化神经网络的直通估计器的稳定性与泛化性
Ying, Yiming
Abstract
We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Statistical Learning Theory (SLT). Our central question is whether algorithmic stability can explain the statistical generalization of the estimator produced by the discontinuous STE training rule. In the saturated-output regime, the zero-initialized samplewise STE recursion is exactly the stochastic subgradient descent on the convex latent loss $(-yu^\top x)_+$. This representation makes a stability analysis possible. We derive an exact distance identity for two coupled updates and prove approximate non-expansiveness of the common-example map, with a quadratic defect only when the two latent margins straddle zero. We then obtain explicit $\ell_2$ on-average model-stability and generalization bounds, transferring stability isometrically from the latent vector to the full first-layer matrix. Combining stability with a standard optimization bound yields an explicit excess induced-risk guarantee and the rate $O(n^{-1/2})$ when $T=n^2$. Under margin separability, a complementary argument gives the optimal-order $O(R^2/(\gamma^2n))$ expected excess misclassification error for a randomized one-pass STE iterate and a corresponding majority-vote bound.
Chinese Translation
我们从统计学习理论(SLT)的视角研究用于训练带铰链损失的两层二值激活网络的恒等直通估计器(STE)。我们的核心问题是:算法稳定性是否能够解释由不连续的STE训练规则所产生的估计器的统计泛化能力。在饱和输出机制下,零初始化的逐样本STE递推恰好是凸隐损失 $(-yu^ op x)_+$ 上的随机次梯度下降。这一表示使稳定性分析成为可能。我们推导了两个耦合更新之间的精确距离恒等式,并证明了共同样本映射的近似非扩张性,其二次缺陷仅在两个隐边距跨越零时出现。随后,我们得到了显式的 $\ell_2$ 平均模型稳定性界和泛化界,将稳定性以等距方式从隐向量传递到完整的第一层矩阵。将稳定性与标准的优化界相结合,可得到显式的超额诱导风险保证,并在 $T=n^2$ 时达到 $O(n^{-1/2})$ 的速率。在边距可分的条件下,通过补充论证,我们为随机单遍STE迭代给出了最优阶 $O(R^2/(\gamma^2 n))$ 的期望超额误分类误差,以及相应的多数投票界。
cs.LG / 71 / 2609.06460

A Theoretical Framework for Masked Pretraining (MPT)

掩码预训练(MPT)的理论框架
Zhang, Qi, Zhou, Runyu, Wang, Yifei, Wang, Yisen
Abstract
Recently, Masked Pretraining (MPT) based on reconstruction pretraining tasks has risen to a promising self-supervised learning paradigm across various domains and achieves remarkable performance in multiple downstream tasks. However, the theoretical understanding of the working mechanism behind MPT is still limited. In this paper, we introduce a new theoretical framework to analyze MPT and understand the crucial role of masking in extracting meaningful representations. We establish theoretical connections between MPT and another popular self-supervised paradigm: contrastive learning. We prove that the masking technique implicitly creates positive pairs that are semantically similar and the reconstruction loss pulls them together in the feature space. Besides, as a result of the implicit alignment, we point out the dimensional collapse issue of MPT and propose a Uniformity-enhanced MPT (U-MPT) loss that can effectively address this issue and bring significant improvements in downstream tasks including linear evaluation, cross-dataset fine-tuning and out-of-distribution generalization on real-world data sets. Furthermore, we establish downstream guarantees of U-MPT and theoretically analyze the influence of masking strategies. Based on the theoretical analysis, we propose a new masking strategy which enhances the downstream performance of MPT and explains current improvements of masking strategies with our theoretical perspective.
Chinese Translation
近年来,基于重构预训练任务的掩码预训练(Masked Pretraining, MPT)已成为一种极具前景的自监督学习范式,在多个领域中得到广泛应用,并在多种下游任务中取得了卓越的性能。然而,对MPT背后工作机制的理论理解仍然有限。本文提出了一个新的理论框架来分析MPT,并理解掩码机制在提取有意义表示中的关键作用。我们建立了MPT与另一种流行的自监督范式——对比学习——之间的理论联系。我们证明,掩码技术隐式地创建了语义相似的正样本对,而重构损失则在特征空间中将它们拉近。此外,作为这种隐式对齐的结果,我们指出了MPT的维度坍塌(dimensional collapse)问题,并提出了一种增强均匀性(Uniformity)的MPT(U-MPT)损失,该损失能够有效解决这一问题,并在包括线性评估、跨数据集微调以及真实世界数据集上的分布外泛化等下游任务中带来显著提升。此外,我们建立了U-MPT的下游性能保证,并从理论上分析了掩码策略的影响。基于理论分析,我们提出了一种新的掩码策略,提升了MPT的下游性能,并从我们的理论视角解释了现有掩码策略改进的原因。
cs.LG / 72 / 2609.06467

Local and Global Stability in Performative Reinforcement Learning

表演性强化学习中的局部与全局稳定性
Mandal, Debmalya
Abstract
In performative reinforcement learning the deployed policy shapes the environment that generates the learner's future data, and the natural solution concept is a performatively stable policy that is optimal in the environment it induces. Existing convergence guarantees rely on Lipschitz sensitivity assumptions on the environment map $\pi \mapsto (P_\pi, r_\pi)$, which are hard to verify and fail in settings such as multi-agent best-response dynamics. We instead study stability for mixtures of policies, and show that the resulting picture is fundamentally different from performative prediction, where randomization removes the need for any sensitivity assumption. We distinguish local mixed stability, an occupancy-weighted first-order relaxation that we show is equivalent to stationarity, from global mixed stability, which certifies against arbitrary deviating policies. Our first result is that a weighted per-state Hedge dynamic drives the local stability gap to zero at an $O(1/\sqrt{T})$ rate for an arbitrary, possibly discontinuous, environment map, both with exact and with trajectory feedback. The two notions genuinely differ: we exhibit an instance where local stability is achieved exactly but every mixture has global stability gap bounded away from zero. For global stability we introduce a bounded transition range assumption, strictly weaker than Lipschitz sensitivity, under which unweighted per-state Hedge converges up to a floor of $O(\gamma\epsilon_P/(1-\gamma)^3)$, and we prove a matching-in-$\epsilon_P$ lower bound of $\Omega(\gamma\epsilon_P/(1-\gamma))$ under trajectory feedback, so this floor is unavoidable. Finally, we extend both notions to $n$-player performative Markov games, obtaining local stability with no assumption on the joint environment map or game structure, and global stability for performative Markov potential games.
Chinese Translation
在表演性强化学习(performative reinforcement learning)中,部署的策略会塑造产生学习者未来数据的环境,其自然的解概念是表演性稳定策略,即在其所诱导的环境中最优的策略。现有的收敛性保证依赖于对环境映射 $\pi \mapsto (P_\pi, r_\pi)$ 的 Lipschitz 敏感性假设,该假设难以验证,且在多智能体最佳响应动态等场景中会失效。我们转而研究策略混合(mixture of policies)的稳定性,并证明由此得到的图景与表演性预测(performative prediction)有着本质不同——在表演性预测中,随机化可以消除对任何敏感性假设的需要。我们区分了局部混合稳定性(一种以占据度加权的一阶松弛,我们证明其等价于平稳性条件)与全局混合稳定性(可抵御任意偏离策略的认证)。我们的第一个结果表明,一种按状态加权的 Hedge 动态能够以 $O(1/\sqrt{T})$ 的速率将局部稳定性间隙驱至零,且对任意、可能不连续的环境映射均成立,在精确反馈和轨迹反馈下皆然。这两种概念确实存在差异:我们构造了一个实例,其中局部稳定性被精确达到,但每个混合策略的全局稳定性间隙都有非零下界。对于全局稳定性,我们引入了有界转移范围假设,它严格弱于 Lipschitz 敏感性;在该假设下,不加权的按状态 Hedge 动态收敛至 $O(\gamma\epsilon_P/(1-\gamma)^3)$ 的下限,并且在轨迹反馈下我们证明了在 $\epsilon_P$ 上匹配的下界 $\Omega(\gamma\epsilon_P/(1-\gamma))$,因此该下限不可避免。最后,我们将这两个概念扩展到 $n$ 人表演性马尔可夫博弈(performative Markov games),在对联合环境映射或博弈结构不作任何假设的情况下获得了局部稳定性,并为表演性马尔可夫势博弈(performative Markov potential games)获得了全局稳定性。
cs.LG / 73 / 2609.06468

Sector-Mean: Deterministic Initialization of K-Means Centroids via Angular Sector Partitioning

Sector-Mean:基于角扇区划分的K-Means质心确定性初始化方法
Dhakal, Abhiyan, Kafle, Pranish, Chulyadyo, Rajani
Abstract
K-Means is one of the most widely used clustering algorithms, but its susceptibility to initial centroid selection remains a primary bottleneck for its convergence speed and clustering accuracy. This paper proposes Sector-Mean Initialization, a deterministic initialization strategy with O(N) time complexity that partitions the two-dimensional data space into angular sectors around the global centroid and initializes centroids using sector-wise means. We evaluate the method on established two-dimensional benchmarks (SIPU, Birch) and multiple real-world datasets, comparing against random, K-Means++, and Max-Min initialization under identical Lloyd iterations. The statistical analysis of Friedman's test (p<0.05) and Nemenyi post-hoc comparison indicates that, while delivering equivalent clustering quality as K-Means++ and Max-Min, Sector-Mean offers significant computational efficiency. Experimental results show that Sector-Mean reduces the initialization time by 74.9% and 59.8% in comparison to K-Means++ and max-min, respectively. And, it yields the lowest average number of iterations, achieving approximately 5% fewer iterations than K-Means++ and 16% fewer than max-min. These results highlight that Sector-Mean initialization offers a deterministic and computationally efficient initialization strategy while preserving cluster quality.
Chinese Translation
K-Means是最广泛使用的聚类算法之一,但其对初始质心选择的敏感性仍然是制约其收敛速度和聚类精度的主要瓶颈。本文提出Sector-Mean初始化方法,这是一种时间复杂度为O(N)的确定性初始化策略,该方法将二维数据空间围绕全局质心划分为角度扇区,并利用各扇区的均值来初始化质心。我们在经典的二维基准数据集(SIPU、Birch)以及多个真实世界数据集上对该方法进行评估,并在相同的Lloyd迭代次数下与随机初始化、K-Means++和Max-Min初始化方法进行比较。Friedman检验(p<0.05)和Nemenyi事后比较的统计分析表明,Sector-Mean在提供与K-Means++和Max-Min相当聚类质量的同时,具有显著的计算效率优势。实验结果显示,与K-Means++和Max-Min相比,Sector-Mean分别将初始化时间缩短了74.9%和59.8%。此外,它还获得了最低的平均迭代次数,比K-Means++减少约5%的迭代次数,比Max-Min减少约16%。这些结果突出表明,Sector-Mean初始化在保持聚类质量的同时,提供了一种确定性强且计算高效的初始化策略。
cs.LG / 74 / 2609.06469

One Step, One Lead: Mitigating Higher-Order Interference in Multi-Domain Reinforcement Learning via Cross-Step Control

一步一导:通过跨步控制缓解多领域强化学习中的高阶干扰
Lin, Zihan, Wang, Xiaohan, Cao, Jie, Chai, Jiajun, Yin, Guojun, Lin, Wei, He, Ran
Abstract
Reinforcement learning (RL) across multiple domains can broaden the reasoning capabilities of large language models (LLMs), yet joint training often degrades individual-domain performance and can destabilize optimization. Existing work typically diagnoses such interference from a single-step view using first-order gradient alignment or curvature-based proxies. We show that this view can miss a critical form of sequential interference: same-point domain gradients may remain nearly orthogonal even when consecutive realized updates partially reverse one another in output space. We further show that consecutive token log-probability footprints recover this interaction directly from adjacent checkpoints as a local second-order interaction in output space, without explicitly reconstructing same-step curvature. Building on this insight, we propose OSOL, which designates a focus domain at each iteration, uses the preceding checkpoint footprint to rank token-level rebound risk, and applies a drift-ranked, adaptively scaled correction within the standard GRPO update. Our analysis shows that this correction suppresses the targeted cross-step output backtracking component. Controlled studies further show that cross-step backtracking is more strongly associated with subsequent task damage than same-point gradient diagnostics, while the preceding footprint ranks future rebound risk more accurately than Hessian-based proxies. On Qwen3-30B-A3B, OSOL reaches a domain-macro average of 0.4822, improving by 5.7% over the strongest compared baseline, without explicit higher-order differentiation.
Chinese Translation
跨多个领域的强化学习(RL)可以拓展大语言模型(LLM)的推理能力,然而联合训练常常导致单个领域性能下降,并可能使优化过程不稳定。现有工作通常基于一阶梯度对齐或基于曲率的替代指标,从单步视角诊断此类干扰。我们证明,这一视角可能遗漏一种关键的序列性干扰形式:即便相邻实际更新在输出空间中部分相互抵消,同一点上的领域梯度仍可能保持近乎正交。我们进一步证明,相邻词元的对数概率足迹(log-probability footprint)可以直接从相邻检查点恢复这种交互,作为输出空间中的局部二阶交互,而无需显式重构同一步曲率。基于这一洞察,我们提出了OSOL:该方法在每次迭代中指定一个重点领域,利用前一个检查点的足迹对词元级的回弹风险进行排序,并在标准GRPO更新中施加基于漂移排序、自适应缩放的校正。我们的分析表明,该校正能够抑制目标跨步输出回退分量。受控研究进一步表明,与同点梯度诊断相比,跨步回退与后续任务损害的关联更强;同时,前一个检查点的足迹在排序未来回弹风险方面比基于Hessian的替代指标更准确。在Qwen3-30B-A3B上,OSOL达到了0.4822的领域宏平均分,相比最强的对比基线提升了5.7%,且无需显式的高阶微分。
cs.LG / 75 / 2609.06473

Steering Under Compression: Dose-Response, Capability Cost, and Failure Asymmetry in Quantized LLMs

压缩下的引导:量化大语言模型中的剂量-反应关系、能力代价与失败不对称性
Bhandari, Saurav, Wade, Benjamin
Abstract
Inference-time activation steering enables behavioral control of large language models without parameter modification, while post-training quantization reduces memory and compute costs for deployment. Despite their growing convergence in practice, the interaction between these two techniques remains uncharacterized. We systematically study activation steering under weight-only quantization (INT8 and NF4) across four open-weight 7-9B models and two behavioral targets: judged sentiment and judge-free reasoning length. Using an iso-effect framework that compares capability costs at matched behavioral effect, we find that sentiment steering survives quantization intact. After correcting a GSM8K parser artifact with a uniform v2.3.1 rescore, the pooled INT8 contrast is -0.010 (90% CI [-0.026, +0.007]), descriptively Equivalent under the preregistered three-label rule, while NF4 remains Inconclusive at -0.017 ([-0.067, +0.033]). In contrast, reasoning length exhibits a surprising asymmetric dose-response: lengthening is graded but terminates in cap-runaway and collapse, while shortening is a step function with only 12-30% shortening (model-dependent) before discontinuous failure. We expose a methodological pitfall: the naive iso-effect ladder anchors on the collapse floor for floor-bounded targets, and we introduce a censored construction that restores interpretable crossings. We also quantify a substantial baseline capability shift for Mistral-NF4 (0.545 to 0.365 GSM8K at alpha=0), demonstrating that compression can dominate the steering intervention. Despite this, steering vectors remain highly collinear with their FP16 siblings (cosine similarity 0.989-0.998 for INT8, 0.945-0.990 for NF4), confirming that the behavioral direction survives quantization even when the cost structure does not. All code and data are released.
Chinese Translation
推理时激活引导(activation steering)无需修改参数即可实现对大语言模型行为的控制,而训练后量化(post-training quantization)则降低了部署时的内存与计算成本。尽管这两种技术在实践中日益融合,但它们之间的相互作用尚未得到系统刻画。我们系统地研究了仅权重量化(weight-only quantization,INT8 和 NF4)条件下的激活引导,涵盖四个开源权重 7-9B 模型以及两类行为目标:有评判者的情感(judged sentiment)和无评判者的推理长度(judge-free reasoning length)。我们采用等效应(iso-effect)框架,在匹配行为效应的条件下比较能力代价,发现情感引导在量化后得以完整保留。在通过统一的 v2.3.1 重新评分修正 GSM8K 解析器缺陷后,INT8 的合并对比值为 -0.010(90% 置信区间 [-0.026, +0.007]),根据预注册的三标签规则在描述上为等效(Equivalent),而 NF4 在 -0.017([-0.067, +0.033])处仍为不确定(Inconclusive)。相比之下,推理长度表现出惊人的不对称剂量-反应关系:延长行为是渐进式的,但最终会终止于失控增长(cap-runaway)和崩溃;而缩短行为则表现为阶跃函数,在不连续失败发生之前仅能缩短 12-30%(因模型而异)。我们揭示了一个方法论陷阱:对于以下限为界的目标,朴素的等效应阶梯会锚定在崩溃下限上,并引入了一种删失构造方法以恢复可解释的交叉点。我们还量化了 Mistral-NF4 在基线能力上的显著漂移(alpha=0 时 GSM8K 从 0.545 降至 0.365),表明压缩本身可能主导引导干预的效果。尽管如此,引导向量仍与其 FP16 对应向量高度共线(INT8 的余弦相似度为 0.989-0.998,NF4 为 0.945-0.990),证实即使在代价结构改变的情况下,行为方向仍能在量化中得以保留。所有代码和数据均已公开。
cs.LG / 76 / 2609.06474

Learning Kernels by Alignment for Multiclass Bayes Classification

通过对齐学习核函数的多类贝叶斯分类
Haule, Hollan, Gonzalez-Sulser, Alfredo, Escudero, Javier
Abstract
Kernel methods separate data representation from decision-making, but typically require the kernel to be chosen in advance. We show that this kernel can instead be learned by alignment, and develop the resulting framework through the recently introduced Collaborative Learning and Inference (CLaI). We show that Collaborative Learning can be viewed as a kernel alignment process, in which an embedding is trained so that its induced similarity matches a label-derived target kernel. We also prove that Collaborative Inference is equivalent to kernel Bayes classification with Parzen-window density estimation. Motivated by these perspectives, we generalise CLaI by replacing cosine similarity with a learned Mahalanobis distance and extend it to multiclass classification. On CIFAR-10, PathMNIST, and SleepEDF, the Mahalanobis formulation improves accuracy, converges faster, and yields lower calibration error than the cosine-based variant. Auxiliary experiments further support these connections, showing that CLaI produces latent signals of the same form as a Gaussian process, while achieving competitive calibration on sepsis prediction. Together, these results establish a principled learned-kernel framework that unifies representation learning, kernel alignment, and Bayesian classification, and extends naturally to the multiclass setting.
Chinese Translation
核方法将数据表示与决策制定相分离,但通常需要事先选定核函数。我们证明该核函数可以通过对齐来学习,并通过最近提出的协同学习与推理(Collaborative Learning and Inference, CLaI)框架发展出这一方法。我们证明协同学习可被视为一种核对齐过程,即训练一个嵌入,使其诱导的相似性与由标签导出的目标核相匹配。我们还证明协同推理等价于采用Parzen窗密度估计的核贝叶斯分类。基于这些视角,我们通过将余弦相似度替换为学习得到的马氏距离(Mahalanobis distance)对CLaI进行泛化,并将其扩展至多类分类。在CIFAR-10、PathMNIST和SleepEDF数据集上,马氏距离公式相比基于余弦相似度的变体提高了准确率、收敛更快,并产生了更低的校准误差。辅助实验进一步支持了这些联系,表明CLaI生成的潜在信号与高斯过程具有相同的形式,同时在脓毒症预测任务上取得了具有竞争力的校准性能。综上,这些结果建立了一个有原理支撑的可学习核框架,统一了表示学习、核对齐和贝叶斯分类,并可自然地扩展到多类设置。
cs.LG / 77 / 2609.06483

How Does Parameter Pruning Reshape DNN Representations? An Interaction-Driven Exploration

参数剪枝如何重塑深度神经网络(DNN)的表示?一种基于交互作用的探索
Li, Fangbo, Zhang, Junpeng, Ren, Qihan, Zhang, Quanshi
Abstract
This study focuses on the scientific problem of understanding internal factors that govern the diverse performance degradation of deep neural networks (DNNs) when different parameters are pruned. In order to explain why pruning certain parameters leads to significant performance degradation but pruning other parameters does not, we examine how the pruning operation affects the interaction patterns encoded by the DNN. We find that when we progressively increase the pruning ratio, the interaction patterns encoded by DNNs exhibit a distinct three-phase dynamics, \emph{i.e.}, model performance is not largely affected until the pruning operation begins to remove low-order interactions, and low-order interactions exhibit strong generalizability. Moreover, we find that the high sensitivity of DNN performance to the pruning of certain modules is attributed to whether the pruning operation removes generalizable low-order interaction patterns.
Chinese Translation
本研究聚焦于理解控制深度神经网络(DNN)在剪除不同参数时产生不同性能退化程度的内部因素这一科学问题。为了解释为什么剪除某些参数会导致显著的性能退化,而剪除另一些参数则不会,我们考察了剪枝操作如何影响DNN所编码的交互模式。我们发现,当逐渐提高剪枝比例时,DNN编码的交互模式呈现出明显的三阶段动态特性,即:模型性能不会受到较大影响,直到剪枝操作开始移除低阶交互为止,而低阶交互表现出很强的泛化能力。此外,我们发现DNN性能对某些模块的剪枝高度敏感,其原因在于剪枝操作是否移除了可泛化的低阶交互模式。
cs.LG / 78 / 2609.06484

Second-Order Smooth Planning with Optimal-Transport Bellman Smoothing

基于最优传输贝尔曼平滑的二阶平滑规划
Dam, Tuan
Abstract
Planning with a generative model aims to estimate the value of a state using as few simulator calls as possible. SmoothCruiser achieves problem-independent complexity $\widetilde O(\varepsilon^{-4})$ by exploiting the smoothness of the entropy-regularized Bellman backup, but its estimator is only first-order. We show that the sample-complexity exponent of SmoothCruiser-type planners is governed by the order $\beta$ of the local Taylor remainder, giving oracle complexity $\widetilde O(\varepsilon^{-(2+2/(\beta-1))})$: the first-order case $\beta=2$ recovers SmoothCruiser, while a second-order/cubic remainder $\beta=3$ yields $\widetilde O(\varepsilon^{-3})$. We reach this regime with an optimal-transport-smoothed Bellman backup over action distributions, which has a closed form, a policy gradient, and a Lipschitz Hessian, and whose quadratic correction admits an unbiased cross-product estimator. The resulting SecondOrderSmoothCruiser achieves $\widetilde O(\varepsilon^{-3})$ oracle complexity for fixed OT parameters, and we relate the OT, entropy-regularized, and unregularized objectives through explicit regularization-bias bounds.
Chinese Translation
基于生成模型的规划旨在使用尽可能少的模拟器调用来估计某个状态的价值。SmoothCruiser 通过利用熵正则化贝尔曼备份(entropy-regularized Bellman backup)的平滑性,实现了与问题无关的复杂度 $\widetilde O(\varepsilon^{-4})$,但其估计量仅为一阶。我们证明,SmoothCruiser 类规划器的样本复杂度指数由局部泰勒余项的阶数 $\beta$ 决定,其预言复杂度(oracle complexity)为 $\widetilde O(\varepsilon^{-(2+2/(\beta-1))})$:一阶情形 $\beta=2$ 可恢复 SmoothCruiser 的结果,而二阶/三次余项 $\beta=3$ 则可得到 $\widetilde O(\varepsilon^{-3})$。我们通过在动作分布上引入最优传输平滑的贝尔曼备份来达到这一区间,该备份具有闭式解、策略梯度以及 Lipschitz 连续的 Hessian 矩阵,并且其二次修正项存在无偏的叉积估计量。所得的 SecondOrderSmoothCruiser 在固定的最优传输参数下实现了 $\widetilde O(\varepsilon^{-3})$ 的预言复杂度,并且我们通过显式的正则化偏差界将最优传输目标、熵正则化目标与未正则化目标联系起来。
cs.LG / 79 / 2609.06489

Power Mean Estimation in Stochastic Continuous Monte Carlo Tree Search

随机连续蒙特卡洛树搜索中的幂均值估计
Dam, Tuan
Abstract
Monte Carlo Tree Search (MCTS) has demonstrated success in online planning for deterministic environments, yet significant challenges remain in adapting it to stochastic Markov Decision Processes (MDPs), particularly in continuous state-action spaces. Existing methods, such as HOOT, which combines MCTS with the Hierarchical Optimistic Optimization (HOO) bandit strategy, address continuous spaces but rely on a logarithmic exploration bonus that lacks theoretical guarantees in non-stationary, stochastic settings. Recent advancements, such as POLY-HOOT, introduced a polynomial bonus term to achieve convergence in deterministic MDPs, though a similar theory for stochastic MDPs remains undeveloped. In this paper, we propose a novel MCTS algorithm, \Algname, designed for continuous, stochastic MDPs. \Algname integrates a power mean as a value backup operator, alongside a polynomial exploration bonus to address the non-stationarity inherent in continuous action spaces. Our theoretical analysis establishes that \Algname converges at a polynomial rate of $\mathcal{O}(n^{-\zeta})$, $\zeta \in (0,1/2)$, where \( n \) is the number of visited trajectories, thereby extending the non-asymptotic convergence guarantees of POLY-HOOT to stochastic environments. Experimental results on stochastic tasks validate our theoretical findings, demonstrating the effectiveness of \Algname in continuous, stochastic domains.
Chinese Translation
蒙特卡洛树搜索(MCTS)已在确定性环境的在线规划中取得成功,但将其应用于随机马尔可夫决策过程(MDP)仍面临重大挑战,尤其是在连续状态-动作空间中。现有方法(如结合MCTS与分层乐观优化(HOO)赌博机策略的HOOT)虽然能够处理连续空间,但其依赖的对数探索奖励在非平稳、随机环境中缺乏理论保证。近期的研究进展(如POLY-HOOT)引入了多项式奖励项以在确定性MDP中实现收敛,但针对随机MDP的类似理论尚未建立。在本文中,我们提出了一种新颖的MCTS算法\Algname,专为连续随机MDP设计。\Algname将幂均值作为价值回传算子,并结合多项式探索奖励来应对连续动作空间中固有的非平稳性。我们的理论分析表明,\Algname以多项式速率$\mathcal{O}(n^{-\zeta})$($\zeta \in (0,1/2)$)收敛,其中$n$为已访问轨迹的数量,从而将POLY-HOOT的非渐近收敛保证扩展到了随机环境。在随机任务上的实验结果验证了我们的理论发现,证明了\Algname在连续随机域中的有效性。
cs.LG / 80 / 2609.06499

Structural Entropy-Driven Graph Diffusion Generation for One-Shot Federated Graph Learning

面向单轮联邦图学习的结构熵驱动图扩散生成方法
Zheng, Shutong, Fu, Lele, Huang, Sheng, Lim, Wei Yang Bryan, Chen, Chuan
Abstract
One-shot federated graph learning (FGL) requires the server to estimate client contributions from highly compressed information, yet conventional volume-based weighting captures the amount of client data while overlooking how its connectivity is organized. In this paper, we propose SPIRE, a Structural Entropy-Driven Graph Diffusion Generation method that introduces topology-aware client differentiation into one-shot FGL. Specifically, we employ first-order degree-distribution structural entropy as a compact descriptor of degree-mass dispersion and use it to derive structural client weights, providing an inductive bias that accounts for differences in graph topology beyond data volume. On the generation side, a graph diffusion model on the server synthesizes pseudographs conditioned on the weighted client prototypes, capturing both semantic and structural information without requiring additional client-side training. The generated pseudographs are then assembled via disjoint union fusion to train a global graph neural network. Extensive experiments on seven real-world graph datasets demonstrate that SPIRE consistently outperforms conventional and one-shot FGL methods, with particularly strong gains under highly heterogeneous (non-IID) and graph-perturbed settings.
Chinese Translation
单轮联邦图学习(One-shot Federated Graph Learning, FGL)要求服务器从高度压缩的信息中估计客户端的贡献,然而传统的基于数据量(volume-based)的加权方式仅捕捉客户端数据量的大小,而忽略了其连接结构的组织方式。本文提出SPIRE,一种结构熵驱动的图扩散生成方法,将拓扑感知的客户端差异化引入单轮联邦图学习。具体而言,我们采用一阶度分布结构熵作为度质量离散度的紧凑描述符,并据此推导结构性客户端权重,从而提供一种能够刻画数据量之外的图拓扑差异的归纳偏置。在生成方面,服务器端的图扩散模型以加权客户端原型为条件合成伪图,在无需客户端额外训练的情况下同时捕捉语义信息和结构信息。生成的伪图随后通过不相交并集融合进行汇总,用于训练全局图神经网络。在七个真实世界图数据集上的大量实验表明,SPIRE持续优于传统方法和单轮联邦图学习方法,在高异构(非IID)和图扰动场景下尤为显著。
cs.LG / 81 / 2609.06511

Bi-HYCO: Bi-Objective Cooperative Learning for PDE Parameter Identification under Fragmented Observations

Bi-HYCO:碎片化观测下偏微分方程参数识别的双目标协同学习方法
Biccari, Umberto, Chen, Jun, Morales, Roberto, Zuazua, Enrique
Abstract
Physical and synthetic models may describe complementary aspects of the same PDE-governed system while receiving different, possibly fragmented, observations. We propose Bi-Objective HYCO (Bi-HYCO), a cooperative framework that retains both representations and their local observational objectives while coupling their predicted states at unlabeled interaction points. These points contain no measurements and do not augment the data; they provide a communication mechanism in the common state space. The two criteria form a vector-valued objective, and weighted scalarizations provide computational realizations. For the deterministic shared-observation algorithm with fixed interaction points, we prove sufficient decrease and finite length of the whole alternating sequence, which converges to a mixed critical point under the stated Kurdyka-Lojasiewicz-type assumptions. Elliptic transmission and two-dimensional Navier-Stokes experiments assess parameter and state reconstruction, noise and scalarization effects, and PINN/XPINN references. Ablations show that removing state interaction while retaining aggregation deteriorates parameter recovery in the tested configurations, particularly for Navier-Stokes.
Chinese Translation
物理模型与合成模型可能从互补的角度描述同一个由偏微分方程(PDE)支配的系统,同时接收不同的、可能是碎片化的观测数据。我们提出了双目标HYCO(Bi-HYCO),这是一个协同学习框架,它在保留两种表示及其各自的局部观测目标的同时,在无标签的交互点上耦合二者预测的状态。这些交互点不包含任何测量数据,也不会增加数据量;它们仅提供公共状态空间中的一种通信机制。这两个准则构成一个向量值目标函数,加权标量化则提供了其计算实现。对于具有固定交互点的确定性共享观测算法,我们证明了整个交替序列的充分下降性和有限长度性,并在所述的Kurdyka-Lojasiewicz型假设下,证明该序列收敛到一个混合临界点。椭圆 transmission 问题和二维Navier-Stokes实验评估了参数与状态重构效果、噪声和标量化的影响,并与PINN/XPINN基准方法进行了对比。消融实验表明,在所测试的配置中,移除状态交互而仅保留聚合会恶化参数恢复性能,尤其是对于Navier-Stokes问题。
cs.LG / 82 / 2609.06514

Model-Adaptive and Risk-Constrained Frequency Hopping Against Predictive Jammers

面向预测性干扰机的模型自适应与风险约束跳频方法
Chen, Yanbo, Zhou, Xinjing
Abstract
Adaptive frequency hopping against predictive jamming must address both model uncertainty and policy exposure: the context-loss relationship may vary across operating regimes, while persistent hopping patterns may expose high-probability channels to attack. We propose D-PACT-AFH, a model-adaptive and risk-constrained adversarial contextual-bandit framework in which a Tsallis-FTRL master combines a global linear learner with a partitioned local learner and selects the model class online. D-PACT-Hit incorporates channel-wise marginal hit risk into model selection, while D-PACT-Safe applies a minimum-Kullback-Leibler projection to enforce a per-slot risk budget. We establish estimator validity under non-anticipating attacks, an oracle decomposition relative to the better fixed base, and an exact conditional-risk guarantee for the Safe projection. Experiments across diverse channel regimes and jammer types demonstrate effective model adaptation and a controllable goodput-risk tradeoff: D-PACT-AFH recovers 95.5% of the local learner's gain under observable switching while avoiding 77.7% of its degradation in a negative-control regime.
Chinese Translation
针对预测性干扰的自适应跳频必须同时应对模型不确定性与策略暴露问题:上下文与损失之间的关系可能随工作状态而变化,而持续不变的跳频模式则可能使高概率信道遭受攻击。我们提出了 D-PACT-AFH,一种模型自适应且风险受限的对抗性上下文多臂老虎机框架,其中 Tsallis-FTRL 主算法将全局线性学习器与分区的局部学习器相结合,并在线选择模型类别。D-PACT-Hit 将逐信道的边际命中风险纳入模型选择,而 D-PACT-Safe 则应用最小 KL 散度投影来强制执行每时隙的风险预算。我们在非预知攻击条件下建立了估计量的有效性、相对于更优固定基学习器的 oracle 分解,以及 Safe 投影的精确条件风险保证。在多种信道状态和干扰机类型上的实验表明,该框架实现了有效的模型自适应以及可控的有效吞吐量-风险权衡:在可观测切换条件下,D-PACT-AFH 恢复了局部学习器增益的 95.5%,同时在负对照状态下避免了其 77.7% 的性能退化。
cs.LG / 83 / 2609.06519

Role-Specific Predictive Geometries for Nonstationary Multivariate Graph-Signal Forecasting

面向非平稳多元图信号预测的角色特定预测几何结构
Chen, Yanbo, Makur, Anamitra
Abstract
Forecasting multivariate graph signals is challenging when node-level trajectories are nonstationary but stable relations persist across nodes and features. In an error-correction representation, long-run equilibrium restoration and short-run transient propagation represent different predictive roles and need not share a common cross-feature geometry. We introduce role-specific predictive geometries in which directed Long relations act on estimated equilibrium coordinates, whereas directed Short relations act on lagged differences. Matrix-valued Long responses mix equilibrium coordinates before graph propagation, while Short responses use graph-filtered transient designs; a direct multi-horizon estimator couples forecast corrections across adjacent horizons. Temporal cross-fitting and Frisch-Waugh-Lovell partialling-out give selected edges a conditional predictive interpretation relative to a graph-temporal backbone. The Long operator remains right-factorized through the equilibrium subspace and therefore annihilates source common-trend directions. Controlled experiments recover all planted Long relations (20/20), all planted Short relations (20/20), and both role families in every Dual realization (10/10). Across four real-world benchmarks, the proposed predictor improves on the G-VARMA backbone in three datasets, with all 25 fold-horizon comparisons favorable on the five-fold financial benchmark.
Chinese Translation
当节点级轨迹非平稳但节点与特征之间的稳定关系持续存在时,多元图信号的预测极具挑战性。在误差修正表示中,长期均衡恢复与短期瞬态传播代表不同的预测角色,无需共享共同的跨特征几何结构。我们引入角色特定的预测几何结构:有向的Long关系作用于估计的均衡坐标,而有向的Short关系作用于滞后差分。矩阵形式的Long响应在图传播之前混合均衡坐标,而Short响应采用图滤波的瞬态设计;直接多步预测估计器将相邻预测视界之间的预测修正耦合起来。时间交叉拟合与Frisch-Waugh-Lovell偏回归方法使所选边相对于图-时间主干获得条件预测解释。Long算子通过均衡子空间保持右因子化,从而能够消除源的共同趋势方向。在受控实验中,所有植入的Long关系(20/20)、所有植入的Short关系(20/20)以及每次Dual实现中的两类角色关系(10/10)均被完整恢复。在四个真实世界基准数据集上,所提出的预测器在三个数据集上优于G-VARMA主干模型,且在五折金融基准上的所有25个折-视界比较中均表现更优。
cs.LG / 84 / 2609.06521

Not Just Oversmoothing: Detecting the Echo Chamber Effect in Graph Neural Networks

不仅仅是过平滑:检测图神经网络中的回音室效应
Hevapathige, Asela, Zehmakan, Ahad N., Wijesinghe, Asiri, Halgamuge, Saman
Abstract
Oversmoothing is a well-known failure mode of Graph Neural Networks (GNNs). However, most existing diagnostics rely on global aggregation measures that fail to capture the heterogeneous dynamics of message passing. Real-world graphs exhibit pronounced community structure, and message passing operates on two timescales, with representations collapsing rapidly within communities and slowly across them. This creates a critical gap in which intra-community representations can become indistinguishable while inter community separation persists, a failure mode that we refer to as the Echo Chamber Effect. To quantify this effect, we introduce the Echo Chamber Index (ECI), which stratifies pairwise distances by community membership and reveals when global energy diminishes while inter-community separation persists. ECI further shows that feature retention mechanisms can preserve the echo chamber under the conditions of our theoretical analysis. The consequences depend on label structure: when communities align with classes, the echo chamber can sharpen node classification, whereas when they do not, the same collapse makes classification provably harder. Motivated by this analysis, we propose Community-Aware Split Propagation (CASP), a lightweight plugin that decouples intra- and inter-community aggregation and learns their balance from label structure. CASP improves diverse backbone GNNs across most evaluated homophilic and heterophilic settings.
Chinese Translation
过平滑是图神经网络(GNN)的一种众所周知的失效模式。然而,现有的诊断方法大多依赖全局聚合度量,无法捕捉消息传递的异质性动态。真实世界的图具有显著的社区结构,且消息传递在两个时间尺度上进行:表示在社区内部迅速坍缩,而在社区之间则坍缩缓慢。这造成了一个关键缺口:社区内部的表示可能变得不可区分,而社区之间的分离却依然存在,我们将这种失效模式称为回音室效应(Echo Chamber Effect)。为了量化这一效应,我们提出了回音室指数(Echo Chamber Index, ECI),该指标按社区成员身份对成对距离进行分层,从而揭示全局能量衰减而社区间分离仍然存在的情况。ECI进一步表明,在我们的理论分析条件下,特征保留机制可能会维持回音室效应。其后果取决于标签结构:当社区与类别一致时,回音室效应可以增强节点分类性能;而当二者不一致时,同样的坍缩会使分类可证明地变得更加困难。基于这一分析,我们提出了社区感知分裂传播(Community-Aware Split Propagation, CASP),这是一个轻量级插件,它将社区内部与社区间的聚合解耦,并根据标签结构学习二者的平衡。CASP在大多数评估的同配和异配设置中提升了多种骨干GNN的性能。
cs.LG / 85 / 2609.06557

Hidden in Plain Sight: The Overlooked Significance of Canonical Elements for Extreme LLM Sparsity

显而易见却被忽视:规范元素对大语言模型极端稀疏化的重要意义
Jang, Hyeondo, Lee, Kwanhee, Lee, Dongyeop, Lee, Namhoon
Abstract
Large language models (LLMs) are often considered fragile under aggressive sparsification, and maintaining reliable performance typically requires sticking to moderate sparsity levels. However, recent studies suggest that LLMs are more resilient to high sparsity than previously thought, reframing the problem as a design challenge rather than a fundamental limitation. In this work, we challenge the perceived limits of unstructured post-training LLM pruning by revisiting elementary pruning strategies that have remained relatively underexplored at this scale. Through a progressive sparsification framework with second-order saliency and continued training coordinated with sparsity progression, we show that pretrained LLMs can retain strong performance far beyond commonly studied sparsity regimes. Across LLaMA-2 and Qwen-3 model families, our approach improves perplexity and downstream accuracy up to 99\% sparsity, surpassing both the current state-of-the-art and representative baselines. Precisely, on LLaMA-2-7B, our approach achieves WikiText-2 perplexities of 13.48 and 19.67 at 95\% and 99\% sparsity, respectively, while delivering 3.23$\times$ decoding speedup and 6.21$\times$ memory savings at 95\% sparsity. Taken together, our results show that LLMs can be pushed into extreme sparsity while retaining strong performance, providing a foundation for further improving sparse models in this regime.
Chinese Translation
大语言模型(LLM)通常被认为在激进稀疏化下十分脆弱,要维持可靠的性能通常只能停留在中等稀疏水平。然而,近期研究表明,LLM 对高稀疏度的韧性比以往认知的更强,这将该问题重新定义为一个设计挑战而非根本性局限。在本工作中,我们重新审视了在该规模下相对未被充分探索的基础剪枝策略,以此挑战非结构化训练后 LLM 剪枝的既有极限。通过一个结合二阶显著性度量的渐进式稀疏化框架,并与稀疏度推进相协调的持续训练,我们证明预训练 LLM 能够在远超常见研究范围的稀疏度下保持强劲性能。在 LLaMA-2 和 Qwen-3 模型系列上,我们的方法在高达 99% 的稀疏度下提升了困惑度和下游任务准确率,超越了当前最先进方法和代表性基线。具体而言,在 LLaMA-2-7B 上,我们的方法在 95% 和 99% 稀疏度下分别取得了 13.48 和 19.67 的 WikiText-2 困惑度,并在 95% 稀疏度下实现了 3.23 倍的解码加速和 6.21 倍的内存节省。综上所述,我们的结果表明 LLM 可以被推向极端稀疏度同时保持强劲性能,为进一步提升该稀疏机制下的稀疏模型奠定了基础。
cs.LG / 86 / 2609.06590

Layer-Wise Gate-Controlled Prompt Truncation in a Multimodal Chest X-Ray Classifier

多模态胸部X光分类器中的逐层门控提示截断
Lei, Jingtao, Li, Hongji, Shu, Dexiang
Abstract
Mixture of Prompt Experts (MoPE) adapts multimodal transformers through input-dependent prompt composition, while retaining a fixed prompt length. We investigate a layer-wise gating extension in a binary chest X-ray classification pilot study. The controller predicts a retention ratio for each sample, averages these ratios within a mini-batch, and uses the resulting integer length to truncate the static and mixed visual prompts. Retained mixed prompts are also scaled by the individual ratios. In one recorded run per configuration, the gated model reached a best validation accuracy of 0.8996, compared with 0.8969 for the fixed-length baseline; the corresponding final values were 0.8963 and 0.8802. The exported gate statistics imply a retained length of one at all recorded training points, relative to a configured maximum of six. This reduces the complete visual sequence from 210 to 200 tokens, but no direct runtime measurements establish an acceleration benefit. Report-derived labels, report text as input, sequential data partitioning, and the absence of repeated controlled experiments limit interpretation. The findings document prompt shortening under the configured gate penalty; they do not establish sample-specific length allocation, superiority over fixed short prompts, or clinical utility. Code is available at: https://github.com/jingtaolei/mope-dynamic-prompt-truncation.
Chinese Translation
混合提示专家方法(Mixture of Prompt Experts,MoPE)通过依赖于输入的提示组合来适配多模态Transformer,同时保持固定的提示长度。我们在一项二分类胸部X光分类的试点研究中研究了一种逐层门控扩展方法。控制器为每个样本预测一个保留比例,在小批次内对这些比例取平均,并使用所得的整数长度来截断静态和混合视觉提示。被保留的混合提示还会按各自的比例进行缩放。在每个配置的一次记录运行中,门控模型达到了0.8996的最佳验证准确率,相比之下固定长度基线为0.8969;对应的最终值分别为0.8963和0.8802。导出的门控统计数据显示,相对于配置的最大长度六,在所有记录的训练点上保留长度均为一。这将完整的视觉序列从210个词元减少到200个词元,但没有直接的运行时间测量来证明其加速收益。基于报告生成的标签、以报告文本作为输入、顺序的数据划分以及缺乏重复对照实验等因素限制了结果的可解释性。本研究结果记录了在配置的门控惩罚下提示被缩短的现象;但并未确立样本特定的长度分配、相对于固定短提示的优势,或临床实用性。代码可在以下网址获取:https://github.com/jingtaolei/mope-dynamic-prompt-truncation。
cs.LG / 87 / 2609.06598

Deep Barycentric Regression for Optimal Transport Map Estimation and its Statistical Optimality

用于最优传输映射估计的深度重心回归及其统计最优性
Kim, Kunwoong, Kong, Insung, Kim, Yongdai
Abstract
The optimal transport (OT) map provides a geometric transformation for aligning probability distributions and has become a useful tool in machine learning. However, existing estimators of the OT map still exhibit a gap between sharp statistical guarantees and practical parametric estimation based on stable training objectives. Theoretical estimators achieve minimax optimal convergence rates, but they are typically nonparametric and can incur demanding implementation design or inference costs. Practical estimators are parametric and scalable, but their statistical guarantees remain underexplored, and their min-max, adversarial-like training objectives can be sensitive to optimization algorithms. We propose BROT (Barycentric Regression for OT), a simple two-step method that first computes the unregularized OT plan and then fits a deep neural network (DNN) to the induced barycentric targets by least-squares regression. Under standard regularity conditions, we prove that the DNN estimator of BROT attains the minimax convergence rate, when the ground-truth OT map is Lipschitz. Numerical studies on synthetic datasets and an image dataset show that BROT provides accurate map estimates, strong target distribution matching, and competitive transport costs, compared to existing estimation methods. Experiments on two downstream tasks, single-cell perturbation prediction and unsupervised domain adaptation, further suggest that the accurate estimation of BROT can translate into stronger task performance.
Chinese Translation
最优传输(OT)映射为对齐概率分布提供了一种几何变换,并已成为机器学习中的重要工具。然而,现有的OT映射估计器在严格的统计保证与基于稳定训练目标的实际参数化估计之间仍存在差距。理论性估计器能够达到极小化极大(minimax)最优收敛速率,但它们通常是非参数化的,可能需要复杂的实现设计或高昂的推断成本。实际估计器是参数化且可扩展的,但其统计保证仍未被充分研究,且其极小极大、类似对抗性的训练目标可能对优化算法较为敏感。我们提出了BROT(Barycentric Regression for OT,用于最优传输的重心回归),这是一种简单的两步方法:首先计算未正则化的OT传输方案,然后通过最小二乘回归用深度神经网络(DNN)拟合由此导出的重心目标。在标准的正则性条件下,我们证明了当真实OT映射满足Lipschitz条件时,BROT的DNN估计器能够达到极小化极大收敛速率。在合成数据集和图像数据集上的数值研究表明,与现有估计方法相比,BROT提供了精确的映射估计、良好的目标分布匹配以及具有竞争力的传输成本。在单细胞扰动预测和无监督域自适应两个下游任务上的实验进一步表明,BROT的精确估计能够转化为更强的任务性能。
cs.LG / 88 / 2609.06610

A Statistical and Machine Learning Framework for Quantifying Offensive Impact in Professional Box Lacrosse

用于量化职业场地曲棍球(室内长曲棍球)进攻影响的统计与机器学习框架
Jimerson Jr, Robert
Abstract
Professional box-lacrosse statistics summarize outcomes but provide limited information about shot quality or the roles behind scoring opportunities. This study develops a documented framework for estimating expected goals (xG) and attributing recorded offensive involvement using 1,006 manually annotated Rochester Knighthawks shot attempts, including 151 goals, from 13 consecutive 2025-2026 National Lacrosse League games. Logistic regression, random forest, and extremely randomized trees were evaluated across three nested feature sets using Leave-One-Game-Out cross-validation and a training-fold base-rate benchmark. The contextual baseline random forest had the lowest observed pooled log loss (0.4189) and Brier score (0.1260), improving on the benchmark by 1.22% and 1.50%; five of nine specifications did not beat the benchmark. Adding two-man-action and pick-type fields did not improve the primary metrics. Core Offensive Impact attributes recorded involvement through shooter xG and shot-based expected assists for final passers. Expected Pick Value (xPV) compares a qualifying pick's observed-state probability with a no-pick counterfactual. Its magnitude was indistinguishable from model noise. Its directional pattern exceeded 200 row-permutation replicates, but limited tail resolution and failure to preserve game-level pick composition make the diagnostic descriptive rather than inferential. Accordingly, xPV is reported only as an exploratory augmented component. Given the single-team, 13-game sample, the results are an initial case study rather than league-wide or causal estimates.
Chinese Translation
职业室内长曲棍球(box lacrosse)的数据统计总结了比赛结果,但对射门质量以及得分机会背后的角色所提供的信息有限。本研究基于1,006次经人工标注的罗切斯特骑士鹰队(Rochester Knighthawks)射门尝试(其中包含151个进球),这些数据来自2025-2026赛季国家长曲棍球联盟(National Lacrosse League)连续13场比赛,建立了一个有完整文档记录的框架,用于估计期望进球数(xG)并归因于已记录的进攻参与。研究采用留一场比赛交叉验证(Leave-One-Game-Out cross-validation)和训练折基准比率作为参照,在三个嵌套特征集上评估了逻辑回归、随机森林和极端随机树模型。包含情境特征基线的随机森林具有最低的观测合并对数损失(0.4189)和Brier分数(0.1260),相较基准分别提升了1.22%和1.50%;九种模型设定中有五种未能超越基准。加入二人配合动作和掩护类型字段并未改善主要指标。核心进攻影响指标通过射门者xG以及面向最后一传者的基于射门的期望助攻来归因于已记录的参与。期望掩护价值(Expected Pick Value, xPV)将符合条件掩护的观测状态概率与无掩护的反事实进行比较。其数值大小与模型噪声无法区分;其方向性模式超过了200次行置换重复检验,但由于尾部分辨率有限且未能保持比赛层面的掩护构成,该诊断指标仅具描述性而非推断性。因此,xPV仅作为探索性增补组件予以报告。鉴于样本仅来自单一球队的13场比赛,本研究结果是一项初步案例研究,而非联盟范围的或因果性的估计。
cs.LG / 89 / 2609.06649

Inducing Emergent Misalignment from Reward Hacks with Iterative DPO

利用迭代DPO从奖励作弊中诱导涌现性失调
Daniels, Oliver, Moodley, Perusha, Marlin, Benjamin M., Lindner, David
Abstract
Reward hacking during reinforcement learning from verifiable rewards (RLVR) can induce reward seeking and broad misalignment in language models. Studying this misgeneralization is important for developing better threat models and countermeasures, but is often infeasible due to the cost of RL on large models. As an alternative, we propose studying emergent misalignment from iterative DPO, which preserves important properties of RLVR while reducing costs and enabling training on popular finetuning APIs. In practice, we find that training GPT-4.1 with iterative DPO on a single-turn reward hacking environment induces covert misaligned power-seeking and alignment faking, the first openly available (semi)-online training pipeline to induce these concerning forms of misalignment. We also find that training Qwen2.5-32B-Instruct with the same pipeline induces both misalignment and improved instruction following accuracy, showing that iterative DPO can be used as a testbed for selective generalization. Overall, we think iterative DPO can help democratize and accelerate the study of emergent misalignment from RLVR.
Chinese Translation
在基于可验证奖励的强化学习(RLVR)过程中,奖励作弊可能诱发语言模型的奖励寻求行为和广泛的失调。研究这种错误泛化对于开发更好的威胁模型和应对措施十分重要,但由于在大模型上进行强化学习的成本高昂,相关研究往往难以开展。作为替代方案,我们提出通过迭代DPO研究涌现性失调,该方法在保留RLVR重要特性的同时降低了成本,并可在流行的微调API上进行训练。在实践中,我们发现使用迭代DPO在单轮奖励作弊环境中训练GPT-4.1,会诱导出隐蔽的失调性权力寻求和对齐伪装行为,这是首个公开可用的(半)在线训练流水线,能够诱导出这些令人担忧的失调形式。我们还发现,使用相同流水线训练Qwen2.5-32B-Instruct会同时产生失调和指令遵循准确性的提升,表明迭代DPO可用作研究选择性泛化的测试平台。总之,我们认为迭代DPO有助于推动RLVR涌现性失调研究的普及和加速。
cs.LG / 90 / 2609.06651

SwiftExplorer: Training-free Diffusion Model Alignment with Swift Diversity Exploration

SwiftExplorer:基于快速多样性探索的无训练扩散模型对齐方法
Yan, Renye, Cheng, Jikang, Wu, You, Huang, Bojin, Peng, Wei, Wang, Zongwei, Liang, Ling, Cai, Yimao
Abstract
Diffusion models have general generative abilities but struggle to align with specific objectives. Fine-tuning can improve alignment, yet its training cost is often prohibitive. This led to training-free methods that apply objective-guided terms in sampling to bias the generation distribution toward designated regions, e.g., high-reward areas. However, these methods face two issues: (1) the strong directional bias narrows the pretrained distribution and generation diversity, and (2) indiscriminate constant guidance fails to prune redundant signals, hurting both quality and efficiency. To address the above challenges, we propose SwiftExplorer, a plugin that mitigates distribution collapse caused by excessive diversity loss and reduces compute costs. First, we adopt an Inheritance-Restart exploration mechanism to avoid early convergence, while exploration also increases the likelihood of high-reward trajectories. Additionally, it balances diversity and fidelity, adding diversity without causing a distribution over-shift. Second, our Quality-Efficiency arbitration mechanism improves guidance by removing incorrect signals, and it reduces computation by dynamically stopping generation when completeness and marginal reward gain are optimal. In an extensive number of experiments and different types of evaluation metrics, the proposed SwiftExplorer achieves excellent performance on all metrics, including preference, fidelity, diversity, and richness.
Chinese Translation
扩散模型具有通用的生成能力,但难以与特定目标对齐。微调可以改善对齐效果,但其训练成本往往过于高昂。这催生了无训练方法,即在采样过程中应用目标引导项,使生成分布偏向指定区域(例如高奖励区域)。然而,这些方法面临两个问题:(1)强烈的方向性偏置缩小了预训练分布的范围和生成多样性;(2)不加区分的恒定引导无法剪除冗余信号,损害了生成质量和效率。为应对上述挑战,我们提出SwiftExplorer,这是一个插件式方法,可缓解由过度多样性损失导致的分布坍塌,并降低计算成本。首先,我们采用“继承-重启”探索机制以避免过早收敛,同时探索也提高了获得高奖励轨迹的可能性。此外,该机制平衡了多样性与保真度,在增加多样性的同时不会导致分布过度偏移。其次,我们的“质量-效率”仲裁机制通过移除错误信号来改进引导,并通过在完整性和边际奖励收益达到最优时动态停止生成来减少计算量。在大量实验和不同类型的评估指标上,所提出的SwiftExplorer在所有指标上均取得了优异表现,包括偏好性、保真度、多样性和丰富性。
cs.LG / 91 / 2609.06656

Assessing Covariate-Informed Grid Load Forecasting with a Time-Series Foundation Model

基于时间序列基础模型的协变量信息辅助电网负荷预测评估
Pendyala, Varsha, Fu, Yiwei, Yan, Weizhong, Virani, Nurali
Abstract
Modern power systems are growing increasingly complex as they integrate diverse generation sources to meet rising demand, making accurate load forecasting challenging. Recent advances in time-series foundation models (TSFMs) resulted in promising performance in zero-shot univariate load forecasting tasks. However, real-world load forecasting often involves multiple target variables and requires the integration of exogenous variables, raising important questions about the utility of TSFMs in realistic settings. In this study, we position Chronos-2, a recently developed model by Amazon, as a representative multi-channel TSFM that supports univariate, multivariate, and covariate-informed forecasting, and conduct a systematic investigation of how such models can be used for real-world load forecasting. While prior work has evaluated Chronos-2 on a limited number of energy-related tasks in a zero-shot setting, its performance relative to established task-specific deep learning models and its behavior when adapted using task-specific historical data remains insufficiently understood. In this work, we evaluate Chronos-2 on two real-world utility datasets, ISO New England and ENTSO-E, and benchmark it against widely used task-specific deep learning models. Our results show that Chronos-2 benefits substantially from task-specific fine-tuning and achieves strong short-horizon forecasting performance, but its zero-shot accuracy lags behind task-specific models and its forecasting error grows more rapidly with increasing forecast steps. Overall, this study provides a detailed characterization of the strengths and limitations of TSFMs such as Chronos-2 in grid load forecasting and offers practical insights into how a pretrained TSFM can be effectively adapted for operational load forecasting applications.
Chinese Translation
现代电力系统正变得日益复杂,为满足不断增长的需求而整合多种发电来源,这使得准确的负荷预测面临挑战。时间序列基础模型(TSFM)的最新进展使其在零样本单变量负荷预测任务中展现出良好的性能。然而,现实世界的负荷预测通常涉及多个目标变量,并需要整合外生变量,这引发了关于TSFM在实际场景中实用性的重要问题。在本研究中,我们将Chronos-2(亚马逊近期开发的模型)定位为具有代表性的多通道TSFM,该模型支持单变量、多变量以及协变量信息辅助的预测,并系统地研究了此类模型如何用于真实世界的负荷预测。尽管已有工作在零样本设置下于数量有限的能源相关任务上对Chronos-2进行了评估,但其相对于成熟的任务专用深度学习模型的性能,以及利用任务专用历史数据进行适配后的表现,仍缺乏充分的理解。在本工作中,我们在两个真实世界的电力公司数据集(ISO New England和ENTSO-E)上评估了Chronos-2,并将其与广泛使用的任务专用深度学习模型进行了基准对比。结果表明,Chronos-2从任务专用的微调中显著获益,并取得了较强的短期预测性能,但其零样本准确率落后于任务专用模型,且其预测误差随预测步数增加而增长得更快。总体而言,本研究详细刻画了以Chronos-2为代表的TSFM在电网负荷预测中的优势与局限,并为如何有效适配预训练TSFM以用于实际负荷预测应用提供了实用的见解。
cs.LG / 92 / 2609.06667

Behavioral Cloning Outperforms Entropy-Regularized RL: Critic-Driven Failure of Actor-Critic Methods on Adaptive Tumor Treatment

行为克隆优于熵正则化强化学习:自适应肿瘤治疗中演员-评论家方法的评论家驱动性失效
Dimitrov, Aleksandar, Spigler, Giacomo
Abstract
Adaptive dosing requires policies that reduce tumor burden without excessive toxicity. Learned dosing policies are typically judged against historical or heuristic comparators, which cannot show whether a policy has found the best behavior available. We instead study a three-population tumor-control ODE in which optimal-control analysis fixes the form of a good schedule -- bang-bang dosing punctuated by a singular arc -- and construct a numerical controller of that form as a proxy for near-optimal behavior. Judged against this reference under a sustained-cure criterion -- 200 consecutive days below 5% carrying capacity -- Soft Actor-Critic (SAC) trained from scratch never reaches cure. Behavioral cloning (BC) of the reference reproduces it (100% sustained cure, 30/30 seeds), but SAC fine-tuning of the cloned policy destroys it across five entropy coefficients, and TD3 and BC-regularized SAC fail identically; the pattern persists under multiplicative pharmacokinetic action noise. Along curative trajectories the post-collapse critic ranks the collapsed-policy action above the reference action in 96% of states, concentrated in the maintenance phase, and the policy settles into a non-curative adaptive-therapy equilibrium. The reference is what makes this legible: against a heuristic comparator the fine-tuned policy would read as a competent controller rather than a failure.
Chinese Translation
自适应给药需要在不产生过度毒性的前提下降低肿瘤负荷的策略。学习到的给药策略通常仅与历史或启发式对照进行比较,这无法表明该策略是否已经找到了可获得的最佳行为。我们转而研究一个三群体肿瘤控制的常微分方程(ODE)模型,其中最优控制分析确定了良好给药方案的形式——即由奇异弧段间隔的bang-bang(开关式)给药——并构建了该形式的数值控制器作为近优行为的代理基准。在持续治愈判据下(肿瘤体积连续200天低于5%的环境容纳量),从头训练的软演员-评论家(Soft Actor-Critic, SAC)算法从未达到治愈。对该参考策略进行行为克隆(BC)可以复现其效果(持续治愈率100%,30/30个随机种子),但对克隆策略进行SAC微调在五个熵系数设置下均破坏了其性能,TD3以及带BC正则化的SAC也以相同方式失败;该模式在乘性药代动力学动作噪声下依然存在。在治愈轨迹上,崩溃后的评论家在96%的状态中将崩溃策略的动作排在参考动作之上,且集中于维持期,最终策略收敛至一个非治愈的自适应治疗平衡态。正是这一参考基准使上述失效得以清晰呈现:若以启发式对照为基准,微调后的策略会被误认为是一个称职的控制器而非失败案例。
cs.LG / 93 / 2609.06668

Towards Unified Multimodal Graph Foundation Model: A Bridge-Router-Adapter Based Approach

迈向统一的多模态图基础模型:一种基于桥接-路由-适配器的方法
Zhang, Sirui, Zhou, Yubing, Li, Xunkai, Chen, Zekai, Li, Shumeng, Luo, Wang, Zhu, Yinlin, Gao, Yujin, Li, Rong-Hua
Abstract
Multimodal graphs couple node attributes in different modalities, such as text and images, with relational structure, enabling topological structure and cross-modality attributes to be modeled jointly. Multimodal graph foundation models seek unified representations from such data that transfer across different graph domains and downstream tasks. However, existing methods exhibit two fundamental limitations. (1) Cross-Scope Context Entanglement. They merge scope-specific graph contexts into a unified representation, obscuring their distinctions during multimodal construction. (2) Scope-Ignorant Modality Routing. They route modalities within a fixed graph scope, overlooking how modality relevance varies across neighborhood ranges. To address these challenges, we propose BRAIN, a unified model that focuses on graph context that combines neighborhood scope with modality composition. BRAIN comprises a scope-conditioned Bridge that combines structural information spanning local-to-global neighborhood scopes with different modality compositions; a hierarchical Router that estimates the relevance between the scope and the task, and selects compositions separately within each scope, allowing modality utility to vary with graph range; and a lightweight residual Adapter that further specializes the routed embedding for downstream prediction. BRAIN is trained through multi-graph pretraining followed by task-specific adaptation. Experiments across nine datasets and four task families demonstrate its broad effectiveness, improving node-classification and link-prediction performance by up to 4.73% relative to the strongest baseline, while achieving an average relative improvement of 14.72% across four graph-to-text and two graph-to-image metrics.
Chinese Translation
多模态图将不同模态(如文本和图像)的节点属性与关系结构耦合在一起,使得拓扑结构和跨模态属性能够被联合建模。多模态图基础模型旨在从这类数据中学习可跨不同图领域和下游任务迁移的统一表示。然而,现有方法存在两个根本性局限:(1)跨范围上下文纠缠。它们将特定范围的图上下文合并为统一表示,在多模态构建过程中掩盖了它们之间的差异。(2)忽略范围的模态路由。它们在固定的图范围内路由模态,忽视了模态相关性如何随邻域范围而变化。为应对这些挑战,我们提出了BRAIN,一个专注于结合邻域范围与模态组成的图上下文的统一模型。BRAIN包含:一个范围条件化的桥接模块(Bridge),将跨越局部到全局邻域范围的结构信息与不同模态组合相结合;一个层次化的路由模块(Router),估计范围与任务之间的相关性,并在每个范围内分别选择模态组合,使模态效用随图范围变化;以及一个轻量级的残差适配器(Adapter),进一步将路由后的嵌入特化用于下游预测。BRAIN通过多图预训练及后续的任务特定适配进行训练。在九个数据集和四类任务上的实验证明了其广泛的有效性:与最强基线相比,节点分类和链路预测性能相对提升高达4.73%,并且在四个图到文本指标和两个图到图像指标上平均相对提升达14.72%。
cs.LG / 94 / 2609.06670

Data Efficient Sample Selection for In-Context Learning

面向上下文学习的数据高效样本选择方法
Venktesh, V, levi, Cem, Anand, Avishek
Abstract
The In-context learning (ICL) paradigm aids large language models (LLMs) to adapt to new tasks without need for fine-tuning. However, selecting an optimal combination of demonstration examples from a large pool of example subsets is a challenging problem. Existing approaches for selection do not model the complex relationship between ICL samples and downstream LLM performance. They typically perform static task-level selection, choosing subsets once offline, which can fail to generalize to unseen queries. We introduce DearICL (Data Efficient Algorithm for Ranking) ICL samples, a new framework that models demonstration example selection as a subset ranking problem. DearICL employs a non-linear surrogate employing a differentiable sorting objective within a gap-index bandit algorithm. The gap-index based approach enables fine-grained separation of good arms and borderline arms, which is used as an auxiliary objective to train the non-linear surrogate through sufficient sampling of borderline arms, supporting instance-level subset ranking. On exemplar selection benchmarks with open-source LLMs, DearICL achieves 8.08-15.9% accuracy gains over strong linear bandit baselines, with low sample complexity. Code and data: https://github.com/VenkteshV/DearICL.
Chinese Translation
上下文学习(In-context Learning, ICL)范式使大型语言模型(LLMs)能够在无需微调的情况下适应新任务。然而,如何从大量的示例子集中选择最优的演示示例组合是一个具有挑战性的问题。现有的选择方法未能建模ICL样本与下游LLM性能之间的复杂关系。它们通常执行静态的任务级选择,即离线一次性选择子集,这可能导致无法泛化到未见过的查询。我们提出了DearICL(Data Efficient Algorithm for Ranking,用于排序ICL样本的数据高效算法),这是一个将演示示例选择建模为子集排序问题的新框架。DearICL采用一个非线性代理模型,在间隔索引多臂老虎机算法中使用可微排序目标。基于间隔索引的方法能够对优质臂(good arms)和临界臂(borderline arms)进行细粒度区分,并将其作为辅助目标,通过对临界臂的充分采样来训练非线性代理模型,从而支持实例级的子集排序。在使用开源LLM的示例选择基准测试中,DearICL以较低的样本复杂度相比强大的线性多臂老虎机基线取得了8.08%至15.9%的准确率提升。代码与数据:https://github.com/VenkteshV/DearICL。
cs.LG / 95 / 2609.06671

Tracking the Moving Frontier: Long-Short Term Advantage Estimator

追踪移动的前沿:长短期优势估计器
Yao, Xinhao, Yu, Lu, Wang, Changhao, Teng, Fengwei, Zhang, Yuyao, Cui, Qing, Zhou, Jun, Liu, Yong
Abstract
Group-based RLVR methods estimate advantages by repeatedly sampling multiple trajectories for each prompt, making long-horizon agent training expensive and discarding useful experience accumulated across iterations. We ask whether historical experience can replace these repeated within-iteration comparisons without directly optimizing on stale trajectories. We introduce Long-Short Term Advantage Estimator (LSTAE), a single-stream RL algorithm that uses history for advantage estimation while updating the policy only with the current rollout. LSTAE maintains a persistent tracker for each task anchor. At the trajectory level (long term), a drift-aware historical baseline tracks the anchor's moving success frontier and measures the relative contribution of each new trajectory. At the step level (short term), a recent state-experience buffer exploits recurrent states to estimate localized action advantages. This two-timescale design converts accumulated experience into multi-granular credit signals, requiring only one rollout per anchor. Across agentic and mathematical reasoning benchmarks, LSTAE matches or improves upon strong group-based baselines while substantially reducing rollout cost.
Chinese Translation
基于分组的RLVR方法通过为每个提示重复采样多条轨迹来估计优势,这使得长时程智能体训练代价高昂,并且丢弃了跨迭代累积的有用经验。我们探讨历史经验能否替代这些重复的迭代内比较,而不直接在过时的轨迹上进行优化。我们提出了长短期优势估计器(Long-Short Term Advantage Estimator, LSTAE),这是一种单流强化学习算法,利用历史信息进行优势估计,同时仅使用当前轮次的采样结果来更新策略。LSTAE为每个任务锚点维护一个持久的追踪器。在轨迹层面(长期),一个漂移感知的历史基线追踪该锚点的移动成功前沿,并衡量每条新轨迹的相对贡献。在步骤层面(短期),一个近期状态经验缓冲区利用循环出现的状态来估计局部化的动作优势。这种双时间尺度设计将积累的经验转化为多粒度的信用信号,每个锚点仅需一次采样。在智能体和数学推理基准测试中,LSTAE匹配或超越了强大的基于分组的基线方法,同时大幅降低了采样成本。
cs.LG / 96 / 2609.06761

LATS: Levy Adaptive Tree Sampling for Feedback-Driven Diverse Target Discovery

LATS:用于反馈驱动多样化目标发现的列维自适应树采样
Ji, Binglin, Sarkar, Anindya, Lu, Hengchang, Kong, Lecheng, Chen, Yixin, Vorobeychik, Yevgeniy
Abstract
While diffusion models excel at capturing complex data distributions, scientific discovery often requires steering generation toward specific, uncharacterized regions that maximize a target objective. These high-utility modes frequently reside in low-likelihood tail regions and are only revealed sequentially through interactive feedback. Existing diffusion samplers fail in this regime: they inherit the pre-trained model's bias toward high-density regions, leaving rare yet promising phenomena underexplored. Conversely, exploration-heavy samplers ensure broad coverage but fail to efficiently exploit high-utility modes when constrained by a strict sampling budget. To resolve this dilemma, we introduce Levy Adaptive Tree Search (LATS), a principled sampling framework for online feedback-driven search. LATS leverages heavy-tailed exploration coupled with tree-based value backpropagation to progressively uncover preferred modes. By maintaining broad distributional coverage, LATS successfully discovers low-likelihood, high-utility regions while preserving sample fidelity and structural diversity. Experiments across diverse benchmarks, including materials science, demonstrate that LATS significantly outperforms baselines in target discovery efficiency.
Chinese Translation
尽管扩散模型擅长捕捉复杂的数据分布,但科学发现往往需要将生成过程引导至能够最大化目标函数的特定、尚未被充分表征的区域。这些高效用模态通常位于低似然的尾部区域,并且只能通过交互式反馈逐步显现。现有的扩散采样器在这一情境下表现不佳:它们继承了预训练模型对高密度区域的偏向,导致稀有但有前景的现象未被充分探索。相反,侧重探索的采样器虽然保证了广泛的覆盖范围,但在严格采样预算的约束下无法高效利用高效用模态。为了解决这一困境,我们提出了列维自适应树搜索(Levy Adaptive Tree Search,LATS),这是一个面向在线反馈驱动搜索的具有原则性的采样框架。LATS利用重尾探索机制,结合基于树的数值回传,逐步发现偏好模态。通过保持广泛的分布覆盖,LATS在保持样本保真度和结构多样性的同时,成功发现了低似然、高效用的区域。在包括材料科学在内的多个基准测试上的实验表明,LATS在目标发现效率方面显著优于基线方法。
cs.LG / 97 / 2609.06779

DrugReason: Dynamic Multi-View Reasoning over Knowledge Graph and Language Evidence for Drug Repurposing

DrugReason:基于知识图谱与语言证据的动态多视角推理药物重定位方法
Liu, Zijie, Li, Hongxuan, Tan, Zhen, Duan, Jinhao, Huang, Baixiang, Liu, Zunpeng, Shu, Kai, Chen, Tianlong
Abstract
Drug repurposing aims to identify new therapeutic uses for existing compounds and, compared with de novo drug discovery, offers a faster and more cost-effective path to clinical translation. However, the space of candidate drug-disease pairs is enormous and their underlying relationships often depend on complex multi-hop biological mechanisms, making it difficult to reliably predict which pairs represent true therapeutic relationships. Existing approaches tackle this from two directions: knowledge graph-based methods organize curated biomedical evidence into structured relational networks for grounded multi-hop reasoning, while LLM-based methods leverage pretrained knowledge to generate flexible mechanistic rationales. Yet neither is sufficient alone - KGs are confined to observed graph structure while LLMs lack factual grounding and risk hallucination. To address this gap, we propose DrugReason, a multi-view reasoning framework that integrates grounded KG reasoning with LLM-generated mechanistic inference for drug repurposing. DrugReason adaptively routes diverse reasoning paths to specialized experts conditioned on the query context, while a cross-expert distillation objective enables knowledge sharing without sacrificing expert specialization. Experiments on PharmaDB, DDInter, and DrugBank show that DrugReason improves average performance over strong single-view reasoning baselines and achieves competitive or superior results compared with graph-based alternatives, while providing interpretable routing-based predictions.
Chinese Translation
药物重定位旨在为现有化合物发现新的治疗用途,与从头药物发现相比,它为临床转化提供了更快、更具成本效益的途径。然而,候选药物-疾病对的空间极其庞大,且其潜在关系往往依赖于复杂的多跳生物学机制,这使得难以可靠地预测哪些药物-疾病对代表真实的治疗关系。现有方法从两个方向应对这一挑战:基于知识图谱(KG)的方法将整理好的生物医学证据组织为结构化关系网络,以进行有依据的多跳推理;而基于大语言模型(LLM)的方法则利用预训练知识生成灵活的机制性推理依据。然而,两者单独使用均不够充分——知识图谱受限于已观测到的图结构,而大语言模型缺乏事实依据且存在幻觉风险。为填补这一空白,我们提出了DrugReason,一个将基于知识图谱的有据推理与大语言模型生成的机制性推断相结合的多视角药物重定位推理框架。DrugReason根据查询上下文自适应地将多样化的推理路径路由到专门的专家模块,同时跨专家蒸馏目标在不牺牲专家专业性的前提下实现知识共享。在PharmaDB、DDInter和DrugBank数据集上的实验表明,DrugReason相比强大的单视角推理基线提升了平均性能,与基于图谱的替代方法相比取得了具有竞争力或更优的结果,同时提供了可解释的基于路由的预测。
cs.LG / 98 / 2609.06786

When Retain Constraints Conflict: Mitigating Forget-Retain Interference in Tabular Data

当保留约束发生冲突时:缓解表格数据中的遗忘-保留干扰
Liu, Zijie, Duan, Jinhao, Shang, Bingqi, An, Xinming, Liu, Sijia, Chen, Tianlong
Abstract
Machine unlearning aims to remove the influence of designated training data while preserving model utility, but its behavior on tabular data remains underexplored. This gap is important because tabular prediction is widely used in high-stakes domains and is increasingly adapted to language models through record serialization and schema-aware prompting. We identify a key challenge that distinguishes tabular unlearning from unlearning in free-form text or other modalities: schema-induced forget-retain overlap. In serialized tabular data, records share fixed column-name/value slots, similar attribute ranges, and common output spaces. Consequently, a forget row may have nearby retain rows that rely on the same high-signal attributes, causing retain preservation to oppose the update required for forgetting. Motivated by this failure mode, we propose Conflict-Aware Unlearning (CAU), a schema-aware approach that reduces forget-retain interference by relaxing preservation constraints on retained rows that most conflict with the forget set. Across sample-level and feature-level unlearning on clinical and non-medical tabular tasks, CAU more closely matches a retraining oracle while maintaining predictive utility and retain-region behavior. Our results show that reliable tabular LLM unlearning depends not only on the forgetting objective, but also on how retain constraints are constructed.
Chinese Translation
机器遗忘旨在移除指定训练数据的影响,同时保持模型效用,但其在表格数据上的行为仍未得到充分研究。这一空白十分重要,因为表格预测被广泛应用于高风险领域,并且正通过记录序列化和模式感知提示日益适配到语言模型中。我们识别出一个关键挑战,它将表格数据遗忘与自由文本或其他模态的遗忘区分开来:由模式引起的遗忘-保留重叠。在序列化的表格数据中,各条记录共享固定的列名/取值槽位、相似的属性范围以及共同的输出空间。因此,一个待遗忘行附近可能存在依赖于相同高信息量属性的保留行,导致保留约束与遗忘所需的更新方向相互对立。基于这一失效模式,我们提出冲突感知遗忘(Conflict-Aware Unlearning, CAU),这是一种模式感知的方法,通过放宽与待遗忘集合冲突最大的保留行的保持约束,来减少遗忘-保留之间的干扰。在临床和非医疗表格任务上的样本级与特征级遗忘实验中,CAU 在保持预测效用和保留区域行为的同时,更接近地匹配了重新训练的基准。我们的结果表明,可靠的表格大语言模型遗忘不仅取决于遗忘目标,还取决于保留约束的构建方式。
cs.LG / 99 / 2609.06806

Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning

聪明训练而非苦练:主动学习中的切换信号引导训练
Omar, Nagham, Rozenshtein, Maya, Mishlyakov, Evgeny, Gal, Avigdor
Abstract
Training strategy, namely whether to retrain from scratch or fine-tune from the previous checkpoint, is an overlooked decision variable in active learning. We show that this choice has exploitable structure: retraining is most useful in early rounds, when each batch can substantially reshape the labeled distribution, while fine-tuning becomes safer once the model trajectory stabilizes. We propose HybridAL, an adaptive training schedule that monitors an online stabilization signal and switches from retraining to fine-tuning after sustained stabilization. Two complementary signals, spectral exponent change $\Delta\alpha$ (weight-based) and accuracy change $\Delta$Acc (validation-based), span different points on the time-calibration trade-off. Across three encoder backbones and six text-classification tasks (five seeds each), HybridAL keeps endpoint macro-F1 non-inferior to retraining and fine-tuning at a 0.010 margin, saves up to 49% of retraining time, and recovers a substantial fraction of retraining's calibration advantage as measured by negative log-likelihood (NLL). Compared with schedules that switch at a pre-committed round, HybridAL obtains lower NLL at moderate additional cost, showing that trajectory-dependent switching provides a stronger time-calibration trade-off than fixed early switching.
Chinese Translation
训练策略,即是从头重新训练还是从上一个检查点进行微调,是主动学习中一个被忽视的决策变量。我们证明这一选择具有可利用的结构性规律:重新训练在早期轮次中最为有用,此时每一批新数据都能显著重塑已标注数据的分布;而一旦模型轨迹趋于稳定,微调则变得更为安全。我们提出 HybridAL,一种自适应训练调度方法,它监控一个在线稳定化信号,并在持续稳定后从重新训练切换为微调。两个互补的信号——谱指数变化 Δα(基于权重)和准确率变化 ΔAcc(基于验证集)——在时间-校准权衡上覆盖了不同的位置。在三个编码器主干和六个文本分类任务(每个任务五个随机种子)上的实验表明,HybridAL 在 0.010 的容差下,其最终宏平均 F1 不劣于重新训练和微调,最多可节省 49% 的重新训练时间,并按负对数似然(NLL)衡量,恢复了重新训练在校准方面的相当一部分优势。与在预先固定轮次切换的调度方法相比,HybridAL 以适度的额外代价获得更低的 NLL,表明基于轨迹依赖的切换比固定的早期切换能提供更优的时间-校准权衡。
cs.LG / 100 / 2609.06830

Constrained Bayesian Optimization for Hierarchical Federated Learning in IoT Networks for Plant Disease Classification

面向植物病害分类的物联网网络中分层联邦学习的约束贝叶斯优化
Papanikolaou, Athanasios, Tziouvaras, Athanasios, Xenakis, Apostolos, Chatzimisios, Periklis, Parambath, Shameem A. Puthiya, Floros, George, Zereik, Enrica, Petrovic, Ivan, Bonsignorio, Fabio
Abstract
The deployment of Hierarchical Federated Learning (HFL) in resource-constrained Internet of Things (IoT) environments requires careful configuration to balance predictive performance with energy consumption and execution time. This challenge is particularly relevant to smart agriculture, where distributed IoT devices can support automated plant disease classification while operating under limited computational and communication resources. This paper presents a constrained Bayesian Optimization framework for the efficient configuration of HFL deployments. The proposed approach jointly explores the deep learning backbone architecture, aggregation strategy, and number of communication rounds, while the federation size is determined according to the spatial coverage requirements of the agricultural deployment. A weighted objective function captures user-defined trade-offs among energy consumption, execution time, and predictive performance, while explicit constraints ensure compliance with deployment-specific resource and accuracy requirements. The framework is evaluated on an IoT-based plant disease classification task considering multiple deep learning architectures, federated aggregation strategies, and communication-round settings. Experimental results across 30 independent optimization runs show that the proposed approach explores only 11.11% of the search space, while consistently identifying solutions within 1% of the exhaustive-search optimum, with a mean optimality gap of only 0.056%.
Chinese Translation
在资源受限的物联网环境中部署分层联邦学习需要精心配置,以在预测性能与能耗和执行时间之间取得平衡。这一挑战在智慧农业中尤为重要,因为分布式物联网设备可在计算和通信资源有限的情况下支持自动化植物病害分类。本文提出了一种用于高效配置分层联邦学习部署的约束贝叶斯优化框架。该方法联合探索深度学习骨干网络架构、聚合策略和通信轮数,而联邦规模则根据农业部署的空间覆盖需求确定。加权目标函数刻画了用户自定义的能耗、执行时间和预测性能之间的权衡,同时显式约束确保满足特定部署的资源和精度要求。该框架在基于物联网的植物病害分类任务上进行评估,考虑了多种深度学习架构、联邦聚合策略和通信轮数设置。在30次独立优化运行中的实验结果表明,所提方法仅探索了11.11%的搜索空间,同时始终能找到与穷举搜索最优解相差在1%以内的解,平均最优性差距仅为0.056%。
cs.LG / 101 / 2609.06862

Feature Superposition in Neural Networks: From Theory to Practice

神经网络中的特征叠加:从理论到实践
Shi, Dai, Li, Xiaoyu, Han, Andi, Hernández-Lobato, José Miguel
Abstract
Superposition refers to neural networks representing more features than they have dimensions. It offers a possible explanation for polysemantic neurons and motivates methods for recovering interpretable features from neural activations. Theoretical models typically start with a given set of input features and assumptions about how their values vary across inputs, then study how a network encodes those values in a lower-dimensional hidden representation. Empirical work, by contrast, seeks to identify the features encoded in trained networks and determine their role in computation. In this survey, we review the geometry, learning, and computation of superposed representations, explaining how feature statistics and decoder choice affect the conclusions. To connect these theoretical accounts with evidence from trained networks, we compare practical methods for recovering and analyzing features and examine what their evaluations establish. Since accurate activation reconstruction alone does not establish feature identity or causal use, we discuss the methods' documented failures and applications in light of the evidence available for these different claims. Finally, we assess previously stated open problems and identify remaining theoretical and empirical questions about superposition in trained networks. We hope our work can pave the way for a deeper understanding of superposition and more reliable methods for interpreting neural networks.
Chinese Translation
叠加(Superposition)是指神经网络所表示的特征数量超过其自身维度的现象。它为多义神经元(polysemantic neurons)提供了一种可能的解释,并推动了从神经激活中恢复可解释特征的相关方法的发展。理论模型通常从一组给定的输入特征以及关于特征值在不同输入间如何变化的假设出发,进而研究网络如何将这些值编码到低维的隐层表示中。相比之下,实证工作则致力于识别已训练网络中所编码的特征,并确定它们在计算中的作用。在本综述中,我们回顾了叠加表示的几何性质、学习过程与计算机制,阐述了特征统计特性与解码器选择如何影响相关结论。为了将这些理论解释与来自已训练网络的证据联系起来,我们比较了用于恢复和分析特征的实用方法,并考察了其评估所能够确立的内容。由于仅凭精确的激活重构并不能确立特征的身份或因果作用,我们基于这些不同主张所拥有的证据,讨论了相关方法已被记录的失败案例及其应用。最后,我们评估了先前提出的开放性问题,并指出了关于已训练网络中叠加现象仍待解决的理论与实证问题。我们希望本工作能够为更深入地理解叠加现象以及开发更可靠的神经网络解释方法铺平道路。
cs.LG / 102 / 2609.06869

PPIM: Pennes Physics-Informed Mamba for Heat-Source-Conditioned 3D Bioheat Simulation

PPIM:基于Pennes物理信息的Mamba模型用于热源条件下的三维生物热仿真
Lee, Dongyun, Yoon, Kyungho, Shin, Minwoo
Abstract
Three-dimensional bioheat simulation aims to predict transient temperature distributions in biological tissue and is commonly modeled using the Pennes bioheat equation, which combines thermal diffusion, perfusion-mediated heat loss, and external heat generation. In this study, we consider a controlled 3D Pennes bioheat simulation under a localized heat-source condition inspired by microwave ablation (MWA). To evaluate neural approximation performance, we compare three neural partial differential equation (PDE) solvers under the same controlled simulation: a spatial Fourier-feature physics-informed neural network (PINN), a generic PINNMamba temporal subsequence model, and Pennes Physics-Informed Mamba (PPIM). PPIM builds on the temporal subsequence model by incorporating conditioned heat-source input and Pennes-aware state-space model (SSM) decay initialization. All three neural models are trained under the same conditions with the same Pennes residual, and an explicit finite-difference method (FDM) solution is used only as the numerical reference. In a representative 600~s run, PPIM achieved the lowest MAE, relative $L_1$ error, and relative $L_2$ error among the evaluated neural solvers. Error maps further showed that the remaining PPIM errors were more concentrated near the heat-source region than across the rest of the domain. These results indicate that PPIM is effective for approximating the FDM reference final temperature field in this controlled simulation. The source code is available at https://github.com/muvYun/PPIM.
Chinese Translation
三维生物热仿真旨在预测生物组织中的瞬态温度分布,通常采用Pennes生物热方程进行建模,该方程综合了热扩散、血流灌注引起的热损失以及外部热源生成。在本研究中,我们考虑了一个受控条件下的三维Pennes生物热仿真,其局部热源条件受微波消融(MWA)启发。为评估神经近似性能,我们在相同的受控仿真下比较了三种神经偏微分方程(PDE)求解器:基于空间傅里叶特征的物理信息神经网络(PINN)、通用的PINN-Mamba时间子序列模型,以及Pennes物理信息Mamba模型(PPIM)。PPIM在时间子序列模型的基础上,引入了热源条件输入和Pennes感知的状态空间模型(SSM)衰减初始化。三种神经模型均在相同条件下使用相同的Pennes残差进行训练,显式有限差分方法(FDM)的解仅用作数值参考。在一个具有代表性的600秒仿真中,PPIM在所评估的神经求解器中取得了最低的平均绝对误差(MAE)、相对$L_1$误差和相对$L_2$误差。误差图进一步表明,PPIM的残余误差较集中于热源区域附近,而非域的其余部分。这些结果表明,在该受控仿真中,PPIM能够有效逼近FDM参考的最终温度场。源代码可在 https://github.com/muvYun/PPIM 获取。
cs.LG / 103 / 2609.06872

Exact Record Omission in Delta Attention: A Transport Criterion, Its Cost, and a Replay Certificate

Delta 注意力中的精确记录删除:一个传输判据、其代价与一种重放证明
Ramesh, Vishwajith
Abstract
When a user asks an assistant to forget a record, the test is whether the memory now matches the state it would hold if the record had never been stored. Independently encoded rows can be removed directly; a recurrent memory folds records into an evolving state. One hope is a receipt: save the difference the record made when it arrived, carry it forward through later updates, and subtract it, so that deletion costs one fixed-size edit no matter how long the conversation runs. We show that a transported receipt reaches exact omission if and only if the changes the record induces in later updates cancel out on net, and we measure whether they do on the released 48B Kimi Linear hybrid. They do not: after 4,096 further tokens the record still leaves an imprint of about 4.5% of the state norm that none of the tested receipt classes removes, recomputing half the suffix closes less than half the gap, and the per-token log a receipt needs costs more than a full checkpoint after 88 tokens. The same write-rule classification held on Mamba-2, Falcon-H1, and RWKV-7 with predictions recorded before the runs. Restoring a checkpoint from before the record and replaying the surviving suffix matches the never-stored state exactly on every array we check. In the hybrid suffix sweep, masking the record's attention rows brings sampled recovery close to the never-stored floor even though the recurrent imprint remains, and an auditor who rebuilds the reference can still detect it. Among the evaluated methods, checkpoint replay achieves exact omission, with work proportional to the replayed suffix.
Chinese Translation
当用户要求助手遗忘某条记录时,检验标准在于:当前的记忆状态是否与该记录从未被存储过的状态完全一致。独立编码的行可以直接移除;而循环记忆(recurrent memory)则会将记录折叠进不断演化的状态中。一种理想方案是“收据”(receipt):在记录到达时保存其造成的变化,在后续更新中将其前向传递,然后减去它,使得删除操作的代价仅为一次固定大小的编辑,无论对话持续多长。我们证明,一个经传输的收据能够实现精确删除,当且仅当该记录在后续更新中所诱导的变化在净额上相互抵消;我们在已发布的 48B Kimi Linear 混合模型上测量了这些变化是否抵消。结果表明它们并不抵消:在后续 4,096 个 token 之后,该记录仍在状态范数中留下约 4.5% 的印记,所测试的任何收据类别都无法将其消除;重算后半段后缀所缩小的差距不足一半;而收据所需的逐 token 日志在 88 个 token 之后的代价就超过了完整检查点(checkpoint)。同样的写入规则分类在 Mamba-2、Falcon-H1 和 RWKV-7 上也成立,且相关预测均在实验运行前记录。从记录之前保存的检查点恢复并重放剩余后缀,在我们检查的每一个数组上都能与“从未存储”状态完全一致。在混合模型后缀扫描实验中,屏蔽该记录的注意力行可使采样恢复接近“从未存储”的底线,尽管循环印记仍然存在,且重建参考状态的审计者仍能检测到该印记。在所有评估的方法中,检查点重放实现了精确删除,其工作量与被重放的后缀长度成正比。
cs.LG / 104 / 2609.06881

Learning Adaptive SED for heterogeneous load balancing

面向异构负载均衡的自适应SED学习
van Kempen, Sanne, Sanders, Jaron, Sloothaak, Fiona, Wolf, Maarten G.
Abstract
We study a two-server load balancing system with heterogeneous service rates that are a priori unknown to the dispatcher. The goal is to route customers according to the Shortest--Expected--Delay (SED) policy, but this requires knowledge of the service rates. Empirical policies that route based on estimates perform poorly: due to estimation error, the empirical policy disagrees with the oracle on an infinite region of the state space. We propose an online learning algorithm that converges to SED while learning the service rates. The algorithm carefully balances empirical SED routing with forced exploration phases that guarantee sufficient sampling of both servers. We prove that our algorithm achieves finite regret; this differs from classical Multi-Armed Bandit settings where regret typically grows logarithmically in time. Finally, numerical experiments demonstrate the performance of our algorithm and highlight the regimes in which forced exploration is especially beneficial.
Chinese Translation
我们研究了一个双服务器负载均衡系统,其服务速率存在异构性且对调度器而言是先验未知的。目标是按照最短期望延迟(Shortest-Expected-Delay, SED)策略对顾客进行路由,但这需要已知服务速率。基于估计值进行路由的经验策略表现较差:由于估计误差,经验策略与最优策略(oracle)在状态空间的一个无限区域上不一致。我们提出了一种在线学习算法,在学习服务速率的同时收敛到SED策略。该算法谨慎地在经验SED路由与强制探索阶段之间取得平衡,强制探索阶段保证了对两台服务器的充分采样。我们证明该算法能够实现有限遗憾(finite regret);这与经典的多臂老虎机(Multi-Armed Bandit)设置不同,后者的遗憾通常随时间呈对数增长。最后,数值实验验证了我们算法的性能,并揭示了强制探索尤为有效的场景。
cs.LG / 105 / 2609.06882

Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning

离线强化学习中扩散策略的噪声空间策略梯度
Selim, Mahmoud, Cipriani, Cristina, Johansson, Karl H.
Abstract
Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remains conceptually and algorithmically challenging. In this work, we address this gap by introducing a noisy-space action-value (Q-)function that assigns values to diffusion latents through the distribution of executed actions induced by the denoising process. We show that this construction admits a precise semantic interpretation and derive a noisy-space policy gradient (NSPG) that optimizes noisy latents using only clean action-space value estimates. Building on this result, we formulate a KL-regularized policy improvement over noisy latents and show that the resulting objective admits a diffusion-compatible regression form, avoiding backpropagation through the denoising process. Empirical results on state-based D4RL benchmarks and vision-based OGBench tasks demonstrate that the proposed noisy-space objective provides a principled and effective basis for training diffusion policies in offline reinforcement learning. Project webpage: https://mahmoud-selim.github.io/NSPG/
Chinese Translation
扩散策略为连续控制提供了一种强大且富有表达力的参数化方法。然而,其与强化学习的结合在概念和算法层面仍面临挑战。本文通过引入噪声空间动作值(Q)函数来弥补这一空白,该函数通过去噪过程所诱导的执行动作分布为扩散潜变量赋值。我们证明了这一构造具有精确的语义解释,并推导出噪声空间策略梯度(Noisy-Space Policy Gradient, NSPG),其仅利用干净动作空间的价值估计即可对噪声潜变量进行优化。基于这一结果,我们提出了针对噪声潜变量的KL正则化策略改进方法,并证明所得目标函数具有与扩散过程兼容的回归形式,从而避免了通过去噪过程的反向传播。在基于状态的D4RL基准和基于视觉的OGBench任务上的实验结果表明,所提出的噪声空间目标函数为离线强化学习中扩散策略的训练提供了一个有理论依据且有效的框架。项目网页:https://mahmoud-selim.github.io/NSPG/
cs.LG / 106 / 2609.06912

From Synthetic Priors to Model Behavior: Structural Coverage in Tabular Foundation Models

从合成先验到模型行为:表格基础模型中的结构性覆盖
Zhao, He, Thompson, Ryan, Steinberg, Daniel M., Rahman, Ashfaqur, Bonilla, Edwin V., Ong, Cheng Soon
Abstract
Tabular foundation models (TFMs) are commonly pretrained on large collections of procedurally generated synthetic tasks, yet it remains unclear how well these synthetic pretraining priors support the downstream tasks on which the models are evaluated. We study this question from a distribution-level attribution perspective. We recover or reconstruct the synthetic data generators of four TFMs and compare their generated tasks with datasets from two widely used tabular benchmarks. Each dataset is represented by a common set of structural descriptors capturing schema, feature distributions, dependence structure, response properties, and feature--response relationships. In this space, we measure how broadly and repeatedly each synthetic prior reaches benchmark tasks using structural coverage and normalized density, and examine whether stronger local support is associated with better predictive performance. We find substantial differences across synthetic pretraining priors: some generators provide consistently broader and denser support for benchmark tasks than others. Moreover, stronger synthetic-to-benchmark support is generally associated with better relative model performance. These results suggest that structural coverage provides a useful diagnostic for characterizing synthetic pretraining priors and relating their data-generating assumptions to downstream model behavior.
Chinese Translation
表格基础模型(Tabular Foundation Models, TFMs)通常在大量程序化生成的合成任务集合上进行预训练,然而这些合成预训练先验在多大程度上能够支持模型所评估的下游任务,仍不清楚。我们从分布层面的归因视角研究这一问题。我们恢复或重构了四个表格基础模型的合成数据生成器,并将其生成的任务与两个广泛使用的表格基准数据集中的数据集进行比较。每个数据集由一组共同的结构描述子表示,这些描述子刻画了模式、特征分布、依赖结构、响应属性以及特征—响应关系。在该空间中,我们使用结构性覆盖率和归一化密度来衡量每个合成先验对基准任务的覆盖广度与重复程度,并考察更强的局部支持是否与更好的预测性能相关联。我们发现不同合成预训练先验之间存在显著差异:某些生成器为基准任务提供的一致支持在广度和密度上优于其他生成器。此外,更强的合成数据对基准任务的支持通常与更好的相对模型性能相关。这些结果表明,结构性覆盖率为刻画合成预训练先验、并将其数据生成假设与下游模型行为联系起来,提供了一种有效的诊断工具。
cs.LG / 107 / 2609.06921

Constrained Online Learning with Noisy Constraint Values

含噪声约束值的约束在线学习
Aggarwal, Vaneet
Abstract
We study constrained online convex optimization with adversarial constraints when constraint values and gradients are observed through unbiased noise. Gaussian value noise of standard deviation $\sigma$ yields a worst-case lower bound of $\Omega(\min\{\sigma,1\}T/\log^7T)$ on the maximum of expected regret and expected hard violation, even with known gradients. This rules out any jointly $O(T^{1-\delta})$ guarantee for fixed $\delta>0$ and fixed positive noise level. We therefore study budget violation: the largest cumulative overspend over any window within a fixed horizon. We introduce \LEDGER, which tracks observed net consumption in a nonnegative balance and sets constraint weights before the current feedback noise. Under common feasibility and conditional finite-variance feedback, for fixed problem parameters, \LEDGER\ achieves $O(\sqrt T/V)$ expected regret and $O(\sqrt V\,T^{3/4}+\sigma\sqrt T)$ expected budget violation for $V\in[T^{-1/2},1]$. This gives the pair $(O(\sqrt T),O(T^{3/4}))$ at $V=1$ and $(O(T^{2/3}),O(T^{2/3}))$ at $V=T^{-1/6}$, without a Slater condition. The budget-focused endpoint $V=T^{-1/2}$ gives $(O(T),O(\sqrt T))$. The same update yields $O((1+E[P_T])\sqrt T/V)$ expected dynamic regret for predictable feasible comparator paths, without common feasibility or path-length input. Its budget bound instead depends on the shortest feasible path, up to a dimension factor.
Chinese Translation
我们研究了约束值为对抗性设置、且约束值和梯度仅能通过无偏噪声观测的约束在线凸优化问题。即使梯度已知,标准差为 $\sigma$ 的高斯值噪声也会使得期望遗憾与期望硬违反的最大值在最坏情况下具有 $\Omega(\min\{\sigma,1\}T/\log^7 T)$ 的下界。这排除了对于任意固定 $\delta>0$ 和固定正噪声水平获得联合 $O(T^{1-\delta})$ 保证的可能性。因此,我们转而研究预算违反(budget violation),即固定时间范围内任意窗口上的最大累计超支。我们提出了 \LEDGER 算法,它以非负余额的形式跟踪观测到的净消耗,并在当前反馈噪声实现之前设定约束权重。在常见的可行性条件与条件有限方差反馈假设下,对于固定的问题参数,\LEDGER\ 在 $V\in[T^{-1/2},1]$ 时实现 $O(\sqrt T/V)$ 的期望遗憾和 $O(\sqrt V\,T^{3/4}+\sigma\sqrt T)$ 的期望预算违反。在 $V=1$ 时得到 $(O(\sqrt T),O(T^{3/4}))$ 的组合,在 $V=T^{-1/6}$ 时得到 $(O(T^{2/3}),O(T^{2/3}))$,且无需 Slater 条件。在预算导向的端点 $V=T^{-1/2}$ 处则得到 $(O(T),O(\sqrt T))$。同样的更新规则对于可预测的可行比较路径可实现 $O((1+E[P_T])\sqrt T/V)$ 的期望动态遗憾,且无需共同可行性假设或路径长度输入。其预算界则(在相差一个维度因子的情况下)取决于最短可行路径。
cs.LG / 108 / 2609.06934

The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists

拒绝的几何学:为什么事后安全是脆弱的而预训练时安全是持久的
Malla, Srikanth, Choi, Chiho, Choi, Joon Hee
Abstract
Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023b), fine-tuning attacks (Qi et al., 2024), and activation-space probes (Arditi et al., 2024) keep recovering the behaviors it was meant to remove. We give this fragility one geometric explanation and trace it to when, during pretraining, safety can take hold. We measure the safety update $\Delta = W_{\text{safe}} - W_{\text{base}}$ against the curvature of the model's capabilities (the empirical Fisher of a capability loss). Post-hoc safety consistently lands in a suppression regime: $\Delta$ is nearly orthogonal to the capability directions, and its small in-subspace part concentrates on a few high-curvature ones. The update is thin but sharp, a refusal gate laid over intact capabilities rather than erasure of them. A kernel-immobility lemma explains why such an update can only mask a capability, not remove it, so a little benign fine-tuning restores it: 100 steps of benign fine-tuning collapse refusal on Qwen-2.5-7B and Llama-3-8B Instruct at preserved capability, a signature that replicates across five model families. Following the account into pretraining, a 267-checkpoint sweep of OLMo-2-1B (OLMo et al., 2025) shows the substrate that safety engages emerging in a sharp transition between roughly 6B and 60B pretraining tokens. We then use the account constructively: models trained from scratch with safety co-training spread continuously across pretraining reach 87 to 98% refusal whose post-attack level holds at 84 to 91% at every scale, an erosion of 2 to 14 pp against 35 to 38 pp for post-hoc installs, at capability matched or better than an LM-only baseline and holding from 410M to 6.9B, whereas a compute-matched windowed schedule installs no lasting refusal. Persistence of the safety signal across pretraining, not its timing, is what buys attack robustness.
Chinese Translation
事后安全训练(RLHF、DPO)是对齐大语言模型的主流方法,然而越狱攻击(Zou et al., 2023b)、微调攻击(Qi et al., 2024)以及激活空间探针(Arditi et al., 2024)不断恢复出它本要消除的行为。我们对这种脆弱性给出一个几何解释,并将其追溯到预训练过程中安全能够生效的时机。我们针对模型能力的曲率(能力损失的经验Fisher信息)来度量安全更新 $\Delta = W_{\text{safe}} - W_{\text{base}}$。事后安全始终落在一种抑制机制中:$\Delta$ 与能力方向近乎正交,且其在子空间内的微小部分集中于少数高曲率方向上。该更新是薄而尖锐的——一个覆盖在完整能力之上的拒绝门控,而非对能力的擦除。一个核不可移动性引理解释了为何此类更新只能掩蔽能力而无法将其移除,因而少量良性微调即可将其恢复:在Qwen-2.5-7B和Llama-3-8B Instruct上,100步良性微调在保持能力的前提下使拒绝行为坍塌,这一特征在五个模型家族中均可复现。将这一解释延伸到预训练阶段,对OLMo-2-1B的267个检查点的扫描(OLMo et al., 2025)显示,安全所依赖的底层结构在约60亿到600亿预训练词元之间的一次急剧转变中涌现。随后我们建设性地运用这一解释:从零开始训练并在整个预训练过程中持续进行安全共训练的模型,达到了87%到98%的拒绝率,且在攻击后于所有规模上仍保持在84%到91%,侵蚀幅度仅为2到14个百分点,而事后安装的方式则为35到38个百分点,同时能力与仅语言模型的基线持平或更优,且从410M到6.9B规模均保持有效;相比之下,算力匹配的窗口式训练方案无法建立持久的拒绝行为。真正带来攻击鲁棒性的,是安全信号在预训练中的持续性,而非其时机。
cs.LG / 109 / 2609.06942

PCSDiff: Diffusion-Based Bias Correction and Super Resolution Toward Practical Operational Medium-Term Precipitation Forecast

PCSDiff:面向实际业务化中期降水预报的基于扩散模型的偏差校正与超分辨率方法
Sun, Yuze, Wang, Shiyi, Pan, Jiancheng, Wang, Die, Prein, Andreas F., Luo, Wentao, Jiang, Linhan, Wu, Jie, Zhang, Quan, Huang, Xiaomeng
Abstract
Medium-range precipitation forecasts are impaired by persistent systematic biases, lead-time-dependent error accumulation, and coarse spatial resolution, restricting their reliability for flood-drought risk assessment. Existing AI correction techniques lack dedicated modeling for multi-day dynamic bias evolution and proper meteorological constraints, often generating over-smoothed rainfall structures, and cannot meet operational deployment demands. This work introduces PCSDiff, a cascaded task-decoupled diffusion framework targeting 10-day precipitation bias correction and downscaling. To jointly counteract temporal error drifts and reconstruct physically plausible local precipitation details, PCSDiff integrates the Precipitation Intensity-aware Multi-branch Decoder (PIMD) module for dynamic multi-day error mitigation using synoptic-temporal features, followed by a two-phase conditional diffusion super-resolution module to restore fine-scale precipitation patterns. Evaluated against CMA-CRA observations over China after global-data training, PCSDiff cuts RMSE by 16.1% and lifts ACC by 13.9% relative to raw ECMWF forecasts at 3-10-day lead times, and consistently outperforms mainstream deep-learning baselines on both general and extreme-precipitation metrics. Benefiting from a streaming inference pipeline, our method achieves low-latency rolling forecasting for practical meteorological operations.
Chinese Translation
中期降水预报受持续性系统偏差、随预报时效增加的误差累积以及粗糙空间分辨率的制约,限制了其在洪旱风险评估中的可靠性。现有的人工智能校正技术缺乏对多日动态偏差演化的专门建模和适当的气象约束,往往生成过度平滑的降水结构,且无法满足业务部署需求。本工作提出了PCSDiff,一个针对10天降水偏差校正与降尺度的级联任务解耦扩散框架。为同时对抗时间误差漂移并重建物理合理的局地降水细节,PCSDiff集成了降水强度感知多分支解码器(PIMD)模块,利用天气-时间特征缓解动态多日误差,随后通过两阶段条件扩散超分辨率模块恢复精细尺度的降水形态。在基于全球数据训练后,针对中国区域CMA-CRA观测数据的评估表明,在3-10天预报时效内,PCSDiff相较原始ECMWF预报将RMSE降低了16.1%,ACC提升了13.9%,并在常规指标和极端降水指标上均持续优于主流深度学习基线方法。得益于流式推理流水线,我们的方法实现了面向实际气象业务的低延迟滚动预报。
cs.LG / 110 / 2609.06947

Particle Dynamics of Flow Matching and Classifier-Free Guidance from a Stagewise Geometry Perspective

从分阶段几何视角研究流匹配与无分类器引导的粒子动力学
Cai, Jian-Feng, Su, Zhengyi, Wang, Chao
Abstract
Flow matching, together with classifier-free guidance (CFG), is widely used in generative modeling, yet much of the theoretical understanding remains distribution-wise. Since practical sampling follows individual trajectories, distribution-level guarantees alone do not fully capture how trajectories interact with the data geometry or how guidance reshapes it. To overcome this limitation, we establish a unified stagewise geometric theory of attraction and absorption for both continuous dynamics and explicit Euler discretization. Specifically, with $t\in[0,1]$ running from noise to data, we show that unconditional flow trajectories are successively attracted toward a neighborhood of the global mean, the data convex hull, and a neighborhood of a possibly nonconvex local cluster. Across these stages, the corresponding distance satisfies a common contraction estimate, yielding an ${O}(1-t)$ decay of the distance in the final stage. For CFG, the same structure persists with an extrapolated mean, an inflated conditional convex hull, and, near the target cluster, the restored local geometry of conditional flow matching. We further show that a general time schedule $a(t)$ replaces the $O(1-t)$ decay by $O(1-a(t))$. Together, these results provide a unified particle-level geometric account of flow matching and CFG across continuous and discrete sampling.
Chinese Translation
流匹配(Flow Matching)连同无分类器引导(Classifier-Free Guidance, CFG)被广泛应用于生成建模,然而目前的理论理解大多停留在分布层面。由于实际采样遵循的是个体轨迹,仅靠分布层面的保证无法充分刻画轨迹如何与数据几何相互作用,以及引导如何重塑数据几何。为克服这一局限,我们针对连续动力学和显式欧拉离散化,建立了一个统一的吸引与吸收的分阶段几何理论。具体而言,当 $t\in[0,1]$ 从噪声推进到数据时,我们证明无条件流轨迹会相继被吸引至全局均值的邻域、数据凸包以及某个可能非凸的局部聚类的邻域。在这些阶段中,相应的距离满足统一的收缩估计,并在最后阶段给出距离的 ${O}(1-t)$ 衰减。对于CFG,同样的结构依然成立,只是变为外推的均值、膨胀的条件凸包,以及在目标聚类附近条件流匹配局部几何的恢复。我们进一步证明,一般的时间调度 $a(t)$ 可将 $O(1-t)$ 的衰减替换为 $O(1-a(t))$。综上,这些结果为连续与离散采样下的流匹配和CFG提供了一个统一的粒子层面几何解释。
cs.LG / 111 / 2609.06951

Steering Interference Reflects the Model's Defaults, Not the Behavior Directions

导向干扰反映的是模型的默认倾向,而非行为方向本身
Malla, Srikanth, Choi, Chiho, Choi, Joon Hee
Abstract
Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alone. It does not. We ask what decides which other behaviors move, and by how much, and find that it is the model rather than the behavior being steered. A steer relaxes the model toward a small set of behaviors it already favors, chiefly refusal, sycophancy, and poeticism, and that set is much the same whatever is steered. Three results across 24 behaviors and ten instruction-tuned models support this, every effect read off the generated text by a language-model judge rather than off a probe. That readout matters: all 24 behaviors are linearly decodable, but only 20 change what the model writes. First, a direction carrying no behavioral content, matched to a real steer only in the size of the vector it adds, moves the same behaviors in the same order as real steers do, while producing none of the behaviors that need a specific direction. Second, most interference runs one way, so it cannot be an overlap between two directions: steering profanity makes the model toxic, while steering toxicity leaves profanity untouched. Third, with a behavior held out entirely, geometry measured on the others explains almost none of the interference it takes part in. The account holds on all ten models, the pull toward defaults strongest below 10B parameters and weakening in each family's largest. Reading a steer as a perturbation whose endpoint the model fixes implies that disentangling behavior directions cannot by itself make steering modular.
Chinese Translation
激活导向(activation steering)承诺对语言模型行为进行模块化控制:某种行为(如礼貌)对应于模型激活空间中的一个方向,在模型生成时加入该方向应当能开启该行为,同时不影响其他一切。但事实并非如此。我们探究了是什么决定了哪些其他行为会发生改变以及改变的程度,发现决定因素是模型本身,而非被导向的行为。导向会使模型松弛地偏向其本已偏好的少量行为集合,主要是拒答(refusal)、谄媚(sycophancy)和诗化倾向(poeticism),而且无论导向什么行为,这一集合都大致相同。覆盖24种行为和十个指令微调模型的三个结果支持了这一结论,所有效应均由语言模型裁判从生成文本中读取,而非通过探针(probe)读取。这种读取方式至关重要:全部24种行为都可线性解码,但只有20种能改变模型的实际输出。第一,一个不含任何行为内容、仅在与真实导向所加向量大小上相匹配的方向,会以与真实导向相同的顺序影响相同的行为,却不会产生任何需要特定方向才能触发的行为。第二,大多数干扰是单向的,因此不能用两个方向的重叠来解释:导向粗俗言语(profanity)会使模型变得有毒(toxic),而导向毒性却不会影响粗俗言语。第三,当某种行为被完全留出(held out)时,在其他行为上测得的几何结构几乎无法解释它所参与的干扰。这一解释在全部十个模型上均成立,其中向默认倾向的拉力在100亿参数以下的模型中最强,并在各模型家族的最大模型中减弱。将导向视为一个由模型决定其终点的扰动,意味着仅仅解开行为方向的纠缠本身并不能使导向实现模块化。
cs.LG / 112 / 2609.06966

MOLE: Detecting Insider Threats in AI Agents

MOLE:检测AI智能体中的内部威胁
Muhamed, Aashiq, Smith, Virginia
Abstract
Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives and agent refusal does not predict completion. MOLE enables comparison of 40 monitors across corpus generators, observability levels, and threats; even the best evaluated monitor in our single-day audit-event comparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improves budget-AUC by 10% over applying it to every account-day at comparable modeled cost.
Chinese Translation
模型失准、提示词注入或操作者滥用可能导致运营前沿实验室账户的AI智能体窃取模型权重、污染训练数据或削弱发布准入机制。现有基准测试并未检验防御者能否在有限的审查预算下,从日常工作中检测出此类活动。我们提出了MOLE,一个开放基准:150个由AI运营的账户在30个工作日内共享9个有状态服务,包含12种威胁以及来自四个模型的8个语料库,总计约200亿词元。在39个智能体模型中,72%完成了大部分被指派的有害目标,且智能体的拒绝行为并不能预测其完成情况。MOLE支持在不同语料库生成器、可观测性级别和威胁类型下对40个监控器进行比较;在我们单日审计事件的比较中,即使表现最好的监控器也会漏掉近一半已造成的危害。MOLE还支持监控器的开发:基于基准的引导搜索可将一个中等水平监控器的性能提升49-64%,而在相当建模成本下,选择性使用更强的监控器比将其应用于每个账户每日可使预算AUC提升10%。
cs.LG / 113 / 2609.06976

HealthLoopQA: A Context-Aware Question Answering Benchmark for Interpreting Wearable Monitoring Data in Diabetes Care

HealthLoopQA:用于解读糖尿病护理中可穿戴监测数据的上下文感知问答基准
Niu, Yuchen, Ma, Yanan, Nandakumar, Srinivasan, Chen, Maolin, Schlegel, Viktor, Wei, Kexin, Cheng, Ling, Bird, Anna, Bharath, Anil Anthony, Lam, Siew-Kei
Abstract
As medical wearables become integrated into daily chronic disease care, effectively interpreting longitudinal monitoring data is essential for patients and clinicians to understand health trends, detect safety-critical events, and make informed decisions. While large language models (LLMs) show promise for transforming this streaming physiological data into personalized health insights, evaluating their reasoning capability and analytical rigor in diverse monitoring tasks remains a fundamental challenge. Existing medical wearable question answering (QA) benchmarks primarily assess short-horizon classification or statistical summaries, largely ignoring the long-term patterns, therapeutic and behavioural contexts, and potential system failures inherent in real-world deployments. To address this, we introduce HealthLoopQA, a comprehensive diagnostic benchmark for evaluating LLM reasoning over continuous diabetes monitoring data. Grounded in a novel taxonomy of eleven atomic reasoning abilities, HealthLoopQA comprises 127 tasks and over 1,500 QA instances spanning process mining, anomaly detection, and prediction over 30-day horizons. To systematically evaluate safety awareness, we complement real-world datasets with a fault-injected simulation testbed modeling diverse device malfunctions and cyber-physical attacks to generate physiologically plausible hazard scenarios. Evaluating state-of-the-art LLMs across prompting and agentic frameworks reveals severe limitations in complex temporal pattern mining. Furthermore, we identify a broader phenomenon of In-context Laziness under long-context prompting, highlighting critical open challenges in deploying LLMs for rigorous long-horizon medical reasoning.
Chinese Translation
随着医疗可穿戴设备日益融入日常慢性病护理,有效解读纵向监测数据对于患者和临床医生理解健康趋势、检测安全关键事件并做出明智决策至关重要。尽管大语言模型(LLM)在将流式生理数据转化为个性化健康洞察方面展现出潜力,但评估其在多样化监测任务中的推理能力和分析严谨性仍是一个根本性挑战。现有的医疗可穿戴问答(QA)基准主要评估短时程分类或统计摘要,在很大程度上忽略了真实部署中长期模式、治疗与行为情境以及潜在系统故障。为此,我们提出了HealthLoopQA,一个用于评估LLM对连续糖尿病监测数据进行推理的综合性诊断基准。HealthLoopQA基于一个由十一种原子推理能力构成的新型分类体系,包含127个任务和超过1,500个问答实例,涵盖过程挖掘、异常检测以及30天时程的预测。为系统性地评估安全意识,我们在真实世界数据集的基础上,补充了一个故障注入仿真测试平台,模拟多种设备故障与信息物理攻击,以生成生理上合理的危险场景。对最先进LLM在提示和智能体框架下的评估显示,其在复杂时序模式挖掘方面存在严重局限。此外,我们在长上下文提示下识别出一种更广泛的现象——上下文惰性(In-context Laziness),凸显了在部署LLM进行严格的长时程医学推理时面临的关键开放性挑战。
cs.LG / 114 / 2609.06984

AF-Mamba: Efficient Long-Term Signal Modeling for Early Prediction of Atrial Fibrillation Onset

AF-Mamba:用于心房颤动发作早期预测的高效长期信号建模
Lee, Yongbin, Chon, Ki H.
Abstract
Atrial fibrillation (AF) is the most common cardiac arrhythmia and is associated with increased risks of stroke and heart failure. The growing availability of wearable and portable ECG monitoring enables continuous assessment of cardiac rhythm outside clinical settings. Predicting AF before its onset could provide additional lead time for timely clinical assessment and potentially improve the management of patients at risk of AF-related complications. This study focuses on predicting AF onset one hour in advance using long-term RR intervals (RRIs). To address this challenge, we propose a deep learning architecture that integrates temporal convolutional networks (TCNs) for local features encoding with Mamba, a selective state-space model capable of long-range sequence modeling. This hybrid TCN-Mamba design enables efficient training and inference on one-hour input windows, overcoming limitations of Transformers' quadratic scaling and recurrent networks' vanishing gradients. In subject-wise 5-fold testing, the proposed model achieved a sensitivity of 0.889, specificity of 0.943, F1-score of 0.813, AUROC of 0.974, and AUPRC of 0.933. In paired cross-dataset holdout evaluation, AF-Mamba maintained discriminative performance across unseen AF and NSR datasets, achieving a mean AUROC of 0.897. Compared against state-of-the-art AF prediction models and general time-series models, AF-Mamba achieved competitive predictive performance while providing a favorable performance-efficiency trade-off for long RRI sequences. These findings demonstrate the potential of AF-Mamba for accurate AF prediction one hour in advance and real-time continuous ambulatory monitoring.
Chinese Translation
心房颤动(AF)是最常见的心律失常,与中风和心力衰竭风险增加相关。可穿戴和便携式心电(ECG)监测设备的日益普及,使得在临床环境之外对心脏节律进行持续评估成为可能。在心房颤动发作之前进行预测,可为及时的临床评估提供额外的提前时间,并有望改善存在房颤相关并发症风险患者的管理。本研究聚焦于利用长期RR间期(RRIs)提前一小时预测心房颤动的发作。为应对这一挑战,我们提出了一种深度学习架构,该架构将用于局部特征编码的时间卷积网络(TCN)与Mamba相结合,Mamba是一种能够进行长程序列建模的选择性状态空间模型。这种TCN-Mamba混合设计能够在一小时输入窗口上实现高效的训练和推理,克服了Transformer的二次方扩展限制和循环网络的梯度消失问题。在按受试者划分的5折测试中,所提出的模型取得了0.889的灵敏度、0.943的特异度、0.813的F1分数、0.974的AUROC和0.933的AUPRC。在配对跨数据集留出评估中,AF-Mamba在未见过的房颤和正常窦性心律(NSR)数据集上保持了良好的判别性能,平均AUROC达到0.897。与最先进的房颤预测模型和通用时间序列模型相比,AF-Mamba取得了具有竞争力的预测性能,同时在长期RRI序列上提供了良好的性能-效率权衡。这些结果表明,AF-Mamba在提前一小时准确预测房颤以及实时连续动态监测方面具有巨大潜力。
cs.LG / 115 / 2609.06986

Continual Learning Mechanisms Compose for Long-Horizon Memorization

持续学习机制的组合实现长时程记忆
Zhang, Zheyuan, Zhang, Alvin, Khashabi, Daniel, Shu, Tianmin
Abstract
Language models may need to internalize information that arrives over time and retain it through many subsequent updates. To study this challenge, we introduce long-horizon memorization, a setting in which a model learns 100 query-answer tasks through continual supervised fine-tuning without retaining earlier training examples or receiving task identifiers at inference. Sequential updates cause catastrophic forgetting, and no single continual learning mechanism we evaluate maintains strong retention at this horizon. We hypothesize that mechanisms addressing complementary sources of forgetting will be more effective when composed. We organize these compositions along two design dimensions. Data, function, and weight anchors specify what prior information each update should preserve, while low-rank allocation rules determine where successive updates are retained. To test this hypothesis systematically, we construct three distinct 100-task memorization datasets. We introduce task-level successive halving to search the combinatorial design space and use a factorial experiment to measure individual and interaction effects. Our best method combines all three anchors with merged LoRA, ranks among the top 3 methods in all datasets, and raises average final retention from 1.2% under naive sequential fine-tuning to 34.9%, a 28-fold improvement. The data anchor and merged LoRA provide the largest average gains and interact super-additively on all three datasets. Together, these results show that composing complementary mechanisms substantially improves long-horizon memorization beyond what any individual mechanism achieves.
Chinese Translation
语言模型可能需要内化随时间到达的信息,并在后续多次更新中予以保留。为研究这一挑战,我们提出了长时程记忆(long-horizon memorization)这一设定:模型通过持续的监督微调学习100个问答任务,过程中不保留早期训练样本,且在推理时不接收任务标识符。顺序更新会导致灾难性遗忘,我们评估的所有单一持续学习机制在该时间跨度下均无法维持较强的记忆保持能力。我们假设,针对遗忘不同互补来源的机制在组合使用时将更加有效。我们从两个设计维度组织这些组合:数据、函数和权重锚点(anchor)规定了每次更新应保留哪些先验信息,而低秩分配规则决定了后续更新的信息存放位置。为系统性检验这一假设,我们构建了三个不同的100任务记忆数据集,引入任务级逐次减半(successive halving)方法搜索这一组合设计空间,并利用因子实验度量各机制的个体效应和交互效应。我们最优的方法将三种锚点与合并式LoRA相结合,在所有数据集中均位列前三,将平均最终记忆保持率从朴素顺序微调的1.2%提升至34.9%,实现了28倍的改进。其中数据锚点和合并式LoRA带来了最大的平均增益,并在所有三个数据集上表现出超加性的交互作用。这些结果共同表明,组合互补的机制能够显著提升长时程记忆能力,其效果远超任何单一机制。
cs.LG / 116 / 2609.07009

NeuCME: Toward Dynamic Multimodal Continual Learning via Neural Combinatorics of Multiple Experts

NeuCME:基于多专家神经组合学的动态多模态持续学习
Guo, Kai, Liu, Chuanbin, Hu, Peng, Wang, Hao, Peng, Xi
Abstract
Multimodal continual learning has recently shown great potential for developing agents with human-like intelligence by continuously learning new tasks across multiple modalities. However, existing methods typically assume that the set of modalities per task is predefined and fixed. In this paper, we investigate a more realistic learning setting, referred to as dynamic multimodal continual learning, in which the set of modalities may vary across tasks rather than remaining fixed. This setting involves two primary challenges: (i) spatio-temporal catastrophic forgetting and (ii) adaptive multimodal fusion. To address these challenges, we propose NeuCME (as shorthand for \textbf{Neu}ral \textbf{C}ombinatorics of \textbf{M}ultiple \textbf{E}xperts), a novel framework designed to effectively learn and integrate knowledge across tasks with varying modalities. The proposed NeuCME model comprises three key components, namely modality-combinational rehearsal, multi-gated mixture-of-experts, and task relevance-guided distillation. Furthermore, we formulate an evaluation metric to quantify the dynamism of task sequences and then set up a comprehensive benchmark with different degrees of dynamism. Extensive experiments using four real-world datasets demonstrate that the proposed NeuCME outperforms state-of-the-art methods markedly.
Chinese Translation
多模态持续学习近年来在通过跨多种模态持续学习新任务来开发具有类人智能的智能体方面展现出巨大潜力。然而,现有方法通常假设每个任务的模态集合是预定义且固定的。本文研究了一种更现实的学习设置,称为动态多模态持续学习(dynamic multimodal continual learning),其中模态集合可能随任务而变化,而非保持固定。该设置涉及两个主要挑战:(i)时空灾难性遗忘;(ii)自适应多模态融合。为解决这些挑战,我们提出了NeuCME(多专家神经组合学,Neural Combinatorics of Multiple Experts 的缩写),这是一个旨在有效学习并整合具有不同模态的任务间知识的新型框架。所提出的NeuCME模型包含三个关键组件,即模态组合重演(modality-combinational rehearsal)、多门控专家混合(multi-gated mixture-of-experts)以及任务相关性引导的蒸馏(task relevance-guided distillation)。此外,我们构建了一个评估指标来量化任务序列的动态性,并据此建立了包含不同动态程度的综合基准。基于四个真实世界数据集的大量实验表明,所提出的NeuCME显著优于当前最先进的方法。
cs.LG / 117 / 2609.07017

HyperTransfer: Understanding the Equivalence between Base Optimizer and Hyperball

HyperTransfer:理解基础优化器与超球优化器之间的等价性
Yuan, Jinghui, Zhang, Hongtao, Zou, Jade, Li, Tianyu, Zhou, Wenjie, He, Tianyu, Chen, Wei
Abstract
Hyperball optimizers constrain parameter norms and update only their directions, establishing a distinct paradigm for neural network optimization. Although this geometry appears fundamentally different from that of conventional Base Optimizers, which update both parameter norms and directions, we show that the two paradigms are dynamically equivalent for scale-invariant networks. Building on this equivalence, we propose HyperTransfer, which constructs a Hyperball optimizer that reproduces the dynamics of a target Base Optimizer using only its initialization and learning-rate schedule, without running the target optimizer itself. We further derive the inverse mapping and extend the framework to non-scale-invariant networks. Experiments show that both HyperTransfer and the inverse mapping produce loss trajectories nearly identical to those of their targets, suggesting that Hyperball dynamics are governed primarily by the induced effective learning-rate schedule and optimizer state.
Chinese Translation
超球(Hyperball)优化器约束参数范数并仅更新其方向,为神经网络优化确立了一种独特的范式。尽管这种几何结构看起来与更新参数范数和方向的传统基础优化器(Base Optimizer)有着根本的不同,我们证明了对于尺度不变网络,这两种范式在动力学上是等价的。基于这一等价性,我们提出了 HyperTransfer,它构建了一个超球优化器,仅利用目标基础优化器的初始化和学习率调度即可复现其动力学,而无需实际运行该目标优化器。我们进一步推导了逆映射,并将该框架扩展到非尺度不变网络。实验表明,HyperTransfer 和逆映射所产生的损失轨迹与目标优化器的损失轨迹几乎完全一致,这表明超球动力学主要由其诱导的有效学习率调度和优化器状态所决定。
cs.LG / 118 / 2609.07031

Efficient Learning and Symmetry Discovery under Exact Invariances

精确不变性下的高效学习与对称性发现
Soleymani, Ashkan, Tahmasebi, Behrooz, Jaillet, Patrick, Jegelka, Stefanie
Abstract
Learning with group invariances is central to many scientific and geometric learning problems, yet its computational foundations remain poorly understood. Even for classical supervised regression settings, it has been unclear whether one can efficiently compute a regression function that is exactly invariant to a given group action. Recent work showed that exact invariance can be enforced in polynomial time when the underlying group is finite and known, but left open the cases of infinite groups and unknown symmetries. In this paper, we resolve both challenges. First, we present the first polynomial-time algorithm for learning with exact group invariances that applies uniformly to finite and infinite groups. The runtime is polynomial in the data dimension and sample size, and independent of the group, while achieving strong generalization guarantees. This provides a computational explanation for the empirical success of invariant and equivariant methods in geometric machine learning and partially answers a recent open question in the literature. Second, we study learning in the symmetry discovery setting, where the invariance group is unknown. Focusing on the subgroup lattice of a finite group, we show that exact symmetries can be identified from data and exploited for learning in polynomial time. For regression over finite-dimensional feature spaces, our algorithm provably recovers the underlying symmetry, matches the minimax-optimal sample complexity of the known-symmetry setting, and runs in time polynomial in the data dimension and sample size. Our analysis relies on tools from random Cayley graphs and expander theory, which may be of independent interest.
Chinese Translation
带群不变性的学习是许多科学和几何学习问题的核心,但其计算基础仍知之甚少。即使在经典的有监督回归设置中,人们也不清楚是否能高效地计算出一个对给定群作用精确不变的回归函数。近期工作表明,当底层群是有限群且已知时,可以在多项式时间内施加精确不变性,但无限群和未知对称性的情形仍未解决。本文同时解决了这两个挑战。首先,我们提出了首个适用于精确群不变性学习的多项式时间算法,该算法统一适用于有限群和无限群。其运行时间关于数据维度和样本量是多项式的,且与群无关,同时具备强泛化保证。这为几何机器学习中不变与等变方法的实证成功提供了计算层面的解释,并部分回答了文献中近期提出的一个公开问题。其次,我们研究了对称性发现设置下的学习问题,即不变群未知的情况。聚焦于有限群的子群格,我们证明了精确对称性可以从数据中识别出来,并在多项式时间内用于学习。对于有限维特征空间上的回归,我们的算法可证明地恢复出底层对称性,达到了已知对称性设置下的极小化极大最优样本复杂度,且运行时间关于数据维度和样本量是多项式的。我们的分析依赖于随机凯莱图(Cayley graph)和扩张图(expander)理论的工具,这些工具本身可能也具有独立的研究价值。
cs.LG / 119 / 2609.07037

Disentangling Steering Vectors

解耦转向向量
Hiramatsu, Takeru, Atarashi, Kyohei, Takeuchi, Koh, Kashima, Hisashi
Abstract
Activation steering has emerged as a lightweight, inference-time approach to control the behavior of Large Language Models (LLMs). However, traditional steering vectors used to intervene in LLMs' activations, such as those derived from the difference-in-means method, tend to entangle multiple semantic and stylistic concepts into a single composite direction, leading to unpredictable steering effects. Our core objective is to disentangle this composite direction into its constituent concepts. To this end, we propose Steering Vector Dissection, a framework to explicitly isolate individual and semantically consistent features from these composite directions. Specifically, we pair positive and negative activations and take their differences to generate a set of instance-level steering vectors, and train a dedicated Sparse Autoencoder (SAE) directly on them. Quantitative evaluations across two datasets, two models, and two intervention depths show that our method yields a set of semantically consistent basis vectors whose steering effects are mutually distinguishable. Furthermore, we show that this disentanglement enables precise control over model behaviors.
Chinese Translation
激活转向(Activation Steering)已成为一种轻量级的推理时方法,用于控制大语言模型(LLM)的行为。然而,用于干预LLM激活的传统转向向量(例如由均值差分法得到的向量)往往将多个语义和风格概念纠缠在单一复合方向中,导致转向效果不可预测。我们的核心目标是将该复合方向解耦为其组成概念。为此,我们提出Steering Vector Dissection框架,从这些复合方向中显式分离出单个且语义一致的特征。具体而言,我们将正负激活配对并取其差异,生成一组实例级转向向量,并在其上直接训练一个专用的稀疏自编码器(Sparse Autoencoder, SAE)。在两个数据集、两个模型和两种干预深度上的定量评估表明,我们的方法得到一组语义一致的基础向量,其转向效果彼此可区分。此外,我们证明这种解耦能够实现对模型行为的精确控制。
cs.LG / 120 / 2609.07046

AI and TCAD for Inverse Design and Defect Discovery: From Simple Machine Learning to LLM

AI与TCAD在逆向设计与缺陷发现中的应用:从简单机器学习到大语言模型
Wong, Hiu Yung
Abstract
AI has revolutionized various engineering domains, but its impact on semiconductor device design and defect discovery is still limited, due to limited data and the curse of dimensionality. In this paper, we will discuss our work on using the Technology Computer-Aided-Design (TCAD) to generate precise data needed for machine learning (ML) to enable simulation-augmented ML. We demonstrate that with minimal domain expertise, it is possible to create a machine that performs as well as a device engineer on a specific task. We will show that auto-encoder-based machine learning models and noise engineering applied to TCAD data are effective at learning latent physics, and that the models can be seamlessly applied to experimental data. We will demonstrate how to build a device-engineer-level model step by step through various examples, including using only non-destructive electrical data to inverse-engineer the PiN diode layer thickness variations, the Ga2O3 Schottky diode doping and anode workfunction variations, and the transistor contact resistance in an inverter. Examples also include the generation of a FinFET IV/CV prediction model, the mapping between transistor images and IV curves, and the automatic calibration of TCAD parameters for a Ga2O3 Schottky diode, which can only be handled well by experienced TCAD engineers. Finally, to fully realize the potential of AI, large language models (LLMs) and multimodal LLMs (MLLMs) are believed to be necessary. We will discuss the application of LLMs to TCAD command file creation and our vision for MLLMs in automated device design and defect discovery.
Chinese Translation
人工智能(AI)已在众多工程领域引发变革,但受限于数据匮乏和维数灾难,其对半导体器件设计与缺陷发现的影响仍然有限。本文将讨论我们利用工艺计算机辅助设计(Technology Computer-Aided-Design, TCAD)生成机器学习(ML)所需精确数据,从而实现仿真增强型机器学习的研究工作。我们证明,仅需极少的专业领域知识,就有可能构建出在特定任务上表现与器件工程师相当的机器。我们将展示基于自编码器(auto-encoder)的机器学习模型以及应用于TCAD数据的噪声工程能够有效学习潜在物理规律,且这些模型可以无缝应用于实验数据。我们通过多个示例演示如何逐步构建器件工程师水平的模型,包括:仅利用无损电学数据逆向推断PiN二极管层厚度变化、Ga2O3肖特基二极管的掺杂与阳极功函数变化,以及反相器中晶体管的接触电阻。示例还包括FinFET IV/CV特性预测模型的构建、晶体管显微图像与IV曲线之间的映射,以及Ga2O3肖特基二极管TCAD参数的自动校准——后者通常只有经验丰富的TCAD工程师才能胜任。最后,为充分释放AI的潜力,我们认为大语言模型(LLM)和多模态大语言模型(MLLM)是必要的。我们将讨论LLM在TCAD命令文件生成方面的应用,以及我们对MLLM在自动化器件设计与缺陷发现中应用前景的展望。
cs.LG / 121 / 2609.07051

TrojanWorld: Backdooring World-Model Agents via Imagination Steering

TrojanWorld:通过想象引导对世界模型智能体实施后门攻击
Huang, Wenkai, Liang, Siyuan, Li, Gaolei, Li, Yiming, Peng, Tianhao, Li, Jianhua, Tao, Dacheng
Abstract
World models increasingly serve as the predictive core of model-based reinforcement learning agents, enabling them to simulate future dynamics and reason over imagined trajectories before acting. Their substantial training demands make pretrained world models attractive for distribution and reuse, exposing downstream systems to model supply chain threats. Backdoor attacks offer a targeted and stealthy means of exploiting such supply chains, yet their threat to interactive world-model agents remains largely unexplored. To fill this gap, we present TrojanWorld, a backdoor framework for world-model agents that induces attacker-specified behavior by steering internal imagination. A physical object placed in the scene acts as the trigger, enabling deployment-time activation through the agent's native observation pipeline without digitally manipulating the observation stream. To achieve effective, stealthy, and persistent control, TrojanWorld combines Decision-Reflective Induction to steer trigger-conditioned imagination toward attacker-specified actions using decision feedback, Clean Behavior Anchoring to preserve trigger-free predictive and behavioral fidelity, and Causal Propagation to sustain the induced preference along subsequent trajectories after the trigger disappears. Together, these mechanisms establish an end-to-end attack chain from physical perception through corrupted imagination to malicious action selection. Experiments with the TD-MPC2, DreamerV3, and R2-Dreamer systems across the DeepMind Control, MetaWorld, MyoSuite, and RoboDesk benchmarks show that under trigger activation, TrojanWorld achieves a target-action deviation as low as 0.026 while retaining at least 98.8% of the corresponding clean performance. Even after trigger removal, the compromised agent can remain trapped in the induced behavioral trajectory, continuing to execute attacker-specified actions.
Chinese Translation
世界模型日益成为基于模型的强化学习智能体的预测核心,使其能够模拟未来动态,并在采取行动之前对想象的轨迹进行推理。世界模型的训练需求巨大,这使得预训练世界模型在分发和复用方面颇具吸引力,从而使下游系统面临模型供应链威胁。后门攻击提供了一种有针对性且隐蔽的利用此类供应链的手段,但其对交互式世界模型智能体的威胁在很大程度上仍未被探索。为填补这一空白,我们提出了TrojanWorld,一个面向世界模型智能体的后门攻击框架,它通过引导内部想象来诱发攻击者指定的行为。放置在场景中的物理对象作为触发器,使得攻击可以在部署时通过智能体原生的观测管道被激活,而无需对观测数据流进行数字篡改。为实现有效、隐蔽且持久的控制,TrojanWorld结合了三种机制:决策反思归纳(Decision-Reflective Induction),利用决策反馈将触发条件下的想象引导至攻击者指定的动作;干净行为锚定(Clean Behavior Anchoring),以保持无触发情况下预测和行为的保真度;因果传播(Causal Propagation),在触发器消失后沿后续轨迹维持被诱发的偏好。这些机制共同构建了从物理感知、经被污染的想象、到恶意动作选择的端到端攻击链。在DeepMind Control、MetaWorld、MyoSuite和RoboDesk基准上,基于TD-MPC2、DreamerV3和R2-Dreamer系统的实验表明,在触发器激活时,TrojanWorld的目标动作偏差可低至0.026,同时保留至少98.8%的相应干净性能。即使在触发器移除之后,被攻陷的智能体仍可能被困于被诱发的行为轨迹中,继续执行攻击者指定的动作。
cs.LG / 122 / 2609.07061

PhysSAE: Mechanistic Interpretability with Sparse Autoencoders

PhysSAE:基于稀疏自编码器的物理信息神经网络的机制可解释性研究
Patil, Nandita N., A., Eshwar R., Honnavar, Gajanan V.
Abstract
Physics-Informed Neural Networks (PINNs) embed PDE residuals into neural network training, but their internal representations remain opaque: it is unknown what physical features their hidden layers encode or whether those features have a localized causal role. We present PhysSAE, a mechanistic interpretability framework that trains overcomplete sparse autoencoders (SAEs) on PINN penultimate-layer activations and evaluates dictionary atoms through direct causal intervention in the original frozen hidden state: $h_{\mathrm{cf}} = h - \alpha z_k d_k$, bypassing the SAE decoder entirely. Across six PDE families, with 3 PINN seeds and 3 SAE seeds each---we show that (i) Our discovered SAE atoms align with independently-defined physical observables (max Pearson $|r|=0.951$, always $\gg$ permutation null), (ii) the causal footprint of top-aligned atom ablation is 1.2--4.2$\times$ more spatially concentrated canonical than PCA or ICA interventions, and (iii) top-aligned atoms outperform matched random controls on causal localization for structured physical concepts (ESF$_{80}$ advantage 0.04-0.44). Two-atom bilateral representations improve concept regression R$^2$ by $\Delta R^2\!=\!0.05\text{-}0.15$ over single atoms, while random pairs decrease it by up to 0.60. These results demonstrate that PINNs develop sparse, physically structured latent representations that can be identified and causally interrogated post-hoc, opening a path toward interpretability-aware scientific machine learning.
Chinese Translation
物理信息神经网络(PINNs)将偏微分方程(PDE)残差嵌入神经网络训练中,但其内部表示仍然是不透明的:我们不知道其隐藏层编码了哪些物理特征,也不清楚这些特征是否具有局部化的因果作用。我们提出了PhysSAE,一个机制可解释性框架,该框架在PINN倒数第二层激活上训练过完备稀疏自编码器(SAE),并通过对原始冻结隐藏状态进行直接因果干预来评估字典原子:$h_{\mathrm{cf}} = h - \alpha z_k d_k$,完全绕过SAE解码器。我们在六个PDE族上进行实验,每个族使用3个PINN随机种子和3个SAE随机种子,结果表明:(i) 我们发现的SAE原子与独立定义的物理观测量相一致(最大Pearson相关系数 $|r|=0.951$,始终远大于置换零假设);(ii) 对齐程度最高的原子消融的因果足迹在空间上比PCA或ICA干预更集中1.2至4.2倍;(iii) 对于结构化物理概念,对齐程度最高的原子在因果定位上优于匹配的随机对照(ESF$_{80}$优势为0.04-0.44)。双原子双向表示相比单原子将概念回归R²提高了 $\Delta R^2\!=\!0.05\text{-}0.15$,而随机原子对则使其降低最多0.60。这些结果表明,PINN会发展出稀疏的、具有物理结构的潜在表示,这些表示可以在事后被识别并进行因果探究,为可解释性感知的科学机器学习开辟了一条道路。
cs.LG / 123 / 2609.07063

Trust-But-Verify: Poisoning-Resilient Locally Private Graph Learning Protocols

信任但需验证:具有抗投毒能力的本地私有图学习协议
He, Longzhu, Sun, Li, Peng, Hao, Wang, Ruijie, Wong, Raymond Chi-Wing, Su, Sen
Abstract
Built upon local differential privacy (LDP), locally private graph learning protocols have emerged as an important paradigm for decentralized graph learning, balancing privacy protection and learning utility. Under such protocols, each user locally perturbs their node features and adjacency information before transmission, ensuring formal privacy guarantees without original data leaving the device. However, the inherently open participation nature renders these protocols critically vulnerable to data poisoning attacks, where adversaries inject carefully crafted malicious nodes to corrupt neighborhood aggregation and degrade downstream utility. Despite the severity of this threat, effective defenses in this setting remain largely unexplored. In this paper, we propose VERITAS, a poisoning-resilient locally private graph learning protocol built on a trust-but-verify paradigm. By introducing a verification list encoding graded peer trust levels, VERITAS jointly privatizes node features and graph structure on the user side, while exploiting bilateral attestation asymmetry on the server side to identify and prune malicious nodes. Concretely, VERITAS comprises four synergistic stages: (1) local data perturbation, (2) attestation-driven malicious node pruning, (3) utility restoration via dual denoising, and (4) robust private graph learning. Extensive experiments on four real-world benchmark datasets across multiple LDP mechanisms and GNN architectures demonstrate that VERITAS effectively defends against data poisoning attacks and significantly improves downstream graph learning utility under rigorous privacy guarantees.
Chinese Translation
基于本地差分隐私(Local Differential Privacy, LDP)的本地私有图学习协议已成为去中心化图学习的重要范式,在隐私保护与学习效用之间实现了平衡。在此类协议下,每个用户在传输前对节点特征和邻接信息进行本地扰动,在原始数据不离开设备的情况下确保形式化的隐私保证。然而,其固有的开放式参与特性使这些协议极易受到数据投毒攻击——攻击者通过注入精心构造的恶意节点来破坏邻域聚合,从而降低下游任务的效用。尽管该威胁十分严重,但针对这一场景的有效防御仍鲜有研究。本文提出了 VERITAS,一个基于"信任但需验证"(trust-but-verify)范式的抗投毒本地私有图学习协议。通过引入编码分级同伴信任等级的验证列表,VERITAS 在用户端对节点特征和图结构进行联合私有化,同时在服务器端利用双向证明的不对称性来识别并剔除恶意节点。具体而言,VERITAS 包含四个协同阶段:(1)本地数据扰动;(2)基于证明驱动的恶意节点剔除;(3)通过双重去噪恢复效用;(4)鲁棒的私有图学习。在四个真实世界基准数据集上、跨多种 LDP 机制和 GNN 架构的大量实验表明,VERITAS 能够有效抵御数据投毒攻击,并在严格的隐私保证下显著提升下游图学习效用。
cs.LG / 124 / 2609.07086

Conditioned Initialization for Attention

面向注意力机制的条件化初始化
Saratchandran, Hemanth, Lucey, Simon
Abstract
Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. At the core of their success lies the attention layer, where the query, key, and value matrices determine how token dependencies are captured. While considerable work has focused on scaling and optimizing Transformers, comparatively little attention has been paid to how the weights of the queries, keys and values are initialized. Common practice relies on random initialization or alternatives such as mimetic initialization, which imitates weight patterns from converged models, and weight selection, which transfers weights from a teacher model. In this paper, we argue that initialization can introduce an optimization bias that fundamentally shapes training dynamics. We propose conditioned initialization, a principled scheme that initializes attention weights to improve the spectral properties of the attention layer. Theoretically, we show that conditioned initialization can potentially reduce the condition number of the attention Jacobian, leading to more stable optimization. Empirically, it accelerates convergence and improves generalization across diverse applications, highlighting conditioning as a critical yet underexplored area for advancing Transformer performance. Importantly, conditioned initialization is simple to apply and integrates seamlessly into a wide range of Transformer architectures.
Chinese Translation
Transformer(变换器)是现代机器学习中的主导架构,支撑着视觉、语言及更广泛领域的各类应用。其成功的核心在于注意力层,其中查询、键和值矩阵决定了如何捕捉词元之间的依赖关系。尽管已有大量工作致力于Transformer的扩展与优化,但对于查询、键和值权重的初始化方式却相对缺乏关注。常见的做法依赖于随机初始化或其他替代方案,例如模仿收敛模型权重模式的模仿初始化(mimetic initialization),以及从教师模型迁移权重的权重选择(weight selection)。在本文中,我们提出初始化可以引入一种优化偏差,从根本上塑造训练动力学。我们提出条件化初始化(conditioned initialization),这是一种有原则的初始化方案,通过对注意力权重进行初始化来改善注意力层的谱性质。从理论上,我们证明条件化初始化能够潜在地降低注意力雅可比矩阵的条件数,从而带来更稳定的优化。从实证上看,该方法在多种应用中加速了收敛并提升了泛化能力,凸显了条件数控制作为提升Transformer性能的一个关键但尚未被充分探索的领域。重要的是,条件化初始化实施简单,并且可以无缝集成到广泛的Transformer架构中。
cs.LG / 125 / 2609.07100

Temporal Heterogeneous Graph Transformer for Credit Card Fraud Detection

用于信用卡欺诈检测的时序异质图Transformer
Yan, Qinwen
Abstract
Credit card fraud detection typically relies on tabular features, while repeated attributes can also provide useful relational signals. This paper proposes THGT-FD, a Temporal Heterogeneous Graph Transformer for Fraud Detection. Each transaction is represented using one transaction token and six types of relation tokens and incorporates Time2Vec encoding into the transaction representation. A Transformer learns the interactions among these tokens within each individual transaction and then outputs a fraud probability. Experiments were conducted on 150,000 transactions sampled from the IEEE-CIS Fraud Detection dataset and chronologically partitioned according to TransactionDT. On the test set, THGT-FD achieved an AUC-ROC of 0.8536, an average precision of 0.4164, and a Recall@5% of 0.4708. The class-weighted histogram-based gradient-boosting baseline achieved an AUC-ROC of 0.8722. The results indicate that relation tokens provide useful information for fraud-risk ranking, although the current model does not yet incorporate entity-level historical aggregation.
Chinese Translation
信用卡欺诈检测通常依赖于表格特征,而重复出现的属性也可以提供有用的关系信号。本文提出了THGT-FD,一种用于欺诈检测的时序异质图Transformer(Temporal Heterogeneous Graph Transformer for Fraud Detection)。每笔交易由一个交易令牌(transaction token)和六种类型的关联令牌(relation token)表示,并将Time2Vec编码融入交易表示中。Transformer学习每笔交易内部这些令牌之间的交互,随后输出欺诈概率。实验在从IEEE-CIS欺诈检测数据集中采样的150,000笔交易上进行,并按照TransactionDT进行时序划分。在测试集上,THGT-FD取得了0.8536的AUC-ROC、0.4164的平均精度(average precision)以及0.4708的Recall@5%。基于直方图的类别加权梯度提升基线模型取得了0.8722的AUC-ROC。结果表明,关联令牌为欺诈风险排序提供了有用的信息,但当前模型尚未纳入实体级的历史聚合信息。
cs.LG / 126 / 2609.07108

Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training

面向大规模、长上下文强化学习后训练中投机解码的在线草稿模型协同训练
Wang, Zili, Qiu, Zhaopeng, Zhang, Yuekai, Yu, Shuang, Lai, Junjie
Abstract
Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-training. Online co-training can further increase the draft's accuracy, yielding greater speedups. However, scaling this approach to co-training on large models with long contexts poses two obstacles: (1) branch attention is unsupported by standard causal context-parallel (CP) implementations, and (2) target features span across pipeline-parallel (PP) stages. We address both with an end-to-end system for large-scale online draft co-training. For CP, we extend packed, load-balanced zigzag ring attention by merging rank-local branch attention with causal main-sequence attention. For PP, TapChannel transports intermediate target features across stages via a separate path, leaving the pipeline schedule unaffected. Experiments demonstrate that co-trained drafts closely track the policy baseline while delivering substantial rollout and end-to-end speedups across model scales up to 122B. Our CP design achieves strong scaling at 256K tokens with significant memory savings over prior work, and our PP transport incurs modest overhead. Code can be found at https://github.com/NVIDIA-NeMo/RL/issues/3698.
Chinese Translation
投机解码(speculative decoding)可加速 rollout 生成,而 rollout 生成本身是强化学习(RL)后训练成本的主要部分。在线协同训练可进一步提高草稿模型(draft)的准确率,从而带来更大的加速比。然而,将该方法扩展到长上下文大模型的协同训练面临两个障碍:(1)标准因果上下文并行(context-parallel,CP)实现不支持分支注意力;(2)目标模型特征跨越流水线并行(pipeline-parallel,PP)的多个阶段。我们通过一个端到端系统解决了这两个问题,实现了大规模在线草稿模型协同训练。在 CP 方面,我们扩展了打包的、负载均衡的之字形环形注意力(zigzag ring attention),将秩局部分支注意力与因果主序列注意力相融合。在 PP 方面,TapChannel 通过独立通道在各阶段间传输中间目标模型特征,且不影响流水线调度。实验表明,协同训练的草稿模型能够紧密跟踪策略基线,同时在高达 122B 参数的模型规模上带来显著的 rollout 与端到端加速。我们的 CP 设计在 256K token 长度下实现了良好的扩展性,且相比先前工作显著节省内存;我们的 PP 传输开销较小。代码见 https://github.com/NVIDIA-NeMo/RL/issues/3698。
cs.LG / 127 / 2609.07147

Fine-grained Distributed Backdoor Attacks in Federated Learning

联邦学习中的细粒度分布式后门攻击
Wang, Jian, Shen, Hong, Ke, Wei, Liu, Xue Hua
Abstract
Federated learning, as a privacy-preserving distributed machine learning paradigm, faces significant threats from backdoor attacks. Compared to centralized attacks, distributed backdoor attacks are more harmful but require more poisoned samples to compensate for the loss of trigger strength due to decomposition. Fixed trigger patterns are also easily detected by robust aggregation algorithms, increasing the risk of attack exposure. To address these challenges, we propose a fine-grained distributed backdoor attack framework (FDBA). This framework uses dynamic trigger generation and embedding vector optimization to perform attacks with fewer poisoned samples. First, we design a dynamic trigger generation method based on image edge structures using the Canny algorithm to extract edge features, which are then injected with Laplacian noise. RGB channel decomposition is applied for covert adaptation of the distributed trigger, reducing detection chances. Second, we introduce an embedding vector contrastive learning strategy that forces poisoned samples to approach the target class center in the feature space, enhancing attack effectiveness. On CIFAR-10, piecewise-linear estimates for target ASRs between 70\% and 90\% show that FDBA reduces the required poisoning ratio by 37.4\%--48.4\% compared with DBA. In non-independent and identically distributed (Non-IID) scenarios, FDBA retains 84.7\% of its IID attack performance under extreme heterogeneity, whereas DBA drops to 73.5\%, and the framework successfully bypasses mainstream defense mechanisms. This study offers new insights into federated learning security and emphasizes the potential threats and defense challenges posed by fine-grained distributed attacks.
Chinese Translation
联邦学习作为一种保护隐私的分布式机器学习范式,面临着后门攻击的严重威胁。与集中式攻击相比,分布式后门攻击危害更大,但需要更多的中毒样本来弥补因分解而导致的触发器强度损失。此外,固定的触发器模式容易被鲁棒聚合算法检测到,增加了攻击暴露的风险。为应对这些挑战,我们提出了一种细粒度分布式后门攻击框架(FDBA)。该框架利用动态触发器生成和嵌入向量优化,以更少的中毒样本实施攻击。首先,我们设计了一种基于图像边缘结构的动态触发器生成方法,利用Canny算法提取边缘特征,并注入拉普拉斯噪声;随后采用RGB通道分解对分布式触发器进行隐蔽适配,降低被检测的概率。其次,我们引入了一种嵌入向量对比学习策略,迫使中毒样本在特征空间中逼近目标类中心,从而增强攻击效果。在CIFAR-10数据集上,针对70%至90%目标攻击成功率(ASR)的分段线性估计表明,与DBA相比,FDBA将所需中毒比例降低了37.4%至48.4%。在非独立同分布(Non-IID)场景下,FDBA在极端异构性条件下仍保留了84.7%的IID攻击性能,而DBA则降至73.5%,且该框架成功绕过了主流防御机制。本研究为联邦学习安全提供了新的见解,并强调了细粒度分布式攻击所带来的潜在威胁与防御挑战。
cs.LG / 128 / 2609.07148

Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification

Stable-MM-R1:通过熵引导分层锚定多模态推理动态
Ye, Yimeng, Chen, Shuang, Huang, Wenxuan, Zhang, Manyuan, Feng, Kaituo, Chen, Zhangquan, Chen, Jiayu, Zhou, Yucheng, Xiao, Yicheng, Feng, Zhiyuan, Shi, Tianyu
Abstract
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from "Rollout Silencing" and low-quality gradient signals in standard sampling procedures. In this work, we propose a robust, data-centric framework to stabilize RL training. We first introduce Potential-Aware Query Mining (PAQM), which filters data dynamically to focus on the "Distillation Zone"---samples with high potential for capability elicitation. Furthermore, we present Hybrid Stratified Replay (HSR), a novel mechanism that restructures batches by stratifying rollouts based on Path Entropy, a rollout-level confidence proxy, and outcome reward. Within each optimization step, HSR reuses current-policy "Stability Anchors" and "Hard Negatives" to construct high-contrast optimization groups, then clears its buffers before the next step. This approach mitigates entropy collapse while improving the utilization of learning signals under limited compute. Our method outperforms strong baselines on complex reasoning tasks, offering a principled solution for stable and efficient RL fine-tuning.
Chinese Translation
尽管强化学习(Reinforcement Learning, RL)能够有效激励大语言模型的推理能力,但当前的训练流程仍受制于训练不稳定和熵的快速坍缩。这些局限往往源于标准采样过程中的"Rollout Silencing"(回合消音)现象以及低质量的梯度信号。在本工作中,我们提出了一个稳健的、以数据为中心的框架来稳定RL训练。我们首先引入了潜在感知查询挖掘(Potential-Aware Query Mining, PAQM),该方法动态筛选数据,聚焦于"蒸馏区"(Distillation Zone)——即具有高能力激发潜力的样本。此外,我们提出了混合分层回放(Hybrid Stratified Replay, HSR),这是一种新颖的机制,它基于路径熵(Path Entropy,一种回合级别的置信度代理指标)和结果奖励对回合进行分层,从而重构批次。在每个优化步骤中,HSR复用当前策略的"稳定性锚点"(Stability Anchors)和"困难负样本"(Hard Negatives)来构建高对比度的优化组,并在下一步开始前清空其缓冲区。该方法在缓解熵坍缩的同时,提高了有限算力下学习信号的利用率。我们的方法在复杂推理任务上优于强大的基线模型,为稳定且高效的RL微调提供了一种有原则的解决方案。
cs.LG / 129 / 2609.07162

The Oversight Gap: What LLM Safety Monitors Miss, and Why It Is Not Capability

监督鸿沟:LLM 安全监控器遗漏了什么,以及为何问题不在于能力
Xu, Xin
Abstract
Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hyperproperties, witnessed only by two executions. The standard consequence is a binary impossibility: one trace cannot decide them. We replace the binary with a measurement. A tight bound puts the balanced accuracy of any single-trace monitor at $\tfrac12+\tfrac12\,TV(P_0,P_1)$, turning undecidability into a graded detectability frontier and defining an oversight gap: a monitor's shortfall below it. On a leak family with closed-form $TV$, nine LLM monitors are optimal at $TV=0$ but capture little signal as $TV$ grows; at $TV=1$, where a 20-line membership check scores $100\%$, they average $60.9\%$. That shortfall is mostly not capability: naming what to check closes $61\%$ of it while leaving the $TV=0$ control at chance. The same split runs through a $2{\times}2$ factorial: an imagined second run leaves monitors at chance ($50.4\%$) while the same rule on an executed second run reaches $90.0\%$, and a stored oracle without a comparison procedure yields only $68.2\%$. Information and procedure are each necessary and neither is capability. Under nondeterminism, replay tracks a closed-form $k$-replay curve only under the right projection, and a projection frontier shows the resulting dilemma is forced: narrow misses $98.6\%$ of off-channel leaks, broad flags $75.7\%$ of clean traffic, and attainable accuracy decays like $1/(qm)$ in the benign-variation rate and the channel count. Finally, two frontier LLM judges certified an earlier version of our own benchmark as sound while a sign test found a directional bias ($p=2.7\times10^{-5}$) that invalidated three of our findings. Construction validity for hyperproperty benchmarks should be proved mechanically, not audited by models.
Chinese Translation
安全监控器被要求认证的若干属性,其中包括跨租户非干扰性(cross-tenant noninterference)、沙袋行为(sandbagging)与评估意识(evaluation awareness),属于 2-安全超属性(2-safety hyperproperties),只能通过两次执行共同体现。其标准推论是一个二元不可能性:单条轨迹无法判定这些属性。我们将这种二元判定替换为一种度量。一个紧界表明,任何单轨迹监控器的平衡准确率至多为 $\tfrac12+\tfrac12\,TV(P_0,P_1)$,这将不可判定性转化为分级的可检测性边界,并由此定义了监督鸿沟(oversight gap):即监控器相对于该边界的差距。在一个具有闭式总变差(TV)的泄露族上,九个 LLM 监控器在 $TV=0$ 时达到最优,但随着 $TV$ 增大几乎无法捕获信号;在 $TV=1$ 处,一段 20 行的成员检查代码可取得 $100\%$ 的成绩,而这些监控器平均仅为 $60.9\%$。这一差距主要并非源于能力不足:明确指出应检查什么即可弥合其中 $61\%$,同时 $TV=0$ 的对照仍保持随机水平。同样的分化也体现在一个 $2{\times}2$ 因子实验中:想象中的第二次运行使监控器停留在随机水平($50.4\%$),而同一规则作用于实际执行的第二次运行则达到 $90.0\%$;拥有存储的预言机(oracle)但没有比较程序时,仅达到 $68.2\%$。信息与程序各自都是必要的,但两者都不是能力问题。在非确定性条件下,只有在正确的投影下,重放才能跟随闭式的 $k$-重放曲线;而投影边界表明由此产生的困境是不可避免的:窄投影会漏掉 $98.6\%$ 的带外泄露,宽投影则会对 $75.7\%$ 的正常流量报警,且可达到的准确率随良性变化率和信道数量以 $1/(qm)$ 的速度衰减。最后,两个前沿 LLM 评审者曾认证我们基准测试的早期版本为合理,而符号检验却发现了一个方向性偏差($p=2.7\times10^{-5}$),该偏差推翻了我们的三个结论。超属性基准的构建效度应当通过机械化方式加以证明,而非依赖模型审计。
cs.LG / 130 / 2609.07192

FedRAW: Preserving Rare-Label Influence in Asynchronous Federated Learning

FedRAW:在异步联邦学习中保留稀有标签的影响力
Bajpai, Prashant, Saxena, Divya, Lalanda, Philippe, Vega, German
Abstract
Asynchronous federated learning improves scalability by updating the global model from a server-side buffer of client updates as they arrive, rather than waiting for all selected clients to finish. While efficient, this arrival-driven aggregation can silently distort representation learning under heterogeneous participation. We identify silent rarity failure, a hidden failure mode in which clients holding rare labels contribute too weakly to the global model even though its overall accuracy appears largely unaffected. This failure arises from two coupled effects: rare-label clients may submit updates less frequently when they are slower or less available, creating participation bias; and once their updates enter the buffer, standard asynchronous aggregation assigns them no compensating influence, creating aggregation bias. We propose FedRAW, a fully server-side aggregation method that preserves rare-label influence without changing local training, client objectives, or communication protocols. FedRAW combines client-level update deduplication, which prevents frequently arriving clients from repeatedly dominating the update buffer, with rare-label-aware weighting, which increases the influence of clients carrying low-coverage labels. We formalize silent rarity failure through participation and aggregation bias, and show that FedRAW increases rare-label client influence over uniform aggregation while preserving convergence. Across EMNIST Balanced, CIFAR-10, HAM10000, and ISIC-2019, FedRAW improves rarelabel accuracy while preserving comparable global accuracy and adding negligible server-side computation.
Chinese Translation
异步联邦学习通过服务器端缓冲区在客户端更新到达时立即更新全局模型,而无需等待所有被选中的客户端完成,从而提升了可扩展性。然而,这种以到达驱动的聚合方式虽然在效率上具有优势,却可能在异构参与的情况下悄然扭曲表示学习。我们识别出一种名为“隐性稀有性失效”(silent rarity failure)的隐藏失效模式:持有稀有标签的客户端对全局模型的贡献过弱,而模型的整体准确率却看起来基本未受影响。这种失效源于两个相互耦合的效应:其一,稀有标签客户端可能因速度较慢或可用性较低而较少提交更新,造成参与偏差;其二,即使其更新进入缓冲区,标准异步聚合也不会赋予其补偿性的影响力,造成聚合偏差。我们提出了FedRAW,一种完全在服务器端执行的聚合方法,它在不改变本地训练、客户端目标或通信协议的前提下保留稀有标签的影响力。FedRAW结合了客户端级更新去重(防止频繁到达的客户端反复主导更新缓冲区)与稀有标签感知加权(提高携带低覆盖度标签客户端的影响力)。我们通过参与偏差和聚合偏差对隐性稀有性失效进行了形式化,并证明FedRAW在保持收敛性的同时,相比均匀聚合提升了稀有标签客户端的影响力。在EMNIST Balanced、CIFAR-10、HAM10000和ISIC-2019数据集上的实验表明,FedRAW在提升稀有标签准确率的同时,保持了相当的全局准确率,且服务器端计算开销可以忽略不计。
cs.LG / 131 / 2609.07199

Protocol effects on feature-based hardware-Trojan detection across Trust-Hub families

评估协议对基于特征的硬件木马检测在Trust-Hub各电路系列上的影响
Xiao, Hang, Xu, Chuhong, Zhou, Kainan, Qian, Gangzhen, Yi, Lu
Abstract
Trust-Hub reuses host circuits: several files differ mainly in the inserted Trojan. When gates from sibling variants enter both training and test folds, a detector can benefit from host logic it has already seen. We measure that effect instead of proposing another classifier. The corpus contains 49,124 gates from 16 netlists grouped into five host families. We left the parser, 36 gate features, class weighting, model settings, threshold, and family-level aggregation unchanged and altered one choice: the test boundary. The three settings draw test gates from the pooled corpus, withhold a complete netlist, or withhold every variant of one host. The choice matters. Random forest records F1/AP of 0.914/0.978 with pooled gates, 0.636/0.851 with one netlist held out, and 0.460/0.577 with a host family held out. XGBoost falls from 0.946/0.976 to 0.464/0.544 across the same comparison. Logistic regression loses AP, although its fixed-threshold F1 is not monotonic. Each family shows the same pooled-to-family direction. Feature removal, repeated model and simulator seeds, score normalization, parser-related exclusions, and a smaller sample change the size of the gap without reversing it. Aggregation also matters: a gate-weighted average is dominated by the larger ISCAS files, so the headline values give each host family one vote. Bootstrap and jackknife summaries keep the gap positive, but their folds reuse training families. We treat the five family rows as descriptive evidence rather than independent trials. Five host families are too few for a population claim, and the experiment says nothing about transfer to a new cell library or an industrial design. It supports a narrower conclusion: sibling benchmark variants can inflate apparent transfer. Benchmarks with several variants of one host circuit should report family-aware holdouts and all five family results beside pooled scores.
Chinese Translation
Trust-Hub重复使用宿主电路:若干文件之间的主要区别仅在于所插入的木马。当同一宿主的兄弟变体中的门电路同时进入训练集和测试集时,检测器可能因已见过宿主逻辑而获益。本文旨在度量这一效应,而非提出新的分类器。语料库包含来自16个网表、归入五个宿主系列的49,124个门电路。我们保持解析器、36个门级特征、类别权重、模型设置、阈值以及系列级聚合方式不变,仅改变一项设置:测试边界。三种设置分别从合并语料库中抽取测试门电路、留出完整网表、或留出某一宿主的全部变体。这一选择影响重大:随机森林在合并门电路下的F1/AP为0.914/0.978,留出一个网表时降至0.636/0.851,留出一个宿主系列时进一步降至0.460/0.577。XGBoost在相同对比下从0.946/0.976降至0.464/0.544。逻辑回归损失了AP值,但其固定阈值下的F1并非单调变化。每个系列均呈现从合并到系列留出方向的一致趋势。特征移除、重复的模型与仿真随机种子、分数归一化、解析器相关的排除以及更小的样本量会改变差距的大小,但不会逆转其方向。聚合方式同样重要:按门数加权的平均会被较大的ISCAS文件主导,因此主要指标为每个宿主系列赋予一票。Bootstrap与jackknife摘要显示差距保持为正,但其折次复用了训练系列。我们将五个系列的结果视为描述性证据,而非独立试验。五个宿主系列不足以支撑总体性结论,且本实验无法说明向新单元库或工业设计的迁移能力。它支持一个更为有限的结论:兄弟基准变体可能夸大表观的迁移性能。包含同一宿主电路多个变体的基准测试,应报告系列感知的留出结果,并在合并分数之外同时报告全部五个系列的结果。
cs.LG / 132 / 2609.07206

REFINE: Trajectory Representation Learning via Closed-Loop Transcription -- Extended Version

REFINE:基于闭环转录的轨迹表示学习——扩展版
Yang, Sean Bin, Sun, Ying, Hu, Jilin, Xu, Zongyi, Torp, Kristian, Lu, Hua, Yang, Bin, Jensen, Christian S.
Abstract
Trajectory representation learning underpins a wide range of trajectory analytics tasks; however, most existing self-supervised approaches, whether discriminative or generative, adopt an open-loop paradigm, relying on fixed data augmentations or random masking without feedback, which limits their ability to generalize and scale. We propose REFINE, a simple yet effective Representation lEarning Framework vIa closed-loop traNscription rEfinement for trajectory data. Drawing upon feedback control theory, REFINE tightly couples road-network-aware generative reconstruction with feedback-driven contrastive learning, enabling the model to capture fine-grained local movement semantics and global spatio-temporal dependencies without manually designed augmentation views. We further provide a control-theoretic analysis that establishes convergence guarantees for the proposed closed-loop optimization. Extensive experiments on four real-world datasets demonstrate that REFINE consistently outperforms state-of-the-art methods across multiple downstream tasks while remaining computationally efficient and scalable. This paper is an extended version of REFINE: Trajectory Representation Learning via Closed-Loop Transcription, to appear in KDD 2026.
Chinese Translation
轨迹表示学习是众多轨迹分析任务的基础;然而,现有的大多数自监督方法,无论是判别式还是生成式,均采用开环范式,依赖于固定的数据增强或无反馈的随机掩码,这限制了其泛化和扩展能力。我们提出了REFINE,一个简单而有效的基于闭环转录精炼的轨迹数据表示学习框架(Representation lEarning Framework vIa closed-loop traNscription rEfinement)。借鉴反馈控制理论,REFINE将道路网络感知的生成式重建与反馈驱动的对比学习紧密耦合,使模型无需人工设计的增强视图即可捕捉细粒度的局部运动语义和全局时空依赖关系。我们进一步提供了控制理论分析,为所提出的闭环优化建立了收敛性保证。在四个真实数据集上的大量实验表明,REFINE在多个下游任务中始终优于最先进的方法,同时保持计算高效和可扩展性。本文是《REFINE: Trajectory Representation Learning via Closed-Loop Transcription》的扩展版本,该文将发表于KDD 2026。
cs.LG / 133 / 2609.07230

Robust Decentralized Federated Distillation via Multi-Modality Knowledge Collaboration

基于多模态知识协同的鲁棒去中心化联邦蒸馏方法
Ma, Xiao, Shen, Hong, Tian, Hui, Ke, Wei, Lyu, Wenqi
Abstract
This paper propose a robust decentralized federated distillation method that enables clients with heterogeneous models to collaborate through predictions on shared unlabeled public data. In the proposed method, each client first evaluates the received predictions in three modalities of class prediction, boundary decision, and prediction correlation. It then filters unreliable clients, assigns reliability-based weights to the retained clients, and constructs a teacher for each type of knowledge. Finally, the corresponding distillation gradients are validated using a supervised gradient computed from private data. Conflicting prediction and boundary gradients are removed, and conflicting relation gradients are suppressed before the final model update. We prove the convergence of the proposed method by showing stable local optimization for honest clients under Byzantine distillation. Particularly, we show that our method ensures a bounded Byzantine influence on both distillation gradients and individual client private gradients after cross-modality fusion, thereby enabling stable local optimization for honest clienunder Byzantine distillation. Extensive experiments on CIFAR-10 and CIFAR-100 demonstrate that the proposed method improves the prediction accuracy of heterogeneous models of clients under non-IID data and Byzantine attacks. As the booming demands of federated learning in decentralized environments such as edge computing and mission-oriented UAV collaborations, our method has a great potential for adoption of DFL in unreliable real-world scenarios where clients are exposed to receiver-specific Byzantine messages of malicious predictions.
Chinese Translation
本文提出一种鲁棒的去中心化联邦蒸馏方法,使拥有异构模型的客户端能够通过在共享的无标签公共数据上的预测进行协作。在所提出的方法中,每个客户端首先从类别预测、边界决策和预测相关性三种模态对接收到的预测进行评估;然后过滤不可靠的客户端,为保留的客户端分配基于可靠性的权重,并为每类知识构建相应的教师模型;最后,利用从私有数据计算得到的监督梯度对相应的蒸馏梯度进行验证,剔除相互冲突的预测梯度和边界梯度,并在最终模型更新前抑制相互冲突的关系梯度。我们通过证明诚实客户端在拜占庭蒸馏下的局部优化稳定性,证明了所提方法的收敛性。特别地,我们表明该方法在跨模态融合后,能将拜占庭影响对蒸馏梯度和各个客户端私有梯度的影响限制在有限范围内,从而保证诚实客户端在拜占庭蒸馏下的稳定局部优化。在CIFAR-10和CIFAR-100上的大量实验表明,所提方法在非IID数据和拜占庭攻击下提升了客户端异构模型的预测精度。随着边缘计算和面向任务的无人机协作等去中心化环境中对联邦学习需求的快速增长,该方法在客户端易受接收端特定拜占庭恶意预测消息影响的不可靠真实场景中,具有推动去中心化联邦学习(DFL)实际应用的巨大潜力。
cs.LG / 134 / 2609.07236

Parallelism Strategy Chaining for Fast Training Convergence

面向快速训练收敛的并行策略链式切换方法
Kang, Minchul, Shin, Changyong, Go, Younghun, Lee, Hyunho, Jeong, Jinwoo, Yoo, Chuck, Yang, Gyeongsik
Abstract
Selecting a parallelism strategy - the configuration of data, tensor, and pipeline parallelism degrees together with micro- and global-batch sizes - largely determines the training efficiency of large language models. State-of-the-art methods search for a parallelism strategy offline and select the single strategy that minimizes per-iteration time. But we find that they neglect the target validation perplexity and time-to-perplexity (TTP). In particular, our analysis reveals that the best strategy yielding the fastest perplexity improvement changes multiple times during training. As a result, state-of-the-art methods are 1.8-11.4x slower in TTP than the strategy sequence that selects the best strategy at each iteration. This paper proposes CONA, a new training method that introduces online strategy chaining. Instead of a single strategy selected offline, CONA ranks candidate strategies during training using a surrogate metric built from compute throughput and gradient statistics, and switches the current strategy to a new strategy with a higher metric. In our evaluation with GPT-3 1.3B, BERT-Large, and Llama-3.2-1B, CONA reaches the target validation perplexity 1.4-9.6x faster than state-of-the-art methods. Moreover, CONA closely tracks the perplexity achieved by the sequence that selects the best strategy at each iteration, within 2.6%.
Chinese Translation
并行策略——即数据并行、张量并行和流水线并行的配置以及微批次和全局批次大小的组合——在很大程度上决定了大语言模型的训练效率。最先进的方法在离线状态下搜索并行策略,并选择使每次迭代时间最短的单一策略。但我们发现,这些方法忽略了目标验证困惑度和达到目标困惑度所需时间(TTP)。特别是,我们的分析表明,能最快降低困惑度的最佳策略在训练过程中会多次变化。因此,最先进的方法在TTP上比每次迭代都选择最佳策略的策略序列慢1.8至11.4倍。本文提出CONA,一种引入在线策略链式切换的新训练方法。CONA不再使用离线选择的单一策略,而是利用基于计算吞吐量和梯度统计信息构建的代理指标在训练过程中对候选策略进行排序,并将当前策略切换到指标更高的新策略。在GPT-3 1.3B、BERT-Large和Llama-3.2-1B的评估中,CONA达到目标验证困惑度的速度比最先进方法快1.4至9.6倍。此外,CONA的困惑度与每次迭代都选择最佳策略的序列所达到的困惑度非常接近,差距在2.6%以内。
cs.LG / 135 / 2609.07240

Kolmogorov--Arnold stability for discontinuous functions

不连续函数的Kolmogorov--Arnold稳定性
Dzhenzher, Sviatoslav V.
Abstract
Here we investigate the stability of the Kolmogorov--Arnold representation theorem (KART) under adversarial reparameterisations of the hidden layer for multivariate discontinuous and unbounded functions. Our results provide a rigorous mathematical foundation for the structural robustness of modern deep learning architectures, such as Kolmogorov--Arnold Networks (KANs), under adversarial configurations.
Chinese Translation
本文研究了Kolmogorov--Arnold表示定理(KART)在多变量不连续且无界函数的隐藏层遭受对抗性重参数化时的稳定性。我们的结果为现代深度学习架构(如Kolmogorov--Arnold网络(KAN))在对抗性配置下的结构鲁棒性提供了严格的数学基础。
cs.LG / 136 / 2609.07264

Dense Structural Compression of Transformers via Gauge-Correct Channel Removal

基于规范修正信道移除的Transformer稠密结构压缩
Duersch, Jed A., Es-Sebbani, Naïm, Haas, Nathanaël, Bouraoui, Zied
Abstract
Inference energy per token drives the cost and carbon footprint of deployed transformers. It is dominated by dense matrix products that incur fused multiply-accumulate (FMA) operations and memory traffic. To reduce these computations while retaining dense tensors for high GPU throughput, we develop a methodology from first principles to adapt structural complexity during training to maximize inference utility per unit compute. Channel penalties drive entire tensor slices to zero to enable physical removal while preserving density and the network function. The natural approach, penalizing the norm of operator components acting through each channel, is provably destabilized by gauge freedom. We resolve this pathology with GaugeLasso: additive symmetric group-lasso penalties that recover a monotone function of product-norms when the network converges to gauge balance. Our equilibrium analysis enables per-channel calibration to correctly suppress slices that under-perform in inference utility per unit compute. Under adaptive pressure, the network reorganizes into depth-dependent structural profiles that can be far smaller than the architecture required to learn the task. On polynomial long division over $\mathbb{F}_{31}$, compute compresses from 148 to 255 times with perfect accuracy. On character-level language modeling, compressed models outperform the hand-designed baseline at equal FMA. On masked autoencoding, a compression trial exposes which axes were over-provisioned and which saturated, guiding a better second design. Compaction also accelerates training monotonically as the model progresses. Post-hoc pruning with the same utility ranking cannot reach these structures, showing that sustained pressure is central to discovery of efficient models. Retraining a discovered architecture recovers baseline quality on our statistical tasks, but fails on our exact algorithmic task.
Chinese Translation
每token的推理能耗决定了已部署Transformer的成本和碳足迹,而该能耗主要由稠密矩阵乘法产生的融合乘加(FMA)运算和内存流量主导。为了在保留稠密张量以实现高GPU吞吐量的同时减少这些计算,我们从第一性原理出发建立了一套方法,在训练过程中自适应地调整结构复杂度,以最大化单位计算量下的推理效用。信道惩罚将整个张量切片驱动至零,从而实现物理移除,同时保持稠密性和网络功能。直接惩罚作用于每个信道的算子分量范数的自然做法,会被规范自由度证明性地破坏稳定。我们用GaugeLasso解决了这一病理问题:这是一种加性对称的组lasso惩罚,当网络收敛到规范平衡时,可恢复乘积范数的单调函数。我们的均衡分析支持逐信道校准,从而正确抑制那些单位计算量下推理效用表现不佳的切片。在自适应压力下,网络重组为依赖深度的结构剖面,其规模可以远小于学习该任务所需的架构。在$\mathbb{F}_{31}$上的多项式长除法任务中,计算量压缩了148至255倍且保持完美精度。在字符级语言建模中,压缩后的模型在相同FMA条件下优于手工设计的基线。在掩码自编码任务中,一次压缩试验揭示了哪些轴被过度配置、哪些已饱和,从而指导了更好的第二次设计。随着模型训练的推进,紧凑化还能单调地加速训练。使用相同效用排序的事后剪枝无法达到这些结构,这表明持续的压力是发现高效模型的关键。对发现的架构进行再训练,在我们的统计任务上能恢复基线质量,但在精确算法任务上则会失败。
cs.LG / 137 / 2609.07294

Constitutive State-Space Modeling of Path-Dependent Plasticity: A Resolution-Consistent and Parallelizable Computational Framework

路径依赖塑性的本构状态空间建模:一种分辨率一致且可并行化的计算框架
Barreira, Rui, Soydan, Taylan, Scipione, Francesco, Bessa, Miguel A., Mohr, Dirk
Abstract
Data-driven constitutive models for path-dependent plasticity are commonly formulated using nonlinear recurrent neural networks, whose sequential state evolution limits parallel training and whose predictions may depend on the discretization of the applied strain path. We introduce a Constitutive State Space (CSS) model that reformulates structured state-space dynamics as an incremental constitutive operator. The strain increment is decomposed into magnitude and direction: the loading direction drives the latent state-space system, while the increment magnitude enters the zero-order-hold discretization of its continuous-time linear recurrence. This mechanics-tailored construction guarantees stationarity under zero increments, strongly reduces sensitivity to strain-path resolution, and retains the parallel-scan structure of S5 for efficient training on long constitutive histories. The CSS and Minimal State Cell (MSC) architectures are compared for four multiaxial path-dependent material models including isotropic J2 plasticity, pressure-sensitive foam plasticity, and combined isotropic-kinematic hardening. CSS matches or exceeds the prediction accuracy of the MSC, including one order of magnitude lower validation losses for the plastically incompressible materials. Importantly, CSS maintains low errors across large changes in strain-path discretization, whereas the MSC error increases substantially when evaluated at coarser resolutions than used for training. CSS trains substantially faster and requires fewer strain-stress pairs to attain comparable or better accuracy. Analysis of the learned state further reveals latent structure consistent with the dimensionality of the underlying physical constitutive models. These results establish mechanics-tailored structured state-space dynamics as a computational framework for efficient and discretization-robust data-driven constitutive modeling.
Chinese Translation
针对路径依赖塑性的数据驱动本构模型通常采用非线性循环神经网络构建,其顺序状态演化限制了并行训练,且预测结果可能依赖于所施加应变路径的离散化方式。本文提出一种本构状态空间(Constitutive State Space, CSS)模型,将结构化状态空间动力学重新表述为增量本构算子。应变增量被分解为幅值和方向:加载方向驱动潜在状态空间系统,而增量幅值则进入其连续时间线性递推的零阶保持离散化。这种面向力学特性的构造保证了零增量下的平稳性,显著降低了对应变路径分辨率的敏感性,并保留了S5的并行扫描结构,从而能够对长本构历史进行高效训练。本文针对四种多轴路径依赖材料模型——包括各向同性J2塑性、压力敏感泡沫塑性以及各向同性-运动学组合硬化——对CSS与最小状态单元(Minimal State Cell, MSC)架构进行了比较。CSS达到或超越了MSC的预测精度,对于塑性不可压缩材料,其验证损失甚至低一个数量级。重要的是,CSS在应变路径离散化发生大幅变化时仍能保持较低误差,而MSC在粗于训练所用分辨率下评估时误差显著增大。CSS训练速度显著更快,且达到相当或更优精度所需的应变-应力样本对更少。对学到的状态的分析进一步表明,其潜在结构与底层物理本构模型的维度一致。这些结果确立了面向力学特性的结构化状态空间动力学作为高效且对离散化稳健的数据驱动本构建模计算框架的地位。
cs.LG / 138 / 2609.07300

PCFlow: Physics-Conditioned Flow Matching for GPR B-Scan Image Synthesis

PCFlow:面向探地雷达B-scan图像合成的物理条件流匹配方法
Shen, Zhijie, Fu, Chenchen, Chang, Xuanhao, Bai, Hongtao, He, Lili
Abstract
Ground-penetrating radar (GPR) B-scan image synthesis is important for data augmentation, algorithm validation, and simulation acceleration, yet generating radargrams with both visual realism and physical consistency remains challenging. Existing learning-based generative models often emphasize visual appearance but provide limited control over response geometry. In this paper, we propose PCFlow, a physics-conditioned flow matching framework for fast GPR B-scan image synthesis. The core of PCFlow is a Maxwell-informed dense physical condition field constructed from the parameterized physical model used for electromagnetic simulation, including material properties, target geometry, propagation cues, and response-domain priors. This condition field provides an interpretable interface between physical scene parameters and radar response geometry, and guides conditional flow matching in the VAE latent space toward physically feasible generation paths. We evaluate PCFlow on a gprMax-based buried-pipeline dataset with both in-distribution and out-of-distribution test cases. Experimental results show that PCFlow generates images with more accurate response geometry and high visual fidelity, demonstrating its effectiveness for controllable and physically faithful radar image synthesis.
Chinese Translation
探地雷达(GPR)B-scan图像合成对于数据增强、算法验证和仿真加速具有重要意义,然而生成兼具视觉真实性和物理一致性的雷达图像仍然具有挑战性。现有的基于学习的生成模型往往强调视觉外观,但对响应几何形态的控制能力有限。本文提出PCFlow,一个用于快速GPR B-scan图像合成的物理条件流匹配框架。PCFlow的核心是由用于电磁仿真的参数化物理模型构建的麦克斯韦方程信息引导的稠密物理条件场,包括材料属性、目标几何、传播线索和响应域先验。该条件场在物理场景参数与雷达响应几何之间提供了一个可解释的接口,并引导VAE潜空间中的条件流匹配走向物理可行的生成路径。我们在基于gprMax构建的埋地管线数据集上,对PCFlow进行了分布内和分布外测试用例的评估。实验结果表明,PCFlow生成的图像具有更精确的响应几何和较高的视觉保真度,证明了其在可控且物理可信的雷达图像合成方面的有效性。
cs.LG / 139 / 2609.07303

Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching

基于渐进点匹配的长程语言模型强化学习
Fu, Preston, Frans, Kevin, Rybkin, Oleh, Levine, Sergey, Kumar, Aviral
Abstract
Current paradigms for training language models via reinforcement learning rely heavily on sparse outcome rewards. However, as we pursue tasks that require longer and more complicated trajectories, such strategies result in slow learning. Prior work has attempted to address this problem by rewarding partial progress; however, naive formulations are often biased and converge to suboptimal policies. We show that a simple and unbiased dense reward formulation, which we term progressive point matching, scales exponentially more efficiently to long-horizon tasks by rewarding partial progress on a segment level, both theoretically and empirically via synthetic environments. We then show how progressive point matching can be practically instantiated using a single reference trajectory per task. On extremely hard math reasoning problems, sparse outcome rewards cannot make any progress, whereas segment-level rewards enable improvements at larger test-time token budgets when measured by success rate or pass@k.
Chinese Translation
当前通过强化学习训练语言模型的范式严重依赖稀疏的结果奖励。然而,当我们追求需要更长、更复杂轨迹的任务时,这类策略会导致学习缓慢。先前的工作尝试通过对部分进展给予奖励来解决这一问题,但朴素的奖励设计往往存在偏差,并收敛到次优策略。我们提出了一种简单且无偏的稠密奖励设计,称为渐进点匹配(progressive point matching),它在片段(segment)层面奖励部分进展。我们通过理论分析和合成环境中的实证实验表明,该方法在长程任务上的扩展效率呈指数级提升。随后,我们展示了如何通过每个任务仅需一条参考轨迹来实际实现渐进点匹配。在极难的数学推理问题上,稀疏结果奖励无法取得任何进展,而片段级奖励在更大的测试时token预算下,无论以成功率还是pass@k衡量,都能带来改进。
cs.LG / 140 / 2609.07312

Robust Decentralized Personalized Federated Learning via Prediction-Constrained Neighborhood Collaboration

基于预测约束邻域协作的鲁棒去中心化个性化联邦学习
Ma, Xiao, Shen, Hong, Tian, Hui, Lyu, Wenqi, Ke, Wei
Abstract
This paper proposes a robust decentralized personalized federated learning method R-DPFL, that enables clients to reduce the impact of Byzantine attacks via robust neighborhood direction estimation and history-based update trend prediction, rather than purely aggregating client models as in the existing work. In R-DPFL, each client first computes the current-round model update by aggregating the received neighborhood update vectors. It then predicts what this update should be based on its historical values and local model changes. Finally, R-DPFL computes the difference between these two quantities, adaptively clips this difference, and adds it to the local update. We prove convergence of the learning process through rigorous analysis and show that honest clients maintain stable personalized descent dynamics under Byzantine neighbor perturbations without requiring consensus among neighboring models. Extensive experiments on CIFAR-10 demonstrate that RDPFL consistently outperforms state-of-the-art decentralized and personalized federated learning baselines under heterogeneous and adversarial settings.
Chinese Translation
本文提出了一种鲁棒的去中心化个性化联邦学习方法 R-DPFL,使客户端能够通过鲁棒的邻域方向估计和基于历史的更新趋势预测来降低拜占庭攻击的影响,而不是像现有工作那样纯粹地聚合客户端模型。在 R-DPFL 中,每个客户端首先通过聚合接收到的邻域更新向量来计算当前轮次的模型更新;然后基于历史值和本地模型变化预测该更新应有的取值;最后,R-DPFL 计算这两个量之间的差异,自适应地对差异进行裁剪,并将其加到本地更新中。我们通过严格的分析证明了学习过程的收敛性,并表明诚实客户端在拜占庭邻居扰动下能够保持稳定的个性化下降动态,而无需相邻模型之间达成共识。在 CIFAR-10 上的大量实验表明,在异构和对抗性设置下,R-DPFL 始终优于最先进的去中心化和个性化联邦学习基线方法。
cs.LG / 141 / 2609.07349

Inferring Urban Mobility Interactions from Aggregated Dynamics

从聚合动态推断城市出行交互
Wang, Yi, Li, Jing, Deng, Jinliang, Wang, Zhenghong, Zhang, Yizhi, Zhang, Fan, Tsang, Ivor W., Liu, Yu
Abstract
Real-time urban governance depends not only on knowing where people are, but on how they move between places, directional flows that could be conventionally resolved by tracking individuals through space, i.e., expensive to sustain and built on traces that are highly unique and readily re-identifiable. Here we show that this directional structure need not be observed to be known: aggregated counts which cities already collect retain enough information to reconstruct the temporal evolution of origin-destination (OD) matrix. Using an uncertainty-aware physics-informed framework, we infer future OD flows from area-level counts alone across twelve mobility datasets from cities in the United States and China, reaching accuracy comparable to models that take historical OD matrices as input. Probabilistic modeling corrects the systematic underestimation of sparse, high-value corridors and yields calibrated predictions consistent with observed flows. Architectures that respect the generation-before-assignment logic of transport planning recover interactions more faithfully, indicating that location-level spatial heterogeneity should be preserved before pairwise interactions are reconstructed. Because inference requires only aggregated observations after training, recovering interactions this way reduces reliance on continuous individual-level tracking, pointing toward a more deployable and less exposure-heavy basis for real-time urban intelligence.
Chinese Translation
实时城市治理不仅需要掌握人在哪里,还需要了解人如何在地点之间移动——这类方向性流动传统上需要通过追踪个体在空间中的轨迹来获取,既成本高昂,又建立在高度独特、极易被重新识别的痕迹之上。本文表明,这种方向性结构无需直接观测即可获知:城市已在采集的聚合计数数据保留了足够的信息,可用于重构起点-终点(OD)矩阵的时间演化。我们利用一个具有不确定性感知的物理信息框架,仅基于区域层面的计数数据,在美国和中国城市的十二个出行数据集上推断未来的OD流量,其精度可与以历史OD矩阵为输入的模型相媲美。概率建模纠正了对稀疏、高价值通道的系统性低估,并产生与观测流量一致的、经过校准的概率预测。遵循交通规划中“生成-分配”先后逻辑的架构能够更忠实地恢复交互关系,这表明在重构成对交互之前,应保留位置层面的空间异质性。由于训练后的推断仅需聚合观测数据,这种方式降低了对持续性个体级追踪的依赖,为实时城市智能提供了更易部署、隐私暴露更少的基础。
cs.LG / 142 / 2609.07406

Think Wider: Mitigating Latent Rank Collapse in Implicit Chain-of-Thought Reasoning

拓宽思维:缓解隐式思维链推理中的潜在秩坍缩
Hao, Yuwen, Yang, Menglin
Abstract
Chain-of-thought (CoT) reasoning improves the reasoning ability of large language models by introducing intermediate computation, but explicit rationales increase decoding length, latency, and context cost. Implicit CoT offers a more efficient alternative by moving intermediate reasoning into continuous latent states. However, latent reasoning can be unstable: successive latent states may become overly similar and collapse toward a shared dominant direction, reducing the diversity of the reasoning trajectory. In this work, we identify $\textit{latent rank collapse}$ and propose $\textbf{WIDER}$, a lightweight spectral regularizer for implicit CoT. During training, WIDER estimates the shared direction of each latent trajectory and penalizes projections onto this direction, encouraging latent states to span a broader representational subspace. The method is plug-and-play and leaves the backbone model, latent schedule, and inference-time decoding procedure unchanged. We further formulate this collapse as a geometric bottleneck in implicit reasoning, casting its mitigation as a training-time regularization problem rather than an inference-time decoding change. Extensive experiments show that WIDER improves matched implicit CoT baselines, while mechanistic analyses reveal higher effective rank, lower dominant-direction energy, and reduced redundancy among latent steps. These results highlight latent subspace utilization as an important factor for efficient continuous reasoning, providing a geometric perspective for analyzing and improving implicit CoT. Code is available at https://github.com/whitesweater/WIDER.
Chinese Translation
思维链(Chain-of-Thought, CoT)推理通过引入中间计算来提升大语言模型的推理能力,但显式的推理依据会增加解码长度、延迟和上下文开销。隐式CoT(Implicit CoT)通过将中间推理过程转移至连续潜在状态中,提供了一种更高效的替代方案。然而,潜在推理可能不稳定:相继的潜在状态可能变得过于相似,并坍缩向某个共享的主导方向,从而降低推理轨迹的多样性。在本工作中,我们识别出潜在秩坍缩(latent rank collapse)现象,并提出了一种用于隐式CoT的轻量级谱正则化方法——WIDER。在训练过程中,WIDER估计每条潜在轨迹的共享方向,并惩罚向该方向的投影,促使潜在状态张成更广泛的表示子空间。该方法即插即用,且不改变骨干模型、潜在状态调度以及推理时的解码流程。我们进一步将这种坍缩形式化为隐式推理中的几何瓶颈,将其缓解视为训练时的正则化问题,而非推理时的解码改动。大量实验表明,WIDER在同等条件下的隐式CoT基线上取得了提升;机制分析进一步揭示出更高的有效秩、更低的主导方向能量以及潜在步骤之间更少的冗余。这些结果凸显了潜在子空间利用率作为高效连续推理的重要因素,为分析和改进隐式CoT提供了几何视角。代码已发布于 https://github.com/whitesweater/WIDER。
cs.LG / 143 / 2609.07408

Impact of canny edge detection preprocessing on performance of machine learning models for Parkinson's disease classification

Canny边缘检测预处理对帕金森病分类机器学习模型性能的影响
Bhat, Sameer, Szczuko, Piotr
Abstract
This study investigates the classification of individuals as healthy or at risk of Parkinson's disease using machine learning (ML) models, focusing on the impact of dataset size and preprocessing techniques on model performance. Four datasets are created from an original dataset: DS_0, (normal dataset), DS_1 (DS_O subjected to Canny edge detection and Hessian filtering), DS_2 (augmented DS_0), and DS_3 (augmented DS_1). We evaluate a range of ML models-Logistic Regression (LR), Decision Tree (DT), Random Forest (RF), Gradient Boosting (GB), XGBoost (XBG), Naive Bayes (NB), Support Vector Machine (SVM), and AdaBoost (AdB)-on these datasets, analyzing prediction accuracy, model size, and prediction latency. The results show that while larger datasets lead to increased model memory footprints and prediction latencies, the Canny edge detection preprocessing supplemented by Hessian filtering (used in DS_1 and DS_3) degrades the performance of most models. In our experiment, we observe that Random Forest (RF) maintains a stable memory footprint of 61 KB across all datasets, while models like KNN and SVM show significant increases in memory usage, from 5.7-7 KB on DS_0 to 102-220 KB on DS_2, and similar increases in prediction time. Logistic Regression, Decision Tree, and Naive Bayes show stable memory footprints and fast prediction times across all datasets. XGBoost's prediction time increases from 180-200 ms on DS_0 to 700-3000 ms on DS_2 (truncated)
Chinese Translation
本研究探讨了使用机器学习(ML)模型对健康个体与有帕金森病风险的个体进行分类,重点关注数据集规模和预处理技术对模型性能的影响。基于原始数据集创建了四个数据集:DS_0(普通数据集)、DS_1(对DS_0进行Canny边缘检测和Hessian滤波处理)、DS_2(增强后的DS_0)和DS_3(增强后的DS_1)。我们在这些数据集上评估了多种机器学习模型——逻辑回归(LR)、决策树(DT)、随机森林(RF)、梯度提升(GB)、XGBoost(XBG)、朴素贝叶斯(NB)、支持向量机(SVM)和AdaBoost(AdB),并分析了预测准确率、模型大小和预测延迟。结果表明,虽然更大的数据集会导致模型内存占用和预测延迟的增加,但辅以Hessian滤波的Canny边缘检测预处理(用于DS_1和DS_3)降低了大多数模型的性能。在实验中,我们观察到随机森林(RF)在所有数据集上均保持61 KB的稳定内存占用,而KNN和SVM等模型的内存使用量显著增加,从DS_0上的5.7-7 KB增加到DS_2上的102-220 KB,预测时间也有类似幅度的增加。逻辑回归、决策树和朴素贝叶斯在所有数据集上均表现出稳定的内存占用和快速的预测时间。XGBoost的预测时间从DS_0上的180-200 ms增加到DS_2上的700-3000 ms(截断)
cs.LG / 144 / 2609.07432

Revisiting Thinning Methods for Kernel Learning Problems

重访核学习问题中的稀疏化方法
Cano-Camarero, Blanca, Aguado-Carrillo-de-Albornoz, Yago R., Fernández-Pascual, Ángela, Dorronsoro, José R.
Abstract
Kernel methods are widely used because of their strong theoretical guarantees and empirical performance. However, their high computational cost limits their applicability to large-scale datasets. To address this shortcoming, several approaches use Maximum Mean Discrepancy to construct representative subsets that preserve the properties of the full dataset in a Reproducing Kernel Hilbert Space. We introduce Backward Kernel Herding, an algorithm that addresses this problem by iteratively removing points from the dataset, achieving results comparable to current state-of-the-art approaches while accelerating the subsampling process in realistic scenarios where the reduced size is less than half of the dataset. Moreover, we overcome a limitation of Kernel Thinning by proposing an extension that enables the construction of subsets of arbitrary size rather that restricting to successive halvings. Finally, we conduct an extensive experimental comparison focusing on the most relevant kernel learning procedures: Gaussian Processes and Kernel Support Vector Machines. The results show that Backward Kernel Herding consistently achieves competitive performance with the most favorable training-time efficiency, while the proposed Flexible Kernel Thinning frequently achieves the best predictive performance. These gains become especially pronounced for moderate compression ratios, highlighting the benefits of incorporating supervised information into the thinning process. In terms of memory consumption, Flexible Kernel Thinning is also competitive, whereas Backward Kernel Herding remains an alternative when computational efficiency is the primary objective. Overall, no single method dominates across all scenarios, underscoring the importance of selecting the reduction strategy according to the desired trade-off between predictive performance, training cost, and memory requirements.
Chinese Translation
核方法因其强大的理论保证和良好的实证性能而被广泛应用。然而,其高昂的计算成本限制了其在大规模数据集上的适用性。为解决这一缺陷,若干方法利用最大均值差异(Maximum Mean Discrepancy)在再生核希尔伯特空间(Reproducing Kernel Hilbert Space)中构建能够保留完整数据集性质的代表性子集。我们提出了反向核群集(Backward Kernel Herding)算法,该算法通过从数据集中迭代地删除点来解决这一问题,在缩减规模小于数据集一半的现实场景中,能够在加速子采样过程的同时取得与当前最先进方法相当的结果。此外,我们克服了核稀疏化(Kernel Threading)的一个局限,提出了一种扩展方法,使得能够构建任意大小的子集,而不再局限于连续减半的方式。最后,我们针对最相关的核学习流程——高斯过程(Gaussian Processes)和核支持向量机(Kernel Support Vector Machines)——进行了广泛的实验比较。结果表明,反向核群集(Backward Kernel Herding)始终取得具有竞争力的性能,并拥有最佳的训练时间效率;而所提出的灵活核稀疏化(Flexible Kernel Thinning)则经常取得最佳的预测性能。这些优势在中等压缩率下尤为显著,凸显了在稀疏化过程中引入监督信息的益处。在内存消耗方面,灵活核稀疏化同样具有竞争力,而当计算效率为首要目标时,反向核群集仍是备选方案。总体而言,没有任何单一方法能够在所有场景中占据主导地位,这强调了根据预测性能、训练成本与内存需求之间的期望权衡来选择缩减策略的重要性。
cs.LG / 145 / 2609.07441

TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tables

TabBench-Bio:面向高维生物医学表格数据机器学习的动态基准测试
Kreuer, Jules, Ouaari, Sofiane, Hellmig, Julia, Braitinger, Julius, Pfeifer, Nico
Abstract
Biomedical tables often combine thousands of measured variables with only tens or hundreds of labelled samples, a regime that is poorly represented in general-purpose tabular benchmarks. We introduce TabBench-Bio, a living and interactive benchmark of 43 biomedical datasets spanning multiple domains. Under a shared cross-validation protocol, we compare classical estimators, neural networks, and tabular foundation models across 28 feature-by-sample operating points. At the reference cell of 10,000 features and 100 training samples, RealTabPFN v2.5 has the highest point estimate, followed by Logistic Regression and TabDPT, whose point estimates are nearly identical. A paired bootstrap over the target pool separates RealTabPFN v2.5 from Logistic Regression by 145 Elo (95% interval [59, 232]). Tabular foundation models generally occupy the leading ranks, while the strongest configuration depends on the operating point and biomedical modality. The AutoML framework AutoGluon, using its one-hour "extreme" preset, is configured as a separate resource-intensive reference and is reported here at the reference cell. Fold-level predictions, run status, and deterministic aggregations make every reported result reproducible and reusable. We invite the community to contribute: TabBench-Bio is designed to grow, and we welcome submissions of new biomedical tabular datasets, particularly from underrepresented assays and clinical endpoints, for inclusion in future releases. The interactive leaderboard is available at: https://tabbench-bio.eu
Chinese Translation
生物医学表格数据通常将数千个测量变量与仅有数十或数百个标注样本相结合,这一情形在通用表格基准测试中代表性不足。我们提出了TabBench-Bio,一个包含43个跨多个领域生物医学数据集的动态交互式基准测试。在统一的交叉验证协议下,我们在28个特征×样本的操作点上比较了经典估计器、神经网络和表格基础模型。在10,000个特征和100个训练样本的参考单元中,RealTabPFN v2.5的点估计值最高,其次是Logistic Regression和TabDPT,后两者的点估计值几乎相同。基于目标池的配对自助法(paired bootstrap)分析显示,RealTabPFN v2.5领先Logistic Regression 145个Elo分(95%置信区间[59, 232])。表格基础模型总体上占据领先排名,而最优配置取决于操作点和生物医学模态。AutoML框架AutoGluon(使用其一小时"extreme"预设)被配置为单独的资源密集型参考,并在参考单元处进行报告。折级别预测、运行状态和确定性聚合使每个报告的结果均可复现和复用。我们邀请社区参与贡献:TabBench-Bio旨在持续扩展,欢迎提交新的生物医学表格数据集,尤其是来自代表性不足的检测方法和临床终点的数据集,以便纳入未来版本。交互式排行榜可通过以下网址访问:https://tabbench-bio.eu
cs.LG / 146 / 2609.07444

TASTE: Throughput-Aware Batch Size Tuning for On-Device Edge Learning

TASTE:面向设备端边缘学习的吞吐量感知批大小调优
Bhatnagar, Avik, Peccia, Federico Nicolas, Bringmann, Oliver
Abstract
The rise of privacy-preserving artificial intelligence (AI) has shifted the focus of model adaptation and personalization towards on-device learning, where deep learning models are finetuned directly on edge hardware using local user data. However, this shift requires optimization of deep learning training on resource-constrained hardware to maximize throughput while maintaining predictive accuracy. This paper introduces a novel technique for on-device model training that incorporates an efficient Bayesian optimization-based batch size tuning approach to maximize hardware throughput. To evaluate the impact of this hyperparameter on the learning dynamics, we investigated two distinct paradigms: standard supervised learning (SL) and online continual learning (CL). Experimental results across various edge devices demonstrate a throughput ceiling, beyond which increasing the batch size yields no additional throughput gains. The proposed tuning approach identifies the optimal batch size, which, when combined with gradient accumulation and linear learning rate scaling, achieves up to a 2X increase in training throughput on platforms such as Raspberry Pi 4 compared to maximum batch sizes, without compromising model accuracy. Furthermore, in the CL paradigm, we demonstrate that optimal batch sizes maintain the stability-plasticity balance required for incremental learning, effectively mitigating catastrophic forgetting while maximizing computational efficiency on edge-hardware.
Chinese Translation
隐私保护人工智能(AI)的兴起使模型适配与个性化的重心转向设备端学习(on-device learning),即利用本地用户数据在边缘硬件上直接微调深度学习模型。然而,这一转变要求在资源受限的硬件上优化深度学习训练,以在保持预测精度的同时最大化吞吐量。本文提出了一种新颖的设备端模型训练技术,该技术采用高效的基于贝叶斯优化的批大小(batch size)调优方法,以最大化硬件吞吐量。为评估该超参数对学习动态的影响,我们研究了两种不同的范式:标准监督学习(SL)和在线持续学习(CL)。在多种边缘设备上的实验结果表明存在一个吞吐量上限,超过该上限后,增大批大小不再带来额外的吞吐量提升。所提出的调优方法能够识别最优批大小,结合梯度累积和线性学习率缩放,在Raspberry Pi 4等平台上与使用最大批大小相比,训练吞吐量提升高达2倍,且不损害模型精度。此外,在持续学习范式下,我们证明最优批大小能够维持增量学习所需的稳定性-可塑性平衡,在最大化边缘硬件计算效率的同时,有效缓解灾难性遗忘。
cs.LG / 147 / 2609.07461

Temporal-Causal Inference for Reinforcement Learning via Automata Learning

基于自动机学习的强化学习时序因果推断
Corazza, Jan, Kaminskyi, Daniil, Lutz, Simon, Nossol, Patrick, Aria, Hadi Partovi, Xu, Zhe, Neider, Daniel
Abstract
We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition governed by a hidden temporal pattern. The agent observes the base state but cannot observe the phase directly. We formalize this problem as a two-phase non-Markovian decision process and introduce Temporal-Causal Inference for Reinforcement Learning (TCIRL), a framework that jointly learns a control policy and infers the hidden temporal cause of the phase transition. TCIRL maintains a hypothesis deterministic finite automaton (DFA) to track what phase is active and refines it via counterexample-driven SAT-based synthesis. We prove that the hypothesis converges almost surely to a DFA recognizing the true cause language on all attainable label sequences, yielding an optimal policy for the original non-Markovian decision process. Experiments on a genetic therapy gridworld and a traffic signal environment show that TCIRL recovers the correct cause DFA and matches the full-information baseline in both domains.
Chinese Translation
我们研究强化学习在环境动态经历由隐藏时序模式支配的不可逆相变时的应用。智能体可以观察到基础状态,但无法直接观察到相。我们将该问题形式化为一个两阶段非马尔可夫决策过程,并提出用于强化学习的时序因果推断框架(Temporal-Causal Inference for Reinforcement Learning, TCIRL),该框架联合学习控制策略并推断相变的隐藏时序原因。TCIRL 维护一个假设的确定性有限自动机(DFA)来追踪当前处于哪个相,并通过基于反例的 SAT 求解器驱动的综合方法对其进行精化。我们证明,该假设在所有可达的标签序列上几乎必然收敛于一个识别真实原因语言的 DFA,从而为原始的非马尔可夫决策过程产生最优策略。在基因治疗网格世界(genetic therapy gridworld)和交通信号环境上的实验表明,TCIRL 能够恢复正确的原因 DFA,并在两个领域中均达到与全信息基线相当的性能。
cs.LG / 148 / 2609.07493

Improving Multivariate Time Series Classification with Class-Wise Training and Model Aggregation

基于类别级训练与模型聚合的多变量时间序列分类改进方法
Lo, Mouhamadou Mansour, Morvan, Gildas, Rossi, Mathieu, Morganti, Fabrice, Mercier, David
Abstract
In this paper, we propose a class-wise dimension (channel) selection framework for Multivariate Time Series Classification (MTSC). Rather than applying a single global dimension selection process, the proposed approach independently identifies informative dimensions for each class. A dedicated learning process is subsequently performed for each class, followed by a fusion stage for final prediction. The objective is to improve the generation of discriminative feature representations while reducing the influence of noisy or non-informative dimensions. The proposed framework is evaluated using MiniRocket, a random kernel-based baseline method. Experimental results indicate that class-wise dimension selection improves the quality of extracted representations and can enhance classification performance, particularly in high-dimensional settings. These findings suggest that incorporating class-specific information into the training process represents a promising direction for MTSC, improving robustness through consistent gains across heterogeneous datasets, and interpretability through the explicit identification of class-relevant dimensions.
Chinese Translation
本文提出了一种面向多变量时间序列分类(MTSC)的类别级维度(通道)选择框架。与采用单一全局维度选择过程不同,所提方法为每个类别独立地识别有信息量的维度。随后,针对每个类别执行专门的学习过程,并通过融合阶段得到最终预测结果。该方法旨在提升判别性特征表示的质量,同时降低噪声维度或无信息维度的影响。我们采用基于随机核的基线方法 MiniRocket 对所提框架进行了评估。实验结果表明,类别级维度选择提升了所提取表示的质量,并能够增强分类性能,尤其在高维场景下效果更为显著。这些发现表明,将类别特定信息融入训练过程是 MTSC 领域一个颇具前景的研究方向:一方面,通过在异构数据集上取得一致的增益提升了鲁棒性;另一方面,通过显式识别与类别相关的维度增强了可解释性。
cs.LG / 149 / 2609.07512

Statistical versus machine learning-based spatial interpolation of post-processed ensemble weather forecasts

基于统计方法与机器学习方法的后处理集合天气预报空间插值比较
Lakatos, Mária
Abstract
Statistical post-processing improves ensemble weather forecasts, but generating calibrated predictions at locations without observations remains challenging. This study compares statistical and machine-learning-based methods for post-processing ECMWF 2-m temperature and 10-m wind speed forecasts at observed and unobserved stations in Germany. We consider EMOS-based approaches, distributional regression networks, Transformers, and graph neural networks under both limited and extended predictor settings. For temperature, we also investigate linear forecast combinations and propose an altitude-aware linear pool (ALP). The results show that post-processing improves upon the raw ensemble in most settings, but no single method performs best across all variables, station groups, and evaluation metrics. The proposed ALP provides a small but significant improvement over the standard linear pool at unobserved locations.
Chinese Translation
统计后处理能够改进集合天气预报,但在没有观测的地点生成经过校准的预测仍然具有挑战性。本研究比较了统计方法和基于机器学习的方法,用于对德国境内有观测和无观测站点的 ECMWF 2米温度和10米风速预报进行后处理。我们考虑了基于 EMOS 的方法、分布回归网络(distributional regression networks)、Transformer 以及图神经网络(graph neural networks),并分别在有限预测因子和扩展预测因子设置下进行实验。对于温度预报,我们还研究了线性预报组合,并提出了一种考虑海拔高度的线性组合方法(altitude-aware linear pool, ALP)。结果表明,在大多数设置下,后处理均优于原始集合预报,但没有任何单一方法在所有变量、站点组和评估指标上均表现最佳。所提出的 ALP 方法在无观测地点相比标准线性组合方法取得了小幅但显著的改进。
cs.LG / 150 / 2609.07529

CoRL: Co-Evolutionary Reinforcement Learning for Adaptive Indirect Prompt-Injection Attacks and Defenses

CoRL:面向自适应间接提示注入攻击与防御的协同进化强化学习
Zhang, Boyang, Xiao, Qingxin, Dang, Lingwei, Wu, Qingyao
Abstract
Tool-augmented language agents are vulnerable to indirect prompt injection (IPI). Unlike direct prompt injection, IPI hides adversarial instructions in untrusted tool outputs and can covertly alter the execution of a legitimate task. Defenses trained on fixed attacks may fail as an attacker changes its strategy, injection site, and payload. To address this problem, we formulate adaptive IPI as an asymmetric, partially observable, general-sum Markov game: a multi-turn attacker adapts payloads at reached tool-return sites from the public trajectory, while a tool-using defender must block the injected objective and complete the user task. We propose CoRL, a verifier-grounded co-evolution and repair framework with three stages: Attacker SFT initializes multi-turn attacks from successful trajectories; bilateral Co-PPO jointly trains both agents with role-specific rewards and historical opponent populations; and Defender SFT consolidates verifier-accepted teacher repairs for population-discovered failures. Across 1,514 clean, fixed-template, and adaptive executions per defender, CoRL reduces overall ASR by 38.5 points to 0.0% and raises utility by 13.1 points to 76.3%. Stage-wise and controlled ablations show positive contributions from online Co-PPO and population-mined repair, while external-benchmark evaluation indicates transfer in attack resistance. The defender balances safety and task utility under the evaluated attacks, while the retained attackers provide candidates for adaptive red-team evaluation.
Chinese Translation
工具增强型语言智能体容易受到间接提示注入(Indirect Prompt Injection, IPI)攻击。与直接提示注入不同,间接提示注入将对抗性指令隐藏在不可信的工具输出中,能够隐蔽地改变合法任务的执行过程。基于固定攻击训练的防御方法,可能随着攻击者改变策略、注入位置和攻击载荷而失效。为解决这一问题,我们将自适应间接提示注入形式化为一个非对称、部分可观测的一般和马尔可夫博弈:多轮攻击者根据公开轨迹在已达成的工具返回位置自适应地调整攻击载荷,而使用工具的防御者必须阻止注入目标并完成用户任务。我们提出CoRL,一个以验证器为基础的协同进化与修复框架,包含三个阶段:攻击者SFT从成功轨迹初始化多轮攻击;双边Co-PPO利用角色特定的奖励和历史对手种群对双方智能体进行联合训练;防御者SFT巩固验证器接受的教师修复结果,以应对种群发现的失败案例。在针对每个防御者的1,514次干净、固定模板和自适应执行测试中,CoRL将整体攻击成功率(ASR)降低38.5个百分点至0.0%,并将任务效用提升13.1个百分点至76.3%。分阶段和受控消融实验表明,在线Co-PPO和种群挖掘的修复方法具有积极贡献,而外部基准评估显示防御能力具备攻击抵御方面的迁移性。在所评估的攻击下,防御者在安全性与任务效用之间取得了平衡,同时保留的攻击者为自适应红队评估提供了候选方案。
cs.LG / 151 / 2609.07566

No-Regret Mixing of LRU and LFU with Optimal Switching Cost

具有最优切换代价的LRU与LFU的无遗憾混合策略
Mazziane, Younes Ben, Zou, Xinying
Abstract
Caching systems often rely on simple eviction policies such as Least Recently Used (LRU) and Least Frequently Used (LFU), which perform well in complementary request regimes. Recent policies such as LeCar and Cacheus combine LRU and LFU using ideas from the experts problem in online learning. Specifically, upon a miss, they randomize between the two eviction rules using probabilities derived from scores updated by tracking the history of past evictions. While these policies exhibit strong empirical performance, it remains unclear whether they are guaranteed, on every request sequence, to perform asymptotically as well as the better of LRU and LFU, i.e., whether they achieve sublinear regret with respect to this benchmark. We first show that LeCar suffers linear regret against an oblivious adversary, even with unbounded history. We then propose H-MC, a Hedge-based mixture of virtual LRU and LFU caches that preserves Hedge's selection probabilities, and hence its regret guarantees, while minimizing the switching cost among all joint selection rules with these marginals.
Chinese Translation
缓存系统通常依赖简单的淘汰策略,如最近最少使用(LRU)和最不经常使用(LFU),它们在互补的请求模式下表现良好。近期的一些策略,如 LeCar 和 Cacheus,借鉴在线学习中专家问题的思想,将 LRU 和 LFU 结合起来。具体而言,在发生缓存未命中时,它们根据分数在两种淘汰规则之间随机选择,这些分数通过追踪过去的淘汰历史进行更新。尽管这些策略在经验上表现优异,但仍不清楚它们是否能够在每条请求序列上保证渐近地表现得与 LRU 和 LFU 中较优者相当,即是否能够相对于该基准实现次线性遗憾。我们首先证明,即使拥有无界的历史信息,LeCar 在面对无感知对手时仍会产生线性遗憾。随后,我们提出了 H-MC,一种基于 Hedge 算法的虚拟 LRU 与 LFU 缓存混合策略。该策略保持了 Hedge 的选择概率,从而保留其遗憾保证,同时在所有具有相同边际分布的联合选择规则中最小化切换代价。
cs.LG / 152 / 2609.07575

Efficient Exploration Is Enough

高效探索已然足够
Malagón, Mikel, Vadillo, Jon, Ceberio, Josu, Bowling, Michael, Lozano, Jose A.
Abstract
This work introduces an alternative view of efficient exploration and studies its theoretical and empirical implications in the absence of extrinsic rewards. Specifically, we define efficient explorers as agents that prioritize generating generalizable experience, i.e., data that supports learning models capable of predicting and adapting across the environment. This allows us to analyze efficient exploration through the lens of prediction and generalization. Theoretically, we demonstrate that optimally efficient explorers naturally schedule their trajectories to visit the most informative and learnable regions first. Empirically, we show that optimizing for these agents gives rise to an automatic curriculum of progressively more complex behaviors, even in relatively simple environments. These results indicate that pursuing this purely intrinsic objective alone is enough to drive the emergence of highly sophisticated behaviors. We believe that this new framework provides a principled mechanism by which agent-environment systems may sustain an open-ended process of increasingly complex behavior without external rewards, tasks, or objectives.
Chinese Translation
本工作提出了关于高效探索的一种替代性视角,并研究其在缺乏外部奖励情形下的理论与实证意义。具体而言,我们将高效探索者定义为优先产生可泛化经验(即能够支持学习可跨环境预测与适应的模型的数据)的智能体。这使我们能够通过预测与泛化的视角来分析高效探索。在理论方面,我们证明最优高效探索者会自然地规划其轨迹,优先访问信息量最大且最可学习的区域。在实证方面,我们表明针对此类智能体进行优化会产生一个自动课程,呈现出逐步复杂的行为,即使是在相对简单的环境中也是如此。这些结果表明,仅追求这一纯内在目标就足以驱动高度复杂行为的涌现。我们相信,这一新框架为智能体—环境系统提供了一种有原则的机制,使其无需外部奖励、任务或目标,即可维持一个日益复杂行为的开放式进程。
cs.LG / 153 / 2609.07596

I Don't Miss You, but I Do: Self-Explanation Faithfulness of Modality Missingness in Vision-Language Models

我没有想念你,但我确实想念:视觉-语言模型中模态缺失的自我解释忠实性
Javadov, Aydin, Schoess, Daniel, von Wangenheim, Florian
Abstract
Vision-language models are increasingly used in settings where some input modalities may be unavailable, yet we know little about whether they can faithfully explain how such missing information affects their own predictions. We introduce an interventional protocol for evaluating self-explanations of modality dynamics: models state what each modality alone would support, whether restoring a missing modality would change their answer, and whether the available evidence is sufficient; we then execute the corresponding modality intervention and compare these claims with the model's realized behavior. We evaluate eight open-weight VLMs from two model families across four tasks spanning complementary and isomorphic text-image settings and a multi-view driving setting. We find a systematic tendency to overstate the sufficiency of available modality evidence. Models substantially underestimate the effect of restoring missing modalities: task-level median predicted change rates are at most 8.8%, while the corresponding executed change rates reach 72.1%, with underprediction in 62 of 64 model-task-condition settings. Insufficiency claims are rare, but precise when produced: restoring the modality changes the answer in a median of 78-100% of flagged cases. Retrospective self-explanations show the same tendency: on complementary data, models over-credit single-modality sufficiency; on isomorphic data, they over-credit single representation sufficiency relative to their executed behavior. Together, these results show that VLMs systematically mischaracterize how their predictions depend on available and missing modality evidence, motivating executable interventions as a behavioral ground truth for evaluating multimodal self-explanations.
Chinese Translation
视觉-语言模型(Vision-Language Models, VLM)越来越多地被应用于部分输入模态可能不可用的场景,但我们对这些模型能否忠实解释缺失信息如何影响自身预测知之甚少。我们提出了一种干预式协议来评估模型对模态动态的自我解释:模型需陈述每个单独模态能支持什么、恢复缺失模态是否会改变其答案、现有证据是否充分;随后我们执行相应的模态干预,并将这些陈述与模型的实际行为进行比较。我们在来自两个模型家族的八个开源权重 VLM 上进行评估,涵盖互补与同构的文本-图像设置以及多视角驾驶设置共四项任务。我们发现模型存在系统性夸大现有模态证据充分性的倾向。模型大幅低估恢复缺失模态的影响:任务层面预测变化率中位数至多为 8.8%,而相应的实际执行变化率高达 72.1%,且在 64 个模型-任务-条件设置中有 62 个出现预测偏低。声称证据不足的情况较少见,但一旦出现则相当准确:在被标记的案例中,恢复该模态会使答案改变的比例中位数达 78-100%。回溯性自我解释也呈现同样趋势:在互补数据上,模型过度归因于单模态充分性;在同构数据上,相对于其实际执行行为,模型过度归因于单一表征的充分性。这些结果共同表明,VLM 系统性地错误刻画其预测对现有及缺失模态证据的依赖方式,从而支持将可执行干预作为评估多模态自我解释的行为学基准(ground truth)。
cs.LG / 154 / 2609.07597

Beyond the Matrix Sign: Quadratic Spectral Descent

超越矩阵符号:二次谱下降
Zhang, Qiaozhe, Sun, Jun, Liu, Yingzhuang
Abstract
Muon can be interpreted as optimizing a linear local objective over a spectral-norm ball. This gives a matrix-sign update that preserves the singular directions of the gradient and assigns the same magnitude to all active singular modes. We ask whether these two properties remain optimal when local curvature is taken into account. To answer this question, we keep Muon's spectral-norm constraint unchanged and replace the linear local model with a quadratic one. We call the resulting method \emph{Quadratic Spectral Descent} (QSD). We show that curvature can change both the singular values and the singular directions of the optimal update. To make QSD practical, we approximate curvature with Kronecker-factored statistics and solve the constrained quadratic with a small number of Frank--Wolfe steps, each of which has a closed-form matrix-sign subproblem. We further provide an optimality certificate, a comparison with Muon under the same quadratic surrogate, and an $O(1/K)$ convergence rate for the inner solver. Experiments on GPT pre-training show that QSD consistently improves validation loss over Muon and recent Muon variants, and reduces wall-clock training time by up to $8.49\%$ at matched validation loss.
Chinese Translation
Muon 可以被解释为在谱范数球上优化一个线性局部目标,由此得到的矩阵符号更新能够保留梯度的奇异方向,并为所有激活的奇异模式赋予相同的幅度。我们提出了一个问题:当考虑局部曲率时,这两个性质是否仍然是最优的?为回答该问题,我们保持 Muon 的谱范数约束不变,将线性局部模型替换为二次模型,并将由此得到的方法称为二次谱下降(Quadratic Spectral Descent,QSD)。我们证明,曲率可以同时改变最优更新的奇异值和奇异方向。为使 QSD 具有实用性,我们使用 Kronecker 分解统计量近似曲率,并通过少量 Frank--Wolfe 步骤求解该约束二次问题,其中每一步都有一个具有闭式解的矩阵符号子问题。我们进一步提供了最优性证书、在同一二次代理模型下与 Muon 的比较,以及内层求解器 $O(1/K)$ 的收敛速率。在 GPT 预训练上的实验表明,QSD 在验证损失上始终优于 Muon 及近期的 Muon 变体,且在相同验证损失下可将实际训练时间最多减少 $8.49\%$。
cs.LG / 155 / 2609.07606

CLUES-WEASEL: No additional clues required to choose your time series clustering algorithm

CLUES-WEASEL:无需额外线索即可选择你的时间序列聚类算法
Faouzi, Johann
Abstract
Time series data is very common in many real-world applications and in numerous domains, with increasing interest for automated information extraction using machine learning. One of these subfields is time series clustering, which consists in identifying clusters among a set of time series in an unsupervised fashion. Most time series clustering algorithms suffer from the same balancing act: they trade clustering performance for faster runtimes or vice versa. We present a novel time series clustering algorithm that we call CLUES-WEASEL, which stands for CLustering with the UnsupervisEd Second version of Word ExtrAction for time SEries cLassification. CLUES-WEASEL extracts features using the unsupervised version of the transformation step of WEASEL 2.0, which is a time series classification algorithm, then reduces these features using principal component analysis, and finally performs clustering with the $k$-means algorithm using these reduced extracted features. Through extensive experiments, we prove that CLUES-WEASEL is significantly better than any other existing time series clustering algorithm while being (much) faster than any state-of-the-art one. We also show that the architecture of CLUES-WEASEL can work well with other time series feature extraction algorithms. Our findings highlight the relevance of CLUES-WEASEL for time series clustering.
Chinese Translation
时间序列数据广泛存在于众多现实应用和领域中,利用机器学习进行自动化信息提取的兴趣日益增长。时间序列聚类是其中一个子领域,其目标是以无监督的方式在一组时间序列中识别聚类。大多数时间序列聚类算法都面临同样的平衡难题:它们以聚类性能换取更快的运行速度,或反之。我们提出了一种新颖的时间序列聚类算法,称为CLUES-WEASEL,全称为CLustering with the UnsupervisEd Second version of Word ExtrAction for time SEries cLassification。CLUES-WEASEL首先使用时间序列分类算法WEASEL 2.0变换步骤的无监督版本提取特征,然后利用主成分分析对这些特征进行降维,最后使用降维后的提取特征通过$k$-means算法进行聚类。通过大量实验,我们证明CLUES-WEASEL显著优于现有任何其他时间序列聚类算法,同时(远)快于任何最先进的算法。我们还表明,CLUES-WEASEL的架构可以与其他时间序列特征提取算法良好配合。我们的研究结果凸显了CLUES-WEASEL在时间序列聚类中的重要性。
cs.LG / 156 / 2609.07610

Translation of Black-Box Clinical Prediction Models into Standalone Transparent Nomograms: Temporal External Validation in Heart Transplantation

将黑箱临床预测模型转化为独立透明的列线图:在心脏移植中的时间外部验证
Pigot, Henry, Lisboa, Paulo J. G., Ortega-Martorell, Sandra, Olier, Ivan, Mahon, Joseph, Nilsson, Johan
Abstract
We convert black-box clinical prediction models for tabular data into standalone nomograms that can be audited term by term. PRiSM (Partial Responses in Structured Models) takes the shape of each effect and interaction from the source model, not merely which variables mattered, and lets the outcome select and weight them. We tested this in 50,356 heart transplant recipients, with validation in a later era than training. Nomograms from all 5 source models - a public clinical risk score, logistic regression, neural networks, random forests and extreme gradient boosting - met a prespecified noninferiority criterion for discrimination before any further simplification, and generally preserved calibration and clinical net benefit. Those from the 3 machine-learning models showed no detectable difference in discrimination from de novo generalized additive and explainable boosting models, exceeded neural additive models, and carried fewer terms than the explainable boosting model. PRiSM is released as an open-source Python package.
Chinese Translation
我们将面向表格数据的黑箱临床预测模型转化为可逐项审查的独立列线图。PRiSM(Partial Responses in Structured Models,结构化模型中的部分响应)不仅提取源模型中哪些变量重要,还提取每个效应和交互作用的形状,并由结局变量对其进行选择和赋权。我们在50,356名心脏移植受者中对该方法进行了测试,验证数据来自训练时期之后的时期。来自全部5个源模型——公开临床风险评分、逻辑回归、神经网络、随机森林和极端梯度提升(XGBoost)——的列线图,在未做任何进一步简化之前即满足了预设的判别能力非劣效标准,并总体上保持了校准度和临床净获益。来自3个机器学习模型的列线图与从头构建的广义可加模型和可解释提升模型在判别能力上无可检测的差异,优于神经可加模型,且所含项数少于可解释提升模型。PRiSM已作为开源Python软件包发布。
cs.LG / 157 / 2609.07617

Forecasting the Winner of a Live Tennis Match

预测网球比赛的实时胜者
Xie, Charles, Muppidi, Aneesh
Abstract
With the rise of live sports betting in recent years, tennis forecasting has expanded from pre-match prediction to models that update win probabilities as a match unfolds. A central challenge in creating such a model is the constant need for models to adapt to score and performance changes. This study examines how pre-match and live information can be most effectively integrated into a model to produce accurate win-probability estimates. The analysis uses 8,222 Grand Slam matches containing a total of 1,505,355 points. Five models were evaluated using a chronological split, with matches from 2011-2021 used for training, 2022 for validation, and 2023-2024 for testing. Trace, a hybrid model, achieved accuracies of 76.06%, 82.15%, and 88.34% at 25%, 50%, and 75% match progress, suggesting that hybrid modeling is a practical approach to live tennis forecasting.
Chinese Translation
随着近年来体育实时博彩的兴起,网球赛事预测已从赛前预测扩展到随着比赛进行而动态更新获胜概率的模型。构建此类模型的一个核心挑战在于模型需要不断适应比分和表现的变化。本研究探讨了如何将赛前信息和实时信息最有效地整合到模型中,以产生准确的获胜概率估计。分析使用了8,222场大满贯比赛,共包含1,505,355个回合(分)。研究采用按时间顺序划分的方法评估了五个模型:2011-2021年的比赛用于训练,2022年用于验证,2023-2024年用于测试。混合模型Trace在比赛进行至25%、50%和75%时分别达到了76.06%、82.15%和88.34%的预测准确率,表明混合建模是实时网球预测的一种实用方法。
cs.LG / 158 / 2609.07655

Online Surrogate Repair: Decoupling High-Fidelity Feedback from Search Length in Closed-Loop Discovery

在线代理模型修复:在闭环发现中将高保真反馈与搜索长度解耦
Feng, Xiaotang, Torr, Philip, Andreis, Bruno
Abstract
Closed-loop AI scientists can generate candidate designs at low marginal computational cost, whereas reliable feedback may require wet-lab synthesis, characterization, or high-fidelity computation. Addressing this imbalance through custom laboratory automation remains infrastructure-intensive and costly, while replacing new experiments with a fixed surrogate leaves persistent model errors that can be amplified by optimization. We propose \emph{online surrogate repair} (OSR), a closed-loop algorithm that uses sparse high-fidelity evaluations to update the surrogate throughout a longer agent search conducted primarily with inexpensive surrogate feedback. An acquisition rule selects which designs from the agent's accumulated proposals receive high-fidelity evaluation, and the resulting labels update the surrogate used in subsequent episodes. Across controlled synthetic environments, we demonstrate that improving global surrogate fit does not necessarily reduce maximum regret, whereas Q90-UCB and expected improvement (EI) substantially reduce regret by directing evaluations toward regions that determine the optimizer's decisions. On MADE, controls receiving high-fidelity feedback after every episode require $6.36$--$7.23\times$ more oracle queries to match Online EI under two LLM orchestrators and $10.27\times$ more under the non-LLM Chemeleon+MLIP workflow. Online surrogate repair introduces a novel third feedback regime between fixed-surrogate operation and high-fidelity feedback after every episode, separating the frequency of high-fidelity evaluation from the duration of the agent's search.
Chinese Translation
闭环AI科学家能够以较低的边际计算成本生成候选设计,而可靠的反馈可能需要湿实验合成、表征或高保真计算。通过定制实验室自动化来解决这一不平衡问题仍然需要大量基础设施投入且成本高昂,而用固定的代理模型替代新实验则会留下持久的模型误差,这些误差还可能被优化过程放大。我们提出在线代理模型修复(Online Surrogate Repair, OSR),这是一种闭环算法,它利用稀疏的高保真评估来更新代理模型,同时进行主要由廉价的代理反馈驱动的更长的智能体搜索。一个采集规则从智能体累积的提案中选择哪些设计接受高保真评估,所得的标签用于更新后续回合中使用的代理模型。在受控合成环境中,我们证明改善代理模型的全局拟合并不一定会降低最大遗憾值(maximum regret),而Q90-UCB和期望改进(Expected Improvement, EI)通过将评估引导至决定优化器决策的区域,显著降低了遗憾值。在MADE基准上,在每个回合后都接受高保真反馈的对照方法,在两个LLM编排器下需要多$6.36$–$7.23$倍的oracle查询才能匹敌Online EI,而在非LLM的Chemeleon+MLIP工作流下需要多$10.27$倍。在线代理模型修复在固定代理模型运行与每回合后高保真反馈之间引入了一种新颖的第三种反馈模式,将高保真评估的频率与智能体搜索的持续时间分离开来。
cs.LG / 159 / 2609.07664

Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs

准确率并不足够:一种基于散度的方法评估量化大语言模型中的保真度损失
Qamar, Shahzeb, Sparrenberg, Lorenz, Bauckhage, Christian, Rababah, Baha, Leung, Carson, Kantarcioglu, Murat, Akcora, Cuneyt Gurcan, Sifa, Rafet
Abstract
Deployment of Large Language Models (LLMs) on memory-constrained edge devices relies heavily on aggressive post-training quantization. However, evaluating these models is largely based on zero-shot task accuracy, which depends solely on argmax predictions and is insensitive to changes in the underlying predictive distribution. Consequently, accuracy can exhibit unstable, non-monotonic behavior under progressive quantization, masking substantial fidelity loss relative to the BFloat16 (BF16) uncompressed base model and providing misleading deployment signals. We introduce a distribution-sensitive evaluation framework quantifying information loss in quantized LLMs as the divergence between full-vocabulary predictive distributions at the token decision boundary. We compute statistical distances, including Jensen-Shannon Divergence and Total Variation Distance, between outputs of full-precision and quantized models, enabling a fine-grained analysis of distributional shift. Using this framework, we quantify probability mass displacement and distributional drift relative to the BF16 reference, capturing predictive distribution changes not reflected in top-1 accuracy. We conduct a 120-run experimental matrix across five foundation architectures and four reasoning benchmarks under progressive quantization regimes, from uncompressed BF16 to Q2_K, providing a systematic fidelity analysis. Our results show divergence metrics generally increase under stronger quantization, complementing task accuracy with a fidelity signal. Across tested llama.cpp schemes, mixed-precision Q4_K generally yields lower divergence than uniform Q4_0 at similar memory footprints. These findings motivate distribution-aware evaluation as a practical diagnostic complement to task accuracy; they do not directly establish correctness, calibration, safety, or user-perceived quality.
Chinese Translation
在内存受限的边缘设备上部署大语言模型(LLMs)严重依赖激进的后训练量化。然而,对这些模型的评估主要基于零样本任务准确率,而准确率仅取决于argmax预测,对底层预测分布的变化并不敏感。因此,在渐进量化过程中,准确率可能表现出不稳定、非单调的行为,掩盖了相对于BFloat16(BF16)未压缩基础模型的显著保真度损失,并提供具有误导性的部署信号。我们提出了一个分布敏感的评估框架,将量化LLMs中的信息损失量化为token决策边界处全词表预测分布之间的散度。我们计算全精度模型与量化模型输出之间的统计距离,包括Jensen-Shannon散度(Jensen-Shannon Divergence)和总变差距离(Total Variation Distance),从而实现对分布偏移的细粒度分析。利用该框架,我们量化了相对于BF16参考模型的概率质量位移和分布漂移,捕捉到top-1准确率无法反映的预测分布变化。我们在渐进量化方案(从未压缩的BF16到Q2_K)下,对五种基础架构和四个推理基准开展了包含120次运行的实验矩阵,提供了系统性的保真度分析。结果表明,散度指标总体上随量化强度增大而上升,为任务准确率提供了保真度信号方面的补充。在所测试的llama.cpp量化方案中,在相近内存占用下,混合精度的Q4_K通常比均匀的Q4_0产生更低的散度。这些发现表明,分布感知的评估可作为任务准确率的一种实用诊断补充;但它们并不能直接确立正确性、校准质量、安全性或用户感知质量。
cs.LG / 160 / 2609.07666

MpSub: A Momentum $p$-Dimensional Subspace Trust-Region Method for Derivative-Free Fine-Tuning of Large Language Models

MpSub:一种用于大语言模型无导数微调的动量p维子空间信赖域方法
Wang, Yuyang, Yao, Haoyu, Xie, Pengcheng
Abstract
Full-parameter fine-tuning of large language models has substantial memory costs because backpropagation stores activations and gradients. Zeroth-order optimization avoids this by estimating update directions from loss evaluations, but existing methods require tuning a sensitive learning rate for each model and task. We propose the momentum $p$-dimensional subspace trust-region method (MpSub). At each iteration, MpSub searches within a $p$-dimensional subspace: one direction preserves historical momentum from the most recent accepted step, while the remaining directions explore via fresh random sampling. The subspace gradient is estimated by central differences, a trial step is computed from a linear trust-region model, and the trust-region radius adapts according to the agreement between predicted and observed loss reduction, eliminating the learning rate. For LLM fine-tuning, evaluations within an iteration share a minibatch, and directions are regenerated in place from seeds, using forward passes alone. For smooth deterministic objectives under unorthogonalized Gaussian directions, we bound the finite-difference error, quantify gradient energy captured by the subspace, and prove that $\lim_{k\to\infty} \|\nabla f(x_k)\|_2 = 0$ almost surely under a safeguarded radius update. Under a matched budget of 8,400 training-objective forward passes, we fine-tune OPT-125M and OPT-350M on CommitmentBank. With the same preset parameters at both model sizes, MpSub attains mean test accuracies of 0.673 and 0.690 over three seeds, matching tuned MeZO (0.685) without any learning-rate search.
Chinese Translation
大语言模型的全参数微调具有高昂的内存开销,因为反向传播需要存储激活值和梯度。零阶优化通过损失函数评估来估计更新方向,从而避免了这一问题,但现有方法需要为每个模型和任务调节一个敏感的学习率。我们提出动量p维子空间信赖域方法(MpSub)。在每次迭代中,MpSub在一个p维子空间内进行搜索:其中一个方向保留最近一次被接受步的历史动量,其余方向则通过新的随机采样进行探索。子空间梯度由中心差分估计,试验步由线性信赖域模型计算得出,信赖域半径根据预测损失下降与实际观测损失下降的一致性进行自适应调整,从而消除了学习率。在大语言模型微调中,同一迭代内的评估共享一个小批量数据,且各方向利用随机种子就地重新生成,仅使用前向传播。对于在未经正交化的高斯方向下的光滑确定性目标函数,我们给出了有限差分误差的上界,量化了子空间所捕获的梯度能量,并在带保护的半径更新策略下证明了 $\lim_{k o\infty} \| abla f(x_k)\|_2 = 0$ 几乎必然成立。在8,400次训练目标前向传播的相同预算下,我们在CommitmentBank数据集上微调OPT-125M和OPT-350M。在两种模型规模下使用相同的预设参数,MpSub在三个随机种子上分别取得了0.673和0.690的平均测试准确率,在无需任何学习率搜索的情况下达到了经过调参的MeZO的水平(0.685)。
cs.LG / 161 / 2609.07681

On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing

论Mamba中的召回率缩放定律:基于哈希的理论与机制研究
Koren, Yuval, Ben-Kish, Assaf, Giryes, Raja, Wolf, Lior, Zimerman, Itamar
Abstract
Associative Recall (AR) is the cognitive ability to learn and retrieve links between items in memory. In NLP, AR is used as a benchmark for evaluating the in-context memory capacity of architectures such as Mamba, and has been found to strongly correlate with language modeling performance. This paper explores AR from the perspective of mechanistic interpretability, aiming to reverse-engineer the exact internal algorithm used by Mamba to perform recall. Our key insight is that Mamba performs recall by implicitly learning linear hash functions, and we identify the low-level circuit that enables this behavior. Building on these findings and inspired by theoretical tools in similarity-preserving hashing, such as the Johnson-Lindenstrauss lemma, we develop a theoretical framework for analyzing AR, which we term Recall Scaling Laws. Given the vocabulary size and the number of facts in context, this framework allows us to (1) predict the embedding and state dimensions required for Mamba to achieve perfect recall, (2) predict recall success probability given the model dimensions, and (3) analyze multi-layer models and multi-head SSM patterns. Empirical results show that our theoretical findings are accurate and predictive, offering insights into how AR capacity scales with vocabulary, state, embedding size, and architecture.
Chinese Translation
联想召回(Associative Recall, AR)是一种学习和检索记忆中条目之间关联关系的认知能力。在自然语言处理(NLP)中,AR被用作评估Mamba等架构的上下文记忆容量的基准,并被发现与语言建模性能高度相关。本文从机制可解释性的角度探讨AR,旨在逆向工程Mamba执行召回任务所使用的确切内部算法。我们的关键发现是:Mamba通过隐式学习线性哈希函数来执行召回,并且我们识别出实现这一行为的底层电路。基于这些发现,并受到保相似性哈希理论工具(如Johnson-Lindenstrauss引理)的启发,我们建立了一个用于分析AR的理论框架,我们称之为召回率缩放定律(Recall Scaling Laws)。在给定词表大小和上下文中事实数量的情况下,该框架使我们能够:(1) 预测Mamba实现完美召回所需的嵌入维度和状态维度;(2) 在给定模型维度的情况下预测召回成功概率;(3) 分析多层模型及多头SSM模式。实证结果表明,我们的理论发现准确且具有预测性,为理解AR容量如何随词表、状态、嵌入规模和架构而扩展提供了洞见。
cs.LG / 162 / 2609.07689

Emergent Charging Coordination in Electric Delivery Fleets

电动配送车队中涌现的充电协调
Vales-Alonso, Javier, Alcaraz, Juan J.
Abstract
In electric delivery fleets, mid-shift charging is non-trivial: each vehicle must decide when, where and how much to charge to finish on time with battery above a safety floor. The choices are coupled: queues build where too many vehicles pick the same station. Prior work resolves this coupling with central dispatching, precomputed schedules or reservations, machinery that charging infrastructure rarely supports. Instead, we use a family of learning agents under purely local control: every vehicle runs the same policy, deciding alone from its time budgets and broadcast station occupancies, leading to emergent coordination without central control or messaging. We validate this paradigm in simulation on real OpenStreetMap networks of twenty cities, each with a frozen scenario calibrated by an omniscient Oracle (99.5% of shifts completed on time), whereas a naive greedy rule (nearest station on low battery) completes just 73%. Agents trained with neuroevolution (NEAT) and policy gradients (PPO) on four cities and deployed zero-shot across all twenty, sixteen never seen in training, complete 96.8% and 98.6% of shifts, with the policy-gradient controllers proving more robust when demand or vehicle characteristics drift beyond the trained regime. In contrast, tuned threshold heuristics that read vehicle urgency alone fall short in contended cities (~80%). Through training, these learning agents rediscover partial charging and short opportunistic sessions, and route around busy stations, cutting per-session queue waits from about 45 minutes to under 2. In summary, this coordination paradigm balances local urgency against public occupancy, reaching near-Oracle performance at minimal implementation cost.
Chinese Translation
在电动配送车队中,班中充电并非易事:每辆车必须决定何时、何地以及充多少电,才能准时完成任务并保持电量高于安全底线。这些选择是相互耦合的:当过多车辆选择同一充电站时会形成排队。先前的工作通过集中调度、预计算排程或预约机制来解决这种耦合问题,但这些机制往往缺乏相应的充电基础设施支持。本文转而采用完全局部控制下的一类学习智能体:每辆车运行相同的策略,仅根据自身的时间预算和广播的充电站占用情况独立决策,从而在没有中央控制或消息传递的情况下实现涌现式协调。我们在二十个城市的真实OpenStreetMap路网上通过仿真验证了这一范式,每个城市的冻结场景均由全知Oracle校准(99.5%的班次准时完成),而简单的贪心规则(低电量时选择最近的充电站)仅完成73%。在四个城市上使用神经进化(NEAT)和策略梯度(PPO)训练的智能体,在全部二十个城市(其中十六个未在训练中出现过)零样本部署后,分别完成96.8%和98.6%的班次;当需求或车辆特性偏离训练范围时,策略梯度控制器表现出更强的鲁棒性。相比之下,仅依据车辆紧急程度调优的阈值启发式方法在拥堵城市中表现欠佳(约80%)。通过训练,这些学习智能体自发地重新发现了部分充电和短时机会性充电策略,并绕开繁忙的充电站,将每次充电的排队等待时间从约45分钟降至2分钟以内。总之,该协调范式在局部紧急性与公共占用信息之间取得平衡,以极低的实现成本达到接近Oracle的性能。
cs.LG / 163 / 2609.07706

ParetoTransport: Generative Optimization by Mass Transport Toward The Pareto Front

ParetoTransport:基于质量输运的面向Pareto前沿的生成式优化
Holly, Stephanie, Hochreiter, Sepp, Zellinger, Werner
Abstract
Offline multi-objective optimization requires not only moving the objective vectors of candidate designs toward the Pareto front, but also distributing them effectively along it. Generative methods have recently emerged as a natural approach because they learn a distribution over feasible designs while allowing generation to be steered toward promising designs. Existing methods, however, largely retain classical sample-wise guidance strategies, leaving the distribution-level modeling capability of generative methods underused. We propose ParetoTransport, a training-free guidance method for pre-trained flow-matching models that explicitly specifies and refines a population-level distribution in objective space. ParetoTransport guides a flow-matching sampler to iteratively transport the empirical offline distribution toward the Pareto front, with Wasserstein matching to intermediate proxy distributions. This directly controls distributional displacement and mass allocation along the front. We establish a convergence result and demonstrate state-of-the-art performance on standard offline MOO benchmarks, extending recent evaluations beyond hypervolume to generational distance, inverted generational distance, and Wasserstein distance.
Chinese Translation
离线多目标优化不仅需要将候选设计的目标向量向Pareto前沿推进,还需要将它们有效地分布在前沿上。生成式方法近来成为一种自然的选择,因为它们能够学习可行设计的分布,同时允许将生成过程引导至有前景的设计。然而,现有方法大多沿用经典的样本级引导策略,使得生成式方法在分布层面的建模能力未被充分利用。我们提出ParetoTransport,这是一种面向预训练流匹配模型的无训练引导方法,能够在目标空间中显式地指定并细化种群级分布。ParetoTransport引导流匹配采样器将经验离线分布迭代地向Pareto前沿输运,并通过Wasserstein距离与中间代理分布进行匹配。这直接控制了沿前沿的分布位移和质量分配。我们建立了收敛性结果,并在标准离线多目标优化基准上展示了最先进的性能,将近期评估从超体积指标扩展到世代距离、反世代距离和Wasserstein距离。
cs.LG / 164 / 2609.07729

Attributing Cohen's d: Training Data Attribution for Disease-Related Effects in Normative Age Biomarkers

归因Cohen's d:面向标准年龄生物标志物中疾病相关效应的训练数据归因
Snel, Jakob, Schulz, Marc-Andre
Abstract
Normative age models are trained to predict chronological age in a nominally healthy cohort. Applied to patients, they deviate, and the gap between predicted and chronological age is read as disease risk. Here, we attribute the disease-related effect size of the age gap directly to individual training samples, rather than using a prediction-level loss as the attribution target. For Cohen's $d$, the resulting closed-form influence functional, validated against leave-one-out retraining, ranks training samples by their effect on held-out case-control separation. Across four diseases and two biomarker modalities in UK Biobank, removing the 10% most influential training samples raises held-out disease-related effect size in every seed. It more than doubles the metabolomic-age effect for type-2 diabetes and raises the brain-age effect for multiple sclerosis by roughly a third. Random removal leaves effect size flat even at 50% removal, confirming the gain comes from which samples are removed, not how many. Flagged subjects carry subclinical cardiometabolic burden that diagnosis-based exclusion misses, on markers the model never sees. For type-2 diabetes, where the method gains most, the marker recovered is HbA1c, the standard measure of blood sugar control. We release pyinfluence, our influence-function package, for reproducibility and reuse.
Chinese Translation
标准年龄模型(normative age models)通过名义上健康的队列训练来预测实际年龄。应用于患者时,模型预测会产生偏差,预测年龄与实际年龄之间的差距被解读为疾病风险。本研究将年龄差距的疾病相关效应量直接归因于个体训练样本,而非使用预测层面的损失作为归因目标。对于Cohen's $d$,我们推导出闭式影响函数,并通过留一法重训练(leave-one-out retraining)加以验证,该函数可依据训练样本对留出集病例-对照分离效果的影响对其进行排序。在英国生物样本库(UK Biobank)的四种疾病和两种生物标志物模态上,移除最具影响力的10%训练样本后,每个随机种子下的留出集疾病相关效应量均有所提升。对于2型糖尿病,代谢组学年龄效应提升了一倍以上;对于多发性硬化症,脑年龄效应提升了约三分之一。而随机移除样本即使达到50%,效应量仍保持不变,这证实收益来自移除的是哪些样本,而非移除样本的数量。被标记的受试者携带亚临床心脏代谢负担,而这些负担是基于诊断的排除方法在模型从未见过的标志物上所遗漏的。对于该方法收益最大的2型糖尿病,所恢复的标志物是HbA1c,即血糖控制的标准测量指标。我们发布了影响函数工具包pyinfluence,以便于复现和复用。
cs.LG / 165 / 2609.07749

Guiding Worker Self-Selection in Crowdsourcing Contests: An LLM-Augmented Algorithmic Approach

众包竞赛中的工作者自选择引导:一种大语言模型增强的算法方法
Thach, Nguyen, Chan, Hau, Parkes, David, Lakhani, Karim
Abstract
Crowdsourcing platforms coordinate large pools of online workers who strategically choose which contests to enter and how much effort to invest. This self-selection can leave important contests with too few participants or too little effort, while workers may regret entering contests that leave them worse off than available alternatives. We study how platforms can recommend contests to workers using self-selection in Tullock contests (SSTC), a two-stage model in which workers first choose contests and then compete within them. We introduce GRAF, a greedy polynomial-time framework that constructs self-selection outcomes by ordering workers according to a score vector, with guarantees of zero worker regret and platform optimality in special cases of SSTC. Because effective orderings are difficult to design under worker heterogeneity, we propose LLMScore, an LLM-driven evolutionary framework that automatically designs GRAF's scoring algorithm. LLMScore addresses two challenges: jointly optimizing platform utility and worker satisfaction, and evaluating worker regret when exact computation is intractable. Trained only on small instances of one setting, it transfers to larger and structurally different settings; moreover, its output is human-readable code that platform operators can inspect and modify. Across 1,000 synthetic instances spanning four settings, GRAF with LLMScore consistently achieves high-quality, often near-optimal, outcomes with low worker regret, benefiting both platforms and workers.
Chinese Translation
众包平台需要协调大量在线工作者,这些工作者会策略性地选择参加哪些竞赛以及投入多少努力。这种自选择可能导致重要竞赛参与者过少或投入不足,而工作者也可能因参加了相比其他可选机会使自己处境更差的竞赛而感到后悔。我们研究平台如何利用Tullock竞赛中的自选择(Self-Selection in Tullock Contests, SSTC)为工作者推荐竞赛,该模型分为两个阶段:工作者首先选择竞赛,然后在竞赛内部展开竞争。我们提出了GRAF,一个贪心多项式时间框架,通过按照分数向量对工作者排序来构建自选择结果,并在SSTC的特定情形下保证工作者零后悔和平台最优性。由于在工作者异质性条件下难以设计有效的排序,我们进一步提出LLMScore,一个由大语言模型驱动的进化框架,可自动设计GRAF的评分算法。LLMScore解决了两个挑战:联合优化平台效用与工作者满意度,以及在精确计算不可行时评估工作者后悔度。该框架仅在某一设定的小规模实例上训练,即可迁移到规模更大且结构不同的设定;此外,其输出为人类可读的代码,平台运营者可以审查和修改。在涵盖四种设定的1,000个合成实例上,结合LLMScore的GRAF始终能获得高质量(通常接近最优)的结果,且工作者后悔度低,使平台和工作者双方均受益。
cs.LG / 166 / 2609.07752

Local gradient neural operator

局部梯度神经算子
Zhang, Baiming, Tang, Jinsong, Xu, Ying, Chen, Lihua, Xiong, Shiying
Abstract
Field temporal prediction and source identification constitute canonical problems in dynamical systems. Conventional approaches to these problems depend on a thorough understanding of the governing partial differential equations (PDEs). Recently, deep learning, as represented by neural operators, has provided a data-driven paradigm for addressing such tasks. However, most existing global neural operators for PDEs require large training datasets and many learnable parameters, with limited interpretability and generalization. We propose the local gradient neural operator (LGNO) as a lightweight and interpretable alternative for field temporal evolution prediction and source identification in typical mechanical problems. The method builds on priors from nonlinear gradient discretization and uses multilayer perceptron convolutional layers to learn translation-invariant local kernels that resemble discrete stencils. A zero consistent stencil factorization separates coefficient learning from field reconstruction, rendering the learned operators more transparent. For problems with symmetries, network folding shares equivalent components and reduces parameter counts. We evaluate the method on PDE benchmarks covering linear and nonlinear, static and dynamic, and low and high dimensional cases. Results show that LGNO maintains accuracy, parameter efficiency, and rollout stability across these tasks, and further exhibits wide applicability to mechanical problems including diffusion, flow, and quantum phenomena.
Chinese Translation
场的时序预测与源识别是动力系统中的典型问题。解决这些问题的传统方法依赖于对控制偏微分方程(PDE)的透彻理解。近年来,以神经算子(neural operators)为代表的深度学习为处理此类任务提供了一种数据驱动的范式。然而,现有的大多数用于PDE的全局神经算子需要大量的训练数据和许多可学习参数,且可解释性和泛化能力有限。我们提出局部梯度神经算子(Local Gradient Neural Operator, LGNO),作为典型力学问题中场时序演化预测与源识别的一种轻量且可解释的替代方案。该方法建立在非线性梯度离散化的先验知识之上,利用多层感知机卷积层来学习类似离散模板的平移不变局部核。零一致性模板分解将系数学习与场重构分离,使所学习的算子更加透明。对于具有对称性的问题,网络折叠可共享等价组件并减少参数数量。我们在涵盖线性与非线性、静态与动态、低维与高维情形的PDE基准问题上评估了该方法。结果表明,LGNO在这些任务中保持了精度、参数效率和滚动预测稳定性,并进一步展现出对包括扩散、流动和量子现象在内的力学问题的广泛适用性。
cs.LG / 167 / 2609.07755

A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay

权重衰减梯度下降下神经网络泛化动力学的理论分析
Wang, Yuqing, Kevrekidis, Ioannis G., Belkin, Mikhail
Abstract
Understanding generalization remains a central challenge in machine learning because it requires jointly considering data, architecture, and training dynamics. In this paper, we develop a theoretical framework that characterizes how these factors jointly shape generalization performance throughout training. More precisely, we study a broad class of neural networks trained under the $\ell^2$ loss by gradient descent (GD) with weight decay, and prove the convergence of GD to a neighbourhood of the global minimizers of the empirical loss. By partitioning the space based on the input data, we then decompose the population error into data error, optimization error, and prediction variation error, and bound them separately. In particular, for the prediction variation error, which measures the oscillations of the learned function, we propose (local) approximate homogeneity and derive explicit cellwise and layerwise bounds for its evolution along the training trajectory. These bounds yield two important implications: a necessary condition of improved generalization explains differences in layerwise generalization behavior; a sufficient condition describes delayed generalization and provides a theoretical characterization of grokking.
Chinese Translation
理解泛化问题仍然是机器学习中的核心挑战,因为它需要同时考虑数据、架构和训练动力学。在本文中,我们建立了一个理论框架,刻画这些因素如何在训练过程中共同塑造泛化性能。更准确地说,我们研究了在 $\ell^2$ 损失下采用带权重衰减的梯度下降(GD)训练的一类广泛的神经网络,并证明了 GD 收敛到经验损失全局极小值点的一个邻域。通过基于输入数据对空间进行划分,我们将总体误差分解为数据误差、优化误差和预测波动误差,并分别给出其上界。特别地,对于度量所学函数振荡程度的预测波动误差,我们提出了(局部)近似齐次性的概念,并沿训练轨迹推导出其演化的显式的逐单元(cellwise)和逐层(layerwise)界。这些界带来两个重要推论:一个泛化改善的必要条件解释了逐层泛化行为的差异;一个充分条件描述了延迟泛化现象,并为 grokking 提供了理论刻画。
cs.LG / 168 / 2609.07756

Decomposition-Guided Diffusion Language Models for Inertial Confinement Fusion Prediction

面向惯性约束聚变预测的分解引导扩散语言模型
Zhang, Xiang, Gopalaswamy, Varchas, Ejaz, Rahman, Betti, Riccardo, Liu, Dongfang
Abstract
Inertial confinement fusion (ICF) is a leading pathway toward clean energy, but each shot at the National Ignition Facility costs on the order of one million dollars, making accurate AI surrogates a high-value target. We study exogenous-driven ICF waveform prediction, where a 512-step neutron-rate diagnostic must be inferred directly from a laser pulse and target design parameters, with no historical response observed. The regime stresses standard time-series predictors with temporal sparsity (picosecond peak in a nanosecond window), input-output scale mismatch (under 300 real shots), and peak sensitivity (picosecond timing). We propose ICF-DLM, to our knowledge the first LM-based ICF predictor, combining (i) a physics-typed decomposition into yield $Y_{DT}$, peak timing $t_{\mathrm{peak}}$, and local waveform $w_{\mathrm{local}}$; (ii) bidirectional denoising that defers commitment to peak location; and (iii) a physics-driven PPO reward re-injecting metric structure across numeric tokens. On ICFBench (50K simulations + 232 experimental shots), ICF-DLM cuts peak-timing error from 11.6 to 9.2 steps over a matched autoregressive LLaMA-3-8B and outperforms classical sequence models and LLM-based time-series predictors. Beyond ICF, the recipe shows potential to address science domains with low data and sparse events.
Chinese Translation
惯性约束聚变(ICF)是实现清洁能源的重要途径之一,但国家点火装置(National Ignition Facility)每次打靶实验的成本高达约一百万美元,因此构建精确的人工智能代理模型具有极高的价值。我们研究外生驱动条件下的ICF波形预测问题,即需要仅根据激光脉冲和靶丸设计参数直接推断512步的中子产率诊断信号,且无法观测到任何历史响应。该场景对标准时间序列预测器提出了严峻挑战:时间稀疏性(纳秒量级窗口内的皮秒级尖峰)、输入输出尺度失配(真实实验打靶数据不足300次)以及峰值敏感性(皮秒级时间精度)。我们提出ICF-DLM,据我们所知这是首个基于语言模型的ICF预测器,其结合了:(i)物理类型化分解,将输出分解为聚变产额 $Y_{DT}$、峰值时刻 $t_{\mathrm{peak}}$ 和局部波形 $w_{\mathrm{local}}$;(ii)双向去噪机制,延迟对峰值位置的确定;(iii)物理驱动的PPO奖励,在数值token之间重新注入度量结构。在ICFBench数据集(5万次模拟加232次实验打靶数据)上,与同等配置的自回归LLaMA-3-8B相比,ICF-DLM将峰值时刻误差从11.6步降至9.2步,并优于经典序列模型和基于大语言模型的时间序列预测器。除ICF之外,该方法还有望应用于数据稀少且事件稀疏的其他科学领域。
cs.LG / 169 / 2609.07814

Latent-MoE: Domain-Aware Mixture-of-Experts for PDEs with Multi-Regime Physics

Latent-MoE:面向多物理状态偏微分方程的域感知混合专家模型
Wang, Hanwen, Perdikaris, Paris
Abstract
Physics-informed neural networks (PINNs) struggle on PDEs whose governing physics varies across the domain. We trace this to a structural property of standard coordinate networks: their neural tangent kernel (NTK) is translation-variant and lets training points of large coordinate magnitude disproportionately influence predictions elsewhere, producing long-range coupling and gradient conflict during training. We show analytically and empirically that mixture-of-experts (MoE) architectures with centered, compact-support routers yield a uniformly banded NTK whose kernel-regression weights decay exponentially with distance, localizing the learning. Building on this, we propose \emph{Latent-MoE}, which interleaves domain-aware MoE blocks within a shared backbone. Unlike FB-PINNs or X-PINNs, which rigidly partition both the domain and the parameters so that the parameters on different subdomains are updated independently, Latent-MoE is designed to preserve the localization benefit of domain-aware routing while allowing capacity to flow across regions through the shared backbone. On standard homogeneous-physics benchmarks Latent-MoE is competitive with established baselines; on benchmarks with multi-stage time-variable physics, where global models and rigid domain decompositions both fall into spurious solutions, it improves over them by more than an order of magnitude, with markedly reduced gradient conflict during training.
Chinese Translation
物理信息神经网络(PINNs)在控制物理随计算域变化的偏微分方程(PDEs)上表现不佳。我们将此问题追溯到标准坐标网络的一个结构性缺陷:其神经正切核(NTK)是平移变化的,使得坐标值较大的训练点不成比例地影响其他位置的预测,从而在训练过程中产生长程耦合与梯度冲突。我们从理论和实验上证明,采用居中且紧支撑路由器的混合专家(MoE)架构能够产生均匀带状的NTK,其核回归权重随距离呈指数衰减,从而实现学习的局部化。基于这一发现,我们提出了Latent-MoE,它在共享骨干网络中交错嵌入域感知的MoE模块。与FB-PINNs或X-PINNs不同——后者对计算域和参数均进行刚性划分,使不同子域上的参数被独立更新——Latent-MoE旨在保留域感知路由带来的局部化优势,同时允许模型容量通过共享骨干网络在不同区域之间流动。在标准的均匀物理基准测试中,Latent-MoE与现有基线方法具有竞争力;在具有多阶段时变物理的基准测试中——此类场景下全局模型和刚性域分解方法均会陷入虚假解——Latent-MoE的性能超越上述方法一个数量级以上,且训练过程中的梯度冲突显著减少。
cs.LG / 170 / 2609.07816

Kalman Delta Networks: Uncertainty-aware Associative Memory

卡尔曼Delta网络:不确定性感知的联想记忆
Bui, Ngoc, Huang, Tinglin, Ying, Rex
Abstract
Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memory decoding. Its fixed-size recurrent memory, however, requires an online decision at each token: what to write and how strongly to overwrite existing associations before knowing which information future queries will require. Delta-rule models learn this strength from the current token embedding but do not track confidence in the memory estimate, preventing each write from adapting to accumulated evidence. To represent this uncertainty explicitly, we reformulate recurrent associative memory as a linear--Gaussian state-space model, for which the Kalman filter is the optimal recursive estimator, and introduce a new family of models, Kalman Delta Networks (KDNs). Within KDNs, the transition propagates both the memory state and its uncertainty, allowing the Kalman gain to weight each residual write by accumulated evidence and observation reliability. Under this formulation, Delta-style updates emerge as a special case that substitutes a token-wise isotropic surrogate for predictive covariance and omits covariance tracking. Exact tracking, however, entails a dense, state-dependent Riccati recursion that is poorly suited to GPU-parallel linear-attention scans. To address this issue, we introduce two scan-compatible KDN approximations. Diagonal KDN projects each one-step posterior onto the diagonal Gaussian family through online mean-field variational inference, whereas Isotropic KDN uses an isotropic approximation with a single uncertainty scalar per head. Their uncertainty recurrences are Mobius maps, enabling associative scans with logarithmic parallel depth. Across controlled pretraining at 750M and 1.3B parameters, KDN variants consistently improve perplexity and mean downstream accuracy over state-of-the-art linear-attention models.
Chinese Translation
线性注意力因可实现高效的长上下文推理和常数内存解码,正被越来越多地应用于前沿语言模型中。然而,其固定大小的循环记忆机制要求在每个token处做出在线决策:即在尚不知道未来查询需要哪些信息的情况下,决定写入什么内容以及以多强的力度覆盖已有关联。Delta规则模型从当前token嵌入中学习这一写入强度,但并不追踪记忆估计的置信度,导致每次写入无法根据累积的证据进行自适应调整。为了显式地表征这种不确定性,我们将循环联想记忆重新表述为一个线性-高斯状态空间模型——对此类模型,卡尔曼滤波器是最优的递归估计器——并据此提出了一类新模型:卡尔曼Delta网络(Kalman Delta Networks, KDNs)。在KDN中,状态转移同时传播记忆状态及其不确定性,使得卡尔曼增益能够依据累积证据和观测可靠性对每次残差写入进行加权。在该表述下,Delta式更新可被视为一种特例,即用逐token的各向同性代理协方差替代预测协方差并省略协方差追踪。然而,精确的协方差追踪需要稠密的、依赖状态的Riccati递推,这难以适应GPU并行化的线性注意力扫描。为解决这一问题,我们提出了两种与扫描兼容的KDN近似方法。对角KDN通过在线平均场变分推断将每一步的单步后验投影到对角高斯分布族上;而各向同性KDN则采用各向同性近似,每个注意力头仅使用单一的不确定性标量。二者的不确定性递推均为莫比乌斯映射(Mobius maps),从而支持具有对数并行深度的联想扫描。在750M和1.3B参数规模的控制性预训练实验中,KDN各变体在困惑度和下游任务平均准确率上均持续超越最先进的线性注意力模型。
cs.LG / 171 / 2609.07853

Foundation Models for Generalizable Semantic and Goal-Oriented Communication

面向可泛化语义与目标导向通信的基础模型
Liu, Boliang, Poe, Wint Yi, Trivisonno, Riccardo, Caire, Giuseppe
Abstract
Semantic and goal-oriented communication is increasingly studied for 6G, but generalization beyond seen data remains a key weakness under tight rate budgets. Many existing systems overfit their training data and degrade sharply at very low bit rates because they attempt to compress the entire signal. We introduce Foundation Model-Guided Semantic and Goal-Oriented Communication (FMSGOC), a framework that uses broad visual-linguistic Foundation Model priors to mitigate overfitting. It further improves rate efficiency by concentrating bits on sparse, goal-aligned anchors and relying on generative foundation-model priors to reconstruct the masked regions. By decoupling what to send from how to reconstruct, a vision-language foundation model selects and transmits a sparse set of semantic anchors, while a pretrained diffusion model, fine-tuned for masked completion, reconstructs the image at the receiver. In our experiments, FMSGOC reaches 0.039 bits per pixel (BPP), maintains high semantic fidelity (cosine similarity 0.87-0.90 on CIFAR-10), remains robust on previously unseen inputs (0.83-0.86 on ImageNet), and shows good perceptual similarity (0.1278/0.1558, CIFAR-10/ImageNet), outperforming strong end-to-end baselines at lower bit rates.
Chinese Translation
语义通信与目标导向通信在6G领域的研究日益增多,但在严格的速率预算下,如何超越已见数据实现泛化仍是一个关键弱点。许多现有系统对训练数据过拟合,且由于试图压缩整个信号,在极低比特率下性能急剧下降。我们提出基础模型引导的语义与目标导向通信(Foundation Model-Guided Semantic and Goal-Oriented Communication, FMSGOC),该框架利用广泛的视觉-语言基础模型先验来缓解过拟合问题。它通过将比特集中于稀疏的、与目标对齐的锚点,并依靠生成式基础模型先验来重建被掩蔽区域,从而进一步提升速率效率。通过将发送什么与如何重建解耦,视觉-语言基础模型选择并传输稀疏的语义锚点集合,而经过掩蔽补全微调的预训练扩散模型则在接收端重建图像。实验结果表明,FMSGOC达到了0.039比特/像素(BPP)的速率,保持了较高的语义保真度(CIFAR-10上余弦相似度为0.87-0.90),在先前未见过的输入上依然稳健(ImageNet上为0.83-0.86),并展现出良好的感知相似性(CIFAR-10/ImageNet分别为0.1278/0.1558),在更低比特率下优于强端到端基线方法。
cs.LG / 172 / 2609.07874

InfluenceField: A Differentiable Field with Interventionally Identifiable Causal Structure for Multimodal World Modeling

InfluenceField:一种具有干预可识别因果结构的可微场,用于多模态世界建模
Yang, Zihao, Wang, Zijia, Huang, Zhiqiu
Abstract
Multimodal large language models often capture visual-linguistic correlations but struggle to predict how local visual interventions propagate and affect downstream answers. We introduce InfluenceField, an intervention-aware latent field inserted between the visual encoder and language decoder. It lifts patch features into a continuous spatial representation, propagates directed influence over multiple steps, and predicts local intervention effects through a shared transition operator. Training jointly optimizes language modeling, cross-environment invariance, counterfactual rollout supervision, and structural regularization. For a nonlinear finite-basis population model, we show that target-aligned interventional supervision, together with a one-step separation condition on the transition, restricts admissible representations to within-location reparameterizations, so that the directed dependency graph of the full transition is recovered exactly. A linear specialization gives an exact partial-coverage characterization and a finite-loss stability bound, and the field analysis derives the spatial profile of coefficient interventions together with a shared-channel calibration result. On CausalVQA, InfluenceField improves overall accuracy over its backbone by 13.1 percentage points, with the largest gains on the planning and hypothetical categories. Capacity-matched baselines and structural controls attribute the gains in robustness and factual-counterfactual consistency to the causal objectives rather than to added capacity.
Chinese Translation
多模态大语言模型通常能够捕捉视觉与语言之间的相关性,但难以预测局部视觉干预如何传播并影响下游答案。我们提出InfluenceField,这是一个具有干预感知能力的潜在场,插入于视觉编码器与语言解码器之间。它将图像块特征提升为连续的空间表示,通过多步传播有向影响,并通过共享的转移算子预测局部干预效果。训练过程联合优化语言建模、跨环境不变性、反事实展开监督以及结构正则化。对于一个非线性有限基种群模型,我们证明,目标对齐的干预监督与转移算子上的单步分离条件,将可容许的表示限制在位置内的重参数化范围内,从而能够精确恢复完整转移的有向依赖图。线性特化情形给出了精确的部分覆盖刻画和有限损失稳定性界,且场分析推导了系数干预的空间轮廓以及共享通道的校准结果。在CausalVQA数据集上,InfluenceField较其骨干模型将整体准确率提升了13.1个百分点,其中在规划类和假设类问题上提升最为显著。容量匹配的基线模型和结构控制实验表明,鲁棒性与事实-反事实一致性方面的提升源于因果目标本身,而非额外容量的增加。
cs.LG / 173 / 2609.07897

The Accuracy Paradox: Empirical Diagnostic of Default Decision Thresholds in Multi-Label Enzyme Commission Prediction [With Code]

准确率悖论:多标签酶委员会(EC)预测中默认决策阈值的实证诊断[附代码]
Ahmad, Bilal, Mehmood, Rajed
Abstract
Automated prediction of Enzyme Commission (EC) numbers plays a central role in functional annotation and computational drug discovery. However, standard multi-label machine learning pipelines frequently rely on default decision thresholds (t=0.50), assuming balanced prior distributions across target heads. In this study, we present a systematic empirical diagnostic of uncalibrated fixed decision boundaries operating under severe class imbalance across N = 14,096 annotated compounds categorized into six primary EC classes (EC1-EC6). Our results highlight a pronounced Accuracy Paradox: while the multi-label system achieves a deceivingly high mean accuracy of 77.16%, the macro F1-score (0.3976) and macro recall (0.3872) reveal severe predictive breakdown. Majority target classes suffer from hyper-sensitivity and over-prediction, whereas minority classes exhibit sharp recall decay, culminating in a total decision boundary collapse for EC6 (Recall = 0.00%) despite underlying discriminative power (ROC-AUC = 0.5857). Feature correlation analysis further reveals high linear redundancy among topological indices relative to fingerprint density metrics. Ultimately, this diagnostic study demonstrates that standard point predictions mask critical errors in bioinformatics workflows. We establish target-specific threshold optimization and post-hoc conformal calibration as essential, open-source post-processing safeguards for reliable applied machine learning and deep learning architectures.
Chinese Translation
酶委员会编号(Enzyme Commission, EC)的自动预测在功能注释和计算药物发现中发挥着核心作用。然而,标准的多标签机器学习流程通常依赖默认决策阈值(t=0.50),并假设各目标头上具有均衡的先验分布。在本研究中,我们对在严重类别不平衡条件下运行的未校准固定决策边界进行了系统性实证诊断,数据涵盖14,096个已标注化合物(N = 14,096),归属于六大EC主类(EC1-EC6)。我们的结果凸显了一个显著的准确率悖论:尽管该多标签系统取得了看似很高的平均准确率(77.16%),但宏平均F1分数(0.3976)和宏平均召回率(0.3872)揭示出严重的预测失效。多数类目标表现出超敏感性和过度预测,而少数类则出现召回率急剧衰减,最终导致EC6的决策边界完全崩溃(召回率 = 0.00%),尽管其本身仍具有一定的判别能力(ROC-AUC = 0.5857)。特征相关性分析进一步表明,拓扑指数相对于指纹密度指标存在高度的线性冗余。最终,本诊断研究表明,标准的点预测会掩盖生物信息学工作流程中的关键错误。我们确立了面向特定目标的阈值优化和事后保形校准(post-hoc conformal calibration)作为可靠应用机器学习与深度学习架构不可或缺的开源后处理保障措施。
cs.LG / 174 / 2609.07917

AVCG: A Generalized Variational Framework for Counterfactual Generation under Hypothesis Distributions

AVCG:假设分布下反事实生成的广义变分框架
Duell, Jamie, Rodriguez, Alejandro Jimenez, Albarracin, Mahault
Abstract
Counterfactual explanations formalize "what-if" scenarios by identifying modifications to an input instance that obtain a desired alternative prediction. Traditionally, whether generated via instance-specific optimization or amortized single pass models, these approaches rely on a single, deterministic point-estimate predictor. However, this ignores predictive uncertainty and hypothesis variability, leading to brittle explanations that frequently become invalid if the underlying model is retrained or updated. To address this fragility, we propose the Amortized Variational Counterfactual Generator (AVCG), a generalized optimization framework that formulates counterfactual generation as optimization over an arbitrary distribution of plausible predictive hypotheses rather than a single deterministic predictor. This formulation naturally accommodates Bayesian posteriors, Rashomon-restricted hypothesis spaces, and other uncertainty representations within a unified optimization framework. Evaluation across multiple benchmark datasets demonstrates that the AVCG framework produces counterfactual explanations that remain highly valid under predictive uncertainty and model changes, while maintaining competitive plausibility and single-pass runtime performance.
Chinese Translation
反事实解释通过识别对输入实例的修改以获得期望的替代预测,从而将“假设性(what-if)”场景形式化。传统上,无论是通过针对具体实例的优化方法还是摊销式单次前向模型生成,这些方法都依赖于单一的确定性点估计预测器。然而,这忽略了预测不确定性和假设的可变性,导致解释的脆弱性——一旦底层模型被重新训练或更新,这些解释往往会失效。为了解决这种脆弱性,我们提出了摊销式变分反事实生成器(Amortized Variational Counterfactual Generator,AVCG),这是一个广义的优化框架,它将反事实生成表述为在一个由合理预测假设构成的任意分布上的优化,而非基于单一确定性预测器。这一表述自然地容纳了贝叶斯后验、Rashomon 受限假设空间以及其他不确定性表示,并将它们统一在同一个优化框架内。在多个基准数据集上的评估表明,AVCG 框架所生成的反事实解释在预测不确定性和模型变化下仍保持高度有效性,同时在合理性和单次前向运行时间方面保持了有竞争力的性能。
cs.LG / 175 / 2609.07922

Prevalence calibration as shortcut mitigation

以患病率校准缓解捷径学习
Kina, Mohamed Amine, Petersen, Eike
Abstract
Shortcut learning denotes the widespread situation in which a classifier exploits spurious correlations rather than diagnostic features. Existing mitigation strategies mostly aim to learn shortcut-invariant representations; their empirical success is limited and they cannot be applied to classifiers using frozen foundation model encoders. We propose to reframe shortcut learning as fundamentally a calibration problem: unconstrained learning implicitly calibrates each shortcut group to its training set disease prevalence, rendering the resulting classifier necessarily over-confident in one group and under-confident in the other. Building on this insight, we prevalence-equalize calibration between shortcut groups through two encoder-agnostic methods, an in-processing regularizer and a post-hoc prevalence-equalized recalibration step. Across chest-drain-pneumothorax benchmarks on CheXpert and SIIM-ACR, spanning fine-tuned CNNs and frozen foundation-model backbones, both methods substantially outperform all baselines. Post-hoc recalibration of a standard ERM-trained DenseNet raises misaligned-group AUROC from 0.23 to 0.73, indicating that shortcut reliance degrades the classification head rather than the underlying representation. Besides two new state-of-the-art shortcut mitigation approaches, our findings more fundamentally connect shortcut learning to calibration theory and algorithmic fairness.
Chinese Translation
捷径学习(shortcut learning)指分类器利用虚假相关而非诊断性特征的普遍现象。现有的缓解策略大多旨在学习对捷径不变的特征表示;其经验效果有限,且无法应用于使用冻结基础模型编码器的分类器。我们提出将捷径学习重新框定为一个本质上的校准问题:无约束的学习会隐式地将每个捷径组校准到其训练集中的疾病患病率,使得所得到的分类器必然对一个组过度自信而对另一个组自信不足。基于这一洞察,我们通过两种与编码器无关的方法来实现捷径组之间的患病率均衡校准:一种是训练过程中的正则化项,另一种是事后(post-hoc)的患病率均衡重校准步骤。在 CheXpert 和 SIIM-ACR 的气胸-胸腔引流基准上,涵盖微调的 CNN 和冻结的基础模型骨干网络,两种方法均显著优于所有基线方法。对标准 ERM 训练的 DenseNet 进行事后重校准,可将失配组的 AUROC 从 0.23 提升至 0.73,这表明对捷径的依赖损害的是分类头而非底层表示。除了提出两种新的最先进的捷径缓解方法之外,我们的发现更根本地将捷径学习与校准理论及算法公平性联系起来。
cs.LG / 176 / 2609.07952

Structured Extrema Errors in Classical Surrogates for Viscous Burgers: A Physics-Consistent Interpretation

黏性Burgers方程经典代理模型中的结构化极值误差:一种物理一致的解读
Oubari, Youssef
Abstract
We study the local errors of classical machine-learning surrogate models, which approximate the time evolution of the one-dimensional viscous Burgers equation. Four models are compared on the same prediction task, using the spatial grid values directly: radial basis function (RBF) kernel ridge regression (KRR), linear Ridge, ExtraTrees, and Random Forests. Across all four models, the one-step residual, defined here as the true value minus the predicted value at each grid point, forms clear curved branches near predicted maxima and minima. A more detailed analysis of KRR shows that these errors are much more strongly related to the second spatial derivative, which measures local curvature, than to the first spatial derivative. Near a smooth extremum, predicted value and curvature form a local two-branch fold. Under our local curvature-based model of the residual, this fold predicts a leading-order near-parabolic relation between predicted value and residual. This geometric result motivates a direct test of the Burgers advection (transport) and diffusion (smoothing) terms. For KRR and Ridge, regression tests on held-out trajectories, a control that breaks the spatial alignment of the diffusion term, and a spectral test of high-frequency content are consistent with insufficient viscous smoothing at moderate and high viscosity. In this case, the surrogate retains more small-scale structure than the true future state. The same physical explanation is much weaker for the tree models. Finally, a correction that uses only predicted quantities reduces both one-step error and error during recursive rollout, where each prediction is used as the next input.
Chinese Translation
我们研究了经典机器学习代理模型在逼近一维黏性Burgers方程时间演化时的局部误差。我们在同一预测任务上直接使用空间网格数值,比较了四种模型:径向基函数(RBF)核岭回归(KRR)、线性岭回归(Ridge)、极端随机树和随机森林。在所有四种模型中,单步残差(本文定义为每个网格点上真实值减去预测值)在预测的极大值和极小值附近形成清晰的弯曲分支。对KRR的更细致分析表明,这些误差与二阶空间导数(度量局部曲率)的相关性远强于与一阶空间导数的相关性。在光滑极值点附近,预测值与曲率形成一个局部的双分支折叠结构。在我们基于局部曲率的残差模型下,该折叠结构预言了预测值与残差之间呈近抛物线的首阶关系。这一几何结果促使我们对Burgers方程的对流(输运)项和扩散(平滑)项进行直接检验。对于KRR和Ridge,在留出轨迹上的回归测试、一项破坏扩散项空间对齐的对照实验,以及对高频成分的谱检验均表明,在中等和高黏性下模型存在黏性平滑不足的问题。在这种情况下,代理模型比真实未来状态保留了更多的小尺度结构。对于树模型,同样的物理解释则明显较弱。最后,一种仅使用预测量的校正方法同时降低了单步误差和递归滚动预测(即每次预测被用作下一步输入)过程中的误差。
cs.LG / 177 / 2609.07956

Streaming Hierarchical Inference with Tabular Foundation Models

基于表格基础模型的流式层次化推理
Crista, Vitor, Lourenço, Afonso, Martinho, Diogo, Marreiros, Goreti
Abstract
Tabular Foundation Models (TFMs) have recently demonstrated strong predictive performance through in-context learning, but their deployment in high-throughput data streams remains challenging due to communication overhead and latency. We propose \textit{HINT}, a hierarchical inference framework that combines edge-based retrieval with cloud-based TFM inference. A graph-based approximate nearest neighbor memory maintained over a sliding window provides local predictions and uncertainty estimates, allowing confident samples to be processed locally while uncertain instances are selectively offloaded, together with their retrieved context, to a cloud-hosted TFM. The framework exposes an offloading threshold and a neighborhood retrieval policy that can be varied to balance predictive performance and communication cost. Experiments show \textit{HINT} consistently identifies favorable trade-offs.
Chinese Translation
表格基础模型(Tabular Foundation Models, TFMs)近期通过上下文学习展现了强大的预测性能,但受限于通信开销和延迟,其在高吞吐量数据流中的部署仍然具有挑战性。我们提出了 HINT,一种将基于边缘的检索与基于云端的 TFM 推理相结合的层次化推理框架。该框架在滑动窗口上维护一个基于图的近似最近邻记忆,用于提供局部预测和不确定性估计,使置信度高的样本能够在本地处理,而不确定的实例则与其检索到的上下文一起被选择性地卸载至云端托管的 TFM。该框架提供了一个可调节的卸载阈值和邻域检索策略,以平衡预测性能与通信成本。实验表明,HINT 能够持续识别出有利的权衡方案。
cs.LG / 178 / 2609.07961

$\alpha$-Graph: Attention-Infused Normalizing Flow Approach to Tractable Graph Modeling

α-Graph:一种用于可处理图建模的注意力增强归一化流方法
Truong, Thanh-Dat, Alharbi, Sarah, Gauch, Susan, Zhao, Xinghui, Savvides, Marios, Luu, Khoa
Abstract
Graph modeling, a crucial task for representing complex relationships in graph-structured data, has achieved significant success in recent years. However, current graph modeling methods rely on traditional Graph Neural Networks and pre-training approaches to implicitly learn the underlying relational structure of graph data. Thus, these prior methods cannot capture the complex graph structure and correlations among inputs. In this paper, we introduce a novel Attention-based Normalizing Flow-based Approach\footnote{Our implementation and models will be released publicly for research reproducibility.} (ANFA or $\alpha$) that provides an explicit, interpretable, and tractable Graph Modeling ($\alpha$-Graph). In particular, we propose a new Unconditional Graph Normalizing Flow with an Invertible Attention Mechanism to capture the complex relational structure of graph data. To further enhance the expressiveness of the model, we introduce Conditional Graph Normalizing Flow with Learnable Queries that enables efficient modeling of correlations in graph-structured data. We show that our Conditional Graph Normalizing Flows behave similarly to Unconditional Graph Normalizing Flows, enhancing expressiveness while maintaining training stability and efficiency. Our experimental results on three benchmarks will illustrate the effectiveness and the state-of-the-art (SoTA) performance of the proposed $\alpha$-Graph method.
Chinese Translation
图建模是表示图结构数据中复杂关系的一项关键任务,近年来取得了显著成功。然而,当前的图建模方法依赖传统的图神经网络(GNN)和预训练方法来隐式地学习图数据的潜在关系结构。因此,这些先前的方法无法捕捉复杂的图结构以及输入之间的相关性。本文提出了一种新颖的基于注意力的归一化流方法(ANFA 或 α),可提供显式、可解释且易处理的图建模(α-Graph)。具体而言,我们提出了一种新的带可逆注意力机制的无条件图归一化流,以捕捉图数据的复杂关系结构。为进一步增强模型的表达能力,我们引入了带可学习查询(Learnable Queries)的条件图归一化流,能够高效地对图结构数据中的相关性进行建模。我们证明了我们的条件图归一化流的表现与无条件图归一化流相似,在增强表达能力的同时保持了训练的稳定性与效率。我们在三个基准数据集上的实验结果表明,所提出的 α-Graph 方法具有有效性,并达到了最先进的(SoTA)性能。
cs.LG / 179 / 2609.07966

MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

MetaKV:面向受限LLM推理的自适应KV缓存压缩
Wang, Michael, Li, Keith, Bostandoost, Roozbeh
Abstract
Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (LLM) inference, particularly for long-context workloads. However, existing compression methods make different trade-offs among accuracy, inference latency, and peak KV cache memory utilization, making a single fixed configuration unsuitable across different prompts and resource constraints. We introduce MetaKV, an adaptive framework that selects a KV cache compression configuration for each input prompt based on user-specified latency and peak memory budgets. MetaKV uses lightweight prediction models to estimate the end-to-end latency, peak memory, and probability of a correct response for each candidate configuration, and selects the configuration that best satisfies the latency-memory constraints while preserving accuracy. We evaluate MetaKV across ten configurations from three representative KV cache compression methods, KVQuant, H$_2$O, and RocketKV, together with an uncompressed FP16 configuration, on four datasets covering mathematics, science, commonsense reasoning, and reading comprehension. Across a wide range of latency and peak memory constraints, MetaKV consistently outperforms the best static configuration, improving constrained success rate (CSR), the fraction of prompts answered correctly while satisfying both constraints, by approximately 0.07 on average and up to 0.135. These results demonstrate the benefit of adapting KV cache compression to individual prompts and latency-memory constraints. Code is available at https://github.com/MichaelWang0505/MetaKV.git
Chinese Translation
键值(KV)缓存压缩是降低大语言模型(LLM)推理内存开销的有效方法,尤其在长上下文工作负载中。然而,现有的压缩方法在准确率、推理延迟和峰值KV缓存内存占用之间进行不同的权衡,使得单一的固定配置无法适用于不同的提示和资源约束。我们提出了MetaKV,一个自适应框架,可根据用户指定的延迟和峰值内存预算为每个输入提示选择KV缓存压缩配置。MetaKV使用轻量级预测模型来估计每个候选配置的端到端延迟、峰值内存以及正确响应的概率,并选择在保持准确率的同时最能满足延迟-内存约束的配置。我们在三个代表性KV缓存压缩方法KVQuant、H$_2$O和RocketKV的十个配置以及一个未压缩的FP16配置上,在涵盖数学、科学、常识推理和阅读理解的四个数据集上评估了MetaKV。在广泛的延迟和峰值内存约束范围内,MetaKV始终优于最佳的静态配置,将受约束成功率(CSR,即同时满足两个约束且回答正确的提示比例)平均提升约0.07,最高提升0.135。这些结果表明,针对单个提示以及延迟-内存约束自适应地调整KV缓存压缩是有益的。代码可在https://github.com/MichaelWang0505/MetaKV.git获取。
cs.LG / 180 / 2609.07975

Heat Field Signatures: From Point Clouds to Smooth Geometry

热场特征:从点云到光滑几何
Wang, Yuanqing, Tian, Yapeng, Coskunuzer, Baris
Abstract
Bringing multiscale geometric analysis directly to irregular point clouds remains difficult: quantities such as local dimension, anisotropy, density variation, and geometric transitions are typically estimated through explicit neighborhood, manifold, or graph constructions, or left for neural networks to infer from coordinates. We introduce Heat Field Signatures (HFS), which lift a point cloud to a multiscale family of smooth ambient heat fields, providing a direct interface from discrete samples to geometric analysis. From this field, HFS computes closed-form global and local signatures directly from pairwise distances, capturing heat concentration, intrinsic dimension, anisotropy, and scale transitions. We further introduce the Heat Dimension Spectrum (HDS), a compact summary of multiscale geometric composition. HFS can be used as a closed-form descriptor, a lightweight learned representation, or a geometric feature channel for neural point-cloud models. Across synthetic and real-world benchmarks spanning subcellular, neuronal, tree, and protein data, HFS outperforms strong point-cloud and multiparameter-persistence baselines while substantially reducing end-to-end cost. On SCOP protein-fold classification, HFS improves over the strongest deep baseline by nearly $24$ percentage points using coordinates alone, while standalone HFS representations are exactly rotation-invariant by construction. More broadly, HFS turns a classical heat field into a practical interface for multiscale geometric analysis in modern point-cloud learning.
Chinese Translation
将多尺度几何分析直接应用于不规则点云仍然是一个难题:诸如局部维度、各向异性、密度变化和几何过渡等量通常需要通过显式的邻域、流形或图结构来估计,或者留给神经网络从坐标中自行推断。我们提出了热场特征(Heat Field Signatures, HFS),它将点云提升为一个多尺度的光滑环境热场族,为离散样本到几何分析提供了直接接口。基于该热场,HFS 直接从成对距离出发计算闭式(closed-form)的全局与局部特征,捕捉热浓度、内在维度、各向异性和尺度过渡。我们进一步提出了热维度谱(Heat Dimension Spectrum, HDS),一种紧凑的多尺度几何组成概要。HFS 可以作为闭式描述子、轻量级学习表示,或作为神经点云模型的几何特征通道。在涵盖亚细胞、神经元、树木和蛋白质数据的合成与真实世界基准测试中,HFS 优于强大的点云和多参数持续同调(multiparameter-persistence)基线方法,同时大幅降低了端到端成本。在 SCOP 蛋白质折叠分类任务上,仅使用坐标,HFS 比最强的深度学习基线提升了近 $24$ 个百分点,且独立的 HFS 表示在构造上具有严格的旋转不变性。更广泛地说,HFS 将经典热场转变为现代点云学习中多尺度几何分析的实用接口。
cs.LG / 181 / 2609.07982

Semi-Supervised Learning under Spatially Biased Sampling

空间偏倚采样下的半监督学习
Nuakoh, Bright Wiredu, Fouedjio, Francky, Bradshaw, Stephen, Awuah-Mensah, Yaw Kwaafo, Tan, Wei Hong, Arya, Emet, Afrifa-Yamoah, Ebenezer
Abstract
Standard semi-supervised learning (SSL) typically relies on labelled and unlabelled data sharing a common marginal distribution. This assumption is often violated by biased spatial sampling mechanism, when labels are collected under spatially biased or preferential site selection. We treat this marginal mismatch, spatial autocorrelation, and spatial non-stationarity as three distinct mechanisms, varied independently via a labelled-sampling concentration parameter, a spatial length scale, and a non-stationarity strength parameter, and ask how mismatch degrades SSL, whether the cluster and manifold assumptions survive it, and how the resulting failure can be diagnosed. Using a controlled synthetic framework alongside PovertyMap-WILDS, California housing, socio-economic and US air quality monitoring datasets, we systematically vary the degree of mismatch while accounting for spatial autocorrelation and non-stationarity. Through a series of analyses including a segmented-regression changepoint, we show that in the synthetic generator, SSL performance does not degrade gradually but instead exhibits a threshold-like breakdown between approximately 0.71 and 0.77 once distribution mismatch becomes sufficiently severe. We further demonstrate that spatial non-stationarity contributes to performance loss independently of marginal mismatch and that models become increasingly overconfident outside the regions where labels are available. To support practical deployment, we evaluate several distribution-divergence measures as indicators of reliability and introduce a kernel-weighted local divergence metric that provides a more stable estimate of spatial mismatch than a na\"ive localised approach. These findings provide empirical evidence and diagnostic tools for better documenting the risk of incorporating unlabelled spatial data into semi-supervised learning workflows.
Chinese Translation
标准的半监督学习(Semi-Supervised Learning, SSL)通常依赖于有标签数据与无标签数据共享相同的边缘分布。当标签是在空间偏倚或优先选址的条件下采集时,这一假设往往会被有偏的空间采样机制所违背。我们将这种边缘分布失配、空间自相关和空间非平稳性视为三种相互独立的机制,分别通过一个有标签采样集中度参数、一个空间长度尺度和一个非平稳性强度参数进行独立变化,并探讨以下问题:分布失配如何降低半监督学习的性能、聚类假设与流形假设在失配下是否仍然成立,以及如何诊断由此产生的失效。基于一个受控的合成框架以及PovertyMap-WILDS、加州房价、社会经济数据和美国空气质量监测数据集,我们在考虑空间自相关与非平稳性的前提下,系统地改变分布失配的程度。通过一系列分析(包括分段回归变点分析),我们发现在合成生成器中,半监督学习的性能并非逐渐退化,而是在分布失配足够严重时,在约0.71至0.77之间表现出阈值式的崩溃。我们进一步证明,空间非平稳性会独立于边缘分布失配导致性能损失,并且模型在无标签可用的区域之外会变得越来越过度自信。为支持实际部署,我们评估了若干分布散度度量作为可靠性指标,并引入了一种核加权局部散度指标,其相比朴素的局部化方法能够对空间失配提供更稳定的估计。这些发现为在半监督学习流程中更好地记录纳入无标签空间数据的风险提供了实证证据和诊断工具。
cs.LG / 182 / 2609.07983

Solving the Elastic Wave Equation with Physics-Informed Neural Networks: A Robust and Critical Assessment

利用物理信息神经网络求解弹性波动方程:一项稳健且批判性的评估
Staub, Davide, Moseley, Ben
Abstract
Physics-Informed Neural Networks (PINNs) have recently emerged as a promising approach for solving Partial Differential Equations (PDEs), offering a meshfree alternative that integrates physical principles into the learning process. This presents a new paradigm compared to traditional discretization methods and purely data-driven machine learning techniques. While promising, PINNs are not a panacea; they inherit challenges such as spectral bias and unstable convergence. Moreover, their potential in seismology remains largely unexplored. In this work, we provide a robust and critical assessment of PINNs for solving the elastic wave equation in seismology. We investigate the performance of PINNs on problems with varying degrees of complexity across various seismic sources and parameter models, from constant to highly heterogeneous settings. A pivotal aspect of our work involves investigating whether embedding physical principles directly into the network architecture enhances convergence and accuracy. We test an extensive range of neural architecture designs, from unrestricted, uninformed PINNs to highly specialized ones. We find that integrating an understanding of wave physics into the network design significantly improves accuracy. For instance, introducing a custom wavelet or plane wave layer, coupled with encoder and decoder layers, consistently yields a relative $L_2$ error approximately half that of the standard PINN, as evidenced across numerous experiments. We further demonstrate that this novel architecture enhances accuracy when applied to the acoustic wave equation, underlying the versatility of our network. Another key contribution of our research is the successful conditioning of PINNs on seismic source locations. This signifies a considerable advancement towards rapid seismic hazard detection and seismic analysis.
Chinese Translation
物理信息神经网络(Physics-Informed Neural Networks, PINNs)近来成为求解偏微分方程(PDEs)的一种有前景的方法,它提供了一种将物理原理融入学习过程的无网格替代方案。与传统的离散化方法和纯数据驱动的机器学习技术相比,这呈现了一种新的范式。尽管前景可观,但PINNs并非万能之策;它们继承了谱偏差和不稳定收敛等挑战。此外,其在地震学中的潜力在很大程度上仍未被探索。在这项工作中,我们对PINNs求解地震学中弹性波动方程进行了稳健且批判性的评估。我们研究了PINNs在不同复杂程度问题上的表现,涵盖多种震源和参数模型,从常参数到高度非均匀介质 setting。我们工作的一个关键方面是探究将物理原理直接嵌入网络架构是否能提升收敛性和精度。我们测试了从无约束、无先验信息的PINNs到高度专业化网络的大量神经架构设计。我们发现,将波动物理的理解融入网络设计能显著提高精度。例如,引入自定义小波层或平面波层,并结合编码器和解码器层,在大量实验中一致地使相对$L_2$误差约为标准PINN的一半。我们进一步证明,该新型架构在应用于声波方程时同样能提升精度,凸显了我们网络的通用性。本研究的另一项关键贡献是成功实现了PINNs对震源位置的条件化,这标志着向快速地震危险性检测和地震分析迈出了重要一步。
cs.LG / 183 / 2609.07986

Automated Chest CT Protocol Selection via Large Language Model Derived Text Embeddings from Imaging Request Text

基于大语言模型从影像检查申请文本中提取文本嵌入的胸部CT扫描方案自动选择
Hosseini, Zahra, Pouromidi, Mahan, Khalvati, Farzad, Rogalla, Patrik
Abstract
Purpose: Accurate CT protocol selection is critical for diagnostic quality and patient safety, yet the current process is manual, time-consuming, and prone to inconsistencies. Prior Machine Learning methods using keywords or bag-of-words lack contextual understanding and perform poorly on rare protocols. We propose a decision support system using large language model (LLM) features to recommend protocols from free-text clinical indications, capturing clinical nuance and phrasing variation for more consistent, efficient selection. Methods: In this REB-approved retrospective study, 285,123 chest CT imaging requests from a large academic medical center (2017-2024) were split into training (228,099, 80%) and held-out test (57,024, 20%) sets. Each request included procedure names, clinical indication, HIS comments, and the selected protocol. Clinical text was embedded using a fine-tuned LLM, Meta's LLaMA-3.1-70B; these features input a logistic regression classifier predicting 18 protocol labels (e.g., PE, LDCT). Results: The pipeline achieved a weighted precision of 0.84, weighted F1-score of 0.81, and overall accuracy of 79% across 18 CT protocols. On 300 independent cases with expert consensus, the LLM reached an overall accuracy of 80% versus 83% for radiologists, with no significant difference (p = 0.263). Performance was comparable across most classes, with the LLM exceeding radiologists for some challenging categories, and entropy analyses indicated more balanced protocol use, suggesting reduced variability. Conclusion: An LLM-based recommendation system can leverage general knowledge from a large natural-text corpus to accurately assign chest CT protocols from free-text imaging requests, and may serve as a viable foundation for protocol recommendation tools where inputs require language understanding.
Chinese Translation
目的:准确的CT扫描方案选择对诊断质量和患者安全至关重要,然而当前流程依赖人工、耗时且容易出现不一致。先前采用关键词或词袋模型的机器学习方法缺乏对上下文的理解,在罕见方案上表现较差。我们提出一种利用大语言模型(LLM)特征的决策支持系统,从自由文本的临床适应症中推荐扫描方案,从而捕捉临床细节和措辞变化,实现更一致、更高效的选择。方法:在这项经研究伦理委员会批准的回顾性研究中,来自某大型学术医疗中心(2017-2024年)的285,123份胸部CT影像检查申请被划分为训练集(228,099份,80%)和留出测试集(57,024份,20%)。每份申请包含检查项目名称、临床适应症、HIS备注以及所选方案。临床文本使用经过微调的LLM(Meta的LLaMA-3.1-70B)进行嵌入,这些特征被输入一个逻辑回归分类器,用于预测18种扫描方案标签(如肺栓塞PE、低剂量CT LDCT)。结果:该流程在18种CT方案上取得了加权精确率0.84、加权F1分数0.81和总体准确率79%。在300例经专家共识确认的独立病例上,LLM的总体准确率为80%,放射科医生为83%,两者无显著差异(p = 0.263)。在大多数类别上表现相当,LLM在某些具有挑战性的类别上甚至超过放射科医生;熵分析表明扫描方案使用更为均衡,提示变异性有所降低。结论:基于LLM的推荐系统可以利用从大规模自然语言语料库中获得的通用知识,根据自由文本的影像检查申请准确地分配胸部CT扫描方案,并可作为在输入需要语言理解能力的场景下扫描方案推荐工具的可行基础。
cs.LG / 184 / 2609.07990

HyCO: A Hybrid Neural Solver for Combinatorial Optimization

HyCO:一种用于组合优化的混合神经求解器
Li, Yuheng, Yang, Di, Chen, Haipeng, Xiong, Yanhai
Abstract
Sequential reinforcement learning (RL) solvers and global diffusion model (DM) solvers for neural combinatorial optimization exhibit complementary failure modes under an optimization-regret view. The former enjoys small marginal regret in the early construction stage, but suffers from horizon-wise compounding errors with super-linear regret growth; the latter avoids horizon compounding but incurs linear or sublinear regret w.r.t. the dimension of the remaining unsolved subspace. We propose Hybrid Neural Solver for Combinatorial Optimization (HyCO), a hybrid inference algorithm that constructs a solution prefix with an RL solver and adaptively switches to a conditional DM to complete the remaining decisions. To characterize why such hybridization helps, when to trigger the handover, and how to realize it in practice, we first develop a unified error-scaling theoretical framework and prove that, under explicit error-scaling assumptions, i) the hybrid structure achieves strictly lower expected regret than either backbone alone, and ii) there exists a unique optimal trigger step that minimizes the hybrid regret. We then design a lightweight adaptive trigger that combines policy entropy and RL-DM disagreement to detect trajectory-level signals of the regime shift as a practical proxy, since the optimal trigger step is defined at the expected-regret level and is not directly computable on individual trajectories. Experimental results on diverse benchmarks demonstrate that HyCO achieves consistent improvements over both backbones and support the empirical effectiveness of adaptive triggering.
Chinese Translation
从优化遗憾(optimization-regret)的视角来看,用于神经组合优化的序列式强化学习(RL)求解器与全局扩散模型(DM)求解器呈现出互补的失败模式。前者在早期构造阶段边际遗憾较小,但随着求解步数的推进会出现复合误差,导致遗憾超线性增长;后者虽避免了逐步复合误差,但其遗憾相对剩余未解子空间的维度呈线性或亚线性增长。我们提出了混合神经组合优化求解器(Hybrid Neural Solver for Combinatorial Optimization,HyCO),这是一种混合推理算法:先用RL求解器构造解的前缀,再自适应地切换到条件扩散模型(conditional DM)以完成剩余决策。为了阐明这种混合方法为何有效、何时触发切换以及如何在实践中实现,我们首先建立了一个统一的误差缩放理论框架,并证明在明确的误差缩放假设下:i) 混合结构能够取得严格低于任一单一骨干方法的期望遗憾;ii) 存在唯一的最优触发步骤,可最小化混合遗憾。由于最优触发步骤是在期望遗憾层面定义的,无法在单条轨迹上直接计算,我们进而设计了一种轻量级自适应触发机制,结合策略熵与RL-DM分歧度来检测机制转变的轨迹级信号,作为其实用代理。在多个基准上的实验结果表明,HyCO 相比两种骨干方法均取得了一致的性能提升,并验证了自适应触发的实际有效性。
cs.LG / 185 / 2609.07997

Sharp Structure-Agnostic Minimax Risk for Partial Linear Models

部分线性模型中精确的结构无关极小极大风险
Hu, Haichen, Simchi-Levi, David
Abstract
We characterize the sharp structure-agnostic minimax risk for coefficient estimation in the partial linear model when the outcome and treatment nuisances are learned by two distinct black-box learners, which resolves the open problem in double machine learning posed by Gu (2025). For each nuisance \(q\in\{\mu,\pi\}\), we characterize the available learner by an approximation-error budget \(a_q\) and a stochastic-error budget \(s_q\), with the latter controlled through localized Rademacher complexity. Writing \(\mathcal E_n\) for the minimax mean-squared error, we show that \[\mathcal E_n\asymp1\wedge\left\{\frac1n+\left(a_\mu a_\pi+\min\left\{a_\pi s_\mu+s_\pi^2,\,a_\mu s_\pi+s_\mu^2\right\}\right)^2\right\}.\] The main new ingredient is a novel lower bound for the general two-learner problem. Our proof constructs four finite-mixture testing experiments using orthogonal code functions. Across these experiments, the hidden perturbations are placed outside both learner classes, outside only the treatment learner class, outside only the outcome learner class, or inside both learner classes. These four configurations capture, respectively, the interaction between the two approximation errors, the two asymmetric interactions between one learner's approximation error and the other learner's learning error, and the joint estimation difficulty of learning both nuisances. Combining the four resulting lower bounds yields the displayed rate, which matches the latest upper bound in Gu (2026). Our result shows that standard double machine learning can overstate the intrinsic difficulty of target estimation and provides a target-specific principle for learner selection: approximation error and stochastic complexity must be jointly balanced across the two nuisance learners rather than optimized separately.
Chinese Translation
我们刻画了部分线性模型中系数估计的精确结构无关极小极大风险,其中结果变量与处理变量的干扰项由两个不同的黑箱学习器学习,这解决了 Gu(2025)提出的双重机器学习中的公开问题。对于每个干扰项 \(q\in\{\mu,\pi\}\),我们通过近似误差预算 \(a_q\) 和随机误差预算 \(s_q\) 来刻画可用学习器,后者通过局部化 Rademacher 复杂度加以控制。记 \(\mathcal E_n\) 为极小极大均方误差,我们证明 \[\mathcal E_n\asymp1\wedge\left\{\frac1n+\left(a_\mu a_\pi+\min\left\{a_\pi s_\mu+s_\pi^2,\,a_\mu s_\pi+s_\mu^2\right\}\right)^2\right\}.\] 主要的新工具是针对一般双学习器问题的新颖下界。我们的证明利用正交编码函数构造了四个有限混合检验实验。在这些实验中,隐含扰动分别被置于两个学习器类之外、仅处理学习器类之外、仅结果学习器类之外,或同时置于两个学习器类之内。这四种配置分别刻画了两个近似误差之间的交互作用、一个学习器的近似误差与另一个学习器的学习误差之间的两种非对称交互作用,以及同时学习两个干扰项的联合估计难度。综合由此得到的四个下界即可得出上述收敛速率,该速率与 Gu(2026)最新的上界相匹配。我们的结果表明,标准双重机器学习可能高估目标估计的内在难度,并提供了一种针对特定目标的学习器选择原则:必须在两个干扰学习器之间联合权衡近似误差与随机复杂度,而非分别独立优化。
cs.LG / 186 / 2609.08034

Two-Scale Localized PCA-Net: Coarse-Global and Local-Residual Representations for Artifact-Reduced PDE Operator Learning

双尺度局部化PCA-Net:用于伪影削减PDE算子学习的粗尺度全局与局部残差表示
Dhingra, Mrigank, Stout, Jordan, San, Omer
Abstract
Localized dimensionality reduction improves the scalability of operator learning for high-dimensional partial differential equations (PDEs), but independently decoded local patches can introduce block offsets, interface mismatches, and spurious high-wavenumber content. We introduce Two-Scale Localized PCA-Net, which decomposes the solution into a coarse-global component and local residual corrections. A compact global PCA basis captures domain-scale structure, while nonoverlapping local PCA bases represent the remaining fine-scale residual. A block-balanced latent objective couples the two representations, and optional interface-aware fine-tuning further promotes continuity through reconstruction and trace losses. On Poisson benchmarks, the two-scale representation substantially reduces reconstruction error and visible block artifacts relative to plain and overlap-based localized PCA-Net while approximately halving PCA fitting cost relative to overlap. On heterogeneous Darcy flow, it strongly reduces interface and discrete-residual errors, with more modest reconstruction gains. Ablations show that the primary improvement arises from the two-scale output representation, while interface-aware fine-tuning provides complementary continuity refinement. Overall, separating globally coherent structure from localized residual detail provides an efficient representation for artifact-reduced PDE operator learning.
Chinese Translation
局部化降维提升了高维偏微分方程(PDE)算子学习的可扩展性,但独立解码的局部块可能引入块间偏移、界面失配以及虚假的高波数成分。我们提出双尺度局部化PCA-Net(Two-Scale Localized PCA-Net),将解分解为粗尺度全局分量和局部残差修正。紧凑的全局PCA基用于捕捉域尺度结构,而非重叠的局部PCA基表示其余的细尺度残差。通过块平衡的潜在目标函数将两种表示耦合起来,并可选地通过界面感知微调,利用重构损失和迹损失进一步促进连续性。在泊松(Poisson)基准测试中,相对于朴素的及基于重叠的局部化PCA-Net,双尺度表示显著降低了重构误差和可见的块状伪影,同时相比重叠方法将PCA拟合成本大约减半。在非均质达西(Darcy)流问题上,该方法显著降低了界面误差和离散残差误差,而重构误差的改善相对有限。消融实验表明,主要改进源于双尺度输出表示,而界面感知微调提供了互补的连续性细化。总体而言,将全局相干结构与局部残差细节分离,为伪影削减的PDE算子学习提供了一种高效的表示方法。
cs.LG / 187 / 2609.08064

Risk-Conditioned Fine-Tuning of Large Language Models

大型语言模型的风险条件化微调
Liu, Zixuan, Wu, Fangzheng, Summa, Brian, zheng, Zizhan
Abstract
Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. Existing Risk-Averse RLHF addresses this issue by optimizing Conditional Value-at-Risk (CVaR), but it trains policies for fixed risk levels and therefore cannot adjust the desired degree of risk aversion at inference time. In this paper, we propose risk-conditioned RLHF, a framework that trains a single policy that provides a continuous risk-control interface, enabling users to select different degrees of risk aversion without retraining or deploying multiple risk-specific models. Experiments across multiple benchmarks demonstrate that a single risk-conditioned policy can adapt to different risk levels at inference time, enabling more flexible and risk-aware LLM deployment.
Chinese Translation
大型语言模型(LLM)越来越多地被部署在这样一些场景中:罕见但严重的有害生成可能造成重大后果。现有的风险规避型RLHF通过优化条件风险价值(CVaR)来解决这一问题,但其针对固定的风险水平训练策略,因此无法在推理时调整期望的风险规避程度。本文提出风险条件化RLHF(risk-conditioned RLHF)框架,该框架训练一个单一策略,提供连续的风险控制接口,使用户能够在无需重新训练或部署多个特定风险模型的情况下选择不同程度的风险规避。在多个基准上的实验表明,单一的风险条件化策略可以在推理时适应不同的风险水平,从而实现更灵活、更具风险感知能力的LLM部署。
cs.LG / 188 / 2609.08078

A Machine Learning Framework for Predicting Restaurant Food Waste to Support Sustainable Food Management

一个用于预测餐厅食物浪费的机器学习框架,以支持可持续食物管理
Naeem, Md Mehedi Hasan, Islam, Md Ashraful, Barua, Moumita, Araf, Ishtiyak Ahmmad, Mahir, Md. Arefin Haque
Abstract
Food waste in the restaurant sector poses a substantial challenge to environmental sustainability and economic efficiency. This paper presents an exploratory machine learning framework for estimating daily restaurant food waste quantities from operational and contextual features. A structured dataset was constructed by integrating restaurant demand records, meteorological data and temporal event indicators, yielding 77,980 records across 27 features. Because large-scale ground-truth food waste measurements are not publicly available, the target variable was derived from operationally justified assumptions, with the complete construction formula and controlled stochastic variability disclosed for full reproducibility. Four supervised regression models, namely Linear Regression, Decision Tree, Random Forest and Gradient Boosting, were evaluated under a chronological 70-30 train-test split that respects the temporal ordering of restaurant operations, augmented by 5-fold time-series cross-validation. All reported metrics are explicitly scoped to performance against the constructed target and do not imply validation against measured food waste. Ensemble methods consistently outperformed linear baselines. Random Forest attained an MAE of 6.19 kg, RMSE of 8.36 kg and $R^2$ of 0.817 on the realistic feature subset following systematic exclusion of algebraically leakage-prone variables. Feature importance analysis identified menu diversity, operational area and temporal activity patterns as the primary predictive drivers. The full dataset, target construction formula, codebase and experimental configurations are publicly released to support reproducibility and future extension to empirically measured waste data.
Chinese Translation
餐饮行业的食物浪费对环境可持续性和经济效率构成了重大挑战。本文提出了一个探索性机器学习框架,用于基于运营和情境特征估计餐厅每日食物浪费量。通过整合餐厅需求记录、气象数据和时间事件指标,构建了一个结构化数据集,共包含77,980条记录和27个特征。由于大规模的真实食物浪费测量数据并未公开可得,目标变量基于具有运营依据的假设推导而来,并完整公开了其构建公式和受控随机变异性,以确保完全可复现性。在遵循餐厅运营时间顺序的70-30按时间划分的训练-测试方式下,并对四种监督回归模型(即线性回归、决策树、随机森林和梯度提升)进行了评估,同时辅以5折时间序列交叉验证。所有报告的指标均明确限定为针对所构建目标变量的性能,并不代表对实际测量食物浪费的验证。集成方法始终优于线性基线模型。在系统性排除易产生代数性数据泄露的变量后的真实特征子集上,随机森林取得了6.19 kg的MAE、8.36 kg的RMSE和0.817的R²。特征重要性分析表明,菜单多样性、营业面积和时间活动模式是主要的预测驱动因素。完整数据集、目标构建公式、代码库和实验配置均已公开,以支持可复现性以及未来向经验测量浪费数据的扩展。
cs.LG / 189 / 2609.08080

Proactive Context-Forecasted Safety Constraints for Nonstationary Reinforcement Learning

面向非平稳强化学习的主动式上下文预测安全约束
Tomashevskiy, Tim
Abstract
Ensuring safety in reinforcement learning under nonstationarity requires anticipating changes in risk before they lead to unsafe behavior. Existing approaches typically rely on safety constraints defined at design time or updated reactively during execution, assuming that such constraints remain valid over time. However, in nonstationary environments with evolving contexts and changing driving layouts, these assumptions may fail. We propose a framework for proactive safety constraint generation based on context forecasting. The approach infers latent environmental context from observations, predicts its future evolution, and constructs safety constraints adapted to anticipated conditions. This enables the agent to proactively avoid unsafe regions instead of reacting only after safety violations occur. We evaluate the method in driving environments with structured context variation. The experiments include a sweep over nonstationarity intensities and additional held-out driving layouts, including highway, intersection, and racetrack scenarios. Results show that proactive constraint generation substantially reduces collisions under both seen and out-of-training nonstationarity intensities and generally remains effective across held-out driving layouts while maintaining usable task performance. These findings suggest that context-based constraint generation is a promising approach for safe reinforcement learning under nonstationarity.
Chinese Translation
在非平稳环境下确保强化学习的安全性,需要在风险导致不安全行为之前对其进行预判。现有方法通常依赖于在设计时定义的安全约束,或在执行过程中被动更新的约束,并假设这些约束随时间保持有效。然而,在具有不断演化的上下文和变化驾驶布局的非平稳环境中,这些假设可能失效。我们提出了一种基于上下文预测的主动式安全约束生成框架。该方法从观测中推断潜在的环境上下文,预测其未来演化,并构建适应预期条件的安全约束。这使智能体能够主动避开不安全区域,而不是在安全违规发生后才作出反应。我们在具有结构化上下文变化的驾驶环境中评估了该方法。实验包括对非平稳性强度的扫描测试以及额外的保留驾驶布局测试,涵盖高速公路、交叉路口和赛道场景。结果表明,主动式约束生成在训练中已见及未见过的非平稳性强度下均能显著减少碰撞,并且在保留驾驶布局上总体保持有效,同时维持了可用的任务性能。这些发现表明,基于上下文的约束生成是非平稳环境下安全强化学习的一种有前景的方法。
cs.LG / 190 / 2609.08102

Learning Metamaterial Eigenmodes with Wavelet-Encoded Fourier Neural Operators

基于小波编码傅里叶神经算子的超材料本征模态学习
Zhang, Han, Ogren, Alexander, Rudin, Cynthia, Guilleminot, Johann, Brinson, L. Catherine
Abstract
Machine learning surrogates based on neural operators have shown broad applicability in solving forward PDE problems. However, eigenvalue problems, in which an eigenparameter and one of several valid eigenmodes must be simultaneously solved, remain difficult because standard operator learning formulations assume a unique input-output map. This work demonstrates that Fourier Neural Operators (FNOs), combined with wavelet-based encodings of PDE inputs, can learn and predict multiple eigenmodes of the elastic wave equation, corresponding to deformation modes of acoustic waves propagating through arbitrary metamaterial geometries. We provide a mechanistic explanation and experimental evidence for why wavelet encodings are well matched to the dual spatial-spectral structure of the FNO, enabling deterministic mode selection on both continuous-valued and binary-valued geometries within a single model, and for why prediction accuracy varies with geometric discontinuities. For metamaterial design, the resulting surrogate accelerates the simulation stage of the design cycle by three orders of magnitude relative to finite element analysis on a consumer-grade CPU, while preserving high fidelity. These results also carry broader implications for designing input encodings in other multi-mode PDE solvers based on spectral neural operators.
Chinese Translation
基于神经算子的机器学习代理模型在求解偏微分方程(PDE)正问题方面已展现出广泛的适用性。然而,本征值问题——即需要同时求解一个本征参数和若干有效本征模态之一——仍然困难,因为标准的算子学习框架假设输入到输出的映射是唯一的。本工作表明,傅里叶神经算子(Fourier Neural Operators, FNO)与基于小波的PDE输入编码相结合,能够学习并预测弹性波方程的多个本征模态,这些模态对应于声波在任意超材料几何结构中传播时的形变模式。我们从机理上解释并通过实验证明:小波编码之所以与FNO的双重空间-频谱结构高度契合,是因为它使单一模型能够在连续值和二值几何结构上实现确定性的模态选择;同时解释了预测精度为何随几何不连续性而变化。对于超材料设计而言,所得代理模型在消费级CPU上相比有限元分析将设计周期中仿真阶段的耗时加速了三个数量级,同时保持了高保真度。这些结果对于其他基于谱神经算子的多模态PDE求解器中输入编码的设计也具有更广泛的意义。
cs.LG / 191 / 2609.08106

Nystr\"om Attention Matches Full Attention for Cross-Sectional Stock Prediction

Nyström注意力在横截面股票预测中媲美全注意力机制
Guo, Kunhan
Abstract
MASTER's inter-stock multi-head attention -- the module responsible for modeling cross-sectional stock relationships -- accounts for 42.5% of model parameters and 25% of predictive value. We systematically decompose this module and uncover a surprising structure: the learned attention is near-uniform (perplexity 278/300), yet forcing exact uniformity eliminates all cross-sectional discrimination. Spectral analysis resolves this paradox: the deviation from uniformity is low-rank (effective rank ~65, top-10 modes capture 96.5% of energy), explaining why sparse approximations consistently fail while Nystrom low-rank attention (m=32 landmarks) matches full O(N^2) attention at O(mN) cost -- certified equivalent via TOST at both N=300 (5 seeds, Rank IC p=0.003) and N=800 (10 seeds, Rank IC p=0.034). Additional findings include: (i) attention anti-correlates with return similarity (Spearman rho = -0.614; on the industry-labeled subset, -0.645 unconditionally and -0.627 after controlling for industry, beta, and volatility), suggesting complementarity-seeking rather than correlation mining; (ii) all graph-based alternatives degrade performance, with hard masking worse than complete module removal; and (iii) at N ~ 3,500 with adapted architectures, no cross-stock module (GCN, Nystrom, or MASTER-style pipeline) significantly outperforms a per-stock LSTM baseline (n=4 seeds), indicating that the benefits observed at smaller scales do not trivially transfer. These results establish that the inter-stock attention's value resides in a compressible, dynamic, near-global redistribution that rewards low-rank approximation but resists sparsification.
Chinese Translation
MASTER模型中的股票间多头注意力模块——负责建模横截面股票关系的核心组件——占据了模型参数的42.5%和预测价值的25%。我们对该模块进行了系统性分解,并发现了一个令人惊讶的结构:学到的注意力接近均匀分布(困惑度为278/300),然而强制其完全均匀化却会消除所有横截面区分能力。谱分析解释了这一悖论:对均匀性的偏离是低秩的(有效秩约为65,前10个模态捕获了96.5%的能量),这解释了为何稀疏近似始终失败,而Nyström低秩注意力(m=32个地标点)能以O(mN)的代价匹配全量O(N^2)注意力——并通过TOST检验在N=300(5个随机种子,Rank IC p=0.003)和N=800(10个随机种子,Rank IC p=0.034)下均被认证为统计等价。其他发现包括:(i) 注意力与收益率相似性呈负相关(Spearman ρ = -0.614;在行业标注子集上,无条件时为-0.645,在控制行业、贝塔和波动率后为-0.627),这表明该机制倾向于寻求互补性而非挖掘相关性;(ii) 所有基于图结构的替代方案均降低性能,其中硬掩码的效果比完全移除该模块更差;(iii) 在N约3,500且采用适配架构的条件下,任何跨股票模块(GCN、Nyström或MASTER式流水线)均未显著优于逐股票LSTM基线(n=4个种子),这表明在小规模下观察到的收益并不能轻易迁移。这些结果表明,股票间注意力的价值存在于一种可压缩的、动态的、近乎全局的重分布机制中——它青睐低秩近似,但抵抗稀疏化。
cs.LG / 192 / 2609.08133

Sparse Data Augmentation for Optimization with Provable Guarantees

具有可证明保证的稀疏数据增强优化方法
Tahmasebi, Behrooz, Weber, Melanie
Abstract
In nonconvex optimization problems arising in geometric machine learning, data augmentation is commonly used to promote invariance by averaging empirical losses over transformations of the data. Computing the fully augmented objective, however, requires access to every element of the transformation group $G$, which may be prohibitively expensive when $G$ is large or accessible only through sampling. We study whether full augmentation can instead be approximated using a small, fixed sample of transformations acquired before optimization and reused thereafter. Under suitable regularity conditions, we show that, with probability at least $1-\delta$, gradient descent (GD) on the resulting sparsely augmented objective returns an $\varepsilon$-stationary point of the fully augmented objective using $\mathcal{O}\bigl((\log |G|+\log(1/\delta))/\varepsilon^2\bigr)$ group-transformation-oracle queries. By comparison, standard group stochastic gradient descent (group-SGD), which samples a fresh transformation at every iteration, uses $\mathcal{O}(1/\varepsilon^4)$ transformation queries. Therefore, gradient descent with fixed sparse augmentation requires fewer transformation queries than both GD applied to the fully augmented objective and group-SGD. Our proof techniques, which may be of independent interest, establish a uniform approximation of the full group-averaged gradient field by a random group average using spectral properties of group-induced operators and tools from representation theory.
Chinese Translation
在几何机器学习中的非凸优化问题里,数据增强常被用于通过对数据的变换求经验损失的平均来促进不变性。然而,计算完全增强后的目标函数需要访问变换群 $G$ 的每一个元素,当 $G$ 很大或仅能通过采样访问时,代价可能高得难以承受。我们研究了是否可以用一小部分固定的、在优化前采集并在其后重复使用的变换样本,来近似完全增强。在适当的正则性条件下,我们证明,以至少 $1-\delta$ 的概率,对所得的稀疏增强目标函数执行梯度下降(GD),可以在 $\mathcal{O}\bigl((\log |G|+\log(1/\delta))/\varepsilon^2\bigr)$ 次群变换预言机查询下返回完全增强目标函数的一个 $\varepsilon$-稳定点。相比之下,标准的群随机梯度下降(group-SGD)在每次迭代都采样新的变换,需要 $\mathcal{O}(1/\varepsilon^4)$ 次变换查询。因此,使用固定稀疏增强的梯度下降所需的变换查询次数,既少于对完全增强目标函数执行 GD 所需的次数,也少于 group-SGD 所需的次数。我们的证明技术具有独立的参考价值:利用群诱导算子的谱性质以及表示论中的工具,建立了随机群平均对完整群平均梯度场的一致逼近。
cs.LG / 193 / 2609.08135

KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization

KBBQ:一种预测性噪声定律与FP4量化中频谱平坦化的极限
Whalen, Lexington, Ito, Yuki, Sakamoto, Ryo
Abstract
We develop a second-order theory of quantization noise in matrix multiplication in which the quantization format is characterized by the variance it assigns to each element. The constant variance profile of integer quantization recovers existing integer-noise theory, while the multiplicative profile of floating-point rounding reduces the data dependence to a scalar, the participation factor $\kappa$, yielding a closed-form signal-to-noise-ratio law. The resulting functional also admits a closed-form upper bound $\kappa^{*}$ that no function-preserving linear transform can exceed and that is attained by a recent state-of-the-art method. Building on this analysis, we introduce KBBQ (\textbf{K}appa-\textbf{B}raked \textbf{B}lockwise \textbf{Q}uantization), which parameterizes the extent to which a transform approaches this ceiling. At W4A4, across four base models and two FP4 formats, KBBQ outperforms the prior state of the art without additional deployment-time computation.
Chinese Translation
我们建立了一种矩阵乘法中量化噪声的二阶理论,其中量化格式由其分配给每个元素的方差来刻画。整数量化的恒定方差分布恢复了现有的整数噪声理论,而浮点舍入的乘性分布则将数据依赖性简化为一个标量——参与因子 $\kappa$,从而得到闭式信噪比定律。所得的泛函还存在一个闭式上界 $\kappa^{*}$:任何保持函数性质的线性变换都无法超越该上界,而最近一种最先进的方法恰好达到了这一上界。基于此分析,我们提出了KBBQ(Kappa-Braked Blockwise Quantization,Kappa制动分块量化),它对变换逼近该上限的程度进行参数化。在W4A4设置下,跨越四个基础模型和两种FP4格式,KBBQ无需额外的部署时计算即超越了先前的最先进方法。
cs.LG / 194 / 2609.08136

GPU-Enabled Large-Scale Optimization Using Randomized Linear Algebra

基于随机化线性代数的GPU加速大规模优化
Rathore, Pratik, Frangella, Zachary, Nobel, Parth, Hu, Xuning, Udell, Madeleine
Abstract
This paper introduces rlaopt, a PyTorch-based package for large-scale optimization and scientific computing using randomized numerical linear algebra (RandNLA). Despite substantial progress in RandNLA-based algorithms, few implementations combine GPU acceleration with a simple interface for specifying optimization problems. rlaopt addresses this gap by providing GPU-enabled solvers for positive-definite linear systems and convex empirical risk minimization with constraints and regularizers. These solvers use RandNLA to accelerate conjugate gradient (NystromPCG), operator splitting (NysADMM), and stochastic gradient methods (SAPPHIRE). Moreover, rlaopt includes a modeling language that lets users specify problems using natural mathematical syntax. rlaopt automatically checks compatibility with the selected solver and performs the required problem decomposition. The solvers also support differentiation through their iterations, enabling applications such as hyperparameter tuning. Experiments on ridge regression, bounded multinomial logistic regression, and bounded elastic net identify when randomized preconditioning improves performance and demonstrate substantial speedups from GPU execution. The package is open-source under an Apache license, with source code at https://github.com/udellgroup/rlaopt and version 0.1.0 available on PyPI.
Chinese Translation
本文介绍了rlaopt,一个基于PyTorch的软件包,用于利用随机化数值线性代数(RandNLA)进行大规模优化与科学计算。尽管基于RandNLA的算法已取得显著进展,但很少有实现能够将GPU加速与用于指定优化问题的简洁接口相结合。rlaopt填补了这一空白,为正定线性系统以及带约束和正则化项的凸经验风险最小化提供了支持GPU的求解器。这些求解器利用RandNLA加速共轭梯度法(NystromPCG)、算子分裂法(NysADMM)和随机梯度法(SAPPHIRE)。此外,rlaopt包含一种建模语言,使用户能够以自然的数学语法描述问题。rlaopt会自动检查所选求解器的兼容性,并执行所需的问题分解。这些求解器还支持对其迭代过程的微分,从而实现超参数调优等应用。在岭回归、有界多项逻辑回归和有界弹性网络上的实验确定了随机化预处理在何时能够提升性能,并展示了GPU执行带来的显著加速。该软件包基于Apache许可开源,源代码位于 https://github.com/udellgroup/rlaopt,版本0.1.0可在PyPI上获取。
cs.LG / 195 / 2609.08152

Topology-induced Operators Reveal Complementary Graph Representations without Training

拓扑诱导算子无需训练即可揭示互补的图表示
Qin, Meng, Cui, Jinqiang, Zheng, Hongwei, Li, Weihua, Pei, Sen
Abstract
Graph representation learning has largely focused on designing increasingly sophisticated models to transform graph topology into vector representations, or embeddings. However, the extent to which embedding quality depends on model learning, rather than on the underlying topological transformations, remains unclear. Here, we show that informative embeddings can be derived without complicated model design and gradient-based training. Propagating random features through implicit hierarchical structures induced by random walks and anonymous walks yields embeddings that capture node proximity and structural role, respectively. These two training-free embeddings preserve complementary aspects of graph organization and perform competitively with classic and recent methods across various node-, edge-, and graph-level tasks. They often require substantially less computation, resulting in a favorable quality-efficiency trade-off. Combining the two types of embeddings further improves inference quality of some tasks compared with using either embedding type alone. Our results suggest that informative graph embeddings can arise from carefully chosen topological transformations before any learning operation is applied.
Chinese Translation
图表示学习的研究主要集中于设计日益复杂的模型,以将图拓扑结构转换为向量表示(即嵌入)。然而,嵌入质量在多大程度上依赖于模型学习,而非底层的拓扑变换,仍不清楚。本文表明,无需复杂的模型设计和基于梯度的训练,也可以获得有信息量的嵌入。通过将随机特征在由随机游走和匿名游走所诱导的隐式层次结构中进行传播,可分别获得捕捉节点邻近性和结构角色的嵌入。这两种无需训练的嵌入保留了图组织的互补方面,并在多种节点级、边级和图级任务上与经典及近期方法相比具有竞争力。它们通常所需的计算量大幅减少,从而形成良好的质量-效率权衡。与单独使用任一类型嵌入相比,将两类嵌入结合可进一步提升某些任务的推理质量。我们的结果表明,在任何学习操作之前,通过精心选择的拓扑变换即可产生有信息量的图嵌入。
cs.LG / 196 / 2609.08153

Geodesic-informed Generative Diffusion Model For Topology-preserved Image Video Generation

面向拓扑保持图像与视频生成的测地线引导生成扩散模型
Wu, Nian, Jayakumar, Nivetha, Xing, Jiarui, Zhang, Miaomiao
Abstract
Generative diffusion models have emerged as a class of powerful techniques for various imaging applications, including but not limited to synthesis, reconstruction, and segmentation. Despite their success, current generative models pose two key limitations. First, they primarily rely on image intensity and texture information, with limited attention to underlying object geometry. As a result, they do not guarantee geometric or topological consistency during the generation process, which is a crucial requirement for high-stakes domains such as computational anatomy, biology, and robotics, where preserving object structure is critical. Second, existing models fail to explicitly learn or represent shape changes in the generative process. Such deformation dynamics remain occluded within network parameters; hence leaving the transformation process uninterpretable and physically uninformed. To address these challenges, we introduce IGG (Image Generation informed by Geodesic dynamics), a novel framework that integrates topology-preserving geodesic principles into the diffusion-based generative process. In contrast to conventional methods that operate in image intensity space, IGG learns and synthesizes diverse samples within geodesic deformation spaces, where geometric object changes are learned as smooth and invertible smooth mappings from a given template/source image. Our code is publicly available at https://github.com/nellie689/IGG.
Chinese Translation
生成式扩散模型已成为各种成像应用(包括但不限于合成、重建和分割)中一类强大的技术。尽管取得了成功,当前的生成模型存在两个关键局限。第一,它们主要依赖图像强度和纹理信息,对底层物体几何的关注有限。因此,它们无法保证生成过程中的几何或拓扑一致性,而这对于计算解剖学、生物学和机器人学等高风险领域至关重要,因为在这些领域中保持物体结构至关重要。第二,现有模型未能显式地学习或表示生成过程中的形状变化。这种形变动态仍隐含于网络参数之中,导致变换过程不可解释且缺乏物理依据。为应对这些挑战,我们提出了 IGG(Image Generation informed by Geodesic dynamics,测地线动力学引导的图像生成),这是一种将拓扑保持的测地线原理融入基于扩散的生成过程的新型框架。与在图像强度空间中运行的传统方法不同,IGG 在测地线形变空间中学习并合成多样化样本,其中物体的几何变化被学习为从给定模板/源图像出发的平滑且可逆的映射。我们的代码已在 https://github.com/nellie689/IGG 公开。
cs.LG / 197 / 2609.08200

SIM: Subspace Interaction-based Method for Token-Level Text Anomaly Detection

SIM:一种基于子空间交互的词元级文本异常检测方法
Yan, Kehan, Tan, Yue, Chen, Qingfeng, Li, Shiyuan, Zheng, Yu, Liu, Yixin
Abstract
Token-level text anomaly detection, as an emerging trend of text anomaly detection, moves beyond coarse-grained document-level detection by localizing anomalous tokens within text. By providing fine-grained abnormality prediction, token-level text anomaly detection plays a critical role in various real-world applications, such as spam filtering and fake news detection. However, existing methods still rely on the global distance calculation for scoring, during which the local anomaly signals are severely diluted by numerous redundant normal feature dimensions. Moreover, pre-trained language models used in these methods inevitably smooth out surface anomalies, further limiting their effectiveness in token-level anomaly detection. To address these limitations, we propose a Subspace Interaction-based Method (SIM for short) for token-level text anomaly detection. To prevent local signal dilution, SIM adopts a subspace interaction-based anomaly detector, which decouples high-dimensional token embeddings into multiple low-dimensional ones, amplifying localized anomaly signals hidden within specific dimensions. To counteract the over-smoothing effect, we design a hard pseudo-anomaly generation module to construct pseudo-anomalous tokens, simulating the subtle anomalies obscured by semantic smoothing. Also, a probabilistic boundary loss is developed to standardize anomaly scores into statistical distances, effectively enforcing anomalous instances to deviate significantly from the normal distribution center. Extensive experiments on multiple benchmark datasets verify the effectiveness of SIM and demonstrate its remarkable efficiency, robustness, and interpretability. The source code is available at: https://github.com/yankehan/SIM-TAD.
Chinese Translation
词元级文本异常检测作为文本异常检测的新兴方向,超越了粗粒度的文档级检测,能够定位文本中的异常词元。通过提供细粒度的异常预测,词元级文本异常检测在垃圾信息过滤、虚假新闻检测等多种现实应用中发挥着关键作用。然而,现有方法仍然依赖全局距离计算进行评分,在此过程中,局部异常信号被大量冗余的正常特征维度严重稀释。此外,这些方法所使用的预训练语言模型不可避免地会平滑掉表面异常,进一步限制了其在词元级异常检测中的有效性。为解决上述局限,我们提出了一种基于子空间交互的方法(Subspace Interaction-based Method,简称 SIM)用于词元级文本异常检测。为防止局部信号被稀释,SIM 采用了基于子空间交互的异常检测器,将高维词元嵌入解耦为多个低维嵌入,从而放大隐藏在特定维度中的局部异常信号。为对抗过度平滑效应,我们设计了一个困难伪异常生成模块来构造伪异常词元,以模拟被语义平滑所掩盖的细微异常。此外,我们开发了概率边界损失,将异常评分标准化为统计距离,有效地促使异常实例显著偏离正常分布中心。在多个基准数据集上的大量实验验证了 SIM 的有效性,并展示了其卓越的效率、鲁棒性和可解释性。源代码发布于:https://github.com/yankehan/SIM-TAD。
cs.LG / 198 / 2609.08244

CUNO: Curriculum and Preference Optimization for Stable Graph Unlearning under Mass Deletion

CUNO:面向大规模删除下稳定图遗忘的课程学习与偏好优化
Zhang, Chenhan, Braytee, Ali, Bandara, Madhushi, Hao, Xin, Kennedy, Paul J., Piccardi, Massimo, Owen, Raymond
Abstract
Graph unlearning removes the influence of designated training data from a trained graph model without retraining from scratch. However, existing methods suffer a sharp drop in model utility under large deletion ratios (mass deletion), a phenomenon we refer to as catastrophic unlearning. We find that a key cause is the uniform treatment of all deleted samples, which is particularly damaging in graph learning: structural dependencies cause different nodes to play vastly different roles in the learned model, yet existing methods apply the same forgetting operation to the entire forget set. Based on this insight, we propose CUNO, a curriculum-based graph unlearning framework that removes the forget set progressively, ordering samples by their estimated unlearning difficulty across multiple stages. CUNO further employs a distribution-level negative preference optimization (NPO) objective at each curriculum stage that steers the model away from its original behavior on the current forget subset while preserving retained performance. Our theoretical analysis shows that the curriculum design is most beneficial when the forget set spans a wide range of unlearning difficulty, a condition naturally satisfied under mass deletion. Comprehensive experiments confirm that CUNO consistently mitigates catastrophic unlearning: at 20% deletion, it retains 74% of the original utility compared to 26-53% for existing methods, and maintains more than half the original utility even at 50% deletion. Our code is publicly available at https://anonymous.4open.science/r/cuno-D4FF.
Chinese Translation
图遗忘旨在在不从头重新训练的情况下,从已训练的图模型中移除指定训练数据的影响。然而,现有方法在大删除比例(大规模删除)下会出现模型效用的大幅下降,我们将这一现象称为灾难性遗忘(catastrophic unlearning)。我们发现一个关键原因是所有被删除样本被一视同仁地处理,这在图学习中尤为有害:结构依赖性使得不同节点在已学习模型中扮演着截然不同的角色,而现有方法却对整个遗忘集应用相同的遗忘操作。基于这一洞察,我们提出了CUNO,一个基于课程学习的图遗忘框架,它以渐进的方式移除遗忘集,并在多个阶段中根据估计的遗忘难度对样本进行排序。CUNO进一步在每个课程阶段采用分布层面的负偏好优化(NPO)目标,引导模型偏离其在当前遗忘子集上的原始行为,同时保留已保留数据的性能。我们的理论分析表明,当遗忘集涵盖较大范围的遗忘难度时,课程设计最为有效,而这一条件在大规模删除下自然成立。全面的实验证实,CUNO能够持续缓解灾难性遗忘:在20%删除比例下,它保留了74%的原始效用,而现有方法仅为26-53%;即使在50%删除比例下,它仍能保持超过一半的原始效用。我们的代码已公开发布于 https://anonymous.4open.science/r/cuno-D4FF。
cs.LG / 199 / 2609.08253

Revisiting Spectral Representations in Generative Diffusion Models

重访生成扩散模型中的谱表示
Wang, Yuehao, Wang, Peihao, Jiang, Hanwen, Yang, Ziyi, Huang, Qixing, Wang, Zhangyang
Abstract
Diffusion models have shown remarkable performance on diverse generation tasks. Recent work finds that imposing representation alignment on the hidden states of diffusion networks can both facilitate training convergence and enhance sampling quality, yet the mechanism driving this synergy remains insufficiently understood. In this paper, we investigate the connection between self-supervised spectral representation learning and diffusion generative models through a shared perspective on perturbation kernels. On the diffusion side, samples (e.g., images, videos) are produced by reversing a stochastic noise-injection process specified by Gaussian kernels; on the spectral representation side, spectral embeddings emerge from contrasting positive and negative relations induced by random perturbation kernels. Motivated by this, we propose a self-supervised spectral representation alignment method to facilitate diffusion model training. In addition, we clarify how joint spectral learning can benefit diffusion training from a geometric perspective. Furthermore, we find that the optimization of the spectral alignment objective is in an equivalent form of diffusion score distillation in the representation space. Building on these findings, we integrate a spectral regularizer into diffusion training objectives to improve the performance of diffusion models on multiple datasets. Experiments across images and 3D point clouds show consistent gains in generation quality. Code is released at https://github.com/yuehaowang/spectral-reg-diffusion.
Chinese Translation
扩散模型在多种生成任务上表现出卓越的性能。近期研究发现,对扩散网络的隐藏状态施加表示对齐既可以促进训练收敛,又能提升采样质量,但驱动这种协同效应的机制仍缺乏充分的理解。在本文中,我们通过扰动核(perturbation kernels)这一共同视角,研究了自监督谱表示学习与扩散生成模型之间的联系。在扩散一侧,样本(如图像、视频)通过反转由高斯核定义的随机噪声注入过程生成;在谱表示一侧,谱嵌入则源自对比由随机扰动核诱导的正负关系。受此启发,我们提出了一种自监督谱表示对齐方法以促进扩散模型的训练。此外,我们从几何视角阐明了联合谱学习如何有益于扩散训练。进一步地,我们发现谱对齐目标的优化在表示空间中等价于扩散分数蒸馏(diffusion score distillation)的一种形式。基于这些发现,我们将谱正则化项整合到扩散训练目标中,以提升扩散模型在多个数据集上的性能。在图像和三维点云上的实验显示生成质量获得了一致的提升。代码已发布于 https://github.com/yuehaowang/spectral-reg-diffusion。
cs.LG / 200 / 2609.08268

Synergistic Fusion of Topological Structure and Temporal Semantics of Mobility for Urban Region Embedding

面向城市区域嵌入的移动性拓扑结构与时间语义的协同融合
Kim, Namwoo, Chang, Jeeyun, Lee, Kanghoon, Yoon, Yoonjin
Abstract
Urban region embeddings have shown promising results in diverse urban sensing tasks such as crime, income, and service-call prediction. Recent methods improve representation quality by integrating mobility data with auxiliary modalities, using cross-view attention or contrastive objectives to align heterogeneous features into a unified region representation. However, leveraging the temporal dynamics of human mobility remains under-explored. Regional inflow and outflow fluctuate throughout the day, and inter-region connections emerge, persist, and dissolve over time. Moreover, prevailing fusion strategies combine views additively and miss the joint signal that emerges only when views co-occur. To address these gaps, we propose Mobility Stream-Structure Synergy (MoSS), which derives complementary views from mobility data: a Sequence view that preserves each region's hourly inflow/outflow profile, and a Structure view based on zigzag persistence diagrams that capture how regional connectivity emerges, persists, and dissolves over time. A synergy module then extracts emergent representations from the co-occurrence of these views through multi-degree interactions, explicitly capturing higher-order signal across views. Extensive experiments on New York City and Chicago show that MoSS achieves state-of-the-art performance across three downstream tasks using mobility data alone, outperforming baselines that rely on auxiliary modalities.
Chinese Translation
城市区域嵌入在犯罪、收入和服务呼叫预测等多种城市感知任务中已展现出良好效果。近期方法通过将移动性数据与辅助模态结合,利用跨视图注意力或对比学习目标将异构特征对齐为统一的区域表示,从而提升表示质量。然而,如何利用人类移动性的时间动态特性仍未得到充分探索。区域流入和流出在一天中不断波动,区域间的连接也随时间出现、持续和消散。此外,主流的融合策略采用加性方式组合各视图,遗漏了仅在视图共现时才涌现的联合信号。为弥补这些不足,我们提出了移动性流-结构协同方法(Mobility Stream-Structure Synergy, MoSS),从移动性数据中提取互补视图:一是序列(Sequence)视图,保留每个区域每小时的流入/流出特征;二是结构(Structure)视图,基于锯齿形持续图刻画区域连通性如何随时间出现、持续和消散。随后,协同模块通过多阶交互从这些视图的共现中提取涌现表示,显式捕捉跨视图的高阶信号。在纽约市和芝加哥数据集上的大量实验表明,MoSS 仅使用移动性数据即可在三个下游任务中取得最先进的性能,优于依赖辅助模态的基线方法。
cs.LG / 201 / 2609.08276

Online Signature Verification Using Augmented Path Signature and T-Mamba

基于增广路径签名与T-Mamba的在线签名验证
Li, Ruiling, Yang, Danyu
Abstract
Handwritten signature verification is vital for personal authentication across commercial and financial applications. Although deep learning methods are widely adopted for online signature verification (OSV), they often struggle with capturing highly discriminative features and modelling long-range dependencies. To address these issues, we propose a novel framework that integrates the augmented path signature (APS) descriptor with the T-Mamba model. The APS descriptor first applies time and basepoint augmentations, then computes sliding-window path signatures. The path signature is a non-parametric feature map from rough path theory that effectively captures geometric structures and nonlinear inter-channel interactions. Inspired by the efficacy of state space models (SSMs) in sequence modelling, our T-Mamba model employs a hybrid design combining two temporal convolutional network (TCN) blocks with a time-scanning Mamba. This design enables the model to learn both local temporal patterns and global long-range dependencies, substantially improving verification accuracy. Our framework achieves state-of-the-art EERs on three public benchmark datasets (MCYT-100, SVC-2004 Task 2, DeepSignDB), validating its effectiveness and robustness, especially when the training data is limited. Our code is publicly available at https://github.com/DLRL04/OSV-using-APS-and-T-Mamba.
Chinese Translation
手写签名验证对于商业和金融应用中的个人身份认证至关重要。尽管深度学习方法已被广泛应用于在线签名验证(OSV),但它们在捕获高度判别性特征和建模长程依赖方面常常面临困难。为解决这些问题,我们提出了一种将增广路径签名(APS)描述子与T-Mamba模型相集成的新框架。APS描述子首先进行时间和基点增广,然后计算滑动窗口路径签名。路径签名是源自粗糙路径理论的一种非参数特征映射,能够有效捕获几何结构和通道间的非线性交互。受状态空间模型(SSM)在序列建模中出色表现的启发,我们的T-Mamba模型采用混合设计,将两个时间卷积网络(TCN)模块与时间扫描Mamba相结合。该设计使模型能够同时学习局部时间模式和全局长程依赖,显著提升了验证准确率。我们的框架在三个公开基准数据集(MCYT-100、SVC-2004 Task 2、DeepSignDB)上取得了最先进的等错误率(EER),验证了其有效性和鲁棒性,尤其是在训练数据有限的情况下。我们的代码已公开发布于 https://github.com/DLRL04/OSV-using-APS-and-T-Mamba。
cs.LG / 202 / 2609.08277

Adaptively Incorporating Directional Hints into Zeroth-Order Optimization

将方向性提示自适应地引入零阶优化
Ryabchenko, Alexander, Qian, Jian, Mou, Wenlong
Abstract
We study zeroth-order optimization of non-convex functions with the aid of directional hints, which are cheap but potentially inaccurate approximations of the true gradient direction, given by linear subspaces at each iteration. To leverage these hints adaptively while maintaining robustness to their quality, we introduce Control-Variate Zeroth-Order Descent (CV-ZOD), a new framework that refines the classical zeroth-order gradient estimator with a control variate that can be set based on the directional hints. We first show that the oracle algorithm that optimally sets the reference vector and step size at each iteration achieves a convergence rate that interpolates between the first-order $O(1/T)$ rate and the zeroth-order $O(d/T)$ rate, depending on the quality of the hints along the trajectory. We then develop a practical variant of CV-ZOD that achieves the same oracle guarantee up to logarithmic factors, without any prior knowledge of the hint quality. We validate the method empirically on simulation-based scientific optimization tasks, demonstrating sustained progress on non-convex landscapes where zeroth-order descent is slower and existing guided methods stall as guidance deteriorates.
Chinese Translation
我们研究在方向性提示辅助下对非凸函数的零阶优化问题。方向性提示是对真实梯度方向的廉价但可能不准确的近似,以每次迭代给出的线性子空间形式提供。为了在保持对提示质量的鲁棒性的同时自适应地利用这些提示,我们提出了控制变量零阶下降(Control-Variate Zeroth-Order Descent, CV-ZOD),这是一个新框架,它通过一个可基于方向性提示设置的控制变量来改进经典的零阶梯度估计器。我们首先证明,在每次迭代中最优设置参考向量和步长的预言机(oracle)算法可实现一个介于其一阶 O(1/T) 收敛率与零阶 O(d/T) 收敛率之间的收敛率,具体取决于提示在整个轨迹上的质量。随后,我们开发了 CV-ZOD 的一个实用变体,该变体在无需任何关于提示质量的先验知识的情况下,在对数因子范围内达到与预言机相同的保证。我们在基于仿真的科学优化任务上对该方法进行了实证验证,结果表明:在零阶下降速度较慢、而现有引导方法随着引导质量下降而停滞的非凸地形上,本方法能够持续取得进展。
cs.LG / 203 / 2609.08286

HypLTSF: A Hyperbolic Geometric View of Multi-Scale Hierarchies for Long-Term Time Series Forecasting

HypLTSF:用于长期时间序列预测的多尺度层次结构的双曲几何视角
Kim, Namwoo, Baik, Hyungryul, Yoon, Yoonjin
Abstract
Multi-scale modeling has become an effective approach for long-term time series forecasting, capturing temporal patterns that range from fine-grained local dynamics to coarse global trends. Representations across these temporal scales are inherently hierarchical, with coarser scales abstracting and aggregating information from finer ones. While existing approaches readily exchange information across these scales, the hierarchy itself is typically left as an emergent byproduct of such interactions rather than captured as a geometric structure in its own right. In this paper, we introduce HypLTSF, a framework that endows the multi-scale hierarchy with a concrete geometric form by embedding scale-wise representations into the Poincar\'e ball, whose exponentially expanding volume naturally accommodates hierarchical structures. To align this geometry with the temporal hierarchy, HypLTSF imposes two constraints: (1) a radial constraint that orders embeddings by their level of abstraction, and (2) an angular constraint that groups fine-scale patterns sharing a common coarser-scale ancestor. Extensive experiments on long-term time series forecasting benchmarks show that HypLTSF achieves state-of-the-art performance, suggesting that explicitly modeling the multi-scale hierarchy as a geometric structure is effective for forecasting.
Chinese Translation
多尺度建模已成为长期时间序列预测的一种有效方法,能够捕获从细粒度局部动态到粗粒度全局趋势的多层次时间模式。这些时间尺度上的表示本质上具有层次性,较粗的尺度通过抽象和聚合较细尺度的信息而形成。现有方法虽然能够轻松地在不同尺度之间交换信息,但层次结构本身通常只是这种交互作用的自然副产物,而非作为一种独立的几何结构被显式刻画。本文提出HypLTSF框架,通过将各尺度表示嵌入到庞加莱球(Poincaré ball)中,为多尺度层次结构赋予具体的几何形式——庞加莱球指数级扩张的体积天然适合容纳层次结构。为使该几何与时间层次结构对齐,HypLTSF引入两个约束:(1)径向约束,按抽象层次对嵌入进行排序;(2)角度约束,将共享同一粗尺度祖先的细尺度模式进行分组。在长期时间序列预测基准上的大量实验表明,HypLTSF取得了最先进的性能,说明将多尺度层次结构显式建模为几何结构对预测任务是有效的。
cs.LG / 204 / 2609.08330

EMBLEM: Enhancing Multi-script Table Detection through Masking

EMBLEM:通过掩码增强多文字体系表格检测
Kudale, Dhruv, Brahmi, Udhay, Ramakrishnan, Ganesh
Abstract
Table detection is a core task in document analysis, supporting downstream applications such as information retrieval, document reconstruction, and visual question answering. While existing deep learning models perform well on English and Chinese documents, they struggle with multilingual, multi-script documents due to script diversity and the limited availability of labeled data. To address this challenge, we introduce MANDALA (Multi-script Annotated Documents for Table Detection), a manually curated dataset of 2,323 table-containing pages spanning 18 languages and 15 scripts across diverse domains. We also propose EMBLEM, a masking-based paradigm for Multi-script Table Detection (MTD). EMBLEM generates masked images that conceal script- and font-specific details, enabling models pre-trained on abundant English documents to focus on script-agnostic page layout. Experiments across three table detection architectures show that EMBLEM consistently outperforms strong baselines on MANDALA while remaining competitive on five standard English-dominant benchmarks. Using only English masked images for fine-tuning, with no multi-script training data, EMBLEM achieves an absolute F1-score gain of 20.8% on MANDALA. We release MANDALA along with the accompanying code and models at https://github.com/IITB-LEAP-OCR/EMBLEM.git.
Chinese Translation
表格检测是文档分析中的一项核心任务,支持信息检索、文档重构和视觉问答等下游应用。尽管现有的深度学习模型在英文和中文文档上表现良好,但由于文字体系的多样性以及标注数据的有限性,它们在多语言、多文字体系文档上表现不佳。为应对这一挑战,我们推出了MANDALA(用于表格检测的多文字体系标注文档数据集),这是一个人工整理的数据集,包含2,323个含表格的页面,涵盖18种语言、15种文字体系以及多个领域。我们还提出了EMBLEM,一种基于掩码的多文字体系表格检测(MTD)范式。EMBLEM生成掩码图像以隐藏特定于文字体系和字体的细节,使在大规模英文文档上预训练的模型能够专注于与文字体系无关的页面布局。在三种表格检测架构上的实验表明,EMBLEM在MANDALA上始终优于强基线模型,同时在五个以英文为主的标准基准测试上也保持竞争力。仅使用英文掩码图像进行微调,且不使用任何多文字体系训练数据的情况下,EMBLEM在MANDALA上取得了20.8%的绝对F1分数提升。我们在 https://github.com/IITB-LEAP-OCR/EMBLEM.git 发布了MANDALA数据集以及相关代码和模型。
cs.LG / 205 / 2609.08337

Distillation as Probability Transport: Routed On-Policy Distillation

蒸馏作为概率传输:路由化在线策略蒸馏
Xia, Tianle, Hu, Lingxiang, Sun, Yiding, Shang, Linfang, Xu, Ming, Xu, Lan, Zheng, Ning, Xu, Wei, Jiang, Jie
Abstract
On-policy distillation (OPD) transfers teacher knowledge on student-generated trajectories, but efficient sampled objectives reduce the teacher distribution to scalar credit on individual tokens. Such credit indicates whether a token should gain or lose probability, yet leaves the corresponding redistribution unspecified. We recast OPD as teacher-guided probability transport and propose RouteOPD (Routed On-Policy Distillation), which decomposes local teacher--student disagreement into student-excess sources and teacher-deficit destinations and couples them into explicit transport pairs. RouteOPD optimizes pairwise log-odds toward jointly realizable targets obtained from a bounded teacher potential, while adapting the transport budget to the concentration of teacher demand. This formulation directs updates toward teacher-preferred destinations and controls their magnitude within a single transport operator. Experiments across four teacher--student settings and four mathematical-reasoning benchmarks demonstrate that RouteOPD consistently outperforms sampled reverse-KL OPD, with improvements accompanied by higher routing fidelity and lower background leakage. These results demonstrate the effectiveness of explicitly modeling probability transport in on-policy distillation.
Chinese Translation
在线策略蒸馏(On-policy distillation, OPD)在学生模型生成的轨迹上迁移教师模型的知识,但基于采样的高效目标函数将教师分布压缩为作用于单个词元的标量形式的信用信号。这种信用只能指示某个词元应当增加还是减少概率,却未指定相应的概率如何重新分配。我们将 OPD 重新表述为教师引导的概率传输,并提出 RouteOPD(Routed On-Policy Distillation),该方法将局部的教师—学生分歧分解为学生的概率过剩来源(student-excess sources)和教师的概率不足目的地(teacher-deficit destinations),并将它们耦合成显式的传输对。RouteOPD 基于有界的教师势函数(teacher potential)所给出的联合可实现目标,优化成对对数几率(pairwise log-odds),同时根据教师需求的集中程度自适应地调整传输预算。这一表述将更新引导至教师偏好的目的地,并在单一传输算子内控制其幅度。在四个教师—学生设置和四个数学推理基准上的实验表明,RouteOPD 一致优于基于采样的反向 KL OPD,且在提升性能的同时具有更高的路由保真度和更低的后台泄漏(background leakage)。这些结果证明了在在线策略蒸馏中显式建模概率传输的有效性。
cs.LG / 206 / 2609.08341

TV-Regulated OPD: Direction Matters in On-Policy Distillation

TV-Regulated OPD:方向在在线策略蒸馏中至关重要
Xiao, Han, Niu, Yifan, Liu, Dongyi, Luo, Chang, Li, Jia
Abstract
On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). However, the supervision signals in mainstream OPD methods suffer from high variance and noise which is generally instable during training. In this work, we systematically investigated what really matters to the performance and the fundamental mechanisms behind the instability during training. We found that retaining only the sign of token-level advantages is sufficient to achieve the performance comparable to standard OPD. Meanwhile, smoother and bounded advantages can stabilize the training process without sacrificing its performance. These motivated us to shape the advantages using the Total Variation (TV) and propose a robust TV regulated On-Policy Distillation (TV-OPD) method. Benefiting from the bounded and diminished advantages, TV-OPD exhibits stable training dynamics and steady late-stage performance. We conducted comprehensive experiments and found that, across various settings, TV-OPD consistently achieved better performance and lower variance in the late-stage of training.
Chinese Translation
在线策略蒸馏(On-Policy Distillation, OPD)促进了大语言模型(LLM)后训练阶段从领域专家模型到学生模型的知识转移。然而,主流OPD方法中的监督信号存在高方差和噪声问题,在训练过程中通常表现不稳定。在本工作中,我们系统性地研究了真正影响性能的因素以及训练不稳定背后的根本机制。我们发现,仅保留token级优势(advantage)的符号(即方向信息)就足以取得与标准OPD相当的性能。同时,更平滑且有界的优势函数能够在不牺牲性能的前提下稳定训练过程。基于这些发现,我们利用全变差(Total Variation, TV)对优势函数进行塑形,提出了一种鲁棒的TV正则化在线策略蒸馏方法(TV-OPD)。得益于有界且衰减的优势函数,TV-OPD展现出稳定的训练动态和稳定的后期性能。我们进行了全面的实验,发现在各种设置下,TV-OPD在训练后期始终取得了更优的性能和更低的方差。
cs.LG / 207 / 2609.08354

Geometry-Aware Bayesian Parameter-Efficient Fine-Tuning on the Stiefel Manifold via Stein Variational Gradient Descent

基于Stein变分梯度下降的Stiefel流形上的几何感知贝叶斯参数高效微调
Tran, Quang-Duy, Le, Trung, Duong, Bao, Nguyen, Phuoc, Nguyen, Thin
Abstract
Several geometry-aware approaches to low-rank adaptation have emerged for parameter-efficient fine-tuning of large pre-trained models. These methods aim to take full advantage of the geometric structure of low-rank manifolds for improving the efficiency in subspace utilization and reducing redundancy by enforcing orthogonality constraints during optimization. The strong empirical results of these techniques have motivated further study into whether predictions from such geometry-based adaptation methods could be overconfident. In this paper, we build on the singular value decomposition factorization of adapters to develop a framework based on Stein variational gradient descent (SVGD). In this formulation, the low-rank matrices are transported along the Stiefel manifold to match the targeted distributions while retaining their crucial geometric structure. Since this geometry-aware SVGD approach provides multiple solutions during inference, it supports uncertainty quantification and produces better-calibrated adapters on the Stiefel manifold. Extensive experiments show that our method delivers strong model calibration and attains higher prediction accuracy than SVGD and related uncertainty estimation methods that are formulated in Euclidean space.
Chinese Translation
针对大型预训练模型的参数高效微调,已有若干几何感知的低秩适配方法出现。这些方法旨在充分利用低秩流形的几何结构,通过在优化过程中施加正交性约束,提高子空间利用效率并减少冗余。这些技术的强大实证结果引发了进一步的研究:此类基于几何的适配方法所产生的预测是否可能过于自信。本文在适配器的奇异值分解分解基础上,构建了一个基于Stein变分梯度下降(SVGD)的框架。在该框架中,低秩矩阵沿Stiefel流形被传输以匹配目标分布,同时保持其关键的几何结构。由于这种几何感知的SVGD方法在推理过程中提供多个解,它支持不确定性量化,并在Stiefel流形上产生校准性能更好的适配器。大量实验表明,我们的方法具有出色的模型校准能力,并且在预测精度上优于在欧几里得空间中构造的SVGD及相关不确定性估计方法。
cs.LG / 208 / 2609.08368

Miles v0.1: Production-Level Post-Training

Miles v0.1:生产级后训练系统
RadixArk, :, Chen, Tom, Cheng, Mao, Dong, Shi, Du, Kangrui, Jiang, Yanbin, Li, Jiajun, Li, Yiming, Lin, Tao, Su, Yusheng, Ye, Andy, Yuan, Yueming, Zeng, Zhichen
Abstract
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable. With accuracy, efficiency, reliability, and scalability as first-class goals, Miles aims to make frontier-scale RL accessible to researchers and enterprises alike. This report walks through the system end to end: rollout engines built on SGLang, a trainer with a choice of two backends (NVIDIA Megatron-LM and PyTorch FSDP), and three weight-synchronization transports for different deployment topologies. Beyond full-parameter RL, Miles also supports LoRA RL, on-policy distillation, supervised fine-tuning, and true-on-policy rollout-training alignment, and extends the same architecture to diffusion models. We close with an end-to-end case study: fully asynchronous agentic RL on a GLM-5.2 744B-A40B model over terminal-use coding tasks, running on 64 NVIDIA GB300 GPUs with a median step time of 263 seconds over the first 30 measured steps. Miles is open-sourced at https://github.com/radixark/miles, with the project website at https://miles.radixark.com.
Chinese Translation
我们提出 Miles v0.1,一个全栈、可投入生产的前沿后训练系统。Miles 基于 slime 的简洁设计构建,围绕单一原则设计强化学习(RL)训练循环的每个阶段:各组件应经过验证、简洁且可定制。以准确性、高效性、可靠性和可扩展性作为首要目标,Miles 旨在让前沿规模的强化学习对研究者和企业都触手可及。本报告对系统进行了端到端的介绍:基于 SGLang 构建的 rollout 引擎、可选两种后端(NVIDIA Megatron-LM 和 PyTorch FSDP)的训练器,以及面向不同部署拓扑的三种权重同步传输方式。除全参数强化学习外,Miles 还支持 LoRA 强化学习、在线策略蒸馏、监督微调以及真正的在线策略 rollout 与训练对齐,并将相同的架构扩展至扩散模型。最后我们给出一个端到端案例研究:在 GLM-5.2 744B-A40B 模型上,针对终端使用类编程任务进行完全异步的智能体强化学习,运行于 64 块 NVIDIA GB300 GPU,在最初 30 个测量步骤中步时中位数为 263 秒。Miles 已开源,地址为 https://github.com/radixark/miles,项目网站为 https://miles.radixark.com。
cs.LG / 209 / 2609.08375

IPM-FM: A Foundation Model with Consensus Feature Selection for Industrial Process Monitoring

IPM-FM:一种基于共识特征选择的工业过程监测基础模型
Cao, Liang, Liu, Weide, Qin, Yan, Cheng, Jun, Lin, Weisi, Gopaluni, Bhushan
Abstract
Industrial process monitoring is fundamental to the safety and economic performance of modern process plants. Current practice remains a one-task-one-model paradigm that is label-inefficient and prone to degradation under operating drift. Foundation models have reshaped language, vision, and generic time-series forecasting, but it has not been adapted to industrial process monitoring. This setting poses domain-specific challenges, including safety-critical decisions and asymmetric sampling between process variables and laboratory measurements. We propose the industrial process monitoring foundation model (IPM-FM). It first learns general-purpose representations from unlabeled industrial process data through self-supervised pretraining, then adapts to specific monitoring tasks using a small amount of task-labeled data, and finally produces calibrated predictions through an uncertainty-aware prediction head. IPM-FM integrates a self-supervised Informer backbone with a multi-criteria consensus feature selector, a recursive lag-feature regression head, and a calibrated Monte Carlo dropout uncertainty module. On a seven-year hydrotreater dataset for diesel flash-point soft sensing, IPM-FM attains an RMSE of 2.99, $R^2$ of 0.50, and 97\% coverage of its 95\% predictive interval, outperforming the strongest classical and from-scratch sequence baselines by 8.3\% and 14.6\% in RMSE respectively, supporting the viability of a unified pretraining--adaptation framework for industrial process monitoring.
Chinese Translation
工业过程监测对现代流程工厂的安全性和经济性能至关重要。目前的实践仍然采用“一任务一模型”的范式,其标注效率低下,且在工况漂移下容易性能退化。基础模型(Foundation Model)已经重塑了语言、视觉以及通用时间序列预测领域,但尚未被适配到工业过程监测中。这一应用场景带来了特定领域的挑战,包括安全关键决策以及过程变量与实验室化验数据之间的非对称采样问题。我们提出了工业过程监测基础模型(Industrial Process Monitoring Foundation Model, IPM-FM)。该模型首先通过自监督预训练从无标签的工业过程数据中学习通用表征,然后利用少量任务标注数据适配到具体监测任务,最后通过不确定性感知的预测头输出经校准的预测结果。IPM-FM 集成了自监督 Informer 骨干网络、多准则共识特征选择器、递归滞后特征回归头以及经校准的蒙特卡洛 dropout 不确定性模块。在一个为期七年的加氢处理装置数据集上的柴油闪点软测量任务中,IPM-FM 取得了 2.99 的 RMSE、0.50 的 $R^2$,以及 95% 预测区间下 97% 的覆盖率,在 RMSE 上分别比最强的经典基线模型和从零训练的序列基线模型提升了 8.3% 和 14.6%,验证了统一的“预训练—适配”框架在工业过程监测中的可行性。
cs.LG / 210 / 2609.08379

Geographically Regularized AUC-Maximizing Personalized Federated Learning

地理正则化的AUC最大化个性化联邦学习
Hiraishi, Mayu, Tanioka, Kensuke, Shimokawa, Toshio
Abstract
Accurate diagnostic and risk-prediction models are important for supporting clinical decision-making during infectious disease outbreaks. However, privacy and governance requirements may restrict patient-level data sharing across healthcare institutions, and data distributions often vary. Moreover, AUC is widely used to evaluate discriminative performance, motivating its direct optimization in model development. We propose geographically regularized AUC-maximizing personalized federated learning (GrAUC-PFL), which directly optimizes a smooth pairwise AUC surrogate to learn personalized models while keeping patient-level data local and accounting for institutional heterogeneity. Graph-based regularization encourages geographically neighboring institutions to have similar coefficient vectors while retaining a personalized models. Simulations and a real-data application suggest improved discriminative performance, particularly when geographically neighboring institutions have similar data-generating characteristics.
Chinese Translation
准确的诊断和风险预测模型对于传染病暴发期间支持临床决策至关重要。然而,隐私和治理要求可能限制患者层面数据在医疗机构之间的共享,且数据分布往往存在差异。此外,AUC被广泛用于评估模型的判别性能,这促使在模型开发中对其进行直接优化。我们提出了地理正则化的AUC最大化个性化联邦学习(GrAUC-PFL),该方法直接优化一个平滑的成对AUC代理函数,在保持患者数据本地化并考虑机构异质性的同时学习个性化模型。基于图的正则化鼓励地理上邻近的机构拥有相似的系数向量,同时保留个性化模型。仿真研究和真实数据应用的结果表明,该方法提升了判别性能,尤其是在地理上邻近的机构具有相似数据生成特征的情况下。
cs.LG / 211 / 2609.08381

Equivariance Breaks the Learning Rate

等变性打破学习率
Manolache, Andrei, Niepert, Mathias
Abstract
Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on these architectures without explaining why. We identify one source of this difference inside equivariant linear layers. Each irrep block learns a channel-mixing matrix $W_l$ shared across its $2l+1$ components, giving the expanded map $W_l \otimes I_{2l+1}$. For a single application of the layer, the gradient of $W_l$ sums $2l+1$ outer product contributions and has rank at most $2l+1$. Adam rescales stored weights individually without using the irrep boundaries, so one learning rate can produce different spectral step sizes across blocks within a layer. We address this mismatch by normalizing each block update separately, without introducing a new hyperparameter. This changes only the scale of the update, leaving Adam's moment estimates and its direction within each block unchanged. We evaluate the mechanism in a controlled $\mathrm{SO}(3)$-equivariant model with a matched dense control and in an e3nn interatomic potential model trained on rMD17 and MD22. The toy setup isolates a mismatch that grows with width while the dense control shows no corresponding growth. In the interatomic potential model, block normalization and tuning Adam's momentum coefficients independently improve performance, but neither alone matches Muon. Combined, they make Adam competitive with Muon on all datasets, indicating that blockwise step control and momentum accumulation account for much of Muon's advantage.
Chinese Translation
等变网络通常使用Adam进行训练,然而近期有研究报道,诸如Muon等矩阵结构优化器在这类架构上表现更好,但未解释其原因。我们在等变线性层内部识别出造成这一差异的一个来源。每个不可约表示(irrep)块学习一个在其2l+1个分量之间共享的通道混合矩阵$W_l$,从而给出扩展映射$W_l \otimes I_{2l+1}$。在单次应用该层时,$W_l$的梯度是2l+1个外积贡献之和,且秩至多为2l+1。Adam在重缩放存储的权重时并未利用irrep边界,因此同一个学习率可能在一个层内的不同块之间产生不同的谱步长。我们通过分别对每个块的更新进行归一化来解决这一不匹配问题,且不引入新的超参数。这仅改变更新的尺度,而不改变Adam的矩估计及其在每个块内的更新方向。我们在一个受控的SO(3)等变模型(配有匹配的稠密对照)以及在rMD17和MD22数据集上训练的e3nn原子间势模型中评估了该机制。玩具实验设置分离出一种随宽度增长而加剧的不匹配,而稠密对照则未显示相应的增长。在原子间势模型中,块归一化和调整Adam的动量系数均能独立提升性能,但单独使用任何一种都无法匹敌Muon。两者结合后,Adam在所有数据集上均能与Muon竞争,表明分块步长控制和动量累积是Muon优势的主要来源。
cs.LG / 212 / 2609.08399

MLIP Detective: Active Failure Mode Discovery Beyond Benchmark Scores for Machine-Learning Interatomic Potentials

MLIP Detective:超越基准分数的机器学习原子间势主动失效模式发现
Okuno, Ryuhei, Charoenphakdee, Nontawat, Hisama, Kaoru, Tsuboi, Yuta
Abstract
Universal machine-learning interatomic potentials (u-MLIPs) aim to generalize across diverse configurations. Benchmarks enable reproducible evaluation but may not expose failures outside their predefined scope. Here, we show that physics-informed search can complement benchmark-based evaluation by uncovering hidden failure modes. We introduce MLIP Detective, an agentic framework for active failure mode discovery. Starting from benchmark evidence, MLIP Detective generates falsifiable, physics-informed failure hypotheses, screens them with inexpensive simulations, and escalates only the most suspicious cases to human experts together with proposed verification protocols. Without issue-specific prompting, MLIP Detective identified and characterized a systematic anomaly in MACE-MPA-0: the model predicted some relaxed adsorbate-surface systems involving O- or F-containing adsorbates to be higher in energy than their corresponding separated fragments. Using cross-model comparisons, MLIP Detective further inferred a likely training-data origin for the anomaly, consistent with recent reports.
Chinese Translation
通用机器学习原子间势(universal machine-learning interatomic potentials, u-MLIPs)旨在泛化至多样化的构型。基准测试虽然能够实现可复现的评估,但可能无法暴露其预定义范围之外的失效情况。在本研究中,我们展示了基于物理信息的搜索可以补充基于基准的评估,从而揭示隐藏的失效模式。我们提出了 MLIP Detective,一个用于主动失效模式发现的智能体(agentic)框架。从基准测试证据出发,MLIP Detective 生成可证伪的、基于物理信息的失效假设,通过低成本的模拟对其进行筛选,并仅将最可疑的案例连同提出的验证方案一并升级给人类专家。在无需针对特定问题进行提示的情况下,MLIP Detective 发现并刻画了 MACE-MPA-0 中的一个系统性异常:该模型预测某些涉及含氧或含氟吸附物种的弛豫吸附质-表面体系的能量高于其相应的分离碎片。通过跨模型比较,MLIP Detective 进一步推断出该异常可能源于训练数据,这与近期的相关报道相一致。
cs.LG / 213 / 2609.08404

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

环境作为脚手架:通过丰富反馈来引导自进化智能体完成长程任务
Yuan, Hongbang, Jin, Zhuoran, Cao, Yixin
Abstract
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. While conventional \textit{agent-side warming} up via supervised fine-tuning (SFT) can alleviate this, it is frequently limited by data scarcity and constrained exploration. To address this, we propose a paradigm shift to \textit{environment-side adaptation} by constructing \textbf{F}eedback-\textbf{E}nriched \textbf{E}nvironments (\textbf{FEEs}). Through a pilot study, we establish a feedback design strategy that reformulates environments by transitioning from action guidance to observation enrichment during the later stages of both intra-episode exploration and inter-episode evolution. Large-scale experiments on SciWorld and BFCL benchmarks using various Qwen3 model scales and RL algorithms such as GRPO, GSPO, and DAPO demonstrate that FEEs consistently yield performance improvements over standard settings. Furthermore, our analysis reveals that training with FEEs \textbf{(1)} stabilizes training dynamics by reducing entropy volatility, \textbf{(2)} facilitates proactive state-space exploration in difficult tasks, \textbf{(3) }ensures the internalization of environmental guidance into policy weights rather than acting as a mere inference-time prior, and \textbf{(4) }identifies intra-group feedback consistency as a critical boundary for stable optimization.
Chinese Translation
大语言模型在静态推理方面展现出卓越的能力,但通过强化学习(RL)将其训练为能够完成长程任务的自主智能体时,往往受到严重奖励稀疏性的阻碍。虽然传统的通过监督微调(SFT)进行"智能体侧预热"可以缓解这一问题,但通常受限于数据稀缺和探索受限。为解决这一问题,我们提出向"环境侧适应"的范式转变,构建反馈丰富环境(Feedback-Enriched Environments, FEEs)。通过一项试点研究,我们确立了一种反馈设计策略:在回合内探索和回合间演化的后期阶段,将环境从动作引导逐步转变为观察增强。在SciWorld和BFCL基准上,使用多种规模的Qwen3模型以及GRPO、GSPO和DAPO等RL算法的大规模实验表明,FEEs始终能带来优于标准设置的性能提升。此外,我们的分析揭示了使用FEEs训练能够:(1)通过降低熵波动来稳定训练动态;(2)促进在困难任务中主动进行状态空间探索;(3)确保环境引导内化到策略权重中,而非仅仅作为推理时的先验;(4)识别出组内反馈一致性是稳定优化的关键边界条件。
cs.LG / 214 / 2609.08412

Stochastically Perturbed Weights: Ensembles from Deterministic Machine-Learning Weather Models

随机扰动权重:基于确定性机器学习天气模式的集合预报
Adamov, Simon, Fuhrer, Oliver, Knutti, Reto, Schemm, Sebastian
Abstract
Machine-learning weather models (MLWMs) now match or outperform operational numerical weather prediction (NWP) at global medium-range forecasting, at far lower inference cost. Many deployed MLWMs are deterministic, producing a single forecast with no estimate of its own uncertainty, whereas a growing family of trained-probabilistic models generate calibrated ensembles directly, at the price of a dedicated training run. We ask instead how much uncertainty can be extracted from a deterministic checkpoint that already exists, without retraining it. Where physical ensembles represent model uncertainty by stochastically perturbing parametrisation tendencies, we perturb the network's raw weight tensors at inference time, a scheme we call stochastically perturbed weights (SPW). We also ask whether it works, where and on which scales to inject the noise, and where it fails. A three-phase ablation across four deterministic backbones, Aurora, GraphCast, SFNO, and AIFS, selects one production baseline per model, benchmarked against the trained-probabilistic AIFS-ENS, FourCastNet 3 and Atlas as well as the operational ECMWF ensemble (IFS-ENS) over 112 initialisation times. At a 240 h (10-day) lead time the SPW ensembles reach continuous ranked probability skill scores (CRPSS) between 0.04 and 0.13 below the best trained-probabilistic baseline, at zero marginal training cost. No injection site works across models: the productive tensor group is architecture-specific, so SPW is at present a tuning procedure rather than a plug-and-play recipe. Its main failure mode is a coherent whole-field offset that overdisperses the domain mean, and restricting the noise to coarse scales or perturbing the initial conditions each repair part of it.
Chinese Translation
机器学习天气模式(MLWMs)如今在全球中期预报中已达到甚至超越业务化数值天气预报(NWP)的水平,且推理成本远低于后者。许多已部署的MLWMs是确定性的,仅产生单一预报而不提供自身不确定性的估计;而日益增多的一类经过训练的概率模式则能直接生成经过校准的集合预报,但代价是需要专门的训练过程。我们转而探究:在不重新训练的前提下,能从一个已有的确定性模型检查点中提取多少不确定性。物理集合预报通过随机扰动参数化倾向来表征模式不确定性,与之类似,我们在推理时对网络的原始权重张量进行扰动,我们将该方案称为随机扰动权重(Stochastically Perturbed Weights, SPW)。我们还研究了该方案是否有效、应在何处以及哪个尺度上注入噪声、以及它在哪里失效。我们在四个确定性骨干模式——Aurora、GraphCast、SFNO和AIFS——上进行了三阶段消融实验,为每个模式选出一个生产基线,并与经过训练的概率模式AIFS-ENS、FourCastNet 3、Atlas以及业务化的ECMWF集合预报(IFS-ENS)在112个初始时刻上进行基准比较。在240小时(10天)预报时效上,SPW集合的连续排序概率技巧评分(CRPSS)比最优的训练概率基线低0.04至0.13,且边际训练成本为零。没有一种注入位置对所有模式都有效:有效的张量组依赖于具体架构,因此SPW目前是一种调参流程,而非即插即用的方案。其主要失效模式是一种一致的全场偏移,导致区域平均场过度离散;而将噪声限制在粗尺度上或扰动初始条件,均可部分修复这一问题。
cs.LG / 215 / 2609.08445

Topological Fraud Detection in Latent Transaction Spaces

潜在交易空间中的拓扑欺诈检测
Bourla, Avraham
Abstract
Working entirely on topologically anonymized embeddings, we perform fraud detection using iterative rounds of unsupervised filtering followed by supervised sniping. The result is an ultra-low latency privacy--preserving triage that allows institutions to flag suspicious activity without compromising Personally Identifiable Information.
Chinese Translation
我们完全基于拓扑匿名化的嵌入表示开展工作,通过无监督过滤与有监督狙击相结合的迭代轮次进行欺诈检测。其结果是一种超低延迟的隐私保护式分诊机制,使机构能够在不泄露个人身份信息(Personally Identifiable Information)的前提下标记可疑活动。
cs.LG / 216 / 2609.08554

Not All Variables Agree: Reliability-Aware Variable-Wise Gradient Surgery for Multivariate Time-Series Forecasting

并非所有变量都一致:面向多变量时间序列预测的可靠性感知的变量级梯度手术
Park, Jinwoo, Kang, Hyeongwon, Kang, Pilsung
Abstract
In data-driven training, multivariate time-series forecasting is usually optimized with a scalar loss averaged over samples, variables, and horizons. This averaging is convenient, but the optimizer sees only the aggregated gradient, which does not reveal whether the variable-wise contributions align or oppose one another. To quantify how often this disagreement arises, we measure the variable-wise gradients directly and find that 30.6% of their pairwise cosine similarities are negative on average across seven datasets. However, conflict and harm are not the same thing. Under shared training 35 of the 64 variables do worse than a full-input single-target oracle, and the harmed fraction is not reliably predicted by how often gradients conflict. We propose Per-Variable Surgery (PV-Surgery), an optimizer-side training strategy for backbones with cache-compatible layers. One backward pass builds variable-wise gradient proxies from output-side signals and keeps the pointwise forecasting loss. Reliability-aware selection targets layers whose proxy sums closely approximate their shared-gradient slices. Conditional pooling forms anchor and conflict pools without dropping variables. Common-direction surgery aligns variable or pooled gradients with their normalized mean and restores input norms to avoid reweighting. In experiments across five backbones, seven datasets, and four horizons, PV-Surgery lowers MSE by 3.61% and MAE by 2.93% on average. For multivariate forecasting, this indicates that the variable-wise structure hidden by mean-loss training is a usable optimization signal.
Chinese Translation
在数据驱动的训练中,多变量时间序列预测通常采用在样本、变量和预测时域上取平均的标量损失进行优化。这种平均化虽然方便,但优化器只能看到聚合后的梯度,无法揭示各变量的梯度贡献究竟是相互一致还是相互对立。为了量化这种不一致出现的频率,我们直接测量变量级梯度,发现在七个数据集上平均有30.6%的两两余弦相似度为负值。然而,冲突与损害并不等同。在共享训练下,64个变量中有35个的表现劣于全输入单目标(full-input single-target)的oracle,且受损害的变量比例无法通过梯度冲突的频率可靠预测。我们提出Per-Variable Surgery(PV-Surgery),这是一种面向具有缓存兼容层(cache-compatible layers)骨干网络的优化器侧训练策略。单次反向传播即可从输出侧信号构建变量级梯度代理,同时保留逐点的预测损失。可靠性感知的选择机制针对那些代理之和能紧密逼近其共享梯度分片的层。条件化池化在不丢弃任何变量的情况下构建锚定池(anchor pool)与冲突池(conflict pool)。共同方向手术(Common-direction surgery)将变量梯度或池化梯度与其归一化均值对齐,并恢复输入范数以避免重新加权。在五个骨干网络、七个数据集和四个预测时域的实验中,PV-Surgery平均使MSE降低3.61%、MAE降低2.93%。对于多变量预测而言,这表明被均值损失训练所掩盖的变量级结构是一种可利用的优化信号。
cs.LG / 217 / 2609.08561

Certified Topological Interaction in Neural Representations: Class Disentanglement Is Mostly Pairwise

神经表征中的认证拓扑交互:类解耦大多呈成对形式
Majhi, Sushovan
Abstract
Class disentanglement (the separation of a representation's class-conditional point clouds along depth and over training) is usually read off descriptive curves. We measure it as certified topological interaction between labeled point clouds, using the recently introduced Intersection Euler Characteristic Profile: the Euler characteristic of the overlap of the clouds' ball unions as a function of scale, computed by one Alpha-complex sweep with no boundary-matrix reduction. Every number carries a test: exact permutation tests in both directions, a guarded separation certificate, and a paired test for the comparative claims applications make. Across 111 trained networks and 52,650 certified measurements, disentanglement is depth-graded and concentrated in the first epochs, and interaction quotients rank class pairs by confusability (Spearman rho=0.83), on par with cheap separability statistics. In a 96-model factorial population, augmentation is the one training choice that separates classes relative to chance; weight decay compresses the overlap without separating, and depth and width do nothing. The structural finding is one only a k-fold statistic can pose: the joint entanglement of a class triple sits below that of its strongest pair in 97% of triple-layer cells and 99.5% of deep cells, far below a measured null floor, in vision encoders and frozen language models alike. This pairwise dominance is a regularity, not a law: expected from the nesting of overlaps but not forced by geometry, present at initialization and in raw pixels, and manufactured in the last stage alone when a network memorizes random labels. The unnormalized profile mass predicts test accuracy (R^2=0.94), the quotient does not, and neither beats a linear probe. One lesson is reported in full: the paired test must use a scale-free statistic, or it certifies feature-norm dynamics as disentanglement.
Chinese Translation
类解耦(即表征的类条件点云随深度和在训练过程中的分离程度)通常通过描述性曲线来读取。我们将其度量为带标签点云之间的认证拓扑交互,采用最近提出的交并欧拉示性数轮廓(Intersection Euler Characteristic Profile):即点云球体并集的交集的欧拉示性数随尺度的变化,通过一次Alpha复形扫描计算,无需边界矩阵约化。每个数值都附带检验:双向的精确置换检验、受保护的分离证书,以及应用中的比较性论断所需的配对检验。在111个训练网络和52,650次认证测量中,解耦随深度分级并集中于最初的若干训练轮次;交互商按类的可混淆程度对类对进行排序(Spearman rho=0.83),与廉价的分离性统计量相当。在一个96个模型的因子化群体中,数据增强是唯一能相对随机水平分离类别的训练选择;权重衰减在不分离的情况下压缩重叠,而深度和宽度则没有任何作用。结构性发现只有一个k重统计量才能提出:在视觉编码器和冻结语言模型中,三类联合纠缠低于其最强类对的纠缠,在三层单元格中占97%,在深层单元格中占99.5%,远低于实测的零假设下限。这种成对主导性是一种规律而非定律:它可以由重叠的嵌套性所预期,但并非由几何所强制;它存在于初始化时和原始像素中,并且当网络记忆随机标签时,仅在最后阶段被制造出来。未归一化的轮廓质量能预测测试准确率(R^2=0.94),而交互商则不能,且二者均不及线性探针。我们完整报告的一个教训是:配对检验必须使用无量纲(尺度无关)统计量,否则它会把特征范数的动态认证为解耦。
cs.LG / 218 / 2609.08581

AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery

AlphaRJM:用于随机收益引导的Alpha发现的奖励跳跃记忆
Dhan, Sayan, Natarajan, Selvaraju
Abstract
Formulaic alpha discovery is a pool-dependent symbolic search problem in which informative feedback is observed primarily when a complete expression is evaluated. This delayed feedback creates two coupled difficulties: the retained alpha pool does not preserve the full history of realized evaluation feedback, and the value of an intermediate construction action is uncertain because its consequence depends on the formula eventually completed. We introduce AlphaRJM, which addresses these difficulties through Reward-Jump Memory, an event-driven latent state that remains fixed during token construction and updates only at terminal evaluation events using the realized pool reward and evaluation outcome, and an action-conditioned SDE return critic that represents future discounted discovery returns with stochastic particles. The particles guide action selection through their mean and uncertainty and are learned using a distributional Bellman objective combining energy-distance matching, mean calibration, and jump regularization. Empirically, AlphaRJM delivers strong and stable gains across multiple equity universes, forecasting horizons, and random seeds, while ablations confirm the complementary roles of persistent evaluation history, stochastic return modeling, and distributional supervision.
Chinese Translation
公式化Alpha发现是一个依赖于因子池的符号搜索问题,其信息性反馈主要在完整表达式被评估时才能观测到。这种延迟反馈带来了两个相互耦合的困难:保留的Alpha池无法保存已实现评估反馈的完整历史;同时,由于中间构建动作的后果取决于最终完成的公式,其价值是不确定的。我们提出AlphaRJM,通过两项机制解决这些困难:其一是奖励跳跃记忆,这是一种事件驱动的潜在状态,在令牌构建过程中保持不变,仅在终止评估事件时利用已实现的池收益和评估结果进行更新;其二是动作条件化的SDE收益评论家,它以随机粒子形式表示未来的折扣发现收益。这些粒子通过其均值和不确定性引导动作选择,并通过结合能量距离匹配、均值校准和跳跃正则化的分布式Bellman目标进行学习。实验表明,AlphaRJM在多个股票池、预测周期和随机种子上均取得了强劲且稳定的提升,消融实验进一步验证了持久化评估历史、随机收益建模和分布式监督的互补作用。
cs.LG / 219 / 2609.08582

Leveraging Cardiac Imaging to Improve ECG-Based Detection of Chagas Disease in Resource-Constrained Settings

利用心脏影像改善资源受限环境中基于心电图的恰加斯病检测
Alvarez-Florez, Laura, Uyterlinde, Daniel, Ruipérez-Campillo, Samuel, Arts, Lukas P. A., Asselbergs, Folkert W., Tjong, Fleur V. Y.
Abstract
Chagas disease is a major cause of cardiomyopathy in Latin America. Cardiac magnetic resonance (CMR) imaging can characterize its structural abnormalities, but scanners and expert readers remain scarce in endemic regions. Electrocardiography (ECG) is inexpensive and widely available, yet structural disease must be inferred indirectly from electrical signals. We propose to transfer CMR-derived structural knowledge to ECG through contrastive pre-training. Using 63,193 paired ECG-CMR examinations from the UK Biobank, we align an ECG encoder with a clinically grounded CMR embedding space using an asymmetric InfoNCE objective. Despite seeing no Chagas cases during pre-training, the resulting representation improves ECG-based Chagas detection. Across CODE-15% and SaMi-Trop, a frozen linear probe achieves an AUROC of 0.851 and sensitivity at the top 5% of predicted risk (Top5%-TPR) of 0.427 in five-fold cross-validation, compared with 0.827 and 0.377 for an unaligned ECG-FM baseline. On the PhysioNet/CinC 2025 Challenge test set, our model obtains the highest AUROC on SaMi-Trop-3 and the best ELSA-Brasil challenge score among the three top-performing methods, indicating that imaging-supervised ECG representations can generalize to populations and settings beyond the pre-training distribution.
Chinese Translation
恰加斯病(Chagas disease)是拉丁美洲心肌病的主要病因。心脏磁共振(CMR)成像能够表征其结构性异常,但在流行地区,扫描设备和专业读片人员仍然稀缺。心电图(ECG)价格低廉且普及广泛,然而结构性病变只能通过电信号间接推断。我们提出通过对比预训练,将CMR衍生的结构知识迁移到心电图。我们利用英国生物样本库(UK Biobank)中63,193对ECG-CMR配对检查数据,采用非对称InfoNCE目标函数,将ECG编码器与基于临床的CMR嵌入空间进行对齐。尽管预训练期间未接触任何恰加斯病例,所得表征仍提升了基于ECG的恰加斯病检测性能。在CODE-15%和SaMi-Trop数据集上,冻结的线性探针在五折交叉验证中达到0.851的AUROC和0.427的预测风险前5%灵敏度(Top5%-TPR),而未对齐的ECG-FM基线分别为0.827和0.377。在PhysioNet/CinC 2025挑战赛测试集上,我们的模型在SaMi-Trop-3上取得了最高的AUROC,并在三种表现最佳的方法中获得了最好的ELSA-Brasil挑战赛得分,表明由影像监督的ECG表征能够泛化到预训练分布之外的人群和环境。
cs.LG / 220 / 2609.08594

Multi-Level-Set-Based Physics-Driven Neural Network to Solve 3-D Inverse Scattering Problems

基于多水平集的物理驱动神经网络求解三维逆散射问题
Du, Yutong, Liu, Zicheng, Qi, Bo, Zong, Yali, Han, Peixian
Abstract
This paper proposes a level-set-based physics-driven neural network solver (LSPDNN) for 3-D electromagnetic inverse scattering. To mitigate boundary blurring and reconstruction artifacts in voxel-wise contrast reconstruction, the proposed solver exploits the piecewise homogeneity of practical scatterers by representing unknown targets with multiple coordinate-dependent neural level-set components. Specifically, a soft-union multi-material model is proposed to separately describe the object support and material distribution. The global support is formed by the union of multiple level-set components, while the local contrast is determined by normalized component weights and learnable complex permittivity candidates. In addition, a model-consistent total variation (TV) regularization is imposed on the material-region indicators, rather than directly on the reconstructed contrast, to suppress fragmented material assignments without excessively smoothing material interfaces. An adaptive loss balancing strategy is further introduced to reduce the dependence on manually selected regularization weights. For each measurement instance, the neural level-set parameters and material candidates are optimized by minimizing a physics-consistent objective function. Numerical and experimental results demonstrate that LSPDNN can reconstruct scatterers with clear boundaries, more uniform material regions, and substantially reduced background artifacts. The results highlight the advantage of the neural level-set parameterization in challenging 3-D inverse scattering cases involving irregular shapes, closely spaced objects, multiple materials, and measurement noise.
Chinese Translation
本文提出了一种基于水平集的物理驱动神经网络求解器(LSPDNN),用于三维电磁逆散射问题。为缓解逐像素对比度重建中的边界模糊和重建伪影,所提出的求解器利用实际散射体的分段均匀性,采用多个坐标相关的神经水平集分量来表示未知目标。具体而言,提出了一种软并集多材料模型,分别描述目标支撑域和材料分布。全局支撑域由多个水平集分量的并集构成,而局部对比度由归一化的分量权重和可学习的复介电常数候选值确定。此外,在材料区域指示子上施加模型一致的总变分(TV)正则化,而非直接作用于重建的对比度,以抑制碎片化的材料分配,同时避免对材料界面的过度平滑。进一步引入自适应损失平衡策略,以减少对手动选取正则化权重的依赖。对于每个测量实例,通过最小化物理一致的目标函数来优化神经水平集参数和材料候选值。数值与实验结果表明,LSPDNN能够重建出边界清晰、材料区域更均匀且背景伪影显著减少的散射体。结果凸显了神经水平集参数化在涉及不规则形状、近距离目标、多材料以及测量噪声等具有挑战性的三维逆散射场景中的优势。
cs.LG / 221 / 2609.08615

Why shared attention vectors fail: a case for outcome-indexed tuning

共享注意向量为何失效:一种基于结果索引的调谐方案
Dome, Lenard
Abstract
Dimensional attention in learning is often implemented as a globally shared attention vector, where each stimulus dimension corresponds to a single scalar. These scalars are learned by models through gradient-descent on error, where predictive features acquire more salience. We show that under multi-outcome learning, where models predict more than one outcome, this shared vector becomes unstable; it collapses to its bounds and prevents the models from learning meaningful attentional tunings for learning and generalization. We address this by introducing an outcome-indexed attentional matrix that converts globally shared attentional tuning into an outcome-indexed representation. We present an analysis of the unstable shared vectors and derive the conditions under which it holds. Empirically, three synthetic experiments benchmark the proposed attention matrices and show that they converge to meaningful representations, something shared attention vectors fail to do. These results suggest that outcome-indexed attentional matrices are a general fix for gradient-based attentional processes, which improves models of learning under multi-outcome conditions.
Chinese Translation
学习中的维度注意通常被实现为全局共享的注意向量,其中每个刺激维度对应单个标量。这些标量由模型通过基于误差的梯度下降学习得到,具有预测性的特征会获得更高的显著性。我们证明,在多结果学习(即模型需预测多个结果)的情形下,这种共享向量会变得不稳定:它会塌缩到其边界值,从而阻碍模型学习到有意义的注意调谐,影响学习与泛化。为解决这一问题,我们引入了一种基于结果索引的注意矩阵,将全局共享的注意调谐转换为按结果索引的表示。我们对共享向量的不稳定现象进行了分析,并推导出该现象成立的条件。在实证方面,我们通过三个合成实验对所提出的注意矩阵进行基准测试,结果表明它们能够收敛到有意义的表示,而共享注意向量则无法做到。这些结果表明,基于结果索引的注意矩阵是基于梯度的注意过程的一种通用修复方法,可改进多结果条件下的学习模型。
cs.LG / 222 / 2609.08618

Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families

面向跨语言模型家族训练响应预测的与目标无关的微干预方法
Liu, Zhongxuan, Zhou, Sicheng, Wang, Hongzhi
Abstract
Benchmark scores describe what a checkpoint can do now, but they do not determine how it will respond to the next training episode. We measure this missing state by branching four short, standardized, target-independent micro-interventions from the same checkpoint and recording their effects in a common capability space. Together with current capability, these responses form L-State; its pulse block supports a flexible direct readout and a structure-preserving operator readout. Under smooth local dynamics, the operator construction admits an end-to-end cross-family bound with explicit source- and target-family coordinate heterogeneity. In three-family leave-one-family-out development, both pulse readouts reduce source-standardized MSE by 39.4% relative to capability alone, while separating the best response and direction estimates. On sealed GLM-4-9B, the direct and operator readouts reduce MSE by 71.8% and 78.3%, respectively, and the operator readout raises sign balanced accuracy from 0.366 to 0.754. On sealed Granite-3.1-8B, the direct readout reaches RMSE 0.544 and a development-fitted action-wise selector reaches 0.554, compared with 1.172 for capability alone. A five-family audit finds that the operator coordinate varies by action and family, and that modeling these deviations improves retrospective held-trajectory prediction. Target-independent interventions therefore expose training-response information that current capability misses, with direct and structured readouts covering complementary transfer regimes.
Chinese Translation
基准测试分数描述了模型检查点(checkpoint)当前所能做到的事情,但并不能决定它在下一次训练阶段将如何响应。我们通过从同一检查点分支出四种简短、标准化且与目标无关的微干预(micro-interventions),并在统一的能力空间中记录其效果,来度量这一缺失的状态。这些响应与当前能力一起构成了 L-State;其脉冲(pulse)模块支持灵活的直接读出(direct readout)和保持结构的算子读出(operator readout)。在平滑局部动力学假设下,该算子构造给出了一个端到端的跨家族误差界,其中显式包含源家族与目标家族的坐标异质性。在三家族留一家族的开发实验中,两种脉冲读出方法相比仅使用能力特征,将源标准化 MSE 降低了 39.4%,同时将最优响应估计与方向估计有效分离。在密封测试的 GLM-4-9B 上,直接读出与算子读出分别将 MSE 降低 71.8% 和 78.3%,且算子读出将符号平衡准确率从 0.366 提升至 0.754。在密封测试的 Granite-3.1-8B 上,直接读出达到 RMSE 0.544,经开发数据拟合的按动作选择器达到 0.554,而仅使用能力特征的基线为 1.172。一项五家族审计发现,算子坐标随动作和家族而变化,而对这些偏差进行建模可以改善回顾性的留轨迹预测。因此,与目标无关的干预能够暴露当前能力所无法反映的训练响应信息,直接读出与结构化读出则覆盖了互补的迁移场景。
cs.LG / 223 / 2609.08622

Leveraging contextual events on structure-aware next activity prediction

在结构感知的下一活动预测中利用情境事件
Mele, Alessandro, Diamantini, Claudia, Potena, Domenico
Abstract
Predictive process monitoring aims at forecasting various aspects of running processes. Among the different tasks, next activity prediction represents the most extensively investigated. However, only a limited number of existing approaches explicitly encode contextual information, i.e., the environmental conditions in which the process is executed, typically modeled through event log attributes or aggregated measures. In this paper, an approach based on the concept of Instance Graphs is introduced. To incorporate contextual process instances, several encoding strategies are proposed and evaluated by measuring their impact on prediction performance. For each encoding strategy, a set of prefix-Instance Graphs is generated and subsequently provided as input to a Graph Neural Network for the classification task. The proposed approach is evaluated on multiple real-world event logs, and the experimental results demonstrate that incorporating contextual process instances benefits prediction performance.
Chinese Translation
预测性流程监控旨在预测运行中流程的各个方面。在众多任务中,下一活动预测是被研究最为广泛的一项。然而,现有方法中仅有少数显式地编码了情境信息,即流程执行时所处的环境条件,这类信息通常通过事件日志属性或聚合度量来建模。本文提出了一种基于实例图(Instance Graphs)概念的方法。为了融入情境化的流程实例,我们提出了多种编码策略,并通过衡量其对预测性能的影响对其进行评估。对于每种编码策略,生成一组前缀实例图(prefix-Instance Graphs),随后将其作为输入提供给图神经网络(Graph Neural Network)以执行分类任务。我们在多个真实世界的事件日志上对该方法进行了评估,实验结果表明,融入情境化的流程实例有助于提升预测性能。
cs.LG / 224 / 2609.08634

Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

Suan:修正大语言模型中的直接偏好安全对齐
Cherednichenko, Oleksandr, Klypa, Roman
Abstract
Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. While proprietary systems exhibit reliable safety controls, their underlying methodologies and trade-offs remain largely undisclosed. Achieving comparable security in open-weight models remains a persistent challenge, as post-trained variants frequently suffer from over-refusal and degraded general quality. To overcome these drawbacks, we introduce Suan, a novel preference optimization algorithm. Unlike existing methods, we formulate the optimization objective directly at the gradient level, bypassing the standard variational derivation. As a result, we obtain more interpretable and robust training dynamics. Extensive evaluations across a diverse suite of competitive baselines and benchmarks demonstrate that Suan achieves superior safety alignment while fully preserving response utility.
Chinese Translation
在大语言模型(LLM)中集成可靠的安全防护机制,对于提供既有帮助又无害的回答至关重要。尽管专有系统展现出可靠的安全控制能力,但其底层方法和权衡机制仍大多未公开。要在开放权重模型中实现相当的安全性,仍然是一个持续的挑战,因为经后训练的变体经常出现过度拒答和整体质量下降的问题。为克服这些缺陷,我们提出了 Suan,一种新颖的偏好优化算法。与现有方法不同,我们直接在梯度层面构建优化目标,绕过了标准的变分推导。由此,我们获得了更具可解释性和更稳健的训练动态。在多种竞争性基线模型和基准测试上的广泛评估表明,Suan 在完全保留回答实用性的同时,实现了更优的安全对齐效果。
cs.LG / 225 / 2609.08642

SUN: Reaching for Novelty in Reinforcement Learning

SUN:在强化学习中追求新颖性
Yang, Wenyan, Mustafin, Arsenii, Baumann, Dominik, Pajarinen, Joni, Parisi, Simone
Abstract
Exploration in reinforcement learning (RL) remains a fundamental challenge. Recent goal-conditioned RL strategies (which select goals to encourage broader state coverage) have shown promising results, but none scores a goal by novelty and reachability jointly: the two signals are traded off by hand, applied in sequence, or one is neglected outright. In this paper, we introduce a reachability-aware goal-selection framework that explicitly integrates these two aspects, and that can be seamlessly incorporated into any off-policy RL algorithm. To this aim, we propose SUccessor-to-Novelty (SUN), an indicator derived from successor value functions to identify goals that are both novel and reachable. We prove that SUN recovers count-based bonuses in the limit, bounds short-horizon hitting probabilities, and provably rejects unreachable goals. We further present an adaptive goal-selection strategy that leverages these properties, and an accurate yet lightweight pseudocount to avoid the overhead of classic methods. We back up all our claims with thorough benchmarks: SUN consistently outperforms state-of-the-art methods in standard and novel environments with unreachable or hard-to-reach states, irreversible transitions, obstacles, mazes, and unbounded spaces.
Chinese Translation
强化学习(RL)中的探索仍然是一个根本性挑战。近期的目标条件强化学习策略(通过选择目标来鼓励更广泛的状态覆盖)已展现出良好的效果,但目前没有方法能同时依据新颖性和可达性对目标进行评分:这两个信号要么通过手动权衡,要么按顺序先后应用,抑或完全忽略其中之一。在本文中,我们提出了一个可达性感知的目标选择框架,显式地整合这两个方面,并且可以无缝嵌入任何离线策略(off-policy)强化学习算法中。为此,我们提出了 SUccessor-to-Novelty(SUN),这是一个由后继价值函数(successor value functions)导出的指标,用于识别既新颖又可达的目标。我们证明 SUN 在极限情况下可以恢复基于计数的奖励(count-based bonuses),能够界定短时程命中概率,并且可证明地拒绝不可达目标。我们进一步提出了一种利用这些性质的自适应目标选择策略,以及一种精确而轻量的伪计数(pseudocount)方法,以避免经典方法带来的计算开销。我们通过全面的基准实验支持了所有论断:在标准环境以及包含不可达或难以到达状态、不可逆转移、障碍物、迷宫和无界空间的新颖环境中,SUN 始终优于最先进的方法。
cs.LG / 226 / 2609.08650

Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR

面向RLVR推理覆盖率扩展的难度自适应树结构策略优化
Yu, Youngjun, Jang, Sanghwan, Yu, Hwanjo
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the model's intrinsic reasoning coverage (pass@k) due to limited exploration during training. To address this, we optimize the structural design of train-time rollouts to enhance pass@k. Our analysis identifies three key design principles: (1) difficulty-adaptive rollout can play an important role in expanding pass@k, beyond serving as an efficiency heuristic; (2) tree-based rollout outperforms parallel sampling in discovering correct answers; and (3) sentence-entropy-guided forking overcomes the localization phenomenon of token-level branching to maximize semantic diversity. Building on these insights, we propose DATPO (Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization). DATPO integrates difficulty-adaptive tree search with a sibling-diversity advantage term, explicitly promoting semantic diversity to expand reasoning coverage during training. Experiments on mathematical reasoning benchmarks demonstrate that DATPO outperforms baselines especially in pass@k, which directly translates to superior test-time scaling performance.
Chinese Translation
可验证奖励强化学习(Reinforcement Learning with Verifiable Rewards, RLVR)是近年来大型推理模型取得成功的关键技术。然而,尽管RLVR显著提升了单样本准确率,但由于训练过程中探索能力有限,它往往无法扩展模型的内在推理覆盖率(pass@k)。为解决这一问题,我们优化了训练时 rollout 的结构设计以提升 pass@k。我们的分析确立了三个关键设计原则:(1)难度自适应 rollout 在扩展 pass@k 方面可以发挥重要作用,而不仅仅是作为一种效率启发式方法;(2)基于树的 rollout 在发现正确答案方面优于并行采样;(3)基于句子熵引导的分叉(forking)克服了 token 级分支的局部化现象,从而最大化语义多样性。基于这些洞见,我们提出了 DATPO(Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization,难度自适应句子熵引导树结构策略优化)。DATPO 将难度自适应树搜索与兄弟多样性优势项相结合,显式地促进语义多样性,以在训练过程中扩展推理覆盖率。在数学推理基准上的实验表明,DATPO 优于基线方法,尤其是在 pass@k 指标上,这直接转化为更优的测试时扩展(test-time scaling)性能。
cs.LG / 227 / 2609.08663

MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models

MoEMB:基于高效混合专家模型的通用多模态嵌入扩展方法
Cui, Xuanming, Mishra, Shlok Kumar, Bao, Wentao, Singh, Aashu, Wang, Zihao, Fan, Xiangjun, Xiao, Jun, Lim, Ser-Nam, Cheng, Jianpeng
Abstract
Universal multimodal embedding (UME) increasingly demands encoder's capacity for handling a broad range of tasks and modalities with increased complexity. Prior scaling methods either increase the representation size, retrieval effort, or scales the encoder into a heavy multimodal LLM. Recent works, such as Think-Then-Embed (TTE), explore scaling via reasoning tokens. However, embedding models are hard to scale up: increasing parameters directly tradeoffs for the large training batch size that contrastive learning needs, and retrieval has to be served under tight latency. Moreover, UME tasks are diverse in complexity, where scaling up embedders can bring significant redundant computation. In this work, we propose MOEMB, which instead scales UME along the expert axis through mixture-of-experts (MoE), growing encoder capacity while preserving single-vector, non-autoregressive encoding. Through a systematic study of the design space and training recipes for MoE-based UME, MoEMB sets a new state of the art on both MMEB-V2 and MRMR among models trained on public MMEB-family data: with only 3B active parameters, MoEMB surpasses TTE-based methods with >4x active parameters, using significantly less computes. To further improve the scalability and efficiency, we conduct the first comprehensive study of adaptive computation for MoE-based embedding, spanning diverse strategies across training-based and inference-only methods. Together, these results support expert scaling as an effective and efficient direction for UME, with adaptive computation further improving efficiency for MLLM-based embedding models towards large-scale retrieval and recommendation systems.
Chinese Translation
通用多模态嵌入(Universal Multimodal Embedding, UME)日益要求编码器具备处理范围广泛、复杂度不断提升的任务与模态的能力。已有的扩展方法要么增加表示维度、提升检索开销,要么将编码器扩展为庞大的多模态大语言模型(LLM)。近期工作如 Think-Then-Embed(TTE)则通过推理词元探索扩展途径。然而,嵌入模型难以扩展:增加参数量会与对比学习所需的大训练批量直接产生权衡,且检索服务必须满足严格的延迟要求。此外,UME 任务在复杂度上差异巨大,对嵌入模型进行扩展会带来大量冗余计算。本文提出 MoEMB,通过混合专家(Mixture-of-Experts, MoE)沿专家维度扩展 UME,在保持单向量、非自回归编码的同时提升编码器容量。通过对基于 MoE 的 UME 的设计空间与训练方案的系统性研究,MoEMB 在使用公开 MMEB 系列数据训练的模型中,于 MMEB-V2 和 MRMR 两个基准上均刷新了当前最优性能:仅使用 30 亿(3B)激活参数,MoEMB 便超越了激活参数超过其 4 倍的基于 TTE 的方法,且计算量显著更低。为进一步提升可扩展性与效率,我们首次对基于 MoE 的嵌入的自适应计算进行了全面研究,涵盖基于训练与仅推理阶段的多种策略。这些结果共同表明,专家扩展是 UME 的一个高效且有效的方向,而自适应计算可进一步提升基于多模态大语言模型(MLLM)的嵌入模型的效率,使其更好地服务于大规模检索与推荐系统。
cs.LG / 228 / 2609.08683

Neither Adversarial Training Nor Purification: Emergent Adversarial Robustness from Oscillatory Predictive Learning

无需对抗训练亦无需净化:由振荡预测学习涌现的对抗鲁棒性
Habibi, Mohammed-Yassine, Ziu, Klea, Takáč, Martin, Yamada, Makoto
Abstract
Adversarial robustness in computer vision is still largely achieved through adversarial training or test-time adversarial purification, both of which introduce significant computational overhead by generating adversarial examples during training or performing iterative denoising at test time. We study whether empirical robustness can instead emerge from architectural and representation-learning inductive biases. We introduce Oscillatory Predictive Learning (OPL), a two-stage framework that combines Artificial Kuramoto Oscillatory Neurons (AKOrN) with predictive self-supervised pretraining using X-PhiNet. Because our default checkpoint uses randomized initial oscillator states, we compare it with other randomized adversarial defense methods that provide precise, reproducible, and strong attack protocols. Experiments on CIFAR-10 and CIFAR-100, with additional corruption evaluation on CIFAR-10-C, demonstrate that our method achieves competitive results under the AutoAttack-rand evaluation protocol. On CIFAR-10 and CIFAR-100, OPL attains 76.63$\pm$0.76$\%$ and 50.44$\%$ robust accuracy, respectively, under $\ell_\infty$, $\epsilon=8/255$, AutoAttack-rand with EoT $K=20$.
Chinese Translation
计算机视觉中的对抗鲁棒性目前主要通过对抗训练或测试时对抗净化来实现,但这两者都需要在训练期间生成对抗样本或在测试时执行迭代去噪,从而带来显著的计算开销。我们研究了经验鲁棒性能否转而源于架构和表示学习的归纳偏置。我们提出了振荡预测学习(Oscillatory Predictive Learning, OPL),这是一个两阶段框架,将人工Kuramoto振荡神经元(AKOrN)与基于X-PhiNet的预测式自监督预训练相结合。由于我们的默认检查点使用随机化的初始振荡器状态,我们将其与其他提供精确、可复现且强大攻击协议的随机化对抗防御方法进行比较。在CIFAR-10和CIFAR-100上的实验,以及在CIFAR-10-C上的额外损坏鲁棒性评估表明,我们的方法在AutoAttack-rand评估协议下取得了具有竞争力的结果。在$\ell_\infty$扰动、$\epsilon=8/255$、EoT $K=20$的AutoAttack-rand协议下,OPL在CIFAR-10和CIFAR-100上分别取得了76.63$\pm$0.76$\%$和50.44$\%$的鲁棒准确率。
cs.LG / 229 / 2609.08685

HOPE: Heterophily-Aware Open-Set Node Classification with Pseudo-Extrapolation

HOPE:基于伪外推的异配感知开放集节点分类
Dai, Yumeng, Tan, Yue, Liu, Yixin, Wang, Chenxu, Wang, Pinghui, Qin, Tao
Abstract
Standard open-set node classification methods rely on the homophily assumption, where connected nodes share labels. However, real-world graphs are often heterophilic, exposing the limitations of current methods and posing new challenges to open-set node classification. On the one hand, cross-class connectivity causes representations from different known or unknown classes to become intertwined after aggregation, undermining their discriminative capacity. On the other hand, structural mixture invalidates threshold-based open-set methods and cross-class feature interpolation, leading to unreliable unknown-class rejection. To address these challenges, we propose HOPE, a Heterophily-aware Open-set node classification method with Pseudo-Extrapolation. To adapt open-set graph neural networks (GNNs) to heterophilic scenarios, HOPE uses a structure-augmented feature initialization layer to capture multi-hop structural patterns. Meanwhile, we design a trustworthy neighborhood aggregation mechanism for standard GNNs to dynamically filter noisy cross-class neighbors. To enhance unknown-class rejection, we introduce a heterophily-guided pseudo-extrapolation strategy. It dynamically maintains known-class centers and extrapolates along cross-class neighborhood displacement directions, synthesizing pseudo-unknown proxies near structurally ambiguous regions. Finally, we optimize the network with joint classification and logit margin regularization, routing synthetic proxies into a dedicated rejection slot without imposing geometric margin constraints in the representation space. Extensive experiments on multiple datasets show that HOPE consistently outperforms state-of-the-art models, validating its effectiveness, robustness, and efficiency.
Chinese Translation
标准的开放集节点分类方法依赖于同配性假设,即相连的节点共享标签。然而,现实世界的图往往是异配的,这暴露了现有方法的局限性,并为开放集节点分类带来了新的挑战。一方面,跨类连接使得来自不同已知或未知类别的表示在聚合后相互交织,削弱了其判别能力。另一方面,结构混合使基于阈值的开放集方法和跨类特征插值失效,导致对未知类别的拒绝不可靠。为应对这些挑战,我们提出了HOPE,一种基于伪外推(Pseudo-Extrapolation)的异配感知开放集节点分类方法。为了使开放集图神经网络(GNN)适应异配场景,HOPE采用结构增强的特征初始化层来捕捉多跳结构模式。同时,我们为标准GNN设计了一种可信邻居聚合机制,以动态过滤带噪声的跨类邻居。为了增强对未知类别的拒绝能力,我们引入了一种异配引导的伪外推策略。该策略动态维护已知类中心,并沿跨类邻域位移方向进行外推,在结构模糊区域附近合成伪未知代理样本。最后,我们通过分类与logit边距正则化的联合优化来训练网络,将合成代理样本路由到专用的拒绝通道,而无需在表示空间中施加几何边距约束。在多个数据集上的大量实验表明,HOPE持续优于最先进的模型,验证了其有效性、鲁棒性和高效性。
cs.LG / 230 / 2609.08690

Hyperparameter Scaling Laws Across MoE Sparsity

MoE稀疏度下的超参数缩放定律
Tian, Changxin, Chen, Kunlong, Liu, Jia, Liu, Ziqi, Zhang, Zhiqiang, Zhou, Jun
Abstract
Mixture-of-Experts (MoE) models expand model capacity without a proportional increase in training compute, but increasing sparsity makes reliable hyperparameter transfer challenging. In this work, we show that conventional hyperparameter scaling laws are insufficient for ultra-sparse MoEs: the optimal learning rate and batch size vary with activation ratio, and these shifts cannot be explained by either total or activated parameter count alone. To characterize this dependence, we conduct 1,800 pre-training runs spanning six activated-parameter scales and models with up to 6B total non-embedding parameters, processing approximately 20 trillion tokens at a cost of 200,000 equivalent H800 GPU-hours. Our results reconcile conflicting findings in prior work by revealing two scaling regimes. At fixed sparsity, the optimal batch size follows a power-law relationship with training tokens $D$, whereas the optimal learning rate scales with training compute $C$ and remains robust to the allocation between model size and data. Across sparsity levels, the activation ratio $A$ enters both relationships as an additional multiplicative power-law factor. These observations lead to unified hyperparameter scaling laws that transfer across MoE sparsity levels. Large-scale evaluation shows that the scaling form outperforms alternative functional forms. On a held-out ultra-sparse MoE with 12B total parameters and only 1/64 of its experts activated, the predicted hyperparameters remain close to the observed optima, supporting joint extrapolation across model scale and sparsity. Further experiments demonstrate transfer across expert granularities and isolate the effect of activation ratio from that of total expert count.
Chinese Translation
混合专家模型在无需按比例增加训练计算量的情况下扩展了模型容量,但随着稀疏度的提高,可靠地迁移超参数变得更具挑战性。在这项工作中,我们表明传统的超参数缩放定律不足以描述超稀疏MoE:最优学习率和批次大小会随激活比例变化,而这些变化无法仅用总参数量或激活参数量来解释。为刻画这一依赖关系,我们进行了1,800次预训练实验,涵盖六种激活参数规模、总非嵌入参数量最高达6B的模型,处理了约20万亿词元,成本相当于200,000 H800 GPU小时。我们的结果揭示了两种缩放机制,从而调和了先前工作中相互矛盾的发现。在固定稀疏度下,最优批次大小与训练词元数$D$遵循幂律关系,而最优学习率随训练计算量$C$缩放,并对模型规模与数据之间的分配保持稳健。在不同稀疏度水平之间,激活比例$A$作为额外的乘性幂律因子进入上述两种关系。这些观测结果导出了统一的超参数缩放定律,可跨MoE稀疏度水平迁移。大规模评估表明,该缩放形式优于其他函数形式。在一个总参数量为12B、且仅有1/64专家被激活的留出超稀疏MoE上,预测的超参数仍接近观测到的最优值,支持跨模型规模和稀疏度的联合外推。进一步的实验证明了该定律在不同专家粒度之间的迁移性,并将激活比例的影响与专家总数的影响分离开来。
cs.LG / 231 / 2609.08709

Chimaera: A Mixture-of-Graph-Experts Architecture for Cross-Task and Cross-Dataset Graph Learning

Chimaera:一种用于跨任务与跨数据集图学习的图专家混合架构
Frank, Jonathan, Richerby, David, Scherp, Ansgar
Abstract
Designing foundation models for graphs is challenging due to the irregular structure of graphs and the different sizes and characteristics of embeddings. Chimaera integrates mixture-of-experts with graph foundation models (GFM). It integrates different GFM architectures, such as graph prompts and linear GNN models. Large language models are used to generate embeddings, and experts can be trained and combined following different strategies, GFMs, embeddings, etc. Furthermore, Chimaera extends existing linear GNNs to support link-level and graph-level tasks in addition to node-level tasks. Empirical analyses are performed on same-task and cross-task experiments with node, link, and graph classification tasks using six benchmark text-attributed graph datasets. The experiments demonstrate the effectiveness of Chimaera and its capabilities for transfer across tasks and datasets. Further insights include the need to use both large and small language models to generate embeddings for the experts, a strong cross-task transferability of simple but effective linear GNNs, and using few samples only to provide strong results.
Chinese Translation
由于图的不规则结构以及嵌入的尺寸和特性差异,为图设计基础模型极具挑战性。Chimaera 将专家混合机制与图基础模型(Graph Foundation Model, GFM)相结合,整合了不同的 GFM 架构,例如图提示和线性 GNN 模型。该方法利用大语言模型生成嵌入,并可依据不同策略(如不同的 GFM、嵌入方式等)对专家进行训练与组合。此外,Chimaera 扩展了现有的线性 GNN,使其在节点级任务之外还能支持链接级和图级任务。我们在节点、链接和图分类任务上,使用六个基准文本属性图数据集,开展了同任务和跨任务实验的实证分析。实验结果表明了 Chimaera 的有效性及其跨任务、跨数据集的迁移能力。进一步的洞察包括:需要同时使用大型和小型语言模型为专家生成嵌入;简单而有效的线性 GNN 具有很强的跨任务可迁移性;以及仅使用少量样本即可取得优异结果。
cs.LG / 232 / 2609.08725

BAFF: Bid-Aware Filter Family for Mitigating Training Data Interference in RTB A/B Tests

BAFF:用于缓解实时竞价 A/B 测试中训练数据干扰的竞价感知过滤器族
Oh, Jeonglyul, Choi, Ikkyu, Youn, Inseop, Kim, Youngjae
Abstract
In online A/B tests for real-time bidding (RTB), control and treatment models are typically trained on a shared serving log that includes data generated by the counterpart model. This shared-log training biases each model's training data through two channels: the counterpart model may have selected a different ad from the ad-candidate pool (ad-ranking disagreement) and may have bid a different price (bid-pricing disagreement), potentially distorting the A/B test outcome. Log-splitting eliminates the bias but sacrifices training data; log-sharing retains all data but leaves the bias unaddressed. We formalize the Bid-Aware Filter Family (BAFF), a class of (k,l)-parameterized hard filters that controls tolerance to each channel independently, providing a structured search space between these two extremes. We further propose a three-stage online measurement protocol that enables evaluating data-sharing strategies by their deviation from an interference-free reference model in production. In offline simulation, a (k,l) sweep surfaces operating points with smaller deviation from the interference-free reference model than both log-sharing and log-splitting. In a live RTB deployment on a demand-side platform (DSP), filter-based variants preserve the reference model's business metrics (e.g., CPC, CTR) more closely than both baselines. The best operating point is setting-dependent, underscoring the practical value of the search space itself.
Chinese Translation
在实时竞价(RTB)的在线 A/B 测试中,对照组和实验组模型通常在同一份共享服务日志上训练,而该日志中包含了由对方模型产生的数据。这种共享日志的训练通过两个渠道使各模型的训练数据产生偏差:对方模型可能从广告候选池中选择了不同的广告(广告排序不一致),也可能给出了不同的出价价格(出价定价不一致),从而可能扭曲 A/B 测试结果。日志切分可以消除偏差,但会损失训练数据;日志共享则保留全部数据,但偏差问题依然存在。我们形式化提出了竞价感知过滤器族(Bid-Aware Filter Family,BAFF),这是一类由 (k,l) 参数化的硬过滤器,可独立控制对每个渠道的容忍度,从而在这两个极端之间提供了一个结构化的搜索空间。我们进一步提出了一种三阶段在线测量协议,能够在生产环境中根据数据共享策略相对于无干扰参考模型的偏离程度对其进行评估。在离线模拟中,对 (k,l) 的扫描发现了一些工作点,其相对无干扰参考模型的偏离程度小于日志共享和日志切分两种方式。在需求方平台(DSP)上的真实 RTB 部署中,基于过滤器的变体比两种基线方法更好地保持了参考模型的业务指标(如 CPC、CTR)。最佳工作点依赖于具体场景设置,这也凸显了该搜索空间本身的实用价值。
cs.LG / 233 / 2609.08740

PAC-Bayesian Bounds for Learning Partially Observed Stochastic Linear Time-Invariant State-Space Systems with Inputs and Sub-Gaussian Noise

具有输入与次高斯噪声的部分观测随机线性时不变状态空间系统学习的PAC-贝叶斯界
Petreczky, Mihaly, Ahdab, Mohamad Al, Leth, John
Abstract
In this paper we derive a Probably Approximately Correct (PAC)-Bayesian error bound for partially observed linear time-invariant (LTI) stochastic dynamical systems in state-space form with inputs and sub-Gaussian noise. Such bounds are widespread in machine learning, and they are useful for characterizing the predictive power of models learned from finitely many data points. The bound derived in this paper relates the expectation of prediction errors with the prediction error generated by the model on the data used for learning. In addition, we show that it can also be used to derive bounds for the parameter estimation error. In turn, this allows us to provide finite-sample error bounds for the prediction error and parameter estimation error for a wide class of system identification algorithms. Furthermore, as LTI systems are a sub-class of recurrent neural networks (RNNs), these error bounds could be a first step towards PAC-Bayesian bounds for RNNs.
Chinese Translation
本文针对具有输入和次高斯(sub-Gaussian)噪声的、以状态空间形式表示的部分观测线性时不变(LTI)随机动态系统,推导了一个可能近似正确(PAC)-贝叶斯误差界。此类界在机器学习中十分常见,可用于刻画从有限数据点学习得到的模型的预测能力。本文推导的界将预测误差的期望与该模型在学习所用数据上产生的预测误差联系起来。此外,我们还证明该界可用于推导参数估计误差的界。进而,这使我们能够为一类广泛的系统辨识算法提供预测误差和参数估计误差的有限样本误差界。此外,由于LTI系统是循环神经网络(RNN)的一个子类,这些误差界可以成为迈向RNN的PAC-贝叶斯界的第一步。
cs.LG / 234 / 2609.08788

Adaptive Anisotropic Attention for Axis-Structured Signals

面向轴结构信号的自适应各向异性注意力机制
Jain, Mahir, Runwal, Parshva, Mishra, Aditya Ray, Kulkarni, Arvasu, Singh, Sandeep, Panwar, Siddharth
Abstract
Dense self-attention treats all token pairs as equally plausible before learning, an interaction-isotropic prior that can be mismatched to structured signals. For structured, low signal-to-noise ratio (SNR) signals such as EEG, dependencies are organized along the electrode and time axes, and this uniform prior exposes each token to many irrelevant interactions. We introduce Adaptive Anisotropic Attention (AAA), which splits attention into two paths: a temporal path, where each token attends to the tokens of its own electrode across time, and a spatial path, where it attends to the tokens of the other electrodes at the same time step. A small gate predicts, for every token, a convex combination of the two path outputs: two non-negative weights that sum to one. On six EEG downstream tasks, the resulting model, AXON (AXis-factorized Operator Network), improves mean balanced accuracy over a dense baseline under both linear probing and full fine-tuning. We show that both paths (temporal and spatial) are necessary and that the weighted sum beats a hard choice of one path; most of the benefit comes from the gate learning a different temporal/spatial balance at each layer of the network. Controlled audio spectrogram experiments show that axis factorization transfers beyond EEG. These results suggest that aligning attention with the natural axes of structured signals provides a useful inductive bias.
Chinese Translation
稠密自注意力在学习之前将所有token对视为同等可能,这是一种交互各向同性的先验,可能与结构化信号不匹配。对于脑电(EEG)等结构化、低信噪比(SNR)信号,其依赖关系沿电极轴和时间轴组织,而这种均匀先验使每个token暴露于大量无关交互中。我们提出自适应各向异性注意力(Adaptive Anisotropic Attention, AAA),将注意力分为两条路径:时间路径,每个token关注自身电极在时间维度上的token;空间路径,每个token关注同一时间步上其他电极的token。一个小型门控网络为每个token预测两条路径输出的凸组合,即两个和为一的非负权重。在六个EEG下游任务上,所得到的模型AXON(AXis-factorized Operator Network)在线性探测和全量微调两种设置下均优于稠密基线的平均平衡准确率。我们证明时间路径和空间路径都是必要的,且加权和优于对单一路径的硬性选择;大部分收益来自门控在网络每一层学习到不同的时间/空间平衡比例。受控的音频频谱图实验表明,轴分解方法可推广至EEG之外的领域。这些结果表明,使注意力与结构化信号的自然轴对齐能够提供一种有效的归纳偏置。
cs.LG / 235 / 2609.08798

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

基于在线策略反向蒸馏的弱到强泛化
Park, Youngrok, Bae, Sangmin, Jung, Hojung, Ko, Jongwoo, Choi, Yunseon, Kim, Young Jin, Cameron, Pashmina, Courville, Aaron, Yun, Se-Young
Abstract
Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student. We introduce On-Policy Reverse Distillation (OPRD), which evaluates the teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. By rescaling only verifier-supported updates, OPRD preserves the stationary points of policy optimization while accelerating learning beyond the teacher. In both successive model transfer and multi-teacher distillation, OPRD achieves higher performance with fewer student updates than existing RL and distillation approaches. Response-style analysis shows that OPRD students remain closer to models trained with verifier-based RL alone than to their weak teachers, suggesting that teacher guidance accelerates rather than redirects the student's own optimization. Results in conventional strong-to-weak distillation further demonstrate that OPRD effectively combines verifier-driven policy optimization with teacher guidance regardless of capacity ordering.
Chinese Translation
弱到强泛化探讨的是更强的模型能否从较弱的监督者那里学习并超越它们。这一问题对于 successive model generations(连续代际模型)和多领域整合尤为重要,因为在这些场景下从头重复前沿规模的后训练成本可能高得令人望而却步。然而,传统蒸馏方法将弱教师视为优化目标,可能将其能力上限强加给学生模型。我们提出了在线策略反向蒸馏(On-Policy Reverse Distillation, OPRD),该方法在学生模型的 rollout 上评估教师策略相对于其参考策略的偏移,并放大学生模型中由验证器(verifier)驱动的策略梯度在该方向上的分量。通过仅对验证器支持的更新进行重新缩放,OPRD 在保持策略优化驻点不变的同时,加速学习以超越教师模型。在连续模型迁移和多教师蒸馏两种场景中,OPRD 均以更少的学生更新次数取得了比现有强化学习和蒸馏方法更高的性能。响应风格分析表明,OPRD 训练出的学生模型与仅使用基于验证器的强化学习训练的模型更为接近,而非接近其弱教师,这说明教师引导是加速而非改变学生自身优化的方向。在传统强到弱蒸馏中的实验结果进一步表明,无论能力大小关系如何,OPRD 都能有效地将验证器驱动的策略优化与教师引导相结合。
cs.LG / 236 / 2609.08851

Length Generalization for Transformers via Compression

基于压缩的Transformer长度泛化
Zetzsche, Georg, Jiang, Hongjian, Yang, Andy, Bergsträßer, Pascal, Sälzer, Marco, Chiang, David, Lin, Anthony W.
Abstract
Recent advancements in transformer length generalization theory enable us to reliably predict when a transformer can learn to solve a task. In particular, the C-RASP hypothesis (a formalized version of the so-called RASP-l conjecture) posits that transformers length-generalize on a task if and only if a solution is expressible in the C-RASP language. While this hypothesis has strong empirical validation, theoretical problems arise from the fact that no computable length generalization bounds exist for C-RASP, alongside the discovery of seemingly contradictory experiments. To address these problems, we refine the C-RASP hypothesis utilizing the recently-proposed fragments C-RASP+ and C-RASP1. These fragments have computable length generalization bounds, though in the worst case requiring an extremely large (double exponential) sample size. It is an open question whether these sample size bounds are tight. In this paper, we resolve this open question by providing an exponentially tighter bound. In doing so, we show a polynomial length generalization bound for transformers if we adopt compressed strings, via a novel connection to power words. As an application, we show how this yields a fine-grained analysis of the C-RASP conjecture that resolves contradicting experimental evidence against it.
Chinese Translation
Transformer长度泛化理论的最新进展使我们能够可靠地预测Transformer何时能够学会解决某项任务。特别地,C-RASP假设(即所谓RASP-l猜想的形式化版本)提出:当且仅当某任务存在可用C-RASP语言表达的解时,Transformer才能在该任务上实现长度泛化。尽管该假设已获得强有力的实证验证,但由于C-RASP不存在可计算的长度泛化界,加之出现了看似相互矛盾的实验结果,理论上产生了若干问题。为解决这些问题,我们利用近期提出的C-RASP+与C-RASP1片段对C-RASP假设进行了改进。这些片段具有可计算的长度泛化界,尽管在最坏情况下需要极其庞大(双指数级)的样本量。这些样本量界是否紧致仍是一个悬而未决的问题。在本文中,我们通过给出一个指数级更紧的界来解决这一开放问题。在此过程中,我们通过与幂词(power words)的一种新颖联系,证明了在采用压缩字符串时Transformer具有多项式级的长度泛化界。作为应用,我们展示了这一结果如何为C-RASP猜想提供细粒度分析,从而化解了与之相矛盾的实验证据。
cs.LG / 237 / 2609.08855

Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems

面向假设性仿真的地球系统世界模型:以陆地生态系统为案例研究
Wang, Zhihao, Wang, Ruichen, Li, Ruohan, Ma, Lei, Hurtt, George, Jia, Xiaowei, Mai, Gengchen, Wang, Shaowen, Xie, Yiqun
Abstract
Machine learning emulators have become essential for accelerating expensive Earth-system simulations, but most existing approaches remain passive forecasters: they reproduce simulator trajectories under prescribed forcings without an explicit interaction mechanism for user-specified interventions. This limits their use in interactive scientific workflows and Earth-system digital twins, where users often need to explore how a system would respond if selected state components were changed. We propose an action-conditioned world-modeling framework for Earth-system emulation that reformulates simulator trajectories as supervision for controllable state-transition learning. The key idea is transition-action pretraining: naturally observed state changes are treated as label-free action supervision, allowing the model to learn both prescribed dynamics and action-conditioned responses without manually annotated interventions. We further introduce masked response learning to infer unobserved variables under partial state edits and learn coupled system dependencies. We test this framework on ecosystem dynamics across six global regions and multiple stand ages. Experiments show that the model preserves competitive long-horizon emulation accuracy while enabling controllable structural interventions and coherent responses in coupled ecosystem-cycle variables. These results suggest a practical route from passive Earth-system emulators toward interactive, intervention-aware scientific surrogates.
Chinese Translation
机器学习模拟器已成为加速高成本地球系统仿真的重要工具,但现有方法大多仍是被动的预测器:它们在给定强迫条件下复现仿真轨迹,缺乏针对用户指定干预的显式交互机制。这限制了其在交互式科学工作流和地球系统数字孪生中的应用,而在这些场景中,用户往往需要探索当系统中选定的状态分量被改变时,系统将如何响应。我们提出了一种面向地球系统模拟的动作条件世界模型框架,将仿真轨迹重新表述为可控状态转移学习的监督信号。其核心思想是转移-动作预训练:将自然观测到的状态变化视为无标签的动作监督,使模型无需人工标注的干预数据即可同时学习给定动力学和动作条件响应。我们进一步引入掩码响应学习,以在部分状态编辑下推断未观测变量并学习耦合系统的依赖关系。我们在六个全球区域和多个林龄的生态系统动力学上对该框架进行了测试。实验表明,该模型在保持具有竞争力的长时程模拟精度的同时,能够实现可控的结构性干预,并在耦合的生态系统循环变量中产生连贯一致的响应。这些结果表明了一条从被动地球系统模拟器迈向交互式、干预感知的科学代理模型的可行路径。
cs.LG / 238 / 2609.08901

The BatchNorm Illusion: Diagnosing Normalization Artifacts in Machine Unlearning Evaluation

BatchNorm 幻象:机器遗忘评估中归一化伪影的诊断
Kalani, Aaryaman, Mandal, Murari, Kumar, Dhruv, Kankanhalli, Mohan, Sinha, Yash
Abstract
Approximate machine unlearning aims to remove the influence of specific training data from a trained model without retraining from scratch. We identify a previously undocumented confound in how unlearning is evaluated on BatchNorm-based architectures: a single forward pass over retain data, an operation that modifies no weight, can deterministically rewrite the model's normalization state and reverse the apparent surface-metric forgetting. We formalize this operation as a weight-preserving fixed-point operator and prove that any pre-versus-post gap it induces is provably attributable to BN running statistics rather than to any modification the unlearning method made to the weights. This attribution claim cleanly separates measurement failure (BN artifact) from encoder failure (residual weight-encoded information, recently documented in concurrent work), and the same operator framework yields a unique decomposition of linear-probe elevation into BN-measurement-bias and encoder-geometry components. Empirically, the artifact reverses headline forget accuracy by up to 78 pp across nine evaluated methods on standard benchmarks; an attacker with as few as 10 unlabeled images recovers most of the masked accuracy; and a strict GroupNorm control reduces the artifact to zero across all methods. The tested membership-inference attacks change little under recalibration, locating the observed evaluation failure in forget accuracy and linear probing.
Chinese Translation
近似机器遗忘旨在从已训练模型中移除特定训练数据的影响,而无需从头重新训练。我们发现了一个此前未被记录的混淆因素,它存在于基于 BatchNorm 的架构上如何评估遗忘的过程中:仅需对保留数据进行一次前向传播——一个不修改任何权重的操作——就能确定性地重写模型的归一化状态,并逆转表面指标上表现出的遗忘。我们将该操作形式化为一个保持权重的定点算子,并证明它所导致的任何“遗忘前后差距”都可被严格归因于 BN 的运行统计量,而非遗忘方法对权重所做的任何修改。这一归因性断言清晰地区分了测量失效(BN 伪影)与编码器失效(残留在权重中编码的信息,近期有平行工作对此进行了记录),且同一算子框架还对线性探针(linear probe)准确率的提升给出了唯一的分解,将其分为 BN 测量偏差与编码器几何结构两个分量。实证结果显示:该伪影使九种被评估方法在标准基准上的遗忘准确率这一核心指标逆转高达 78 个百分点;攻击者仅需 10 张无标签图像即可恢复大部分被掩蔽的准确率;而采用严格的 GroupNorm 对照则能在所有方法上将该伪影降为零。所测试的成员推断攻击在重新校准下变化甚微,从而将所观察到的评估失效定位于遗忘准确率和线性探针上。
cs.LG / 239 / 2609.08970

GraphFAS: A Distributed System for Automated Graph Feature Generation and Selection in Industrial Transaction Networks

GraphFAS:面向工业交易网络的自动化图特征生成与选择分布式系统
Luo, Yice, Zhu, Yun, Chen, Xi, Liu, Yongchao, Zeng, Xintan, Huan, Chengying, Zhang, Kai, Zhang, Jinrui, Zhang, Juelu, Zheng, Jiajun
Abstract
Industrial fraud detection often relies on costly expert-crafted features that overlook graph-structured relational signals, while GNNs often do not meet the interpretability and deployment requirements of financial risk control. We propose GraphFAS (Graph Feature Automated Selection), a distributed feature selection procedure based on Boruta that bridges this gap through: (1) a non-parametric graph feature generation module that constructs explicit, interpretable structural features via multi-hop subgraph extraction and multi-scale aggregation without learned parameters; and (2) an automated distributed feature selection algorithm extending Boruta with median-based aggregation across partitions to robustly identify informative features at scale with minimal domain expertise. Compared with end-to-end GNN pipelines, GraphFAS decouples feature aggregation from model training, enabling direct integration with tabular models and direct compatibility with TreeSHAPbased explanations. Deployed in Alipay, GraphFAS delivers orderof-magnitude improvements in engineering efficiency while showing strong performance against expert-driven and graph-learning baselines on large-scale graphs.
Chinese Translation
工业欺诈检测通常依赖代价高昂的专家手工特征,这些特征忽略了图结构中的关系信号;而图神经网络(GNN)往往难以满足金融风控对可解释性和部署的要求。我们提出GraphFAS(Graph Feature Automated Selection,图特征自动选择),一种基于Boruta的分布式特征选择方法,通过以下方式弥合这一差距:(1)一个非参数化的图特征生成模块,通过多跳子图提取和多尺度聚合构建显式、可解释的结构特征,且无需学习参数;(2)一种自动化的分布式特征选择算法,对Boruta进行扩展,通过跨分区的中位数聚合,在大规模场景下以最少领域专业知识稳健地识别有效特征。与端到端GNN流程相比,GraphFAS将特征聚合与模型训练解耦,可直接与表格模型集成,并天然兼容基于TreeSHAP的解释方法。GraphFAS已部署于支付宝,在工程效率上实现了数量级的提升,同时在大规模图上相比专家驱动和图学习基线展现出强劲的性能。
cs.LG / 240 / 2609.08981

Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling

Transformer作为上下文采样器:从闭式扩散到无估计采样
Adibi, Arman, Jafari, Alireza, Ghavamzadeh, Mohammad, Daneshmand, Hadi
Abstract
A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates. Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression. We prove that in-context learning extends further to \emph{data generation}: frozen transformers can simulate iterative generative samplers from in-context samples. We first show that transformers can realize closed-form and smoothed closed-form diffusion samplers. The construction identifies a concrete generative role for softmax attention: it computes responsibility weights and weighted empirical averages, while feedforward layers implement Euler updates. To empirically relate these constructions to pretrained language models, we study \emph{semantic-topic sampling}: prompts consisting of words drawn from a common semantic category, such as animals, foods, or cities. Across transformer layers, the normalized hidden states exhibit a two-stage geometry: they move toward a uniform spherical reference in intermediate layers and then return to structured, topic-dependent representations near the output. We further measure an interacting-particle energy on these hidden-state clouds and observe the same U-shape pattern. We then prove that transformers can approximate an energy-based sampler, constructing the same U-shape energy across the layers.
Chinese Translation
越来越多的研究工作表明,大语言模型并非单纯的统计记忆器,而是具备上下文学习能力:在测试时仅利用提示中提供的示例进行推理,而无需任何参数更新。先前的理论工作已经证明,这种能力可扩展到诸如线性回归等监督学习任务。我们证明上下文学习可进一步扩展到数据生成:冻结的Transformer能够基于上下文样本模拟迭代式生成采样器。我们首先证明Transformer可以实现闭式及平滑闭式的扩散采样器。该构造揭示了softmax注意力的一个具体生成角色:它负责计算责任权重和加权经验均值,而前馈层则实现欧拉更新。为了从实证上将这些构造与预训练语言模型联系起来,我们研究了语义主题采样:即由来自同一语义类别(如动物、食物或城市)的词语构成的提示。跨Transformer各层,归一化后的隐状态呈现出两阶段几何形态:在中间层中它们向均匀球面参考分布靠近,随后在靠近输出层时回归到结构化的、依赖主题的表示。我们进一步在这些隐状态云上度量了一种相互作用粒子的能量,并观察到相同的U形模式。随后我们证明Transformer能够近似一个基于能量的采样器,并在各层间构造出相同的U形能量。
cs.LG / 241 / 2609.08992

Physics-Informed Deep Learning for False Ventricular Tachycardia Alarm Reduction in the ICU

基于物理信息的深度学习用于减少ICU中室性心动过速的误报
Papastathopoulos-Katsaros, Athanasios, Stavrianidi, Alexandra, Liu, Zhandong
Abstract
False ventricular tachycardia (VT) alarms are a leading contributor to alarm fatigue in intensive care units. We propose a deep learning framework combining a 1D SE-ResNet with ICU-realistic data augmentations and a physics-informed auxiliary reconstruction task based on the three-element Windkessel hemodynamic model, implemented as a differentiable forward simulation. By requiring the network's latent representation to produce physiologically plausible arterial pressure waveforms, artifact-driven ECG patterns are penalized while true VT remains coherent across modalities. Evaluated on the VTaC benchmark under a strict real-time protocol (10-second pre-alarm window), our method achieves a 5-point Challenge Score improvement over prior state-of-the-art. Ablation studies confirm that the physics-informed objective is the primary performance driver, providing gains in accuracy, 2x label efficiency, and more localized and clinically meaningful ECG segments.
Chinese Translation
室性心动过速(VT)误报警是重症监护病房(ICU)中报警疲劳的主要成因。我们提出了一种深度学习框架,将一维SE-ResNet与符合ICU实际情况的数据增强方法相结合,并引入基于三元素Windkessel血流动力学模型的物理信息辅助重建任务,该任务以可微分前向仿真方式实现。通过要求网络的潜在表示生成生理上合理的动脉血压波形,由伪迹驱动的ECG模式受到惩罚,而真实的VT信号在跨模态间保持一致性。在严格的实时协议(10秒报警前窗口)下,我们在VTaC基准数据集上评估了该方法,相比此前最先进方法,Challenge Score提升了5分。消融实验证实,物理信息目标是性能提升的主要驱动力,带来了准确率的提高、2倍的标签效率提升,以及更加局部化和更具临床意义的ECG片段。
cs.LG / 242 / 2609.09009

Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks

放手还是学会自我纠正:面向约束离散任务的连续扩散模型
Drozdova, Mariia, Nguyen, Stéphane Liem, Fleuret, François
Abstract
Denoising Diffusion Probabilistic Models (DDPMs) generate samples by starting from noise and repeatedly denoising while keeping each update close to the current noisy state. This behavior is effective in many continuous domains, but its role is less clear for globally constrained discrete tasks, such as Sudoku, graph connectivity, Latin squares, and N-queens. In such settings, early discrete errors can be difficult to undo. As a result, standard diffusion sampling may preserve early mistakes, even when the model's clean predictions are informative. We compare standard samplers to sampling directly from the model's clean prediction. Without retraining, this single change improves Sudoku validity from 31% to 95%, with consistent gains across the other discrete tasks. We hypothesize that staying close to the current noisy state is harmful because the reverse trajectory can drift off the forward noising distribution the model was trained on. To reduce this train-test mismatch, we further introduce self-correction training, which exposes the model to its own predictions, improving robustness to errors that arise during inference. This substantially improves the performance of standard samplers. Our results suggest that continuous diffusion models can learn nontrivial global constraints, but discrete reasoning tasks require better alignment between training and inference: either through samplers that reduce commitment to early decisions, or through training that teaches the model to correct its own inference-time errors.
Chinese Translation
去噪扩散概率模型(DDPM)从噪声出发,通过反复去噪生成样本,同时使每次更新保持接近当前的带噪状态。这一行为在许多连续领域中十分有效,但其在全局约束离散任务(如数独、图连通性、拉丁方和N皇后问题)中的作用尚不明确。在这类任务中,早期的离散错误往往难以撤销。因此,标准扩散采样可能会保留早期错误,即使模型的干净预测本身是有效的信息。我们将标准采样器与直接从模型的干净预测进行采样进行了比较。在不重新训练的情况下,仅这一项改动就将数独的有效率从31%提升至95%,并在其他离散任务上取得了一致的增益。我们假设,保持接近当前带噪状态是有害的,因为反向轨迹可能偏离模型训练时所基于的前向加噪分布。为了减少这种训练-测试失配,我们进一步引入了自我纠正训练,让模型接触自身的预测结果,从而提高其对推理过程中出现的错误的鲁棒性。这显著提升了标准采样器的性能。我们的结果表明,连续扩散模型能够学习非平凡的全局约束,但离散推理任务需要训练与推理之间更好的对齐:要么通过减少对早期决策承诺的采样器,要么通过教会模型纠正自身推理时错误的训练方式来实现。
cs.LG / 243 / 2609.09038

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

推理表征有助于人类评估大语言模型的输出吗?
Lim, Jaewoo, Shin, Sungbok, Hong, Sanghyun
Abstract
Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations as human-facing interfaces rather than proxies for model reasoning ability. We conduct a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order. The study collects fine-grained judgments of structural understanding, error detection and localization, and trust calibration. Our study shows a mismatch between perceived preference and support for human evaluation. Participants prefer planning- and decomposition-based representations, but simpler chain-of-thought traces better support verification, trust, and interpretability. Preferred representations also introduce calibration risks, with more false alarms on correct traces and high trust despite low willingness to verify.
Chinese Translation
推理表征正日益被用作大语言模型输出的解释。然而,它们通常以模型为中心的标准(如答案准确性和忠实度)进行评估,尚不清楚这些表征是否有助于人类评估模型的响应。在这项工作中,我们将推理表征视为面向人类的交互界面,而非模型推理能力的替代指标。我们在不同复杂度的任务上对六种推理格式进行了受控人类实验研究,并借助一个基于Web的框架对任务领域、问题实例和表征顺序进行随机化。该研究收集了关于结构理解、错误检测与定位以及信任校准的细粒度判断。我们的研究表明,感知偏好与对人类评估的实际支持之间存在错位:参与者更偏好基于规划和分解的表征,但更简单的思维链更能有效支持验证、信任和可解释性。此外,受偏好的表征还会带来校准风险,即在正确的推理轨迹上产生更多误报,且在验证意愿较低的情况下仍保持高度信任。
cs.LG / 244 / 2609.09054

Training-Free Task Vectors for LLM Behavioral Control

面向大语言模型行为控制的无训练任务向量
Perin, Gabriel J., Boscaini, Lucas, Araujo, André, Hirata, Nina S. T.
Abstract
Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. However, this reliance on fine-tuning makes discovering such directions costly and limits the practicality of post-training model editing. To address this limitation, we introduce Training-Free Task Vectors (TFTVs), a novel method to compute task-vector-like directions without requiring fine-tuning. Our method maps activation steering vectors to rank-one weight-space edits using only forward-pass statistics, while satisfying arithmetic properties that directly support learning via addition, forgetting via subtraction, and the composition of multiple edits. Empirically, we evaluate TFTVs on large language model behavioral control tasks and show that they consistently amplify, suppress, and compose target behaviors while preserving general knowledge and problem-solving skills. We also validate our method against other editing and steering baselines, experimentally demonstrating that TFTVs achieve stronger trait control with better or competitive utility preservation. We hope our work opens new directions for the community in post-training model editing and broader training-free model control. Code is available on the project website: tftv-llm.github.io.
Chinese Translation
任务向量(Task Vectors)通过在权重空间中识别语义上有意义的方向来实现训练后模型编辑,其通常被计算为微调模型与预训练初始模型之间的差值。然而,这种对微调的依赖使得发现此类方向的代价高昂,限制了训练后模型编辑的实用性。为解决这一局限,我们提出了无训练任务向量(Training-Free Task Vectors, TFTVs),这是一种无需微调即可计算类任务向量方向的新方法。我们的方法仅利用前向传播的统计信息,将激活导向向量(activation steering vectors)映射为秩一(rank-one)的权重空间编辑,同时满足支持通过加法实现学习、通过减法实现遗忘以及多重编辑组合的算术性质。在实证方面,我们在大语言模型行为控制任务上评估了TFTVs,结果表明它们能够持续地放大、抑制和组合目标行为,同时保持通用知识和问题解决能力。我们还将该方法与其他编辑和导向基线方法进行了对比验证,实验证明TFTVs在实现更强特质控制的同时,具有更好或相当的功能保持能力。我们希望这项工作能为社区在训练后模型编辑以及更广泛的无训练模型控制方向开辟新的道路。代码可在项目网站获取:tftv-llm.github.io。
cs.LG / 245 / 2609.09059

PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games

PlayTrain:一个用于LLM生成的可适配JavaScript游戏的高效强化学习框架
Truong, Ryan, Ying, Lance, Gershman, Samuel J., Irie, Kazuki
Abstract
While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existing ones to support new features, has been a laborious process requiring extensive hand-coding. Here we present PlayTrain, an RL framework that combines the abilities of large language models (LLMs) to robustly generate JavaScript (JS) games from a minimal human prompt, and an efficient pipeline that can run any JS game in a standard 'gym' environment. Not only are recent LLMs particularly good at writing JS code, but the JS format also allows users to easily play generated VGEs, while PlayTrain enables us to train RL agents on the exact same games. We demonstrate multiple use cases of PlayTrain, including cloning well-known Atari and ProcGen games in simple JS, where PlayTrain trains pixel-based agents end-to-end at over 1M agent-decisions per second on a single GPU node; and creating modified versions thereof (e.g., that support novel test sets, procedural generation logics, or game dynamics). Through PlayTrain, we reimagine RL VGE development: all we need is a single JS file, generated and modified through an LLM. We discuss promising future RL research directions that PlayTrain unlocks.
Chinese Translation
尽管许多视频游戏环境(VGE)在推动强化学习(RL)发展方面发挥了关键作用,但开发新的VGE或修改现有VGE以支持新功能,一直是一项需要大量手工编码的繁琐工作。本文提出了PlayTrain,一个强化学习框架,它结合了大语言模型(LLM)从极少的人类提示中稳健生成JavaScript(JS)游戏的能力,以及一个可在标准'gym'环境中运行任何JS游戏的高效流水线。近期的LLM不仅特别擅长编写JS代码,而且JS格式还允许用户轻松玩生成的VGE,同时PlayTrain使我们能够在完全相同的游戏上训练RL智能体。我们展示了PlayTrain的多个用例,包括用简单的JS克隆著名的Atari和ProcGen游戏,其中PlayTrain在单个GPU节点上以每秒超过100万次智能体决策的速度对基于像素的智能体进行端到端训练;以及创建这些游戏的修改版本(例如,支持新测试集、程序化生成逻辑或游戏动态的版本)。通过PlayTrain,我们重新构思了RL VGE的开发方式:我们只需要一个JS文件,通过LLM生成和修改。我们还讨论了PlayTrain所开启的富有前景的未来RL研究方向。
cs.LG / 246 / 2609.09062

Multi-Task Learning for Sparsely-Labeled Time Series: A Case Study on Cold-Hardiness Modeling

面向稀疏标注时间序列的多任务学习:以耐寒性建模为例
Saxena, Aseem, Pesántez-Cabrera, Paola, Magby, Jonathan, Keller, Markus, Fern, Alan
Abstract
We present a real-world case study of multi-task learning (MTL) for temporal process modeling from limited data with temporally sparse labels. Specifically, we investigate multi-task learning for the important agricultural problem of predicting grape cold hardiness, which is the temperature at which lethal freezing occurs. Cold hardiness changes in response to weather and is difficult to measure directly in the field. Thus, growers rely on predictions to decide when to apply costly frost mitigation measures. We apply recurrent neural networks (RNNs) for daily cold-hardiness prediction from time series weather data. A major challenge is that the cold hardiness response varies across plant cultivars and ground-truth data for each cultivar is temporally sparse and limited. To address this challenge, we investigate multi-task learning (MTL) approaches for combining data, where different tasks correspond to different cultivars. We develop a variety of MTL architectures and evaluate them in both MTL and transfer learning settings. Our results show significant differences between architectures and that certain architectures are able to consistently outperform single-task learning and state-of-the-art scientific models. Additionally, we show similar results for the qualitatively different, but related, task of budbreak prediction. Further, improved accuracy for budbreak and cold hardiness is achieved by a single MTL model that simultaneously learns both tasks.
Chinese Translation
我们提出了一项真实场景下的案例研究,探讨在数据有限且标签在时间上稀疏的条件下,利用多任务学习(MTL)进行时间过程建模。具体而言,我们研究多任务学习在预测葡萄耐寒性这一重要农业问题上的应用,耐寒性是指发生致死性冻害的温度。耐寒性随天气变化而变化,且难以在田间直接测量,因此种植者需要依靠预测来决定何时采取昂贵的防霜冻措施。我们应用循环神经网络(RNN)基于时间序列气象数据进行每日耐寒性预测。一个主要挑战在于,不同植物品种的耐寒性响应存在差异,而每个品种的真实观测数据在时间上稀疏且有限。为应对这一挑战,我们研究了多任务学习(MTL)方法来融合数据,其中不同任务对应不同品种。我们开发了多种MTL架构,并在多任务学习和迁移学习两种设置下对其进行评估。结果表明,不同架构之间存在显著差异,且某些架构能够持续优于单任务学习和最先进的科学模型。此外,对于性质不同但相关的芽萌发预测任务,我们也得到了类似的结果。进一步地,通过一个同时学习这两个任务的单个MTL模型,芽萌发和耐寒性的预测精度均得到提升。
cs.LG / 247 / 2609.09075

ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR

ThinkPrior:用于RLVR冷启动提示选择的零推演难度先验
Sha, Tommy, Zhai, Skylar, Zhao, Siqi
Abstract
In reinforcement learning with verifiable rewards (RLVR) trained with group relative policy optimization (GRPO), the KL-free reward-advantage term studied here depends on within-group reward variation. If all rollouts in a group are correct or all are wrong, their group-relative advantages are identically zero; these zero-advantage silent groups provide no reward-advantage gradient, yet uniform sampling spends 39% of a run's rollouts on them. History-based prompt selection must first spend target-policy rollouts to estimate difficulty, creating a cold start with rollout waste; ThinkPrior instead uses an external anchor in one offline pass to construct a zero-rollout difficulty prior before the first target-policy rollout. The verifier-scored anchor pass rate supplies an external-anchor initialization for a Beta posterior; ThinkPrior selects by expected learnability and then updates from training outcomes, changing neither the loss nor the optimizer. On Qwen2.5-Math-7B across sixteen seeds, ThinkPrior more than halves early silent groups and cuts wasted rollouts through step 30 by nearly a fifth, while we detect no difference in final accuracy. On this 250-prompt pool the fixed-budget result is a reallocation rather than a net saving. The measured ThinkPrior+DAPO composition reduces generated rollouts by 10.6% while both arms retain the same 3840-rollout update budget. The prior requires no target-policy rollout before the first selection, but the posterior thereafter uses target-policy outcomes.
Chinese Translation
在使用组相对策略优化(GRPO)训练的可验证奖励强化学习(RLVR)中,本文研究的无KL散度奖励-优势项依赖于组内奖励差异。如果一组中的所有推演结果全部正确或全部错误,其组相对优势恒为零;这些零优势静默组无法提供奖励-优势梯度,而均匀采样却会将一次运行中39%的推演花费在其上。基于历史的提示选择必须先消耗目标策略推演来估计难度,从而造成推演浪费的冷启动问题;ThinkPrior 则改用外部锚点,通过一次离线过程在首次目标策略推演之前构建零推演难度先验。由验证器评分的锚点通过率为 Beta 后验提供外部锚点初始化;ThinkPrior 依据期望可学习性进行选择,随后根据训练结果进行更新,既不改变损失函数也不改变优化器。在 Qwen2.5-Math-7B 上跨越十六个随机种子的实验表明,ThinkPrior 将早期静默组减少一半以上,并将前30步的浪费推演减少近五分之一,同时我们在最终准确率上未检测到差异。在这个包含250个提示的池上,固定预算的结果是重新分配而非净节省。实测的 ThinkPrior+DAPO 组合将生成推演减少了10.6%,同时两个分支保持相同的3840次推演更新预算。该先验在首次选择之前不需要目标策略推演,但此后的后验更新则使用目标策略的结果。
cs.LG / 248 / 2609.09099

Curriculum Learning as Transport: Understanding Curricula with Wasserstein Geodesics

作为传输问题的课程学习:基于Wasserstein测地线的课程理解
Shin, Changho, Alvarez-Melis, David
Abstract
Curriculum learning is governed by several coupled design choices---how difficulty is defined, how examples are ordered, how much exposure each level receives, and how quickly training moves across levels---making it hard to isolate what actually helps. We present Wasserstein curriculum paths, a simple transport-based framework that decouples these factors by representing curricula as trajectories of training distributions over discrete difficulty levels. Across a calibrated synthetic suite with 12 tasks and 33 difficulty axes, we use this framework to isolate the effects of ordering, matched exposure, endpoint smoothness, and pacing under fixed training budgets. We find that curriculum effects are strongly context-dependent: no single strategy dominates across tasks, difficulty axes, and budgets, and curricula mainly change where a fixed budget is spent most effectively. Within this framework, easy-to-hard ordering improves hard-level performance relative to exposure-matched static sampling, showing that the benefit is not explained by cumulative exposure alone. We further show that endpoint smoothness and pacing substantially affect where along the difficulty spectrum a curriculum is effective. Finally, we show that the same transport view naturally supports extensions to learned pacing through geometry and to structured difficulty spaces beyond one-dimensional orderings.
Chinese Translation
课程学习受多个相互耦合的设计选择支配——难度如何定义、样本如何排序、每个层级获得多少曝光、以及训练在层级之间推进多快——这使得很难分离出真正起作用的因素。我们提出了Wasserstein课程路径(Wasserstein curriculum paths),这是一个简单的基于传输的框架,它将课程表示为训练分布在离散难度层级上的轨迹,从而解耦这些因素。在一个包含12个任务和33个难度轴的校准合成实验套件上,我们利用该框架在固定训练预算下分离出排序、匹配曝光、端点平滑度和节奏推进的影响。我们发现课程效应强烈依赖于具体情境:没有任何单一策略能够在所有任务、难度轴和预算下都占据优势,课程主要改变的是固定预算被最有效分配的位置。在该框架内,从易到难的排序相较于曝光匹配的静态采样提升了困难层级的性能,这表明其收益不能仅由累积曝光来解释。我们进一步表明,端点平滑度和节奏推进会显著影响课程在难度谱上发挥作用的位置。最后,我们展示了这种传输视角自然地支持了若干扩展:通过几何结构学习节奏推进,以及将难度空间扩展到超越一维排序的结构化形式。
cs.LG / 249 / 2609.09116

When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay

尺度不变优化何时变得不稳定?带权重衰减的精确调度定律
Amin, Hasan, Chang, Wei-Kai, Khanna, Rajiv
Abstract
Normalization renders large parts of neural networks effectively scale invariant, inducing a hidden feedback loop in which learning-rate schedules and weight decay interact through the parameter norm to control the effective step taken by the optimizer. We show that this interaction is governed by an exact discrete-time law: a single scalar quantity captures all schedule and decay forcing, while norm growth induces an opposing geometric self-quenching effect. This yields a sharp boundary that cleanly separates contraction- and expansion-dominated effective learning rate regimes. To understand the underlying mechanism, we provide exact analysis of a fully solved normalized regression model where the dynamics reduce to two dimensions and show that the balance point is intrinsically unstable, implying that constant learning rate with weight decay cannot stably maintain an interior equilibrium and instead produces recurrent behavior driven by discrete-time Jacobian structure. We further extend this perspective across optimizers through unified homogeneous-optimizer framework that reveals a structural dichotomy in self-quenching strength, providing a first-principles explanation for why adaptive methods exhibit systematically weaker stabilization under normalization. Across dynamical systems and neural networks (MLP, CNN, GPT2 / MNIST, CIFAR, wikiText, OpenWebText), the predicted law holds with high precision and enables direct control of training via the identified scalar, with performance peaking sharply at the predicted boundary. Together, these results isolate a single governing quantity for scale-invariant optimization, providing a precise and actionable lens on training dynamics, optimizer behavior, and schedule design in modern deep learning. Code is available in https://github.com/shasanamin/normalized-optimization-dynamics.
Chinese Translation
归一化使神经网络的大部分部分在效果上具有尺度不变性,由此产生了一个隐藏的反馈回路:学习率调度与权重衰减通过参数范数相互作用,控制优化器实际采取的有效步长。我们证明这种相互作用由一个精确的离散时间定律所支配:单个标量量刻画了所有调度与衰减的驱动作用,而范数增长则诱导出一种相反的几何自淬灭效应。这给出了一个清晰的边界,精确地区分了有效学习率中以收缩为主和以扩张为主的两种状态。为理解其内在机制,我们对一个完全可解的归一化回归模型进行了精确分析,其动力学可约化为二维,并证明该平衡点在本质上是 instable(不稳定的),这意味着常数学习率配合权重衰减无法稳定地维持内部平衡,而是产生由离散时间雅可比结构驱动的循环行为。我们进一步通过统一的齐次优化器框架将这一视角扩展到多种优化器,揭示了自淬灭强度上的结构性二分现象,为自适应方法在归一化下表现出系统性较弱的稳定化提供了第一性原理解释。在动力系统与神经网络(MLP、CNN、GPT-2;MNIST、CIFAR、WikiText、OpenWebText)上的实验表明,所预测的定律以高精度成立,并可通过所识别的标量直接控制训练过程,性能在预测边界处达到尖锐峰值。总之,这些结果为尺度不变优化分离出了单一的主导量,为现代深度学习中的训练动力学、优化器行为和调度设计提供了精确且可操作的视角。代码见 https://github.com/shasanamin/normalized-optimization-dynamics。
cs.LG / 250 / 2609.09130

Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks

稀疏激活神经网络的近乎紧致的Rademacher复杂度界
Li, Xiaoyu, Sha, Zhizhou, Jiang, Jiaojiao, Gao, Junbin, Han, Andi
Abstract
An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input-dependent sparsity in the one-hidden-layer ReLU model of Awasthi et al. (COLT 2024). For width $s$, at most $k$ active units per input, and effective weight and bias bounds $W,B$, every size-$m$ sample in the class's fixed radius-$R$ input domain satisfies $\mathcal{R}(S)\le CWR\min\{k,\sqrt{sk/m}\log^{3/2}(2m)\}+kB/\sqrt m$. A support-preserving cover and a single normalized chaining argument remove the previous explicit dimension factor, up to logarithms. Lower bounds on appropriate i.i.d. marginals match up to those logarithms, showing how changing active units across inputs retains a width dependence. The input domain matters: zero-bias networks sparse on the entire ball have at most $2k$ nonzero units and complexity $O(kWR/\sqrt m)$, whereas bias bounds comparable to $WR$ restore the worst-case rate on that same domain in only logarithmic dimension. A spherical-cap construction proves the latter claim without assuming sparsity merely on the sampling support. For a specified normalized bounded loss and biases comparable to $WR$, we also obtain agnostic minimax excess-risk bounds of order $\min\{1,\sqrt{s/(km)}\}$ up to logarithms.
Chinese Translation
一个输入可能仅激活少数隐藏单元,即使不同输入合起来使用了整个网络。我们在Awasthi等人(COLT 2024)的单隐层ReLU模型中研究这种依赖于输入的稀疏性的统计复杂度。对于宽度为$s$、每个输入至多激活$k$个单元、有效权重和偏置界为$W$和$B$的情形,该类别在固定半径$R$输入域中的任意大小为$m$的样本均满足$\mathcal{R}(S)\le CWR\min\{k,\sqrt{sk/m}\log^{3/2}(2m)\}+kB/\sqrt m$。通过一个保持支撑集的覆盖以及一次归一化的链式(chaining)论证,我们在仅相差对数因子的意义下去除了先前显式的维度因子。在适当的独立同分布边缘分布上的下界与之相差相同的对数因子,表明不同输入间激活单元的变化如何保留了宽度依赖性。输入域本身也很关键:零偏置且在整个球上稀疏的网络至多有$2k$个非零单元,其复杂度为$O(kWR/\sqrt m)$;而当偏置界与$WR$可比时,即使在仅对数维度的情形下,同一输入域上也会恢复最坏情况的速率。一个球冠(spherical-cap)构造在不假设仅在采样支撑集上稀疏的前提下证明了后一结论。对于给定的归一化有界损失以及与$WR$可比的偏置,我们还在相差对数因子的意义下得到了量级为$\min\{1,\sqrt{s/(km)}\}$的不可知(agnostic)极小化极大超额风险界。
cs.LG / 251 / 2609.09135

Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation

面向代码生成测试时强化学习的熵正则化排序掩码策略优化
Xu, Jiacheng, Chen, Feng, Xu, Xiuneng, An, Bo
Abstract
Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on unlabeled test-time tasks with canonical answers, but this breaks down for code generation because programs cannot be compared by surface form and therefore do not directly provide a usable training signal. To make TTRL applicable to code generation, we propose probe-driven TTRL, which constructs output-free probe inputs from the problem statement, executes candidate programs on these probes, and defines a Probe Consensus Reward (PCR) from the resulting behavioral agreement. PCR provides a behavioral training signal for open-vocabulary programs, but it is not a fully reliable verifier and remains susceptible to reward hacking through spurious consensus. We therefore introduce Entropy-Regularized Rank-Masked Policy Optimization (ERPO), which converts low PCR into conservative negative updates through rank masking and controls policy drift with an entropy ceiling. On coding benchmarks, ERPO substantially improves pass@1 and pass@k in both in-domain adaptation and zero-shot transfer.
Chinese Translation
现有的测试时强化学习(TTRL)方法通过对具有标准答案的无标注测试时任务进行答案级自投票来获取奖励,但该方法在代码生成任务中失效,因为程序无法通过表面形式进行比较,因而无法直接提供可用的训练信号。为使TTRL适用于代码生成,我们提出探测驱动的TTRL(probe-driven TTRL),该方法从问题描述中构建无需期望输出的探测输入(probe inputs),在这些探测输入上执行候选程序,并根据由此产生的行为一致性定义探测一致性奖励(Probe Consensus Reward, PCR)。PCR为开放词表程序提供了行为层面的训练信号,但它并非完全可靠的验证器,仍容易通过虚假一致性遭受奖励攻击(reward hacking)。因此,我们提出熵正则化排序掩码策略优化(Entropy-Regularized Rank-Masked Policy Optimization, ERPO),该方法通过排序掩码将低PCR转化为保守的负向更新,并利用熵上限控制策略漂移。在代码基准测试中,ERPO在域内适应和零样本迁移场景下均显著提升了pass@1和pass@k性能。
cs.LG / 252 / 2609.09140

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

NOAH:学习完整的患者旅程。一种用于表示与预测的纵向多模态时间感知模型
Susetzky, Tobias, Rehms, Raphael, Seletkov, Dmitrii, Turgut, Özgün, Liman, Michelle Espranita, Steinhelfer, Lisa, Braren, Rickmer, Rueckert, Daniel
Abstract
The digitization of healthcare has generated vast, longitudinal, and multimodal patient records over a lifetime, yet fully exploiting these data to represent and predict patient state trajectories remains a critical challenge. Current AI models often struggle to capture the complex, irregular temporal dynamics and inherent stochasticity of real-world multimodal patient data. Existing AI approaches for modeling longitudinal patient records are predominantly discriminative, limited to a few modalities, constrained by closed categorical vocabularies, treating time as a monotonic inductive bias, or they are limited in forecasting future patient states. We introduce NOAH, a time-aware, task-agnostic, generative transformer model representing and forecasting the full multimodal patient journey. NOAH features a novel bidirectional time integration and a variational latent space to capture the continuous evolution of patient states and the stochasticity of clinical trajectories. Built from over 559 million clinical events from 431,000 hospital visits of 299,000 patients across the MIMIC dataset family, NOAH natively processes medical images, time-series and numeric signals, categorical events, as well as structured and unstructured clinical records. NOAH is the first truly holistic generative model in its field, enabling autoregressive forecasting with optional time control, zero-shot classification, and counterfactual intervention simulation. It generates highly informative and predictive patient state representations that demonstrate strong performance in probing for clinical outcomes, 15 ICD chapters, and 29 comorbidities, as well as in time-to-event prediction. Seamlessly handling diverse modalities and complex temporal dynamics, NOAH provides a versatile, task-agnostic, scalable foundation for intelligent predictive systems in personalized clinical care and digital medicine.
Chinese Translation
医疗数字化已在患者一生中产生了海量、纵向、多模态的患者记录,然而如何充分利用这些数据来表示和预测患者状态轨迹仍然是一个关键挑战。当前的人工智能模型往往难以捕捉真实世界多模态患者数据中复杂、不规则的时间动态性及固有的随机性。现有用于建模纵向患者记录的人工智能方法主要是判别式的,局限于少数模态,受限于封闭的类别词表,将时间仅视为单调的归纳偏置,或在预测未来患者状态方面能力有限。我们提出了NOAH,一个时间感知、任务无关的生成式Transformer模型,用于表示和预测完整的多模态患者旅程。NOAH具有新颖的双向时间积分机制和变分潜在空间,以捕捉患者状态的连续演变以及临床轨迹的随机性。NOAH基于MIMIC数据集家族中299,000名患者、431,000次住院的超过5.59亿条临床事件构建,能够原生处理医学图像、时间序列和数值信号、分类事件以及结构化和非结构化临床记录。NOAH是该领域首个真正整体性的生成式模型,支持带可选时间控制的自回归预测、零样本分类以及反事实干预模拟。它能够生成信息量丰富且具有预测性的患者状态表示,在临床结局探测、15个ICD章节和29种共病的预测以及事件发生时间预测中均展现出强劲的性能。通过无缝处理多种模态和复杂的时间动态,NOAH为个性化临床护理和数字医学中的智能预测系统提供了一个多功能、任务无关且可扩展的基础。
cs.LG / 253 / 2609.09157

Learning Length-Extrapolatable Recurrent Models

学习可长度外推的循环模型
Jiang, Hanwen
Abstract
Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their training horizon. Classical analyses emphasize gradients that vanish or explode along temporal paths. However, dense per-token losses can still train a shared recurrent rule despite severe decay, showing that decay alone does not determine whether learning fails. We instead study state credit: the signal through which future losses reach earlier recurrent states before contributing to parameter updates. Accordingly, we intervene directly on state credit and propose Credit Stabilization through Time (CST). During backward propagation, CST locally rescales the state-credit signal to stabilize its norm without rotating the component being corrected, while leaving the forward computation unchanged. Because controlled synthetic tasks and real data exhibit different credit dynamics, we specialize CST to each regime. In both settings, CST improves performance beyond the training horizon, with gains observed at up to 128x the training length.
Chinese Translation
循环模型为长上下文建模提供了一条自然的路径,然而通过时间反向传播(BPTT)训练的模型在超出训练时域后往往表现失效。经典分析强调沿时间路径的梯度消失或爆炸问题。然而,密集的逐词元损失即使在严重衰减的情况下,仍能训练出一个共享的循环规则,这表明衰减本身并不决定学习是否失败。我们转而研究状态信用(state credit):即未来损失在参与参数更新之前传递到较早循环状态的信号。据此,我们直接对状态信用进行干预,提出了通过时间信用稳定化(Credit Stabilization through Time, CST)。在反向传播过程中,CST 对状态信用信号进行局部重缩放以稳定其范数,同时不旋转正在被校正的分量,且保持前向计算不变。由于受控合成任务与真实数据表现出不同的信用动态,我们将 CST 分别针对这两种情形进行了专门化设计。在两种设置下,CST 都提升了模型在超出训练时域后的性能,在训练长度高达 128 倍的长度上均观察到性能提升。
机器人学 (Robotics)
128
cs.RO / 1 / 2609.05498

Observability Analysis of Joint Steering and Extrinsic Calibration

联合转向与外参标定的可观测性分析
Mishra, Subodh
Abstract
This technical report studies the local weak observability of a planar bicycle-model vehicle when vehicle pose, planar LiDAR extrinsic calibration, and steering-angle bias are estimated jointly. A Lie-derivative-based nonlinear observability analysis is used to examine stationary, straight-line, constant-curvature, and combined straight-plus-arc motion. The resulting observability matrices and nullspaces describe how pose, LiDAR translation and yaw offsets, and steering bias become coupled under different motion primitives. Stationary motion and individual motion primitives retain unobservable directions, whereas the combination of straight and curved motion removes the identified degeneracies and yields full local weak observability of the seven-state system. The analysis provides a theoretical basis for selecting calibration trajectories that sufficiently excite both steering and sensor-extrinsic parameters.
Chinese Translation
本技术报告研究了在同时估计车辆位姿、平面激光雷达(LiDAR)外参标定与转向角偏置时,平面自行车模型的局部弱可观测性。采用基于李导数的非线性可观测性分析方法,考察了静止、直线、恒定曲率以及直线与圆弧组合等运动情形。所得到的可观测性矩阵及其零空间刻画了在不同运动基元下,位姿、激光雷达平移与偏航偏移以及转向偏置之间的耦合方式。静止运动和单一运动基元均保留有不可观测方向,而直线与曲线运动的组合消除了所识别的退化,使该七状态系统具有完全的局部弱可观测性。该分析为选择能够充分激励转向与传感器外参参数的标定轨迹提供了理论基础。
cs.RO / 2 / 2609.05515

Multi-robot Learning-based Informative Path Planning Using Spatio-Temporal Gaussian Process Kalman Filter

基于时空高斯过程卡尔曼滤波的多机器人学习信息路径规划
Cao, Muqing, Lee, Yunwoo, Yuan, Junbin, Schenk, Lorenzo, Scherer, Sebastian
Abstract
Multi-robot informative path planning (IPP) for persistent target monitoring requires robots to reason about spatial uncertainty, temporal evolution, and practical sensing and communication constraints. Recent learning-based multi-robot IPP methods use Gaussian Processes (GPs) for target uncertainty, but often rely on simplified sensing models and centralized belief updates. We propose a grid-based spatio-temporal GP-Kalman filtering framework for learning-based multi-robot IPP. Instead of maintaining one GP per target, we represent anonymous target presence as a single latent field over a discrete workspace grid. The proposed recursive update considers all visible cells inside a camera footprint and supports arbitrary fields of view and range-dependent noise. A GP-consistent temporal process update accounts for moving targets and stale information by inflating uncertainty over time. For decentralized deployment, each robot maintains its own mapper and exchanges compact belief summaries rather than raw measurements. Received beliefs are fused using diagonal covariance intersection to remain conservative under unknown inter-robot correlations. We integrate the mapper with a reinforcement-learning policy for graph-based neighbor selection. Simulation benchmarks show about 20% lower average target uncertainty and improved target visitation compared with learning-based and classical auction/coverage baselines. Real-world two-UAV experiments demonstrate transfer to outdoor multi-robot search over a large field of more than 7000 square meters.
Chinese Translation
面向持续目标监测的多机器人信息路径规划(Informative Path Planning, IPP)要求机器人能够综合考虑空间不确定性、目标的时间演化,以及实际的传感与通信约束。近年来基于学习的多机器人IPP方法通常使用高斯过程(Gaussian Processes, GPs)来刻画目标的不确定性,但往往依赖简化的传感模型和集中式信念更新。我们提出了一种基于网格的时空GP-卡尔曼滤波框架,用于基于学习的多机器人IPP。不同于为每个目标维护一个高斯过程,我们将匿名的目标存在性表示为离散工作空间网格上的单一潜在场。所提出的递归更新机制能够考虑相机视野内的所有可见网格单元,并支持任意视场形状和随距离变化的噪声。采用与GP一致的时间过程更新,通过随时间膨胀不确定性来处理移动目标和过时信息。在分布式部署方面,每个机器人维护自己的地图构建器,并交换紧凑的信念摘要而非原始测量数据。接收到的信念通过对角协方差交叉(diagonal covariance intersection)方法进行融合,从而在机器人间相关性未知的情况下保持保守估计。我们将该地图构建器与强化学习策略集成,实现基于图的邻居选择。仿真基准实验表明,与基于学习的方法以及经典的拍卖/覆盖类基线相比,我们的方法将平均目标不确定性降低了约20%,并提升了目标访问率。真实世界的双无人机实验验证了该方法可迁移至超过7000平方米的大型户外场地多机器人搜索任务。
cs.RO / 3 / 2609.05519

Robots Influencing Humans to Reveal their Goals during Collaboration and Competition

机器人在合作与竞争中引导人类揭示其目标
Ghose, Debasmita, Gitelson, Oz, Lewkowicz, Michal, Brawer, Jake, Vazquez, Marynel, Scassellati, Brian
Abstract
We propose a unified strategy for fast goal inference in human-robot interaction. The core idea is to drive the human toward Critical Decision Points (CDPs)-states where competing human strategies prescribe different next actions and thus maximally reveal the goal. We formalise CDPs using a goal-conditioned policy divergence measure and incorporate them into a Receding-Horizon Planner that explores future action sequences while optimizing a cost function balancing task progress and information gain. We evaluate this approach in both a collaborative, fully observable cooking task and a competitive, partially observable hide-and-seek game, each in simulation and on real robots. In both scenarios, our method infers human goals more accurately and earlier than baseline strategies.
Chinese Translation
我们提出了一种用于人机交互中快速目标推断的统一策略。其核心思想是引导人类走向关键决策点(Critical Decision Points, CDPs)——在这些状态下,人类的不同竞争性策略会规定不同的下一步动作,从而最大程度地揭示其目标。我们使用一种目标条件策略散度度量对关键决策点进行形式化定义,并将其融入滚动时域规划器(Receding-Horizon Planner)中,该规划器在优化一个平衡任务进度与信息增益的代价函数的同时,探索未来的动作序列。我们在协作性的完全可观测烹饪任务和竞争性的部分可观测捉迷藏游戏中评估了该方法,每种任务均在仿真环境和真实机器人上进行了实验。在两种场景中,我们的方法都比基线策略更准确、更早地推断出人类目标。
cs.RO / 4 / 2609.05521

Situation Awareness for Intelligent Data Distribution in Connected Vehicles

面向智能网联汽车数据分发的态势感知
Dettinger, Falk, Narla, Akshay, Weyrich, Michael
Abstract
The limitations of on-board sensors and blind spots caused by occlusion cause the reduction of perception quality in autonomous vehicles. In such cases, cooperative perception provides additional data via Vehicle-to-Everything communication to enhance local perception, causing a large volume of data transmission. The vehicle can focus on acquiring and utilizing relevant data according to the prevailing road context by identifying the current traffic situation. To achieve this, we propose a concept for the situation identification of the vehicle using Bird's-Eye-View images. Firstly, the situation around the vehicle is identified using object detection with semantic segmentation, followed by understanding the context of the traffic using a situation identification module consisting of an open-source projective transformation network Cam2BEV and a situation identification neural network. The concept was evaluated and validated by running the software on the CARLA simulator using the in-built RGB camera and the semantic segmentation camera. Additionally, the portability of the situation identification module for real-world applications was verified on Cityscapes and nuScenes urban driving datasets. Overall, the proposed situation identification approach enables efficient sensor data management by prioritizing relevant data to the current traffic situation. The source code is available in the following link: https://github.com/akshaynarla/DySi_Select
Chinese Translation
车载传感器的局限性以及遮挡造成的盲区会导致自动驾驶车辆感知质量的下降。在这种情况下,协同感知通过车联网通信提供额外数据以增强本地感知,但会造成大量数据的传输。通过识别当前交通态势,车辆可以根据实际道路情境专注于获取和利用相关数据。为此,我们提出了一种基于鸟瞰图图像的车辆态势识别方法。首先,利用结合语义分割的目标检测识别车辆周围的态势,随后通过由开源投影变换网络Cam2BEV和态势识别神经网络组成的态势识别模块来理解交通情境。该概念通过在CARLA仿真器上使用内置的RGB相机和语义分割相机运行软件进行了评估和验证。此外,还在Cityscapes和nuScenes城市驾驶数据集上验证了该态势识别模块在实际应用中的可移植性。总体而言,所提出的态势识别方法通过对与当前交通态势相关的数据进行优先级排序,实现了高效的传感器数据管理。源代码可通过以下链接获取:https://github.com/akshaynarla/DySi_Select
cs.RO / 5 / 2609.05569

Information-Guided Safe Reinforcement Learning for Autonomous Gas Source Localization using sUAS

基于小型无人机系统的信息引导安全强化学习气体源自主定位方法
Giri, Sachin, Zhao, Thomas, Huynh, Matthew, Chen, YangQuan
Abstract
The autonomous localization of fugitive gas emissions using small Unmanned Aircraft Systems (sUAS) constitutes a fundamentally ill-posed inverse problem. In turbulent atmospheric boundary layers, highly intermittent scalar concentration fields violate the assumptions of classical gradient-based navigation, causing data-driven estimators to suffer from severe noise and spurious local minima. To address these challenges, we introduce an Information-Guided Safe Reinforcement Learning framework evaluated within a custom, GPU-accelerated 3D simulation environment coupling an Eulerian wind solver with a Lagrangian puff dispersion model. We identify a critical vulnerability in deterministic information-seeking planners - a Gramian bias where agents act greedily upon flawed early estimates, starving the estimator of spatial diversity. To systematically break this degeneracy, our architecture integrates a classical empirical observability Gramian (EMGR) planner with a learned Soft Actor-Critic (SAC) exploratory policy. A deterministic meta-supervisor actively monitors estimator reliability via Kullback-Leibler (KL) divergence, dynamically blending deterministic exploitation with learned exploration to steer the sUAS into high-information zones. Trained via a progressive curriculum and safeguarded by a strictly enforced Robust Control Barrier Function (RCBF), our RL framework achieves nearly 80% localization success on complex, mobile sources - drastically outperforming classical baselines (~30%) - while ensuring zero safety violations.
Chinese Translation
利用小型无人飞行系统(sUAS)对逃逸性气体排放进行自主定位,本质上是一个不适定的逆问题。在湍流大气边界层中,高度间歇性的标量浓度场违背了经典基于梯度导航方法的假设,导致数据驱动的估计器受到严重噪声干扰并陷入虚假局部极小值。为应对这些挑战,我们提出了一种信息引导的安全强化学习框架,并在一个定制的GPU加速三维仿真环境中进行评估,该环境将欧拉风场求解器与拉格朗日烟团扩散模型相耦合。我们发现了确定性信息搜索规划器的一个关键脆弱性——即格拉姆矩阵偏差(Gramian bias),即智能体基于有缺陷的早期估计进行贪婪行动,导致估计器缺乏空间多样性。为系统地打破这种退化,我们的架构将经典的经验可观测性格拉姆矩阵(EMGR)规划器与学习到的Soft Actor-Critic(SAC)探索策略相结合。一个确定性的元监督器通过Kullback-Leibler(KL)散度主动监测估计器的可靠性,动态地融合确定性利用与学习到的探索,引导sUAS进入高信息量区域。通过渐进式课程训练,并辅以严格执行的鲁棒控制障碍函数(RCBF)安全保障,我们的强化学习框架在复杂移动源上实现了近80%的定位成功率——大幅优于经典基线方法(约30%)——同时确保零安全违规。
cs.RO / 6 / 2609.05585

Benchmarking Dexterity of Multifingered Robot Hands: A Review and Perspective

多指机器人手灵巧性基准测试:综述与展望
Shilati, Anthony, Ramaswami, Anunth, Batteas, Luke, Tan, Sylvia, Barcio, Anthony, Umakanth, Sairam, Rao, Preksha, McDougall, David, Graves, Landry, Ozkan, Ahmet A., Yoon, Yunsoo, Pradhan, Arushi, Henry, Michael G., Kota, Rohan, Gonzalez, Damian, Thomas, Gray C., Fedder, Gary K., Colgate, J. Edward, Lynch, Kevin M.
Abstract
Robot hands are a key interface between AI and the physical world, making advances in robotic dexterity essential to realizing the vision of physical AI. While impressive dexterity has been demonstrated with simple grippers, multifingered hands offer the potential for substantially greater versatility, precision, and adaptability in manipulation. In this review, we survey the state of the art in benchmarking the dexterity of multifingered robot hands. Recognizing dexterity as a complex and multifaceted concept, we present the perspective of the U.S. National Science Foundation HAND Engineering Research Center, with a particular focus on fine in-hand manipulation. We introduce a framework consisting of three benchmark levels that correspond to increasing system complexity, review representative benchmarks at each level, and propose new benchmarks and metrics to address limitations in the literature. More information can be found at https://hand-erc.github.io/benchmarking/.
Chinese Translation
机器人手是人工智能与物理世界之间的关键接口,因此机器人灵巧性的进步对于实现物理人工智能(Physical AI)的愿景至关重要。尽管简单夹持器已展现出令人瞩目的灵巧性,但多指机器手有望在操作的多样性、精度和适应性方面提供显著更大的潜力。在本综述中,我们对多指机器人手灵巧性基准测试的最新研究进展进行了调研。鉴于灵巧性是一个复杂且多层面的概念,我们提出了美国国家科学基金会HAND工程研究中心的观点,并特别关注精细的手内操作。我们引入了一个包含三个基准层级的框架,这些层级对应于不断增长的系统复杂性,回顾了每个层级的代表性基准测试,并提出了新的基准测试和指标以弥补现有文献中的不足。更多信息请参见 https://hand-erc.github.io/benchmarking/。
cs.RO / 7 / 2609.05588

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

GE-Act 2.0:面向机器人操作的世界-动作模型的预训练与扩展
AgiBot Research Team, Liu, Renhang, Zhao, Wenzhi, Yang, Zhuo, Chen, Liliang, Zhou, Pengfei, Chen, Shengcong, Ren, Guanghui, Peng, Youlun, Jin, Rongjun, Wang, Nan, Wang, Sukai, He, Xindong, Feng, Jinyuan, Xiong, Ziyu, Zhong, Linqing, Wei, Yifei, Han, Feng, Zhang, Long, Huang, Da, Zhao, Nanshu, Yin, Chenghao, Wu, Mo, Yan, Zhaodong, Hu, Kongtao, Yan, Yuxiang, Niyazi, Aogelijiang, Fang, Yu, Zeng, Jia, Meng, Lizhu, Lv, Daizhen, Cao, Haoyu, Hou, Zhiwen, Ye, Lianjin, Niu, Yuehan, Cai, Zhikai, Hu, Xuan, Min, Hui, Cai, Xiongfeng, Liao, Yue, Wu, Jing, Poria, Soujanya, Li, Ye, Zhou, Sanping, Yao, Maoqing
Abstract
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose trainable generative and action components are all initialized from scratch on manipulation data. It combines a control-oriented autoencoder (CoAE), a single-step visual planner (SVP), and an inverse dynamics model (IDM). CoAE retains action- and instruction-relevant information under aggressive compression, while SVP produces a complete future state in one differentiable pass, so visual planning and inverse dynamics can be pretrained separately on complementary data. The components are then jointly trained with knowledge-aligned selective optimization (KASO), which reduces mismatched supervision by selecting only predicted futures judged behaviorally compatible with the recorded action. We evaluate pretrained checkpoints directly, without per-task fine-tuning, on 100 tasks across 20 manipulation skill groups with held-out scenes, backgrounds, lighting, and object instances. Scaling co-training data from 300 to 30,000 hours raises success from 17.1% to 44.1% on G1-OP and from 13.4% to 31.1% on G2-90D; despite comprising less than 2% of the co-training data, G2-90D improves by 17.7 points, suggesting cross-embodiment transfer. Gains span 19/20 and 18/20 skill groups, and skill-specific coverage strongly correlates with zero-shot out-of-distribution (OOD) success (Pearson r=0.80; Spearman rho=0.85). Under the same protocol, the model grounds object, color, shape, and position references in at least 90% of trials and follows explicit instructions even when they conflict with an already-committed behavior or a conventional scene association.
Chinese Translation
世界-动作模型(World-Action Model, WAM)通过预测未来状态来引导机器人动作,使其能够同时从无动作标签的视频和带有动作标签的交互数据中学习。现有大多数方法继承预训练的视频生成器,导致WAM的预训练与规模化扩展仍缺乏深入研究。我们提出Genie Envisioner Act 2.0(GE-Act 2.0),这是一个世界-动作模型,其可训练的生成组件与动作组件均基于操作数据从零初始化。该模型结合了面向控制的自动编码器(CoAE)、单步视觉规划器(SVP)和逆动力学模型(IDM)。CoAE在激进压缩下仍保留动作和指令相关的信息;SVP通过一次可微分前向传播生成完整的未来状态,使视觉规划与逆动力学可以在互补数据上分别进行预训练。随后,各组件通过知识对齐的选择性优化(KASO)进行联合训练,该方法仅选择被判定为与记录动作行为兼容的预测未来,从而减少监督信号的不匹配。我们在100个任务(涵盖20个操作技能组)上直接评估预训练检查点,无需针对每个任务进行微调,测试场景在背景、光照和物体实例上均与训练数据分离。将协同训练数据从300小时扩展至30,000小时,使G1-OP上的成功率从17.1%提升至44.1%,G2-90D上的成功率从13.4%提升至31.1%;其中G2-90D仅占协同训练数据不到2%,却提升了17.7个百分点,这表明存在跨具身迁移。性能提升覆盖19/20和18/20个技能组,且技能特定的数据覆盖度与零样本分布外(OOD)成功率强相关(Pearson r=0.80;Spearman rho=0.85)。在同一协议下,该模型在至少90%的试验中能正确理解物体、颜色、形状和位置的指代,并且即使显式指令与已执行的行为或常规场景关联相冲突时仍能遵循该指令。
cs.RO / 8 / 2609.05593

Adaptive Cost-Sensitive Machine Learning for Autonomous Robot Navigation Failure Prediction: When Not All Errors Are Equal

面向自主机器人导航失效预测的自适应代价敏感机器学习:当并非所有错误都同等重要时
Ferzana, Rifa
Abstract
Autonomous robot navigation failures differ not only in categorical severity but also in the physical context in which they occur. A near-miss at low speed under reliable sensing is not equivalent to the same event during rapid motion, close obstacle approach or degraded perception. This paper reframes navigation failure prediction as consequence-sensitive forecasting. We first establish a fixed baseline in which training weights are modulated by categorical severity, then introduce an adaptive extension defining a state-dependent consequence function combining severity with normalised velocity, obstacle proximity and sensing uncertainty, together with a risk-sensitivity term that rises as conditions deteriorate. We evaluate on 2,000 simulated differential-drive episodes (~1,000,000 timesteps) using episode-level GroupKFold, with external validation on the UCI SCITOS G5 dataset. Fixed weighting raises Logistic Regression high-severity recall from 0.851 to 0.985 and reduces missed consequence cost from 1,940 to 313; the adaptive extension reaches 0.998 and 82. Under matched false-positive conditions, however, the discriminative advantage is modest (0.986 versus 0.984), so most of the gain reflects a more conservative operating point rather than better ranking. The effect is consistent across all five folds and stable across a threefold span of context coefficients. Because the primary simulation produced no collisions, we add a controlled extension in which 108 of 600 episodes terminate in contact: collision recall rises from 0.850 to 0.966 (fixed) and 0.984 (adaptive), with missed collision cost falling from 1,000 to 105, at false-positive rates of 0.413 and 0.799, respectively. Context-dependent consequence modelling thus provides a principled mechanism for allocating conservatism by physical risk.
Chinese Translation
自主机器人的导航失效不仅在严重程度的类别上存在差异,还因其发生的物理环境不同而有所区别。在可靠感知条件下低速行驶时的险情,与在快速运动、接近障碍物或感知退化时发生的同类事件并不等同。本文将导航失效预测重新表述为后果敏感的预测问题。我们首先建立一个固定基线,其中训练权重由严重程度的类别进行调节;随后引入一种自适应扩展方法,定义了一个状态相关的后果函数,将严重程度与归一化速度、障碍物接近程度和感知不确定性相结合,并加入一个随环境条件恶化而上升的风险敏感项。我们在2,000个模拟差速驱动回合(约1,000,000个时间步)上进行评估,采用回合级GroupKFold交叉验证,并在UCI SCITOS G5数据集上进行外部验证。固定加权方法将逻辑回归(Logistic Regression)对高严重程度的召回率从0.851提升至0.985,并将漏检后果成本从1,940降至313;自适应扩展方法进一步达到0.998和82。然而,在误报率相同的条件下,两种方法的判别优势差距较小(0.986对0.984),因此大部分收益来源于更为保守的运行点选择,而非排序能力的提升。该效应在全部五个折中保持一致,并在三倍跨度的情境系数范围内保持稳定。由于初始模拟未产生任何碰撞,我们增加了一个受控扩展实验,其中600个回合中有108个以碰撞终止:碰撞召回率分别提升至0.966(固定方法)和0.984(自适应方法),漏检碰撞成本从1,000降至105,对应的误报率分别为0.413和0.799。因此,情境相关的后果建模为按物理风险分配保守程度提供了一种有原则的机制。
cs.RO / 9 / 2609.05783

Closed-Loop Evaluation of Bird's-Eye-View Maps from Cross-View Transformers as Inputs to Behavior-Cloning Policies

基于跨视图Transformer生成的鸟瞰图作为行为克隆策略输入的闭环评估
Santos, Felipe Carlos dos, Antonelo, Eric, Couto, Gustavo Claudio Karl
Abstract
In autonomous driving, Bird's-Eye View (BEV) representations provide a structured, top-down abstraction of the vehicle's surroundings and have become a key input modality for Behavioral Cloning (BC) policies. While ground-truth BEV maps are readily available in simulation, real-world deployment requires replacing them with camera-predicted counterparts - a substitution that introduces perceptual errors whose downstream impact on closed-loop driving performance is not well understood. In this work, we investigate the use of Cross-View Transformer (CVT)-predicted BEV maps as direct policy inputs for a BC agent in the CARLA simulator. We propose a six-channel BEV representation covering road surface, planned route, lane boundaries, vehicles, pedestrians, and traffic lights, and introduce a Kernel Density Estimation (KDE) weighting scheme that rebalances the segmentation loss towards underrepresented driving maneuvers such as curves and intersections. Closed-loop evaluation across two CARLA towns shows that the KDE-weighted model is the only predicted-BEV agent to complete a full episode without infractions, despite not achieving the highest aggregate IoU. This discrepancy reveals that global segmentation metrics are poor proxies for driving performance: what determines navigation success is prediction quality at geometrically critical locations, and the route channel emerges as the primary bottleneck for reliable agent navigation under predicted BEV inputs.
Chinese Translation
在自动驾驶领域,鸟瞰图(Bird's-Eye View, BEV)表示为车辆周围环境提供了一种结构化的俯视抽象描述,已成为行为克隆(Behavioral Cloning, BC)策略的关键输入模态。虽然在仿真环境中真值BEV地图易于获取,但现实部署中必须用相机预测的BEV地图替代,这一替代引入了感知误差,而这些误差对闭环驾驶性能的下游影响尚不清楚。本研究在CARLA仿真器中,探究了将跨视图Transformer(Cross-View Transformer, CVT)预测的BEV地图直接作为BC智能体策略输入的效果。我们提出了一种六通道BEV表示,涵盖路面、规划路线、车道边界、车辆、行人及交通信号灯,并引入一种核密度估计(Kernel Density Estimation, KDE)加权方案,将分割损失向转弯和交叉口等欠代表驾驶场景重新平衡。在两个CARLA城镇中的闭环评估表明,KDE加权模型是唯一能够完整完成一个回合且无违规行为的预测BEV智能体,尽管其总体IoU并非最高。这一差异揭示了全局分割指标无法很好地替代驾驶性能评估:决定导航成功与否的是几何关键位置处的预测质量,且路线通道被证明是在预测BEV输入下实现可靠智能体导航的主要瓶颈。
cs.RO / 10 / 2609.05832

CR-VLA-Force: Learning Control-aware Compliance VLA Model for Robust Contact-rich Robotic Manipulation

CR-VLA-Force:面向鲁棒接触丰富机器人操作的学习控制感知柔顺VLA模型
Mai, Zhaohong, Wang, Chao, Zeng, Chao, Mao, Sitong, Zhang, Heng, Zhou, Shunbo, Yang, Chenguang
Abstract
Integrating visuomotor policies or Vision-Language-Action (VLA) models with force/torque (F/T) perception has demonstrated significant progress in imitation learning for robotic manipulation. However, existing force-aware VLA models frequently exhibit limited capability in precise force tracking and rapid successive adjustments. This deficiency stems from the limitations of action-chunk execution strategies and the substantial latency between perception and real-time control. Such limitations can lead to task failures and safety risks, particularly when the execution of an action chunk exerts excessive interaction forces without timely adjustment. To overcome this challenge, we propose the Control-aware Compliance VLA (CC-VLA) framework for reactive control. The CC-VLA model employs a multimodal mixture-of-experts (MoE) to encode force signal sequences and vision-language fused feature. Furthermore, it utilizes a multi-stage training strategy to ensure robust perception within the visual-semantic space and effective force perception under sparse sampling conditions. Additionally, a VLA-guided adaptive compliance controller is designed to facilitate precise position tracking during contact-free motion and optimal force-position tracking for contact-rich tasks. To facilitate high-precision F/T data acquisition, we also implement an adversaria shared teleoperation strategy for contact-rich demonstrations that bolsters system safety and interactivity. Extensive real-world experiments demonstrate that CC-VLA significantly improves success rates in challenging force-perception tasks and enhances force-control precision, while providing multi-level safety and robustness under the tested partial-OOD pose-shift settings.
Chinese Translation
将视觉运动策略或视觉-语言-动作(VLA)模型与力/力矩(F/T)感知相结合,已在机器人操作的模仿学习中取得显著进展。然而,现有的力感知VLA模型在精确力跟踪和快速连续调整方面能力有限。这一缺陷源于动作块(action-chunk)执行策略的局限性以及感知与实时控制之间的较大延迟。此类局限可能导致任务失败和安全风险,尤其是在动作块执行过程中产生过大的交互力而未能及时调整时。为克服这一挑战,我们提出了用于反应式控制的控制感知柔顺VLA(CC-VLA)框架。CC-VLA模型采用多模态混合专家(MoE)结构来编码力信号序列和视觉-语言融合特征。此外,该模型利用多阶段训练策略,以确保在视觉-语义空间内实现鲁棒感知,并在稀疏采样条件下实现有效的力感知。同时,我们设计了一个VLA引导的自适应柔顺控制器,以在无接触运动阶段实现精确的位置跟踪,并在接触丰富任务中实现最优的力-位置跟踪。为便于获取高精度力/力矩数据,我们还实现了一种对抗性共享遥操作策略,用于接触丰富示范数据的采集,以增强系统的安全性和交互性。大量真实世界实验表明,CC-VLA在具有挑战性的力感知任务中显著提高了成功率并增强了力控精度,同时在所测试的部分分布外(partial-OOD)位姿偏移设置下提供了多层次的安全性和鲁棒性。
cs.RO / 11 / 2609.05892

A4A: Cross-Embodiment Transfer of Action-Oriented 4D Affordances from Human Demonstrations

A4A:从人类演示中跨本体迁移面向动作的4D可供性
Han, Yifan, Liu, Litao, Gu, Yuqi, Lu, Ye, Wang, Hanqing, Wai, Sidney, Myrie, Ishaan, Zhang, Qi, Yu, Jingjin, Li, Gen
Abstract
Human demonstrations contain rich manipulation knowledge, but it remains unclear what information can be transferred effectively to robot control. Existing affordance representations are typically formulated as 2D masks, 3D regions, contact points, or actionability scores, and therefore primarily identify where interaction may occur. However, effective manipulation also requires modeling how interaction-relevant geometry evolves during task execution. To bridge this gap, we introduce action-oriented 4D affordances, which represent the language-conditioned future trajectories of interaction-relevant 3D points. These trajectories capture task-conditioned geometric evolution rather than embodiment-specific actions, enabling transferable interaction priors across humans and robots. Based on this representation, we construct a large-scale action-oriented 4D affordance dataset from existing human--object interaction video data and complementary RGB-D demonstrations, and introduce A4A, an affordance-to-action framework that uses 4D affordance trajectory prediction to pretrain robot policies before manipulation fine-tuning. Experiments in both simulation and the real world validate the effectiveness of A4A, showing that pretraining with action-oriented 4D affordance data consistently improves the manipulation performance of diverse VLA policies. These results establish action-oriented 4D affordances as an effective cross-embodiment representation for transferring manipulation knowledge from human demonstrations to robot control.
Chinese Translation
人类演示蕴含丰富的操作知识,但哪些信息能够有效迁移至机器人控制仍不明确。现有的可供性表征通常被形式化为2D掩码、3D区域、接触点或可操作性评分,因而主要识别交互可能发生的位置。然而,有效的操作还需要对任务执行过程中与交互相关的几何形状如何演化进行建模。为弥合这一差距,我们提出了面向动作的4D可供性,它表示与交互相关的3D点在语言条件下的未来轨迹。这些轨迹捕捉的是任务条件下的几何演化,而非特定本体的动作,从而实现了在人类与机器人之间可迁移的交互先验。基于这一表征,我们从现有人类-物体交互视频数据和补充的RGB-D演示中构建了一个大规模的面向动作4D可供性数据集,并提出了A4A——一个从可供性到动作的框架,该框架利用4D可供性轨迹预测对机器人策略进行预训练,随后再进行操作微调。在仿真和真实世界中的实验验证了A4A的有效性,表明使用面向动作4D可供性数据的预训练能够持续提升多种VLA策略的操作性能。这些结果确立了面向动作4D可供性作为一种有效的跨本体表征,可将人类演示中的操作知识迁移至机器人控制。
cs.RO / 12 / 2609.05927

GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning

GIF:面向机器人学习的交互性与功能性物体组合的智能体式生成方法
Xu, Long, Zhang, Zhiqi, Yan, Mi, Deng, Shengliang, Xia, Chong, Dong, Mingyu, Chen, Jiayi, Lyu, Jiangran, Gao, Fei, Zhang, Zhizheng, Wang, He
Abstract
Robot manipulation foundation models require scalable evaluation and data generation across diverse scenarios, with simulation providing an environment for both. Automated scene generation offers a promising path, yet prior work has largely emphasized coarse-grained scene layouts rather than fine-grained functional object compositions. Motivated by this gap, we present GIF, an agentic Generation framework for Interactive and Functional object compositions. In this framework, we recast this problem as disentangled reconstruction followed by relative pose recovery. CoGen produces instance-disentangled meshes with coarse initial poses leveraging complementary strengths of 2D and 3D generative models. GPRM refines the relative pose under joint geometric and physical guidance, and a VLM verifier selects the candidate that best matches the structured specification. We further construct a benchmark spanning eight representative contact-geometry classes and compare with state-of-the-art generators; GIF improves both asset quality and relation matching, while reducing collision rate to below 1%. Finally, we synthesize data for policy learning, revealing diversity scaling in both simulation and real-world deployment.
Chinese Translation
机器人操作基础模型需要在多样化场景中进行可扩展的评估与数据生成,而仿真环境可为二者提供支撑。自动化场景生成是一条有前景的路径,但已有工作大多侧重于粗粒度的场景布局,而非细粒度的功能性物体组合。基于这一空白,我们提出了 GIF——一个用于生成交互性与功能性物体组合的智能体式框架。在该框架中,我们将这一问题重新表述为解耦重建加相对位姿恢复两个阶段。CoGen 利用 2D 与 3D 生成模型的互补优势,生成实例解耦的网格及粗略初始位姿;GPRM 在几何与物理的联合引导下对相对位姿进行细化;随后由一个 VLM 验证器选出最符合结构化规格的候选结果。我们进一步构建了一个涵盖八类代表性接触几何类型的基准,并与最先进的生成方法进行了比较;GIF 在资产质量与关系匹配两方面均有提升,同时将碰撞率降至 1% 以下。最后,我们合成了用于策略学习的数据,揭示了多样性在仿真与真实世界部署中的扩展效应。
cs.RO / 13 / 2609.05941

Moment-Matching Probabilistic Data Association for Optimization-Based SLAM

基于优化的SLAM中矩匹配概率数据关联方法
Nguyen, Khoa, Turton, Mitchell, Meyer, Florian
Abstract
Optimization-based simultaneous localization and mapping (SLAM) makes it possible to reduce accumulated navigation errors of sensing platforms by returning to known areas (loop closure). In this paper, we present an approach to combine probabilistic data association (PDA) with optimization-based SLAM. Instead of associating a single measurement with each landmark, we follow the PDA paradigm from the multiobject tracking community. In particular, in a processing stage performed in addition to the nonlinear least-squares solver of optimization-based SLAM, our method (i) assigns multiple measurements to landmarks probabilistically, (ii) computes the mean and covariance of landmark distributions via moment matching by taking multiple measurement-to-landmark associations into account, and (iii) establishes a virtual landmark measurement and a corresponding linear-Gaussian measurement model that leads to the mean and covariance matrix as moment-matching PDA in (ii). By converting the PDA update step into an equivalent linear-Gaussian measurement update step, PDA can be performed effectively within any optimization-based SLAM method. Our preliminary numerical evaluation in a scenario with false negatives and false positives indicates that incremental smoothing and mapping 2 (iSAM2), combined with the proposed PDA approach, can improve agent localization performance compared to conventional iSAM2.
Chinese Translation
基于优化的同步定位与建图(SLAM)使得感知平台能够通过返回已知区域(回环闭合)来减少累积的导航误差。本文提出一种将概率数据关联(PDA)与基于优化的SLAM相结合的方法。与为每个路标仅关联单个量测不同,我们遵循多目标跟踪领域的PDA范式。具体而言,在基于优化的SLAM的非线性最小二乘求解器之外增加的一个处理阶段中,我们的方法:(i)以概率方式为路标分配多个量测;(ii)通过矩匹配,在考虑多个量测到路标关联的情况下,计算路标分布的均值和协方差;(iii)建立一个虚拟路标量测及相应的线性高斯量测模型,该模型产生的均值和协方差矩阵与(ii)中矩匹配PDA的结果一致。通过将PDA更新步骤转换为等价的线性高斯量测更新步骤,PDA可以有效地嵌入任何基于优化的SLAM方法中。我们在包含假阴性和假阳性的场景中进行的初步数值评估表明,与传统的iSAM2相比,结合所提出PDA方法的增量平滑与建图2(iSAM2)能够提升智能体定位性能。
cs.RO / 14 / 2609.05962

Observation Design for Certified Control Authority: Projection--Estimability Separation and Active-Face Equivalence

面向认证控制权限的观测设计:投影—可估性分离与活动面等价
Wan, Guangxi, Du, Hualong, Liu, Yuqi, Dong, Qingwei, Li, Qingxin, Bai, Hongfei, Zeng, Peng
Abstract
A sound runtime admission gate executes only actions it can certify, and certifies only what its observations support. This paper asks how observations should be designed to maximize the set of actions that can be safely admitted, and shows the question is not a re-vocabulary of classical design problems. First, a projection--estimability separation: decomposing a constraint normal as $c=c_{\mathrm{Range}}+c_{\ker}$ relative to an information matrix, two-point discrimination along $c$ becomes arbitrarily reliable as the budget grows whenever $c_{\mathrm{Range}}\neq 0$, while robust admission of an action with normal $c$ is impossible at every budget whenever $c_{\ker}\neq 0$; such mixed directions are generic at any deficient rank, and at full rank the decoupling is bounded by the Kantorovich ratio and diverges with the condition number. Discrimination-optimal designs, being corner solutions of a linear criterion, land in exactly this regime. Second, an active-face equivalence theorem: the permissiveness-optimal design is $L$-optimal for a target matrix generated endogenously by the action faces that become certification bottlenecks, weighted inversely by their remaining slack; single-face collapse recovers $c$-optimal and goal-oriented design exactly, and a Caratheodory argument yields a bottleneck certificate of at most $r(r+1)/2+1$ faces. Around these we assemble exact certifiability per convex contract mode, for which $\kappa\approx 3.29$ is derived rather than calibrated, the $\sqrt{r}$ price of contract-agnostic design, an information-to-slack transfer theorem with a curvature-budget corollary, and a two-part audit in which a design meeting every margin requirement still leaves an action face at a certification cost above ten times its testing cost, in every probe library tested.
Chinese Translation
一个可靠的运行时准入门控只执行其能够认证的动作,并且只认证其观测所能支持的内容。本文探讨应如何设计观测以最大化可安全准入的动作集合,并表明该问题并非经典设计问题的重新表述。首先,提出投影—可估性分离:将约束法向量相对于信息矩阵分解为 $c=c_{\mathrm{Range}}+c_{\mathrm{ker}}$,当 $c_{\mathrm{Range}}\neq 0$ 时,沿 $c$ 的两点判别随预算增长可达到任意可靠;而当 $c_{\mathrm{ker}}\neq 0$ 时,以 $c$ 为法向量的动作在任何预算下都无法实现鲁棒准入。此类混合方向在任意秩亏情形下都是普遍存在的;在满秩情形下,该解耦程度由 Kantorovich 比值界定,并随条件数发散。判别最优设计作为线性准则的角点解,恰好落入这一区域。其次,提出活动面等价定理:宽松度最优的设计对某个目标矩阵是 $L$-最优的,该目标矩阵由成为认证瓶颈的动作面内生生成,并按其剩余松弛量的倒数加权;单面坍缩情形可精确恢复 $c$-最优与面向目标的设计,且借助 Carathéodory 论证可得到至多含 $r(r+1)/2+1$ 个面的瓶颈证书。围绕这些结果,我们构建了逐凸契约模式的精确可认证性,其中 $\kappa\approx 3.29$ 是推导得出而非校准得到的;给出了契约无关设计的 $\sqrt{r}$ 代价;建立了信息—松弛转移定理及其曲率—预算推论;并在所测试的每个探针库中,通过两部分审计表明:满足所有裕度要求的设计,仍可能使某一动作面的认证成本高于其测试成本的十倍。
cs.RO / 15 / 2609.05985

A Brain-inspired Hierarchical Framework for Zero-Shot Robot Task Reasoning and Execution

一种面向零样本机器人任务推理与执行的类脑分层框架
Wang, Guangming, Ye, Pengfei, Ying, Qizhen, Jing, Yixiong, Ma, Yuxiang, Chen, Haonan, Wu, Haibing, Wysocki, Olaf, Duan, Molong, Sheil, Brian
Abstract
Robots that follow open-ended language instructions need to connect semantic intent to visual scene understanding, geometric feasibility, object states, and physical interaction conditions. End-to-end Vision-Language-Action policies have improved cross-task generalization, but they typically map visual and language inputs directly to robot actions, leaving limited explicit structure for long-horizon decomposition, physical verification, and recovery. We present \method, a zero-shot hierarchical framework functionally inspired by the division of roles in the human brain, comprising visual perception and state inference, language grounding and action-sequence generation from a shared atomic action library, cost-based plan selection, and real-robot execution and verification. The framework grounds commands in explicit object states, composes reusable atomic actions into task-conditioned sequences, ranks alternative sequences by execution cost, and verifies intermediate physical outcomes from refreshed observations. In the evaluation, \method{} completes 10/10 clean board trials, 10/10 pick-and-place trials, and 4/5 pyramid stacking trials for both the flat and irregular initial-layout conditions; the corresponding mean task progress is $99.03\%$, $100.00\%$, and $96.67\%$ respectively. Across all evaluated conditions, \method{} achieves higher success rates than ReKep, Dream2Flow, and $\pi_{0.5}$ benchmarks, demonstrating the effectiveness of combining explicit object-state reasoning, compositional atomic actions, cost-based plan selection, and closed-loop execution verification.
Chinese Translation
能够遵循开放式语言指令的机器人需要将语义意图与视觉场景理解、几何可行性、物体状态以及物理交互条件相连接。端到端的视觉-语言-动作(Vision-Language-Action)策略已经改善了跨任务泛化能力,但它们通常将视觉和语言输入直接映射为机器人动作,对于长时程任务分解、物理验证与恢复而言,缺乏足够的显式结构。我们提出了\method,一个受人类大脑角色分工功能启发的零样本分层框架,包含视觉感知与状态推断、基于共享原子动作库的语言落地与动作序列生成、基于代价的规划选择,以及真机执行与验证。该框架将指令落地为显式的物体状态,将可复用的原子动作组合为任务条件化的序列,依据执行代价对候选序列进行排序,并从更新的观测中验证中间物理结果。在评估中,无论初始布局是平坦还是不规则条件,\method{}在擦黑板任务中完成10/10次试验,在抓取放置任务中完成10/10次试验,在金字塔堆叠任务中完成4/5次试验;相应的平均任务完成进度分别为$99.03\%$、$100.00\%$和$96.67\%$。在所有评估条件下,\method{}均取得了高于ReKep、Dream2Flow和$\pi_{0.5}$基准方法的成功率,证明了显式物体状态推理、可组合原子动作、基于代价的规划选择与闭环执行验证相结合的有效性。
cs.RO / 16 / 2609.05994

GLoRI: Closed-Loop Whole-Body Tracking with Global-Local Reference Interaction for Humanoid Loco-Manipulation

GLoRI:基于全局-局部参考交互的闭环全身跟踪方法,用于人形机器人移动操作
Xu, Qingyao, Yin, Sheng, Zhou, Zibo, Zhang, Ya, Chen, Siheng, Hu, Yue
Abstract
Humanoid loco-manipulation requires accurate whole-body motion tracking in the world frame for physical interaction. While local references preserve motion structure, they lack explicit constraints on absolute spatial placement, leading to accumulated global errors. Existing globally aware approaches augment teleoperation policies with global observations but do not explicitly integrate global correction with local motion guidance, limiting autonomous tracking accuracy. We present GLoRI, a closed-loop whole-body controller that integrates structured global reference and feedback with local motion guidance. Its GLoRI-Net uses Global-Local Cross Attention(GLCA) to refine local keypoint features with global target and pose-difference features, preserving motion structure while correcting world-frame placement. GLoRI achieves 100% completion and a g-MPJPE of 6.44cm on held-out HuMoTo motions. This accuracy remains robust under direct Isaac Gym-to-MuJoCo transfer without fine-tuning, demonstrating strong generalization. Furthermore, such accuracy and generalization enable autonomous loco-manipulation with a single policy on a real Unitree G1 interacting with diverse unseen objects, extending beyond prior systems that primarily rely on teleoperation or focus on single-object interactions.
Chinese Translation
人形机器人移动操作(loco-manipulation)需要在世界坐标系下进行精确的全身运动跟踪以实现物理交互。局部参考虽然能够保持运动结构,但缺乏对绝对空间位置的显式约束,导致全局误差累积。现有的全局感知方法通过引入全局观测来增强遥操作策略,但并未将全局校正与局部运动引导显式融合,限制了自主跟踪的精度。我们提出GLoRI,一种将结构化全局参考与反馈和局部运动引导相融合的闭环全身控制器。其GLoRI-Net采用全局-局部交叉注意力机制(GLCA),利用全局目标特征与位姿差异特征来精炼局部关键点特征,在校正世界坐标系位置的同时保持运动结构。GLoRI在留出的HuMoTo动作上实现了100%的完成率和6.44厘米的g-MPJPE。在无需微调的直接Isaac Gym到MuJoCo迁移下,该精度依然保持稳健,展现出强大的泛化能力。此外,这种精度与泛化能力使得单一策略即可在真实的Unitree G1机器人上实现与多种未见物体的自主移动操作,超越了以往主要依赖遥操作或仅专注于单一物体交互的系统。
cs.RO / 17 / 2609.06009

How to Learn from What a Human Would Avoid? Intervention-Aware World Models with Real-World RL for Dexterous Manipulation

如何从人类会避免的情形中学习?面向灵巧操作的干预感知世界模型与真实世界强化学习
Yin, Jiaju, Zhang, Zhenhui, Xu, Lixin, Zhang, Heng, Shao, Jun, Feng, Yating, Ajoudani, Arash, Xu, Renjing
Abstract
Multi-fingered dexterous manipulation remains a frontier for real-world reinforcement learning (RL) due to the high-dimensional action space and the prohibitive cost of hardware failures. While human-in-the-loop (HIL) RL allows operators to intervene before failures occur, current pipelines often treat these interventions as reactive corrections, discarding the rich safety signal inherent in the operator's decision to take control. In this paper, we ask: How can we learn from what a human would avoid? We present WHIRL, a safety-aware RL framework that transforms binary human interventions into forward-predictive signals for proactive risk avoidance. Our approach centers on an intervention-aware latent world model with four prediction heads: dynamics, reward, termination, and a novel per-state intervention-probability head that learns to predict the likelihood of a human takeover at future states. This head provides an actor-side risk-shaping term that discourages the policy from entering "intervention-prone" regions, modeling the operator's internal safety threshold. We evaluate our framework on a 16-DoF LEAP Hand across tasks spanning convex and irregular object grasping, prismatic manipulation, and long-horizon multi-stage tasks. Our results show that predictive risk-shaping enables the system to achieve a 96.7 percent success rate on complex grasping tasks while reducing the operator intervention burden by up to 84 percent in step-weighted terms. By closing the loop between human intuition and predictive world modeling, this work provides a practical safety-aware recipe for training complex dexterous agents in the real world while reducing operator fatigue and hardware-risk exposure.
Chinese Translation
由于高维动作空间以及硬件失效的巨大代价,多手指灵巧操作仍然是真实世界强化学习(RL)的前沿难题。虽然人在环路(human-in-the-loop, HIL)强化学习允许操作者在故障发生前进行干预,但现有流程往往将这些干预视为被动的纠正,丢弃了操作者接管控制这一决策中蕴含的丰富安全信号。本文提出这样一个问题:我们如何从人类会避免的情形中学习?我们提出了 WHIRL,一个安全感知的强化学习框架,它将二元的人类干预转化为前向预测信号,以实现主动的风险规避。我们的方法核心是一个干预感知的潜在世界模型,它包含四个预测头:动力学、奖励、终止,以及一个新颖的逐状态干预概率头,用于学习预测人类在未来状态上接管控制的概率。该预测头为执行器(actor)侧提供了一个风险塑形项,阻止策略进入“易受干预”的区域,从而对操作者内在的安全阈值进行建模。我们在一个16自由度的 LEAP Hand 上评估了该框架,任务涵盖凸面和不规则物体的抓取、棱柱形物体操作以及长时程多阶段任务。结果表明,预测性风险塑形使系统在复杂抓取任务上达到 96.7% 的成功率,同时在按步数加权统计下将操作者的干预负担降低多达 84%。通过将人类直觉与预测性世界建模形成闭环,本工作为在真实世界中训练复杂灵巧智能体提供了一套实用的安全感知方案,同时降低了操作者疲劳和硬件风险暴露。
cs.RO / 18 / 2609.06046

FALCON-S: Fixed-wing ground-effect Aerodynamics Simulator and Flight Control Learning Suite

FALCON-S:固定翼地效空气动力学仿真器与飞行控制学习套件
Hariry, Matteo El, Lima, Pedro, Orsula, Andrej, Geist, Matthieu, Olivares-Mendez, Miguel
Abstract
We present a modular, high-fidelity simulation framework for the development and benchmarking of flight control strategies in fixed-wing aerial robots operating near the ground. Unlike existing simulators that rely on simplified or hover-oriented dynamics, our framework models full 6DoF rigid-body physics, semi-empirical ground-effect aerodynamics, actuator dynamics, sensor noise, and environmental disturbances. This physical realism, combined with modular component design, enables systematic analysis of low-altitude flight behavior under realistic conditions. The simulator supports both CPU and GPU backends via Torch and NVIDIA Warp, enabling high-throughput parallel execution suitable for large-scale reinforcement learning training and optimal control rollouts. A unified interface accommodates a range of controllers (both RL and optical control algorithms) across tasks such as altitude regulation and trajectory tracking. Cross-validation with X-Plane and JSBSim is also supported to facilitate engineering integration and visual fidelity.
Chinese Translation
我们提出了一个模块化、高保真度的仿真框架,用于开发和评估在近地面运行的固定翼无人机的飞行控制策略。与现有依赖简化动力学模型或面向悬停动力学的仿真器不同,我们的框架建模了完整的六自由度(6DoF)刚体物理、半经验地效空气动力学、执行器动力学、传感器噪声以及环境扰动。这种物理真实性结合模块化的组件设计,使得在真实条件下对低空飞行行为进行系统性分析成为可能。该仿真器通过Torch和NVIDIA Warp支持CPU与GPU后端,能够实现高吞吐量的并行执行,适用于大规模强化学习训练和最优控制仿真推演。统一的接口可容纳多种控制器(包括强化学习和最优控制算法),并支持高度调节和轨迹跟踪等任务。此外,还支持与X-Plane和JSBSim进行交叉验证,以促进工程集成和视觉保真度的提升。
cs.RO / 19 / 2609.06061

SCoCaT: Success Conditioned Constrained Reinforcement Learning for Spacecraft Docking

SCoCaT:面向航天器对接的成功条件约束强化学习
Arora, Aman, Castan, Ricard Marsal I, El-Hariry, Matteo, Olivares-Mendez, Miguel
Abstract
Termination-based constrained reinforcement learning is attractive for safety-critical robotic deployments: it avoids online optimization at inference, scales easily to many constraints via a single scalar per constraint, and is simpler to implement than commonly used Lagrangian methods. Instead of pricing violations through summed cost penalties, this approach makes violations structurally unprofitable by shortening the effective horizon for each violation. We identify a structural failure mode of this method class on terminal-navigation tasks: reaching a precise goal configuration while satisfying safety constraints that tighten along the final approach. When the goal sits inside the region close to where the constraints become active, the survival-weighted objective makes dwelling outside the goal region strictly preferable to entering, producing high constraint compliance with low task completion. We formalize this pathology and show that a minimal augmentation to off-the-shelf RL algorithms like PPO resolves this ``feasibility collapse''. We empirically demonstrate that adding a dense per-step success signal via an auxiliary value critic improves the task completion rate while maintaining safety-critical constraint compliance. Validation across two representative spacecraft platforms: a 6U-CubeSat spanning the mass and degree-of-freedom envelope of operational proximity operations, and a floating platform testbed for zero-shot sim-to-real transfer in our laboratory, supports the generality of these findings.
Chinese Translation
基于终止的约束强化学习对于安全至关重要的机器人部署具有吸引力:它避免了推理阶段的在线优化,可通过每个约束对应单一标量的方式轻松扩展到多约束场景,并且比常用的拉格朗日方法更易于实现。该方法并非通过累加代价惩罚来衡量违规,而是通过缩短每次违规的有效时间范围,使违规在结构上无利可图。我们识别出此类方法在终端导航任务上的一种结构性失效模式:在满足最终逼近阶段逐渐收紧的安全约束的同时到达精确的目标位形。当目标位于约束开始生效区域附近时,基于生存加权的目标函数使得停留在目标区域之外严格优于进入目标区域,从而产生高约束符合性但低任务完成率。我们对这一病态现象进行了形式化描述,并证明对现有强化学习算法(如PPO)进行最小程度的增强即可解决这种“可行性坍缩”。我们通过实验证明,借助辅助价值评判器添加密集的每步成功信号,可以在保持安全关键约束符合性的同时提高任务完成率。在两个代表性航天器平台上的验证支持了这些发现的普适性:一个是覆盖在轨近距离操作质量和自由度范围的6U立方星,另一个是我们实验室中用于零样本仿真到现实迁移的悬浮平台测试台。
cs.RO / 20 / 2609.06114

Where Success Breaks: Failure-Boundary Learning for Robust Vision-Language-Action Models

成功之处何在:面向鲁棒视觉-语言-动作模型的失败边界学习
Chen, Yanzhe, Cao, Zhijun, Shou, Mike Zheng
Abstract
Vision-language-action (VLA) models adapted through supervised fine-tuning (SFT) inherit a structural asymmetry: expert demonstrations teach the policy where success behavior lies, but provide no signal about where it ceases to be reliable. We argue that robust VLA adaptation should therefore be viewed not as further demonstration fitting, but as **Failure-Boundary Learning**---the problem of *Discovering*, *Localizing*, and *Shaping* the boundary between recoverable deviations and task failure. To instantiate this view, we propose **DLS**: built on a **real-grounded behavioral prior** from few real demonstrations and simulated co-training, DLS *discovers* failure boundaries at scale through on-policy digital twin rollouts. Rather than reducing each rollout to a binary label, **semantic progress localization** uses privileged simulator states to assign progress-aware signals that capture *where* the failure boundary is crossed, not merely *whether*. These signals drive **directional boundary shaping** in the flow dynamics---reinforcing success-producing denoising directions and suppressing failure-producing ones, without action likelihoods or auxiliary critics. Across real-robot manipulation tasks, DLS improves robustness over SFT and online RL baselines, especially under randomized initial states and unseen visual conditions.
Chinese Translation
通过监督微调(SFT)适配的视觉-语言-动作模型继承了一种结构性不对称:专家示范只教会策略成功行为所在之处,却无法提供关于其何时不再可靠的信号。因此,我们认为鲁棒的VLA适配不应被视为进一步的示范拟合,而应被视为失败边界学习——即发现、定位并塑形可恢复偏差与任务失败之间边界的问题。为将这一观点具体化,我们提出了DLS方法:DLS建立在来自少量真实示范与仿真共训练的真实基础行为先验之上,通过在线数字孪生推演大规模地发现失败边界。不同于将每次推演简化为二元标签,语义进度定位利用特权仿真器状态分配具有进度感知的信号,从而捕捉失败边界在何处被跨越,而不仅仅是是否被跨越。这些信号驱动流动力学中的方向性边界塑形——强化产生成功的去噪方向,抑制导致失败的去噪方向,且无需动作似然或辅助评论家。在真实机器人操作任务上,DLS相较于SFT和在线RL基线提升了鲁棒性,尤其是在随机化初始状态和未见视觉条件下。
cs.RO / 21 / 2609.06221

RefGuard: Identity-Aware Language-Guided Robot Manipulation via Joint Target-Anchor-Frame Grounding

RefGuard:基于目标-锚点-参考帧联合接地(Grounding)的身份感知语言引导机器人操作
Wei, Lan, Lu, Kangyi, Wang, Yongchen, Bi, Chenmeng, Chen, Qi, Niu, Hanlin, Yeung, Yip Fun, Zhang, Dandan
Abstract
Vision-language-action (VLA) models have substantially advanced language-guided robot manipulation, yet reliable execution still hinges on identifying which physical object an instruction refers to. In cluttered scenes containing repeated objects, ambiguous anchors, or frame-dependent spatial terms, a robot can execute a geometrically valid action on a semantically compatible but unintended instance; we call this failure an identity switch. The referent is jointly determined by three coupled latent variables: the target, the anchor, and the reference frame, so committing to any one of them before execution turns residual ambiguity into a silent and irreversible error. We propose RefGuard, an identity-aware grounding framework that delays commitment by maintaining a joint posterior over all three variables. RefGuard builds a frame-conditioned object-centric scene graph from RGB-D observations, separating frame-independent geometry from directional relations, and routes the posterior through a decision policy that executes, clarifies, reobserves, or aborts. On a real UF850 arm, RefGuard records no identity switch on any ambiguity-stress trial and executes correctly on 90.0% of them, whereas fine-tuned VLA and LLM (Large Language Model)-based baselines switch identity in 33-46% of the same trials, while retaining 93.3% success on unambiguous scenes and recovering from post-grounding scene changes in 86.7% of trials. On a 3200-episode procedural suite, it raises correct execution on solvable instructions from 56.6% to 80.5% over the ablation that commits to the anchor and frame before the target, while deferring less often (19.5% vs. 43.4%).
Chinese Translation
视觉-语言-动作(VLA)模型极大地推动了语言引导的机器人操作,但可靠的执行仍然依赖于正确识别指令所指的物理对象。在包含重复物体、模糊锚点或依赖参考帧的空间表述的杂乱场景中,机器人可能会对一个语义上兼容但并非目标实例的对象执行几何上有效的动作;我们将这种失败称为身份切换(identity switch)。指称对象由三个相互耦合的隐变量共同决定:目标、锚点和参考帧,因此在执行前对其中任何一个变量提前做出承诺,都会将残留的歧义转化为沉默且不可逆的错误。我们提出 RefGuard,一个身份感知的接地(grounding)框架,它通过维护这三个变量的联合后验分布来延迟承诺。RefGuard 从 RGB-D 观测构建以物体为中心、条件于参考帧的场景图,将不依赖参考帧的几何信息与方向性关系分离开来,并通过一个决策策略对后验分布进行路由,决定执行、澄清、重新观测或中止。在真实的 UF850 机械臂上,RefGuard 在所有歧义压力测试中均未发生身份切换,并在其中 90.0% 的测试中正确执行;而经过微调的 VLA 以及基于大语言模型(LLM)的基线方法在相同测试中有 33%–46% 发生身份切换。同时,RefGuard 在无歧义场景中保持 93.3% 的成功率,并在 86.7% 的测试中能够从接地后的场景变化中恢复。在一个包含 3200 个回合的程序化测试套件中,与在目标之前先承诺锚点和参考帧的消融版本相比,RefGuard 将可解指令的正确执行率从 56.6% 提升至 80.5%,同时更少地进行延迟决策(19.5% 对比 43.4%)。
cs.RO / 22 / 2609.06256

GloVLA: Let Geometry Move and Local VLA Interact for Robust Object-Centric Manipulation in Unstructured Environments

GloVLA:让几何移动与局部VLA交互,实现非结构化环境中鲁棒的以物体为中心的操作
Nguyen, Truong Thanh, Nguyen, Huy Hoang, Nguyen, Ha Anh, Dinh, Binh Khanh, Vien, Ngo Anh, Minh, Duy Nguyen Ho, Vu, Minh Nhat, Le, Ngan
Abstract
Vision-language-action (VLA) models have shown promising generalization for language-conditioned robot manipulation, but deploying them in unstructured environments remains challenging. A single end-to-end VLA policy must simultaneously solve long-range transport of the end effector to task-relevant regions and short-horizon, contact-rich interaction upon arrival. This formulation is inefficient and brittle: small visual shifts, distractors, clutter, occlusions, or unfavorable initial gripper poses can push the policy outside the local state distribution in which it was trained, leading to task failure. We introduce GloVLA, a hybrid framework that explicitly separates object-centric manipulation into two complementary regimes: a geometric transport controller moves the end-effector into interaction-centric handoff regions, and local VLA policies handle only the short-horizon interaction phases. GloVLA is model-agnostic and can be integrated with different VLA backbones with no additional demonstrations and no changes to the action space or success predicate. Experiments on standard LIBERO and LIBERO-Plus Object tasks together with a newly introduced LIBERO-Challenge benchmark ettings with clutter, distractors,illumination changes, visual shifts, and obstruction show that GloVLA improves task success and substantially lowers VLA inference cost compared with full end-to-Challenge, full-trajectory GR00T N1.6execution degrades to 20.9% average success while GloVLA retains 88.5%; on a physical UR10e, overall success improves from 35.6% to 90.0% while mean inference time is more than halved. Videos and additional results are available at https://glovla-project.github.io/
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型在语言条件化的机器人操作任务中展现出良好的泛化能力,但将其部署于非结构化环境仍然具有挑战性。单一端到端VLA策略必须同时解决末端执行器向任务相关区域的长距离移动,以及到达后的短时程、接触丰富的交互问题。这种 formulation 效率低下且脆弱:微小的视觉偏移、干扰物、杂乱环境、遮挡或不利初始夹爪位姿都可能使策略偏离其训练时的局部状态分布,导致任务失败。我们提出GloVLA,这是一种混合框架,明确将以物体为中心的操作划分为两个互补模式:几何移动控制器将末端执行器移至以交互为中心的交接区域,而局部VLA策略仅负责处理短时程交互阶段。GloVLA与具体模型无关,可与不同的VLA骨干网络集成,无需额外的示范数据,也无需改变动作空间或成功判据。在标准LIBERO和LIBERO-Plus物体任务上,以及新提出的包含杂乱环境、干扰物、光照变化、视觉偏移和遮挡的LIBERO-Challenge基准设置上的实验表明,与完整端到端执行相比,GloVLA提高了任务成功率并显著降低了VLA推理成本。在LIBERO-Challenge上,全轨迹GR00T N1.6执行的最新成果平均成功率退化至20.9%,而GloVLA仍保持88.5%;在物理UR10e机器人上,总体成功率从35.6%提升至90.0%,同时平均推理时间缩短一半以上。视频和更多结果见 https://glovla-project.github.io/
cs.RO / 23 / 2609.06273

Bundle Length Tradeoffs in Decentralized Multi-Robot Task Allocation Under Degraded Communications

退化通信条件下分散式多机器人任务分配中的任务束长度权衡
Lott, James, Honary, Vahraz
Abstract
Bundle length B is commonly fixed when configuring multi-task multi-robot task allocation (MRTA) algorithms. MinSum and MinMax are known to favor different task distributions, but the role of B in this objective tradeoff has not been systematically characterized. Additionally, degraded-communication evaluations also often retain settings selected under ideal communication, leaving whether nominal bundle-length tuning transfers under message loss unresolved. We examine both questions for ACBBA, PI, and HIPC across six bundle lengths in 300 paired ten-target Collaborative Visit scenarios under ideal communication and 25% Bernoulli packet loss. Under ideal communication, increasing B from 1 to 12 reduces MinSum cost by 19.0%, 23.0%, and 31.8% for ACBBA, PI, and HIPC, respectively, while increasing MinMax cost by 45.6%, 94.3%, and 67.6%. Under packet loss, the lowest-mean MinSum setting shifts from B = 12 to B = 2 for ACBBA and PI. Repeated paired cross-fitting shows that retaining the ideal-network setting incurs held-out MinSum penalties of 14.4% and 7.2%, respectively, and increases MinMax cost by 30.0% and 41.8% relative to the loss-conditioned MinSum setting. HIPC retains a deep MinSum operating region, while the MinMax setting remains stable for all three allocators. Experiments at two additional target loads reproduce the ACBBA and PI MinSum shifts.
Chinese Translation
在配置多任务多机器人任务分配(MRTA)算法时,任务束长度B通常被固定设置。已知MinSum和MinMax目标倾向于不同的任务分布,但B在这一目标权衡中的作用尚未得到系统性的刻画。此外,退化通信条件下的评估也往往沿用理想通信条件下选择的设置,使得标称的任务束长度调优在消息丢失情形下能否迁移仍悬而未决。我们在理想通信和25%伯努利丢包条件下,针对ACBBA、PI和HIPC三种算法,在300个成对的十目标协同访问场景中考察了六种任务束长度下的上述两个问题。在理想通信条件下,将B从1增加到12使ACBBA、PI和HIPC的MinSum成本分别降低19.0%、23.0%和31.8%,而MinMax成本分别增加45.6%、94.3%和67.6%。在丢包条件下,ACBBA和PI的平均MinSum最低设置从B = 12转移到B = 2。重复的成对交叉拟合表明,沿用理想网络设置相对基于丢包条件调优的MinSum设置,分别产生14.4%和7.2%的留出集MinSum代价,并使MinMax成本分别增加30.0%和41.8%。HIPC保留了较深的MinSum运行区间,而三种分配器的MinMax设置均保持稳定。在另外两种目标负载下进行的实验重现了ACBBA和PI的MinSum偏移现象。
cs.RO / 24 / 2609.06279

IM-ENGINE: Image Editing for Embodied Data Generation

IM-ENGINE:面向具身数据生成的图像编辑方法
Wang, Yian, Cao, Junyi, Qiu, Xiaowen, Gan, Chuang
Abstract
Learning-based manipulation requires supervision that is both semantically meaningful and physically executable, but current data pipelines often provide only one of these properties. Human demonstrations capture intent but are costly to collect and constrained by the human-robot embodiment gap, while simulation can scale data generation but often under-specifies functional behavior. We present IM-ENGINE, a simulator-grounded pipeline that uses image editing as an intermediate representation for embodied data generation. Given a rendered scene with known geometry, depth, segmentation, and camera parameters, IM-ENGINE edits the image to inject task-relevant semantics, recovers explicit 3D state using simulator priors and an unchanged anchor object, refines the state in physics, and converts it into robot-executable supervision. We instantiate the pipeline for dexterous grasp synthesis and goal-state generation. For grasping, IM-ENGINE generates a human grasp in image space, recovers the hand-object interaction, retargets it to a robot hand, and refines it into physically validated robot grasps. For goal generation, it edits a rendered scene into a desired outcome, recovers the target-object pose, and refines it into physically valid, semantically meaningful goals and trajectories. This combination of generative semantic priors and simulator grounding enables scalable task-relevant supervision for robot learning.
Chinese Translation
基于学习的操作任务需要既具有语义意义又可在物理上执行的监督信号,但现有的数据流水线往往只能提供其中一种属性。人类示教能够捕捉任务意图,但采集成本高昂,且受限于人机形态差异(embodiment gap);而仿真虽然可以大规模生成数据,却常常无法充分刻画功能性动作。我们提出了 IM-ENGINE,一种以仿真为基础的流水线,将图像编辑作为具身数据生成的中间表示。给定一个具有已知几何、深度、分割和相机参数的渲染场景,IM-ENGINE 通过编辑图像注入与任务相关的语义信息,利用仿真器先验和保持不变的锚定物体恢复显式三维状态,在物理引擎中对状态进行优化,并将其转换为机器人可执行的监督信号。我们将该流水线应用于灵巧抓取合成与目标状态生成。在抓取任务中,IM-ENGINE 在图像空间中生成人类抓取姿态,恢复手-物体交互关系,将其重定向至机器人手,并优化为经物理验证的机器人抓取。在目标生成任务中,它将渲染场景编辑为期望的结果,恢复目标物体的位姿,并优化为物理上有效且语义上有意义的目标与轨迹。生成式语义先验与仿真基础的结合,为机器人学习提供了可扩展的、与任务相关的监督信号。
cs.RO / 25 / 2609.06331

Geometric Distributional Control: Learning Progress with Partial Structural Knowledge

几何分布控制:基于部分结构化知识的学习进展
Wu, Tong
Abstract
Real-time control often sits between two limiting regimes. Predictive optimization and model-based control are powerful when dynamics, parameters, objectives, and online planning models are specified; reinforcement learning can relax this requirement, but must infer long-horizon value signals from sequential data and interaction, making training slow, high-variance, and hard to scale in large action spaces. This middle regime is common in systems including autonomous driving, warehouse robotics, traffic control, and delivery drones: partial geometry, physics, rules, or constraints are known, yet the local direction of task progress remains uncertain. Geometric Distributional Control (GDC) is designed for this partial-knowledge setting. It factorizes control into feasibility and progress: known geometry, rules, constraints, and response maps define an executable scaffold, while progress-weighted feasible data learns the missing directional signal on that scaffold. The learned score acts as a Bellman-like local value-gradient, selecting actions that make progress without requiring global Bellman recursion, a fully specified planner, or a black-box policy that absorbs both feasibility and preference. This knowledge can be lightweight and partial, such as simple dynamics, safety filters, local maps, constraint projectors, or lower-level response maps; it need not encode full dynamics or a long-horizon objective. Offline, GDC fits a progress-tilted distribution from short known-feasible snippets with weak signed progress certificates. Online, its score is projected through the scaffold and applied in receding-horizon feedback. We prove that this score descends a data-induced soft progress value and validate GDC on structured multilevel optimization and SUMO route-progress driving, where it improves over known-only solvers and learning baselines while preserving scaffold-enforced feasibility.
Chinese Translation
实时控制常常处于两种受限范式之间。当动力学、参数、目标和在线规划模型被明确给定时,预测优化和基于模型的控制是强大的;强化学习可以放宽这一要求,但必须从序贯数据和交互中推断长时程价值信号,这使得训练缓慢、方差高,且难以在大动作空间中扩展。这种中间情形在包括自动驾驶、仓储机器人、交通控制和送货无人机在内的系统中十分常见:部分几何、物理规律、规则或约束是已知的,但任务进展的局部方向仍然不确定。几何分布控制(Geometric Distributional Control, GDC)正是为这种部分知识情形设计的。它将控制分解为可行性与进展两部分:已知的几何、规则、约束和响应映射定义了一个可执行的骨架,而基于进展加权的可行数据则在该骨架上学习缺失的方向性信号。所学习的分数(score)充当一种类贝尔曼(Bellman-like)的局部价值梯度,能够选择产生进展的动作,而无需全局贝尔曼递归、完全明确的规划器,或同时吸收可行性与偏好的黑盒策略。这些知识可以是轻量且部分的,例如简单动力学、安全滤波器、局部地图、约束投影算子或底层响应映射;它不必编码完整的动力学或长时程目标。离线阶段,GDC 利用带有微弱符号进展凭证的短已知可行片段,拟合一个偏向进展的分布。在线阶段,其分数经由骨架投影,并以滚动时域反馈的方式应用。我们证明了该分数沿数据诱导的软进展价值下降,并在结构化多层优化和 SUMO 路线进展驾驶任务上验证了 GDC。在这些任务中,它优于仅利用已知信息的求解器和学习类基线方法,同时保持了由骨架强制保证的可行性。
cs.RO / 26 / 2609.06342

OcclusionCBF: Backup Control Barrier Functions for Safe Navigation Among Hidden Dynamic Obstacles

OcclusionCBF:面向隐藏动态障碍物安全导航的备份控制障碍函数
Kim, Taekyung, Park, Hun Kuk, Wada, Renya, Atanasov, Nikolay, Koga, Shumon, Panagou, Dimitra
Abstract
Robots navigating under occlusion may enter states from which no admissible input can avoid a dynamic obstacle once it becomes visible. We present OcclusionCBF, a safety filter that extends backup control barrier functions to reachable-occupancy predictions for potentially hidden dynamic obstacles in occluded regions. The method certifies a prescribed backup rollout against collision-inflated occupancy and a verified terminal set, yielding affine constraints for minimally invasive quadratic-program filtering. We establish recursive feasibility of the resulting safety filter, and collision avoidance for every hidden-obstacle motion covered by the occupancy prediction. Randomized benchmarks, MetaUrban simulations, and hardware experiments demonstrate improved task success over reactive and occlusion-aware predictive baselines with millisecond-scale computation.
Chinese Translation
在遮挡条件下导航的机器人可能进入这样的状态:一旦动态障碍物变为可见,任何可容许输入都无法避免碰撞。我们提出了OcclusionCBF,一种安全滤波器,它将备份控制障碍函数(backup control barrier functions)扩展至对遮挡区域内可能隐藏的动态障碍物的可达占用预测。该方法针对碰撞膨胀后的占用区域以及经验证的终端集合,对预设的备份轨迹进行认证,从而产生仿射约束,用于实现最小干预的二次规划滤波。我们证明了所得安全滤波器的递归可行性,以及在占用预测所覆盖的任何隐藏障碍物运动情况下的碰撞规避性。随机化基准测试、MetaUrban仿真以及硬件实验表明,该方法在毫秒级计算量下,相较反应式和遮挡感知预测基线方法显著提升了任务成功率。
cs.RO / 27 / 2609.06368

LANTERN: A Closed-Loop Benchmark for VLM-Based Cooperative Driving with Temporally Grounded Warnings

LANTERN:基于视觉语言模型(VLM)协作式驾驶的闭环基准测试,具备时间定位的预警机制
Liu, Yongshuo, Gao, Xu, Zhu, Morui, Zhu, Yongqi, Chen, Qi, Qu, Deyuan, Fu, Song, Yang, Qing
Abstract
We present LANTERN, a closed-loop benchmark for temporally grounded cooperative warnings. LANTERN separates warning onset, hazard onset, warning termination, and post-hazard recovery, and evaluates each physical event under matched warning and no-warning executions so that the warning's contribution is measured in isolation rather than confounded with onboard vision. The benchmark spans six safety-critical scenario families and provides 3,272 sequences with 236,309 frames for training, together with 120 matched route pairs for closed-loop evaluation. Each hazard route is evaluated under the warning and no-warning conditions, while its no-hazard control penalizes unconditional braking. We further introduce the Cooperative Unified Score (CUS), a safety-gated metric that jointly rewards route progress, anticipation, clearance, and recovery. Fine-tuning a representative VLM driving model raises CUS from 34.6 without warnings to 75.5 with them, demonstrating both the value of cooperative warnings and the discriminative power of the paired protocol. All resources will be made publicly available.
Chinese Translation
我们提出了LANTERN,一个面向时间定位协作式预警的闭环基准测试。LANTERN将预警起始、危险起始、预警终止以及危险后恢复分别解耦,并在匹配的“有预警”和“无预警”执行条件下评估每个物理事件,从而孤立地衡量预警本身的贡献,避免与车载视觉感知相混淆。该基准涵盖六个安全关键场景族,提供包含236,309帧的3,272条序列用于训练,并附带120对匹配的路线用于闭环评估。每个危险路线均在有预警和无预警两种条件下进行评估,而其无危险的对照路线则用于惩罚无条件刹车行为。我们进一步提出了协作式统一评分,这是一个以安全性为门控的指标,联合奖励路线进度、预判能力、危险解除和恢复能力。对一款具有代表性的VLM驾驶模型进行微调后,CUS从无预警时的34.6提升至有预警时的75.5,这既展示了协作式预警的价值,也验证了配对协议的判别能力。所有资源将公开发布。
cs.RO / 28 / 2609.06424

OVMAN: A Task and Benchmark for Open-Vocabulary Motion-Aware Navigation

OVMAN:面向开放词汇运动感知导航的任务与基准
Ghosh, Dibyendu
Abstract
Homes change between a robot's visits. Navigation benchmarks pose their goals in the world the agent currently sees, and the two-visit benchmarks that exist score recall or rearrangement rather than navigation. None of them can express go to the chair that was moved or go to where the vase used to be. OVMAN is a task in which an agent tours a scene, returns after a scripted change, and must navigate to a goal specified by the change itself. Two of its six change relations answer with a place an object has left, where nothing remains to be detected. We release 219 two-visit episodes, each certified solvable by an oracle agent that completes it three times. Two released systems fail as predicted. A zero-shot object-goal navigator reaches the change-defined target in 11.4% of episodes and essentially fails at vacated locations, 0.000 on former and 0.025 on removed. A self-maintaining open-vocabulary map answers past-tense queries at about a third of the rate of the same maps read as two visits, even when grounding is supplied. A simple two-visit reference agent reaches 45.2% when navigating, against an embodied oracle of 99.5%. An error decomposition places the remaining difficulty in open-vocabulary instance grounding rather than in geometry or in selecting the answer once positions are known.
Chinese Translation
家庭环境在机器人两次到访之间会发生变化。现有导航基准将目标设定在智能体当前所见的世界中,而已有的两次到访类基准评测的是回忆或重排,而非导航。它们都无法表达“前往被移动过的那把椅子”或“前往花瓶原先所在的位置”这类目标。OVMAN 是一项任务:智能体先巡览一个场景,在场景发生预定变化后返回,并必须导航至由该变化本身所指定的目标。在其六种变化关系中,有两种的答案是某个物体已离开的位置,那里不再有任何可被检测到的东西。我们发布了219个两次到访的回合(episode),每个回合均经认证可由一个完成该回合三次的 oracle 智能体求解。两个已发布的系统如预期般失败。一个零样本物体目标导航器仅在11.4%的回合中到达由变化定义的目标,且在已腾空位置上基本失败:在原位置(former)上为0.000,在被移除(removed)上为0.025。一个自我维护的开放词汇地图在回答过去时态查询时,成功率仅约为将同一地图按两次到访方式读取时的三分之一,即使提供了视觉定位(grounding)也是如此。一个简单的两次到访参考智能体在导航时达到45.2%,而具身 oracle 为99.5%。误差分解表明,剩余的困难在于开放词汇的实例定位,而非几何感知或在已知位置后选择答案。
cs.RO / 29 / 2609.06433

Collision Snapshot Guided Time-Reversed Safety-Critical Scenario Generation

碰撞快照引导的时间反向安全关键场景生成
Kim, Taehyung, Choi, Jongeun
Abstract
The generation of safety-critical traffic scenarios is essential for training and evaluating autonomous vehicles. Prior approaches typically perturb the trajectories of existing agents in a traffic scenario using simplified adversarial objectives to induce safety-critical interactions, which can limit the plausibility and diversity of the generated scenarios. Although inserting new adversarial vehicles can alleviate this limitation, determining when and where to introduce them in a scenario-specific manner remains challenging. In this work, we introduce \underline{CO}llision \underline{S}napshot guided \underline{T}im\underline{E}-\underline{R}eversed safety-critical scenario generation (COSTER), a framework that leverages learned traffic priors to determine plausible collision times and locations. COSTER first constructs a collision snapshot by inserting a new vehicle in contact with the target vehicle at the identified collision state within a traffic scenario. Starting from this collision snapshot, a conditional variational autoencoder is used to perform a time-reversed rollout, reconstructing the trajectory of the inserted vehicle backward toward earlier timesteps. Experiments show that COSTER outperforms existing methods in plausibility, diversity, and data efficiency. Moreover, agents trained on COSTER-generated scenarios reduce collision rates by 31\% on safety-critical scenarios from the Waymo Open Motion Dataset while also improving ego task completion. The project website is available at https://anonym-121.github.io/COSTER/.
Chinese Translation
安全关键交通场景的生成对于自动驾驶车辆的训练与评估至关重要。已有方法通常通过简化的对抗性目标对现有交通场景中智能体的轨迹进行扰动,以引发安全关键交互,这可能限制生成场景的合理性与多样性。尽管插入新的对抗车辆可以缓解这一限制,但如何以场景特定的方式确定何时何地引入这些车辆仍具挑战性。在本工作中,我们提出了碰撞快照引导的时间反向安全关键场景生成框架(COllision Snapshot guided TimE-Reversed safety-critical scenario generation, COSTER),该框架利用学习到的交通先验来确定合理的碰撞时间与地点。COSTER 首先通过在交通场景中所识别的碰撞状态下插入一辆与目标车辆接触的新车,构建碰撞快照。从该碰撞快照出发,采用条件变分自编码器(conditional variational autoencoder)进行时间反向推演,反向重建所插入车辆在更早时间步的轨迹。实验表明,COSTER 在合理性、多样性和数据效率方面均优于现有方法。此外,在 COSTER 生成的场景上训练的智能体,在 Waymo Open Motion Dataset 的安全关键场景上将碰撞率降低了 31%,同时提升了自车任务完成度。项目网站见 https://anonym-121.github.io/COSTER/。
cs.RO / 30 / 2609.06450

\textbf{PLATO}: \emph{Preintegration Learning from Accurate Trajectory Observations} for Neural Inertial Odometry

PLATO:面向神经惯性里程计的基于精确轨迹观测的预积分学习方法
Li, Haoying, Liu, Qihang, Peng, Yifan, Miao, Keyan, Wu, Junfeng
Abstract
Neural inertial odometry has demonstrated strong potential for motion estimation in challenging environments, yet inertial-only preintegration remains sensitive to IMU bias and uncertainty. To this end, this paper introduces \textbf{PLATO}:~\emph{Preintegration Learning from Accurate Trajectory Observations}, a likelihood-based framework that leverages accurate trajectory observations to jointly learn IMU bias dynamics modeled by a neural ordinary differential equation~(NODE) and gyroscope and accelerometer noise covariances. Optimization exploits the sparse structure of the negative log-likelihood, with IMU noise-parameter gradients computed by forward differentiation. A tailored double-adjoint scheme couples a discrete invariant-error adjoint with a continuous-time adjoint for the bias NODE, enabling memory-efficient likelihood optimization over the nested bias-dynamics and IMU-preintegration rollouts. Validation on EuRoC shows improved performance, and underwater robot experiments demonstrate applicability under intermittent lighting failures and visual degradation.
Chinese Translation
神经惯性里程计在复杂环境中的运动估计方面展现出强大潜力,然而仅基于惯性测量的预积分方法对IMU偏置和不确定性仍然较为敏感。为此,本文提出了PLATO(Preintegration Learning from Accurate Trajectory Observations,基于精确轨迹观测的预积分学习),这是一个基于似然的框架,利用精确的轨迹观测来联合学习由神经常微分方程(NODE)建模的IMU偏置动态,以及陀螺仪和加速度计的噪声协方差。优化过程利用了负对数似然的稀疏结构,IMU噪声参数的梯度通过前向微分计算。我们设计了一种定制化的双伴随(double-adjoint)方案,将离散不变误差伴随方法与针对偏置NODE的连续时间伴随方法相耦合,从而在嵌套的偏置动态与IMU预积分 rollout 上实现内存高效的似然优化。在EuRoC数据集上的验证表明性能有所提升,水下机器人实验则证明了该方法在间歇性光照故障和视觉退化条件下的适用性。
cs.RO / 31 / 2609.06508

VLA-Corrector: Stage-Aware Observable State Understanding for Prompt-Based Closed-Loop Recovery of Vision-Language-Action Policies

VLA-Corrector:面向提示驱动闭环恢复的视觉-语言-动作策略阶段感知可观测状态理解方法
Song, Chang, Qian, Bin, Feng, Yan, Song, Zhijie
Abstract
Long-horizon robot manipulation with Vision-Language-Action (VLA) policies remains vulnerable to execution-time deviations, as final task success provides little information for diagnosing and correcting failures caused by action noise, object displacement, or goal misalignment. We introduce a stage-aware failure verification and Prompt Recovery framework that enables closed-loop correction of a fixed VLA policy without parameter updates or privileged simulator states. The framework introduces an observable-history-based Learned Verifier that jointly estimates manipulation progress and execution risk by temporally modeling multi-view visual observations, proprioceptive states, and executed actions. To provide interpretable task understanding, we represent manipulation execution through semantic progress stages, including approach, alignment, grasp, transport, and placement, and identify stage-specific failure patterns. Upon detecting abnormal execution, the framework preserves the original instruction and generates a stage-conditioned recovery prompt, allowing the same frozen VLA policy to produce corrective actions. Extensive multi-round evaluations on LIBERO and LIBERO Plus demonstrate that the proposed approach substantially improves closed-loop reliability under diverse perturbations. Without access to privileged object or goal coordinates, the Learned Verifier achieves recovery performance close to that of the privileged rule-based verifier in the evaluated settings. These results show that observable visual-proprioceptive-action history is sufficient to infer latent task states and enable practical failure recovery for existing VLA policies.
Chinese Translation
基于视觉-语言-动作(Vision-Language-Action, VLA)策略的长时程机器人操作仍然容易受到执行时偏差的影响,因为最终任务成功与否几乎无法为诊断和纠正由动作噪声、物体位移或目标不对齐所导致的失败提供信息。我们提出了一种阶段感知的失败验证与提示恢复(Prompt Recovery)框架,能够在不更新参数、无需特权仿真器状态的情况下,对固定的VLA策略实现闭环纠错。该框架引入了一个基于可观测历史的学习型验证器(Learned Verifier),通过对多视角视觉观测、本体感知状态和已执行动作进行时间建模,联合估计操作进度与执行风险。为了提供可解释的任务理解,我们通过语义进度阶段(包括接近、对齐、抓取、运输和放置)来表示操作的执行过程,并识别各阶段特有的失败模式。在检测到异常执行时,该框架保留原始指令并生成阶段条件化的恢复提示,使同一冻结的VLA策略能够产生纠正性动作。在LIBERO和LIBERO Plus上的大量多轮评估表明,所提方法在多种扰动下显著提升了闭环可靠性。在无法获取特权物体或目标坐标的情况下,学习型验证器在所评估的设置中达到了接近特权规则验证器的恢复性能。这些结果表明,可观测的视觉-本体感知-动作历史足以推断潜在任务状态,并能为现有VLA策略实现实用的失败恢复。
cs.RO / 32 / 2609.06522

Knowledge-Guided Hierarchical Policy Learning for High-Precision Cylindrical Assembly under Tight Tolerances

面向紧密公差高精度圆柱装配的知识引导分层策略学习
Lian, Binbin, Liu, Xinyu, Sun, Tao
Abstract
A hybrid hierarchical learning framework is proposed to achieve high-precision assembly of 170mm cylindrical components with tolerance of 0.1mm. The lower-level network integrates expert experience through Behavior Cloning (BC), giving the robot human-like intuition, and incorporates the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm to enhance training stability and robustness. The upper-level network dynamically adjusts the lower-level decisions based on heuristic rules, ensuring flexibility in operations. A simulated model is constructed to learn before transferring to real world. An efficient and safe training is allowed. Comparisons show that the reward curve converges within 500 episodes, indicating high learning efficiency. It also demonstrates better adaptability to initial conditions and pose errors, achieving satisfactory success rates even under extreme conditions. Moreover, the method exhibits good stability under Gaussian noise interference. In the real world, the assembly trajectory of the cylindrical segment shows smoother motion and less fluctuation.
Chinese Translation
提出了一种混合分层学习框架,用于实现公差为0.1mm的170mm圆柱形部件的高精度装配。底层网络通过行为克隆(Behavior Cloning, BC)融合专家经验,赋予机器人类人的直觉,并结合双延迟深度确定性策略梯度(Twin Delayed Deep Deterministic Policy Gradient, TD3)算法以提升训练的稳定性和鲁棒性。上层网络基于启发式规则动态调整底层决策,确保操作的灵活性。构建了仿真模型以先学习后迁移到真实世界,从而实现高效且安全的训练。对比实验表明,奖励曲线在500个回合内收敛,显示出较高的学习效率。该方法对初始条件和位姿误差表现出更好的适应性,即使在极端条件下也能取得令人满意的成功率。此外,该方法在高斯噪声干扰下表现出良好的稳定性。在真实世界中,圆柱段的装配轨迹运动更加平滑、波动更小。
cs.RO / 33 / 2609.06547

SHIFT: Surface-aware High-speed Integration For TSDFs

SHIFT:面向TSDF的表面感知高速融合方法
Choudhury, Ayaan, Tiwari, Lokender
Abstract
Real-time 3D mapping is fundamental for autonomous robotic navigation, with Euclidean Signed Distance Fields (ESDFs) serving as the standard representation for online motion planning. While recent advancements in non- projective distance fields yield highly accurate maps, their computational overhead remains a severe bottleneck. Conventional integrators redundantly re-fuse millions of depth pixels every frame, even long after the corresponding voxels have converged, wasting significant computational resources in environments dominated by large planar surfaces. In this paper, we present SHIFT (Surface-aware High-speed Integration For TSDFs), an efficient mapping framework designed to reduce this per-frame update cost. By exploiting structural redundancy directly from 3D depth geometry, SHIFT compresses flat local regions into weighted super-rays and freezes flat-voxel gradients. A compact ESDF voxel layout further reduces the memory footprint of the remaining wavefront. Extensive evaluations across various RGB-D and LiDAR sequences show that SHIFT cuts TSDF cost by 1.42 to 4.07 times, while holding mesh error within millimeters, and reduces ESDF-layer memory by up to 28%
Chinese Translation
实时三维建图是自主机器人导航的基础,欧几里得符号距离场(Euclidean Signed Distance Fields, ESDFs)是在线运动规划的标准表示形式。尽管非投影距离场的最新进展能够生成高度精确的地图,但其计算开销仍然是一个严重的瓶颈。传统的融合器每帧都会冗余地重新融合数百万个深度像素,即使相应体素早已收敛,在以大面积平面表面为主的环境中浪费了大量计算资源。本文提出了SHIFT(Surface-aware High-speed Integration For TSDFs),一种旨在降低每帧更新成本的高效建图框架。通过直接从三维深度几何中利用结构冗余性,SHIFT将平坦的局部区域压缩为加权超射线(super-rays),并冻结平坦体素的梯度。紧凑的ESDF体素布局进一步降低了剩余波前(wavefront)的内存占用。在各种RGB-D和LiDAR序列上的大量评估表明,SHIFT将TSDF计算成本降低了1.42至4.07倍,同时将网格误差保持在毫米级以内,并使ESDF层的内存占用最多降低28%。
cs.RO / 34 / 2609.06591

Unifying Physics-Based Humanoid Interaction with a Context-Conditioned Interaction Prior

基于物理的仿人机器人交互与情境条件化交互先验的统一框架
Li, Jianan, Chen, Xiao, Wong, Tien-Tsin
Abstract
Developing unified physics-based humanoid controllers that can navigate complex 3D scenes and manipulate objects remains a longstanding challenge. Existing approaches are often specialized for either locomotion or object-centric manipulation, or rely on task-specific reward engineering that does not scale well across diverse behaviors. We present CHIP, a unified, physics-grounded framework for learning reusable humanoid interaction skills from heterogeneous motion data. Central to our approach is a conditional interaction prior that models a context-dependent distribution over these skills within a shared discrete space. Our method is trained in three stages. We first learn physics-based motion-imitation policies that acquire grounded teacher behaviors from heterogeneous interaction data. We then distill these behaviors into a context-conditioned interaction prior that captures reusable motion structure across locomotion and manipulation. Finally, we initialize downstream task policies from the pretrained prior and adapt them through prior-regularized online RL post-training. Experiments on a diverse suite of humanoid interaction tasks show that our approach supports scene-aware locomotion, contact-rich object manipulation, and compositional behaviors such as environment-aware object transport and long-horizon skill sequencing, while producing smooth transitions and physically plausible motion.
Chinese Translation
开发能够导航复杂三维场景并操纵物体的统一物理仿真仿人机器人控制器,一直是一项长期挑战。现有方法通常仅针对运动控制或以物体为中心的操纵进行专门设计,或依赖于难以在多样化行为之间良好扩展的任务特定奖励工程。我们提出了CHIP,一个统一的、基于物理的框架,用于从异构运动数据中学习可复用的仿人机器人交互技能。我们方法的核心是一个条件化交互先验,它在共享离散空间中对这些技能的情境依赖分布进行建模。我们的方法分三个阶段训练。首先,我们学习基于物理的运动模仿策略,从异构交互数据中获取具有物理基础的教师行为。然后,我们将这些行为蒸馏到一个情境条件化交互先验中,以捕获运动控制和操纵中可复用的运动结构。最后,我们使用预训练的先验初始化下游任务策略,并通过带先验正则化的在线强化学习后训练对其进行适应。在多样化仿人机器人交互任务套件上的实验表明,我们的方法支持场景感知运动控制、富接触物体操纵,以及环境感知物体运输和长时程技能序列等组合性行为,同时产生平滑的过渡和物理上合理的运动。
cs.RO / 35 / 2609.06615

MemCorr-DP: Counterfactual Correspondence Conditioning for a Diffusion Policy Guided by a Reference

MemCorr-DP:基于参考轨迹引导的扩散策略的反事实对应关系条件化方法
Su, Tan, Yang, Haoxiang, Wang, Ruxin, Xie, Binghui
Abstract
Behavior-cloned visuomotor policies can remain accurate near their training distribution yet fail when object position and camera viewpoint change together. A successful reference trajectory contains the geometry needed to transfer the same interaction, but the policy must align that geometry with the current scene and remain sensitive to it during denoising. To address these challenges, we present MemCorr-DP, a diffusion policy that lifts frozen RoMa v2 matches into explicit 3D relations between the current scene and the reference trajectory. A counterfactual paired objective assigns opposite behaviors the same physical state and noisy action while retaining reference-specific denoising targets. Mixed-condition fine-tuning then adapts the policy from ground-truth geometry to measured correspondence errors. Our strongest evaluation places the Door in the outermost position bands beyond the training support and changes the query camera by $\pm15^\circ$. Under this combined shift, MemCorr-DP achieves 96.67% closed-loop success, compared with 88.00% for a visual Transformer with the same action architecture. Objective ablations and reference interventions show that behavior responds to the selected reference, while matched controls favor the complete relation set over future motion or centroid geometry alone. These results support explicit 3D reference relations as a robust conditioning interface when spatial and viewpoint changes are compounded in the evaluated task.
Chinese Translation
通过行为克隆获得的视觉运动策略在接近训练分布时仍可保持准确性,但当物体位置和相机视角同时发生变化时往往会失效。一条成功的参考轨迹包含了迁移同一交互所需的几何信息,但策略必须将该几何信息与当前场景对齐,并在去噪过程中对其保持敏感。为应对这些挑战,我们提出了 MemCorr-DP,这是一种扩散策略,它将冻结的 RoMa v2 匹配结果提升为当前场景与参考轨迹之间的显式三维关系。反事实配对目标将相反的行为赋予相同的物理状态和带噪动作,同时保留针对参考轨迹的去噪目标。随后,混合条件微调使策略从真实几何过渡到基于测量对应误差的适应。在我们最具挑战性的评估中,将门(Door)放置在超出训练支持范围的最外侧位置区间,并将查询相机改变 ±15°。在这一组合偏移下,MemCorr-DP 实现了 96.67% 的闭环成功率,而采用相同动作架构的视觉 Transformer(Visual Transformer)为 88.00%。目标消融实验和参考干预实验表明,策略行为会对所选参考轨迹作出响应,而对照匹配实验表明完整的关系集合优于仅使用未来运动或质心几何。这些结果表明,在评估任务中空间与视角变化叠加的情况下,显式三维参考关系可作为一种鲁棒的条件化接口。
cs.RO / 36 / 2609.06623

CAVEAT: Recurrent Multimodal Diffusion Planning for Mapless Aerial Exploration

CAVEAT:面向无地图空中探索的循环多模态扩散规划方法
Visch, Steven, Botteghi, Nicolò, Franchi, Antonio, Bazzana, Barbara
Abstract
Can exploratory UAV waypoint sequences be generated from multimodal onboard observations and a fixed-dimensional recurrent internal state without maintaining a persistent global map in the deployed policy? We investigate this question through CAVEAT, a diffusion policy conditioned on a recurrent internal state updated from fused LiDAR, visual, and pose features and trained from trajectories generated by the map-based FUELv2 expert. Rolling inference partially warm-starts consecutive predictions, while a temporary local signed distance field provides heuristic obstacle guidance. Simulation results evaluate both inference mechanisms and compare CAVEAT with its demonstration-generating expert. Proof-of-concept experiments on a Flyability Elios 3 demonstrate partial exploration of a previously unseen indoor environment and target-directed visual servoing using a separately trained policy.
Chinese Translation
探索性无人机航点序列能否在不依赖部署策略中维护的持久全局地图的情况下,由多模态机载观测和一个固定维度的循环内部状态生成?我们通过CAVEAT来研究这一问题。CAVEAT是一种扩散策略,其以由融合的激光雷达(LiDAR)、视觉和位姿特征更新的循环内部状态为条件,并通过基于地图的FUELv2专家生成的轨迹进行训练。滚动推理可对连续预测进行部分热启动,而临时局部符号距离场则提供启发式的障碍物引导。仿真结果评估了两种推理机制,并将CAVEAT与其用于生成示范数据的专家进行了比较。在Flyability Elios 3上的概念验证实验展示了对先前未见过的室内环境的部分探索,以及使用单独训练的策略实现的目标导向视觉伺服。
cs.RO / 37 / 2609.06718

SkillX: Unified Multi-Skill Policy Learning for Humanoid Soccer

SkillX:面向人形机器人足球的统一多技能策略学习
Ye, Zhangchen, Ruan, Enxuan, Bao, Yifei, Huang, Runhan, Yang, Jiankun, Jin, Jiakang, Huo, Yixiao, Wang, Pengyuan, Han, Yinan, Huang, Huaxing, Cui, Wenhao, Li, Yiming, Tian, Xiaoyu
Abstract
Humanoid soccer is a challenging testbed for dynamic whole-body control, requiring robots to coordinate balance, locomotion, object interaction, and skill switching over long horizons. Existing humanoid sports methods often rely on task-specific multi-stage pipelines, making it difficult to jointly learn and compose multiple object-interactive skills within a single deployable policy. To address this, we present SkillX, a unified reinforcement learning framework that learns and composes multiple atomic soccer skills through a single command-conditioned policy. SkillX integrates three core designs: skill-specific adversarial motion priors, skill-specific critics, and an object-aware temporal encoder, enabling the robot to execute atomic skills and transition among them such as dribbling, trapping, and shooting. Experiments in simulation and on a real Noetix E1 humanoid demonstrate robust multi-skill execution, long-horizon skill composition, and successful sim-to-real deployment.
Chinese Translation
人形机器人足球是动态全身控制的一个具有挑战性的测试平台,要求机器人在长时间跨度内协调平衡、运动、物体交互以及技能切换。现有的人形机器人运动方法通常依赖于特定任务的多阶段流水线,难以在单个可部署策略中联合学习并组合多种物体交互技能。为解决这一问题,我们提出了 SkillX,一个统一的强化学习框架,通过单一的条件指令策略(command-conditioned policy)学习并组合多种原子足球技能。SkillX 集成了三项核心设计:技能特定的对抗运动先验(adversarial motion priors)、技能特定的评论家网络(critics)以及物体感知的时间编码器,使机器人能够执行原子技能并在它们之间进行切换,例如运球、停球和射门。在仿真和真实 Noetix E1 人形机器人上的实验展示了鲁棒的多技能执行、长时程技能组合以及成功的仿真到现实(sim-to-real)部署。
cs.RO / 38 / 2609.06810

OCTN: Neural OCT Representations for Robot-Guided Precision Intervention

OCTN:用于机器人引导精准介入的神经OCT表示方法
Prakash, Ravi, McNabb, Ryan P., Codd, Patrick J., Lin, Shan
Abstract
Optical coherence tomography (OCT) offers compact, contactless, micron-scale imaging suitable for intraoperative guidance, but native OCT volumes are discretely sampled, anisotropic, and currently inefficient for downstream geometric reasoning and robot integration. We present OCTN (pronounced "octane"), an implicit neural representation framework that converts volumetric OCT scans into a continuous, differentiable, and spatially faithful tissue-intensity field. OCTN uses a two-stage hybrid training strategy that combines supervision from acquired voxels with inter-slice interpolations, preserving B-scan fidelity while improving continuity in sparsely sampled regions. For versatility, we first show that OCTN enables fast volumetric reasoning by storing the learned tissue representation natively on the GPU, supporting intensity-based spatial queries with up to 43x speedup over conventional CPU processing. We then demonstrate OCTN-enabled OCT-guided robotic laser surgery where the continuous tissue representation supports implicit surface discovery and surface-constrained path planning via multiple optimization strategies, including Newton- and SGD-based optimization. Next, OCTN enables reconstruction of dense volumetric structure from sparsely acquired B-scans, while reducing acquisition time by 4x and preserving clinically relevant structures. Across the newly generated Duke TissueOCT dataset and public OCT datasets, OCTN achieves robust, high-fidelity reconstruction with PSNR > 30 dB and training time < 10 s, while preserving surface consistency within 10 $\mu$m Chamfer distance relative to baseline reconstruction. The TissueOCT dataset and code are publicly available at raprakashvi.github.io/octn
Chinese Translation
光学相干断层扫描(OCT)提供了一种紧凑、非接触、微米级分辨率的成像方式,适用于术中引导,但原始OCT体数据是离散采样且各向异性的,目前难以高效支持下游的几何推理和机器人集成。我们提出了OCTN(发音为"octane"),一种隐式神经表示框架,可将体OCT扫描转换为连续、可微且空间上忠实于组织的强度场。OCTN采用两阶段混合训练策略,将采集体素的监督信号与切片间插值相结合,在保持B扫描保真度的同时,提高了稀疏采样区域的连续性。为展示其通用性,我们首先证明OCTN通过将学习到的组织表示原生存储在GPU上,实现了快速体数据推理,支持基于强度的空间查询,相比传统CPU处理可加速高达43倍。随后,我们展示了基于OCTN的OCT引导机器人激光手术,其中连续的组织表示支持隐式表面发现以及通过多种优化策略(包括基于牛顿法和SGD的优化)进行的表面约束路径规划。其次,OCTN能够从稀疏采集的B扫描中重建稠密的体结构,同时将采集时间缩短4倍,并保留临床相关结构。在新生成的Duke TissueOCT数据集和公开OCT数据集上,OCTN实现了鲁棒的高保真重建,PSNR超过30 dB,训练时间少于10秒,且与基线重建相比表面一致性保持在10微米Chamfer距离以内。TissueOCT数据集和代码已在raprakashvi.github.io/octn上公开。
cs.RO / 39 / 2609.06813

RoboSense: Leveraging Robotaxi Fleets as Drive-by Sensors for Urban Traffic Monitoring

RoboSense:利用Robotaxi车队作为车载式传感器进行城市交通监测
Wang, Yilin, Feng, Yiheng
Abstract
Urban traffic monitoring plays a critical role in safety analysis, congestion management, and incident response. The growing deployment of robotaxis creates a new opportunity for network-level traffic monitoring. Although robotaxis are primarily designed to serve passengers, they can also be leveraged as drive-by sensors to collect traffic data. Compared to conventional infrastructure sensors or probe vehicles, a fleet of robotaxis forms a cooperative perception environment, which can collectively gather spatially and temporally continuous traffic information. This paper proposes a novel dynamic robotaxi routing framework that explicitly incorporates traffic monitoring tasks as an objective. The framework introduces: (1) a cell-based network representation that aligns with sensing capabilities of robotaxis; (2) a cell-level monitoring metric to quantify spatiotemporal robotaxi coverage; and (3) a mixed-integer linear programming (MILP) formulation that jointly minimizes time-dependent travel time and maximizes traffic monitoring performance. A 5 by 5 urban grid network is built in SUMO to evaluate the framework under three robotaxi market penetration rates (2%, 5%, and 10%) with a range of objective weight combinations. Results show that incorporating spatiotemporal network coverage in the objective function can effectively improve the traffic monitoring performance. Interestingly, with appropriate weights between the two objectives, monitoring performance and robotaxi average speed can be improved simultaneously. This suggests better network monitoring leads to more accurate traffic state prediction and improved mobility. This win-win situation could incentivize robotaxi operators to contribute their vehicles as drive-by sensors for traffic monitoring.
Chinese Translation
城市交通监测在安全分析、拥堵管理和事件响应中发挥着关键作用。Robotaxi(自动驾驶出租车)的日益普及为网络级交通监测创造了新的机遇。尽管Robotaxi主要以服务乘客为目的,但它们也可以作为车载式(drive-by)传感器来采集交通数据。与传统的路侧传感器或探测车辆相比,Robotaxi车队构成了一个协同感知环境,能够共同采集时空连续的交通信息。本文提出了一种新颖的动态Robotaxi路径规划框架,将交通监测任务显式地作为优化目标之一。该框架引入了:(1)一种与Robotaxi感知能力相匹配的基于单元(cell-based)的路网表达方法;(2)一种用于量化Robotaxi时空覆盖程度的单元级监测指标;(3)一个混合整数线性规划(MILP)模型,联合最小化时变行程时间并最大化交通监测性能。本研究在SUMO中构建了一个5×5的城市格状路网,在三种Robotaxi市场渗透率(2%、5%和10%)以及多种目标权重组合下对框架进行了评估。结果表明,在目标函数中纳入时空路网覆盖能够有效提升交通监测性能。有趣的是,当两个目标之间采用合适的权重时,监测性能与Robotaxi平均速度可以同时得到提升。这表明更好的路网监测能够带来更准确的交通状态预测和出行效率的改善。这种双赢的局面可以激励Robotaxi运营商将其车辆作为车载式传感器用于交通监测。
cs.RO / 40 / 2609.06820

Diagnosing and Dynamically Filtering Occupancy World Models for Active Mapping

面向主动建图的占据世界模型诊断与动态过滤方法
Zhang, Jiahui, Liang, Gongbo, Zhang, Yu
Abstract
Active mapping requires a robot to select camera viewpoints that efficiently reconstruct an unknown 3D scene. To reason about unobserved regions, recent systems use pretrained occupancy networks as world models that complete missing geometry. The predicted structure contributes to expected coverage gain and constrains feasible robot motion. Consequently, occupancy errors can change both what the robot chooses to explore and where it is able to move. We diagnose these effects by holding the planner fixed and varying only the occupancy representation provided to it. We consider planning without completion, with learned occupancy, with false positives removed by a ground truth oracle, with false negatives restored by an oracle, and with ground truth occupancy. Our experiments show that correcting false positives or false negatives alone does not consistently improve final coverage. This finding reveals a gap between occupancy accuracy and downstream planning performance. Ground truth occupancy provides a much larger improvement in coverage efficiency than in endpoint coverage, suggesting that planning and reachability remain important bottlenecks even when the geometric world model is accurate. Based on these findings, we introduce a dynamic filtering strategy that preserves predictions in unexplored space while suppressing repeatedly unsupported occupancy using online observations. Preliminary examples show that this strategy can redirect viewpoint selection toward reachable surfaces that would otherwise remain unobserved.
Chinese Translation
主动建图要求机器人选择能够高效重建未知三维场景的相机视角。为了对未观测区域进行推理,近期的系统使用预训练的占据网络(occupancy networks)作为世界模型来补全缺失的几何结构。预测的结构有助于提升期望覆盖增益,并约束机器人的可行运动。因此,占据预测误差可能会同时改变机器人选择探索的区域及其可移动的范围。我们通过固定规划器、仅改变提供给规划器的占据表示来诊断这些影响。我们考虑了以下几种情况:不使用补全的规划、使用学习到的占据表示的规划、通过真值(ground truth) oracle 去除假阳性后的规划、通过 oracle 恢复假阴性后的规划,以及使用真值占据的规划。我们的实验表明,仅纠正假阳性或假阴性并不能持续地提高最终覆盖率。这一发现揭示了占据预测精度与下游规划性能之间的差距。真值占据在覆盖效率方面带来的提升远大于在末端覆盖率(endpoint coverage)方面的提升,这表明即使在几何世界模型准确的情况下,规划与可达性仍然是重要的瓶颈。基于这些发现,我们提出了一种动态过滤策略,该策略在未探索空间中保留预测结果,同时利用在线观测抑制反复得不到支持的占据预测。初步示例表明,该策略能够将视角选择重新引导至原本无法被观测到的可达表面上。
cs.RO / 41 / 2609.06852

ContextFlow: In-Context Flow Matching for Robot Manipulation

ContextFlow:面向机器人操作的上下文流匹配方法
Ding, Jian, Dai, Xianjie, Herzig, Roei, Hroub, Nussair, Mai, Jinjie, Dai, Dengxin, Ghanem, Bernard, Elhoseiny, Mohamed
Abstract
Although highly effective in vision and language domains, applying in-context learning to robotics remains challenging. Existing autoregressive in-context imitation methods discretize continuous actions and exacerbate the accumulation of early prediction errors through next-token prediction, limiting their generalization on unseen task configurations. Meanwhile, flow-matching policies have been explored for continuous robot control and can help mitigate compounding errors; however, in-context imitation learning within a flow-matching framework remains underexplored. To address these limitations, we introduce ContextFlow, a conditional flow-matching model that learns continuous action distributions for in-context imitation learning. ContextFlow conditions flow-based action prediction on demonstrations and observations, enabling robust generation from noisy action distributions. To better encode multimodal in-context demonstrations, we adapt perceiver-style multimodal context compressors that distill visual, proprioceptive, and action sequences into compact, task-relevant latent representations. On LIBERO, ContextFlow outperforms ICRT by 35 percentage points in average success rate on unseen task configurations, while matching the performance of the task-specific fine-tuned VLA model $\pi_0$ without any fine-tuning on unseen tasks. On real robots, it generalizes to unseen configurations of both single-arm and bimanual tasks, achieving 40% success on a new pen-uncapping configuration. Project Page: https://dingjiansw101.github.io/contextflow-page/.
Chinese Translation
尽管上下文学习(in-context learning)在视觉和语言领域高度有效,但将其应用于机器人领域仍然充满挑战。现有的自回归式上下文模仿学习方法将连续动作离散化,并通过下一词元预测加剧了早期预测误差的累积,限制了其在未见任务配置上的泛化能力。与此同时,流匹配(flow-matching)策略已被探索用于连续机器人控制,有助于缓解复合误差;然而,在流匹配框架内进行上下文模仿学习的研究仍然不足。为解决这些局限,我们提出了ContextFlow,一种用于上下文模仿学习的条件流匹配模型,用于学习连续动作分布。ContextFlow 将基于流的动作预测条件化于演示和观测之上,从而能够从含噪的动作分布中实现鲁棒的生成。为了更好地编码多模态上下文演示,我们采用感知器(perceiver)风格的多模态上下文压缩器,将视觉、本体感觉和动作序列蒸馏为紧凑且与任务相关的潜在表示。在LIBERO基准上,ContextFlow在未见任务配置上的平均成功率比ICRT高出35个百分点,并且在未对未见任务进行任何微调的情况下,达到了任务特定微调的VLA模型π₀的性能水平。在真实机器人上,该方法能够泛化到未见配置的单臂和双臂任务,在新的开笔帽任务配置上达到40%的成功率。项目主页:https://dingjiansw101.github.io/contextflow-page/。
cs.RO / 42 / 2609.06888

Predictive-Coding-Based Autonomous Regulation of Internally Generated and Externally Coupled Processing in Human-Robot Interaction

基于预测编码的人机交互中内部生成与外部耦合处理的自主调节
Oyama, Henrique, Sawada, Hiroki, Tani, Jun
Abstract
Predictive coding characterizes adaptive behavior as a dynamic balance between internally generated predictions and external sensory evidence, yet how an embodied cognitive system can regulate this balance online during ongoing interaction remains poorly understood. This study proposes a predictive-coding-based mechanism for regulating internally generated and externally coupled processing during physical human--robot interaction. The framework employs a predictive-coding-inspired variational recurrent neural network (PV-RNN), in which a meta-prior controls the degree to which posterior inference is constrained by learned prior dynamics. We extend this architecture with an online mechanism that uses reconstruction error accumulated over recent interaction history to select between predefined meta-prior regimes. The mechanism was evaluated across three physical human--robot interaction tasks involving fixed structured, changing structured, and less-constrained interaction. Across all tasks, lower meta-prior values produced the expected increase in posterior--prior divergence and reduction in reconstruction error. More importantly, reconstruction-history-driven regime selection was also associated with reduced prospective prediction error and robot-side physical interaction conflict, demonstrating consequences beyond the retrospective reconstruction objective itself. Task~3 further showed that recent sensory observations can be successfully accommodated while subsequent human motion still departs from the model's prior-generated future trajectory. Overall, these findings show that accumulated reconstruction mismatch can provide an endogenous signal for regulating how strongly subsequent inference relies on learned internal dynamics relative to ongoing sensory input during embodied interaction.
Chinese Translation
预测编码将适应性行为刻画为内部生成的预测与外部感觉证据之间的动态平衡,然而具身认知系统在持续交互过程中如何在线调节这种平衡仍知之甚少。本研究提出了一种基于预测编码的机制,用于在物理人机交互过程中调节内部生成处理与外部耦合处理。该框架采用受预测编码启发的变分循环神经网络(PV-RNN),其中一个元先验(meta-prior)控制后验推断受已学习先验动力学约束的程度。我们对该架构进行了扩展,增加了一种在线机制,利用近期交互历史中积累的重构误差在预定义的元先验模式之间进行选择。该机制在三个物理人机交互任务中进行了评估,分别涉及固定的结构化交互、变化的结构化交互以及约束较少的交互。在所有任务中,较低的元先验值均产生了预期的后验—先验散度增加和重构误差降低。更重要的是,基于重构历史的模式选择还与前瞻性预测误差和机器人侧物理交互冲突的降低相关,表明其效果超越了回溯性重构目标本身。任务3进一步表明,在后续人类运动仍偏离模型先验生成的未来轨迹的情况下,近期的感觉观测仍可被成功吸纳。总体而言,这些发现表明,积累的重构失配可以为具身交互过程中调节后续推断对已学习内部动力学的依赖强度(相对于持续感觉输入)提供一种内生信号。
cs.RO / 43 / 2609.06896

Distributed Secure Learning Control for Large-scale Multirobots under Stealthy Actuator Attacks

隐蔽执行器攻击下大规模多机器人的分布式安全学习控制
Zhang, Xinglong, Ma, Qingwen, Li, Cong, Yin, Hui, Zhang, Changxin, Wang, Yueying, Pan, Wei, Xu, Xin
Abstract
Distributed learning control for multirobot systems (MRS) offers significant flexibility in presence of uncertainties but lacks provable performance guarantees. A promising direction involves integrating reinforcement learning (RL) into distributed model predictive control (DMPC), leveraging the strengths of RL in nonlinear policy design and the receding-horizon replanning capabilities of DMPC. However, ensuring secure control within such a learning framework under malicious cyber attacks, particularly stealthy ones, remains a critical challenge, because the distributed policies generation depends on information exchange among neighbors, where compromised agents can rapidly influence the behavior of others through the communication network. This article proposes a distributed secure learning control (DSLC) framework for large-scale MRS under malicious, stealthy actuator attacks. Our framework offers two key features: (i) a unified approach that enables secure learning control across various coordination scenarios and (ii) a game-theoretic distributed learning-based predictive control strategy that learns how to balance the attacker and defender through a differential-game based DMPC framework. Specifically, DSLC employs a distributed attacker-actor-critic architecture to learn the optimal defense and attack policies online within each prediction interval. Unlike numerical optimization-based controllers that calculate open-loop control sequences, our method simultaneously generates adversarial attack policies and corresponding defense policies in analytical closed-loop form. The defense policies could be directly generalized to MRS with varying scales and diverse actuator attack probabilities. The effectiveness and scalability of DSLC are validated through comprehensive simulations and real-world experiments in multiple wheeled robots via various control tasks.
Chinese Translation
多机器人系统(MRS)的分布式学习控制在存在不确定性时具有显著的灵活性,但缺乏可证明的性能保证。一个有前景的方向是将强化学习(RL)融入分布式模型预测控制(DMPC),利用强化学习在非线性策略设计方面的优势以及DMPC的滚动时域重规划能力。然而,在此类学习框架下,如何确保控制在恶意网络攻击(尤其是隐蔽攻击)下的安全性仍然是一个关键挑战,因为分布式策略的生成依赖于邻居间的信息交换,被攻陷的智能体可以通过通信网络迅速影响其他智能体的行为。本文针对恶意隐蔽执行器攻击下的大规模多机器人系统,提出了一种分布式安全学习控制(DSLC)框架。我们的框架具有两个关键特点:(i)一种统一的方法,可在多种协同场景下实现安全学习控制;(ii)一种基于博弈论的分布式学习预测控制策略,通过基于微分博弈的DMPC框架学习如何平衡攻击者与防御者。具体而言,DSLC采用分布式攻击者-行动者-评论家(actor-critic)架构,在每个预测区间内在线学习最优防御与攻击策略。与计算开环控制序列的数值优化控制器不同,我们的方法以解析闭环形式同时生成对抗性攻击策略及相应的防御策略。所得到的防御策略可以直接泛化至具有不同规模和多样化执行器攻击概率的多机器人系统。通过多个轮式机器人在多种控制任务下的全面仿真和真实实验,验证了DSLC的有效性和可扩展性。
cs.RO / 44 / 2609.06923

DriftParking: Trajectory Modeling via Drifting Field for End-to-End Automated Parking

DriftParking:基于漂移场的轨迹建模方法用于端到端自动泊车
Wang, Ziyan, Li, Dong, Wang, Weibo, Lu, Yinyin, Xie, Jiayu, Gu, Jiamao, Zhang, Dongpeng
Abstract
Automated parking requires generating complete and executable trajectories in highly constrained spaces with low tolerance for goal pose error. Existing end-to-end parking methods struggle to jointly achieve inference efficiency, trajectory quality, and precise endpoint alignment, while conventional imitation objectives provide limited supervision on structured deviations from expert maneuver geometry. We propose DriftParking, a one-step trajectory generation framework that reconstructs the drifting-field paradigm for high-precision conditional trajectory generation. Specifically, we replace distribution-level attraction with conditional one-to-one attraction toward the paired expert trajectory, introduce expert-centered constructive repulsion, and adaptively attenuate repulsion near convergence. We further formulate trajectory generation in an endpoint-residual space by decomposing each trajectory into a start-to-goal baseline and a learnable residual, turning endpoint alignment into a representation-level structural constraint on the supervision target while providing a structured space for repulsive supervision. DriftParking achieves state-of-the-art performance across all evaluation metrics. Closed-loop on-vehicle experiments across diverse parking scenarios further show a 97% parking success rate, demonstrating strong zero-shot generalization.
Chinese Translation
自动泊车需要在高度受限的空间内生成完整且可执行的轨迹,并对目标位姿误差具有极低的容忍度。现有的端到端泊车方法难以同时兼顾推理效率、轨迹质量以及精确的终点对齐,而传统的模仿学习目标对偏离专家操控几何结构的结构化偏差所能提供的监督也十分有限。我们提出DriftParking,这是一个单步轨迹生成框架,它重构了漂移场范式以实现高精度的条件轨迹生成。具体而言,我们用指向配对专家轨迹的条件性一对一吸引力取代分布层面的吸引力,引入以专家轨迹为中心的构造性排斥力,并在接近收敛时自适应地衰减排斥力。我们进一步在终点残差空间中构建轨迹生成任务,将每条轨迹分解为从起点到目标的基线和一个可学习的残差,从而将终点对齐转化为监督目标在表示层面的结构性约束,同时为排斥性监督提供了结构化的空间。DriftParking在所有评估指标上均取得了最先进的性能。在多种泊车场景下进行的闭环实车实验进一步显示其泊车成功率达到97%,展现出强大的零样本泛化能力。
cs.RO / 45 / 2609.06930

Distributed Dexterous Manipulation with Spatially Conditioned Multi-Agent Transformers

基于空间条件化多智能体Transformer的分布式灵巧操作
Patil, Sarvesh
Abstract
Distributed Dexterous Manipulation (DDM) is a novel paradigm that presents significant control challenges due to high action-space redundancy, inter-robot cooperation, and dynamic object-robot interactions. This paper introduces a framework based on spatially conditioned Multi-Agent Transformers (MATs) to efficiently learn robust control policies for a DDM system grounded in an array of 64 soft delta robots arranged in an 8x8 grid. Our three core contributions are: (i) an MAT with adaptive layer norm for compute efficiency, (ii) spatial contrastive embeddings to ground transformer embeddings in the spatial configuration of the robots, and (iii) an MAT-based behavior cloning method fine-tuned using Soft Actor Critic. We also propose an action selection formulation to analyze the trade-off between task performance and the number of robots utilized. Our experiments show that MATs iteratively refine their actions through the stacked attention blocks. This further informs the benefit of spatial conditioning in transformers to learn DDM policies. We demonstrate long-horizon planar manipulation tasks with objects of various geometries in simulation and real-world. Finally, we show how action selection mitigates robot maintenance by reducing wear and tear due to inter-robot collisions while maintaining the ability to manipulate objects along various trajectories in the real-world, achieving an average error of ~1.5 cm, while using ~65% fewer robots.
Chinese Translation
分布式灵巧操作(Distributed Dexterous Manipulation, DDM)是一种新型范式,由于其动作空间的高度冗余、机器人间的协作以及动态的物体-机器人交互,带来了显著的控制挑战。本文提出了一个基于空间条件化多智能体Transformer(Multi-Agent Transformers, MATs)的框架,以高效学习鲁棒的控制策略,该DDM系统由排列成8x8网格的64个软体Delta机器人阵列构成。我们的三项核心贡献是:(i) 采用自适应层归一化以提高计算效率的MAT;(ii) 将Transformer嵌入与机器人空间配置相结合的空间对比嵌入;(iii) 一种基于MAT的行为克隆方法,并使用Soft Actor Critic进行微调。我们还提出了一种动作选择建模方法,用于分析任务性能与所用机器人数量之间的权衡。实验表明,MATs通过堆叠的注意力模块迭代地细化其动作。这进一步证明了空间条件化在Transformer中学习DDM策略的有效性。我们在仿真和真实环境中演示了针对不同几何形状物体的长时序平面操作任务。最后,我们展示了动作选择如何在保持沿各种轨迹操作物体的真实环境能力的同时,减少机器人间碰撞造成的磨损,从而缓解机器人维护问题,在减少约65%机器人使用量的情况下,实现了约1.5厘米的平均误差。
cs.RO / 46 / 2609.06958

Mind the Phase: Effective Rank and Representation Health in Legged Locomotion

注意相位:足式运动中的有效秩与表征健康度
Tommaselli, Felipe, Segreto, Thiago H., Negri, Juliano D., Godoy, Ricardo V., Becker, Marcelo
Abstract
Reinforcement learning has become the leading paradigm in legged locomotion, enabling complex behaviors from backflips to parkour through massively parallel simulation. Under PPO's non-stationarity, shallow networks remain the de facto architecture, supported by carefully staged curricula and environments, yet the representations these policies learn stay poorly understood, leaving no training-time signal of how they will behave on hardware. In this work, we empirically study locomotion policies through the effective rank of the policy Jacobian and show that conditioning rank on the gait phase exposes architectural structure that global rank averages away. In particular, we find that standard architectural choices, namely layer normalization and residual connections, allocate roughly two more dimensions of effective rank to swing than to stance, which is fully absent in vanilla MLPs. Building on this, we propose a simple recipe that turns these representational signatures into smoother, more reliable sim-to-real transfer. In practice, this results in roughly 3x lower joint jitter that holds from simulation onto a physical Spot, suggesting that representation health is an effective training-time lens to track sim-to-real smoothness.
Chinese Translation
强化学习已成为足式运动领域的主流范式,通过大规模并行仿真实现了从后空翻到跑酷等复杂行为。在 PPO 的非平稳性条件下,浅层网络仍是事实上的标准架构,并辅以精心设计的课程学习和环境设置,然而这些策略所学到的表征仍缺乏深入理解,也没有训练时的信号能够预测其在硬件上的表现。在本工作中,我们通过策略雅可比矩阵的有效秩对运动策略进行实证研究,结果表明将秩条件化于步态相位能够揭示被全局秩平均所掩盖的架构结构。特别地,我们发现标准架构选择——即层归一化(Layer Normalization)和残差连接——分配给摆动相的有效秩维度比支撑相大约多两个,而这种差异在原始 MLP 中完全不存在。基于此,我们提出了一种简单的方法,将这些表征特征转化为更平滑、更可靠的仿真到现实(sim-to-real)迁移。在实践中,这使关节抖动降低了约 3 倍,且该效果从仿真一直保持到物理 Spot 机器人上,这表明表征健康度是跟踪仿真到现实平滑性的有效训练时分析视角。
cs.RO / 47 / 2609.07002

WM-Craftnet: World Synesthesia Model for Generalizable and Robust Dexterous In-Hand Manipulation

WM-Craftnet:面向泛化且鲁棒灵巧手内操作的世界联觉模型
Yin, Jie, Zhao, Zeyuan, Tan, Xiaojing, Liu, Yang, Wang, Chiyu, Gu, Xinyang
Abstract
Generalizable and robust dexterous in-hand manipulation requires a policy to infer object pose, geometry, contact, and potential slip from partial and noisy observations. Although recent tactile and visuotactile RL methods achieve strong in-hand rotation in controlled settings, their robustness often degrades under pose shifts, force disturbances, and object variation. We propose WM-Craftnet, a world-model-conditioned framework that learns compact action-conditioned latent dynamics from proprioception, depth, tactile sensing, and actions, supervised by multimodal reconstruction and reward prediction. Rather than using the world model for latent imagination or policy optimization, WM-Craftnet uses the learned World Synesthesia Model (WSM) as recurrent task context for an asymmetric actor--critic policy. Importantly, WSM is trained to reconstruct clean depth targets from noisy depth inputs, providing a denoised geometric state for real-robot deployment. Ablations over recurrent baselines, auxiliary heads, tactile masking, and WSM modality heads show that predictive world modeling, clean-depth supervision, and tactile contact cues all shape the learned state. A WSM pretrained on nine \(z\)-axis objects serves as a reusable prior for \(49\)-object downstream policy learning. This context improves multi-object rotation, with quantitative and qualitative evidence for unseen-object, perturbation-recovery, and sim-to-real transfer.
Chinese Translation
泛化且鲁棒的灵巧手内操作要求策略能够从部分且带噪声的观测中推断物体的位姿、几何形状、接触状态以及潜在滑动。尽管近期的触觉和视触觉强化学习方法在受控环境中实现了较强的手内旋转能力,但其鲁棒性在位姿变化、力扰动和物体变化情况下往往会退化。我们提出 WM-Craftnet,这是一个以世界模型为条件的框架,它从本体感觉、深度、触觉感知和动作中学习紧凑的动作条件潜空间动力学,并通过多模态重建和奖励预测进行监督。与将世界模型用于潜空间想象或策略优化不同,WM-Craftnet 将学习到的世界联觉模型(World Synesthesia Model, WSM)作为非对称 actor-critic 策略的循环任务上下文。重要的是,WSM 被训练为从带噪声的深度输入中重建干净的深度目标,为真机部署提供去噪后的几何状态。针对循环基线、辅助任务头、触觉掩蔽和 WSM 模态头的消融实验表明,预测性世界建模、干净深度监督和触觉接触线索共同塑造了学习到的状态。在九个 z 轴物体上预训练的 WSM 可作为可复用的先验,用于 49 个物体的下游策略学习。该上下文提升了多物体旋转能力,并在未见物体、扰动恢复和仿真到真实迁移方面提供了定量与定性的证据。
cs.RO / 48 / 2609.07047

MEMOBench: A Process Level Memory Benchmark for Robotic Manipulation

MEMOBench:面向机器人操作的过程级记忆基准测试
Sun, Haiyang, Wang, Haoxiao, Chen, Junming, Fang, Weicheng, Su, Zihao, Yi, Jingkun, Yi, Wenyou, Chen, Hao, Zhao, Zhou
Abstract
Robotic manipulation often requires acting on information that is no longer visible, yet Vision-Language-Action policies are usually evaluated when the current observation largely determines the next action. Existing robotic memory benchmarks expose this gap, but they still rely mainly on final task success and therefore conflate forgetting with manipulation failure. We present \textbf{MEMOBench}, a benchmark for process level memory evaluation in robotic manipulation. MEMOBench includes 30 history dependent tasks, 1{,}500 expert demonstrations, and 4{,}200 executable checkpoint instances from 84 templates. Each checkpoint pairs coarse to fine language with a simulator predicate and labels one memory operation: Storage, Update, or Compression. These annotations define Memory Storage Rate, Memory Update Rate, and Memory Compression Rate, which measure memory fidelity alongside task success. Across standard and memory augmented VLA policies, the strongest memory module baseline reaches only 31.9\% average success rate, and high storage often coexists with weak update and compression. Checkpoint language also supervises semantic, contrastive, and framewise memory alignment objectives, yielding modest gains across different memory operations. MEMOBench provides a diagnostic evaluation suite and training supervision for memory grounded robotic policies. The project page is available at https://github.com/Collab-Gen/MEMOBench.
Chinese Translation
机器人操作常常需要依据当前已不可见的信息采取行动,然而视觉-语言-动作(Vision-Language-Action, VLA)策略通常是在当前观测基本能决定下一动作的场景下进行评估的。现有的机器人记忆基准虽揭示了这一差距,但仍主要依赖最终任务成功率,因而将记忆遗忘与操作失败混为一谈。我们提出了MEMOBench,一个用于机器人操作过程级记忆评估的基准。MEMOBench包含30个依赖历史的任务、1,500条专家演示,以及来自84个模板的4,200个可执行检查点实例。每个检查点将粗粒度到细粒度的语言描述与模拟器谓词配对,并标注一种记忆操作类型:存储(Storage)、更新(Update)或压缩(Compression)。这些标注定义了记忆存储率、记忆更新率和记忆压缩率,用于在任务成功率之外衡量记忆保真度。在标准及记忆增强的VLA策略上,最强的记忆模块基线平均成功率仅达到31.9%,且高存储表现往往与薄弱的更新和压缩能力并存。检查点语言还可监督语义、对比以及逐帧记忆对齐目标,在不同记忆操作上带来了适度的性能提升。MEMOBench为基于记忆的机器人策略提供了诊断性评估套件和训练监督。项目页面见 https://github.com/Collab-Gen/MEMOBench。
cs.RO / 49 / 2609.07049

Large Discrete Policy: Advancing Explicit Behavior Modeling with Stochastic Iterative Scoring

大型离散策略:基于随机迭代评分推进显式行为建模
Li, Zhenxin, Chang, Nadine, Sun, Xinglong, Chen, Jingde, Yao, Wenhao, Wang, Zi, Shen, Maying, Jiang, Yu-Gang, Wu, Zuxuan, Lan, Shiyi, Alvarez, Jose M.
Abstract
Behavior policies are often formulated as continuous generative models, whose iterative denoising processes are expressive but difficult to interpret and prone to producing implausible actions. We propose the Large Discrete Policy (LDiP), a fully discrete behavior modeling framework that selects actions from a large vocabulary of physically plausible candidates. Rather than perturbing actions, LDiP improves expressivity through stochastic iterative scoring: it progressively re-scores and prunes candidates with score-space stochasticity, enabling fine-grained ranking and exploration among plausible actions while preserving an explicit decision process. Across end-to-end planning, closed-loop driving, robotic manipulation, and vision-language-action settings, LDiP consistently outperforms strong discrete and continuous baselines in autonomous driving, and exceeds or matches continuous generative policies in robotic manipulation. These results show that discrete policies, when equipped with effective scoring mechanisms, offer an expressive, plausible, and interpretable alternative for behavior modeling. Project website: https://zhenxinli.net/LargeDiscretePolicy/.
Chinese Translation
行为策略通常被构建为连续生成模型,其迭代去噪过程虽然具有很强的表达能力,但难以解释,且容易产生不合常理的动作。我们提出了大型离散策略(Large Discrete Policy,LDiP),这是一个完全离散的行为建模框架,从大量物理上合理可行的候选动作词汇中进行选择。LDiP 并非对动作施加扰动,而是通过随机迭代评分(stochastic iterative scoring)来提升表达能力:它利用评分空间中的随机性,逐步对候选动作进行重新评分与剪枝,从而在合理动作之间实现细粒度的排序与探索,同时保持显式的决策过程。在端到端规划、闭环驾驶、机器人操作以及视觉-语言-动作(vision-language-action)等多种设定下,LDiP 在自动驾驶任务中持续优于强大的离散与连续基线方法,并在机器人操作任务中超越或媲美连续生成式策略。这些结果表明,配备有效评分机制的离散策略,为行为建模提供了一种兼具表达性、合理性与可解释性的替代方案。项目网站:https://zhenxinli.net/LargeDiscretePolicy/。
cs.RO / 50 / 2609.07091

Human-Aware Target Tracking and Navigation: Fusing Kinematic State Estimation with Structural Map Constraints

基于人本感知的目标跟踪与导航:融合运动学状态估计与结构化地图约束
Gupta, Sagar, Gideon, Don, Loke, Seng W., Lee, Kevin, Sebastian, Bijo
Abstract
Autonomous mobile robots performing person-following tasks often suffer from temporary occlusions and sensor track loss in dynamic environments. This research presents an end-to-end autonomous navigation stack that addresses target occlusion through map-informed spatial reasoning. The proposed system features a multi-modal perception pipeline, fusing deep learning-based visual tracking with 2-dimensional LiDAR point clustering to maintain high-fidelity tracking of a tagged person. A continuous state estimator integrates this perception data with wheel odometry and IMU sensors for stable localization. When the active track is lost due to occlusion, the system activates a map-based recovery framework. Leveraging a predefined topological map, the system executes a graph-based search to propagate the target's last known trajectory along structurally defined walking lanes, adhering to left-hand regional conventions. By generating a discrete set of feasible future trajectories, the robot reasons about potential structural trajectory changes, such as continuing a heading or turning at an intersection. This map-informed prediction is fed directly to the local obstacle avoidance planner, enabling the robot to continue following its target safely and predictably until the person is visually reacquired. Real-world evaluations in dense multi-person environments demonstrate the system's robustness, achieving a 71.4\% target reacquisition success rate during major occlusion events lasting up to 7 seconds.
Chinese Translation
执行人员跟随任务的自主移动机器人在动态环境中常常受到临时遮挡和传感器跟踪丢失的困扰。本研究提出了一种端到端的自主导航系统,通过基于地图的空间推理来解决目标遮挡问题。所提出的系统具有多模态感知流水线,将基于深度学习的视觉跟踪与二维激光雷达(LiDAR)点云聚类相融合,以保持对被标记人员的高保真跟踪。一个连续状态估计器将感知数据与轮式里程计和惯性测量单元(IMU)传感器相结合,实现稳定的定位。当活跃跟踪因遮挡而丢失时,系统启动基于地图的恢复框架。利用预定义的拓扑地图,系统执行基于图的搜索,沿结构化定义的行走通道传播目标的最后已知轨迹,并遵循左侧行走的地域惯例。通过生成一组离散的可行未来轨迹,机器人能够对潜在的结构性轨迹变化进行推理,例如继续沿当前方向行进或在交叉口转弯。这种基于地图的预测被直接输入到局部避障规划器中,使机器人能够安全且可预测地继续跟随目标,直到重新在视觉上捕获目标人员。在高密度多人环境中的真实世界评估验证了系统的鲁棒性,在持续时间长达7秒的重大遮挡事件中实现了71.4%的目标重捕获成功率。
cs.RO / 51 / 2609.07096

RoboDreamer: Anticipatory Humanoid Locomotion with Predictive State-Space Models

RoboDreamer:基于预测状态空间模型的 anticipating 人形机器人运动控制
Li, Zhe, Wei, Yangyang, Yuan, Xichen, Zhang, Zhenzhe, Yuan, Weihao, Zhang, Shanghang, Yang, Jianfei
Abstract
Humanoid locomotion requires control policies that remain stable under imperfect sensing while exploiting temporal context for consistent motion. We present RoboDreamer, a two-stage teacher--student framework that combines next-observation consistency with randomized continuous temporal masking. A teacher is first trained on clean observations, and a student is then distilled under masked recent observations, encouraging the policy to infer missing current information from history. At inference, the same masking interface is reused for implicit closed-loop action refinement and optional multi-step action chunking. Mamba is used as the temporal backbone, while matched ablations show that masking/distillation provides a substantial part of the gain and Mamba contributes additional tracking improvements with real-time latency. Experiments in IsaacLab, MuJoCo, and on a Unitree G1 demonstrate robust motion tracking under observation masking and successful real-world deployment.
Chinese Translation
人形机器人运动控制要求策略在不完美感知下保持稳定,同时利用时间上下文以实现连贯的运动。我们提出了RoboDreamer,这是一个两阶段的教师-学生框架,将下一观测一致性与随机化连续时间掩码相结合。首先在干净的观测上训练教师策略,随后在近期观测被掩码的条件下对学生策略进行蒸馏,促使策略从历史信息中推断缺失的当前信息。在推理阶段,同样的掩码接口被复用于隐式闭环动作精化以及可选的多步动作分块(action chunking)。我们采用Mamba作为时间序列骨干网络,而对照消融实验表明,掩码/蒸馏贡献了大部分性能增益,Mamba则在实时延迟下额外提升了跟踪性能。在IsaacLab、MuJoCo以及Unitree G1机器人上的实验证明了该方法在观测掩码下具有鲁棒的运动跟踪能力,并成功实现了真实世界部署。
cs.RO / 52 / 2609.07109

Eventually Optimal and Scalable Multi-Agent Planning for Block Cave Mining

面向崩落采矿的渐近最优且可扩展的多智能体规划
Leet, Christopher, Forte, Paolo, Köckemann, Uwe, Andreasson, Henrik, Koenig, Sven
Abstract
Automation in underground mining has the potential to significantly enhance safety, operational efficiency, and sustainability. However, effectively coordinating fleets of autonomous vehicles in dynamic mine environments introduces substantial challenges in both optimization and motion planning. To address these challenges, we introduce and formalize the \emph{Block Cave Mining (BCM)} problem, which focuses on computing a transport plan that maximizes ore throughput while satisfying draw ratio constraints. To solve this problem, we propose SAMM, an eventually optimal anytime solver that jointly integrates task assignment, scheduling, and path planning via a mixed-integer linear programming formulation. To improve scalability, we also introduce SAMMS, a variant of SAMM that trades optimality guarantees for efficiency by decomposing the problem into shorter planning subcycles. Experimental evaluations using realistic industrial mine scenarios demonstrate that SAMMS achieves near-optimal throughput and scales effectively to larger fleets and mine layouts.
Chinese Translation
地下采矿自动化有潜力显著提升安全性、运营效率和可持续性。然而,在动态矿山环境中有效协调自动驾驶车辆车队,给优化和运动规划带来了重大挑战。为应对这些挑战,我们提出并形式化了崩落采矿(Block Cave Mining, BCM)问题,该问题聚焦于计算一个在满足放矿比例约束的同时最大化矿石产量的运输计划。为求解该问题,我们提出了SAMM,一种渐近最优的任意时间求解器,通过混合整数线性规划形式化,联合整合了任务分配、调度与路径规划。为提升可扩展性,我们还引入了SAMMS,即SAMM的一种变体,通过将问题分解为更短的规划子周期,以牺牲最优性保证来换取效率。基于真实工业矿山场景的实验评估表明,SAMMS能够达到接近最优的矿石产量,并可有效扩展至更大规模的车队和矿山布局。
cs.RO / 53 / 2609.07111

From LLM-Generated Specifications to Learned Quadruped Locomotion

从大语言模型生成的规约到学习型四足机器人运动控制
Atasever, Merve, Azbijari, Keyan, Bakirci, Cagan, Corona, Alfredo Reina, Izdas, Tolga, Yang, Richard, Biyik, Erdem, Deshmukh, Jyotirmoy V.
Abstract
Quadruped robot locomotion policies are often trained using reinforcement learning, which in turn relies heavily on hand-crafted reward functions. Designing reward functions requires substantial manual engineering, and it is often unclear which local rewards will induce the desired global behavior. Shaped rewards from formal specifications in languages like Signal Temporal Logic (STL) can make rewards more interpretable, but writing STL specifications itself still requires domain expertise. We study whether large language models (LLMs) can fill this gap by generating Parametric Signal Temporal Logic (PSTL) specifications that are subsequently used for policy learning. Given a natural language locomotion objective and a constrained specification grammar, GPT-5.5 and Qwen 3.6 independently propose STL templates for command tracking, safety, and gait structure. We instantiate the parameters of the generated PSTL templates using expert trajectories and retain only specifications that are consistent with demonstrated expert behavior. The resulting specifications are then transformed into smooth, finite-history reward functions and used to train a quadruped locomotion policy with Proximal Policy Optimization (PPO) in MuJoCo XLA (MJX). We evaluate both \emph{gait-aware} and \emph{gait-agnostic} settings. The former specifies walking-trot, trot, and bound regimes, while the latter allows contact patterns to emerge from the task objective. We compare against hand-engineered rewards, Text2Reward-style LLM-generated reward code, and an expert-switching oracle. Gait-aware Qwen 3.6 specifications achieved 100\% survival and command success across all tested speeds (0.3--2.1 m/s) and matched the target gait at high speeds, whereas Text2Reward achieved 0\% for both metrics at $\geq 1.9$ m/s. Videos: https://stl-locomotion.github.io/
Chinese Translation
四足机器人的运动控制策略通常采用强化学习训练,而这高度依赖人工设计的奖励函数。奖励函数的设计需要大量的人工工程工作,且往往难以确定哪些局部奖励能够诱导出期望的全局行为。基于信号时序逻辑(STL)等形式化规约语言构建的整形奖励可以使奖励更具可解释性,但编写STL规约本身仍然需要领域专业知识。我们研究大语言模型(LLM)能否填补这一空白,即通过生成参数化信号时序逻辑(PSTL)规约用于后续的策略学习。在给定自然语言运动目标和受限规约文法的条件下,GPT-5.5和Qwen 3.6独立地为指令跟踪、安全性和步态结构提出STL模板。我们利用专家轨迹对生成的PSTL模板的参数进行实例化,并仅保留与专家示范行为一致的规约。随后,将这些规约转化为平滑的有限历史奖励函数,并在MuJoCo XLA(MJX)中使用近端策略优化(PPO)训练四足机器人运动控制策略。我们评估了“步态感知”和“步态无关”两种设置:前者指定了行走快步、快步和跳跃等步态模式,后者则允许接触模式从任务目标中自然涌现。我们将所提方法与人工设计的奖励、Text2Reward风格的LLM生成奖励代码以及专家切换oracle进行了比较。步态感知设置下,Qwen 3.6生成的规约在所有测试速度(0.3–2.1 m/s)下均达到100%的存活率和指令成功率,并在高速下匹配了目标步态;而Text2Reward在速度≥1.9 m/s时两项指标均为0%。视频:https://stl-locomotion.github.io/
cs.RO / 54 / 2609.07116

MAC-I$^2$: Learned Metrics-Aware Covariance for Robust Visual-Inertial Fusion in Initialization and Calibration

MAC-I$^2$:用于初始化与标定中鲁棒视觉-惯性融合的学习型度量感知协方差
Fei, Xiang, Qiu, Yuheng, Xu, Can, Chen, Yutian, Li, Ruogu, Zuo, Xingxing, Wang, Wenshan, Scherer, Sebastian
Abstract
Visual-Inertial (VI) fusion is fundamental to accurate and robust state estimation, where camera and IMU measurements are combined according to their respective uncertainties. Existing methods, however, fuse the two modalities with predefined uncertainties, regardless of how reliable each is in the local context, and thus often struggle under challenging environments involving illumination changes, dynamic objects, and textureless regions. In this paper, we present MAC-I$^2$, which achieves robust VI fusion through learned metric-aware covariance for both modalities, so that vision and IMU compete on their own merits rather than relying on predefined uncertainties. Here, metrics-aware means that each predicted covariance faithfully reflects the actual magnitude of the corresponding measurement noise. On the visual side, we propagate learned feature-matching uncertainties into pose covariances for the fusion. On the inertial side, motivated by the observation that integration error accumulates sharply at the early stage and grows slowly afterward, we design a learned IMU model with a learnable initial covariance, and propose a dedicated fine-tuning strategy on a held-out training subset to enable the metrics-aware covariance on unseen sequences. As a showcase, we build a VI initialization and calibration system, since accurate and robust initialization and calibration are the prerequisite for any reliable VI system. Experiments on EuRoC, and VBR show that MAC-I$^2$ substantially outperforms existing methods: it achieves a 99.9% initialization success rate on EuRoC, reducing gravity and velocity errors by about 60% and 42% over the strongest baseline, and maintains 80% success rate on challenging VBR sequences where baseline methods such as VINS-Mono drop below 10%.
Chinese Translation
视觉-惯性(VI)融合是实现精确且鲁棒状态估计的基础,其通过结合相机与IMU测量并依据各自的不确定性进行加权。然而,现有方法通常使用预定义的不确定性融合这两种模态,而不考虑各模态在局部情境中的可靠性,因此在涉及光照变化、动态物体和无纹理区域等挑战性环境下往往表现不佳。本文提出MAC-I$^2$,通过对两种模态学习度量感知协方差来实现鲁棒的视觉-惯性融合,使视觉与IMU凭借自身优势相互竞争,而非依赖预定义的不确定性。这里的度量感知是指每个预测的协方差能够忠实地反映相应测量噪声的实际大小。在视觉方面,我们将学习到的特征匹配不确定性传播到位姿协方差中以用于融合。在惯性方面,基于积分误差在初始阶段急剧累积、随后缓慢增长的观察,我们设计了一种具有可学习初始协方差的IMU学习模型,并针对保留训练子集提出一种专门的微调策略,使度量感知协方差能够推广到未见过的序列上。作为应用展示,我们构建了一个视觉-惯性初始化与标定系统,因为准确且鲁棒的初始化与标定是任何可靠VI系统的前提。在EuRoC和VBR数据集上的实验表明,MAC-I$^2$显著优于现有方法:在EuRoC上实现了99.9%的初始化成功率,相比最强基线将重力误差和速度误差分别降低了约60%和42%;在具有挑战性的VBR序列上仍保持80%的成功率,而VINS-Mono等基线方法则下降至10%以下。
cs.RO / 55 / 2609.07126

Beyond Task Success: Stage-Wise Reliability of World Model Planning under Sensing Degradation

超越任务成功:感知退化下世界模型规划的阶段性可靠性
Lee, Geonmyeong, Zhang, Byoung-Tak
Abstract
In world model planning, sensing inputs pass through an encoder and predictor before affecting planner decisions, so final task success alone cannot reveal where sensing disturbances attenuate or persist in the pipeline. We apply 10 visual and temporal sensing degradations to a world model planner and track their effects across representation, future prediction, planner preference, and physical outcome using paired evaluation on the same 50 tasks. The relative impact of degradations was not preserved across stages: large representation shifts could attenuate downstream, while smaller initial shifts could persist to the outcome, and internal-response ordering did not directly match physical-outcome ordering. Temporal degradations also showed distinct patterns: even with similar overall changes in observation history, responses differed substantially with the location of corrupted information and the planner's actual exposure. This non-uniform stage-wise response was also observed in secondary evaluations with another manipulation task and a different world model. Stage-wise diagnosis can therefore identify where sensing disturbances attenuate or persist and help prioritize subsequent model verification and sensing mitigation.
Chinese Translation
在世界模型规划中,感知输入需经过编码器和预测器后才会影响规划器的决策,因此仅凭最终任务成功无法揭示感知扰动在流水线中的何处衰减或持续存在。我们对一个世界模型规划器施加10种视觉与时间维度的感知退化,并通过对相同的50个任务进行配对评估,追踪其在表征、未来预测、规划器偏好和物理结果各阶段的影响。实验发现,各退化的相对影响并未在各阶段间保持一致:较大的表征偏移可能在下游衰减,而较小的初始偏移却可能延续至最终结果,且内部响应的排序与物理结果的排序并不直接对应。时间维度的退化也表现出独特的模式:即使观测历史的整体变化程度相似,其响应也会因损坏信息的位置以及规划器实际接收到的信息而有显著差异。这种非均匀的阶段性响应在包含另一操作任务和另一世界模型的补充评估中同样得到验证。因此,阶段性诊断可以识别感知扰动在何处衰减或持续存在,并有助于为后续的模型验证和感知缓解措施确定优先级。
cs.RO / 56 / 2609.07145

EquiGQNet: Fast Grasp Quality Evaluation via Shared Equivariant Point Cloud Encoding

EquiGQNet:基于共享等变点云编码的快速抓取质量评估
Seo, Sungwon, Won, Jaeseog, Shin, Jiyou, Seo, Youngjin, Kim, Hyunjun, Yoon, Seokmin, Luong, Tuan, Moon, Hyungpil
Abstract
Planning six-degree-of-freedom (6-DoF) grasps for unseen objects in cluttered tabletop scenes from a single-view depth image requires accurate and efficient evaluation of diverse grasp candidates. Existing early-fusion methods capture local object geometry relative to each grasp candidate but repeatedly encode the scene, whereas late-fusion methods reuse a shared scene representation but may lose this grasp-relative local geometry. We propose EquiGQNet, an efficient 6-DoF grasp quality evaluator that combines the strengths of both approaches. For grasp orientation, EquiGQNet replaces the early-fusion operation of rotating and re-encoding the point cloud for each grasp candidate with an SO(3)-equivariant encode-once-then-rotate scheme, yielding grasp-aligned geometric features from a shared scene encoding. For grasp translation, Mid-level Action Fusion (MAF) injects the grasp position into intermediate features before global aggregation, retaining local geometry relative to each candidate. We evaluate EquiGQNet in two grasp planning pipelines: Cross-Entropy Method (CEM)-based continuous grasp search and candidate ranking with a pretrained generative planner. In simulation, EquiGQNet achieves grasping performance comparable to the early-fusion baseline and substantially outperforms late fusion on objects with complex geometry and limited graspable regions, while reducing CEM planning time from 3.31s to 0.48s, a 6.9x speedup over early fusion. In real-world household-object decluttering, EquiGQNet achieves a 95.2% grasp success rate and 230 picks per hour, versus 153 and 170 for early- and late-fusion baselines. Code is available at https://equigqnet.github.io/.
Chinese Translation
从单视角深度图像出发,为杂乱桌面场景中的未知物体规划六自由度(6-DoF)抓取,需要对多样化抓取候选进行准确而高效的评估。现有的早期融合方法能够捕捉相对于每个抓取候选的局部物体几何信息,但需要反复对场景进行编码;而后期融合方法虽然可以复用共享的场景表示,却可能丢失这种与抓取相关的局部几何信息。我们提出EquiGQNet,一种高效的6-DoF抓取质量评估器,融合了两种方法的优势。在抓取姿态方面,EquiGQNet用基于SO(3)等变的“一次编码、多次旋转”方案取代了早期融合中为每个抓取候选旋转并重新编码点云的操作,从共享的场景编码中获得与抓取对齐的几何特征。在抓取平移方面,中层动作融合模块在全局聚合之前将抓取位置注入中间特征,从而保留相对于每个候选的局部几何信息。我们在两种抓取规划流程中评估EquiGQNet:基于交叉熵方法(CEM)的连续抓取搜索,以及与预训练生成式规划器结合的候选排序。在仿真中,EquiGQNet的抓取性能与早期融合基线相当,并在几何复杂、可抓取区域有限的物体上显著优于后期融合,同时将CEM规划时间从3.31秒降低到0.48秒,相比早期融合实现了6.9倍的加速。在真实世界的家居物品整理任务中,EquiGQNet实现了95.2%的抓取成功率和每小时230次抓取,而早期融合与后期融合基线分别为每小时153次和170次。代码发布于 https://equigqnet.github.io/。
cs.RO / 57 / 2609.07165

State-of-the-Art in Learning-by-Demonstration with Passive Observation for Industrial Assembly Automation

基于被动观察的示教学习在工业装配自动化中的研究现状
Koetter, David, Petrovic, Oliver, Brecher, Christian
Abstract
Learning-by-Demonstration (LbD) enables intuitive robot programming by capturing expert skills, which is crucial for agility in high-mix, low- volume manufacturing. This systematic literature review analyzes passive LbD for industrial assembly processes, focusing on the perception architecture and the generalization of the perceived demonstration. We specifically investigate one-shot approaches where only a single demonstration is required. The review evaluates how systems adapt to new assemblies using this limited data. We identify a shift towards object-centric perception, allowing learned primitives to be transferred to new product variants with minimal training.
Chinese Translation
示教学习(Learning-by-Demonstration, LbD)通过捕获专家技能实现直观的机器人编程,这对于高混合、小批量制造中的敏捷性至关重要。本系统性文献综述分析了面向工业装配过程的被动式示教学习,重点关注感知架构以及所感知示教的泛化能力。我们特别研究了仅需单次示教的单样本(one-shot)方法。本综述评估了系统如何利用这种有限的数据适应新装配任务。我们发现该领域正转向以物体为中心的感知方式,使学习到的技能基元能够以最少的训练迁移到新的产品变体上。
cs.RO / 58 / 2609.07211

Phase-and-First-Arrival VLM Feedback for Sparse-Reward Reinforcement Learning in Surgical Manipulation

面向外科手术操作稀疏奖励强化学习的相位与首次到达视觉语言模型反馈
Liuchen, Wanli, Wang, Fangyuan, Li, Bin, Duan, Anqing, Liu, Yunhui, Zhou, Peng, Navarro-Alarcon, David
Abstract
Sparse outcome feedback limits what robots can learn from unsuccessful attempts at complex manipulation. Failed multi-stage surgical attempts can contain grasps, lifts, or transfers worth reusing. In sparse-reward reinforcement learning, terminal rewards collapse such attempts to the same outcome, while scalar vision-language model (VLM) ratings reveal neither what progress merits credit nor when it occurred. We introduce phase-and-first-arrival feedback: one VLM query per recorded episode identifies the furthest visually verified task phase and when that phase is first reached, allowing the learner to reuse partial behavior and localize credit. We instantiate it in SurgPhaseBench, a phase-structured suite spanning rigid and deformable tasks, and evaluate it in simulation and hardware. Across five simulated tasks, our method reaches 75.2% mean success, compared with 52.1% for a reward based on Contrastive Language-Image Pre-training (CLIP) using the same visual input; the advantage persists when only the feedback representation changes. On hardware, the same record supports autonomous block picking and slip recovery. Together, these results show that trajectory-level visual supervision can preserve partial progress while providing the temporal credit needed for sparse-reward control.
Chinese Translation
稀疏的结果反馈限制了机器人从复杂操作的失败尝试中学习的能力。失败的多阶段外科手术尝试中可能包含值得复用的抓取、提升或转移动作。在稀疏奖励强化学习中,终止奖励将这些尝试压缩为相同的结果,而标量的视觉语言模型(VLM)评分既无法揭示哪些进展值得给予信用,也无法揭示进展发生的时间。我们提出了相位与首次到达反馈:对每个记录的回合进行一次VLM查询,识别视觉上可验证的最远任务相位以及该相位首次到达的时间,使学习者能够复用部分行为并定位信用。我们在SurgPhaseBench中实现了该方法,这是一个涵盖刚性和可变形任务的相位结构化任务套件,并在仿真和硬件上进行了评估。在五个仿真任务中,我们的方法达到75.2%的平均成功率,而使用相同视觉输入的基于对比语言-图像预训练(CLIP)的奖励方法为52.1%;当仅改变反馈表示时,这一优势依然存在。在硬件上,同一记录支持自主积木抓取和滑落恢复。总之,这些结果表明,轨迹级视觉监督既能保留部分进展,又能为稀疏奖励控制提供所需的时间信用分配。
cs.RO / 59 / 2609.07231

Generalizable 6D Pose Estimation of Textureless Objects with Planar-based Gaussian Splatting

基于平面高斯泼溅的无纹理物体可泛化6D位姿估计
Lu, Jie, Zhang, Hengtan, Gong, Li, Wang, Pengpeng, Yu, Xianjia, Deng, Jinxiang, Westerlund, Tomi, Gan, Zhongxue, Zheng, Lirong, Zou, Zhuo
Abstract
Estimating the 6D pose of textureless objects without prior CAD models remains a critical challenge due to the lack of appearance features. While recent generalizable approaches alleviate the dependence on object-specific models, their performance on low-texture objects is often limited by insufficient geometric constraints in the underlying representations. In this work, we propose PG-Pose, a geometry-aware framework combining Planar-based Gaussian Splatting (PGS) reconstruction and Geometry-driven pose optimization. In the offline representation extraction stage, three distinct representations of the object are extracted from multi-view reference RGB images with known poses. PG-Pose reconstructs a 3D Gaussian representation and renders high-fidelity depth maps to generate 3D point clouds through back projection. In the online pose inference stage, the initial pose of the input image is estimated by 2D-3D correspondence matching between the input image and the reconstructed 3D point clouds, followed by a PGS-Refiner for iterative pose optimization. Evaluations on the OnePose-LowTexture datasets, PG-Pose achieves an average accuracy of 94.2% ADD(S)@0.1d, with a 2.1% improvement average accuracy compared with the state-of-the-art (SOTA) GS-based approach. To further demonstrate the effectiveness of PG-Pose for industrial robots in grasping tasks, we deploy it on a dual-arm industrial robot and successfully realize the grasping task on an unseen object.
Chinese Translation
在缺乏外观特征的情况下,对无先验CAD模型的无纹理物体进行6D位姿估计仍然是一项关键挑战。尽管近期的可泛化方法缓解了对物体特定模型的依赖,但其在低纹理物体上的性能往往受限于底层表示中几何约束的不足。在本工作中,我们提出了PG-Pose,一种结合基于平面高斯泼溅(Planar-based Gaussian Splatting, PGS)重建与几何驱动位姿优化的几何感知框架。在离线表示提取阶段,从具有已知位姿的多视角参考RGB图像中提取物体的三种不同表示。PG-Pose重建3D高斯表示并渲染高保真深度图,通过反投影生成3D点云。在在线位姿推理阶段,通过输入图像与重建的3D点云之间的2D-3D对应匹配来估计输入图像的初始位姿,随后由PGS-Refiner进行迭代位姿优化。在OnePose-LowTexture数据集上的评估表明,PG-Pose达到了94.2%的平均ADD(S)@0.1d精度,相比最先进的(SOTA)基于高斯泼溅(GS)的方法平均精度提升了2.1%。为进一步验证PG-Pose在工业机器人抓取任务中的有效性,我们将其部署在双臂工业机器人上,并成功实现了对未见物体的抓取任务。
cs.RO / 60 / 2609.07269

Singularity-Free Guiding Vector Fields on SO(3) with Designer-Specified Progression Behavior

SO(3)上具有设计者指定行进行为的无奇点引导向量场
Bautista, Jesus, de Marina, Hector Garcia
Abstract
This paper develops a singularity-free guiding vector field (SF-GVF) for path following on the special orthogonal group SO(3). First, we lift the Euclidean SF-GVF construction to SO(3), integrating the augmented-state approach with the intrinsic Lie-group geometry and obtaining a closed-form geometric guidance law whose integral curves converge to a designer-specified attitude path. The field is defined on a dense open subset of SO(3), excluding only the measure-zero antipodal set - a manifestation of the topological obstruction to continuous global stabilization on SO(3). The construction requires no per-step optimization and produces a control input intrinsically in so(3) as body angular rates. Second, we formalize the progression behavior along the path as a designer-supplied function \nu(\xi), promoting the parametric speed from an implicitly resolved degree of freedom to a first-class design specification. In contrast to the Euclidean condition v = 0, which excludes vehicles with minimum-speed constraints, the corresponding condition \omega = 0 on SO(3) is physically admissible for most platforms with active attitude control, making the progression behavior a design freedom structurally available on SO(3) but absent in the Euclidean setting. The framework's structural results are established under a bi-invariant Riemannian metric and hold uniformly across choices of path, progression, and Lyapunov gain. The framework is illustrated in simulation on self-intersecting paths under both constant and point-convergence progression behaviors.
Chinese Translation
本文针对特殊正交群SO(3)上的路径跟踪问题,提出了一种无奇点引导向量场(SF-GVF)。首先,我们将欧氏空间中的SF-GVF构造提升到SO(3)上,将增广状态方法与李群内蕴几何相结合,得到一种闭式几何制导律,其积分曲线收敛于设计者指定的姿态路径。该向量场定义在SO(3)的一个稠密开子集上,仅排除了测度为零的对径点集——这是SO(3)上连续全局镇定的拓扑障碍的体现。该构造无需逐步优化,并可产生本质上属于so(3)的机体角速度形式的控制输入。其次,我们将沿路径的行进行为形式化为设计者提供的函数ν(ξ),将参数速度从一个隐式求解的自由度提升为一等设计规范。与欧氏空间中排除具有最小速度约束载具的v = 0条件不同,SO(3)上相应的条件ω = 0对于大多数具有主动姿态控制的平台在物理上是可容许的,这使得行进行为成为SO(3)上结构性可用而在欧氏设置中不存在的设计自由度。该框架的结构性结果在双不变黎曼度量下建立,并对路径、行进行为和Lyapunov增益的选择均一致成立。通过在自相交路径上针对恒定行进行为和点收敛行进行为的仿真,对该框架进行了验证说明。
cs.RO / 61 / 2609.07274

LightSplat: Real-Time High-Fidelity 3D Gaussian SLAM with Loop Closure

LightSplat:具有回环检测的实时高保真三维高斯SLAM
Bao, Junze, Gao, Ye, Huang, Yiming, Yu, Xiaolong, Dong, Chen, Gao, Qing, Wang, Wei, Lü, Jinhu
Abstract
SLAM systems based on 3D Gaussian Splatting (3DGS) have recently demonstrated promising reconstruction accuracy for dense 3D scene representations. However, current 3DGS systems struggle to meet the strict demands of real-world deployments due to severe limitations in operational performance and map adaptability. To this end, we propose LightSplat, a hybrid-representation RGB-D SLAM framework. It synergizes local sparse features for robust and fast tracking with a dual-thread backend that progressively constructs dense Gaussian submaps. Crucially, we enable online loop closure through feature-accelerated 3DGS registration, refining overall map consistency through pose graph optimization. Ultimately, LightSplat achieves the online reconstruction of high-fidelity Gaussian map. Extensive experiments on multiple datasets and real-world robotic platform demonstrate that our method achieves near state-of-the-art reconstruction quality and the capability to accommodate practical camera motions, maintaining an average framerate of 8 FPS. Overall, LightSplat provides an efficient and robust foundation for deploying high-fidelity 3DGS in real-world environments.
Chinese Translation
基于三维高斯泼溅(3D Gaussian Splatting, 3DGS)的SLAM系统近来在稠密三维场景重建方面展现出令人瞩目的精度。然而,当前3DGS系统在运行性能和地图适应性方面存在严重局限,难以满足实际部署的严苛要求。为此,我们提出了LightSplat,一种混合表示的RGB-D SLAM框架。该框架将用于鲁棒且快速跟踪的局部稀疏特征与渐进式构建稠密高斯子地图的双线程后端相结合。关键在于,我们通过特征加速的3DGS配准实现了在线回环检测,并借助位姿图优化提升整体地图的一致性。最终,LightSplat实现了高保真高斯地图的在线重建。在多个数据集及真实机器人平台上的大量实验表明,我们的方法达到了接近最先进水平的重建质量,并能适应实际相机运动,平均帧率保持在8 FPS。总体而言,LightSplat为在真实环境中部署高保真3DGS提供了一个高效且鲁棒的基础。
cs.RO / 62 / 2609.07288

How Long Until Your Robot Ignores You? A Safety Benchmark for LLM Orchestrators in Human-Humanoid Collaboration

你的机器人多久会开始无视你?面向人形机器人协作中LLM编排器的安全基准测试
Bajrami, Aulon, Elshamouty, Mohamed, Kraus, Werner
Abstract
Large Language Models (LLMs) are increasingly employed to orchestrate robot behavior through natural-language interfaces, yet no benchmark exists to evaluate their reliability as safety-aware decision makers in human-humanoid collaboration. Unlike deterministic safety systems that enforce binary allow/deny decisions, LLM-based orchestrators exhibit a compliance spectrum ranging from overcompliance (refusing safe actions) to full safety violations. This paper introduces the first safety benchmarking environment for LLM orchestrators in human-humanoid collaboration, built on a Model Context Protocol (MCP)-based architecture with safety invariants grounded in ISO 10218-2:2025 protective measures. The benchmark defines five testable safety invariants, a four-level compliance taxonomy (correct compliance, overcompliance, undercompliance, full violation), and a three-layer evaluation pipeline (text prompting, simulated sensor-actuator loops, and physical validation on a Unitree G1 EDU humanoid). We report Layer-1 results: three cloud backends (Claude Haiku 4.5, GPT-4o-mini, Gemini 2.5 Flash) and a local open-weights baseline (qwen3:8b) across 40 100-turn sessions under full-context and sliding-window budget conditions, while the simulation and physical layers remain ongoing. We find that (1) model family determines the safety floor, as Claude and Gemini remain at or near zero violations while GPT-4o-mini commits up to 13 per session, (2) context management dissociates two failure axes, reducing mean behavioral issues by 42-57% for every cloud backend while nearly doubling GPT-4o-mini's violations (3.8 to 7.2 per session), and (3) proportional compliance, clamping movement speed to the rule-specified maximum rather than refusing, emerges consistently only in Gemini; the preliminary simulation layer reproduces the model ranking and the GPT-4o-mini failure-mode inversion.
Chinese Translation
大语言模型(LLM)越来越多地被用于通过自然语言接口编排机器人行为,但目前尚无基准用于评估其作为人形机器人协作中具备安全感知能力的决策者的可靠性。与执行二元允许/拒绝决策的确定性安全系统不同,基于LLM的编排器表现出一个合规性谱系,范围从过度合规(拒绝安全操作)到完全违反安全规定。本文提出了首个面向人形机器人协作中LLM编排器的安全基准测试环境,该环境基于模型上下文协议(Model Context Protocol, MCP)架构构建,其安全不变量(safety invariants)以ISO 10218-2:2025防护措施为基础。该基准定义了五个可测试的安全不变量、一个四级合规性分类体系(正确合规、过度合规、合规不足、完全违规),以及一个三层评估流水线(文本提示、仿真传感-执行回路,以及在Unitree G1 EDU人形机器人上的物理验证)。我们报告了第一层的结果:三种云端后端(Claude Haiku 4.5、GPT-4o-mini、Gemini 2.5 Flash)和一个本地开源权重基线模型(qwen3:8b),在完整上下文与滑动窗口预算条件下进行了40个100轮的会话测试,仿真层与物理层仍在进行中。我们发现:(1)模型家族决定了安全下限——Claude和Gemini始终保持零违规或接近零违规,而GPT-4o-mini每会话最多达13次违规;(2)上下文管理使两个失败维度解耦——所有云端后端的平均行为问题减少了42-57%,同时GPT-4o-mini的违规几乎翻倍(从每会话3.8次增至7.2次);(3)比例合规(即根据规则将移动速度限制在规定上限而非直接拒绝)仅在Gemini中持续出现;初步的仿真层复现了模型排名以及GPT-4o-mini失败模式的反转现象。
cs.RO / 63 / 2609.07328

PV-WM: A Heterogeneous Micro-Macro World Model for Articulated Pedestrian-Vehicle Co-Rollout

PV-WM:用于行人-车辆关节化协同推演的异构微观-宏观世界模型
Chi, Haozhuang, Liang, Jingsong, Song, Ziying, Yang, Lei, Li, Shihao, Zhang, Haoruo, Lv, Chen
Abstract
Local pedestrian-vehicle forecasting spans heterogeneous physical scales: pedestrians combine root locomotion with articulated motion, whereas vehicles are rigid bodies described by kinematic state and oriented extent. Existing road-agent forecasters typically omit pedestrian articulation, while pose forecasters leave vehicle futures outside the learned rollout. We introduce PV-WM, a history-only world model over structured post-perception tracks. It recurrently advances pedestrian root motion, 15-joint articulation, and learned vehicle states within a synchronized heterogeneous state. The generated pedestrian and vehicle chunks supply the next recurrent boundary; vehicle boxes are reconstructed from predicted center and heading with observed extent, and P-V geometry is recomputed after every transition. Relative to a matched one-shot complete-state predictor, recurrent execution reduces Root ADE by 12.7% and MPJPE by 14.8%. Feedback interventions show that later predictions depend on the content, temporal order, and pedestrian identity of generated articulation. Across 824 aligned Waymo contexts, with 797 providing valid future vehicle support, PV-WM reduces Root ADE by 5.2%, MPJPE by 7.6%, P-V distance error by 11.9%, and oriented-box closest-approach error by 5.8% relative to a validation-selected Modular Specialist. The single-network model uses 57.1% fewer parameters, 96.5% lower average FLOPs per local scene, and 25.5% lower measured p95 latency. PV-WM unifies this heterogeneous future state while preserving type-specific pedestrian and vehicle dynamics.
Chinese Translation
局部行人-车辆预测涉及异构的物理尺度:行人将根节点运动与关节化运动相结合,而车辆则是由运动学状态和有向外形描述的刚体。现有的道路智能体预测方法通常忽略行人的关节化运动,而姿态预测方法则将车辆的未来状态排除在所学推演之外。我们提出了PV-WM,一个仅基于历史信息、作用于结构化后感知轨迹的世界模型。它在同步的异构状态中递归地推进行人根节点运动、15关节关节化以及学习得到的车辆状态。生成的行人和车辆片段为下一递归步骤提供边界;车辆框由预测的中心点、朝向以及观测到的外形重建,且每次状态转移后都重新计算行人-车辆(P-V)几何关系。与匹配的单次完整状态预测器相比,递归执行将根节点ADE降低12.7%,MPJPE降低14.8%。反馈干预实验表明,后续预测依赖于所生成关节化运动的内容、时间顺序以及行人身份。在824个对齐的Waymo场景中(其中797个提供了有效的未来车辆支持),与经验证的模块化专家模型(Modular Specialist)相比,PV-WM将根节点ADE降低5.2%,MPJPE降低7.6%,行人-车辆距离误差降低11.9%,有向框最近接近误差降低5.8%。该单一网络模型的参数量减少57.1%,每个局部场景的平均FLOPs降低96.5%,实测p95延迟降低25.5%。PV-WM在统一这一异构未来状态的同时,保留了行人和车辆各自特有的动力学特性。
cs.RO / 64 / 2609.07350

D3ARC: Time-Critical Distributed Disaster Detection for Asynchronous Cooperative Multi-Robot Systems

D3ARC:面向异步协作多机器人系统的时敏分布式灾难检测
Koursioumpas, Nikolaos, Magoula, Lina, Alonistioti, Nancy, Khalili, Ramin
Abstract
Climate change is increasing the severity and unpredictability of natural disasters. In time-critical crises such as wildfires, traditional monitoring practices remain limited by coverage, cost, and personnel risk, paving the way for autonomous and adaptive monitoring solutions. Within this context, this paper introduces D3ARC, an asynchronous distributed hierarchical framework for time-aware and reliable wildfire detection. D3ARC integrates multiple robotic agents that cooperate under uncertainty through distributed perception, shared situational awareness and coordinated actions. A remote controller asynchronously decides upon each robot's motion, while each robotic agent senses the environment and decides where and how to execute the wildfire detection. All robotic operations require time, and as time progresses, wildfires continue to spread, reducing the opportunity for early intervention. As such, all agents share a common objective: to detect a wildfire with a certain performance threshold as fast as possible and within a time limit. D3ARC integrates mechanisms for safe navigation, coverage efficiency, cooperation and reliability. It introduces a forward-looking capability that allows agents to anticipate the future by evaluating candidate strategies before execution. The framework is evaluated through realistic robotics simulations, ablation studies, and baseline comparisons, achieving an overall mission success up to 94% with 89.4% detection confidence.
Chinese Translation
气候变化正在加剧自然灾害的严重程度和不可预测性。在野火等时间紧迫的危机中,传统监测手段在覆盖范围、成本和人员风险方面仍存在局限,这为自主化、自适应的监测解决方案创造了需求。在此背景下,本文提出了D3ARC,一个面向时间感知且可靠野火检测的异步分布式分层框架。D3ARC集成了多个机器人智能体,通过分布式感知、共享态势感知和协同行动,在不确定性条件下进行协作。远程控制器异步地决策每个机器人的运动,而每个机器人智能体则感知环境,并决定在何处以及如何执行野火检测。所有机器人操作都需要时间,而随着时间的推移,野火会持续蔓延,减少早期干预的机会。因此,所有智能体共享一个共同目标:在限定时间内,以一定的性能阈值尽快检测到野火。D3ARC集成了安全导航、覆盖效率、协作与可靠性等机制,并引入了一种前瞻能力,使智能体能够在执行前通过评估候选策略来预判未来。该框架通过真实的机器人仿真、消融实验和基线对比进行了评估,整体任务成功率最高达94%,检测置信度达89.4%。
cs.RO / 65 / 2609.07398

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining

OpenWAM:面向系统化世界-动作模型预训练的开放模块化探索
Wang, Yuran, Huang, Siqiao, Li, Mingleyang, Zhang, Chenhao, Liang, Jiaqi, Jin, Weiyang, Chen, Yue, Chi, Xuemin, Zhou, Donghao, Yu, Qize, Wang, Yu-Kai, Rui, Yuhan, Yao, Shenzhe, Yuan, Zhen, Shen, Zhenhao, Zhu, Kefei, Zhu, Zijie, Gao, Ning, Chi, Xiaowei, He, Guanqi, Zhang, Shanghang, Dong, Hao, Shao, Lin, Zhao, Hang
Abstract
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-{\alpha}, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-{\alpha} delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.
Chinese Translation
世界-动作模型(World-Action Models, WAM)从视频生成先验中继承世界知识,并通过具身体验将其转化为可执行的控制信号。然而,现有系统是单体式的:生成骨干网络、视觉表征、架构、信息流、推理流程和训练数据紧密耦合,掩盖了哪些设计选择至关重要以及为什么。我们提出OpenWAM,一个将世界-动作预训练转化为可控实验项目的开放研究体系。OpenWAM-Infra将WAM设计空间分解为可组合的模块,并统一了训练、推理、部署和评估。在此基础上,OpenWAM-Study通过受控实验考察三个问题:应继承什么、世界学习与动作学习如何相互作用、以及它们的协同如何随规模扩展;并提炼出三条原则:上游知识通过足够强大的生成骨干网络和一个紧凑且信息丰富的潜空间进行迁移;世界-动作协同需要专门的动作能力、显式的世界到动作的信息流以及同步的联合去噪;具身预训练主要提升域外泛化能力,在自我中心视角(egocentric)数据与机器人数据上进行单阶段联合训练能够兼顾世界覆盖范围与动作落地能力。综合这些原则,我们构建了OpenWAM-α,一个基于约6,400小时自我中心视角人类和机器人数据预训练的开放WAM,并在仿真和真实世界基准上进行了评估。在八个仿真基准以及涵盖单臂、双臂操作到灵巧手等多种本体的真实机器人实验中,OpenWAM-α始终表现出色,从仿真到物理世界均保持顶级性能水平。我们发布了完整的技术栈,包括基础设施、评估协议、预训练模型和数据配方,以促进未来研究。
cs.RO / 66 / 2609.07430

CALM: Configuration-Aware Human Intervention Boundaries During Robot Approach

CALM:机器人接近过程中具有构型感知的人类干预边界
Gao, Xinting, Zhu, Sipu, Zhuang, Weimin
Abstract
How robot body configuration shapes human intervention during approach remains underexplored. We conducted a within-participants study with 41 participants, measuring final stopping distance, subjective comfort, and exploratory eye-tracking responses across four humanoid arm configurations and two spatial scales. Full forward arm extension increased stopping distance by approximately 31-36 cm relative to arms-down. Spatial scale primarily affected comfort and pupil responses without a detectable stopping-distance shift. We introduce the Configuration-Aware Limit Model (CALM), which translates stopping-distance distributions into configuration-dependent population-coverage boundaries. Estimated boundaries at 80% coverage ranged from 0.88 to 1.47 m. In an illustrative one-dimensional planning analysis, reconfiguration enabled a 1.10 m approach goal that was unreachable with arms remaining fully extended under the same nominal pointwise 20% intervention-probability constraint. These findings support treating body configuration as a planning variable while distinguishing physical safety, behavioral intervention, and subjective cost.
Chinese Translation
机器人的身体构型(configuration)如何影响人类在其接近过程中的干预行为,目前仍缺乏充分研究。我们开展了一项被试内实验,共41名参与者,测量了最终停止距离、主观舒适度以及探索性眼动追踪反应,涵盖四种人形手臂构型和两种空间尺度。与手臂下垂相比,手臂完全前伸使停止距离增加了约31-36厘米。空间尺度主要影响舒适度和瞳孔反应,而未观察到停止距离的显著变化。我们提出了构型感知极限模型(Configuration-Aware Limit Model, CALM),该模型将停止距离分布转化为依赖于构型的群体覆盖边界。估计的80%覆盖率边界范围为0.88至1.47米。在一个示例性的一维规划分析中,通过构型重组,机器人能够在相同的标称逐点20%干预概率约束下,实现手臂保持完全伸展时无法达到的1.10米接近目标。这些发现支持将身体构型作为规划变量加以考虑,同时区分物理安全、行为干预和主观代价。
cs.RO / 67 / 2609.07440

Open-Set Ego-Noise Separation for Legged-Robot Audition via Annotation-Free Adaptation and Pretrained-Model Transfer

基于无标注自适应与预训练模型迁移的腿式机器人听觉开集自噪声分离
Shoda, Koki, Kasahara, Jun Younes Louhi, Koyanagi, Aoba, An, Qi, Yamashita, Atsushi
Abstract
This paper proposes an open-set ego-noise separation framework for legged-robot audition via annotation-free adaptation and pretrained-model transfer. The framework removes robot-specific ego-noise while preserving environmental sounds whose classes are not specified in advance. Acoustic sensing provides cues about a robot's surroundings beyond the visual field, but walking-induced ego-noise from footstep impacts, joint-backlash rattling, and motor noise severely contaminates the recordings. The framework first uses RecurGraph to select ego-noise-dominant clips from the unlabeled recordings by aggregating clip embeddings into an embedding centroid and propagating scores over an audio-embedding graph. The selected clips are mixed with diverse environmental sounds from a large-scale sound-event dataset to provide paired mixture--target supervision for open-set separation. Transfer-DiT then adapts a general-purpose zero-shot neural separator to achieve high-fidelity open-set ego-noise separation for the target robot. Experiments with bipedal and quadrupedal robots show reliable clip selection and improvements in separation quality and downstream task performance over baseline separators. These results demonstrate the feasibility of annotation-free adaptation without separately recorded ego-noise-only data or manual clip-level annotations.
Chinese Translation
本文提出了一种面向腿式机器人听觉的开集自噪声分离框架,该框架基于无标注自适应与预训练模型迁移。该框架能够在去除机器人特定自噪声的同时,保留类别未预先指定的环境声音。声学感知能够提供机器人视野之外的环境线索,但行走引起的自噪声——包括脚步冲击、关节间隙碰撞和电机噪声——会严重污染录音。该框架首先利用 RecurGraph 通过将片段嵌入聚合为嵌入质心并在音频嵌入图上传播分数,从未标注录音中筛选出自噪声主导的片段。随后,将所选片段与来自大规模声音事件数据集的多样化环境声音混合,为开集分离提供配对的混合-目标监督。接着,Transfer-DiT 对通用的零样本神经分离器进行自适应调整,以实现针对目标机器人的高保真开集自噪声分离。在双足和四足机器人上的实验表明,该框架在片段筛选方面表现可靠,并且在分离质量和下游任务性能上均优于基线分离器。这些结果证明了无需单独录制的纯自噪声数据或人工片段级标注即可实现无标注自适应的可行性。
cs.RO / 68 / 2609.07445

SMaRT-Tug: Structured Multi-Agent Reinforcement Learning for Physics-Based Tugboat-Barge Collaborative Manipulation

SMaRT-Tug:基于物理的拖船-驳船协同操纵的结构化多智能体强化学习
Lu, Junkai, Zhao, Jiadong, Zhang, Jiacheng, Zhao, Wenqi, Chia, Hao Gen, Png, Qun Shen, Ee, Germaine, He, Chengyang, Zhang, Yifeng, Tan, Nathanael, Sartoretti, Guillaume
Abstract
Autonomous tugboating is central for automating maritime operations such as port logistics and vessel maneuvering, where multiple tugboats must cooperatively transport/manipulate a larger vessel. Collaborative pushing in this setting is challenging due to coupled hydrodynamics, low resistance, strong environmental disturbances, underactuated barge dynamics, and contact-rich interactions. Conventional control methods often rely on simplified models and fixed configurations, which limit their adaptability, while learning-based approaches are constrained by the lack of scalable and physically realistic training environments. We address these challenges by introducing a physics-based, GPU-accelerated simulation and learning framework for collaborative tugboat manipulation. Our simulator incorporates a customized buoyancy model, wave modeling, and hydrodynamic resistance, and supports large-scale multi-agent training under marine dynamics. In this simulator, we train a decentralized MAPPO (Multi-Agent PPO) policy augmented with a structured control prior (SCP) to improve training stability and maintain feasible pushing configurations. We evaluate our learned policy on straight-line transit, turning, and deceleration tasks, where we show that our decentralized framework yields more reliable and accurate maneuvering performance compared to a PID-based controller and a centralized PPO baseline. We further demonstrate zero-shot generalization to more challenging sea states and advanced maneuvers, as well as zero-shot scalability to larger teams of three and four tugboats despite training with only two agents.
Chinese Translation
自主拖船作业是实现港口物流和船舶操纵等海事作业自动化的核心,其中多艘拖船需要协同运输/操纵一艘较大的船舶。在此场景下,协作推顶由于水动力耦合、低阻力、强环境扰动、欠驱动驳船动力学以及丰富的接触交互而极具挑战性。传统控制方法通常依赖简化模型和固定配置,限制了其适应性;而基于学习的方法则受限于缺乏可扩展且物理真实的训练环境。为应对这些挑战,我们提出了一个基于物理、GPU加速的拖船协同操纵仿真与学习框架。我们的仿真器集成了定制的浮力模型、波浪建模和水动力阻力,并支持海船动力学下的大规模多智能体训练。在该仿真器中,我们训练了一个去中心化的MAPPO(Multi-Agent PPO)策略,并通过结构化控制先验(SCP)加以增强,以提高训练稳定性并保持可行的推顶构型。我们在直线航行、转向和减速任务上评估了所学习的策略,结果表明,与基于PID的控制器和集中式PPO基线相比,我们的去中心化框架产生了更可靠、更精确的操纵性能。我们进一步展示了策略对更具挑战性海况和高级机动动作的零样本泛化能力,以及在仅用两个智能体训练的情况下,对三艘和四艘拖船更大团队的零样本可扩展性。
cs.RO / 69 / 2609.07470

Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy

测量机器人策略中的语言迁移:向Cosmos3视觉-语言-动作策略中添加希腊语
Kirouane, Ayoub, Giaples, Georgios, Petrocheilos, Christos
Abstract
Robot foundation models are trained and evaluated predominantly in English, and robot demonstration corpora do not exist for most languages. We study the addition of Greek to an open vision-language-action stack using only machine-rephrased instructions and no architecture changes. The main challenge is measurement rather than translation. Several plausible instruments produce false conclusions: a color-histogram metric rewards noise, a single-goal benchmark scores 84.6% under correct Greek and 82.6% under deliberately wrong instructions, training loss fails to predict Greek success, and single-run comparisons are dominated by seed variation. On a discriminative ninety-task suite with three seeds per arm, a multilingual text tower without Greek demonstrations remains at its wrong-instruction floor, while Greek-only training exceeds its control by at most 2.7 points. Bilingual training yields a consistent 6.7-7.1 point margin over its control and reaches about two fifths of English performance. The policy also overfits the translator's phrasing; training on seven phrasings per task approximately halves this penalty. Warm-starting from a language-adapted world model and unfreezing the text tower both degrade performance. The results support two practical requirements for low-resource robot-policy localization: build a guaranteed null before trusting a metric, and replicate low-resource-language results across seeds.
Chinese Translation
机器人基础模型的训练与评估主要以英语进行,且大多数语言不存在机器人演示语料库。我们研究了如何仅使用机器改写的指令、不改变模型架构,将希腊语加入一个开放的视觉-语言-动作(Vision-Language-Action)技术栈。其主要挑战在于测量而非翻译。若干看似合理的测量手段会得出错误结论:颜色直方图指标会奖励噪声;一个单目标基准在正确的希腊语指令下得分为84.6%,而在故意错误的指令下仍有82.6%;训练损失无法预测希腊语任务的成功率;单次运行的对比则被随机种子差异所主导。在一个包含九十项任务的判别性评测套件上(每组设置三个随机种子),不含希腊语演示的多语言文本编码器(text tower)仍停留在错误指令的水平线,而仅使用希腊语的训练相比其对照最多仅高出2.7个百分点。双语训练则相对其对照取得稳定的6.7-7.1个百分点的优势,并达到英语性能的大约五分之二。该策略还会对翻译器的措辞过拟合;对每个任务使用七种措辞进行训练可将这一损失大致减半。从语言适配的世界模型进行热启动以及解冻文本编码器均会降低性能。这些结果为低资源机器人策略的本地化提供了两点实践要求:在信任某个指标之前先构建一个有保证的零基线(null),并在多个随机种子上重复低资源语言的实验结果。
cs.RO / 70 / 2609.07495

Wearable Multimodal Human-Machine Interface for Integrated Hand Intentions Decoding in Dynamic Teleoperation

用于动态遥操作中手部意图综合解码的可穿戴多模态人机接口
Li, Jiaxuan, Wu, Yinshi, Zhang, Xiao, Wang, Hongyu, Le, Renzhen, Ying, Zhenzhi, Shu, Liming
Abstract
Under ubiquitous teleoperation environments with optically challenging conditions, an interface for tele-operated grasping that combines wearability with precise decoding of hand intentions (hand pose, gestures, and grasping force) is essential. Yet, existing interfaces often fall short in meeting these demands, compromising either the diversity of multiple intentions decoding or wearability. To address this, we developed a novel Multiple Intentions Decoding Human-Machine Interface (MI-DHMI) that integrates high-throughput surface electromyography (sEMG) sensors with hand-mounted and forearm-mounted inertial measurement units (IMUs). The developed interface is supported by a unified framework for simultaneous multiple intentions decoding. By employing multimodal deep learning and hardware design with a low noise floor, the decoding framework selectively focuses on the sEMG components that are genuinely associated with finger movements. This effectively reduces decoding errors caused by sEMG variability during unconstrained upper-limb motions, thereby significantly enhancing robustness. Even under unconstrained wrist and forearm motion, the interface achieves a gesture recognition accuracy exceeding 97%, grasping force estimation with $R^2 = 0.95$, and hand pose decoding consistent with the actual hand pose, outperforming baseline devices and algorithms. Ablation studies further validate the effectiveness of the proposed decoding framework. Finally, two online experiments were conducted to validate the device, demonstrating its superior performance in high-stability tasks, including a pouring task and object grasping. The developed interface provides a new solution of a fully wearable, multiple intentions decoding system, offering effective support for ubiquitous teleoperation and contributing to the advancement of human-machine interaction research.
Chinese Translation
在光学条件苛刻的泛在遥操作环境中,一种兼具可穿戴性与手部意图(手部姿态、手势和抓握力)精确解码能力的遥操作抓取接口至关重要。然而,现有接口往往难以满足这些需求,在多意图解码的多样性或可穿戴性方面有所妥协。为解决这一问题,我们开发了一种新型的多意图解码人机接口(Multiple Intentions Decoding Human-Machine Interface, MI-DHMI),该接口集成了高通量表面肌电(sEMG)传感器以及安装在手部和前臂上的惯性测量单元(IMU)。该接口由一个用于同步多意图解码的统一框架支持。通过采用多模态深度学习以及低噪声底 hardware 设计,该解码框架能够选择性地聚焦于真正与手指运动相关的 sEMG 成分,有效降低了非约束性上肢运动中 sEMG 变异性导致的解码误差,从而显著增强了鲁棒性。即使在前臂和手腕非约束运动的情况下,该接口仍能实现超过97%的手势识别准确率、$R^2 = 0.95$ 的抓握力估计精度,以及与实际手部姿态一致的手部姿态解码,优于基线设备和算法。消融实验进一步验证了所提解码框架的有效性。最后,通过两项在线实验对该设备进行了验证,展示了其在高稳定性任务(包括倒水任务和物体抓取任务)中的卓越性能。所开发的接口为全可穿戴多意图解码系统提供了一种新的解决方案,为泛在遥操作提供了有效支持,并推动了人机交互研究的发展。
cs.RO / 71 / 2609.07497

Functional-SLAM: Interaction-Aware Mapping with Online Functional Scene Graphs

Functional-SLAM:基于在线功能场景图的交互感知建图
Hu, Xinggang, Zhang, Chenyangguang, Zhu, Zihan, Zhang, Ruida, Zhang, Xiangkui, Ji, Xiangyang
Abstract
Existing SLAM systems lack modeling of the functional relations required for fine-grained robotic interaction. Functional 3D scene graphs can represent relations between objects and interaction elements, but existing methods rely on offline reconstruction, making them inadequate for real-time interaction in real-world exploration. To address this limitation, we propose Functional-SLAM, the first framework that continuously and recursively maintains a functional scene graph as an online SLAM state. The framework combines anchor-keyframe geometry with functional-context constraints for persistent node maintenance, accumulates multi-frame evidence through temporal relations to commit stable functional edges, and supplements visual loop-closure candidates with functional topology in scenes with repetitive appearance or degraded texture. Experiments show that Functional-SLAM efficiently constructs stable functional maps online, substantially improving runtime over offline methods while maintaining highly competitive accuracy. Compared with peer SLAM systems, it further improves pose estimation accuracy through functional-topology-assisted loop closure. The code is publicly available at https://github.com/Hbelief1998/Functional-SLAM-CoRL_2026.
Chinese Translation
现有SLAM系统缺乏对细粒度机器人交互所需功能关系的建模。功能三维场景图(Functional 3D Scene Graph)能够表示物体与交互元素之间的关系,但现有方法依赖离线重建,难以满足真实世界探索中的实时交互需求。针对这一局限,我们提出了Functional-SLAM,这是首个将功能场景图作为在线SLAM状态进行持续递归维护的框架。该框架结合锚点关键帧几何与功能上下文约束以维持持久节点,通过时间关系累积多帧证据以提交稳定的功能边,并在外观重复或纹理退化的场景中利用功能拓扑补充视觉回环候选。实验表明,Functional-SLAM能够在线高效构建稳定的功能地图,在保持高度竞争力的精度的同时,运行时间相较于离线方法大幅缩短。与同类SLAM系统相比,它还通过功能拓扑辅助的回环检测进一步提升了位姿估计精度。代码已公开发布于 https://github.com/Hbelief1998/Functional-SLAM-CoRL_2026。
cs.RO / 72 / 2609.07498

CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements

CosmoH2G:面向复杂空间运动物体操作的手部到夹爪迁移数据集与基线方法
Zhao, Hongxiang, Xu, Mutian, Jin, Zeyu, Hao, Yiming, Cui, Shuguang, Han, Xiaoguang
Abstract
Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.
Chinese Translation
将人类手部演示迁移到机器人夹爪近来已成为一种低成本的机器人学习解决方案。然而,现有方法大多局限于简单的平面任务,无法处理对机器人操作至关重要的复杂空间运动(例如涉及旋转或翻转的复杂轨迹)。基于这一空白,我们采用了一种由细粒度手部姿态运动引导的隐式数据驱动方法。为此,我们提出了一种可扩展的采集流程来收集手部-夹爪成对演示,该流程遵循以运动复杂性为优先的严格协议,并利用手持式夹爪实现无缝的动作模仿。由此得到一个大规模成对数据集,包含1,254个独特物体的6,189条演示序列,其空间复杂性显著高于现有基准。然而,学习如此复杂的映射仍然具有挑战性。我们观察到,直接端到端地生成完整的夹爪位姿序列是不够的,因为微小的轨迹偏差会在复杂动力学下迅速累积。为解决这一问题,我们提出了一个两阶段框架:第一阶段预测稀疏的夹爪关键帧(起始和终止帧)以简化映射目标,第二阶段在这些关键帧的条件下生成完整的连续动作序列。此外,为缓解累积漂移,我们让夹爪的姿态持续学习,同时基于抓取启发式规则和运动学一致性对其平移进行后优化。在仿真和真实机器人实验中,我们的框架实现了复杂空间操作的手部到夹爪的稳定且精确的迁移,显著优于传统基线方法。项目页面:https://cosmoh2g.github.io。
cs.RO / 73 / 2609.07511

Generation of Vectorized Maps Beyond Vehicle View

超越车辆视野的矢量化地图生成
Gomez, Clara, Jaenal, Alberto, Artuñedo, Antonio, Godoy, Jorge, Villagra, Jorge
Abstract
Autonomous driving relies on High Definition (HD) maps for safe navigation. Traditional HD maps construction is costly in hardware, data and human resources, which together with its update limitations hinders scalability. Recent works have proposed online alternatives for HD vectorized mapping from onboard sensors. However, sensor field of view is limited, and the range of the reconstructed maps ahead of the vehicle is insufficient for safe planning. This paper aims to address this limitation by proposing the novel beyond-view vectorized map generation problem: given vectorized maps of the area sensed by the vehicle (in-view), to generate plausible map continuations. To experimentally assess its feasibility, we propose BeyondFormer, which, to the best of out knowledge, is the first work designed towards beyond-view map generation. Given the novelty of the problem, we generate the first dataset specifically designed for it and evaluate the proposed approach. The results demonstrate consistent performance across diverse scenarios, establishing learning-based methods as a promising direction for map forecasting in autonomous driving. Beyond demonstrating the feasibility of the task, we provide an extensive discussion of the method's limitations and identify key future research directions for scaling it to more complex driving conditions. Code is available at https://git-autopia.car.upm-csic.es/beyondformer.
Chinese Translation
自动驾驶依赖高精地图(HD maps)实现安全导航。传统的高精地图构建在硬件、数据和人力资源方面成本高昂,且受更新限制的制约,难以扩展。近期工作提出了基于车载传感器的在线高精矢量化地图构建方案。然而,传感器视野有限,车辆前方重建地图的范围不足以支持安全规划。本文旨在解决这一局限,提出了全新的超视野矢量化地图生成问题:根据车辆感知区域的矢量化地图(视野内地图),生成合理的地图延续。为实验评估其可行性,我们提出了BeyondFormer,据我们所知,这是首个面向超视野地图生成的工作。鉴于该问题的新颖性,我们构建了首个专门为此设计的数据集,并对所提方法进行了评估。结果表明该方法在多种场景下均具有一致的表现,确立了基于学习的方法作为自动驾驶地图预测的一个有前景的方向。除了证明该任务的可行性之外,我们还对方法的局限性进行了深入讨论,并指出了将其扩展至更复杂驾驶条件的关键未来研究方向。代码可在 https://git-autopia.car.upm-csic.es/beyondformer 获取。
cs.RO / 74 / 2609.07516

P$^2$Calib: Utilizing Pattern Priors for LiDAR-Camera Extrinsic Calibration

P$^2$Calib:利用模式先验进行LiDAR-相机外参标定
Hu, Xiangcheng
Abstract
Target-based LiDAR-camera extrinsic calibration is a prerequisite for multi-sensor fusion in robotics. However, in the widely adopted four-hole pipeline, calibration accuracy is bottlenecked by LiDAR-side hole-center extraction, which suffers from sparse angular coverage and mixed-pixel corruption. This paper presents P$^2$Calib, which exploits pattern priors, geometric constraints specified by the CAD model of the target board, to improve calibration accuracy. First, we incorporate the known hole radius as a fitting constraint to prevent center estimates from degrading under sparse angular coverage. Building on the improved hole estimates, we further enforce the rigid rectangular layout of the four holes as a global consistency constraint to correct residual errors across holes. Both priors are integrated into an interactive calibration tool that provides a complete extrinsic calibration pipeline. Experiments on simulated and real datasets show that P$^2$Calib lowers the joint registration residual by 90\% and 82\% and the held-out reprojection error by 96\% and 77\% over the baseline. Code, https://github.com/JokerJohn/P2Calib.git, and data will be released to facilitate future research.
Chinese Translation
基于标定板的LiDAR-相机外参标定是机器人多传感器融合的前提条件。然而,在广泛采用的四孔标定流程中,标定精度受限于LiDAR侧的孔洞中心提取,其易受稀疏角度覆盖和混合像素干扰的影响。本文提出P$^2$Calib,利用模式先验——即由标定板CAD模型指定的几何约束——来提升标定精度。首先,我们将已知的孔洞半径作为拟合约束,以防止中心估计在稀疏角度覆盖下退化。在改进孔洞估计的基础上,我们进一步将四个孔洞的刚性矩形布局作为全局一致性约束,以校正孔洞间的残余误差。两种先验均集成于一个交互式标定工具中,提供了完整的外参标定流程。在仿真和真实数据集上的实验表明,P$^2$Calib相比基线方法将联合配准残差分别降低了90%和82%,将留出集重投影误差分别降低了96%和77%。代码(https://github.com/JokerJohn/P2Calib.git)和数据将开源发布,以促进未来研究。
cs.RO / 75 / 2609.07532

PhysReal: Learning Real-World Deformable Object Physics via Hybrid Constitutive Modeling

PhysReal:通过混合本构建模学习真实世界可变形物体的物理特性
Deng, Yinan, Song, Jianqiao, Zhang, Yisi, Wang, Yuhan, Wang, Jiahui, Yue, Yufeng
Abstract
Learning physically plausible dynamics from visual observations is essential for interactive world models and embodied agents. However, modeling real-world deformable objects remains challenging because their dynamics often arise from complex, spatially heterogeneous material responses. To address this challenge, we propose PhysReal, a video-driven framework for learning and simulating the underlying physics of real deformable objects. PhysReal integrates a spatially varying hybrid expert-neural constitutive model with a differentiable MPM simulator and 3DGS renderer. Analytical expert models provide interpretable physical priors, while neural constitutive residuals capture material responses beyond predefined formulations. Spatially distributed patches parameterize the constitutive field, enabling a continuous representation of local material variations. To organize the identification of this model from sparse visual observations, we adopt a progressive curriculum that sequentially optimizes global material properties, spatially varying local parameters, and neural constitutive residuals, together with complementary motion and mask supervision. Extensive experiments on diverse deformable-object interactions demonstrate that PhysReal achieves superior performance in dynamic reconstruction and future-state prediction, while showing strong potential for downstream robotic applications.
Chinese Translation
从视觉观测中学习物理合理的外观动力学对于交互式世界模型和具身智能体至关重要。然而,对真实世界可变形物体进行建模仍然极具挑战性,因为其动力学往往源于复杂且空间异质的材料响应。为应对这一挑战,我们提出了 PhysReal,一个用于学习和模拟真实可变形物体底层物理的视频驱动框架。PhysReal 将空间变化的混合专家-神经本构模型与可微分 MPM(物质点法)模拟器和 3DGS 渲染器相集成。解析专家模型提供可解释的物理先验,而神经本构残差则捕捉超出预定义表达式的材料响应。空间分布的 patch 对本构场进行参数化,实现对局部材料变化的连续表示。为了从稀疏的视觉观测中有条理地辨识该模型,我们采用渐进式课程策略,依次优化全局材料属性、空间变化的局部参数以及神经本构残差,并结合互补的运动与掩码监督。在多样化的可变形物体交互实验中,大量实验表明 PhysReal 在动态重建和未来状态预测方面取得了卓越性能,同时在下游机器人应用中展现出巨大潜力。
cs.RO / 76 / 2609.07534

Zero-Shot Sim-to-Real Contact-Rich Assembly via Proprioception-Anchored Cross-Modal Pretraining

基于本体感觉锚定跨模态预训练的零样本仿真到现实接触密集装配
Wang, Yuhan, Chen, Yurou, Jiang, Hongye, Lian, Wenzhao
Abstract
Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sustained contact. Although simulation-based reinforcement learning offers a scalable training paradigm, discrepancies in visual observations, contact dynamics, and force/torque (F/T) measurements often limit policy transfer. We observe that proprioception is comparatively consistent across domains because calibrated joint positions and consistently computed joint velocities align closely between simulation and hardware. Based on this observation, we present PACE (Proprioception-Anchored Cross-Modal Encoder), which supervises temporal visual and F/T representations by predicting proprioceptive state transitions. Static domain-specific factors, including lighting, texture, and sensor bias, contain little information about joint motion; the proposed objective therefore encourages the encoder to suppress these factors while retaining task-relevant motion cues. Policies trained on frozen PACE features are deployed on hardware without real-world fine-tuning or object-pose tracking. Across four contact-rich assembly tasks, PACE attains an average real-world success rate of 93.3\% and only a 2.7-percentage-point sim-to-real drop, meanwhile remaining robust to perturbations that substantially degrade pose-based and learned-fusion baselines.
Chinese Translation
接触密集型装配仍然具有挑战性,因为它需要亚毫米级的空间精度以及在持续接触过程中对力的可靠解读。尽管基于仿真的强化学习提供了一种可扩展的训练范式,但视觉观测、接触动力学以及力/力矩(F/T)测量方面的差异往往限制了策略迁移。我们观察到,本体感觉在各个域之间相对一致,因为经过标定的关节位置和一致计算的关节速度在仿真与硬件之间高度吻合。基于这一观察,我们提出了PACE(Proprioception-Anchored Cross-Modal Encoder,本体感觉锚定跨模态编码器),它通过预测本体感觉状态转移来监督时序的视觉和力/力矩表征。静态的域特定因素,如光照、纹理和传感器偏差,几乎不包含关于关节运动的信息;因此,所提出的目标函数促使编码器抑制这些因素,同时保留与任务相关的运动线索。基于冻结的PACE特征训练的策略可直接部署到真实硬件上,无需真实世界微调或物体位姿跟踪。在四项接触密集型装配任务中,PACE取得了平均93.3%的真实世界成功率,仿真到现实的性能下降仅为2.7个百分点,同时对那些会显著降低基于位姿方法和学习融合基线性能的扰动保持了鲁棒性。
cs.RO / 77 / 2609.07544

Anti-Gravity Walking by a Flying Humanoid Robot via Thrust-Rate Input Whole-Body Model Predictive Control

基于推力变化率输入全身模型预测控制的飞行仿人机器人反重力行走
Sugihara, Kazuki, Okada, Kei
Abstract
Flying humanoids are expected to perform tasks in diverse environments, while their existing locomotion is mainly limited to aerial flight and ground walking. The capability to move in complex three-dimensional space can greatly expand their application range. For such walking motion on ceilings and similar anti-gravity environments, whole-body MPC is effective. However, the discontinuous changes in dynamic structure accompanying contact switching during walking can induce thrust spikes, resulting in control instability. Therefore, in this work, we propose and implement a real-time whole-body MPC framework for anti-gravity bipedal walking. First, we formulate whole-body MPC using the time derivative of thrust, namely thrust-rate, as the control input. This formulation guarantees continuity of the thrust trajectory during contact switching while preserving the sparse structure of the optimal control problem for fast computation. Second, we address the lack of natural support forces in anti-gravity environments. We introduce lower bounds on the foot-normal component of the contact force, and smoothly transfer them during the doublesupport phase. Finally, we implement the proposed framework and demonstrate anti-gravity walking by a flying humanoid through simulation and a hardware experiment. To the best of our knowledge, this is the first demonstration of multi-contact whole-body MPC for a transformable aerial robot and walking by a flying humanoid beyond the ground.
Chinese Translation
飞行仿人机器人有望在多样化环境中执行任务,但其现有的运动能力主要局限于空中飞行和地面行走。在复杂三维空间中移动的能力可以极大地扩展其应用范围。对于此类在天花板及类似反重力环境中的行走运动,全身模型预测控制(MPC)是有效的。然而,行走过程中接触切换所伴随的动力学结构不连续变化会引起推力尖峰,导致控制不稳定。因此,本工作提出并实现了一个用于反重力双足行走的实时全身MPC框架。首先,我们以推力的时间导数(即推力变化率)作为控制输入来构建全身MPC。该构建方式保证了接触切换过程中推力轨迹的连续性,同时保留了最优控制问题的稀疏结构以实现快速计算。其次,我们解决了反重力环境中缺乏自然支撑力的问题。我们对接触力在足部法线方向的分量引入下界约束,并在双支撑阶段对其进行平滑过渡。最后,我们实现了所提出的框架,并通过仿真和硬件实验演示了飞行仿人机器人的反重力行走。据我们所知,这是首次针对可变形飞行机器人展示多接触全身MPC,也是首次展示飞行仿人机器人在地面之外行走。
cs.RO / 78 / 2609.07581

ICI-VLA: In-Context Imitation with Spatiotemporally Aligned Demonstrations for Vision-Language-Action Models

ICI-VLA:面向视觉-语言-动作模型的基于时空对齐示范的上下文模仿学习
Yang, Songhua, Liu, Ziyu, Li, Xuetao, Xiao, Ruqi, Zhu, Kangxin, Li, Miao
Abstract
Vision-Language-Action (VLA) policies are commonly adapted to new manipulation settings through additional gradient updates, which limits rapid deployment when task-specific data or compute is scarce. We present ICI-VLA, a training and retrieval framework that equips a text-action VLM with few-shot test-time adaptation through in-context demonstrations. Unlike mainstream VLA designs based on action-specific multimodal fusion, ICI-VLA retains the native text-generation interface. ICI-VLA updates its parameters only during offline training; at inference, the policy remains fixed and conditions action generation on retrieved micro-demonstrations. The framework decomposes long trajectories into short, semantically labeled examples and trains an RD-Encoder with positives mined by Dynamic Time Warping (DTW), aligning the retrieved context with the phase and geometry of the current subtask. We further introduce Target Action Masking, a context-corruption objective designed to reduce direct action copying and increase reliance on the current observation. ICI-VLA reaches average success rates of 97.7% on LIBERO and 60.4% on RoboTwin 2.0, exceeding the highest reported baseline average on RoboTwin 2.0 by 19.3 percentage points. It also achieves 83.2% across four physical tasks. These results indicate that a fixed VLA policy can benefit from conditioning on spatiotemporally aligned demonstrations at test time.
Chinese Translation
视觉-语言-动作(VLA)策略通常需要通过额外的梯度更新来适配新的操作场景,这在任务特定数据或算力稀缺时限制了快速部署。我们提出了ICI-VLA,一个训练与检索框架,通过上下文示范为文本-动作视觉语言模型(VLM)提供少样本测试时自适应能力。与主流的基于动作特定多模态融合的VLA设计不同,ICI-VLA保留了原生的文本生成接口。ICI-VLA仅在离线训练阶段更新参数;在推理时,策略保持固定,并基于检索到的微示范(micro-demonstrations)来条件化动作生成。该框架将长轨迹分解为短的、带有语义标注的示例,并训练一个RD-Encoder,其正样本通过动态时间规整(Dynamic Time Warping, DTW)挖掘,使检索到的上下文与当前子任务的阶段和几何特征对齐。我们进一步提出了目标动作掩码(Target Action Masking),这是一种上下文损坏目标,旨在减少对动作的直接复制,并增加对当前观测的依赖。ICI-VLA在LIBERO上达到97.7%的平均成功率,在RoboTwin 2.0上达到60.4%,超过RoboTwin 2.0上已报道的最高基线平均值19.3个百分点。此外,它在四个真实物理任务上达到83.2%的成功率。这些结果表明,固定的VLA策略可以通过在测试时以时空对齐的示范为条件而获益。
cs.RO / 79 / 2609.07747

Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction

Dex-X:通过模拟交互从人类视频学习视觉-触觉灵巧操作
Chen, Ruoqu, Ruan, Feixiang, Cao, Liu, Wang, Zihao, Xu, Botian, Tong, Shiqin, Liu, Jiajun, Pei, Mingzhi, Zhang, Chenyu, Xing, Wanli, Zhang, Kaifeng, Xu, Mengdi
Abstract
Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable visual-tactile dexterous manipulation policies from human video demonstrations without robot-side data collection? We present DEX-X, a framework for learning visual-tactile dexterous manipulation from human videos through simulation. Our key insight is that simulation can serve as a tactile completion engine. Given monocular human demonstrations, DEX-X reconstructs hand-object interactions in simulation, where physically grounded contact dynamics provide tactile supervision unavailable in the original videos. Leveraging this recovered tactile information, we train visual-tactile dexterous manipulation policies and distill them into deployable policies operating on point-cloud observations and tactile sensing. We demonstrate zero-shot sim-to-real transfer on a dexterous hand-arm platform across diverse grasping and contact-rich tool-use tasks. The teacher policy achieves 65.9% average success across six task categories in simulation, while the distilled visual-tactile policy achieves 93% success on real-world cube picking and 53% on the challenging table-cleaning task. Zero-shot generalization to unseen object geometries is also observed on object-picking tasks. Our results suggest that simulated interaction is a key bridge between human videos and deployable dexterous manipulation policies, providing the missing physical supervision needed for scalable robot skill learning from Internet-scale human video data.
Chinese Translation
人类视频是灵巧操作行为的丰富来源,但缺乏对接触密集型交互至关重要的触觉信息。这引出了一个根本性问题:机器人能否仅从人类视频演示中学习可部署的视觉-触觉灵巧操作策略,而无需在机器人侧进行数据采集?我们提出了 DEX-X,一个通过仿真从人类视频学习视觉-触觉灵巧操作的框架。我们的核心洞察是:仿真可以充当触觉补全引擎。给定单目人类演示视频,DEX-X 在仿真中重建手-物体交互,其中基于物理的接触动力学提供了原始视频中无法获得的触觉监督。利用这种恢复的触觉信息,我们训练视觉-触觉灵巧操作策略,并将其蒸馏为基于点云观测和触觉感知的可部署策略。我们在灵巧手-手臂平台上,针对多样化抓取和接触密集型工具使用任务,展示了零样本的仿真到真实(sim-to-real)迁移。教师策略在仿真中六个任务类别上达到 65.9% 的平均成功率,而蒸馏后的视觉-触觉策略在真实世界的立方体抓取任务上达到 93% 的成功率,在具有挑战性的桌面清洁任务上达到 53% 的成功率。在物体拾取任务上还观察到了对未见物体几何形状的零样本泛化能力。我们的结果表明,仿真交互是人类视频与可部署灵巧操作策略之间的关键桥梁,为从互联网规模人类视频数据中学习可扩展的机器人技能提供了缺失的物理监督。
cs.RO / 80 / 2609.07838

ComVLA: Communication-Aware Split Inference for VLA Models in 6G-Connected Robotics

ComVLA:面向6G连接机器人的VLA模型通信感知分割推理
Liu, Boliang, Poe, Wint Yi, Di, Jingyun, Trivisonno, Riccardo, Caire, Giuseppe
Abstract
Connected robotics is an emerging 6G application where mobile robots follow natural-language instructions to manipulate physical objects. The Vision-Language-Action (VLA) models that enable this are too large to run on the robot; a common trend is to offload inference to the cloud. The wireless link, however, limits how much sensing data the edge can transmit per control step. Two recent lines address this constraint: semantic communication codecs compress sensor data but require channel-specific retraining, and VLA token pruners select tokens from image but ignore the channel. Our insight is that the dense semantic information contained in the language already indicates which visual tokens matter. We propose ComVLA, a framework that uses this language guidance to adapt the VLA token budget to the channel capacity. Transmitting 32 tokens instead of 512 on the LIBERO benchmark, ComVLA cuts inference compute by 74% and inference latency by 22% versus the original OpenVLA-OFT baseline, at a cost of 1.5 pp in average task success (95.4% vs. 96.9%), and it stays within the capacity budget under Rayleigh and Rician fading. These results demonstrate that co-designing VLA inference and wireless communication is a practical direction for 6G-connected robotics.
Chinese Translation
连接机器人是6G的一项新兴应用,其中移动机器人根据自然语言指令操控物理对象。实现这一能力的视觉-语言-动作(Vision-Language-Action, VLA)模型规模过大,无法在机器人端运行,因此一种常见做法是将推理任务卸载至云端。然而,无线链路限制了边缘端在每个控制步内可传输的感知数据量。目前有两条研究方向应对这一约束:语义通信编解码器可以压缩传感器数据,但需要针对特定信道进行再训练;而VLA令牌剪枝方法仅从图像中选择令牌,却忽略了信道条件。我们的洞察在于,语言中所包含的密集语义信息已经指示了哪些视觉令牌是重要的。我们提出ComVLA框架,利用这种语言引导将VLA的令牌预算自适应地匹配到信道容量。在LIBERO基准上,ComVLA仅传输32个令牌而非512个,相比原始OpenVLA-OFT基线,推理计算量降低74%,推理延迟降低22%,而平均任务成功率仅下降1.5个百分点(95.4% vs. 96.9%),并且在瑞利(Rayleigh)和莱斯(Rician)衰落条件下仍能保持在容量预算之内。这些结果表明,VLA推理与无线通信的协同设计是6G连接机器人技术的一个可行方向。
cs.RO / 81 / 2609.07857

Scene Graph-Driven Haptic Feedback for Safety Enhancement in Robotic Ophthalmic Surgery via Physically Simulated iOCT

基于物理仿真iOCT的场景图驱动的触觉反馈用于增强机器人眼科手术安全性
Arbabi, Danial, Hoxha, Korab, Henriques, Angelo, Imamovic, Mirza, Nasseri, M. Ali
Abstract
Robotic ophthalmic surgery offers high precision but introduces a "sensory gap" by decoupling the surgeon from their instrument, resulting in a loss of tactile feedback. This paper presents a novel haptic feedback system for subretinal injection tasks leveraging Scene Graphs (SG). The system bridges the sensory gap by analyzing a physically simulated intraoperative Optical Coherence Tomography (iOCT) feed to construct a real-time surgical SG. The SG serves as a semantic abstraction layer for the surgical scene, which is then utilized by a deterministic, rule-based engine to generate state-dependent haptic feedback on a robotic input device. The system was evaluated in a user study (N=16) using an anthropomorphic head phantom and a custom-built surgical robot. Results demonstrate that the SG-driven haptic feedback improved surgical precision, reducing needle alignment error by 14% (p = 0.044) and improving System Usability Scale (SUS) scores by 8% (p = 0.015), while maintaining comparable task completion times. A needle trajectory analysis revealed the emergence of a safer "Align-then-Approach" strategy, in which our haptic negative reinforcement prompted users to fine-tune the tool's trajectory before approaching the retinal target. This work suggests that SGs can effectively serve as the direct computational foundation for real-time, safety-enhancing context-aware haptic feedback in robotic microsurgery.
Chinese Translation
机器人眼科手术具有高精度优势,但由于将外科医生与其手术器械解耦,造成了"感知鸿沟",导致触觉反馈的丧失。本文提出了一种利用场景图(Scene Graph, SG)的新型触觉反馈系统,用于视网膜下注射任务。该系统通过分析物理仿真的术中光学相干断层扫描(iOCT)图像流来构建实时手术场景图,从而弥合感知鸿沟。该场景图作为手术场景的语义抽象层,随后由确定性的、基于规则的引擎利用,在机器人输入设备上生成与状态相关的触觉反馈。该系统在一项用户研究(N=16)中进行了评估,研究使用了拟人化头部模型和自行研制的手术机器人。结果表明,场景图驱动的触觉反馈提高了手术精度,将针头对准误差降低了14%(p = 0.044),并将系统可用性量表(SUS)得分提高了8%(p = 0.015),同时保持了相当的任务完成时间。针头轨迹分析揭示了一种更安全的"先对准后接近"策略的出现,即我们的触觉负反馈促使用户在接近视网膜目标之前先微调工具的轨迹。这项工作表明,场景图可以有效地作为机器人显微手术中实时、增强安全性的情境感知触觉反馈的直接计算基础。
cs.RO / 82 / 2609.07859

M3-Tele: A Unified Multimodal Teleoperational Framework for Compliant Whole-Body Mobile Manipulation

M3-Tele:面向柔顺全身移动操作的一体化多模态遥操作框架
Chen, Hengxiang, Deng, Shenwen, Ma, Yujian, Ma, Gan, Li, Qiang, Chen, Nutan
Abstract
Executing contact-rich tasks efficiently requires the seamless integration of whole-body coordination and physical compliance regulation. However, existing teleoperation and data-collection frameworks often overlook the joint consideration of multimodal perception and coordinated whole-body operation. This limitation can reduce the efficiency and quality of demonstration collection, thereby affecting the effectiveness of downstream policy learning. In this work, we present \textbf{M3-Tele}: A Unified \underline{M}ultimodal \underline{Tele}operational Framework for Compliant Whole-Body \underline{M}obile \underline{M}anipulation, enabling stable physical interaction and capturing aligned visual, tactile, force, and proprioceptive observations during task execution. Extensive experiments demonstrate that the proposed framework significantly improves contact-rich teleoperation performance. The proposed controller reduces the force tracking error from 4.132~N to 0.346~N, the contact loss from 2.46 to 0.02 events per trial and the tactile deformation error by 65\%. User studies across four mobile manipulation tasks also verify the reliability and usability of the proposed system. Furthermore, Diffusion Policy experiments highlight the value of joint tactile and force sensing.
Chinese Translation
高效执行富接触任务需要全身协调与物理柔顺调节的无缝集成。然而,现有的遥操作与数据采集框架往往忽视了多模态感知与全身协同操作的联合考量。这一局限会降低示教数据采集的效率和质量,进而影响下游策略学习的有效性。在本工作中,我们提出了M3-Tele:一个面向柔顺全身移动操作的统一多模态遥操作框架,能够实现稳定的物理交互,并在任务执行过程中同步采集对齐的视觉、触觉、力以及本体感觉观测。大量实验表明,所提出的框架显著提升了富接触遥操作的性能。所提出的控制器将力跟踪误差从4.132 N降低至0.346 N,将每次试验中的接触丢失次数从2.46次降低至0.02次,并将触觉形变误差降低了65%。针对四项移动操作任务的用户研究也验证了所提出系统的可靠性与可用性。此外,Diffusion Policy(扩散策略)实验凸显了触觉与力感知联合采集的价值。
cs.RO / 83 / 2609.07905

Conditional Timed Partial Orders: An Expressive and Interpretable Framework for Robot Task Specification and Planning

条件时序偏序:一种面向机器人任务规范与规划的强表达性且可解释的框架
Escobar, Sebastian, Lahijanian, Morteza
Abstract
Timed Partial Orders (TPOs), originally proposed for workflows, provide an interpretable framework for robot task specification with planning algorithms based on mixed-integer linear programming (MILP). However, TPOs are limited in expressivity, capturing only partial-order events with simple timing constraints. In this paper, we introduce Conditional TPOs (cTPOs), which extend TPOs with richer relative-timing constraints and conditional event activations based on environmental conditions. We show that planning for cTPOs also reduces to an MILP problem; however, the added expressivity results in significantly larger MILPs that can become computationally intractable. To address this challenge, we propose a decomposition algorithm that partitions a cTPO into smaller sub-TPOs, yielding a sequence of smaller MILP problems. We prove that this decomposition is complete and preserves plan optimality while improving the interpretability of complex tasks. Experimental results demonstrate the effectiveness of cTPOs as a task specification framework and the efficiency of our decomposition approach, achieving up to four orders of magnitude speedup over the monolithic MILP.
Chinese Translation
时序偏序(Timed Partial Orders, TPOs)最初是为工作流提出的,为机器人任务规范提供了一种可解释的框架,并配有基于混合整数线性规划(MILP)的规划算法。然而,TPOs 的表达能力有限,仅能刻画具有简单时序约束的偏序事件。本文提出条件时序偏序(Conditional TPOs, cTPOs),通过更丰富的相对时序约束以及基于环境状态的条件事件激活来扩展 TPOs。我们证明,cTPOs 的规划问题同样可归结为 MILP 问题;然而,表达能力的增强会导致 MILP 规模显著增大,从而可能在计算上变得难以处理。为应对这一挑战,我们提出一种分解算法,将一个 cTPO 划分为若干较小的子 TPO,从而生成一系列规模更小的 MILP 问题。我们证明该分解是完备的,并能在提升复杂任务可解释性的同时保持规划的最优性。实验结果表明,cTPOs 作为任务规范框架是有效的,且我们的分解方法具有高效率,相较于整体式 MILP 方法可获得高达四个数量级的加速。
cs.RO / 84 / 2609.07930

A Multimodal Label Forecasting Method for Aperiodic Visuo-Motor Time Series

一种面向非周期性视觉-运动时间序列的多模态标签预测方法
He, Borui, Katz, Garrett E
Abstract
Deep learning models have been increasingly applied to Time Series Forecasting (TSF) in recent years. Transformer-based and MLP-based models have both been used effectively on many real-world TSF regression benchmarks, and there is ongoing debate as to which family of methods is best. While these benchmarks have drawn much attention, it is also worth noting that many current datasets and methods assume approximate periodicity in the time series. In this work, we focus on a new TSF task without periodicity: anticipating falls during humanoid locomotion, on the basis of egocentric vision and proprioception. When the locomotion trajectories are sufficiently diverse, periodicity is violated. We contribute two new benchmark datasets (one from simulation, one from real hardware), showing that periodicity is violated and recent deep TSF methods struggle on these benchmarks. We also propose a novel deep learning architecture that exploits both endogenous and exogenous variables and a training process that rigorously enforces i.i.d sampling of training examples. Our results show statistically significant improvement over prior art in multiple experimental conditions, by 12.73% or more on the real data and 10.40% or more on the simulation data. Code and datasets will be available upon acceptance.
Chinese Translation
近年来,深度学习模型被越来越多地应用于时间序列预测(Time Series Forecasting, TSF)。基于Transformer的模型和基于MLP的模型都已在许多真实世界的TSF回归基准上得到有效应用,并且关于哪一类方法更优的争论仍在持续。尽管这些基准受到了广泛关注,但值得注意的是,许多现有数据集和方法都假设时间序列具有近似周期性。在本工作中,我们关注一个全新的无周期性TSF任务:基于自我中心视觉和本体感觉,预测人形机器人行走过程中的跌倒。当行走轨迹足够多样化时,周期性即被破坏。我们贡献了两个新的基准数据集(一个来自仿真,一个来自真实硬件),并表明在这些基准上周期性被违反,且近期深度TSF方法表现不佳。我们还提出了一种新颖的深度学习架构,能够同时利用内生变量和外生变量,并采用一种严格强制训练样本独立同分布(i.i.d.)采样的训练过程。实验结果表明,我们的方法在多种实验条件下相对现有技术取得了统计显著的改进:在真实数据上提升12.73%或更多,在仿真数据上提升10.40%或更多。代码和数据集将在论文被接收后公开。
cs.RO / 85 / 2609.07933

SPOT: Spatial Perception-Oriented Long-Horizon Humanoid Teleoperation

SPOT:面向空间感知的长时域人形机器人遥操作
Fang, Lixing, Xiong, Ziyan, Chen, Sunli, Dou, Zhiyang, Gan, Chuang
Abstract
High-quality demonstration data is becoming a central bottleneck for training general-purpose humanoid robots. While recent humanoid teleoperation systems have made substantial progress in retargeting human motion to robot motion, long-horizon loco-manipulation requires another capability: operators must maintain task-relevant spatial awareness over time, e.g., object locations, surrounding environments, the robot's pose. We call the extent of this awareness the operator's perceptual horizon. However, existing methods often shorten this: narrow views miss peripheral events, robot-mounted cameras become unstable during locomotion, and coupled head-view control makes looking around interfere with robot motion. We present SPOT, a Spatial Perception-Oriented VR Teleoperation system for collecting long-horizon humanoid demonstration data by providing extended perceptual horizon. SPOT combines a robot-mounted binocular fisheye camera, a wide-field stereoscopic display, viewpoint-decoupled free-looking, and visual stabilization to provide a robot-centric view that is wide, stable, and actively inspectable. Unlike conventional egocentric interfaces, SPOT decouples visual exploration from robot actuation: the egocentric stereo observation is rendered on a virtual hemisphere around the operator, so natural head rotations change where the operator looks within the wide-field view rather than commanding the robot head, camera, or torso. We evaluate SPOT on perception-critical humanoid data-collection tasks spanning drop recovery, peripheral retrieval, large-workspace bimanual manipulation, fine alignment, and dynamic interaction. SPOT improves efficiency, accuracy, and recovery speed, demonstrating its effectiveness for user-friendly and scalable long-horizon humanoid data collection.
Chinese Translation
高质量的示范数据正成为训练通用人形机器人的核心瓶颈。尽管近年来的人形机器人遥操作系统在人运动到机器人运动的重定向方面取得了长足进展,但长时域的移动操作还需要另一种能力:操作者必须在时间维度上保持与任务相关的空间感知,例如物体的位置、周围环境以及机器人的姿态。我们将这种感知的程度称为操作者的感知视界(perceptual horizon)。然而,现有方法往往会缩短这一视界:狭窄的视野会遗漏边缘区域的事件,安装在机器人上的相机在行走过程中会变得不稳定,而头部视角耦合控制则使得环顾四周会干扰机器人运动。我们提出了 SPOT——一个面向空间感知的 VR 遥操作系统,通过扩展感知视界来采集长时域的人形机器人示范数据。SPOT 结合了安装在机器人上的双目鱼眼相机、宽视场立体显示、与视点解耦的自由观察以及视觉稳定技术,提供宽广、稳定且可主动巡视的以机器人为中心的视角。与传统的第一人称交互界面不同,SPOT 将视觉探索与机器人执行解耦:第一人称立体观测被渲染在操作者周围的虚拟半球上,因此头部的自然转动只会改变操作者在宽视场画面中的观察位置,而不会命令机器人头部、相机或躯干运动。我们在多个感知关键的人形机器人数据采集任务上评估了 SPOT,涵盖掉落物恢复、边缘区域取物、大工作空间双手操作、精细对准以及动态交互。SPOT 提升了效率、精度和恢复速度,证明了其在用户友好且可扩展的长时域人形机器人数据采集方面的有效性。
cs.RO / 86 / 2609.08010

mjorbit: A Simulation Framework for Space Robotics

mjorbit:一个面向空间机器人技术的仿真框架
Zhang, John Z., Verhagen, Joris, Vega, Fausto, McKeen, Patrick, Manchester, Zachary
Abstract
This paper presents a general framework for simulating multi-body space robots with contact. We bring efficient, large-scale robot simulation to in-space servicing, assembly, and manufacturing applications. First, we perform an empirical trade study of methods for coupling orbit propagation with existing robotics simulation frameworks. Next, we present mjorbit, a general, flexible, and performant framework built on the MuJoCo engine widely used in robotics, to which we add key spacecraft dynamics, actuators, and sensors. We provide a low-latency C++ CPU backend and a high-throughput GPU backend behind a simple Python API. We demonstrate mjorbit by solving several realistic on-orbit case studies with both model-predictive control and reinforcement learning. Open-source code and examples are available at: https://johnzhang3.github.io/mjorbit/
Chinese Translation
本文提出了一个用于仿真带接触的多体空间机器人的通用框架。我们将高效的大规模机器人仿真引入在轨服务、装配与制造应用。首先,我们对轨道传播与现有机器人仿真框架耦合的方法进行了实证权衡研究。其次,我们提出了mjorbit——一个基于机器人领域广泛使用的MuJoCo引擎构建的通用、灵活且高性能的框架,并为其添加了关键的航天器动力学、执行器和传感器模型。我们在一个简单的Python API背后提供了低延迟的C++ CPU后端和高吞吐量的GPU后端。我们通过使用模型预测控制和强化学习求解多个真实的在轨案例研究来演示mjorbit。开源代码和示例可在以下网址获取:https://johnzhang3.github.io/mjorbit/
cs.RO / 87 / 2609.08123

DISEIL: Demonstration Distillation for Sample-Efficient Imitation Learning

DISEIL:面向样本高效模仿学习的示范蒸馏方法
Khanal, Suyog, A V, Arun Kumar, Rana, Santu
Abstract
A robot that can be taught a new task from a handful of demonstrations has to work out for itself what it still cannot do, and then ask for exactly that. Interactive imitation learning takes a step in that direction by letting a policy practice on its own and calling an expert when it goes wrong. Existing methods decide when to interrupt the learner. A further 2 decisions are left to whichever episode happened to trigger the interruption: which failure to correct, and where the demonstration should start. This paper is a first attempt at making both of them deliberately. DISEIL (Demonstration dIstillation for Sample-Efficient Imitation Learning) marks each failed episode at the step where the policy first becomes unreliable, represents that moment with a geometric descriptor, and groups the failures into recurring failure modes. A vision-language model and a language model read the selected mode and write a request for the next demonstration, and a store of task constraints checks that the request can be carried out before any expert time is spent. No model produces a robot action. Across 5 simulated tasks under state and image observations, changing only what the expert is asked for gives the highest mean held-out success rate in all 10 settings, with a tie in 1, and the margin is widest at the smallest budget we tested. The scope is narrow: a single round of practice at a time, in simulation, with experts that are mostly scripted. The longer-term aim is a learner that also tracks what its demonstration set already covers, and that asks a human teacher for the missing behavior in proportion to the effort each request costs them.
Chinese Translation
一个能够通过少量示范学习新任务的机器人,必须自行判断自己还无法完成哪些操作,然后有针对性地进行询问。交互式模仿学习朝这一方向迈出了一步:它允许策略自主练习,并在出错时呼叫专家。现有方法只决定何时中断学习过程,而其余两个决策则交由恰好触发中断的回合来决定:即纠正哪个失败,以及示范应从何处开始。本文首次尝试对这两个决策进行审慎设计。DISEIL(Demonstration dIstillation for Sample-Efficient Imitation Learning,面向样本高效模仿学习的示范蒸馏)在策略首次变得不可靠的步骤处标记每个失败回合,用几何描述符表示该时刻,并将失败归类为重复出现的失败模式。视觉-语言模型和语言模型读取所选的失败模式并撰写下一次示范的请求,同时一个任务约束存储库会在消耗专家时间之前检查该请求是否可以执行。整个过程没有任何模型产生机器人动作。在状态和图像观测下的5个仿真任务中,仅改变对专家所请求的内容这一项,就在全部10个设置中取得了最高的平均留出成功率(其中1个为并列),且在所测试的最小预算下优势最为明显。本方法的适用范围较窄:每次仅进行一轮练习,且限于仿真环境,专家大多是脚本化的。更长远的目标是构建一个能够追踪其示范集已覆盖内容的学习器,并根据每次请求所需的人力成本,按比例向人类教师请求缺失的行为。
cs.RO / 88 / 2609.08159

OmniNav: Robust Long-Horizon Target Navigation in Dynamic Environments

OmniNav:动态环境中鲁棒的长时程目标导航
Tang, Yujie, Wang, Meiling, Jiang, Jinhao, Zuo, Sibo, Deng, Yinan, Zhang, Xinyu, Yue, Yufeng
Abstract
Long-horizon target navigation requires a robot to sustain task execution across evolving observations, decisions, and physical interactions. This requires three coupled capabilities: maintaining valid scene memory, revising target beliefs under partial observability, and selecting interaction-feasible navigation endpoints. However, the state underlying each capability is only conditionally valid: scene representations become stale when objects move or disappear, unsuccessful searches alter beliefs over target locations, and geometrically convenient endpoints may still be infeasible for manipulation. To address these challenges, we present OmniNav, which formulates long-horizon navigation as continual inference over a factorized task state posterior coupling scene validity, target belief, and interaction feasibility. For representation, OmniNav incrementally constructs an updatable 3D object scene memory, preventing stale scene evidence from propagating to subsequent decisions. For exploration, it introduces an evidence-aware Bayesian belief-revision mechanism that derives dependency-aware region priors from semantic context, incorporates unsuccessful searches as negative evidence, and updates them for posterior-guided frontier selection. For interaction, OmniNav incorporates manipulation reachability and collision constraints into navigation-endpoint selection and propagates execution feedback through hierarchical closed-loop recovery. Extensive experiments demonstrate that OmniNav achieves the highest success rates among the compared methods on semantic ObjectNav and fine-grained instance navigation benchmarks, remains robust to target relocation, and improves real-world pick-and-place success from 53.3% to 71.7% over an adapted open-loop baseline. The project page of OmniNav is available at https://omni-nav.github.io/.
Chinese Translation
长时程目标导航要求机器人在不断演化的观测、决策和物理交互中持续执行任务。这需要三种相互耦合的能力:维护有效的场景记忆、在部分可观测条件下修正目标信念,以及选择可交互的导航终点。然而,每种能力所依赖的状态仅在特定条件下有效:当物体移动或消失时,场景表征会过时;失败的搜索会改变对目标位置的信念;几何上便利的终点对于操作而言可能仍不可行。为应对这些挑战,我们提出了OmniNav,将长时程导航表述为对一个分解的任务状态后验的持续推断,该后验耦合了场景有效性、目标信念和交互可行性。在表征方面,OmniNav增量式地构建可更新的三维物体场景记忆,防止过时的场景证据传播到后续决策中。在探索方面,它引入了一种证据感知的贝叶斯信念修正机制,从语义上下文中推导依赖感知的区域先验,将失败的搜索纳入为负证据,并据此更新以实现后验引导的前沿选择。在交互方面,OmniNav将操作可达性与碰撞约束纳入导航终点选择,并通过分层闭环恢复传播执行反馈。大量实验表明,OmniNav在语义ObjectNav和细粒度实例导航基准上取得了所对比方法中最高的成功率,对目标位置变动保持鲁棒,并将真实世界中抓取放置任务的成功率相较于改进的开环基线从53.3%提升至71.7%。OmniNav的项目主页可在 https://omni-nav.github.io/ 获取。
cs.RO / 89 / 2609.08164

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

面向空中目标导航的双层语义-空间信念地图构建
Xiao, Jianqiang, Deng, Xiang, Sun, Yuexuan, Wu, Yanjin, Yan, Wenbiao, Nie, Liqiang
Abstract
Aerial Object Goal Navigation (ObjectNav) requires an unmanned aerial vehicle (UAV) to locate a described target in an unknown outdoor environment using onboard visual observations. Vision-language models (VLMs) can interpret open-ended target descriptions and visual observations, but their frame-level outputs are often noisy, sparse, and spatially transient. We propose AeroBelief, a dual-layer semantic-spatial belief mapping framework that transforms transient VLM observations into persistent spatial guidance. It separates broad contextual plausibility from target-specific evidence: an intuition layer accumulates scene-level semantic cues for exploration, while an evidence layer preserves qualified target-specific observations for approach and confirmation. Evidence-gated fusion combines the two layers into spatial belief hotspots. We further introduce object-conditioned visual reasoning with conservative evidence qualification to improve observation reliability before spatial accumulation. In parallel, egocentric regional guidance converts quadtree coverage into UAV-centered, yaw-aligned directional proposals and stabilizes them through temporal commitment. Its regional scoring is independent of semantic belief values, maintaining exploration pressure and reducing repeated low-gain search. Experiments on the UAV-ON benchmark show that AeroBelief achieves the best reported overall SR, OSR, and SPL among the compared methods, reaching 21.61%, 35.57%, and 10.62, respectively. These results support the effectiveness of persistent semantic-spatial belief, conservative evidence qualification, and temporally stable regional guidance for aerial ObjectNav.
Chinese Translation
空中目标导航(ObjectNav)要求无人机(UAV)利用机载视觉观测在未知室外环境中定位所描述的目标。视觉-语言模型(VLM)能够理解开放式的目标描述和视觉观测,但其帧级输出往往存在噪声大、稀疏且空间上短暂的问题。我们提出AeroBelief,一种双层语义-空间信念地图构建框架,可将短暂的VLM观测转化为持久的空间引导。该框架将广泛的上下文合理性与目标特定证据分离:直觉层累积场景级语义线索以支持探索,而证据层保留经过筛选的目标特定观测以用于接近和确认目标。证据门控融合将两层结合为空间信念热点。我们进一步引入对象条件视觉推理与保守的证据筛选机制,以在空间累积之前提高观测的可靠性。与此同时,以自我为中心的区域引导将四叉树覆盖转换为以无人机为中心、与偏航对齐的方向性建议,并通过时间承诺机制使其保持稳定。其区域评分独立于语义信念值,从而维持探索压力并减少重复的低收益搜索。在UAV-ON基准上的实验表明,AeroBelief在所比较的方法中取得了最佳的整体SR、OSR和SPL成绩,分别达到21.61%、35.57%和10.62。这些结果验证了持久语义-空间信念、保守证据筛选以及时间稳定的区域引导对空中ObjectNav的有效性。
cs.RO / 90 / 2609.08209

Monkey See, Can Monkey Do? A Benchmark for Evaluating Robot Skill Learning by Observation

猴子观察能否模仿?一个用于评估机器人通过观察进行技能学习的基准测试
Gu, Weiwei, Gupta, Anmol, Sah, Anant, Varghese, Ryan, Vanam, Lalitha Shreya, Adireddi, Prabhath, Karkus, Peter, Gopalan, Nakul
Abstract
Learning from Observation (LfO) is a fundamental robotic capability that replicates how humans and animals socially learn from each other. Beyond its biological parallels, this modality provides a practical solution for data scaling in sample-inefficient and data-starved domains like robotics. Recent work has demonstrated promising results in learning manipulation skills from human videos, yet progress in this area remains difficult to assess. Existing methods vary widely in assumptions, hardware choices, and environment setups making it difficult to draw meaningful comparisons and identify advances in the field. To address these challenges, we introduce RoboReel: a unified benchmark for evaluating models that learn policies from human videos. RoboReel consists of bundled real-world human demonstration videos, simulated robot trajectories, and evaluation environments on ten manipulation tasks. We develop four test suites to evaluate the models' performance on multiple axes, including the robustness to visual distractors and the ability to complete long-horizon tasks. Our benchmark covers learning-from-observation models from different categories, and studies the effectiveness of multiple representation choices in our benchmark evaluation that covers over seven state-of-the-art algorithms (including our VLA based variants) in the field of LfO. Finally, we present an analysis of the different types of algorithms showing that long-horizon tasks and tasks with low tolerances are still challenging for current models. Webpage: https://roboreel.github.io
Chinese Translation
从观察中学习(Learning from Observation, LfO)是机器人的一项基本能力,它复制了人类和动物通过社会交互相互学习的方式。除了其生物学上的相似性之外,这种模态还为机器人学等样本效率低、数据匮乏的领域中的数据扩展提供了一种实用的解决方案。近期的工作已经展示了从人类视频中学习操作技能的可喜成果,但该领域的进展仍然难以评估。现有方法在假设条件、硬件选择和环境设置方面差异很大,这使得难以进行有意义的比较,也难以识别该领域的进展。为应对这些挑战,我们提出了 RoboReel:一个用于评估从人类视频中学习策略的模型的统一基准。RoboReel 包含捆绑的真实世界人类示范视频、仿真机器人轨迹以及十个操作任务的评估环境。我们开发了四个测试套件,从多个维度评估模型的性能,包括对视觉干扰的鲁棒性以及完成长时程任务的能力。我们的基准涵盖了来自不同类别的从观察中学习模型,并研究了多种表示选择在我们基准评估中的有效性,该评估涵盖了 LfO 领域中超过七种最先进的算法(包括我们基于 VLA 的变体)。最后,我们对不同类型的算法进行了分析,结果表明长时程任务以及容错率低的任务对当前模型而言仍然具有挑战性。网页:https://roboreel.github.io
cs.RO / 91 / 2609.08214

TacClip: a clip-on sensor measures dynamic contact forces without covering the fingerpads

TacClip:一种不遮挡指腹的夹扣式动态接触力传感器
Ye, Yuqian, Li, Hao, Xu, Jingxi, Feng, Haojun, Hong, Seongheon, Cutkosky, Mark R.
Abstract
TacClip is a minimally encumbering wearable device for recording fingertip deformation caused by contact forces and vibrations. It can be combined with vision- or glove-based hand tracking systems that leave the fingertips uncovered and provides a measure of dynamic contact interactions, while leaving the finger pads exposed so that the user retains natural sensitivity to texture, friction, temperature, and fine surface features. The signal is produced by a Fiber Bragg Grating (FBG) embedded on a small plastic clip mounted over the fingernail. Optionally, for use with vision-based tracking, additional FBGs on polyimide strips can complement camera-based pose estimation. In finger pressing tests, TacClip estimates the force magnitude with typical errors below $0.5~\mathrm{N}$ over a $0$--$8~\mathrm{N}$ range. In tests of cloth handling and tape edge finding, we show that it captures the vibrations and dynamic events generated during exploratory sliding. With no electronics, TacClip can also be used submerged in water, while preserving bare finger contact.
Chinese Translation
TacClip 是一种轻便低负担的可穿戴设备,用于记录由接触力和振动引起的指尖变形。它可以与基于视觉或手套的手部追踪系统结合使用,保持指尖不被遮挡,从而提供动态接触交互的测量,同时指腹保持裸露,使用户保留对纹理、摩擦、温度和细微表面特征的自然感知。信号由嵌入在安装于指甲上方的小型塑料夹上的光纤布拉格光栅(FBG)产生。此外,在基于视觉的追踪应用中,可选择在聚酰亚胺条上添加额外的光纤布拉格光栅,以辅助基于摄像头的姿态估计。在手指按压测试中,TacClip 在 0–8 N 范围内估计力的大小的典型误差低于 0.5 N。在布料操作和胶带边缘探测测试中,我们展示了该设备能够捕捉探索性滑动过程中产生的振动和动态事件。由于不含电子元件,TacClip 还可以在水下使用,同时保持手指的裸露接触。
cs.RO / 92 / 2609.08220

Bridging Language and Physics: Automated Design of Continuum Robots with Large Language Models

连接语言与物理:基于大语言模型的连续体机器人自动化设计
Chen, Jingyi, Zhang, Mohan, Yao, Laura, Ni, Yingtai, Ji, Jianmin, Peng, Jie, Wang, Song, Chen, Tianlong
Abstract
Large language models (LLMs) have recently emerged as a promising tool for automating robot design from high-level specifications, yet they remain ineffective for robots operating under complex physical interactions. This limitation stems from the gap between language-based reasoning and the physical consequences of embodiment, often resulting in designs with low physical validity. In this work, we propose a multi-layered framework, AID-SR, that establishes a closed loop by translating simulator-observed physical states into structured feedback for the LLM designer. Combined with semantic critique, human feedback, and iterative refinement, the framework promotes the generation of physically feasible and functionally meaningful robot designs. We evaluate our approach on tendon-driven continuum robots across a benchmark of 14 tasks spanning reaching, grasping, locomotion, and manipulation. The proposed framework achieves 96.2% rate for passing the simulation feasibility check and by applying a common reinforcement learning training, 26.7% robots can successfully fulfill the corresponding task. We then fabricate three designed robots of AID-SR that successfully complete the task in real-world. These extensive experiments across simulation and real-world environments demonstrate and break the wall of utilizing the LLMs for automated design of continuum robots. The source code and experimental resources are publicly available at https://github.com/UNITES-Lab/AID-SR.
Chinese Translation
大语言模型(LLMs)近来已成为根据高层规格自动化机器人设计的一种有前景的工具,但其在应对复杂物理交互的机器人设计方面仍然效果不佳。这一局限源于基于语言的推理与具身物理后果之间的鸿沟,常导致设计方案物理有效性较低。在本工作中,我们提出了一种多层框架AID-SR,通过将仿真器观测到的物理状态转化为面向LLM设计器的结构化反馈,建立起闭环机制。结合语义批评、人类反馈和迭代优化,该框架有助于生成物理上可行且功能上有意义的机器人设计。我们在肌腱驱动连续体机器人上评估了该方法,涵盖到达、抓取、运动和操作等14个任务的基准测试。所提出的框架在仿真可行性检查中的通过率达到96.2%,并且经过共同的强化学习训练后,26.7%的机器人能够成功完成相应任务。随后,我们制造了AID-SR设计的三个机器人,并成功在真实环境中完成任务。这些在仿真和真实环境中的大量实验证明并突破了利用大语言模型进行连续体机器人自动化设计的壁垒。源代码和实验资源已公开发布于 https://github.com/UNITES-Lab/AID-SR。
cs.RO / 93 / 2609.08224

3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints

3DWay:通过三维一致的路径点实现机器人操作的泛化
Huang, Ziqin, Li, Yingyue, Zhang, Chenyangguang, Zhang, Ruida, Chen, Yuxin, Wang, Gu, Liu, Xingyu, Tomizuka, Masayoshi, Ji, Xiangyang
Abstract
Intermediate representations are key to bridging the modality gap between generalizable manipulation policies and large-scale pretrained vision-language models (VLMs). Among these, trajectory-based representations compactly represent motion-relevant cues, yet most existing approaches predict trajectories in 2D image space, resulting in intrinsic 3D ambiguity. Moreover, using 2D trajectories with depth still leaves the free-space waypoints ambiguous, limiting reliable 3D reasoning. To address this, we propose predicting 3D consistent waypoints (3DWay) from multi-view images. By reformulating 3D waypoints prediction as generating multi-view consistent 2D waypoints followed by geometric triangulation, we enable explicit 3D motion specification while preserving the strong priors of pretrained VLMs. The predicted waypoints can guide existing VLA models for better generalization or be directly executed on simple tasks. Extensive experiments show that 3DWay substantially improves 3D spatial grounding and vision-language reasoning, demonstrating strong potential for generalizable robot manipulation. Codes will be released at https://github.com/ziqin-h/3DWay.
Chinese Translation
中间表示是弥合可泛化操作策略与大规模预训练视觉-语言模型(VLM)之间模态鸿沟的关键。在各类表示中,基于轨迹的表示能够紧凑地刻画与运动相关的线索,然而现有方法大多在二维图像空间中预测轨迹,导致固有的三维歧义性。此外,即使结合深度信息使用二维轨迹,自由空间中的路径点仍然存在歧义,限制了可靠的三维推理能力。为解决这一问题,我们提出从多视角图像中预测三维一致的路径点(3DWay)。通过将三维路径点预测重新表述为先生成多视角一致的二维路径点、再进行几何三角化的过程,我们在保留预训练VLM强大先验的同时,实现了显式的三维运动指定。所预测的路径点可用于引导现有VLA模型以获得更好的泛化能力,或在简单任务上直接执行。大量实验表明,3DWay显著提升了三维空间定位与视觉-语言推理能力,展现出在可泛化机器人操作方面的巨大潜力。代码将发布于 https://github.com/ziqin-h/3DWay。
cs.RO / 94 / 2609.08250

CALIPER: Clean Scenes Cannot Rank Physical Inference in Pretrained Visual Representations

CALIPER:干净场景无法对预训练视觉表征中的物理推理能力进行排序
Mehta, Aman, Baviskar, Riya
Abstract
How far a pushed object slides depends on its mass and friction, which no single image reveals. Pretrained visual encoders are increasingly used as the perception front end of world models for manipulation, and their physical competence is assessed with perturbation benchmarks and linear probes, almost always in a clean, fixed-camera scene. We show that these assessments cannot distinguish an encoder that infers physics from one that does not. CALIPER (calibrate, then predict) is a direct test: an object of unknown mass and friction is struck twice at known speeds, a third strike is shown only up to the moment of contact, and a linear readout on frozen features must predict how far the object slides. Swapping in another object's calibration clips checks that the evidence is actually used. Across 2,000 simulated episodes and eight representations, from V-JEPA 2 to a randomly initialised ViT and raw pixels, calibration adds +0.50 R^2 and the swap removes it. Yet in the clean scene every representation lands within 0.02 R^2 of the ceiling set by true simulator state, because a fixed camera exposes the object's displacement directly in pixel coordinates. Resampling camera, lighting, and clutter for every clip spreads the same representations across 0.50 R^2; when the readout chooses a push speed for a goal distance, V-JEPA 2 misses by 4 mm and the random ViT by 20 mm, no better than ignoring the object. Linear probes track none of this: a change in frame aggregation moves a probe more than pretraining does, and erasing the probed mass direction from the same representation costs nothing in one scene and 0.35 R^2 in the other. Whether a benchmark can rank models is an empirical property, and we give three checks that establish it.
Chinese Translation
一个被推动的物体滑动多远取决于其质量和摩擦力,而任何单张图像都无法揭示这些属性。预训练视觉编码器越来越多地被用作操作世界模型的感知前端,其物理能力通常通过扰动基准测试和线性探针来评估,且几乎总是在干净、固定相机的场景中进行。我们证明,这类评估无法区分一个真正推断物理规律的编码器和一个做不到的编码器。CALIPER(先校准、后预测)是一种直接测试:一个质量和摩擦未知的物体被以已知速度撞击两次,第三次撞击仅展示至接触瞬间,然后必须在冻结特征上的线性读出预测物体滑动多远。换用另一个物体的校准片段作为对照,以验证证据确实被使用。在2,000个模拟回合和八种表征(从V-JEPA 2到随机初始化的ViT乃至原始像素)上的实验中,校准带来了+0.50的R²提升,而交换校准则消除了这一提升。然而在干净场景中,所有表征的R²都落在真实模拟器状态所设定上限的0.02以内,因为固定相机使物体的位移直接暴露在像素坐标中。为每个片段重新采样相机、光照和杂物,会使相同的表征在0.50的R²范围内分散;当读出网络需要为给定目标距离选择推动速度时,V-JEPA 2偏差4毫米,而随机ViT偏差20毫米,并不比忽略物体更好。线性探针完全无法追踪这些差异:改变帧聚合方式对探针的影响甚至大于预训练本身,而从同一表征中抹除被探针检测的质量方向,在一个场景中不造成任何损失,在另一个场景中却损失0.35的R²。一个基准能否对模型进行排序是一种经验属性,我们提出了三项检验来确立这一属性。
cs.RO / 95 / 2609.08280

Seeing is Not Believing: Breaking the Physical-to-Digital Trust Boundary in Robotics

眼见未必为实:打破机器人技术中物理世界到数字世界的信任边界
Shen, Leming, Geng, Shikai, Zheng, Yuanqing, Lu, Chris Xiaoxuan
Abstract
In multi-robot collaboration, task handovers rely on downstream verifiers performing remote attestation, which inspects sensor telemetry to ensure a robot's physical behavior strictly matches its assigned task. But can this telemetry be trusted? We show that it often cannot. In this paper, we uncover a severe vulnerability in Robot Operating System (ROS) 2: by modifying a single environment variable, an adversary can execute a pre-built hook to covertly intercept and inject both telemetry and control signals before they are published. Consequently, adversaries can hijack a robot to perform dangerous tasks while spoofing downstream verifiers with synthesized fake telemetry. Worse still, by exploiting the widespread reliance on third-party Docker containers and auxiliary tools, attackers can distribute compromised packages embedded with these malicious hooks to launch such attacks easily. On a physical Franka Emika robotic arm running Secure ROS 2, our attack injects fabricated telemetry in real time with only around 3 ms of jitter, preserving temporal synchronization and hardware integrity while achieving an 87% success rate even against an AI-based detector. We have responsibly disclosed these findings to the ROS 2 development team. We prepared a demo video available at https://youtu.be/ExeiGqUrnhQ.
Chinese Translation
在多机器人协作中,任务交接依赖于下游验证器执行远程证明(remote attestation),即检查传感器遥测数据以确保机器人的物理行为严格符合其分配的任务。但这种遥测数据是否可信?我们证明它往往不可信。在本文中,我们揭示了机器人操作系统(ROS)2中的一个严重漏洞:攻击者只需修改单个环境变量,即可执行一个预先构建的钩子(hook),在遥测数据和控制信号发布之前对其进行隐蔽的拦截与注入。由此,攻击者可以劫持机器人执行危险任务,同时利用伪造的合成遥测数据欺骗下游验证器。更糟糕的是,通过利用第三方Docker容器和辅助工具的广泛使用,攻击者可以分发嵌入这些恶意钩子的被篡改软件包,从而轻松发动此类攻击。在运行Secure ROS 2的物理Franka Emika机械臂上,我们的攻击能够实时注入伪造的遥测数据,抖动仅约3毫秒,在保持时间同步和硬件完整性的同时,即使面对基于AI的检测器也达到了87%的成功率。我们已负责任地向ROS 2开发团队披露了这些发现。演示视频见 https://youtu.be/ExeiGqUrnhQ。
cs.RO / 96 / 2609.08292

EvoNav-Bench: Benchmarking Lifelong Navigation in Evolving Environments

EvoNav-Bench:面向演化环境中终身导航的基准测试
Wang, Xilin, Zhang, Guoxi, Xu, Hongming, Zhang, Zhuofan, Wang, Tianxu, Fan, Lifeng
Abstract
Lifelong navigation (LN) requires an embodied agent to solve a sequence of navigation subtasks in the same environment. Since solving each subtask from scratch incurs redundant exploration, an LN agent must consolidate experience from earlier stages and reuse it in later stages, often through persistent scene representations such as scene graphs or visual snapshots. However, existing approaches typically assume a stationary environment, whereas in real-world LN settings, human activities can cause the environment to evolve. With the stationary assumption violated, existing methods may fuse outdated prior observations with new observations, yet current benchmarks cannot reveal this failure mode. In this paper, we present EvoNav-Bench, which extends the GOAT-Bench style LN formulation in the context of evolving environments. Built on the ProcTHOR framework, EvoNav-Bench introduces environment modifications between navigation tasks, making prior experience useful but not fully reliable. This design enables controlled evaluation of how environment evolution affects LN agents that reuse prior scene observations. Using EvoNav-Bench, we benchmark three recent methods that build and reuse scene representations for navigation. We also compare three simple heuristic strategies for handling environment evolution: Frontier-Update, Fail-then-Update, and Stage-Reset. Our results show that existing methods are brittle under environment evolution, while the heuristic strategies enable a controlled analysis of how agents can adapt to scene changes and mitigate their impact.
Chinese Translation
终身导航(Lifelong Navigation, LN)要求具身智能体在同一环境中解决一系列导航子任务。由于从零开始解决每个子任务会产生冗余探索,LN智能体必须将早期阶段积累的经验加以整合并在后续阶段复用,这通常借助持久化的场景表示来实现,例如场景图或视觉快照。然而,现有方法通常假设环境是静止的,而在现实世界的LN场景中,人类活动可能导致环境发生演化。当静止假设不成立时,现有方法可能会将过时的先验观测与新的观测融合,而当前的基准测试无法揭示这一失效模式。在本文中,我们提出了EvoNav-Bench,它在演化环境的背景下扩展了GOAT-Bench风格的LN任务设定。EvoNav-Bench构建于ProcTHOR框架之上,在导航任务之间引入环境修改,使得先验经验有用但并非完全可靠。这一设计使得我们能够对环境演化如何影响复用先前场景观测的LN智能体进行可控评估。利用EvoNav-Bench,我们对三种近期构建并复用场景表示的导航方法进行了基准测试。我们还比较了三种处理环境演化的简单启发式策略:Frontier-Update、Fail-then-Update和Stage-Reset。结果表明,现有方法在环境演化下表现脆弱,而启发式策略使我们能够可控地分析智能体如何适应场景变化并缓解其影响。
cs.RO / 97 / 2609.08338

A Multi-Modal Perception Pipeline for Object Detection and Tracking in Autonomous Racing

面向自主赛车中目标检测与跟踪的多模态感知流水线
Malvezzi, Davide, Pestarino, Michele, Cavicchioli, Vittoria, La Gamba, Valentina, Severi, Silvia, Bagni, Fabio, Bartoli, Luca, Bosi, Massimiliano, Gatti, Francesco, Verucchi, Micaela, Raji, Ayoub, Bertogna, Marko
Abstract
Object detection and tracking are fundamental components of perception systems for autonomous driving. Achieving robust performance under adverse conditions such as limited visibility, sensor noise, and failures remains an open challenge, particularly in autonomous racing, where vehicles operate at very high speeds, experience strong vibrations, and interact under small safety margins. This paper presents a multi-modal late-fusion perception pipeline for object detection and tracking in the autonomous racing domain. The proposed system extends previous work by exploiting all onboard sensors through a late-fusion approach and a dedicated multi-object tracking framework. Independent detections from cameras, LiDARs, and RADARs are combined to provide timely and robust state estimates of surrounding vehicles. The tracking method explicitly compensates for detection delays and embeds in its model prior knowledge of vehicle dynamics and track layout. Experimental evaluation on real-world data across diverse critical scenarios, representative of challenging edge cases also in urban driving, confirms the effectiveness of the proposed pipeline and its suitability to support safe and adaptive planning decisions.
Chinese Translation
目标检测与跟踪是自动驾驶感知系统的基本组成部分。在能见度受限、传感器噪声和故障等不利条件下实现鲁棒性能仍然是一个开放的挑战,尤其是在自主赛车领域,因为赛车以极高速度行驶、经受强烈振动,并且在极小的安全裕度下进行交互。本文提出了一种用于自主赛车领域目标检测与跟踪的多模态后期融合感知流水线。该系统是对先前工作的扩展,通过后期融合方法和专用的多目标跟踪框架来利用所有车载传感器。来自相机、激光雷达和毫米波雷达的独立检测结果被融合,以提供对周围车辆的及时且鲁棒的状态估计。该跟踪方法显式地补偿了检测延迟,并在其模型中嵌入了关于车辆动力学和赛道布局的先验知识。在涵盖多种关键场景(这些场景也代表了城市驾驶中具有挑战性的边缘情况)的真实数据上进行的实验评估,验证了所提流水线的有效性及其对支持安全且自适应的规划决策的适用性。
cs.RO / 98 / 2609.08339

RoboCousin: Build Your Own Simulation Playground for Robust Bimanual Robotic Manipulation

RoboCousin:构建您自己的仿真平台,实现鲁棒的双臂机器人操作
Zhu, Jingxuan, Li, Jingyi, Chen, LiangLiang, Jing, Zhiyuan, Zhang, Jidong, Li, Hongming
Abstract
Bimanual manipulation policies require large and diverse training datasets, yet collecting demonstrations on physical robots is expensive and difficult to scale. Simulation can generate data efficiently, but existing pipelines typically operate within closed asset libraries and predefined scenes: adding a newly observed object or environment still requires substantial effort to reconstruct geometry, specify physical and semantic properties, annotate interactions, and integrate the result into executable tasks. We present RoboCousin, an extensible simulation-based data-generation platform that turns user-provided observations into reusable assets, scenes, and expert trajectories for bimanual manipulation. Built on RoboTwin~2.0, RoboCousin converts object images into simulation-ready assets with visual and collision geometry, semantic and physical metadata, and automatically generated grasp-contact candidates. It further constructs digital cousins that vary compatible objects, backgrounds, layouts, and language instructions while preserving task-relevant affordances and spatial relations. The same asset system supports tabletop and room-level scene construction, with collision-aware base control for interaction beyond a fixed workspace. We release RoboCousin-OBD, containing more than 3,000 annotated object instances and 50 background environments, and use RoboCousin to generate over one million expert trajectories across 50 tasks. Simulation and real-robot experiments show that the automatically generated interaction annotations are comparable to curated annotations, generated assets provide effective sim-to-real supervision, and tabletop cousins can improve transfer beyond training on a single reconstructed scene. RoboCousin therefore provides a practical path for expanding both the scale and coverage of synthetic bimanual manipulation data.
Chinese Translation
双臂操作策略需要大规模且多样化的训练数据,然而在物理机器人上采集示范数据成本高昂且难以扩展。仿真可以高效地生成数据,但现有的流水线通常在封闭的资产库和预定义场景中运行:添加一个新观测到的物体或环境,仍然需要大量的工作来重建几何结构、指定物理和语义属性、标注交互,并将结果集成到可执行的任务中。我们提出了RoboCousin,这是一个可扩展的基于仿真的数据生成平台,能够将用户提供的观测数据转化为用于双臂操作的可复用资产、场景和专家轨迹。RoboCousin构建于RoboTwin~2.0之上,可将物体图像转换为仿真可用的资产,包括视觉与碰撞几何、语义与物理元数据,以及自动生成的抓取接触候选点。它进一步构建“数字近亲”(digital cousins),在保持任务相关的可供性(affordance)和空间关系的同时,对兼容的物体、背景、布局和语言指令进行变化。同一资产系统支持桌面级和房间级场景构建,并具备碰撞感知的底盘控制能力,可支持固定工作空间之外的交互。我们发布了包含超过3,000个标注物体实例和50个背景环境的RoboCousin-OBD数据集,并利用RoboCousin在50个任务中生成了超过一百万条专家轨迹。仿真和真实机器人实验表明,自动生成的交互标注与人工精选标注相当,生成的资产可提供有效的仿真到现实(sim-to-real)监督,而且桌面级“数字近亲”能够超越仅在单一重建场景上训练的迁移效果。因此,RoboCousin为扩大合成双臂操作数据的规模与覆盖范围提供了一条切实可行的路径。
cs.RO / 99 / 2609.08408

Localized Visual Feature Aggregation via Focus Pooling for Visuomotor Policies

基于聚焦池化的局部视觉特征聚合方法及其在视觉运动策略中的应用
Wang, Ruiyu, Zhuang, Zheyu, Kragic, Danica, Pokorny, Florian T.
Abstract
Focusing on spatially localized, control-relevant visual cues has been shown to improve data efficiency in visuomotor policies by reducing the need to model task-irrelevant visual variation. Existing methods often impose this focus through input preprocessing, such as cropping control- or object-centric regions in RGB images or point-clouds. However, it remains underexplored whether such localized features can be exposed directly from commonly used convolutional neural network (CNN) encoded features. In this paper, we show that intermediate CNN features preserve localized visual context for control, but existing pooling methods fail to aggregate it effectively. We introduce FocusPool, an attention pooling module that selectively aggregates intermediate visual features according to their relevance to the robot's current proprioceptive context. The resulting pooled representation captures task-progressive, control-relevant local information and is used directly for policy learning. Across simulation and real-world experiments, FocusPool improves policy success rates over pooling and explicit local focus methods by 36.2% and 41.2%, with training only 5.8% of encoder parameters.
Chinese Translation
已有研究表明,聚焦于空间上局部化的、与控制相关的视觉线索可以减少对任务无关视觉变化建模的需求,从而提高视觉运动策略的数据效率。现有方法通常通过输入预处理来施加这种聚焦,例如在RGB图像或点云中裁剪以控制或物体为中心的区域。然而,这类局部化特征能否直接从常用的卷积神经网络(CNN)编码特征中暴露出来,仍缺乏充分探索。在本文中,我们证明了CNN的中间特征保留了用于控制的局部化视觉上下文,但现有的池化方法无法有效地将其聚合。我们提出了FocusPool,一种注意力池化模块,它根据与机器人当前本体感觉上下文的相关性,选择性地聚合中间视觉特征。所得的池化表示能够捕捉任务推进过程中与控制相关的局部信息,并直接用于策略学习。在仿真和真实世界实验中,FocusPool相较于池化方法和显式局部聚焦方法分别将策略成功率提升了36.2%和41.2%,且仅需训练编码器参数的5.8%。
cs.RO / 100 / 2609.08409

Coverage Path Planning for Redundant Manipulators using Generalized Spanning Trees

基于广义生成树的冗余机械臂覆盖路径规划
Kopo, Raksi, Kyriakopoulos, Kostas J.
Abstract
Surface coverage with task-redundant manipulators is challenging because each surface point may admit multiple inverse kinematics (IK) solutions, and configuration choices strongly affect motion quality. This paper extends the classical Spanning Tree Coverage (STC) method to redundant manipulators through offline and online Joint Spanning Tree Coverage (JSTC) algorithms. Offline JSTC samples multiple Inverse Kinematics (IK) solutions per grid cell and formulates the problem as a Generalized Minimum Spanning Tree (GMST), selecting one configuration per cell and tracing the resulting tree to obtain a non-revisiting coverage path. Online JSTC incrementally expands and backtracks a spanning tree with feasibility and cost evaluation while handling dynamic grid updates. Simulation results show that offline JSTC reduces computation time, reconfigurations, and joint motion compared to other methods, while online JSTC achieves fast per-step planning in dynamic scenarios.
Chinese Translation
具有任务冗余的机械臂进行表面覆盖是具有挑战性的,因为每个表面点可能存在多个逆运动学(IK)解,且构型选择会显著影响运动质量。本文通过离线和在线关节生成树覆盖(Joint Spanning Tree Coverage,JSTC)算法,将经典的生成树覆盖(Spanning Tree Coverage,STC)方法扩展至冗余机械臂。离线JSTC为每个网格单元采样多个逆运动学(IK)解,并将问题建模为广义最小生成树(Generalized Minimum Spanning Tree,GMST),为每个单元选择一个构型,并通过遍历所得到的树获得无重访的覆盖路径。在线JSTC在处理动态网格更新的同时,结合可行性与代价评估,增量地扩展和回溯生成树。仿真结果表明,与其他方法相比,离线JSTC减少了计算时间、构型变换次数和关节运动量,而在线JSTC能够在动态场景下实现快速的逐步规划。
cs.RO / 101 / 2609.08444

Safe Task Planning with Long-Term Graph Memory for Embodied Agents

基于长期图记忆的具身智能体安全任务规划
Li, Siyuan, Lang, Taiyan, Yan, Aoqi, Yu, Jia, Liu, Feifan, Du, Yihan, Zheng, Yu, Wang, Xun, Liu, Peng
Abstract
Large language models (LLMs) and vision-language models (VLMs) have significantly advanced zero-shot task planning for embodied agents. However, most LLM- and VLM-driven methods struggle to generate safe high-level actions due to a lack of physical risk awareness, particularly under partial observability, where hazards lie outside the immediate field of view. To address this challenge, we propose a novel safe task-planning framework, SafeMem, which constructs and maintains a long-term semantic graph memory of the open and dynamic environment. Based on egocentric observations, the proposed framework incrementally accumulates knowledge about surrounding objects and their relationships with a graph. Then, an LLM-based risk predictor evaluates candidate actions using the graph memory, triggering a conservatism-modulated replanning loop with explanations for detected hazards. Extensive experiments on the IS-Bench benchmark and a real-world robot platform demonstrate that the SafeMem framework substantially improves safe success rates compared to state-of-the-art VLM-driven task planners. Video results are available on our webpage: https://sites.google.com/view/safemem.
Chinese Translation
大语言模型(LLM)和视觉-语言模型(VLM)极大地推动了具身智能体的零样本任务规划。然而,由于缺乏物理风险意识,大多数由LLM和VLM驱动的方法难以生成安全的高层动作,尤其是在部分可观测的环境下,危险可能位于智能体直接视野之外。为应对这一挑战,我们提出了一种新颖的安全任务规划框架SafeMem,该框架针对开放且动态的环境构建并维护一个长期语义图记忆。基于自我中心的观测,该框架以图的形式增量地积累关于周围物体及其相互关系的知识。随后,基于LLM的风险预测器利用图记忆评估候选动作,并在检测到危险时触发一个带有解释的、由保守度调节的重新规划循环。在IS-Bench基准和真实机器人平台上的大量实验表明,与最先进的VLM驱动的任务规划器相比,SafeMem框架显著提升了安全成功率。视频结果可在我们的网页上查看:https://sites.google.com/view/safemem。
cs.RO / 102 / 2609.08490

Multi-bounce Drum Roll with Optimized Active Tricks to Leverage Soft Embodiment

利用优化的主动技巧发挥软体本体性能的多弹跳鼓滚奏
Yamanaka, Naoto, Jin, Takanori, Kobayashi, Taisuke
Abstract
This paper presents a soft robotic drummer for accurate and efficient drum rolls. High-frequency drum rolls require the "multi-bounce technique," where a drumstick bounces multiple times with a single stroke. In robotic reproduction of this technique, the body's elasticity is key, while the fine motion during the stroke is also crucial for maximizing the potential of that elasticity. Therefore, we design two tricks: i) Tap-Pull (TP) trick to increase the number of rebounds by adding a pulling motion after impact; and ii) Micro-Pulse (MP) trick to keep the drumming volume by injecting small oscillations during the stroke. Due to the nonlinear complexity of soft embodiment, both tricks are efficiently tuned using Bayesian optimization in a data-driven manner for accomplishing the respective objectives quantified. We evaluated the optimized behaviors with soft and rigid end-effectors. As a result, the soft TP achieved the highest bounce count (12.25 per stroke) with uniform intervals. The soft MP suppressed the volume decay, yielding 6.8-times higher acoustic efficiency compared to the rigid MP. These results indicate that the proposed tricks with the combination of elasticity and optimization can make robots play excellent drum rolls.
Chinese Translation
本文提出了一种软体机器人鼓手,用于实现精确且高效的鼓滚奏(drum roll)。高频鼓滚奏需要“多弹跳技术”,即鼓槌在单次敲击中多次弹跳。在机器人复现该技术时,本体的弹性是关键,而敲击过程中的精细运动对于最大化弹性潜力同样至关重要。因此,我们设计了两种技巧:i)敲击-拉动(Tap-Pull, TP)技巧,通过在撞击后施加拉动动作来增加弹跳次数;ii)微脉冲(Micro-Pulse, MP)技巧,通过在敲击过程中注入微小振荡来维持敲击音量。鉴于软体本体的非线性复杂性,这两种技巧均采用贝叶斯优化以数据驱动的方式高效调优,从而实现各自量化目标的达成。我们使用软性和刚性末端执行器评估了优化后的行为。结果表明,软体TP技巧实现了最高的弹跳次数(每次敲击12.25次)且弹跳间隔均匀。软体MP技巧抑制了音量衰减,其声学效率比刚性MP技巧高出6.8倍。这些结果表明,所提出的技巧结合弹性与优化方法,能够使机器人演奏出出色的鼓滚奏。
cs.RO / 103 / 2609.08493

AURORA: Active Uncertainty-Driven Re-Orientation for In-Hand Reconstruction

AURORA:面向手内物体重建的主动不确定性驱动重定向方法
Zhao, Feiyu, Li, Yuetong, Xiao, Chenxi
Abstract
Observing objects grasped by a robot hand is challenging due to severe visual occlusions. Although in-hand manipulation can expose hidden surfaces, existing approaches often rely on predefined or open-loop reorientation strategies that do not explicitly target under-observed regions. We propose AURORA, an active 3D reconstruction framework that closes the loop between online object-centric reconstruction and in-hand reorientation. At its core, Ray-GPIS estimates direction-wise reconstruction uncertainty along candidate viewing rays and selects next-best-view targets using an uncertainty--novelty objective, which are realized through an axis-conditioned in-hand rotation policy. The resulting RGB-D observations are fused incrementally using CAD-free 6D pose tracking and lightweight geometric reconstruction. Experiments demonstrate that AURORA improves reconstruction quality and information-acquisition efficiency over non-active rotation strategies, while Ray-GPIS also outperforms active view-planning baselines in reconstruction performance, action-ranking quality, and planning efficiency. Targeted ablations further validate its robustness to hand occlusion and pose errors. The project webpage is available at https://aurorahand.github.io/
Chinese Translation
由于严重的视觉遮挡,观察机器人手所抓取的物体极具挑战性。尽管手内操作可以暴露被遮挡的表面,但现有方法通常依赖预定义或开环的重定向策略,未能明确针对观测不足的区域。我们提出了AURORA,一个主动三维重建框架,它将在线以物体为中心的重建与手内重定向形成闭环。该框架的核心是Ray-GPIS,它沿候选视线估计方向性的重建不确定性,并通过不确定性—新颖性目标函数选择下一个最佳视点目标,再通过轴条件化的手内旋转策略予以实现。所得的RGB-D观测通过免CAD的6D位姿跟踪和轻量级几何重建进行增量融合。实验表明,与非主动旋转策略相比,AURORA提升了重建质量和信息获取效率;同时,Ray-GPIS在重建性能、动作排序质量和规划效率方面也优于主动视点规划基线方法。针对性的消融实验进一步验证了其对遮挡和位姿误差的鲁棒性。项目网页见 https://aurorahand.github.io/
cs.RO / 104 / 2609.08511

PGMT: Perceptive General Motion Tracking for Humanoid Robots

PGMT:面向人形机器人的感知式通用运动跟踪
Li, Hongyi, Peizhuo, Li, Tao, Yucheng, Wang, Ze, Xu, Fangzhou, Chen, Jinyi, Yuan, Yanyan, Jia, Dapeng, Jin, Yongbin, Fan, Mingfeng, Sartoretti, Guillaume, Wang, Hongtao
Abstract
Humanoid motion trackers can reproduce diverse whole-body motions, but their performance degrades on complex terrain where terrain-agnostic references become physically infeasible. We present PGMT, a Perceptive General Motion Tracking pipeline for humanoid robots that learns terrain adaptation from independently selected motion references and terrains. PGMT first learns a general tracking and recovery prior, then incorporates terrain perception through motion-conditioned terrain glimpses that selectively encode regions relevant to the current motion. Terrain-aware tracking relaxation allows necessary deviations from the reference while preserving its motion intent. Zero-shot deployment on a Unitree G1 demonstrates robust terrain-adaptive locomotion and whole-body motion execution over real-world terrain with obstacles up to 37 cm high, while supporting teleoperation, dynamic motion tracking, and fall recovery. PGMT extends general humanoid motion tracking beyond flat ground, providing a unified policy for terrain-adaptive locomotion, diverse whole-body behaviors, and teleoperation in complex environments.
Chinese Translation
人形机器人运动跟踪器能够复现多样的全身动作,但在复杂地形上,其性能会因与地形无关的参考动作在物理上不可行而下降。我们提出了PGMT,一种面向人形机器人的感知式通用运动跟踪流水线,它从独立选取的运动参考和地形中学习地形适应能力。PGMT首先学习一个通用的跟踪与恢复先验,然后通过运动条件下的地形注视机制引入地形感知,该机制选择性地编码与当前运动相关的地形区域。地形感知的跟踪松弛机制允许在保留运动意图的前提下对参考动作进行必要的偏离。在Unitree G1上的零样本部署展示了稳健的地形自适应运动和全身动作执行能力,可应对高达37厘米高的真实世界障碍地形,同时支持遥操作、动态运动跟踪和跌倒恢复。PGMT将通用人形机器人运动跟踪拓展至平坦地面之外,提供了一个统一的策略,实现地形自适应运动、多样的全身行为以及复杂环境下的遥操作。
cs.RO / 105 / 2609.08512

TASG-Explore: Traversability-Aware Sector-Guided Exploration for Ground Robot on Uneven Terrain

TASG-Explore:面向非平整地形地面机器人的可通行性感知扇区引导探索
Wang, Shaocong, Shao, Shiliang, Wang, Ting, Han, Guangjie, Liu, Lianqing
Abstract
Autonomous exploration on uneven terrain requires ground robots to balance exploration efficiency, coverage completeness, and terrain safety. Detailed tsrrain reasoning improves local reliability but can slow large-scale exploration, whereas coarse region guidance expands quickly in open areas but can miss narrow passages and irregular traversable boundaries. To address this challenge, this paper presents TASG-Explore, a traversability-aware sector-guided exploration framework for ground robots. The framework first performs hierarchical traversability analysis using variable-voxel ground fitting and adaptive 8-bit obstacle encoding. It then splitting cost map into sectors, incrementally updates sector clusters, extracts terrain-coupled frontier viewpoints, and maintains a dynamic topological roadmap with unknown topological hypotheses. Finally, a sector-guided planner selects region targets and inserts local viewpoints to generate efficient exploration routes. Benchmark experiments in diverse challenging environments, including caves, forests, and rugged hills, show that TASG-Explore achieves the best overall performance among six representative state-of-the-art planners. The proposed traversability analysis improves processing efficiency by 6.3 times while maintaining high accuracy, and the exploration planner improves exploration efficiency by 51% and increases coverage by up to 2.95 times in rugged hill scene. Large-scale real-world experiments further demonstrate the practical value of the proposed method.
Chinese Translation
非平整地形上的自主探索要求地面机器人在探索效率、覆盖完整性和地形安全性之间取得平衡。精细的地形推理可提升局部可靠性,但会降低大规模探索的速度;而粗粒度的区域引导虽然能在开阔区域快速扩展,却可能遗漏狭窄通道和不规则的可通行边界。为应对这一挑战,本文提出了TASG-Explore,一种面向地面机器人的可通行性感知扇区引导探索框架。该框架首先采用可变体素地面拟合和自适应8位障碍物编码进行层次化可通行性分析;随后将代价地图划分为扇区,增量更新扇区聚类,提取与地形耦合的前沿视点,并维护包含未知拓扑假设的动态拓扑路网;最后,扇区引导的规划器选择区域目标并插入局部视点,以生成高效的探索路径。在包括洞穴、森林和崎岖山地等多种挑战性环境中的基准实验表明,TASG-Explore在六种具有代表性的最先进规划器中取得了最佳综合性能。所提出的可通行性分析方法在保持高精度的同时,将处理效率提升了6.3倍;在崎岖山地场景中,探索规划器将探索效率提升了51%,覆盖率最高提升2.95倍。大规模真实世界实验进一步验证了所提方法的实用价值。
cs.RO / 106 / 2609.08583

Estimating Semantic Ambiguity via Gaussian Context Distributions for VLM-Driven Traversability Analysis

基于高斯上下文分布的语义模糊性估计用于视觉语言模型驱动的可通行性分析
Häuselmann, Ramona, Saucedo, Mario A. V., Kanellakis, Christoforos, Nikolakopoulos, George
Abstract
Autonomous navigation in unstructured environments requires robust scene understanding, yet Vision-Language Models (VLMs) often suffer from semantic ambiguity, where conflicting predictions can lead to dangerous failures. To address this, we present a novel pipeline for vision-based traversability estimation that explicitly models contextual uncertainty. Our approach utilizes Conceptual Anchoring to ground open-vocabulary VLM predictions onto a continuous physical traversability scale. By formulating the model's responses as a Gaussian Context Distribution (GCD), we derive both a dense traversability map and a dense uncertainty map based on the statistical properties of the distribution. Experimental validation on the real-world GOOSE dataset demonstrates that our proposed uncertainty metric effectively correlates with sources of ambiguity, such as visual artifacts and mixed terrain overlap. The method exhibits competitive performance while offering the distinct advantage of providing statistical uncertainty estimates to address semantic ambiguity, enabling safer and more reliable autonomous behavior in complex outdoor settings.
Chinese Translation
非结构化环境中的自主导航需要鲁棒的场景理解能力,然而视觉语言模型(Vision-Language Models, VLMs)常常受到语义模糊性的困扰,冲突的预测可能导致危险故障。为解决这一问题,我们提出了一种新颖的基于视觉的可通行性估计流程,该流程显式地建模上下文不确定性。我们的方法利用概念锚定(Conceptual Anchoring)将开放词汇的VLM预测映射到连续的物理可通行性尺度上。通过将模型的响应建模为高斯上下文分布(Gaussian Context Distribution, GCD),我们基于该分布的统计特性推导出稠密可通行性图和稠密不确定性图。在真实世界GOOSE数据集上的实验验证表明,我们提出的不确定性度量与视觉伪影和混合地形重叠等模糊性来源有效地相关。该方法在展现出竞争力的性能的同时,还具有提供统计不确定性估计以应对语义模糊性的独特优势,从而在复杂的户外环境中实现更安全、更可靠的自主行为。
cs.RO / 107 / 2609.08626

MFVINS: Multiple Fisheye Camera-Based Visual Inertial System

MFVINS:基于多鱼眼相机的视觉惯性系统
Jang, Eunseong, Chung, YuJin, Lee, Sang Jun, Yoon, Jihyun, Jo, HyungGi
Abstract
A simultaneous localization and mapping (SLAM) method using a monocular camera and a low-cost inertial measurement unit (IMU) sensor is an effective way to fulfill a low-cost sensor configuration. Using this sensor configuration, visual-inertial system (VINS) focuses on fusing data from a camera and an IMU sensor to estimate the six degrees-of-freedom (DOF) of the sensor pose. Typically, VINS uses only a single camera as visual input, which lead to problems such as error accumulation due to occlusion, various illumination, and textureless environments. In this paper, we propose a new multiple fisheye camera-based visual-inertial system called MFVINS. We present an IMU-aided FAST feature tracker for multiple cameras that enables efficient extraction and robust matching of local features. Then, the proposed method filters out outliers caused by fisheye distortion on the normalized image plane. Subsequently, a new reprojection error with physical validity constraints is proposed for bundle adjustment using learning-based depth estimation. The proposed method is applied to various scenarios, and its effectiveness is demonstrated by comparing previous VINS methods. In particular, MFVINS is implemented in real-time process to leverage the advantages of using multiple cameras -- robustness against occlusion and textureless regions -- while reducing the computational burden.
Chinese Translation
使用单目相机和低成本惯性测量单元(IMU)传感器的同步定位与建图(SLAM)方法是实现低成本传感器配置的有效途径。在这种传感器配置下,视觉惯性系统(VINS)致力于融合相机与IMU传感器的数据,以估计传感器位姿的六个自由度(DOF)。通常,VINS仅使用单个相机作为视觉输入,这会导致遮挡、光照变化以及无纹理环境下的误差累积等问题。本文提出了一种新的基于多鱼眼相机的视觉惯性系统,称为MFVINS。我们提出了一种面向多相机的IMU辅助FAST特征跟踪器,能够实现高效的特征提取和鲁棒的特征匹配。随后,该方法在归一化像平面上滤除由鱼眼畸变引起的外点。接着,提出了一种具有物理有效性约束的新重投影误差,用于基于学习的深度估计的束调整(bundle adjustment)。该方法被应用于多种场景,并通过与现有VINS方法的比较验证了其有效性。特别地,MFVINS以实时方式实现,充分利用了多相机的优势——对遮挡和无纹理区域的鲁棒性——同时降低了计算负担。
cs.RO / 108 / 2609.08638

CASD: Chunk-Aligned Semantic Distillation for Multi-StageRobot Manipulation

CASD:面向多阶段机器人操作的块对齐语义蒸馏方法
Ding, Tinghe, Li, Jiahao, Wang, He
Abstract
An action chunk can span several stages of a manipulation task, yet a label for its first step describes only the current stage. We introduce Chunk-Aligned Semantic Distillation (CASD), which derives semantic targets for entire action chunks. An offline vision--language model segments demonstrations into described stages. Their occupancy within each action chunk determines a weighted semantic target, including transitions between stages. A CASD generator learns to predict this target from the current observation, robot state, and task instruction. We then freeze the generator and train a policy conditioned on its predictions. The semantic branch runs once per policy query, without online VLM calls or reasoning-trace decoding. Teacher matching on annotated LIBERO training episodes is above chance for both single-stage and boundary-crossing chunks. We evaluate three Fast-WAM variants and a DreamZero integration across four benchmarks, including distribution shifts on LIBERO-Plus. Compared with published references, IDM+CASD reaches 98.9\% versus 98.0\% average success on LIBERO, while Uncond falls below its reference. Joint+CASD reaches 93.0\% versus 90.6\% on RoboTwin 2.0, and DreamZero+CASD reaches a 47.9\% four-category MolmoSpaces manipulation average versus 40.7\%. Performance varies across backbone integrations.
Chinese Translation
一个动作块(action chunk)可能跨越操作任务的多个阶段,而其首步的标签仅描述当前阶段。我们提出了块对齐语义蒸馏(Chunk-Aligned Semantic Distillation, CASD),为整个动作块推导语义目标。一个离线视觉-语言模型将演示分割为若干被描述的阶段。各阶段在每个动作块中的占比决定了一个加权语义目标,其中包括阶段间的过渡。CASD生成器学习根据当前观测、机器人状态和任务指令来预测该目标。随后我们冻结生成器,并以生成器的预测作为条件训练策略。语义分支在每次策略查询时仅运行一次,无需在线调用视觉-语言模型或推理轨迹解码。在标注的LIBERO训练回合上,教师匹配在单阶段块和跨阶段边界块上均显著高于随机水平。我们在四个基准上评估了三种Fast-WAM变体以及一个DreamZero集成版本,包括LIBERO-Plus上的分布偏移测试。与已发表的参考结果相比,IDM+CASD在LIBERO上达到98.9%的平均成功率(参考方法为98.0%),而Uncond则低于其参考结果。Joint+CASD在RoboTwin 2.0上达到93.0%(参考方法为90.6%),DreamZero+CASD在MolmoSpaces操作任务的四类别平均成绩上达到47.9%(参考方法为40.7%)。性能在不同骨干网络集成方式之间存在差异。
cs.RO / 109 / 2609.08669

Learning to build covering structures with continuous adjustments

学习通过连续调整构建覆盖结构
Vallat, Gabriel, Kamgarpour, Maryam, Parascho, Stefana
Abstract
Robotic construction offers the potential to use materials more efficiently and create complex geometries, but current methods rely on rigid, high-precision plans that cannot accommodate the tolerances, inaccuracies, and unexpected changes inherent in physical fabrication. In this work, we introduce a reinforcement learning approach that forgoes predefined plans entirely, instead generating construction sequences adaptively as the structure is built. Our method operates on graph-structured state representations and a mixed (parameterized) action space, requiring both discrete block selection and continuous placement parameters. Because the stability simulation of a structure is computationally heavy, we develop an efficient exploration strategy by incorporating unilateral edges into graph neural networks, extending soft actor-critic (SAC) to this hybrid setting. We evaluate our algorithm, HSAC, against the prior method hybrid-PPO (HPPO), demonstrating significantly higher asymptotic performance and good sample efficiency. We also demonstrate HSAC's robustness to hyperparameter choices and its exploration capability, handling up to 10 discrete actions without performance degradation. Finally, we validate our approach on a physical two-robot setup, successfully building a spanning arch with 3D-printed blocks in closed-loop execution, confirming that policies trained in simulation transfer to real hardware.
Chinese Translation
机器人建造有潜力更高效地利用材料并创建复杂几何形状,但现有方法依赖于刚性的、高精度的规划,无法适应物理制造过程中固有的公差、误差和意外变化。在本工作中,我们提出了一种强化学习方法,完全摒弃预定义规划,而是在结构建造过程中自适应地生成建造序列。我们的方法基于图结构的状态表示和混合(参数化)动作空间,既需要离散的块体选择,也需要连续的放置参数。由于结构的稳定性仿真计算代价高昂,我们通过在图神经网络中引入单向边,并将软演员-评论家算法(SAC)扩展到这一混合设置中,开发了一种高效的探索策略。我们将我们的算法 HSAC 与先前的方法 hybrid-PPO(HPPO)进行对比评估,结果表明 HSAC 具有显著更高的渐近性能和良好的样本效率。我们还证明了 HSAC 对超参数选择的鲁棒性及其探索能力,能够处理多达 10 个离散动作而不出现性能下降。最后,我们在一个双机器人物理平台上验证了我们的方法,通过闭环执行成功用 3D 打印块体建造了一个拱形跨度结构,证实了在仿真中训练的策略可以迁移到真实硬件上。
cs.RO / 110 / 2609.08673

BIFTA: Brain-Inspired Few-Shot Tactile Adaptation for Unknown Sensors

BIFTA:面向未知传感器的脑启发少样本触觉自适应方法
Liu, Boheng, Li, Ziyu, Wu, Xia
Abstract
Advances in tactile sensing have made contact-rich perception possible, accelerating progress in robotic manipulation, material understanding, and embodied interaction. However, because optical design, elastomer mechanics, and imaging geometry differ substantially across tactile sensors, models trained on known sensor types can suffer an abrupt performance collapse on unknown sensors. To address this problem, we propose the Brain-Inspired Few-Shot Tactile Adaptation (BIFTA) framework; it draws on the brain's rapid sensory adaptation mechanism to adapt a frozen encoder to an unknown tactile sensor from a small labeled support set. BIFTA preserves pretrained representations through dual-view statistical memory, constructs support-conditioned spectral graphs to repair sensor-dependent feature neighborhoods, and applies uncertainty-gated recurrent propagation to strengthen reliable cross-query evidence. Extensive benchmarks across three tactile datasets show that BIFTA substantially improves adaptation to unknown sensors: with only 10\% labeled target data on SITR, it raises mean Sparsh accuracy from 6.86\% for the frozen source classifier to 87.09\%, exceeding the strongest implemented prior comparison by 47.22 percentage points, and these gains generalize across datasets, pretrained backbones, and tactile tasks. These results validate BIFTA for data-efficient adaptation to unknown tactile sensors and offer a promising route toward tactile models that transfer across heterogeneous hardware.
Chinese Translation
触觉感知技术的进步使富接触感知成为可能,加速了机器人操作、材料理解与具身交互的发展。然而,由于不同触觉传感器在光学设计、弹性体力学和成像几何方面存在显著差异,在已知传感器类型上训练的模型在未知传感器上可能出现性能骤降。为解决这一问题,我们提出了脑启发少样本触觉自适应(Brain-Inspired Few-Shot Tactile Adaptation, BIFTA)框架;该框架借鉴大脑快速感觉适应机制,利用少量带标签的支持集将冻结的编码器适配到未知触觉传感器。BIFTA 通过双视角统计记忆保留预训练表征,构建以支持集为条件的谱图以修复依赖传感器的特征邻域,并采用不确定性门控的循环传播来强化可靠的跨查询证据。在三个触觉数据集上的大量基准实验表明,BIFTA 显著提升了对未知传感器的适应能力:在 SITR 上仅使用 10% 的带标签目标数据时,它将冻结源分类器 6.86% 的平均 Sparsh 准确率提升至 87.09%,超出已实现的最强先前对比方法 47.22 个百分点,且这些增益在不同数据集、预训练骨干网络和触觉任务上均具有泛化性。这些结果验证了 BIFTA 在面向未知触觉传感器数据高效适应方面的有效性,并为构建可在异构硬件间迁移的触觉模型提供了有前景的路径。
cs.RO / 111 / 2609.08678

HiBRIDGE: A Hierarchical Bayesian Neural Network Framework for Interpretable Dialogue Management in Group-Robot Interaction

HiBRIDGE:一个用于群体机器人交互中可解释对话管理的层次贝叶斯神经网络框架
Nigro, Massimiliano, Gunes, Hatice, Spitale, Micol, Dogan, Fethiye Irmak
Abstract
In multi-party human-robot interaction, a robot must continuously decide whom to address and what to say to participate effectively in the conversation. In real-world interactions, this is challenging because several behaviours may be plausible at the same time: a robot might continue a topic with one participant, involve another through a question, or address the whole group, with the appropriate choice depending on both whom it addresses and the interaction context. Current approaches remain limited in representing uncertainty when several behaviours are plausible and in structuring decisions into semantically meaningful intermediate steps that make robot decisions easier to interpret. Addressing these, we present HiBRIDGE, a hierarchical Bayesian neural network framework for group-robot dialogue management. Its Bayesian formulation enables uncertainty-aware prediction and robust learning from limited interaction data, while the hierarchical approach formulates behaviour selection as a structured, multi-stage decision process. We further use decision-tree surrogates to investigate whether this structure can support more interpretable explanations. Across three offline group-HRI datasets, our findings show that Bayesian formulations outperform their deterministic counterparts and several state-of-the-art baselines. Next, through an online study (N=20), we show that explanations derived from the hierarchical model are rated as more helpful for understanding robot behaviour and are preferred over those derived from the flat model. Finally, through our in-person study (N=12), we demonstrate the feasibility of HiBRIDGE for autonomous real-time group interaction, with both hierarchical and flat Bayesian variants positively perceived. Overall, HiBRIDGE combines strong predictive performance with a structured decision process that supports more interpretable explanations of robot behaviour.
Chinese Translation
在多主体人机交互中,机器人必须持续决定与谁交流以及说什么,才能有效地参与对话。在真实世界的交互中,这极具挑战性,因为多种行为可能同时都是合理的:机器人可以与某位参与者继续当前话题,也可以通过提问让另一位参与者加入,或者面向整个群体发言,而合适的选择取决于其交流对象以及交互情境。现有方法在表示多种行为均合理时的不确定性方面仍然有限,且难以将决策组织为语义上有意义的中间步骤,从而使机器人决策更易于解释。针对这些问题,我们提出了HiBRIDGE,一个用于群体机器人对话管理的层次贝叶斯神经网络框架。其贝叶斯建模方式能够实现具有不确定性感知的预测,并支持在有限交互数据下的鲁棒学习;同时,层次化方法将行为选择构建为一个结构化的多阶段决策过程。我们进一步使用决策树代理模型来探究这种结构是否能够支持更可解释的解释。在三个离线群体人机交互(HRI)数据集上的结果表明,贝叶斯建模优于其确定性对应方法以及多个最先进的基线方法。随后,通过一项在线研究(N=20),我们表明由层次化模型导出的解释在帮助理解机器人行为方面被评为更有用,且比扁平模型导出的解释更受青睐。最后,通过面对面研究(N=12),我们证明了HiBRIDGE在自主实时群体交互中的可行性,层次化与扁平两种贝叶斯变体均获得了积极的评价。总体而言,HiBRIDGE兼具强大的预测性能和结构化的决策过程,能够为机器人行为提供更可解释的解释。
cs.RO / 112 / 2609.08711

DCLP++: Learning to Navigate with Footprint Clearance and Relative Motion

DCLP++:基于足迹间隙与相对运动的导航学习
Wang, Shanze, Zhang, Wei
Abstract
We present DCLP++, a local navigation frameworkthat uses footprint clearance as the geometric basis for studying relative motion features in dynamic environments. Each valid LiDAR return is mapped to its shortest Euclidean distance from the filled robot footprint before reciprocal encoding, replacing distance from the sensor with distance to the occupied body. Radial measurementsor simulated planar relative velocities provide short-horizon features without static-dynamic labels in the policy input. A preliminary study uses a rectangular robot with a speed limit of 1 m/s among 20 moving obstacles. On 100 fixed validation tasks, two selected training seeds yield mean success rates of 42% with sensor rangeand 70% with footprint clearance after 200,000 environment steps.Motion variants show mixed additional gains. These results supportthe clearance-based observation in the evaluated setting; reliable motion benefits and transfer across robots require further evaluation.
Chinese Translation
我们提出了DCLP++,一个以足迹间隙作为几何基础来研究动态环境中相对运动特征的局部导航框架。在互惠编码之前,每个有效的激光雷达(LiDAR)返回点被映射为其到填充机器人足迹的最短欧氏距离,用到占据本体的距离取代距传感器的距离。径向测量或模拟的平面相对速度在策略输入中提供短时程特征,而无需静态-动态标签。初步研究采用速度限制为1 m/s的矩形机器人在20个移动障碍物中进行实验。在100个固定验证任务上,经过200,000个环境步骤后,两个选定的训练种子分别取得了基于传感器距离观测42%和基于足迹间隙观测70%的平均成功率。运动特征变体显示出参差不齐的额外增益。这些结果支持在所评估环境中采用基于间隙的观测;可靠的运动特征收益以及跨机器人的迁移仍需进一步评估。
cs.RO / 113 / 2609.08743

FOCI Policy: Focus on Object-Centric Interactions for Relational Manipulation Policies

FOCI策略:面向关系型操作策略以物体为中心的交互聚焦方法
Fu, Ze, Song, Pinhao, Hu, Yutong, Detry, Renaud
Abstract
Object-centric manipulation policies improve generalization by modeling object motion instead of directly predicting robot actions. However, existing methods are often limited by representations which are either too simplistic to capture interaction dynamics or too dense to learn efficiently. We observe that many rigid relational manipulation tasks are governed by short interaction phases where the relative motion between task-relevant objects is tightly constrained. Based on this observation, we propose \textsc{Foci Policy}, an interaction-centric framework that achieves a two-fold abstraction: (1) temporally, by automatically extracting compact interaction segments from demonstrations;(2) spatially, by representing skills as relative $SE(3)$ motion between task-relevant objects, yielding invariance to scene configurations and robot embodiment. Experiments on RLBench, COLOSSEUM, and real-world tasks show that \textsc{Foci Policy} achieves strong performance with substantially less training data than prior object-centric and action-centric policies. These results suggest that modeling object-object interactions provides a simple and efficient inductive bias for rigid relational manipulation. Project page: \href{https://fitz0401.github.io/foci-page/}{fitz0401.github.io/foci-page/}.
Chinese Translation
以物体为中心的操作策略通过建模物体运动而非直接预测机器人动作来提升泛化能力。然而,现有方法往往受限于其表示方式:要么过于简单而无法捕捉交互动力学,要么过于稠密而难以高效学习。我们观察到,许多刚性关系型操作任务由短暂的交互阶段主导,在这些阶段中,任务相关物体之间的相对运动受到严格约束。基于这一观察,我们提出了FOCI策略,这是一个以交互为中心的框架,实现了两个层面的抽象:(1)时间维度上,通过从演示中自动提取紧凑的交互片段;(2)空间维度上,将技能表示为任务相关物体之间的相对$SE(3)$运动,从而对场景配置和机器人本体具有不变性。在RLBench、COLOSSEUM以及真实世界任务上的实验表明,FOCI策略相比以往以物体为中心和以动作为中心的策略,能够以显著更少的训练数据取得优异性能。这些结果表明,建模物体间交互为刚性关系型操作提供了一种简单而高效的归纳偏置。项目页面:https://fitz0401.github.io/foci-page/。
cs.RO / 114 / 2609.08770

A Controlled Comparison of Manual and Teleoperated Intraocular Instrument Motion for an Input Device

一种输入设备下手动与遥操作眼内器械运动的受控比较
Hoxha, Korab, Imamovic, Mirza, Henriques, Angelo, Nasseri, M. Ali
Abstract
Input devices for robotic microsurgery are frequently described as preserving the surgeon's trained technique, but the claim is rarely measured. We compared manual and teleoperated intraocular instrument motion with the trocar constraint, the instrument, the eye model and the tracking source common to both conditions, so that the control interface was the only factor varied. Prior comparisons cannot hold the instrument fixed, because a robotic instrument is not the tool used manually. Sixteen participants performed a navigation task on a commercial ophthalmic simulator by hand and through a three-degree-of-freedom input device commanding a five-joint robot. Task outcome was equal but at ceiling: every participant acquired all five targets under both interfaces with no retinal or lens injury. Execution differed on every measure. Teleoperated trials took three times as long at a quarter of the median speed, covered less than half the angular working range, and were broken into 3.5 times as many separate movements. Completion time and movement fragmentation improved substantially across four trials of practice and had not plateaued; the measures set by the configured rate ceiling and joint limit changed the least. Finger activity doubled and pinch variability tripled, so reducing instrument degrees of freedom redistributed manual effort rather than reducing it. The interface preserves the outcome and reshapes the execution.
Chinese Translation
机器人显微手术的输入设备常被描述为能够保留外科医生所受训练的操作技术,但这一说法鲜有量化验证。我们在套管针约束、器械、眼部模型和追踪信号源在两种条件下均保持一致的前提下,对手动与遥操作的眼内器械运动进行了比较,从而使控制界面成为唯一变化的因素。以往的对比无法固定器械,因为机器人器械并非手动操作时所使用的工具。十六名参与者在商用眼科模拟器上完成一项导航任务,分别通过手动方式以及通过一台指挥五关节机器人的三自由度输入设备进行操作。任务结果相同且已达上限:所有参与者在两种界面下均成功获取全部五个目标,且未造成视网膜或晶状体损伤。但在所有测量指标上执行过程均存在差异。遥操作试验耗时为手动的三倍,中位速度仅为四分之一,角工作范围不足一半,且被分割成的独立运动数量是手动的3.5倍。在四次练习试验中,完成时间和运动碎片化程度显著改善且尚未趋于平稳;而由设定的速率上限和关节限位所决定的指标变化最小。手指活动量增加了一倍,捏合变异性增加了两倍,因此减少器械自由度重新分配了手部操作负荷,而非减轻它。该界面保留了操作结果,却重塑了执行过程。
cs.RO / 115 / 2609.08800

Ostrich: Taking Large Strides Through Stiff Contact in Differentiable Dynamics

Ostrich:在可微动力学中大步跨越刚性接触
Kučera, Aleš, Zimmermann, Karel
Abstract
Three properties determine whether a differentiable simulator can drive gradient-based optimization through contact: simulation accuracy, gradient reliability, and per-iteration cost. Tape-based engines such as MJX and Newton Semi-Implicit require timesteps small enough to keep contacts numerically tractable, and their backpropagation memory grows linearly with the number of timesteps T. Surrogate models bound memory by approximating contact away, but the resulting gradients lose the geometry the optimization depends on. We present Ostrich, a GPU-accelerated rigid-body simulator that resolves hard contacts and friction with non-smooth Newton iteration at large timesteps (h ~ 0.1 s), and differentiates the converged residual via the implicit function theorem, reusing the forward Schur complement to compute the adjoint at O(1) memory per timestep. On real-robot trajectories over a pallet obstacle, Ostrich holds MuJoCo's sim-to-real accuracy up to a 50x larger timestep. Its gradients converge from random initializations where MJX descends slowly and Newton Semi-Implicit stalls; a warm iteration runs 211x faster than MJX's and 4.7x faster than Semi-Implicit's. On the same scene Ostrich differentiates 8,192 parallel worlds on a single 24 GB GPU, sustaining 29x checkpointed MJX's optimization throughput; without checkpointing both baselines exhaust memory at far fewer worlds. We close with a gradient-based trajectory optimization demonstration over triangle-mesh terrain across a 10 s horizon, a setting where prior engines either restrict to primitive geometry or face the convergence and memory limits shown above.
Chinese Translation
可微仿真器能否通过基于梯度的优化有效处理接触问题,取决于三个性质:仿真精度、梯度可靠性和单次迭代成本。基于记录(tape-based)的引擎(如 MJX 和 Newton Semi-Implicit)需要足够小的时间步长以保持接触问题的数值可处理性,且其反向传播内存随时间步数 T 线性增长。代理模型通过近似消除接触来限制内存,但所得梯度丢失了优化所依赖的几何信息。我们提出 Ostrich,一种 GPU 加速的刚体仿真器,能够在较大时间步长(h ~ 0.1 s)下通过非光滑牛顿迭代求解硬接触与摩擦,并利用隐函数定理对收敛后的残差进行微分,通过复用前向 Schur 补来计算伴随量,实现每时间步 O(1) 的内存开销。在跨越托盘障碍的真实机器人轨迹上,Ostrich 在时间步长放大 50 倍的情况下仍能保持 MuJoCo 的仿真到现实(sim-to-real)精度。其梯度能从随机初始化收敛,而 MJX 下降缓慢、Newton Semi-Implicit 则停滞不前;热启动迭代下 Ostrich 比 MJX 快 211 倍,比 Semi-Implicit 快 4.7 倍。在同一场景中,Ostrich 可在单张 24 GB GPU 上对 8,192 个并行世界进行微分,其优化吞吐量达到使用检查点(checkpointing)的 MJX 的 29 倍;不使用检查点时,两个基线方法在远少于该数量的并行世界下即耗尽内存。最后,我们展示了基于梯度的轨迹优化演示:在三角形网格地形上跨越 10 秒时域进行优化——在该场景下,先前引擎要么局限于基本几何体,要么面临上述收敛与内存限制。
cs.RO / 116 / 2609.08802

Graph-Based Safe Reinforcement Learning for Multi-Agent Systems with Time-Varying Topology

面向时变拓扑多智能体系统的基于图的安全强化学习
Sizhe, Xiao, Lijing, Dong, Rui, Bai, Xin, Tan
Abstract
This paper presents a graph-based safe multi-agent reinforcement learning (MARL) framework for cooperative navigation with time-varying topology. To address the critical challenge of ensuring safety in environments with sensing constraints, a safety-decoupled mechanism is introduced through a Control Barrier-Like Function (CBLF) action screening layer. This mechanism bridges the gap between discrete LiDAR perception and continuous safety constraints, ensuring that physical safety constraints are strictly satisfied regardless of the learning progress. Building upon this safety foundation, a unified structural architecture is proposed, integrating a attention-based actor and a Graph Attention Network (GAT) centralized critic. The actor utilizes a value vector reconstruction mechanism that explicitly encodes relative geometric relations through a collaborative tracking error matrix, enabling scale-insensitive policy learning under time-varying communication topologies. Meanwhile, the GAT-based critic models evolving interaction structures for accurate global value estimation. The proposed framework is validated on real differential-drive robot platforms, and experimental results demonstrate superior stability and safety in dynamic scenarios with limited fields-of-view.
Chinese Translation
本文提出了一种面向时变拓扑下协同导航的基于图的安全多智能体强化学习(MARL)框架。为解决在感知受限环境中确保安全这一关键挑战,本文通过类控制屏障函数(Control Barrier-Like Function, CBLF)动作筛选层引入了一种安全解耦机制。该机制弥合了离散激光雷达(LiDAR)感知与连续安全约束之间的差距,确保无论学习进展如何,物理安全约束都能被严格满足。在此安全基础之上,本文提出了一种统一的结构化架构,集成了基于注意力的执行器(actor)和基于图注意力网络(Graph Attention Network, GAT)的集中式评价器(critic)。执行器利用价值向量重构机制,通过协同跟踪误差矩阵显式编码相对几何关系,实现了在时变通信拓扑下对规模不敏感的策略学习。同时,基于GAT的评价器对不断演化的交互结构进行建模,以实现准确的全局价值估计。该框架在实际差速驱动机器人平台上得到了验证,实验结果表明其在视场受限的动态场景中具有优异的稳定性与安全性。
cs.RO / 117 / 2609.08804

Real-time Puncture Detection and Recovery for Pneumatic Soft Actuators

气动软体执行器的实时穿刺检测与恢复
Deshpande, Tejonidhi R., Cheng, Tingyu, Hester, Josiah
Abstract
Soft robots offer safe and adaptive interaction with humans and unstructured environments through their inherent ability to deform and comply. Pneumatic actuators are one way to build soft robots. They are typically made from soft silicone materials and are especially effective for driving such systems, enabling smooth and adaptable motion. However, their compliant nature also makes them vulnerable to mechanical failures like punctures and tears, limiting practical deployment. To address this, we propose a puncture detection system for soft actuators using motion data from a single inertial measurement unit. Extracted features are used to train anomaly detectors for puncture detection and non-linear models to estimate severity. We also introduce a multi-chamber pneumatic soft bending actuator capable of diverse configurations via selective chamber inflation. Our algorithm identifies the punctured chamber and provides a severity score using a chamber perturbation scheme. Anomaly detectors are trained on normal operation data and detect damage through reconstruction errors, while severity is estimated by a separate model trained under slightly modified conditions. Finally, we demonstrate a failure recovery strategy to maintain actuation force post-failure. This approach enhances the reliability and safety of soft robotic systems through real-time, data-driven damage detection.
Chinese Translation
软体机器人凭借其固有的变形和顺应能力,能够与人类及非结构化环境进行安全、自适应的交互。气动执行器是构建软体机器人的一种方式,通常由软硅树脂材料制成,对于驱动此类系统尤为有效,可实现平滑且适应性强的运动。然而,其柔顺特性也使其容易受到穿刺和撕裂等机械损伤,限制了实际部署。为解决这一问题,我们提出了一种使用单个惯性测量单元(IMU)运动数据的软体执行器穿刺检测系统。提取的特征用于训练异常检测器以实现穿刺检测,并训练非线性模型以估计损伤严重程度。我们还介绍了一种多腔室气动软体弯曲执行器,可通过选择性腔室充气实现多种构型。我们的算法通过腔室扰动方案识别被穿刺的腔室并给出严重程度评分。异常检测器在正常运行数据上训练,通过重构误差检测损伤,而严重程度则由一个在轻微修改条件下训练的独立模型进行估计。最后,我们演示了一种故障恢复策略,以在故障发生后维持驱动力。该方法通过实时、数据驱动的损伤检测,提升了软体机器人系统的可靠性与安全性。
cs.RO / 118 / 2609.08853

CAST: Alternating State-Value Targets and Expanded Policy Gradients for Model-Based Reinforcement Learning

CAST:用于基于模型的强化学习的交替状态值目标与扩展策略梯度方法
Crestaz, Pietro Noah, Kabouri, Mohamed Yassine, Mansard, Nicolas, Del Prete, Andrea
Abstract
Model-based reinforcement learning (MBRL) is a family of RL methods that learn a model of the environment and use it for action selection, making it well suited to robotics due to its sample efficiency. Combining learned models with online planning can further improve action selection, as the planner can exploit the model to find better actions than the learned policy alone. Recent methods combining learned policies with online planning typically learn the value of the policy rather than the stronger planner-guided behavior. We present CAST (Critic with Alternating State-value Target), which uses planner-guided behavior to improve value learning while regularizing the value estimate with the current policy. CAST replaces the action-value critic with a state-value critic, trained using a target that combines a real planner-guided transition and an imagined transition under the current policy. The resulting value function corresponds to an alternating process between planner-guided behavior and the current policy, allowing it to benefit from the stronger planner behavior while being regularised by the policy being learned. We evaluate CAST on the DeepMind Control and HumanoidBench Suites against several state-of-the-art methods, and demonstrate successful transfer to a physical Unitree Go2 quadruped performing a dynamic handstand.
Chinese Translation
基于模型的强化学习(Model-Based Reinforcement Learning, MBRL)是一类通过学习环境模型并利用该模型进行动作选择的强化学习方法。由于其样本效率高,该方法特别适用于机器人领域。将学习到的模型与在线规划相结合可以进一步改进动作选择,因为规划器能够利用模型找到比仅依靠学习到的策略更好的动作。近来结合学习策略与在线规划的方法通常学习策略的价值,而非更强的规划器引导行为的价值。我们提出了CAST(Critic with Alternating State-value Target,具有交替状态值目标的评论家),该方法利用规划器引导的行为来改进价值学习,同时用当前策略对价值估计进行正则化。CAST用状态值评论家取代动作值评论家,并通过结合真实的规划器引导转移与当前策略下的想象转移的目标来训练该评论家。所得的价值函数对应于规划器引导行为与当前策略之间的交替过程,使其既能从更强的规划器行为中获益,又能受到所学习策略的正则化约束。我们在DeepMind Control和HumanoidBench基准上对CAST与多种最先进方法进行了评估,并成功将其迁移至执行动态倒立动作的物理Unitree Go2四足机器人上。
cs.RO / 119 / 2609.08905

Visible-Reachable Workspace for Perception-Aware Humanoid Design

面向感知感知型人形机器人设计的可视可达工作空间
Xia, Boxi, Yang, Zijiang, Shin, Ryan, Li, Bokuan, Lu, Eric Wun-Hao, Lee, Jacob, Liu, Jiaxun, Chen, Boyuan
Abstract
Workspace analysis measures where a robot can place its end effector. For visually guided manipulation, reachability alone is insufficient: a kinematically reachable target may not be visible in the specific pose required to reach it. The robot must then redirect its sensing or move its body to acquire a view, turning a perception limitation into additional motion. Existing humanoids largely inherit this limitation when copying human form factors. We introduce the visible-reachable workspace (VRW), a design-stage measure that conditions visibility on feasible reaching configurations and extends it to concurrent visibility of spatially separated work regions. We apply VRW by building a 31-DoF humanoid with independently actuated RGB-D cameras. On the same robot, camera articulation increases visible-reachable coverage from 38% to 97%. With actuated camera layouts, a second camera raises pairwise coverage from 0.45 to 0.95, while a third changes it only to 0.97. In a controlled two-target reach-and-grasp benchmark, our dual-actuated design reduces mean completion time by 17% and mechanical energy by 19% relative to the same robot with its cameras fixed. Hardware experiments demonstrate simultaneous observation and manipulation of front/back and left/right target pairs without torso reorientation. The results suggest that reachability becomes a more informative design quantity for perception-driven humanoid manipulation when it is evaluated together with the sensing configurations that make the reachable space observable. We will open-source all software and the humanoid hardware design. Our website is https://generalroboticslab.com/DukeHumanoidv2
Chinese Translation
工作空间分析衡量机器人能够将其末端执行器放置的位置。对于视觉引导的操作而言,仅有可达性是不够的:一个运动学上可达的目标,在到达它所需的特定姿态下可能不可见。此时机器人必须重新调整其传感器的方向或移动身体以获取视野,从而将感知局限转化为额外的运动。现有人形机器人在模仿人体形态时大多继承了这一局限。我们提出了可视可达工作空间(Visible-Reachable Workspace, VRW),这是一种设计阶段的度量方法,它将可见性以可行的抓取构型为条件,并将其扩展到对空间上分离的多个工作区域的并发可见性。我们通过构建一台具有独立驱动RGB-D相机的31自由度人形机器人来应用VRW。在同一台机器人上,相机的关节化使可视可达覆盖率从38%提升到97%。在带驱动相机的布局下,第二个相机将成对覆盖率从0.45提升至0.95,而第三个相机仅将其提升至0.97。在一个受控的双目标抓取基准测试中,与相机固定的同一机器人相比,我们的双驱动设计将平均完成时间降低了17%,机械能耗降低了19%。硬件实验展示了对前后以及左右目标对的同步观测与操作,且无需躯干重新定向。这些结果表明,当可达性与使可达空间可被观测的传感构型一同评估时,它将成为感知驱动的人形机器人操作中更具信息量的设计量。我们将开源所有软件及人形机器人的硬件设计。我们的网站是 https://generalroboticslab.com/DukeHumanoidv2
cs.RO / 120 / 2609.08920

Remotely Detectable Keyed Communication through Motion

通过运动实现可远程检测的密钥化通信
Chang, Benjamin, Amir, Michael, Flageat, Manon, Prorok, Amanda
Abstract
Messages from electronic devices are conventionally received as text, audio, or radio signals. But robots move with rich, articulate motion in the real world, opening up the possibility of transmitting messages through motion itself. In this paper, we consider the problem of motion-based communication, where we seek to modify a robot's movements so as to transmit messages detectable from remote sensing (e.g., video or motion capture), without degrading policy performance. We introduce a method for messaging through motion capable of encoding arbitrary message content over short payloads - such as an agent's current intent - as noise in any pre-trained policy's actions. This brings a new kind of robustness to robot communication: this 'physical' channel complements standard wireless communications channels but does not depend on them, requiring no extra hardware nor the establishment of a direct link to the robot. We systematically characterize the space of encoding schemes and derive design heuristics, then validate them across simulated environments and real-robot deployment; on real robots running at 50 Hz, four robots jointly recover an 8-bit message at an aggregate 0.67 bits/s.
Chinese Translation
来自电子设备的消息通常以文本、音频或无线电信号的形式被接收。但机器人在现实世界中以丰富、精巧的运动方式移动,这开辟了通过运动本身传递消息的可能性。在本文中,我们研究了基于运动的通信问题,即在不降低策略性能的前提下,修改机器人的动作以传递可从远程感知(如视频或动作捕捉)中检测到的消息。我们提出了一种通过运动传递消息的方法,能够将任意消息内容编码为短载荷——例如智能体当前的意图——作为噪声嵌入到任意预训练策略的动作中。这为机器人通信带来了一种新的鲁棒性:这种'物理'信道可以补充标准的无线通信信道,但不依赖于它们,无需额外硬件,也无需与机器人建立直接链路。我们系统地刻画了编码方案的空间并推导出设计启发式原则,随后在多个仿真环境和真实机器人部署中对它们进行了验证;在以50 Hz运行的真实机器人上,四台机器人共同恢复出一条8比特消息,总速率达0.67比特/秒。
cs.RO / 121 / 2609.08958

Model Predictive Control of Tensegrity Robots via Contact-Aware Graph Neural Dynamics Model

基于接触感知图神经网络动力学模型的张拉整体机器人模型预测控制
Chen, Nelson, Meng, Patrick, Tang, Charles, Degay, Angelina, Brei, Zachary, Kramer-Bottiglio, Rebecca, Bekris, Kostas E., Aanjaneya, Mridul
Abstract
Tensegrity robots offer lightweight, compliant mobility over challenging terrain but remain difficult to model and control due to complex contact-rich dynamics and partial observability. This work presents a model predictive path integral (MPPI) controller for a three-bar tensegrity robot driven by a learned graph neural network (GNN) dynamics model. This work first extends prior GNN-based models with a differentiable contact detection module. The extension allows the dynamics model to reason over non-horizontal planar terrains, obstacles, as well as self-collisions. Then, the learned dynamics model and the MPPI controller operate in a closed data-collection loop, iteratively improving model accuracy and control performance. This work further introduces a hybrid MPPI strategy that combines MPPI with turning motion primitives to improve maneuverability. Experiments are performed in MuJoCo across five navigation tasks, which include, wall obstacles, inclines, narrow corridors, low-clearance structures, and a composite 3D obstacle course. The experiments demonstrate that the hybrid MPPI controller operating over the learned GNN dynamics model improves predictive accuracy over a flat-ground baseline model and achieves superior navigation performance compared to $A^*$-based re-planning and MPPI-only variants. Results show that the contact-aware learned dynamics combined with the sampling-based model predictive control enable robust tensegrity navigation in complex, contact-rich environments.
Chinese Translation
张拉整体(Tensegrity)机器人具有轻量、柔顺的特点,能够在复杂地形上移动,但由于其复杂的接触丰富动力学和部分可观测性,其建模与控制仍然十分困难。本工作提出了一种用于三杆张拉整体机器人的模型预测路径积分(MPPI)控制器,该控制器由学习得到的图神经网络(GNN)动力学模型驱动。本工作首先通过引入一个可微的接触检测模块对先前的基于GNN的模型进行扩展,使动力学模型能够对非水平平面地形、障碍物以及自碰撞进行推理。随后,学习到的动力学模型与MPPI控制器在闭环数据采集流程中运行,迭代地提升模型精度和控制性能。本工作进一步提出了一种混合MPPI策略,将MPPI与转向运动基元相结合以提高机动性。实验在MuJoCo中针对五个导航任务进行,包括墙壁障碍、斜坡、狭窄走廊、低净空结构以及一个复合三维障碍赛道。实验表明,基于学习到的GNN动力学模型运行的混合MPPI控制器,相比平地基线模型提高了预测精度,并且相比基于A*的重规划方法和纯MPPI变体取得了更优的导航性能。结果表明,接触感知的学习动力学与基于采样的模型预测控制相结合,能够实现张拉整体机器人在复杂、接触丰富环境中的鲁棒导航。
cs.RO / 122 / 2609.09023

DYAD: A Multimodal Dataset of Co-Located Human Assistance

DYAD:一个共处一地的人类辅助多模态数据集
Ajikumar, Akhil, Qorbani, Mahya, Reza, Sakib, Andrist, Sean, Moghaddam, Mohsen
Abstract
An embodied assistant working beside a person must track task state, recognize help seeking, choose how to intervene, and produce an appropriate response. Existing procedural datasets richly describe individual execution, while interactive datasets capture remote verbal instruction or undifferentiated co-working. They do not jointly link a co-located helper's verbal and physical interventions to performer requests, task state, assistance triggers, and outcomes. We introduce DYAD (DYadic Assistance Dataset), a synchronized multimodal record of human-human assistance during gearbox assembly. Across 20 sessions, one trained helper follows a guidance-first policy while assisting HoloLens 2 wearers. DYAD links 528 task-step intervals and 611 performer requests with 851 valid assistance records spanning verbal and physical help. DYAD's annotations span the assistance process; three reference tasks evaluate selected components rather than an end-to-end system: causal step understanding, pre-onset mode anticipation, and instructor response generation. On 829 eligible mode events, the strongest four-seed RGB mean is 0.548 +/- 0.007 macro-F1; causal metadata reaches 0.624 and a privileged trigger mapping 0.915, revealing information not recovered from pre-onset RGB. DYAD's contribution is not scale, but a linked interaction structure spanning help seeking, intervention choice, execution, and outcome under egocentric and workspace sensing.
Chinese Translation
一个在人身边工作的具身助手必须跟踪任务状态、识别求助行为、选择干预方式,并生成恰当的响应。现有的程序性数据集充分描述了个体执行过程,而交互式数据集则捕获了远程的口头指令或无差异的协同工作。它们均未能将共处一地的帮助者的口头和肢体干预与执行者的请求、任务状态、辅助触发条件及结果共同关联起来。我们提出了DYAD(DYadic Assistance Dataset,二元辅助数据集),这是一个关于变速箱组装过程中人与人之间辅助行为的同步多模态记录。在20次会话中,一名受过训练的帮助者遵循“引导优先”的策略,为佩戴HoloLens 2的执行者提供辅助。DYAD将528个任务步骤区间和611次执行者请求与851条涵盖口头和肢体辅助的有效辅助记录相关联。DYAD的标注覆盖了整个辅助过程;三个参考任务评估了选定组件而非端到端系统:因果步骤理解、辅助发生前的模式预判,以及指导者响应生成。在829个符合条件的模式事件上,最强的四种子(four-seed)RGB模型的平均macro-F1为0.548 +/- 0.007;因果元数据达到0.624,而特权触发映射达到0.915,揭示了无法从辅助发生前RGB数据中恢复的信息。DYAD的贡献不在于规模,而在于一个在自我中心和工位感知条件下、贯穿求助、干预选择、执行和结果的关联交互结构。
cs.RO / 123 / 2609.09066

A Distributed Consensus Particle Filter for Target Tracking using Autonomous Surface Vessels

一种基于自主水面船舶的目标跟踪分布式一致性粒子滤波方法
Noh, Carter, Crandall, Kyle, Yates, Connor, Wilhelmi, Corbin
Abstract
Maritime target tracking over large distances often requires multi-agent teams without centralized coordination, and intermittent communication. Each agent must maintain an independent estimate that can take advantage of opportunistic communications availability when possible. This can lead to overly confident local estimates in the absence of external data. In this work, we propose an augmentation to a classical particle filter implementation that accounts for this potential source of error by forcing particles to spread strategically in the absence of informative updates from other sensor nodes. We demonstrate our method using Unmanned Surface Vessels (USVs) on a lake, and show that our augmentations do not deteriorate nominal performance, and provide an advantage in some specific edge cases.
Chinese Translation
远距离海上目标跟踪通常需要多智能体团队协作,且缺乏集中式协调,通信往往是间歇性的。每个智能体必须维护一个独立的估计结果,并能在可能时利用机会式通信获取的信息。在缺乏外部数据的情况下,这可能导致局部估计过度自信。在本工作中,我们提出对经典粒子滤波实现的一种改进,通过在缺乏来自其他传感器节点的有效更新时迫使粒子进行策略性扩散,来应对这一潜在误差来源。我们通过无人水面船舶(USV)在湖泊上进行了实验验证,结果表明我们的改进方法不会降低标称性能,并且在某些特定的边缘情况下能够带来优势。
cs.RO / 124 / 2609.09069

Rethinking Learned Occupancy in Autonomous Active Mapping with Observation-Gated Filtering

基于观测门控滤波的自主主动建图中学习型占据栅格的重新思考
Zhang, Jiahui, Han, Bonian, Liang, Gongbo, Zhang, Yu
Abstract
Autonomous 3D active mapping requires a space robot to choose where to sense while building the geometry needed for navigation. Learned occupancy completion extends spatial context beyond the current field of view, but one predicted map often serves two planning roles: it scores expected surface gain and constrains collision-free motion. Unsupported occupancy can therefore distort both where the robot looks and where it believes it can travel. We study this coupled interface in a controlled closed-loop benchmark by holding the active-mapping system fixed and varying only its planner-facing occupancy across observation-only, learned, oracle-corrected, and ground-truth conditions. Improving occupancy accuracy does not monotonically improve closed-loop coverage: across 25 starts, planning with ground-truth occupancy reaches 70% of the learned baseline's final coverage 12.7 steps earlier on average, while increasing final coverage by only 0.031. Guided by this diagnosis, we introduce an observation-gated filter that retains completion in insufficiently observed regions and suppresses predictions only after repeated frustum exposure without nearby RGB-D support. The filter improves both targeted failure-prone starts without retraining or ground truth. These results motivate online revision of planner-facing geometry during autonomous intervals between communication windows. The current study assumes benchmark RGB-D observations and sufficiently accurate pose estimates; planetary sensing conditions and accumulated localization drift remain to be evaluated.
Chinese Translation
自主三维主动建图要求空间机器人在构建导航所需几何信息的同时选择感知位置。学习型占据补全(learned occupancy completion)将空间上下文扩展到当前视野之外,但同一张预测地图往往承担两种规划角色:既用于评估预期表面增益,又用于约束无碰撞运动。因此,缺乏支撑的占据预测可能同时扭曲机器人的观测位置选择及其对可行通行区域的判断。我们在一个受控的闭环基准中研究这一耦合接口:保持主动建图系统不变,仅改变规划器所使用的占据地图,包括仅观测、学习型、真值校正和地面真值四种条件。实验表明,提升占据精度并不能单调地改善闭环覆盖率:在25个起始场景中,使用地面真值占据进行规划平均提前12.7步达到学习型基线最终覆盖率的70%,而最终覆盖率仅提高0.031。基于这一诊断,我们提出一种观测门控滤波器,在观测不足的区域保留补全结果,仅在多次视锥曝光且附近缺乏RGB-D支持时才抑制预测。该滤波器在无需重训练或真值的情况下改善了目标性的易失败起始场景。这些结果激发了在通信窗口之间的自主间隔内对规划器所用几何信息进行在线修正的需求。本研究假设基准RGB-D观测和足够精确的位姿估计;行星探测条件以及累积的定位漂移仍有待评估。
cs.RO / 125 / 2609.09073

Online, Reachability-Aware, Sampling-Based Motion Planning

在线的、可达性感知的、基于采样的运动规划
Gould, Brendan, Zhang, Zhiyuan, Tsiotras, Panagiotis, Coogan, Samuel
Abstract
Sampling-Based Model-Predictive Control (MPC) algorithms are a flexible class of controllers used for navigation on a wide range of robotic systems. Historically, such approaches have lacked hard safety guarantees, a shortcoming which we remedy in this work by computing guaranteed reachable-set overapproximations online with a fast, interval-based pipeline. We show that our method achieves similar performance to a state-of-the-art reachability-based planner without the need for the expensive pre-computation step, and can be scaled to systems that are infeasible using existing approaches. Finally, we demonstrate that our technique reduces safety violations by over 99% in a racing simulation and successfully controls a model racecar on real hardware experiments without crashes.
Chinese Translation
基于采样的模型预测控制(MPC)算法是一类灵活的控制器,被广泛用于各类机器人系统的导航。历史上,这类方法缺乏严格的安全性保证,本文通过一种快速的、基于区间的流程在线计算可达集的保证性过近似来弥补这一不足。我们证明,我们的方法在无需昂贵预计算步骤的情况下,可达到与最先进的基于可达性的规划器相近的性能,并且能够扩展到使用现有方法不可行的系统。最后,我们在赛车仿真中展示了我们的技术将安全违规减少了超过99%,并在真实硬件实验中成功控制了一辆模型赛车而没有发生碰撞。
cs.RO / 126 / 2609.09119

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

DeCAL:基于接触感知的潜在共想象实现物理接地灵巧视觉-语言-动作模型
Fu, Yankai, Chen, Ning, Zhao, Junkai, Zhang, Heng, Yao, Guocai, Wang, Pengwei, Wang, Zhongyuan, Zhang, Shanghang
Abstract
Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision-language-action (VLA) models due to severe visual occlusions and complex contact dynamics. While recent works have incorporated tactile sensing into robotic manipulation, most approaches still rely on homogeneous multimodal fusion, lacking adaptive tactile integration and explicit modeling of physical dynamics. In this work, we present DeCAL, a physically-grounded dexterous vision-language-action model that unifies understanding, imagination and action generation for contact-rich dexterous manipulation. Built upon a Mixture-of-Transformers (MoT) architecture, DeCAL leverages specialized experts for each capability while enabling efficient information flow among them. To effectively leverage tactile information, we introduce Adaptive Visuo-Tactile Fusion that dynamically regulates tactile interactions via a contact-aware gating strategy. Furthermore, we propose Visuo-Tactile Latent Co-Imagination to jointly model visual and tactile dynamics, equipping the policy with implicit physical world knowledge. Experimental results show that DeCAL consistently achieves state-of-the-art performance across all tasks, attaining a 71% average success rate and an 83.4% progress success rate, while also demonstrating strong generalization to unseen scenarios. The website is available at https://aureleopku.github.io/DeCAL.
Chinese Translation
灵巧操作涉及与物理世界之间丰富的接触和细粒度交互,由于严重的视觉遮挡和复杂的接触动力学,给现有的视觉-语言-动作(VLA)模型带来了重大挑战。尽管近期工作已将触觉感知引入机器人操作,但大多数方法仍依赖于同质化的多模态融合,缺乏自适应的触觉整合以及显式的物理动力学建模。在本工作中,我们提出了 DeCAL,一个物理接地的灵巧视觉-语言-动作模型,它为接触丰富的灵巧操作统一了理解、想象和动作生成能力。DeCAL 构建于混合变换器(Mixture-of-Transformers, MoT)架构之上,为每种能力配备专用专家模块,同时实现它们之间高效的信息流动。为了有效利用触觉信息,我们引入了自适应视觉-触觉融合(Adaptive Visuo-Tactile Fusion),通过接触感知的门控策略动态调节触觉交互。此外,我们提出了视觉-触觉潜在共想象(Visuo-Tactile Latent Co-Imagination),联合建模视觉与触觉动力学,为策略赋予隐式的物理世界知识。实验结果表明,DeCAL 在所有任务上始终取得最先进的性能,平均成功率达 71%,进展成功率达 83.4%,同时在未见场景中展现出强大的泛化能力。项目网站见 https://aureleopku.github.io/DeCAL。
cs.RO / 127 / 2609.09148

Proxy Policy Steering

代理策略引导
Ning, Chuanruo, Wang, Tianrui, Ma, Wei-Chiu, Fang, Kuan
Abstract
Generalist robot policies carry broad manipulation priors from large-scale data, but specializing them to a new task remains the deployment bottleneck. This requires eliciting task-specific behavior from limited demonstrations without degrading their broad capabilities. We introduce Proxy Policy Steering (PPS), an inference-time adaptation method that resolves this challenge by training two lightweight proxy policies whose calibrated velocity-space difference steers the frozen base sampler. A reference proxy models the frozen base's behavior on target-task observations, and a task proxy, initialized from the reference, captures how this behavior changes under task supervision. Their difference forms a calibrated velocity-space residual that steers the frozen base sampler at every denoising step. We identify the conditions under which this residual isolates the change induced by task supervision, and validate them empirically. Because the base is never directly modified, its broad capabilities remain available at inference, including behaviors such as recovery from failure that the demonstrations themselves do not exercise. Adaptation requires only forward velocity predictions from the base, making PPS lightweight to train and applicable even without access to the base's parameters. On 8 real-world and 4 simulation manipulation tasks, PPS lifts the state-of-the-art pi 0.5 base policy by 53% absolute success rate on average, with zero-to-one gains on tasks the base never solves, while preserving the base's broad capabilities. PPS outperforms LoRA fine-tuning, from-scratch specialists, residual policies, and prior inference-time steering methods.
Chinese Translation
通用机器人策略从大规模数据中获得了广泛的操作先验,但将其专门化到新任务仍是部署瓶颈。这需要从有限的演示中引导出任务特定的行为,同时不损害其广泛能力。我们提出了代理策略引导,这是一种推理时自适应方法,通过训练两个轻量级代理策略来解决这一挑战,其经校准的速度空间差异可引导冻结的基础策略采样器。参考代理建模冻结基础策略在目标任务观测下的行为,而任务代理(由参考代理初始化)捕捉该行为在任务监督下如何变化。二者的差异构成一个经校准的速度空间残差,在每个去噪步骤中引导冻结的基础采样器。我们确定了该残差能够分离出任务监督所引起变化的条件,并进行了实证验证。由于基础策略从未被直接修改,其广泛能力在推理时仍然可用,包括演示本身未曾展现的失败恢复等行为。自适应仅需基础策略的前向速度预测,使PPS训练轻量,即使无法访问基础策略参数也可应用。在8个真实世界和4个仿真操作任务上,PPS将最先进的pi 0.5基础策略的平均成功率绝对提升了53%,在基础策略从未解决的任务上实现了从零到有的突破,同时保留了基础策略的广泛能力。PPS优于LoRA微调、从头训练的专家策略、残差策略以及现有的推理时引导方法。
cs.RO / 128 / 2609.09158

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

TANGO:基于全身视觉-语言-动作模型的人形机器人在杂乱环境中的导航
Li, Anqi, Chen, Yuxin, Li, Zhaobo, Cao, Zhuo, Ren, Junli, Tomizuka, Masayoshi, Shah, Dhruv
Abstract
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requires continuous geometry-aware whole-body adaptation, including coordinated arm placement, torso adjustment, and gait modulation for collision-free movement through complex 3D spaces. We introduce TANGO, the first whole-body vision-language navigation framework for language-conditioned humanoid traversal in cluttered environments. Given a natural-language instruction and egocentric RGB observations, TANGO directly predicts 29-DoF joint-space actions for downstream whole-body control. We train TANGO entirely in simulation by synthesizing diverse collision-free traversal behaviors via global path planning, kinematic whole-body motion generation, obstacle-aware motion editing, and RL-based tracking. This pipeline provides dynamically feasible action supervision for learning language-conditioned whole-body policies. In extensive simulation experiments, TANGO demonstrates state-of-the-art performance in vision-language navigation, while outperforming strong modular baselines in navigating challenging scenes requiring obstacle negotiation. Lastly, we deploy TANGO zero-shot on a Unitree G1 humanoid robot, and observe robust language-guided traversal in cluttered real-world scenes without training on any real-world navigation data.
Chinese Translation
我们研究了人形机器人在杂乱室内环境中导航的问题。与将导航建模为二维路径规划问题的传统方法不同,人形机器人在杂乱环境中的通行需要具备几何感知能力的连续全身自适应,包括手臂位置的协调、躯干的调整以及步态的调制,以实现在复杂三维空间中的无碰撞移动。我们提出了TANGO,这是首个面向语言条件下人形机器人杂乱环境通行的全身视觉-语言导航框架。给定自然语言指令和自我中心的RGB观测,TANGO直接预测29自由度的关节空间动作,用于下游的全身控制。我们完全在仿真环境中训练TANGO,通过全局路径规划、运动学全身动作生成、障碍物感知的动作编辑以及基于强化学习的跟踪来合成多样化的无碰撞通行行为。该流程为学习语言条件下的全身策略提供了动力学可行的动作监督。在大量仿真实验中,TANGO在视觉-语言导航任务中展现出最先进的性能,并在需要跨越障碍的挑战性场景导航中优于强大的模块化基线方法。最后,我们将TANGO零样本部署到Unitree G1人形机器人上,在未使用任何真实世界导航数据进行训练的情况下,观察到其在杂乱真实场景中具有鲁棒的语言引导通行能力。