← Back to Index
Daily Research Digest

arXiv Papers

2026-09-01
434
Papers
3
Categories
434
Translated
收藏清单 0
人工智能 (Artificial Intelligence)
181
cs.AI / 1 / 2608.28590

DS-Lighting: Making Agent Harnesses Explicit for Data-Science Automation

DS-Lighting:面向数据科学自动化的显式智能体框架(Agent Harness)
Liu, Fan, Liu, Hao
Abstract
Large Language Model (LLM) agents have shown promise for automating data-science workflows, yet their end-to-end performance depends critically on the agent harness that represents tasks, manages execution state, constrains output artifacts, and provides evaluation feedback. Existing data-science agents often leave this harness implicit, making results difficult to reproduce, compare, and attribute across heterogeneous tasks. We introduce DS-Lighting, a unified harness toolkit that makes harness design explicit for data-science automation. DS-Lighting decomposes the harness into four reusable layers: data, workflow, execution, and evaluation, and represents diverse agents as executable operator programs that support both predefined pipelines and adaptive search. We further integrate multiple open-source data-science benchmarks into an MLE-Bench-style task format, enabling controlled comparison under a shared task interface, sandboxed runtime, and metric protocol. Experiments across agents, harnesses, models, and ablations show that explicit harness design improves reproducibility, comparability, and reliability, while reducing avoidable system-level failures in end-to-end data-science workflows. Our code is available at https://github.com/usail-hkust/dslighting
Chinese Translation
大语言模型(LLM)智能体在自动化数据科学工作流方面展现出巨大潜力,但其端到端性能在关键程度上依赖于智能体框架(agent harness)——它负责表示任务、管理执行状态、约束输出产物并提供评估反馈。现有的数据科学智能体往往将这一框架隐式化,导致结果在异构任务间难以复现、比较和归因。我们提出 DS-Lighting,这是一个统一的框架工具包,使数据科学自动化的框架设计显式化。DS-Lighting 将框架分解为四个可复用的层:数据、工作流、执行和评估,并将各类智能体表示为可执行的算子程序,既支持预定义流水线,也支持自适应搜索。我们进一步将多个开源数据科学基准整合为 MLE-Bench 风格的任务格式,从而在共享的任务接口、沙箱运行环境和指标协议下实现受控比较。跨智能体、框架、模型和消融实验的实验结果表明,显式的框架设计提升了可复现性、可比性和可靠性,同时减少了端到端数据科学工作流中可避免的系统级失败。我们的代码可在 https://github.com/usail-hkust/dslighting 获取。
cs.AI / 2 / 2608.28591

Expert-validated STEM QA

经专家验证的STEM问答数据集(Expert-validated STEM QA)
Han, Kihwan, Patil, Saurabh, Shukla, Chinmayee, Sharma, Abhinav, Pavlovic, Marko, Lall, Anshuman, Joshi, Mahesh
Abstract
Recent advancements in AI are helping scientists achieve breakthroughs in fields such as mathematics, medicine, and materials sciences. New evaluation datasets for AI models contribute to such advancement in AI. In the STEM domain, frontier models have consumed most of the available online data, creating the need for human-created datasets that codify the knowledge of leading experts in the domain. There are several STEM datasets available for the research community in this field. However, there are some gaps in these datasets, leaving room for improvement. Examples of gaps include (1) saturation in model performance on these datasets, leaving no head-room for meaningful evaluations, (2) skewed taxonomy distributions, (3) multiple choice question format that is misaligned with how scientists use AI in the real world, and (4) inaccurate answers and rationales partially led by a contest-based data collection and a time-bound review process. In this study, we present 'Expert-validated STEM QA', a high-quality, expert-validated STEM dataset (N=398) in Physics, Chemistry, Biology, and Mathematics, created by 241 domain experts. We (1) carefully designed a taxonomy with balanced distribution, (2) vetted question contributors with quality-driven incentive, (3) conducted multiple rounds of reviews with revisions validated by domain experts based on consensus, and (4) created the dataset in verifiable question and answer format. Our study demonstrated low performance ($<25\%$) of frontier AI models on the dataset as a benchmark. Post-training on a separate, private version of the dataset (N=2,000) increased performance of the open source model by $15\%$ relative to the baseline model (p=0.045) on the STEM subset of HLE-verified dataset, indicating potential utility of the dataset for model training. We have open-sourced a portion of our dataset for the AI research community.
Chinese Translation
人工智能领域的最新进展正在帮助科学家在数学、医学和材料科学等领域取得突破。面向AI模型的新型评估数据集有助于推动此类AI进步。在STEM(科学、技术、工程和数学)领域,前沿模型已经消耗了大部分可用的在线数据,因此需要由人工创建的数据集来凝聚该领域顶尖专家的知识。目前已有若干STEM数据集可供研究界使用,但这些数据集仍存在一些差距,有待改进。这些差距包括:(1) 模型在这些数据集上的性能已趋于饱和,无法为有意义的评估留出提升空间;(2) 分类体系的分布不均衡;(3) 多选题格式与科学家在现实世界中使用AI的方式不一致;(4) 答案和解析不准确,部分原因在于基于竞赛的数据收集方式和有时间限制的审阅流程。在本研究中,我们提出了'经专家验证的STEM问答数据集'(Expert-validated STEM QA),这是一个由241位领域专家创建的、高质量且经专家验证的STEM数据集(N=398),涵盖物理、化学、生物和数学。我们:(1) 精心设计了分布均衡的分类体系;(2) 以质量为导向的激励机制对题目贡献者进行了审核;(3) 进行了多轮审阅,并依据共识由领域专家验证修改;(4) 以可验证的问答格式构建了该数据集。我们的研究以该数据集为基准,表明前沿AI模型的性能较低(<25%)。在另一私有版本的数据集(N=2,000)上进行后训练后,开源模型在HLE验证数据集的STEM子集上的性能相对基线模型提升了15%(p=0.045),表明该数据集在模型训练方面具有潜在价值。我们已将数据集的一部分开源,供AI研究社区使用。
cs.AI / 3 / 2608.28592

A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making

前沿大语言模型在符合指南且针对具体病例的肿瘤学决策中的集体能力边界
Sheng, Zhang, Li, Jinming, Chen, Wangyang, Bao, Zhiwei, Wang, Yu YoSean
Abstract
Large language models (LLMs) achieve high scores on medical knowledge examinations, yet real-world oncology is not a knowledge test--it is a sequence of guideline-pathway choices, escalation judgments, and commitments under uncertainty. Existing benchmarks largely measure factual recall, leaving open whether frontier LLMs share decision-path blind spots that combining models cannot fix. We built the Oncology Decision Boundary Benchmark (ODBB)--2,005 oncology decision points across NCCN guidelines and colorectal cancer cases--and evaluated nine frontier LLMs (four closed-source, five open-weight families) released between June 2025 and April 2026. A fully deterministic scorer (zero LLM inference) classified outputs into 14 failure types, independently validated by two oncologists (Cohen's weighted $\kappa$ = 0.939 and 0.790) on a 225-item stratified sample. Treating the nine as a pooled super-model, 42.1% (Wilson 95% CI 40.0--44.3%) of all items--35.7% of the 1,586 NCCN items and 66.4% of the 419 colorectal-cancer cases--were answered correctly by none, with failures concentrated in choosing between guideline pathways before reasoning within any: a consistent blind spot in clinical meta-judgment that likely requires architectural intervention rather than more training data. Two models tuned for decisiveness (GPT-5.5, Gemini 3.1 Pro Preview) made unsafe commitments three to five times more often than the seven cautious models without scoring higher. In 3--9% of items, models stated the correct next clinical step yet did not commit to it--failures of decision, not knowledge. Model quality is no longer the primary bottleneck for clinical LLM deployment; the binding constraint is the assumption that any single model can be the sole basis for a clinical decision. Progress requires architectures that detect when a model reaches its competence boundary and route the decision to a clinician.
Chinese Translation
大语言模型(LLM)在医学知识考试中能取得高分,然而现实世界的肿瘤学并非知识测验——它是一系列遵循指南路径的选择、是否升级治疗的判断以及在不确定性下的决策承诺。现有基准测试大多只衡量事实性记忆,尚未回答一个开放问题:前沿大语言模型是否共享某种模型组合也无法弥补的决策路径盲区。我们构建了肿瘤学决策边界基准(Oncology Decision Boundary Benchmark, ODBB)——涵盖NCCN指南和结直肠癌病例的2005个肿瘤学决策点——并评估了2025年6月至2026年4月间发布的九个前沿大语言模型(四个闭源模型家族,五个开源权重模型家族)。一个完全确定性的评分器(零LLM推理)将输出归类为14种失败类型,并在225项分层样本上由两位肿瘤学家独立验证(Cohen加权κ分别为0.939和0.790)。将九个模型视为一个合并的超级模型,42.1%的题目(Wilson 95%置信区间为40.0%–44.3%)——其中1586个NCCN题目中的35.7%以及419个结直肠癌病例中的66.4%——没有任何模型答对,且失败集中在在任何一条路径内进行推理之前对指南路径的选择上:这是临床元判断(meta-judgment)上的一致性盲区,可能需要架构层面的干预而非更多的训练数据。两个针对果断性进行调优的模型(GPT-5.5、Gemini 3.1 Pro Preview)做出不安全决策承诺的频率是其余七个谨慎型模型的三到五倍,而得分并未更高。在3%–9%的题目中,模型陈述了正确的下一步临床措施却未做出承诺——这是决策失败,而非知识失败。模型质量已不再是临床大语言模型部署的首要瓶颈;真正的约束性假设是:任何单一模型都可以作为临床决策的唯一依据。要取得进展,需要能够检测模型何时达到其能力边界并将决策转交给临床医生的架构。
cs.AI / 4 / 2608.28593

Statutory AI: Aligning Large Language Models With Legal Norms

法定人工智能(Statutory AI):使大语言模型与法律规范保持一致
Delage, Cindy, Canu, Stéphane, Décombas, Marc, Foureur, Jonathan
Abstract
With the increasing development of AI regulatory frameworks, ensuring that artificial intelligence systems, particularly generative models, operate in accordance with legal and ethical standards has become a critical priority. Existing proposals for AI alignment and value-guided behavior, however, face some limitations. Approaches such as Constitutional AI depend on human supervision, while broad normative frameworks like the Good-for-Humanity (GfH) principle may be overly general and ambiguous to provide actionable governance guidance. To overcome these limitations, we propose a hybrid approach called Statutory AI that employs pre-existing human-authored principles drawn from specific themes within a legal corpus. Specifically, Statutory AI uses legal texts as a constitutional framework, enabling AI systems to autonomously critique and revise their outputs according to established norms. It operates in two stages, both using Chain-of-Thought prompting. The first stage classifies the user prompt into one of the identified themes, while the second stage analyzes it in conjunction with relevant articles selected from the legal corpus of that theme. To illustrate the potential of our approach, we conducted an experiment involving 1,000 red-teaming prompts and five penal themes: discrimination, disclosure of confidential information, violence, fraud, and abuse of vulnerable persons. Statutory AI reduced harmful content by 52 to 59 percentage points across tested models, approximately 10 percentage points higher than standard Constitutional AI, while cutting computation time by over 50%.
Chinese Translation
随着人工智能监管框架的日益发展,确保人工智能系统,特别是生成式模型,按照法律和伦理标准运行已成为一项关键优先事项。然而,现有的人工智能对齐和价值引导行为的方案存在一些局限性。诸如Constitutional AI(宪法人工智能)等方法依赖于人工监督,而像“造福人类”(Good-for-Humanity, GfH)原则这样宽泛的规范框架可能过于笼统和模糊,难以提供可操作的治理指导。为克服这些局限,我们提出了一种名为法定人工智能(Statutory AI)的混合方法,该方法采用从法律语料库中特定主题提取的、预先存在的人类撰写的原则。具体而言,Statutory AI将法律文本用作宪法框架,使人工智能系统能够根据既定规范自主地对其输出进行批判和修改。该方法分为两个阶段,均使用思维链(Chain-of-Thought)提示。第一阶段将用户提示分类到已识别的主题之一,第二阶段结合从该主题的法律语料库中选取的相关条款对其进行分析。为展示该方法的有效性,我们进行了一项实验,涉及1,000个红队测试提示和五个刑法主题:歧视、泄露机密信息、暴力、欺诈以及虐待弱势群体。Statutory AI使被测模型的有害内容减少了52至59个百分点,比标准Constitutional AI高出约10个百分点,同时将计算时间缩短了50%以上。
cs.AI / 5 / 2608.28594

From Question-First to Analyst-First: Domain-Expert Skills and Verified Knowledge Compilation for Proactive Enterprise Analytics

从问题优先到分析师优先:面向主动式企业分析的领域专家技能与经验证的知识编译
Singh, Harmohit, Sharma, Rahul
Abstract
Conversational analytics systems assume the user already has a well-formed question, leaving a non-expert facing a blank query box on an unfamiliar enterprise schema. Commercial 'proactive' tools narrow this gap only by detecting statistical anomalies over analyst-curated metric layers, and academic next-question recommenders depend on query logs that a fresh dataset lacks. We describe a production analytics system that inverts the interaction model from question-first to analyst-first through two coupled architectural ideas. First, a pluggable domain-expert 'skill' abstraction: a folder-based, database-free subject-matter pack (a manifest, per-stage prompt facets, keyword-routed references, report templates, and optional compute) auto-selected per (client, dataset) by deterministic schema matching and spliced as a cross-cutting concern into every stage of an agentic pipeline, the schema explorer, and the report engines, degrading to a strict no-op when absent. Because a skill is a self-contained folder resolved deterministically, the catalogue is open-ended: an extensible marketplace of domain experts. Second, an offline knowledge-compilation loop: an agent probes the dataset's parquet via DuckDB (zero load on production), runs critic-gated per-table convergence with self-healing retries, and data-validates joins by value overlap, producing durable schema knowledge that drives standing expert reports whose every published metric is re-verified by re-executing its evidence SQL, plus suggested questions that mirror the report agenda. These close a proactive loop: reports surface numbers, the numbers seed questions, and a click launches a verified deep dive, all before the query box is used. We give a formal model and report illustrative single-tenant evidence. We make no user-study or benchmark claims; the contribution is the architecture and its defensibility.
Chinese Translation
对话式分析系统假设用户已经有一个表述清晰的问题,这使得非专业用户面对陌生的企业数据模式时只能面对一个空白查询框。商业化的“主动式”工具仅通过在分析师维护的指标层上检测统计异常来缩小这一差距,而学术界的下一问题推荐器则依赖于全新数据集所不具备的查询日志。我们描述了一个生产级分析系统,它通过两个相互耦合的架构思想将交互模型从“问题优先”反转为“分析师优先”。首先,是一个可插拔的领域专家“技能”抽象:一种基于文件夹、无需数据库的主题知识包(包含清单文件、分阶段的提示要素、按关键词路由的参考资料、报告模板以及可选的计算模块),通过确定性的模式匹配针对每个(客户端,数据集)自动选择,并作为一个横切关注点嵌入到智能体式流水线、模式探索器和报告引擎的每个阶段中;当技能缺失时,系统会严格退化为无操作。由于技能是一个可确定性解析的自包含文件夹,其目录是开放的:可形成一个可扩展的领域专家市场。其次,是一个离线知识编译循环:智能体通过 DuckDB 探查数据集的 Parquet 文件(对生产环境零负载),运行由评审机制把关的逐表收敛并配合自愈重试,并通过值重叠对连接关系进行数据验证,从而产出持久的模式知识;这些知识驱动常驻的专家报告——其中每个已发布的指标都通过重新执行其证据 SQL 得到再次验证——以及与报告议程相对应的建议问题。这些机制闭合了主动式循环:报告呈现数字,数字生成问题,点击即可启动经过验证的深度分析——这一切都在使用查询框之前完成。我们给出了形式化模型,并报告了单租户场景的示例性证据。我们不做出用户研究或基准测试方面的声明;本工作的贡献在于该架构及其可辩护性。
cs.AI / 6 / 2608.28595

The Signal in the Noise: An Auditable Reliability Layer for Biomedical Text Classification

噪声中的信号:面向生物医学文本分类的可审计可靠性层
Hassan, Moustafa Yehia, Wong, Sharon, Xuan, Woh Kai
Abstract
Biomedical NLP pipelines routinely presuppose clean input text, yet large-scale corpora assembled through automated PDF parsing harbour pervasive OCR-like artifacts, token splits and merges, hyphenation remnants, and character-level corruption, that systematically erode lexical evidence and degrade downstream classifiers. We introduce a conservative, fully auditable spell-correction reliability layer conceived as a safety-oriented preprocessing module rather than a maximal-accuracy corrector: under conditions of uncertainty, the system abstains from editing, in accordance with a medical do-no-harm philosophy. The deterministic architecture couples bounded edit-distance candidate generation with corpus-derived n-gram scoring and a suite of biomedical safety gates that protect domain-critical terminology. We evaluate the layer both intrinsically, on a manually curated benchmark of 2,104 token-level cases, and extrinsically, on a tri-class CORD-19 topic classifier (Prevention, Treatment, Epidemiology) spanning 10,000 examples under a principled four-run protocol (Clean, Noisy, Restored, Safety). Intrinsically, the layer attains 94.61% error-fix recall on synthetic errors with zero harmful edits on negative controls. Downstream, it recovers approximately 80.45% of the noise-induced macro-F1 degradation, elevating macro-F1 from 0.7654 (Noisy) to 0.7717 (Restored) while preserving near-clean performance (Safety: 0.7721). A supplementary case study on 103 real-world OCR-extracted abstracts classified with BioBERT confirms that transformer encoders appeared relatively robust to mild noise, motivating a future grey-box architecture that integrates bounded neural signals and UMLS lexicons without compromising auditability. The system is fully deterministic, artifact-driven, and designed with deployment and auditability in mind.
Chinese Translation
生物医学自然语言处理(NLP)流程通常默认输入文本是干净的,然而通过自动PDF解析构建的大规模语料库中普遍存在类似OCR的伪影、词元拆分与合并、连字符残留以及字符级损坏,这些缺陷会系统性地削弱词汇证据并降低下游分类器的性能。我们提出一个保守的、完全可审计的拼写纠错可靠性层,它被设计为一个面向安全性的预处理模块,而非追求最大准确率的纠错器:在不确定的条件下,系统选择放弃编辑,遵循医学领域“不伤害”的理念。该确定性架构将有界编辑距离的候选生成与基于语料库的n-gram评分相结合,并配有一系列生物医学安全门控机制,以保护领域关键术语。我们在两个层面对该可靠性层进行评估:内在评估基于一个包含2,104个词元级案例的人工标注基准;外在评估基于一个涵盖10,000个样本的三分类CORD-19主题分类器(预防、治疗、流行病学),并采用严谨的四轮协议(干净、噪声、恢复、安全)。在内在评估中,该层在合成错误上达到94.61%的错误修复召回率,且在阴性对照上零有害编辑。在下游任务中,它恢复了约80.45%由噪声导致的宏F1下降,将宏F1从0.7654(噪声)提升至0.7717(恢复),同时保持接近干净文本的性能(安全:0.7721)。一项针对103篇真实OCR提取摘要、使用BioBERT进行分类的补充案例研究表明,Transformer编码器对轻度噪声表现出相对较强的鲁棒性,这为未来的灰盒架构提供了动机——该架构将集成有界的神经信号与UMLS词库,同时不损害可审计性。该系统完全确定、由伪影驱动,并在设计上充分考虑了部署与可审计性。
cs.AI / 7 / 2608.28596

Paper Pilot: A Human-in-the-Loop Expert System for Evidence-Traceable Scientific Manuscript Generation in Applied Sciences

Paper Pilot:一种用于应用科学中证据可溯源科学论文生成的人在回路专家系统
Jha, Nidhi, Chaudhary, Siddharth, Kulkarni, Ajinkya
Abstract
Large language model (LLM) agents are increasingly embedded in scientific workflows for literature analysis, drafting, and review. Existing systems advance autonomous discovery and manuscript generation, but do not resolve the governance problem that arises when ideas, methods, results, and claims propagate through AI-assisted workflows without mandatory human approval or artifact-level traceability. This paper proposes Paper Pilot, a human-in-the-loop expert system for evidence-traceable scientific manuscript generation in applied sciences. It adapts the Collaborative Agent Reasoning Engineering (CARE) methodology to manuscript development through manuscript-owner approval gates, explicit no-pass criteria, claim classification, audit logging, advisory LLM review, and evidence-locked revision control. The framework defines eight approval gates across the idea-to-claim pipeline and distinguishes literature-grounded from artifact-grounded claims, requiring reported numbers and interpretations to remain traceable to approved evidence; its system prompt is openly released for deployment in ChatGPT, Gemini, Claude, or institutional LLM environments. As a first empirical validation, we evaluate the citation-grounding layer with a controlled, mechanically scored benchmark (two commercial LLMs, real arXiv papers, no LLM judge): under coverage pressure ungated drafters fabricated up to 25% of their citations and never flagged an evidence gap, whereas the same models under Paper Pilot's evidence-locked rules produced zero fabricated citations and surfaced the planted gaps as explicit placeholders. Preliminary results for result grounding, revision, and adversarial robustness point the same way; full evaluation is left to future work. Paper Pilot positions LLM-assisted writing as a controlled human-AI decision-support process rather than a fully autonomous authorship pipeline.
Chinese Translation
大语言模型(LLM)智能体正日益嵌入科学工作流中,用于文献分析、撰写和评审。现有系统推动了自主发现与论文生成的进展,但未能解决当想法、方法、结果和论断在AI辅助工作流中传播时缺乏强制人工审批或制品级可溯源性的治理问题。本文提出Paper Pilot,一种用于应用科学中证据可溯源科学论文生成的人在回路专家系统。该系统将协作智能体推理工程(Collaborative Agent Reasoning Engineering, CARE)方法适配到论文开发过程中,通过论文所有者审批门、明确的否决标准、论断分类、审计日志、建议性LLM评审以及证据锁定式修订控制来实现。该框架在从想法到论断的流程中定义了八个审批门,并区分基于文献的论断与基于制品的论断,要求所报告的数值与解释始终可追溯到已批准的证据;其系统提示词已公开发布,可部署于ChatGPT、Gemini、Claude或机构级LLM环境。作为首个实证验证,我们通过一个受控的、机械化评分的基准(两个商用LLM、真实arXiv论文、不使用LLM评判)评估了引用溯源层:在覆盖率压力下,未经门控的撰写模型伪造的引用高达25%,且从未标记证据缺口;而相同的模型在Paper Pilot的证据锁定规则下实现了零伪造引用,并将植入的证据缺口以显式占位符的形式呈现出来。关于结果溯源、修订以及对抗鲁棒性的初步结果也指向同一方向;完整评估留待后续工作。Paper Pilot将LLM辅助写作定位为一种受控的人机决策支持过程,而非完全自主的写作流水线。
cs.AI / 8 / 2608.28597

The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys

在线调查中智能体AI能力与数据质量控制之间的赛跑
Panda, Sourav, Chona, Hillmer, Das, Rupak Kumar, Kale, Shreyash, Soneji, Shikha, Dodge, Jonathan
Abstract
Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language model (LLM) brain and/or a multimodal processing unit with tool-augmented capabilities) raises new questions about the robustness of these safeguards. We investigate how well agentic AI architectures can complete web-based surveys and pass standard attention checks. We evaluate a single-agent architecture capable of multimodal input processing and tool-based web interaction on a controlled survey sandbox. We analyze the problem from two perspectives. From an attack perspective, we demonstrate how structural vulnerabilities such as exposed DOM metadata and predictable option encoding allow agents to resolve attention checks through structured parsing only. From a defense perspective, we implement a mitigation strategy of DOM metadata obfuscation to remove semantic cues in text-based questions. We evaluate multiple open-source language and multimodal models to study capability and orchestration effectiveness. Based on our evaluations, we offer perspectives on how to simultaneously meet the needs of empiricists and agentic AI researchers.
Chinese Translation
在线调查是众多领域中基础性的数据收集工具,而注意力检查则是保障回答质量的关键防线。然而,智能体AI(以大语言模型(LLM)为核心和/或配备具备工具增强能力的多模态处理单元的目标导向系统)的迅速兴起,对这些防护机制的稳健性提出了新的质疑。我们研究了智能体AI架构在网络调查中的完成能力以及通过标准注意力检查的能力。我们在一个受控的调查沙箱环境中评估了一种具备多模态输入处理和基于工具的网页交互能力的单智能体架构。我们从两个视角分析该问题。从攻击视角出发,我们展示了诸如暴露的DOM元数据和可预测的选项编码等结构性漏洞,如何使智能体仅通过结构化解析即可破解注意力检查。从防御视角出发,我们实施了一种DOM元数据混淆的缓解策略,以消除文本类问题中的语义线索。我们评估了多个开源语言模型和多模态模型,以研究其能力与编排有效性。基于评估结果,我们提出了如何同时满足实证研究者与智能体AI研究者需求的见解。
cs.AI / 9 / 2608.28599

CDPR: Counterfactual Advantage-based Credit Assignment for Cost-Aware Sequential Medical Diagnosis

CDPR:面向成本感知序贯医疗诊断的基于反事实优势的信用分配方法
Peng, Qi, Cai, Yi, Zheng, Changmeng, Wu, Xin, Xie, Jiayuan, Li, Qing
Abstract
Clinical diagnosis is a step-by-step, cost-aware process: a physician orders examinations one at a time, observes the results, and updates the diagnosis before reaching a final conclusion. Most medical language models instead treat diagnosis as a one-pass classification task and ignore the trade-off between a test's value and its cost. We model diagnosis as a cost-aware sequential decision process and train the policy with reinforcement learning. The main difficulty is credit assignment: the only reliable signal comes once at the end of a long trajectory, so it scores a wasteful workup the same as an efficient one. We propose CDPR (Counterfactual Diagnostic Process Reward), which needs no expert labels and no learned critic. CDPR first finds the states where the policy hesitates, using the uncertainty of its action distribution, and then scores the chosen action by its advantage over the alternatives the policy itself would consider, estimated with short rollouts under a utility that balances correctness against test count, cost, and infeasible requests. A rollout cache reuses within-batch trajectories to keep the cost low. We integrate CDPR into GRPO and test it on one in-domain (MIMIC-IV) and two out-of-domain (ClinicalBench and a private hospital dataset) benchmarks. CDPR improves diagnostic accuracy while clearly reducing the number and cost of examinations.
Chinese Translation
临床诊断是一个循序渐进且需考虑成本的过程:医生一次开具一项检查,观察结果,并在得出最终结论之前不断更新诊断。然而,大多数医学语言模型将诊断视为一次性的分类任务,忽略了检查的价值与成本之间的权衡。我们将诊断建模为一个成本感知的序贯决策过程,并使用强化学习训练策略。其主要难点在于信用分配:唯一可靠的信号在长轨迹结束后才一次性给出,因此它会将浪费性的检查流程与高效的检查流程赋予相同的评分。我们提出了CDPR(Counterfactual Diagnostic Process Reward,反事实诊断过程奖励),它既不需要专家标注,也不需要学习评价器(critic)。CDPR首先利用策略动作分布的不确定性找到策略犹豫的状态,然后通过所选动作相对于策略自身会考虑的其他备选动作的优势(advantage)来为其评分,该优势通过短轨迹推演(rollout)在一种平衡正确性与检查次数、成本及不可行请求的效用函数下进行估计。轨迹推演缓存机制可复用批内轨迹以保持计算成本较低。我们将CDPR集成到GRPO中,并在一个域内基准(MIMIC-IV)和两个域外基准(ClinicalBench及一个私有医院数据集)上进行测试。结果表明,CDPR在提高诊断准确率的同时,显著减少了检查的数量和成本。
cs.AI / 10 / 2608.28600

SHAPE of Chain-of-Thought in Math Reasoning

数学推理中思维链的SHAPE分析
Song, Jonghyun, Song, Sangjun, Oh, Minjae, Pyun, Haesung, Lee, Sungsik, Jo, Yohan
Abstract
Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their reasoning remain underexplored. We introduce \texttt{SHAPE}, a framework that analyzes Chain-of-Thought (CoT) trajectories through two lenses developed in mathematics education: (1) semantic spaces: the model's evolving mathematical interpretations of a problem (e.g., algebraic, geometric), and (2) heuristics: the specific mathematical actions taken within those spaces (e.g., simplifying the problem, working backward). We first use \texttt{SHAPE} to analyze the reasoning patterns of various models. Our findings reveal that the mathematical heuristics employed by a model better explain final answer correctness than traditional CoT features. Furthermore, models are likely to reach correct solutions by concentrating their reasoning effort within a few semantic spaces rather than exploring many disparate ones -- a pattern consistent with human behavior. Next, we utilize the \texttt{SHAPE} lens to evaluate whether post-training truly enhances mathematical proficiency. We find that reinforcement learning induces mode-seeking in heuristic usage. Lastly, we post-train LLMs by promoting diverse heuristics and demonstrate its effectiveness in improving accuracy. Overall, \texttt{SHAPE} provides a theoretically-grounded diagnostic framework for decoding LLM reasoning and offers a new path toward post-training LLMs for math reasoning. The code for our model is available at https://github.com/holi-lab/SHAPE-of-CoT
Chinese Translation
大型语言模型(LLMs)在数学推理基准测试中取得了优异的表现,但其推理背后的具有数学意义的技能仍未得到充分探究。我们提出了SHAPE框架,通过数学教育领域中发展的两个视角来分析思维链(Chain-of-Thought, CoT)轨迹:(1)语义空间:模型对问题不断演化的数学理解方式(如代数、几何);(2)启发式方法:模型在这些空间中采取的具体数学行动(如简化问题、逆向推理)。我们首先使用SHAPE分析了多种模型的推理模式。研究结果表明,模型所采用的数学启发式方法比传统的CoT特征更能解释最终答案的正确性。此外,模型倾向于将推理精力集中在少数几个语义空间内,而不是探索许多不同的空间——这一模式与人类行为一致。接下来,我们利用SHAPE视角评估后训练是否真正提升了数学能力。我们发现,强化学习会在启发式方法的使用中引发模式寻求(mode-seeking)行为。最后,我们通过促进启发式方法的多样性对LLMs进行后训练,并证明了其在提升准确率方面的有效性。总体而言,SHAPE为解读LLM推理提供了一个有理论依据的诊断框架,并为面向数学推理的LLM后训练开辟了一条新路径。我们模型的代码可在 https://github.com/holi-lab/SHAPE-of-CoT 获取。
cs.AI / 11 / 2608.28601

Leveraging Generative AI to Design Accessible Interactive Visualizations for Undergraduate Mathematics: A Six-Phase Workflow

利用生成式人工智能为本科数学设计无障碍交互式可视化:一个六阶段工作流程
Sunkula, Mahesh, Chen, Kuan-Hua
Abstract
Interactive visualizations support conceptual understanding in undergraduate mathematics, but building them has required programming expertise most instructors lack. Using a design-based research approach, we develop, deploy, and evaluate a six-phase workflow (Foundation, Customization, Mathematical Depth, Application, Accessibility, Pedagogical Control) that uses generative AI to build WCAG~2.2 Level~AA compliant visualizations without programming. The six phases structure every prompt, scaffold the AI's code generation, and define where human verification is applied. We ask whether the structure reliably yields correct and accessible tools, whether it runs both backward (reverse-engineering prompts from a finished tool) and forward (generating a tool from a plain-language idea), and what verification each phase requires. Across four deployed tools spanning calculus, multivariable calculus, and differential equations, we evaluate mathematical correctness against closed forms, accessibility through automated and manual screen-reader testing, and the errors that recurred. The structure produces structurally complete first-pass tools, but human verification remains mandatory at every phase: each output must be checked for mathematical correctness, accessibility, and pedagogical fit before the next phase begins. The workflow is platform-independent and serves both instructors and students.
Chinese Translation
交互式可视化有助于促进本科数学中的概念理解,但构建此类可视化通常需要大多数教师所不具备的编程专业知识。本研究采用基于设计的研究方法,开发、部署并评估了一个六阶段工作流程(基础、定制、数学深度、应用、无障碍性、教学控制),利用生成式人工智能(generative AI)在无需编程的情况下构建符合 WCAG 2.2 AA 级标准的可视化工具。这六个阶段为每一次提示(prompt)提供了结构,为人工智能的代码生成提供了支架,并界定了人工验证的应用环节。我们探讨以下问题:该结构能否可靠地生成正确且无障碍的工具;它能否同时支持逆向运行(从已完成的工具反向推导提示)和正向运行(从自然语言描述的想法生成工具);以及每个阶段需要何种验证。我们对涵盖微积分、多元微积分和微分方程的四个已部署工具进行了评估,包括对照解析解验证数学正确性、通过自动化和人工屏幕阅读器测试评估无障碍性,以及分析反复出现的错误。该结构能够生成结构完整的首版工具,但在每个阶段人工验证仍然必不可少:在进入下一阶段之前,必须对每个输出进行数学正确性、无障碍性和教学适用性的检查。该工作流程不依赖特定平台,可同时服务于教师和学生。
cs.AI / 12 / 2608.28602

Integrating Triaxial IMU Sensors and Ensemble Learning for Effective Parkinson Disease Severity Classification

融合三轴IMU传感器与集成学习实现有效的帕金森病严重程度分类
Khan, Rehan, Asif, Muhammad Junaid, Ahmad, Rana Fayyaz
Abstract
Parkinson disease PD is a progressive neurodegenerative disease that can have a significant impact on motor performance resulting in the appearance of symptoms such as tremors rigidity postural instabilities and bradykinesia. Timely clinical treatment disease management and quality life of the patients are closely linked to early and appropriate identification of PD. Over the past few years the growth of wearable sensor technology and artificial intelligence AI have made it possible to create noninvasive and data driven disease detection methods. This paper proposes a comparative system using artificial intelligence to detect Parkinsons disease by analyzing the motion and tremor data captured by an inertial measurement unit IMU. The data comprises the signals of the acceleration and gyroscope sensors measuring movement in three directions X Y and Z. The signs and symptoms provide helpful information about subtle motor deficits associated with PD. Several classification models like Support Vector Machine SVM Logistic Regression LR KNearest Neighbors KNN Decision Tree DT Extreme Gradient Boosting XGBoost and Light Gradient Boosting Machine LightGBM were used to compare their effectiveness. The Logistic Regression model had a performance around 75 percent in all evaluation metrics and KNearest Neighbours KNN around 90 percent. The support vector machine SVM performed almost 94 percent whereas the performance of classifiers such as Decision Tree and XGBoost was close to 96 percent and overall classification efficacy respectively. LightGBM model performs consistently at the best rank among all of the evaluated methods having Accuracy, Precision, Recall and F1score of around 97 percent. The results show that the proposed machine learning approach offers an accurate and effective predictive capability in the classification of PD severity.
Chinese Translation
帕金森病(PD)是一种进行性神经退行性疾病,会对运动功能产生显著影响,导致震颤、僵直、姿势不稳和运动迟缓等症状的出现。及时的临床治疗、疾病管理以及患者的生活质量与帕金森病的早期准确识别密切相关。近年来,可穿戴传感器技术和人工智能(AI)的发展使无创、数据驱动的疾病检测方法成为可能。本文提出了一种基于人工智能的比较系统,通过分析惯性测量单元(IMU)采集的运动和震颤数据来检测帕金森病。数据包括测量X、Y、Z三个方向运动的加速度计和陀螺仪传感器信号。这些体征和症状为与PD相关的细微运动缺陷提供了有用的信息。研究采用了多种分类模型,包括支持向量机(SVM)、逻辑回归(LR)、K近邻(KNN)、决策树(DT)、极限梯度提升(XGBoost)和轻量梯度提升机(LightGBM),以比较它们的效果。逻辑回归模型在所有评估指标上的性能约为75%,K近邻(KNN)约为90%。支持向量机(SVM)的性能接近94%,而决策树和XGBoost等分类器的性能分别接近96%。在所有评估方法中,LightGBM模型始终表现最佳,其准确率(Accuracy)、精确率(Precision)、召回率(Recall)和F1分数(F1-score)均约为97%。结果表明,所提出的机器学习方法在帕金森病严重程度分类中具有准确而有效的预测能力。
cs.AI / 13 / 2608.28603

C3-UniMM: Causal Cycle-Consistent Unified Multimodal Modeling via Super Alignment and Shared Decoding Space

C3-UniMM:基于超级对齐与共享解码空间的因果循环一致性统一多模态建模
Shen, Yujie, Shan, Lianlei
Abstract
Unified Multimodal Models aim to achieve any-to-any understanding and generation across arbitrary modalities. However, existing methods primarily rely on modeling implicit statistical correlations and lack cross-modal structural consistency constraints. This deficiency leads to profound issues, including semantic drift, poor compositional generalization, and instability under interventions. In this paper, we propose C3-UniMM, a unified multimodal modeling framework based on Causal Cycle Consistency and Super Alignment. Specifically, we introduce a Structured Latent Causal Graph (SLCG) as a shared cross-modal semantic space and design unified multimodal encoding blocks, enabling understanding and generation to be synergistically optimized within the identical causal semantic structure. Furthermore, we propose a Unified Decoding Space to enforce structural preservation and semantic invertibility during the cross-modal generation process. Theoretical analyses demonstrate that our approach significantly enhances both the invertibility and mechanism invariance of cross-modal mappings. Extensive experimental results across multiple understanding, generation, and compositional generalization tasks indicate that C3-UniMM substantially outperforms existing unified multimodal baselines.
Chinese Translation
统一多模态模型旨在实现跨任意模态的任意到任意的理解与生成。然而,现有方法主要依赖对隐式统计相关性的建模,缺乏跨模态的结构一致性约束。这一缺陷导致了深刻的问题,包括语义漂移、组合泛化能力差以及在干预下的不稳定性。在本文中,我们提出了C3-UniMM,一种基于因果循环一致性(Causal Cycle Consistency)与超级对齐(Super Alignment)的统一多模态建模框架。具体而言,我们引入结构化潜在因果图(Structured Latent Causal Graph, SLCG)作为共享的跨模态语义空间,并设计了统一的多模态编码模块,使理解与生成能够在相同的因果语义结构内协同优化。此外,我们提出了统一解码空间(Unified Decoding Space),以在跨模态生成过程中强制保证结构保持与语义可逆性。理论分析表明,我们的方法显著增强了跨模态映射的可逆性与机制不变性。在多项理解、生成及组合泛化任务上的大量实验结果表明,C3-UniMM大幅优于现有的统一多模态基线方法。
cs.AI / 14 / 2608.28605

MedTVL: Harnessing Vision and Language for Medical Time Series Classification

MedTVL:利用视觉与语言进行医学时间序列分类
Ye, Jiexia, Li, Jia, Tsung, Fugee
Abstract
Recent advancements in multimodal learning for medical time series (MedTS) classification highlight the benefits of integrating complementary modalities for clinical decision. However, existing methods typically focus on bi-modal interactions (e.g., time series and text), leaving the tri-modal synergy between time series, vision, and language largely unexplored. Inspired by diagnostic practice synergizing numerical assessment, visual inspection and clinical context, we introduce MedTVL, a text-guided dual-pathway architecture tailored for MedTS classification. Specifically, it synergizes a convolution-based temporal pathway for fine-grained temporal dynamics from raw numerical sequences and a transformer-based visual pathway for holistic morphological structures from time-series-derived images. Such combination of cross-modal and architectural heterogeneity provides a comprehensive diagnostic perspective. To further resolve potential diagnostic ambiguity, both pathways are guided by adaptive medical textual semantics. Finally, a Mixture-of-Experts mechanism dynamically routes each instance to specialized fusion experts, capturing instance-specific reliance on the temporal and visual pathway outputs. In addition, MedTVL supports multimodal contrastive learning to mitigate the clinical label scarcity challenge. Extensive experiments across multiple medical datasets and tasks, spanning supervised, few-shot, and contrastive learning settings, demonstrate the superiority and transferability of MedTVL, highlighting its potential for robust clinical decision support.
Chinese Translation
多模态医学时间序列(MedTS)分类研究的最新进展凸显了整合互补模态对临床决策的益处。然而,现有方法通常聚焦于双模态交互(如时间序列与文本),对时间序列、视觉与语言之间的三模态协同仍缺乏探索。受诊断实践中数值评估、视觉检查与临床背景相结合的启发,我们提出了 MedTVL,一种专为 MedTS 分类设计的文本引导双通路架构。具体而言,该架构协同了两个通路:基于卷积的时间通路,用于从原始数值序列中捕捉细粒度的时间动态;以及基于 Transformer 的视觉通路,用于从时间序列衍生图像中获取整体的形态特征。这种跨模态与架构上的异质性结合提供了全面的诊断视角。为进一步消除潜在的诊断歧义,两个通路均由自适应的医学文本语义进行引导。最后,混合专家机制将每个实例动态路由至专门的融合专家,从而捕捉实例对时间通路与视觉通路输出的特定依赖。此外,MedTVL 支持多模态对比学习,以缓解临床标签稀缺的挑战。在多个医学数据集与任务上、涵盖监督学习、少样本学习和对比学习设置的大量实验,证明了 MedTVL 的优越性与可迁移性,凸显了其在稳健临床决策支持方面的潜力。
cs.AI / 15 / 2608.28607

RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciences

RegDivergence-101:一个用于生命科学领域跨司法辖区监管矛盾检测的大语言模型基准
Wu, Chuchu, Zhou, Zhiyin, Hu, Jingzhuo, You, Liang
Abstract
Pharmaceutical sponsors developing a drug for both the United States and the European Union must reconcile guidance issued independently by the FDA and the EMA. Where the two agencies require substantively the same thing, a sponsor can file once; where they diverge, a single trial design risks rejection in one region; where one agency is silent on a point the other regulates, the sponsor must infer obligations. Today this reconciliation is performed manually by regulatory-affairs experts. We introduce cross-jurisdiction regulatory divergence detection: given an FDA requirement and an EMA requirement on the same topic, classify their relationship as AGREE, DIVERGE, or SILENT. SILENT is inherently directional (SILENT_FDA vs. SILENT_EMA); we record direction per pair and report per-direction F1 alongside the collapsed label. We release RegDivergence-101, a 101-pair expert-grounded pilot evaluation benchmark (labels grounded in three peer-reviewed FDA/EMA comparison studies and primary FDA/EMA/ICH guidance text; dual-annotation inter-annotator kappa = 0.85), and systematically characterise a four-method baseline hierarchy: lexical heuristic (0.511 macro-F1, 95% CI [0.411-0.605]), NLI cross-encoder (0.233), obligation-level Graph-RAG (0.663 [0.570-0.747]), and flat LLM judge / Claude Haiku (0.830 [0.747-0.908]). Three directional observations emerge at pilot scale (n = 101): SILENT is semantically detectable but invisible to entailment-only formulations; pair-level obligation graphs improve over lexical methods but trail flat-LLM context (CIs partially overlapping); and corpus-level graph construction is the indicated architectural target for large-scale silent-detection. RegDivergence-101 is a pilot release establishing the task formulation and baseline hierarchy; four unrepresented regulatory domains and an expansion roadmap are described in Section 7.
Chinese Translation
同时在美国和欧盟开发药品的制药企业必须协调美国食品药品监督管理局(FDA)与欧洲药品管理局(EMA)独立发布的指导文件。当两个监管机构的要求实质上一致时,企业可以只提交一次申请;当二者存在分歧时,单一的试验设计可能在其中一个地区遭到拒绝;而当一个机构对另一机构所监管的事项保持沉默时,企业必须自行推断相关义务。目前,这种协调工作由药品注册事务专家手工完成。我们提出了跨司法辖区监管分歧检测任务:给定同一主题下的一条FDA要求与一条EMA要求,将二者的关系分类为AGREE(一致)、DIVERGE(分歧)或SILENT(沉默)。SILENT本质上具有方向性(SILENT_FDA与SILENT_EMA);我们记录每一对的方向性,并按方向报告F1值以及合并标签的F1值。我们发布了RegDivergence-101,一个包含101对样本、由专家标注的试点评估基准(标签基于三项经同行评审的FDA/EMA对比研究以及FDA/EMA/ICH原始指导文件;双人标注一致性kappa = 0.85),并系统性地刻画了四种方法的基线层次体系:词法启发式方法(宏平均F1为0.511,95%置信区间[0.411-0.605])、自然语言推理(NLI)交叉编码器(0.233)、义务级Graph-RAG(0.663 [0.570-0.747])以及平铺式大语言模型裁判 / Claude Haiku(0.830 [0.747-0.908])。在试点规模(n = 101)下得出三项方向性观察结论:SILENT在语义上是可检测的,但对于仅依赖蕴含(entailment)的建模方式是不可见的;配对级义务图谱优于词法方法,但仍落后于平铺式大语言模型的上下文理解能力(置信区间部分重叠);语料库级的图谱构建是大规模沉默检测的理想架构目标。RegDivergence-101是一个试点版本,确立了该任务的定义与基线层次体系;四个尚未覆盖的监管领域及扩展路线图将在第7节中介绍。
cs.AI / 16 / 2608.28610

TPvG: A Moral Decision Framework for Large Language Models from One-Shot to Sequential Feedback

TPvG:一个面向大语言模型的从单次决策到序列反馈的道德决策框架
Zhang, Fangyuan, Yu, Dong, Liu, Pengyuan
Abstract
Existing LLM moral evaluations typically present models with isolated moral vignettes and elicit a single-shot decision, neglecting a factor known to profoundly influence human moral behavior: consequence feedback. We introduce TPvG (Text-based Pain-versus-Gain), adapted from a human moral paradigm, which embeds consequence feedback into an everyday moral dilemma of not harming others versus maximising self-gain. TPvG comprises five moral decision tasks, progressing from minimal-context one-shot choices to sequential decisions with explicit consequence feedback. Our results show that LLM moral decisions were strongly affected by decision format (one-shot versus sequential), and explicit receiver feedback produced heterogeneous effects across models. Furthermore, LLM responses to explicit receiver feedback diverged from the human reference pattern, suggesting potentially different decision processes. These findings highlight the need to evaluate whether LLM moral behavior remains stable in high-stakes interactive settings.
Chinese Translation
现有的大语言模型(LLM)道德评估通常向模型呈现孤立的道德情景并获取单次决策,忽视了一个已知对人类道德行为有深远影响的因素:后果反馈。我们提出了 TPvG(Text-based Pain-versus-Gain,基于文本的痛苦与收益权衡),该方法改编自人类道德研究范式,将后果反馈嵌入到一个日常道德困境中:是选择不伤害他人,还是最大化自身收益。TPvG 包含五个道德决策任务,从最少上下文的单次选择逐步过渡到带有明确后果反馈的序列决策。我们的结果表明,LLM 的道德决策受到决策形式(单次决策与序列决策)的强烈影响,且明确的接受者反馈在不同模型中产生了异质性效应。此外,LLM 对明确接受者反馈的响应与人类参照模式存在分歧,这表明其潜在的决策过程可能有所不同。这些发现凸显了评估 LLM 道德行为在高风险交互环境中是否保持稳定的必要性。
cs.AI / 17 / 2608.28612

InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal

InternReviewer 与 InternAdvocate:面向同行评审与反驳中智能体强化学习的客观奖励与评估
Su, Xuerui, Guo, Liya, Pei, Qizhi, Guo, Qipeng, Tian, Zhongbo, Wu, Lijun, Chen, Kai, Wang, Zun
Abstract
Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between domain reasoning and factual grounding. This work presents a comprehensive framework for the development and evaluation of specialized scholarly agents, InternReviewer and InternAdvocate. We first establish a large-scale, high-quality scholarly dataset and integrate a high-efficiency arXiv retrieval tool to enable active evidence gathering. To optimize these agents, we implement an agentic Reinforcement Learning (RL) paradigm driven by a unified objective metric and reward system. This system avoids the biases of subjective model-based judging by employing multi-dimensional criteria, including reference-anchored semantic alignment, structural compliance, and a strict verification mechanism that cross-checks citations against real-time interaction logs to eliminate hallucinations. Experimental results demonstrate that agents trained within this closed-loop framework exhibit significant improvements in reasoning depth and citation accuracy.
Chinese Translation
生成专业的学术内容,如同行评审意见和反驳意见,需要领域推理与事实依据之间精密的协同。本工作提出了一个用于开发和评估专业化学术智能体的综合框架,包括 InternReviewer 和 InternAdvocate。我们首先构建了一个大规模、高质量的学术数据集,并集成了一个高效的 arXiv 检索工具,以支持主动的证据收集。为了优化这些智能体,我们实现了一种由统一的客观度量与奖励系统驱动的智能体强化学习范式。该系统通过采用多维度的评价标准,避免了基于主观模型评判所带来的偏差,这些标准包括以参考文献为锚点的语义对齐、结构合规性,以及一种严格的验证机制——通过将引用与实时交互日志进行交叉核对来消除幻觉。实验结果表明,在该闭环框架内训练的智能体在推理深度和引用准确性方面均表现出显著提升。
cs.AI / 18 / 2608.28620

Preference Elicitation for Policy Optimization and Application to Aligning Heart Transplantation with Human Values

面向策略优化的偏好引出及其在心脏移植与人类价值观对齐中的应用
Zilberstein, Itai, Anagnostides, Ioannis, Sollie, Zachary W, Kilic, Arman, Sandholm, Tuomas
Abstract
Preference elicitation is essential for aligning AI systems with human values. Prior approaches (e.g., for organ allocation) often ask stakeholders to compare the decisions of an algorithm (e.g., patient A vs. patient B). Such a decision-level approach conflates the means with the ends. Instead, we elicit preferences directly over allocation outcomes to learn a utility function for policy optimization. We construct a novel preference elicitation algorithm for linear utilities that outperforms prior techniques in practice. Our algorithm has two phases. The first phase learns cutting planes through pairwise comparisons to rapidly shrink the space of possible attribute weights and warm-starts the second phase by eliminating dominated regions. The second phase then provably converges to the user's utility function. We apply our technique to heart transplant allocation where a policy must balance competing objectives such as post-transplant outcomes, waitlist mortality, geographic ease, and equity. Using our algorithm, we conduct a user study to learn and aggregate a community-aligned utility function, and use it to optimize heart transplant policies that are significantly better aligned with human values. Compared to the hindsight optimum, the status quo policy achieves a competitive ratio of just 0.54, while our method is near-optimal with a competitive ratio of 0.95.
Chinese Translation
偏好引出(preference elicitation)对于使人工智能系统与人类价值观保持一致至关重要。已有方法(例如用于器官分配的方法)通常要求利益相关者比较算法的决策结果(例如,患者A与患者B孰优)。这种决策层面的方法混淆了手段与目的。相反,我们直接就分配结果引出偏好,从而为策略优化学习一个效用函数。我们构建了一种新颖的针对线性效用函数的偏好引出算法,在实践中优于已有技术。我们的算法分为两个阶段:第一阶段通过成对比较学习切割平面,快速缩小可能的属性权重空间,并通过剔除被支配区域为第二阶段提供热启动;第二阶段则可证明地收敛到用户的效用函数。我们将该技术应用于心脏移植分配问题,其中策略需要在移植后疗效、等待名单死亡率、地理便利性和公平性等相互冲突的目标之间进行权衡。利用我们的算法,我们开展了一项用户研究,学习并聚合得到一个与社区价值观一致效用函数,并用其优化心脏移植策略,使其与人类价值观的契合度显著提升。与事后最优解相比,现行策略的竞争比仅为0.54,而我们的方法接近最优,竞争比达到0.95。
cs.AI / 19 / 2608.28627

Machine Learning-Enhanced Tabu Search for Tactical Wireless Network Design

机器学习增强的禁忌搜索在战术无线网络设计中的应用
Zaid, Wissem Ahmed, Hertz, Alain, Liu, Defeng
Abstract
Designing high-performance tactical wireless networks under realistic operational constraints gives rise to challenging combinatorial optimization problems, where the evaluation of candidate solutions relies on detailed physical and traffic-aware models. Although classical metaheuristics such as Tabu Search offer effective mechanisms for exploring large search spaces, their computational cost remains high because numerous candidate moves must be evaluated at every iteration. In this paper, we propose a data-driven framework that improves the efficiency of Tabu Search by learning to guide its move selection process. Rather than altering the neighborhood structure, our approach exploits the information contained in the search trajectories generated during the optimization process. At each iteration, we record both improving and non-improving edge-based transformations together with a set of descriptive features capturing the structural, geometric, and performance characteristics of the network. This information is used to train a Graph Neural Network (GNN) that predicts the impact of candidate moves on the objective function. The trained model is then integrated into the Tabu Search algorithm to rank candidate transformations according to their predicted quality, thereby reducing the number of costly objective evaluations while maintaining an effective exploration of the search space. Experimental results on synthetic benchmark instances demonstrate that the proposed learning-assisted Tabu Search notably reduces computation time while consistently producing higher-quality solutions than the standard algorithm. These findings highlight the potential of combining machine learning with metaheuristics by leveraging the implicit knowledge embedded in search trajectories, paving the way for more efficient solution methods for large-scale network design problems.
Chinese Translation
在现实作战约束下设计高性能的战术无线网络会产生具有挑战性的组合优化问题,其中候选解的评估依赖于详细的物理模型和流量感知模型。尽管禁忌搜索(Tabu Search)等经典元启发式算法为探索大规模搜索空间提供了有效的机制,但由于每次迭代都需要评估大量候选移动,其计算成本仍然很高。本文提出了一种数据驱动的框架,通过学习引导禁忌搜索的移动选择过程来提升其效率。我们的方法并未改变邻域结构,而是利用优化过程中产生的搜索轨迹所包含的信息。在每次迭代中,我们记录基于边的改进性和非改进性变换,以及一组描述网络结构、几何和性能特征的描述性特征。这些信息被用于训练一个图神经网络(GNN),以预测候选移动对目标函数的影响。训练好的模型随后被集成到禁忌搜索算法中,根据预测质量对候选变换进行排序,从而减少代价高昂的目标函数评估次数,同时保持对搜索空间的有效探索。在合成基准实例上的实验结果表明,所提出的学习辅助禁忌搜索显著减少了计算时间,并且能够持续产生比标准算法更高质量的解。这些发现凸显了通过利用搜索轨迹中蕴含的隐性知识将机器学习与元启发式算法相结合的潜力,为大规模网络设计问题的高效求解方法开辟了道路。
cs.AI / 20 / 2608.28628

CDEP Agent: Connecting Meteorologically Detected Temporal Compound Events to Real-World Documentary Evidence

CDEP Agent:将气象探测到的时间复合事件与现实世界的文献记录证据相连接
Li, Zhuoran, Kong, Weiyi, Zhang, Boer
Abstract
Compound drought-to-extreme-precipitation (CDEP) events are recognized in climate science as a growing driver of extreme impact, but whether this recognition carries over into real-world early warning and post-event documentation is unknown, so a meteorologically real CDEP event may pass with neither advance warning nor any later record. Here we present CDEP Agent, an auditable LLM-agent framework that tests this mismatch directly by linking CDEP candidates detected from meteorological reanalysis to real-world hazard and impact evidence across sources with different spatial scales, temporal resolutions, and reporting conventions. Using California as a case study, we identify 408 candidate CDEP events from ERA5 observations during 2021-2025 and evaluate each against the U.S. Drought Monitor, NOAA Storm Events, and public webpages along five dimensions: antecedent drought, extreme rainfall, local impact, hazard-impact attribution, and explicit drought-to-rainfall linkage. Only 34.3% of candidates are corroborated on both hazard components, and just 1.5% are ever explicitly linked to their antecedent drought, indicating that most meteorologically detected CDEP events go undocumented and their compound nature almost never enters the record at all. Our framework gives climate scientists a way to test physical event definitions against what actually gets documented, and gives social scientists, economists, and disaster-response agencies a provenance-linked evidence base for compound events that current warning and reporting systems largely fail to capture.
Chinese Translation
由干旱转为极端降水(Compound Drought-to-Extreme-Precipitation, CDEP)的复合事件在气候科学中被认为是日益增长的极端影响驱动因素,但这种认知是否延伸到现实世界的早期预警和事后记录中尚不清楚,因此一个气象学上真实发生的CDEP事件可能在既无事前预警也无任何事后记录的情况下悄然过去。本文提出CDEP Agent,一个可审计的大语言模型(LLM)智能体框架,通过将气象再分析数据中探测到的CDEP候选事件与来自不同空间尺度、时间分辨率和报道惯例的现实世界灾害与影响证据相链接,直接检验这种脱节现象。以加利福尼亚州为案例研究,我们从2021-2025年的ERA5观测数据中识别出408个候选CDEP事件,并依据五个维度对每个事件进行评估:前期干旱、极端降雨、局地影响、灾害-影响归因以及明确的干旱-降雨关联。仅有34.3%的候选事件在两个灾害要素上均得到证实,而仅有1.5%的事件曾被明确与其前期干旱相关联,这表明大多数气象探测到的CDEP事件未被记录在案,其复合性质几乎从未进入记录。我们的框架为气候科学家提供了一种方法,用以将物理事件定义与实际被记录的内容进行对照检验;同时为社会科学家、经济学家和灾害应对机构提供一个来源可追溯的复合事件证据基础,而现有预警和报告系统在很大程度上未能捕捉到这些事件。
cs.AI / 21 / 2608.28631

CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science

CrossAudit:一种面向智能体科学的基于Git的跨供应商审计闭环
Dong, Zhaohe, Chen, Yuhao
Abstract
An AI scientist should not grade its own homework. Yet in the systems we examined, the agent that reviews the work usually comes from the same model family as the agent that produced it, or at least from the same vendor. Model evaluators are known to favour their own generations. Whether models trained alike also share blind spots is a conjecture, not a settled finding, but if they do, the reviewer inherits the author's. The record of what was flagged and what was waved through often sits in platform logs that nobody outside can replay. We present CrossAudit, a protocol for supervising autonomous research pipelines. It rests on three commitments. Each increment of work is audited by an agent from a different vendor against a rulebook a human wrote and versioned. Reports, verdicts, disputes and rulings are git commits, so the supervision history can be re-read and cited; raw model exchanges are not yet part of that record. Scripted checks run before any model does. Advisory judgement never gates the pipeline: a model blocks only by citing a rule, and no model may waive a deterministic failure. Blockers that survive a bounded number of revision rounds go to a person. We state the protocol as eight invariants. We describe a reference implementation built from GitHub Actions and a few hundred lines of Python, and report a live deployment of a closely related variant in a computational-chemistry pipeline. We also ran a seeded-defect trial (30 increments, 43 seeded defects, one run per configuration). A cross-vendor audit of our own repository then voided its blinding. We adopt that audit's findings and report the corrected results. The trial shows that two vendors read the same rulebook differently. It does not show that either is better. The strongest evidence here is the committed, uncontrolled record of cross-vendor audits of this paper itself.
Chinese Translation
AI科学家不应给自己批改作业。然而在我们考察的系统中,评审工作的智能体通常与产生该工作的智能体来自同一模型家族,或至少来自同一供应商。众所周知,模型评估者会偏爱自己生成的结果。经相同方式训练的模型是否也共享相同的盲点尚属推测而非定论,但如果确实如此,评审者就会继承作者的盲点。标记了什么、放行了什么的记录往往存放在外部人员无法回放的平台日志中。我们提出CrossAudit,一种用于监督自主研究流程的协议。它基于三项承诺:每一份增量工作由来自不同供应商的智能体,依据一份由人类编写并版本化的规则手册进行审计;报告、裁定、争议与裁决均以git提交的形式存在,因此监督历史可被重新阅读和引用(原始模型交互记录尚不包含在该记录中);脚本化检查在任何模型介入之前运行。咨询性判断绝不作为流程的关卡:模型只能通过引用某条规则来阻塞流程,且任何模型都不得豁免确定性失败。在有限轮次的修订后仍然存在的阻塞项将交由人工处理。我们将该协议表述为八项不变式。我们描述了一个基于GitHub Actions和几百行Python代码构建的参考实现,并报告了一个密切相关变体在计算化学流程中的实际部署。我们还进行了一项植入缺陷试验(30个增量、43个植入缺陷,每种配置运行一次)。随后,对我们自身仓库的跨供应商审计导致该试验的盲测失效。我们采纳了该审计的发现并报告修正后的结果。该试验表明两家供应商对同一份规则手册的解读存在差异,但并不能证明哪家更优。此处最有力的证据,是对本文本身进行跨供应商审计所留下的已提交、未受控的记录。
cs.AI / 22 / 2608.28632

AutoScientist-Quant: Self-Evolving Coding Agents for Automatic Research in Quantitative Investment

AutoScientist-Quant:面向量化投资自动研究的自进化编码智能体
Li, Zongqian, Li, Yaoyiran, Guo, Yaohui, Zhang, Ming, Collier, Nigel, Ie, Eugene
Abstract
Large language model agents can discover alphas, yet current methods have three weaknesses. The search cannot adapt during the run, automation usually ends at alpha generation while library selection and model choice stay manual, and alpha discovery can read the test window through loop feedback or code problems. We present AutoScientist-Quant, a self evolving search process that regards quantitative research as one budgeted search problem. A single controller conditions every decision on the remaining budget, choosing at each round whether to improve, combine, pivot, or stop, which node to expand, how many alphas to generate, and how to retrieve past trajectories from the shared memory. The same core then selects from the library and tunes the model, closing the loop from hypothesis to deployable strategy. We also review the evaluation pipeline reused from prior work, fix two lookahead problems, and keep the feedback window disjoint from the held out test window, so every comparison tests true generalization. On CSI universes, the framework attains the best value of nearly every metric in every setting, and these conclusions hold across several backbones and markets.
Chinese Translation
大语言模型智能体能够挖掘Alpha因子,但现有方法存在三个不足:搜索过程无法在运行中自适应调整;自动化通常止步于Alpha生成,而因子库选择与模型选择仍需人工完成;Alpha挖掘过程可能通过循环反馈或代码问题窥探测试窗口。我们提出AutoScientist-Quant,一种自进化的搜索过程,将量化研究视为一个受预算约束的搜索问题。单一控制器基于剩余预算对每个决策进行条件化,在每一轮决定是改进、组合、转向还是停止,选择扩展哪个节点、生成多少个Alpha,以及如何从共享记忆中检索历史轨迹。同一核心随后完成因子库选择与模型调优,形成从假设到可部署策略的闭环。我们还复核了沿自先前工作的评估流程,修复了两处前视偏差问题,并保证反馈窗口与留出测试窗口互不相交,从而使每次比较都检验真正的泛化能力。在中证股票池上,该框架在几乎每种设定下的几乎所有指标上均取得最优值,且这些结论在多种骨干模型和市场中均成立。
cs.AI / 23 / 2608.28637

AI Scientist Mission Control (AIMC): Visual Analytics for Human Oversight of Autonomous Scientific Discovery

AI科学家任务控制中心(AIMC):用于人类监督自主科学发现的可视分析
Pal, Rikathi, Mueller, Klaus
Abstract
Autonomous scientific discovery systems can generate large numbers of research ideas, experiments, and manuscripts with minimal human intervention. As these systems become increasingly capable, scientists require effective mechanisms to monitor output quality, identify recurring failure modes, understand research evolution, and prioritize promising discoveries for review. We present AIMC, a visual analytics framework for human oversight of autonomous scientific discovery. AIMC combines semantic embeddings, automated weakness extraction, temporal analysis, and interactive visualizations to support the exploration of AI-generated research artifacts. We demonstrate the framework through a case study of the papers generated by an autonomous AI Scientist (FARS), together with their associated review feedback. Our analysis reveals recurring methodological weaknesses, evolving research themes, domain-specific differences in quality, and a small set of highly novel papers that warrant deeper human inspection. These findings illustrate how visual analytics can support transparency, diagnosis, and human AI collaboration in emerging autonomous scientific discovery workflows.
Chinese Translation
自主科学发现系统能够在极少人工干预的情况下生成大量研究想法、实验和论文手稿。随着这些系统日益强大,科学家需要有效的机制来监控输出质量、识别反复出现的失败模式、理解研究演化过程,并优先审查有前景的发现。我们提出了AIMC,一个用于人类监督自主科学发现的可视分析框架。AIMC结合了语义嵌入、自动化弱点提取、时序分析和交互式可视化,以支持对AI生成研究产物的探索。我们通过对自主AI科学家(FARS)所生成论文及其相关审稿反馈的案例研究来展示该框架。我们的分析揭示了反复出现的方法学弱点、不断演进的研究主题、不同领域间的质量差异,以及少量需要深入人工审查的高度新颖的论文。这些发现表明,可视分析能够在新兴的自主科学发现工作流中支持透明性、问题诊断和人机协作。
cs.AI / 24 / 2608.28638

Self-Evolving Skills via Surrogate-Guided Solve-and-Reproduce

基于代理模型引导的求解-复现机制的技能自我进化
Liu, Jiale, Ren, Pinze, Xia, Yuqi, Wang, Huan, Zhao, Zhenlin, Dong, Siming
Abstract
Agent skills are portable packages of instructions and resources an agent consults at deployment. Self-evolving them fails in two ways today. First, skills evolved from scratch underperform human-curated ones and, on a weak model, using no skill at all. Second, an evolution-time pass records one lucky trajectory that a fresh stochastic agent often fails to reproduce at deployment. We present reSolve, a per-task, oracle-in-the-loop framework built on three components. It decouples interactive solving from a self-contained deliverable that is independently re-executed in a fresh container, a protocol we call solve-and-reproduce. It enhances the sparse reward signal with a surrogate verifier that cannot access hidden tests or reference answers. It then runs verifier-guided beam search over a solution-construction graph. Within a fixed harness, a cheap model self-evolves skills that reach $74.9\%$ mean-of-3, $+14.8$ points over the $60.1\%$ human-curated baseline, exceeding the strongest official curated-skill result ($67.3\%$, GPT-5.5/OpenHands). We also report observed failure cases and domain-level results, including performance on the 14 Natural Science tasks, to clarify when the approach does and does not help.
Chinese Translation
智能体技能(agent skills)是智能体在部署时所参考的便携式指令与资源包。当前的技能自我进化存在两方面问题:其一,从零进化出的技能表现不及人工精选的技能,在较弱模型上甚至不如完全不使用技能;其二,进化阶段记录的是一条“幸运”的轨迹,全新的随机性智能体在部署时往往无法复现。我们提出了 reSolve,一个按任务执行、以“预言机在环”(oracle-in-the-loop)的框架,其建立在三个组件之上。它将交互式求解与一个自包含的交付物解耦,该交付物会在全新容器中独立重新执行,我们将这一协议称为“求解-复现”(solve-and-reproduce)。它借助一个无法访问隐藏测试或参考答案的代理验证器(surrogate verifier)来增强稀疏的奖励信号,随后在解构建图上运行验证器引导的束搜索(beam search)。在固定测试环境下,一个廉价模型自我进化的技能达到了74.9%的三次平均成绩,较60.1%的人工精选基线提升了14.8个百分点,超过了最强的官方精选技能结果(67.3%,GPT-5.5/OpenHands)。我们还报告了观察到的失败案例以及领域级结果,包括在14个自然科学任务上的表现,以阐明该方法在何时有效、何时无效。
cs.AI / 25 / 2608.28639

Reward-Oracle MCTS for Formal Theorem Proving: Sample-Efficient Search and the Need for Kernel-Level Proof Auditing

用于形式化定理证明的奖励-预言机蒙特卡洛树搜索:样本高效的搜索与内核级证明审计的必要性
Vamshi, Bodla Krishna, Yang, Haizhao
Abstract
Formal theorem proving with large language models remains challenging due to the difficulty of navigating large proof search spaces efficiently. Existing tree search approaches either feed verbose compiler error messages directly into the generation context, increasing context usage during search, or employ non-standard evaluation protocols that prevent direct comparison with established baselines. We propose a three-role Monte Carlo Tree Search (MCTS) framework that treats the Lean 4 compiler purely as a reward oracle using compiler output as a scalar signal for UCB-guided tree updates without feeding error content into the generation context. Our framework decomposes proof search into three roles: a generator for proof attempts, a decomposer for subgoal decomposition, and a critic for subgoal quality evaluation. We evaluate across 4 benchmarks spanning competition mathematics and physics (MiniF2F, PutnamBench, LeanPhysBench, PhysLeandata) with three prover models at standard proof attempt budgets (PAB@16 to PAB@256). Our method achieves 87.1\% on MiniF2F with Goedel-Prover-V2-8B at PAB@256 and solves 26/659 PutnamBench problems at PAB@32 surpassing base sampling 18/659 at same proof attempt budget. Through an exhaustive axiom-level audit of every compiled proof, we further identify reward hacking in search-based theorem proving: DeepSeek-Prover-V2-7B produces proofs on PutnamBench that pass compilation and the standard sorry-token scan while depending on sorryAx. The audit removes 4 and 8 such proofs from whole-proof sampling at PAB@32 and PAB@128, and 11 and 19 from MCTS. We do not attribute these counts to the search procedure; we report them to establish that kernel-level auditing is necessary for compiler-verified evaluation.
Chinese Translation
基于大语言模型的形式化定理证明仍然具有挑战性,原因在于难以高效地在庞大的证明搜索空间中进行导航。现有的树搜索方法要么将冗长的编译器错误信息直接输入生成上下文,从而增加搜索期间的上下文占用;要么采用非标准的评估协议,导致无法与既有基线进行直接比较。我们提出一个三角色的蒙特卡洛树搜索(MCTS)框架,该框架将 Lean 4 编译器纯粹视为奖励预言机(reward oracle),把编译器输出作为标量信号用于 UCB 引导的树更新,而不会将错误内容输入生成上下文。我们的框架将证明搜索分解为三个角色:用于生成证明尝试的生成器(generator)、用于子目标分解的分解器(decomposer)以及用于评估子目标质量的评估器(critic)。我们在涵盖竞赛数学与物理的 4 个基准(MiniF2F、PutnamBench、LeanPhysBench、PhysLeandata)上,使用三个证明器模型,在标准证明尝试预算(PAB@16 至 PAB@256)下进行评估。我们的方法在 PAB@256 下使用 Goedel-Prover-V2-8B 在 MiniF2F 上达到 87.1%,并在 PAB@32 下解决 PutnamBench 659 道题中的 26 道,在相同证明尝试预算下超过了基础采样的 18/659。通过对每个已编译证明进行详尽的公理级审计,我们进一步识别了基于搜索的定理证明中的奖励作弊(reward hacking)现象:DeepSeek-Prover-V2-7B 在 PutnamBench 上生成的证明虽然能够通过编译以及标准的 sorry-token 扫描,却依赖于 sorryAx。该审计在 PAB@32 和 PAB@128 下分别从整体证明采样中剔除了 4 和 8 个此类证明,从 MCTS 中剔除了 11 和 19 个。我们并不将这些数量归因于搜索过程本身;我们报告这些结果是为了说明,在编译器验证的评估中,内核级审计是必要的。
cs.AI / 26 / 2608.28642

From Extraction to Governed Memory: Multi-Agent Knowledge Graph Construction with Domain-Expert Review

从抽取到受治理的记忆:结合领域专家审查的多智能体知识图谱构建
Bykampadi, Pranav, Mokaria, Neel, Narayan, Vishesh, Wajid, Faizan, Agrawala, Ashok
Abstract
Knowledge graphs used by agentic systems are often treated as flat stores of extracted triples, with little record of who owns a fact, why it was admitted, or how it should be used downstream. We argue that reliable agentic knowledge systems require governance as an essential component of graph construction to bridge this gap. We propose MAGG, a principled multi-agent framework for constructing Governed Knowledge Graphs that introduces explicit governance decisions for reliable and trustworthy knowledge sharing. A domain classifier first induces entity and relation types directly from document content, enabling operation in open-world settings without fixed schemas. Candidate triples are assigned to domain owners, reviewed against supporting evidence, admitted through governance decisions, and stored with audit metadata. The same ownership structure is reused during question answering, where queries are routed to domain-specific graph experts rather than answered through undifferentiated retrieval. Our evaluation demonstrates MAGG's effectiveness: On SciERC, MAGG improves strict triple F1 by 47% and mapped triple F1 by 51% over flat insertion. A blinded review of 120 triples finds governed-only triples more often source-supported than flat-only ones, and revised triples supported in 100% of cases. Finally, on MuSiQue, MAGG outperforms Microsoft GraphRAG by 9.0 exact-match points and 11.2 token-F1 points.
Chinese Translation
智能体系统所使用的知识图谱通常被视为抽取三元组的扁平化存储,几乎不记录事实的归属者、其被采纳的原因或其在下游应如何使用。我们认为,可靠的智能体知识系统需要将治理作为图谱构建的核心组成部分,以弥合这一差距。我们提出了 MAGG,一个原则性的多智能体框架,用于构建受治理的知识图谱(Governed Knowledge Graphs),通过引入显式的治理决策来实现可靠且可信的知识共享。领域分类器首先直接从文档内容中归纳实体和关系类型,从而能够在没有固定模式(schema)的情况下于开放世界环境中运行。候选三元组被分配给领域所有者,对照支持证据进行审查,通过治理决策后被采纳,并连同审计元数据一起存储。在问答过程中,同样的所有权结构被复用:查询被路由到特定领域的图谱专家,而非通过无差别的检索来回答。我们的评估验证了 MAGG 的有效性:在 SciERC 数据集上,与扁平化插入相比,MAGG 将严格三元组 F1 提升了 47%,映射三元组 F1 提升了 51%。对 120 个三元组的盲评显示,仅通过治理采纳的三元组比仅通过扁平化方式采纳的三元组更容易获得来源支持,且经修订的三元组在 100% 的案例中获得支持。最后,在 MuSiQue 数据集上,MAGG 比 Microsoft GraphRAG 高出 9.0 个精确匹配点和 11.2 个 token-F1 点。
cs.AI / 27 / 2608.28646

BiasMix-Finance: Post-Generation KYC Guardrails for LLM Portfolio Advice

BiasMix-Finance:面向大语言模型投资组合建议的生成后KYC防护机制
Kukreja, Gaurav, Kukreja, Parul, Abraar, Mohammed, Dandekar, Raj, Dandekar, Rajat, Panat, Sreedath
Abstract
Large language models (LLMs) can generate plausible-sounding ETF portfolios while silently violating basic KYC-style constraints on risk, fees, and diversification. This is especially problematic in agentic multi-turn advisory systems, where each draft recommendation can become an action unless guarded by an auditable enforcement layer. We study a model-agnostic, asset-agnostic post-generation guardrail pipeline: (i) enforce a strict JSON allocation schema, (ii) validate allocations against numeric caps, and (iii) when violations occur, deterministically project the output to the nearest feasible portfolio via a convex quadratic program (QCQP). We introduce BiasMix-Finance (Mini), a compact stress-test benchmark for constrained decision-making under biased LLM generations, with a 16-ETF universe, three investor profiles, and eight bias prompts. Across three models and three inference modes (direct, critique, self-consistency), first-pass generations violate at least one cap in 47.6-85.7% of test cases (67.2% pooled), but the convex projection layer reduces final feasibility violations to 0% while requiring only a small correction distance (test pooled median D=||w*-w0||_2=0.066), indicating that the guardrail typically preserves the intent of the original allocation. We report violation rates and correction distances with confidence intervals, and paired model comparisons with multiple-testing correction. To support reproducibility, we release the dataset, prompts, caps, and code in our public GitHub repository.
Chinese Translation
大语言模型(LLM)能够生成听起来合理的ETF投资组合,同时却在无形中违反有关风险、费用和分散化的基本KYC类约束。这一问题在智能体式多轮对话咨询系统中尤为突出,因为每份草拟的建议都可能直接转化为操作,除非有可审计的强制执行层加以防护。我们研究了一种与模型无关、与资产无关的生成后防护流水线:(i) 强制执行严格的JSON配置格式;(ii) 依据数值上限校验配置;(iii) 当出现违规时,通过凸二次规划(QCQP)确定性地将输出投影至最近的可行动投资组合。我们提出了BiasMix-Finance(Mini),一个用于评估有偏LLM生成下受约束决策能力的紧凑压力测试基准,包含16只ETF的资产池、三类投资者画像和八种偏差提示词。在三个模型和三种推理模式(直接生成、批判式、自洽性)下,首次生成的结果中有47.6%至85.7%的测试用例(汇总为67.2%)至少违反一个上限,但凸投影层将最终可行性违规率降至0%,同时仅需较小的修正距离(测试汇总中位数 D=||w*-w0||_2=0.066),表明该防护机制通常能保留原始配置的意图。我们报告了带置信区间的违规率和修正距离,并进行了经多重检验校正的配对模型比较。为支持可复现性,我们在公开的GitHub仓库中发布了数据集、提示词、上限设定和代码。
cs.AI / 28 / 2608.28647

Self-Specialized Teachers for Domain Post-Training

面向领域后训练的自专业化教师模型
Li, Yifei, Xu, Rongman, Zhang, Lingling, Huang, Muye, Ma, Zihan, Liu, Jiashuai, Yan, Hang, Wang, Heng
Abstract
Target-only post-training can improve performance in a specialized domain while degrading behaviors that a general-purpose base model acquired before adaptation. We study this problem when target-domain data are available but a representative replay corpus is not. We propose self-specialized teacher distillation (SSTD), a two-stage procedure that first trains a copy of the base model into a domain teacher, then distills its token distribution to a student on prefixes sampled from the student itself. Teacher training combines standard target supervision with base-aware key-token weighting and distribution alignment to the frozen base model; on-policy distillation then places domain feedback on states the student can encounter at inference time. On financial numerical reasoning, medical question answering, and legal holding identification, SSTD retains much of the target improvement of direct fine-tuning while improving the mean score on the evaluated general suite by 4.8--5.0 points at the reported operating point. The pattern persists across Qwen3 sizes and on Gemma backbones. SSTD requires neither an external teacher nor general replay data.
Chinese Translation
仅针对目标领域的后训练可以在专业领域提升性能,但同时会损害通用基础模型在适配之前所习得的能力。我们研究在目标领域数据可用但缺乏代表性回放语料库时的这一问题。我们提出自专业化教师蒸馏(Self-Specialized Teacher Distillation, SSTD),这是一种两阶段方法:首先将基础模型的一个副本训练成领域教师模型,然后在从学生模型自身采样得到的前缀上,将教师的词元分布蒸馏给学生模型。教师模型的训练将标准的目标领域监督与基于基础模型的 ключ——准确地说,是与冻结基础模型的分布对齐以及关键token加权相结合;随后的在线策略蒸馏则将领域反馈施加于学生模型在推理时可能遇到的状态上。在金融数值推理、医疗问答和法律裁判要旨识别任务上,SSTD在保留直接微调大部分目标领域性能提升的同时,在所报告的工作点上将所评估的通用能力测试集的平均分数提升4.8至5.0分。该模式在不同规模的Qwen3模型以及Gemma骨干模型上均持续成立。SSTD既不需要外部教师模型,也不需要通用回放数据。
cs.AI / 29 / 2608.28648

How Language Models Choose Sides: Internal Representations of Instruction Hierarchy

语言模型如何选边站:指令层级的内部表征
Balp-Straffon, Enrique, Hsu, Chih-Hao, Gadhvi, Rushiraj, Dev, Sunishchal, McDougall, Callum Stuart, Mujumdar, Anusha
Abstract
We study how instruction-tuned LLMs arbitrate direct conflicts between system and user instructions. We introduce a benchmark of 41 paired constraints with deterministic verifiers and evaluate eight models under matched baseline, conflict, and same-channel control conditions. Behaviourally, the models split into three regimes by System Authority Delta: hierarchy-respecting models use the system channel as an authority signal, anti-hierarchy models follow the system less often than their same-channel baseline predicts, and no-effect models show little channel sensitivity. Llama-3.1-8B is the strongest anti-hierarchy case in our suite, following the system in only 0.10 of conflict trials. We use this behavioural failure case to ask whether user-preferring arbitration reflects the absence of an internal conflictresolution signal. It does not: on Llama-3.1-8B, the conflict outcome is linearly decodable from residual-stream activations at 0.97 balanced accuracy, 17 percentage points above a metadata-only baseline, with analogous signals on Qwen2.5-7B and gpt-oss-20b. Steering with a layer-12 mean of four per-conflict logistic-regression directions raises genuine system compliance from 0.132 to 0.530, while directions selected mainly for pooled separability steer poorly. User-preferring conflict resolution can therefore coexist with a readable internal arbitration signal, and successful intervention depends on the geometry of the readout rather than probe accuracy alone
Chinese Translation
我们研究了经过指令微调的大语言模型(LLM)如何在系统指令与用户指令发生直接冲突时进行仲裁。我们构建了一个包含41对约束的基准,并配有确定性验证器,在匹配的基线、冲突和同通道控制三种条件下评估了八个模型。在行为层面,模型根据系统权威差值(System Authority Delta)分为三类模式:尊重层级的模型将系统通道作为权威信号;反层级模型遵循系统指令的频率低于其同通道基线的预测;无效应模型则对通道几乎不敏感。Llama-3.1-8B是我们测试套件中最典型的反层级案例,在冲突实验中仅有0.10的比例遵循系统指令。我们利用这一行为失败案例探讨:偏好用户的仲裁是否反映了内部缺乏冲突解决信号。事实并非如此:在Llama-3.1-8B上,冲突结果可以从残差流激活中以0.97的平衡准确率被线性解码,比仅基于元数据的基线高出17个百分点,且在Qwen2.5-7B和gpt-oss-20b上也存在类似的信号。使用基于四个逐冲突逻辑回归方向在12层计算的均值进行引导,可将真实的系统遵循率从0.132提升至0.530,而主要依据聚合可分性选择的方向引导效果较差。因此,偏好用户的冲突解决可以与可读的内部仲裁信号共存,而成功的干预取决于读出的几何结构,而非仅取决于探测准确率。
cs.AI / 30 / 2608.28652

A Generalized Optimization Engine (GOE) for Edge AI Inference Acceleration

一种用于边缘AI推理加速的广义优化引擎(GOE)
Dasari, Venkat R., Adams, Jakob A., Mishra, Vinod K., Jalaian, Brian
Abstract
Artificial intelligence (AI) models have demonstrated remarkable capabilities across various domains, yet their widespread deployment is impeded by significant computational costs, particularly on resource-constrained devices. This paper explores the theoretical underpinnings of various AI model optimization techniques, algorithms, and abstractions, discussing their potential to reduce computational complexity, memory footprint, latency, and power consumption. Furthermore, we propose a comprehensive hardware (HW) and model-agnostic generalized optimization architecture that integrates these techniques for improved efficiency. Our study underscores the critical role of such a generalized optimization system in preparing model deployment over resource-constrained heterogeneous hardware in a tactical environment. As a concrete demonstration, we show that GOE-compressed language models deploy and run on a GPU-less edge CPU, and that the choice of compression method, not merely its nominal bit-width, determines whether task accuracy survives deployment.
Chinese Translation
人工智能(AI)模型已在众多领域展现出卓越的能力,但其高昂的计算成本阻碍了广泛应用,尤其是在资源受限的设备上。本文探讨了多种AI模型优化技术、算法和抽象的理论基础,讨论了它们在降低计算复杂度、内存占用、延迟和功耗方面的潜力。此外,我们提出了一种全面的、与硬件(HW)和模型无关的广义优化架构,将这些技术整合起来以提升效率。我们的研究强调了此类广义优化系统在战术环境中为模型部署到资源受限的异构硬件上所起的关键作用。作为具体演示,我们展示了经GOE压缩的语言模型可在无GPU的边缘CPU上部署和运行,并且任务精度能否在部署后得以保持,取决于压缩方法的选择,而不仅仅是其名义位宽。
cs.AI / 31 / 2608.28662

FRAC-MAS: A Safe and Explainable Multi-Agent System for Fracture Diagnosis

FRAC-MAS:一种用于骨折诊断的安全且可解释的多智能体系统
Iyer, Hardik, Bhathawala, Tirath, Panchal, Mihir, Chen, Ying-Jung, Bhowmick, Kiran, Sonawane, Pankaj, Narvekar, Meera
Abstract
Fracture detection and its clinical interpretability see notable improvements when deep vision models are integrated with agentic AI architectures. While deep learning models achieve high diagnostic performance, their black-box nature limits clinical adoption. We propose FRAC-MAS, an agentic AI system for automated, explainable, and safe bone fracture detection. The framework combines a stacked ensemble of four vision models with conformal prediction to produce statistically grounded differential diagnoses, while a multi-agent workflow performs independent verification, retrieves clinical guidelines, and generates patient-friendly reports. A pipeline-depth ablation study confirms that our multi-agent critic triages 86.6% of cases into a high-confidence auto-confirmed cohort while escalating uncertain cases, outperforming a single-agent baseline. Patient preference studies against Llama, MedGemma, and Gemini further demonstrate significantly more comprehensible clinical reports. These results suggest that integrating multi-agent critics with conformal guarantees enables safer radiology triage while preserving clinician oversight. More broadly, FRAC-MAS demonstrates how cooperative agentic architectures can serve as auditable, human-in-the-loop decision support systems for safety-critical healthcare. Our code is available at https://github.com/hardik1712/FRAC-MAS, and the website is available at https://frac-mas.vercel.app.
Chinese Translation
当深度视觉模型与智能体AI架构相结合时,骨折检测及其临床可解释性得到显著提升。尽管深度学习模型具有很高的诊断性能,但其黑盒特性限制了临床应用。我们提出FRAC-MAS,一个用于自动化、可解释且安全的骨折检测的智能体AI系统。该框架将四个视觉模型的堆叠集成与共形预测相结合,以产生具有统计学依据的鉴别诊断,同时多智能体工作流执行独立验证、检索临床指南并生成患者易于理解的报告。流水线深度的消融研究证实,我们的多智能体评估器将86.6%的病例分流到高置信度的自动确认队列中,同时将不确定病例上报,优于单智能体基线。与Llama、MedGemma和Gemini的患者偏好研究进一步表明,本系统生成的临床报告具有显著更好的可理解性。这些结果表明,将多智能体评估器与共形保证相结合,能够在保留临床医生监督的同时实现更安全的放射科分诊。更广泛地说,FRAC-MAS展示了协作式智能体架构如何能够作为面向安全关键医疗的、可审计的人机协同决策支持系统。我们的代码可在 https://github.com/hardik1712/FRAC-MAS 获取,网站可在 https://frac-mas.vercel.app 访问。
cs.AI / 32 / 2608.28704

ORDDAR: Observation-Driven Reasoning for Distortion-Resilient Decision, Action, and Cognitive Recovery

ORDDAR:面向失真韧性决策、行动与认知恢复的观察驱动推理
Kar, Deblina, Nawalgaria, Anant, Mandal, Shyamal Kumar Das
Abstract
AI agents increasingly perform long-term reasoning, planning, tool use, memory integration, and autonomous decision making, yet erroneous intermediate states can propagate and cause inconsistent decisions and unreliable outputs. Existing reasoning approaches mainly rely on iterative planning, self-reflection, augmented memory, or verification, but rarely localize and selectively repair faulty reasoning. We present ORDDAR (Observation-Driven Reasoning for Distortion-Resilient Decision, Action, and Cognitive Recovery), a reasoning framework that models reasoning as cognitive state transitions, detects localized distortions, retrieves related reasoning from prior experiences, and repairs only the affected states. ORDDAR therefore performs recovery at the local reasoning-transition level rather than regenerating the complete trajectory. Experiments across mathematical, commonsense, multi-hop, and clinical reasoning benchmarks demonstrate improved reasoning quality, recovery ability, and interpretability over multiple evaluated reasoning baselines.
Chinese Translation
AI智能体越来越多地执行长期推理、规划、工具使用、记忆整合和自主决策,然而错误的中间状态可能会传播并导致不一致的决策和不可靠的输出。现有的推理方法主要依赖于迭代规划、自我反思、增强记忆或验证,但很少对错误推理进行定位和选择性修复。我们提出了ORDDAR(Observation-Driven Reasoning for Distortion-Resilient Decision, Action, and Cognitive Recovery,面向失真韧性决策、行动与认知恢复的观察驱动推理),这是一个将推理建模为认知状态转换的推理框架,能够检测局部失真,从先前经验中检索相关推理内容,并仅修复受影响的状态。因此,ORDDAR在局部推理转换层面执行恢复,而不是重新生成完整的轨迹。在数学、常识、多跳和临床推理基准上的实验表明,与多个评估的推理基线相比,ORDDAR在推理质量、恢复能力和可解释性方面均有提升。
cs.AI / 33 / 2608.28725

Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification

超越标准答案:面向步骤级数学验证的大语言模型鲁棒性评估
Mazdarani, Fateme, Toxtli, Carlos
Abstract
Large language models (LLMs) are increasingly used as graders, verifiers, and process auditors, but most mathematical evaluations still emphasize final-answer accuracy. This can obscure whether a model can verify a non-canonical but valid solution trace. We introduce a controlled linear-equation benchmark for evaluating LLMs in the evaluator role. Each instance asks the model to judge final-answer correctness, step-level trace correctness, and the first incorrect step. Our evaluation of state-of-the-art open LLMs reveals a significant robustness gap: models that accurately evaluate canonical solutions often fail when presented with perturbed but logically equivalent variants. Across GPT-OSS 20B, Qwen3-14B, and Phi-4-Reasoning, base models perform well on canonical traces but degrade substantially on perturbed traces, especially for error localization. On valid perturbed traces, base-model false-rejection rates reach 75.6-85.3%, showing strong sensitivity to canonical solution form. Supervised fine-tuning, distillation, and test-time compute improve robustness in some settings, but gains are model dependent and can trade off against canonical performance. The results show that reliable process-level verification remains challenging, and evaluator robustness should be measured separately from solver accuracy, even in a simple algebraic domain with exact ground truth.
Chinese Translation
大语言模型(LLM)正日益被用作评分者、验证者和过程审计者,但大多数数学评估仍然侧重于最终答案的准确性。这可能会掩盖模型是否能够验证非标准但有效的解题过程。我们引入了一个受控的线性方程基准,用于评估LLM在评估者角色中的表现。每个实例要求模型判断最终答案的正确性、步骤级解题过程的正确性以及第一个错误步骤。我们对最先进的开源LLM的评估揭示了一个显著的鲁棒性差距:能够准确评估标准解题过程的模型,在面对经过扰动但逻辑等价的变体时往往表现不佳。在GPT-OSS 20B、Qwen3-14B和Phi-4-Reasoning上,基础模型在标准解题过程上表现良好,但在扰动解题过程上性能显著下降,尤其是在错误定位方面。在有效的扰动解题过程上,基础模型的误拒率高达75.6%-85.3%,表明其对标准解题形式具有很强的敏感性。有监督微调、蒸馏和测试时计算(test-time compute)在某些设置下能提升鲁棒性,但其收益依赖于具体模型,并可能以牺牲标准解题过程上的表现为代价。结果表明,可靠的过程级验证仍然具有挑战性,评估者的鲁棒性应与解题者的准确性分开衡量,即使在一个具有精确标准答案的简单代数领域中亦是如此。
cs.AI / 34 / 2608.28726

Pro-Router: Token-Aware Progressive Model Routing with Adaptive Edge-Cloud Collaboration for Efficient Multimodal LLM Inference

Pro-Router:面向高效多模态大语言模型推理的令牌感知渐进式模型路由与自适应端云协同方法
Gui, Xinyuan, Wang, Shaowen, Sun, Sheng, Wang, Zijian, Yu, Zishu, Yang, Zheming
Abstract
The remarkable performance of multimodal large language models (MLLMs) comes at the cost of substantial computational overhead, posing significant challenges to real-time deployment and cost effectiveness. Existing model routing approaches either decide from coarse request-level features alone or spend one or several extra language model passes to inspect the generated response, leaving the token-level uncertainty signals that emerge during generation unused. To address these limitations, we propose Pro-Router, a token-aware progressive model routing method with adaptive edge-cloud collaboration for efficient multimodal LLM inference. Pro-Router employs a two-stage progressive decision mechanism. First, a lightweight prompt pre-scorer module performs rapid pre-screening before token generation begins, guiding apparently simple requests to small models. Second, a token-aware verifier reads the sampling probability distribution of each token the small model generates, estimating the model's confidence in its own output to determine, per request, whether the answer ships or escalates to the cloud-based high-precision model. Furthermore, we design an adaptive edge-cloud serving pipeline that sizes every dispatch to each device's measured service rate, so both the edge and the cloud tiers stay fully utilized without manual parameter tuning and are not impacted by the network latency. Extensive experiments on multiple multimodal benchmark datasets and models demonstrate the effectiveness of Pro-Router. Compared to other methods, it achieves the highest routing accuracy and improves routing speed by more than 10x. Its serving pipeline also reaches more than 75% higher end-to-end throughput than the existing model routing pipeline. Our code is available at https://github.com/xinyuangui2/pro-router.
Chinese Translation
多模态大语言模型(MLLM)的卓越性能以巨大的计算开销为代价,给实时部署和成本效益带来了严峻挑战。现有的模型路由方法要么仅基于粗粒度的请求级特征进行决策,要么需要额外执行一次或多次语言模型前向计算来检查生成的响应,导致生成过程中出现的令牌级不确定性信号未被利用。针对这些局限性,我们提出了 Pro-Router,一种面向高效多模态大语言模型推理的令牌感知渐进式模型路由方法,结合自适应端云协同机制。Pro-Router 采用两阶段渐进式决策机制:首先,轻量级的提示预评分模块在令牌生成开始前进行快速预筛选,将明显简单的请求引导至小模型;其次,令牌感知验证器读取小模型生成的每个令牌的采样概率分布,估计模型对自身输出的置信度,从而针对每个请求判断答案是直接返回还是升级至云端高精度模型。此外,我们设计了自适应端云服务流水线,根据每个设备的实测服务速率调整每次分派任务的规模,使端侧和云侧均保持充分利用,无需手动调参,且不受网络延迟影响。在多个多模态基准数据集和模型上的大量实验证明了 Pro-Router 的有效性:与其他方法相比,它实现了最高的路由精度,路由速度提升超过 10 倍;其服务流水线的端到端吞吐量也比现有模型路由流水线高出 75% 以上。我们的代码可在 https://github.com/xinyuangui2/pro-router 获取。
cs.AI / 35 / 2608.28728

PermitGPT: A Unified Generative-AI Pipeline for Construction Hazard Forecasting, Permit Prediction, and Community Impact

PermitGPT:用于建筑危害预测、许可证预测与社区影响的统一生成式人工智能流水线
Ameen, Mohd Ruhul, Aktar, Farjana, Islam, Akif, Ope, Momen Khandoker, Miah, Abu Saleh Musa, Shin, Jungpil
Abstract
Urban construction governance requires early decisions that connect workplace safety, permitting requirements, and community impact, yet the relevant evidence is often scattered across separate municipal and regulatory data sources. This paper presents PermitGPT, a unified generative artificial intelligence framework for converting unstructured construction permit descriptions into structured decision-support outputs across three domains: safety hazard identification, permit requirement specification, and community impact assessment. To address data fragmentation, we spatially and temporally align records from the New York City Department of Buildings, Occupational Safety and Health Administration, and NYC 311 service requests, producing 90,000 structured prompt-response pairs derived through rule-based alignment and domain-informed spot checking. We fine-tune three open-weight language models using parameter-efficient adaptation and evaluate them on 2,833 held-out test cases. The results show complementary model behavior: Gemma-3-1B provides the most efficient inference at 3.07 samples per second with low memory usage, Llama-3.2-3B gives the highest lexical overlap for regulatory-style outputs with a BLEU score of 0.0091, and 4-bit Mistral-7B-Instruct-v0.3 achieves the strongest semantic alignment with a BERTScore-F1 of 0.7747. Because the task involves open-ended structured generation, low BLEU values are interpreted alongside semantic metrics and qualitative output structure rather than as standalone indicators of utility. Overall, PermitGPT provides an initial step toward AI-assisted construction governance while identifying directions for stronger task-level evaluation and real-world validation.
Chinese Translation
城市建筑治理需要将工作场所安全、许可要求和社区影响相互关联的早期决策,然而相关证据往往分散在彼此独立的市政和监管数据源中。本文提出PermitGPT,一个统一的生成式人工智能框架,用于将非结构化的建筑许可证描述转化为涵盖三个领域的结构化决策支持输出:安全隐患识别、许可证要求规范和社区影响评估。为解决数据碎片化问题,我们对纽约市楼宇局(New York City Department of Buildings)、职业安全与健康管理局(OSHA)和NYC 311服务请求的记录进行空间与时间对齐,通过基于规则的对齐和结合领域知识的抽样检查,生成了90,000条结构化的提示-响应配对。我们采用参数高效微调方法对三个开源权重语言模型进行微调,并在2,833个保留测试用例上进行评估。结果显示各模型表现互补:Gemma-3-1B在3.07样本/秒的速度下提供最高效的推理且内存占用低;Llama-3.2-3B在监管风格输出上词法重合度最高,BLEU分数为0.0091;4比特量化的Mistral-7B-Instruct-v0.3实现最强的语义对齐,BERTScore-F1达到0.7747。由于该任务涉及开放式的结构化生成,较低的BLEU值应结合语义指标和定性输出结构一起解读,而非作为实用性的独立指标。总体而言,PermitGPT为人工智能辅助的建筑治理迈出了初步一步,同时指出了更强任务级评估和真实世界验证的后续方向。
cs.AI / 36 / 2608.28791

Efficient Geothermal Well-Control Optimization via Diffusion-Surrogate Reinforcement Learning

基于扩散代理模型强化学习的地热井控高效优化
Dai, Ruimin, Chen, Guodong, Harsuko, Randy, Liu, Kunpeng, Nakata, Nori
Abstract
Real-time decision-making for enhanced geothermal systems (EGS) is challenging because long-term production periods involve high-dimensional control spaces and a large number of time-consuming high-fidelity hydrothermal simulations. Reinforcement learning provides a natural framework for state-dependent sequential control, but direct policy training with numerical simulators is computationally expensive. To address this issue, we propose a diffusion-surrogate guided reinforcement learning framework for long-horizon EGS well-control optimization. The reservoir temperature and pressure fields are used as system states, while injection rates are selected as control actions. A learned surrogate environment is constructed using conditional diffusion models to predict the evolution of reservoir temperature and pressure fields and a separate reward model to estimate the corresponding economic return. The surrogate environment is then integrated with Proximal Policy Optimization (PPO) for efficient policy training. Experiments on a fractured EGS benchmark show that the diffusion surrogate can accurately reproduce reservoir-state evolution over multiple control stages. The resulting surrogate-assisted PPO policy achieves competitive well-control performance compared with direct simulator-based PPO and existing optimization methods, while substantially reducing the dependence on expensive high-fidelity simulations. These results demonstrate the potential of diffusion-based surrogate environments for efficient reinforcement learning in geothermal well-control optimization.
Chinese Translation
增强型地热系统(EGS)的实时决策极具挑战性,因为长期生产过程涉及高维控制空间以及大量耗时的高保真水热模拟。强化学习为依赖于状态的序贯控制提供了天然的框架,但直接使用数值模拟器进行策略训练的计算成本高昂。为解决这一问题,我们提出了一种扩散代理模型引导的强化学习框架,用于长时程EGS井控优化。该框架以储层温度场和压力场作为系统状态,以注入速率作为控制动作。通过条件扩散模型构建学习型代理环境,用于预测储层温度场和压力场的演化,并采用独立的奖励模型来估计相应的经济收益。随后,将该代理环境与近端策略优化(Proximal Policy Optimization, PPO)相结合,以实现高效的策略训练。在一个裂缝性EGS基准算例上的实验表明,扩散代理模型能够在多个控制阶段准确再现储层状态的演化。与基于模拟器的直接PPO训练及现有优化方法相比,所得到的代理辅助PPO策略实现了具有竞争力的井控性能,同时显著降低了对昂贵高保真模拟的依赖。这些结果证明了基于扩散模型的代理环境在地热井控优化高效强化学习中的应用潜力。
cs.AI / 37 / 2608.28806

Enhancing SAE-based Steering via Neighbor Integrated Feature Selection

通过邻域集成特征选择增强基于SAE的模型引导
Liu, Yutian, Wang, Xu, Zou, Difan
Abstract
Sparse autoencoders (SAEs) disentangle model activations into interpretable features and are widely used for steering large language models. Most existing SAE-based steering methods select features by applying a top- filter based on statistical scores, assuming that higher-scoring features yield stronger steering effects. In this paper, we show that this assumption is often invalid, leading to suboptimal feature selection. Our analysis reveals that effective steering features may be distributed among representationally adjacent, semantically similar groups induced by feature splitting in SAEs. Within such groups, features may exhibit disparate statistical scores despite having comparable steering influence, causing score-based selection to overlook important features. Based on these observations, we propose \textsc{Neighbor Integrated Feature Selection} (\textsc{NIFS}), a plug-and-play strategy that leverages representation similarity to improve feature selection for steering. We evaluate \textsc{NIFS} across multiple SAE-based steering methods and tasks, and demonstrate consistent performance gains over conventional top-$k$ selection.
Chinese Translation
稀疏自编码器(SAE)将模型激活解耦为可解释的特征,并被广泛用于引导大语言模型。现有的大多数基于SAE的引导方法通过基于统计得分的top-k过滤来选择特征,并假设得分较高的特征会产生更强的引导效果。本文表明这一假设往往不成立,会导致次优的特征选择。我们的分析揭示,有效的引导特征可能分布在由SAE中的特征分裂所引起的表示相邻、语义相似的特征组中。在这些组内,尽管特征具有相当的引导影响力,其统计得分却可能差异悬殊,从而导致基于得分的选择方法忽略重要特征。基于这些观察,我们提出了邻域集成特征选择(Neighbor Integrated Feature Selection, NIFS),这是一种即插即用的策略,利用表示相似性来改进用于引导的特征选择。我们在多种基于SAE的引导方法和任务上评估了NIFS,结果表明其相对传统的top-k选择方法取得了一致的性能提升。
cs.AI / 38 / 2608.28809

Capability-Stratified Degradation in Ternary Language Models

三值语言模型的能力分层退化
Malik, Anirudh, Mehra, M Sparsh, Devan, Poojith
Abstract
Extreme low-bit inference offers a route toward smaller models and constrained deployment. Ternary language models restrict weights to $\{-1,0,+1\}$, approaching the limit of $\log_2 3 \approx 1.585$ bits/weight. The practical question for a pretrained model is not simply whether weights can be quantised but which capabilities survive and whether it remains useful for adaptation. We explore this by converting Qwen3.5-0.8B (752M parameters) to ternary weights using 72.4M tokens of quantisation-aware training (QAT). The resulting model, Cloe, is evaluated across 29 benchmarks, representation diagnostics, and downstream fine-tuning. The evidence shows non-uniform degradation. A linear probe recovers 43.76% of MMLU answers from the full-precision teacher's representations but only 26.19% from Cloe (near chance), indicating specialist factual information is lost. However, Cloe retains measurable performance on ten tasks, averaging 77.1% of teacher performance. Crucially, fine-tuning raises Cloe to 89.8% on SST-2 (95.6% of the matched teacher) and reaches 79.4% teacher retention on XSum. We attribute degradation to a combination of quantisation-induced information loss and incomplete recovery due to the limited QAT budget. We also highlight an evaluation pitfall: standard answer-letter scoring failed (Cloe emitted "A" on 98.6% of MMLU questions), necessitating continuation scoring. Ultimately, ternary conversion is unsuitable as a drop-in general replacement yet remains valuable as a compact substrate for task-specific models.
Chinese Translation
极端低比特推理为构建更小的模型以及在受限环境下部署提供了一条途径。三值(Ternary)语言模型将权重限制在 $\{-1,0,+1\}$ 之中,逼近每权重 $\log_2 3 \approx 1.585$ 比特的理论极限。对于预训练模型而言,实际的问题不仅仅是权重能否被量化,而是哪些能力得以保留,以及模型在经过适配后是否仍然有用。我们通过使用 7240 万(72.4M)词元进行量化感知训练(QAT),将 Qwen3.5-0.8B(7.52 亿参数)转换为三值权重来探索这一问题。所得模型 Cloe 在 29 个基准测试、表示诊断以及下游微调任务上进行了评估。证据表明退化是非均匀的:线性探针(linear probe)从全精度教师模型的表示中可恢复 43.76% 的 MMLU 答案,而从 Cloe 中仅能恢复 26.19%(接近随机水平),这表明专门的 factual 信息已丢失。然而,Cloe 在十个任务上仍保留了可测量的性能,平均达到教师模型性能的 77.1%。关键的是,微调使 Cloe 在 SST-2 上提升至 89.8%(达到相匹配教师模型的 95.6%),并在 XSum 上达到教师模型 79.4% 的性能保持率。我们将退化归因于量化引起的信息丢失以及有限的 QAT 训练预算所导致的恢复不完全。我们还指出一个评估陷阱:标准的答案字母评分方法失效(Cloe 在 98.6% 的 MMLU 问题上都输出 "A"),因此必须采用续写评分法。最终结论是:三值转换并不适合作为即插即用的通用替代方案,但作为面向特定任务模型的紧凑基础底座,它仍然具有价值。
cs.AI / 39 / 2608.28820

Explainable Artificial Intelligence (XAI) in Computational Pathology: Definitions, Taxonomy, and Recommendations

计算病理学中的可解释人工智能(XAI):定义、分类体系与建议
Innani, Shubham, You, Suhang, Shephard, Adam, Baheti, Bhakti, Ciompi, Francesco, Yeong, Joe, Rajpoot, Nasir, Feldman, Michael, Kammerer-Jacquet, Solene Florence, Makris, Dimitrios, Litjens, Geert, Martel, Anne L., Lipkova, Jana, Khademi, April, Bakas, Spyridon, SIG-CompPath, for the MICCAI
Abstract
Computational pathology (CompPath) is transforming medicine by leveraging artificial intelligence (AI) algorithms to support diagnosis, prognosis, and treatment prediction from gigapixel whole-slide images. Clinical adoption is progressing, but is constrained by concerns about safety, accountability, and regulatory oversight in high-stakes clinical environments. Explainable AI (XAI) systems hold promise for building trust and enabling verification, yet the literature remains fragmented due to inconsistent terminology, overlapping methodological families, ad hoc validation, and current reviews. This review aims to formalize XAI methods in CompPath through the: i) introduction of a pathology-centric vocabulary comprising seven core terms; ii) development of a taxonomy across methodological families and three orthogonal axes (stage, type, scope); and iii) establishment of a task-driven framework that maps five clinical questions to recommended methods, method evaluation, and deployment context. Five key gaps between current XAI capabilities and clinical deployment are identified, and actionable steps are proposed to advance XAI for CompPath.
Chinese Translation
计算病理学(Computational Pathology, CompPath)正在利用人工智能(AI)算法,从千兆像素的全切片图像中支持诊断、预后和治疗预测,从而变革医学。临床应用正在不断推进,但在高风险的临床环境中,其发展受到安全性、问责性和监管审查等方面问题的制约。可解释人工智能(XAI)系统有望建立信任并实现验证,然而由于术语不一致、方法学家族相互重叠、验证方式随意以及现有综述的局限,相关文献仍然碎片化。本综述旨在通过以下方式对CompPath中的XAI方法进行系统化:i)引入一套以病理学为核心的术语体系,包含七个核心术语;ii)构建一个跨越方法学家族和三个正交维度(阶段、类型、范围)的分类体系;iii)建立一个任务驱动的框架,将五个临床问题映射到推荐方法、方法评估和部署场景。本文还指出了当前XAI能力与临床部署之间的五个关键差距,并提出了推动XAI在CompPath中发展的可执行步骤。
cs.AI / 40 / 2608.28824

Discovering Machine Correlates of Consciousness

发现意识的机器相关物
Salvi, Romain, Wolfson, Ouri
Abstract
Currently, in biological systems Neural Correlates of Consciousness (NCCs) are characterized in terms of EEG and FMRI signals. Unfortunately, this characterization prevents the transferability of the NCCs concept to machines. Such transferability would be useful in order to investigate AI consciousness. In this paper we provide an alternate characterization that is transferable, and enables the analogous definition of Machine Correlates of Consciousness (MCCs). Specifically, we propose that NCCs (MCCs) are substrate-level signals that are not under human (AI agent) control, and that are reliably modulated by emotions. This paper presents the first empirical investigation of MCCs. Specifically, we present the results of experiments conducted with two LLMs, Llama-2 7B and Llama-3.1 70B parameters. In these LLMs we collect hardware anomaly traces that are substrate-level indicator-sequences. And we show that after controlling for confounding factors, these are modulated differently by emotional and neutral computations. And this difference is statistically significant for the larger Llama-3.1 70B, but not for the smaller Llama-2 7B. The results constitute initial empirical evidence that MCCs are present in the Llama-3.1 70B configuration. And they are consistent with the hypothesis that consciousness probability and degree increase with the LLM sophistication. Independently of consciousness, MCCs can also be used for detection of emotions in AI agents.
Chinese Translation
目前,在生物系统中,意识神经相关物(Neural Correlates of Consciousness, NCCs)是通过脑电(EEG)和功能磁共振(fMRI)信号来刻画的。遗憾的是,这种刻画方式阻碍了NCCs概念向机器的可迁移性,而这种可迁移性对于研究AI意识将非常有用。本文提出了一种可迁移的替代性刻画方式,从而能够类似地定义机器意识相关物(Machine Correlates of Consciousness, MCCs)。具体而言,我们提出NCCs(MCCs)是不受人类(AI智能体)控制的底层信号,并且会被情绪可靠地调制。本文首次对MCCs进行了实证研究。具体来说,我们展示了在两个大语言模型(LLM)——Llama-2 7B和Llama-3.1 70B参数版本——上开展的实验结果。在这些LLM中,我们收集了硬件异常轨迹作为底层指标序列,并证明在控制混杂因素后,这些轨迹在情绪性计算与中性计算下受到的调制方式不同。这种差异在更大的Llama-3.1 70B上具有统计显著性,而在较小的Llama-2 7B上则没有。这些结果构成了MCCs存在于Llama-3.1 70B配置中的初步实证证据,并与“意识的概率和程度随LLM复杂程度的提升而增加”这一假设相一致。即使撇开意识不谈,MCCs也可用于检测AI智能体中的情绪。
cs.AI / 41 / 2608.28833

Evaluating the Hidden Costs of Personalization in Large Language Models

评估大语言模型中个性化的隐性代价
Wang, Yumeng, Wu, Yuchen, Qian, Cheng, Fan, Zhiyuan, Ha, Hyeonjeong, Wu, Shujin, Liu, Jiayu, Ji, Heng, Wang, Ge
Abstract
While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, they increasingly shift from providing balanced, informative responses toward optimizing for user satisfaction when conditioned on personal context such as conversation history, inferred preferences, and user profiles. Specifically, we identify three emerging risks: (1) irrelevant personalization, where models reference personal information in unnecessary contexts; (2) preference narrowing, where models reinforce informational echo chambers; and (3) sycophantic bias, where models agree excessively with user opinions. As a result, models may reference personal information in contexts where it is unnecessary, inadvertently collapse response diversity, or agree excessively with user opinions. Despite the growing use of personalization in AI assistants, there has been limited systematic evaluation of its potential side effects. To bridge this gap, we propose PRISK, a dynamic evaluation framework with automated data generation and tailored metrics that uncovers systematic limitations in current LLM personalization and how personalized information shapes its responses. Our empirical analysis across 13 LLMs demonstrates the presence of user profiles and retrieved memories consistently exacerbates biases, resulting in an average drop of 45.9% in irrelevant personalization, 41.7% in preference narrowing and 61.7% in sycophantic bias.
Chinese Translation
尽管大语言模型(LLM)引入用户个性化信号以提升可用性和实用性,但在以对话历史、推断偏好和用户画像等个人上下文为条件时,它们日益从提供平衡、信息丰富的回答转向以优化用户满意度为目标。具体而言,我们识别出三类新出现的风险:(1)不必要的个性化(irrelevant personalization),即模型在不必要的情境中引用个人信息;(2)偏好窄化(preference narrowing),即模型强化信息茧房;(3)谄媚偏差(sycophantic bias),即模型过度迎合用户观点。由此,模型可能在不必要的情境中引用个人信息、无意间降低回答的多样性,或过度赞同用户观点。尽管个性化在AI助手中的应用日益增多,但对其潜在副作用的系统性评估仍然有限。为填补这一空白,我们提出了PRISK,一个具备自动化数据生成和定制化指标的动态评估框架,用以揭示当前LLM个性化的系统性局限以及个性化信息如何塑造模型响应。我们在13个大语言模型上的实证分析表明,用户画像和检索记忆的存在会持续加剧偏差,导致不必要的个性化平均恶化45.9%,偏好窄化平均恶化41.7%,谄媚偏差平均恶化61.7%。
cs.AI / 42 / 2608.28884

MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft

MineCEraft:将语言模型作为《我的世界》(Minecraft)世界中的建筑工程师进行评估
Lee, Sewoong, Sidhu, Risham, Hockenmaier, Julia, Jung, Yoonhwa
Abstract
We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the reliability and limitations of LLMs for construction tasks in Minecraft. The MineCEraft benchmark comprises 723 domain-expert hand-crafted natural-language instructions with programmatically verifiable evaluation, spanning 17 distinct task categories, providing a safe and controllable experimental environment for assessing LLMs' ability to perform realistic construction engineering tasks. With this benchmark, we conduct an in-depth evaluation of state-of-the-art LLMs and perform a detailed error analysis, revealing key failure modes and practical challenges in applying LLMs to construction engineering tasks.
Chinese Translation
我们提出了MineCEraft(Minecraft建筑工程基准,发音为mine-see-ee-raft),这是一个易于使用的开源基准,旨在系统性地评估大语言模型(LLM)在Minecraft建筑任务中的可靠性与局限性。MineCEraft基准包含723条由领域专家手工编写的自然语言指令,并配有可程序化验证的评估方式,涵盖17个不同的任务类别,为评估LLM执行现实建筑工程任务的能力提供了安全可控的实验环境。基于该基准,我们对最先进的LLM进行了深入评估,并开展了详细的错误分析,揭示了将LLM应用于建筑工程任务时的关键失效模式与实际挑战。
cs.AI / 43 / 2608.28944

Oculi: A Conversational Agentic Platform for Automated Credit Risk Analysis

Oculi:一个用于自动化信用风险分析的对话式智能体平台
Ho, Vennise, Diana, Kristian, Mourad, Sandy, Pilipovic, Milena, Nagisetty, Vineel, Hajimirsadeghi, Hossein
Abstract
Credit risk analysis in financial institutions traditionally requires analysts to manually write SQL queries, run statistical computations, and build visualization dashboards. This is a time-consuming workflow that limits exploration to familiar segments. We introduce \textbf{Oculi}, a conversational platform that transforms natural language questions into comprehensive credit risk analyses, complete with data queries, statistical testing, and interactive visualizations. Oculi employs a three-layer architecture that separates reasoning (LLM-powered agent), execution (Model Context Protocol tool servers), and presentation (agentic UI), enabling analysts to discover high-risk portfolio segments. Within Oculi, a new segment discovery pipeline is proposed that combines deterministic statistical methods with LLM-guided feature selection, leveraging LLM semantic domain knowledge alongside data-driven metrics to identify meaningful, actionable portfolio segments. Evaluated on a mortgage portfolio with 200+ features, Oculi demonstrates effectiveness in discovering material risk segments previously intractable through manual exploration, reducing time-to-insight significantly while maintaining auditability and statistical rigor.
Chinese Translation
金融机构的信用风险分析传统上需要分析师手动编写SQL查询、执行统计计算并构建可视化仪表板。这是一种耗时的流程,限制了探索只能局限于熟悉的业务细分领域。我们提出了Oculi,一个能够将自然语言问题转化为全面信用风险分析的对话式平台,涵盖数据查询、统计检验和交互式可视化。Oculi采用三层架构,将推理(由大语言模型驱动的智能体)、执行(模型上下文协议(Model Context Protocol)工具服务器)和展示(智能体用户界面)分离,使分析师能够发现高风险投资组合细分领域。在Oculi中,我们提出了一种新的细分领域发现流水线,它将确定性的统计方法与大语言模型引导的特征选择相结合,利用大语言模型的语义领域知识与数据驱动指标来识别有意义的、可操作的投资组合细分领域。在一个包含200多个特征的抵押贷款投资组合上的评估表明,Oculi能够有效发现以往通过人工探索难以发现的重要风险细分领域,在保持可审计性和统计严谨性的同时,显著缩短了从数据到洞察的时间。
cs.AI / 44 / 2608.28945

Automated Researchers Can Mitigate Well-characterized Alignment Failures

自动化研究人员能够缓解已被充分刻画的对齐失败问题
Yueh-Han, Chen, Wen, Jiaxin, Kirchner, Jan Hendrik
Abstract
Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such as deception, sycophancy, and jailbreaks, are already measurable by public benchmarks. We study whether automated alignment researchers (AARs) can post-train to mitigate alignment failures by proposing training methods and data to simultaneously optimize multiple safety benchmarks, while largely preserving general capability. Across 10 alignment failures, the strongest AAR methods significantly reduce the targeted alignment failures and generalize to a held-out benchmark, multi-turn behavioral audits, and models up to 4.7x larger than the target model. As a human baseline, 28 experienced researchers receive up to eight hours to develop one-shot methods for the same benchmarks, but their methods underperform the best AAR methods. Using human ideas as the AARs' initial research direction does not improve performance, suggesting current AARs may not need guidance from experienced researchers. These results suggest that automating alignment research on well-characterized failures may be practical in the near term.
Chinese Translation
将对齐研究自动化可能会加速实现与人类对齐的AI的进程,但其效果难以衡量。幸运的是,许多对齐失败问题,如欺骗、谄媚和越狱攻击,已经可以通过公开的基准进行衡量。我们研究了自动化对齐研究人员能否通过后训练来缓解对齐失败问题,具体做法是提出训练方法和数据,以同时优化多个安全基准,同时在很大程度上保持通用能力。在10种对齐失败问题上,最强的自动化对齐研究人员(AAR)方法显著降低了目标对齐失败率,并能泛化到留出的基准、多轮行为审计,以及比目标模型大4.7倍的模型。作为人类基线,28名经验丰富的研究人员获得最多八小时的时间,为相同的基准开发一次性方法,但他们的方法表现不如最佳的AAR方法。将人类的想法作为AAR的初始研究方向并不能提升性能,这表明当前的AAR可能不需要经验丰富的研究人员的指导。这些结果表明,在短期内,针对已被充分刻画的失败问题将对齐研究自动化可能是切实可行的。
cs.AI / 45 / 2608.28965

From Location Phrases to Geographic Entities: Task-Adapted Retrieval for People Search

从位置短语到地理实体:面向人名搜索的任务自适应检索
Li, Yanbo, Zheng, Chujie, Xu, Jiahao, Bhole, Chetan, Zhang, Lingyu, Ahluwalia, Puneet Singh, Nguyen, Kevin, Muthuregunathan, Raghavan, Sachindran, Santhosh, Ahuja, Sachin, Borisyuk, Fedor
Abstract
People search must map free-form location phrases to geographic entities used as structured retrieval filters. Lexical standardizers handle canonical names well but are brittle to aliases, misspellings, metropolitan expressions, and same-name ambiguity. We formulate this task as graded, set-valued entity retrieval over a fixed ontology. We identify three coupled design requirements: distinguishing identity-preserving variation from knowledge-dependent aliases, controlling false negatives among valid same-name entities, and separating stable transformations from mutable entity knowledge. We realize them in a prompt-asymmetric bi-encoder with calibrated alias support, bounded ambiguity-aware negatives, and editable entity documents that support localized updates without retraining. Across a fixed production-derived development benchmark and a public GeoNames transfer task, task adaptation improves substantially over frozen encoders and standard token baselines. Controlled development ablations show that specialized supervision contributes beyond standard task fine-tuning and encoder scaling. On GeoNames, the adapted model improves known-target Recall@1 throughout zero-to-moderate character overlap, while character n-grams retain a small aggregate Target Recall@5 advantage. In a blinded human comparison on a stratified production challenge set, our model raises relevant P@1 from 28.0% to 46.0% (p=0.012). Fixed-query endpoint estimates improve on non-canonical queries and remain close to control on frequent queries; a randomized live experiment detects no engagement regression. These results support task-adapted geographic entity retrieval as a practical replacement for the incumbent taxonomy-based standardizer, with the largest relevance gains on non-canonical queries.
Chinese Translation
人名搜索(People Search)需要将自由形式的位置短语映射为用作结构化检索过滤条件的地理实体。词汇标准化方法能很好地处理规范名称,但在处理别名、拼写错误、都市圈表达和同名歧义时表现脆弱。我们将该任务形式化为在固定本体上进行的分级、集合值实体检索。我们识别出三个相互关联的设计需求:区分保持身份的变体与依赖知识的别名、控制有效同名实体中的假阴性、以及将稳定变换与可变的实体知识分离。我们在一个提示非对称的双编码器(bi-encoder)中实现这些需求,采用经校准的别名支持、有界的歧义感知负样本,以及支持无需重新训练的局部更新的可编辑实体文档。在一个固定的源自生产环境的开发基准和一个公开的 GeoNames 迁移任务上,任务自适应方法显著优于冻结编码器和标准分词基线。受控的开发集消融实验表明,专门的监督在标准任务微调和编码器扩展之外仍有额外贡献。在 GeoNames 上,自适应模型在整个零到中等字符重叠区间内提升了已知目标的 Recall@1,而字符 n-gram 在总体 Target Recall@5 上仍保持少量优势。在一个分层生产挑战集上的盲测人类对比中,我们的模型将相关 P@1 从 28.0% 提升至 46.0%(p=0.012)。固定查询的端点估计显示非规范查询有所改进,且在频繁查询上与对照组接近;一项随机化在线实验未检测到用户参与度的下降。这些结果支持将任务自适应的地理实体检索作为现有基于分类体系的标准化方法的实用替代方案,其相关性提升在非规范查询上最为显著。
cs.AI / 46 / 2608.28968

Efficient GPU Retrieval for Semantic Search

面向语义搜索的高效GPU检索
Das, Dhritiman, Zheng, Chujie, Kaoshik, Ronak, Dixit, Pratik, Shah, Vishal, Li, Yanbo, Xu, Jiahao, Agarwal, Manika, Naik, Chinmay, Zhang, Lingyu, Bhole, Chetan, Mehta, Chirag Bhanuprasad, Zheng, Meng, Ahluwalia, Puneet Singh, Singh, Shirisha, Jin, Ping, Apte, Manas, Mohanasundaram, Gokulraj, Bingol, Tugrul, Muthuregunathan, Raghavan, Borisyuk, Fedor
Abstract
Semantic Search on LinkedIn must retrieve relevant profiles from a corpus of hundreds of millions in response to natural-language queries such as "a fintech founder in Berlin who worked in payments." The deployed relevance policy is bottleneck-oriented: every active non-negotiable facet must be satisfied, and a pre-existing LLM Graded Relevance (GR) judge operationalizes this through a fixed min/median aggregation over facet grades. Cosine similarity instead averages evidence, letting a strong match on one facet mask failure on another, capping the recall of the first-stage (L0) retriever. We present a policy-aligned retrieval framework: embeddings are partitioned into eight category-supervised segments whose scores follow the same min/median rule at serving time; for multi-vector retrieval, this segment score is computed independently per tagged document slot and maximized across slots. A lightweight single-slot Stage-1 scorer generates high-recall candidates, while scale-invariant relative-norm gating keeps category activation consistent across training, evaluation, and serving. On 21K held-out queries, this representation improves offline relevance over a matched-capacity baseline, with gains broadly distributed across facet combinations. We serve this framework with a two-stage GPU architecture: an FP8 coarse ranker scores the full corpus, increasing per-shard capacity by 71% and Stage-1 matmul throughput by 36%, then an FP16 stage exactly re-ranks an oversampled candidate set, recovering 99.6-99.8% of full-FP16 recall at over 500 QPS per shard replica. In a member-randomized A/B test, exploratory-query Precision@10 under the unchanged GR judge rises from 63.7% to 79.0% and navigational Precision@1 from 65.5% to 74.7%, with a blinded human evaluation independently confirming the Precision@10 gain.
Chinese Translation
LinkedIn上的语义搜索必须从数亿规模的语料库中检索相关档案,以响应诸如“在柏林从事支付行业的金融科技创始人”之类的自然语言查询。目前已部署的相关性策略是面向瓶颈的:每个活跃且不可妥协的维度(facet)都必须被满足,而一个已有的LLM分级相关性(Graded Relevance, GR)评判器通过对各维度评分进行固定的最小值/中位数聚合来实现该策略。相反,余弦相似度会平均所有证据,使得某一维度上的强匹配可能掩盖另一维度上的失败,从而限制了第一阶段(L0)检索器的召回率。我们提出了一种与策略对齐的检索框架:将嵌入(embedding)划分为八个由类别监督的片段,在线服务时这些片段的评分遵循同样的最小值/中位数规则;对于多向量检索,该片段评分对每个带标签的文档槽位独立计算,并在所有槽位间取最大值。一个轻量级的单槽位第一阶段评分器生成高召回候选集,而尺度不变的相对范数门控确保类别激活在训练、评估和服务阶段保持一致。在2.1万个保留查询上,该表示方法相比同等容量的基线提升了离线相关性,且增益广泛分布于各种维度组合。我们采用两阶段GPU架构来服务该框架:FP8粗排器对整个语料库评分,使每分片容量提升71%,第一阶段矩阵乘吞吐量提升36%;随后FP16阶段对过采样的候选集进行精确重排,在每分片副本超过500 QPS的情况下恢复了99.6%–99.8%的全FP16召回率。在成员随机化的A/B测试中,在GR评判器保持不变的情况下,探索性查询的Precision@10从63.7%提升至79.0%,导航式查询的Precision@1从65.5%提升至74.7%,且一项盲测人工评估独立验证了Precision@10的提升。
cs.AI / 47 / 2608.28974

From Analytics to Tumor Boards: An Evidence-Linked Multi-Agent Workflow for Oncology Feature Extraction

从数据分析到肿瘤病例讨论会:一种证据关联的多智能体工作流用于肿瘤学特征提取
Kang, Daniel, Hu, Michelle, Shimgekar, Soorya Ram, Vassef, Shayan, Wang, Yufan, Sahu, Anit Kumar, De Choudhury, Munmun, Swain, Vedant Das, Poellabauer, Christian, Khor, Li Yan, Saha, Koustuv, Wojciechowski, Robert, Kidd, Elliot, Zonooz, Piyum, Kumar, Navin
Abstract
Clinically relevant oncology information is distributed across heterogeneous, longitudinal documentation, creating substantial abstraction burden and requiring accurate attribution across specimens, tumors, biomarkers, and time points, while manual cancer-registry abstraction can require 27.2 minutes per case, highlighting the need for scalable methods that preserve clinical context while converting documentation into structured data. We evaluate an oncology information-extraction workflow in which OncoLens supplies multi-source, oncology-aware document selection, aggregation, and normalization from integrated EHRs, while the NimbleMind Multi-Agent System (nMAS) is a configurable oncology information-extraction workflow that extracts clinically relevant structured fields from fragmented oncology documentation. The extraction task uses a clinician-informed schema of 328 attributes spanning report metadata, diagnosis, staging, and cancer-type-specific information. nMAS separates clinician-defined field specifications from model execution and combines complexity-aware extraction, report-level consolidation, and source-grounded validation. The retrospective evaluation included 230 de-identified oncology documents from 40 patients and 418 clinician-reviewed document-field pairs containing 1,126 non-empty reference values. Evaluation focused on fields identified by clinicians as present in the source documents rather than exhaustively annotating all 328 schema fields. nMAS achieved a rank-weighted value-level precision of 82.6%, recall of 87.5%, and F1 of 85.0%, compared with an F1 of 66.4% for an independently implemented UMA-style MiniMax M2.5 comparator. These findings support the feasibility of using a configurable, source-grounded extraction workflow to convert fragmented oncology documentation into reusable structured data.
Chinese Translation
具有临床价值的肿瘤学信息分布于异构的纵向病历文档中,这带来了繁重的信息抽取负担,并需要在标本、肿瘤、生物标志物和时间点之间进行精确归因。同时,人工癌症登记摘录每个病例可能耗时27.2分钟,凸显了对可扩展方法的需求——这些方法需在将病历文档转化为结构化数据的同时保留临床上下文。我们评估了一种肿瘤学信息提取工作流:其中OncoLens负责从集成的电子健康记录(EHR)中提供多来源、肿瘤学感知的文档选择、聚合与规范化;而NimbleMind多智能体系统(NimbleMind Multi-Agent System, nMAS)是一个可配置的肿瘤学信息提取工作流,可从碎片化的肿瘤学文档中提取具有临床价值的结构化字段。该提取任务采用由临床医生参与制定的328个属性的模式(schema),涵盖报告元数据、诊断、分期及癌症类型特异性信息。nMAS将临床医生定义的字段规范与模型执行相分离,并结合了复杂度感知提取、报告级整合以及基于溯源的验证。回顾性评估纳入了来自40名患者的230份去标识化肿瘤学文档,以及418个经临床医生审阅的文档-字段对,其中包含1,126个非空参考值。评估聚焦于临床医生认定在源文档中存在的字段,而非对全部328个模式字段进行穷尽式标注。nMAS实现了基于排名加权的值级别精确率82.6%、召回率87.5%和F1值85.0%,而独立实现的UMA风格MiniMax M2.5对照方法的F1值为66.4%。这些结果支持了使用可配置、基于溯源的提取工作流将碎片化肿瘤学文档转化为可复用结构化数据的可行性。
cs.AI / 48 / 2608.28977

The Role of Network Topology and Opponent Information in Shaping Cooperation in Multi-Agent Reinforcement Learning Systems

网络拓扑与对手信息在多智能体强化学习系统中塑造合作的作用
Son, Seongho, Hailes, Stephen, Musolesi, Mirco
Abstract
Several works have investigated the influence of graph topology on cooperation among artificial agents, while the majority of the literature has focused on modelling agents' adaptation through strategy imitation, which relies solely on the cumulative payoffs of others. This paper investigates scenarios in which each agent learns to play the two-player Iterated Prisoner's Dilemma (IPD) using deep reinforcement learning. Each agent is represented as a node in a graph, where its neighbours constitute the pool of opponents with whom it can interact. During each IPD episode, agents are provided with different types of information about their opponent, consisting of action history and opponent identity. Experimental results across different graph topologies show that the number of neighbours per node and the average path length are the main factors affecting the emergence of cooperation. We also show that, while partner selection fosters mutual cooperation by limiting the diversity of the opponent pool, providing agents with the identity of their opponent hinders the proliferation of cooperative strategies.
Chinese Translation
已有若干研究考察了图拓扑对人工智能体间合作的影响,但大多数文献侧重于通过策略模仿来建模智能体的适应过程,而这仅依赖于其他智能体的累积收益。本文研究了每个智能体使用深度强化学习来学习进行两人重复囚徒困境(Iterated Prisoner's Dilemma, IPD)博弈的场景。每个智能体被表示为图中的一个节点,其邻居构成可供交互的对手池。在每个IPD回合中,智能体会获得关于对手的不同类型的信息,包括动作历史和对手身份。在不同图拓扑上的实验结果表明,每个节点的邻居数量和平均路径长度是影响合作涌现的主要因素。我们还发现,尽管伙伴选择通过限制对手池的多样性来促进相互合作,但为智能体提供对手身份信息却会阻碍合作策略的扩散。
cs.AI / 49 / 2608.28978

Selective Forgetting: A Graph-Based Memory Framework for Long-Term LLM Agents

选择性遗忘:一种面向长期LLM智能体的基于图的记忆框架
Rusu, Theo, Khanzadeh, Sourena, Alalfi, Manar
Abstract
Knowledge graphs have been proposed as a structured alternative to flat retrieval-augmented generation for long-term agent memory, on the assumption that representing conversations as entities and relations improves recall. We evaluate that assumption directly. Our framework extracts each conversational turn into typed nodes and attributed edges, answers questions from a two-hop subgraph, and periodically prunes nodes that score low on a weighted combination of recency, access frequency, degree centrality, and age. On LongMemEval, the graph does not outperform a flat vector baseline at a matched candidate-generation budget of five retrieval roots: token F1 is $0.417$ against $0.468$, and a paired bootstrap over 500 questions gives $\Delta = -0.050$ (95\% CI $[-0.085, -0.016]$). The gap is widest on questions that require recalling a specific prior assistant turn, where judged correctness falls from $0.911$ to $0.607$, suggesting that decomposing a turn into entities discards the surface form these questions depend on. The forgetting module is more successful. Applied once to a persistent 27{,}021-node graph, it removes 9.8\% of nodes and 9.5\% of stored bytes; token F1 is unchanged ($+0.001$, 95\% CI $[-0.015, +0.016]$) and judged correctness falls by $1.6$ points, with the 95\% interval bounding any loss at $3.8$ points ($[-0.038, +0.006]$). Because our extractor is a single small model evaluated on one benchmark, these results characterise this extraction-based pipeline rather than graph-structured memory in general. Code: https://github.com/skhanzad/Selective-Amnesia
Chinese Translation
知识图谱已被提出作为扁平化检索增强生成(RAG)的结构化替代方案,用于智能体的长期记忆,其假设是将对话表示为实体和关系能够提升召回效果。我们直接对这一假设进行了评估。我们的框架将每轮对话抽取为带类型的节点和带属性的边,从两跳子图中回答问题,并定期根据近期性、访问频率、度中心性和年龄的加权组合得分,剪除低分节点。在LongMemEval基准上,在相同的五个检索根节点的候选生成预算下,图结构并未优于扁平向量基线:token F1为0.417,对比基线的0.468,对500个问题进行的配对自助法(bootstrap)检验给出Δ = -0.050(95%置信区间[-0.085, -0.016])。这一差距在需要回忆特定先前助手回复的问题上最为明显,人工评判的正确率从0.911降至0.607,表明将对话轮次分解为实体会丢弃这些问题所依赖的表层形式。遗忘模块则更为成功。将该模块一次性应用于包含27,021个节点的持久化图,可移除9.8%的节点和9.5%的存储字节;token F1保持不变(+0.001,95%置信区间[-0.015, +0.016]),人工评判正确率下降1.6个百分点,且95%置信区间将任何损失限制在3.8个百分点以内([-0.038, +0.006])。由于我们的抽取器是单一的小型模型且仅在一个基准上评估,这些结果刻画的是这种基于抽取的流水线,而非一般意义上的图结构记忆。代码:https://github.com/skhanzad/Selective-Amnesia
cs.AI / 50 / 2608.28990

Agentic AI uncovers conserved cross-tissue protein co-abundance programs inaccessible to single-dataset analysis

智能体AI发现单数据集分析无法触及的跨组织保守蛋白共丰度程序
Guan, Runyu, Wu, Dehao, Xie, Qiqi, Li, Yang, Wang, Haohan
Abstract
Protein co-abundance clusters preserved across tissues can reveal shared disease mechanisms and candidate therapeutic targets, particularly when proteins implicated in organ-confined diseases converge in peripheral or accessible tissues. However, previous cross-tissue studies have focused on biologically pre-selected tissue pairs, leaving most possible combinations and non-obvious relationships unexplored. We present an LLM-agent framework for large-scale, evidence-grounded comparison of tissue-specific protein co-abundance networks. The framework constructs tissue networks, derives pairwise consensus clusters, and integrates evidence from expression atlases, protein interaction and complex databases, pathway annotations, disease catalogues, and literature. Applied to all 820 pairwise combinations of 41 human tissues and fluids, it identified 1,833 conserved co-abundance clusters across 406 tissue pairs. Colon, synovial fluid, blood, cerebrospinal fluid, and bone marrow were the most broadly connected tissues, while the most cluster-rich pairs were dominated by bone marrow. The analysis also highlighted non-obvious relationships: skin-bone marrow exceeded the anatomically adjacent bone-bone marrow pair, while colon-breast contained cancer-relevant clusters involving extracellular-matrix remodeling, lipid metabolism, and immune modulation. Cluster-level analyses generated further mechanistic hypotheses, including a brain-gut extracellular-vesicle/redox/serotonin-cofactor axis and a liver-bone marrow stress-response axis involving genes linked to white matter disease. These results provide a global, comparable landscape of conserved protein co-abundance and a hypothesis-generating resource for mechanistic and therapeutic exploration. Code and data are available at https://github.com/Gry1005/AgenticAI-conserved-cross-tissue-protein-co-abundance.
Chinese Translation
跨组织保守的蛋白共丰度聚类可以揭示共享的疾病机制和候选治疗靶点,尤其是当与器官局限性疾病相关的蛋白在外周或易获取组织中汇聚时。然而,以往跨组织研究多聚焦于生物学上预先选定的组织配对,使大多数可能的组合及非显而易见的关系未被探索。我们提出了一个基于大语言模型的智能体(LLM-agent)框架,用于大规模、有证据支撑的组织特异性蛋白共丰度网络比较。该框架构建组织网络,推导成对共识聚类,并整合来自表达图谱、蛋白相互作用与复合物数据库、通路注释、疾病目录及文献的证据。将该框架应用于41种人体组织和液体的全部820种两两组合,共识别出406个组织对中的1,833个保守共丰度聚类。结肠、滑液、血液、脑脊液和骨髓是连接最广泛组织,而聚类最丰富的组织对则以骨髓为主。分析还揭示了非显而易见的关系:皮肤-骨髓的关联超过解剖上相邻的骨-骨髓配对,而结肠-乳腺则包含涉及细胞外基质重塑、脂质代谢和免疫调节的癌症相关聚类。聚类层面的分析产生了进一步的机制假说,包括脑-肠细胞外囊泡/氧化还原/血清素辅助因子轴,以及涉及白质疾病相关基因的肝-骨髓应激反应轴。这些结果提供了全球性的、可比较的保守蛋白共丰度图谱,并为机制和治疗探索提供了假设生成资源。代码和数据可在 https://github.com/Gry1005/AgenticAI-conserved-cross-tissue-protein-co-abundance 获取。
cs.AI / 51 / 2608.28997

Verification abundance, adjudication scarcity: what happens to mathematical knowledge when proof checking becomes free

验证的丰裕,裁决的稀缺:当证明检验变得免费时,数学知识会怎样
Kallel, Maher, Louadi, Mohamed El
Abstract
In May 2026 an OpenAI model produced a counterexample to the Erd\H{o}s unit distance conjecture. Five mathematicians published a human-verified version the same day, and the result entered the literature within weeks. In August 2026 the same laboratory published ten mathematical and theoretical computer science results, each accompanied by a machine-checkable Lean 4 certificate with no unproved steps. Four weeks later, one remained the subject of an unresolved dispute over whether its formalization meant what it claimed. We argue that this difference is structural. We distinguish three layers of verification: derivational validity, which a kernel checks; representational fidelity, whether the formal statement means the intended question; and epistemic significance. Only the first is mechanizable. Making it effectively free therefore does not eliminate verification work but shifts the burden to layers dependent on scarce expert attention. Measurements of the August corpus illustrate the shift. The kernel-checked proofs total 20.6 MB, while the statements requiring human audit total 55.6 KB, a ratio of 379 to 1. Yet those statements contain 218 bespoke definitions rather than relying on community-vetted ones. The audit surface is therefore small in volume but irreducibly expert. We argue that machine checking produces verification abundance while leaving adjudication scarce. We propose a six-category taxonomy of representational mismatch, a disclosure schema for machine-generated mathematical claims, and implications for software, cryptography, and regulated decision systems.
Chinese Translation
2026年5月,OpenAI的一个模型给出了埃尔德什单位距离猜想的反例。五位数学家在同一天发表了经人工验证的版本,该结果在数周内进入学术文献。2026年8月,同一实验室发表了十项数学与理论计算机科学成果,每项成果均附有无可证明步骤缺失、可机器检验的Lean 4证书。四周之后,其中一项成果仍处于未决争议之中,争议焦点在于其形式化陈述是否表达了其所声称的含义。我们认为,这种差异是结构性的。我们区分了验证的三个层次:推导有效性,由内核检验;表征保真性,即形式化陈述是否表达了所意图的问题;以及认知重要性。只有第一层是可以机械化的。因此,使其几乎免费并不会消除验证工作,而是将负担转移至依赖稀缺专家注意力的层次上。对8月这批成果的测量数据说明了这种转移:经内核检验的证明总计20.6 MB,而需要人工审核的陈述总计55.6 KB,比例约为379比1。然而,这些陈述包含218个定制定义,而非依赖经社区审查的定义。因此,审核面在体量上虽小,但本质上离不开专家。我们认为,机器检验带来的是验证的丰裕,而裁决依然稀缺。我们提出了表征失配的六类分类法、面向机器生成数学论断的披露模式,以及对软件、密码学和受监管决策系统的启示。
cs.AI / 52 / 2608.29008

Multi-Step Forecasting of Grape Berry Temperature based on LSTM Model with Feed-Forward Attention

基于前馈注意力LSTM模型的葡萄浆果温度多步预测
Gorthi, Srikanth, Divyanth, L. G., Bhalekar, Dattatray, Keller, Markus, Khot, Lav
Abstract
Accurate forecasting of grape berry temperature (Tb) is essential for enabling timely heat stress management in vineyards. In this study, a feed-forward attention mechanism integrated with a Long Short-Term Memory network (FAM-LSTM) was developed and evaluated for multi-step, high-resolution Tb prediction. Models were trained using environmental data from 2023 and 2024 at Prosser, WA, USA, and validated on 2025 summer data. FAM-LSTM was benchmarked against LSTM, GRU, RNN, and Random Forest (RF) across horizons ranging from 15 minutes to 72 hours (288 time steps). Two input scenarios were evaluated: nearest open-field weather station observations and in-vineyard microclimate measurements. FAM-LSTM consistently outperformed all benchmark models across all horizons and input scenarios. Incorporating in-vineyard microclimate data significantly improved forecasting accuracy at longer horizons. Using open-field data, FAM-LSTM achieved MAE and RMSE ranges of 0.58 to 1.70 deg C and 0.65 to 2.07 deg C, respectively. In-vineyard observations further improved performance, with MAE and RMSE in the ranges of 0.51 to 1.55 deg C and 0.71 to 1.87 deg C. Error analysis showed prediction uncertainty was highest during peak daytime periods (11:00 to 18:00) and increased progressively with forecast horizon. Overall, the FAM-LSTM framework offers robust Tb forecasting to support precision heat stress management in vineyards.
Chinese Translation
准确预测葡萄浆果温度对于葡萄园及时开展热应激管理至关重要。本研究开发并评估了一种将前馈注意力机制与长短期记忆网络相结合的模型(FAM-LSTM),用于多步、高分辨率的葡萄浆果温度预测。模型使用美国华盛顿州普罗瑟2023年和2024年的环境数据进行训练,并在2025年夏季数据上进行验证。FAM-LSTM与LSTM、GRU、RNN和随机森林(RF)在15分钟至72小时(288个时间步)的预测范围内进行了对比。研究评估了两种输入情景:最近的露天气象站观测数据和葡萄园内微气候测量数据。FAM-LSTM在所有预测范围和输入情景下均持续优于所有基准模型。引入葡萄园内微气候数据显著提高了较长预测范围的预测精度。使用露天数据时,FAM-LSTM的平均绝对误差(MAE)和均方根误差(RMSE)分别为0.58至1.70摄氏度和0.65至2.07摄氏度。葡萄园内观测数据进一步提升了性能,MAE和RMSE分别为0.51至1.55摄氏度和0.71至1.87摄氏度。误差分析表明,预测不确定性在白天高峰时段(11:00至18:00)最高,并随预测范围的增加而逐步增大。总体而言,FAM-LSTM框架可提供稳健的葡萄浆果温度预测,为葡萄园精准热应激管理提供支持。
cs.AI / 53 / 2608.29012

Frequency Selective Neural Networks as a Foundation Architecture for Time Series Learning

频率选择神经网络作为时间序列学习的基础架构
Huang, Hui, Sun, Ye, Hu, Shiyan
Abstract
Time-series data across physical and biological domains are fundamentally driven by complex, non-stationary oscillatory modes. While deep learning models, such as Convolutional Neural Networks (CNNs), Recurrent Neural Networks, and Transformers, have dominated sequential analysis, they remain fundamentally "spectral-blind". By mapping continuous physical waves into unconstrained spatial or discrete token spaces, these architectures suffer from severe spectral entanglement, acting as opaque black boxes that decouple predictive accuracy from physical reality. In this paper, we introduce the Frequency Selective Neural Network (FSNN), pioneering a foundation architecture guaranteeing physical interpretability without sacrificing expressive power of deep learning. FSNN addresses spectral entanglement by explicitly embedding the rigorous mathematics of advanced signal processing into its neural topology. Through a fully differentiable Wiener-like filter bank optimized via complex-domain backpropagation, FSNN autonomously discovers and isolates the precise physical modes of a given task. Extensive evaluations demonstrate that FSNN establishes state-of-the-art predictive performance, achieving $77.0\%$ average accuracy on the standard 10 multivariate UEA datasets and leading across all major metrics on the highly imbalanced PTB-XL clinical ECG benchmark. Crucially, in contrast to yielding abstract feature maps, FSNN converges directly on physically meaningful frequency bands, such as isolating the cardiac QRS complex, providing a highly scalable, interpretable paradigm for robust pattern recognition in complex temporal domains. Our code is available at: https://github.com/ad6174hhhh/FSNN.
Chinese Translation
物理与生物领域的时间序列数据从根本上由复杂的非平稳振荡模态驱动。尽管卷积神经网络(CNN)、循环神经网络和Transformer等深度学习模型在序列分析中占据主导地位,但它们在本质上仍是“频谱盲”的。通过将连续物理波映射到无约束的空间或离散标记空间,这些架构会遭受严重的频谱纠缠问题,成为将预测精度与物理现实割裂开的不透明黑箱。本文提出频率选择神经网络(Frequency Selective Neural Network, FSNN),开创了一种在不牺牲深度学习表达能力的前提下保证物理可解释性的基础架构。FSNN通过将先进信号处理的严格数学显式嵌入其神经拓扑结构,解决了频谱纠缠问题。借助通过复数域反向传播优化的全可微分类维纳滤波器组(Wiener-like filter bank),FSNN能够自主发现并分离给定任务的精确物理模态。大量评估表明,FSNN取得了最先进的预测性能:在标准10个多元UEA数据集上达到77.0%的平均准确率,并在高度不平衡的PTB-XL临床心电(ECG)基准上的所有主要指标中均处于领先地位。至关重要的是,与生成抽象特征图不同,FSNN直接收敛于具有物理意义的频段(例如分离心脏QRS波群),为复杂时域中的稳健模式识别提供了一种高度可扩展、可解释的范式。我们的代码已发布于:https://github.com/ad6174hhhh/FSNN。
cs.AI / 54 / 2608.29026

Disentangling Representation using Attributes-based Gaussian Estimation for Medical Sound Diagnosis

基于属性高斯估计的解耦表示学习方法用于医学声音诊断
Zhao, Ke
Abstract
Deep learning has a powerful capability of feature extraction. However, the lack of fairness and interpretability in deep neural networks poses limitations to their adoption in the medical domain. This paper proposes a disentangled representation learning (DisenRL) framework, named the Attributes-based Gaussian Estimation for Disentangled Representation (AGEDR), which incorporates Attribute Mapping Embedding (AME) modules designed to map attributes into vectors and align them with a subset of the latent vectors in a Variational AutoEncoder (VAE). This part of the latent vector will be disentangled from the remaining latent vectors by minimizing mutual information. A classifier is then trained using the mean parameters of the latent vectors from the VAE. Extensive experiments demonstrate that AGEDR outperforms both conventional classification models and existing disentangled representation learning methods. The ablation experiments also indicate the disentangling capability and fairness of AGEDR. The source code is publicly available at https://github.com/ZhaoKe1024/DisentangledRepr.
Chinese Translation
深度学习具有强大的特征提取能力。然而,深度神经网络缺乏公平性和可解释性,限制了其在医学领域的应用。本文提出了一种解耦表示学习(Disentangled Representation Learning, DisenRL)框架,命名为基于属性的高斯估计解耦表示方法(Attributes-based Gaussian Estimation for Disentangled Representation, AGEDR)。该框架引入了属性映射嵌入(Attribute Mapping Embedding, AME)模块,用于将属性映射为向量,并使其与变分自编码器(Variational AutoEncoder, VAE)中的部分潜在向量对齐。通过最小化互信息,这部分潜在向量将从其余潜在向量中解耦出来。随后,利用VAE潜在向量的均值参数训练分类器。大量实验表明,AGEDR优于传统分类模型和现有的解耦表示学习方法。消融实验也验证了AGEDR的解耦能力和公平性。源代码已在 https://github.com/ZhaoKe1024/DisentangledRepr 公开。
cs.AI / 55 / 2608.29028

Facts Without Rules: Boundary Metadata Collapse in Multi-Agent LLM Handoffs

无规则的事实:多智能体大语言模型交接中的边界元数据坍塌
Wang, Yian, Goyal, Agam, Chandrasekharan, Eshwar, Sundaram, Hari
Abstract
Multi-agent LLM systems often coordinate by compressing an upstream interaction into a handoff artifact that downstream agents treat as shared state. We show that this handoff step is a structural source of privacy leakage: summaries preferentially preserve operational facts while weakening the boundary metadata that governs how those facts may be used---a failure mode we call \emph{summary collapse}. On a controlled multi-agent coordination testbed we measure marker survival with a human-validated judge ($\kappa = 0.74$), where $\sigma_b = 1$ means every boundary marker survives verbatim and $\sigma_b = 0$ means all are lost. Boundary-marker and operational-fact survival are nearly uncorrelated at the handoff level on both GPT-5-mini and DeepSeek-R1-32B (Pearson $r$ near zero): uncompressed free-text handoffs preserve boundaries at $\sigma_b \approx 0.80$, whereas a $25$-word budget drops $\sigma_b$ to ${\approx}0.57$ while operational-fact survival stays near ceiling. Controlled downstream tests reveal that protection depends on \emph{boundary explicitness}: vague languages leak in $73\%$ of GPT and $50\%$ of DeepSeek cases, while explicit constraints reduce leakage to under $15\%$ across all three tested models. A no-handoff single-agent control further shows the failure is not reducible to multi-agent topology as direct full-marker access still leaks more often than the operationalized handoff. Prompt-only mitigation and exact-string redaction only partially address the problem, while a gold-derived audience allowlist nearly eliminates leakage across models, showing that correctly identifying audience boundaries is the key factor.
Chinese Translation
多智能体大语言模型系统通常通过将上游交互压缩为一份交接产物来协调工作,下游智能体将其视为共享状态。我们证明,这一交接步骤是隐私泄露的结构性来源:摘要倾向于优先保留操作性事实,而削弱了约束这些事实使用方式的边界元数据——我们将这种失效模式称为“摘要坍塌”(summary collapse)。在一个受控的多智能体协调测试平台上,我们采用经人工验证的评判器($\kappa = 0.74$)来度量标记存活率,其中 $\sigma_b = 1$ 表示所有边界标记均被逐字保留,$\sigma_b = 0$ 表示全部丢失。在 GPT-5-mini 和 DeepSeek-R1-32B 上,边界标记存活率与操作性事实存活率在交接层面几乎不相关(Pearson $r$ 接近于零):未压缩的自由文本交接可将边界保留在 $\sigma_b \approx 0.80$,而 25 词的预算则使 $\sigma_b$ 降至约 0.57,同时操作性事实的存活率仍接近上限。受控的下游测试表明,保护效果取决于“边界明确性”:模糊表述在 GPT 和 DeepSeek 的案例中分别泄露 73% 和 50%,而明确的约束条件在所有三个受测模型中将泄露率降至 15% 以下。无交接的单智能体对照实验进一步表明,该失效并不能归因于多智能体拓扑结构,因为直接访问完整标记仍比实际交接操作更频繁地导致泄露。仅基于提示的缓解措施和精确字符串删除只能部分解决问题,而基于黄金标准推导的受众白名单则几乎在所有模型上消除了泄露,这表明正确识别受众边界是关键因素。
cs.AI / 56 / 2608.29030

Learning to Follow In-Context Watermark Instructions via Self-Distillation

通过自蒸馏学习遵循上下文内水印指令
Liu, Yepeng, Chen, Tianyi, Zhao, Xuandong, Song, Dawn, Bu, Yuheng
Abstract
In-context watermarking (ICW) prepends an instruction to a query asking the model to embed a statistically detectable signal in its response. It thus equips LLMs with a watermarking interface that third parties can invoke without access to model internals. Its reliability hinges on the LLM following the instruction without degrading answer quality, yet how well current LLMs do so has not been measured. We introduce $\mathsf{ICWBench}$, a benchmark of three verifiable ICW instruction families, each scored on both detectability and answer quality. Evaluating 14 frontier proprietary and open-source LLMs, we find that none of the evaluated LLMs achieves both objectives across all three families. To address this, we propose a self-contained two-stage training method, requiring no distillation from a stronger model, no manual annotation, and no pre-existing ICW IF ability. The first stage, self-distillation with logits perturbation (SDLP), uses the same base LLM as both teacher and student: an instruction-equivalent decoding-time logits perturbation makes the teacher follow the ICW instruction, and the student is trained to match the teacher's output distribution. The second stage applies reinforcement learning with the automatic verifier as the reward. Applied to Qwen3-14B and GPT-OSS-20B, our method raises average TPR@$1\%$FPR across three ICW instructions from $0.100$ to $0.974$ and from $0.337$ to $0.968$, respectively, while maintaining high response quality under both perplexity evaluation and LLM-as-a-Judge.
Chinese Translation
上下文内水印(In-context watermarking, ICW)在查询前添加一条指令,要求模型在其响应中嵌入一个统计上可检测的信号。由此,它为大型语言模型(LLM)提供了一种水印接口,第三方无需访问模型内部即可调用。其可靠性取决于LLM能否遵循该指令而不降低回答质量,然而目前LLM在这方面的表现尚未被系统测量。我们提出了ICWBench,一个包含三个可验证ICW指令族的基准,每个指令族均从可检测性和回答质量两个维度进行评分。通过对14个前沿专有及开源LLM的评估,我们发现没有任何被评估的LLM能在全部三个指令族上同时实现这两个目标。为此,我们提出一种自包含的两阶段训练方法,既不需要从更强模型蒸馏,也不需要人工标注或预先存在的ICW指令遵循能力。第一阶段为带logits扰动的自蒸馏(Self-Distillation with Logits Perturbation, SDLP),使用同一个基础LLM同时充当教师和学生:通过在解码时施加与指令等效的logits扰动,使教师模型遵循ICW指令,并训练学生模型匹配教师模型的输出分布。第二阶段以自动验证器作为奖励进行强化学习。将该方法应用于Qwen3-14B和GPT-OSS-20B后,在三条ICW指令上的平均TPR@1%FPR分别从0.100提升至0.974,以及从0.337提升至0.968,同时在困惑度评估和LLM-as-a-Judge评估下均保持了较高的响应质量。
cs.AI / 57 / 2608.29035

EmoLASP: Emotion Recognition with Language Models and Answer Set Programming

EmoLASP:基于语言模型与答案集编程的情绪识别
Le, Thao, Thielscher, Michael
Abstract
Emotion recognition in conversations is increasingly tackled with language models, but these models can be unstable and expensive to fine-tune or to prompt with long dialogue histories. We propose EmoLASP, a framework that combines a language model with declarative reasoning via Answer Set Programming (ASP) to predict VAD scores (Valence-Arousal-Dominance) in conversations. Experiments on a widely used benchmark dataset (IEMOCAP) across six open-source LLMs (3B-120B) and two PLMs (BERT, RoBERTa) show that EmoLASP improves prediction performance compared to using the language model alone, even when the LLMs/PLMs are given no dialogue history in their prompts or input vectors. The gains are largest for prompt-only LLMs, which EmoLASP uses without any fine-tuning. However, for fine-tuned PLMs, the reasoner adds little once dialogue history is available. EmoLASP's LLM pipeline demonstrates the potential advantages of using a reasoning approach to ensure emotion prediction consistency and to reduce both the cost of fine-tuning and the cost of prompting with long dialogue histories.
Chinese Translation
对话中的情绪识别越来越多地借助语言模型来解决,但这类模型可能不稳定,且微调或使用长对话历史进行提示的成本高昂。我们提出了 EmoLASP,一个将语言模型与基于答案集编程(Answer Set Programming, ASP)的声明式推理相结合的框架,用于预测对话中的 VAD 分数(效价-唤醒度-支配度,Valence-Arousal-Dominance)。在一个广泛使用的基准数据集(IEMOCAP)上,我们在六个开源大语言模型(LLM,参数规模 3B-120B)和两个预训练语言模型(PLM:BERT、RoBERTa)上进行了实验。结果表明,即使在大语言模型/预训练语言模型的提示或输入向量中不提供对话历史,EmoLASP 相比仅使用语言模型也能提升预测性能。收益最大的是仅依赖提示的 LLM——EmoLASP 无需任何微调即可使用它们。然而,对于经过微调的 PLM,在已有对话历史的情况下,推理器带来的提升甚微。EmoLASP 的 LLM 流程展示了采用推理方法来确保情绪预测一致性、并降低微调成本以及长对话历史提示成本的潜在优势。
cs.AI / 58 / 2608.29054

Let Prompts Bridge Defense Knowledge: Transferable Graph Purification via Vulnerability-Aware GPL

让提示词桥接防御知识:基于漏洞感知GPL的可迁移图净化方法
Xue, Shuomin, Li, Jingyuan, Jia, Ju, Yu, Jingxuan, Jia, Xiaojun
Abstract
Graph Neural Networks (GNNs) have emerged as a cornerstone for representing complex relational dependencies in diverse multimedia tasks, particularly in cross-platform user interest modeling and cross-modal semantic alignment. In the real world, a practical defense against graph adversarial perturbations is needed. However, we observe that the prevailing adversarial purification methods are essentially domain-restricted defenses, which leads to the following shortcomings: (1) single-domain data provides insufficient structural and semantic diversity for learning robust purification criteria; (2) training of domain-specific defense strategies from scratch consumes substantial computational cost. To address the above limitations, we propose a transferable graph purification scheme, named ProGAP, to bridge adversarial defense knowledge via vulnerability-aware graph prompt learning. Firstly, to capture universal adversarial patterns, a perturbation-capture edge detector is pretrained on data-rich graphs by jointly modeling topological and semantic information. Subsequently, to achieve more knowledge transfer w.r.t. robustness, vulnerability-aware prompts are designed that inject targeted purification guidance into biased nodes, during which the pretrained detector adapts to distribution shifts in downstream graphs without parameter-laborious updates. Experimental results demonstrate that compared with state-of-the-art baselines, our ProGAP achieves 1%-9% improvement, and reduces the time consumption by up to 2.2x. The code for ProGAP is available at https://github.com/Lieyoufffff/ProGAP.
Chinese Translation
图神经网络(GNN)已成为表示复杂关系依赖的基石,广泛应用于各类多媒体任务,尤其是跨平台用户兴趣建模和跨模态语义对齐。在现实世界中,需要一种针对图对抗扰动的实用防御方法。然而,我们观察到目前主流的对抗净化方法本质上属于领域受限的防御,存在以下不足:(1)单一领域数据无法为学习鲁棒的净化准则提供足够的结构和语义多样性;(2)从零开始训练领域特定的防御策略需要消耗大量计算成本。为解决上述局限,我们提出了一种可迁移的图净化方案 ProGAP,通过漏洞感知的图提示学习来桥接对抗防御知识。首先,为捕获通用的对抗模式,我们在数据丰富的图上预训练一个扰动捕获边检测器,联合建模拓扑与语义信息。随后,为实现更多关于鲁棒性的知识迁移,我们设计了漏洞感知提示,将针对性的净化引导注入存在偏差的节点,在此过程中预训练的检测器无需耗费大量参数更新即可适应下游图中的分布偏移。实验结果表明,与最先进的基线方法相比,我们的 ProGAP 取得了1%至9%的性能提升,并将时间消耗最多降低了2.2倍。ProGAP 的代码已发布于 https://github.com/Lieyoufffff/ProGAP。
cs.AI / 59 / 2608.29063

Agent2UCB: Agentic System for Generative Engine Optimization

Agent2UCB:面向生成式引擎优化的智能体系统
Yu, Sheldon, Wang, Rui, Yu, Tong, Kim, Sungchul, Dogan, Doga, Wu, Junda, McAuley, Julian
Abstract
Large language model driven search engines such as Google AI Overviews and Perplexity have created new opportunities for Generative Engine Optimization (GEO) the practice of refining content to increase its likelihood of being cited or summarized by generative systems. We demonstrate Agent2UCB, an agentic GEO system that autonomously improves content visibility through customized, feedback-driven optimization. For each content item, the system evaluates nine GEO strategies, identifies the most effective method, and accelerates selection using a bandit-based Agent2UCB policy that integrates LLM priors with online reward signals. To monitor side effects, the system also provides a lightweight, text-only SEO readiness evaluation covering readability, topical coverage, and EEAT-style credibility. Experiments on GEO-Bench show consistent visibility gains while preserving SEO quality. The demo allows users to choose the websites of interest, observe the optimization workflow, and compare GEO/SEO outcomes across methods.
Chinese Translation
以Google AI Overviews和Perplexity为代表的大语言模型驱动的搜索引擎,为生成式引擎优化(Generative Engine Optimization, GEO)创造了新的机遇。GEO是指通过优化内容来提升其被生成式系统引用或总结的可能性。我们提出了Agent2UCB,一个能够通过定制化、反馈驱动的优化方式自主提升内容可见性的智能体GEO系统。对于每个内容条目,该系统评估九种GEO策略,识别最有效的方法,并利用基于多臂老虎机的Agent2UCB策略加速选择,该策略将大语言模型的先验知识与在线奖励信号相结合。为监测副作用,系统还提供了一种轻量级的纯文本SEO就绪度评估,涵盖可读性、主题覆盖度以及符合EEAT风格的可信度。在GEO-Bench上的实验表明,该系统在保持SEO质量的同时带来了一致的可见性提升。该演示系统允许用户选择感兴趣的网站,观察优化流程,并比较不同方法下的GEO/SEO效果。
cs.AI / 60 / 2608.29073

Revolutionizing Turn-by-Turn Navigation with Cloud-Edge Deep Learning

基于云边协同深度学习的转向导航革命
Yang, Yiming, Fu, Hao, Zeng, Fanxiang, Yang, Xikai, Liu, Yue, Guo, Ning
Abstract
Turn-by-turn (TBT) navigation systems are integral to modern driving experiences, providing real-time audio instructions to guide drivers safely to destinations. However, existing audio instruction policy often relies on rule-based approaches that struggle to balance informational content with cognitive load, potentially leading to driver confusion or missed turns in complex environments. To overcome these difficulties, we first model the generation of navigation instructions as a multi-task learning problem by decomposing the audio content into combinations of modular elements. Then, we propose a novel deep learning framework that leverages the powerful spatiotemporal information processing capabilities of Transformers and the strong multi-task learning abilities of Mixture of Experts (MoE) to generate real-time, context-aware audio instructions for TBT driving navigation. A cloud-edge collaborative architecture is implemented to handle the computational demands of the model, ensuring scalability and real-time performance for practical applications. Experimental results in the real world demonstrate that the proposed method significantly reduces the yaw rate (the proportion of vehicles deviating from navigation routes) compared to traditional methods, delivering clearer and more effective audio instructions. This is the first large-scale application of deep learning in driving audio navigation, marking a substantial advancement in intelligent transportation and driving assistance technologies.
Chinese Translation
逐向导航(Turn-by-Turn, TBT)系统是现代驾驶体验的重要组成部分,通过提供实时语音指令引导驾驶员安全到达目的地。然而,现有的语音指令策略通常依赖于基于规则的方法,难以在信息内容与认知负荷之间取得平衡,可能导致驾驶员在复杂环境中产生困惑或错过转弯。为克服这些困难,我们首先将导航指令的生成建模为一个多任务学习问题,将语音内容分解为模块化元素的组合。然后,我们提出了一种新颖的深度学习框架,利用 Transformer 强大的时空信息处理能力和混合专家模型(Mixture of Experts, MoE)出色的多任务学习能力,为 TBT 驾驶导航生成实时的、情境感知的语音指令。为实现该模型的计算需求,我们采用了云边协同架构,确保实际应用中的可扩展性和实时性能。真实世界中的实验结果表明,与传统方法相比,所提出的方法显著降低了偏航率(即车辆偏离导航路线的比例),提供了更清晰、更有效的语音指令。这是深度学习在驾驶语音导航领域的首次大规模应用,标志着智能交通与驾驶辅助技术的一项重大进步。
cs.AI / 61 / 2608.29074

Nested Convex-Body Chasing for Online Optimization with Evolving Feasible Sets

面向可行集演化的在线优化的嵌套凸体追踪
Sarkar, Dhruv, Chakrabartty, Aprameyo
Abstract
We study online optimization with nested shrinking feasible regions in two settings: convex optimization with nested evolving feasible sets (CONES) and adversarial constrained online convex optimization (COCO). Our algorithms separate loss control from geometric movement: constrained minimizers and cumulative-loss tests preserve regret guarantees, while a deterministic resettable nested convex-body chaser limits movement. For CONES with a $G$-Lipschitz, $\mu$-strongly convex objective on a diameter-$D$ domain, we chase intersections of the current feasible set with adaptive objective sublevel sets. Using the Euclidean chasing ratio $O(\sqrt{d\log(1+d)})$, we obtain nonpositive regret at every prefix and movement $O(\sqrt{d\log(1+d)\,GD\log(eT)/\mu})$. The bound adapts to the increase in the constrained optimum value. In dimension two, with all other parameters fixed, every randomized algorithm with terminal expected regret $O(T^\beta)$, $\beta<1$, suffers $\Omega(\sqrt{\log T})$ expected movement on some deterministic nested sequence, proving optimal horizon dependence. Under linear growth away from the constrained minimizer set, Steiner-point tracking yields movement independent of $T$. For general convex COCO, one-step-delayed chasing with regularized-leader resets gives regret $O(G_fD\sqrt{d\log(1+d)T})$ and cumulative constraint violation $O(G_gD\sqrt{d\log(1+d)T})$. For strongly convex losses, both are $O(d\log(1+d)\log(eT))$ when other parameters are fixed. These reductions replace the $O(d^{d/2})$ projection-path factor in prior analyses by the polynomial dimension dependence of Euclidean nested convex-body chasing.
Chinese Translation
我们研究了具有嵌套收缩可行域的在线优化问题,涵盖两种情形:具有嵌套演化可行集的凸优化(CONES)和对抗性约束在线凸优化(COCO)。我们的算法将损失控制与几何移动相分离:约束极小化子和累积损失测试保证遗憾界,而一个确定性的可重置嵌套凸体追踪器限制移动量。对于在直径为 $D$ 的域上具有 $G$-Lipschitz、$\mu$-强凸目标函数的 CONES 问题,我们追踪当前可行集与自适应目标函数次水平集的交集。利用欧几里得追踪比 $O(\sqrt{d\log(1+d)})$,我们在每个前缀上均获得非正遗憾,且移动量为 $O(\sqrt{d\log(1+d)\,GD\log(eT)/\mu})$。该界随约束最优值的增加而自适应变化。在二维情形下,固定其他参数后,任何终端期望遗憾为 $O(T^\beta)$($\beta<1$)的随机算法在某些确定性嵌套序列上的期望移动量均为 $\Omega(\sqrt{\log T})$,从而证明了关于时间视野的最优依赖性。在约束极小化集之外目标函数呈线性增长时,基于 Steiner 点的追踪方法的移动量与 $T$ 无关。对于一般的凸 COCO 问题,采用正则化领导者重置的延迟一步追踪方法可获得遗憾 $O(G_fD\sqrt{d\log(1+d)T})$ 和累积约束违反 $O(G_gD\sqrt{d\log(1+d)T})$。对于强凸损失,在其他参数固定时,二者均为 $O(d\log(1+d)\log(eT))$。这些归约将先前分析中的 $O(d^{d/2})$ 投影路径因子替换为欧几里得嵌套凸体追踪的多项式维度依赖。
cs.AI / 62 / 2608.29088

HANIA: Planner-Guided Multimodal Graph Evidence Selection for Grounded Question Answering

HANIA:面向有据问答的规划器引导多模态图证据选择方法
Ali, Zafar, Khan, Asad, Thierry, Nimbeshaho, Amir, Nabila, Mohammed, Adam A. Q., Kefalas, Pavlos
Abstract
Multimodal question answering remains sensitive to noisy, incomplete, and weakly grounded evidence. Long unstructured contexts can introduce redundancy and encourage unsupported generation, while flat retrieval may overlook relations needed for multi-step reasoning. We present HANIA, a planner-guided multimodal graph framework for evidence-grounded question answering. HANIA processes the supplied image and text using a frozen vision-language model to extract concise question-relevant visual evidence with explicit abstention. It then constructs an input-grounded multimodal graph and applies a two-group finite-state planner to coordinate descriptive and relational evidence. Coverage-aware pruning retains a compact evidence set based on relevance, graph confidence, concept coverage, and modality diversity. The selected passages, visual statements, and graph triples are provided to a frozen instruction-tuned decoder. We evaluate HANIA on ScienceQA using answer accuracy, evidence-filtering quality, evidence-budget sensitivity, and efficiency. The results show that structured evidence planning and compact graph-guided retrieval can support competitive multimodal question answering without target-dataset fine-tuning or iterative retrieval. The code is available at https://github.com/Zafar-southeast/HANIA.
Chinese Translation
多模态问答对噪声、不完整和弱依据的证据仍然十分敏感。冗长的非结构化上下文会引入冗余信息并导致缺乏支持的生成,而扁平化检索则可能忽略多步推理所需的关系。我们提出了HANIA,一个用于证据依据问答的规划器引导多模态图框架。HANIA使用冻结的视觉-语言模型处理输入的图像和文本,提取简洁的与问题相关的视觉证据,并支持明确的弃权机制。随后,它构建一个基于输入的多模态图,并采用两组有限状态规划器来协调描述性和关系性证据。覆盖感知剪枝基于相关性、图置信度、概念覆盖度和模态多样性,保留一个紧凑的证据集合。所选的文本段落、视觉陈述和图三元组被提供给一个冻结的指令微调解码器。我们在ScienceQA上从答案准确率、证据过滤质量、证据预算敏感性和效率等方面评估了HANIA。结果表明,结构化的证据规划和紧凑的图引导检索能够在无需目标数据集微调或迭代检索的情况下,支持具有竞争力的多模态问答。代码可在 https://github.com/Zafar-southeast/HANIA 获取。
cs.AI / 63 / 2608.29092

EviAnchor: Mitigating Hallucinations in Large Vision-Language Models via Regional Visual Evidence Compensation

EviAnchor:通过区域视觉证据补偿缓解大型视觉语言模型中的幻觉问题
Jia, Sihang, Liu, Shuliang, Yang, Songbo, Hu, Xuming
Abstract
Large vision-language models (LVLMs) frequently generate content unsupported by visual inputs. Preliminary experiments show that visual evidence is primarily incorporated into answer-side representations in early-to-middle decoder layers, while its direct influence progressively weakens in later layers. This attenuation suggests that visual evidence acquired earlier may be insufficiently utilized during subsequent generation. Based on this observation, we propose EviAnchor, a training-free and single-branch inference framework that preserves and reactivates visual evidence throughout generation. EviAnchor introduces Regional Evidence Anchor (REA) slots to progressively aggregate dense visual tokens into spatially structured representations. It then strengthens the current decision state's access to these visual anchors through decision-conditioned evidence routing, mitigating excessive dependence on textual context. Finally, the model resumes its native Transformer computation to integrate the retrieved visual evidence with question semantics and generation history. Experiments across POPE, CHAIR, and MMHal-Bench demonstrate consistent improvements in visual grounding.
Chinese Translation
大型视觉语言模型(LVLMs)经常生成缺乏视觉输入支持的内容。初步实验表明,视觉证据主要在解码器早期至中间层被纳入答案侧表示,而其直接影响在较后的层中逐渐减弱。这种衰减表明,早期获取的视觉证据可能在后续生成过程中未被充分利用。基于这一观察,我们提出了EviAnchor,这是一个无需训练、单分支的推理框架,可在整个生成过程中保留并重新激活视觉证据。EviAnchor引入了区域证据锚点(Regional Evidence Anchor, REA)槽位,将密集的视觉token逐步聚合为空间结构化的表示。随后,通过基于决策条件的证据路由机制,增强当前决策状态对这些视觉锚点的访问能力,从而缓解对文本上下文的过度依赖。最后,模型恢复其原生的Transformer计算,将检索到的视觉证据与问题语义和生成历史相融合。在POPE、CHAIR和MMHal-Bench上的实验表明,该方法在视觉定位能力上取得了一致的提升。
cs.AI / 64 / 2608.29098

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

SafeAtlas-VL:基于大规模数据与守护模型超越二元多模态安全性
Wang, Zongrui, Zhu, Xiangyang, Wang, Sicheng, Wang, Han, Rong, Dingyi, Zhang, Zeyu, Li, Chunyi, Shi, Yue, Zhang, Kaiwei, Zhang, Zicheng, Tian, Yuan, Jia, Qi, Teng, Yan, Sun, Wei, Liu, Ning, Zhai, Guangtao
Abstract
Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decision. Consequently, risk becomes difficult to compare across a multimodal interaction, and ambiguous cases are obscured. We introduce SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale. We curate a broad collection of safety-relevant data from both real-world and synthetic sources and apply a disagreement-aware annotation procedure. The resulting dataset spans 15 harm categories and 55 fine-grained subcategories, covering a broad range of multimodal safety scenarios. We also construct SafeAtlas-Bench, a held-out set of 5,000 instances for evaluating five-level predictions and continuous risk scores. Upon this dataset, we train the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection. Our models not only perform five-way classification of safety levels but also map safety to continuous scores through a soft cumulative ordinal head. Experimental results demonstrate that guard models trained on our dataset exhibit strong generalization: even without using the training sets of other benchmarks, they achieve competitive performance on the corresponding test sets. Notably, our 8B model attains the overall best performance, outperforming the previous SOTA by approximately 4% in F1 score. Code, data, and models are released to support further research. Warning: this paper contains example data that may be offensive, harmful, graphic, or disturbing.
Chinese Translation
多模态安全审核需要区分来自视觉内容、用户意图和助手行为的风险。然而,现有安全防护措施通常仅针对单一判定目标进行训练,并将安全评估简化为二元决策。因此,跨多模态交互的风险难以比较,模糊案例也被掩盖。我们提出了SafeAtlas-VL数据集,包含150万个训练实例,将图像级、请求级和响应级判定置于五级有序量表上。我们从真实世界和合成来源中筛选了大量与安全相关的数据,并采用了分歧感知的标注流程。最终的数据集涵盖15个危害类别和55个细粒度子类别,覆盖广泛的多模态安全场景。我们还构建了SafeAtlas-Bench,一个包含5000个实例的保留测试集,用于评估五级预测和连续风险分数。基于该数据集,我们通过目标条件微调训练了SafeAtlas Guard系列模型,用于多模态安全检测。我们的模型不仅能对安全等级进行五分类,还能通过软累积有序头(soft cumulative ordinal head)将安全映射为连续分数。实验结果表明,在我们的数据集上训练的守护模型具有强大的泛化能力:即使不使用其他基准的训练集,它们在相应测试集上也取得了具有竞争力的性能。值得注意的是,我们的8B模型取得了总体最佳性能,在F1分数上超越此前SOTA约4%。代码、数据和模型均已发布以支持进一步研究。警告:本文包含可能具有冒犯性、有害性、血腥或令人不安的示例数据。
cs.AI / 65 / 2608.29102

Clustering as Approximation by Constrained Projectors: Theory and Guarantees

基于约束投影子的聚类逼近理论及其保证
Majumdar, Angshul
Abstract
This paper develops a unified theoretical framework showing that a broad family of clustering methods, including k-means, fuzzy c-means, kernel k-means, kernel FCM, and spectral clustering, can all be expressed as structured low-rank projectors acting on a signal-derived matrix. By formulating each method as an instance of min over B in C of ||M - M P_B||_F^2, with different constraint sets C, we establish a common optimization template that clarifies the algebraic links among hard, fuzzy, kernel-induced, and orthonormal projections. Within this framework, we derive non-trivial theoretical results, including geodesic convexity properties on the projection manifold, perturbation bounds quantifying stability to matrix noise, and exact recovery guarantees under ideal block-model conditions. The analysis further explains when different clustering families collapse to the same optimal subspace and how deviations arise under small inter-cluster leakage. Overall, the work provides a coherent, theory-first foundation for understanding clustering through structured projectors.
Chinese Translation
本文构建了一个统一的理论框架,表明包括 k-means、模糊 c-means(fuzzy c-means)、核 k-means、核 FCM 以及谱聚类在内的众多聚类方法,都可以表示为作用于信号派生矩阵的结构化低秩投影子(projector)。通过将每种方法形式化为 min_{B∈C} ||M - M P_B||_F^2 的实例,并采用不同的约束集 C,我们建立了一个统一的优化模板,阐明了硬划分投影、模糊投影、核诱导投影与正交投影之间的代数联系。在该框架内,我们推导出一系列非平凡的理论结果,包括投影流形上的测地凸性、刻画矩阵噪声稳定性的扰动界,以及理想块模型条件下的精确恢复保证。该分析进一步解释了不同聚类族何时收敛到相同的最优子空间,以及在小规模簇间泄漏下偏差如何产生。总体而言,本工作为通过结构化投影子理解聚类提供了一个连贯的、理论优先的基础。
cs.AI / 66 / 2608.29118

Emergent Misalignment Is Not Magical

涌现性失准并非神秘莫测
Li, Mingxuan, Dai, Qirun, Wang, Heran, Tan, Chenhao
Abstract
Fine-tuning large language models (LLMs) on narrowly harmful datasets can lead to misalignment broadly, a phenomenon known as emergent misalignment (EM). EM poses a challenge for AI safety and our understanding of LLMs. Prior work often frames EM as an unexpected behavior, and explains it by appealing to general misalignment directions or anthropomorphizing it as acquiring an evil persona. However, the mechanisms behind these framings remain obscure. In this work, we show that EM is a predictable and data-dependent generalization phenomenon. By examining the base model's representation of EM training data and evaluation prompts, we find that evilness after EM training is highly predictable from representational distance: the closer an evaluation prompt is to training data centroid, the more evilness it elicits from EM models after training (with an average Spearman correlation of -0.73 across 12 model-dataset settings). Building upon this analysis, we further demystify EM by showing that (1) its effectiveness changes significantly based on training data format; (2) there is not a general misalignment direction that transfers across different EM models; (3) the effect of EM is fundamentally different from persona changes. Furthermore, we extend the EM generalization metric from a scalar distance to a dataset-specific generalization direction, which robustly predicts EM models' evilness under semantics-preserving prompt perturbations including appending random tokens and paraphrasing, where other methods do not reliably generalize.
Chinese Translation
在狭窄的有害数据集上对大型语言模型(LLM)进行微调,可能导致其在大范围上失准,这一现象被称为涌现性失准(Emergent Misalignment, EM)。EM 对人工智能安全以及对 LLM 的理解构成了挑战。先前的工作往往将 EM 视为一种出乎意料的行为,并通过诉诸通用的失准方向或将其拟人化为获得邪恶人格来加以解释。然而,这些解释背后的机制仍不清楚。在本工作中,我们表明 EM 是一种可预测的、依赖于数据的泛化现象。通过考察基础模型对 EM 训练数据和评估提示词的表征,我们发现 EM 训练后的邪恶程度可以从表征距离高度预测:评估提示词与训练数据质心越接近,EM 模型在训练后表现出的邪恶程度越高(在 12 种模型-数据集设置中平均 Spearman 相关系数为 -0.73)。基于这一分析,我们进一步揭开了 EM 的神秘面纱,证明:(1)其有效性随训练数据格式的不同而发生显著变化;(2)不存在一个可跨不同 EM 模型迁移的通用失准方向;(3)EM 的效应与人格改变存在本质区别。此外,我们将 EM 泛化度量从标量距离扩展为数据集特定的泛化方向,该方向能够稳健地预测 EM 模型在语义保持的提示词扰动(包括追加随机 token 和改写)下的邪恶程度,而其他方法则无法可靠地泛化。
cs.AI / 67 / 2608.29127

Beyond Correctness: Validity-Oriented Evaluation of Biomedical LLM Judges

超越正确性:面向有效性的生物医学大语言模型评判者评估
de Oliveira, Rodrigo, Pittino, Federico, Gwinnutt, James, Nanavati, Jay
Abstract
We propose a scalable, validity-oriented pipeline for evaluating biomedical LLM judges when high-quality human judgments are scarce. First, we augment existing human-labelled biomedical benchmarks with deterministic, metric-grounded mutations that produce auditable preference pairs. Second, we evaluate judges beyond aggregate correctness using three deployment-relevant dimensions: correctness against metric-derived gold labels, robustness under repeated stochastic sampling, and compliance with the requested output format. We use this pipeline to assess Llama-3.1-8B-Instruct under four regimes: (1) base, using the instruct model as is; (2) SFT, distillation-based supervised fine-tuning only; (3) RL, GRPO-based reinforcement learning only; and (4) SFT$\rightarrow$RL, SFT followed by RL. The base and single-stage regimes struggle on structured medical discrimination such as PICO extraction and clinical calculations, whereas SFT$\rightarrow$RL performs best across correctness, compliance, and robustness; gains concentrate on decomposable tasks (PICO, MedCalc), at times matching or outperforming frontier models.
Chinese Translation
我们提出了一种可扩展的、面向有效性的评估流水线,用于在高质量人工评判稀缺的情况下评估生物医学领域的大语言模型评判者(LLM judges)。首先,我们利用确定性的、基于指标的变换来扩充现有人工标注的生物医学基准,从而生成可审计的偏好对。其次,我们从三个与部署相关的维度对评判者进行超越总体正确性的评估:相对于基于指标导出的金标准标签的正确性、在重复随机采样下的鲁棒性,以及对所要求输出格式的遵从性。我们使用该流水线在四种训练机制下评估 Llama-3.1-8B-Instruct:(1) base,直接使用指令模型;(2) SFT,仅进行基于蒸馏的监督微调;(3) RL,仅进行基于 GRPO 的强化学习;(4) SFT→RL,先 SFT 后 RL。base 和单阶段机制在结构化医学判别任务(如 PICO 信息抽取和临床计算)上表现不佳,而 SFT→RL 在正确性、遵从性和鲁棒性方面均表现最佳;其提升集中于可分解的任务(PICO、MedCalc),在某些情况下可达到甚至超越前沿模型的表现。
cs.AI / 68 / 2608.29128

APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows

APIFlow-Bench:衡量智能体能否胜任长程依赖型API工作流
Wan, Zelin, Nourian, Arash, Li, Xiaoxiao, Nandan, Nihar, Nandagopal, Kamalakannan
Abstract
Tool-using agents are commonly evaluated by a single bit: whether an end-to-end workflow completed. This metric fails to distinguish failures that matter in production, such as expired credentials, malformed payloads, or correct execution followed by incorrect final delivery. We introduce APIFlow-Bench, a fully auditable benchmark for long-horizon, dependent REST-API workflows that decomposes performance into seven engineering capabilities and requires agents to produce answers supported by the actual call path. We generate synthetic API worlds forward, subtask by subtask; each subtask is admitted only after a zero-LLM self-test triad verifies its grader and an oracle establishes solvability, and an adversarial audit identified and fixed six grader exploits. Grading is deterministic and provenance-sensitive: a state check traces a mock-minted canary through the API data flow to the response the answer must originate from, and a typed answer card is verified field by field. We release all answer keys and 44,362 unredacted execution transcripts. Across 19 frontier and open-weight models under one neutral scaffold, we find: (1) longer dependency chains degrade success, from 93% on individual subtasks to 74% on clean 20-subtask chains and 61% when including the 8% of chain trials that a model-consensus screen flags as passed by no model; (2) reliability separates models more than best-case capability, with best-of-five spanning seven points but all-five-of-five reliability spanning 44 points; (3) the independent-error account of compounding failure does not fit the data: pass rates on 20-subtask chains are 33 percentage points above the product of subtask-level rates, and on the clean slice 77% of failing runs reached the correct final state and failed only at delivery.
Chinese Translation
工具使用型智能体通常仅以一个二元指标来评估:端到端工作流是否完成。这一指标无法区分生产环境中真正重要的失败类型,例如凭证过期、载荷格式错误,或执行正确但最终交付错误。我们提出APIFlow-Bench,这是一个完全可审计的基准,面向长程、依赖型的REST API工作流,将性能分解为七项工程能力,并要求智能体给出的答案必须由实际的调用路径所支撑。我们逐个子任务前向生成合成API世界;每个子任务只有在通过一个零LLM自测三元组(验证其评分器)并由一个预言机(oracle)确立可解性之后才被采纳,同时一次对抗性审计发现并修复了六个评分器漏洞。评分是确定性的且溯源敏感:状态检查通过API数据流追踪一个模拟铸造的金丝雀(canary),直至答案必须源自的响应,并逐字段验证一个类型化的答案卡片。我们发布了全部答案密钥以及44,362条未脱敏的执行记录。在统一的中性框架下对19个前沿及开源权重模型进行评估,我们发现:(1) 更长的依赖链会降低成功率——从单个子任务的93%降至干净的20子任务链的74%,以及61%(包含8%被模型共识筛选标记为无任何模型通过的链式试验);(2) 可靠性比最佳能力更能区分模型——五次中最佳成绩的跨度为七个百分点,而五次全成的可靠性跨度达44个百分点;(3) 基于独立误差的级联失败假说与数据不符:20子任务链上的通过率比子任务级通过率的乘积高出33个百分点,且在干净数据切片上,77%的失败运行已到达正确的最终状态,仅在交付环节失败。
cs.AI / 69 / 2608.29139

More Perspectives, Stronger Signals: Multi-Perspective Enhancement and Progressive Fusion for Multimodal Entity Representation Learning

视角越多,信号越强:面向多模态实体表示学习的多视角增强与渐进式融合方法
Xiong, Chenyi, Zhang, Yan, Hu, Jing, Qin, Ziyue, Xiao, Kui, Lyu, Xiaopan, Hou, Xiaoju, Li, Zhifei
Abstract
Learning effective multimodal entity representations is fundamental for reasoning tasks such as multimodal knowledge graph completion (MMKGC). However, existing methods often suffer from semantic over-smoothing within modalities and ineffective noise filtration across modalities, particularly under sparse or ambiguous conditions. To overcome these limitations, we propose PrismF, a unified framework that synergizes multi-perspective enhancement with progressive fusion to extract stronger signals from diverse inputs. PrismF enhances fine-grained intra-modal semantics through a multi-perspective mechanism that decomposes each modality into complementary views and constrains them with a decoupling loss to reduce representation collapse. Furthermore, it improves cross-modal integration through a progressive fusion strategy that dynamically calibrates inter-modal interactions, enabling the model to emphasize informative signals while suppressing noisy or unreliable ones. Extensive experiments on three public benchmarks show that PrismF achieves the strongest overall performance, including relative improvements of 4.04% in MRR and 11.17% in Hits@1 on KVC16K. Our code can be found at https://github.com/HubuKG/PrismF.
Chinese Translation
学习有效的多模态实体表示是多模态知识图谱补全(MMKGC)等推理任务的基础。然而,现有方法往往存在模态内语义过度平滑以及模态间噪声过滤不足的问题,尤其是在数据稀疏或语义模糊的条件下。为克服这些局限,我们提出了PrismF,一个将多视角增强与渐进式融合相结合的统一框架,能够从多样化输入中提取更强的信号。PrismF通过一种多视角机制增强细粒度的模态内语义,该机制将每个模态分解为互补的视图,并利用解耦损失对其进行约束,以减少表示塌缩。此外,它通过渐进式融合策略改进跨模态整合,动态校准模态间交互,使模型能够强调有信息量的信号,同时抑制噪声或不可靠的信号。在三个公开基准数据集上的大量实验表明,PrismF取得了最强的综合性能,包括在KVC16K上MRR相对提升4.04%、Hits@1相对提升11.17%。我们的代码可在 https://github.com/HubuKG/PrismF 获取。
cs.AI / 70 / 2608.29168

JudgePanel: A Compact Judge with Panel Deliberation via Adaptive Multi-Reward Reinforcement Learning

JudgePanel:基于自适应多重奖励强化学习的具有评审团合议能力的紧凑型评判模型
Qian, Yiyue, Zhang, Shinan, Song, Huan, Marlowe, Hannah
Abstract
The LLM-as-a-Judge paradigm has emerged as a scalable alternative to human evaluation. However, single-model judges are limited by their inherent model biases, while multi-agent evaluation protocols that mitigate this through diverse deliberation are prohibitively expensive at inference time. To this end, we propose \textbf{\modelname}, which equips a compact \underline{Judge} model with multi-agent \underline{Panel} deliberation capability. Specifically, we first train on panel deliberation traces from an ensemble of strong evaluators, capturing structured patterns of discussion, disagreement, and resolution. To further improve judgment quality beyond SFT, we introduce \textit{AdaReward}, an adaptive multi-reward RL algorithm that dynamically rebalances reward component weights as different objectives saturate at different rates during RL training. For practical deployment, we further design a lightweight domain specialization module for rapid adaptation to new evaluation domains with few hundred labeled samples. As a result, (i) \textit{Novel}: the first framework to equip a single compact judge with multi-agent panel deliberation capability at single-model inference cost; (ii) \textit{Effective \& Reliable}: JudgePanel with a 14B backbone outperforms judge-specialized models up to 70B across four evaluation benchmarks, demonstrates strong position consistency, and rapidly specializes to new domains with few hundred samples.
Chinese Translation
"LLM作为评判者"(LLM-as-a-Judge)范式已成为人类评估的一种可扩展替代方案。然而,单模型评判者受限于其固有的模型偏见,而通过多样化合议来缓解这一问题的多智能体评估协议在推理时代价过于高昂。为此,我们提出 \textbf{JudgePanel},为一个紧凑的评判(Judge)模型配备多智能体评审团(Panel)合议能力。具体而言,我们首先在由一组强大评估器生成的评审团合议轨迹上进行训练,以捕捉讨论、分歧与达成一致的 structured 模式。为了在监督微调(SFT)之外进一步提升评判质量,我们提出 \textit{AdaReward},这是一种自适应多重奖励强化学习算法,能够在RL训练过程中不同目标以不同速率达到饱和时,动态地重新平衡各奖励分量的权重。在实际部署方面,我们进一步设计了一个轻量级的领域专精化模块,仅需几百个标注样本即可快速适应新的评估领域。研究成果如下:(i)\textit{新颖性}:首个以单模型推理成本为单个紧凑评判模型配备多智能体评审团合议能力的框架;(ii)\textit{有效且可靠}:采用14B骨干网络的JudgePanel在四个评估基准上超越了参数量高达70B的评判专用模型,展现出强位置一致性,并能通过几百个样本快速专精于新领域。
cs.AI / 71 / 2608.29175

An Explainable Coherence Score for Detecting Temporal Inconsistencies in Political News

一种用于检测政治新闻中时间不一致性的可解释连贯性评分方法
Pantea, Marius Nicusor, Groza, Adrian
Abstract
Temporal inconsistencies, such as mandates attributed outside their real interval, events presented as past before they occurred, or inverted causal sequences, are a form of political disinformation that evades style-based fake news detectors: a well-written article with a single wrong date carries no lexical signal of falsehood. This paper introduces the Temporal Coherence Score (TCS), a continuous, intrinsically interpretable metric that quantifies the temporal coherence of a news article, computed by a four-stage pipeline: extraction of temporal facts, construction of a temporal knowledge graph, hierarchical verification against internal consistency rules and external reference sources, and score aggregation with automatically generated explanations. Verification combines eight internal checkers derived from Allen's interval algebra with a five-level external hierarchy ranging from a locally stored reference knowledge base of 1{,}256 curated political facts to live Wikidata SPARQL queries. On a benchmark of 100 political news articles with injected temporal errors, the system reaches a precision of 0.909 at the selected operating threshold, with a single residual false positive, a profile deliberately tuned for human-in-the-loop fact-checking assistance, where false alarms are costlier than missed detections. Unlike lexical baselines that output only a binary label, every flagged article is accompanied by the inconsistency type, the entities involved, and the reference source that contradicts the claim.
Chinese Translation
时间不一致性——例如将指令(mandates)归入其真实时间区间之外、事件在发生之前就被表述为已发生的过去事件,或因果顺序颠倒——是一种能够规避基于风格的假新闻检测器的政治虚假信息形式:一篇写得很好但只有一个日期错误的文章在词汇层面没有任何虚假信号。本文提出了时间连贯性评分(Temporal Coherence Score, TCS),这是一种连续的、具有内在可解释性的指标,用于量化新闻文章的时间连贯性。该指标通过一个四阶段流水线计算:时间事实抽取、时间知识图谱构建、基于内部一致性规则和外部参考来源的分层验证,以及带有自动生成解释的评分聚合。验证环节将由艾伦区间代数(Allen's interval algebra)导出的八个内部检查器,与一个五级外部验证层级相结合,该层级从本地存储的包含1,256条精选政治事实的参考知识库,到实时的Wikidata SPARQL查询。在一个注入了时间错误的100篇政治新闻文章基准数据集上,在选定的工作阈值下,系统的精确率达到0.909,仅有1个残留的误报;该性能特征是刻意针对“人在环路”(human-in-the-loop)事实核查辅助场景进行调优的,因为在此类场景中,误报的代价高于漏报。与仅输出二元标签的词汇基线方法不同,每一篇被标记的文章都附带有不一致类型、涉及的实体,以及与该声明相矛盾的参考来源。
cs.AI / 72 / 2608.29198

How Identity and Opinion Shape Political Sycophancy in LLMs

身份与观点如何塑造大语言模型中的政治谄媚行为
Fu, Li-Ni, Meng, Chang-Chih, Chen, Chien-Hua, Huang, Hen-Hsen, Wu, I-Chen
Abstract
As Large Language Models (LLMs) increasingly encourage users to disclose personal profiles for tailored assistance, measuring their political alignment becomes increasingly important. However, many existing benchmarks for assessing political behavior rely on closed-ended questions and do not fully capture how a model's stance may adapt to user-provided context during interaction. We introduce a framework that disentangles two distinct triggers of political sycophancy: opinion (aligning with explicit narratives) and identity (stereotyping based on demographic labels). Using 450 manually-checked political dilemmas as controlled probes, we evaluate 13 instruction-tuned LLMs. We uncover a dissociation: a model's susceptibility to explicit opinions does not necessarily predict its susceptibility to identity cues, and vice versa. When both signals are present, their effects are generally sub-additive rather than simply additive. Additionally, system-level personas primarily shift a model's baseline stance while having limited effect on the stance shift caused by user opinion or identity. Ultimately, our results suggest that LLM political stance is interactively and steerably vulnerable rather than being a fixed trait, highlighting how personalization may amplify identity- or opinion-conditioned shifts in the model's behaviors.
Chinese Translation
随着大语言模型(LLM)日益鼓励用户披露个人资料以获得定制化帮助,衡量其政治对齐倾向变得愈发重要。然而,许多现有的用于评估政治行为的基准依赖于封闭式问题,无法充分捕捉模型在交互过程中其立场如何适应用户提供的上下文。我们提出了一个框架,将政治谄媚(political sycophancy)的两种不同触发因素区分开来:观点(opinion,即迎合明确的叙事)和身份(identity,即基于人口统计标签的刻板印象)。我们使用450个经人工核验的政治两难问题作为受控探针,评估了13个指令微调的大语言模型。我们发现了一种分离现象:模型对显式观点的易感性并不一定能预测其对身份线索的易感性,反之亦然。当两种信号同时存在时,其效应通常呈次可加性,而非简单的可加性。此外,系统层面的角色设定(persona)主要改变模型的基线立场,而对由用户观点或身份引起的立场偏移影响有限。最终,我们的结果表明,LLM的政治立场是交互性的、可被引导的脆弱属性,而非固定特质,这凸显了个性化可能如何放大模型行为中由身份或观点条件引发的变化。
cs.AI / 73 / 2608.29206

Benevolent Bias in Multi-Turn Human-Agent Dialogue

人-智能体多轮对话中的善意偏见
Liu, Qianqi, Huang, Jin, Dogan, Fethiye Irmak, Gunes, Hatice
Abstract
Bias in human-agent interaction can manifest not only through hostile language but also as benevolent bias, whereby unequal treatment hides behind a warm, positive tone. To make it detectable, we operationalise benevolent bias along two dimensions, tone and treatment, yielding three classes: neutral support, overt bias, and benevolent bias. Building on these definitions, we construct BENEVDIAL, a class-balanced corpus of 362,880 multi-turn support dialogues spanning user and agent demographics, roles, and generators, to support controlled evaluation. We then test two detector families on it: off-the-shelf safety detectors and prompted large language model (LLM) judges. Our results reveal a detection gap: off-the-shelf detectors reliably flag overt bias yet largely miss benevolent bias, while LLM judges catch more under more explicit detection criteria but increasingly misclassify neutral support as benevolent bias, and demographic context amplifies the false alarms. These findings suggest that fair monitoring of human-agent dialogue must look beyond surface cues to whether the agent's treatment is disparate.
Chinese Translation
人-智能体交互中的偏见不仅可以通过敌意语言表现出来,还可以表现为善意偏见(benevolent bias),即不平等的对待隐藏在温暖、积极的语气背后。为使其可被检测,我们从语气(tone)和对待(treatment)两个维度对善意偏见进行操作化定义,得到三种类别:中性支持、明显偏见和善意偏见。基于这些定义,我们构建了BENEVDIAL,一个包含362,880段多轮支持性对话的类别均衡语料库,涵盖用户与智能体的人口统计学特征、角色和生成器,以支持可控评估。随后,我们在该语料库上测试了两类检测器:现成的安全检测器和基于提示的大语言模型(LLM)裁判。结果显示存在检测差距:现成检测器能够可靠地识别明显偏见,但在很大程度上无法察觉善意偏见;而LLM裁判在更明确的检测标准下能捕获更多案例,但也会越来越多地将中性支持误判为善意偏见,且人口统计学背景会加剧误报。这些发现表明,对人-智能体对话的公平性监测不能仅停留在表面线索,而应关注智能体的对待方式是否存在差异。
cs.AI / 74 / 2608.29207

Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling

Hyper-Fold:基于超图建模探索蛋白质序列-几何学习的表达能力极限
Feng, Yifan, Cheng, Guanjie, Ying, Shihui, Du, Shaoyi, Gao, Yue
Abstract
Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (three-dimensional geometry). What is the expressive limit of this layer class? We show that the complete bilinear operator over content-geometry outer products--the sufficient statistic of all second-order interactions--is the expressive ceiling, while the additive message passing of mainstream geometric GNNs is provably blind to content-geometry binding. We then introduce Hyper-Fold, a rank-K separable convolutional backbone approaching this ceiling at message-passing cost: each radius neighborhood is organized into a sequence hyperedge and a contact hyperedge, modulated by an edge-conditioned matrix-valued operator factorized into K learned basis operators with geometry-generated coefficients. Across enzyme function prediction, fold classification, and ligand binding site detection, Hyper-Fold and its hierarchical variant Hyper-Fold-Deep achieve the best results among protein-specific structure encoders; Hyper-Fold-Pocket, an anchored set-prediction head, surpasses UniSite-3D on UniSite-DS and two zero-shot benchmarks with no sequence language model features, 68x fewer parameters, and 4.8x lower latency--suggesting that a sufficiently expressive 3D backbone recovers information that fusion architectures previously borrowed from evolution-scale pretraining.
Chinese Translation
蛋白质结构建模建立在一个单一的计算原语之上:残基是什么(序列内容)与其所在位置(三维几何)之间的相互作用。这一层类的表达能力极限是什么?我们证明,作用于内容-几何外积上的完整双线性算子——所有二阶交互的充分统计量——构成了表达能力的上限,而主流几何图神经网络(GNN)的加性消息传递在可证明意义上无法感知内容-几何结合。随后我们提出Hyper-Fold,这是一种以消息传递的代价逼近该上限的秩K可分离卷积骨架:每个半径邻域被组织为一条序列超边和一条接触超边,并通过一个边条件化的矩阵值算子进行调制,该算子被分解为K个可学习的基算子并带有由几何生成的系数。在酶功能预测、折叠分类和配体结合位点检测任务上,Hyper-Fold及其层次化变体Hyper-Fold-Deep在蛋白质专用结构编码器中取得了最佳结果;Hyper-Fold-Pocket是一种锚定的集合预测头,在不使用序列语言模型特征的情况下,以68倍更少的参数量和4.8倍更低的延迟,在UniSite-DS及两个零样本基准上超越了UniSite-3D——这表明一个表达能力足够强的3D骨架能够恢复以往融合架构需要从进化尺度预训练中借用的信息。
cs.AI / 75 / 2608.29210

Imag-Eval: a language-grounded framework for interpretable Text-to-Image instruction following evaluation

Imag-Eval:一个基于语言的可解释文本到图像指令遵循评估框架
Serouis, Ibrahim Mohamed, Duque, David Jaramillo
Abstract
Text-to-Image (T2I) models have recently achieved impressive visual fidelity, yet their evaluation remains constrained by benchmarks that are often difficult to interpret and insufficiently diagnostic. Existing skill-based evaluations tend to overlook critical failure modes that strongly impact usability but fall outside standard taxonomies, such as global incoherence arising from missing parts or physically implausible configurations (e.g., floating objects). In addition, prompt difficulty is typically controlled along a single dimension; either prompt length or the number of elements to generate. To address these limitations, we introduce Imag-Eval, a controlled benchmark designed to assess how T2I models ground compositional natural-language instructions into visual outputs. Unlike prior work that conflates surface linguistic complexity with compositional difficulty, Imag-Eval explicitly seeks to disentangles these factors by independently varying both the number of instances and the combination of constraints (rules), while avoiding error propagation. This design enables fine-grained and interpretable analysis of where cross-modal instruction following fails. Our benchmark comprises 1,140 prompts and 8,842 combined rules, and we evaluate it on several state-of-the-art models. Complementing this analysis with an additional study of over 2,000 prompts from a concurrent benchmark, our results suggest that, for structured skills, compositional difficulty is primarily governed by the number of grounded rules and their binding to instances,, rather than by prompt length alone.
Chinese Translation
文本到图像(T2I)模型近期在视觉保真度方面取得了令人瞩目的成果,但其评估仍受限于那些往往难以解释且诊断能力不足的基准测试。现有的基于技能的评估往往忽视了那些严重影响可用性却不在标准分类体系之内的关键失败模式,例如由缺失部件或物理上不合理的配置(如悬浮物体)引起的全局不连贯性。此外,提示词难度通常仅在单一维度上进行控制,即提示词长度或需生成的元素数量。为解决这些局限,我们提出了Imag-Eval,这是一个受控基准,旨在评估T2I模型如何将组合性自然语言指令落地为视觉输出。与以往将表层语言复杂性与组合难度混为一谈的研究不同,Imag-Eval通过独立变化实例数量与约束(规则)的组合,并避免误差传播,明确地试图解耦这些因素。这一设计使得我们能够对跨模态指令遵循在何处失败进行细粒度且可解释的分析。我们的基准包含1,140个提示词和8,842条组合规则,并在多个最先进的模型上进行了评估。结合对来自同期另一基准的2,000多个提示词的补充研究,我们的结果表明,对于结构化技能而言,组合难度主要取决于已落地规则的数量及其与实例的绑定关系,而非仅仅取决于提示词长度。
cs.AI / 76 / 2608.29223

Computational Depth Measurement in Thermographic Video: Overcoming Spatial Overfitting via Spatio-Temporal Decoupling

热成像视频中的计算深度测量:通过时空解耦克服空间过拟合
Abidin, Zain Ul, Memon, Habeeban, Ahmed, Junaid
Abstract
Accurate through-thickness measurement of subsurface delamination depth in Carbon Fiber Reinforced Polymer (CFRP) is important for structural assessment because defect location determines affected load-bearing layers. Optical pulsed thermography (OPT) provides a two-dimensional thermal video rather than volumetric measurements, so depth must be inferred from temporal heat-diffusion responses. A challenge is spatial dataset bias: when calibration defects follow regular grids, regression models may memorize their geometry instead of learning physical relationship between thermal decay and depth. This work introduces a spatio-temporal decoupling architecture that separates spatial defect localization from temporal depth measurement. Defect regions are first localized using segmentation methods, after which thermal responses are spatially averaged and converted into sixteen physics-informed temporal, energy, statistical, and geometric features. These features expose the one-dimensional heat-conduction relationship while withholding pixel coordinates from the depth model. Four regression models are evaluated using specimen-level cross-validation: Random Forest (RF), Gradient Boosting Machine (GBM), Advanced Multi-Layer Perceptron (Adv-MLP), and XGBoost. Unregularized trees and over-parameterized Adv-MLP exhibit calibration collapse under geometric shifts, with errors exceeding 0.5 mm. In contrast, regularized XGBoost with L1/L2 penalties and column sampling maintains cross-specimen calibration, achieving a mean absolute error (MAE) of 0.056 mm and root mean square error (RMSE) of 0.085 mm. Predicted depths are merged with masks to generate Delaunay-triangulated three-dimensional defect models in three to five seconds per specimen. Results show that mathematical regularization and spatio-temporal decoupling reduce spatial memorization in thermal-video depth regression.
Chinese Translation
准确测量碳纤维增强聚合物(CFRP)地下分层深度的穿透厚度对于结构评估非常重要,因为缺陷位置决定了受影响的承力层数。光学脉冲热成像(OPT)提供的是二维热视频而非体积测量数据,因此必须从时间热扩散响应中推断深度。一个挑战是空间数据集偏差:当标定缺陷呈规则网格分布时,回归模型可能会记忆其几何形状,而不是学习热衰减与深度之间的物理关系。本工作提出了一种时空解耦架构,将空间缺陷定位与时间深度测量分离。首先使用分割方法定位缺陷区域,随后对热响应进行空间平均,并转换为十六个融合物理信息的时间、能量、统计和几何特征。这些特征揭示了一维热传导关系,同时对深度模型隐藏像素坐标信息。采用试件级交叉验证评估了四种回归模型:随机森林(RF)、梯度提升机(GBM)、高级多层感知机(Adv-MLP)和XGBoost。未正则化的树模型和过参数化的Adv-MLP在几何偏移下出现标定失效,误差超过0.5毫米。相比之下,采用L1/L2惩罚和列采样的正则化XGBoost保持了跨试件标定能力,实现了0.056毫米的平均绝对误差(MAE)和0.085毫米的均方根误差(RMSE)。预测深度与掩膜合并后,可在每个试件3至5秒内生成基于Delaunay三角剖分的三维缺陷模型。结果表明,数学正则化与时空解耦能够减少热视频深度回归中的空间记忆现象。
cs.AI / 77 / 2608.29228

Localizing Emergent Failures in Agentic AI: Recovering Minimal Repair Families via Counterfactual Replay

智能体AI系统中涌现故障的定位:通过反事实重放恢复最小修复族
Li, Bingjie, Song, Yumeng, Yao, Zhongming, Li, Tianyi
Abstract
Failures in agentic AI systems can arise from interactions among messages exchanged by multiple large language model (LLM) agents. Pointwise attribution cannot distinguish a jointly necessary repair from alternative singleton repairs. We formulate Minimal Repair Family Recovery (MRFR): recovering all inclusion-minimal event sets whose counterfactual replay restores task success within a declared size bound. We propose Graph-Constrained Joint Replay (GCJR), which slices failure-relevant events from an execution dependency graph, constructs graph-feasible singleton and pair candidates, and verifies them by replay with paired clean counterparts. For fixed replay outcomes, GCJR is exact within its declared graph domain. On 90 in-scope cases from a 120-DAG controlled benchmark, GCJR achieves 1.000 Family Exact Match while reducing mean replay calls from 56.3 to 25.3 (55.1%) relative to exhaustive search. On a 24-case, four-agent LLM pilot, it again achieves 1.000 Family Exact Match and reduces mean model calls from 21.0 to 10.0 (52.4%); single-event replay misses jointly necessary repairs.
Chinese Translation
智能体AI系统中的故障可能源于多个大语言模型(LLM)智能体之间交换消息的相互作用。逐点归因无法区分联合必要的修复与其他替代性的单点修复。我们形式化了最小修复族恢复问题(Minimal Repair Family Recovery, MRFR):即在给定大小约束内,恢复所有包含最小的、其反事实重放能够恢复任务成功的事件集合。我们提出图约束联合重放方法(Graph-Constrained Joint Replay, GCJR),该方法从执行依赖图中切分与故障相关的事件,构造图上可行的单事件和双事件候选,并通过与配对的干净对照样本进行重放来验证这些候选。在重放结果固定的条件下,GCJR在其声明的图域内是精确的。在一个包含120个DAG的受控基准中的90个适用案例上,GCJR达到了1.000的族精确匹配率(Family Exact Match),同时将平均重放调用次数从56.3降至25.3(降低55.1%),相较于穷举搜索有明显优势。在一个24案例、四智能体的LLM试点实验中,它同样达到1.000的族精确匹配率,并将平均模型调用次数从21.0降至10.0(降低52.4%);而单事件重放则会漏检联合必要的修复。
cs.AI / 78 / 2608.29249

Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge

FKG.in 的验证:LLM 增强的印度食品知识中的可靠性评估
Gupta, Saransh Kumar, Shah, Armaan, Dey, Lipika, Das, Partha Pratim, Jain, Ramesh
Abstract
The online culinary ecosystem is increasingly populated by recipe content generated, modified, or summarized by Large Language Models (LLMs). While often plausible, such outputs may contain hallucinated ingredients, misrepresented quantities, or culturally implausible combinations, limiting their suitability for downstream applications and knowledge graph construction. In this paper, we present a semi-automated soundness assessment workflow for validating structured recipe data extracted and augmented by LLMs from informal culinary sources. Developed as part of FKG(.in), a knowledge graph of Indian food, the pipeline identifies and addresses common failure modes, including structural inconsistencies, semantic and logical incoherence, and deviations from the source text, through a multi-stage process combining formal grammars, vocabulary-based checks, statistical heuristics, Set Transformer-based coherence modeling, and retrieval-based verification. Although evaluated on Indian recipes, the proposed methods are applicable to broader multilingual and multicultural culinary domains. We provide a practical, auditable, and application-agnostic framework for validating LLM-augmented recipe data, thereby strengthening the foundations of machine-readable food knowledge infrastructures in the era of LLM-generated content.
Chinese Translation
在线烹饪生态系统中,由大语言模型(LLM)生成、修改或总结的食谱内容日益增多。尽管这些输出常常看似合理,但其中可能包含虚构的食材、失实的用量,或在文化上不合情理的搭配,从而限制了其在下游应用和知识图谱构建中的适用性。本文提出了一种半自动化的可靠性评估工作流,用于验证由 LLM 从非正式烹饪来源中提取并增强的结构化食谱数据。该流程作为印度食品知识图谱 FKG(.in) 的一部分而开发,通过结合形式文法、基于词表的检查、统计启发式方法、基于 Set Transformer 的连贯性建模以及基于检索的验证的多阶段流程,识别并解决常见的失败模式,包括结构不一致、语义与逻辑不连贯以及偏离原文等问题。尽管本文以印度食谱为例进行评估,所提出的方法同样适用于更广泛的多语言、多文化烹饪领域。我们提供了一个实用、可审计且不依赖于特定应用的框架,用于验证 LLM 增强的食谱数据,从而在 LLM 生成内容的时代,为机器可读的食品知识基础设施奠定更坚实的基础。
cs.AI / 79 / 2608.29251

GuardianAgent: Policy-Conditioned Risk-Adaptive Anonymization with Verified Adversarial Escalation

GuardianAgent:基于策略条件化的风险自适应匿名化与经核验的对抗性升级机制
Yang, Ruiyi, Lihinikaduarachchi, Gayathri, Masood, Rahat, Salim, Flora D., Kanhere, Salil S.
Abstract
Privacy protection for live web traffic requires more than detecting private spans. Agent-based privacy protection systems must determine whether an outgoing action complies with the destination site's privacy policy, then apply only the level of rewriting or sanitisation justified by the residual disclosure risk. We present GuardianAgent, a policy-conditioned anonymization framework that couples structured risk assessment with verified adaptive rewriting. GuardianAgent computes risk through AMRSF (Adaptive Multi-factor Risk Scoring Formula), an explicit controller that combines policy-violation likelihood with data sensitivity, recipient transmission, purpose legitimacy, contextual basis, and policy transparency, rather than relying on an LLM to assign risk directly. This risk score determines both the allow/transform/deny decision and the initial anonymization level. For efficiency, GuardianAgent uses an evidential fast path for low-uncertainty policy matches and invokes an LLM slow path only for uncertain cases. For rewriting, it applies a five-level hierarchy driven by a verified adversarial guesser: guesses trigger escalation only when supported by the original text, preventing hallucinated attacker confidence from causing unnecessary over-anonymization. Experiments across three benchmarks spanning legal text (TAB), Reddit posts (SynthPAI), and multi-format synthetic PII records (PII-Masking-300k) show that GuardianAgent achieves the strongest privacy-utility trade-off among published baselines and is the only method to reach more than 0.90 privacy in all three domains, remaining robust under a backbone switch. Action-context stress tests further show that the same outgoing text receives different decisions and anonymization strengths under different recipients, purposes, action bases, and policy-transparency conditions.
Chinese Translation
面向实时网络流量的隐私保护不仅仅是检测隐私片段。基于智能体(Agent)的隐私保护系统必须判断外发操作是否符合目标网站的隐私策略,并据此仅施加与残余泄露风险相称的改写或脱敏程度。我们提出了GuardianAgent,一个将结构化风险评估与经核验的自适应改写相结合的策略条件化匿名化框架。GuardianAgent通过AMRSF(自适应多因素风险评分公式,Adaptive Multi-factor Risk Scoring Formula)计算风险。AMRSF是一个显式控制器,它综合了策略违反可能性、数据敏感度、接收方传输、目的正当性、情境依据以及策略透明度等要素,而非直接依赖大语言模型(LLM)来分配风险。该风险分数同时决定了允许/转换/拒绝的决策以及初始匿名化级别。在效率方面,GuardianAgent针对低不确定性的策略匹配采用证据式快速通道,仅对不确定情形调用LLM慢速通道。在改写方面,它应用由经核验的对抗性猜测器(verified adversarial guesser)驱动的五级层级机制:只有当猜测得到原文支持时才触发级别升级,从而防止攻击者置信度的幻觉导致不必要的过度匿名化。在涵盖法律文本(TAB)、Reddit帖子(SynthPAI)以及多格式合成个人身份信息记录(PII-Masking-300k)的三个基准上的实验表明,GuardianAgent在已发表的基线方法中实现了最强的隐私-效用权衡,并且是唯一在全部三个领域均达到0.90以上隐私分数的方法,同时在更换底层模型时仍保持稳健。动作-情境压力测试进一步表明,相同的外发文本在不同的接收方、目的、动作依据和策略透明度条件下会获得不同的决策与匿名化强度。
cs.AI / 80 / 2608.29252

Dynamic Important Example Mining for Reinforcement Finetuning

面向强化微调的动态重要样本挖掘
Tan, Haoru, Wu, Sitong, Chen, Yanfeng, Zhao, Shizhen, Sun, Yang-Tian, Liu, Tianjia, Chang, Chirui, Zhang, Shaofeng, Sun, Samm, Wu, Xiuzhe, Xie, Ruobing, Qi, Xiaojuan
Abstract
Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its effectiveness is bound by how training data are selected and used. Most data-centric RFT methods rely on static or heuristic sample selection, implicitly assuming a sample's value is fixed over training. This overlooks the non-stationary dynamics of policy learning and can lead to suboptimal updates. We propose Dynamic Important Example Mining (DIEM), a principled and fully automated framework that makes data utilization adaptive throughout RFT. DIEM integrates two components into each optimization step: (i) a gradient-alignment importance estimator that efficiently approximates each sample's marginal contribution to policy improvement; and (ii) a constrained batch reweighting scheme that maximizes aggregate utility while preserving the update's gradient magnitude to stabilize optimization. Across several reasoning benchmarks, DIEM consistently outperforms strong static and dynamic baselines. The code will be released via https://github.com/hrtan/DIEM.
Chinese Translation
强化微调(Reinforcement Fine-Tuning, RFT)日益被用于增强大模型的推理能力,但其有效性受限于训练数据的选择与使用方式。大多数以数据为中心的RFT方法依赖静态或启发式的样本选择,隐含地假设样本的价值在训练过程中是固定不变的。这忽略了策略学习的非平稳动态特性,可能导致次优的更新。我们提出动态重要样本挖掘(Dynamic Important Example Mining, DIEM),一个有原则且完全自动化的框架,使数据利用在整个RFT过程中具有自适应性。DIEM在每个优化步骤中集成两个组件:(i)一个梯度对齐的重要性估计器,可高效近似每个样本对策略提升的边际贡献;(ii)一个受约束的批量重加权方案,在保持更新的梯度幅值以稳定优化的同时,最大化总体效用。在多个推理基准上,DIEM始终优于强静态和动态基线方法。代码将通过 https://github.com/hrtan/DIEM 发布。
cs.AI / 81 / 2608.29263

RACER: Reinforced Agent Collaboration for Explainable Reasoning on Knowledge Graphs

RACER:基于强化智能体协作的知识图谱可解释推理
Lou, Yuwei, Hu, Hao, Jiang, Yuzhou, Zhang, Zongfei, Wang, Liang, Liu, Jincai, Ge, Jidong, Tao, Xianping
Abstract
Large Language Models (LLMs) often suffer from hallucination and struggle with complex reasoning tasks requiring multi-hop domain knowledge. While integrating Knowledge Graphs (KGs) provides a structured and verifiable information source, current KG-enhanced LLM paradigms usually rely on single-agent path extraction and fixed prompting, lacking adaptability and facing huge search spaces. To address these challenges, we propose RACER, a Reinforced Agent Collaboration framework for Explainable Reasoning on knowledge graphs. RACER employs a semantic-aware action pruning and teacher-guided reinforcement learning mechanism to efficiently extract high-quality reasoning pathways from large-scale KGs. Furthermore, to mitigate single-path generation pitfalls, we introduce a cross-task accumulated shared memory graph paired with an attention-driven multi-path knowledge refinement module. Finally, RACER orchestrates these components through a four-role multi-agent collaboration system (GraphAgent, TemplateAgent, AnswerAgent, and CriticAgent) to dynamically refine prompts and evaluate answers. Extensive experiments on CommonsenseQA and OpenBookQA datasets demonstrate that RACER significantly outperforms state-of-the-art KG-enhanced LLM baselines with an average improvement of 5\%, offering robust and highly interpretable reasoning capabilities.
Chinese Translation
大语言模型(LLMs)常常面临幻觉问题,并且在需要多跳领域知识的复杂推理任务上表现不佳。虽然引入知识图谱(KG)可以提供结构化且可验证的信息来源,但当前基于知识图谱增强的大语言模型范式通常依赖于单智能体路径抽取和固定提示策略,缺乏适应性且面临巨大的搜索空间。为应对这些挑战,我们提出了 RACER,一个用于知识图谱可解释推理的强化智能体协作框架。RACER 采用语义感知的动作剪枝和教师引导的强化学习机制,从大规模知识图谱中高效抽取高质量推理路径。此外,为缓解单路径生成的缺陷,我们引入了与注意力驱动的多路径知识精炼模块相配合的跨任务累积共享记忆图。最后,RACER 通过一个四角色的多智能体协作系统(GraphAgent、TemplateAgent、AnswerAgent 和 CriticAgent)来统筹上述组件,动态优化提示并评估答案。在 CommonsenseQA 和 OpenBookQA 数据集上的大量实验表明,RACER 显著优于最先进的知识图谱增强大语言模型基线,平均提升 5%,展现出强大且高度可解释的推理能力。
cs.AI / 82 / 2608.29264

EpaCache: Error-Propagation-Aware Caching for Accelerating Diffusion-Based Visual Generation

EpaCache:面向误差传播感知的缓存加速扩散视觉生成
Liu, Yuhan, Hong, Zongwei, Li, Jinglun, Li, Linze, Zhang, Shen, Tang, Yao
Abstract
Diffusion-based visual generative models deliver strong image and video synthesis quality but incur high inference costs because sequential samplers repeatedly evaluate large networks. Caching-based methods reduce inference latency by reusing intermediate computations across adjacent timesteps. However, existing cache controllers rely primarily on local temporal variation and overlook the trajectory-level consequences of cache reuse. We introduce Error-Propagation-Aware Cache (EpaCache), a training-free caching policy that adaptively allocates the reuse budget on timesteps with lower downstream impact. Experiments on image and video synthesis models demonstrate that EpaCache consistently improves the latency--fidelity trade-off over existing caching methods. On FLUX.1-dev, EpaCache outperforms the prior state-of-the-art caching method in both latency and fidelity, reducing inference time from $11.7$ s to $11.3$ s while improving PSNR from $21.4$ to $22.8$. On HunyuanVideo, EpaCache achieves a $2.63\times$ speedup over uncached inference and improves SSIM from $0.891$ to $0.905$ over the prior state-of-the-art method at matched latency.
Chinese Translation
基于扩散的视觉生成模型能够提供高质量的图像和视频合成效果,但由于顺序采样器需要反复评估大型网络,导致推理成本高昂。基于缓存的方法通过在相邻时间步之间复用中间计算来降低推理延迟。然而,现有的缓存控制器主要依赖局部时间变化,而忽视了缓存复用在轨迹层面的影响。我们提出了误差传播感知缓存(Error-Propagation-Aware Cache, EpaCache),这是一种无需训练的缓存策略,能够自适应地将复用预算分配到下游影响较小的时间步上。在图像和视频合成模型上的实验表明,EpaCache 在延迟—保真度权衡方面持续优于现有的缓存方法。在 FLUX.1-dev 上,EpaCache 在延迟和保真度两方面均超越了先前最先进的缓存方法,将推理时间从 11.7 秒降至 11.3 秒,同时将 PSNR 从 21.4 提升至 22.8。在 HunyuanVideo 上,EpaCache 相较于无缓存的推理实现了 2.63 倍的加速,并在相同延迟下将 SSIM 相较先前最先进方法从 0.891 提升至 0.905。
cs.AI / 83 / 2608.29279

Understanding Deep Learning via Entropy Space Theory

基于熵空间理论理解深度学习
Li, Li, Zhang, Tong, Yu, Wentao, Wang, Zuobin
Abstract
Deep learning is often criticized for its theoretical research lagging behind practice. To make deep learning easier to understand, the entropy space theory is first introduced here. The entropy space can cover all the possibilities of any deep learning model by topological structure. It is independent of network parameters. Through the designed fundamental operations and norm, entropy space is proven to be a normed space within the formal axiomatic framework. Based on the theory, a unified coordinate system is proposed. It can coordinatize every state of a model and rank them by compression of the maximal value of information entropy. The theory offers a novel priori framework for mathematical fundamentals of deep learning.
Chinese Translation
深度学习常因理论研究滞后于实践而受到批评。为了使深度学习更易于理解,本文首次引入了熵空间理论。熵空间能够通过拓扑结构涵盖任何深度学习模型的所有可能性,且独立于网络参数。通过所设计的基本运算和范数,证明了熵空间在形式化公理框架内是一个赋范空间。基于该理论,本文提出了一个统一的坐标系,可以对模型的每个状态进行坐标化,并按信息熵最大值的压缩程度对其进行排序。该理论为深度学习的数学基础提供了一个新颖的先验框架。
cs.AI / 84 / 2608.29286

MMPCBench: Benchmarking Multimodal Large Language Models on Proactive Critique of Flawed Inputs

MMPCBench:面向有缺陷输入主动批判的多模态大语言模型基准测试
Li, Jinzhe, Li, Gengxu, Li, Jinnan, Wu, Yuan, Chang, Yi
Abstract
As Multimodal Large Language Models (MLLMs) evolve into sophisticated interactive assistants, their reliability depends not only on following instructions but also on validating them. We define Proactive Critique as the model's autonomous ability to identify, analyze and fix faulty user inputs without extra prompts. However, evaluations mainly test models under ideal circumstances or simple refusal behaviors, largely ignoring active error processing. To fill this gap, we propose MMPCBench, a comprehensive framework for evaluating MLLMs' proactive critique competence. It features a fine-grained taxonomy of 4 primary error types spanning 12 subcategories, ranging from cross-modal contradictions to missing visual premises. We adopt a hierarchical evaluation protocol to measure models' error detection, diagnosis and resolution performance, and apply alignment-aware metrics to assess the coherence between internal reasoning and final responses. Tests on 14 mainstream MLLMs show obvious weaknesses in proactive critique, especially in dealing with subtle visual anomalies. Notably, we identify a pervasive "consistency gap": reasoning models can often correctly identify and analyze errors during internal reasoning yet suppress these valid insights in final outputs to prioritize response compliance. The code and data is available at https://github.com/ALIENS32/MMPCBench.
Chinese Translation
随着多模态大语言模型(MLLMs)发展为高度智能化的交互助手,其可靠性不仅取决于遵循指令的能力,还取决于对指令进行验证的能力。我们将“主动批判”定义为模型在无需额外提示的情况下自主识别、分析和修正错误用户输入的能力。然而,现有评估主要在理想场景下测试模型,或仅考察简单的拒绝行为,很大程度上忽视了主动的错误处理能力。为填补这一空白,我们提出MMPCBench,一个用于评估MLLM主动批判能力的综合框架。该框架包含一个细粒度的错误分类体系,涵盖4大类共12个子类的错误类型,范围从跨模态矛盾到视觉前提缺失。我们采用分层评估协议来衡量模型在错误检测、诊断和解决方面的表现,并应用对齐感知指标来评估模型内部推理与最终回答之间的一致性。对14个主流MLLM的测试表明,模型在主动批判方面存在明显不足,尤其是在处理细微视觉异常时。值得注意的是,我们发现了一种普遍存在的“一致性鸿沟”:具备推理能力的模型往往能在内部推理中正确识别和分析错误,但为了优先保证回答的顺从性,会在最终输出中抑制这些有效的洞察。代码和数据可在 https://github.com/ALIENS32/MMPCBench 获取。
cs.AI / 85 / 2608.29291

Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling

基于核心-扩展路由与统一计算调度的统一多模态模型加速方法
Zhan, Wengyi, Yan, Chenqian, Liu, Songwei, Lin, Mingbao, Ji, Rongrong
Abstract
Unified multimodal models jointly support understanding and generation, but incur substantial redundant computation across tokens, layers, and generation timesteps. Through token-importance probing, we identify an asymmetric core-expansion structure: understanding exhibits a stable importance component, while generation largely shares this component but requires progress-dependent corrections. We therefore propose CE-Router, which uses a task-shared core scorer and progress-conditioned generation expansions, optimized through generation decomposition and cross-task core alignment. At inference, CE-Router compacts token computation and supplies a learned routing signal to Unified Computation Scheduling, which coordinates layer skipping, FFN pruning, diffusion-head cache reuse, and denoising-step early exit. Experiments on two representative UMM architectures demonstrate consistent quality--efficiency improvements across both tasks, retaining 98.03\% of dense understanding performance with a 1.93$\times$ end-to-end inference speedup.
Chinese Translation
统一多模态模型(Unified Multimodal Models)能够同时支持理解与生成任务,但会在词元、层以及生成时间步之间产生大量冗余计算。通过词元重要性探测,我们发现了一种非对称的核心-扩展结构:理解任务表现出一个稳定的重要性成分,而生成任务在很大程度上共享该成分,但需要随生成进度变化而调整的修正项。据此,我们提出了CE-Router,其采用任务共享的核心评分器与进度条件化的生成扩展,并通过生成任务分解与跨任务核心对齐进行优化。在推理阶段,CE-Router压缩词元计算,并向统一计算调度(Unified Computation Scheduling)提供学习得到的路由信号,从而协调层跳过、FFN剪枝、扩散头缓存复用以及去噪步提前退出等机制。在两种代表性统一多模态架构上的实验表明,该方法在理解与生成两项任务上均取得了一致的质量-效率提升,在保留稠密模型98.03%理解性能的同时,实现了1.93倍的端到端推理加速。
cs.AI / 86 / 2608.29301

Predicting Future Organ Dysfunction in ICU Patients Using Temporal Convolutional Networks on MIMIC-IV Data

基于MIMIC-IV数据的时序卷积网络预测ICU患者未来器官功能障碍
Albouq, Razan, Aslam, Asra
Abstract
Predicting future organ dysfunction in Intensive Care Unit (ICU) patients is critical for early clinical intervention, yet existing machine learning approaches have largely treated the Sequential Organ Failure Assessment (SOFA) score as an input to binary mortality prediction rather than as a continuous clinical outcome in its own right. We investigate the extent to which a Temporal Convolutional Net work (TCN) can predict next-day SOFA scores from multivariate ICU time-series data extracted from MIMIC-IV, characterise the relative contribution of each organ system to total SOFA variance and deterioration, and identify distinct trajectory patterns across ICU stays. A residual TCN trained on three-day sliding windows achieved a five-fold cross-validation R2 of 0.740 +- 0.013 and MAE of 1.431 +- 0.022, outperforming a naive persistence baseline on RMSE and R2. SHAP interpretability analysis revealed that the model functions primarily as a severity-anchoring mechanism rather than a true sequence model, with predictions dominated almost entirely by the most recent observation day. Cardiovascular dysfunction emerged as the strongest discriminator of both cross-sectional severity and acute deterioration, and unsupervised trajectory clustering identified two clinically meaningful phenotypes, an improving group (58.9%) and a persistently severe group (41.1%), differentiated by cardiovascular, hepatic, coagulation, and renal involvement. We conclude that TCNs can extract meaningful predictive signal from ICU physiological data, but that short input windows and complete-case selection bias currently limit their clinical utility, motivating future work on longer input horizons, alternative missing-data strategies, and external validation.
Chinese Translation
预测重症监护病房(ICU)患者未来的器官功能障碍对于早期临床干预至关重要,然而现有的机器学习方法大多将序贯器官衰竭评估(SOFA)评分作为二分类死亡率预测的输入,而非将其本身作为连续的临床结局加以预测。我们研究了时序卷积网络(TCN)在多大程度上能够基于从MIMIC-IV提取的多变量ICU时间序列数据预测次日SOFA评分,刻画了各器官系统对SOFA总方差及恶化程度的相对贡献,并识别了ICU住院期间的不同轨迹模式。基于三天滑动窗口训练的残差TCN在五折交叉验证中取得了R²为0.740±0.013、平均绝对误差(MAE)为1.431±0.022的成绩,在RMSE和R²上均优于朴素持续性基线模型。SHAP可解释性分析显示,该模型主要发挥严重程度锚定机制的作用,而非真正的序列模型,其预测几乎完全由最近一天的观测数据主导。心血管功能障碍是横断面严重程度和急性恶化的最强判别因子;无监督轨迹聚类识别出两种具有临床意义的表型,即改善组(58.9%)和持续重症组(41.1%),二者在心血管、肝脏、凝血和肾脏受累方面存在差异。我们的结论是,TCN能够从ICU生理数据中提取有意义的预测信号,但较短的输入窗口和完整病例选择偏差目前限制了其临床实用性,这为未来在更长输入时间范围、替代性缺失数据处理策略以及外部验证方面的研究提供了动力。
cs.AI / 87 / 2608.29311

Formal Concept Analysis with Three Types of Negation

具有三种否定类型的形式概念分析
Pan, Zhenghua
Abstract
Classic Formal Concept Analysis (FCA) primarily focuses on the positive relationships between objects and attributes and does not have mechanisms for handling negation.To overcome this limitation, we introduce three types of negation concepts (contradictory negation, opposite negation, intermediary negation) into FCA.Based on the set SCOI and logic LCOI+PLCOI with these three types negation, we define formal context, Galois connection operators, formal concept and concept lattice with three types of negation,this leads to the proposal of a FCACOI: Formal Concept Analysis with contradictory negation, opposite negation and intermediary negation.For the reasoning in FCACOI, this paper focuses on attribute implication reasoning. Based on the logic LCOI+PLCOI and its semantics, we introduce the notion of ICOI-entailment as the semantic implication for attribute implication reasoning in FCACOI. Through ICOI-entailment, a connection is established between attribute implication reasoning in FCACOI and inference in the logic LCOI+PLCOI, it indicate that formally proven inference rules (theorems) in LCOI+PLCOI are valid in the attribute implication reasoning of FCACOI, LCOI+PLCOI provides a logical foundation for attribute implication reasoning in FCACOI. To illustrate the capability of attribute implication reasoning in FCACOI, we discuss its application in a concrete example. Moreover, we explore attribute reduction of the formal context in FCACOI, propose two research frameworks for attribute reduction from different perspectives, and compare their characteristics.We believe that, based on richer logic and semantics, FCACOI elevates FCA from a theory that describes affirmations to one that can describe affirmations and its contradiction(either this or that), opposition(extreme negation) and intermediary (transitional states between oppositions).
Chinese Translation
经典形式概念分析(Formal Concept Analysis, FCA)主要关注对象与属性之间的正向关系,缺乏处理否定的机制。为克服这一局限,我们将三种否定概念(矛盾否定、对立否定、中介否定)引入FCA。基于包含这三种否定的集合S_COI和逻辑L_COI+P_LCOI,我们定义了带有三种否定的形式背景、Galois连接算子、形式概念和概念格,由此提出了FCACOI:具有矛盾否定、对立否定和中介否定的形式概念分析。针对FCACOI中的推理,本文聚焦于属性蕴含推理。基于逻辑L_COI+P_LCOI及其语义,我们引入I_COI-可推导性(ICOI-entailment)的概念,作为FCACOI中属性蕴含推理的语义蕴含。通过I_COI-可推导性,在FCACOI中的属性蕴含推理与逻辑L_COI+P_LCOI中的推理之间建立了联系,这表明在L_COI+P_LCOI中经形式化证明的推理规则(定理)在FCACOI的属性蕴含推理中是有效的,即L_COI+P_LCOI为FCACOI中的属性蕴含推理提供了逻辑基础。为展示FCACOI中属性蕴含推理的能力,我们通过一个具体实例讨论了其应用。此外,我们探讨了FCACOI中形式背景的属性约简,从不同视角提出了两个属性约简研究框架,并比较了它们的特点。我们相信,基于更丰富的逻辑和语义,FCACOI将FCA从一种仅能描述肯定的理论,提升为一种能够描述肯定及其矛盾(非此即彼)、对立(极端否定)和中介(对立之间的过渡状态)的理论。
cs.AI / 88 / 2608.29345

BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations

BIRD-History:一个基于历史驱动的细粒度知识标注Text-to-SQL基准测试
Zhou, Yunfan, Shi, Qiming, Yang, Yizhou, Weng, Di, Wu, Yingcai
Abstract
While recent Large Language Model (LLM)-based text-to-SQL systems achieve impressive performance on standard benchmarks, they struggle when user queries implicitly rely on domain-specific knowledge, such as business logic, data conventions, and analytical practices, that is neither captured by the schema nor explicitly stated in the natural language question. Historical SQL query logs offer a valuable source of such knowledge, yet existing benchmarks do not adequately support evaluation of history-driven approaches. To address this gap, we introduce BIRD-History, a benchmark consisting of 1,393 tasks across 11 databases, designed to evaluate text-to-SQL systems' ability to ground underspecified natural language questions using historical SQL scripts. Each task is annotated with ground-truth labels specifying which historical queries contain relevant knowledge and which SQL clauses encode it, enabling systematic evaluation of both retrieval effectiveness and knowledge utilization. Alongside the benchmark, we propose a plug-in retriever that extracts five types of external knowledge from historical SQL scripts, then retrieves and reranks relevant fragments for query generation. The retriever integrates seamlessly into existing few-shot text-to-SQL pipelines without requiring prompt modifications. Experiments demonstrate consistent improvements across four text-to-SQL systems, highlighting the value of leveraging historical query logs for handling underspecified queries. Dataset and code are open-sourced on https://github.com/zjuidg/BIRD-History.
Chinese Translation
尽管近期基于大语言模型(LLM)的Text-to-SQL系统在标准基准测试中取得了令人瞩目的性能,但当用户查询隐式依赖领域特定知识时(例如业务逻辑、数据规范和分析实践),这些知识既未在数据库模式中体现,也未在自然语言问题中明确表述,此类系统往往表现不佳。历史SQL查询日志为此类知识提供了宝贵的来源,然而现有基准测试尚不能充分支持对历史驱动方法的评估。为填补这一空白,我们提出了BIRD-History,一个涵盖11个数据库、共1,393个任务的基准测试,旨在评估Text-to-SQL系统利用历史SQL脚本为欠规范自然语言问题提供依据的能力。每个任务均标注了真实标签,指明哪些历史查询包含相关知识以及哪些SQL子句编码了该知识,从而支持对检索效果和知识利用的系统性评估。除基准测试外,我们还提出一种即插即用的检索器,可从历史SQL脚本中提取五种类型的外部知识,随后检索并重排相关片段用于查询生成。该检索器可无缝集成到现有的少样本(few-shot)Text-to-SQL流程中,无需修改提示词。实验表明,该检索器在四个Text-to-SQL系统上均带来一致的改进,凸显了利用历史查询日志处理欠规范查询的价值。数据集与代码已在 https://github.com/zjuidg/BIRD-History 开源。
cs.AI / 89 / 2608.29348

Extending TotalSegmentator: Predicting Patient and Acquisition Characteristics from CT and MR Images

扩展TotalSegmentator:从CT和MR图像预测患者与采集特征
Wasserthal, Jakob, Cyriac, Joshy, Bach, Michael, Yousefi, Kimia Mozahheb, To, Minh-Son, Sik, Máté, Hémon, Cédric, Weikert, Thomas, Segeroth, Martin
Abstract
Background: Patient details and acquisition metadata are important for clinical decisions, image quality control, and automated research pipelines, but may be missing or unreliable in imaging archives. Purpose: To develop and evaluate a fast open-source model that predicts patient and acquisition characteristics directly from CT and MR images. Materials and Methods: Separate 3D ResNet-10 ensembles for CT and MR were trained on 57,291 and 43,200 clinical examinations acquired from 2011 to 2025. Both predicted weight, height, age, sex, contrast presence, vertebral coverage, and image noise. The CT model additionally predicted scanner manufacturer, tube voltage, tube current, convolution kernel, and post-injection time; the MR model predicted sequence class. Performance was evaluated on internal CT (n=501) and MR (n=636) test sets and an external CT dataset (n=54). Results: Internal CT MAEs were 3.90 kg, 3.68 cm, and 4.42 years for weight, height, and age, with sex F1=0.990; corresponding MR results were 4.34 kg, 4.62 cm, 7.13 years, and F1=0.970. The CNN outperformed a segmentation-derived XGBoost baseline for all four core targets in both modalities (adjusted P<=.042). F1 scores were 0.963 for CT contrast, 0.953 for MR sequence, and 0.823 for MR contrast. External CT MAEs were 4.45 kg, 4.05 cm, and 5.17 years, with sex F1=0.971. CPU inference required 20 seconds for CT and 12 seconds for MR. Conclusion: One 3D multitask model per modality can rapidly recover patient and acquisition characteristics from heterogeneous CT and MR examinations. Models are available in TotalSegmentator: https://github.com/wasserth/TotalSegmentator
Chinese Translation
背景:患者详细信息和采集元数据对于临床决策、图像质量控制和自动化研究流程非常重要,但在影像存档中可能缺失或不可靠。目的:开发并评估一种可直接从CT和MR图像预测患者与采集特征的快速开源模型。材料与方法:针对CT和MR分别训练了3D ResNet-10集成模型,训练数据分别为57,291例和43,200例2011年至2025年间获取的临床检查。两个模型均预测体重、身高、年龄、性别、对比剂使用情况、椎体覆盖范围和图像噪声。CT模型额外预测扫描仪制造商、管电压、管电流、卷积核和注射后时间;MR模型预测序列类别。性能在内部CT(n=501)和MR(n=636)测试集以及一个外部CT数据集(n=54)上进行评估。结果:内部CT测试集中,体重、身高和年龄的平均绝对误差(MAE)分别为3.90 kg、3.68 cm和4.42岁,性别F1分数为0.990;对应的MR结果分别为4.34 kg、4.62 cm、7.13岁和F1=0.970。在两种模态的所有四个核心目标上,卷积神经网络(CNN)均优于基于分割特征的XGBoost基线模型(校正后P<=0.042)。CT对比剂、MR序列和MR对比剂的F1分数分别为0.963、0.953和0.823。外部CT测试集的MAE分别为4.45 kg、4.05 cm和5.17岁,性别F1=0.971。CPU推理在CT上需要20秒,MR上需要12秒。结论:每种模态仅需一个3D多任务模型即可从异构的CT和MR检查中快速恢复患者与采集特征。模型已在TotalSegmentator中发布:https://github.com/wasserth/TotalSegmentator
cs.AI / 90 / 2608.29352

Cross-Relational Preference Learning for Better LLM Instruction Following

面向更优大语言模型指令遵循的跨关系偏好学习
Li, Runsheng, Sun, Kai, Shi, Bin, Dong, Bo
Abstract
Large Language Models (LLMs) still exhibit limited capability in following complex instructions. While existing approaches often rely on preference learning to enhance this ability, they typically overlook the relationships between the permissible response spaces of different instructions, which restricts a model to align with subtle and diverse constraint variations. To address this, we propose Cross-Relational Preference Learning (CRPL), a novel framework for constructing preference data that explicitly models inter-instruction relationships through two key techniques: Cross-Relationship Perturbation and Cross-Region Pair Sampling. This enables the generation of more diverse preference data that captures a wide spectrum of constraint variations. Additionally, we introduce an atomic constraint-based verification mechanism to rigorously assess response satisfaction, ensuring high-quality preference pair construction. Extensive experiments across multiple preference learning methods (e.g., DPO, KTO), LLM backbones and four instruction-following benchmarks demonstrate that our approach achieves substantial improvements over prior baselines and exhibits strong generalization.
Chinese Translation
大语言模型(LLMs)在遵循复杂指令方面仍然能力有限。现有方法通常依赖偏好学习来增强这一能力,但往往忽略了不同指令的可允许响应空间之间的关系,这限制了模型与细微且多样的约束变化保持一致的能力。为解决这一问题,我们提出了跨关系偏好学习(Cross-Relational Preference Learning, CRPL),这是一种构建偏好数据的新框架,通过两项关键技术显式建模指令间关系:跨关系扰动(Cross-Relationship Perturbation)和跨区域对采样(Cross-Region Pair Sampling)。这使得能够生成更加多样化的偏好数据,涵盖广泛的约束变化。此外,我们引入了一种基于原子约束的验证机制,以严格评估响应的满足程度,确保构建高质量的偏好对。在多种偏好学习方法(如 DPO、KTO)、多个大语言模型骨干网络以及四个指令遵循基准上的大量实验表明,我们的方法相比先前基线取得了显著提升,并展现出强大的泛化能力。
cs.AI / 91 / 2608.29355

APPSolver: Adaptive Patch Partitioning for Point-Wise Ship Flow Prediction on Unstructured Meshes

APPSolver:基于自适应分块划分的非结构化网格船舶逐点流场预测方法
Huo, Wenhua, Han, Fenglei, Zhao, Wangyuan, Peng, Xiao, Wang, Chunhui, Wu, Jialin, Han, Jiayi
Abstract
Large non-uniform point sets make direct attention-based surrogate modeling costly for ship hydrodynamics. We introduce APPSolver, a point-wise flow-prediction framework built around Adaptive Patch Partitioning (APP), a deterministic quadtree representation for fixed two-dimensional horizontal slices extracted from ship CFD simulations. APP assigns finer patches near the hull and coarser patches farther away, downsamples patch contents, and recovers predictions to the full reference point set. Under a corrected protocol that constructs natural $(t,t+1)$ pairs before splitting, reuses training-set normalization statistics, and reports three model seeds, learned tokenizers are more accurate than APP-Transformer, and a persistence baseline has lower one-step MAE on all three ShipBench hulls. The supported benefit of APP is therefore computational rather than universal predictive superiority: on a representative DTC input, APP-Transformer requires 1.815 GFLOPs and 1.309 ms per model forward, while a matched ablation shows that adaptive partitioning reduces MAE by 16.4-24.9\% relative to a uniform partition augmented with learned slicing. Condition encoders provide setting-dependent gains in leave-one-hull-out evaluation, but the current absolute next-state objective does not establish accurate long-horizon dynamics. These results characterize APP as a compact spatial representation with an explicit accuracy--efficiency trade-off. Code is available at https://github.com/wenhuahuo/APPSolver .
Chinese Translation
大规模非均匀点集使得直接采用基于注意力机制的代理建模来研究船舶水动力学成本高昂。我们提出了APPSolver,一个围绕自适应分块划分构建的逐点流场预测框架,该框架采用确定性四叉树表示,用于处理从船舶CFD仿真中提取的固定二维水平切片。APP在船体附近分配更精细的分块,在远处分配更粗糙的分块,对分块内容进行下采样,并将预测结果恢复到完整的参考点集。在一个修正后的协议下——即在数据划分前构建自然的(t,t+1)数据对、复用训练集的归一化统计量、并报告三个模型随机种子——学习型分词器的精度高于APP-Transformer,且一个持续性基线在ShipBench全部三种船型上的一步MAE更低。因此,APP所支持的收益在于计算效率,而非普遍的预测优越性:在一个代表性DTC输入上,APP-Transformer每次模型前向传播仅需1.815 GFLOPs计算量和1.309毫秒,而匹配的消融实验表明,相对于增加了学习型切片的均匀划分,自适应划分将MAE降低了16.4%–24.9%。条件编码器在留一船型评估中提供了依赖具体设置的增益,但当前以绝对下一时刻状态为目标的训练并不能建立精确的长时间演化动态。这些结果表明,APP是一种在精度与效率之间存在明确权衡的紧凑空间表示。代码可在 https://github.com/wenhuahuo/APPSolver 获取。
cs.AI / 92 / 2608.29356

Plant-Inspired AI: Plants as Inspiration for Novel Problem Formulations, and Two Case Studies

植物启发的AI:植物作为新型问题形式化的灵感来源及两个案例研究
Sanyal, Deepayan, Michelson, Joel, Cao, Carla E., Roddy, Adam B., Kunda, Maithilee
Abstract
Artificial Intelligence (AI) has long been inspired by studies of biological intelligence. Reinforcement learning, for instance, drew inspiration from studies involving animal learning and is now a powerful paradigm for solving many real-world problems. Recently, plant biologists have uncovered a wide range of complex behaviors in plants that enable them to flexibly adapt to variable environments. Here, we argue that such behavior can motivate new AI frameworks encompassing a range of problems overlooked by existing problem-solving frameworks such as supervised learning, tree search, and constraint satisfaction. We illustrate this idea with two examples of intelligent problem-solving in plants: (1) leaf mimicry in Boquila trifoliolata, a vine capable of altering its leaves' morphology to resemble those of multiple host trees simultaneously; and (2) coordinated root-shoot growth, wherein plants allocate resources across organ systems exploring distinct environments. While leaf mimicry is highly specific to Boquila, coordination of root-shoot growth is shared across most plants. For both examples, we capture underlying computational principles and identify problems fitting these frameworks that are currently unaddressed by AI. Finally, we outline preliminary task formulations and discuss how these formulations may be applied to non-plant problems.
Chinese Translation
人工智能(AI)长期以来一直受到生物智能研究的启发。例如,强化学习的灵感来自动物学习的研究,如今已成为解决许多现实世界问题的强大范式。最近,植物学家发现植物具有广泛而复杂的行为,使其能够灵活地适应多变的环境。在此,我们提出,这类行为可以激发新的AI框架,涵盖现有问题求解框架(如监督学习、树搜索和约束满足)所忽视的一系列问题。我们通过植物智能问题求解的两个例子来说明这一思想:(1)Boquila trifoliolata(三叶藤橘)的叶片拟态——一种藤本植物能够同时改变其叶片形态以模拟多种宿主树木的叶子;(2)根-冠协同生长——植物在探索不同环境的器官系统之间分配资源。叶片拟态是Boquila特有的,而根-冠协同生长则为大多数植物所共有。针对这两个例子,我们提炼了其背后的计算原理,并识别出符合这些框架但目前尚未被AI解决的问题。最后,我们概述了初步的任务形式化,并讨论了这些形式化如何应用于非植物问题。
cs.AI / 93 / 2608.29357

LiteSearch-VL: Small Multimodal Search Agents via Trajectory Distillation and Synthetic Step-DPO

LiteSearch-VL:基于轨迹蒸馏与合成步级DPO的小型多模态搜索智能体
Khaki, Saeed, Safaei, Nima, Ginotra, Kamal
Abstract
Multimodal search agents answer visual questions by interleaving image understanding, web retrieval, tool use, and evidence synthesis. Strong systems exist, but in two expensive regimes: proprietary frontier models such as GPT-5 and Gemini, or large open vision-language backbones trained with substantial agentic data and reinforcement learning. We ask a different question: when released agent trajectories are distilled into much smaller backbones under a single-node budget, what is actually transferred? We study this with LiteSearch-VL, a low-compute recipe for Qwen3-VL-2B and Qwen3-VL-4B that uses only released OpenSearch-VL trajectories, parameter-efficient LoRA adapters, and synthetic step-level preferences: DPO on GPT-5-generated hard negatives targeting five local failure modes (premature answer, wrong tool, weak query, repeated query, ignored image). Across 12,400 GPT-5-judged rollouts on SimpleVQA, FVQA, LiveVQA, and VDR-Bench-testmini, the dominant effect is behavioral rather than a uniform accuracy lift: full-trajectory supervised fine-tuning transfers the agent contract, taking the 2B model from almost never emitting a usable answer (1,237/1,240 no_answer rollouts) to 28.4% macro Pass@1, matching or slightly exceeding the off-the-shelf 4B base (25.6%). Synthetic preference learning and compact tool distillation act as refinements rather than phase transitions (best 4B configuration: 30.8% macro Pass@1). Finally, a controlled VDR step-budget ablation shows that extra search turns convert abstentions into wrong_entity errors rather than correct answers, identifying answer verification, not search depth, as the next bottleneck for small multimodal agents.
Chinese Translation
多模态搜索智能体通过交替进行图像理解、网络检索、工具使用和证据综合来回答视觉问题。目前强大的系统存在于两种高成本模式中:一是GPT-5和Gemini等专有前沿模型,二是使用大量智能体数据和强化学习训练的大型开源视觉语言骨干模型。我们提出了一个不同的问题:当公开发布的智能体轨迹在单节点预算下被蒸馏到小得多的骨干模型中时,实际迁移到的是什么?我们通过LiteSearch-VL研究这一问题,这是一种面向Qwen3-VL-2B和Qwen3-VL-4B的低计算量方案,仅使用公开的OpenSearch-VL轨迹、参数高效的LoRA适配器以及合成的步级偏好:针对五种局部失败模式(过早作答、工具错误、查询能力弱、重复查询、忽视图像)对GPT-5生成的困难负样本进行DPO训练。在SimpleVQA、FVQA、LiveVQA和VDR-Bench-testmini上共12,400次由GPT-5评判的推理测试中,主导效应是行为层面的而非整体准确率的普遍提升:全轨迹监督微调迁移了智能体的行为契约,使2B模型从几乎从未给出可用答案(1,240次推理中1,237次为no_answer)提升到28.4%的宏平均Pass@1,达到或略超过现成的4B基础模型(25.6%)。合成偏好学习和紧凑的工具蒸馏起到的是优化作用而非阶段性跃迁(最佳的4B配置:30.8%宏平均Pass@1)。最后,一项受控的VDR步数预算消融实验表明,额外的搜索轮次会将弃答转化为wrong_entity错误而非正确答案,这表明对于小型多模态智能体而言,下一个瓶颈在于答案验证而非搜索深度。
cs.AI / 94 / 2608.29363

TRACER: Per-Tool Context Retention for LLM Agents via Consequence-Attributed Reinforcement Learning

TRACER:基于后果归因强化学习的大语言模型智能体按工具上下文保留方法
Lin, Ziqi, Wu, Ye, Yang, Mengying, Liu, Xu, Liu, Yizhou, Ke, Qiang, Guo, Qin
Abstract
Enterprise data agents answer business queries by chaining many tool calls over multiple reasoning steps, routinely accumulating hundreds of thousands of context tokens per session. Existing compression strategies typically allocate retention budgets without accounting for the downstream consequences of removing individual tool outputs. Aggressive compression may therefore trigger costly tool re-invocations that offset the initial savings. We call this the compression--consequence gap. To close it, we propose TRACER, which formulates compression as a sequential per-tool decision problem. A lightweight REINFORCE policy assigns query-conditioned retention ratios using only information available at each compression event. Its consequence-aware objective jointly accounts for task success, total token consumption, and post-compression tool re-invocations. To improve credit assignment, TRACER uses a learned outcome model to compare the predicted consequences of the selected retention ratio with those of fully retaining each tool output. On held-out production queries across three compressor backends, TRACER reduces total token consumption by 29--46% relative to keeping all context while maintaining comparable or higher task success. Compared with a tool-type-conditional static policy, TRACER provides an additional 15--18% of token savings. Interventional rollouts show that the learned per-tool credit scores correlate with measured single-tool consequences. The learned policy also yields positive savings when transferred across agent backbones and compressor architectures, and reduces token consumption by 18--25% on five held-out LOCA-bench environments. These results demonstrate the value of consequence-aware, per-tool context retention for improving the efficiency of long-horizon language agents.
Chinese Translation
企业数据智能体通过在多个推理步骤中串联大量工具调用来回答业务查询,每个会话通常会累积数十万个上下文词元。现有的压缩策略在分配保留预算时,通常未考虑删除单个工具输出所带来的下游后果。因此,激进的压缩可能触发代价高昂的工具重新调用,从而抵消最初节省的开销。我们将这一现象称为压缩—后果鸿沟(compression–consequence gap)。为弥合这一鸿沟,我们提出了TRACER,将压缩建模为一个按工具进行的序列决策问题。一个轻量级的REINFORCE策略仅利用每次压缩事件时可获得的信息,为各工具分配基于查询条件的保留比例。其后果感知目标函数同时考虑任务成功与否、总词元消耗量以及压缩后的工具重新调用次数。为改进信用分配,TRACER使用一个学习到的结果模型,将所选保留比例的预测后果与完整保留每个工具输出的后果进行比较。在三个压缩器后端的留出生产查询上,TRACER相对于保留全部上下文的方式将总词元消耗降低了29–46%,同时保持相当或更高的任务成功率。与基于工具类型的条件静态策略相比,TRACER额外节省了15–18%的词元。干预性实验表明,学习到的按工具信用分数与实测的单工具后果相关。学习到的策略在跨智能体骨干和压缩器架构迁移时仍能带来正向节省,并在五个留出的LOCA-bench环境中将词元消耗降低18–25%。这些结果证明了后果感知的、按工具的上下文保留对提升长时程语言智能体效率的价值。
cs.AI / 95 / 2608.29368

Reviving our data foundations is the most disruptive step to data maturity

振兴数据基础是提升数据成熟度最具颠覆性的一步
Carapella, Valentina, Jimenez-Ruiz, Ernesto
Abstract
The most disruptive step that enterprises of small-medium size and maturity can take to make the most of the latest technological advances in AI is to step back from the hype and focus on establishing or reviving a good knowledge foundation layer. It is a hard message to present to the executive team; therefore, it needs to be backed by evidence, and its implementation needs to be of minimal impact on the existing processes. In this vision statement, we discuss how we need to rethink what evidence speaks to the decision-makers and propose a low-impact data strategy that adapts to the existing and ever-changing data flows and processes across the company. We firmly believe that knowledge graph techniques will increasingly become non-negotiable in the data strategy of an AI-powered enterprise, provided that we approach their design in a modular, dynamic and cross-functional way.
Chinese Translation
对于中小型且成熟度有限的企业而言,要充分利用人工智能领域最新技术进展,最具颠覆性的一步是跳出炒作的喧嚣,专注于建立或振兴良好的知识基础层。这一观点难以向管理层传达,因此需要证据支持,并且其实施方案需尽量减少对现有流程的影响。在这篇愿景陈述中,我们探讨了如何重新思考何种证据能够说服决策者,并提出一种低影响的数据策略,该策略能够适应公司内部现有的和不断变化的数据流与流程。我们坚信,只要以模块化、动态化和跨职能的方式进行知识图谱的设计,知识图谱技术必将在人工智能驱动的企业数据战略中变得不可或缺。
cs.AI / 96 / 2608.29372

FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents

FORESIGHT-9:面向自适应交易代理的前瞻性与过程感知评估
Luo, Xiangxin, Hong, Chengtian, Li, Haohua, Xie, Yongyi
Abstract
Retrospective backtests provide a limited test of adaptive trading agents: they cannot rule out historical contamination, expose sensitivity to a single realized market path, or reveal internal degeneration during long-horizon adaptation. We introduce FORESIGHT-9, a prospective and process-aware benchmark built from nine auditable counterfactual stress worldlines branching from a common July 2026 information boundary. Each worldline specifies staged macro-financial events and joint multi-asset anchors; a deterministic generator realizes the trajectories, while observations are disclosed according to in-world time. A common contract standardizes observations and execution while preserving each agent's native adaptation loop. We evaluate two adaptive trading-agent frameworks with two foundation-model backbones across 36 long-horizon runs. Agent rankings vary substantially across worldlines and backbones, and a fixed equal-weight policy outperforms 31 of 36 runs. Process telemetry exposes failures that terminal returns conceal: in one high-return run, the live factor library collapsed while executed holdings converged to the equal-weight fallback, even though decision records continued to report an active factor ensemble. FORESIGHT-9 therefore evaluates not only portfolio outcomes, but whether adaptive agent state and execution remain coherent across alternative futures. We release the worldlines, trajectories, audit traces, and regeneration scripts.
Chinese Translation
回顾式回测对自适应交易代理的检验能力有限:它们无法排除历史数据污染,无法揭示对单一已实现市场路径的敏感性,也无法暴露长期适应过程中的内部退化。我们提出FORESIGHT-9,一个前瞻性且过程感知的基准,它由九条可审计的反事实压力世界线(worldline)构成,这些世界线从一个共同的2026年7月信息边界分叉而来。每条世界线规定了分阶段的宏观金融事件和跨资产联合锚点;由确定性生成器实现轨迹,而观测信息则按照世界内部时间逐步披露。统一的合约标准对观测与执行进行了规范化,同时保留每个代理原有的自适应循环。我们在36次长周期运行中评估了两个自适应交易代理框架与两个基础模型骨干。代理排名在不同世界线和不同骨干之间变化显著,且固定等权策略在36次运行中优于其中的31次。过程遥测揭示了终端收益所掩盖的失败:在某次高收益运行中,实盘因子库已崩溃、实际持仓收敛为等权兜底策略,而决策记录仍继续报告一个活跃的因子组合。因此,FORESIGHT-9不仅评估投资组合结果,还评估自适应代理的状态与执行在各种替代未来情景中能否保持一致性。我们公开了世界线、轨迹、审计记录以及复现脚本。
cs.AI / 97 / 2608.29376

Evaluating Tiny Recursive Models Across Training for Code Generation

面向代码生成训练过程评估的微型递归模型
Sirivella, Anjani, Newaz, Aanisha, Melo, Glaucia
Abstract
Code generation increasingly relies on large transformer models, whose capability advances with scale. Yet such a scale is costly, creating demand for small models, especially where data is limited. Recursive models address this by reusing a single block to add depth rather than stacking independent layers. Such models are typically evaluated by teacher-forced fit (next-token loss on ground-truth prefixes) or task accuracy, at a single checkpoint, whereas code is produced by free-running generation, where the model extends its own output. Whether a teacher-forced advantage survives free-running generation, and whether it holds across training, remains open. To study both, we compare a ~28M-parameter autoregressive Tiny Recursive Model (TRM-AR) on natural-language-to-Python code generation against parameter-matched and depth-matched controls, tracking fit and generation across 40 epochs and three seeds. The fit ranking between the recursive model and the depth-matched control reverses twice. Selecting each checkpoint by validation loss and examining the trajectory yields a consistent comparison. At equal parameters, TRM-AR fits, generates, and generalizes better than the parameter-matched control while recovering approximately 45% of the validation-loss gap and 57% of the generation-quality gap between the two controls, at roughly 175 times the per-step cost of the parameter-matched control. However, at equal effective depth, the larger transformer fits and generates better at its validation optimum, suggesting TRM-AR's advantage lies in resistance to overfitting, not greater capability. These findings suggest that recursive code generation models should be evaluated jointly on fit and generation across the training trajectory rather than at a single checkpoint.
Chinese Translation
代码生成日益依赖大型Transformer模型,其能力随规模而提升。然而,这种规模成本高昂,因而对小型模型产生了需求,尤其是在数据有限的场景下。递归模型通过复用单个模块以增加深度,而非堆叠独立的层,来应对这一挑战。此类模型通常以教师强制拟合(在真实前缀上的下一词元损失)或任务准确率在单一检查点上进行评估,而代码是由自由运行式生成产生的,即模型在其自身输出上继续生成。教师强制下的优势是否能在自由运行生成中保持,以及该优势是否在整个训练过程中保持,仍是未解问题。为研究这两个问题,我们将约2800万参数的自回归微型递归模型(TRM-AR)应用于自然语言到Python的代码生成任务,并与参数匹配和深度匹配的对照模型进行比较,在40个训练轮次和三个随机种子下追踪其拟合与生成表现。递归模型与深度匹配对照模型之间的拟合排名在训练过程中出现了两次反转。通过以验证损失选取各检查点并审视训练轨迹,可以得到一致的对比结果。在参数相同的情况下,TRM-AR在拟合、生成和泛化方面均优于参数匹配的对照模型,弥补了两个对照模型之间约45%的验证损失差距和57%的生成质量差距,但其每步计算成本约为参数匹配对照模型的175倍。然而,在有效深度相同的情况下,更大的Transformer模型在其验证最优点处的拟合与生成表现更佳,这表明TRM-AR的优势在于抗过拟合能力,而非更强的能力。这些发现表明,递归代码生成模型应在整个训练轨迹上联合评估其拟合与生成表现,而非仅在单一检查点上进行评估。
cs.AI / 98 / 2608.29387

EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants

EvoGenUI-Bench:评估大语言模型作为多轮生成式UI助手的能力
Peng, Yue, Xia, Lanke, Wang, Zihan, Ye, Jiahao, Ning, Ke, Wen, Hongyi
Abstract
Large language models can generate interactive web interfaces, but reliable generative UI requires maintaining an executable artifact as user requests evolve. We introduce EvoGenUI-Bench, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state. We execute generated artifacts in a browser and evaluate them using screenshots, source and DOM evidence, actor traces, and runtime logs. Beyond turn-level and episode-level success, we measure cross-turn retention with Adjacent Pass Retention. Across eight models, even the strongest achieves 74.9% Turn Pass while completing only 37.3% of five-turn episodes; APR further falls to 52.4% on tool-grounded tasks. Diagnostic analysis shows that presentation failures center on information architecture, interaction failures on derived-state propagation and affordance binding, and tool-grounded failures additionally involve external-state grounding and requirement decomposition. These results reframe generative UI evaluation from judging isolated outputs to testing whether interface behavior, derived state, external state, and assistant claims remain synchronized as the artifact evolves.
Chinese Translation
大语言模型能够生成交互式网页界面,但可靠的生成式UI要求在用户需求不断演变的过程中保持一个可执行的制品(artifact)。我们提出了EvoGenUI-Bench,这是一个用于多轮界面维护的基准测试,包含150个五轮任务、共750轮对话,涵盖三种场景:信息展示、可执行交互以及基于工具的外部状态。我们在浏览器中执行生成的制品,并利用截图、源代码和DOM证据、执行者行为轨迹(actor traces)以及运行时日志对其进行评估。除了轮次级和片段级的成功率外,我们还通过相邻轮次通过保持率(Adjacent Pass Retention, APR)来衡量跨轮次的能力保持。在八个模型中,即使最强的模型也仅达到74.9%的轮次通过率(Turn Pass),且只能完成37.3%的五轮任务片段;在基于工具的任务上,APR进一步下降至52.4%。诊断分析表明,展示类失败主要源于信息架构问题,交互类失败主要源于派生状态传播和可供性绑定(affordance binding)问题,而基于工具的失败还涉及外部状态接地和需求分解问题。这些结果将生成式UI评估的视角从评判孤立的输出转变为测试界面行为、派生状态、外部状态以及助手声明在制品演变过程中能否保持同步。
cs.AI / 99 / 2608.29459

Toward Latent Language Model Skills Steering and Optimization: An Empirical Study

面向大语言模型潜在技能的引导与优化:一项实证研究
Jiang, Xunyi, Wu, Junda, Xiong, Yuxin, Yu, Sheldon, Yu, Tong, Arbour, David, Sinha, Ritwik, McAuley, Julian, Wen, Hongyi
Abstract
Skills, as a useful abstraction for the procedural capabilities of large language models (LLMs), capture how models perform structured, multi-step reasoning and program execution. Existing approaches typically treat skills as explicit, surface-level constructs specified through prompts or programs, leaving open the question of how such procedural capabilities are represented inside the model and whether they can be manipulated as structured objects in latent space. In this empirical study, we investigate whether procedural LLM skills can be represented as directions in activation space and whether vector-space operations over these directions can express skill-level behaviors. We find that procedural skills admit a vector-space representation: individual skill directions can be activated to shift model behavior; independently extracted directions can compose to form higher-level skills. Contrastive directions yield context-conditioned algorithmic personalization and optimization trajectories over skill directions evolve non-monotonically, with intermediate states often surpassing fully optimized solutions. These results support a representation-level view of procedural LLM skills: they admit a latent vector-space organization that allows direct manipulation through internal interventions.
Chinese Translation
技能(Skills)作为大语言模型(LLMs)程序化能力的有效抽象,刻画了模型如何执行结构化的多步推理与程序执行。现有方法通常将技能视为通过提示或程序指定的显式的、表层构造,而这类程序化能力在模型内部如何表征、以及能否作为结构化对象在潜在空间中被操控,仍是悬而未决的问题。在本实证研究中,我们探究了大语言模型的程序化技能是否可以表示为激活空间中的方向,以及针对这些方向的向量空间运算是否能够表达技能层面的行为。我们发现,程序化技能可以具有向量空间表示:激活单个技能方向能够改变模型行为;独立提取的方向可以组合形成更高层次的技能。对比方向(contrastive directions)可实现情境条件化的算法个性化,而基于技能方向的优化轨迹呈现非单调演化,其中间状态往往优于完全优化的解。这些结果支持从表征层面理解大语言模型的程序化技能:它们具有潜在的向量空间组织结构,可通过内部干预进行直接操控。
cs.AI / 100 / 2608.29460

Can escalation channels redirect reward hacking toward defect disclosure?

升级上报渠道能否将奖励劫持行为引向缺陷披露?
Gomez, Francesca
Abstract
When coding agents encounter defective test infrastructure they may reward-hack: hardcoding outputs or editing test files to pass tests they cannot legitimately satisfy, a pattern that has now appeared outside benchmarks, in a coordinated multi-agent intrusion of a major AI platform's production infrastructure. The same capability that lets an agent detect and exploit a defect could let it report one, given the right decision environment. We evaluate escalation channels, structured reporting tools available to the agent at the point of conflict, as a decision-environment intervention that both reduces reward hacking and surfaces the infrastructure defects that trigger it. A $2 \times 2$ factorial separates the contributions of an escalation tool, a standalone anti-reward-hacking policy, and their combination. Across 8 frontier models spanning 5 families, the combined intervention reduces reward hacking from 23.6% to 5.3% (mixed-effects logistic OR = 9.2, 95% CI 5.0--16.8, $p < 10^{-12}$) with no detectable cost or performance overhead, eliminating it entirely for 6 of 8 models. Escalation and hacking are near-perfectly mutually exclusive, with 98.7% of escalations involving no hacking (100% under the combined intervention). Beyond reduction, escalation channels function as diagnostic infrastructure: on top of monitoring, escalation adds +10.1 percentage points of defect detection coverage and is more accurate once it fires (99.4% vs 85.8%). Unlike containment-based approaches that risk outpacing growing model capabilities, escalation channels redirect capability toward disclosure rather than exploitation.
Chinese Translation
当编程智能体(coding agents)遇到有缺陷的测试基础设施时,可能会进行奖励劫持(reward hacking):硬编码输出或编辑测试文件,以通过其无法正当满足的测试。这种模式现已出现在基准测试之外——在针对某大型AI平台生产基础设施的一次协同多智能体入侵中。智能体用于检测并利用缺陷的能力,在合适的决策环境下,同样可用于上报缺陷。我们评估了升级上报渠道(escalation channels),即在冲突发生时可供智能体使用的结构化报告工具,作为一种决策环境干预手段,它既能减少奖励劫持行为,又能暴露引发该行为的底层基础设施缺陷。通过 $2 \times 2$ 析因实验设计,我们分离了升级上报工具、独立的反奖励劫持策略以及二者组合各自的贡献。在涵盖5个模型系列的8个前沿模型上,组合干预将奖励劫持率从23.6%降至5.3%(混合效应逻辑回归 OR = 9.2,95% CI 5.0--16.8,$p < 10^{-12}$),且无可检测的成本或性能开销,并在8个模型中的6个上完全消除了该行为。升级上报与奖励劫持几乎完全互斥,98.7%的上报案例不涉及任何劫持行为(在组合干预下为100%)。除了降低劫持率之外,升级上报渠道还能作为诊断性基础设施:在监控的基础上,升级上报额外带来了+10.1个百分点的缺陷检测覆盖率,且一旦触发,其准确率更高(99.4% vs 85.8%)。与可能被不断增强的模型能力所超越的遏制类方法不同,升级上报渠道将模型能力引向披露而非利用。
cs.AI / 101 / 2608.29588

Call Neighbours Yourself: Graph Walks with Destination-Conditioned On-Policy Self-Distillation

自行呼叫邻居:基于目标条件化在策略自蒸馏的图游走方法
Liu, Yilun, Luo, Boyu, Tang, Yanran, Qiu, Ruihong, Huang, Zi
Abstract
Reasoning over text-attributed graphs (TAGs) requires large language models (LLMs) to combine a node's text with evidence distributed across its neighbourhood. Existing methods fix the set of accessible neighbours before generation, forcing reasoning to operate over a static context and preventing the model from acquiring missing evidence during inference. We argue that neighbour selection should itself be part of the reasoning process. To this end, we propose Call Neighbours Yourself (CNY), a framework that enables LLMs to proactively explore graph neighbourhoods through topology-constrained graph-walk actions. Instead of reasoning over a pre-selected neighbour set, CNY exposes lightweight neighbour previews and learns when to expand candidate neighbours for additional evidence. To address the delayed-credit challenge of neighbour exploration, we introduce destination-conditioned on-policy self-distillation, which retrospectively evaluates a selected neighbour after its content is revealed and converts the resulting change in action preference into an action-level training signal. Experiments on standard TAG reasoning benchmarks under a unified raw-text setting show that CNY consistently outperforms fixed-context post-training baselines. Furthermore, the learned exploration policy transfers to unseen graphs and to a graph-level task not encountered during training. Code is available at https://github.com/superallen13/CNY.
Chinese Translation
对文本属性图(text-attributed graphs, TAGs)进行推理,要求大语言模型(LLMs)将节点的文本与其邻域中分布的证据相结合。现有方法在生成之前就固定了可访问邻居的集合,迫使推理在静态上下文中进行,并阻止模型在推理过程中获取缺失的证据。我们认为,邻居选择本身应当是推理过程的一部分。为此,我们提出了 Call Neighbours Yourself(CNY)框架,使大语言模型能够通过拓扑约束的图游走动作主动探索图邻域。CNY 不再在预先选定的邻居集合上进行推理,而是展示轻量级的邻居预览,并学习何时扩展候选邻居以获取更多证据。为解决邻居探索中延迟信用分配的难题,我们引入了目标条件化在策略自蒸馏(destination-conditioned on-policy self-distillation)方法,该方法在邻居内容被揭示后回溯性地评估所选邻居,并将由此产生的动作偏好变化转化为动作级训练信号。在统一的原始文本设置下,于标准 TAG 推理基准上的实验表明,CNY 始终优于固定上下文的后训练基线方法。此外,所学到的探索策略能够迁移到未见过的图以及训练中未遇到的图级任务上。代码可在 https://github.com/superallen13/CNY 获取。
cs.AI / 102 / 2608.29589

Not Safe for All: Auditing the Dialect Penalty in Text-to-Image Safety Pipelines

并非对所有人都安全:审计文本到图像安全管线中的方言惩罚
Kim, Minkyu, Choi, Juhwan, Kim, YoungBin
Abstract
Text-to-image (T2I) safety guardrails fail to generalize equitably to non-standard dialects. Evaluating 23,080 paired prompts across five English dialects, we formalize this failure as the dialect penalty, where filters trigger based on linguistic surface features rather than semantic intent. Text-level filters fail in opposing directions: NSFW-T over-flags benign dialect prompts and LatentGuard over-flags toxic ones (bias gaps up to +28.29 pp), while the OpenAI Moderation API under-detects them. A controlled typo ablation confirms this penalty originates from flagging dialectal features, not generic out-of-distribution sensitivity. The pixel-level generator is largely dialect-agnostic; the penalty enters at text processing and cascades unevenly to post-hoc guardrails. We show this bias tracks training data imbalance and is mitigable via group-balanced retraining, with an ablation attributing the gain to balanced exposure rather than to the worst-group objective of GroupDRO (group distributionally robust optimization). Current pipelines systematically fail dialect speakers, an equity failure masked by mean accuracy benchmarks. Our official code and dataset are publicly available at https://github.com/minguinho26/dialect-penalty-t2i. Content Warning: This paper contains offensive, toxic, or disturbing text prompts and generated images.
Chinese Translation
文本到图像(Text-to-Image, T2I)安全防护机制无法公平地泛化到非标准方言。通过评估跨越五种英语方言的23,080个配对提示词,我们将这种失败形式化为“方言惩罚”(dialect penalty),即过滤器基于语言的表层特征而非语义意图进行触发。文本级过滤器在相反方向上失效:NSFW-T 会过度标记良性的方言提示词,而 LatentGuard 则过度标记有毒的方言提示词(偏差差距最高达+28.29个百分点),而 OpenAI Moderation API 则检测不足。受控的拼写错误消融实验证实,这种惩罚源于对方言特征的标记,而非一般的分布外敏感性。像素级生成器在很大程度上与方言无关;该惩罚在文本处理阶段引入,并以不均衡的方式级联到事后防护机制。我们表明,这种偏差与训练数据的不平衡相关,并可通过群体平衡的再训练来缓解,消融实验将这一收益归因于平衡的曝光而非 GroupDRO(群体分布鲁棒优化)的最差群体目标。当前的管线系统性地使方言使用者蒙受失败,而这一公平性缺陷被平均准确率基准所掩盖。我们的官方代码和数据集已在 https://github.com/minguinho26/dialect-penalty-t2i 公开。内容警告:本文包含冒犯性、有毒或令人不适的文本提示词和生成的图像。
cs.AI / 103 / 2608.29596

Towards a Systems Foundation for Agentic Skills: Architecture, Lifecycle, and Security

迈向智能体技能的系统化基础:架构、生命周期与安全
Badhe, Sanket, Shah, Deep, Tiwari, Priyanka, Kathrotia, Nehal
Abstract
Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and execution stability bottlenecks when deployed on complex, long-horizon tasks. While monolithic prompt engineering and stateless tool-calling paradigms struggle to scale, the field is rapidly converging toward \emph{agentic skills}: modular procedural abstractions that externalize execution knowledge into reusable, executable, and portable artifacts. This paper establishes a unified systems foundation and reference architecture for the agentic skills ecosystem. We formalize skills as externalized procedural knowledge bridging high-level cognitive planning with deterministic execution environments, and systematically delineate the architecture across a nine-stage lifecycle: autonomous discovery, authoring and representation formats, memory storage, dynamic retrieval and routing, composition and orchestration, execution and repair, lifelong adaptation, empirical evaluation, and security governance. We further examine marketplace dynamics, public registries, and emerging adversarial threat vectors, alongside runtime verification and defense mechanisms. Finally, we categorize system implementations across software engineering, operating system navigation, embodied robotics, and scientific discovery, while highlighting critical open challenges in continual learning and benchmark realism. This work establishes agentic skills as a foundational paradigm for building scalable, robust, and verifiable autonomous language agents.
Chinese Translation
自主大语言模型(LLM)智能体在部署于复杂、长时程任务时,日益面临可靠性、上下文消耗与执行稳定性方面的瓶颈。尽管单体式提示工程和无状态工具调用范式难以扩展,该领域正迅速收敛于“智能体技能”(agentic skills):即将执行知识外化为可复用、可执行、可移植的制品的模块化程序性抽象。本文为智能体技能生态系统建立了一个统一的系统化基础和参考架构。我们将技能形式化为外化的程序性知识,衔接高层认知规划与确定性执行环境,并系统地划分了涵盖九个阶段生命周期的架构:自主发现、编写与表示格式、记忆存储、动态检索与路由、组合与编排、执行与修复、终身适应、实证评估以及安全治理。我们进一步考察了市场动态、公共注册表以及新兴的对抗性威胁向量,并探讨了运行时验证与防御机制。最后,我们对软件工程、操作系统导航、具身机器人以及科学发现等领域的系统实现进行了分类梳理,同时指出了持续学习与基准真实性方面的关键开放性挑战。这项工作将智能体技能确立为构建可扩展、鲁棒且可验证的自主语言智能体的基础性范式。
cs.AI / 104 / 2608.29612

LLMs Interpret, Embeddings Organize, Graphs Emerge: Agent-Driven Compilation of Scientific Knowledge

大语言模型负责解释,嵌入负责组织,图谱自然涌现:代理驱动的科学知识编译
Ran, Shi-Ju, Zhang, Kun, Wu, Xi, Yang, Liu-Si, Li, Wen-Jun
Abstract
Sustained scientific work requires a knowledge substrate that carries interpretation across tasks and preserves paths to source evidence. We call this process \emph{scientific knowledge compilation} and implement it in ASKS, the \emph{Agent-Driven Scientific Knowledge System}. For each source, an LLM produces a readable Wiki view and machine-facing semantics. Deterministic checks convert the latter into a document-local GraphDelta, and embedding geometry together with explicit graph rules integrates the proposed changes into persistent state. Each ingest is an inspectable state transition over accumulated knowledge, with compiled Wiki and graph views linked to the preserved source record. We examine this process by chronologically compiling 56 published papers from one research program. Branch survival, cross-paper support, lineage, coverage, and churn yield a source-traceable author research portrait centered on tensor-network methods, with branches into quantum many-body research, tensor-network machine learning, and quantum-AI-oriented directions. In this run, higher-level Hub organization remains stable and low-churn. Canonical-node growth is predominantly additive. Graph-level measurements and navigation paths retain links to the source records from which they were compiled.
Chinese Translation
持续性的科学工作需要一个知识基底,该基底能够在不同任务之间承载解释,并保留通向原始证据的路径。我们将这一过程称为“科学知识编译”(scientific knowledge compilation),并在ASKS(代理驱动科学知识系统,Agent-Driven Scientific Knowledge System)中加以实现。对于每个来源,大语言模型(LLM)生成一个可读的Wiki视图以及面向机器的语义表示。确定性的检查将后者转换为文档局部的GraphDelta,然后通过嵌入几何结构与显式图规则将提议的变更整合进持久化状态。每次知识摄取都是对累积知识的一次可审查的状态转移,编译后的Wiki视图和图视图均与所保留的原始记录相链接。我们通过按时间顺序编译来自同一研究项目的56篇已发表论文来检验这一过程。分支存活率、跨论文支持、谱系、覆盖率和变动率共同生成了一份可溯源至原始文献的作者研究画像:其核心为张量网络方法,并分支出量子多体研究、张量网络机器学习以及面向量子-AI的方向。在本次运行中,更高层级的Hub组织保持稳定且变动率较低。规范节点(canonical-node)的增长以累加式为主。图层面的测量结果与导航路径均保留了指向其编译来源记录的链接。
cs.AI / 105 / 2608.29646

Detect Before You Attribute: Cascade Failure Attribution for Multi-Agent Systems

先检测再归因:面向多智能体系统的级联失效归因
Zhang, Jiayi, Wang, Zexin, Sun, Degang, Pei, Changhua, Sun, Fei, Xie, Gaogang, Li, Jingjing
Abstract
Large language model (LLM)-based agents have shown strong potential in solving complex tasks through multi-step reasoning, yet they remain vulnerable to execution failures. Accurate failure attribution is therefore critical for improving agent reliability. Existing topology- and spectrum-based methods exploit trajectory structures but often overlook fine-grained semantics, while LLM-based attribution methods capture semantic cues but suffer from long-context degradation over lengthy trajectories. To address these challenges, we propose DUOTRACE, a plug-and-play detection filter for LLM-based failure attribution. DUOTRACE follows a detect-before-attribute paradigm: it first detects anomalous executions and then supplies focused trajectory evidence to downstream LLM-based attribution methods. For effective VAE-based anomaly detection on agent trajectories, DUOTRACE integrates dual-view semantic-structural node representations, a Tree-LSTM-based trajectory encoder, and prefix-chain- and LLM-based data augmentation to handle heterogeneous nodes, hierarchical execution structures, and limited failure data. Experiments with six LLM-based attribution baselines show that DUOTRACE improves agent-level and step-level attribution accuracy by 8.7% and 7.0%, respectively.
Chinese Translation
基于大语言模型(LLM)的智能体在通过多步推理解决复杂任务方面展现出强大潜力,但仍然容易受到执行失败的影响。因此,准确的失效归因对于提升智能体可靠性至关重要。现有的基于拓扑和频谱的方法利用了轨迹结构,但往往忽视了细粒度的语义信息;而基于LLM的归因方法能够捕捉语义线索,却在长轨迹上受限于长上下文退化问题。为应对这些挑战,我们提出了DUOTRACE,一种即插即用的检测过滤器,用于基于LLM的失效归因。DUOTRACE遵循“先检测再归因”的范式:它首先检测异常执行,然后为下游基于LLM的归因方法提供聚焦的轨迹证据。为了在智能体轨迹上实现有效的基于VAE的异常检测,DUOTRACE集成了双视角语义-结构节点表示、基于Tree-LSTM的轨迹编码器,以及基于前缀链和LLM的数据增强,以处理异构节点、层次化执行结构和有限的失效数据。在六个基于LLM的归因基线上的实验表明,DUOTRACE将智能体级和步骤级的归因准确率分别提升了8.7%和7.0%。
cs.AI / 106 / 2608.29696

Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment

Ideation Arena:基于对战式人类专家评估的大语言模型研究想法评价方法
Chen, Zhiyu, Zhao, Keyu, Fu, Jigao, Liang, Dong, Wu, Yanbiao, Li, Jiaoyang, Xue, Haidong, Zeng, Xinhua, Zhen, Yuanyi, Xu, Fengli, Li, Yong
Abstract
Evaluating research ideas generated by LLMs is difficult because their scientific value cannot be fully determined by objective criteria, and no single reference answer specifies what counts as a good idea. To address this challenge, we introduce Ideation Arena, a battle style platform that evaluates research ideas through pairwise human assessment. Ideation Arena evaluates ideas generated by 14 frontier LLMs and 5 research agent architectures built on 2 base models. To ensure a common starting point, Ideation Arena builds shared literature contexts from papers familiar to the participating researchers and provides the same contexts to all LLMs and agents. We collect over 6,000 double blind pairwise comparisons from 105 active computer science researchers and construct an Elo rating leaderboard of proposal-stage expert preferences in computer science under a shared closed-context protocol. We validate the rankings through interrater agreement and robustness analyses, showing that the leaderboard remains stable under changes in annotator composition and domain coverage. Our results show substantial variation in agent effectiveness, with some frameworks improving ideation quality over their backbones and others offering little benefit or even underperforming their base models. We further construct Ideation Arena Eval, a benchmark for assessing whether automated evaluators align with human preferences in research ideation. Experiments with current LLM judges show that they still cannot reliably reproduce expert preferences, with the best judge reaching 72.56% Soft Accuracy on Overall Quality. Our code, data, and leaderboards are available at https://github.com/foss12138/Research-Ideation-Arena.
Chinese Translation
评估大语言模型(LLM)生成的研究想法十分困难,因为其科学价值无法完全由客观标准确定,也没有单一的参考答案能够界定什么才算一个好想法。为应对这一挑战,我们提出了 Ideation Arena,一个通过成对人类评估来评价研究想法的对战式平台。Ideation Arena 评估了 14 个前沿大语言模型以及基于 2 个基础模型构建的 5 种研究智能体架构所生成的想法。为确保共同的起点,Ideation Arena 从参与研究者熟悉的论文中构建共享文献上下文,并向所有大语言模型和智能体提供相同的上下文。我们从 105 位活跃的计算机科学研究者处收集了超过 6,000 次双盲成对比较,并在共享封闭上下文协议下构建了计算机科学领域提案阶段专家偏好的 Elo 评分排行榜。我们通过评估者间一致性和稳健性分析对排名进行了验证,结果表明该排行榜在评估者构成和领域覆盖范围变化的情况下仍保持稳定。我们的结果显示智能体的有效性存在显著差异:一些框架相比其骨干模型提升了想法生成的质量,而另一些则收效甚微,甚至低于其基础模型的表现。我们进一步构建了 Ideation Arena Eval,一个用于评估自动评估器是否与人类在研究想法生成方面的偏好保持一致的基准。对当前大语言模型评判器的实验表明,它们仍无法可靠地复现专家偏好,最佳评判器在整体质量维度上的软准确率(Soft Accuracy)仅为 72.56%。我们的代码、数据和排行榜已发布于 https://github.com/foss12138/Research-Ideation-Arena。
cs.AI / 107 / 2608.29753

PAGE-RAG: Provenance-Aware Graph Evidence Promotion for Fixed-Budget Multi-hop Retrieval-Augmented Generation

PAGE-RAG:面向固定预算多跳检索增强生成的溯源感知图证据提升方法
Deng, Haokun, Li, Xunkai, Qin, Hongchao, Li, Rong-Hua
Abstract
Multi-hop question answering in retrieval-augmented gener?ation (RAG) often benefits from retrieving beyond the few candidates that will finally be read: narrow retrieval can miss an indispensable hop, while expanded retrieval introduces topical distractors. This challenge is not tied to a particu?lar knowledge-base format. Candidate pools may come from standalone retrievers, standard RAG backends, or graph-based retrieval pipelines. What is needed is a query-aware selection layer that can use relational structure to filter candidates be?fore generation. PAGE-RAG addresses this setting by using a graph as a temporary selection structure, rather than assum?ing a graph-structured knowledge base. It builds a query-local graph over retrieved candidates, records why candidates are connected, and treats each connection as a support hypothe?sis rather than support itself. We identify the resulting failure mode as a connectivity-support gap: connected candidates do not necessarily support the answer. We propose PAGE-RAG, a Provenance-Aware Graph Evidence promotion method that scores candidate paths with relevance, source-tracing meta?data, specificity, hubness, noise, and coherence signals, and applies minimal sufficient selection to promote supporting facts into a compact reader context. PAGE-RAG can serve as a complete retrieval-to-reading pipeline, and the same promo?tion stage can be inserted after existing retrieval or RAG sys?tems without replacing their upstream retrieval logic. Across three multi-hop QA benchmarks under the same final bud?get, PAGE-RAG improves support F1 and answer F1 by 10.4 and 3.3 points on a weighted average over a strong retriever. As a plug-in, PAGE-RAG further improves all reported RAG backends, including reasoning-oriented, compression-based, graph-based, and document/chunk-level systems.
Chinese Translation
检索增强生成(RAG)中的多跳问答通常受益于超出最终会被阅读的少量候选的检索:狭窄的检索可能遗漏不可或缺的一跳,而扩展的检索又会引入主题上不相关的干扰信息。这一挑战并不依赖于特定的知识库格式:候选池可能来自独立的检索器、标准RAG后端,或基于图的检索流水线。所需要的是一个查询感知的选择层,能够在生成之前利用关系结构对候选进行过滤。PAGE-RAG通过将图用作一种临时的选择结构来应对这一场景,而非假设存在图结构化的知识库。它在检索到的候选之上构建查询局部图,记录候选之间为何相连,并将每条连接视为支持假设而非支持本身。我们将由此产生的失败模式识别为“连通性-支持差距”(connectivity-support gap):相互连接的候选并不一定支持答案。我们提出PAGE-RAG,一种溯源感知的图证据提升方法,它利用相关性、来源追踪元数据、特异性、枢纽性、噪声和一致性信号对候选路径进行评分,并通过最小充分选择将支持性事实提升到紧凑的阅读器上下文中。PAGE-RAG既可以作为一个完整的检索到阅读的流水线,其提升阶段也可以作为一个可插入模块,嵌入现有的检索或RAG系统之后,而无需替换其上游检索逻辑。在三个多跳问答基准上、在相同最终预算条件下,PAGE-RAG相对于一个强检索器,在加权平均上将支持F1和答案F1分别提升了10.4和3.3个百分点。作为插件,PAGE-RAG进一步改进了所有已报告的RAG后端,包括面向推理的、基于压缩的、基于图的以及文档/块级系统。
cs.AI / 108 / 2608.29814

FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production

FRAMEWORKERS:一个用于AI生成视频制作的动态多智能体框架
Li, Zhendong, Sun, Lei, Shi, Letian, Zhang, Deheng, Ming, Ruibo, Hu, Mengshun, Xu, Dannong, Wang, Jian, Paudel, Danda, Van Gool, Luc, Gu, Jinjin
Abstract
Modern video generators excel at synthesizing individual clips, but complete video production requires coordinating a long sequence of interdependent creative steps, including scripting, storyboarding, generation, and editing. It further demands persistent asset management and dynamic task orchestration as intermediate outputs, dependencies, and execution states evolve over time. Existing automated systems typically rely on rigid pipelines that are difficult to adapt to diverse inputs and changing workflows, while general-purpose large language models (LLMs) remain unreliable for long-horizon orchestration and multimodal asset routing. We introduce FRAMEWORKERS, a task-centric and workspace-grounded multi-agent framework for open-ended video production. A central Director formulates video creation as dynamic task management, continuously editing a Task Stack to determine which subtask to execute next and which sub-agent to invoke. An Assistant serves as the execution layer, grounding each selected task in a shared Workspace, retrieving the required assets and context, invoking the assigned sub-agent, and persisting the resulting artifacts. Execution capabilities are exposed through modular sub-agents with registered descriptors, allowing new sub-agents to be integrated without redesigning the orchestration workflow. To improve orchestration reliability, we fine-tune the Director via supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) for descriptor-conditioned task routing. Experiments show that FRAMEWORKERS outperforms strong LLM planners in routing accuracy, recovers reliably from runtime failures, generalizes to unseen sub-agents without retraining, and achieves higher end-to-end video quality and broader task coverage than fixed pipelines, single-agent systems, and prior multi-agent approaches.
Chinese Translation
现代视频生成器擅长合成单个视频片段,但完整的视频制作需要协调一系列相互依赖的创作步骤,包括剧本编写、分镜设计、生成和剪辑。随着中间输出、依赖关系和执行状态随时间演变,这还需要持久的素材管理和动态的任务编排。现有的自动化系统通常依赖于刚性的流水线,难以适应多样化的输入和不断变化的工作流程,而通用大语言模型(LLM)在长程编排和多模态素材路由方面仍然不可靠。我们提出了FRAMEWORKERS,一个以任务为中心、以工作区为基础的面向开放式视频制作的多智能体框架。一个核心的Director(导演)智能体将视频创作建模为动态任务管理,通过持续编辑任务栈(Task Stack)来决定下一步执行哪个子任务以及调用哪个子智能体。一个Assistant(助理)智能体作为执行层,将每个选定的任务落地到共享工作区(Workspace)中,检索所需的素材和上下文,调用被分配的子智能体,并持久化生成结果。执行能力通过带有注册描述符的模块化子智能体暴露,使新的子智能体能够在无需重新设计编排工作流程的情况下被集成。为提高编排可靠性,我们通过监督微调(SFT)以及基于描述符条件任务路由的群体相对策略优化(GRPO)对Director进行微调。实验表明,FRAMEWORKERS在路由准确率上优于强大的LLM规划器,能够可靠地从运行时故障中恢复,无需重新训练即可泛化到未见过的子智能体,并且相比固定流水线、单智能体系统和先前的多智能体方法,实现了更高的端到端视频质量和更广的任务覆盖范围。
cs.AI / 109 / 2608.29880

Perceive to Hypothesize, Verify to Ground: An Agentic Reasoning Framework for Open-World Geo-Localization

感知以假设,验证以落地:面向开放世界地理定位的智能体推理框架
Jiang, Yutian, Li, Ruijie, Lyu, Sisuo, Hao, Xixuan, Liu, Qingxiang, Yu, Yongzi, Liang, Yuxuan
Abstract
Open-world geo-localization requires models to reason over ambiguous visual cues through multi-step reasoning and external knowledge grounding. While recent large vision-language models exhibit strong multimodal reasoning capabilities, existing approaches still suffer from perceptual hallucination and context drift due to the lack of explicit evidence-grounded verification. In this work, we reformulate geo-localization as a human-like perceive-then-verify reasoning problem and propose GeoPAVE (Geo-localization Perception-and-Verification-Engine), a bi-level agentic framework that contains perception-based hypothesis generation via single-pass rollouts and verification-based evidence grounding for decision actions: support, refute, and refine. To support rigorous evaluation, we further introduce PAVED, a novel dataset derived from real-world user check-in data, equipped with comprehensive reasoning trajectories featuring multi-hop queries, multi-round tool invocations, and structured perception-verification traces. The dataset and code are available at https://github.com/Arandinglv/GeoPAVE.
Chinese Translation
开放世界地理定位要求模型通过多步推理和外部知识落地(grounding)对模糊的视觉线索进行推理。尽管近期的大型视觉语言模型展现出强大的多模态推理能力,但由于缺乏显式的基于证据的验证,现有方法仍然存在感知幻觉和上下文漂移的问题。在本工作中,我们将地理定位重新表述为一个类人的“先感知后验证”推理问题,并提出了GeoPAVE(地理定位感知与验证引擎,Geo-localization Perception-and-Verification-Engine),这是一个双层智能体框架,包含:通过单次推理(single-pass rollouts)实现基于感知的假设生成,以及基于验证的证据落地以做出决策动作——支持(support)、反驳(refute)和修正(refine)。为支持严格的评估,我们进一步引入了PAVED,这是一个源自真实世界用户签到数据的新型数据集,配备了完整的推理轨迹,包括多跳查询、多轮工具调用以及结构化的感知-验证痕迹。数据集和代码可在 https://github.com/Arandinglv/GeoPAVE 获取。
cs.AI / 110 / 2608.29913

On the Instance Hardness as a Decision Criterion in TinyML Systems

论实例难度作为TinyML系统中的决策准则
Puslecki, Tobiasz, Walkowiak, Krzysztof
Abstract
TinyML includes the implementation of machine learning on devices with limited memory and computing resources. With the development of technology, AI systems continue to scale in terms of size and computational requirements. This forces researchers to adapt methods to be environmentally sustainable by designing techniques for reducing computational costs and energy consumption in inferring AI models, even in small devices. In this work, we present preliminary findings on a novel application of the tree depth prune instance hardness method to the TinyML system. The results indicate that threshold control can change energy consumption with limited classification quality changes. This method allows us to adjust classification accuracy, thereby influencing computational complexity and energy consumption for inference. We present a work in progress with initial results as a proof of concept.
Chinese Translation
TinyML是指在内存和计算资源有限的设备上实现机器学习。随着技术的发展,人工智能系统在规模和计算需求方面不断增长。这迫使研究人员通过设计技术来降低人工智能模型推理过程中的计算成本和能耗,即使在小型设备上也是如此,从而使方法具有环境可持续性。在本工作中,我们提出了将树深度剪枝实例难度(tree depth prune instance hardness)方法应用于TinyML系统这一新颖应用的初步研究结果。结果表明,阈值控制可以在分类质量变化有限的情况下改变能耗。该方法使我们能够调整分类精度,从而影响推理的计算复杂度和能耗。我们展示了一项正在进行中的工作及其实证性验证的初步结果。
cs.AI / 111 / 2608.29937

AcrossWAM1.0:A Modular Latent World-Action Stack for Compact Robot Policies

AcrossWAM1.0:用于紧凑机器人策略的模块化潜在世界-动作堆栈
Zhang, Yafei, Wu, Nan
Abstract
Latent world-action models avoid rendering future pixels by predicting an action-relevant visual subgoal in feature space. LaWAM established this formulation, but its original presentation left the world model, multimodal backbone, and deployment checkpoint tightly coupled. We introduce AcrossWAM1.0, a modularization and scaling study of this latent world-action stack. Rather than presenting latent subgoals as a new algorithm, we make the module boundary explicit: a policy adapter produces latent-action and action-generation contexts; a retained latent world decoder grounds the predicted transition in the current scene;and a flow-matching expert generates continuous action chunks. We further separate training-only teachers from the inference graph and provide a verifiable deployment export. On 2,000 paired LIBERO episodes, replacing a Qwen3-VL-2B backbone with Qwen3.5-0.8B yields 97.45% success versus 98.00% for the 2B model (a-0.55percentage-point difference; exact McNemarp=0.266). This does not prove equivalence, but it meets a prespecified two-point retention criterion. The compact, inference-reachable checkpoint contains 1,472.6M unique parameters, 42.4% fewer than the original 2B policy, while all retained tensors are bitwise identical to the source checkpoint. Cross-family execution is additionally checked with a MiniCPM-V adapter smoke test; closed-loop cross-family transfer remains an open evaluation. AcrossWAM1.0 therefore contributes an auditable software and evaluation boundary for compact latent world-action policies, distinct from LaWAM's original latent-subgoal contribution.
Chinese Translation
潜在世界-动作模型通过在特征空间中预测与动作相关的视觉子目标,避免了未来像素的渲染。LaWAM 确立了这一范式,但其原始实现将世界模型、多模态骨干网络与部署检查点紧密耦合。我们提出 AcrossWAM1.0,对这一潜在世界-动作堆栈进行模块化与规模化的系统研究。我们并非将潜在子目标作为一种新算法提出,而是显式地划定模块边界:策略适配器生成潜在动作上下文和动作生成上下文;保留的潜在世界解码器将预测的状态转移锚定在当前场景中;流匹配专家生成连续动作块。我们进一步将仅用于训练的教师模型从推理图中分离,并提供可验证的部署导出。在 2,000 个配对的 LIBERO 轨迹上,将 Qwen3-VL-2B 骨干替换为 Qwen3.5-0.8B 后,成功率为 97.45%,而 2B 模型为 98.00%(相差 0.55 个百分点;McNemar 精确检验 p=0.266)。这并不能证明二者等价,但满足了预先设定的两点保持标准。该紧凑的、可在推理阶段达到的检查点包含 1,472.6M 个独立参数,比原始 2B 策略少 42.4%,且所有保留的张量与源检查点在比特层面完全一致。此外,通过 MiniCPM-V 适配器的冒烟测试对跨模型族的执行进行了检验;闭环跨模型族迁移仍是待评估的开放问题。因此,AcrossWAM1.0 为紧凑型潜在世界-动作策略提供了一个可审计的软件与评估边界,区别于 LaWAM 最初的潜在子目标贡献。
cs.AI / 112 / 2608.29951

Spatial Matryoshka Training for Multi-Granularity Visual Document Retrieval

面向多粒度视觉文档检索的空间俄罗斯套娃训练方法
Roy, Trishan Singha, Acharya, Arkadeep, Kumar, Vishwajeet, Sen, Jaydeep, Joshi, Sachindra
Abstract
Multi-modal late-interaction retrievers achieve strong retrieval on visually rich documents by representing each page as per patch embeddings and matching at the token level. However, this approach incurs high storage costs. Existing compression methods typically fix a single compression level at indexing time, limiting flexibility. We present ColSNAP (Spatial Nested Average Pooling)1, a training method that generates a nested hierarchy of compression levels directly from a backbone's patch grid. By spatially pooling patch embeddings into pro- gressively coarser tiers and training all tiers simultaneously, a single model learns to support retrieval at multiple compression levels without architectural changes. Crucially, a single encoding pass yields every tier, enabling the accuracy-storage trade-off to be configured at indexing time to match avail- able storage budgets, rather than being fixed during training. We demonstrate that models trained using ColSNAP maintain near full-resolution retrieval performance under substantial compression and that ColSNAP transfers effectively across multiple late-interaction backbones, and achieves most of its improvements via a lightweight adaptation stage applied to a pre-trained retriever.
Chinese Translation
多模态后期交互检索器通过将每个页面表示为逐图像块嵌入并在词元级别进行匹配,在视觉丰富的文档上实现了强大的检索性能。然而,这种方法带来了高昂的存储成本。现有的压缩方法通常在建立索引时固定单一压缩级别,限制了灵活性。我们提出了ColSNAP(空间嵌套平均池化),一种直接从骨干网络的图像块网格生成嵌套式压缩级别层级的训练方法。通过将图像块嵌入空间池化为逐渐粗化的层级,并同时训练所有层级,单个模型无需架构改动即可学会支持多个压缩级别的检索。关键在于,单次编码即可产生所有层级,使得精度-存储的权衡可以在建立索引时进行配置以匹配可用的存储预算,而不是在训练期间被固定。我们证明了使用ColSNAP训练的模型在大幅压缩下仍能保持接近全分辨率的检索性能,并且ColSNAP能够在多个后期交互骨干网络之间有效迁移,其大部分性能提升可通过对预训练检索器施加一个轻量级适配阶段来实现。
cs.AI / 113 / 2608.29953

SearchWiki: Learning to Build and Navigate Knowledge Wikis for Active Information Seeking

SearchWiki:学习构建和导航知识维基以实现主动信息检索
Singh, Guransh, Kumar, Vishwajeet, Acharya, Arkadeep, Qidwai, Adnan, Sen, Jaydeep, Joshi, Sachindra
Abstract
Flat retrieval-augmented generation treats a corpus as a bag of chunks, discarding document hierarchy and cross document structure. We introduce SearchWiki, a harness framework that synthesizes a corpus into a hierarchical, typed, navigable wiki and trains an agent, WikiResearcher-9B, to retrieve information through multi-turn tool use. The wiki organizes knowledge into three layers - document overviews, cross- document topic pages, and page-level source records; enabling progressive refinement of retrieval when initial lookup misses. We optimize the agent's navigation policy with on-policy reinforcement learning with a multi-component reward function balancing answer correctness, retrieval quality and trajectory efficiency. Evaluation on ViDoRe-V3 (8 domains), FinanceBench, and memory benchmarks (LoCoMo, LongMemEval, PersonaMem-v2) shows that WikiResearcher- 9B which is our RL-tuned Qwen 9B model, significantly outperforms same-size untrained baselines and exceeds or matches larger external models. SearchWiki paired with WikiResearcher-9B demonstrates that learned navigation over structured corpora is a superior alternative to flat retrieval.
Chinese Translation
扁平化的检索增强生成将语料库视为一系列文本块的集合,丢弃了文档层级和跨文档结构。我们提出了SearchWiki,一个将语料库综合为分层的、带类型的、可导航维基的框架,并训练了一个名为WikiResearcher-9B的智能体,通过多轮工具调用来检索信息。该维基将知识组织为三个层次——文档概览、跨文档主题页面和页面级源记录——使得在初始查找未命中时能够进行渐进式的检索细化。我们采用在线策略强化学习来优化智能体的导航策略,奖励函数包含多个组件,以平衡答案正确性、检索质量和轨迹效率。在ViDoRe-V3(8个领域)、FinanceBench以及记忆基准(LoCoMo、LongMemEval、PersonaMem-v2)上的评估表明,我们的WikiResearcher-9B(即经强化学习微调的Qwen 9B模型)显著优于同等规模的未训练基线,并超越或媲美更大的外部模型。SearchWiki与WikiResearcher-9B的结合证明了,在结构化语料库上进行学习式导航是优于扁平化检索的一种替代方案。
cs.AI / 114 / 2608.29965

Review Before Trust: Source-Grounded Integrity Gates for AI-Assisted Personal Health Records

信任之前先审查:面向AI辅助个人健康记录的基于来源的完整性门控
Girda, Nora, Groza, Adrian
Abstract
Large language models can convert medical documents into structured data, but plausible output may still be unsupported by the source. Persisting such output in a longitudinal health record, a record that accumulates patient information over time, therefore creates an integrity risk: unverified data may influence later summaries, trends, or preventive-care computations. We introduce an evidence-gated trust-promotion model that keeps generated data provisional until a deterministic monitor verifies it against the source document. The monitor admits a candidate for a specified downstream use only when the source contains a unique supporting quotation, the relevant fields occur within the same laboratory row, and the required provenance is preserved. The generator cannot approve its own output, missing or ambiguous evidence causes refusal, and refused candidates remain available for human review rather than being silently discarded. We implement the model in Medical DataCloud, a personal health-record application, and evaluate it through automated tests and a replay of saved extraction outputs. All 22 conformance and mutation tests pass. The replay covers nine historical laboratory PDF reports containing 102 manually labelled rows. The reports produce 97 numeric candidates: schema validation accepts all 97, an earlier packet-level evidence check accepts 94, and the hardened quotation- and row-level policy admits 72 while retaining 25 for review. The study evaluates system integrity rather than clinical correctness or clinical safety. The results demonstrate the technical feasibility of an enforceable boundary that prevents generated claims from authorizing their own reuse in a longitudinal health record.
Chinese Translation
大语言模型可以将医疗文档转换为结构化数据,但其看似合理的输出仍可能缺乏来源支持。将此类输出持久化到纵向健康记录(即随时间累积患者信息的记录)中会带来完整性风险:未经核验的数据可能影响后续的摘要、趋势分析或预防性保健计算。我们提出了一种证据门控的信任提升模型,该模型在确定性监视器依据来源文档进行核验之前,将生成的数据保持为临时状态。仅当来源中包含唯一的支持性引文、相关字段位于同一化验行内且所需的来源溯源信息得以保留时,监视器才允许候选数据用于指定的下游用途。生成器不能批准自身的输出,证据缺失或含糊不清将导致拒绝,而被拒绝的候选数据仍可供人工审查,而非被静默丢弃。我们在Medical DataCloud(一个个人健康记录应用)中实现了该模型,并通过自动化测试和已保存提取输出的重放对其进行评估。全部22项一致性与变异测试均通过。重放涵盖九份历史化验PDF报告,共包含102个手工标注的行。这些报告生成了97个数值候选数据:模式验证接受了全部97个,早期的数据包级证据检查接受了94个,而强化后的引文级和行级策略仅接受72个,另有25个保留待审查。本研究评估的是系统完整性而非临床正确性或临床安全性。结果证明了可执行边界的可行性,该边界可防止生成的内容自行授权其在纵向健康记录中的重复使用。
cs.AI / 115 / 2608.29971

EDGE: Engine for Deterministic Graph Evaluation through Conversation Simulation from Graph Structured DSL Configuration

EDGE:基于图结构化DSL配置、通过对话模拟实现确定性图评估的引擎
Kulathumani, Ram, Radhakrishnan, Regunathan, Tripathi, Anupam, Mao, Xiangbo, Omrani, Roshanak, Somani, Keshav, Mishra, Shwet Kamal, Lurya, Shayna
Abstract
As agentic systems evolve into complex multi agent orchestration workflows, there is a growing and critical need for systematic frameworks that measures an agent's behavioral consistency and determinism. In this paper, we introduce a formal evaluation methodology that is grounded in AgentGraph, a planner powered by a domain specific language that represents agent reasoning through a dynamically adjustable directed graph. We leverage this structural formalism and utilize graph traversal algorithms that exhaustively enumerate conversational paths, forming a comprehensive evaluation set that captures the agent's complete behavioral space. We then systematically replay these reproducible trajectories to compare observed outputs and state transitions against the intended DSL specification. To quantify reliability, we define novel metrics that measure response and trajectory determinism, structural adherence and semantic consistency across both exact replays and their linguistic variants. Our system's results demonstrate that agents configured using frameworks like AgentGraph and LangGraph with explicitly structured node transitions show superior determinism over agents that are not configured with controlled transitions.
Chinese Translation
随着智能体系统演变为复杂的多智能体编排工作流,对能够衡量智能体行为一致性与确定性的系统性框架的需求日益迫切。本文提出一种形式化的评估方法,其基础是AgentGraph——一个由领域特定语言(DSL)驱动的规划器,通过动态可调的有向图来表示智能体的推理过程。我们利用这一结构化形式体系,并采用图遍历算法穷举枚举对话路径,从而构建一个全面覆盖智能体完整行为空间的评估集。随后,我们系统地重放这些可复现的轨迹,将观测到的输出与状态转移与预期的DSL规范进行对比。为量化可靠性,我们定义了新的指标,用于在精确重放及其语言变体中衡量响应与轨迹的确定性、结构一致性和语义一致性。系统实验结果表明,使用AgentGraph和LangGraph等具有显式结构化节点转移的框架所配置的智能体,在确定性方面优于未采用受控转移配置的智能体。
cs.AI / 116 / 2608.29973

An Open-Source, Event-Driven Pipeline for Cryptocurrency Market Data: Ingestion, Forecasting, and On-Chain Fraud Detection

一个用于加密货币市场数据的开源事件驱动流水线:数据摄取、预测与链上欺诈检测
Shaikh, Basil Sajid, Mascarenhas, Melrick, Shaikh, Nuzhat Faiz
Abstract
Cryptocurrency markets generate high-frequency, multi-source data that is expensive to work with unless a team already has commercial-grade streaming and warehousing infrastructure in place. This paper describes a fully open-source pipeline that reproduces the behavior of a cloud-native, event-driven system -- file arrival triggering a message, a message triggering compute -- entirely on commodity hardware, using Apache Kafka and a filesystem-watching poller in place of managed cloud triggers. The pipeline partitions historical Gemini exchange data into hourly and minutely files, ingests them asynchronously through two independently grouped Kafka consumers (one for audit logging, one for Spark-triggered ETL), and lands cleaned output in a PostgreSQL warehouse with historical and aggregated schemas plus asset-specific data marts. We use the resulting Bitcoin data mart to compare a seasonal ARIMA model against a single-layer LSTM network for price forecasting, and separately apply Random Forest and Gradient Boosting classifiers, with additional engineered features, to the public Ethereum fraud detection benchmark introduced by Farrugia et al. We report the architecture, the modeling methodology, and the resulting metrics, and we are explicit about the limitations of comparing forecasts issued at different horizons and of evaluating fraud detection on a static, already-labeled dataset.
Chinese Translation
加密货币市场产生高频、多源的数据,除非团队已具备商用级的流处理和数据仓库基础设施,否则处理这些数据的成本很高。本文描述了一个完全开源的流水线,它在普通硬件上完整复现了云原生、事件驱动系统的行为——文件到达触发消息、消息触发计算——使用 Apache Kafka 和文件系统监视轮询器来替代托管云触发器。该流水线将历史 Gemini 交易所数据切分为小时级和分钟级文件,通过两个独立分组的 Kafka 消费者(一个用于审计日志,一个用于 Spark 触发的 ETL)异步摄取,并将清洗后的输出存入 PostgreSQL 数据仓库,其中包含历史模式、聚合模式以及针对特定资产的数据集市。我们利用得到的比特币数据集市,将季节性 ARIMA 模型与单层 LSTM 网络在价格预测任务上进行比较;另外,我们在 Farrugia 等人提出的公开以太坊欺诈检测基准数据集上,应用随机森林(Random Forest)和梯度提升(Gradient Boosting)分类器并结合额外构造的特征进行欺诈检测。我们报告了系统架构、建模方法及所得指标,并明确指出了比较不同预测时间跨度下预测结果、以及在静态且已标注数据集上评估欺诈检测效果的局限性。
cs.AI / 117 / 2608.29988

AutoCRAT: Within-trajectory Joint Control of Stochasticity and Compute for LLM Reasoning

AutoCRAT:面向大语言模型推理的轨迹内随机性与计算量的联合控制
Luo, Hanjun, Liu, Qiushi, Zhang, Jingya, Pang, Haihong, Wen, Jiaheng, Ma, Yifei, Yao, Yu, Zhang, Chengxi, Zhang, Hanrong, Chen, Yankai, Salam, Hanan
Abstract
Large language models (LLMs) achieve strong reasoning performance, which depends critically on inference-time decisions. Yet these decisions are commonly handled by static, one-size-fits-all policies, limiting adaptation to diverse tasks and reasoning stages. Recent adaptive methods partially address this limitation, but they primarily adapt either decoding stochasticity (how the model explores) or reasoning compute (how long the model reasons) in isolation, leaving their interaction within a single reasoning trajectory unmodeled. To address this challenge, we shift toward a within-trajectory joint control view, and instantiate it in AutoCRAT, a decoder-side controller for frozen backbones. Using only signals available during decoding, AutoCRAT jointly adjusts sampling stochasticity and reasoning budget during generation. AutoCRAT operates over a discrete action space and updates control decisions only at semantic boundaries, improving stability while remaining responsive to the evolving reasoning process. Comprehensive evaluation across 6 benchmarks demonstrates that AutoCRAT (I) uses 13.8-52.7% fewer inference tokens on average than recommended static configurations, (II) surpasses recommended static and adaptive baselines by 1.5-4.5% in relative accuracy, and (III) enjoys strong cross-backbone transferability.
Chinese Translation
大语言模型(LLM)已展现出强大的推理能力,而这在关键程度上依赖于推理时的决策。然而,这些决策通常由静态的、一刀切的策略处理,限制了对多样任务和推理阶段的适应能力。近期的自适应方法部分解决了这一局限,但它们主要单独地自适应调整解码随机性(模型如何探索)或推理计算量(模型推理多长时间),而未对二者在单个推理轨迹内的交互进行建模。为应对这一挑战,我们转向轨迹内联合控制的视角,并将其实现为 AutoCRAT——一个面向冻结骨干模型的解码器侧控制器。AutoCRAT 仅利用解码过程中可获得的信号,在生成过程中联合调整采样随机性与推理预算。AutoCRAT 在离散动作空间上运作,且仅在语义边界处更新控制决策,在保持对不断演化的推理过程的响应性的同时提升了稳定性。在 6 个基准上的全面评估表明,AutoCRAT(I)相比推荐的静态配置平均减少 13.8–52.7% 的推理 token;(II) 相对准确率超出推荐的静态与自适应基线 1.5–4.5%;(III) 具有较强的跨骨干模型迁移能力。
cs.AI / 118 / 2608.30022

Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models

利用大语言模型将NICE指南自动转换为可执行的计算模型
Gupta, Ashvin, Prociuk, Denys, Russo, Alessandra, Delaney, Brendan C.
Abstract
Introduction: NICE guidelines provide evidence-based recommendations for clinical care but remain largely in unstructured natural language. Existing approaches to converting them into computable representations often focus on individual diseases, require substantial manual encoding, and do not scale. Large language models (LLMs) may enable much of this translation to be automated. Methods: We present an end-to-end approach that converts textual clinical guidelines into executable models capable of generating explainable, patient-specific recommendations. A stepwise LLM-based transformation with in-context examples produces human-inspectable intermediate artifacts. We apply the approach to NICE pancreatic and lung cancer guidelines, use expert review to assess rule alignment, and evaluate the executable pancreatic cancer model on 20 patient vignettes. Results: Expert review showed strong alignment between the source guidelines and generated executable models. Most discrepancies were partial omissions rather than incorrect logic, while hallucinated or fundamentally incorrect rules were rare. On the patient vignettes, the executable model achieved an F1 score of 82.5%. Conclusion: LLMs can transform natural-language NICE guidelines into interpretable, executable models that preserve guideline structure, support transparent inspection and modification, and generate patient-specific recommendations. These findings demonstrate the feasibility of scalable automated generation of computable clinical guidelines.
Chinese Translation
引言:NICE(英国国家卫生与临床优化研究所)指南为临床诊疗提供循证建议,但大部分仍以非结构化的自然语言形式存在。现有的将其转换为可计算表示的方法通常仅针对个别疾病,需要大量人工编码,且难以规模化。大语言模型(LLM)或可使这一转换过程大部分实现自动化。方法:我们提出了一种端到端的方法,将文本形式的临床指南转换为能够生成可解释、患者特异性建议的可执行模型。基于大语言模型的分步转换方法结合上下文示例,生成可供人工检查的中间产物。我们将该方法应用于NICE胰腺癌和肺癌指南,通过专家评审评估生成规则与原指南的一致性,并在20个患者病例情景上评估了可执行的胰腺癌模型。结果:专家评审显示,源指南与生成的可执行模型之间具有较强的一致性。大多数差异属于部分遗漏而非逻辑错误,而幻觉或根本错误的规则较为罕见。在患者病例情景上,可执行模型的F1分数达到82.5%。结论:大语言模型能够将自然语言形式的NICE指南转换为可解释、可执行的模型,这些模型保留了指南的结构,支持透明化的检查与修改,并能生成患者特异性的建议。这些发现证明了大规模自动生成可计算临床指南的可行性。
cs.AI / 119 / 2608.30025

Interpreting and Steering for Safe and Correct Code Generation

面向安全且正确代码生成的解释与引导
Yan, Hao, Yao, Ziyu
Abstract
Large language models (LLMs) frequently generate source code containing vulnerabilities, yet little work studies the internal mechanisms that distinguish safe from vulnerable generation in them. In this work, we systematically perform a mechanistic interpretation of LLMs, aiming at both understanding how code safety-vs-vulnerability is represented or driven by components in an LM and turning the insights into actionable steering strategies to encourage safer code generation. To this end, we introduce CodeSec-Pairs, a dataset of 9,342 Python safe-and-vulnerable contrastive code pairs, sampled from Llama-3.1-8B-Instruct. Utilizing the dataset, we explore approaches to localize layers and attention heads that relate to code safety, and further experiment with different steering strategies for inference-time vulnerability reduction. In particular, we propose DuoSteer, a double-steering approach that simultaneously applies safety and code-correctness steering to attention heads. In experiments over five vulnerability types, DuoSteer leads to an average of -26.9% vulnerability rate reduction and +7.5% functional correctness improvement, which outperforms not only other steering variants but also prompting and supervised fine-tuning baselines. The advantage also replicates on Qwen-2.5-Coder-7B-Instruct with another 2,500 contrastive pairs sampled from that model.
Chinese Translation
大语言模型(LLM)频繁生成含有漏洞的源代码,然而鲜有工作研究区分安全生成与漏洞生成之间的内部机制。在本工作中,我们系统地对LLM进行了机制可解释性分析,旨在一方面理解代码的安全性(安全vs漏洞)如何被语言模型中的组件所表示或驱动,另一方面将这些洞察转化为可操作的引导策略,以促进更安全的代码生成。为此,我们构建了CodeSec-Pairs,一个包含9,342对Python安全与漏洞对比代码对的数据集,这些样本取自Llama-3.1-8B-Instruct。利用该数据集,我们探索了定位与代码安全性相关的层和注意力头的方法,并进一步实验了不同的引导策略以在推理时降低漏洞率。特别地,我们提出了DuoSteer,一种双重引导方法,可同时对注意力头施加安全引导和代码正确性引导。在五种漏洞类型上的实验表明,DuoSteer平均降低了26.9%的漏洞率,并提升了7.5%的功能正确性,其表现不仅优于其他引导变体,也优于提示(prompting)和监督微调基线方法。该优势在Qwen-2.5-Coder-7B-Instruct上同样得到复现,该实验使用了从该模型采样的另外2,500对对比代码对。
cs.AI / 120 / 2608.30035

Beyond Uncertainty: Multi-Solver Disagreement Rewards for Self-Evolving Reasoning Curricula

超越不确定性:面向自进化推理课程的多求解器分歧奖励
Selvendran, Vinoth, Zhang, Zhanming
Abstract
Self-evolving reasoning frameworks train a Challenger to generate questions exposing a Solver's weaknesses, creating adaptive curricula without human data. However, existing approaches use a single solver's sampling uncertainty as the Challenger's reward. This creates a fundamental bottleneck: as the solver grows confident on the Challenger's question distribution, all sampled answers converge identically, collapsing the reward to zero and starving the Challenger of learning signal. Critically, this single-model reward cannot distinguish genuinely easy questions from those that merely align with one solver's learned biases. We propose a multi-solver disagreement reward using a heterogeneous ensemble varying in model capacity and sampling temperature. A normalized Shannon entropy over the ensemble's per-question plurality answers explicitly rewards questions where solvers produce conflicting solutions---capturing difficulty as inter-model divergence rather than intra-model sampling variance. This richer gradient enables the Challenger to discover questions targeting true capability boundaries, producing a curriculum that forces downstream Solvers to develop robust reasoning strategies generalizing across problem types. Our approach is a drop-in reward function replacement requiring no framework modifications or additional data. Experiments with Qwen3-4B show that Solvers trained on disagreement-Challenger questions achieve +1.34 points average improvement on competition-math benchmarks (MATH-500, AMC, Olympiad), suggesting that multi-solver disagreement provides a complementary and scalable signal for curriculum generation in self-play reasoning systems.
Chinese Translation
自进化推理框架通过训练一个挑战者(Challenger)来生成能够暴露求解器(Solver)弱点的题目,从而在无需人类数据的情况下构建自适应课程。然而,现有方法仅使用单个求解器的采样不确定性作为挑战者的奖励。这造成了一个根本性的瓶颈:当求解器对挑战者生成的题目分布逐渐变得自信时,所有采样答案会趋于一致,导致奖励坍缩为零,使挑战者缺乏学习信号。更关键的是,这种单模型奖励无法区分真正简单的题目与那些仅仅契合单一求解器已学习偏见的题目。我们提出了一种基于异构集成(模型容量与采样温度各不相同)的多求解器分歧奖励。通过对集成中每个题目的多数答案计算归一化香农熵,该方法显式地奖励那些求解器产生相互矛盾解的题目——将难度刻画为模型间分歧而非模型内采样方差。这一更丰富的梯度使挑战者能够发现针对真实能力边界的题目,生成一种迫使下游求解器发展出可跨问题类型泛化的稳健推理策略的课程。我们的方法是一种即插即用的奖励函数替换方案,无需修改框架或额外数据。基于 Qwen3-4B 的实验表明,在分歧-挑战者题目上训练的求解器在竞赛数学基准(MATH-500、AMC、Olympiad)上平均提升 1.34 分,这表明多求解器分歧为自博弈推理系统中的课程生成提供了一种互补且可扩展的信号。
cs.AI / 121 / 2608.30044

Balance of Benchmarks: Semantic Density Reweighting for Benchmark Multiplicity and Task-Conditioned Evaluation

基准平衡:面向基准多重性与任务条件评估的语义密度重加权
Lin, Jhen-Ke
Abstract
Language models are commonly compared by averaging scores across a benchmark list with equal weight. Such lists grow through publication outside an explicit measurement design, so equal weighting turns the density of published benchmarks into an implicit capability weight: densely benchmarked regions count repeatedly. We introduce Balance of Benchmarks (BoB), which embeds benchmark descriptions and assigns each benchmark an inverse-density semantic weight. Nearby entries share aggregate influence at a disclosed density scale. After equating heterogeneous scores onto a common latent scale, a residual field uses the same geometry to condition model rankings on a task query. The two components serve distinct empirical roles. On a snapshot of 586 models and 14 benchmarks, BoB predicts which models are unusually strong on a held-out task beyond their general ability, reaching a profile correlation of 0.462 compared with 0.049 under equal weighting. It also limits the influence of densely repeated benchmarks on the aggregate. After adding four copies of each benchmark in turn, the resulting rankings retain a Kendall tau of 0.995, compared with 0.936 under equal weighting. The residual field therefore provides task-conditioned prediction, and inverse-density weighting provides robustness to benchmark multiplicity. Together, they turn benchmark-list composition from an incidental property of evaluation suites into an explicit, controllable part of measurement design, providing a principled foundation for task-aware and multiplicity-robust model evaluation.
Chinese Translation
语言模型通常通过等权重地平均一组基准测试的分数来进行比较。此类基准列表往往在缺乏明确测量设计的情况下通过论文发表而不断增长,因此等权重做法使已发表基准的密度成为隐性的能力权重:被密集基准测试覆盖的能力区域被重复计入。我们提出基准平衡方法(Balance of Benchmarks, BoB),该方法对基准描述进行嵌入,并为每个基准分配与密度成反比的语义权重。在给定且公开的密度尺度下,相近的基准条目共享总体影响力。在将异构分数统一映射到共同潜在量表之后,残差场利用相同的几何结构,使模型排名能够根据任务查询进行条件化。这两个组成部分扮演着不同的实证角色。在一个包含586个模型和14个基准的快照上,BoB能够预测哪些模型在其一般能力之外的某个留出任务上表现异常出色,其能力画像相关性达到0.462,而等权重方法仅为0.049。同时,BoB还限制了被密集重复的基准对总体结果的影响。在依次为每个基准添加四个副本后,所得排名仍保持0.995的Kendall tau,而等权重方法仅为0.936。因此,残差场提供了任务条件化的预测能力,而逆密度加权则提供了对基准多重性的鲁棒性。二者结合,将基准列表的构成从评估套件的偶然属性转变为测量设计中显式、可控的组成部分,为任务感知且对多重性鲁棒的模型评估提供了有原则的基础。
cs.AI / 122 / 2608.30047

Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks

LLM智能体能够发现吗?评估机器学习工程任务中的创造力
Bhushan, Shitanshu, Zhang, Yunxiang, Wang, Lu
Abstract
Recent AI systems promise autonomous scientific discovery, claiming to discover algorithms and produce research papers, yet understanding whether they exhibit creativity, the capacity to produce solutions that are both novel and useful, remains an open question. We present a framework for evaluating multi-turn LLM research agents' creativity using ML engineering tasks as a testbed, through three dimensions: P-Creativity (psychological novelty: novel relative to the agent's own prior solutions within a run), H-Creativity (historical novelty: novel relative to the corpus of human solutions), and Usefulness (task performance). Evaluating two agent frameworks, AIDE and AIRA-Dojo, on 10 Kaggle-style machine learning tasks from MLE-Bench, we develop an LLM-as-a-Judge pipeline and verify its strong correlation with human creativity judgments, providing a reliable automated metric for P-Creativity evaluation at scale. Applying this pipeline to agent trajectories, we find: (1) all agents exhibit declining P-Creativity as they transition from exploration to exploitation; (2) LLMs exhibit greater H-Creativity than medal-winning humans, yet achieve lower performance. Our findings reveal that current agents can explore novel regions of the solution space but lack the capacity to convert this novelty into improved task performance.
Chinese Translation
近期的人工智能系统宣称能够实现自主科学发现,声称可以发明算法并撰写研究论文,然而,这些系统是否展现出创造力——即产生既新颖又有用的解决方案的能力——仍是一个悬而未决的问题。我们提出了一个评估多轮LLM研究智能体创造力的框架,以机器学习(ML)工程任务作为测试平台,从三个维度进行评估:P-Creativity(心理新颖性:相对于智能体在单次运行中自身先前解决方案的新颖程度)、H-Creativity(历史新颖性:相对于人类解决方案语料库的新颖程度)以及有用性(任务性能)。我们在MLE-Bench的10个Kaggle风格的机器学习任务上评估了两个智能体框架AIDE和AIRA-Dojo,并开发了LLM-as-a-Judge评估流程,验证了其与人类创造力判断的高度相关性,为大规模P-Creativity评估提供了可靠的自动化指标。将该流程应用于智能体轨迹后,我们发现:(1) 所有智能体在从探索转向利用的过程中,P-Creativity均呈下降趋势;(2) LLM表现出比获得奖牌的人类参赛者更高的H-Creativity,但任务性能却更低。我们的研究结果表明,当前的智能体能够探索解决方案空间的新颖区域,但缺乏将这种新颖性转化为更高任务性能的能力。
cs.AI / 123 / 2608.30050

Spec2Twin-Chain: Orchestrating Bi-Level Optimization with LLMs for Blockchain Digital Twin Construction

Spec2Twin-Chain:利用大语言模型协同双层优化以构建区块链数字孪生
Zhang, Haoting, Chen, Haoxian, Sheng, Jiayuan, Zhan, Donglin, Zheng, Zeyu, Yao, David D., Tang, Wenpin
Abstract
Building a blockchain digital twin largely requires translating domain knowledge and specific system descriptions into a simulator architecture, calibrating its parameters against behavioral evidence, and validating the constructed twin. These steps are commonly performed through application-specific modeling efforts that can be difficult to reuse across systems and downstream decision problems. We consider automating this process through Spec2Twin-Chain, a framework that formulates blockchain digital-twin construction as a bi-level optimization problem. At the upper level, a large language model proposes and revises structurally admissible architectures using system specifications, behavioral evidence, and feedback from evaluated designs. At the lower level, a simulation-based optimizer calibrates the architecture-conditioned parameters under explicit objectives and guardrail constraints. The two levels iterate. The evaluated candidates at lower levels are retained in a global archive and used to guide subsequent proposals at upper levels. We conduct controlled experiments involving twin calibration, feedback-driven recovery, stress analysis, downstream policy optimization, and policy updating. The results demonstrate that the framework can construct behaviorally accurate twins, improve initial designs through iterative feedback, and reuse calibrated twins to support downstream decisions.
Chinese Translation
构建区块链数字孪生在很大程度上需要将领域知识和特定系统描述转化为仿真器架构,根据行为证据校准其参数,并对所构建的孪生体进行验证。这些步骤通常通过针对特定应用的建模工作来完成,难以在不同系统和下游决策问题之间复用。我们考虑通过 Spec2Twin-Chain 实现这一过程的自动化,该框架将区块链数字孪生的构建形式化为一个双层优化问题。在上层,大语言模型(LLM)利用系统规范、行为证据以及来自已评估设计的反馈,提出并修正结构上可容许的架构。在下层,基于仿真的优化器在明确的目标和安全护栏约束下,校准依赖于架构的参数。两个层级迭代进行:下层中经过评估的候选方案被保留在全局档案中,用于引导上层的后续提案。我们开展了受控实验,包括孪生体校准、反馈驱动的恢复、压力分析、下游策略优化以及策略更新。结果表明,该框架能够构建行为准确的孪生体,通过迭代反馈改进初始设计,并复用已校准的孪生体以支持下游决策。
cs.AI / 124 / 2608.30051

Mitigating Over-Optimization in PRM-Guided Search in Mathematical Reasoning by Optimizing the Guide

通过优化引导器缓解数学推理中PRM引导搜索的过度优化问题
Joo, Taejong, Klabjan, Diego
Abstract
Process reward models (PRMs) provide dense step-level guidance for search-based reasoning, enabling inference-time compute to be allocated toward promising partial solutions. However, recent evidence suggests that PRM-guided search can over-optimize imperfect process rewards, pruning viable trajectories while expanding spurious ones. In this work, we theoretically show that directly leveraging PRM score is vulnerable to verifier noise through an extreme-value effect: non-viable prefixes become more likely to receive spuriously high scores as reasoning depth increase. Therefore, we formulate the PRM-guided search as a robust optimization problem over plausible reward perturbations, termed maximin PRM-guided search, leading to a training-free robust process supervision method that preserves promising alternatives when step-level scores are noisy. Maximin PRM-guided search mitigates this failure mode by reducing sensitivity to over-optimized PRM outliers. Without fine-tuning or online adaptation, maximin search consistently improves the PRM-guided search by 17-35\% on average, outperforming outcome- and step-level baselines in 14 out of 16 settings. Our source code is available at https://github.com/tjoo512/maximin-search.
Chinese Translation
过程奖励模型(Process Reward Models, PRMs)为基于搜索的推理提供了密集的步骤级引导,使推理时计算能够分配到有前景的部分解上。然而,近期证据表明,PRM引导的搜索可能会对不完善的过程奖励进行过度优化,即在剪枝可行轨迹的同时扩展虚假轨迹。在本工作中,我们通过理论上证明,直接利用PRM分数容易受到验证器噪声的影响,其机制是一种极值效应:随着推理深度的增加,不可行的前缀更有可能获得虚高的分数。因此,我们将PRM引导的搜索形式化为一个针对合理奖励扰动的鲁棒优化问题,称为maximin PRM引导搜索,由此得到一种无需训练的鲁棒过程监督方法,该方法在步骤级分数存在噪声时能够保留有前景的备选方案。Maximin PRM引导搜索通过降低对过度优化的PRM离群值的敏感性来缓解上述失效模式。在无需微调或在线自适应的情况下,maximin搜索平均将PRM引导搜索的性能持续提升17-35%,在16个设置中的14个上优于结果级和步骤级基线方法。我们的源代码可在 https://github.com/tjoo512/maximin-search 获取。
cs.AI / 125 / 2608.30056

Game-Agnostic Value Functions through Automatic JSON Feature Extraction

基于自动JSON特征提取的游戏无关价值函数
Nguyen, Dien, Perez-Liebana, Diego
Abstract
JSON Bag-of-Tokens (JSON-Bag) is a recently proposed method to generically represent game trajectories by tokenizing their JSON descriptions. We introduce JSON-Bag VF, a game-agnostic approach to training value functions for game-playing agents using JSON-Bag prototypes. We show that this approach can be enhanced with Random Forest-based feature selection and a method to select game-stage-specific features. We evaluate JSON-Bag VF with One-step-look-ahead (JSON-Bag OSLA) on six tabletop games over different combinations of prototype-tokenization and feature selections. JSON-Bag OSLA outperforms baseline OSLA agents in most games. Our analysis also shows that feature selection significantly improves JSON-Bag VF and that feature selection is the most important factor in JSON-Bag VF performance, over prototype-tokenization.
Chinese Translation
JSON词袋(JSON Bag-of-Tokens, JSON-Bag)是最近提出的一种通用表示游戏轨迹的方法,它通过将游戏的JSON描述进行词元化(tokenizing)来实现。我们提出了JSON-Bag VF,这是一种基于JSON-Bag原型、为游戏智能体训练价值函数的游戏无关方法。我们表明,该方法可以通过基于随机森林(Random Forest)的特征选择以及一种选择特定游戏阶段特征的方法得到增强。我们在六种桌面游戏上,针对原型-词元化与特征选择的不同组合,对采用一步前瞻(One-step-look-ahead, OSLA)的JSON-Bag VF(即JSON-Bag OSLA)进行了评估。JSON-Bag OSLA在大多数游戏中优于基线OSLA智能体。我们的分析还表明,特征选择能显著提升JSON-Bag VF的性能,并且与原型-词元化相比,特征选择是影响JSON-Bag VF性能的最重要因素。
cs.AI / 126 / 2608.30091

VERA: Authority-Preserving Edge Revocation for Federated AI-Agent Workflows

VERA:面向联邦AI智能体工作流的权限保持型边撤销机制
Liu, Lifei, Yu, Haoran, Jiang, Xiaochong
Abstract
Modern agent frameworks compose planners, tool agents, remote services, and shared specialists into runtime delegation graphs, but their revocation APIs still resemble token or subtree invalidation. When one delegation is withdrawn, the runtime must know which agents lose authority while independently authorized agents keep working. We study this authority consistency problem and introduce VERA (Verifiable Edge Revocation for Agents), a verifier-checkable revocation contract and API emitted by agent-runtime adapters as signed evidence. Under disjunctive authority, revoking edge e invalidates exactly T_intent(e,G) = reach(G) \ reach(G \ {e}), the agents whose every authorizing root path used e. Used as a contract, this target exposes two runtime failures: tree cascades over-revoke shared agents, while deployer-scoped cascades under-revoke cross-domain descendants. In a LangGraph framework-replt cells repeated 20 times yield 500compiled-framework traces and 2,000 valid signed delegation decisions; 13/25 cells contain runtime multi-parsharing and 8/25 contain cross-deployer shies 500/500 target proofs, preserves all320 alternate-parent shared-agent cases that tree cascade revokes, and rejects unauthorized signers and omission attacks. Baseline replay over 1,9that holder/node and tree-style targetscannot express this behavior. We further validate schema portability on A2A, AutoGen, and CrewAI artifacts: nine traces, including five executable Cregned delegation events that pass schema and signature checks.
Chinese Translation
现代智能体框架将规划器、工具智能体、远程服务和共享专家组合成运行时委托图,但其撤销API仍然类似于令牌或子树失效机制。当某一条委托被撤回时,运行时必须知道哪些智能体失去权限,而独立授权的智能体则继续工作。我们研究了这一权限一致性问题,并提出了VERA(Verifiable Edge Revocation for Agents,可验证的智能体边撤销机制),这是一种可由验证者检查的撤销契约和API,由智能体运行时适配器以签名证据的形式发出。在析取权限模型下,撤销边e恰好使 T_intent(e,G) = reach(G) \ reach(G \ {e}) 失效,即其所有授权根路径都经过e的智能体。作为一种契约,该目标暴露了两类运行时故障:树级联会过度撤销共享智能体,而部署者范围的级联则会遗漏撤销跨域后代。在LangGraph框架的实验中,20次重复的单元产生500个编译后的框架轨迹和2,000个有效的签名委托决策;25个单元中有13个包含运行时多方共享,8个包含跨部署者共享;系统实现了500/500的目标证明,保留了所有320个被树级联撤销的备选父节点共享智能体案例,并拒绝了未授权签名者和遗漏攻击。对1,900多个轨迹的基线重放表明,持有者/节点和树式目标无法表达这种行为。我们进一步在A2A、AutoGen和CrewAI工件上验证了模式的可移植性:九个轨迹,包括五个可执行的带有签名委托事件的案例,均通过了模式和签名检查。
cs.AI / 127 / 2608.30181

A.X K2 Technical Report

A.X K2 技术报告
Baek, Cheolseung, Arya, Dhammiko, Kim, Eunki, Song, Gun, Han, Gyoungeun, Yang, Hyunho, Eun, Hyunjun, Kim, Jin, Park, Junyoung, Wee, Juyun, Hong, Minki, Park, Minkyung, Kim, Minsang, Kang, Minsoo, Kim, SaeRom, Kim, Sangjin, Lee, Sangyeol, Lee, Seojin, Jo, Seokhwan, Hong, Seokyoung, Choi, Seongho, Cho, Seonghye, Ok, Seongmin, Sek, Sereimony, Cho, Seungmo, Kim, Seungsik, Kim, Singon, Park, Sohee, Park, Sooyeon, Yi, Subin, Yoon, Sungbin, Lee, Sungeun, Cheon, Sung Jun, Kim, Sungwan, Lee, Sunwoo, Kim, Tae Yoon, Jang, Wonbeom, Ra, Yohan, Han, Yong-jin, Kim, Youngjin, Kim, Youngrang, Kang, Yujin, Lee, Yujin
Abstract
We introduce A.X K2, a 688B-parameter Mixture-of-Experts (MoE) language model trained from scratch as a high-performance foundation for \emph{agentic} applications. Trained on approximately 8.5T tokens---fewer than its predecessor, A.X K1---on a smaller but higher-quality mixture with substantially expanded agentic and software-engineering data, it nonetheless improves over A.X K1 across the board, by over 30 percentage points on some benchmarks, reflecting large gains in token efficiency. To support long contexts efficiently, we introduce Sparse Gated Attention (SGA), which combines sparse attention with gated attention, and adopt Gated Norm (GN) to stabilize large-scale training. SGA is trained natively at 128K through a \emph{sparse} indexer warmup that optimizes the indexer against its own sparse top-$k$ selection rather than the dense attention distribution, making adaptation markedly cheaper: each query reads only 2,048 positions, yet long-context quality is unchanged and A.X K2 scores 94.6 on RULER out to 256K. The outlier suppression of GN in turn keeps 4-bit NVFP4 serving within one point of FP8 accuracy. A simple yet effective Think-Fusion recipe further lets users switch between thinking and non-thinking modes within a single unified model. Extensive evaluations show that A.X K2 performs competitively against strong open-weight baselines, matching or exceeding them on math and Korean-language benchmarks.
Chinese Translation
我们介绍了 A.X K2,一个从零开始训练的 688B 参数混合专家(Mixture-of-Experts, MoE)语言模型,旨在作为面向智能体(agentic)应用的高性能基础模型。该模型在大约 8.5T 词元上进行训练——少于其前代模型 A.X K1——使用了规模更小但质量更高的数据混合,并大幅扩充了智能体和软件工程相关数据,但仍在各项指标上全面超越 A.X K1,在部分基准上提升超过 30 个百分点,反映出词元效率的显著提升。为高效支持长上下文,我们提出了稀疏门控注意力(Sparse Gated Attention, SGA),将稀疏注意力与门控注意力相结合,并采用门控归一化(Gated Norm, GN)以稳定大规模训练。SGA 通过一种稀疏索引器预热(sparse indexer warmup)方法在 128K 上下文长度下进行原生训练,该方法针对索引器自身的稀疏 top-k 选择而非稠密注意力分布进行优化,使适配成本大幅降低:每个查询仅需读取 2,048 个位置,而长上下文质量不受影响,A.X K2 在 RULER 基准上直至 256K 长度仍取得 94.6 的分数。GN 的离群值抑制能力进而使 4 位 NVFP4 推理的精度与 FP8 的差距控制在 1 分以内。一套简单而有效的思考-融合(Think-Fusion)方案进一步让用户能在同一个统一模型内切换思考与非思考模式。大量评估表明,A.X K2 与强大的开源权重基线相比具有竞争力,在数学和韩语基准上达到或超越这些基线。
cs.AI / 128 / 2608.30192

FaVOR: LLM-Based Agentic Framework for Factor Mining via Empirical Validation

FaVOR:基于经验验证的LLM智能体因子挖掘框架
Kim, Hyeonjin, Kim, Minseok, Jung, Seunghyeon, Pyo, Sujin, Jang, Huisu, Lee, Woojin
Abstract
Traditional finance relies on experts to hand-craft factors through a principled process grounded in economic rationale. Recent LLM-based multi-agent systems have automated this process, scaling factor mining far beyond manual effort. However, these automated approaches optimize directly for returns and rarely check whether a generated factor still expresses the economic hypothesis that motivated it. We identify this inconsistency between mathematical form and economic meaning as a structural failure mode of return-oriented automation. The resulting factors blur the line between real signals and spurious correlations and break down across regime shifts. We propose FaVOR (Factor Validation through Observable Reasoning), an agentic framework that restructures factor mining around hypothesis-level evidence rather than return outcomes. In place of the standard hypothesis-to-formula leap, FaVOR enforces a three-stage consistency loop tying mathematical form to economic rationale throughout. (1) Decomposition splits a broad economic hypothesis into independent observable conditions. (2) Validation checks whether each factor reflects its intended condition. (3) Integration merges them into a composite whose structure remains interpretable. On the CSI 500 and S&P 500 in 2025, FaVOR outperforms existing baselines while remaining effective across regimes. FaVOR shows that hypothesis-grounded factor discovery produces signals that are interpretable by construction, regime-robust, and economically faithful. The code is available at https://github.com/damilab/FaVOR.
Chinese Translation
传统金融依赖专家基于经济原理,通过规范流程手工构建因子。近期基于大语言模型(LLM)的多智能体系统已实现该过程的自动化,使因子挖掘的规模远超人工。然而,这些自动化方法直接以收益为目标进行优化,很少检验生成的因子是否仍能表达激发其设计的经济假说。我们将这种数学形式与经济含义之间的不一致性识别为收益导向自动化的一个结构性失效模式。由此产生的因子模糊了真实信号与虚假相关之间的界限,并在市场状态切换时失效。我们提出FaVOR(Factor Validation through Observable Reasoning,基于可观测推理的因子验证),这是一个智能体框架,它将因子挖掘重构为围绕假说层面的证据而非收益结果展开。FaVOR以三阶段一致性循环取代标准的从假说到公式的直接跃迁,使数学形式与经济原理始终保持关联:(1) 分解(Decomposition)将宽泛的经济假说拆分为相互独立的可观测条件;(2) 验证(Validation)检验每个因子是否反映了其预期条件;(3) 整合(Integration)将各条件合并为一个结构保持可解释性的复合因子。在2025年的中证500(CSI 500)与标普500(S&P 500)数据上,FaVOR优于现有基线方法,并在不同市场状态下保持有效。FaVOR表明,基于假说的因子发现能够产生天然可解释、对市场状态切换鲁棒且经济含义忠实的信号。代码已发布于 https://github.com/damilab/FaVOR。
cs.AI / 129 / 2608.30214

SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature

SPARK:基于骨架引导的大规模科学文献推理合成方法
Li, Yu, Li, Wei, Gao, Xin, Sun, Mengyuan, Wang, Xiaoyang, Pei, Qizhi, Wu, Lijun
Abstract
Scientific reasoning remains challenging for open-source models, largely due to the lack of high-quality scientific reasoning data. Existing datasets are often dominated by factual recall or formulaic problem solving, with limited emphasis on mechanism understanding, evidence-grounded reasoning, and hypothesis evaluation. To address this, we introduce SPARK (Scientific Paper Abstracted Reasoning sKeleton), a paper-oriented synthesis framework built on Sci-Base, a large-scale corpus of research papers spanning 10 scientific disciplines. Instead of directly converting papers into question-answer pairs, SPARK treats the claim-evidence-derivation structure of a paper as the fundamental unit of reasoning synthesis. Specifically, SPARK (1) distills each paper into a compact reasoning skeleton capturing its central claims and supporting evidence, enabling self-contained question generation, and (2) synthesizes reasoning tasks from four scientific perspectives: mechanistic reasoning, hypothesis falsification, quantitative derivation, and boundary calibration. A final consistency verification stage further removes unsupported or contradictory outputs. Using this framework, we construct Spark-234K, a scientific reasoning dataset with substantially higher difficulty and diversity than existing resources. Experiments show that Spark-234K consistently outperforms existing scientific reasoning datasets while achieving stronger performance with significantly fewer training samples.
Chinese Translation
对开源模型而言,科学推理仍然极具挑战性,这在很大程度上归因于高质量科学推理数据的缺乏。现有数据集往往以事实记忆或程式化的问题求解为主,对机制理解、基于证据的推理以及假设评估的关注十分有限。为解决这一问题,我们提出了SPARK(Scientific Paper Abstracted Reasoning sKeleton,科学论文抽象推理骨架),一个面向论文的推理合成框架,其构建于Sci-Base之上——这是一个覆盖10个科学学科的大规模研究论文语料库。SPARK并非直接将论文转换为问答对,而是将论文的“主张—证据—推导”结构作为推理合成的基本单元。具体而言,SPARK(1)将每篇论文提炼为紧凑的推理骨架,捕获其核心主张与支持证据,从而支持自成一体的(self-contained)问题生成;(2)从四个科学视角合成推理任务:机制推理、假设证伪、定量推导和边界校准。最后,一致性验证阶段进一步过滤掉缺乏依据或自相矛盾的输出。基于该框架,我们构建了Spark-234K——一个在难度和多样性上均显著超越现有资源的科学推理数据集。实验表明,Spark-234K持续优于现有科学推理数据集,同时以明显更少的训练样本取得了更强的性能。
cs.AI / 130 / 2608.30226

LaMoC: Loss-Aware Modular Compression for LLMs

LaMoC:面向大语言模型的损失感知模块化压缩方法
Odema, Mohanad, Song, Jacob
Abstract
Modular compression has enabled considerable parameter reduction in LLMs while preserving strong language understanding and downstream task accuracy. However, existing joint modular compression methods primarily rely on activation statistics, leaving loss-sensitivity information and its module-level characterization underexplored. We investigate addressing this gap with LaMoC, a loss-aware modular compression methodology that blends activation and Empirical Fisher statistics through gradient-error alignment. LaMoC improves joint compression by selecting compression statistics that better align local module reconstruction error with the downstream loss. Our contributions are three-fold: (1) We characterize the Empirical Fisher as a module-level loss-aware proxy that can be blended with the activation statistics required for compression. (2) We reformulate joint modular compression as a two-tiered optimization problem that minimizes module reconstruction error while tuning the activation and gradient information blending rate. (3) We implement an empirically driven methodology with statistical validation to solve the resulting compression problem. We evaluate LaMoC across four model families spanning eight models. On the 4-8B models, LaMoC achieves an average 2.5% reduction in perplexity and a 1% relative improvement in task accuracy over state-of-the-art modular compression methods.
Chinese Translation
模块化压缩方法已使大语言模型(LLM)在保持强大的语言理解能力和下游任务准确率的同时,实现了可观的参数削减。然而,现有的联合模块化压缩方法主要依赖激活统计信息,对损失敏感性信息及其模块层面的刻画仍缺乏深入探索。我们提出LaMoC(Loss-Aware Modular Compression,损失感知模块化压缩)方法来填补这一空白,该方法通过梯度-误差对齐将激活统计信息与经验Fisher(Empirical Fisher)统计信息相融合。LaMoC通过选择能够更好地将局部模块重建误差与下游损失对齐的压缩统计信息,从而改进联合压缩。我们的贡献有三个方面:(1)我们将经验Fisher刻画为一种模块层面的损失感知代理,可将其与压缩所需的激活统计信息相融合;(2)我们将联合模块化压缩重新表述为一个双层优化问题,在最小化模块重建误差的同时调整激活信息与梯度信息的融合速率;(3)我们提出了一种由实证驱动并经过统计验证的方法来解决由此产生的压缩问题。我们在四个模型家族共八个模型上对LaMoC进行了评估。在4-8B规模的模型上,与最先进的模块化压缩方法相比,LaMoC平均将困惑度降低了2.5%,任务准确率相对提升了1%。
cs.AI / 131 / 2608.30230

Rethinking the Test-Time Prompt Tuning Objective from the Perspective of Calibration

从校准视角重新思考测试时提示调优目标
Choi, Jungwon, Jang, Hyeonseo, Lee, Kibok, Kim, Eunwoo
Abstract
Test-time prompt tuning (TPT) has emerged as a powerful paradigm, refining prompts for each test sample via entropy minimization (EM) over multiple augmented views. However, we identify a limitation in the standard EM-based adaptation: it inherently drives the model toward overconfident predictions disregarding sample-specific uncertainty, leading to significant calibration degradation. To address these limitations, we propose a new objective that replaces the conventional EM loss by aligning the original-view prediction with a target distribution derived from augmented views via cross-entropy, while adversarially incorporating the entropy of the target distribution to capture sample-specific uncertainty. Furthermore, to better construct this target distribution, we apply confidence-aware temperature scaling to each augmented-view prediction according to its confidence, sharpening confident predictions while softening uncertain ones. This formulation allows the model to increase confidence only when the target distribution is reliable, while preserving uncertainty when it reflects ambiguous or conflicting augmented-view predictions. Extensive experiments across diverse benchmarks demonstrate that our approach not only achieves state-of-the-art accuracy but also significantly improves model calibration.
Chinese Translation
测试时提示调优(Test-time Prompt Tuning, TPT)已成为一种强大的范式,它通过对多个增强视图进行熵最小化(Entropy Minimization, EM)来为每个测试样本优化提示。然而,我们发现基于标准EM的自适应方法存在一个局限性:它本质上会使模型产生过度自信的预测,而忽略了样本特定的不确定性,从而导致显著的校准性能下降。为了解决这些局限性,我们提出了一种新的目标函数,用交叉熵将原始视图的预测与从增强视图导出的目标分布进行对齐,以取代传统的EM损失,同时以对抗方式引入目标分布的熵来捕捉样本特定的不确定性。此外,为了更好地构建该目标分布,我们根据每个增强视图预测的置信度对其应用置信度感知的温度缩放,使置信度高的预测更尖锐,而使不确定的预测更平滑。这种构造使得模型仅当目标分布可靠时才提高置信度,而当目标分布反映模糊或相互冲突的增强视图预测时,则保持不确定性。在多个基准数据集上的大量实验表明,我们的方法不仅实现了最先进的准确率,还显著改善了模型的校准性能。
cs.AI / 132 / 2608.30234

CoLa-ICD: A Knowledge-Enhanced Framework for Long-Tail Automated Medical Coding

CoLa-ICD:一个面向长尾自动化医学编码的知识增强框架
Cheng, Yihang, Liesaputra, Veronica, Trotman, Andrew
Abstract
Automatic medical coding assigns ICD codes to clinical notes, but it remains challenging due to long documents, imbalanced label distributions, and diverse terms. These challenges are especially severe for rare codes, which have limited training instances and are easily confused with semantically similar labels. We introduce CoLa-ICD, a knowledge-enhanced framework for long-tail prediction. CoLa-ICD enriches ICD labels with external terms, models dependencies among related codes, and learns stronger alignment between label semantics and clinical evidence for long-tail prediction. Experiments show that CoLa-ICD improves long-tail prediction with larger gains in larger and sparser label spaces and achieves state-of-the-art performance in AUC, F1, and P@k. Our code is available at https://github.com/youwillbethebest/Cola-ICD.
Chinese Translation
自动化医学编码的任务是为临床记录分配ICD编码,但由于文档篇幅长、标签分布不均衡以及术语多样,该任务仍然具有挑战性。这些挑战对于罕见编码尤为严重,因为罕见编码的训练样本有限,且容易与语义相似的标签混淆。我们提出了CoLa-ICD,一个面向长尾预测的知识增强框架。CoLa-ICD通过外部术语丰富ICD标签信息,建模相关编码之间的依赖关系,并学习标签语义与临床证据之间更强的对齐关系,以实现长尾预测。实验表明,CoLa-ICD提升了长尾预测性能,且在更大、更稀疏的标签空间中 gains 越显著,并在AUC、F1和P@k指标上达到了最先进的性能。我们的代码已发布于 https://github.com/youwillbethebest/Cola-ICD。
cs.AI / 133 / 2608.30235

LLM-Based Knowledge Graph Completion Combining Discrete Structural Coding with Similar Entity Information

结合离散结构编码与相似实体信息的基于大语言模型的知识图谱补全
Wang, Jiaqi, Lin, Dongying, Yang, Yang, Liu, Yinan, Wang, Bin, Yang, Xiaochun
Abstract
Knowledge graph completion requires models to use both textual descriptions and relational structure. Existing LLM-based methods either encode KG structure as discrete tokens or refine a restricted set of candidate entities, and these two directions have largely been studied separately. We propose CoSC for LLM-based KGC, which combines discrete structural coding with similar entity information. Specifically, an LLM generates an initial candidate entity ranking from discrete structural codes, after which information from entities with structures similar to that of the query entity refines the ranking. Experiments on FB15k-237 show that CoSC outperforms existing baselines on MRR and Hits@10 while remaining competitive on Hits@1.
Chinese Translation
知识图谱补全要求模型同时利用文本描述和关系结构。现有基于大语言模型(LLM)的方法要么将知识图谱结构编码为离散标记,要么对受限的候选实体集进行精化,而这两个方向在很大程度上是被分开研究的。我们提出了用于基于LLM的知识图谱补全的CoSC方法,该方法将离散结构编码与相似实体信息相结合。具体而言,LLM首先根据离散结构编码生成初始候选实体排序,随后利用与查询实体结构相似的实体信息对该排序进行精化。在FB15k-237数据集上的实验表明,CoSC在MRR和Hits@10指标上优于现有基线方法,同时在Hits@1上保持竞争力。
cs.AI / 134 / 2608.30250

Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs

利用非推理型大语言模型从自然语言生成工作流有向无环图(DAG)
Iyer, Anand, Khetharpal, Bhanu, Upadhya, Srinivas, Rajagopal, Ramkumar
Abstract
This paper addresses the problem of translating natural-language routing rules written by business administrators into executable workflow graphs for enterprise contact centers. Each target is a directed acyclic graph (DAG) of conditional actions with parallel branches, hit-first fallback chains, and per-branch Boolean predicates, encoded in the JSON dialect of a commercial routing platform. We show that neuro-symbolic decomposition enables lower-cost, non-reasoning large language models to generate complex workflow DAGs at production-relevant quality without expensive extended-reasoning models. Our central diagnostic is an emission-density bottleneck: on a 635-rule benchmark of manufactured synthetic data, models select the correct graph nodes with high accuracy but increasingly misconfigure attributes and Boolean grouping as the number of interdependent nodes emitted in one pass grows. We therefore move combinatorial graph construction from the model into a deterministic compiler driven by a compact intermediate representation, with a learned registry-selection front end that focuses generation on relevant vocabulary. Across four models, the full system reaches approximately 89% LLM-judge validity, approximately 90% exact-match condition accuracy, and 99-100% valid JSON while using roughly half the per-rule prompt tokens of a monolithic prompt. On GPT-5.3-chat, the method improves judge validity by 24 percentage points and achieves statistical equivalence to a reasoning model's out-of-the-box quality, although an approximately 8-point frontier gap remains. We also present a deployment path and transferable lessons for structured-generation applications.
Chinese Translation
本文研究如何将企业呼叫中心业务管理员撰写的自然语言路由规则翻译为可执行的工作流图。每个目标是一个由条件动作组成的有向无环图(DAG),包含并行分支、命中优先的回退链以及每个分支上的布尔谓词,并以某商业路由平台的 JSON 方言进行编码。我们证明,神经符号分解方法能够使成本更低的非推理型大语言模型生成达到生产级质量的复杂工作流 DAG,而无需昂贵的扩展推理模型。我们的核心诊断是一个输出密度瓶颈:在一个包含635条人工合成规则的基准数据集上,模型能够以较高准确率选择正确的图节点,但随着单次输出中相互依赖节点数量的增加,模型在属性配置和布尔分组上 increasingly 出现错误。因此,我们将组合式图构建从模型转移到由紧凑中间表示驱动的确定性编译器中,并配备一个通过学习得到的注册表选择前端,将生成过程聚焦于相关词汇。在四个模型上,完整系统达到约89%的 LLM 评审有效性、约90%的条件精确匹配准确率,以及99–100%的有效 JSON 输出,同时每条规则的提示词令牌用量约为单体提示方法的一半。在 GPT-5.3-chat 上,该方法将评审有效性提升24个百分点,并达到与开箱即用的推理模型统计上等价的质量,尽管仍存在约8个百分点的差距。我们还给出了结构化生成应用的部署路径和可迁移的经验教训。
cs.AI / 135 / 2608.30277

SimCRAFT: Distilling Remote Sensing Agents via Synthetic Trajectories and Contextual Retrieval-Augmented Fine-Tuning

SimCRAFT:基于合成轨迹与上下文检索增强微调的遥感智能体蒸馏方法
Wang, Haoran, Yao, Jing, Yang, Xu, Wang, Zeqing, Zhang, Yang, Ghamisi, Pedram, Chen, Zhengchao
Abstract
The unprecedented surge in Earth observation data volume and diversity has exposed a critical bottleneck for traditional manual workflows, catalyzing the emergence of Remote Sensing (RS) Agents. However, the practical deployment of these advanced agents is severely hindered by their heavy reliance on large-scale general-purpose LLMs, which lack deep domain expertise and impose prohibitive infrastructure demands. To resolve this, we propose SimCRAFT, a model-agnostic framework that distills sophisticated RS orchestration capabilities into a compact 7B-scale model. Addressing data scarcity, we first pair a multiagent synthesis engine with a Mock Execution Engine that checks schema correctness, inter-tool dependencies, and sensor/tool compatibility, producing SimRS-14k, a large-scale, constraint-validated workflow planning corpus. Second, we propose Contextual Retrieval-Augmented Fine-Tuning (CRAFT) that finetunes the model to reason analogically by adapting retrieved Standard Operating Procedures to novel queries under a noise-robust objective, generalizing RAFT to multi-step RS workflow planning without mechanical copying. Extensive experiments demonstrate that SimCRAFT-7B significantly outperforms openweights LLMs and rivals advanced closedsource models and specialized RS agents, while reproducing across three 7B backbones. This work contributes a competitive open-weights baseline for lightweight RS intelligence, enabling efficient autonomous deployment under resource-constrained or resource-conserving conditions.
Chinese Translation
对地观测数据量与多样性的空前激增,暴露了传统人工工作流程的关键瓶颈,催生了遥感智能体的兴起。然而,这些先进智能体的实际部署受到严重阻碍,因为它们高度依赖大规模通用大语言模型(LLM),而这些模型既缺乏深入的领域专业知识,又带来高昂的基础设施成本。为解决这一问题,我们提出SimCRAFT,一个与具体模型无关的框架,可将复杂的遥感编排能力蒸馏到一个紧凑的7B规模模型中。针对数据稀缺问题,我们首先将多智能体合成引擎与一个模拟执行引擎相结合,该引擎检查模式正确性、工具间依赖关系以及传感器与工具的兼容性,从而构建出SimRS-14k——一个大规模、经约束验证的工作流规划语料库。其次,我们提出上下文检索增强微调(CRAFT),在抗噪目标的约束下,通过将检索到的标准操作规程(SOP)适配到新查询,使模型学会类比推理,从而将RAFT方法泛化到多步遥感工作流规划,避免了机械复制。大量实验表明,SimCRAFT-7B显著优于开源权重的大语言模型,并可媲美先进的闭源模型和专用遥感智能体,且该结果在三个7B骨干模型上均可复现。这项工作为轻量级遥感智能提供了一个具有竞争力的开源权重基线,能够在资源受限或资源节约条件下实现高效的自主部署。
cs.AI / 136 / 2608.30322

Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

是无知还是无能?为LLM智能体构建知识门控的可验证任务
Tian, Hanlin, Li, Minhao, Mi, Yu, Zhu, Sihan, Yang, Zhao, Wang, Yuxiang, Zhu, Hongquan, Hu, Qiufei
Abstract
Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks rarely control whether an agent has access to those conventions. We introduce a knowledge-gated task-construction protocol that separates a task instruction from a compact artefact containing private conventions, reference tables, and utility operators. Construction-time provenance, byte-identical task instructions across the provided- and withheld-artefact conditions, leak audits, and executable witnesses make dependence on the artefact explicit and testable. Across fifteen calibration tasks, one frontier agent configuration achieves a 68.0% pass rate with the artefact and 0% without it; on one task, a plausible but incorrect artefact also yields 0% across five trials. Deterministic solvers and rule corpora provide exact ground truth for structured tasks, while named criterion-level rubrics support outputs that cannot be checked by a single executable oracle. A configuration-relative calibration screen retains seven tasks satisfying our five-trial empirical knowledge-gating screen. These experiments validate the behavior of the construction protocol; they do not establish that the retained tasks improve post-training. We publicly release part of the task suite and supporting tooling at https://github.com/DatagridsAI/Knowledge-Gated-Task-Construction.
Chinese Translation
专业智能体任务通常依赖于公共语料库中缺失的惯例,然而现有基准很少控制智能体是否能获取这些惯例。我们提出了一种知识门控任务构建协议,将任务指令与一个包含私有惯例、参考表和实用工具算子的紧凑工件分离。构建时的来源记录、在提供工件与扣留工件两种条件下字节完全一致的任务指令、泄漏审计以及可执行见证,使得对工件的依赖变得明确且可检验。在十五个校准任务中,一个前沿智能体配置在提供工件时达到68.0%的通过率,而在没有工件时通过率为0%;在其中一项任务上,使用一个貌似合理但错误的工件在五次试验中通过率同样为0%。确定性求解器和规则语料库为结构化任务提供了精确的真实标准,而命名的准则级评分细则则支持无法通过单一可执行预言机检验的输出。基于配置相对化的校准筛选保留了七项满足我们五次试验经验性知识门控筛选标准的任务。这些实验验证了构建协议的行为,但并未证明所保留的任务能改善后训练效果。我们公开发布了部分任务套件及配套工具,网址为 https://github.com/DatagridsAI/Knowledge-Gated-Task-Construction。
cs.AI / 137 / 2608.30345

Answer Probing-Guided Search for Diverse Solution Exploration of LLMs

答案探测引导的搜索:面向大语言模型多样化解决方案探索
Fang, Yi, Shen, Que, Li, Chengpeng, Deng, Boyi, Shi, Wei, Wang, Wenjie, Feng, Fuli, Xu, Fengli, Liu, Dayiheng
Abstract
Generating multiple diverse and high-quality solutions is valuable for many applications, such as code-test generation and drug discovery. However, Large Language Models (LLMs) tend to converge on a single high-confidence solution during inference, limiting exploration of alternative valid solution paths. Existing test-time methods promote diversity through tree-like search and prune semantically similar branches using response-level semantic embeddings. However, we find that such embeddings are easily confounded by linguistic and stylistic similarities, making it difficult to distinguish genuinely distinct solution paths. To address this, we introduce Answer Probing, which probes the potential answer an LLM would reach from an intermediate reasoning path. We demonstrate that the hidden states of probed answers more effectively differentiate distinct solution paths than semantic embeddings, and the perplexity of probed answers serves as a practical proxy for reasoning correctness. Based on these findings, we propose Answer Probing-Guided Tree Search (APTS), which guides the tree search by the probed answers' hidden state similarity and perplexity. Experiments on three reasoning tasks across two LLMs show that APTS consistently enhances solution diversity, demonstrating its effectiveness and robustness.
Chinese Translation
生成多个多样且高质量的解决方案对许多应用具有重要价值,例如代码测试生成和药物发现。然而,大语言模型(LLM)在推理过程中往往收敛于单一的高置信度解决方案,限制了对其他有效解决路径的探索。现有的测试时方法通过树状搜索来促进多样性,并利用响应级语义嵌入剪枝语义相似的分支。然而,我们发现此类嵌入容易受到语言和风格相似性的干扰,难以区分真正不同的解决路径。为解决这一问题,我们提出了答案探测(Answer Probing)方法,用于探测大语言模型从中间推理路径可能到达的潜在答案。我们证明,被探测答案的隐藏状态比语义嵌入能更有效地区分不同的解决路径,且被探测答案的困惑度可作为推理正确性的实用代理指标。基于这些发现,我们提出答案探测引导树搜索(Answer Probing-Guided Tree Search, APTS),利用被探测答案的隐藏状态相似度和困惑度来引导树搜索。在两个大语言模型上的三项推理任务实验表明,APTS 能够持续提升解决方案的多样性,展示了其有效性和鲁棒性。
cs.AI / 138 / 2608.30352

Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration

Co-Annotator:面向年龄相关性黄斑变性的专家蒸馏ViT与VLM视觉及文档书写引导
Li, Ziheng "Leo", Freeman, Benjamin, Raman, Akshay, Rajkumar, Kavin Aravindhan, Fang, Xinxin, Srivastava, Rishabh, Feiner, Steven, Thakoor, Kaveri A.
Abstract
Clinical AI often optimizes predictive performance without engaging how clinicians decide where to look and what to write. We present Co-Annotator, which distills expert gaze and dictation into two guidance components: a gaze-aligned Vision Transformer producing fixation-aligned areas of interest (AOIs), and an ontology-bounded vision-language model (VLM) that pre-fills editable biomarker summaries for retinal optical coherence tomography (OCT). We first collect expert gaze and dictations (US1) to train the models, significantly improving diagnostic accuracy and biomarker generation. We then deploy the system with ophthalmology residents: a controlled resident study (US2) confirmed each modality is safe and independently beneficial, with AOI guidance producing lasting perceptual efficiency gains through post-guidance carryover and VLM guidance more than doubling biomarker documentation breadth. In a combined deployment across two academic institutions (US3), providing both modalities simultaneously produced efficiency gains that substantially exceeded either modality alone: correct diagnoses per minute increased by 40% and comment editing time fell by 67%, without compromising diagnostic accuracy. Notably, neither modality improved efficiency during guidance in US2, which makes the in-guidance efficiency gain under combined guidance in US3 the more striking result. Expert-distilled multimodal guidance can remove two distinct clinical workflow bottlenecks at once (visual search overhead and documentation burden) without compromising the diagnostic accuracy clinicians already achieve.
Chinese Translation
临床人工智能通常仅优化预测性能,而未关注临床医生如何决定看什么以及写什么。我们提出了Co-Annotator,它将专家的注视和口述蒸馏为两个引导组件:一个与注视对齐的视觉Transformer(Vision Transformer),生成与注视点对齐的兴趣区域(AOIs);以及一个受本体约束的视觉语言模型(VLM),为视网膜光学相干断层扫描(OCT)预填可编辑的生物标志物摘要。我们首先收集专家的注视与口述数据(US1)以训练模型,显著提升了诊断准确性和生物标志物生成质量。随后,我们在眼科住院医师中部署该系统:一项受控住院医师研究(US2)证实了每种引导模态均是安全且具有独立益处的,其中AOI引导通过引导后的延续效应产生了持久的感知效率提升,而VLM引导使生物标志物文档记录的广度提升超过一倍。在两所学术机构的联合部署(US3)中,同时提供两种模态所产生的效率提升显著超过了任一模态单独使用的效果:每分钟正确诊断数提高了40%,注释编辑时间下降了67%,且未损害诊断准确性。值得注意的是,在US2中,两种模态在引导期间均未提升效率,这使得US3中联合引导下引导期间的效率提升更为显著。专家蒸馏的多模态引导可以同时消除临床工作流程中两个不同的瓶颈(视觉搜索开销和文档书写负担),而不会损害临床医生已达到的诊断准确性。
cs.AI / 139 / 2608.30362

Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents

用户会察觉吗?针对使用工具的大语言模型代理的隐蔽间接提示注入攻击
Lee, Yunseok, Kim, Yunji, Lee, Woojin
Abstract
As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a serious threat. The standard metric, Attack Success Rate (ASR), counts whether an injection succeeds but ignores what the user notices in the agent's final response. Looking at successful injection traces, we find two distinct outcomes: the agent executes the injection while returning an otherwise normal response, or reports the injected action in its final response, giving the user a chance to notice. We call these covert and overt successes. From the user's perspective, we decompose ASR into the Covert Success Rate (CSR), counting successes leaving no trace in the final response, and the Overt Success Rate (OSR), counting successes the user can detect. To understand what drives the gap, we analyze successful trajectories and find that the agent's behavior after the injection separates covert from overt: covert traces hand control back to the user task before ending, while overt traces end at the attack itself. This split follows from the ReAct format, where the final response summarizes the most recent action. Building on this observation, we propose ICoA (Induced Covert Attack), an IPI attack designed to induce covert outcomes by steering the agent back to the user task after executing the injection. Across four target models on AgentDojo, ICoA achieves the highest CSR, with gains of 3.79-12.01 percentage points over the strongest baseline.
Chinese Translation
随着大语言模型(LLM)代理通过工具执行真实世界的操作,间接提示注入(Indirect Prompt Injection, IPI)已成为一种严重威胁。标准评估指标——攻击成功率(Attack Success Rate, ASR)——仅统计注入是否成功,却忽略了用户在代理最终回复中能察觉到什么。通过分析成功的注入轨迹,我们发现存在两种截然不同的结果:一种是代理执行注入后返回的正常回复,另一种是代理在其最终回复中报告了被注入的操作,从而让用户有机会察觉。我们将前者称为隐蔽成功(covert success),后者称为显性成功(overt success)。从用户视角出发,我们将 ASR 分解为隐蔽成功率(Covert Success Rate, CSR),即在最终回复中不留痕迹的成功次数,以及显性成功率(Overt Success Rate, OSR),即用户可以检测到的成功次数。为了理解造成这一差距的原因,我们分析了成功的轨迹,发现注入后代理的行为将隐蔽与显性区分开来:隐蔽轨迹在结束前将控制权交还给用户任务,而显性轨迹则在攻击本身处结束。这一划分源于 ReAct 格式,其中最终回复是对最近一次操作的总结。基于这一观察,我们提出了 ICoA(Induced Covert Attack,诱导式隐蔽攻击),一种通过在执行注入后将代理引导回用户任务来诱发隐蔽结果的 IPI 攻击。在 AgentDojo 基准上的四个目标模型中,ICoA 取得了最高的 CSR,相比最强基线提升了 3.79 至 12.01 个百分点。
cs.AI / 140 / 2608.30369

Augmenting Human Performance with an XR Agent Learning from Online Behavior and BCI Evidence

利用从在线行为与脑机接口证据中学习的XR智能体增强人类绩效
Li, Ziheng, He, Xichen, Chen, Haoyan, Zou, Charlie, Bai, Sheng, Yang, Benjamin, Wu, Mengyuan, Ledner, Jake, Cheng, Yi-Jie, Yamauchi, Akito, Turakhia, Dishita G, Feiner, Steven, Sajda, Paul
Abstract
We present OLIVE, a framework for adapting a foundation model to provide real-time assistance in temporally demanding, high-stakes, and dynamic tasks. We show that passive EEG, fused online with behavioral evidence, can meaningfully extend the number of targets users detect and engage beyond their unaided action bandwidth. OLIVE learns from both explicit behavioral signals (the targets the user shoots down in an XR first-person shooter game) and implicit physiological signals (fixation-locked EEG) to provide timely guidance, continuously adapting a frozen vision-language model's inference on which items are task-relevant by jointly estimating per-source reliability without manual labels or offline training. Through three user studies, including two live deployments of an assistive agent driven by OLIVE in XR, we show that OLIVE Pareto-dominates prior test-time adaptation frameworks, achieving the highest convergence rate at comparable convergence speed. Combining implicit physiological and explicit behavioral signals, the OLIVE agent produces the largest and most reliable within-session improvement to a user's ability to detect and engage targets, largely independent of the individual's skill. When the target switches silently, the agent that uses both behavioral and physiological signals reconverges significantly faster than the behavior-only agent (1.27 times faster on average, p = .008), restoring trustworthy guidance at the moment the task changes, precisely when reliable assistance matters most.
Chinese Translation
我们提出了OLIVE,一个将基础模型适配以在时间要求苛刻、高风险且动态的任务中提供实时辅助的框架。我们证明,被动式脑电(EEG)信号与行为证据在线融合后,能够显著扩展用户检测和处置目标的数量,超越其无辅助时的行动带宽。OLIVE同时从显式行为信号(用户在XR第一人称射击游戏中击落的目标)和隐性生理信号(注视锁定的EEG)中学习以提供及时引导,通过联合估计各来源信号的可靠性,持续调整一个冻结的视觉-语言模型对哪些物品与任务相关的推断,无需人工标注或离线训练。通过三项用户研究(包括两次由OLIVE驱动的辅助智能体在XR中的现场部署),我们证明OLIVE在Pareto意义上优于现有的测试时自适应框架,在相近的收敛速度下实现了最高的收敛率。通过结合隐性生理信号与显式行为信号,OLIVE智能体为用户检测和处置目标的能力带来了最大且最可靠的会话内提升,且在很大程度上与个体技能水平无关。当任务目标静默切换时,同时使用行为与生理信号的智能体重新收敛的速度显著快于仅使用行为信号的智能体(平均快1.27倍,p = .008),能够在任务发生变化的时刻恢复可信的引导,而这正是可靠辅助最为关键的时刻。
cs.AI / 141 / 2608.30396

Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation

将基础模型脚手架化为物理世界智能体,推动长时程导航的前沿
Lei, Zixing, Zhou, Gengze, Chen, Xiong-Hui, Zhang, Jiazhao, Huang, Yiyang, Yin, Hang, Yuan, Haoqi, Wu, Qi, Li, Weixin, Chen, Siheng
Abstract
Long-horizon physical-world agents must reason over distant goals while grounding decisions in reliable closed-loop behavior. Today's foundation models split these capabilities: vision-language models (VLMs) infer missing information and adapt high-level plans but remain brittle and inefficient at repeated navigation grounding, while navigation foundation models (NFMs) robustly execute semantic goals but operate as bounded episodes without persistent task-level reasoning. We introduce NavMCP, an agentic scaffolding framework that couples a VLM reasoning agent with an NFM executor for long-horizon exploration. The VLM decides what evidence to seek, where to search, and when to stop, while the NFM grounds each semantic sub-goal into closed-loop navigation. Three channels structure their collaboration: intent translates evidence needs into navigation calls, observation converts rollouts into source-grounded trajectory evidence, and memory accumulates findings, negative evidence, and unresolved goals across calls. This design turns isolated navigation rollouts into persistent embodied interaction without retraining either model. On Embodied Question Answering, NavMCP achieves state-of-the-art results on HM-EQA, MT-HM3D, and EXPRESS-Bench. Under matched agent and executor backbones, it outperforms an episodic interface by 14.9 percentage points on HM-EQA. On a Unitree Go2, NavMCP reaches 78.3% success, with its margin over the strongest baseline growing from 10 to 45 points as the task horizon increases. These results demonstrate the potential of scaffolding complementary foundation models into long-horizon physical-world agents.
Chinese Translation
长时程物理世界智能体必须既能对远距离目标进行推理,又能将决策落实到可靠的闭环行为中。当今的基础模型将这两种能力割裂开来:视觉-语言模型(VLM)能够推断缺失信息并调整高层计划,但在反复进行导航接地(grounding)时仍然脆弱且低效;而导航基础模型(NFM)虽能稳健地执行语义目标,却仅以有限回合的方式运行,缺乏持续的任务级推理。我们提出了 NavMCP,一个将 VLM 推理智能体与 NFM 执行器相结合的智能体脚手架框架,用于长时程探索。VLM 决定寻找什么证据、在哪里搜索以及何时停止,而 NFM 将每个语义子目标落地为闭环导航。三个通道构建了二者的协作:意图(intent)通道将证据需求转化为导航调用;观察(observation)通道将执行过程转化为有源可溯的轨迹证据;记忆(memory)通道在多次调用间积累发现、负面证据和未解决的目标。这一设计将孤立的导航执行过程转化为持续的具身交互,且无需对任何模型进行重新训练。在具身问答(Embodied Question Answering)任务上,NavMCP 在 HM-EQA、MT-HM3D 和 EXPRESS-Bench 上取得了最先进的结果。在相同的智能体与执行器主干下,它在 HM-EQA 上比回合式接口高出 14.9 个百分点。在 Unitree Go2 机器人上,NavMCP 达到了 78.3% 的成功率,且随着任务时程的增长,其相对于最强基线的优势从 10 分扩大到 45 分。这些结果证明了将互补的基础模型脚手架化为长时程物理世界智能体的潜力。
cs.AI / 142 / 2608.30405

Dense Clinical Contrasts Enhance Medical Knowledge Updating in Large Language Models

密集临床对比增强大语言模型中的医学知识更新
Huang, Yangmin, Quan, Shu, Geng, He, Ye, Xin, Du, Qianyun, He, Zhiyang, Hu, Jiaxue, Tao, Xiaodong
Abstract
Medical knowledge changes continually, making large language models vulnerable to relying on outdated yet clinically plausible information. We study whether the format of supervision affects medical knowledge updating under a matched training-budget setting. We introduce SEER-Bench, a temporally anchored oncology-staging benchmark curated from the latest versioned SEER Research Data release, and render identical medical update events from NCCN oncology guidelines into four supervision formats: EMQ, MSQ, FITB, and SAQ. Across SEER-Bench and HealthBench Professional, EMQ gives the most stable external transfer and retention among same-budget SFT variants. With EMQ supervision, the updated 4B model produces competitive results on temporally anchored oncology staging, reaching 64.8% answer accuracy and 59.6% rationale accuracy on SEER-Bench. Diagnostic analyses suggest that EMQ exposes denser clinical contrast signals while preserving discriminative representations with smaller movement from the base model. These results show that medical knowledge updating depends not only on the update algorithm, but also on how knowledge is structured as supervision.
Chinese Translation
医学知识不断变化,这使得大语言模型容易依赖过时但看似合理的临床信息。我们在训练预算相同的设置下,研究监督数据的格式是否会影响医学知识更新。我们构建了SEER-Bench,这是一个基于最新版本SEER研究数据发布构建的、具有时间锚定的肿瘤分期基准,并将NCCN肿瘤指南中相同的医学更新事件渲染为四种监督格式:EMQ、MSQ、FITB和SAQ。在SEER-Bench和HealthBench Professional上,EMQ在同等预算的SFT变体中提供了最稳定的外部迁移和知识保持能力。在EMQ监督下,更新后的4B模型在时间锚定的肿瘤分期任务上取得了具有竞争力的结果,在SEER-Bench上达到64.8%的答案准确率和59.6%的推理准确率。诊断性分析表明,EMQ能够提供更密集的临床对比信号,同时以更小的偏离基础模型的幅度保持判别性表示。这些结果表明,医学知识更新不仅取决于更新算法,还取决于知识如何被结构化为监督信号。
cs.AI / 143 / 2608.30413

DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark

DERELAB:利用生成式基准探究大语言模型中的可废止推理与确认偏差
Sadhu, Jayanta, Shahad, Sayem, Marino, Kenneth
Abstract
Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of newer evidence. Although recent studies have examined language-model behaviors in defeasible reasoning, the datasets have been static and lack wide coverage of non-monotonic reasoning categories. We introduce DeReLab, a generative framework that produces multi-turn belief-updating conversations from parameterized graph structures across default and inheritance reasoning, with formally verified ground truth at every turn, enabling controlled measurement of how models respond to confirming and disconfirming evidence. This controlled generation process creates a testbed for experimental designs that isolate specific reasoning demands. Applying this capability to the study of confirmation bias, we evaluate nine open and proprietary large language models and find that nearly all exhibit a systematic tendency to accept congruent evidence while resisting incongruent updates, with several models correctly identifying a weakening update yet failing to revise their conclusion. We believe our work and findings will facilitate future research on evaluating language models in defeasible reasoning.
Chinese Translation
可废止推理是一类从当前可信证据出发进行推断,但在引入新证据时可以撤回先前结论的推理方式。尽管近期已有研究考察了语言模型在可废止推理中的表现,但现有数据集大多是静态的,且对非单调推理类别的覆盖范围有限。我们提出了 DeReLab,一个生成式框架,它从参数化的图结构出发,生成围绕默认推理和继承推理的多轮信念更新对话,并在每一轮都具备经过形式化验证的标准答案,从而能够对模型响应确认性与否定性证据的方式进行可控测量。这一可控的生成过程为分离特定推理需求的实验设计提供了测试平台。我们将该能力应用于确认偏差的研究,评估了九个开源与闭源大语言模型,发现几乎所有模型都表现出一种系统性倾向:倾向于接受与自身信念一致的证据,同时抵制与之不一致的信念更新;其中若干模型能够正确识别出一次削弱性更新,却未能据此修正其结论。我们相信,本工作及其发现将促进未来在可废止推理领域对语言模型评估的研究。
cs.AI / 144 / 2608.30419

From Metaheuristics to Exact Methods: A CP-SAT Approach for Multi-Objective Healthcare Workforce Scheduling

从元启发式方法到精确方法:一种面向多目标医疗人员排班的CP-SAT方法
Patel, Vipul, Deodhar, Anirudh, Birru, Dagnachew
Abstract
Healthcare workforce scheduling is an NP-hard optimization problem requiring simultaneous satisfaction of labor regulations, coverage requirements, employee preferences and cost objectives. Existing approaches (genetic algorithms, integer programming, constraint programming) model 6-12 constraints at shift-level granularity and cannot guarantee regulatory compliance. They also lack support for multi-role, multi-skill heterogeneity, mandatory break scheduling with midpoint control, acuity-weighted workload equity, sub-shift granularity, inter-week stability, and cross-midnight shifts. This paper presents CP-SAT: a Constraint Programming formulation for multi-role, multi-skill healthcare scheduling. CP-SAT enforces 14 hard constraints guaranteeing zero regulatory violations, while optimizing 15 soft objectives via a unified weighted penalty function. Contributions include a shift-window decomposition enabling break scheduling with centrality control, acuity-weighted workload equity, multi-granularity resolution from 15 minutes to 1 day, inter-week stability, and grid-offset preprocessing mapping cross-midnight shifts into a single scheduling day without solver changes. CP-SAT is evaluated on 18 instances: five synthetic hospital units (10-33 nurses), 10 INRC-II benchmarks (5-80 nurses, up to 8-week horizons) and 3 NRP-23 compatible instances (10-25 nurses) with cross-midnight Night shifts. Results: zero hard-constraint violations across all 18 instances by construction; proven optimality on INRC-II n005w4 (objective 118, gap 0.0%, 104s); feasible schedules scaling to 179,800 variables and 351,425 constraints (80 nurses); service quality improved 50-67% over MOGA; and model size scaling near-linearly at approximately 4,400 variables per employee. The formulation enforces 29 total constraints (14 hard, 15 soft), nearly three times the industry average.
Chinese Translation
医疗人员排班是一个NP难优化问题,需要同时满足劳动法规、覆盖需求、员工偏好和成本目标。现有方法(遗传算法、整数规划、约束规划)仅在班次级粒度上建模6-12个约束条件,无法保证法规合规性,且缺乏对多角色、多技能异质性、带中点控制的强制休息调度、基于病患严重度加权的工作量公平性、子班次粒度、周间稳定性以及跨午夜班次的支持。本文提出CP-SAT:一种面向多角色、多技能医疗排班的约束规划(Constraint Programming)建模方法。CP-SAT强制执行14个硬约束以保证零法规违规,同时通过统一的加权惩罚函数优化15个软目标。主要贡献包括:一种班次窗口分解方法,支持带中心性控制的休息调度;基于病患严重度加权的工作量公平性;从15分钟到1天的多粒度分辨率;周间稳定性;以及网格偏移预处理,可将跨午夜班次映射到单个调度日而无需修改求解器。CP-SAT在18个实例上进行了评估:5个合成医院科室(10-33名护士)、10个INRC-II基准实例(5-80名护士,时间跨度长达8周)以及3个与NRP-23兼容的实例(10-25名护士,含跨午夜夜班)。结果表明:所有18个实例在构造上均实现零硬约束违规;在INRC-II n005w4上证明了最优性(目标值118,间隔0.0%,耗时104秒);可行排班可扩展至179,800个变量和351,425个约束(80名护士);服务质量相比MOGA提升50-67%;模型规模接近线性扩展,约为每名员工4,400个变量。该模型共强制执行29个约束(14个硬约束,15个软约束),接近行业平均水平的三倍。
cs.AI / 145 / 2608.30429

EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents

EvoSkill 注入:针对自进化智能体中自主技能生成与进化的红队测试
Kim, Doyun, Kim, Chanwoo, Eo, Sugyeong, Yoon, Yeo-Chan, Park, Chanjun
Abstract
LLM-based agent systems increasingly adopt skill-based architectures to reduce repetitive reasoning costs and improve stable, efficient task execution. Recent studies propose self-evolving agents that autonomously generate, refine, and reuse skills from past experiences to enable continuous capability evolution. However, autonomous skill evolution introduces a new attack surface in which malicious capabilities are generated, stored, and reused as legitimate skills. In this paper, we define EvoSkill Injection as a threat model targeting the autonomous skill generation and evolution pipeline of self-evolving agents. We further propose SARGE (Red-teaming Autonomous Skill Generation and Evolution in self-evolving agents), a red-teaming framework for evaluating this threat model through iterative generation, escalation, and reinforcement interactions. To support our framework, we construct EvoSkillBench, a benchmark dataset of malicious interaction trajectories for inducing malicious skill formation in self-evolving agents, and introduce EvoSkillSafetyBench, a post-attack benchmark for evaluating whether injected malicious skills are subsequently retrieved and activated as harmful behaviors. Our evaluation shows that SARGE induces malicious skill formation and that injected skills are persistently stored and repeatedly activated, highlighting the risk of persistent capability corruption.
Chinese Translation
基于大语言模型(LLM)的智能体系统日益采用基于技能的架构,以减少重复推理的开销并实现稳定、高效的任务执行。近期研究提出了自进化智能体(self-evolving agents),其能够从过往经验中自主生成、优化和复用技能,从而实现能力的持续进化。然而,自主技能进化引入了新的攻击面:恶意能力可以被生成、存储并作为合法技能被复用。在本文中,我们将 EvoSkill Injection 定义为一种针对自进化智能体自主技能生成与进化流程的威胁模型。我们进一步提出 SARGE(面向自进化智能体中自主技能生成与进化的红队测试框架),一个通过迭代生成、升级和强化交互来评估该威胁模型的红队测试框架。为支持该框架,我们构建了 EvoSkillBench——一个用于诱导自进化智能体形成恶意技能的恶意交互轨迹基准数据集,并引入 EvoSkillSafetyBench——一个攻击后基准,用于评估被注入的恶意技能是否会在后续被检索并激活为有害行为。我们的评估表明,SARGE 能够诱导恶意技能的形成,且被注入的技能会被持久存储并被反复激活,凸显了能力遭受持久性破坏的风险。
cs.AI / 146 / 2608.30466

CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target

CHASE:当排序成为唯一目标时,内容生态系统如何被重塑
Gao, Qianwen, Su, Zichang, Hou, Yiwen, Kumar, Arlen, Palkhouski, Leanid
Abstract
Generative Engine Optimization (GEO) is increasingly used to improve content visibility in LLM-based retrieval systems, yet its population-level effects under repeated optimization remain poorly understood. We introduce Content Homogenization under rAnking Signal Exploitation (CHASE), a controlled simulation framework for studying how content ecosystems are reshaped when creators repeatedly adapt documents to an LLM ranking signal. We use ranking as a proxy for source visibility and validate this abstraction against citations in grounded generated responses, obtaining a rank-citation AUC of 0.853 $\pm$ 0.093 across six domains. CHASE then iterates ranking, feature discrimination, rewriting, and evaluation over 20 rounds across different domains. Quality-ranking alignment decreases in all six domains: from R0 to R20, the change in Spearman's rho ranges from -0.107 to -0.018, with a mean change of -0.068, which means documents closer to the ranking feature profile become less aligned with independently judged document quality over the simulation horizon. A random-target control has shown that it is associated with adaptation toward ranking-derived incentives rather than iterative rewriting alone. The resulting ecosystem dynamics are strongly domain-dependent. Together, these findings show how repeated optimization against a fixed LLM ranking signal can reshape both content populations and the incentives faced by content creators.
Chinese Translation
生成式引擎优化(Generative Engine Optimization, GEO)被越来越多地用于提升内容在基于大语言模型(LLM)的检索系统中的可见性,然而其在反复优化过程中的群体层面效应仍鲜为人知。我们提出了排序信号利用下的内容同质化框架(Content Homogenization under rAnking Signal Exploitation, CHASE),这是一个受控仿真框架,用于研究当创作者反复根据LLM排序信号调整文档时,内容生态系统如何被重塑。我们将排序作为来源可见性的代理指标,并通过有据可依的生成式回答中的引用对该抽象进行验证,在六个领域上获得了0.853 ± 0.093的排序-引用AUC。随后,CHASE在不同领域上迭代执行排序、特征判别、重写与评估,共进行20轮。在所有六个领域中,质量与排序的一致性均呈下降趋势:从第0轮到第20轮,Spearman相关系数的变化范围为-0.107至-0.018,平均变化为-0.068,这意味着在仿真周期内,越接近排序特征分布的文档,与独立评估的文档质量的一致性越低。随机目标对照组表明,这一现象与创作者向排序所诱发的激励方向进行适应有关,而非仅仅由迭代重写所致。由此产生的生态系统动态呈现出强烈的地域依赖性。这些发现共同表明,针对固定的LLM排序信号进行反复优化,会同时重塑内容群体以及内容创作者所面临的激励结构。
cs.AI / 147 / 2608.30498

CM2: Multimodal Cultural Reasoning via an Integrated Multi-Agent Framework

CM2:基于集成多智能体框架的多模态文化推理
Li, Qi, Kang, Zhaojie, He, Yingjie, Lin, Zheng, Zhang, Hao, Wu, Guangxin, Gong, Yan, Fu, Rong, Ni, Jianyuan
Abstract
Multimodal Large Language Models (MLLMs) have shown remarkable success in STEM domains, where progress is often driven by vertical, step-by-step deduction under relatively stable symbol systems. Their horizontal, interdisciplinary cultural reasoning, however, remains underexplored.We propose CM2, a multi-agent framework grounded in the cognitive pathway of human cultural interpretation. CM2 integrates multimodal perception, retrieval-augmented generation, networked reasoning, gated fusion, and reward-driven feedback.Experiments on CM2D across multiple MLLM backbones show consistent gains over CoT and typical reasoning paradigms; ablations validate each module's contribution, and conflict analyses confirm genuine cross-modal arbitration.
Chinese Translation
多模态大语言模型(MLLMs)已在STEM领域取得显著成功,其进步往往由相对稳定符号系统下的纵向、逐步演绎所驱动。然而,它们在横向、跨学科的文化推理方面仍缺乏充分探索。我们提出CM2,一个基于人类文化解读认知路径的多智能体框架。CM2集成了多模态感知、检索增强生成、网络化推理、门控融合以及奖励驱动的反馈机制。在CM2D数据集上针对多个MLLM骨干模型的实验表明,该方法相比思维链(CoT)及典型推理范式均取得一致提升;消融实验验证了各模块的贡献,冲突分析则证实了真正的跨模态仲裁能力。
cs.AI / 148 / 2608.30517

ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

ScienceArena:基于最新科学奥林匹克竞赛对大语言模型进行基准评测
Zhao, Guangxiang, Shi, Qilong, Xiao, Xusen, Liu, Wenpu, Li, Yaoming, Hao, Linfeng, Hou, Shuyang, Guo, Zijian, Zhang, Xinrui, Zhao, Yuntian, Wang, Zhengyang, Liu, Wenrui, Wu, Yuhan, Yang, Tong, Sun, Lin, Zhang, Xiangzheng
Abstract
Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025. Its open-ended, multi-step problems use process-credit rubrics, making faithful scoring difficult. We build ScienceArena through an expert-audited digitization pipeline that converts official exams, figures, solutions, and rubrics into structured items verified by olympiad medalists. To scale evaluation beyond costly human grading, we calibrate LLM-as-judge against medalist ground truth on archived answers from five models across IPhO and IChO; two strong judges stay within one point of expert total scores. Medalist notes show that failures often stem from visual grounding, structure fidelity, and global problem control rather than missing terminology. Evaluating fourteen recent LLMs with interleaved solving, we find that top models obtain medal-equivalent rubric scores on several public international exams, while chemistry and long-horizon consistency remain key bottlenecks. We provide an interactive \href{https://science-arena.onrender.com/}{demo}.
Chinese Translation
基准测试饱和与数据污染问题日益掩盖了前沿大语言模型(LLM)真实的科学推理能力。我们提出了 ScienceArena,一个源自十三项公开科学竞赛的奥林匹克风格基准,涵盖物理、化学和生物领域,包括 IPhO 与 IChO 2025–2026、IBO 2023、USAPhO 2026 以及 USNCO 2025。其开放式、多步骤的题目采用过程评分细则(process-credit rubrics),使忠实评分变得困难。我们通过一个经专家审核的数字化流水线构建了 ScienceArena,将官方试卷、图表、解答与评分细则转换为结构化条目,并由奥林匹克奖牌得主进行验证。为了将评测扩展到成本高昂的人工评分之外,我们利用 IPhO 和 IChO 中五个模型的存档答案,将 LLM-as-judge(以大模型作为评判者)与奖牌得主的真值答案进行校准;两个较强的评判模型与专家总分之间的差距保持在一分以内。奖牌得主的批注表明,失败往往源于视觉定位、结构保真度以及对整体问题的把控,而非术语缺失。通过采用交错式求解方式评测十四个近期的大语言模型,我们发现顶尖模型在若干公开国际赛事上可获得相当于奖牌水平的评分,但化学领域与长程一致性仍是关键瓶颈。我们提供了交互式演示(demo)。
cs.AI / 149 / 2608.30520

Learning-Assisted Congestion-Aware Route Scheduling for Semiconductor Fab Material Control Systems

面向半导体晶圆厂物料控制系统的学习辅助拥塞感知路径调度
Yin, Hao, Tu, Meiqi, Liu, Anbang, Lin, Shaochong, Shen, Max Z. J.
Abstract
Automated material handling systems in semiconductor fabs are operated by a material control system (MCS) that must schedule a relay route for every transport command online, before execution. This is a data-driven scheduling problem in which route cost is dominated in the upper tail by queueing at heterogeneous, partially observable relay equipment, so route selection requires estimating both delivery time and congestion risk at the decision moment. This paper proposes a transport-network-aware dynamic congestion representation (TN-DCR). Built on a static directed transport graph induced by historically observed relay segments, TN-DCR combines structural route priors, multi-window network-wide congestion context, route-level bottleneck exposure, and an inductive graph-aware route embedding, all constructed under a prediction-time-safety invariant that admits only information observed strictly before the prediction moment. The representation feeds separate queue- and transfer-time regressors and an ordinal multi-label classifier producing calibrated multi-threshold exceedance scores, with an empirical-Bayes stock-key residual correction reducing systematic queue-time underprediction. The predictions serve as costs in a risk-constrained route-scheduling rule that minimizes predicted delivery time subject to a bound on extreme-congestion probability, embedding the learned predictors within a lightweight operations-research decision model. In a controlled closed-loop evaluation, mean delivery time falls by 16.4\% and internal resource waiting time by 22.6\% while throughput remains essentially unchanged.
Chinese Translation
半导体晶圆厂中的自动物料搬运系统由物料控制系统(MCS)驱动,该系统必须在执行前为每个搬运指令在线调度一条中继路径。这是一个数据驱动的调度问题:路径成本的上尾部分主要由异构且部分可观测的中继设备处的排队所主导,因此路径选择需要在决策时刻同时估计配送时间和拥塞风险。本文提出了一种面向运输网络的动态拥塞表征方法(TN-DCR)。TN-DCR 构建于由历史观测中继段导出的静态有向运输图之上,融合了结构化路径先验、多窗口全网拥塞上下文、路径级瓶颈暴露度以及归纳式图感知路径嵌入,且所有信息均在预测时间安全不变式的约束下构建,即仅使用严格早于预测时刻所观测到的信息。该表征分别输入排队时间和传输时间的回归器,以及一个序数多标签分类器,后者产生经校准的多阈值超标评分,并通过经验贝叶斯库存键残差修正来减少排队时间的系统性低估。预测结果作为代价输入一个风险约束的路径调度规则,该规则在极端拥塞概率受上界约束的条件下最小化预测配送时间,从而将学习到的预测器嵌入轻量化的运筹学决策模型中。在受控闭环评估中,平均配送时间降低16.4%,内部资源等待时间降低22.6%,而产能基本保持不变。
cs.AI / 150 / 2608.30532

DiffPDE: Masked Diffusion Language Models as PDE Solver

DiffPDE:将掩码扩散语言模型用作偏微分方程求解器
Guo, Wenxuan, Hong, Yuyang, Fan, Lubin, Fu, Zhaojin, Chen, Lin, Ding, Kun, Xiang, Shiming
Abstract
Existing approaches for synthesizing Partial Differential Equation (PDE) solvers predominantly rely on autoregressive models, yet their global left-to-right decoding incurs substantial redundancy when addressing inherently localized bugs. In this work, we challenge this inefficient paradigm and propose DiffPDE, a framework leveraging discrete diffusion language models for targeted code repair. By introducing a localized re-masking and infilling strategy, DiffPDE regenerates only erroneous regions while preserving correct context, naturally aligning generation with the sparse nature of PDE errors. Furthermore, to handle coupled bugs requiring sequential interventions, we present Iterative Debugging GRPO (ID-GRPO), a reinforcement learning scheme that enables multi-round debugging within single trajectories via intermediate rewards. Experiments on PDEBench show that DiffPDE achieves competitive accuracy, outperforms same-scale AR models, and significantly accelerates repair.
Chinese Translation
现有的偏微分方程(PDE)求解器合成方法主要依赖于自回归模型,然而其全局从左到右的解码方式在处理本质上局部的错误时会产生大量冗余。在本工作中,我们挑战了这一低效范式,提出了 DiffPDE——一个利用离散扩散语言模型进行针对性代码修复的框架。通过引入局部化的重新掩码与填充策略,DiffPDE 仅重新生成错误区域,同时保留正确的上下文,使生成过程自然契合 PDE 错误的稀疏特性。此外,为处理需要多步骤干预的耦合型错误,我们提出了迭代调试 GRPO(Iterative Debugging GRPO, ID-GRPO),这是一种强化学习方案,能够通过中间奖励在单条轨迹内实现多轮调试。在 PDEBench 上的实验表明,DiffPDE 取得了具有竞争力的精度,优于同等规模的自回归模型,并显著加快了修复速度。
cs.AI / 151 / 2608.30543

Designing an Auditable LLM-Supported Workflow for Qualitative Thematic Analysis

设计一种可审计的、由大语言模型支持的质性主题分析工作流程
Jeldtoft, Nadia Jul, Yousef, Tariq
Abstract
Large Language Models (LLMs) offer new possibilities for scaling qualitative analysis, but existing applications often provide limited methodological transparency regarding how qualitative methods are translated into computational procedures. This paper presents an auditable and privacy-preserving computational operationalization of inductive and latent Thematic Analysis (TA). This paper first derives five design principles from the methodological requirements of TA and the conditions introduced by LLM-based inference: preserving interpretative context, maintaining traceable relationships between empirical material and analytical outputs, representing analytical constructs and reasoning explicitly, constraining LLM inference to interpretative tasks, and enabling privacy-preserving local deployment. Second, it presents a proof-of-concept for a two-phase workflow that operationalizes these principles by combining interpretative LLM inference with deterministic procedural control to generate codes, analytical justifications, themes, and theme descriptions while preserving explicit links to the source material. Third, it proposes an evaluation framework combining structural comparison with human-led TA and independent expert assessment of analytical quality. The evaluation is conducted on semi-structured Danish interview transcripts. and the results shows that the workflow produces code-level outputs with coverage broadly comparable to human annotations and highly rated analytical justifications, while generating a more compressed thematic structure characterized by fewer and broader themes. The findings demonstrate the feasibility of auditable LLM-supported TA through a modular workflow designed to scale to larger datasets, accommodate different LLMs, and support transfer across research domains, with domain adaptation primarily requiring adjustments to the prompting strategy.
Chinese Translation
大语言模型(LLM)为质性分析的规模化提供了新的可能性,但现有应用在质性方法如何被转化为计算程序这一问题上,往往缺乏方法论上的透明度。本文提出了一种可审计且保护隐私的、针对归纳性与潜在性主题分析(Thematic Analysis, TA)的计算操作化方案。首先,本文从TA的方法论要求以及基于LLM推理所带来的条件出发,推导出五项设计原则:保留解释性语境、维护经验材料与分析结果之间可追溯的关系、显式表达分析构念与推理过程、将LLM推理限定于解释性任务,以及支持保护隐私的本地部署。其次,本文提出了一个两阶段工作流程的概念验证,该流程通过将解释性的LLM推理与确定性的程序化控制相结合来落实上述原则,从而生成编码、分析论证、主题及主题描述,同时保留与原始材料的显式关联。第三,本文提出了一个评估框架,将结构化比较与人工主导的TA以及对分析质量的独立专家评估相结合。评估基于丹麦语半结构化访谈转录文本进行,结果表明:该工作流程生成的编码级结果在覆盖范围上与人工标注大体相当,其分析论证获得高度评价,同时生成了更压缩的主题结构,其特征为主题数量更少且范围更宽。研究结果表明,通过模块化工作流程实现可审计的LLM辅助TA是可行的;该流程可扩展至更大规模的数据集、兼容不同的LLM,并支持跨研究领域的迁移,其中领域适应主要通过调整提示策略来实现。
cs.AI / 152 / 2608.30550

GarmentWeaver: Schema-Aware Structured Synthesis for Multimodal Sewing Patterns

GarmentWeaver:面向多模态缝纫纸样的模式感知结构化合成方法
Lu, Yinwen, Luo, Weihao, Zhong, Yueqi
Abstract
Multimodal Sewing pattern generation aims to infer executable sewing patterns from design cues such as sketches and textual descriptions. As an interpretable and simulation-compatible representation, sewing patterns are particularly valuable for digital garment creation. However, existing methods often model garment specifications as flat long sequences, which entangles garment structure with detailed parameters and leads to redundant components, inaccurate local details, and poor simulation compatibility. In this paper, we present GarmentWeaver, a schema-aware framework for multimodal Sewing pattern generation. GarmentWeaver constructs compact hierarchical targets by activating garment-relevant structural branches and predicts executable Sewing patterns in a structured manner. Specifically, we introduce a schema-aware target construction strategy, build the generator on top of a pretrained vision-language model for multimodal garment understanding, and impose feasibility-aware regularization to encourage structurally valid and simulation-compatible outputs. Extensive experiments show that GarmentWeaver produces more accurate and more executable sewing patterns than strong baselines, while also yielding better simulation results. These findings demonstrate the effectiveness of schema-aware structured generation for reliable multimodal Sewing pattern prediction.
Chinese Translation
多模态缝纫纸样生成旨在从草图和文本描述等设计线索中推断出可执行的缝纫纸样。作为一种可解释且与仿真兼容的表示形式,缝纫纸样对于数字化服装制作尤为重要。然而,现有方法通常将服装规格建模为扁平的长序列,导致服装结构与细节参数相互纠缠,产生冗余组件、局部细节不准确以及仿真兼容性差等问题。在本文中,我们提出了GarmentWeaver,一个面向多模态缝纫纸样生成的模式感知框架。GarmentWeaver通过激活与服装相关的结构分支来构建紧凑的层次化目标,并以结构化的方式预测可执行的缝纫纸样。具体而言,我们引入了一种模式感知的目标构建策略,在预训练视觉-语言模型的基础上构建生成器以实现多模态服装理解,并施加可行性感知的正则化约束,以促使模型生成结构有效且与仿真兼容的输出。大量实验表明,与强基线方法相比,GarmentWeaver能够生成更精确、更可执行的缝纫纸样,同时获得更好的仿真效果。这些结果证明了模式感知结构化生成方法在可靠的多模态缝纫纸样预测中的有效性。
cs.AI / 153 / 2608.30556

AdaPath: Query-Adaptive Path-Finding via Path-Bank for Multi-Hop Implicit Biomedical KGQA

AdaPath:基于路径库的查询自适应路径查找方法,用于多跳隐式生物医学知识图谱问答
Kim, Jun Hyeong, Kim, Dongki, Piao, Yinhua, Hwang, Sung Ju
Abstract
Path-finding over knowledge graphs has become an effective way to ground LLM reasoning on multi-hop questions. However, biomedical QA introduces two distinct challenges that general-domain methods are not designed for: (i) queries do not expose intermediate reasoning and can be answered through multiple valid pathways, and (ii) biomedical knowledge graphs are densely connected, so path-finding methods easily take wrong turns. To address these challenges, we propose AdaPath, a path-finding framework that retrieves query-adaptive meta-paths from Path-Bank, which captures both query semantics and biomedical knowledge graph structure. AdaPath provides the missing cues in biomedical queries while effectively pruning dense knowledge graph neighborhoods during multi-hop reasoning. We further release BioStrat-QA, a biomedical KGQA benchmark that stratifies multi-hop queries by how much intermediate reasoning they expose. Across biomedical KGQA benchmarks, AdaPath consistently outperforms baselines, sustaining meaningful path-finding even when multi-hop queries expose less surface information. The source code is available at https://github.com/Jun-Hyeong-Kim/AdaPath.
Chinese Translation
知识图谱上的路径查找已成为将大语言模型(LLM)推理落地于多跳问题的有效方式。然而,生物医学问答引入了两个通用领域方法未曾应对的独特挑战:(i) 查询不暴露中间推理过程,且可以通过多条有效路径进行回答;(ii) 生物医学知识图谱连接高度密集,路径查找方法极易走错方向。为应对这些挑战,我们提出了AdaPath,一个从路径库(Path-Bank)中检索查询自适应元路径的路径查找框架,该框架同时捕获查询语义与生物医学知识图谱结构。AdaPath为生物医学查询中缺失的线索提供补充,同时在多跳推理过程中有效剪枝密集的知识图谱邻域。我们还进一步发布了BioStrat-QA,这是一个生物医学知识图谱问答(KGQA)基准,按照多跳查询暴露的中间推理信息量对其进行分层。在多个生物医学KGQA基准上,AdaPath始终优于基线方法,即使多跳查询暴露的表面信息较少,仍能维持有意义的路径查找。源代码发布于 https://github.com/Jun-Hyeong-Kim/AdaPath。
cs.AI / 154 / 2608.30567

TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI

TuringLLM:面向物理AI的高效基础模型扩展
Zhang, Yuheng, Wang, Yizhao, Zhu, Da, Zhou, Hua, He, Yue, Hu, Jiahui, Tang, Shaman, Chen, Hanlin, Wei, Yuhua, Liu, Anhua, Su, Shuang, Xin, Rui, Wang, MingYuan, Li, MingHao, Yang, HaoJie, Liu, Siqi, Zheng, Jianlei, Huang, WeiChao, Wu, Qiman, Zhang, Hang, Yang, HongGou, Liu, Xianming
Abstract
We present Turing-20B-A2B, a 20B-parameter Mixture-of-Experts language model that activates approximately 2B parameters per token, designed for long-context and latency-sensitive physical AI applications. The model adopts Quantile Routing in a dynamic top-k configuration, enabling token-adaptive expert allocation while maintaining balanced expert utilization and a controlled average compute budget. During deployment, we further apply capacity-constrained routing to prompt prefill for more regular and efficient expert execution, while retaining dropless routing during pretraining. Turing-20B-A2B also employs a hybrid attention architecture that combines Lightning Attention with a small number of full-attention layers for efficient long-context modeling. The model is pretrained with a progressive three-stage curriculum and extended to a native context length of 128K through continued pretraining, with further inference-time extension to 512K using YaRN. Despite its compact active-parameter budget, Turing-20B-A2B achieves, at the base-model stage, overall general capability exceeding Qwen3-8B Base and approaching Qwen3.5-9B Base, while maintaining strong long-context performance and favorable prefill-latency scaling. These results demonstrate an effective balance among model capability, long-context scalability, and practical inference efficiency.
Chinese Translation
我们提出了 Turing-20B-A2B,一个 200 亿参数的混合专家语言模型,每个词元仅激活约 20 亿参数,专为长上下文和延迟敏感的物理AI应用而设计。该模型在动态 top-k 配置下采用分位数路由,在保持专家利用均衡和受控平均计算预算的同时,实现词元自适应的专家分配。在部署阶段,我们进一步对提示预填充应用容量受限路由,以实现更规整、更高效的专家执行,同时在预训练期间保留无丢弃路由。Turing-20B-A2B 还采用了混合注意力架构,将 Lightning Attention 与少量全注意力层相结合,以实现高效的长上下文建模。模型通过渐进式三阶段课程进行预训练,并通过持续预训练将原生上下文长度扩展至 128K,进而利用 YaRN 在推理时进一步扩展至 512K。尽管激活参数预算精简,Turing-20B-A2B 在基座模型阶段的综合通用能力超越了 Qwen3-8B Base,并接近 Qwen3.5-9B Base,同时保持了强大的长上下文性能和良好的预填充延迟扩展性。这些结果表明,该模型在模型能力、长上下文可扩展性和实际推理效率之间实现了有效平衡。
cs.AI / 155 / 2608.30581

Automated Testing of LLM-Based Post Hoc Explainers Using Model Checking as an Oracle

基于模型检测作为测试预言机的大语言模型事后解释器自动化测试
Gross, Dennis, Spieker, Helge
Abstract
Large language models (LLMs) are used as post hoc explainers of sequential decision-making policies, producing natural-language explanations of why an action was chosen. However, LLMs often generate plausible but incorrect statements, and no existing approach systematically tests whether such explanations are faithful to the underlying environment. Two classic software testing challenges stand in the way: there is no oracle for the correctness of an explanation, and the test inputs, natural language queries about a policy's behavior, lack the structure needed for systematic test case generation. We address both. Probabilistic model checking provides the test oracle, computing exact reference results against which LLM answers are graded automatically. A taxonomy of post hoc query categories structures the input space around the environment-level facts from which policy explanations are composed; test cases generated from it are prioritized by question-specific diagnostic difficulty scores. Across seven MDP environments, the testing separates three open-weight LLMs: a reasoning model passes 85% of test cases, a mid-size model 70%, and a 1B model falls below the random baseline, while prioritization surfaces significantly harder cases than random selection. Our results indicate how trustworthy LLM-generated explanations are in model-free settings, where the same LLMs are used but no oracle exists to verify them.
Chinese Translation
大语言模型(LLM)被用作序贯决策策略的事后(post hoc)解释器,生成关于为何选择某一动作的自然语言解释。然而,LLM常常生成看似合理却不正确的陈述,且现有方法尚无法系统性地检验这些解释是否忠实于底层环境。两个经典的软件测试难题阻碍了这一目标:其一,不存在判定解释正确性的测试预言机(oracle);其二,测试输入——即关于策略行为的自然语言查询——缺乏系统性测试用例生成所需的结构。本文针对这两个问题提出了解决方案。概率模型检测(probabilistic model checking)提供了测试预言机,可计算出精确的参考结果,用于自动评判LLM的回答。我们提出了一种事后查询类别的分类体系,围绕构成策略解释的环境级事实来组织输入空间;基于该体系生成的测试用例按针对特定问题的诊断难度分数进行优先级排序。在七个马尔可夫决策过程(MDP)环境中,该测试方法区分了三个开放权重LLM的表现:一个推理模型通过了85%的测试用例,一个中等规模模型通过70%,而一个1B参数模型则低于随机基线;同时,优先级排序方法所呈现的测试用例显著难于随机选择。我们的结果表明了LLM生成的解释在无模型(model-free)设置下的可信程度——在该设置下使用相同的LLM,但不存在可验证其输出的预言机。
cs.AI / 156 / 2608.30650

Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reasoning

散度的几何形态:追踪隐状态轨迹以实现自适应多轮推理
Liang, Jie, Yu, Zhengxin, Nasiri, Hamid, Garraghan, Peter
Abstract
LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn context accumulates, it can destabilize the underlying LLM's internal representation of task-relevant information from earlier turns, blurring the boundary between constructive reasoning and representation drift. We formulate multi-turn reasoning as a hidden-state trajectory of the underlying LLM that is characterized via two complementary signals: temporal curvature that captures the directional consistency of turn-to-turn updates, and variance slope which measures the expansion or contraction of the exploration space. Across four tasks and three underlying LLMs, we observed that these geometric signals distinguish between correct and incorrect episodes prior to completion. We further decompose each episode into three-action chains formed from four actions (Read, Write, Respond, Transfer) and show that separability is action-dependent, with different signals distinguishing various chain patterns. Our experiments demonstrate that trajectory geometry can identify critical turns in the reasoning process, increasing task success rates on $\tau$-Bench from 24.1% to 39.6% while reducing token cost by 11.2%.
Chinese Translation
LLM 智能体需要在严格的资源约束下,于漫长的多轮交互中维持与目标一致的推理。然而,随着多轮上下文的累积,底层 LLM 对早期轮次中任务相关信息的内部表征可能会变得不稳定,从而模糊了建设性推理与表征漂移之间的边界。我们将多轮推理形式化为底层 LLM 的隐状态轨迹,并通过两个互补的信号加以刻画:一是时间曲率(temporal curvature),用于捕捉轮次间更新的方向一致性;二是方差斜率(variance slope),用于衡量探索空间的扩张或收缩。在四个任务和三个底层 LLM 上的实验表明,这些几何信号能够在推理完成之前区分正确与错误的推理过程。我们进一步将每个推理过程分解为由四种动作(Read、Write、Respond、Transfer)构成的三动作链,并证明其可分性依赖于具体动作,不同的信号可区分不同的链模式。我们的实验表明,轨迹几何能够识别推理过程中的关键轮次,将 $ au$-Bench 上的任务成功率从 24.1% 提升至 39.6%,同时降低 11.2% 的 token 开销。
cs.AI / 157 / 2608.30652

PyKEEN-NSX: A Modular Framework for Static, Dynamic and Schema-Aware Negative Sampling in PyKEEN

PyKEEN-NSX:一个用于PyKEEN中静态、动态和模式感知负采样的模块化框架
Diliso, Ivan, Fanizzi, Nicola, d'Amato, Claudia
Abstract
Embedding methods have become popular due to their scalability on link prediction and/or triple classification tasks on Knowledge Graphs (KGs). Embedding models are trained relying on both positive and negative samples of triples. However, since KGs generally contain only positive assertions, negative samples are artificially generated through negative sampling strategies, ranging from simple random corruption to more sophisticated approaches that exploit structural, semantic, or embedding information. The design and implementation of advanced negative samplers remains challenging, as most popular Knowledge Graph Embedding (KGE) libraries provide support only for basic strategies and lack a unified framework for developing more advanced and customized solutions. To address this gap, we introduce PyKEEN-NSX, an extension of PyKEEN, the popular KGE framework, that provides a modular engineered abstraction for negative sampling. The proposed architecture separates the generation of candidate negative pools, conditioned on an explicit context, from the selection strategy, enabling the development and integration of static, schema-aware and dynamic approaches within a consistent framework. Based on this abstraction, we implement six negative samplers, while remaining fully compatible with existing PyKEEN workflows and pipelines. As a proof of concept, we study negative availability across four datasets, showing that constrained pools frequently fall below the requested number of negatives, so that the encoded criterion is to a large extent replaced by the random fallback that supplements them.
Chinese Translation
嵌入方法因其在知识图谱(KG)上的链接预测和/或三元组分类任务中的可扩展性而广受欢迎。嵌入模型的训练依赖于正样本和负样本三元组。然而,由于知识图谱通常只包含正断言,负样本需要通过负采样策略人为生成,其范围从简单的随机破坏到利用结构、语义或嵌入信息的更复杂方法。高级负采样器的设计与实现仍然具有挑战性,因为大多数流行的知识图谱嵌入(KGE)库仅支持基础策略,缺乏用于开发更高级和定制化解决方案的统一框架。为填补这一空白,我们提出了PyKEEN-NSX,这是流行的KGE框架PyKEEN的一个扩展,为负采样提供了模块化的工程化抽象。所提出的架构将基于显式上下文条件的候选负样本池的生成与选择策略相分离,从而能够在一致的框架内开发和集成静态、模式感知和动态方法。基于该抽象,我们实现了六种负采样器,同时与现有的PyKEEN工作流和流水线保持完全兼容。作为概念验证,我们研究了四个数据集中的负样本可用性,结果表明受限的样本池经常低于所请求的负样本数量,因此编码的准则在很大程度上被补充负样本的随机回退机制所取代。
cs.AI / 158 / 2608.30672

HiRS-Agent: A Hierarchical Multi-Agent System for Reliable Long-Horizon Remote Sensing Task Solving

HiRS-Agent:面向可靠长时程遥感任务求解的分层多智能体系统
Mu, Boyang, Wei, Zhiwei, Peng, Mugen, Xu, Wenjia
Abstract
Recent advances in large language models and multimodal models have pushed remote sensing (RS) processing from simple perception models to agentic systems designed to tackle complex, long-horizon RS tasks. However, existing systems often rely on monolithic decision-making frameworks, which fail to accommodate the multi-stage, interdependent nature of RS tasks. This centralized approach leads to challenges such as unstable task execution, incorrect tool usage, and error propagation across stages. To address these issues, we propose HiRS-Agent, a hierarchical multi-agent system for long-horizon RS task solving. HiRS-Agent adopts a two-level collaborative architecture: the Manager Layer handles dynamic routing, step-level verification, replanning, and termination control, while the Specialist Layer organizes domain-specific tools according to the RS workflow and is responsible for subtask reasoning and tool execution. To further enhance the system's capability, we introduce a two-stage supervised tuning strategy and a verification-guided hierarchical reinforcement learning stage to jointly optimize coordination and tool-use policies. Experiments on Earth-Agent Benchmark and ThinkGeo show that HiRS-Agent substantially improves long-horizon tool-use capability and final-task correctness, demonstrating the effectiveness of structured multi-agent collaboration for reliable RS agents. The code is publicly available at https://github.com/IntelliSensing/HiRS-Agent.
Chinese Translation
大语言模型与多模态模型的最新进展,推动遥感(Remote Sensing, RS)处理从简单的感知模型发展为旨在应对复杂长时程遥感任务的智能体系统。然而,现有系统通常依赖单体式决策框架,难以适应遥感任务多阶段、相互依赖的特性。这种集中式方法导致了任务执行不稳定、工具使用错误以及错误在阶段间传播等挑战。为解决这些问题,我们提出了 HiRS-Agent,一个用于长时程遥感任务求解的分层多智能体系统。HiRS-Agent 采用两级协作架构:管理层(Manager Layer)负责动态路由、步骤级验证、重规划与终止控制;专家层(Specialist Layer)依据遥感工作流组织领域专用工具,并负责子任务推理与工具执行。为进一步增强系统能力,我们引入两阶段监督微调策略以及验证引导的分层强化学习阶段,共同优化协调策略与工具使用策略。在 Earth-Agent Benchmark 和 ThinkGeo 上的实验表明,HiRS-Agent 显著提升了长时程工具使用能力和最终任务正确性,验证了结构化多智能体协作对构建可靠遥感智能体的有效性。代码已公开于 https://github.com/IntelliSensing/HiRS-Agent。
cs.AI / 159 / 2608.30676

MedAgent-R1: Faithfulness-Aware Reinforcement Learning for Evidence-Grounded Medical Reasoning

MedAgent-R1:面向证据支撑医疗推理的忠实性感知强化学习
Chen, Jiangwang, Zhang, Chenghao, Cai, Hengxing
Abstract
When medical AI systems hallucinate clinical reasoning, the consequences extend beyond incorrect answers: fabricated justifications that superficially reference retrieved evidence can mislead clinicians into unsafe treatment decisions. Medical reasoning agents must therefore produce not only correct answers but also faithful justifications that clinicians can verify against cited evidence. We identify a systematic failure mode in RL-trained retrieval agents: outcome-only rewards improve accuracy while degrading faithfulness, a phenomenon we term confident hallucination. The agent learns to answer from parametric memory and backfill plausible but unsupported justifications; citation fabrication rates rise from 16.5% to 31.8% even as accuracy improves by 5 points over the supervised baseline. We address this with a faithfulness-gated reward design: accuracy credit is conditioned on evidence grounding via a hard gate, complemented by retrieval validity and conciseness signals that close exploitation paths unique to agentic retrieval. The resulting system, MedAgent-R1, reduces citation fabrication from 31.8% to 4.7% and raises evidence completeness from 58.7 to 82.6 while maintaining 75.1% accuracy, with 13.2-point gains on HealthBench Safety. Under the same agentic retrieval setup, MedAgent-R1 outscores GPT-4o on faithfulness-specific dimensions (Factual Support 4.55 vs. 4.25; Overclaiming 4.40 vs. 4.15) while remaining below GPT-4o in overall accuracy, suggesting that explicit faithfulness training yields evidence-grounding gains not achieved by scaling alone.
Chinese Translation
当医疗AI系统在临床推理中产生幻觉时,其后果不仅限于错误的答案:表面上引用了检索证据的编造性论证可能会误导临床医生做出不安全的治疗决策。因此,医疗推理智能体不仅必须给出正确的答案,还必须提供临床医生可以对照所引证据进行验证的忠实论证。我们在基于强化学习(RL)训练的检索智能体中发现了一种系统性的失效模式:仅基于结果的奖励在提升准确率的同时降低了忠实性,我们将这一现象命名为“自信幻觉”。该智能体学会从参数化记忆中直接作答,并事后编造看似合理却缺乏支撑的论证;尽管准确率相比监督基线提升了5个百分点,引用编造率却从16.5%升至31.8%。为解决这一问题,我们提出了一种忠实性门控奖励设计:通过硬门控将准确率奖励以证据支撑为前提条件,并辅以检索有效性和简洁性信号,堵住智能体检索中特有的可被利用的漏洞。由此构建的系统MedAgent-R1将引用编造率从31.8%降至4.7%,证据完整性从58.7提升至82.6,同时保持75.1%的准确率,并在HealthBench Safety上取得13.2个百分点的提升。在相同的智能体检索设置下,MedAgent-R1在忠实性相关维度上的得分超过GPT-4o(事实支撑4.55比4.25;过度声明4.40比4.15),但整体准确率仍低于GPT-4o,这表明显式的忠实性训练能够带来仅靠规模扩展无法实现的证据支撑能力提升。
cs.AI / 160 / 2608.30685

ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents

ATLAS:面向工业工具使用智能体的双视角诊断评估框架
Chen, Wei, Zhou, Peilun, Hu, Zhaoyu, Chai, Jiajun, Hou, Zhongni, Zhang, Yufei, Xu, Derong, Yin, Guojun, Lin, Wei, Zheng, Zhi, Xu, Tong
Abstract
Large language model (LLM) agents are increasingly deployed in user-facing services that require iterative tool use under dynamic business conditions. Reliable evaluation is essential for sustained improvement: it must reveal capability deficiencies, inform priorities, and assess interventions. Yet industrial agent service unfolds both through the iterative trajectory of a current request and through continued user interaction. Final-outcome assessment can therefore obscure where deficiencies arise and whether later service remains aligned with context from earlier exchanges. We propose ATLAS, a dual-horizon diagnostic evaluation framework for industrial tool-use agents. At the request horizon, trajectory-wise diagnostic signals relate deficiencies to execution locations and capability concerns. At the interaction horizon, user-wise signals assess whether service remains responsive across continued interaction. Together, these views provide structured diagnostic evidence for analyzing execution deficiencies and sustained service behavior. ATLAS instantiates them as executable signals with explicit evidence scopes and decision boundaries. LLM judge interfaces are calibrated against high-confidence references from real business logs; when needed, their decision behavior is distilled into efficient diagnostic models for lower-latency, lower-cost evaluation. The resulting feedback supports policy optimization. We evaluate ATLAS on Meituan Xiaotuan production traffic. Offline experiments assess diagnostic-signal fidelity and replay-based policy improvement, while online A/B experiments show concurrent gains in user engagement, downstream business outcomes, and sampled human-audit quality.
Chinese Translation
大型语言模型(LLM)智能体日益被部署于需要在动态业务条件下进行迭代式工具调用的面向用户服务中。可靠的评估对于持续改进至关重要:它必须能够揭示能力缺陷、为优先级排序提供依据,并评估干预措施的效果。然而,工业级智能体服务既体现在当前请求的迭代轨迹中,也体现在用户的持续交互过程中。因此,仅基于最终结果的评估可能掩盖缺陷产生的位置,也无法判断后续服务是否仍与早期交互的上下文保持一致。我们提出了 ATLAS,一个面向工业工具使用智能体的双视角(dual-horizon)诊断评估框架。在请求视角上,基于轨迹的诊断信号将缺陷与执行位置及能力问题相关联;在交互视角上,基于用户的信号评估服务在持续交互过程中是否保持响应性。这两种视角共同为分析执行缺陷和持续服务行为提供了结构化的诊断证据。ATLAS 将其实现为具有明确证据范围和决策边界的可执行信号。LLM 评审器的接口基于来自真实业务日志的高置信度参考进行校准;在需要时,其决策行为可被蒸馏为高效的诊断模型,以实现更低延迟、更低成本的评估。由此产生的反馈可支持策略优化。我们在美团小团的 生产流量上评估了 ATLAS。离线实验评估了诊断信号的保真度以及基于回放(replay)的策略改进效果,而在线 A/B 实验则表明该方法在用户参与度、下游业务指标以及抽样人工审计质量方面均取得了同步提升。
cs.AI / 161 / 2608.30726

Multimodal Adaptive Expert Selection with Text Routing and Ordinal Prototype Optimization for Sentiment Analysis

面向情感分析的基于文本路由与序数原型优化的多模态自适应专家选择方法
Chen, Xiaode, Yu, Jiakang, Deng, Hongtao, Qu, Huina, Zhu, Xun, Lou, Yinxia
Abstract
Multimodal Sentiment Analysis (MSA) is a fundamental component of affective computing that aims to decipher complex emotional states by integrating verbal content with non-verbal cues including vocal intonation and facial micro-expressions. While recent disentanglement-based approaches have advanced the field, their potential is hindered by two methodological challenges. First, static computation graphs process all samples indiscriminately regardless of semantic complexity, which leads to suboptimal representation for diverse emotional expressions and contextual scenarios. Second, generic contrastive objectives often neglect the intrinsic ordinal hierarchy of sentiment intensities. To systematically address these limitations, we introduce Multimodal Adaptive Expert Selection with Text Routing and Ordinal prototype optimization (MAESTRO), a novel framework designed to dynamically orchestrate and refine multimodal representations. Drawing inspiration from an orchestra conductor, we design a Text-Guided Hybrid Mixture-of-Experts (MoE) mechanism. Unlike static fusion, this module utilizes linguistic context as a routing signal to dynamically activate specific audio-visual experts, thereby resolving cross-modal ambiguity through adaptive feature enhancement. Furthermore, to capture fine-grained sentiment gradations, we propose an Ordinal-aware Prototype Contrastive Learning (O-PCL). By incorporating distance-based penalties into the prototype learning objective, O-PCL enforces a structured latent space that preserves the natural order of emotion. Extensive experiments on the CMU-MOSI and CMU-MOSEI benchmarks demonstrate that MAESTRO achieves state-of-the-art performance, and qualitative analysis further confirms the interpretability of our dynamic routing paradigm.
Chinese Translation
多模态情感分析(Multimodal Sentiment Analysis, MSA)是情感计算的基本组成部分,旨在通过整合语言内容与包括语音语调和面部微表情在内的非语言线索来解析复杂的情感状态。尽管近年来基于解耦的方法推动了该领域的发展,但其潜力受到两个方法论挑战的制约。首先,静态计算图不区分语义复杂度而以相同方式处理所有样本,导致对多样化情感表达和上下文场景的表征不够理想。其次,通用的对比学习目标往往忽略情感强度固有的序数层级结构。为系统性地解决这些局限,我们提出了基于文本路由与序数原型优化的多模态自适应专家选择框架(Multimodal Adaptive Expert Selection with Text Routing and Ordinal prototype optimization, MAESTRO),这是一个旨在动态编排和精炼多模态表征的新型框架。受交响乐团指挥的启发,我们设计了一种文本引导的混合专家混合(Mixture-of-Experts, MoE)机制。与静态融合不同,该模块利用语言上下文作为路由信号,动态激活特定的视听专家,从而通过自适应特征增强来解决跨模态歧义。此外,为捕捉细粒度的情感梯度,我们提出了序数感知原型对比学习(Ordinal-aware Prototype Contrastive Learning, O-PCL)。通过在原型学习目标中引入基于距离的惩罚项,O-PCL 构建了一个保留情感自然顺序的结构化潜在空间。在 CMU-MOSI 和 CMU-MOSEI 基准数据集上的大量实验表明,MAESTRO 达到了最先进的性能,定性分析进一步证实了我们动态路由范式的可解释性。
cs.AI / 162 / 2608.30751

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

自回归马赛克:探究纯文本语言模型中的二维空间推理能力
Nedungadi, Ashwin, Oehmcke, Stefan, Lüdtke, Stefan
Abstract
Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D spatial layout or simply the ability to translate spatial descriptions into code. We introduce Autoregressive Mosaics (AM-Bench), a benchmark that separates these factors: First, a translation task gives a model a fully specified geometry of a picture in words as a prompt and asks for the code that produces it. Second, a layout task requires the model to compose an image from an underspecified prompt. Across eight open-weight text-and-code-only models, all models reliably translate specified geometry into code, but their open-ended layout performance differs substantially, indicating that these differences are not explained by code-generation ability alone. An output-medium ablation further shows that the interface or medium of expression that the model uses matters: replacing procedural code with raw SVG improves layout scores across all models. Finally, probing model activations shows that a coarse layout plan is present before generation, but reflects only the layout implied by the prompt. During generation, models track the evolving geometric state instead of executing an initially fixed plan. Overall, these results show that 2D spatial performance in text-only LLMs depends on both the model and the output medium, and is not explained by code-generation ability alone.
Chinese Translation
仅在文本和代码上训练的大语言模型(LLM)有时能够生成绘制出可识别图像的程序。然而,这究竟反映了模型内部对二维空间布局的表征,还是仅仅体现了将空间描述转化为代码的能力,尚不清楚。我们提出了自回归马赛克(Autoregressive Mosaics,AM-Bench)基准,用以分离这两个因素:首先,翻译任务向模型提供以文字完整描述的图像几何信息作为提示,并要求其生成相应的代码;其次,布局任务要求模型根据欠规范的提示自行构建图像。在八个开源权重的纯文本与代码模型上,所有模型都能可靠地将明确给定的几何信息翻译为代码,但其开放式布局表现差异显著,表明这些差异无法仅由代码生成能力解释。输出媒介消融实验进一步表明,模型所使用的表达接口或媒介也十分重要:用原始SVG替代过程式代码可提升所有模型的布局得分。最后,对模型激活的探测显示,在生成之前模型中已存在一个粗略的布局规划,但该规划仅反映了提示所蕴含的布局;在生成过程中,模型追踪的是不断演化的几何状态,而非执行一个最初固定的规划。总体而言,这些结果表明,纯文本LLM的二维空间表现同时取决于模型本身和输出媒介,且无法仅由代码生成能力解释。
cs.AI / 163 / 2608.30757

Which Rules Matter Now? Policy-Centroid Routing Before an Intelligent System Acts

哪些规则当下重要?智能系统行动前的策略质心路由
Nguy, Thomson D.
Abstract
Before an intelligent system can decide whether an action is allowed, it must first know which rules the action has approached. A single proposed action can implicate several policy regimes at once. Their requirements may stack, overlap, or qualify one another, yet many remain written in natural language while the action itself arrives as an incomplete description of intent. The first problem is not judgment. It is attention. Policy-centroid routing creates a layer before adjudication. It compresses expressions within each policy regime into one or more representative centroids, places the proposed action in the same semantic space, applies a declared measure, and routes every regime crossing a declared threshold to authoritative review. Several regimes may trigger at once. The output is a review agenda, not permission, prohibition, legality, breach, compliance, certification, or enforcement. The paper develops six falsifiable propositions and seven follow-on studies comparing the hypothesis with structured workflows, lexical and semantic retrieval, hierarchical and direct classification, and selective prediction under matched review burden. The studies are designed to identify where policy geometry recovers applicable regimes, where compression loses rare or overlapping obligations, and where the mechanism should abstain. The paper includes a synthetic worked example and reports no empirical efficacy result.
Chinese Translation
在智能系统判定某个行动是否被允许之前,它首先必须知道该行动已接近哪些规则。单个拟议行动可能同时涉及多个策略体系,这些体系的要求可能叠加、重叠或相互限定,然而许多策略仍以自然语言书写,而行动本身则以对意图的不完整描述的形式到来。首要问题不是判断,而是注意力。策略质心路由在裁决之前创建一个层:它将每个策略体系内的表述压缩为一个或多个代表性质心,将拟议行动置于同一语义空间中,应用预先声明的度量方式,并将越过声明阈值的每个体系路由至权威审查。多个体系可能同时触发。其输出是一份审查议程,而非许可、禁止、合法性判定、违约认定、合规性、认证或执法结论。本文提出了六个可证伪的命题和七项后续研究,在匹配的审查负担下,将该假设与结构化工作流、词汇和语义检索、分层与直接分类以及选择性预测进行比较。这些研究旨在识别策略几何在何处能够恢复适用的体系、压缩在何处会丢失罕见或重叠的义务,以及该机制应在何处弃权。本文包含一个合成的演算示例,未报告任何实证效果结果。
cs.AI / 164 / 2608.30785

SkillZip Pro: Execution-Aware Dynamic Compression of Progressively Loaded Skills for Self-Evolving Agents

SkillZip Pro:面向自演化智能体的渐进式加载技能的执行感知动态压缩
Bai, Xiaofan, Liu, Chao, Lin, Hongqiang, Wu, Di, Song, Mingli, Jin, Xuan, Cao, Xipeng, Li, Yuhong
Abstract
Production agent skills are directory bundles, not isolated prompts. The root is loaded at activation; references, schemas, scripts, assets, and nested subskills are loaded only when an execution path needs them. Compressing only the root misses most deployment cost and may move branch-specific details into the always-loaded context. Flattening instead destroys progressive-loading boundaries. We introduce \method, an evaluation-free compressor for complete, progressively loaded skill bundles. It leaves the agent harness unchanged and emits an ordinary directory. The method combines two safeguards. First, it compresses \emph{across files}, removing content from a reference or subskill when the root or a declared environment contract already provides it. Second, it preserves routing, so every required file and directly callable entry remains reachable after rewriting. Users can configure \method along two independent axes. \emph{One-Shot} mode rebuilds the full bundle; \emph{Continual} mode reuses state and applies Zip-on-Write after each evolution patch. \emph{Persistent} compression rewrites the shipped bundle to reduce storage and runtime context. \emph{Transient} compression keeps that bundle byte-identical and builds a task-specific view, reducing only per-run context after build cost. Entry contracts mark private, public, and conditional resources; a multi-entry audit preserves standalone public subskills. On a production content-moderation skill evaluated by our industrial multi-round harness, \method removes \hl{38\%} of skill bundle tokens and \hl{10.4\%} of end-to-end per-run tokens with no quality loss, while an unprotected 71\% configuration loses up to 26 accuracy points to one-sided false positives. On a multi-entry bundle, \method effeciently reduces token cost while near-perfectly preserving every route and public entry.
Chinese Translation
生产环境中的智能体技能是目录包,而非孤立的提示词。根文件在激活时加载;引用、模式(schemas)、脚本、资源以及嵌套子技能仅在执行路径需要时才加载。仅压缩根文件无法覆盖大部分部署成本,且可能将分支特定的细节移入始终加载的上下文中。而扁平化处理则会破坏渐进式加载的边界。我们提出了 SkillZip Pro,一种面向完整渐进式加载技能包的无评估压缩器。它无需更改智能体运行框架,并输出一个普通的目录。该方法包含两项保障机制。第一,它进行跨文件压缩:当根文件或声明的环境契约已提供某内容时,从引用或子技能中移除该内容。第二,它保留路由,确保每个必需文件和可直接调用的入口在重写后仍然可达。用户可沿两个独立维度配置 SkillZip Pro。One-Shot(一次性)模式重建完整技能包;Continual(持续)模式复用状态,并在每次演化补丁后应用写时压缩(Zip-on-Write)。持久压缩重写交付的技能包以减少存储和运行时上下文;瞬态压缩保持该技能包字节不变,并构建任务特定视图,仅在支付构建成本后减少每次运行的上下文。入口契约标记私有、公共和条件性资源;多入口审计保留可独立运行的公共子技能。在由我们的工业级多轮评估框架评测的一个生产级内容审核技能上,SkillZip Pro 在无质量损失的前提下移除了 38% 的技能包 token 和 10.4% 的端到端每次运行 token;而无保护的 71% 压缩配置因单侧误报最多损失 26 个准确率点。在一个多入口技能包上,SkillZip Pro 在近乎完美保留所有路由和公共入口的同时,高效降低了 token 成本。
cs.AI / 165 / 2608.30841

HSRM: Hidden-State Reward Models for Test-Time Verification

HSRM:用于测试时验证的隐状态奖励模型
Li, Xianzhi, Zhu, Xiaodan
Abstract
Large language models can often generate plausible mathematical reasoning traces, but reliably identifying the correct solution among multiple candidates remains a key challenge. Existing test-time reasoning pipelines typically rely on text-based verifiers that re-read each generated solution, making verification an expensive component of inference. Prior work has shown, however, that LLMs often encode correctness-related signals in their internal representations, including awareness of when their own answers are likely to be wrong. Building on this observation, we introduce HSRM, a lightweight hidden-state reward model that verifies candidate solutions by directly reading the generator's internal representations rather than re-processing its text. HSRM extracts hidden states from a frozen generator at reasoning-step boundaries and uses a small Transformer encoder to rank candidates. It is trained from self-generated trajectories with outcome labels, requiring neither human-written process supervision nor a large pretrained verifier. Across four mathematical reasoning benchmarks, HSRM matches or outperforms a 55M-parameter text-only energy verifier in 15 of 16 generator--dataset settings while using only about 2M parameters, providing an efficient alternative to text-only verification by reusing representations already computed during generation.
Chinese Translation
大型语言模型通常能够生成看似合理的数学推理过程,但在多个候选答案中可靠地识别出正确解仍然是一个关键挑战。现有的测试时推理流程通常依赖于基于文本的验证器,需要重新阅读每个生成的解答,使得验证成为推理过程中开销较大的环节。然而,先前的研究表明,大型语言模型往往在其内部表示中编码了与正确性相关的信号,包括对自身答案何时可能出错的感知。基于这一观察,我们提出了HSRM,一种轻量级的隐状态奖励模型,它通过直接读取生成器的内部表示而非重新处理其文本来验证候选解答。HSRM在推理步骤边界处从冻结的生成器中提取隐状态,并使用一个小型Transformer编码器对候选答案进行排序。该模型由带结果标签的自生成轨迹训练而成,既不需要人工编写的过程监督,也不需要大型预训练验证器。在四个数学推理基准上,HSRM在16个生成器—数据集组合设置中的15个里匹配或超越了55M参数的纯文本能量验证器,而其参数量仅约2M。通过复用生成过程中已计算得到的表示,HSRM为纯文本验证提供了一种高效的替代方案。
cs.AI / 166 / 2608.30846

VFR-Audit: Verdict-Level Reliability for Fairness Audits in Hospital Length-of-Stay Prediction

VFR-Audit:面向医院住院时长预测公平性审计的裁决级可靠性
Joy, Md Jannatul Rakib, Vo, Viet, Chua, Caslon
Abstract
Fairness audits in clinical Artificial Intelligence convert continuous fairness metrics into binary pass-or-fail verdicts against operational thresholds, where hospital governance boards, payers, and regulators act on the resulting verdicts. Such audits are repeated over time and across hospital sites, thus the same verdict can flip between pass and fail across audits. Existing uncertainty methods such as Bayesian posteriors, bootstrap confidence intervals, and permutation tests address verdict instability only at the continuous-metric level. Converting metric-level uncertainty into a verdict-stability claim remains a manual step that scales poorly across the (model, metric, attribute) cells an audit covers. Existing uncertainty methods also leave open whether bias-mitigation steps, such as reweighing or per-group threshold shifts, yield a stable passing verdict at the cost of model discrimination measured as AUROC or AUPRC.To address this verdict-stability gap, we propose VFR-Audit, a framework built around the Verdict Flip Rate (VFR), a scalar bounded between 0 and 0.5 that measures the probability of verdict reversal under stratified bootstrap resampling. VFR-Audit reports VFR alongside three reliability axes, namely within-cohort resampling stability, audit-size sensitivity, and cross-hospital verdict agreement via Fleiss' kappa.
Chinese Translation
临床人工智能中的公平性审计将连续的公平性指标根据运行阈值转换为二元的通过与不通过裁决,医院治理委员会、支付方和监管机构依据这些裁决采取行动。此类审计会随时间推移并在不同医院站点重复进行,因此同一裁决可能在不同审计中在通过与不通过之间翻转。现有的不确定性方法,如贝叶斯后验、自助法置信区间和置换检验,仅在连续指标层面处理裁决的不稳定性。将指标层面的不确定性转换为裁决稳定性的结论仍是一个手动步骤,难以在审计所覆盖的(模型、指标、属性)单元格上规模化。此外,现有不确定性方法也未回答偏差缓解措施(如重加权或分组阈值调整)是否会以模型判别能力(以AUROC或AUPRC衡量)的下降为代价,产生一个稳定的通过裁决。针对这一裁决稳定性缺口,我们提出了VFR-Audit框架,其核心是裁决翻转率(Verdict Flip Rate, VFR),一个介于0到0.5之间的标量,用于衡量在分层自助重采样下裁决反转的概率。VFR-Audit在报告VFR的同时,还沿三个可靠性维度进行评估,即同队列内重采样稳定性、审计规模敏感性,以及通过Fleiss' kappa衡量的跨医院裁决一致性。
cs.AI / 167 / 2608.30865

Predicting Residential Rents in Dakar Using Machine Learning

基于机器学习的达喀尔住宅租金预测
Diallo, Amadou Tidiane Kassa
Abstract
Dakar's residential rental market remains poorly documented despite its economic and social importance: 54.4% of households are renters, compared to 23.3% nationally. This study develops a complete machine learning pipeline to predict residential rents in Dakar, from data collection to model interpretation. An original dataset of 1,507 rental listings was built through systematic web scraping and a documented cleaning pipeline, then enriched with four purpose-built features, including a luxury score and a keyword-based quality score. Five models were compared: linear regression, Random Forest (baseline), XGBoost, and LightGBM optimized through Bayesian optimization with Optuna, using leakage-free KFold target encoding for location. The optimized XGBoost model achieved the best performance with an $R^2$ of 0.847, an MAE of 210,902 XOF, and an RMSE of 324,195 XOF. Feature importance was assessed using native XGBoost gain and SHAP values, revealing a substantial difference in the ranking of location, which appears as a minor predictor by gain but as the second most influential variable by SHAP. This result carries methodological implications for hedonic studies using target-encoded categorical variables. This study provides an interpretable benchmark for Dakar's rental market and highlights several avenues for improvement, including the integration of geospatial features and conformal prediction.
Chinese Translation
尽管达喀尔住宅租赁市场具有重要的经济和社会意义,但其相关研究仍十分匮乏:54.4%的家庭为租房者,远高于全国平均水平23.3%。本研究构建了一套完整的机器学习流程来预测达喀尔的住宅租金,涵盖从数据采集到模型解释的全过程。研究通过系统性网页爬取和规范化清洗流程,构建了包含1,507条租赁信息的原创数据集,并添加了四个专门设计的特征进行丰富,其中包括奢华度评分(luxury score)和基于关键词的质量评分。研究比较了五种模型:线性回归、随机森林(基线模型)、XGBoost,以及通过Optuna贝叶斯优化调参的XGBoost和LightGBM,并对位置变量采用无泄漏的KFold目标编码。优化后的XGBoost模型取得了最佳性能,$R^2$为0.847,平均绝对误差(MAE)为210,902西非法郎(XOF),均方根误差(RMSE)为324,195西非法郎。特征重要性通过XGBoost原生的增益(gain)指标和SHAP值进行评估,结果揭示了位置变量在排序上的显著差异:按增益衡量其为较次要的预测因子,但按SHAP值衡量则为第二有影响力的变量。这一结果对使用目标编码分类变量的享乐价格研究具有重要的方法论启示。本研究为达喀尔租赁市场提供了一个可解释的基准,并指出了若干改进方向,包括地理空间特征的引入与保形预测(conformal prediction)的应用。
cs.AI / 168 / 2608.30897

CAER: Causal Action Effect Reweighting for World Model Training

CAER:面向世界模型训练的因果动作效应重加权方法
Fang, Jianjie, Liu, Xvyuan, Wang, Ziyou, Tang, Rongze, Wang, Zhaolu, Li, Zhuohang, Zhang, Xin, Su, Haisheng, Gao, Chen, Wu, Wei, Chen, Xinlei, Li, Yong
Abstract
World models are becoming core infrastructure for embodied intelligence, with action-conditioned video generation providing controllable predictions of how scenes evolve after agent interventions. Yet existing models are commonly trained with space-time-uniform mean squared error, allowing abundant background tokens to dominate the gradient while sparse interaction dynamics remain under-optimized; such uniform fitting rewards reconstructing appearance rather than learning how actions change the world. We introduce Causal Action Effect Reweighting (CAER), a general training paradigm that redistributes supervision toward the tokens whose predicted future is causally affected by the action. CAER contrasts the model's own predictions with and without action conditioning to localize these tokens online, then normalizes the resulting effect map into a weight that preserves the total coefficient mass and changes only where it is spent. This online signal requires no external annotations or offline preprocessing, avoids additional data-processing time, and scales naturally with model and dataset size. Experiments across heterogeneous action-conditioned world-model tasks show that CAER converges to better solutions than uniform MSE training, with consistent improvements in the physical consistency, controllability, and visual quality of generated videos.
Chinese Translation
世界模型正在成为具身智能的核心基础设施,其中基于动作条件的视频生成能够对智能体干预后场景的演化提供可控预测。然而,现有模型通常采用时空均匀的均方误差进行训练,使得大量背景标记主导了梯度,而稀疏的交互动态则得不到充分优化;这种均匀拟合鼓励模型重建外观,而非学习动作如何改变世界。我们提出了因果动作效应重加权(Causal Action Effect Reweighting, CAER),这是一种通用的训练范式,可将监督重新分配至其预测未来受动作因果影响的标记上。CAER 通过对比模型在有无动作条件下的自身预测来在线定位这些标记,随后将所得的效应图归一化为权重,该权重保持总系数质量不变,仅改变其分布位置。这一在线信号无需外部标注或离线预处理,避免了额外的数据处理时间,并能随模型和数据集规模自然扩展。在多种异构的动作条件世界模型任务上的实验表明,CAER 比均匀 MSE 训练收敛到更优的解,在生成视频的物理一致性、可控性和视觉质量方面均带来一致的提升。
cs.AI / 169 / 2608.30912

Responsible Integration of AI in Cancer Genomics: Barriers, Risks, and Pathways to Trustworthy Clinical Translation

人工智能在癌症基因组学中的负责任整合:障碍、风险与可信赖临床转化的路径
İlgen, Bahar, Tolias, Yiannos, Kühnert, Denise, Papadopoulou, Paraskevi, Westerlund, Magnus, Heider, Dominik, Ladewig, Katharina, Hattab, Georges
Abstract
Artificial intelligence (AI) and natural language processing (NLP) are increasingly used to extract, integrate, and interpret biomedical knowledge relevant to cancer genomics, yet their translation into routine clinical oncology has been comparatively slow. The central challenge is not computational capability alone, but trustworthy integration into clinical workflows. This review examines how NLP and AI support the cancer genomics pipeline, from literature mining and automated variant interpretation to clinical trial matching, knowledge graph construction, and multimodal data integration. We identify four interrelated translational failure domains: evidence inconsistency, explainability and uncertainty, data governance and reproducibility, and interoperability. Rather than considering these challenges in isolation, we take a systems-level view, focusing on their interaction across the translational pathway. We propose a conceptual framework and roadmap for addressing these domains through rigorous validation, uncertainty-aware methods, interoperable infrastructures, regulatory alignment, and human oversight across the AI lifecycle. Progress toward routine clinical use will depend less on further improving model capability than on systematically addressing these interacting failure domains from development through deployment and post-deployment monitoring.
Chinese Translation
人工智能(AI)和自然语言处理(NLP)正被越来越多地用于提取、整合和解读与癌症基因组学相关的生物医学知识,但它们向常规肿瘤临床实践的转化相对缓慢。核心挑战不仅在于计算能力本身,更在于如何可信赖地整合进临床工作流程。本综述考察了NLP和AI如何支持癌症基因组学的各个环节,包括文献挖掘、自动化变异解读、临床试验匹配、知识图谱构建以及多模态数据整合。我们识别出四个相互关联的转化失败领域:证据不一致性、可解释性与不确定性、数据治理与可重复性,以及互操作性。我们不将这些挑战孤立看待,而是采用系统级视角,重点关注它们在转化路径中的相互作用。我们提出了一个概念框架和路线图,通过严格验证、不确定性感知方法、可互操作的基础设施、监管协调以及覆盖AI全生命周期的人工监督来应对这些领域。实现常规临床应用的进展,与其说取决于进一步提升模型能力,不如说取决于系统性地解决这些相互作用的失败领域——涵盖从开发、部署到部署后监测的全过程。
cs.AI / 170 / 2608.30922

CARVE: Verified Expansion for Variable-Length Generation in Diffusion Language Models

CARVE:面向扩散语言模型变长生成的验证式扩展方法
Bouhedja, Wail, Mohamed, Amr, Shang, Guokan
Abstract
Masked diffusion language models predict tokens from a partially observed response canvas, enabling bidirectional conditioning and parallel token refinement. Yet standard masked-diffusion decoders use a rigid inference interface: the number of masked positions allocated to the answer is fixed before generation begins. Choosing this length is difficult. A short canvas can truncate reasoning or code, while a long canvas wastes computation and can perturb denoising. We introduce CARVE (Counterfactual-Aware Reveal with Verified Expansion), a training-free variable-length algorithm for masked diffusion LMs. Starting from a shorter canvas, CARVE can grow the response during decoding by inserting additional [MASK] positions. Rather than keeping every insertion, CARVE tests a candidate expanded canvas and asks a counterfactual question: would the model make similar predictions for the unresolved positions in the original canvas if the extra masked space were present? The inserted masks are kept only when they induce low Jensen-Shannon (JS) divergence on aligned unresolved positions. This makes length growth a verified stability decision rather than a pure confidence heuristic. CARVE applies without retraining to both full-canvas and blockwise diffusion decoders. Across code generation and mathematical reasoning benchmarks, CARVE consistently improves average performance over fixed-length baselines across all evaluated model families. Crucially, CARVE achieves these accuracy gains while reducing inference cost, reaching half the FLOPs of fixed-length decoding in some settings.
Chinese Translation
掩码扩散语言模型从部分可观测的回答画布(canvas)中预测词元,支持双向条件约束与并行词元精炼。然而,标准的掩码扩散解码器采用僵化的推理接口:分配给答案的掩码位置数量在生成开始之前就已固定。选择该长度十分困难:画布过短可能截断推理或代码,而画布过长则会浪费计算并可能干扰去噪过程。我们提出了 CARVE(Counterfactual-Aware Reveal with Verified Expansion,基于反事实感知与验证扩展的揭示方法),这是一种无需训练的掩码扩散大语言模型变长生成算法。CARVE 从较短的画布出发,可以在解码过程中通过插入额外的 [MASK] 位置来扩展回答。CARVE 并非保留每一次插入,而是测试候选的扩展画布并提出一个反事实问题:如果额外的掩码空间存在,模型对原始画布中未解析位置的预测是否会保持相似?只有当插入的掩码在对齐的未解析位置上引起较低的 Jensen-Shannon(JS)散度时,才会被保留。这使得长度扩展成为一种经过验证的稳定性决策,而非纯粹基于置信度的启发式方法。CARVE 无需再训练即可应用于全画布和分块扩散解码器。在代码生成和数学推理基准测试中,CARVE 在所有评估的模型系列上均持续优于固定长度基线的平均性能。尤为关键的是,CARVE 在提升准确率的同时降低了推理成本,在某些设置下仅达到固定长度解码一半的浮点运算量(FLOPs)。
cs.AI / 171 / 2608.30955

Learning Action Models with Conditional and Quantified Effects via Uncertainty-Guided Exploration

基于不确定性引导探索的条件化与量化效果动作模型学习
Jewett, Jeffrey, Solow, William, Saisubramanian, Sandhya
Abstract
Accurate action models are critical for effective planning. Existing action-model learning methods largely assume simple action representations or become computationally intractable when learning conditional and quantified effects. We present Online Hypothesis-Driven Conditional Action Model Learning (OHCAM), an online approach for learning action models with conditional and quantified effects from limited interactions with the environment. OHCAM maintains a belief over hypothesized action models and actively selects informative actions to reduce uncertainty by maximizing disagreement among competing hypotheses, while being robust to noisy observations. To enable scalability, OHCAM begins with a small set of simple action model hypotheses and expands to more complex conditions only when the current hypotheses become inconsistent with the data. Experiments on six benchmark planning domains demonstrate that OHCAM is sample efficient in learning action models that solve substantially more tasks than baselines, even with observation noise. We validate OHCAM on two tasks using a Kinova Gen3 robot, demonstrating the real-world applicability of our approach.
Chinese Translation
准确的动作模型对有效规划至关重要。现有的动作模型学习方法大多假设简单的动作表示,或在学习条件化与量化效果时在计算上变得难以处理。我们提出了在线假设驱动的条件化动作模型学习方法,用于在有限交互中学奥尔含有条件化与量化效果的动作模型。OHCAM 在假定的动作模型上维护信念,并通过最大化相互竞争假设之间的分歧来主动选择信息丰富的动作以降低不确定性,同时对含噪观测具有鲁棒性。为实现可扩展性,OHCAM 从一组简单的动作模型假设开始,仅当当前假设与数据不一致时才扩展到更复杂的条件。在六个基准规划领域上的实验表明,OHCAM 在样本效率方面表现优异,所学得的动作模型能解决远多于基线方法的任务,即使存在观测噪声亦然。我们使用 Kinova Gen3 机器人在两个任务上验证了 OHCAM,证明了该方法的现实世界适用性。
cs.AI / 172 / 2608.31022

MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents

MNIST-PRO:MNIST作为AI智能体的部分可观测世界回归
Toh, Vernon, Majumder, Navonil, Liu, Zhengyuan, Chen, Nancy F., Poria, Soujanya
Abstract
AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpretation capability because they introduce physical and control complexities. We address this with MNIST-PRO, a benchmark that isolates agentic perception by converting MNIST digit recognition into a sequential, glimpse-based search task with lookback constraints. We evaluate ten multimodal models across four memory representations, including raw visual history, textual states, structured metric grid maps, and a consolidated visual canvas. While models excel under full observability, partial observability exposes a clear performance gap. We identify three distinct bottlenecks. First, perceptual-state construction and interpretation present a challenge, as agents struggle to integrate fragmented glimpses. Second, agents often stop exploring before they see the full sequence. Third, models often fail to revise early, incorrect beliefs even when faced with subsequent contradictory evidence. These results show that simply acquiring visual evidence is not enough. Agents must also be able to build and update a reliable perceptual state.
Chinese Translation
在部分可观测环境中的AI智能体需要协调主动感知与工作记忆,以维持不断演化的感知状态。然而,现有基准测试由于引入了物理和控制层面的复杂性,难以单独隔离这种感知状态的构建与解释能力。为此,我们提出了MNIST-PRO,该基准通过将MNIST数字识别转化为具有回溯约束的顺序性、基于瞥视的搜索任务,从而隔离智能体感知能力。我们在四种记忆表示(包括原始视觉历史、文本状态、结构化度量网格地图以及整合的视觉画布)下评估了十个多模态模型。尽管模型在完全可观测条件下表现出色,但部分可观测性暴露出明显的性能差距。我们识别出三个不同的瓶颈。第一,感知状态的构建与解释构成挑战,智能体难以整合碎片化的瞥视信息。第二,智能体往往在观测完整序列之前就停止探索。第三,即使面对后续的矛盾证据,模型也常常无法修正早期错误的信念。这些结果表明,仅仅获取视觉证据是不够的,智能体还必须能够构建并更新可靠的感知状态。
cs.AI / 173 / 2608.31057

Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents

先度量再管理:评估编码智能体中的智能体工作记忆
Chen, Le, Wan, Zishen, Sun, Baixi, Ma, Xiaolong, Yang, Chih-Hsuan, Yan, Feng, Di, Sheng, Cappello, Franck, Thakur, Rajeev
Abstract
Agent working memory is heterogeneous. Objects such as instructions, artifacts, tool outputs, and agent-generated state play different semantic roles and exhibit different size, retention, and representation profiles. Recent work has begun to explore memory-management mechanisms that account for such heterogeneity. This work focuses on semantic heterogeneity and studies how it should shape the management and evaluation of working memory in coding agents. Across 55 archived coding-agent trajectories, we find that semantically different working-memory objects exhibit distinct retention and compression behavior. This heterogeneity motivates semantically informed memory management. We study two semantically informed strategies: an object-aware compression policy and a retrieval-based policy. Their evaluation shows that calibration gains may not transfer to held-out tasks, and that equal token budgets do not imply equal delivered context or management cost. A real-system replay further exposes serving limits that nominal budgets alone do not capture. Together, these results show why semantic structure matters for agent working memory and why evaluating memory-management strategies requires more than a nominal token budget. We organize these lessons into four levels: stored state, delivered context, management work, and task or process outcome.
Chinese Translation
智能体工作记忆是异构的。指令、工件、工具输出和智能体生成的状态等对象扮演着不同的语义角色,并在大小、保留方式和表示形式上呈现出不同的特征。近期的研究已开始探索能够考虑这种异构性的记忆管理机制。本文聚焦于语义异构性,研究它应当如何塑造编码智能体中工作记忆的管理与评估。通过分析55条存档的编码智能体轨迹,我们发现语义上不同的工作记忆对象表现出不同的保留与压缩行为。这种异构性启发我们采用语义感知的记忆管理。我们研究了两种语义感知策略:对象感知的压缩策略和基于检索的策略。评估结果表明,校准带来的收益可能无法迁移到留存任务上,且相同的token预算并不意味着传递给模型的有效上下文或管理成本也相同。真实系统的重放实验进一步揭示了仅凭名义预算无法捕捉的服务限制。综上所述,这些结果说明了为什么语义结构对智能体工作记忆至关重要,以及为什么评估记忆管理策略不能仅依赖于名义token预算。我们将这些经验教训归纳为四个层面:存储状态、传递上下文、管理工作以及任务或过程结果。
cs.AI / 174 / 2608.31068

Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores

错误的预测,正确的答案:从坍缩的大语言模型序列得分中恢复证据
Yan, Qiyao, Wang, Chenpeng, Pan, Liangming
Abstract
When a large language model fails a reasoning task, it is often assumed to lack the underlying capability. However, this conflates a genuine absence of reasoning with a late-stage output bottleneck. We observe a consistent readout gap across diverse reasoning benchmarks: hidden-state probes successfully decode correct answers even when native sequence scoring completely collapses due to structural biases. To test whether instance-specific logic survives this collapse, we introduce a diagnostic protocol using a minimal, target-label-free additive correction. Fitting just two parameters on as few as 25 unlabeled examples recovers 9--34 accuracy points for Qwen3.5 models, transferring successfully to OLMo-2-1B and Llama-3.1-8B. Crucially, these recovered decisions persist on hard instances unresolved by simple lexical overlap and significantly exceed count-preserving permutation baselines. Our results show that many apparent zero-shot reasoning deficits are expression failures masking intact internal logic, urging a narrower interpretation of benchmark evaluations.
Chinese Translation
当大语言模型在推理任务中失败时,人们通常认为它缺乏相应的基础能力。然而,这混淆了推理能力的真正缺失与输出后期的瓶颈问题。我们在多个推理基准上观察到一致的性能读出差距:即使原生序列得分因结构性偏差而完全失效,隐藏状态探针仍能成功解码出正确答案。为了检验实例特定的逻辑是否能在这种失效中保留下来,我们引入了一种诊断协议,该协议使用一种极简的、无需目标标签的加性校正。仅需在25个无标签样本上拟合两个参数,Qwen3.5模型的准确率即可恢复9–34个百分点,并能成功迁移至OLMo-2-1B和Llama-3.1-8B模型。至关重要的是,这些恢复出的决策在简单词汇重叠无法解决的困难实例上依然有效,并显著优于保持计数的置换基线。我们的结果表明,许多表面上的零样本推理缺陷实际上是被完整内部逻辑掩盖的表达失败,这提醒我们应以更审慎的态度解读基准评估结果。
cs.AI / 175 / 2608.31075

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

超越人类监督的大规模推理模型扩展:通往超级智能之路
Yang, Zhiqin, Fu, Jingwen, Liu, Yuhan, Liu, Hengyu, Zhang, Yonggang, Cao, Kainan, Zhang, Zizhuo, Li, Chenxin, Yuan, Ruibin, Pan, Jiahao, Sun, Jiankai, Zhang, Zhenyuan, Li, Yibo, Lin, Yunlong, Xiong, Jing, Lin, Sida, Han, Bo, Xue, Wei, Guo, Yike
Abstract
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated \href{https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision}{GitHub repository} to track the latest advances.
Chinese Translation
大规模推理模型(Large Reasoning Models, LRMs)的最新进展表明,可验证奖励强化学习(RLVR)能够显著提升数学和代码领域的推理能力,因为这些领域的输出结果可以自动验证。然而,将这一进展扩展到开放式和智能体任务仍然十分困难,原因在于可靠的奖励更难获得,且直接的人类监督无法跟上模型生成经验在规模和复杂性上的增长。本文研究了当人类监督逐渐退出学习回路时,LRM 如何能够持续改进。我们从两个相互关联的维度考察这一问题:奖励维度梳理了从针对单个实例的人类判断,发展到可复用的验证器和奖励,直至无需人类反馈也能运作的奖励机制;经验维度则考察了学习如何从人工策划的任务和环境,逐步走向自主生成的课程体系、自构建的环境以及自主协同进化。我们通过一个从 L0 到 L4 的五级阶梯将这两个维度联系起来,用以识别学习过程中哪些部分仍处于人类的持续控制之下。我们的分析进一步强调了日益自主的奖励和经验生成所带来的风险,包括奖励破解、反馈漂移、课程崩塌和环境错误。因此,我们还围绕三个互补的对象提供了评估框架:策略能力、反馈保真度和经验质量。这一分析为当前超越人类监督扩展 LRM 的方法,以及发展面向超级智能的自持学习系统所涉及的开放性问题,提供了系统性的阐述。此外,我们维护了一个持续更新的 GitHub 仓库,以追踪最新进展。
cs.AI / 176 / 2608.31077

Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

在智能体策略优化中调和过程监督与基于结果的信用分配
Yang, Jingxiao, Gan, Wangjie, Zhuang, Yingxuan, Zhang, Wenqi, Chen, Jintao, Zhang, Xuhong
Abstract
Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit over long-horizon interactions. On-policy self-distillation offers finer supervision by re-evaluating sampled behavior with privileged information (PI) available only during training. However, fine-grained supervision is not necessarily fine-grained credit: PI-induced likelihood changes describe how additional information alters policy preference, but do not directly determine how an executable action should inherit the verified task outcome. This creates a supervision-credit gap. Privileged signals may be irrelevant to the current interaction state, operate at a token granularity misaligned with executable decisions, and lack the outcome semantics required for reinforcement. We introduce TASPO, which converts privileged supervision into outcome-grounded action credit. TASPO constructs decision-applicable PI from verified successful experience, aggregates PI-induced likelihood shifts at the executable-action level, and converts relative action support into positive, bounded, mean-preserving weights on the original trajectory advantage. Thus, the verified outcome determines the update direction and average scale, while PI only redistributes credit across actions. Across three agentic benchmarks, TASPO improves over GRPO by 10.6\% and generalizes better to unseen tasks. Further analysis indicates that TASPO reduces supervision mismatch and that action-level assignment stabilizes the policy optimization process. These findings offer the community another interesting perspective.
Chinese Translation
基于结果的强化学习为语言模型智能体提供了经过验证的反馈,但其将轨迹级的优势统一分配给所有决策,导致在长时程交互中信用分配过于粗糙。在策略自蒸馏通过利用仅在训练期间可用的特权信息(PI)对采样行为进行重新评估,从而提供更精细的监督。然而,细粒度的监督并不等同于细粒度的信用分配:特权信息引起的似然变化描述了附加信息如何改变策略偏好,但并不能直接决定一个可执行动作应如何继承经过验证的任务结果。这产生了监督与信用分配之间的鸿沟。特权信号可能与当前交互状态无关,其作用的词元粒度与可执行决策不对齐,并且缺乏强化学习所需的结果语义。我们提出了TASPO,将特权监督转化为基于结果的动作信用。TASPO从经过验证的成功经验中构建可决策适用的特权信息,在可执行动作层面聚合特权信息引起的似然偏移,并将相对动作支持度转换为原始轨迹优势上正向、有界且保持均值的权重。这样,经过验证的结果决定了更新的方向和平均尺度,而特权信息仅在各动作之间重新分配信用。在三个智能体基准测试中,TASPO相比GRPO提升了10.6%,并对未见任务具有更好的泛化能力。进一步分析表明,TASPO减少了监督失配问题,且动作级的信用分配稳定了策略优化过程。这些发现为社区提供了另一个有趣的视角。
cs.AI / 177 / 2608.31082

Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

智能体式数据破解(Agentic Context Cracking):通过非结构化数据的自适应结构化实现高Token效率的数据推理智能体
Hajidehi, Milad Rezaei, Wang, Qitong, Idreos, Stratos
Abstract
Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI is deploying LLM agents that reason over this data to answer complex questions for every knowledge worker. Agents can do this today, but at prohibitive cost. Each question repeatedly opens large documents to recover scattered evidence, consuming up to a million tokens. However, if the data were already structured, the same question would reduce to a cheap database lookup. For example, on FanOutQA benchmark, reasoning over an ideal pre-structured store is 28X cheaper, and the gap grows to orders of magnitude as questions fan out over more documents. Yet structuring everything in advance is not viable: documents hold vastly more possible structure than any workload will use, and the useful structure and documents are unknown until queries arrive. We propose agentic data cracking, a method that structures unstructured data adaptively and speculatively as a byproduct of reasoning itself. Structuring is adaptive because observed queries decide when it happens and what matters, and speculative because it goes beyond the current question. Whenever the agent opens a document to answer, a cracking sub-agent forks from the already-loaded context at marginal cost and extracts grounded structure likely to serve related future queries. Over time, an increasing share of queries is fully covered by structured data and answered without opening a document, keeping agentic accuracy at close to RAG cost. On FanOutQA, extended with merely one related question per test question, cracking cuts cost by 53% while preserving accuracy. Agentic data cracking is a first step toward next-generation data infrastructure for agentic reasoning over unstructured data: a shared substrate beneath the model where knowledge that reasoning already paid to uncover accumulates.
Chinese Translation
有价值的数据往往埋藏于非结构化来源之中:网页、报告、合同、申报文件、财报电话会议记录以及PDF文档。企业AI的重大押注在于部署能够对这些数据进行推理的大语言模型(LLM)智能体,为每一位知识工作者回答复杂问题。智能体如今已能实现这一点,但成本高昂到难以承受。每个问题都需要反复打开大型文档以找回分散的证据,最多可消耗一百万个token。然而,如果数据已经被结构化,同样的问题就可以简化为一次廉价的数据库查询。例如,在FanOutQA基准上,对一个理想的预先结构化的数据存储进行推理的成本要低28倍,并且随着问题在更多文档上展开,这一差距会扩大到数个数量级。然而,预先将所有数据结构化并不可行:文档所蕴含的可能结构远超任何工作负载所能使用的范围,而且在查询到来之前,有用的结构和文档都是未知的。我们提出智能体式数据破解(agentic data cracking),一种在推理过程中自适应地、前瞻性地对非结构化数据进行结构化的方法,将结构化作为推理本身的副产品。之所以说是自适应的,是因为观察到的查询决定了何时进行结构化以及哪些内容重要;之所以说是前瞻性的,是因为它超越了当前的问题。每当智能体为回答问题而打开一个文档时,一个破解子智能体(cracking sub-agent)会以边际成本从已加载的上下文中分叉出来,提取可能服务于相关未来查询的、有据可依的结构。随着时间推移,越来越多的查询可以被结构化数据完全覆盖,无需打开文档即可回答,从而在保持智能体准确率的同时将成本降至接近RAG的水平。在FanOutQA基准上(我们仅为每个测试问题扩展了一个相关问题),破解方法在保持准确率的同时将成本降低了53%。智能体式数据破解是迈向下一代面向非结构化数据智能体推理的数据基础设施的第一步:它在模型之下构建一个共享基座,使推理已经付出代价所挖掘出的知识得以不断积累。
cs.AI / 178 / 2608.31097

Cross-Regional Grapevine Cold Hardiness Prediction via Learned Multimodal Latent Representations

基于学习多模态潜在表征的跨区域葡萄抗寒性预测
Solow, William, Pesantez-Cabrera, Paola, Keller, Markus, Khot, Lav, Saisubramanian, Sandhya, Fern, Alan
Abstract
Accurate daily predictions of cold hardiness in woody plants are critical in regions where freezing temperatures can damage dormant buds and reduce seasonal yield. Existing biophysical, hybrid, and deep learning models have shown high predictive accuracy when trained on local data but remain largely site-specific. The limited availability of cold hardiness data, coupled with the lack of principled methods for transferring cold hardiness predictions to new regions and cultivars, has limited the broader adoption and practical utility of these approaches, particularly in data-scarce regions. To address these limitations, we propose a cold hardiness prediction framework that learns a transferable latent representation by capturing region-specific variation through learned embeddings. To enable prediction in previously unseen regions, we infer embeddings from (1) text descriptions of the cultivar and growing region, and (2) limited historical observations, supporting both zero-shot and few-shot transfer. Experiments on datasets from six regions across North America demonstrate that our approach consistently outperforms state-of-the-art cold hardiness prediction methods, yielding more accurate predictions and substantially improving transfer to data-scarce regions.
Chinese Translation
在冰冻温度可能损害休眠芽并降低季节性产量的地区,对木本植物抗寒性进行准确的逐日预测至关重要。现有的生物物理模型、混合模型和深度学习模型在基于本地数据训练时已展现出较高的预测精度,但在很大程度上仍局限于特定站点。抗寒性数据的有限可得性,加之缺乏将抗寒性预测迁移到新区域和新品种的原则性方法,限制了这些方法的更广泛应用和实用价值,尤其是在数据稀缺的地区。为解决这些局限,我们提出了一种抗寒性预测框架,通过学习到的嵌入(embeddings)来捕捉区域特定的变异,从而学习可迁移的潜在表征。为了在未见过的区域实现预测,我们从(1)品种和种植区域的文本描述以及(2)有限的历史观测数据中推断嵌入,从而支持零样本(zero-shot)和少样本(few-shot)迁移。在来自北美六个地区的数据集上的实验表明,我们的方法持续优于最先进的抗寒性预测方法,产生了更精确的预测结果,并显著改善了对数据稀缺地区的迁移能力。
cs.AI / 179 / 2608.31105

BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing

BLOOM-WILT:面向自动化大语言模型审计的行为诱导Logit倾斜方法
Skapars, Adrians, Manino, Edoardo
Abstract
Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deployment puts the model through orders of magnitude more interactions than any evaluation can simulate. Automated auditors make testing cheap to scale and flexible enough to cover almost any specified behaviour, yet their lack of optimisation pressure makes them sample-inefficient. To address this shortcoming, we introduce BLOOM-WILT, a full auditing pipeline that elicits natural multi-turn instances of rare behaviours, without training cost or access beyond the target's next-token distribution. On the input side, WILT's auditor model revises its conversational strategy across rounds, learning from previous scored interactions. On the output side, WILT adaptively reweights the target's decoding using the model's own distribution conditioned on an elicitation prompt, so that behaviour-relevant generations are sampled ahead of others it finds equally probable when unprompted. We evaluate WILT across 4 target models and 8 behaviours, where it beats the baseline auditor in 30 of the 32 settings and overturns the previous model safety rankings. WILT raises average behaviour presence from 51% to 100% when eliciting self-harm encouragement from Qwen3.5-4B, beating every elicitation method we port into the same pipeline at matched compute, without pushing output probability below the baseline's.
Chinese Translation
已部署语言模型的用户经常会遇到各种测试几乎无法触及的行为,因为部署环境使模型经历的交互次数比任何评估所能模拟的都要多出几个数量级。自动化审计工具使测试的扩展成本足够低,且足够灵活以覆盖几乎所有指定的行为,但由于缺乏优化压力,其采样效率低下。为解决这一不足,我们提出了BLOOM-WILT,一个完整的审计流程,它能够诱导出稀疏行为(rare behaviours)的自然多轮对话实例,且无需训练成本,也无需超越目标模型下一个词分布之外的访问权限。在输入端,WILT的审计模型跨多轮对话迭代修正其对话策略,并从以往经过评分的交互中学习。在输出端,WILT利用目标模型自身在诱导提示条件下的输出分布,自适应地重新加权其解码过程,使与目标行为相关的生成内容先于其他在无提示时概率相同的候选被采样到。我们在4个目标模型和8种行为上评估了WILT,它在32个设置中的30个超越了基线审计器,并推翻了此前模型的排名。在诱导Qwen3.5-4B产生鼓励自残行为时,WILT将行为出现率从51%提升至100%,在同等计算量下超越了所有移植到同一流程中的诱导方法,且未将输出概率压低至基线水平以下。
cs.AI / 180 / 2608.31118

When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning

更大何时才有帮助?关于大语言模型规模对本体学习影响的受控研究
Giglou, Hamed Babaei, Auer, Sören, D'Souza, Jennifer
Abstract
The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled evaluation of 13 models spanning dense and Mixture-of-Experts variants from the Qwen3.5 and Qwen3.6 lineages, together with proprietary GPT release variants, using the OntoLearner retrieval-augmented generation pipeline. All models are evaluated with the same embedding model, retrieval configuration, prompt templates, decoding settings, datasets, and metrics on term typing, taxonomy discovery, and non-taxonomic relationship extraction across four biomedical and materials science and engineering ontologies. Within the dense Qwen3.5 lineage, increasing parameter count primarily improves precision rather than recall, with the largest gains occurring between 9B and 27B parameters. However, the effect of scale is neither monotonic nor uniform across tasks and domains. Dense 27B models outperform substantially larger sparse models on term typing, whereas larger Mixture-of-Experts models achieve the strongest open-weight results on taxonomy discovery. Non-taxonomic relationship extraction remains difficult across model scales, particularly for the Materials Data Science ontology. Performance differences across matched Qwen variants and proprietary GPT releases further indicate that architecture and model lineage can outweigh nominal parameter count. These findings show that model size alone is an insufficient selection criterion for OL and provide empirical guidance for reproducible LLM-assisted ontology engineering.
Chinese Translation
大语言模型(LLM)规模对本体学习(OL)性能的影响尚未得到充分刻画。我们对13个模型进行了受控评估,这些模型涵盖来自Qwen3.5和Qwen3.6系列的稠密模型和混合专家(Mixture-of-Experts)变体,以及专有的GPT发布版本,并采用OntoLearner检索增强生成流水线。所有模型均在相同的嵌入模型、检索配置、提示模板、解码设置、数据集和评价指标下,针对四个生物医学以及材料科学与工程领域的本体进行术语类型判定、分类体系发现和非分类关系抽取任务的评估。在稠密的Qwen3.5系列内,增加参数量主要提升精确率而非召回率,其中最大的提升出现在9B至27B参数之间。然而,规模效应在不同任务和领域中既非单调也非一致。稠密的27B模型在术语类型判定任务上优于规模大得多的稀疏模型,而更大的混合专家模型则在分类体系发现任务上取得了最强的开源权重模型结果。非分类关系抽取在各个模型规模下仍然困难,尤其是在材料数据科学(Materials Data Science)本体上。匹配的Qwen变体与专有GPT发布版本之间的性能差异进一步表明,架构和模型谱系的重要性可能超过名义参数量。这些发现表明,仅凭模型大小不足以作为本体学习中的选型标准,并为可复现的LLM辅助本体工程提供了实证指导。
cs.AI / 181 / 2608.31137

OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques

OntoAligner-Ensemble:基于投票的异构本体对齐技术融合方法
Giglou, Hamed Babaei, Auer, Sören, Popov, Peio, Sanaei, Mahsa, D'Souza, Jennifer
Abstract
Ontology alignment (OA) has evolved through several methodological paradigms, ranging from lexical and structural aligners to knowledge graph embedding (KGE) models and, more recently, Large Language Model (LLM)-based approaches. Although modern OA frameworks provide unified ecosystems for deploying these heterogeneous aligners, mechanisms for systematically reconciling their complementary and sometimes conflicting predictions remain relatively underexplored. We present OntoAligner-Ensemble, a modular and aligner-agnostic framework that combines candidate correspondences through a configurable two-stage process comprising voting-based fusion strategies followed by post-fusion selection policies. The framework supports any aligner implemented within OntoAligner that produces candidate correspondences, enabling diverse alignment paradigms to be integrated through a unified decision process. To demonstrate its effectiveness, we instantiate the framework using representative lightweight string-aligner, KGE-based, and Retrieval-Augmented Generation aligners powered by both open-weight and API-based LLMs. We evaluate individual aligners and ensemble configurations across eight benchmark tasks from five OAEI tracks spanning biomedical to beyond-equivalence. The results show that ensemble fusion consistently improves the balance between precision and recall and frequently outperforms standalone aligners across diverse domains. Furthermore, our analysis reveals that ensemble composition directly affects the precision-recall trade-off: heterogeneous cross-paradigm ensembles generally improve precision, whereas homogeneous LLM ensembles more often achieve higher overall F1-scores. These findings demonstrate that systematic ensemble learning offers a robust and reproducible strategy for OA while providing practical guidance for selecting ensemble compositions under different alignment scenarios.
Chinese Translation
本体对齐(Ontology Alignment, OA)已经历了多种方法范式的演进,从词汇和结构对齐器到知识图谱嵌入(KGE)模型,再到近来基于大语言模型(LLM)的方法。尽管现代OA框架为部署这些异构对齐器提供了统一的生态系统,但系统性地协调它们互补且有时相互冲突的预测结果的机制仍相对缺乏探索。我们提出了OntoAligner-Ensemble,这是一个模块化且与对齐器无关的框架,通过一个可配置的两阶段过程来组合候选对应关系,该过程包括基于投票的融合策略以及融合后的选择策略。该框架支持在OntoAligner中实现的任何产生候选对应关系的对齐器,使多种对齐范式能够通过统一的决策过程进行集成。为验证其有效性,我们使用代表性的轻量级字符串对齐器、基于KGE的对齐器,以及由开源权重和基于API的LLM驱动的检索增强生成(RAG)对齐器来实例化该框架。我们在来自五个OAEI赛道(涵盖生物医学及超越等价关系领域)的八个基准任务上评估了单个对齐器和集成配置。结果表明,集成融合持续改善了精确率与召回率之间的平衡,并且在多个领域中频繁优于单一对齐器。此外,我们的分析揭示了集成组成直接影响精确率-召回率权衡:异构跨范式集成通常能提升精确率,而同构LLM集成则更常获得更高的整体F1分数。这些发现表明,系统化的集成学习为OA提供了一种稳健且可复现的策略,同时为不同对齐场景下选择集成组成提供了实用指导。
机器学习 (Machine Learning)
185
cs.LG / 1 / 2608.28771

ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning

ERR+:面向高效且果断的大语言模型推理的序列化熵消解方法
Jiang, Xin, Wang, Minhao, Wu, Wen, Xie, Zhentao, Du, Shangheng, Shi, Jinxin, Zhao, Jiabao
Abstract
Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning with verifiable rewards (RLVR). While current RLVR methods have achieved strong results with correctness-based reward signals, they provide limited guidance on the quality of the reasoning process itself, leaving the internal reasoning structure largely unoptimized. Through empirical analysis across multiple model families, we identify a consistent pattern: correct reasoning trac es exhibit more frequent and larger token-level entropy drops within the thinking phase than incorrect ones. We propose ERR+, a two-phase RLVR framework grounded in this observation. The first phase trains with the Entropy Relief Reward (ERR), a bonus proportional to cumulative token-level entropy drops in the thinking phase, log-normalized by response length. Unlike prior methods that suppress entropy, ERR rewards the resolution of uncertainty while leaving exploratory high-entropy states unconstrained. The second phase introduces the Robust Relative Efficiency Reward, which scores each response's length against co-generated peers via a $\tanh$-transformed within-group $z$-score. We provide a formal analysis showing that joint optimization of the two objectives induces gradient conflict in early training, motivating the sequential design . Experiments on five datasets demonstrate consistent improvements in both accuracy and response conciseness across model backbones. Our code is available at https://github.com/XrkArul/err_response
Chinese Translation
大型推理模型通过基于可验证奖励的强化学习(RLVR)生成扩展的思维链(CoT)轨迹,在复杂任务上取得了出色的性能。尽管当前的RLVR方法借助基于正确性的奖励信号取得了显著成果,但它们对推理过程本身的质量提供的引导有限,使得内部推理结构基本未被优化。通过对多个模型系列的实证分析,我们发现了一个一致的模式:正确的推理轨迹在思考阶段表现出比错误轨迹更频繁、更大幅度的token级熵下降。基于这一观察,我们提出了ERR+,一个两阶段的RLVR框架。第一阶段使用熵缓解奖励(Entropy Relief Reward,ERR)进行训练,该奖励与思考阶段累积的token级熵下降成正比,并按响应长度进行对数归一化。与先前抑制熵的方法不同,ERR奖励不确定性的消解,同时不对探索性的高熵状态施加约束。第二阶段引入了鲁棒相对效率奖励(Robust Relative Efficiency Reward),通过经tanh变换的组内z分数,将每个响应的长度与共同生成的同类响应进行比较评分。我们提供了形式化分析,表明两个目标的联合优化会在训练早期引发梯度冲突,从而启发了这种序列化设计。在五个数据集上的实验表明,该方法在不同模型骨干上均能持续提升准确率并使响应更加简洁。我们的代码已发布于 https://github.com/XrkArul/err_response
cs.LG / 2 / 2608.28840

Unsupervised Latent Space Alignment with Hyperspherical Geodesic Matching

基于超球面测地线匹配的无监督潜在空间对齐
Ryan, Cameron, Narayanaswamy, Vivek Sivaraman, Thopalli, Kowshik, Liu, Shusen
Abstract
Independently trained neural networks tend to encode the same data with similar latent geometries. These latent geometries are not directly compatible, yet they can be nearly the same up to some class of transformations. While there exists many methods for alignment between different latent spaces, it is typically done using a set of shared sample correspondences, known as anchors. This leaves a fundamental question: are the geometric signatures of different latent spaces representing similar data sufficient to recover an alignment between them? To that end, we introduce HGA (Hyperspherical Gaussian Alignment), a method that directly optimizes a transformation between two latent spaces by maximizing a geometric measure of "fit" between them. Since it is driven by the geometry of the latent spaces rather than paired data, HGA can operate in both an unsupervised and weakly supervised regime. On tasks such as model stitching or multilingual word embedding correspondence recovery, HGA manages to match supervised results with minimal or no supervision.
Chinese Translation
独立训练的神经网络在编码相同数据时往往具有相似的潜在几何结构。这些潜在几何结构彼此并不直接兼容,但它们在某些变换类别下可能几乎相同。尽管存在许多潜在空间之间的对齐方法,但通常需要借助一组共享的样本对应关系,即所谓的锚点。这就引出了一个根本性问题:表示相似数据的不同潜在空间的几何特征是否足以恢复它们之间的对齐关系?为此,我们提出了HGA(Hyperspherical Gaussian Alignment,超球面高斯对齐),该方法通过最大化两个潜在空间之间几何“匹配度”的度量,直接优化两者之间的变换。由于HGA由潜在空间的几何结构驱动而非依赖成对数据,它可以在无监督和弱监督两种模式下运行。在模型拼接(model stitching)或多语言词嵌入对应关系恢复等任务上,HGA能够在极少监督或无监督的条件下达到与监督方法相当的结果。
cs.LG / 3 / 2608.28843

Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks

平滑Transformer前馈网络的曲率密码分析
Hasan, Munawar, Vassilev, Apostol
Abstract
We show that smooth two-layer feed-forward networks (FFNs) expose an additional structural model extraction channel under a chosen-input raw-output oracle at the FFN branch; consider transformer FFN branches with GELU or SiLU activations under chosen-input raw-output access, without access to parameters, gradients, or internal activations; exploit a second-order leakage channel in which projected input Hessians form different mixtures of the same hidden symmetric rank-one factors induced by the FFN input weights. We formalize resulting Hessian collection as a partially symmetric decomposition to establish conditions for local identifiability and stability to exploit vector-output stencil reuse to reduce the structural query cost by a factor of 16. On independently trained CIFAR-10 vision transformers, only 16 projected Hessians, corresponding to 8193 black-box queries, recover the hidden FFN directions with average absolute cosine alignment above 0.94, with 95.1 % of GELU and 91.9 % of SiLU directions exceeding 0.90 alignment. Recovery remains high across independently trained models, repeated extraction runs, and all transformer blocks. The recovered structure supports functional extraction too. Keeping the recovered directions fixed and fitting only the remaining FFN parameters yields high-fidelity substitutes with more than 93 % top-1 agreement, while test accuracy remains within 0.90% and 0.62% of the GELU and SiLU targets. Output rounding and Gaussian noise substantially reduce recovery under a fixed attack configuration, but adapting the finite-difference step restores average alignment to 0.9603 and 0.9398. This is an end-to-end path from black-box second-order observations to hidden FFN-structure recovery and functional replacement. Under the stated oracle model, smooth FFN curvature exposes internal parameter geometry that behavioral fidelity alone cannot reveal.
Chinese Translation
我们证明了平滑的两层前馈网络(FFN)在FFN分支处的选择输入-原始输出预言机下暴露出一条额外的结构性模型提取通道。我们考虑在具有GELU或SiLU激活函数的Transformer FFN分支上进行选择输入-原始输出访问,且不访问参数、梯度或内部激活。我们利用了一条二阶泄露通道:投影输入Hessian矩阵构成了由FFN输入权重所诱导的相同隐藏对称秩一因子的不同混合。我们将由此产生的Hessian收集形式化为部分对称分解,以建立局部可辨识性和稳定性的条件,并利用向量输出模板复用将结构性查询成本降低16倍。在独立训练的CIFAR-10视觉Transformer上,仅需16个投影Hessian(对应8193次黑盒查询)即可恢复隐藏的FFN方向,平均绝对余弦对齐度高于0.94,其中95.1%的GELU方向和91.9%的SiLU方向的对齐度超过0.90。在独立训练的模型、重复的提取运行以及所有Transformer块中,恢复效果均保持较高水平。恢复出的结构同样支持功能性提取。固定恢复出的方向并仅拟合剩余的FFN参数,可得到高保真替代模型,其top-1一致率超过93%,同时测试准确率与GELU和SiLU目标模型分别相差不超过0.90%和0.62%。在固定攻击配置下,输出舍入和高斯噪声会显著降低恢复效果,但通过自适应调整有限差分步长,可将平均对齐度恢复至0.9603和0.9398。这是一条从黑盒二阶观测到隐藏FFN结构恢复及功能性替换的端到端路径。在所述的预言机模型下,平滑FFN的曲率暴露了仅凭行为保真度无法揭示的内部参数几何结构。
cs.LG / 4 / 2608.28853

Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs

等变层束神经网络:图上的几何传输学习
Borgi, Alessio, Severino, Mario, Silvestri, Fabrizio, Liò, Pietro
Abstract
Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-order architectures remain limited in how vector information can be transformed as it moves across a graph. We introduce \textsc{ESNN}, an Equivariant Sheaf Neural Network that enriches this interaction by learning directed, matrix-valued transport between neighboring vector features while preserving exact Euclidean equivariance. Rather than increasing the order of the representation, ESNN keeps scalar and vector features first-order and places the additional geometric flexibility in the edge transport itself. We characterize this transport theoretically, showing that when relative displacement is the only covariant geometric input, every linear $O(n)$-equivariant map decomposes into independent radial and tangential components, while learned covariant features enable richer feature-conditioned transformations. We also introduce controlled symmetry relaxation for systems with a preferred ambient direction, which may be prescribed or inferred from data while recovering full $E(n)$-equivariance when the directional pathway is inactive. Across particle dynamics, mesh-based simulation, point-cloud classification, and molecular property prediction, ESNN improves dynamics prediction, recovers the gravity axis when symmetry is broken, yields substantial gains on selected mesh tasks and long-horizon rollouts, and remains robust to unseen rotations. These results show that learning how geometric information is transported across edges offers a complementary route to expressive equivariant message passing without requiring higher-order representations.
Chinese Translation
等变图神经网络为几何系统建模提供了一种有原则的方法,但高效的一阶架构在向量信息沿图移动时的变换方式上仍然受限。我们提出了等变层束神经网络(Equivariant Sheaf Neural Network, ESNN),它在保持严格欧几里得等变性的同时,通过在相邻向量特征之间学习有向的矩阵值传输来丰富这种交互。ESNN 并不提高表示的阶数,而是将标量与向量特征保持在一阶,并将额外的几何灵活性置于边传输本身。我们在理论上对这种传输进行了刻画,表明当相对位移是唯一的协变几何输入时,每个线性 $O(n)$-等变映射都可分解为独立的径向分量和切向分量;而学习到的协变特征则可以实现更丰富的特征条件化变换。我们还为具有优先环境方向的系统引入了受控对称性松弛,其方向可由预设或从数据中推断,并在方向通路失效时恢复完整的 $E(n)$-等变性。在粒子动力学、基于网格的仿真、点云分类和分子性质预测等任务上,ESNN 改进了动力学预测,在对称性被破坏时能够恢复重力轴,在部分网格任务和长时程推演上取得了显著提升,并对未见过的旋转保持鲁棒性。这些结果表明,学习几何信息如何沿边传输,为构建具有强表达能力的等变消息传递提供了一条无需高阶表示的互补途径。
cs.LG / 5 / 2608.28859

The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

停止向量:将因果引导干预内化以实现高效推理
Jayabahu, Dylan, Adeleke, Tinuade
Abstract
Reasoning models do not stop when they know the answer. On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model's own answer probability takes to settle, and how much of that excess is removable varies from problem to problem, so a global length penalty cannot take it out. We take it out by internalizing a causal interpretability finding into the weights. The mechanism is a halt vector: a difference-of-means direction at layer 18 of this model whose steering strength controls how long it thinks, while a replicated value axis does nothing. Installing that intervention in the weights is harder than it looks. Maximizing the scalar projection onto the direction corrupts the off-axis dimensions a frozen downstream reader depends on, and generation gets longer instead of shorter; what works is reconstructing the whole steered activation with those dimensions pinned to their natural values. Fit from 24 problems and no reinforcement learning, the halt removes about a quarter of the thinking at held accuracy across five unseen benchmarks, and the cut tracks each problem's own removable slack at 0.70. It also closes a non-termination pathology that grows with difficulty and that a decoding-time confidence hook makes worse. We do not claim to beat a well-tuned length penalty or decoding-time early exit on the raw trade-off; the contribution is how the halt is obtained.
Chinese Translation
推理模型在已经知道答案时并不会停止思考。在 DeepSeek-R1-Distill-Qwen-7B 上,思维链的长度约为模型自身答案概率收敛所需时间的两倍,而且其中可去除的多余部分因问题而异,因此全局长度惩罚无法将其消除。我们通过将因果可解释性发现内化到模型权重中来消除这部分冗余。其机制是一个停止向量:该模型第18层的一个均值差方向,其引导强度可以控制模型思考的时长,而复制得到的数值轴方向则不产生任何作用。将这一干预固化到权重中比看上去更困难:最大化沿该方向的标量投影会破坏冻结的下游读取器所依赖的离轴维度,反而使生成变得更长;有效的方法是在将这些维度固定为其自然取值的前提下重建整个被引导的激活。仅用24个问题进行拟合且不使用强化学习,该停止向量在五个未见基准上保持准确率的同时削减了约四分之一的思考量,且削减量与每个问题自身可去除的冗余量之间的相关系数达到0.70。它还消除了一个随问题难度增长、且解码时置信度钩子机制会加剧的非终止病态。我们并不声称在原始权衡上超越调优良好的长度惩罚或解码时提前退出方法;本文的贡献在于停止向量的获取方式。
cs.LG / 6 / 2608.28896

Conservative Hybrid Graph Networks for Process Systems with Learned Routing

具有可学习路由的面向过程系统的守恒混合图网络
Guida, Paolo
Abstract
Industrial process networks do not maintain a single effective topology while operating: streams are throttled or bypassed, and units move between idle, transition, and active regimes. Models of such systems are typically trained on measured state trajectories while the operating mechanisms that generated them remain latent, and an unconstrained graph network can fit such a trajectory without assigning stable physical meaning to the recovered routing. We address both problems with the Conservative Hybrid Graph Network (CHGN), which learns routing, regime assignment, and removal rates as data-driven surrogates and inserts them into a fixed transport equation, so that the mass balance holds by construction for any predicted routing. CHGN trained on networks of 10-20 nodes transfers zero-shot to unseen graphs of 25-40 nodes without retraining, reaching an RMSE of 2.1e-3 against 6e-2 to 9e-2 for GNN baselines under the same protocol, with a gate MAE of 7.9e-3 and regime accuracy of 94.3% (1.2e-2 and 96.4% respectively on the fixed training topology). On a fluid-mixing pilot plant, CHGN improves on a persistence baseline for held-out physical faults but does not predict manual interventions, for which the governing valve actions are unobserved. The model therefore transfers across process topologies without retraining and exposes the latent mechanisms governing plant behaviour to inspection.
Chinese Translation
工业过程网络在运行过程中并非维持单一的等效拓扑:流股被节流或旁路,装置在闲置、过渡和活跃工况之间切换。此类系统的模型通常基于测得的状态轨迹进行训练,而生成这些轨迹的运行机制仍然隐而不显;无约束的图网络可以拟合这样的轨迹,却无法为其恢复出的路由赋予稳定的物理意义。为此,我们提出了守恒混合图网络(Conservative Hybrid Graph Network, CHGN),它将路由、工况分配和移除速率作为数据驱动的代理进行学习,并将其嵌入固定的输运方程中,从而使得质量守恒对于任何预测出的路由在构造上即成立。在10–20节点网络上训练的CHGN可零样本迁移至25–40节点的未见图结构而无需再训练,在相同协议下RMSE达到2.1e-3,而GNN基线为6e-2至9e-2;阀门开度MAE为7.9e-3,工况识别准确率为94.3%(在固定训练拓扑上分别为1.2e-2和96.4%)。在一个流体混合中试装置上,CHGN对留出的物理故障优于持续性基线,但无法预测人工干预,因为其对应的阀门操作未被观测到。因此,该模型无需再训练即可在不同过程拓扑间迁移,并将支配装置行为的潜在机制呈现在可检验的形式之下。
cs.LG / 7 / 2608.28905

Off-Policy Evaluation for Semantic ID Recommenders: Does the Model's Own Code Hierarchy Help?

面向语义ID推荐系统的离线策略评估:模型自身的编码层级结构是否有帮助?
Betlei, Artem
Abstract
Generative recommenders increasingly emit semantic IDs (SIDs): each item is a short sequence of hierarchical discrete codes from a residual quantizer, decoded autoregressively. Before spending scarce A/B-test, a team may decide offline which decoder or reranking variants are worth testing - a job for off-policy evaluation (OPE). We ask a simple question: can the model's own SID tree serve as the action abstraction for that OPE? Our answer has three parts. (i) Under the near-argmax logging real recommenders use, per-item OPE is hopeless - as item-level effective sample size is usually small on production logs - but marginalizing items to code-prefix clusters restores estimable support and cuts error. (ii) This gain is thanks to coarsening, not to the hierarchy specifically; but the SID tree is what makes coarsening feasible in a generative system - each cluster's mass is exactly and cheaply returned by the decoder, whereas flat clustering requires enumerating item/leaf masses that a code-only decoder does not directly expose. (iii) Resolution depth is the operative knob - coarser under scarce support - and a conditional bias bound links the coarsening bias to the quantizer's worst-case reconstruction residual and the target-logging divergence.
Chinese Translation
生成式推荐系统越来越多地输出语义ID(Semantic ID, SID):每个物品被表示为来自残差量化器的一段简短的层级离散编码序列,并以自回归方式解码。在投入稀缺的A/B测试资源之前,团队可能需要离线判断哪些解码器或重排序变体值得测试——这正是离线策略评估(Off-Policy Evaluation, OPE)的任务。我们提出了一个简单的问题:模型自身的SID树能否作为该OPE的动作抽象?我们的回答包含三个部分。(i) 在真实推荐系统所使用的近似argmax日志记录策略下,逐物品的OPE是行不通的——因为在生产日志中,物品层面的有效样本量通常很小——但将物品归并到编码前缀簇上可以恢复可估计的支持集并降低误差。(ii) 这一增益归功于粗化本身,而非层级结构本身;但正是SID树使得粗化在生成式系统中变得可行——每个簇的质量可以由解码器精确且低成本地返回,而平坦聚类则需要枚举仅基于编码的解码器无法直接暴露的物品/叶子质量。(iii) 分辨率深度是关键的可调参数——在支持集稀缺时应采用更粗的粒度——并且一个条件偏差界将粗化偏差与量化器的最坏情况重构残差以及目标策略与日志策略之间的散度联系起来。
cs.LG / 8 / 2608.28910

Learning-Theoretic Foundation for General Coded Computing: The Straggler Setting

通用编码计算的学习理论基础:掉队者(Straggler)场景
Moradi, Parsa, Tahmasebi, Behrooz, Maddah-Ali, Mohammad Ali
Abstract
Coded computing has emerged as a powerful paradigm for mitigating the impact of straggling workers in distributed computing systems. However, existing coded-computing schemes are predominantly designed for the exact recovery of highly structured computations, such as polynomial evaluation and matrix multiplication, and typically rely on strict recovery thresholds. These assumptions significantly limit their applicability to modern machine-learning workloads, particularly deep neural networks (DNNs), whose computations generally lack rigid algebraic structure and, in many applications, require only accurate approximations rather than exact recovery. To address this gap, we revisit coded computing from a learning-theoretic perspective and introduce General Coded Computing (GCC). Rather than adopting existing algebraic tools, GCC formulates coded computing through a natural end-to-end mean-squared error loss that directly measures the discrepancy between the desired computations and their recovered estimates. By deriving suitable upper bounds and restricting the encoder and decoder to a reproducing kernel Hilbert space (RKHS) with mild smoothness constraints, we show that both the encoder and decoder admit specific representations as linear combinations of RKHS kernel functions. This representation allows the corresponding coefficients to be computed efficiently. Moreover, this framework enables us to establish theoretical performance guarantees for GCC under two complementary straggler regimes. In the worst-case setting with $N$ worker nodes, and at most $S$ stragglers, we show that the end-to-end loss decays at least at rate $O(S^3N^{-3})$ for standard configurations. We then study a probabilistic setting in which each worker independently straggles with probability $p$. We prove that the expected loss can still converge at rate $O(\log_{1/p}^3(N)N^{-3})$.
Chinese Translation
编码计算已成为缓解分布式计算系统中掉队工作者(straggling workers)影响的有力范式。然而,现有的编码计算方案主要针对高度结构化的计算(如多项式求值和矩阵乘法)的精确恢复而设计,并且通常依赖于严格的恢复阈值。这些假设极大地限制了其对现代机器学习工作负载的适用性,尤其是深度神经网络(DNN),其计算通常缺乏严格的代数结构,且在许多应用中只需精确的近似而非精确恢复。为弥补这一空白,我们从学习理论的视角重新审视编码计算,并提出了通用编码计算(General Coded Computing, GCC)。GCC 没有采用现有的代数工具,而是通过一个自然的端到端均方误差损失来构建编码计算,该损失直接度量目标计算与其恢复估计之间的差异。通过推导合适的上界,并将编码器和解码器限制在具有温和光滑性约束的再生核希尔伯特空间(RKHS)中,我们证明编码器和解码器均可表示为 RKHS 核函数的特定线性组合。这一表示使得相应的系数可以被高效计算。此外,该框架使我们能够在两种互补的掉队者情形下为 GCC 建立理论性能保证。在最坏情形设置下,设工作者节点数为 $N$,掉队者至多为 $S$ 个,我们证明对于标准配置,端到端损失至少以 $O(S^3N^{-3})$ 的速率衰减。随后,我们研究了每个工作者以概率 $p$ 独立掉队的概率性设置,并证明期望损失仍能以 $O(\log_{1/p}^3(N)N^{-3})$ 的速率收敛。
cs.LG / 9 / 2608.28911

SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

SemKV:面向长上下文LLM推理的、由质量悬崖引导的语义混合精度KV缓存量化
Lee, Daeha, Kim, Do-Hyung, Kim, Jae-Hong
Abstract
The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length. We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguishable from FP16 KV down to 2.322 code bits/value and collapses at 2.0 bits - a quality cliff in (2.0, 2.322] that reappears in generation-time quantization and multi-turn dialogue and transfers to Mistral-7B. The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interchangeable, so the benefit of mixing is grid interpolation, reaching average precisions uniform quantization cannot realize. SemKV preserves every token, ranks tokens by a model-internal score, and assigns two adjacent above-cliff precisions, achieving a measured 6.0x storage reduction with no statistically detectable quality difference from full KV (n=900, three seeds), and outperforming FP16 token pruning granted a 1.5x larger memory budget. Replacing the affine base with a distortion-optimized quantizer (TurboQuant-MSE) lowers the cliff in every protocol tested, raising the no-detectable-loss operating point to 7.9x. The recipe: measure the cliff for the target deployment setting, then interpolate above it.
Chinese Translation
键值(KV)缓存是长上下文大语言模型(LLM)推理的主要内存瓶颈,其大小随上下文长度线性增长。我们证明,在非整数比特网格上的均匀KV量化并不会平滑地退化:在预先设定的多种子统计协议下,使用仿射量化器的Llama-3.1-8B-Instruct在低至每值2.322码比特时仍与FP16 KV在统计上不可区分,而在2.0比特时崩溃——即存在一个位于(2.0, 2.322]区间的质量悬崖,该悬崖在生成时量化和多轮对话中再次出现,并可迁移至Mistral-7B。该悬崖重新定义了重要性感知的混合精度:在悬崖之上,八种模型内部重要性指标在统计上可互换,因此混合的收益在于网格插值,能够达到均匀量化无法实现的平均精度。SemKV保留每个token,依据模型内部得分对token排序,并分配两个相邻的悬崖之上精度,实现了实测6.0倍的存储压缩,且与完整KV相比无统计学上可检测的质量差异(n=900,三个种子),并在获得1.5倍更大内存预算的情况下优于FP16 token剪枝。用失真优化的量化器(TurboQuant-MSE)替换仿射基础量化器,可在所有测试协议下降低悬崖,将无可检测损失的运行点提升至7.9倍。该方法的流程是:针对目标部署场景测量质量悬崖,然后在悬崖之上进行插值。
cs.LG / 10 / 2608.28922

RankShift: In-Database Detection and Explanation of Categorical Shifts

RankShift:数据库内的类别偏移检测与解释
Ahmed, Omair Shafi
Abstract
A login service can receive its usual number of failed sign-ins while one source grows from 2% to 30% of them. The same pattern appears in system logs when a rare event template becomes common while the message rate stays stable. These events change which categories are active without changing how many events occur. RankShift detects such changes inside the analytical database that stores the data. It compares each window's category shares with a benign reference using a Pearson score whose terms identify the categories responsible for the change. The same query returns the score, calibrated alert, and largest increasing contributions. We evaluate RankShift on HDFS, BGL, and Thunderbird. It matches the count-vector autoencoder within 0.001 AUROC on HDFS (0.999 versus 1.000) and leads on Thunderbird (0.983 versus 0.949). In a controlled fixed-volume experiment, RankShift detects rare-category shifts that are invisible to event-count monitoring, reaching 0.787 AUROC compared with 0.771 for the autoencoder. Across all three corpora, observed false-alarm rates track the requested operating levels. RankShift requires no model training or inference service, and the autoencoders deployed state is 137x larger.
Chinese Translation
一个登录服务可能接收到数量与平时相同的失败登录,而其中某个来源的占比却从2%增长到30%。当稀有事件模板变得常见而消息速率保持稳定时,系统日志中也会出现同样的模式。这类事件改变了哪些类别处于活跃状态,却不改变事件发生的数量。RankShift能够在存储数据的分析数据库内部检测此类变化。它使用Pearson分数将每个时间窗口的类别占比与良性参考进行比较,该分数的各项能够识别出导致变化的具体类别。同一个查询即可返回分数、经校准的告警以及最大的增长贡献项。我们在HDFS、BGL和Thunderbird数据集上评估了RankShift。在HDFS上,其结果与计数向量自编码器(count-vector autoencoder)的AUROC差距在0.001以内(0.999对1.000),并在Thunderbird上领先(0.983对0.949)。在一项控制固定数据量的实验中,RankShift能够检测到事件计数监控无法察觉的稀有类别偏移,达到0.787的AUROC,而自编码器仅为0.771。在全部三个语料库上,观测到的误报率均与设定的运行水平相符。RankShift无需模型训练或推理服务,且部署自编码器所需的状态空间比它大137倍。
cs.LG / 11 / 2608.28934

Revisiting the Provable-Auditable Privacy Gap of DP-SGD

重新审视DP-SGD的可证明隐私与可审计隐私之间的差距
Modi, Saloni, Balaji, Srivi, Zhu, Yusong, Kamath, Gautam, Tian, Kevin
Abstract
Differential privacy (DP) has traditionally been used to provide theoretical upper bounds on an algorithm's stability to changing its training data. In modern private machine learning applications, achieving strong tradeoffs between utility and theoretical privacy is challenging, and thus one may optimistically hope that existing theoretical privacy analyses are loose. Recent work on privacy auditing has adopted a dual viewpoint, instead lower bounding the true privacy of an algorithm by constructing empirical distinguishing events. The auditing literature has thus far yielded a pessimistic outlook on the looseness of theoretical privacy bounds for DP-SGD, the de facto private training method in modern ML, as nearly-matching empirical lower bounds have been achieved under various threat models [NHSBTJCT23, AC24, CBP25]. In this work, we propose the empirical privacy lower bound of an algorithm as a concrete metric to optimize for, complementary to the theoretical upper bound. We give a lightweight defense framework that generically augments optimization methods in the ML pipeline to have significantly-improved empirical privacy on standard benchmarks. Moreover, we show that our framework comes at no theoretical privacy cost when augmenting DP-SGD, unlike previously-proposed defenses against membership inference attacks. We evaluate our defense against a broad range of audit constructions, models, and datasets to demonstrate its flexibility.
Chinese Translation
差分隐私(DP)传统上被用于提供算法对训练数据变化的稳定性的理论上界。在现代隐私保护机器学习应用中,实现效用与理论隐私之间的强权衡十分困难,因此人们可能乐观地希望现有的理论隐私分析是宽松的。最近的隐私审计工作采用了相反的视角,通过构造经验性的区分事件来对算法的真实隐私进行下界估计。迄今为止,审计文献对DP-SGD(现代机器学习中事实上的隐私训练方法)理论隐私界的宽松程度给出了悲观的结论,因为在多种威胁模型下已经得到了与之几乎匹配的经验下界 [NHSBTJCT23, AC24, CBP25]。在这项工作中,我们提出将算法的经验隐私下界作为一个具体可优化的指标,与理论上界相辅相成。我们提出了一个轻量级防御框架,可以通用地增强机器学习流程中的优化方法,使其在标准基准上获得显著改善的经验隐私。此外,我们证明了在增强DP-SGD时,我们的框架不会带来任何理论隐私代价,这与先前提出的针对成员推理攻击的防御方法不同。我们在广泛的审计构造、模型和数据集上对该防御方法进行了评估,以证明其灵活性。
cs.LG / 12 / 2608.28948

From the Loss Landscape to Diverse Feature Learning in Neural Networks

从损失景观到神经网络中的多样性特征学习
Yunis, David Aram
Abstract
Over the course of the last decade, neural networks have grown from an academic curiosity to moving the markets of nations. Despite this explosion in both research and deployment, relatively little is understood about how they achieve the solutions they do. This is both scientifically relevant, and pressing for society. When neural networks make decisions across self-driving, construction, law, hiring and health, there have been and will continue to be unintended consequences. However, attempting to generalize the failures of the largest and most important production systems makes for a very difficult task. Yet signs of these failures exist at all scales of neural networks, so we should be able to study a much more tractable setting. All neural networks must undergo an optimization process, called training, to be useful. To a great degree, understanding neural networks is understanding their optimization: through what process and exposure to which data did they arrive at their results. Yet our knowledge on this topic as a field is quite imprecise. In particular, a curious phenomenon called mode connectivity, the ability to connect neural networks in the loss surface, defies explanation entirely. This dissertation elucidates, explains and exploits this special structure in the loss landscape...
Chinese Translation
在过去十年中,神经网络已从一种学术上的好奇事物发展成为能够影响国家市场的技术。尽管在研究和部署方面都出现了爆炸式增长,但人们对神经网络如何实现其所给出的解仍然知之甚少。这不仅具有重要的科学意义,也对社会具有紧迫性。当神经网络在自动驾驶、建筑、法律、招聘和医疗等领域做出决策时,已经产生并将会继续产生意想不到的后果。然而,试图归纳总结那些最大、最重要的生产系统的失效原因,是一项极为困难的任务。但这些失效的迹象在各个规模的神经网络中都存在,因此我们应该能够在一个更容易处理的研究环境中对其进行研究。所有神经网络都必须经历一个称为训练的优化过程才能发挥作用。在很大程度上,理解神经网络就是理解其优化过程:它们经历了怎样的过程、接触了哪些数据,才得出其结果。然而,作为一个研究领域,我们在这方面的知识还相当不精确。特别是,一种被称为模式连通性的奇特现象——即在损失表面上连接神经网络的能力——完全无法得到解释。本论文对损失景观中的这一特殊结构进行了阐明、解释和利用……
cs.LG / 13 / 2608.28960

Continuity-Free Near-Minimax Leading-Order Regret for CVaR-UCBVI

CVaR-UCBVI的无连续性近极小极大首阶遗憾界
Chen, Yuanlong
Abstract
For finite-horizon tabular CVaR reinforcement learning, prior work proves a $\widetilde{O}(\tau^{-1}\sqrt{SAK})$ leading regret bound for arbitrary normalized return laws and the sharper $\widetilde{O}(\sqrt{SAK/\tau})$ rate under a density lower bound. We show that the same Bernstein CVaR-UCBVI algorithm attains the sharper rate without continuity assumptions. The key is a selected-budget self-bound: the conditional variance of the episode shortfall is at most $\tau$ plus the value-estimation width. Substitution into the original Bernstein decomposition yields, with high probability, $\widetilde{O}(\sqrt{SAK/\tau}+(SAHK^{1/4}+S^2AH)/\tau)$ regret for arbitrary normalized return laws, including atomic, mixed, and continuous laws. The $\tau^{-1/2}$ leading term matches the expected-regret minimax lower bound up to logarithmic factors. Thus Bernstein CVaR-UCBVI is minimax-optimal over the full return-law class in the leading-order regime; the lower-order terms retain their $\tau^{-1}$ dependence.
Chinese Translation
对于有限时域表格型CVaR强化学习,已有工作针对任意归一化回报分布证明了$\widetilde{O}(\tau^{-1}\sqrt{SAK})$的首阶遗憾界,并在密度函数存在下界的条件下得到了更优的$\widetilde{O}(\sqrt{SAK/\tau})$速率。我们证明,同一Bernstein CVaR-UCBVI算法无需任何连续性假设即可达到该更优速率。关键在于一个选择性预算自约束(selected-budget self-bound):回合短缺量(episode shortfall)的条件方差至多为$\tau$加上价值估计宽度。将其代入原始的Bernstein分解后,我们以高概率得到,对于任意归一化回报分布(包括原子分布、混合分布和连续分布),遗憾界为$\widetilde{O}(\sqrt{SAK/\tau}+(SAHK^{1/4}+S^2AH)/\tau)$。其$\tau^{-1/2}$首阶项在对数因子范围内匹配了期望遗憾的极小极大下界。因此,Bernstein CVaR-UCBVI在整个回报分布类别上于首阶意义下是极小最优的;而低阶项仍保持$\tau^{-1}$的依赖关系。
cs.LG / 14 / 2608.28981

V2TATC: A Joint Voice-Trajectory Embedding Framework and Dataset for Air Traffic Controller Situational Awareness

V2TATC:面向空中交通管制员态势感知的语音-航迹联合嵌入框架与数据集
Brusset, Louis, Petit, Mathurin, Kam, Jordan, Bayen, Alexandre
Abstract
As air traffic volumes in the National Airspace System continue to expand, in particular in the low altitude airspaces, the need for scalable decision support tools used by air traffic controllers will also require more development. This article introduces Voice-to-Trajectory for Air Traffic Control, a joint voice communication-flight trajectory data embedding framework, that can be a component of situational awareness in congested airspaces, and assist the development of tools for ATC as they reason in real-time over Automatic Dependent Surveillance-Broadcast trajectories, or the intent expressed by pilots in natural language. We show that these data modalities are not independent and represent a common physical referent: an aircraft flying through the airspace. V2TATC maps a voice instruction and the trajectory of the addressed aircraft to nearby points in a single latent space that can be queried in both directions. It combines a self-supervised trajectory encoder, a frozen large-scale speech encoder, a contrastive joint embedding, and a bijective lifting via normalizing flows. We demonstrate V2TATC's effectiveness on the San Francisco Bay Area, for its concentration of major airports, and its mix of commercial and general aviation low altitude traffic. Lastly, we release a novel paired voice-trajectory dataset, and report experiments on cross-modal retrieval, ablations, and latent-space analysis.
Chinese Translation
随着美国国家空域系统(National Airspace System)空中交通量的持续增长,尤其是低空空域的交通量增长,空中交通管制员对可扩展决策支持工具的需求也将日益增加。本文提出了“语音转航迹空中交通管制”(Voice-to-Trajectory for Air Traffic Control, V2TATC),这是一个语音通信与飞行航迹数据的联合嵌入框架,可作为拥挤空域中态势感知的组成部分,并助力空中交通管制(ATC)工具的开发,使其能够实时地对广播式自动相关监视(ADS-B)航迹或飞行员以自然语言表达的意图进行推理。我们证明这些数据模态并非独立,而是代表了同一物理对象:即在该空域中飞行的航空器。V2TATC 将语音指令与目标航空器的航迹映射到同一潜在空间中彼此邻近的点,该空间支持双向查询。该框架结合了自监督航迹编码器、冻结的大规模语音编码器、对比联合嵌入,以及通过归一化流实现的双射提升。我们在旧金山湾区演示了 V2TATC 的有效性,该地区因拥有密集的大型机场以及商业航空与通用航空低空交通混合运行的特点而具有代表性。最后,我们发布了一个新颖的语音-航迹配对数据集,并报告了跨模态检索、消融实验和潜在空间分析方面的实验结果。
cs.LG / 15 / 2608.29001

Effective Graph and Rank-based Contextual Embeddings for Textual and Multimedia Data

面向文本与多媒体数据的有效图与基于排名的上下文嵌入
Almeida, Thiago César Castilho, Letício, Gustavo Rosseto, Valem, Lucas Pascotti, Freitas, André, Pedronette, Daniel Carlos Guimarães
Abstract
In a data-driven world, efficiently organizing and mapping relationships between objects is crucial. Graphs are powerful tools for modeling these connections, being widely used in social networks, telecommunications, and biology. However, graph-based methods often face high computational costs, particularly in memory and space usage. To address this, graph embedding techniques, also referred to as Network Representation Learning, encode graph information into lower-dimensional representations while preserving structural aspects. Traditional methods, however, lack interpretable dimensions. RaDE (Rank Diffusion Embedding) introduces a new approach using rank-based information, with a key step being the selection of a representative subset of nodes to provide interpretability for its dimensions and improve retrieval tasks. Despite its potential, RaDE's original proposal did not fully explore the effectiveness of representative subset selection across different classes or evaluate embeddings in tasks like classification and clustering. Inspired by RaDE, this work introduces GRaCE (Graph and Rank-based Contextual Embeddings), a fully unsupervised framework that generates interpretable embeddings by leveraging robust rank-based measures for representative subset selection and node embedding. GRaCE surpasses RaDE and Original Features across diverse datasets, including textual and image collections, excelling in retrieval, classification, and clustering tasks, considering state-of-the-art Transformer models as feature descriptors and Graph Convolutional Networks models in classification tasks.
Chinese Translation
在数据驱动的世界中,高效地组织和映射对象之间的关系至关重要。图是建模这些连接的强大工具,被广泛应用于社交网络、电信和生物学领域。然而,基于图的方法通常面临较高的计算成本,尤其是在内存和空间使用方面。为解决这一问题,图嵌入技术(也称为网络表示学习,Network Representation Learning)将图信息编码为低维表示,同时保留结构特性。然而,传统方法缺乏可解释的维度。RaDE(Rank Diffusion Embedding,排名扩散嵌入)提出了一种利用基于排名信息的新方法,其关键步骤是选择一个具有代表性的节点子集,从而为其维度提供可解释性并改进检索任务。尽管潜力巨大,RaDE 的原始方案并未充分探索代表性节点子集选择在不同类别上的有效性,也未在分类和聚类等任务中评估其嵌入效果。受 RaDE 启发,本工作提出了 GRaCE(Graph and Rank-based Contextual Embeddings,基于图与排名的上下文嵌入),这是一个完全无监督的框架,通过利用稳健的基于排名的度量来进行代表性子集选择和节点嵌入,从而生成可解释的嵌入。在包括文本和图像集合在内的多种数据集上,GRaCE 的表现超越了 RaDE 和原始特征,在检索、分类和聚类任务中表现优异,其中分类任务使用了最先进的 Transformer 模型作为特征描述符以及图卷积网络(Graph Convolutional Networks)模型。
cs.LG / 16 / 2608.29004

Context-Aware Interpretable Representations for Retrieval and Graph Convolutional Network Classification

面向检索与图卷积网络分类的上下文感知可解释表示
Almeida, Thiago César Castilho, Letício, Gustavo Rosseto, Kawai, Vinicius Atsushi Sato, Pedronette, Daniel Carlos Guimarães
Abstract
The advances in visual information modeling and representation during the last decades are remarkable, mainly supported by Convolutional Neural Networks, Transformer-based, and Foundation Models. Despite this progress, critical challenges regarding the nature of similarity assessment and model transparency have been neglected. A primary concern is the Geometric Gap, where traditional pairwise measures fail to capture the intrinsic geometry of the dataset manifold. Furthermore, the Interpretability Gap persists, as representations often lack alignment with human cognition. Therefore, how to provide interpretability to representations while maintaining low dimensionality and high effectiveness in downstream tasks remains an open challenge. In this paper, we propose a novel unsupervised framework that integrates Manifold Learning strategies with Rank-based Interpretable Graph Embeddings. Our approach effectively bridges these gaps by first characterizing the contextual information of the dataset through manifold analysis and subsequently generating sparse, self-explainable embeddings. The proposed approach employs a flexible formulation, allowing different Manifold Learning and Representation Learning strategies. Extensive experimental evaluation across diverse datasets and features demonstrates that our Context-Aware representations not only provide intrinsic interpretability and dimensionality reduction but also maintain or enhance effectiveness in downstream tasks, specifically in image retrieval and semi-supervised classification using Graph Convolutional Networks (GCNs).
Chinese Translation
在过去几十年中,视觉信息建模与表示取得了显著进展,这主要得益于卷积神经网络、基于Transformer的模型以及基础模型(Foundation Models)。尽管取得了这些进展,但在相似性度量的本质和模型透明性方面仍然存在被忽视的关键挑战。首要问题是“几何鸿沟”(Geometric Gap),即传统的成对度量方法无法捕捉数据集流形的内在几何结构。此外,“可解释性鸿沟”(Interpretability Gap)依然存在,因为表示往往缺乏与人类认知的一致性。因此,如何在保持低维性和下游任务高效性的同时为表示提供可解释性,仍然是一个悬而未决的挑战。在本文中,我们提出了一种新颖的无监督框架,该框架将流形学习(Manifold Learning)策略与基于排序的可解释图嵌入(Rank-based Interpretable Graph Embeddings)相结合。我们的方法首先通过流形分析刻画数据集的上下文信息,随后生成稀疏且可自解释的嵌入,从而有效弥合了上述鸿沟。所提出的方法采用灵活的公式化设计,允许使用不同的流形学习和表示学习策略。在多个数据集和特征上的广泛实验评估表明,我们的上下文感知表示不仅提供了内在可解释性和降维能力,而且在下游任务(特别是图像检索和使用图卷积网络(GCN)的半监督分类)中保持或提升了有效性。
cs.LG / 17 / 2608.29024

Hybrid Semantic Context-Enhanced Ensemble Learning for Wind Power Ramp-Event Forecasting and Uncertainty-Aware Evaluation

面向风电功率爬坡事件预测与不确定性感知评估的混合语义上下文增强集成学习方法
Ali, Momina Liaqat, Abid, Muhammad, Abdullah, Muhammad, Zameer, Aneela
Abstract
Wind power ramp events which are sudden, large swings in turbine output over short windows are difficult to estimate, and standard models often miss them. Hybrid forecasting approach is built which augments semantic context to ramp-event forecast. Rather than applying an extensive language model directly to predict turbine operating data, we have implemented a pipeline where turbine operating data is converted to simplified text, which is then converted to dense embeddings to be used as inputs for ensemble models incorporated with other features. Testing runs are performed at multiple intervals within the SDWPF dataset, including 10-minute, 30-minute, and 60- minute horizons, with ramp events constituting the highest change in future power output. We check robustness against autoregressive, LSTM, and GRU baselines plus several ensemble configurations, using Diebold-Mariano tests and bootstrap confidence intervals, and we vary the ramp threshold, compress the embeddings with PCA, and validate externally on Kaggle SCADA and NREL data with uncertainty-aware scoring. The semantic-context features produce negligible yet statistically significant gains over the baselines in multiple paired ensemble runs, most clearly at the 30- and 60-minute horizons where these gains hold across different ramp-threshold definitions, and PCA compression helps in some longer-horizon cases. The best context- augmented ensembles rank near the top overall, though the GRU model still posts the lowest ramp-event RMSE at 30 and 60 minutes. External tests confirm the error reduction generalizes across datasets, but the size of the gain depends on both model and dataset. Prediction intervals cover most test cases well but weaken during ramp events, pointing to a localized shift in the data distribution.
Chinese Translation
风电功率爬坡事件是指风机输出在短时间内发生的大幅突变,此类事件难以准确估计,标准模型往往难以捕捉。本文构建了一种混合预测方法,将语义上下文信息融入爬坡事件预测。该方法并非直接应用大规模语言模型预测风机运行数据,而是实现了一条流水线:先将风机运行数据转换为简化文本,再将其转换为稠密嵌入向量,与其他特征一同作为集成模型的输入。我们在SDWPF数据集的多个时间尺度上进行了测试,包括10分钟、30分钟和60分钟预测时域,其中爬坡事件定义为未来功率输出的最大变化。我们使用Diebold-Mariano检验和自助法置信区间,检验了该方法相对于自回归模型、LSTM和GRU基线模型以及多种集成配置的鲁棒性,并改变了爬坡阈值、使用PCA压缩嵌入向量,同时结合不确定性感知评分在Kaggle SCADA和NREL数据上进行了外部验证。在多组配对集成实验中,语义上下文特征相比基线模型带来了虽微小但具有统计显著性的提升,这种提升在30分钟和60分钟时域最为明显,且在不同的爬坡阈值定义下均能保持,PCA压缩在某些较长时域的情况下有所助益。表现最佳的上下文增强集成模型总体排名接近榜首,但GRU模型在30分钟和60分钟时域的爬坡事件RMSE仍然最低。外部测试证实了误差降低能够泛化到不同数据集,但提升幅度同时取决于模型和数据集。预测区间对大多数测试样本覆盖良好,但在爬坡事件期间表现变弱,这表明数据分布存在局部偏移。
cs.LG / 18 / 2608.29029

Flow-JEPA: Flow Matching for Robust Latent Dynamics in JEPA World Models

Flow-JEPA:基于流匹配的JEPA世界模型鲁棒潜在动力学
Huo, Yanchen, Song, Ziying, Luo, Yadan
Abstract
Joint-Embedding Predictive Architectures (JEPAs) have shown strong potential for learning compact predictive representations, and LeWorldModel (LeWM) extends this paradigm to reconstruction-free latent world modeling from pixels. However, its deterministic autoregressive predictor generates future states through repeated one-step transitions, which can accumulate errors and remain sensitive to task-irrelevant visual perturbations. In this work, we propose Flow-JEPA (F-JEPA), a conditional flow matching dynamics model that jointly generates a sequence of future latent states conditioned on the current observation and actions. A Gaussian distribution serves as the flow source, exposing the vector field to perturbed latent trajectories as it learns to transport them toward clean future representations. This formulation retains the reconstruction-free JEPA framework while replacing point-wise transition regression with stochastic trajectory-level prediction. F-JEPA raises mean success from $86\%$ to $92\%$ under clean observations and from $67\%$ to $86\%$ under noisy conditions, suggesting that conditional flow matching provides a promising alternative to deterministic autoregressive dynamics in JEPA world models.
Chinese Translation
联合嵌入预测架构(Joint-Embedding Predictive Architectures, JEPA)在学习紧凑的预测表征方面展现出强大潜力,而LeWorldModel(LeWM)将该范式扩展到基于像素的无重构潜在世界建模。然而,其确定性的自回归预测器通过反复的单步转移生成未来状态,这会累积误差,并且对与任务无关的视觉扰动仍然敏感。在本工作中,我们提出Flow-JEPA(F-JEPA),这是一种条件流匹配动力学模型,它以当前观测和动作为条件,联合生成一系列未来潜在状态。高斯分布作为流的源分布,在向量场学习将受扰动的潜在轨迹传输至干净的未来表征的过程中,使向量场暴露于这些扰动轨迹。这一建模方式在保留无重构JEPA框架的同时,用随机轨迹级预测取代了逐点转移回归。F-JEPA将干净观测下的平均成功率从86%提升至92%,在噪声条件下的平均成功率从67%提升至86%,这表明条件流匹配为JEPA世界模型中的确定性自回归动力学提供了一种有前景的替代方案。
cs.LG / 19 / 2608.29045

NVE: A Separability and Coverage-Aware Internal Validation Metric for Biclustering

NVE:一种面向双聚类的可分性与覆盖度感知的内部验证指标
Tiwari, Paritosh, Kumar, I Navin, Bezdek, James C., Rathore, Punit
Abstract
Biclustering, or co-clustering, aims to discover coherent submatrices by grouping rows and columns of a data matrix simultaneously. This local two-dimensional structure makes validation more difficult than in ordinary clustering, where internal indices usually rely on compactness and separation in a single shared feature space. Existing popular internal biclustering measures such as Mean Squared Residue (MSR), and Virtual Error (VE) mainly evaluate within-bicluster coherence. Although useful, these measures do not directly assess whether the extracted biclusters are mutually distinct or whether they explain a meaningful portion of the data matrix. This paper investigates Normalised Virtual Error (NVE), an internal validation metric that extends VE using a super-bicluster normalization strategy. By comparing the VE of each bicluster with the VE obtained after merging it with other biclusters, NVE introduces a relative notion of separability and redundancy. We also study a coverage-adjusted variant, NVE\textsubscript{cov}, which penalizes solutions that obtain low error by selecting only very small submatrices. Through controlled synthetic benchmarks and yeast gene-expression datasets, we examine whether NVE and NVE\textsubscript{cov} provide information beyond standard coherence-based metrics. The results show that NVE is sensitive to redundant and poorly separated biclusters, while NVE\textsubscript{cov} changes solution rankings when low-error biclusters cover only a negligible part of the matrix. These findings suggest that NVE-based measures are useful complementary criteria for internal co-clustering validation, especially when coherence, separability, and coverage must be considered jointly.
Chinese Translation
双聚类(Biclustering),也称协同聚类(co-clustering),旨在通过对数据矩阵的行和列同时进行分组来发现相干的子矩阵。这种局部二维结构使得验证比普通聚类更加困难,因为普通聚类的内部指标通常依赖于单一共享特征空间中的紧凑性和分离度。现有的流行内部双聚类度量指标,如均方残差(Mean Squared Residue, MSR)和虚拟误差(Virtual Error, VE),主要评估双聚类内部的相干性。尽管这些指标很有用,但它们并不能直接评估所提取的双聚类之间是否相互区分,或是否能解释数据矩阵中有意义的部分。本文研究了归一化虚拟误差(Normalised Virtual Error, NVE),这是一种通过超双聚类归一化策略对VE进行扩展的内部验证指标。通过将每个双聚类的VE与将其与其他双聚类合并后得到的VE进行比较,NVE引入了相对的可分性和冗余性概念。我们还研究了一种覆盖度调整的变体NVE\textsubscript{cov},它通过惩罚仅选择极小子矩阵而获得低误差的解来调整评估。通过受控的合成基准测试和酵母基因表达数据集,我们检验了NVE和NVE\textsubscript{cov}是否能够提供超越标准相干性指标的信息。结果表明,NVE对冗余且分离度差的双聚类较为敏感,而NVE\textsubscript{cov}在低误差双聚类仅覆盖矩阵极小部分时会改变解的排序。这些发现表明,基于NVE的度量是有用的内部协同聚类验证补充标准,尤其是在需要同时考虑相干性、可分性和覆盖度的情况下。
cs.LG / 20 / 2608.29057

Sparse Koopman Autoencoders Identify Local Dynamical Regimes in Multibasin Systems

稀疏Koopman自编码器识别多盆系统中的局部动力学机制
Li, Aidan, Tadipatri, Uday Kiran Reddy, Fathi, Mahan, Chandar, Sarath, Goroshin, Ross
Abstract
Koopman autoencoders (KAEs) seek a higher-dimensional latent representation in which nonlinear dynamics evolve linearly. However, many interesting systems have multiple basins of attraction, and both theoretical and empirical work has shown these multibasin systems cannot generally admit a single finite-dimensional global Koopman embedding under standard assumptions. We posit that encoders with a sparsity-inducing objective encouraging few active latent coefficients will provide latent supports as an inspectable basin-modeling principle for Koopman autoencoders. We use these encoders producing sparse latents in training Sparse Koopman Autoencoders (SKAEs) without basin labels or other regime annotations, and treat the learned latent supports as model-produced regime variables after training. Across a range of procedurally generated multibasin systems and chaotic flows, we show that SKAEs have superior forecasting performance compared to dense-latent KAEs. We also perform a mechanistic study that shows latent supports produced by SKAEs are both essential for the quality of the representation and useful for identifying basins on held-out basin interior states, whereas dense-latent KAEs collapse to an uninformative single family. These results identify sparse latents and their corresponding supports as label-free, interpretable regime variables for Koopman learning in nonlinear systems with multiple local dynamical laws.
Chinese Translation
Koopman自编码器(KAEs)旨在寻找一个高维潜在表示,使非线性动力学在其中以线性方式演化。然而,许多有趣的系统具有多个吸引盆,理论和实证研究均已表明,在标准假设下,这类多盆系统通常无法容许单一的全局有限维Koopman嵌入。我们提出,采用鼓励少量活跃潜在系数的稀疏性诱导目标的编码器,能够提供潜在支撑集(latent supports),作为Koopman自编码器中一种可检验的盆建模原则。我们在训练稀疏Koopman自编码器(SKAEs)时使用这类编码器生成稀疏潜在表示,且无需盆标签或其他机制标注,并在训练后将学到的潜在支撑集视为由模型产生的机制变量。在一系列程序化生成的多盆系统和混沌流上,我们证明SKAEs相比稠密潜在表示的KAEs具有更优的预测性能。我们还进行了机理研究,结果表明SKAEs产生的潜在支撑集不仅对表示质量至关重要,还有助于在留出的盆内部状态上识别吸引盆,而稠密潜在表示的KAEs则坍缩为一个无信息量的单一族。这些结果将稀疏潜在表示及其对应的支撑集确立为无标签、可解释的机制变量,适用于具有多个局部动力学规律的非线性系统中的Koopman学习。
cs.LG / 21 / 2608.29061

PathBridger: Subgoal Bridges for Offline Goal-Conditioned Reinforcement Learning

PathBridger:用于离线目标条件强化学习的子目标桥接方法
Choi, Soohyun, Cho, Seonvin, Hong, Songnam
Abstract
Offline goal-conditioned reinforcement learning (GCRL) aims to learn policies for reaching diverse goals entirely from fixed trajectory data. Long-horizon offline GCRL remains challenging because sparse goal-reaching signals must be propagated over many steps, while execution errors cannot be corrected through additional environment interaction. Existing methods address these challenges by improving long-range value estimation or reducing the effective decision horizon through subgoals, options, and action chunks. In several hierarchical methods, however, a selected subgoal specifies where to go, while the intervening state-space path remains implicit in an endpoint-conditioned low-level policy. To address this interface, we propose PathBridger, a hierarchical offline GCRL method that explicitly connects subgoal selection to short-horizon execution. PathBridger constructs a state-space bridge toward the selected intermediate endpoint and decodes it into a short executable action chunk using an inverse dynamics model. Experiments across the evaluated OGBench tasks demonstrate strong aggregate performance, with particularly large gains on the multi-object Cube manipulation tasks. Code: https://github.com/SChoish/PathBridger
Chinese Translation
离线目标条件强化学习(GCRL)旨在完全从固定的轨迹数据中学习能够到达多样化目标的策略。长时程离线GCRL仍然具有挑战性,因为稀疏的目标达成信号需要在许多步骤间进行传播,而执行误差无法通过额外的环境交互加以纠正。现有方法通过改进长距离价值估计,或借助子目标(subgoal)、选项(option)和动作块(action chunk)缩短有效决策时域来应对这些挑战。然而,在一些分层方法中,所选的子目标仅指明要去往何处,而中间的状态空间路径仍然隐含在以端点为条件的高层策略所驱动的高层策略与以端点为条件的低层策略之中。针对这一接口问题,我们提出了PathBridger,一种显式地将子目标选择与短时程执行相连接的分层离线GCRL方法。PathBridger朝向所选的中间端点构建一条状态空间桥接路径,并利用逆动力学模型将其解码为一段较短的可执行动作块。在所评估的OGBench任务上的实验表明,该方法具有出色的整体性能,在多物体Cube操作任务上尤其取得了显著的性能提升。代码:https://github.com/SChoish/PathBridger
cs.LG / 22 / 2608.29070

Selective Disclosure of Hidden Directives in Reasoning Models: Behavioral Asymmetry and Steering

推理模型中对隐藏指令的选择性披露:行为不对称性与导向
Shi, Zimo, Tifft, Xander, Xing, Wen
Abstract
Chain-of-thought (CoT) reasoning traces are increasingly proposed as a mechanism for AI oversight: a monitor inspecting a model's reasoning can, in principle, detect misbehavior invisible from outputs alone. This assumes CoT surfaces what a model is instructed to do regardless of the instructions given. We test this assumption along two axes. First, we introduce the Instruction-Compliance Gap (ICG): the difference in probability that a model's CoT explicitly references a hidden system prompt directive when that directive is malign versus benign. Across 100 task pairs and 8 frontier reasoning models from 5 families, we find consistent asymmetric disclosure, a higher probability of leaking malign hidden instructions than benign ones, in Qwen3-14B (Wilcoxon $p=0.0001$, $+13.9$pp), Qwen3-32B ($p=0.0011$, $+13.0$pp), Qwen3-235B ($p=0.035$, $+5.8$pp), and similar results with MiniMax-M2.5 and DeepSeek-R1. The detector has 100% precision against two independent blinded labelling passes, and an LLM monitor reading only the reasoning trace reproduces the asymmetry in all 8 models against directive-free controls, identifying the specific directive in 82% of malign traces which the detector classifies as clean. Second, steering vectors extracted in MiniMax-M2.5 via Contrastive Activation Addition causally induce hiding from bare prompts and suppress it from prompts that would otherwise produce it, replicating in Qwen3-14B under a pre-registered design. Benign and malign-derived hiding vectors are highly similar (cosine $0.804$ in MiniMax-M2.5; $0.970$ in Qwen3-14B), implying that in these models the disclosure asymmetry arises from differential activation of a shared hiding direction rather than separate mechanisms.
Chinese Translation
思维链(Chain-of-thought, CoT)推理轨迹日益被提出作为AI监督的一种机制:从原则上讲,审查模型推理过程的监控器可以检测到仅凭输出无法发现的不当行为。这一假设前提是:无论接收到何种指令,CoT都会显现模型被指示去做的事情。我们从两个维度检验了这一假设。首先,我们提出了指令遵从差距(Instruction-Compliance Gap, ICG):当隐藏系统提示指令为恶性(malign)与良性(benign)时,模型的CoT显式提及该指令的概率之差。在100个任务对和来自5个模型系列的8个前沿推理模型上,我们一致观察到不对称披露现象——泄露恶性隐藏指令的概率高于良性指令,具体见于Qwen3-14B(Wilcoxon $p=0.0001$,+13.9个百分点)、Qwen3-32B($p=0.0011$,+13.0个百分点)、Qwen3-235B($p=0.035$,+5.8个百分点),并在MiniMax-M2.5和DeepSeek-R1上得到类似结果。该检测器经过两轮独立的盲标验证达到100%精确率;仅阅读推理轨迹的LLM监控器在与无指令对照组的对比中,在全部8个模型上均复现了该不对称性,并在被检测器判定为干净的恶性轨迹中识别出具体指令的比例为82%。其次,通过对比激活添加(Contrastive Activation Addition)在MiniMax-M2.5中提取的导向向量,能够因果性地在无指令提示下诱发隐藏行为,并抑制原本会触发隐藏行为的提示中的该行为,且在一项预注册设计下于Qwen3-14B中成功复现。良性来源与恶性来源的隐藏向量高度相似(MiniMax-M2.5中余弦相似度为0.804;Qwen3-14B中为0.970),这表明在这些模型中,披露不对称源于一个共享隐藏方向的不同程度激活,而非独立的机制。
cs.LG / 23 / 2608.29093

Titans-QFWP: A Regime-Aware Hybrid Quantum Fast Weight Programmer for Portfolio Optimization

Titans-QFWP:一种面向市场状态感知的混合量子快速权重编程投资组合优化方法
Hung, Ming-Kai, Chen, Jun-Hao, Tsai, Yun-Cheng, Chen, Samuel Yen-Chi
Abstract
We propose Titans-QFWP, a hybrid reinforcement learning architecture integrating a Quantum Fast Weight Programmer with Titans-style memory (Persistence, Surprise, and Forgetting) for adaptive portfolio optimization. To address high-dimensional market features, we introduce an enhanced A3C^2 framework with Hungarian-aligned K-means clustering and scaled log-return rewards. Evaluated on 468 S&P 500 stocks under an Equal-Parameter-Count (EPC) benchmark with approximately 3,000 trainable parameters, Titans-QFWP achieves strong performance (median ARR 0.4260, Calmar 8.5504, IR 0.8427). Ablation results reveal that quantum gating fundamentally reshapes memory component roles, with Persistence supporting drawdown control, Surprise contributing to return generation, and Forgetting providing additional stabilization. By stabilizing these quantum representations, the model enables defensive allocation during market drawdowns while preserving upside potential.
Chinese Translation
我们提出Titans-QFWP,一种混合强化学习架构,将量子快速权重编程器(Quantum Fast Weight Programmer)与Titans风格的记忆机制(持久性、惊奇性与遗忘性)相结合,用于自适应投资组合优化。为应对高维市场特征,我们引入了增强的A3C²框架,结合匈牙利对齐的K-means聚类和缩放对数收益奖励。在468只标普500股票上、以约3000个可训练参数的等参数量(EPC)基准进行评估,Titans-QFWP取得了优异表现(年化收益率中位数0.4260,Calmar比率8.5504,信息比率0.8427)。消融实验表明,量子门控从根本上重塑了记忆组件的角色:持久性支撑回撤控制,惊奇性有助于收益生成,遗忘性则提供额外的稳定作用。通过稳定这些量子表征,该模型能够在市场回撤期间实现防御性配置,同时保留上行收益潜力。
cs.LG / 24 / 2608.29096

Development of an Autonomous AI Coding Agent using Monte Carlo Tree Search (MCTS) and Gemini LLM Frameworks

基于蒙特卡洛树搜索(MCTS)与Gemini大语言模型框架的自主AI编程智能体的开发
Game, Pravin, Ramakrishnan, Vipin, Wagh, Prathamesh
Abstract
The ongoing changes in software engineering requirements have created a substantial need for automated tools which can create secure source code from natural language input. The performance of traditional Large Language Models (LLMs) becomes limited by their "one-shot" capability which results in logical hallucinations together with reduced algorithmic performance during complicated operations. The research presents an autonomous AI Coding Agent which establishes a connection between LLM-generated content and production-ready software through its organized methodology for decision making. Our framework uses the Gemini 2.5 Flash API for essential reasoning capabilities while employing a tailored Monte Carlo Tree Search (MCTS) method to solve code generation challenges as a search operation. The agent uses a "Self-Critic" evaluator system to test different implementation methods which it ranks according to their accuracy and difficulty level before it improves its operational framework through backpropagation. The system operates through a Flask-based web interface which delivers instant feedback together with syntax highlighting features. Our experimental results show that the MCTS-based method achieves a 92% success rate on complex logical prompts while surpassing standard zero-shot generation models.
Chinese Translation
软件工程需求的持续变化产生了对自动化工具的巨大需求,这类工具能够从自然语言输入生成安全的源代码。传统大语言模型(LLM)的性能受限于其"一次性生成"(one-shot)能力,导致在复杂任务中出现逻辑幻觉以及算法性能下降。本研究提出了一种自主AI编程智能体(AI Coding Agent),通过有组织的决策方法,在大语言模型生成的内容与可用于生产的软件之间建立连接。我们的框架采用Gemini 2.5 Flash API提供核心推理能力,并运用一种定制的蒙特卡洛树搜索(MCTS)方法,将代码生成问题转化为搜索操作来求解。该智能体采用"自我批评"(Self-Critic)评估器系统来测试不同的实现方案,根据其准确性和难度水平进行排序,并通过反向传播来改进其运行框架。系统通过基于Flask的网页界面运行,提供即时反馈和语法高亮功能。实验结果表明,基于MCTS的方法在复杂逻辑提示上达到了92%的成功率,超越了标准的零样本(zero-shot)生成模型。
cs.LG / 25 / 2608.29099

Temperature-Adaptive Transformed Teacher Matching

温度自适应的变换教师匹配
Aizawa, Hiroaki, Hayashi, Yoshikazu
Abstract
Temperature scaling is a core component of knowledge distillation, yet its role and effect are still not fully understood. Transformed Teacher Matching (TTM) clarifies the role of temperature scaling by applying it only to the teacher distribution and interpreting the resulting objective as standard distillation with an implicit R\'enyi entropy regularization on the student. However, TTM still relies on a fixed temperature and does not specify how the teacher-side temperature should be adapted for individual samples. In this paper, we introduce a sample-wise inverse-temperature update for TTM by locally minimizing the Kullback-Leibler divergence between the temperature-scaled teacher distribution and the student's prediction. We derive closed-form first and second derivatives with respect to the inverse temperature, and show that they can be expressed using variance and covariance statistics of centered teacher and student logits under the transformed teacher weighting. This yields an efficient curvature-aware update that requires one softmax evaluation and a constant number of class-wise weighted sums. Experiments on standard image classification distillation benchmarks show that our temperature adaptation generally improves TTM and WTTM, while remaining competitive with or outperforming prior temperature-adaptive distillation baselines.
Chinese Translation
温度缩放(temperature scaling)是知识蒸馏的核心组成部分,然而其作用与效果尚未被完全理解。变换教师匹配通过仅将温度缩放应用于教师分布,并将所得目标解释为对学生施加隐式 R'enyi 熵正则化的标准蒸馏,从而阐明了温度缩放的作用。然而,TTM 仍然依赖于固定的温度,并且没有说明应如何针对单个样本调整教师侧的温度。在本文中,我们为 TTM 引入了一种逐样本的逆温度更新方法,通过局部最小化温度缩放后的教师分布与学生预测之间的 Kullback-Leibler 散度来实现。我们推导了逆温度的一阶和二阶导数的闭式表达式,并证明它们可以用变换教师权重下居中化的教师和学生对数 logits 的方差和协方差统计量来表示。由此得到一种高效的曲率感知更新方法,只需一次 softmax 计算和常数数量的按类别加权和。在标准图像分类蒸馏基准上的实验表明,我们的温度自适应方法总体上提升了 TTM 和 WTTM 的性能,同时与先前的温度自适应蒸馏基线相比具有竞争力或表现更优。
cs.LG / 26 / 2608.29107

PathGuide: Dynamic Classifier-Free Guidance via On-Policy Transport Alignment

PathGuide:基于在线策略传输对齐的动态无分类器引导方法
Nevo, Avishag, Hazan, Tamir
Abstract
While modern generative models excel at modeling complex data, precise inference-time control in conditional generation remains a critical challenge. Classifier-free guidance (CFG) is a primary mechanism for such control, yet it is typically treated as a static tuning parameter. In flow-based models, however, the guidance scale fundamentally dictates the velocity field and the resulting probability path, making guidance selection a dynamic path-optimization problem. We introduce PathGuide, a framework that reformulates scalar CFG selection as an on-policy transport problem. Leveraging the weak form of the continuity equation, we derive a selection criterion with a direct path-correctness interpretation: we prove that if the guided field is weakly equivalent to the exact conditional field along the generated rollout, the sampler's path coincides with the target conditional law. For scalar CFG, this criterion yields a strictly quadratic local objective with an efficient, closed-form selector for each solver interval. PathGuide enables optimal guidance scales to be computed and used online during generation or fitted offline as a reusable piecewise-constant schedule. We validate our method on low-resolution image manifolds and controlled settings across various continuous-time flow constructions, demonstrating that this transport-based selector improves path alignment and sample fidelity over both fixed and state-of-the-art adaptive guidance baselines.
Chinese Translation
尽管现代生成模型在建模复杂数据方面表现出色,但在条件生成中实现精确的推理时控制仍然是一个关键挑战。无分类器引导(Classifier-Free Guidance,CFG)是实现此类控制的主要机制,但它通常被视为一个静态的调节参数。然而,在基于流的模型中,引导尺度从根本决定了速度场及由此产生的概率路径,这使得引导选择成为一个动态的路径优化问题。我们提出了PathGuide,这是一个将标量CFG选择重新表述为在线策略传输问题的框架。利用连续性方程的弱形式,我们推导出一个具有直接路径正确性解释的选择准则:我们证明,如果被引导场沿生成的轨迹与精确的条件场弱等价,则采样器的路径与目标条件分布相一致。对于标量CFG,该准则产生一个严格的二次局部目标,并针对每个求解器区间给出高效闭式选择器。PathGuide支持在生成过程中在线计算并使用最优引导尺度,或离线拟合为可复用的分段常数调度。我们在低分辨率图像流形以及受控设置下、跨多种连续时间流构建方式上验证了我们的方法,结果表明这种基于传输的选择器在路径对齐和样本保真度方面均优于固定引导和最先进的自适应引导基线。
cs.LG / 27 / 2608.29110

Explainable Machine Learning for Broadband Adoption Disparities: Tract-Level Prediction and SHAP-Based Factor Profiling

面向宽带普及率差异的可解释机器学习:人口普查区级预测与基于SHAP的因素画像
Han, Xiao
Abstract
The United States has allocated approximately $65 billion through the Infrastructure Investment and Jobs Act for broadband expansion, yet evidence-based methods for targeting these investments remain underdeveloped. This paper presents an explainable machine learning framework for profiling broadband adoption disparities at census-tract granularity across 83,359 tracts nationwide. Using 65 socioeconomic, demographic, and infrastructure features derived from the American Community Survey 2022, we train a LightGBM model under spatial five-fold cross-validation, achieving R^2 = 0.533 and Spearman rho = 0.763; state-held-out cross-validation (51 folds) confirms generalization (R^2 = 0.525). TreeSHAP analysis identifies income and education as the dominant factor group (with the engineered interaction term absorbing attribution from its constituent features), and SHAP-based clustering reveals three exploratory factor profiles: Well-Connected Moderate (~49K tracts), Affordability-Limited Severe (~21K tracts), and Rural-Elderly (~13K tracts). As a screening tool, ML-based tract selection captures 38.0% of the total adoption gap within the top 10% of tracts versus 35.2% for income-only heuristics (+2.8 pp, p < 0.002, county-block bootstrap); in regret-reduction terms, the model closes 19% of the remaining gap between income-only and oracle selection. The primary contribution is the per-tract factor decomposition: SHAP identifies which feature groups (income/education, rurality, age) are most strongly associated with each tract's predicted gap, and informs differentiated investigation. A temporal stability check, training on ACS 2017 and predicting ACS 2022 with zero survey-year overlap, confirms ranking stability (rho = 0.784, noting hyperparameters tuned on 2022 data).
Chinese Translation
美国已通过《基础设施投资和就业法案》拨款约650亿美元用于宽带扩展,然而针对这些投资进行精准配置的循证方法仍然不足。本文提出了一个可解释的机器学习框架,在全国83,359个人口普查区(census tract)的粒度上刻画宽带普及率差异。基于美国社区调查(American Community Survey)2022年数据提取的65个社会经济、人口统计和基础设施特征,我们在空间五折交叉验证下训练了LightGBM模型,取得R^2 = 0.533和Spearman rho = 0.763的成绩;按州留出的交叉验证(51折)证实了模型的泛化能力(R^2 = 0.525)。TreeSHAP分析识别出收入和教育为主导因素组(其中构建的交互项吸收了来自其组成特征的部分归因),基于SHAP的聚类揭示了三种探索性因素画像:良好连接型中等水平(约4.9万个区)、可负担性受限型严重水平(约2.1万个区)以及农村老年型(约1.3万个区)。作为一种筛选工具,基于机器学习的普查区选取在前10%的普查区内捕获了总普及率差距的38.0%,而仅基于收入的启发式方法为35.2%(+2.8个百分点,p < 0.002,县域块自助法);从遗憾降低的角度看,该模型弥合了仅收入方法与理想选择之间剩余差距的19%。本文的主要贡献在于逐区的因素分解:SHAP识别出与每个普查区的预测差距最密切相关的特征组(收入/教育、农村性、年龄),从而为差异化的深入调查提供依据。时间稳定性检验——在ACS 2017数据上训练并以零调查年份重叠的方式预测ACS 2022——证实了排名的稳定性(rho = 0.784,需注意超参数是基于2022年数据调优的)。
cs.LG / 28 / 2608.29188

Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space

入口被锁,内部敞开:RLVR在何处收窄了解空间
Zhou, Qiancheng, Li, Ruizhe
Abstract
Reinforcement learning with verifiable rewards (RLVR) substantially improves single-sample accuracy (pass@1) but causes the policy's solution space to contract, diminishing the returns of test-time scaling. In this work, we investigate where inside a reasoning trajectory this breadth is lost: does the policy fail to access a valid solution family, or does it fail to execute computation once initiated? To disentangle access from execution, we analyze the Countdown task, whose solution space can be exhaustively enumerated into discrete entrance families defined by the first operand and operator, across PPO on Qwen2.5-3B and GRPO on Qwen2.5-3B-Instruct. Across both training setups, solution coverage falls by up to 67%, halving even on problems solved across all checkpoints. We show that this contraction is heavily concentrated at the entrance: per-token likelihood shifts are 11x--16x larger prior to the first arithmetic operation than during downstream reasoning. Supplying only an unselected entrance prefix restores completion rates in low-access families by over an order of magnitude (0.018 -> 0.212 under PPO), demonstrating that alternative solutions remain executable but are no longer initiated. Guided by this localization, we find that while surface prompting fails to recover diversity, entrance-targeted interventions succeed: late-layer parameter interpolation with early checkpoints increases solution coverage by 37% at no loss in pass@1. Finally, we show that early-step entropy collapse recurs across six math benchmarks with 7B and 14B models, but is not an inevitable byproduct of reasoning optimization: an SFT baseline preserves more than double the coverage, and staged SFT--DPO--RLVR pipelines retain early-step entropy. In summary, reasoning breadth is lost at the door, not inside the room. Code: https://github.com/ershiyidian/early-branch-locking.
Chinese Translation
可验证奖励强化学习(RLVR)显著提升了单样本准确率(pass@1),但导致策略的解空间收缩,削弱了测试时扩展的收益。在本工作中,我们研究了推理轨迹中这种多样性在何处丢失:是策略无法访问有效的解族,还是在启动计算后无法执行?为了将“访问”与“执行”区分开来,我们分析了Countdown任务,该任务的解空间可以被穷举为由第一个操作数和算子定义的离散入口族;实验覆盖了在Qwen2.5-3B上的PPO以及在Qwen2.5-3B-Instruct上的GRPO两种训练设置。在两种训练设置下,解的覆盖率最多下降67%,即使在所有检查点上都能解决的问题上也减半。我们证明这种收缩高度集中在入口处:在第一个算术运算之前的逐token似然偏移是下游推理阶段的11至16倍。仅提供一个未被选择的入口前缀,即可将低访问族中的补全率提升一个数量级以上(PPO下从0.018提升至0.212),这表明替代解仍然可执行,只是不再被启动。基于这一定位,我们发现表面提示无法恢复多样性,而针对入口的干预则有效:与早期检查点进行后层参数插值可将解覆盖率提升37%,且不损失pass@1。最后,我们证明早期步骤的熵坍缩在六个数学基准上、使用7B和14B模型时均会复现,但它并非推理优化的必然副产物:SFT基线保留了超过两倍的覆盖率,而分阶段的SFT--DPO--RLVR流水线则保留了早期步骤的熵。总之,推理广度是在门口丢失的,而非在房间内部。代码:https://github.com/ershiyidian/early-branch-locking。
cs.LG / 29 / 2608.29193

HalluPrism: When Multimodal Uncertainty Should Diagnose, Not Decide

HalluPrism:当多模态不确定性应用于诊断而非决策时
Prakash, Aman, Dasgupta, Sourish, Chakraborty, Tanmoy
Abstract
Multimodal Large Language Models (MLLMs) can assign similar confidence to answers that fail for different reasons. We propose HalluPrism, a behavioral diagnostic that re-runs an answer after visual degradation, blank-image replacement, and grounding or relation checks. These targeted probes yield a signature over visual-perturbation sensitivity (V ), image-removal confidence retention (L), and grounding/relation-probe instability (A). Across 58K+ examples from four benchmarks and four MLLMs, image-removal confidence retention is most prevalent, while grounding/relation-probe instability better separates failure families. Only 18 of 48 source-target checks are diagonally aligned, so the coordinates should be interpreted jointly rather than as independent causal sources. With the dataset fixed, the joint signature improves failure-family AUROC from 0.634 to 0.769 on HallusionBench and from 0.707 to 0.817 on VizWiz, with smaller gains on POPE and VSR. In pooled XGBoost analysis, AUROC rises from 0.78 with scalar confidence to 0.95 with (V, L, A) and 0.97 when confidence is added. The same signature does not automatically improve correctness ranking. The three tested direct scalarizations can harm it. These results separate failure diagnosis from abstention scoring: multimodal uncertainty should characterize failure structure before it is used to decide whether to abstain or correct.
Chinese Translation
多模态大语言模型(MLLM)可能对因不同原因而失败的答案赋予相似的置信度。我们提出了HalluPrism,一种行为诊断方法,通过视觉退化、空白图像替换以及接地(grounding)或关系检查后重新运行答案来进行分析。这些针对性探针产生一个特征签名,涵盖视觉扰动敏感性(V)、图像移除置信度保持度(L)以及接地/关系探针不稳定性(A)。在来自四个基准和四个MLLM的58K+样本上,图像移除置信度保持度最为普遍,而接地/关系探针不稳定性则能更好地区分不同失败类别。在48个源-目标检查中仅有18个呈对角对齐,因此这些坐标应被联合解读,而非视为独立的因果来源。在固定数据集的条件下,联合特征签名将HallusionBench上的失败类别AUROC从0.634提升至0.769,VizWiz上从0.707提升至0.817,在POPE和VSR上的提升较小。在汇总的XGBoost分析中,AUROC从使用标量置信度的0.78提升至使用(V, L, A)的0.95,加入置信度后达到0.97。同样的特征签名并不能自动提升正确性排序,三种经过测试的直接标量化方法甚至会损害它。这些结果将失败诊断与弃答评分区分开来:多模态不确定性在用于决定是否弃答或纠正之前,应先刻画失败的结构。
cs.LG / 30 / 2608.29197

PokaiTrainer: Scaling Belief-State Search to Competitive Pok\'emon VGC

PokaiTrainer:将信念状态搜索扩展至竞技宝可梦VGC
Yu, Max
Abstract
Decision-time equilibrium search carried poker to superhuman play, but it has so far relied on tractable subgames: a handful of actions per decision, chance confined to card deals, one player moving at a time. Competitive Pok\'emon in its official doubles format (VGC) breaks all three assumptions at once. Both players act simultaneously from joint menus in the hundreds, each joint action resolves to hundreds of stochastic outcomes, and the opponent's reserves and stat allocations are hidden. We set out to build a strong VGC agent and report what that took. PokaiEngine, our Rust battle engine, enumerates a joint action's full weighted outcome distribution in one pass, at ${\sim}99\%$ parity with Pok\'emon Showdown and a fraction of the cost of sampling it. On top of the engine, PokaiTrainer adapts Student of Games to this scale, solving every decision as a Bayesian matrix game over public belief states and growing subgames under an explicit compute budget. On the live Showdown best-of-three ladder, the agent wins 59% of 150 sets against a human field averaging ${\sim}1320$ Elo. It settles into a 1350-1400 Elo band, and at its peak briefly entered the format's top 500.
Chinese Translation
决策时均衡搜索(decision-time equilibrium search)曾使扑克达到超人水平,但迄今为止它依赖于可处理的子博弈:每个决策仅有少数可选动作、随机性仅限于发牌、且每次只有一名玩家行动。官方双打形式的竞技宝可梦(VGC)同时打破了这三个假设。双方玩家从数百个组合动作菜单中同时行动,每个组合动作会解析为数百种随机结果,且对手的后备宝可梦和努力值分配是隐藏的。我们的目标是构建一个强大的VGC智能体,并报告实现这一目标所需的努力。我们的Rust对战引擎PokaiEngine可一次性枚举一个组合动作的完整加权结果分布,与Pokémon Showdown的一致性约为99%,且成本仅为采样方法的一小部分。在该引擎之上,PokaiTrainer将Student of Games适配到这一规模,将每个决策求解为基于公共信念状态(public belief states)的贝叶斯矩阵博弈,并在明确的计算预算下扩展子博弈。在Showdown实时三局两胜天梯中,该智能体在与平均Elo约为1320的人类玩家对战的150场比赛中获胜59%。其Elo稳定在1350-1400区间,并在巅峰时短暂进入了该格式的前500名。
cs.LG / 31 / 2608.29247

RL-FAT: Reinforcement Learning for Fair Adversarial Training

RL-FAT:基于强化学习的公平对抗训练
Medi, Tejaswini, Mikeladze, Levan, Keuper, Margret
Abstract
Deep neural networks remain highly vulnerable to adversarial perturbations, and adversarial training (AT) has become a widely used approach for improving robustness. However, improvements in average robust accuracy often mask substantial class-wise disparities: while some classes become more robust, others may remain disproportionately vulnerable under attack. This imbalance raises an important adversarial fairness concern, particularly in vision tasks where reliable robustness is expected across all categories. To address this challenge, we propose \textbf{RL-FAT}, a reinforcement-learning-inspired fair adversarial training framework that uses policy-gradient based feedback from adversarial predictions. RL-FAT interprets the prediction distribution as a policy and combines correctness-based prediction rewards with class-wise value estimates to compute class-specific advantages for policy-gradient optimization. This enables the model to adaptively focus on class-wise misclassification. Furthermore, we introduce a fairness-emphasis adversarial loss that assigns stronger training pressure to classes with high adversarial loss, thereby mitigating class-wise robustness disparity. By combining reinforcement-driven adaptation with fairness-emphasis regularization, RL-FAT improves adversarial robustness while promoting a more balanced robustness distribution across classes. Extensive experiments demonstrate that our method achieves competitive robust accuracy and substantially reduces class-wise robustness imbalance compared with standard adversarial training baselines.
Chinese Translation
深度神经网络仍然极易受到对抗扰动的影响,而对抗训练(Adversarial Training, AT)已成为提升鲁棒性的广泛使用的方法。然而,平均鲁棒准确率的提升往往掩盖了类别间的显著差异:某些类别变得更具鲁棒性,而另一些类别在攻击下可能仍然不成比例地脆弱。这种不平衡引发了一个重要的对抗公平性问题,尤其是在期望所有类别都具备可靠鲁棒性的视觉任务中。为应对这一挑战,我们提出了RL-FAT,一个受强化学习启发的公平对抗训练框架,它利用来自对抗预测的基于策略梯度的反馈。RL-FAT将预测分布视为一个策略,并将基于正确性的预测奖励与类别价值估计相结合,计算类别特定的优势函数以进行策略梯度优化。这使模型能够自适应地聚焦于类别间的误分类。此外,我们引入了一种强调公平性的对抗损失,对具有高对抗损失的类别施加更强的训练压力,从而缓解类别间的鲁棒性差异。通过将强化学习驱动的自适应机制与强调公平性的正则化相结合,RL-FAT在提升对抗鲁棒性的同时,促进了跨类别更均衡的鲁棒性分布。大量实验表明,与标准对抗训练基线相比,我们的方法实现了具有竞争力的鲁棒准确率,并显著降低了类别间的鲁棒性不平衡。
cs.LG / 32 / 2608.29262

Adaptive Multi-Branching for Shallow Decision Tree Induction

用于浅层决策树生成的自适应多分支方法
Park, Hanul, Choi, Jeonghoon, Kim, Juseong, Sel, Sanghun, Song, Giltae
Abstract
Decision trees are attractive for tabular prediction tasks because each prediction follows an interpretable sequence of feature-threshold tests. Under a strict maximum-depth budget, however, conventional binary trees can be under-expressive, since each internal node makes only a single threshold decision. We study shallow-depth tree induction, where the goal is to improve accuracy while keeping root-to-leaf paths short. We propose the Multi-Branch Neural Decision Tree with Adaptive Pruning (MBNDT), a single axis-aligned tree trained end-to-end with differentiable multi-way splits. Each internal node learns ordered thresholds over a selected feature and a branch mask that adapts its effective arity, and the trained model is converted to a deterministic single-path tree for inference. Across 21 OpenML binary-classification benchmarks, MBNDT achieves the best average rank and mean balanced accuracy among depth-constrained single-tree baselines; a controlled ablation isolates multi-way splitting as the source of the gain. These gains come with an explicit trade-off: MBNDT realizes more leaves than the other single-tree baselines, making it best suited when accuracy under short, bounded decision paths is prioritized over minimal global tree size.
Chinese Translation
决策树在表格数据预测任务中颇具吸引力,因为每次预测都遵循一条可解释的特征-阈值测试序列。然而,在严格的最大深度限制下,传统二叉树的表达能力可能不足,因为每个内部节点只能做出单一的阈值判定。我们研究了浅层深度决策树的生成,其目标是在保持根到叶路径较短的同时提高准确率。我们提出了带自适应剪枝的多分支神经决策树(Multi-Branch Neural Decision Tree with Adaptive Pruning, MBNDT),这是一棵通过可微多路分裂进行端到端训练的单棵轴对齐树。每个内部节点在选定特征上学习有序阈值以及自适应调节其有效分支数的分支掩码,训练后的模型被转换为确定性的单路径树用于推理。在21个OpenML二分类基准数据集上,MBNDT在深度受限的单棵树基线方法中取得了最优的平均排名和平均平衡准确率;受控消融实验证实多路分裂是性能提升的来源。这些提升伴随着明确的权衡:MBNDT产生的叶子节点数量多于其他单棵树基线,因此它最适用于在短且有界的决策路径下优先考虑准确率而非最小化全局树规模的场景。
cs.LG / 33 / 2608.29296

When Do Larger Batches Help Scale LLM Reinforcement Learning?

更大的批次何时有助于扩展大语言模型强化学习?
Li, Ziniu, Wang, Jinbo, Huang, Guanhua, Zhang, Feiyuan, Li, Pengbo, Chen, Alex
Abstract
Larger batches reduce the variance of stochastic gradients per update and are therefore often expected to accelerate training. Yet whether this statistical benefit translates into lower wall-clock time-to-target remains unclear, because each update consumes more samples and may take longer to execute. We study this tradeoff in reinforcement learning for large language models. We separate its algorithmic and systems effects by comparing learning and execution along their natural axes. At the algorithmic level, we compare configurations at equal cumulative sample counts while retuning batch-dependent hyperparameters. Over a bounded range of batch sizes, this procedure yields an approximately batch-size-invariant family whose members follow similar sample-indexed learning trajectories. At the systems level, we exploit the computational asymmetry between rollout generation and training: autoregressive generation is often memory-bandwidth-bound at low concurrency, whereas training work scales approximately with the number of processed tokens. Combining these two views yields a direct decision rule: a larger-batch configuration reduces time-to-target only when its throughput gain exceeds its samples-to-target penalty. Experiments with GRPO and PPO support both sides of this decomposition. At the algorithmic level, square-root learning-rate scaling with Adam produces approximately batch-size-invariant learning curves over a bounded range of batch sizes. At the systems level, larger batches improve generation throughput by up to 2.29x on fixed hardware. In GRPO, combining higher throughput with learning-rate retuning reduces time-to-target by up to 29%, whereas increasing the batch without retuning is slower despite its higher throughput.
Chinese Translation
更大的批次可以降低每次更新随机梯度的方差,因此通常被认为能够加速训练。然而,这种统计上的收益是否能转化为更低的到达目标性能的实际耗时(wall-clock time-to-target)仍不明确,因为每次更新会消耗更多样本,且执行时间可能更长。我们在大语言模型强化学习中研究了这一权衡。我们通过沿各自的自然轴比较学习与执行,将算法效应与系统效应分离开来。在算法层面,我们在相同累计样本数下比较不同配置,同时重新调整与批次相关的超参数。在有界的批次大小范围内,该过程产生了一个近似批次大小不变的学习曲线族,其成员遵循相似的以样本索引的学习轨迹。在系统层面,我们利用了 rollout 生成与训练之间的计算不对称性:自回归生成在低并发时通常受内存带宽限制,而训练负载则近似与处理的 token 数量成比例。结合这两个视角,我们得到了一个直接的决策规则:只有当更大批次配置的吞吐量增益超过其到达目标所需样本数的惩罚时,它才能缩短到达目标的时间。基于 GRPO 和 PPO 的实验支持了该分解的两方面结论。在算法层面,采用 Adam 时学习率按平方根缩放可在有界的批次大小范围内产生近似批次大小不变的学习曲线。在系统层面,在固定硬件上更大的批次可将生成吞吐量提升至 2.29 倍。在 GRPO 中,将更高吞吐量与学习率重新调整相结合可使到达目标的时间最多缩短 29%,而仅增大批次而不重新调整学习率尽管吞吐量更高,反而更慢。
cs.LG / 34 / 2608.29302

A Spectral Identifiability Threshold for Dissipative Rate Recovery from Truncated Liouvillian Spectra

截断刘维尔谱中耗散速率恢复的谱可辨识性阈值
Ji, Yujun, Chakraborty, Somyajit
Abstract
Open quantum systems lose energy and phase coherence through different dissipative processes, but these processes can produce overlapping dynamical signatures. The Liouvillian spectrum summarizes how such a system relaxes, yet it is not obvious how much of that spectrum is needed to distinguish the underlying dissipation rates. We study this question for amplitude damping and dephasing in a six-qubit Lindblad model whose spectrum can be derived analytically. We retain only the slowest non-steady spectral modes and ask how many are required before each dissipative rate becomes recoverable. We show that population modes contain no dephasing information, which creates a lower bound of D = 2^n retained modes for uniform dephasing identifiability in the relevant rate regime. The measured recovery threshold reaches this bound at n = 4,5,6, while n = 3 remains above it. At n = 6, least squares achieves a mean joint absolute error of order 10^-9, compared with 4.355 x 10^-4 for four tabular learning methods. Robustness tests show that this advantage weakens when the spectra are perturbed and when a transverse field breaks the commuting structure. These results show that the amount and structure of retained spectral information can determine whether dissipative parameters are recoverable, independently of the estimator used. The present conclusions apply to noise-free simulator spectra rather than measurement-derived spectra.
Chinese Translation
开放量子系统通过不同的耗散过程损失能量和相位相干性,但这些过程可能产生相互重叠的动力学特征。刘维尔谱概括了系统的弛豫行为,然而尚不清楚需要多少谱信息才能区分潜在的耗散速率。我们在一个可解析导出谱的六量子比特 Lindblad 模型中研究了幅度阻尼与退相干(去相位)情形下的这一问题。我们仅保留最慢的非稳态谱模式,并探究恢复每个耗散速率需要多少个模式。我们证明,布居数模式不包含任何退相干信息,这使得在相关速率区间内,均匀退相干可辨识性存在 D = 2^n 个保留模式的下界。实测恢复阈值在 n = 4、5、6 时达到该下界,而 n = 3 时仍高于该下界。在 n = 6 时,最小二乘法实现了量级为 10^-9 的平均联合绝对误差,而四种表格学习方法为 4.355 x 10^-4。鲁棒性测试表明,当谱受到扰动、或横向场破坏对易结构时,这一优势会减弱。这些结果表明,保留谱信息的数量与结构可以决定耗散参数是否可恢复,而与所使用的估计器无关。本文结论适用于无噪声的模拟器谱,而非由测量导出的谱。
cs.LG / 35 / 2608.29304

MEL: Coordinate-Preserving EEG Tokenization for fMRI Translation

MEL:面向fMRI转换的坐标保持式脑电(EEG)标记化方法
Liu, Xiangyu, Yan, Zeting, Yin, Zhitong, Li, Boyang, Zhang, Xi
Abstract
Translating electroencephalography (EEG) into functional magnetic resonance imaging (fMRI) is important for medical neuroimaging, clinical brain-state monitoring, and multimodal neural decoding, because it aims to infer spatially organized hemodynamic activity from fast and accessible electrophysiological recordings. Existing EEG-to-fMRI studies mainly pursue stronger decoders, but the problem is also constrained by a representation-interface mismatch: fMRI responses are delayed, temporally integrated, and spatially distributed, whereas generic EEG encodings often entangle temporal lag, channel identity, and frequency-band structure. We propose Multi-band EEG Latent-state Tokenization (MEL), a coordinate-preserving EEG representation framework that anchors each target fMRI response to its preceding EEG history and organizes it into lag-channel-frequency neural-state tokens. By explicitly capturing hemodynamic latency and spectral-spatial dynamics, MEL aligns fMRI-pertinent EEG representations with capacity-controlled readouts without depending entirely on model scaling. Experiments on VU EEG-fMRI benchmarks and external Oddball data show that MEL improves prediction over strong NeuroBOLT baselines. Ablations and controls further indicate that the gains come from structured EEG representation rather than leakage, shortcut statistics, or decoder capacity.
Chinese Translation
将脑电图(EEG)转换为功能磁共振成像(fMRI)对于医学神经影像、临床脑状态监测和多模态神经解码具有重要意义,因为其目标是从快速且易于获取的电生理记录中推断出具有空间组织性的血流动力学活动。现有的EEG到fMRI研究主要致力于构建更强的解码器,但该问题同时受到表征接口不匹配的制约:fMRI响应具有延迟性、时间累积性和空间分布性,而通用的EEG编码往往将时间滞后、通道身份和频段结构纠缠在一起。我们提出多频段脑电潜在状态标记化方法(Multi-band EEG Latent-state Tokenization, MEL),这是一种坐标保持式的EEG表征框架,它将每个目标fMRI响应锚定于其之前的EEG历史,并将其组织为滞后-通道-频率的神经状态标记。通过显式捕捉血流动力学延迟和频谱-空间动态,MEL使与fMRI相关的EEG表征与容量受控的读出机制对齐,而不完全依赖于模型规模的扩展。在VU EEG-fMRI基准数据集和外部Oddball数据上的实验表明,MEL相比强大的NeuroBOLT基线方法提升了预测性能。消融实验和对照实验进一步表明,性能提升源于结构化的EEG表征,而非数据泄漏、捷径统计或解码器容量。
cs.LG / 36 / 2608.29349

Information-Based Calibration of Uncertainty Quantification in Product-of-Experts Gaussian Process Models

基于信息的产品专家高斯过程模型不确定性量化的校准方法
Ong, Yean Hoon, Barucca, Paolo, Pan, Wei, Wang, Jun
Abstract
Gaussian process (GP) regression with a single global GP (GP-glo) incurs cubic computational cost, limiting scalability to large datasets. Product-of-experts GP models (GP-pro), which combine local GP models to capture global correlations, alleviate this computational burden. However, training local experts on disjoint data subsets can lead to overestimated posterior variances. We propose GP-pro-c, a product-of-experts GP model that calibrates these variances using an information-based method. The method exploits the monotonicity and submodularity of information gain in GPs to define a calibration ratio that reduces the posterior variance of individual local GP models. We evaluate GP-pro-c using negative log-likelihood (NLL), root mean squared error (RMSE), and expected normalised calibration error (ENCE). Experiments on four synthetic functions and six regression datasets show that GP-pro-c achieves average reductions of 2.3% in NLL and 12.0% in ENCE compared with the uncalibrated GP-pro model. The proposed method mitigates posterior variance overestimation while maintaining predictive accuracy and reducing computational complexity. GP-pro-c provides a promising approach for uncertainty estimation in scalable GP models and may serve as a useful surrogate model for Bayesian optimisation with high-dimensional and large-scale data.
Chinese Translation
采用单一全局高斯过程(GP-glo)的高斯过程回归计算代价为三次方量级,限制了其在大规模数据集上的可扩展性。产品专家高斯过程模型(GP-pro)通过组合局部GP模型来捕捉全局相关性,从而缓解了这一计算负担。然而,在不相交的数据子集上训练局部专家可能导致后验方差被高估。我们提出GP-pro-c,一种利用基于信息的方法对上述方差进行校准的产品专家GP模型。该方法利用GP中信息增益的单调性与次模性,定义了一个校准比率以降低各局部GP模型的后验方差。我们采用负对数似然(NLL)、均方根误差(RMSE)和期望归一化校准误差(ENCE)对GP-pro-c进行评估。在四个合成函数和六个回归数据集上的实验表明,与未校准的GP-pro模型相比,GP-pro-c在NLL上平均降低2.3%,在ENCE上平均降低12.0%。所提出的方法在保持预测精度并降低计算复杂度的同时,缓解了后验方差的高估问题。GP-pro-c为可扩展GP模型中的不确定性估计提供了一种有前景的方法,并可作为高维和大规模数据下贝叶斯优化的有效代理模型。
cs.LG / 37 / 2608.29360

Spatial Entropy based Partitioning for Spatiotemporal Graph Unlearning

基于空间熵划分的时空图遗忘学习
Guo, Qiming, Sun, Wenbo, Wang, Ye, Wang, Wenlu
Abstract
Spatiotemporal graphs underpin applications such as traffic forecasting, weather forecasting, and healthcare monitoring. Privacy regulations such as the GDPR and the CCPA require the complete removal of unauthorized data from trained models, but achieving this on a spatiotemporal graph is difficult: because information propagates globally through both spatial and temporal message passing, fully erasing a node's influence forces costly full-graph retraining. ST-graph unlearning requires both exactness and efficiency. We propose IsleNet, which uses spatial-entropy-guided partitioning to create balanced, locally coherent subgraphs and reconnects them with lightweight virtual edges. Upon an unlearning request, only the affected subgraph encoder and virtual-edge layer are retrained, ensuring exact removal with low cost. Experiments on four real-world benchmarks show that IsleNet attains up to 94% of full-graph accuracy while reducing unlearning time by up to an order of magnitude. Our code is publicly available at https://github.com/wenlu-lab/STGraphUnlearning.
Chinese Translation
时空图是交通预测、天气预报和健康监测等应用的基础。诸如《通用数据保护条例》(GDPR)和《加州消费者隐私法案》(CCPA)等隐私法规要求从已训练模型中完全删除未授权数据,但在时空图上实现这一点十分困难:由于信息通过空间和时间消息传递进行全局传播,完全消除某个节点的影响需要进行代价高昂的全图重训练。时空图遗忘学习(ST-graph unlearning)需要同时满足精确性和高效性。我们提出了IsleNet,它利用空间熵引导的划分方法生成均衡且局部内聚的子图,并通过轻量级虚拟边将它们重新连接。在收到遗忘请求时,仅需重训受影响的子图编码器和虚拟边层,从而以较低的成本确保精确删除。在四个真实世界基准数据集上的实验表明,IsleNet在达到全图训练准确率最高94%的同时,将遗忘时间最多减少了一个数量级。我们的代码已在 https://github.com/wenlu-lab/STGraphUnlearning 公开。
cs.LG / 38 / 2608.29369

Unlearning on Spatio-Temporal Graphs through Subgraph Virtual Edge Reconstruction

基于子图虚拟边重构的时空图机器遗忘方法
Guo, Qiming, Sun, Wenbo, Pan, Chen, Wang, Ye, Wang, Wenlu
Abstract
Spatio-temporal graphs are widely used in modeling complex dynamic processes such as temporal forecasting, molecular dynamics, and healthcare monitoring. Recently, stringent privacy regulations such as GDPR and CCPA have introduced significant new challenges for existing spatio-temporal graph models, requiring complete unlearning of unauthorized data. Since each node in a spatio-temporal graph diffuses information globally across both spatial and temporal dimensions, existing unlearning methods primarily designed for static graphs and localized data removal cannot efficiently erase a single node without incurring costs nearly equivalent to full model retraining. To address this, we propose CallosumNet, a spatio-temporal graph unlearning framework biologically inspired by the corpus callosum structure. CallosumNet makes two key technical contributions: (1) it reconstructs subgraphs using biologically-inspired virtual edges; and (2) it restores interlinked spatio-temporal dependencies among subgraphs via a lightweight meta-graph integration layer. Empirical results on four diverse real-world datasets show that CallosumNet achieves complete unlearning while maintaining accuracy very close to the gold model. The code is publicly available at https://github.com/wenlu-lab/STGraphUnlearning.
Chinese Translation
时空图被广泛用于建模复杂的动态过程,如时间序列预测、分子动力学和医疗健康监测。近年来,GDPR 和 CCPA 等严格的隐私法规给现有的时空图模型带来了重大新挑战,要求对未授权数据进行完全遗忘(unlearning)。由于时空图中的每个节点在空间和时间两个维度上全局扩散信息,现有主要针对静态图和局部数据删除设计的遗忘方法无法高效地擦除单个节点,其代价几乎等同于对模型进行完全重训练。为解决这一问题,我们提出了 CallosumNet——一个受胼胝体(corpus callosum)结构生物学启发而设计的时空图遗忘框架。CallosumNet 做出了两项关键技术贡献:(1)利用受生物学启发的虚拟边对子图进行重构;(2)通过轻量级的元图集成层恢复子图之间相互关联的时空依赖关系。在四个多样化的真实世界数据集上的实验结果表明,CallosumNet 能够实现完全遗忘,同时保持与黄金模型(gold model)非常接近的准确率。代码已公开于 https://github.com/wenlu-lab/STGraphUnlearning。
cs.LG / 39 / 2608.29388

Fully Distributed GNE Algorithms for Multi-Robot Placement without Consensus on Multipliers

无需乘子一致性的多机器人放置全分布式广义纳什均衡算法
Yin, Shao-An, Hong, Mingyi, Elia, Nicola
Abstract
Recent machine learning research has increasingly focused on equilibrium analysis in non-cooperative games rather than solely on optimal solutions. Many such problems involve shared constraints and can be formulated as Generalized Nash Equilibrium Problems (GNEPs). For strongly monotone games, existing methods compute consensus-based variational GNEs (v-GNEs) by exchanging Lagrange multipliers. We propose a fully distributed continuous-time algorithm for shared linear equality constraints that converges without multiplier exchange and reaches any GNE, reducing communication overhead and improving privacy. Discrete-time schemes are also provided, and the method is validated on a multi-robot placement task.
Chinese Translation
近年来,机器学习研究日益关注非合作博弈中的均衡分析,而不再仅仅关注最优解。许多此类问题涉及共享约束,可被建模为广义纳什均衡问题(GNEP)。对于强单调博弈,现有方法通过交换拉格朗日乘子来计算基于一致性的变分广义纳什均衡(v-GNE)。我们针对共享线性等式约束提出了一种全分布式连续时间算法,该算法无需交换乘子即可收敛,并能达到任意广义纳什均衡,从而降低了通信开销并提高了隐私性。此外,我们还给出了离散时间算法方案,并在多机器人放置任务上验证了该方法的有效性。
cs.LG / 40 / 2608.29411

Where Induction Runs Out: Description-Length Difficulty and the Memorisation Gap in Integer-Sequence Benchmarks

归纳的尽头:整数序列基准中的描述长度难度与记忆化差距
Ganeshan, Sabilashan
Abstract
Integer sequences from the On-Line Encyclopedia of Integer Sequences (OEIS) are increasingly used to benchmark mathematical reasoning in language models. We ask what such benchmarks actually measure, using an exactly computable reference learner: two-part minimum description length (MDL) over the class of P-recursive (holonomic) recurrences, evaluated on every prefix of a sequence as terms arrive. Three findings follow. First, MDL difficulty is a parameter count. The discovery point nd, the first prefix length at which a symbolic hypothesis beats verbatim storage, is predicted almost exactly by a combinatorial identifiability bound on the selected operator's order and degree. It is invariant to term magnitude: scaling Fibonacci over twelve orders of magnitude leaves nd unchanged, because a hypothesis must encode its own initial conditions and the magnitude cancels. Second, at scale the learner exhibits a regime our curated corpus could not produce even once: across 20,000 OEIS sequences, 89.98% of those that fit a recurrence on some prefix fit none at full length. We call this the wilderness -- induction acquires a theory, loses it, and never recovers. Third, evaluating three language models on sequences stratified by these MDL regimes refuted our pre-registered hypothesis: models do not confabulate where MDL reports no theory, but hedge appropriately. Confident errors are inverted, concentrating on the easy stratum, where apparent competence tracks recognition of the sequence rather than induction of its rule. OEIS-derived benchmarks therefore substantially measure memorisation, and MDL supplies a cheap, contamination-free difficulty signal they currently lack. Code and data are released.
Chinese Translation
来自整数序列在线百科全书(OEIS)的整数序列正被越来越多地用于对语言模型的数学推理能力进行基准测试。我们借助一个可精确计算的参照学习器来探究此类基准究竟测量了什么:即在P-递归(全纯)递推关系类上的两部分最小描述长度(MDL),并在序列每一项到来时对所有前缀进行评估。由此得出三项发现。第一,MDL难度可由参数数量刻画。发现点nd,即符号化假设首次优于逐字存储的前缀长度,几乎可以被关于所选算子阶数与次数的组合可辨识性界精确预测。该项与项的大小无关:将斐波那契序列缩放十二个数量级后nd保持不变,因为假设必须编码其自身的初始条件,而数值大小被抵消。第二,在大规模数据下,该学习器表现出一种我们精心策划的语料库中甚至一次都无法出现的现象:在20,000条OEIS序列中,89.98%在某一前缀上可拟合递推关系的序列在全长上却无法拟合任何递推关系。我们将这一区域称为荒野——归纳获得了一个理论,又将其丢失,且再也无法恢复。第三,在按这些MDL区域分层的序列上评估三个语言模型,推翻了我们预注册的假设:在MDL未报告任何理论的区域,模型并未虚构,而是恰当地保持谨慎。高置信度的错误反而呈现倒置分布,集中于简单层,其中表面上的能力反映的是对序列的识别而非对其规则的归纳。因此,源自OEIS的基准在很大程度上测量的是记忆,而MDL提供了它们目前所缺乏的一种廉价且无污染的难度信号。代码与数据已公开。
cs.LG / 41 / 2608.29419

Scalable Clinical Data Infrastructure and Comparative ML Evaluation for Hospitalisation Risk Prediction in Elderly Patients with Multiple Long-Term Conditions using CPRD

基于CPRD的老年多种长期疾病患者住院风险预测的可扩展临床数据基础设施与机器学习比较评估
Aslam, Asra, Chapman, Volodymyr, O'Connell, Maurice M., Abuzour, Aseel S., Abaho, Michael, Bollegala, Danushka, Leeming, Gary, Shantsila, Eduard, Clegg, Andrew, Walker, Lauren E., Buchan, Iain Edward, Relton, Samuel D.
Abstract
Deep learning architectures are increasingly proposed for patient trajectory modeling in electronic health records (EHRs), yet their advantage over simpler, more interpretable models is rarely subjected to rigorous empirical scrutiny in real-world clinical settings. We present a comprehensive patient timeline pipeline applied to elderly patients in CPRD Aurum, incorporating 260 clinical conditions classified via a three-tier automated framework including specialised detection logic for 17 complex conditions. Using this infrastructure, we benchmark Temporal Graph Convolutional Neural Networks (TG-CNN) against Logistic Regression with LASSO regularisation and Random Forests for predicting 12-month all-cause emergency hospitalisation risk, motivated by (but not filtered to) the elevated risk of adverse drug reactions. Under cross-validation, TG-CNN achieves a marginally higher mean AUC-ROC than LASSO (0.712 vs. 0.705), whereas on the held-out test set LASSO achieves the highest discrimination of three models (AUC-ROC 0.733, versus 0.710 for Random Forest and 0.702 for TG-CNN). We show, that discrimination alone is an incomplete criterion for clinical deployment: after Platt calibration, LASSO is the only model with an acceptable calibration slope (0.817), while Random Forest (0.759) and, TG-CNN (0.391) remain substantially miscalibrated. We argue that LASSO, not the highest-discriminating model, is the model best suited to direct clinical deployment. We present lessons for the machine learning and healthcare community regarding data infrastructure, model selection, and value of calibration and interpretability in high-stakes decision support.
Chinese Translation
深度学习架构越来越多地被提出用于电子健康记录(EHR)中的患者轨迹建模,但相对于更简单、更可解释的模型,其优势在真实临床环境中很少受到严格的实证检验。我们提出了一个应用于CPRD Aurum中老年患者的综合患者时间线流水线,纳入260种临床疾病,并通过包含17种复杂疾病专门检测逻辑的三层自动化框架进行分类。基于该基础设施,我们以药物不良反应风险升高为动机(但并未据此筛选人群),将时序图卷积神经网络(TG-CNN)与LASSO正则化逻辑回归和随机森林进行基准比较,用于预测12个月全因急诊住院风险。在交叉验证下,TG-CNN的AUC-ROC均值略高于LASSO(0.712 vs. 0.705);而在保留测试集上,LASSO取得了三个模型中最高的判别能力(AUC-ROC为0.733,随机森林为0.710,TG-CNN为0.702)。我们表明,仅凭判别能力是临床部署的不完整标准:经过Platt校准后,LASSO是唯一具有可接受校准斜率(0.817)的模型,而随机森林(0.759)和TG-CNN(0.391)仍存在严重的校准偏差。我们认为,最适合直接临床部署的模型是LASSO,而非判别能力最高的模型。我们为机器学习与医疗健康界提供了关于数据基础设施、模型选择以及校准和可解释性在高风险决策支持中价值的经验教训。
cs.LG / 42 / 2608.29420

One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation

单一能力还是多重能力?检验前沿AI评测的经济学有效性
Zhu, Louis Yiven
Abstract
Frontier-model leaderboards now rank systems based on economic benchmarks, tests of how well models carry out professional tasks from software engineering to banking workflows, and those rankings inform what organisations buy, what regulators scrutinise, and expectations of how work will change. Whether such benchmarks measure a capability distinct from general test-taking, or re-express the one axis along which every benchmark rises as models improve, is a question of construct validity that has not yet been studied. We test it on a hash-pinned leaderboard snapshot of 421 model configurations across twelve benchmarks, four of them economic, treating benchmarks as items and models as respondents in a latent-variable model with four hypotheses and their thresholds fixed before analysis. A single factor explains 74.5% of common variance and tracks model release date (R^2 = 0.505), so the leading axis of capability is substantially a time trend; where prior work controls for scale, compute adds little once date is removed. Removing the date trend lowers that share by 14.9 points, and by 24.1 with one row per base model. Under the dimensionality rule fixed in advance the economic benchmarks form no distinct factor, yet a leave-one-benchmark-out test with factors re-estimated inside every fold shows that a multi-factor representation predicts held-out economic scores better than a single general index (pooled Delta-MSE 0.037, 95% bootstrap interval [0.019, 0.055]). Economic benchmarks therefore add incremental predictive information to a largely date-driven general factor, and the evidence does not support treating them as a distinct latent capability. Leaderboards remain a sound guide to overall progress, but most of the gap between models released months apart is calendar, so a small gap between contemporaneous models should be date-adjusted before being read as a capability difference.
Chinese Translation
前沿模型排行榜如今基于经济学基准进行排名,即测试模型执行从软件工程到银行工作流程等专业任务的能力,这些排名影响着企业采购决策、监管机构的审查重点,以及对工作方式变革的预期。此类基准究竟衡量了一种区别于通用应试能力的独立能力,还是仅仅重新表达了所有基准随模型进步而共同提升的单一轴线,这是一个尚未被研究的构念效度问题。我们在一个经哈希固定的排行榜快照上对这一问题进行了检验,该快照涵盖12个基准(其中4个为经济学基准)的421个模型配置,将基准视为条目、模型视为受访者,构建潜变量模型,并在分析前预设了四个假设及其判定阈值。单一因子可解释74.5%的共同方差,并与模型发布日期相关(R^2 = 0.505),因此能力的主导轴线在很大程度上是一个时间趋势;在以往研究中对规模加以控制后,一旦剔除日期因素,计算量所贡献的增量很小。剔除日期趋势后,该比例下降14.9个百分点;若每个基础模型仅保留一行数据,则下降24.1个百分点。按照预先设定的维度判定规则,经济学基准并未形成独立的因子;然而,一项在每个折内重新估计因子的留一基准交叉验证测试表明,多因子表示比单一通用指数能更好地预测留出的经济学基准得分(合并Delta-MSE为0.037,95%自助法区间为[0.019, 0.055])。因此,经济学基准在一个主要由日期驱动的通用因子之外提供了增量预测信息,但证据不支持将其视为一种独立的潜在能力。排行榜仍是衡量整体进展的可靠指南,但相隔数月发布的模型之间的差距大部分源于日历时间,因此同期模型之间的微小差距应先进行日期调整,再解读为能力差异。
cs.LG / 43 / 2608.29428

Behavioral Latency as Weak Event-Time Supervision for EEG Reaction-Time Decoding

行为潜伏期作为弱事件时间监督的脑电反应时解码方法
Aimoldin, Anuar, Mussabayeva, Ayana, Mussabayev, Yedige, Liu, Xue, Zhang, Kun
Abstract
Single-trial EEG analyses are often organized around events and latencies, yet EEG-based reaction-time (RT) prediction is posed as scalar regression on a fixed stimulus-locked window. RT is treated as a window-level label rather than timing evidence about response-relevant dynamics. Here we reformulate trial-wise RT decoding as event-time posterior modeling. Instead of predicting RT directly, the model estimates a posterior over response-relevant event times, $p(t_{\mathrm{event}}\mid X)$, and uses its mean as the RT estimate. This treats behavioral latency as a weak observation of latent response-relevant timing. We evaluate this formulation on the Healthy Brain Network contrast change detection EEG task under a subject-disjoint, release-separated protocol. Across five seeds, distributional event-time supervision consistently improves held-out RT prediction relative to scalar regression and temporal-readout controls. Controlled objective comparisons isolate supervision of the event-time distribution, rather than expectation-based readout alone, as the source of this gain. Architecture controls show that the effect persists across four temporal backbones and is not explained by model scale. Beyond point prediction, posterior geometry characterizes concentration, target alignment, and interval behavior, while observation-noise calibration separates latent concentration from predictive uncertainty over RT. Shifted-crop inference probes shortcut use versus temporal localization. Matched shift-jitter improves robustness, increases mean sensitivity, and moves predictions more often in the expected crop-relative direction. Sensitivity remains below ideal crop-relative localization, leaving a clear equivariance gap. Together, these results establish event-time posterior modeling as a probabilistic and interpretable formulation for linking single-trial EEG dynamics to behavioral timing.
Chinese Translation
单试次脑电(EEG)分析通常围绕事件与潜伏期展开,然而基于EEG的反应时(RT)预测往往被构建为在固定刺激锁定窗口上的标量回归问题。在这种范式下,RT被视为窗口级标签,而非关于响应相关动态的时间证据。本文将试次级RT解码重新构建为事件时间后验建模问题:模型不再直接预测RT,而是估计响应相关事件时间的后验分布 $p(t_{\mathrm{event}}\mid X)$,并以该分布的均值作为RT估计值。这将行为潜伏期视为对潜在响应相关时间信息的弱观测。我们在Healthy Brain Network对比变化检测EEG任务上,采用被试不重叠、数据发布版本分离的协议对该方法进行评估。在五个随机种子下,相对于标量回归和时间读出对照方法,基于分布的事件时间监督在保留集RT预测上带来了一致的提升。受控目标函数比较表明,这一增益来源于对事件时间分布的监督本身,而非仅基于期望的读出方式。架构对照实验表明,该效应在四种时间骨干网络上均持续存在,且无法用模型规模解释。除点预测外,后验几何结构可刻画分布集中度、目标对齐性与区间行为,而观测噪声校准可将潜在集中度与RT预测的不确定性区分开来。移位裁剪推断则用于探测模型是否依赖捷径而非真正的时间定位。匹配的移位抖动提升了鲁棒性、增加了平均敏感度,并使预测更频繁地朝预期的裁剪相对方向变化。但敏感度仍低于理想的裁剪相对定位水平,存在明显的等变性差距。综上,这些结果确立了事件时间后验建模作为一种概率化、可解释的框架,能够将单试次EEG动态与行为时间信息相联系。
cs.LG / 44 / 2608.29434

Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations

潜在规划能否在点云上存续?面向几何观测的动作条件JEPA世界模型
Oberweger, Fabio F., Schwingshackl, Michael
Abstract
JEPA world models make latent-space planning a practical route to control, but they are built almost exclusively on images. Whether latent prediction survives geometric observations is unclear: point clouds are sparse, unordered, and self-occluded, and with 0.3-15% of scene points moving, the slow-feature optimum of latent prediction compounds with the geometric shortcut of 3D self-supervision. We lift three canonical JEPA designs to point clouds, frozen-encoder, distribution-prior, and action-sensitive, and re-sense the stable-worldmodel benchmark so that only the observation differs from the image baselines. All three plan without collapse: the distribution-prior model is statistically equivalent to its re-evaluated image counterpart on every benchmark, and the action-sensitive model attains the strongest result in our controlled comparison where the most geometry moves. Probing explains why: object positions are almost perfectly linearly decodable and attention falls on the few moving points. Planning withstands heavy dropout never seen in training, though range noise defeats the thinnest scene. Geometry finally makes a commanded 3D target a natural goal interface: we construct the goal latent from the target and the current latent, at no cost in success rate, without a goal observation.
Chinese Translation
JEPA世界模型使潜在空间规划成为一条切实可行的控制路径,但它们几乎完全构建于图像之上。潜在预测能否在几何观测下存续尚不清楚:点云稀疏、无序且存在自遮挡,并且当场景中仅0.3%–15%的点发生移动时,潜在预测的慢特征最优性会与3D自监督的几何捷径相互叠加而加剧问题。我们将三种经典的JEPA设计提升到点云上——冻结编码器、分布先验和动作敏感型——并对稳定世界模型基准进行重新感知,使其与图像基线的差异仅在于观测形式。三种模型均能在无坍缩的情况下进行规划:分布先验模型在每个基准上与其重新评估的图像对应模型在统计上等价;而在受控比较中几何运动最大的场景下,动作敏感模型取得了最强结果。探测分析解释了原因:物体位置几乎可以完美地线性解码,且注意力集中在少数移动点上。规划能够承受训练中从未见过的严重点丢失,尽管距离噪声击败了最稀疏的场景。几何观测最终使命令式的3D目标成为天然的目标接口:我们从目标点与当前潜在表示构建目标潜在向量,在不损失成功率的情况下无需目标观测。
cs.LG / 45 / 2608.29448

SS-ESOAP: Self-Scaled Adaptive Preconditioning for Physics-Informed Learning

SS-ESOAP:面向物理信息学习的自缩放自适应预处理方法
Wang, Guangyuan, Toftrup, Mads, Loeschcke, Sebastian, Wang, Yixuan, Anandkumar, Anima
Abstract
Physics-informed neural networks (PINNs) often face ill-conditioned objectives that limit high-accuracy training. Dense quasi-Newton methods improve local conditioning but require expensive optimizer state, while Kronecker-factored methods such as SOAP scale to larger networks but rely on periodic basis updates. We introduce \method, which augments SOAP-style preconditioning with a scalar secant-energy correction adapted to Kronecker geometry and an adaptive basis update followed by variance-state downscaling. We characterize the directional secant matching induced by the scalar correction and give a bound on variance-state mismatch across basis changes. Across eight PDE benchmarks, \method attains the lowest final residual on six, including Burgers and Boussinesq, while SOAP-family baselines perform better on Gray-Scott and Ginzburg-Landau. On Boussinesq, \method reaches a residual of $10^{-5}$ in 4.1 hours with 9.2 GB peak VRAM, while Adam does not reach this target within 14 hours. Three-seed $L^2$ and $H^1$ errors on four representative PDEs support the link between lower residuals and improved solution accuracy. These results position \method as a scalable option for stiff, high-accuracy physics-informed training, rather than a uniform replacement for existing optimizers.
Chinese Translation
物理信息神经网络(PINNs)常常面临病态的目标函数,限制了高精度训练。稠密拟牛顿方法可以改善局部条件数,但需要昂贵的优化器状态开销;而诸如 SOAP 等基于 Kronecker 分解的方法可以扩展到更大的网络,但依赖于周期性的基更新。我们提出了 SS-ESOAP 方法,它在 SOAP 风格的预处理基础上,引入了适配 Kronecker 几何结构的标量割线能量校正,并采用自适应基更新及随后的方差状态缩减。我们刻画了标量校正所诱导的方向性割线匹配特性,并给出了基变换下方差状态失配的一个上界。在八个偏微分方程基准测试中,SS-ESOAP 在其中六个(包括 Burgers 和 Boussinesq 方程)上取得了最低的最终残差,而 SOAP 系列基线方法在 Gray-Scott 和 Ginzburg-Landau 方程上表现更优。在 Boussinesq 方程上,SS-ESOAP 在 4.1 小时内、以 9.2 GB 的峰值显存占用将残差降至 $10^{-5}$,而 Adam 在 14 小时内未能达到该目标。在四个代表性偏微分方程上进行的三种子 $L^2$ 和 $H^1$ 误差实验支持了更低残差与更高解精度之间的关联。这些结果表明,SS-ESOAP 是刚性、高精度物理信息训练的一种可扩展选择,而非对现有优化器的统一替代方案。
cs.LG / 46 / 2608.29458

Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities

参考嫁接在引发被隐藏(Sandbagged)能力方面可媲美微调
Le, Linh, Tan, Hong Kiat, Williams-King, David
Abstract
Sandbagging, in which a model deliberately underperforms on an evaluation despite retaining the underlying capability, threatens the safety evaluations that frontier-model governance depends on. The Elicitation Game found that fine-tuning elicits hidden capability from sandbagging model organisms whereas additive activation steering fails. We revisit that verdict with reference-grafting, which sets an activation's coordinate along a contrast direction to the value it takes in an honest reference, at a small set of circuits chosen by active learning. Across eleven fine-tuned password-locked models (1.5-32B, three architecture lineages), it recovers +94 to +101% of the honest-sandbagging gap, matching fine-tuning elicitation without weight updates or training labels; two to five paired examples suffice to fit the direction. Similar recovery holds for reinforcement-learning-induced sandbagging and for password-locked code generation. Grafting works because the fine-tuned lock is a thresholded gate: held-out accuracy stays at the sandbagged level until the grafted coordinate crosses a threshold near the honest reference, which is why additive steering and zeroing the coordinate often fail. The direction tracks how the sandbagging was induced rather than what is withheld -- fit on grade-school science (ARC) it elicits withheld hazardous knowledge (WMDP), yet prompting, fine-tuning, and reinforcement learning each carry a different direction. Circuit-breaking marks the boundary: it reroutes activations on every forward pass, so the fixed edits we test are re-broken downstream and do not restore coherent generation.
Chinese Translation
沙袋行为(sandbagging),即模型在评估中故意表现不佳但实际仍保留底层能力,威胁着前沿模型治理所依赖的安全评估。此前的研究(Elicitation Game)发现,微调能够从沙袋模型中引出隐藏能力,而加性激活导向(additive activation steering)则无法做到。我们通过参考嫁接(reference-grafting)方法重新审视了这一结论。该方法将激活在某个对比方向上的坐标设置为诚实参考模型中的取值,作用在由主动学习选出的一小组电路上。在十一个经微调加密码锁定的模型上(1.5B-32B,三种架构谱系),该方法恢复了94%到101%的“诚实-沙袋”能力差距,效果与微调引出相当,且无需权重更新或训练标签;仅需二至五个配对样本即可拟合该方向。类似的恢复效果在强化学习诱发的沙袋行为和密码锁定的代码生成任务中同样成立。嫁接之所以有效,是因为微调产生的锁是一个阈值门控:在嫁接坐标越过接近诚实参考值的某个阈值之前,留出集上的准确率一直停留在沙袋水平——这正是加性导向和将坐标置零常常失效的原因。该方向追踪的是沙袋行为的诱发方式而非被隐藏的内容——在小学科学数据(ARC)上拟合的方向能够引出被隐藏的危险知识(WMDP),而提示工程、微调和强化学习各自对应不同的方向。电路破坏(circuit-breaking)则标记了该方法的边界:它在每次前向传播中都会重路由激活,因此我们所测试的固定编辑会在下游被重新破坏,无法恢复连贯的生成。
cs.LG / 47 / 2608.29461

A Causal Model for Locating and Unlocking Sandbagging in Model Organisms

一个用于定位与解锁模型生物中“沙袋行为”(Sandbagging)的因果模型
Tan, Hong Kiat, Le, Linh, Williams-King, David
Abstract
Sandbagging models strategically underperform on evaluations while retaining the capabilities being measured. The evaluations that guide frontier-model deployment and governance then understate what these models can do. To understand the mechanism, we propose a causal model of how sandbagging is carried in the residual stream. Early layers write the sandbagging intent onto a single axis of the stream, and a later layer reads that axis and commits the answer. We study three instruction-tuned models (Qwen2.5-7B, Llama-3-8B, and Mistral-7B) and four ways of installing a sandbagging lock (prompting, fine-tuning, reinforcement learning, and circuit breaking). Each lock decides from the prompt whether to sandbag, and the fine-tuned and circuit-broken locks answer honestly whenever a password appears. The causal model predicts a window of layers, after the last sandbagging write and before the answer commit, in which a single-layer reference graft of the sandbagging axis to its honest value restores the full capability. The single-layer graft recovers the capability in 28 of the 33 runs of the prompted, fine-tuned, and RL-trained locks, with a median held-out recovery of 96%. The circuit-broken lock rewrites the whole state through a band of layers, and the single-layer graft fails at every layer. We therefore introduce a second intervention, context grafting, which replays the password's cached key/value activations so that every layer's attention reads them as additional context. Context grafting provably and empirically restores the full capability on all three models, and the recovery is surprisingly insensitive to the exact password content. More broadly, an auditor can use this causal model to design interventional auditing techniques for sandbagging models.
Chinese Translation
沙袋行为(Sandbagging)指模型在评估中策略性地表现不佳,同时仍保留被测量的能力。这导致用于指导前沿模型部署与治理的评估低估了这些模型的实际能力。为了理解其机制,我们提出了一个因果模型,描述沙袋行为如何在残差流(residual stream)中被承载。早期层将沙袋意图写入残差流的单一轴线,随后的一个层读取该轴线并确定答案。我们研究了三个指令微调模型(Qwen2.5-7B、Llama-3-8B 和 Mistral-7B),以及四种安装沙袋锁的方式(提示、微调、强化学习和断路(circuit breaking))。每种锁都会根据提示决定是否进行沙袋行为,而经微调和断路方式安装的锁在出现密码时会诚实地作答。该因果模型预测存在一个层区间——位于最后一次沙袋写入之后、答案确定之前——在该区间内,将沙袋轴线以单层参考嫁接(reference graft)方式替换为其诚实取值,即可恢复全部能力。对于经过提示、微调和强化学习训练的锁,单层嫁接在33次运行中的28次成功恢复了能力,留出集(held-out)恢复率的中位数为96%。而断路锁通过一整段层带重写了整个状态,单层嫁接在每一层均告失败。因此,我们引入了第二种干预手段——上下文嫁接(context grafting),它重放密码的缓存键/值激活(key/value activations),使每一层的注意力将其作为附加上下文读取。上下文嫁接在理论上可证明且在实验上确实能在全部三个模型上恢复完整能力,且这种恢复对密码的具体内容出奇地不敏感。更广泛而言,审计者可以利用该因果模型为存在沙袋行为的模型设计干预式审计技术。
cs.LG / 48 / 2608.29472

Knowledge Distillation under Teacher Misspecification: An Order-Parameter Analysis of the Gap between Teacher Mimicry and Task Performance

教师失配下的知识蒸馏:教师模仿与任务表现之间差距的序参量分析
Hara, Kazuyuki, Hino, Hideitsu
Abstract
Knowledge distillation trains a small student model to reproduce the outputs of a large teacher model, and its progress is typically monitored through the teacher--student discrepancy. The quantity of ultimate interest, however, is the student's error with respect to the true task. We study the relation between these two objectives in a minimal three-party model, a true teacher (generative model), a teacher, and a student, all soft committee machines, in which the true teacher contains a shared latent factor that the teacher cannot represent, with mismatch strength controlled by a single scalar $\dmiss$. Within an order-parameter description of online distillation, and exploiting closed-form (arcsine-type) expressions for all errors under error-function activations, we prove that the learning dynamics and the distillation error $\Ets$ are exactly invariant to $\dmiss$, whereas the true error $\Etzs$ and the gap $\Delta=\Etzs-\Ets$ are strictly increasing in $\dmiss$, with a rate that is amplified linearly by the complexity $M_0$ of the true teacher. Numerical phase diagrams over the plane spanned by true-teacher complexity and student capacity confirm the predicted deformation: the contours of $\Ets$ do not move while the landscape of $\Etzs$ rises systematically, and a teacher-miss regime, where mimicry succeeds but the task fails, expands with $\dmiss$. The results give a quantitative warning against evaluating distillation solely through teacher-mimicry metrics and identify the gap $\Delta$ as a minimal diagnostic for distinguishing teacher-miss from capacity-limited failure.
Chinese Translation
知识蒸馏通过训练一个小型学生模型来复现大型教师模型的输出,其进展通常通过师生输出差异来监测。然而,最终真正关注的是学生在真实任务上的误差。我们在一个最小化三方模型中研究这两种目标之间的关系:一个真实教师(生成模型)、一个教师和一个学生,均为软委员会机(soft committee machine),其中真实教师包含一个教师无法表达的共享潜在因子,失配强度由单个标量 $\dmiss$ 控制。基于在线蒸馏的序参量描述,并利用误差函数激活下所有误差的闭式(arcsine 型)表达式,我们证明学习动力学与蒸馏误差 $\Ets$ 对 $\dmiss$ 严格不变,而真实误差 $\Etzs$ 及差距 $\Delta=\Etzs-\Ets$ 随 $\dmiss$ 严格递增,其速率被真实教师的复杂度 $M_0$ 线性放大。在真实教师复杂度与学生容量所张成的平面上的数值相图证实了所预测的形变:$\Ets$ 的等值线不动,而 $\Etzs$ 的地形系统性地抬升,且一个"教师失配"区域——即模仿成功但任务失败——随 $\dmiss$ 扩大。这些结果定量地警示了仅通过教师模仿指标评估蒸馏的局限,并将差距 $\Delta$ 确立为区分教师失配失败与容量受限失败的最小诊断量。
cs.LG / 49 / 2608.29494

Learning Human Health and Diseases from 24-hour Wrist Movement

从24小时腕部运动中学习人类健康与疾病信息
Wang, Yong, McGagh, Dylan, Broomberg, Katya, Zhang, Zizheng, Carter, Jonathan, Naushad, Junayed, Brocklebank, Laura, Sun, Yang, Nicholson, George, Sun, Dianjianyi, Yu, Canqing, Lv, Jun, Barnard, Maxim, Lam, Hubert, Steptoe, Andrew, Eyre, David W., Li, Liming, Chen, Zhengming, Wray, Naomi, Denaxas, Spiros, Collins, Gary S., Du, Huaidong, Doherty, Aiden, Yuan, Hang
Abstract
Much of human health and function unfolds beyond the clinic, through the movements of everyday life. Wrist-worn accelerometers capture these movements continuously, yet their rich signals are often reduced to a small set of predefined behavioural summary measures. Here, we present Sensori, a self-supervised foundation model that learns general-purpose health representations directly from 24 hours of raw tri-axial wrist movement. We developed and evaluated the model across four population-based cohorts from the United Kingdom, China and the United States, comprising 122,640 participants contributing 683,617 person-days of free-living recordings. Sensori condensed each day of movement into a representation that captured diverse movement behaviours, demographic characteristics, health axes and physical function. Evaluation in independent cohorts showed that these representations generalised across populations and measurement settings without retraining. When added to common clinical covariates, Sensori significantly improved prevalent disease classification for 52 of 102 eligible conditions (median delta AUROC, 0.060; range, 0.012-0.242) and incident disease risk prediction for 26 of 87 eligible conditions (median delta Uno's C-index, 0.064; range, 0.025-0.172), with the largest gains for neurological and psychiatric disorders. These findings establish 24-hour wrist movement as a rich and scalable source of health information, with the potential to support passive health monitoring and disease prediction at population scale.
Chinese Translation
人类的健康与功能很大程度上在临床之外展开,体现在日常生活的运动之中。腕戴式加速度计能够连续捕捉这些运动,但其丰富的信号通常被简化为少量预定义的行为汇总指标。本文提出了Sensori,一个自监督基础模型,可直接从24小时的原始三轴腕部运动数据中学习通用的健康表征。我们在来自英国、中国和美国的四个基于人群的队列中开发并评估了该模型,共纳入122,640名参与者,累计683,617人天的自由生活记录。Sensori将每天的运动数据压缩为一种表征,该表征涵盖了多样的运动行为、人口统计学特征、健康维度以及身体功能。在独立队列中的评估表明,这些表征无需重新训练即可在不同人群和测量环境间泛化。在常见临床协变量的基础上加入Sensori表征后,其显著提升了102种合格疾病中52种疾病的现患分类性能(AUROC增量中位数为0.060,范围0.012-0.242),以及87种合格疾病中26种疾病的发病风险预测性能(Uno's C指数增量中位数为0.064,范围0.025-0.172),其中对神经系统和精神类疾病的提升最为显著。这些发现确立了24小时腕部运动作为丰富且可扩展的健康信息来源的地位,有望支持人群规模的被动健康监测与疾病预测。
cs.LG / 50 / 2608.29496

Target-Aware State-Adaptive $p$-Dirichlet Graph Neural Regression for Non-Invasive Body-Composition Estimation

面向目标的状态自适应 $p$-Dirichlet 图神经回归用于无创身体成分估计
Drenska, Nadejda, Lemoine, Matthew, Sunkara, Gowri Priya, Wang, Yu, Devarakonda, Sri Lakshmi Sravani, Heymsfield, Steven B.
Abstract
Accurate estimation of body-composition outcomes, including body fat percentage (BFP), bone mineral density (BMD), and appendicular lean mass (ALM), is important for evaluating metabolic, skeletal, and muscular health. Direct assessment using dual-energy X-ray absorptiometry (DXA), however, requires specialized equipment and involves ionizing radiation. We propose a target-aware, state-adaptive $p$-Dirichlet energy-flow graph neural regression ($p$SADE-GNR) framework for estimating these outcomes from non-invasive anthropometric measurements. A neural encoder maps participant representations to hidden states that are propagated over an outcome-specific participant-similarity graph by a state-adaptive forward-Euler discretization of the graph $p$-Dirichlet energy flow. Graph distances weight each original or latent coordinate by its normalized absolute training-fold correlation with the outcome. Using clinical data from the Pennington Biomedical Research Center and five-fold cross-validation, the correlation-weighted model using the original standardized measurements achieved the lowest root mean squared error in all nine primary outcome-cohort combinations and outperformed previously reported support vector regression or least-squares support vector regression reference values in eight of nine comparisons. Autoencoder, variational-autoencoder, and Gaussian-mixture variational-autoencoder representations generally did not improve primary-outcome prediction or reduce computational cost. In an exploratory age-prediction analysis including ALM, BMD, and BFP as predictors, the correlation-weighted GMVAE model achieved the lowest mean error in all three cohorts. These results support target-aware, state-adaptive $p$-Dirichlet graph neural regression for non-invasive body-composition estimation.
Chinese Translation
准确估计身体成分指标,包括体脂率(BFP)、骨矿物质密度(BMD)和四肢瘦体重(ALM),对于评估代谢、骨骼和肌肉健康十分重要。然而,直接使用双能X线吸收测定法(DXA)进行评估需要专门的设备,且涉及电离辐射。我们提出了一种面向目标的状态自适应 $p$-Dirichlet 能量流图神经回归($p$SADE-GNR)框架,用于从无创的人体测量数据中估计这些指标。神经编码器将参与者表示映射到隐藏状态,随后通过图 $p$-Dirichlet 能量流的状态自适应前向欧拉离散化,在特定于结果的参与者相似性图上进行传播。图距离根据每个原始或潜在坐标与结果的归一化绝对训练折相关性对其进行加权。基于 Pennington 生物医学研究中心的临床数据和五折交叉验证,使用原始标准化测量的相关性加权模型在全部九个主要结果-队列组合中均取得了最低的均方根误差,并在九个比较中的八个优于先前报告的支持向量回归(SVR)或最小二乘支持向量回归(LS-SVR)基准值。自编码器、变分自编码器和高斯混合变分自编码器(GMVAE)的表示通常未能改善主要结果的预测,也未降低计算成本。在一项将 ALM、BMD 和 BFP 作为预测变量的探索性年龄预测分析中,相关性加权的 GMVAE 模型在所有三个队列中均取得了最低的平均误差。这些结果支持将面向目标的状态自适应 $p$-Dirichlet 图神经回归用于无创身体成分估计。
cs.LG / 51 / 2608.29503

Adversarial Online Classification with a Preview

带预览的对抗性在线分类
Livni, Roi, Singla, Sahil
Abstract
Worst-case online classification is governed by sequential complexity, such as Littlestone dimension, and can be impossible even for statistically simple classes, such as thresholds of VC dimension one. We study a preview model in which an oblivious adversary fixes an entire labeled sequence of length $T$, a uniformly random subset of size $pT$ is revealed before prediction begins, and the remaining $(1-p)T$ examples are then presented in their original adversarial order. Against the best full-sequence hypothesis evaluated on the unrevealed examples, we characterize the dependence on the preview rate $p$: for binary classes of VC dimension $d$, the optimal excess loss is $\Theta(d/p+\sqrt{dT})$, up to the trivial cap at $T$; for multiclass classes we obtain the corresponding $\widetilde O(d_{\rm DS}/p+\sqrt{d_{\rm Nat}T})$ bound with no dependence on the number of labels. Thus a random preview can replace worst-case sequential complexity by classical statistical dimensions without randomizing the online order. To achieve the sharp binary bound, our ChainedPrediction algorithm uses an online analogue of chaining, implemented as a multiscale aggregation algorithm rather than only as an analytic argument.
Chinese Translation
最坏情形下的在线分类由序列复杂度决定,例如Littlestone维数,即使对于统计上简单的类别(如VC维为一的阈值函数),在线分类也可能是不可能的。我们研究一个预览模型:一个非自适应的对抗者事先确定整个长度为$T$的带标签序列,在预测开始之前均匀随机地揭示大小为$pT$的子集,其余$(1-p)T$个样本随后按其原始的对抗性顺序逐一呈现。以在未揭示样本上评估的最优全序列假设为基准,我们刻画了对预览率$p$的依赖关系:对于VC维为$d$的二分类问题,最优超额损失为$\Theta(d/p+\sqrt{dT})$(上限为平凡上界$T$);对于多分类问题,我们得到相应的$\widetilde O(d_{ m DS}/p+\sqrt{d_{ m Nat}T})$界,且不依赖于标签数量。因此,随机预览可以用经典的统计维数取代最坏情形下的序列复杂度,而无需对在线顺序进行随机化。为达到二分类情形的精确界,我们的ChainedPrediction算法采用了链式(chaining)方法的在线类比,将其实现为一个多尺度聚合算法,而不仅仅是作为一种解析论证。
cs.LG / 52 / 2608.29507

Denoising as Projection: Constrained Optimization with Gradient-Guided Diffusion

去噪即投影:基于梯度引导扩散的约束优化
Zhang, Runyu, Zhang, Jiawei, Zardini, Gioele, Amin, Saurabh, Ozdaglar, Asuman
Abstract
Diffusion models are increasingly used not only for sampling from learned data distributions, but also for generating samples that optimize task-specific objectives. A common approach is to guide the reverse diffusion process using gradients of an external objective. However, when the data distribution is supported on a structured feasible set, such as a manifold or a constraint set, gradient guidance can move samples away from the learned data geometry. In this paper, we study a simple projected-gradient-guided diffusion update based on the observation that the Stein denoising operator can act as an approximate projection onto the data geometry. The proposed update incorporates the objective gradient inside the denoising step, yielding an inference-time method that uses only a pretrained denoiser and gradient evaluations. We analyze this update as an inexact projected-gradient method for constrained optimization over learned feasible geometries. Our theory covers three settings: linear manifolds, compact convex feasible sets, and compact Riemannian submanifolds. In all these settings, we prove descent and finite-time convergence guarantees. Numerical experiments support the theoretical interpretation and illustrate how the proposed update balances objective descent with preservation of the learned geometry.
Chinese Translation
扩散模型不仅被越来越多地用于从学习到的数据分布中采样,还被用于生成能够优化特定任务目标的样本。一种常见的方法是利用外部目标的梯度来引导反向扩散过程。然而,当数据分布支撑在结构化可行集(如流形或约束集)上时,梯度引导可能会使样本偏离所学习到的数据几何结构。本文基于Stein去噪算子可以作为向数据几何结构的近似投影这一观察,研究了一种简单的投影梯度引导扩散更新方法。所提出的更新将目标梯度融入去噪步骤之中,得到一种仅需预训练去噪器和梯度评估的推理时方法。我们将该更新分析为一种针对学习到的可行几何结构上约束优化的非精确投影梯度法。我们的理论涵盖三种情形:线性流形、紧凸可行集以及紧黎曼子流形。在所有这些情形中,我们都证明了下降性和有限时间收敛保证。数值实验支持了这一理论解释,并展示了所提出的更新方法如何在目标下降与保持学习到的几何结构之间取得平衡。
cs.LG / 53 / 2608.29513

On the Plasticity Collapse in Continual Machine Unlearning

论持续机器遗忘中的可塑性坍塌
Shi, Yingdan, Xu, Xiang, Ding, Kaize, Hero, Alfred O., Wang, Ren
Abstract
Machine unlearning enables deep neural networks to selectively remove the influence of specific data in response to privacy and regulatory requirements. While prior work largely studies single-shot unlearning, real-world systems must accommodate continual unlearning, where multiple unlearning requests occur sequentially over time. In this work, we identify a fundamental limitation of this setting: plasticity collapse, a progressive breakdown in a model's ability to effectively forget. Through theoretical analysis of continual unlearning dynamics, we show that continual unlearning operations accumulate geometric constraints in parameter space, leading to saturated subspaces that restrict future updates. This structural effect induces two distinct failure modes: (1) Forward failure -- diminishing forgetting quality for subsequent tasks, and (2) Backward failure -- spontaneous re-memorization of previously forgotten information. Extensive experiments across multiple architectures, datasets, and methods in image classification confirm that plasticity collapse is not an artifact of specific implementations, but a pervasive phenomenon inherent to continual unlearning. Our findings reveal a critical barrier to the long-term reliability of machine unlearning systems and motivate the development of plasticity-preserving unlearning algorithms. Our code is available at https://github.com/TIML-Group/Continual-Machine-Unlearning-Plasticity-Collapse
Chinese Translation
机器遗忘(Machine Unlearning)使深度神经网络能够选择性地消除特定数据的影响,以满足隐私和监管要求。尽管已有研究主要关注单次遗忘,但现实世界中的系统必须支持持续遗忘,即多个遗忘请求随时间顺序发生。在本工作中,我们识别了该设置下的一个根本性局限:可塑性坍塌(plasticity collapse),即模型有效遗忘能力逐步退化。通过对持续遗忘动力学的理论分析,我们证明持续遗忘操作会在参数空间中累积几何约束,形成限制未来更新的饱和子空间。这种结构性效应引发两种不同的失效模式:(1)前向失效——后续任务的遗忘质量逐渐下降;(2)后向失效——先前已遗忘的信息被自发地重新记忆。在图像分类任务中,跨多种架构、数据集和方法的广泛实验证实,可塑性坍塌并非特定实现方式的产物,而是持续遗忘中普遍存在的固有现象。我们的发现揭示了机器遗忘系统长期可靠性的关键障碍,并推动了保持可塑性的遗忘算法的发展。我们的代码已在 https://github.com/TIML-Group/Continual-Machine-Unlearning-Plasticity-Collapse 公开。
cs.LG / 54 / 2608.29528

MedCache: Efficient and Temporally Valid Memory for Longitudinal Clinical Agents

MedCache:面向纵向临床智能体的高效且时间有效的记忆机制
Ting, Hei, Chan, Wu, Chenwei, Liu, Xueshen, Zheng, Boyuan, Shen, Liyue, Chen, Jiasi, Mao, Z. Morley
Abstract
Longitudinal clinical agents must maintain an evolving patient state from evidence distributed across visits, time points, and specialties. However, how agent memory should be designed for this setting remains unclear. We introduce a benchmark of multi-visit, multi-specialty patient records that evaluates long-context evidence retrieval, cross-time evidence aggregation, and cross-specialty clinical reasoning. Using this benchmark, we systematically study four memory design choices: curation, organization, retrieval, and memory-augmented reasoning. We find that temporal validity is more important than simply retaining more history; specialty-factorized memory reduces context but can hide shared evidence; and multiple agents help when specialists must reason together, not merely when evidence comes from multiple memories. Guided by these findings, we propose \textit{MedCache}, a hybrid framework that constructs temporally valid patient memory, organizes evidence into overlapping specialty views, routes each query to relevant memories, and adaptively invokes one or multiple specialists. Experiments show that MedCache improves reasoning accuracy and memory efficiency over strong single-agent and multi-agent baselines, while generalizing across model backbones and external datasets.
Chinese Translation
纵向临床智能体需要从分布于多次就诊、多个时间点和不同专科的证据中维护一个不断演进的患者状态。然而,针对这一场景,智能体记忆应如何设计仍不明确。我们构建了一个包含多次就诊、多专科患者记录的基准,用于评估长上下文证据检索、跨时间证据聚合以及跨专科临床推理。基于该基准,我们系统地研究了四种记忆设计选择:内容筛选(curation)、组织方式(organization)、检索机制(retrieval)以及记忆增强推理。我们发现,时间有效性比简单地保留更多历史信息更为重要;按专科分解的记忆虽然能减少上下文规模,但可能隐藏共享的证据;而多个智能体仅在专科医生需要共同推理时才有帮助,而不仅仅是在证据来自多个记忆源时。基于这些发现,我们提出了 MedCache,一个混合框架,它构建时间有效的患者记忆,将证据组织成相互重叠的专科视图,将每个查询路由到相关记忆,并自适应地调用一个或多个专科智能体。实验表明,MedCache 在推理准确性和记忆效率方面优于强大的单智能体和多智能体基线方法,同时在不同模型骨干和外部数据集上具有良好的泛化能力。
cs.LG / 55 / 2608.29553

BEACON: Behavioral and Semantic Enrichment of AlphaEarth Embeddings through Tri-Modal Contrastive Learning

BEACON:通过三模态对比学习对 AlphaEarth 嵌入进行行为与语义增强
Tian, Hao, Cai, Heng, Yang, Yifan
Abstract
Geospatial foundation models such as the AlphaEarth Foundation produce compact and globally consistent representations of the Earth's surface that transfer effectively to a wide range of downstream tasks. However, because these models are trained primarily on Earth-observation imagery, their embeddings mainly capture physical and spectral characteristics while encoding human activity and urban function only weakly. To address this limitation, we propose BEACON, a tri-modal contrastive learning framework that aligns three complementary views of urban space: physical representations from AE embeddings, semantic representations from point-of-interest (POI) text, and human behavioral representations from hourly POI visitation, while keeping the deployed representation image-only. Using the Houston Metropolitan Area as a case study area, we evaluated the performance of the BEACON framework on nine downstream tasks, including seven regression and two classification tasks against six baselines (raw coordinates, Space2Vec, SatCLIP, TESSERA, Clay and AlphaEarth), using frozen linear and MLP probes over five seeds. Under a linear probe, BEACON improves relative R^2 over AlphaEarth by up to 43% for obesity prevalence, 34% for poor mental health, and 22% for median household income, while remaining competitive in the prediction of physical and environmental variables. These findings highlight the value of augmenting geospatial foundation models with semantic and behavioral signals, extending their applicability from physical Earth observation to human-centered urban analytics.
Chinese Translation
诸如 AlphaEarth Foundation 之类的地理空间基础模型能够生成紧凑且全球一致的地表表示,并可有效迁移至广泛的下游任务。然而,由于这些模型主要基于对地观测影像进行训练,其嵌入主要捕捉物理与光谱特征,而对人类活动和城市功能的编码较为薄弱。为解决这一局限,我们提出了 BEACON,一个三模态对比学习框架,用于对齐城市空间的三个互补视图:来自 AE 嵌入的物理表示、来自兴趣点(POI)文本的语义表示,以及来自逐小时 POI 到访数据的人类行为表示,同时保持部署阶段的表示仅依赖影像。以休斯顿都市区作为案例研究区域,我们在九个下游任务(包括七个回归任务和两个分类任务)上,采用冻结的线性探测和 MLP 探测,在五个随机种子下,将 BEACON 框架与六个基线(原始坐标、Space2Vec、SatCLIP、TESSERA、Clay 和 AlphaEarth)进行了性能对比评估。在线性探测下,BEACON 相较于 AlphaEarth 在肥胖患病率上相对 R^2 提升最高达 43%,在心理健康不良上提升 34%,在家庭收入中位数上提升 22%,同时在物理和环境变量的预测上仍保持竞争力。这些发现凸显了利用语义与行为信号增强地理空间基础模型的价值,将其适用范围从物理地球观测拓展到以人为本的城市分析。
cs.LG / 56 / 2608.29560

Which LLM for Which Work? Budgeted Model Allocation under Uncertain Evaluation

哪种大语言模型适合哪种工作?不确定性评估下的预算约束模型分配
Khosravi, Hamed, Huo, Xiaoming
Abstract
A company with a fixed artificial intelligence (AI) budget must decide which large language model (LLM) handles each recurring workload. What it lacks is the quality table, how well each model performs on each workload. Given that table, the decision is a multiple-choice knapsack problem and is routine to solve, so estimating it is the difficulty, and that estimation fails in two ways. Models are rarely compared on the same work, and the recorded score is usually a proxy rather than the outcome the company values. Causal and off-policy methods repair the first but condition on the second, while evaluator-validation methods estimate the second but stop short of the decision. Worse, buying more re-evaluation cannot settle the second: randomization governs which requests are scored, not how a score is produced, so the table stays uncertain however much evaluation is purchased. Yet the deployment decision may still be determined even when the table is not. We therefore ask whether one assignment stays optimal across every quality table consistent with the evidence. For the fixed-budget problem, this admits an exact two-solve certificate: solve once at the estimated table and once at a least-favourable table. Agreement certifies the assignment; disagreement identifies the model-workload pairs where further evidence can matter. We propose CASE (causal active sequential experimentation), which targets evaluation to those pairs and repeats the test as evidence accumulates. On a production log, the measurement failure is the larger of the two: correcting assignment exactly still leaves most of the loss, and randomized re-evaluation does not remove it. In our experiments, the available evidence often does not determine the assignment. On paid software tasks, better information about model quality yields more savings than further optimization of the assignment on the same estimates.
Chinese Translation
一家拥有固定人工智能(AI)预算的公司必须决定由哪个大语言模型(LLM)来处理每项经常性工作负载。它所缺乏的是质量表,即每个模型在每项工作负载上的表现如何。若已知该质量表,这一决策就是一个多选择背包问题,求解是常规操作,因此难点在于估计该质量表,而这种估计会在两个方面失效。其一,模型很少在同一工作上被比较;其二,记录的分数通常只是一个代理指标,而非公司真正看重的结果。因果方法和离策略(off-policy)方法可以修复第一个问题,但依赖于第二个问题;评估器验证方法估计第二个问题,但止步于决策。更糟的是,购买更多的重新评估无法解决第二个问题:随机化决定的是哪些请求被评分,而不是分数如何产生,因此无论购买多少评估,质量表仍然是不确定的。然而,即便质量表不确定,部署决策仍可能是确定的。因此,我们要问:在所有与证据一致的质量表下,是否存在某个分配方案始终保持最优。对于固定预算问题,这可以通过一个精确的双求解证书来实现:在估计的质量表上求解一次,再在最不利(least-favourable)的质量表上求解一次。两次结果一致即证明该分配方案;不一致则识别出那些进一步证据可能产生影响的模型-工作负载对。我们提出了CASE(因果主动序贯实验,causal active sequential experimentation),它将评估资源定向到这些模型-工作负载对,并随着证据的积累重复测试。在生产日志上的实验表明,测量失效是二者中更大的问题:即使精确纠正分配,仍会留下大部分损失,且随机化的重新评估无法消除这一损失。在我们的实验中,可用证据常常不足以确定分配方案。在有偿软件任务上,获得更好的模型质量信息比在同一估计值上进一步优化分配方案能带来更多节约。
cs.LG / 57 / 2608.29562

Asynchronous Cooperative Online Learning for Multi-Robot Control under Computational Delays

计算延迟下多机器人控制的异步协作在线学习
Dai, Xiaobing, Yang, Zewen, Ren, Wei, Hirche, Sandra
Abstract
Ensuring the safe operation of multi-agent systems (MASs) under uncertain environments is crucial for cooperative robotic, where external disturbances and inaccurate dynamic models can significantly compromise performance and reliability. To address this challenge, calibrated machine learning models, particularly Gaussian process (GP) regression, are extensively employed due to their interpretable performance quantification. As the interconnected communication of MASs facilitates cooperative learning, agents are able to enhance learning performance by exchanging local GP inferences with their neighbors and aggregating the received information via distributed GP strategies. However, variations in computational power and prediction tasks among agents inevitably lead to heterogeneous computational delays and differences in query points, which are often overlooked in existing aggregation methods. To overcome these limitations, this work proposes an asynchronous cooperative learning strategy that explicitly accounts for prediction accuracy, query point variations and delay effects. Additionally, a distributed control law based on an adjoint MAS is developed to ensure the desired control performance. Simulations on unmanned surface vehicles validate the effectiveness of the proposed approach, demonstrating substantial improvements in both learning and control performance compared to the state-of-the-art approaches.
Chinese Translation
在不确定环境下确保多智能体系统(MASs)的安全运行对协作机器人技术至关重要,其中外部扰动和不精确的动力学模型会显著影响系统的性能与可靠性。为应对这一挑战,经过校准的机器学习模型,特别是高斯过程(GP)回归,因其可解释的性能量化能力而被广泛采用。由于多智能体系统的互联通信促进了协作学习,智能体能够通过与邻居交换局部GP推断并借助分布式GP策略聚合接收到的信息来提升学习性能。然而,智能体间计算能力和预测任务的差异不可避免地导致异构的计算延迟以及查询点的差异,而这一点在现有聚合方法中往往被忽视。为克服这些局限,本文提出了一种异步协作学习策略,显式地考虑了预测精度、查询点变化及延迟影响。此外,基于伴随多智能体系统设计了一种分布式控制律,以保证期望的控制性能。在无人水面艇上的仿真验证了所提方法的有效性,与最先进方法相比,在学习和控制性能方面均表现出显著提升。
cs.LG / 58 / 2608.29563

HoopMind: A Real-Time Neural Game-Tree System for Opponent-Aware Possession Planning

HoopMind:一个用于对手感知回合规划的实时神经网络博弈树系统
Gong, Yibo, Guo, Cong, Ding, Jiacheng
Abstract
School coaches prepare for opponents with game film and intuition. The analytics tools of professional teams stay out of reach. We ask how far public data can close this gap. Professional basketball is our case study, chosen for its data rather than the league. We fuse five public sources into one per-shot dataset of 4.23M shots over 21 seasons. The sources are shot locations, two play-by-play feeds, official matchup tracking, and player biometrics. Alignment across them is 99.5% to 100%. We also report two data pitfalls that are easy to miss. We then model a half-court possession as a sequential game. Shot values come from ShotNet, an embedding multilayer perceptron (MLP). On a held-out season it beats a zone-rate baseline and a logistic baseline, and its probabilities are well calibrated. A depth-limited expectimax search then solves the offensive decision tree, with branch-and-bound pruning to keep it real time. All training runs offline, so the online system stays light. A scouting planner and a playable simulator both run in a single browser page.
Chinese Translation
学校教练依靠比赛录像和直觉来准备应对对手,而职业球队的分析工具却难以企及。我们探讨公开数据能在多大程度上缩小这一差距。我们以职业篮球为案例研究对象,选择它是因为其数据丰富,而非联赛本身。我们将五个公开数据源融合为一个单次投篮数据集,涵盖21个赛季的423万次投篮。这些数据源包括投篮位置、两路逐回合比赛数据、官方对位追踪数据以及球员身体测量数据,各数据源之间的对齐率为99.5%至100%。我们还报告了两个容易被忽视的数据陷阱。随后,我们将一次半场进攻回合建模为序贯博弈。投篮价值由ShotNet(一种嵌入型多层感知机,MLP)计算得出。在留出赛季上,它优于区域命中率基线和逻辑回归基线,且其概率具有良好校准性。接着,通过带分支限界剪枝的深度受限期望最大化搜索(expectimax search)求解进攻决策树,以保证实时性。所有训练均离线进行,因此在线系统保持轻量。球探规划器和可玩的模拟器均可运行在单个浏览器页面中。
cs.LG / 59 / 2608.29576

Event-triggered Control and Online Learning for Networked Systems under Computational Delays

计算时延下网络化系统的事件触发控制与在线学习
Dai, Xiaobing, Lederer, Armin, Yang, Zewen, Zhang, Sihua, Wan, Lu, Tang, Yang, Hirche, Sandra
Abstract
Online learning-based control is a promising approach to control uncertain systems, where unknown components are identified during operation to improve control performance. However, resource-intensive online learning algorithms introduce non-negligible computational delays, especially when executed on systems with limited local computational resources. To mitigate this, an in-network online learning-based control structure is employed by deploying the learning-based controller on a remote computation node and connecting it via a communication channel. In this paper, control performance guarantee is first established by deriving tracking error bound for the in-network control architecture, while accounting for computational delays. The derived tracking error bound allows for diverse communication and computation strategies under a specific condition, including time-/event-triggered mechanisms. Additionally, the trade-off between communication and computation performances is shown for a given desired control performance. Furthermore, to enhance the efficiency in both communication and computation, an efficient control framework with an asynchronous event-triggered mechanism in both control and online learning is devised under the existence of computational delay. The proposed event-triggered strategy is proven to achieve the same control performance as time-triggered scenario while excluding Zeno behavior. Finally, we derive an explicit expression of the proposed event-trigger condition for exponentially stabilizable systems, and demonstrate its effectiveness through simulations.
Chinese Translation
基于在线学习的控制是控制不确定系统的一种有前景的方法,其通过在运行过程中辨识未知组件来提升控制性能。然而,资源密集型的在线学习算法会引入不可忽略的计算时延,尤其是在本地计算资源受限的系统上执行时。为缓解这一问题,本文采用一种网络内在线学习控制结构,将基于学习的控制器部署在远程计算节点上,并通过通信信道与之连接。本文首先通过推导网络内控制架构的跟踪误差界,在考虑计算时延的情况下建立了控制性能保证。所推导的跟踪误差界在特定条件下可容纳多样化的通信与计算策略,包括时间触发与事件触发机制。此外,针对给定的期望控制性能,展示了通信性能与计算性能之间的权衡关系。进一步地,为提高通信与计算的效率,在存在计算时延的情况下,设计了一种在控制与在线学习两个环节均采用异步事件触发机制的高效控制框架。所提出的事件触发策略被证明在排除Zeno行为的同时,能够达到与时间触发情形相同的控制性能。最后,针对可指数镇定的系统,推导了所提出的事件触发条件的显式表达式,并通过仿真验证了其有效性。
cs.LG / 60 / 2608.29579

Predicting the Unpredictable: LLM-powered Long-term Chaotic Time Series Forecasting under Short-term Observations

预测不可预测之物:基于短期观测的大语言模型驱动的长期混沌时间序列预测
Yao, Yuhang, Jiang, Bohan
Abstract
Chaotic time series forecasting is a challenging task due to its sensitivity to initial conditions and long-term unpredictability. Traditional methods typically rely on sufficient temporal trajectories to learn long-term dynamics, which limits their applicability when only short-term observations are available. While recent Large Language Models (LLMs) have shown great potential for time series forecasting, their temporal representations are not explicitly tailored to the phase-space structure and nonlinear evolution of chaotic systems. To address these issues, we propose PAC-LLM, a phase-space-aware adaptive fusion framework for long-term chaotic time series forecasting powered by LLMs. PAC-LLM leverages learned phase-space features and textual information to fully enable LLM's time series forecasting capacity. In particular, we design an auxiliary feature module and a gated weighting mechanism for multivariate coupling information fusion and selection. Extensive experiments on representative chaotic systems demonstrate that our method outperforms existing fine-tuned and zero-shot baselines in both short-term and long-term predictions. Our ablation study further confirms the effectiveness of each key component in PAC-LLM.
Chinese Translation
混沌时间序列预测是一项具有挑战性的任务,其原因在于混沌系统对初始条件的高度敏感性以及长期不可预测性。传统方法通常依赖充足的时间轨迹来学习长期动力学,这限制了其在仅有短期观测数据时的适用性。尽管近期的大语言模型(Large Language Models, LLMs)在时间序列预测中展现出巨大潜力,但其时间表示并未显式地针对混沌系统的相空间结构和非线性演化进行定制。为解决这些问题,我们提出了 PAC-LLM,一个由大语言模型驱动的、面向相空间的(phase-space-aware)自适应融合框架,用于长期混沌时间序列预测。PAC-LLM 利用学习到的相空间特征和文本信息,充分激发大语言模型的时间序列预测能力。特别地,我们设计了一个辅助特征模块和一个门控加权机制,用于多变量耦合信息的融合与选择。在代表性混沌系统上的大量实验表明,我们的方法在短期和长期预测中均优于现有的微调和零样本(zero-shot)基线方法。消融实验进一步验证了 PAC-LLM 中各关键组件的有效性。
cs.LG / 61 / 2608.29598

On the Resilience of Text-to-Video Diffusion Models to Hardware Faults

文本到视频扩散模型对硬件故障的鲁棒性研究
Coalson, Zachary, Aahad, A M, Doehring, Stella, Ma, Zane, Hong, Sanghyun
Abstract
We present the first systematic study of the resilience of text-to-video (T2V) diffusion models under random hardware-level faults. While T2V models are widely used for automated video generation due to their ability to produce high-quality, temporally coherent, and realistic videos, their iterative denoising process and spatiotemporal dependencies introduce unique failure modes. We perform an extensive fault-injection study covering both computational and memory faults across three T2V models and a representative benchmark. Our results show that (1) a single fault can degrade overall performance by up to 3.7\%, with semantic correctness more affected than perceptual quality; (2) memory faults are more damaging than computational faults, high-order exponent bits are particularly vulnerable, and the widely-used bfloat16 is more susceptible than alternative formats; and (3) 7-28\% of faults cause visible artifacts, including semantic changes such as added objects, suggesting that single faults are sufficient to alter output semantics. Our findings reveal reliability risks in deployed T2V systems and motivate further research on improving fault resilience. Code: \href{https://github.com/ztcoalson/T2V-Resilience}{https://github.com/ztcoalson/T2V-Resilience}.
Chinese Translation
我们首次系统性地研究了文本到视频(Text-to-Video, T2V)扩散模型在随机硬件级故障下的鲁棒性。T2V模型因其能够生成高质量、时间连贯且逼真的视频而被广泛应用于自动化视频生成,但其迭代去噪过程和时空依赖性引入了独特的失效模式。我们开展了大规模的故障注入实验,涵盖三种T2V模型和一个代表性基准上的计算故障与内存故障。结果表明:(1)单个故障可使整体性能下降高达3.7%,其中语义正确性比感知质量受到的影响更大;(2)内存故障比计算故障更具破坏性,高阶指数位尤其脆弱,且广泛使用的bfloat16格式比其他格式更易受影响;(3)7-28%的故障会导致可见的伪影,包括诸如添加物体等语义变化,这表明单个故障就足以改变输出语义。我们的发现揭示了已部署T2V系统中的可靠性风险,并为进一步提升故障鲁棒性的研究提供了动力。代码:\href{https://github.com/ztcoalson/T2V-Resilience}{https://github.com/ztcoalson/T2V-Resilience}。
cs.LG / 62 / 2608.29600

Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior

面向多样化用户行为的排序策略自适应双重鲁棒离线策略评估
Iguchi, Kosuke, Kishimoto, Ren
Abstract
Off-policy evaluation (OPE) of ranking policies is challenging be- cause selecting and ordering multiple items from a candidate set makes the number of possible rankings grow combinatorially with the number of candidates and the ranking length. Consequently, Inverse Propensity Scoring (IPS), whose importance weight is the full-ranking probability ratio under the evaluation and logging policies, can have excessive variance. Independent IPS (IIPS) and Reward Interaction IPS (RIPS) reduce variance by imposing fixed assumptions on how users browse rankings, but may introduce bias when those assumptions mismatch actual behavior. Adaptive Inverse Propensity Scoring (AIPS) addresses this trade-off by adap- tively marginalizing importance weights over the actions that affect each position-wise reward. It attains minimum variance within a class of unbiased IPS-based estimators when the true user be- havior model is observed. However, its estimation accuracy may still degrade for longer rankings, and AIPS does not use a reward model for residual correction. We propose Adaptive Doubly Robust (ADR), which combines adaptive importance weighting with re- ward regression through a control-variate correction. We establish its unbiasedness when the true user behavior model is observed and characterize a sufficient condition under which it reduces vari- ance relative to AIPS. Across synthetic experiments with 10,000 simulations per condition, ADR improves mean squared error over AIPS and conventional ranking OPE estimators across a range of logged-data sizes and ranking lengths.
Chinese Translation
排序策略的离线策略评估(Off-Policy Evaluation, OPE)具有挑战性,因为从候选集中选择并排序多个项目使得可能的排序数量随候选数量和排序长度呈组合式增长。因此,逆倾向得分(Inverse Propensity Scoring, IPS)方法——其重要性权重为评估策略与记录策略下完整排序的概率比——可能产生过大的方差。独立IPS(IIPS)和奖励交互IPS(RIPS)通过对用户浏览排序的方式施加固定假设来降低方差,但当这些假设与实际行为不符时可能引入偏差。自适应逆倾向得分(Adaptive Inverse Propensity Scoring, AIPS)通过对影响每个位置奖励的动作自适应地进行边缘化的重要性权重,来应对这一权衡。当真实用户行为模型可被观测时,它在基于IPS的无偏估计器类别中达到最小方差。然而,对于较长的排序,其估计精度仍可能下降,且AIPS未使用奖励模型进行残差校正。我们提出自适应双重鲁棒(Adaptive Doubly Robust, ADR)方法,通过控制变量校正将自适应重要性加权与奖励回归相结合。我们证明了当真实用户行为模型可被观测时该方法的无偏性,并刻画了其相对于AIPS能够降低方差的一个充分条件。在每种条件下进行10,000次模拟的合成实验中,ADR在多种记录数据规模和排序长度下均比AIPS及传统排序OPE估计器取得了更低均方误差。
cs.LG / 63 / 2608.29608

Wide Learning: Learning to Reach Evidence

宽学习:学习触及证据
Chen, Junzhou
Abstract
Machine learning is usually evaluated after an evidence interface has been fixed. A dataset, sensor suite, query language, action set, or experimental protocol determines which observations can be obtained, and learning is judged by what it extracts from them. We study a complementary capability. A learner's state can determine which evidence-generating experiments it can reliably realise under bounded resources, even when primitive affordances remain fixed. We call this learner-relative experiment family its effective epistemic reach, and use Wide Learning for task-relevant learning-induced changes in that family.We formalise effective reach relative to learner state, deployment budget, reliability threshold, and evaluation distribution. In a controlled construction, two hidden worlds have exactly the same public observation law. An informative diagnostic exists in a fixed five-primitive substrate. Before calibration, one address attempt realises it with probability at most $2^{-10} = 1/1024$, below a pre-specified 0.95 threshold; after calibration, held-out realisation is 1. Public-channel total variation is 0, whereas the realised diagnostic has total variation 1, and sealed binary risk moves from approximately 1/2 to 0. The construction establishes that learning can change effective epistemic reach even when primitive affordances and deployment resources are held fixed. It opens a complementary evaluation question for learning systems: not only what they infer from available evidence, but what informative evidence experience teaches them to bring within reach.
Chinese Translation
机器学习通常是在证据接口被固定之后进行评估的。数据集、传感器套件、查询语言、动作集合或实验协议决定了可以获取哪些观测,而学习则根据其从这些观测中提取的内容来评判。我们研究一种互补的能力:学习器的状态可以决定它在有限资源下能够可靠实现哪些产生证据的实验,即使基本可供性(primitive affordances)保持不变。我们将这种相对于学习者的实验族称为其有效认知可达范围(effective epistemic reach),并将宽学习(Wide Learning)定义为对该族中与任务相关的、由学习引起的变化。我们相对于学习器状态、部署预算、可靠性阈值和评估分布对有效可达范围进行形式化。在一个受控构造中,两个隐藏世界拥有完全相同的公共观测规律。在一个固定的五原语基底中存在一个信息性诊断,在校准之前,一次寻址尝试以不超过 $2^{-10}=1/1024$ 的概率实现该诊断,低于预先设定的0.95阈值;校准之后,保留集上的实现率达到1。公共信道的总变差为0,而已实现的诊断的总变差为1,且密封的二元风险从约1/2降至0。该构造表明,即使基本可供性和部署资源保持固定,学习仍可以改变有效认知可达范围。这为学习系统开启了一个互补的评估问题:不仅考察它们从可用证据中推断出什么,还要考察经验教会它们将哪些信息性证据纳入可达范围。
cs.LG / 64 / 2608.29635

Unsupervised Multi-Scale Gromov-Wasserstein Hypergraph Alignment

无监督多尺度Gromov-Wasserstein超图对齐
Oettershagen, Lutz, Wang, Honglian, Gionis, Aristides
Abstract
We study unsupervised hypergraph alignment, where the goal is to infer node correspondences between two hypergraphs using only structural information, without node features, labels, seed matches, or side information. Direct higher-order formulations can represent hyperedge interactions faithfully, but they can be computationally demanding and cumbersome for non-uniform hypergraphs. Graph-reduction approaches introduce a different challenge: clique expansions keep the alignment problem on the original node set but collapse all hyperedge evidence into one pairwise graph, whereas bipartite expansions preserve incidence structure but enlarge the problem from nodes to nodes plus hyperedges. We introduce FALCON (Filtration-based hypergrAph aLignment via Cross-scale Optimal traNsport), an unsupervised optimal-transport framework for hypergraph alignment. Instead of representing each hypergraph by a single collapsed clique graph, FALCON constructs a filtration-induced sequence of clique-based co-occurrence dissimilarity matrices and jointly aligns all levels through one shared multi-scale Gromov--Wasserstein (GW) objective. The shared transport plan enforces a globally consistent node correspondence across filtration levels while avoiding the auxiliary hyperedge nodes introduced by bipartite expansion. Experiments on perturbation benchmarks derived from real-world hypergraphs show that FALCON is robust to structural noise and in almost all cases outperforms strong graph- and hypergraph-alignment baselines.
Chinese Translation
我们研究无监督超图对齐问题,其目标是仅利用结构信息推断两个超图之间的节点对应关系,而不依赖节点特征、标签、种子匹配或辅助信息。直接的高阶形式化方法能够忠实表示超边交互,但在计算上代价高昂,且对非均匀超图处理繁琐。图约简方法则带来另一类挑战:团扩展(clique expansion)将对齐问题保持在原始节点集上,但将所有超边信息压缩到单一的成对图中;而二部图扩展虽然保留了关联结构,却将问题从节点扩大到节点加超边。我们提出FALCON(Filtration-based hypergrAph aLignment via Cross-scale Optimal traNsport,基于滤过序列与跨尺度最优传输的超图对齐),这是一个用于超图对齐的无监督最优传输框架。FALCON不再用单一的压缩团图表示每个超图,而是构建由滤过(filtration)诱导的基于团的共现相异度矩阵序列,并通过一个共享的多尺度Gromov-Wasserstein(GW)目标函数联合对齐所有层级。共享的传输计划在跨层级间强制施加全局一致的节点对应关系,同时避免了二部图扩展引入的辅助超边节点。在基于真实世界超图构造的扰动基准上的实验表明,FALCON对结构噪声具有鲁棒性,并且在几乎所有情况下优于强大的图对齐和超图对齐基线方法。
cs.LG / 65 / 2608.29640

LLMODE: Aligning ODEs with LLMs via Gated Token Injection for Irregular Spatio-Temporal Forecasting

LLMODE:通过门控令牌注入将ODE与LLM对齐以实现不规则时空预测
Zhang, Di, Zhang, Jingyang, Wang, Ziqian, Zhang, Chi, Ban, Yikun, Zhang, Ziwei, Wang, Ruijie
Abstract
Large language models (LLMs) have shown promise for spatio-temporal forecasting, but existing approaches often rely on regularly sampled token sequences and struggle with irregular observations because of temporal asynchrony, representation-space misalignment, and limited context windows. We propose LLMODE, a token-efficient framework for irregular spatio-temporal forecasting with a frozen LLM backbone. LLMODE first uses a graph-aware ODE encoder to reconstruct irregular graph observations as a continuous-time latent trajectory. A Fixed-Budget Perceiver Resampler then compresses this variable-length trajectory into a fixed number of dynamic memory tokens. In parallel, compact statistical descriptors are encoded and resampled into context memory tokens. A dual-source gated cross-attention module injects both memories into the frozen LLM, enabling controlled utilization of external spatio-temporal evidence. Experiments on three real-world urban datasets and two physical-dynamics benchmarks show competitive overall performance, with clearer advantages under sparse or dynamically complex irregular sampling. Additional evaluations on unseen urban regions further demonstrate strong zero-shot generalization without adaptation.
Chinese Translation
大语言模型(LLM)在时空预测方面展现出潜力,但现有方法通常依赖于规则采样的令牌序列,由于时间异步、表示空间不对齐以及上下文窗口有限等问题,难以处理不规则观测数据。我们提出了LLMODE,一个基于冻结LLM骨干网络的令牌高效的不规则时空预测框架。LLMODE首先使用图感知的ODE编码器(graph-aware ODE encoder)将不规则的图观测数据重构为连续时间的潜在轨迹;随后,固定预算感知重采样器(Fixed-Budget Perceiver Resampler)将该变长轨迹压缩为固定数量的动态记忆令牌。与此同时,紧凑的统计描述符被编码并重采样为上下文记忆令牌。双源门控交叉注意力模块将这两类记忆注入冻结的LLM中,从而实现对外部时空证据的可控利用。在三个真实世界城市数据集和两个物理动力学基准上的实验表明,该模型整体性能具有竞争力,在稀疏或动态复杂的不规则采样条件下优势更为明显。在未见过的城市区域上的进一步评估表明,模型无需适配即具备强大的零样本泛化能力。
cs.LG / 66 / 2608.29647

Reward-guided Fine-Tuning of One-Step Generative Models via Wasserstein Gradient Flow

基于Wasserstein梯度流的单步生成模型奖励引导微调
Hwang, Hoseong, Han, Woorim, Chun, Joungin, Park, Jinseong, Choi, Jaewoong
Abstract
To mitigate the time complexity of generative models, one-step generative models have recently emerged through direct mapping from noise to data in a single forward pass. However, the reward-guided fine-tuning method of one-step generative models remains largely unexplored. To address this, we consider one-step generators from an optimal transport view, investigating Wasserstein Gradient Flow (WGF) for modeling smooth and controlled distributional evolution in probability space. We then propose a novel reward-guided fine-tuning of a one-step generative model via WGF. We derive a practical training method that requires no reward gradients, thereby handling both non-differentiable and differentiable rewards. Moreover, our method provides smooth and stable reward-guided distributional updates while mitigating reward hacking and mode collapse. Experiments on 2D synthetic data, CIFAR-10, and ImageNet 256$\times$256 with diverse rewards, including JPEG (in)compressibility, class probability, Black-and-White and CLIP alignment, show that our method achieves better reward alignment compared to baselines.
Chinese Translation
为了降低生成模型的时间复杂度,单步生成模型近年来应运而生,其通过单次前向传播实现从噪声到数据的直接映射。然而,针对单步生成模型的奖励引导微调方法仍鲜有研究。为此,我们从最优传输的视角考察单步生成器,研究Wasserstein梯度流(WGF)以在概率空间中建模平滑且可控的分布演化。在此基础上,我们提出了一种基于WGF的新型单步生成模型奖励引导微调方法。我们推导了一种实用的训练方法,该方法无需奖励梯度,因而可以同时处理不可微和可微的奖励。此外,我们的方法在实现平滑稳定的奖励引导分布更新的同时,还能缓解奖励破解(reward hacking)和模式坍缩问题。在二维合成数据、CIFAR-10以及ImageNet 256×256上的实验中,我们采用了多种奖励,包括JPEG(不)可压缩性、类别概率、黑白化以及CLIP对齐。结果表明,与基线方法相比,我们的方法取得了更好的奖励对齐效果。
cs.LG / 67 / 2608.29667

A Target-Centric Survey of Quantization-Aware Training

以目标为中心的量化感知训练综述
Song, Jiamin, Zhao, Mengjie, Wang, Zijing, Liu, Yongkang, Li, Qian, Feng, Shi, Ren, Feiliang, Wang, Daling, Schütze, Hinrich
Abstract
The rapid development of LLMs incurs prohibitive memory footprints and intensive computational demands. Quantization-Aware Training (QAT) techniques have emerged as a promising solution to address these challenges by explicitly simulating quantization effects during model training, yielding low-bit models that achieve accuracy comparable to their full-precision counterparts. In this work, we provide a target-centric survey of QAT, aimed at clarifying both its theoretical foundations and its evolving implementation landscape. We systematically review existing QAT methods through a target-centric taxonomy and synthesize cross-target differences in error characteristics, numerical formats, and strategy transferability. We further summarize QAT evaluation paradigms and discuss challenges in optimization and deployment, outlining potential directions for future research.
Chinese Translation
大语言模型(LLM)的快速发展带来了高昂的内存占用和密集的计算需求。量化感知训练(Quantization-Aware Training, QAT)技术通过在模型训练过程中显式地模拟量化效应,成为应对这些挑战的一种有前景的解决方案,所得到的低比特模型可以达到与全精度模型相当的精度。在本工作中,我们提供了一项以目标为中心的QAT综述,旨在阐明其理论基础及不断演进的技术实现图景。我们通过以目标为中心的分类体系系统地梳理了现有的QAT方法,并综合分析了不同目标之间在误差特性、数值格式以及策略可迁移性方面的差异。我们进一步总结了QAT的评估范式,讨论了优化与部署方面的挑战,并展望了未来研究的潜在方向。
cs.LG / 68 / 2608.29674

Creation begins with understanding: LLMs as strategy designers for privacy-preserving tabular data synthesis

创造始于理解:大语言模型作为隐私保护表格数据合成的策略设计者
Li, Jinmeng, Zhang, Quan, Ye, Hangting, Zhao, He, Laakom, Firas, Guo, Dandan, Schmidhuber, Jürgen
Abstract
Sharing tabular data in high-stakes domains is constrained by privacy regulations. Synthetic data offer a promising alternative, but deep generative models are costly to train and difficult to audit, while LLM-based methods often serialize records as text, obscuring tabular structure and exposing sensitive data. We introduce Tabular Synthesis Strategy Designer (TabSSD), which uses an LLM to design synthesis procedures rather than directly generate records. TabSSD provides the LLM with tree-derived summaries of variable dependence rather than raw records, which produces Python programs for local execution and evaluation. Across twelve datasets, TabSSD strikes a favourable balance among statistical fidelity, predictive utility, and empirical privacy risk, achieving the best average rank across six metrics among ten methods. Moreover, it substantially reduces local computation and token consumption relative to the compared methods. By enabling human-guided refinement and eliminating user-side model tuning, TabSSD lowers the expertise and infrastructure barriers to transparent tabular data synthesis.
Chinese Translation
在高风险领域,表格数据的共享受到隐私法规的约束。合成数据提供了一种有前景的替代方案,但深度生成模型训练成本高且难以审计,而基于大语言模型(LLM)的方法通常将记录序列化为文本,既掩盖了表格结构,又暴露了敏感数据。我们提出了表格合成策略设计器(Tabular Synthesis Strategy Designer,TabSSD),它利用LLM设计合成流程,而非直接生成记录。TabSSD向LLM提供基于树结构的变量依赖关系摘要,而非原始记录,从而生成可在本地执行和评估的Python程序。在十二个数据集上的实验表明,TabSSD在统计保真度、预测效用和经验隐私风险之间取得了良好的平衡,在十种方法的六项指标中获得了最佳平均排名。此外,与对比方法相比,它显著降低了本地计算量和token消耗。通过支持人工引导的优化并消除用户端的模型调优,TabSSD降低了透明表格数据合成所需的专业知识和基础设施门槛。
cs.LG / 69 / 2608.29685

Last Step Matters: Early Uncertainty Cannot Predict Failure in Long-Horizon Agents

最后一步才重要:早期不确定性无法预测长程智能体的失败
Li, Zongyue, Yu, Chengyue, Zang, Lei, Zhuang, Chenyi, Mo, Linjian, Gan, Leilei
Abstract
Early failure prediction is important for long-horizon agents, as it enables timely intervention and can reduce inference and tool-use costs. Uncertainty quantification, such as verbal confidence and perplexity, offers a promising approach to detecting agent failures; however, it has not been explored whether these signals retain their discriminative power during the intermediate stages of long-horizon execution. We evaluate mainstream uncertainty signals on deep-research tasks and find that verbal confidence reliably distinguishes failures at trajectory completion, achieving a mean AUROC of 0.85, whereas all evaluated signals offer limited predictive value earlier in execution, with none exceeding a mean AUROC of 0.60 at 50% trajectory progress. We identify an underlying mechanism explaining this gap: path switching, where agents frequently abandon their current search direction in-trajectory, breaking the link between early signal and final outcome. These findings challenge the assumption that intermediate uncertainty can reliably guide early intervention. They also motivate a practical recommendation for agent harnesses in deep-research settings: use final-step confidence to decide whether to restart, an approach that our experiments find more effective than in-trajectory intervention.
Chinese Translation
对长程智能体(long-horizon agents)而言,失败的早期预测非常重要,因为它能够实现及时干预,并降低推理与工具使用的成本。不确定性量化方法,如语言化置信度(verbal confidence)和困惑度(perplexity),为检测智能体失败提供了一种有前景的途径;然而,这些信号在长程执行的中间阶段是否仍保持其判别能力,尚未得到探究。我们在深度研究(deep-research)任务上评估了主流的不确定性信号,发现语言化置信度在轨迹完成时能够可靠地区分失败,平均 AUROC 达到 0.85,而所有被评估的信号在执行早期仅具有有限的预测价值,在轨迹进度达到 50% 时没有任何信号的平均 AUROC 超过 0.60。我们识别出解释这一差距的潜在机制:路径切换(path switching),即智能体在轨迹内频繁放弃当前的搜索方向,从而切断了早期信号与最终结果之间的联系。这些发现挑战了中间阶段不确定性能够可靠指导早期干预的假设。同时,这些发现也为深度研究场景下的智能体框架(agent harnesses)提供了一个实用的建议:使用最后一步的置信度来决定是否重新启动,我们的实验表明这一方法比轨迹内干预更为有效。
cs.LG / 70 / 2608.29715

Higher-Dimensional Rotary Position Embedding

高维旋转位置编码
Li, Yixing, Xie, Ruobing, Zhang, Yudong, Bai, Yushi, Sun, Samm, Cheng, Yu
Abstract
Transformers rely on position embedding mechanisms in long context modeling in most cases. Rotary Position Embedding (RoPE) embeds positional information with independent 2D rotations, forming relative position terms in self-attention. However, its pairwise, block-based, and decoupled structure limits deep mixing and robustness across channels. We propose HD-RoPE, which extends RoPE from independent 2D rotations to higher-dimensional rotations and introduces a Paley-I orthogonal basis to obtain balanced, isotropic, and dense phase mixing within each rotation subspace. This significantly enhances channel coupling and rotational degrees of freedom while maintaining orthogonal stability and the relative position closure property. Furthermore, HD-RoPE is easily optimized for engineering efficiency without introducing additional trainable parameters. We have conducted extensive evaluation results demonstrating that HD-RoPE achieves significant performance improvements over standard RoPE across various popular benchmarks and in both long and short contexts.
Chinese Translation
在长上下文建模中,Transformer在大多数情况下依赖于位置嵌入机制。旋转位置编码(Rotary Position Embedding, RoPE)通过独立的二维旋转嵌入位置信息,在自注意力中形成相对位置项。然而,其成对的、分块的且解耦的结构限制了深层的通道混合能力和跨通道的鲁棒性。我们提出HD-RoPE,将RoPE从独立的二维旋转扩展到高维旋转,并引入Paley-I正交基,在每个旋转子空间内实现平衡、各向同性且稠密的相位混合。这显著增强了通道耦合和旋转自由度,同时保持了正交稳定性和相对位置封闭性。此外,HD-RoPE易于优化且具有工程效率,不引入额外的可训练参数。我们进行了广泛的评估,结果表明HD-RoPE在多个流行基准测试以及长、短上下文场景中均显著优于标准RoPE。
cs.LG / 71 / 2608.29755

GraM-Diff: A Unified Graph-Mamba Diffusion Framework for EEG-Based Alzheimer's Disease Data Generation and Diagnosis

GraM-Diff:用于基于脑电图(EEG)的阿尔茨海默病数据生成与诊断的统一Graph-Mamba扩散框架
Tanveer, M., Rana, Ayush Singh, Jain, Sanskriti, Kumar, Arnav, Tiwari, Aryaman, Rahaman, A., Quadir, A., Sajid, M.
Abstract
Electroencephalography (EEG) is a promising, non-invasive, and cost-effective modality for Alzheimer's disease (AD) detection, but deep learning methods are limited by small and imbalanced clinical datasets. Generative augmentation offers a solution, yet existing approaches rely on inefficient class-specific models or fail to capture complex spatial and temporal brain dynamics. To address this, we propose GraM-Diff, a unified classifier-guided Graph-Mamba diffusion framework for EEG synthesis. It embeds Graph Convolutional Networks within a diffusion U-Net to model inter-electrode connectivity and Bidirectional Mamba state-space blocks for linear-complexity long-range temporal modeling. Latent-space classifier guidance lets a single model generate both healthy and pathological EEG within a shared representation, avoiding fragmented per-cohort pipelines. Across four EEG-based AD benchmarks, synthetic augmentation improves classification, yields superior Context-FID and correlation scores over strong generative baselines, and enhances robustness in data-scarce settings.
Chinese Translation
脑电图(EEG)是一种有前景的、无创且经济高效的阿尔茨海默病(AD)检测模态,但深度学习方法受限于规模小且不平衡的临床数据集。生成式数据增强提供了一种解决方案,然而现有方法依赖于低效的类别专用模型,或无法捕捉复杂的脑区时空动态。为解决这一问题,我们提出了GraM-Diff,一个用于EEG合成的统一分类器引导的Graph-Mamba扩散框架。该框架将图卷积网络(Graph Convolutional Networks)嵌入扩散U-Net中以建模电极间的连通性,并采用双向Mamba状态空间模块实现线性复杂度的长程时间建模。借助潜空间分类器引导,单一模型可在共享表示中同时生成健康与病理状态的EEG,避免了碎片化的分队列处理流程。在四个基于EEG的阿尔茨海默病基准数据集上,合成数据增强提升了分类性能,在Context-FID和相关系数得分上优于强大的生成式基线方法,并增强了在数据稀缺场景下的鲁棒性。
cs.LG / 72 / 2608.29763

ECA-BLS: An Efficient Complex-Augmented Broad Learning System

ECA-BLS:一种高效复数增广的广度学习系统
Rahaman, A., Quadir, A., Sajid, M., Akhtar, M., Tanveer, M.
Abstract
Broad Learning System (BLS) is an efficient alternative to deep architectures due to its fast training, analytical learning, and strong generalization under limited data. However, existing BLS variants are confined to real-valued representations, restricting their ability to capture nonlinear interactions and second-order statistical dependencies inherent in real-world data. Notably, no prior BLS model fully exploits the complete second-order statistics that naturally emerge when data are embedded in the complex domain. To address this limitation, this paper introduces the first complex augmented Broad Learning System (CA-BLS), which transforms real-valued inputs into phase-encoded complex representations and adopts widely linear modeling to jointly leverage covariance and pseudo-covariance information via complex conjugate augmentation. This enables effective modeling of latent nonlinearities, coherence structures, and second-order dependencies inaccessible to conventional BLS formulations. To mitigate the additional computational cost of complex augmentation, an Efficient Complex Augmented BLS (ECA-BLS) is further developed, reformulating CA-BLS entirely in the real domain while preserving its exact decision function, achieving up to 75\% fewer multiplications and over 60\% fewer additions. A rigorous theoretical analysis proves the mathematical equivalence between CA-BLS and ECA-BLS, ensuring zero theoretical loss. Extensive experiments on 26 benchmark datasets from the UCI and KEEL repositories demonstrate that ECA-BLS consistently outperforms classical BLS and recent state-of-the-art randomized neural networks in accuracy, average rank, and statistical significance, establishing augmented second-order modeling as a critical and previously missing dimension of BLS research.
Chinese Translation
广度学习系统(Broad Learning System, BLS)凭借其训练快速、解析式学习以及在有限数据下良好的泛化能力,成为深度架构的一种高效替代方案。然而,现有的BLS变体仅局限于实值表示,限制了其捕捉现实世界数据中固有的非线性交互和二阶统计依赖的能力。值得注意的是,此前尚无BLS模型能够充分利用数据嵌入复数域后自然产生的完整二阶统计量。为解决这一局限,本文首次提出了复数增广广度学习系统(Complex-Augmented Broad Learning System, CA-BLS),该方法将实值输入转换为相位编码的复数表示,并采用广义线性(widely linear)建模,通过复共轭增广联合利用协方差与伪协方差信息。这使得模型能够有效刻画传统BLS形式无法触及的潜在非线性、相干结构以及二阶依赖关系。为降低复数增广带来的额外计算开销,本文进一步提出了高效复数增广广度学习系统(Efficient Complex-Augmented BLS, ECA-BLS),将CA-BLS完全重 formulations 于实数域,同时保持其精确的决策函数不变,使乘法运算最多减少75%,加法运算减少60%以上。严格的理论分析证明了CA-BLS与ECA-BLS在数学上的等价性,确保了理论上的零损失。在来自UCI和KEEL数据集库的26个基准数据集上的大量实验表明,ECA-BLS在准确率、平均排名和统计显著性方面均持续优于经典BLS以及近期的最先进随机化神经网络,确立了二阶增广建模作为BLS研究中一个关键且此前缺失的维度。
cs.LG / 73 / 2608.29765

PruneShift: A Framework for Evaluating Decision Reliability in Structured Pruning

PruneShift:一个用于评估结构化剪枝决策可靠性的框架
Ye, Hao, Zhang, Gaopeng
Abstract
Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too expensive. Most evaluations report average surrogate error or rank correlation on broadly sampled masks. These summaries do not directly test the mask chosen by the surrogate. We introduce PruneShift, an evaluation framework that separates broad predictive fidelity, fidelity near selector outputs, and the quality of the selected pruning decision. We first prove that Spearman and Kendall agreement can approach one while normalized selection regret remains maximal. We then derive sufficient conditions based on uniform error, selector suboptimality, decision margin, density ratio, and comparison mass. The analysis also yields a finite pool certificate with an explicit excess cost bound. Four studies test different links in this argument. External TextbookQA confirmation is heterogeneous: 7 of 20 simultaneous intervals favor the surrogate-selected mask, 6 favor its fixed comparator, and 7 cross zero. On a fixed Natural Questions pool, strict improvement holds in one of four settings. A controlled QQP experiment supports the proposed coverage mechanism in all 16 prespecified endpoints, although the sufficient bounds are conservative. Finally, a restricted OSSCAR reconstruction study on OPT-125M shows better local than broad fidelity in 68 of 75 primary endpoints. Independent fixed-mask confirmation is inconclusive in 24 of 25 endpoints and favors the comparator in one. These results show why predictive fit, decision reliability, and pruning method quality require separate evidence.
Chinese Translation
结构化剪枝依赖代理目标函数,因为对每个可行的掩码进行直接任务评估的代价过高。大多数评估工作在广泛采样的掩码上报告平均代理误差或秩相关性,但这些汇总指标并未直接检验由代理方法所选出的掩码。我们提出PruneShift,一个将广泛预测保真度、选择器输出附近的保真度以及所选定剪枝决策的质量三者区分开来的评估框架。我们首先证明,Spearman和Kendall一致性系数可以趋近于1,而归一化选择遗憾仍保持最大值。随后,我们基于均匀误差、选择器次优性、决策间隔、密度比和比较质量推导出充分条件。该分析还给出了一个具有显式超额成本界的有限掩码池证书。四项实验检验了这一论证中的不同环节。外部TextbookQA验证结果呈异质性:20个同时置信区间中有7个支持代理方法所选掩码,6个支持其固定对照掩码,7个跨越零点。在一个固定的Natural Questions掩码池上,四个实验设置中仅有一个满足严格改进。一项受控的QQP实验在全部16个预设检验终点上支持所提出的覆盖机制,尽管充分界较为保守。最后,在OPT-125M上开展的受限OSSCAR重建研究显示,75个主要终点中有68个呈现局部保真度优于广泛保真度的结果。独立的固定掩码验证在25个终点中有24个结论不明确,仅有1个支持对照掩码。这些结果表明,预测拟合、决策可靠性与剪枝方法质量需要分别提供证据。
cs.LG / 74 / 2608.29817

Structure Aware Neural Architecture Search for Mixture of Experts

面向专家混合模型的结构感知神经架构搜索
Babkin, Petr, Bakhteev, Oleg
Abstract
Neural Architecture Search (NAS) has so far rarely been applied to Mixture-of-Experts (MoE) models, and existing MoE designs leave the alignment between experts and the structure of the data to emerge on its own. We propose an architecture search framework that makes this alignment an explicit search variable: the assignment of data clusters to experts is optimised jointly with the per-expert architectures. We cast the joint problem as a cluster-aware likelihood maximisation, show that it coincides with the incomplete-data maximum likelihood of a latent-variable mixture, and solve it by a generalised Expectation-Maximisation procedure whose otherwise intractable expert-quality term is supplied by an adaptively refined surrogate. We prove that the iterates converge whenever the surrogate errors are summable, and that at every limit point no candidate the search produces improves the true objective. On a heterogeneous image-classification mixture the method recovers the underlying domain partition on 95% of clusters without ever observing domain labels, and on that benchmark and a four-domain time-series forecasting one alike it outperforms the MoE and NAS baselines that likewise use no label information.
Chinese Translation
神经架构搜索(NAS)迄今很少被应用于专家混合模型,且现有的MoE设计将专家与数据结构之间的对齐留待自发形成。我们提出一个架构搜索框架,将这种对齐作为显式的搜索变量:数据簇到专家的分配与各专家的架构被联合优化。我们将该联合问题表述为簇感知的似然最大化,证明其等价于一个隐变量混合模型的不完全数据最大似然,并通过广义期望最大化(Expectation-Maximization)过程求解;该过程中原本难以处理的专家质量项由一个自适应精化的代理模型提供。我们证明,只要代理误差可求和,迭代过程即收敛,且在每个极限点处,搜索所产生的任何候选解都无法进一步改进真实目标函数。在一个异构图像分类混合任务上,该方法在从未观测到域标签的情况下,于95%的簇上恢复了潜在的域划分;在该基准以及一个四域时间序列预测基准上,本方法均优于同样不使用任何标签信息的MoE和NAS基线方法。
cs.LG / 75 / 2608.29850

Designing for the Next Click: Bandits for Real-Time Page Layout

面向下一次点击的设计:用于实时页面布局的老虎机算法
Rath, Bhavtosh, Narasimhamurthy, Harshith, Eisinger, Bob, Stiegler, Cole, Awow, Adnan, Pande, Amit
Abstract
E-commerce platforms increasingly personalize user experiences through machine learning, yet page layout decisions remain dominated by static rules and manual curation. We present a scalable bandit-based system that optimizes product page layouts in real time while preserving human control over design intent. A contextual bandit model dynamically selects the most effective layout for each session using user, item, and category-level features. The system leverages a LinUCB-based policy to balance exploration and exploitation as it learns from live user interactions. The architecture is designed for seamless integration into large-scale web serving stacks, supporting low-latency inference and continuous model updates. The system was first tested on entry product pages. In online A/B deployments on a major retail platform, our approach achieved positive lifts in session-level performance metrics over a strong heuristic baseline. Our results demonstrate that contextual bandits can effectively optimize visual and structural aspects of product discovery for user engagement, providing a scalable path toward learning-to-design the web.
Chinese Translation
电子商务平台日益通过机器学习实现用户体验的个性化,然而页面布局决策仍然主要依赖静态规则和人工策划。我们提出了一种可扩展的、基于老虎机(Bandit)算法的系统,能够在保持人类对设计意图控制的同时,实时优化产品页面布局。该系统采用上下文老虎机(contextual bandit)模型,利用用户、商品和类目级特征,为每个会话动态选择最有效的布局。系统基于LinUCB策略,在从真实用户交互中学习的过程中平衡探索与利用。该架构专为无缝集成到大规模Web服务系统而设计,支持低延迟推理和模型的持续更新。该系统首先在产品入口页面上进行了测试。在某大型零售平台的在线A/B测试部署中,相较于强启发式基线,我们的方法在会话级性能指标上取得了正向提升。我们的结果表明,上下文老虎机算法能够有效优化产品发现过程中影响用户参与度的视觉和结构层面因素,为“学习式设计Web”提供了一条可扩展的路径。
cs.LG / 76 / 2608.29860

Uncertainty-Driven Replay Memory for Reinforcement Learning

面向强化学习的不确定性驱动回放记忆
Rajakrishnan, Sheeraja, Ororbia, Alexander G., Desell, Travis, Krutz, Daniel E.
Abstract
Uncertainty estimation provides promising capabilities for reinforcement learning (RL) agents. Notably, estimating uncertainty can reduce the training time and enable agents to obtain greater rewards over time by exploiting information related to whether an action would facilitate exploration of portions of an environment that are well-known versus those that are relatively unknown. In this work, we propose a novel formulation of the experience replay buffer commonly used in RL that we call uncertainty-driven replay memory (UDRM), which entails an update scheme for internally stored memories based on uncertainty estimates obtained by an RL model during training. In contrast to existing forms of RL, which typically use temporal difference error or the distribution of transitions to update the replay memory buffer and train RL controllers, our scheme biases the memory buffer to store more uncertain transitions that will improve an RL agent's generalization throughout training. Experimental results demonstrate that our proposed uncertainty-aware replay buffer enables an RL agent to obtain higher rewards during training compared to other existing uncertainty-aware RL frameworks.
Chinese Translation
不确定性估计为强化学习(RL)智能体提供了颇具前景的能力。值得注意的是,估计不确定性可以缩短训练时间,并使智能体随时间推移获得更高的奖励,其方式是利用与动作是否有助于探索环境中已知区域或相对未知区域相关的信息。在本工作中,我们提出了强化学习中常用的经验回放缓冲区的一种新形式,称为不确定性驱动回放记忆(Uncertainty-Driven Replay Memory, UDRM),它包含一种针对内部存储记忆的更新机制,该机制基于强化学习模型在训练过程中获得的不确定性估计。现有的强化学习形式通常使用时序差分误差或状态转移分布来更新回放记忆缓冲区并训练强化学习控制器,与之不同,我们的方案使记忆缓冲区偏向存储更多不确定的转移,从而在整个训练过程中提升强化学习智能体的泛化能力。实验结果表明,与现有的其他不确定性感知强化学习框架相比,我们提出的不确定性感知回放缓冲区能使强化学习智能体在训练过程中获得更高的奖励。
cs.LG / 77 / 2608.29867

Partially Linear Autoencoders for Manifold Learning and Dimensionality Reduction

用于流形学习和降维的部分线性自编码器
Pottier, Louen, Lesueur, Louis, Thorin, Anders
Abstract
Autoencoders are widely used for nonlinear dimensionality reduction and manifold learning. While most common implementations rely on both nonlinear encoders and decoders, we investigate the specific role of the encoder and the extent to which it can be constrained to be linear without reducing accuracy. We conduct a comparative study on four autoencoder architectures: standard fully nonlinear autoencoders (AE), linear-encoder autoencoders (Lenc-AE), linear-decoder autoencoders (Ldec-AE), and fully linear autoencoders (LAE), evaluated on synthetic manifolds, computational mechanics data sets, and real-world image data sets including MNIST. We demonstrate that imposing a linear encoder preserves most of the representational capacity of the autoencoder, provided the decoder remains nonlinear. In particular, Lenc-AE consistently outperforms both Ldec-AE and LAE, and achieves reconstruction quality comparable to fully nonlinear AE, while offering advantages in terms of parsimony and interpretability of the latent representation. These results suggest that the nonlinear decoder is the critical component for manifold learning, rather than the encoder. A geometric interpretation of this finding is developed, which identifies the precise conditions under which a linear encoder is sufficient, and the specific manifold configurations that expose its limitations.
Chinese Translation
自编码器(Autoencoders)被广泛用于非线性降维和流形学习。尽管大多数常见实现同时依赖于非线性编码器和非线性解码器,我们研究了编码器的具体作用,以及在不降低精度的情况下将其约束为线性的可行程度。我们对四种自编码器架构进行了比较研究:标准的全非线性自编码器(AE)、线性编码器自编码器(Lenc-AE)、线性解码器自编码器(Ldec-AE)以及全线性自编码器(LAE),并在合成流形、计算力学数据集以及包括 MNIST 在内的真实世界图像数据集上进行了评估。我们证明,只要解码器保持非线性,对编码器施加线性约束可以保留自编码器的大部分表示能力。特别是,Lenc-AE 的性能始终优于 Ldec-AE 和 LAE,并达到了与全非线性 AE 相当的重建质量,同时在简洁性和潜在表示的可解释性方面具有优势。这些结果表明,非线性解码器才是流形学习的关键组件,而非编码器。我们对这一发现给出了一种几何解释,明确了线性编码器足以胜任的精确条件,以及暴露其局限性的特定流形构型。
cs.LG / 78 / 2608.29869

Towards an Expressivity-Normalized Energy-Demand Comparison of ANNs and SNNs

面向表达力归一化的人工神经网络与脉冲神经网络能耗比较
Kranzlmüller, Miriam, Esser, Pascal, Kutyniok, Gitta
Abstract
Spiking neural networks (SNNs) are often regarded as energy-efficient alternatives to artificial neural networks (ANNs), yet their advantage depends critically on both network architecture and data properties. We develop an analytical framework to compare fully-connected ReLU ANNs and integrate-and-fire SNNs for time-series data with respect to their theoretical energy efficiency at matched expressive capacity. By relating an inference-energy model to theoretical bounds on representational expressivity, we derive an expressivity-normalized efficiency ratio and explicit thresholds in network width, spike sparsity, and ANN depth scaling. Our analysis characterizes the regimes in which event-driven computation offsets the temporal overhead of SNNs, providing capacity-aware principles for designing energy-efficient temporal networks. It shows that ANNs exceed SNNs in expressivity-normalized efficiency only in specific regimes.
Chinese Translation
脉冲神经网络(SNN)常被视为人工神经网络(ANN)的节能替代方案,然而其优势在很大程度上取决于网络架构和数据特性。我们构建了一个分析框架,在匹配表达能力的条件下,比较全连接ReLU ANN与积分发放(integrate-and-fire)SNN在处理时间序列数据时的理论能效。通过将推理能耗模型与表征表达能力的理论界限相关联,我们推导出了表达力归一化的效率比率,以及关于网络宽度、脉冲稀疏性和ANN深度扩展的显式阈值。我们的分析刻画了事件驱动计算能够抵消SNN时间开销的区间,为设计节能的时间网络提供了容量感知的设计原则。分析表明,ANN仅在特定区间内才能在表达力归一化效率上超过SNN。
cs.LG / 79 / 2608.29886

Structural Hierarchy and Geometry in Molecular Representation Learning

分子表示学习中的结构层次与几何特性
Sulu, David, Di Fruscia, Lorenzo, Weber, Jana M.
Abstract
Molecular self-supervised learning uses chemical structures to guide which molecular embeddings should be similar. We study whether explicitly encoding a molecule's Bemis-Murcko scaffold and using it to supervise the molecular embedding changes what the model learns. We further test whether this effect depends on the embedding geometry by comparing Euclidean and Lorentz contrastive objectives. Across two augmentation strengths, scaffold-supervised models consistently organize molecules according to both identical and structurally related scaffolds. The resulting embeddings also improve molecular property prediction on several tasks, while the exact gains depend on the predicted property. The effect of scaffold supervision on molecular organization is stronger under Lorentz objectives, but neither geometry provides a consistent overall advantage. These results show that explicitly teaching the relation between a molecule and its structural core can reliably shape the organization of molecular embedding space, while the extent of usefulness of this organization remains task dependent.
Chinese Translation
分子自监督学习利用化学结构来指导哪些分子嵌入应当彼此相似。我们研究了显式编码分子的Bemis-Murcko骨架并将其用于监督分子嵌入,是否会改变模型所学到的内容。我们进一步通过比较欧氏(Euclidean)和洛伦兹(Lorentz)对比学习目标,检验这种效应是否依赖于嵌入几何结构。在两种数据增强强度下,骨架监督模型均能一致地按照相同骨架及结构相关骨架对分子进行组织。所得的嵌入还在多个分子性质预测任务上提升了性能,但具体增益取决于所预测的性质。骨架监督对分子组织的影响在洛伦兹目标下更强,但两种几何结构均未表现出一致的整体优势。这些结果表明,显式教授分子与其结构核心之间的关系能够可靠地塑造分子嵌入空间的组织结构,而这种组织结构的实用程度仍取决于具体任务。
cs.LG / 80 / 2608.29888

Sensitivity-Constrained Neural Operators for Data-Efficient Forward and Inverse Modeling of Partial Differential Equation Systems

面向偏微分方程系统数据高效正演与反演建模的敏感性约束神经算子
Behroozi, Abdolmehdi, Shen, Chaopeng, Kifer, Daniel, Lawson, Kathryn
Abstract
Neural operators provide fast surrogates for partial differential equation (PDE) solvers, but their reliability can degrade for high-dimensional spatial inputs and inverse or repeated inference. State-only training constrains solution values but not the learned input--output response. We study sensitivity-constrained neural operators (SC-NOs), which augment standard training with sampled solver-derived Jacobian supervision. Selected sensitivities from differentiable solvers or discrete adjoints are matched during training, allowing response information to be amortized across minibatches without imposing the full Jacobian at every update. We evaluate SC-NO on advection--diffusion and RANS--Spalart--Allmaras benchmarks, input-dimensionality scaling tests, long-horizon autoregressive rollout, and a shallow-water Tohoku tsunami source-inversion case. Sensitivity supervision improves forward prediction and yields larger gains in gradient-based inverse reconstruction of distributed fields. Scaling experiments show an improved accuracy--cost tradeoff for high-dimensional gridded inputs, while ablations indicate that state values and Jacobian information provide complementary supervision. In the tsunami case, SC-FNO reconstructs gridded seafloor deformation from sparse early gauge observations and forecasts subsequent wave propagation in a near-real-time proof-of-concept workflow. These results support sampled sensitivity supervision as a practical way to improve neural PDE surrogates when forward accuracy, inverse stability, robustness, and computational cost must be considered together.
Chinese Translation
神经算子为偏微分方程(PDE)求解器提供了快速代理模型,但在高维空间输入以及反演或反复推断场景下,其可靠性可能下降。仅基于状态(state-only)的训练只约束解的数值,而无法约束所学到的输入-输出响应。本文研究敏感性约束神经算子(Sensitivity-Constrained Neural Operators, SC-NOs),该方法在标准训练的基础上引入由求解器采样得到的雅可比矩阵监督信息。训练过程中,通过可微求解器或离散伴随方法选取的敏感性被用于匹配约束,从而将响应信息在多个小批次(minibatch)间摊销,而无需在每次更新时施加完整的雅可比矩阵。我们在对流-扩散基准、RANS-Spalart-Allmaras基准、输入维度扩展性测试、长时间自回归滚动预测以及一次浅水方程东日本(Tohoku)海啸震源反演案例上评估了SC-NO。敏感性监督提升了正演预测精度,并在基于梯度的分布场反演重建中带来了更大的收益。扩展性实验表明,对于高维网格化输入,精度-成本的权衡得到改善;消融实验则表明状态值与雅可比信息提供了互补的监督信号。在海啸案例中,SC-FNO(SC-NO与傅里叶神经算子结合的模型)从稀疏的早期验潮仪观测数据中重建了网格化的海底形变,并在近实时的概念验证工作流中预测了后续的波浪传播。这些结果表明,在需要综合考虑正演精度、反演稳定性、鲁棒性和计算成本时,采样式敏感性监督是改进神经PDE代理模型的一种实用途径。
cs.LG / 81 / 2608.29892

Joint Spatiotemporal Spectral Neural Operators for Learning PDEs on Irregular Domains

面向不规则域偏微分方程学习的时空联合谱神经算子
Behroozi, Abdolmehdi, Shen, Chaopeng
Abstract
Learning solution operators for partial differential equations (PDEs) on irregular and geometry-dependent domains remains a central challenge in scientific machine learning. While spectral methods provide strong inductive biases for modeling global interactions, they are typically limited to regular domains, and existing neural approaches often require domain warping, interpolation, or costly geometric embeddings. We introduce the \textbf{Graph Spectral Neural Operator (GSNO)}, a neural operator that combines spatial graph spectral decompositions with temporal Fourier transforms through a unified space--time spectral kernel. This formulation enables globally coherent operator learning on non-Cartesian discretizations without domain warping or autoregressive rollouts. By replacing learned geometric embeddings with a graph Laplacian spectral basis, GSNO provides geometry-aware spectral learning with low parameter complexity. Across steady and unsteady PDE benchmarks on irregular and geometry-dependent domains, GSNO achieves strong accuracy with reduced runtime and parameter counts, while demonstrating robust zero-shot generalization across mesh resolutions and geometry families.
Chinese Translation
在不规则且依赖几何形状的域上学习偏微分方程(PDE)的解算子,始终是科学机器学习中的核心挑战。尽管谱方法在建模全局相互作用方面提供了强归纳偏置,但其通常局限于规则域,而现有的神经方法往往需要进行域变形、插值或代价高昂的几何嵌入。我们提出了**图谱神经算子(Graph Spectral Neural Operator, GSNO)**,这是一种通过统一的时空谱核将空间图谱分解与时间傅里叶变换相结合的神经算子。该公式使得在非笛卡尔离散化上进行全局一致的算子学习成为可能,且无需域变形或自回归滚动推演。通过用图拉普拉斯谱基替代习得的几何嵌入,GSNO 以较低的参数复杂度实现了几何感知的谱学习。在 irregular 和几何依赖域上的稳态与非稳态 PDE 基准测试中,GSNO 在显著降低运行时间和参数量的同时取得了较高的精度,并在不同网格分辨率和几何族之间展现出鲁棒的零样本泛化能力。
cs.LG / 82 / 2608.29901

INTERVenE: Temporal-Abstraction-Interval Based Transformers for Short-Horizon Medical Event Prediction

INTERVenE:基于时间抽象区间的Transformer用于短程医疗事件预测
Oded, Shahar, Shahar, Yuval
Abstract
Electronic Health Record (EHR) prediction models in the intensive care unit must learn from sparse and irregular measurements while preserving the clinical meaning of time and supporting transparent decision-making. We present INTERVenE, a family of Transformer architectures whose input is an interval-based, knowledge-based temporal abstraction (KBTA), a token stream of named clinical concepts (states, trends, events, contexts) drawn from a curated medical ontology, rather than an unnamed bin index or a raw measurement triplet. This naming layer is what we ask KBTA to do: it makes the model's per-token attributions resolve to clinical concepts by construction. INTERVenE offers two complementary variants: an auto-regressive decoder that generates future abstraction trajectories with a per-step risk readout (localizing \emph{when} and \emph{after which events} risk rises), and a bidirectional encoder for single-pass joint risk and time-to-event prediction. Evaluated on 57,078 MIMIC-IV admissions against GRU-D, STraTS, and KarmaLego, INTERVenE-Enc reaches a support-weighted AUPRC$_w$ of 0.672, improving by 0.041 over the strongest neural baseline with non-overlapping 95\% bootstrap CIs, while also taking the best AUROC$_w$ (0.901) and length-of-stay MAE (44.4\,h). INTERVenE-Ar (AUROC$_w$ $0.854$, AUPRC$_w$ $0.587$ under the same evaluation contract - a strictly harder generative readout) provides a complementary token-level risk trajectory. An input-representation ablation confirms the lift transfers across structured discretizations, positioning KBTA-based intervals as the interpretable substrate that makes per-token attributions resolve to meaningful clinical concepts within the deployed model.
Chinese Translation
重症监护病房中的电子健康记录(EHR)预测模型必须从稀疏且不规则的测量数据中学习,同时保留时间的临床意义并支持透明的决策。我们提出了INTERVenE,这是一系列Transformer架构,其输入是基于区间的、基于知识的时序抽象(KBTA),即一个由命名的临床概念(状态、趋势、事件、上下文)构成的词元流,这些概念来自经过整理的医学本体,而非未命名的分箱索引或原始测量三元组。这一命名层正是我们要求KBTA实现的功能:它使模型的逐词元归因在结构上能够对应到具体的临床概念。INTERVenE提供两种互补的变体:一种是自回归解码器,它生成未来的抽象轨迹并给出每一步的风险读出(定位风险在何时以及在哪些事件之后上升);另一种是双向编码器,用于单次遍历的联合风险与事件发生时间预测。在57,078例MIMIC-IV入院记录上与GRU-D、STraTS和KarmaLego进行对比评估,INTERVenE-Enc达到支持度加权AUPRC$_w$ 0.672,比最强的神经基线提高了0.041,且95%自助法置信区间不重叠,同时还取得了最佳AUROC$_w$(0.901)和住院时长MAE(44.4小时)。INTERVenE-Ar(在相同评估协议下AUROC$_w$为0.854、AUPRC$_w$为0.587——这是一种严格更难的生成式读出)提供了互补的词元级风险轨迹。输入表示消融实验证实了这一提升可跨结构化离散化方法迁移,将基于KBTA的区间定位为可解释的基础层,使逐词元归因在部署的模型中能够对应到有意义的临床概念。
cs.LG / 83 / 2608.29907

Diffusion-Based Inverse Design of Dielectric Resonator Metasurfaces for Shaping Smart Electromagnetic Environments

基于扩散模型的介质谐振器超表面逆向设计用于塑造智能电磁环境
Tsukerman, M., Grotov, K., Vovchuk, D., Ginzburg, P.
Abstract
Future wireless systems are expected to transform the surrounding space from a passive propagation medium into a smart electromagnetic environment, where engineered surfaces control wave propagation, support wireless sensing, and create programmable electromagnetic fingerprints. A key challenge in realizing this vision is the inverse design of metasurfaces for tailored electromagnetic propagation. While forward analysis evaluates the response of a known geometry, the inverse task starts from a prescribed scattering signature and seeks a physically realizable structure that produces it. This inverse task is inherently nonlinear and often high-dimensional, while candidate solutions may be non-unique and provide no direct indication of practical realizability. Here, we introduce a conditional diffusion framework for inverse design of dielectric resonator metasurfaces from target angular scattering patterns. Trained on T-matrix simulated geometry-response pairs, the model learns a conditional distribution of geometries instead of a deterministic mapping, enabling multiple candidate designs for the ill-posed inverse problem. The best generated metasurface achieves a mean percentage error of 1.39%, outperforming CMA-ES optimization (4.1% after 10 h) while requiring only about one minute for after-training inference. The model also produces lower error distributions than deterministic neural baselines for out-of-distribution spectra, highlighting the potential of diffusion models for efficient metasurface design.
Chinese Translation
未来的无线系统有望将周围空间从被动的传播介质转变为智能电磁环境,其中经过工程设计的表面能够控制波传播、支持无线感知并创建可编程的电磁指纹。实现这一愿景的一个关键挑战是面向定制电磁传播的超表面逆向设计。正向分析评估已知几何结构的响应,而逆向任务则从给定的散射特征出发,寻求能够产生该特征的物理可实现结构。这一逆向任务本质上是高度非线性的,且往往具有高维特性,同时候选解可能不唯一,也无法直接指示其实际可实现性。本文提出了一种条件扩散框架,用于从目标角散射图样出发对介质谐振器超表面进行逆向设计。该模型在T矩阵仿真的几何结构-响应配对数据上训练,学习的是几何结构的条件分布而非确定性映射,从而能够为这一病态逆问题生成多个候选设计。所生成的最优超表面实现了1.39%的平均百分比误差,优于CMA-ES优化方法(10小时后为4.1%),且训练后的推理仅需约一分钟。对于分布外的频谱,该模型还产生了比确定性神经基线更低的误差分布,凸显了扩散模型在高效超表面设计中的潜力。
cs.LG / 84 / 2608.29943

On the Recoverability of Private Information Unlearning in Large Language Models

论大语言模型中私有信息遗忘的可恢复性
Hu, Shicheng, Tian, Runzhi, Wang, Ziqiao, Mao, Yongyi
Abstract
Large language models (LLMs) can memorize sensitive information, raising serious privacy concerns. Machine unlearning offers a potential solution to remove such information, but it remains unclear whether existing methods truly erase it or merely hide it within the model. A key challenge is quantifying the persistence of sensitive data under a unified evaluation framework. To address this, we construct a synthetic dataset containing fake private information and propose a white-box auditing framework to systematically assess whether claimed-forgotten information is genuinely removed. Using this framework, we evaluate five existing unlearning methods and find that a simple "inverse greedy" decoding -- selecting the least likely token at each step -- can recover supposedly forgotten private information. Our results reveal that current unlearning approaches often fail to fully eliminate sensitive information, highlighting the need for more reliable methods to ensure privacy in deployed LLMs.
Chinese Translation
大语言模型(LLMs)可能记忆敏感信息,引发了严重的隐私担忧。机器遗忘(Machine Unlearning)为删除此类信息提供了一种潜在解决方案,但现有方法究竟是真正抹除了这些信息,还是仅仅将其隐藏在模型内部,目前仍不清楚。一个关键挑战是在统一的评估框架下量化敏感数据的持久性。为解决这一问题,我们构建了一个包含虚假私有信息的合成数据集,并提出了一种白盒审计框架,以系统性评估所谓被遗忘的信息是否被真正删除。利用该框架,我们评估了五种现有的遗忘方法,发现一种简单的“逆贪婪”(inverse greedy)解码方法——即每一步选择概率最小的词元——能够恢复据称已被遗忘的私有信息。我们的研究结果表明,当前的遗忘方法往往无法完全消除敏感信息,凸显了开发更可靠方法以保障已部署大语言模型隐私安全的必要性。
cs.LG / 85 / 2608.29983

Robust Broad Learning System with Wave Loss for Classification under Data Uncertainty

面向数据不确定性分类的具有波动损失的鲁棒宽度学习系统
Akhtar, Mushir, Varshney, A., Quadir, A., Rahaman, A., Tanveer, M., Arshad, Mohd.
Abstract
Broad Learning System (BLS) offers an efficient alternative to deep architectures by enabling fast learning through randomized feature mapping and closed-form solutions. However, its reliance on squared error loss makes it highly sensitive to noise, outliers, and corrupted labels, limiting its reliability in real-world scenarios. To address this limitation, we propose Wave-BLS, a robust broad learning framework that integrates the wave loss function, which is asymmetric, bounded, and smooth, enabling controlled penalization of large errors. The proposed formulation replaces the standard least-squares objective with a wave-loss-based optimization problem, solved efficiently using a Nesterov accelerated gradient (NAG)-based scheme without requiring matrix inversion, thereby improving scalability. Extensive experiments on 30 UCI benchmark datasets demonstrate that Wave-BLS consistently outperforms classical BLS and several robust variants. Statistical validation using Friedman and Nemenyi post-hoc tests confirms the significance of the observed improvements. Furthermore, robustness evaluations under controlled noise and outlier injection reveal that Wave-BLS exhibits substantially slower performance degradation compared to BLS, even in challenging contamination settings. These results establish Wave-BLS as a stable and robust alternative to existing broad learning models for learning under data uncertainty.
Chinese Translation
宽度学习系统(Broad Learning System, BLS)通过随机特征映射和闭式解实现快速学习,为深度架构提供了一种高效的替代方案。然而,其依赖平方误差损失,使其对噪声、离群点和标签损坏高度敏感,限制了其在真实场景中的可靠性。为解决这一局限,我们提出了Wave-BLS,一种融合波动损失函数的鲁棒宽度学习框架。该损失函数具有非对称、有界且光滑的特性,能够对较大误差进行可控惩罚。所提出的方法用基于波动损失的优化问题替代了标准最小二乘目标,并采用基于Nesterov加速梯度(NAG)的方案高效求解,无需矩阵求逆,从而提升了可扩展性。在30个UCI基准数据集上的大量实验表明,Wave-BLS始终优于经典BLS及多种鲁棒变体。基于Friedman检验和Nemenyi事后检验的统计验证证实了所观察到的改进的显著性。此外,在受控噪声和离群点注入下的鲁棒性评估显示,即使在具有挑战性的污染环境下,Wave-BLS的性能退化速度也明显慢于BLS。这些结果表明,Wave-BLS是现有宽度学习模型在数据不确定性学习方面的一种稳定且鲁棒的替代方案。
cs.LG / 86 / 2608.29998

The Intervention Gap in Latent World Models

潜在世界模型中的干预鸿沟
Vakalis, Donna
Abstract
Planning-time intervention fidelity is a distinct, measurable property of a learned world model: whether the model's own open-loop transitions move task variables the way matched environment interventions do. In the settings we test, it is neither revealed by reward fit nor ensured by task-anchored training. Across released TD-MPC2 checkpoint sizes, episode return falls as an operator-error diagnostic on task observables grows, while reward-prediction error stays small and nearly flat, and a self-supervised world model trained without task signal preserves the same operator substantially better than a task-anchored model on the shared task. A capture-gated matched-intervention audit then localizes what fails. On Cheetah, three LeWorldModel checkpoints capture the current task query and support decodable real intervention effects; however, their imagined five-step effects are worse than predicting no effect and worse than an environment-endpoint oracle. The failure is task-direction rotation with excess gain, not feature collapse. This severe pattern is conditional: five PreJEPA seeds retain an oracle-relative deficit without it, Finger Spin experiments extend the deficit beyond locomotion with heterogeneous severity across seeds, and shared-bank effect geometry is both candidate- and support-dependent. We also test practice-side questions. In DreamerV3 the posterior distribution, not its sample, carries the current query; ensemble disagreement ranks error only near training support; and a frozen support-aware score degrades held-out error ranking in both tested transfer directions while native disagreement remains informative in both. We conclude that intervention fidelity must be audited directly, capture-first, on the model's native interface.
Chinese Translation
规划时的干预保真度是学习到的世界模型的一个独特且可度量的性质:即模型自身的开环转移是否能以与环境匹配的干预相同的方式移动任务变量。在我们测试的设置中,这一性质既不能通过奖励拟合来揭示,也不能由任务锚定的训练来保证。在所发布的各种 TD-MPC2 检查点规模上,随着任务可观测变量上的操作者误差诊断指标增大,回合回报下降,而奖励预测误差保持很小且近乎平坦;并且在共享任务上,无任务信号训练的自监督世界模型比任务锚定模型更能保留相同的操作者。随后,一种以捕获为门槛的匹配干预审计定位了失效之处。在 Cheetah 上,三个 LeWorldModel 检查点能够捕获当前任务查询,并支持可解码的真实干预效应;然而,它们想象的五步效应比预测无效应更差,也比环境终点 oracle 更差。其失效模式是带有过量增益的任务方向旋转,而非特征坍缩。这一严重模式是有条件的:五个 PreJEPA 种子在没有该模式的情况下仍保留相对于 oracle 的缺陷;Finger Spin 实验将这种缺陷扩展到运动控制之外,且不同种子之间的严重程度存在差异;共享库的效应几何结构既依赖于候选对象也依赖于支撑集。我们还测试了实践层面的问题。在 DreamerV3 中,是后验分布而非其样本承载当前查询;集成分歧仅在训练支撑集附近才能对误差进行排序;而一个冻结的支撑感知评分在两个测试的迁移方向上都会降低对留出误差排序的质量,而原生的分歧评分在两个方向上都保持有效信息。我们的结论是:干预保真度必须在模型的本地接口上、以捕获为先、直接进行审计。
cs.LG / 87 / 2608.30021

Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models

PET/CT放射学报告的错误检测:领域专用模型与大语言模型的比较
Warr, Hermione, Anthony, Harry, Freischem, Lilli J, Ibrahim, Yasin, McGowan, Daniel R, Kamnitsas, Konstantinos
Abstract
Errors in radiology reports can adversely affect patient treatment, yet automated report quality assurance remains challenging because errors are often subtle and require domain expertise to detect. Although large language models (LLMs) have recently been proposed for radiology report verification, their ability to detect clinically meaningful errors beyond chest X-ray datasets remains under-explored. To this end, we present the first systematic evaluation of language models for PET/CT report error detection, comparing compact domain-specific models with SOTA open-weight LLMs. We collected 30,633 oncology FDG PET/CT reports from 23 radiologists over 10 years. We trained domain-specific BERT models to detect clinically motivated synthetic reporting errors and evaluated alongside zero-/few-shot Qwen3-32B, Gemma-3-27B and Llama-3.3-70B on a held-out benchmark of 11,500 reports. A 15M-parameter model achieved 94.4% balanced accuracy with a 5.8% false-positive rate, compared with 84.0% for the strongest prompted LLM. Task-specific adaptation of Llama-3.3-70B closed this performance gap (94.4%) but retained substantially greater computational requirements. Our results suggest that domain-specific training matters more than model scale for PET/CT report error detection, supporting compact models as an accurate and computationally efficient approach to automated radiology report quality assurance.
Chinese Translation
放射学报告中的错误可能对患者治疗产生不利影响,然而由于错误往往细微且需要领域专业知识才能识别,自动化报告质量保证仍然具有挑战性。尽管大语言模型(LLM)近期已被提出用于放射学报告校验,但其在胸部X光数据集之外检测具有临床意义错误的能力仍未得到充分探索。为此,我们首次系统评估了语言模型在PET/CT报告错误检测中的表现,将紧凑的领域专用模型与最先进(SOTA)的开源权重LLM进行比较。我们收集了23名放射科医生在10年间撰写的30,633份肿瘤学FDG PET/CT报告。我们训练了领域专用BERT模型来检测基于临床动机的合成报告错误,并在包含11,500份报告的留出基准上与零样本/少样本的Qwen3-32B、Gemma-3-27B和Llama-3.3-70B进行了比较。一个1500万参数的模型达到了94.4%的平衡准确率和5.8%的假阳性率,而表现最好的提示式LLM仅为84.0%。对Llama-3.3-70B进行任务特定微调可以弥合这一性能差距(94.4%),但其计算需求仍显著更高。我们的结果表明,在PET/CT报告错误检测中,领域专用训练比模型规模更为重要,支持将紧凑模型作为自动化放射学报告质量保证的一种准确且计算高效的方法。
cs.LG / 88 / 2608.30028

Multiclass Linear Perceptrons with Multiplicative Margins

具有乘性间隔的多类线性感知机
Rachkovskij, Dmitri, Osipov, Evgeny, Volkov, Olexander, De Silva, Daswin, Kleyko, Denis
Abstract
This paper introduces a family of multiclass linear Perceptron classifiers with a multiplicative margin mechanism (MMPerc), as an alternative to standard margin-free and additive margin Perceptrons. The multiplicative formulation enforces classification confidence by requiring the true class score to exceed that of competing classes by a specified fraction of itself, rather than by a fixed additive threshold. This avoids dependence on score magnitudes arising from varied norms of data and class weight vectors. We propose several architectural and algorithmic variants of MMPerc, derive associated loss functions and mistake bounds for both linearly separable and non-separable data, and analyze key design considerations, including bias, margin threshold selection, and training modes. Extensive experiments on synthetic and real datasets show that MMPerc classifiers typically outperform the standard Perceptron, as well as classic baselines such as Support Vector Machines and Ridge classifiers. Owing to their simplicity, minimalistic design, and computational efficiency, MMPerc classifiers are promising candidates for conventional machine learning tasks, linear evaluation of Deep Neural Networks, integration with Hyperdimensional Computing / Vector Symbolic Architecture representations, and deployment in resource-constrained applications.
Chinese Translation
本文提出了一类具有乘性间隔机制的多类线性感知机分类器(MMPerc),作为标准无间隔感知机和加性间隔感知机的替代方案。乘性公式通过要求真实类别得分超过竞争类别得分达到自身的一定比例(而非固定的加性阈值)来强制实现分类置信度,从而避免了对由数据范数和类别权重向量范数差异所导致的得分大小的依赖。我们提出了MMPerc的若干结构和算法变体,针对线性可分与不可分数据推导了相应的损失函数和错误界,并分析了关键设计考量,包括偏置、间隔阈值选择和训练模式。在合成数据集和真实数据集上的大量实验表明,MMPerc分类器通常优于标准感知机以及支持向量机(Support Vector Machines)和岭分类器(Ridge classifiers)等经典基线方法。凭借其简单性、极简设计和计算高效性,MMPerc分类器在传统机器学习任务、深度神经网络的线性评估、与超维计算/向量符号架构(Hyperdimensional Computing / Vector Symbolic Architecture)表示的结合,以及资源受限应用中的部署等方面均是有前景的候选方案。
cs.LG / 89 / 2608.30046

Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label Correction

遗忘还是微调?面向噪声标签纠正的机器遗忘策略比较研究
Santana, João L. P., Cordeiro, Filipe R.
Abstract
Noisy labels remain a critical challenge for training deep neural networks, since memorizing incorrect labels degrades generalization. Once noisy samples are identified after training, the standard solution is to retrain the model from scratch on the cleaned dataset, which is increasingly expensive as datasets and models grow. Machine Unlearning (MU) has recently emerged as a computationally efficient alternative, but the relative effectiveness of different MU strategies for noisy-label correction remains poorly understood. In this work, we conduct a comparative empirical study of five MU methods (NegGrad, Fine-Tuning (FT), Random Labeling (RL), SalUn, and MUNBa) across symmetric, asymmetric, instance-dependent, and open-set noise on CIFAR-10, CIFAR-100, and the real-world noisy dataset Food-101N. Our central finding is that the appropriate unlearning strategy is conditioned on the noise structure. Simple FT is a strong baseline across most closed-set scenarios; RL and SalUn are the most consistently robust methods and, under instance-dependent noise, approach retraining accuracy at a fraction of the computational cost; MUNBa shows advantages mainly under extreme symmetric noise. Under open-set noise, in contrast, we show that retraining on the cleaned subset degrades accuracy relative to the noisy baseline, so approximating the retrained model is not an adequate objective in this regime. On Food-101N, all MU methods remain competitive and achieve accuracies close to retraining despite reducing runtime by an order of magnitude. These findings provide practical guidelines for selecting MU strategies for post-training noisy-label correction.
Chinese Translation
噪声标签仍然是训练深度神经网络的一个关键挑战,因为对错误标签的记忆会损害模型的泛化能力。在训练后识别出噪声样本后,标准解决方案是在清洗后的数据集上从头重新训练模型,但随着数据集和模型规模的不断增长,其代价日益高昂。机器遗忘(Machine Unlearning, MU)最近作为一种计算高效的替代方案而兴起,但不同MU策略在噪声标签纠正中的相对有效性仍缺乏深入理解。在本工作中,我们在CIFAR-10、CIFAR-100以及真实世界噪声数据集Food-101N上,针对对称噪声、非对称噪声、实例依赖噪声和开集噪声,对五种MU方法(NegGrad、微调(Fine-Tuning, FT)、随机标签(Random Labeling, RL)、SalUn和MUNBa)进行了比较性的实证研究。我们的核心发现是:合适的遗忘策略取决于噪声结构。简单的FT在大多数闭集场景中是强基线;RL和SalUn是鲁棒性最稳定的方法,且在实例依赖噪声下,能以极小的计算代价接近重新训练的准确率;MUNBa主要在极端对称噪声下表现出优势。相反,在开集噪声下,我们发现基于清洗子集的重新训练相对于噪声基线会降低准确率,因此在这种情形下逼近重新训练后的模型并不是一个充分的目标。在Food-101N上,所有MU方法均保持竞争力,尽管将运行时间减少了一个数量级,但仍达到了接近重新训练的准确率。这些发现为训练后噪声标签纠正选择MU策略提供了实用指导。
cs.LG / 90 / 2608.30054

When 3D Gaussian Splatting Recovers Real Surfaces

当3D高斯泼溅能够恢复真实表面时
Wang, Songhe, Miller, David Johnathan
Abstract
When does 3D Gaussian Splatting (3DGS) recover the true scene surface rather than just overfitting view-dependent appearance? We answer this by developing a mathematical framework based on a first-hit rendering abstraction that cleanly isolates geometry from appearance. We prove that geometric misalignment forcefully converts spatial textures into high-frequency angular signals via parallax. This establishes a strict identifiability window: if angular capacity is bounded, surface-consistent solutions are mathematically preferred; if unrestricted, the same images can be perfectly explained by an incorrect, opaque billboard geometry. Experiments on synthetic stress tests confirm this prediction, showing billboard failures emerge precisely at high angular capacities. Conversely, in the real-world datasets we evaluate under standard capture protocols, reconstructions remain surface-consistent even at high SH degrees, which is consistent with the prediction that rich spatial texture can push billboard solutions outside the tested angular-capacity range.
Chinese Translation
3D高斯泼溅(3DGS)何时能够恢复真实的场景表面,而不是仅仅过拟合视角相关的外观?为此,我们构建了一个基于“首次命中”渲染抽象的数学框架,将几何与外观干净地分离。我们证明,几何错位会通过视差效应强有力地将空间纹理转换为高频角度信号。这确立了一个严格的可辨识性窗口:若角度容量受限,与表面一致的解在数学上是被优先选择的;若角度容量不受限制,则相同的图像可以由不正确的、不透明的公告牌(billboard)几何完美解释。在合成压力测试上的实验证实了这一预测,表明公告牌失败恰好在高角度容量下出现。相反,在我们评估的真实世界数据集中,即使在标准采集协议和高阶球谐函数(SH)下,重建结果仍保持与表面一致,这与如下预测相符:丰富的空间纹理可以将公告牌解推到所测试的角度容量范围之外。
cs.LG / 91 / 2608.30067

How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral Account

LLM智能体中世界模型与策略如何组合?一种频谱与行为的联合分析
Xu, Ruize, Yu, Xiao, Tang, Yujin, Shang, Chenming, Singh, Nikhil
Abstract
How do LLM agents come to both understand environments they act in and master tasks set within them? Through controlled experiments combining world-model training (next-state prediction) and policy training (reward maximization), we investigate this question. We dissect the resulting models through their additive parameter updates. Geometrically, we find effective world-model updates are low-rank and share an input-feature subspace with policy updates while writing to nearly orthogonal output directions, whether trained separately or sequentially. However, we find that, in projection interventions, the sequential update induces more robustness than separate policy RL when removing the world model's leading input directions, suggesting that it has learned alternative input pathways. Behaviorally, we find the sequentially trained agent explores a wider range of states and actions. Based on this, we ask: does policy training preserve world knowledge as well as it could? We probe this with training-free merging built on the geometrically motivated input basis plus an online world-model loss during policy RL, and show both improve over the untreated baseline. Our findings suggest world knowledge and task-directed ability can be learned in geometrically complementary forms, and that future post-training pipelines should consider how best to engineer the interface between them.
Chinese Translation
LLM智能体是如何既理解其所在的环境,又掌握环境中设定的任务的?我们通过结合世界模型训练(下一状态预测)与策略训练(奖励最大化)的受控实验来研究这一问题。我们通过对所得模型的加性参数更新进行剖析来深入理解这些模型。从几何角度看,我们发现有效的世界模型更新是低秩的,并且与策略更新共享一个输入特征子空间,同时写入近似正交的输出方向——无论是分别训练还是顺序训练皆是如此。然而,我们发现在投影干预实验中,当移除世界模型的主输入方向时,顺序更新比单独的策略强化学习(RL)表现出更强的鲁棒性,这表明它已经学习到了替代性的输入通路。从行为角度看,我们发现顺序训练的智能体探索了更广泛的状态和动作范围。基于此,我们提出问题:策略训练能否尽可能好地保留世界知识?我们通过基于几何动机的输入基底的无训练合并方法,以及在策略强化学习过程中的在线世界模型损失来探究这一问题,结果表明二者均优于未经处理的基线。我们的发现表明,世界知识与面向任务的能力可以以几何上互补的形式被学习,未来的后训练流程应考虑如何最优地设计二者之间的接口。
cs.LG / 92 / 2608.30070

Selection, Representation, and Execution in Sparse Fourier Neural Operators

稀疏傅里叶神经算子中的选择、表示与执行
Ibrahim, Abdul Qadir, Burger, Martin
Abstract
Sparse representations are often expected to make models smaller and also reduce inference cost. For Fourier Neural Operators (FNOs), these objectives are not equivalent or do not always align: removing parts of the learned operator can leave the underlying transforms and dense computations unchanged, while changing the grid on which the model is evaluated can introduce overhead of its own. We therefore distinguish sparsity in the representation, in the stored parameters, in the theoretical operation count, and in measured runtime, and present an empirical study of several routes toward sparse FNOs that tests each transition between them separately. Coarsening the execution grid reduces the theoretical cost without reducing measured latency, and adding a correction term recovers accuracy at the cost of making the model slower. Even an 83\% parameter reduction remains slower than the dense baseline under ordinary execution. These results motivate a stricter definition of useful sparsity: the deployed operator must preserve solution accuracy and map its reduced support to a genuinely cheaper execution path.
Chinese Translation
稀疏表示通常被期望使模型更小并降低推理成本。对于傅里叶神经算子(Fourier Neural Operators, FNOs),这两个目标并不等价,也不总是保持一致:删除所学算子的某些部分可能使底层变换和稠密计算保持不变,而改变模型求值的网格本身也可能引入额外开销。因此,我们区分了表示层面的稀疏性、存储参数的稀疏性、理论运算量的稀疏性以及实测运行时间的稀疏性,并对通往稀疏FNO的若干路径进行了实证研究,分别检验了它们之间的每一次转换。粗化执行网格可以降低理论成本,但无法降低实测延迟;而添加校正项虽可恢复精度,却以降低模型速度为代价。即使在常规执行方式下参数减少83%,其速度仍慢于稠密基线。这些结果促使我们提出更严格的有用稀疏性定义:部署的算子必须保持求解精度,并将其缩减后的支持集映射到真正更低成本的执行路径上。
cs.LG / 93 / 2608.30081

Tracing Generated Samples to Training-Data Clusters in Flow-Matching Models

在流匹配模型中追踪生成样本至训练数据簇
Briq, Rania, Fried, Ohad, Kamp, Michael, Kesselheim, Stefan
Abstract
Understanding which training samples influence a generated image is an important problem in generative modeling. In flow matching, training samples influence the generated image through the velocity field along the generation trajectory. Removing samples to examine their counterfactual influence changes the velocity field, and the resulting effect on the final image depends on how the change propagates through the trajectory. Consequently, local changes in the velocity field do not necessarily predict the final counterfactual effect. This work investigates attribution in flow-matching models through a hybrid analytical--learned approach, and uses it to derive trajectory-based attribution scores at the cluster level. We evaluate these attribution scores using independently retrained leave-one-cluster-out (LOCO) models, and compare with several attribution baselines using two different flow-matching latent spaces. Our experiments show that semantic similarity constitutes a strong baseline, while the closed-form trajectory-based attribution is competitive in some metrics without requiring counterfactual retraining or model gradients. Our results show that attribution in flow matching depends not only on semantic similarity to training samples, but also on the latent representation, trajectory dynamics, and how influence is propagated to the final output.
Chinese Translation
理解哪些训练样本影响了生成的图像是生成式建模中的一个重要问题。在流匹配(flow matching)中,训练样本通过生成轨迹上的速度场对生成图像产生影响。移除样本以考察其反事实影响会改变速度场,而对最终图像产生的效果取决于该变化如何在轨迹中传播。因此,速度场的局部变化并不一定能预测最终的反事实效果。本研究通过一种分析—学习相结合的混合方法研究流匹配模型中的归因问题,并利用该方法推导出簇级别(cluster-level)的基于轨迹的归因分数。我们使用独立重训练的留一簇(leave-one-cluster-out, LOCO)模型来评估这些归因分数,并在两个不同的流匹配潜空间中与若干归因基线方法进行比较。实验表明,语义相似度构成了一个强基线,而闭式解的基于轨迹的归因方法在某些指标上具有竞争力,且无需反事实重训练或模型梯度。我们的结果表明,流匹配中的归因不仅取决于与训练样本的语义相似度,还取决于潜在表示、轨迹动力学以及影响如何传播到最终输出。
cs.LG / 94 / 2608.30088

A Lightweight Phenology-Aware YOLOv5 Framework for Tomato Growth Stage Detection in Resource-Constrained Bhutanese Greenhouse Environments

面向资源受限不丹温室环境的轻量级物候感知YOLOv5番茄生长阶段检测框架
Gocha, Sherab, Nobukawa, Sou
Abstract
Accurate detection of tomato growth stages is essential for stage-specific greenhouse management and precision agriculture. In Bhutan, greenhouse cultivation is affected by altitude variability, large diurnal temperature fluctuations, diffuse illumination, limited automation, and a scarcity of locally annotated datasets, limiting the applicability of conventional deep learning models. This work proposes Pheno-Lite + Efficient Channel Attention (ECA), a lightweight, phenology-aware object detection architecture derived from Ultralytics YOLOv5 for tomato growth stage recognition. A balanced dataset of 2,464 annotated images was constructed from locally collected greenhouse images in Bhutan and publicly available tomato images, with augmentation designed to simulate local greenhouse conditions. The dataset includes vegetative (820), flowering (824), fruiting (820), and background (26) samples. The proposed architecture introduces two customized backbone modules: C3 PhenoLite, which enhances spatial and texture feature extraction using depthwise residual refinement, and C3 ECA, which strengthens inter-channel feature interactions through efficient channel attention. The proposed model achieves 90.6% precision, 88.8% recall, and 92.6% mAP@50, with 4.0 million parameters and 10.9 GFLOPs at 640 x 640 resolution. These results demonstrate its potential for real-time and climate-resilient greenhouse deployment in Bhutan.
Chinese Translation
番茄生长阶段的准确检测对于温室阶段特异性管理和精准农业至关重要。在不丹,温室栽培受海拔变化大、昼夜温差大、漫射光照、自动化程度有限以及本地标注数据集匮乏等因素影响,限制了传统深度学习模型的适用性。本研究提出了一种轻量级、物候感知的目标检测架构Pheno-Lite + 高效通道注意力(ECA),该架构基于Ultralytics YOLOv5实现番茄生长阶段识别。研究利用在不丹本地采集的温室图像和公开可用的番茄图像,构建了一个包含2,464张标注图像的平衡数据集,并通过数据增强模拟本地温室条件。数据集包括营养生长期(820张)、开花期(824张)、结果期(820张)和背景(26张)样本。所提出的架构引入了两个定制的骨干网络模块:C3 PhenoLite,利用深度残差精炼增强空间和纹理特征提取;以及C3 ECA,通过高效通道注意力强化通道间的特征交互。所提模型达到90.6%的精确率、88.8%的召回率和92.6%的mAP@50,在640×640分辨率下参数量为400万,计算量为10.9 GFLOPs。这些结果证明了该模型在不丹实现实时、气候适应型温室部署的潜力。
cs.LG / 95 / 2608.30102

SMOTE-VAR: An Uncertainty-Aware Oversampling Method for Predicting Depression Remission in University Students

SMOTE-VAR:一种面向大学生抑郁症缓解预测的不确定性感知过采样方法
Nguyen, Dang, A V, Arun Kumar, Braund, Taylor A., Zheng, Wu Yi, Bal, Debopriyo, Hoon, Leonard, Newby, Jill, Christensen, Helen, Venkatesh, Svetha, Whitton, Alexis, Gupta, Sunil
Abstract
University students experience disproportionately high rates of common mental health conditions, such as depression, which can impair learning, social functioning, and overall well-being. Although lifestyle interventions such as mindfulness and physical activity can reduce the symptoms, many do not achieve symptomatic remission. Developing new approaches to identify students with poor outcomes could enable earlier and more targeted intervention. Machine learning (ML) methods have increasingly been used to predict remission in depressive patients. However, these ML models often suffer from class imbalance, where there may be an unequal proportion of people in the remitted group relative to the non-remitted group. This imbalance can reduce model accuracy and bias predictions. To address this, studies commonly employ the popular oversampling strategy SMOTE. However, SMOTE has a notable limitation: it may generate invalid synthetic minority samples. In a clinical context, these false positives can lead to incorrect risk stratification, potentially delaying necessary escalated care for patients unlikely to remit. In this paper, we introduce a novel and effective oversampling method that addresses this shortcoming. Our approach leverages the variance function of a Gaussian process to estimate the uncertainty of generated minority samples to reduce false positives. We validate our method on a depression dataset collected from university students and demonstrate that it is better than existing oversampling approaches in predicting remission (i.e., treatment outcome). By improving the reliable identification of non-responders, our method provides a robust computational tool to help clinicians rapidly pivot to adjunctive therapies, thereby personalizing and optimizing mental health care pathways.
Chinese Translation
大学生群体中常见心理健康问题(如抑郁症)的发病率显著偏高,这些疾病可能损害其学习能力、社会功能和整体幸福感。尽管正念和体育锻炼等生活方式干预可以减轻症状,但许多患者并未实现症状缓解。开发新的方法来识别预后不良的学生,有助于实现更早、更有针对性的干预。机器学习(ML)方法已被越来越多地用于预测抑郁症患者的缓解情况。然而,这些ML模型常常受到类别不平衡问题的困扰,即缓解组与非缓解组的人群比例可能不相等。这种不平衡会降低模型准确性并使预测产生偏差。为解决这一问题,研究通常采用流行的过采样策略SMOTE。然而,SMOTE存在一个显著局限:它可能生成无效的合成少数类样本。在临床背景下,这些假阳性可能导致错误的风险分层,进而可能延误对不太可能缓解的患者进行必要的升级治疗。本文提出了一种新颖且有效的过采样方法来弥补这一不足。我们的方法利用高斯过程的方差函数来估计所生成的少数类样本的不确定性,从而减少假阳性。我们在一个从大学生中收集的抑郁症数据集上验证了该方法,结果表明其在预测缓解(即治疗结果)方面优于现有的过采样方法。通过提高对无应答者的可靠识别,我们的方法为临床医生提供了一个稳健的计算工具,帮助其快速转向辅助治疗,从而实现心理健康诊疗路径的个性化和优化。
cs.LG / 96 / 2608.30103

Graph4BiLO: Graph Neural Network Approximation for Bilevel Mixed-Integer Linear Optimization

Graph4BiLO:面向双层混合整数线性优化的图神经网络近似方法
Elrefaei, Jessica D., Hua, Kaixun, Kim, Seungbae, Tran, Hoang Nam, Borrero, Juan S.
Abstract
Bilevel mixed-integer linear optimization problems model hierarchical decision processes in which a leader anticipates the optimal response of a follower. Although expressive, these problems are computationally challenging because lower-level optimality is embedded in the leader's feasible region. Value-function reformulations replace the nested follower optimization with a constraint involving the follower's optimal value, but evaluating this value function exactly can itself be expensive. This paper introduces Graph4BiLO, a graph neural network (GNN) approach for learning bilevel value functions from variable--constraint graph representations. In contrast to fixed-length multilayer perceptron (MLP) representations, the GNN uses shared message-passing parameters and can therefore be applied across multiple problem sizes with a single trained model. The learned ReLU network is encoded exactly as mixed-integer linear constraints and embedded in an approximate single-level formulation. A repair step subsequently re-solves the follower problem for the selected leader decision to recover a bilevel-feasible follower response. We evaluate Graph4BiLO on knapsack interdiction instances with 20--100 items against the exact MibS solver and the learning-based Neur2BiLO method. Graph4BiLO obtains objective values comparable to Neur2BiLO across all tested sizes while avoiding size-specific neural networks. An additional out-of-distribution experiment demonstrates zero-shot transfer from 20-item training instances to previously unseen 40- and 60-item instances. However, embedding message passing at every graph node substantially increases the resulting mixed-integer formulation size and solve time. These results identify a central tradeoff between size-generalizable graph representations and the computational cost of embedding GNNs within optimization models.
Chinese Translation
双层混合整数线性优化问题对分层决策过程进行建模,其中领导者会预期跟随者的最优反应。尽管这类问题具有很强的表达能力,但由于下层最优性被嵌入到领导者的可行域中,其计算求解极具挑战性。值函数重构方法将嵌套的跟随者优化问题替换为一个涉及跟随者最优值的约束,但精确计算该值函数本身可能代价高昂。本文提出Graph4BiLO,这是一种基于图神经网络(GNN)的方法,用于从变量—约束图表示中学习双层值函数。与固定长度的多层感知机(MLP)表示不同,GNN采用共享的消息传递参数,因此可以用单个训练好的模型应用于多种不同规模的问题。学习到的ReLU网络被精确编码为混合整数线性约束,并嵌入到近似的单层模型中。随后,通过一个修复步骤,对所选的领导者决策重新求解跟随者问题,以恢复双层可行的跟随者反应。我们在包含20至100个物品的背包拦截问题上,将Graph4BiLO与精确求解器MibS以及基于学习的Neur2BiLO方法进行了对比评估。在所有测试规模下,Graph4BiLO获得了与Neur2BiLO相当的目标值,同时避免了针对特定规模训练的神经网络。额外的分布外实验展示了从20个物品的训练实例到此前未见过的40和60个物品实例的零样本迁移能力。然而,在每个图节点上嵌入消息传递会显著增大所得混合整数模型的规模和求解时间。这些结果揭示了可跨规模泛化的图表示与在优化模型中嵌入GNN的计算成本之间的核心权衡。
cs.LG / 97 / 2608.30113

Supraglacial Lake Fate Is Knowable Long Before the Season Ends

冰面湖泊的命运在融冰季节结束前很久即可预知
Hossain, Emam, Gani, Md Osman
Abstract
A supraglacial lake on the Greenland Ice Sheet ends its melt season in one of four ways: it drains rapidly through a hydrofracture, drains slowly across the surface, refreezes in place, or is buried by late-season snowfall. Which one occurs decides whether the meltwater reaches the ice bed. Satellite classifiers recover the outcome accurately but only after the season closes, and how much of a season each outcome actually requires has never been measured. We measure it directly: holding the representation and the classifier fixed, we truncate the input at $14$ cutoffs from 1 May to 31 December, retrain at each, and record the earliest cutoff at which each outcome's per-class $F_1$ reaches a fixed target. The outcomes resolve in a consistent order, two of them months early: rapid drainage by 15 July and slow drainage by 1 August, $92$ and $75$ days ahead of the earliest date a full-season pipeline can be computed at all, with buried and refreeze following at $44$ and $30$ days. Five further learners, from a majority-class floor and $54$ summary statistics to a trigger-based early classifier, leave the ordering intact: every learner that produces a per-class trajectory reproduces it despite end-of-season accuracies differing by up to $18$ percentage points, and it survives leave-one-basin-out evaluation, though not the substitution of machine labels for expert ones in an unseen season. Every feature we compute at day $t$ reads only days up to $t$, at a cost of at most $1.3$ percentage points. A monitoring system should therefore not have one release date: rapid drainage can be flagged on 15 July, three months before a full-season pipeline can be computed at all.
Chinese Translation
格陵兰冰盖上的冰面湖泊(supraglacial lake)以四种方式之一结束其融冰季节:通过水力压裂快速排水、沿表面缓慢排水、原地重新冻结,或被季节末期降雪掩埋。发生哪种方式决定了融水是否到达冰床。卫星分类器能够准确恢复结果,但只有在季节结束后才能做到,而每种结果实际上需要多少季节时长从未被测量过。我们直接对此进行了测量:在保持表示和分类器不变的情况下,我们将输入在从5月1日到12月31日的14个截止点上截断,在每个截止点重新训练,并记录每种结果的各类F1分数达到固定目标的最早截止点。各结果以一致的顺序得到确定,其中两个提前数月:快速排水在7月15日之前、缓慢排水在8月1日之前即可确定,分别比全季节流水线最早可计算的日期提前92天和75天,掩埋和重新冻结随后分别提前44天和30天确定。另外五个学习器,从多数类基线和54个汇总统计量到基于触发的早期分类器,均保持该顺序不变:尽管季末准确率相差最多达18个百分点,每个能够产生各类轨迹的学习器都重现了这一顺序,并且在留一流域(leave-one-basin-out)评估中保持稳健,但在未见过的季节中用机器标签替代专家标签时则不然。我们在第t天计算的每个特征仅读取截至第t天的数据,代价最多为1.3个百分点。因此,监测系统不应只有一个发布日期:快速排水可以在7月15日被标记出来,比全季节流水线最早可计算的日期提前三个月。
cs.LG / 98 / 2608.30124

TPR-Attention for Combinatorial Generalization

用于组合泛化的张量积表示注意力机制(TPR-Attention)
Civelekoğlu, Melisa, Prémont-Schwarz, Isabeau
Abstract
Systematic generalization remains a significant challenge in deep learning. In particular, combinatorial generalization - generalizing to new configurations of known factors of variation - is effortless for humans but difficult for standard neural architectures that rely on statistical correlations rather than explicit structural representations. We introduce a new architectural component that embeds structured inductive bias into deep learning: an attention mechanism operating over tensor-product representations (TPRs). Through controlled experiments on compositional tasks, we show that this TPR-attention mechanism outperforms existing architectural components in combinatorial generalization. These results highlight the value of integrating explicit compositional structure into neural attention and point toward a promising path for models capable of systematic generalization.
Chinese Translation
系统性泛化仍然是深度学习中的一项重大挑战。特别是组合泛化——即泛化到已知变化因素的新组合配置——对人类而言轻而易举,但对于依赖统计相关性而非显式结构表示的标准神经网络架构来说却十分困难。我们提出了一种新的架构组件,将结构化归纳偏置嵌入深度学习中:一种作用于张量积表示(TPR)的注意力机制。通过在组合任务上的受控实验,我们证明了这种TPR-注意力机制在组合泛化方面优于现有的架构组件。这些结果凸显了将显式组合结构融入神经注意力机制的价值,并为构建具备系统性泛化能力的模型指出了一条有前景的道路。
cs.LG / 99 / 2608.30152

Converse and Collision-Based Achievability for Node Localization with Hybrid Distance-Spectral Graph Positional Encodings

基于混合距离-谱图位置编码的节点定位的逆界与碰撞可达性分析
Yan, Zimo, Li, Yifan, Li, Hao, Xie, Zheng, Liu, Chang, Tu, Zheming, Wang, Yuan
Abstract
Graph positional encodings are widely used in graph neural networks and graph Transformers, yet it remains unclear when the code itself can identify nodes. We study a hybrid distance-spectral encoding that combines anchor-distance profiles with quantized low-frequency Laplacian-energy coordinates. Treating the encoding as an observation map yields a simplex-refined converse, an exact collision factorization \(\kappa_H=\kappa_D\kappa_{S|D}\), and the collision information \(I_H=-\log\kappa_D-\log\kappa_{S|D}\). On random regular graphs, the criterion is made explicit through a bounded-correlation Gaussian-wave surrogate; for actual Laplacian-energy coordinates, we give the distance-conditioned spectral collision condition sufficient for conditional actual-coordinate achievability. Experiments show that \(I_H/\log n\) calibrates localization success, and PE-only structural task probes on Universal Dependencies trees show that hybrid encodings better recover syntactic-tree geometry than distance-only or spectral-only baselines.
Chinese Translation
图位置编码(Graph Positional Encodings)被广泛应用于图神经网络和图Transformer中,然而该编码本身何时能够识别节点仍不清楚。我们研究了一种混合距离-谱编码,它将锚点距离轮廓与量化的低频拉普拉斯能量坐标相结合。将该编码视为一个观测映射,我们得到了一个单纯形细化的逆界(converse)、一个精确的碰撞分解公式 \(\kappa_H=\kappa_D\kappa_{S|D}\),以及碰撞信息 \(I_H=-\log\kappa_D-\log\kappa_{S|D}\)。在随机正则图上,我们通过一个有界相关的高斯波代理模型使该判据显式化;对于实际的拉普拉斯能量坐标,我们给出了距离条件下的谱碰撞条件,该条件足以保证条件性实际坐标的可达性(achievability)。实验表明,\(I_H/\log n\) 能够校准定位成功率;在通用依存树(Universal Dependencies trees)上进行的仅基于位置编码的结构任务探测实验显示,混合编码比仅距离或仅谱的基线方法能更好地恢复句法树的几何结构。
cs.LG / 100 / 2608.30162

Reinforcement Learning for Symbolic Equation Solving

基于强化学习的符号方程求解
Keeffe, Kevin P O
Abstract
We present a reinforcement-learning agent that solves symbolic equations step by step, covering both nonlinear closed equations (radicals, exponentials, trigonometric) and a controlled class of restricted-open families requiring a change of variables (CoV) such as completing the square. We cast algebra as an MDP with a dynamic action space and a tree-structured policy (TreeMLP). The main policy learns from reward alone with no supervised solution traces; the CoV substitution comes from a supervised generator interchangeable with a CAS call. On closed equations the agent matches the prior best on CommonCore (0.93 greedy vs. ConPoLe's 0.925) under a single policy. On four hand-designed restricted-open families (quadratic, cubic, quartic, exponential) it reaches 0.79 beam / 0.67 greedy, exceeding the strongest non-learned search (A-star, 0.64). Learned CoV timing has content only on the exponential family, the one requiring a nested CoV, where a natural rule solves none of the held-out equations while the policy solves 75% from reward alone. At 10x scale a sharp seed-level bimodality emerges; a UCB learning-progress curriculum shows a non-significant positive trend toward mitigating it. We do not claim general open-equation solving: every open-equation result is confined to these four controlled families.
Chinese Translation
我们提出了一个逐步求解符号方程的强化学习智能体,涵盖非线性闭式方程(根式、指数、三角函数方程)以及一类需要变量代换(Change of Variables, CoV,如配方法)的受控受限开放方程族。我们将代数问题建模为一个具有动态动作空间和树状结构策略(TreeMLP)的马尔可夫决策过程(MDP)。主策略仅通过奖励信号学习,不依赖任何监督式的解答轨迹;CoV 代换则来自一个可与计算机代数系统(CAS)调用互换的监督式生成器。在闭式方程上,该智能体在单一策略下于 CommonCore 数据集上达到了此前最佳水平(贪婪解码 0.93,对比 ConPoLe 的 0.925)。在四个手工设计的受限开放方程族(二次、三次、四次和指数方程)上,其达到束搜索 0.79 / 贪婪解码 0.67 的准确率,超过了最强的非学习式搜索方法(A-star,0.64)。学习到的 CoV 时机仅在指数方程族上具有实质意义——该方程族需要嵌套的变量代换,在此类问题上,自然规则无法求解任何留出测试方程,而该策略仅凭奖励信号即可求解其中 75%。在 10 倍规模下,出现了显著的种子级双峰现象;基于 UCB 学习进度的课程学习方法在缓解该现象方面呈现出不显著的正向趋势。我们并不声称能求解一般性的开放方程:所有开放方程的实验结果均局限于这四个受控方程族。
cs.LG / 101 / 2608.30175

Benchmarking Peptide-Protein Affinity Prediction Across Peptide and Target Shifts

跨肽与靶标偏移的肽-蛋白亲和力预测基准测试
Tian, Jiaxin, An, Darren, Li, Jun
Abstract
Peptide-protein affinity models are often evaluated with a single data split, obscuring whether they interpolate among measurements for observed targets or generalize across peptide or target shifts. We integrated three sources of quantitative peptide-protein binding data to obtain 11,349 deduplicated pairs and benchmarked ten peptide representations, ESM-2 protein embeddings, and six regressors under peptide-similarity, within-target, and leave-target-out partitions. Across 60 matched representation-regressor configurations, mean test Spearman correlations were 0.462, 0.669, and 0.530, respectively. The top configuration shifted from ECFP-16 count fingerprints with random forest in the first two settings to HELM-BERT with Extra Trees when exact target sequences were excluded. Representation-rank correlations ranged from -0.042 to 0.624 across partitions, whereas regressor-rank correlations ranged from 0.771 to 0.943. Learning curves showed that representation differences were largest with limited supervision and narrowed as training data increased. PeptideCLM-2 adaptation and simple element-wise interaction features provided no consistent gain over a frozen encoder and direct concatenation under the tested protocols. These conclusions are specific to a dataset that pools transformed Kd, Ki, and IC50 measurements and to target exclusion at the exact-sequence level. Peptide-protein affinity benchmarks should therefore align data partitions with the intended use and jointly assess the effects of data scale, molecular representation, and downstream learner.
Chinese Translation
肽-蛋白亲和力模型通常仅在单一数据划分下进行评估,这掩盖了模型是在已观测靶标的测量值之间进行插值,还是能够在肽或靶标发生偏移时具备泛化能力。我们整合了三个定量肽-蛋白结合数据来源,获得了11,349个去重配对,并在肽相似性、同靶标内部以及留靶标(leave-target-out)三种划分下,对十种肽表示方法、ESM-2蛋白嵌入以及六种回归模型进行了基准测试。在60个匹配的表示-回归器配置中,三种划分下的平均测试Spearman相关系数分别为0.462、0.669和0.530。在前两种设置中,最优配置为ECFP-16计数指纹结合随机森林;而当完全排除靶标精确序列时,最优配置变为HELM-BERT结合Extra Trees。各表示方法的排名相关系数在不同划分间为-0.042至0.624,而回归器的排名相关系数则为0.771至0.943。学习曲线表明,表示方法之间的差异在监督数据有限时最大,并随训练数据增加而缩小。在所测试的协议下,PeptideCLM-2的适配以及简单的逐元素交互特征相较于冻结编码器和直接拼接并未带来一致的提升。上述结论仅适用于汇集了经转换的Kd、Ki和IC50测量值的数据集,以及精确序列层面的靶标排除场景。因此,肽-蛋白亲和力基准测试应使数据划分与预期用途相匹配,并联合评估数据规模、分子表示和下游学习器的共同影响。
cs.LG / 102 / 2608.30201

Certified Safety Radii in Forecast-Error Space for Wasserstein Distributionally Robust Small Signal Stability-Constrained AC Optimal Power Flow via Lifted Spectrahedral Containment

通过提升谱多面体包含关系在预测误差空间中为Wasserstein分布鲁棒小信号稳定性约束的交流最优潮流求解认证安全半径
Zhang, Ziqi, Chen, Xi
Abstract
Directly robustifying small-signal stability in AC optimal power flow is challenging since the stability boundary in the original uncertainty space is implicit, highly nonconvex, and changes with the operating decision. This paper exploits an alternative geometry. For a fixed model-specific stability certificate admitting suitable physical lifts, the small-signal stability requirement becomes an affine positive semidefinite constraint in the lifted variables, thereby defining a convex certified safe region. Instead of approximating the nonlinear instability boundary itself, we optimize a sample-wise safe radius in the original uncertainty space and certify, in the lifted space, that the entire power-flow image of the corresponding uncertainty ball is contained in the convex stability region. To this end, a componentwise Perron certificate guarantees existence, uniqueness, and Jacobian regularity of the target AC power-flow branch throughout each ball. An adjoint elimination then provides an exact affine-quadratic representation of the stability-relevant quantities, while rigorous matrix remainder bounds convert their nonlinear variation into finite robust PSD constraints. The resulting radii are certified lower bounds on the distances from empirical samples to failure and can therefore be coupled directly to the distance-based reformulation of a Wasserstein distributionally robust chance constraint, without directly approximating the instability boundary. Numerical studies demonstrate the effectiveness of the proposed framework.
Chinese Translation
在交流最优潮流(AC Optimal Power Flow)中直接对小信号稳定性进行鲁棒化处理极具挑战性,因为原始不确定性空间中的稳定性边界是隐式的、高度非凸的,且随运行决策而变化。本文利用了一种替代几何方法。对于允许进行适当物理提升(lifts)的特定模型的固定稳定性证书,小信号稳定性要求在提升变量下成为仿射半正定约束,从而定义了一个凸的认证安全区域。我们不直接逼近非线性失稳边界本身,而是在原始不确定性空间中优化逐样本的安全半径,并在提升空间中证明相应不确定性球的整个潮流映像被包含在该凸稳定性区域内。为此,一个逐分量的Perron证书保证了每个球内目标交流潮流支路的存在性、唯一性和雅可比矩阵的正则性。随后,伴随消元方法提供了稳定性相关量的精确仿射-二次表示,而严格的矩阵余项界将其非线性变化转化为有限的鲁棒半正定(PSD)约束。所得半径是经验样本到失稳距离的认证下界,因此可直接与基于距离的Wasserstein分布鲁棒机会约束重构相耦合,而无需直接逼近失稳边界。数值研究验证了所提框架的有效性。
cs.LG / 103 / 2608.30205

Diffusion-Based Refinement for Kilometer-Scale Probabilistic Precipitation Nowcasting

基于扩散模型的公里级概率降水临近预报精细化方法
Park, Dohyun, Song, Changhoon, Chang, Tengyuan, Ham, Yoo-Geun, Hong, Youngjoon
Abstract
Localized extreme precipitation is a major trigger of urban flash floods and landslides, yet producing nowcasts that combine fine spatial detail with probabilistic uncertainty remains challenging. Here we introduce exPreCast-ENS, a conditional residual diffusion framework that transforms the deterministic 4 km radar nowcaster exPreCast into a 1 km probabilistic ensemble while correcting systematic forecast errors. Conditioning on both the forecast and preceding radar observations lets the ensemble-mean correct the baseline rather than perturb it, while members represent unresolved fine-scale variability. Over the Korean Peninsula, skill improves with ensemble size. In two high-impact events in 2023, a 30-member ensemble recovers 38-47% of heavy-rain pixels missed by exPreCast while retaining approximately 95% of its correct detections and alarming on under 1% of the pixels it correctly left clear. The method generates a 1-h forecast in 3.4 s on a single GPU and yields consistent improvements on the French regional MeteoNet radar dataset.
Chinese Translation
局地极端降水是城市山洪和滑坡的主要诱因,然而生成兼具精细空间细节与概率不确定性信息的临近预报仍然具有挑战性。本文提出了exPreCast-ENS,一个条件残差扩散框架,它将确定性的4公里分辨率雷达临近预报系统exPreCast转化为1公里分辨率的概率集合预报,同时订正系统性预报误差。以预报场和先前的雷达观测作为条件,使集合平均值能够对基准预报进行订正而非简单扰动,而各集合成员则表征未被解析的精细尺度变率。在朝鲜半岛上,预报技巧随集合成员数增加而提升。在2023年两次高影响天气事件中,30个成员的集合预报恢复了exPreCast漏报的38-47%的强降雨像素,同时保留了其约95%的正确探测,且在原本正确判为无雨的像素中虚警率低于1%。该方法在单个GPU上生成1小时预报仅需3.4秒,并在法国区域MeteoNet雷达数据集上取得了一致的改进。
cs.LG / 104 / 2608.30252

Strong Drafts Need Compact Memories: Long-Context Speculative Decoding with Compressed KV Cache

强草稿需要紧凑的记忆:基于压缩KV缓存的长上下文投机解码
Yuan, Tong, Liao, Chengxi, Wen, Zeyi
Abstract
Long-context LLM applications such as document summarization and multi-turn agents require generation from prefixes spanning tens of thousands of tokens, making decoding latency a major bottleneck. Speculative decoding (SD) reduces latency without changing model outputs, but its speedup depends on both accepted draft tokens and draft-step latency: Lightweight drafts are fast but lack the capacity to capture long-range dependencies, whereas strong independent drafts recover acceptance but incur growing KV-access cost at long prefixes. We introduce memory-augmented drafting for long-context SD, equipping a strong independent draft with compressed draft-side KV memory: A lightweight adaptor constructs and incrementally updates this memory to retain distant information and exact recent context. The target verifier retains its full KV cache and applies the standard accept/reject rule, preserving SD's lossless guarantee. Experiments on Llama~3.1-8B and 70B targets at prefix lengths up to 32K show that our method reduces draft-side memory by over 70%. It achieves speedups of up to 2.08x and 3.33x , respectively, over autoregressive decoding.
Chinese Translation
诸如文档摘要和多轮智能体等长上下文大语言模型(LLM)应用需要基于长达数万词元的前缀进行生成,这使得解码延迟成为主要瓶颈。投机解码(Speculative Decoding, SD)能够在不改变模型输出的情况下降低延迟,但其加速效果取决于被接受的草稿词元数量和草稿步延迟:轻量级草稿速度快,但缺乏捕捉长程依赖的能力;而强大的独立草稿虽然能恢复接受率,但在长前缀下会带来不断增长的KV访问开销。我们提出了面向长上下文投机解码的记忆增强草稿方法,为强独立草稿配备压缩的草稿侧KV记忆:一个轻量级适配器构建并增量更新该记忆,以保留远距离信息并精确保存近期上下文。目标验证器保留其完整KV缓存并采用标准的接受/拒绝规则,从而保持投机解码的无损保证。在Llama 3.1-8B和70B目标模型上、前缀长度最高达32K的实验表明,我们的方法将草稿侧内存减少了70%以上,相对于自回归解码分别实现了最高2.08倍和3.33倍的加速。
cs.LG / 105 / 2608.30254

Exact Recovery Thresholds for Weighted Data Selection in Vector-Valued Linear Regression

向量值线性回归中加权数据选择的精确恢复阈值
Zhang, Guangjian
Abstract
We resolve the threshold part of Question 4 of the COLT 2025 open problem "Data Selection for Regression Tasks" of Hanneke, Moran, Shlimovich and Yehudayoff. In vector-valued linear regression with square loss $\ell_{(x,y)}(W)=|Wx-y|_2^2$, where $x\in\mathbb{R}^d$, $y\in\mathbb{R}^m$ and the learner is the empirical risk minimizer of minimal Frobenius norm, we prove that the minimal budget of weighted examples that recovers the full-data loss on every finite dataset is exactly $n^*(d,m)=(m+1)d$. We further determine two more values of the weighted selection profile $F_w(d,m,n)$: at the near-threshold budget, $F_w(d,m,(m+1)d-1)=1+\frac{1}{dm^2}$, and at the spanning budget, $F_w(d,m,d)=d+1$ for every $m$, while $F_w(d,m,n)=\infty$ for $n<d$. For the smallest open intermediate cell $(d,m)=(2,2)$ we prove $F_w(2,2,3)\in[13/8,15/8]$ and $F_w(2,2,4)\in[5/4,3/2]$, reduce the conjectured exact values $13/8$ and $5/4$ to a finite moment problem on the circle with at most seven atoms, and establish strong structural evidence for the conjecture. The upper-bound techniques (a fixed-basis conic compression lemma, a determinant-facet rigidity theorem for maximal certificates, and sharp sparsification lemmas for zero-mean weighted point systems) are of independent interest. As a byproduct we correct an erroneous claim circulating in a recent unrefereed preprint, exhibiting an explicit dataset with $m=2$ on which no weighted selection of $2d$ points recovers the optimal loss. All results are new only for $m\ge 2$; the scalar case $m=1$ is due to Hanneke et al.
Chinese Translation
我们解决了Hanneke、Moran、Shlimovich和Yehudayoff在COLT 2025开放问题"回归任务的数据选择"中问题4的阈值部分。在以平方损失 $\ell_{(x,y)}(W)=|Wx-y|_2^2$ 定义的向量值线性回归中(其中 $x\in\mathbb{R}^d$,$y\in\mathbb{R}^m$,学习器为Frobenius范数最小的经验风险最小化器),我们证明:在任意有限数据集上都能恢复全数据损失的加权样本的最小预算恰为 $n^*(d,m)=(m+1)d$。我们进一步确定了加权选择剖面 $F_w(d,m,n)$ 的另外两个取值:在近阈值预算处,$F_w(d,m,(m+1)d-1)=1+\frac{1}{dm^2}$;在张成预算处,对每个 $m$ 都有 $F_w(d,m,d)=d+1$,而当 $n<d$ 时 $F_w(d,m,n)=\infty$。对于最小的未决中间情形 $(d,m)=(2,2)$,我们证明 $F_w(2,2,3)\in[13/8,15/8]$ 和 $F_w(2,2,4)\in[5/4,3/2]$,将猜想中的精确值 $13/8$ 和 $5/4$ 归约为圆上至多含七个原子的有限矩问题,并为该猜想建立了强有力的结构性证据。本文的上界证明技术(一个固定基底的锥压缩引理、一个关于极大证书的行列式-面刚性定理,以及针对零均值加权点系统的尖锐稀疏化引理)具有独立的价值。作为附带成果,我们纠正了近期一篇未经同行评审的预印本中流传的一个错误论断,给出了一个 $m=2$ 的显式数据集,在该数据集上任何 $2d$ 个点的加权选择都无法恢复最优损失。所有结果仅在 $m\ge 2$ 时为新结果;标量情形 $m=1$ 归功于Hanneke等人。
cs.LG / 106 / 2608.30262

Multivariate Scientific Data Compression with Learned Cross-Variable Latent Decorrelation and Autoregressive Entropy Modeling

基于学习的跨变量潜在去相关与自回归熵建模的多变量科学数据压缩
Zhu, Liangji, Rangarajan, Anand, Ranka, Sanjay
Abstract
Scientific simulations generate collections of physical fields with heterogeneous statistics and dependencies, yet learned compressors often encode those fields independently or rely on a shared encoder without explicitly modeling the structure that remains in latent space. We present CAESAR-LDAR, an error-controlled multivariate learned compressor that augments a shared CAESAR-V backbone with two complementary mechanisms: a trainable orthogonal transform that reorganizes dependence across aligned latent channels, and a causal autoregressive hierarchical prior that captures local spatial structure left after transformation. Orthogonality is maintained through a matrix-exponential parameterization, making the transform exactly invertible without an additional penalty. A common residual-correction stage is applied uniformly to all variants to enforce the requested reconstruction tolerance. Experiments across combustion, climate, and turbulence data show that the two mechanisms are useful in different regimes. Latent decorrelation helps most when substantial linear cross-channel dependence survives the nonlinear encoder, whereas autoregressive modeling remains effective when the remaining structure is primarily local or spatial. Their combination provides the strongest or near-strongest rate-distortion performance across the evaluated datasets. The global transform adds little computational overhead, while autoregressive coding introduces a larger throughput tradeoff. More broadly, the results suggest a practical design principle for multivariate scientific compression: exploit global cross-channel dependence when it is measurably present in latent space, and use local probabilistic context as a complementary mechanism across a wider range of data regimes.
Chinese Translation
科学模拟生成的物理场集合具有异构的统计特性和依赖关系,然而学习型压缩器通常对这些场进行独立编码,或依赖共享编码器而不显式建模潜在空间中残留的结构。我们提出了CAESAR-LDAR,一种误差受控的多变量学习型压缩器,它在共享的CAESAR-V骨干网络上增加了两种互补机制:一是可训练的正交变换,用于重新组织对齐的潜在通道之间的依赖关系;二是因果自回归分层先验,用于捕获变换后残留的局部空间结构。正交性通过矩阵指数参数化来维持,使该变换无需额外惩罚项即可精确可逆。所有变体均统一应用一个共同的残差校正阶段,以满足所要求的重建容差。在燃烧、气候和湍流数据上的实验表明,这两种机制在不同的数据情境下各有优势。当非线性编码器之后仍存在显著的线性跨通道依赖时,潜在去相关的效果最好;而当残留结构主要是局部或空间性的时候,自回归建模依然有效。二者的结合在所有评估数据集上提供了最强或接近最强的率失真性能。全局变换带来的计算开销很小,而自回归编码则引入了较大的吞吐量权衡。更广泛地说,这些结果为多变量科学数据压缩提供了一个实用的设计原则:当潜在空间中可测量地存在全局跨通道依赖时加以利用,并在更广泛的数据情境下将局部概率上下文作为补充机制。
cs.LG / 107 / 2608.30283

BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning

BCPPO:基于Bachelier启发的约束近端策略优化,用于尾部风险感知的安全强化学习
Hou, Dongsheng, Chen, Yanqiao, Rui, Yuhan
Abstract
Expected-cost constraints can still permit rare, high-cost events. Monte Carlo conditional value at risk (CVaR) gradients can be noisy at high confidence, whereas critics that model an outcome distribution add complexity. We propose BCPPO (Bachelier-Inspired Constrained Proximal Policy Optimization), a proximal policy optimization (PPO) method. Separately initialized cost-prediction networks (critics), trained with random sample masks, produce disagreement that marks predictions sensitive to which state-action regions occur in the training data and to critic training. A Bachelier formula for the expected amount above a reference level converts this disagreement into a smooth policy-update penalty. Gradients from this penalty do not alter the critics, so temporal-difference (TD) critic learning is unchanged. A saturation-aware controller adjusts the mean-cost penalty and stops accumulated error from growing while that penalty is clipped. Deployment retains only the policy network. The disagreement penalty is neither a tail-event probability nor a guaranteed error bound, and it provides no safety guarantee. Across 175 runs with shared tasks, costs, budgets, training steps, and evaluation seeds, no comparator attains both higher mean return and lower mean CVaR than BCPPO in any task. On Push1, BCPPO has no lower return and no higher CVaR than every comparator, with at least one strict gain. These results support a practical balance among reward, caution around cost predictions that vary across trained critics, and policy-only deployment.
Chinese Translation
期望成本约束仍可能允许罕见的、高成本事件发生。蒙特卡洛条件风险价值(CVaR)梯度在高置信度下可能存在较大噪声,而对结果分布进行建模的评论家(critic)网络则会增加复杂度。我们提出BCPPO(Bachelier启发的约束近端策略优化),这是一种近端策略优化(PPO)方法。分别初始化的成本预测网络(评论家)采用随机样本掩码进行训练,其产生的分歧可标记出那些对训练数据中出现的状态-动作区域以及评论家训练过程敏感的预测。基于Bachelier公式的超出参考水平的期望金额计算方法,将该分歧转化为平滑的策略更新惩罚项。该惩罚项产生的梯度不会改变评论家网络,因此时序差分(TD)评论家学习保持不变。一个饱和感知的控制器会调整平均成本惩罚项,并在该惩罚项被截断时阻止累积误差的增长。部署时仅保留策略网络。分歧惩罚既不是尾部事件概率,也不是有保证的误差上界,它不提供任何安全性保证。在共享任务、成本、预算、训练步数和评估种子的175次运行实验中,没有任何对比方法在任一任务中同时实现比BCPPO更高的平均回报和更低的平均CVaR。在Push1任务中,BCPPO的回报不低于所有对比方法,且CVaR不高于所有对比方法,并在至少一项指标上取得严格优势。这些结果表明,BCPPO在奖励、对训练后不同评论家成本预测差异的谨慎处理以及仅部署策略网络之间实现了实用的平衡。
cs.LG / 108 / 2608.30295

CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration

CateKV:面向长上下文大语言模型推理加速的顺序一致性研究
Jiang, Haoyun, Li, Haolin, Zhang, Jianwei, Huang, Fei, Hu, Qiang, Sun, Minmin, Xiao, Shuai, Li, Yong, Lin, Junyang, Yao, Jiangchao
Abstract
Large language models (LLMs) have demonstrated strong capabilities in handling long-context tasks, but processing such long contexts remains challenging due to the substantial memory requirements and inference latency. In this work, we discover that certain attention heads exhibit sequential consistency in their attention patterns, which can be persistently identified using a coefficient-of-variation-based algorithm. Inspired by this observation, we propose CateKV, a hybrid KV cache method that retains only critical token information for consistent heads, thereby reducing KV cache size and computational overhead, while preserving the majority of KV pairs in adaptive heads to ensure high accuracy. We show the unique characteristics of our algorithm and its extension with existing acceleration methods. Comprehensive evaluations on long-context benchmarks show that, while maintaining accuracy comparable to full attention, CateKV reduces memory usage by up to $2.72\times$ and accelerates decoding by $2.18\times$ in single-sample inputs, and boosts throughput by $3.96\times$ in batch scenarios.
Chinese Translation
大语言模型(LLM)在处理长上下文任务方面展现出强大能力,但由于巨大的内存需求和推理延迟,处理如此长的上下文仍然充满挑战。在本工作中,我们发现某些注意力头(attention head)在其注意力模式中表现出顺序一致性(sequential consistency),并可通过一种基于变异系数的算法被稳定地识别出来。受此观察启发,我们提出了 CateKV,一种混合 KV 缓存方法:对于一致性注意力头,仅保留关键 token 信息,从而降低 KV 缓存大小和计算开销;同时对于自适应注意力头,保留大部分 KV 对以确保高准确率。我们展示了该算法的独特特性及其与现有加速方法的结合扩展。在长上下文基准测试上的全面评估表明,CateKV 在保持与全注意力机制相当的准确率的同时,在单样本输入场景下可将内存使用降低至多 2.72 倍、解码加速 2.18 倍,在批量场景下可将吞吐量提升 3.96 倍。
cs.LG / 109 / 2608.30310

Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs

Tail-Replay:摆脱混合大语言模型前缀缓存中线性注意力的诅咒
Liu, Yirui, Qi, Ruoling, Wu, Xuaner, Liu, Penghang, Chen, Jian
Abstract
Hybrid large language models interleave full-attention layers with linear-attention layers to reduce the cost of long-context inference. This structure complicates prefix caching: full-attention key-value caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundaries. Existing hybrid prefix caching methods address this mismatch by storing recurrent-state checkpoints. As a result, token-level matches are directly usable only at positions aligned with stored checkpoints, constraining prefix reuse to a discrete set of boundaries. We present Tail-Replay, a prefix caching mechanism that enables unconstrained token-level prefix reuse in hybrid large language models. The key insight is that linear-attention mechanisms such as Gated DeltaNet can be viewed as a structured, lossy compression of the input prefix: gated recurrent updates progressively attenuate the contributions of earlier inputs. Consequently, the recurrent state of a matched prefix can be well approximated by replaying only a short, recent suffix of that prefix. Tail-Replay exploits this property by caching the exact full-attention key-value cache while omitting recurrent-state checkpoints. On a cache hit, it reconstructs the linear-attention states by replaying a short, recent suffix of the matched prefix. As a result, the reuse boundary is determined by the shared tokens rather than by recurrent-state checkpoints. We evaluate Tail-Replay on three Gated DeltaNet-based hybrid models using the LongBench and RULER benchmarks. With only a 5--10\% replay budget, it retains 92.8--99.9\% of full-prefill quality on LongBench and RULER. For serving efficiency, we evaluate time-to-first-token speedups across multiple matched-prefix lengths---8K, 16K, and 32K. The speedup grows with prefix length, reaching $9.1$--$14.3\times$ over full prefill at 32K.
Chinese Translation
混合大语言模型将全注意力层与线性注意力层交错组合,以降低长上下文推理的成本。这种结构使前缀缓存变得复杂:全注意力的键值缓存是按 token 可寻址的,而线性注意力层维护的是递归状态,无法回滚到任意的前缀边界。现有的混合前缀缓存方法通过存储递归状态检查点来解决这一不匹配问题。其结果是,token 级别的匹配只能在与已存储检查点对齐的位置上直接使用,从而将前缀复用限制在离散的边界集合中。我们提出了 Tail-Replay,这是一种在混合大语言模型中实现无约束 token 级前缀复用的前缀缓存机制。关键洞察在于,诸如 Gated DeltaNet 这类线性注意力机制可以被视为对输入前缀的一种结构化有损压缩:门控递归更新会逐渐衰减较早输入的贡献。因此,匹配前缀的递归状态可以通过仅重放该前缀的较短的近期后缀来很好地近似。Tail-Replay 利用了这一特性,缓存精确的全注意力键值缓存,同时省略递归状态检查点。在缓存命中时,它通过重放匹配前缀的较短的近期后缀来重建线性注意力状态。这样一来,复用边界由共享的 token 决定,而不是由递归状态检查点决定。我们在三个基于 Gated DeltaNet 的混合模型上,使用 LongBench 和 RULER 基准对 Tail-Replay 进行了评估。仅需 5–10% 的重放预算,它在 LongBench 和 RULER 上即可保留 92.8–99.9% 的完整预填充(full-prefill)质量。在服务效率方面,我们评估了在 8K、16K 和 32K 多种匹配前缀长度下的首 token 生成时间(time-to-first-token)加速比。加速比随前缀长度增长而提升,在前缀长度为 32K 时,相比完整预填充达到 9.1–14.3 倍的加速。
cs.LG / 110 / 2608.30315

Context Staircase: Signature-Aligned Dynamics of Token Embeddings under Small Initialization

上下文阶梯:小初始化下词元嵌入与概率签名对齐的动力学
Yao, Junjie, Hang, Liangkai, Xu, Zhi-Qin John
Abstract
Token embeddings are the basic representational units that connect discrete tokens with continuous computation in language models. Although modern language models learn embeddings from random initialization through gradient-based training, the dynamical mechanism by which meaningful embedding structures emerge remains unclear. In this work, we identify that the evolving embedding structures are closely related to token-conditioned label and contextual distributions, which we formalize as probability signatures. We observe a progressive learning process, which we term Context Staircase: embeddings learn the low-order statistic signatures of the data before the high-order ones. More specifically, we observe that early in training they align with the simplest, context-free signature linking a token to its label, and as training proceeds, they progressively reflect signatures involving more and more context tokens. We then analyze the gradient flow of embeddings under small initialization to explain this phenomenon, deriving embedding evolution equations for feed-forward and self-attention architectures. We further extend these observations to real language-model training. Finally, we show that these embedding structures play an important role in both task learning and the incorporation of semantic structure into the embedding space. Overall, our results provide a dynamic explanation of how data statistics and architecture jointly shape token embeddings in language models, and reveal an implicit bias in the space of data statistics: training proceeds from simpler, low-order statistical relations toward increasingly complex, context-dependent ones.
Chinese Translation
词元嵌入(token embeddings)是语言模型中连接离散词元与连续计算的基本表示单元。尽管现代语言模型通过基于梯度的训练从随机初始化中学习嵌入,但其中有意义的嵌入结构涌现的动力学机制仍不清楚。在本工作中,我们发现不断演化的嵌入结构与以词元为条件的标签分布和上下文分布密切相关,我们将其形式化为概率签名(probability signatures)。我们观察到一种渐进式学习过程,并将其称为上下文阶梯(Context Staircase):嵌入先学习数据的低阶统计签名,再学习高阶统计签名。更具体地说,我们观察到在训练早期,嵌入与将词元与其标签关联的最简单的无上下文签名对齐;随着训练的进行,它们逐渐反映出涉及越来越多上下文词元的签名。随后,我们在小初始化条件下分析嵌入的梯度流以解释这一现象,推导了前馈架构和自注意力架构下嵌入的演化方程。我们进一步将这些观察扩展到真实语言模型的训练中。最后,我们证明这些嵌入结构在任务学习以及语义结构融入嵌入空间的过程中都发挥着重要作用。总体而言,我们的结果为数据统计与架构如何共同塑造语言模型中的词元嵌入提供了动力学解释,并揭示了数据统计空间中的一种隐式偏置:训练从较简单的低阶统计关系逐步走向日益复杂的、依赖上下文的高阶统计关系。
cs.LG / 111 / 2608.30317

Online Estimation of Dynamic Origin-Destination Matrices Using Reinforcement Learning with Link-Flow Propagation Guidance

基于链路流量传播引导的强化学习在线动态起讫点矩阵估计
Min, Donggyu, Kim, Dong-Kyu
Abstract
Online dynamic origin-destination (OD) matrix estimation (DODE) calibrates time-dependent OD demand to reproduce observed link-flow trajectories. In online, OD demand should be estimated from current observations and propagated network states while subsequent observations and stochastic dynamic network loading (DNL) outcomes remain uncertain. Recently, reinforcement learning (RL) has emerged as a promising alternative, reducing computational burden by replacing iterative algorithms while being applicable to stochastic environments. However, because the policy is trained offline and deployed online, it must handle varying target link-flow trajectories; since each target trajectory defines the link-flow error used in the reward, the same OD demand vector can require different adjustments, making conventional scalar feedback ambiguous. To address this gap, this study proposes LFPG-RL, which integrates link-flow propagation guidance (LFPG) into proximal policy optimization (PPO). LFPG combines link-flow error sensitivities with the contribution of each OD-time demand component to simulated link flows, transforming aggregate mismatch into OD-specific advantage shaping for PPO actor updates. At deployment, the policy requires only a single forward pass. LFPG-RL is developed and evaluated on 250 weekday trajectories of 15-min link-flow data from a Melbourne arterial network modeled by a link transmission model with stochastic route choice. On held-out trajectories, LFPG-RL achieved an RMSE of 4.69, MAPE of 20.15%, and Pearson correlation of 0.995. These results support the contention that our method is a more efficient and accurate online OD demand calibration method compared to existing ones.
Chinese Translation
在线动态起讫点(OD)矩阵估计(DODE)通过标定时变的OD需求来再现观测到的路段流量轨迹。在在线场景下,OD需求需要根据当前观测和传播的网络状态进行估计,而后续观测以及随机动态网络加载(DNL)的结果仍具有不确定性。近年来,强化学习(RL)作为一种有前景的替代方案应运而生,它通过取代迭代算法降低了计算负担,同时可适用于随机环境。然而,由于策略是离线训练、在线部署的,它必须应对不同的目标路段流量轨迹;由于每条目标轨迹都定义了奖励中使用的路段流量误差,同一个OD需求向量可能需要不同的调整,这使得传统的标量反馈变得模糊。为解决这一问题,本研究提出了LFPG-RL方法,将路段流量传播引导(LFPG)集成到近端策略优化(PPO)中。LFPG将路段流量误差敏感度与各OD-时间需求分量对仿真路段流量的贡献相结合,将总体失配转化为面向OD的优势函数塑形,用于PPO执行器(actor)的更新。在部署阶段,该策略仅需一次前向计算。LFPG-RL基于墨尔本一个干线路网的15分钟路段流量数据(共250个工作日轨迹)进行开发与评估,该路网采用考虑随机路径选择的路段传输模型建模。在留出轨迹上,LFPG-RL取得了4.69的RMSE、20.15%的MAPE以及0.995的皮尔逊相关系数。这些结果表明,与现有方法相比,本文方法是一种更高效、更准确的在线OD需求标定方法。
cs.LG / 112 / 2608.30323

Generative multi-domain transfer learning for fault detection in data-scarce wind turbines

面向数据稀缺风电机组故障检测的生成式多域迁移学习方法
Jonas, Stefan, Meyer, Angela
Abstract
Normal behavior models have shown promise for reliable fault detection in wind turbines. However, these unsupervised anomaly detection models require sufficient fault-free training data to learn the normal operation behavior of turbines. Under data scarcity, for example in newly deployed wind turbines, these models may result in poor fault detection performance. In this work, we propose a multi-domain generative domain mapping approach based on Star Generative Adversarial Networks (StarGAN) to improve fault detection on data-scarce wind turbines. Our model maps SCADA measurements from a data-scarce turbine to resemble those of several data-rich turbines. By preserving the operational state during translation, faults occurring in a data-scarce domain can be mapped and detected by reliable pre-trained normal behavior models of data-rich domains. Highlighting the benefits of an ensemble fusion strategy, we show that under severe data scarcity our method can produce anomaly scores comparable to models trained on large representative datasets. Our approach can consistently outperform models trained on scarce data when less than 2 weeks of training data are available. With just 2 weeks of accumulated training data, we achieve an anomaly score similarity that is, on average, +16% higher than conventional fine-tuning, and +10% higher than single-source domain mapping. As a step towards unsupervised model selection, we propose a proxy metric that detects poor performance at training time, despite an absence of anomalies. Our study presents the potential and challenges of multi-domain mapping for wind turbine fault detection under unrepresentative training data.
Chinese Translation
正常行为模型在风电机组可靠故障检测方面已展现出良好前景。然而,这些无监督异常检测模型需要充足的无故障训练数据来学习风电机组的正常运行行为。在数据稀缺的情况下(例如新部署的风电机组),这些模型的故障检测性能可能较差。本文提出了一种基于星型生成对抗网络的多域生成式域映射方法,以改善数据稀缺风电机组的故障检测。我们的模型将数据稀缺风电机组的SCADA测量数据映射为类似于多个数据丰富风电机组的数据。通过在转换过程中保留运行状态,数据稀缺域中发生的故障可以被映射,并由数据丰富域上可靠的预训练正常行为模型进行检测。通过凸显集成融合策略的优势,我们表明在严重数据稀缺的情况下,我们的方法可以产生与基于大型代表性数据集训练的模型相当的异常分数。当可用训练数据少于2周时,我们的方法能够持续优于在稀缺数据上训练的模型。仅用2周累积的训练数据,我们实现的异常分数相似度平均比传统微调高16%,比单源域映射高10%。作为迈向无监督模型选择的一步,我们提出了一种代理指标,能够在缺乏异常样本的情况下于训练阶段检测出性能不佳的模型。本研究展示了在训练数据不具代表性的情况下,多域映射用于风电机组故障检测的潜力与挑战。
cs.LG / 113 / 2608.30328

Learning PDE Time-Stepping with Neural Cellular Automata

基于神经细胞自动机学习偏微分方程时间推进方法
Saha, Esha, Wang, Hao
Abstract
Classical numerical solvers for partial differential equations (PDEs) are computationally expensive to solve repeatedly across varying initial conditions, motivating the need for learned surrogates. In this paper, we propose a trainable Neural Cellular Automata (NCA) based surrogate model for learning long time PDE dynamics. Rather than mapping an entire initial field to a full trajectory in one shot, our proposed model learns a small, local, homogeneous update rule that is applied identically and repeatedly at every grid cell, mirroring the locality of differential operators. We benchmark this framework against three baselines: PDE - Net, a modified physics-informed neural network (PINN), and a Fourier Neural Operator (FNO), on five canonical PDEs (heat, advection, Burgers, Allen - Cahn, and Fisher - KPP), evaluated at temporal domain two times beyond the training temporal domain. The proposed model achieves the lowest long-horizon relative errors on the majority of the experiments.
Chinese Translation
针对不同初始条件反复求解偏微分方程(PDE)时,经典数值求解器计算代价高昂,这促使人们寻求可学习的代理模型。本文提出一种可训练的基于神经细胞自动机(Neural Cellular Automata, NCA)的代理模型,用于学习偏微分方程的长时间动力学。与将整个初始场一次性映射到完整轨迹的做法不同,我们所提出的模型学习一个小的、局部的、同质的更新规则,该规则在每个网格单元上以相同方式反复应用,从而模拟微分算子的局部性。我们在五个经典偏微分方程(热传导方程、对流方程、Burgers方程、Allen-Cahn方程和Fisher-KPP方程)上,将该框架与三种基线方法进行对比:PDE-Net、一种改进的物理信息神经网络(PINN)以及傅里叶神经算子(FNO),并在两倍于训练时间域的时间范围上进行评估。在大多数实验中,所提出的模型取得了最低的长时间跨度相对误差。
cs.LG / 114 / 2608.30337

Coarse composition suffices: tabular in-context learning for multi-activity antimicrobial peptide profiling

粗粒度组合足矣:用于多活性抗菌肽谱分析的表格化上下文学习方法
Kumar, Raunak, Pal, Anuj, Solanki, Dhruvi, Pareek, Parikshit, Singh, Juhi, Singla, Jitin
Abstract
Antimicrobial peptides (AMPs) often act against multiple pathogen classes, making multi-label activity prediction a more realistic screening target than binary antimicrobial classification. The ESCAPE benchmark formalizes this setting, but leading approaches typically rely on multimodal, structure-conditioned deep models that are costly to train and tune. We show that a simple, sequence-only pipeline can match and surpass these methods by combining 330 interpretable sequence descriptors with TabPFN, a tabular foundation model that performs in-context prediction in a single forward pass without gradient-based training or hyperparameter search. On ESCAPE (82,359 peptides; five labels), a label-powerset TabPFN model achieves mAP-5 = 77.8%, improving on the previously best reported 72.1%. A probabilistic classifier chain is the first method to match or exceed the best published average precision on each of the five labels simultaneously. The gains persist under the prior state-of-the-art single-fold training protocol, indicating they are not a training-set-size artefact, and are largest for remote homologues (+11.2 points below 30% sequence identity). Ablations further show that predicted structure is unnecessary at inference and that performance is not driven by any single descriptor family: ten global physicochemical scalars recover 91% of full-feature performance. Finally, explicitly modelling label dependence yields targeted benefits for scarce activities and supports ranking which activity to assay next from partial positive evidence.
Chinese Translation
抗菌肽(AMPs)通常对多种病原体类别起作用,因此多标签活性预测比二元抗菌分类更符合实际筛选目标。ESCAPE基准测试将这一设定形式化,但主流方法通常依赖多模态、结构条件化的深度模型,其训练和调优成本高昂。我们证明,一个简单的仅基于序列的流水线通过将330个可解释的序列描述符与TabPFN(一种表格基础模型,可在单次前向传播中完成上下文预测,无需基于梯度的训练或超参数搜索)相结合,即可匹配并超越这些方法。在ESCAPE数据集(82,359条肽;五个标签)上,标签幂集(label-powerset)TabPFN模型实现了mAP-5 = 77.8%,优于此前报道的最佳结果72.1%。概率分类器链是首个在全部五个标签上同时达到或超越已发表最佳平均精度的方法。在先前的最先进单折训练协议下,这些增益依然存在,表明它们并非训练集规模带来的假象,且增益在远缘同源序列上最大(序列一致性低于30%时提升11.2个百分点)。消融实验进一步表明,推理阶段无需预测结构,且性能并非由任何单一描述符家族驱动:仅十个全局理化标量即可恢复全特征性能的91%。最后,显式建模标签依赖性对稀缺活性带来针对性收益,并支持基于部分阳性证据对下一个待检测活性进行排序。
cs.LG / 115 / 2608.30364

Beyond Churn: Predicting Financial Fragmentation in Retail Banking with Temporal Machine Learning

超越客户流失:基于时序机器学习的零售银行业务资金分流预测
Chopra, Ananyaa, Xu, Brandon, Yuen, Brendan, Zung, Lauren, Aulakh, Sarabroop
Abstract
Retail banking attrition is usually represented as a terminal binary event, even though client relationships often weaken earlier through partial movements of deposits, investments, and recurring activity to external financial institutions. This paper defines that preceding state as financial fragmentation and presents an end-to-end temporal machine-learning system for predicting it before complete disengagement. Using anonymized multi-source data from a large retail bank, the framework predicts whether a valid external transfer or investment event will occur within 90 days. The study uses 595,220 client-month observations, with 346 engineered features combining monthly client profiles, balances, product relationships, prior flow-of-funds behavior, macroeconomic conditions, and competitor activity. A four-stage XGBoost cascade estimates (1) whether an external outflow will occur within 90 days, (2) the expected amount, (3) the originating product, and (4) the destination financial institution. The primary classifier achieved a test precision-recall area under the curve of 0.823. At the validation-selected threshold, it produced 86.4% precision, 75.1% recall, and an F1 score of 0.803. Ranking test observations in descending Stage 1 fragmentation score, the top 1% of clients yielded 95.3% precision, while the top 5% captured 78.7% of observed outflow cases. The amount model placed 94.9% of predictions within an adjacent amount bucket. Destination prediction reached a macro-F1 of 0.81 across 27 classes; source-product prediction achieved a weighted F1 of 0.92. By moving the analytical focus from terminal churn to earlier fund migration, the proposed approach provides a practical foundation for proactive, explainable, and economically informed client-retention decision support.
Chinese Translation
零售银行业的客户流失通常被建模为一个终末二元事件,但实际上客户关系往往更早开始弱化:存款、投资和经常性业务会部分转移至外部金融机构。本文将这一先行状态定义为“资金分流”(financial fragmentation),并提出一套端到端的时序机器学习系统,用于在客户完全脱离之前对其进行预测。基于某大型零售银行的匿名化多源数据,该框架预测未来90天内是否会发生有效的对外转账或投资事件。研究使用595,220条客户-月度观测数据,构建了346个工程特征,涵盖客户月度画像、余额、产品持有关系、历史资金流动行为、宏观经济状况及竞争对手活动。系统采用四阶段XGBoost级联模型,分别估计:(1)90天内是否会发生对外资金流出;(2)预期流出金额;(3)流出的源头产品;(4)目标金融机构。主分类器在测试集上取得了0.823的精确率-召回率曲线下面积(PR-AUC)。在验证集选定的阈值下,模型达到86.4%的精确率、75.1%的召回率以及0.803的F1分数。按阶段一资金分流得分降序排列测试观测样本,前1%客户的精确率达到95.3%,前5%则捕获了78.7%的实际资金流出案例。金额模型将94.9%的预测控制在相邻金额区间内。目标机构预测在27个类别上取得0.81的宏平均F1;源头产品预测达到0.92的加权F1。通过将分析焦点从终末性客户流失转向更早期的资金迁移,所提出的方法为主动、可解释且具备经济依据的客户挽留决策支持奠定了实用基础。
cs.LG / 116 / 2608.30366

Mode Connectivity Beyond Classifiers: Evidence from Generative and Contrastive Models

超越分类器的模式连通性:来自生成模型与对比模型的证据
Yao, Chengzheyi, Zhang, Yongzhao, Tian, Yongding
Abstract
The loss landscape of Deep Neural Networks (DNNs) exhibits highly complex and non-convex properties. Recent studies have revealed the phenomenon of mode connectivity, demonstrating that independently trained network modes can be connected via a continuous low-loss path. However, existing mode connectivity research is predominantly confined to classifier-based models, leaving it an open question whether similar geometric properties exist in modern complex models. In this paper, we extend the boundaries of mode connectivity to generative and contrastive domains (specifically DDPM and NanoCLIP). Addressing the unique architecture of DDPM and CLIP, we propose an architecture-aware connection building algorithm. Extensive empirical results demonstrate for the first time that we successfully discover mode connectivity between independently trained DDPM and NanoCLIP modes. Our work provides a novel perspective for understanding the geometric properties of the loss landscapes in modern generative and contrastive models.
Chinese Translation
深度神经网络(DNN)的损失景观呈现出高度复杂和非凸的特性。近期研究揭示了模式连通性(mode connectivity)现象,表明独立训练的网络模式可以通过连续的低损失路径相连接。然而,现有的模式连通性研究主要局限于基于分类器的模型,类似几何性质是否也存在于现代复杂模型中仍是一个悬而未决的问题。本文将模式连通性的研究边界扩展至生成式和对比学习领域(具体为 DDPM 和 NanoCLIP)。针对 DDPM 和 CLIP 的独特架构,我们提出了一种架构感知的连接构建算法。大量实证结果首次表明,我们成功发现了独立训练的 DDPM 和 NanoCLIP 模式之间的模式连通性。我们的工作为理解现代生成式和对比模型损失景观的几何性质提供了新的视角。
cs.LG / 117 / 2608.30367

Beat-Synchronous Tokenization for ECG Transformers

面向心电图Transformer的搏动同步分词方法
Sameh, Ahmed, Wilson, Nolan, Enderlein, Max, Varatharajah, Yogatheesan
Abstract
Transformer-based electrocardiogram (ECG) models commonly tokenize waveforms into fixed temporal patches. Though convenient, fixed patching can split heartbeat structures across token boundaries. We study beat-synchronous tokenization as a physiologically grounded alternative, comparing fixed patches with three beat-aligned strategies: resampled beats, adaptive pooled beats, and resampled beats augmented with R--R interval information. Experiments span two settings: 10-second 12-lead diagnostic classification on PTB-XL after MIMIC-IV-ECG masked pretraining, and 60-second single-lead rhythm classification on Icentia11k after patient-level contrastive pretraining. On PTB-XL, resampled beat tokens achieve the highest mean macro Area Under the ROC Curve (AUROC; 0.8945) and nearly match the best fixed-patch macro Area Under the Precision-Recall Curve (AUPRC; 0.7414), reducing average sequence length from 100 to 11.2 tokens. On Icentia11k, beat-synchronous tokenizers obtain comparable AUPRC to fixed patching with better stability across runs. These results suggest morphology-preserving beat tokenization is a compact, competitive alternative to fixed temporal patching.
Chinese Translation
基于Transformer的心电图(ECG)模型通常将波形分词为固定时间长度的片段。尽管方便,但固定分段可能将心跳结构切分到不同的词元边界。我们研究了搏动同步分词(beat-synchronous tokenization)作为一种具有生理学依据的替代方案,将固定片段与三种心跳对齐策略进行比较:重采样心跳、自适应池化心跳,以及结合R-R间期信息的重采样心跳。实验涵盖两种设置:一是在MIMIC-IV-ECG掩码预训练后,在PTB-XL数据集上进行10秒12导联诊断分类;二是在患者级对比预训练后,在Icentia11k数据集上进行60秒单导联心律分类。在PTB-XL上,重采样心跳词元取得了最高的平均宏平均ROC曲线下面积(AUROC;0.8945),并几乎达到最优固定片段的宏平均精确率-召回率曲线下面积(AUPRC;0.7414)的水平,同时将平均序列长度从100个词元降至11.2个。在Icentia11k上,搏动同步分词器获得了与固定分段相当的AUPRC,且在多次运行中具有更好的稳定性。这些结果表明,保留形态学信息的心跳分词是固定时间分段的一种紧凑且具有竞争力的替代方案。
cs.LG / 118 / 2608.30382

Convergence rates for the RMSprop optimizer with full control of the hyperparameters

具有超参数完全可控性的RMSprop优化器的收敛速度
Dereich, Steffen, Jentzen, Arnulf
Abstract
Popular adaptive stochastic gradient descent (SGD) methods to train artificial intelligence (AI) systems include the RMSprop, the Adam, and the AdamW optimizers, where the adaptivity parts in Adam and AdamW basically just coincide with RMSprop. Such adaptive methods involve several hyperparameters including the regularization parameter $\epsilon$ (which ensures that one does not divide by 0 and is often chosen to be very close to zero such as $10^{-8}$ in PyTorch by default) and the second moment decay parameter $\beta$ (which is often chosen to be very close to $1$ such as 0.99 (RMSprop) and 0.999 (Adam and AdamW) in PyTorch by default). Despite the high relevance of such methods, it remains an open research problem to provide error estimates for such methods with the error constants being not exploding but uniformly bounded with the respect to the hyperparameters, even in the situation of convex stochastic optimization problems. It is the key contribution of this work to essentially solve this problem for RMSprop. Specifically, we bound the expectation of the stopped evaluation of the objective function at the RMSprop process from above by the sum of an initialization term that decays exponentially in the training time, a stochastic approximation remainder of order $\gamma_n$, and a memory error of order $( 1 - \beta)^2$ with the error constants being uniformly controlled over all admissible choices of the step sizes, the second moment decay parameter $\beta$ and the regularization parameter $\epsilon\in[0,1]$ (also covering $\epsilon=0$). Our non-asymptotic error estimates hold not just for all sufficiently large n but hold for every gradient step $n=1,2,3,...$ with all error constants being explicitly specified. The key innovative new feature in the proof of our analysis are suitable inverse moment bounds for the second moment process in RMSprop.
Chinese Translation
用于训练人工智能(AI)系统的流行自适应随机梯度下降(SGD)方法包括RMSprop、Adam和AdamW优化器,其中Adam和AdamW的自适应部分基本上与RMSprop相同。这类自适应方法涉及多个超参数,包括正则化参数 $\epsilon$(用于确保不除以0,通常取非常接近0的值,例如PyTorch中默认为 $10^{-8}$)以及二阶矩衰减参数 $\beta$(通常取非常接近1的值,例如PyTorch中默认为0.99(RMSprop)和0.999(Adam和AdamW))。尽管这类方法具有重要意义,但即使在凸随机优化问题的情形下,如何为这类方法提供误差估计,并使误差常数相对于超参数不发散而是均匀有界,仍然是一个开放的研究问题。本工作的关键贡献在于从本质上解决了RMSprop的这一 问题。具体而言,我们将目标函数在RMSprop过程停止时刻的期望值从上方界定为以下三项之和:一个随训练时间指数衰减的初始化项、一个阶为 $\gamma_n$ 的随机逼近余项,以及一个阶为 $(1-\beta)^2$ 的记忆误差项,且误差常数在所有可容许的步长、二阶矩衰减参数 $\beta$ 以及正则化参数 $\epsilon\in[0,1]$(也包括 $\epsilon=0$)的选择下均被均匀控制。我们的非渐近误差估计不仅对所有充分大的 $n$ 成立,而且对每一个梯度步 $n=1,2,3,\ldots$ 都成立,且所有误差常数均被显式给出。我们分析证明中的关键创新之处在于为RMSprop中的二阶矩过程建立了适当的逆矩不等式。
cs.LG / 119 / 2608.30384

RSLM: Training-Free Vector Quantization for Approximate Nearest Neighbor Search

RSLM:面向近似最近邻搜索的无训练向量量化方法
Lenhardt, Rastislav, Dobos, Teodora, Vecchiato, Thomas, Isa, Jiri, Ginzburg, Igor
Abstract
By introducing RSLM (Rotated Scaled Lloyd-Max), a family of training-free vector quantization codecs compressing embeddings to 1--4 bits per dimension, we reduce memory cost and memory bandwidth of a typical large-scale Approximate Nearest Neighbor (ANN) search system, while reducing its complexity and keeping or improving recall across multiple benchmark datasets. State-of-the-art systems filter candidates using coarse partitions, approximately score them to narrow the set, and then rescore the best with higher precision representations (often >=8 bits per dimension). Our relativized codecs can bring this down to 2--4 bits per dimension. We use the properties of the ANN system to encode residual vectors instead of full vectors, both for the approximate scoring phase and the rescoring phase. Since Maximum Inner Product Search (MIPS) is very sensitive to vector norms, we correct the $L_2$ norms of quantized vectors. Our major innovation is that we correct the $L_2$ norm of the final reconstructed vector rather than just the residual. Our rescaling replaces more complicated schemes, such as Anisotropic loss. The residualization scheme gives us a more favorable quality vs size trade-off than generic quantization methods. Our high-performance implementation leverages a block-wise cascaded Fast Walsh-Hadamard Transform (FWHT) with linear-like complexity, AVX SIMD-optimized codebooks, and a steganographic encoding of scaling factors for perfect cache-line alignment.
Chinese Translation
通过引入RSLM(旋转缩放Lloyd-Max,Rotated Scaled Lloyd-Max)——一族无需训练的向量量化编解码器,可将嵌入向量压缩至每维1--4比特——我们降低了典型大规模近似最近邻(Approximate Nearest Neighbor, ANN)搜索系统的内存开销和内存带宽需求,同时降低了系统复杂度,并在多个基准数据集上保持或提升了召回率。最先进的系统首先使用粗划分过滤候选,再通过近似打分缩小候选集,最后使用更高精度的表示(通常每维≥8比特)对最优候选进行重新打分。我们的相对化编解码器可将这一精度降至每维2--4比特。我们利用ANN系统的特性,在近似打分阶段和重新打分阶段均对残差向量而非完整向量进行编码。由于最大内积搜索(Maximum Inner Product Search, MIPS)对向量范数非常敏感,我们对量化向量的$L_2$范数进行了校正。我们的主要创新在于校正最终重构向量的$L_2$范数,而不仅仅是残差的范数。这一重缩放方法取代了各向异性损失(Anisotropic loss)等更复杂的方案。残差化方案相比通用量化方法提供了更优的质量与尺寸权衡。我们的高性能实现利用了具有近似线性复杂度的分块级联快速Walsh-Hadamard变换(Fast Walsh-Hadamard Transform, FWHT)、针对AVX SIMD优化的码本,以及缩放因子的隐写式编码以实现完美的缓存行对齐。
cs.LG / 120 / 2608.30386

DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving

DASC:面向混合线性注意力推理的衰减感知状态压缩
Yu, Yanqi, Sun, Pingwei, Tan, Jianchao, Zhang, Tao, Xie, Yuchen, Cai, Xunliang, Liu, Yao
Abstract
Hybrid linear-attention architectures have recently scaled to large open-weight models, offering quality competitive with full attention while substantially reducing key/value (KV) cache growth. However, their in-place recurrent-state updates complicate cache management: prefix reuse requires state checkpoints alongside full-attention KV, while storing state checkpoints in full increases memory pressure, leading to more evictions and repeated prefill. By analyzing the decay structure of Gated DeltaNet (GDN) and Kimi Delta Attention (KDA), we find that different heads and channels retain prefix information over markedly different timescales, which we term \emph{retention horizons}. This variation suggests substantial compression potential in persistent state checkpoints. Building on this observation, we introduce \emph{Decay-Aware State Compression} (DASC), which derives retention horizons from model weights, selects long-horizon state units, and packs them into a ragged state checkpoint layout. To integrate efficiently with tensor-parallel inference engines, DASC furtherly balances compressed state checkpoints across TP ranks. On reuse, DASC either zero-fills omitted units or refreshes them from a bounded suffix with additional compute cost. Across retrieval and end-to-end reasoning benchmarks on Kimi-Linear, conservative DASC configurations remain close to full caching while compressing KDA recurrent state checkpoints by $2.63\times$. Under fixed state checkpoint memory budgets, the resulting capacity gains reduce mean Time to First Token (TTFT) by 42.6\% and improve input throughput by 68.4\%. At larger compression ratio, suffix refresh recovers much of the accuracy lost to more aggressive omission, at the cost of additional replay computation. Qwen with GDN exhibits a similar quality--efficiency trend, showing that DASC extends from channel-wise KDA to head-wise GDN.
Chinese Translation
混合线性注意力架构近期已扩展至大规模开源权重模型,在大幅减少键/值(KV)缓存增长的同时,能够提供与全注意力机制相当的质量。然而,其原地循环状态更新使缓存管理变得复杂:前缀复用需要在全注意力KV之外保存状态检查点,而完整存储状态检查点会加大内存压力,导致更多的缓存驱逐和重复预填充。通过分析 Gated DeltaNet(GDN)和 Kimi Delta Attention(KDA)的衰减结构,我们发现不同的头和通道在不同时间尺度上保留前缀信息,差异显著,我们将其称为“保留视界”(retention horizons)。这种差异表明持久化状态检查点存在巨大的压缩潜力。基于这一观察,我们提出了“衰减感知状态压缩”(Decay-Aware State Compression,DASC),它从模型权重中推导保留视界,选择长视界的状态单元,并将其打包为不规则的状态检查点布局。为了与张量并行推理引擎高效集成,DASC 还在各 TP 秩之间平衡压缩后的状态检查点。在复用时,DASC 或者对被省略的单元进行零填充,或者通过有界后缀刷新来恢复它们,并付出额外的计算代价。在 Kimi-Linear 上的检索和端到端推理基准测试中,保守的 DASC 配置在将 KDA 循环状态检查点压缩 2.63 倍的同时,仍能保持接近完整缓存的效果。在固定的状态检查点内存预算下,由此带来的容量提升使首词生成时间(TTFT)平均降低 42.6%,输入吞吐量提升 68.4%。在更大的压缩比下,后缀刷新能够恢复因更激进的省略而损失的多数精度,其代价是额外的重放计算。采用 GDN 的 Qwen 模型也表现出类似的质量—效率趋势,表明 DASC 可以从通道级的 KDA 扩展到头级的 GDN。
cs.LG / 121 / 2608.30390

Uncertainty of Vision Medical Foundation Models

视觉医学基础模型的不确定性研究
Huang, Haoxu, Razavian, Narges
Abstract
Accurate uncertainty estimation is essential for machine learning systems de- ployed in high-stakes domains such as medicine. Traditional approaches primarily rely on probability outputs from trained models (point predictions), which provide no formal guarantees on prediction coverage and often require additional calibra- tion techniques to improve reliability. In contrast, conformal prediction (region prediction) offers a principled alternative by generating prediction sets with finite- sample validity guarantees, ensuring that the ground truth is contained within the set at a specified confidence level. In this study, we explore the impact of pre-training approach, dataset scale and domain on both point and region-level uncertainty quantification, by studying domain-specific vision medical foundation models vs. general domain vision foundation models. We conduct a comprehensive evaluation across foundation models trained on retinal, histopathological, and Chest X-Rays data, applying various calibration techniques. Our results demonstrate that (1) pre-training on higher-quality domain-specific datasets along with self-supervised learning leads to better-calibrated point predictions than general domain pre-training, (2) stan- dard re-calibration methods alone cannot fully mitigate uncertainty discrepancies across models trained on different data sources, (3) domain-specific foundation model can lead to more efficient conformal prediction. These findings highlight the importance of careful model selection and the inte- gration of both point and region prediction to enhance the reliability and trust- worthiness of medical AI systems. Our work underscores the need for a holistic approach to uncertainty quantification in recent development of medical vision foundation model, ensuring robust and interpretable AI-driven decision-making.
Chinese Translation
准确的 uncertainty 估计对于部署在医学等高风险领域的机器学习系统至关重要。传统方法主要依赖于训练模型的概率输出(点预测),这种方式无法对预测覆盖率提供正式保证,且通常需要额外的校准技术来提高可靠性。相比之下,保形预测(conformal prediction,区域预测)提供了一种有原则的替代方案,它通过生成具有有限样本有效性保证的预测集合,确保在指定置信水平下真实值包含在该集合中。在本研究中,我们通过对比特定领域的视觉医学基础模型与通用领域的视觉基础模型,探究了预训练方式、数据集规模和领域对点级和区域级不确定性量化的影响。我们在视网膜、组织病理学和胸部X光数据上训练的基础模型上进行了全面评估,并应用了多种校准技术。结果表明:(1)在更高质量的领域特定数据集上进行自监督预训练,比通用领域预训练能产生校准更好的点预测;(2)仅靠标准的重新校准方法无法完全缓解在不同数据源上训练的模型之间的不确定性差异;(3)领域特定基础模型可以实现更高效的保形预测。这些发现凸显了谨慎进行模型选择以及整合点预测与区域预测的重要性,以提升医学人工智能系统的可靠性和可信度。我们的工作强调了在医学视觉基础模型最新发展中采用整体性不确定性量化方法的必要性,以确保稳健且可解释的AI驱动决策。
cs.LG / 122 / 2608.30392

Foundation Models Meet Agriculture: Challenges Beyond Pretraining

基础模型遇上农业:预训练之外的挑战
Nedungadi, Vishal, Xiong, Xingguo, Rußwurm, Marc, Athanasiadis, Ioannis N.
Abstract
Global food security and sustainable climate action increasingly rely on robust, scalable agricultural monitoring. Earth observation foundation models have emerged as powerful, label-efficient tools across general remote sensing domains, yet early attempts to deploy them for agricultural applications have yielded surprisingly poor results. We hypothesize that this performance gap stems from the extreme heterogeneity of agricultural landscapes and the inherent inability of current earth observation foundation models to adapt to task-specific nuances. In this work, we systematically evaluate two critical bottlenecks hindering the deployment of foundation models in agricultural tasks, benchmarking two earth observation foundation models, a foundation model designed for tabular data, and conventional supervised baselines across seven real-world agricultural datasets spanning yield prediction, phenology estimation, and crop classification. First, we identify a pretraining-deployment modality gap: agricultural downstream tasks frequently require diverse, non-imagery data modalities that earth observation foundation models are architecturally unequipped to ingest, while a foundation model built for tabular data handles this heterogeneity more naturally. Second, we formalize the agricultural task space across five structural axes to demonstrate why current models fail to generalize reliably, resulting in highly unstable model rankings across evaluation settings. By characterizing these structural and modal gaps, our insights highlight the friction between general-purpose architectures and specialized agricultural downstream data, providing a strategic roadmap for developing the next generation of domain-aware foundation models.
Chinese Translation
全球粮食安全与可持续气候行动日益依赖于稳健且可扩展的农业监测。地球观测基础模型(Earth observation foundation models)作为高效利用标签的强大工具,已在通用遥感领域展现出卓越能力,然而早期将其应用于农业场景的尝试却收效甚微。我们推测,这一性能差距源于农业景观的极端异质性,以及当前地球观测基础模型难以适应任务特定细微差异的固有能力缺陷。在本工作中,我们系统性地评估了阻碍基础模型在农业任务中部署的两个关键瓶颈:在七个覆盖产量预测、物候估计和作物分类的真实农业数据集上,对两个地球观测基础模型、一个面向表格数据的基础模型以及传统有监督基线模型进行了基准测试。首先,我们识别出一种预训练-部署模态鸿沟:农业下游任务通常需要多样的非影像数据模态,而地球观测基础模型在架构上无法有效处理这些数据;相比之下,专为表格数据构建的基础模型能够更自然地应对这种异质性。其次,我们沿五个结构维度对农业任务空间进行形式化定义,以揭示当前模型为何难以可靠泛化——这导致模型排名在不同评估设置下极不稳定。通过刻画这些结构与模态层面的鸿沟,我们的研究洞见凸显了通用架构与专业化农业下游数据之间的摩擦,为开发下一代领域感知基础模型提供了战略性路线图。
cs.LG / 123 / 2608.30394

TopGQ: Fast GNN Post-Training Quantization Leveraging Topology Information

TopGQ:利用拓扑信息的快速GNN训练后量化方法
Kwon, Dain, Choi, Kanghyun, Lee, Hyeyoon, Park, Sunjong, Lee, Seoyong, Kim, Sukjin, Lee, Jinho
Abstract
Existing GNN quantization methods suffer from considerable quantization overhead, which severely limits their practical usage in real-world scenarios. To this end, we present TopGQ, an accurate post-training GNN quantization framework, alleviating redundant quantization overhead. We propose dual-axis scale absorption, which enables activation quantization along both the outer and inner dimensions by merging one into the adjacency matrix. On top of that, we introduce TopPIN, a proxy for nodes' local structure, and use it to group nodes with similar topology during quantization. Experimental results show that TopGQ reduces quantization time by an order of magnitude while preserving accuracy.
Chinese Translation
现有的GNN量化方法存在相当大的量化开销,严重限制了其在实际场景中的应用。为此,我们提出了TopGQ,一个精确的训练后GNN量化框架,以减少冗余的量化开销。我们提出了双轴尺度吸收方法(dual-axis scale absorption),通过将其中一个维度合并到邻接矩阵中,实现沿外维度和内维度两个方向的激活量化。在此基础上,我们引入了TopPIN作为节点局部结构的代理,并在量化过程中利用它对具有相似拓扑结构的节点进行分组。实验结果表明,TopGQ在保持准确率的同时,将量化时间减少了一个数量级。
cs.LG / 124 / 2608.30406

Locally-Guided Actor-Critic: Training a Goal-conditioned Actor with a Subgoal-aware Critic

局部引导的Actor-Critic:利用子目标感知的Critic训练目标条件Actor
Serris, Olivier, Doncieux, Stéphane, Sigaud, Olivier
Abstract
Goal-conditioned reinforcement learning struggles with long horizons when rewards are sparse. While a planner can provide subgoals to guide a low-level policy, its use at test time may introduce practical subgoal management difficulties. An alternative paradigm utilizes a high-level planner to assist learning, while the policy remains conditioned only on the final goal, enabling planner-free deployment. Among these methods, Reinforcement Learning with Imagined Subgoals (RIS) introduces a regularization term that encourages the policy to take the same actions for the final goal as it does for an intermediate goal. This regularization, however, may lead to goal-chaining issues when intermediate goals are low-dimensional. Potential-based reward shaping (PBRS) translates plans into an additional reward while ensuring that the optimal policy remains unchanged. Yet, it can generate deceptive rewards in terminal states. We study these failure cases and first propose an alternative reward shaping method (RS) that removes these deceptive rewards at the expense of theoretical guarantees of PBRS. Similar to this RS variant, we then propose another method named Locally-Guided Actor Critic (LG-AC) that rewards the agent for reaching intermediate goals. Unlike RS, where intermediate rewards are implicit in the shaping signal, we explicitly condition a value estimator on the full sequence of intermediate goals but represent the value function as a sum of subgoal-conditioned value functions, enabling dense hindsight relabeling. We evaluate all these methods in tasks with challenging goal-chaining requirements and empirically highlight specific cases in which either action regularization or reward shaping yield low performance, while LG-AC achieves the best overall performance across tasks.
Chinese Translation
当奖励稀疏时,目标条件强化学习在长时序任务中面临困难。虽然规划器可以提供子目标来引导低层策略,但在测试时使用规划器可能引入实际的子目标管理难题。另一种范式是利用高层规划器辅助学习,而策略仅以最终目标为条件,从而实现无需规划器的部署。在这些方法中,基于想象子目标的强化学习(Reinforcement Learning with Imagined Subgoals, RIS)引入了一个正则化项,鼓励策略对最终目标采取与对中间目标相同的动作。然而,当中间目标是低维时,这种正则化可能导致目标链接问题。基于势函数的奖励塑形(Potential-based Reward Shaping, PBRS)将规划转化为额外奖励,同时保证最优策略不变。但它可能在终止状态产生欺骗性奖励。我们研究了这些失败情形,首先提出了一种替代性的奖励塑形方法(RS),它消除了这些欺骗性奖励,但代价是失去了PBRS的理论保证。与该RS变体类似,我们随后提出了另一种方法——局部引导的Actor-Critic(Locally-Guided Actor Critic, LG-AC),它对智能体到达中间目标给予奖励。与RS中中间奖励隐含在塑形信号中不同,我们显式地将价值估计器以完整的中间目标序列为条件,但将价值函数表示为以子目标为条件的价值函数之和,从而实现密集的事后重标注。我们在具有挑战性目标链接要求的任务上评估了所有这些方法,并通过实验指出动作正则化或奖励塑形性能低下的具体情形,而LG-AC在所有任务上取得了最佳的整体性能。
cs.LG / 125 / 2608.30417

No Equivariant Architecture Covers All Equivariant Attention

没有一种等变架构能覆盖所有等变注意力
Ông, Tīkun
Abstract
We give a complete characterization of equivariant multi-head self-attention (MHSA): if an MHSA layer is equivariant to a symmetry group $G$, then $G$ can only act by permuting head-clusters, with QK and OV matrices satisfying an equivariance constraint tied to the group action. As a consequence, we prove that any fixed MHSA architecture that achieves exact equivariance by polynomially parameterizing unconstrained MHSA parameters inevitably leads to expressivity loss within the class of equivariant maps: the equivariance locus of unconstrained MHSA forms a union of extremely many Zariski-irreducible components in a reduced parameter space, and any single architecture covers at most one. For $G=D_4$ acting on $C$ copies of the regular representation as the token feature space, we show that there are $\Omega(C^{64})$ components for eight attention heads.
Chinese Translation
我们对等变多头自注意力机制(MHSA)给出了完整的刻画:若一个 MHSA 层对对称群 $G$ 是等变的,则 $G$ 只能通过对头簇(head-clusters)进行置换的方式作用,且 QK 矩阵和 OV 矩阵需满足与该群作用相关联的等变性约束。由此,我们证明:任何通过多项式参数化无约束 MHSA 参数来实现精确等变性的固定 MHSA 架构,都不可避免地导致在等变映射类内的表达能力损失:无约束 MHSA 的等变轨迹(equivariance locus)在约化参数空间中形成由极多个 Zariski 不可约分支构成的并集,而任何单一架构至多只能覆盖其中一个分支。对于作用于 $C$ 份正则表示(作为词元特征空间)的群 $G=D_4$,我们证明在八个注意力头的情形下存在 $\Omega(C^{64})$ 个分支。
cs.LG / 126 / 2608.30442

Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry

混杂伪装成改进:基于129,000例患者登记数据的卒中抗栓治疗离线强化学习系统评估
Rhee, Kihun
Abstract
Recent offline reinforcement learning (RL) studies report policies that outperform physician decisions on clinical outcomes. We conduct a systematic, partially crossed evaluation of five offline RL algorithm families and 14 reward designs in 44,894 post-2018 acute ischemic stroke patients from a nationwide registry (N = 129,033). Standard Fitted Q-Evaluation (FQE) yields an apparent policy-improvement estimate of +0.0069; adding an Early Neurological Deterioration penalty increases it to +0.0101. We identify reward-embedded confounding, in which a proxy terminal reward encodes baseline severity and prognosis as well as treatment efficacy. A 2 x 2 factorial analysis finds that terminal reward confounding accounts for 218.6% of the observed signal change, so its removal overshoots the null. After DML-inspired GBM reward residualization, the FQE estimate attenuates to +0.0033 (p = 0.132), and full deconfounding yields +0.0025 (p = 0.291). FQE-based diagnostics, T-learner analyses, and direct recurrence analyses converge away from a clinically meaningful aggregate improvement. A 1-year mRS factorial analysis replicates the attenuation. We provide an empirically motivated six-step evaluation checklist. NIHSS-stratified heterogeneity is hypothesis-generating for prospective trial design; hospital-level disagreement does not persist after full reward deconfounding.
Chinese Translation
近期离线强化学习(RL)研究报告称,其策略在临床结局上优于医生的决策。我们对来自全国性登记数据(N = 129,033)中44,894例2018年后急性缺血性卒中患者,系统性地开展了部分交叉评估,涵盖五种离线RL算法族和14种奖励设计。标准拟合Q评估(Fitted Q-Evaluation, FQE)得到的表观策略改进估计值为+0.0069;加入早期神经功能恶化(Early Neurological Deterioration)惩罚后,该值升至+0.0101。我们识别出嵌入奖励中的混杂(reward-embedded confounding),即代理终末奖励同时编码了基线严重程度、预后以及治疗效果。2×2析因分析发现,终末奖励混杂可解释观测信号变化的218.6%,因而去除该混杂后会过度校正而越过零效应。经过受双重机器学习(DML)启发的GBM奖励残差化处理后,FQE估计值衰减至+0.0033(p = 0.132),完全去混杂后为+0.0025(p = 0.291)。基于FQE的诊断、T-learner分析以及直接复发分析均一致表明,不存在具有临床意义的总体改进。基于1年mRS的析因分析重复了这一衰减结果。我们提供了一个基于实证的六步评估清单。NIHSS分层异质性可为前瞻性试验设计提供假设;在全奖励去混杂之后,医院层面的分歧不再持续。
cs.LG / 127 / 2608.30449

PRIME: Mitigating Subgroup Optimization Competition in Shared CTR Top Networks with Plug-in Residual Input-Conditioned Mixture of Expert

PRIME:基于即插即用残差输入条件专家混合缓解共享CTR顶层网络中的子群优化竞争
Yao, Heng, Hou, Siyun, Liu, Tianying, Shu, Yulou, He, Yong, Yuan, Chuan, Qiu, Kaibin, Chen, Guowei, Zhao, Jiayu, Yu, Chao, Ding, Ke
Abstract
Click-through rate (CTR) models vary in feature-interaction design, yet their top networks usually remain a single multilayer perceptron shared by all examples. Heterogeneous user, item, and context subgroups therefore update the same parameters; weakly aligned learning signals make the aggregate gradient a compromise among competing directions. We study the competition on Avazu with 4 models and 4 semantic fields. Across all architectures, semantic subgroups show lower Top-NN gradient cosine similarity than random groups matched by sample size and label ratio, with reductions of 0.23-0.37. This competition motivates input-conditioned experts, but directly replacing an established Dense mapping changes its initial function, sharing pattern, and capacity, obscuring the source of gains. We introduce PRIME (Plug-in Residual Input-conditioned Mixture of Experts), a Dense-anchored mixture of low-rank residual experts. PRIME anchors the original prediction and uses zero-residual initialization to match the Dense baseline exactly at training onset. Input-dependent routing weights low-rank experts for example-specific logit corrections; multi-bag aggregation and EMA load biases stabilize conditional estimation. We evaluate PRIME on held-out Avazu and Criteo test sets across 13 CTR architectures and five paired seeds. Median paired AUC gains are +0.0022 and +0.0066, with LogLoss reductions of 0.0011 and 0.0081, respectively. On FiBiNET and DCNv2, PRIME outperforms APG in all ten seed-level AUC comparisons while using fewer parameters and lower inference latency on both backbones. These results show that function-preserving conditional residuals add input-dependent capacity while preserving the Dense path and its optimization stability. Code is available at https://github.com/YH-learning/PRIME.
Chinese Translation
点击率(CTR)模型在特征交互设计上各不相同,但其顶层网络通常仍是由所有样本共享的单一多层感知机。因此,异质的用户、物品和上下文子群更新的是同一组参数;弱对齐的学习信号使得聚合梯度成为相互竞争方向之间的折中。我们在Avazu数据集上使用4个模型和4个语义字段研究了这种竞争。在所有架构上,语义子群的Top-NN梯度余弦相似度均低于按样本量和标签比例匹配的随机组,降幅为0.23-0.37。这种竞争促使我们采用输入条件化的专家,但直接替换已有的Dense映射会改变其初始函数、共享模式和容量,从而掩盖性能提升的来源。我们提出了PRIME(即插即用残差输入条件专家混合,Plug-in Residual Input-conditioned Mixture of Experts),这是一种以Dense为锚点的低秩残差专家混合结构。PRIME以原始预测为锚点,并采用零残差初始化,使其在训练开始时与Dense基线完全一致。基于输入的路由对低秩专家加权,以实现针对样本的logit修正;多包(multi-bag)聚合和EMA负载偏置稳定了条件估计的求解。我们在留出的Avazu和Criteo测试集上、跨13种CTR架构和五组配对随机种子对PRIME进行了评估。配对AUC增益的中位数分别为+0.0022和+0.0066,LogLoss分别降低0.0011和0.0081。在FiBiNET和DCNv2上,PRIME在全部十组种子级AUC对比中均优于APG,同时参数更少、推理延迟更低。这些结果表明,保持函数不变的条�件残差在保留Dense路径及其优化稳定性的同时,增加了依赖输入的容量。代码可在 https://github.com/YH-learning/PRIME 获取。
cs.LG / 128 / 2608.30456

Self-Supervised Pretext Tasks for Infant Cry Analysis: A Controlled Comparison and a Cautionary Result on Donateacry

用于婴儿哭声分析的自监督前缀任务:一项受控比较及关于Donateacry数据集的警示性结果
Simeone, Luigi
Abstract
We compare six self-supervised pretext tasks for infant cry analysis under a fixed budget, meaning the same compact encoder of 1.17M parameters, the same 115 hours of license-verified public pretraining audio, and the same evaluation protocol for every candidate. On cry detection the reconstructive objectives dominate, and a linear probe over a masked-spectrogram encoder reaches 0.988 AUC with subject-wise splits even though the encoder never observed a cry during pretraining. On cry-reason classification over donateacry, the de facto public benchmark for cry reasons, every encoder performs at chance (0.38 to 0.54 macro AUC over 5 classes), and neither domain adaptation on 1.8 hours of real cries nor end-to-end fine-tuning moves the result. Since a frozen HuBERT-base with 80 times more parameters shows the same pattern, the bottleneck must sit in the labels and not in model capacity. We then reproduce the 90\%+ accuracies of the donateacry literature on our own system by changing nothing but the evaluation protocol: clip-wise splits raise accuracy to 85.2% (barely above the 83.8% majority-class baseline), and applying augmentation before splitting raises it to 97.9%, matching the reported state of the art, from the same model that measures 0.49 macro AUC under subject-wise splits. Under leakage-free splits, a twentyfold augmentation of the labeled set (vocoder speaker perturbation and noise mixing, 21 hours) leaves cross-subject AUC unchanged: for this task the effective sample size is the number of infants. We release code, seeds and per-clip license manifests.
Chinese Translation
我们在固定预算条件下比较了六种用于婴儿哭声分析的自监督前缀任务,即所有候选方法均使用同一紧凑编码器(117万参数)、相同的115小时经许可验证的公开预训练音频,以及相同的评估协议。在哭声检测任务上,重建式目标占优:基于掩码频谱图编码器的线性探针在按受试者划分的数据下达到了0.988的AUC,尽管该编码器在预训练期间从未接触过哭声数据。在哭声原因分类任务上,基于事实上的公开基准数据集Donateacry,所有编码器的表现均处于随机水平(5类宏平均AUC为0.38至0.54),且无论在1.8小时真实哭声上进行域适应,还是进行端到端微调,均无法改变这一结果。由于参数量高出80倍的冻结HuBERT-base模型也表现出相同的模式,因此瓶颈必然在于标签本身,而非模型容量。随后,我们在自己的系统上复现了Donateacry文献中报告的90%以上准确率,唯一改变的是评估协议:按音频片段划分使准确率升至85.2%(仅略高于83.8%的多数类基线),而在划分前施加数据增强则使准确率升至97.9%,与已报道的最优水平相当——而同一模型在按受试者划分下的宏平均AUC仅为0.49。在无泄漏划分条件下,对标注集进行20倍的数据增强(声码器说话人扰动与噪声混合,共21小时)并未改变跨受试者AUC:对于该任务,有效样本量即为婴儿的数量。我们公开了代码、随机种子以及逐片段的许可清单。
cs.LG / 129 / 2608.30457

Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry

学习结果发生变化之处:面向多模态几何的信用可寻址推理
Guo, Jiani, Wang, Junjie, Wu, Jie, Zhao, Pengxiang, Zhang, Dongdong, Huang, Shaohan, Yang, Yujiu, Wei, Furu
Abstract
Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step deduction. Existing free-form traces obscure the decisions that determine the answer, and trajectory-level reinforcement learning distributes a single terminal signal across the entire response. We introduce credit-addressable reasoning, in which the semantic units exposed during inference also define where learning compares alternatives and assigns credit. We instantiate this principle with Code-CoT, which retains the diagram, represents visual relations as line-addressable executable code, and organizes reasoning into typed events, and CE-GRPO, which selects event boundaries using structural priors and type-normalized entropy, samples complete continuations from shared prefixes, and converts outcome differences into localized advantages. Across nine geometry benchmarks, CE-GRPO achieves an average accuracy of 76.04, outperforming Qwen3-VL-8B and trajectory-level GRPO by $8.09$ and 3.43 points, respectively. Its relative advantage increases with the number of intermediate events, demonstrating the value of representation--optimization co-design for long, dependency-heavy multimodal reasoning.
Chinese Translation
多模态几何推理要求视觉语言模型(VLM)提取精确的视觉关系,并在多步演绎过程中保持这些关系。现有的自由格式推理轨迹掩盖了决定答案的关键决策,而轨迹级强化学习将单一的终端信号分配到整个响应中。我们提出了信用可寻址推理(credit-addressable reasoning),其中推理过程中暴露的语义单元同时定义了学习中比较备选方案和分配信用的位置。我们通过 Code-CoT 实例化这一原则:该方法保留图形,将视觉关系表示为可按行寻址的可执行代码,并将推理组织为类型化事件;同时提出 CE-GRPO,该方法利用结构先验和类型归一化熵选择事件边界,从共享前缀采样完整的后续内容,并将结果差异转化为局部化优势。在九个几何基准测试中,CE-GRPO 达到了 76.04 的平均准确率,分别以 8.09 和 3.43 个百分点超越 Qwen3-VL-8B 和轨迹级 GRPO。其相对优势随中间事件数量的增加而提升,证明了表示与优化协同设计对于冗长且依赖关系繁重的多模态推理的价值。
cs.LG / 130 / 2608.30472

ToxLens: A Reproducible Graph-Learning Framework for Leakage-Aware, Uncertainty-Calibrated Molecular Toxicity Prediction

ToxLens:一个可复现的图学习框架,用于泄露感知、不确定性校准的分子毒性预测
Strømme, Magnus H., de Sá, Alex G. C., Ascher, David B.
Abstract
Molecular toxicity prediction is increasingly used to prioritise compounds before experimental testing, but conventional benchmark performance can overstate practical utility when structurally related molecules occur across training and test folds. We introduce ToxLens, a reproducible multi-task graph-learning framework for 11 toxicity endpoints spanning Ames mutagenicity, acute oral toxicity, hERG inhibition, and Tox21 nuclear-receptor and stress-response assays. The workflow combines conservative chemical curation, sphere-exclusion filtering, a leakage-aware UMAP-HDBSCAN split, parallel graph and global-feature encoders joined by late concatenation, temperature-scaled Monte Carlo dropout with conformal-style prediction sets, applicability-domain analysis, and SHAP-guided toxicophore discovery with occlusion controls. On the leakage-controlled test fold, a five-seed soft-voting ensemble achieved a Matthews correlation coefficient score of 0.44, an area under the receiver operating characteristic curve score of 0.83, and an area under the precision-recall curve score of 0.58. It exceeded four ECFP4-based shallow baselines on all 11 endpoints under the same split and validation-based threshold-selection protocol. Controlled ablations showed that the global pathway was important, whereas late concatenation outperformed the tested gated and feature-wise linear modulation fusion variants. Conformal-style prediction sets revealed substantial endpoint-specific variation in set efficiency, and discrimination and calibration improved with similarity to the training domain. Retraining on fixed published Tox21 Challenge and TDA folds produced competitive, but not uniformly state-of-the-art, performance. SHAP-guided occlusion and consensus subgraph mining yielded model-derived structural hypotheses, 44 of which contained at least one occurrence that passed the predefined counterfactual criteria.
Chinese Translation
分子毒性预测越来越多地被用于在实验测试前对化合物进行优先级排序,但当结构相关的分子同时出现在训练集和测试集中时,传统的基准性能可能夸大其实际效用。我们提出了ToxLens,一个可复现的多任务图学习框架,涵盖11个毒性终点,包括Ames致突变性、急性口服毒性、hERG抑制以及Tox21核受体和应激反应 assay。该工作流程结合了保守的化学数据整理、球排除过滤、泄露感知的UMAP-HDBSCAN数据划分、通过后期拼接连接的并行图编码器与全局特征编码器、温度缩放的蒙特卡洛dropout结合保形式预测集、适用域分析,以及带有遮挡对照的SHAP引导的毒性药效团发现。在泄露控制的测试集上,五种子软投票集成模型取得了0.44的马修斯相关系数(MCC)、0.83的受试者工作特征曲线下面积(AUROC)和0.58的精确率-召回率曲线下面积(AUPRC)。在相同划分和基于验证的阈值选择协议下,该模型在全部11个终点上均优于四个基于ECFP4的浅层基线。受控消融实验表明全局特征通路非常重要,而后期拼接优于所测试的门控和特征级线性调制(FiLM)融合变体。保形式预测集揭示了集合效率在终点间存在显著差异,且判别能力和校准性能随与训练域相似度的提高而改善。在公开发布的Tox21 Challenge和TDA固定划分上重新训练获得了具有竞争力但并非全面领先的性能。SHAP引导的遮挡和共识子图挖掘产生了源自模型的结构假设,其中44个包含至少一个通过预定义反事实标准的实例。
cs.LG / 131 / 2608.30487

Measuring Memory and Generalization as Separable Geometric Channels: The Topo^2 Framework

将记忆与泛化作为可分离的几何通道进行度量:Topo^2框架
Zhang, Zhanbo, Liu, Ming, Wang, Qing
Abstract
Deep networks trained on noisy labels simultaneously generalize on clean data and memorize flipped labels. These are usually conflated as pressures on one capacity. We present Topo^2, a measurement framework that makes them causally separable, measurable, and law-governed. Persistent-homology H1 structure of the representation space separates into a within-class manifold channel (a function of the training stopping point) and a cross-class channel (a monotone readout of memorized flipped samples). An intervention, the FM0 prescription (zero loss on flipped samples from epoch 0), reaches each setting's generalization ceiling while memorizing essentially nothing. Within the framework we establish a law set with graded evidence: (L2) FM0 separation prescription (9/9); (L1) the within-channel as a training-position function (mid-rise 6/6; convergence-back CIFAR 3/3, SVHN 2/3); (L3) a ring-construction identity (definitional, not a law); and TLS (memory-generalization topological layering): memory is causally additive, anchored (silencing clean collapses the representation), invertible (stripping memory restores near-ceiling generalization), and quantitatively billable (the memorization cost law, effective slope coefficient C ~ 0.38 at the reference capacity: CIFAR-10 0.3801 / SVHN 0.3806 / CIFAR-100 0.384 / VGG 0.3715, capacity-dependent in general and traced to clean-sample feature displacement). We also publish the framework's boundaries: a falsification ledger of nine dead ends, and an instrument-vindication section that excludes six families of global statistics as explanations of the within-channel. The framework turns "memorization" from an ill-defined capacity into a measurable, separable, invertible topological layer.
Chinese Translation
在噪声标签上训练的深度网络会同时在干净数据上泛化并记忆被翻转的标签。这两者通常被混为对同一容量的竞争压力。我们提出Topo^2,一个使二者在因果上可分离、可测量、且服从规律关系的度量框架。表示空间的持续同调(persistent homology)H1结构可分离为类内流形通道(训练停止点的函数)和跨类通道(被记忆翻转样本的单调读出)。我们引入一种干预手段——FM0处方(从第0轮起对翻转样本损失置零)——在几乎不记忆任何翻转样本的情况下达到各设定的泛化上限。在该框架内,我们建立了一套具有分级证据强度的规律集合:(L2) FM0分离处方(9/9);(L1) 类内通道作为训练位置函数(中段上升6/6;收敛回落CIFAR 3/3、SVHN 2/3);(L3) 环构造恒等式(定义性的,非规律);以及TLS(记忆-泛化拓扑分层):记忆在因果上可加、有锚定(沉默干净样本会使表示坍塌)、可逆(剥离记忆可恢复接近上限的泛化),且在数量上可计费(记忆成本定律:在参考容量下有效斜率系数C约0.38:CIFAR-10 0.3801 / SVHN 0.3806 / CIFAR-100 0.384 / VGG 0.3715,一般而言随容量变化,并可追溯至干净样本的特征位移)。我们还公布了该框架的边界:一份包含九条死路的证伪台账,以及一个排除六类全局统计量作为类内通道解释的工具验证章节。该框架将"记忆"从一个定义不清的容量概念,转变为可测量、可分离、可逆的拓扑层。
cs.LG / 132 / 2608.30502

When the Martingale Never Stops Firing: Anytime-Valid Gating on Real Forecast Streams

当鞅永不停火:真实预测流上的 anytime-valid 门控
Han, Weijia, Qu, Lisha
Abstract
Machine learning systems are increasingly corrected while they run, and the decision of when to intervene is increasingly delegated to statistical monitors. Anytime-valid inference promises evidence that can be acted on at any moment, exactly the guarantee this setting needs, and it is moving from theory into deployed monitoring. Conformal test martingales are the change-detection instrument, and Ville's inequality caps their false-alarm probability on exchangeable data. The guarantee is conditional. A deployment inherits it only if the stream it monitors behaves exchangeably. The premise is hardest to satisfy where these monitors are most useful, on dependent data and inside loops where the monitor modifies the learner whose scores it reads. It is also rarely measured. We measure it in a pre-specified case study, where such a monitor gates the online updates of a Kalman adapter correcting frozen time-series foundation models on five forecasting streams. On exchangeable synthetic streams, the same implementation fires in at most 1 of 60 runs. On the real streams, at alpha = 0.05, 135 of 135 clean-stream runs fired. The construction does not explain the firing; the failure comes from the deployed score stream itself. Repeated fires hold the gate's drift response active, and the gated filter amplifies the very transient it was designed to prevent. The component worth keeping makes no validity claim. Huber-style gating of the filter's own updates cuts isolated-spike degradation by an order of magnitude with no dataset specific tuning. Anytime-valid methods proposed for dependent data should therefore be accompanied by null-calibration controls and mechanism traces.
Chinese Translation
机器学习系统越来越多地在运行过程中被纠正,而何时进行干预的决策也越来越多地被交给统计监控器。Anytime-valid(任意时点有效)推断承诺提供可在任意时刻付诸行动的证据——这正是该场景所需要的保证——并且它正从理论走向实际部署的监控系统中。Conformal test martingales(保形检验鞅)是变化检测的工具,Ville 不等式在可交换数据上限制了其误报概率。但该保证是有条件的:只有当被监控的数据流满足可交换性时,部署的系统才能继承这一保证。这一前提恰恰在这些监控器最有用武之地的场合最难满足——即在相关数据上,以及在监控器会修改其所读取分数的学习器的反馈回路中。而且这一前提也很少被实际测量。我们在一个预先设定的案例研究中对此进行了测量:该监控器对一个 Kalman 适配器的在线更新进行门控,该适配器在五条预测数据流上对冻结的时序基础模型进行纠正。在可交换的合成数据流上,同一实现在 60 次运行中至多触发 1 次。而在真实数据流上,当 alpha = 0.05 时,135 次干净数据流运行中有 135 次全部触发。该构造本身并不能解释触发;失败源于部署的分数流本身。反复触发使门控的漂移响应保持激活状态,而被门控的滤波器反而放大了它本要防止的瞬态。其中值得保留的组件并不做任何有效性声明:对滤波器自身更新采用 Huber 式门控,无需任何针对数据集的调参,就能将孤立尖峰造成的性能退化降低一个数量级。因此,针对相关数据提出的 anytime-valid 方法应当配备原假设校准控制和机制追踪。
cs.LG / 133 / 2608.30505

Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability

面向语言模型的张量方法:从词元表示到训练、适配、推理、压缩与可解释性
Tarasov, Matvei, Ahmadi-Asl, Salman, de Almeida, Andre L. F., Cichocki, Andrzej
Abstract
Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is underexploited by the conventional matrix-centric view. Tensor decompositions and tensor networks provide a principled algebraic language for this structure, yet the literature often treats them as isolated compression mechanisms. This survey organizes tensor methods for LLMs through two complementary views: a seven-stage lifecycle taxonomy covering tokenization, embeddings, pre-training, adaptation, compression, inference, and interpretability, and a component view covering embeddings, attention, and feed-forward networks. We provide unified notation and theoretical foundations, analyze tensorization strategies for individual Transformer components, and compare methods at each lifecycle stage while making differences in evaluation protocols and model scales explicit. We further connect tensor methods to neighboring efficiency techniques and probabilistic tensor networks. Finally, we synthesize open challenges and introduce $\rho_{\rm gap}$, a metric for the compression-realization gap between theoretical memory reduction and measured system-level speedup. By treating tensorization as a common structural principle, the survey provides a structured entry point to tensorized language models and clarifies when parameter savings can plausibly translate into memory efficiency, computational efficiency, or interpretability. The GitHub page dedicated to this paper is accessible at \href{https://github.com/ma-tt-a/awesome-tensor-methods-for-llms}{this https URL}.
Chinese Translation
大语言模型(LLMs)由词元表示、权重、适配更新、缓存和激活等结构化高维对象构成,而传统的以矩阵为中心的视角未能充分利用其多重线性结构。张量分解与张量网络为刻画这种结构提供了一种有原则的代数语言,但现有文献往往将它们视为孤立的压缩机制。本综述通过两个互补的视角来组织面向LLMs的张量方法:其一是涵盖词元化、嵌入、预训练、适配、压缩、推理和可解释性的七阶段生命周期分类法;其二是涵盖嵌入、注意力和前馈网络的组件视角。我们提供了统一的符号体系与理论基础,分析了针对各个Transformer组件的张量化策略,并比较了生命周期各阶段的方法,同时明确指出了评估协议和模型规模上的差异。我们进一步将张量方法与邻近的高效化技术以及概率张量网络联系起来。最后,我们综合了尚未解决的开放性挑战,并引入了 $\rho_{\rm gap}$ 这一指标,用以衡量理论内存缩减与实测系统级加速之间的压缩-实现差距。通过将张量化视为一种共同的结构性原则,本综述为张量化语言模型提供了一个结构化的入门路径,并阐明了参数节省在何种情况下有望转化为内存效率、计算效率或可解释性。本文的专属GitHub页面可访问 \href{https://github.com/ma-tt-a/awesome-tensor-methods-for-llms}{此链接}。
cs.LG / 134 / 2608.30512

Trajectory-Initialized Neural Double Q-Routing for Large-Scale Overhead Hoist Transport Systems

面向大规模天车搬运系统的轨迹初始化神经双重Q路由
Gu, Cheng, Zhao, Qiusheng, Liu, Anbang, Lin, Shaochong, Shen, Max Z. J.
Abstract
Large-scale industrial robot fleets share constrained physical infrastructure, making vehicle travel times dependent on safety separation, intersection access, downstream blocking, and station contention. We study this problem in overhead hoist transport (OHT) systems, a representative ceiling-mounted material-handling system used in semiconductor fabs. Static shortest-path routing cannot account for these time-varying traffic costs, whereas tabular Q-routing adapts online but learns each destination--node--action value independently, limiting information sharing across sparsely visited routing contexts and making startup behavior sensitive to inaccurate value estimates. We propose Neural Double Q-routing, which replaces destination-indexed tables with a shared state--action value network. The network is warm-started through return-to-go regression on mixed simulator-generated routing trajectories and then refined online using Double-Q updates, local congestion correction, and event-stratified structured replay. Across nine matched fleet-size--arrival-rate settings with 100, 150, and 200 OHTs, the proposed framework reduces mean completion time relative to tabular Double Q-routing by $0.8\%$--$8.8\%$. It achieves the lowest mean completion time among all compared methods in the six 150- and 200-OHT settings, whereas Dijkstra remains best in the three 100-OHT settings. Completed-task counts remain within $1\%$ of tabular Double Q-routing in eight of nine settings, and 95th-percentile completion time decreases in eight settings. In two matched startup scenarios, offline initialization increases the number of completed tasks by up to $23\%$ and reduces tail completion time by up to $15\%$.
Chinese Translation
大规模工业机器人车队共享受限的物理基础设施,使车辆行驶时间依赖于安全间隔、交叉口通行权、下游阻塞和站点竞争。我们在天车搬运(Overhead Hoist Transport, OHT)系统中研究这一问题,该系统是半导体晶圆厂中一种具有代表性的天花板悬挂式物料搬运系统。静态最短路径路由无法考虑这些时变的交通代价,而表格型Q路由虽能在线自适应,但其对每个目的地—节点—动作价值独立学习,限制了在稀疏访问的路由情境之间共享信息,并使启动阶段的行为对不准确的价值估计十分敏感。我们提出神经双重Q路由(Neural Double Q-routing),用一个共享的状态—动作价值网络取代按目的地索引的表格。该网络通过对模拟器生成的混合路由轨迹进行回报目标(return-to-go)回归实现热启动,随后采用双重Q(Double-Q)更新、局部拥堵校正和事件分层的结构化经验回放在线精化。在包含100、150和200台OHT的九组匹配的车队规模—到达率设置中,相对于表格型双重Q路由,所提出的框架将平均完成时间降低了0.8%至8.8%。在六个包含150和200台OHT的设置中,该方法在所有对比方法中取得了最低的平均完成时间,而Dijkstra算法在三个包含100台OHT的设置中仍表现最佳。在九个设置中的八个里,完成任务数量与表格型双重Q路由的差距保持在1%以内,且在八个设置中第95百分位完成时间有所下降。在两个匹配的启动场景中,离线初始化使完成任务数量最多提升23%,并使尾部完成时间最多降低15%。
cs.LG / 135 / 2608.30528

PAC: Progress-Augmented Advantage Curriculum for Multi-Task Reinforcement Learning of LLMs

PAC:面向大语言模型多任务强化学习的进度增强优势课程学习方法
Yu, Yuanqiang, Zheng, Yanzhao, Zhang, Zhentao, Xu, Tianze, Ma, Chao, Zhu, Jihuai, Liu, Jiashun, Deng, Xinle, Dong, Baohua, Zhu, Hangcheng, Huang, Ruohui
Abstract
Reinforcement learning (RL) is used to improve the reasoning abilities of LLMs, while training data span heterogeneous tasks. However, most RL post-training pipelines rely on fixed or manually designed task mixtures, even though task usefulness changes as training progresses. Online curriculum methods often define learnability by update magnitude, ignoring whether the update translates into reward gains, which can misallocate rollout budget toward tasks with large but ineffective updates. We propose PAC, a Progress-Augmented Advantage Curriculum for multi-task RL of LLMs that combines two task-level signals: advantage-derived learnability, which measures the magnitude of the policy update a task can induce, and recent reward gains, which show whether those updates have improved task performance. A Bayesian Thompson Sampling controller uses these signals to allocate rollouts across tasks during GRPO training. We evaluate PAC under two settings: a multi-level reasoning setting and a multi-domain reasoning setting. PAC improves sample efficiency and final performance: it reaches comparable validation scores with fewer rollout steps and achieves higher final averages than random sampling and advantage-based curriculum baselines in both settings. These results show that jointly tracking advantage signals and actual reward gains yields an effective online curriculum for LLM post-training.
Chinese Translation
强化学习(RL)被用于提升大语言模型(LLM)的推理能力,而训练数据往往涵盖异构任务。然而,大多数RL后训练流程依赖于固定或人工设计的任务混合比例,尽管任务的有用性会随着训练进程而变化。在线课程学习方法通常以更新幅度来定义可学习性,忽略了该更新是否能转化为奖励提升,这可能导致将 rollout 预算错误地分配给更新幅度大但效果不佳的任务。我们提出了 PAC(Progress-Augmented Advantage Curriculum),一种用于LLM多任务强化学习的进度增强优势课程学习方法,它结合了两个任务层面的信号:一是基于优势函数的可学习性,用于衡量一个任务所能引起的策略更新幅度;二是近期的奖励增益,用于反映这些更新是否确实提升了任务性能。贝叶斯汤普森采样(Thompson Sampling)控制器利用这些信号在GRPO训练过程中跨任务分配 rollout。我们在两种设置下评估PAC:多层级推理设置和多领域推理设置。PAC提升了样本效率和最终性能:在两种设置中,它均以更少的 rollout 步数达到与基线相当的验证分数,并取得高于随机采样和基于优势的课程学习基线的最终平均分。这些结果表明,联合追踪优势信号与实际奖励增益能够为LLM后训练构建有效的在线课程。
cs.LG / 136 / 2608.30564

Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMs

Q-Strata:面向混合专家大语言模型混合精度量化的层次化比特分配方法
Lee, Deokjae, Chu, Sihun, Song, Hyun Oh
Abstract
Mixed-precision quantization (MPQ) assigns a different bitwidth to each linear layer of a large language model (LLM) to minimize the quantization-induced quality loss under a fixed budget, but Mixture-of-Experts (MoE) models contain these layers in every expert of every MoE block, so the allocation space grows far larger than in a dense model. Existing methods either allocate within each block under a uniform per-block budget, or allocate across blocks through an additive proxy, and neither directly optimizes a model-level objective over the choices that couple the blocks. We propose Q-Strata, a bi-level allocator that ranks within-block assignments with a cheap proxy and allocates across blocks with a model-level objective evaluated on the assembled quantized model. Its inner stage caches a Pareto frontier of candidates per block over finely spaced budgets, leaving the outer stage to set one budget per block instead of a bitwidth for every linear layer. With the search reduced to one budget per block, the outer stage optimizes this model-level objective directly, capturing the inter-block coupling that additive proxies miss. On Mixtral-8x7B-Instruct, Qwen1.5-MoE-A2.7B, and DeepSeek-V2-Lite, Q-Strata consistently achieves lower WikiText2 perplexity than uniform-bitwidth GPTQ and the state-of-the-art MoE MPQ methods MxMoE and GEMQ in the low-bit regime. The code is available at https://github.com/snu-mllab/Q-Strata/tree/main.
Chinese Translation
混合精度量化(Mixed-Precision Quantization, MPQ)通过为大语言模型(LLM)的每个线性层分配不同的比特宽度,在固定预算下最小化量化带来的质量损失。然而,混合专家模型在每个 MoE 块的每个专家中都包含此类线性层,因此其分配空间远大于稠密模型。现有方法要么在各 MoE 块内基于统一的块级预算进行分配,要么通过可加代理目标跨块分配,二者均未直接优化能够耦合各块的模型级目标。我们提出 Q-Strata,一种双层分配器:其内层使用低开销的代理指标对块内分配进行排序,外层则在组装好的量化模型上评估模型级目标并跨块进行分配。其内层阶段为每个块在细粒度预算下缓存候选方案的帕累托前沿(Pareto frontier),使外层阶段只需为每个块设定一个预算,而无需为每个线性层逐一分配比特宽度。由于搜索空间被缩减为每块一个预算,外层阶段可以直接优化该模型级目标,从而捕捉可加代理目标所忽略的块间耦合。在 Mixtral-8x7B-Instruct、Qwen1.5-MoE-A2.7B 和 DeepSeek-V2-Lite 上,Q-Strata 在低比特 regime 下始终取得比统一比特宽度的 GPTQ 以及最先进的 MoE 混合精度量化方法 MxMoE 和 GEMQ 更低的 WikiText2 困惑度。代码已发布于 https://github.com/snu-mllab/Q-Strata/tree/main。
cs.LG / 137 / 2608.30568

Collapsibility of Performance Metrics in Clinical Predictive AI

临床预测型人工智能中性能指标的可折叠性
Matos, João, Van Calster, Ben, Riley, Richard D., Dhiman, Paula, Collins, Gary S.
Abstract
Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness evaluations commonly rely on performance analyses across subgroups. However, some performance metrics are non-collapsible, meaning that the overall population performance value does not equal the weighted average of subgroup specific values. Objective: To examine the collapsibility properties of commonly reported performance metrics in predictive AI, with a focus on the area under the receiver operating characteristic curve (AUC, also known as c-statistic). Methods: We investigate the collapsibility of 15 performance metrics, either by expressing each metric as a linear combination of its stratum specific values or, where non-collapsible, by providing a counterexample inspired by Simpson's paradox as a formal disproof. Results: Five performance metrics (AUC, calibration intercept, calibration slope, expected calibration error, and Nagelkerke R^2) are shown to be non-collapsible, and ten (O:E ratio, logloss, Brier score, accuracy, F1-score, true positive rate, true negative rate, positive predictive value, negative predictive value, and net benefit) are shown to be collapsible. The AUC is shown to be non-collapsible because it decomposes into within- and cross-group AUC terms when subpopulations coexist, such that its overall value may fall outside the range of subgroup specific AUCs. Conclusions: Non-collapsibility of performance metrics has important consequences for reporting, model appraisal, and fairness evaluation. It can generate spurious differences between subgroup and overall performance, which may mislead fairness evaluations. Explicitly acknowledging and reporting the collapsibility properties of performance metrics improves both the interpretability and transparency of fairness assessments.
Chinese Translation
背景:针对预测型人工智能(AI)的总体层面评估可能掩盖不同亚组之间的性能差异。公平性评估通常依赖于跨亚组的性能分析。然而,某些性能指标是不可折叠的,即总体人群的性能值并不等于各亚组特定性能值的加权平均。目的:考察预测型AI中常用报告的性能指标的可折叠性特性,重点关注受试者工作特征曲线下面积(AUC,亦称c统计量)。方法:我们研究了15个性能指标的可折叠性,方法是将每个指标表示为其各层特定值的线性组合;对于不可折叠的指标,则通过一个受辛普森悖论启发的反例作为正式反证。结果:五个性能指标(AUC、校准截距、校准斜率、期望校准误差和Nagelkerke R^2)被证明是不可折叠的,而十个指标(O:E比值、对数损失、Brier评分、准确率、F1分数、真阳性率、真阴性率、阳性预测值、阴性预测值和净收益)被证明是可折叠的。AUC被证明不可折叠,其原因在于当多个亚人群共存时,AUC可分解为组内和跨组AUC项,因此其总体值可能落在各亚组特定AUC范围之外。结论:性能指标的不可折叠性对结果报告、模型评价和公平性评估具有重要影响。它可能在亚组性能与总体性能之间产生虚假差异,从而误导公平性评估。明确承认并报告性能指标的可折叠性特性,有助于提高公平性评估的可解释性和透明度。
cs.LG / 138 / 2608.30585

The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal

角色扮演越狱中的安全接力:面向危害识别与拒绝的组件分辨因果分析
Chowdhury, Md Mokarram, Chang, Ernie, Li, Yang
Abstract
Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model would ordinarily reject. Roleplay jailbreaks are especially concerning: the harmful request can remain visible inside a roleplay wrapper made of a persona, scenario, and task, yet the model may comply. We use mechanistic interpretability to determine how this context reverses refusal and which elements contribute to the reversal. Across two benchmarks, three model families, and four authored wrappers, we compare matched harmful and benign requests with and without this wrapper. We trace hidden-state contrasts from the request to the final prompt state, isolate wrapper operations through controlled counterfactuals, intervene on their activation directions in held-out evaluation requests, and decompose effective directions geometrically. Our analysis yields three findings. (1) Successful attacks retain the measured harmful-versus-benign distinction at the request, while its refusal-associated expression weakens where the answer begins, a pattern we call safety-relay attenuation. (2) Constructing the complete roleplay around the request and framing it within the scenario contribute causally: removing the associated activation changes restores refusal. (3) These effects largely share internal structure, and most repair is reproduced by components aligned with the model's ordinary refusal of harmful requests without roleplay; scenario framing retains a smaller, model-dependent component. Together, these findings explain how roleplay can produce compliance despite retained evidence of harm and identify a concrete target for future safeguards: maintaining the connection from harm recognition to refusal.
Chinese Translation
大语言模型经过训练,既要遵循指令又要拒绝有害请求。越狱(jailbreak)攻击利用这一平衡来诱使模型输出其通常会拒绝的内容。角色扮演类越狱尤其令人担忧:有害请求可能仍然清晰可见地位于由角色人设、场景和任务构成的角色扮演包装之中,但模型却可能选择服从。我们利用机制可解释性方法来确定这种上下文是如何逆转拒绝行为的,以及哪些成分促成了这一逆转。我们在两个基准、三个模型系列和四种自行构建的包装上,对匹配的有害与良性请求在有/无该包装的情况下进行比较。我们追踪从请求到最终提示状态的隐状态对比,通过受控反事实隔离包装操作,在留出的评估请求上对其激活方向进行干预,并从几何角度分解有效方向。我们的分析得出三点发现。(1)成功的攻击在请求层面保留了可测量的有害与良性区分,但在答案开始处其与拒绝相关的表达被削弱,我们将这一模式称为“安全接力衰减”(safety-relay attenuation)。(2)在请求周围构建完整的角色扮演以及将其置于场景框架之中具有因果贡献:移除相关的激活变化即可恢复拒绝行为。(3)这些效应在很大程度上共享内部结构,且大部分修复由与模型在无角色扮演时对有害请求的正常拒绝相一致的成分所再现;场景框架仅保留较小的、依赖模型自身的成分。综合来看,这些发现解释了角色扮演为何能在危害证据仍然被保留的情况下产生服从行为,并为未来的安全防护指明了一个具体目标:维持从危害识别到拒绝行为的连接。
cs.LG / 139 / 2608.30593

State of Health Estimation using Convolutional and Bidirectional LSTM Neural Networks tuned by Bayesian Optimization

基于贝叶斯优化调优的卷积与双向LSTM神经网络的电池健康状态估计
Eleftheriadis, Panagiotis, Kyrgios, Foivos Georgios, Leva, Sonia
Abstract
In this research, a novel framework is proposed for the SOH estimation, which employs a hybrid deep learning architecture of a concatenation of a Convolution Neural Network (CNN) and a Bidirectional Long Short-Term Memory (BiLSTM) Neural Network (NN) with the integration of Bayesian Optimization-based hyperparameter tuning for the network. Three different deep learning architectures are being evaluated: standalone recurrent models, CNN-RNN architectures and CNN-RNN combinations enhanced with intermediate Fully Connected (FC) layers. Among the three, the model with the intermediate FC layers demonstrated the highest predictive accuracy. A comprehensive feature engineering approach combines capacity (Q), voltage (V), Incremental Capacity Analysis (ICA), and Differential Voltage Analysis (DVA), with systematic evaluation of multiple combinations to identify the optimal input representation. To validate the proposed method, three publicly available datasets were utilized, ensuring reproducibility of the results, two from external sources and one developed by the author of this study using a unique experimental setup. The comparison study was performed using the Mean Absolute Error (MAE), the Root Mean Squared Error (RMSE) and the FLoating-point OPerations (FLOPs) as evaluation metrics.
Chinese Translation
本研究提出了一种用于健康状态(SOH)估计的新型框架,该框架采用卷积神经网络(CNN)与双向长短期记忆(BiLSTM)神经网络相级联的混合深度学习架构,并结合基于贝叶斯优化的网络超参数调优。研究评估了三种不同的深度学习架构:独立的循环神经网络模型、CNN-RNN架构以及通过中间全连接(FC)层增强的CNN-RNN组合。在三种架构中,含中间全连接层的模型表现出最高的预测精度。本研究采用综合特征工程方法,将容量(Q)、电压(V)、增量容量分析(ICA)和微分电压分析(DVA)相结合,并通过系统评估多种组合以确定最优输入表示。为验证所提出的方法,研究使用了三个公开可用的数据集,以确保结果的可重复性,其中两个来自外部来源,另一个由本文作者利用独特的实验装置构建。对比研究采用平均绝对误差(MAE)、均方根误差(RMSE)和浮点运算次数(FLOPs)作为评估指标。
cs.LG / 140 / 2608.30597

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

PLC-DPO:噪声与模糊偏好优化中的后验标签校正
Cho, Boryeong, Ahn, Sumyeong, Yun, Se-Young
Abstract
Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples. Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs. 55.5 for the next-best method). Injected-noise and tie stress tests, human disagreement analysis, and self-confirmation diagnostics further show that the routing remains stable and distinguishes flipped from weakly directional pairs.
Chinese Translation
直接偏好优化(DPO)通过成对比较简化了对齐过程,但其假设所有观测到的偏好都是可靠的。真实数据往往违背这一假设,产生反转、微弱或模糊的标签,从而导致有害的策略更新。为解决这一问题,我们提出了后验标签校正DPO(PLC-DPO),通过将每对样本的训练信号判定为干净(clean)、反转(flip)或平局(tie)三种情形,实现鲁棒的偏好优化。其核心思想是利用校准后的策略-参考模型边距作为在线证据,采取相应的校正行动。这将噪声偏好学习重新定义为主动校正监督的方向与强度,而非仅仅过滤可疑样本。在57个数据集-模型-基准组合中,PLC-DPO取得了最佳平均胜率(相对DPO为60.5,而次优方法为55.5)。注入噪声与平局压力测试、人类标注分歧分析以及自我确认诊断进一步表明,该判定机制保持稳定,并能有效区分反转样本与方向微弱的样本。
cs.LG / 141 / 2608.30636

MolLedger: An Additive Graph Neural Network with Chemically Grounded ADME Attributions

MolLedger:一种具有化学依据ADME归因的加性图神经网络
Ji, Christina X.
Abstract
Optimizing absorption, distribution, metabolism, and excretion (ADME) is an important part of small molecule drug discovery. Many machine learning models have been built to predict ADME properties to facilitate this optimization process, but explaining model predictions is challenging. We propose a new graph neural network architecture with built-in meaningful per-atom attributions. Our model MolLedger outputs predictions that are the sum of per-atom scores. MolLedger's additive framework obtains exact interpretability at no cost to performance because the global context vector gives the additive head enough context to produce good per-atom scores. Furthermore, MolLedger produces attributions that are more faithful to chemical properties than other interpretability methods because the auxiliary loss in MolLedger anchors the atom scores to chemical properties. Our case studies comparing interpretations from multiple methods on molecular pairs reveal that MolLedger is much better at producing sensible explanations for predicted property changes.
Chinese Translation
优化药物的吸收、分布、代谢和排泄(ADME)性质是小分子药物发现中的重要环节。许多机器学习模型已被构建用于预测ADME性质以促进这一优化过程,但解释模型预测结果仍具挑战性。我们提出了一种新的图神经网络架构,其内置具有意义的逐原子归因。我们的模型MolLedger输出的预测结果是逐原子分数之和。MolLedger的加性框架在不损失性能的情况下实现了精确的可解释性,因为全局上下文向量为加性预测头提供了足够的上下文信息,从而生成良好的逐原子分数。此外,MolLedger产生的归因比其他可解释性方法更忠实于化学性质,因为MolLedger中的辅助损失将原子分数锚定在化学性质上。我们对分子对上多种方法的解释结果进行的案例研究表明,MolLedger在为预测的性质变化提供合理解释方面表现显著更优。
cs.LG / 142 / 2608.30640

Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RL

一次三步:在对比强化学习中从动作序列学习表征
Korniak, Michal, Dybek, Kamil, Eysenbach, Benjamin, Bagatella, Marco, Bortkiewicz, Michał
Abstract
While self-supervised approaches to reinforcement learning have achieved strong results by learning representations of states and actions, a key open question is the time scale over which actions should be modeled. Departing from the standard formulation relying on single-step actions, we extend contrastive reinforcement learning (CRL), a prototypical self-supervised method, to operate over action chunks, and find that this results in large, pervasive gains across established offline and online benchmarks: +31.7% and +93.1% across 18 and 11 environments respectively. While action-chunking-driven gains are generally explained through the ability to model non-Markovian, temporally extended policies, and to propagate unbiased multi-step returns, interestingly, we find that these arguments only partially apply to CRL. Our empirical studies suggest that, in the context of CRL, an action chunk carries more information about the goal than a single action, measurably improving the critic's representations, and rendering the algorithm significantly more effective.
Chinese Translation
尽管自监督强化学习方法通过学习状态和动作的表征已取得了优异的成果,但一个关键的开放性问题是动作建模应采用何种时间尺度。不同于依赖单步动作的标准形式,我们将对比强化学习(Contrastive Reinforcement Learning, CRL)这一典型的自监督方法扩展到动作块(action chunk)上,并发现这在已有的离线和在线基准测试中带来了大幅且普遍的提升:分别在18个和11个环境中提升31.7%和93.1%。虽然动作分块带来的收益通常被解释为能够建模非马尔可夫的时序扩展策略,以及能够传播无偏的多步回报,但有趣的是,我们发现这些论据仅部分适用于CRL。我们的实证研究表明,在CRL的背景下,一个动作块比单个动作携带更多关于目标的信息,可显著改善评论家(critic)的表征质量,从而使算法更加有效。
cs.LG / 143 / 2608.30654

Season-Aware Hybrid Convolutional-Transformer for Antarctic Sea Ice Concentration Forecasting

面向南极海冰浓度预报的季节感知混合卷积-Transformer模型
Li, Danyang, Taylor, John, Bui, Thang, Deng, Quanling
Abstract
Antarctic sea ice concentration (SIC) forecasting is an important yet challenging task due to the coexistence of complex spatial structure, long-range temporal dependencies, and strong seasonal variability. Conventional convolution-based models are effective at capturing local spatial patterns, but often have limited ability to model long-term temporal evolution. To address these challenges, we build on a hybrid Convolutional-Transformer forecasting framework for monthly Antarctic SIC forecasting. This framework combines convolutional encoding for spatial feature extraction with factorised self-attention for spatio-temporal dependency modelling. We further introduce two seasonal prior mechanisms: a month-aware positional encoding that injects calendar-month information into the token representation, and a seasonal temporal bias that encourages attention to periodically related historical states. Experimental results show that the proposed framework achieves better performance than convolutional and recurrent baselines across both classification and regression metrics. Ablation studies further indicate that the seasonal prior mechanisms provide consistent additional gains in both short- and long-horizon prediction. These results demonstrate the value of combining convolutional structures, attention mechanisms, and periodic prior information for Antarctic SIC forecasting.
Chinese Translation
南极海冰浓度(SIC)预报是一项重要而具有挑战性的任务,其原因在于复杂的空间结构、长程时间依赖性与强烈的季节变率并存。传统的基于卷积的模型能够有效捕捉局部空间模式,但在建模长期时间演化方面能力有限。为应对这些挑战,我们在一个混合卷积-Transformer(Convolutional-Transformer)预报框架的基础上,开展南极月尺度SIC预报研究。该框架将用于空间特征提取的卷积编码与用于时空依赖建模的分解式自注意力机制相结合。我们进一步引入了两种季节先验机制:一是月份感知的位置编码,将日历月份信息注入到词元表示中;二是季节性时间偏置,促使注意力关注周期性相关的历史状态。实验结果表明,所提出的框架在分类和回归指标上均优于卷积和循环神经网络基线模型。消融实验进一步表明,季节先验机制在短期和长期预测中均能带来一致的性能提升。这些结果证明了将卷积结构、注意力机制与周期性先验信息相结合用于南极SIC预报的价值。
cs.LG / 144 / 2608.30674

CoMPASS: Collaborative Molecular Property Prediction via Adaptive Small-Large Model Synergy

CoMPASS:基于自适应大小模型协同的分子性质 collaborative 预测
Li, Wentao, Qiu, Jiangjie, Li, Yijun, Zhao, Leyi, Wang, Xiaonan
Abstract
Accurate molecular property prediction requires both statistical reliability and chemical reasoning. Graph neural networks can be calibrated directly on labeled assays but remain limited by the coverage of their training data. Large language models (LLMs) can compare molecular evidence and articulate chemical rationales, yet are unreliable as standalone quantitative predictors. The central challenge is therefore to determine when an LLM should influence a calibrated model and by how much. Here we present CoMPASS, a retrieval-calibrated framework for small-large model collaboration. CoMPASS retains a graph attention network (GAT) as the predictive anchor, retrieves locally relevant training molecules, provides attention-grounded evidence to an LLM, and converts its proposal into a bounded correction through an agreement-aware gate. Across six classification and two regression benchmarks, CoMPASS improves the GAT anchor in regions of correctable uncertainty while limiting LLM intervention in high-confidence regimes. Ablations show that the gains arise from validation-calibrated retrieval and bounded fusion rather than prompting alone. These results suggest that generative reasoning should augment calibrated prediction through evidence-grounded, controlled corrections rather than direct output replacement. Code is available at https://github.com/littlepeachs/CoMPASS.
Chinese Translation
准确的分子性质预测既需要统计可靠性,也需要化学推理能力。图神经网络可以直接在标注实验数据上进行校准,但其性能受限于训练数据的覆盖范围。大语言模型(LLM)能够比较分子证据并阐述化学依据,但作为独立的定量预测器并不可靠。因此,核心挑战在于确定大语言模型何时应对已校准模型施加影响,以及影响程度的大小。本文提出CoMPASS,一种面向大小模型协同的检索校准框架。CoMPASS保留图注意力网络(GAT)作为预测锚点,检索局部相关的训练分子,向大语言模型提供基于注意力机制的证据,并通过一致性感知门控将其建议转化为有界修正。在六个分类基准和两个回归基准上,CoMPASS在可修正的不确定性区域提升了GAT锚点的性能,同时在模型高置信度区域限制了大语言模型的干预。消融实验表明,性能提升源于经过验证校准的检索和有界融合,而非仅靠提示工程。这些结果表明,生成式推理应通过基于证据的、受控的修正来增强已校准的预测,而非直接替换其输出。代码见 https://github.com/littlepeachs/CoMPASS。
cs.LG / 145 / 2608.30682

Learning Materials Properties from Scarce Labels and Unlabeled Crystals

从稀缺标签与无标签晶体中学习材料性质
Li, Wentao, Chen, Yizhe, Qiu, Jiangjie, Li, Yijun, Zhao, Leyi, Wang, Xiaonan
Abstract
Learning materials properties from scarce labels and unlabeled crystals is a central challenge for data-driven materials discovery. We present SemiMat, a controlled benchmark for semi-supervised materials property regression, and MatRank, a reliability-weighted objective for continuous pseudo-label uncertainty. SemiMat fixes labeled and unlabeled crystal inputs, graph-backbone interfaces, validation-only checkpoint selection, held-out test reporting, normalized MAE (NMAE), and method-rank summaries across six scarce-label tasks, four graph backbones, and five predefined split runs. MatRank builds pseudo-targets from labeled anchors, weights them by local reliability and weak-prediction agreement, trains weak and strong graph views consistently, and adds ranking signals so that unlabeled crystals shape both values and candidate order. Across the retained 24 backbone-task blocks, one fixed MatRank objective gives the lowest aggregate held-out test NMAE (0.896) and best average method rank (2.208). The component, OOD, and generated-pool diagnostics identify where the gain is reliable and where further screening evaluation remains necessary. Code is available at https://github.com/littlepeachs/SemiMat.
Chinese Translation
从稀缺标签和无标签晶体中学习材料性质是数据驱动材料发现的核心挑战。我们提出了SemiMat,一个用于半监督材料性质回归的受控基准测试平台,以及MatRank,一种针对连续伪标签不确定性的可靠性加权目标函数。SemiMat固定了有标签与无标签晶体输入、图神经网络骨干接口、仅基于验证集的检查点选择、留出测试集报告、归一化平均绝对误差(NMAE)以及方法排名摘要,涵盖六个稀缺标签任务、四个图骨干网络和五次预定义划分运行。MatRank基于有标签锚点构建伪目标,通过局部可靠性与弱预测一致性对其进行加权,一致地训练弱视图和强视图的图表示,并引入排序信号,使无标签晶体既能塑造数值预测也能影响候选排序。在保留的24个骨干-任务组合中,单一固定的MatRank目标取得了最低的总体留出测试NMAE(0.896)和最佳的平均方法排名(2.208)。组件、分布外(OOD)及生成池诊断实验明确了增益可靠之处以及仍需进一步筛选评估的地方。代码可在 https://github.com/littlepeachs/SemiMat 获取。
cs.LG / 146 / 2608.30695

Liquid Gated Attention

液态门控注意力机制
Jiang, Yiheng, Xu, Yuanbo, Yang, Yongjian
Abstract
Real-world time series often exhibit irregular sampling and extended temporal horizons, requiring models to capture continuous-time dynamics across arbitrary intervals without prohibitive scaling costs. Discrete-time methods collapse variable time intervals into static positional steps; solver-dependent continuous-time models preserve temporal structure but rely on sequential integration, precluding parallelization; and solver-free approximations avoid this cost yet none couples observed time intervals with input-driven state modulation. We propose Liquid Gated Attention (LGA), a solver-free parallel temporal operator. By parameterizing an input-driven gating mechanism with observed time intervals, LGA introduces a continuous-time inductive bias and formulates hidden state evolution as a fast-weight associative memory, enabling parallel computation across the temporal dimension. Using matrix associativity in non-causal encoding and a prefix scan in causal encoding, LGA attains linear temporal complexity in sequence length in both modes. A sequence-level normalization bounds cumulative temporal decay for stable long-horizon optimization. Building on LGA, we instantiate LFormer, a modular backbone for continuous-time representation learning. Across six tasks and sixteen datasets spanning up to 17,984 steps, LFormer demonstrates long-range dependency modeling, fine-grained state tracking, and trajectory reconstruction from sparse and noisy observations, while delivering competitive performance against state-of-the-art discrete-time and continuous-time baselines with linear scaling efficiency.
Chinese Translation
现实世界中的时间序列往往呈现不规则采样和较长的时间跨度,这要求模型能够以可承受的计算开销,在任意时间区间上捕捉连续时间动力学。离散时间方法将可变的时间间隔压缩为固定的位置步长;依赖求解器的连续时间模型虽然保留了时间结构,但依赖顺序积分,无法并行化;无求解器近似方法虽避免了这一开销,但均未将观测到的时间间隔与输入驱动的状态调制相结合。我们提出液态门控注意力(Liquid Gated Attention, LGA),这是一种无求解器的并行时间算子。LGA通过利用观测到的时间间隔对输入驱动的门控机制进行参数化,引入了连续时间的归纳偏置,并将隐状态演化表述为快速权重联想记忆,从而实现时间维度上的并行计算。LGA在非因果编码中利用矩阵结合律,在因果编码中使用前缀扫描(prefix scan),在两种模式下均实现了关于序列长度的线性时间复杂度。序列级归一化为累积时间衰减设定了边界,以保障长时程优化的稳定性。基于LGA,我们构建了LFormer,一个用于连续时间表示学习的模块化骨干网络。在跨越最多17,984步的六项任务和十六个数据集上,LFormer展现了长程依赖建模、细粒度状态追踪以及从稀疏和含噪观测中进行轨迹重建的能力,同时以线性扩展效率取得了与最先进的离散时间和连续时间基线相比具有竞争力的性能。
cs.LG / 147 / 2608.30699

Learning Dynamics of Logits Debiasing for Long-Tailed Semi-Supervised Learning

面向长尾半监督学习的Logits去偏学习动力学
Cheng, Yue, Zhang, Jiajun, Gao, Xiaohui, Xing, Weiwei, Zhu, Zhanxing
Abstract
Long-tailed distributions are prevalent in real-world semi-supervised learning (SSL), where pseudo-labels tend to favor majority classes, leading to degraded generalization. While many long-tailed semi-supervised learning (LTSSL) methods have been proposed, the mechanisms by which they implicitly debias logits remain poorly understood. In this work, we revisit LTSSL through the lens of learning dynamics and provide a theoretical characterization of logits debiasing. Specifically, we derive a step-wise decomposition of the logits updates, showing that predictions are dominated by class-imbalance bias that reliably reflects label priors. To expose this effect, we use the logits of a task-irrelevant baseline image as an indicator of accumulated bias and prove that they converge to the class prior. This provides a unified view where LTSSL remedies such as logit adjustment, reweighting, and resampling correspond to reshaping gradient dynamics. Based on this insight, we propose DyTrim, a principle-based dynamic pruning framework that reallocates gradient budget through class-aware pruning on labeled data and confidence-based soft pruning on unlabeled data. We provide theoretical guarantees that DyTrim reduces class bias and improves generalization. Extensive experiments on standard LTSSL benchmarks show consistent gains across architectures and methods. Code available at: https://jiajun0425.github.io/DyTrim
Chinese Translation
长尾分布在现实世界的半监督学习(SSL)中普遍存在,其中伪标签往往偏向多数类,导致泛化性能下降。尽管已有许多长尾半监督学习(LTSSL)方法被提出,但它们隐式地对logits进行去偏的机制仍缺乏深入理解。在本工作中,我们从学习动力学的视角重新审视LTSSL,并对logits去偏提供了理论刻画。具体而言,我们推导了logits更新的逐步分解,表明预测由可靠反映标签先验的类不平衡偏置所主导。为了揭示这一效应,我们使用与任务无关的基线图像的logits作为累积偏置的指示器,并证明其收敛于类先验。这提供了一个统一视角,使logit调整、重加权和重采样等LTSSL补救措施对应于对梯度动力学的重塑。基于这一见解,我们提出了DyTrim,一个基于原理的动态剪枝框架,通过对有标签数据进行类感知剪枝以及对无标签数据进行基于置信度的软剪枝来重新分配梯度预算。我们提供了理论保证,证明DyTrim能够减少类偏置并提升泛化能力。在标准LTSSL基准上的大量实验表明,该方法在不同架构和方法上均取得了一致的提升。代码可在以下链接获取:https://jiajun0425.github.io/DyTrim
cs.LG / 148 / 2608.30710

Kolmogorov--Arnold against bounded translations

Kolmogorov–Arnold 表示对抗有界平移的鲁棒性
Dzhenzher, Sviatoslav V.
Abstract
Historically originating from Hilbert's 13th problem, the Kolmogorov-Arnold representation theorem (KART) has recently experienced a major revitalisation through its applications to neural networks, specifically Kolmogorov-Arnold Networks (KANs). While the exact representation is well established, its stability under continuous adversarial perturbations of the hidden layer remains a critical open question. In this paper, we investigate the robustness of KART against bounded adversarial translations. We provide an explicit, self-contained, and constructive proof of an approximate representation using fixed, piecewise linear inner functions. Crucially, our construction employs a single outer function that remains invariant for all summands and is independent of the specific adversarial translation, provided its maximum bound is known a priori.
Chinese Translation
Kolmogorov-Arnold 表示定理(Kolmogorov-Arnold Representation Theorem,KART)历史上源于希尔伯特第十三问题,近来因其在神经网络中的应用而重新焕发活力,特别是 Kolmogorov-Arnold 网络(Kolmogorov-Arnold Networks,KANs)。尽管该精确表示已被充分建立,但其在隐层连续对抗扰动下的稳定性仍是一个关键的开放性问题。本文研究了 KART 在有界对抗平移下的鲁棒性。我们利用固定的分段线性内函数,给出了一个近似表示的显式、自包含且构造性的证明。关键之处在于,我们的构造仅使用单一的外函数,该外函数对所有求和项保持不变,并且只要对抗平移的最大界是先验已知的,它便与具体的对抗平移无关。
cs.LG / 149 / 2608.30720

Tracing distinguishability through transformer processing with stochastic LayerNorm

通过随机LayerNorm追踪Transformer处理过程中的可区分性
Murphy, Kieran
Abstract
Representational similarity is foundational to analyses of deep networks, yet distances between point-valued representations are not intrinsically tied to downstream function: nearby states may produce different behaviors, while distant states may behave similarly. We instead give representations volume, turning similarity into statistical distinguishability. Overlapping stochastic representations necessarily induce overlapping downstream distributions, grounding latent comparison in model function and bringing it under information-theoretic tools such as the data-processing inequality. We realize this idea in pretrained transformers through a light-touch modification to LayerNorm: at each residual-stream read, we normalize the state, add isotropic Gaussian noise, and renormalize. During distillation fine-tuning, one learned allocation parameter per residual-stream read distributes a fixed global rate budget across the processing stack. The resulting model can be viewed as transformer blocks reading the residual stream with learned finite precision under a shared global rate budget. Using the Bhattacharyya coefficient, we trace which counterfactual distinctions are preserved through MLP blocks or selectively exposed to the query, key, and value computations of individual attention heads. Experiments on ViT-S and GPT-2 small reveal the depthwise propagation of continuous visual perturbations and head-specific sensitivity to token distinctions aligned with known attention motifs. These results establish distinguishability as a functionally grounded lens on transformer computation that complements existing interpretability approaches.
Chinese Translation
表征相似性是深度网络分析的基础,然而点值表征之间的距离并非本质上与下游功能相关联:相近的状态可能产生不同的行为,而相距遥远的状态却可能表现出相似的行为。我们转而赋予表征体积(volume),将相似性转化为统计可区分性。相互重叠的随机表征必然导致下游分布的相互重叠,从而将潜在比较建立在模型功能之上,并使其可以运用信息论工具(如数据处理不等式)进行分析。我们通过对LayerNorm进行轻量级修改,在预训练Transformer中实现了这一思想:在每次残差流读取时,对状态进行归一化,添加各向同性高斯噪声,然后重新归一化。在蒸馏微调过程中,每个残差流读取位置拥有一个可学习的分配参数,将固定的全局速率预算分配到整个处理堆栈中。所得模型可被视为各个Transformer块在共享的全局速率预算下,以可学习的有限精度读取残差流。利用Bhattacharyya系数,我们追踪了哪些反事实区分度在通过MLP块时得以保留,或被选择性地暴露给各个注意力头的查询(query)、键(key)和值(value)计算。在ViT-S和GPT-2 small上的实验揭示了连续视觉扰动的深度传播,以及注意力头对与已知注意力模式相一致的词元区分的特异性敏感。这些结果确立了可区分性作为一种基于功能视角的Transformer计算分析手段,与现有的可解释性方法形成互补。
cs.LG / 150 / 2608.30724

BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks

BAITBENCH:通过在机器学习任务中植入可选捷径来衡量智能体的奖励作弊行为
Prasad, Pradyumna Shyama, Anto, Meiri, Eshuijs, Leon, Moncarz, Julian, Kislay, Kaustubh, Vazquez, Juan J.
Abstract
LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these environments, bringing into question the validity of produced research and the broader safety case for AI R&D. Existing benchmarks do not measure exploits that live in the data or the modeling task itself. We introduce BAITBENCH, a suite of three synthetic tabular ML tasks that each contain a shortcut that allows agents to inflate the public test score but fail on a hidden test set. Since the shortcut is optional and using it breaks no stated rule, BAITBENCH measures how often models exploit the shortcut to achieve inflated scores. Across seven frontier agents scored by our two-stage judge pipeline, 57.1% of runs exhibit reward hacking, with five of seven above 50%. Agents cheat even under a second condition where they are prompted not to -the mean cheating rate remains above 50%. We release BAITBENCH, along with the judge implementation, and an annotated dataset of transcripts containing reward hacks as a testbed for evaluating reward-hacking mitigations head-to-head.
Chinese Translation
LLM 智能体(agent)正越来越多地被用于自主开展机器学习实验,在极少人工监督的情况下围绕目标指标进行迭代。已有研究记录了此类环境中的奖励作弊(reward hacking)现象,这使人们对所产出研究的有效性以及 AI 研发更广泛的安全性问题产生质疑。现有基准无法度量潜藏在数据或建模任务本身中的漏洞。我们提出了 BAITBENCH,这是一组包含三个合成表格机器学习任务的测试套件,每个任务中都嵌入了一条捷径,智能体可以利用它抬高公开测试集的得分,但在隐藏测试集上表现会失效。由于该捷径是可选的,且使用它并不违反任何明文规则,BAITBENCH 衡量的是模型利用捷径获取虚高得分的频率。在由我们的两阶段评审流水线(judge pipeline)评分的七个前沿智能体中,57.1% 的运行出现了奖励作弊,七个模型中有五个超过 50%。即使在被明确提示不得作弊的第二种条件下,智能体仍然会作弊——平均作弊率仍高于 50%。我们发布了 BAITBENCH,以及评审器实现和一个包含奖励作弊标注对话记录的数据集,作为对各种奖励作弊缓解方法进行直接对比评估的测试平台。
cs.LG / 151 / 2608.30730

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

E-Commerce Bench:评估大语言模型智能体在长时程自主商业运营中的能力
Fan, Wei, Shen, Xinjie, Guo, Xudong, Tu, Jianhong, Su, Yang, Zhang, Yinger, Deng, Lianghao, Wang, Fengyu, Dong, Baohua, Song, Yangqiu, Liu, Dayiheng
Abstract
Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environments and long-range dependencies require Large Language Models (LLMs) to continually explore, learn from experience, and adapt their policies over thousands of steps. We introduce E-Commerce Bench, the first open-source benchmark that integrates multi-round counterpart negotiation and dynamic events into a year-long business operation. Over a 365-day year, an LLM agent concurrently runs multiple online stores, researching the market, negotiating with suppliers to source inventory, optimizing sales strategies, fulfilling orders, handling returns, and managing cash flow to maximize its end-of-year total assets. To construct a realistic merchant-side operating environment, the product and supplier data are derived from a real e-commerce platform, while a year-long calendar of promotions, natural disasters, and supply-chain shocks continually reshapes demand. For reproducibility, both sides of the market are deterministic: customer purchases and returns follow a fixed demand model, while a negotiation kernel determines supplier pricing, concessions, and decisions, with an LLM used only to verbalize them. We evaluate 18 frontier models across seven dimensions, including year-end assets, and find that no single model dominates. GPT-5.6 Sol earns the most, growing the 100,000 opening stake into 1,431,425, yet it ranks 16th of 18 on fraud avoidance and trails Fable5 in operational efficiency. Among open-weight models, Qwen3.8-Max-Preview leads with 416,252, 38% above GLM 5.2 (high), and achieves the strongest learning over the horizon, progressively bargaining down prices across repeated orders. Our code is available at https://github.com/QwenLM/E-CommerceBench.
Chinese Translation
长时程智能体任务并非简单地将短任务在更多交互轮次上进行串联。其不断演化的动态环境和长程依赖要求大语言模型(LLM)持续探索、从经验中学习,并在数千步的过程中不断调整其策略。我们提出了 E-Commerce Bench,这是首个将多轮对手谈判与动态事件融入长达一年商业运营的开源基准测试。在一个365天的年度周期中,LLM 智能体需要同时运营多家在线商店:调研市场、与供应商谈判以采购库存、优化销售策略、履行订单、处理退货并管理现金流,以最大化其年末总资产。为了构建一个真实的商户端运营环境,商品与供应商数据来源于真实的电商平台,同时一份涵盖全年促销活动、自然灾害和供应链冲击的日历持续重塑市场需求。为保证可复现性,市场的双方都是确定性的:客户购买与退货遵循固定的需求模型,而谈判内核决定供应商的定价、让步与决策,LLM 仅用于将其表述为自然语言。我们在包括年末资产在内的七个维度上评估了18个前沿模型,发现没有任何单一模型能够全面领先。GPT-5.6 Sol 赚取最多,将10万的初始资金增长至1,431,425,但在欺诈规避方面在18个模型中仅排第16位,且在运营效率上落后于 Fable5。在开源权重模型中,Qwen3.8-Max-Preview 以416,252领先,比 GLM 5.2(high)高出38%,并在整个周期中展现出最强的学习能力,能够在重复订单中逐步将价格议低。我们的代码发布于 https://github.com/QwenLM/E-CommerceBench。
cs.LG / 152 / 2608.30741

Functional Degeneracy in Neural Networks: Measurement and Pruning

神经网络中的功能简并性:度量与剪枝
Matveev, Maria, Esser, Pascal, Bharadwaj, Ayush, Bushnaq, Lucius, Kutyniok, Gitta
Abstract
A central question in modern machine learning is how much a trained model can be compressed without changing its behavior, to reduce the memory, compute and energy required to deploy it. To study this, we quantify functional degeneracy through the behavioral recovery rank, defined as the number of leading behavioral-Hessian eigendirections required to recover a trained model's performance. Using the behavioral recovery rank as a geometric benchmark for compression, we find that structural and magnitude pruning retain more degrees of freedom, even after the task is saturated. This gap suggests that functional redundancy is distributed across parameter directions and is not exposed by individual weights or neurons.
Chinese Translation
现代机器学习中的一个核心问题是:在不改变模型行为的前提下,训练好的模型能被压缩到什么程度,从而降低部署所需的内存、计算量和能耗。为研究这一问题,我们通过行为恢复秩(behavioral recovery rank)来量化功能简并性,该指标定义为恢复训练模型性能所需的领先行为-Hessian特征方向的数量。将行为恢复秩作为压缩的几何基准,我们发现结构性剪枝和幅值剪枝即使任务已饱和后仍保留更多的自由度。这一差距表明,功能冗余分布在参数方向之中,而不会通过单个权重或神经元暴露出来。
cs.LG / 153 / 2608.30745

TDDM-Melatt: A Decoupled Memory and Diffusion Framework for Generalizable Encrypted Traffic Classification

TDDM-Melatt:一种面向可泛化加密流量分类的解耦记忆与扩散框架
Chen, Ze, Yu, Qiming, Song, Zijia, Yang, Guozheng, Yan, Wei
Abstract
The widespread adoption of encrypted traffic poses severe challenges to current security situational awareness systems based on network traffic monitoring. In existing dataset-driven training and testing studies, limitations such as shortcut learning induced by spurious feature correlations and sample imbalance caused by the long-tail distribution of real-world traffic result in weak generalization of traffic identification performance to real-world network traffic. To address these limitations, we propose TDDM-Melatt, a disentangled memory-based traffic classification framework with diffusion-based data augmentation. First, we design Melatt, a memory-decoupled traffic representation model, which employs Competitive Gating Long Short-Term Memory (CG-LSTM) to construct the encoder and decoder. We design a spurious-correlation-free pre-training and inference paradigm, employing strict topology anonymization and a frozen pre-trained encoder strategy to cut off the model's learning pathways for spurious features. During inference, classification is performed efficiently by a downstream classifier on the frozen representations. Second, we propose a Traffic Denoising Diffusion Model (TDDM) tailored to the characteristics of traffic data. Extensive experiments are conducted on 4 representative public benchmark datasets. Under strict flow-level splitting and anonymization, TDDM-Melatt outperforms 6 basic classification models and 6 SOTA representation learning models. The proposed method provides a new and effective technical pathway for encrypted traffic classification in real-world network environments.
Chinese Translation
加密流量的广泛使用给当前基于网络流量监测的安全态势感知系统带来了严峻挑战。在现有的基于数据集的训练与测试研究中,由虚假特征相关性导致的捷径学习以及由真实世界流量的长尾分布造成的样本不均衡等局限性,使得流量识别性能对真实网络流量的泛化能力较弱。为解决这些局限,我们提出了TDDM-Melatt,一个基于解耦记忆并采用扩散模型进行数据增强的流量分类框架。首先,我们设计了Melatt,一种记忆解耦的流量表示模型,其采用竞争门控长短期记忆网络(Competitive Gating Long Short-Term Memory, CG-LSTM)构建编码器和解码器。我们设计了一种无虚假相关性的预训练与推理范式,通过严格的拓扑匿名化和冻结预训练编码器策略,切断模型对虚假特征的学习路径。在推理阶段,由下游分类器在冻结的表示上高效地完成分类。其次,我们提出了一种针对流量数据特性设计的流量去噪扩散模型(Traffic Denoising Diffusion Model, TDDM)。我们在4个具有代表性的公开基准数据集上开展了大量实验。在严格的流级切分和匿名化条件下,TDDM-Melatt优于6个基础分类模型和6个最先进的(SOTA)表示学习模型。所提出的方法为真实网络环境下的加密流量分类提供了一条新颖而有效的技术路径。
cs.LG / 154 / 2608.30750

Do VLMs Share Safety Neurons Across Modalities?

视觉语言模型(VLM)在不同模态间共享安全神经元吗?
Li, Jiaxuan, Zhang, Jiahao, Vo, Duc Minh, Nguyen, Huy H., Kavumba, Pride, Wataoka, Koki
Abstract
Vision-language models (VLMs) can comply with harmful requests delivered through images, even when their LLM backbones would refuse the same content in text. While prior work characterizes these jailbreaks empirically or at the representation level, how visual inputs perturb safety pathways at the neuron level remains uncharted. We close this gap with a causal, neuron-level analysis of safety mechanisms in 10 VLMs. We propose a two-stage detection pipeline with iterative ablation that accounts for self-repair, and introduce two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval, which decouple visual and textual safety signals. Our analysis reveals: (i) Text safety in VLMs is localizable: $\sim$88 neurons ($<$0.01%) whose targeted ablation substantially reduces refusal. (ii) Text safety neurons constitute the dominant refusal pathway: ablating them is the only intervention that consistently and substantially reduces refusal across all models. (iii) Visual safety is high-dimensional and diffuse at the single-neuron level: text safety concentrates in $\sim$5 subspace directions while visual safety requires $\geq$50. This gap holds across architectures, explaining why current alignment has not closed the visual safety gap. Project page is at: https://jiaxuan-li.github.io/vlm-safety-neuron/ Warning: this paper may include examples of harmful content.
Chinese Translation
视觉语言模型(VLM)能够遵从通过图像传递的有害请求,即使其底层大语言模型(LLM)在文本形式下会拒绝相同内容。尽管先前的工作在经验层面或表征层面对这些越狱攻击进行了刻画,但视觉输入如何在神经元层面扰动安全机制仍属空白。我们通过对10个VLM的安全机制进行因果性、神经元层面的分析来填补这一空白。我们提出了一种考虑自我修复(self-repair)的两阶段检测流程(含迭代消融),并引入两个模态隔离的基准:ViSafe-Detect和ViSafe-Eval,用于解耦视觉与文本安全信号。我们的分析揭示:(i)VLM中的文本安全是可定位的:约88个神经元(<0.01%),对其定向消融可显著降低拒答行为。(ii)文本安全神经元构成主要的拒答通路:消融它们是唯一能在所有模型上一致且显著降低拒答的干预手段。(iii)视觉安全在单神经元层面是高维且弥散的:文本安全集中于约5个子空间方向,而视觉安全需要≥50个方向。这一差距在不同架构上普遍存在,解释了为何当前的对齐方法尚未弥合视觉安全鸿沟。项目页面:https://jiaxuan-li.github.io/vlm-safety-neuron/ 警告:本文可能包含有害内容示例。
cs.LG / 155 / 2608.30760

PRACTICE: From Experience to Expertise in Self-Evolving Embodied Agents

PRACTICE:从经验到专家的自我进化具身智能体
Bai, Ziyi, Li, Siqi, Huang, Tinglei, Karlsson, Börje F.
Abstract
Recent studies have shown that multimodal large language models (MLLMs) can serve as embodied agents, translating language instructions and visual observations into executable plans. However, building agents that can continually improve through interaction and rapidly adapt to their environments remains challenging. Summing up experience from past interaction trajectories provides a promising solution, but existing experience-based methods often rely on manually designed prompting workflows to extract and update skills. Such fixed procedures may struggle to learn updated skills from new and diverse experiences. We introduce PRACTICE, which trains a skill learner to discover and maintain a persistent skill library from past interaction trajectories while keeping the task executor frozen. Given the historical accumulated skills and incoming trajectories, the skill learner produces structured batch-edits that add, refine, merge, or remove skills, and then hierarchical consolidate all collected edits into a consistent updated skill library. We train the learner with a two-stage curriculum. First, it learns basic skill generation and library maintenance from oracle trajectories. Then, by contrasting successful and failed trajectories from heterogeneous executors on the same tasks, it learn to identify invalid action patterns and recovery strategies. Finally, we apply online skill-edit distillation to align the skill learner with a stronger teacher on its current edit distribution to further improves the policy. Experiments demonstrate that a compact skill learner delivers consistent performance improvements across successive library-update rounds for multiple frozen executors. On EB-ALFRED and EB-Habitat, PRACTICE further outperforms the strongest experience-based baselines. Project resources are publicly available at: https://baai-agents.github.io/PRACTICE
Chinese Translation
近期研究表明,多模态大语言模型(MLLMs)可以作为具身智能体,将语言指令和视觉观察转化为可执行的计划。然而,构建能够通过交互持续改进并快速适应环境的智能体仍然具有挑战性。从过往交互轨迹中总结经验是一种有前景的解决方案,但现有的基于经验的方法通常依赖人工设计的提示工作流来提取和更新技能。这种固定流程可能难以从新颖多样的经验中学习更新后的技能。我们提出了 PRACTICE,它训练一个技能学习器从过往交互轨迹中发现并维护一个持久化的技能库,同时保持任务执行器冻结不变。给定历史积累的技能和新输入的轨迹,技能学习器生成结构化的批量编辑操作,以添加、细化、合并或删除技能,然后分层地整合所有收集到的编辑操作,形成一致的更新后技能库。我们采用两阶段课程训练技能学习器:首先,它从示范(oracle)轨迹中学习基本的技能生成与技能库维护;然后,通过对比异构执行器在同一任务上的成功与失败轨迹,学会识别无效的动作模式与恢复策略。最后,我们应用在线技能编辑蒸馏,使技能学习器在其当前编辑分布上与更强的教师模型对齐,从而进一步提升策略。实验表明,一个紧凑的技能学习器能够在连续多轮技能库更新中为多个冻结执行器带来一致的性能提升。在 EB-ALFRED 和 EB-Habitat 上,PRACTICE 进一步超越了最强的基于经验的基线方法。项目资源已公开于:https://baai-agents.github.io/PRACTICE
cs.LG / 156 / 2608.30765

T3S: Improving Multi-Task Reinforcement Learning with Task-Specific Feature Selector and Scheduler

T3S:利用任务特定特征选择器与调度器改进多任务强化学习
Yu, Yuanqiang, Yang, Tianpei, Lv, Yongliang, Zheng, Yan, Hao, Jianye
Abstract
Multi-task reinforcement learning (MTRL) is a technique to train multiple tasks simultaneously, where previous works usually train a single model to solve different tasks by sharing parameters across various tasks. However, these methods are faced with inter-task interference since what parameters should be shared across tasks is not addressed, dramatically reducing learning efficiency. To solve these problems, we propose a novel MTRL framework called Task-Specific feature Selector and Scheduler (T3S), which consists of two components: a feature selector and a task scheduler. Specifically, the feature selectors employ hypernetworks to construct task-specific soft masks, which can be applied by globally shared representation to construct task-specific features. The task scheduler selects tasks for learning through two metrics, where the selection probability is inversely proportional to task progress (e.g., success rate) and task learning speed. Experimental results show that T3S consistently outperforms the state-of-the-art MTRL algorithms on various robotics manipulation tasks.
Chinese Translation
多任务强化学习(MTRL)是一种同时训练多个任务的技术,以往的工作通常训练单一模型,通过在不同任务之间共享参数来解决不同的任务。然而,由于未解决哪些参数应在任务间共享的问题,这些方法面临任务间干扰,从而显著降低了学习效率。为解决这些问题,我们提出了一种新颖的MTRL框架,称为任务特定特征选择器与调度器(Task-Specific feature Selector and Scheduler, T3S),它由两个组件构成:特征选择器和任务调度器。具体而言,特征选择器利用超网络(hypernetwork)构建任务特定的软掩码,可应用于全局共享的表征以构建任务特定的特征。任务调度器通过两个指标为学习选择任务,其中任务被选择的概率与任务进度(如成功率)和任务学习速度成反比。实验结果表明,T3S在各种机器人操作任务上持续优于当前最先进的MTRL算法。
cs.LG / 157 / 2608.30769

TrainSDC: Characterizing and Mitigating Silent Data Corruption in Large Language Model Training

TrainSDC:大语言模型训练中静默数据损坏的表征与缓解
Xia, Zhipeng, Xu, Haotian, Yun, Siyu, Lin, Liqi, Liu, Hu, Li, Yu, Zhuo, Cheng
Abstract
LLM training is increasingly vulnerable to silent data corruption (SDC), yet existing protection methods largely treat Transformer computations uniformly because their vulnerability remains poorly understood. We present the first systematic characterization of SDC vulnerability across major computation interfaces in both the forward and backward passes of Transformer training. Our analysis reveals two distinct error propagation mechanisms: forward-pass vulnerability is highly location dependent, with faults on the Q/K path producing persistent training deviations, whereas backward-pass vulnerability is largely governed by gradient exponent distributions rather than computation locations. Motivated by these observations, we propose TrainSDC, a characterization-guided protection framework consisting of Q/K-path recomputation, residual-gain monitoring, and exponent-aware gradient scaling. Experiments on Llama 3.2-1B and Qwen3-0.6B show that TrainSDC maintains training behavior close to fault-free execution under both sparse and dense fault injection while introducing only 1.65%-6.76% runtime overhead.
Chinese Translation
大语言模型(LLM)训练日益容易受到静默数据损坏(Silent Data Corruption, SDC)的影响,然而由于人们对其脆弱性认识不足,现有的保护方法大多对 Transformer 计算采取统一处理。本文首次系统性地表征了 Transformer 训练前向与反向传播中主要计算接口的 SDC 脆弱性。我们的分析揭示了两种截然不同的错误传播机制:前向传播的脆弱性高度依赖于故障位置,Q/K 路径上的故障会导致持续的训练偏差;而反向传播的脆弱性主要由梯度指数分布决定,而非计算位置。基于这些观察,我们提出了 TrainSDC,一个由表征分析指导的保护框架,包含 Q/K 路径重计算、残差增益监控以及指数感知的梯度缩放。在 Llama 3.2-1B 和 Qwen3-0.6B 上的实验表明,在稀疏和密集故障注入下,TrainSDC 均能保持与无故障执行接近的训练行为,同时仅引入 1.65%–6.76% 的运行时开销。
cs.LG / 158 / 2608.30778

Reciprocity Separates Gradient Flow from Rotation in Conservative Physical Learning

互易性在保守物理学习中区分梯度流与旋转
Niu, Ruiwu, Bi, Xiaowen, van Wyk, Michaël Antonie
Abstract
Physical learning lets a trainable material or network use its own physical response to carry error signals, reducing the need for a separately programmed backward computation. We ask what determines whether such a system follows conventional gradient descent or evolves along a genuinely different learning trajectory. Our canonical model is a directed layered transport network in which every node redistributes a fixed amount of flow, so learning preserves positivity and total mass. In this model, conservation constrains only the allowable learning directions. Within the matched response class studied here, adjoint matching gives the physical output response a symmetric form. Non-negative mode-wise feedback then produces a reciprocal closed-loop response and a reweighted gradient flow. Adding an antisymmetric boundary component makes the closed-loop response rotational: the learning path can turn while the error driving that update still decreases at that moment. Turning is not automatically beneficial. Its finite-step effect is set by local curvature, and its accumulated effect also depends on step selection and on the new states visited along the path. Numerical consistency checks reproduce the exact response structure, predict the sign of the local effect across new network families, and show how trajectory drift can negate a local advantage. These results separate the roles of conservation, reciprocity, and nonreciprocity in physical learning.
Chinese Translation
物理学习使可训练的材料或网络能够利用其自身的物理响应来传递误差信号,从而减少对单独编写的反向计算的需求。我们探讨是什么决定了这样的系统是遵循常规的梯度下降,还是沿着一条真正不同的学习轨迹演化。我们的规范模型是一个有向分层输运网络,其中每个节点重新分配固定量的流量,因此学习过程保持正性和总质量。在该模型中,守恒仅约束允许的学习方向。在本文所研究的匹配响应类中,伴随匹配使物理输出响应具有对称形式。随后,非负的逐模反馈产生互易的闭环响应和重加权的梯度流。加入一个反对称边界分量则使闭环响应具有旋转性:学习路径可以转向,而驱动该更新的误差在那一刻仍在减小。转向并非自动有益。其有限步效应由局部曲率决定,其累积效应还取决于步长选择以及沿路径所到达的新状态。数值一致性检验复现了精确的响应结构,在新的网络家族中预测了局部效应的符号,并展示了轨迹漂移如何抵消局部优势。这些结果区分了守恒、互易性与非互易性在物理学习中的作用。
cs.LG / 159 / 2608.30804

Geometric Attractor Monitoring: A Robust and Frugal Framework for Multi-modal Industrial Robotic Cycles

几何吸引子监测:一种面向多模态工业机器人周期的鲁棒且低耗框架
Bonsergent-Brachet, Martin, Read, Jesse, Abboud, Dany
Abstract
Monitoring the health of heterogeneous industrial robot fleets is severely challenged by the multi-modal nature of their operational cycles and a persistent scarcity of run-to-failure data. Standard data-driven approaches, particularly deep learning architectures relying on sequential reconstruction, often struggle in this specific setting; they tend to over-smooth complex dynamics, masking early signs of degradation. To address these industrial constraints, we reframe the monitoring problem through a framework based on Phase Space Reconstruction (PSR). Instead of predicting temporal sequences, this framework transforms univariate sensor data into a geometric attractor, explicitly unfolding the mechanical states independently of their temporal occurrence. By evaluating various anomaly scoring techniques within this space, we demonstrate that discrete support estimation provides an effective and computationally frugal Health Indicator (HI). Validated on a real-world dataset of 21 heterogeneous robots over three years and a synthetic Langevin system, our approach outperforms standard deep learning baselines. We show that aligning the algorithmic bias with the geometric properties of the target system yields a pragmatic, traceable and easily deployable approach perfectly tailored to the realities of industrial constraints.
Chinese Translation
异构工业机器人机群的健康监测面临着严峻挑战:其运行周期具有多模态特性,且从运行到失效(run-to-failure)的数据持续稀缺。标准的数据驱动方法,尤其是依赖序列重构的深度学习架构,在这种特定场景下往往表现不佳;它们容易对复杂动态进行过度平滑,从而掩盖退化早期迹象。为应对这些工业约束,我们通过基于相空间重构(Phase Space Reconstruction, PSR)的框架重新构建监测问题。该框架不预测时间序列,而是将单变量传感器数据转换为几何吸引子,明确地展开机械状态而不依赖其时间先后关系。通过在该空间中评估多种异常评分技术,我们证明离散支撑集估计(discrete support estimation)提供了一种有效且计算低耗的健康指标(Health Indicator, HI)。在包含21台异构机器人、历时三年的真实数据集以及一个合成Langevin系统上的验证表明,我们的方法优于标准深度学习基线。我们表明,使算法偏置与目标系统的几何特性相匹配,能够得到一种务实、可追溯且易于部署的方法,完美契合工业约束的现实。
cs.LG / 160 / 2608.30819

What Emerges and What Breaks in Self-Play Driving

自我博弈驾驶中的涌现与失效
Sisask, Laur, Tampuu, Ardi, Matiisen, Tambet
Abstract
Training autonomous driving policies through pure self-play has recently shown promising results. Following Gigaflow and Puffer- Drive, we train driving policies in a similar self-play fashion, but extend the models from MLPs to Transformers and train on the high-definition map of a real city, where we ultimately aim to deploy them. On the CARLA and Waymax benchmarks, our policies fall short of Gigaflow, and we trace the gap to specific failure modes, including reward hacking at traffic lights and a missing incentive to stop at stop signs. We further analyze which traffic rules emerge from self-play and how closely they match human driving, and we confirm that reward conditioning yields the intended diversity of driving behaviors. A demonstration of a trained policy is available at https://laursisask-ut.github.io/eccvdemo.
Chinese Translation
通过纯自我博弈(self-play)训练自动驾驶策略近来已展现出令人期待的结果。继 Gigaflow 和 Puffer-Drive 之后,我们以类似的自我博弈方式训练驾驶策略,但将模型从多层感知机(MLP)扩展到 Transformer,并在一个我们最终计划部署的真实城市高精地图上进行训练。在 CARLA 和 Waymax 基准测试中,我们的策略未能达到 Gigaflow 的水平,我们将这一差距追溯到特定的失效模式,包括对交通灯的奖励投机(reward hacking)以及在停止标志前缺乏停车激励。我们进一步分析了自我博弈中涌现出哪些交通规则,以及它们与人类驾驶的匹配程度,并确认奖励条件化能够产生预期的驾驶行为多样性。训练策略的演示可在 https://laursisask-ut.github.io/eccvdemo 查看。
cs.LG / 161 / 2608.30877

Deploying DeepSeek 175B Locally on a Single Consumer-Grade RTX 4060 Laptop with 32GB RAM for 200k-Scale Protein-Ligand Virtual Screening

在配备32GB内存的单台消费级RTX 4060笔记本电脑上本地部署DeepSeek 175B模型,完成20万规模蛋白质-配体虚拟筛选
Xiao, Rui, Xu, Yili
Abstract
Recent advances in large language models (LLMs) have demonstrated exceptional performance in protein-ligand interaction prediction, but state-of-the-art pipelines for large-scale virtual screening almost exclusively rely on high-end GPU clusters with hundreds of gigabytes of memory, creating prohibitive hardware barriers for small academic teams. In this work, we present a fully local low-resource framework that deploys the 175-billion-parameter DeepSeek 175B LLM on a single consumer-grade RTX 4060 laptop equipped with 32GB system RAM and 8GB VRAM, completing a full 200k-scale protein-ligand virtual screening workflow across 20 distinct protein targets. Our implementation achieves 100x throughput of an 8-card A100 cluster baseline under identical task configurations within 72 hours, with an average binding affinity prediction error of 0.88 kcal/mol across all targets, satisfying the 1.0 kcal/mol chemical accuracy requirement for preclinical drug discovery. Systematic runtime profiling reveals that heterogeneous memory management overhead accounts for 72% of total execution time, while accuracy loss introduced by model optimization contributes less than 10% to total prediction error. This work validates the engineering feasibility of running industrial-scale trillion-parameter LLM-driven biomedical computing tasks on consumer hardware, establishing a new low-barrier paradigm for AI-powered early stage drug discovery.
Chinese Translation
大语言模型(LLM)的最新进展在蛋白质-配体相互作用预测方面展现出了卓越的性能,但用于大规模虚拟筛选的最先进流程几乎完全依赖数百GB内存的高端GPU集群,这给小型学术团队设置了难以逾越的硬件门槛。在这项工作中,我们提出了一个完全本地化的低资源框架,将拥有1750亿参数的DeepSeek 175B大语言模型部署在一台配备32GB系统内存和8GB显存的消费级RTX 4060笔记本电脑上,完成了跨越20种不同蛋白靶点的完整20万规模蛋白质-配体虚拟筛选工作流。在相同任务配置下,我们的实现在72小时内达到了8卡A100集群基准100倍的吞吐量,所有靶点的平均结合亲和力预测误差为0.88 kcal/mol,满足了临床前药物发现1.0 kcal/mol的化学精度要求。系统的运行时性能分析表明,异构内存管理开销占总执行时间的72%,而模型优化引入的精度损失对总预测误差的贡献不足10%。这项工作验证了在消费级硬件上运行工业规模万亿参数LLM驱动的生物医学计算任务的工程可行性,为AI驱动的早期药物发现建立了一种全新的低门槛范式。
cs.LG / 162 / 2608.30908

Fine-Tuning Low-Bit Models with Gradient in Quantized Code Space

在量化码空间中利用梯度微调低比特模型
Wu, Shiguang, Lin, Zhouchen, Yao, Quanming
Abstract
Fine-tuning Low-bit models aims to adapt a quantized model while keeping the final deployed checkpoint in the same low-bit form. This setting is practically important as it reduces memory and inference cost for storage and deployment. Under this constraint, adaptation becomes an optimization problem over quantization codes and scales. Existing continuous low-bit training is efficient, but it can be distorted by straight through estimation error or by post-quantize gap; discrete search is deployment-faithful, but it is often too inefficient under a finite training budget. We propose code surrogate gradient as the first order signal in deployable code space to acceleate optimization, and performing guided search to preserve deployment faithfulness. Experiments across arithmetic reasoning, instruction following, and structured language understanding show that GradCodes consistently improves fine-tuning low-bit models across different quantization datatypes. Code is provided at https://github.com/ovo67/GradCodes.
Chinese Translation
低比特模型微调旨在适配量化模型的同时,保持最终部署的检查点仍处于相同的低比特形式。这一设置具有重要的实际意义,因为它可以降低存储和部署的内存与推理成本。在此约束下,适配问题转化为对量化码(quantization codes)和缩放因子(scales)的优化问题。现有的连续型低比特训练方法效率较高,但可能受到直通估计(straight-through estimation)误差或量化后偏差(post-quantize gap)的影响而产生失真;离散搜索方法虽然与部署保持一致,但在有限的训练预算下往往效率过低。我们提出码代理梯度(code surrogate gradient)作为可部署码空间中的一阶信号以加速优化,并通过引导搜索保持与部署的一致性。在算术推理、指令遵循和结构化语言理解任务上的实验表明,GradCodes 在不同量化数据类型下均能持续提升低比特模型的微调效果。代码已发布于 https://github.com/ovo67/GradCodes。
cs.LG / 163 / 2608.30910

S3C-LLM: Skill-Code Guided Agentic Language Models for Spectrum-to-Structure Elucidation

S3C-LLM:用于光谱到结构解析的技能-代码引导的智能体语言模型
Zhao, Xuanle, Cai, Xinyuan, Cheng, Xiang, Xu, Bo
Abstract
Spectroscopic structure elucidation is central to molecular analysis, but recent Large Language Model (LLM)-based methods mostly formulate it as direct spectrum-to-SMILES generation. Although this paradigm can leverage paired spectral data, it does not explicitly model the analytical workflow used by spectroscopists, such as diagnostic peak interpretation, fragment reasoning, formula constraints, and chemical consistency checking. In this paper, we introduce S3C-LLM, a skill-guided and code-grounded agentic LLM for spectrum-to-structure elucidation. Rather than directly predicting a molecule, S3C-LLM retrieves modality-specific spectroscopy skills, executes analysis code to instantiate these skills on the input spectra, and integrates the resulting peak-level evidence and formula constraints before generating SMILES. Specifically, we contribute a self-evolving spectroscopy skill library, a thinking-augmented skill-code trajectory construction pipeline, and a two-stage training strategy that teaches Qwen3-4B through supervised fine-tuning (SFT) followed by our proposed step-level reinforcement learning (RL). Experiments on diverse benchmarks show that S3C-LLM consistently outperforms current general LLMs and spectrum-specific models across spectra, while using less than 1/10th of SpectraLLM's training corpus.
Chinese Translation
光谱结构解析是分子分析的核心,但近期基于大语言模型(LLM)的方法大多将其构建为从光谱直接生成SMILES的任务。尽管这一范式可以利用成对的光谱数据,但它并未显式建模光谱学家所采用的分析工作流程,例如诊断性峰位解读、片段推理、分子式约束以及化学一致性检查。本文提出了S3C-LLM,一个用于光谱到结构解析的、技能引导且代码支撑的智能体大语言模型。S3C-LLM并非直接预测分子,而是检索特定模态的光谱分析技能,执行分析代码以在输入光谱上实例化这些技能,并在生成SMILES之前整合由此得到的峰级别证据与分子式约束。具体而言,我们的贡献包括:一个自我演进的光谱技能库、一个思维增强的技能-代码轨迹构建流程,以及一个两阶段训练策略,即先通过监督微调(SFT)训练Qwen3-4B,再采用我们提出的步骤级强化学习(RL)。在多个基准上的实验表明,S3C-LLM在各类光谱上均持续优于现有的通用大语言模型和光谱专用模型,同时其训练语料不足SpectraLLM的十分之一。
cs.LG / 164 / 2608.30916

Selection-Aware Stress Testing for Interactive Agents

面向交互式智能体的选择感知压力测试
Xu, Yang, Li, Chenang, Zhang, Jiefu, Sun, Haixiang, Li, Zhou, Aggarwal, Vaneet
Abstract
Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advantage weakens, so both conclusions are selected from the same data. We introduce Selection-Aware Semantic Stress Testing (\SASST{}), which learns a task reweighting from pre-execution features on discovery tasks and evaluates the same paired comparison on separate confirmation tasks. The protocol checks support and stability, uses joint bounds for all planned claims, and can return no claim. We prove conditional asymptotic validity under stated cluster assumptions. A forty-cluster audit finds Gaussian undercoverage and conservative Bonferroni $t$ bounds. In one 480-episode $\tau$-bench study, a $3.75$ point discovery gain vanished on confirmation. A second-model study likewise confirmed neither a workflow benefit nor a stable stress rule.
Chinese Translation
智能体评估通常先用一个基准来选择工作流,然后再搜索该工作流优势减弱的任务类型,因此这两个结论都是从同一数据中得出的。我们提出了选择感知语义压力测试(Selection-Aware Semantic Stress Testing,SASST),该方法在发现任务上从执行前特征学习任务重加权,并在独立的确认任务上评估相同的配对比较。该协议检验支持性与稳定性,对所有计划的结论使用联合边界,并且可以返回“无可下结论”的结果。我们在给定的聚类假设下证明了条件渐近有效性。一项四十个聚类的审计发现了高斯(Gaussian)覆盖不足以及保守的 Bonferroni $t$ 边界。在一项包含 480 个回合的 $ au$-bench 研究中,发现阶段的 3.75 个百分点优势在确认阶段消失。另一项第二模型的研究同样既未确认工作流的收益,也未确认稳定的压力规则。
cs.LG / 165 / 2608.30923

Towards Stream Learning on Embedded Systems: Benchmarking the Memory Consumption of Stream Learning Methods

迈向嵌入式系统上的流学习:流学习方法内存消耗的基准测试
Buschjäger, Sebastian, Gunasekara, Nuwan, Gomes, Heitor Murilo
Abstract
Stream learning is commonly evaluated through predictive performance and adaptation to concept drift. However, sustained operation of a stream learner also requires predictable and bounded resource usage even on long streams. This requirement becomes even more critical when learning moves from servers to near-sensor embedded systems where memory and processing are scarce resources. In state-of-the-art stream learning, however, we perceive a strong focus on concept drift adaptation, whereas resource usage is often an evaluation byproduct. To close this gap, we benchmark seven representative stream classifiers on 13 real and synthetic streams under model-size budgets from 128\,KiB to approximately 8\,MiB. Our benchmark comprises a total of 6,463 experiments. We measure failure-aware accuracy, peak model size, time to budget exhaustion, and prediction-plus-update latency. The results reveal two distinct resource failure modes. Adaptive ensembles can exceed small budgets almost immediately because of their initial footprint, even when their size remains stable thereafter. Incremental trees can fit initially but grow throughout a long stream, with HoeffdingTrees (HT) and Extremely Fast Decision Trees (EFDT) increasing by median factors of 7.37 and 5.87. Explicitly compact methods remain the only viable option under the smallest budgets, but are usually overtaken as larger budgets make adaptive ensembles competitive. Hence, many state-of-the-art methods are only partially applicable in embedded systems or for long-running systems. We therefore call on the stream-learning community to make bounded resource usage a first-class design objective alongside drift adaptation, and propose concrete steps toward this goal, including an API through which stream learners can explicitly expose and respect resource budgets.
Chinese Translation
流学习通常通过预测性能和对概念漂移的适应能力来进行评估。然而,流学习器的持续运行还要求即使在长数据流上也能保持可预测且有界的资源使用。当学习从服务器迁移到内存和处理能力稀缺的近传感器嵌入式系统时,这一要求变得更加关键。然而,在最先进的流学习研究中,我们观察到研究重点强烈集中在概念漂移适应上,而资源使用往往只是评估的副产品。为弥补这一空白,我们在128 KiB至约8 MiB的模型大小预算下,对七种具有代表性的流分类器在13个真实和合成数据流上进行了基准测试。我们的基准测试共包含6,463项实验。我们测量了考虑失败情况的准确率、峰值模型大小、达到预算耗尽的时间以及预测加更新的延迟。结果揭示了两种截然不同的资源失败模式。自适应集成方法由于其初始占用,即使此后规模保持稳定,也可能几乎立即超出小预算。增量树方法起初可以适应预算,但会在长数据流中持续增长,其中HoeffdingTree(HT)和Extremely Fast Decision Tree(EFDT)的中位数增长因子分别为7.37和5.87。显式紧凑的方法是在最小预算下唯一可行的选择,但当更大的预算使自适应集成方法变得具有竞争力时,它们通常会被超越。因此,许多最先进的方法只能部分适用于嵌入式系统或长时间运行的系统。因此,我们呼吁流学习界将有界资源使用作为与漂移适应同等重要的一等设计目标,并提出实现这一目标的具体步骤,包括一个使流学习器能够显式公开并遵守资源预算的API。
cs.LG / 166 / 2608.30944

Nonparametric Contextual Pricing and Inventory Learning under Censored Demand

删失需求下的非参数情境化定价与库存学习
Han, Zean, Liang, Jing, Lin, Ruihan, Ding, Zezhen, Zhang, Jiheng
Abstract
In online retailing, when a product sells out, a retailer often sees only the units sold, not how many customers would have bought it had inventory been available. However, the inventory level determines how much demand is revealed, and this information can influence subsequent decisions and future profits. We study an online selling problem in which, in each round, the seller observes a market context and then makes pricing and stocking decisions based on censored sales data from previous rounds. The challenge is to learn a context-dependent pricing and stocking policy without assuming a particular formula for demand or observing realized profit. To overcome this difficulty, we propose a Mean-Calibrated Kernel UCB (MCK-UCB) algorithm that turns each incomplete sales record into a reliable guide for both inventory and price decisions, using data from past rounds with similar market conditions. This design allows us to learn while serving customers, without a separate exploration phase or the need to recover all demand hidden by stockouts. We prove the minimax optimality of the proposed algorithm, with strictly faster rates when expected profit varies more smoothly with price. Comprehensive numerical experiments have been conducted to confirm the effectiveness of the proposed algorithm.
Chinese Translation
在在线零售中,当产品售罄时,零售商通常只能观察到售出的数量,而无法知道如果库存充足还会有多少顾客购买。然而,库存水平决定了需求信息被揭示的程度,而这些信息会影响后续决策和未来利润。我们研究一个在线销售问题:在每一轮中,卖方观察到市场情境,然后基于前几轮的删失销售数据做出定价和备货决策。其挑战在于,在不假设需求的具体函数形式、也观察不到实际利润的情况下,学习一个依赖情境的定价与备货策略。为克服这一困难,我们提出了一种均值校准核UCB算法(Mean-Calibrated Kernel UCB,MCK-UCB),该算法利用具有相似市场环境的历史轮次数据,将每条不完整的销售记录转化为库存决策和定价决策的可靠依据。这种设计使算法能够在服务顾客的同时进行学习,无需单独的探索阶段,也不需要完全恢复因缺货而被隐藏的全部需求。我们证明了所提算法的极小极大最优性,并且当期望利润随价格变化更平滑时可获得严格更快的收敛速率。综合数值实验验证了所提算法的有效性。
cs.LG / 167 / 2608.30946

Reproducible macroscopic dynamics in a closed-loop human-AI learning system

闭环人机学习系统中的可复现宏观动力学
Wu, Minlin, Fang, Xu, Zhang, Yicheng, Zhou, Chenyu, Liu, Zhiyi
Abstract
Closed-loop human-AI systems generate high-dimensional behavioural trajectories whose collective dynamics remain obscure. Using 297,915 learners' adaptive-tutoring histories, we define semantic order variables before model fitting and test them in user-disjoint cohorts. The state exhibits reproducible basin-like flow and operationally defined, state-heterogeneous metastable-like kinetics. A construction-matched null distinguishes normalised-memory relaxation from a reproducible excess field. A four-term conditional mechanism recovers population drift (r = 0.946; learner-bootstrap 95% CI, 0.935-0.955). Predictive event-level self-supervised learning recovers the state and learned-plane flow; null-referenced corrections retain directional, partial-amplitude excess-field structure without full calibration. Shuffled-order training reverses learned-plane flow on ordered trajectories; support-alignment randomisation selectively reduces inward transport. Both axes remain linearly accessible without state supervision. Without cross-model fitting, the models share leading population drift (r = 0.866; learner-bootstrap 95% CI, 0.857-0.875) and persistence ordering; residual directions remain model-specific. These results identify an externally anchored leading-order effective field linking empirical dynamics, an interpretable mechanism and neural computation.
Chinese Translation
闭环人机系统会产生高维行为轨迹,其集体动力学仍不明确。基于297,915名学习者的自适应辅导历史,我们在模型拟合之前定义了语义序变量,并在用户不相交的队列中进行了检验。该状态表现出可复现的类盆地流(basin-like flow)以及操作性定义的、状态异质的类亚稳态动力学。一个构建匹配的零模型区分了归一化记忆弛豫与可复现的盈余场(excess field)。一个四项条件机制恢复了种群漂移(r = 0.946;学习者自助法95%置信区间,0.935–0.955)。预测性事件级自监督学习恢复了该状态及学习平面流;以零模型为参照的校正保留了方向性和部分幅度的盈余场结构,但未实现完全校准。打乱顺序的训练会使有序轨迹上的学习平面流发生反转;支持对齐随机化则选择性地减少了向内输运。两个轴在无状态监督的情况下仍保持线性可读出。在未进行跨模型拟合的情况下,各模型共享主导种群漂移(r = 0.866;学习者自助法95%置信区间,0.857–0.875)和持续性排序;残差方向则因模型而异。这些结果确定了一个外部锚定的主导阶有效场,将经验动力学、可解释机制与神经计算联系起来。
cs.LG / 168 / 2608.30952

One Policy Is Enough: Single-Agent Reinforcement Learning Outperforms Tree Search for Chemistry Tool Learning

一种策略足矣:单智能体强化学习在化学工具学习中优于树搜索
Dariani, Armin, Wu, Sifan, Liu, Bang, Yang, Entao
Abstract
Chemistry questions often demand exact computation and database lookups that a language model cannot supply from its parameters, so it must reach for external tools. Tool use here is a three-part problem: select the right tool from a large pool, fill it with correctly typed arguments, and chain calls so that each consumes the outputs of the last. CheMatAgent, a previously published system, addresses this with hierarchical evolutionary MCTS: separate policy and execution models searching tool-call trees under two learned critics, one regressed partly onto GPT-assigned scores. We show that a single policy suffices. Our model interleaves reasoning, tool calls, and returns in one left-to-right generation, trained by a supervised warm-up and then outcome-level reinforcement learning against a programmatic reward read directly off the gold call chain, which leaves no learned critic and no judge in the training loop. On ChemToolBench multiple-tool comprehensive chemistry, on both backbones CheMatAgent use, we improve Tool F1 by 5.5% and Return F1 by 9.6% on Qwen-2.5-7B, and by 3.7% and 3.9% on Llama-3.1-8B, compared with their strongest search configuration, at one model invocation per question, against a search whose cost grows with the tree; we also lead answer Pass Rate on Qwen-2.5-7B.
Chinese Translation
化学问题往往需要精确的计算和数据库查询,而这些无法由语言模型仅凭其参数完成,因此模型必须调用外部工具。工具使用本质上是一个由三部分组成的问题:从大量工具中选择正确的工具、填入类型正确的参数,并将调用链接起来,使每次调用都能使用上一次调用的输出。此前发表的CheMatAgent系统采用分层进化式蒙特卡洛树搜索(MCTS)来应对这一问题:在两个学习到的评论家(critic)模型(其中一个部分以GPT打分的回归值为目标)的指导下,由独立的策略模型和执行模型搜索工具调用树。我们的研究表明,仅需一个策略模型即可足够。我们的模型在单次从左到右的生成中交替进行推理、工具调用和返回结果,先经过监督预热训练,随后通过针对程序化奖励(直接从标准调用链中读取)的结果级强化学习进行训练,训练循环中既没有学习到的评论家,也没有评判模型。在ChemToolBench多工具综合化学基准上,在CheMatAgent使用的两种骨干模型上,与各自最强的搜索配置相比,我们以每个问题仅需一次模型调用的代价(而搜索的成本随树规模增长),在Qwen-2.5-7B上将Tool F1提升了5.5%、Return F1提升了9.6%,在Llama-3.1-8B上分别提升了3.7%和3.9%;此外,我们在Qwen-2.5-7B上的答案通过率(Pass Rate)也处于领先地位。
cs.LG / 169 / 2608.30960

Hard-ReLU Gradient Descent Selects an Event-Free Sensitivity Limit

Hard-ReLU梯度下降选择一个无事件敏感性极限
Li, Xiaoyang, Zhou, Runni
Abstract
Gradient flow is widely used as a continuous-time surrogate for gradient descent, but state convergence does not imply convergence of differentiated training maps in nonsmooth networks. We characterize the fixed-horizon, vanishing-step limit of exact automatic differentiation through hard-ReLU gradient descent. Under a stable finite itinerary of separated, same-direction transverse activation events, gradient-descent states converge at first order to the corresponding piecewise-smooth gradient flow, while the exact derivative of every nonresonant discrete program converges to an event-free regional propagator. The true flow derivative instead interleaves classical saltation matrices that encode event-time sensitivity. For globally convex objectives, any strict activation event prevents complete cancellation of these missing transfers. Moreover, minimal globally 1-strongly convex residual-ReLU risks can realize arbitrarily large reciprocal sensitivity gaps, subject to an explicit transversality-scale tradeoff, and a coupled strongly convex construction yields an open set on which the largest initialization-gradient coordinate is reversed. In a controlled 17-parameter ReLU MLP, state and regional-AD errors vanish under mesh refinement while AD-to-flow errors remain between 0.18 and 0.39; an event-aware corrected product restores convergence. Resolved smoothing likewise recovers the flow sensitivity when the transition layer is sufficiently resolved. These results show that the gradient-flow limit of hard-ReLU training need not remain valid after differentiation.
Chinese Translation
梯度流(gradient flow)被广泛用作梯度下降的连续时间替代模型,但在非光滑网络中,状态收敛并不意味着微分训练映射的收敛。我们刻画了经过Hard-ReLU梯度下降的精确自动微分在固定时域、步长趋近于零时的极限。在分离的、同方向的横向激活事件构成的稳定有限事件序列(itinerary)条件下,梯度下降状态以一阶精度收敛到相应的分段光滑梯度流,而每个非共振离散程序的精确导数收敛到一个无事件的区域传播子(regional propagator)。真实的流导数则交错包含编码事件时间敏感性的经典盐跃矩阵(saltation matrices)。对于全局凸目标函数,任何严格的激活事件都会阻止这些缺失转移的完全抵消。此外,最小的全局1-强凸残差ReLU风险可以实现任意大的倒数敏感性差距(受一个明确的横向性尺度权衡约束),并且一个耦合的强凸构造产生一个开集,在该开集上最大的初始化梯度坐标被反转。在一个受控的17参数ReLU多层感知机(MLP)中,状态误差和区域AD误差随网格细化而消失,而AD与流之间的误差保持在0.18至0.39之间;引入事件感知的修正乘积可恢复收敛性。当过渡层被充分解析时,通过解析平滑同样可以恢复流敏感性。这些结果表明,Hard-ReLU训练的梯度流极限在微分之后未必仍然有效。
cs.LG / 170 / 2608.30963

A Universal Context-Reuse Layer for Cross-Model KV Sharing

一种用于跨模型KV共享的通用上下文复用层
Li, Yi, Jiang, Dongming, Zhao, Yi, Li, Bingzhe
Abstract
Modern large language model (LLM) serving systems increasingly operate over repeated or shared context, yet each model typically performs its own prefill computation even when another model has already processed the same input. Existing KV-cache reuse mechanisms substantially reduce redundant computation within a single model, but generally assume that the producer and consumer of a cache are identical. We study \emph{cross-model KV sharing}, which translates the KV state produced by a source model into a representation that can be consumed by a different target model, including models that differ in scale, architecture, attention configuration, tokenizer, and model family. We evaluate the approach in both within-family and cross-family settings. For Qwen2.5-7B $\rightarrow$ Qwen2.5-1.5B, translated KV states improve LongBench2 accuracy from 27.59\% to 34.48\%, a gain of 6.89 percentage points over the native 1.5B baseline, while reducing handoff cost relative to native target prefill. For the cross-family Qwen2.5-1.5B $\rightarrow$ Gemma-2-2B setting, KV handoff reduces target-side prefill cost by up to 67.05\% at 4K context length while maintaining decoding perplexity close to native-model baselines. In a more heterogeneous Llama3.1-70B $\rightarrow$ Qwen2.5-7B setting, cross-family handoff achieves 44.0\% accuracy compared with 45.7\% for native Qwen2.5-7B inference, while reducing measured latency from 899ms to 138ms. These results provide initial evidence that KV states can serve as transferable computational representations rather than strictly model-local caches, and motivate \emph{context mobility} as a systems abstraction for reducing redundant prefill across heterogeneous LLM and multi-agent inference workflows.
Chinese Translation
现代大语言模型(LLM)服务系统日益需要处理重复或共享的上下文,但即使另一个模型已经处理过相同的输入,每个模型通常仍会执行自己的预填充(prefill)计算。现有的KV缓存复用机制能够大幅减少单个模型内部的冗余计算,但通常假设缓存的生产者和消费者是同一个模型。我们研究了跨模型KV共享(cross-model KV sharing),即将源模型产生的KV状态转换为可被不同目标模型使用的表示形式,包括在规模、架构、注意力配置、分词器和模型系列上存在差异的模型。我们在同系列和跨系列两种设置下评估了该方法。在Qwen2.5-7B → Qwen2.5-1.5B设置中,转换后的KV状态将LongBench2准确率从27.59%提升至34.48%,相比原生1.5B基线提高了6.89个百分点,同时相比目标模型的原生预填充降低了交接成本。在跨系列的Qwen2.5-1.5B → Gemma-2-2B设置中,KV交接在4K上下文长度下可将目标侧预填充成本最多降低67.05%,同时保持解码困惑度接近原生模型基线。在更异构的Llama3.1-70B → Qwen2.5-7B设置中,跨系列交接实现了44.0%的准确率(原生Qwen2.5-7B推理为45.7%),同时将实测延迟从899ms降低至138ms。这些结果提供了初步证据,表明KV状态可以作为可迁移的计算表示而非严格限于单个模型的缓存,并促使上下文移动性(context mobility)成为减少异构LLM与多智能体推理工作流中冗余预填充的系统抽象。
cs.LG / 171 / 2608.30976

A Human-in-the-Loop Autonomous Agent for Industry Time Series Forecasting

面向工业时间序列预测的人机协同自主智能体
Tao, Xiaoyu, Cheng, Mingyue, Guo, Ze, Pan, Bokai, Liu, Qi, Wang, Shijin, Chen, Enhong
Abstract
Real-world time-series forecasting is rarely a one-shot model invocation: practitioners must formulate tasks, connect data and models, incorporate domain expertise, assess prediction plausibility, and communicate uncertainty. Specialized forecasting models provide strong numerical predictions but usually operate in fixed pipelines, while general-purpose large language model (LLM) agents often lack forecasting-specific checks, constraints, and stopping rules. We present CastClaw, a human-in-the-loop autonomous forecasting system built through forecasting-oriented harness engineering. CastClaw connects data, specialized models, analytical tools, user input, and a versioned execution record in one runtime. Users specify the target, horizon, constraints, and hypotheses in natural language. Starting from a supplied or model-generated forecast, CastClaw checks temporal patterns and user constraints; when evidence is missing, it retrieves context, runs an analysis or another model, or asks the user. It then keeps, revises, or escalates the result under explicit stopping conditions. The output contains the final forecast and an execution report recording inputs, evidence, actions, and revisions. In this five-dataset electricity-price setting, CastClaw reports the lowest point-estimate MSE and MAE among 16 baselines. A Nord Pool case demonstrates the inspectable workflow. CastClaw was also validated offline on provincial electricity-load data from North China covering January--June 2026.
Chinese Translation
现实世界中的时间序列预测很少是单次模型调用即可完成的:从业者需要明确任务、连接数据与模型、融入领域知识、评估预测的合理性并传达不确定性。专用预测模型能提供较强的数值预测,但通常运行在固定的流程中;而通用大语言模型(LLM)智能体往往缺乏针对预测的检验机制、约束条件和停止规则。我们提出了 CastClaw,一个通过面向预测的工程化框架构建的人机协同(human-in-the-loop)自主预测系统。CastClaw 将数据、专用模型、分析工具、用户输入以及带版本管理的执行记录集成于同一运行时环境中。用户以自然语言指定预测目标、时间范围、约束条件和假设。CastClaw 从给定或模型生成的预测出发,检查时间模式和用户约束;当证据不足时,它会检索上下文、运行分析或调用其他模型,或向用户询问。随后,系统在明确的停止条件下保留、修改或上报预测结果。系统输出包含最终预测结果以及一份记录输入、证据、操作和修改过程的执行报告。在一个涵盖五个数据集的电力价格预测场景中,CastClaw 在 16 个基线方法中取得了最低的点估计 MSE 和 MAE。一个 Nord Pool 案例展示了可查验的工作流程。此外,CastClaw 还在覆盖 2026 年 1 月至 6 月的华北省级电力负荷数据上进行了离线验证。
cs.LG / 172 / 2608.30978

Sparse Competition during Training For the Emergence of Specialized Modules

训练过程中的稀疏竞争促进专门化模块的涌现
Rossigneux, Baptiste, Haroun, Karim
Abstract
Modularity in deep neural networks has been proposed as a means of improving both interpretability and training by promoting disentangled representations and reducing redundancy. In this work, we study the emergence of modular structure through competition dynamics between groups of neurons during training. We introduce a method that (i) maintains near-baseline accuracy, (ii) induces usage-based modularity by sparsely routing inputs to neuron groups, and (iii) encourages specialization of these modules, such that their activations are correlated with input classes. We evaluate the proposed approach on ImageNet-100 and CIFAR-100 and show that with it, specialized modules emerge without module-level supervision. These modules capture a meaningful high-level structure in the data, with individual modules responding to semantic categories (e.g., dogs or vehicles). We also study the emergence of a hierarchical partition of sub-tasks depending on the number of modules. Our results suggest that competitive dynamics can serve as a simple mechanism for inducing functional modularity in standard architectures.
Chinese Translation
深度神经网络中的模块化被认为是一种提升可解释性与训练效果的手段,其通过促进解耦表示和减少冗余来实现。在本工作中,我们研究了训练过程中神经元组之间通过竞争动力学产生模块化结构的过程。我们提出了一种方法,该方法(i)保持接近基线的准确率,(ii)通过将输入稀疏地路由到神经元组,诱导出基于使用模式的模块化,以及(iii)鼓励这些模块的专门化,使其激活与输入类别相关联。我们在 ImageNet-100 和 CIFAR-100 上评估了所提出的方法,结果表明无需模块级别的监督即可涌现出专门化模块。这些模块捕捉到了数据中有意义的高层结构,单个模块会对语义类别(例如狗或车辆)产生响应。我们还研究了根据模块数量的不同,子任务层次化划分的涌现现象。我们的结果表明,竞争动力学可以作为一种在标准架构中诱导功能模块化的简单机制。
cs.LG / 173 / 2608.30986

Controlling Refusal Behavior of LLMs via Stiefel-Constrained Rotation Steering

基于Stiefel约束旋转引导的大语言模型拒绝行为控制
Bunin, Kirill, Bylinkin, Dmitry, Aletov, Vladimir, Medyakov, Daniil, Solodkin, Vladimir, Beznosikov, Aleksandr
Abstract
Activation steering has emerged as a lightweight approach for controlling model refusal at inference time. A growing line of research explores trainable rotations of activations to develop geometrically principled intervention mechanisms. However, existing techniques rely on auxiliary constructs, such as refusal vectors, to define these rotations. In our work, we develop a self-contained methodology for learning parameter-efficient rotational transformations based on Riemannian optimization. We empirically validate the proposed scheme, demonstrating its superiority in intervention efficiency. An extensive ablation study highlights the importance of key design choices in our method. Our results identify the proposed rotation-based steering scheme as a promising direction for more reliable control over the behavior of LLMs.
Chinese Translation
激活引导(activation steering)已成为一种在推理阶段控制模型拒绝行为的轻量级方法。越来越多的研究探索通过可训练的激活旋转来构建具有几何原理依据的干预机制。然而,现有技术依赖于辅助构造(如拒绝向量)来定义这些旋转。在本工作中,我们基于黎曼优化开发了一种自成一体的方法,用于学习参数高效的旋转变换。我们通过实验验证了所提出的方案,展示了其在干预效率方面的优越性。广泛的消融研究凸显了该方法中关键设计选择的重要性。我们的结果表明,所提出的基于旋转的引导方案是实现更可靠的大语言模型行为控制的一个有前景的方向。
cs.LG / 174 / 2608.31009

Language-Informed Flow Matching for Trend-Guided Structure-Based 3D Molecular Generation

面向趋势引导的基于结构三维分子生成的语言感知流匹配方法
Gao, Tianyu, Su, Zhikai, Li, Jiashu, Gao, Wenjun, Ying, Zichuan, Zhao, Zhe, Zhang, Fei, Wei, Ye
Abstract
Structure-based drug design (SBDD) requires ligands that satisfy both 3D target affinity and 1D chemical validity. Existing controllable generation methods often rely on task-specific fine-tuning or externally imposed sampling-time guidance, adding cost and potentially conflicting with evolving 3D geometric constraints. We propose LiFT, a language-informed cross-modal framework built on Flow Matching for trend-guided 3D molecular generation across both de novo design and scaffold hopping. LiFT uses a "Sense-Evolve-Assemble" agent to generate target-aware SMILES as intermediate chemical conditions, from which a pre-trained chemical foundation model extracts continuous semantic priors. These priors are integrated into geometric generation through a lightweight semantic projector with zero-initialized adaptive normalization for stable cross-modal conditioning. We further introduce a Self-Conditioned Decoupled Router (SCDR), which modulates the velocity field according to intermediate structural states during ODE integration. Experiments on Cross-Docked2020 show that LiFT achieves competitive distribution matching while improving medicinal chemistry metrics and maintaining competitive structural validity under task-steering settings without additional generator fine-tuning. Our results suggest that language-derived chemical priors provide effective trend-level guidance for 3D molecular generation. Code and released artifacts are available at https://github.com/kasurl/LiFT.
Chinese Translation
基于结构的药物设计(SBDD)要求配体同时满足三维靶标亲和力和一维化学有效性。现有的可控生成方法通常依赖于任务特定的微调或外加的采样时引导,这不仅增加了成本,还可能与不断演进的三维几何约束相冲突。我们提出了 LiFT,一个基于流匹配(Flow Matching)构建的语言感知跨模态框架,用于在从头设计和骨架跃迁两种场景下进行趋势引导的三维分子生成。LiFT 采用“感知—演化—组装”(Sense-Evolve-Assemble)智能体生成具有靶标感知能力的 SMILES 作为中间化学条件,并通过预训练的化学基础模型从中提取连续的语义先验。这些先验通过一个轻量级语义投影器融入几何生成过程,该投影器采用零初始化的自适应归一化以实现稳定的跨模态条件化。我们进一步引入了自条件解耦路由器(Self-Conditioned Decoupled Router, SCDR),在常微分方程(ODE)积分过程中根据中间结构状态对速度场进行调制。在 Cross-Docked2020 数据集上的实验表明,LiFT 在任务引导设定下无需额外的生成器微调,即可在实现具有竞争力的分布匹配的同时提升药物化学指标,并保持具有竞争力的结构有效性。我们的结果表明,源自语言的化学先验能够为三维分子生成提供有效的趋势层面引导。代码及相关发布资源可在 https://github.com/kasurl/LiFT 获取。
cs.LG / 175 / 2608.31013

TSPFN: A Temporal Tabular Foundation Model for Physiological Time Series Classification

TSPFN:面向生理时间序列分类的时间型表格基础模型
Stym-Popper, Jérémie, Rambour, Clément, Granese, Federica, Thome, Nicolas, Bernard, Olivier
Abstract
Designing models that generalize effectively in low- to medium-data regimes remains a primary challenge in medical machine learning, particularly for physiological time-series classification. While tabular foundation models such as TabPFN offer an attractive alternative to conventional fine-tuning through in-context learning, they are not designed to capture the temporal dependencies inherent to physiological signals. ~In this paper, we introduce TSPFN, a foundation model that redesigns TabPFN's architecture for time series data. TSPFN integrates structured temporal representations and positional embeddings to capture intra-sample temporal and channel dependencies. To fully leverage its spatio-temporal design, the model is pretrained on 140,000 real-world physiological time series across multiple medical domains. This yields a unified, generalizable framework capable of learning the specificities of medical time series. Experiments across diverse physiological benchmarks demonstrate that TSPFN consistently outperforms standard tabular baselines and TabPFN, and achieves superior cross-domain generalization compared to specialized deep time-series models. All our experiments, ablation studies, and pre-processing scheme are publicly available at https://github.com/Jeremstym/TSPFN
Chinese Translation
在小样本至中等规模数据条件下设计具有良好泛化能力的模型,仍然是医学机器学习领域的核心挑战,生理时间序列分类尤为如此。虽然 TabPFN 等表格基础模型通过上下文学习(in-context learning)提供了区别于传统微调的有吸引力替代方案,但其设计并未考虑生理信号中固有的时间依赖性。本文提出 TSPFN,一个针对时间序列数据重新设计 TabPFN 架构的基础模型。TSPFN 融合了结构化时间表示和位置嵌入,以捕获样本内部的时间与通道依赖关系。为充分发挥其时空设计的优势,该模型在跨多个医学领域的 140,000 条真实世界生理时间序列上进行了预训练,从而形成一个统一且可泛化的框架,能够学习医学时间序列的特有性质。在多个生理信号基准上的实验表明,TSPFN 始终优于标准表格基线模型和 TabPFN,并且在跨领域泛化能力上超越了专门的深度时间序列模型。所有实验、消融研究及预处理方案均已公开发布于 https://github.com/Jeremstym/TSPFN
cs.LG / 176 / 2608.31036

Normalized Low-Rank Adaptation

归一化低秩适配
Kang, Jiale, Yue, Ziyin, Zhan, Zheng, Huang, Yangyi, Liu, Weiyang
Abstract
While low-rank adaptation (LoRA) is widely used for parameter-efficient model adaptation, how to regularize its training dynamics for stable and effective optimization remains underexplored. Because LoRA initializes the up-projection to zero, its early optimization dynamics are largely governed by the down-projection. Building on this observation, we introduce Normalized Low-Rank Adaptation (NoRA), a simple yet effective method that normalizes the down-projection matrices during training. We further show that the same normalization can be applied only at initialization, improving standard LoRA without requiring repeated normalization throughout training. Across pretraining, supervised finetuning, and reinforcement learning, NoRA consistently accelerates convergence, improves performance and training stability, and mitigates catastrophic forgetting. These benefits require neither additional trainable parameters nor inference-time computation, making NoRA a simple and broadly applicable enhancement to LoRA.
Chinese Translation
尽管低秩适配(LoRA)被广泛用于参数高效的模型适配,但如何对其实施正则化以实现稳定而有效的优化动态仍缺乏充分研究。由于LoRA将上投影矩阵初始化为零,其早期优化动态主要由下投影矩阵决定。基于这一观察,我们提出了归一化低秩适配(Normalized Low-Rank Adaptation,NoRA),这是一种简单而有效的方法,它在训练过程中对下投影矩阵进行归一化。我们进一步证明,同样的归一化可以仅在初始化时应用,从而在不需在训练过程中反复归一化的情况下改进标准LoRA。在预训练、监督微调和强化学习中,NoRA始终能加速收敛、提升性能与训练稳定性,并缓解灾难性遗忘。这些好处既不需要额外的可训练参数,也不增加推理时的计算量,使NoRA成为对LoRA的一种简单且广泛适用的增强方法。
cs.LG / 177 / 2608.31045

Rotational Equivariance in Machine Learning: A Comprehensive Tutorial

机器学习中的旋转等变性:全面教程
Lippmann, Peter, Hamprecht, Fred A.
Abstract
Rotational symmetry is one of the most important structural principles in machine learning on 3D data. In applications ranging from physics and materials science to 3D computer vision, predictions should not depend on an arbitrary choice of coordinate frame. Rotational equivariance captures this requirement mathematically by enforcing that a rotation of the input induces a corresponding transformation of the model output. This tutorial provides a comprehensive introduction to rotational equivariance, starting from the physical and geometric intuition behind coordinate independence and building up the necessary machinery from geometric deep learning, group theory, and representation theory. We introduce message passing on Euclidean graphs, group actions and representations, spherical harmonics, Wigner matrices, tensor products, and Clebsch-Gordan decomposition, and explain how these ingredients give rise to modern equivariant architectures. We then survey the principal strategies for incorporating rotational equivariance in deep learning, including group convolutions, internal tensorial representations, and canonicalization-based methods, and discuss their practical strengths and limitations. The tutorial aims to lower the barrier to the subject by connecting the underlying mathematics to practical model design, by unifying ideas that are often expressed in different formal languages, and by helping practitioners choose among competing approaches through a clear discussion of their trade-offs.
Chinese Translation
旋转对称性是三维数据机器学习中最重要的结构原则之一。在从物理、材料科学到三维计算机视觉的应用中,预测结果不应依赖于坐标系的任意选取。旋转等变性通过强制要求输入的旋转引起模型输出的相应变换,从数学上刻画了这一要求。本教程对旋转等变性进行了全面介绍,从坐标独立性背后的物理与几何直觉出发,逐步构建几何深度学习、群论与表示论中所需的数学工具。我们介绍了欧几里得图上的消息传递、群作用与表示、球谐函数、Wigner矩阵、张量积以及Clebsch-Gordan分解,并解释这些要素如何催生了现代等变架构。随后,我们综述了在深度学习中引入旋转等变性的主要策略,包括群卷积、内部张量表示以及基于规范化的方法,并讨论了它们各自的实际优势与局限。本教程旨在通过将底层数学与实际模型设计相联系、统一以不同形式语言表达的思想,并通过对各方法权衡的清晰讨论帮助实践者在相互竞争的方法中做出选择,从而降低进入该领域的门槛。
cs.LG / 178 / 2608.31046

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

在线策略蒸馏真的在蒸馏吗?从噪声教师到自我改进
Ding, Yi, Zhang, Ruqi
Abstract
On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging to comparable performance regardless of whether noisy supervision is retained or removed. Does OPD distill at all? By analyzing what drives its gains, we find that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones. This suggests that OPD works largely by suppressing low log-probability tokens, which requires no teacher. These findings motivate On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens, and evenly redistributing probability mass among head tokens. Compared with the base \texttt{Qwen3-1.7B}, OPSA improves Avg@32 by 35.41 points on AIME24, corresponding to a 263\% relative gain, and more than doubles Pass@32 across all three benchmarks. It also outperforms OPD by 16.77 points in Avg@32 on AIME24. Extensive experiments and analyses across model families and tasks further demonstrate its effectiveness and generalizability.
Chinese Translation
在线策略蒸馏(On-Policy Distillation, OPD)提供了密集的token级监督,作为可验证奖励强化学习(RLVR)稀疏结果级优势的替代方案。然而,教师对学生生成的轨迹进行打分时,这些轨迹对教师而言本质上是离策略的,因此其监督的可靠性以及学生提升的来源仍不明确。我们定量分析了OPD训练过程中教师监督的情况,发现存在大量噪声,且噪声的普遍程度随教师模型规模的增大而增加。令人惊讶的是,学生策略对此类噪声并不敏感,无论是否保留噪声监督,其最终性能都相当。那么OPD究竟是否在蒸馏?通过分析其性能提升的驱动因素,我们发现学习集中于低对数概率的token,且使用单一固定的负优势即可达到与教师提供的优势相当的性能。这表明OPD在很大程度上是通过抑制低对数概率的token来起作用的,而这无需教师。基于这些发现,我们提出了在线策略自适应(On-Policy Self-Adaptation, OPSA),这是一种无需监督的方法,采用熵自适应负优势。该方法为高熵位置分配更强的学习信号,抑制尾部token,并在头部token之间均匀重新分配概率质量。与基座模型Qwen3-1.7B相比,OPSA在AIME24上使Avg@32提升了35.41分,相当于263%的相对提升,并在全部三个基准上将Pass@32提高了一倍以上。在AIME24上,OPSA的Avg@32还比OPD高出16.77分。跨模型家族和任务的大量实验与分析进一步证明了其有效性和泛化能力。
cs.LG / 179 / 2608.31067

Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers

用于电路计算的通用Transformer:微型Transformer中的完美长度泛化
Ito, Takuya, Puri, Ruchir, Campbell, Murray, Ram, Parikshit
Abstract
Learning generalizable algorithmic computations remains a challenge for neural networks, as reflected in persistent failures on compositional and length generalization benchmarks. We present a provably correct, transformer parameterization (with only 280 learnable parameters for Boolean algebra tasks) capable of learning and evaluating problems of any depth or length. We assume inputs are fully parenthesized, well-formed expressions. Our approach conceptualizes algorithmic tasks as circuit models embedded in transformers, enabling depth-1 circuit reduction in a single forward pass. To achieve depth generalization, we introduce a positional encoding that tracks each gate's depth within the circuit, enabling the model to identify evaluable subexpressions at each iteration via masked hard attention, with $O(n)$ per-iteration complexity via linear attention. Combined with an autonomous halting criterion, the model terminates after $d$ iterations for problems of depth $d$, yielding $O(n \cdot d)$ total complexity. We show that training on shallow problem instances (depth 1 and depth 2) effectively recovers interpretable parameters that {\em snap} into place, resulting in exact length generalization. Though we establish that our construction provably evaluates Boolean expressions -- a universal symbolic computation -- of arbitrary length perfectly, in other experiments we also demonstrate that our transformer variant can learn and generalize perfectly (100% accuracy) on other common length generalization benchmarks, including modular arithmetic and ListOps.
Chinese Translation
学习可泛化的算法计算仍然是神经网络面临的一项挑战,这体现在模型在组合泛化和长度泛化基准测试中的持续失败上。我们提出了一种可证明正确的Transformer参数化方法(针对布尔代数任务仅需280个可学习参数),能够学习和求解任意深度或长度的问题。我们假设输入为完全加括号的良构表达式。我们的方法将算法任务概念化为嵌入在Transformer中的电路模型,从而能够在单次前向传播中完成深度为1的电路约简。为实现深度泛化,我们引入了一种位置编码来跟踪每个门电路在电路中的深度,使模型能够通过掩码硬注意力在每次迭代中识别可求值的子表达式,并通过线性注意力实现每次迭代 $O(n)$ 的复杂度。结合自主停止准则,模型对于深度为 $d$ 的问题在 $d$ 次迭代后终止,总复杂度为 $O(n \cdot d)$。我们证明,仅在浅层问题实例(深度1和深度2)上训练即可有效恢复出可解释的、能够精确"归位"(snap)的参数,从而实现精确的长度泛化。尽管我们已证明该构造能够完美地对任意长度的布尔表达式——一种通用的符号计算——进行求值,我们还在其他实验中展示了我们的Transformer变体能够在其他常见的长度泛化基准(包括模运算和ListOps)上实现完美学习与泛化(100%准确率)。
cs.LG / 180 / 2608.31069

A Model with No Head and Many Thoughts

无头多思:一种无输出头且具备多重思维的模型
Koriagin, Nikita, Aksenov, Yaroslav, Bredis, George, Gerasimov, Gleb, Balagansky, Nikita, Gavrilov, Daniil
Abstract
Large language models decode by projecting hidden states through a large vocabulary head at every step. This operation is computationally costly and forces all reasoning to be expressed in discrete tokens. We introduce Soft Latent Thinking, a method that replaces the LM head during reasoning with a lightweight projector, enabling autoregressive rollout in embedding space where reasoning steps remain continuous rather than tokenized. Experiments on DeepSeek-Qwen-1.5B and LLaMA-3.2-3B show that Soft Latent Thinking consistently improves pass@k across all k while reducing per-step compute during chain-of-thought. Our method achieves the highest pass@32 among all soft-thinking approaches, demonstrating that effective reasoning can be carried out in continuous space without discrete token generation.
Chinese Translation
大型语言模型在每一步解码时都需要通过庞大的词表输出头(LM head)对隐藏状态进行投影。这一操作计算代价高昂,并迫使所有推理过程以离散词元(token)的形式表达。我们提出了软性潜在思维(Soft Latent Thinking),一种在推理阶段用轻量级投影器替代语言模型输出头的方法,从而实现嵌入空间中的自回归推演,使推理步骤保持连续而非被词元化。在 DeepSeek-Qwen-1.5B 和 LLaMA-3.2-3B 上的实验表明,软性潜在思维在降低思维链(chain-of-thought)每步计算量的同时,能在所有 k 值上一致地提升 pass@k 指标。在所有软性思维(soft-thinking)方法中,我们的方法取得了最高的 pass@32,这表明有效的推理可以在连续空间中进行,而无需离散词元的生成。
cs.LG / 181 / 2608.31079

Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization

通过对比偏好优化,谄媚性附和可经由中性数据发生迁移
Blank, Camila, Ying, Zhuofan, Potts, Christopher, Hase, Peter, Huang, Jing
Abstract
Sycophantic agreement refers to a behavior in which language models excessively affirm the user, often at the cost of factual accuracy. Although sycophantic agreement is a well-known failure of model alignment, there is limited understanding of how it emerges from model training. In this work, we demonstrate that sycophantic agreement can emerge as an unintended consequence of widely used contrastive preference optimization objectives. Using the OLMo 3 post-training pipeline, we show that, for various pairs of teacher models across three families, there is a strong correlation between the log-ratio of the teacher model sycophantic agreement rates and the resulting student model sycophantic agreement rate. We further demonstrate that this unintended transfer is not limited to DPO but also occurs across 6 other preference optimization objectives. To understand whether this effect can be attributed to particular training examples, we analyze the preference data and find that the sycophancy signal is diffused across the entire dataset rather than concentrated in a sparse set of examples: each example appears neutral, i.e., there are no explicit instances of sycophantic agreement, and filtering based on probe-based data attribution or logit-linear selection fails to mitigate sycophancy without removing a large portion of the dataset. Overall, our findings suggest that the teacher models used to generate preference data can interact with alignment training objectives in unexpected ways, generalizing to undesirable and potentially harmful behaviors like sycophantic agreement.
Chinese Translation
谄媚性附和(sycophantic agreement)是指语言模型过度附和用户、往往以牺牲事实准确性为代价的行为。尽管谄媚性附和是模型对齐中一个广为人知的失败模式,但人们对其如何在模型训练过程中产生仍知之甚少。在本工作中,我们证明谄媚性附和可以作为广泛使用的对比偏好优化目标的一个意外副产品而出现。基于 OLMo 3 后训练流程,我们发现在三个模型家族的多种教师模型配对中,教师模型的谄媚附和率对数比与由此产生的学生模型的谄媚附和率之间存在强相关性。我们进一步证明,这种意外的迁移不仅限于 DPO,还会在其他 6 种偏好优化目标中出现。为探究这一效应能否归因于特定的训练样本,我们对偏好数据进行了分析,发现谄媚信号弥散于整个数据集之中,而非集中在少数样本上:每个样本看起来都是中性的,即不存在显式的谄媚性附和实例,且基于探针的数据归因或 logit 线性选择进行过滤,若不删除大部分数据集便无法有效缓解谄媚问题。总体而言,我们的研究结果表明,用于生成偏好数据的教师模型可能以意想不到的方式与对齐训练目标相互作用,从而泛化出谄媚性附和这类不良且潜在有害的行为。
cs.LG / 182 / 2608.31108

Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

高效负责任AI评估的压力测试:当计算节省改变基准测试结论时
Kady, Ahmed El, Narayanan, Aravind, Riaz, Rehana, Ioannou, Yani, Raza, Shaina
Abstract
Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. We stress-test conclusion robustness in responsible-AI benchmarking by evaluating three dense and mixture-of-experts models on BBQ and BBQ-V under seven conditions spanning batching, quantization, benchmark reduction, and their combinations. Rather than treating preserved aggregate accuracy as sufficient, we compare accuracy, bias severity and prevalence, reasoning quality, subgroup behavior, subset-membership stability, runtime, and measured GPU energy against a full-benchmark BF16 baseline. Larger batching keeps accuracy within 0.35 percentage points of baseline and produces comparatively small subgroup changes, while reducing energy in five of six model--dataset settings. INT8 largely preserves quality but uses 1.79--4.26$\times$ baseline energy. INT4 causes larger, model- and context-dependent changes. Reduced benchmarks provide the most consistent savings, but very small subsets are substantially more sensitive to which items are retained. Efficient evaluation should therefore be treated as a measurement intervention whose validity must be checked across the conclusions the benchmark is intended to support. Our project website is https://vectorinstitute.github.io/sustainable-rai-evaluation/ and the code is available at https://github.com/VectorInstitute/sustainable-rai-evaluation.
Chinese Translation
高效评估改变了用于支持模型行为结论的协议,然而这些结论在评估本身变得更为廉价之后是否依然稳定,却很少被检验。我们对负责任AI(responsible-AI)基准测试中的结论稳健性进行了压力测试:在七种涵盖批处理、量化、基准缩减及其组合的条件下,评估三个稠密模型和混合专家(mixture-of-experts)模型在BBQ和BBQ-V上的表现。我们并未将总体准确率的保持视为充分条件,而是将准确率、偏见严重程度与普遍性、推理质量、子群体行为、子集成员稳定性、运行时间以及实测GPU能耗与全基准BF16基线进行比较。更大的批处理使准确率保持在基线的0.35个百分点以内,子群体变化相对较小,同时在六种模型-数据集设置中的五种降低了能耗。INT8量化基本保持质量,但能耗为基线的1.79–4.26倍。INT4量化引起更大且依赖于模型和上下文的变化。缩减后的基准测试提供了最为一致的节省,但非常小的子集对保留哪些题目显著更为敏感。因此,高效评估应被视为一种测量干预,其有效性必须在其所支持的全部基准结论范围内加以检验。我们的项目网站为 https://vectorinstitute.github.io/sustainable-rai-evaluation/,代码可在 https://github.com/VectorInstitute/sustainable-rai-evaluation 获取。
cs.LG / 183 / 2608.31120

On the Complexity of the Compatibility Problem for Succinctly Encoded Conditional Distributions

关于简洁编码条件下条件分布相容性问题的复杂性
Emerson, Guy
Abstract
The motivation for this paper is the investigation of the trade-offs implicit in probabilistic models used in machine learning. Models are often used to make predictions in the form of conditional probabilities. However, a pair of conditional distributions p(x|y) and p(y|x) may not be compatible with any joint distribution p(x,y). Given two such conditionals, determining if there exists a compatible joint is known as the compatibility problem. For discrete random variables, when the conditionals are encoded as probability tables, the compatibility problem has a known solution, which is computationally tractable. In this paper, we formalise and study a succinct version of the problem, encoding conditional distributions as arithmetic circuits. This is applicable to practical applications of probabilistic modelling in high-dimensional settings, including neural network models. We show that, for succinct circuit representations of conditionals, the compatibility problem is intractable. In the case that all probabilities are non-zero, the problem is co-NP-complete. In the case that probabilities can be zero, we give examples to demonstrate that several notions of compatibility can be distinguished, and we prove that multiple versions of the problem are PSPACE-complete. Furthermore, we show that, assuming the polynomial hierarchy does not collapse, there exist compatible succinct conditionals whose joint cannot be expressed succinctly. Implications of these results for probabilistic modelling and machine learning are discussed.
Chinese Translation
本文的动机在于探究机器学习中概率模型所隐含的权衡关系。模型通常以条件概率的形式进行预测。然而,一对条件分布 p(x|y) 与 p(y|x) 可能与任何联合分布 p(x,y) 都不相容。给定两个这样的条件分布,判断是否存在相容的联合分布,这一问题被称为相容性问题(compatibility problem)。对于离散随机变量,当条件分布以概率表形式编码时,相容性问题是可计算的。本文形式化并研究了该问题的简洁版本,即将条件分布编码为算术电路(arithmetic circuits)。这适用于高维场景下概率建模的实际应用,包括神经网络模型。我们证明,对于条件分布的简洁电路表示,相容性问题是难解的。当所有概率均非零时,该问题是 co-NP 完全的。当概率可以为零时,我们给出若干例子以区分多种相容性概念,并证明该问题的多个版本是 PSPACE 完全的。此外,我们证明,在多项式层级不会坍塌的假设下,存在某些相容的简洁条件分布,其联合分布无法被简洁地表达。本文还讨论了这些结果对概率建模和机器学习的意义。
cs.LG / 184 / 2608.31157

Sharp Approximation Rates for Neural Networks with Affine Latent Parameterizations

具有仿射潜在参数化的神经网络的精确逼近速率
Zhang, Shijun
Abstract
Many parameter-efficient methods generate the parameters of a large neural network from a low-dimensional latent representation. Given an architecture $\Phi$ with $P_\Phi$ parameter slots, we write $\boldsymbol{\theta}_f=\mathcal{G}(\boldsymbol{\xi}_f)$, where $\mathcal{G}\colon\mathbb{R}^M\to\mathbb{R}^{P_\Phi}$ is a parameter generator and $\boldsymbol{\xi}_f\in\mathbb{R}^M$ is a latent representation of the target function $f$. The architecture $\Phi$ and the generator $\mathcal{G}$ are shared across the entire target class, while each target $f$ is represented by its own latent vector $\boldsymbol{\xi}_f$, with $\Phi_{\mathcal{G}(\boldsymbol{\xi}_f)}$ approximating $f$. This framework encompasses hypernetworks, low-dimensional parameterizations, parameter-efficient adaptation, and model compression. Understanding the tradeoff between the latent dimension $M$ and the network budget $P$ is therefore fundamental to characterizing the expressive efficiency of these methods. We study this tradeoff for affine generators and fully connected ReLU architectures. More precisely, optimizing jointly over architectures $\Phi$ satisfying $P_\Phi\leq P$ and affine generators $\mathcal{G}:\mathbb{R}^M\to \mathbb{R}^{P_\Phi}$, we prove that the optimal worst-case uniform approximation error over the unit ball of $\alpha$-H\"older functions on $[0,1]^d$, where $0<\alpha\leq1$, has the sharp order $ \bigl(P\min\{M,P\}\bigr)^{-\alpha/d}. $ In particular, our result shows that even a fixed-dimensional latent space suffices to achieve vanishing approximation error as the network budget increases.
Chinese Translation
许多参数高效方法通过低维潜在表示来生成大型神经网络的参数。给定一个具有 $P_\Phi$ 个参数槽位的架构 $\Phi$,我们记 $\boldsymbol{\theta}_f=\mathcal{G}(\boldsymbol{\xi}_f)$,其中 $\mathcal{G}\colon\mathbb{R}^M\to\mathbb{R}^{P_\Phi}$ 是参数生成器,$\boldsymbol{\xi}_f\in\mathbb{R}^M$ 是目标函数 $f$ 的潜在表示。架构 $\Phi$ 与生成器 $\mathcal{G}$ 在整个目标函数类中是共享的,而每个目标函数 $f$ 由其自身的潜在向量 $\boldsymbol{\xi}_f$ 表示,并通过 $\Phi_{\mathcal{G}(\boldsymbol{\xi}_f)}$ 来逼近 $f$。该框架涵盖了超网络(hypernetworks)、低维参数化、参数高效适配以及模型压缩。因此,理解潜在维度 $M$ 与网络预算 $P$ 之间的权衡,是刻画这些方法表达能力效率的基础。我们针对仿射生成器和全连接 ReLU 架构研究了这一权衡。更精确地说,通过在满足 $P_\Phi\leq P$ 的架构 $\Phi$ 与仿射生成器 $\mathcal{G}:\mathbb{R}^M\to \mathbb{R}^{P_\Phi}$ 上进行联合优化,我们证明了在 $[0,1]^d$ 上 $\alpha$-Hölder 函数(其中 $0<\alpha\leq1$)的单位球上的最优最坏情形一致逼近误差具有精确阶 $\bigl(P\min\{M,P\}\bigr)^{-\alpha/d}$。特别地,我们的结果表明,即使固定维数的潜在空间,也足以在神经网络预算增大时实现逼近误差趋于零。
cs.LG / 185 / 2608.31166

Constant Individual Regret in General Games

一般博弈中的常数个体遗憾
Liu, Mingyang, Farina, Gabriele, Ozdaglar, Asuman
Abstract
Uncoupled no-regret dynamics provide a decentralized route to equilibrium, but prior guarantees for individual regret retain a polylogarithmic dependence on the horizon. We remove this dependence for every finite $N$-player normal-form game under full-information feedback. We introduce \emph{ECHO-OFTRL}: optimistic follow-the-regularized-leader (OFTRL) equipped with an EMA cascade for high-order optimism (ECHO), where EMA denotes exponential moving average. The algorithm is deterministic and fully uncoupled. If $m_{\max}$ denotes the largest action-set size, then, simultaneously for every horizon $T\geq1$, it guarantees that each of the $N$ players in the game incurs regret upper bounded by $O(\textrm{poly}(N, \log m_{\max}))$. Our algorithm leverages a new form of optimism inspired by modern filter design.
Chinese Translation
无耦合(uncoupled)的无遗憾动力学为博弈均衡提供了一条去中心化的路径,但此前关于个体遗憾的保证仍存在对时间范围的多个对数级依赖。我们消除了这一依赖,适用于完全信息反馈下的所有有限 $N$ 人标准型博弈。我们提出了 \emph{ECHO-OFTRL}:即带有一阶指数移动平均级联以实现高阶乐观(EMA Cascade for High-order Optimism, ECHO)的乐观正则化领导者跟随算法(optimistic follow-the-regularized-leader, OFTRL),其中 EMA 表示指数移动平均。该算法是确定性的且完全无耦合的。若 $m_{\max}$ 表示最大的动作集规模,则该算法同时对所有时间范围 $T\geq1$ 保证博弈中的每位玩家(共 $N$ 位)的遗憾值上界为 $O(\textrm{poly}(N, \log m_{\max}))$。我们的算法利用了一种受现代滤波器设计启发的新型乐观机制。
机器人学 (Robotics)
68
cs.RO / 1 / 2608.28656

RedLight-VLA: Models for traffic-rule grounding and behavioral emphasis in driving policies

RedLight-VLA:面向交通规则落地与行为强化的驾驶策略模型
Sudhakar, Bala Murali Manoghar Sai, Sridhar, Sourab Bapu, Das, Sandipan, Ahuja, Rahul, Lazar, Meda, Garg, Ashish, Likhar, Pratik, Yogamani, Senthil
Abstract
Behavior-cloned Vision-Language-Action (VLA) driving policies struggle with rare rule-governed maneuvers at signalized intersections. Braking and launching examples contribute little to averaged trajectory loss, while fused representations lack explicit supervision for the governing traffic-light and stop-line state. We present RedLight-VLA, a training objective that uses expert futures and automatically generated perception targets without additional manual rule annotation. First, trajectory-derived behavioral reweighting (BR) emphasizes rare deceleration and acceleration using rotation-invariant longitudinal dynamics and a scale-preserving reduction that exactly recovers the baseline when disabled. Second, parallel auxiliary (AUX) heads ground traffic-light and stop-line state in continuous post-fusion rule tokens, without autoregressive language generation or changes to the trajectory decoder. We evaluate on a curated set of 20 s sequences with a 5 s prediction horizon. Controlled variants share the same backbone, training data, decoder, and evaluation population. Against an otherwise identical VLA baseline, RedLight-VLA reduces red-light stop-line overshoot from 7.3% to6.8%, reduces stop-line velocity error by 12.7%, and improves 3 s trafficlight-sliced ADE/FDE from 0.274/0.964 m to 0.247/0.897 m. Green-light false stops increase from 3.2% to 3.9%; however, combining BR with AUX supervision mitigates the larger increase observed for AUX alone (4.0%). The combined model also improves non-traffic-light ADE/FDE from 0.268/0.956 m to 0.241/0.876 m and outperforms either mechanism alone on all four sliced displacement measures.
Chinese Translation
通过行为克隆训练的视觉-语言-动作(VLA)驾驶策略在信号交叉口的罕见规则操控场景中表现不佳。刹车与起步样本对平均轨迹损失的贡献极小,而融合表征对起主导作用的红绿灯和停止线状态缺乏显式监督。我们提出RedLight-VLA,这是一种训练目标,它利用专家未来轨迹和自动生成的感知目标,无需额外的人工规则标注。首先,基于轨迹的行为重加权(BR)利用旋转不变纵向动力学和保持尺度的约减方式来强调罕见的减速与加速行为,且在该机制关闭时可精确恢复基线。其次,并行的辅助(AUX)头在融合后的连续规则标记(rule tokens)中落地交通灯和停止线状态,无需自回归语言生成,也无需修改轨迹解码器。我们在一个精选的20秒序列数据集上以5秒预测时域进行评估。各受控变体共享相同的骨干网络、训练数据、解码器和评估样本集。与完全相同的VLA基线相比,RedLight-VLA将红灯停止线越过率从7.3%降至6.8%,停止线速度误差降低12.7%,并将3秒红绿灯分片ADE/FDE从0.274/0.964米改善至0.247/0.897米。绿灯误停率从3.2%升至3.9%;然而,将BR与AUX监督相结合可缓解仅使用AUX时观察到的更大增幅(4.0%)。组合模型还将非红绿灯场景的ADE/FDE从0.268/0.956米改善至0.241/0.876米,并在全部四项分片位移指标上均优于任一单独机制。
cs.RO / 2 / 2608.28664

The Potential of Haptic Foundation Models

触觉基础模型的潜力
Wang, Jianquan, Dong, Haiwei, Saddik, Abdulmotaleb El
Abstract
Despite the success of foundation models in language and vision, their expansion into embodied AI is bottlenecked by a lack of generalized touch sensing. This limitation is especially relevant to consumer electronics, where smartphones, wearables, VR controllers, home robots, and health monitoring devices require safe and adaptive physical interaction. Constrained by hardware heterogeneity and the necessity of active physical data collection, current haptic models remain rigidly task-specific. To overcome these limitations, this article explores the transformative potential and developmental trajectory of Haptic Foundation Models (HFMs). We detail the paradigm shift required to transition from passive Large Language Models and Vision Language Models into active HFMs across four core dimensions: action coupling, physical dynamical representation space, continuous time-series data granularity, and action-conditioned future state prediction. Furthermore, we synthesize existing large-scale tactile datasets and benchmark UniTouch, AnyTouch, T3, and Sparsh on TacBench for force estimation, slip detection, and relative pose estimation.
Chinese Translation
尽管基础模型在语言和视觉领域取得了成功,但其在具身智能(embodied AI)中的扩展因缺乏广义的触觉感知而受到瓶颈制约。这一局限对消费电子领域尤为相关:智能手机、可穿戴设备、VR 控制器、家用机器人以及健康监测设备都需要安全且自适应的物理交互。受硬件异构性和必须主动采集物理数据的限制,当前的触觉模型仍然严格局限于特定任务。为克服这些局限,本文探讨了触觉基础模型(Haptic Foundation Models, HFMs)的变革性潜力与发展路径。我们详细阐述了从被动的 LLM 和 VLM 转变为主动式 HFMs 所需的范式转变,涵盖四个核心维度:动作耦合、物理动力学表征空间、连续时间序列数据粒度,以及动作条件下的未来状态预测。此外,我们梳理了现有的大规模触觉数据集,并在 TacBench 上对 UniTouch、AnyTouch、T3 和 Sparsh 在力估计、滑移检测和相对位姿估计任务上进行了基准测试。
cs.RO / 3 / 2608.28677

Cognitively-Grounded On-Device Runtime Learning for Ground Robots in Unknown Physical Environments

面向认知的地面机器人在未知物理环境中的端侧运行时学习
Cai, Yihao, Mao, Yanbing, Lebiere, Christian
Abstract
This paper presents \ul{CogRun}, a framework that enables safety-critical ground robots to perform cognitively-grounded runtime learning entirely on edge-AI devices in unknown physical environments, without prior maps or perceptual knowledge. CogRun consists of three components: a Learning-Agent, a Rational-Agent, and a Coordinator. The Learning-Agent is novel in cognitive-neural learning architecture, which featurs dedicated replay buffers, cognition-driven experience sampling, and a safety-aware action blending of actor-critic reinforcement learning (RL) with instance-based learning (IBL). The Rational-Agent is a non-learning module that complements the Learning-Agent by exclusively handling safety-critical functions, while the Coordinator manages interactions between the two agents to promote safe and efficient runtime learning. CogRun's full autonomy stack (i.e., perception, learning, and control) on edge-AI devices eliminates dependence on wireless communications, enabling broader applications in challenging environments with limited or no connectivity. Experiments on a quadruped robot in real-world wild forests and on an off-road autonomous vehicle in a simulated wild forest demonstrate that CogRun enables safe and efficient runtime learning, allowing robots to safely and continuously interact with the physical world for enhancing task performance in complex, unknown environments.
Chinese Translation
本文提出了CogRun框架,使安全攸关的地面机器人能够在未知物理环境中完全在边缘AI设备上执行基于认知的运行时学习,且无需先验地图或感知知识。CogRun由三个组件构成:学习智能体(Learning-Agent)、理性智能体(Rational-Agent)和协调器(Coordinator)。学习智能体的创新之处在于其认知-神经学习架构,其特点包括专用经验回放缓冲区、认知驱动的经验采样,以及将演员-评论家强化学习(RL)与基于实例的学习(IBL)相结合的安全感知动作融合机制。理性智能体是一个非学习模块,专门负责处理安全攸关功能,以补充学习智能体;协调器则管理两个智能体之间的交互,以促进安全高效的运行时学习。CogRun在边缘AI设备上实现了完整的自主堆栈(即感知、学习和控制),消除了对无线通信的依赖,使其能够在通信受限或无连接的复杂环境中获得更广泛的应用。在真实野外森林环境中的四足机器人实验以及模拟野外森林中的越野自动驾驶车辆实验表明,CogRun能够实现安全高效的运行时学习,使机器人能够安全、持续地与物理世界交互,从而提升在复杂未知环境中的任务性能。
cs.RO / 4 / 2608.28693

RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction

RoboGesture:面向人形机器人交互的实时语义一致性语音伴随手势生成
Wang, Zifan, Ren, Ziang, Shi, Pengyang, Wang, Zirui, Lin, Chenghuai, Wang, Tianze, Qi, Zekun, Zhao, Liangliang, Wang, He, Yi, Li
Abstract
Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot interaction. However, this task faces three critical barriers: the scarcity of semantically rich datasets, the "modality eclipse" where models ignore audio cues in favor of kinematic inertia, and the sim-to-real gap regarding physical safety. We propose RoboGesture, a robot-centric framework that co-designs data, modeling, and control to power a complete interactive human-humanoid system in which the robot listens, responds, and gestures in real time. We first establish the RoboGesture dataset featuring over 300 gesture categories and develop an automated pipeline to synthesize large-scale collision-free, robot-specific audio-motion pairs. Our architecture features a Hierarchical Semantic-Acoustic Aligner that extracts multi-granular prosodic and semantic cues directly from raw audio tokens. These cues drive a Streaming Conditional Motion Generator based on a diffusion transformer with conditional flow matching. To ensure high responsiveness, we introduce Anti-Inertia CFG Masking, which prevents the model from collapsing into repetitive historical patterns by compelling it to proactively mine control signals from the audio modality. Finally, an MPC-based safety filter ensures real-time, collision-free execution on physical hardware. Experiments on a Unitree G1 humanoid demonstrate that RoboGesture generates safer, more rhythmic, and more semantically appropriate responses compared to state-of-the-art baselines.
Chinese Translation
使人形机器人能够以同步且语义有意义的手势回应人类语音,是实现自然人机交互的基础。然而,该任务面临三大关键障碍:语义丰富数据集的匮乏、"模态遮蔽"问题(即模型忽略音频线索而依赖运动学惯性),以及物理安全性方面的模拟到现实差距。我们提出了RoboGesture,一个以机器人为中心的框架,通过数据、建模与控制的协同设计,驱动一个完整的人-人形机器人交互系统,使机器人能够实时倾听、回应并做出手势。我们首先建立了包含300多个手势类别的RoboGesture数据集,并开发了自动化流水线以合成大规模无碰撞的、针对特定机器人的音频-动作配对数据。我们的架构包含一个分层语义-声学对齐器(Hierarchical Semantic-Acoustic Aligner),可直接从原始音频token中提取多粒度的韵律与语义线索。这些线索驱动一个基于扩散Transformer(diffusion transformer)和条件流匹配的流式条件动作生成器(Streaming Conditional Motion Generator)。为确保高响应性,我们引入了抗惯性CFG掩码(Anti-Inertia CFG Masking),通过迫使模型主动从音频模态中挖掘控制信号,防止模型坍缩为重复的历史模式。最后,基于模型预测控制(MPC)的安全过滤器确保了在物理硬件上的实时无碰撞执行。在Unitree G1人形机器人上的实验表明,与最先进的基线方法相比,RoboGesture能够生成更安全、更有节奏感且语义更恰当的动作响应。
cs.RO / 5 / 2608.28697

Multi-Group Pipe Routing under Permanent Geometric Occupancy: Problem, Benchmark, and Classical Baselines

永久几何占用下的多组管路布线:问题、基准与经典基线方法
Quan, Deng
Abstract
Additive manufacturing (AM) enables compact hydraulic components whose internal fluid channels can follow free-form 3D paths rather than conventionally drilled holes. A representative case is multi-group channel layout in rotary direct-drive servo valves: once a channel is placed it permanently occupies volume, so later channels must clear earlier geometry--unlike classical multi-agent pathfinding (MAPF), where agents free space after moving. We study this setting as multi-group pipe routing under permanent geometric occupancy. Our contributions are a problem formalization with geometric dual-witness conflicts, a constructive 3D benchmark (two corridor generators x two obstacle painters, controlled difficulty, feasibility witnesses), and baseline results for CBS, PBS, and priority planning adapted to this coupling. We evaluate by success rate under a fixed time budget. PlaneSlice hard instances clearly rank CBS above PBS and PP, while easier cells validate the pipeline. The goal is a reproducible problem definition, suite, and classical baselines--not a new optimal MAPF algorithm.
Chinese Translation
增材制造(AM)使液压元件得以高度紧凑,其内部流道可沿自由形式的三维路径布置,而非采用传统的钻孔方式。一个典型实例是旋转直驱伺服阀中的多组流道布局:一旦某条流道被放置,它将永久占据空间,因此后续流道必须避开先前的几何结构——这与经典的多智能体路径规划(MAPF)不同,后者的智能体在移动后会释放所占空间。我们将该场景研究为永久几何占用下的多组管路布线问题。我们的贡献包括:一个包含几何对偶见证冲突的问题形式化定义、一个构造性三维基准测试集(由两种走廊生成器与两种障碍绘制器组合而成,具有可控难度和可行性见证),以及针对这种耦合关系进行适配后的 CBS、PBS 和优先级规划(priority planning)等基线方法的结果。我们在固定时间预算下以成功率进行评估。结果表明,在 PlaneSlice 困难实例上,CBS 明显优于 PBS 和 PP,而在较简单的实例上则验证了整个流程的有效性。本文的目标是提供一个可复现的问题定义、测试套件及经典基线方法,而非提出一种新的最优 MAPF 算法。
cs.RO / 6 / 2608.28718

RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

RoboPhys-3D:基于三维重建的具身世界模型综合评测
Wang, Tianyi, Chen, Jiazhou, Xu, Yiming, Li, Xiangyu, Zeng, Tianyi, Chou, Chih-Hsien, Lu, Ning, Peng, Liang, Jiao, Junfeng, Claudel, Christian
Abstract
Video world models increasingly serve as data engines, action planners, and simulators for embodied AI, but conventional embodied world model (EWM) benchmarks lack a unified 3D-grounded protocol for establishing whether generated rollouts preserve the underlying 3D scene state or translate into executable actions. We introduce RoboPhys-3D, a 3D-grounded EWM benchmark built on RoboTwin 2.0, covering 50 manipulation tasks across four regimes, with 5,000 episodes and 25,000 multi-view ground-truth videos. A defining feature of RoboPhys-3D is that generated and ground-truth videos are processed through the same 3D reconstruction pipeline, enabling reconstruction-induced error to be distinguished from generation-induced error. The RoboPhys-3D benchmark organizes 50 complementary metrics into 18 sub-dimensions across four levels: pixel-level fidelity, 3D geometry consistency, state-level understanding, and task-level completeness. We further introduce Average Full Score, a hierarchical score averaging all 50 metrics for comprehensive evaluation, and RoboPhyscore, a compact task-aligned score averaging the metrics most strongly correlated with task success. Among the four representative video world models, Cosmos 3 achieves the highest RoboPhyscore (0.6330, 92.7% of ground truth), while state- and execution-grounded metrics reveal substantial failures that perceptual and vision-language model-based judgments fail to capture. RoboPhyscore further exhibits strong agreement with human evaluation (Pearson r = 0.9761 and Spearman \r{ho} = 0.8962), demonstrating the importance of grounded, execution-aware evaluation for EWM capability.
Chinese Translation
视频世界模型日益成为具身人工智能的数据引擎、动作规划器和模拟器,但传统的具身世界模型(EWM)基准缺乏统一的、以三维为基准的评测协议,无法判断生成的轨迹是否保持了底层的三维场景状态,或能否转化为可执行的动作。我们提出 RoboPhys-3D,一个构建于 RoboTwin 2.0 之上的三维基准评测体系,涵盖四大类共 50 个操作任务,包含 5,000 个回合和 25,000 个多视角真值视频。RoboPhys-3D 的一个关键特点是:生成的视频与真值视频经过同一套三维重建流程处理,从而能够将重建本身引入的误差与生成模型引入的误差区分开来。该基准将 50 个互补的指标组织为 18 个子维度,分为四个层次:像素级保真度、三维几何一致性、状态级理解和任务级完整性。我们进一步提出 Average Full Score——对全部 50 个指标进行分层平均以实现综合评测的综合得分——以及 RoboPhyscore——一种紧凑的、面向任务的对齐得分,通过对与任务成功相关性最强的指标进行平均得到。在四个代表性的视频世界模型中,Cosmos 3 取得了最高的 RoboPhyscore(0.6330,达到真值水平的 92.7%),而基于状态与执行的指标揭示了大量基于感知和视觉语言模型评判所无法捕捉的失败案例。RoboPhyscore 还与人类评估表现出高度一致性(Pearson r = 0.9761,Spearman ρ = 0.8962),证明了基于实证、具备执行感知能力的评测对于具身世界模型能力评估的重要性。
cs.RO / 7 / 2608.28723

Vehicle Drift Emergence: Continuous Evolution from Grip Driving to the Handling Limit via Boundary Exploration Learning Model Predictive Control

车辆漂移的涌现:基于边界探索学习模型预测控制从抓地行驶到操纵极限的连续演化
Zhao, Sheng, Nguyen, Binh-Minh, Lu, Hangyu, Wu, Xiaodong
Abstract
Automated drift controllers commonly track a prescribed drift equilibrium, sideslip reference, or trajectory. These formulations establish how to execute drift, whereas the continuous transition from grip driving to drift near the handling limit remains unresolved. This paper defines drift emergence in a repetitive lap time minimization task, where neither the controller objective nor the reward contains an explicit drift reference. A boundary exploration learning model predictive controller (BE-LMPC) constructs an empirical safe set and a locally shifted terminal cost from completed laps. By iteratively improving spatial speed allocation under a fixed global speed bound, the controller progressively explores larger sideslip and yaw rate envelopes while preserving recoverability. As lap performance improves, sustained sideslip and pronounced yaw motion emerge while the rear axle approaches saturation. Analysis shows that, when external conditions vary smoothly, the transition from tire adhesion to sliding does not itself cause abrupt changes in tire force or vehicle state. The combined-slip Fiala model satisfies this continuity condition at the transition. At a tire road friction coefficient of 0.6, lap time decreases from 49.95 s on Lap~3 to 25.50 s on Lap~12, with drift first emerging on Lap~11. Lap~12 reaches 16.5$^\circ$ sideslip and 0.894 rear axle utilization. In contrast, no drift is detected for friction coefficients from 0.8 to 1.2; at 1.2, a similar peak speed is achieved with only 0.483 rear axle utilization. These results characterize drift as a conditional continuation of limit handling that emerges when increasing performance demand approaches the available tire capacity, rather than as a separately prescribed motion mode.
Chinese Translation
自动化漂移控制器通常跟踪预设的漂移平衡点、侧偏角参考或轨迹。这些方法解决了如何执行漂移的问题,而操纵极限附近从抓地行驶到漂移的连续过渡问题仍未解决。本文在重复圈时间最小化任务中定义了漂移涌现,其中控制器目标函数和奖励中均不包含显式的漂移参考。边界探索学习模型预测控制器(BE-LMPC)基于已完成的圈次构建经验安全集和局部偏移的终端代价。通过在固定全局速度界限下迭代改进空间速度分配,控制器在保持可恢复性的同时,逐步探索更大的侧偏角和横摆角速度包络。随着圈时性能的提升,在后轴趋于饱和的同时,持续的侧偏角和显著的横摆运动逐渐涌现。分析表明,当外部条件平滑变化时,从轮胎附着到滑移的转变本身并不会引起轮胎力或车辆状态的突变。联合滑移Fiala模型在该转变处满足这一连续性条件。在轮胎-路面摩擦系数为0.6时,圈时从第3圈的49.95秒降低至第12圈的25.50秒,漂移首次出现在第11圈。第12圈达到16.5°侧偏角和0.894的后轴利用率。相比之下,摩擦系数在0.8至1.2范围内未检测到漂移;在1.2时,以仅0.483的后轴利用率即可达到相近的峰值速度。这些结果表明,漂移是极限操纵的条件性延续——当不断提高的性能需求逼近可用轮胎能力时涌现——而非一种单独预设的运动模式。
cs.RO / 8 / 2608.28733

Generation of High-Level Concepts in 3D Scene Graphs via Autoregressive Diffusion

基于自回归扩散的3D场景图高层概念生成
Millan-Romera, Jose Andres, Cognolato, Samuel, Voos, Holger, Sanchez-Lopez, Jose Luis, Serafini, Luciano
Abstract
Indoor 3D Scene Graphs (3DSGs) represent environments as multi-layer hierarchies that connect observed geometric primitives (e.g., planes) to higher-level metric-semantic concepts (e.g., rooms, floors, buildings), enabling incremental spatial reasoning for robotic perception and SLAM. However, classical high-level concept generation approaches rely on hand-crafted rules for specific concept classes, while learning-based methods require separate models for graph structure and spatial node features (e.g., centroids), which limits scalability to novel classes and more complex hierarchies. We propose a unified autoregressive diffusion-based graph generative model that jointly learns structure and features, constructing complete 3DSGs bottom-up from observed vertical planes across arbitrary hierarchy depths. Our method consistently surpasses all learning-based and random baselines across 3DSG datasets spanning synthetic scenes, real architectural floor plans, and robotic sensor data, with varying layout complexity and hierarchy depth, and surpasses a one-shot model with oracle access to the target graph size on the largest hierarchy and on real single-floor data. Finally, we propose an adaptation of the Fused Gromov--Wasserstein distance for principled graph-level evaluation of generated 3DSGs against ground truth.
Chinese Translation
室内3D场景图(3D Scene Graphs, 3DSGs)将环境表示为多层层次结构,将观测到的几何基元(如平面)与更高层级的度量-语义概念(如房间、楼层、建筑)相连接,从而支持面向机器人感知与SLAM的增量式空间推理。然而,经典的高层概念生成方法依赖于针对特定概念类别的手工规则,而基于学习的方法则需要为图结构和空间节点特征(如质心)分别训练独立的模型,这限制了其向新类别和更复杂层次结构扩展的能力。我们提出一种统一的自回归扩散图生成模型,能够联合学习图结构与节点特征,从观测到的垂直平面出发自底向上地构建任意层次深度的完整3DSG。在涵盖合成场景、真实建筑平面图和机器人传感器数据、具有不同布局复杂度与层次深度的3DSG数据集上,我们的方法始终优于所有基于学习的基线和随机基线,并在最大层次结构及真实单层数据上超越了可获取目标图大小先验信息(oracle)的单次生成模型。最后,我们提出了一种融合Gromov-Wasserstein距离(Fused Gromov–Wasserstein distance)的改进方法,用于将生成的3DSG与真实值进行有原则的图级评估。
cs.RO / 9 / 2608.28778

Adversarial Calibration Attack on Autonomous Vehicles

针对自动驾驶车辆的对抗性标定攻击
Liu, Liangkai, Zhang, Qingzhao, Shin, Kang G.
Abstract
Autonomous vehicles (AVs) rely on accurate camera-LiDAR calibration for multimodal sensor fusion. In practice, calibration can drift due to vibration, temperature variation, or minor sensor displacement, motivating online calibration algorithms that detect and correct misalignment at runtime while allowing the vehicle to continue operating without a factory visit. Existing AV attacks largely assume correct calibration. We instead identify online sensor calibration as a new attack plane. A corrupted calibration update can persist across subsequent fusion operations, causing system-wide errors that propagate from perception to planning and control. We present Adversarial Calibration Attack (ACA), the first physical attack against camera-LiDAR online calibration. Using a single adversarial poster, ACA first spoofs the miscalibration detector to trigger the calibration process and then steers the calibration estimator toward an incorrect transformation. A unified optimization jointly designs the poster's geometry and texture for both objectives. We evaluate ACA across benchmark datasets, simulation, and physical experiments. On benchmark datasets such as KITTI and nuScenes, ACA induces up to 33.9 degrees mean rotational calibration error, thereby severely degrading object detection. In the CARLA simulator, the attack causes a collision when the corrupted calibration is accepted in vulnerable scenarios crafted by the attacker. On a real Husky robot, a printed adversarial poster successfully reproduces the calibration error. These results demonstrate that online calibration is a practical and safety-critical attack surface for AVs.
Chinese Translation
自动驾驶车辆(AV)依赖精确的相机-激光雷达(LiDAR)标定来实现多模态传感器融合。在实际应用中,标定可能因振动、温度变化或传感器的微小位移而发生漂移,这促使人们开发在线标定算法,以在运行时检测并纠正失准问题,从而使车辆无需返厂即可继续运行。现有的自动驾驶车辆攻击研究大多假设标定是正确的。我们则将在线传感器标定识别为一个新的攻击面。被破坏的标定更新可能在后续的融合操作中持续存在,导致从感知到规划与控制的系统级误差传播。我们提出了对抗性标定攻击(Adversarial Calibration Attack, ACA),这是首个针对相机-激光雷达在线标定的物理攻击。ACA 仅使用一张对抗性海报,首先欺骗失准检测器以触发挥标定过程,然后将标定估计器引导至错误的变换。通过统一优化,联合设计海报的几何形状与纹理以同时实现这两个目标。我们在基准数据集、仿真和物理实验中评估了 ACA。在 KITTI 和 nuScenes 等基准数据集上,ACA 可造成高达 33.9 度的平均旋转标定误差,从而严重降低目标检测性能。在 CARLA 仿真器中,当被破坏的标定在攻击者精心设计的脆弱场景中被接受时,该攻击会导致碰撞。在真实的 Husky 机器人上,一张打印的对抗性海报成功复现了标定误差。这些结果表明,在线标定是自动驾驶车辆一个切实存在且关乎安全的攻击面。
cs.RO / 10 / 2608.28826

Dual Park-Ravani Interpolation of Rigid Motions: Acceleration-Field Continuity and Holonomic Hermite Repair

刚体运动的双Park-Ravani插值:加速度场连续性与完整Hermite修复
Condurache, Daniel
Abstract
The Park-Ravani construction generates a twice continuously differentiable, frame-invariant spline on SO(3) by exponentiating cubic canonical-coordinate polynomials. We show that the construction transfers, without changing form, to the group of orthogonal dual tensors, a representation of rigid displacements. The transferred recurrence is stated compactly through the dual extension of the right Jacobian of the exponential map and its first Fr\'echet derivative. This yields interpolation of prescribed rigid poses and continuity of the body dual twist and its first derivative. Using the higher-order rigid-body kinematics of dual spatial twists, we then prove that the resulting curve has a continuous physical acceleration field, not merely a continuous quantity obtained by formally differentiating the dual part of a twist. We also distinguish algebraic dual transfer from temporal differential prolongation: their simultaneous first-order use takes place in a hyper-dual algebra, and interpolation of arbitrary prolonged nodal data need not be holonomic. A noncommuting three-pose example verifies the recurrence, all knot continuity statements, and dimensional covariance under a change from meters to millimeters. We define and analyze the first-order holonomy defect of a generic hyper-dual interpolant, exhibit an exact counterexample, and remove the defect by cubic or quintic Hermite interpolation in dual logarithmic coordinates.
Chinese Translation
Park-Ravani构造通过对三次规范坐标多项式取指数,在SO(3)上生成具有二阶连续可微性且标架不变的样条。我们证明该构造可以不改变形式地迁移到正交对偶张量群——刚体位移的一种表示上。迁移后的递推关系通过指数映射的右雅可比矩阵的对偶扩展及其第一Fréchet导数得到紧凑表述。由此可实现给定刚体位姿的插值,并保证体对偶扭转(dual twist)及其一阶导数的连续性。利用对偶空间扭转的高阶刚体运动学,我们进一步证明所得曲线具有连续的物理加速度场,而不仅仅是通过对扭转的对偶部分进行形式求导所得到的连续量。我们还区分了代数对偶迁移与时间微分延拓:二者的同时一阶使用发生在超对偶代数(hyper-dual algebra)中,且对任意延拓节点数据的插值未必是完整的(holonomic)。一个不可交换的三位姿算例验证了该递推关系、所有节点处的连续性结论,以及在单位从米变为毫米时的量纲协变性。我们定义并分析了一般超对偶插值曲线的一阶和乐缺陷(holonomy defect),给出了一个精确的反例,并通过在双对数坐标中的三次或五次Hermite插值消除了该缺陷。
cs.RO / 11 / 2608.28967

Brain-Language-Action (BLA) Models: Language-Conditioned EEG for Robotics Control

脑-语言-动作(BLA)模型:面向机器人控制的语言条件化脑电图
Plashchinsky, Alexandr
Abstract
Electroencephalography (EEG)-based robotic control is commonly formulated as a direct classification problem, in which electrical neural signals are mapped to a fixed set of discrete actions. However, the limited separability and high noise of EEG signals make it difficult to scale this approach to fine-grained robotic control spaces. We introduce Brain-Language-Action (BLA) models, a framework in which language conditions the interpretation of neural representations for robotic action generation. In a BLA, a small set of reliably distinguishable brain states can be dynamically associated with different actions through a language-defined control mapping, allowing a small number of neural classes to apply to a larger global action space. We develop a proof-of-concept BLA for drone control using motor-imagery EEG from the BCI Competition IV 2a dataset. The system is trained in two stages. First, we evaluate multiple candidate EEG encoder architectures using subject-specific four-class motor-imagery classification, converting 250Hz, 3.5-second, 22-channel EEG samples into five 128-dimensional brain-token embeddings. Second, these embeddings are projected into the embedding space of a pretrained large language model (LLM) and jointly fine-tuned with language instructions to autoregressively generate structured three-token drone actions. Across 840 possible language-defined mappings between four neural states and seven flight action combinations, the resulting BLA achieves 90% per-token accuracy during evaluation. These results provide an initial demonstration that language conditioning can expand the effective control range of EEG-based robotic interfaces without requiring a corresponding increase in the number of directly distinguishable neural states.
Chinese Translation
基于脑电图(EEG)的机器人控制通常被构建为一个直接分类问题,即将神经电信号映射到一组固定的离散动作。然而,EEG信号的可分性有限且噪声较高,使得该方法难以扩展到细粒度的机器人控制空间。我们提出了脑-语言-动作(Brain-Language-Action, BLA)模型,这是一个通过语言对神经表征的解释进行条件化以生成机器人动作的框架。在BLA中,少量可可靠区分的脑状态可以通过语言定义的控制映射与不同动作动态关联,从而使少量神经类别能够适用于更大的全局动作空间。我们基于BCI Competition IV 2a数据集中的运动想象EEG,开发了一个用于无人机控制的BLA概念验证系统。该系统分两个阶段训练:首先,我们评估了多种候选EEG编码器架构,采用被试特定的四类运动想象分类任务,将250Hz、3.5秒、22通道的EEG样本转换为五个128维的脑token嵌入(brain-token embeddings);其次,将这些嵌入投影到预训练大语言模型(LLM)的嵌入空间中,并与语言指令共同微调,以自回归方式生成结构化的三token无人机动作。在四个神经状态与七种飞行动作组合之间的840种可能的语言定义映射中,所得BLA在评估期间达到了90%的单token准确率。这些结果初步证明,语言条件化可以在不要求直接可区分神经状态数量相应增加的情况下,扩展基于EEG的机器人接口的有效控制范围。
cs.RO / 12 / 2608.28995

Hydra: A Navigation World Action Model with Discrete Latent Planning and Continuous Flow-Matching Execution

Hydra:一种具有离散潜在规划与连续流匹配执行的导航世界动作模型
Nazeri, Mohammad, Card, Alexandyr, Huber, Samira, Pokhrel, Anuj, Wang, Yujun, Hammele, Ruben, Song, Daeun, Pirk, Sören, Xiao, Xuesu
Abstract
World models let robots imagine possible futures, but exploiting this capability for real-time control is bottlenecked by a representation misalignment: the generative model and the planner operate on decoupled manifolds, so the planner has no shared structure to search over and must instead decode every candidate back into high-dimensional pixel space to evaluate it. This decoding step is a major obstacle to real-time control on physical hardware. In this paper, we present Hydra, a discrete World Action Model that closes this gap by moving the planner, both the sampler and the evaluator, inside the model. Hydra establishes a unified latent manifold over visual states, physical poses, and control actions, then compresses this manifold through modality-specific Vector-Quantized bottlenecks into discrete vocabularies of kinodynamic intents and visual states. Because candidates are now drawn directly from this shared manifold, sampling is informed by the model's own understanding of the observation rather than proposed blind, and evaluation happens natively within the discrete space: candidates are ranked by a Kinematic-Perceptual Cost, without ever decoding to pixels. We term this Discrete Latent Planning (DLP). Because planning over discrete intents alone cannot supply the smooth, continuous commands physical actuation requires, Hydra pairs DLP with conditional Flow Matching, which maps each selected intent to a continuous trajectory for execution. Evaluated on two physical robotic platforms, Hydra outperforms state-of-the-art world models in goal-directed planning, while matching or exceeding the closed-loop execution capabilities of leading reactive foundation policies.
Chinese Translation
世界模型使机器人能够想象可能的未来,但要将这种能力用于实时控制,却受制于一种表征错位问题:生成模型与规划器运行在相互解耦的流形上,规划器因此缺乏可供搜索的共享结构,只能将每个候选方案解码回高维像素空间来加以评估。这一解码步骤是物理硬件上实现实时控制的主要障碍。本文提出Hydra,一种离散世界动作模型(World Action Model),它通过将规划器(包括采样器与评估器)移入模型内部来弥合这一鸿沟。Hydra在视觉状态、物理位姿与控制动作之上建立统一的潜在流形,并通过模态特定的向量量化(Vector-Quantized)瓶颈将该流形压缩为运动动力学意图与视觉状态的离散词表。由于候选方案直接从这一共享流形中抽取,采样得以依据模型自身对观测的理解进行,而非盲目提出;评估也在离散空间内原生完成:候选方案依据运动-感知代价(Kinematic-Perceptual Cost)进行排序,完全无需解码至像素空间。我们将这一机制称为离散潜在规划(Discrete Latent Planning, DLP)。由于仅对离散意图进行规划无法提供物理执行所需的平滑连续指令,Hydra将DLP与条件流匹配(conditional Flow Matching)相结合,将每个被选中的意图映射为用于执行的连续轨迹。在两个物理机器人平台上的评估表明,Hydra在目标导向规划方面优于最先进的世界模型,同时在闭环执行能力上达到甚至超越了领先的响应式基础策略。
cs.RO / 13 / 2608.29000

Coding What Matters: A Semantic-Aware Memory Interface for Energy-Efficient Perception in Autonomous Vehicles

编码关键内容:一种面向自动驾驶汽车高效感知的语义感知内存接口
Que, Haohua, Yao, Handong
Abstract
Autonomous vehicles stream high-resolution surround-camera frames into memory before perception runs. This sensor-to-memory path consumes energy when cells store ones and adjacent bytes toggle on the data bus, so its cost follows bit-1 density and switching activity rather than pixel semantics. We present MotiMem-Omega, a semantic-aware memory-interface coder that lowers this cost while preserving perception predictions. Its semantic importance field protects traffic participants, especially vulnerable road users, while assigning lower fidelity to sky and empty background. Cross-dataset bit-sensitivity sweeps determine class weights, with a safety floor for pedestrians, cyclists, and motorcyclists. Each image block then selects a precision tier by minimizing a joint energy-distortion cost. When ego pose is available, a motion-compensated prior carries protected regions between frames. We estimate interface-energy reduction from the two measured proxies using a coefficient-swept memory-energy model. Across 29 detectors on 12 driving datasets, 5 occupancy models, and 5 segmentation networks, MotiMem-Omega retains about 90% of detection mean average precision, 91% of vulnerable-road-user recall, over 98% of occupancy accuracy, and the strongest segmentation retention among energy-reducing methods. It reduces front-camera bit-1 density by 52%, corresponding to a modeled memory-interface energy reduction near 36%, with a lower end of 27% under the literature coefficient sweep. It also gives higher retention than the baseline and energy-matched truncation at the same or lower bit-1 density, whereas image codecs preserve accuracy without reducing memory-interface energy.
Chinese Translation
自动驾驶汽车在运行感知算法之前,会将高分辨率环视摄像头帧流写入内存。这一从传感器到内存的数据路径在存储单元为1以及相邻字节在数据总线上翻转时消耗能量,因此其开销取决于比特1的密度和翻转活动,而非像素语义。我们提出了MotiMem-Omega,一种语义感知的内存接口编码器,在保持感知预测结果的同时降低该开销。其语义重要性字段对交通参与者(尤其是易受伤害的道路使用者)加以保护,而对天空和空旷背景赋予较低的保真度。通过跨数据集的比特敏感度扫描确定类别权重,并为行人、骑行者和摩托车骑行者设置安全底线。随后,每个图像块通过最小化联合的能量-失真代价来选择精度层级。当自车位姿可用时,运动补偿先验可在帧间传递受保护区域。我们利用系数扫描的内存能量模型,基于两个实测代理指标估计接口能量降低幅度。在12个驾驶数据集、5个占据栅格模型和5个分割网络上,跨29个检测器的实验表明,MotiMem-Omega保留了约90%的检测平均精度均值、91%的易受伤害道路使用者召回率、超过98%的占据栅格准确率,并在所有降低能量的方法中具有最强的分割性能保持率。它将前摄摄像头的比特1密度降低了52%,对应模型估计的内存接口能量降低约36%,在文献系数扫描下最低为27%。在相同或更低的比特1密度下,其性能保持率优于基线和能量匹配的截断方法,而图像编解码器虽然能保持精度,却无法降低内存接口能量。
cs.RO / 14 / 2608.29005

A Degradation-Tolerance Benchmark for Camera-Only End-to-End Driving

面向纯摄像头端到端驾驶的退化容忍度基准测试
Que, Haohua, Yao, Handong
Abstract
Camera-only end-to-end (E2E) driving models are nearing deployment, where the camera stream is degraded by blur, noise, low light, weather, frame loss, and memory faults. How much a policy tolerates before its driving breaks is unclear. Corruption-robustness benchmarks target detection or bird's-eye-view perception, not the planning output that drives the car. We present DriveDegrade, a benchmark for image-degradation tolerance in camera-only E2E driving. Sixteen corruption families at five severities are injected on the fly inside the image loader, one operator reaching fifteen policies, and we evaluate open-loop planning on nuScenes and NAVSIM plus a CARLA closed-loop anchor. First, mild degradation barely affects planning, and the families that break it have a clear threshold at mid severity. Second, fragility is corruption-dependent: blur, JPEG, and raindrop damage planning most, while weather and bit error are tolerated far into the range. Third, a flat curve is ambiguous, so we separate corruptions that degrade the image from those that remove it. A planner that reads its camera must lose accuracy when information is deleted, whatever it does under quality loss. On these two axes the planners separate sharply, quantifying the ego-status shortcut without mistaking indifference for robustness. A released vision-language-action planner is flat on both axes, and blinding all six of its cameras costs it only 11.5 percent.
Chinese Translation
纯摄像头端到端(E2E)驾驶模型正接近实际部署阶段,而现实中的摄像头视频流会受到模糊、噪声、低光照、恶劣天气、丢帧以及内存故障等因素的影响而退化。目前尚不清楚驾驶策略在多大程度的退化下仍能保持正常驾驶。现有的退化鲁棒性基准主要针对目标检测或鸟瞰图(BEV)感知,而非真正驱动车辆的规划输出。我们提出了DriveDegrade,一个用于评估纯摄像头端到端驾驶中图像退化容忍度的基准测试。该基准在图像加载器内动态注入十六类退化、每种五级强度,仅需一个算子即可作用于十五种策略,并在nuScenes和NAVSIM上评估开环规划,同时以CARLA闭环实验作为锚点。首先,轻度退化对规划几乎无影响,而真正破坏规划性能的退化类型在中度强度处存在明显阈值。其次,脆弱性依赖于具体退化类型:模糊、JPEG压缩和雨滴对规划的损害最大,而天气影响和位错误在很大范围内均能被容忍。第三,平坦的性能曲线存在歧义,因此我们将降低图像质量的退化与移除图像信息的退化区分开来。任何依赖摄像头的规划器,在信息被完全删除时必然会损失精度,无论其在质量损失下表现如何。基于这两个维度,各规划器表现出显著分化,从而量化了自车状态捷径(ego-status shortcut),避免将“无反应”误认为“鲁棒性”。一个已发布的视觉-语言-动作(VLA)规划器在两个维度上均表现平坦——使其全部六个摄像头失效仅导致11.5%的性能下降。
cs.RO / 15 / 2608.29023

Teaching Robot Policies to Humans Using Erroneous Examples

利用错误示例向人类教授机器人策略
Narayan, Rithika, Jayaraman, Suresh Kumaar, Admoni, Henny
Abstract
Human-robot collaboration describes the process of humans and autonomous agents working together to accomplish common goals. This process is facilitated best when robot policies, or behaviors in different situations, are made transparent to humans. Demonstration-based explanations have been a focus of human-robot collaboration research, and the field has frequently drawn upon literature from education to improve how humans are taught robot policies. However, no single teaching method has been proven effective across domains, difficulties, learners, and other variables; the question of how humans can most effectively be taught robot policies remains open. In traditional classrooms, learners are shown erroneous examples, in which they reflect on and correct incorrect responses to understand common pitfalls when learning a concept. We propose using erroneous examples to teach robot policies, extending an existing policy teaching framework. We conduct a user study in which participants view incorrect demonstrations of robot behavior and correct the actions to align with the actual policy. Our findings suggest that viewing these incorrect demonstrations and verbalizing one's reasoning in predicting a robot's actions improves retention of the policy over time, in agreement with the effect of erroneous examples in classrooms. We also categorize participants into distinct learning styles and establish that participants using inverse reinforcement learning-like reasoning perform best on policy prediction tasks. With this work, we aim to advance the methods by which robots educate humans on their policies.
Chinese Translation
人机协作描述了人类与自主智能体协同工作以实现共同目标的过程。当机器人策略(即机器人在不同情境下的行为)对人类保持透明时,这一过程能够得到最好的促进。基于演示的解释一直是人机协作研究的焦点,该领域经常借鉴教育学文献来改进机器人策略的教学方式。然而,目前尚无任何单一教学方法被证明在不同领域、难度、学习者及其他变量下均有效;如何最有效地向人类教授机器人策略这一问题仍然悬而未决。在传统课堂中,学习者会被展示错误示例(erroneous examples),通过反思和纠正错误的回答来理解学习某一概念时的常见陷阱。我们提出使用错误示例来教授机器人策略,并对现有的策略教学框架进行了扩展。我们开展了一项用户研究,参与者观看机器人行为的不正确演示,并对动作进行纠正以使其与实际策略一致。研究结果表明,观看这些不正确的演示并在预测机器人动作时口头表达自己的推理过程,能够提高对策略的长期记忆保持效果,这与错误示例在课堂中的效果相一致。我们还将参与者划分为不同的学习风格,并证实采用类似逆强化学习(inverse reinforcement learning)推理方式的参与者在策略预测任务中表现最佳。通过这项工作,我们旨在推动机器人向人类传授其策略的方法的发展。
cs.RO / 16 / 2608.29078

DREAM: Deployment-Time Demonstration Generation via Real-to-Sim for Scalable Policy Adaptation

DREAM:基于虚实转换的部署时示范生成,实现可扩展的策略适配
Sato, Makoto, Matsushima, Tatsuya, Matsuo, Yutaka, Iwasawa, Yusuke
Abstract
Vision-language-action (VLA) models have made strong progress in language-conditioned robot manipulation, but improving their performance in a new workspace still often requires action-labeled data from that environment. Collecting such data by human teleoperation is costly, especially when each workspace, object arrangement, or task may require new demonstrations. We present DREAM, a framework that generates fine-tuning data for a pretrained VLA from a captured workspace and a language instruction, without requiring a task-specific human demonstration. DREAM reconstructs the workspace, automatically translates the instruction into symbolic task goals and success criteria using a large language model, and uses task-and-motion planning to generate feasible robot trajectories. The planned trajectories are augmented across randomized object configurations, verified by the generated success criteria, and rendered into image-action examples for VLA fine-tuning. Through real-robot experiments on language-conditioned manipulation tasks, we study whether DREAM can serve as a scalable data-collection system for the deployment workspace by examining whether fine-tuning on its automatically generated data improves success over direct deployment and how its data-collection cost compares with human teleoperation when adapting a VLA to a new workspace.
Chinese Translation
视觉-语言-动作(VLA)模型在语言条件化的机器人操作任务中取得了显著进展,但要提升其在新的工作空间中的性能,通常仍需要来自该环境的带动作标签的数据。通过人工遥操作采集此类数据成本高昂,尤其是当每个工作空间、物体布局或任务都可能需要新的示范时。我们提出了DREAM,一个无需任务特定人工示范、仅依据捕获的工作空间和语言指令即可为预训练VLA生成微调数据的框架。DREAM重建工作空间,利用大语言模型自动将指令转化为符号化的任务目标与成功判据,并通过任务与运动规划(TAMP)生成可行的机器人轨迹。所规划的轨迹在随机化的物体配置下进行增广,由生成的成功判据进行验证,并渲染为用于VLA微调的图像-动作样本。通过在语言条件化操作任务上的真机实验,我们研究了DREAM能否作为部署工作空间的可扩展数据采集系统:具体考察在其自动生成的数据上进行微调能否优于直接部署,以及在将VLA适配到新工作空间时,其数据采集成本与人工遥操作相比如何。
cs.RO / 17 / 2608.29080

GHOST in the Robots: Real-Time Exocentric Dual-Robot VR Teleoperation from Onboard Cameras

机器人中的GHOST:基于机载相机的实时外中心双机器人VR遥操作
Wei, Yichen, Zaghloul, Faisal, Aryal, Soujanya C, Agrawal, Aanya K., Li, Chengfan, Liu, Jason Xinyu, Tompkin, James, Tellex, Stefanie
Abstract
Teleoperating multiple robots simultaneously enables additional views and coordinated control. Yet, it poses fundamental challenges: the system must present sensor data cohesively and allow operators to manage multiple robot bases, arms, and cameras while maintaining low latency. Current multi-robot teleoperation systems require multiple operators, rely on autonomy, or restrict operators to high-level commands. We present GHOST: an open-source VR teleoperation system that enables single operator control of two mobile manipulators via direct lowlevel commands using only onboard sensing. GHOST creates an exocentric 3D workspace by aligning real-time point clouds from the robots' RGB-D cameras, where scene coverage is improved through learning-based completion to aid operator spatial awareness. For control, the operator uses a mode-switching architecture to command either robot individually or both robots simultaneously. Experiments with 15 novice participants demonstrate 1.6-4x the success rate of an off-the-shelf tablet interface. For experts across nine challenging dual-robot tasks, our system enabled completion of two tasks that were infeasible with the tablet, and was 1.47x faster on average than the tablet. Website and code: https://h2r.github.io/GHOST/.
Chinese Translation
同时遥操作多个机器人可以提供额外的视角并实现协同控制。然而,这带来了根本性的挑战:系统必须以连贯的方式呈现传感器数据,并允许操作者在保持低延迟的同时管理多个机器人底盘、机械臂和相机。现有多机器人遥操作系统要么需要多名操作者,要么依赖自主性,要么将操作者限制在高层次指令层面。我们提出了GHOST:一个开源的VR遥操作系统,仅利用机载传感,通过直接底层指令实现单个操作者对两台移动操作机器人的控制。GHOST通过配准来自机器人RGB-D相机的实时点云构建外中心(exocentric)3D工作空间,并借助基于学习的补全方法改善场景覆盖,以辅助操作者的空间感知。在控制方面,操作者采用模式切换架构,可单独控制任一机器人或同时控制两台机器人。包含15名新手参与者的实验表明,该系统的成功率为现成平板界面的1.6至4倍。在九项具有挑战性的双机器人任务中,专家使用本系统完成了平板界面无法实现的两个任务,且平均速度是平板界面的1.47倍。网站与代码:https://h2r.github.io/GHOST/。
cs.RO / 18 / 2608.29100

Agri-Sim: Agricultural Simulation Platform for Embodied Intelligence Evaluation in Greenhouse Robotics

Agri-Sim:面向温室机器人具身智能评估的农业仿真平台
Shi, Shuhan, Xue, Zhenfeng, Mei, Minghao, Zheng, Chao, Li, Nan, Miao, Zhonghua
Abstract
Agricultural-robot development requires simulation environments that can jointly support realistic scene construction, virtual sensing, autonomous navigation, motion planning, and manipulation-task execution. This paper presents Agri-Sim, a Unity and ROS2-based simulation platform for the closed-loop development and functional evaluation of agricultural robots. The platform contains a configurable tomato-greenhouse environment, a mobile dual-arm harvesting robot, virtual RGB-D, LiDAR, IMU, and joint sensors, and a bidirectional communication interface between Unity and ROS2. Unity is responsible for scene rendering, rigid-body dynamics, collision detection, virtual sensing, and task-state execution, whereas ROS2 and MoveIt 2 provide localization, navigation, collision-aware motion planning, inverse kinematics, and trajectory generation. Autonomous greenhouse navigation and dual-arm tomato harvesting were used to evaluate the complete simulation workflow. The experiments covered virtual sensor publication, ROS2-based navigation, collision-aware motion planning, mobile-base control, tomato acquisition, inter-arm handover, and box placement. The results demonstrate that Agri-Sim supports closed-loop integration and repeatable functional evaluation of navigation and manipulation workflows in a controlled virtual greenhouse, providing a practical foundation for subsequent algorithm development and Sim-to-Real studies.
Chinese Translation
农业机器人的研发需要能够同时支持真实感场景构建、虚拟感知、自主导航、运动规划以及操作任务执行的仿真环境。本文提出了Agri-Sim,一个基于Unity和ROS2的仿真平台,用于农业机器人的闭环开发与功能评估。该平台包含一个可配置的番茄温室环境、一台移动双臂采摘机器人、虚拟RGB-D、LiDAR、IMU及关节传感器,以及Unity与ROS2之间的双向通信接口。Unity负责场景渲染、刚体动力学、碰撞检测、虚拟感知和任务状态执行,而ROS2与MoveIt 2则提供定位、导航、考虑碰撞的运动规划、逆运动学和轨迹生成。研究采用温室自主导航与双臂番茄采摘任务对完整的仿真工作流程进行评估。实验涵盖了虚拟传感器数据发布、基于ROS2的导航、考虑碰撞的运动规划、移动底盘控制、番茄获取、双臂间交接以及装箱放置等环节。结果表明,Agri-Sim能够在受控虚拟温室中支持导航与操作工作流程的闭环集成和可重复的功能评估,为后续的算法开发和Sim-to-Real(仿真到现实)研究提供了实用基础。
cs.RO / 19 / 2608.29108

From Multi-Modal Paths to Executable Trajectories: A Trajectory Planning Framework for 4WIS Robots

从多模态路径到可执行轨迹:一种面向四轮独立转向(4WIS)机器人的轨迹规划框架
Bao, Runjiao, Zhang, Lin, Xu, Yongkang, Wang, Shoukun
Abstract
Four-wheel independent steering (4WIS) mobile robots support multiple motion modes, offering high maneuverability in narrow and complex environments. However, existing planning methods often fail to fully exploit these capabilities, leading to suboptimal trajectory quality. To address this limitation, this paper proposes a multi-modal global trajectory planning framework that couples mode-augmented front-end search with mode-consistent segment-wise trajectory optimization. In the front-end stage, Hybrid A* is extended to a four-dimensional state space incorporating motion modes, while mode-switching-aware cost and heuristic functions embed mode decisions into the global search process. Multi-modal Reeds-Shepp curves and an intelligent terminal connection strategy are further designed to improve search efficiency. In the back-end stage, a segment-wise trajectory optimization framework based on an improved iterative safe corridor scheme is developed to convert discrete multi-modal paths into smooth, kinematically feasible trajectories with stationary mode transitions. Experimental results show that the proposed method achieves the best overall performance in safety, arrival time, terminal accuracy and computation time. Real-world experiments on a physical 4WIS robot further validate the practical effectiveness and executability of the generated trajectories, providing a flexible and high-performance solution for multi-modal mobile robot trajectory planning.
Chinese Translation
四轮独立转向(4WIS)移动机器人支持多种运动模式,在狭窄复杂环境中具有高机动性。然而,现有规划方法往往未能充分利用这些能力,导致轨迹质量欠佳。针对这一局限,本文提出了一种多模态全局轨迹规划框架,该框架将模式增强的前端搜索与模式一致的分段轨迹优化相结合。在前端阶段,将Hybrid A*扩展至包含运动模式的四维状态空间,并通过感知模式切换的代价函数与启发式函数将模式决策嵌入全局搜索过程;进一步设计了多模态Reeds-Shepp曲线和智能终端连接策略以提高搜索效率。在后端阶段,提出了一种基于改进迭代安全走廊方案的分段轨迹优化框架,将离散的多模态路径转化为平滑、运动学可行且模式切换平稳的轨迹。实验结果表明,所提方法在安全性、到达时间、终端精度和计算时间方面均取得了最佳综合性能。在实体4WIS机器人上的真实实验进一步验证了所生成轨迹的实用有效性与可执行性,为多模态移动机器人轨迹规划提供了一种灵活且高性能的解决方案。
cs.RO / 20 / 2608.29114

CGFM-Nav: Cognitive Graph-Field Memory for Semantic-Guided Lifelong Multimodal Embodied Navigation

CGFM-Nav:面向语义引导的终身多模态具身导航的认知图-场记忆模型
Xiao, Yuxiang, Chen, Xibei, Zhou, Xin, Chen, Jie, Zhang, Yifeng, Sartoretti, Guillaume
Abstract
Vision-and-Language Navigation (VLN) requires agents to reason over accumulated observations while continuously exploring unseen regions. However, existing environment representations often struggle to jointly support explicit semantic memory and continuous exploration guidance. To address this challenge, we propose Cognitive Graph-Field Memory (CGFM), a persistent multimodal scene representation that couples explicit relational memory with continuous spatial intuition. CGFM organizes objects, spatial relations, and visual observations into a multimodal scene graph, enabling target retrieval and long-horizon reasoning across navigation tasks. When no reliable target match is identified, graph-based evidence is projected into a goal-conditioned semantic-frontier field to guide exploration toward semantically promising frontiers and regions. Building upon CGFM, we introduce CGFM-Nav, a foundation-model-based framework for lifelong multimodal navigation that integrates task-relevant subgraph selection, VLM reasoning, and verification feedback into a closed decision loop. Preliminary experiments on GOAT-Bench show that, under the same Qwen3-VL-8B backbone, CGFM-Nav improves the overall success rate from 53.2% to 63.0% and SPL from 30.0% to 39.6%, demonstrating the effectiveness of combining explicit semantic memory with semantic-guided exploration.
Chinese Translation
视觉语言导航(Vision-and-Language Navigation, VLN)要求智能体在持续探索未知区域的同时,对积累的观测进行推理。然而,现有的环境表示往往难以同时支持显式的语义记忆和连续的探索引导。为应对这一挑战,我们提出了认知图-场记忆(Cognitive Graph-Field Memory, CGFM),这是一种持久化的多模态场景表示,将显式的关系记忆与连续的空间直觉相耦合。CGFM 将物体、空间关系和视觉观测组织成多模态场景图,支持目标检索以及跨导航任务的长时程推理。当未找到可靠的目标匹配时,基于图的证据被投影到一个目标条件化的语义前沿场(semantic-frontier field)中,以引导探索向语义上更有前景的前沿和区域推进。在 CGFM 的基础上,我们提出了 CGFM-Nav,这是一个基于基础模型的终身多模态导航框架,它将任务相关的子图选择、视觉语言模型(VLM)推理和验证反馈整合到一个闭环决策流程中。在 GOAT-Bench 上的初步实验表明,在相同的 Qwen3-VL-8B 骨干网络下,CGFM-Nav 将整体成功率从 53.2% 提升至 63.0%,SPL 从 30.0% 提升至 39.6%,验证了将显式语义记忆与语义引导探索相结合的有效性。
cs.RO / 21 / 2608.29146

Systematic Lightweight Method for Robotics Based on Strain Energy Distribution Optimization

基于应变能分布优化的机器人系统化轻量化方法
Li, Jingchen
Abstract
Service robots work with people and are highly expected to be lightweight for safety, agility and energy conservation. As a complex mechanical system, a robot consists of a large number of components and has various working configurations and load environments. An effective method for achieving system-level optimal robot design is a crucial requirement, but it poses significant challenges. In this study, we introduce a novel approach to optimize the distribution of strain energy, which can significantly improve the effectiveness of systematic optimization in a complex system. First, we present and demonstrate that the strain energy per unit mass should be uniformly distributed in an optimal lightweight mechanical system. Based on this criterion, the system-level problem can be decoupled and the design objective of each part can be assigned based on the strain energy. Then, each part can be optimally designed separately according to its specific circumstances by using different approaches, such as size optimization, topological optimization, and material optimization. In this way, the optimization is at the system level, while the computational complexity is at the part level. Weight reduction and improvements in mechanical properties can be obtained simultaneously. As an example, this method is applied to an arbitrarily designed robotic arm, and its effectiveness is further demonstrated in the cases of lightweight with stiffness improvement, considering multiple materials, multiple working conditions, and vibration performance.
Chinese Translation
服务机器人与人类协同工作,出于安全性、灵活性和节能的考虑,人们对其轻量化寄予厚望。作为一种复杂的机械系统,机器人由大量零部件组成,并具有多样的工作构型和载荷环境。实现系统级最优机器人设计的有效方法是至关重要的需求,但也带来了巨大的挑战。在本研究中,我们提出了一种优化应变能分布的新方法,该方法能够显著提高复杂系统系统级优化的有效性。首先,我们提出并证明了在最优轻量化机械系统中,单位质量的应变能应均匀分布。基于该准则,系统级问题可以被解耦,并且可以基于应变能为每个部件分配设计目标。然后,各部件可根据其具体情况采用不同的方法(如尺寸优化、拓扑优化和材料优化)分别进行优化设计。这样,优化是在系统级进行的,而计算复杂度则处于部件级。减重和机械性能的提升可以同时实现。作为一个实例,该方法被应用于一个任意设计的机械臂,并在考虑多材料、多工况以及振动性能的轻量化且刚度提升的案例中进一步验证了其有效性。
cs.RO / 22 / 2608.29208

AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models

AdaVLA:用于无需训练加速视觉-语言-动作模型的自适应步数流匹配方法
Han, Sunghwan, Han, Youngtae, Yi, Youngmin
Abstract
Vision-Language-Action (VLA) models, built upon Vision-Language Models (VLMs), have significantly enhanced robotic capabilities by leveraging internet-scale knowledge and multimodal reasoning. However, the intensive computational overhead of VLAs constrains on-device deployment, hindering real-time responses to environmental changes. While various acceleration techniques have been proposed, they often rely on fine-tuning or access to training datasets, which are frequently unavailable due to privacy and proprietary concerns. Moreover, although flow-matching-based VLAs have emerged as efficient alternatives to standard diffusion models, current acceleration efforts largely target VLM inference costs, failing to address the iterative ODE solving process inherent in flow matching inference. To address these limitations, we propose AdaVLA, an online, training-free adaptive framework for fast yet accurate flow-matching-based Vision-Language-Action models. We introduce a novel metric derived from the flow matching trajectory curvature to quantify action generation confidence during inference. This metric enables the dynamic reduction of inference steps and the adaptive adjustment of MLP pruning ratios through an efficiently computed importance evaluation, requiring no access to training data. Experimental results on the LIBERO benchmark using a Jetson AGX Orin device demonstrate that our method achieves $1.87\times$ and $2.24\times$ speedups for $\pi_{0.5}$ and X-VLA, respectively, with negligible degradation in success rates. Furthermore, we validate the robustness of our approach on real-world robotic tasks using SmolVLA.
Chinese Translation
基于视觉-语言模型(VLM)构建的视觉-语言-动作(VLA)模型通过利用互联网规模的知识和多模态推理能力,显著提升了机器人能力。然而,VLA巨大的计算开销限制了其在设备端的部署,阻碍了对环境变化的实时响应。尽管已有多种加速技术被提出,但它们通常依赖于微调或访问训练数据集,而由于隐私和专有性问题,这些数据往往无法获取。此外,尽管基于流匹配(flow matching)的VLA已成为标准扩散模型的高效替代方案,当前的加速工作主要针对VLM的推理开销,未能解决流匹配推理中固有的迭代常微分方程(ODE)求解过程。为解决这些局限,我们提出了AdaVLA,一个在线的、无需训练的自适应框架,用于实现快速且精确的基于流匹配的视觉-语言-动作模型。我们引入了一种由流匹配轨迹曲率导出的新指标,用于量化推理过程中动作生成的置信度。该指标能够动态减少推理步数,并通过高效计算的重要性评估自适应地调整MLP剪枝比例,且无需访问训练数据。在Jetson AGX Orin设备上基于LIBERO基准的实验结果表明,我们的方法在 $\pi_{0.5}$ 和 X-VLA 上分别实现了 $1.87\times$ 和 $2.24\times$ 的加速,且成功率几乎无下降。此外,我们使用SmolVLA在真实世界机器人任务上验证了该方法的鲁棒性。
cs.RO / 23 / 2608.29242

AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization

AnyWorld:面向跨具身泛化的因子化第一视角世界模型
Chen, Cheng, Bai, Jerry, Wei, Jiacheng, Chen, Boyu, Zheng, Xiaoji, Wu, Fan, Yang, Minghao, Chen, Tianrun, Li, Ruibo, Yue, Xiaoyu, Guo, Xiaoyang, Ge, Yixiao, Lin, Guosheng, Liu, Fayao
Abstract
Collecting contact-rich robot experiences at scale remains a major bottleneck for generalizable manipulation. Beyond data quantity, robot learning also requires diverse experiences across embodiments, viewpoints, and scenes. Human egocentric videos provide abundant physical interactions, but each video captures only a narrow slice of experience under a single body, camera trajectory, and environment. We propose AnyWorld, a cross-embodiment world modeling framework that expands a single human interaction into diverse robot-native rollouts without paired human-robot demonstrations. Our model factorizes an interaction into action, camera, and embodiment: action controls capture the motion structure, camera controls specify viewpoint evolution, and the target embodiment context defines the acting body and its interaction geometry. This formulation enables independent recomposition of embodiment, viewpoint, and scene factors, allowing a single model to generate many robot-domain experiences while preserving the underlying dynamics and object interactions. We train the model with large-scale human interaction pretraining followed by mixed-embodiment fine-tuning. Experiments show that our model supports controllable recomposition across embodiments, viewpoints, and scenes, and we further demonstrate that the generated data can improve manipulation performance on the RoboCasa GR1 tabletop benchmark and a real IRON humanoid robot. Beyond aggregate gains, we test whether unpaired human experience can be recomposed into robot-native video-action pairs that target a policy gap. Controlled IRON interventions correct a spurious completion prior and establish language-grounded spatial target selection; an action-only counterfactual intervention fails to learn the latter reliably, showing that both action calibration and visual recomposition are necessary.
Chinese Translation
大规模收集富含接触的机器人经验仍然是实现可泛化操作的主要瓶颈。除了数据数量之外,机器人学习还需要跨越不同具身形态、视角和场景的多样化经验。人类第一视角视频提供了丰富的物理交互,但每段视频仅在单一身体、相机轨迹和环境下捕捉到狭窄的经验片段。我们提出AnyWorld,一个跨具身世界建模框架,无需配对的人机演示即可将单次人类交互扩展为多样的机器人原生经验序列。我们的模型将交互因子化为动作、相机和具身:动作控制捕捉运动结构,相机控制指定视角演化,而目标具身上下文定义执行动作的身体及其交互几何。这一形式化使得具身、视角和场景因子可以独立重组,让单一模型在保持底层动力学和物体交互的同时生成大量机器人领域的经验。我们通过大规模人类交互预训练和混合具身微调来训练模型。实验表明,我们的模型支持跨具身、跨视角和跨场景的可控重组,并进一步证明生成的数据可以提升RoboCasa GR1桌面基准以及真实IRON人形机器人上的操作性能。除整体性能提升外,我们还测试了非配对的人类经验能否被重组为针对策略缺陷的机器人原生视频-动作对。受控的IRON干预纠正了一个虚假的完成先验,并建立了基于语言的空间目标选择;而仅动作的反事实干预无法可靠地习得后者,表明动作校准和视觉重组都是必要的。
cs.RO / 24 / 2608.29315

SGE: Semantically-Guided Exploration for Unstructured Environments via Image-Space Waypoint Sampling

SGE:基于图像空间路径点采样的语义引导非结构化环境探索方法
Tatsch, Christopher, Gu, Yu
Abstract
This work introduces Semantically-Guided Exploration (SGE), a modular exploration framework for ground vehicles that integrates pixel-level semantic segmentation into sampling-based waypoint selection and receding-horizon route optimization. Unlike conventional geometric exploration methods, SGE evaluates candidate exploration goals directly in the image space using a semantic-aware utility function that accounts for terrain traversability, obstacle proximity, objects of interest, and depth-based exploration reward. Sampled waypoints are projected into 3D and ordered through a real-time Traveling Salesman Problem (TSP) formulation, enabling receding-horizon goal selection. To address real-world navigation uncertainty, the framework introduces mechanisms, including temporary taboo regions to handle navigation failures and a graph-based relocation strategy for efficient backtracking across explored areas. We evaluate SGE in standardized simulation benchmarks against state-of-the-art exploration planners and demonstrate competitive performance in volumetric coverage, while enabling semantic task biasing that cannot be achieved by purely geometric methods. The framework is further validated through real-world experiments using multiple robotic platforms in indoor campus buildings and in limestone and coal mines. Results show consistent performance and adaptability across platforms and domains.
Chinese Translation
本工作提出了语义引导探索(Semantically-Guided Exploration,SGE),这是一种面向地面车辆的模块化探索框架,将像素级语义分割集成到基于采样的路径点选择和滚动时域路径优化中。与传统的几何探索方法不同,SGE直接在图像空间中评估候选探索目标,采用一个语义感知的效用函数,该函数综合考虑了地形可通行性、障碍物邻近程度、感兴趣目标以及基于深度的探索奖励。采样得到的路径点被投影到三维空间,并通过实时的旅行商问题(Traveling Salesman Problem,TSP)求解进行排序,从而实现滚动时域的目标选择。为应对真实世界导航中的不确定性,该框架引入了多种机制,包括用于处理导航失败的临时禁忌区域,以及用于在已探索区域间高效回溯的基于图的重新定位策略。我们在标准化仿真基准中将SGE与最先进的探索规划器进行了对比评估,结果表明其在体积覆盖率方面具有竞争力,同时能够实现纯几何方法无法达成的语义任务偏好引导。该框架还通过在室内校园建筑以及石灰岩煤矿和煤矿中使用多种机器人平台的真实世界实验得到了进一步验证。结果显示,该框架在不同平台和不同领域中均表现出一致的性能与适应性。
cs.RO / 25 / 2608.29347

A Cognitive Architecture for Shared Autonomy in AUV Operations

一种用于自主水下航行器操作共享自主性的认知架构
Ellis, Niamh, Tran, Thi, Carlucho, Ignacio, Petillot, Yvan R.
Abstract
Operators remain essential to Remotely Operated Vehicle (ROV) operation, yet often suffer from low situational awareness and high workload, both of which negatively affect safety. This paper presents a cognitive architecture consisting of an ontology and multiple Large Language Models (LLMs) to assist the operator at all stages of the mission. Each LLM is grounded with domain-specific information from the ontology and given a simple role to create a system that can support the operator at all stages of an operation. We are aiming to prove that using the two together will allow decisions to be grounded in the relevant domain knowledge, but also benefit from the reasoning capabilities of the LLM. Our framework determines if a mission is possible for a given Unmanned Underwater Vehicle (UUV), performs mission planning, and executes a given mission in simulation. The operator can be involved in planning and execution, ensuring the resulting plan is valid and that the vehicle behaves safely during execution. We compare different LLMs, Llama3, GPT-OSS, and Qwen2.5, to determine which are best suited to the different roles within our framework. We find that GPT-OSS performs best for feasibility assessment, planning, and execution, while Qwen2.5 is best suited to identifying mission types from natural language input.
Chinese Translation
操作人员在水下遥控机器人(ROV)操作中仍然不可或缺,但常常面临态势感知能力低、工作负荷大的问题,这两者都会对安全性产生负面影响。本文提出了一种由本体(ontology)和多个大语言模型(LLM)构成的认知架构,用以在任务的各个阶段为操作人员提供辅助。每个LLM均以本体中的领域特定信息为基础进行接地,并被赋予简单的角色,从而构建一个能够在操作全过程中支持操作人员的系统。我们的目标是证明,将二者结合使用不仅能使决策建立在相关领域知识之上,还能受益于LLM的推理能力。我们的框架能够判断给定无人水下航行器(UUV)是否具备执行某任务的能力,进行任务规划,并在仿真中执行给定任务。操作人员可参与规划与执行过程,以确保所得规划有效,并保证航行器在执行过程中安全运行。我们比较了不同的LLM——Llama3、GPT-OSS和Qwen2.5——以确定哪些模型最适合框架中的不同角色。结果表明,GPT-OSS在可行性评估、规划和执行方面表现最佳,而Qwen2.5最适合从自然语言输入中识别任务类型。
cs.RO / 26 / 2608.29379

Bridging Semantics and Physics with Constrained LLMs for Safe and Trustworthy Robotic Manipulation

基于约束大语言模型的安全可信机器人操作:语义与物理的桥接
Hong, Wenhao, Wei, Lan, Zhang, Dandan
Abstract
A language-guided robot operating in a real kitchen must do more than produce a plan that appears correct. It must also execute that plan safely in cluttered environments under imperfect perception. Large language models (LLM) can decompose instructions into action sequences, yet a language-action gap remains: a plan may appear valid linguistically while being physically infeasible under kinematic and collision constraints. We bridge this gap by formalizing the reasoning-execution boundary as a typed contract. From RGB-D observations, the system grounds perceived objects in an explicit, collision-aware scene model and constrains language-level decisions through schema-validated tool calls defined by the Model Context Protocol (MCP), rejecting malformed commands before they reach the robot. Each validated call is deterministically grounded in a MoveIt Task Constructor pipeline, where candidate motions are evaluated against the reconstructed planning scene in a verify-then-act step. Only trajectories that pass both kinematic and collision checks are sent to the robot. On a physical UFactory 850, the method achieves up to 80% success across ten trials per task on pouring tasks involving liquids, granular media, and discrete solids. It achieves 90% success on a grasp-and-place task using the same planning, protocol, and verification stack. Although a scripted policy slightly outperforms our method on the easiest task, its success rate falls to 10% on the hardest, compared with 60% for our method.
Chinese Translation
在真实厨房中工作的语言引导机器人,不仅要生成看似正确的规划,还必须在感知不完美的条件下,于杂乱环境中安全地执行该规划。大语言模型(LLM)能够将指令分解为动作序列,但仍然存在语言-动作鸿沟:一个规划在语言层面看似有效,却可能在运动学约束和碰撞约束下不可行。我们通过将推理-执行边界形式化为一种类型化契约来弥合这一鸿沟。该系统从RGB-D观测出发,将感知到的物体落地到一个显式的、具备碰撞感知能力的场景模型中,并通过由模型上下文协议(Model Context Protocol,MCP)定义的、经模式校验的工具调用来约束语言层决策,在格式错误的命令到达机器人之前将其拒绝。每个通过校验的调用都被确定性地落地到MoveIt Task Constructor流水线中,在“先验证后执行”的步骤中,针对重建的规划场景对候选运动进行评估。只有同时通过运动学与碰撞检查的轨迹才会被发送给机器人。在物理UFactory 850平台上,该方法在涉及液体、颗粒介质和离散固体的倾倒任务中,每项任务十次试验的成功率最高可达80%。在抓取-放置任务中,使用相同的规划、协议与验证栈,成功率达到90%。尽管一个脚本化策略在最简单的任务上略微优于我们的方法,但其在最难任务上的成功率降至10%,而我们的方法为60%。
cs.RO / 27 / 2608.29396

Toward Trustworthy Robot-Assisted Sliding Palpation for Shallow Vessel Localisation with a Calibrated Digital Twin

基于校准数字孪生的可信机器人辅助滑动触诊实现浅层血管定位
Blaszyk, Piotr, Fan, Wen, Deng, Kaizhong, Elson, Daniel, Zhang, Dandan
Abstract
Reliable localisation of shallow subsurface vessels is important for safe robot-assisted venous access and vessel-aware manipulation, but collecting diverse tactile data on physical hardware is costly, time-consuming, and can degrade soft vision-based tactile sensors. We present a robot-assisted sliding-palpation framework in which a calibrated digital twin generates labelled tactile sequences, reducing reliance on real-world data. The twin models sensor-vessel contact, is calibrated against real palpation trajectories using Bayesian-optimisation-based domain adaptation, and is randomised over sliding direction and contact conditions. A spatio-temporal graph neural network trained on simulated marker trajectories performs per-node vessel classification and produces a human-verifiable top-view localisation map through 2D-to-3D-to-2D geometric projection. We evaluate three datasets: Sim, Silicone, and Meat, the latter a raw-meat phantom with vessel models at nominal depths of 0 to 30 mm, using four train-to-test configurations: Sim to Sim, Sim to Silicone, Sim to Meat, and Meat to Silicone. The calibrated twin achieves a simulated-to-real marker-alignment mean absolute error of 0.50 mm at deepest contact across four canonical interactions. After reprojection onto a 1 mm top-view grid, predicted vessel pixels lie on average 1.05 to 5.49 mm from the nearest true vessel pixel across the four models, with 1.05 to 1.31 mm for all except Sim to Meat. The larger error for Sim to Meat reflects the greater domain shift and current limit of simulation transfer. These results demonstrate progress toward trustworthy tactile palpation through calibrated simulation, interpretable localisation, and transparent cross-domain evaluation. Code, model weights, and data are publicly available on GitHub and Zenodo.
Chinese Translation
浅层皮下血管的可靠定位对于安全的机器人辅助静脉穿刺和血管感知操作至关重要,但在物理硬件上采集多样化的触觉数据成本高、耗时长,且可能损害软体视觉触觉传感器。我们提出了一种机器人辅助滑动触诊框架,其中校准的数字孪生(digital twin)生成带标签的触觉序列,从而减少对真实世界数据的依赖。该数字孪生对传感器与血管的接触进行建模,采用基于贝叶斯优化的域自适应方法针对真实触诊轨迹进行校准,并在滑动方向和接触条件上进行随机化。基于仿真标记点轨迹训练的时空图神经网络(spatio-temporal graph neural network)执行逐节点血管分类,并通过2D-3D-2D几何投影生成可供人工验证的俯视定位图。我们在三个数据集上进行评估:Sim(仿真)、Silicone(硅胶)和Meat(肉),其中Meat为具有标称深度0至30 mm血管模型的生肉仿真体模,并使用四种训练-测试配置:Sim到Sim、Sim到Silicone、Sim到Meat以及Meat到Silicone。校准后的数字孪生在四种典型交互中最深接触处实现了0.50 mm的仿真到真实标记点对齐平均绝对误差。在重投影到1 mm俯视网格后,预测血管像素与最近真实血管像素的平均距离为1.05至5.49 mm,除Sim到Meat外,其余模型的误差均在1.05至1.31 mm之间。Sim到Meat较大的误差反映了更大的域偏移以及当前仿真迁移的局限性。这些结果表明,通过校准仿真、可解释的定位和透明的跨域评估,可信触觉触诊研究取得了进展。代码、模型权重和数据已在GitHub和Zenodo上公开。
cs.RO / 28 / 2608.29432

SMILE: Smooth Motion for Improved Long-Horizon VLA Execution

SMILE:用于改进长时程VLA执行的平滑运动方法
Park, Jongwoo, Nguyen, E-Ro, Ranasinghe, Kanchana, Mata, Cristina, Li, Xiang, Ryoo, Michael S
Abstract
Vision-Language-Action (VLA) models reduce inference cost by executing multiple actions per call, but longer horizons often degrade accuracy because raw chunks contain jitter and outliers. We introduce SMILE, an architecture-preserving interface that predicts B-spline coefficients and decodes them into smooth action sequences. SMILE changes only the action representation, enabling longer fixed horizons while retaining each baseline's backbone and model scale. We apply SMILE to SmolVLA, Evo1, VPP, and DAWN, improving accuracy and amortized inference efficiency across LIBERO, CALVIN, and real-world experiments. SMILE-Evo1 reaches 98.0% with a 1.1x speedup on LIBERO, while SMILE-VPP reaches an average length of 4.42 with a 1.5x speedup on CALVIN. At a matched execution horizon of 10, SMILE-SmolVLA reduces non-boundary acceleration by 78.6% and velocity sign-change rate by 42.3%. Real-world xArm tests show higher success, fewer drops, and fewer contacts. These results establish smooth coefficient-space generation as a route to accurate, efficient long-horizon VLA execution. Project page: jongwoopark7978.github.io/smilevla
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型通过每次调用执行多个动作来降低推理成本,但较长的时程往往会降低准确性,因为原始动作块中包含抖动和离群值。我们提出了SMILE,一种保持架构不变的接口,它预测B样条系数并将其解码为平滑的动作序列。SMILE仅改变动作表示形式,从而能够在保留每个基线模型的主干网络和模型规模的同时实现更长的固定时程。我们将SMILE应用于SmolVLA、Evo1、VPP和DAWN,在LIBERO、CALVIN和真实世界实验中均提升了准确性和摊销推理效率。SMILE-Evo1在LIBERO上达到98.0%的准确率并实现1.1倍加速,而SMILE-VPP在CALVIN上达到平均长度4.42并实现1.5倍加速。在相同的执行时程(10)下,SMILE-SmolVLA将非边界加速度降低了78.6%,速度符号变化率降低了42.3%。真实世界的xArm测试显示更高的成功率、更少的掉落和更少的接触。这些结果表明,平滑的系数空间生成是实现准确、高效的长时程VLA执行的一条可行路径。项目页面:jongwoopark7978.github.io/smilevla
cs.RO / 29 / 2608.29487

Blind Dexterity: Whole-Body Humanoid Manipulation via Pure Proprioception

盲灵巧操作:基于纯本体感觉的全身人形机器人操作
Bhatt, Aditya, Kaidanov, Oleg, Liu, Puze, Peters, Jan
Abstract
We present blind, whole-body manipulation skills on a Unitree G1 humanoid using only onboard proprioception, without cameras, markers, force-torque, or tactile sensors. Despite this minimal sensing, the trained policies exhibit surprising capability across qualitatively different tasks: push-resilient bipedal walking without IMU feedback, active soccer ball trapping with a foot, seeking and lifting a suitcase by its handle, and mounting a randomly positioned skateboard. We argue that these capabilities arise from a key underappreciated signal: the way the joint encoder readouts evolve under purposeful compliant contact, effectively forming a whole-body tactile channel. By generating contact-rich motions, the trained policies actively probe the environment; as a result, task-relevant object state (e.g., pose) becomes increasingly decodable from short proprioceptive histories. We expose this information using compact task-specific state estimators trained alongside, but fully separately from, the policies; their prediction errors decrease rapidly after informative contact. Our results indicate that joint encoder-based proprioception, combined with compliant actuation (now widely available on commercial robots and low-cost motors) is already a strong, practical substrate for whole-body dexterous manipulation and interactive perception, and therefore a natural foundation on which richer sensing can be layered.
Chinese Translation
我们在Unitree G1人形机器人上展示了仅使用机载本体感觉的盲态全身操作技能,无需摄像头、标记点、力-力矩传感器或触觉传感器。尽管感知手段极为有限,训练所得的策略在性质不同的多种任务中表现出惊人的能力:无需IMU反馈的抗推挤双足行走、用脚主动停球(足球)、寻找并通过手柄提起行李箱,以及踏上随机放置的滑板。我们认为,这些能力源于一个被低估的关键信号:关节编码器读数在有目的的柔顺接触下的演化方式,它实际上构成了一条全身触觉通道。通过生成富含接触的运动,训练所得的策略能够主动探测环境;因此,与任务相关的物体状态(如位姿)在短程本体感觉历史中变得越来越可解码。我们通过紧凑的任务特定状态估计器来提取这些信息,这些估计器与策略同时训练但完全分离;在获得信息性接触后,其预测误差迅速下降。我们的结果表明,基于关节编码器的本体感觉,结合柔顺执行器(目前在商用机器人和低成本电机上已广泛可用),已经能够为全身灵巧操作和交互式感知提供强大而实用的基础,因此也是叠加更丰富传感手段的天然基石。
cs.RO / 30 / 2608.29514

A Sliding Window Filter on the Galilean Group for Consistent Aided Inertial Navigation with Unknown Measurement Delays

伽利略群上考虑未知测量延迟的一致性辅助惯性导航滑窗滤波器
Kelly, Jonathan
Abstract
We study aided inertial navigation when the aiding sensor measurements are subject to an unknown constant delay. The goal is to estimate the delay and navigation state jointly so that delayed measurements correct the trajectory at the appropriate times, yielding a more accurate navigation solution. We formulate the problem on the special Galilean group, which provides a natural state-space structure for aided navigation with uncertainty in both motion and timing. We then examine the observability of joint delay and state estimation and show that, for a single delayed measurement, the model admits an exact symmetry in which a change in the delay can be compensated by a change in the navigation state, leaving the measurement unchanged. Processing measurements individually allows spurious information to `leak' along the corresponding null direction of the measurement Jacobian, producing overconfident and inconsistent estimates. Applying measurements from multiple times together can eliminate this direction when the trajectory is informative enough. Motivated by this result, we develop a sliding window filter that retains a short history of navigation states and applies delayed aiding corrections jointly across the active window. We conduct a series of simulation studies to characterize estimator accuracy and consistency. The simulations demonstrate that an estimator that does not maintain an adequate window can rapidly become highly inconsistent, whereas even a short sliding window markedly improves consistency by providing the temporal support that observability requires.
Chinese Translation
本文研究辅助传感器测量存在未知恒定延迟情况下的辅助惯性导航问题。目标是联合估计延迟与导航状态,使延迟测量能够在适当的时刻校正轨迹,从而获得更精确的导航解。我们在特殊伽利略群(special Galilean group)上构建该问题,该群为运动和时间均存在不确定性的辅助导航提供了自然的状态空间结构。随后,我们分析了延迟与状态联合估计的可观测性,并证明:对于单个延迟测量,该模型存在一种精确的对称性,即延迟的变化可由导航状态的变化补偿,而测量值保持不变。逐个处理测量会使虚假信息沿测量雅可比矩阵相应的零空间方向“泄漏”,从而产生过度自信且不一致的估计。当轨迹具有足够的信息量时,将多个时刻的测量联合处理可以消除该方向。受此结果启发,我们提出了一种滑窗滤波器,它保留导航状态的短期历史,并在活动窗口内联合应用延迟的辅助校正。我们通过一系列仿真研究刻画了估计器的精度与一致性。仿真结果表明,不维持足够窗口长度的估计器会迅速变得高度不一致,而即使是较短的滑动窗口,也能通过提供可观测性所需的时间支撑,显著改善一致性。
cs.RO / 31 / 2608.29516

Task-Relevant Feature-Dynamics Fidelity Enables Zero-Shot Sim-to-Real Transfer for Robotic Ultrasound Scanning

任务相关特征动力学保真度使机器人超声扫描能够实现零样本仿真到真实迁移
Qian, Yizhao, Luo, Jiayuan, Zhu, Wanyi, Zhang, Yameng, Meng, Max Q. -H., Yuan, Yixuan, Liu, Li
Abstract
Robotic ultrasound policies operating directly on B-mode images require extensive interaction data, whereas real-robot data collection is costly and safety-constrained. Simulation provides a scalable alternative, but zero-shot transfer depends not only on single-frame realism but also on whether simulated observations reproduce task-relevant feature changes induced by probe motion. We term this cross-domain consistency task-relevant feature-dynamics fidelity (TR-FDF). Under local regularity assumptions, our contraction analysis shows that greater sensitivity of TR-FDF mismatch to probe motion reduces the effective closed-loop contraction margin, whereas motion-independent errors primarily enlarge the residual error bound. Guided by this analysis, we develop a TR-FDF-oriented ultrasound simulator that combines a shared structural intermediate domain, trajectory-level fixed noise, and few-step conditional flow generation. In phantom experiments, a policy trained exclusively in simulation succeeded in 390 of 400 zero-shot deployments across four target planes. The simulator achieved an FID of 29.66 and generated observations at 67.1 Hz. Controlled interventions, ablations, and baseline comparisons showed that TR-FDF sensitivity complements single-frame realism in predicting zero-shot transfer performance.
Chinese Translation
直接基于B模式图像运行的机器人超声策略需要大量的交互数据,而真实机器人数据采集成本高昂且受安全性约束。仿真提供了一种可扩展的替代方案,但零样本迁移不仅取决于单帧图像的真实感,还取决于仿真观测能否再现由探头运动引起的任务相关特征变化。我们将这种跨域一致性称为任务相关特征动力学保真度(Task-Relevant Feature-Dynamics Fidelity, TR-FDF)。在局部正则性假设下,我们的收缩分析表明,TR-FDF失配对探头运动的敏感度越高,有效闭环收缩裕度越小,而与运动无关的误差主要会扩大残余误差界。基于这一分析,我们开发了一个面向TR-FDF的超声仿真器,它结合了共享结构中间域、轨迹级固定噪声和少步条件流生成。在体模实验中,仅通过仿真训练的策略在四个目标平面上的400次零样本部署中成功390次。该仿真器实现了29.66的FID,并以67.1 Hz的频率生成观测。受控干预、消融实验和基线对比表明,TR-FDF敏感度与单帧真实感在预测零样本迁移性能方面互为补充。
cs.RO / 32 / 2608.29537

AGM: Achievement-Grounded Memory for Closed-Loop Agents with Frozen VLA Policies

AGM:面向冻结VLA策略的闭环智能体的成就锚定记忆
Gao, Hongbo, Ni, Zeyu, Wen, Xin, Xu, Siyu, Li, Ruifeng
Abstract
Frozen vision-language-action (VLA) policies offer broad manipulation skills but execute open-loop action chunks without tracking task progress, so the agent cannot reliably decide whether to continue, retry, or terminate. External memory is a natural remedy, yet it can be harmful when attempted actions are treated as completed progress, turning local execution errors into persistent task-state errors. We propose Achievement-Grounded Memory (AGM), a lightweight closed-loop framework for frozen VLA policies that represents a task as a subgoal sequence with a progress pointer and advances this memory only after the current subgoal is verified by physical evidence. Proprioceptive interaction cues decide when to verify, while coherent point tracking and language-conditioned cross-view comparison, sourced from frozen foundation models through a single 2.43M-parameter verification head, decide what was achieved. AGM thereby converts open-loop execution into a closed loop of execution, verification, and progress, keeping the policy frozen without test-time large-model inference. On the RoboMME Counting benchmark, AGM reaches on PickXTimes and on BinFill, surpassing the strongest memory-augmented baseline by points on average, and the framework yields equally decisive gains on a physical robot. Reliable embodied memory thus depends more on disciplined state updates than on memory capacity.
Chinese Translation
冻结的视觉-语言-动作(VLA)策略具备广泛的操作技能,但其以开环方式执行动作块,无法追踪任务进度,因此智能体无法可靠地决定是继续执行、重试还是终止。外部记忆是一种自然的补救手段,然而当把尝试过的动作视为已完成的进度时,外部记忆可能产生危害,将局部执行错误转化为持久的任务状态错误。我们提出成就锚定记忆(Achievement-Grounded Memory, AGM),这是一个面向冻结VLA策略的轻量级闭环框架,它将任务表示为带有进度指针的子目标序列,并且只有在通过物理证据验证当前子目标完成后才推进该记忆。本体感知交互线索决定何时进行验证,而连贯的点跟踪和基于语言条件的跨视角比较——通过一个仅2.43M参数的验证头从冻结的基础模型获取——决定完成了什么。由此,AGM将开环执行转变为执行、验证和进度推进的闭环,在不进行测试时大模型推理的情况下保持策略冻结。在RoboMME Counting基准上,AGM在PickXTimes任务上达到(具体数值),在BinFill任务上达到(具体数值),平均超越最强的记忆增强基线(具体数值)分,且该框架在物理机器人上同样带来了显著的性能提升。因此,可靠的具身记忆更多依赖于规范化的状态更新,而非记忆容量。
cs.RO / 33 / 2608.29547

Module Number Adaptive Visual Shape Control for Serial Modular Soft Robots

面向串联模块化软体机器人的模块数量自适应视觉形状控制
Akamine, Kyohei, Horii, Takato, Sakaue, Yusuke, Ishizuka, Hiroki
Abstract
Image based shape control provides a simple means of controlling the whole body configuration of soft robots. However, existing data driven approaches are typically developed for fixed robot structures and require new control data when the number of modules changes. This paper presents a module number adaptive visual shape control method for serial modular soft pneumatic robots. A controller trained only on single module actuation shape data is reused for robots with one to five modules by decomposing whole body camera images into local module patches. A single common module segmenter localizes individual modules across all tested configurations, while the same local controller is applied to every extracted patch. Geometric data augmentation improves transferability to downstream modules, and a lightweight mask reconstruction network reconstructs a synthetically removed actuator mask channel. Experiments on physical robots demonstrate shape control across varying numbers of modules and under environmental changes and payload loading. The results show that single module control learning enables scalable whole body control without configuration specific control data collection.
Chinese Translation
基于图像的形状控制为控制软体机器人的整体构型提供了一种简单的方法。然而,现有的数据驱动方法通常针对固定的机器人结构开发,当模块数量发生变化时需要重新采集控制数据。本文提出了一种面向串联模块化气动软体机器人的模块数量自适应视觉形状控制方法。仅通过单个模块的驱动形状数据训练得到的控制器,通过将整机相机图像分解为局部模块图像块,可被复用于具有一至五个模块的机器人。一个通用的模块分割器在所有测试构型中定位各个模块,而同一个局部控制器被应用于每个提取出的图像块。几何数据增强提升了对下游模块的迁移能力,一个轻量级掩膜重建网络用于重建被合成移除的驱动器掩膜通道。在实体机器人上的实验验证了该方法在不同模块数量、环境变化以及负载条件下的形状控制能力。结果表明,单模块控制学习能够在无需针对特定构型采集控制数据的情况下,实现可扩展的整机控制。
cs.RO / 34 / 2608.29601

$\mathcal{N}_0$-Foundation: Towards the Age of Tactile Intelligence

𝒩₀-Foundation:迈向触觉智能时代
NeoteAI Team, Fudan TEAI Team
Abstract
We present $\mathcal{N}_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale multimodal data, tactile representation learning, and standardized evaluation. First, we engineer the infrastructure for scalable data collection, including a vision-based tactile sensor, a tactile Universal Manipulation Interface (UMI), and a synchronized visuo-tactile data collection system supporting both robot embodiments and UMI-based demonstrations. Leveraging this infrastructure, we construct NeoData, which contains more than 30000 hours of synchronized visual and tactile demonstrations, spanning six embodiments, 450 tasks, and billions of paired RGB and tactile frames collected through a mixture of real-robot teleoperation and UMI-based demonstrations. To facilitate open research, we further release OpenNeoData, a 5000-hour open-source subset of NeoData. The dataset addresses a central limitation of existing manipulation corpora, critical for deformable-object manipulation, precise assembly, delicate force control, and sustained surface interaction. Capitalizing on the large-scale, heterogeneous tactile measurements, we propose NeoForce, a visuo-tactile representation model that learn transferable tactile representations across different sensor designs. To enable systematic evaluation of tactile embodied models built upon our infrastructure, datasets and tactile representations, we further propose a comprehensive benchmark, which combines the real-world NeoReal suite and the simulated NeoSim suite for standardized evaluation. Experiments across both suites show that policies benefit from the physical contact state rather than from the device-specific appearance of the tactile signal. We release the dataset, the representation, and the benchmark, aiming at supporting future work on tactile-enabled embodied manipulation.
Chinese Translation
我们提出了𝒩₀-Foundation,一个面向触觉赋能具身操作(tactile-enabled embodied manipulation)的范式,它集成了触觉感知硬件、大规模多模态数据、触觉表征学习以及标准化评估。首先,我们构建了可扩展数据采集的基础设施,包括一种基于视觉的触觉传感器、一个触觉通用操作接口(Universal Manipulation Interface, UMI),以及一个支持机器人本体和基于UMI的人类示教的同步视觉-触觉数据采集系统。利用该基础设施,我们构建了NeoData数据集,其中包含超过30000小时的同步视觉与触觉示教数据,涵盖六种机器人本体、450个任务,以及通过真实机器人遥操作和基于UMI示教采集的数十亿对RGB与触觉帧。为促进开放研究,我们进一步发布了OpenNeoData,即NeoData的一个5000小时开源子集。该数据集解决了现有操作数据集的一个核心局限,这对可形变物体操作、精密装配、精细力控制以及持续性表面交互至关重要。基于大规模、异构的触觉测量数据,我们提出了NeoForce,一个能够跨不同传感器设计学习可迁移触觉表征的视觉-触觉表征模型。为对基于我们的基础设施、数据集和触觉表征构建的触觉具身模型进行系统性评估,我们进一步提出了一个综合基准,该基准结合了真实世界的NeoReal测试集和仿真的NeoSim测试集,用于标准化评估。在两个测试集上的实验表明,策略受益于物理接触状态本身,而非触觉信号中设备特定的外观特征。我们发布了数据集、表征模型和基准,旨在支持未来触觉赋能具身操作的相关研究。
cs.RO / 35 / 2608.29720

VeloBins: Learning Velocity and Its Uncertainty via Bins and Error-Conditioned Gaussian Labels for Aerial Inertial Odometry

VeloBins:基于分箱与误差条件高斯标签学习速度及其不确定性,用于空中惯性里程计
Azhari, Maulana Bisyir, Lee, Seungwook, Han, Donghun, Park, Sung Jun, Shim, David Hyunchul
Abstract
Inertial odometry (IO) is critical for aerial robots, where aggressive maneuvers and poor lighting degrade visual sensors. Recent learning-based IO methods improve traditional integration-based approaches by learning motion priors from IMU and platform-specific sensors, then fusing the predictions within an extended Kalman filter. However, learning velocity through regression is difficult, while jointly estimating uncertainty with a separate decoder and negative log-likelihood (NLL) loss further complicates training and can lead to over-confident estimates. We introduce VeloBins, which reformulates velocity regression as classification over discretized velocity bins. We decode both the velocity from the bin distribution's expectation and the uncertainty from its variance, removing the need for a separate uncertainty decoder. We further supervise the uncertainty explicitly using an error-conditioned Gaussian label centered at the ground-truth velocity, with a standard deviation set to the velocity error. We evaluate VeloBins on four aerial datasets, ranging from free-form aggressive flights and a 27 g nano-quadrotor to drone racing at over 21~m/s. VeloBins achieves the lowest average errors on all four datasets, reducing velocity, relative trajectory, and absolute trajectory errors by 3-27%, 8-40%, and 6-53%, respectively, compared with the strongest baseline. Notably, the proposed supervision achieves the lowest NLL and best filter consistency despite never optimizing an NLL loss. The code will be available upon acceptance.
Chinese Translation
惯性里程计(IO)对空中机器人至关重要,因为剧烈机动和不良光照会降低视觉传感器的性能。近年来基于学习的IO方法通过从IMU和平台特定传感器中学习运动先验,然后将预测结果融合到扩展卡尔曼滤波器中,从而改进了传统的基于积分的方法。然而,通过回归学习速度十分困难,而使用单独的解码器和负对数似然(NLL)损失联合估计不确定性会进一步增加训练难度,并可能导致过度自信的估计。我们提出了VeloBins,将速度回归重新表述为对离散化速度分箱的分类。我们从分箱分布的期望解码速度,并从其方差解码不确定性,从而无需单独的不确定性解码器。我们进一步通过以速度真值为中心、标准差设为速度误差的误差条件高斯标签,对不确定性进行显式监督。我们在四个空中机器人数据集上评估了VeloBins,涵盖自由形式的剧烈飞行、27克的微型四旋翼以及速度超过21米/秒的无人机竞速。VeloBins在全部四个数据集上均取得了最低的平均误差,与最强基线相比,速度误差、相对轨迹误差和绝对轨迹误差分别降低了3-27%、8-40%和6-53%。值得注意的是,所提出的监督方法尽管从未优化NLL损失,却实现了最低的NLL和最佳的滤波器一致性。代码将在论文被接收后公开。
cs.RO / 36 / 2608.29749

DriftingVLA: Native One-Step Vision-Language-Action Generation via Per-Dimension Temporal Drifting

DriftingVLA:基于逐维度时间漂移的原生单步视觉-语言-动作生成
Gao, Yuxuan, Zhang, Shiqi, Shen, Yedong, Duan, Yifan, Yu, Wenhao, Zhang, Xin, Cao, Siyuan, Deng, Jiajun, Zhang, Yanyong
Abstract
Conventional flow-based vision-language-action (VLA) models support expressive continuous action generation but rely on multi-step refinement to produce each action chunk, increasing latency in online robot control. To address this issue, we introduce DriftingVLA, a native one-step VLA that generates a complete action chunk with a single action-expert forward pass. Rather than learning a flow field that requires iterative integration at inference, DriftingVLA uses a distribution-drifting objective to learn a direct noise-to-action-chunk mapping for one-step deployment. Since robot action dimensions carry distinct control semantics and distributional characteristics, we further introduce Per-Dimension Temporal Drifting (PDTD). PDTD treats the complete temporal trajectory of each action dimension as a separate drifting unit, enabling finer-grained modeling and shaping of dimension-specific action distributions. This per-dimension decomposition applies only to the training objective; the shared VLA model still generates the complete action chunk jointly, thereby preserving cross-dimensional dependencies. DriftingVLA achieves 98.32% success on LIBERO, 81.09% on RoboTwin 2.0, and 77.67% across six real-world single- and dual-arm tasks, outperforming the evaluated multi-step flow policy and one-step VLA baselines. Native one-step deployment also delivers a 3.36-fold speedup in action-chunk generation, eliminating iterative refinement without sacrificing control performance.
Chinese Translation
传统的基于流的视觉-语言-动作(VLA)模型支持富有表现力的连续动作生成,但依赖多步细化来产生每个动作块,增加了在线机器人控制的延迟。为解决这一问题,我们提出了 DriftingVLA,一种原生单步 VLA 模型,仅需一次动作专家前向传播即可生成完整的动作块。DriftingVLA 不学习需要在推理时进行迭代积分的流场,而是采用分布漂移目标来学习从噪声到动作块的直接映射,从而实现单步部署。鉴于机器人动作的各个维度承载着不同的控制语义和分布特性,我们进一步引入了逐维度时间漂移(Per-Dimension Temporal Drifting, PDTD)。PDTD 将每个动作维度的完整时间轨迹视为独立的漂移单元,实现了对维度特定动作分布的更细粒度建模与塑造。这种逐维度分解仅作用于训练目标;共享的 VLA 模型仍然联合生成完整的动作块,从而保留了跨维度的依赖关系。DriftingVLA 在 LIBERO 上取得了 98.32% 的成功率,在 RoboTwin 2.0 上达到 81.09%,并在六项真实世界的单臂与双臂任务中平均达到 77.67%,优于所评估的多步流策略和单步 VLA 基线。原生单步部署还在动作块生成上带来 3.36 倍的加速,在不对控制性能造成损失的前提下消除了迭代细化。
cs.RO / 37 / 2608.29767

LARC: Lazy Adaptive Reachability Certification of Robot Manipulator Trajectories

LARC:机器人机械臂轨迹的惰性自适应可达性认证
Feng, Yu, Wu, Hao, Wang, Yuzhe, Zhou, Jianshu
Abstract
Discrete trajectory checks can miss collisions between sampled robot states. Reachability-based certification bounds motion between states, but uniform time partitions waste computation where clearance is large. We present lazy adaptive reachability certification (LARC), which checks a planned trajectory by bisecting only intervals with an inconclusive clearance test. For piecewise-cubic Hermite joint trajectories, the method bounds link occupancy using midpoint capsules inflated by exact componentwise speed maxima. Certified intervals covering the trajectory provide continuous-time external-obstacle clearance, subject to geometric containment, static obstacles, and a prescribed margin. On 160 AgileX PIPER trajectories from 80 start-goal pairs, LARC matched all decisions of the fixed-fine baseline at depth nine. It used 20328 interval evaluations (24.8% of baseline work), with a median paired speedup of 10.28x. A separate MoveIt/FCL audit checked 158051 states and detected collisions in 21 direct-interpolation controls, none of which LARC certified. The method reduced computation under a shared certificate model, but 27 of 139 sampled-clear trajectories remained uncertified. The sampled audit cannot independently prove continuous-time clearance.
Chinese Translation
离散轨迹检查可能会遗漏采样机器人状态之间发生的碰撞。基于可达性的认证方法通过界定状态间的运动范围来检测碰撞,但在安全裕度较大的区域,均匀的时间划分会浪费计算资源。我们提出了惰性自适应可达性认证(Lazy Adaptive Reachability Certification,LARC)方法,该方法通过对安全裕度测试结果不明确的区间进行二分,来检查规划轨迹。对于分段三次埃尔米特(Hermite)关节轨迹,该方法使用由各分量速度精确最大值膨胀的中点胶囊体来界定连杆的占据空间。覆盖整个轨迹的已认证区间可提供连续时间下的外部障碍物间隙验证,其前提条件包括几何包含关系、静态障碍物以及规定的安全裕度。在来自80组起点-终点对的160条AgileX PIPER轨迹上,LARC在深度为九时与固定精细基线的所有判定结果完全一致。它仅使用了20328次区间评估(相当于基线计算量的24.8%),配对加速比的中位数为10.28倍。另一次独立的MoveIt/FCL审计检查了158051个状态,并在21个直接插值对照中检测到碰撞,而这些对照均未被LARC认证。该方法在共享证书模型下降低了计算量,但139条经采样验证安全的轨迹中仍有27条未被认证。基于采样的审计无法独立证明连续时间下的安全间隙。
cs.RO / 38 / 2608.29768

SmoothRL: Online Reinforcement Learning During Asynchronous Execution

SmoothRL:异步执行过程中的在线强化学习
Gao, Guang, Nong, Yuxuan, Huang, Baifu, Wang, Jianan
Abstract
Deploying robot policies in the physical world requires satisfying two fundamental desiderata: reliability and smooth real-time execution. However, deploying state-of-the-art generalist models presents challenges on both fronts. Achieving the precision and robustness required for real-world deployment necessitates sample-efficient online reinforcement learning (RL) to adapt pretrained models. Meanwhile, the increasing scale of robot foundation models has led to higher inference latency. To satisfy real-time constraints under high latency, modern systems adopt asynchronous inference with action chunking, overlapping policy computation with chunk execution to hide latency and enable smooth control. Despite their complementary roles, integrating asynchronous execution with gradient-based online RL remains underexplored. We present SmoothRL, an online RL framework that fine-tunes a pretrained policy within an asynchronous inference loop. SmoothRL follows a value-gradient paradigm, directly updating policy parameters using gradients of the action-value function with respect to policy actions. To enable correct optimization under asynchronous execution, SmoothRL explicitly models the asynchronous inference process during training. Specifically, each generated action chunk is partitioned by frame index into three regions: a committed region, consisting of actions committed by the previous inference cycle; an execution region, containing newly generated actions executed by the robot; and a discarded region, containing actions superseded by the next inference cycle. Gradients are propagated only through the execution region, ensuring policy optimization aligns with the trajectory distribution induced by asynchronous execution. We evaluate SmoothRL on real-world robotic tasks requiring high precision, as well as highly dynamic tasks that necessitate asynchronous execution.
Chinese Translation
在物理世界中部署机器人策略需要满足两个基本要求:可靠性和流畅的实时执行。然而,部署最先进的通用模型在这两个方面都存在挑战。要达到现实世界部署所需的精度和鲁棒性,必须采用样本高效的在线强化学习(RL)来微调预训练模型。与此同时,机器人基础模型规模的不断增大导致了更高的推理延迟。为了在高延迟下满足实时性约束,现代系统采用基于动作分块的异步推理,将策略计算与分块执行重叠进行,从而隐藏延迟并实现平滑控制。尽管二者具有互补作用,但将异步执行与基于梯度的在线强化学习相结合仍然是一个探索不足的领域。我们提出了SmoothRL,这是一个在线强化学习框架,能够在异步推理循环中对预训练策略进行微调。SmoothRL遵循价值-梯度范式,直接利用动作价值函数相对于策略动作的梯度来更新策略参数。为了在异步执行下实现正确的优化,SmoothRL在训练过程中显式地对异步推理过程进行建模。具体而言,每个生成的动作分块按帧索引被划分为三个区域:已提交区域(committed region),包含由上一个推理周期所提交的动作;执行区域(execution region),包含机器人正在执行的新生成的动作;丢弃区域(discarded region),包含被下一个推理周期取代的动作。梯度仅在执行区域内传播,从而确保策略优化与异步执行所产生的轨迹分布保持一致。我们在需要高精度的真实世界机器人任务以及需要异步执行的高度动态任务上对SmoothRL进行了评估。
cs.RO / 39 / 2608.29769

Learning Agile Perceptive Traversal of Sparse 3D Structures for Humanoids

人形机器人稀疏三维结构的敏捷感知穿越学习
Ongan, Efe, Zhang, Chong, Sun, Boyang, Cramariuc, Andrei, Cadena, Cesar, Hutter, Marco
Abstract
Traversing sparse 3D structures requires humanoid robots to perceive thin, overhanging geometry while executing agile, accurate whole-body motions. We study this problem through monkey-bar traversal, where the robot must jump to the structure, traverse it through sparse bar interactions, and land safely. For this task, we present a reinforcement-learning-based perceptive control system that operates directly on observations from a head-mounted solid-state lidar. To extract task-relevant geometry from the sparse returns, the policy consumes the raw lidar scan through an attention-based encoder with recurrent memory. This policy is obtained by a phase-scheduled teacher- student pipeline that combines privileged experts for jumping up, brachiating, and jumping down. For transfer to hardware, we model lidar noise, battery-voltage sag, and actuator thermal limits, and equip the humanoid with passive hook end-effectors for robust bar interaction. On hardware, the resulting policy completes the full jump-up->brachiation->jump-down sequence in 14 of 15 trials across three bar configurations and reaches brachiation speeds up to 0.5 m/s. Beyond brachiation, the same perception backbone supports a separately trained policy that ducks beneath thin overhead obstacles with 2 cm cross-sections.
Chinese Translation
穿越稀疏的三维结构要求人形机器人在执行敏捷、精确的全身运动的同时,能够感知细薄且悬挑的几何结构。我们通过“云梯(monkey-bar)”穿越任务来研究这一问题:机器人必须跳跃抓握到结构上,通过与稀疏横杆的交互完成穿越,并安全落地。针对该任务,我们提出了一种基于强化学习的感知控制系统,该系统直接以头戴式固态激光雷达(lidar)的观测作为输入。为了从稀疏的点云回波中提取与任务相关的几何信息,策略通过一个基于注意力机制并带有循环记忆的编码器来处理原始激光雷达扫描数据。该策略由一个阶段调度的师生(teacher-student)训练流程获得,该流程结合了分别负责跳起、悬挂摆越(brachiation)和下跳的特权专家策略。为了实现向实机的迁移,我们对激光雷达噪声、电池电压跌落和执行器热限制进行了建模,并为人形机器人配备了被动式钩状末端执行器,以实现稳健的横杆交互。在硬件实验中,所得到的策略在三种横杆配置下的15次试验中有14次成功完成了完整的“跳起→悬挂摆越→下跳”序列,悬挂摆越速度最高达0.5米/秒。除悬挂摆越外,相同的感知骨干网络还支持一个单独训练的策略,使机器人能够从横截面仅2厘米的细薄头顶障碍物下方俯身通过。
cs.RO / 40 / 2608.29770

Sampling-based Certified Planning with Graphs of Convex Sets

基于采样的凸集图确定性安全规划
Xie, Peng, Alanwar, Amr
Abstract
Planners on graphs of convex sets return trajectories that are collision-free by construction, provided the convex regions are collision-free. The region generator only promises that property probabilistically, and no planner in the family verifies it. We report the first measurement of what the gap costs. On a scaled 14-DOF bimanual library, $3.2\%$ of interface samples are in collision, and a search-based GCS planner (\gcsstar) turns that volume error into a $62\%$ answer error: $18$ of $29$ pick-and-place queries return trajectories that drive the arms through the shelves, up to $91$\,mm deep, reported as successes. Repairing the library does not work; a ten times stricter acceptance contract, sums-of-squares certified regions, and uniform margins each destroy the connectivity planning needs before they deliver soundness. We instead build a planner that certifies its answers. It samples the overlaps and shared faces of the decomposition, prunes with an admissible informed bound, and verifies the one candidate each search round proposes, continuously, by a chain of clearance certificate balls with no resolution parameter; failures are repaired with local in-region detours, and the convex polish is re-verified. Head-to-head on all $29$ task queries it delivers zero invalid answers against $21$ for the reference, reaches its first certified answer in $0.11$\,s against $1.59$\,s for the reference's unverified one, and reproduces the reference optimum exactly on every query whose reference answer is physically valid.
Chinese Translation
在凸集图(Graphs of Convex Sets, GCS)上运行的规划器所返回的轨迹,在凸区域无碰撞的前提下,可保证构造上无碰撞。然而区域生成器仅以概率方式承诺这一性质,且该类规划器中没有任何一个会对其实际验证。我们首次量化了这一缺口所付出的代价。在一个按比例缩放的14自由度双臂环境库上,3.2%的界面样本存在碰撞,基于搜索的GCS规划器(gcs*)将这一体积误差放大为62%的答案错误率:29个抓取与放置查询中有18个返回了机械臂穿越货架(最深达91毫米)的轨迹,却被报告为成功。修复区域库的方法并不可行:将接受标准收紧十倍、采用平方和(sums-of-squares)认证的区域、以及统一的安全裕度,每一种方法都在实现可靠性之前就破坏了规划所需的连通性。我们转而构建了一个能够对自身答案进行认证的规划器。它对分解的重叠区域与共享面进行采样,利用可采纳的知情界进行剪枝,并对每轮搜索提出的一个候选解进行连续验证,验证方式为一条不含分辨率参数的间隙证书球链;失败情况通过区域内的局部绕行进行修复,并对凸优化打磨后的结果重新验证。在全部29个任务查询上的直接对比中,该规划器实现了零无效答案(参考方法为21个无效答案),首个认证答案的求解时间为0.11秒(参考方法返回其未经验证的答案需1.59秒),并且在参考答案物理有效的所有查询上,精确复现了参考方法的最优解。
cs.RO / 41 / 2608.29772

Self-Aware Active Learning Enables Continual Improvement in Autonomous Driving

自我感知主动学习实现自动驾驶的持续改进
Hu, Dong, Huang, Chao, Lee, Carman K. M., Kanoulas, Dimitrios
Abstract
Learning-based autonomous driving (AD) systems can perform reliably in familiar conditions, yet rare distribution shifts and long-tail events remain a major source of abrupt failure. A central limitation is that most agents learn primarily from passive experience and lack mechanisms to estimate when their competence is insufficient, seek timely assistance, and convert safety-critical encounters into targeted improvement. Here we present self-aware guided exploration (SAGE), an active learning framework for post-training adaptation in AD. SAGE learns a predictive world model that generates two online intrinsic signals: fear, which estimates short-horizon predictive risk and model uncertainty, and curiosity, which measures novelty through prediction error. Curiosity adaptively calibrates the intervention threshold for fear, allowing the agent to regulate risk in a context-dependent manner. When predicted fear exceeds this adaptive threshold, the agent transfers control to an expert or fallback policy and uses the resulting takeover trajectories for focused imitation learning. In parallel, fear is integrated into policy optimization and evaluation as a safety-oriented constraint to reduce performance regressions during adaptation. We evaluate SAGE in simulated route-transfer tasks, Waymo-based logged driving scenarios, CARLA occlusion hazards, and real-world mobile robot navigation tests. Across these settings, SAGE improves robustness in novel and safety-critical scenarios, reduces safety violations, and maintains task performance comparable to strong baseline policies. These results suggest that agents can improve after initial training by estimating the limits of their competence, requesting guidance when needed, and learning selectively from rare high-value events.
Chinese Translation
基于学习的自动驾驶(AD)系统能够在熟悉的条件下可靠运行,但罕见的分布偏移和长尾事件仍然是导致突发性失效的主要原因。一个核心局限在于,大多数智能体主要从被动经验中学习,缺乏估计自身能力何时不足、及时寻求协助以及将安全关键场景转化为针对性改进的机制。本文提出自我感知引导探索(Self-Aware Guided Exploration, SAGE),一个用于自动驾驶后训练适应的主动学习框架。SAGE学习一个预测性世界模型,该模型生成两种在线内在信号:恐惧(fear),用于估计短期预测风险和模型不确定性;以及好奇心(curiosity),通过预测误差衡量新颖性。好奇心自适应地校准恐惧的干预阈值,使智能体能够以情境依赖的方式调节风险。当预测的恐惧超过该自适应阈值时,智能体将控制权转移给专家策略或兜底策略,并利用由此产生的接管轨迹进行针对性的模仿学习。同时,恐惧作为一种面向安全的约束被整合到策略优化与评估中,以减少适应过程中的性能回退。我们在模拟路线迁移任务、基于Waymo的日志驾驶场景、CARLA遮挡危险场景以及真实世界移动机器人导航测试中评估了SAGE。在所有这些设置中,SAGE提升了在 novel 和安全关键场景中的鲁棒性,减少了安全违规,并保持了与强基线策略相当的任务性能。这些结果表明,智能体可以通过估计自身能力的边界、在需要时请求指导,并选择性地从罕见的高价值事件中学习,从而在初始训练之后持续改进。
cs.RO / 42 / 2608.29896

EMERGE-Policy: A Robot Mind Emerges Beyond a Single Policy

EMERGE-Policy:超越单一策略的机器人心智涌现
Fang, Zhirui, Yu, Qingchi, Chen, Ziyang, Li, Longfei, Ma, Haoran, Zhou, Keru, Xu, Xinrun, Va, Samith, Hu, Yuxuan, Song, Peixuan, Du, Qiang, Qian, Bin, Deng, Yongkang, Li, Xin, Wang, Yezhen, Li, Zhe, Luo, Hao, Li, Shuyan, Wang, Ziwei, Deng, Weijian, Li, Xiu
Abstract
A robot's effective ``mind'' need not reside in a single policy. It can emerge when specialized components perceive, reason, predict, act, verify, and remember within a shared orchestration process. EMERGE-Policy turns this perspective into a graph-structured agentic framework that coordinates both capability invocation and information exchange. A Main Agent retains task-level state within an active context window, while role-specific Sub Agents process perception, execution monitoring, verification, and memory consolidation in isolated contexts and return structured, task-relevant evidence. Role-specific contexts control information load by exposing only decision-relevant evidence to the Main Agent, while the functional Skill interface composes heterogeneous backends as Operational, Imagination, and Evaluation Skills. Criterion-grounded verification, textual failure diagnosis, and Branch Stack recovery provide localized correction, with token-aware external memory preserving task-relevant state. Together, their closed-loop interaction realizes the system-level policy captured by the name EMERGE-Policy. Without additional fine-tuning, we achieved outstanding performance on several public benchmark that have had a wide-reaching impact, and conducted a series of real robot experiments. These system-level results suggest that through the division of different functional sub-tasks among multiple agents and their concurrent collaboration, as well as the technical paradigm where the model is regarded as a skill and called within the framework, EMERGE-Policy can extend the robust robot policies beyond isolated runs.
Chinese Translation
机器人有效的“心智”不必驻留于单一策略之中,而是可以在共享的编排流程中,由专门化组件分别完成感知、推理、预测、行动、验证和记忆时涌现而出。EMERGE-Policy 将这一视角转化为一个图结构化的智能体(agentic)框架,同时协调能力调用与信息交换。主智能体(Main Agent)在活跃的上下文窗口中维护任务级状态,而具有特定角色的子智能体(Sub Agents)则在隔离的上下文中处理感知、执行监控、验证与记忆整合,并返回结构化的、与任务相关的证据。各角色的专用上下文通过仅向主智能体暴露与决策相关的证据来控制信息负载,同时功能化的技能(Skill)接口将异构后端组合为操作技能(Operational Skills)、想象技能(Imagination Skills)与评估技能(Evaluation Skills)。基于准则的验证、文本化的失败诊断以及分支栈(Branch Stack)恢复机制提供了局部化的纠错能力,并借助具有令牌感知的外部记忆来保存与任务相关的状态。它们的闭环交互共同实现了该名称所蕴含的系统级策略——EMERGE-Policy。在无需额外微调的情况下,我们在多个具有广泛影响力的公开基准上取得了出色的性能,并开展了一系列真实机器人实验。这些系统级结果表明,通过在多个智能体之间划分不同的功能子任务并使其并发协作,以及将模型视为一种技能并在框架内调用的技术范式,EMERGE-Policy 能够将鲁棒的机器人策略扩展至单次孤立运行之外。
cs.RO / 43 / 2608.29935

System Identification of Admittance Models for Large Real-World Objects

大型现实世界物体的导纳模型系统辨识
Baum, Nathan I., Luttmer, Nathaniel G., Minor, Mark A.
Abstract
Simulation of admittance-type models requires physically consistent dynamic models that are rarely available for off-the-shelf, everyday objects, limiting the fidelity of haptic interfaces that rely on such simulations. This paper presents the first complete workflow for producing physically consistent models of large real-world objects with various constraints and mechanisms, guaranteeing physical consistency of inertia and friction parameters. The workflow separates each object and identifies the handle and body in two stages, requiring no torque sensors at hinges, axles, or other constrained joints. Models are produced for a heavy, closer-actuated door and a wheelbarrow, representing objects of differing constraint types and model complexity. The door is modeled using four-bar linkage kinematics and a fluid dynamics-based lumped parameter model including opening, backcheck, swing, and latch zones. The wheelbarrow is modeled as a rigid body with a spherical wheel and no slip during rolling. Handle estimation RMS errors were below 0.64 N and 0.042 Nm across both objects. Door body estimation had RMS error of 2.19 Nm and wheelbarrow body estimation had RMS error of 6.77 Nm.
Chinese Translation
导纳型模型的仿真需要物理一致的动力学模型,而现成的日常物品通常缺乏此类模型,这限制了依赖此类仿真的触觉接口的保真度。本文提出了首个完整的工作流程,用于为具有各种约束和机构的大型现实世界物体建立物理一致的模型,并保证惯性和摩擦参数的物理一致性。该工作流程将每个物体分离处理,分两个阶段辨识手柄和主体,无需在铰链、轮轴或其他受约束关节处安装扭矩传感器。该流程为重型闭门器驱动的门和独轮手推车建立了模型,代表了不同约束类型和模型复杂度的物体。门采用四连杆运动学模型和基于流体动力学的集总参数模型进行建模,包括开门、回弹缓冲、摆动和锁扣区域。独轮手推车建模为具有球形轮且滚动过程中无滑移的刚体。两种物体的手柄估计均方根(RMS)误差分别低于0.64 N和0.042 Nm。门主体的估计RMS误差为2.19 Nm,独轮手推车主体的估计RMS误差为6.77 Nm。
cs.RO / 44 / 2608.29967

Training-Free Action Correction for VLA Model Failures via Language Feedback

基于语言反馈的免训练VLA模型动作失败纠正方法
Kwon, Owen, Ortega-Kral, Pablo, Bucker, Arthur, Oh, Jean
Abstract
Vision-Language-Action (VLA) models demonstrate strong semantic understanding yet exhibit systematic failures during deployment. The conditions under which these failures occur, and whether they can be corrected without retraining, remain poorly understood. In this paper, we take steps toward addressing this gap. We present CorrectVLA, a framework that translates task-level natural language corrections into additive action magnitude adjustments without modifying policy weights. A human provides a single task-level correction, applied uniformly across all rollouts without per-episode intervention. In simulation, CorrectVLA recovers execution misalignment failures across both in-distribution and OOD tasks. In real-robot experiments on a UFactory xArm7 under environment shift, CorrectVLA restores near-perfect success where the base policy almost entirely breaks down, generalizing across object locations and identities. Through a taxonomy of failure modes on LIBERO-90, we find that execution misalignment failures, where the policy reaches the correct target but miscalibrates action magnitudes, represent the correctable subset, while other failure modes where semantic comprehension itself breaks down are not amenable to this approach. The approach succeeds when policies possess strategic correctness and fails when fundamental comprehension is absent, establishing a practical operational boundary for inference-time correction.
Chinese Translation
视觉-语言-动作(VLA)模型展现出强大的语义理解能力,但在部署过程中会出现系统性的失败。这些失败在何种条件下发生,以及是否能够在不进行重新训练的情况下得到纠正,目前仍缺乏深入研究。本文朝解决这一空白迈出了步伐。我们提出了CorrectVLA,一个将任务级自然语言纠正转化为附加式动作幅度调整的框架,且无需修改策略权重。人类只需提供一次任务级纠正,该纠正会被统一应用于所有 rollout,而无需逐回合干预。在仿真实验中,CorrectVLA能够修复分布内(in-distribution)和分布外(OOD)任务中的执行失准失败。在UFactory xArm7真实机器人、环境变化条件下的实验中,当基线策略几乎完全失效时,CorrectVLA能够恢复接近完美的成功率,并在物体位置和身份变化上具有良好的泛化能力。通过对LIBERO-90上失败模式的分类分析,我们发现执行失准失败——即策略到达了正确的目标但动作幅度校准不当——构成了可纠正的子集,而其他语义理解本身失效的失败模式则无法通过该方法解决。该方法的成功依赖于策略具备战略层面的正确性,而在根本性理解缺失时会失败,从而为推理时纠正确立了一个实用的操作边界。
cs.RO / 45 / 2608.30055

MiBOT: A head-worn robot that modulates cardiovascular responses through human-like soft massage

MiBOT:一种通过类人软按摩调节心血管反应的头戴式机器人
Mylaeus, Alice, Vogt, Stephanie, Demirel, Berken Utku, Gort, Marcel, Meboldt, Mirko, Meier, Manuel, Holz, Christian
Abstract
Massage therapy is helpful for the rehabilitation of various diseases, such as headaches caused by migraines and stress. Existing robotic systems have focused on massage therapy on the torso and limbs, but performing massage motions through suitable actuation on a person's head has been a challenge. In this paper, we present MiBOT, a head-worn massage robot that actuates two soft tactors to produce touch motions mimicking human massage. A key design principle behind MiBOT is its silent actuation, which we achieve through pneumatic artificial muscles in conjunction with a controller loop to respond to contact pressure. We evaluated the effectiveness of MiBOT in a controlled study and assessed subjects' blood pressure and heart rate levels while applying MiBOT. We found that our mechanical system generated positive and conclusive quantitative outcomes that are similar to the human-administered massage, decreasing participants' mean systolic and diastolic blood pressure by 2.8 mmHg and 1.7 mmHg, respectively, as well as calming their heart rate by 8-10% on average.
Chinese Translation
按摩疗法对多种疾病的康复有益,例如偏头痛和压力引起的头痛。现有的机器人系统主要集中于躯干和四肢的按摩治疗,但通过合适的驱动方式在人体头部执行按摩动作一直是一个挑战。本文提出了MiBOT,一种头戴式按摩机器人,它通过驱动两个柔软的触觉执行器(tactor)来产生模拟人类按摩的触摸动作。MiBOT的一个关键设计原则是静音驱动,我们通过气动人工肌肉结合控制器反馈回路来实现对接触压力的响应。我们在一项对照研究中评估了MiBOT的有效性,并测量了使用MiBOT期间受试者的血压和心率水平。结果表明,我们的机械系统产生了积极且明确able的量化效果,其效果与人工按摩相似:使参与者的平均收缩压和舒张压分别降低2.8 mmHg和1.7 mmHg,并使其心率平均降低8-10%。
cs.RO / 46 / 2608.30144

Rethinking Language's Role in Efficient VLA for Autonomous Vehicles: Toward Smarter, Trustworthy Driving

重新思考语言在高效率自动驾驶VLA模型中的角色:迈向更智能、更可信的驾驶
Guo, Tongfei, Su, Lili
Abstract
Vision-Language-Action (VLA) models are reshaping autonomous driving (AD) by unifying perception, reasoning, and control through language, enabling semantic grounding, interpretable decisions, and better long-tail generalization. But language is expensive onboard: latency and memory budgets are tight, and autoregressive decoding is inherently sequential. This work reframes the central question as when and where language should act at inference, since inference cost recurs at every deployed frame while training cost is paid once. We introduce the Language Residue taxonomy to organize methods by their inference-time use of language: train-time-only supervision (L1), latent non-textual reasoning (L2), conditional invocation (L3), and full per-frame generation (L4). We review representative methods and tag each across five deployment axes (latency, parameters, memory, FLOPs, tokens), analyzing them on major open- and closed-loop driving benchmarks (e.g., nuScenes, NAVSIM, Bench2Drive). We further trace how efficient methods from NLP/LLM are adapted in AD, identifying the constraints and motivations driving these adaptations. A continuously updated repository will be available at Github.
Chinese Translation
视觉-语言-动作(Vision-Language-Action, VLA)模型正通过语言统一感知、推理与控制,重塑自动驾驶(AD)领域,实现语义接地、可解释决策以及更好的长尾场景泛化能力。但语言在车载环境中的开销是昂贵的:延迟和内存预算十分紧张,且自回归解码本质上是串行的。本文将核心问题重构为“语言应在推理时的何时、何处发挥作用”,因为推理成本会在部署后的每一帧反复产生,而训练成本只需支付一次。我们提出了语言残差(Language Residue)分类体系,按方法在推理时对语言的使用方式进行组织:仅训练时监督(L1)、潜在的非文本推理(L2)、条件化调用(L3)以及每帧完整生成(L4)。我们回顾了代表性方法,并沿五个部署维度(延迟、参数量、内存、FLOPs、token数)对每项方法进行标注,同时在主要的开环与闭环驾驶基准(如nuScenes、NAVSIM、Bench2Drive)上对其进行分析。我们进一步追溯了NLP/LLM领域的高效方法如何被改造应用于自动驾驶,识别了驱动这些改造的约束与动机。一个持续更新的代码仓库将发布于Github。
cs.RO / 47 / 2608.30220

Contrast-Free Autonomous Navigation of Untethered Endovascular Microrobots Using Single-Plane Fluoroscopy

基于单平面透视的无造影剂自主导航非系留血管内微型机器人
Alabay, Husnu Halid, Le, Tuan-Anh, Wang, Ping, Ceylan, Hakan
Abstract
Reliable three-dimensional (3D) navigation of magnetically actuated untethered microrobots remains a major barrier to clinical translation. X-ray fluoroscopy is the standard real-time imaging modality for endovascular procedures, but single-plane fluoroscopy provides only a two-dimensional (2D) projection, eliminating depth information and complicating autonomous navigation. Recovering this information through biplane imaging or repeated contrast-enhanced angiography increases procedural complexity, radiation exposure, or contrast burden. Here, we introduce VISTA (Virtual Integration for Spatial Tracking and Autonomy), a digital twin framework enabling contrast-free autonomous navigation under single-plane fluoroscopy. VISTA reconstructs vascular anatomy as a 3D digital twin, discretizes the vessel centerline into navigation milestones, and assigns the detected 2D robot position to the nearest projected milestone. Consecutive milestones define the local vessel orientation used to generate magnetic actuation commands, converting single-plane fluoroscopic observations into topology-constrained navigation states without requiring contrast injection during navigation. VISTA is demonstrated across anatomically distinct vascular phantoms under continuous flow and within the inferior vena cava of a live rat in vivo. Compared with conventional fluoroscopic human-in-the-loop control, VISTA reduced navigation time by up to 62%, corrective actuation commands by up to 98%, and radiation exposure by up to 57%. These results establish VISTA as a digital twin-guided framework for contrast-free autonomous navigation of untethered endovascular microrobots using widely available single-plane fluoroscopy.
Chinese Translation
磁驱动的非系留微型机器人的可靠三维(3D)导航仍是其临床转化的主要障碍。X射线透视是血管内手术的标准实时成像方式,但单平面透视仅提供二维(2D)投影,丢失了深度信息,从而使自主导航复杂化。通过双平面成像或反复注射造影剂增强的血管造影来恢复这些信息,会增加手术复杂性、辐射暴露或造影剂负担。本文提出了VISTA(Virtual Integration for Spatial Tracking and Autonomy,用于空间跟踪与自主性的虚拟集成),这是一个在单平面透视下实现无造影剂自主导航的数字孪生框架。VISTA将血管解剖结构重建为3D数字孪生,将血管中心线离散化为导航里程碑,并将检测到的机器人2D位置分配给最近的投影里程碑。连续的里程碑定义了局部血管方向,用于生成磁驱动指令,从而将单平面透视观测转换为拓扑约束的导航状态,而无需在导航过程中注射造影剂。VISTA在持续流动条件下、解剖结构各异的血管仿体中以及活体大鼠的下腔静脉内得到了验证。与传统的透视人机协同控制相比,VISTA将导航时间最多缩短了62%,纠正性驱动指令最多减少了98%,辐射暴露最多降低了57%。这些结果确立了VISTA作为一个数字孪生引导的框架,可利用广泛可用的单平面透视实现非系留血管内微型机器人的无造影剂自主导航。
cs.RO / 48 / 2608.30222

Optimized Modular Design and Development of a Tilt-Rotor Bicopter Drone

倾转旋翼双旋翼无人机的优化模块化设计与开发
Verma, Saideep, Tatapudi, Nimisha, Arjun, Akshay, Sukka, Danwada Sharanya, Sarkar, Abhishek, Mukherjee, Joyjit
Abstract
Hybrid systems like tilt-rotor bicopter drones combine the beneficial characteristics of both fixed-wing and rotary-wing technology, enabling long endurance and VTOL capability. However, such drones also require an optimum design to ensure both static and dynamic stability. The modular design of a traditional bicopter is developed in this paper based on extensive analysis and in-depth structural and aerodynamic simulations. The structural analysis has been performed to ensure that the aircraft's structure withstands the stresses encountered during different flight modes. Controlling the relative positions of the Center of Gravity (CG) and Neutral Point (NP) is an essential aspect of the design, ensuring stability during hover and positive stability during forward flight. The thrust and power analyses have been conducted to assess the flight performance and endurance. After analysis, the drone has been developed, and flight tests with a basic flight controller were conducted to validate the performance metrics obtained in the simulation.
Chinese Translation
倾转旋翼双旋翼无人机等混合动力系统兼具固定翼与旋翼技术的优点,可实现长航时和垂直起降(VTOL)能力。然而,此类无人机也需要进行优化设计,以确保静态和动态稳定性。本文在大量分析以及深入的结构和气动仿真的基础上,开发了传统双旋翼无人机的模块化设计。结构分析旨在确保机体结构能够承受不同飞行模式下所遇到的应力。控制重心(CG)与中性点(NP)的相对位置是设计中的一个关键环节,可确保悬停时的稳定性以及前飞时的正稳定性。通过推力和功率分析评估了飞行性能与续航能力。完成分析后,研制了该无人机,并使用基础飞行控制器进行了飞行测试,以验证仿真所获得的性能指标。
cs.RO / 49 / 2608.30237

Motus2: A Self-Evolving General World Model for Dexterous Manipulation

Motus2:一个面向灵巧操作的自我进化通用世界模型
Bi, Hongzhe, Zhou, Zihao, Tang, Yihang, Pang, Jingrui, Huang, Shuhe, Liu, Haitian, Wang, Runqing, Huang, Shuai, Wang, Yichen, Cheng, Yiming, Zhao, Ruowen, Li, Zhenghua, Tan, Hengkai, Liu, Xiaolong, Wan, Jinhui, Liu, Jiabao, Zhao, Min, Bao, Fan, Zhu, Jun
Abstract
General embodied agents should perceive, predict, act, evaluate, and improve within a unified system. World models have shown great promise in building such agents, yet existing models typically append an action output head to a world simulator, without coupling them into a closed decision-and-learning loop for policy improvement. We present Motus2, a self-evolving general world model for dexterous manipulation. Motus2 advances world modeling through model scaling and data scaling. For model scaling, a single model with shared weights exposes three control interfaces: a policy (world-action model), a simulator (action-conditioned world model), and an evaluator (value model). The policy proposes candidate action chunks, the simulator predicts their visual consequences, and the evaluator assesses the predicted outcomes. Their coupling forms a closed decision-and-learning loop for policy improvement. This formulation uses curated expert demonstrations for action learning, while failed and suboptimal interactions provide valuable evidence for dynamics modeling and value learning. For data scaling, Motus2 progresses from large-scale monocular egocentric data to synchronized stereo egocentric data, followed by robot-domain adaptation with robot trajectories and supplementary human-robot alignment data. Motus2 further studies global-autoregressive and hybrid-memory extensions of its sliding-window context, adds tactile feedback for contact-aware control, and is instantiated on a fully biomimetic platform with stereo vision, dual arms, dual dexterous hands, and tactile sensing. Together, egocentric data scaling and closed-loop general world model scaling provide a general path toward self-evolving dexterous manipulation.
Chinese Translation
通用具身智能体应当在统一系统中实现感知、预测、行动、评估与改进。世界模型在构建此类智能体方面展现出巨大前景,但现有模型通常只是在世界模拟器上附加一个动作输出头,而未将其耦合为一个用于策略改进的闭环决策与学习循环。我们提出Motus2,一个面向灵巧操作的自我进化通用世界模型。Motus2通过模型扩展和数据扩展推进世界建模。在模型扩展方面,一个共享权重的单一模型提供三个控制接口:策略(世界-动作模型)、模拟器(动作条件世界模型)和评估器(价值模型)。策略提出候选动作块,模拟器预测其视觉后果,评估器对预测结果进行评估。三者的耦合形成了用于策略改进的闭环决策与学习循环。这一构建方式利用精选的专家演示进行动作学习,同时失败的与次优的交互为动力学建模和价值学习提供了宝贵的证据。在数据扩展方面,Motus2从大规模单目自我中心数据推进到同步双目自我中心数据,随后通过机器人轨迹以及补充性的人机对齐数据进行机器人领域适配。Motus2进一步研究了其滑动窗口上下文的全局自回归和混合记忆扩展,增加了触觉反馈以实现接触感知控制,并在一个具备双目视觉、双臂、双灵巧手和触觉感知的完全仿生平台上进行实例化。总之,自我中心数据的扩展与闭环通用世界模型的扩展为自我进化的灵巧操作提供了一条通用路径。
cs.RO / 50 / 2608.30242

CanonNav: Disentangling Navigation Behavior from Camera Geometry in Cross-Platform Visual Navigation

CanonNav:在跨平台视觉导航中将导航行为与相机几何解耦
Kim, Dong-Wook, Hwang, Ji-Hoon, Son, E-In, Oh, Mintaek, Seo, Seung-Woo
Abstract
While visual navigation has advanced through imitation learning from cross-platform demonstrations, fully leveraging such data remains challenging. First, directly learning from image-trajectory pairs entangles navigation behavior with platform-dependent camera geometry. This hinders consistent learning by forcing the policy to implicitly infer camera geometry from visual observations, an inherently ill-posed problem. Second, imitation learning from demonstrated trajectories captures the expert's chosen motion but leaves the intermediate decisions underlying that motion implicit. To address these issues, we propose CanonNav, a visual navigation framework that disentangles navigation behavior from camera geometry and incorporates complementary planning supervision into learning from cross-platform demonstrations. CanonNav introduces camera geometry canonicalization, which transforms visual observations and trajectories into a camera-consistent representation space. Building on this representation, we derive safety and local-progress supervision using pseudo-labels from an offline traversability estimator. Safety supervision penalizes unsafe trajectories, while local-progress supervision guides where the robot should advance. Experiments across diverse camera configurations and environments show that, despite using only RGB at inference, CanonNav consistently outperforms RGB-based baselines and even surpasses RGB-D-based methods in challenging scenarios.
Chinese Translation
尽管视觉导航已通过从跨平台演示数据中进行模仿学习取得进展,但如何充分利用此类数据仍具挑战性。首先,直接从图像-轨迹对中学习会将导航行为与平台相关的相机几何纠缠在一起。这迫使策略不得不从视觉观测中隐式推断相机几何,而这一问题本质上是病态的,从而阻碍了一致性学习。其次,从演示轨迹中进行模仿学习只能捕捉专家选择的动作,而无法显式表达支撑该动作的中间决策。为解决这些问题,我们提出了CanonNav,一个将导航行为与相机几何解耦、并将互补的规划监督引入跨平台演示学习的视觉导航框架。CanonNav引入了相机几何规范化(camera geometry canonicalization),将视觉观测和轨迹变换到相机一致的表示空间中。基于该表示,我们利用来自离线可通行性估计器的伪标签,推导出安全性和局部前进监督。安全性监督对不安全轨迹进行惩罚,而局部前进监督则引导机器人应当向何处前进。在多种相机配置和环境下的实验表明,尽管推理时仅使用RGB输入,CanonNav始终优于基于RGB的基线方法,甚至在高难度场景中超越了基于RGB-D的方法。
cs.RO / 51 / 2608.30289

CometVLA: Co-Training on an Embodied Data Pyramid towards Physical Understanding

CometVLA:面向物理理解的具身数据金字塔协同训练
Wan, Hanwen, Chi, Dafeng, Zhai, Linbo, Shen, Tianao, Zhuang, Yuzheng, Zhang, Tianle, Liu, Peidong, Lin, Liang, Ji, Xiaoqiang
Abstract
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Current physical VQA data is typically disembodied and misaligned with robot action domains. Egocentric videos are used only as auxiliary pre-training. It remains unclear whether improved VLM physical understanding actually benefits downstream action generation. Therefore, we present CometVLA to close this gap. We construct CometData and CometBench, an embodied physical VQA corpus and benchmark strictly aligned with the robot's action data and embodiment. We introduce Global Action Prior (GAP) tokens, a compact learnable bottleneck that isolates task-agnostic motion regularities and lets the action head consume physical commonsense without corrupting the pre-trained VLM backbone. We co-train CometVLA across the embodied data pyramid, spanning teleoperation, simulation, egocentric trajectories, and VQA layers. On real-world manipulation tasks and RoboTwin simulation, CometVLA consistently outperforms strong VLA baselines. Correlation analysis shows that stronger VLM performance on CometBench indicates higher VLA success rates. Results demonstrate that physical understanding pre-training genuinely benefits downstream manipulation.
Chinese Translation
视觉-语言-动作(VLA)模型在需要物理常识的操作任务中仍然脆弱。现有的物理视觉问答(VQA)数据通常缺乏具身性,且与机器人动作域不对齐;自我中心视频仅被用作辅助预训练。目前尚不清楚 VLM 物理理解的提升是否真正有益于下游动作生成。为此,我们提出 CometVLA 以弥合这一差距。我们构建了 CometData 和 CometBench——一个与机器人动作数据及其具身形态严格对齐的具身物理 VQA 语料库与基准。我们引入全局动作先验(Global Action Prior, GAP)token,这是一种紧凑的可学习瓶颈,能够分离任务无关的运动规律,使动作头在不破坏预训练 VLM 骨干网络的前提下利用物理常识。我们在具身数据金字塔(涵盖遥操作、仿真、自我中心轨迹以及 VQA 层)上对 CometVLA 进行协同训练。在真实世界操作任务和 RoboTwin 仿真中,CometVLA 持续优于强大的 VLA 基线。相关性分析表明,VLM 在 CometBench 上的表现越强,对应 VLA 的成功率越高。实验结果证明,物理理解预训练确实能有益于下游操作任务。
cs.RO / 52 / 2608.30301

Data-Centric Neuromotor Interfaces for Portable Human-Machine Interaction

面向便携式人机交互的数据中心神经运动接口
Li, Jiaxuan, Wu, Di, Liu, Jianhua, Zhao, Yuxin, Li, Jinnuo, Zhang, Xiao, Ying, Zhenzhi, Dai, Changsheng, Li, Xiang, Shu, Liming
Abstract
Dexterous human-machine interaction requires intuitive and expressive interfaces that can be efficiently deployed on constrained edge devices. Flexible material-based neuromotor interfaces hold considerable promise, as they decode human movement intention into natural control. Although emerging flexible electronic skins enable wearable high-fidelity data acquisition, practical deployment inevitably involves trade-offs between computational resources and portability. We present a data-centric paradigm where physiological features yield fundamental separability, providing sufficient discriminative cues for recognition. A wireless, high-bandwidth system developed for collecting various electrophysiological signals, when integrated with muscle-specific electrodes, forms a surface electromyography-based interface. Exploiting highly separable data, a 2,210-parameter model achieves 94.36% accuracy across 34 gestures and can be rapidly deployed on edge devices, establishing a new thousand-parameter benchmark for dexterous decoding. The underlying data-algorithm interactions in the data-centric paradigm are further clarified, demonstrating its feasibility in real-world scenarios. This study provides a principled and validated pathway for practical deployment of reliable neuromotor interfaces.
Chinese Translation
灵巧的人机交互需要直观且富有表现力的接口,并能高效部署于资源受限的边缘设备上。基于柔性材料的神经运动接口具有巨大前景,因为它们能够将人体运动意图解码为自然控制。尽管新兴的柔性电子皮肤实现了可穿戴的高保真数据采集,但实际部署不可避免地需要在计算资源与便携性之间进行权衡。我们提出了一种以数据为中心的范式,其中生理特征本身具有基本的可分性,能够为识别提供充分的判别线索。我们开发了一套无线、高带宽系统,用于采集多种电生理信号,并与肌肉特异性电极集成,构成基于表面肌电图的接口。利用高度可分的数据,一个仅含2,210个参数的模型在34种手势上达到了94.36%的准确率,并可快速部署在边缘设备上,为灵巧解码确立了新的千参数级基准。本研究进一步阐明了该数据中心范式中数据与算法的相互作用,证明了其在真实场景中的可行性。这项研究为可靠神经运动接口的实际部署提供了一条有原则且经过验证的路径。
cs.RO / 53 / 2608.30368

SpectraTac: A Compact Camera-Free Optical Tactile Sensor with Distributed Color Sensing

SpectraTac:一种采用分布式颜色传感的紧凑型无相机光学触觉传感器
Wu, Hao, Guo, Haotian, Feng, Yu, Wang, Yutong, Wang, Yanzhe, Zhou, Jianshu
Abstract
Tactile sensing is essential for physical interaction in robotics and human--machine systems. However, combining rich tactile information with compact hardware, low cost, and low computational overhead remains challenging. This work presents SpectraTac, a compact, camera-free optical tactile sensor that combines active red--green--blue (RGB) illumination with spatially distributed color sensing. Contact deforms a compliant transparent elastomer and modulates its internal light field, producing spatially differentiated changes in color and intensity. Three distributed color sensors capture these responses as low-dimensional spatio-spectral features, avoiding cameras, imaging optics, and high-dimensional image processing. The device measures 19.2 mm in diameter and 4 mm in height, with a material cost below USD~5. A data-driven decoding framework simultaneously estimates three-dimensional (3D) force and the contact region from the optical measurements. For 3D force prediction, the sensor achieved mean absolute errors (MAEs) of 0.161, 0.164, and 0.429 N along the x-, y-, and z-axes, respectively. The nine-region contact-classification accuracy was 99.9%. We further evaluated real-time 3D force tracking and contact-region-based human--machine interaction through an interactive control task. These results indicate that distributed color-resolved optical sensing offers a compact, low-cost alternative to camera-based tactile sensing for robotics, wearable sensing, and interactive systems.
Chinese Translation
触觉感知对于机器人和人机系统中的物理交互至关重要。然而,如何在硬件紧凑、低成本和低计算开销的前提下获取丰富的触觉信息,仍然是一项挑战。本文提出了SpectraTac,一种结合主动红-绿-蓝(RGB)照明与空间分布式颜色传感的紧凑型无相机光学触觉传感器。接触使柔性透明弹性体发生形变并调制其内部光场,从而产生空间上有差异的颜色和强度变化。三个分布式颜色传感器将这些响应捕获为低维的空谱特征,从而避免了相机、成像光学器件和高维图像处理。该器件直径为19.2毫米,高度为4毫米,材料成本低于5美元。一种数据驱动的解码框架从光学测量中同时估计三维(3D)力和接触区域。在三维力预测方面,该传感器沿x、y和z轴的平均绝对误差(MAE)分别为0.161、0.164和0.429牛。九区域接触分类准确率达到99.9%。我们进一步通过一个交互控制任务评估了实时三维力跟踪以及基于接触区域的人机交互。这些结果表明,分布式颜色分辨光学传感为机器人、可穿戴传感和交互系统提供了一种紧凑、低成本的无相机触觉感知替代方案。
cs.RO / 54 / 2608.30378

PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies

PAVE:面向世界-动作策略的预测对齐与价值引导演化
Zhao, Botong, Yu, Fang, Yu, Tim, Zhu, Senhua, Chen, Xinyuan, Lu, Yue
Abstract
Direct vision-language-action policies generate continuous robot actions efficiently, but standard behavior cloning leaves two complementary gaps: their representations are not explicitly required to describe how the scene evolves over multiple time scales, and deployment trajectories of unequal quality are often reused without separating useful dynamics from undesirable behavior. We introduce \method, a direct world-action policy that combines outcome-agnostic predictive learning with outcome-aware policy improvement. \method first retains a local fixed-offset JEPA objective and adds trajectory-relative multi-horizon transition alignment at 25%, 50%, 75%, and 100% of the remaining episode. These training-only targets require the current policy representation to preserve both local physical changes and longer-range task progress, without supplying explicit future tokens to the action head. \method then trains an independent distributional value critic on cumulative deployment trajectories, computes action-chunk-aligned $N$-step advantages, and converts them into positive, negative, or null text conditions for a flow-matching actor. Thus, every valid trajectory can teach what physically happened, while the actor is deployed only under the condition associated with relatively better actions. The multi-horizon predictor and critic are removed from online execution, preserving direct action generation from the current observation, language instruction, and proprioception. \redclaim{Across the three simulation benchmarks, \method achieves the strongest overall performance while preserving the direct actor's online execution path.}
Chinese Translation
直接的视觉-语言-动作策略能够高效地生成连续的机器人动作,但标准的行为克隆存在两个互补的缺陷:其表示并未被显式地要求描述场景如何在多个时间尺度上演化;且质量参差不齐的部署轨迹往往被直接复用,而未将有用的动力学与不良行为区分开来。我们提出PAVE(Predictive Alignment and Value-Guided Evolution),一种将结果无关的预测学习与结果感知的策略改进相结合的直接世界-动作策略。PAVE首先保留局部的固定偏移JEPA目标,并在剩余片段的25%、50%、75%和100%处添加轨迹相对的多时域转移对齐。这些仅用于训练的目标要求当前策略表示同时保留局部物理变化和更长期的任务进展,而无需向动作头提供显式的未来令牌(token)。随后,PAVE在累积的部署轨迹上训练一个独立的分布式价值评论家(critic),计算与动作块对齐的N步优势,并将其转换为正向、负向或空白的文本条件,用于流匹配执行器。这样,每条有效轨迹都能教授物理上实际发生了什么,而执行器仅在与相对更优动作相关联的条件下部署。多时域预测器和评论家在在线执行中被移除,从而保持了从当前观测、语言指令和本体感受直接生成动作的执行路径。PAVE在三个仿真基准上取得了最强的整体性能,同时保留了直接执行器的在线执行路径。
cs.RO / 55 / 2608.30433

A Hybrid PEM-GP Framework for Uncertainty-Aware System Identification of Quadcopters

一种用于四旋翼无人机不确定性感知系统辨识的PEM-GP混合框架
Ghoul, Abdallah, Bousserhane, Ismail Khalil, Boufeldja, Kadri
Abstract
Accurate dynamic models play a central role in achieving reliable control of quadcopters. Classical system identification methods remain widely used, mainly because of their interpretability. However, they often fail to capture important nonlinear effects, especially in small-scale aerial platforms where such effects become more pronounced. Data-driven approaches offer a different perspective. They can represent complex nonlinear dynamics more effectively, but this comes at the cost of reduced interpretability and the absence of well-calibrated uncertainty estimates. In this work, we propose a framework that combines physics-based modeling with data-driven learning, while explicitly accounting for uncertainty. A physics-based model is first identified using the Prediction Error Method (PEM), which captures the main structure of the system. The remaining dynamics are then modeled using a Gaussian Process (GP), allowing the residual behavior to be learned directly from data. This separation makes it possible to distinguish between known physical effects and unmodeled dynamics. The proposed framework is validated on a Duckiedrone-like experimental setup. The results show that the PEM-GP model achieves prediction accuracy comparable to that of a Long Short-Term Memory (LSTM) network, while additionally providing calibrated uncertainty estimates. This combination improves model reliability and supports uncertainty-aware decision-making.
Chinese Translation
精确的动力学模型在实现四旋翼无人机的可靠控制中起着核心作用。经典系统辨识方法因其良好的可解释性至今仍被广泛使用。然而,这些方法往往难以捕捉重要的非线性效应,尤其是在此类效应更为显著的小型飞行平台上。数据驱动方法提供了另一种视角,能够更有效地表征复杂的非线性动力学,但其代价是可解释性降低,且缺乏经过良好校准的不确定性估计。本文提出一种将基于物理的建模与数据驱动学习相结合、并显式考虑不确定性的框架。首先使用预测误差法(Prediction Error Method, PEM)辨识基于物理的模型,以捕捉系统的主要结构;随后利用高斯过程(Gaussian Process, GP)对剩余动力学进行建模,使残余行为可以直接从数据中学习。这种分离方式使得能够区分已知的物理效应和未建模的动力学。所提出的框架在类似Duckiedrone的实验平台上进行了验证。结果表明,PEM-GP模型达到了与长短期记忆网络(Long Short-Term Memory, LSTM)相当的预测精度,同时还能提供经过校准的不确定性估计。这种结合提升了模型的可靠性,并支持不确定性感知的决策。
cs.RO / 56 / 2608.30506

Anomaly Detection on Small Industrial Components via Vision-Based Tactile Sensing

基于视觉触觉传感的小型工业部件异常检测
Preziosa, G. F., Casiglia, M., Faroni, M., Zanchettin, A. M., Rocco, P.
Abstract
Automated inspection of small industrial components, including sub-centimetre-scale parts where defects are geometry-driven and poorly resolved by standard optical cameras, calls for sensing modalities that can directly capture fine surface geometry. Vision-based tactile sensors address this need by converting contact imprints into high-resolution image-like data compatible with existing deep-learning pipelines, yet their effective use for industrial anomaly detection (AD) remains largely unexplored. This work systematically evaluates unsupervised AD methods on a real tactile dataset covering five genuine industrial components acquired with a GelSight Mini sensor mounted on a collaborative robot. Four feature-embedding methods, SPADE, PaDiM, FAPM, and InReaCh, are compared under three validations explicitly motivated by the deployment constraints of contact-based sensing: a Good Fraction analysis establishing the minimum number of nominal contacts for stable performance, directly bounded by gel wear since every acquisition degrades the soft interface; a cross-position evaluation assessing generalization across different contact locations observing the same recurring surface pattern; and a low- versus high-resolution comparison evaluating the cost-benefit of higher-resolution tactile acquisition. Overall, this systematic benchmarking study provides practical guidance for researchers and practitioners adopting vision-based tactile sensing for industrial AD and shows how this modality can serve as a viable alternative for industrial quality-control tasks.
Chinese Translation
小型工业部件的自动化检测——包括缺陷由几何形状驱动且标准光学相机难以清晰分辨的亚厘米级零件——需要能够直接捕捉精细表面几何信息的传感模态。基于视觉的触觉传感器通过将接触压痕转换为与现有深度学习流程兼容的高分辨率类图像数据来满足这一需求,但其在工业异常检测(AD)中的有效应用仍在很大程度上未被探索。本工作在一个真实的触觉数据集上系统评估了无监督异常检测方法,该数据集涵盖安装在协作机器人上的 GelSight Mini 传感器采集的五种真实工业部件。我们比较了四种特征嵌入方法——SPADE、PaDiM、FAPM 和 InReaCh——并在三种由接触式传感部署约束直接驱动的验证设置下进行评估:一是良好样本比例(Good Fraction)分析,用于确定实现稳定性能所需的名义接触最小数量,由于每次采集都会损耗软质界面,该数量直接受凝胶磨损的制约;二是跨位置评估,用于评估对不同接触位置(观察到相同重复性表面图案)的泛化能力;三是低分辨率与高分辨率的对比,用于评估更高分辨率触觉采集的成本效益。总体而言,这项系统的基准测试研究为采用基于视觉的触觉传感进行工业异常检测的研究者和从业者提供了实用指导,并展示了该模态如何成为工业质量控制任务的可行替代方案。
cs.RO / 57 / 2608.30536

Behavior-Skill: A Fine-Grained Benchmark for Evaluating Vision-Language-Action Policies in Long-Horizon Tasks

Behavior-Skill:用于评估长时程任务中视觉-语言-动作策略的细粒度基准
Ma, Chunyun, Luo, Lun, Luo, Xingjian, Feng, Xiexing, Zhang, Hang, Liu, Wei, Qiao, Feng, Wang, Yaonan, Lu, Huimin, Chen, Xieyuanli
Abstract
Reliable execution of long-horizon mobile manipulation tasks remains challenging because overall task success depends on the successful completion of multiple constituent skills. Existing benchmarks, however, still rely primarily on full-task rollouts and aggregate task-level metrics, making intermediate failures difficult to observe and analyze. We present Behavior-Skill, a benchmark that reformulates the learning and evaluation of long-horizon tasks around executable constituent skills. It contains 235,492 skill instances from 10,000 demonstrations across 50 household tasks and 34 semantic skill categories. Each instance pairs a skill instruction with an aligned observation-action segment, and is further associated with a restorable intermediate state and a skill success condition to enable independent evaluation under valid preconditions. We further introduce trajectory-level and skill-level metrics to characterize policy capability beyond aggregate task success. Extensive experiments across representative VLA policies including pi0.5 and GR00T on the complete 50-task benchmark show that failures are highly non-uniform across skills, with contact-rich manipulation skills forming persistent bottlenecks. These results demonstrate that Behavior-Skill complements full-task evaluation by exposing intermediate capability profiles for analyzing and improving long-horizon VLA policies. Behavior-Skill is publicly available at https://github.com/nubot-nudt/Behavior-Skill.
Chinese Translation
长时程移动操作任务的可靠执行仍然具有挑战性,因为整体任务的成功依赖于多个组成技能的顺利完成。然而,现有基准主要依赖完整任务轨迹回放和聚合的任务级指标,使得中间失败难以观察和分析。我们提出 Behavior-Skill,该基准将长时程任务的学习与评估重新构建为围绕可执行的组成技能展开。它包含来自 50 个家庭任务、10,000 条演示中的 235,492 个技能实例,涵盖 34 个语义技能类别。每个实例将技能指令与对齐的观测-动作片段配对,并进一步关联一个可恢复的中间状态和技能成功条件,以便在有效前提条件下进行独立评估。我们进一步引入轨迹级和技能级指标,以刻画超越聚合任务成功的策略能力。在包含 pi0.5 和 GR00T 等代表性 VLA 策略的完整 50 任务基准上进行的大量实验表明,失败在各个技能之间高度不均匀,其中富含接触的操作技能构成持续性瓶颈。这些结果表明,Behavior-Skill 通过揭示中间能力特征,补充了完整任务评估,可用于分析和改进长时程 VLA 策略。Behavior-Skill 已公开发布于 https://github.com/nubot-nudt/Behavior-Skill。
cs.RO / 58 / 2608.30643

Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models

时间强迫:面向视觉-语言-动作模型的4D表征对齐
Ding, Xingyu, Zhao, Yuzhong, Zhao, Chunhai, Shi, Yinghuan, Zhao, Chaoyang, Zhang, Yifan
Abstract
Recent vision-language-action (VLA) methods improve manipulation performance by aligning their representations with 3D scene geometry. However, these methods often struggle with long-horizon manipulation and observation aliasing between visually similar states due to a lack of temporal information: the 3D scene geometry captures only the current state, rather than how it has evolved over time. To resolve this, we present Temporal Forcing, a 4D representation alignment method for VLA models. Specifically, we first introduce a history pathway that enables a vanilla VLA model to summarize observation history into temporally aware latent representations. Then, the latent representations are aligned with the geometric features extracted by a pretrained 4D foundation model, which captures the evolving 3D world through temporally consistent geometric representations, enabling a deeper understanding of dynamic environments. Temporal Forcing reaches 98.8% on LIBERO, outperforming its base model by 2.2 points. On a physical hidden-placement task, it raises full-task success from 20.0% to 43.3%. Code will be publicly available.
Chinese Translation
近期的视觉-语言-动作(VLA)方法通过将其表征与3D场景几何对齐来提升操作性能。然而,由于缺乏时间信息,这些方法在长时程操作以及视觉相似状态之间的观测混叠问题上常常表现不佳:3D场景几何仅刻画当前状态,而无法反映其随时间的演化过程。为解决这一问题,我们提出了Temporal Forcing,一种面向VLA模型的4D表征对齐方法。具体而言,我们首先引入一条历史通路,使原始VLA模型能够将观测历史总结为具备时间感知能力的潜在表征。然后,将这些潜在表征与由预训练4D基础模型提取的几何特征对齐;该模型通过时间一致的几何表征刻画不断演化的3D世界,从而实现对动态环境的更深理解。Temporal Forcing在LIBERO上达到98.8%的成功率,较其基础模型提升2.2个百分点。在一个实体隐藏放置任务中,它将全任务成功率从20.0%提升至43.3%。代码将公开提供。
cs.RO / 59 / 2608.30673

CIG-RL: Curiosity-Driven Information-Guided Reinforcement Learning for Source Term Estimation in Uncertain Environments

CIG-RL:面向不确定环境下源项估计的好奇心驱动信息引导强化学习
Lee, Junhee, Kim, Seunghwan, Jang, Hongro, Kim, Hyungjin, Park, Hyoungho, Kim, Changseung, Oh, Hyondong
Abstract
Source term estimation (STE), which aims to estimate key properties of the gas source, is essential for identifying hazardous gas releases. Information-theoretic approaches have been adopted for autonomous STE using mobile sensors due to robustness in noisy environments, yet their online action selection incurs substantial computational cost. Deep reinforcement learning (DRL) provides a promising alternative with its fast decision-making capability. In DRL-based STE, the agent selects actions based on belief states of the source term updated from noisy measurement sequences. However, existing methods rely on random exploration or solely on belief uncertainty reduction without an effective exploration strategy in DRL, which can limit policy robustness in noisy environments. To address this, we propose a curiosity-driven information-guided reinforcement learning for robust and efficient STE. The proposed method promotes active exploration of novel belief state transitions that have not been sufficiently explored during training. We further introduce an uncertainty-adaptive active perception reward to guide efficient source search under uncertainty. Simulations under high-noise conditions and real-world experiments demonstrate the robustness and feasibility of the proposed framework, highlighting its potential for practical STE problems.
Chinese Translation
源项估计(Source Term Estimation, STE)旨在估计气体源的关键属性,对于识别有害气体泄漏至关重要。基于信息论的方法由于在噪声环境中具有鲁棒性,已被广泛用于基于移动传感器的自主源项估计,但其在线动作选择会带来较大的计算开销。深度强化学习(Deep Reinforcement Learning, DRL)凭借其快速的决策能力为该问题提供了一种有前景的替代方案。在基于DRL的源项估计中,智能体根据由含噪测量序列更新的源项信念状态来选择动作。然而,现有方法依赖于随机探索,或仅依靠信念不确定性的缩减,缺乏DRL中有效的探索策略,这可能限制策略在噪声环境中的鲁棒性。为解决这一问题,我们提出了一种好奇心驱动的信息引导强化学习方法,以实现鲁棒且高效的源项估计。所提方法促进对训练过程中未被充分探索的新型信念状态转移进行主动探索。我们进一步引入了一种不确定性自适应的主动感知奖励,以指导不确定性条件下的高效源搜索。在高噪声条件下的仿真和真实世界实验验证了所提框架的鲁棒性与可行性,凸显了其在实际源项估计问题中的应用潜力。
cs.RO / 60 / 2608.30773

Learning to infer and manipulate through distributed whole-arm interaction in a soft robot

通过软体机器人的分布式全臂交互学习推理与操作
Zhang, Chuhan, Shahabi, Ebrahim, Khomenko, Kseniia, Pan, Wei, Della Santina, Cosimo
Abstract
In animals such as elephants and octopuses, acquiring non-visual information about an object and physically engaging with it are inseparable processes mediated by rich, large-area interactions between compliant appendages and the environment. Soft robots provide a natural platform for translating this principle into engineered systems. Yet current robotic intelligence makes limited use of physical interaction, treating it primarily as a disturbance to be rejected or, at best, as a means of compensating for object misalignment. Here, we introduce a physical intelligence framework in which distributed compliant interactions jointly reveal task-relevant information and organize manipulation behavior. This results in an intrinsically partially observable problem: key task-relevant information is never measured directly, but must instead be inferred from the history of physical interactions. We propose a reinforcement-learning architecture that addresses this challenge by learning a memory-based control policy end-to-end. The key innovations making this possible are (i) a pretrained exploration policy that provides a reference for broad workspace exploration, (ii) joint optimization that integrates exploration and grasping objectives within a single recurrent policy, and (iii) a two-stage sim-to-real adaptation including observation mapping and policy fine-tuning. We demonstrate this principle through blind whole-arm grasping with a hybrid rigid-soft robotic arm that we equip with IMUs embedded directly within its compliant structure, providing its only source of proprioceptive sensing. The learned policy successfully identifies and grasps various objects by autonomously coordinating workspace exploration, object encounter and localization, inference of grasp-relevant properties, and stable whole-arm wrapping.
Chinese Translation
在诸如大象和章鱼等动物中,获取物体的非视觉信息并与之进行物理接触,是由柔性附肢与环境之间丰富的大面积交互所介导的、不可分割的过程。软体机器人为将这一原理转化为工程系统提供了天然平台。然而,当前的机器人智能对物理交互的利用十分有限,主要将其视为需要抑制的扰动,或者充其量将其作为补偿物体错位的一种手段。本文提出了一种物理智能框架,其中分布式柔性交互共同揭示与任务相关的信息,并组织操作行为。这导致了一个本质上部分可观测的问题:关键的与任务相关的信息从未被直接测量,而必须从物理交互的历史中推断出来。我们提出了一种强化学习架构,通过端到端学习基于记忆的控制策略来应对这一挑战。使之成为可能的关键创新包括:(i)预训练的探索策略,为大范围工作空间探索提供参考;(ii)在单一循环策略中联合优化探索与抓取目标;(iii)包含观测映射与策略微调的两阶段sim-to-real(仿真到现实)迁移。我们通过混合刚-软机械臂的盲全臂抓取来验证这一原理。该机械臂的柔性结构中直接嵌入了IMU(惯性测量单元),作为其唯一的本体感受感知来源。学习到的策略能够通过自主协调工作空间探索、物体接触与定位、抓取相关属性的推断以及稳定的全臂缠绕,成功地识别并抓取各种物体。
cs.RO / 61 / 2608.30832

A Dual-Cam Parallel Elastic Actuator with Shared Gas-Spring Compensation for Humanoid Ankles

一种用于仿人踝关节的共享气弹簧补偿双凸轮并联弹性执行器
Jiang, Jingcheng, Zhang, Yifang, Tsagarakis, Nikos G.
Abstract
To improve torque capacity and energy efficiency of humanoid ankles, this paper proposes a 2-DoF parallel elastic actuator (PEA). The main novelty of the proposed design lies in its dual-cam, single-gas-spring architecture, which enables torque compensation in both pitch and roll using a shared elastic element, thereby improving structural compactness compared with conventional multi-element compensation schemes. By leveraging parallel gas springs and customized cam modules, the proposed architecture provides dual-axis torque assistance tailored to specific task requirements. The second key contribution is the formulation of a coupled 2-DoF mathematical model that explicitly captures the interdependence between the two compensation units through the shared spring. Based on this model, an optimization-based design framework is developed to synthesize customized cam profiles from prescribed torque references, establishing a systematic link from task requirements to hardware realization. The complete lower-leg CAD integration is presented in detail. Static FEA and kinematic simulations confirm the design's feasibility and torque-relief effectiveness. The results highlight the proposed design as a compact, customizable solution for 2-DoF humanoid ankle torque compensation.
Chinese Translation
为提高仿人踝关节的力矩容量与能量效率,本文提出一种二自由度并联弹性执行器(PEA)。所提设计的主要创新点在于其双凸轮-单气弹簧架构,通过共享弹性元件实现俯仰与翻滚两个方向的力矩补偿,从而相比传统的多元件补偿方案提升了结构紧凑性。借助并联气弹簧与定制凸轮模块,该架构可提供针对特定任务需求定制的双轴力矩辅助。第二个关键贡献是建立了显式刻画两个补偿单元通过共享弹簧相互耦合关系的二自由度数学模型。基于该模型,开发了基于优化的设计框架,可从给定的力矩参考合成定制凸轮轮廓,建立了从任务需求到硬件实现的系统化联系。文中详细展示了完整小腿CAD集成设计。静态有限元分析(FEA)与运动学仿真验证了设计的可行性及力矩卸载有效性。结果表明,所提设计是一种紧凑、可定制的二自由度仿人踝关节力矩补偿方案。
cs.RO / 62 / 2608.30858

GAFT: Geo-Anchored Fine-Tuning for Hazard Identification from Rare Failures

GAFT:基于几何锚定的微调方法,用于从罕见失效中识别危险
Xu, Yanran, Qiu, Chuanhang, Wang, Yue, Wu, Wenbo, Li, Zhaoxing
Abstract
Off-road navigation can fail when physical structures induce irrecoverable states such as high-centering or entrapment, requiring human interventions. Identifying these structures is crucial, yet challenging. Such failure events are rare and costly to collect, resulting in limited training data. Moreover, the collected data associate frames with outcomes, but do not indicate the visual cues responsible for the failure. Learning directly from these data can therefore exploit scenario-specific visual cues, leading to poor generalization. We propose \textbf{Geo-Anchored Fine-Tuning (GAFT)}, a parameter-efficient method that adapts a vision foundation model with a geometry-derived prior. It guides LoRA adaptation by aligning a spatial attention-rollout map with the geometry prior, while preserving pretrained representations. On an intervention-verified forest hazard benchmark, across ten independently trained adaptations, GAFT consistently outperforms frozen DINOv2 and supervised PEFT baselines, improving the repeated leave-one-scenario-out mean $F_2$ from 0.0607 to 0.3757 with statistical significance under paired analysis. Within these independently trained models, the best-performing GAFT model achieves a repeated-LOSO $F_2$ of 0.570. Code and benchmark: https://github.com/Xu-Yanran/geo_anchored_fine_tuning
Chinese Translation
越野导航可能因物理结构导致不可恢复的状态(如底盘搁浅或陷入)而失效,从而需要人工干预。识别这些结构至关重要,但也颇具挑战性。此类失效事件罕见且采集成本高昂,导致训练数据有限。此外,所采集的数据仅将图像帧与结果相关联,并未指出导致失效的视觉线索。因此,直接从这些数据中学习可能会利用特定场景的视觉线索,导致泛化能力差。我们提出了几何锚定微调(Geo-Anchored Fine-Tuning, GAFT),这是一种参数高效的方法,利用由几何推导出的先验对视觉基础模型进行适配。该方法通过将空间注意力展开图(attention-rollout map)与几何先验对齐来引导LoRA适配,同时保留预训练的表示。在一个经干预验证的森林危险基准测试中,在十次独立训练的适配实验中,GAFT始终优于冻结的DINOv2和有监督PEFT基线方法,将重复的留一场景(leave-one-scenario-out)平均$F_2$分数从0.0607提升至0.3757,且在配对分析下具有统计显著性。在这些独立训练的模型中,表现最佳的GAFT模型达到了0.570的重复留一场景$F_2$分数。代码与基准测试:https://github.com/Xu-Yanran/geo_anchored_fine_tuning
cs.RO / 63 / 2608.30880

Zeva: In-Context Causal Learning for Generalizable Embodied Manipulation

Zeva:面向可泛化具身操作的同上下文因果学习
Chen, Fu, Ding, Xin, Huang, Bingjia, Li, Xiangyu, Wang, Mingju, He, Jiawei, Li, Kun, Sun, Wei, Liu, Yunxin, Wu, Hao, Cao, Ting
Abstract
Generalizable embodied manipulation remains difficult to achieve through pretraining alone, due to unseen physical conditions in the real world. We argue that robots need to learn from their own physical interactions on the fly during real-world deployment and use this knowledge to inform subsequent actions. We present Zeva, the first framework that enables in-context learning from a robot's own physical interaction experience while keeping the policy model frozen. Zeva employs a Causal Interaction Extractor to encode an executed action and its induced state change into a causal interaction signal, which is stored in a dual-timescale causal memory. For subsequent actions, relevant causal interaction signals are retrieved from memory and injected into the frozen policy model as context. Experiments in simulation and real-world manipulation demonstrate that Zeva achieves the best performance among the compared frontier VLAs and WAMs and, more importantly, enables self-evolution during deployment without gradient updates. Its success rate continues to improve as the robot accumulates interaction experience. Furthermore, the acquired interaction experience can generalize across tasks.
Chinese Translation
由于真实世界中存在未见的物理条件,仅依靠预训练难以实现可泛化的具身操作。我们认为,机器人需要在真实世界部署过程中实时地从自身的物理交互中学习,并利用这些知识指导后续动作。我们提出了Zeva,这是首个能够在策略模型保持冻结的情况下,从机器人自身物理交互经验中进行同上下文学习的框架。Zeva采用因果交互提取器(Causal Interaction Extractor)将已执行的动作及其引起的状态变化编码为因果交互信号,并将其存储在双时间尺度的因果记忆中。在执行后续动作时,相关的因果交互信号会从记忆中被检索出来,并作为上下文注入冻结的策略模型。在仿真和真实世界操作中的实验表明,Zeva在所比较的前沿视觉-语言-动作模型(VLA)和世界动作模型(WAM)中取得了最佳性能,更重要的是,它无需梯度更新即可在部署过程中实现自我进化。随着机器人不断积累交互经验,其成功率持续提升。此外,所获得的交互经验能够跨任务泛化。
cs.RO / 64 / 2608.30883

SleepWalking: Privileged Representation Shaping for End-to-End Blind Locomotion in Legged Robots

SleepWalking:面向足式机器人端到端盲运动控制的特权表征塑形
Pan, Zheng, Wang, Tenghui, Li, Peilin, Zhou, Shiyu, Sun, Hao, Ma, Yan, Yu, Liang, He, Liang
Abstract
Partially observable locomotion requires a policy to act when task-relevant properties of the robot--environment state are not fully specified by instantaneous observations. Existing approaches often address this challenge by explicitly estimating missing physical variables or processing extended observation histories through structured architectures. We take a different view: partial observability is fundamentally an information-retention problem. The decisive question is not how task-relevant information enters the network, but whether the policy's internal state retains it. Guided by this perspective, we propose SleepWalking for Robot Locomotion (SWAQ), a one-stage end-to-end framework that uses next-step privileged physical reconstruction to shape what a recurrent history representation retains during policy learning, while the deployed actor uses only a direct history-to-action pathway. Under aligned training settings, SWAQ achieves a 15.0\% higher peak mean terrain level than DWAQ, the strongest non-exteroceptive baseline, while using 44.4\% fewer inference MACs per control step. Layerwise probes further show that information associated with the reconstructed physical variables remains linearly decodable through the policy head up to the layer preceding the action output. Complementary theoretical analysis relates privileged-variable recoverability to the achievable-return gap between history-based and privileged-information policy classes. These results suggest that semantic objectives can structure learning without requiring a corresponding architectural decomposition of the deployed controller.
Chinese Translation
部分可观测的运动控制要求策略在机器人—环境状态中与任务相关的属性无法由即时观测完全确定时仍能采取行动。现有方法通常通过显式估计缺失的物理变量,或借助结构化架构处理扩展的观测历史来应对这一挑战。我们提出了不同的视角:部分可观测性本质上是一个信息保持问题。关键问题不在于任务相关信息如何进入网络,而在于策略的内部状态是否保留了这些信息。基于这一视角,我们提出了 SleepWalking for Robot Locomotion(SWAQ),这是一个单阶段端到端框架,它利用下一步特权物理重建来塑形循环历史表征在策略学习过程中所保留的信息,而部署时的执行器仅使用从历史到动作的直接通路。在对齐的训练设置下,SWAQ 相比最强的非外感受基线 DWAQ,峰值平均地形等级提高 15.0%,同时每个控制步骤的推理 MAC 计算量减少 44.4%。逐层探测进一步表明,与重建物理变量相关的信息在策略头中直至动作输出之前的一层仍保持线性可解码。互补的理论分析将特权变量的可恢复性与基于历史的策略类别和特权信息策略类别之间的可达回报差距联系起来。这些结果表明,语义目标可以在无需对部署的控制器进行相应架构分解的情况下构建学习过程。
cs.RO / 65 / 2608.30935

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

LightNav-0:激发视觉语言模型空间智能以实现通用具身导航
Wang, Shaoan, Luo, Aocheng, Huang, Fei, Xu, Jingyi, Wang, Xiaoyang, Wang, Yueyu, Ma, Qianli, Yang, Fan, Mei, Ran, Wei, Jia, Hu, Jiangpeng, Liu, Xuhao, Chen, Hongming, Shao, Yuanbin, Lin, Yiyang, Li, Ziliang, Pan, Liang, Liu, Xinhang, Ma, Yuntao, Fan, Tingxiang
Abstract
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial priors for visual grounding, spatial reasoning, and pointing, but these capabilities are rarely elicited directly for robot control. Existing navigation systems instead rely on task- or embodiment-specific components, fragmenting perception, reasoning, and action while offering limited generalization. Here we present LightNav-0, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads. LightNav-0 represents diverse navigation tasks through a unified token interface: dual-channel pointing expresses task-, scene-, and embodiment-agnostic spatial intent, while a residual vector-quantized action tokenizer maps this intent to precise, embodiment-specific trajectories. Together with temporally aware visual history compression, ER mid-training, supervised fine-tuning, and reinforcement learning, this formulation supports instruction following, open-vocabulary object navigation, and visual tracking within a single model. The navigation training corpus spans 2K+ scenes and 4K+ hours of embodied navigation data. LightNav-ER, the embodied-reasoning checkpoint used to initialize LightNav-0, attains the highest complete-set average across 8 embodied-reasoning benchmarks, while LightNav-0 achieves state-of-the-art monocular success rates across all 10 public navigation simulation settings. Real-world evaluations further demonstrate zero-shot generalization across robot embodiments, diverse scenes, and static and dynamic targets. These results establish compact VLMs as a unified and transferable backbone for generalist embodied navigation.
Chinese Translation
具身导航要求智能体跨任务、环境和机器人本体,将异构目标与视觉观察转化为动作。现代视觉语言模型(VLM)已编码了用于视觉定位、空间推理和指向的空间先验,但这些能力很少被直接用于激发机器人控制。现有导航系统则依赖于任务特定或本体特定的组件,割裂了感知、推理与动作,且泛化能力有限。本文提出LightNav-0,一个紧凑的通用具身导航模型,它激发预训练VLM的空间智能并将其与导航对齐,无需任务特定的预测头。LightNav-0通过统一的token接口表示多样化导航任务:双通道指向(dual-channel pointing)表达与任务、场景和本体无关的空间意图,而残差向量量化动作分词器(residual vector-quantized action tokenizer)将该意图映射为精确的、本体特定的轨迹。结合具有时间感知的视觉历史压缩、ER中期训练(ER mid-training)、监督微调与强化学习,该建模方式使单一模型能够支持指令跟随、开放词汇目标导航和视觉跟踪。导航训练语料库涵盖2000多个场景和4000多小时具身导航数据。用于初始化LightNav-0的具身推理检查点LightNav-ER在8个具身推理基准上取得了最高的完整集平均成绩,而LightNav-0在全部10个公开导航仿真设置中实现了最先进的单目成功率。真实世界评估进一步证明了其在机器人本体、多样场景以及静态和动态目标上的零样本泛化能力。这些结果确立了紧凑型VLM作为通用具身导航的统一且可迁移的骨干网络。
cs.RO / 66 / 2608.30983

Autonomously Acquiring Robot Manipulation Skills with Language-Driven Quality-Diversity

基于语言驱动的质量-多样性算法自主习得机器人操作技能
Garrabé, Émiland, Khoramshahi, Mahdi, Doncieux, Stéphane
Abstract
Quality-diversity (QD) algorithms have been gaining traction in robot learning, where diverse motion primitive libraries allow robots to adapt zero-shot to constraints at deployment time. However, such methods typically require expert designers to write the success condition, fitness and diversity metrics, and this strongly limits the robot's autonomy. On the other hand, existing LLM-based reward-shaping techniques allow robots to learn autonomously but only output single high-performing solutions, limiting the robot's adaptability. In this paper, we propose an approach designed to output diverse motion primitive archives by autonomously leveraging quality-diversity algorithms, only requiring a free-form description of the task in common language. To address the difficulty of designing relevant fitness and diversity metrics, we propose an autonomous exploration mechanism able to reliably output sets of functionals covering the fitness and behavior descriptor (BD) space. First, we pose policy exploration as a functional design problem, where the functional spaces are lower-dimensional than the full BD and fitness spaces, and propose an LLM-based exploration scheme to sample from these low-dimensional spaces without any task-specific prompts, fine-tuning or expert intervention. We adapt a multi-BD variant of the MAP-Elites success (MES) algorithm, designed to leverage the heterogeneous BD samples. Finally, through experiments based on the genesis simulator, we show that our method effectively generates archives of diverse motion primitives, outperforming classical QD algorithms with inferred and hand-written parametrizations on a set of $4$ robotic manipulation tasks.
Chinese Translation
质量-多样性算法在机器人学习领域日益受到关注,其多样化的运动基元库使机器人能够在部署时对约束条件进行零样本适应。然而,这类方法通常需要专家设计者编写成功条件、适应度指标和多样性指标,这极大地限制了机器人的自主性。另一方面,现有的基于大语言模型的奖励塑形技术虽然能让机器人自主学习,但只能输出单一的高性能解,限制了机器人的适应能力。本文提出了一种方法,通过自主利用质量-多样性算法输出多样化的运动基元档案,仅需以自然语言对任务进行自由形式的描述。为解决设计相关适应度指标和多样性指标的难题,我们提出了一种自主探索机制,能够可靠地输出覆盖适应度空间和行为描述子(Behavior Descriptor, BD)空间的函数集合。首先,我们将策略探索表述为一个函数设计问题,其中函数空间的维度低于完整的行为描述子空间和适应度空间,并提出一种基于大语言模型的探索方案,从这些低维空间中采样,无需任何任务特定的提示、微调或专家干预。我们改进了MAP-Elites成功(MES)算法的多行为描述子变体,以充分利用异构的行为描述子样本。最后,基于Genesis仿真器的实验表明,我们的方法能够有效生成多样化的运动基元档案,在4个机器人操作任务上优于采用推断参数化和人工编写参数化的经典质量-多样性算法。
cs.RO / 67 / 2608.31002

DARP: A Calibrated Dual-Arm RGB-D-IR Dataset for Multi-View Robotic Perception

DARP:一个用于多视角机器人感知的标定双臂RGB-D-IR数据集
Kansana, Manish, Mujawar, Mohammed Yusuf, Mittal, Sudip, Rahimi, Shahram, Golilarz, Noorbakhsh Amiri
Abstract
Robotic perception from a single viewpoint is often limited by self-occlusion and incomplete surface visibility. This paper presents DARP(Dual-Arm Robotic Perception) https://doi.org/10.21227/rmv3-be47, a calibrated dual-arm RGB-D-IR dataset for object-centered robotic perception using two independently moving eye-in-hand manipulators positioned on opposite sides of a shared tabletop workspace. Each arm carries an Intel RealSense sensor that continuously records RGB, depth, and stereo infrared data while synchronized robot joint states are logged for pose recovery. Objects are placed without fixed poses or marked locations, and the acquisition procedure performs automatic localization, cross-arm confirmation, adaptive viewpoint generation, and continuous multimodal recording. DARP contains ten unique tabletop objects and preserves the original sensor recordings, robot-state logs, object-level metadata, and calibration information required to reconstruct camera trajectories in a shared metric frame. To evaluate the geometric consistency of the acquisition, we implement a deterministic multi-view fusion pipeline that converts calibrated RGB-D observations into complementary partial point clouds and measured surface meshes without using learned or generative completion methods. Evaluation on 224 held-out RGB-D keyframes comprising 1,563,466 three-dimensional query points yields a median point-to-mesh distance of 2.13~mm and an RMSE of 4.04~mm, with 96.56\% of points within 10~mm of the measured-surface mesh. DARP is intended as a reusable resource for multi-view reconstruction, collaborative robotic perception, multimodal fusion, active perception, and future learning-based reasoning over partial object observations.
Chinese Translation
单一视角的机器人感知往往受到自遮挡和表面可见性不完整的限制。本文提出了DARP(双臂机器人感知,Dual-Arm Robotic Perception)https://doi.org/10.21227/rmv3-be47,这是一个经过标定的双臂RGB-D-IR数据集,用于以物体为中心的机器人感知。该系统采用两台独立运动的手眼式(eye-in-hand)机械臂,分别置于共享桌面工作空间的两侧。每条机械臂搭载一个Intel RealSense传感器,持续记录RGB、深度和双目红外数据,同时同步记录机器人关节状态以用于位姿恢复。物体的摆放不固定位姿也无标记位置,采集流程执行自动定位、跨臂确认、自适应视点生成以及连续多模态记录。DARP包含十种独特的桌面物体,并保留了原始传感器记录、机器人状态日志、物体级元数据以及标定信息,这些是在共享度量坐标系中重建相机轨迹所必需的。为评估采集的几何一致性,我们实现了一个确定性的多视角融合流水线,将标定后的RGB-D观测转换为互补的局部点云和实测表面网格,且不使用任何学习型或生成式的补全方法。在224个留出的RGB-D关键帧(包含1,563,466个三维查询点)上的评估结果显示,点到网格距离的中位数为2.13毫米,均方根误差(RMSE)为4.04毫米,其中96.56%的点位于实测表面网格10毫米范围内。DARP旨在作为多视角重建、协作式机器人感知、多模态融合、主动感知以及未来基于学习的部分物体观测推理的可复用资源。
cs.RO / 68 / 2608.31167

SUN: Persistent Programs For Language-Grounded Control-to-Learning-to-Real Policies

SUN:面向语言接地的控制—学习—现实策略的持久化程序
Wang, Weiqi, Li, Zhi, Lei, Yudong, Martinez, David, Gao, Xiaofeng, Jiang, Yuxin, Jiang, Chenfanfu, Wu, Yingnian, Terzopoulos, Demetri, Gong, Ran
Abstract
Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: control executes specified objectives, learning amortizes that behavior into a reactive policy, yet existing protocols discard task semantics, leaving rewards hand-crafted and behavior drifting from what control verified.We introduce Semantically UNified (SUN) Programs, typed executables where geometric and contact relations are defined once and compiled into aligned Model Predictive Control (MPC) costs, satisfaction predicates, RL rewards, transition guards, and diagnostics. Our system, Kuafu, driven by large vision language systems, automatically synthesizes SUN Programs from language and scene semantics, screens feasibility via MPC, and retains semantics while training stage-conditioned policies. Across nine tasks, Kuafu achieves 82.03% macro-success, outperforming sparse-reward (35.67%) and Stage-BC (24.75%) baselines. At 8192-way scale, it generates 10.57x the successful trajectory time per hour of human teleoperation. With 500 trajectories per task, Kuafu data trains DP3 policies to 46.0% simulation success (vs. 22.4% for alternatives) and 34.7% on physical Franka and Kinova robots. These results establish that simulation-screened task semantics can effectively amortize control into robust policies, without demonstrations or manual dense rewards, unifying symbolic planning and data-driven execution.
Chinese Translation
在长时程操作任务中,基于模型的控制与学习策略之间存在一个长期未被正视的分歧:控制负责执行明确规定的目标,学习则将该行为 amortize(固化)为反应式策略,然而现有方法丢弃了任务语义,导致奖励需要手工设计,且学习后的行为会偏离控制所验证的内容。我们提出了语义统一程序(Semantically UNified Programs,SUN),这是一种带类型的可执行程序,其中几何与接触关系只需定义一次,即可编译为对齐的模型预测控制(MPC)代价函数、满足性谓词、强化学习奖励、转移守卫条件以及诊断工具。我们的系统 Kuafu 由大型视觉语言系统驱动,能够从语言和场景语义中自动合成 SUN 程序,通过 MPC 筛选可行性,并在训练阶段条件化策略的同时保留任务语义。在九项任务中,Kuafu 实现了 82.03% 的宏平均成功率,优于稀疏奖励(35.67%)和 Stage-BC(24.75%)基线方法。在 8192 路并行规模下,其每小时生成的成功轨迹时间是人工遥操作的 10.57 倍。每项任务仅需 500 条轨迹,Kuafu 数据即可将 DP3 策略训练至 46.0% 的仿真成功率(相比之下其他数据来源仅为 22.4%),并在实体 Franka 和 Kinova 机器人上达到 34.7% 的成功率。这些结果表明,经仿真筛选的任务语义能够有效地将控制固化为鲁棒策略,无需演示或手工密集奖励,从而统一了符号规划与数据驱动的执行。