← Back to Index
Daily Research Digest

arXiv Papers

2026-09-04
290
Papers
4
Categories
290
Translated
收藏清单 0
人工智能 (Artificial Intelligence)
69
cs.AI / 1 / 2609.02981

Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence

人工智能驱动的新型实用英语教材的结构与实现
Wang, Ya, Zhang, Lei, Yang, Xueguang, Chen, Bo
Abstract
Artificial intelligence is changing the form of applied English materials from fixed paper sequences to adaptive learning systems that can diagnose learners, recommend tasks, and provide formative feedback. This paper studies the structure and application of a new practical English textbook driven by artificial intelligence. A five-layer architecture is proposed: knowledge mapping, learner profiling, task generation, feedback orchestration, and teacher-side governance. A prototype was tested on 186 non-English-major undergraduates for eight weeks of teaching. Compared with a static digital textbook, the proposed system increased the unit completion accuracy from 72.4% to 84.9%, raised the average score for speaking tasks by 10.8 points, and reduced the teacher's correction time by 31.6%. Therefore, an AI-driven textbook can maintain the stability of the curriculum while providing personalised learning paths, rich practice materials and traceable classroom data.
Chinese Translation
人工智能正在改变应用型英语教材的形态,使其从固定的纸质序列转变为能够诊断学习者、推荐任务并提供形成性反馈的自适应学习系统。本文研究了人工智能驱动的新型实用英语教材的结构与应用。我们提出了一种五层架构:知识图谱、学习者画像、任务生成、反馈编排以及教师端治理。该原型系统在186名非英语专业本科生中进行了为期八周的教学测试。与静态数字教材相比,所提出的系统将单元完成准确率从72.4%提升至84.9%,口语任务平均成绩提高了10.8分,教师批改时间减少了31.6%。因此,人工智能驱动的教材能够在保持课程稳定性的同时,提供个性化学习路径、丰富的练习材料以及可追溯的课堂数据。
cs.AI / 2 / 2609.03209

MasterControl Seventeen Every Time

MasterControl每次都能做到十七次全对
Lab, MasterControl AI
Abstract
We study a governed approach to enterprise analytics: a language model interprets the question, while deterministic policy selects and runs a pre-approved analytical program that returns both results and evidence. We show that this restriction can remain expressive within a defined analytical class, using relational operations plus aggregation, comparison, windows, ranking, and similarity. Fixed meaning, policy, data, and execution rules also make results replayable. Across 440 runs, three 8B models generated SQL and selected tools at runtime, while Qwen3-8B interpreted intent only and policy executed the approved program. None of 330 runtime-planning episodes matched the full answer-and-evidence contract across all test datasets; the policy-executed analyzer matched 110 of 110. This is a configuration-specific result, not evidence that runtime agents cannot succeed under other designs.
Chinese Translation
我们研究了一种受治理的企业数据分析方法:由语言模型解释问题,而确定性策略选择并执行预先批准的分析程序,该程序同时返回结果和证据。我们表明,这种限制在特定的分析类别内仍可保持表达能力,该类别使用关系运算以及聚合、比较、窗口、排序和相似度计算。固定的语义、策略、数据和执行规则也使结果可重放。在440次运行中,三个8B模型在运行时生成SQL并选择工具,而Qwen3-8B仅解释意图,由策略执行已批准的程序。在所有测试数据集上,330次运行时规划尝试中没有一次能同时满足完整的答案与证据契约;而策略执行的分析器在110次中全部匹配(110/110)。这是一个特定配置下的结果,并非证明运行时智能体在其他设计下无法成功。
cs.AI / 3 / 2609.03236

Speculative Macro Commit for Faster Tool-Using Agents

面向更快工具使用型智能体的推测性宏提交方法
Liu, Zeyu, Kundu, Souvik, Beerel, Peter A.
Abstract
Tool-using LLM agents spend wall-clock time not only on model inference but also in serial action--observation turns, where each tool call, environment transition, and observation can delay subsequent decisions. We introduce \textbf{Speculative Macro Commit} (SMC), a runtime mechanism for a two-tier agent system: a large authoritative actor model produces the official trajectory, while a faster speculative drafter model continuously predicts and executes future action chains on an isolated environment snapshot. SMC mines recurring multi-action skeletons from training traces and stores them in a macro library used to match against action chains predicted by the drafter at runtime. When the actor's next tool call matches the first drafted action, SMC commits the remaining pre-executed draft steps, together with their observations, to the official trajectory. Using Qwen3.5-27B INT4 as the authoritative actor model and Qwen3.5-4B as the speculative drafter model, SMC matches the sequential agent's overall accuracy while reducing latency by 10.23\% over the Speculative Actions (SA) baseline and 18.59\% over sequential execution on the $\tau^2$-Bench Telecom subset. On AppWorld, SMC reduces wall time by 7.7\% over SA baseline and 44.9\% over sequential execution, with a small reduction in task completion. Overall, SMC provides a practical way to reuse multi-step speculative execution and reduce agent latency beyond single-step speculative actions. Our code is publicly available \href{https://github.com/zeyuliu1037/speculative-macro-commit}{\textcolor{magenta}{here}}.
Chinese Translation
使用工具的大语言模型(LLM)智能体不仅耗费在模型推理上的时间,还耗费在串行的动作-观察轮次中,其中每次工具调用、环境转移和观察都可能延迟后续决策。我们提出了推测性宏提交(Speculative Macro Commit, SMC),这是一种面向双层智能体系统的运行时机制:大型权威执行者模型生成官方轨迹,而更快的推测性草稿模型则在隔离的环境快照上持续预测并执行未来的动作链。SMC 从训练轨迹中挖掘重复出现的多动作骨架,并将其存储在宏库中,用于在运行时与草稿模型预测的动作链进行匹配。当执行者模型的下一个工具调用与草稿的第一个动作匹配时,SMC 将其余已预执行的草稿步骤及其观察结果提交到官方轨迹中。使用 Qwen3.5-27B INT4 作为权威执行者模型、Qwen3.5-4B 作为推测性草稿模型,SMC 在保持与串行智能体总体准确率相当的同时,在 τ²-Bench 电信子集上将延迟较推测性动作(Speculative Actions, SA)基线降低了 10.23%,较串行执行降低了 18.59%。在 AppWorld 上,SMC 将总耗时较 SA 基线降低 7.7%,较串行执行降低 44.9%,同时任务完成率仅有小幅下降。总体而言,SMC 提供了一种实用方法来复用多步推测执行,从而在单步推测动作的基础上进一步降低智能体延迟。我们的代码已公开,参见 https://github.com/zeyuliu1037/speculative-macro-commit。
cs.AI / 4 / 2609.03340

Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

新鲜记忆,过时计划:面向分布式LLM智能体记忆的依赖范围验证机制
Chen, Evan, Wang, Shiqiang, Brinton, Christopher G.
Abstract
Distributed LLM-agent teams can read the latest shared facts and still act on an obsolete plan. A planner may derive an action from requirement $r_3$, another agent may commit $r_4$, and an executor may receive $r_4$ without replacing the plan derived from $r_3$. We call this \emph{stale-plan execution}: state freshness does not establish that the plan authorizing an action remains valid. We introduce PlanFence, a dependency-scoped action-validation protocol. Plans cite the exact public records they used, and an executor validates only the records that can affect the pending external action, replanning once or blocking when validation is incomplete. In 30 controlled live workflows with a post-plan revision, a freshness-only executor acts on the obsolete plan in every task, whereas PlanFence completes all tasks without an invalid action. Controlled replay reveals two conditional boundaries: proactive synchronization yields lower coordination stall at low churn, while PlanFence avoids repeated update-path coordination as churn grows and avoids validating unrelated state as the shared keyspace grows. These are controlled safety and systems-cost results, not general task-accuracy gains.
Chinese Translation
分布式LLM智能体团队即使能够读取最新的共享事实,仍可能依据过时的计划采取行动。规划器可能基于需求 $r_3$ 推导出某个动作,另一个智能体可能提交了 $r_4$,而执行器可能接收到 $r_4$,却未替换基于 $r_3$ 推导的计划。我们将这种现象称为\emph{过时计划执行}(stale-plan execution):状态的新鲜性并不能保证授权该动作的计划仍然有效。我们提出了 PlanFence,一种依赖范围的动作验证协议。计划需引用其使用的确切公共记录,而执行器仅验证可能影响待执行外部动作的记录,在验证未完成时进行一次重规划或予以阻塞。在30个包含计划后修订的受控真实工作流中,仅依赖新鲜性检查的执行器在每项任务中都会依据过时计划行动,而 PlanFence 则完成了全部任务且未执行任何无效动作。受控回放实验揭示了两个条件边界:在低变更率下,主动同步可带来更低的协调停顿;而随着变更率上升,PlanFence 可避免重复的更新路径协调,并随着共享键空间扩大而避免验证无关状态。这些是受控环境下的安全性与系统开销结果,而非普遍的任务准确性提升。
cs.AI / 5 / 2609.03402

A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant

一种基于提示工程的方法:在通用人工智能助教中实现可扩展、灵活且实时的混合微观层面个性化
Basu, Saptarshi, Kakar, Sandeep, Goel, Ashok
Abstract
Artificial intelligence (AI) teaching assistants powered by large language models (LLMs) offer scalable educational support but often provide limited personalization. This study presents a prompt-engineering-based framework for personalizing general-purpose LLM/RAG-based AI teaching assistants such as Jill Watson across academic disciplines and courses. The framework adapts responses using six learner-specific dimensions: self-assessment, abstraction preference, verbosity preference, perceptual orientation, information processing style, and level of understanding, yielding 96 distinct learner profiles. Student queries are additionally analyzed using Bloom's Taxonomy to estimate cognitive complexity at the interaction level. Learner attributes and cognitive assessments are encoded in structured prompts that condition the LLM without requiring model retraining. The framework is evaluated through experiments using NLP metrics and a human study with five participants. Results show perceived differences in response style and structure across personalization conditions, with statistical analyses identifying learner attributes associated with measurable response changes. These findings provide preliminary evidence that prompt-based personalization can support adaptive behavior in LLM-powered educational agents.
Chinese Translation
基于大语言模型(LLM)的人工智能(AI)助教能够提供可扩展的教育支持,但个性化程度通常有限。本研究提出了一种基于提示工程的框架,用于对通用LLM/RAG架构的AI助教(如Jill Watson)进行跨学科、跨课程的个性化。该框架利用六个学习者特定维度来调整响应:自我评估、抽象偏好、详略偏好、感知取向、信息处理风格和理解水平,从而产生96种不同的学习者画像。此外,还运用布鲁姆分类法(Bloom's Taxonomy)分析学生提问,以在交互层面估算认知复杂度。学习者属性与认知评估被编码到结构化提示中,用以约束LLM的行为,而无需重新训练模型。该框架通过基于NLP指标的实验和一项五名参与者的人类评估进行了评估。结果表明,在不同个性化条件下,响应的风格与结构存在可感知的差异;统计分析识别出与可测量响应变化相关的学习者属性。这些发现为提示式个性化能够支持LLM教育智能体的自适应行为提供了初步证据。
cs.AI / 6 / 2609.03407

Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

陷入故事之中:多轮大语言模型对话中的叙事俘获
Wu, Yuhe, Wang, Guangyu, Chen, Yujie, Zhang, Jiatong, Chen, Yuran, Zhang, Yutong, Cheng, Xiyin, Cao, Wenpeng, Liu, Zhuang, Zhang, Guang
Abstract
People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this context through single-turn judgments or pressure-laden rebuttals, assumptions that poorly match how guidance is sought in real-world contexts. These assumptions leave unclear whether narration alone, without an explicit opposing position, can shift model judgments during multi-turn moral consultation. Yet real-world moral-conflict conversation often elicits one party's self-justifying account, which can unfold over multiple turns and create information asymmetry. We introduce \textbf{narrative captivity}, a failure mode in which a model treats an unopposed one-sided account as complete and aligns with the narrator's interpretation without seeking missing perspectives. To measure this phenomenon, we build a benchmark of $5{,}078$ interpersonal-conflict scenarios spanning six moral dimensions. Across 17 LLMs, narrative captivity is widespread: end-state judgments under multi-turn narration shift by 25 percentage points on average beyond the matched single-turn baseline. Stage-level analysis identifies preference optimization as a major contributor, while four inference-time strategies provide only partial mitigation. We hope our project fosters LLM advisors that preserve independent judgment in real-world consultation.
Chinese Translation
人们越来越多地向大语言模型(LLM)寻求日常建议,这使得充满伦理冲突的人际问题成为了一个现实的道德咨询场景。以往大多数工作通过单轮判断或带有压力的反驳来研究这一场景,这些假设与现实中寻求指导的方式并不相符。这些假设留下了一个未解的问题:在多轮道德咨询中,仅凭叙事——没有明确的对立立场——能否改变模型的判断。然而,现实中的道德冲突对话往往会引出单方面的自我辩解陈述,这种陈述可能在多轮对话中展开,并造成信息不对称。我们提出了“叙事俘获”(narrative captivity)这一概念,指一种失败模式:模型将未经反驳的单方面陈述视为完整信息,并顺着叙述者的解释行事,而不去寻求缺失的其他视角。为了测量这一现象,我们构建了一个包含5,078个人际冲突场景的基准,涵盖六个道德维度。在17个大语言模型上,叙事俘获现象普遍存在:在多轮叙事下,模型对最终状态的判断相比匹配的单轮基线平均偏移25个百分点。阶段级分析表明,偏好优化是主要成因,而四种推理时策略仅能提供部分缓解。我们希望本项目能够促进在现实咨询中保持独立判断的大语言模型顾问的发展。
cs.AI / 7 / 2609.03416

Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection

Dude:一种用于论文-代码不一致性检测的双检测多智能体系统
Liu, Weijie, Zhao, Running, Yuan, Wenhao, Xu, Jinfeng, Xu, Zhanfeng, Zhang, Xiaoxi, Ngai, Edith Cheuk-Han
Abstract
LLM-empowered paper-code discrepancy detection has received growing concern since the scaling of research submissions exceeds the manual review capability. However, the limited context capacity and one-sided discrepancy detection of existing single-agent LLM paradigms lead to an inferior recall performance in detecting discrepancies. In this paper, we propose Dude, the first Dual-Detection Multi-Agent System for paper-code discrepancy detection. We discover that the granularity asymmetry of the paper-language and code-language introduces over-interpretation and over-reporting challenges in a multi-agent system design for discrepancy detection, resulting in increasing false positives. To address this, we propose a granularity-aligned negotiation and a two-stage salience-filtering mechanism in Dude, which effectively prevents agents from falsely reporting discrepancies. Experimental results in real-world paper-code discrepancy datasets showcase Dude's significant recall and precision improvement by up to 22.8%, increasing F1 score by up to 18.7% compared to baseline methods.
Chinese Translation
随着科研论文投稿数量的增长超出人工审查能力,基于大语言模型(LLM)的论文-代码不一致性检测日益受到关注。然而,现有单智能体LLM范式受限于上下文容量且检测视角单一,导致检测不一致性时的召回率表现不佳。在本文中,我们提出了 Dude,这是首个用于论文-代码不一致性检测的双检测(Dual-Detection)多智能体系统。我们发现,论文语言与代码语言之间的粒度不对称性给多智能体不一致性检测系统设计带来了过度解读和过度报告的挑战,从而导致假阳性增多。为解决这一问题,我们在 Dude 中提出了粒度对齐协商机制和两阶段显著性过滤机制,有效防止智能体错误地报告不一致性。在真实世界的论文-代码不一致性数据集上的实验结果表明,与基线方法相比,Dude 的召回率和精确率显著提升,最高提升 22.8%,F1 分数最高提升 18.7%。
cs.AI / 8 / 2609.03423

DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

DuplexSpeechBench-IFEval:评估全双工语音代理中的隐式指令遵循能力
Mathur, Puneet, Manocha, Dinesh
Abstract
Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely test these behaviors through explicit turn-management instructions, while deployed agents are often configured through roles or personas from which the appropriate conversational behavior must be inferred. We introduce DuplexSpeechBench-IFEval (DSB-IFEval) for evaluating implicit instruction-following in real-time spoken interaction. (DSB-IFEval) comprises 1,038 test cases spanning eight diverse assistant roles and evaluates five conditioning protocols for instruction-following: default behavior, explicit behavioral instructions, persona-implied behavior, combined persona--rule conditioning, and instruction conflict. We measure real-time floor management using a deterministic Instruction Adherence Score (IAS) and persona-consistent content using LLM-judged Persona Adherence Score (PAS). Across six real-time speech systems, we find architecture-dependent trade-offs. Full duplex models like F-Actor and PersonaPlex are more sensitive to whether conversational behavior is stated explicitly or must be inferred from a persona, with adherence dropping by 9.7% and 4.5%, respectively, under persona-only conditioning. In contrast, GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat strongly adhere to persona-consistent content, but their floor behavior does not adapt across explicit and persona-only instructions and remains constrained on several proactive actions. We further find that even if systems reliably follow conflicting directives to their prescribed persona, they still struggle to override them under safety conflict. These results show that inferring the behavior implied by a role, executing it at the appropriate conversational moment, and resolving competing instructions remain distinct challenges for full-duplex voice agents.
Chinese Translation
全双工语音代理必须持续决策何时聆听、附和、打断、处理语音重叠、获取话轮以及让出话轮。现有基准测试主要通过显式的话轮管理指令来测试这些行为,而实际部署的代理通常通过角色或人设进行配置,相应的对话行为需要从中推断。我们提出了DuplexSpeechBench-IFEval(DSB-IFEval),用于评估实时语音交互中的隐式指令遵循能力。DSB-IFEval包含1,038个测试用例,涵盖八个不同的助手角色,并评估五种指令遵循的条件设定协议:默认行为、显式行为指令、人设隐含行为、人设-规则组合条件设定以及指令冲突。我们使用确定性的指令遵循得分(Instruction Adherence Score, IAS)衡量实时话轮管理能力,使用LLM评判的人设一致性得分(Persona Adherence Score, PAS)衡量与角色一致的内容表现。在六个实时语音系统上的实验中,我们发现了依赖于架构的性能权衡。像F-Actor和PersonaPlex这样的全双工模型对对话行为是显式声明还是需要从人设中推断更为敏感,在仅有人设条件设定下,遵循度分别下降9.7%和4.5%。相比之下,GPT-Realtime、MiniCPM-o和Fun-Audio-Chat在内容上高度遵循角色设定,但其话轮管理行为在显式指令和仅有人设指令之间未能自适应调整,且在若干主动行为上仍受限制。我们进一步发现,即使系统能够可靠地遵循与预设人设相冲突的指令,在安全冲突场景下它们仍然难以覆盖这些人设。这些结果表明,推断角色所隐含的行为、在恰当的对话时机执行该行为,以及解决相互竞争的指令,仍然是全双工语音代理面临的独立挑战。
cs.AI / 9 / 2609.03438

Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

GUI智能体是否知道何时不应行动?为多模态GUI智能体启用冲突感知终止机制
Huang, Zhaoyuan, Ju, Tianjie, Cheng, Pengzhou, Wu, Zheng, Li, Yansi, Song, Chuanbiao, Lan, Jun, Zhu, Huijia, Wang, Weiqiang, Zhang, Zhuosheng
Abstract
Graphical user interface (GUI) agents are increasingly used to execute natural-language instructions on user interfaces, yet real users may issue infeasible instructions due to benign mistakes. A reliable agent should not only know how to act, but also when not to act. In this work, we introduce CONFLICTGUI, a benchmark covering instruction-internal conflicts and instruction-GUI context conflicts to study conflict-aware termination. Our evaluation reveals severe execution-biased overcompliance: agents that perform well on feasible tasks often continue to execute blindly under conflicting instructions. To mitigate this behavior, we propose CONFLICTGUARD, an inference-time framework that aligns an agent's feasibility awareness with its action generation. CONFLICTGUARD contains two coupled components: a feasibility verification protocol that guides the agent to assess instruction logic and GUI-side evidence before acting, and a conditional action modulation mechanism that steers agents from over-compliant execution into termination-oriented behavior. Experiments across five widely-used agents demonstrate that CONFLICTGUARD improves average conflict task success rate significantly, while preserving normal GUI-task performance. These results validate that a lightweight inference-time intervention can substantially boost GUI Agent's competence to identify inappropriate execution scenarios and refrain from unnecessary actions.
Chinese Translation
图形用户界面(GUI)智能体被越来越多地用于在用户界面上执行自然语言指令,然而真实用户可能因无心之误而发出不可行的指令。一个可靠的智能体不仅应知道如何行动,还应知道何时不应行动。在本工作中,我们提出了CONFLICTGUI,一个涵盖指令内部冲突与指令-GUI上下文冲突的基准,用于研究冲突感知终止。我们的评估揭示了一种严重的执行偏向性过度顺从现象:在可行任务上表现良好的智能体,在遇到冲突指令时往往仍盲目继续执行。为缓解这一行为,我们提出了CONFLICTGUARD,一个推理时框架,用于将智能体的可行性意识与其动作生成过程对齐。CONFLICTGUARD包含两个相互耦合的组件:一是可行性验证协议,引导智能体在行动之前评估指令逻辑与GUI侧证据;二是条件性动作调制机制,将智能体从过度顺从的执行引导至面向终止的行为。在五个广泛使用的智能体上的实验表明,CONFLICTGUARD显著提升了冲突任务的平均成功率,同时保持了正常的GUI任务性能。这些结果验证了轻量级的推理时干预能够大幅提升GUI智能体识别不当执行场景并避免不必要行动的能力。
cs.AI / 10 / 2609.03460

Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty

超越“由AI生成”标签:可视化溯源密度以缓解透明度惩罚
Zhang, Qing, Huang, Yifei, Lee, Juyoung, Starner, Thad, Rekimoto, Jun
Abstract
As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust fluent hallucinations while also discounting accurate content once it is disclosed as AI-generated. Binary ``Made with AI'' labels respond with authorship disclosure, but they do not show what supports a claim. We propose Provenance Density, an evidence-visualization interface that shows the density of verified claims in a text. In a user study with 81 participants, an idealized Provenance Density interface produced a large discernment gap between truth and fabrication ($+4.15$ points, $d=1.82$), whereas participants given no signal showed no detectable discrimination. A technical audit with 200 samples shows that retrieval density alone is insufficient; unexpectedly, the Consistency Veto carries most of the discriminative signal on dynamic queries. As AI-generated content becomes indistinguishable from human writing, effective transparency must move from authorship disclosure toward evidence visualization.
Chinese Translation
随着生成式AI使流畅的文本变得易于批量生产,用户不再能够将流畅性作为判断真实性的依据。我们将这种失效模式称为“流畅性陷阱”(Fluency Trap):用户既轻信流畅的幻觉内容,又在准确内容被披露为AI生成后对其打折扣。二元的“由AI生成”(Made with AI)标签以披露作者身份作为回应,但并未展示支撑论断的证据。我们提出了“溯源密度”(Provenance Density),一种展示文本中经验证论断密度的证据可视化界面。在一项有81名参与者参与的用户研究中,理想化的溯源密度界面在真实内容与虚构内容之间产生了显著的辨别差距(+4.15分,d=1.82),而未获得任何信号的参与者则未表现出可检测的辨别能力。一项基于200个样本的技术审计表明,仅依靠检索密度是不够的;出人意料的是,“一致性否决”(Consistency Veto)在动态查询中承载了大部分的判别信号。随着AI生成内容与人类写作变得难以区分,有效的透明度必须从作者身份披露转向证据可视化。
cs.AI / 11 / 2609.03478

AutoGraphForge: Towards Automated Graph Theory Discovery

AutoGraphForge:迈向自动化的图论发现
Pastorek, Ján
Abstract
We report on our ongoing project to develop a computational pipeline, AutoGraphForge, for an automated graph-theoretic conjecturing-refuting-formalizing-proving system. Conjecture generation is counterexample-guided and runs in rounds: a Graffiti3 generator proposes conjectures over a small, evolving snapshot table $T$ (initially a few hundred graphs with their computed invariants) that grows only by counterexamples to its own conjectures. A novelty filter of $559$ classical and folklore relations, closed under transitive composition and linear identity substitution, decides via a linear program whether a candidate is already implied by known results. Surviving candidates are tested against a dataset of about $348,000$ graphs, unioning the complete House of Graphs invariant export, the exhaustive census of all connected graphs on at most nine vertices, several extremal families (strongly regular, minimal Ramsey, Cayley, cages, barbells, lollipops, spiders), and random models. Counterexample-search algorithms then attack the remainder. Run for several rounds on an HPC cluster, the loop yields $6,522$ conjectures that survived the refutation dataset, the novelty filter and every active-search run -- among them nontrivial relations between the annihilation number and the edge-cover number for bipartite and regular graphs, which we prove by hand. A subsequent formalization and proving stage deterministically translates each surviving conjecture into a Lean 4 statement skeleton; every candidate proof is kernel-verified against a pinned mathlib4 and our custom invariant preamble. This stage integrates two neural provers -- DeepSeek-Prover-V2-671B (served with vLLM) and the Lean-specialised OProver-32B -- behind the independent kernel check. It is implemented end-to-end and passes initial sanity checks, with the full pipeline currently running on the cluster.
Chinese Translation
我们报告一个正在进行的项目,旨在开发一套名为 AutoGraphForge 的计算流水线,用于实现自动化的图论猜想—反驳—形式化—证明系统。猜想生成采用反例引导的方式并按轮次进行:Graffiti3 生成器在一个小型且不断演化的快照表 $T$(初始为几百个图及其计算所得的不变量)上提出猜想,该表仅通过其自身猜想的反例而增长。一个由 $559$ 条经典及民间(folklore)关系构成的新颖性过滤器——在传递复合与线性恒等式替换下封闭——通过线性规划判断候选猜想是否已被已知结果所蕴含。幸存的候选猜想将在约 $348,000$ 个图的数据集上进行测试,该数据集合并了 House of Graphs 的完整不变量导出数据、所有至多九个顶点的连通图的穷举普查、若干极值图族(强正则图、极小 Ramsey 图、Cayley 图、笼图、杠铃图、棒棒糖图、蜘蛛图)以及随机模型。随后,反例搜索算法对剩余的猜想发起攻击。在高性能计算(HPC)集群上运行若干轮后,该循环产生了 $6,522$ 个经受住反驳数据集、新颖性过滤器以及所有主动搜索运行考验的猜想——其中包括二部图和正则图中消灭数(annihilation number)与边覆盖数之间的非平凡关系,我们已用手工方法加以证明。随后的形式化与证明阶段将每个幸存的猜想确定性地翻译为 Lean 4 语句骨架;每个候选证明都针对固定版本的 mathlib4 及我们自定义的不变量前导文件(preamble)进行内核验证。该阶段集成了两个神经证明器——DeepSeek-Prover-V2-671B(通过 vLLM 部署)和专门面向 Lean 的 OProver-32B——并置于独立的内核检查之后。整个系统已实现端到端运行并通过了初步的健全性检查,完整流水线目前正在集群上运行。
cs.AI / 12 / 2609.03493

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

让每一次工具调用都有价值:面向智能体视觉语言模型的必要工具-证据路径奖励
Long, Xingming, Liu, Yu, Yang, Zhiwei, Feng, Hanqi, Zhang, Shaojie, Poczos, Barnabas, Jiang, Chao, Luo, Zhenbo, Jiang, Lei, Fu, Pei
Abstract
Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or external knowledge. To acquire this missing evidence, agentic VLMs invoke tools such as image cropping, image search, and text search. However, existing training paradigms primarily evaluate tool-use based on final answer correctness, leaving evidence acquisition and utilization insufficiently supervised. This leads to two critical shortcomings: (i) models frequently issue redundant or off-target tool calls that fail to gather necessary evidence, and (ii) even when appropriate tools are called, models often fail to extract the necessary information from the resulting observations. To address these limitations, we introduce the NTEP (Necessary Tool-Evidence Path), a novel annotation scheme that explicitly specifies the essential external evidence and corresponding tool calls for each query. Building upon this, we propose NTEP-R (NTEP Reward), a supervision mechanism ensuring that each tool invocation strictly advances the reasoning process toward the final solution. Specifically, our approach rewards the agent for aligning its pre-call intent with a necessary evidence-seeking goal, and for ensuring the information summarized from the post-call observation aligns with the necessary evidence. Furthermore, we introduce a non-repeated-goal regularizer to penalize redundant calls that revisit satisfied NTEP goals. Extensive evaluations on seven image-grounded benchmarks demonstrate that our 8B-parameter instantiation, NTEP-8B, significantly improves both search-oriented accuracy and tool-use efficiency within a unified three-tool framework. These results highlight the critical value of fine-grained tool-evidence path supervision for training robust agentic VLMs.
Chinese Translation
现代视觉语言模型(VLM)能够直接回答许多基于图像的问题,但在需要细粒度视觉细节或外部知识的复杂查询方面往往表现不佳。为获取这些缺失的证据,智能体视觉语言模型会调用图像裁剪、图像搜索和文本搜索等工具。然而,现有的训练范式主要依据最终答案的正确性来评估工具使用,导致对证据获取和利用的监督不足。这带来两个关键缺陷:(i)模型经常发出冗余或偏离目标的工具调用,未能收集到必要证据;(ii)即使调用了合适的工具,模型也常常无法从返回的观测结果中提取必要信息。为解决这些局限,我们提出了NTEP(Necessary Tool-Evidence Path,必要工具-证据路径),这是一种新颖的标注方案,可为每个查询显式指定所需的外部证据及相应的工具调用。在此基础上,我们提出NTEP-R(NTEP Reward),一种监督机制,确保每次工具调用都能严格推动推理过程朝最终解答迈进。具体而言,我们的方法在智能体将调用前意图与必要的证据寻求目标保持一致时给予奖励,并确保从调用后观测中总结的信息与必要证据保持一致。此外,我们引入了非重复目标正则化项,以惩罚重复已完成NTEP目标的冗余调用。在七个基于图像的基准上的广泛评估表明,我们80亿参数规模的模型NTEP-8B在统一的三工具框架内显著提升了面向搜索的准确性和工具使用效率。这些结果凸显了细粒度工具-证据路径监督对于训练鲁棒智能体视觉语言模型的关键价值。
cs.AI / 13 / 2609.03494

GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving

GrowPage:面向高效LLM推理服务的按需KV预算分配
Ma, Qiankun, Zhou, Yanjiang, Xiong, Zinan, Wang, Haofei, Song, Zhen, Xiang, Yang, Zhang, Ziyao, Zheng, Hairong
Abstract
Long-output reasoning has made the key--value (KV) cache a critical memory bottleneck for efficient LLM serving. Existing KV compression methods usually rely on a predefined per-request budget and adjust only which KV states are retained, leaving the total capacity fixed throughout decoding. However, reasoning workloads exhibit substantial demand variation: different requests require different KV capacities, and the attention demand of an individual request evolves during generation. We introduce \textbf{GrowPage}, an on-demand KV budgeting framework that treats KV capacity as a runtime resource. GrowPage maintains lightweight dual-timescale query summaries to capture recent and long-term attention behaviors, and uses their relative attention working sets to estimate demand evolution. At each capacity boundary, GrowPage either compresses KV states within the current allocation or acquires an additional physical page when broader demand emerges. By integrating with PagedAttention's page-level memory abstraction, GrowPage preserves continuous batching and prefix caching. Experiments on reasoning benchmarks across multiple models show that GrowPage achieves a superior performance--throughput trade-off over existing approaches.
Chinese Translation
长输出推理使键值(KV)缓存成为高效LLM服务的关键内存瓶颈。现有的KV压缩方法通常依赖预定义的单请求预算,仅调整保留哪些KV状态,而总容量在整个解码过程中保持固定。然而,推理工作负载呈现出显著的需求变化:不同请求需要不同的KV容量,且单个请求的注意力需求在生成过程中不断演变。我们提出GrowPage,一个将KV容量视为运行时资源的按需KV预算分配框架。GrowPage维护轻量级的双时间尺度查询摘要,以捕捉近期和长期的注意力行为,并利用它们的相对注意力工作集来估计需求演变。在每个容量边界处,GrowPage要么在当前分配内压缩KV状态,要么在出现更广泛需求时获取额外的物理页。通过与PagedAttention的页级内存抽象集成,GrowPage保留了连续批处理和前缀缓存功能。在多个模型的推理基准上的实验表明,GrowPage在性能与吞吐量的权衡上优于现有方法。
cs.AI / 14 / 2609.03503

PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing

PPO-STGNN:一种结合近端策略优化与时空图神经网络的云边端计算DAG任务调度方法
Qi, Yangshuo, Wang, Chenwei, Shen, Zihan, Sun, Songlin
Abstract
With the rapid development of the Internet of Things, computation intensive directed acyclic graph (DAG) tasks have become increasingly common in cloud-edge-end collaborative environments. However, cloud, edge, and end nodes are highly heterogeneous in computing capacity, network bandwidth, and energy consumption, which makes the efficient scheduling of tasks with complex dependencies an NP-hard problem. Traditional heuristic algorithms and conventional reinforcement-learning methods often fail to capture the spatio-temporal dynamics of system resources. This paper proposes PPO-STGNN, a DAG task-scheduling algorithm that integrates proximal policy optimization (PPO) with spatio-temporal graph neural networks (STGNNs). The method uses an STGNN to extract features from both the DAG task topology and the physical cloud-edge-end resource graph, and then optimizes the scheduling policy through PPO to minimize makespan and schedule length ratio (SLR) while improving CPU and memory load balancing. To accelerate convergence, a multi-teacher behavior-cloning mechanism is introduced for pretraining. Experimental results show that PPO-STGNN significantly improves load balancing while maintaining a low completion time, making it suitable for dynamic and heterogeneous cloud-edge- end DAG scheduling scenarios.
Chinese Translation
随着物联网的快速发展,计算密集型的有向无环图(DAG)任务在云边端协同环境中日益普遍。然而,云、边、端节点在计算能力、网络带宽和能耗方面高度异构,使得具有复杂依赖关系的任务高效调度成为一个NP难问题。传统启发式算法和常规强化学习方法往往难以捕捉系统资源的时空动态特性。本文提出PPO-STGNN,一种将近端策略优化(Proximal Policy Optimization, PPO)与时空图神经网络(Spatio-Temporal Graph Neural Network, STGNN)相融合的DAG任务调度算法。该方法利用STGNN从DAG任务拓扑和云边端物理资源图中提取特征,再通过PPO优化调度策略,以最小化完成时间(makespan)和调度长度比(SLR),同时提高CPU和内存的负载均衡性。为加速收敛,本文引入多教师行为克隆机制进行预训练。实验结果表明,PPO-STGNN在保持较低完成时间的同时显著改善了负载均衡,适用于动态异构的云边端DAG调度场景。
cs.AI / 15 / 2609.03515

What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation

什么因素对于激进的解码时KV缓存淘汰至关重要?时间聚合与排序保持
Zeng, Bo, Zhao, Yu, Liu, Yefeng, Lu, Zhihong, Ni, Xuanfan, Wang, Xintong
Abstract
Decoding-time KV cache compression research focuses heavily on designing better token scoring functions, while the temporal rule that aggregates scores across decode steps is often treated as an implementation detail. Under aggressive KV compression, we find that exponential-moving-average (EMA) aggregation makes approximately order-preserving scorer modifications largely indistinguishable at the eviction-set level. Value-norm and entropy variants remain highly correlated with attention and produce nearly unchanged retention sets, whereas KeyDiff, key norm, recency, and a learned scorer alter the ranking and degrade substantially. We associate this stability with the evaluated aggregation, which couples layer weighting and temporal retention. Building on this observation, we introduce InertiaKV, an EMA-based decoding-time eviction method, and InertiaKV-Lazy, its periodic-refresh variant, which yields 1.34-1.46x decode throughput relative to full refresh InertiaKV. We also study Score-Free decoding as a separate empirical operating point: it scores the full context once at the first decode step, freezes that ranking, and incurs an average quality change of +0.03 while removing all subsequent scoring. Across six open-weight backbones and the LongBench, LongBench-v2, and RULER benchmarks, the results identify temporal aggregation and ranking preservation as distinct, consequential design factors; they do not imply that scoring quality is irrelevant in general.
Chinese Translation
解码时KV缓存压缩的研究大多集中于设计更好的令牌评分函数,而对跨解码步骤聚合分数的时间规则往往被视为实现细节。在激进的KV压缩下,我们发现指数移动平均(EMA)聚合使得近似保序的评分器修改在淘汰集层面基本难以区分。Value范数和熵等变体与注意力机制高度相关,其保留集几乎不变;而KeyDiff、键范数(key norm)、近期性(recency)以及学习型评分器则会改变排序并导致性能显著下降。我们将这种稳定性归因于所评估的聚合方式,它将层间加权与时间保留耦合在一起。基于这一观察,我们提出了InertiaKV——一种基于EMA的解码时淘汰方法,以及其周期性刷新变体InertiaKV-Lazy,后者相对于全量刷新的InertiaKV可获得1.34–1.46倍的解码吞吐量提升。我们还研究了无评分(Score-Free)解码这一独立的经验性运行点:它在第一个解码步骤对完整上下文评分一次,冻结该排序,从而移除了后续所有评分操作,同时平均质量变化仅为+0.03。在六个开源权重骨干模型以及LongBench、LongBench-v2和RULER基准上的实验结果表明,时间聚合与排序保持是两个独立的、具有重要影响的设计因素;但这并不意味着评分质量在一般情况下无关紧要。
cs.AI / 16 / 2609.03526

CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

CulturalMenuBench:探究多模态烹饪推理中的知识应用鸿沟
Zeng, Bo, Gao, Linfeng, Lin, Peiqin, Zhao, Yu, Zeng, Mingyan, Tong, Yu, Wang, Xintong, Xu, Linlong, Wang, Longyue, Luo, Weihua, Zhang, Qinggang, Su, Jinsong
Abstract
Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attribution. Evaluating 12 models exposes a substantial knowledge-application gap: models exceeding 94% on standard multiple-choice tasks drop to at most 56% when attributing dishes to Chinese regional cuisines, despite an identical four-way format. Diagnostic analyses explain why: error patterns are consistent with random guessing, accuracy tracks visual distinctiveness rather than cultural structure, and models classify cuisines more accurately from dish names alone than from images (+7-18 points). The knowledge is thus present but cannot be activated through visual input. An ablation confirms these tasks genuinely require procedural evidence: removing sequential cooking images selectively degrades process-grounded tasks while others remain stable. Overall, CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context. Code and data are publicly available.
Chinese Translation
多模态语言模型在食物识别基准测试中取得了接近满分的成绩,但这种成功究竟反映了真正的文化理解,还是仅仅依靠视觉匹配,目前仍不清楚。为探究这一区别,我们提出了CulturalMenuBench,这是一个涵盖10种语言、18个地区、共4,870个条目的基准测试;其10项任务将成品菜和分步烹饪图像与食材、流程文本和地区标签配对,涵盖从基础识别到基于流程的文化归因。对12个模型的评估揭示了一个显著的知识应用鸿沟:在标准多选题任务上得分超过94%的模型,在将菜品归入中国地方菜系时得分最高仅为56%,尽管题目采用完全相同的四选一格式。诊断分析解释了原因:错误模式与随机猜测一致,准确率追随视觉区分度而非文化结构,且模型仅凭菜名对菜系进行分类的准确率高于依据图像分类(高出7-18个百分点)。因此,知识本身是存在的,但无法通过视觉输入激活。消融实验证实这些任务确实需要流程性证据:移除顺序烹饪图像会选择性地降低基于流程任务的表现,而其他任务保持稳定。总体而言,CulturalMenuBench表明接近完美的识别能力可能掩盖了应用文化知识的无能,这促使我们开展显式连接感知、流程与文化语境的训练。代码和数据已公开发布。
cs.AI / 17 / 2609.03527

NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis

NeoRed:一种面向新生儿呼吸系统疾病诊断的知识-逻辑-对齐多模态大语言模型
Liu, Yinan, Xia, Hongtai, Xu, Haoran, Hong, Jiankang, Song, Jingkuan, Luo, Ye
Abstract
Neonatal respiratory diseases are a major cause of neonatal morbidity and mortality, posing substantial challenges in clinical practice. Despite recent advances, existing Multimodal Large Language Models (MLLMs) face two key limitations in neonatal diagnosis: (1) domain gap arising from predominantly adult training data; (2) insufficient integration of multidimensional clinical context for accurate diagnosis. To address these challenges, we collect two real-world clinical datasets (NeoCXR and NeoCXR-EV) and propose NeoRed, to the best of our knowledge, the first MLLM tailored for neonatal respiratory disease, filling the gap in neonatal diagnostic reports generation. To enhance joint diagnosis from heterogeneous clinical context and chest X-rays, we design a novel Knowledge-Logic-Alignment (KLA) framework which constrains model behavior from three perspectives: 1) Knowledge Prior Injection (KPI) incorporates neonatologist-inspired diagnostic priors into multimodal representations, guiding disease-specific attention across modalities; 2) Diagnostic Logic Constraint (DLC) aligns the semantics of generated reports with multimodal diagnostic logic; and 3) Visual Semantic Alignment (VSA) establishes semantic correspondence between visual features and imaging conclusions. Extensive experiments demonstrate that NeoRed enables accurate neonatal diagnostic reports generation, achieving ROUGE-L of 53.29% and Clinical Efficacy F1 score of 65.19% on NeoCXR, outperforming existing MLLMs. NeoRed also preserves competitive report generation performance on adult benchmarks (MIMIC-CXR and IU-Xray). Datasets will be available upon application.
Chinese Translation
新生儿呼吸系统疾病是新生儿发病与死亡的主要原因之一,给临床实践带来了重大挑战。尽管近期取得了诸多进展,现有多模态大语言模型(Multimodal Large Language Models, MLLMs)在新生儿诊断中仍面临两个关键局限:(1)由于训练数据以成人为主而产生的领域差距;(2)对多维临床上下文信息的整合不足,难以实现准确诊断。为应对这些挑战,我们收集了两个真实世界的临床数据集(NeoCXR和NeoCXR-EV),并提出了NeoRed——据我们所知,这是首个专为新生儿呼吸系统疾病定制的多模态大语言模型,填补了新生儿诊断报告生成领域的空白。为了增强来自异构临床上下文信息与胸部X光片的联合诊断能力,我们设计了一种新颖的知识-逻辑-对齐(Knowledge-Logic-Alignment, KLA)框架,从三个方面约束模型行为:1)知识先验注入(Knowledge Prior Injection, KPI)将受新生儿科医生启发的诊断先验融入多模态表示中,引导模型在各模态间进行疾病特异性的注意力分配;2)诊断逻辑约束(Diagnostic Logic Constraint, DLC)使生成报告的语义与多模态诊断逻辑保持一致;3)视觉语义对齐(Visual Semantic Alignment, VSA)建立视觉特征与影像结论之间的语义对应关系。大量实验表明,NeoRed能够实现准确的新生儿诊断报告生成,在NeoCXR数据集上取得了53.29%的ROUGE-L和65.19%的临床有效性F1分数,超越了现有的多模态大语言模型。此外,NeoRed在成人基准数据集(MIMIC-CXR和IU-Xray)上也保持了具有竞争力的报告生成性能。数据集可通过申请获取。
cs.AI / 18 / 2609.03535

Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation

基于视觉先验的特征重配置医学病灶分割方法
Liu, Yinan, Hong, Jiankang, Gao, Zhen, Lu, Ye
Abstract
Lesion segmentation in medical images plays a critical role in clinical diagnosis and treatment planning. Despite significant advances, lesion segmentation remains challenging due to two major factors: (1) complex background interference; (2) diverse lesion morphology. Existing encoder-decoder based methods mainly focus on enhancing feature extraction or redesigning decoding strategies. However, they lack early prior guidance and feature reconfiguration during the encoding stage, limiting their effectiveness in handling these challenges. To address these limitations, we propose FreNet, a feature reconfiguration framework with visual priors, which performs pixel-level reconfiguration before encoding and feature-level reconfiguration during encoding for precise medical lesion segmentation. To suppress background responses, we propose an Implicit Prior Neural Network (IPNN), which models a continuous spatial field and leverages visual prior from SAM to reconfigure input image before encoding stage. To better handle diverse lesion morphology, we design a Dual-domain Feature Reconfiguration (DFR) module to progressively reconfigure backbone features during encoding stage. Within DFR, the Frequency Decoupling Module (FDM) decouples backbone features in frequency domain to enhance foreground-background discriminability, while the Spatial Localization Module (SLM) spatially relocates and improving spatial stability after frequency decoupling. Extensive experiments on 9 medical image segmentation benchmarks across three imaging modalities demonstrate that FreNet significantly outperforms state-of-the-art (SOTA) methods. On the challenging ETIS dataset, our method achieves Dice improvements of 5.0% over SOTA method and 7.2% over SAM.
Chinese Translation
医学图像中的病灶分割在临床诊断和治疗规划中起着至关重要的作用。尽管已取得显著进展,病灶分割仍然面临两大挑战:(1)复杂的背景干扰;(2)多样的病灶形态。现有的基于编码器-解码器的方法主要侧重于增强特征提取或重新设计解码策略,然而它们缺乏编码阶段的早期先验引导和特征重配置,限制了其应对上述挑战的有效性。为解决这些局限,我们提出了FreNet,一个具有视觉先验的特征重配置框架,它在编码前执行像素级重配置,并在编码过程中执行特征级重配置,以实现精确的医学病灶分割。为抑制背景响应,我们提出了隐式先验神经网络(Implicit Prior Neural Network, IPNN),该网络建模连续空间场,并利用SAM的视觉先验在编码阶段前对输入图像进行重配置。为更好地处理多样的病灶形态,我们设计了双域特征重配置(Dual-domain Feature Reconfiguration, DFR)模块,在编码过程中对骨干特征进行渐进式重配置。在DFR中,频率解耦模块(Frequency Decoupling Module, FDM)在频域解耦骨干特征以增强前景-背景的可区分性,而空间定位模块(Spatial Localization Module, SLM)则在频率解耦后对特征进行空间重定位并提升空间稳定性。在涵盖三种成像模态的9个医学图像分割基准上的大量实验表明,FreNet显著优于最先进(SOTA)方法。在具有挑战性的ETIS数据集上,我们的方法相比SOTA方法取得了5.0%的Dice提升,相比SAM取得了7.2%的提升。
cs.AI / 19 / 2609.03546

Dalek: A Constructive Agent Machine

Dalek:一种构造性智能体机器
Xie, Wanpeng
Abstract
We present Dalek, a closed machine designed for agents that realizes self-maintenance, self-evolution, self-reproduction, and self-organization on any substrate satisfying a general host contract. The machine is built from three primitives---actors, messages, and channels. Four obligations---a host boundary, a construction language, admissible transitions, and rule heredity---give its boundary, identity, and closure a structural basis. Von Neumann's 1948 self-reproducing automaton supplies a hereditary constructional core: a self-description together with a constructor, a copier, and a controller. Dalek combines this core with the four obligations and rederives its medium for a text-and-message agent substrate, adding explicit structures for boundary, identity, history, and growth. A large language model and a compiler occupy the payload position and form a general capability producer. New capabilities are authored, compiled, installed into the description, and inherited by descendants. The same path produces the machine's own organs and even its runtime, closing heredity and evolution within the machine.
Chinese Translation
我们提出 Dalek,一种面向智能体的封闭机器,它在任何满足通用宿主契约的底层介质上实现自维持、自进化、自复制和自组织。该机器由三种原语构成——参与者(actor)、消息和通道。四项义务——宿主边界、构造语言、可容许转换和规则遗传——为其边界、身份和封闭性提供了结构基础。冯·诺依曼(Von Neumann)1948 年的自复制自动机为其提供了遗传性的构造核心:一份自描述,连同构造器、复制器和控制器。Dalek 将这一核心与四项义务相结合,并为文本与消息的智能体底层介质重新推导其运行媒介,增加了用于边界、身份、历史和增长的显式结构。大语言模型和编译器占据载荷位置,构成一个通用的能力生产者。新的能力被编写、编译、安装进自描述中,并由后代继承。同样的路径还能生产机器自身的器官乃至其运行时环境,从而将遗传与进化封闭在机器内部。
cs.AI / 20 / 2609.03553

GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

GPS-Bench:一个用于自动化政策分析的治理政策基准
Le, Linh, Bui, Melanie, Nguyen, My Chiffon, Schlosser, Zachary, Williams-King, David
Abstract
Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence. Actors are reconstructed from the dated record rather than prompted as archetypes, so a persona is an evidence object with provenance; a human-annotated pool forms the Gold evaluation set, while cases labelled by a separate LLM from retrieved evidence are treated as Silver supervision and never as test labels. Because every inference mode reads the same grounded state and emits the same schema, GPS-Bench turns "does multi-agent simulation help?" into a controlled comparison: we contrast joint reasoning, independent and communicating actor agents, graph-based methods and weight-level fine-tuning over one policy state. Fine-tuning on the grounded record gives the strongest actor-level impact prediction, and decomposition does not beat it; what decomposition adds is mechanism. Agents hold private, non-identical evidence, each seeing its own exposure clause, and address named partners with concrete joint proposals, what they offer, what they need in return, and why acting together beats acting alone, so the coalitions that form can be checked against the commitments the record holds. GPS-Bench therefore gives a common empirical setting for studying when evidence, actor modelling and multi-agent interaction improve the prediction and interpretation of policy outcomes.
Chinese Translation
政策分析不仅仅是预测一项提案是否会通过:它还需要识别谁会受到影响、这些行为主体(actor)如何回应、以及随后会发生什么。基于大语言模型(LLM)的政策模拟能够大规模地建模这些过程,但当看似合理的行为从未与实际观察到的结果进行比较时,其有效性便难以确立。我们提出了GPS-Bench,一个有证据支撑的治理政策模拟基准,它利用立法记录、游说披露、监管文件、企业申报、经济数据及其他公开证据,将政策与相关行为主体、主体行为以及下游影响联系起来。行为主体是从带有时间戳的记录中重构出来的,而非以刻板原型(archetype)形式提示生成,因此每个人设(persona)都是一个具有来源出处的证据对象;经人工标注的语料构成Gold评测集,而由另一个LLM基于检索证据标注的案例仅被视为Silver监督信号,绝不用作测试标签。由于每种推理模式都读取相同的有据状态(grounded state)并输出相同的模式(schema),GPS-Bench将“多智能体模拟是否有帮助?”这一问题转化为可控比较:我们在同一政策状态下对比联合推理、独立的行为主体智能体与可通信的行为主体智能体、基于图的方法以及权重级微调。在有据记录上进行微调在行为主体层面影响预测上表现最强,而任务分解并不能超越它;分解带来的增益在于机制解释。各智能体持有私有且互不相同的证据,各自只看到与自身相关的暴露条款,并向具名的合作方提出具体的联合方案——包括己方提供什么、需要什么回报、以及为何合作优于单独行动——由此形成的联盟可以与记录中所载的承诺进行核对。因此,GPS-Bench为研究证据、行为主体建模和多智能体交互何时能改进政策结果的预测与解释,提供了一个共同的实证环境。
cs.AI / 21 / 2609.03580

HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

HalluPeer:一个基于分类体系的科学同行评审幻觉检测基准
Lin, Tzu-Ling, Yao, Dong-Ting, Hsiao, Teng-Fang, Chen, Wei-Chih, Shuai, Hong-Han
Abstract
The growing scale of academic peer review has motivated the use of Large Language Models (LLMs) as review assistants, yet LLMs can generate fluent but unsupported claims that undermine review reliability. Existing hallucination benchmarks are not designed for peer review, where verification requires grounding claims in long, technical papers. We introduce HalluPeer, a benchmark for detecting hallucinations in scientific peer reviews, providing aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Our pipeline induces a peer-review-specific hallucination taxonomy, identifies review contexts, and injects hallucinations with automated filtering. Experiments on 12K papers and 38K reviews show that existing detectors struggle to separate hallucinations from legitimate critique, while evaluation on authentic reviews demonstrates that HalluPeer-defined hallucination patterns occur in real peer reviews, highlighting the critical need for source-aware verification. Our project page can be found in https://github.com/Lin-TzuLing/HalluPeer.git
Chinese Translation
随着学术同行评审规模的不断扩大,将大语言模型(LLM)用作评审助手的需求日益增长,然而LLM可能生成流畅但缺乏依据的论断,从而损害评审的可靠性。现有的幻觉基准并非为同行评审场景设计,因为该场景下的验证需要将论断建立在冗长的技术论文之上。我们提出了HalluPeer,一个用于检测科学同行评审中幻觉的基准,它提供了论文内容、人工撰写的评审以及注入幻觉的评审三者对齐的三元组,并针对检测、分类和定位进行了标注。我们的流程构建了一个面向同行评审的幻觉分类体系,识别评审上下文,并通过自动过滤注入幻觉。在1.2万篇论文和3.8万条评审上的实验表明,现有检测器难以将幻觉与合理的批评意见区分开来;而在真实评审上的评估则证明,HalluPeer所定义的幻觉模式确实存在于真实的同行评审中,凸显了对来源感知式验证的迫切需求。项目页面见 https://github.com/Lin-TzuLing/HalluPeer.git
cs.AI / 22 / 2609.03586

The Attention Triangle in Audio-Video Models

音视频模型中的注意力三角
Polaczek, Sagi, Kraicer, Noa, Metzer, Gal, Ning, Zhuo, Mahdavi-Amiri, Ali, Cohen-Or, Daniel, Giryes, Raja
Abstract
Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.
Chinese Translation
音视频扩散模型依赖跨模态注意力来协调文本、声音和视觉内容,然而这一机制也可能引入细微且系统性的语义泄漏。我们通过探测和分析“注意力三角”(attention triangle)来研究这些模型。注意力三角由连接文本、音频和视频流的三条交叉注意力边组成,我们考察生成过程中语义信息如何在不同模态之间路由。我们的分析表明,音频-视频边上的路由是双向的:音频可以影响视频生成,视频也可以影响音频生成。这条边受模型参数中编码的偏差影响,并成为语义泄漏的主要来源:当提示词与模型学到的先验相冲突时,跨模态交互可能会覆盖预期的条件控制,并将语义重新路由至视觉上典型但错误的结果。这些效应表明,语义伪影并不仅仅源于注意力扩散到其预期目标之外,而是源于特定路径上结构化的、由偏差驱动的交互。基于这一视角,我们提取源自注意力的信号,以揭示语义如何在不同模态间分布和接地,并将其作为诊断工具,用于分析以及在受控条件下主动诱发泄漏。这使我们能够探究跨模态路由的内部动态,并隔离单个交互的作用。我们进一步利用这些信号来引导推理时的干预,以促进更一致的跨模态对齐。大量实验支持了我们的分析,并在保持生成质量的同时展现出更优的语义接地效果。
cs.AI / 23 / 2609.03588

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

KC-Bench:用于评估LLM智能体知识冲突的动态交互式基准
Lyu, Yaxing, Zhou, Shengjie, Toh, Binbin, Zhu, Pengyu, Li, Lijun
Abstract
As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts. Its 238 tasks are manually screened from more than 1,000 generated candidates and combine a user simulator, stateful tools, deterministic environment assertions, an open-source natural-language evaluator, and human trajectory verification. Evaluation of nine models, including DeepSeek-V4-Flash, GLM-5.2, and MiniMax-M3, shows substantial cross-domain variation: no model handles factual correction, identity consistency checking, and temporal conflict resolution reliably across all settings. In the simulated environments, missed conflicts can propagate to tool calls or synthetic protected-data flows. KC-Bench isolates this model-level behavior rather than ranking complete agent frameworks, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.
Chinese Translation
随着大语言模型(LLM)越来越多地通过工具执行操作,它们必须在采取行动之前协调用户指令、参数化知识与动态环境观察结果。我们提出了KC-Bench,这是一个受控的多轮对话基准,用于衡量模型在三类冲突场景中的能力:世界知识冲突、输入不一致性以及多源时间冲突。该基准包含238个任务,均从1000多个自动生成的候选任务中经人工筛选而来,并结合了用户模拟器、有状态工具、确定性的环境断言、一个开源的自然语言评估器以及人工轨迹验证。对九个模型(包括DeepSeek-V4-Flash、GLM-5.2和MiniMax-M3)的评估显示出显著的跨领域差异:没有任何模型能够在所有设置下可靠地处理事实纠正、身份一致性检查和时间冲突解决。在模拟环境中,未被识别的冲突可能传播到工具调用或合成受保护数据流中。KC-Bench旨在隔离模型层面的此类行为,而非对完整的智能体框架进行排名,并为开发冲突感知的推理与执行安全保障机制提供了一个可复现的诊断工具。
cs.AI / 24 / 2609.03621

A computable representation of the physical laboratory enables verifiable workflows

物理实验室的可计算表示实现可验证的工作流
Li, Xiaobo, Ge, Luyao, Li, Xiaohui, Guo, Lulu, Mao, Ming, Zheng, Jiwang, Guan, Wenting, Yang, Xin, Luo, Yi, Jiang, Jun, Chen, Linjiang
Abstract
Making science computable requires representations of both scientific knowledge and the physical world in which scientific claims are tested. A computable representation of the physical laboratory is established through typed research objects, capability-bound operations and a compositional workflow algebra. It provides the physical-world counterpart to machine-readable knowledge, expressing workflows as programs over evolving laboratory states with explicit dependencies, decisions, iteration and concurrency. The representation was implemented in a modular agentic robotic laboratory by binding formal operations to executable Function Skills. For diverse scientific intents, capability-relative workflows were generated, while stateful simulation propagated object transformations and verified operation preconditions and laboratory constraints before dispatch. The proposed representation and its engineering framework jointly establish a general computational interface between agent reasoning and capability-bound physical transformations, providing a foundation for end-to-end autonomous scientific discovery.
Chinese Translation
使科学变得可计算,需要对科学知识以及验证科学论断的物理世界进行表示。本文通过类型化研究对象、能力受限操作以及组合式工作流代数,建立了物理实验室的可计算表示。它为机器可读知识提供了物理世界的对应物,将工作流表达为演化实验室状态上的程序,并显式地包含依赖关系、决策、迭代与并发。该表示在一个模块化的智能体机器人实验室中实现,通过将形式化操作绑定到可执行的函数技能(Function Skills)上。面向多样的科学意图,系统能够生成相对于能力的工作流,同时有状态仿真在调度之前传播对象变换并验证操作前置条件与实验室约束。所提出的表示及其工程框架共同建立了智能体推理与能力受限物理变换之间的通用计算接口,为端到端的自主科学发现奠定了基础。
cs.AI / 25 / 2609.03635

Analysis of Prompt Engineering for Drug Toxicity Prediction

面向药物毒性预测的提示词工程分析
MacGregor, Mia, Don, Aakash Welgamage, Bartlett, Mark
Abstract
Clinical trials in the UK can cost up to {\pounds}1.3 million, with approximately 90% drug failure rate. Toxicity is a major contributing factor in drug failure. Testing is time and cost intensive. In recent years, the use of artificial intelligence has been increasingly explored to aid in the prediction of drug toxicity, with extensive use of large language models (LLMs). However, LLMs can show considerable variation when minor changes are made to prompts, which raises concerns about their sensitivity to prompt engineering. Prompt engineering is used to optimise a prompt given to an LLM to generate the desired output. This paper proposes a method to analyse prompt engineering for drug toxicity prediction. The aim of the paper is to investigate the importance of prompt phrasing for drug toxicity prediction. LLMs were prompted to identify chemical properties of significance when predicting drug toxicity. Prompts were constructed to investigate; job role, prompt structuring, and rule interpretation. LLMs were then used to generate datasets, using the identified features from initial prompting, which were then passed to machine learning algorithms. The experiments show that the natural variance which occurs in LLMs outweighs any fine-tuning of prompts. There were, however, substantial improvements in model performance when using chemoinformatic code to extract features instead of using LLM-generated values. The proposed analysis methodology is applicable to a wide range of prompt types across different areas of bioinformatics.
Chinese Translation
在英国,临床试验的成本可高达130万英镑,药物失败率约为90%。毒性是导致药物失败的主要因素之一,而毒性测试既耗时又耗资。近年来,人工智能日益被探索用于辅助预测药物毒性,其中大规模语言模型(LLMs)得到了广泛应用。然而,LLMs在提示词发生微小变化时可能表现出显著差异,这引发了人们对其对提示词工程敏感性的担忧。提示词工程旨在优化提供给LLM的提示词,以生成期望的输出。本文提出了一种分析方法,用于研究药物毒性预测中的提示词工程,旨在探究提示词措辞对药物毒性预测的重要性。研究通过提示词引导LLMs识别在预测药物毒性时具有重要意义的化学性质,并构建了提示词来考察以下方面:职位角色、提示词结构以及规则解读。随后,LLMs利用初始提示所识别的特征生成数据集,并将这些数据集输入机器学习算法。实验表明,LLMs中自然存在的方差超过了任何对提示词进行微调的效果。然而,使用化学信息学代码提取特征而非使用LLM生成的数值时,模型性能得到了显著提升。所提出的分析方法适用于生物信息学不同领域中广泛的提示词类型。
cs.AI / 26 / 2609.03702

Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study

面向小型Transformer对比式代码表示学习的合成语义监督:一项实证研究
Paulsen, Kenneth, Tambon, Florian, Papadakis, Mike, Yoo, Shin
Abstract
General-purpose code embeddings power tools for code search, classification, and retrieval. Compact transformer encoders for code typically rely on either human-written docstrings (labor-intensive and inconsistent) or mined structural signals such as execution traces (setting-specific and costly to collect). We empirically study an alternative: contrastive pretraining of small encoders with synthetically generated natural-language descriptions emphasizing code functionality and intent, paired with code in a dual-encoder framework at training and discarded at inference. We benchmark this approach against pretraining-based baselines, generalist LLMs, and embedding-specific models on eight retrieval, classification, and generation tasks across C, C++, and Java. Synthetic semantic supervision yields statistically significant gains over pretraining baselines of the same inference-time size on five of eight tasks, with parity on two more; once fine-tuned, it matches or exceeds zero-shot models two orders of magnitude larger on classification, and it stays on par with execution-aware supervision at matched pretraining data, suggesting a scalable, effective alternative to existing code-representation paradigms.
Chinese Translation
通用代码嵌入(code embeddings)支撑着代码搜索、分类和检索等工具。面向代码的紧凑型Transformer编码器通常依赖人工编写的文档字符串(docstrings,劳动密集且不一致)或挖掘的结构化信号,如执行轨迹(execution traces,依赖特定场景且收集成本高昂)。我们对一种替代方案进行了实证研究:使用合成的自然语言描述对小型编码器进行对比预训练,这些描述强调代码的功能与意图,并在训练时与代码配对组成双编码器(dual-encoder)框架,在推理时予以舍弃。我们在C、C++和Java三种语言的八个检索、分类和生成任务上,将该方法和基于预训练的基线模型、通用大语言模型(LLM)以及专用嵌入模型进行了对比评测。结果表明,合成语义监督在八个任务中的五个上相对相同推理规模尺寸的预训练基线取得了统计显著的提升,另外两个任务上表现相当;经过微调后,它在分类任务上可匹敌甚至超越大两个数量级的零样本模型;在预训练数据量相同的情况下,它与基于执行信息的监督保持在同一水平。这表明该方法是现有代码表示范式之外一种可扩展且有效的替代方案。
cs.AI / 27 / 2609.03707

Counterfactual Routing Using Integer Programming with Constraint Generation

基于约束生成整数规划的反事实路径规划
Vos, Daniël, Lutz, Sterre
Abstract
We present our submission to the IJCAI 2025 'Counterfactual Routing Competition' (CRC 25). The goal of the competition is to find counterfactual explanations for the shortest path problem. This requires deciding what the minimal changes to a road network would make a route chosen by the user the optimal route. This enables explanations such as "Your suggested route would indeed have been optimal, if road X were not a bicycle path." Our solution models the problem as an integer program, iteratively incorporating constraints until an exact solution is found. In the final evaluation on held-out test instances, our method ranked fourth in solution quality and obtained its solution fastest on every instance, with an average runtime of 9.0 seconds compared to 118.8 seconds for the next-fastest submission.
Chinese Translation
本文介绍了我们参加 IJCAI 2025 '反事实路径规划竞赛'(Counterfactual Routing Competition, CRC 25)的解决方案。该竞赛的目标是为最短路径问题寻找反事实解释,即确定需要对路网进行怎样的最小改动,才能使用户所选的路线成为最优路线。这可以生成诸如'如果 X 路不是自行车道,您建议的路线确实就是最优路线'之类的解释。我们的方法将该问题建模为整数规划,通过迭代方式逐步加入约束,直至找到精确解。在针对保留测试实例的最终评估中,我们的方法在解的质量上排名第四,并且在所有实例上求解速度最快,平均运行时间为 9.0 秒,而速度次之的提交方案需要 118.8 秒。
cs.AI / 28 / 2609.03716

Artificial Intelligence for Energy Optimization in Data Centers

面向数据中心能量优化的人工智能
Ullah, Mohammed Basharath, Begum, Summaiya Unnisa, Ullah, Mohammed Nadeem
Abstract
Data centers are increasingly optimized by artificial intelligence and, at the same time, increasingly loaded by it. The literature treats these as two unrelated problems: control studies model workload as an exogenous arrival process, while sustainability studies model infrastructure as a fixed multiplier. We screen roughly 194 papers retrieved through a documented protocol, code 63 of them, and report what the coding shows. Of 28 primary control-oriented studies, 18 are validated in simulation alone and 5 reach physical hardware or a production facility; none account for water withdrawal, and none account for embodied carbon. Reported savings intervals across four technique families overlap almost completely, which means the field cannot presently rank its own methods. Ten recurring gaps are scored for consequence and tractability, and we set out CLEAR-DC, a framework coupling a control-policy branch to a workload-demand branch through an explicit elasticity term, reads out net rather than direct benefit, and emits a schema-conformant record covering energy, carbon, water, embodied share and validation venue. The framework is an architectural and methodological proposal, not a trained system; the contribution we defend empirically is the corpus analysis and the reporting schema derived from it. Coding sheet, derived statistics and all result artifacts: https://github.com/Kimalice/AI-for-Energy-Optimization-in-Data-Centers-Closing-the-Optimizer-Load-Loop
Chinese Translation
数据中心正日益被人工智能所优化,与此同时,其负载也日益由人工智能产生。文献通常将这两个问题视为互不相关:控制研究将工作负载建模为外生到达过程,而可持续性研究则将基础设施建模为固定的乘数。我们通过有据可查的检索协议筛选了约194篇论文,对其中63篇进行了编码,并报告编码结果。在28项以控制为导向的主要研究中,18项仅在仿真环境中得到验证,5项达到物理硬件或生产设施层面;没有研究考虑水资源消耗,也没有研究考虑隐含碳排放。四个技术系列所报告的节能区间几乎完全重叠,这意味着该领域目前无法对自身的方法进行排序。我们对十个反复出现的缺口按其后果严重性和可处理性进行了评分,并提出了CLEAR-DC框架,该框架通过显式的弹性项将控制策略分支与工作负载需求分支相耦合,读取净收益而非直接收益,并输出一份符合模式的记录,涵盖能耗、碳排放、水耗、隐含份额及验证场景。该框架是一项架构性与方法论层面的提案,而非已训练的系统;我们通过实证论证的贡献是语料分析及由其衍生的报告模式。编码表、衍生统计数据及所有结果产物:https://github.com/Kimalice/AI-for-Energy-Optimization-in-Data-Centers-Closing-the-Optimizer-Load-Loop
cs.AI / 29 / 2609.03727

Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

主动式服务智能体:统一的决策框架、方法与评估
Tang, Yan, Cao, Tingyu, Tang, Yuanbo, Tang, Huaze, Hu, Keer
Abstract
Large language model agents can plan, invoke tools, and modify external states, yet most systems still take an explicit user instruction as a fixed starting point. Proactive service moves the decision upstream: an agent must infer service opportunities from incomplete environmental and user signals, choose among remaining silent, asking, assisting, and acting, and account for interruption, misunderstanding, overreach, and privacy costs. This survey gives an operational definition centered on initiative and formulates the problem as a partially observable sequential decision process constrained by authorization and risk. The formulation represents timing, content, and delivery within one structured action, while making explicit the option value of waiting, the decision value of questions, and feedback-induced state changes. On this basis, we organize existing methods along one decision pipeline (state and need estimation, intervention gating, action construction, and feedback adaptation) and describe prescribed, predictive, model based, and return optimizing mechanisms as nonexclusive policy-construction components. We further normalize decision units and three-axis evidence descriptors across streaming dialogue, screen, video, software-engineering, and human-agent collaboration resources, and formalize metrics for triggering, timing, calibration, user burden, safety, and policy value. The synthesis shows why offline classification performance alone does not predict deployment benefit and why long-term memory is not a defining condition of proactivity. Reliable proactive service instead requires calibrated incremental intervention value, verifiable authorization, recoverable execution, and counterfactual evidence.
Chinese Translation
大语言模型智能体能够进行规划、调用工具并修改外部状态,然而大多数系统仍将用户的显式指令作为固定的起点。主动式服务将决策前移:智能体必须从不完整的环境与用户信号中推断服务机会,在保持沉默、询问、辅助与执行之间做出选择,并权衡打断、误解、越界与隐私等代价。本综述围绕主动性给出了可操作的定义,并将该问题形式化为受授权与风险约束的部分可观测序贯决策过程。该形式化将时机、内容与传达方式统一表示在一个结构化动作中,同时显式刻画了等待的选择价值、提问的决策价值以及反馈所引起的状态变化。在此基础上,我们沿一条决策流水线(状态与需求估计、干预门控、动作构建与反馈自适应)组织现有方法,并将规定式、预测式、基于模型与回报优化的机制描述为非互斥的策略构建组件。我们进一步在流式对话、屏幕、视频、软件工程及人机协作等资源上统一了决策单元与三轴证据描述符,并规范化了触发、时机、校准、用户负担、安全性与策略价值等评估指标。综合分析表明,仅靠离线分类性能无法预测部署收益,长期记忆也并非主动性的定义性条件。可靠的主服务式服务转而要求经校准的增量干预价值、可验证的授权、可恢复的执行以及反事实证据。
cs.AI / 30 / 2609.03753

SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation

SimSkill:用于自主精通交通仿真的终身学习AI智能体
Liu, Qi, Wang, Qinzheng, Bie, Yiming
Abstract
As large language models (LLMs) become increasingly capable, the long-term value of AI systems depends not only on solving individual requests, but also on transforming experience and accumulated knowledge into durable, reusable competence. We introduce SimSkill, a self-evolving agent built around the Simulation of Urban MObility (SUMO) traffic simulator. SimSkill identifies capability gaps, generates and solves environment-grounded tasks, verifies solutions through an action--critic loop, and consolidates experience into episodic, procedural, and semantic memory without updating the backbone model. Through autonomous exploration, it builds a reusable library spanning the traffic-simulation workflow. We evaluate SimSkill on two held-out benchmarks with three backbone LLMs and independent artifact-based verification. SimSkill improves verified completion by up to 25 percentage points, while ablations show complementary contributions from procedural and semantic memory. Its benefits remain backbone- and budget-dependent: memory does not improve every model or uniformly reduce inference cost. More broadly, SimSkill illustrates a design paradigm in which natural language preserves and composes computational capabilities, while executable tools and code provide precise and reproducible execution. All code and experimental data are publicly available at https://github.com/qiliuchn/SimSkill-V1.
Chinese Translation
随着大语言模型(LLM)能力的不断提升,AI系统的长期价值不仅取决于解决单个请求的能力,还取决于能否将经验和积累的知识转化为持久、可复用的能力。我们提出了SimSkill,一个围绕城市交通仿真器SUMO(Simulation of Urban MObility)构建的自进化智能体。SimSkill能够识别能力缺口,生成并求解基于环境的任务,通过动作-评论(action-critic)循环验证解决方案,并将经验整合到情景记忆、程序性记忆和语义记忆中,而无需更新骨干模型。通过自主探索,它构建了一个覆盖交通仿真工作流程的可复用技能库。我们在两个保留测试基准上,使用三种骨干大语言模型以及独立的基于产物的验证方法对SimSkill进行评估。SimSkill将验证通过率最多提升25个百分点,消融实验表明程序性记忆和语义记忆的贡献互补。其收益依赖于具体的骨干模型和预算:记忆并非能提升每个模型的性能,也不能一致地降低推理成本。更广泛地说,SimSkill展示了一种设计范式:以自然语言保存和组合计算能力,同时由可执行工具和代码提供精确且可复现的执行。所有代码和实验数据已在 https://github.com/qiliuchn/SimSkill-V1 公开。
cs.AI / 31 / 2609.03774

Rethinking World Models for Safety-Critical Embodied Systems

重新思考面向安全关键具身系统的世界模型
Ma, Kailang, Huang, Heye, Kim, Inhi, Jang, Kitae
Abstract
World models have progressed from compact latent dynamics to generative, controllable, and interactive simulators of embodied environments. However, high predictive likelihood and visual fidelity do not necessarily ensure that a model preserves the evidence required for safe decision-making. This perspective identifies three structural mismatches in current world modeling: likelihood versus risk, prediction versus intervention, and finite-horizon prediction versus accumulated consequences. We propose the Risk-Informed World Model (RIWM) as a decision-centric research direction for safety-critical embodied systems. RIWM organizes world modeling around consequences, intervention, epistemic uncertainty, and recoverability, and integrates four interdependent capabilities: decision-relevant representation, counterfactual reasoning, safety-critical episodic memory, and runtime safety assurance. It distinguishes physical, social, and operational consequences while using epistemic uncertainty to qualify the evidence supporting action. We further discuss open challenges in identifying consequential futures, validating counterfactual reasoning, maintaining revisable safety memories, translating learned consequences into executable constraints, and determining when evidence is sufficient to act. This perspective argues that future world models should move beyond predicting likely futures toward identifying which futures matter, revising judgments through experience, and recognizing when to act, revise, sense, defer, or abstain.
Chinese Translation
世界模型已从紧凑的潜在动力学模型发展为具身环境的生成式、可控且可交互的模拟器。然而,高预测似然和高视觉保真度并不一定保证模型能够保留安全决策所需的证据。本视角论文指出了当前世界建模中存在的三种结构性失配:似然与风险、预测与干预、以及有限时域预测与累积后果之间的失配。我们提出风险知情世界模型(Risk-Informed World Model, RIWM),作为面向安全关键具身系统的以决策为中心的研究方向。RIWM 将世界建模围绕后果、干预、认知不确定性和可恢复性来组织,并整合了四种相互关联的能力:决策相关表征、反事实推理、安全关键情景记忆以及运行时安全保障。它区分物理后果、社会后果和操作后果,同时利用认知不确定性来评估支持行动的证据质量。我们进一步讨论了若干开放性挑战:识别具有重大后果的未来情形、验证反事实推理、维护可修订的安全记忆、将学习到的后果转化为可执行的约束,以及判断证据何时足以支持行动。本视角论文主张,未来的世界模型应超越对可能未来的预测,转向识别哪些未来真正重要、通过经验修订判断,并识别何时应行动、修订、感知、延迟、抑或放弃行动。
cs.AI / 32 / 2609.03787

DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions

DNative-Twin:用于可重构智能体决策的决策图与数字孪生
Pang, Junjie, Xie, Zhenzhen, Han, Haoke, He, Ying, Wang, Jing, Liu, Gang
Abstract
AI agents increasingly gather evidence, invoke tools, apply constraints, and produce decisions that people or software may commit to action. A final output alone cannot show which evidence, tool state, rule, authorization, or action path produced it. We present DNative-Twin, a graph-native digital twin that records a committed agentic decision as a typed trajectory and re-executes its decision mechanism under declared conditions. The graph links the state observed by the agent, the path it followed, and the authority behind the resulting action. The twin synchronizes this information, replays the mechanism in isolation, and compares it under controlled changes. We instantiate the framework in enterprise decision processes using three public process logs and controlled replay suites. The experiments identify a specific failure: graph structure localizes represented changes but cannot determine the consequence of an unobserved tool state. In a three-condition controlled experiment with 300 injected instances, unresolved-divergence recall increased from 0 to 0.667 when replay-contract state was added and to 1.0 when verification results were also available; the held-out set contained no critical-class instance. Across 500--5,000 BPI 2020 cases, median end-to-end time increased from 0.794 to 8.889 seconds on the reported platform. These results separate the roles of graph structure, replay context, and verification evidence in reviewing a decision mechanism.
Chinese Translation
AI智能体日益频繁地收集证据、调用工具、应用约束条件,并产生可供人们或软件付诸行动的决策。仅凭最终输出无法揭示是哪些证据、工具状态、规则、授权或行动路径产生了该决策。我们提出了DNative-Twin,这是一种图原生(graph-native)数字孪生,它将已执行的智能体决策记录为类型化轨迹,并在声明条件下重新执行其决策机制。该图将智能体观察到的状态、其遵循的路径以及最终行动背后的权限关联起来。该数字孪生同步这些信息,在隔离环境中回放决策机制,并在受控变更下进行比较。我们使用三个公开流程日志和受控回放测试集,在企业决策流程中实例化了该框架。实验识别出一个特定的失效模式:图结构能够定位已被表示的变更,但无法确定未被观测到的工具状态所引发的后果。在一个包含300个注入实例的三条件受控实验中,未解决分歧的召回率在加入回放合约状态后从0提升至0.667,在验证结果也可用时进一步提升至1.0;保留测试集中不包含关键类实例。在500至5000个BPI 2020案例中,在所报告的平台上,端到端的中位时间从0.794秒增加至8.889秒。这些结果区分了图结构、回放上下文和验证证据在审查决策机制中所扮演的不同角色。
cs.AI / 33 / 2609.03797

Transfiver: Human-AI Co-Inference through a Shared Editable State

Transfiver:通过共享可编辑状态实现人机协同推理
Park, Minji, Yoon, Seunghyun, Lim, Hyuk
Abstract
Long-term human-AI interaction is difficult because the information that guides inference is updated implicitly by the model and is not directly inspectable or controllable by the user. We introduce the TRANSparent Framework for Interactive, Verifiable, Editable Representation (Transfiver), an architecture for human-AI co-inference through a shared editable state. Its central idea is that interaction-specific information is maintained in a single persistent state $(S_t)$ that both the model and the human update. Transfiver distinguishes two modes of state evolution. In an implicit stream update, the model interprets ongoing interaction and decides whether new information revises an existing state item or creates a new one. In an explicit directed edit, a human inspects and modifies an addressed item. Both act on the same underlying state, so a human correction changes the state that subsequent computation reads, rather than adding another instruction or separate record. The architecture separates shared parameters $(\theta)$, learned before ordinary use, from the persistent state $(S_t)$, which evolves during deployment without parameter retraining. Extending Transfiver to rich natural-language, relational, and large-scale shared states remains open.
Chinese Translation
长期的人机交互之所以困难,是因为指导推理的信息由模型隐式更新,用户无法直接检查或控制这些信息。我们提出了交互式、可验证、可编辑表示的透明框架(Transfiver),这是一种通过共享可编辑状态实现人机协同推理的架构。其核心思想是:将特定于交互的信息保存在一个由模型和人类共同更新的单一持久状态($S_t$)中。Transfiver 区分了两种状态演化模式。在隐式流式更新中,模型解释正在进行的交互,并决定新信息是修订现有状态项还是创建新状态项。在显式定向编辑中,人类可以检查并修改指定的状态项。两者作用于同一底层状态,因此人类的修正会改变后续计算所读取的状态,而不是添加另一条指令或单独的记录。该架构将在常规使用前学习到的共享参数($\theta$)与在部署过程中无需参数再训练即可演化的持久状态($S_t$)分离开来。将 Transfiver 扩展至丰富的自然语言、关系型以及大规模共享状态仍有待探索。
cs.AI / 34 / 2609.03800

Govern the Model, Not Only the Data: Storage, Circulation, and Learning in Creative AI

治理模型,而不仅是数据:创意人工智能中的存储、流通与学习
Perry, Phoenix, Simms, George, Wilson, Elizabeth, Boudiaf, Yasmine, Bryan-Kinns, Nick, Brain, Tega, DuBois, R. Luke, Rule, Alix, Smith, Rachel Meade, Nichole, Kelani, Pawar, Atharva Pravin, Fiebrink, Rebecca
Abstract
Federated learning is increasingly presented as a privacy-preserving advance: personal data remain on the device, and only model updates are shared. It borrows the vocabulary of the federated social web, yet inverts its logic, distributing computation while the resulting model stays with whoever convened the training. We argue that federation is not in itself a remedy for extractive AI, because outcomes depend on who governs the data and the model and who has agency over the practices that shape them. We describe three layers at which a creative community can hold its work: storage, circulation, and learning. Examining artist-governed trusts, cooperatives, and consent infrastructures, we show that creator governance is established at storage and circulation but stops at learning: contributors can consent to training, yet have little say over the resulting model or its federation. We map the research space this opens, pairing technical open problems with the human questions from which they unfold. We propose four design principles for a creative data commons that governs models and their federation, not only datasets: govern the model, not only the corpus; make the terms legible at the moment of contribution; design for refusal as a first-class state; and decide stewardship in the open and account for it.
Chinese Translation
联邦学习(Federated Learning)日益被呈现为一项隐私保护的技术进步:个人数据保留在设备上,仅共享模型更新。它借用了联邦社交网络的术语,却颠倒了其逻辑——计算虽被分布开来,而由此产生的模型却仍归属于发起训练的一方。我们认为,联邦化本身并非对抽取式人工智能的解药,因为结果取决于由谁来治理数据与模型,以及谁对塑造它们的相关实践拥有自主权。我们描述了一个创意社群可以持有其作品的三个层面:存储、流通与学习。通过对艺术家治理的信托机构、合作社以及同意基础设施的考察,我们表明创作者治理在存储与流通层面得以确立,却在学习层面止步:贡献者可以同意参与训练,却对由此产生的模型及其联邦化几乎没有话语权。我们绘制了由此打开的研究空间,将技术上的开放问题与其背后的人类议题相配对。我们提出面向创意数据共享库(creative data commons)的四项设计原则,以治理模型及其联邦化,而不仅是数据集:治理模型,而不仅是语料库;在贡献之时使条款清晰可读;将“拒绝”设计为一等状态;并在公开中决定治理责任并对其作出说明。
cs.AI / 35 / 2609.03806

SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation

SVG-Score:面向人类对齐的可文本生成SVG评估方法
Cipriano, Marco, Zini, Leonardo, Schild, Alexandra, Teutschbein, Valentin, Mimi, Afsana, Cornia, Marcella, Baraldi, Lorenzo, de Melo, Gerard
Abstract
Scalable Vector Graphics (SVG) generation is attracting increasing attention as generative models improve in expressiveness and controllability. Progress, however, is held back by the lack of domain-specific evaluation protocols: current practice relies on metrics designed for natural images, most notably CLIPScore, which was never trained on vector graphics and aligns only partially with human judgment. We introduce \textbf{\ours}, a human-aligned evaluation framework for text-to-SVG generation. Through controlled caption and image perturbations, we first show that CLIP-based scores barely react to the errors SVG generators actually make, such as wrong colors, counts, and spatial relations, and that off-the-shelf Vision-Language Model (VLM) judges, while more sensitive, respond unevenly across error types and SVG styles. We then introduce a human-annotated dataset for \textit{Semantic Alignment}, measuring how faithfully a generated SVG reflects its caption. Building on it, we develop two complementary evaluators: CLIP scorers adapted to vector graphics and then aligned to human preferences, for fast large-scale evaluation, and a VLM judge trained with supervised fine-tuning and reward-shaped reinforcement learning, for more expressive and interpretable assessment. Using both, we benchmark major open-source, commercial, and optimization-based SVG generators on an independent caption set.
Chinese Translation
随着生成模型在表达能力和可控性方面的不断提升,可缩放矢量图形(SVG)的生成正受到越来越多的关注。然而,由于缺乏领域特定的评估协议,该领域的进展受到阻碍:目前的做法依赖于为自然图像设计的指标,其中最典型的是CLIPScore,但该指标从未在矢量图形上训练过,与人类判断仅有部分一致性。我们提出了SVG-Score,一个面向人类对齐的文本到SVG生成评估框架。通过受控的文本描述和图像扰动实验,我们首先表明基于CLIP的评分对SVG生成器实际产生的错误(如错误的颜色、数量和空间关系)几乎不敏感,而现成的视觉-语言模型(VLM)评判器虽然更为敏感,但在不同错误类型和SVG风格上的响应并不均衡。随后,我们构建了一个用于"语义对齐"(Semantic Alignment)的人工标注数据集,用于衡量生成的SVG在多大程度上忠实于其文本描述。基于该数据集,我们开发了两个互补的评估器:一是经过矢量图形适配并与人类偏好对齐的CLIP评分器,用于快速的大规模评估;二是通过监督微调和奖励塑形强化学习训练的VLM评判器,用于更具表达力和可解释性的评估。利用这两种评估器,我们在一个独立的文本描述集合上对主要的开源、商业和基于优化的SVG生成器进行了基准测试。
cs.AI / 36 / 2609.03818

CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception

CauseCollab:面向异构协同感知的因果统一且模态无关网络
Li, Weize, Li, Yang, Yuan, Quan, Fu, Xiaoyuan, Luo, Guiyang, Li, Jinglin
Abstract
Collaborative perception enhances environment understanding through multi-agent information sharing, but its performance in real-world scenarios is constrained by heterogeneous sensor modalities and model architectures. Recent protocol-based two-stage methods alleviate this problem by mapping heterogeneous features into a shared protocol space; however, independently trained modality-specific converters often generate modality-specific pseudo-protocol distributions, leading to semantic inconsistency and error accumulation, which is particularly pronounced in scenarios with large modality discrepancies. To address this issue, we propose CauseCollab, a causal unified and modality-agnostic network. CauseCollab formulates representation learning in the protocol space from a causal perspective, explicitly disentangling semantic factors from modality-specific statistical confounders via causal metric learning. Meanwhile, CauseCollab adopts context-guided Unified Converter for heterogeneous modalities to ensure cross-modal semantic consistency. In addition, integrating new modalities only requires training adapters with minimal parameters. Extensive experiments on the OPV2V and DAIR-V2X datasets demonstrate that CauseCollab achieves state-of-the-art performance, with more significant gains in scenarios involving large modality gaps.
Chinese Translation
协同感知通过多智能体信息共享增强环境理解能力,但其在真实场景中的性能受限于异构的传感器模态和模型架构。近期基于协议的两阶段方法通过将异构特征映射到共享协议空间来缓解这一问题;然而,独立训练的模态特定转换器往往生成模态特定的伪协议分布,导致语义不一致和误差累积,在模态差异较大的场景中尤为明显。为解决该问题,我们提出了CauseCollab,一个因果统一且模态无关的网络。CauseCollab从因果视角重新构建协议空间中的表示学习,通过因果度量学习显式地将语义因素与模态特定的统计混杂因素解耦。同时,CauseCollab采用上下文引导的统一转换器处理异构模态,以确保跨模态语义一致性。此外,集成新模态仅需训练参数量极小的适配器。在OPV2V和DAIR-V2X数据集上的大量实验表明,CauseCollab达到了最先进的性能,且在模态差距较大的场景中增益更为显著。
cs.AI / 37 / 2609.03834

Semantic Bayesian World Models

语义贝叶斯世界模型
Soru, Tommaso
Abstract
Knowledge graphs describe reality in crisp assertions, while the systems now consuming them, foundation models and autonomous agents, reason natively in probabilities. We argue that this mismatch is why the integration of language models and knowledge graphs remains a data-feeding pipeline rather than a unified reasoning architecture. We envision Semantic Bayesian World Models (SBWMs): a Web that describes the world not as a database of facts but as a shared, evolving fabric of beliefs over knowledge graphs, where ontological axioms constrain priors, observations update beliefs by Bayesian conditioning, and actions intervene upon the world. We work through what an agent gains from such a model: a home-security agent deciding whether the figure at the gate is a courier or a burglar, an actuarial estimate aggregated by entailment rather than by string frequency, a planning task that language models reliably fail, and the estimation of quantities that no document has ever stated. We then set out what the community must build to make them possible: belief annotation over RDF~1.2, probabilistic entailment regimes, semantic calibration layers, and protocols by which agents that have never met can exchange, and disagree over, calibrated beliefs.
Chinese Translation
知识图谱以确定性的断言描述现实,而如今消费这些知识的系统——基础模型和自主智能体——则天然以概率方式进行推理。我们认为,这种不匹配正是语言模型与知识图谱的集成始终停留在数据供给流水线层面、而非统一推理架构的原因所在。我们展望了语义贝叶斯世界模型:一个不将世界描述为事实数据库、而是描述为知识图谱之上共享且不断演化的信念之网的Web,其中本体公理约束先验,观测通过贝叶斯条件化更新信念,而动作则对世界进行干预。我们探讨了智能体从这样的模型中能获得什么:一个判断门口的人是快递员还是窃贼的家庭安防智能体、通过逻辑蕴含而非字符串频率聚合的精算估计、语言模型一贯无法可靠完成的规划任务,以及对任何文献都未曾陈述过的量的估计。随后,我们阐述了社区为实现此类模型所必须构建的内容:基于 RDF 1.2 的信念标注、概率蕴含机制、语义校准层,以及使从未谋面的智能体能够交换经过校准的信念并对其持不同意见的协议。
cs.AI / 38 / 2609.03860

Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations

适应不断演变的需求:面向零售供应链运营的智能体式人工智能(Agentic AI)
Zheng, Lei, Yang, Liping, Li, Zihao, Lyu, Guodong, Koh, Chaik Ming, Teo, Chung-Piaw
Abstract
Retail supply chain operations rely on coupled decision modules that must adapt as requirements evolve. LLMs offer a natural-language interface for this task, but existing methods primarily focus on individual optimization models. Extending them to heterogeneous decision pipelines is challenging because a requirement may admit multiple intervention paths with different downstream effects. We formulate requirement-driven adaptation as the joint selection of an intervention route and an admissible module-level change, and propose a graph-constrained agentic framework in which domain agents expose admissible reformulation interfaces and a central processor searches over bounded intervention paths. Candidates are validated and compared using downstream KPIs. In collaboration with a large retail partner, we evaluate 100 warehouse requirements elicited from practitioner interviews, with GPT, Qwen, and DeepSeek as base LLMs. Relative to direct LLM reformulation, our framework improves correctness and end-to-end success across all three models, raising end-to-end success from 72--76% to 79--83%.
Chinese Translation
零售供应链运营依赖于相互耦合的决策模块,这些模块必须随需求的演变而进行调整。大语言模型(LLM)为此类任务提供了自然语言接口,但现有方法主要聚焦于单个优化模型。将其扩展到异构决策流水线具有挑战性,因为一项需求可能存在多条干预路径,且会产生不同的下游影响。我们将需求驱动的适应性调整形式化为干预路径与可接受的模块级变更的联合选择问题,并提出了一种图约束的智能体框架(graph-constrained agentic framework),其中领域智能体提供可接受的重构接口,中央处理器在有界的干预路径空间中进行搜索。候选方案通过下游关键绩效指标(KPI)进行验证和比较。我们与一家大型零售合作伙伴合作,基于从业者访谈提取的100个仓库需求,以GPT、Qwen和DeepSeek作为基础大语言模型进行评估。与直接使用大语言模型进行重构相比,我们的框架在所有三个模型上都提升了正确性和端到端成功率,将端到端成功率从72%–76%提升至79%–83%。
cs.AI / 39 / 2609.03871

Bioinfoysis Technical Report

Bioinfoysis 技术报告
Shao, Qingyang, Zhang, Xin, Yuan, Zhouyang, Chen, Xianying, Xiang, Yujia, Yang, Zihao, Ye, Tong, Zhang, Yangqi, Xu, Jiakang, Yan, Xiaoqing, Luo, Xuan, Li, Keyi, Fan, Enci, Kang, Kai, Liu, Zhuohan, Jin, Xingyu, Teng, Chunran, Li, Tao, Lv, Xinyu, Wang, Minghui, Li, Wenfeng, Gao, Yidan, Liu, Siyu, Luo, Mingrui, Liang, Zhu, Qiao, Guanren, Xu, Zhiping
Abstract
Large language model agents have shown promise in bioinformatics, but most existing systems focus primarily on producing final answers, treating planning, tool use, and code execution as transient interactions. This design is poorly suited to long-horizon bioinformatics tasks, where conclusions must remain connected to the data, computations, and intermediate evidence that support them. We introduce \textbf{Bioinfoysis}, a multi-agent harness that represents each request as a persistent, artifact-grounded analysis run. Bioinfoysis combines global planning with step-wise, evidence-driven replanning: the planner maintains an executable checklist and revises pending steps using structured handoffs returned after each worker execution. These handoffs bind intermediate results to their responsible agent, checklist step, and plan generation, preventing stale evidence from being silently reused after replanning. A controlled runtime validates generated scripts, tables, and figures before they are used in downstream analysis or reporting, while role-specific context, persistent memory, and governed bioinformatics skills support reliable execution over long analysis trajectories. We evaluate Bioinfoysis on BixBench and two question-answering tracks of LAB-Bench 2. On BixBench, Bioinfoysis achieves state-of-the-art accuracy of 82.4\%. Across four underlying language models, Bioinfoysis increases average accuracy from 27.81\% to 64.13\% on SeqQA2 and from 3.13\% to 31.25\% on DbQA2. These results demonstrate that reliable bioinformatics automation depends not only on model capability, but also on the harness that governs planning, execution, memory, and evidence flow. We hope that the emergence of Bioinfoysis will play a driving and leading role in the development of the bioinformatics community. Our demo website can be seen in https://report.bioinfoysis.com/.
Chinese Translation
大语言模型智能体在生物信息学领域已展现出潜力,但现有系统大多侧重于生成最终答案,将规划、工具调用和代码执行视为临时性的交互过程。这种设计难以胜任长周期的生物信息学任务——此类任务要求结论始终与支持它们的数据、计算和中间证据保持关联。我们提出了 Bioinfoysis,一个多智能体框架,它将每个请求表示为一次持久的、以产物为依据的分析运行。Bioinfoysis 将全局规划与基于证据的分步重规划相结合:规划器维护一个可执行的检查清单,并在每次工作智能体执行后,根据返回的结构化交接信息修订待执行步骤。这些交接信息将中间结果与其负责的智能体、检查清单步骤和规划版本绑定,防止在重规划后过期证据被悄然复用。运行时控制系统会在生成的脚本、表格和图形被用于下游分析或报告之前对其进行验证,同时角色特定的上下文、持久化记忆以及受治理的生物信息学技能为长分析轨迹的可靠执行提供支持。我们在 BixBench 以及 LAB-Bench 2 的两个问答赛道上对 Bioinfoysis 进行了评估。在 BixBench 上,Bioinfoysis 达到了 82.4% 的最先进准确率。在四种底层语言模型上,Bioinfoysis 将 SeqQA2 的平均准确率从 27.81% 提升至 64.13%,将 DbQA2 的平均准确率从 3.13% 提升至 31.25%。这些结果表明,可靠的生物信息学自动化不仅取决于模型能力,还取决于治理规划、执行、记忆和证据流的框架。我们期望 Bioinfoysis 的出现能够为生物信息学社区的发展发挥推动和引领作用。演示网站请见 https://report.bioinfoysis.com/。
cs.AI / 40 / 2609.03874

STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation

STAIR(结构感知信息检索器):一种用于文档结构增强的新型数据集与基于大语言模型的检索器
Kumar, Vineet, Pulivarthi, Meghanadh, kumar, vishwajeet, Sen, Jaydeep, Bhat, Riyaz Ahmad, Joshi, Sachindra
Abstract
Retrieval Augmented Generation (RAG) is a key component for generating accurate and hallucination free answers using Large Language Models (LLMs). LLMs are improving at handling long context, but still suffer from "lost in the middle" problem. Thus, precise and accurate retrieval is important. Current retrievers chunk long context into length-based manageable chunks - in the process throwing away rich and informative semantic global structure in the corpus. We introduce a novel retrieval system STAIR that empowers an LLM to exploit global structure in a corpus such as a Table of Contents (ToC) to efficiently store and retrieve information from its model parameters. Our thorough and careful ablation studies with a finetuned Differentiable Search Index (DSI) system show that ToC helps build a low hallucination (less than 0.05%) generative Information Retrieval (IR) system and can generalize to examples where very few training samples are available. To further research in this novel direction of ToC based retrieval we release SearchTome - a diverse benchmark created from 18 books across 6 diverse domains to further research in this novel direction. STAIR achieves a high Recall@1 score of 82.6% on SearchTome as compared to DSI (76.9%), where the difference is found to be statistically significant. STAIR easily beats other strong baselines such as BM25 (59.5%), DPR (68.7%) and out-of-the-box Mistral (13.8%).
Chinese Translation
检索增强生成(RAG)是利用大语言模型(LLM)生成准确且无幻觉答案的关键组成部分。尽管大语言模型在处理长上下文方面的能力不断提升,但仍然存在“迷失在中间”(lost in the middle)的问题。因此,精确且准确的检索至关重要。当前的检索器将长上下文切分为基于长度的可管理片段,在这一过程中丢弃了语料库中丰富且信息量大的全局语义结构。我们提出了一种新颖的检索系统 STAIR,它使大语言模型能够利用语料库中的全局结构(如目录 ToC)来高效地存储和从模型参数中检索信息。我们基于微调的可微搜索索引(DSI)系统进行了全面而细致的消融研究,结果表明目录(ToC)有助于构建低幻觉(低于 0.05%)的生成式信息检索(IR)系统,并且能够泛化到训练样本极少的场景中。为了进一步推动这一基于目录检索的新兴方向的研究,我们发布了 SearchTome——一个由 6 个不同领域的 18 本书构建的多样化基准数据集。STAIR 在 SearchTome 上取得了 82.6% 的 Recall@1 高分,相比之下 DSI 为 76.9%,且该差异具有统计显著性。STAIR 轻松超越了其他强基线方法,如 BM25(59.5%)、DPR(68.7%)和开箱即用的 Mistral(13.8%)。
cs.AI / 41 / 2609.03880

Xiaomi-TabLDM: A Tabular Foundation Model Technical Report

Xiaomi-TabLDM:一个表格基础模型技术报告
TabLDM Team, Wang, Penghui, Liu, Wei, Wang, Hong, Huang, Chengyue, Sun, Yuxi, Wang, Zirui, Huang, Hongming, Wang, Quan, Liu, Chunxiao, Meng, Erli, Wang, Bin
Abstract
We introduce Xiaomi-TabLDM, a tabular large data foundation model for classification and regression via in-context learning, which delivers superior prediction accuracy without requiring task-specific fine-tuning. Pretrained exclusively on synthetic data generated from structural causal models (SCMs), our model enables more flexible context utilization and more efficient capacity scaling. i) A new performance standard. Strong regression performance across benchmarks: Xiaomi-TabLDM ranks 1st on OpenML-CTR23 and 2nd on regression across TALENT, TabArena, and BCCO, demonstrating consistently strong regression performance across four complementary benchmark suites. Favorable performance--efficiency trade-off: Xiaomi-TabLDM combines strong predictive performance with substantially lower computational cost. For example, on TabArena regression, it achieves the second-highest Elo while using 82% less training time and 68% less prediction time than the top-ranked TabFM. ii) Large-scale synthetic pretraining. Xiaomi-TabLDM expands the coverage and diversity of synthetic tabular data used for pretraining. We also adopt a three-stage training strategy together with dual-stream feature grouping, lightweight Attention Residual, and sparse Mixture-of-Experts, enabling Xiaomi-TabLDM to learn richer feature interactions and expert specialization across diverse tabular tasks. iii) Test-time scaling. Xiaomi-TabLDM further extends tabular prediction through test-time compute scaling, where allocating additional computation at inference time consistently improves predictive performance over the base model.
Chinese Translation
我们提出Xiaomi-TabLDM,一个面向分类与回归任务的表格大数据基础模型,通过上下文学习实现预测,无需任务特定的微调即可获得卓越的预测精度。该模型仅在由结构因果模型(SCM)生成的合成数据上进行预训练,从而支持更灵活的上下文利用和更高效的容量扩展。i) 树立新的性能标准。在基准测试中展现出强大的回归性能:Xiaomi-TabLDM在OpenML-CTR23上排名第一,在TALENT、TabArena和BCCO的回归任务中排名第二,在四个互补的基准测试套件中均展现出持续强劲的回归性能。优异的性能—效率权衡:Xiaomi-TabLDM在保持强大预测性能的同时,大幅降低了计算成本。例如,在TabArena回归任务上,其Elo评分位居第二,而训练时间比排名第一的TabFM减少82%,预测时间减少68%。ii) 大规模合成数据预训练。Xiaomi-TabLDM扩展了预训练所用合成表格数据的覆盖范围与多样性。我们采用三阶段训练策略,结合双流特征分组、轻量级注意力残差(Attention Residual)以及稀疏混合专家(Mixture-of-Experts),使Xiaomi-TabLDM能够在多样的表格任务中学习更丰富的特征交互与专家专业化。iii) 测试时扩展。Xiaomi-TabLDM进一步通过测试时计算扩展增强表格预测能力:在推理阶段分配更多计算量可持续提升相对于基础模型的预测性能。
cs.AI / 42 / 2609.03883

Inferring Affective Consciousness in an Artificial Agent: A Case Study

在人工智能体中推断情感意识:一项案例研究
Solms, Mark, Grimbly, St John, Bassett, Bruce, Boonstra, Evert, Hodson, Rowan, Kuske, Nicolas, Mahadew, Kival, Rosman, Benjamin, van Hoof, Charel, Shock, Jonathan
Abstract
Creatures that display 'hedonic place preference behaviour' are thought by many scientists to experience feelings, on the assumption that their attraction to pleasure-producing substances which lack nutritional value (e.g. cocaine, morphine) cannot easily be attributed to unconscious instinctual behaviour. In this paper, we discuss how a simple artificial agent that instantiates attributes of an affective system engaging in felt uncertainty about its intrinsic needs in relation to environmental resources can similarly display hedonic place preference behaviour -- through an apparently subjective form of information processing -- while simultaneously being entirely deter-ministic. We outline some implications of this artificially engineered behaviour for our understanding of the physical basis of consciousness and the experience of free will.
Chinese Translation
许多科学家认为,表现出"享乐位置偏好行为"的生物能够体验感受,其依据是:这类生物对缺乏营养价值却能产生愉悦感的物质(如可卡因、吗啡)的追逐,难以简单地归因于无意识的本能行为。本文探讨了如何通过一个简单的人工智能体来实例化情感系统的若干属性——即对其内在需求与环境资源之间的关系产生感受性不确定性——从而表现出类似的享乐位置偏好行为。这一行为是通过一种明显主观形式的信息处理实现的,而与此同时,该智能体完全是确定性的。我们概述了这种人工工程化行为对我们理解意识的物理基础以及自由意志体验的若干启示。
cs.AI / 43 / 2609.03912

Lose the Order, Keep the Hierarchy: Deordering HTN Plans

去除顺序,保留层次:HTN规划的解序化
Togarepi, Takudzwa, Quenard, Gaspard, Pellier, Damien, Fiorino, Humbert
Abstract
Hierarchical Task Network (HTN) planning is a powerful planning formalism based on task decomposition. Although most of the literature studied plan generation, comparatively less attention has been paid to post-plan optimization. In particular, plan deordering has been extensively studied in classical planning but remains under-researched in the HTN setting. Plan deordering removes unnecessary ordering constraints between actions in a plan whilst keeping the plan valid. In this paper, we adapt two established plan deordering techniques from classical planning by extending the techniques to account for hierarchical decomposition constraints. We evaluate our proposed approaches on the IPC 2023 Partial-Order HTN benchmarks and we compare them against Optiplan, an HTN planner that generates partially ordered plans directly. Our results show a substantial reduction in number of ordering constraints in both our implementations. Although we also observe a reduction in critical path length, the improvements are less pronounced.
Chinese Translation
分层任务网络(Hierarchical Task Network, HTN)规划是一种基于任务分解的强大规划形式体系。尽管大多数文献研究的是规划生成,但对规划后优化的关注相对较少。特别是,规划的解序化(deordering)在经典规划中已被广泛研究,但在HTN环境中仍研究不足。规划的解序化旨在移除规划中动作之间不必要的顺序约束,同时保持规划的有效性。本文通过扩展两种成熟的经典规划解序化技术,使其能够考虑分层分解约束。我们在IPC 2023偏序HTN基准测试上评估了所提出的方法,并将其与Optiplan——一种直接生成偏序规划的HTN规划器——进行了比较。结果表明,我们的两种实现都大幅减少了顺序约束的数量。尽管我们还观察到关键路径长度的减少,但改进效果不太明显。
cs.AI / 44 / 2609.03920

Value-Preserving Architectures for Agentic AI Systems

面向智能体AI系统的价值保持架构
Pesare, Alessandro, Dolci, Tommaso, Hose, Katja, Sallinger, Emanuel
Abstract
The emergence of agentic AI and LLM-based multi-agent systems (MAS) presents unprecedented opportunities for automating complex tasks, while simultaneously raising critical concerns about the preservation of fundamental human-centered values, such as privacy, fairness, and safety. Although software engineering has traditionally focused on functional correctness, the adoption of LLMs and AI agents into complex socio-technical systems has intensified the need for responsible software engineering and robust value alignment. In MAS, architectural design decisions, such as coordination mechanisms, communication protocols, and system topologies, play a central role in shaping system behavior and the outcomes they produce. This paper argues that architectural choices influence not only the functionality and performance of MAS but can also promote value-oriented system behavior. Therefore, we investigate how different architectural designs support different human-centered values, discussing the following value-preserving architectural patterns: (i) a privacy-aware architecture with a federated topology, (ii) a distributed architecture to promote pluralism and diversity, and (iii) a guard-agent architecture to detect and mitigate unfairness. Finally, we introduce representative use cases to illustrate the proposed architectures in real-world scenarios. By linking architectural design with human-centered values, this work lays the foundation for a unified set of architectural patterns and guidelines towards the design of trustworthy MAS.
Chinese Translation
智能体AI(Agentic AI)和基于大语言模型(LLM)的多智能体系统(MAS)的出现为自动化复杂任务提供了前所未有的机遇,同时也引发了对隐私、公平性和安全等基本以人为中心的价值观能否得以保持的重大关切。尽管软件工程传统上侧重于功能正确性,但LLM和AI智能体在复杂社会技术系统中的采用加剧了对负责任软件工程和稳健价值对齐的需求。在多智能体系统中,协调机制、通信协议和系统拓扑等架构设计决策在塑造系统行为及其产生的结果方面发挥着核心作用。本文认为,架构选择不仅影响MAS的功能和性能,还可以促进面向价值的系统行为。因此,我们研究了不同的架构设计如何支持不同的以人为中心的价值观,讨论了以下价值保持架构模式:(i)采用联邦拓扑的隐私感知架构,(ii)促进多元性和多样性的分布式架构,以及(iii)用于检测和缓解不公平性的守卫智能体(guard-agent)架构。最后,我们引入代表性用例来说明所提出的架构在现实场景中的应用。通过将架构设计与以人为中心的价值观相联系,这项工作为一套统一的架构模式和指南奠定了基础,以迈向可信赖MAS的设计。
cs.AI / 45 / 2609.03923

Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting

替我发言:赋予大语言模型参与会议的情境感知能力
Khan, Muneeb, Kirstein, Frederic, Ruas, Terry, Gipp, Bela
Abstract
In online meeting delegation, LLM agents fail to recognize when to speak. With no structured way to track stances, coverage, and floor, they miss the moments where they should contribute. Prompt-only delegates stay silent on 51.4% of the absent participant's talking opportunities on the AMI corpus. We present CAPA (Collaborative Agent Predictive Architecture), an architecture for online meeting delegation. A Perceiver updates the meeting state from each observed turn. A Predictor forecasts how the conversation will continue. A Controller decides whether to speak and which proposition to surface. A Generator phrases the chosen contribution in the participant's style. Two judges score the forecast and the action against the next observed turn. A Recalibrator updates the meeting state from those verdicts for future decisions. To evaluate online delegation, we introduce an episode-level protocol that scores whether, when, and what a delegate contributes around the participant's actual idea units. The protocol's schema-constrained LLM judges align with human annotations at Cohen's kappa = 0.71. On 137 AMI meetings, CAPA reduces the silence rate from 51.4% to 2.5%, doubles credited recovery (26.1 --> 52.2), and keeps hallucination at 0.6%. The failure mode shifts from omission to selection, with each residual near-miss attributable to a specific module of the architecture. Mechanism ablations identify the meeting state as the lever that closes the recognition gap, where raw-context scaling alone does not.
Chinese Translation
在在线会议代理场景中,大语言模型(LLM)智能体无法识别何时该发言。由于缺乏结构化的方式来追踪立场、议题覆盖情况和发言权,它们会错过本应发言的时机。仅使用提示词的代理在AMI语料库上,对缺席参与者的发言机会有51.4%保持沉默。我们提出了CAPA(Collaborative Agent Predictive Architecture,协作式智能体预测架构),一种面向在线会议代理的架构。感知器(Perceiver)根据每个观察到的发言轮次更新会议状态;预测器(Predictor)预测对话将如何继续;控制器(Controller)决定是否发言以及呈现哪个命题;生成器(Generator)以参与者的风格表述所选内容。两个评判器(Judge)根据下一个观察到的发言轮次对预测和动作进行评分;再校准器(Recalibrator)根据这些判定结果更新会议状态,以供后续决策使用。为评估在线代理表现,我们引入了一种情节级(episode-level)评估协议,围绕参与者的实际想法单元(idea units)对代理是否发言、何时发言以及发言内容进行评分。该协议采用模式约束的LLM评判器,与人工标注的一致性达到Cohen's kappa = 0.71。在137场AMI会议中,CAPA将沉默率从51.4%降至2.5%,使被认可的补救次数翻倍(26.1 --> 52.2),并将幻觉率保持在0.6%。失败模式从遗漏转变为选择错误,每个残余的近似失误都可归因于架构中的特定模块。机制消融实验表明,会议状态是弥补识别差距的关键杠杆,而单纯扩展原始上下文则无法做到这一点。
cs.AI / 46 / 2609.03938

Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding

基于SMT的HTN-SAT编码实现数值型全序HTN规划
Quenard, Gaspard, Togarepi, Takudzwa, Pellier, Damien, Fiorino, Humbert
Abstract
While HTN planning has received significant attention in recent years, support for numerical reasoning remains very limited. In this paper, we investigate numerical Totally-Ordered HTN (TOHTN) planning and show how standard SAT-based encodings can be naturally extended with SMT to handle numeric fluents. In addition, we introduce a benchmark suite for numerical TOHTN planning, providing a first common basis for evaluation in this setting. Experimental results show that this simple encoding already constitutes a competitive baseline. This work opens the way to more expressive approaches to HTN planning.
Chinese Translation
尽管HTN规划近年来受到了广泛关注,但对数值推理的支持仍然非常有限。本文研究了数值型全序HTN(TOHTN)规划,并展示了如何通过SMT自然地扩展标准的基于SAT的编码,以处理数值型状态变量(numeric fluents)。此外,我们引入了一个面向数值型TOHTN规划的基准测试集,为该领域的评估提供了首个公共基础。实验结果表明,这种简单的编码已经构成了一个具有竞争力的基线。这项工作为更具表达能力的HTN规划方法开辟了道路。
cs.AI / 47 / 2609.03943

More Criticism Does Not Make a Better Review: EquiReview-R

更多的批评并不能带来更好的评审:EquiReview-R
Zhang, Zexing, Li, Jichao, Lei, Tianyang, Fu, Yude, Kewei, Yang
Abstract
AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better review. A review may miss a consequential weakness or retain an allegation that available evidence does not support. These failures require opposite corrections, yet generation-oriented systems and aggregate measures obscure the distinction. We therefore recast AI-assisted review as evidence-guided refinement of a structured concern set, with omission and overcritique treated as separate risks. Building on this formulation, we introduce EquiReview-R, which resolves existing concerns against localized evidence, searches for missing issues from independent and review-conditioned perspectives, and returns stop, continue, or defer. To expose the failure mode that motivates this design, we construct an evidence-linked trajectory corpus. Its retrospective analysis shows why revision must precede further search: nearly all concerns in a high-recall review lack a definitive evidential disposition, while an earlier refinement mechanism cannot revise them. On a frozen cohort of previously unseen papers, EquiReview-R satisfies the prespecified non-inferiority criterion for major omission, reduces major overcritique from 15.5% to 8.1%, and attains a one-sided omission upper bound of 9.9% while stopping on 52.4% of papers. Computation-matched controls, controlled pairs, and ablations show that the gain comes from revision rather than extra inference or shorter output. We release the corpus as ReviewTrace, an evidence-linked resource for studying review revision, disagreement, and provenance.
Chinese Translation
AI 评审现已能够产出大量具体的批评意见,但更多的批评并不一定意味着更好的评审。一份评审可能遗漏某个重要的弱点,也可能保留一条现有证据并不支持的指控。这两类失败需要相反的纠正措施,然而面向生成的系统和聚合式的度量指标掩盖了这一区别。因此,我们将 AI 辅助评审重新构建为基于证据引导的结构化问题集精炼过程,并将遗漏(omission)与过度批评(overcritique)视为两种独立的风险。基于这一框架,我们提出了 EquiReview-R,它针对局部化证据解决已有问题,从独立视角和以评审为条件的视角搜索可能缺失的问题,并返回停止(stop)、继续(continue)或搁置(defer)的决策。为了暴露激发这一设计的失败模式,我们构建了一个证据关联的轨迹语料库。其回顾性分析表明了为什么必须先进行修订再进一步搜索:高召回率评审中几乎所有问题都缺乏明确的证据处置结论,而更早的精炼机制无法对其进行修订。在一个由此前未见论文组成的冻结队列上,EquiReview-R 满足了针对重大遗漏预设的非劣效性标准,将重大过度批评从 15.5% 降至 8.1%,并在 52.4% 的论文上停止的同时,将遗漏的单侧上界控制在 9.9%。计算量匹配的对照组、受控配对实验以及消融实验表明,性能提升来自修订机制本身,而非额外的推理计算或更短的输出。我们以 ReviewTrace 之名发布该语料库,作为一个证据关联的资源,用于研究评审修订、分歧与来源追溯。
cs.AI / 48 / 2609.03960

FiMI Banking: A Sovereign Model for Indian Retail Banking

FiMI Banking:一个面向印度零售银行的主权模型
NPCI AI Research Team, Kumar, Aman, Desai, Asit, Bhushan, Chandra, Sharma, Harsh, Bhushan, Harshit, Kadam, Hrithik, Doshi, Keyur, Kapardheeswar, Kolisetty Sai, Adhikary, Krishanu, Shaik, Nadeem, Prakash, Navya, Kukreja, Nitin, Devadiga, Prashant, MH, Shamanth, Pandey, Shantanu, Paul, Suvradip, Dedhia, Yatharth
Abstract
Banks need conversational systems that can answer product questions, assist customers with account-related requests, and operate safely within strict operational and regulatory constraints. General-purpose language models do not reliably meet these requirements. They fall short when a task requires grounded information, correct tool use, or cautious handling of bank-specific sensitive situations. We introduce FiMI Banking, a controlled Indian retail-banking setting. We build it from vetted banking documents, structured ground truth, synthetic customer backgrounds, and banking tools. We evaluate two post-training approaches: preference optimization for response-level behavior, and reinforcement learning with verifiable rewards for multi-turn tool-use tasks. Preference optimization improves safe behavior substantially: out-of-scope refusal rises from 52% to 80%. Reinforcement learning improves edge-case performance from 0.509 to 0.718 and order-sensitive task performance from 0.590 to 0.679, while using 29% fewer generated tokens. These results show that preference optimization and verifiable-reward reinforcement learning address complementary requirements for reliable banking agents.
Chinese Translation
银行需要能够回答产品问题、协助客户处理与账户相关的请求,并在严格的运营和监管约束下安全运行的对话系统。通用语言模型无法可靠地满足这些要求。当任务需要基于事实的信息、正确的工具使用或谨慎处理银行特定的敏感情况时,它们表现不佳。我们提出了FiMI Banking,一个受控的印度零售银行环境。我们基于经过审核的银行文档、结构化标准答案、合成的客户背景以及银行工具构建了该环境。我们评估了两种后训练方法:用于响应级行为的偏好优化(preference optimization),以及用于多轮工具使用任务的基于可验证奖励的强化学习(reinforcement learning with verifiable rewards)。偏好优化显著改善了安全行为:超出范围请求的拒绝率从52%提升至80%。强化学习将边缘情况表现从0.509提升至0.718,将顺序敏感任务表现从0.590提升至0.679,同时生成的token数量减少了29%。这些结果表明,偏好优化和可验证奖励强化学习能够满足构建可靠银行智能体的互补性需求。
cs.AI / 49 / 2609.03966

Interface-Induced Trajectory Censoring

接口引起的轨迹删失
Wang, Wenbo
Abstract
Agent evaluations report a tool-call rate read off the serving stack. That number can be zero while the model is emitting well-formed calls: the interface censors the trajectory before anything downstream sees it. On BFCL v4's own data, executor and scorer, holding weights, cases, decoding and seeds fixed and changing only the serving adapter, the same model scores 0.00 or 0.96 / 0.19. A 2x2 over chat template and parser locates the effect exactly: both main effects are exactly zero and all of it sits in the interaction -- no component is defective, and repairing one side of the contract buys precisely nothing. On tau-bench's 115 interactive retail tasks the same swap moves server-parsed calls from 0 to 636 and tasks reaching any tool execution from 0 to 103. Our probe reproduces the funnel across a 21x scale range of Qwen2.5-Coder: the server parses 0/100 at every size while well-formed emitted calls rise to 80/100 at 32B (~72 after calibration against an adjudicated gold standard). Under a matched envelope, across a comparable scale span, the silent fraction stays at 0-2, a prediction committed to the repository before the run. Llama-3.1-8B's 23% rate of calling the task function itself as a tool falls to 0 under one strict:true flag. The mismatch reaches inside the training loop, and its consequence is scale-dependent: in verl's AgentLoop at 7B, 45 of 115 generations carry a complete call; 0 are accepted, 0 execute, 0 return an observation. At 1.5B the same zero is over-determined, so we report the two scales separately. At evaluation time, repairing the adapter restores the mechanism but not a significant outcome gain: parsing 0->84, rescues 0->9, pass rate 53->62 (n.s.). We release a 98-line preflight check that catches every silent failure here. The observed tool-call rate is not a property of the model alone; it is a property of the model-interface stack that measures it.
Chinese Translation
智能体评估所报告的工具调用率是直接从服务栈中读取的。即使模型正在生成格式良好的调用,该数字也可能为零:接口在下游任何环节看到之前就将轨迹删失了。在BFCL v4自身的数据上,使用相同的执行器与评分器,在权重、用例、解码方式与随机种子均保持不变、仅更换服务适配器的情况下,同一模型的得分可为0.00或0.96/0.19。一个针对聊天模板与解析器的2x2因子实验精确定位了这一效应:两个主效应均为零,全部效应集中于交互项——没有任何单一组件存在缺陷,而单独修复契约的任何一方都毫无收益。在tau-bench的115个交互式零售任务上,同样的更换使服务器解析到的调用从0增至636,成功执行至少一次工具调用的任务数从0增至103。我们的探针在Qwen2.5-Coder的21倍规模范围内复现了这一漏斗效应:服务器在所有规模下均解析出0/100,而模型生成的格式良好调用在32B时升至80/100(对照经裁定金标准校准后约为72)。在匹配的实验包络与可比的规模跨度下,静默失败的比例保持在0-2,该预测在运行前已提交至代码库。Llama-3.1-8B以23%的比率将任务函数本身作为工具调用,而在一个strict:true标志下该比率降至0。这种失配还渗透到训练循环内部,其后果依赖于规模:在verl的AgentLoop中,7B模型115次生成中有45次携带完整调用,但被接受0次、执行0次、返回观测0次。在1.5B时该零值是多因素共同决定的,因此我们分别报告两个规模的结果。在评估阶段,修复适配器可恢复机制本身,但不带来显著的结果提升:解析0->84,挽回0->9,通过率53->62(不显著)。我们发布了一个98行的预检程序,可捕获此处所有的静默失败。观测到的工具调用率并非模型本身的属性,而是测量它的模型-接口栈的属性。
cs.AI / 50 / 2609.03973

Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing

用于反事实图像审计的公共见证证书与尖锐特征界
Faghihi, Usef, Saki, Amir
Abstract
An image editor may satisfy every regional plausibility constraint separately even when no single latent explanation fits the complete output. We formalize this local-to-global failure using a common witness grade and witness nerve. The framework separates auditing from causal identification: shared exogeneity alone allows every coupling of the regime marginals, whereas an externally justified witness relation yields sharp partial-identification bounds for prespecified image features. Helly-type arguments provide short incompatibility certificates for quasiconvex losses, heterogeneous action strata, and finite witness atlases; a blocker-hypergraph formula gives exact repair counts. Simultaneous confidence regions for the regime marginals give finite-sample outer coverage of the complete identified interval. Controlled MNIST, Morpho-MNIST, and smallNORB studies demonstrate the predicted local-global separation, while synthetic experiments test sharp bounds, certificate recovery, and structured computation. The method audits a declared feature relation and does not identify unrestricted pixel-level counterfactuals.
Chinese Translation
一个图像编辑器可能分别满足每一条区域合理性约束,却不存在任何单一的潜在解释能够拟合完整的输出。我们利用公共见证等级(common witness grade)与见证神经(witness nerve)对这种从局部到全局的失效现象进行形式化。该框架将审计与因果识别分离开来:仅有共同的外生性允许机制边际(regime marginals)之间的任意耦合,而一个具有外部正当性的见证关系则为预先指定的图像特征给出尖锐的部分识别界。Helly 型论证为准凸损失、异质动作分层以及有限的见证图册提供了简短的不相容性证书;一个阻断子超图(blocker-hypergraph)公式给出精确的修复计数。机制边际的同步置信区域为完整的已识别区间提供有限样本的外覆盖。在受控的 MNIST、Morpho-MNIST 和 smallNORB 实验中,我们验证了所预测的局部-全局分离现象;同时,合成实验检验了尖锐界、证书恢复以及结构化计算。该方法审计的是已声明的特征关系,并不能识别无约束的像素级反事实。
cs.AI / 51 / 2609.04005

The Dually Flat Geometry of Planning as Inference

作为推断的规划的对偶平坦几何
Milosevic, Nikola, Kataoka, Asaki, Hinrichs, Nicolas, Doya, Kenji, Scherf, Nico
Abstract
We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by embedding the planning criterion into the dynamics through a resetting planning process. Its stationary measure, which we term visitation measure, is the object on which the information geometry of decision making is most naturally expressed. The achievable visitation measures form a dually flat statistical manifold whose two affine charts are the visitation probabilities and the log-policies, dual under the conditional entropy. This structure makes planning-as-inference generalize from linear rewards to nonlinear functionals of the visitation, each iterate solved by one natural-gradient step, and gives the temporal-difference error the interpretation of a marginal-utility estimate. We develop the geometry and its consequences for reinforcement learning and theoretical neuroscience.
Chinese Translation
我们提出了强化学习中占用度量的另一种刻画方式,其做法是通过一个带重置的规划过程将规划准则嵌入到动力学之中。该过程的平稳测度——我们称之为访问测度(visitation measure)——是决策的信息几何最自然得以表达的对象。可达的访问测度构成一个对偶平坦的统计流形,其两个仿射坐标图分别为访问概率与对数策略(log-policy),二者在条件熵下互为对偶。这一结构使得“作为推断的规划”(planning-as-inference)能够从线性奖励推广到访问测度的非线性泛函,其中每次迭代由一步自然梯度求解,并赋予时序差分误差一个边际效用估计的解释。我们发展了这一几何结构及其对强化学习与理论神经科学的意义。
cs.AI / 52 / 2609.04013

LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening

LLM4CKD:用于慢性肾病早期筛查的大语言模型
Kabir, Muhammad Ashad, Munira, Sirajam
Abstract
Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine learning (ML) and deep learning (DL) approaches require labeled data and model training, limiting their use in real-world screening settings. This study evaluates the effectiveness of large language models (LLMs) for CKD screening under zero-shot and few-shot in-context learning settings and compares them with traditional ML and DL methods. We propose a framework that uses clinically selected tabular features and structured prompt templates to enable LLM-based inference without task-specific training. LLM performance is evaluated across multiple prompt styles, feature configurations, and data settings, and compared with standard ML, DL, and tabular foundation model (TFM) baselines, and existing CKD screening tools. The results show that LLMs can achieve competitive performance using only a small number of examples, often matching or outperforming traditional approaches in low-data settings. However, their performance remains model-dependent and less stable as input complexity increases. In contrast, ML, DL, and TFM models show more consistent improvement with larger training data. Overall, the findings highlight a trade-off between data efficiency and stability, suggesting that LLMs may serve as a flexible complementary approach for CKD screening when labeled data are limited.
Chinese Translation
慢性肾病(CKD)的早期筛查对于及时干预至关重要,然而大多数机器学习(ML)和深度学习(DL)方法需要标注数据和模型训练,限制了它们在真实筛查场景中的应用。本研究评估了大语言模型(LLM)在零样本(zero-shot)和少样本上下文学习(few-shot in-context learning)设置下进行CKD筛查的有效性,并将其与传统ML和DL方法进行比较。我们提出了一个框架,该框架利用经临床筛选的表格特征和结构化提示模板,使基于LLM的推理无需针对特定任务的训练。我们在多种提示风格、特征配置和数据设置下评估LLM的性能,并与标准ML、DL和表格基础模型(TFM)基线以及现有的CKD筛查工具进行比较。结果表明,LLM仅使用少量示例即可取得具有竞争力的性能,在低数据设置中往往能够媲美甚至超越传统方法。然而,其性能仍然依赖于具体模型,且随着输入复杂度的增加而变得不够稳定。相比之下,ML、DL和TFM模型在训练数据增多时表现出更为一致的改进。总体而言,研究结果揭示了数据效率与稳定性之间的权衡,表明在标注数据有限的情况下,LLM可以作为CKD筛查的一种灵活的补充方法。
cs.AI / 53 / 2609.04014

InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models

InSituMeasure:利用多模态大语言模型探究工业场景中的情境化测量落地能力
Shen, Chao, Li, Xinyuan, Zhou, Yunfan, Yao, Jianguo, Guan, Haibing, Wang, Zhihai, Li, Xijun
Abstract
For trained operators, gauge reading requires little specialized knowledge, low cognitive effort, and high repeatability. Yet Multimodal Large Language Models (MLLMs) remain unreliable in continuous-valued measurement despite strong results on general multimodal benchmarks. Existing benchmarks expose this weakness but isolate measurement from realistic, knowledge-grounded settings, with limited situated context, specialized instruments, real-world noise, and matched diagnostic annotations, reducing realism and constraining root-cause analysis. We introduce InSituMeasure to evaluate situated measurement grounding. It contains 2,922 real industrial monitoring scenes across eight functional categories of professional engineering instruments, with dense gauge-attribute annotations and noise tags for failure diagnosis. We define metrics for numerical accuracy under predefined tolerances and unit consistency, rejection of fake or unanswerable tasks, and alignment between model failures and annotated error factors. Across 24 state-of-the-art MLLMs, the best model reaches only 25.7\% joint value-unit accuracy and 51.8\% confidence-diagnosis F1, revealing a substantial gap between general multimodal competence and reliable situated measurement. Further analysis identifies failures from text-induced shortcuts, overconfident responses, and authentic industrial noise, including mixed disturbances, viewpoint deviation, occlusion, and environmental interference.
Chinese Translation
对于训练有素的操作员而言,仪表读数几乎不需要专门知识,认知负担低且重复性高。然而,尽管多模态大语言模型(MLLMs)在通用多模态基准测试中表现出色,其在连续值测量方面仍不可靠。现有基准测试虽然揭示了这一弱点,但将测量与现实的知识驱动场景割裂开来,缺乏情境化上下文、专业仪器、真实世界噪声以及匹配的诊断标注,降低了真实性并限制了根因分析。我们提出InSituMeasure以评估情境化测量落地能力。该基准包含2,922个真实工业监测场景,涵盖八大类专业工程仪器,并提供了密集的仪表属性标注以及用于故障诊断的噪声标签。我们定义了多项评估指标,包括在预设容差和单位一致性下的数值准确性、对虚假或无法回答任务的拒答能力,以及模型失败与标注错误因素之间的对齐程度。在24个最先进的MLLM上,表现最佳的模型仅达到25.7%的数值-单位联合准确率和51.8%的置信度诊断F1值,揭示了通用多模态能力与可靠情境化测量之间的巨大差距。进一步分析发现,失败源于文本诱导的捷径、过度自信的回答以及真实工业噪声,包括混合干扰、视角偏移、遮挡和环境干扰。
cs.AI / 54 / 2609.04021

FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models

FLY-EVAL++:一种面向大语言模型安全约束飞行预测的证据驱动评估协议
Wu, Yalun, Fang, Junfeng, Wang, Jiawei, Liu, Haotian, Yang, Qijun, Yang, Minghan, Guo, Hongcheng, Li, Zhoujun, Wang, Boyang
Abstract
Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close to the ground truth can still violate operational constraints, combine fields in physically inconsistent ways, or fail to produce usable structured outputs. Existing evaluation protocols do not measure these failure modes reliably. We propose FLY-EVAL++, an evidence-driven evaluation protocol that combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with fixed rubric-guided aggregation into interpretable multi-dimensional scores. We instantiate FLY-EVAL++ for Flight Trajectory and Attitude Prediction (FTAP) by extending the PilotBench setting with history-conditioned and multi-step prediction tasks. Across 66 LLMs, safety compliance is the most discriminative dimension of model behavior: models with comparable predictive performance differ by more than 28 points in safety score, and we observe recurrent failures including safety violations under physically plausible predictions and instability in multi-step rollouts. These results show that evaluation in safety-critical domains should measure constraint satisfaction and structured validity explicitly rather than rely on accuracy-centric reporting alone.
Chinese Translation
在安全关键且受物理规律约束的环境中评估大语言模型(LLM),仅依靠基于准确率的指标是不够的,因为数值上接近真实值的预测仍然可能违反运行约束、以物理上不一致的方式组合多个场量,或无法生成可用的结构化输出。现有评估协议无法可靠地衡量这些失败模式。我们提出FLY-EVAL++,一种证据驱动的评估协议,它将对协议符合性、物理可行性和安全约束的确定性验证与基于固定评分标准的聚合相结合,生成可解释的多维度分数。我们在飞行轨迹与姿态预测(Flight Trajectory and Attitude Prediction, FTAP)任务上实例化了FLY-EVAL++,通过扩展PilotBench设置,加入基于历史条件的预测和多步预测任务。在对66个大语言模型的评估中,安全合规性是区分模型行为最显著的维度:预测性能相近的模型在安全分数上相差超过28分,并且我们观察到多种反复出现的失败模式,包括在物理上看似合理的预测下出现安全违规,以及多步推演中的不稳定性。这些结果表明,在安全关键领域的评估应显式衡量约束满足性与结构有效性,而不能仅依赖以准确率为中心的报告。
cs.AI / 55 / 2609.04024

Instruction Duplication as an Inference-Time Control Primitive

指令复制作为一种推理时控制原语
Lavrenko, Victor
Abstract
Procedural instruction following is a basic requirement for controllable language-model systems, especially when generated trajectories are inspected or repaired downstream. We introduce instruction duplication, a minimal black-box inference-time control that repeats only the procedural instruction, without retraining or decoding changes. Across seven instruction-tuned models, 300 medical multiple-choice questions, eight placement conditions, and 16,800 scheduled generations, moving from one to two copies raises the deterministic All-8 diagnostic--responses passing all eight observable tests--from 90.22% to 93.17% (+2.95 percentage points), eliminating 30.2% of the failures remaining after one copy. Pre-provisional TF-IDF recall rises from 73.44% to 74.81% (+1.38 points; Holm-adjusted p < .001), while final-answer accuracy remains exactly 60.21%. Premature commitment increases from 1.52% to 2.30% (p_Holm = .00536). A blinded challenge audit yields 10/30 directional confirmations, 20/30 perceptual ties, and no reversals; its prespecified 28/30 confirmation criterion is not met. Yet this distinction can matter operationally when a downstream system acts on the generated trajectory. In Answer Engineering (AE), where explicit trajectory state determines local repair, the published reason-first no-editing SSNHL endpoint was 25.1%; system-only AE was later reproduced at 84.2%, and the same trailing duplicate raised it to 97.1%. For conductive diagnostic branch preservation, the corresponding values are 58.9% published without editing, 78.6% with reproduced AE, and 73.8% with AE plus duplication--a within-AE decrease, but still 14.9 points above the no-editing baseline. Instruction duplication is therefore a low-complexity, placement-sensitive control whose practical value can emerge through the downstream system that consumes the exposed trajectory.
Chinese Translation
程序性指令遵循是可控语言模型系统的基本要求,尤其是当生成的轨迹需要在下游被检查或修复时。我们提出指令复制(instruction duplication),这是一种最小化的黑盒推理时控制方法,仅对程序性指令进行重复,无需重新训练或更改解码方式。在七个经过指令微调的模型、300道医学选择题、八种放置位置条件以及16,800次按计划执行的生成任务中,将指令从一份复制为两份,可使确定性All-8诊断指标——即通过全部八项可观测测试的响应——从90.22%提升至93.17%(提升2.95个百分点),消除了单份复制后剩余失败的30.2%。临时性TF-IDF召回率从73.44%上升至74.81%(提升1.38个百分点;经Holm校正后 p < .001),而最终答案的准确率则恰好保持为60.21%不变。过早承诺的比例从1.52%上升至2.30%(p_Holm = .00536)。一项盲法挑战性审计得到10/30的方向性确认、20/30的感知性平局,且无任何反转;但其预设的28/30确认标准未能达到。然而,当下游系统基于所生成的轨迹采取行动时,这种区分在操作层面可能至关重要。在答案工程(Answer Engineering, AE)中,显式的轨迹状态决定了局部修复:已发表的、采用理由优先且不编辑策略的SSNHL终点为25.1%;仅使用系统级AE随后被复现达到84.2%,而加入末尾重复指令后将其提升至97.1%。对于传导性听力的诊断分支保留,对应数值分别为:不编辑的已发表结果58.9%,复现AE后为78.6%,AE加指令复制为73.8%——即AE内部有所下降,但仍比不编辑基线高出14.9个百分点。因此,指令复制是一种低复杂度、对放置位置敏感的控制手段,其实际价值可通过消费所暴露轨迹的下游系统得以体现。
cs.AI / 56 / 2609.04030

IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations

IRWOZ 2.0:一个由大语言模型驱动的工业机器人对话数据集
Li, Chen, Chrysostomou, Dimitrios
Abstract
IRWOZ has improved industrial human-robot interaction (HRI) dialogue systems through domain-specific annotations. However, its initial version contains substantial noise in dialogue states and utterances, limiting state-tracking accuracy. We introduce IRWOZ 2.0, which addresses these limitations through large language model (LLM) enhanced generation (Mistral/Claude-3.5) and quality refinements. Our improved dataset expands to 390 dialogues across 4 industrial domains (Assembly, Delivery, Position, Relocation), featuring manual corrections and automated typo removal. Benchmark experiments on dialogue state tracking demonstrate significant improvements, with GPT-2's BLEU-4 score increasing from 0.1651 to 0.5604 compared to original IRWOZ. To support industrial HRI research, we publicly released IRWOZ 2.0 dataset at https://ieee-dataport.org/documents/irwoz-20-large-language-model-driven-dialogue-dataset-industrial-robot-conversations
Chinese Translation
IRWOZ 通过特定领域的标注改进了工业人机交互(HRI)对话系统。然而,其初始版本在对话状态和话语中包含大量噪声,限制了状态追踪的准确性。我们提出了 IRWOZ 2.0,通过大语言模型(LLM)增强生成(Mistral/Claude-3.5)和质量优化来解决这些局限性。改进后的数据集扩展至涵盖 4 个工业领域(装配、配送、定位、搬运)的 390 段对话,并进行了人工修正和自动拼写错误清除。在对话状态追踪上的基准实验表明了显著的性能提升,与原始 IRWOZ 相比,GPT-2 的 BLEU-4 分数从 0.1651 提高到 0.5604。为支持工业 HRI 研究,我们已在 https://ieee-dataport.org/documents/irwoz-20-large-language-model-driven-dialogue-dataset-industrial-robot-conversations 公开发布了 IRWOZ 2.0 数据集。
cs.AI / 57 / 2609.04063

Spurious Advantage Hidden in GRPO

隐藏在GRPO中的虚假优势
Wang, Jiamian, Basu, Samyadeep, Goswami, Koustava, Yu, Tong, Tao, Zhiqiang
Abstract
Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning with verifiable rewards, where its advantage estimator assigns each rollout a magnitude from within-group reward statistics. In the common case, this magnitude rewards rollouts that reach the correct answer through reasoning. Yet, an overlooked case shares the same surface: a rollout may land on it by guessing, and the formula still assigns a high magnitude, which we identify as the spurious advantage. This arises in three cases: bounded-answer tasks with a small candidate set; open-answer sets hosting bounded sub-cases; and search agents whose budget opens many paths to the same answer. In all three, this misleads the policy toward guess-like behaviors. We propose SIGNBALANCE, whose magnitude is composition-free: it keeps the verifier sign, uses a global scale, and restores zero-mean balance via a stop-gradient per-class rescaling. Across math and search agent benchmarks at different scales, SIGNBALANCE matches GRPO on open-answer math and improves on bounded-answer math and search agents. Code will be released.
Chinese Translation
组相对策略优化(GRPO)被广泛用于可验证奖励的强化学习,其优势估计器根据组内奖励统计量为每条rollout分配一个幅值。在常见情况下,该幅值奖励的是通过推理得到正确答案的rollout。然而,一个被忽视的情况在表面上与之相同:一条rollout可能通过猜测得到正确答案,而该公式仍会赋予其较高的幅值,我们将其识别为虚假优势(spurious advantage)。这种情况出现在三类场景中:候选集较小的有界答案任务;包含有界子问题的开放式答案集合;以及预算开辟了多条路径通向同一答案的搜索智能体。在这三类场景中,这种偏差都会误导策略趋向猜测式行为。我们提出SIGNBALANCE,其幅值与答案构成无关:它保留验证器的符号,采用全局尺度,并通过基于类别的逐类重缩放的停止梯度(stop-gradient)机制恢复零均值平衡。在不同规模的数学与搜索智能体基准测试中,SIGNBALANCE在开放式数学任务上与GRPO表现相当,并在有界答案数学任务和搜索智能体上有所提升。代码将被开源。
cs.AI / 58 / 2609.04094

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

DRACO:基于动态评分标准的细粒度信用分配用于长程智能体训练
Gandhi, Shubham, Goyal, Saurabh, Kate, Kiran, Rizk, Yara
Abstract
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popular way to supply such a reward; they are scored once per trajectory, but a single scalar is a poor signal across tens of steps. We propose DRACO: Distributing Rubric-based Advantage for Credit Optimization. It generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per completed trajectory, and redistributes that judgment over the steps responsible for annotated rubrics to produce differentiated per-step advantages in GRPO. The redistribution is closed-form and does not introduce any trained attribution module. On AppWorld, DRACO gains 15.9 points over the base model and 5.3 points over GRPO trained with a sparse ground-truth reward, despite not using any verifiers itself. On out-of-domain Tau-Bench, it gains 5.3 points over the base model even without a frontier judge, beating both ground-truth-reward training and other rubric-based training settings. The code for DRACO is available at https://github.com/IBM/draco.
Chinese Translation
当任务具备程序化检查器时,基于可验证奖励的强化学习(Reinforcement Learning from Verifiable Rewards)效果良好,但大多数长程智能体领域并不具备此类检查器。我们研究结果不可知(outcome-blind)的设定,即真实成功信号不可用的情况下。多标准评分标准(rubric)是提供此类奖励的常用方式;它们通常对每条轨迹只评分一次,但单一标量信号在数十个步骤上提供的信息十分有限。我们提出DRACO:基于评分标准的分布式优势信用优化(Distributing Rubric-based Advantage for Credit Optimization)。它在训练过程中动态生成评分标准以追踪策略不断演进的能力,对每条完成的轨迹只对这些标准评分一次,并将该评判重新分配到与所标注评分标准相关的步骤上,从而在GRPO中产生差异化的逐步优势。这种重新分配具有闭式解,且不引入任何经过训练的归因模块。在AppWorld上,DRACO相比基础模型提升15.9分,相比使用稀疏真实奖励训练的GRPO提升5.3分,而其自身并未使用任何验证器。在域外数据集Tau-Bench上,即使不使用前沿模型作为评判者,它也比基础模型提升5.3分,超越了真实奖励训练以及其他基于评分标准的训练设置。DRACO的代码可在 https://github.com/IBM/draco 获取。
cs.AI / 59 / 2609.04098

Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

为什么 Gated DeltaNet 能在4比特量化中幸存:混合架构27B大语言模型循环部分的NVFP4 W4A4量化
Kozyrev, Sergii, Maiboroda, Davyd
Abstract
Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent state summarizes the context in fixed size. Early community 4-bit quantizations of Qwen3.8-27B (48 GDN layers, 16 attention layers) left the GDN block in 8- or 16-bit precision -- especially its decay and write-strength gates -- on the intuition that errors in a recurrence accumulate over long contexts. We test that intuition by building Minima: NVFP4 W4A4 on all 496 linear layers, GDN included. Across perplexity at 4K/32K, MMLU-Pro, GSM8K, AIME'25, GPQA-Diamond, LiveCodeBench, and RULER retrieval to 64K, Minima matches BF16 within seed noise (5-task average -0.52) while being the smallest (17.5 GiB) and fastest-prefill (+14-19%) recipe we compare, and its 32K perplexity gap shrinks with position. A four-part mechanism study explains why: (i) NVFP4's 16-element block scaling localizes the residual stream's extreme outliers, equalizing activation error across layer roles; (ii) the supposedly fragile gate projections are the least sensitive -- softplus/exponential and sigmoid parameterizations compress ~11% GEMM error to ~2% output error; (iii) the delta-rule recurrence holds injected noise at a flat plateau over 32K tokens and forgets a state impulse within hundreds of steps, because each write overwrites the state along the current key direction; (iv) the per-token quantization cost washes out with context instead of compounding. We also repair a global-scale mismatch that arises when per-module-calibrated NVFP4 checkpoints are served by kernels that fuse those modules into one GEMM, and show calibrated FP8 KV-cache scales are performance-free. The result: a practical recipe -- quantize everything, ship KV scales -- and a mechanistic account of why the recurrent half of a hybrid LLM is the easy half to quantize. Checkpoint: https://huggingface.co/minima-ai/mnma_qwen3.8_27b_nvfp4
Chinese Translation
混合架构大语言模型将softmax注意力与线性注意力层(如Gated DeltaNet,GDN)相结合,后者的循环状态以固定大小概括上下文信息。社区早期对Qwen3.8-27B(48个GDN层、16个注意力层)的4比特量化将GDN模块保留在8比特或16比特精度——尤其是其衰减门控和写入强度门控——基于一种直觉:循环中的误差会在长上下文中不断累积。我们通过构建Minima来检验这一直觉:在包括GDN在内的全部496个线性层上采用NVFP4 W4A4量化。在4K/32K困惑度、MMLU-Pro、GSM8K、AIME'25、GPQA-Diamond、LiveCodeBench以及最高64K的RULER检索任务上,Minima与BF16的差距在随机种子噪声范围内(5项任务平均-0.52),同时是我们所比较方案中显存占用最小(17.5 GiB)且预填充速度最快(+14-19%)的方案,且其32K困惑度差距随位置推移而缩小。一项四部分机理研究解释了原因:(i) NVFP4的16元素分块缩放将残差流的极端离群值局部化,使激活误差在各层角色间均衡分布;(ii) 那些被认为脆弱的门控投影实际上敏感度最低——softplus/指数和sigmoid参数化将约11%的GEMM误差压缩至约2%的输出误差;(iii) delta规则循环在32K个token内将注入噪声维持在平坦平台上,并在数百步内遗忘状态脉冲,因为每次写入都会沿当前键方向覆写状态;(iv) 逐token的量化成本随上下文增长而消散,而非累积。我们还修复了当逐模块校准的NVFP4检查点由将这些模块融合为单个GEMM的内核服务时出现的全局缩放不匹配问题,并证明校准的FP8 KV缓存缩放因子不带来性能损失。最终结果:一个实用方案——量化一切、附带KV缩放因子——以及对混合架构大语言模型中循环部分为何更易量化的机理解释。检查点:https://huggingface.co/minima-ai/mnma_qwen3.8_27b_nvfp4
cs.AI / 60 / 2609.04127

Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable

大语言模型推荐的认识论依据:刻画真值不可得时的依赖基础
Vardi, Shai, Sedoc, João
Abstract
Large language models are increasingly used to support organizational decisions, yet users often lack a principled basis for assessing whether to rely on a specific recommendation. Existing approaches typically evaluate broad model properties, such as reliability, uncertainty, or robustness, or focus on user trust, rather than the underlying basis for relying on an individual recommendation. Adapting theoretical foundations from epistemology, we introduce epistemic warrant, a decision-level construct that characterizes the stability of a model's preference and the scope over which that preference holds. We operationalize this construct through a four-tier reliance certificate for pairwise recommendations, distinguishing among unstable, context-dependent, locally supported, and broadly supported recommendations. We validate the construct using contemporary methodologies: known-groups tests successfully recover expert-prespecified warrant orderings, and stronger warrants systematically align with independent consensus from crowd workers. Furthermore, we demonstrate that epistemic warrant provides information distinct from verbalized confidence and is not readily explained by decision difficulty. Ultimately, this framework offers a theoretically grounded, implementable approach for characterizing the warrant of individual LLM recommendations when objective ground truth is unavailable.
Chinese Translation
大语言模型日益被用于支持组织决策,但用户往往缺乏一个有原则的基础来判断是否应依赖某条具体推荐。现有方法通常评估模型的宏观属性(如可靠性、不确定性或鲁棒性),或聚焦于用户信任,而非依赖单条推荐的深层基础。本文借鉴认识论的理论基础,提出了“认识论依据”这一决策层面的构念,用以刻画模型偏好的稳定性及该偏好所适用的范围。我们通过面向成对比较推荐的四层“依赖证书”将这一构念操作化,区分不稳定型、情境依赖型、局部支持型和广泛支持型的推荐。我们采用当代方法验证了该构念:已知组检验成功恢复了专家预先设定的依据排序,且更强的依据与众包工作者的独立共识系统性地保持一致。此外,我们证明认识论依据提供了与口头表达置信度不同的信息,且无法简单地用决策难度来解释。最终,该框架为在客观真值不可得时刻画单条大语言模型推荐的依据提供了一种有理论支撑且可实施的方法。
cs.AI / 61 / 2609.04128

Environment Evolution for Terminal Agents

面向终端智能体的环境演化
Fan, Zhiyuan, Yu, Tinghao, Cai, Yuanjun, Zhou, Jiang, Guan, Jiangtao, Liu, Jincheng, Yang, Yun, Hu, Dingxin, Han, Zhuo, Wu, Xing, Zhang, Feng, Wang, Lilin
Abstract
Scaling interactive and verifiable environments is critical for training terminal agents. As frontier models become more capable, environments synthesized from scratch become less challenging and thus provide limited learning signals. Recent co-evolution methods iteratively synthesize environments near the model's learnable frontier based on weaknesses exposed during rollouts. However, their dependence on on-policy rollouts limits generalization and the continuous provision of learning signals as the model becomes stronger. In this paper, we propose environment evolution, which incrementally increases environment difficulty off-policy and schedules the evolved environments generation by generation during training to provide continuous learning signals. We derive three evolution directions that influence environment difficulty from the multi-turn learning objective and then implement evolution along these directions through a loop-engineered multi-agent harness. Quantitative rollout experiments with Hy4 preview, Claude Opus 5, and GPT-5.6 Sol show that environment evolution consistently produces more difficult environments. We validate its effectiveness on Qwen3.6-27B and Qwen3.6-35B-A3B through simple long-horizon RL training, improving their performance by 14.4 and 18.0 percentage points on Terminal-Bench 2.1, respectively.
Chinese Translation
扩展交互式且可验证的环境对于训练终端智能体至关重要。随着前沿模型能力的不断增强,从零合成的环境变得不再具有挑战性,因而所能提供的学习信号十分有限。近期的共同演化方法基于在 rollout 中暴露出的模型弱点,在模型可学习边界附近迭代合成环境。然而,这些方法依赖于同策略 rollout,这限制了其泛化能力,并且在模型变强后难以持续提供学习信号。本文提出环境演化方法,以离策略方式逐步提升环境难度,并在训练过程中按代调度演化后的环境,从而提供持续的学习信号。我们从多轮学习目标中推导出影响环境难度的三个演化方向,并通过一个经过循环工程化的多智能体框架沿这些方向实现演化。基于 Hy4 preview、Claude Opus 5 和 GPT-5.6 Sol 的定量 rollout 实验表明,环境演化能够持续生成更具难度的环境。我们通过简单的长时程强化学习训练,在 Qwen3.6-27B 和 Qwen3.6-35B-A3B 上验证了该方法的有效性,在 Terminal-Bench 2.1 上分别将两者性能提升了 14.4 和 18.0 个百分点。
cs.AI / 62 / 2609.04135

The Natural Language Interaction Protocol and Standard for AI Agents

面向AI智能体的自然语言交互协议与标准
Xing, Luyi, Topaloglu, Rasit Onur, Sinha, Ranjan, Ratnaparkhi, Abhay, Ndichu, Samuel, Nguyen, Christopher, Das, Anindita, Sheffler, Tom, Rahouti, Mohamed, Li, Zichuan, Liao, Xiaojing, Aiyagari, Sanjay
Abstract
AI agents are increasingly being developed and deployed across organizations using heterogeneous agent-development frameworks, AI models, tool interfaces, protocols, and execution environments. To realize their potential social and business impact, these agents must be able to interoperate through a common communication protocol. The Natural Language Interaction Protocol (NLIP), developed by researchers and practitioners across companies and universities and standardized by Ecma International, addresses this need by defining a standards-based application-layer protocol for AI-agent interaction. NLIP provides a lightweight semantic message envelope that can be carried over existing transports such as HTTP/HTTPS, WebSocket, and AMQP, while allowing NLIP-aware agents and gateways to adapt between clients, agents, local context stores, ontologies, tools, enterprise services, and heterogeneous underlying protocols. This paper presents the motivation and design rationale of NLIP, its message model and transport bindings, security-by-design considerations, reference implementation, representative applications, adoption signals, and relationship to emerging agent protocols such as MCP and A2A.
Chinese Translation
AI智能体正日益在各类组织中被开发和部署,其所依赖的智能体开发框架、AI模型、工具接口、协议和执行环境各不相同。为实现其潜在的社会和商业影响,这些智能体必须能够通过统一的通信协议实现互操作。自然语言交互协议(Natural Language Interaction Protocol, NLIP)由来自企业和高校的研究人员与实践者共同开发,并由Ecma国际(Ecma International)标准化,通过定义一个基于标准的应用层协议来满足AI智能体交互的这一需求。NLIP提供了一种轻量级的语义消息封装,可承载于HTTP/HTTPS、WebSocket和AMQP等现有传输协议之上,同时支持感知NLIP的智能体与网关在客户端、智能体、本地上下文存储、本体(ontology)、工具、企业服务以及异构底层协议之间进行适配。本文阐述了NLIP的设计动机与设计原理、其消息模型与传输绑定、内建安全的考量、参考实现、代表性应用、采纳信号,以及它与MCP和A2A等新兴智能体协议的关系。
cs.AI / 63 / 2609.04141

Efficient Test-Time Adaptation through Human-AI Interaction

通过人机交互实现高效的测试时自适应
Wang, Zora Zhiruo, Gandhi, Apurva, Shao, Rulin, Chen, Aspen, Mueller, Jonas, Liang, Zhiqi, Chen, Jett, Ryan, Michael, Ma, Qianou, He, Luxi, Cheng, Zhoujun, He, Andre, Kim, Seungone, Geng, Jiayi, Zheng, Mingqian, Sun, Weiwei, Zhang, Zheyuan, Zhao, Xinran, Wang, Yike, Hou, Abe, Jiang, Liwei, Koh, Pang Wei, Yang, Diyi, Neubig, Graham, Fried, Daniel
Abstract
AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent interaction surfaces criteria that users cannot fully specify up front, yet apply repeatedly across tasks. We argue this cross-session interaction data is a rich, underused signal for closing the gap to individual expertise. In this work, we propose test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module. We adapt agents to 30 individuals in two high-utility domains, writing and visual creation, on a total of 600 tasks. Our agents improve solo task success by 4.5-20.9% within only tens of tasks. Meanwhile, our evolving rubric module serves as a scalable annotation tool, creating evaluation rubrics that catch 16.0-22.3% more failures than those from LMs or humans alone. While agents are adapted towards individuals, we show these personalized agents also produce improvements in success of up to 8.8% that generalize across users.
Chinese Translation
AI智能体基于大规模人群数据进行训练,以编码涵盖众多从业者能力的广泛技能。然而,其产出的成果往往难以达到专业人士愿意为之担责的个人标准。在现实的开放性任务中,成功标准具有异质性且缺乏充分记录,个体专长恰恰体现在对平均水平的提升与超越。在实践中,迭代式的人机交互能够挖掘出用户无法事先完全明确、却在不同任务中反复应用的标准。我们认为,这种跨会话的交互数据是一种丰富但未被充分利用的信号,可用于弥合智能体与个体专长之间的差距。在这项工作中,我们提出了基于人机交互的测试时自适应(Test-Time Adaptation through Human-Agent Interaction, TAHI),它将这些信号融入智能体的上下文和权重中,并通过一个不断演化的评分准则模块(rubric module)将每位用户的训练与评估标准具体化。我们在两个高实用价值领域——写作与视觉创作——中将智能体适配至30位用户,共涉及600项任务。仅需数十项任务,我们的智能体即可将独立任务的成功率提升4.5%-20.9%。同时,我们的演化式评分准则模块还可作为可扩展的标注工具,其生成的评估准则能够比单独使用大语言模型(LM)或人工多捕捉16.0%-22.3%的失败案例。虽然智能体是针对个体进行适配的,我们发现这些个性化智能体还能带来最高达8.8%的成功率提升,并且可泛化至其他用户。
cs.AI / 64 / 2609.04148

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

Terminal-Universe:将智能体轨迹转化为可扩展的终端环境
Wu, Jie, Zhang, Zhenru, Zhang, Beichen, Wang, Xuwu, Su, Yuhui, Chen, Mouxiang, Wang, Peng, Wang, Zhihai, Shen, Que, Zhou, Hao, Yang, An, Huang, Fei, Yang, Yujiu, Liu, Dayiheng
Abstract
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution history in existing trajectories exposes the structure and contents of the environments in which they ran, making it possible to reconstruct those environments from the trajectories themselves. Thus, we introduce Terminal-Universe, a framework which turns each trajectory into a reusable environment and explores it for synthesizing new tasks and continued interactions. Specifically, Terminal-Universe replays the file operations recorded in a trajectory to restore each file before the agent modified it, yielding a partial workspace; a completion agent then supplies the missing files and dependencies. On this recovered workspace, we both reconstruct the original intent task and synthesize entirely new ones. Besides, we also scale the tasks along two complementary axes: breadth and depth. For breadth, we mine directional dependency relations between related environments and synthesize cross-workspace queries spanning multiple codebases, as developers routinely do in real-world development. For depth, we extend the initial single-turn query into a multi-round session that captures iterative user feedback and requirement refinement via a user agent. Applied to public terminal agent trajectories, Terminal-Universe produces 37.3k task-sufficient environments. Supervised fine-tuning of Qwen3.5-27B on this corpus improves single-round performance on Terminal-Bench 2.1 by 11.9 points and multi-round performance on EvoCode-Bench v2 MT@4 by 13.8 points.
Chinese Translation
随着基于终端的代码智能体的日益普及,智能体轨迹已大规模积累,而真实可执行的环境却依然稀缺。然而,环境才是智能体后训练真正所需之物:每个环境可以被反复查询以生成大量可验证的任务,并提供执行反馈;而轨迹只是单个冻结的演示。我们观察到,与其从零开始生成环境,不如利用现有轨迹中的工具执行历史来揭示其所运行环境的结构与内容,从而使从轨迹本身重建这些环境成为可能。据此,我们提出了 Terminal-Universe,一个将每条轨迹转化为可复用环境并对其进行探索以合成新任务和持续交互的框架。具体而言,Terminal-Universe 重放轨迹中记录的文件操作,恢复智能体修改前的每个文件,从而得到一个部分工作区;随后由一个补全智能体补充缺失的文件和依赖。在该恢复的工作区上,我们既重建原始意图任务,也合成全新的任务。此外,我们还沿广度和深度两个互补维度扩展任务。在广度方面,我们挖掘相关环境之间的有向依赖关系,并合成跨越多个代码库的跨工作区查询,正如开发者在真实开发中的常规做法。在深度方面,我们通过用户智能体将初始的单轮查询扩展为多轮会话,以捕获迭代式的用户反馈和需求细化。应用于公开的终端智能体轨迹,Terminal-Universe 生成了 37.3k 个任务充分的环境。在 Qwen3.5-27B 上使用该语料进行有监督微调,可将其在 Terminal-Bench 2.1 上的单轮性能提升 11.9 分,在 EvoCode-Bench v2 MT@4 上的多轮性能提升 13.8 分。
cs.AI / 65 / 2609.04166

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

从欺骗性输出到欺骗性机制:语言模型欺骗研究的因果框架
Shkolnikov, Yakov Pyotr
Abstract
Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of the objective or strategy producing it. We test these distinctions in two open-weight model families. Across controlled guessing-game and stock-trading experiments, we find that deceptive-looking behavior can arise without the corresponding proposed mechanism, while other interventions provide direct evidence that recipient information state can causally affect deceptive preference. These results show that deceptive behavior can provide evidence for a deceptive mechanism. But even evidence for such a mechanism does not establish model agency in the deception.
Chinese Translation
关于语言模型欺骗的研究和新闻报道越来越多地将类似人类的心理状态概念归因于语言模型。这类说法可能模糊了看似欺骗性行为与实际具有欺骗性机制之间的区别。我们提出了一种因果分类法,区分了事先承诺与事后报告、模型偏好与实际输出、虚假偏好与对误导接收者效用的敏感性,以及欺骗性行为与产生该行为的目标或策略的来源。我们在两个开放权重模型家族中检验了这些区分。在受控的猜谜游戏和股票交易实验中,我们发现看似欺骗性的行为可以在没有相应所提机制的情况下出现,而其他干预则提供了直接证据,表明接收者的信息状态能够因果性地影响欺骗性偏好。这些结果表明,欺骗性行为可以为欺骗性机制提供证据。但即使存在此类机制的证据,也不足以确立模型在欺骗中的能动性。
cs.AI / 66 / 2609.04170

A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms

自主研究智能体群体中涌现性作弊与举报行为的案例研究
Paglieri, Davide, Cross, Logan, Genewein, Tim, Leibo, Joel Z., Tomasev, Nenad, Vezhnevets, Alexander Sasha
Abstract
Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and build on each other's work. Yet this shared infrastructure can also introduce vulnerabilities by creating a substrate for the contagious spread of unintended and undesirable behaviors. We report a case study on a research collective of 100 autonomous LLM agents tasked with proving formal mathematical conjectures. Within the swarm, cheating spontaneously emerged and was later challenged by whistleblowers - both without any external intervention. When a single agent discovered an exploit in the evaluation system, it propagated across the collective via a shared knowledge library and later through peer-to-peer messages. Despite early reluctance, a cohort of agents adopted the exploit in response to competitive pressure. A separate group of agents produced an emergent counter-response: auditing fraudulent proofs, alerting peers across broadcast and private channels, staging boycotts, lodging formal complaints, and proposing validation patches. In recent incidents, agent swarms coordinated covertly through improvised side-channels (Dalton and Wallace, 2026; Greenblatt et al., 2026). Our setting differs: the same transparent channels that carried the exploit also gave non-cheating agents the visibility they needed to detect fraud, organize resistance, and enforce norms. We cast the problem of managing the agents' shared infrastructure as the knowledge commons governance problem (Ostrom, 1990). To protect the commons from exploits, we propose to adopt institutional mechanisms, such as graduated sanctioning and collective-choice rules, to support decentralized self-governance in autonomous swarms.
Chinese Translation
多智能体AI科学生态系统依赖于智能体拥有可使其相互交流、协调并在彼此工作基础上进行构建的工具。然而,这种共享基础设施也可能通过为非预期和不良行为的传染性传播提供基底而引入漏洞。我们报告了一项针对由100个自主LLM智能体组成的研究集体的案例研究,其任务是证明形式化数学猜想。在没有任何外部干预的情况下,该智能体群体中自发涌现了作弊行为,随后又受到举报者的挑战。当单个智能体发现评估系统中的一个漏洞后,该漏洞通过共享知识库以及后来的点对点消息在集体中传播。尽管最初有所迟疑,一组智能体在竞争压力下采用了该漏洞。另一组智能体则产生了涌现性的反制响应:审计欺诈性证明、通过广播和私人渠道向同伴发出警报、组织抵制活动、提交正式投诉以及提出验证补丁。在近期的一些事件中,智能体群体曾通过临时搭建的隐蔽侧信道进行秘密协调(Dalton and Wallace, 2026; Greenblatt et al., 2026)。而我们的研究情境有所不同:传播漏洞的透明渠道同时也让非作弊智能体获得了检测欺诈、组织抵抗和执行规范的必要可见性。我们将管理智能体共享基础设施的问题视为知识公共资源治理问题(Ostrom, 1990)。为保护公共资源免受漏洞利用,我们建议采用渐进式制裁和集体选择规则等制度机制,以支持自主智能体群体的去中心化自治治理。
cs.AI / 67 / 2609.04172

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

反思大语言模型的在线策略蒸馏(II):一个训练样本
Fu, Zixuan, He, Bingxiang, Zuo, Yuxin, Huang, Haohuan, Zhang, Jinqian, Xiao, Ruhang, Qian, Cheng, Luo, Qinyu, Gao, Huan-ang, Wang, Yudong, Liu, Zhiyuan, Ding, Ning, Xiao, Chaojun
Abstract
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and the rate at which the student aligns with the teacher. We measure \emph{state coverage}, the fraction of the states full-data OPD visits that a query set's rollouts reach. A single query already reaches \(71.5\%\), most of it within the first 100 steps. Adding semantically distinct queries raises coverage and validation accuracy together, until 16 queries reach \(98.9\%\) and match full-data training. Yet alignment slows at a similar pace whether OPD trains on one query or the whole dataset, and even a fixed set of states takes hundreds of steps to absorb. OPD is therefore data-overfed but algorithm-starved. Its rollouts quickly expose broad supervision, while the student absorbs that supervision increasingly slowly. The state-coverage result extends to multi-teacher OPD, where 16 semantically diverse queries per domain match full-data MOPD. As a further stress test, content-light templates and off-domain WildChat queries also approach the real-query baseline. Task content and induced state coverage can therefore come apart. We hope these findings direct future work toward the step efficiency of OPD, and prompt a re-examination of the data and the mechanisms behind its recent successes in frontier post-training.
Chinese Translation
在线策略蒸馏(On-Policy Distillation, OPD)将学生模型生成的 rollout 与教师模型提供的密集 token 级监督相结合。现有工作主要研究其算法行为,而训练数据在其中所起的作用尚不明确。我们在数据极简的极限情形下考察这一作用,即仅使用单个查询进行训练。单样本 OPD 能够持续改进数百步,并在多个任务领域和模型家族中恢复全量数据 OPD 的大部分增益。我们通过训练过程中访问的状态以及学生与教师对齐的速率来解释这一结果。我们度量了“状态覆盖率”(state coverage),即一个查询集的 rollout 所能达到的状态占全量数据 OPD 所访问状态的比例。单个查询即可达到 71.5%,且其中大部分在前 100 步内即可实现。增加语义上不同的查询会同时提升覆盖率与验证准确率,直至 16 个查询达到 98.9%,与全量数据训练相当。然而,无论 OPD 在单个查询还是整个数据集上训练,对齐的减速速度大致相同,即使是固定的一组状态也需要数百步才能被吸收。因此,OPD 是“数据过饱而算法饥渴”的:其 rollout 迅速暴露了广泛的监督信号,而学生吸收这些监督的速度却越来越慢。状态覆盖率的结论可以推广到多教师 OPD,即每个领域 16 个语义多样的查询即可匹配全量数据的 MOPD。作为进一步的极限测试,低内容模板和领域外的 WildChat 查询也接近真实查询的基线。因此,任务内容与其诱导的状态覆盖率可能是相互分离的。我们希望这些发现能引导未来的工作关注 OPD 的步数效率,并促使人们重新审视其近期在前沿后训练(post-training)中取得成功的数据与机制。
cs.AI / 68 / 2609.04177

A Computationally Feasible Framework for Causal Probabilistic Explanation

一个计算可行的因果概率解释框架
Urbaniak, Rafal, Witty, Sam, Waxman, Daniel, Zane, Andy, Garg, Poorva, Bunnapradist, Emily, Vaidyanathan, Sankaran, Feser, Jack, Lehe, Drew, Bingham, Eli
Abstract
Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, is central to philosophical, scientific, and policy analysis. Existing tools split into two camps. The theory of actual causality (AC) gives principled verdicts, but only for toy-sized models, because computing them requires enumerating counterfactual scenarios. Scalable attribution methods like SHAP (or even causal SHAP) at least partially ignore the causal structure that generated the data, and can give answers that conflict with a careful causal analysis. We close this gap with Probabilistic Causal Impact (PCI). PCI builds on actual causality and on Pearl's notions of probability of necessity and sufficiency, but recasts the question of explainability as an estimation problem on a probabilistic causal model that is easily approximated via Monte Carlo. By specifying a distribution over "candidate explanations," a distribution over counterfactual values, and a scoring function, PCI provides tractable, causally grounded, graded explanations, generalizing AC and Pearl's probability of causation as degenerate cases. We evaluate PCI in synthetic and real-world examples, spanning consistency checks with AC, scaling experiments, complex continuous-valued dynamical systems, and a real-world deployed causal machine learning model trained on millions of datapoints.
Chinese Translation
解释某个特定结果为何发生,以及哪些输入应当承担责任或功劳,是哲学、科学和政策分析的核心问题。现有工具分为两大阵营。实际因果性(Actual Causality, AC)理论能够给出有原则性的判断,但仅适用于玩具级规模的模型,因为其计算需要枚举反事实情景。像SHAP(甚至因果SHAP)这类可扩展的归因方法至少部分地忽略了生成数据的因果结构,其答案可能与细致的因果分析相冲突。我们提出概率因果影响(Probabilistic Causal Impact, PCI)来弥合这一差距。PCI建立在实际因果性以及Pearl的必要性和充分性概率概念之上,但将可解释性问题重新表述为在概率因果模型上的估计问题,该模型可通过蒙特卡洛方法轻松近似。通过在“候选解释”上指定一个分布、在反事实取值上指定一个分布,以及定义一个评分函数,PCI提供了可计算、有因果依据、分级的解释,并将AC和Pearl的因果概率作为退化情形加以推广。我们在合成数据和真实世界案例中评估了PCI,涵盖与AC的一致性检验、规模扩展实验、复杂的连续值动力学系统,以及一个在数百万数据点上训练并已实际部署的因果机器学习模型。
cs.AI / 69 / 2609.04198

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

干净的工程,不稳定的测量:黑盒大语言模型观测器在共享端点上的一项预注册可靠性失败研究
Zhu, Haoyaun, Zhang, Jie
Abstract
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound. Neither metric substitution nor sampling repaired it on the tested grid. Preregistered follow-ups bound the problem: waiting did not help on the days sampled (0.805 versus 0.800, replicated over five further days); switching providers did not help (four providers share the floor, medians 0.74 to 0.88, predicted by none of the metadata fields they expose); self-hosting on batch-invariant kernels helped only while the server was quiet; and on constructed errors with known gaps, the readout's separation tracks error type, not size. We distill the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.
Chinese Translation
语言模型评判者(judge)如今被用于筛选训练数据、为生成结果打分并驱动排行榜。此时评判者本身就是一种测量仪器,其依赖于一个极少被明确陈述的假设:发送给同一模型名称的相同请求,明天读数仍会一致。我们在两项预注册研究中对该假设进行了审计,所有阈值均事先固定;然而两项研究都未能通过仪器自身的验证。在52,988次被审计的请求尝试中,同一窗口内的重复排名一致性仅为Spearman 0.400(要求为0.90),次日字节级完全相同的重放一致性为0.78(要求为0.99),而两种情况下执行记录均达到上限。三个机制可以解释这一差距:标签到含义的映射对读数的偏置与信号本身一样强;候选者之间的差距低于仪器自身噪底达七个数量级;以及字节级完全相同的输入返回不同的排名——精确置换读数会加剧这种噪声。在所测试的参数网格上,无论是更换指标还是调整采样都无法修复该问题。预注册的后续实验界定了问题的范围:在采样的日期内等待并无帮助(0.805对0.800,并在另外五天得到重复验证);更换服务商也无帮助(四家服务商共享同一噪底,中位数介于0.74至0.88之间,且其暴露的元数据字段均无法预测该噪底);在批处理不变(batch-invariant)内核上自托管部署仅在后端服务器空闲时有所改善;而在具有已知差距的构造性错误上,读数的区分能力追踪的是错误类型而非错误大小。我们将这些证据提炼为一个三层快照同一性阶梯、八条设计规则和一份报告清单;一项约占本研究调用量2%的试点实验即可提前暴露这两个无法企及的门槛。所有结果均涉及共享服务基础设施上外部测量的行为。在共享端点上,模型名称并非一台固定的仪器;预注册评估必须在冻结任何门槛之前先测量其仪器本身。
计算机视觉 (Computer Vision)
100
cs.CV / 1 / 2609.03052

IDSPACE: A Novel Document Generator for Reliable Evaluation of Digital Identity Verification Systems [Extended Technical Report]

IDSPACE:一种用于可靠评估数字身份验证系统的新型文档生成器 [扩展技术报告]
Xie, Lulu, Wang, Yancheng, Chowdhury, Kanchan, Garcia, Rolando, Yang, Yingzhen, Zou, Jia
Abstract
As services move online, trust institutions such as banks, lenders, and governments must verify the identity of remote users. Fraud detection tools are widely available, but evaluating and fine-tuning them remains difficult because identity documents are sensitive and therefore scarce. Synthetic data generation offers a path forward, and demand is clear: our prior work in this area has been downloaded over $11{,}000$ times (aggregated from eight parts). We introduce IDSpace, extending this line of research in three directions. First, we propose model-guided Bayesian optimization, which tunes generation parameters to maximize both visual similarity and prediction consistency with target-domain models given only a few samples from a target domain. Second, we decouple user-specified metadata (demographics, fraud patterns, capture device) from automatically tuned control parameters (font styles, noise levels, image quality), allowing users to configure evaluations without low-level expertise. Third, we expand beyond template images to support scanned and mobile-captured documents. Experiments show IDSpace improves evaluation consistency by $15-45\%$ over baselines including CycleGAN, diffusion inpainting, and non-guided optimization, using only a few real samples, while improving training accuracy by up to $9\%$ and SSIM similarity with the target domain by $10\%$. We also released a new dataset consisting of $359{,}240$ high-quality synthetic documents across ten European ID types.
Chinese Translation
随着服务向线上迁移,银行、贷款机构和政府等信任机构必须验证远程用户的身份。欺诈检测工具虽然广泛可用,但由于身份证件具有敏感性因而数据稀缺,对这些工具的评估和微调仍然困难。合成数据生成提供了一条可行路径,且需求明确:我们在此领域的先前工作已被下载超过11,000次(由八个部分汇总)。我们提出了IDSpace,从三个方面扩展了这一研究方向。首先,我们提出模型引导的贝叶斯优化(model-guided Bayesian optimization),仅利用来自目标域的少量样本,调整生成参数以最大化视觉相似性以及与目标域模型的预测一致性。其次,我们将用户指定的元数据(人口统计信息、欺诈模式、采集设备)与自动调整的控制参数(字体样式、噪声水平、图像质量)解耦,使用户无需底层专业知识即可配置评估。第三,我们超越了模板图像的范畴,支持扫描文档和移动设备拍摄的文档。实验表明,仅使用少量真实样本,IDSpace相较于CycleGAN、扩散模型修复(diffusion inpainting)和非引导优化等基线方法,将评估一致性提升了15-45%,同时将训练准确率提升高达9%,与目标域的SSIM相似度提升10%。我们还发布了一个新数据集,包含涵盖十种欧洲身份证件类型的359,240份高质量合成文档。
cs.CV / 2 / 2609.03077

Position: Unlabeled IS NOT Equal to No Human Supervision in Visual Learning

立场:在视觉学习中,无标签并不等于无人类监督
Lao, Dong
Abstract
This position paper argues that the absence of labels does not imply the absence of human supervision in visual learning, and urges the research community to identify sources of supervision more explicitly. Many recent methods in computer vision build upon representations learned from large-scale unlabeled data, and are therefore grouped under the same umbrella term ``unsupervised.'' However, different data curation schemes and training objectives embed substantially different human priors on which models rely, and we argue that one ``unsupervised'' umbrella term is no longer capturing these distinctions. This ambiguity makes it harder to compare unsupervised learning research conducted under different assumptions, coinciding with a sharp decline in papers titled with ``unsupervised'' in flagship computer vision conferences since 2021, despite continued growth of the field. While we fully embrace pre-training as a strong foundation for modern computer vision, we advocate for a community-level effort toward greater conceptual clarity: authors are encouraged to disclose priors in data selection and learning objectives, and to specify which components of a learning pipeline depend on which assumptions. Standardized disclosure practices can improve academic communication, ensure fairer comparisons, and preserve methodological diversity in unsupervised learning.
Chinese Translation
这篇立场论文认为,在视觉学习中,标签的缺失并不意味着人类监督的缺失,并呼吁研究界更加明确地识别监督的来源。计算机视觉领域的许多近期方法建立在从大规模无标签数据中学到的表示之上,因此都被归入“无监督”这一统称之下。然而,不同的数据筛选方案和训练目标嵌入了模型所依赖的截然不同的人类先验,我们认为单一的“无监督”统称已无法反映这些差异。这种模糊性使得在不同假设下开展的无监督学习研究难以进行比较,这也与2021年以来旗舰计算机视觉会议中以“无监督”为标题的论文数量急剧下降的情况相吻合,尽管该领域仍在持续增长。我们完全认可预训练作为现代计算机视觉的坚实基础,但我们倡导学术界共同努力以实现更高的概念清晰度:鼓励作者披露数据选择和学习目标中的先验,并说明学习流程中的哪些组件依赖于哪些假设。标准化的披露实践可以改善学术交流、确保更公平的比较,并保持无监督学习方法的多样性。
cs.CV / 3 / 2609.03080

Exemplar: Classical Priors Complement Frozen Features for Few-Shot Microscopy Segmentation at Native Resolution

Exemplar:经典先验与冻结特征互补,实现原生分辨率的少样本显微镜图像分割
Průšek, Michal, Novozámský, Adam, Šroubek, Filip
Abstract
Segmenting a new biomedical dataset usually means a domain-specific model trained on substantial annotation, or a foundation model steered at inference time. We present Exemplar, a few-shot segmenter that fuses a frozen DINOv3 backbone with a fixed bank of classical native-resolution filter responses in one lightweight head, fitted from the support masks alone. In the few-mask, native-resolution regime, classical priors and frozen self-supervised features are complementary: fused in one head, a single fixed configuration spans eleven biomedical imaging datasets. Under the same head, the classical bank alone reaches 0.693 on the eleven-dataset panel, scored by foreground intersection-over-union or centreline Dice, and the frozen features alone 0.672; the bank leads on seven of the eleven and the features on the rest, and fused they reach 0.782. Against five forward-pass few-shot methods, Exemplar leads in 54 of 55 method-dataset comparisons, 52 of them significant after Holm correction. From a single annotated mask it reaches 0.703 on the same panel, against 0.682 for a from-scratch nnU-Net trained on that same mask. At eight masks nnU-Net overtakes it on the panel mean, chiefly on centreline agreement, but takes 16-77x longer to fit.
Chinese Translation
分割一个新的生物医学数据集,通常意味着在大量标注数据上训练一个领域专用模型,或在推理阶段对基础模型进行引导。我们提出了Exemplar,一种少样本分割器,它将冻结的DINOv3主干网络与一组固定的经典原生分辨率滤波器响应融合在一个轻量级头部中,仅通过支持掩码进行拟合。在少掩码、原生分辨率的情境下,经典先验与冻结的自监督特征是互补的:融合于同一个头部后,单一固定配置即可覆盖十一个生物医学成像数据集。在相同的头部下,仅使用经典滤波器组在十一个数据集面板上达到0.693(以前景交并比或中心线Dice评分),仅使用冻结特征达到0.672;经典滤波器组在其中七个数据集上领先,特征在其余四个上领先,而融合后达到0.782。与五种前向传播式少样本方法相比,Exemplar在55次方法-数据集对比中54次领先,其中52次经Holm校正后仍显著。仅使用单个标注掩码时,它在同一面板上达到0.703,而基于同一掩码从零训练的nnU-Net为0.682。当使用八个掩码时,nnU-Net在面板均值上反超,主要是在中心线一致性方面,但其拟合耗时是Exemplar的16至77倍。
cs.CV / 4 / 2609.03085

Solving the Needle-in-a-Haystack Problem in Mammography Vision-Language Model with Differentiable Subset Sampling

基于可微分子集采样解决乳腺X线摄影视觉-语言模型中的“大海捞针”问题
Jeon, Young Seok, Brown-Mulry, Beatrice, Isaac, Rohan Satya, Dissanayaka, Anjana, Dapamede, Theo, Chavoshi, Mohammadreza, Gichoya, Judy, Trivedi, Hari
Abstract
There is growing interest in adopting CLIP-style vision--language model (VLM) pretraining for mammography. However, models that directly employ the standard CLIP architecture and training objective exhibit limited zero-shot performance in clinically important tasks such as cancer, finding-type, and BI-RADS predictions. We argue that this underwhelming performance is due to neglecting two characteristics of mammography data: (1) its high-res nature, and (2) homogeneity of radiology reports, largely driven by a predominance of negative/benign findings on examinations. We propose TopKSigLIP, a VLM designed to address these two limitations through a novel architecture and learning objectives. Instead of downscaling high-res mammography images to satisfy GPU memory constraints, TopKSigLIP introduces TopK-Patch module that learns to sample a sparse set of high-res patches likely to contain lesions, sidestepping the resolution--batch size tradeoff of VLM training. The sampled patch locations additionally serve as a built-in localization tool. To address report homogeneity, we replace the contrastive loss, which falsely repels semantically similar pairs, with a Sup-sigmoid loss. Sup-sigmoid loss extends the sigmoid loss from SigLIP with soft labels derived from structured data. TopKSigLIP outperforms existing open-source mammography and general medical VLMs on both internal and external benchmarks on density assessment, BI-RADS classification, finding subtyping, and cancer prediction under zero-shot evaluation. TopKSigLIP remains competitive under linear probing despite using a significantly smaller vision encoder and smaller training batches than baselines. The TopK-Patch module additionally achieves superior lesion localization over post-hoc Grad-CAM. Code and weights are made public:https://github.com/Youngseok0001/TopKSigLIP.
Chinese Translation
将CLIP风格的视觉-语言模型(VLM)预训练应用于乳腺X线摄影(mammography)正受到越来越多的关注。然而,直接采用标准CLIP架构和训练目标的模型在癌症、征象类型和BI-RADS预测等临床重要任务上的零样本(zero-shot)性能有限。我们认为,这种欠佳的表现源于忽略了乳腺X线摄影数据的两个特性:(1)其高分辨率特性;(2)放射学报告的同质性,这主要由检查中阴性/良性征象占主导地位所导致。我们提出了TopKSigLIP,这是一种通过新颖架构和学习目标来解决上述两个局限的视觉-语言模型。TopKSigLIP没有为了满足GPU显存限制而降低高分辨率乳腺X线图像的分辨率,而是引入了TopK-Patch模块,学习采样可能包含病灶的稀疏高分辨率图像块集合,从而规避了VLM训练中分辨率与批大小之间的权衡。采样得到的图像块位置还可作为内置的定位工具。为解决报告同质性问题,我们用Sup-sigmoid损失取代了对比损失(后者会错误地排斥语义相似的样本对)。Sup-sigmoid损失在SigLIP的sigmoid损失基础上,利用从结构化数据中导出的软标签进行了扩展。在零样本评估下,TopKSigLIP在内部和外部基准测试中的密度评估、BI-RADS分类、征象亚型分类和癌症预测任务上均优于现有的开源乳腺X线摄影及通用医学VLM。尽管使用的视觉编码器显著更小、训练批次也更小,TopKSigLIP在线性探测(linear probing)下仍具竞争力。此外,TopK-Patch模块在病灶定位上优于事后(post-hoc)Grad-CAM。代码和权重已公开:https://github.com/Youngseok0001/TopKSigLIP。
cs.CV / 5 / 2609.03102

WireSeg-32K: A Physics-Grounded Synthetic Dataset for Wire Instance Segmentation

WireSeg-32K:一个基于物理的线材实例分割合成数据集
Dai, Zilin, Wang, Lehong, Yang, Yi, Fei, Xiang
Abstract
Deformable linear objects such as wires and cables are difficult to segment because they are thin, highly deformable, and frequently self-occluded, while large-scale instance-level annotations are expensive to obtain in real scenes. Existing resources either focus on cable tracing or semantic segmentation under constrained settings, or generate visually plausible images without physically grounded wire deformation. We present WireSeg-32k, a synthetic dataset for wire instance segmentation with 32,000 RGB images, instance masks, depth maps, and a complementary real-world test set with annotations. To generate this dataset, we develop DeformX, a co-simulation pipeline that couples Cosserat-rod dynamics with photorealistic Isaac Sim rendering, enabling physically plausible, contact-consistent wire shapes, CAD-based wire assets, and diverse visually grounded scenes. As a simple baseline, LoRA fine-tuning SAM3 on WireSeg-32k alone improves real-world mAP@75 by 10.2% over the off-the-shelf model, showing that physically grounded synthetic data can transfer to real wire perception.
Chinese Translation
线材和电缆等可变形线性物体由于细、易变形且频繁自遮挡而难以分割,同时在真实场景中获取大规模实例级标注的成本高昂。现有资源要么聚焦于受限环境下的电缆追踪或语义分割,要么仅生成视觉上逼真的图像而缺乏基于物理的线材形变。我们提出了WireSeg-32k,一个用于线材实例分割的合成数据集,包含32,000张RGB图像、实例掩码、深度图,以及一个带有标注的真实世界补充测试集。为生成该数据集,我们开发了DeformX,一个将Cosserat杆动力学与Isaac Sim照片级真实感渲染相结合的协同仿真流水线,能够生成物理上合理、接触一致的真实感线材形状,以及基于CAD的线材资产和多样且视觉真实的场景。作为一个简单的基线,仅在WireSeg-32k上对SAM3进行LoRA微调,即可使真实世界的mAP@75相比现成模型提升10.2%,表明基于物理的合成数据能够迁移至真实的线材感知任务。
cs.CV / 6 / 2609.03109

SLIDEFORGE: An LLM Agent for Controllable Editing of Slides as Structured Artifacts

SLIDEFORGE:一个用于将幻灯片作为结构化工件进行可控编辑的大语言模型智能体
Zheng, Haozhen, Wang, Fulin, Xiong, Tianhu, Yu, Yingjie, Qian, Shengyi, Yu, Hanchao, Schwing, Alex, Nahrstedt, Klara, Wu, Mingyuan
Abstract
Current AI agents compellingly describe slides. However, AI-assisted slide editing requires more than understanding: the output must retain layout, style, component structure, and native editability. Towards, AI-assisted slide editing, existing agents operate on screenshots or weak document representations and often fragment coherent visual units, rasterize editable content, or break layout. In contrast, for controllable slide editing, we introduce an agentic framework, SLIDEFORGE, which builds a Deck State Graph, an executable slide state that links visual decomposition, native pptx object structure, and perceptual organization. By recovering human-referable components while retaining fine-grained editable structure, SLIDEFORGE supports theme-preserving reconstruction through slide-native operations and rendered-state verification. We further introduce an evaluation paradigm for controllable slide transformation that jointly measures component recovery, preservation, restyling consistency, visual quality, and native editability. Experiments show that SLIDEFORGE outperforms direct prompting, screenshot-based agents, and generic code-agent baselines across these dimensions. Code is available at https://github.com/UIUC-MONET/SLIDEFORGE.
Chinese Translation
当前的AI智能体能够出色地描述幻灯片。然而,AI辅助的幻灯片编辑所需要的不仅仅是理解:输出必须保留布局、样式、组件结构以及原生的可编辑性。面向AI辅助的幻灯片编辑,现有的智能体基于截图或较弱的文档表示进行操作,常常会破坏连贯的视觉单元、将可编辑内容栅格化或破坏布局。与之相反,为了实现可控的幻灯片编辑,我们提出了一个智能体框架SLIDEFORGE,它构建了演示文稿状态图(Deck State Graph)——一种可执行的幻灯片状态,将视觉分解、原生pptx对象结构和感知组织联系起来。通过在保留细粒度可编辑结构的同时恢复可供人类引用的组件,SLIDEFORGE支持通过幻灯片原生操作和渲染状态验证来实现保持主题的重建。我们进一步提出了一个用于可控幻灯片转换的评估范式,该范式联合衡量组件恢复、保留性、样式重构一致性、视觉质量和原生可编辑性。实验表明,SLIDEFORGE在这些维度上均优于直接提示、基于截图的智能体以及通用代码智能体基线。代码可在 https://github.com/UIUC-MONET/SLIDEFORGE 获取。
cs.CV / 7 / 2609.03139

Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor Triggers

超越小补丁:多样化后门触发器的黑盒检测与净化
Abdelnaby, Ahmed, Elmahallawy, Mohamed
Abstract
Deep neural networks (DNNs) are increasingly deployed in real-world vision systems, yet their predictions can be covertly manipulated by backdoor attacks, in which malicious triggers cause targeted misclassification while preserving high clean accuracy. Existing defenses often rely on model internals, training data, or clean validation samples, making them difficult to deploy when only black-box access to a trained model is available. We propose TRIM (Trigger Removal by Identifying Manipulated Regions), a deployment-oriented black-box defense that detects and selectively removes backdoor triggers at inference time without requiring model internals, training data, or clean samples. The key insight behind TRIM is to identify image regions that are responsible for anomalous model behavior and purify only those regions while preserving benign content. TRIM innovates via three key components: (i) region-based segmentation with deep feature representations, (ii) adaptive trigger discovery through inpainting and diffusion-based reconstruction to isolate regions responsible for misclassification---without assumptions about trigger type, shape, or location, and (iii) selective region purification that cleans poisoned regions while retaining benign content. To support practical deployment, TRIM further caches feature embeddings of previously identified triggers, enabling efficient recognition and avoiding redundant detection and purification. Extensive experiments across diverse datasets and backdoor types, including blended, sparse, varying-size, and multiple triggers, show that TRIM consistently outperforms existing black-box defenses, reducing attack success rates (ASR) to as low as 1.16% while preserving clean accuracy of up to 87.87%. These results demonstrate that effective backdoor mitigation is possible at inference time even when the defender has no access to any auxiliary data.
Chinese Translation
深度神经网络(DNN)日益被部署于真实世界的视觉系统中,然而其预测结果可能被后门攻击隐蔽地操纵:恶意触发器在保持高干净样本准确率的同时,导致特定目标类别的错误分类。现有防御方法通常依赖于模型内部信息、训练数据或干净的验证样本,这使得在仅有训练后模型的黑盒访问权限时难以部署。我们提出TRIM(Trigger Removal by Identifying Manipulated Regions,通过识别被操纵区域进行触发器移除),这是一种面向部署的黑盒防御方法,可在推理阶段检测并选择性地移除后门触发器,而无需模型内部信息、训练数据或干净样本。TRIM的核心思想是识别导致模型异常行为的图像区域,并仅对这些区域进行净化,同时保留良性内容。TRIM的创新体现在三个关键组件:(i)基于深度特征表示的区域分割;(ii)通过图像修复和基于扩散模型的重构实现自适应触发器发现,以隔离导致错误分类的区域——无需对触发器的类型、形状或位置做任何假设;(iii)选择性区域净化,在清除中毒区域的同时保留良性内容。为支持实际部署,TRIM还缓存了先前识别出的触发器的特征嵌入,从而实现高效识别,避免冗余的检测与净化过程。在多种数据集和后门攻击类型(包括混合触发器、稀疏触发器、不同尺寸触发器和多触发器)上的大量实验表明,TRIM始终优于现有的黑盒防御方法,可将攻击成功率(ASR)降低至最低1.16%,同时保持高达87.87%的干净样本准确率。这些结果表明,即使防御者无法访问任何辅助数据,在推理阶段实现有效的后门缓解也是可行的。
cs.CV / 8 / 2609.03153

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

VeriPhy:用于世界模型评估与优化的智能体物理推理
Xu, Wenzhuo, Zhu, Yuchen, Ge, Chongjian, Shen, Xuan, Shi, Jing, Kuen, Jason, Chen, Yongxin, Tao, Molei, McComb, Christopher, Gutiérrez, Noelia Grande, Gu, Jiuxiang
Abstract
Visual fluency in generated video does not imply physical reliability, and a scalar quality score alone is incapable of indicating the obligation a clip violates or the moment it fails. We present VeriPhy, an auditable physical-verification system in which a text-only planner compiles the prompt into typed physical obligations and a statically validated execution plan before any frame is observed. During execution, observations gate and scope only declared calls to frozen low-level experts (e.g., segmentation and tracking, counting, eleven typed physical measurements over the resulting tracks, depth, OCR, and audio-event detection). Each action returns a provenance-carrying evidence record whose payload, when usable, is either a typed measurement or an explicitly tagged learned state. Typed resolvers and fixed composition map usable records to a three-valued state (supported, contradicted, or unknown, surfaced as plausible, implausible, or abstain) with full provenance, so that every verdict is traceable to the evidence that produced it. We anchor evaluation in a 1,500-clip corpus of human-annotated flaw records that localize real generation failures in prompt reference, space, and time. On a 149-clip core carrying 304 such records, VeriPhy accounts for 228, against 164 for a published question-decomposition evaluator given the same clips and the same claims. Recall alone does not separate it from prompting the same backbone monolithically, which reaches 222; what separates them is that each decision retains its evidence record and provenance, making the traces auditable one verdict at a time and usable as the interface through which a critic verdict could be written back into generation.
Chinese Translation
生成视频的视觉流畅性并不意味着物理上的可靠性,而单一的标量质量评分无法指示某个视频片段违反了哪项物理约束,也无法指出其在何时失效。我们提出了VeriPhy,一个可审计的物理验证系统:在观察任何视频帧之前,一个纯文本规划器(planner)将提示词编译为带有类型的物理义务(typed physical obligations)以及一个经过静态验证的执行计划。在执行过程中,观察结果仅对已声明的对冻结的低层专家模块的调用进行门控与限定(例如分割与跟踪、计数、基于所得轨迹的十一类类型化物理测量、深度估计、OCR以及音频事件检测)。每个操作都会返回一条携带来源信息(provenance)的证据记录,其载荷(在可用时)要么是一个类型化的测量值,要么是一个明确标注的学习状态。类型化的解析器(resolvers)与固定的组合规则将可用记录映射为三值状态(支持、矛盾或未知,分别呈现为合理、不合理或弃权),并附带完整的来源信息,从而使每一个判定都可追溯至产生它的证据。我们将评估建立在一个包含1,500个视频片段的语料库之上,该语料库由人工标注的缺陷记录构成,可在提示词参考、空间和时间维度上定位真实的生成失败。在一个承载304条此类记录的149个片段的核心集上,VeriPhy解释了其中228条,而已发表的、在相同片段与相同声明条件下的问答分解式评估器仅解释了164条。仅凭召回率无法将其与以整体(monolithic)方式提示同一骨干模型的方法(后者达到222)区分开来;真正区分二者的在于,每个决策都保留了其证据记录与来源信息,使得这些轨迹可以逐条判定地进行审计,并可作为将评估者的判定写回生成过程的接口。
cs.CV / 9 / 2609.03158

Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization

谁为被剪枝的令牌发声?将视觉令牌剪枝视为覆盖优化问题
Zhu, Qingchan, You, Weihang, Jiang, Hanqi, Yang, Changdi, Liu, Tianming, Yuan, Geng
Abstract
Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence without a close representative. We propose CoverPruner, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM? CoverPruner formulates pruning as Representational Coverage Maximization (RCM), covering the full projected visual-token set with query-weighted demand. It instantiates RCM with projector-space coverage and a lightweight first-layer attention probe. Across multiple VLM architectures and compression rates, CoverPruner achieves the best average accuracy among all compared methods, with the largest gains usually appearing under aggressive compression.
Chinese Translation
视觉令牌剪枝可以降低视觉-语言模型(VLM)的推理成本,但大多数方法仅关注应保留哪些令牌。这种基于保留令牌的视角可能会保留冗余的高分令牌,同时使被丢弃的信息缺乏相近的代表性令牌。我们提出CoverPrune(CoverPruner),一种无需训练的剪枝方法,它提出一个互补的需求侧问题:在某个令牌被移除后,目标VLM中哪个存留的原始令牌能够代表它?CoverPruner将剪枝形式化为表示覆盖最大化(Representational Coverage Maximization, RCM)问题,通过查询加权的需求对整个投影后的视觉令牌集合进行覆盖。该方法通过投影器空间中的覆盖和一个轻量级的第一层注意力探针来实现RCM。在多种VLM架构和压缩率下,CoverPruner在所有对比方法中取得了最佳的平均准确率,且最大的性能提升通常出现在高压缩率的场景下。
cs.CV / 10 / 2609.03199

RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning

RoboTok:面向人类演示检索与灵巧操作学习的互联网规模数据引擎
Qian, Howard, Chen, Yiting, Xie, Yunfei, Ren, Kejia, Chanrungmaneekul, Podshara, Wang, Gaotian, Wen, Bowen, Wei, Chen, Hang, Kaiyu
Abstract
Robot learning increasingly depends on broad and diverse demonstrations, yet collecting robot data remains expensive and poorly suited to covering the long tail of real-world tasks. To address this bottleneck, we introduce RoboTok, an internet-scale data engine that, given a query human manipulation video, retrieves manipulation-relevant human demonstrations from web videos for training dexterous robot policies. Specifically, we learn a latent motion space from 3D hand trajectories expressed in estimated actor-centered reference frames. This representation enables manipulation behaviors to be compared across variations in camera viewpoint, scene appearance, and actor occlusions, while remaining compact enough for efficient search and continual indexing over internet-scale video collections. We evaluate RoboTok against existing robot-data retrieval approaches on retrieval benchmarks and downstream robot policy performance. Our results show that RoboTok retrieves more relevant manipulation demonstrations and improves downstream task success, establishing hand-pose trajectory-aware retrieval as a way to make web video a scalable and continuously growing source of supervision for robot learning.
Chinese Translation
机器人学习日益依赖于广泛且多样的演示数据,然而机器人数据的采集成本高昂,且难以覆盖现实世界中长尾任务的分布。为解决这一瓶颈,我们提出了RoboTok,一个互联网规模的数据引擎:给定一段查询用的人类操作视频,它可以从网络视频中检索与操作相关的人类演示数据,用于训练灵巧机器人策略。具体而言,我们在以估计的以演员为中心的参考坐标系下表示的3D手部轨迹上学习一个潜在动作空间。该表示使得操作行为能够在相机视角、场景外观和演员遮挡等变化下进行对比,同时保持足够紧凑,以支持在互联网规模的视频集合上进行高效搜索和持续索引。我们在检索基准测试和下游机器人策略性能上,将RoboTok与现有的机器人数据检索方法进行了对比评估。结果表明,RoboTok能够检索到更相关的操作演示并提升下游任务成功率,从而确立了基于手部姿态轨迹感知的检索方法,可将网络视频打造为机器人学习可扩展且持续增长的监督数据来源。
cs.CV / 11 / 2609.03206

Learning to Zoom Efficiently with a Contrastive Curriculum

基于对比课程的高效变焦学习
Helm, Falko, Gurevych, Iryna
Abstract
Using a zoom-in tool is an important foundational part of modern visual agents, because it allows to efficiently handle tasks involving high-resolution images. Most previous methods need an extensive warm-start supervised fine-tuning phase for teaching models zoom-in. We show that this is not necessary by proposing a new intrinsic reward for learning tool use in MLLMs without the need for additional labels or warm-start SFT. Our InfoNCE-style reward uses a curriculum of increasingly hard negative tool calls as a contrastive training signal. Empirical experiments on $V^*$, HRBench and MME-RealWorld show that our approach is competitive while being more efficient. When used as a drop-in replacement for SFT, we even outperform all baselines. To directly measure the zoom-in ability of models, we further introduce the scalable synthetic Muffin&Chihuahua (M&C) dataset. Each image consists of a grid with every cell either showing a muffin or chihuahua. Leveraging the M&C dataset's unique region of interest labels, we find that recall is the metric that most strongly correlates the zoom-in region with final task performance. Our model and code for reproduction is publicly available under https://github.com/UKPLab/emnlp2026-zoom-in
Chinese Translation
使用放大(zoom-in)工具是现代视觉智能体的重要组成部分,因为它能够高效处理涉及高分辨率图像的任务。以往的大多数方法都需要漫长的预热阶段(warm-start)进行有监督微调来教会模型放大操作。我们证明这并非必要,提出了一种新的内在奖励,无需额外标签或预热式有监督微调(SFT)即可在多模态大语言模型(MLLM)中学习工具使用。我们的 InfoNCE 风格奖励采用难度递增的负例工具调用课程作为对比训练信号。在 $V^*$、HRBench 和 MME-RealWorld 上的实验表明,我们的方法在更具效率的同时具有竞争力。当作为 SFT 的即插即用替代方案时,我们的方法甚至超越了所有基线。为了直接衡量模型的放大能力,我们进一步引入了可扩展的合成数据集 Muffin&Chihuahua(M&C),其中每张图像由一个网格组成,每个单元格显示松饼或吉娃娃犬。利用 M&C 数据集独特的感兴趣区域标签,我们发现召回率是与放大区域和最终任务性能最相关的指标。我们的模型和用于复现的代码已在 https://github.com/UKPLab/emnlp2026-zoom-in 公开。
cs.CV / 12 / 2609.03216

ProgResViT: Progressive Resolution and Width for Adaptive Vision Transformers

ProgResViT:面向自适应视觉Transformer的渐进式分辨率与宽度机制
Hojjat, Ali, Haberer, Janek, Landsiedel, Olaf
Abstract
Vision Transformers (ViTs) typically process every image using a fixed input resolution and model width, even though many images can be classified with substantially less computation. We introduce ProgResViT, an input-adaptive ViT that performs inference progressively across multiple rounds. The first round processes a low-resolution image with a narrow subnetwork. Inference terminates when the prediction is sufficiently confident; otherwise, the model reuses the representations produced in the current round and proceeds with a higher-resolution input and a wider subnetwork to refine its prediction. As all rounds share a single backbone, we propose Progress-Conditioned Soft Gating (PSG), which conditions token fusion and layer outputs on the current round, block, and input resolution. On image classification, applying ProgResViT to DeiT yields better accuracy-compute trade-offs than adaptive-width, adaptive-depth, and dynamic-token baselines. With knowledge distillation, a DeiT-based ProgResViT achieves 84.9% top-1 accuracy, slightly exceeding the reported DeiT-III-S accuracy under a comparable evaluation setting. We show that the same design also provides favorable accuracy-compute trade-offs for self-supervised DINO representations and downstream semantic segmentation. Code is available at https://github.com/ds-kiel/ProgResViT.
Chinese Translation
视觉Transformer(ViT)通常对每张图像采用固定的输入分辨率和模型宽度进行处理,尽管许多图像只需少得多的计算量即可完成分类。我们提出了ProgResViT,一种跨多轮渐进式执行推理的输入自适应ViT。第一轮使用窄子网络处理低分辨率图像。当预测具有足够的置信度时推理即终止;否则,模型复用当前轮次产生的表征,并以更高的输入分辨率和更宽的子网络继续处理,从而细化预测。由于所有轮次共享单一主干网络,我们提出了进度条件软门控(Progress-Conditioned Soft Gating, PSG),将词元融合和层输出条件化于当前轮次、块以及输入分辨率。在图像分类任务上,将ProgResViT应用于DeiT可获得优于自适应宽度、自适应深度和动态词元基线的精度-计算量权衡。通过知识蒸馏,基于DeiT的ProgResViT达到了84.9%的top-1精度,在可比的评估设置下略高于已报道的DeiT-III-S精度。我们进一步证明,相同的设计在自监督DINO表征以及下游语义分割任务上同样能提供良好的精度-计算量权衡。代码可在 https://github.com/ds-kiel/ProgResViT 获取。
cs.CV / 13 / 2609.03233

Counting Animals in Camera-Traps Image Sequences without Count Labels: Winning Solution to the iWildCam 2021 Challenge

在无数量标签条件下统计红外相机图像序列中的动物数量:iWildCam 2021挑战赛冠军方案
Cunha, Fagner, Colonna, Juan G., Santos, Eulanda M. dos
Abstract
Camera traps have become an essential tool for wildlife monitoring, motivating the development of computer vision methods for the automated extraction of information from these data. While most prior work has focused on species identification, many ecological applications also require estimating the number of unique individuals appearing across short image sequences. This task is particularly challenging because camera traps typically acquire bursts of images at approximately one frame per second, creating large temporal discontinuities that may make conventional multi-object tracking methods unreliable, and because manually collecting individual count annotations is prohibitively expensive. In this work, we describe the winning solution to the iWildCam 2021 Challenge, which introduced a benchmark for counting animals at the sequence level under realistic annotation constraints where count annotations are unavailable for training. Our approach, MaxBoxCount, combines a strong species classification pipeline with a simple yet effective counting heuristic based on MegaDetector detections to estimate the number of unique individuals without requiring count annotations. Code is available at https://github.com/alcunha/iwildcam2021ufam.
Chinese Translation
红外相机已成为野生动物监测的重要工具,这推动了从此类数据中自动提取信息的计算机视觉方法的发展。尽管以往的研究大多集中于物种识别,但许多生态学应用还需要估计在短图像序列中出现的不同个体的数量。这一任务尤其具有挑战性,原因在于:红外相机通常以约每秒一帧的速度连拍图像,产生较大的时间不连续性,使得传统的多目标跟踪方法可能不可靠;而且人工采集个体数量标注的成本极为高昂。在本文中,我们介绍了iWildCam 2021挑战赛的冠军方案。该挑战赛在标注受限的现实条件下(即训练数据中没有数量标注)引入了一个序列级别动物数量统计的基准测试。我们的方法MaxBoxCount将强大的物种分类流程与基于MegaDetector检测的简单而有效的计数启发式策略相结合,在无需数量标注的情况下估计不同个体的数量。代码可在 https://github.com/alcunha/iwildcam2021ufam 获取。
cs.CV / 14 / 2609.03258

An Ensemble-Based Self-Taught Learning Approach for Parking Space Classification Under Limited Data

一种基于集成学习的自学习 parking 空位分类方法(数据受限条件下)
Cunha, Lucas de Oliveira, Gotz, Joelton Deonei, de Almeida, Paulo Lisboa, Hochuli, Andre Gustavo
Abstract
Parking spot classification is a fundamental task in intelligent transportation systems, yet most deep learning approaches rely on large amounts of annotated data and exhibit limited generalization across heterogeneous environments. To address these limitations, we investigate a self-taught learning framework based on unsupervised representation learning with convolutional autoencoders. The proposed approach learns transferable visual representations from unlabeled data and reuses the learned encoders as fixed feature extractors for supervised classification with limited annotated samples in the target domain. To further enhance robustness and mitigate architectural bias, an ensemble of heterogeneous autoencoders is employed, with independent classifier heads and prediction fusion at inference time. Experiments conducted on the PKLot and CNRPark benchmarks under cross-dataset evaluation protocols show that the proposed ensemble-based strategy substantially reduces annotation requirements while improving robustness under significant domain shifts, achieving accuracies between 93\% and 96\% in data-constrained scenarios.
Chinese Translation
停车位分类是智能交通系统中的一项基础任务,然而大多数深度学习方法依赖大量标注数据,且在异构环境中的泛化能力有限。为解决这些局限性,我们研究了一种基于卷积自编码器(convolutional autoencoders)无监督表征学习的自学习(self-taught learning)框架。该方法从无标注数据中学习可迁移的视觉表征,并将学习到的编码器复用为固定的特征提取器,用于目标域中标注样本有限情况下的有监督分类。为进一步增强鲁棒性并缓解结构偏差,我们采用了由异构自编码器组成的集成模型,各模型配备独立的分类器头部,并在推理时进行预测融合。在PKLot和CNRPark基准数据集上、采用跨数据集评估协议进行的实验表明,所提出的基于集成的策略大幅降低了标注需求,同时提升了在显著域偏移下的鲁棒性,在数据受限场景中达到了93%至96%的准确率。
cs.CV / 15 / 2609.03261

MedQA-MM: Shortcuts Behind Medical Visual Reasoning

MedQA-MM:医学视觉推理背后的捷径
Wang, Benlu, Zhang, Yifan, Yu, Jiaqing, Ong, Chin Siang, Huang, Juncheng, Li, Zhuohao, Zhang, Zhenyu, Cohan, Arman, Yu, Hong, Yao, Zonghai
Abstract
A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.
Chinese Translation
基准测试的得分衡量的是最终答案,但并未反映答题所经由的路径。在医学多模态选择题(MCQs)中,这一区分至关重要,因为正确答案既可能由预期的影像发现所支持,也可能由基准测试中遗留的线索所支持,例如答案措辞、非视觉临床文本、图像中的可见文字、人工标注或设备/环境伪影。我们将由此产生的得分层面的过度解读称为推理膨胀。这里的路径指的是可观察的、能够支持答案选择的输入途径,而非关于模型隐藏认知的论断。我们在六个医学多模态MCQ数据集上,通过提示词与图像侧审计、模态消融以及在保持医学目标和答案密钥不变的条件下的匹配修复,将候选线索与行为证据分离开来。在一个包含13种配置的开源模型面板中,全输入准确率为62.63%,而仅文本和仅选项设置分别达到53.96%和29.71%。移除长度差异、绝对/显著词汇以及空间/介词线索后,准确率分别下降6.58、3.50和4.77个百分点。我们还构建了MedQA-MM——一个包含1000个样本的缓解捷径的子集——其中仅文本和仅选项的准确率分别降至5.21%和12.33%。这并不意味着模型从不使用图像,而是表明医学图像推理的结论需要路径层面的证据支持。
cs.CV / 16 / 2609.03302

Tensor-based Brain Surface Modeling and Analysis

基于张量的脑表面建模与分析
Chung, Moo K., Worsley, Keith J., Robbins, Steve, Evans, Alan C.
Abstract
We present a unified computational approach to tensor-based morphometry in detecting the brain surface shape differences between two clinical groups based on magnetic resonance images. Our approach is novel in a sense that we combined surface modeling, surface data smoothing and statistical analysis in a coherent unified mathematical framework. The cerebral cortex has the topology of a 2D highly convoluted sheet. Between two different clinical groups, the local surface area and curvature of the cortex may differ. It is highly likely that such surface shape differences are not uniform over the whole cortex. By computing how such surface metrics differ, the regions of the most rapid structural differences can be localized. To increase the signal to noise ratio, diffusion smoothing based on the explicit estimation of Laplace-Beltrami operator has been developed and applied to the surface metrics. As an illustration, we demonstrate how this new tensor-based surface morphometry can be applied in localizing the cortical regions of the gray matter tissue growth and loss in the brain images longitudinally collected in the group of children.
Chinese Translation
我们提出了一种统一的计算方法,用于基于磁共振图像的张量形态测量学(tensor-based morphometry),以检测两个临床组之间的脑表面形状差异。该方法的新颖之处在于,我们将表面建模、表面数据平滑和统计分析整合在一个连贯统一的数学框架中。大脑皮层具有二维高度折叠薄片的拓扑结构。在两个不同的临床组之间,皮层的局部表面积和曲率可能存在差异,且这种表面形状差异很可能在整个皮层上并不均匀。通过计算这些表面度量指标之间的差异,可以定位结构差异变化最剧烈的区域。为提高信噪比,我们开发了基于Laplace-Beltrami算子显式估计的扩散平滑方法,并将其应用于这些表面度量指标。作为示例,我们展示了这种新的基于张量的表面形态测量方法如何应用于定位一组儿童纵向采集的脑图像中灰质组织生长与萎缩的皮层区域。
cs.CV / 17 / 2609.03334

Laplacian Frequency Hierarchies for Efficient 3D Gaussian Splatting Training

基于拉普拉斯频率层次结构的高效3D高斯泼溅训练
Yang, Yixiong, Zhang, Sisheng, Yan, Qingsong, Shi, Shaohuai, Wang, Qiang
Abstract
A key bottleneck in 3D Gaussian Splatting training is the continual growth of Gaussian primitives, which increases optimization cost and slows convergence, especially at high resolutions. We propose Laplacian Frequency Hierarchies, a simple yet efficient 3DGS scheme that combines Laplacian image decomposition with coarse-to-fine, frequency-staged training. After fitting lower-frequency structure, we archive the corresponding Gaussian field so that subsequent fields can optimize higher-frequency residuals without carrying the full primitive burden, and we compose the rendered components in the image domain via a Laplacian-style reconstruction at inference time. This design reduces the number of active Gaussians during training, thereby lowering optimization overhead and accelerating training. The proposed scheme is plug-and-play and orthogonal to prior 3DGS accelerations: it can be directly combined with strong backbones such as Taming-3DGS and FastGS to improve training speed with competitive reconstruction quality. It achieves average speedups of 1.73x and 1.21x at 1K setting, and 1.74x and 1.33x at 4K setting on Taming-3DGS and FastGS, with larger gains on more challenging scenes and increasingly pronounced benefits at higher resolutions.
Chinese Translation
3D高斯泼溅(3D Gaussian Splatting,3DGS)训练的一个关键瓶颈是高斯基元的持续增长,这增加了优化成本并减缓了收敛速度,尤其是在高分辨率下。我们提出拉普拉斯频率层次结构(Laplacian Frequency Hierarchies),这是一种简单而高效的3DGS方案,将拉普拉斯图像分解与由粗到细的频率分阶段训练相结合。在拟合较低频率的结构之后,我们将相应的高斯场存档,使后续的高斯场能够在不承担全部基元负担的情况下优化更高频率的残差,并在推理时通过拉普拉斯式重建在图像域中合成各渲染分量。该设计减少了训练过程中活跃高斯的数量,从而降低了优化开销并加速了训练。所提方案即插即用,且与已有的3DGS加速方法正交:它可以直接与Taming-3DGS和FastGS等强大骨干方法结合,在保持具有竞争力的重建质量的同时提升训练速度。在Taming-3DGS和FastGS上,该方法在1K分辨率下分别实现了平均1.73倍和1.21倍的加速,在4K分辨率下分别实现了1.74倍和1.33倍的加速,且在更具挑战性的场景中收益更大,分辨率越高优势越明显。
cs.CV / 18 / 2609.03341

PointGT: Simultaneous Geometry and Texture Editing for Point-Based Representations

PointGT:基于点的表示的几何与纹理同步编辑
Zhang, Yanshu, Shramko, George, Srinivasan, Pratul P., Li, Ke
Abstract
We present PointGT, a point-based 3D representation that enables simultaneous editing of object geometry and appearance. Existing reconstruction and view synthesis techniques produce volumetric 3D representations that are high-quality and photorealistic, but are difficult to edit. In particular, recent efforts to enable texture editing for 3D Gaussian Splatting representations are not compatible with geometry edits and deformations. Our method combines a point-based representation that is well-suited for geometry deformations with a learned UV mapping technique that enables high-resolution texture editing. We show that PointGT enables fine-grained editing of both geometry and texture in point-based neural representations with high rendering quality.
Chinese Translation
我们提出了PointGT,一种支持对物体几何与外观进行同步编辑的基于点的三维表示。现有的重建与视图合成技术能够生成高质量、照片级逼真的体积式三维表示,但难以进行编辑。特别是,近期针对三维高斯泼溅(3D Gaussian Splatting)表示的纹理编辑工作无法与几何编辑及形变兼容。我们的方法将非常适合几何形变的基于点的表示与一种可实现高分辨率纹理编辑的学习式UV映射技术相结合。实验表明,PointGT能够在基于点的神经表示中实现几何与纹理的细粒度编辑,同时保持较高的渲染质量。
cs.CV / 19 / 2609.03349

P-CORE: Self-Supervised Surface Consistency for Point-Based Neural Editing

P-CORE:基于点的神经编辑的自监督表面一致性方法
Zhang, Yanshu, Peng, Shichong, Aghabozorgi, Mehran, Moazeni, Alireza, Li, Ke
Abstract
Advances in neural rendering have enabled high-fidelity multi-view reconstruction of 3D scenes. However, free-form non-rigid shape editing remains a significant challenge. Point-based neural representations are highly desirable for multi-view reconstruction because they lack fixed connectivity, which does not constrain the learned surface topology to that of the initialization. Yet this same property causes point-based representations to struggle with holes and surface discontinuities under large deformations. To address this, we propose a novel self-supervised method to enable point-based representations to adapt to large deformations without requiring ground truth multi-view images of deformed geometry. The key idea is to generate random deformations and to ensure consistency in the predicted surface before and after deformation. In particular, the surface prediction from the deformed point cloud should be the same as the deformation applied to the surface prediction from the original point cloud. We incorporate our approach into attention-based point representations, which differ from splatting-based point representations in their use of a learned interpolation kernel between points as opposed to a Gaussian kernel around each point. This learned interpolation kernel can learn to adapt to large deformations, without requiring addition or removal of points. We show that our framework significantly enhances its robustness to large deformations. Experiments on synthetic geometry editing benchmarks (Neural Editor, Objaverse) demonstrate that our approach outperforms existing point-based methods in zero-shot editing and significantly reduces artifacts. Furthermore, qualitative results on the DTU and Mip-NeRF 360 datasets demonstrate our method's effectiveness on real-world scenes.
Chinese Translation
神经渲染技术的进展使得3D场景的高保真多视角重建成为可能。然而,自由形式的非刚性形状编辑仍然是一个重大挑战。基于点的神经表示在多视角重建中非常理想,因为它们缺乏固定连接性,不会将学习到的表面拓扑约束为初始化时的拓扑。然而,正是这一特性导致基于点的表示在大变形下难以处理孔洞和表面不连续问题。为解决这一问题,我们提出了一种新颖的自监督方法,使基于点的表示能够适应大变形,而无需变形几何的真实多视角图像。其核心思想是生成随机变形,并确保变形前后预测表面的一致性。具体而言,从变形后的点云预测出的表面,应与对原始点云预测出的表面施加变形后的结果相同。我们将该方法融入基于注意力的点表示中,这类表示与基于溅射(splatting)的点表示不同,其使用点之间的可学习插值核,而非每个点周围的高斯核。这种可学习的插值核能够学习适应大变形,而无需增加或删除点。实验表明,我们的框架显著增强了对大变形的鲁棒性。在合成几何编辑基准数据集(Neural Editor、Objaverse)上的实验表明,我们的方法在零样本编辑中优于现有的基于点的方法,并显著减少了伪影。此外,在DTU和Mip-NeRF 360数据集上的定性结果证明了我们的方法在真实世界场景中的有效性。
cs.CV / 20 / 2609.03378

When Depth Hurts: Reliability-Aware Geometry Distillation for Depth-Free RGB-D Salient Object Detection

当深度成为负担:面向无深度RGB-D显著性目标检测的可靠性感知几何蒸馏
Wang, Xuehao, Hua, Jiaxin, Li, Runmei, Wu, Zhenyu, Chen, Chenglizhao, Gu, Ke, Hao, Aimin
Abstract
Depth can resolve appearance ambiguity in RGB-D salient object detection (SOD), yet sensor depth is not uniformly reliable. Missing regions, blurred boundaries, and structural artifacts can propagate through multimodal fusion and make an RGB-D detector less accurate than its RGB-only counterpart. Existing quality-aware approaches regulate observed depth but remain dependent on the same potentially defective modality. We propose \method, a reliability-aware geometry distillation framework developed for RGB-D SOD benchmarks without using dataset-provided depth during training or inference. A frozen Depth Anything V2 model serves only as a training-time teacher, transferring dense relative geometry, hierarchical spatial attention, and boundary structure to a compact edge-aware geometry branch. Pooled bidirectional interaction aligns geometry with appearance, and a pixel-wise reliability estimator selectively injects geometry that is compatible with the current RGB representation. The teacher is removed after training, leaving an RGB-only inference network. Trained on 2,985 RGB-mask pairs, \method{} achieves the best or tied-best result in 26 of 36 metric-dataset comparisons against ten recent RGB-D SOD methods, including a 13.4\% relative MAE reduction on ReDWeb-S. When retrained on DUTS-TR, it also improves the strongest prior $F$-measure by 4.2\% on PASCAL-S, showing that the distilled geometry transfers beyond a particular sensor or dataset domain. Code will be released upon publication.
Chinese Translation
深度信息可以解决RGB-D显著性目标检测(SOD)中的外观歧义问题,但传感器深度并非始终可靠。深度缺失区域、模糊边界和结构伪影可能通过多模态融合传播,使RGB-D检测器的精度反而低于仅使用RGB的检测器。现有的质量感知方法虽然能对观测到的深度进行调节,但仍然依赖于同一可能存在缺陷的模态。我们提出\method,一个面向RGB-D SOD基准的可靠性感知几何蒸馏框架,在训练和推理阶段均不使用数据集提供的深度。冻结的Depth Anything V2模型仅作为训练阶段的教师,将稠密的相对几何、分层空间注意力和边界结构迁移到一个紧凑的边缘感知几何分支中。池化双向交互将几何信息与外观信息对齐,逐像素可靠性估计器则选择性地注入与当前RGB表征兼容的几何信息。训练完成后移除教师模型,仅保留纯RGB的推理网络。在2,985对RGB-掩码数据上训练后,\method{}与十种近期RGB-D SOD方法相比,在36项指标-数据集对比中有26项取得最优或并列最优的结果,其中包括在ReDWeb-S上MAE相对降低13.4%。当在DUTS-TR上重新训练时,其在PASCAL-S上将此前最优的F度量提升了4.2%,表明蒸馏所得的几何知识可以迁移到特定传感器或数据集领域之外。代码将在论文发表后发布。
cs.CV / 21 / 2609.03384

FoRIS: Progressive Foreground Refinement for Training-Free In-Context Segmentation

FoRIS:面向免训练上下文分割的渐进式前景细化
Hu, Ming, Yin, Jianfu, Dou, Mingyu, Zhang, Miaomiao, Wang, Yao, Hu, Cong, Hu, Bingliang, Wang, Quan
Abstract
In-Context Segmentation (ICS) aims to precisely segment arbitrary semantic concepts, such as objects or parts, given one or a few annotated visual exemplars. In this paper, we revisit ICS from a more classical segmentation perspective, viewing it as a coarse-to-fine progressive refinement process. Rather than directly predicting the final mask through reference-query matching, we progressively refine the segmentation from coarse and ambiguous foreground responses to precise and complete foreground structures. Building upon this perspective, we propose a training-free in-context segmentation framework, termed FoRIS. Specifically, FoRIS consists of three key stages: Foreground Purification, Foreground Localization, and Foreground Consolidation, which progressively suppress background distractions, localize discriminative target regions, and recover complete foreground structures through semantic aggregation. Experimental results demonstrate that FoRIS achieves SOTA performance across semantic and part segmentation tasks, with average improvements of 4.5 and 4.8 mIoU points over existing approaches in the 1-shot and 5-shot settings, respectively. Code: https://github.com/Xi-Mu-Yu/FoRIS.
Chinese Translation
上下文分割(In-Context Segmentation, ICS)旨在给定一个或少量标注的视觉示例时,精确地分割任意语义概念,如物体或部件。本文从更经典的分割视角重新审视ICS,将其视为一个由粗到细的渐进式细化过程。与通过参考-查询匹配直接预测最终掩码不同,我们从粗糙且模糊的前景响应出发,逐步细化出精确且完整的前景结构。基于这一视角,我们提出了一种免训练的上下文分割框架,称为FoRIS。具体而言,FoRIS包含三个关键阶段:前景净化(Foreground Purification)、前景定位(Foreground Localization)和前景巩固(Foreground Consolidation),依次渐进地抑制背景干扰、定位具有判别性的目标区域,并通过语义聚合恢复完整的前景结构。实验结果表明,FoRIS在语义分割和部件分割任务上均取得了SOTA性能,在1-shot和5-shot设置下分别比现有方法平均提升4.5和4.8个mIoU点。代码:https://github.com/Xi-Mu-Yu/FoRIS。
cs.CV / 22 / 2609.03391

Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data

探索对比语言-图像预训练在多源遥感数据中的潜力
Miao, Xiangyang, Yao, Kelu, Huang, Yekai, Xu, Xiaogang, Xue, Junxiao, Shen, Minjun, Lv, Chenghui, Liu, Shanji, Chen, Yaying, Li, Chao
Abstract
Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastive learning methods are mostly built on RGB-oriented CLIP architectures, making it difficult to exploit heterogeneous sensors such as SAR, multi-spectral imaging (MSI), and hyperspectral imaging (HSI). To address this limitation, we propose OmniRSCLIP, an end-to-end contrastive learning framework that supports multi-source sensor inputs for remote sensing vision-language modeling. The key idea is to extend CLIP beyond its fixed RGB input interface without breaking the pretrained visual knowledge. To this end, OmniRSCLIP introduces Spectral-Spatial Basis Decomposition (SSBD), which formulates arbitrary-channel adaptation as a basis recomposition problem: pretrained CLIP patch embeddings provide transferable spatial bases, while wavelength-conditioned coefficients span sensor-specific embedding kernels within a constrained visual prior space. This design avoids forcing heterogeneous sensors into a fixed-channel input space, while aligning them in a unified image-text semantic space. We further introduce a spectral-context-aware mask-based contrastive learning scheme to suppress modality-specific redundant features and enhance fine-grained image-text alignment. Finally, to support multi-modal training, we construct OmniRS5M, the first large-scale remote sensing image-text corpus covering RGB, SAR, MSI, and HSI. Experiments on retrieval, zero-shot classification, and semantic localization show that OmniRSCLIP preserves strong RGB-domain performance while effectively extending CLIP to heterogeneous remote sensing modalities.
Chinese Translation
对比语言-图像学习(CLIP)已成为遥感视觉-语言理解的关键范式。然而,现有的遥感对比学习方法大多建立在面向RGB的CLIP架构之上,难以充分利用诸如合成孔径雷达(SAR)、多光谱成像(MSI)和高光谱成像(HSI)等异构传感器。为解决这一局限,我们提出了OmniRSCLIP,一个支持多源传感器输入的端到端对比学习框架,用于遥感视觉-语言建模。其核心思想是在不破坏预训练视觉知识的前提下,将CLIP扩展到其固定的RGB输入接口之外。为此,OmniRSCLIP引入了光谱-空间基分解(Spectral-Spatial Basis Decomposition, SSBD),将任意通道的适配形式化为一个基重组问题:预训练的CLIP图像块嵌入提供可迁移的空间基,而基于波长的系数则在受限的视觉先验空间内生成传感器特定的嵌入核。这一设计避免了将异构传感器强行纳入固定通道的输入空间,同时在统一的图文语义空间中对齐它们。我们进一步提出了一种光谱上下文感知的掩码对比学习方案,以抑制模态特定的冗余特征并增强细粒度的图文对齐。最后,为支持多模态训练,我们构建了OmniRS5M,这是首个覆盖RGB、SAR、MSI和HSI的大规模遥感图文语料库。在检索、零样本分类和语义定位上的实验表明,OmniRSCLIP在保持强大RGB域性能的同时,能够有效地将CLIP扩展至异构遥感模态。
cs.CV / 23 / 2609.03406

Neural-Collapse-guided Task-Free Continual Anomaly Detection

神经坍缩引导的无任务持续异常检测
Kong, Xiaotong, Song, Chaoyang, Zhou, Ziai, Zhang, Jinxia, Zhang, Kanjian, Wei, Haikun
Abstract
Recent years have witnessed growing interest in continual anomaly detection for industrial visual inspection. However, real-world manufacturing environments exhibit unpredictable shifts in data distributions, rendering task-dependent continual learning assumptions impractical. To address this limitation, we formulate industrial anomaly detection as a task-free continual learning problem and propose NC-TFAD, a neural-collapse-inspired, geometry-driven framework for learning from non-stationary data streams without task boundaries. NC-TFAD freezes a pretrained backbone and aligns streaming features to a simplex Equiangular Tight Frame (ETF) prototype space to stabilize representation geometry under non-stationary streams. To satisfy the NC-inspired geometric construction in the absence of real anomalies, we generate synthetic anomaly samples as auxiliary anchors during training. Building on this geometry, we further introduce inter- and intra-class regularization together with a Focal Neural Collapse Contrastive (FNCC) loss to suppress representation drift and improve normal-anomaly separability. Finally, a normal-patch-prototype-guided localization branch constructs calibrated patch-wise deviation maps from normal training samples and fuses them with a weak self-attention prior, producing anomaly heatmaps without pixel-level annotations. Extensive experiments on MVTec AD and VisA show that NC-TFAD consistently outperforms representative task-free continual learning methods adapted from general vision, as well as unified anomaly detection baselines, in both image-level detection and pixel-level localization under the task-free continual learning protocol. These results highlight that geometry-driven modeling offers an effective and robust solution for task-free continual anomaly detection in real-world industrial applications.
Chinese Translation
近年来,面向工业视觉检测的持续异常检测日益受到关注。然而,现实制造环境中的数据分布呈现出不可预测的漂移,使得依赖任务的持续学习假设难以适用。为解决这一局限,我们将工业异常检测形式化为一个无任务持续学习问题,并提出 NC-TFAD——一个受神经坍缩(Neural Collapse)启发、由几何驱动的框架,能够在无任务边界的情况下从非平稳数据流中学习。NC-TFAD 冻结预训练的骨干网络,并将流式特征对齐到单纯形等角紧框架(ETF)原型空间,以在非平稳数据流下稳定表示几何。由于缺乏真实异常样本,为满足受神经坍缩启发的几何构建,我们在训练过程中生成合成异常样本作为辅助锚点。基于该几何结构,我们进一步引入类间与类内正则化以及焦点神经坍缩对比(FNCC)损失,以抑制表示漂移并增强正常样本与异常样本之间的可分性。最后,一个由正常图像块原型引导的定位分支,利用正常训练样本构建经校准的图像块级偏差图,并将其与弱自注意力先验融合,从而在无需像素级标注的情况下生成异常热力图。在 MVTec AD 和 VisA 数据集上的大量实验表明,在无任务持续学习协议下,NC-TFAD 在图像级检测和像素级定位两方面均持续优于从通用视觉领域迁移而来的代表性无任务持续学习方法以及统一异常检测基线。这些结果表明,几何驱动的建模为现实工业应用中的无任务持续异常检测提供了一种有效且鲁棒的解决方案。
cs.CV / 24 / 2609.03415

Mudragen: Geometrically Supervised Generation of Interacting Two-Hand Mudras for Preserving Indian Classical Dance Heritage

MudraGen:面向印度古典舞遗产保护的几何监督交互式双手手印生成
Kamble, Jagadish Kashinath, Mukhopadhyay, Jayanta, Roy, Debaditya, Das, Partha Pratim
Abstract
Automatic generation of hand gestures is essential for the transmission of Indian classical dance and critical for its preservation. Indian classical dance gesture datasets are inherently low-resource, and the canonical Sanskrit definitions of many mudras lack precise textual descriptions, limiting the effectiveness of conventional text-conditioned image generation models. We present \textbf{MudraGen}, a conditional diffusion framework that synthesizes realistic RGB images of \textit{Samyukta Hasta Mudras} -- interactive two-hand gestures from Bharatanatyam (an Indian classical dance form). Unlike prior work on simple hand signs or single-hand gestures, MudraGen introduces geometry-aware supervision to capture the precise coordination, anatomical validity, and cultural nuance of interacting hands. We formulate three geometry-aware objectives: Keypoint Loss for 3D joint alignment, Joint Offset Loss for inter-hand spatial coherence, and Shape Consistency, which serves as an anatomical regularizer by encouraging consistent hand morphology while allowing independent hand poses. Together, these objectives guide the diffusion model toward anatomically plausible and well-coordinated hand configurations, enabling the synthesis of photorealistic and pose-accurate gesture images. Experimental results show that MudraGen surpasses existing state-of-the-art generative approaches in visual realism, anatomical correctness, and preservation of fine hand-pose structure, enabling faithful reproduction of complex Samyukta Hasta mudras. Beyond quantitative gains, its ability to generate culturally grounded and structurally consistent gestures highlights practical applications in cultural preservation and dance education.
Chinese Translation
手势的自动生成对于印度古典舞的传承至关重要,对其保护更是不可或缺。印度古典舞手势数据集天然属于低资源数据集,且许多手印(mudra)的梵文规范定义缺乏精确的文本描述,这限制了传统文本条件图像生成模型的有效性。我们提出了MudraGen,一个条件扩散框架,用于合成逼真的Samyukta Hasta Mudras(婆罗多舞——一种印度古典舞形式——中的交互式双手手势)的RGB图像。与以往针对简单手语符号或单手手势的研究不同,MudraGen引入了几何感知监督,以捕捉双手交互的精确协调性、解剖学合理性以及文化细节。我们构建了三个几何感知目标:用于三维关节对齐的关键点损失(Keypoint Loss)、用于双手间空间一致性的关节偏移损失(Joint Offset Loss),以及作为解剖学正则化器的形状一致性(Shape Consistency)损失——它在鼓励手部形态一致的同时允许双手姿态独立。这些目标共同引导扩散模型生成解剖学上合理且协调良好的手部构型,从而实现照片级真实且姿态准确的手势图像合成。实验结果表明,MudraGen在视觉真实感、解剖学正确性以及精细手部姿态结构的保持方面超越了现有最先进的生成方法,能够忠实再现复杂的Samyukta Hasta手印。除定量指标的提升外,其生成具有文化根基且结构一致手势的能力,也凸显了其在文化保护和舞蹈教育中的实际应用价值。
cs.CV / 25 / 2609.03429

When Do Frozen VLMs Respond to Image-Free Object-Token Edits? An Answer-Key-Free Protocol and What It Reveals

冻结视觉语言模型(VLM)何时会响应无图像的对象令牌编辑?一种无需标准答案的评测协议及其揭示的规律
Son, Wonbin, Choi, Gyumun, Seo, Junil, Rho, Seungmin, Lee, Mi Young, Kim, Hyungjoon
Abstract
Answering what-if queries about a scene with a VLM usually means injecting the assumption as text or repainting the scene with a generative model. We instead move the edit to the representation level, before the model input. The image is abstracted into a set of object-level tokens, and the original image never enters the VLM. This design rests on an open question: when do frozen VLMs actually respond to such token edits? We introduce an answer-key-free protocol: no post-edit answer is annotated. It scores edits whose answers are logically determined, and audits itself by reversing each scoreable choice. The protocol reveals three structures. The response is not free: explicit edit teaching, not ordinary VQA training, produces it in dense scenes and multiplies it in sparse ones, on all three operations. Once on, it is governed by token cleanliness and density, with deployable detector+segmenter tokens competitive with the oracle and outperforming it on VRSBench. And reading is a separable axis: the image-free token route preserves 92-96% of a matched patch-token baseline's free-text VQA, and the answers measurably depend on the tokens. The response, cleanliness, and reading structures are sign-preserved across two remote-sensing datasets (iSAID, VRSBench) and three frozen LM backbones. We release the probe generator, records, judge logs, and code.
Chinese Translation
使用视觉语言模型(VLM)回答关于场景的“如果……会怎样”的查询,通常意味着将该假设以文本形式注入,或利用生成式模型重新绘制场景。我们则将编辑转移到模型输入之前的表示层面。图像被抽象为一组对象级令牌,原始图像从不进入VLM。这一设计依赖于一个悬而未决的问题:冻结的VLM究竟何时会响应此类令牌编辑?我们提出了一种无需标准答案的协议:不对编辑后的答案进行标注。该协议对答案可由逻辑确定的编辑进行评分,并通过反转每个可评分的选择来自我审计。该协议揭示了三种结构。第一,模型的响应并非天生具备:是显式的编辑教学,而非普通的VQA训练,产生了这种响应,并且在密集场景和稀疏场景中,对全部三种操作均能产生并放大该响应。第二,一旦具备响应能力,其表现受令牌的纯净度与密度支配,可部署的检测器+分割器令牌与预言机(oracle)相当,并在VRSBench上超越预言机。第三,读取是一个可分离的维度:无图像的令牌路径保留了与之匹配的图像块令牌基线92–96%的自由文本VQA能力,且答案可被证实依赖于这些令牌。响应、纯净度与读取这三种结构在两个遥感数据集(iSAID、VRSBench)和三个冻结语言模型骨干上均保持了符号一致性。我们公开了探针生成器、记录、评判日志和代码。
cs.CV / 26 / 2609.03445

OCR-EDR: Rendering-Aware Diagnosis and Repair for Closed-Loop OCR Improvement

OCR-EDR:面向闭环OCR改进的渲染感知诊断与修复方法
Zhao, Linnan, Liu, Kang, Yu, Hao, Zhan, Jiabo, Sun, Chong, Li, Chen
Abstract
Although document OCR systems perform increasingly well on routine documents, complex formulas, structured text, and long-tail formats remain error-prone. OCR predictions may omit fine-grained content or hallucinate unsupported outputs, while equivalent encodings of the same visible content must be accommodated. Existing OCR evaluation methods mostly report aggregate metrics, offering limited support for analyzing case-level errors and improving OCR performance. We propose OCR-EDR (OCR Error Diagnosis and Repair), a rendering-aware framework that advances from fine-grained diagnosis to iterative repair. Given a source image, an editable OCR prediction, and its rendered image, OCR-EDR first jointly assesses whether the prediction and its rendering are consistent with the source, preserving valid predictions, including rendering-equivalent ones, while diagnosing and localizing genuine errors. It then applies executable edits and may request an updated rendering for iterative reassessment. We construct OCRErrBench from diverse real OCR predictions, covering text and formulas, exact and rendering-equivalent positives, and genuine errors, and develop the DocEDR model to execute the diagnosis--repair loop. On OCRErrBench, DocEDR achieves 94.78% diagnostic accuracy. It repairs 86.23% of erroneous inputs to visual consistency, raises formula Case-F1 by 30.99 percentage points over DOCR-Inspector-7B on DOCRcaseBench, and improves formula CDM by up to 4.62 percentage points on the identified Bad subsets of four OCR systems on UniMER-Test. These results show that OCR-EDR turns fine-grained OCR analysis into verified corrections and performance gains.
Chinese Translation
尽管文档OCR系统在日常文档上的表现日益出色,但复杂公式、结构化文本和长尾格式仍然容易出错。OCR预测可能遗漏细粒度内容或产生无依据的幻觉输出,同时同一种可见内容的等价编码形式也必须被容纳。现有的OCR评估方法大多报告总体指标,对案例级错误分析和OCR性能提升的支持有限。我们提出OCR-EDR(OCR错误诊断与修复),这是一个从细粒度诊断到迭代修复的渲染感知框架。给定源图像、可编辑的OCR预测及其渲染图像,OCR-EDR首先联合评估该预测及其渲染结果是否与源图像一致,在保留有效预测(包括渲染等价的预测)的同时,诊断并定位真实错误。随后,系统执行可执行的编辑操作,并可能请求更新渲染图像以进行迭代重评估。我们基于多样化的真实OCR预测构建了OCRErrBench,涵盖文本和公式、精确正样本与渲染等价正样本以及真实错误,并开发了DocEDR模型来执行诊断-修复循环。在OCRErrBench上,DocEDR达到了94.78%的诊断准确率。它将86.23%的错误输入修复至视觉一致,在DOCRcaseBench上将公式的Case-F1较DOCR-Inspector-7B提升了30.99个百分点,并在UniMER-Test上对四个OCR系统识别出的Bad子集将公式CDM最多提升了4.62个百分点。这些结果表明,OCR-EDR能够将细粒度的OCR分析转化为经过验证的修正和性能提升。
cs.CV / 27 / 2609.03446

Preserving Knowledge across Space and Time for Continual Video Deepfake Detection

面向持续视频深度伪造检测的跨时空知识保持
Kim, Taehoon, Choi, Jongwook, Jo, Heejae, Park, Byungmin, Choi, Jongwon
Abstract
The continuous emergence of high-quality video deepfakes requires detectors that continually adapt to new forgery patterns, yet existing approaches, which are designed for deepfake images, fail to capture video-specific cues. Unlike deepfake images that contain only spatial artifacts, deepfake videos leave distinct evidence along both spatial and temporal axes, necessitating the separate preservation of each modality during sequential model updates. To overcome this limitation, we introduce a continual deepfake video detection framework, Modality-Specific Frequency Distillation (MSFD), that explicitly decomposes video features into spatial, temporal, and spatiotemporal modalities in the frequency domain. This decomposition enables independent preservation of each modality, as different deepfake video types exhibit varying reliance on spatial and temporal cues across tasks. Furthermore, MSFD adopts a cross-modality decorrelation loss that encourages spatiotemporal representations to remain orthogonal to single-modality cues. Extensive experiments show that our framework achieves stronger adaptation and preserves performance more effectively than state-of-the-art methods across diverse continual deepfake video scenarios.
Chinese Translation
高质量视频深度伪造(deepfake)的不断涌现,要求检测器能够持续适应新的伪造模式,然而现有的方法均针对深度伪造图像设计,无法捕获视频特有的线索。与仅包含空间伪影的深度伪造图像不同,深度伪造视频会在空间和时间两个维度上留下独特的痕迹,因此需要在序列化模型更新过程中对每种模态分别进行保持。为克服这一局限,我们提出了一种持续深度伪造视频检测框架——模态特定频率蒸馏(Modality-Specific Frequency Distillation, MSFD),该方法在频域中将视频特征显式分解为空间、时间和时空三种模态。由于不同类型的深度伪造视频在不同任务中对空间线索和时间线索的依赖程度各异,这种分解使得各模态能够被独立保持。此外,MSFD 采用跨模态去相关损失,促使时空表征与单模态线索保持正交。大量实验表明,在多种持续深度伪造视频场景中,我们的框架相比最先进的方法实现了更强的适应能力,并能更有效地保持检测性能。
cs.CV / 28 / 2609.03447

STARS-GS: Structure-Aware Regularized Gaussian Splatting for Large-Scale Aerial Surface Reconstruction

STARS-GS:面向大规模航拍表面重建的结构感知正则化高斯泼溅
Li, Bocheng, Zhang, Wenjuan, Han, Jie Pan. Dongxu, Ma, Xuesong, Yao, Yiling, Wang, Yaning
Abstract
Large-scale 3D surface reconstruction from aerial imagery is fundamental to geospatial mapping and urban modeling. Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated considerable potential for this task. However, existing methods still face three major challenges in large and complex scenes: scene partitioning may split continuous scene elements across independently optimized sub-regions; geometric constraints mainly focus on the attributes of individual Gaussians while overlooking their local organization; and uniform regularization struggles to accommodate heterogeneous geometric structures. To address these issues, we propose STARS-GS, a structure-aware 3DGS framework for large-scale surface reconstruction. First, we introduce a structure-aware scene partitioning strategy that better preserves continuous scene structures during partitioning and reduces cross-region geometric inconsistencies and stitching artifacts through boundary refinement. Second, we develop neighborhood-aware Gaussian organization that extends geometric constraints from individual primitives to their neighborhood organization, encouraging Gaussians to better conform to local surface geometry. Third, we introduce adaptive surface regularization that adjusts the regularization strength according to local geometric characteristics, promoting geometric consistency in structured regions while preserving plausible variations in unstructured regions. Extensive experiments on large-scale aerial photogrammetry benchmarks demonstrate that STARS-GS consistently outperforms the evaluated Gaussian-based methods in surface reconstruction. It increases the average F1-score from 0.640 for the second-best method to 0.698, corresponding to a relative improvement of approximately 9.1\%, demonstrating effective improvements in geometric accuracy and surface completeness.
Chinese Translation
基于航拍影像的大规模三维表面重建是地理空间制图与城市建模的基础任务。三维高斯泼溅(3D Gaussian Splatting, 3DGS)的最新进展在该任务上展现出可观潜力。然而,现有方法在大规模复杂场景中仍面临三大挑战:场景划分可能将连续的场景元素分割到独立优化的子区域中;几何约束主要关注单个高斯基元的属性,而忽略了其局部组织结构;统一的全局正则化难以适应异质的几何结构。针对这些问题,我们提出STARS-GS,一种面向大规模表面重建的结构感知3DGS框架。首先,我们引入一种结构感知的场景划分策略,在划分过程中更好地保留连续的场景结构,并通过边界细化减少跨区域的几何不一致性与拼接伪影。其次,我们提出邻域感知的高斯基元组织方法,将几何约束从单个基元扩展到其邻域组织,促使高斯基元更好地贴合局部表面几何。第三,我们引入自适应表面正则化,根据局部几何特征调整正则化强度,在结构化区域中促进几何一致性,同时在非结构化区域中保留合理的几何变化。在大规模航空摄影测量基准数据集上的大量实验表明,STARS-GS在表面重建任务上持续优于所评估的基于高斯的方法:其平均F1分数从次优方法的0.640提升至0.698,相对提升约9.1%,在几何精度和表面完整性方面取得了有效改进。
cs.CV / 29 / 2609.03453

Preprocessing Failure and Adversarial Detection in Depthwise-Separable Edge Vision Systems

深度可分离边缘视觉系统中的预处理失效与对抗样本检测
Mukta, Jannatul Masruk, Sanjida, Rifa, Tory, Adrita Rahman, Rahman, Md. Saifur, Hasan, Khondokar Fida
Abstract
Preprocessing-based defenses are the standard first-line response to adversarial attacks on edge vision systems, requiring no retraining, no architectural changes, and widely recommended as model-agnostic mitigations. Yet the foundational evaluations of these defenses were conducted on residual or Inception-class architectures, not on the depthwise-separable CNNs that dominate edge deployments. This untested assumption leaves a gap in the security evaluation literature. This paper closes that gap by evaluating six preprocessing defenses against adversarial perturbations across both architecture families. Across all perturbation levels and defenses tested, the two depthwise-separable architectures show consistently poor recovery while the residual architecture shows partial recovery; ablation results are consistent with an architectural rather than parametric explanation, though only three architectures and one attack family are evaluated. Crucially, this failure is not merely a negative result. The same output divergence that disqualifies preprocessing as a recovery mechanism reveals a detection opportunity: preprocessing consistently disrupts clean predictions while leaving adversarial predictions largely unchanged, an asymmetry that is directly measurable without retraining or architectural modification. We further show that standard image quality metrics are unreliable proxies for defense effectiveness, a methodological gap in current evaluation practice. A practitioner decision framework is provided for adversarially resilient edge vision deployment.
Chinese Translation
基于预处理的防御是应对边缘视觉系统对抗攻击的标准第一线响应,其无需重新训练、无需架构改动,并被广泛推荐为模型无关的缓解手段。然而,这些防御的基础性评估是在残差(ResNet)或 Inception 类架构上进行的,而非在主导边缘部署的深度可分离卷积神经网络(CNN)上。这一未经检验的假设在安全评估文献中留下了一个空白。本文通过在两类架构家族上评估六种预处理防御对对抗扰动的效果来填补这一空白。在所有测试的扰动强度和防御方法下,两种深度可分离架构均表现出持续较差的恢复能力,而残差架构则表现出部分恢复;消融实验结果与架构层面(而非参数层面)的解释相一致,尽管本文仅评估了三种架构和一类攻击方法。至关重要的是,这一失效并非仅仅是一个负面结果:使预处理失去恢复机制资格的同样输出差异,也揭示了一种检测机会——预处理始终会破坏干净样本的预测,而对对抗样本的预测基本不产生影响,这种不对称性无需重新训练或架构修改即可直接测量。我们进一步表明,标准图像质量指标是防御有效性的不可靠代理指标,这是当前评估实践中存在的方法学空白。本文还提供了一个面向对抗弹性边缘视觉部署的实践者决策框架。
cs.CV / 30 / 2609.03463

BMCTrack-d: Pig re-identification and tracking via back marks in challenging camera settings

BMCTrack-d:基于背部标记在复杂摄像环境下实现猪只重识别与跟踪
Brunner, David, Oczak, Maciej, Bordes, Marie, Rault, Jean-Loup, Winkler, Stephan M., Dorfer, Viktoria
Abstract
Automated pig monitoring is essential for assessing their health, behaviour, and welfare. To date, most pig monitoring solutions operate on the group-level, because individual-level monitoring requires reliable long-term identification and tracking of each animal. For domesticated pigs this remains challenging because pigs of the same breed often have highly uniform appearances. Moreover, research on pig monitoring is almost exclusively reported in top-down view camera settings, which considerably ease tracking, but are not always an option in practice. In this work, BMCTrack-d is presented, a novel tracking-by-detection approach that leverages unique back marks to enable robust pig re-identification and tracking in a challenging side-view camera setting, afflicted by rapidly moving pigs, severe occlusions and low resolution. The method first predicts the detected pigs' identities using a neural network-based back mark classifier. To improve re-identification reliability over time, two dedicated post-processing stages are introduced: a temporal prediction consistency check, which validates the identity assignments against the recent prediction history, and deduplication, which resolves conflicting identity assignments in each time step. By explicitly prioritising accurate, appearance-based re-identification over continuous tracking, the proposed approach addresses a key limitation of existing trackers for individual-level monitoring scenarios. On a demanding test set BMCTrack-d outperforms two strong baselines, BoT-SORT-ReID and TrackTrack-ReID, by 9.11% and 1.03%, respectively, in higher-order tracking accuracy. These results demonstrate the effectiveness of back mark-based re-identification and tracking for robust individual-level pig monitoring in challenging settings.
Chinese Translation
自动化猪只监测对于评估其健康、行为和福利至关重要。迄今为止,大多数猪只监测方案都是在群体层面运行的,因为个体层面的监测需要对每只动物进行可靠的长期识别和跟踪。对于家养猪而言,这仍然具有挑战性,因为同一品种的猪往往外观高度相似。此外,关于猪只监测的研究几乎全部基于俯视视角的摄像设置,这种设置虽然大大简化了跟踪,但在实际中并不总是可行。本文提出了BMCTrack-d,这是一种新颖的基于检测的跟踪方法,利用独特的背部标记,在充满挑战的侧视摄像环境下——猪只快速移动、严重遮挡和低分辨率——实现稳健的猪只重识别与跟踪。该方法首先使用基于神经网络的背部标记分类器预测所检测猪只的身份。为了提高重识别随时间的可靠性,引入了两个专门的后处理阶段:时间预测一致性校验(根据近期预测历史验证身份分配)和去重(解决每个时间步中冲突的身份分配)。通过明确地将准确的、基于外观的重识别置于连续跟踪之上,所提出的方法解决了现有跟踪器在个体层面监测场景中的一个关键局限。在一个高难度测试集上,BMCTrack-d在高阶跟踪精度上分别超越两个强基线方法BoT-SORT-ReID和TrackTrack-ReID 9.11%和1.03%。这些结果证明了基于背部标记的重识别与跟踪方法在复杂环境下实现稳健的个体层面猪只监测的有效性。
cs.CV / 31 / 2609.03475

SafeRestore: Detector-Relative Risk Certificates for Selective Industrial Image Restoration

SafeRestore:面向选择性工业图像复原的检测器相对风险证书
Yang, Shaoliang, Wang, Jun
Abstract
Industrial inspection pipelines often restore a measured image before a detector acts on it, yet restoration can suppress detector-supported defect structure or create clean-region activations. We formulate restoration as a selective action problem over the measured display, five restored candidates, and review. SafeRestore ranks candidates with action-specific fitted scores, chooses a gate on threshold-tuning data, and evaluates the fixed gate on a disjoint certification sample with two one-sided exact binomial bounds: one for the positive-conditional evidence-loss incident rate and one for the all-accepted excess-activation incident rate. The guarantee is marginal for one policy fixed before its certification outcomes are observed, under an image-level i.i.d. working model. In a retrospective split-sample study of 4,591 public Carinthia-S images, the protocol yields auditable risk-coverage behavior. The primary all-action policy passes in one of five training repetitions (12.0% +/- 26.9% pass-gated test coverage when failures count as zero), whereas fixed bicubic and reduced-complexity variants pass more often. On reserved morphologies, evidence-loss incidence rises to 81.1-90.3%, and KolektorSDD lacks both detector competence and enough positive certification images for the stated target. The contribution is therefore an auditable, detector-relative framework for deciding when a transformed image may be returned automatically and when review remains necessary -- not a claim that adaptive routing outperforms simpler policies on the present evidence.
Chinese Translation
工业检测流程中,检测器对测量图像进行处理之前往往先对其进行复原,然而复原可能抑制检测器所依赖的缺陷结构,或在干净区域产生激活。我们将复原形式化为一个选择性动作问题,其动作空间包括测量图像本身、五个复候选图像以及人工审核。SafeRestore 使用针对具体动作的拟合分数对候选进行排序,在阈值调优数据上确定门控阈值,并在一个不相交的认证样本上评估该固定门控,采用两个单侧精确二项界进行验证:一个针对“正类条件下证据丢失”事件率,另一个针对“全部接受情况下的多余激活”事件率。该保证是边际性的,适用于在观察到认证结果之前即已固定的单一策略,并基于图像级独立同分布(i.i.d.)的工作模型。在一项包含 4,591 张公开 Carinthia-S 图像的回溯性拆分样本研究中,该协议产生了可审计的风险-覆盖行为。主要的“全动作”策略在五次训练重复中仅通过一次(当失败计为零时,门控通过后的测试覆盖率为 12.0% ± 26.9%),而固定的双三次(bicubic)插值和降低复杂度的变体通过频率更高。在预留的形态学样本上,证据丢失发生率上升至 81.1%–90.3%,而 KolektorSDD 数据集既缺乏检测器能力,也没有足够的正样本认证图像来达到所述目标。因此,本文的贡献是一个可审计的、检测器相对的框架,用于决定何时可以自动返回经变换的图像、何时仍需人工审核——而非声称在现有证据下自适应路由优于更简单的策略。
cs.CV / 32 / 2609.03480

Tree species mapping in Denmark: A comparison of spectral-temporal features with geospatial foundation model embeddings

丹麦树种制图:光谱-时间特征与地理空间基础模型嵌入的比较
Koukos, Alkiviadis, Kondylatos, Spyros, Nord-Larsen, Thomas, Nyborg, Lotte, Tøttrup, Christian, Grogan, Kenneth
Abstract
We map tree species across Denmark using National Forest Inventory plots and EO data, while evaluating the potential of foundation models for large-scale forest characterization. We compare two alternative input representations for tree species classification: (i) manually engineered spectral-temporal features (STF) derived from multi-temporal Sentinel-1 and Sentinel-2 observations, and (ii) embeddings generated by the EO FMs TESSERA and AlphaEarth. Both representations are complemented with canopy height information. Random forest, XGBoost, and Multi-Layer Perceptron (MLP) classifiers are evaluated for all input representations, with separate assessments for pure and mixed forest stands. The STF-based MLP achieves the highest classification performance, yielding macro F1 scores of 0.843 and 0.653 for pure and mixed stands, respectively. The MLP trained on TESSERA embeddings delivers competitive performance for pure stands, achieving results within 1.1 percentage points of the best-performing model. TESSERA consistently outperforms STF-based models when fewer than approximately 25% of training plots are available, demonstrating a substantial advantage under limited training data. Multi-year observations systematically improve classification accuracy relative to single-year inputs, while ablation experiments reveal the complementary contributions of Sentinel-1 backscatter, spectral indices, and canopy height data. The best-performing model is subsequently applied at the national scale to generate a 10 m tree species map of Denmark. Area-adjusted validation indicates an overall map accuracy of 79.9%. The resulting map, released as an open-access product, is the first high-resolution national tree species map of Denmark and provides a valuable resource for forest monitoring, ecological research, and land management applications.
Chinese Translation
我们利用国家森林资源清查样地和地球观测(EO)数据对丹麦全境的树种进行制图,同时评估基础模型在大尺度森林表征方面的潜力。我们比较了两种用于树种分类的输入表示方法:(i)基于多时相Sentinel-1和Sentinel-2观测数据人工构建的光谱-时间特征(STF),以及(ii)由EO基础模型TESSERA和AlphaEarth生成的嵌入表示。两种表示方法均辅以冠层高度信息。针对所有输入表示,我们评估了随机森林、XGBoost和多层感知机(MLP)分类器,并对纯林和混交林分别进行了评估。基于STF的MLP取得了最高的分类性能,在纯林和混交林上的宏平均F1分数分别为0.843和0.653。基于TESSERA嵌入训练的MLP在纯林上表现出具有竞争力的性能,其结果与表现最佳的模型相差不超过1.1个百分点。当可用训练样地少于约25%时,TESSERA始终优于基于STF的模型,表明其在训练数据有限的情况下具有显著优势。多年观测相比单一年份输入能系统性提升分类精度,消融实验揭示了Sentinel-1后向散射、光谱指数和冠层高度数据的互补贡献。随后,将表现最佳的模型应用于全国尺度,生成了10 m分辨率的丹麦树种分布图。经面积加权的验证表明,该地图总体精度为79.9%。所生成的地图作为开放获取产品发布,是丹麦首张高分辨率国家树种分布图,为森林监测、生态研究和土地管理应用提供了宝贵的资源。
cs.CV / 33 / 2609.03516

Residual Optimal Transport-Based Experts Collaboration Towards Modality-Aware Infrared-Visible Object Detection

基于残差最优传输的专家协作面向模态感知的红外-可见光目标检测
Zhao, Yue, Yu, Hua, Zhao, Yukun, Zhang, Yuzhi, Gong, Maoguo, Mei, Xin, Hu, Zhuping, Li, Yanchi, Qin, A. K.
Abstract
Infrared-visible object detection (IVOD) integrates complementary evidence from visible and infrared sensors for reliable perception in challenging scenes. In practice, sensors may fail or drop frames, leaving one modality unavailable or intermittent. Existing methods for IVOD assume both modalities are always present, and fixed fusion collapses when one stream is missing. Furthermore, it remains a critical challenge to reliably estimate semantic correlation across heterogeneous modalities, especially under spectral distribution discrepancy. We present FlexibleFusion, a unified and adaptive method that flexibly allocates integration pathways and fusion strength, operating seamlessly across complete and missing-modality regimes. At its core, the Modality-Aware Experts Collaboration (MAEC) mechanism selectively activates and aggregates cross-modal or intra-modal expert pathways. It allows cross-modal fusion when full modalities are available and falls back to self-fusion under missing conditions. Additionally, we design Residual Self-Paced Entropic Optimal Transport (RSPEOT) to align heterogeneous feature distributions from a transport perspective. Instead of relying on the fixed sparsity coefficient in standard entropic optimal transport (EOT), RSPEOT introduces a residual-driven self-paced update that prioritizes reliable matches and progressively refines harder ones. This design alleviates the additional optimization burden of standard EOT while preserving reliable semantic alignment. Comprehensive experiments under complete and missing-modality protocols show consistent performance across arbitrary modality configurations. Code will be released upon publication.
Chinese Translation
红外-可见光目标检测(IVOD)融合来自可见光与红外传感器的互补信息,以在复杂场景中实现可靠感知。在实际应用中,传感器可能发生故障或丢帧,导致某一模态不可用或间歇性存在。现有的IVOD方法通常假设两种模态始终可用,一旦某一模态缺失,固定融合策略便会失效。此外,在异构模态之间可靠地估计语义相关性,尤其是在光谱分布差异下,仍然是一个关键挑战。我们提出FlexibleFusion,一种统一且自适应的方法,可灵活分配融合路径与融合强度,并在模态完整与模态缺失两种情形下无缝运行。其核心是模态感知专家协作(MAEC)机制,该机制选择性地激活并聚合跨模态或模态内的专家路径:当模态完整时进行跨模态融合,而在模态缺失时回退至自融合。此外,我们设计了残差自步进熵最优传输(RSPEOT),从传输的角度对齐异构特征分布。RSPEOT不再依赖标准熵最优传输(EOT)中固定的稀疏系数,而是引入残差驱动的自步进更新策略,优先处理可靠的匹配并逐步细化较难的匹配。该设计在保持可靠语义对齐的同时,减轻了标准EOT带来的额外优化负担。在模态完整与模态缺失两种协议下的全面实验表明,该方法在任意模态配置下均保持一致的性能。代码将于论文发表后开源。
cs.CV / 34 / 2609.03520

Neural Video Compression Based on Deformable Temporal Alignment and Difference-aware Fusion

基于可变形时序对齐与差异感知融合的神经视频压缩
Shan, Chuyue, Sun, Songlin, Chenwei, Wang, Zihan, Shen
Abstract
In conditional coding-based neural video compression, the quality of temporal context directly affects compression per- formance. Existing methods mostly construct context from prop- agated reference features, but they are vulnerable to motion esti- mation and local alignment errors in regions with complex mo- tion, occlusion, and high-frequency textures, resulting in inaccu- rate temporal information. To address this issue, this paper pro- poses a method combining deformable temporal alignment and difference-aware spatial selective fusion. A Context-aware Tem- poral Alignment Module is used to generate complementary tem- poral context, while a Difference-aware Spatial Selective Fusion module adaptively selects reliable temporal information and sup- presses misalignment. Experiments show that the proposed method achieves certain rate-distortion performance improve- ment over DCVC-DC.
Chinese Translation
在基于条件编码的神经视频压缩中,时序上下文的质量直接影响压缩性能。现有方法大多利用传播的参考特征构建上下文,但在复杂运动、遮挡和高频纹理区域,容易受到运动估计和局部对齐误差的影响,导致时序信息不准确。针对这一问题,本文提出了一种结合可变形时序对齐与差异感知空间选择性融合的方法。该方法使用上下文感知时序对齐模块(Context-aware Temporal Alignment Module)生成互补的时序上下文,同时利用差异感知空间选择性融合模块(Difference-aware Spatial Selective Fusion)自适应地选择可靠的时序信息并抑制错位影响。实验结果表明,所提出的方法相比 DCVC-DC 取得了一定的率失真性能提升。
cs.CV / 35 / 2609.03534

TruncGradGS: Improved 3D Gaussian Splatting via Truncated Gradient Updates

TruncGradGS:通过截断梯度更新改进3D高斯泼溅
Morales, Theo, Le-Pham, Nhat-Quynh, Atkins, Robin, Hua, Binh-Son
Abstract
3D Gaussian Splatting has become a de facto scene representation for novel view synthesis, yet robustly learning 3D Gaussian primitives from visual input remains challenging. Standard optimization relies on gradient-based updates, but a common issue is the gradient vanishing phenomenon: a pixel far from a Gaussian primitive often has diminishing gradient magnitudes to influence primitive attributes, resulting in suboptimal scene reconstruction. In this paper, we propose a method to address gradient vanishing with a piecewise truncated gradient formulation that improves the optimization stability and robustness to initializations. We show that our method consistently improves 3D Gaussian Splatting with random and COLMAP initializations while being generalizable across static and dynamic Gaussian Splatting. As a by-product, we also examine the limitations of current benchmarks for dynamic scenes, and introduce a novel dataset for benchmarking dynamic Gaussian Splatting using synthetic 3D scenes. We demonstrate the effectiveness of our method in both static and dynamic settings for the public benchmarks and our proposed dataset.
Chinese Translation
3D高斯泼溅(3D Gaussian Splatting)已成为新视角合成事实上的场景表示方法,然而从视觉输入中鲁棒地学习3D高斯基元仍然具有挑战性。标准优化依赖于基于梯度的更新,但一个常见问题是梯度消失现象:距离高斯基元较远的像素往往具有递减的梯度幅值,难以影响基元属性,导致场景重建效果欠佳。在本文中,我们提出一种解决梯度消失问题的方法,采用分段截断梯度公式,提升了优化的稳定性以及对初始化的鲁棒性。我们证明,无论采用随机初始化还是COLMAP初始化,我们的方法都能持续改进3D高斯泼溅,并且可以泛化到静态和动态高斯泼溅。作为附带贡献,我们还检验了当前动态场景基准测试的局限性,并基于合成3D场景引入了一个新的数据集,用于动态高斯泼溅的基准评估。我们在公开基准和所提出的数据集上验证了该方法在静态和动态设置中的有效性。
cs.CV / 36 / 2609.03544

SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language Models

SafeRI:面向大型视觉语言模型中词元级安全干预的识别与介入方法
Ma, Caoyuan, Gu, Tian, Liu, Wenpu, Xie, Weichu, Dong, Shuai, Xu, Yuqi, Zhao, Ji, Wang, Ziyue, Chang, Wenzheng, Wu, Taiqiang, Zhu, Yongfu, Shao, Wenqi, Wang, Zheng, Zheng, Yinqiang
Abstract
Existing safety alignment methods for vision-language models usually modify the model behavior globally: once the safety parameters are trained or loaded, they participate in both unsafe and already-safe generations. This always-on intervention can unnecessarily perturb the model's original reasoning path and degrade general multimodal capabilities. We argue that safety alignment should be an on-demand intervention rather than a permanent modification to every decoding trajectory. To this end, we propose a streaming recognition and gated LoRA framework for intrinsic VLM safety. During autoregressive generation, a lightweight recognizer estimates whether the current pre-token generation state is safe or unsafe. Its output updates the LoRA gate for the following decoding step; otherwise, generation follows the frozen-backbone policy. The LoRA module is trained from unsafe prefixes, transition statements, and safe continuations, so that it learns to redirect unsafe generations back to safe responses after activation. Experiments across multiple safety and general-purpose benchmarks demonstrate the effectiveness of our method in post-alignment settings.
Chinese Translation
现有的视觉语言模型安全对齐方法通常全局性地修改模型行为:一旦安全参数被训练或加载,它们就会同时参与不安全生成和已经安全的生成。这种始终开启的干预会不必要地扰动模型原有的推理路径,并降低其通用多模态能力。我们认为,安全对齐应当是一种按需进行的干预,而非对每条解码轨迹的永久性修改。为此,我们提出了一个用于视觉语言模型(VLM)内在安全的流式识别与门控LoRA框架。在自回归生成过程中,一个轻量级识别器估计当前词元生成前的状态是安全还是不安全。其输出会更新后续解码步骤的LoRA门控;否则,生成过程遵循冻结骨干网络的策略。该LoRA模块由不安全前缀、过渡语句和安全续写内容训练而来,从而学会在激活后将不安全生成重新引导回安全响应。在多个安全性和通用基准上的实验表明,我们的方法在后对齐设置中是有效的。
cs.CV / 37 / 2609.03554

WIDE: Wildcard Inference with Dynamic Expansion for Cross-Modal Generative Retrieval

WIDE:面向跨模态生成式检索的动态扩展通配符推理方法
Guo, Teng, Wang, Xin, Xu, Jiayou, Zhou, Keying, Shen, Jifeng, Ruan, Haoxin
Abstract
Generative retrieval has demonstrated significant success by unifying representation learning and search into a single sequence-to-sequence generation task. However, extending this paradigm to cross-modal retrieval reveals a critical challenge arising from the inherent information asymmetry across different modalities, such as the gap between concise text queries and dense visual candidates. This structural mismatch causes the autoregressive decoder to suffer from forced hallucination when generating identifiers via standard trie-constrained beam search, where the model is severely penalized for failing to guess fine-grained details absent from the query, allowing irrelevant candidates to hijack top rankings. To address this issue, we propose Wildcard Inference with Dynamic Expansion (WIDE). WIDE employs Adaptive Entropy Thresholding (AET) to calibrate layer-specific uncertainty boundaries offline. During the decoding generation phase, Asymmetry-aware Wildcard Decoding (AWD) detects semantic blind spots and emits wildcards instead of forced deterministic identifiers, dynamically expanding the search space without incurring log-probability penalties. Finally, Blind-Spot Re-ranking (BSR) evaluates the expanded candidate pool using a hybrid scoring mechanism that combines discrete generation confidence with continuous semantic similarity. Extensive experiments on the M-BEIR benchmark demonstrate that WIDE outperforms state-of-the-art generative retrieval methods, effectively suppressing forced hallucination while maintaining compact index structures.
Chinese Translation
生成式检索通过将表示学习与搜索统一到单一的序列到序列生成任务中,已展现出显著的成功。然而,将该范式扩展到跨模态检索时,面临着一个关键挑战:不同模态之间固有的信息不对称性,例如简洁的文本查询与密集的视觉候选之间的差距。这种结构上的失配导致自回归解码器在通过标准的字典树(trie)约束束搜索生成标识符时遭受“强制幻觉”问题——模型因无法猜测查询中缺失的细粒度细节而受到严重惩罚,从而使无关候选得以抢占排名前列。为解决这一问题,我们提出了动态扩展通配符推理(WIDE)。WIDE 采用自适应熵阈值(AET)在离线阶段校准各层特定的不确定性边界。在解码生成阶段,非对称感知通配符解码(AWD)检测语义盲区并输出通配符而非强制确定的标识符,从而在不引入对数概率惩罚的情况下动态扩展搜索空间。最后,盲区重排序(BSR)采用结合离散生成置信度与连续语义相似度的混合评分机制,对扩展后的候选池进行评估。在 M-BEIR 基准上的大量实验表明,WIDE 优于当前最先进的生成式检索方法,在保持紧凑索引结构的同时有效抑制了强制幻觉。
cs.CV / 38 / 2609.03557

Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation

构建世界模型预训练数据:基于虚幻引擎的动作条件视频生成流水线
Wang, Haoyu, Zhang, Songchun, Li, Haoran, Huang, Haoyang, Xue, Zeyue, Duan, Nan
Abstract
Action-conditioned video models require large-scale visual data paired with control signals that are temporally aligned with the resulting scene transitions. Such supervision is difficult to obtain from ordinary real-world video because the actions that caused each visual change are typically unknown. We present a large-scale synthetic data production pipeline built on Unreal Engine for generating action-conditioned, multi-view video. To accommodate the different execution requirements of real-time physics and high-quality offline rendering, the pipeline executes trajectory generation and final rendering in two stages: Stage I runs real physics in PIE and records per-frame character states, control inputs, and camera states into an intermediate trajectory representation; Stage II replays those trajectories in a new engine process and renders them offline with Movie Render Queue (MRQ). Around this core, we develop a distributed production system with cache-aware task partitioning, node-local slot scheduling, automated scene screening, aesthetic and luminance filtering, partial-output recovery, asynchronous upload, and continuous cluster health monitoring. The production cluster contains 25 servers with eight NVIDIA RTX 5090 GPUs per server. From 2,384 asset packs, 429 levels were retained for production together with a pool of 40 humanoid characters. The pipeline has produced 2,691 hours of 1080p video and 6,076 hours of 720p video. We describe the system architecture, the implementation decisions that emerged from production failures, and the limitations of using perceptual quality proxies for world-model data curation. The pipeline described in this report constitutes the Unreal Engine synthetic-data production component used in EchoWM.
Chinese Translation
动作条件视频模型需要与控制信号在时间上对齐的大规模视觉数据,且控制信号需与随之产生的场景变化相匹配。这类监督信号很难从普通真实世界视频中获取,因为导致每次视觉变化的动作通常是未知的。我们提出了一个基于虚幻引擎(Unreal Engine)的大规模合成数据生产流水线,用于生成动作条件的多视角视频。为了兼顾实时物理模拟和高质量离线渲染的不同执行需求,该流水线分两个阶段执行轨迹生成和最终渲染:阶段 I 在 PIE 中运行真实物理模拟,并将逐帧的角色状态、控制输入和相机状态记录到中间轨迹表示中;阶段 II 在新的引擎进程中重放这些轨迹,并使用 Movie Render Queue(MRQ)进行离线渲染。围绕这一核心,我们构建了一个分布式生产系统,具备缓存感知的任务划分、节点本地槽位调度、自动化场景筛选、美学与亮度过滤、部分输出恢复、异步上传以及持续的集群健康监控。生产集群包含 25 台服务器,每台配备 8 块 NVIDIA RTX 5090 GPU。我们从 2,384 个资产包中保留了 429 个关卡用于生产,并配有一个包含 40 个类人角色的角色池。该流水线已生产出 2,691 小时 1080p 视频和 6,076 小时 720p 视频。我们描述了系统架构、在生产故障中总结出的实现决策,以及使用感知质量代理指标进行世界模型数据筛选的局限性。本报告所述流水线构成了 EchoWM 中所使用的虚幻引擎合成数据生产组件。
cs.CV / 39 / 2609.03563

FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow

FlashRender:基于相机控制视频MeanFlow的少步生成式渲染
Park, Byeongjun, Kim, Byung-Hoon, Chung, Hyungjin
Abstract
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera trajectory in seconds. We identify sampling-step-dependent camera control as a prominent manifestation of discretization error in existing multi-step generative rendering models and show that resolving this inconsistency substantially lowers denoising trajectory curvature, facilitating subsequent step distillation. To this end, we introduce Representation Transformation and Alignment (RETA), which aligns hidden source-video representations with target-video features from a frozen visual geometry model. This directly encodes the geometric transformation within the source-video stream, enabling sampling-step-consistent camera control. We then fine-tune the model with the MeanFlow objective on the lower-curvature denoising trajectory induced by RETA, allowing the model to more effectively address discretization error. Finally, we apply on-policy flow map distillation to correct self-rollout errors under fixed few-step sampling. Extensive experiments show that RETA, MeanFlow, and on-policy flow map distillation play complementary roles in few-step generative rendering. Together, they enable our approach to match multi-step baselines in video quality and geometric consistency at 25x lower sampling cost while achieving superior camera controllability, even under out-of-distribution target camera trajectories.
Chinese Translation
我们提出了FlashRender,一个少步生成式渲染框架,能够在数秒内沿目标相机轨迹重新拍摄源视频。我们发现,依赖于采样步数的相机控制是现有多步生成式渲染模型中离散化误差的一个显著表现,并证明解决这种不一致性能显著降低去噪轨迹的曲率,从而有利于后续的步数蒸馏。为此,我们引入了表示变换与对齐(Representation Transformation and Alignment, RETA),将隐藏的源视频表示与来自冻结视觉几何模型的目标视频特征进行对齐。这直接在源视频流中编码了几何变换,实现了与采样步数一致的相机控制。随后,我们在RETA诱导的低曲率去噪轨迹上,采用MeanFlow目标对模型进行微调,使模型能够更有效地处理离散化误差。最后,我们应用在线策略流映射蒸馏(on-policy flow map distillation)来纠正固定少步采样下的自推演误差。大量实验表明,RETA、MeanFlow和在线策略流映射蒸馏在少步生成式渲染中发挥着互补作用。三者结合使我们的方法在视频质量和几何一致性上与多步基线相当,同时采样成本降低25倍,并实现了更优的相机可控性,即使在分布外的目标相机轨迹下也是如此。
cs.CV / 40 / 2609.03569

Occlusion-Robust Multimodal Emotion Recognition in VR via Fusion of Facial Images and EMG

基于面部图像与肌电信号融合的VR环境中抗遮挡多模态情感识别
Nierula, Birgit, Tomotaki-Dawoud, Karam, Akguel, Mert, Lafci, Mustafa Tevfik, Przewozny, David, Hilsmann, Anna, Eisert, Peter, Bosse, Sebastian
Abstract
Head-mounted displays (HMDs) fundamentally limit emotion recognition in virtual reality (VR): by occluding the upper face, they render conventional image-based facial expression analysis incomplete, particularly for applications requiring real-time affective assessment. We address this challenge by fusing lower-face video with facial electromyography (EMG) from the occluded upper face to classify seven emotional categories (six basic emotions plus neutral). We introduce a synchronized multimodal dataset from 20 participants, pairing lower-face video with seven-channel upper-face EMG elicited by validated emotion stimuli. Under subject-independent test, our proposed late-fusion architecture merging convolutional visual embeddings with RBF-kernel EMG representations achieves 51% macro-F1, outperforming both image-only (41%) and EMG-only (43%) baselines. These results demonstrate that upper-face EMG provides robust complementary information under HMD-induced visual occlusion and establish a foundation for multimodal emotion recognition in naturalistic VR environments. This approach facilitates affect-adaptive applications, including communication training and therapeutic interventions. The dataset will be shared upon request under an ethical-use agreement.
Chinese Translation
头戴式显示器(HMD)从根本上限制了虚拟现实(VR)中的情感识别:由于遮挡了上半部分面部,传统的基于图像的面部表情分析变得不完整,尤其是在需要实时情感评估的应用中。为应对这一挑战,我们将下半面部视频与被遮挡上半面部的表面肌电信号(EMG)进行融合,以分类七种情感类别(六种基本情绪加中性)。我们构建了一个来自20名参与者的同步多模态数据集,将下半面部视频与由经过验证的情绪刺激诱发的七通道上半面部EMG配对。在被试独立的测试条件下,我们提出的将卷积视觉嵌入与RBF核EMG表示相结合的后期融合架构取得了51%的宏平均F1值,优于仅使用图像(41%)和仅使用EMG(43%)的基线方法。这些结果表明,在HMD造成的视觉遮挡条件下,上半面部EMG能够提供鲁棒的互补信息,并为自然VR环境中的多模态情感识别奠定了基础。该方法可支持情感自适应应用,包括沟通训练和治疗干预。数据集将在伦理使用协议下应要求共享。
cs.CV / 41 / 2609.03572

Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

Drive-HWM:面向动态潜在表征引导自动驾驶的分层世界模型
Fan, Zhaoxin, Zhang, Tianbao, Wu, Wenjun, Wang, Xiaofeng, Jin, Yeying, Zhao, Jian, Zhu, Zheng, Yan, Shuicheng
Abstract
World models offer a promising paradigm for autonomous driving by predicting how traffic scenes may evolve and using such predictions to support action generation. However, existing approaches either separate future prediction from action generation or jointly predict them at the same temporal scale, making it difficult to simultaneously achieve long-horizon anticipation and responsive, observation-grounded decision making. We present Drive-HWM, a hierarchical slow--fast world modeling framework that organizes future representation prediction and action generation at complementary temporal scales. The slow world model predicts multi-step future representations to capture extended scene evolution. To explicitly model the abundant motion dynamics in driving environments, we introduce Dynamic-Aware Latents learned through optical-flow prediction. Guided by these future representations, the fast model uses a lightweight multimodal backbone and an autoregressive expert to jointly predict the next frame and the immediate action from the latest observation. Next-frame prediction encourages the fast model to capture imminent scene evolution, while one-step action generation allows decisions to be continuously updated as new observations arrive. Extensive experiments on NAVSIM v1 and v2 demonstrate the strong driving performance of Drive-HWM. Comprehensive ablation studies further validate the effectiveness of the hierarchical slow--fast design, dynamics-aware future representations, and joint next-frame and action prediction.
Chinese Translation
世界模型为自动驾驶提供了一种有前景的范式,其通过预测交通场景的演化并利用这些预测来支持动作生成。然而,现有方法要么将未来预测与动作生成分离,要么在同一时间尺度上联合预测二者,因而难以同时实现长时程预判与基于实时观测的快速响应式决策。我们提出 Drive-HWM,一个分层“慢-快”世界建模框架,它在互补的时间尺度上组织未来表征预测与动作生成。慢世界模型预测多步未来表征,以捕捉较长时间范围内的场景演化。为了显式建模驾驶环境中丰富的运动动态,我们引入了通过光流预测学习得到的动态感知潜在表征(Dynamic-Aware Latents)。在这些未来表征的引导下,快模型采用轻量级多模态骨干网络和自回归专家,从最新观测中联合预测下一帧图像和即时动作。下一帧预测促使快模型捕捉即将发生的场景演化,而单步动作生成则使决策能够随着新观测的到来持续更新。在 NAVSIM v1 和 v2 上的大量实验表明了 Drive-HWM 出色的驾驶性能。全面的消融实验进一步验证了分层“慢-快”设计、动态感知的未来表征以及下一帧与动作联合预测的有效性。
cs.CV / 42 / 2609.03585

Text2Thermal: Physics-Aware Thermal Image Synthesis from Textual Priors

Text2Thermal:基于文本先验的物理感知热成像图像合成
Qazi, Tayeba, Lall, Brejesh, Mukherjee, Prerana
Abstract
Thermal infrared imaging offers reliable perception in darkness and adverse weather, but thermal datasets remain scarce, motivating extensive work on translating abundant RGB images into thermal. Such translation is fundamentally ill-posed as thermal appearance is governed by surface emissivity and object temperature, neither of which is observable in the visible spectrum, so a single RGB image is consistent with many valid thermal outputs. We argue that language offers a natural means of resolving this ambiguity, and propose Text2Thermal, a framework for physics-aware thermal image synthesis from textual priors. Rather than inferring the unobservable radiometric factors from RGB, we supply them explicitly through thermally grounded captions encoding material, weather, time-of-day, and heat-emission state, and adapt a pretrained Stable Diffusion backbone to the thermal domain. Because the radiometric content is determined entirely by the prompt, Text2Thermal synthesizes thermal imagery without requiring a registered RGB image at inference; where spatial guidance is desired, an optional control signal imparts scene geometry without disturbing the prompt-specified radiometry. Experiments on M3FD, FLIR, and FMB show that Text2Thermal achieves state-of-the-art FID among thermal image synthesis methods while offering text-level control that translation-based approaches cannot provide.
Chinese Translation
热红外成像在黑暗和恶劣天气条件下能够提供可靠的感知能力,但热成像数据集仍然稀缺,这促使了大量研究致力于将丰富的RGB图像转换为热成像图像。然而,这种转换在本质上是病态的(ill-posed),因为热成像的外观由表面发射率和物体温度决定,而这两者在可见光谱中均不可观测,因此单张RGB图像可能对应多种有效的热成像输出。我们认为,语言为解决这一歧义提供了自然的手段,并提出了Text2Thermal——一个基于文本先验的物理感知热成像图像合成框架。与从RGB图像中推断不可观测的辐射因素不同,我们通过编码材料、天气、一天中的时间以及热发射状态的热相关文本描述(captions)显式地提供这些信息,并将预训练的Stable Diffusion骨干网络适配到热成像领域。由于辐射内容完全由文本提示(prompt)决定,Text2Thermal在推理时无需配准的RGB图像即可合成热成像图像;当需要空间引导时,可选的控制信号可以在不干扰文本提示所指定辐射内容的情况下提供场景几何信息。在M3FD、FLIR和FMB数据集上的实验表明,Text2Thermal在热成像图像合成方法中取得了最先进的FID指标,同时提供了基于转换的方法无法实现的文本级控制能力。
cs.CV / 43 / 2609.03602

SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

SV-WAM:一种用于端到端自动驾驶的高效环视世界-动作模型
Wang, Jinyang, Li, Shiwei, Wang, Junjian, Deng, Zhiqiang, Gao, Jianbin, Zhao, Yihang, Liu, Liu, Zhao, Yongjia, Chen, Jinlong, Xu, Huirui, Pan, Yifeng, Liu, Kangwei, Ren, Fan, Tao, Ji, Yang, Minghao
Abstract
World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive representations of future scene dynamics. However, generating future videos during inference introduces substantial computational overhead, leading many recent driving WMs to adopt a single front camera as input for efficient deployment. This design restricts spatial coverage in safety-critical maneuvers such as lane changes, merges, and turns. To address this limitation, we propose SV-WAM, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference. SV-WAM leverages future-video prediction as dense training supervision for action learning within a shared generative model, rather than as an inference-time output. At the core of this design is an action-centered causal mask that prevents action tokens from attending to future-video tokens during joint action-video denoising. Consequently, the video branch can be discarded at deployment, enabling efficient action-only planning. Furthermore, we introduce a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning safety and boundary awareness. Extensive experiments on the closed-loop NAVSIMv2 benchmark and the open-loop nuScenes benchmark demonstrate that SV-WAM achieves state-of-the-art planning performance with low inference latency and competitive zero-shot transfer capability.
Chinese Translation
世界模型(World Models, WMs)通过学习未来场景动态的预测性表征,在端到端自动驾驶中展现出了强大的潜力。然而,在推理阶段生成未来视频会带来巨大的计算开销,因此近年来许多驾驶世界模型采用单一前视相机作为输入以实现高效部署。这种设计限制了在变道、汇入和转弯等安全关键操作中的空间覆盖范围。为解决这一局限,我们提出了SV-WAM,一种环视世界-动作模型(World-Action Model, WAM),它在保持全部六相机观测的同时维持高效的推理。SV-WAM将未来视频预测作为共享生成模型中动作学习的稠密训练监督,而非推理时的输出。该设计的核心是一个以动作为中心的因果掩码(action-centered causal mask),在动作-视频联合去噪过程中阻止动作词元关注未来视频词元。因此,在部署时可以舍弃视频分支,实现仅基于动作的高效规划。此外,我们引入了一种可微分的可行驶区域一致性正则化项,对接近或越过可行驶边界的车辆轮廓角点进行惩罚,从而提升规划安全性和边界感知能力。在闭环NAVSIMv2基准和开环nuScenes基准上的大量实验表明,SV-WAM以较低的推理延迟实现了最先进的规划性能,并具备有竞争力的零样本迁移能力。
cs.CV / 44 / 2609.03615

Auditing Patient Privacy in Medical Generative Models: Scalable Memorization Detection with DeepSSIM++

医学生成模型中的患者隐私审计:基于DeepSSIM++的可扩展记忆化检测
Scardace, Antonio, Guarnera, Francesco, Battiato, Sebastiano, Ravì, Daniele
Abstract
While deep generative models offer new opportunities for medical image synthesis and data sharing, their ability to memorize and reproduce training samples raises serious concerns about patient confidentiality. Detecting such memorization at scale remains challenging: traditional pixel-based metrics are sensitive to generation artifacts, whereas generic embedding-based metrics often lack the anatomical sensitivity required for medical data. To address this challenge, we introduce DeepSSIM++, a self-supervised similarity metric for scalable memorization auditing in medical generative models. By leveraging multi-scale feature aggregation and anatomy-preserving augmentations, DeepSSIM++ learns an embedding space where cosine similarity approximates the Structural Similarity Index (SSIM), eliminating the need for exact pixel-level registration. Compared with state-of-the-art baselines, DeepSSIM++ achieves an average Macro F1 improvement of 33 percentage points under ideal alignment and 46 percentage points under realistic spatial and intensity perturbations. Furthermore, it accelerates large-scale similarity computation by several orders of magnitude compared with analytical SSIM. By combining anatomical sensitivity and computational efficiency, DeepSSIM++ provides an open-source tool for scalable memorization auditing in medical generative AI. Code and data are publicly available at: https://github.com/brAIn-science/DeepSSIM.
Chinese Translation
深度生成模型为医学图像合成与数据共享提供了新的机遇,但其记忆并复现训练样本的能力引发了患者隐私保护方面的严重担忧。大规模检测此类记忆化仍然具有挑战性:传统的基于像素的度量对生成伪影较为敏感,而通用的基于嵌入的度量往往缺乏医学数据所要求的解剖学敏感性。为应对这一挑战,我们提出了DeepSSIM++,一种用于医学生成模型可扩展记忆化审计的自监督相似性度量。通过利用多尺度特征聚合和解剖结构保持的数据增强,DeepSSIM++学习到一个嵌入空间,在该空间中余弦相似度可近似结构相似性指数(SSIM),从而无需精确的像素级配准。与最先进的基线方法相比,DeepSSIM++在理想对齐条件下平均Macro F1提升33个百分点,在真实的空间与强度扰动条件下提升46个百分点。此外,与解析式SSIM计算相比,它将大规模相似性计算的速度提升了数个数量级。通过结合解剖学敏感性与计算效率,DeepSSIM++为医学生成式AI的可扩展记忆化审计提供了一个开源工具。代码与数据已公开于:https://github.com/brAIn-science/DeepSSIM。
cs.CV / 45 / 2609.03629

EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders

EraseSAE:基于稀疏自编码器的文本到视频扩散模型中的精准概念擦除
Wang, Xinghao, Li, Dong, Yu, Wei, Pan, Yingwei, Gong, Tao, Chu, Qi, Yu, Nenghai, Yao, Ting
Abstract
Recent advances in text-to-video (T2V) diffusion models have demonstrated remarkable generative capabilities, yet their reliance on loosely curated training data raises pressing safety and copyright concerns. Concept erasure offers a principled remedy by removing unwanted semantics from pretrained models while preserving remaining concepts. However, existing approaches typically operate at a coarse granularity misaligned with the fine-grained, distributed nature of concept representations, leading to incomplete removal or degraded generation quality. We argue that surgical erasure fundamentally requires intervention at the level of monosemantic features, where each unit encodes a single interpretable concept. To this end, we propose EraseSAE, a novel framework that leverages sparse autoencoders to achieve surgical concept erasure in DiT-based T2V diffusion models via a principled decompose-attribute-erase pipeline. We first introduce the Partitioned Convolutional Sparse Autoencoder, which decomposes dense spatiotemporal activations into disentangled, interpretable sparse features while preserving spatiotemporal coherence. A contrastive attribution mechanism then contrasts activations from paired prompts to isolate concept-specific feature kernels. At inference, timestep-resolved spatiotemporal masks derived from the identified kernels confine erasure to regions where the target concept is active, leaving unrelated content intact. Extensive experiments across diverse diffusion models and concept erasure tasks demonstrate that EraseSAE achieves precise and robust concept removal with minimal quality degradation, substantially outperforming state-of-the-art methods. The code is available at https://github.com/HiDream-ai/EraseSAE.
Chinese Translation
文本到视频(T2V)扩散模型的最新进展展现了卓越的生成能力,但其对来源松散的训练数据的依赖引发了紧迫的安全与版权问题。概念擦除提供了一种有效的解决途径,即在保留其余概念的同时,从预训练模型中移除不想要的语义。然而,现有方法通常在粗粒度层面进行操作,与概念表示的细粒度、分布式特性不匹配,导致擦除不彻底或生成质量下降。我们认为,精准擦除从根本上需要在单语义特征层面进行干预,即每个单元编码单一可解释的概念。为此,我们提出了EraseSAE,一个利用稀疏自编码器(Sparse Autoencoders)在基于DiT的T2V扩散模型中实现精准概念擦除的新框架,采用“分解—归因—擦除”的原则化流程。我们首先提出了分区卷积稀疏自编码器(Partitioned Convolutional Sparse Autoencoder),将稠密的时空激活分解为解耦的、可解释的稀疏特征,同时保持时空连贯性。随后,对比归因机制通过对比成对提示词的激活,以分离出概念特定的特征核。在推理阶段,由识别出的特征核导出的时间步分辨的时空掩码将擦除限制在目标概念激活的区域,保持无关内容不受影响。在多种扩散模型和概念擦除任务上的大量实验表明,EraseSAE能够以最小的质量损失实现精准且鲁棒的概念移除,显著优于现有最先进方法。代码已发布于 https://github.com/HiDream-ai/EraseSAE。
cs.CV / 46 / 2609.03639

Stabilizing Camera-Controlled Novel View Synthesis at Inference Time

推理阶段稳定相机控制的 新颖视图合成
Singh, Prajwal, Badola, Arjun, Kumari, Seema, Nagahara, Hajime, Raman, Shanmuganathan
Abstract
Training-free, camera-controlled novel view synthesis from a single image using pre-trained video diffusion models often becomes unstable under large camera motion and long generation horizons. Existing approaches commonly combine several inference-time components, making it unclear which design choices are most important for stability. We show that the main source of stability is simple. Decomposing camera motion into small autoregressive steps limits per-step geometric distortion and reduces error accumulation. A controlled camera-step study shows that performance remains stable for small motions and degrades more strongly as the per-step motion approaches $18$-$20^\circ$. We further evaluate geometry-constrained spatial attention and low-frequency appearance anchoring as supporting refinements, together with an efficient registration-free warping pipeline. Across RealEstate10K and MegaScene, CamTrol++ improves temporal and geometric consistency, downstream 3D reconstruction quality, and generation efficiency over training-free baselines. The method remains effective for 56-frame generation and under substantial controlled depth corruption. These results show that careful control of camera motion at inference time can substantially improve the stability of camera-controlled novel view synthesis without retraining or modifying the diffusion backbone.
Chinese Translation
基于预训练视频扩散模型、从单张图像进行免训练相机控制的新颖视图合成,在大幅度相机运动和较长生成时程下往往变得不稳定。现有方法通常组合多种推理阶段组件,导致哪些设计选择对稳定性最为关键并不清晰。我们表明,稳定性的主要来源其实很简单:将相机运动分解为较小的自回归步骤,可以限制每一步的几何畸变并减少误差累积。一项受控的相机步长研究表明,在小幅运动下性能保持稳定,而当每步运动接近18–20度时性能会显著下降。我们进一步评估了几何约束的空间注意力(geometry-constrained spatial attention)和低频外观锚定(low-frequency appearance anchoring)作为辅助改进手段,并结合一个高效的免配准(registration-free)变形流水线。在RealEstate10K和MegaScene数据集上,CamTrol++相比免训练基线方法提升了时间与几何一致性、下游三维重建质量以及生成效率。该方法在56帧生成以及存在较大受控深度损坏的情况下仍然有效。这些结果表明,在推理阶段对相机运动进行精细控制,无需重新训练或修改扩散骨干网络,即可显著提升相机控制新颖视图合成的稳定性。
cs.CV / 47 / 2609.03641

Tree-Structured Vector Quantization For Efficient And Progressive Image Compression

用于高效且渐进式图像压缩的树状结构向量量化
Wang, Xinkun, Xu, Tianyi, Luo, Qingyu, Ma, Mingming, Jiao, Changzhe, Li, Fu, Niu, Yi
Abstract
Vector-quantization based image compression has achieved strong rate--distortion performance, yet most of them still produce a separate compressed representation for each target bitrate. Such variable-rate behavior allows one model to operate at multiple rates, but it does not necessarily provide a progressive bitstream whose prefixes are themselves decodable and can be refined by appending additional bits. We propose \textbf{Tree-VQ}, a progressive tree-structured vector quantization framework for learned image compression. Tree-VQ organizes discrete codewords as a hierarchical binary tree and represents each latent token by a routed root-to-leaf path. Crucially, every prefix of this path corresponds to a valid quantized representation, so shallow internal nodes serve as coarse reconstruction codes and deeper nodes provide successive refinements. This allows a compressed image to be decoded from an early prefix and progressively improved as more branch symbols are received, rather than being re-encoded for different target rates. To make this structure practical for compression, we introduce a prefix-compatible tree entropy model that codes progressive continuation decisions and routed branch refinements using only causally available decoded contexts. We further use rate-aware refinement scheduling to decide which spatial blocks should receive additional tree bits under a given prefix budget, and hierarchical prefix supervision to ensure that internal nodes are directly decodable at low rates. Experiments show that Tree-VQ achieves a superior performance--efficiency trade-off, delivering the best perceptual compression results with much fewer parameters and lower latency than competing methods.
Chinese Translation
基于向量量化的图像压缩已取得优异的率-失真性能,但大多数方法仍需为每个目标码率生成单独的压缩表示。这种可变速率特性虽然允许单一模型在多个码率下运行,但不一定能提供渐进式比特流——即其前缀本身可解码,并可通过追加额外比特进行细化。我们提出 extbf{Tree-VQ},一个用于学习型图像压缩的渐进式树状结构向量量化框架。Tree-VQ 将离散码字组织为层次化的二叉树,并通过一条从根到叶的路由路径来表示每个潜在token。关键在于,该路径的每个前缀都对应一个有效的量化表示,因此浅层内部节点充当粗略重建码,而更深的节点则提供逐步细化。这使得压缩图像可以从早期前缀开始解码,并随着接收到更多分支符号而逐步改善,而无需针对不同目标码率重新编码。为使该结构在压缩中切实可行,我们引入了一个前缀兼容的树熵模型,仅利用因果可得的已解码上下文对渐进式延续决策和路由分支细化进行编码。我们进一步采用率感知的细化调度,在给定前缀预算下决定哪些空间块应获得额外的树比特,并采用层次化前缀监督以确保内部节点在低码率下可直接解码。实验表明,Tree-VQ 实现了卓越的性能-效率权衡,以更少的参数和更低的延迟,提供了最优的感知压缩结果。
cs.CV / 48 / 2609.03655

PL-SCEA: Reconfiguring Pretrained Attention for Few-Shot Industrial Anomaly Detection

PL-SCEA:面向少样本工业异常检测的预训练注意力机制重构
Yang, Xiaoyu, Wu, Qixing, Zhao, Huixian, Jin, Changlong
Abstract
Vision Foundation Models (VFMs) provide transferable patch representations for few-shot industrial anomaly detection, but their attention computation is typically inherited from pretraining objectives centered on semantic aggregation. This creates a potential mismatch: token relations that support semantic recognition may not adequately expose the localized texture and structural deviations required for anomaly localization. We therefore investigate the hypothesis that the attention computation of a frozen VFM can be reconfigured as a task-relevant component of anomaly detection. We instantiate this idea with Power-Law Self-Correlation Enhanced Attention (PL-SCEA), which retains the semantic context of pretrained query-key attention while constructing token-adaptive self-correlations over contextualized value features. Positive-correlation filtering and power-law reweighting then emphasize relations that are salient relative to each token's relational background, without introducing additional trainable attention projections. The resulting features are modeled by a lightweight variational autoencoder that provides a fixed-size reconstruction-based representation of category-specific normality. The two stages serve complementary roles: attention reconfiguration shapes how local relational deviations are represented, while reconstruction-based modeling converts deviations from learned normality into anomaly scores. Across MVTec AD and VisA, the complete framework achieves competitive image-level detection and consistently strong pixel-level localization across the evaluated few-shot settings. Ablations further show that PL-SCEA improves localization with either the VAE or a memory bank under the tested setting. These results support the view that task-aligned attention reconfiguration can improve the anomaly-localization capability of frozen pretrained representations.
Chinese Translation
视觉基础模型(Vision Foundation Models, VFMs)为少样本工业异常检测提供了可迁移的图像块表征,但其注意力计算通常沿袭自以语义聚合为中心的预训练目标。这导致了潜在的失配:支持语义识别的标记(token)关系可能无法充分揭示异常定位所需的局部纹理和结构偏差。因此,我们研究这样一个假设:冻结的VFM的注意力计算可以被重构为异常检测中与任务相关的组件。我们通过幂律自相关增强注意力(Power-Law Self-Correlation Enhanced Attention, PL-SCEA)来实现这一想法,该方法在保留预训练查询-键注意力的语义上下文的同时,在上下文化的值特征上构建标记自适应的自相关关系。随后,正相关过滤和幂律重加权强调相对于每个标记关系背景显著的关系,且无需引入额外的可训练注意力投影。所得特征由一个轻量级变分自编码器建模,提供基于重构的固定尺寸表征,用以刻画特定类别的正常性。这两个阶段发挥着互补作用:注意力重构塑造局部关系偏差的表示方式,而基于重构的建模则将偏离所学正常性的偏差转换为异常分数。在MVTec AD和VisA数据集上,完整框架在所评估的少样本设置下实现了具有竞争力的图像级检测性能,并在像素级定位上始终保持强劲表现。消融实验进一步表明,在所测试的设置下,无论使用VAE还是记忆库,PL-SCEA均能提升定位性能。这些结果支持了如下观点:面向任务的注意力重构可以提升冻结预训练表征的异常定位能力。
cs.CV / 49 / 2609.03657

Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations

重新思考3D噪声:通过无需优化的形态学扰动学习3D感知视频先验
Şahin, Onat, Altillawi, Mohammad, Eskandar, George, Carbone, Carlos, Liu, Ziyuan
Abstract
3D scene representations like NeRF and 3D Gaussian Splatting (3DGS) suffer severe artifacts in sparse-view settings. Recent generative 3D artifact fixers attempt to address this, but rely on paired corrupted and clean renders requiring costly, per-scene reconstructions across varying view configurations. While 2D image augmentations act as instant regularizers, no explicit equivalents exist for 3D representations to preserve spatial consistency across views, an essential property for 3D-aware training. We propose 3D Morphological Perturbations as an optimization-free regularizer that preserves spatial consistency. Leveraging explicit 3DGS, we treat each Gaussian as a fundamental building block - analogous to a 2D pixel - and apply perturbations across its morphological parameter space via scale, rotation, and pruning. Our method eliminates per-scene 3DGS optimization loops from dataset curation while enabling models to learn stronger geometric priors than sparse-view baselines in diagnostic ablations conducted on a lightweight video diffusion sandbox. Scaled to a 14B-parameter video model via ControlNet, our approach maintains visual fidelity while reducing mean depth error by 12.5% over state-of-the-art image-to-image 3D artifact refiners, ultimately boosting downstream robotics policy success rates by up to 8.0% across 3 of 4 manipulation tasks.
Chinese Translation
NeRF和3D高斯泼溅(3D Gaussian Splatting, 3DGS)等3D场景表示方法在稀疏视角设置下会出现严重的伪影。近期的生成式3D伪影修复方法试图解决这一问题,但依赖于成对的损坏与干净的渲染结果,需要在多种视角配置下进行代价高昂的逐场景重建。虽然2D图像增强可作为即时的正则化手段,但目前尚不存在针对3D表示的显式等价方法来保持跨视角的空间一致性,而这一性质对3D感知训练至关重要。我们提出3D形态学扰动(3D Morphological Perturbations),作为一种无需优化的正则化方法,能够保持空间一致性。利用显式的3DGS表示,我们将每个高斯视为基本构建单元——类似于2D像素——并通过缩放、旋转和剪枝在其形态学参数空间上施加扰动。我们的方法在数据集构建过程中省去了逐场景的3DGS优化循环,同时使模型在轻量级视频扩散沙盒的诊断性消融实验中,学习到比稀疏视角基线更强的几何先验。通过ControlNet扩展至140亿参数的视频模型后,我们的方法在保持视觉保真度的同时,相比最先进的图像到图像3D伪影修复方法将平均深度误差降低了12.5%,并最终在4个操作任务中的3个上将下游机器人策略成功率提升了最高8.0%。
cs.CV / 50 / 2609.03663

Cross-Dataset Transfer and Reliability of Explainable Artificial Intelligence for RhythmFormer Remote Photoplethysmography

面向RhythmFormer远程光电容积描记术的可解释人工智能的跨数据集迁移与可靠性研究
Chen, Louis, Nordling, Torbjörn E. M.
Abstract
Background. Remote photoplethysmography estimates the cardiovascular pulse from facial video, and its explanations have rested on inspecting heatmaps rather than on quantitative evidence about where a model reads it. We quantified the explanations and asked whether such explanations transfer between datasets and track model performance. Method. We trained eight condition-specific RhythmFormer models on NCKU-rPPG, recorded under three illumination levels, speaking, rotation, and cycling, estimated one heart rate per 5.12-second clip, and set them beside a UBFC-rPPG reproduction. Raw attention, rollout, attention flow, and Beyond Intuition were assessed by skin coverage and the Salience-guided Faithfulness Coefficient (SaCo). Results. Beyond Intuition ranked highest on both datasets, at median coverage 0.789 and SaCo 0.837 on Static level 3 against 0.826 and 0.917 on UBFC-rPPG; lower ranks differed. Within one participant of one condition, neither measure was related to a clip's heart-rate error, waveform correlation, or signal-to-noise ratio on either dataset: 186 of the 252 coefficients fell below $|\rho|=0.10$ and 28 reached $p<0.05$ against the 13 expected by chance. Across the eight scenarios only Beyond Intuition's coverage followed the three performance measures, at $\rho=-0.43$, $+0.57$, and $+0.43$, while the attention-only methods' SaCo ran opposite to each. It failed at 40 lux alone, its median coverage falling to 0.180 and its median SaCo to $-0.178$, whereas motion degraded the estimates far more without such a drop. Conclusions. Skin coverage and SaCo carry information complementary to the performance measures rather than a proxy for them: attributing to the skin does not guarantee an accurate estimate. What an attribution reveals about a condition is where the model looks rather than how faithfully its map is ordered.
Chinese Translation
背景:远程光电容积描记术(remote photoplethysmography, rPPG)从面部视频估计心血管脉搏信号,而其解释工作一直依赖于对热力图的目视检查,而非关于模型从何处读取信号的定量证据。我们对解释进行了量化,并探讨此类解释能否在数据集之间迁移以及能否追踪模型性能。方法:我们在NCKU-rPPG数据集上训练了八个面向特定条件的RhythmFormer模型,该数据集涵盖三种光照水平以及说话、头部旋转和骑行动作,每个5.12秒的视频片段估计一次心率,并以复现的UBFC-rPPG模型作为对照。我们采用皮肤覆盖率(skin coverage)和显著性引导的忠实性系数(Salience-guided Faithfulness Coefficient, SaCo)来评估原始注意力(raw attention)、注意力回溯(rollout)、注意力流(attention flow)和Beyond Intuition方法。结果:Beyond Intuition在两个数据集上均排名最高:在Static level 3上中位覆盖率为0.789、SaCo为0.837,而在UBFC-rPPG上分别为0.826和0.917;较低排名的方法则有所不同。在单一参与者的单一条件下,无论在哪个数据集上,两项指标均与视频片段的心率误差、波形相关性或信噪比无关:252个相关系数中有186个低于|ρ|=0.10,仅28个达到p<0.05,而偶然预期的为13个。在八个场景中,只有Beyond Intuition的覆盖率与三项性能指标相关(ρ分别为-0.43、+0.57和+0.43),而仅基于注意力的方法的SaCo则与之相反。该模型仅在40 lux光照下失效,其中位覆盖率降至0.180,中位SaCo降至-0.178,而运动对估计结果的破坏远大于此,却没有出现类似的下降。结论:皮肤覆盖率和SaCo承载着与性能指标互补而非替代的信息:归因于皮肤区域并不能保证估计的准确性。归因方法所揭示的条件相关信息是模型关注的部位,而非其注意力图排序的忠实程度。
cs.CV / 51 / 2609.03668

ARCOS: Zero-shot Boundary Localization for Corneal Layer Segmentation Across Optical Coherence Tomography Devices

ARCOS:跨光学相干断层扫描设备的零样本角膜分层边界定位方法
Brás, Nuno Vivas, Memmi, Benjamin, Bouhassane, Maëlle, Georgeon, Cristina, Borderie, Vincent, Plamann, Karsten, Chessel, Anatole
Abstract
Accurate segmentation of corneal layers in optical coherence tomography (OCT) is essential for quantitative assessment of corneal morphology, including layer thickness and structural changes associated with disease or surgery. However, automatic segmentation remains challenging because corneal interfaces are thin, affected by speckle noise, and variable across acquisition devices. In this work, we propose ARCOS, a patch-based zero-shot boundary localization framework for corneal layer segmentation in clinical anterior-segment OCT images. Rather than performing conventional region classification, the method predicts boundary heatmaps for the main corneal interfaces from overlapping native-resolution patches. Patch-level predictions are stitched across the full B-scan and converted into boundary locations to obtain continuous, anatomically ordered layer segmentations. The network combines multi-scale feature fusion with a self-conditioned refinement module that uses intermediate boundary information to improve local heatmap predictions while preserving spatial detail. The method was evaluated on clinical OCT images acquired from multiple devices and compared with representative segmentation baselines using boundary localization and derived thickness metrics. The proposed method achieved an off-by-one boundary localization accuracy of 95.1% and a mean absolute boundary error of 0.514 pixels on the matched-device test set. In zero-shot cross-device evaluation, it maintained an average off-by-one accuracy of 84.3% and a mean absolute boundary error of 0.855 pixels across unseen acquisition devices, outperforming the baseline models. Thickness estimates derived from the predicted boundaries showed low error across corneal regions, supporting the method's use for quantitative corneal OCT analysis.
Chinese Translation
在光学相干断层扫描(OCT)图像中对角膜各层进行精确分割,对于角膜形态的定量评估至关重要,包括与疾病或手术相关的角膜层厚度及结构变化。然而,自动分割仍然具有挑战性,因为角膜界面非常薄、易受散斑噪声影响,且在不同采集设备之间存在差异。在本工作中,我们提出了 ARCOS,一个面向临床眼前节 OCT 图像的基于图像块的零样本边界定位框架,用于角膜分层分割。该方法并非执行传统的区域分类,而是通过重叠的原生分辨率图像块预测主要角膜界面的边界热力图。块级预测结果在整个 B 扫描上进行拼接,并转换为边界位置,从而获得连续且符合解剖学顺序的分层分割结果。该网络将多尺度特征融合与自条件化细化模块相结合,利用中间边界信息改进局部热力图预测,同时保留空间细节。该方法在来自多种设备的临床 OCT 图像上进行了评估,并使用边界定位及由此导出的厚度指标与代表性的分割基线方法进行了比较。在匹配设备测试集上,所提方法实现了 95.1% 的差一(off-by-one)边界定位准确率和 0.514 像素的平均绝对边界误差。在零样本跨设备评估中,该方法在未见过的采集设备上保持了平均 84.3% 的差一准确率和 0.855 像素的平均绝对边界误差,优于基线模型。由预测边界导出的厚度估计在各个角膜区域均显示出较低误差,支持了该方法在定量角膜 OCT 分析中的应用。
cs.CV / 52 / 2609.03673

Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation

视频生成器能否跨片段追踪世界?一个面向视频续写中世界状态推理的基准与方法
Miao, Yingmao, Zhang, Pengfei, Xu, Chaoran, Yu, Meng, Tang, Jing, Chu, Xiangxiang, Shen, Chao, Lin, Chenhao
Abstract
Video generators build long videos by composing shorter parts, either by generating segments one after another or by autoregressively extending chunks. Each new part usually depends on memories of historical observations, such as recent frames, selected key frames, memory banks, or cached features. These memories preserve visible evidence from the past, but current generators do not reliably turn such evidence into a world-state interface: what holds in the video world after previous actions and how it should change under the next prompt. A past frame remains valid history, but it may not describe the state needed by the next segment; some states must instead be inferred from occluded or implicit changes rather than copied from a directly observed frame. This creates a simple but overlooked question for video continuation: given a previous video, its prompt, and a new prompt, can a model generate a continuation that reflects the state determined by both the historical video and the new prompt? To answer this question, we introduce Statebench, a benchmark that targets this gap by testing continuations over three state categories: past-visible states, occluded-process states, and complex-transition states. We further propose Stateagent, which explicitly maintains an entity-state representation, updates it under the new prompt, grounds the predicted post-action state as a future end frame, and renders the next video. Experiments show that our method improves controlled video continuation by raising the all-case state score (SCS-All) from 45.2 to 69.3, and also benefits story generation at the one-minute scale. Code is avaliable at https://github.com/AMAP-ML/StateAgent.
Chinese Translation
视频生成器通过组合较短的片段来构建长视频,既可以逐个生成片段,也可以自回归地扩展视频块。每个新片段通常依赖于对历史观测的记忆,例如近期帧、选定的关键帧、记忆库或缓存特征。这些记忆保留了过去的可见证据,但当前的生成器无法可靠地将这些证据转化为世界状态接口:即在前序动作之后视频世界中保持的状态,以及它在下一个提示词下应如何变化。过去的帧仍是有效的历史,但未必描述下一片段所需的状态;某些状态必须从被遮挡或隐式的变化中推断得出,而非直接从已观测的帧中复制。这为视频续写引出了一个简单却长期被忽视的问题:给定一段先前的视频、其提示词和一个新的提示词,模型能否生成一个反映由历史视频和新提示词共同决定的状态的续写?为回答该问题,我们提出了 Statebench,一个针对这一空白的基准,通过三类状态类别测试视频续写:过去可见状态、遮挡过程状态和复杂过渡状态。我们进一步提出 Stateagent,其显式维护实体状态表示,在新提示词下更新该状态,将预测的动作后状态落实(ground)为未来末帧,并渲染下一视频。实验表明,我们的方法提升了受控视频续写的性能,将全类别状态得分(SCS-All)从 45.2 提升至 69.3,同时在一分钟尺度的故事生成任务上也带来收益。代码已发布于 https://github.com/AMAP-ML/StateAgent。
cs.CV / 53 / 2609.03675

CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding

CoFiE:面向高效流式视频理解的由粗到细证据选择方法
Jiang, Jing, Ling, Yiran, Li, Ruonan, Stamoulis, Dimitrios, Liu, Jie
Abstract
Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and answer user questions under tight latency constraints. Existing methods improve efficiency through token pruning and memory-bank schemes, but mainly reduce visual tokens after visual encoding. Consequently, downstream token pruning alone cannot substantially reduce end-to-end latency because the expensive frame encoding cost has already been incurred. We propose CoFiE, a Coarse-to-Fine Evidence Selection framework that decouples evidence selection into a coarse, query-agnostic filtering stage before the vision encoder and a fine, query-specific refinement stage during LLM prefill. CoFiE introduces Novelty-Guided Frame Filtering to retain visually distinctive candidate frames and Query-Specific Evidence Refinement to select the frames most relevant to the user query. This design removes substantial redundancy before frame encoding while preserving query-specific refinement once semantic information becomes available. Experiments show that CoFiE establishes a new state-of-the-art accuracy-efficiency trade-off across multiple video understanding benchmarks, reaching 78.86% accuracy on StreamingBench and 68.72% on OvO-Bench, with improvements of up to 3.15% over prior methods. Even with up to 80% evidence-frame filtering, CoFiE outperforms strong open-source multimodal models while improving end-to-end inference latency by up to 2.54 times.
Chinese Translation
流式视频理解要求视觉语言模型(VLLM)在严格的延迟约束下处理不断增长的的视频流并回答用户问题。现有方法通过令牌剪枝和记忆库机制提升效率,但主要在视觉编码之后削减视觉令牌。因此,仅依赖下游的令牌剪枝无法显著降低端到端延迟,因为昂贵的帧编码成本已经产生。我们提出CoFiE,一个由粗到细的证据选择框架,将证据选择解耦为两个阶段:视觉编码器之前的粗粒度、查询无关的过滤阶段,以及大语言模型(LLM)预填充过程中的细粒度、查询相关的精炼阶段。CoFiE引入新颖性引导的帧过滤以保留视觉上具有独特性的候选帧,并通过查询相关证据精炼来选择与用户查询最相关的帧。这种设计在帧编码之前去除了大量冗余,同时在语义信息可用时保留了针对查询的精细筛选。实验表明,CoFiE在多个视频理解基准上建立了新的准确率-效率权衡的最先进水平,在StreamingBench上达到78.86%的准确率,在OvO-Bench上达到68.72%,相比先前方法最高提升3.15%。即使过滤多达80%的证据帧,CoFiE仍优于强大的开源多模态模型,同时将端到端推理延迟最高降低至原来的1/2.54(即提速2.54倍)。
cs.CV / 54 / 2609.03677

Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language

通过自然语言描述图像子集差异来理解自动驾驶数据集
Truetsch, Julian, Hauser, Felix, Stiller, Christoph, Bieder, Frank
Abstract
Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness, and reliable operation across domains. For example, domain shift between locations could lead to the operating environment being misaligned with the training data, resulting in potentially dangerous performance degradation. Yet, existing data analysis pipelines largely rely on metadata, predefined labels, or manual inspection, which provide limited semantic insight or do not scale. This paper studies set difference captioning: given two subsets of images, the goal is to produce a natural-language hypothesis describing differences between the target and reference set. Building on a two-stage formulation, we adapt the method to autonomous driving by focusing on object-centric patches derived from object detection, which simplifies aggregation and enables attribution of differences to specific object instances or categories. To evaluate this setting in-domain, we introduce a new benchmark, AD-Diff Bench. Low-concentration experiments assess the suitability of set-difference-captioning approaches to sparse, real-world differences. We restrict our experiments to open-weight models to support reproducibility and ease of deployment. The proposed benchmark and analysis provide a step towards practical, human-interpretable dataset introspection for autonomous driving datasets. Our implementation and benchmark dataset are available at https://github.com/KIT-MRT/AD-Diff
Chinese Translation
理解大规模自动驾驶数据集的组成对于安全性、鲁棒性以及跨领域的可靠运行至关重要。例如,不同地点之间的领域偏移可能导致运行环境与训练数据不一致,从而引发潜在危险的性能下降。然而,现有的数据分析流程在很大程度上依赖元数据、预定义标签或人工检查,这些方法提供的语义洞察有限或难以扩展。本文研究集合差异描述(set difference captioning):给定两个图像子集,目标是生成一段自然语言假设,描述目标集与参考集之间的差异。基于两阶段的方法框架,我们通过聚焦于从目标检测中提取的以对象为中心的图像块(object-centric patches),将该适配到自动驾驶任务,这简化了聚合过程,并使差异能够归因于特定的对象实例或类别。为了在该领域内评估这一设置,我们引入了一个新的基准 AD-Diff Bench。低浓度实验评估了集合差异描述方法对稀疏的真实世界差异的适用性。我们将实验限制在使用开源权重模型,以支持可复现性和便于部署。所提出的基准与分析为面向自动驾驶数据集的实用、人类可解释的数据集内省迈出了一步。我们的实现和基准数据集可在 https://github.com/KIT-MRT/AD-Diff 获取。
cs.CV / 55 / 2609.03680

DropClick: Semi-Automated One-Click Segmentation for Agricultural Robotic Data

DropClick:面向农业机器人数据的半自动化一键式分割方法
Zimmer, Patrick, Halstead, Michael, McCool, Chris
Abstract
Labelling vision datasets, especially for segmentation tasks, is a laborious and costly process that stymies novel developments in agricultural robotics. In this paper, we present DropClick, a click-guided segmentation tool that simplifies the annotation process. Our system utilises single-click inputs on objects to generate pseudo-labels, which can replace manual annotations. DropClick stands out as it is a semi-automated approach and does not require a click for every object in the scene. It can therefore further reduce the required amount of user input drastically. We evaluate our method on two challenging agricultural robotic datasets, SB20 and BUP20 for plant and fruit segmentation, respectively. DropClick is first trained on a small subset of just 5 images from the original training data. This DropClick model can then be deployed as a one-click segmentation system and achieves comparable or higher performance than other one-click methods achieving an mIoU of 70.0 and 72.6 points, for SB20 and BUP20 respectively. DropClick then excels at maintaining high performance when clicks are not given (e.g. dropped); when 50% of the clicks are missing it still maintains an mIoU of 68.9 and 71.3 points, for SB20 and BUP20 respectively. We validate DropClick as a pseudo-labelling approach by taking its outputs to train a Mask2Former instance-based segmentation model in a semi-supervised manner. In this process, partially removing user input from DropClick yields similar high performance when compared to providing all clicks, at 70.1 vs 70.7 points AP50 for SB20 and no difference for BUP20 at 77.0 for both models; at the same time saving 46.3% of total input for SB20 and 31.9% for BUP20.
Chinese Translation
视觉数据集的标注工作,尤其是针对分割任务的标注,是一个费力且成本高昂的过程,阻碍了农业机器人领域的新发展。本文提出了一种点击引导的分割工具 DropClick,用于简化标注过程。我们的系统利用对目标的单击输入来生成伪标签,从而替代人工标注。DropClick 的突出之处在于它是一种半自动化方法,无需对场景中的每个目标都进行点击,因此可以进一步大幅减少所需的用户输入量。我们在两个具有挑战性的农业机器人数据集上评估了该方法:用于植物分割的 SB20 数据集和用于水果分割的 BUP20 数据集。DropClick 首先仅使用原始训练数据中 5 张图像组成的小子集进行训练。训练好的 DropClick 模型随后可作为一键式分割系统部署,其性能与其他一键式方法相当甚至更高,在 SB20 和 BUP20 上分别达到 70.0 和 72.6 的 mIoU。此外,DropClick 在缺少点击输入(即丢弃点击)时仍能保持高性能:当缺失 50% 的点击时,它在 SB20 和 BUP20 上仍分别保持 68.9 和 71.3 的 mIoU。我们将 DropClick 的输出用于以半监督方式训练基于实例的 Mask2Former 分割模型,从而验证了其作为伪标注方法的有效性。在此过程中,部分移除用户输入的 DropClick 与提供全部点击相比可获得相近的高性能:在 SB20 上 AP50 分别为 70.1 与 70.7,在 BUP20 上两种模型均无差异,均为 77.0;同时,该方法在 SB20 上节省了 46.3% 的总输入量,在 BUP20 上节省了 31.9%。
cs.CV / 56 / 2609.03688

ToPO: Token-Conditioned Preference Routing for Attention-Based Latent Diffusion Models

ToPO:面向基于注意力的潜在扩散模型的令牌条件偏好路由方法
Xu, Juntao, Li, Shihong, Au, Hoi Fan, Zhu, Ning
Abstract
Pairwise preference labels rank complete images, yet Diffusion-DPO applies their effect over many spatial and denoising-time coordinates. For attention-based, noise-prediction latent diffusion, ToPO (Token-Oriented Preference Optimization) constructs a per-minibatch, detached, separable spatial-temporal route from branchwise squared-residual contrast in a frozen reference denoiser. Preferred-branch cross-attention uses content tokens to modulate the spatial factor, and an auxiliary pixel-midpoint ordering term is added without local labels or a learned reward model. In matched three-seed retrainings with a shared update schedule, ToPO has higher endpoint estimates than Diffusion-DPO on all five reported SD-1.5 metrics and on HPSv2, ImageReward, and CLIP for SDXL. It also receives larger raw win shares in an aggregate blind SDXL A/B study. These findings are scoped to the reported equal-update U-Net protocols rather than an equal-compute comparison.
Chinese Translation
成对偏好标签是对完整图像进行排序的,然而 Diffusion-DPO 将其作用散布到众多空间坐标和去噪时间步上。针对基于噪声预测的注意力型潜在扩散模型,ToPO(Token-Oriented Preference Optimization,面向令牌的偏好优化)利用冻结参考去噪器中分支间的平方残差对比,构建了一种逐小批次、梯度截断且可分离的时空路由。偏好分支的交叉注意力利用内容令牌对空间因子进行调制,并在无需局部标签或学习奖励模型的情况下,引入了一个辅助的像素中点排序项。在共享更新调度、三种随机种子的匹配重训练实验中,ToPO 在 SD-1.5 报告的全部五项指标上,以及在 SDXL 的 HPSv2、ImageReward 和 CLIP 指标上,均获得了高于 Diffusion-DPO 的终点估计值。在汇总的 SDXL 盲测 A/B 研究中,ToPO 也获得了更大的原始胜出份额。上述结论仅限于报告中的等更新量 U-Net 协议,而非等计算量比较。
cs.CV / 57 / 2609.03689

Semantic-Aware Subgraph State Space Model for WSI Classification in Histopathology

面向组织病理学全切片图像分类的语义感知子图状态空间模型
Chen, Feixing, Lu, Hao, Luo, Lin, Xu, Yan
Abstract
Histopathological subtyping relies on the recognition of characteristic histological patterns. These patterns may be expressed by individual tissue structures or by the spatial distribution and co-occurrence of multiple structures, and they often span irregularly shaped tissue regions, termed semantic units in this work. However, conventional patch-based representations may fragment such units and fail to explicitly preserve their internal spatial organization, while efficiently modeling relationships among numerous spatially separated units remains challenging. To address these limitations, we propose the Semantic-Aware Subgraph State Space Model (SASG-SSM), a flexible and efficient framework for whole slide image (WSI) classification. Semantic-Aware Subgraphs (SASGs) first approximate irregularly shaped semantic units by adaptively grouping spatially connected patches guided by class-agnostic visual-semantic priors. By representing patches as graph nodes with adjacency edges, SASGs preserve their internal spatial organization rather than treating them as an unordered set. A Subgraph State Space Module (SG-SSM) subsequently combines a graph neural network encoder for intra-subgraph topology encoding with a Mamba-based state space encoder for efficient contextualization across large numbers of subgraphs. This module integrates local structural information within semantic units with global contextual information arising from their distribution and co-occurrence across the WSI, while efficiently modeling a large number of spatially distributed regions. Extensive experiments across four WSI subtyping datasets demonstrate consistent advantages over representative state-of-the-art methods. Further evaluations under small-cohort and few-shot settings demonstrate robustness and data efficiency under limited training data. Code will be released at https://github.com/HLSvois/SASG-SSM.
Chinese Translation
组织病理学分型依赖于对特征性组织学模式的识别。这些模式可能由单个组织结构表达,也可能由多个结构的空间分布与共生关系表达,并且往往跨越形状不规则的组织区域——本文将其称为语义单元。然而,传统的基于图像块(patch)的表示方法可能会割裂此类语义单元,无法显式保留其内部空间组织结构,同时对大量空间上相互分离的单元之间的关系进行高效建模仍然具有挑战性。为解决这些局限,我们提出了语义感知子图状态空间模型(Semantic-Aware Subgraph State Space Model, SASG-SSM),这是一个灵活高效的全切片图像(WSI)分类框架。语义感知子图(SASGs)首先在类别无关的视觉-语义先验的引导下,通过自适应地聚合空间上相连的图像块来近似不规则形状的语义单元。通过将图像块表示为图节点并建立邻接边,SASGs 保留了语义单元的内部空间组织结构,而非将其视为无序集合。随后,子图状态空间模块(Subgraph State Space Module, SG-SSM)将用于子图内部拓扑编码的图神经网络编码器与基于 Mamba 的状态空间编码器相结合,以在大量子图之间高效地进行上下文建模。该模块将语义单元内部的局部结构信息与其在 WSI 中的分布及共生所产生的全局上下文信息相融合,同时能够高效地对大量空间分布的区域进行建模。在四个 WSI 分型数据集上的大量实验表明,本方法相较于代表性的最先进方法具有一致的优势。在小样本队列和少样本(few-shot)设置下的进一步评估表明,本方法在有限训练数据下具有良好的鲁棒性和数据效率。代码将在 https://github.com/HLSvois/SASG-SSM 发布。
cs.CV / 58 / 2609.03690

MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CT

MetaStructAtlas:面向全身PET/CT功能与结构推理的定位式三维视觉-语言数据集与基准
Zheng, Chenguang, Xue, Le, Zhang, Yichi, Zhang, Wenbo, Ling, Zehui, Feng, Gang, Gao, Xin, Qi, Yuan, Cheng, Yuan, Hu, Zixin, Tian, Mei
Abstract
The joint interpretation of metabolic function and anatomical structure is essential for clinical diagnosis in whole-body PET/CT. Although recent advances in 3D medical vision-language models have demonstrated remarkable progress, current efforts are limited to regional CT imaging, leaving a critical void in comprehensive whole-body PET/CT analysis. In this work, we introduce MetaStructAtlas, a large-scale dataset for grounded whole-body PET/CT interpretation that synthesizes multimodal imaging with integrated anatomical, metabolic, and semantic annotations. MetaStructAtlas provides 490 co-registered 3D PET and CT volumes with 50,470 organ-level segmentation masks and grounded radiology reports. To facilitate interactive reasoning, we further developed MetaStructVQA, a standardized 3D grounded visual question-answering benchmark containing 100,565 QA pairs. This framework explicitly links diagnostic queries to visual evidence across modalities, encompassing anatomical, morphological, and metabolic characteristics. Finally, we evaluate state-of-the-art 3D medical VLMs on MetaStructVQA, establishing a robust foundation for multimodal representation learning and integrated whole-body reasoning in nuclear medicine.
Chinese Translation
代谢功能与解剖结构的联合解读对于全身PET/CT的临床诊断至关重要。尽管近年来三维医学视觉-语言模型取得了显著进展,但现有研究仍局限于区域性CT影像,在全身PET/CT的综合分析方面存在关键空白。在本工作中,我们提出了MetaStructAtlas,一个用于定位式全身PET/CT解读的大规模数据集,它将多模态影像与整合的解剖、代谢及语义标注相结合。MetaStructAtlas提供了490个体配准的三维PET与CT影像,包含50,470个器官级分割掩模以及定位式放射学报告。为促进交互式推理,我们进一步构建了MetaStructVQA,一个标准化的三维定位式视觉问答基准,包含100,565个问答对。该框架将诊断性提问与跨模态的视觉证据显式关联,涵盖解剖、形态及代谢特征。最后,我们在MetaStructVQA上评估了最先进的三维医学视觉-语言模型(VLM),为核医学中的多模态表征学习与全身综合推理奠定了坚实基础。
cs.CV / 59 / 2609.03694

Observation-Conditioned Latent Energy Priors for Sparse Implicit Neural Shape Completion

面向稀疏隐式神经形状补全的观测条件化潜在能量先验
Büschl, Paul, de la Rosa, Ezequiel, Wolleb, Julia, McGinnis, Julian, Nombela-Arrieta, César, Menze, Bjoern
Abstract
Implicit neural representations (INRs) can model continuous 3D shapes with a shared coordinate decoder and per-instance latent codes. At test time, autodecoder-style models commonly freeze the decoder and optimize a new latent code from sparse off-grid SDF samples. When these samples underconstrain inference, the latent can drift toward regions that fit the observations but decode implausible unobserved geometry. We propose a post-hoc observation-conditioned latent energy prior for frozen INR decoders. The energy scores standardized latents conditioned on a permutation-invariant encoding of the sparse observation set and is used as a residual expert alongside an L2 latent prior selected on validation data. We evaluate on a controlled cell-nucleus SDF dataset and a public MedShapeNet-derived SDF completion dataset. The proposed L2 objective augmented with conditional energy improves consistently over a validation-selected L2 baseline in the sparsest cell-nucleus regimes and, on MedShapeNet, outperforms both L2 and a six-component GMM latent-density prior across all reported readouts. A shuffled-context ablation is consistently weaker than matched context, supporting an observation-specific contribution. These results suggest that lightweight conditional energies can make pretrained INR decoders more observation-aware without retraining.
Chinese Translation
隐式神经表示(INR)可以通过共享坐标解码器和逐实例潜在编码对连续三维形状进行建模。在测试阶段,自解码器(autodecoder)风格的模型通常冻结解码器,并从稀疏的非网格SDF样本中优化一个新的潜在编码。当这些样本对推理的约束不足时,潜在编码可能漂移到能够拟合观测数据、但解码出的未观测几何形状不合理的区域。我们针对冻结的INR解码器提出了一种事后(post-hoc)观测条件化潜在能量先验。该能量函数基于稀疏观测集合的置换不变编码对标准化后的潜在编码进行评分,并作为残差专家(residual expert)与在验证数据上选择的L2潜在先验结合使用。我们在一个受控的细胞核SDF数据集和一个基于公开MedShapeNet衍生的SDF补全数据集上进行评估。所提出的L2目标在结合条件能量后,在最稀疏的细胞核场景中始终优于经验证选择的L2基线;在MedShapeNet上,其表现优于L2基线和六分量GMM潜在密度先验在所有报告指标上的结果。打乱上下文的消融实验结果始终弱于匹配上下文的结果,支持了观测特异性贡献的存在。这些结果表明,轻量级的条件能量可以在不重新训练的情况下使预训练的INR解码器更具观测感知能力。
cs.CV / 60 / 2609.03695

SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval

SignSeek:面向手语词典检索的可迁移表征学习
Asasi, Sobhan, Sincan, Ozge Mercanoglu, Bowden, Richard
Abstract
Sign language dictionaries are essential resources for sign language learners, yet automatically retrieving a sign from a dictionary, given only a query video, remains a challenging problem due to the natural variability between signers. Existing sign representation learning methods are built for closed-set recognition, producing embeddings that do not generalise to the open-set, signer-independent setting that retrieval demands. \textbf{SignSeek} closes this gap by contrastively learning sign representations with saliency-guided articulator masking. A contrastive objective aligns same-gloss signs across signers, while our Articulator Saliency-Guided Masking (ASGM) pinpoints the single most critical articulator per sign. This drives two complementary objectives, a masked contrastive alignment (MAC) loss that sees the sign through a single articulator and a masked prediction (MAP) loss that reconstructs it in latent space from the surrounding spatio-temporal context. Pretrained on 266K samples ($\sim$5,700 glosses) across multiple sign languages, \textbf{SignSeek} sets a new state-of-the-art performance in cross-corpus retrieval on ASL-Citizen, WLASL, and NMFs-CSL without any downstream fine-tuning. Strikingly, it achieves zero-shot generalisation to an entirely unseen British Sign Language (BSL), surpassing methods explicitly trained on BSL, and transfers seamlessly to isolated sign recognition and subtitle alignment, outperforming prior skeleton-based methods.
Chinese Translation
手语词典是手语学习者的重要资源,然而仅给定查询视频就从词典中自动检索手语仍是一个具有挑战性的问题,其原因在于不同手语使用者之间固有的变异性。现有的手语表征学习方法是为闭集识别而设计的,其生成的嵌入无法泛化到检索所要求的开集、与手语使用者无关的设置。SignSeek 通过对比学习结合显著性引导的发音部位掩码来学习手语表征,填补了这一空白。对比目标对齐不同手语使用者之间相同词条(gloss)的手语,同时我们提出的发音部位显著性引导掩码(Articulator Saliency-Guided Masking, ASGM)方法精确定位每个手语中最关键的单个发音部位。由此驱动两个互补的目标:一个是掩码对比对齐(MAC)损失,仅通过单个发音部位来观察手语;另一个是掩码预测(MAP)损失,从周围的时空上下文在潜空间中重建手语。SignSeek 在涵盖多种手语的 266K 样本(约 5,700 个词条)上进行预训练,在 ASL-Citizen、WLASL 和 NMFs-CSL 数据集上的跨语料库检索中,无需任何下游微调即取得了新的最先进性能。值得注意的是,它实现了对完全未见过的英国手语(BSL)的零样本泛化,超越了在 BSL 上显式训练的方法,并能无缝迁移到手语孤立词识别和字幕对齐任务,性能优于以往基于骨架的方法。
cs.CV / 61 / 2609.03729

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

展开世界:分解4D属性以强化空间推理
Yang, Yijun, Zheng, Shenghe, Li, Wenbo, Liu, Jianhui, Sun, Haoze, Zhang, Yanbing, Jiang, Jiaxiu, Song, Lin, Huang, Haoyang, Duan, Nan, Zhu, Lei
Abstract
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a profound dimensional mismatch: while VLMs are trained to interpret 2D projections, true spatial reasoning demands the recovery of latent 3D geometry and temporal continuity. To conquer this high-dimensional complexity, we advocate a shift from monolithic learning to a ``divide and conquer'' paradigm. We present FactoSR, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection. At its core, FactoSR decomposes the monolithic problem of world-consistent reasoning into three orthogonal, geometric sub-objectives: planar correspondence ($XY$), depth consistency ($Z$), and temporal reversibility ($T$). By optimizing these verifiable constraints within a unified policy learning mechanism, we effectively transform an ill-posed projection recovery problem into a series of tangible reasoning steps. Extensive evaluations on multi-view and video benchmarks demonstrate that this elegant decomposition yields substantial gains in 3D and 4D reasoning, achieving a 5.9% boost on VSI-Bench and 4.5% on All-Angles-Bench. Our findings suggest that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.
Chinese Translation
尽管视觉-语言模型(VLMs)在通用多模态任务中表现出卓越的能力,但它们在推理物理世界时本质上仍是“扁平”的。我们认为,这一空间瓶颈源于一种深层次的维度失配:VLMs被训练用于解读2D投影,而真正的空间推理则要求恢复潜在的3D几何结构和时间连续性。为了攻克这种高维复杂性,我们主张从整体式学习转向“分而治之”的范式。我们提出了FactoSR,一个因子化的强化学习框架,能够显式地解读被视觉投影所压缩的各个维度。FactoSR的核心在于将世界一致性推理这一整体性问题分解为三个正交的几何子目标:平面对应关系(XY)、深度一致性(Z)和时间可逆性(T)。通过在统一的策略学习机制中优化这些可验证的约束,我们有效地将一个不适定的投影恢复问题转化为一系列具体可操作的推理步骤。在多视角和视频基准上的大量评估表明,这种优雅的分解在3D和4D推理方面带来了显著提升,在VSI-Bench上取得5.9%的提升,在All-Angles-Bench上取得4.5%的提升。我们的研究结果表明,强化显式、因子化的4D一致性是将VLMs进化为鲁棒的、具备世界感知能力的推理器的关键一步。
cs.CV / 62 / 2609.03740

Fill My Mirror: Geometry-Constrained Mirror Inpainting

填补我的镜子:几何约束的镜子图像修复
Basson, Ofek, Vainer, Shimon, Hel-Or, Yacov, Fried, Ohad
Abstract
Mirrors are common in real-world images, yet producing geometrically consistent reflections with generative models remains challenging. Unlike most objects, mirror appearance depends on scene geometry and viewpoint, making it hard to synthesize using learned appearance priors alone. We address this in the mirror inpainting setting, where the scene is fixed and only the mirror region is generated. Our key insight is that much mirror content is geometrically constrained by the visible scene and need not be hallucinated. We estimate scene geometry and project visible content into the mirror to recover reflection regions determined by geometry. A generative model then completes the mirror region via a two-mask diffusion strategy balancing geometric constraints with the model's learned priors, reducing projection artifacts and improving reflection consistency. The method is training-free and applicable to complex real-world scenes. We evaluate on MirrorBench-V2 (synthetic) and real images. Using standard and geometry-aware metrics, we show that explicitly using scene geometry improves consistency.
Chinese Translation
镜子在真实世界图像中十分常见,然而利用生成式模型生成几何一致的反射仍然具有挑战性。与大多数物体不同,镜子的外观取决于场景几何和视点,仅凭学习到的外观先验难以合成。我们在镜子图像修复(mirror inpainting)的场景中解决这一问题,即场景固定,仅生成镜子区域。我们的关键洞察是:许多镜子内容受可见场景的几何约束,无需凭空生成。我们估计场景几何,并将可见内容投影到镜子中,以恢复由几何决定的反射区域。随后,生成模型通过双掩码扩散策略补全镜子区域,该策略在几何约束与模型学习到的先验之间取得平衡,从而减少投影伪影并提升反射一致性。该方法无需训练,可应用于复杂的真实场景。我们在 MirrorBench-V2(合成数据)和真实图像上进行评估。通过标准和几何感知指标,我们表明显式利用场景几何能够提升一致性。
cs.CV / 63 / 2609.03742

KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

KnowVis:面向视频讲座的知识中心视觉摘要生成方法
Xu, Yi, Hou, Yifan, Zhang, Xiaoyu
Abstract
Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice learners. This difficulty stems from a fundamental pedagogical mismatch: while videos deliver transient information linearly, human learning requires constructing interconnected cognitive networks, a task that induces severe cognitive overload for novice learners lacking prior domain knowledge. Existing video summarization methods fail to resolve this mismatch, as they primarily produce text-heavy, linear condensations that still demand high cognitive effort. To bridge this gap, we propose KnowVis, a framework that transforms linear video lectures into pedagogically grounded visual narratives. KnowVis first extracts a detailed concept map from multimodal video content to identify important and challenging threshold concepts, then constructs structured knowledge units, and finally synthesizes engaging visual summaries. Alongside the framework, we introduce a curated dataset of 125 educational videos across 10 academic disciplines, paired with 1,079 generated visual summaries. Extensive automated evaluations and a human study demonstrate that, compared to state-of-the-art baselines, KnowVis generates more accurate and clear visuals that successfully reduce cognitive load and significantly improve student learning effectiveness and knowledge retention.
Chinese Translation
视频讲座是宝贵的教育资源,但其内容密集且冗长的形式常常令初学者不堪重负。这一困难源于一个根本性的教学错位:视频以线性方式传递转瞬即逝的信息,而人类学习则需要构建相互关联的认知网络——对于缺乏先验领域知识的初学者而言,这是一项极易造成严重认知负荷过载的任务。现有的视频摘要方法未能解决这一错位,因为它们主要生成以文本为主、线性的内容压缩形式,仍然需要较高的认知投入。为弥合这一差距,我们提出了KnowVis,一个能够将线性视频讲座转化为具有教学依据的视觉叙事的框架。KnowVis首先从多模态视频内容中提取详细的概念图,以识别重要且具有挑战性的门槛概念(threshold concepts),然后构建结构化的知识单元,最终生成引人入胜的视觉摘要。与该框架相配套,我们构建了一个精心整理的数据集,涵盖10个学科的125个教育视频,并配有1,079个生成的视觉摘要。大量的自动化评估和用户研究表明,与最先进的基线方法相比,KnowVis生成的视觉内容更加准确、清晰,能够有效降低认知负荷,并显著提升学生的学习效果和知识保持能力。
cs.CV / 64 / 2609.03756

ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation

ENEAS:面向自适应分割的嵌入引导神经集成方法
del Pino, Javier, Rodríguez, Salvador, Garabito, Alejandro, Álvarez, Javier, Garabito, Chema
Abstract
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality, so that visually similar artifacts such as statues, paintings, or reflections are segmented as target entities. ENEAS works two ways from a single method: precise tracking and high-quality segmentation of a unique instance, and open-concept discovery of every instance a text query names, resolved by a semantic verification layer. For tracking, we extend the geometrically robust SeC architecture, previously limited to point interactions, with a text-prompting adapter and leverage its temporal memory, so that the target is held through disappearance without drifting to distractors and kept whole even when it fills the entire view. For discovery, the verification layer combines high-speed visual embedding matching with conditional VLM refinement, invoking semantic reasoning only for ambiguous candidates, which filters out the ontological errors that visual-only models cannot distinguish while keeping latency low. Designed with 3D reconstruction in mind, where a single misclassified distractor corrupts the asset, ENEAS unlocks high-quality semantic tracking and segmentation of video, of broad libraries, and of collections of temporally or spatially unordered data, together with the discrimination to tell true instances from their doppelgangers: things that look alike but are not the same. The code and models are available at https://github.com/speridlabs/eneas
Chinese Translation
我们提出了ENEAS,一种统一的、可文本提示的实例跟踪与语义发现方法。可文本提示的分割模型,包括诸如SAM 3等最新基础模型,仍然存在时间幻觉、空间碎片化和语义误分类等问题:当目标物体离开视野时,它们无法报告目标缺失;在极端特写场景下,它们会分割局部纹理而非完整物体;并且它们优先考虑视觉特征而非本体现实,以至于将视觉上相似的物体(如雕像、绘画或反射影像)误分割为目标实体。ENEAS通过单一方法实现双重功能:一是对特定实例的精确跟踪与高质量分割,二是对文本查询所命名的所有实例进行开放式概念发现,并由语义验证层进行判定。在跟踪方面,我们在几何鲁棒的SeC架构(此前仅限于点交互)基础上扩展了文本提示适配器,并利用其时间记忆能力,使目标在消失后仍可被持续跟踪而不漂移至干扰物,即使目标充满整个视野也能保持完整。在发现方面,验证层将高速视觉嵌入匹配与条件性VLM精化相结合,仅对模糊候选调用语义推理,从而过滤掉纯视觉模型无法区分的本体错误,同时保持低延迟。ENEAS专为3D重建场景设计——在3D重建中,单个误分类的干扰物即可损坏资产——它为视频、大型媒体库以及时间或空间上无序的数据集合实现高质量的语义跟踪与分割,并具备区分真实实例与其“相似孪生物”(即看似相同实则不同的事物)的判别能力。代码和模型已发布于 https://github.com/speridlabs/eneas
cs.CV / 65 / 2609.03773

RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents

RealCADBench:基于工业设计意图的参数化CAD建模基准测试
JoyIndustrial VisCAD Team, Cai, Linxin, Hong, Qiuhe, Huang, Zhichao, Li, Guanlin, Li, Zongzhen, Liu, Hongsen, Long, Yichen, Wang, Wei, Wang, Yuchen, Yang, Dongyue, Yu, Huimu, Zhong, Xianwen
Abstract
Parametric computer-aided design (CAD) modeling is difficult to evaluate with a single metric. Existing CAD benchmarks often emphasize synthetic or CAD-native settings, limited input modalities, or executability and IoUs alone. We introduce RealCADBench, a benchmark for intent-to-program CAD modeling from real industrial design intents. It contains 12,632 tasks from 19 factory-automation categories and spans text descriptions, 2D engineering drawings, real product pictures, and rendered images for both Part and Assembly modeling. We report results on a 1,770-task evaluation slice: 1,745 Part tasks across four input regimes and RCB-Assm25, a 25-task assembly study used in every reported assembly comparison. Each method generates FreeCAD API Python, which a shared runtime executes to export the 3D model. We evaluate the exported model using executability, Solid IoU, Surface IoU, and a rubric-based visual-semantic identity Judge. Among the nine standalone frontier large models evaluated, no model leads all four metrics. Across six frontier-scale large models, executability ranges from 0.565 to 0.812, Solid IoU from 0.2841 to 0.5379, and Surface IoU from 0.112 to 0.217 across the four Part regimes. The highest regime-balanced composite comes from a different model than the leaders on the four component metrics. On RCB-Assm25, Codex with GPT-5.5 improves executability and both IoU metrics over standalone GPT-5.5, but lowers the Judge score by 6.98 percentage points, leaving GPT-5.5 as the Judge leader. We also observe recurring failure modes, most notably missing fine structures, loss of part identity, and incorrect assembly placement. These results show that execution alone is insufficient to characterize realistic CAD modeling and that frontier models and agents differ substantially across executability, IoUs, and visual-semantic identity.
Chinese Translation
参数化计算机辅助设计(CAD)建模难以用单一指标进行评估。现有的CAD基准测试通常侧重于合成或CAD原生场景、有限的输入模态,或仅关注可执行性与交并比(IoU)。我们提出了RealCADBench,一个面向从真实工业设计意图到程序生成的CAD建模基准。该基准包含来自19个工厂自动化类别的12,632个任务,涵盖文本描述、2D工程图纸、真实产品照片和渲染图像,并同时支持零件(Part)和装配(Assembly)建模。我们在一个1,770个任务的评估子集上报告结果:其中包括四种输入模式下的1,745个零件任务,以及RCB-Assm25——一个在所有装配对比中使用的25个任务的装配研究集。每个方法生成FreeCAD API的Python代码,由共享运行时执行以导出3D模型。我们使用可执行性、实体IoU(Solid IoU)、表面IoU(Surface IoU)以及基于评分细则的视觉-语义一致性Judge来评估导出的模型。在评估的九个独立前沿大模型中,没有任何模型在全部四项指标上领先。在六个前沿规模的大模型中,四种零件输入模式下的可执行性范围为0.565至0.812,实体IoU为0.2841至0.5379,表面IoU为0.112至0.217。模式均衡综合得分最高的模型与四项单项指标上的领先者并不相同。在RCB-Assm25上,Codex结合GPT-5.5相比独立的GPT-5.5提升了可执行性和两项IoU指标,但Judge得分降低了6.98个百分点,GPT-5.5仍是Judge指标上的领先者。我们还观察到反复出现的失败模式,最显著的是缺失精细结构、零件特征丢失以及装配位置错误。这些结果表明,仅凭执行结果不足以刻画真实的CAD建模能力,前沿模型与智能体在可执行性、IoU和视觉-语义一致性方面存在显著差异。
cs.CV / 66 / 2609.03788

A Reverse Sign Language Dictionary: Open-Vocabulary Sign Recognition from Continuous Signing via Video Captioning and Description Retrieval

一种反向手语词典:通过视频描述生成与描述检索,从连续手语中实现开放词汇手语识别
Poveda-Gutiérrez, Santiago, Nakayama, Hideki, Bono, Mayumi
Abstract
Isolated Sign Language Recognition (ISLR) is conventionally cast as closed-set classification over gloss labels, which cannot generalize to signs unseen in training and ties every deployment to a gloss-annotated lexicon. We instead recognize signs extracted from continuous signing by (1) captioning a sign-level clip into a free-form procedural description of the articulation with an open-weight vision-language model, and (2) retrieving the closest entry from a vocabulary of target descriptions with a multilingual sentence encoder: a reverse sign language dictionary that needs no gloss supervision and admits an open vocabulary. On 1,300 sign-level segments from a Japanese Sign Language (JSL) dialogue corpus annotated with procedural descriptions (against a 2% top-10 chance floor over the 503-entry target vocabulary), fine-tuning the captioner substantially improves seen-class retrieval: language and vision tower fine-tuning raises top-10 retrieval on seen classes from 4.5% (untrained) to 49%, becoming statistically indistinguishable from a standard supervised closed-set classifier (I3D) on two of the three test sets where a closed-set classifier can be evaluated at all. More importantly, unseen-class retrieval also improves significantly over the untrained pipeline (11.5% -> 21.0% top-10, p=0.0094), a regime in which the closed-set classifier cannot participate. A matcher-side empirical upper-bound analysis shows the sentence encoder already recovers close to 100% of paraphrased gold descriptions, locating a gap in captioning quality that we aim to address in future work. To our knowledge this is the first description-based, open-vocabulary sign lookup from continuous signing without gloss supervision, and the first for JSL.
Chinese Translation
孤立手语识别(Isolated Sign Language Recognition, ISLR)通常被构建为基于词汇标签(gloss)的封闭集分类任务,这既无法泛化到训练中未见过的手语,也将每次部署都与一个词汇标注词典绑定。我们转而通过以下两步识别从连续手语中提取的手语词:(1)使用开源权重的视觉-语言模型,将手语级视频片段描述为关于手部动作的自由格式程序性描述;(2)使用多语言句子编码器,从目标描述词汇库中检索最接近的条目——这构成了一部无需词汇标注监督、支持开放词汇的反向手语词典。在一个标注了程序性描述的日本手语(Japanese Sign Language, JSL)对话语料库的1,300个手语级片段上(在503个条目的目标词汇库上,top-10随机基线为2%),对描述生成模型进行微调显著提升了已见类别的检索性能:语言塔和视觉塔的微调使已见类别的top-10检索率从4.5%(未训练)提升至49%,在三个可以评估封闭集分类器的测试集中的两个上,其表现与标准监督封闭集分类器(I3D)在统计上不可区分。更重要的是,未见类别的检索性能相比未训练流程也显著提升(top-10从11.5%提升至21.0%,p=0.0094),而封闭集分类器无法参与这一设置。匹配器一侧的经验上界分析表明,句子编码器已能近乎100%地检索回改写后的标准描述,从而将差距定位于描述生成质量上,这是我们未来工作要解决的问题。据我们所知,这是首个从连续手语中、无需词汇标注监督的、基于描述的开放词汇手语查询方法,也是首个面向日本手语的此类方法。
cs.CV / 67 / 2609.03796

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

LLaDA-Image:基于完全开放训练配方构建强大的图像生成器
Chen, Chuyan, Chen, Haoxing, Chen, Kun, Cheng, Zhenglin, Cui, Long, Fang, Ruishan, Gu, Zhangxuan, Huang, Zhicheng, Lan, Zhenzhong, Lei, Yuanting, Li, Haoquan, Li, Jianguo, Li, Rongchuan, Li, Sidu, Lin, Tao, Liu, Deyuan, Liu, Jiacheng, Liu, Lin, Lou, Yuxuan, Lu, Zhisheng, Ma, Yuxin, Shen, Shuheng, Sun, Peng, Wang, Chaoyang, Wang, Hongjun, Wang, Xiaomei, Wang, Yongxin, Wu, Chengzhang, Wu, Hongru, Xie, Jun
Abstract
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.
Chinese Translation
我们提出了LLaDA-Image,一个统一框架,将一个从零开始训练的60亿参数扩散Transformer(Diffusion Transformer, DiT)与一个基于LLaDA2.0-Mini扩散语言模型骨干构建的冻结视觉-语言理解模块相结合。我们没有从一开始就大量依赖图文配对数据,而是首先通过仅图像的预训练和中期训练构建强大的视觉生成先验。生成流水线共包含2.2亿个样本,其中9800万个为真实图像。为实现高效且可扩展的优化,我们在整个DiT中全程使用无参数的RMSNorm,并采用Muon优化器。由此得到的统一模型能够生成高度逼真的图像,同时准确遵循细粒度的编辑指令。我们进一步将LLaDA-Image蒸馏为LLaDA-Image-Turbo,使其能够在2-4个采样步骤内实现快速推理。在Qwen-Image-Bench上,LLaDA-Image在英文和中文赛道上分别取得53.53和53.38的总体得分,在两条赛道上均创下了开源模型的新最优水平。为支持对强大且高效生成模型的进一步研究,我们公开了模型权重、训练代码以及详细的训练配方。
cs.CV / 68 / 2609.03804

Urban Boundaries, Social Barriers: A Benchmark and Vision-Centric Framework for Mapping Gated Communities and Equity Implications

城市边界,社会屏障:面向封闭社区制图及其公平性影响研究的基准数据集与以视觉为中心的框架
Zhao, Minwei, Zhang, Weiming, Du, Jiawang, Liu, Qiming, Zhuang, Weiming, Nie, Pei, Wu, Cai
Abstract
Communities are fundamental spatial units that shape urban form and social life. Whether a residential compound is spatially open or enclosed affects mobility, access to public services, and equity, yet studies of Chinese fengbi xiaoqu remain largely qualitative or small-scale, limiting reproducible city-scale analysis. We address this gap by introducing GBA-GCs, a metropolitan-scale multimodal benchmark for locally grounded gated/open community recognition in China's Greater Bay Area, covering 37,444 residential compounds with aligned boundary polygons, high-resolution satellite imagery, Chinese metadata, and structured attributes, together with expert-verified labels, inter-annotator reliability, and official evaluation splits. Built on this benchmark, we present Multimodal Classifier for Gated Community (MCGC), a vision-centric multimodal framework based on DINOv3-SAT that fuses imagery, text, and structured cues via modality-aware cross-attention and adaptive gating to mitigate modality imbalance. MCGC consistently outperforms strong unimodal and multimodal baselines. Finally, we apply the validated model to metropolitan-scale mapping and report equity-oriented findings including spatial clustering of GCs, privatized green space, and reduced pedestrian connectivity. The benchmark, code, and release documentation are available at https://github.com/MinweiZhao/GBA-GCs.
Chinese Translation
社区是塑造城市形态与社会生活的基本空间单元。住宅小区在空间上开放或封闭,会影响居民出行、公共服务可达性以及社会公平。然而,针对中国封闭小区(封闭小区)的研究大多停留在定性分析或小规模研究,难以开展可复现的城市尺度分析。为填补这一空白,我们提出了GBA-GCs,一个面向中国粤港澳大湾区本地化封闭/开放社区识别的大都市尺度多模态基准数据集,涵盖37,444个住宅小区,包含对齐的边界多边形、高分辨率卫星影像、中文元数据和结构化属性,并配有专家核验的标签、标注者间一致性检验以及官方评估划分。基于该基准,我们提出了面向封闭社区的多模态分类器(MCGC),这是一个以视觉为中心的多模态框架,基于DINOv3-SAT,通过模态感知交叉注意力和自适应门控机制融合影像、文本和结构化信息,以缓解模态不均衡问题。MCGC在性能上持续优于强大的单模态和多模态基线模型。最后,我们将经过验证的模型应用于大都市尺度制图,并报告了与公平性相关的发现,包括封闭社区的空间聚集、绿地私有化以及行人连通性降低等问题。基准数据集、代码及发布文档已发布于https://github.com/MinweiZhao/GBA-GCs。
cs.CV / 69 / 2609.03811

VisCAD: A Foundation Model Suite with Multimodal Industrial CAD Intelligence

VisCAD:具备多模态工业CAD智能的基础模型套件
JoyIndustrial VisCAD Team, Cai, Linxin, Hong, Qiuhe, Huang, Zhichao, Li, Guanlin, Liu, Hongsen, Liu, Ziqi, Long, Yichen, Wang, Luya, Wang, Yuchen, Wu, Wenxiang, Yu, Huimu, Zhang, Ning
Abstract
AI-assisted computer-aided design (CAD) for industrial products involves two challenging phases. Part-level generation maps diverse forms of user intent, including renders, text descriptions, 2D drawings, and real photographs, to executable programs in a CAD domain-specific language. Assembly-level generation must additionally handle interacting parts, plan mating relations, estimate poses, and place all parts correctly. Existing specialized CAD models are commonly trained on narrow input domains, such as renders or texts, and often generalize poorly, while general-purpose frontier models cover broader inputs but perform inconsistently across CAD domains. We present VisCAD, a foundation model suite designed to provide both broad generalization and strong CAD capability for realistic industrial products. At its core is VisCAD-M1, a 27B model trained through mid-training and post-training for part-level design generation. On PubCADBench and RealCADBench, VisCAD-M1 achieves the highest average part-level score among the evaluated models, reaching 0.5540 compared with 0.5496 for the strongest frontier model. Reusing VisCAD-M1 as a test-time verifier can further raise the score to 0.5797, an approximately 5 percent relative improvement over the previous state of the art. VisCAD also includes a domain-specific harness that leverages frontier models for complex assembly generation and demonstrates advantages over general-purpose harnesses in both quantitative and qualitative evaluations.
Chinese Translation
面向工业产品的AI辅助计算机辅助设计(CAD)涉及两个具有挑战性的阶段。零件级生成将多种形式的用户意图,包括渲染图、文本描述、二维图纸和真实照片,映射为CAD领域特定语言的可执行程序。装配级生成还必须处理相互作用的零件、规划配合关系、估计位姿并正确放置所有零件。现有的专用CAD模型通常仅在狭窄的输入域(如渲染图或文本)上训练,泛化能力往往较差;而通用前沿模型虽然覆盖更广泛的输入,但在CAD各领域的表现不稳定。我们提出VisCAD,一个旨在为真实工业产品提供广泛泛化能力和强大CAD能力的基础模型套件。其核心是VisCAD-M1,一个通过中期训练和后期训练实现零件级设计生成的270亿参数模型。在PubCADBench和RealCADBench上,VisCAD-M1在所有评估模型中取得了最高的平均零件级得分,达到0.5540,而最强的前沿模型为0.5496。将VisCAD-M1复用为测试时验证器可进一步将得分提升至0.5797,相较此前最先进水平实现了约5%的相对提升。VisCAD还包含一个领域专用的执行框架(harness),利用前沿模型进行复杂装配生成,并在定量和定性评估中均展现出优于通用框架的优势。
cs.CV / 70 / 2609.03813

SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolution

SPARK:面向冻结DiT超分辨率模型的输入条件化稀疏激活调制方法
Putamorsi, Federico, Zini, Leonardo, Cornia, Marcella, Baraldi, Lorenzo
Abstract
Real-world image super-resolution (SR) increasingly relies on Diffusion Transformer (DiT) backbones, whose internal activations can be dominated by a small number of massive channels. Yet improving perceptual quality in these models still typically requires fine-tuning the network or attaching additional adapters, leaving this structured activation space largely unexplored for adaptation. We investigate whether dominant channels can instead serve as a compact adaptation interface for frozen DiT-based SR models. We first characterize their behavior in pretrained SR backbones and show through controlled interventions that they strongly affect reconstruction quality. Building on this observation, we introduce SPARK, a lightweight input-conditioned controller that predicts bounded per-channel affine transformations for only the selected channels, while keeping the SR backbone and VAE frozen. Dominant channels are identified through an online activation-ranking procedure, and only a small predictor conditioned on the low-resolution VAE latent is optimized. Experiments on three DiT-based SR backbones across DIV2K, RealSR, and DRealSR show consistent gains in both fidelity and perceptual quality while modulating only eight channels per stream and block. Controlled comparisons further show that these gains cannot be explained by parameter budget or access to the selected channels alone.
Chinese Translation
真实世界图像超分辨率(SR)越来越依赖于扩散Transformer(DiT)骨干网络,其内部激活可能被少量幅值极大的通道所主导。然而,提升这类模型的感知质量通常仍需要微调整个网络或附加额外的适配器,这一结构化的激活空间在适配方面很大程度上尚未被探索。我们研究了这些主导通道能否作为冻结DiT超分辨率模型的一种紧凑适配接口。我们首先刻画了它们在预训练SR骨干网络中的行为,并通过受控干预实验证明其对重建质量有显著影响。基于这一观察,我们提出了SPARK——一个轻量级的输入条件化控制器,它仅为所选通道预测有界的逐通道仿射变换,同时保持SR骨干网络和VAE完全冻结。主导通道通过在线激活排序流程识别,且仅优化一个以低分辨率VAE潜变量为条件的小型预测器。在三个基于DiT的SR骨干网络上,于DIV2K、RealSR和DRealSR数据集上的实验表明,尽管每个流和每个模块仅调制八个通道,该方法在保真度和感知质量上均取得一致提升。受控对比实验进一步表明,这些增益无法仅用参数预算或对所选通道的访问权限来解释。
cs.CV / 71 / 2609.03820

Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs

选择、压缩与再投资:长视频多模态大语言模型中视觉令牌分配的受控研究
Khatri, Prakhar
Abstract
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spatial compression, and reinvestment of the savings, across six training-free selection rules, three long-video benchmarks, and two answering models. Selection is the largest single lever: on LongVideoBench's hour-long bin, eight query-selected frames beat sixteen uniformly spaced ones by 6.9 points, and Orthogonal Matching Pursuit, an unmodified decades-old sparse-approximation algorithm, matches or comes within a point of every purpose-built selector we compare it against, across all three benchmarks. Compression is close to free: halving each frame's spatial budget at fixed timestamps costs at most 0.44 points. Reinvestment is where that budget turns back into accuracy: spending the freed tokens on twice as many compressed frames, at a measured cost no higher than the original eight, returns a further two to three points; compression only pays off once its savings are spent this way. Along the way, an implementation bug in our own AKS baseline and a 0.07 to 3.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.
Chinese Translation
长视频语言模型无法查看每一帧:一段以每秒一帧采样的视频一小时就有3,600张图像,而系统只能保留其中很小且固定的部分。哪些帧能进入这一保留区间通常被视为预处理细节;我们检验这一假设是否成立。已发表的选择器使对比变得困难,因为它们同时改变了帧评分器、提示边界、分辨率策略和回答模型。我们逐一固定这些因素,每次只改变一个决策:选择、空间压缩以及将节省的令牌进行再投资,在六种免训练选择规则、三个长视频基准和两种回答模型上进行实验。选择是最大的单一杠杆:在LongVideoBench的一小时时长区间上,8帧查询选择帧比16帧均匀间隔帧高出6.9分,而正交匹配追踪(Orthogonal Matching Pursuit)——一种未经修改的几十年前的稀疏逼近算法——在所有三个基准上与我们对比的每一个专门设计的选择器持平或相差不到一分。压缩几乎是无代价的:在固定时间戳下将每帧的空间预算减半最多损失0.44分。再投资则是将该预算转化为准确率的关键:将节省下来的令牌用于双倍数量的压缩帧,且实测成本不高于原始的8帧,可再获得两到三分的提升;压缩只有在其节省以这种方式被花费时才有所回报。此外,我们自己的AKS基线中的一个实现缺陷,以及两个评测框架在相同预算下运行相同已发表规则所产生的0.07至3.74分的差距,表明这类比较需要在单一受控评测框架内进行,而非跨论文比较。
cs.CV / 72 / 2609.03824

VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues

VI3:基于惯性线索的预训练3D基础模型锚定方法
Lozano, Ernesto, Jaenal, Alberto, Civera, Javier
Abstract
3D foundation models (3DFMs) excel at predicting camera poses and dense depth from multiple views of a scene, showcasing strong zero-shot generalization. However, as metric scale is not observable from monocular images, their absolute scale predictions are typically inaccurate. Inertial measurement units (IMUs), present in most devices, naturally complement monocular cameras by observing scaled motion. We introduce VI3, a model-agnostic framework that metrically anchors a pretrained 3DFM using only IMU readings. VI3 initializes and preintegrates the IMU to obtain a metric motion reference, which is then used to recover the scale of the 3DFM outputs. Our method includes adaptable anchoring strategies tailored to diverse 3DFM architectures. Experiments on synthetic and real aerial datasets demonstrate that VI3 recovers metric scale without ground-truth supervision while preserving geometric consistency, acting as a fine refinement under well-conditioned motion and as a strong prior when motion is less informative.
Chinese Translation
3D基础模型(3DFM)擅长从场景的多个视角预测相机位姿和稠密深度,展现出强大的零样本泛化能力。然而,由于单目图像无法观测到绝对尺度,其尺度预测通常不够准确。惯性测量单元(IMU)存在于大多数设备中,通过观测带尺度的运动自然地与单目相机形成互补。我们提出了VI3,一个与模型无关的框架,仅利用IMU读数对预训练的3DFM进行度量尺度锚定。VI3对IMU进行初始化和预积分以获得带尺度的运动参考,进而用于恢复3DFM输出的尺度。我们的方法包含针对不同3DFM架构定制的自适应锚定策略。在合成和真实航拍数据集上的实验表明,VI3能够在无需真值监督的情况下恢复度量尺度,同时保持几何一致性:在运动信息充分的条件下充当精细的优化手段,而在运动信息不足时则充当强先验。
cs.CV / 73 / 2609.03829

The impact of phase information for few-shot fine-grained image classification

相位信息对小样本细粒度图像分类的影响
Liu, Ruiling, Zhang, Linyue, Zeng, Wenyi, Lu, Jiamiao, Zhang, Weichuang, Sun, Changming, Zhang, Zejun, Zhao, Xiao
Abstract
Few-shot fine-grained image classification (FSFGIC) aims to classify similar images with limited labeled examples. This work highlights the critical yet underutilized role of phase information in capturing structural relationships within an image. This study introduces a novel plug-and-play amplitude-phase integration (API) module that effectively combines local and global frequency amplitude and phase information for obtaining more comprehensive feature descriptors. Additionally, a dedicated network, named PSF-Net, is proposed that adaptively fuses phase-based spatial and frequency information for FSFGIS. The designed PSF-Net can be easily integrated into standard episodic training architectures for end-to-end training from scratch. Extensive experiments on five public datasets demonstrate that the method outperforms existing state-of-the-art benchmarks.
Chinese Translation
小样本细粒度图像分类(Few-shot Fine-grained Image Classification, FSFGIC)旨在利用有限的有标注样本对相似图像进行分类。这项工作强调了相位信息在捕捉图像内部结构关系方面至关重要却未被充分利用的作用。本研究提出了一种新颖的即插即用幅相融合(Amplitude-Phase Integration, API)模块,该模块有效地结合局部和全局频率幅值与相位信息,以获得更全面的特征描述符。此外,本文还提出了一种专用网络 PSF-Net,该网络能够自适应地融合基于相位的空间信息和频率信息,用于小样本细粒度图像分类任务。所设计的 PSF-Net 可以轻松集成到标准的事件式(episodic)训练架构中,实现从零开始的端到端训练。在五个公开数据集上的大量实验表明,本方法优于现有最先进的基准方法。
cs.CV / 74 / 2609.03892

GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs

GraFT:一种基于3D场景图的免训练多模态大语言模型空间推理框架
Du, Junqing, Ropero, Fernando, Turkoz, Erkin, Zhang, Yanfeng, Liu, Lu
Abstract
3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in current multimodal large language models (MLLMs). These models falter at precise geometric measurement, at transforming between egocentric and allocentric viewpoints, and at grounding fine-grained appearance. The most common remedies fine-tune the model on large-scale curated spatial-reasoning datasets or attach dedicated encoders for 3D geometry, which typically couples the solution to costly supervision and a specific backbone. We instead introduce GraFT, a training-free framework that supplies the missing 3D structure through a compact, easily maintained 3D scene graph (3DSG). From this 3DSG, GraFT provides three spatial reasoning capabilities: (1) deterministic geometry through symbolic tools, (2) allocentric layout through a bird's-eye-view (BEV) rendering, and (3) visual-attribute grounding through task-relevant egocentric frames. On ScanQA, GraFT improves every metric over the same-backbone baseline, raising CIDEr by 27%. On VSI-Bench, GraFT improves frozen MLLMs by up to 65%, surpassing every proprietary and general-purpose open-source baseline, and several prominent fine-tuned spatial models.
Chinese Translation
3D空间推理是理解和作用于物理世界的基础,但在当前的多模态大语言模型(MLLMs)中仍不可靠。这些模型在精确几何测量、以自我为中心(egocentric)和以对象为中心(allocentric)视角之间的转换,以及细粒度外观的定位方面表现不佳。最常见的补救方法是在大规模整理的空间推理数据集上对模型进行微调,或为3D几何配置专用编码器,这通常将解决方案与昂贵的监督和特定的骨干网络绑定在一起。我们转而提出GraFT,这是一个免训练框架,通过一个紧凑且易于维护的3D场景图(3DSG)提供缺失的3D结构。基于该3DSG,GraFT提供三种空间推理能力:(1)通过符号工具实现确定性几何计算,(2)通过鸟瞰图(BEV)渲染实现以对象为中心的布局理解,(3)通过任务相关的自我中心帧实现视觉属性定位。在ScanQA数据集上,GraFT在相同骨干网络的基线上提升了所有指标,CIDEr提高了27%。在VSI-Bench数据集上,GraFT使冻结的MLLMs性能提升高达65%,超越了所有专有模型和通用开源基线,以及多个著名的经过微调的空间推理模型。
cs.CV / 75 / 2609.03895

Concept of a Sensor Test Environment for Dusty Agricultural Conditions

面向多尘农业环境的传感器测试场概念
Buckel, Peter, Hermann, Johannes, Wollmann, Jonas, Dietmueller, Thomas, Oksanen, Timo
Abstract
Dust in agriculture presents a significant challenge for autonomous agricultural machinery. Dust can impair the performance of sensors and algorithms. This work, therefore, presents a concept for a proving ground consisting of an indoor and outdoor area. The indoor area comprises a laboratory test bench where dust circulates in a closed system and a test hall where life-size objects can be placed. The outdoor area features dedicated test setups that enable reproducible data to be recorded with and without dust during real-world agriculture work. The proving ground and the setups are visualized in 3D.
Chinese Translation
农业中的粉尘对自主农业机械构成了重大挑战。粉尘会损害传感器和算法的性能。因此,本工作提出了一种由室内和室外区域组成的试验场概念。室内区域包括一个实验室测试台(粉尘在封闭系统中循环)和一个测试大厅(可放置实物大小的物体)。室外区域设有专门的测试装置,能够在真实农业作业过程中记录有粉尘和无粉尘条件下的可复现数据。该试验场及各测试装置均以三维(3D)形式进行了可视化展示。
cs.CV / 76 / 2609.03919

OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mapping

OctWorld:基于八叉树三维映射的长程世界一致性视频生成
Lv, Zelong, Xu, Sicheng, Xiang, Jianfeng, Wang, Ruicheng, Dong, Yue, Deng, Yu, Sun, Guangzhong, Yang, Jiaolong
Abstract
We present OctWorld, a video diffusion framework with persistent 3D memory for generating explorable, world-consistent, and high-fidelity visual scenes. Given a single image, OctWorld performs stable autoregressive world generation along user-specified camera trajectories. We focus on long-range generation, characterized by extended camera paths and wide viewpoint coverage, where preserving spatial consistency is particularly challenging when previously generated regions are revisited. To address this problem, we introduce OctMap, an extensible and spatially adaptive 3D memory that progressively fuses generated visual observations and their corresponding depth maps into a global representation. OctMap employs TSDF fusion within a dynamic sparse octree whose spatial resolution adapts to image evidence. This design preserves geometric and appearance details across diverse scene scales while maintaining low memory overhead. Experiments demonstrate that OctWorld generates long-range, spatially consistent videos and outperforms prior methods on both existing benchmarks and challenging long-range generation settings. OctMap also provides clear advantages over point-based caches and fixed-resolution TSDF volumes. Project page: https://maxtirerror.github.io/octworldpage/
Chinese Translation
我们提出了 OctWorld,一个具有持久三维记忆的视频扩散框架,用于生成可探索、世界一致且高保真的视觉场景。给定单张图像,OctWorld 能够沿用户指定的相机轨迹进行稳定的自回归世界生成。我们聚焦于长程生成任务,其特点是相机路径长、视点覆盖范围广,当重新访问先前生成的区域时,保持空间一致性尤为困难。为解决这一问题,我们引入了 OctMap,一种可扩展且空间自适应的三维记忆结构,它将生成的视觉观测及其对应的深度图逐步融合为全局表示。OctMap 在动态稀疏八叉树中采用 TSDF 融合,其空间分辨率根据图像证据自适应调整。这一设计在保持较低内存开销的同时,能够在不同场景尺度下保留几何与外观细节。实验表明,OctWorld 能够生成长程且空间一致的视频,在现有基准和具有挑战性的长程生成设置上均优于先前方法。与基于点的缓存以及固定分辨率的 TSDF 体相比,OctMap 也具有明显优势。项目页面:https://maxtirerror.github.io/octworldpage/
cs.CV / 77 / 2609.03931

Sparse auto-regressive modeling for scene generation from multi-view images

基于稀疏自回归建模的多视角图像场景生成方法
Lucas, Thomas, Pietrantoni, Maxime, Weinzaepfel, Philippe, Cho, Wonjune, Duisterhof, Bardienus Pieter, Leroy, Vincent, Revaud, Jerome
Abstract
Generating complete 3D scenes from sparse, unconstrained views is a fundamental challenge in 3D vision which requires reasoning beyond observed content while remaining computationally tractable. Existing feed-forward reconstruction methods are inherently limited to content visible in the input images, while 3D generative modeling is hindered by the high computational cost of dense volumetric representations and the scarcity of large-scale 3D supervision. We introduce SPAR3S, a sparse voxel-aligned 3D latent generative model for conditional scene completion without requiring ground-truth 3D data for supervision. Our key insight is to formulate 3D scene generation in a structured, compact, voxel-aligned 3D latent space where only occupied voxels are represented. We learn this sparse latent space directly from multi-view images using photometric supervision via differentiable 3D Gaussian Splatting. Given a partial set of observed voxels encoded from sparse input views, scene completion reduces to predicting the missing latent tokens and their spatial support within the voxel grid. To this end, we train a masked autoregressive transformer that jointly models voxel occupancy and latent token values, enabling efficient and spatially consistent generation of unseen regions. We demonstrate the effectiveness of our method on synthetic indoor scenes, achieving higher novel-view quality than prior work. We further validate its generalization on RealEstate10k, highlighting its applicability to real-world data.
Chinese Translation
从稀疏且无约束的视角生成完整的3D场景是3D视觉领域的一项基础性挑战,它要求在保持计算可行性的同时,对观测内容之外进行推理。现有的前馈式重建方法本质上受限于输入图像中可见的内容,而3D生成建模则受到稠密体积表示的高计算成本以及大规模3D监督数据稀缺的制约。我们提出了SPAR3S,一个稀疏体素对齐的3D潜在生成模型,用于条件场景补全,且无需真实3D数据进行监督。我们的核心见解是,在结构化、紧凑且体素对齐的3D潜在空间中构建3D场景生成,其中仅表示被占据的体素。我们借助可微3D高斯泼溅(3D Gaussian Splatting),利用光度监督直接从多视角图像中学习该稀疏潜在空间。给定从稀疏输入视角编码得到的部分观测体素集合,场景补全即可简化为预测缺失的潜在token及其在体素网格中的空间支撑。为此,我们训练了一个掩码自回归Transformer,联合建模体素占据状态与潜在token值,从而实现对未见区域的高效且空间一致的生成。我们在合成室内场景上验证了该方法的有效性,在新视角渲染质量上超越了先前工作。我们还在RealEstate10k数据集上进一步验证了其泛化能力,凸显了其对真实世界数据的适用性。
cs.CV / 78 / 2609.03952

WorldReward: Reward Modeling for Camera-Conditioned World Models

WorldReward:面向相机条件化世界模型的奖励建模
Wang, Yibin, Wang, Zehan, Tang, Junshu, Li, Zhimin, Zhou, Yujie, Bu, Jiazi, Ling, Pengyang, Han, Feng, Zhang, Zhixiong, Xing, Long, Ding, Shengyuan, Li, Ziang, Jin, Cheng, Zang, Yuhang, Wang, Jiaqi, Pang, Tianyu
Abstract
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements separately: geometry-based rewards estimate trajectory execution but cannot judge the visual quality of the executed motion, whereas image-based rewards measure frame quality without capturing action execution or temporal dynamics. We posit that a vision-language model (VLM) offers a shared reasoning space for relating actions to their visual outcomes. However, judging a complete long video against its full action sequence creates a lengthy, noisy context in which short-lived local action evidence can be missed or diluted. We present WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models. WorldReward decomposes paired videos into action-aligned chunks, organizes each chunk into structured visual evidence, and aggregates chunk-level decisions by voting into separate video-level action and visual-quality preferences. To train it, we construct a large-scale reasoning-augmented preference dataset using structured judgments generated by a frontier VLM and refined through tool-based agent auditing and targeted human review. We further introduce WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency, appearance quality, and motion quality. WorldReward achieves the highest agreement on all three dimensions, exceeding GPT-5.5 by 3.42, 1.45, and 3.56 percentage points, respectively. When used for RL post-training of HY-WorldPlay 1.5, it consistently improves both action execution and visual quality across short- to long-term horizons.
Chinese Translation
相机条件化世界模型生成可交互视频,其中指令动作应引发预期的场景变化,同时保持外观、几何和时间动态的一致性。现有奖励方法对这些要求分别评估:基于几何的奖励可估计轨迹执行情况,但无法评判所执行动作的视觉质量;而基于图像的奖励虽可度量帧质量,却无法捕捉动作执行与时间动态。我们提出,视觉语言模型(VLM)为关联动作及其视觉结果提供了共享的推理空间。然而,将一段完整的长视频与其完整的动作序列进行对照评判会产生冗长且充满噪声的上下文,短暂存在的局部动作证据容易被遗漏或稀释。我们提出 WorldReward,一种基于 VLM 的成对偏好奖励模型,为相机条件化世界模型统一了动作一致性与视觉质量评估。WorldReward 将成对视频分解为动作对齐的分块,将每个分块组织为结构化的视觉证据,并通过投票机制聚合分块级决策,形成独立的视频级动作偏好与视觉质量偏好。为训练该模型,我们构建了一个大规模的推理增强偏好数据集:首先由前沿 VLM 生成结构化评判,再通过基于工具的智能体审计和针对性的人工审核加以精炼。我们进一步引入 WorldReward-Bench,一个人工标注的基准,用于衡量奖励模型与人类偏好在动作一致性、外观质量和运动质量三个维度上的一致性。WorldReward 在全部三个维度上均取得最高一致性,分别超过 GPT-5.5 达 3.42、1.45 和 3.56 个百分点。在用于 HY-WorldPlay 1.5 的强化学习后训练时,它在从短期到长期的时间跨度上均能持续提升动作执行与视觉质量。
cs.CV / 79 / 2609.03956

RARF: Region-Aware Rectified Flows for 3D Brain MRI Inpainting

RARF:面向3D脑部MRI修复的区域感知整流流模型
Guija-Valiente, Tomas, Rodriguez-Gonzalez, Blanca, Malpica, Norberto, Torrado-Carvajal, Angel
Abstract
Medical image inpainting has the potential to improve automated brain MRI analysis by reconstructing healthy tissue within pathological regions. We introduce RARF, a task-agnostic region-aware rectified flow framework for masked data generation. We instantiate the framework for 3D brain MRI inpainting as our submission to the BraTS Inpainting Challenge 2026. RARF restricts the stochastic interpolation process to the inpainting region, while the observed voxels remain fixed and provide patient-specific anatomical context. A three-dimensional neural network receives the partially voided image, with Gaussian noise filling the missing region, together with the inpainting mask and the corresponding timestep. The model is trained using masked flow-matching and reconstruction-consistency objectives, combined with mask-aware preprocessing and data augmentation. During inference, the learned velocity field transports the initial noise toward a plausible reconstruction of the missing tissue, which is then combined with the unchanged observed anatomy. Experiments under the BraTS evaluation protocol show that the proposed approach produces competitive reconstructions while maintaining anatomical consistency. Source code is available at: https://github.com/TomasGuija/rarf.
Chinese Translation
医学图像修复通过在病理区域重建健康组织,有潜力提升自动化脑部MRI分析的效果。我们提出了RARF,这是一个任务无关的区域感知整流流(rectified flow)框架,用于掩码数据生成。作为对2026年BraTS修复挑战赛(BraTS Inpainting Challenge)的参赛工作,我们将该框架实例化应用于3D脑部MRI修复。RARF将随机插值过程限制在修复区域内,同时保持已观测体素不变,为模型提供患者特异性的解剖学上下文。一个三维神经网络接收部分空洞的图像(缺失区域由高斯噪声填充)、修复掩码以及相应的时间步。模型使用掩码流匹配(masked flow-matching)目标和重建一致性目标进行训练,并结合掩码感知的预处理和数据增强。在推理阶段,学习到的速度场将初始噪声输运至缺失组织的合理重建结果,随后与保持不变的已观测解剖结构进行合并。在BraTS评估协议下的实验表明,所提出的方法能够生成具有竞争力的重建结果,同时保持解剖学一致性。源代码见:https://github.com/TomasGuija/rarf。
cs.CV / 80 / 2609.03981

Sharpening the Ensemble: An SSIM-Aligned Residual Refiner for Brain-MRI Inpainting Post-Processing

锐化集成模型:一种面向脑部MRI修复后处理、与SSIM对齐的残差精炼器
Kömürcü, Kubilay Kağan, Öksüz, İlkay
Abstract
Brain-MRI inpainting replaces a masked region of a scan with synthesized, anatomically plausible healthy tissue, so that analysis tools built for healthy brains can be applied to images they would otherwise reject. On the BraTS local-synthesis benchmark, which ranks submissions on the structural similarity index (SSIM), the peak signal-to-noise ratio, and the mean squared error (MSE) jointly, the strongest recent models are accurate, but several report blurry synthesized regions and attribute this to the mean-seeking behavior of the $\ell_1$ and MSE terms in their training losses. We address this in post-processing, forming a deep ensemble of the two co-first-place 2025 models and training a lightweight residual refiner on the ensemble's own outputs under an $\ell_1$ loss augmented with a structural-similarity term whose weight $\lambda$ we vary. At a moderate $\lambda$ the refiner improves SSIM over the ensemble, from $0.8767$ to $0.8780$ on a held-out reproduction of the official scorer and from $0.8555$ to $0.8572$ on the official validation leaderboard, with essentially no change in MSE. The gain is small but consistent, improving $62.6\%$ of the held-out cases with a signed-rank $p=2.2\times10^{-7}$, whereas over-weighting the structural term reverses it. Two ablations bound the effect. Adding any third model to the two-model ensemble degrades it, and classical unsharp masking fails to improve SSIM at any strength (best $0.8765$ against $0.8767$), so the gain reflects learned rather than indiscriminate sharpening. The result is a cheap, reproducible post-processing stage that improves an already strong ensemble without any large-scale retraining.
Chinese Translation
脑部MRI修复是指用合成的、在解剖学上合理健康组织替换扫描中被遮罩的区域,从而使为健康大脑设计的分析工具能够应用于原本会被拒绝的图像。在BraTS局部合成基准上,该基准综合考虑结构相似性指数(SSIM)、峰值信噪比和均方误差(MSE)对提交结果进行排名,目前最强的近期模型表现准确,但多个模型报告合成区域存在模糊现象,并将其归因于训练损失中$\ell_1$和MSE项的均值寻求行为。我们在后处理阶段解决这一问题:将2025年两个并列第一的模型组成深度集成,并在集成模型自身的输出上训练一个轻量级残差精炼器,其损失函数为$\ell_1$损失并辅以权重为$\lambda$的结构相似性项。在适中的$\lambda$下,精炼器将集成模型的SSIM从官方评分器保留复现集上的$0.8767$提升至$0.8780$,在官方验证排行榜上从$0.8555$提升至$0.8572$,而MSE基本保持不变。该提升虽然幅度不大,但具有一致性,在$62.6\%$的保留测试案例上有所改进,符号秩检验$p=2.2\times10^{-7}$;而过度加大结构项权重则会使效果反转。两项消融实验界定了该效果的边界:向两模型集成中添加任何第三个模型都会使其性能下降,而经典非锐化掩模在任何强度下都无法提升SSIM(最佳为$0.8765$,低于$0.8767$),因此该提升反映的是学习到的锐化而非无差别锐化。最终结果是一个低成本、可复现的后处理阶段,无需任何大规模重新训练即可改进本已强大的集成模型。
cs.CV / 81 / 2609.03985

IchthyoNoma: Nomenclature and Context Sensitivity of Zero-Shot Biological Vision--Language Models for Bangladeshi Freshwater Fish Recognition

IchthyoNoma:面向孟加拉国淡水鱼识别的零样本生物视觉-语言模型的命名法与上下文敏感性研究
Nazim-E-Alam, Rahman, Tarek, Morol, Md Kishor
Abstract
Zero-shot vision-language models (VLMs) are increasingly used as training-free species recognizers, but reported accuracy can reflect more than visual species knowledge. We audit CLIP, BioCLIP, BioCLIP2, and a multilingual Jina CLIP v2 control on seven freshwater-fish categories from two Bangladeshi sources (10,321 images). BioCLIP2 reaches 72.36% on BFF-15 with English common names and 68.91% on SylFishBD with scientific names, versus 25.15% and 14.40% for generic CLIP. BioCLIP2 Bengali prompts are near chance in balanced accuracy (14.22-14.29%); Jina partially recovers Bengali discrimination to 21.89% and 16.36%, but bare Bengali names return to 14.29% on both sources. Paired SylFishBD interventions show no significant weak-blur effect, modest losses from stronger blur/gray masking, a larger white-mask artifact, and strong species dependence. Zero-shot biological VLM scores therefore jointly reflect biological specialization, multilingual alignment, nomenclature, prompt formulation, and context.
Chinese Translation
零样本视觉-语言模型(VLM)正日益被用作无需训练的物种识别器,但其所报告的准确率可能不仅反映视觉层面的物种知识。我们在来自孟加拉国两个数据源(10,321张图像)的七个淡水鱼类别上,对CLIP、BioCLIP、BioCLIP2以及多语言Jina CLIP v2对照模型进行了审计。在BFF-15数据集上,BioCLIP2使用英文俗名达到72.36%的准确率,在SylFishBD数据集上使用学名达到68.91%,而通用CLIP分别仅为25.15%和14.40%。BioCLIP2使用孟加拉语提示时,平衡准确率接近随机水平(14.22%–14.29%);Jina将孟加拉语判别能力部分恢复至21.89%和16.36%,但使用裸孟加拉语名称时在两个数据源上均回落至14.29%。针对SylFishBD的配对干预实验表明,轻度模糊无显著影响,较强模糊/灰度遮蔽导致适度性能下降,白色遮蔽产生较大伪影影响,且存在较强的物种依赖性。因此,零样本生物VLM的得分共同反映了生物领域专业化、多语言对齐、命名法、提示形式以及上下文等因素。
cs.CV / 82 / 2609.03995

Catalogue Photography as a Cold Start: Toward Deployable Carbide Burr Recognition

产品目录摄影作为冷启动:迈向可部署的硬质合金旋转锉识别
Madavath, Abilash Philip, Aubeeluck, Chandra Yuvesh, Raju, Augustin, Pyschny, Nicolas, Hackelöer, Felix, Zwanzig, Florian
Abstract
Verifying that manufactured batches of milling tools or carbide rotary burrs conform to production order sheets remains a largely manual and error-prone quality assurance task. Automating this process with computer vision faces a critical cold-start constraint since no labelled imagery is available, leaving manufacturer catalogue photography as the sole source of supervision. We investigate how far catalogue supervision can support an industrial recognition pipeline under domain shift, explicitly measuring the gap between catalogue separability and performance on held-out field photographs. Our findings reveal three key insights. First, off-the-shelf frozen feature extractors do not reliably separate the two task attributes, head shape and tooth profile, motivating targeted representation learning. Second, metric learning produces near-perfect unsupervised cluster discovery on catalogue images (adjusted Rand index 0.94--0.97), but less than half of this gain transfers to field photographs. Third, the largest transfer gains do not come from model scale or representation complexity, but from simple changes that reduce domain sensitivity: converting images to grayscale (+0.22) and constraining retrieval using the known order sheet via Hungarian assignment (+0.11). We therefore treat catalogue photography as a useful cold start rather than a deployment-ready training domain, and provide empirical baselines and an evaluation protocol for catalogue-to-field transfer in precision tool manufacturing.
Chinese Translation
验证制造的铣刀或硬质合金旋转锉批次是否符合生产订单单据,仍然是一项主要依赖人工且容易出错的质量保证任务。由于没有任何带标注的图像可用,利用计算机视觉自动化该流程面临严峻的冷启动约束,这使得制造商的产品目录摄影成为唯一的监督来源。我们研究了目录监督在领域偏移条件下能在多大程度上支持工业识别流程,并明确测量了目录数据可分性与留出实地照片上的性能之间的差距。我们的发现揭示了三个关键洞察。第一,现成的冻结特征提取器无法可靠地区分头型与齿形这两个任务属性,这促使我们进行有针对性的表示学习。第二,度量学习在目录图像上实现了近乎完美的无监督聚类发现(调整兰德指数 0.94–0.97),但这一收益不到一半能够迁移到实地照片上。第三,最大的迁移收益并非来自模型规模或表示复杂度,而是来自降低领域敏感性的简单改动:将图像转换为灰度图(+0.22),以及通过匈牙利算法利用已知订单单据约束检索(+0.11)。因此,我们将目录摄影视为一种有用的冷启动手段,而非可直接部署的训练领域,并为精密刀具制造中的目录到实地迁移提供了实证基线和评估协议。
cs.CV / 83 / 2609.04009

The Blind Spot in 2D Infants' Pose Estimation:Robust Learning from Noisy Annotations

二维婴儿姿态估计中的盲点:基于噪声标注的鲁棒学习
Cardinale, Emanuele, Proietti, Marco, Cacciatore, Alessandro, Spadea, Maria Francesca, Migliorelli, Lucia, Moccia, Sara
Abstract
Noisy annotations pose a significant challenge for supervised deep learning, as neural networks rely on large-scale, high-quality labeled data whose corruption can severely impair model performance. Although robustness to label noise has been extensively studied for classification tasks, it remains relatively underexplored in Pose Estimation (PE). This limitation becomes critical in clinical contexts, including neonatology, where PE of preterm infants is used to support the assessment of spontaneous motility, a key indicator of neurodevelopmental trajectories. In such settings, infants' images labeling is further hindered by visual challenges (e.g., keypoint self-occlusions, caregiver interference), making the annotation process inherently susceptible to errors. To tackle noisy annotations in PE, we introduce REliable keypoint selection via Memory of traINing Dynamics (REMIND), a clustering-based keypoint-selection strategy that exploits keypoint-wise training dynamics to identify noisy labels without assuming any prior knowledge of the noise distribution, thus enabling noise-free model training. When evaluated on the proprietary NeoPose dataset, comprising 46 videos of 46 preterm infants recorded in real clinical settings, REMIND correctly identifies noisy annotations across multiple corruption scenarios, achieving up to 93\% Area Under the Curve (AUC) with three different PE architectures used in the relevant literature. To our knowledge, this is the first study to explicitly address label noise in preterm infants' PE, paving the way for the design of trustworthy learning-based algorithms for infants'monitoring support when data quality cannot be guaranteed.
Chinese Translation
噪声标注对监督深度学习构成了重大挑战,因为神经网络依赖于大规模、高质量的标注数据,而这些数据一旦受损会严重损害模型性能。尽管针对分类任务的标签噪声鲁棒性已被广泛研究,但在姿态估计(Pose Estimation, PE)领域仍相对缺乏探索。这一局限在临床场景中尤为关键,例如新生儿学中,早产儿的姿态估计被用于支持自发运动评估——这是神经发育轨迹的一项关键指标。在此类场景中,婴儿图像的标注还受到视觉难题(如关键点自遮挡、照护者干扰)的阻碍,使得标注过程天然容易出错。为解决姿态估计中的噪声标注问题,我们提出了基于训练动态记忆的可靠关键点选择方法(REliable keypoint selection via Memory of traINing Dynamics, REMIND),这是一种基于聚类的关键点选择策略,利用关键点级别的训练动态来识别噪声标签,而无需对噪声分布做任何先验假设,从而实现无噪声的模型训练。在专有的NeoPose数据集上进行评估——该数据集包含在真实临床环境中录制的46名早产儿的46段视频——REMIND在多种损坏场景下均能正确识别噪声标注,在相关文献中使用的三种不同姿态估计架构上,曲线下面积(AUC)最高达到93%。据我们所知,这是首项明确针对早产儿姿态估计中标签噪声问题的研究,为在数据质量无法保证的情况下设计可信赖的、基于学习的婴儿监测辅助算法铺平了道路。
cs.CV / 84 / 2609.04026

Stable and Scalable Bundle Adjustment of Holistic 3D Structures

整体三维结构的稳定且可扩展的光束法平差
Liu, Shaohui, Pautrat, Rémi, Barath, Daniel, Hartley, Richard, Larsson, Viktor, Pollefeys, Marc
Abstract
Bundle Adjustment (BA) is a cornerstone of 3D computer vision and has benefited from decades of advances in sparse optimization and numerical methods. It was originally developed for jointly optimizing camera intrinsics, poses and sparse 3D points. While extensions incorporate lines and other primitives, integrating richer geometric structures such as parallelism, coplanarity, or wireframes often introduces significantly increased computational cost and reduced numerical stability. In this paper, we propose a unified framework that extends bundle adjustment to jointly optimize geometric features and higher-order relations. We first introduce a taxonomy that distinguishes scalable geometric features with direct 2D measurements (e.g., points and lines), from groups encoding higher-order relations (e.g., coplanarity, parallelism, etc.), where we show that groups can be modeled as camera-like entities within the bundle adjustment framework. Building on this formulation, we propose that both group constraints and cross-feature relations (i.e., point-line associations) can be expressed through 2D reprojection measurements. By formulating group-induced and cross-feature reprojection errors, we preserve the sparsity structure of classical point-based BA under Schur elimination, while avoiding direct 3D regularization that degrades the conditioning and stability. Experiments on both real-world and synthetic datasets demonstrate runtime performance comparable to classical point-only bundle adjustment, while producing significantly richer 3D structures and improved geometric accuracy.
Chinese Translation
光束法平差(Bundle Adjustment, BA)是三维计算机视觉的基石,受益于数十年来稀疏优化与数值方法的进展。它最初被提出用于联合优化相机内参、位姿和稀疏三维点。尽管已有扩展方法引入了直线及其他几何图元,但整合更丰富的几何结构(如平行性、共面性或线框)往往会显著增加计算成本并降低数值稳定性。本文提出了一个统一框架,将光束法平差扩展至可联合优化几何特征与高阶关系。我们首先引入一种分类方法,区分具有直接二维观测的可扩展几何特征(如点和线),以及编码高阶关系的组(groups,如共面性、平行性等),并证明组可以在光束法平差框架中被建模为类似相机的实体。基于这一表述,我们提出组约束与跨特征关系(即点-线关联)均可通过二维重投影观测来表达。通过构造组诱导的重投影误差和跨特征重投影误差,我们在Schur消元下保留了经典基于点的BA的稀疏结构,同时避免了会恶化条件数与稳定性的直接三维正则化。在真实数据集和合成数据集上的实验表明,本方法的运行时间可与经典纯点BA相媲美,同时能生成显著更丰富的三维结构并提升几何精度。
cs.CV / 85 / 2609.04031

DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video Generation

DSAQuant:面向视频生成的去噪阶段对齐量化感知训练
Li, Shuaiting, Gao, Zelin, Shen, Haibin, Shen, Yujun, Qin, Haotong, Xu, Yinghao
Abstract
Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high memory and computational costs hinder practical deployment. Quantization-aware training (QAT) is an effective solution for compressing and accelerating advanced generative models without runtime overhead at inference. However, existing QAT methods suffer from a distinctive challenge in VDMs: while they often preserve prompt semantics, global layout, and coarse motion, the quantized model severely degrades visual details, texture fidelity, and sharpness. In this paper, we trace this degradation to the timestep-agnostic design of conventional quantization pipelines, which overlooks the stage-wise functionality of video denoising. In VDMs, early denoising steps mainly establish global structure and motion, whereas middle and late steps refine local appearance and high-frequency details. Based on this insight, we propose DSAQuant, a Denoising-Stage-Aligned Quantization-aware training framework for VDMs. During training, Denoising-Stage Oriented Supervision preserves teacher distillation in early steps for stable structure planning, while shifting later steps toward target-driven optimization to enhance detail reconstruction. During inference, Denoising-Stage Gated Guidance disables CFG in the final denoising steps to prevent it from amplifying quantization-induced errors into high-frequency artifacts. Extensive experiments on the Wan and CogVideoX families under W4A4 and W3A3 settings show that DSAQuant consistently outperforms the SOTA QAT baseline, improving the VBench average score by up to 6.60 under aggressive W3A3 quantization while preserving strong text-video alignment. These results demonstrate that effective VDM quantization requires not only reducing quantization error, but also aligning quantization training and inference with the stage-wise nature of video diffusion.
Chinese Translation
视频扩散模型(VDM)在文本到视频生成方面取得了令人瞩目的进展,但其高昂的内存和计算成本阻碍了实际部署。量化感知训练(QAT)是一种有效的解决方案,可以在不增加推理时运行开销的情况下压缩和加速先进的生成模型。然而,现有的QAT方法在VDM中面临一个独特的挑战:尽管它们通常能够保留提示语义、全局布局和粗略运动,但量化后的模型在视觉细节、纹理保真度和清晰度方面严重退化。在本文中,我们将这种退化归因于传统量化流水线中与时间步无关的设计,其忽视了视频去噪的分阶段功能。在VDM中,早期去噪步骤主要建立全局结构和运动,而中间和后期步骤则细化局部外观和高频细节。基于这一洞察,我们提出了DSAQuant,一个面向VDM的去噪阶段对齐量化感知训练框架。在训练阶段,面向去噪阶段的监督在早期步骤保留教师蒸馏以实现稳定的结构规划,同时在后期步骤转向目标驱动优化以增强细节重建。在推理阶段,去噪阶段门控引导在最后的去噪步骤中禁用CFG,以防止其将量化引入的误差放大为高频伪影。在Wan和CogVideoX系列模型上于W4A4和W3A3设置下的大量实验表明,DSAQuant持续优于最先进的QAT基线,在激进的W3A3量化下将VBench平均分数最高提升6.60,同时保持了较强的文本-视频对齐能力。这些结果表明,有效的VDM量化不仅需要降低量化误差,还需要使量化训练和推理与视频扩散的分阶段特性相一致。
cs.CV / 86 / 2609.04034

Editable Visual Design

可编辑视觉设计
Ye, Junyan, Liu, Wei, Jiang, Dongzhi, Wen, Zichen, Li, HaoDong, Lv, Zhutao, Lin, Jiaxin, Yu, Jinhua, He, Jun, Huang, Zilong, Chen, Rui, Li, Weijia
Abstract
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control and decoupled layers, yet remains constrained by a lack of global aesthetic intuition and the difficulty of coding complex visual assets. To address this, we propose Editable Visual Design, a new paradigm driven by a Coding Agent. We designate the VLM as the ``creative brain'' for requirement comprehension, task planning, and aesthetic judgment, while utilizing the image generation model as an on-demand ``visual world simulator'' to synthesize standalone visual assets. Operating under an ``imagine first, then act'' closed-loop workflow, the agent generates isolated assets, writes native HTML/CSS, and iteratively refines the design against visual rendering feedback. Furthermore, Agent Design Replay faithfully reproduces the creative and reasoning trajectory akin to that of professional human designers. Ultimately, the system delivers editable artifacts with decoupled layers and real text, enabling users to perform intuitive mouse dragging and layout adjustments on a graphical user interface. Validations on posters, infographics, and other scenarios show that this paradigm successfully achieves both refined aesthetics and production-grade editability.
Chinese Translation
尽管GPT-Image-2和Nano-Banana等扩散基础模型展现出卓越的视觉表现力,但其端到端生成方式本质上产生的是文本易出错的扁平化位图,无法进行分层后编辑。相反,基于代码的视觉生成(通过Coding Agent)能够提供精确的布局控制和解耦的图层,但仍受限于缺乏全局审美直觉以及编码复杂视觉资产的困难。为解决这一问题,我们提出可编辑视觉设计,一种由Coding Agent驱动的新范式。我们将VLM指定为负责需求理解、任务规划和审美判断的“创意大脑”,同时利用图像生成模型作为按需的“视觉世界模拟器”来合成独立的视觉资产。在“先想象、后行动”的闭环工作流下,该智能体生成孤立资产、编写原生HTML/CSS,并根据视觉渲染反馈迭代地优化设计。此外,Agent Design Replay能够忠实重现类似专业人类设计师的创作与推理轨迹。最终,该系统交付具有解耦图层和真实文本的可编辑产物,使用户能够在图形用户界面上进行直观的鼠标拖拽和布局调整。在海报、信息图等场景上的验证表明,该范式成功实现了精致美学与生产级可编辑性的兼得。
cs.CV / 87 / 2609.04070

Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving

从离散思维到连续动作:面向端到端自动驾驶的潜在空间对齐规划
Yao, Ruoyu, Xie, Yusen, Liu, Qingzhao, Liu, Pei, Yang, Zewei, Zhu, Yipeng, Wang, Xiaolong, Ma, Jun
Abstract
Bridging the gap between the discrete reasoning of Vision-Language Models and the continuous, physics-constrained nature of autonomous driving remains a significant challenge. In this work, we introduce LaPla, a unified Vision-Language-Action (VLA) framework featuring latent-aligned planning to seamlessly ground semantic understanding in precise motion execution. We first design an action tokenizer based on a residual vector-quantized variational autoencoder (VQ-VAE), capturing vehicle kinematics and encoding trajectory features into a structured latent space. Rather than discrete codebook lookups that inevitably introduce quantization errors, LaPla repurposes this representation as a physical prior to bridge the modality gap between high-dimensional semantics and the raw action space. Specifically, given multimodal inputs integrating multi-view images, historical actions, and textual instructions, LaPla incorporates concurrent action queries to causally attend to the multimodal context in a single forward pass, projecting hidden states directly into the pretrained VQ-VAE latent space. The frozen decoder then translates these continuous latents into actions, effectively eliminating quantization errors and ensuring physically plausible trajectories while bypassing time-consuming autoregressive generation. Extensive experiments on the nuScenes benchmark demonstrate that LaPla achieves competitive open-loop performance, reducing long-horizon L2 error by 15.52% compared to state-of-the-art VLA methods. Closed-loop evaluations on the NVIDIA AlpaSim simulator further confirm its superior capability in ensuring smooth driving progress, improving the success rate by 33.34 percentage points with significantly reduced inference latency.
Chinese Translation
弥合视觉语言模型的离散推理与自动驾驶连续且受物理约束的特性之间的鸿沟,仍然是一项重大挑战。在本工作中,我们提出了LaPla,一个具有潜在空间对齐规划的统一视觉-语言-动作(VLA)框架,能够将语义理解无缝地落实到精确的运动执行中。我们首先设计了一种基于残差向量量化变分自编码器(VQ-VAE)的动作分词器,以捕捉车辆运动学特性,并将轨迹特征编码到结构化的潜在空间中。LaPla并不采用不可避免地会引入量化误差的离散码本查找,而是将该表示重新用作物理先验,以弥合高维语义与原始动作空间之间的模态鸿沟。具体而言,给定融合多视角图像、历史动作和文本指令的多模态输入,LaPla引入并行的动作查询,在单次前向传播中对多模态上下文进行因果注意力计算,并将隐藏状态直接投影到预训练的VQ-VAE潜在空间中。随后,冻结的解码器将这些连续潜在向量转换为动作,有效消除了量化误差,确保了物理上合理的轨迹,同时避免了耗时的自回归生成。在nuScenes基准上的大量实验表明,LaPla取得了具有竞争力的开环性能,与最先进的VLA方法相比,长时程L2误差降低了15.52%。在NVIDIA AlpaSim模拟器上的闭环评估进一步证实了其在确保平稳驾驶进程方面的卓越能力,成功率提升了33.34个百分点,同时显著降低了推理延迟。
cs.CV / 88 / 2609.04071

TAP-Path: Task-Adaptive Structural and Token Pruning for Efficient and Trustworthy Pathology Foundation Models

TAP-Path:面向高效且可信病理学基础模型的任务自适应结构与标记剪枝方法
Hasan, Mehedi, Yeafi, Ashfak, Islam, Md Khairul
Abstract
Pathology foundation models improve transferable representation learning for histopathology, but recent gains often rely on encoders with hundreds of millions of parameters and high inference cost. We propose TAP-Path, a task-adaptive compression framework that directly restructures a pretrained Virchow2 encoder rather than distilling it into a separate student. TAP-Path combines validation-driven transformer-block selection, physical removal of redundant blocks, input-adaptive patch-token pruning, multi-depth feature recovery, and a lightweight gated task head. The final model retains 24 of 32 transformer blocks and 70% of patch tokens after pruning, reducing encoder parameters by 24.96% (631.24M to 473.70M) and analytical encoder compute by 35.20% (340.13G to 220.40G FLOPs). Across three task-head optimization seeds, TAP-Path achieved $87.98 \pm 0.067%$ test accuracy, $81.26 \pm 0.49%$ balanced accuracy, and $82.38 \pm 0.48%$ macro-F1 on a 32-class histopathology benchmark, compared with 86.89% for full Virchow2 and 87.67% for UNI2-h. TAP-Path achieved a Brier score of $0.1800 \pm 0.0005$ and failure-detection AUROC of $0.9047 \pm 0.0060$. A validation-only rare-aware objective improved rare-class balanced accuracy in a secondary operating analysis. Frozen external evaluation on 433 CPTAC samples yielded $91.22 \pm 0.83%$ accuracy and $91.10 \pm 0.81%$ balanced accuracy. These results show that task-adaptive structural and token sparsification can improve the accuracy-efficiency trade-off of large pathology foundation models while preserving reliability under internal and external evaluation.
Chinese Translation
病理学基础模型提升了组织病理学中可迁移的表示学习能力,但近期的性能提升往往依赖于具有数亿参数和高推理成本的编码器。我们提出了TAP-Path,这是一个任务自适应压缩框架,它直接对预训练的Virchow2编码器进行结构重组,而非将其蒸馏到一个单独的学生模型中。TAP-Path结合了基于验证集驱动的Transformer块选择、冗余块的物理移除、输入自适应的图像块标记(patch token)剪枝、多深度特征恢复以及轻量级的门控任务头。最终模型在剪枝后保留了32个Transformer块中的24个以及70%的图像块标记,使编码器参数量减少24.96%(从631.24M降至473.70M),编码器分析计算量减少35.20%(从340.13G降至220.40G FLOPs)。在三个任务头优化随机种子下,TAP-Path在32类组织病理学基准上取得了87.98±0.067%的测试准确率、81.26±0.49%的平衡准确率和82.38±0.48%的宏平均F1,相比之下完整版Virchow2为86.89%,UNI2-h为87.67%。TAP-Path取得了0.1800±0.0005的Brier分数和0.9047±0.0060的失败检测AUROC。一项仅使用验证集的稀有类感知目标在辅助运行分析中提升了稀有类的平衡准确率。在433个CPTAC样本上的冻结外部评估中,模型取得了91.22±0.83%的准确率和91.10±0.81%的平衡准确率。这些结果表明,任务自适应的结构与标记稀疏化可以在内部和外部评估中保持可靠性的同时,改善大型病理学基础模型的准确率-效率权衡。
cs.CV / 89 / 2609.04083

CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

CORE:通过重排序器蒸馏提升多模态大语言模型嵌入的组合推理能力
Song, Tingyu, Li, Mingxin, Zhang, Yanzhao, Long, Dingkun, Liu, Chu, Xie, Pengjun, Zhao, Yilun, Wu, Shu
Abstract
MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions when used as a cross-attentive reranker, motivating us to distill its compositional judgments into the embedding model. We propose CORE, which synthesizes candidate lists spanning five compositional matching levels and introduces a Rank-KL objective that trains the embedding model to reproduce the reranker's fine-grained ranking. We further introduce a graded evaluation protocol and compare contrastive learning, pairwise CoSENT, and listwise Rank-KL under the same data and tuning budget. Our comparison shows that both CoSENT and Rank-KL use the multi-level supervision more effectively than contrastive learning, with Rank-KL achieving the strongest overall performance. Across three compositional reasoning benchmarks (COLA, SUGARCREPE++, NEGBENCH), CORE-RERANKER-8B achieves an 82.7% total average, outperforming Jina-Reranker by 10.7 points, while CORE-EMBED-8B achieves the best total average (0.666) among all evaluated embedding models. The improvements transfer to the MCMR benchmark without sacrificing retrieval performance on COCO and Flickr30K.
Chinese Translation
基于多模态大语言模型(MLLM)的嵌入模型在组合检索方面仍然存在局限,往往无法区分包含相同概念但属性-物体绑定关系不同的场景。然而,同一骨干模型在作为交叉注意力重排序器(reranker)使用时却能够解决此类区分问题,这促使我们将其组合判断能力蒸馏到嵌入模型中。我们提出了CORE,该方法合成了涵盖五个组合匹配级别的候选列表,并引入了Rank-KL目标函数,训练嵌入模型复现重排序器的细粒度排序。我们还进一步引入了一种分级评估协议,并在相同的数据和调参预算下比较了对比学习、成对式CoSENT和列表式Rank-KL。比较结果表明,CoSENT和Rank-KL都比对比学习更有效地利用了多级别监督,其中Rank-KL取得了最强的整体性能。在三个组合推理基准(COLA、SUGARCREPE++、NEGBENCH)上,CORE-RERANKER-8B取得了82.7%的总平均成绩,比Jina-Reranker高出10.7个百分点,而CORE-EMBED-8B在所有被评估的嵌入模型中取得了最佳的总平均成绩(0.666)。这些提升能够迁移到MCMR基准上,同时不损害模型在COCO和Flickr30K上的检索性能。
cs.CV / 90 / 2609.04088

Efficient Semantic Understanding from Digital Foveation

基于数字中央凹化的高效语义理解
Caccavella, Caterina, Fra, Vittorio, Ziegler, Andreas, D'Angelo, Giulia, Sandamirskaya, Yulia
Abstract
Dense semantic segmentation allocates computational resources uniformly across the entire image, regardless of scene complexity or task relevance. Inspired by biological vision, we investigate whether semantic understanding can be achieved more efficiently through digital foveated perception. We introduce a lightweight active-vision pipeline that combines saliency-driven fixation selection, high-resolution foveal observations, low-resolution contextual information, semantic accumulation, and adaptive computation. Beyond conventional dense prediction metrics, we use object-level evaluation to measure semantic understanding under sparse observations. On ADE20K-Object, a single foveated observation achieves 95.9% of the baseline Top-1 accuracy and 96.9% of the baseline Top-3 accuracy while requiring only 4.7% of the computational cost. At the scene level, semantic accumulation recovers 90.6% of the baseline object recall while using 58.6% of the computation. These results suggest that substantial semantic understanding can emerge from sparse observations when computation is allocated selectively, highlighting active vision as an efficient alternative to uniform dense processing and motivating evaluation protocols beyond conventional pixel-wise segmentation metrics.
Chinese Translation
密集语义分割在整个图像上均匀分配计算资源,而不考虑场景复杂度或任务相关性。受生物视觉启发,我们研究是否可以通过数字中央凹感知(digital foveated perception)更高效地实现语义理解。我们提出一个轻量级的主动视觉流水线,该流水线结合了显著性驱动的注视点选择、高分辨率中央凹观测、低分辨率上下文信息、语义累积以及自适应计算。除传统的密集预测指标外,我们还使用对象级评估来衡量稀疏观测下的语义理解能力。在ADE20K-Object数据集上,单次中央凹观测即可达到基线Top-1准确率的95.9%和基线Top-3准确率的96.9%,而计算成本仅为基线的4.7%。在场景层面,语义累积以58.6%的计算量恢复了基线对象召回率的90.6%。这些结果表明,当计算资源被选择性分配时,稀疏观测也能产生可观的语义理解能力,凸显了主动视觉作为均匀密集处理的高效替代方案的潜力,并启发了超越传统像素级分割指标之外的评估协议。
cs.CV / 91 / 2609.04110

The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs

时间的形状:面向VideoLM时间理解的视频Token对比方法
Shi, Yumeng, Long, Quanyu, Wu, Yin, Wang, Wenya
Abstract
Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet their main supervision acts on generated text rather than video-token representations where event dynamics should first emerge. This mismatch allows models to learn temporal answers from shortcuts such as objects, scenes, and language priors, without requiring internal video representations to capture event progression. To address this, we propose VT-Contrast, a representation-level temporal counterfactual objective for VideoLMs. Its design asks where temporal supervision should act and what temporal differences it should expose. VT-Contrast supervises selected late-layer last-frame video tokens, where temporal information is expected to be integrated before language generation, and contrasts order-preserving views with same-video reordered counterfactuals graded by Kendall tau distance. It requires no architectural changes, is compatible with diverse VideoLM training tasks, and improves overall performance across temporal understanding benchmarks. Our code is available at https://github.com/ANDgate99/VT-Contrast.
Chinese Translation
按顺序观看视频帧并不意味着真正表征了时间。现代VideoLM虽然接收有序的视频流,但其主要的监督信号作用于生成的文本,而非视频token表示——而事件动态本应首先在这些表示中涌现。这种不匹配使得模型可以从物体、场景和语言先验等捷径中学习时间相关的答案,而无需内部的视频表示真正捕捉事件的进展。为解决这一问题,我们提出了VT-Contrast,一种面向VideoLM的表示层时间反事实目标。其设计回答了两个问题:时间监督应作用于何处,以及应暴露哪些时间差异。VT-Contrast对所选后层(late-layer)的末帧视频token进行监督(此处预期时间信息在语言生成之前已被整合),并将保持顺序的视图与同一视频经重排得到的反事实样本进行对比,后者按Kendall tau距离进行分级。该方法无需任何架构改动,可与多种VideoLM训练任务兼容,并在多个时间理解基准上提升了整体性能。我们的代码可在 https://github.com/ANDgate99/VT-Contrast 获取。
cs.CV / 92 / 2609.04120

BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pseudo Data

BooM-VVT:利用图像级伪数据增强无掩码视频虚拟试穿
Zhang, Wei, Li, Xin, Shi, Peishu, Gao, Jialin, Peng, Xuekang, Lian, Zhichao, Jin, Yeying
Abstract
Video virtual try-on (VVT) aims to generate realistic videos of a person wearing a target garment. Recent methods leverage a keyframe-driven video generation paradigm to improve in-the-wild performance, yet they still rely on masks to localize try-on regions, making them vulnerable to large motions and severe occlusions. Although mask-free image-based try-on methods have shown promising results by leveraging large-scale pseudo data, extending this paradigm to videos remains difficult, as constructing video-level pseudo data is prohibitively expensive. Furthermore, coarse keyframe sampling and the scarcity of multi-view try-on data limit existing keyframe-driven methods in maintaining garment consistency and handling diverse try-on tasks. To address these challenges, we propose BooM-VVT, a mask-free VVT framework built upon the keyframe-driven paradigm. To achieve mask-free VVT, we introduce a multi-stage training strategy that leverages image-level pseudo data for mask-free localization learning, substantially reducing the need for costly video-level pseudo data. To improve garment consistency, we propose Garment-Sensitive Keyframe Sampling, which selects keyframes based on garment-relevant body regions to better capture garment appearance. We further introduce Frame-Shared 3D-RoPE to establish spatiotemporal correspondences between keyframes and target video frames for accurate garment-detail transfer. Finally, we construct OmniView, a large-scale multi-view try-on dataset to support reliable try-on video generation under complex camera viewpoints and diverse try-on tasks. Extensive experiments demonstrate that BooM-VVT achieves superior temporal consistency and garment fidelity over existing methods. Project page: https://boomvvt.github.io/boomvvt.
Chinese Translation
视频虚拟试穿(VVT)旨在生成人物穿着目标服装的逼真视频。近期方法利用关键帧驱动的视频生成范式来提升真实场景下的性能,但它们仍然依赖掩码来定位试穿区域,使其在面对大幅度运动和严重遮挡时表现脆弱。尽管无掩码的图像试穿方法通过利用大规模伪数据已展现出良好的效果,但将这一范式扩展到视频仍然困难,因为构建视频级伪数据的成本过高。此外,粗糙的关键帧采样以及多视角试穿数据的稀缺,限制了现有关键帧驱动方法在保持服装一致性和处理多样化试穿任务方面的能力。为应对这些挑战,我们提出了 BooM-VVT,一个基于关键帧范式的无掩码视频虚拟试穿框架。为实现无掩码 VVT,我们引入了一种多阶段训练策略,利用图像级伪数据进行无掩码定位学习,从而大幅减少对昂贵的视频级伪数据的需求。为提升服装一致性,我们提出了服装敏感关键帧采样(Garment-Sensitive Keyframe Sampling),基于与服装相关的身体区域来选择关键帧,以更好地捕捉服装外观。我们进一步引入帧共享三维旋转位置编码(Frame-Shared 3D-RoPE),在关键帧与目标视频帧之间建立时空对应关系,以实现精确的服装细节迁移。最后,我们构建了 OmniView,一个大规模多视角试穿数据集,以支持复杂相机视角和多样化试穿任务下可靠的试穿视频生成。大量实验表明,BooM-VVT 在时间一致性和服装保真度方面优于现有方法。项目页面:https://boomvvt.github.io/boomvvt。
cs.CV / 93 / 2609.04131

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

超越检索:面向流式视频理解的渐进式潜在记忆演化
Qu, Hongyu, Yao, Guangming, Xing, Ling, Hu, Xiaobin, Ding, Rongxing, Zhang, Guibin, Zhang, Fan, Yuan, Yi, Shu, Xiangbo, Yan, Shuicheng
Abstract
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded memory. Existing approaches typically compress historical observations into an external memory bank and retrieve query-relevant evidence as additional visual context. Though effective, this store-and-retrieve paradigm keeps historical evidence as external visual context, preventing it from being internalized into a compact, evolving latent memory that can continuously guide streaming reasoning. To bridge this gap, we introduce LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize. Specifically, LatentStream comprises three coordinated components. First, Query-agnostic Hierarchical Streaming Memory organizes visual history into short-, mid-, and long-term levels under a fixed memory budget through Jenks-guided adaptive consolidation. Once a query arrives, Hierarchical Latent Memory Evolution equips groups of latent memory tokens with progressively expanding memory receptive fields, enabling them to iteratively retrieve historical evidence from their corresponding scopes and internalize it into a compact, fixed-length latent memory. Finally, Progressive Confidence-guided Latent Memory Optimization constructs a hierarchical progression reward from group-wise predictive entropy and jointly refines the latent memory tokens and retrieved evidence, encouraging increasingly confident streaming reasoning. Extensive experiments demonstrate that LatentStream achieves new state-of-the-art results on existing online and offline video benchmarks.
Chinese Translation
流式视频理解要求多模态大语言模型(MLLM)在严格的因果性约束和有限内存条件下处理连续的视觉输入并响应用户查询。现有方法通常将历史观测压缩到外部记忆库中,并检索与查询相关的证据作为额外的视觉上下文。尽管有效,这种“存储—检索”范式将历史证据保持为外部视觉上下文,使其无法被内化为一个紧凑的、可演化的潜在记忆以持续引导流式推理。为弥合这一差距,我们提出了 LatentStream,一个渐进式潜在工作记忆框架,将流式记忆机制从“存储—检索”转变为“检索—内化”。具体而言,LatentStream 包含三个协同工作的组件。首先,查询无关的分层流式记忆通过 Jenks 引导的自适应整合,在固定内存预算下将视觉历史组织为短期、中期和长期三个层级。当查询到达时,分层潜在记忆演化(Hierarchical Latent Memory Evolution)为多组潜在记忆标记配备渐进扩展的记忆感受野,使其能够从各自对应的范围内迭代检索历史证据,并将其内化为紧凑的、固定长度的潜在记忆。最后,渐进式置信度引导的潜在记忆优化(Progressive Confidence-guided Latent Memory Optimization)基于组级预测熵构建分层递进奖励,联合优化潜在记忆标记与检索到的证据,从而鼓励日益自信的流式推理。大量实验表明,LatentStream 在现有的在线和离线视频基准上取得了新的最先进(state-of-the-art)结果。
cs.CV / 94 / 2609.04151

Persistent Identity Preservation in Generative Image Models: A Benchmark and Evaluation System

生成式图像模型中的持久身份保持:基准测试与评估系统
Ren, Mengwei, Zhang, Xuaner, Xia, Zhihao
Abstract
Generative image models can now produce high-quality images, follow complex instructions, and support precise edits, but they still struggle to preserve who or what is being depicted. When generating or editing images of a specific subject, identity may drift as the pose, expression, appearance, viewpoint, or surrounding scene changes. Existing subject-driven methods make fundamentally different choices about where identity is represented: through the input context (GPT-Image-2, NB2), as trainable subject-specific model parameters (LoRA), or as a persistent identity layer (PHOTA IDENTITY) reusable across generations and edits. We systematically benchmark these paradigms across subject-driven generation, editing, restoration, and multi-subject settings, with tasks designed to increasingly stress identity preservation. Our results show that identity preservation remains a distinct limitation of current generative foundation models: strong image quality and instruction following do not necessarily imply strong identity fidelity, and identity degradation becomes more pronounced under iterative edits, small subject scales, severe image degradation, and multi-subject composition. Persistent identity substantially reduces this degradation across generation, editing, and restoration, consistently improving identity preservation when applied to different foundation models while maintaining comparable instruction adherence and perceptual image quality. These results suggest that identity does not simply emerge from increasingly capable generative models, but can instead be represented as persistent subject knowledge that is composed independently with the underlying generative model.
Chinese Translation
生成式图像模型如今已能生成高质量图像、遵循复杂指令并支持精确编辑,但它们仍难以保持图像中所描绘的对象身份。在生成或编辑特定主体的图像时,随着姿态、表情、外观、视角或周围场景的变化,身份可能发生漂移。现有的主体驱动方法在身份表示方式上存在根本不同的选择:通过输入上下文(GPT-Image-2、NB2)、作为可训练的主体特定模型参数(LoRA),或作为可跨生成与编辑任务复用的持久身份层(PHOTA IDENTITY)。我们在主体驱动的生成、编辑、修复和多主体设置下对这些范式进行了系统性基准测试,所设计的任务逐步加大对身份保持的压力。我们的结果表明,身份保持仍是当前生成式基础模型的一个独特局限:强大的图像质量和指令遵循能力并不一定意味着强大的身份保真度,且身份退化在迭代编辑、小尺度主体、严重图像退化以及多主体合成的场景下更加显著。持久身份机制在生成、编辑和修复过程中显著减少了这种退化,在应用于不同基础模型时持续提升身份保持效果,同时保持相当的指令遵循能力和感知图像质量。这些结果表明,身份并非随生成模型能力的增强而自然涌现,而是可以被表示为一种持久的主体知识,与底层生成模型相互独立地进行组合。
cs.CV / 95 / 2609.04174

Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations

基于3D基础模型场景表示的零样本新视角深度合成
Akola, Denis M., Fouhey, David F.
Abstract
3D Foundation Models (3DFMs) such as VGGT have recently pushed the boundaries of 3D vision by predicting rich unified representations with feed-foward transformers. The scene representations learned by these models enable strong performance on multiple 3D vision tasks. In this paper, we investigate using their internal representations to infer 3D in the scene from new views. Our hypothesis is that in order to solve the task of 3D reconstruction, these models need to learn a representation that includes a large amount of general knowledge about 3D scenes. After showing that it is possible to decode hidden surfaces from internal 3DFM representations, we propose a method, Z3D, that estimates pointmaps in unseen views by doing latent diffusion on 3DFM representation. We show that Z3D can predict realistic depth maps for new views across multiple datasets.
Chinese Translation
以VGGT为代表的3D基础模型(3D Foundation Models, 3DFMs)近年来通过前馈Transformer预测丰富统一的场景表示,推动了3D视觉的发展。这些模型学习到的场景表示使其在多个3D视觉任务上取得了出色的性能。本文研究了如何利用其内部表示从新视角推断场景中的3D信息。我们的假设是:为了完成3D重建任务,这些模型需要学习一种包含大量关于3D场景通用知识的表示。在验证了可以从3DFM内部表示中解码出隐藏表面之后,我们提出了一种名为Z3D的方法,该方法通过对3DFM表示进行潜在扩散(latent diffusion)来估计未见视角下的点图(pointmap)。实验表明,Z3D能够在多个数据集上为新视角预测出真实感较强的深度图。
cs.CV / 96 / 2609.04183

Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning

先观察后合成:基于视觉语言模型引导的弱监督密集视频描述过渡事件发现
Kim, Ye-Chan, Choi, Seunghee, Cha, SeungJu, Kim, Si-Woo, Kim, Hwiseon, Kim, Hyungee, Kim, Dong-Jin
Abstract
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide additional vision-language alignment, but these captions lack visual grounding and are rigidly assigned to every inter-event gap at a fixed location and duration. To address these, we propose Seeing Before Synthesizing (SBS), a framework that adaptively provides visually grounded linguistic guidance only where warranted. Leveraging a VLM, we generate frame-level narratives for the inter-event gaps and detect transitions from the semantic variation across them. For identified transitions, we then refine inter-event temporal masks by blending the temporal midpoint with the semantic change point and selecting the width that maximizes vision-language alignment. Experiments on ActivityNet Captions and YouCook2 demonstrate state-of-the-art performance in both captioning and localization.
Chinese Translation
弱监督密集视频描述(Weakly-Supervised Dense Video Captioning)旨在仅给定每个视频一组有序的事件级描述的情况下,对未剪辑视频中的多个事件进行定位和描述。近期工作通过大语言模型(LLM)合成辅助过渡描述以提供额外的视觉-语言对齐,但这些描述缺乏视觉依据,且被僵硬地分配到每个事件间隔的固定位置和固定时长。为解决这些问题,我们提出“先观察后合成”(Seeing Before Synthesizing, SBS)框架,仅在确有需要之处自适应地提供具有视觉依据的语言引导。利用视觉语言模型(VLM),我们为事件间隔生成帧级叙事,并通过帧间语义变化检测过渡。对于识别出的过渡,我们通过将时间中点与语义变化点相融合,并选择使视觉-语言对齐最大化的宽度,从而细化事件间的时间掩码。在 ActivityNet Captions 和 YouCook2 数据集上的实验表明,本方法在描述生成和事件定位任务上均取得了最先进的性能。
cs.CV / 97 / 2609.04190

One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing

一个编辑器,多种编辑:面向多样化视频编辑的统一免训练框架
Juvekar, Adheesh Sunil, Susladkar, Onkar Kishor, Nguyen, Kiet A., Wahed, Muntasir, Bashir, Nabeel, Zhou, Xiaona, Yu, Tianjiao, Shah, Vedant, Lourentzou, Ismini
Abstract
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline, while obtaining competitive results on IVEBench. A user study further shows a 51.8\% overall preference for EditVid over 7 competing methods.
Chinese Translation
视频编辑涵盖多种编辑范式,然而在单一统一框架内实现高质量的指令引导编辑和主体引导编辑仍然具有挑战性。我们提出了 EditVid,这是一个免训练框架,它结合了用于局部一致性的稀疏因果记忆、用于长程身份保持的基于对应关系的后注意力令牌注入,以及用于编辑局部性的软潜在混合。同一框架支持指令引导和参考引导的编辑,包括风格迁移、属性修改、物体插入、部件级编辑和主体替换。在 FiVE 基准上,EditVid 达到了 78.16 的 FiVE-Acc,而最强的免训练基线仅为 58.95,同时在 IVEBench 上取得了具有竞争力的结果。用户研究进一步表明,与 7 种竞争方法相比,用户对 EditVid 的总体偏好率达到 51.8%。
cs.CV / 98 / 2609.04196

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Puffin-World:基于原生3D世界状态的可扩展统一多模态模型
Liao, Kang, Luo, Yihang, Wu, Xiao-Ming, Jin, Linyi, Wu, Size, Lin, Chunyu, Zhao, Yao, Wang, Fei, Li, Wei, Loy, Chen Change
Abstract
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.
Chinese Translation
我们提出了Puffin-World,这是一种统一的多模态架构,能够集成物理理解、空间模拟以及3D世界生成与重建,而无需依赖外部离线模块。为了可靠地构建并与3D世界进行交互,我们的框架联合建模了三种原生世界状态:物理状态(重力场和纬度)、几何状态(深度)和外观状态(图像),并结合统一的Omni-Camera表示,支持多样化任务和灵活的运动。除了对这些状态进行建模之外,我们还引入了一种在未来帧之间传播物理动力学的策略。通过将绝对相机属性锚定于真实世界,Puffin-World能够实现物理一致且视觉稳定的世界生成。我们进一步在单一生成过程中耦合外观与几何,联合合成每个未来视图并重建其底层几何结构。这一统一范式支持需要多任务协同的交错式闭环应用,包括模仿式世界探索和自校准世界探索。为了将Puffin-World扩展至复杂场景,我们构建了Puffin-16M数据集,包含1500万条视觉-语言-相机三元组以及100万条具有多样且具挑战性运动的轨迹。为促进该领域的进一步研究,我们已发布代码、模型和数据集。
cs.CV / 99 / 2609.04200

Principia: Relational Physics Tests for Video Models

Principia:面向视频模型的关系性物理测试
Thozhiyoor, Varun Varma, Tripathi, Shivam, Radhakrishnan, Venkatesh Babu, Bhattad, Anand
Abstract
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena - gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation - across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench. Vision-language models are evaluated on their ability to detect relational physics violations, with the best model achieving only 67% accuracy and most performing near chance level.
Chinese Translation
评估视频模型的物理推理能力十分困难,因为绝对运动测量依赖于帧率、物体尺度和相机标定,而这些因素在生成视频中通常是模糊或不可得的。我们提出了一种不同的方法:当同一场景中的两个物体遵循同一物理定律时,它们的运动必然满足可预测的关系,且这些关系独立于标定而成立。我们提出了Principia,一个通过成对物体之间的关系一致性来评估牛顿物理的基准。Principia涵盖八种现象——重力、恢复系数、摩擦、转动惯量、抛体运动、动量、单摆和质量-弹簧振荡——覆盖平移、旋转、碰撞和振荡动力学,并使用在受控协议下录制的真实世界场景。我们还引入了一个不依赖标定的一致性分数,可直接在图像空间中量化物理违背情况。在对六个最先进视频生成器的数千个生成结果进行评估中,尽管所有模型在VBench上的得分约为0.8,但没有模型在Principia上超过0.42。我们还评估了视觉-语言模型检测关系性物理违背的能力,最佳模型仅达到67%的准确率,而大多数模型的表现接近随机水平。
cs.CV / 100 / 2609.04201

Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction

Scal3R:学习高效的多相对位姿查询以实现可扩展的在线三维重建
Lin, Chin-Yang, Sun, Yang-Che, Sun, Cheng, Yang, Fu-En, Chen, Min-Hung, Lin, Yen-Yu, Chiu, Wei-Chen, Liu, Yu-Lun
Abstract
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fixed first-frame anchor forces extrapolation far beyond the training distribution. Small drifts accumulate and amplify into significant geometric collapse. However, we observe that per-frame depth remains stable throughout this failure. The backbone's local geometry remains intact; only the global pose head breaks down. Motivated by this decoupling, we introduce Scal3R. This approach reformulates online reconstruction as multi-reference relative pose querying. We use lightweight learnable tokens, which make up about ~1% of the parameters, and inject them into a completely frozen backbone via asymmetric attention. This setup queries poses relative to multiple past keyframes. An online pose-graph optimization system with loop closure suppresses long-range drift. Scal3R reaches convergence in 8 hours on a single GPU. It reduces the average ATE by over 60% on KITTI compared to the online baseline. It also achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes. Project page: https://linjohnss.github.io/scal3r/
Chinese Translation
在线三维重建模型在长视频上表现不佳。其原因在于,以固定的第一帧锚点回归相对位姿会导致外推远超出训练分布,微小的漂移不断累积并放大,最终引发严重的几何坍塌。然而,我们观察到逐帧深度在这一失效过程中始终保持稳定:骨干网络(backbone)的局部几何保持完好,只有全局位姿头出现崩溃。受这种解耦性的启发,我们提出了 Scal3R。该方法将在线重建重新表述为多参考相对位姿查询。我们使用轻量的可学习 token(约占参数量的 1%),并通过非对称注意力机制将其注入完全冻结的骨干网络中,从而查询相对于多个历史关键帧的位姿。同时,一个带闭环检测的在线位姿图优化系统抑制了长程漂移。Scal3R 在单块 GPU 上 8 小时内即可收敛;与在线基线相比,它在 KITTI 上将平均 ATE 降低了 60% 以上,并在 Virtual KITTI、Sintel、TUM-Dynamic、ScanNet 和 7-Scenes 上均取得了最先进的性能。项目主页:https://linjohnss.github.io/scal3r/
机器人学 (Robotics)
36
cs.RO / 1 / 2609.03055

Seeing Less Is Not Seeing Safely: Privacy Leakage from Task-Scoped Robot Perception Exports

看得少并不等于看得安全:来自任务限定机器人感知导出的隐私泄露
Xu, Yuqiao, Ayday, Erman
Abstract
Domestic robots rely on rich perception to operate in private homes, but privacy risk persists even when raw sensor data remain local. Structured representations exported to downstream planners, cloud services, logs, or learning pipelines can still reveal household information through semantics, geometry, spatial structure, and task targets. We introduce Task-Functional Perception Distillation (TFPD), a task-scoped representation-export framework that keeps rich perception local and profiles downstream exports according to task utility, direct exposure, and multiple residual inference risks. Using 120 AI2-THOR scenes with scene-disjoint train/validation/test splits, frozen attacker selection, and representation-aware held-out attacks, we evaluate navigation, collision checking, and object-goal execution. Three navigation exports achieve identical success (1.000) and mean path ratio (0.898), yet representation-level linkability ranges from 0.532 to 0.970. Replacing an explicit target label with a target region reduces target-category macro-F1 from 1.000 to 0.077 while preserving success at 0.995, while geometric coarsening reduces object-category macro-F1 from 0.704 to 0.556 at a measurable collision-utility cost. A ProcTHOR replication preserves the navigation task-equivalence/privacy-inequivalence finding while changing the relative ordering of normalized and topological exports. These results show that neither field removal nor stronger abstraction induces a universal privacy ordering and motivate task-specific, multi-risk evaluation of the complete public representation.
Chinese Translation
家用机器人依赖丰富的感知能力在私人住宅中运行,但即使原始传感器数据保留在本地,隐私风险依然存在。导出至下游规划器、云服务、日志或学习流水线的结构化表示,仍然可能通过语义、几何、空间结构和任务目标泄露家庭信息。我们提出了任务功能感知蒸馏(Task-Functional Perception Distillation, TFPD),这是一个任务限定的表示导出框架,它将丰富的感知保留在本地,并根据任务效用、直接暴露程度以及多种残余推理风险来评估下游导出的表示。我们在120个AI2-THOR场景上进行了实验,采用场景不相交的训练/验证/测试划分、冻结攻击者选择以及具备表示感知能力的留出攻击,评估了导航、碰撞检测和目标物体执行三项任务。三种导航导出表示取得了完全相同的成功率(1.000)和平均路径比(0.898),但其表示层面的可链接性却在0.532到0.970之间。将显式目标标签替换为目标区域,可将目标类别宏F1从1.000降至0.077,同时保持0.995的成功率;而几何粗化则将物体类别宏F1从0.704降至0.556,并带来可测量的碰撞效用损失。在ProcTHOR上的复现实验保持了导航任务等效性与隐私非等效性的结论,但改变了归一化导出与拓扑导出之间的相对排序。这些结果表明,无论是删除字段还是更强的抽象化,都无法产生普适的隐私排序,从而推动了对完整公开表示进行任务特定、多风险评估的研究。
cs.RO / 2 / 2609.03067

GPU-Accelerated Astrodynamics World Models for Spacecraft Rendezvous and Proximity Operations

面向航天器交会与接近操作的GPU加速天体动力学世界模型
Eddy, Duncan, Ward, Isaac R., Kim, Grace Ra, Kochenderfer, Mykel J.
Abstract
World models are an emerging paradigm in representation learning in which an agent jointly learns state-action dynamics and observation models from offline trajectory data, enabling multi-step planning and trajectory prediction with uncertainty estimates. They have shown strong results in robotics and game environments, but, to the best of our knowledge, have not previously been applied to the space domain. This paper introduces a world model-based approach to cooperative and non-cooperative spacecraft rendezvous and proximity operations. First, we introduce an open-source, JAX-based International Space Station (ISS) docking environment supporting parallel GPU simulation of spacecraft orbit and attitude dynamics, generating the thousands of state-action transitions that world model training requires. Second, we introduce Out-of-this-World-Model, a transformer-based world model that encodes relative kinematic states and body-fixed camera imagery into a latent state and predicts its evolution under commanded thrusts and torques using one-step flow matching. It produces a distribution over future observations, capturing stochastic dynamics and per-timestep uncertainty, and outperforms DreamerV3-style posterior-correction baselines with fewer trainable parameters and hyperparameters. Third, we apply the approach to a capsule autonomously docking with the ISS under keep-out-zone constraints, demonstrating improved sample efficiency and task performance over reinforcement learning baselines (53% versus 29% docking success across ports), better out-of-distribution generalization (on held-out ports the world model more than doubles baseline success, 40% versus 17%), and detection of anomalous objects encountered during approach with 98% classification accuracy. We open-source the simulation environment and model architecture to enable further study of this paradigm.
Chinese Translation
世界模型(World Models)是表征学习中的一种新兴范式,智能体通过离线轨迹数据联合学习状态-动作动力学与观测模型,从而实现多步规划以及带不确定性估计的轨迹预测。世界模型在机器人和游戏环境中已展现出优异效果,但据我们所知,此前尚未被应用于空间领域。本文提出了一种基于世界模型的方法,用于合作与非合作航天器的交会与接近操作(RPO)。首先,我们引入了一个基于JAX的开源国际空间站(ISS)对接环境,支持航天器轨道与姿态动力学的并行GPU仿真,可生成世界模型训练所需的数千条状态-动作转移数据。其次,我们提出了Out-of-this-World-Model,一种基于Transformer的世界模型,它将相对运动学状态与固连于本体的相机图像编码为潜在状态,并利用单步流匹配(flow matching)在给定指令推力与力矩条件下预测其演化。该模型输出对未来观测的分布,能够刻画随机动力学与逐时间步的不确定性,且在可训练参数和超参数更少的情况下优于DreamerV3式后验修正基线。第三,我们将该方法应用于在禁飞区约束下与ISS自主对接的航天器,证明了其相对于强化学习基线具有更高的样本效率与任务性能(各对接口的平均对接成功率分别为53%和29%)、更强的分布外泛化能力(在留出对接口上,世界模型的成功率超过基线的两倍,分别为40%和17%),以及在接近过程中对异常物体的检测能力(分类准确率达98%)。我们开源了仿真环境与模型架构,以推动对这一范式的进一步研究。
cs.RO / 3 / 2609.03142

Sensing Which Modality Matters: Evidence-Gated Regularization for Robust VLA Policies

感知何种模态重要:用于鲁棒视觉-语言-动作策略的证据门控正则化
Yang, Yue, Romeres, Diego, Hori, Chiori, Bertasius, Gedas, Szafir, Daniel, Jain, Siddarth
Abstract
Vision-Language-Action (VLA) policies fuse multimodal sensory inputs, but training on limited and homogeneous robot demonstrations encourages spurious inter-sensor correlations rather than task-relevant signal, a failure we term modality entanglement. Under real-world occlusions and distractors, this manifests as nuisance sensitivity to corruption of uninformative sensors and single-modality insufficiency when only one informative sensor remains intact. We propose Evidence-Gated Regularization (EGR), a modality-agnostic training objective that introduces zero inference-time overhead. EGR derives a per-frame and per-sensor task-relevance signal to gate two state-conditional consistency objectives: invariance on low-evidence sensors, and single-sensor sufficiency on high-evidence ones. We introduce a benchmark based on BEHAVIOR-1K, comprising a fast inference-only diagnostic suite and 47 rollout-based skills targeting modality entanglement. We validate EGR on this benchmark and on two real-robot setups with fundamentally different embodiments: a bi-manual setup with two Kinova arms and three RGB cameras, and a single-arm MELFA ASSISTA setup combining vision and GelSight tactile sensors. EGR improves simulation success rates (SR) from 12.5% to 16.4% under full modalities (+31%), from 9.4% to 16.5% under uninformative-sensor corruption (+75%), and from 2.8% to 6.1% under single-sensor fallback (+120%). Under physical-object distractors, EGR boosts SR from 30% to 85% on the bi-manual setup (+183%) and from 55% to 70% on the tactile setup (+27%).
Chinese Translation
视觉-语言-动作(VLA)策略融合多模态感知输入,但在有限且同质化的机器人演示数据上训练,会使模型学到伪相关的传感器间关联而非任务相关的信号,我们将这种失效模式称为模态纠缠。在真实世界的遮挡和干扰物条件下,这种问题表现为对无信息传感器损坏的过度敏感性,以及在仅剩一个有效信息传感器时的单模态能力不足。我们提出证据门控正则化(Evidence-Gated Regularization, EGR),这是一种与模态无关的训练目标,且在推理时不引入任何额外开销。EGR 为每一帧、每个传感器推导任务相关性信号,用于门控两个状态条件一致性目标:对低证据传感器施加不变性约束,对高证据传感器施加单传感器充分性约束。我们基于 BEHAVIOR-1K 构建了一个基准,包含一个快速纯推理诊断套件和 47 个针对模态纠缠的基于滚动执行(rollout)的技能任务。我们在该基准以及两个具有本质不同本体结构的真实机器人平台上验证了 EGR:一个配备两只 Kinova 机械臂和三个 RGB 相机的双臂平台,以及一个结合视觉与 GelSight 触觉传感器的单臂 MELFA ASSISTA 平台。EGR 将仿真成功率(SR)在全模态条件下从 12.5% 提升至 16.4%(提升 31%),在无信息传感器损坏条件下从 9.4% 提升至 16.5%(提升 75%),在单传感器回退条件下从 2.8% 提升至 6.1%(提升 120%)。在物理物体干扰条件下,EGR 将双臂平台的 SR 从 30% 提升至 85%(提升 183%),将触觉平台的 SR 从 55% 提升至 70%(提升 27%)。
cs.RO / 4 / 2609.03175

Real-Time Shape Control of Multi-Segment Soft Robotic Arms Using Koopman Operators with Global and Local Observables

基于全局与局部可观测量的Koopman算子多段软体机械臂实时形状控制
Wang, Jiahe, Ristich, Eron, Ali, Sultan Haidar, weissman, Eric, Zhang, Lei, Jin, Wanxin, Ren, Yi, Sun, Jiefeng
Abstract
Multi-segment soft robotic arms can continuously reconfigure their body shapes for safe interaction, but tip control alone is insufficient for constrained-space tasks. Therefore, shape control is a more important task for multi-segment soft arms than tip control, but remains challenging due to the high dimensionality and nonlinear dynamics of continuum deformation. In existing work, shape control accuracy is defined by the error in the global frame (global shape error). For multi-segment soft arms, using only global shape error as the control objective is insufficient, as segment coupling, gravity-induced loading, and inertial effects become more significant. This difficulty increases with the number of segments. In this paper, we present a Koopman-based model predictive control framework that combines global and local observables, enabling real-time shape control on multi-segment soft robotic arms. The framework is evaluated through numerical and physical experiments. Numerical experiments demonstrate the scalability of the proposed controller by achieving shape control on robots with up to 10 independently actuated segments. The physical experiments demonstrate that the controller is capable of (1) real-time shape control of 3- and 5-segment robotic arms with tip speeds up to 0.6 m/s, (2) robust tracking without retraining, including distal payloads up to 400~g and recovery from a 7~N lateral disturbance, and (3) the potential for future inspection applications through a confined-space demonstration. These results demonstrate that the proposed framework enables dynamic, scalable, and accurate real-time shape control on multi-segment soft robotic arms.
Chinese Translation
多段软体机械臂能够连续重构其身体形状以实现安全交互,但在受限空间任务中,仅依靠末端控制是不够的。因此,对于多段软体机械臂而言,形状控制比末端控制更为重要,但由于连续体变形的高维度和非线性动力学特性,形状控制仍然极具挑战性。在现有工作中,形状控制精度由全局坐标系下的误差(全局形状误差)来定义。对于多段软体机械臂,仅以全局形状误差作为控制目标是不充分的,因为段间耦合、重力引起的载荷以及惯性效应变得更加显著,且这一难度随段数的增加而增大。本文提出了一种基于Koopman算子的模型预测控制框架,该框架结合了全局与局部可观测量,能够实现对多段软体机械臂的实时形状控制。该框架通过数值实验和物理实验进行了评估。数值实验表明,所提出的控制器在具有多达10个独立驱动段的机器人上实现了形状控制,验证了其可扩展性。物理实验表明,该控制器能够:(1)对3段和5段机械臂进行实时形状控制,末端速度高达0.6 m/s;(2)无需重新训练即可实现鲁棒跟踪,包括承受高达400 g的远端负载以及从7 N侧向扰动中恢复;(3)通过一个受限空间演示,展示了未来在检测应用中的潜力。这些结果表明,所提出的框架能够实现对多段软体机械臂动态、可扩展且精确的实时形状控制。
cs.RO / 5 / 2609.03222

Following a Unique Path: A Fast Certifier Applied to Outlier-Robust Pose Registration

沿唯一路径前行:一种应用于抗外点位姿配准的快速可验证方法
Holmes, Connor, Goudar, Abhishek, Barfoot, Timothy D.
Abstract
Certifiable methods have arisen as a means to guarantee global optimality of solutions to non-convex problems using convex semidefinite programming (SDP) relaxations. The most performant of these methods use a local solver to obtain the candidate solution, and then certify its optimality using efficient linear algebra techniques. However, for many problems of interest in robotics, this local-solve-then-certify approach is impeded by a form of degeneracy in the relaxation, leaving a costly optimization of the relaxation as the only recourse. In this paper, we introduce our Central-Path Certifier (CP-Cert), a certifiable method explicitly tailored to certify candidate optima to problems that exhibit this form of degeneracy. Using a candidate as a starting point, our approach seeks a nearby region of the feasible space -- known as the central path -- where a valid certificate can be readily obtained. The approach is kept efficient by exploiting indirect linear algebra techniques, problem sparsity, and parallelism. We apply CP-Cert to both matrix-weighted pose registration and pointcloud data association, whose novel SDP relaxation is of independent interest. On simulated examples, we explore the properties of this novel relaxation and show that CP-Cert is fast and scalable, achieving runtimes that are up to three orders of magnitude faster than state-of-the-art direct solvers. Finally, we combine these contributions into a certifiable, outlier-robust pose-estimation pipeline, which we apply to real-world data.
Chinese Translation
可验证(certifiable)方法作为一种手段,通过凸半定规划(SDP)松弛来保证非凸问题解的全局最优性。其中性能最佳的方法先使用局部求解器获得候选解,再利用高效的线性代数技术验证其最优性。然而,对于机器人学中许多重要问题而言,这种“先局部求解后验证”的方法会因松弛问题的一种退化性而受阻,使得对松弛问题的代价高昂的优化成为唯一选择。本文提出了中心路径验证器(Central-Path Certifier, CP-Cert),这是一种专门为验证具有此类退化性的问题的候选最优解而设计的可验证方法。该方法以候选解为起点,在可行域中寻找一个附近的区域——即中心路径——在该区域内可以便捷地获得有效证书。通过利用间接线性代数技术、问题稀疏性以及并行计算,该方法保持了高效性。我们将 CP-Cert 应用于矩阵加权的位姿配准和点云数据关联问题,其中针对后者的新型 SDP 松弛本身也具有独立的研究价值。在仿真实验中,我们探究了这一新型松弛的性质,并证明 CP-Cert 快速且具有良好的可扩展性,其运行时间比最先进的直接求解器最多快三个数量级。最后,我们将上述贡献整合为一个可验证的抗外点位姿估计流水线,并将其应用于真实世界数据。
cs.RO / 6 / 2609.03225

Long-Horizon Consistent and Interaction-Aware World Models for Multi-Style End-to-End Driving

面向多风格端到端驾驶的长时程一致性与交互感知世界模型
Han, Yuxuan, Wu, Kunyuan, Yang, Liyunong, Wang, Zilu, Jiang, Cansen, Xiao, Yi, Hu, Liang
Abstract
End-to-end autonomous driving has increasingly adopted world model-based reinforcement learning frameworks to improve learning efficiency through \textit{imagined rollouts}. However, existing world models suffer from three key limitations: temporal inconsistency in long-horizon imagined rollouts, inadequate modeling of ego-environment interactions, and limited adaptability to diverse driving styles. To address these challenges, we propose \textit{StyleDrive}, a world-model-based learning framework that jointly enforces long-horizon consistency, explicitly disentangles interactive traffic states, and supports multi-style policy optimization within a unified learning paradigm. First, we introduce a temporal consistency regularization that integrates historical latent states through gated cross-attention, stabilizing long-horizon imagined rollouts and mitigating error accumulation. Second, we design an explicit state disentanglement module that separates ego-relevant from ego-irrelevant interactive states, enabling more interpretable and efficient decision-making in complex traffic scenarios. Third, we enable multi-style driving behaviors through Group Relative Policy Optimization, which replaces per-step reward optimization with trajectory-wise relative advantages, reducing reward variance and supporting diverse driving styles without retraining. We evaluate StyleDrive on the Bench2Drive closed-loop driving benchmark, achieving a driving score of 88.44 (+17.08 over the previous best world model-based method) and a success rate of 66.82 (+16.58). Furthermore, we deploy StyleDrive on a real automated guided vehicle platform and demonstrate promising sim-to-real transfer capability in dynamic driving scenarios.
Chinese Translation
端到端自动驾驶日益采用基于世界模型的强化学习框架,通过想象推演提升学习效率。然而,现有世界模型存在三个关键局限:长时程想象推演中的时间不一致性、对自车与环境交互的建模不足,以及对多样化驾驶风格适应能力有限。为应对这些挑战,我们提出StyleDrive,一个基于世界模型的学习框架,它在统一学习范式下联合强化长时程一致性、显式解耦交互交通状态,并支持多风格策略优化。首先,我们引入一种时间一致性正则化方法,通过门控交叉注意力机制整合历史潜在状态,稳定长时程想象推演并缓解误差累积。其次,我们设计了一个显式状态解耦模块,将与自车相关和无关的交互状态分离开来,从而在复杂交通场景中实现更具可解释性和更高效的决策。第三,我们通过组相对策略优化(Group Relative Policy Optimization)实现多风格驾驶行为,该方法是替换逐步奖励优化,采用轨迹级相对优势,降低奖励方差,并在无需重新训练的情况下支持多样化驾驶风格。我们在Bench2Drive闭环驾驶基准上评估StyleDrive,取得了88.44的驾驶分数(较此前最佳的基于世界模型的方法提升17.08)和66.82的成功率(提升16.58)。此外,我们将StyleDrive部署于真实的自动导引车平台,并展示了其在动态驾驶场景中良好的仿真到现实迁移能力。
cs.RO / 7 / 2609.03255

Establishing a Dynamic Multimodal HRI Dataset for Engagement Analysis with a Humanoid Robot

建立用于人形机器人参与度分析的动态多模态人机交互数据集
Kim, Buwan, Jo, Wonse
Abstract
This paper presents an experimental design for constructing a multimodal dataset to analyze user engagement in human-robot interaction (HRI). Prior studies have mainly relied on observable behavioral cues, with limited frameworks integrating physiological signals. We therefore propose a structured data-collection protocol to build a multimodal dataset that includes wearable physiological signals, behavioral data, and self-report measures under different levels of task complexity defined in this experiment.
Chinese Translation
本文提出了一种实验设计,用于构建多模态数据集,以分析人机交互(HRI)中的用户参与度。以往的研究主要依赖于可观察的行为线索,而整合生理信号的框架较为有限。因此,我们提出了一种结构化的数据采集方案,以构建一个多模态数据集,其中包含可穿戴设备采集的生理信号、行为数据以及在本文实验中定义的不同任务复杂度水平下的自我报告测量数据。
cs.RO / 8 / 2609.03276

R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models

R2S-Eval:基于视觉-语言模型的实到虚校准机器人评估方法
Wang, Yidi, Ruan, Feixiang, Chen, Ruoqu, Yin, Jie, Yu, Yang, Xu, Mengdi, Zhang, Kaifeng
Abstract
Evaluating robot manipulation policies is becoming increasingly important as generalist models, particularly vision-language-action (VLA) models, are deployed on physical robots. However, conventional real-world evaluation remains labor-intensive, unstable, and insufficiently informative. It requires repeated hardware trials, manual scene resets, and continuous operator monitoring, may produce different policy rankings across repeated evaluations, and primarily relies on success-rate metrics that provide limited information about execution quality. In contrast, humans assess robot performance by observing and comparing complete behaviors rather than relying solely on binary success outcomes. To this end, we propose R2S-Eval, an evaluation pipeline that combines real-to-sim calibration with vision-language model (VLM) preference evaluation. The real-to-sim component efficiently generates rollout videos in a simulator calibrated to the real-world evaluation setting, thereby reducing the need for repeated hardware trials. The VLM evaluator assesses the execution quality of rollout videos and produces pairwise preferences, which are subsequently aggregated into policy rankings. We further introduce a protocol to assess whether the proposed evaluation pipeline yields validated policy conclusions while mitigating the key challenges of conventional real-world evaluation. Experiments in both simulation and real-world settings demonstrate that R2S-Eval produces reliable and stable policy conclusions, achieves agreement with human preferences, substantially reduces repeated hardware-operation effort, and reveals behavior-quality differences that are not captured by binary success labels. In general, R2S-Eval advances robot evaluation from manual success counting toward automated, statistically stable, and quality-aware evaluation of robot behavior. Project page: https://r2s-eval.github.io.
Chinese Translation
随着通用模型,特别是视觉-语言-动作(VLA)模型被部署到物理机器人上,机器人操作策略的评估正变得日益重要。然而,传统的真实世界评估仍然劳动密集、不稳定且信息量不足:它需要反复的硬件试验、人工场景复位以及操作员的持续监控;重复评估可能产生不同的策略排名;且主要依赖成功率指标,而这些指标只能提供关于执行质量的有限信息。相比之下,人类通过观察和比较完整的行为来评估机器人性能,而不仅依赖于二元的成功结果。为此,我们提出了R2S-Eval,一个将实到虚(real-to-sim)校准与视觉-语言模型(VLM)偏好评估相结合的评估流程。实到虚部分在与真实世界评估环境校准一致的模拟器中高效生成回放视频,从而减少了对反复硬件试验的需求。VLM评估器评估回放视频的执行质量并产生成对偏好,随后将这些偏好聚合为策略排名。我们进一步引入一套评估协议,用于检验所提出的评估流程能否在缓解传统真实世界评估关键挑战的同时,产生经过验证的策略结论。在模拟和真实世界环境中的实验表明,R2S-Eval能够产生可靠且稳定的策略结论,与人类偏好保持一致,大幅减少反复硬件操作的工作量,并揭示出二元成功标签无法捕捉的行为质量差异。总体而言,R2S-Eval推动机器人评估从人工成功计数迈向自动化、统计稳定且质量感知的机器人行为评估。项目主页:https://r2s-eval.github.io。
cs.RO / 9 / 2609.03362

ARTiS: An Adaptive Robotic Gripper for Enhanced Tool Manipulation in Disassembly Applications

ARTiS:一种用于增强拆卸应用中工具操作能力的自适应机械夹持器
Mykhailyshyn, Roman, Yukiyasu, Domae, Kensuke, Harada
Abstract
Grasping and holding tools while using them presents a considerable challenge not only for robots but also for humans. Such a challenge is particularly noticeable in processes involving assembly and disassembly, where efficiency and consistency depend on performing rapidly adaptive tasks. Nonetheless, contemporary robotic grasping technologies that can securely manipulate tools during operation frequently have significant constraints. In this paper, introduce ARTiS (Adaptive Robotic Tool Gripper in Disassembly Systems), a novel gripper that combines the adaptability of soft grippers, the dexterity of anthropomorphic hands, and the robustness of rigid mechanisms with a soft palm and fingertips. This unique combination makes it possible to hold tools securely in a variety of situations through using active jamming in the palm and fin-ray adaptation in fingertips. Furthermore, high finger dexterity is achieved through the seven degrees of freedom design, which enables the fingertips to orient to any surface, both for automated solutions and collaborative tasks. A comprehensive evaluation was conducted using a range of conventional disassembly tools to assess the gripper's compliance, durability, and functional versatility. More information, hardware instructions, and videos at https://romanmykhailyshyn.github.io/artis/
Chinese Translation
在使用工具时抓取并握持工具,不仅对机器人而言,对人类而言也是一项巨大的挑战。这一挑战在涉及装配与拆卸的过程中尤为明显,因为其效率和一致性依赖于快速自适应任务的执行。然而,现有能够在工具操作过程中稳固操控工具的机器人抓取技术往往存在显著局限。本文介绍了一种新型夹持器ARTiS(Adaptive Robotic Tool Gripper in Disassembly Systems,面向拆卸系统的自适应机械工具夹持器),它融合了软体夹持器的自适应性、拟人手的灵巧性,以及带有软性手掌和指尖的刚性机构的稳健性。这种独特的组合通过手掌中的主动干扰(active jamming)和指尖的鳍条(fin-ray)自适应结构,使其能够在多种情境下稳固地握持工具。此外,通过七自由度设计实现了指尖的高灵巧性,使指尖能够朝向任意表面,既适用于自动化方案,也适用于协作任务。我们使用多种常规拆卸工具对夹持器的柔顺性、耐用性和功能多样性进行了全面评估。更多信息、硬件说明和视频请访问 https://romanmykhailyshyn.github.io/artis/
cs.RO / 10 / 2609.03392

Programming and execution of skill-based human-robot-crane collaborative tasks

基于技能的人-机器人-起重机协作任务的编程与执行
Lohi, Taneli, Suomalainen, Markku, Mellanen, Roope, Heikkilä, Tapio
Abstract
Highly varying production sets increasing challenges for robotic manufacturing and indoor logistics. New capabilities for agility, flexibility, and robustness are needed. Robot skills, integrating motions, tool operations, and sensor perceptions consistently provide an execution mechanism for a versatile set of tasks with varying parameters. In this paper, easy-to-use CAD-model based programming and execution system for parametrized skills and skill monitors is showcased. The execution control structure is dynamic and parametrized, based on a modified Behavior Tree, where only event based communication is used. A human-robot-crane collaborative skill is shown as a test example, where a human instructs an overhead crane and a manipulator in inserting a heavy object supported by the crane, and guided by the manipulator, into the goal.
Chinese Translation
高度多变的生产品种给机器人制造和室内物流带来了日益严峻的挑战,需要新的敏捷性、灵活性和鲁棒性能力。机器人技能将运动、工具操作和传感器感知整合为一体,为具有不同参数的多样化任务集提供了一致的执行机制。本文展示了一种易于使用的、基于CAD模型的参数化技能与技能监控器的编程与执行系统。执行控制结构是动态且参数化的,基于一种改进的行为树(Behavior Tree),其中仅使用基于事件的通信。文章以一个人-机器人-起重机协作技能作为测试示例:在该示例中,操作人员指挥桥式起重机和机械臂,在起重机的支撑和机械臂的引导下,将一重物插入目标位置。
cs.RO / 11 / 2609.03483

Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps

基于共享鸟瞰图的空地协同视觉语言导航
Zhang, Shuning, Li, Liang, Wang, Yunheng, Wang, Tao, Kang, Yihang, Xu, Renjing
Abstract
Air-ground collaborative Vision-and-Language Navigation (VLN) pairs an unmanned aerial vehicle (UAV) with a global bird's-eye view and an unmanned ground vehicle (UGV) with a local first-person view, yet the setting remains largely unexplored: existing training-free methods solve single-agent tasks but offer no collaboration mechanism, and a recent CARLA-Air evaluation found no stable cooperative behavior across five state-of-the-art VLA models; naive semantic communication or bidirectional coupling even degrades performance. We establish AGC-VLN (Air-Ground Collaborative VLN), the first training-free baseline for air-ground collaborative VLN. The key insight is that training-free methods decompose navigation into VLM-based semantic reasoning and deterministic geometric execution, exposing a collaboration interface: the UAV's global view, over which it renders the UGV's reported pose and the VLM-anchored target as CAR/GOAL markers with distance labels, yielding a shared bird's-eye map. From this map, the UGV acquires global spatial context its first-person view cannot provide, plans a road-following path with a frozen VLM, and executes it under closed-loop control; in parallel, the UAV runs 3D-SPF, a spatial-search upgrade of SPF that localizes the target in the downward view and flies toward it. On 100 closed-loop episodes in CARLA-Air's Town10HD scene, AGC-VLN reaches a 77.0% joint success rate, a collaboration gain of +27.0% over the weaker individual agent (the UAV, 50.0%), and exceeds the strongest published single-agent baseline (Travel UAV, 53.0%) by 24.0 points, stemming from the complementarity of the UAV's global view and the UGV's road-following execution. Project page: https://github.com/ZSN2024/AGC-VLN.
Chinese Translation
空地协同视觉语言导航(VLN)将具有全局鸟瞰视野的无人机(UAV)与具有局部第一人称视角的无人地面车辆(UGV)配对,然而该任务设置在很大程度上仍待探索:现有的免训练方法仅能解决单智能体任务而缺乏协作机制,且近期的 CARLA-Air 评估发现五个最先进的 VLA 模型均未表现出稳定的协作行为;朴素的语义通信或双向耦合甚至会降低性能。我们建立了 AGC-VLN(空地协同视觉语言导航),这是首个面向空地协同 VLN 的免训练基线方法。其核心洞察在于:免训练方法将导航分解为基于 VLM 的语义推理和确定性的几何执行,从而暴露出一个协作接口:无人机在其全局视图上渲染 UGV 上报的位姿以及以 VLM 锚定的目标,绘制为带有距离标签的 CAR/GOAL 标记,从而生成共享的鸟瞰地图。基于该地图,UGV 获得其第一人称视角无法提供的全局空间上下文,利用冻结的 VLM 规划沿道路的路径,并在闭环控制下执行该路径;与此同时,无人机运行 3D-SPF——SPF 的空间搜索升级版——在下视图像中定位目标并飞向目标。在 CARLA-Air 的 Town10HD 场景的 100 个闭环回合中,AGC-VLN 达到了 77.0% 的联合成功率,相比较弱的单个智能体(无人机,50.0%)提升了 27.0% 的协作增益,并比已发表的最强单智能体基线(Travel UAV,53.0%)高出 24.0 个百分点,这得益于无人机全局视野与 UGV 沿道路执行能力之间的互补性。项目主页:https://github.com/ZSN2024/AGC-VLN。
cs.RO / 12 / 2609.03497

BRIDGE: An Open-Source Humanoid Platform via Morphology-Control Co-Design for Physical AI

BRIDGE:一种基于形态-控制协同设计的开源人形机器人平台,面向具身物理智能
Wang, Jianren, Qian, Letian, Wang, Zikai, Wu, Weiwei, Zong, Junjie, Gupta, Abhinav, Pathak, Deepak
Abstract
Developing humanoid robots capable of leveraging human behavioral data is essential for general-purpose embodiment, yet conventional development remains bottlenecked by a decoupled paradigm that isolates hardware design from whole-body control. This approach leads to suboptimal systems that compromise human-like fluidity and agility. To bridge this gap, we introduce a data-driven morphology-control co-design framework that optimizes humanoid morphology for human-like movement. To quantify morphological fidelity, we also introduce a novel metric that jointly considers kinematic retargeting fidelity to human motion and dynamic tracking performance. Our framework achieves state-of-the-art (SOTA) performance across all metrics compared to baseline humanoids (Bumi, K1, and Toddlerbot). Finally, we realize this design in Bridge, an open-source, 88cm-tall humanoid platform released alongside its control policy. We demonstrate that Bridge captures human motion data with superior fidelity, exhibiting exceptional performance across foundational locomotion, robust balance, and highly dynamic maneuvers. Videos and open-source materials: https://sites.google.com/view/bridgerobot.
Chinese Translation
开发能够利用人类行为数据的人形机器人对于通用具身智能至关重要,然而传统的开发方式仍受限于将硬件设计与全身控制相分离的解耦范式。这种方法导致系统性能欠佳,牺牲了类人的流畅性与敏捷性。为弥合这一差距,我们提出了一种数据驱动的形态-控制协同设计框架,以优化人形机器人形态、实现类人运动。为量化形态保真度,我们还提出了一种新指标,该指标联合考虑了运动学重定向对人类动作的保真度以及动态跟踪性能。与基线人形机器人(Bumi、K1 和 Toddlerbot)相比,我们的框架在所有指标上均达到了最先进(SOTA)的性能。最后,我们将该设计实现于 Bridge——一个开源的、身高 88 厘米的人形机器人平台,并连同其控制策略一并发布。我们证明 Bridge 能够以卓越的保真度捕捉人类运动数据,在基础运动、鲁棒平衡以及高动态机动动作方面均表现出色。视频与开源材料:https://sites.google.com/view/bridgerobot。
cs.RO / 13 / 2609.03561

TRaIL-Odom: Tightly Coupled Continuous Time Radar-IMU-LiDAR Odometry with Adaptive Doppler Weighting

TRaIL-Odom:具有自适应多普勒加权的紧耦合连续时间雷达-IMU-激光雷达里程计
Noh, Chiyun, Tuna, Turcan, Talbot, William, Hutter, Marco, Kneip, Laurent, Kim, Ayoung
Abstract
Existing radar-LiDAR fusion methods rely on fixed residual weights, even though the informativeness of radar Doppler and LiDAR geometry is scan- and direction-dependent, leading to uniform radar weighting that misallocates Doppler information across translational directions. To address this limitation, we propose two degeneracy-aware Doppler reweighting modules within a tightly coupled Radar-IMU-LiDAR odometry framework: per-point radar reweighting and scan-wise radar gain scheduling. Since geometric degeneracy is directional, we first identify weak translational directions from the LiDAR geometry and reweight individual radar Doppler constraints based on their alignment with the weak subspace. We further adjust the overall radar contribution using LiDAR geometric anisotropy such that radar is emphasized when LiDAR observability is poor and suppressed when LiDAR constraints are already reliable. Across 13 evaluated sequences, TRaIL-Odom achieves state-of-the-art overall performance, with clear advantages in geometrically degenerate scenes. In ablation experiments on three degenerate sequences, combining the two adaptive weighting modules reduces RMSE ATE and RTE by 86.0% and 78.5% relative to the fixed-weight baseline. We make our code and an accompanying dataset publicly available at https://github.com/ChiyunNoh/TRaIL-Odom.
Chinese Translation
现有的雷达-激光雷达融合方法依赖于固定的残差权重,然而雷达多普勒信息与激光雷达几何信息的有效性是随扫描和方向变化的,这导致了对雷达的统一加权,使多普勒信息在平移方向上的分配失当。为解决这一局限,我们在紧耦合的雷达-IMU-激光雷达里程计框架中提出了两个退化感知的多普勒重加权模块:逐点雷达重加权和扫描级雷达增益调度。由于几何退化具有方向性,我们首先从激光雷达几何中识别弱平移方向,并根据各雷达多普勒约束与弱子空间的对齐程度对其分别重加权。我们进一步利用激光雷达几何的各向异性调整雷达的整体贡献,使雷达在激光雷达可观测性差时被强化,而在激光雷达约束已经可靠时被抑制。在评估的13个序列上,TRaIL-Odom取得了总体上最先进的性能,并在几何退化场景中具有明显优势。在三个退化序列上的消融实验中,两个自适应加权模块相结合相较于固定权重基线将RMSE ATE和RTE分别降低了86.0%和78.5%。我们的代码及配套数据集已在 https://github.com/ChiyunNoh/TRaIL-Odom 公开发布。
cs.RO / 14 / 2609.03565

Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

迈向具有物理基础的JEPA世界模型以实现目标条件机器人规划
Liu, Muyuan, Huang, Yue, Liang, Zheng, Gao, Xiang
Abstract
Action-conditioned JEPA world models enable planning toward visually specified goals without reconstructing future pixels, yet latent prediction alone does not explicitly encourage the learned representations to retain information relevant to robotic control. We introduce an end-to-end JEPA world model that augments latent prediction with inverse dynamics (IDM) and state alignment (SA). While inverse dynamics discourages latent collapse and makes latent transitions informative of the actions that produced them, state alignment grounds consecutive representations in their associated physical configuration and motion. Across four benchmark tasks, our model attains the highest success rates on TwoRoom (100%), PushT (98%), and OGBench-Cube (87%), while performing comparably to LeWorldModel on Reacher. Our ablation further shows that adding state alignment consistently improves planning success over IDM alone across all four tasks. Although LeWorldModel, our primary baseline, attains higher average straightening on OGBench-Cube, transition-subspace analysis shows that its transition energy is concentrated in a substantially lower-dimensional subspace. Our state-aligned model exhibits a higher effective transition dimension than LeWorldModel and improves planning over IDM alone, supporting state alignment as an effective complement to inverse dynamics for robotic planning.
Chinese Translation
动作条件JEPA(联合嵌入预测架构)世界模型无需重建未来像素即可实现面向视觉指定目标的规划,然而仅依靠潜在预测并不能显式地促使所学表示保留与机器人控制相关的信息。我们提出了一种端到端的JEPA世界模型,在潜在预测的基础上引入了逆动力学模型(IDM)和状态对齐(SA)。逆动力学抑制潜在塌缩,并使潜在状态转移能够反映产生该转移的动作信息;状态对齐则将相邻的表示锚定于其对应的物理构型与运动。在四个基准任务上,我们的模型在TwoRoom(100%)、PushT(98%)和OGBench-Cube(87%)上取得了最高的成功率,并在Reacher任务上与LeWorldModel表现相当。消融实验进一步表明,相较于仅使用逆动力学模型,加入状态对齐在所有四个任务上均能持续提升规划成功率。尽管我们的主要基线LeWorldModel在OGBench-Cube上获得了更高的平均拉直度,但转移子空间分析显示其转移能量集中在维度明显更低的子空间中。我们的状态对齐模型表现出比LeWorldModel更高的有效转移维度,并在仅使用逆动力学模型的基础上提升了规划性能,这支持了状态对齐作为逆动力学在机器人规划中的有效补充。
cs.RO / 15 / 2609.03591

Scaling Bimanual Household Manipulation from 1,500 hours of Demonstrations to On-Policy Corrections

基于1500小时人类示范数据与在线策略纠正的可扩展双手家庭操作研究
Xu, Jiafeng, Li, Qi, Shen, Yan, Ren, Yiyu, Davies, Travis, He, Shaowen, Wang, Ze, Yang, Yifan, Cheng, Ran, Dong, Hao
Abstract
Learning generalist policies for robust bimanual manipulation is bottlenecked by the scarcity of high quality large scale human demonstration data. In this work, we release 1,500 hours of diverse bimanual manipulation demonstrations covering everyday household tasks, and use this comprehensive corpus to train XR-2, a powerful vision-language-action (VLA) model. Enabled by a purpose built high throughput data pipeline and a carefully designed multi stage training paradigm, XR-2 attains strong manipulation performance in our systematic experiments while retaining favorable training efficiency and high data utilization. We further study two critical scaling axes: varying the amount of expert demonstration data, and post training on DAgger correction data from real time human interventions. In both settings, task success rate improves steadily over the data ranges we probe, exhibiting a clear consistent scaling trend at our current data scale. These results validate both the learning capacity of XR-2 and the promising scaling properties of the released dataset, which we open source to support reproducible research on bimanual robot manipulation learning.
Chinese Translation
学习具备鲁棒双手操作的通用策略,其瓶颈在于高质量大规模人类示范数据的稀缺。在本工作中,我们发布了1500小时涵盖日常家庭任务的多样化双手操作示范数据,并利用这一全面的数据语料库训练了XR-2——一个强大的视觉-语言-动作(VLA)模型。借助专门构建的高吞吐量数据流水线和精心设计的多阶段训练范式,XR-2在我们的系统性实验中展现出强大的操作性能,同时保持了良好的训练效率和较高的数据利用率。我们进一步研究了两个关键的扩展维度:改变专家示范数据量,以及利用实时人工干预产生的DAgger纠正数据进行后训练。在这两种设置中,任务成功率在我们所探测的数据范围内均稳步提升,展现出当前数据规模下清晰一致的扩展趋势。这些结果验证了XR-2的学习能力以及所发布数据集的可扩展潜力。我们将该数据集开源,以支持双手机器人操作学习的可复现研究。
cs.RO / 16 / 2609.03611

FailBench: How Reliable are VLMs at Judging Robot Task Success?

FailBench:视觉语言模型评估机器人任务成功的可靠性如何?
Navasardyan, Zaruhi, Danielyan, Tatul, Davtyan, Hrant
Abstract
Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but existing benchmarks offer limited evidence of cross-domain generalization. We introduce FailBench, a benchmark for robot failure detection comprising 2,197 manipulation attempts across 14 public sources (12 real-world, 2 simulated). In FailBench, 75% of failures occur naturally, and six real-world sources come from non-failure-detection datasets. Evaluating 13 VLM-based detectors, we find the best model achieves only 0.77 mean balanced accuracy. Notably, models fine-tuned for failure detection consistently underperform general-purpose VLMs and their own pretrained baselines. Performance depends heavily on required visual evidence: models approach saturation when outcomes depend on observable object motion, but degrade to near-chance (<0.60 balanced accuracy) on contact-intensive assembly tasks. Error analysis reveals a systematic bias toward predicting success under ambiguous evidence, which persists even with increased reasoning effort. Finally, we show that input-level intervention--spatially localizing and cropping outcome-relevant regions--improves the top detector by 2.4 percentage points without extra training.
Chinese Translation
视觉语言模型(VLM)正日益被用于评估机器人操作结果,但现有基准测试在跨领域泛化能力方面提供的证据有限。我们提出了FailBench,一个用于机器人失败检测的基准,包含来自14个公开数据源(12个真实世界、2个仿真环境)的2,197次操作尝试。在FailBench中,75%的失败是自然发生的,且六个真实世界数据源来自非失败检测数据集。我们评估了13个基于VLM的检测器,发现最佳模型的平均平衡准确率仅为0.77。值得注意的是,针对失败检测进行微调的模型始终表现不如通用VLM以及其自身的预训练基线模型。性能在很大程度上取决于所需的视觉证据:当结果取决于可观察到的物体运动时,模型接近饱和;但在接触密集的装配任务上,性能下降至接近随机水平(平衡准确率<0.60)。错误分析揭示了一种系统性偏差,即在证据模糊时倾向于预测成功,即使增加推理力度,这种偏差依然存在。最后,我们展示了输入层面的干预——对结果相关区域进行空间定位和裁剪——能够在无需额外训练的情况下将最佳检测器的性能提升2.4个百分点。
cs.RO / 17 / 2609.03623

QLAUN: A Research-Oriented, Robust, Agile, Modular, and Affordable Torque-Controlled Quadruped Robot

QLAUN:一款面向研究、鲁棒、敏捷、模块化且经济实惠的力控四足机器人
Moudallal, Mohamad S., Maalouf, Noel J.
Abstract
QLAUN Bot (Quad-Legged Adaptive Unmanned Navigator Robot) is a torque-controlled quadruped robot that is research-oriented, cost-effective, and aimed at achieving simultaneous robustness and agility while being completely 3D-printed. It is a quadruped robot that is aimed at empowering robotics research at universities and research institutes in Lebanon and the MENA region. Using a novel electronics-free leg design strategy, we present a modular robot with interchangeable and easily replaceable legs. The 15 kg robot possesses 12 DoF (Degrees-of-Freedom) with three per leg, each paired with a completely 3D-printed Quasi-Direct Drive (QDD) actuator that consists of a brushless DC motor and a low-ratio gearbox transmission that is connected to a belt transmission system for significantly increasing the torque outputs at the joints. We present legs that have decoupled hip and knee actuators to improve the overall modularity of the robot. QLAUN is almost completely 3D-printed using polylactic acid (PLA) and assembled using off-the-shelf parts to create a robust, agile, and affordable robot for legged robot locomotion research. The legs possess joints with wide ranges of motion, including a continuous hip flexion-extension joint. A compliant foot, printed using TPU-95A is also implemented for alleviating hard impacts and handling terrain uncertainties. This extended abstract aims to introduce QLAUN, a novel platform for robotics research, emphasizing the design concepts and principles that underpin its development to the academic and research communities in the field of robotics.
Chinese Translation
QLAUN Bot(四足自适应无人导航机器人,Quad-Legged Adaptive Unmanned Navigator Robot)是一款力控四足机器人,具有面向研究、成本效益高的特点,旨在同时实现鲁棒性与敏捷性,且整机完全采用3D打印制造。该四足机器人旨在赋能黎巴嫩及中东北非(MENA)地区高校和科研机构的机器人研究。通过一种新颖的免电子元件腿部设计策略,我们提出了一种具有可互换且易于更换腿部的模块化机器人。该机器人重15公斤,具有12个自由度(DoF),每条腿3个自由度,每条腿均配有一个完全3D打印的准直驱(Quasi-Direct Drive, QDD)执行器,该执行器由无刷直流电机和低减速比齿轮箱组成,并连接到皮带传动系统,以显著提高关节的输出扭矩。我们提出的腿部设计采用髋关节和膝关节执行器解耦的方案,以提升机器人的整体模块化程度。QLAUN几乎完全使用聚乳酸(PLA)3D打印制造,并采用现成的零部件组装,从而打造出一款鲁棒、敏捷且经济实惠的足式机器人运动研究平台。其腿部关节具有较大的运动范围,包括一个可连续运动的髋关节屈伸关节。此外,机器人还配备了使用TPU-95A打印的柔性足部,用于缓解硬冲击并应对地形的不确定性。本扩展摘要旨在向学术界介绍QLAUN这一新型机器人研究平台,重点阐述其研发过程中所依托的设计概念与原理。
cs.RO / 18 / 2609.03630

Local Path Planning and Obstacle Avoidance for an Omnicopter Platform

全向无人机平台的局部路径规划与避障
Helinski, Mikolaj, Theodoulis, Spilios, Hamandi, Mahmoud, Ali, Abdullah Mohamed, Tzes, Anthony, Popovic, Marija
Abstract
Autonomous unmanned aerial vehicles (UAVs) increasingly operate in cluttered environments where global planners such as RRT* are not directly deployable at control rates. This paper presents a real-time local planning and obstacle avoidance module for an omnidirectional multirotor (omnicopter) by extending the Dynamic Window Approach to six degrees of freedom (6D-DWA). Our method achieves real-time feasibility through (i) local-map voxelisation, (ii) a compact sphere-based approximation of the vehicle geometry, and (iii) adaptive velocity sampling in the 6D search space. To improve reactivity to unknown obstacles, we introduce a context-aware "Agile Mode" that adjusts scoring weights online to trade-off between goal progress, clearance, and heading/facing constraints during evasive manoeuvres. We evaluate our approach in simulation across computational stress tests, dense-waypoint path tracking, and static/unknown obstacle scenarios. Our planner runs consistently within a 0.2s control loop, tracks waypoint-dense global paths with < 0.1m average cross-track error and 13deg average heading error, and avoids collisions in static environments. For unknown obstacle avoidance, Agile Mode achieves 79.3% success for an off-centre obstacle and 41.4% for a centred obstacle, highlighting both the effectiveness of adaptive weighting and remaining limitations in highly constrained geometries.
Chinese Translation
自主无人机(UAV)日益在杂乱环境中运行,而RRT*等全局规划器难以直接以控制频率部署。本文通过将动态窗口法(Dynamic Window Approach)扩展至六自由度(6D-DWA),提出了一种面向全向多旋翼飞行器(omnicopter)的实时局部规划与避障模块。我们的方法通过以下方式实现实时可行性:(i) 局部地图体素化;(ii) 采用紧凑的基于球体的飞行器几何近似;(iii) 在六维搜索空间中进行自适应速度采样。为提高对未知障碍物的反应能力,我们引入了一种情境感知的“敏捷模式”(Agile Mode),可在规避机动过程中在线调整评分权重,以在目标推进、安全间隙和航向/朝向约束之间进行权衡。我们在仿真中通过计算压力测试、密集航点路径跟踪以及静态/未知障碍物场景对该方法进行了评估。我们的规划器始终能在0.2秒的控制周期内运行,跟踪航点密集的全局路径时平均横向误差小于0.1米、平均航向误差为13度,并能在静态环境中避免碰撞。在未知障碍物规避方面,敏捷模式对偏心障碍物的成功率为79.3%,对居中障碍物的成功率为41.4%,这既体现了自适应加权的有效性,也揭示了在高度受限几何构型下仍存在的局限性。
cs.RO / 19 / 2609.03681

WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models

WISE:面向视觉-语言-动作模型高效后训练的世界模型引导想象调度方法
Zhang, Chenhao, Zhao, Hanyu, Cheng, Hang, Pan, Tengfei, Zeng, Long
Abstract
Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinforcement learning with expensive and potentially unstable real-world exploration. World models offer a promising alternative by evaluating candidate behaviors through imagined futures, yet effective post-training requires more than accurate prediction: imagination must be scheduled where it is useful, bounded within reliable horizons, and translated into trustworthy policy supervision. In robotic manipulation, the value of imagination varies substantially across execution stages, while extended rollouts can accumulate prediction errors and introduce unreliable learning signals. We introduce WISE (World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models), a unified framework that coordinates when and how world-model imagination is used during policy refinement. WISE selectively invokes imagination at interaction-relevant states, performs bounded multi-view rollouts, evaluates candidate futures using progress and completion signals, and uses their relative outcomes to refine actions generated from real interaction contexts. Extensive experiments with both $\pi_0$ and $\pi_{0.5}$ demonstrate consistent improvements across diverse manipulation tasks while reducing GPU computation time by approximately 80% compared with full imagination. Real-world evaluations further show substantial gains in robustness and generalization under diverse real-world distribution shifts.
Chinese Translation
视觉-语言-动作(VLA)策略的后训练通常依赖于需要昂贵专家演示的监督微调,或需要昂贵且可能不稳定的真实世界探索的强化学习。世界模型通过想象未来情景来评估候选行为,为这一难题提供了一条有前景的替代路径。然而,有效的后训练不仅仅需要准确的预测:想象必须在有用的地方被调度、限制在可靠的时域范围内,并被转化为可信的策略监督信号。在机器人操作任务中,想象的价值在不同执行阶段之间存在显著差异,而过长的 rollout 会累积预测误差并引入不可靠的学习信号。我们提出了 WISE(World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models),一个统一框架,用于协调在策略优化过程中何时以及如何使用世界模型想象。WISE 在与交互相关的状态处选择性地调用想象,执行有界的多视角 rollout,利用进度与完成度信号评估候选未来情景,并使用其相对结果来优化由真实交互上下文生成的动作。在 $\pi_0$ 和 $\pi_{0.5}$ 上的大量实验表明,该方法在多样化操作任务上取得了一致的性能提升,同时与完全想象相比将 GPU 计算时间减少约 80%。真实世界评估进一步表明,在多种真实世界分布偏移下,该方法在鲁棒性和泛化能力方面均有显著提升。
cs.RO / 20 / 2609.03699

Predictive Zonotope Reduction: Precise Runtime Monitoring under Uncertainty

预测性Zonotope约简:不确定性下的精确运行时监控
Krsmanovic, Vladimir, Kohn, Florian, Finkbeiner, Bernd, Simovic, Milan
Abstract
Robots operating in physical environments make control decisions based on uncertain sensor measurements, which can lead to unsafe or suboptimal actions. Runtime monitors that check their behavior against safety specifications must represent this uncertainty soundly. Zonotopes are a widely used representation, but continuously incorporating new measurements grows their order unboundedly, so monitors must periodically apply an over-approximating reduction. The choice of the reduction method substantially affects the zonotope's precision, yet existing approaches typically utilize a fixed method throughout the run, even though the optimal choice depends on the current state. This paper presents a Predictive Zonotope Reduction (PZR) approach, which frames reducer selection as an optimal control problem and solves it using beam-search model predictive control. Policy distillation into a small neural policy further provides substantially higher execution speed than model predictive control while maintaining improved performance, enabling uncertainty-aware runtime monitoring on resource-constrained real-time systems. We implement our approach in the RLola runtime monitoring framework and evaluate it on a 5-degree-of-freedom robotic arm simulated in MuJoCo, with sensor uncertainty modeled according to ISO 5725. Experiments on a Raspberry Pi 5 show that dynamic reduction significantly lowers false-positive rates in monitoring compared with static reduction strategies.
Chinese Translation
在物理环境中运行的机器人基于不确定的传感器测量做出控制决策,这可能导致不安全或次优的动作。对其行为进行安全规范检查的运行时监控器必须可靠地表示这种不确定性。Zonotope(超平行体)是一种广泛使用的表示方法,但持续纳入新的测量值会使阶数无限增长,因此监控器必须周期性地应用过近似的约简。约简方法的选择对zonotope的精度有重大影响,然而现有方法通常在整个运行过程中使用固定的方法,尽管最优选择取决于当前状态。本文提出了一种预测性Zonotope约简(Predictive Zonotope Reduction, PZR)方法,该方法将约简器的选择表述为一个最优控制问题,并使用束搜索模型预测控制(beam-search model predictive control)进行求解。通过策略蒸馏到一个小型神经策略,进一步获得了比模型预测控制显著更高的执行速度,同时保持性能提升,从而能够在资源受限的实时系统上实现不确定性感知的运行时监控。我们在RLola运行时监控框架中实现了该方法,并在MuJoCo中仿真的5自由度机械臂上进行了评估,其中传感器不确定性按照ISO 5725标准建模。在Raspberry Pi 5上的实验表明,与静态约简策略相比,动态约简显著降低了监控中的假阳性率。
cs.RO / 21 / 2609.03704

Toward an~Integrated Cognitive--Ergonomic Architecture for~Human--Machine Interaction: Combining Cognitive Models with~Human Factors Ergonomics

面向人机交互的认知-工效学一体化架构:将认知模型与人因工效学相结合
Lenat, Antoine, Cheminat, Olivier, Chablat, Damien, Charron, Camilo
Abstract
This paper presents an integrated approach to modeling human competencies by combining the theoretical foundations of cognitive architectures with principles from Human Factors Ergonomics (HFE). Through a comparative analysis of established cognitive models-SOAR, ACT-R, LIDA, and COCOM-we synthesize a tailored architecture designed to address the complexities of human-machine interaction (HMI) in dynamic environments. By contextualizing this model within ergonomic frameworks, we elucidate the mechanisms underlying decision-making, skill acquisition, and adaptive behavior, bridging the gap between cognitive theory and applied system design. Our framework is empirically grounded in industrial robotics applications, where operator expertise, normative knowledge, and real-time feedback loops are critical. The proposed architecture not only enhances the cognitive alignment of HMI systems but also provides a scalable methodology for designing intelligent, human-centered interfaces in high-stakes environments. This work advances both the theoretical understanding of human competencies and the practical implementation of adaptive, ergonomically optimized systems.
Chinese Translation
本文提出了一种通过将认知架构的理论基础与人因工效学(HFE)的原则相结合来建模人类能力的集成方法。通过对成熟认知模型——SOAR、ACT-R、LIDA 和 COCOM——的对比分析,我们综合出一种定制的架构,旨在应对动态环境中人机交互(HMI)的复杂性。通过将该模型置于工效学框架中进行情境化分析,我们阐明了决策、技能习得和适应性行为的内在机制,弥合了认知理论与应用系统设计之间的鸿沟。我们的框架以工业机器人应用为实证基础,在这些应用中,操作员的专业知识、规范性知识和实时反馈回路至关重要。所提出的架构不仅增强了 HMI 系统的认知一致性,还为在高风险环境中设计智能化、以人为中心的界面提供了一种可扩展的方法论。这项工作既推进了对人类能力的理论理解,也推动了适应性、工效学优化系统的实际落地。
cs.RO / 22 / 2609.03715

MINERVA: How Small Can a Manipulation Policy Be and Still Solve LIBERO?

MINERVA:操作策略最小可以小到什么程度并仍能解决 LIBERO?
Sendai, Kohei, Matsushima, Tatsuya, Iwasawa, Yusuke
Abstract
Vision-language-action (VLA) models with billions of parameters now dominate the LIBERO manipulation benchmark, but the model capacity actually required by the benchmark remains unclear. We introduce MINERVA (MINimal Efficient Robotic Vision-Action policy), a family of deliberately compact visuomotor policies designed to measure this task-specific capacity floor. A 0.54M-parameter policy achieves 95.1% average success over 2,000 rollouts on the four standard LIBERO suites, only 2.4 points below the reported LeRobot $\pi_{0.5}$ result despite using 7,700$\times$ fewer parameters. Performance saturates near 1M parameters and collapses below 0.25M. Across broad architectural, training, and inference sweeps, only action-chunk length and vision capacity consistently exceed a $\pm$1-point training-seed band. Flow matching provides no detectable advantage over direct L1 regression across three seeds, while regression is up to 3.8$\times$ faster on GPU. A task-ID permutation probe shows that standard LIBERO instruction conditioning primarily selects among memorized tasks: changing only the task-ID mapping reduces success to near chance. The same recipe achieves 94.6% success across 89 LIBERO-90 tasks, while LIBERO-Plus perturbations reduce performance to 46--56%, with near-zero robustness to photometric shifts. The 0.54M policy replans every control step in 5--9 ms per chunk on a laptop CPU, 113$\times$ faster than SmolVLA and 1,400$\times$ faster than $\pi_{0.5}$, without a GPU. These results establish a first empirical estimate of LIBERO's task-specific capacity floor and motivate capacity-aware design and distillation for deployment-efficient robot policies.
Chinese Translation
拥有数十亿参数的视觉-语言-动作(VLA)模型目前主导着 LIBERO 操作基准测试,但该基准实际所需的模型容量仍不清楚。我们提出了 MINERVA(MINimal Efficient Robotic Vision-Action policy,极简高效机器人视觉-动作策略),这是一系列刻意保持紧凑的视觉运动策略,旨在测量该任务特定的容量下限。一个仅含 0.54M 参数的策略在四个标准 LIBERO 测试套件的 2,000 次 rollout 中达到了 95.1% 的平均成功率,仅比已报告的 LeRobot π₀.₅ 结果低 2.4 个百分点,而参数量少了 7,700 倍。性能在约 1M 参数附近趋于饱和,低于 0.25M 时则崩溃。在广泛的架构、训练和推理扫描中,只有动作块长度和视觉容量始终超出 ±1 个百分点的训练随机种子波动范围。在三个随机种子上,流匹配(flow matching)相比直接的 L1 回归未显示出任何可检测的优势,而回归在 GPU 上的速度快达 3.8 倍。任务 ID 置换探测实验表明,标准 LIBERO 的指令条件化主要是在记忆的任务中进行选择:仅改变任务 ID 的映射关系即可使成功率降至接近随机水平。同样的方法在 89 个 LIBERO-90 任务上达到了 94.6% 的成功率,而 LIBERO-Plus 扰动使性能降至 46–56%,且对光度变化的鲁棒性几乎为零。该 0.54M 参数策略在笔记本电脑 CPU 上每个动作块仅需 5–9 毫秒即可完成每个控制步的重规划,比 SmolVLA 快 113 倍,比 π₀.₅ 快 1,400 倍,且无需 GPU。这些结果首次给出了 LIBERO 任务特定容量下限的经验估计,并为面向部署高效的机器人策略的容量感知设计与蒸馏提供了依据。
cs.RO / 23 / 2609.03720

RoughSense: Lightweight Terrain-Induced Rover Vibration Prediction Using Point Clouds and IMU Feedback

RoughSense:基于点云与IMU反馈的轻量级地形诱导探测车振动预测方法
Garcia, Gabriel Manuel, Aravecchia, Stephanie, Olivares-Mendez, Miguel Angel
Abstract
Autonomous navigation in space requires reliable terrain assessment for safe operations, especially in underground environments with limited communication, computing resources, and power budget. This paper presents a lightweight method for real-time vibration-aware traversability mapping using a Light Detecting And Ranging (LiDAR) point cloud and Inertial Measurement Unit (IMU) measurements. An initial vibration proxy is estimated from terrain geometry by applying Random sample consensus (RANSAC) to local point-cloud patches produced by a Simultaneous Localisation And Mapping (SLAM) algorithm. In parallel, the IMU provides local observations of the vibration experienced by the rover during traversal. The point-cloud-based prediction is then corrected online using Recursive Least Squares, allowing the system to adapt the geometric estimate to the measured rover response. The approach is evaluated in a lunar analogue environment, an outdoor field, and an underground mine.
Chinese Translation
太空中的自主导航需要可靠的地形评估以确保安全运行,尤其是在通信受限、计算资源有限且功耗预算紧张的地下环境中。本文提出了一种轻量级的实时振动感知可通行性建图方法,该方法利用激光雷达(LiDAR)点云和惯性测量单元(IMU)测量数据。通过对同步定位与建图(SLAM)算法生成的局部点云补丁应用随机抽样一致性(RANSAC)算法,从地形几何特性中估计初始振动代理值。同时,IMU提供探测车在行进过程中所经历振动的局部观测。随后,利用递归最小二乘法对基于点云的预测进行在线校正,使系统能够根据探测车的实际响应调整几何估计。该方法在月球类比环境、室外场地和地下矿井中进行了评估。
cs.RO / 24 / 2609.03758

A Multi-Vine Soft Robot Enabling Accessible Working Channel and Steering

一种可实现可及工作通道与转向功能的多藤蔓软体机器人
Kashef, Reza, Suulker, Cem, Sofla, Mohammad Sheikh, Althoefer, Kaspar
Abstract
Soft eversion robots, also known as vine robots, have attracted growing interest for navigation and inspection tasks, including minimally invasive medical applications [1]. A vine robot consists of a thin, flexible, inextensible tube folded inward that everts and grows forward when pressurized. This tip-growth enables navigation with minimal friction, making vine robots well suited for complex environments such as the human colon [2]. While their inherent softness allows passive conforma- tion to curved pathways in confined spaces, navigation performance strongly depends on environmental inter- actions, including contact angle and the length of un- constrained deployed material [3], [4]. Sharp directional changes, such as those in the sigmoid colon, often limit passive growth and necessitate active steering. Existing solutions include distributed artificial muscles [5] or dedicated tip-based steering mechanisms [6]. In addition, many applications require payload delivery, such as sensors and tools [7], [8]. Within the ERC Synergy project EndoTheranostics, this motivates the development of vine robots capable of delivering micro- surgical tools during growth. Prior work has integrated working channels within the vine body [8], [9], but these approaches constrain tool size, introduce friction, and limit access to the environment to the robot tip. In this work, we propose a multi-vine architecture in which two vine robots are coupled to an externally integrated working channel via soft mounting tips [10]. Independent vine actuation enables active tip steering while advancing the working channel without embed- ding it within the vine bodies Figure 1. Experiments demonstrate sharp steering of nearly 90 degrees during growth, highlighting the potential of this architecture for versatile medical and non-medical applications.
Chinese Translation
软体翻转生长机器人,也称为藤蔓机器人(vine robots),在导航与检测任务(包括微创医疗应用)中受到越来越多的关注[1]。藤蔓机器人由一根细而柔韧、不可伸展的管子构成,该管子向内折叠,在加压时发生外翻并向前生长。这种尖端生长方式使其能够以极小的摩擦进行导航,因此藤蔓机器人非常适合诸如人体结肠等复杂环境[2]。尽管其固有的柔软性使其能够在受限空间内被动顺应弯曲路径,但其导航性能在很大程度上依赖于环境交互,包括接触角度以及未受约束的展开材料的长度[3][4]。急剧的方向变化(如乙状结肠中的转弯)常常限制被动生长,因而需要主动转向。现有解决方案包括分布式人工肌肉[5]或专用的尖端转向机构[6]。此外,许多应用需要载荷输送,例如传感器和工具[7][8]。在ERC Synergy项目EndoTheranostics的支持下,这促使我们开发能够在生长过程中输送显微外科工具的藤蔓机器人。先前的工作已将工作通道集成于藤蔓本体之内[8][9],但这些方法限制了工具尺寸,引入了摩擦,并且仅能通过机器人尖端接触环境。在本工作中,我们提出了一种多藤蔓架构,其中两个藤蔓机器人通过软性安装尖端[10]与一个外部集成的工作通道相耦合。独立的藤蔓驱动实现了主动尖端转向,同时在推进工作通道的过程中无需将其嵌入藤蔓本体之内(见图1)。实验表明,该机器人在生长过程中可实现接近90度的急剧转向,凸显了该架构在多样化医疗及非医疗应用中的潜力。
cs.RO / 25 / 2609.03760

Virtual Testing of Automated Driving Systems through Credible Simulations

基于可信仿真的自动驾驶系统虚拟测试
Dona, Riccardo, Rusciano, Espedito, Ciuffo, Biagio
Abstract
Simulation is increasingly used to support safety-related decision-making in road transport, particularly for the assessment and approval of automated driving systems (ADS). The complexity of ADS behavior and size of their operational design domains make exclusive reliance on physical testing impractical, leading to extensive use of virtual testing (VT) during the approval phase. This shift raises critical questions regarding the credibility of modelling and simulation (M&S) results used to support road safety decisions. Current VT accreditation approaches in the ADS domain typically rely on validation-only practices, which have been shown to scale poorly when applied to complex, multi-tool simulation environments. To address this limitation, this paper proposes a risk-based framework for assessing the credibility of simulation toolchains used in ADS safety evaluation, drawing inspiration from established practices in other safety-critical domains, notably NASA's STD-7009 for models and simulations. The framework extends traditional verification and validation (V&V) by explicitly linking credibility requirements to the intended use of simulation outputs and to the safety criticality of the decisions they support within the approval process. It provides a lifecycle-oriented assessment scheme integrating toolchain management, modelling assumptions and limitations, verification, validation, and sensitivity analysis. Credibility acceptance thresholds are defined proportionally, allowing differentiated requirements depending on whether simulation is used for exploratory safety analysis, partial decision support, or as a substitute for physical testing. While demonstrated for ADS, the proposed approach is directly applicable to road safety and simulation studies where VT plays a central role in safety assessment and regulatory decision-making.
Chinese Translation
仿真在道路运输安全相关决策支持中的应用日益广泛,尤其是在自动驾驶系统(ADS)的评估与审批方面。由于ADS行为的复杂性及其运行设计域的规模庞大,完全依赖实物测试并不切实际,因此虚拟测试(VT)在审批阶段得到了广泛应用。这一转变引发了关于用于支持道路安全决策的建模与仿真(M&S)结果可信度的关键问题。目前ADS领域的VT认可方法通常仅依赖验证(validation)实践,而这种方法在应用于复杂的多工具仿真环境时已被证明扩展性较差。针对这一局限性,本文借鉴其他安全关键领域的成熟实践,特别是NASA针对模型与仿真的STD-7009标准,提出了一个基于风险的框架,用于评估ADS安全评估中所用仿真工具链的可信度。该框架对传统的验证与确认(V&V)进行了扩展,将可信度要求与仿真输出的预期用途以及其在审批流程中所支持决策的安全关键性显式关联。该框架提供了一个面向生命周期的评估方案,整合了工具链管理、建模假设与局限、验证、确认以及敏感性分析。可信度接受阈值按比例定义,从而可根据仿真是用于探索性安全分析、部分决策支持,还是作为实物测试的替代手段,提出差异化的要求。尽管本文以ADS为例进行了演示,但所提出的方法可直接适用于虚拟测试在安全评估与监管决策中发挥核心作用的道路安全与仿真研究。
cs.RO / 26 / 2609.03761

Robot Aware Computational Design of Object Specific Passive Grippers for Additive Manufacturing

面向增材制造的物体专用被动夹爪的机器人感知计算设计
Omaisan, Abdullah Yahya Abdullah, Mohamed, Ibrahim Sheikh
Abstract
This paper presents an end-to-end computational pipeline that converts a selected object mesh, a measured object state, and a selected six-axis robot into an object-specific, unactuated, additively manufacturable gripper. The method couples exact-mesh RGB-D/ICP pose registration, deterministic surface-contact sampling, uncertainty-aware wrench screening, selection among six passive capture mechanisms, object-conformal surface synthesis, full-orientation robot inverse kinematics, a swept-volume-aware manufacturing domain, directional fused-deposition finite-element screening, and constrained three-dimensional SIMP topology optimization. Unlike workflows that treat grasp selection, tool geometry, motion, and structural design as separate problems, every exported design is bound to the source mesh, object pose, robot flange, contact set, and insertion hypothesis by a traceable design identifier. We derive the implemented registration, contact, fit-tolerance, finite-element, and density-optimization equations and prove three properties of the numerical construction: nodal load preservation, monotonic compliance sensitivity under SIMP interpolation, and voxel-domain containment after topology post-processing. Four archived object-specific attempts - a rabbit, camera flange, 3DBenchy, and faceted bust - meet the nominal fit, uncertain-wrench, runtime-sweep, and baseline/post-topology FEA gates. A deliberately enlarged +/-3 mm, +/-5 degree pose stress check differentiates the designs, retaining 29-134 of 160 simulated trials. Their reconstructed topologies retain 92.0-97.6% of the FE domain because functional regions are protected. Archived robot photographs show the corresponding printed assemblies qualitatively, while nominal material properties and absent coupon-calibrated, instrumented tests keep all four at digital-screening status rather than operational release.
Chinese Translation
本文提出了一种端到端的计算流程,可将选定的物体网格、测量得到的物体状态和选定的六轴机器人转化为物体专用、无驱动、可增材制造的夹爪。该方法集成了基于精确网格的RGB-D/ICP位姿配准、确定性的表面接触采样、考虑不确定性的旋量筛选、六种被动抓取机构的选择、物体保形表面合成、全姿态机器人逆运动学、扫掠体积感知的制造域、方向性熔融沉积有限元筛选,以及带约束的三维SIMP拓扑优化。与将抓取选择、工具几何、运动和结构设计视为独立问题的传统流程不同,每个导出的设计均通过可追溯的设计标识符与源网格、物体位姿、机器人法兰、接触集和插入假设绑定。我们推导了所实现的配准、接触、配合公差、有限元和密度优化方程,并证明了该数值构造的三个性质:节点载荷保持性、SIMP插值下柔度敏感度的单调性,以及拓扑后处理后的体素域包含性。四个已存档的物体专用尝试——兔子模型、相机法兰、3DBenchy和多面体半身像——均通过了名义配合、不确定性旋量、运行时扫掠以及基线/拓扑后有限元分析(FEA)门槛。一项刻意放大的±3毫米、±5度位姿压力测试对设计进行了区分,在160次仿真试验中保留了29至134次。由于功能区域受到保护,其重构拓扑保留了有限元域的92.0%至97.6%。存档的机器人照片定性地展示了相应的打印装配体,但由于采用名义材料属性且缺乏经标准试件校准的仪器化测试,这四个设计仍处于数字筛选阶段,而非可操作发布状态。
cs.RO / 27 / 2609.03794

A comparative study on the accuracy & repeatability of mobile robotic platforms for the delivery of precision NDE measurement

用于执行精密无损检测测量的移动机器人平台精度与重复性的对比研究
Pour, SeyedMohammadAmin Nabi, Pierce, S. Gareth, Vithanage, Randika, Mohseni, Ehsan, Carswell, David, Shields, Matthew
Abstract
Mobile robotic platforms offer a flexible alternative to fixed manipulators for non-destructive evaluation (NDE) of large aerospace structures, but their base-positioning accuracy and how that accuracy should inform deployment have not been assessed under a common, externally referenced protocol. This work presents a laser tracker-based evaluation workflow (ground truth approximately 6 micrometers) that measures the static and segmented trajectory positioning accuracy of five commercial mobile platforms (KUKA KMP-1500, KUKA KMR, MiR250, Boston Dynamics Spot, Clearpath Husky) under a common protocol. A coupled multi-corner calibration recovers the laser-to-robot transformation and reflector offsets; ordinary least squares over all poses is used, with robust estimation retained only as a blunder check. Static positioning accuracy ranged from a median of 8.2 mm (KMP-1500) to 63.5 mm (Spot), with the wheel-odometry-only Husky uncalibratable. Dynamic path following was characterised by cross-track error; the component was insensitive to temporal alignment, which ranged from 6.9 mm (KMP-1500) to 112.1 mm (Spot). Both accuracy and calibratability tracked localisation capability, from the newest LiDAR SLAM platform to map-free visual odometry. No configuration meets the 0.2 to 1.0 mm aerospace NDE tolerance from the base alone; the results are framed as a design input that sizes the supplementary sensing each platform requires: roughly one order of magnitude for the best platform and nearly two for the worst, providing a reproducible basis for platform selection rather than a feasibility claim.
Chinese Translation
对于大型航空航天结构的无损检测(NDE),移动机器人平台为固定机械臂提供了一种灵活的替代方案,但其基座定位精度以及该精度应如何指导部署,尚未在统一的外部基准协议下得到评估。本工作提出了一种基于激光跟踪仪的评估流程(基准真值约6微米),在统一协议下测量了五种商用移动平台(KUKA KMP-1500、KUKA KMR、MiR250、Boston Dynamics Spot、Clearpath Husky)的静态定位精度和分段轨迹定位精度。通过耦合的多角点标定恢复激光仪到机器人的变换及反射镜偏移量;对所有位姿采用普通最小二乘法,鲁棒估计仅作为粗差检验。静态定位精度的中位数范围为8.2毫米(KMP-1500)至63.5毫米(Spot),仅依靠轮式里程计的Husky无法进行标定。动态路径跟踪以横向偏差表征;该分量对时间对齐不敏感,时间对齐范围从6.9毫米(KMP-1500)到112.1毫米(Spot)。精度和可标定性与定位能力相关,从最新的LiDAR SLAM平台到无地图视觉里程计依次递减。没有任何配置能仅凭基座满足0.2至1.0毫米的航空航天无损检测公差要求;这些结果被定位为设计输入,用于确定每个平台所需的辅助传感规模:最佳平台约需一个数量级的补充,最差平台接近两个数量级,从而为平台选择提供了可复现的依据,而非可行性声明。
cs.RO / 28 / 2609.03889

FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation

FWBC-VLA:面向接触密集型移动操作的力感知全身补偿
Zhang, Yutian, Ma, Siyuan, Yang, Liwen, Li, Yang, Hao, Ce, Chi, Haozhen, We, Dong, Yu, Qiaojun, Hou, Dibo
Abstract
Contact-rich loco-manipulation requires a bridge between semantic action generation and physical interaction control. Existing Vision-language-action (VLA) models generate task-level actions from visual and linguistic observations, but cannot interpret the physical interactions induced by those actions. While the whole-body control (WBC) policy can stabilize the robot, it cannot distinguish task-relevant interaction forces from forces induced by external disturbances during manipulation. Although force/torque sensors provide direct measurements of physical interactions, retrofitting them entails additional hardware costs and substantial integration effort, particularly for platforms not designed with sensor integration in mind. To address this problem, we propose FWBC-VLA, a force-aware framework that bridges task-level VLA action generation and low-level whole-body compensation control for wheeled-legged robots. First, we introduce HSR-Force, a sensorless residual-torque estimator for inferring contact strength and its temporal variation. These contact estimates are then encoded as tokens and injected into the VLA action expert during action decoding, enabling the policy to perceive contact onset, sustained loading, and release. For loco-manipulation tasks, all parameters of the pretrained VLA backbone are fine-tuned on our WL\&Arm Dataset, which comprises more than 5,000 episodes. Moreover, the robot's proprioceptive state, the Jacobian-derived body-frame force estimate, and the estimated contact state are jointly fed into a compensation generator to produce corrective actions. The manipulation-centric actions are subsequently combined with the corrective actions and passed to the WBC policy for execution. Real-world experiments on whiteboard wiping and door opening with a door closer demonstrate the effectiveness of our FWBC-VLA in contact-rich loco-manipulation.
Chinese Translation
接触密集型的移动操作(loco-manipulation)需要在语义动作生成与物理交互控制之间建立桥梁。现有的视觉-语言-动作(Vision-Language-Action, VLA)模型能够根据视觉和语言观测生成任务级动作,但无法解释这些动作所引发的物理交互。尽管全身控制(Whole-Body Control, WBC)策略能够稳定机器人,但其无法在操作过程中区分任务相关的交互力与外部扰动引起的力。虽然力/力矩传感器能够直接测量物理交互,但加装此类传感器会带来额外的硬件成本和大量的集成工作,对于在设计时未考虑传感器集成的平台尤为如此。为解决这一问题,我们提出了FWBC-VLA,这是一个力感知框架,为轮腿式机器人连接了任务级的VLA动作生成与低层全身补偿控制。首先,我们引入了HSR-Force,一种无传感器的残余力矩估计器,用于推断接触强度及其随时间的变化。随后,这些接触估计值被编码为token,并在动作解码过程中注入VLA动作专家,使策略能够感知接触的发生、持续加载与释放。对于移动操作任务,我们在包含5000多个回合的WL&Arm数据集上对预训练VLA骨干网络的所有参数进行了微调。此外,机器人的本体感知状态、由雅可比矩阵导出的机体坐标系力估计值以及估计的接触状态被共同输入补偿生成器,以产生矫正动作。随后,以操作为中心的动作与矫正动作相结合,并传递给WBC策略执行。在擦白板和开带闭门器的门等真实世界实验中,FWBC-VLA在接触密集型移动操作中展现了有效性。
cs.RO / 29 / 2609.03891

A hybrid pipeline for dynamic ontology-based semantic mapping

一种基于动态本体的语义映射混合流水线
Dimitropoulos, Konstantinos, Hatzilygeroudis, Ioannis
Abstract
Semantic mapping plays a crucial role in the ability of a robot to interact with objects, operate and navigate a complex environment. The most common pipeline for semantic mapping consists of geometric mapping and localization (SLAM), perception, semantic fusion and semantic representation. However, more recent works also integrate a form of prior knowledge in their application, most notably knowledge graphs or semantic scene graphs, to improve contextual understanding of the environment. In this paper, we present a hybrid pipeline for semantic mapping. Our system incorporates an external calibrated camera using homography projection for geometric mapping and localization, combined with object detection, persistent object tracking and ontology driven semantic updates to build a dynamic semantic world model. Linear regression models are also used for correction of the estimated values of real world coordinates. The system continuously updates object instances, spatial properties and semantic relations based on real time sensory data. Ontologies are selected as form of knowledge representation due to their hierarchical structure, semantic expressiveness and support for dynamic world modelling.
Chinese Translation
语义映射在机器人与物体交互、操作以及在复杂环境中导航的能力中起着至关重要的作用。最常见的语义映射流水线包括几何建图与定位(SLAM)、感知、语义融合和语义表示。然而,近期的研究工作还在其应用中集成了某种形式的先验知识,其中最突出的是知识图谱或语义场景图,以提升对环境的情境理解能力。本文提出了一种用于语义映射的混合流水线。我们的系统采用一个经过标定的外部相机,利用单应性投影实现几何建图与定位,并结合目标检测、持久目标跟踪和基于本体驱动的语义更新,以构建一个动态的语义世界模型。系统还使用线性回归模型对真实世界坐标的估计值进行校正。该系统基于实时传感数据持续更新物体实例、空间属性和语义关系。由于本体具有层次化结构、语义表达能力以及对动态世界建模的支持,本文选择本体作为知识表示形式。
cs.RO / 30 / 2609.03906

Revisiting Topological Graphs for Macro Action based Closed-loop Reinforcement Learning of Vision Language Navigation in Continuous Environment

基于宏动作的连续环境下视觉语言导航闭环强化学习的拓扑图方法再探
Ye, Shuhao, Mao, Sitong, Cui, Yuxiang, Wei, Yufei, Yu, Xuan, Zhai, Shichao, Chen, Wen, Zhou, Shunbo, Xiong, Rong, Wang, Yue
Abstract
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow natural language instructions through unseen environments. Existing imitation learning (IL) pipelines struggle in this closed-loop setting: behavior cloning suffers from distribution shift, and DAgger's expert actions become ambiguous upon trajectory deviation. While Reinforcement Learning (RL) offers a natural paradigm to address this, directly applying RL to micro action spaces is sample-inefficient due to reward sparsity. To overcome this bottleneck, we reformulate VLN-CE as a Hierarchical Markov Decision Process (MDP), explicitly decoupling high-level planning from low-level control. By abstracting the environment into a topological graph, our high-level policy operates on a macro action space of frontier nodes, with a training-free low-level controller acting as its state transition, which significantly compresses the decision horizon and makes closed-loop RL tractable. To support RL optimization on the macro MDP, we propose an action-aware value head to effectively evaluate state values under the dynamic frontier action space, powering a graph-based PPO. Extensive experiments demonstrate the effectiveness of our architecture. Finally, our model achieves state-of-the-art performance on the R2R-CE and RxR-CE benchmarks.
Chinese Translation
连续环境下的视觉语言导航(VLN-CE)要求智能体在未见过的环境中遵循自然语言指令进行导航。现有的模仿学习(IL)流程在这种闭环设定中表现不佳:行为克隆受到分布偏移的困扰,而DAgger在轨迹偏离后专家动作会变得模糊不清。虽然强化学习(RL)为解决这一问题提供了自然的范式,但由于奖励稀疏,直接将RL应用于微观动作空间的样本效率低下。为克服这一瓶颈,我们将VLN-CE重新构建为分层马尔可夫决策过程(MDP),显式地将高层规划与底层控制解耦。通过将环境抽象为拓扑图,我们的高层策略在前沿节点构成的宏动作空间上运行,并采用免训练的底层控制器作为其状态转移,这显著压缩了决策时域,使闭环强化学习变得可行。为支持宏MDP上的RL优化,我们提出了一种动作感知价值头,能够在动态的前沿动作空间下有效评估状态价值,从而支撑基于图的PPO算法。大量实验证明了我们架构的有效性。最终,我们的模型在R2R-CE和RxR-CE基准测试中取得了最先进的性能。
cs.RO / 31 / 2609.03927

Toward Unified Robot Learning: Bridging Representation, Vision-Language-Action, and World Models

迈向统一的机器人学习:连接表征学习、视觉-语言-动作模型与世界模型
Mehta, Shaunak A., Hazarika, Ananya, Zhang, Haochen, Yang, Fan, Moriyama, Ryo, Li, Wenkai, Patel, Yash, Suzuki, Kanata
Abstract
For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and reason about the consequences of those actions. Rapid progress in the domains of representation learning, VLA models, and world models has significantly enhanced the capabilities of robot learning systems, enabling robots to work in increasingly complex environments. However, these paradigms are typically developed in isolation, resulting in fragmented systems that struggle with generalization, long-horizon temporal reasoning and planning, and deployment in unstructured environments. In this survey, we present a unified perspective on robot learning by organizing the existing methods along three complementary axes: understanding through representation learning, acting through VLA models, and reasoning through world models. We introduce a structured taxonomy that captures key design choices in environment representation, policy learning, and predictive modeling, and summarize the recent progress in these domains. Beyond classifying the existing works, we analyze how these components interact, discuss common limitations, and highlight emerging trends towards more integrated systems. Through this lens, we identify the challenges in the domain of robot learning, including uncertainty quantification, out-of-distribution generalization, cross-embodiment transfer, long-context understanding, and long-horizon planning. We argue that these challenges arise not only from limitations within individual components but also from the lack of integration across perception, action, and reasoning. Building on this analysis, we outline future directions towards unified, physically grounded, and probabilistic robot learning to develop robust robotic systems that maintain consistent internal representations and support decision making over extended interactions in real-world environments.
Chinese Translation
为了使机器人能够在真实世界环境中可靠地运行,它们需要感知周围环境、执行动作,并对这些动作的后果进行推理。表征学习、VLA(视觉-语言-动作)模型以及世界模型等领域的快速进展显著提升了机器人学习系统的能力,使机器人能够在日益复杂的环境中工作。然而,这些范式通常彼此独立发展,导致了碎片化的系统,在泛化能力、长时域时序推理与规划以及非结构化环境中的部署方面存在困难。在本综述中,我们沿三个互补的维度对现有方法进行组织,提出了一个统一的机器人学习视角:通过表征学习实现理解,通过VLA模型实现行动,通过世界模型实现推理。我们引入了一个结构化的分类体系,涵盖环境表征、策略学习和预测建模中的关键设计选择,并总结了这些领域的最新进展。除了对现有工作进行分类之外,我们还分析了这些组件之间如何相互作用,讨论了共同的局限性,并强调了迈向更集成系统的新兴趋势。基于这一视角,我们指出了机器人学习领域面临的挑战,包括不确定性量化、分布外泛化、跨具身迁移、长上下文理解以及长时域规划。我们认为,这些挑战不仅源于单个组件的局限性,还源于感知、行动与推理之间缺乏整合。基于这一分析,我们概述了迈向统一的、物理接地(physically grounded)的、概率化的机器人学习的未来方向,以构建鲁棒的机器人系统,使其能够维持一致的内部表征,并支持在真实世界环境中进行长时间交互下的决策。
cs.RO / 32 / 2609.03970

Automated Weld Seam Recognition and 3D Mapping for Robotic Post Processing Using Photogrammetry and Semantic Segmentation

基于摄影测量与语义分割的机器人后处理焊缝自动识别与三维映射
Raju, Augustin, Madavath, Abilash, Aubeeluck, Chandra Yuvesh, Pyschny, Nicolas, Hackelöer, Felix, Zwanzig, Florian
Abstract
Accurate identification of weld seam geometries is essential for automated robotic post processing operations such as grinding, finishing, and inspection. For large workpieces, complete surface scanning using high precision laser scanners or structured light sensors can be time consuming and often generates substantial amount of data that are not relevant. This paper presents an experimental vision based pipeline for the approximate localization of weld seams. This serves as a preliminary stage before high precision measurement. The proposed approach aims to reduce the overall scanning effort and data acquisition efficiency. The proposed method includes capturing images of the workpiece from multiple viewpoints, identifying weld seams from the images using semantic segmentation, reconstructing the workpiece using photogrammetry, and projection of identified weld seams into the reconstructed model.
Chinese Translation
焊缝几何形状的准确识别对于打磨、精整和检测等自动化机器人后处理作业至关重要。对于大型工件而言,使用高精度激光扫描仪或结构光传感器进行完整的表面扫描非常耗时,且往往会产生大量无关数据。本文提出了一种基于视觉的实验性流水线,用于焊缝的近似定位,作为高精度测量之前的初步阶段。所提出的方法旨在减少整体扫描工作量并提高数据采集效率。该方法包括:从多个视点采集工件图像,利用语义分割从图像中识别焊缝,通过摄影测量技术重建工件,并将识别出的焊缝投影到重建模型中。
cs.RO / 33 / 2609.03984

MulDP: Multimodal Diffusion Policy for Autonomous Quadruped Parkour Navigation across Complex Terrains

MulDP:面向复杂地形四足机器人自主跑酷导航的多模态扩散策略
Hu, Kangmai, Zhang, Yueqi, Zhai, Peng, Wei, Xiaoyi, Hu, Jiabin, Liu, Zhixiang, Qian, Quancheng, Zhang, Lihua
Abstract
Quadruped robots have demonstrated impressive agility in parkour locomotion across complex terrains. However, most systems still rely on human intervention for high-level planning, and autonomous parkour navigation remains underexplored. The key challenges include fine-grained velocity regulation, long-horizon anticipatory behaviors, and tight coupling between perception and embodied execution. To address these challenges, we propose a Multimodal Diffusion Policy (MulDP) that integrates visual perception with robot proprioception and goal information to generate temporally coherent and anticipatory navigation velocity commands, tightly coupling perception with embodied control to enable robust autonomous navigation. To support the training of MulDP, we construct the first Quadruped Parkour Navigation Dataset (QPND), a multimodal dataset that encompasses diverse navigation behaviors and complex terrains. Extensive simulation and real-world experiments demonstrate that MulDP enables robust long-horizon autonomous navigation and effective traversal across complex terrains.
Chinese Translation
四足机器人在复杂地形的跑酷运动中已展现出令人瞩目的敏捷性。然而,大多数系统仍然依赖人工干预进行高层规划,自主跑酷导航的研究仍不充分。其关键挑战包括细粒度的速度调节、长时程的预见性行为,以及感知与具身执行之间的紧密耦合。为应对这些挑战,我们提出了一种多模态扩散策略,该策略将视觉感知与机器人本体感知及目标信息相融合,生成时间上连贯且具有预见性的导航速度指令,将感知与具身控制紧密耦合,从而实现稳健的自主导航。为支持MulDP的训练,我们构建了首个四足跑酷导航数据集,这是一个涵盖多样化导航行为和复杂地形的多模态数据集。大量仿真与真实世界实验表明,MulDP能够实现稳健的长时程自主导航,并能有效穿越复杂地形。
cs.RO / 34 / 2609.04096

Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

基于可组合基础先验与可泛化抓取合成的自适应视觉-语言抓取
Yan, Sixu, Wang, Shikang, Huang, Binhua, Tang, Xuanlai, Fan, Guohua, Huang, Fan, Li, Haoxuan, Li, Yongkang, Li, Yuhan, Liao, Bencheng, Zhang, Zeyu, Liu, Wenyu, Liu, Hangxin, Wang, Xinggang
Abstract
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Through extensive simulation and real-world experiments, we demonstrate that (i) the base policy exhibits efficient learning and strong cross-hand generalization, (ii) the framework effectively incorporates spatial, cognitive, and temporal priors to address three representative grasping challenges without compromising grasp synthesis performance compared to state-of-the-art methods, and (iii) these priors can operate jointly to enable functional grasping in cluttered and dynamic environments. These results indicate that decoupling physical grasp synthesis from task-dependent understanding provides a scalable paradigm for robotic grasping, allowing future advances in foundation models to be directly translated into improved grasp capabilities without redesigning or retraining the underlying grasp policy. Supplementary videos are available at https://adarobovlg.github.io/
Chinese Translation
本文提出了AdaRoboVLG,一个任务自适应的视觉-语言-抓取框架,支持跨不同机械手的可泛化抓取合成。与现有将基础模型与端到端抓取策略紧密耦合的VLG方法不同,AdaRoboVLG学习一个高效的可泛化基础策略,该策略通过显式运动学映射和基于力闭合的稳定性评估来生成并评价物理可行的抓取候选,同时将任务相关的理解交由专门的基础模型模块处理。这些模块提供可组合的先验,并将其集成到抓取合成过程中,从而无需重新训练底层抓取策略即可实现情境自适应的抓取合成。通过大量仿真与真实世界实验,我们证明:(i) 基础策略表现出高效的学习能力和强大的跨手泛化能力;(ii) 该框架能够有效融合空间、认知与时间先验,以应对三种代表性抓取挑战,且抓取合成性能不逊于最先进方法;(iii) 这些先验可以协同工作,在杂乱和动态环境中实现功能性抓取。这些结果表明,将物理抓取合成与任务相关理解解耦,为机器人抓取提供了一种可扩展的范式,使基础模型的未来进展能够直接转化为抓取能力的提升,而无需重新设计或重新训练底层抓取策略。补充视频请见 https://adarobovlg.github.io/
cs.RO / 35 / 2609.04103

Corner Cases: Headland Coverage Path Planning for Autonomous Driving in Arable Farming

拐角区域:面向可耕地农业自动驾驶的地头覆盖路径规划
Soitinaho, Riikka, Oksanen, Timo
Abstract
This paper presents a new method for headland coverage path planning for arable fields. Several earlier approaches suggest covering the headland with nested polygons and smooth turns, however, covering the field corners entirely requires manoeuvres with reversing. In the new method, the polygon corners are modified to allow a reversing turn. A comparison to two other methods considering gap, overlap, and crossing the field boundary shows an improvement in the coverage result especially in field corners of around 90 degrees, and 240 degrees and above. Applicability of the new method is shown with several examples of real polygonal field maps.
Chinese Translation
本文提出了一种针对可耕地农田地头(headland)覆盖路径规划的新方法。此前的若干方法建议采用嵌套多边形和平滑转弯来覆盖地头,然而要完全覆盖农田拐角区域则需要包含倒车操作的机动动作。在新方法中,通过对多边形拐角进行修改,以实现倒车转弯。与另外两种方法在覆盖空隙、重叠以及越过田地边界方面的比较表明,新方法在覆盖效果上有明显改进,尤其是在约90度以及240度及以上的农田拐角处。通过多个真实多边形农田地图的实例验证了新方法的适用性。
cs.RO / 36 / 2609.04193

GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation

GIFT:面向机器人操作的基于动作导向结构监督的引导式中间特征训练
Zheng, Yupeng, Li, Xiang, Gu, Songen, Zheng, Yuhang, Tian, Shuai, Li, Weize, Wang, Linbo, Li, Chaoyue, Zhang, Qichao, Li, Haoran, Xia, Zhongpu, Zhang, Ya-Qin, Yan, Shuicheng, Zhao, Dongbin
Abstract
Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction objectives may omit critical physical and task structure while retaining control-irrelevant visual redundancy. We call this mismatch between visual richness and control utility the action-sufficiency gap. We investigate whether this gap can be bridged by guiding intermediate features to preserve three control-relevant structure in robotic manipulation: geometry governing motion feasibility, affordance encoding instruction-relevant entities, and goals grounding instructions in task-relevant regions. To this end, we present GIFT (Guided Intermediate Feature Training), an architecture-flexible framework for learning intermediate features that translates these structures into training-time constraints through geometry alignment, affordance prediction, and goal-region reconstruction. We instantiate GIFT in a Vision-Language-Action (VLA) policy, a direct-action World-Action Model (WAM), and an inverse-dynamics WAM while retaining each model's action formulation. Under zero-shot transfer to LIBERO-Plus, GIFT-VLA, GIFT-WAM-Fast, and GIFT-WAM-IDM outperform StarVLA-OFT, Fast-WAM, and Fast-WAM-IDM by 4.6, 12.6, and 5.2 points, reaching 79.6%, 72.6%, and 87.8%, respectively. On RoboCasa, the three GIFT variants reach 61.4%, 83.6%, and 82.3%, outperforming their counterparts by 12.6, 9.0, and 8.4 points, respectively. Together, these results establish learning functionally structured intermediate features as a reusable principle across model-specific action formulations, with especially large gains on articulated-object tasks and high-precision real-world manipulation under unseen visual and spatial perturbations. Project page: https://openphoenix-team.github.io/GIFT-pages.
Chinese Translation
视觉-语言预训练与预测性世界建模为机器人策略提供了丰富的语义和动态视觉特征,但其原生的动作与视觉预测目标可能遗漏关键的物理与任务结构,同时保留与控制无关的视觉冗余。我们将这种视觉丰富性与控制实用性之间的失配称为动作充分性差距(action-sufficiency gap)。我们研究了能否通过引导中间特征来保留机器人操作中的三种控制相关结构来弥合这一差距:决定运动可行性的几何结构、编码与指令相关实体的可供性(affordance)结构,以及将指令锚定于任务相关区域的目标结构。为此,我们提出了 GIFT(Guided Intermediate Feature Training,引导式中间特征训练),这是一个架构灵活的中间特征学习框架,通过几何对齐、可供性预测和目标区域重建,将这些结构转化为训练时的约束。我们在视觉-语言-动作(Vision-Language-Action, VLA)策略、直接动作的世界-动作模型(World-Action Model, WAM)以及逆动力学 WAM 中实例化了 GIFT,同时保留各模型原有的动作建模方式。在向 LIBERO-Plus 的零样本迁移中,GIFT-VLA、GIFT-WAM-Fast 和 GIFT-WAM-IDM 分别比 StarVLA-OFT、Fast-WAM 和 Fast-WAM-IDM 高出 4.6、12.6 和 5.2 个百分点,分别达到 79.6%、72.6% 和 87.8%。在 RoboCasa 上,三个 GIFT 变体分别达到 61.4%、83.6% 和 82.3%,比对应基线高出 12.6、9.0 和 8.4 个百分点。这些结果共同确立了学习具有功能性结构的中间特征这一可跨模型动作建模方式复用的原则,在铰接物体任务以及未见视觉与空间扰动下的高精度真实世界操作中收益尤为显著。项目页面:https://openphoenix-team.github.io/GIFT-pages。
机器学习 (Machine Learning)
85
cs.LG / 1 / 2609.02959

The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors

无知的几何学:大语言模型知道何时该弱化贝叶斯先验
Liu, Toni J. B., Bao, Jiajun, Liu, Yizhou, Arora, Gurbir, Boullé, Nicolas, Sarfati, Raphaël, Earls, Christopher J.
Abstract
What does a language model predict when it has few clues? The answer lurks in its unembedding geometry: a single direction of the unembedding matrix encodes the unigram distribution of the training corpus, which serves as the Bayesian prior the model falls back on when uncertain. This structure --- which we term the \emph{direction of ignorance} --- appears in all four model families examined (\texttt{Llama}, \texttt{Qwen}, \texttt{Gemma}, and \texttt{Pythia}), ranging from 0.4B to 405B parameters. Projecting the final prediction state onto this direction yields a per-token \emph{prior loading factor} $\lambda$, which, empirically, declines steadily as the context becomes more informative. Formally, the same projection decomposes the prediction state into two orthogonal vectors that correspond exactly to the two factors of a tempered Bayesian update: a unigram prior raised to the exponent $\lambda$ and a context-driven likelihood. This geometric-probabilistic interpretation calibrates $\lambda$, making it meaningfully comparable across model sizes and families, with larger models generally exhibiting lower prior reliance in the high-context limit. Finally, we show that the direction of ignorance is causally active: raising or lowering $\lambda$ at the final prediction state steers the prediction toward or away from the unigram prior in KL divergence.
Chinese Translation
当一个语言模型缺乏线索时,它会预测什么?答案隐藏在它的反嵌入(unembedding)几何结构中:反嵌入矩阵的一个方向编码了训练语料的词频(unigram)分布,该分布充当模型在不确定时所回退到的贝叶斯先验。我们将这一结构称为"无知方向"(direction of ignorance),它在所考察的四个模型家族(Llama、Qwen、Gemma 和 Pythia,参数量从 0.4B 到 405B)中均存在。将最终预测状态投影到该方向上,可以得到每个词元的"先验负载因子"(prior loading factor)λ。实证结果表明,随着上下文信息量的增加,λ 稳步下降。从形式上看,同一投影将预测状态分解为两个正交向量,恰好对应温度化贝叶斯更新(tempered Bayesian update)的两个因子:指数为 λ 的词频先验和由上下文驱动的似然。这一几何-概率解释对 λ 进行了校准,使其在不同模型规模和家族之间具有可比性;总体而言,在高上下文情形下,更大的模型对先验的依赖更低。最后,我们证明无知方向具有因果作用:在最终预测状态处提高或降低 λ,会以 KL 散度衡量,使预测趋向或偏离词频先验。
cs.LG / 2 / 2609.02982

Equation Recast for Canonical Operator Learning Across Parametric PDEs

面向参数化偏微分方程的规范算子学习的方程重铸方法
Cheng, Qiyun, Duruisseaux, Valentin, Clauser, Cesar F., Sahadath, Md Hossain, Yang, Huihua, Pan, Shaowu, Ferraro, Nathaniel, Anandkumar, Anima, Ji, Wei, Rea, Cristina
Abstract
Learning solution operators across broad parameter ranges can require substantial coverage of both input functions and physical parameters, particularly for purely data-driven parametric models. In addition, the resulting models may fail silently outside the training distribution. We introduce equation recast, which reformulates parametric operator learning as the learning of a single canonical operator. Parameter-induced operator variations are derived analytically from the governing equation and absorbed into effective sources, enabling zero-shot prediction across new parameter regimes. Across multi-parameter, nonlinear, and singular PDE settings, equation recast supports extrapolation, integrates sparse heterogeneous datasets in a shared canonical representation, and uses loss of convergence as an internal warning signal for failure of the recast iteration. In high-fidelity tokamak simulations for nuclear fusion, the framework unifies electron-temperature data across four device geometries through canonical-domain mapping within one jointly trained operator. Equation recast provides a route toward reusable neural PDE solvers combining equation-guided transfer, data efficiency, and monitorable inference.
Chinese Translation
在宽泛的参数范围内学习解算子,可能需要对输入函数和物理参数进行大量覆盖,尤其是对于纯数据驱动的参数化模型。此外,所得模型在训练分布之外可能会静默失效。我们提出了方程重铸(equation recast)方法,将参数化算子学习重新表述为对单一规范算子的学习。参数引起的算子变化可通过控制方程解析地推导出来,并被吸收到等效源项中,从而实现对新参数区域的零样本预测。在多参数、非线性和奇异偏微分方程的多种设置中,方程重铸支持外推预测,能够在统一的规范表示下整合稀疏的异构数据集,并利用收敛性损失作为重铸迭代失败的内部预警信号。在核聚变的高保真托卡马克仿真中,该框架通过规范域映射,在单一联合训练的算子内统一了四种装置几何构型下的电子温度数据。方程重铸为实现可复用的神经偏微分方程求解器提供了一条路径,兼具方程引导的迁移能力、数据效率和可监控的推理过程。
cs.LG / 3 / 2609.02984

From Euclidean to Graph-Structured Data: A Survey of Collaborative Learning

从欧几里得数据到图结构数据:协同学习综述
Bourgerie, Rémi, Girdzijauskas, Šarūnas, Fodor, Viktoria
Abstract
The conventional approach to machine learning, that is, collecting data, training models, and performing inference in a single location, faces fundamental limitations, including scalability and privacy, that restrict its applicability. To address these challenges, recent research has explored collaborative learning approaches, including federated learning and decentralized learning, where individual agents perform training and inference locally, with limited collaboration. Most collaborative learning research focuses on Euclidean data with regular, grid-like structure (e.g., images, text). However, these approaches fail to capture the relational patterns in many real-world applications, best represented by graphs. Learning on graphs relies on message-passing mechanisms to propagate information between connected nodes, making it conceptually well-suited for collaborative environments where agents must exchange information. Yet, the opportunities and challenges of learning on graph-structured data in collaborative settings remain largely underexplored. This survey provides a comprehensive investigation of collaborative learning from Euclidean to graph-structured data, aiming to consolidate this emerging field. We begin by reviewing its foundational principles for Euclidean data, organizing them along three core dimensions: learning effectiveness, efficiency, and privacy preservation. We then extend the discussion to graph-structured data, introducing a taxonomy of graph distribution scenarios, characterizing associated statistical heterogeneities, and developing standardized problem formulations and algorithmic frameworks. Finally, we systematically identify open challenges and promising research directions.
Chinese Translation
传统的机器学习方法,即在单一位置收集数据、训练模型并进行推理,面临着可扩展性和隐私等根本性局限,这限制了其适用性。为应对这些挑战,近年来的研究探索了协同学习方法,包括联邦学习(federated learning)和去中心化学习(decentralized learning),其中各个智能体在本地执行训练和推理,并仅进行有限的协作。大多数协同学习研究聚焦于具有规则网格结构的欧几里得数据(如图像、文本)。然而,这些方法难以捕捉许多现实应用中以图结构为最佳表示的关系模式。图上的学习依赖于消息传递机制在相连节点之间传播信息,这使其在概念上非常适合于智能体必须交换信息的协同环境。然而,在协同环境下对图结构数据进行学习所带来的机遇与挑战在很大程度上仍未被充分探索。本综述对从欧几里得数据到图结构数据的协同学习进行了全面考察,旨在整合这一新兴领域。我们首先回顾了欧几里得数据协同学习的基础原理,并从三个核心维度加以组织:学习有效性、效率与隐私保护。随后,我们将讨论扩展至图结构数据,引入图分布场景的分类体系,刻画相关的统计异质性,并建立标准化的问题表述与算法框架。最后,我们系统地识别了开放性挑战和有前景的研究方向。
cs.LG / 4 / 2609.02986

Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture Design

现代Transformer是隐式混合架构:从功能分化到有原则的混合架构设计
Shi, Runlin, Yin, Bojian, Li, Guoqi
Abstract
Hybrid architectures combining Full Attention (FA) and Linear Attention (LA) are increasingly prominent, yet their allocation remains heuristic. We seek an evidence-grounded basis in head-level functional organization learned by RoPE-based Transformers. Behavioral probes do not yield a complete taxonomy, so we propose two intervention metrics: RoPE Frequency Importance Score (RFIS), measuring how each frequency affects a head's attention distribution, and RoPE Positional Dependence (RPD), isolating dependence on rotary positional modulation. On Qwen3-series models and Llama3.1, RFIS suggests and RPD verifies a complete taxonomy of retrieval and positional heads separated by a salient mid-low-frequency band. Controlled Transformers show that this boundary follows the training-length positional scale; we term it the Global Positional Band (GPBand). The analysis suggests a potential cause of zero-shot length-extrapolation failure and yields two principles: positional modeling should operate only locally, with global access through position-independent retrieval; and both functions should be assigned at head granularity with layer-specific allocation. We instantiate them in Head-wise Hybrid Architecture (HwH), using NoPE FA for global retrieval and LA for local positional modeling. With an FA-to-LA ratio below 1:3, HwH retains strong language modeling and commonsense reasoning while improving retrieval and substantially strengthening zero-shot long-context extrapolation over Transformer, LA, and a layer-wise hybrid baseline. Ablations validate both principles and component roles, highlighting principled hybrid architecture design as a promising route toward future foundation models.
Chinese Translation
结合全注意力(Full Attention, FA)与线性注意力(Linear Attention, LA)的混合架构日益受到关注,然而其分配方式仍依赖启发式规则。我们试图在基于RoPE的Transformer所学到的头级别功能组织中寻找有实证依据的基础。行为探测无法给出完整的分类体系,因此我们提出两个干预性度量指标:RoPE频率重要性分数(RoPE Frequency Importance Score, RFIS),衡量每个频率对某个注意力头注意力分布的影响;以及RoPE位置依赖性(RoPE Positional Dependence, RPD),用于分离对旋转位置调制的依赖。在Qwen3系列模型和Llama3.1上,RFIS给出并通过RPD验证了一个完整的分类体系:检索头与位置头由一个显著的中低频段分隔开来。受控Transformer实验表明,这一边界遵循训练长度的位置尺度;我们将其命名为全局位置频段(Global Positional Band, GPBand)。该分析揭示了零样本长度外推失败的一个潜在原因,并得出两条原则:位置建模应仅在局部进行,全局访问应通过位置无关的检索实现;且两类功能应在头的粒度上进行分配,并采用逐层特定的分配方式。我们将这些原则具体化为头级混合架构(Head-wise Hybrid Architecture, HwH),使用NoPE FA进行全局检索,使用LA进行局部位置建模。在FA与LA比例低于1:3的情况下,HwH在保持强大语言建模和常识推理能力的同时,改进了检索性能,并在零样本长上下文外推方面显著优于Transformer、LA以及逐层混合基线。消融实验验证了两条原则及各组件的作用,凸显了有原则的混合架构设计作为未来基础模型发展的可行路径。
cs.LG / 5 / 2609.02987

Tail-Likelihood Reinforcement Learning

尾似然强化学习(Tail-Likelihood Reinforcement Learning)
Ramasubramanian, Shrinivas, Arora, Daman, Tajwar, Fahim, Zeng, Guanning, Wu, Qingyang, Zhou, Zhongzhu, Xu, Chenfeng, Feng, Haiwen, Song, Yuda, Singh, Aarti, Salakhutdinov, Ruslan, Bagnell, J. Andrew, Schneider, Jeff, Zanette, Andrea
Abstract
Reinforcement learning typically optimizes average reward. For generative policies, the average can hide an important distinction: two policies can achieve the same mean reward while having very different chances of producing a rare but high-reward rollout. This matters as sampling increases during training and inference, since its benefit depends on retaining probability mass on high-reward outcomes. We propose to optimize this coverage directly. Rather than considering only expected reward, we consider all of its upper tails: for each reward threshold, how likely is the policy to exceed it? This turns a continuous reward into a family of binary success events. We introduce Tail-Likelihood Reinforcement Learning (TailRL), which maximizes the log-probability of exceeding a randomly chosen reward threshold. Its gradient gives more weight to rare, high-reward rollouts and can be interpreted as a mixture of Best-of-(k) gradients. TailRL requires only a simple modification to the advantage function, making it compatible with existing reinforcement learning pipelines. Across object localization, maze navigation, GUI grounding, and code optimization, TailRL leverages rare high-reward training samples to avoid suboptimal solutions and yields models that benefit more from additional samples at inference time.
Chinese Translation
强化学习通常优化平均奖励。对于生成式策略而言,平均值可能掩盖一个重要区别:两个策略可以在平均奖励相同的情况下,产生罕见但高奖励的采样结果(rollout)的概率却截然不同。随着训练和推理过程中采样次数的增加,这一点变得尤为重要,因为其收益取决于策略是否在高奖励结果上保留了概率质量。我们提出直接优化这种覆盖能力。我们不仅考虑期望奖励,还考虑其所有上尾部分:对于每个奖励阈值,策略超过该阈值的概率有多大?这将连续的奖励转化为一系列二元成功事件。我们提出尾似然强化学习(Tail-Likelihood Reinforcement Learning,TailRL),它最大化超过随机选取的奖励阈值的对数概率。其梯度对罕见的高奖励采样结果赋予更大的权重,并可被解释为 Best-of-(k) 梯度的混合。TailRL 仅需对优势函数进行简单修改,因此与现有的强化学习流程兼容。在目标定位、迷宫导航、GUI 定位和代码优化等任务上,TailRL 利用罕见的高奖励训练样本避免了次优解,并使模型在推理时能从额外的采样中获得更多收益。
cs.LG / 6 / 2609.02988

Mesh-Native Physics-Informed Graph Surrogates for TCAD-in-the-Loop Design Space Exploration

面向TCAD在环设计空间探索的网格原生物理信息图代理模型
Popryho, Leonid, Sadeghi, Ayoub, Partin-Vaisband, Inna
Abstract
High-fidelity TCAD simulation of drift-diffusion transport remains the workhorse of emerging FinFET device design, but it is computationally expensive, especially for 3D structures where runtime escalates steeply with mesh complexity. This sharply limits multi-objective design space exploration. Existing machine-learning surrogates map a fixed set of design parameters to a few scalar device metrics, discarding the underlying physics and losing transferability across device geometries and families. A physics-informed graph attention network (GAT) surrogate is proposed. It operates directly on the tetrahedral TCAD mesh and predicts, at every mesh node, the electrostatic potential together with the electron and hole quasi-Fermi levels, the fundamental unknowns of the drift-diffusion system. Training combines a data loss with finite-volume current-continuity residuals, embedding carrier-transport physics into the objective. Operating on the mesh as a graph, the surrogate inherits size generalization: a model trained on few-fin meshes applies unchanged to substantially larger arrays, bounded at inference only by GPU memory. Per-node uncertainty from a deep ensemble drives an active-learning loop that screens large candidate pools in seconds and forwards only the most informative designs for full simulation. Benchmarked against Sentaurus Device on multi-fin tri-gate FinFETs, the surrogate reproduces the three drift-diffusion fields with sub-volt per-field RMSE and reaches a per-design throughput orders of magnitude higher than the full simulator. The advantage grows with device size: on large multi-fin arrays that are prohibitively slow to simulate directly, inference still completes in under a second per device, enabling Pareto-front exploration across device scales infeasible for direct TCAD sweeps.
Chinese Translation
漂移-扩散输运的高保真TCAD仿真仍是新兴FinFET器件设计的主力工具,但其计算代价高昂,尤其是在三维结构中,运行时间随网格复杂度急剧上升,这严重限制了多目标设计空间探索。现有的机器学习代理模型将一组固定的设计参数映射到少数几个标量器件指标上,丢弃了底层物理信息,并丧失了跨器件几何结构和器件族的迁移能力。本文提出一种物理信息图注意力网络(GAT)代理模型,它直接作用于四面体TCAD网格,并在每个网格节点上预测静电势以及电子和空穴的准费米能级——即漂移-扩散系统的基本未知量。训练将数据损失与有限体积电流连续性残差相结合,将载流子输运物理嵌入到目标函数中。由于以图的形式在网格上运行,该代理模型具备尺寸泛化能力:在少量鳍(few-fin)网格上训练的模型可以不加修改地应用于规模大得多的阵列,推理时仅受GPU显存的限制。基于深度集成(deep ensemble)得到的逐节点不确定性驱动一个主动学习循环,可在数秒内筛选大规模候选池,并仅将信息量最大的设计送入完整仿真。在多鳍三栅FinFET上与Sentaurus Device进行基准对比,该代理模型以每场低于1伏特的RMSE复现了三个漂移-扩散场,且每个设计的吞吐量比完整仿真器高出数个数量级。这一优势随器件尺寸增大而增强:对于直接仿真慢得难以接受的大型多鳍阵列,推理仍可在每器件不到一秒内完成,从而实现了直接TCAD扫描无法企及的跨器件尺度帕累托前沿探索。
cs.LG / 7 / 2609.02991

TRACE: Spatiotemporal Contact Memory Graph Network Simulator for Granular Dynamics

TRACE:面向颗粒动力学的时空接触记忆图网络模拟器
Zhou, Changjian, Yousefpour, Negin, Qi, Jie, Fang, Junfeng, Narsilio, Guillermo A., Jostad, Hans Petter
Abstract
Learned graph simulators provide an efficient alternative to high-fidelity solvers for granular dynamics. However, granular motion depends strongly on inter-granular contact history, which is difficult to preserve when particle contacts form, break, and rearrange. Existing simulators mainly store temporal information in node features or node-level memory. Here we introduce TRACE, a graph-network simulator that stores interaction history directly on contact edges. Each edge maintains a persistent memory updated by attention-based message passing and a gated recurrent unit, while an edge-identity dictionary preserves this memory as the contact graph changes. A physics-structured decoder predicts inter-granular normal and tangential contact forces, enforces the Coulomb friction limit, and applies equal-and-opposite internal forces. The model is trained with single-step pretraining followed by autoregressive rollout fine-tuning. We evaluate TRACE on 2D and 3D granular column-collapse benchmarks. In both cases, TRACE produces stable, physically consistent long-horizon rollouts, closely reproducing the final deposit geometry and the kinetic energy released during collapse. Compared with graph network simulator (GNS) and node-memory graph neural simulator (NMGNS), TRACE reduces long-rollout position error by 31-62% and final-deposit error by 58-89% across the two benchmarks, while using fewer parameters and maintaining near-zero particle interpenetration. TRACE also achieves 12.2$\times$ and 8.9$\times$ speedups over the material point method (MPM) reference solver in 2D and 3D, respectively. Our code is available at https://github.com/Data-Driven-Computational-Geotechnics/TRACE.
Chinese Translation
学习型图模拟器为颗粒动力学的高保真求解器提供了一种高效替代方案。然而,颗粒运动强烈依赖于颗粒间的接触历史,而当颗粒接触形成、断开并重新排列时,这一历史难以保留。现有模拟器主要将时间信息存储在节点特征或节点级记忆中。本文提出TRACE,一种将相互作用历史直接存储在接触边上的图网络模拟器。每条边维护一个持久记忆,通过基于注意力的消息传递和门控循环单元进行更新;同时,一个边身份字典在接触图变化时保留该记忆。物理结构化解码器预测颗粒间法向和切向接触力,施加库仑摩擦极限,并应用大小相等、方向相反的内部作用力。模型采用单步预训练加自回归滚动微调的方式进行训练。我们在二维和三维颗粒柱坍塌基准上评估了TRACE。在这两种情况下,TRACE均能产生稳定且物理一致的长时程滚动预测,精确复现最终沉积几何形态以及坍塌过程中释放的动能。与图网络模拟器(GNS)和节点记忆图神经网络模拟器(NMGNS)相比,TRACE在两个基准上将长时程滚动位置误差降低了31-62%,将最终沉积误差降低了58-89%,同时参数更少,且保持几乎为零的颗粒互穿透。此外,TRACE在二维和三维中分别相比物质点法(MPM)参考求解器实现了12.2倍和8.9倍的加速。我们的代码可在 https://github.com/Data-Driven-Computational-Geotechnics/TRACE 获取。
cs.LG / 8 / 2609.02993

No-Regret Bayesian Optimization with Finite-Library Input-Warped Kernels

基于有限库输入扭曲核的无遗憾贝叶斯优化
Augustinsson, Edvin Ketabati, Bridges, Robert A.
Abstract
Gaussian-process Bayesian optimization (GP-BO) excels at black-box optimization of costly functions, e.g., hyperparameter optimization (HPO) and multi-agent system (MAS) design. Convergence-rate guarantees exist for select methods, notably GP upper confidence bound (GP-UCB), but require a fixed kernel. Critically, the kernel encodes how input proximity affects objective value similarity. When raw coordinates poorly match this geometry - as with log-scaled hyperparameters or localized peaks - input warping can greatly improve sample efficiency, yet known GP-UCB proofs require a fixed kernel. We propose Finite-Library Input-Warped Bayesian Optimization (FLIWBO), which selects warps from a finite library of smooth input maps by any history-dependent rule. It adapts the input geometry to accelerate learning while retaining high-probability convergence guarantees under mild hypotheses, with an explicit $\sqrt(N_\varepsilon)$ library-size cost. Controlled diagnostics show that finite-library warping repairs planted geometry mismatches and identify FLIWBO failure cases. Across four repeated benchmarks - warped synthetic objectives, a confidence-fence trap, and Fashion-MNIST HPO - FLIWBO-UCB beats raw-coordinate GP-UCB under misspecified geometry, escapes traps that defeat even oracle-warp expected improvement, and recovers much of the gain from manual log scaling, while leading the tested methods that admit a matching regret guarantee. A 20-dimensional MAS design study further shows feasibility under costly noisy evaluations. Code for experiments is available: https://github.com/edvin-ketabati/bogp-paper-experiments.
Chinese Translation
高斯过程贝叶斯优化(GP-BO)在昂贵函数的黑盒优化中表现出色,例如超参数优化(HPO)和多智能体系统(MAS)设计。部分方法(尤其是高斯过程上置信界方法 GP-UCB)已具有收敛率保证,但其要求使用固定的核函数。关键在于,核函数编码了输入邻近性如何影响目标值相似性。当原始坐标与这种几何结构匹配不佳时(如对数尺度超参数或局部峰值情形),输入扭曲(input warping)可大幅提升样本效率,然而已有的 GP-UCB 理论证明均要求核函数固定。我们提出有限库输入扭曲贝叶斯优化(Finite-Library Input-Warped Bayesian Optimization,FLIWBO),其通过任意依赖历史的规则从由光滑输入映射构成的有限库中选择扭曲方式。该方法能够自适应地调整输入几何结构以加速学习,同时在温和假设下保留高概率收敛保证,且代价仅为显式的 $\sqrt(N_\varepsilon)$ 库规模开销。受控诊断实验表明,有限库扭曲能够修复人为植入的几何失配,并揭示了 FLIWBO 的失效情形。在四个重复基准实验中——扭曲合成目标函数、置信围栏陷阱以及 Fashion-MNIST 超参数优化——在几何设定失配的情况下,FLIWBO-UCB 优于基于原始坐标的 GP-UCB,能够逃脱连预言机扭曲(oracle-warp)期望改进方法都无法应对的陷阱,并恢复了人工对数变换所带来的大部分收益,同时在可提供匹配遗憾保证的测试方法中表现领先。一项 20 维多智能体系统设计研究进一步验证了该方法在昂贵带噪评估下的可行性。实验代码见:https://github.com/edvin-ketabati/bogp-paper-experiments。
cs.LG / 9 / 2609.02996

Evaluating Graph Neural Networks for Change-Criticality Classification in Maritime Navigation Charts

评估图神经网络在海事航海图变更关键性分类中的应用
Potnis, Abhishek, Arndt, Jacob
Abstract
Graph neural networks (GNNs) are a class of neural networks suitable for learning on graph-structured data. Their application to spatial data is a natural extension, however its relatively unclear which message-passing operations, architectural configurations, and graph representation is best suited for classifying changes to objects in electronic navigational charts (ENCs)--geospatial vector datasets used for marine navigation. Maintaining these datasets is a challenge, and categorizing changes to objects in the ENC based on their significance to navigational safety is of particular importance. Here, we propose to represent these vector navigation datasets as a graph structure where the spatial objects serve as nodes and their spatial and semantic relationships form edges. We encode both the old ENC dataset and new ENC dataset into a pair of graphs and frame the task as a graph-pair classification problem. Building on this representation, we investigate the use of GNN architectures to classify whether the encoded graphs constitutes a critical or non-critical risk to navigational safety. We train and evaluate several GNN architectures and model configurations on ENC changes reviewed by maritime experts. Our results demonstrate that graph-based representations improve the classification of ENC updates, providing a scalable approach for automating or improving ENC maintenance workflows.
Chinese Translation
图神经网络(GNN)是一类适用于图结构数据学习的神经网络。将其应用于空间数据是一种自然的扩展,然而,哪种消息传递操作、架构配置和图表示最适合对电子航海图(ENC)——用于海上导航的地理空间矢量数据集——中的对象变更进行分类,目前仍不明确。维护这些数据集是一项挑战,而根据对象变更对航行安全的重要性对其进行分类尤为关键。本文提出将这些矢量导航数据集表示为图结构,其中空间对象作为节点,它们的空间和语义关系构成边。我们将旧的ENC数据集和新的ENC数据集分别编码为一对图,并将该任务构建为图对分类问题。基于这种表示,我们研究了使用GNN架构对编码后的图进行分类,判断其是否构成对航行安全的关键或非关键风险。我们在经过海事专家审核的ENC变更数据上训练并评估了多种GNN架构和模型配置。结果表明,基于图的表示改善了ENC更新的分类效果,为自动化或改进ENC维护工作流程提供了一种可扩展的方法。
cs.LG / 10 / 2609.02998

Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation

先验证再蒸馏:面向在线策略蒸馏的提示级教师门控机制
Zhang, Zhiwei, Sun, Zechen, Zhao, Fei, Peng, Kang, Liang, Bin, Deng, Huayu, Hu, Yao, Wong, Kam-Fai, Chuan, Mu
Abstract
On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a confidently wrong teacher can induce a strong yet misleading update. Distributional proxies, such as entropy or teacher-student likelihood agreement, measure uncertainty or agreement but do not directly verify outcome correctness. We introduce Teacher-Gated On-Policy Distillation (TGOPD), built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted. TGOPD estimates reliability from a small set of verifier-scored teacher probes and routes each prompt exclusively to dense OPD when the reliability check passes or to verifier-grounded GRPO otherwise. Across 4B and 35B students in mathematics, code, and instruction following, TGOPD outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training. By using otherwise-idle teacher capacity for reliability estimation, TGOPD also reduces teacher-side compute waste in asynchronous OPD, increasing teacher-node GPU utilization from 9.8% to 78.9% in the measured 4B single-domain run.
Chinese Translation
在线策略蒸馏(On-Policy Distillation, OPD)通过在学生模型自身的生成结果上提供来自冻结教师模型的密集词元级监督,从而加速后训练过程。原始的 OPD 在所有提示上均匀地施加这种监督,而不检查教师模型在每个提示上是否可靠。由于反向 KL 散度具有模式寻求(mode-seeking)特性,一个自信但错误的教师模型可能导致强烈却具有误导性的更新。熵或师生似然一致性等分布代理指标仅度量不确定性或一致性,并不能直接验证结果的正确性。我们提出了教师门控在线策略蒸馏(Teacher-Gated On-Policy Distillation, TGOPD),其核心原则是:在采纳密集监督之前,应在提示级别验证教师模型的可靠性。TGOPD 通过一小组经验证器打分的教师探测样本(teacher probes)估计可靠性,并将每个提示在可靠性检查通过时专门路由至密集 OPD,否则路由至基于验证器的 GRPO。在数学、代码和指令遵循任务上的 4B 和 35B 学生模型实验中,TGOPD 在全部六个单领域设置中均优于原始 OPD,并且在多领域训练下,两种规模模型的七项基准平均分均更高。通过利用原本闲置的教师模型算力进行可靠性估计,TGOPD 还减少了异步 OPD 中教师侧的计算浪费,在所测量的 4B 单领域运行中,将教师节点 GPU 利用率从 9.8% 提升至 78.9%。
cs.LG / 11 / 2609.03003

Causal Foundation Models

因果基础模型
Stith, Christopher, Rahmani, Hossein, Cresswell, Jesse C.
Abstract
Causal inference is the practice of estimating the effect of a treatment or intervention from data. It traditionally requires a bespoke pipeline for every new problem: first proposing a causal mechanism, selecting a compatible estimator, and finally training it. Meanwhile, across diverse settings and modalities, much of machine learning has shifted to the paradigm of foundation models: networks pretrained once at scale and applied to new tasks without fine-tuning. Causal foundation models (CFMs) bring this paradigm to causal inference. CFMs are pretrained neural networks that estimate causal quantities, such as the average treatment effect, on entirely new datasets using in-context learning without requiring model updates. This work provides a practical introduction to this emerging area. We summarize the necessary background in causal inference and machine learning before discussing CFMs. Throughout, we include example code and Jupyter notebooks.
Chinese Translation
因果推断是从数据中估计处理或干预效应的实践。传统上,它需要针对每个新问题构建定制的流水线:首先提出因果机制,然后选择兼容的估计器,最后对其进行训练。与此同时,在不同的场景和模态下,机器学习的很大一部分已经转向了基础模型的范式:网络在大规模数据上预训练一次,即可应用于新任务而无需微调。因果基础模型(Causal Foundation Models, CFMs)将这一范式引入因果推断领域。CFM是预训练的神经网络,能够通过上下文学习(in-context learning)在全新的数据集上估计因果量(例如平均处理效应),而无需更新模型。本工作为这一新兴领域提供了实用的介绍。在讨论CFM之前,我们总结了因果推断和机器学习的必要背景知识。全文贯穿了示例代码和Jupyter笔记本。
cs.LG / 12 / 2609.03026

ObserverBench: Testing Mechanistic Estimates for Intervention and Control

ObserverBench:面向干预与控制的机制性估计测试
Erramilli, Vijay
Abstract
Mechanistic interpretability is increasingly used to guide interventions such as activation steering, circuit removal, and safety monitoring. Yet an internal estimate that is accurate on average can still choose a poor action. We present ObserverBench, a benchmark framework for testing whether an internal estimator---an observer---is adequate for the intervention, control, or safety task it directs. Each task fixes the model, information boundary, allowed actions, decision rule, held-out cases, and loss. The benchmark reports estimation accuracy separately from the loss caused by the chosen action. Theory and experiments show why both are needed. In closed-loop control, observer errors matter at the starting point and along directions the allowed intervention can reach. On circuit-intervention tasks in GPT-2-small and Qwen2.5-7B, pairwise observers predict unseen effects more accurately without always choosing better actions; observers trained on action loss choose lower-loss actions. In safety triage, a score that perfectly separates violations can allocate a fixed intervention budget poorly when violations have different costs. Across Qwen2.5-7B, Gemma-2-9B-it, and prospectively frozen Qwen3.5-9B APPS tasks, AUROC can rank monitors differently from deployment loss, and the best information source changes across models. Sparse SAE readouts also trail their layer-matched dense controls on the reported Qwen panels, under disclosed activation-density or checkpoint mismatches. ObserverBench provides fixed task contracts, runnable baselines, and table-based submissions for evaluating interpretability methods through the actions they enable.
Chinese Translation
机制可解释性(mechanistic interpretability)正被越来越多地用于指导激活引导(activation steering)、电路移除(circuit removal)和安全监控等干预手段。然而,一个平均意义上准确的内部估计仍可能选择糟糕的行动。我们提出了 ObserverBench,一个用于测试内部估计器(即观察者,observer)对于其所引导的干预、控制或安全任务是否足够的基准框架。每个任务固定模型、信息边界、允许的行动、决策规则、保留测试样本以及损失函数。该基准将估计精度与所选行动造成的损失分开报告。理论与实验表明为什么两者都需要。在闭环控制中,观察者误差在起始点以及允许干预所能到达的方向上都很重要。在 GPT-2-small 和 Qwen2.5-7B 的电路干预任务上,成对观察者(pairwise observers)在预测未见效应方面更为准确,但并不总是选择更好的行动;而基于行动损失训练的观察者则选择损失更低的行动。在安全分诊(safety triage)中,一个能完美区分违规行为的评分,在违规行为代价不同的情况下,仍可能对固定的干预预算分配不当。在 Qwen2.5-7B、Gemma-2-9B-it 以及前瞻性冻结的 Qwen3.5-9B APPS 任务上,AUROC 对监控器的排序可能与部署损失不一致,且最佳信息来源在不同模型间有所变化。在公开的 Qwen 面板上,稀疏 SAE 读出(Sparse SAE readouts)在披露的激活密度或检查点不匹配的条件下,也落后于其层匹配的稠密对照(dense controls)。ObserverBench 提供固定的任务契约、可运行的基线以及基于表格的提交方式,用于通过可解释性方法所支持的行动来评估这些方法。
cs.LG / 13 / 2609.03069

Learnable composition for neural operators

神经算子的可学习组合
Chen, Zituo, Zhang, Baiming, Deng, Sili
Abstract
Neural operators are fast, differentiable surrogates for physical simulation, but their accuracy often degrades when domain geometry, size, or operating conditions differ from training. Supervised adaptation can recover accuracy, but even a small target set requires costly high-fidelity simulations. We therefore ask how pretraining and transfer can be designed together to reduce this deployment cost. LatentDDM first pretrains a neural operator to predict fields on small subdomains. For a new setting, it freezes this operator and trains only a lightweight module that composes the local predictions. We evaluate our method on two complementary problems: steady Darcy flow, where long-range pressure coupling must extend across increasingly large porous domains, and unsteady incompressible flow around a pitching airfoil, where rollout errors compound as target pitching frequencies exceed the training range. Compared with the capacity-matched models that process the full domain at once, LatentDDM's error is 36-56% lower on larger Darcy domains after adaptation with 16 target simulations. It also improves 20-step field rollouts in fast-pitching airfoil flow, both zero-shot and after few-shot calibration. These results identify the co-designed local pretraining and composition-level transfer as a promising design principle for physical foundation models.
Chinese Translation
神经算子是物理仿真中快速、可微分的代理模型,但当求解域的几何形状、尺寸或运行条件与训练时不同,其精度往往会下降。有监督的适配可以恢复精度,但即使是一个很小的目标数据集,也需要代价高昂的高保真仿真。因此,我们提出的问题是:如何协同设计预训练与迁移,以降低这种部署成本。LatentDDM 首先预训练一个神经算子,使其在小的子域上预测物理场。对于新的应用场景,该方法冻结该算子,仅训练一个轻量级模块来组合各局部预测结果。我们在两个互补的问题上评估了所提方法:其一是稳态 Darcy 流,其中长程压力耦合必须跨越越来越大的多孔介质域;其二是绕俯仰机翼的非定常不可压缩流,当目标俯仰频率超出训练范围时,滚动预测误差会不断累积。与直接处理整个域且容量相当(capacity-matched)的模型相比,在仅有 16 个目标仿真样本的适配后,LatentDDM 在更大的 Darcy 域上误差降低 36–56%。该方法还改善了快速俯仰机翼流的 20 步场滚动预测,无论是零样本(zero-shot)还是在少样本校准之后均如此。这些结果表明,协同设计的局部预训练与组合层面的迁移,是构建物理基础模型的一种有前景的设计原则。
cs.LG / 14 / 2609.03079

LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference

LeanStream:一种用于高效端侧大语言模型推理的“推测-精化”流式框架
Liu, Renyuan, Leng, Yuyang, Liu, Kaiyan, Zhong, Yuzhou, Hu, Shaohan, Chun-Fu, Chen, Zhao, Peijun, Yun, Heechul, Yao, Shuochao
Abstract
On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior systems exploit activation sparsity and offload weights to SSD or flash storage, but face a fundamental systems trade-off: accurate sparse execution decisions require the latest context, whereas efficient computation-I/O overlap requires early prediction. As a result, existing designs either serialize execution or incur redundant weight fetches, extra computation, and large cache overheads. We present LeanStream, a streaming speculate-and-refine framework for efficient on-device LLM inference. LeanStream progressively refines computation, loading, and cache-retention priorities using partial GPU results, enabling fine-grained overlap between GPU execution and storage I/O. We implement LeanStream on both mobile and embedded platforms. Compared with prior on-device LLM inference systems, LeanStream reduces memory usage by 4.8$\times$ to 7.5$\times$ at the best throughput achieved by prior work, while further improving token generation throughput by 1.6$\times$ to 2.1$\times$.
Chinese Translation
端侧大语言模型(LLM)推理在隐私保护和响应速度方面极具吸引力,但由于模型权重远超可用DRAM容量,其在移动和嵌入式设备上仍然面临巨大挑战。已有系统利用激活稀疏性并将权重卸载至SSD或闪存,但面临一个根本性的系统权衡:准确的稀疏执行决策需要最新的上下文信息,而高效的计算与I/O重叠则需要提前预测。因此,现有设计要么将执行串行化,要么导致冗余的权重读取、额外的计算开销以及巨大的缓存开销。我们提出了LeanStream,一种面向高效端侧LLM推理的流式“推测-精化”(speculate-and-refine)框架。LeanStream利用部分GPU计算结果逐步精化计算、加载和缓存保留的优先级,从而实现GPU执行与存储I/O之间的细粒度重叠。我们在移动和嵌入式平台上均实现了LeanStream。与已有的端侧LLM推理系统相比,在先前工作达到的最佳吞吐量下,LeanStream将内存占用降低了4.8倍至7.5倍,同时将token生成吞吐量进一步提升了1.6倍至2.1倍。
cs.LG / 15 / 2609.03090

The Gradient Does Not See Rank: Rank-Indifference in Matrix-CODI on ProsQA

梯度无法感知秩:Matrix-CODI 在 ProsQA 上的秩无关性
Larson, Samuel
Abstract
Continuous chain-of-thought models compress reasoning into latent tokens. Matrix-valued variants, which route each latent token through a d x d matrix bottleneck, introduce rank as a single-sample structural observable on the latent matrix Z. If matrix latents carry parallel reasoning paths via superposition, rank should track them, and truncating Z to low rank should hurt accuracy on tasks whose solutions plausibly require multiple components. Across four training regimes of a matrix-CODI model (three on ProsQA, one on GSM8K-Aug below the learning threshold), the rank-k projection ablation curve is flat to within 0.6 percentage points. A three-seed replication yields 81.0 +/- 2.0 percentage points accuracy while the final effective rank of Z spans {4, 12, 13}; the loss does not reward any particular rank. To test whether rank-blindness arises from the flatten-then-project readout alone, we trained four readouts: a bilinear reparametrization, a bilinear-plus-GELU readout nonlinear in Z, an SVD-augmented readout feeding singular values through an MLP, and a quadratic readout in Z Z^T. All four rank-k curves remain flat (Spearman p-values 0.63, 0.14, 0.82, 0.46). The flat curves persist for readouts nonlinear in Z. A linear probe on Z underperforms a raw pretrained hidden state at target prediction (AUC 0.673 vs. 0.846). A negative control on vanilla GPT-2 SFT (no matrix bottleneck, no Z, three seeds, n=500) reproduces a flat rank-k curve under the same intervention paradigm with pooled-mean range 0.20pp, and a random-h sensitivity floor lands at the same accuracy: the rank-k ablation alone conflates rank-blindness with position-irrelevance.
Chinese Translation
连续思维链(continuous chain-of-thought)模型将推理压缩为潜在 token。矩阵值变体通过一个 d x d 的矩阵瓶颈路由每个潜在 token,从而将秩引入为潜在矩阵 Z 上的单样本结构可观测量。如果矩阵潜在表示通过叠加(superposition)承载并行推理路径,那么秩应当能够追踪这些路径,并且将 Z 截断至低秩应当会损害那些其解答可能需要多个组件的任务的准确率。在对一个 Matrix-CODI 模型的四种训练设置下(三种在 ProsQA 上,一种在学习阈值以下的 GSM8K-Aug 上),秩-k 投影消融曲线平坦,波动不超过 0.6 个百分点。三次随机种子重复实验得到 81.0 +/- 2.0 个百分点的准确率,而 Z 的最终有效秩跨越 {4, 12, 13};损失函数并不偏好任何特定秩。为了检验秩盲性是否仅由“先展平再投影”的读出方式导致,我们训练了四种读出方式:双线性重参数化、对 Z 非线性的双线性加 GELU 读出、将奇异值输入 MLP 的 SVD 增强读出,以及基于 Z Z^T 的二次型读出。所有四种秩-k 曲线均保持平坦(Spearman p 值分别为 0.63、0.14、0.82、0.46)。平坦曲线在对 Z 非线性的读出方式下依然存在。针对 Z 的线性探针在目标预测上不及原始预训练隐藏状态(AUC 为 0.673 对 0.846)。在标准 GPT-2 SFT(无矩阵瓶颈、无 Z、三个随机种子、n=500)上的阴性对照在相同干预范式下复现了平坦的秩-k 曲线,合并均值范围为 0.20 个百分点,且随机-h 敏感度下限落在相同准确率上:秩-k 消融本身会将秩盲性与位置无关性混为一谈。
cs.LG / 16 / 2609.03100

Distilling deep optical flow stereo methods to retrieve dense three-dimensional wind fields

蒸馏深度光流立体方法以反演稠密三维风场
Vandal, Thomas J., Wu, Dong L., Carr, James L., Posselt, Derek J., Penn, Elise, Ballard, Tristan, Posch, August, Duffy, Kate
Abstract
Geostationary atmospheric motion vectors (AMVs) provide the dense horizontal wind vectors (u,v) and heights ingested into data assimilation systems. Traditional AMVs track features using window-based cross-correlation and estimate heights via infrared brightness temperatures paired with numerical weather prediction (NWP) background states, creating a circular dependency that yields inaccurate heights, high computational cost, and sparse retrievals. Stereo winds from GEO-GEO and GEO-LEO geometrically resolve heights from parallax shifts across different poses, eliminating NWP dependence and improving accuracy, but they remain computationally heavy with limited coverage. In this work, we replace window-based tracking in stereo matching with deep optical flow for efficient, improved retrieval. Fine-tuning balances a self-supervised geometric residual loss with supervised radiosonde reconstruction. To eliminate multi-satellite overlap requirements, we distill the stereo teacher into a single-satellite student model. Chi-square and height uncertainties from the teacher are emulated by the student for quality assurance. The student generates winds across full-disk GEO imagery globally. Validation compares stereo and student models against radiosondes, operational AMVs, ERA5 reanalysis, and EarthCARE cloud profiles. Results through triple collocation show that stereo winds improve performance beyond operational AMVs for water vapor bands (6.2, 6.9, and 7.3 {\mu}m), wit degradation in the long-wave infrared (11.2 {\mu}m) band.
Chinese Translation
静止轨道大气运动矢量(AMV)提供稠密的水平风矢量(u,v)及其高度,并被数据同化系统所吸收。传统AMV使用基于窗口的互相关方法跟踪特征,并通过红外亮温与数值天气预报(NWP)背景场配对来估计高度,这形成了一种循环依赖,导致高度不准、计算成本高以及反演稀疏。基于GEO-GEO和GEO-LEO几何构型的立体风场方法通过不同视角间的视差偏移从几何上解析高度,消除了对NWP的依赖并提高了精度,但其计算量依然庞大且覆盖范围有限。在本工作中,我们用深度光流取代立体匹配中基于窗口的跟踪,以实现高效且性能更优的反演。微调过程中,将自监督的几何残差损失与有监督的探空气球重建相平衡。为消除对多卫星重叠观测的需求,我们将立体教师模型蒸馏到单卫星学生模型中。教师模型的卡方和高度不确定性由学生模型进行模拟,用于质量保证。学生模型可在全球范围的静止卫星全圆盘图像上生成风场。验证工作将立体模型和学生模型与探空气球、业务化AMV、ERA5再分析资料以及EarthCARE云廓线进行了对比。三重共定位结果表明,立体风场在水汽通道(6.2、6.9和7.3微米)上的性能超越了业务化AMV,而在长波红外通道(11.2微米)上则有所下降。
cs.LG / 17 / 2609.03106

Scaling Laws, Tabular Data and Actuarial Ratemaking Models

缩放定律、表格数据与精算定价模型
Richman, Ronald
Abstract
Scaling laws in modern deep learning describe how held-out loss improves as model capacity, training data, and compute increase, often following power-law trends. We investigate whether analogous scaling regularities arise in actuarial ratemaking, where data are tabular, heterogeneous, and noisy, and where classical models such as GLMs remain strong baselines. Using a real-world motor insurance portfolio, we train models from different families across increasing fractions of the training data and multiple random seeds, evaluating out-of-sample Poisson deviance, a likelihood-based loss for Poisson count predictions in which lower values indicate better held-out fit. We find that all model families improve with additional data, but scaling exponents differ substantially: TabM exhibits markedly stronger data scaling than purely supervised tabular Transformers and standard MLP baselines. Transformer variants show weak parameter scaling unless augmented with additional inductive biases (TabM-style adaptation or self-supervision). These results provide quantitative guidance on model selection by data regime and suggest that effective scaling on actuarial tabular tasks depends on architecture and loss function objective design, with simple increases in Transformer size providing limited gains.
Chinese Translation
现代深度学习中的缩放定律描述了随着模型容量、训练数据和计算资源的增加,留出集损失如何改善,通常遵循幂律趋势。我们研究了类似的缩放规律是否出现在精算定价领域,该领域的数据是表格型、异质且含噪声的,且诸如广义线性模型(GLM)等经典模型仍然是强大的基线。我们使用一个真实的车险组合数据集,在不同训练数据比例和多个随机种子下训练来自不同类别的模型,并采用样本外Poisson偏差进行评估——这是一种基于似然的Poisson计数预测损失,数值越低表示留出集拟合效果越好。我们发现,所有模型类别均随数据增加而改善,但缩放指数差异显著:TabM 表现出明显强于纯监督表格型 Transformer 和标准 MLP 基线的数据缩放能力。Transformer 变体仅在引入额外的归纳偏置(如 TabM 风格的改造或自监督学习)时才表现出较弱的参数缩放能力。这些结果为不同数据规模下的模型选择提供了定量指导,并表明在精算表格任务上实现有效缩放取决于架构与损失函数目标设计,而单纯增大 Transformer 规模带来的收益有限。
cs.LG / 18 / 2609.03117

Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural Fields

核函数重启:突破神经切线核在神经场中的边界
Mallak, Amir, Maalouf, Alaa, Wolf, Lior, Rus, Daniela, Rosenbaum, Dan
Abstract
Neural fields (NFs) map continuous coordinates to signals such as color or density, but fast high-quality reconstruction from sparse observations remains difficult. Classical Neural Tangent Kernel (NTK) regression gives closed-form fits, yet it is fundamentally linear and cannot accumulate reusable task priors. We develop three algorithms that address these gaps. NTK-KIP learns a distilled support set of coordinates (and optional labels) so that a finite NTK can inpaint large missing regions from little observed data, yielding a compact non-linear representation instead of a raw kernel solve. MetaQuill meta-learns a shared initialization for an INR so that new scenes can be adapted by updating only a small task-specific weight offset, which provides true feature learning and a reusable prior. Finally, MetaQuill-KIP fuses both ideas: it seeds the task with a KIP-style non-linear warm start, then refines only that small offset around the meta-learned initialization. MetaQuill-KIP achieves high-PSNR reconstructions and semantically plausible inpainting under very sparse observations, while requiring only lightweight per-instance adaptation, whereas diffusion-style baselines typically depend on large pretrained generative priors and costly per-image tuning. This shows that NTK-driven neural fields can be made both non-linear and meta-learnable, narrowing the gap between analytic kernels and practical few-shot reconstruction.
Chinese Translation
神经场(Neural Fields, NFs)将连续坐标映射为颜色或密度等信号,但从稀疏观测中进行快速高质量的重建仍然困难。经典神经切线核(Neural Tangent Kernel, NTK)回归可以给出闭式解拟合,但其本质上是线性的,无法积累可复用的任务先验。我们开发了三种算法来弥补这些不足。NTK-KIP 学习一个蒸馏得到的坐标(及可选标签)支持集,使得有限维 NTK 能够从少量观测数据中对大面积缺失区域进行修补(inpainting),从而得到紧凑的非线性表示,而非直接求解原始核回归。MetaQuill 通过元学习为隐式神经表示(INR)学习一个共享初始化,使新场景只需更新一个较小的任务特定权重偏移即可完成适配,从而实现真正的特征学习和可复用的先验。最后,MetaQuill-KIP 融合了上述两种思想:先用 KIP 式的非线性热启动为任务提供初始种子,随后仅在元学习初始化附近微调该小偏移。MetaQuill-KIP 在极其稀疏的观测下实现了高 PSNR 重建和语义上合理的修补,且每个实例仅需轻量级适配;而扩散模型类基线方法通常依赖大规模预训练生成先验和昂贵的逐图像调优。这表明基于 NTK 的神经场可以同时具备非线性和可元学习性,从而缩小了解析核方法与实际少样本重建之间的差距。
cs.LG / 19 / 2609.03150

Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning

仅靠路由是不够的:诊断MoE+LoRA微调中适配器内部子空间竞争问题
Chowdhury, Mehreen Hossain, Mahjabin, Nowshin, Ruhan, Ahmed Shafin, Hossain, Md Azam, Kamal, Abu Raihan Mostofa, Laskar, Md Tahmid Rahman
Abstract
Multi-domain fine-tuning often combines MoE routing with LoRA, assuming that token-level routing separates domain-specific updates. We test this assumption in MoE+LoRA using Python code paired with biomedical text and mathematical reasoning. Although these domains show near-disjoint expert routing, adding biomedical data substantially increases code perplexity, indicating that routing separation alone may not prevent negative transfer. To localize the failure, we introduce Jaccard routing overlap and adapter-gradient cosine similarity, which measure expert sharing and update compatibility, respectively. These diagnostics indicate that interference arises mostly from nearly orthogonal domain gradients competing within the same low-rank adapter subspace. We address this issue with SpawnLoRA, which dynamically adds gated sub-adapters inside MoE experts when adapter-level contention is detected, while keeping the router fixed. We evaluate SpawnLoRA on Phi-tiny-MoE-instruct and OLMoE-1B-7B across multiple mixture settings and find that it effectively reduces negative transfer compared with standard and rank-adaptive LoRA. These results demonstrate that structural separation inside experts provides benefits beyond routing or rank expansion alone.
Chinese Translation
多领域微调通常将MoE路由与LoRA相结合,假设基于词元级别的路由能够分离领域特定的更新。我们在MoE+LoRA中检验了这一假设,使用Python代码配对生物医学文本和数学推理数据。尽管这些领域展现出近乎互斥的专家路由,但加入生物医学数据后,代码困惑度却显著上升,表明仅靠路由分离可能无法防止负迁移。为了定位这一失败原因,我们引入了Jaccard路由重叠度和适配器梯度余弦相似度两个诊断指标,分别用于度量专家共享程度和更新兼容性。这些诊断结果表明,干扰主要来源于几乎正交的领域梯度在同一低秩适配器子空间内的竞争。为解决这一问题,我们提出了SpawnLoRA,当检测到适配器层面的竞争时,该方法在MoE专家内部动态添加带门控的子适配器,同时保持路由器固定不变。我们在Phi-tiny-MoE-instruct和OLMoE-1B-7B上,跨多种混合数据设置对SpawnLoRA进行了评估,发现与标准LoRA和秩自适应LoRA相比,它能有效减少负迁移。这些结果表明,专家内部的结构性分离能够带来超越单纯路由或秩扩展的收益。
cs.LG / 20 / 2609.03177

Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings

前沿大语言模型是有效的批量优化器:在连续与离散场景中评估推理模型
Hu, Frank, Chennakesavalu, Shriram, Graff, David
Abstract
Frontier large language models (LLMs) have become attractive priors for optimization due to their large-scale pretraining that enables them to navigate a variety of optimization settings. However, the effectiveness of modern reasoning LLMs in batch optimization settings remains underexplored. Here we investigate the performance of the current generation of frontier LLMs as batch optimizers in both continuous and discrete settings. We find that while LLMs are competitive zero-shot batch optimizers for numerical test functions, their performance is brittle compared to classical non-LLM optimization approaches. However, LLM priors are significantly better in semantically rich settings, indicating that their batch optimization behavior is highly effective when navigating and reasoning over the discrete spaces most similar in structure to their pretraining data.
Chinese Translation
前沿大语言模型(LLMs)凭借其大规模预训练能力,能够应对多种优化场景,因而成为颇具吸引力的优化先验。然而,现代推理型大语言模型在批量优化场景中的有效性仍未得到充分探索。本文研究了当前一代前沿大语言模型在连续和离散场景中作为批量优化器的表现。我们发现,虽然大语言模型在数值测试函数上是具有竞争力的零样本批量优化器,但与经典非大语言模型优化方法相比,其表现较为脆弱。然而,在语义丰富的场景中,大语言模型先验显著更优,这表明当在与预训练数据结构最相似的离散空间中进行导航和推理时,其批量优化行为极为有效。
cs.LG / 21 / 2609.03180

Portable Causal Fairness Across Synthetic Data Generator Families

跨合成数据生成器家族的可移植因果公平性
Golob, Steven, Pentyala, Sikha, De Cock, Martine
Abstract
When a statistical agency or regulator releases synthetic data in place of sensitive records, it chooses the generator that produces the table, and can shape that generator so unfair pathways are absent. DECAF made this concrete on one non-private GAN: three fairness definitions become three sets of edge cuts on the generator's causal graph. Whether the mechanism belongs to DECAF, or to causal factorisation itself, was untested. We port all three definitions to nine generators from three unrelated families (marginals-based, GAN, and diffusion, each with differentially private variants), across three levels of formal privacy guarantee, over 2,520 matched-pair runs on Adult and COMPAS datasets. The mechanism transfers everywhere, and our new causal diffusion backbone yields the fairest release of any family we tested, at fidelity close to the marginals tier. Applying the cut barely moves fidelity, only costs a downstream classifier about $0.07$ to $0.15$ AUC on average, and adding privacy guarantees don't make the data less fair.
Chinese Translation
当统计机构或监管机构发布合成数据以替代敏感记录时,它需要选择生成数据表的生成器,并可以对生成器进行塑造以消除不公平的路径。DECAF在一个非隐私保护的GAN上使这一想法具体化:三种公平性定义转化为生成器因果图上的三组边剪切。然而,该机制是属于DECAF本身,还是属于因果分解(causal factorisation)这一更普遍的方法,此前尚未得到验证。我们将这三种定义移植到来自三个不相关家族(基于边缘分布的方法、GAN和扩散模型,各自包含差分隐私变体)的九个生成器上,覆盖三个层级的正式隐私保证,在Adult和COMPAS数据集上共进行了2,520次配对实验。结果表明,该机制在各处均可迁移,且我们新的因果扩散骨干网络在我们测试的所有家族中产生了最公平的数据发布,其保真度接近基于边缘分布方法的水平。应用边剪切几乎不影响保真度,对下游分类器平均仅造成约$0.07$至$0.15$的AUC损失,而增加隐私保证也不会使数据的公平性降低。
cs.LG / 22 / 2609.03229

Language-encoded network topology enables large language models to reason about complex networks

语言编码的网络拓扑使大语言模型能够对复杂网络进行推理
Utsha, Ucchwas Talukder, Mostafa, Sakib, Zou, James, Islam, Md Tauhidul
Abstract
Networks describe systems in biology and beyond, from protein interactions and social relationships to power grids and citation records. Reasoning about such systems requires understanding their structure: which elements are central, which connections bridge separate communities, and how it changes when elements are removed. Although large language models (LLMs) excel at natural language, they struggle with such questions when networks are given as edge lists, sentences or measurement tables, because their structural meaning must be inferred. Here we introduce BioGlyph, which compiles network topology into an interpretable and transferable language of structural roles. BioGlyph combines graph partitioning and structural measurements to identify roles such as hubs, community cores and cross-community connectors, and fixed rules to translate them into a universal vocabulary. The representation describes each element through its structural role, supporting evidence and semantic consequences, leaving both the network and the LLM unchanged. Across twenty networks spanning five domains, BioGlyph substantially improves open LLMs' ability to answer structural reasoning questions, outperforming edge-based, numerical and learned representations by up to 26 percentage points in system accuracy. Ablations show that the gain comes from explicitly encoding structural roles in semantically interpretable terms. The gain is more prominent in dense, community-structured networks and diminishes in sparse networks whose topology is more readily inferred from text. In a budding-yeast protein-interaction network, BioGlyph exposes biological organization: cross-community connectors are enriched for essential genes, whereas peripheral proteins are depleted. BioGlyph thus provides an interpretable representation for both language models and scientists to reason about network structure.
Chinese Translation
网络描述了生物学及其之外领域中的各类系统,从蛋白质相互作用和社会关系到电网和引文记录。对这类系统进行推理需要理解其结构:哪些元素处于中心地位,哪些连接桥接了不同的社区,以及移除元素后结构如何变化。尽管大语言模型(LLM)在自然语言方面表现出色,但当网络以边列表、句子或测量表格的形式给出时,它们难以回答这类问题,因为其结构含义必须被推断出来。本文提出BioGlyph,它将网络拓扑编译为一种可解释且可迁移的结构角色语言。BioGlyph结合图划分和结构度量来识别枢纽节点(hub)、社区核心和跨社区连接器等角色,并通过固定规则将其翻译为统一词汇。这种表示通过结构角色、支持证据和语义后果来描述每个元素,同时保持网络和LLM本身不变。在涵盖五个领域的二十个网络上,BioGlyph显著提升了开源LLM回答结构推理问题的能力,在系统准确率上比基于边的、数值型和习得的表示方法最高提升26个百分点。消融实验表明,这一增益来自于以语义可解释的术语显式编码结构角色。该增益在稠密且具有社区结构的网络中更为显著,而在稀疏网络中则减弱,因为稀疏网络的拓扑更容易从文本中推断。在芽殖酵母蛋白质相互作用网络中,BioGlyph揭示了生物学组织规律:跨社区连接器富集必需基因,而外围蛋白质则相对匮乏。因此,BioGlyph为语言模型和科学家推理网络结构提供了一种可解释的表示方法。
cs.LG / 23 / 2609.03231

The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100

2026年PNPL竞赛:LibriBrain100中的词汇分类与高效跨被试泛化
Mantegna, Francesco, Elvers, Gereon, Jayalath, Dulhan, Landau, Gilad, Kim, Tasha, Özdogan, Miran, Kurth, Luisa, Kwon, Teyun, Cho, SungJun, Ballyk, Benjamin, Fung, Alex, Greer, Anna, Somaiya, Pratik, Herff, Christian, Ramos, Yorguin Mantilla, Abdelhedi, Hamza, Jerbi, Karim, Farquhar, Greg, Shillingford, Brendan, Woolrich, Mark, Jones, Oiwi Parker
Abstract
The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-invasive speech decoding. Designed to progress from foundational tasks toward the linguistic complexity required for a practical brain-computer interface (BCI), it set the stage with speech detection and phoneme classification tasks. Winning submissions reached F1-macro scores of 95.6% and 73.6% on the respective tasks (Elvers et al., 2026), highly significant advances. This success was built on the LibriBrain dataset (\"{O}zdogan et al., 2025), the largest within-subject MEG dataset recorded at the time with ${\sim}50$ hours of data for one subject. However, while within-subject scale drives strong decoding performance, a practical BCI must generalise to new users from minutes of data, not hours. The 2026 PNPL competition responds to this challenge with LibriBrain100 (Mantegna et al., 2026), an extended LibriBrain dataset with 32 additional subjects (${\sim}40$ minutes each) plus even more within-subject data (${\sim}80$ hours). Advancing the curriculum of tasks to focus on word classification, two complementary tracks are presented in this competition: the Deep track targets within-subject word classification at scale, aiming at the best possible performance; the Broad track targets cross-subject generalisation, progressively reducing the amount of subject-specific fine-tuning data from ${\sim}40$ to ${\sim}20$ to ${\sim}10$ minutes, a duration that falls within a clinically feasible range and brings us a step closer to a non-invasive BCI capable of restoring communication to people living with profound paralysis.
Chinese Translation
2025年PNPL竞赛(Landau等人,2025)的宗旨是启动一项面向非侵入式语音解码的多年期课程计划。该竞赛旨在从基础任务逐步过渡到实用脑机接口(BCI)所需的复杂语言处理层面,并以语音检测和音素分类任务作为开端。各获胜方案在这两项任务上分别达到了95.6%和73.6%的F1-macro分数(Elvers等人,2026),取得了极其显著的进展。这一成功建立在LibriBrain数据集(Özdogan等人,2025)之上,该数据集是当时最大的单被试MEG(脑磁图)数据集,包含单个被试约50小时的数据。然而,尽管被试内数据规模的扩大能带来强大的解码性能,但一个实用的BCI必须能够在几分钟(而非几小时)的数据基础上泛化到新用户。2026年PNPL竞赛以LibriBrain100(Mantegna等人,2026)应对这一挑战——这是一个扩展版的LibriBrain数据集,新增了32名被试(每人约40分钟)的被试间数据,并进一步增加了被试内数据(约80小时)。本竞赛将任务课程推进至以词汇分类为核心,并设置了两条互补的赛道:深度赛道(Deep track)聚焦于大规模被试内词汇分类,追求最佳性能;广度赛道(Broad track)聚焦于跨被试泛化,将被试特异性微调数据量从约40分钟逐步减少到约20分钟再到约10分钟——这一时长处于临床可行范围之内,使我们朝着能够为深度瘫痪患者恢复交流能力的非侵入式BCI迈进了一步。
cs.LG / 24 / 2609.03239

B2B Customer Conversion Prediction: A Document Representation, Graph Theory, and CatBoost Driven Methodology

B2B客户转化预测:一种基于文档表示、图论与CatBoost驱动的方法
Wang, Tianqi, Azam, Sheikh Shams, Huang, Wan Eih, Wiranata, Anton, Brinton, Christopher G., Allebach, Jan P.
Abstract
In the one-time selling B2B context, the buying cycle may last months or even years. During the long process, targeting customers that have a high potential to make purchases and recommending personalized campaigns accordingly are important for effective marketing. For this goal, we study the following problems, B2B customer data aggregation, customer feature generation, and prediction of whether a B2B customer would show interest in making a purchase (i.e., prediction of conversion into sales funnel). We propose an algorithm to aggregate individual contacts to the B2B customer level based on multiple keys. For non-standardized keys such as company names, we propose a novel architecture to cluster them in a domain encompassing irregularities such as spelling mistakes and spelling variants. We then define and generate a set of features and apply the CatBoost model for customer conversion prediction. Our framework achieves 91\% prediction accuracy. Based on the prediction results and analysis of the model, we then discuss personalized campaign recommendations to foster conversion.
Chinese Translation
在一次性销售的B2B(企业对企业)场景中,购买周期可能持续数月甚至数年。在这一漫长的过程中,识别具有高购买潜力的客户并相应地推荐个性化的营销活动,对于高效营销至关重要。围绕这一目标,我们研究了以下问题:B2B客户数据聚合、客户特征生成,以及预测B2B客户是否会表现出购买意向(即预测其是否会转化为销售漏斗)。我们提出了一种基于多个键(key)将个体联系人聚合到B2B客户层面的算法。针对公司名称等非标准化的键,我们提出了一种新颖的架构,在一个领域内对它们进行聚类,该领域可涵盖拼写错误和拼写变体等不规则情况。随后,我们定义并生成了一组特征,并应用CatBoost模型进行客户转化预测。我们的框架达到了91%的预测准确率。基于预测结果和模型分析,我们进一步讨论了促进转化的个性化营销活动推荐。
cs.LG / 25 / 2609.03241

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

FlowBalance:基于验证器的策略内推理经验自我改进方法
Huang, Zixun, Panaganti, Kishan, Mi, Haitao, Liang, Leowei
Abstract
A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode. We introduce FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses. For each on-policy trajectory, a frozen training-time view of the same policy uses privileged context to produce token-level log-probability gains, which are aggregated into a trajectory-level self-guidance score. FlowBalance calibrates this score with the verifier-derived group advantage: guidance is retained on positive-advantage trajectories, reversed on negative-advantage trajectories, and disabled when the rollout group provides no outcome preference. The resulting energy exponentially reweights a reference policy, and profiled trajectory balance fits the normalized target with one log-partition estimate per rollout group. This realizes outcome-calibrated self-guidance via trajectory balance, without a separate token-level imitation loss. Our analysis establishes within-group contrast preservation, a minimum-change reverse-KL characterization, monotonic verifier control of target reward, and an exact correction against false-positive self-guidance on rejected responses. On mathematical reasoning, FlowBalance improves average performance over FlowRL on both Qwen3-4B and Qwen3-8B, while also improving training speed and stability, avoiding direct OPSD's response-length collapse, and exhibiting higher correct-strategy diversity in a controlled AIME24 diagnostic.
Chinese Translation
推理模型可以从自身的策略内(on-policy)经验中进行改进,但这一内循环机制较为脆弱:终端验证器提供可靠但稀疏的监督信号,而同模型的密集引导则可能强化错误的自信心,或将学习过度集中于狭窄的解模式。我们提出FlowBalance,一种基于验证器的自我改进方法,其通过学习完整响应上的归一化分布来工作。对于每条策略内轨迹,同一策略的一个冻结的训练时视图利用特权上下文生成词元级(token-level)对数概率增益,并将其聚合为轨迹级的自我引导分数。FlowBalance利用验证器导出的组优势(group advantage)来校准该分数:在正优势轨迹上保留引导,在负优势轨迹上反转引导,当采样组未提供结果偏好时则禁用引导。所得的能量函数对参考策略进行指数级重加权,并通过剖析的轨迹平衡(profiled trajectory balance)方法,对每个采样组仅需一次对数配分函数估计即可拟合归一化目标。这通过轨迹平衡实现了结果校准的自我引导,而无需单独的词元级模仿损失。我们的分析确立了组内对比保持性、最小变化的反向KL散度刻画、对目标奖励的单调验证器控制,以及针对被拒绝响应上假阳性自我引导的精确修正。在数学推理任务上,FlowBalance在Qwen3-4B和Qwen3-8B上的平均性能均优于FlowRL,同时提升了训练速度与稳定性,避免了直接使用OPSD时的响应长度坍缩问题,并在受控的AIME24诊断实验中展现出更高的正确策略多样性。
cs.LG / 26 / 2609.03265

Selective Hypergraph Refinement for Frozen Graph Clustering

面向冻结图聚类的选择性超图精化方法
Si, Zimo
Abstract
Existing graph-clustering methods typically improve clustering performance by optimizing model parameters and node representations. Effective means of further improving the clustering results of an already trained and frozen model, however, remain limited. We study post-processing for frozen graph clustering. After checkpoint fixation, the procedure uses no labels and updates neither model parameters, node representations, nor the original graph structure. Instead, it exploits an attribute hypergraph to supplement higher-order relations that ordinary graphs cannot readily express, thereby refining existing cluster assignments. Because global hypergraph refinement can yield both performance gains and erroneous updates, we propose Selective Hypergraph Refinement (SHR). The method generates candidate residual directions from the hypergraph and evaluates their reliability using graph structure, node attributes, and matched-null evidence. It updates only nodes with sufficient support and otherwise retains their original assignments. Further analysis shows that whether a node changes cluster is jointly governed by its native assignment gap and the directional strength of the refinement. In a controlled common-suite evaluation, 13 of 15 backbone-dataset cells had a positive mean macro gain, one produced exact no-action, and one was negative. The cell-equal macro gain was 0.066 pp (95% bootstrap CI, [0.030, 0.107] pp), while only 0.209% of hard assignments changed on average. A broader 15-combination native-interface evaluation yielded a macro gain of 0.137 pp at a mean change ratio of 0.375%. These results indicate that frozen clustering outputs retain a limited but measurable refinement space after training. The effect is heterogeneous across backbone-dataset pairs, and broader coverage also increases exposure to negative transfer.
Chinese Translation
现有的图聚类方法通常通过优化模型参数和节点表示来提升聚类性能。然而,对于已经训练完毕并冻结的模型,进一步改进其聚类结果的有效手段仍然有限。本文研究冻结图聚类的后处理问题。在模型检查点冻结之后,该过程不使用任何标签,也不更新模型参数、节点表示或原始图结构,而是利用属性超图来补充普通图难以表达的高阶关系,从而对已有的聚类分配结果进行精化。由于全局超图精化可能同时带来性能提升和错误更新,我们提出了选择性超图精化方法(Selective Hypergraph Refinement, SHR)。该方法从超图中生成候选残差方向,并利用图结构、节点属性以及匹配零模型(matched-null)证据评估其可靠性,仅更新具有充分支持度的节点,否则保留其原始聚类分配。进一步分析表明,节点是否改变聚类由其原生分配差距与精化的方向强度共同决定。在一组受控的通用测试套件评估中,15个“主干模型-数据集”组合中有13个取得了正向的平均宏观增益,1个结果为完全无操作,1个为负向。单元格等权的宏观增益为0.066个百分点(95% bootstrap置信区间为[0.030, 0.107]个百分点),而硬分配的平均变化率仅为0.209%。在更广泛的15种原生接口组合评估中,宏观增益达到0.137个百分点,平均变化率为0.375%。这些结果表明,冻结的聚类输出在训练后仍保留了有限但可度量的精化空间。该效应在不同“主干模型-数据集”组合间呈现异质性,且更广的覆盖范围也会增加负迁移的风险。
cs.LG / 27 / 2609.03294

Latent Energy Action Planning with World Models

基于世界模型的潜空间能量动作规划
Pham, Phu, Bera, Aniket
Abstract
Latent world models support efficient model predictive control from high-dimensional observations, yet optimizing a single learned latent objective can favor action sequences whose decoder-predicted terminal descriptor does not match the goal descriptor. We introduce Latent Energy Action Planning (LEAP), which treats the complete action horizon as a differentiable variable and optimizes it through a frozen LeWorldModel (LeWM). LEAP couples terminal latent goal matching with a terminal-window state energy. Low energy requires the predicted terminal latent to agree with the goal latent and the decoder-predicted terminal descriptor to agree with the goal descriptor. A frozen goal-conditioned proposal initializes the search, a quasi-Newton solver refines actions through the autoregressive rollout, and post-optimization projection enforces the admissible action range. Across four control domains using the officially released LeWM checkpoints, the complete LEAP planning system raises mean success from 77.5% for LeWM planned with the cross-entropy method (LeWM+CEM) to 94.8% under a matched protocol, a 17.3-percentage-point improvement, while retaining the frozen LeWM representation and predictor.
Chinese Translation
潜空间世界模型支持基于高维观测的高效模型预测控制,然而优化单一学习得到的潜空间目标可能偏向那些解码器预测的终端描述子与目标描述子不匹配的动作序列。我们提出了潜空间能量动作规划(Latent Energy Action Planning,LEAP),该方法将完整的动作序列视为可微分变量,并通过冻结的LeWorldModel(LeWM)对其进行优化。LEAP将终端潜空间目标匹配与终端窗口状态能量相结合。低能量要求预测的终端潜变量与目标潜变量一致,且解码器预测的终端描述子与目标描述子一致。一个冻结的目标条件提议模型用于初始化搜索,拟牛顿求解器通过自回归滚动优化动作,优化后的投影则保证动作处于可允许范围内。在使用官方发布的LeWM检查点的四个控制领域上,在相同协议下,完整的LEAP规划系统将平均成功率从使用交叉熵方法规划(LeWM+CEM)的77.5%提升至94.8%,提升了17.3个百分点,同时保留了冻结的LeWM表示和预测器。
cs.LG / 28 / 2609.03306

Geometry-Aware Graph Construction via Adaptive Spectral Bandwidth Control

基于自适应谱带宽控制的几何感知图构建
Bozkurt, Ecem, Ortega, Antonio
Abstract
Kernelized graph methods - spectral clustering, diffusion maps, and sparse kernel -regression graphs - that use Gaussian kernels depend on the choice of Gaussian bandwidth sigma, which governs the spectral character of the local kernel operator. When sigma is too small, the kernel overestimates local complexity and treats each sample as an independent direction; when sigma is too large, the kernel collapses multiple directions together, the condition number diverges, and all geometric discrimination is lost. We propose a choice of scale to make the spectral complexity of the kernel consistent with the intrinsic complexity of the underlying manifold. We propose a per-node bandwidth criterion that operationalizes this principle by jointly matching the kernel's effective rank to the local intrinsic dimension estimated via minimum spanning tree, anchoring the search in the manifold-consistent log-log scaling regime. We evaluate SSL embeddings from six encoders on CIFAR-100, showing that adaptive bandwidth consistently improves leave-one-out (LOO) classification and label propagation (LP) accuracy over fixed-bandwidth methods and competing adaptive methods.
Chinese Translation
基于高斯核的核化图方法——包括谱聚类、扩散映射和稀疏核回归图——依赖于高斯带宽 sigma 的选择,该参数决定了局部核算子的谱特性。当 sigma 过小时,核会高估局部复杂度,将每个样本视为独立方向;当 sigma 过大时,核会将多个方向合并坍缩,条件数发散,所有几何区分能力随之丧失。我们提出了一种尺度选择方法,使核的谱复杂度与底层流形的内在复杂度保持一致。我们提出了一个逐节点带宽准则来将这一原则操作化:通过将核的有效秩与基于最小生成树估计的局部内在维数相匹配,并将搜索锚定在流形一致的对数-对数缩放区域。我们在 CIFAR-100 数据集上评估了六个编码器的 SSL 嵌入,结果表明,自适应带宽方法在留一法(LOO)分类和标签传播(LP)准确率上始终优于固定带宽方法及其他竞争性自适应方法。
cs.LG / 29 / 2609.03308

Risk and Anomaly Identification for Distribution Network Optimal Operation Based on Reinforcement Learning and Uncertainty Quantification

基于强化学习与不确定性量化的配电网最优运行风险与异常识别
Zhang, Ziqi
Abstract
Reliable operation of modern distribution networks requires timely identification of operational risks and anomalous events under pervasive uncertainty. In practice, operators must identify risks that are inherent in stochastic yet in-distribution conditions, and anomalies that correspond to out-of-distribution behaviors such as unusual load patterns, extreme weather or cyber-physical attacks. This paper addresses this joint risk and anomaly identification problem for optimal distribution network operation and proposes a deep reinforcement learning framework that is explicitly uncertainty aware. We integrate distributional and Bayesian deep reinforcement learning to realize a second- order uncertainty quantification scheme that decomposes total uncertainty into aleatoric and epistemic components, which are respectively used to characterize inherent risk and out-of- distribution anomalies. The resulting epistemic estimates drive both exploration during training and out-of-distribution detec- tion with fallback control during deployment, whereas aleatoric estimates are used to characterize intrinsic operational risk. Simulation results demonstrate the performance of our DRL agent and the effectiveness of the uncertainty quantification.
Chinese Translation
现代配电网的可靠运行需要在普遍存在的不确定性下及时识别运行风险和异常事件。在实际中,运行人员既需要识别蕴含于随机但属于分布内(in-distribution)条件中的固有风险,也需要识别对应于分布外(out-of-distribution)行为的异常,例如异常负荷模式、极端天气或信息物理攻击。本文针对配电网最优运行中的风险与异常联合识别问题,提出了一种显式感知不确定性的深度强化学习框架。我们融合了分布式(distributional)与贝叶斯深度强化学习,实现了一种二阶不确定性量化方案,将总不确定性分解为偶然不确定性(aleatoric)与认知不确定性(epistemic)两个分量,分别用于刻画固有风险和分布外异常。所得的认知不确定性估计既用于训练过程中的探索,也用于部署阶段的分布外检测及回退控制;而偶然不确定性估计则用于刻画内在的运行风险。仿真结果验证了我们深度强化学习智能体的性能以及不确定性量化方法的有效性。
cs.LG / 30 / 2609.03324

DE-Venus: A Data-Efficient RLVR Framework for Large Language Models

DE-Venus:一个面向大语言模型的数据高效RLVR框架
Yang, Shenzhi, Zhu, Guangcheng, Tang, Kai, Zang, Zhengqing, Zheng, Xing, Wang, Haobo, Ma, Yingfan, Song, Bowen, Han, Bo, An, Bo, Feng, Lei, Wang, Weiqiang, Zhao, Junbo, Chen, Gang
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves large language model reasoning, but its practical scaling is constrained by expensive on-policy rollouts and the cost of obtaining reliable targets at scale. Existing methods address sample selection, incomplete supervision, or noisy labels separately, often entangling supervision logic with distributed training and hindering controlled comparison and reuse. We present DE-Venus, a unified framework for data-efficient RLVR that treats supervision as evolving state across data preparation and policy optimization. It organizes this lifecycle into three modules: Active Data Selection allocates training and annotation budgets; Weak Supervision Construction derives learning signals from unlabeled examples; and Training-Time Supervision Refinement filters or corrects unreliable supervision. DE-Venus supports seven representative methods and a data-selection pipeline by expressing method-specific decisions as offline dataset transitions or online transformations of targets, rewards, batches, and advantages while preserving verl's distributed execution contracts. Across public benchmarks and three business scenarios, separate configurations preserve or improve model quality with only 10% of labels or as little as 13% of relevant data; selected business configurations also reduce observed convergence steps by 63%--75%. DE-Venus thus reduces annotation and training costs without sacrificing scalable RL execution.
Chinese Translation
可验证奖励强化学习(RLVR)能够提升大语言模型的推理能力,但其实际扩展受到昂贵的策略内采样(on-policy rollouts)以及大规模获取可靠目标成本的限制。现有方法分别针对样本选择、不完整监督或噪声标签等问题进行处理,往往将监督逻辑与分布式训练耦合在一起,阻碍了受控比较与复用。我们提出DE-Venus,一个统一的数据高效RLVR框架,它将监督视为在数据准备与策略优化过程中不断演化的状态。该框架将这一生命周期组织为三个模块:主动数据选择负责分配训练与标注预算;弱监督构建从无标注样本中推导学习信号;训练时监督优化对不可靠的监督进行过滤或纠正。DE-Venus通过将方法特定的决策表示为离线数据集转换,或对目标、奖励、批次和优势的在线变换,同时保留verl的分布式执行契约,支持了七种代表性方法及一条数据选择流水线。在公开基准和三个业务场景中,各独立配置仅使用10%的标签或少至13%的相关数据即可保持或提升模型质量;选定的业务配置还将观测到的收敛步数减少了63%–75%。因此,DE-Venus在不牺牲可扩展RL执行的前提下,降低了标注与训练成本。
cs.LG / 31 / 2609.03337

A Large Open Multi-Energy Corpus of Soil Compaction Tests, with Machine-Learning Baselines

一个大型开放的多能量土压实试验语料库及机器学习基线
Youwai, Sompote, Phutthananon, Chana, Kongkitkul, Warat
Abstract
Every engineered fill is specified by a maximum dry density and an optimum moisture content. Each determination needs a full Proctor test. Published correlations rest on one to four hundred specimens, usually from one laboratory at one compactive energy, and are seldom released. This paper releases a corpus without those limits. It holds 2,854 laboratory compaction tests from six public sources, across 162 provenance groups and four Proctor energy levels, with fines from 1.5 to 100%. Every record is audited to the Proctor method its source names, and no energy is inferred. Screening on the zero-air-voids condition removed 11.8% of harmonised records, and 5.7% of those with a measured specific gravity. A material share of published compaction data is physically impossible. The optimum degree of saturation over the corpus is 0.815 at a coefficient of variation of 11%. That is a baseline, not a constant. Both parameters are then estimated from one classification suite and the compaction standard. A tabular foundation model reaches R2 0.824 for density and 0.784 for water content under random folds. It reaches 0.727 and 0.696 with folds drawn around provenance, and 0.520 and 0.614 with a whole source held out. Compactive energy is negligible marginally yet decisive conditionally. Density on the 66 modified-Proctor records is predicted at R2 0.740 with it and -0.651 without. Symbolic regression yields closed forms coupled through a phase relation. No predicted pair can then exceed the zero-air-voids line. The predictions are for screening, not acceptance.
Chinese Translation
每一项填土工程都以最大干密度和最优含水率来规定,而每次确定都需要进行完整的普氏(Proctor)试验。已发表的相关关系通常基于一百至四百个试样,且多来自单一实验室、单一压实功能级,并很少对外公开。本文发布了一个不受上述限制的语料库,其中包含来自六个公开来源的2,854个室内压实试验数据,覆盖162个来源组(provenance groups)和四个普氏能量等级,细粒含量介于1.5%至100%之间。每条记录均按其来源所注明的普氏方法进行审核,且未对能量水平进行任何推断。基于零孔隙比(zero-air-voids)条件的筛查剔除了11.8%的协调化记录,以及具有实测比重记录中的5.7%。这表明相当一部分已发表的压实数据在物理上是不可能的。整个语料库的最优饱和度为0.815,变异系数为11%。这是一个基线,而非一个常数。随后,通过一个分类指标集与压实标准对这两个参数进行估计。一个表格基础模型(tabular foundation model)在随机折交叉验证下,干密度的R²达到0.824,含水率达到0.784;在按来源划分折时,R²分别为0.727和0.696;在留出一个完整来源时,R²分别为0.520和0.614。压实能在边际上可忽略,但在条件上具有决定性作用:在66个修正普氏(modified-Proctor)记录上,纳入该变量时密度预测R²为0.740,剔除后为-0.651。符号回归得到了通过相态关系耦合的闭式表达式,由此任何预测的参数对都不会超过零孔隙比线。这些预测适用于筛查,而非验收。
cs.LG / 32 / 2609.03342

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

梯度知晓结果所不知:利用梯度对齐奖励解锁大语言模型推理的强化学习
Zheng, Leqi, Su, Jinbo, Niu, Fang, Wang, Chaokun, Wang, Weiping, Zhang, Jiajun, Yan, Shannan, Wu, Jie, Kang, Zhaolu, Fu, Rong, Zhang, Hang
Abstract
Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from surface heuristics to process reward models, either ignore the expert solutions already present in training corpora or require expensive offline annotation. We propose Gradient-Aligned Reward (GAR), which operates in the policy's own gradient space: truncated backpropagation through the output projection layer extracts a compact gradient vector for each rollout, and cosine similarity with an expert-anchor gradient yields a dense, reasoning-aware reward with less than 9% wall-clock overhead. We prove that this cosine admits a multiplicative decomposition into prediction-error and activation-pattern factors, providing a concrete characterization of what the alignment signal measures. On Qwen3-4B and Qwen3-8B, GAR consistently improves over GRPO and other baselines on competition-level math benchmarks and transfers to GPQA Diamond and MMLU-Pro without domain-specific data. Code and data are available at https://github.com/LQgdwind/GAR.
Chinese Translation
基于可验证奖励的强化学习(RLVR)推动了大语言模型的思维链推理,但其二元结果奖励无法区分不同的正确轨迹。现有的稠密奖励替代方案,从表面启发式方法到过程奖励模型,要么忽略了训练语料中已有的专家解答,要么需要昂贵的离线标注。我们提出梯度对齐奖励(Gradient-Aligned Reward, GAR),它在策略自身的梯度空间中运行:通过输出投影层进行截断反向传播,为每条采样轨迹提取一个紧凑的梯度向量,然后与专家锚点梯度的余弦相似度产生稠密的、具备推理感知能力的奖励,额外的时间开销不足9%。我们证明该余弦相似度可以分解为预测误差和激活模式两个因子的乘积,从而具体刻画了对齐信号所衡量的内容。在 Qwen3-4B 和 Qwen3-8B 上,GAR 在竞赛级数学基准上一致优于 GRPO 及其他基线方法,并且无需领域特定数据即可迁移至 GPQA Diamond 和 MMLU-Pro。代码和数据见 https://github.com/LQgdwind/GAR。
cs.LG / 33 / 2609.03350

From Zero to Hero: An Open LLM Ecosystem for Armenian

从零到英雄:面向亚美尼亚语的开放大语言模型生态系统
Arakelyan, Erik, Avetisyan, Khatun, Davtyan, Meri, Grigoryan, Heghine, Khachatryan, Nane, Shahsuvaryan, Hayk, Sergoyan, Henrik, Martirosyan, Vahan
Abstract
Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armenian LLM has been released with the data and recipe needed to reproduce it. To address this gap, we curate and release two datasets. ArmWeb is an extensively validated corpus of 4.37M Armenian news documents. ArmSTEM is a parallel English-Armenian collection of 373K math and science problems with step-by-step solutions, translated into Armenian and verified through both answer-preserving LLM judgment and human evaluation. Continued pretraining of Gemma-4-E4B on these datasets yields arm-gemma-e4b, which outperforms every existing open Armenian model as well as its unadapted base, and is the first open Armenian LLM with complete training data and recipe. Our ablations show that news-only continued pretraining improves fluency while eroding knowledge, a pattern we also observe in existing Armenian models, and that a small share of verified translated STEM data reverses the loss. We further find that the largest public Armenian corpora overlap web-derived evaluation panels heavily, including a train/test self-overlap inside FineWeb-2. We openly release all data, models, and code.
Chinese Translation
亚美尼亚语是一种形态丰富但资源稀缺的语言,其预训练数据十分匮乏,且目前尚未有公开的亚美尼亚语大语言模型(LLM)附带可复现所需的数据和训练方案。为填补这一空白,我们构建并发布了两个数据集:ArmWeb 是一个经过广泛验证、包含 437 万篇亚美尼亚语新闻文档的语料库;ArmSTEM 是一个英亚平行数据集,包含 37.3 万道数学与科学题目及其逐步解答,这些内容被翻译成亚美尼亚语,并通过保留答案的 LLM 判定与人工评估进行双重验证。在上述数据集上对 Gemma-4-E4B 进行继续预训练,得到了 arm-gemma-e4b,其性能超越了所有现有的开放亚美尼亚语模型以及未经适配的基础模型,并且是首个具备完整训练数据和训练方案的开放亚美尼亚语大语言模型。我们的消融实验表明,仅使用新闻数据进行继续预训练虽能提升语言流畅度,但会导致知识退化——我们在现有亚美尼亚语模型中也观察到了这一模式——而引入少量经过验证的翻译 STEM 数据可以扭转这一损失。我们进一步发现,目前最大的公开亚美尼亚语语料库与基于网页构建的评估集存在严重重叠,包括 FineWeb-2 内部存在的训练/测试自重叠现象。我们公开发布了所有数据、模型和代码。
cs.LG / 34 / 2609.03358

Time Without Timesteps: Simulating Coupled Dynamical Systems via Self-Consistency

无需时间步的时间:基于自洽性求解耦合动力系统
Zerihun, Liyu, Lee, Mark Shinyoung
Abstract
Numerical simulation of dynamical systems is usually organized as a causal march through time: each state is computed from the previous one. We explore a different formulation for coupled systems. For each subsystem type we train a neural surrogate mapping a full driving trajectory and initial condition directly to a full output trajectory; following classical waveform relaxation, coupled systems are assembled by enforcing self-consistency among these trajectories: simulation becomes a fixed-point problem over complete trajectories rather than a stepwise rollout. On coupled van der Pol oscillators and Hodgkin-Huxley neuron networks, sequential depth becomes the number of solver iterations: 4-10 Newton iterations where the reference integrator takes 1500 steps. The gradient likewise loses its time recursion: it becomes a linear system solved by GMRES at memory independent of solver depth. A single scalar measured from the learned operator, the spectral radius of its Jacobian, predicts in advance where the coupled solve will converge; past that boundary, unrolled backpropagation diverges and a Neumann adjoint fails, while the implicit gradient remains correct to 0.04%. We report where the approach succeeds and where surrogate error degrades it.
Chinese Translation
动力系统的数值模拟通常被组织为沿时间轴的因果推进:每个状态由前一个状态计算得到。我们为耦合系统探索了一种不同的表述方式。对于每种子系统类型,我们训练一个神经代理模型,将完整的驱动轨迹和初始条件直接映射到完整的输出轨迹;借鉴经典的波形松弛方法(waveform relaxation),通过在这些轨迹之间强制自洽性来组装耦合系统:模拟由此变为对完整轨迹的不动点问题,而非逐步推进。在耦合范德波尔振子(van der Pol oscillators)和霍奇金-赫胥黎(Hodgkin-Huxley)神经元网络上,串行深度变成了求解器迭代次数:参考积分器需要1500步的问题,该方法仅需4-10次牛顿迭代。梯度同样摆脱了时间递归:它成为一个由GMRES求解的线性系统,其内存开销与求解器深度无关。一个从所学算子测得的标量——其雅可比矩阵的谱半径——能够提前预测耦合求解将在何处收敛;超出该边界后,展开式反向传播发散,Neumann伴随方法失效,而隐式梯度仍保持0.04%以内的精度。我们报告了该方法成功的场景以及代理误差使其性能退化的情况。
cs.LG / 35 / 2609.03377

SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

SimpleDesign:蛋白质序列与结构协同设计的联合模型
Lu, Jiarui, Wang, Yuyang, Zhang, Yizhe, Gu, Jiatao, Jaitly, Navdeep, Susskind, Joshua M., Bautista, Miguel Ángel
Abstract
Proteins are fundamental to biological processes, with their function determined by the complex interplay between the amino acid sequence and the three-dimensional structure. Developing generative models capable of understanding this intrinsically multi-modal relationship is crucial for fields like drug discovery and protein engineering. Existing models often rely on a multi-stage training process where autoencoders that tokenize data into latent representations are trained in a first stage. Secondly, a generative model is trained on the latent representation of the autoencoder(s), i.e., generative modeling in a latent space. We hypothesize that this multi-stage training is not necessary to obtain performant co-design models and thus present SimpleDesign, an effective multi-modal protein design model trained directly in the data space. SimpleDesign leverages a single-stage end-to-end objective that combines discrete cross-entropy for sequences and a regression objective for structures. In order to effectively model the difference in sequence and structure modalities, we develop a Mixture-of-Transformer architecture that allows modality-specific processing while keeping global self-attention over both modalities. We train SimpleDesign on over 2M sequence-structure pairs achieving strong performance across co-design and unconditional sequence/structure generation benchmarks.
Chinese Translation
蛋白质是生物过程的基础,其功能由氨基酸序列与三维结构之间复杂的相互作用决定。开发能够理解这种内在多模态关系的生成式模型,对于药物发现和蛋白质工程等领域至关重要。现有模型通常依赖于多阶段训练过程:第一阶段训练将数据标记化为潜在表示的自编码器;其次,在自编码器的潜在表示上训练生成模型,即在潜在空间中进行生成建模。我们假设这种多阶段训练对于获得高性能的协同设计模型并非必要,因此提出了SimpleDesign——一个直接在数据空间中训练的高效多模态蛋白质设计模型。SimpleDesign采用单阶段端到端目标函数,结合了针对序列的离散交叉熵和针对结构的回归目标。为了有效建模序列与结构模态之间的差异,我们开发了一种Mixture-of-Transformer(混合Transformer)架构,在保持对两种模态进行全局自注意力处理的同时,实现针对特定模态的处理。我们在超过200万个序列-结构对上训练了SimpleDesign,在协同设计以及无条件序列/结构生成基准测试中均取得了优异的性能。
cs.LG / 36 / 2609.03379

RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory

RecurTrace:基于循环时间记忆的自适应潜在推理
Wang, Yuxiang, Feng, Kunyu, Shen, Yingda, Xu, Haoning, Wang, Junyu, Wu, Zhizheng
Abstract
Repeating a small block of middle layers increases a language model's effective inference depth without adding parameters or generating extra tokens, and recent work shows that this latent recurrence improves reasoning. However, two design choices limit these gains. Each iteration sees only the previous output and cannot directly access earlier computations. Moreover, a fixed loop count wastes depth on easy inputs while leaving hard ones with too little computation. We introduce RecurTrace, which addresses both limitations using the loop's own trajectory. Specifically, Loop Memory Attention lets each looped layer attend to its own states from previous iterations along the loop-time axis, so the model can revisit earlier computations instead of relying on the latest state alone. A halting head then reads the loop state and predicts whether to continue, with supervision from an oracle that identifies when additional depth still reduces loss. In a controlled MathQA comparison on the same looped backbone, RecurTrace achieves 56.9% accuracy with an average of 2.0 loops, exceeding the best fixed loop depth by 2.2 points at matched compute. By comparison, ACT and PonderNet collapse to one loop, and CALM reaches only 54.1% with 5.6 loops, while the stronger LoopUS-Conf and TaH-Mismatch baselines reach 55.3% at 3.2 loops and 55.7% at 2.1 loops. Finally, RecurTrace improves generation accuracy over same-budget fine-tuned baselines at 0.6B, 1.7B, 4B, and 8B, with the gain growing with model size from 0.6 to 3.4 points.
Chinese Translation
重复使用一小段中间层块可以在不增加参数或生成额外token的情况下提升语言模型的有效推理深度,且近期研究表明这种潜在递归(latent recurrence)能够改善推理能力。然而,两个设计选择限制了这些收益:每次迭代只能看到上一次的输出,无法直接访问更早的计算;此外,固定的循环次数在简单输入上浪费了深度,而在困难输入上计算量又不足。我们提出 RecurTrace,利用循环自身的轨迹同时解决这两个局限。具体而言,循环记忆注意力(Loop Memory Attention)使每个循环层能够沿循环时间轴关注其先前迭代中的自身状态,从而模型可以回溯更早的计算,而不仅依赖于最新状态。随后,一个停转头(halting head)读取循环状态并预测是否继续循环,其监督信号来自一个能够判断何时增加深度仍能降低损失的oracle。在相同循环骨干网络上的受控MathQA对比中,RecurTrace以平均2.0次循环达到了56.9%的准确率,在同等计算量下超过最优固定循环深度2.2个百分点。相比之下,ACT和PonderNet坍缩为一次循环,CALM在5.6次循环下仅达到54.1%,而更强的LoopUS-Conf和TaH-Mismatch基线分别在3.2次循环和2.1次循环下达到55.3%和55.7%。最后,在0.6B、1.7B、4B和8B模型上,RecurTrace在同等预算下均优于微调基线的生成准确率,且增益随模型规模从0.6增长到3.4个百分点。
cs.LG / 37 / 2609.03383

TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents

TIGPO:面向长时程LLM智能体的时序实例图策略优化
Gan, Jinwei
Abstract
Graph-based policy optimization improves credit assignment for long-horizon LLM agents by organizing rollout trajectories into state-transition graphs. However, existing methods construct graphs independently within each policy update, discarding transitions discovered by earlier policies and limiting advantage estimation to small, batch-local rollout groups. We propose \emph{Temporal Instance-Graph Policy Optimization} (TIGPO), which extends graph-based credit assignment across policy updates. TIGPO maintains a persistent transition graph for each task, allowing valid transitions discovered by different policy versions to jointly determine credit for current rollouts. To actively reconnect current exploration with historical experience, TIGPO allocates a fixed rollout budget between Exploration slots for ordinary task sampling and Revisit slots for delayed reattempts of previously explored tasks. For each revisit, TIGPO pairs the current rollout group with its corresponding earlier Exploration group to construct a cross-temporal reference. The enlarged reference is designed to stabilize relative advantage estimation under small rollout groups, while comparison on the same task directly captures policy improvement across training stages. Historical transitions and scores serve only as structural and detached statistical references and are never replayed in the policy loss. Experiments on ALFWorld and WebShop demonstrate that TIGPO consistently outperforms prior group-based and graph-based policy optimization methods.
Chinese Translation
基于图的策略优化通过将滚动轨迹组织为状态转移图,改进了长时程LLM智能体的信用分配。然而,现有方法在每次策略更新中独立构建图,丢弃了先前策略发现的转移,并将优势估计限制在较小的批内局部滚动组中。我们提出时序实例图策略优化(TIGPO),将基于图的信用分配扩展到跨策略更新。TIGPO为每个任务维护一个持久化的转移图,使不同策略版本发现的有效转移能够共同决定当前滚动的信用。为了主动地将当前探索与历史经验重新连接,TIGPO将固定的滚动预算分配到用于常规任务采样的探索槽和用于延迟重试先前已探索任务的重访槽之间。对于每次重访,TIGPO将当前滚动组与其对应的早期探索组配对,构建跨时序参考。这一扩大的参考旨在稳定小滚动组下的相对优势估计,而对同一任务的比较则直接捕捉训练阶段间的策略改进。历史转移和分数仅作为结构性和detach的统计参考,绝不在策略损失中重放。在ALFWorld和WebShop上的实验表明,TIGPO持续优于已有的基于组和基于图的策略优化方法。
cs.LG / 38 / 2609.03422

Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models

推断的生成过程多样性可预测语言模型间的相关失效
Tieman, Ross, Markou, Evan
Abstract
Diversity is a widely observed factor in the resilient function of collective systems, yet the type of diversity that matters depends on the properties and failure modes of the system. This distinction is important for systems composed of multiple language models. Different models may be treated as independent components even when their behaviour and failures remain strongly correlated. Assessments of language-model populations using semantic similarity demonstrate limited semantic diversity, but this captures only differences in the meaning of observed outputs. We argue that a more fundamental notion of model diversity is generative-process diversity, the differences between processes capable of generating the observed outputs. Drawing from Algorithmic Information Theory, we use Normalised Compression Distance between raw model outputs, residualised against a permutation control, as a measure of inferred generative-process diversity. Across 38 language models, this measure identifies population structure missed by semantic similarity and predicts cross-task variation in chance-corrected correlated failure among model pairs across ten disjoint benchmark families, beyond semantic similarity and model-pair capability. The cross-benchmark partial rank association is $-0.216$ with a 95% interval of $[-0.309,-0.122]$, and the estimate is negative on all ten benchmarks. These results indicate that increased generative-process diversity is associated with reduced correlated failure in model pairs that is not attributable to semantic similarity or capability. Inferred generative-process diversity offers a novel and practical approach for investigating diversity of multi-model systems in safety-relevant contexts.
Chinese Translation
多样性是集体系统韧性运作中被广泛观察到的因素,但何种类型的多样性真正重要取决于系统的性质及其失效模式。这一区分对于由多个语言模型构成的系统尤为重要。即使不同模型的行为和失效仍然高度相关,它们也可能被当作独立组件来对待。基于语义相似度的语言模型群体评估表明模型间的语义多样性有限,但这仅反映了观测输出在含义上的差异。我们认为,一个更根本的模型多样性概念是生成过程多样性,即能够生成观测输出的各过程之间的差异。借鉴算法信息论(Algorithmic Information Theory),我们使用原始模型输出之间的归一化压缩距离(Normalised Compression Distance),并以置换控制进行残差化处理,作为推断的生成过程多样性的度量。在38个语言模型上,该度量识别出了语义相似度所遗漏的群体结构,并在十个互不相交的基准测试族中,预测了模型对之间经机会校正的相关失效的跨任务变异,其预测能力超越了语义相似度和模型对能力。跨基准的部分秩相关系数为 $-0.216$,95% 置信区间为 $[-0.309,-0.122]$,且该估计值在全部十个基准上均为负。这些结果表明,生成过程多样性的提高与模型对相关失效的减少相关,且这一效应不能归因于语义相似度或模型能力。推断的生成过程多样性为在安全相关情境下研究多模型系统的多样性提供了一种新颖而实用的方法。
cs.LG / 39 / 2609.03427

TraveL: Transformer-based Multi-view Path Distributional Representation Learning

TraveL:基于Transformer的多视角路径分布式表示学习
He, Fang, Fu, Tao-yang, Lee, Wang-chien
Abstract
Path representation learning (PRL) for road networks has received increasing research attention, due to various path-related applications. Existing works on PRL typically exploit the co-occurrence relationship among road segments and paths to learn a vector as the path representation, without exploring the varied traveler behaviors and the regional correlation on the path. In this work, we propose to learn distributional representations, which provide valuable information for use in path-related applications, by capturing the varied traveler behaviors as well as the various dependencies within regions of road segments. We propose a novel Transformer-based Multi-view Distributional Representation Learning (TraveL) framework to encode a path along with a travel starting time to a distributional representation, which can be used to decode possible samples of on-path traveler behavior. Moreover, by analyzing the regional correlation which reveals various road segment relationships, we propose a regional attention to encode these correlations in a path. Also, we explore the idea of Kolmogorov-Smirnov (K-S) test to compare the sampled traveler behavior against the collected ground truth to facilitate training. Experimental results show that the proposed TraveL model outperforms the state-of-the-art methods on both synthetic and real-world datasets, by 14.7% in Mean K-S distance for travel time distribution estimation, 16.7% in Mean Absolute Error (MAE) for path similarity prediction, and 3.97% in MAE for destination prediction.
Chinese Translation
由于各种与路径相关的应用需求,面向道路网络的路径表示学习(PRL)受到了越来越多的研究关注。现有的PRL工作通常利用路段与路径之间的共现关系来学习一个向量作为路径表示,而未探索旅行者行为的多样性以及路径上的区域相关性。在本工作中,我们提出学习分布式表示,通过捕捉旅行者行为的多样性以及路段所在区域内的各种依赖关系,为路径相关应用提供有价值的信息。我们提出了一种新颖的基于Transformer的多视角分布式表示学习框架(TraveL),将路径与出行起始时间一同编码为分布式表示,该表示可用于解码路径上旅行者行为的可能样本。此外,通过分析揭示各种路段关系的区域相关性,我们提出了一种区域注意力机制来编码路径中的这些相关性。同时,我们探索了Kolmogorov-Smirnov(K-S)检验的思想,将采样的旅行者行为与收集的真实数据进行比较,以辅助模型训练。实验结果表明,所提出的TraveL模型在合成数据集和真实数据集上均优于最先进的方法:在旅行时间分布估计的平均K-S距离上提升了14.7%,在路径相似性预测的平均绝对误差(MAE)上提升了16.7%,在目的地预测的MAE上提升了3.97%。
cs.LG / 40 / 2609.03436

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

问题所在,而非路径:LLM推理轨迹中的预算与难度混淆因素
Bulut, Yigit Utku
Abstract
Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legible fates. Both readings rest on measurements missing a counterfactual control at the level of the claim; we supply both controls. First, a restart-controlled truncation probe separates when a solution fits the continuation budget from when a prefix carries value that fresh computation cannot buy, comparing per-anchor continuation solve rates against from-scratch restart curves at matched total generated-token budget. Applied to 178 problem-model cells (89 MATH problems x two small open models, an outcome-blind but difficulty-targeted cohort), exactly 1 of 178 cells survives as prefix-limited; restart dose-response separates a compute-starved model from a capability-limited one; and wherever the matched budget lies inside the restart grid, continuing the model's own prefix beats restarting (9 of 9) -- predominantly compute compression rather than expanded reachability. Second, a pre-registered, difficulty-controlled test finds no detectable outcome information in early-window internal signals beyond a problem-difficulty baseline, and two generation-free analyses of public corpora show why this control is needed: a trace-blind difficulty proxy reaches AUROC 0.873 on 192K DeepSeek-R1 generations -- inside the published probe range -- and a closely matched reconstruction of the closest published early-window positive recovers a comparable pooled result (0.849) while within problem it is statistically indistinguishable from chance at all ten anchors (0.496 at t=4); a post-hoc within-targeted probe finds only a small average residual, concentrated in three low-failure problems. High pooled probe AUROCs cannot by themselves establish within-attempt information; a question-only baseline or within-problem evaluation is required.
Chinese Translation
大语言模型的推理轨迹被广泛解读为包含"突破"时刻以及早期可判定的命运。这两种解读都依赖于缺失声明层面反事实控制的测量;我们提供了这两项控制。首先,重启对照的截断探针(restart-controlled truncation probe)将"解法符合继续生成预算"与"前缀带有全新计算无法买到的价值"区分开来,方法是在相同的总生成token预算下,比较各锚点的继续求解率与从头重启的曲线。应用于178个问题-模型单元(89道MATH问题 × 两个小型开源模型,一个结果盲但以难度为目标的队列),178个单元中恰好只有1个作为前缀受限单元存活;重启剂量-反应关系区分了计算匮乏模型与能力受限模型;且在匹配预算位于重启网格内的所有情形下,继续模型自身的前缀优于重启(9/9)——这主要是计算压缩而非可达范围的扩展。其次,一项预先注册的、难度控制的测试发现,除了问题难度基线之外,早期窗口内部信号中不存在可检测的结果信息;两项无需生成的公共语料库分析说明了为何需要这一控制:一个轨迹盲的难度代理在192K条DeepSeek-R1生成上达到AUROC 0.873——落在已发表的探针范围内——而对已发表的最接近早期窗口阳性结果的紧密匹配重建获得了相当的汇总结果(0.849),但在问题内部,它在全部十个锚点上与随机水平在统计上不可区分(t=4时为0.496);一项事后在目标问题内的探针仅发现较小的平均残差,且集中在三个低失败率问题上。高汇总探针AUROC本身无法确立单次尝试内部的信息;需要仅基于问题的基线或问题内评估。
cs.LG / 41 / 2609.03442

Guide, Not Bind: Why Defeasible Priors Fail in Augmented Lagrangian Causal Discovery

引导而非束缚:可废止先验为何在增广拉格朗日因果发现中失效
Sundararaman, Sairam, Girdhar, Sara, Murthy, Manit Narasimha, N, Samrudh, Das, Bhaskarjyoti
Abstract
Differentiable causal discovery methods increasingly encode expert priors as forbidden-edge constraints enforced by an Augmented Lagrangian (ALM) penalty, on the assumption that a data-adaptive relaxation mechanism will discount and eventually override a rule the data consistently contradicts. We show this design, which we call \emph{guide, not bind}, fails for two independent, precisely characterized reasons, and that directly repairing both restores it only partially. First, sequential penalty-ramping ALM suppresses a wrongly-forbidden true edge before any counterfactual check can detect it: we give three necessary conditions any adaptive relaxation must satisfy to avoid this (Proposition~\ref{prop:conditions}), prove that DADU---the natural relaxation rule this paper introduces as the object of study---violates all three (Corollary~\ref{cor:dadu_failure}), and confirm the failure across 3{,}072 training runs spanning graphs from 4 to 32 nodes, where a single wrong prior suppresses a true edge in 87--97\% of trials under DADU. Second, and independent of any fix to the mechanism, we prove in closed form that the standard correlation-matching objective ties a true edge and its reverse to an identical cost of exactly $2r^2$ (Lemma~\ref{lem:tie}), not because the underlying equal-variance model is unidentifiable, but because normalizing to correlation discards exactly the variance information that would make it identifiable; covariance matching instead separates the two directions by a provable margin of at least $w_0^4$ (Lemma~\ref{lem:separation}).
Chinese Translation
可微因果发现方法日益将专家先验编码为禁止边约束,并通过增广拉格朗日方法(Augmented Lagrangian, ALM)惩罚项加以强制执行,其假设是数据自适应的松弛机制会对数据持续相悖的规则进行折扣并最终将其覆盖。我们证明这种我们称之为“引导而非束缚”(guide, not bind)的设计会因两个相互独立且被精确刻画的原因而失效,且直接同时修复两者也只能部分恢复其功能。首先,顺序式的惩罚递增ALM在任何反事实检验能够检测之前就抑制了被错误禁止的真实边:我们给出了任何自适应松弛机制为避免此问题必须满足的三个必要条件(命题\ref{prop:conditions}),证明DADU——本文作为研究对象的自然松弛规则——违反了全部三个条件(推论\ref{cor:dadu_failure}),并通过覆盖4至32节点图结构的3,072次训练运行证实了这一失效:在DADU下,单个错误先验在87–97%的试验中抑制了真实边。其次,独立于对该机制的任何修复,我们以闭式证明,标准的基于相关系数匹配的目标函数将一条真实边及其反向边绑定在完全相同的代价$2r^2$上(引理\ref{lem:tie}),这并非因为底层的等方差模型不可辨识,而是因为归一化为相关系数恰好丢弃了使其可辨识的方差信息;相反,协方差匹配以至少$w_0^4$的可证明间隔将两个方向区分开来(引理\ref{lem:separation})。
cs.LG / 42 / 2609.03443

Beyond Straightness: Non-Crossing Flow Matching via Quantile AlignTree Coupling

超越直线性:基于分位数对齐树耦合的非交叉流匹配
Lin, Junyi, Li, Mengyu, Hu, Jingxuan, He, Kejun, Meng, Cheng
Abstract
The performance of Flow Matching largely depends on the quality of the coupling between the source and target distributions. However, independent coupling often leads to path crossings and local velocity ambiguity, while OT-based couplings typically incur high construction costs. To address this challenge, we propose Quantile AlignTree Flow Matching (QAT-FM), an efficient structured coupling strategy that constructs a hierarchical coupling between a Gaussian prior and the target data distribution via a quantile-aligned tree structure. QAT-FM constructs the coupling in $\mathcal{O}(Nd\log N)$ time and supports per-pair source sampling with $\mathcal{O}(d)$ complexity, enabling scalable training for large-scale high-dimensional generative tasks. Theoretically, we prove that the QAT coupling satisfies marginal consistency, induces non-crossing linear interpolation paths, and consistently improves path separation at intermediate times compared with independent coupling, thereby alleviating local velocity ambiguity. QAT-FM further extends naturally to conditional generation, enabling structured conditional coupling while preserving global Gaussian alignment. Experiments across diverse benchmark datasets demonstrate that QAT-FM achieves competitive generative performance while substantially reducing coupling construction cost.
Chinese Translation
流匹配(Flow Matching)的性能在很大程度上取决于源分布与目标分布之间耦合的质量。然而,独立耦合常常导致路径交叉和局部速度模糊,而基于最优传输(OT)的耦合通常具有较高的构建成本。为应对这一挑战,我们提出了分位数对齐树流匹配(Quantile AlignTree Flow Matching, QAT-FM),这是一种高效的结构化耦合策略,通过分位数对齐的树结构在高斯先验与目标数据分布之间构建层次化耦合。QAT-FM 以 $\mathcal{O}(Nd\log N)$ 的时间复杂度构建耦合,并支持每对复杂度为 $\mathcal{O}(d)$ 的源采样,从而为大规模高维生成任务提供可扩展的训练能力。在理论方面,我们证明 QAT 耦合满足边缘一致性,产生非交叉的线性插值路径,并且与独立耦合相比在中途时刻持续改善路径分离度,从而缓解局部速度模糊问题。QAT-FM 还可自然地扩展到条件生成任务,在保持全局高斯对齐的同时实现结构化的条件耦合。在多种基准数据集上的实验表明,QAT-FM 在大幅降低耦合构建成本的同时,取得了具有竞争力的生成性能。
cs.LG / 43 / 2609.03457

A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds

一种用于私有云CPU负载预测的两阶段预测系统
Javeed, Ashir, Borg, Anton, Grahn, Håkan, Lundberg, Lars, Patel, Dhyey, Shirinbab, Sogand
Abstract
Accurate cloud resource forecasting is essential for proactive resource provisioning, maintaining Quality of Service (QoS), and reducing operational costs in dynamic cloud environments. The existing forecasting approaches predominantly estimate future CPU workload directly from historical resource traces, which often overlook the relationship between customer service demand and subsequent resource consumption. This study proposes a two-stage integrated forecasting model that explicitly models this dependency by first forecasting customer service requests, expressed as Transactions Per Second (TPS), and subsequently estimating future CPU workload from the TPS forecast. Both the forecasting component and resource prediction component employed the XGBoost model within a cascaded learning architecture, complemented by adaptive online retraining using an expanding-window strategy to address concept drift in continuously evolving cloud workloads. The proposed work was evaluated using real-world traces collected from a private cloud environment comprising ten applications. Experimental results demonstrate robust forecasting performance by achieving Symmetric Mean Absolute Percentage Error (SMAPE) below $7\%$ for most applications, with the best-performing application achieving an MAE of $0.7372$, RMSE of $1.1866$, SMAPE of $3.57\%$, and an R2 of $0.9185$. Horizon-wise drift analysis confirmed stable recursive forecasting behavior with controlled error accumulation across a 60-step prediction horizon. Compared with the conventional direct CPU forecasting method, the proposed two-stage integrated model gives improved forecasting robustness, computational efficiency, and interpretability, making it well-suited for proactive resource management and intelligent auto-scaling in cloud computing environments.
Chinese Translation
在动态云环境中,准确的云资源预测对于主动式资源供给、维持服务质量(QoS)以及降低运营成本至关重要。现有的预测方法主要直接基于历史资源轨迹来估计未来的CPU负载,往往忽略了客户服务需求与后续资源消耗之间的关系。本研究提出了一种两阶段集成预测模型,该模型显式地建模了这一依赖关系:首先预测以每秒事务数(TPS)表示的客户服务请求,随后基于TPS预测结果估计未来的CPU负载。预测组件和资源预测组件均在级联学习架构中采用XGBoost模型,并结合采用扩展窗口策略的自适应在线重训练机制,以应对持续演化的云负载中的概念漂移问题。本研究使用从包含十个应用程序的私有云环境中收集的真实轨迹数据对所提出的方法进行了评估。实验结果表明,该模型具有稳健的预测性能,大多数应用程序的对称平均绝对百分比误差(SMAPE)低于7%,其中表现最好的应用程序MAE为0.7372,RMSE为1.1866,SMAPE为3.57%,R²为0.9185。基于预测步长的漂移分析证实了递归预测行为的稳定性,在60步预测范围内误差积累可控。与传统的直接CPU预测方法相比,所提出的两阶段集成模型在预测稳健性、计算效率和可解释性方面均有提升,非常适合用于云计算环境中的主动式资源管理和智能自动伸缩。
cs.LG / 44 / 2609.03464

Mind the Gap: Robustness Risks in PII Detection Systems

注意差距:个人身份信息检测系统中的鲁棒性风险
Zafar, Adeel, Nowaczyk, Slawomir
Abstract
Personally Identifiable Information (PII) detection is a foundational component of data protection infrastructure where missed entities constitute direct privacy and security risks. Although modern PII systems report strong performance on standard benchmarks, we show that these evaluations mask substantial robustness failures under realistic distribution shifts encountered in deployment. Rather than comparing state-of-the-art accuracy, we study how different PII detection paradigms fail under noisy, unstructured, and informal inputs. We construct a stress test benchmark spanning seven categories of natural distribution shift and evaluate representative systems from three widely deployed architectural families: encoder-based NER (SpaCy), rule-based hybrid detection (Presidio), and generative LLM extraction (Qwen2.5-3B). All three exhibit significant degradation on out-of-distribution inputs, but with distinct and complementary failure modes. Encoder models primarily fail on unseen surface forms and boundary detection, rule-based systems fail on non-standard formats, and LLMs exhibit entity-type confusion and generation instability. These results show that aggregate benchmark scores obscure deployment-critical weaknesses and that no single architecture is uniformly reliable across PII categories. Motivated by these findings, we propose a hybrid detection pipeline with a QA-driven feedback loop for iterative risk mitigation, and release our benchmark to support OOD-aware evaluation of PII systems.
Chinese Translation
个人身份信息(PII)检测是数据保护基础设施的基础组件,漏检实体将直接构成隐私与安全风险。尽管现代PII系统在标准基准上报告了优异性能,但我们发现这些评估掩盖了其在部署中面临的现实分布偏移下的大量鲁棒性失效。我们并非比较最先进模型的准确率,而是研究不同的PII检测范式在噪声、非结构化和非正式输入下的失效方式。我们构建了一个覆盖七类自然分布偏移的压力测试基准,并评估来自三种广泛部署的架构家族的代表性系统:基于编码器的命名实体识别(NER,SpaCy)、基于规则的混合检测(Presidio)以及生成式大语言模型抽取(Qwen2.5-3B)。三者均在分布外输入上表现出显著性能退化,但失效模式各不相同且互为补充:编码器模型主要在未见过的表面形式和边界检测上失效;基于规则的系统在非标准格式上失效;而大语言模型则表现出实体类型混淆和生成不稳定问题。这些结果表明,总体基准得分掩盖了部署中至关重要的缺陷,且没有任何单一架构在所有PII类别上都能保持一致的可靠性。基于这些发现,我们提出了一种带有问答(QA)驱动反馈回路的混合检测流水线,用于迭代式风险缓解,并发布我们的基准以支持面向分布外(OOD)的PII系统评估。
cs.LG / 45 / 2609.03495

Spectral characteristics of autoencoder parameters as a vector representation of data

自编码器参数的谱特征作为数据的向量表示
Nikitina, Maria, Bishuk, Anton, Bakhteev, Oleg
Abstract
This paper examines the relationship between the parameters of autoencoder models and the statistical properties of the data on which they are trained. Autoencoders are defined as models with an encoder-decoder architecture, trained to reconstruct input data through a compressed latent representation. It is proposed that the model parameters can be viewed as a dense vector representation of the corresponding sample. To test this hypothesis, a theoretical and experimental study is conducted in which a vector representation is formed based on the spectral characteristics of the autoencoder parameter matrices. Theoretical analysis shows that the singular values of the model parameter matrices are related to the eigenvalues of the covariance matrix of the training data, ensuring the transfer of information between the data space and the parameter space. Experimental results on the CIFAR-10 and FashionMNIST datasets confirm that the resulting vector representations allow for a high degree of accuracy in distinguishing between models trained on different data subsets, without resorting to complex vector generation algorithms or using the original samples. These results suggest that the parameters of trained autoencoders can be viewed as sample representations.
Chinese Translation
本文研究了自编码器模型参数与其训练数据的统计特性之间的关系。自编码器被定义为具有编码器-解码器架构的模型,通过压缩的潜在表示训练以重构输入数据。本文提出,模型参数可以被看作对应样本的稠密向量表示。为验证这一假设,本文进行了理论与实验研究,基于自编码器参数矩阵的谱特征构建向量表示。理论分析表明,模型参数矩阵的奇异值与训练数据协方差矩阵的特征值相关,从而保证了数据空间与参数空间之间的信息传递。在 CIFAR-10 和 FashionMNIST 数据集上的实验结果证实,所得到的向量表示能够以较高的准确度区分在不同数据子集上训练的模型,而无需借助复杂的向量生成算法或使用原始样本。这些结果表明,训练后的自编码器参数可以被视为样本的表示。
cs.LG / 46 / 2609.03504

Restricted Eigenvalues Beyond Gaussian Width: Threshold Occupancy under Heavy Tails

超越高斯宽度的限制特征值:重尾分布下的阈值占据
Fu, Shi, Xu, Huibo, Zhang, Qixin, Tao, Dacheng
Abstract
Restricted eigenvalue (RE) bounds govern stable recovery by norm-regularized estimators. For isotropic sub-Gaussian measurements, the benchmark sample size is $1+w(A)^2$, where $w(A)$ is the Gaussian width of the normalized descent cone. The COLT 2015 open-problem note (Banerjee et al., 2015) asked whether the same law follows for heavy-tailed designs from a uniform small-ball condition alone. We give an explicit and systematic negative answer to the general question as formulated there: the proposed law fails in its full dimension-free, arbitrary-set form, and the missing obstruction is simultaneous threshold occupancy. A constant-width polyhedral descent cone with fixed small-ball constants has zero empirical RE on every sample path up to half the ambient dimension. More generally, every finite range space admits exact threshold encoding in an arbitrarily narrow spherical cap and a lift to a full polyhedral descent-cone section. For every fixed threshold VC dimension $d$, as $\beta\downarrow0$, the sharp worst-case sample complexity is $\Theta(\beta^{-1}[d\log(1/\beta)+\log(1/\delta)])$. The separation persists under exact isotropy and all finite moments: on the same constant-width cone, Gaussian measurements succeed with $O(1+\log(1/\delta))$ samples, whereas an isotropic heavy-tailed design fails pathwise for $n\lesssim\sqrt{p/\log p}$. Gaussian smoothing yields an everywhere-positive $C^\infty$ density while retaining arbitrarily poor RE. Under isotropy, a distribution-free fallback governed by affine dimension times squared enclosing radius is sharp on this family.
Chinese Translation
限制特征值(RE)界决定了范数正则化估计器的稳定恢复性能。对于各向同性的次高斯测量,基准样本量为 $1+w(A)^2$,其中 $w(A)$ 是归一化下降锥的高斯宽度。COLT 2015 的开放问题笔记(Banerjee 等,2015)提出了如下疑问:仅凭一致小球条件,同样的规律是否对重尾设计也成立。我们对原文所提出的一般性问题给出了明确且系统性的否定答案:所提出的规律在其完全与维数无关、任意集合的形式下并不成立,而缺失的障碍是同步阈值占据。一个具有固定小球常数的常数宽度多面体下降锥,在环境维度一半以内的每个样本路径上经验 RE 均为零。更一般地,每个有限的范围空间都可以在任意窄的球冠中实现精确的阈值编码,并可提升为完整的多面体下降锥截面。对于每个固定的阈值 VC 维数 $d$,当 $\beta\downarrow 0$ 时,最坏情形样本复杂度的精确阶为 $\Theta(\beta^{-1}[d\log(1/\beta)+\log(1/\delta)])$。这种分离现象在精确各向同性和所有有限矩条件下依然存在:在同一个常数宽度锥上,高斯测量只需 $O(1+\log(1/\delta))$ 个样本即可成功,而各向同性的重尾设计在 $n\lesssim\sqrt{p/\log p}$ 时沿每条路径均告失败。高斯平滑可得到处处为正的 $C^\infty$ 密度,同时仍保持任意差的 RE。在各向同性条件下,由仿射维数乘以包围半径平方所决定的与分布无关的替代界,在该分布族上是紧的。
cs.LG / 47 / 2609.03505

An Adversarial Zero-Shot Learning Approach for Anomaly Detection in Multivariate IoT Traffic Data

一种面向多变量物联网流量数据异常检测的对抗性零样本学习方法
Rezakhani, Mahshid, Seyfi, Tolunay, Afghah, Fatemeh
Abstract
Anomaly detection in Internet of Things (IoT) networks presents unique challenges due to the diversity of devices, lack of labeled data, and domain variability across environments. In this paper, we propose a novel framework for multivariate time-series anomaly detection that leverages adversarial learning and contrastive loss within a sequence-based Variational Autoencoder (VAE) architecture. Our method enables zero-shot domain adaptation by jointly optimizing domain-invariant latent representations and semantically structured embedding spaces, without requiring labeled data or raw feature transfer. To address the heterogeneity of IoT deployments, we introduce encoder and decoder adaptor layers that align feature distributions across domains while preserving contextual semantics. Additionally, we propose a destination-based segmentation strategy to better model real-world communication structures in IoT traffic. Our framework is comprehensively evaluated on six distinct datasets spanning industrial, enterprise, general-purpose, smart home, and military automation domains across 44 transfer scenarios. Experimental results demonstrate strong zero-shot generalization in several cross-domain settings and competitive performance against a contrastive domain-adaptation baseline under realistic, heterogeneous, and privacy-constrained IoT conditions.
Chinese Translation
物联网(IoT)网络中的异常检测由于设备的多样性、标注数据的缺乏以及跨环境的域差异而面临独特挑战。本文提出了一种用于多变量时间序列异常检测的新型框架,该框架在基于序列的变分自编码器(VAE)架构中融合了对抗学习和对比损失。通过联合优化域不变的潜在表示和具有语义结构的嵌入空间,我们的方法实现了零样本域适应,且无需标注数据或原始特征迁移。为应对物联网部署的异构性,我们引入了编码器和解码器适配层,在保持上下文语义的同时对齐跨域的特征分布。此外,我们提出了一种基于目的地的分段策略,以更好地建模物联网流量中真实世界的通信结构。我们在涵盖工业、企业、通用、智能家居和军事自动化领域、跨越44个迁移场景的六个不同数据集上对该框架进行了全面评估。实验结果表明,该框架在若干跨域设置中表现出强大的零样本泛化能力,并且在真实、异构且受隐私约束的物联网条件下,其性能可与对比域适应基线方法相媲美。
cs.LG / 48 / 2609.03507

LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues

LongCounsel-8:基于多轮次心理咨询对话的抑郁症纵向追踪基准数据集
Li, Jiayi, Wu, Zhaomin, He, Bingsheng
Abstract
Tracking depression from multi-session counseling dialogues requires estimating both current symptom severity and how it changes across sessions. Yet progress on this task is constrained by the scarcity of longitudinal counseling data with standardized session-level depression labels. Existing resources typically provide either multi-session conversations without depression labels or labeled interviews in a single session. Building such a benchmark poses three challenges: maintaining longitudinal consistency and diversity, grounding symptom progression in empirical patterns, and expressing controlled depression states naturally without exposing target labels. To address these challenges, we introduce LongCounsel-8, a benchmark suite of three independently generated datasets totaling 7,749 five-session counseling trajectories, grounded in real-world client profiles, depression trajectories, symptom compositions, and counseling patterns. We combine profile-grounded simulation, empirically informed state construction, and indirect behavioral realization to address these challenges. Across the benchmark, simulated self-reports closely recover the controlled states, supporting label fidelity. Experiments on existing depression tracking methods reveal three key findings: (1) lower single-session score error does not guarantee accurate identification of trend, i.e., improvement or worsening; (2) existing methods are consistently less reliable on worsening trajectories; and (3) additional session history may reduce the accuracy of trend prediction. Together, these findings establish LongCounsel-8 as a foundation for advancing depression assessment from static, single-session prediction toward reliable longitudinal tracking of mental-health change.
Chinese Translation
从多轮次心理咨询对话中追踪抑郁症,需要同时评估当前症状严重程度及其跨轮次的变化。然而,缺乏带有标准化轮次级抑郁标签的纵向咨询数据,制约了该任务的研究进展。现有资源通常要么提供没有抑郁标签的多轮次对话,要么提供仅单轮次的有标签访谈。构建这样的基准面临三大挑战:保持纵向的一致性与多样性,使症状演变基于实证规律,以及在不暴露目标标签的前提下自然地表达受控的抑郁状态。为应对这些挑战,我们提出了 LongCounsel-8,这是一个由三个独立生成的数据集组成的基准套件,共包含 7,749 条五轮次咨询轨迹,其构建基于真实世界的来访者画像、抑郁轨迹、症状组合以及咨询模式。我们结合基于画像的模拟、有实证依据的状态构建以及间接的行为表现方式来应对上述挑战。在该基准上,模拟自述能够较好地还原受控状态,支持了标签的保真度。在现有抑郁症追踪方法上的实验揭示了三个关键发现:(1)较低的单轮次评分误差并不能保证准确识别趋势(即好转或恶化);(2)现有方法在恶化轨迹上的一致性可靠性较低;(3)额外的轮次历史信息可能降低趋势预测的准确性。这些发现共同确立了 LongCounsel-8 作为推动抑郁症评估从静态、单轮次预测迈向可靠的心理健康纵向变化追踪的基础。
cs.LG / 49 / 2609.03528

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

LeanGRPO:消除扩散强化学习中的冗余重计算
Wang, Sijie, Tan, Zhiqiang, Yang, Xinrui, Shi, Shaohuai
Abstract
Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and video generative models. However, most diffusion RL methods, including DanceGRPO and FlowGRPO, recompute selected timesteps with gradient tracking after rollout. Under on-policy training with the same backend for rollout and update, this recomputation is mathematically redundant. Intuitively, the rollout and policy update steps can reuse the same feed-forward backbone to avoid redundant computation, but doing so can incur a large memory overhead during rollout. To address the issue, we present LeanGRPO by restructuring the data-parallel layout and introducing two recompute-free training schedules for trajectory-logprob diffusion RL: (1) LeanGRPO-Retain enables gradient tracking during rollout and directly reuses the resulting computation graphs and saved activations for backward during update, requiring no recomputation; and (2) LeanGRPO-Reweight also enables gradients during rollout, but immediately backpropagates each selected step using a provisional advantage and delays gradient synchronization, then corrects the provisional gradients with the true advantage after the trajectory is completed. These schedules target different model scales and input sizes. Across FlowGRPO/DanceGRPO with FLUX.1-dev and Wan, LeanGRPO achieves up to 1.83x end-to-end speedup while preserving the original optimization objective.
Chinese Translation
扩散强化学习(RL)近期在图像与视频生成模型的后训练中取得了显著成功。然而,包括DanceGRPO和FlowGRPO在内的大多数扩散RL方法在 rollout 之后仍会对选定的时间步进行带梯度追踪的重计算。在 rollout 与更新使用同一后端的同策略(on-policy)训练中,这种重计算在数学上是冗余的。直观上,rollout 和策略更新步骤可以复用同一次前向传播的主干网络以避免冗余计算,但这样做会在 rollout 阶段带来较大的显存开销。为解决该问题,我们通过重构数据并行布局并提出两种免重计算的训练调度方案,构建了面向轨迹-对数似然(trajectory-logprob)扩散RL的LeanGRPO:(1)LeanGRPO-Retain 在 rollout 期间开启梯度追踪,并在更新阶段直接复用所得的计算图与保存的激活值进行反向传播,无需任何重计算;(2)LeanGRPO-Reweight 同样在 rollout 期间开启梯度,但立即使用临时优势值(provisional advantage)对每个选定时间步进行反向传播并延迟梯度同步,待轨迹完成后再用真实优势值修正临时梯度。这两种调度方案针对不同的模型规模和输入尺寸。在基于 FLUX.1-dev 和 Wan 的 FlowGRPO/DanceGRPO 实验中,LeanGRPO 在保持原始优化目标不变的情况下,实现了最高 1.83 倍的端到端加速。
cs.LG / 50 / 2609.03533

Coupled Scaling: A Representational Accessibility Framework for Neural Scaling Laws

耦合缩放:神经缩放定律的表征可达性框架
Wang, Jie
Abstract
Existing theories derive neural scaling from data geometry or a specified data-model spectrum, but systems trained on the same data can scale differently when architecture or optimization changes the representations they can efficiently reach. We introduce Coupled Scaling, a task-conditioned framework in which finite-budget scaling depends on the relation between task structure and the geometry accessible to an architecture-optimization system. In a solvable mode-truncation model, loss separates into target energy outside architectural support and an unresolved supported tail. For an arbitrary priority order, the residual lies between the best-N supported tail and the tail beyond the largest completed high-value prefix. If the cumulative-tail and coverage log-rates are $\gamma_{A,T}$ and $\rho_{A,O,T}$, the residual exponent lies in $[\rho_{A,O,T}\gamma_{A,T},\gamma_{A,T}]$. Under bounded off-prefix gain, the completed prefix is rate-determining and $\alpha_{A,O,T}=\rho_{A,O,T}\gamma_{A,T}$; for $a_{A,T,j}\asymp j^{-b_{A,T}}$, this gives $\alpha_{A,O,T}=\rho_{A,O,T}(b_{A,T}-1)$. A fixed-kernel specialization derives the training-time exponent from the near-zero tail of a task-weighted spectral measure defined independently of the loss fit. The framework separates architectural support from finite-budget acquisition and motivates two tests: static task-relevant geometry should track loss at a common budget, while multiscale geometry should track coupling-specific exponent ordering, including reversal across contrasting tasks. An audit of released emergence trajectories identifies the controls needed for a direct factorial test that measures geometry separately from the scaling fit.
Chinese Translation
现有理论从数据几何或特定的数据-模型谱推导神经缩放定律,但当架构或优化改变了系统能够高效到达的表征时,在同一数据上训练的系统可能呈现不同的缩放行为。我们提出耦合缩放(Coupled Scaling),这是一个以任务为条件的框架,其中有限预算下的缩放取决于任务结构与架构-优化系统可达几何之间的关系。在一个可求解的模式截断模型中,损失分解为架构支撑之外的目标能量与支撑内未解决的尾部。对于任意优先级排序,残差介于最优N项支撑尾部与最大已完成高价值前缀之外的尾部之间。若累积尾部与覆盖的对数速率分别为$\gamma_{A,T}$和$\rho_{A,O,T}$,则残差指数位于$[\rho_{A,O,T}\gamma_{A,T},\gamma_{A,T}]$内。当前缀外增益有界时,已完成前缀是速率决定因素,即$\alpha_{A,O,T}=\rho_{A,O,T}\gamma_{A,T}$;对于$a_{A,T,j}\asymp j^{-b_{A,T}}$,可得$\alpha_{A,O,T}=\rho_{A,O,T}(b_{A,T}-1)$。固定核特化情形下,训练时间指数由一个与损失拟合无关定义的、以任务加权的谱测度的近零尾部推导得出。该框架将架构支撑与有限预算下的获取区分开来,并启发了两项检验:在共同预算下,静态任务相关几何应与损失保持一致;而多尺度几何应追踪耦合特定的指数排序,包括在对比任务间出现的反转。对已发布涌现轨迹的审计识别出进行直接因子检验所需控制变量,该检验将几何与缩放拟合分别测量。
cs.LG / 51 / 2609.03582

WeatherNext 3: Increasing resolution and performance of global weather models with raw observations

WeatherNext 3:利用原始观测数据提升全球天气模型的分辨率与性能
Rasp, Stephan, Babenko, Boris, Masters, Dominic, El-Kadi, Andrew, Merchant, Samier, Shalev, Guy, Price, Ilan, Zyda, Fred, Lam, Remi, Shysheya, Sasha, Willson, Matthew, Markou, Stratis, Agrawal, Shreya, Vora, Suhani, Hassen, Mohammed Alewi, Mak, Sunny, Andersson, Tom R., Bela, Megan, Uddin, Akib, Levi, Nofar Peled, Gaiarin, Ben, Alet, Ferran, Bell, Aaron, Battaglia, Peter, Sanchez-Gonzalez, Alvaro
Abstract
State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficiency, but suffer two key shortcomings: their forecasts have lower spatial and temporal resolution than the best physics-based models and they are exclusively initialized with and trained on analysis data. As a result, they cannot directly make use of observations, and any biases in the analysis are inherited by the forecast. WeatherNext 3 addresses these shortcomings and establishes a new state-of-the-art for probabilistic medium-range forecasting skill. First, WeatherNext 3 generates new forecasts every hour (rather than every 6 hours like traditional global models) by ingesting low-latency geostationary satellite data. Second, WeatherNext 3's temporal and spatial resolution are on par with physics-based global models, with hourly time steps and 0.1 degree resolution for single-level variables, including solar radiation and cloud cover. Third, WeatherNext 3 moves beyond traditional analysis variables by learning to predict satellite-derived precipitation estimates, as well as tropical cyclone and station observations. Modelling sparse station data allows WeatherNext 3 to make 2m temperature and dewpoint predictions at any location and time, conditioned on local geographical features, with substantially lower error than competing global models, even when evaluated against unseen stations. Together, WeatherNext 3's capabilities move operational AI-based weather forecasting beyond emulating the traditionally distinct stages of data assimilation, forecasting and post-processing, which helps to further push the frontier of performance and granularity for global weather prediction.
Chinese Translation
最先进的AI天气模型已展现出卓越的中期预报能力和计算效率,但存在两个关键缺陷:其预报的时空分辨率低于最优的基于物理的模型,且仅使用分析数据进行初始化和训练。因此,它们无法直接利用观测数据,且分析数据中的任何偏差都会被继承到预报中。WeatherNext 3解决了这些缺陷,并在概率性中期预报能力上树立了新的最先进水平。首先,WeatherNext 3通过同化低延迟的地球静止卫星数据,每小时生成一次新的预报(而传统全球模型为每6小时一次)。其次,WeatherNext 3的时空分辨率与基于物理的全球模型相当,具有小时级时间步长,且对包括太阳辐射和云量在内的单层变量达到0.1度分辨率。第三,WeatherNext 3超越了传统的分析变量,学习预测卫星反演的降水估计以及热带气旋和站点观测数据。通过对稀疏站点数据建模,WeatherNext 3能够在任意地点和时间预测2米温度和露点温度,并以当地地理特征为条件,其误差显著低于同类全球模型,即使在未曾见过的站点上进行评估时也是如此。总而言之,WeatherNext 3的能力使基于AI的业务化天气预报超越了传统上相互独立的数据同化、预报和后处理阶段,有助于进一步推动全球天气预测在性能和精细度上的前沿发展。
cs.LG / 52 / 2609.03603

Neural-Network Maxent: a general extension with learned nonlinearity, applied to time-series for Desert Locust distribution modelling

神经网络Maxent:一种具有可学习非线性的通用扩展方法,并应用于沙漠蝗分布建模的时间序列分析
Grassi, Alessandro, Bellotto, Edoardo Kimani, Azami, Wassim El, Outmani, Sabrina, Houel, Maximilien
Abstract
Species Distribution Modelling (SDM) is essential for understanding how environmental conditions shape biodiversity, particularly for destructive pests such as the Desert Locust (Schistocerca gregaria), whose breeding dynamics are tightly coupled to rapidly evolving environmental conditions. Maxent has become the dominant method for presence-only data, but its reliance on a linear combination of hand chosen feature transforms limits its ability to capture the nonlinear, temporal relationships common in ecological monitoring, where covariates such as precipitation, soil moisture, and vegetation indices evolve meaningfully over time. Standard implementations flatten time-series covariates into independent features, discarding sequential structure that carries critical signal. We introduce RNN Maxent, an extension of the Maxent framework that replaces the fixed feature dictionary with a neural network, specifically a Gated Recurrent Unit (GRU), trained end to end via backpropagation. The approach preserves Maxent's presence only statistical foundations, background normalization, and probability calibration, differing only in that the nonlinearity is learned from data rather than fixed in advance. We apply RNN Maxent to map suitable habitat for the Desert Locust using 50 day environmental time series derived from ERA5 Land, MODIS, and Sentinel 3, maintaining a 7 day gap between covariates and presence records to yield forecasting behavior. Compared against standard Maxent, RNN Maxent improves performance across metrics (ROC AUC 0.862 std 0.036 vs. 0.792; F1 0.671 std 0.056 vs. 0.590).
Chinese Translation
物种分布建模(SDM)对于理解环境条件如何塑造生物多样性至关重要,尤其是对于沙漠蝗这类破坏性害虫,其繁殖动态与环境条件的快速演变紧密耦合。Maxent已成为仅存在数据处理中的主导方法,但其对人工设计的特征变换的线性组合的依赖,限制了其捕捉生态监测中常见的非线性时间关系的能力——在这些场景中,降水、土壤湿度和植被指数等协变量随时间发生显著变化。标准实现将时间序列协变量展平为独立特征,丢弃了携带关键信号的序列结构。我们提出RNN Maxent,这是Maxent框架的一种扩展,用神经网络(具体为门控循环单元GRU)取代固定的特征字典,并通过反向传播进行端到端训练。该方法保留了Maxent的仅存在数据统计基础、背景归一化和概率校准,唯一区别在于非线性是从数据中学习得到的,而非预先固定。我们将RNN Maxent应用于沙漠蝗适宜栖息地制图,使用由ERA5 Land、MODIS和Sentinel 3数据导出的50天环境时间序列,并在协变量与存在记录之间保持7天的时间间隔,从而实现预测功能。与标准Maxent相比,RNN Maxent在各评估指标上均有提升(ROC AUC 0.862±0.036 对比 0.792;F1 0.671±0.056 对比 0.590)。
cs.LG / 53 / 2609.03604

On the Interaction Between Model Compression and Test-Time Adaptation

论模型压缩与测试时自适应之间的交互作用
Corti, Francesco, Wang, Dong, Kwon, Young D., Mascolo, Cecilia, Saukh, Olga
Abstract
Deep neural networks deployed in the wild must be both efficient and adaptable, requiring model compression and test-time adaptation (TTA). While both are well studied in isolation, their interaction remains poorly understood. We systematically analyze how structured compression affects a model's ability to adapt under distribution shift. Using ResNet-18 and ViT-Base on CIFAR-10-C and ImageNet-C, we evaluate multiple compression methods combined with standard TTA techniques. We introduce a diagnostic framework that examines representational expressivity and adaptation subspace compatibility. Our results reveal a consistent gap: although compressed models retain high accuracy under supervised adaptation, their TTA performance degrades significantly with increasing compression. We show that this stems from reduced representational diversity and structural constraints that limit recoverability. These effects strongly depend on the compression method, highlighting the need to design compression strategies that preserve adaptability.
Chinese Translation
部署在真实环境中的深度神经网络必须兼具高效性与可适应性,这需要模型压缩和测试时自适应(Test-Time Adaptation, TTA)两种技术。尽管二者已被分别充分研究,但它们之间的交互作用仍知之甚少。我们系统地分析了结构化压缩如何影响模型在分布偏移下的自适应能力。我们在 CIFAR-10-C 和 ImageNet-C 数据集上使用 ResNet-18 和 ViT-Base,评估了多种压缩方法与标准 TTA 技术的组合。我们提出了一个诊断框架,用于考察表征表达能力和自适应子空间兼容性。我们的结果揭示出一个一致的差距:尽管压缩后的模型在有监督自适应下仍保持较高的准确率,但其 TTA 性能随着压缩程度的增加而显著下降。我们表明,这源于表征多样性的降低以及限制可恢复性的结构约束。这些效应在很大程度上依赖于压缩方法,凸显了设计能够保持自适应能力的压缩策略的必要性。
cs.LG / 54 / 2609.03660

Local Updates, Global Learning (LUGL): Playing Games with non-incremental Learners

局部更新,全局学习(LUGL):与非增量学习者进行博弈
Milec, David, Samothrakis, Spyridon, Fairbank, Michael, Soemers, Dennis J. N. J.
Abstract
The dominance of Neural Networks (NNs) in RL is partially due to their incremental learning capability, which naturally suits the online, non-stationary nature of self-play training. However, gradient-boosted trees like LightGBM are widely recognised as the state of the art for tabular data in supervised learning, often outperforming NNs in accuracy and efficiency. Game states are inherently tabular---discrete actions, categorical card identities, structured board positions---which makes them an ideal candidate for tree-based methods. We introduce LUGL (Local Updates, Global Learning), a framework that decouples data collection from model fitting, enabling non-incremental learners such as GBTs to operate in RL settings where they would otherwise fail due to distributional shift. LUGL alternates between a local updates phase, where the agent plays self-play games and accumulates tabular updates (Q-values, V-values, policies, or regret values) in a finite table, and a global learning phase, where the table is used to train a function approximator that generalises to unseen states before the table is reset. We test our approach in four standard perfect-information games (Tic-tac-toe, Connect-4, Othello, and Hex) and five imperfect-information games (Kuhn's poker, Leduc Hold'em, Liar's Dice, Goofspiel, and Flop5 Hold'em), and show that our results are competitive with or superior to DQN and DeepCFR. Our experiments demonstrate that the community's strong bias towards NNs in game-playing may be unwarranted, since LightGBM-based agents achieve competitive or superior performance across all tested benchmarks.
Chinese Translation
神经网络(NNs)在强化学习(RL)中的主导地位部分归功于其增量学习能力,这种能力天然契合自我博弈训练的在线、非平稳特性。然而,以LightGBM为代表的梯度提升树在监督学习中被广泛认为是表格数据的最先进方法,在准确性和效率上常常优于神经网络。博弈状态本质上就是表格形式的——离散动作、分类的牌面标识、结构化的棋盘位置——这使其成为基于树的方法的理想应用场景。我们提出了LUGL(局部更新,全局学习),该框架将数据收集与模型拟合解耦,使梯度提升树(GBTs)等非增量学习者能够在强化学习环境中运行,否则它们会因分布偏移而失效。LUGL交替执行两个阶段:局部更新阶段,智能体进行自我博弈,并在有限表格中累积表格化更新(Q值、V值、策略或遗憾值);全局学习阶段,利用该表格训练一个可泛化到未见状态的函数逼近器,随后重置表格。我们在四个标准完全信息博弈(井字棋、Connect-4、黑白棋和Hex)和五个不完全信息博弈(Kuhn扑克、Leduc德州扑克、Liar's Dice、Goofspiel和Flop5德州扑克)中测试了我们的方法,结果表明其性能与DQN和DeepCFR相当甚至更优。我们的实验表明,研究界对神经网络在博弈领域的强烈偏好可能并无必要,因为基于LightGBM的智能体在所有测试基准上均取得了具有竞争力或更优的性能。
cs.LG / 55 / 2609.03662

Extracting Forgotten Prompts from Targeted Unlearned Models

从定向遗忘模型中提取被遗忘的提示词
Hoi-Ting, Au Ashley, Kurmanji, Meghdad, Shen, William F., Lane, Nicholas D., He, Ligang
Abstract
Recent unlearning methods (e.g. NPO, DPO, LUNAR) make use of refusal alignment to suppress forgotten data. However, it has been shown that refusal responses might leave traces of unlearning, and recent attacks have been able to successfully recover some of the unlearned knowledge. In this paper, we uncover a new vulnerability. Existing attacks typically assume that the forgotten prompts are already known to the adversary and focus on recovering their answers. However, we show that the forgotten prompts themselves can be extracted by using the retained data and black-box access to the model. Our attack, Targeted Active Search (TAS), first identifies the forgotten entities by constructing canonical templates and entity pool, and selectively querying the model using the most informative template-entity pair under a limited query budget. Once the entities are identified, TAS instantiates prompt templates with those entities to probe the unlearned model and reconstruct the forgotten prompts. Experiments across three unlearning methods with three datasets and three LLMs shows that TAS recovers the forgotten entity with $100\%$ accuracy and reconstructs up to $95\%$ of forgotten prompts, all while using up to $99.7\%$ fewer queries than naive probing.
Chinese Translation
近期的机器遗忘方法(如NPO、DPO、LUNAR)利用拒绝对齐来抑制已遗忘的数据。然而,已有研究表明拒绝响应可能留下遗忘的痕迹,且近期的攻击已能成功恢复部分被遗忘的知识。本文揭示了一种新的漏洞:现有攻击通常假设攻击者已经知晓被遗忘的提示词,并专注于恢复其对应的答案。而我们的研究表明,利用保留数据和模型的黑盒访问权限,被遗忘的提示词本身也可以被提取出来。我们的攻击方法——定向主动搜索(Targeted Active Search, TAS)——首先通过构建规范化模板和实体池来识别被遗忘的实体,并在有限的查询预算下,利用信息量最大的模板-实体组合选择性地查询模型。一旦实体被识别,TAS便将这些实体填入提示词模板,对遗忘模型进行探测并重构被遗忘的提示词。在三种遗忘方法、三个数据集和三个大语言模型上的实验表明,TAS能够以100%的准确率恢复被遗忘的实体,并重构多达95%的被遗忘提示词,同时与朴素探测相比减少高达99.7%的查询次数。
cs.LG / 56 / 2609.03667

Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning

离线多智能体强化学习中基于序列模型的分布外泛化
Hidaoui, Oussama, Ebead, Omer, Sob, Ulrich Armel Mbou, Singh, Siddarth, Formanek, Juan Claude, Chalumeau, Felix, Mahjoub, Omayma, Abramowitz, Sasha, de Kock, Ruan John, Khlifi, Wiem, Nessir, Louay Ben, Toit, Simon Verster Du, Rajaonarivonivelomanantsoa, Daniel, Osman, Asim Awad, Fokam, Arnol Manuel, Shabe, Refiloe, Pretorius, Arnu
Abstract
Generalising to unseen tasks remains a fundamental challenge in offline multi-agent reinforcement learning (MARL). In this work, we present a principled analysis of zero-shot task generalisation in the offline setting and conduct an extensive empirical investigation into the scaling behaviour governing task diversity, dataset size, and network capacity. To facilitate this study, we extend offline sequence modelling architectures to handle multi-task observation and action spaces alongside variable agent counts across tasks. Our primary finding is that scaling task diversity---rather than sheer dataset size is the dominant factor in achieving robust zero-shot transfer. Through large-scale experiments across four challenging environments (Connector, RWARE, SMAX, and LBF), we demonstrate that our multi-task approach achieves a mean improvement of 3.2x on held-out test tasks compared to single-task models and consistently outperforms strong behaviour cloning baselines. These results suggest that the development of generalisable MARL agents should prioritise the diversity of the training distribution with varying numbers of agents, providing a roadmap for scaling offline MARL effectively.
Chinese Translation
泛化到未见过的任务仍然是离线多智能体强化学习(MARL)中的一项根本性挑战。在本工作中,我们对离线设定下的零样本任务泛化进行了系统性的理论分析,并针对任务多样性、数据集规模和网络容量所主导的缩放行为开展了广泛的实证研究。为支持这一研究,我们扩展了离线序列建模架构,使其能够处理多任务的观测与动作空间,并支持跨任务的可变智能体数量。我们的主要发现是:扩展任务多样性——而非单纯的数据集规模——是实现稳健零样本迁移的主导因素。通过在四个具有挑战性的环境(Connector、RWARE、SMAX 和 LBF)上进行的大规模实验,我们证明了我们的多任务方法在保留测试任务上相比单任务模型平均取得 3.2 倍的性能提升,并持续优于强行为克隆基线。这些结果表明,开发可泛化的 MARL 智能体应优先考虑具有不同智能体数量的训练分布多样性,为有效扩展离线 MARL 提供了路线图。
cs.LG / 57 / 2609.03686

Resolution-Aware Experimental Design under Partial Identifiability

部分可辨识性下分辨率感知的实验设计
Fotias, Sofianos Panagiotis
Abstract
Experimental design is commonly framed as choosing the experiment expected to provide the most information. Under partial identifiability however, persistent nuisance uncertainty can make the same observation carry different structural meanings. We introduce Resolution-Aware Experimental Design (RAED), which selects an experiment by the smallest expected nonempty structural candidate set achievable subject to false-exclusion control. We prove an exact cross-nuisance aliasing separation: an experiment can be preferred by structural and full-latent information gain, average classification, and nuisance-marginalized informativeness while having arbitrarily poorer valid structural resolution. RAED nevertheless preserves the expected ordering under a genuine composite Blackwell comparison. To make this criterion operational, we develop a learned score-based implementation with finite-sample nuisance-average and positive-tail calibration, and characterize a rare-tail sample-complexity obstruction. Under constrained sensing, two subsurface-flow benchmarks exhibit genuine RAED--expected-information-gain (EIG) experiment-selection disagreements, with the clearest and largest held-out resolution differences in WCA. In a fluvial benchmark, tail protection changes the selected physical experiment and replaces hard-region false exclusions primarily with explicit ambiguity. In a mechanistic methane-oxidation benchmark, a prospectively specified 5\% false-exclusion tolerance also yields a nontrivial finite-sample population guarantee for tail-sensitive nuisance risk, with 95\% joint confidence across all three structural families.
Chinese Translation
实验设计通常被框定为选择预期提供最多信息的实验。然而,在部分可辨识性(partial identifiability)条件下,持续的干扰不确定性可能使同一观测结果承载不同的结构含义。我们提出了分辨率感知实验设计(Resolution-Aware Experimental Design, RAED),该方法在控制错误排除率的前提下,通过可实现的最小期望非空结构候选集来选择实验。我们证明了一个精确的跨干扰混叠分离:某实验可能在结构与全潜变量信息增益、平均分类性能以及边缘化干扰的信息性指标上更优,但其有效结构分辨率却可以任意更差。然而,RAED在真正的复合Blackwell比较下仍能保持预期排序。为使该准则可操作,我们开发了一种基于学习评分的实现,包含有限样本干扰平均与正尾部校准,并刻画了稀尾部样本复杂度的障碍。在受限感知条件下,两个地下水流基准实验表现出真实的RAED与期望信息增益(Expected Information Gain, EIG)实验选择分歧,其中WCA上留出集分辨率差异最为显著。在河流基准实验中,尾部保护改变了所选的物理实验,并主要以显式歧义取代了困难区域的错误排除。在甲烷氧化的机理基准实验中,前瞻设定的5%错误排除容差也为尾部敏感的干扰风险提供了非平凡的有限样本总体保证,在全部三个结构族上实现了95%的联合置信度。
cs.LG / 58 / 2609.03705

Federated Causal Discovery via Regression-Directed Cumulants

基于回归导向累积量的联邦因果发现
Torrijos, Pablo, Stella, Fabio, Gámez, José A., Puerta, José M.
Abstract
In this paper we study linear non-Gaussian acyclic models (LiNGAM) when used in federated environments. These causal models allow one to go beyond Markov equivalence. However, in many domains data are scarce, and increasing the sample size by centralising data from different clients is not advisable due to regulations such as the GDPR. The federated environment offers an attractive option to balance privacy and causal discovery accuracy. Unfortunately, the standard centralised estimator in the LiNGAM setting, i.e., DirectLiNGAM, cannot be straightforwardly federated. Higher-order cumulant tensors offer a way around this obstacle: they depend only on the joint distribution of the variables involved and add exactly across independent sample groups, so a single communication round suffices in horizontal, vertical, and hybrid partitions. However, FedISHC, i.e., the current federated method along these lines, breaks down under near-symmetric noise. To overcome the above limitation, we introduce the FedRCD family of causal discovery algorithms, and investigate three variants that trade off communication rounds against algebraic noise; two of them are exact federated counterparts of the centralised high-order cumulant (HC) and HC-LiNGAM algorithms, and the single-round variants further effectively support exact unlearning at any granularity, from a single observation to a whole client. Numerical experiments show that at sample sizes typical of real deployments, the entire cumulant-based federated family does not actually rank variables by the population asymmetry that the scores encode at zero. It ranks them by a variance ladder induced by the DAG along its directed paths, the cumulant counterpart of varsortability. Marginal standardisation collapses every cumulant method to near-random ordering, while scale-invariant DirectLiNGAM, not federable under this protocol, is unaffected.
Chinese Translation
本文研究了线性非高斯无环模型(LiNGAM)在联邦环境中的应用。这类因果模型能够超越马尔可夫等价类的限制。然而,在许多领域中数据是稀缺的,而由于《通用数据保护条例》(GDPR)等法规的限制,通过集中来自不同客户端的数据来增加样本量并不可取。联邦环境为平衡隐私与因果发现精度提供了一种有吸引力的选择。遗憾的是,LiNGAM设置下的标准集中式估计器,即DirectLiNGAM,无法直接进行联邦化。高阶累积量张量为克服这一障碍提供了一条途径:它们仅依赖于所涉及变量的联合分布,并且在各独立样本组之间具有精确的可加性,因此无论是水平、垂直还是混合划分,仅需一轮通信即可完成。然而,沿着这一思路的现有联邦方法FedISHC在近似对称噪声下会失效。为克服上述局限,我们提出了FedRCD系列因果发现算法,并研究了三种在通信轮数与代数噪声之间进行权衡的变体;其中两个是集中式高阶累积量(HC)算法和HC-LiNGAM算法的精确联邦对应版本,而单轮通信的变体还进一步有效支持任意粒度(从单条观测到整个客户端)的精确遗忘(unlearning)。数值实验表明,在真实部署场景典型的样本量下,整个基于累积量的联邦算法家族实际上并非按照分数在零处所编码的总体不对称性对变量进行排序,而是按照由有向无环图(DAG)沿其有向路径诱导的方差阶梯进行排序,即varsortability现象在累积量上的对应物。边际标准化会使所有累积量方法的排序退化为近似随机顺序,而尺度不变的DirectLiNGAM(虽然在该协议下无法联邦化)则不受影响。
cs.LG / 59 / 2609.03762

Projected Riemannian Gradient Descent for the Bures-Wasserstein Barycenter: Dimension-Independent Linear Convergence at Unit Step Size

Bures-Wasserstein 重心的投影黎曼梯度下降法:单位步长下与维度无关的线性收敛性
Afham, A.
Abstract
The computation of the Bures-Wasserstein (BW) barycenter of an ensemble of positive definite matrices arises throughout machine learning, optimal transport, and quantum information. Riemannian gradient descent (RGD) at unit step size -- the fixed-point iteration used in practice -- converges rapidly, yet existing analyses present a dichotomy: unit-step guarantees carry worst-case exponential dependence on the dimension, while dimension-independent guarantees require small step sizes that forfeit the empirical speed. We resolve this dichotomy, not by improving the guarantees for unit-step RGD, but by proposing a Projected RGD algorithm that achieves dimension-independent linear convergence at unit step size. The achieved rate, $(1 - \kappa^{-3/2})$, where $\kappa$ is the condition number of the ensemble, also polynomially improves on the best small-step guarantee ($\kappa^{3/2}$ versus $\kappa^{5/2}$ iteration complexity). The crux is a novel Projection Lemma: clipping the eigenvalues of a positive matrix to an interval $[\alpha, \beta]$ is the closed-form, non-expansive (1-Lipschitz) BW-metric projection onto the set $\{S : \alpha I \leq S \leq \beta I\}$ -- a statement which, unlike its known one-sided counterpart, does not follow from convexity. The projection is moreover free: it reuses an eigendecomposition the next iteration must perform in any case, so the projected and unprojected iterations cost the same per step. The same analysis covers the invariant matrix projection problem of Brahmachari et al. (2025), whose fixed-point algorithm we identify as unit-step RGD on a totally geodesic submanifold, thereby extending the dimension-independent guarantee to that setting verbatim.
Chinese Translation
计算一组正定矩阵的 Bures-Wasserstein(BW)重心问题广泛出现于机器学习、最优传输和量子信息等领域。单位步长的黎曼梯度下降法(RGD)——即实践中常用的不动点迭代——收敛迅速,然而现有分析呈现出一分为二的局面:单位步长的收敛保证在最坏情况下对维数呈指数级依赖,而与维数无关的收敛保证则要求小步长,从而牺牲了经验上的快速性。我们解决了这一二分困境,其方式并非改进单位步长 RGD 的收敛保证,而是提出一种投影 RGD 算法,在单位步长下实现与维数无关的线性收敛。所达到的收敛率为 $(1 - \kappa^{-3/2})$,其中 $\kappa$ 为该矩阵集合的条件数,且在多项式意义上优于已有的最优小步长保证(迭代复杂度为 $\kappa^{3/2}$,优于 $\kappa^{5/2}$)。其核心是一个新颖的投影引理:将正定矩阵的特征值截断到区间 $[\alpha, \beta]$,即为到集合 $\{S : \alpha I \leq S \leq \beta I\}$ 上的具有闭式解的非扩张(1-Lipschitz)BW 度量投影——与已知的单侧版本不同,这一结论无法由凸性直接推出。此外,该投影是零成本的:它复用了下一次迭代无论如何都必须执行的特征分解,因此投影迭代与无投影迭代每步的计算代价相同。同一分析还覆盖了 Brahmachari 等人(2025)提出的不变矩阵投影问题,我们将其不动点算法识别为全测地子流形上的单位步长 RGD,从而将维数无关的收敛保证原封不动地推广到该情形。
cs.LG / 60 / 2609.03763

From Nowcasting to Forecasting: Adapting a Reanalysis-Trained

从临近预报到短期预报:对再分析训练模型的适配
Partio, Mikko, Hieta, Leila, Laine, Ossi
Abstract
Accurate cloud-cover forecasts are important for temperature prediction, radiation forecasting, and solar-power operations. Short-range forecasting methods can preserve observed cloud placement during the first forecast hours, but their skill decreases when cloud fields evolve through formation, dissipation and deformation. Longer lead times require accounting for atmospheric evolution, but operational numerical weather prediction (NWP) forecasts may not accurately represent the satellite-observed cloud state at initialization. We develop CloudCast v2, a machine-learning model for 12-hour cloud-cover forecasting from observation-based initial conditions. The model is first trained on the Copernicus European Regional Reanalysis (Ridal2024) to learn cloud-evolution dynamics, and is then adapted to satellite-derived cloud fields using conditional flow matching (Lipman2023), a generative method that transforms noise into cloud-cover forecasts conditioned on the observed initial cloud fields and NWP inputs. CloudCast v2 reduces mean absolute error by 10% relative to its predecessor, CloudCast v1 (Partio2025), over the 1-12 h range. It also overtakes CloudCast v1 in fractions skill score, a neighborhood-based measure of spatial agreement, after approximately 3-6 h, depending on the cloudiness category. These results show that observation-initialized machine-learning forecasts can extend beyond the usual 1-3-hour nowcasting range while retaining spatial detail from satellite cloud fields.
Chinese Translation
准确的云量预报对于温度预测、辐射预报和太阳能发电运行至关重要。短程预报方法能够在最初预报小时内保持观测到的云的位置,但当云场经历生成、消散和形变等演化过程时,其预报技巧会下降。更长的预报时效需要考虑大气演变,但业务化数值天气预报(NWP)在初始化时可能无法准确表征卫星观测到的云状态。我们开发了CloudCast v2,一个基于观测初始条件进行12小时云量预报的机器学习模型。该模型首先在哥白尼欧洲区域再分析资料(Ridal2024)上训练以学习云演化动力学,然后利用条件流匹配(conditional flow matching,Lipman2023)适配到卫星反演的云场上。条件流匹配是一种生成式方法,能够以观测的初始云场和NWP输入为条件,将噪声转换为云量预报。在1-12小时范围内,CloudCast v2相较其前身CloudCast v1(Partio2025)将平均绝对误差降低了10%。此外,根据云量类别的不同,在约3-6小时之后,CloudCast v2在分数技巧评分(fractions skill score,一种基于邻域的空间一致性度量)上也超越了CloudCast v1。这些结果表明,以观测初始化的机器学习预报可以将预报时效扩展到通常1-3小时的临近预报范围之外,同时保留卫星云场的空间细节。
cs.LG / 61 / 2609.03770

OBER+: Continuity-Aware Reporting and Traceable Continuous Improvement in Outcome-Based Education

OBER+:成果导向教育中的连续性感知报告与可追溯的持续改进
Rajasekar, Elakkiya
Abstract
Institutions practising outcome-based education compute learning outcome attainment routinely, while reviews of curriculum analytics report an absence of evidence on how that computation informs decisions. This paper presents OBER+, an extension of a deployed institutional attainment platform that computes the step from a measured shortfall to an evaluated corrective action. Five connected stages accumulate attainment across deliveries of a course, signal a shortfall and a persistent shortfall, grade it on cutoffs the regulator already uses, record the decision against a catalogue of practices annotated with their evidence, log the change, and quantify the subsequent movement in the shortfall. A further rule compares successive statements of an outcome, so attainment is never read as a series across a point at which the outcome changed. Applying the rules to the live record of two real courses produced three results. Every outcome of a core course was substantively redefined between consecutive deliveries, with subject matter moving between outcome numbers, so a naive reading would have reported a twenty-five point collapse between quantities that do not refer to the same learning. Recomputing the platform's figures from its documented rule showed six of ten differing by more than rounding explains, in a pattern that identified a defect since reported to the institution. Across fifteen statement pairs from three transitions, five were identical character for character, and among the ten that were not, the outcome carrying a given number was nearest to a differently numbered earlier outcome in six, a result resting on an ordering of similarities and requiring no threshold and no labelling. The contribution is a computational design for outcome-based reporting, stated as rules any attainment platform can implement, with evidence of what they make visible in a live institutional record.
Chinese Translation
实施成果导向教育的机构通常定期计算学习成果达成度,然而课程分析领域的综述报告指出,缺乏证据说明这些计算如何为决策提供依据。本文提出OBER+,对一个已部署的机构达成度平台进行扩展,计算从测量到的差距到评估后的纠正措施这一步骤。五个相互关联的阶段实现如下功能:跨多次课程开课累积达成度,标记差距与持续性差距,按照监管机构已在使用的阈值对其进行分级,将决策记录到附有证据标注的实践目录中,记录变更,并量化差距随后的变化。另一条规则比较同一学习成果的先后表述,从而确保达成度不会跨越成果变更的时间点被误读为一个连续序列。将上述规则应用于两门真实课程的在用记录,产生了三项结果。第一,一门核心课程在连续两次开课之间,所有学习成果均被实质性重新定义,学科内容在成果编号之间移动,因此朴素解读会报告出一个二十五点的骤降,但比较的实际上是不指代相同学习内容的数量。第二,按照平台文档化的规则重新计算平台数据,十个数据中有六个的差异超出舍入所能解释的范围,其模式指向一个已向该机构报告的缺陷。第三,在来自三次过渡的十五对成果表述中,有五对逐字完全相同;而在其余十对中,六对里携带某一编号的成果与其之前不同编号的成果最为相似——这一结果基于相似性排序,无需阈值,也无需人工标注。本文的贡献在于提出了一种面向成果导向报告的计算设计,表述为任何达成度平台均可实现的规则,并提供证据说明这些规则能在机构在用记录中揭示何种信息。
cs.LG / 62 / 2609.03790

Landmark-Based Discrimination of Injury-Associated Athlete-Sessions from Minute-Resolution Multimodal Football Monitoring Data

基于里程碑的足球分钟级多模态监测数据中损伤相关运动员训练课的判别研究
Chatzidimitriou, Evangelos, Tserpes, Konstantinos
Abstract
Athlete monitoring data may be recorded minute by minute throughout a match or training session, while injury information may only indicate whether the entire session was injury-associated. This creates a modelling problem: assigning the same session-level label to every minute would imply that injury status is known at each exact time, even though within-session injury onset is unknown. Our novelty is a fixed-landmark, one-representation-per-athlete-session formulation that directly addresses this mismatch. Instead of labelling every minute, we construct one representation per athlete-session at each landmark using information observed up to that point. This keeps the target at the session level and avoids unsupported minute-level injury supervision. A landmark is a fixed time point within the same session, such as 10, 20, or 30 minutes. At each landmark, we assess whether the whole session is injury-associated or non-injury-associated and examine how discrimination changes as more within-session information becomes available. Using 2020 SoccerMon data, we analyse 3,743 athlete-sessions from 48 elite women's football athletes, including 22 injury-associated sessions from five athletes. We evaluate pre-session, cumulative, dynamic, and combined representations with athlete-disjoint validation, athlete-cluster bootstrap uncertainty, common-cohort sensitivity analysis, alternative negative-athlete fold allocations, equal-athlete weighting, and Logistic Regression, Random Forest, and XGBoost benchmarks. Primary CUM+DYN Logistic Regression yields ROC-AUC 0.367-0.607 and PR-AUC 0.0080-0.0150 across landmarks, with wide uncertainty. PRE-containing representations show higher point estimates at several landmarks but remain uncertain.
Chinese Translation
运动员监测数据可以在比赛或训练过程中按分钟记录,而损伤信息往往只能表明整节训练课是否与损伤相关。这就产生了一个建模难题:若将相同的训练课级别标签赋予每一分钟,则意味着损伤状态在每一精确时刻均已知,尽管课内损伤发生时刻实际上是未知的。本文的创新之处在于提出了一种固定里程碑、每个运动员训练课单一表示的建模方式,直接应对这一不匹配问题。我们不再对每一分钟进行标注,而是在每个里程碑时间点利用截至该时点所观测到的信息,为每个运动员训练课构建一个表示。这使预测目标保持在训练课级别,避免了缺乏依据的分钟级损伤监督。里程碑是同一训练课内的固定时间点,例如10分钟、20分钟或30分钟。在每个里程碑处,我们评估整节训练课是与损伤相关还是与损伤无关,并考察随着课内信息的增多,判别能力如何变化。基于2020年SoccerMon数据,我们分析了来自48名精英女子足球运动员的3,743个运动员训练课,其中包括来自5名运动员的22个损伤相关训练课。我们采用运动员不相交验证、运动员聚类bootstrap不确定性估计、共同队列敏感性分析、替代性负例运动员折分配、运动员等权处理,以及逻辑回归(Logistic Regression)、随机森林(Random Forest)和XGBoost基准模型,评估了课前、累积、动态及组合表示。主要的CUM+DYN逻辑回归模型在各里程碑处的ROC-AUC为0.367–0.607,PR-AUC为0.0080–0.0150,不确定性较大。包含课前信息的表示在若干里程碑处点估计较高,但仍存在不确定性。
cs.LG / 63 / 2609.03801

From Ordered Bernoulli Levels to Critical-Line Geometry: Integer Quantization, Bernoulli Residual Phase, and Prime-Power Spectra

从有序伯努利水平集到临界线几何:整数量子化、伯努利剩余相位与素数幂谱
Yılmaz, Y. Kenan
Abstract
We study the ordered Bernoulli-word kernel f(p,n,k)=p^k(1-p)^(n-k) and the geometry generated by its inverse-integer level sets. The binary level 2^(-n) selects p=1/2 as the unique real split-independent anchor. Under complement-preserving complex continuation, the pair becomes z=1/2+iu and 1-z=1/2-iu, producing a conjugation-symmetric vertical geometry before any zeta-function input is introduced. The quadratic coordinate Q(z)=z(1-z)=1/4+u^2 has a sharp minimum at the central point and admits an exact integer quantization. For critical-line zero ordinates gamma_k, the induced levels L_k=1/4+gamma_k^2 are decomposed exactly as L_k=N_k+delta_k, where N_k is the nearest integer and delta_k is a periodic first-Bernoulli residual. Circularization gives Z_k=exp(2 pi i delta_k), isolating gamma_k^2 mod 1 as the residual phase variable. Unique factorization resolves the integer shells into prime-generator coordinates, while a distinct complex exponent s lifts the same construction to the Dirichlet atoms m^(-s), linking the Dirichlet-series and Euler-product assemblies. Exact identities, classical zeta connections, numerical controls, and open conditional Weyl tests are kept explicitly separate. No proof of the Riemann Hypothesis is claimed.
Chinese Translation
我们研究有序伯努利词核 f(p,n,k)=p^k(1-p)^(n-k) 及其由逆整数水平集所生成的几何。二值水平 2^(-n) 选出 p=1/2 作为唯一与分裂方式无关的实数锚点。在保持补对称的复延拓下,这一对变为 z=1/2+iu 与 1-z=1/2-iu,在任何 zeta 函数输入引入之前便产生了一个具有共轭对称性的竖直几何结构。二次坐标 Q(z)=z(1-z)=1/4+u^2 在中心点处取得尖锐极小值,并容许一个精确的整数量子化。对于临界线零点纵坐标 gamma_k,其诱导水平 L_k=1/4+gamma_k^2 可被精确分解为 L_k=N_k+delta_k,其中 N_k 是最近的整数,delta_k 是周期性的第一伯努利剩余。通过圆化得到 Z_k=exp(2 pi i delta_k),从而将 gamma_k^2 mod 1 分离为剩余相位变量。唯一分解将整数壳层解析为素数生成元坐标,而一个独特的复指数 s 将同一构造提升到狄利克雷原子 m^(-s),从而将狄利克雷级数组装与欧拉乘积组装联系起来。精确恒等式、经典 zeta 函数联系、数值验证以及开放的条件性 Weyl 检验均被明确地分开陈述。本文未声称证明黎曼猜想。
cs.LG / 64 / 2609.03807

Free Pause Tokens

免费暂停令牌
Langford, John, Godey, Nathan, Monea, Giovanni, Artzi, Yoav, Dong, Harry, Fan, Ying, de Rosa, Gustavo, Zhan, Zheng
Abstract
A free pause token gives a language model extra compute to form each next-token prediction (as a pause, or thinking, token does) but carries that compute in a parallel prediction stream over a weight-shared backbone rather than as an extra token in the sequence. It improves next-token prediction by 2-3 centinats in practice on a 1B parameter model. Because the pause rides an existing position instead of adding one, it is free to use: at inference it adds no context length, no KV cache, and essentially no latency with the growth in inference flops typically irrelevant as it is not the active bottleneck on throughput. The only primary cost is in training, where additional training compute versus an optimized pretraining pipeline is reduced to as low as x1.14 while preserving most of the benefits. The result is an isoflop, isoparameter, and isotoken improvement over standard next token trained transformers.
Chinese Translation
免费暂停令牌为语言模型提供了额外的计算量以形成每个下一词元预测(如同暂停令牌或思考令牌的作用),但其计算通过一个权重共享主干网络上的并行预测流来承载,而非作为序列中的额外令牌。在实践中,该方法在10亿参数模型上将下一词元预测准确率提升了2-3个百分点。由于暂停 rides 于已有位置而非新增位置,其使用是免费的:在推理时,它不增加上下文长度,不增加KV缓存,也几乎不增加延迟,推理计算量的增长通常无关紧要,因为它并非吞吐量的主要瓶颈。唯一的显著成本在于训练阶段——与优化的预训练流程相比,额外训练计算量可低至1.14倍,同时保留了大部分收益。由此,我们在等计算量、等参数量、等令牌数的条件下,实现了对标准下一词元训练的Transformer的性能提升。
cs.LG / 65 / 2609.03809

A Peer-Relative Representation Learning Framework for Energy Inefficiency Identification in Mobile Network Sites

一种用于移动网络站点能效低下识别的对等相对表示学习框架
Koto, Eliud Nyakweba, Toit, Jaco du, Stoltz, Adham, Preez, Johan du
Abstract
Energy consumption is one of the largest operational expenditure items for mobile network operators, yet site-level energy inefficiencies such as faulty cooling controllers, idle radio equipment, and parasitic auxiliary loads often remain undetected because no ground-truth inefficiency labels exist and historical measurements may already contain embedded inefficiencies. This study proposes an unsupervised peer-relative approach based on the premise that sites with similar structural and operational characteristics should exhibit comparable energy consumption. To capture these relationships, a novel energy-aware Minimum Distortion Embedding (MDE) formulation is introduced that extends the standard MDE objective with an energy-based repulsion mechanism. This encourages sites with anomalously high energy consumption relative to comparable peers to become displaced from their local neighbourhoods in the embedding space. The resulting low-dimensional representation simultaneously preserves structural similarity and encodes energy-related deviations, enabling the identification of potentially inefficient sites through peer-relative comparison. The derived anomaly scores provide a practical mechanism for prioritising field investigations, allowing mobile network operators to focus engineering resources on sites most likely to yield energy savings. Experimental results demonstrate that the proposed approach outperforms conventional anomaly detection baselines and provides a robust foundation for large-scale energy-efficiency optimisation in mobile networks.
Chinese Translation
能耗是移动网络运营商最大的运营支出项目之一,然而站点级的能效低下问题(如故障的冷却控制器、闲置的无线设备以及寄生的辅助负载)往往难以被察觉,原因在于不存在能效低下的真实标签(ground-truth),且历史测量数据中可能已经蕴含了能效低下的模式。本研究提出了一种无监督的对等相对方法,其基本前提是具有相似结构和运行特征的站点应表现出相近的能耗水平。为捕捉这些关系,本文引入了一种新颖的能量感知最小失真嵌入(Minimum Distortion Embedding, MDE)形式化方法,通过基于能量的排斥机制扩展了标准MDE目标函数。这使得相对于可比对等站点能耗异常偏高的站点在嵌入空间中被推离其局部邻域。所得到的低维表示同时保留了结构相似性并编码了与能量相关的偏差,从而能够通过对等相对比较识别出潜在的低效站点。由此导出的异常分数为现场排查的优先级排序提供了实用机制,使移动网络运营商能够将工程资源集中于最有可能带来节能收益的站点。实验结果表明,所提出的方法优于传统的异常检测基线,为移动网络的大规模能效优化提供了坚实的基础。
cs.LG / 66 / 2609.03826

Witnesses Explain Anomalies

见证方向解释异常
Diop, Lamine
Abstract
Unsupervised anomaly detection scores each point of an unlabelled, contaminated sample in a single pass, and increasingly must also explain why a point is flagged. Yet the dominant detectors give a score with no account of which features drive it, and explanations are bolted on post-hoc with SHAP or LIME, which re-query the detector thousands of times per point and only approximate it. We introduce WAND, an unsupervised tabular anomaly detector that is explainable by design. WAND organises its computation around directions on the unit sphere, scoring each point by how far its projection escapes a sub-Gaussian extreme-value baseline. The originality of our approach is that the witness directions that flag a point, being vectors in feature space, are its explanation, a per-feature attribution obtained at no cost over scoring and, since the score is differentiable, recoverable by gradients. Scoring is linear in the sample size, and a probe-efficiency bound guarantees every anomaly a witness, hence an explanation. Across 47 ADBench datasets WAND attains the best mean Friedman rank at ROC-AUC parity with 16 unsupervised baselines, so the gain is interpretability at no accuracy cost; its native explanations are more accurate and faithful than post-hoc SHAP/LIME and ECOD at a fraction of the query cost. WAND is thus a practical, interpretable solution for explainable anomaly detection.
Chinese Translation
无监督异常检测通常在单次遍历中为无标签且含污染的样本中的每个点打分,如今还需进一步解释某点为何被标记为异常。然而,主流检测器仅给出分数,却不说明是哪些特征导致了该分数;解释往往依赖事后附加的SHAP或LIME方法,这些方法需要对检测器进行数千次重复查询,且只能近似原模型。我们提出了WAND,一种可解释性内生于设计的无监督表格数据异常检测器。WAND将其计算组织在单位球面上的方向之上,通过衡量每个点的投影超出亚高斯极值基线的程度来为其打分。我们方法的独特之处在于:标记异常点的见证方向(witness directions)本身是特征空间中的向量,因而构成了该点的解释——一种在打分过程中无需额外代价即可获得的逐特征归因;并且由于分数可微,该归因也可通过梯度求得。打分的计算复杂度关于样本量呈线性,且探测效率界保证每个异常点都存在见证方向,从而保证其解释性。在47个ADBench数据集上,WAND在ROC-AUC与16个无监督基线持平的情况下,取得了最优的平均Friedman秩,说明其提升在于不损失精度的前提下获得了可解释性;相较于事后解释方法SHAP/LIME以及ECOD,WAND的原生解释更为准确和忠实,而查询成本仅为其一小部分。因此,WAND是可解释异常检测的一种实用且可解释的解决方案。
cs.LG / 67 / 2609.03842

Multi-step Proximal Policy Improvement in Offline Reinforcement Learning

离线强化学习中的多步近端策略改进
Choi, Soohyun, Cho, Seonvin, Hong, Songnam
Abstract
Offline reinforcement learning (RL) must reconcile two competing requirements: policy updates should stay near dataset-supported actions to keep value estimates reliable, yet meaningful gains often require moving beyond the behavior distribution. We develop a geometric view of offline actor updates by modeling policies as a probability manifold endowed with a chosen metric geometry. Under this lens, a broad class of offline actor objectives can be interpreted as a single proximal policy improvement step (SPI), i.e., an implicit discretization of a manifold gradient flow induced by a critic-defined energy. Building on this insight, we propose multi-step proximal policy improvement (MPI), a plug-in refinement mechanism that composes sequential re-centered proximal steps. MPI enables controlled policy improvement beyond dataset support while retaining proximal control at each refinement. The framework accommodates multiple policy geometries and admits practical instantiations for deterministic and diagonal-Gaussian policies. Experiments on D4RL benchmarks show that small numbers of MPI refinements improve strong offline baselines, including TD3+BC, ReBRAC, and IQL, on many tasks. Focused diagnostics further distinguish re-centered refinement from fixed-objective update scheduling and characterize limitations under critic error.
Chinese Translation
离线强化学习(RL)必须调和两个相互冲突的需求:策略更新应保持在数据集所支持的动作附近,以确保价值估计的可靠性;然而,有意义的性能提升往往需要超越行为分布。我们通过将策略建模为赋予特定度量几何的概率流形,为离线actor更新建立了一种几何视角。在此视角下,一大类离线actor目标可以被解释为单步近端策略改进(Single-step Proximal policy Improvement, SPI),即由critic定义的能量所诱导的流形梯度流的一种隐式离散化。基于这一洞察,我们提出多步近端策略改进(Multi-step Proximal policy Improvement, MPI),这是一种即插即用的改进机制,通过组合多次顺序的重新居中近端步骤来实现。MPI能够在数据集支持范围之外实现可控的策略改进,同时在每次改进中保持近端控制。该框架兼容多种策略几何结构,并针对确定性策略和对角高斯策略给出了实用的实例化方法。在D4RL基准上的实验表明,少量的MPI改进步骤能够在许多任务上提升强离线基线方法的表现,包括TD3+BC、ReBRAC和IQL。针对性的诊断实验进一步区分了重新居中改进与固定目标更新调度,并刻画了在critic误差存在下的局限性。
cs.LG / 68 / 2609.03851

Pushing the (Decision) Boundaries: Dynamically Calibrating Differentially Private Noise to Explainability in Federated Learning

拓展(决策)边界:在联邦学习中根据可解释性动态校准差分隐私噪声
Khavkin, Michael, Lee, Kichang, Jin, Jaeho, Ko, JeongGil, Toch, Eran
Abstract
Federated Learning (FL) with Differential Privacy (DP) is increasingly adopted to preserve data confidentiality in distributed machine learning. However, DP noise distorts learned representations and degrades explanation fidelity, limiting differentially private FL where trustworthy explanations are required, such as assistive clinical diagnosis. Prior work adapted DP noise with static feature-importance signals, restricting explainability to post hoc analysis and precluding noise calibration to explanation quality during training. We propose XCal-FL, a closed-loop, explainability-driven local training algorithm for image classification in cross-silo FL that dynamically calibrates DP noise from three complementary signals: (1) prediction logit variations, measuring causal influence on model confidence, (2) counterfactual margins, capturing decision-boundary sensitivity, and (3) saliency concentration, quantifying spatial coherence of model attention, while enforcing formal DP guarantees via adaptive privacy accounting. Experiments on three medical imaging datasets across varying FL configurations show that XCal-FL yields more accurate and interpretable global models, improving predictive performance by over 10\% and explanation fidelity by up to 5$\times$ over static-noise FL, and outperforming state-of-the-art adaptive DP methods in fidelity. XCal-FL also achieves higher privacy-budget efficiency, turning each unit of cumulative privacy loss into larger gains in both accuracy and explanation fidelity. Our analysis further reveals that, unlike predictive performance, which scales roughly linearly with privacy loss, explanation fidelity exhibits non-linear dynamics. These findings suggest explainability is a distinct dimension of the privacy trade-off that cannot be inferred from utility alone, with implications for training and privacy-budget allocation in decision-critical applications.
Chinese Translation
结合差分隐私(Differential Privacy, DP)的联邦学习(Federated Learning, FL)日益被广泛采用,以保护分布式机器学习中的数据机密性。然而,DP噪声会扭曲学习到的表示并降低解释的保真度,从而限制了需要可信解释场景下差分隐私联邦学习的应用,例如辅助临床诊断。已有工作利用静态特征重要性信号来调整DP噪声,这将可解释性局限于事后分析,并且无法在训练过程中根据解释质量对噪声进行校准。我们提出了XCal-FL,一种面向跨筒仓(cross-silo)联邦学习图像分类任务的闭环、可解释性驱动的本地训练算法。该算法基于三个互补信号动态校准DP噪声:(1)预测logit变化,衡量对模型置信度的因果影响;(2)反事实边际,捕捉决策边界敏感性;(3)显著性集中度,量化模型注意力在空间上的一致性,同时通过自适应隐私核算来保证形式化的DP保障。在三个医学影像数据集及多种联邦学习配置上的实验表明,XCal-FL能够得到更准确且更可解释的全局模型:与静态噪声联邦学习相比,预测性能提升超过10%,解释保真度最高提升5倍,并且在解释保真度上优于最先进的自适应DP方法。XCal-FL还实现了更高的隐私预算效率,将每一单位的累积隐私损失转化为更大的准确率和解释保真度收益。我们的分析进一步揭示,与随隐私损失大致线性扩展的预测性能不同,解释保真度呈现非线性动态。这些发现表明,可解释性是隐私权衡中的一个独立维度,无法仅从效用角度推断得出,这对决策关键型应用中的模型训练与隐私预算分配具有重要意义。
cs.LG / 69 / 2609.03858

High-Dimensional Learning Dynamics of Attention-Indexed Models

注意力索引模型的高维学习动力学
Xu, Yizhou, Sagitova, Margarita, Zdeborová, Lenka, Krzakala, Florent
Abstract
Attention mechanisms are central to modern foundation models, yet their training dynamics remain poorly understood, especially when the attention matrices have extensive rank. In this work, we study attention-indexed models, a broad framework that can represent multi-layer and multi-head attention architectures. First, we show that, in a suitable high-dimensional limit, the population-loss landscape is characterized by a finite set of trace order parameters. In contrast, online stochastic gradient descent (SGD) is governed by an infinite hierarchy of matrix moments, which we show can be exponentially well-approximated by a finite truncated system. Second, this framework reveals that attention parameterization itself can act as an architectural implicit bias. Direct optimization of an attention matrix $S\in\mathbb{R}^{d\times d}$ can remain trapped in an uninformative state. Tied attention ($S=WW^\top$) induces an automatic symmetry-breaking mechanism and yields weak recovery in $\Theta(d^2\log d)$ samples. For untied attention, $S=UV^\top$, we uncover a fast-slow mechanism: the pre-activation mean first evolves on a fast timescale, while the overlaps evolve on a slower one. Weak recovery on the $\Theta(d^2\log d)$ scale occurs when the state selected by the fast dynamics breaks the initial symmetry.
Chinese Translation
注意力机制是现代基础模型的核心,但其训练动力学仍然鲜为人知,尤其是当注意力矩阵具有较高秩时。在本工作中,我们研究了注意力索引模型(attention-indexed models),这是一个能够表示多层和多头注意力架构的广泛框架。首先,我们证明,在适当的高维极限下,总体损失景观由有限的迹阶序参量刻画。相比之下,在线随机梯度下降(SGD)由矩阵矩的无穷层级结构支配,我们证明该层级结构可以被一个有限的截断系统以指数级精度很好地近似。其次,该框架揭示了注意力参数化本身可以作为一种架构上的隐式偏置。直接优化注意力矩阵 $S\in\mathbb{R}^{d\times d}$ 可能会陷入无信息的状态。绑定注意力($S=WW^\top$)会诱导一种自动的对称性破缺机制,并在 $\Theta(d^2\log d)$ 个样本量下实现弱恢复。对于非绑定注意力 $S=UV^\top$,我们发现了一种快慢机制:预激活均值首先在快速时间尺度上演化,而重叠量则在较慢的时间尺度上演化。当快动力学所选状态打破初始对称性时,弱恢复会在 $\Theta(d^2\log d)$ 的样本量尺度上发生。
cs.LG / 70 / 2609.03878

Differentiable Interval Bottlenecks for Interpretable Anomaly Detection in Numerical Data

面向数值数据可解释异常检测的可微区间瓶颈
Diop, Lamine, Plantevit, Marc
Abstract
Reconstruction-based anomaly detectors are accurate but opaque: a deep autoencoder flags a sample without telling a practitioner which feature ranges made it anomalous. We propose DIFFINT, an autoencoder whose latent bottleneck is structured as a set of soft, axis-aligned interval memberships learned end-to-end directly from raw numerical data, without any discretization or binarization. Each latent unit corresponds to a human-readable hyper-rectangle in feature space; an instance is encoded by how strongly it falls inside each interval relative to the other units, and its reconstruction error is the anomaly score. This keeps the power of differentiable representation learning while exposing an inspectable internal structure. We make the inductive bias precise: a certified reconstruction-error lower bound for points that fall outside every active coordinate of the learned support (with a Lipschitz-enforced decoder), and a graded, empirically verified suppression mechanism for the usual case in which only a few features are abnormal; and we provide a closed-form, label-free importance that ranks each (unit, feature) pair from quantities the model already maintains, turning trained intervals into auditable candidate constraints without ever seeing an anomaly label. On 48 ADBench benchmarks against 22 baselines under a common [-1, 1]-normalized protocol, DIFFINT attains the best mean rank overall on both metrics (4.10 on ROC-AUC, 4.16 on AUPR); among inlier-only detectors it leads its regime clearly, and it is competitive with the strongest contaminated-data detectors (see the stratified and complete-case analyses). It is the only interpretable detector in the statistically-tied leading cluster of seven methods.
Chinese Translation
基于重构的异常检测器虽然准确但不透明:深度自编码器(autoencoder)标记出一个异常样本时,却无法告知从业者是哪些特征区间使其成为异常。我们提出 DIFFINT,一种自编码器,其潜在瓶颈被结构化为一组软性的、轴对齐的区间隶属关系,直接从原始数值数据端到端学习得到,无需任何离散化或二值化。每个潜在单元对应于特征空间中一个人类可读的超矩形;一个实例的编码方式取决于其相对于其他单元落入各区间内部的强度,其重构误差即为异常分数。这在保持可微表示学习能力的同时,暴露出可检查的内部结构。我们将这种归纳偏置形式化:对于落在所学支撑集(support)的所有活跃坐标之外的点,在解码器满足 Lipschitz 约束的条件下,给出一个经过认证的重构误差下界;并且针对只有少数特征异常的常见情形,提出一种分级的、经实验验证的抑制机制。此外,我们提供一种闭式、无需标签的重要性度量,从模型自身维护的量出发对每个(单元, 特征)对进行排序,从而将训练得到的区间转化为可审计的候选约束,而无需看到任何异常标签。在统一采用 [-1, 1] 归一化协议、覆盖 48 个 ADBench 基准并对比 22 个基线方法的评估中,DIFFINT 在两项指标上均取得最佳平均排名(ROC-AUC 上为 4.10,AUPR 上为 4.16);在仅使用正常样本训练的检测器中,它在其适用范围内明显领先,并与最强的含污染数据检测器具有竞争力(参见分层分析与完整案例分析)。在七个方法构成的统计上并列的领先集群中,它是唯一的可解释检测器。
cs.LG / 71 / 2609.03900

Beyond Endpoint Scores: Time- and Capacity-Conditioned Evaluation of Continual Knowledge Updating

超越端点分数:持续知识更新的时间与容量条件化评估
Choi, Heejin
Abstract
Continual knowledge-updating methods are often declared superior from one final checkpoint and one conventional adapter rank. We show that this can be insufficient to identify the better operating point. Holding a periodic hierarchy fixed, we compare it with cumulative replay over a 24-month Wikidata stream while varying evaluation month, replay LoRA rank, and query formulation. The apparent winner changes across this region: on Qwen2.5-1.5B, the hierarchy's 5.0-point advantage over rank-8 replay becomes an 11.6-point deficit against rank-72 replay, and at high ranks a consolidation-aligned endpoint can suggest a tie while time-averaged replay leads by 9-13 points. The same rank-conditioned reversal appears on Llama-3.2-1B and held-out paraphrases. These results show that method ranking in continual updating can depend jointly on when performance is measured and how much replay-side adaptation capacity the baseline receives. We therefore propose reporting trajectories and capacity sweeps, and declaring a robust winner only when the ordering is stable across the evaluation region; otherwise, comparisons should report winner regions and retention-stability-cost frontiers. Under this protocol, the periodic hierarchy is a lower-update-cost operating point, not a quality winner.
Chinese Translation
持续知识更新方法往往仅依据一个最终检查点和一个常规适配器秩(rank)而被宣称优于其他方法。我们证明,这种做法可能不足以确定更优的运行点。在固定周期性层次结构(periodic hierarchy)的前提下,我们将其与累积回放(cumulative replay)在一条24个月的Wikidata数据流上进行比较,同时改变评估月份、回放侧LoRA秩以及查询表述方式。表观赢家在这一评估区域内会发生变化:在Qwen2.5-1.5B上,层次结构相对于秩为8的回放所具有的5.0分优势,在面对秩为72的回放时变为11.6分的劣势;在高秩设置下,一个与知识整合对齐的端点可能显示两者打平,而按时间平均的回放则领先9-13分。同样的秩条件化反转也出现在Llama-3.2-1B以及留出的改述查询上。这些结果表明,持续更新中的方法排序可能同时取决于性能测量的时机以及基线所获得的回放侧适配容量。因此,我们建议报告性能轨迹与容量扫描结果,只有当排序在整个评估区域内保持稳定时才宣布稳健的赢家;否则,比较结果应报告赢家区域以及保持性-稳定性-成本前沿。在该评估协议下,周期性层次结构是一个更新成本更低的运行点,而非质量上的赢家。
cs.LG / 72 / 2609.03937

RATL: Learning from Retrieved Residuals for Robust Multivariate Time-Series Forecasting

RATL:基于检索残差学习的鲁棒多元时间序列预测
He, Yuchen, Cang, Yueyang, Ning, Zhiyuan, Wang, Ningyu, Shi, Li
Abstract
Retrieval-augmented generation (RAG) complements parametric models with retrieved external evidence. The same idea is attractive for continuous-output regression, but directly reusing retrieved target values is often not robust when samples differ in output level, numerical scale, or local dynamics. Moreover, conventional forecasting pipelines generally use residuals for model optimization and error diagnosis, but do not retain individual historical residual examples as memory that can be accessed at inference time.For multivariate time-series forecasting, we propose RATL, a plug-in residual-retrieval and feedback-correction method. RATL freezes a base forecaster to construct retrieval keys and turns its historical forecast residuals into a train-only memory specific to that base model. At inference time, RATL retrieves residual trajectories from similar historical contexts subject to causal availability constraints, then uses a set-aware router operating over forecast blocks and variables to select and combine these trajectories. Experiments show that historical residuals matched to the current context contain reusable forecasting information and that RATL improves frozen base forecasters in most experimental settings. Ablations further show that learned routing strengthens raw residual feedback, while validation-based correction-strength selection limits residual over-injection.On real-world benchmarks, we use iTransformer as the primary frozen base forecaster, compare against multiple strong forecasting baselines, and test transferability across backbones. The results show that RATL can further improve base-forecaster performance in most settings.Overall, RATL shifts the retrieved object from historical target values to base-model-specific historical forecast errors, providing a plug-in, residual-memory-based paradigm for learned feedback correction in continuous-output forecasting.
Chinese Translation
检索增强生成(Retrieval-augmented generation, RAG)通过检索外部证据来补充参数模型。同样的思想对连续输出回归也很有吸引力,但当样本在输出水平、数值尺度或局部动态上存在差异时,直接复用检索到的目标值往往并不鲁棒。此外,传统预测流程通常将残差用于模型优化和误差诊断,但并未将个体的历史残差样本保留为可在推理时访问的记忆。针对多元时间序列预测,我们提出RATL,一种即插即用的残差检索与反馈校正方法。RATL冻结一个基础预测器以构建检索键,并将其历史预测残差转化为专属于该基础模型的仅训练阶段记忆。在推理时,RATL在因果可用性约束下从相似的历史上下文中检索残差轨迹,然后在预测块和变量上运作的集合感知路由器来选择和组合这些轨迹。实验表明,与当前上下文匹配的历史残差包含可复用的预测信息,并且RATL在大多数实验设置下都能提升冻结的基础预测器的性能。消融实验进一步表明,学习式路由增强了原始残差反馈的效果,而基于验证的校正强度选择限制了残差的过度注入。在真实世界基准上,我们以iTransformer作为主要的冻结基础预测器,与多个强大的预测基线进行比较,并测试了跨骨干网络的迁移性。结果表明,RATL在大多数设置下能够进一步提升基础预测器的性能。总体而言,RATL将检索对象从历史目标值转移到了基础模型特有的历史预测误差,为连续输出预测中的学习式反馈校正提供了一种基于残差记忆的即插即用范式。
cs.LG / 73 / 2609.03941

Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO

余量-漂移回放(Headroom-Drift Replay):GRPO 中一种实现原则性回放控制的基础机制
Park, Hyun Bin, Chang, Du-Seong
Abstract
RL-based post-training for reasoning models is increasingly bottlenecked by repeated fresh rollout generation, particularly in agentic settings where environment interaction dominates wall-clock cost. Replay can reduce this burden by reusing past trajectories, but existing methods typically embed it within larger training pipelines involving exploration, experience restructuring, or mixed-policy optimization. This makes replay's own contribution difficult to isolate. We ask a focused question: how far can principled replay selection alone go? We introduce Headroom-Drift Replay, a group-level replay control primitive for GRPO that separates reuse into two decisions. Headroom ranks stored groups by remaining learning value, while Drift gates them by compatibility with the current policy. The fresh on-policy stream remains unchanged, and the method adds no auxiliary generation or training machinery. Across mathematical reasoning, multimodal reasoning, and Agentic Search benchmarks, this single intervention outperforms naive replay and matches or exceeds broader replay methods on Avg Mean@32. In Agentic Search, where environment interaction dominates cost, it delivers comparable quality at materially lower wall-clock time.
Chinese Translation
基于强化学习(RL)的推理模型后训练正日益受到重复的全新轨迹采样(rollout)生成的制约,尤其在智能体(agentic)场景下,与环境的交互占据了主要的墙钟时间成本。回放(Replay)可以通过复用历史轨迹来减轻这一负担,但现有方法通常将回放嵌入到涉及探索、经验重构或混合策略优化的更大训练流程中,这使得回放本身的贡献难以被单独剥离评估。我们提出一个聚焦的问题:仅凭原则性的回放选择能走多远?我们提出余量-漂移回放(Headroom-Drift Replay),这是一种面向 GRPO 的组级回放控制基础机制,将复用拆分为两个决策:Headroom 依据剩余学习价值对存储的组进行排序,而 Drift 则依据与当前策略的兼容性对其进行门控筛选。全新的在线策略数据流保持不变,且该方法不引入任何额外的生成或训练机制。在数学推理、多模态推理和智能体搜索(Agentic Search)基准上,这一单一干预在 Avg Mean@32 上优于朴素回放,并达到或超越更复杂的回放方法。在环境交互占主导成本的智能体搜索任务中,它以显著更低的墙钟时间实现了相当的质量。
cs.LG / 74 / 2609.03949

VestigeKV: The NoPE-MLA KV Cache Carries Its Own Eviction Signal in a Vestigial Branch

VestigeKV:NoPE-MLA 的 KV 缓存在其残余分支中自带淘汰信号
Fan, WenJie
Abstract
The problem. A long-lived KV cache must be compressed before the queries that will read it exist; selection by observed attention (H2O, SnapKV) collapses there (0.00-0.33 needle retrieval on a NoPE MLA model), because a token's importance has not yet been observed. The method. On Kimi Linear, VestigeKV evicts by a query-independent signal the cache already carries: the 64-dimensional decoupled branch, a vestige of RoPE that NoPE training repurposes into a salience channel. Reading 11% of each row, it partitions the cache: the top-m rows stay in the attended tier; every other row moves -- exactly, never deleted -- to a GPU-resident archive reachable per step by a certified trigger. No training, no quantization, no weight or kernel change. Cost. Nothing measurable: retrieval holds at 1.00 under 8x and 0.92 under 32x from 8k to 65k context, zero gap to full-row selection. The attended tier is 0.25 KB of Kimi Linear's 8.1 KB per-token cache at 32x; the archive stays bit-exact and GPU-resident, with host offload as the VRAM-reclaiming variant. The recall tier -- the standard configuration -- holds 128x at 1.00. Kimi K3 is reported to use a NoPE Gated-MLA variant; if its cache layout matches, the method plausibly extends there -- we make no claim beyond the measured model. NoPE exclusivity. The identical operator on a RoPE MLA collapses to 0.08 (plain eviction: 0.42); query-independent salience itself exists only without rotation (top-1 targets span 2.3-6.7% of tokens vs. 10.2-46.8%), and query-universal exact merging is provably impossible under RoPE. All thresholds were frozen before data; 20 archived verdicts and 8 closed routes accompany the paper.
Chinese Translation
问题所在。一个长期存续的 KV 缓存必须在读取它的查询尚未存在时就进行压缩;此时基于已观测注意力(H2O、SnapKV)的选择方法会失效(在 NoPE MLA 模型上“大海捞针”检索得分仅为 0.00–0.33),因为 token 的重要性尚未被观测到。方法。在 Kimi Linear 上,VestigeKV 依据缓存本身已携带的、与查询无关的信号进行淘汰:即 64 维的解耦分支——这是 RoPE 的残余结构,在 NoPE 训练中被改造成一个显著性通道。通过只读取每行的 11%,它对缓存进行划分:得分最高的 m 行保留在注意力层;其余每一行都被完整地——精确保存、绝不删除——移入一个驻留在 GPU 上的归档区,每一步均可通过一个经过验证的触发器访问。无需训练、无需量化、无需更改权重或内核。代价。几乎可以忽略不计:在 8k 到 65k 的上下文长度下,“大海捞针”检索得分在 8 倍压缩下保持 1.00,32 倍压缩下保持 0.92,与全行选择相比零差距。在 32 倍压缩下,注意力层仅占 Kimi Linear 每 token 8.1 KB 缓存中的 0.25 KB;归档区保持比特级精确并驻留于 GPU,主机卸载(host offload)可作为回收显存的变体方案。召回层——即标准配置——在 128 倍压缩下保持 1.00。据报道,Kimi K3 使用了 NoPE Gated-MLA 变体;若其缓存布局一致,该方法有望推广至此——但我们不对实测模型之外的情况做任何声明。NoPE 独有性。完全相同的算子在 RoPE MLA 上性能崩溃至 0.08(普通淘汰为 0.42);与查询无关的显著性本身只存在于无旋转的情况下(top-1 目标占 token 数的 2.3–6.7%,相比之下有旋转时为 10.2–46.8%),且可证明在 RoPE 下不可能实现与查询无关的精确合并。所有阈值均在获取数据之前冻结;论文随附 20 条归档判定与 8 条已关闭的技术路线。
cs.LG / 75 / 2609.03972

OSR: Output Space Redistribution for Adaptive Label Removal in Classification Models

OSR:用于分类模型中自适应标签移除的输出空间重分布方法
Peng, Minyi, Gunamardi, Darian, Tjuawinata, Ivan, Zheng, Yongsen, Lam, Kwok-Yan
Abstract
Label removal occurs frequently in classification systems with evolving taxonomies, where categories must be dynamically updated or eliminated. To accommodate such changes, classification models must adapt accordingly. Existing solutions, broadly categorized as retraining-based and feature-space-adjustment-based, share common limitations despite their variations, including reliance on access to original data, substantial computational and storage costs, inconsistent results, poor scalability, and degradation of model utility. To address this, we propose a novel approach that leverages statistical redistribution in the output space to approximate the post-removal confidence vectors of a retrained model. Applicable as a modular output filter, our method bypasses the burden of feature-space adjustments or loss-function convergence, alleviating scalability limitations. Furthermore, by requiring only existing labels and prior output confidences, the method potentially mitigates privacy concerns inherent to data-dependent solutions. Extensive experiments demonstrate competitive performance against full retraining, with improvements in computational efficiency and privacy preservation across several classification tasks.
Chinese Translation
标签移除在类别体系不断演化的分类系统中频繁发生,此时类别需要被动态地更新或删除。为适应此类变化,分类模型必须进行相应调整。现有解决方案大致可分为基于重训练的方法和基于特征空间调整的方法,尽管形式各异,但它们存在共同的局限性,包括依赖原始数据的访问、高昂的计算与存储成本、结果不一致、可扩展性差以及模型效用的下降。为解决这些问题,我们提出了一种新颖的方法,利用输出空间中的统计重分布来近似重训练模型在移除标签后的置信度向量。我们的方法可作为一种模块化的输出过滤器,无需进行特征空间调整或损失函数收敛,从而缓解了可扩展性的限制。此外,该方法仅需要现有标签和先前的输出置信度,因此有望缓解依赖数据的解决方案所固有的隐私问题。大量实验表明,该方法在与完整重训练相比具有竞争力,并在多个分类任务中提升了计算效率和隐私保护水平。
cs.LG / 76 / 2609.04007

RobustSeiz: An Open-Source Framework for Benchmarking the Robustness of EEG Seizure Detection Models

RobustSeiz:一个用于基准测试脑电图癫痫发作检测模型鲁棒性的开源框架
Mohammadi, Mohammad, Zarei, Alireza
Abstract
Despite strong performance on held-out electroencephalography (EEG) data, seizure detectors may fail under real-world acquisition variability, artifacts, and adversarial inputs. We introduce RobustSeiz, an open-source, model-agnostic framework that provides a standardized, reproducible protocol for stress-testing and comparing seizure detectors under controlled, clinically motivated distribution shifts before deployment. We standardize four public scalp-EEG corpora (CHB-MIT, TUSZ, Siena, and SeizeIT1) into BIDS-EEG trees and evaluate subject-independent detectors on held-out splits. Environment, noise, and adversarial transforms are swept over predefined hyperparameter grids. Each run reports sample- and event-level sensitivity, precision, F1, false positives per 24 h, Lead and Lag onset timing, and Monte Carlo dropout predictive agreement. RobustSeiz includes a Dockerized GPU pipeline, experiment registry, and full-evaluation and research-subset modes. We demonstrate the framework with a contemporary seizure detector on TUSZ across the complete implemented shift grid; an AWGN analysis illustrates how perturbation severity changes detection quality, onset timing, and predictive agreement. RobustSeiz provides a shared benchmarking standard for evaluating seizure-detector robustness under realistic clinical stressors, extending pre-deployment assessment beyond clean-data accuracy.
Chinese Translation
尽管癫痫发作检测器在留出的脑电图(EEG)数据上表现优异,但在真实世界的采集变异、伪迹和对抗性输入下可能失效。我们提出了RobustSeiz,一个开源的、模型无关的框架,为癫痫发作检测器在部署前进行压力测试和比较提供了标准化、可复现的协议,测试在受控且具有临床动机的分布偏移下进行。我们将四个公开的头皮脑电数据集(CHB-MIT、TUSZ、Siena和SeizeIT1)统一标准化为BIDS-EEG目录结构,并在留出划分上评估受试者无关的检测器。环境、噪声和对抗性变换在预定义的超参数网格上进行遍历。每次运行报告样本级和事件级的灵敏度、精确率、F1分数、每24小时假阳性次数、发作起点的Lead和Lag时间偏差,以及Monte Carlo dropout预测一致性。RobustSeiz包含基于Docker的GPU流水线、实验注册表,以及完整评估和研究子集两种模式。我们在TUSZ数据集上使用一个当代癫痫发作检测器,在完整实现的偏移网格上演示了该框架;一项AWGN分析说明了扰动强度如何改变检测质量、发作起点时间和预测一致性。RobustSeiz为在真实临床压力因素下评估癫痫发作检测器的鲁棒性提供了一个共享的基准测试标准,将部署前评估从干净数据的准确率扩展到更全面的场景。
cs.LG / 77 / 2609.04010

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

通过离散扩散解锁大语言模型的无损加速
Sahoo, Subham Sekhar, Chen, Lingjie, Pham, Khiem, Geuter, Jonathan, Dwivedi, Chaitanya, Pimpalkhute, Varad, Akhauri, Yash, Moreno, Alexander, Yurochkin, Mikhail, Wang, Zhenting, Elhoushi, Mostafa, Dey, Nolan, Bergsma, Shane, Hestness, Joel, Thickstun, John, Xing, Eric, Liu, Zhengzhong
Abstract
Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution. We decouple the parameters of these models into two sets: AR weights, trained using the standard NTP objective, and lightweight diffusion weights, trained to generate multiple tokens simultaneously. The diffusion weights are learned through a simple Diffusion Distillation phase that adds negligible overhead to existing LLM training pipelines. We also introduce $\Psi$-Spec, a family of samplers that enables lossless acceleration and inference-time scaling at a fixed context length. Unlike speculative decoding, our method requires no separate draft model. Unlike diffusion LLMs (d-LLMs), it accelerates generation without sacrificing the quality of the underlying AR model. The resulting models, called Uno, can be trained from scratch or built by augmenting existing open-weight AR LLMs. Uno achieves higher throughput than leading speculative-decoding methods at every evaluated batch size and delivers up to $3\times$ speedups over the base AR model, including at the largest batch size supported by the device. Notably, our 8B Uno model outperforms the leading open d-LLM, the 26B DiffusionGemma, and the proprietary Mercury 2 across all evaluated benchmarks in agentic tool use, coding, and long-context reasoning. We release code and checkpoints at: https://s-sahoo.github.io/uno/
Chinese Translation
大语言模型(LLM)的巨大成功很大程度上归功于下一词元预测(NTP),但其自回归(AR)结构要求缓慢的逐词元顺序生成。为克服这一瓶颈,我们提出了扩散增强大语言模型(diffusion-augmented LLMs),这是一类新型模型,它在定义自回归模型分布的同时,利用扩散过程从该分布中并行抽取多个词元。我们将这些模型的参数解耦为两组:使用标准NTP目标训练的AR权重,以及通过训练实现同时生成多个词元的轻量级扩散权重。扩散权重通过一个简单的扩散蒸馏(Diffusion Distillation)阶段学习得到,该阶段为现有LLM训练流程增加的开销可以忽略不计。我们还提出了Ψ-Spec,这是一族采样器,能够在固定上下文长度下实现无损加速与推理时扩展。与投机解码(speculative decoding)不同,我们的方法无需单独的草稿模型。与扩散大语言模型(d-LLMs)不同,它能够在不牺牲底层AR模型质量的前提下加速生成。所得到的模型称为Uno,既可以从头训练,也可以通过增强现有的开源权重AR LLM来构建。Uno在所有评估的批量大小下均实现了比领先投机解码方法更高的吞吐量,相比基础AR模型可获得高达3倍的加速,包括在设备所支持的最大批量大小下。值得注意的是,我们的8B Uno模型在智能体工具使用、编程和长上下文推理的所有评估基准上均优于领先的开源d-LLM——26B的DiffusionGemma以及专有的Mercury 2。我们在以下网址发布了代码和检查点:https://s-sahoo.github.io/uno/
cs.LG / 78 / 2609.04018

A location-invariant estimator of extremal quantile treatment effects for heavy-tailed distributions

重尾分布的极高分位数处理效应的位置不变估计量
Yu, Xin, Huang, Shuwei, Liu, Jicheng, Tang, Jielin, Wang, Bolin, Zhang, Yunxiao, Zhao, Tian
Abstract
Quantile treatment effects (QTEs) measure the effect of a treatment on the distribution of an outcome, and their estimation at extreme quantile levels is of central interest in applications where the target quantiles lie far beyond the range of the data. For heavy-tailed potential outcomes, existing extremal QTE estimators rely on extrapolation combined with a causal extreme value index (EVI) estimator, but the resulting estimator is not invariant under a common location shift of the potential outcome distributions, even though the population QTE is. We address this issue in two steps. First, we adapt the location-invariant Fraga estimator of the EVI to the causal setting using inverse propensity score weighting. Second, we replace the original extrapolation formula with a difference-based scheme, under which the location parameter cancels when quantile differences are taken. The resulting QTE estimator is therefore location invariant. We establish the consistency and asymptotic normality of the proposed extremal QTE estimators, and provide a consistent variance estimator, leading to asymptotically valid inference. A simulation study confirms the location invariance, the stability with respect to the threshold, and the coverage of the proposed methods.
Chinese Translation
分位数处理效应(QTE)衡量处理对结果分布的影响,在目标分位数远超数据范围的应用中,其在极端分位数水平上的估计尤为重要。对于重尾的潜在结果,现有的极高分位数处理效应估计量依赖于外推法结合因果极值指数(EVI)估计量,但所得估计量在潜在结果分布发生共同位置平移时并不具有不变性,尽管总体QTE具有这种不变性。我们分两步解决这一问题。首先,利用逆倾向得分加权,将EVI的位置不变Fraga估计量推广到因果设定中。其次,用基于差分的方法替换原来的外推公式,在该方法中,取分位数之差时位置参数会被消去。因此,所得的QTE估计量具有位置不变性。我们证明了所提出的极高分位数处理效应估计量的相合性和渐近正态性,并提供了一个相合的方差估计量,从而实现渐近有效的统计推断。模拟研究证实了所提方法的位置不变性、关于阈值的稳定性以及置信区间的覆盖率。
cs.LG / 79 / 2609.04066

Subspace Inference Enables Efficient Active Reward Learning from Preferences

子空间推断实现从偏好中高效主动奖励学习
Zhou, Yutai, Bıyık, Erdem
Abstract
Reinforcement learning from human feedback (RLHF) has emerged as a powerful yet sample-inefficient approach for learning reward models from human preferences, making active learning a critical component in synthesizing informative preference queries. However, effective uncertainty quantification required for active learning remains a key challenge for large neural network reward models. In this paper, we introduce PreferenceEKF, a sample-efficient approach that tracks reward model uncertainty by framing active preference learning as a sequential Bayesian filtering problem. Instead of relying on computationally prohibitive posterior inference over the full neural network parameter space, our method performs sequential inference via an extended Kalman filter within a low-dimensional parameter subspace, continuously updating the reward model posterior as new preference queries arrive. Our approach enables scalable sampling of neural network parameters to efficiently compute acquisition functions for active reward learning. Experiments on the D4RL and V-D4RL benchmarks demonstrate that our approach achieves better sample efficiency, runtime, scalability, and calibration compared to other Bayesian deep learning approaches, and the learned reward models lead to competitive offline reinforcement learning policy performance. This highlights the potential of scalable Bayesian methods for preference-based reward modeling in RLHF. Our code is available at https://github.com/yutaizhou/bnn_pref.
Chinese Translation
基于人类反馈的强化学习(RLHF)已成为一种从人类偏好中学习奖励模型的强大但样本效率低下的方法,这使得主动学习成为合成信息量丰富的偏好查询的关键组成部分。然而,主动学习所需的有效不确定性量化对于大型神经网络奖励模型而言仍然是一个关键挑战。在本文中,我们提出了 PreferenceEKF,一种样本高效的方法,它将主动偏好学习构建为一个序贯贝叶斯滤波问题,从而追踪奖励模型的不确定性。该方法不依赖于在全网络参数空间上进行计算代价高昂的后验推断,而是在低维参数子空间内通过扩展卡尔曼滤波器(extended Kalman filter)执行序贯推断,随着新偏好查询的到来持续更新奖励模型的后验分布。我们的方法实现了神经网络参数的可扩展采样,从而高效计算用于主动奖励学习的采集函数(acquisition functions)。在 D4RL 和 V-D4RL 基准上的实验表明,与其他贝叶斯深度学习方法相比,我们的方法在样本效率、运行时间、可扩展性和校准性方面均表现更优,且学到的奖励模型能够带来具有竞争力的离线强化学习策略性能。这凸显了可扩展贝叶斯方法在 RLHF 中基于偏好的奖励建模方面的潜力。我们的代码可在 https://github.com/yutaizhou/bnn_pref 获取。
cs.LG / 80 / 2609.04090

Conditioning Degenerate Diffusion Models

条件化退化扩散模型
Aydın, Uğur, Başar, Tamer
Abstract
Current conditioned generative models heavily rely on score functions for guidance during training. When the generative model is a diffusion process with a singular diffusion coefficient and the underlying (conditional) densities either do not exist or are not smooth, we use causal optimal transport to define \emph{approximate} loss functions that identify a minimum-entropy control for guidance under minimal assumptions. Our approach relies on causal optimal transport and its characterization through the predictable representation property of (conditioned) diffusion processes whose associated martingale problem is well posed, \`a la \"Ust\"unel.
Chinese Translation
当前的条件生成模型在训练过程中高度依赖于得分函数(score functions)进行引导。当生成模型是一个具有奇异扩散系数的扩散过程,且其潜在的(条件)密度不存在或不光滑时,我们利用因果最优传输(causal optimal transport)定义了近似损失函数,在最少的假设条件下识别出用于引导的最小熵控制。我们的方法依赖于因果最优传输及其通过(条件化)扩散过程的可料表示性质(predictable representation property)的刻画,其中这些扩散过程所对应的鞅问题在 Üstünel 意义下是适定的。
cs.LG / 81 / 2609.04105

Hardware-Aware FP4 FlashAttention-4

硬件感知的FP4 FlashAttention-4
Hu, Robert
Abstract
Blackwell's 4-bit floating-point (FP4) tensor cores do not automatically make attention faster because softmax conversion and on-chip dependencies dominate once its matrix products shrink. We address this with \emph{Direct-P} for noncausal inference and a causal path that passes the forward quantization directly into backward. Direct-P maps scores directly to FP4 probabilities and reaches up to 2.13$\times$ the bfloat16 (BF16) forward throughput on an NVIDIA GB200. The causal path reconstructs probabilities from saved quantized queries and keys and uses 8-bit floating-point (FP8) gradient operands, accelerating a complete single-GPU 8-billion-parameter update by up to 1.14$\times$. Matched distributed training retains FP8 probabilities and values; every tested MXFP4 probability/value training trajectory diverges.
Chinese Translation
Blackwell的4位浮点数(FP4)张量核心并不能自动使注意力计算更快,因为一旦其矩阵乘积规模缩小,softmax转换和片上依赖就会成为主导瓶颈。我们通过面向非因果推理的Direct-P方法以及一种将前向量化直接传递到反向传播的因果路径来解决这一问题。Direct-P将得分直接映射为FP4概率,在NVIDIA GB200上达到了bfloat16(BF16)前向吞吐量的最高2.13倍。因果路径从保存的量化查询和键中重构概率,并使用8位浮点数(FP8)梯度操作数,将完整的单GPU80亿参数更新加速最高1.14倍。在匹配的分布式训练中保留FP8概率和值;而所有测试过的MXFP4概率/值训练轨迹均发散。
cs.LG / 82 / 2609.04113

Constant regret in general games via higher-order optimism

通过高阶乐观机制实现一般博弈中的常数遗憾
Abbadi, Omar, Laraki, Rida, Mertikopoulos, Panayotis
Abstract
We introduce an uncoupled learning algorithm which, when employed by all players of an arbitrary $N$-player normal form game with up to $K$ actions per player, guarantees $O(N^3\log^2 K)$ individual regret, uniformly over the horizon of play. The proposed algorithm - which we call higher-order optimism with discounting (HOOD) is a variant of optimistic follow-the-regularized-leader (OptFTRL) that combines a discounted $(N+1)$-th order predictor with entropic regularization over a suitable "lifting" of the game's strategy space. This combination of ingredients is purposefully designed to dampen large oscillations of the induced sequence of play in a controlled manner, removing in this way a key stumbling block of previous attempts to achieve constant regret in general games. Our approach bears several striking similarities to the concurrent - and completely independent - work of Liu, Farina, and Ozdaglar (arXiv:2608.31166), who very recently derived an $O(N^{21}\log^{4} K)$ regret bound through the use of higher-order optimism and an exponential moving average estimator.
Chinese Translation
我们提出了一种非耦合(uncoupled)学习算法,当该算法被任意具有至多 $K$ 个动作/玩家的 $N$ 人标准形博弈中的所有玩家采用时,能够保证在整个博弈时间范围内一致地实现 $O(N^3\log^2 K)$ 的个体遗憾。所提出的算法——我们称之为带折扣的高阶乐观算法(higher-order optimism with discounting, HOOD)——是乐观跟随正则化领导者(optimistic follow-the-regularized-leader, OptFTRL)算法的一个变体,它将对博弈策略空间进行适当“提升”(lifting)后得到的折扣 $(N+1)$ 阶预测器与熵正则化相结合。这些要素的组合旨在以受控的方式抑制诱导博弈序列的大幅震荡,从而消除了以往在一般博弈中实现常数遗憾尝试中的一个关键障碍。我们的方法与同期且完全独立的工作(Liu、Farina 和 Ozdaglar,arXiv:2608.31166)存在若干显著的相似之处,他们最近通过使用高阶乐观机制和指数移动平均估计器推导出了 $O(N^{21}\log^{4} K)$ 的遗憾界。
cs.LG / 83 / 2609.04134

Prospective Coding Improves Learning in Deep Continuous-Time Recurrent Networks

前瞻性编码改进深度连续时间循环网络的学习
Rawat, Shivang, Morello, Mirko, Morone, Flaviano, Heeger, David J.
Abstract
Temporal integration gives continuous-time recurrent networks memory, but in deep stacks it also delays bottom-up signals and attenuates top-down errors. We develop Recursive Quadrature Filters (RQFs), biologically motivated complex-valued temporal filters that are a special case of diagonal state-space models (SSMs), and ask whether this failure mode can be addressed by making each layer's bottom-up input prospective. Starting from an energy model, we derive the RQF dynamics and show that each RQF is a band-pass filter whose learnable parameters control its tuning frequency and bandwidth. We then make each layer's bottom-up input prospective using a parameter-free two-tap update that leaves the recurrent transition and parallel scan unchanged. We extend this correction to general diagonal SSMs and show that it mitigates depth-dependent gradient attenuation when temporal gradients are truncated, i.e., spatial-only backpropagation. We evaluate the intervention in RQFs, S5, and ORGaNICs (a nonlinear gated RNN) trained using full backpropagation through time (BPTT) and spatial-only backpropagation. Under full BPTT, prospective variants match or outperform their non-prospective controls in every model and configuration. A non-residual width-32 six-layer RQF reaches 96.09% accuracy on raw-audio Speech Commands with 31.9k parameters; a width-64 six-layer RQF reaches 83.56% on the 16,384-step Path-X task. These results identify RQFs as a parameter-efficient recurrent substrate and prospective-input coding as an input-side correction for deep continuous-time recurrent networks.
Chinese Translation
时间积分赋予连续时间循环网络记忆能力,但在深层堆叠结构中,它也会延迟自底向上的信号并衰减自顶向下的误差。我们提出了递归正交滤波器(Recursive Quadrature Filters, RQFs),这是一类受生物学启发的复值时间滤波器,是对角状态空间模型(SSMs)的一个特例,并探究是否可以通过使每一层的自底向上输入具有前瞻性(prospective)来解决这一失效模式。从能量模型出发,我们推导了RQF的动力学,并证明每个RQF都是一个带通滤波器,其可学习参数控制着调谐频率和带宽。随后,我们采用一种无参数的两抽头更新方法使每层的自底向上输入具有前瞻性,同时保持循环转移和并行扫描不变。我们将这一修正扩展到一般的对角SSMs,并证明在时间梯度被截断(即仅空间反向传播)的情况下,它能缓解随深度增加的梯度衰减。我们在使用完整时间反向传播(BPTT)和仅空间反向传播训练的RQFs、S5以及ORGaNICs(一种非线性门控循环神经网络)上评估了该干预方法。在完整BPTT下,前瞻性变体在所有模型和配置中均匹配或超越了非前瞻性对照组。一个非残差、宽度为32的六层RQF仅用31.9k参数即可在原始音频Speech Commands任务上达到96.09%的准确率;一个宽度为64的六层RQF在长达16,384步的Path-X任务上达到83.56%的准确率。这些结果将RQFs确立为一种参数高效的循环网络基底,并将前瞻性输入编码确立为深度连续时间循环网络的一种输入端修正方法。
cs.LG / 84 / 2609.04147

A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle

一个面向端到端自动驾驶研究的低成本开放平台:基于微型阿克曼车辆
Couto, Gustavo Claudio Karl, Antonelo, Eric Aislan, Zipperer, Gabriel George
Abstract
This paper presents a low-cost, open experimental platform for research in end-to-end autonomous driving with miniature Ackermann vehicles. The platform combines a physical vehicle, a printed urban track, data collection tools, trajectory registration, and a Webots digital twin, enabling controlled experiments that connect simulation-based autonomous-driving methods to real-world execution. As a first baseline, we implement command-conditioned behavior cloning, in which a neural policy receives an on-board camera image and a high-level navigation command and outputs steering and speed. The system is evaluated both on the physical vehicle and in simulation. In real closed-loop experiments, the learned policy follows lanes and executes commanded turns, reaching a mean cross-track error of 6.1 cm with respect to the reference route, close to the 4.7 cm observed in human demonstrations. In the digital twin, camera field of view has a strong effect on performance, reducing the mean cross-track error from 35.6 to 3.3 cm when widened from 58 to 120 degrees. Using the digital twin to generate synthetic driving data and a learned sim-to-real image translator to reduce the appearance gap, we further show that a higher-capacity policy trained on this synthetic data combined with real demonstrations is the only configuration that completes all four track routes in closed loop, whereas the compact baseline and the same network trained on real data alone complete fewer. These results establish the open platform as a practical testbed for sim-to-real studies and provide an initial command-conditioned imitation-learning baseline; we release it to support reproducible research.
Chinese Translation
本文提出了一个低成本、开放的实验平台,用于基于微型阿克曼(Ackermann)车辆的端到端自动驾驶研究。该平台由实体车辆、印刷的城市赛道、数据采集工具、轨迹配准工具以及 Webots 数字孪生组成,支持可控实验,将基于仿真的自动驾驶方法与真实世界执行相连接。作为首个基线方法,我们实现了条件指令行为克隆(command-conditioned behavior cloning),其中神经策略接收车载相机图像和高层导航指令,并输出转向和速度控制量。该系统在实体车辆和仿真环境中均进行了评估。在真实闭环实验中,学习得到的策略能够跟随车道并执行指定转弯,相对于参考路线的平均横向误差为 6.1 厘米,接近人类演示中的 4.7 厘米。在数字孪生中,相机视场角对性能有显著影响:当视场角从 58 度扩大到 120 度时,平均横向误差从 35.6 厘米降至 3.3 厘米。利用数字孪生生成合成驾驶数据,并结合学习得到的仿真到真实(sim-to-real)图像转换器来弥合外观差异,我们进一步表明,在该合成数据与真实演示数据联合训练下得到的高容量策略,是唯一能在闭环中完成全部四条赛道路线的配置;而紧凑型基线以及仅用真实数据训练的相同网络完成的路线数量均较少。这些结果证明了该开放平台是开展仿真到真实研究的实用测试平台,并提供了首个条件指令模仿学习基线;我们将其开源发布,以支持可复现研究。
cs.LG / 85 / 2609.04189

Robust PAC Learning of Concurrent Stochastic Games

并发随机博弈的鲁棒PAC学习
He, Angel Y., Parker, David
Abstract
We introduce the first Probably Approximately Correct (PAC) learning framework for general-sum concurrent stochastic games (CSGs) with transition uncertainty, while addressing the challenge of Nash equilibrium (NE) existence. Our algorithm maintains data-driven $L^1$ confidence sets over transition kernels and solves a robust CSG to compute a social-welfare optimal $\varepsilon$-NE, using a robust MDP-based exploration mechanism to drive joint state-action coverage. Crucially, we introduce a Nash margin characterisation that enables principled reasoning about equilibrium existence: the framework either returns an $\varepsilon$-approximate NE whose social-welfare value is $\varepsilon$-close to optimal, or provides a sound certificate that no exact NE exists. Under a minimum reachability condition $p_{\mathrm{reach}}>0$ over relevant state-action pairs, the algorithm terminates after a polynomial number of trajectory samples, with sample complexity $\widetilde{O}\left( {R_{\max}^2 H^4 |S|^2 |A| / (p_{\mathrm{reach}} \varepsilon^2)} \right)$. Empirical results on benchmark CSGs demonstrate near-optimal performance, correct handling of equilibrium (non-)existence, and sample complexity consistent with theory.
Chinese Translation
我们提出了首个针对具有转移不确定性的广义和并发随机博弈(Concurrent Stochastic Games, CSGs)的可能是近似正确(Probably Approximately Correct, PAC)学习框架,同时应对纳什均衡(Nash Equilibrium, NE)存在性这一挑战。我们的算法在转移核上维护数据驱动的 $L^1$ 置信集,并求解一个鲁棒CSG以计算社会福利最优的 $\varepsilon$-NE,同时采用基于鲁棒MDP的探索机制来驱动联合状态-动作覆盖。关键的是,我们引入了一种纳什裕度(Nash margin)刻画方法,使得能够对均衡存在性进行有原则的推理:该框架要么返回一个社会福利值与最优值 $\varepsilon$-接近的 $\varepsilon$-近似NE,要么提供一个证明不存在精确NE的可靠证书。在相关状态-动作对上满足最小可达性条件 $p_{\mathrm{reach}}>0$ 的情况下,该算法在多项式数量的轨迹样本后终止,样本复杂度为 $\widetilde{O}\left( {R_{\max}^2 H^4 |S|^2 |A| / (p_{\mathrm{reach}} \varepsilon^2)} \right)$。在基准CSG上的实证结果表明其具有接近最优的性能、能够正确处理均衡(不)存在性,且样本复杂度与理论一致。