Recent advances in humanoid robotics have produced diverse skills through reinforcement learning, motion imitation, and generative modeling. Yet these capabilities remain siloed because they are built around incompatible representations, interfaces, and controllers. We present CHOREO, a framework for training-free composition of heterogeneous humanoid skills. Our key observation is that, regardless of how a skill is learned, it can ultimately be expressed as an executable motion trajectory. Based on this observation, CHOREO converts each capability into SkillMotion, a unified representation that combines motion states, contacts, semantics, and boundary conditions. Skills are composed through direct continuation, cubic Hermite blending, or validated bridge motions, without retraining source models or updating models at test time. On Unitree G1 in MuJoCo, CHOREO organizes 2,950 admitted SkillMotion assets derived from heterogeneous sources and achieves 95.4\% sequence success across 130 multi-action tasks, including 93.8\% success on eight-action sequences. These results demonstrate that executable trajectories provide a scalable interface for accumulating and composing pretrained humanoid capabilities.
The Dynamic Window Approach (DWA) is widely used for local navigation, but its performance depends strongly on parameters that are typically fixed before navigation. In particular, the appropriate prediction horizon can vary with local free space: longer horizons support efficient motion in open areas, whereas shorter horizons help preserve feasible motions in narrow or cluttered regions. This paper proposes D3DWA, an adaptive DWA framework based on a Dueling Double Deep Q-Network (D3QN), which jointly selects the DWA evaluation weights and prediction horizon from a continuous navigation state at every control step while retaining DWA's trajectory generation and collision checking. In eight simulated environments, including unseen layouts, D3DWA reached every goal. Real-robot experiments further showed that D3DWA completed all three tested configurations, including a constrained case in which the weights-only variant timed out. These results demonstrate the benefit of jointly adapting the evaluation weights and prediction horizon. Additional material is available at https://mertcookimg.github.io/d3dwa/
DeViGrasp: Robust Visual Mobile Grasping for Quadruped Manipulators under Degraded Perception
DeViGrasp:退化感知下四足操作臂的鲁棒视觉移动抓取
Zhou, Liang, Su, Jiaming, Wei, Yancong, Dong, Kangkang, Liu, Houde
Abstract
Quadruped manipulators enable mobile grasping in complex environments, yet their whole-body control policies remain vulnerable to unreliable onboard visual perception. Existing methods are typically developed under relatively reliable observations and have not systematically examined how occlusion, segmentation-mask dropout, depth noise, and target-localization jitter affect grasp reasoning and target tracking. To address this gap, we introduce DeViGrasp-Bench, a benchmark for mobile grasping under degraded vision that incorporates controlled visual degradations, seen and unseen objects, multiple difficulty levels, and complex terrains, and evaluates task success, execution efficiency, and action smoothness. We further propose DeViGrasp-Net, a teacher--student framework that combines state-conditioned grasp reasoning with reliability-aware temporal target estimation. The privileged teacher attends to offline grasp candidates conditioned on object, robot, end-effector, and task states, while the deployable student fuses dual-view segmented-depth observations with current, memory, and recovery target hypotheses through Target Hold Memory and Temporal Memory Attention. DeViGrasp-Net outperforms VBC across degradation levels, unseen objects, and complex terrains, and surpasses an adapted DQ-Net across all evaluated degradation levels. Under the Difficult setting, it achieves a success rate of 62.3\%, improving upon VBC and DQ-Net by 16.1 and 4.3 percentage points, respectively; under the Hard setting, its margin over DQ-Net increases to 10.5 percentage points. Ablation studies confirm the complementary benefits of grasp-aware supervision and reliability-aware temporal memory.
ORDER: A Fictitious-World Benchmark for Domain-Adaptive Embodied AI
ORDER:一个面向领域自适应具身智能的虚构世界基准
Sathi, Sai Krishna Reddy, Tiwari, Anuj
Abstract
Adapting language models to new domains via continual pre-training raises a basic evaluation problem: if the training corpus overlaps with what the model already knows, performance gains cannot be cleanly attributed to new learning rather than pre-existing knowledge. This matters most for knowledge-intensive, task-light (KHTL) robot deployments - pharmaceutical dispensing, hazardous-material handling, facility-specific protocols, where the physical task is simple but the governing rules are proprietary and safety-critical, and where extensive live testing is costly or unsafe. We introduce ORDER (Ontology-driven Decision-making for Embodied Reasoning), a benchmark built on a fictitious world: a 342,069-token synthetic corpus defining a self-consistent physics that cannot appear in any model's pre-training data. ORDER pairs a 500-question knowledge test (ORDER-BENCH) with a harder compositional task, ORDER-SPATIAL: ordering objects for safe manipulation across both familiar and entirely novel scenes. GPT-4.1 without adaptation scores below chance on ORDER-SPATIAL (Kendall's tau = 0.441), showing its priors actively conflict with the invented physics. After continual pre-training, small models improve substantially on both familiar and novel scenes alike evidence of genuine world-model induction rather than memorization. We then carry this through to a robot pipeline: models that answer the knowledge test well often cannot produce valid, executable plans without a further skill-adaptation stage, after which small, fully offline models outperform GPT-4.1 even when GPT-4.1 is given retrieval access to the same rules (Kendall's tau = 0.848 vs. 0.606), on a full perception-to-execution loop demonstrated on a simulated iiwa7 arm with human-in-the-loop correction. Throughout, ORDER-SPATIAL performance, not knowledge-test accuracy is what predicts real plan quality.
Chinese Translation
通过持续预训练将语言模型适配到新领域时,存在一个基本的评估问题:如果训练语料与模型已知内容重叠,性能提升就无法清晰地归因于新知识的学习而非已有知识。这一问题对于知识密集、任务轻量(KHTL)的机器人部署场景最为关键——例如药品配发、危险品处理和机构专用规程——在这些场景中,物理任务本身简单,但约束规则是专有的且关乎安全,而大量的实机测试成本高昂或存在危险。我们提出了ORDER(Ontology-driven Decision-making for Embodied Reasoning,面向具身推理的本体驱动决策),一个建立在虚构世界之上的基准:其包含一个342,069词元的合成语料库,定义了一套自洽的物理体系,不可能出现在任何模型的预训练数据中。ORDER将一个500题的知识测试(ORDER-BENCH)与一个更具挑战性的组合任务ORDER-SPATIAL相结合:在熟悉场景和全新场景中进行面向安全操作物体排序。未经适配的GPT-4.1在ORDER-SPATIAL上的得分低于随机水平(Kendall's tau = 0.441),表明其先验知识与所发明的物理体系存在主动冲突。经过持续预训练后,小模型在熟悉场景和全新场景上均获得显著提升,这证明模型是真正归纳出了世界模型而非死记硬背。随后,我们将这一流程延伸至机器人管线:在知识测试中表现良好的模型,往往仍无法在不经过进一步技能适配阶段的情况下生成有效、可执行的计划;经过技能适配后,小型的完全离线模型超越了GPT-4.1——即使为GPT-4.1提供对相同规则的检索访问(Kendall's tau = 0.848 对 0.606)。整个流程在一个带有遥操作人类在环校正的仿真iiwa7机械臂上,通过完整的感知到执行闭环进行了验证。贯穿全文的核心发现是:真正能预测真实计划质量的是ORDER-SPATIAL上的表现,而非知识测试的准确率。
OJOx: Specification-Conditioned Demonstrations for Embodied AI in Construction
OJOx:面向建筑具身智能的规格条件化示范
Dawod, Mohamed
Abstract
Large-scale egocentric and whole-body human demonstrations are becoming a primary source of data for embodied intelligence. They record what people perceive and do, but rarely the external specification that gave an action its purpose. In construction that omission is consequential: skilled work is directed at project-specific configurations defined in a design model - configurations not yet present in the environment being observed. A mason's transferable competence is not the geometry of one wall but the ability to realise a new geometry from a specification. We introduce the specification-conditioned demonstration: a synchronised record of the physical state a demonstrator perceives, the intended state supplied to them by an external design, and the behaviour connecting the two. We present OJOx, a capture interface that realises this for construction - delivering design geometry to a headset, anchoring it in the physical workspace, rendering it into a demonstrator's stereo passthrough view, and recording that view synchronously with whole-body and hand motion. We report one fully instrumented session - a 33-component wall laid against a specification that changes while the work proceeds - and check the record against the physical scene through an external camera registered independently of the capture. Recorded sessions remain compatible with existing humanoid retargeting infrastructure and replay onto a Unitree G1 in simulation. The result is a data interface for testing whether embodied policies can learn not merely to imitate demonstrated actions, but to act toward specifications absent from their training experience.
When Does Test-Time Physical Diagnosis Pay? A Frozen Policy Buys Evidence It Never Reads
测试时物理诊断何时有价值?冻结策略获取了它从未读取的证据
Zhang, Zhengshu
Abstract
When a robot faces unfamiliar physical conditions, a common approach is to collect evidence about what changed and adapt. For such diagnosis to improve behavior, six ordered empirical conditions must hold: a meaningful reference, identifiability of the physical condition, use of the acquired evidence, decision value, selection value over a fixed alternative, and safe realization. We test this chain in controlled and public environments. It holds end to end in our controlled environments. After transfer to unseen mechanisms, however, it breaks at evidence use. On decisions requiring the full trace, the frozen decoder does not change its choice. A linear model using only trace increments recovers the correct choice on mechanisms excluded from fitting, showing that the trace is informative but unused. The failure is concentrated at the richest evidence level: those decisions fall to chance, while decisions settled with lower-cost evidence remain correct, a split hidden by aggregate accuracy. The same chain can fail at other links in public environments. Successful physical identification therefore guarantees neither evidence use nor useful adaptation; evaluation should identify where the chain breaks rather than rely on recovery accuracy or aggregate performance alone.
World models support autonomous driving by predicting the scene evolution associated with candidate trajectories. Driving dynamics differ in predictability, motivating a distinction between regular evolution and event-induced deviations that call for selective correction. We propose EditWM, a world model that decomposes future prediction into normal evolution and event-driven incremental correction in compact visual feature space. A trajectory-conditioned normal predictor provides the base forecast and is then frozen for correction learning. A correction decoder compares this forecast with observation history and planned actions, producing a bounded feature update whose contribution is regulated by a learned gate. The corrected future features condition trajectory scoring through candidate-specific cross-attention, linking world modeling to plan selection. At inference, EditWM uses only past and current observations, ego state, and candidate trajectories. Across all 12,146 NAVSIM navtest scenes, expert-trajectory-conditioned evaluation shows a 5.35\% reduction in future-feature MSE over Normal, with improvements in 83.54\% of scenes. The system achieves 91.05 EPDMS on a 100-point scale using the official EPDMS evaluator. These results demonstrate improved future-feature prediction and competitive trajectory selection when corrected future representations are integrated into planning.
The rapid progress of artificial intelligence is reshaping robotics and accelerating the adoption of learning-based approaches. While purely data-driven methods have achieved remarkable success in computer vision and natural language processing, robotics remains constrained by limited data, complex real-world interactions, and the need for reliable operation. These challenges have motivated the exploration of physics-embedded robot learning, which embeds physics priors into learning algorithms. By encoding the underlying physical laws and constraints, physics priors can complement limited data with robotics-specific inductive biases, potentially improving generalization, interpretability, and sample efficiency. However, the literature on physics-embedded robot learning remains fragmented across terminology, methodologies, and application domains, making it difficult to assess this growing body of work. This survey reviews physics-embedded robot learning across a broad range of physics priors, robotics applications, and machine learning models, from single-layer perceptrons to generative foundation models. We adopt a unified taxonomy that classifies existing approaches according to their physics embedding: physics-guided inputs, data, and representations; physics-encoded model architectures; and physics-informed training loss functions. Building on this taxonomy, we review methods for robot dynamics learning, trajectory planning, prediction, control, and estimation, together with the corresponding open-source software ecosystem. We identify key open challenges, and outline promising future research directions. Overall, we argue that physics priors provide a particularly relevant robotics-specific inductive bias, complementing rather than replacing data-driven learning, and paving the way toward more generalizable, data-efficient, and trustworthy robotic systems.
Large language models have shown considerable potential for natural-language-driven parametric CAD modeling. However, a fundamental contradiction exists between their probabilistic generation and the deterministic requirements of CAD modeling, resulting in limitations in reliability, design-intent preservation, and geometric validity. Existing methods typically rely on large-scale annotated datasets, lack explicit modeling of design intent, and underutilize the deterministic capabilities of CAD kernels. To address these limitations, we propose ReliCAD, a unified framework that transforms uncertain LLM generation into reliable parametric CAD modeling. Through explicit design-intent modeling, ReliCAD converts user requirements into structured design specifications and explicitly models geometric relations, topological dependencies, and feature construction order. It then generates constraint-aware parametric instructions and invokes the CAD kernel through an Agent-ready API to perform geometric construction and constraint solving. ReliCAD further records runtime evidence and employs a verification-feedback mechanism to assess consistency between the generated model and the design specifications, enabling error localization and iterative repair. Experiments on the public HistCAD generation dataset and our multi-granularity CAD editing dataset demonstrate that ReliCAD significantly outperforms baseline methods, achieving 99.8\% validity rate and 0.8753 IoU. ReliCAD provides a verifiable, repairable, and generalizable approach to natural-language-interactive CAD modeling.
Generalizable robot manipulation requires predicting how a scene will evolve, identifying where interactions are feasible, and determining how to act. Action-labeled robot videos directly supervise control but are costly and limited in diversity, whereas egocentric human videos capture diverse interactions but lack robot actions and differ in embodiment and appearance. We introduce AffordanceWAM, an affordance-aware generative World Action Model that represents object-centric spatiotemporal affordance through Scalar Affordance and Affordance Heatmap, within the generated future World. This representation grounds visual prediction in task-relevant objects and interaction regions for action generation, and provides shared interaction targets across human and robot videos. Built on a pretrained video diffusion Transformer, AffordanceWAM uses separately parameterized World and Action Experts, coupled through Masked Joint Self-Attention, to jointly predict future RGB observations, Scalar Affordance fields, Affordance Heatmaps, and continuous robot actions under a unified flow-matching objective. Human videos supervise all three future-World streams, whereas robot trajectories additionally provide action supervision, enabling transfer without human action labels or retargeting. Experiments on RoboCasa, CALVIN ABC$\rightarrow$D, and real-world manipulation demonstrate consistent gains over RGB-only and robot-data-only baselines. Under fixed robot supervision, RoboCasa performance improves monotonically as affordance-annotated human video scales. These results support affordance as an effective interface for both vision-language-action learning and human-to-robot transfer.
Achieving high task success and efficiency remains a central pursuit in embodied exploration. Existing frameworks typically guide agent behavior through a spatially coarse and indirect assessment of suggestive cues and directions, yet such designs may struggle to judge cue sufficiency and the need for further inspection, leading to early termination or excessive continuation and ultimately reducing task success and efficiency. This work rethinks embodied exploration from a spatially explicit standpoint and introduces the spatial-inspection-guided ACtive Exploration (ACE) framework coupling evidence-grounded perception with exposure-informed movement and establishing a spatially resolved decision paradigm. Evidence-grounded perception strengthens inferential validity by integrating granular localization with focused verification, turning suggestive cues into precise decisive visual support. Exposure-informed movement promotes directional judgment through prospective prioritization and retrospective suppression, selecting promising directions for efficient spatial progress. ACE alleviates the tension between preventing early termination and avoiding unnecessary continuation. Extensive experiments demonstrate that ACE achieves 18.0% higher navigation task success and 10.3% higher exploration efficiency for question answering than prior state-of-the-art baselines, advancing effective and efficient embodied exploration.
Chinese Translation
实现高任务成功率与高效率始终是具身探索的核心追求。现有框架通常通过对提示性线索和方向进行空间上粗略且间接的评估来引导智能体行为,然而此类设计可能难以判断线索是否充分以及是否需要进一步检查,从而导致过早终止或过度延续,最终降低任务成功率与效率。本工作从空间显式的角度重新思考具身探索,提出了空间检查引导的主动探索(spatial-inspection-guided ACtive Exploration,ACE)框架,该框架将基于证据的感知与基于暴露信息的移动相结合,建立了一种空间分辨的决策范式。基于证据的感知通过将细粒度定位与聚焦验证相融合来增强推理有效性,将提示性线索转化为精确且具有决定性的视觉支持。基于暴露信息的移动通过前瞻性优先排序与回溯性抑制来促进方向判断,选择有前景的方向以实现高效的空间推进。ACE缓解了防止过早终止与避免不必要延续之间的矛盾。大量实验表明,ACE的导航任务成功率比此前最先进的基线方法高出18.0%,问答探索效率高出10.3%,推动了高效具身探索的发展。
Over the last decade, behavior trees (BT) have become one of the dominant behavior models for coordinating missions of robotic systems. Yet empirical evidence on BT adoption in real-world contexts remains limited, especially regarding practitioners experiences and practices. Without practitioner-grounded evidence from both academic and industrial settings, the research community risks developing guidelines and tools that are plausible in principle but only partly aligned with real challenges. This scarcity of evidence leads to ad-hoc practices, impeding software reuse, maintenance, and evolution. This paper reports a mixed-methods study combining a technical action research investigation at an automotive company with a survey of 34 robotics practitioners. Our results indicate that BTs improve team communication and the understandability of robotic decision-making logic, reflecting BTs practical value beyond mission coordination. At the same time, practitioners face multiple non-trivial design and integration decisions, complicated by the lack of adequate guidelines and tool support. These decisions span architectural and language choices when implementing BTs and integrating them within ROS, for which we report observed patterns. Additionally, practitioners faced multi-factor granularity decisions and mixed experiences with current libraries. For the most used BT libraries, they reported implementation challenges and documentation gaps. Regarding planning algorithms, practitioners find deciding optimal BT node ordering challenging. Views on scalability remained inconclusive, and limitations in current libraries and practices hinder broader adoption. We conclude with cross-cutting observations and implications for both practitioners adopting BTs in robotic systems and researchers aiming to advance empirical understanding of BT adoption.
VLPSA: Vision-Language-Poisson-Safe Actions for Full-Body Safety of Learned Policies
VLPSA:面向学习策略全身安全的视觉-语言-泊松安全动作
Wilkinson, Meg, Fourney, Emily, Burdick, Joel W., Ames, Aaron D.
Abstract
Vision-language-action (VLA) models enable increasingly general-purpose robotic manipulation, but such learned policies do not provide safety guarantees for collision avoidance---especially in environments outside of training distributions. This work presents Vision-Language-Poisson-Safe Actions (VLPSA), a safety filtering framework that provides full-body safety for VLA policies in cluttered and dynamic environments without retraining. VLPSA synthesizes Poisson Safety Functions (PSF) online from perception data, yielding a Control Barrier Function (CBF) that is enforced through a CBF-QP safety filter over the full body and any grasped object, treated as an extension of the final robot link. To enable real-time deployment while maintaining fine spatial resolution in critical task regions, VLPSA combines dual resolutions of this PSF using Boolean CBF compositions. We evaluate VLPSA on SafeLIBERO against safety-filtering baselines, where it achieves the highest collision avoidance rate among the evaluated methods, increasing collision avoidance from 23.1% for the base $\pi_{0.5}$ policy to 91.2% while surpassing its task success rate. We further deploy VLPSA on a Franka FR3 in cluttered scenes with dynamic obstacles and human interference, demonstrating real-time full-body safety during manipulation tasks.
Tracker-Free Robotic Ultrasound Calibration with a Spherical-Marker Phantom and Threshold-Free Center Localization
基于球形标记体模与免阈值中心定位的无跟踪机器人超声标定
Koo, Kyoungmo, Ma, Guangshen, Wang, Xueding, Draelos, Mark
Abstract
Robotic ultrasound (US) calibration is essential for accurately relating US images to the robot coordinate system, but accurate and automated calibration remains challenging because existing methods often require complex phantoms or external 3D trackers. In this work, we develop a tracker-free robotic US calibration framework using a spherical-marker phantom and a highly automated perception pipeline for sphere-center localization. The proposed threshold-free image-processing method localizes the spherical feature based on each US image's intensity distribution, eliminating hand-tuned intensity thresholds and improving robustness across imaging settings without per-system retuning. Multi-pose observations of the spherical fiducial are then used to estimate the US-to-EE transformation without external tracking or prior localization of the sphere center in the robot base frame. We validate the proposed framework on a robotic US platform through repeated sphere scans and further assess the calibrated system using geometrically distinct phantoms with known CAD models. Across 12 cross-validation folds, single-marker calibration achieved sphere-center accuracy and precision of $1.75\pm0.51$ mm and $0.85\pm0.11$ mm, respectively, compared with $1.72\pm0.51$ mm and $0.84\pm0.12$ mm for three-marker calibration. Across the three reconstruction sets, the single-marker calibration yielded pooled post-registration point-to-surface MAE$\pm$SD values of $0.41\pm0.35$ mm for the cone and $0.50\pm0.41$ mm for the triangular prism, closely matching the three-marker results of $0.40\pm0.35$ mm and $0.48\pm0.40$ mm, respectively.
Cong, Qingzheng, Devillard, Alexis W. M., Dawood, Abu Bakar, Zhang, Xinxin, Fan, Wen, Dei, Neri Niccolò, Suulker, Cem, Althoefer, Kaspar, Burdet, Etienne, Zhang, Dandan
Abstract
This paper presents a stacked two-layer force-sensing resistor (FSR) array designed for robotic fingertips that combines high-resolution pressure mapping with shear-force estimation. A compliant lattice elastomer spacer converts shear loading into a measurable inter-layer displacement, producing relative center-of-pressure (CoP) shifts between layers. A physics-based moment balance links inter-layer CoP displacement to shear force, while an end-to-end CNN--GRU model captures nonlinear effects from load-dependent compression and contact redistribution. This model, with both layers as input, achieves coefficients of determination $R^2 = 0.914$ for $F_x$ and $R^2 = 0.944$ for $F_y$, consistently outperforming single-layer baselines for shear-force estimation. Robotic manipulation experiments show that, for contact-motion tracking, the deep layer tracks the translation and rotation imposed by the robot arm, whereas the superficial layer tracks the slip at the contact surface. Transient changes in the difference between the total pressure responses of the two layers provide the best slip-event detection performance among the tested cues. These results demonstrate that two-layer FSR arrays can provide three-axis force estimation, contact-motion tracking, and slip-event detection beyond conventional normal-force sensing.
Latent Policy Steering: An Efficient and Flexible Framework for Cross-Embodiment Transfer
潜在策略引导:一种高效灵活的跨本体迁移框架
Wang, Yiqi, Verghese, Mrinal, Schneider, Jeff
Abstract
The performance of learned robot visuomotor policies depends heavily on the size and quality of their training data, yet collecting high-quality demonstrations remains costly for robots in the real world. Although large-scale robot and human datasets are increasingly available, embodiment gaps and mismatched action spaces make them difficult to leverage directly. Cross-embodiment transfer, reusing experience from other embodiments to improve learning on a target embodiment, is therefore crucial for scaling robot learning beyond per-robot data collection. In this work, we find that efficient transfer can be achieved by learning from what is shared across embodiments, the visual dynamics of how the world responds to motion, and by effectively exploiting the scarce target-embodiment data at test time. The proposed framework, called Latent Policy Steering (LPS), implements an embodiment-agnostic pretraining phase, which trains an image-based World Model (WM) with optical flow across diverse embodiments. The resulting WM is finetuned on the target embodiment with robot actions. It then steers the base policy toward better actions by searching in the WM's latent space for plans that stay close to the finetuning data. LPS is a policy-agnostic framework: it can flexibly accommodate different policies without having to retrain them. In Robomimic and real-world evaluations, LPS improves the average performance of Diffusion Policy relatively by 16% and 62%, and Pi0.5 by 8% and 14%, with only 50 demonstrations on an unseen target embodiment.
Large language model (LLM) planners can decompose natural-language instructions and select reusable robot skills, but choosing the correct skill does not guarantee successful physical execution. This gap is especially important in humanoid loco-manipulation, where errors during approach, grasping, transport, or placement can invalidate the remainder of a long-horizon plan. We present FRAMES, a failure-aware supervisory framework for the Unitree G1 humanoid that operates above the CEER whole-body controller. A Planner Agent selects subtasks through parameterized mid-level skills, while a vision-language-model-based Monitor Agent evaluates each skill using temporal multi-view observations and structured robot and contact evidence. Detected failures stop the active skill and provide grounded feedback to a Recovery Agent. The framework further includes a Memory Module for reusing prior skill experience, and geometric grounding via depth and segmentation. We independently evaluate the monitoring module of the framework in MuJoCo using 100 trials comprising 50 failed and 50 successful executions across five tasks. The monitor detects 48 of 50 failures, correctly accepts 46 of 50 successful executions, and achieves 94.0% overall accuracy. These results provide initial evidence for the monitoring component, while end-to-end evaluation of the complete recovery loop remains ongoing.
Vision-Language-Action (VLA) models commonly predict action chunks, limiting their ability to react to environmental changes during execution. Existing asynchronous inference methods improve reactivity but typically rely on a fixed inference gap. In this paper, we propose an event-guided dynamic inference strategy that adapts the inference gap according to scene changes observed since the previous inference. Thereby, it simultaneously preserves motion consistency and prompt reactivity. Across static and dynamic real-world settings, our method consistently performs best, averaging 95% success and exceeding the strongest baseline by 55 percentage points. The code will be made publicly available upon acceptance. The project page is available at https://react-when-you-need-to.github.io/.
REBOOT: From Failure to Recovery - A Dataset and Benchmark for Precision Assembly
REBOOT:从失败到恢复——面向精密装配的数据集与基准
Ofori-Ampofo, Nana Yaw Owusu, Kahou, Samira Ebrahimi, Thekinen, Joseph
Abstract
Robot learning policies fail in characteristic ways: they stall in uncertain states, drift during contact-rich alignment, and miss targets by millimetres in precision tasks. Yet training datasets consist largely of successful demonstrations, while real-world benchmarks often reduce performance to binary success. This limits both supervision for recovery and analysis of where failures occur. We introduce REBOOT (Recovery Episode Benchmark for Off-nominal Trajectories), the first robot manipulation benchmark designed around failure as a first-class signal. REBOOT contains 2,160 demonstrations across 18 precision assembly tasks, each decomposed into five shared phases: Align(pick), Engage(pick), Transport, Align(place), and Engage(place), enabling phase-level evaluation beyond terminal success. Failures are introduced across phases and paired with expert recovery trajectories that return the system to a valid continuation state. Tasks are annotated with rotational symmetry, engagement-clearance precision tier, and assembly direction through matched install-remove pairs. Failure episodes are labeled by phase and categorical failure mode, enabling attribution to kinematic stage and tolerance violation. Data includes synchronized RGB-D observations from four viewpoints and grounded natural-language descriptions of phase-level success and failure conditions. Half the dataset contains expert demonstrations; the other half contains recovery demonstrations sampled to reflect failures observed in imitation-learned policy rollouts. We benchmark action-chunked transformer, diffusion, and $\pi_0$-FAST policies using phase-level completion rates, revealing model-specific failure points hidden by binary evaluation. Dataset and code: https://nanayawoa.github.io/REBOOT
Chinese Translation
机器人学习策略会以特定方式失败:它们在不确定状态下停滞,在接触密集的对齐过程中漂移,并在精密任务中偏离目标数毫米。然而,训练数据集主要由成功的演示构成,而现实世界的基准测试往往将性能简化为二元成功与否。这既限制了对恢复行为的监督,也限制了对失败发生位置的分析。我们提出了REBOOT(非正常轨迹恢复片段基准,Recovery Episode Benchmark for Off-nominal Trajectories),这是首个将失败作为一等信号来设计的机器人操作基准。REBOOT包含18个精密装配任务共2,160条演示,每个任务被分解为五个共享阶段:Align(pick)、Engage(pick)、Transport、Align(place)和Engage(place),从而实现超越终端成功的阶段级评估。失败被引入到各个阶段,并配有使系统返回到有效续行状态的专家恢复轨迹。任务通过匹配的安装-拆卸对,标注了旋转对称性、配合间隙精度等级和装配方向。失败片段按阶段和分类失败模式进行标注,从而可将失败归因于运动学阶段和公差违规。数据包括来自四个视角的同步RGB-D观测,以及关于阶段级成功与失败条件的接地自然语言描述。数据集中一半为专家演示;另一半为恢复演示,其采样方式反映了模仿学习策略 rollout 中观察到的失败。我们使用阶段级完成率对动作分块的Transformer、扩散模型以及π0-FAST策略进行基准测试,揭示了被二元评估所掩盖的模型特定失败点。数据集与代码:https://nanayawoa.github.io/REBOOT
Characterizing Wildlife Response to Biomimetic and Conventional Underwater Vehicles
野生生物对仿生与传统水下航行器响应的特征研究
Pham, Huy, Cai, Levi, Girdhar, Yogesh, Rus, Daniela, Patterson, Zach J.
Abstract
Autonomous underwater vehicles (AUVs) are a promising alternative to divers for scalable collection of natural ocean ecology data. However, these robots may disturb local fauna and cause drastic behavioral differences compared to other monitoring techniques, decreasing their value as scientific tools. A promising prospect is to make AUVs that are more biomimetic, with the hope that taking on the form and behavior of a non-predatory animal may reduce adverse responses. We present the first dataset comparing fish disturbance in response to a conventional thruster-driven AUV, a sea turtle inspired flipper-driven AUV, and a diver. Experiments were conducted at a biodiversity hotspot in a Caribbean reef, and images of the scene were analyzed using computer vision to localize and study fish behavior change. Both AUVs cause measurable changes in fish behavior. Although no between-robot differences remain significant after correction for multiple comparisons, point estimates generally favor the biomimetic AUV, motivating larger studies capable of resolving modest effects and further design changes to optimize for disturbance. We also observe larger responses during diver transects than during AUV transects, although this exploratory comparison is based on a small diver sample with several confounds. While conclusions must be taken as preliminary due to operational and experimental limitations, this study provides the community with a first known dataset and benchmarks to quantify the behavioral impacts of biomimetic robots for ecological monitoring in the wild.
Vision-Tactile-Language-Action (VTLA) models have demonstrated clear advantages over Vision-Language-Action (VLA) models in contact-rich manipulation. However, developing VTLA models is severely constrained by the massive amounts of vision-tactile data and computational resources required. To address this bottleneck, we propose VT-Bridge, a lightweight residual adaptation strategy that bridges pretrained foundation VLAs to VTLAs. Rather than training a VTLA model from scratch or modifying the original architecture of a pretrained VLA, VT-Bridge employs an identical lightweight residual-adapter architecture across VLA backbones and uses backbone-specific weights to refine actions at the robot execution frequency. This design substantially lowers the data and training barriers. Specifically, it requires up to 50 vision-tactile demonstrations per task to fine-tune a VLA backbone and train a 0.98M-parameter residual adapter. Experiments with three representative VLA backbones ($\pi_0$, $\pi_{0.5}$, and SmolVLA) across four contact-rich manipulation tasks further demonstrate its consistent effectiveness across VLA architectures. On average, VT-Bridge raises the task completion rate from 11.7% with task-level VLA fine-tuning alone to 62.9%. Together, these findings demonstrate the broad applicability, effectiveness, and accessibility of VT-Bridge for contact-rich manipulation. The project page is available at https://hoxnocha.github.io/vt-bridge-web/.
Projectile interception is a challenging dynamic manipulation problem. Intercepting a thrown object with a robot arm requires reaching a point on the object's path as the object passes through it. Slicing also fixes the blade's velocity and orientation at contact. The goal is therefore a subset of the states of the robot and arrival times that moves as the object falls, and the arm must reach it within its actuator limits in milliseconds. We present FRUITNINJA, an anytime sampling-based planner that grows a tree on the GPU in batches toward the interception manifold. Each edge is an exact cubic whose travel time is found by a parallel search against the arm's dynamics, so every edge satisfies the actuator limits. Plans are ranked by a risk-aware objective over the uncertainty in the object's position and the arm's arrival time. We evaluate on a Franka Research 3 against six baselines in a calibrated real-time simulator, where FRUITNINJA cuts 96.7% of tosses in the open and 68.3% among five obstacles, versus the best baseline's 68.3% and 35.0% respectively.
Chinese Translation
抛射物拦截是一个具有挑战性的动态操作问题。用机械臂拦截抛掷物需要在该物体经过其路径上的某一点时抵达该点。切割还要求刀刃在接触时刻具有特定的速度和姿态。因此,目标是机器人状态和到达时间的一个子集,且该子集随物体下落而移动,机械臂必须在毫秒级时间内、在其执行器限制范围内抵达目标。我们提出了FRUITNINJA,这是一种任意时间(anytime)的基于采样的规划器,它在GPU上以批次方式朝拦截流形方向生长一棵树。每条边是一条精确的三次曲线,其行进时间通过与机械臂动力学进行并行搜索来确定,因此每条边都满足执行器限制。规划方案根据一个风险感知目标进行排序,该目标考虑了物体位置和机械臂到达时间的不确定性。我们在一个经过校准的实时仿真器中,使用Franka Research 3机械臂与六个基线方法进行了对比评估。结果显示,FRUITNINJA在开阔环境中成功切割了96.7%的抛掷物,在五个障碍物环境中成功切割了68.3%,而最佳基线方法分别仅达到68.3%和35.0%。
From Documented Strengths to Force Limits: Material-Informed Robotic Insertion for Construction Assembly
从文献记载的强度到力限制:面向建筑装配的材料信息驱动的机器人插装
He, Lin, Chen, Yanyi, Sun, Haofei, Li, Lingyao, Deng, Min
Abstract
Insertion is a fundamental operation in robotic construction assembly, where variations in material properties and assembly conditions make it difficult to select contact forces that complete the task without exceeding the assembly's capacity. Although construction documents encode engineering knowledge about materials and their conditions, translating this knowledge into load limits for a specific assembly remains difficult. This paper presents SAGE (Source-grounded Assembly Gating and Execution), a system that converts documented material evidence into capacity estimates for robotic insertion. SAGE restricts a large language model (LLM) to extracting tensile and compressive strengths from retrieved passages and tables and records their sources. A response model then interpolates offline finite element (FE) solutions to convert these strengths and the assembly conditions into axial load capacity. For fits with positive clearance, the estimated capacity sets the policy's axial force limit; for interference fits, it is compared with measured support demand to determine admission. On the primary benchmark, SAGE reduces mean capacity error from 80.65\% for direct LLM estimates based on the same evidence to 10.74\%. Without refitting, the mean error remains 8.00\% on 16 additional geometries. Under the assigned support release model, SAGE correctly classifies 59 of 62 scored simulation runs, with only conservative errors. In recorded xArm6 demonstrations, SAGE takes material documents as input and completes physical insertion in 9 of 13 trials. These results show that assigning document interpretation to the LLM and force calculation to an explicit mechanical model produces accurate capacity estimates and traceable insertion decisions.
Chinese Translation
插装是机器人建筑装配中的一项基础操作,材料属性和装配条件的变化使得选择既能在不超过装配承载能力的情况下完成任务的接触力变得困难。尽管建筑文件中编码了关于材料及其状态的工程知识,但将这些知识转化为特定装配的载荷限制仍然困难重重。本文提出了SAGE(基于文献的装配门控与执行系统,Source-grounded Assembly Gating and Execution),该系统将文献记载的材料证据转换为用于机器人插装的承载能力估计。SAGE限制大语言模型(LLM)仅从检索到的段落和表格中提取抗拉强度和抗压强度,并记录其来源。随后,响应模型对离线有限元(FE)解进行插值,将这些强度和装配条件转换为轴向载荷承载能力。对于具有正间隙的配合,估计的承载能力设定策略的轴向力限制;对于过盈配合,则将其与测得的支撑需求进行比较以确定是否准许执行。在主要基准测试中,SAGE将平均承载能力误差从基于相同证据的直接LLM估计的80.65%降至10.74%。在无需重新拟合的情况下,在16个额外几何结构上的平均误差仍保持为8.00%。在指定的支撑释放模型下,SAGE在62次评分仿真运行中正确分类了59次,且仅有保守性误差。在xArm6的实录演示中,SAGE以材料文档作为输入,在13次试验中有9次完成了物理插装。这些结果表明,将文档解读任务分配给LLM、将力计算任务分配给显式力学模型,能够产生准确的承载能力估计和可追溯的插装决策。
Humanoid robots can acquire complex skills by imitating kinematic humanoid motion references, yet reliable references for contact-rich interactions remain difficult to obtain: motion capture deteriorates under occlusion and close physical contact, while retargeting introduces additional contact and geometric inconsistencies. We present HIGenNTO, a framework that synthesizes humanoid-scene interaction motion references by optimizing the initial noise of a pretrained text-conditioned motion model under sparse spatiotemporal and scene constraints. The same formulation satisfies desired contacts, avoids collisions, and maintains stable support while retaining the prior's realism and temporal coherence, generating interaction motions from scratch and composing long-horizon behaviors stage-wise. Across robot-environment and robot-object tasks, HIGenNTO produces motions that can be executed by tracking policies in simulation and used to train depth-conditioned visuomotor policies operating solely from onboard sensing. We deploy these policies on a Unitree G1 across four contact-rich tasks. Finally, the task specifications themselves can be written by a coding agent, which proposes interaction tasks and compiles them into prompt, constraint, and scene programs, authoring three of our eight evaluated tasks and four further behaviors. Together, these results establish a scalable path from high-level task descriptions to physically executable humanoid interactions.
Simultaneous Forward and Inverse Human-in-the-Loop Optimization
同步正向与逆向人在回路优化
Park, Kyeongwon, Collins, Steven H.
Abstract
Subjective user experience is important to human-robot interaction, but the outcomes users value, and how those preferences vary across individuals and contexts, are often unknown. While inverse learning approaches using human data can help identify user rewards, in many assistive settings the experimental costs of executing a control policy, measuring biomechanical or physiological outcomes, and collecting user feedback often limit the number of queries and optimization iterations. Here, we present Simultaneous Forward and Inverse Human-In-the-Loop Optimization (SFIHILO), which efficiently infers individual-specific reward functions from preferences over human outcomes and identifies a final control policy that maximizes the learned reward. SFIHILO bootstraps a forward model to predict user outcomes from control policies, uses this model for active querying to accelerate inverse reward learning, and then optimizes the final control policy without additional user trials. We score candidate policies by their expected reduction in uncertainty across both forward and inverse beliefs, targeting a regional preference boundary to robustly inform this simultaneous learning process. In simulation, we show that SFIHILO was effective across user heterogeneity, outcome dimensionalities, precision requirements, noise levels, and nonstationarity; compared with mutual information approaches, the proposed active querying strategy significantly improved sample efficiency in inverse learning while preserving forward model accuracy. This approach demonstrates the potential to infer latent human goals, enabling more transferable and effective human-robot interaction.
Density-Driven Area Coverage for Nonholonomic Multi-Robot Systems with Safety Guarantee
具有安全保障的非完整多机器人系统密度驱动区域覆盖
Martinez, Julian, Lee, Kooktae
Abstract
Density-Driven Optimal Control (D2OC) provides a principled approach to distributing multi-robot teams over non-uniform spatial distributions. Applying D2OC to nonholonomic robots, however, creates a gap between safety constraints imposed on a reference motion and the physical inputs that determine the actual robot motion. We address this issue by enforcing the safety constraint directly on the robot's physical inputs while preserving the density-driven coverage objective. The proposed framework combines D2OC with a control barrier function safety filter through a feedback-linearizing look-ahead point, allowing safety and actuator limits to be considered together during control. We further derive a safety margin that accounts for the look-ahead geometry, robot footprint, and motion during each control interval. Simulation results show that the proposed method maintains the required physical separation while achieving coverage performance comparable to a conventional reference-tracking approach, which can satisfy safety on the reference motion yet violate the corresponding physical clearance. Experiments on multiple nonholonomic robots in the Robotarium further demonstrate safe execution while driving the robots toward the desired spatial distribution. These results show that enforcing safety directly on the physical inputs can eliminate the mismatch between safety certification and physical robot motion in density-driven multi-robot coverage.
Underwater robot simulation requires diverse environments in which terrain, scene composition, tasks, and currents remain mutually consistent. We present AquaWorld, a world-generation framework that preserves these relationships through shared terrain structure. A language-conditioned plan generates 3D terrain and shared structural references that guide asset placement, task definition, and inflow specification under stochastic variation. The framework incorporates over 10,000 underwater-compatible assets, predicts reusable terrain-conditioned mean-flow fields through a CFD-supervised residual model, and supports conventional underwater vehicles and bio-inspired robotic fish. On 24 paired terrains, structure-consistent randomization produces substantially better cross-factor consistency than independent randomization. In a matched-budget policy-training comparison, structurally coherent randomization achieves a validation success rate 21% higher than independent randomization. In separate physical experiments, a simulation-trained visual navigation policy succeeds in 95% of physical tank trials without updating its perception or control modules. Overall, AquaWorld provides a practical way to generate varied underwater environments while retaining the structural relationships needed for flow simulation and robot learning.
Lightweight and slim manipulators enable safe operation in human living environments. Proximal actuation using remote transmission mechanisms, such as wire-driven or Bowden cables, effectively reduces inertia and arm size by relocating motors near the base and transmitting torque to distal joints. Existing approaches either increase mass through additional components, such as pulleys for direction changes, or suffer from reduced transmission efficiency due to friction losses. Flexible shaft transmission avoids both mass increase and excessive friction losses, but faces increasing angular transmission error due to helical buckling caused by slack generated at joint bending. To address this problem, we propose SHAFT: a Slack-compensating, Helical-buckling-Attenuating Flexible- shaft Transmission mechanism. This mechanism compensates for slack through a proximal tensioner, improving the angular transmission error and efficiency of flexible shaft transmission without increasing the moving mass of the arm section. In a transmission path containing four 90-degree bends, the proposed mechanism demonstrated approximately 30% higher efficiency and approximately 65% lower angular transmission error compared to a flexible shaft transmission without a tensioner. Using this mechanism, we fabricated a 6-Degree-of-Freedom (DoF) arm with a 1-DoF gripper manipulator consisting of a rotary module housing motors with a tensioner, and a 280 g weight for the arm module. The proposed manipulator represents a novel remote actuation system for achieving lightweight construction with high efficiency, contributing to the acceleration of safe robot deployment in human environments.
A Direct Rigid Transmission 2-DoF Wrist Extension for Tendon-Driven Hand
用于肌腱驱动手的直接刚性传动二自由度腕部扩展装置
Pang, Yujie, Sakib, Sadman, Faruque, Mohammad Abdullah Al
Abstract
Dexterous manipulation in confined spaces requires local control of hand orientation. Without a wrist, a dexterous hand must obtain this local orientation through coordinated motion of the robot arm, often involving several joints and a more complex end-effector path. We present CRAFT-Wrist, a concentric 2-DoF wrist extension that mounts between a robot arm and the CRAFT Hand without modifying the hand. Two XC430-T240BB-T servos drive sideways and front-back rotation through short rigid transmissions. Our initial hardware prototype actuates both axes under load, achieving a demonstrated workspace of $\pm20^{\circ}$ in radial--ulnar deviation (left--right) and a front-back range of $+80^{\circ}$ in flexion and $-18^{\circ}$ in extension. We characterize the resulting finger-motor loading across several wrist postures and demonstrate how the wrist supplies local orientation in grasping, nail hammering, and blackboard wiping. Demonstrations and assembly instructions are available at https://craft-wrist.github.io/.
StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies
StateMem:面向视觉-语言-动作策略的单状态残差记忆与自适应推理方法
Li, Wenzhuo, Shi, Qiongfeng, Zhou, Yi
Abstract
Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA) policies primarily rely on current observations, limiting historical information retention. Memory-augmented VLAs, such as MemoryVLA, address this limitation with external memory banks but require explicit storage and retrieval. To address these limitations, we propose StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes. A training-free controller adjusts the routing threshold online, while fast correction compensates for stale prefix features during cache reuse. We evaluate StateMem on LIBERO, RoboMemArena, and real-world manipulation tasks. On LIBERO, StateMem achieves an average success rate of 97.6% and reduces the average VLM prefix refresh rate by 20.25% relative to full refresh. In the Occlusion category of RoboMemArena, StateMem achieves the best performance among single-VLA methods, reaching 21.8% Task Success Rate (TSR) and 44.3% Cumulative Success Rate (CSR). Across six real-world manipulation tasks, it achieves +21% in average success rate.
Efficient coordination under limited communication remains a key challenge in decentralized multi-robot exploration. While centralized approaches benefit from global information sharing, they are often impractical in large-scale or communication-constrained environments. Existing Monte Carlo Tree Search (MCTS)-based approaches, such as Decentralized Monte Carlo Exploration (DMCE), enable decentralized planning by taking peer intent into account. This peer intent is obtained by communicating sequences of planned waypoints with robots within direct communication range. In this work, we extend this idea by introducing Probabilistic Peer Intent (PPI), which converts peer trajectories into a continuous spatial representation of predicted intent and incorporates it into local MCTS action evaluation. We additionally study the effects of sharing peer intent beyond direct communication range by propagating plans over multiple hops. Experiments across multiple simulated environments and team sizes show that PPI and Multi-hop propagation can each improve decentralized exploration, with their relative benefits depending on environment structure and team size. We also demonstrate the real-world deployment of our method on three robots operating in different environment types.
Chinese Translation
在有限通信条件下实现高效协调仍是分散式多机器人探索中的关键挑战。集中式方法虽然受益于全局信息共享,但在大规模或通信受限环境中往往不切实际。现有的基于蒙特卡洛树搜索(MCTS)的方法,如分散式蒙特卡洛探索(Decentralized Monte Carlo Exploration, DMCE),通过考虑同伴意图实现了分散式规划。这种同伴意图通过与直接通信范围内的机器人交换规划路径点序列来获取。在本工作中,我们扩展了这一思想,提出了概率化同伴意图(Probabilistic Peer Intent, PPI),它将同伴轨迹转换为预测意图的连续空间表示,并将其纳入局部MCTS动作评估中。此外,我们研究了通过多跳传播计划,将同伴意图共享扩展到直接通信范围之外的效果。在多种仿真环境和不同团队规模下的实验表明,PPI和多跳传播各自都能提升分散式探索性能,其相对优势取决于环境结构和团队规模。我们还在三种不同环境类型中部署了三个机器人,验证了该方法的真实世界应用效果。
BEACON: Belief-Enabled Adaptive CONtrol for Imitation Learning under Uncertainty
BEACON:面向不确定性下模仿学习的基于信念的自适应控制
Lee, Moonyoung, Bhattacharya, Soumojit, Kantor, George, Kroemer, Oliver
Abstract
Robot manipulation tasks often involve hidden state information that cannot be directly observed and must be inferred through sequential physical interactions. In such partially observable settings, conditioning an imitation learning policy directly on the recent raw observation history leads to poor performance. This is due to state aliasing, wherein identical observations may arise from different hidden states, and the policy receives conflicting action labels for the same input. To enable history-aware disambiguation capability, we propose conditioning a diffusion policy on a structured representation of the hidden states using Bayesian belief that exposes both the current most likely state estimate and the remaining uncertainty. This representation replaces raw history with a structured, compact input, enabling the policy to implicitly modulate between exploratory and exploitative behaviors based on belief uncertainty, without explicit mode switching or reward shaping. We evaluate across two domains with qualitatively different belief representations: a continuous belief for cornstalk gripper alignment via tactile sensing, and a discrete categorical distribution for latched door opening. In both domains, the belief-conditioned policy substantially outperforms the observation-only baseline and approaches privileged ground-truth performance, with ablations illustrating that the policy adapts its exploration behavior depending on the belief uncertainty at inference time.
This paper develops the Rapidly-iterating kinoDynamic Grid (RDG) algorithm, an asymptotically near-optimal kinodynamic motion planning algorithm that produces high quality solutions through rapid iteration. The algorithm leverages a state space grid decomposition to perform node selection, dynamics propagation, and graph revision in constant time complexity with respect to the number of nodes in the trajectory tree. Through a covering ball sequence induction proof, the algorithm is shown to be asymptotically near-optimal and probabilistically complete. Different subsystems of the algorithm are evaluated against common nearest-neighbor search-based methods at generating exploration bias. The RDG algorithm is evaluated through simulated trials in complex, kinodynamic motion planning problem environments up to 10 DOF relative to similar sparse, kinodynamic planning algorithms with optimality guarantees, SST and DIRT. The RDG algorithm outperforms both SST and DIRT in mean final solution quality by up to 104% and 40% respectively. Additionally, RDG maintained a 100% success rate, even on a 10-DOF test case where both SST and DIRT did not.
There are renewed efforts to build a lunar base and house a team of astronauts on the Moon for extended periods. However, important challenges remain, such as a rapid-response fire-hazard management system. In the event of a fire or toxic leak inside an early crewed lunar habitat, a single pressurized cylinder must be handled within minutes. However, suited astronauts and radio-delayed support from Earth cannot guarantee such response times. We describe a prototype fire-fighting and emergency response safety technology that embeds intelligence in the habitat itself: using thin ceramic QR codes affixed every few meters, each carries a local floor map, safe sensor limits, and step-by-step hazard responses. A low-power rover decodes a plate in a matter of seconds, polls the adjoining temperature-gas sensor cluster, and acts immediately, eliminating the need for a global map . In a representative layout, the interior is divided into roughly one code per two square meters. Simulations show that rapid-response firefighting both simplifies the task and avoids costly resource use and indeterminate outcomes. Because knowledge is spread across passive plates, damage is localized, procedures can be updated by replacing a single code, and processor demands remain minimal. However, important challenges remain, such as a rapid-response fire-hazard management system. In the event of a fire or toxic leak inside an early crewed lunar habitat, a single pressurized cylinder must be handled within minutes. However, suited astronauts and radio-delayed support from Earth cannot guarantee such response times. We describe a prototype fire-fighting and emergency response safety technology that embeds intelligence in the habitat itself: using thin ceramic QR codes affixed every few meters, each carries a local floor map, safe sensor limits, and step-by-step hazard responses.
Task-Oriented Co-Design and Optimization of Geared Actuators for Robotic Applications
面向任务的机器人齿轮执行器协同设计与优化
Huang, Xuanyu, Dong, Jianqiang, Zhao, Hang
Abstract
Different tasks performed by legged robots impose distinct torque and speed requirements on actuators. Existing robotic actuators are generally optimized at the component level for metrics such as torque or power density, without explicit task guidance. System-level optimization across components such as motors, gearboxes, and sensors is challenging because of the high computational cost and coupling among mechanical, electrical, and electromagnetic behaviors. Consequently, improvements in individual components may not translate into better robot performance in a specific task. To this end, we present a systematic optimization framework for task-oriented co-design of actuator hardware and control. First, surrogate models are employed to accelerate motor evaluation and support global exploration of the coupled design space. Then, a hierarchical mixed-variable optimization strategy is adopted, combining discrete enumeration with continuous search over dimensions and real-valued indices. These indices are rounded to select admissible values for the remaining discrete choices before each evaluation. Within this search, rated output torque density and task performance are jointly optimized, with Bezier-parameterized joint torque profiles determined for each hardware candidate. Finally, the effectiveness of the proposed framework is validated through actuator fabrication and experiments on a two-degree-of-freedom jumping leg. Based on its measured mass, the fabricated prototype achieves a nominal rated output torque density of 35.7 N m/kg, approximately 60% higher than that of a widely used commercial geared joint actuator, while being 18.6% lighter. Under matched bench conditions, it achieves 12.0% greater jump height at twice-rated torque. Together, these results demonstrate a systematic route from task requirements to actuator design and control.
Expert demonstrations often specify what a robot should do, but not how fast it can do it. Imitation Learning (IL) inherits demonstration timing, while directly accelerating the learned motion can fail when faster execution changes the robot-object dynamics. We study faster-than-demonstration execution as a dynamics-aware control problem and introduce Expert-Play Contouring Control (EPCC), which combines slow expert demonstrations with fast, non-expert play. Expert demonstrations train a latent trajectory generator whose predictions are reparameterized into a time-independent contour of successful task progression, while play trains a world model (WM) of fast-action outcomes. At deployment, our proposed planner uses the WM to optimize actions that makes maximize progress along the expert-derived contour while penalizing deviation from the intended task evolution. Averaged across three visuomotor manipulation tasks, EPCC achieves a $2.0\times$ the throughput of the IL baseline, including $2.2\times$ that of the throughput of the strongest acceleration baseline on a task with interaction-sensitive object dynamics. Our analysis shows that the gains concentrate where faster execution changes robot-object evolution. Together, our results highlight a simple yet effective principle: demonstrations provide task intent, while play data provides the dynamic coverage needed to execute that intent faster.
ARCGym: Benchmarking Deep Reinforcement Learning in Autonomous Robotic Colonoscopy
ARCGym:自主机器人结肠镜检查中深度强化学习的基准测试
Ji, Guanglin, Finocchiaro, Martina, Erleben, Kenny, Yin, Hang
Abstract
Simulations for learning-based autonomous colonoscopic navigation focus mainly on fully actuated capsule robots, failing to capture the contact-rich navigation of long and flexible clinical colonoscopes. We present the Autonomous Robotic Colonoscopy Gym (ARCGym), an open-source reinforcement learning environment and benchmark for image-based navigation in clinically derived deformable colon anatomies. ARCGym supports multiple types of colonoscope robots, spanning capsule robots and flexible endoscopes, with this work focusing on flexible endoscopes including magnetic-driven tip actuation and clinically used proximally translational actuation. This work includes five CT-reconstructed colons representing typical clinical scenarios, a set of clinically meaningful navigation subtasks, and unified success metrics. We introduce a reward combining depth-based lumen alignment with a lumen-visibility score to improve learning under occlusions. Experiments across tasks, robots, and anatomies show that autonomous navigation remains challenging for both magnetic-driven and proximal-insertion flexible robots, with proximal-insertion actuation remaining an open problem.
A wrist-mounted camera for UMI-style data collection must do two jobs: record the manipulation and localize in the scene. Most handheld devices localize online from workspace-facing views crowded by hands and objects, or add dedicated tracking hardware. Room-scale bimanual capture therefore still tends to instrument the operator or the scene for accurate localization. We present KIWI (Kinematic Interface for the Wild), a capture kit whose only electronics are off-the-shelf cameras. Our core system splits the two jobs across the two lenses of a 360-degree camera. The rear lens faces the room and builds a shared metric map that registers both hands, and an optional head camera, in one frame without workspace co-visibility; the front lens records the manipulation, and offline IMU fusion bridges front-lens tracking loss. Through our quick-release plate, the camera module attaches to chopstick grippers, parallel-jaw grippers, hand-wrist mounts, or robot flanges. Across six bimanual recordings, combining the rear and front lenses failed to localize only 0.1% of query frames, whereas front-only bimanual feature alignment failed on 24.8% of frames and lost one recording entirely; against evaluation fiducials, localization error stayed within 4.5 mm. KIWI's recovered poses were sufficiently consistent for the four wrist streams alone to reconstruct the scene as a 3D Gaussian splat.
Chinese Translation
用于UMI风格数据采集的腕部安装相机必须同时完成两项任务:记录操作过程并在场景中定位。大多数手持设备依赖面向工作空间的视野进行在线定位,但该视野常被手和物体遮挡,或需要额外的专用跟踪硬件。因此,房间尺度的双臂数据采集通常仍需对操作者或场景进行标记以实现精确定位。我们提出了KIWI(Kinematic Interface for the Wild),一套唯一电子设备仅为现成相机的采集套件。我们的核心系统将两项任务分配给360度相机的两个镜头:后置镜头朝向房间,构建一个共享的度量地图,将双手以及可选的头戴相机统一配准到同一坐标系中,无需工作空间共视;前置镜头记录操作过程,并通过离线IMU融合弥补前置镜头的跟踪丢失。通过快装板,相机模块可安装在筷式夹爪、平行夹爪、手腕支架或机器人法兰上。在六段双臂数据采集中,结合后置与前置镜头的方案仅有0.1%的查询帧定位失败,而仅使用前置镜头的双臂特征对齐方案在24.8%的帧上失败,且完全丢失了一段采集数据;相对于评估基准标记,定位误差保持在4.5毫米以内。KIWI恢复的位姿具有足够的一致性,仅凭四路腕部数据流即可将场景重建为3D高斯泼溅(3D Gaussian Splatting)。
Commonsense-Grounded Path Planning from Abstract Instructions
基于常识的抽象指令路径规划
Endo, Masafumi, Honda, Kohei, Yonetani, Ryo
Abstract
We present \emph{commonsense ranked search} (CoRS), a novel path planner that turns an abstract instruction into a route that follows commonsense. While existing methods respect the considerations written down in advance, a robot working among people must follow those left unstated too, as with a wet floor that a worker avoids without being told. CoRS leverages large language models (LLMs) and vision-language models (VLMs) as commonsense knowledge to reason about these latent considerations in its planning. Given an abstract instruction (\emph{e.g.}, ``move carefully'') and visual observations of each region in the environment, CoRS derives a consideration for each region, as in ``this wet floor is slippery and worth a detour.'' It then compares the considerations between regions to see which of the two the robot should avoid more, as in ``the crowd is worse than the wet floor.'' These judgments sort the regions into a commonsense ranking, whose costs drive a conventional search that always returns a valid route. We build a benchmark for planning under latent considerations, with three environments, 1350 problems, and five instructions at three levels of abstraction. Experiments show that CoRS discovers the unstated considerations and goes around the ones worth a detour while crossing the rest, a behavior that recent LLM-based planners do not achieve.
Collecting whole-body demonstrations for humanoid manipulation mostly relies on teleoperation, which is costly and hard to scale up. The Universal Manipulation Interface (UMI) provides a scalable data collection paradigm, but end-effector trajectories alone underdetermine humanoid whole-body coordination, which is insufficient for whole-body demonstration collection. Therefore, we introduce Whole-Body UMI (WB-UMI), a task-agnostic, real-time and end-effector conditioned motion generator that decouples whole-body coordination learning from task semantics learning through a shared end-effector interface. A diffusion policy learns from native UMI demonstrations, while WB-UMI learns independently from retargeted motion capture, requiring no body trackers or paired image--whole-body demonstrations during task-specific data collection. In real deployment, an asynchronous hierarchy integrates the diffusion policy, motion generator, and a whole-body controller with latency compensation and measured-state feedback. Real-robot experiments on G1 support real-time closed-loop transfer across four tasks, achieving 90% success in drawer closing, 80% in shelf pick-and-place, 30% in ball toss, and 40% in Loco-PnP, which shows the effectiveness of this hierarchy in transferring native UMI skills to humanoid whole-body manipulation.
Force-conditioned vision-language-action (VLA) policies can respond to contact, but when trained solely on demonstrations, their recovery behavior may be limited by demonstration coverage, and they do not learn from deployment outcomes. Human corrective imitation provides additional recovery examples, but its objective matches local action targets without explicitly optimizing task return. We present ForceRFT, a force-guided residual reinforcement learning framework that learns contact-dependent corrections from human supervision and autonomous task outcomes. A frozen, demonstration-trained SmolVLA-based prior generates force-conditioned action chunks, while a lightweight residual actor refines individual end-effector pose commands using wrist feedback acquired during chunk execution. The decision-time wrist wrench, its temporal change, and the selected base motion condition both residual correction and value estimation. Human corrections supervise the residual actor, while verified autonomous transitions train the twin critics and support value-guided updates to the same actor. Bootstrapping is restricted to autonomous segments, preventing TD credit from crossing human-intervention boundaries. Real-robot experiments on plug insertion, ring-on-peg assembly, and whiteboard wiping show higher autonomous success rates than the evaluated demonstration-trained and residual-imitation baselines. Comparisons with residual imitation support value-guided residual optimization, while plug-insertion ablations indicate the benefit of direct execution-time wrist feedback.
Physical-Touch Observability from Wrist Wrench in Granular Scooping
从腕部力觉观测颗粒物料铲掘中的物理接触可观测性
Lin, Hongyi, Zhang, Song, Liu, Xubo, Liu, Yang
Abstract
Mining and earthmoving are important real-world deployment settings for embodied intelligence. Autonomous transport and driving systems have improved substantially, but loading and scooping still often depend on skilled human operators, exposing personnel and equipment to operational risk. For robotic scooping, pre-contact RGB-D sensing reveals surface geometry but not the resistance, compaction, tool engagement, or load transfer that emerge during interaction. We test whether the current scoop's six-axis wrist force/torque (F/T), or wrist wrench, contains information about final collected volume, and whether that information depends on the correctly paired action-terrain interaction. We call this property physical-touch observability. Using 6,700 real-robot scoops across 67 terrains, we evaluate correctly paired current-scoop F/T against pre-contact prediction and correspondence-breaking controls under terrain-held-out testing. At the retrospective 60% sequence boundary, correctly paired F/T reduces mean absolute error by 14.2% relative to Action-only and by 21.9% relative to cross-terrain mismatched F/T. Engineered signal summaries reproduce the result across model architectures. Together, these findings position wrist wrench not merely as a low-level feedback signal, but as a task-level perceptual modality through which embodied robots can infer hidden physical states during interaction, providing a foundation for response-aware autonomy in mining and other contact-rich tasks.
Vision-language-action models often predict actions from only the current observation, which can leave tasks involving object occlusion or visually identical objects ambiguous without episode history. The usual countermeasure, widening the observation window, turns the horizon into a hyperparameter and lets per-step cost grow with it. We instead capture the episode in a recurrent state. SmoLSTM couples a frozen 256M-parameter SmolVLM backbone to a matrix-memory LSTM control layer in which observation tokens and action queries are unified in a single causal stream that is never reset throughout the entire episode. Recurrent-state storage is therefore O(1) in episode length. A flow-matching action head predicts chunks of 10 end-effector pose deltas and gripper commands at each control step. Our single policy, trained jointly on 7,461 demonstrations across 140 tasks and evaluated on held-out initial states, performs best, reaching 85.1% subgoal coverage and 77.5% full-task success on LIBERO-Mem with 0.04B trainable parameters, surpassing both the benchmark's own object-centric baseline and recent memory-based approaches. Resetting the recurrent state at every control step reduces full-task success to 7.0%, showing that the trained policy relies on context carried between decisions. The same model achieves 79.6% average success on standard LIBERO.
World models allow robots to anticipate action consequences before execution. This capability is especially valuable in excavation, where each scoop reshapes the terrain and affects subsequent actions. Local observations, however, cannot fully reveal the underlying support and material conditions. We present PileBelief, an interaction-driven persistent world model for partially observed excavation that retains physical evidence beyond the visible surface. It combines an observation-conditioned physical prior with world-addressed deformation memory and physical-response memory. Action-aligned reads and gated residual corrections refine terrain-change and outcome predictions. With deployment weights fixed, completed interactions update measured belief, while hypothetical actions advance a separate imagined state. Compared with a current-observation-only baseline, PileBelief reduces five-step joint prediction error by 10.8% and offline action-selection regret by 65.5%. Experiments on Newton/MPM and real excavation datasets further demonstrate improved terrain-change and bucket-volume prediction. Our method enables multi-step prediction and candidate-action ranking from local observations, even when the underlying soil state is unknown. These results identify persistent physical belief as a useful representation for world models of environments that robots continually reshape.
In-hand assembly is constrained by the need to maintain grasps on two separate parts while controlling their relative motion within a single hand. To enable both in-hand assembly and manipulation, we present a reconfigurable dual-opposition architecture. Specifically, to support simultaneous grasping of two parts and coordinated in-hand manipulation, four independently actuated fingers are organized into two virtual finger (VF) oppositions, with their relative configuration controlled by a reconfigurable palm. To describe hand motion and simultaneous two-object grasping configurations, a kinematic model of the fingers and palm and an object-size-conditioned workspace formulation are built. To further evaluate motion performance and assembly capability, finger-joint motion and palm tracking are characterized, and in-hand assembly is demonstrated through tasks involving grasping, alignment, fastening, and pressing. Ablation experiments further demonstrate the importance of finger abduction/adduction and palm reconfiguration for successful in-hand assembly. In simulation, the proposed hand achieves a mean continuous sphere rotation success rate of 98.6% over diameters of 40-230 mm, compared with 73.8% for the LEAP Hand. After policy fine-tuning with external disturbances, the proposed hand achieves 92.8% success under disturbances from multiple directions, compared with 45.2% for the LEAP Hand. Hardware demonstrations further show in-hand rotation of objects of different sizes using policies trained in simulation. Together, these results show that the proposed architecture supports both assembly of two separately held parts and coordinated manipulation of a single object within one hand.
Barrier Certificate Synthesis for Non-Polynomial Robotic Dynamics via Polynomial Lifting
基于多项式提升的非多项式机器人动力学障碍证书综合
Chaubey, Shivam, Verdoja, Francesco, Deka, Shankar, Kyrki, Ville
Abstract
Safe operation of robotic systems requires trajectories to remain within a prescribed safe set under admissible control inputs. Barrier certificates provide such guarantees by certifying a controlled-invariant region within that set. Sum-of-squares optimization offers a systematic way to synthesize such certificates, but its direct application requires polynomial dynamics, excluding common robotic nonlinearities, including trigonometric terms. We address this limitation using exact polynomial lifting, which replaces non-polynomial dynamics with polynomial-augmented dynamics subject to lifting-induced algebraic constraints, preserving nonlinear geometry without approximation. We formulate lifted-domain joint barrier synthesis that computes a certificate with a state-feedback control witness and develop a sampled-data safety filter for zero-order-hold implementation. To assess whether the benefits of lifting persist across synthesis frameworks, we also adapt a sample-guided successive-barrier method to the lifted representation. On coordinated-turn and planar multirotor models, exact lifting improves certified coverage in both methods: at matched sample sizes, lifted successive-barrier synthesis achieves higher coverage with fewer barriers and lower computational cost, while lifted joint barrier synthesis provides higher coverage and lower computational cost than the finest tested piecewise resolution. In closed-loop experiments, the safety filter maintains feasibility and safety across all evaluated trajectories, reduces spatial conservativeness, and requires less intervention for both models.
Online post-training of vision-language-action (VLA) models requires efficient use of robot interaction and reliable policy improvement from continually collected experience. We propose asynchronous Replay-Anchored Policy improvement (RAPolicy), a framework that performs rollout and learning concurrently while grounding both critic and actor updates in replayed behavior. The critic learns chunk-level values from recorded actions and constructs Bellman targets without predicting next actions, reducing computation and dependence on action-value estimates outside replay coverage. The one-step flow actor reuses the initial noise stored during rollout and learns through advantage-weighted conditional likelihood, directly supervising the action mapping used for execution. We evaluate RAPolicy across four single-task settings and one joint five-task setting in the real world, with online training budgets of only 1--2 hours. Starting from policies fine-tuned on just 10 demonstrations per task, RAPolicy rapidly adapts to new single tasks and achieves an average 86.3% success rate. In the joint multi-task setting, RAPolicy improves overall success rate from 52% to 88% while preserving performance on already reliable tasks and improving weaker capabilities. Overall, RAPolicy substantially outperforms the baselines in aggregate task success while requiring fewer human interventions, demonstrating stable policy improvement and high online training efficiency. Project page: https://flyfaerss.github.io/RAPolicy.
Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but many existing methods still rely on direct mappings from language and visual observations to dense actions. This formulation can weaken the semantic reasoning capability inherited from pre-trained Vision-Language Models (VLMs), which are mainly optimized for visual-linguistic understanding rather than low-level control, and becomes fragile under spatial variations, including changes in object positions, scene layouts, robot embodiments, and camera viewpoints. To address these limitations, we propose H-VLA, a hierarchical VLA framework that decouples high-level key-action reasoning from low-level motion generation. H-VLA combines a Key-Action Model for predicting a key-action as the next manipulation subgoal, a Motion Planning Model for generating dense future actions conditioned on the predicted key-action, and a Unified Camera-Centric Action Space for consistent representation across datasets, embodiments, and viewpoints. We further adopt a two-stage training strategy that emphasizes key-action reasoning during pre-training and dense motion generation during fine-tuning. Experiments show that H-VLA achieves strong performance on SimplerEnv, reaching 91% on Google Robot visual matching, 84% on Google Robot variant aggregation, and 81% on WidowX visual matching. On Agilex real-robot tasks, H-VLA improves over the strongest baseline by 10, 47, and 16 percentage points under in-distribution, out-of-distribution position, and out-of-distribution scene/object settings, respectively.
"Dear LLaVA, Please Drive": A Depth-Aware Vision-Language Agent for Closed-Loop Robotic Control
“亲爱的LLaVA,请开车”:一种用于闭环机器人控制的深度感知视觉-语言智能体
Berger, Sebastian, Winter, Katharina, Flohr, Fabian B.
Abstract
Vision-language models (VLMs) provide a compelling foundation for reasoning-driven mobile navigation, offering rich contextual understanding and strong generalization from large-scale pretraining. Most existing navigation frameworks rely on imitation learning and therefore require substantial labeled trajectory data, limiting their scalability and robustness. In this work, we propose a parameter-efficient approach to fine-tune a pretrained VLM for autonomous navigation using an Imperative Learning paradigm. By optimizing against differentiable geometric cost fields rather than labeled trajectories, our model learns to generate collision-free paths exclusively from stereoscopic depth observations. We introduce a unified end-to-end navigation pipeline for natural-language-driven robotic control. This system leverages a shared VLM backbone with task-specific Low-Rank Adaptation (LoRA) modules, effectively bridging the gap from semantic target selection to low-level trajectory planning. Our approach achieves competitive Success weighted by Path Length (SPL) in unseen environments while updating less than 1% of the model's total parameters. Qualitative real-world experiments validate sim-to-real generalization and stable path planning without fine-tuning on real-world data. These results highlight a practical approach for deploying VLM-based agents on mobile robots, enabling high-level semantic navigation without the prohibitive requirement for large-scale, labeled trajectory data.
Reach-and-attach of soft robotic arms with passive suction requires accurate targeting and sufficient contact speed, yet geared actuators can impose a speed limit that improved trajectory tracking alone cannot overcome. This paper proposes an embodied snap controller that separates slow servo-driven preloading from rapid elastic release, enabling a compliant arm to move beyond its direct tendon-driven speed limit. Octopus biology motivates the controller's section-wise organizational prior, rather than reproduction of the octopus nervous system. A learned policy shared across three sections selects preloads, aim, tendon slack, and release timing, determining where, how, and when to load and release the body. The policy is optimized offline using a hardware-validated recurrent model within experimentally supported bounds. Across five optimization seeds and 400 unseen simulated targets, attachment success is $(73\pm4)\%$ at a $5\text{ cm}$ lateral tolerance, and the shared policy reaches the matched centralized controller's mean final reward after a median $17\%$ of the common evaluation budget. Hardware characterization achieves tip speeds of 1.56-1.64 m/s, at least $108\%$ above direct tendon-driven release. In 18 open-loop hardware trials across six placements, 17 exceed the 1 m/s snap threshold and nine retrieve the object, with successful retrieval at five placements. These results demonstrate a practical division of responsibility in the control problem: learned control prepares the body, and passive body mechanics execute the rapid movement needed for dynamic reach-and-attach.
Mobile gas-sensing inspection requires a mobile platform to reliably reach ordered sampling poses. We present MIRA-PRM, a mission-informed probabilistic roadmap that combines geometry-, mission-flow-, and inspection-conditioned sampling with locally adaptive connectivity and evidence-gated refinement. Eight development campaigns comprising 142,448 planner runs evaluated mission scaling, repeated-query adaptation, finite node budgets, target distributions, local map repair, ablation, cross-family comparison, and parameter sensitivity. MIRA-PRM maintained 100% seed-level success as ordered targets increased from two to eight, and achieved 97.67% seed-level success under a 40-node ceiling, compared with 67.17% for PRM and 86.50% for PRM*. Local repair reduced event time by 79.9-81.9% and node additions by 90.8-94.3% relative to rebuilding. On a separate frozen holdout of 12 maps, 48 ordered missions, and three seeds per map-mission unit, MIRA-PRM produced 96/144 validated successes, versus 82/144 for PRM, 89/144 for PRM*, and 70/144 for LE-HG-PRM. On paired common-success runs, MIRA-PRM reduced wall time by 16.89% relative to PRM* and 81.01% relative to LE-HG-PRM, but was 52.23% slower than PRM; paths were 0.54-1.86% longer. All validated paths passed independent full-segment checks and clearance non-inferiority. The results support a reliability-oriented trade-off in structured simulations and motivate external-map and physical mobile-sensing validation.
Prescribed-Time Contracting-Boundary Control of a Tendon-Driven Flexible Arm
腱驱动柔性臂的预设时间收缩边界控制
Lu, Yi, Tang, Chao, Han, Zhiji, Wang, Hongdu
Abstract
This study develops a prescribed-time performance-shaping control method for curvature tracking of a single-segment flexible arm actuated by three antagonistic tendon pairs. A Cartesian curvature representation is introduced to avoid the undefined bending direction at the straight configuration and to establish an explicit six-tendon kinematic mapping. A cubic performance boundary contracts smoothly from an initially admissible error bound to a nonzero terminal accuracy bound within a prescribed time. Based on this boundary, a dual transformation combining static symmetric error scaling and time-varying behavior shaping maps the tracking error into a fixed unit box. The resulting controller guarantees boundary invariance, prescribed-time entry into the terminal accuracy region, and subsequent asymptotic convergence. Numerical evaluations with Python and OpenCR--MuJoCo, together with a supervised reduced-order experiment on a two-section, four-channel platform, provide complementary validation. Across six experimental trials, no violation of the prescribed boundary is observed, and the proposed controller reduces the mean terminal curvature RMSE by 32.5% relative to a matched baseline, with comparable terminal-band entry times. These results support the feasibility of the proposed approach in the reduced-order experimental setting.
Humans can seamlessly adapt to both physical and digital worlds, suggesting that while a digital-to-real gap exists in embodiment, environment and task, human intelligence itself may transfer across this gap. This naturally raises a fundamental question: can the intelligence of vision-language models (VLMs) similarly generalize from the digital world to the physical world for robotic control? We investigate this question through RoboDawn, a human-intuitive interface that exposes robotic control to an agentic VLM through a compact set of discrete translation, rotation, and gripper commands. Using this interface, the VLM controls a robot in a closed loop: it observes the current visual state, reasons about the next action, executes it, and adapts subsequent decisions to the resulting state. Furthermore, we introduce an in-context learning (ICL) scheme that uses a few demonstrations to ground the VLM in both interface usage and task-solving strategies. Experiments on RoboTwin 2.0 C2R and RoboDojo demonstrate that RoboDawn achieves strong performance without task-specific robot training. In the zero-shot setting, RoboDawn outperforms several strong policies trained on benchmarkspecific robot data, while a single in-context demonstration further yields substantial performance gains and establishes state-of-the-art (SOTA) results. On RoboTwin 2.0 C2R, the success rate increases from 53.2% zero-shot to 73.6% one-shot, exceeding the solid baseline {\pi}0.5 (46.0%). Similar gains are observed on RoboDojo, where success rate improves from 35.67% zero-shot to 47.17% one-shot. The same framework also transfers to real-world robots, performing block-in-basket and block stacking on Franka.
An Empirical Study and Open Testbed for Federated Fine-Tuning of Vision-Language-Action Models
视觉-语言-动作模型联邦微调的实证研究与开放测试平台
Duan, Zhekai, Xie, Kevin Ziyang, Tan, Xinyu, Geng, Shikai, Zhou, Chengxu, Kompella, Ramana, Liu, Gaowen, Lu, Chris Xiaoxuan
Abstract
Adapting a pretrained Vision-Language-Action (VLA) model to a new robot, environment, or task requires demonstrations that are collected locally and often discarded. Federated learning is a promising approach to exploiting such distributed demonstrations by learning a shared policy. However, whether it can adapt large pretrained VLAs remains an open question, and a lack of reproducible benchmarks for pretrained VLAs and reusable training frameworks makes existing results difficult to compare. In this paper, we conduct a systematic study of federated fine-tuning of three modern pretrained VLA policies on the 40 simulated tasks of the LIBERO manipulation benchmark, and on six real-world tasks in two real-robot experiments, with demonstrations collected across two and three sites, respectively. Our study analyzes the key choices in this setting, spanning multiple federated parameter scopes, three aggregation algorithms, and evaluation under distribution shift. Based on the study, we derive a series of lessons, including the dominance of the federated scope over the choice of aggregation algorithm and the difficulty of matching centralized fine-tuning on physical robots, where cross-site heterogeneity is stronger than simulation captures. We also highlight opportunities for federated VLA learning, such as the ability to match centralized fine-tuning on heterogeneous data, to remain at least as robust as centralized fine-tuning under distribution shift, and to personalize, with each client federating part of the policy and keeping the rest local, which helps where the policy's pretraining is weak but leaves no usable global model. We open-source \decentvla{}, the model- and runtime-agnostic testbed behind the study, to facilitate future research and fair comparisons in federated VLA learning.
A Horizon-slicing Approach to Minimum Obstacle Displacement Planning for Robot Navigation
一种面向机器人导航的最小障碍物位移规划的视界切片方法
Thomas, Antony, Ferro, Giulio, Mastrogiovanni, Fulvio, Robba, Michela, Baglietto, Marco
Abstract
In this paper, we investigate the Minimum Obstacle Displacement Planning problem from a robot motion planning perspective. The problem involves determining a feasible path to a goal location by displacing movable obstacles when no collision-free path initially exists. We show that this problem is computationally challenging and, in particular, NP-hard when obstacles are modeled as polygons in the plane. Besides an exact formulation of the minimum obstacle displacement problem generalizing other problems in the literature, and the associated optimal solution, this paper proposes an approximate solution that is less intensive from a computational standpoint, and differs from the optimal solution by a fraction of the optimal cost, being able to trade-off between path length and amount of obstacle displacements.
Connectivity-Aware Exploration of Robotic Grasp Spaces
连接性感知的机器人抓取空间探索
Kazanskii, Maksim A
Abstract
Robotic grasping is typically formulated as the problem of identifying successful actions from a space of candidate grasp poses. However, the organization of successful actions within this space has received less attention. We study the multiscale structure of viable robotic grasps in $SE(3)$ and investigate whether this structure can be exploited for more efficient exploration. Using a large-scale grasp dataset, we show that successful grasp sets exhibit heterogeneous and reproducible connectivity structure across objects. We then introduce a connectivity-aware sampling strategy that incrementally explores the currently observed grasp space by prioritizing potential bridges between components, structural frontiers, boundary extensions, and geometric novelty. In controlled reconstruction experiments, the method recovers the connectivity structure of successful grasp sets substantially more efficiently than random sampling and farthest-point sampling. We further evaluate whether connectivity acquired under hidden grasp viability can improve subsequent grasp discovery, and whether structural experience from previously explored objects can be retrieved and transferred to unseen objects. These results suggest that the spatial organization of viable actions provides information relevant to grasp-space exploration beyond the viability of individual candidate actions. More broadly, they motivate structure-aware exploration as a means of exploiting the geometry of viable action spaces in robotic manipulation.
We present a fast numerical method for safe continuous-time motion planning under Temporal Logic (TL) specifications. The method generates smooth continuous trajectories that remain collision-free while robustly satisfying temporal and logical task requirements. A central component of our method is the formulation of nonconvex safety and logic constraints as unions of convex sets where associated discrete decisions are encoded in a joint feasibility graph. This graph representation allows Euclidean projection onto the feasible set and proximal robustness maximization to be reformulated as shortest- and widest-path problems, respectively. Building on this structure, we develop a nonconvex splitting method based on the Alternating Direction Method of Multipliers (ADMM), which decouples smooth spatio-temporal trajectory optimization from nonsmooth discrete constraint handling within the optimization. The resulting algorithm exhibits reliable convergence across benchmarks and scales to large-scale motion-planning problems, providing a 4.7x average speedup over the state of the art on discrete and continuous-time logic problems.
Chinese Translation
我们提出了一种在时序逻辑(Temporal Logic, TL)规范约束下进行安全连续时间运动规划的快速数值方法。该方法生成平滑的连续轨迹,在保持无碰撞的同时鲁棒地满足时序与逻辑任务要求。本方法的核心是将非凸的安全约束与逻辑约束表述为凸集的并集,并将相关的离散决策编码在一个联合可行性图中。这种图表示方法使得向可行集的欧几里得投影和近端鲁棒性最大化问题可分别重新表述为最短路径与最宽路径问题。基于该结构,我们开发了一种基于交替方向乘子法(Alternating Direction Method of Multipliers, ADMM)的非凸分裂方法,该方法在优化过程中将平滑的时空轨迹优化与非平滑的离散约束处理解耦。所得到的算法在多个基准测试中表现出可靠的收敛性,并可扩展至大规模运动规划问题,在离散与连续时间逻辑问题上相比现有最先进方法实现了平均4.7倍的加速。
Anatomy of a Closed-Loop Collapse: A Causal Case Study of a Compressed VLA Policy
闭环崩溃的解剖:一个压缩VLA策略的因果案例研究
Jia, Fengze
Abstract
Compressed manipulation policies can pass offline evaluation while failing in closed-loop execution; this dissociation is established in prior work and is not our claim. We contribute a causal anatomy of one naturally occurring case. An 8-layer distillation of Octo-Base retains 86% of parameters, passes every offline check we applied (0.996 and 1.000 teacher-ratios on the family's own validation metrics), and collapses in closed loop: 0/72 vs. the teacher's 40/72 on a simulated WidowX pick-and-place task. The collapse is structured, not diffuse: early task stages degrade gradually (the student moves the object at 90% of the teacher's rate and grasps at 55%), while transport-to-target fails categorically, at 0% in every training variant. Paired action-trace forensics isolate the signature: a negative, late-heavy $z$ residual, roughly 10x its post-repair magnitude, and persistent across the base distillation and both continuation branches. Four standard therapies fail under matched controls: continued training and in-domain offline data leave success at zero, even though the latter measurably improves marginal action statistics; command-level compensation recovers nothing at any offset, although the same perturbations degrade healthy policies; clamping the symptom in the command channel preserves grasping, yet success stays at floor. A minimal-pair intervention that substitutes half of the training stream with deployment-distribution teacher rollouts, with every other setting held fixed, restores parity with the teacher (18/36 vs. 17/36 held-out), eliminates that signature, and recovers a teacher-like perturbation-response profile. We claim existence, not universality. Operationally, offline gates, including a family's own validation metrics, are insufficient acceptance tests for compressed policies; a few dozen closed-loop trials sufficed to find what they missed.
We study the online deployment of mobile ad hoc networks in unknown orthogonal environments, formalized as the Partially Observable Cooperative Guard Art Gallery Problem. We give a full proof that CADENCE algorithms achieve full coverage while maintaining a connected visibility graph using at most $n/2 + h - 2$ agents in orthogonal worlds with $n$ corners and $h$ holes. We further evaluate deployment-order heuristics that reduce agent count and deployment time in practice.
Chinese Translation
我们研究了移动自组织网络在未知正交环境中的在线部署问题,并将其形式化为部分可观测协作守卫艺术画廊问题(Partially Observable Cooperative Guard Art Gallery Problem)。我们给出了完整的证明:对于拥有 $n$ 个角点和 $h$ 个洞的正交世界,CADENCE 算法能够在保持可见图连通的前提下,使用至多 $n/2 + h - 2$ 个智能体实现完全覆盖。我们进一步评估了在实际中能够减少智能体数量和部署时间的部署顺序启发式方法。
Where to look and how to move? A robot navigating an unmapped environment must do both at once, and the two goals pull against each other. The regions most worth observing are the ones the map knows least about, and those are exactly where the robot cannot trust its collision margins. We resolve this tension by introducing Splat-CBF, an active perception control barrier function that steers the camera toward the next best view while collision avoidance is enforced as a hard constraint. Safety is enforced by a risk-aware control barrier function that turns the Average Value-at-Risk of the Gaussian field into a single smooth hard constraint. Perception is enforced by a second barrier that rewards camera orientations with high expected Fisher information gain near the robot's planned path. The two meet in a quadratic program where safety is hard and perception is soft, with a slack penalty that adapts to how often perception has already been relaxed and how close the robot is to an uncertain region. We verify the method in indoor simulations, a Isaac Kinova manipulator and in experiments on an Ackermann-drive robot. Our results assert that robot navigates faster, gathers more information, and runs faster online than safety-only and perception-only baselines, giving up informative motion only when safety requires it.
While simulation-ready deformable assets are essential for in-silico robotic manipulation tasks, existing generation frameworks typically assess physical plausibility after generation, leaving an object's simulated response unused as feedback for repairing upstream errors. We present DiagGen, an agentic framework that turns a single in-the-wild image into a simulation-ready deformable asset through a generate--simulate--diagnose--refine loop. DiagGen constructs part-aware geometry and material parameters, then uses a VLM (vision-language model)-based agent to select semantically informative regions, probe them in a physics simulator, observe material responses, and route evidence-backed repair cues to the responsible generation stage. Experiments on 40 assets show that diagnostics provides useful repair cues and can moderately improve the quality of generated deformable assets. Finally, we show that unlike assets generated from visual foundation models which may not be simulatable, DiagGen-generated deformables can be directly dropped into a high-fidelity physical simulator for the planning and simulation of contact-rich pick-and-place tasks. The project's website is https://diaggen.github.io/.
Foundation models (FMs) have expanded task and motion planning (TAMP) to manipulation problems specified through language and visual observations. However, incomplete scene knowledge leaves a critical gap between understanding what the task requires and knowing whether the physical scene can actually realize it. We introduce GRAB-TAMP, an FM-based TAMP framework that searches for scene entities required for task completion, grounds functional roles to valid physical objects, and plans only after a complete joint assignment establishes functional sufficiency. We represent the task through functional roles, relations, and assignment constraints, and incrementally inspect the scene while requirements remain unresolved, verifying candidate objects through semantic, geometric, and relational checks. We evaluate GRAB-TAMP across 32 scene variants spanning Kitchen, Living Room, and Workshop domains. Across 200 feasible trials, our approach achieves 54.0% end-to-end success with 67.3% plan goal coverage. Compared with three FM-based TAMP frameworks under the same execution setting, GRAB-TAMP improves end-to-end success by 25.7 percentage points over the mean baseline. Implementation and evaluation code: https://github.com/Narendhiranv04/GRAB-TAMP
Chinese Translation
基础模型(Foundation Models, FMs)已将任务与运动规划(Task and Motion Planning, TAMP)扩展至通过语言和视觉观测指定的操作问题。然而,场景知识的不完全性在“理解任务需求”与“知晓物理场景能否真正实现该需求”之间留下了关键鸿沟。我们提出了GRAB-TAMP,一个基于基础模型的TAMP框架,它搜索完成任务所需的场景实体,将功能角色落地到有效的物理对象上,并在完整的联合分配建立功能充分性之后才进行规划。我们通过功能角色、关系和分配约束来表示任务,并在需求尚未满足时增量式地检查场景,通过语义、几何和关系检验来验证候选对象。我们在涵盖厨房(Kitchen)、客厅(Living Room)和工作间(Workshop)三个领域的32个场景变体上对GRAB-TAMP进行了评估。在200次可行试验中,我们的方法取得了54.0%的端到端成功率和67.3%的规划目标覆盖率。与在相同执行设置下的三种基于基础模型的TAMP框架相比,GRAB-TAMP的端到端成功率较基线平均值提升了25.7个百分点。实现与评估代码:https://github.com/Narendhiranv04/GRAB-TAMP
Verti-WM: A Physics-Aided Exteroceptive World Model for Off-Road Reinforcement Learning
Verti-WM:一种面向越野强化学习的物理辅助外感受世界模型
Pan, Chenhui, Xu, Tong, Xiao, Xuesu
Abstract
Reinforcement learning for off-road navigation requires extensive vehicle-terrain interaction data, which are costly to collect in high-fidelity simulation. World models offer a promising alternative by replacing simulator roll-outs during policy optimization. However, an off-road world model must condition state transitions on exteroceptive terrain information, which proprioception alone does not provide. This challenge is further amplified by the need to model both rigid and deformable terrain, where data-driven and physics-based approaches offer complementary strengths. We propose Verti-WM, a physics-aided exteroceptive world model that recurrently fuses a frozen Transformer for rigid terrain and a neuro-symbolic terramechanics model for deformable terrain. Elevation and semantic observations queried from a supplied map at each predicted pose condition fusion, enabling six-degree-of-freedom rollouts for policy optimization without further simulator access. Verti-WM reduces prediction error by 34.6% and 21.7% over data-driven and physics-based baselines, respectively. Policies trained entirely within Verti-WM achieve comparable task success rates while reducing computation time by 23.6X relative to direct training in the high-fidelity simulator. We further validate Verti-WM using real-world data, enabling policy optimization within learned real-world kinodynamics and achieving a 80% success rate on the Verti-4-Wheeler platform, compared with 40% for direct sim-to-real transfer.
Language-guided object retrieval under partial observability requires deciding whether to gather more evidence, interact with the scene, grasp a candidate, or abstain. We present a closed-loop framework that coordinates these decisions for retrieving a target specified in relation to a reference container. The framework maintains a persistent joint belief over target identity, container relation, and presence through tracked-object, unobserved-target, and target-absent hypotheses. View-conditioned categorical VLM observations update this belief; conformal grasp eligibility and robot feasibility govern commitment, while finite-horizon belief-space planning selects information-gathering actions. Across five different scenarios, our proposed method succeeds in 19/25 simulation episodes versus 12/25 for the best-performing task-adapted baseline and is the only evaluated policy to achieve at least one success in each scenario. Ablations show that cross-view memory improves success under partial occlusion, while the full system does not consistently outperform simplified variants. Real-robot trials demonstrate closed-loop re-observation and autonomous recovery from injected grasp failures, while injected viewpoint failures end in false defer. Experimental results demonstrate the feasibility of coordinating evidence gathering and selective grasp commitment within a unified framework for retrieval under partial observability.
Tip Manipulation in Soft Everting Robots via Wall Retraction and Deployable Fingers
基于管壁回缩与可展开手指的软体倒置生长机器人末端操作
Perez, Nelson Badillo, Pagliarani, Niccolo, Cianchetti, Matteo, Howe, Robert D.
Abstract
Soft everting robots can traverse long, confined paths by continuously growing, yet active interaction remains limited to a single tool fixed at or near the tip, unable to be repositioned on demand and difficult to reconcile with the robot's soft body. We introduce a tip-manipulation and multi-tool deployment strategy for soft everting robots based on wall retraction, implemented with a base roller assembly that independently meters membrane flow in the outer wall while a tail spool regulates growth in the internal tail section. Coordinated wall and tail actuation decouples robot length from membrane-material position, enabling membrane-mounted devices to be transported, exposed, and repositioned at selected locations near the distal tip. We pair this capability with ultralight pleated inflatable fingers integrated into the membrane, fabricated from TPU-coated nylon with an internal airtight bladder. The fingers achieve large bending at low pressures (approximately 100 degrees at 50 kPa in high-pleat designs) and generate blocking forces up to 1.9 N, while remaining limp during transport. The system demonstrates adaptive grasping across diverse household objects (21 to 550 g; 16 to 200 mm), three-dimensional object manipulation and stacking, environmentally braced extension, distal camera panning for confined-space inspection, and controlled sequential payload delivery. These results enable embodied and reversible tip manipulation for soft growing robots in cluttered and tortuous environments.
Recent advances in vision-language-action models have stimulated growing interest in underwater embodied intelligence. However, their reliance on large-scale interaction data limits their applicability underwater, where data collection is costly and scarce. To address this challenge, we present AquaCap, a training-free Code-as-Policy framework for autonomous underwater navigation and manipulation. AquaCap employs a dual-layer agent that translates task instructions and environmental observations into condition-aware plans and executable control programs. Structured perception then provides the agent with semantic, geometric, and reliability-aware observations under degraded underwater conditions. A failure-aware memory diagnoses unsuccessful actions and supports closed-loop replanning and code revision. This design enables online adaptation without task-specific training or parameter updates. AquaCap achieves a 66.43% success rate in simulation. Real-world experiments further demonstrate autonomous grasping and object transport with an ROV, including the manipulation of targets displaced by hydrodynamic disturbances.
3D scene graphs provide semantically rich and hierarchical representations for robot perception. However, existing systems do not maintain uncertainty as an explicit belief or propagate it through the operations that construct and refine the graph. We introduce Probabilistic Scene Graph (PSG), a generalization of the conventional scene graph that represents a posterior over possible graphs, factorized into a discrete graph structure of entities, relations, and semantic attributes, and continuous states that ground them spatially, with uncertainty maintained over both components. Geometry is carried directly by the nodes rather than selected from a separately constructed metric map, so a metric map, where needed, follows from the graph rather than preceding it. We instantiate PSG's probabilistic spatial grounding with hierarchical graphs of Gaussians (HGG): each object primitive is represented by a full-covariance Gaussian under a Normal-Inverse-Wishart belief, and the same parametrization applied recursively within a node yields a geometry graph that resolves its surface at finer resolution. We then build a mapping pipeline that preserves these beliefs throughout graph construction and refinement: a purely graph-based coarse-to-fine alignment registers observations by comparing node beliefs, while a nested Expectation-Maximization and factor-graph optimization jointly refines poses, object parameters, and internal geometry. Across six datasets spanning indoor RGB-D, outdoor LiDAR, and cross-modality deployment, HGG operates at sensor rate with near-constant memory and achieves state-of-the-art object accuracy and zero-shot graph alignment.
Chinese Translation
3D场景图为机器人感知提供了语义丰富且分层的表示。然而,现有系统并未将不确定性作为显式信念加以维护,也未在构建和细化图的操作中对其进行传播。我们提出概率场景图(Probabilistic Scene Graph, PSG),这是对传统场景图的推广,表示可能图上的后验分布,并将其分解为离散的图结构(包括实体、关系和语义属性)与将其在空间中进行定位(grounding)的连续状态,且对两个组成部分均维护不确定性。几何信息直接由节点承载,而非从单独构建的度量地图中选取;因此在需要时,度量地图由图派生而来,而非先于图存在。我们通过分层高斯图(Hierarchical Graphs of Gaussians, HGG)实现PSG的概率空间定位:每个物体基元由正态-逆威沙特(Normal-Inverse-Wishart)信念下的全协方差高斯分布表示,并且对同一参数化在节点内递归应用,得到一个以更细分辨率解析其表面的几何图。随后,我们构建了一个在整个图构建与细化过程中保留这些信念的建图流水线:一种纯基于图的由粗到精对齐方法通过比较节点信念来配准观测,同时嵌套的期望最大化(Expectation-Maximization)与因子图优化联合地细化位姿、物体参数和内部几何。在涵盖室内RGB-D、室外激光雷达(LiDAR)以及跨模态部署的六个数据集上,HGG以传感器频率运行且内存占用近乎恒定,并实现了最先进的物体精度和零样本图对齐性能。
Robot World Models Are Not Invariant to How the Actions Are Written
机器人世界模型对动作的书写方式不具备不变性
Karim, Ahmed, Chlon, Leon
Abstract
A robot policy is trained with one of two action parameterizations: absolute joint targets, or deltas relative to the current state. The choice is a live engineering decision in robot learning, and a world model conditioned on actions inherits it silently. We show the inheritance is catastrophic. A latent dynamics model trained on one parameterization and handed the identical commanded trajectory written in the other collapses: retrieval degrades by 2.6-13.4x across three robot datasets and two morphologies, goal-conditioned action selection falls from 53% to 15%, and on PushT the two beliefs about the same future are near-orthogonal (cos = 0.067, worst case -0.377), so the predictor does not degrade gracefully, it answers a different question. This is not a distribution-shift artifact in the usual sense: the two encodings are mutually reconstructible at R^2 = 0.996 given the joint input, so no information is lost, and we give the test that separates a valid re-parameterization from a lossy summary or a sensor swap. The test rejected three of the four axes we proposed. The defect lives in the action channel, which the invariance literature for visual models does not examine: work there concerns crops, jitter and camera pose, while the parameterization of the commands goes unaudited. The repair is averaging over the two encodings, and where it goes matters. Averaging the objective restores task performance by itself; averaging the outputs, safe for probabilities by concavity, is not available for direction-valued prediction, where the normalized mean can score below every member of the orbit. What objective-averaging leaves behind is the tail: worst-case agreement stays at 0.78, a disagreement penalty closes it to 0.995, and over a latent rollout it is the difference between a worst case that erodes and one that holds. On PushT, averaging alone does not repair the axis.
Scenario MPC with STL Specifications and Pareto-Based Feasibility Repair
基于STL规范与Pareto可行性修复的场景模型预测控制
Wu, Tianhao, Lyu, Yiwei
Abstract
Temporal logic is a formal language for reasoning about system behaviors over time. Signal temporal logic (STL), in particular, has been used to encode spatio-temporal requirements for control synthesis in multi-agent systems, often under the assumption that agents are cooperative and their dynamics are known. However, real-world multi-agent applications, such as autonomous driving, typically involve stochastic and uncontrollable agents. Recent work explored robust control with worst-case or probabilistic formulations, but remains limited in that it either (1) certifies strict satisfaction of STL constraints without addressing feasibility recovery, or (2) relaxes infeasible constraints with ego-centric objectives. In this paper, we propose a model predictive control (MPC) framework that treats feasibility repair as a Pareto optimization problem to explicitly characterize tradeoffs among agent objectives. We further provide a probabilistic certificate on STL violation rate to formally quantify uncertainty under stochastic and uncontrollable agents. The proposed framework is evaluated on two autonomous driving scenarios. Results show that the framework recovers feasible control with demonstrated safe behaviors.
Latent Telepathy: Multi-Robot Communication with Self-Supervised Perceptual Latents
潜在心灵感应:基于自监督感知潜向量的多机器人通信
Wang, Howard, Zheng, Han, Wu, Cathy
Abstract
In a decentralized multi-robot team under partial observability, the fact that decides a robot's next action is often visible only to a teammate. Existing decentralized methods communicate kinematic information, such as position or planned trajectory, which cannot convey what the teammate perceives. Learned communication in multi-agent reinforcement learning (MARL) can carry perceptual content, but the resulting messages are task-coupled and opaque. We propose Latent Telepathy. Each robot broadcasts the perceptual latent vector it already computes for its own use, the output of an encoder trained with a self-supervised joint-embedding predictive objective, frozen, and shared across the team. A teammate learns to act on it from task reward alone. Because the encoder already runs for perception, the message costs no additional computation and a single compact vector of bandwidth. Because the encoder is frozen before any policy is trained, the message means the same thing to every robot, and the receiving robot is never told what it means. We evaluate Latent Telepathy with a content-controlled protocol in which bandwidth, latency, topology and receiver are held fixed and only the message content varies. Broadcasting the latent lets a navigator avoid an occluded hazard in 99.7% of episodes, matching a noiseless hand-designed message. Position and trajectory messages remain at chance, and the raw camera image, 186 times wider, is less reliable than the compressed latent. The result holds from a discrete gridworld to rendered pixels under continuous velocity control, and the encoder decodes the hazard from a physical robot's camera in 102 of 102 live decisions. We also identify a requirement for porting MARL communication results to continuous control, that the decision a message informs must remain reachable by exploration, and show how to restore it.
Trajectory planning and control in field robotics rely on predicting how propulsion and steering affect vehicle motion when contact points undergo slip. For articulated vehicles, the point-contact kinematic model (PCK) accounts for the linkage geometry but neglects the rotational resistance distributed along the contacts. We propose a finite-support quadratic model (FSQ) for single-track, center-articulated vehicles that incorporates this resistance through a quasi-static balance of lateral slip. Our approach generalizes the standard PCK formulation by relaxing the contact-point assumption. An exact reduction of the quadratic slip cost to contact moments gives a compact closed-form solution for real-time prediction of lateral velocity and yaw rate. We evaluate the proposed method in real-world experiments across asphalt, grass, ice, and mixed routes, using more than 7 km of data. For five-second predictions, FSQ reduces the weighted median translation and yaw errors by 53.5% and 68.9%, respectively, relative to PCK.
Vision-language-action (VLA) policies increasingly incorporate structured intermediate supervision beyond action labels. Yet specifying what an intermediate representation should encode leaves open how action prediction learns to depend on it. We introduce \textbf{SCULPT-VLA}, a policy that learns structured control through staged action grounding. Its action-conditioning state comprises complementary factors for task progression, scene dynamics, and spatial grounding. Training first forms these factors with teacher scaffolds, then grounds coarse action prediction through their composition as scaffold inputs are withdrawn. Direct perceptual access is subsequently restored for continuous refinement, combining the learned state with perceptual detail. The curriculum separates learning to condition actions on structure from refining continuous control. Deployment requires neither teachers nor discrete-action autoregression. SCULPT-VLA achieves higher average success than shared-backbone baselines on LIBERO, SimplerEnv-WidowX, and RoboTwin 2.0 Full. On SimplerEnv-WidowX, final success is 83.5\%, versus 71.3\% when Stage-II action learning directly accesses vision and language. Across four physical robot tasks, average success under the tested distribution shifts reaches 58.1\%, compared with 45.6\% for $\pi_{0.5}$. Training ablations and factor-wise interventions support the staged design and show that the learned state continues to contribute to control after direct perceptual access is restored.
Safety-Critical Control under Uncertainty via Adaptive Conformal Quantile Prediction Intervals
基于自适应保形分位数预测区间的不确定性下安全关键控制
Zhou, Hao, Zhang, Yanze, Lyu, Yiwei, Luo, Wenhao
Abstract
Safety-critical control under uncertainty requires uncertainty representations that are both statistically valid (for certifiable performance) and compatible with enforceable safety constraints. However, existing methods often assume particular distributions of uncertainty for provable safety guarantees or establish symmetric and input-agnostic prediction intervals for robust safety, which can lead to misaligned or overly conservative safety constraints in control synthesis. In this paper, we introduce a novel safe control framework with adaptive uncertainty quantification that constructs calibrated and state-dependent prediction intervals to enable high-probability safety guarantees, while improving constrained control performance. The framework leverages adaptive conformal prediction (ACP) and extends it with conformal quantile regression (CQR) to capture distribution-free, asymmetric uncertainty intervals with certifiable probabilistic coverage, and integrates the resulting uncertainty sets into a probabilistic control barrier function formulation to enforce robust safety with reduced conservativeness. This yields uncertainty-aware safe control constraints that can be incorporated within a model predictive control(MPC) framework to provide provably safe behaviors with high probability. Simulation and theoretical results are provided to demonstrate the effectiveness of our approach.
Chinese Translation
不确定性下的安全关键控制要求不确定性表征既具有统计有效性(以保证可认证的性能),又能与可执行的安全约束相兼容。然而,现有方法通常为了获得可证明的安全保证而假设不确定性服从特定分布,或者为鲁棒安全而建立对称且与输入无关的预测区间,这可能导致控制综合中出现安全约束不匹配或过于保守的问题。本文提出了一种新颖的具有自适应不确定性量化的安全控制框架,该框架通过构建经校准且依赖于状态的预测区间,实现高概率安全保证,同时提升受限控制性能。该框架利用自适应保形预测(ACP),并结合保形分位数回归(CQR)加以扩展,以捕获无分布假设的、非对称的不确定性区间并保证可认证的概率覆盖率,进而将所得的不确定性集合融入概率控制障碍函数(Probabilistic Control Barrier Function)框架中,以在降低保守性的前提下强制实现鲁棒安全。由此得到的不确定性感知安全控制约束可嵌入模型预测控制(MPC)框架中,以高概率提供可证明安全的行为。仿真和理论结果验证了该方法的有效性。
Manipulation under time constraints requires both accurate actions and an execution rhythm that matches the evolving scene. This becomes critical when a robot must intercept moving objects or complete a sequence of adjustments before a deadline. Although one-step policies reduce generation cost, their directly predicted action sequences leave temporal allocation implicit. We propose Shared Execution-Clock Drifting (SECD), which makes execution rhythm an explicit part of one-step action generation. Conditioned on an observation and a latent sample, the policy jointly predicts a progress-indexed action curve and a shared monotone clock that maps fixed control times to locations on the curve. Demonstration-derived alignment anchors this decomposition, which is trained jointly through drifting on the decoded actions. The resulting policy retains a fixed-rate control interface and requires one network evaluation. We evaluate SECD across four real-robot tasks with inference on NVIDIA Thor. Across 300 trials, it achieves 77.00% task-averaged success and outperforms the evaluated one-step baselines on every task, including 91% success in cup retrieval from a 16 m/min conveyor and 54% in restoring and folding a crumpled shirt within 90 s. A fixed-clock variant reaches 79% on the same conveyor protocol. Complementary state-based RoboMimic experiments, including cross-seed ablations on Transport and Square, further support the joint design of the temporal representation and demonstration alignment. Project page: https://secd-anonymous-ewn.pages.dev/
In environments with large movable obstacles, detour-only navigation can be inefficient or even infeasible, while obstacle interaction requires reasoning about navigation benefit, feasible placement, and executable manipulation. We present a hierarchical navigation among movable obstacles (NAMO) framework for mobile manipulators. At the high level, the planner identifies key blocking obstacles from reference paths and searches for relocation plans that jointly satisfy geometric, manipulation, and downstream navigation constraints. When direct relocation is hindered by other movable objects, a large language model (LLM) is selectively invoked to infer auxiliary manipulation dependencies, which are then verified by deterministic geometric planning. To execute the resulting relocation goals, we define discrete contact modes on the surfaces of box-shaped obstacles and select contact faces and regions online based on position and orientation errors, enabling straight, side, and corner pushing through contact switching. A recurrent reinforcement-learning policy coordinates the mobile base and manipulator to track tool center point (TCP) targets while preserving end-effector reachability during sustained pushing. Simulation and real-robot experiments demonstrate feasible navigation-manipulation in detour, single- and multi-obstacle relocation, and dependency-constrained scenarios, validating the framework for interactive navigation with large non-graspable obstacles. The open-source project is available at https://cloudytosunny.github.io/NAMO_DCPushing/.
Chinese Translation
在存在大型可移动障碍物的环境中,仅绕行的导航方式可能效率低下甚至不可行,而与障碍物交互则需要综合考虑导航收益、可行放置位置以及可执行的操纵动作。我们提出了一种面向移动操纵机器人(mobile manipulator)的分层式可移动障碍物间导航(Navigation Among Movable Obstacles, NAMO)框架。在高层,规划器从参考路径中识别关键阻挡障碍物,并搜索同时满足几何约束、操纵约束与下游导航约束的搬移规划。当直接搬移受到其他可移动物体阻碍时,系统有选择地调用大语言模型(LLM)来推断辅助操纵依赖关系,并通过确定性的几何规划进行验证。为执行所得到的搬移目标,我们在箱形障碍物表面上定义了离散接触模式,并根据位置和姿态误差在线选择接触面与接触区域,通过接触切换实现直线推、侧推和转角推。一个循环强化学习策略协调移动底盘与机械臂以跟踪工具中心点(TCP)目标,同时在持续推动过程中保持末端执行器的可达性。仿真与真实机器人实验在绕行、单障碍物与多障碍物搬移以及依赖约束场景中展示了可行的导航-操纵能力,验证了该框架在与大型不可抓取障碍物进行交互式导航方面的有效性。开源项目见 https://cloudytosunny.github.io/NAMO_DCPushing/ 。
Underwater human--robot interaction requires gesture commands that are both easy for divers to use and reliable for robots to recognize. We investigate these aspects through a closed-loop diver--robot interaction framework integrating a compact seven-gesture vocabulary, lightweight landmark-based recognition, and command-level interaction logic. We evaluate the framework through a user study and underwater robot experiments in a laboratory tank and a swimming pool. The user study supported the reproducibility of the gestures after brief learning. Recognition analysis further showed that visual similarity was associated with gesture confusion, while intermediate poses during gesture formation introduced temporal ambiguity. Command-level processing mitigated the effects of transient recognition errors on robot execution, reducing unintended triggers and premature task interruptions. These findings show that reliable underwater gesture interaction depends on human usability, gesture recognizability, and execution reliability in underwater interaction.
Socially competent humanoid robots must communicate affect and intent through gesture as well as speech, yet open-ended interaction must become motion that is both expressive and executable on a specific body. This demands semantic flexibility for contextual social intent while preserving deterministic, embodiment-aware robot control. We present EmoPose, a vision-language model (VLM)-guided framework that bridges this gap through an executable semantic interface. Given language, dialogue history, and optional visual context, the VLM selects an ordered gesture plan containing a communicative class, library variant, intensity, and speech anchor. A scalable robot-owned motion library defines the available expressive vocabulary and the source of 14-DoF joint targets. Pose Studio supports automatic trajectory generation, MuJoCo preview, and automatic synchronization of new library entries with the VLM guide; deterministic robot-side modules validate plans, construct trajectories, schedule gestures, and manage queueing and interruption. This division lets the interaction repertoire grow for new social contexts without changing the control interface or delegating raw joint commands to the foundation model. On the EmoPose-Bench, structured GPT-5.5 planning reaches $98.25\pm0.52\%$ on the Easy tier and $76.50\pm0.54\%$ overall, exceeding same-model direct-label prompting. Further tests validate dialogue-context use and ordered multi-action composition. The system completes the nominal MuJoCo suite and realizes all 29 authored variants on the physical Unitree G1. A four-stop laboratory tour demonstrates expressive narration with interruption, camera-grounded dialogue, and navigation.
HEARTH: An Object-Centric RGB-Thermal-3D Dataset for Temperature-Aware Robot Manipulation
HEARTH:一个用于温度感知机器人操作的对象级RGB-热成像-3D数据集
Su, Yuning, Li, Borui, Shi, Yonghao, Liu, Bofei, Yang, Xing-Dong
Abstract
Language-guided manipulation can depend on physical properties that visible appearance does not reveal. Temperature is one such property, but object datasets for robot learning rarely associate measured temperatures with object appearance and geometry. We present HEARTH, an object-centric RGB-thermal-3D dataset of 90 physical objects from 18 everyday categories, comprising 145 captured object states. Our pipeline maps apparent surface temperatures onto reconstructed meshes through camera calibration and pose transfer. The dataset includes raw temperature measurements, camera parameters, RGB-textured meshes, and thermal textures for simulation. We use these assets to construct three LIBERO-derived tasks and collect 1,200 demonstrations for fine-tuning a pretrained vision-language-action (VLA) model, $\pi_{0.5}$. In an ablation study, adding thermal observations to the VLA increases success on temperature-dependent object-selection tasks from 35.0% for the RGB-only baseline to 75.0%. These results demonstrate the utility of HEARTH for training robot policies to follow temperature-related instructions.
Vision-language navigation (VLN) has largely been developed for indoor and terrestrial robots, where language can often be treated as a static goal and motion is approximated by discrete or near-instantaneous actions. These assumptions break down for unmanned surface vehicles (USVs): river navigation requires continuous motion under inertia and limited maneuverability, while long-horizon instructions must be executed through sparse and visually ambiguous maritime landmarks. We introduce RiverVLN, to our knowledge the first benchmark designed for long-horizon USV VLN under continuous riverine motion, and PGT-NAV, a phase-grounded temporal navigation framework for USVs. Rather than directly mapping an entire instruction to motion, PGT-NAV converts it into an ordered sequence of visually verifiable semantic phases and maintains the active phase online through grounded visual and motion evidence. This explicit semantic progress state is fused with visual-motion history and phase-specific grounding to predict six local SE(2) pose increments. The resulting trajectory is executed in a predict-execute-re-observe loop, where the vessel executes toward W3, updates phase and grounding, and replans through a map-based safety layer. Experiments show that PGT-NAV substantially reduces recursive position and heading drift relative to GNM-style and ViNT-style baselines and achieves an average success rate of 0.79 in Unity-ROS closed-loop navigation. Unseen bridge-opening trials and real-world USV experiments further demonstrate that the phase-grounded representation transfers from controlled evaluation to physical USV deployment.
Dynamic rope manipulation is highly sensitive to unknown object dynamics: the same robot motion can produce substantially different responses across ropes, while explicitly identifying the relevant physical properties is difficult. We present RopeFormer, a history-conditioned framework that uses prior task interaction as context for subsequent control. The policy retains cross-trial action-response history while keeping its weights fixed and requires no explicit online rope-parameter estimation. In matched simulation evaluations across sustained single-arm rotation, bimanual rotation, and transient whipping, retaining context improves subsequent control relative to resetting the same checkpoint, with the benefit varying across rope dynamics and observation settings. We further deploy the frozen policies on a Unitree H1-2 with previously unseen physical ropes. From T1 to T3, target-acquisition time decreases by 30.9% for Rope Swing and 33.9% for Rope Twirl, while mean Rope Whip target hits increase from 0.2 to 2.3 out of three. These results show that prior interaction can provide effective control context for dynamic deformable-object manipulation. Robot videos, code, and data are available at https://ropeformer.github.io/.
Non-prehensile manipulation is practical for relocating large, heavy, or geometrically ungraspable objects. Yet, long-horizon pushing of arbitrarily-shaped 3D objects couples three problems: 1) where to push the object so as to approach the target pose, 2) whether each push is stable and reachable, 3) whether subsequent actions remain feasible. We present an object-centric pushing policy within a feedback-guided hierarchical framework. At the low level, a learning-based policy predicts contact actions from a pose- and scale-normalized point cloud, conditioned on a near single-step subgoal. A stability score is applied to evaluate the predicted contacts by a quasi-static sliding-versus-tipping analysis. At the high level, BIT$^*$ first searches for an object path, and the next several subgoals are checked by contact prediction and robot motion planning for future feasibility. Failed motion plans, as feedback, change the local path costs and trigger re-planning. During execution, only the first feasible action is executed. In simulation, we evaluate 22 objects in six different scenes, upon which we also conduct comprehensive ablation studies. Results demonstrate that our method outperforms baselines with a clear margin and can reliably achieve long-horizon object pushing tasks under different situations. We also report quantitative real-robot experiments with a Franka arm and qualitative demonstrations with a mobile manipulator for large and heavy objects, with directly zero-shot sim-to-real transfer.
Bimanual manipulation requires policies that coordinate two arms while adapting their functional roles to scene geometry, object configuration, and task context. Learning such scene-conditioned role adaptation remains challenging, as demonstrations may contain uneven role distributions that limit generalization to underrepresented arm--role configurations. In addition, many bimanual policies predict actions in fixed left- and right-arm action spaces. While this provides a natural parameterization for robot control, it does not explicitly specify how behaviors should transform when functional roles are exchanged across arms. Across different scene initializations, the two arms may follow a similar coordination pattern, but the role-specific behavior assigned to each arm should change with the scene. Therefore, we propose BiRoAD, a Bimanual Role-Adaptive Decomposition framework for learning shared and role-adaptive representations in bimanual policies. Given bimanual trajectory or action-token features, BiRoAD decomposes these features into swap--symmetric and swap--antisymmetric components: the former captures coordination structure invariant to arm exchange, and the latter captures role-specific distinctions that vary consistently with functional role assignment. The two components are then recomposed as residual updates to the original paired arm representations, allowing BiRoAD to serve as a modular feature transformation without changing the policy inputs, imitation-learning objective, or requiring manually defined role labels. Across multiple bimanual manipulation tasks with balanced and imbalanced role distributions, BiRoAD improves robustness across role configurations over corresponding base policies, with notable gains on underrepresented role configurations.
Smartphone GNSS Booster: Centimeter-Level Pedestrian Positioning Using a Portable Signal Re-Radiator
智能手机GNSS增强器:利用便携式信号再辐射器实现厘米级行人定位
Suzuki, Taro
Abstract
High-precision positioning using the global navigation satellite system (GNSS) embedded in smartphones is demanded for sidewalk-level pedestrian navigation and pinpoint location-based applications. However, current smartphone positioning accuracy for pedestrians remains at the meter level. This is mainly because limitations of the compact linearly polarized (LP) antenna in smartphones increase GNSS observation noise and hinder stable carrier phase tracking. In this paper, we propose an external GNSS signal re-radiation system that boosts smartphone GNSS observation quality while still using the smartphone's built-in GNSS receiver and antenna. The system directly connects a compact active helical antenna and a thin passive patch antenna for re-radiation, and is designed to be attached directly to the smartphone. The proposed GNSS booster achieves (1) reduced thermal noise and stable signal tracking via increased received signal strength, (2) improved multipath robustness compared with a LP antenna, and (3) stabilization of antenna phase center variation. Static experiments confirm a substantial improvement in the observation quality of GNSS carrier phase measurements compared with standard smartphone positioning. Furthermore, in pedestrian experiments, the proposed booster enabled stable carrier phase integer ambiguity resolutions, achieving centimeter-level positioning using a smartphone.
Humanoid loco-manipulation faces two prominent limitations: controllers using continuous velocity commands cannot precisely regulate individual footholds, while specialized foothold-tracking modules are difficult to integrate with whole-body manipulation. Furthermore, standard action-based imitation distillation primarily transfers expert actions, without explicitly encouraging a shared representation of heterogeneous skills. This paper introduces STRIDER, a hierarchical multi-gait framework to bridge these gaps. The framework integrates terrain-aware 3D stepping logic, Adversarial Motion Priors (AMP)-based natural walking, and Cartesian upper-body control: its stepping expert selects feasible footholds in the stance-foot frame and generates clearance-aware swing trajectories. To fuse distinct walking and stepping experts into one executable student policy, we propose Latent Distillation Proximal Policy Optimization (LD-PPO), a distillation algorithm augmented with teacher-conditioned latent alignment. By jointly optimizing on-policy reinforcement learning, DAgger-based action reconstruction, and latent alignment, LD-PPO transfers expert actions while encouraging a shared skill representation across heterogeneous modes. Simulation and real-robot evaluations on the TianGong Omni humanoid show that LD-PPO outperforms vanilla distillation-PPO in foothold-tracking and posture-tracking accuracy. Deployed on hardware, STRIDER realizes multi-gait loco-manipulation with accurate foothold and end-effector tracking.
Robots operating in human-centered environments are typically designed to execute explicit instructions, and most robot-learning datasets likewise pair observations with task instructions or low-level actions. Although recent work has begun to explore proactive embodied assistance, existing resources target different settings and action levels, leaving real-world human-centered multimodal decision-making underexplored. We formulate this problem as \textit{Proactive Robot Action Reasoning} (\textit{ProRobo}), an upstream cognitive decision problem in which a robot must determine which action to take based on multimodal human and environmental cues without explicit action instructions. To support ProRobo, we introduce \textit{ProAction}, a real-world multimodal dataset containing 10K samples of visual observations, audio signals, and text inputs across 12 daily-life scenarios in five common scenes. To construct cognitively grounded high-level action supervision, we develop a two-stage human-in-the-loop pipeline that combines appraisal-guided candidate generation with Affective Theory-of-Mind-guided human refinement, explicitly incorporating contextual judgment about human states, urgency, feasibility, and potential risk into action annotation. Based on this supervision, we benchmark representative Multimodal Large Language Models (MLLMs) and introduce \textit{MMC2Act}, a reference model that implicitly learns the mapping from multimodal observations to cognitively grounded high-level actions. Experiments across modality settings, subject-disjoint generalization, cross-dataset transfer, and human evaluation show that general-purpose MLLMs struggle with proactively reasoning high-level actions from multimodal cues, whereas training on \textit{ProAction} substantially improves performance.
Existing wearable exoskeleton architectures are typically constrained by a single mechanical output modality, providing either joint torque around an anatomical joint or linear traction along a limb-training-oriented direction, which limits adaptability to diverse training scenarios. This letter presents a cable-driven switchable actuator (CDSA) that can rapidly switch between torque and tension modes while centralizing all sensing and actuation components at the proximal drive unit. A Coupled Movable Pulley Mechanism (CMPM) provides tension amplification at the distal end-effector, while a bidirectional Cable-Driven Ratchet Mechanism (CDRM) enables mode switching and preload regulation. To eliminate the need for distal instrumentation, multi-source proximal sensors are integrated with a data-driven fusion model to estimate distal output forces. An adaptive dual-mode force control strategy based on iterative learning control (ILC) is further developed. Platform experiments demonstrate transmission efficiencies of $(92.4 \pm 2.0)\%$ and $(96.5 \pm 3.3)\%$ in the torque and tension modes, respectively, along with a tension amplification ratio of $2.77 \pm 0.10$ under tension mode. Tracking tests on simulated knee-joint gait trajectories and short-stroke tension profiles yield stable control, with RMSEs of $(4.52 \pm 0.51)\%$ and $(3.15 \pm 0.19)\%$ of the uncontrolled peak value, respectively. Finally, seated human-coupled experiments validate the system's controllable force generation in both joint-torque and linear-traction application modes.
FeasibleFlow: One-Step Joint Transport of Configuration Feasibility and Trajectories for End-to-End Driving
FeasibleFlow:面向端到端驾驶的构型可行性与轨迹一步式联合传输
Li, Xiang, Wang, Bikun, Xu, Qing, Wang, Jianjun
Abstract
End-to-end autonomous driving maps current observations directly to future trajectories, yet those trajectories must remain valid as the scene evolves. Future state modeling aims to address this temporal mismatch, but general representations often contain information unrelated to ego planning and affect trajectory generation only through auxiliary supervision, static conditioning, or proposal evaluation. We propose FeasibleFlow, a one-step end-to-end generative framework that jointly transports a configuration-space feasibility field and multimodal ego trajectories. Our Asymmetric Joint MeanFlow uses the pathwise Jacobian-vector product in the MeanFlow identity to incorporate field evolution into trajectory transport. Because safety feedback is sparser than progress feedback, we further introduce the Anchor-relative ranker (ARR) and Pareto-ReinFlow to balance safety and progress in candidate selection and generation, respectively. Experiments on the NAVSIM benchmark demonstrate the strong performance of FeasibleFlow and validate both the joint transport of feasibility and trajectories and the proposed safety-first mechanisms.
We present Elevator-VIGS, a visual-inertial 3D Gaussian Splatting SLAM system that keeps tracking and mapping through elevator rides. Inside a moving elevator, the two sensors are in conflict. The camera sees only the robot's motion relative to the elevator, while the IMU senses that motion plus the elevator's motion relative to the world. This conflict is challenging for existing visual-inertial estimators. If vision dominates, the estimator tracks only the robot's motion within the elevator and misses the elevator's rise, and if the conflict remains, the estimator diverges. We observe that the conflict comes from forcing both observations into a single coordinate frame. We instead estimate the robot's pose in the elevator's coordinate frame, and the elevator's motion relative to the world as a per-keyframe transport state, the elevator's rise and vertical velocity, within dense visual-inertial bundle adjustment. Elevator-VIGS detects rides zero-shot with a vision-language model and a depth network, and constrains the transport state at the departure and the arrival. We record real-world and simulated elevator sequences. On these sequences, Elevator-VIGS achieves state-of-the-art tracking and rendering performance. On four elevator-free public benchmarks it keeps the state-of-the-art performance of VIGS-SLAM. Project page: https://ruizhou-cn.github.io/elevator-vigs/.
Task-oriented grasping (TOG) requires robots to grasp functional parts of objects (e.g., the handle of a mug for pouring), yet these affordance regions are frequently occluded in cluttered scenes. Active perception via next-best-view (NBV) planning can resolve such occlusions by moving the camera for more informative observations. However, existing NBV methods typically optimize viewpoints for grasping the target object as a whole without distinguishing which part is task-relevant. A naive adaptation, fully scanning the target object before predicting the affordance, wastes most of the viewpoint budget on task-irrelevant surfaces (e.g., the mug body for pouring). To address this, we propose ATAP, an Affordance-Targeted Active Perception framework that shifts viewpoint planning from exhaustive target scanning to targeted affordance verification. ATAP hypothesizes the occluded target geometry via a generative shape prior and predicts the affordance distribution over the imagined complete surface. In cluttered scenes, severe occlusion can make the location of the hidden affordance ambiguous, leaving multiple locations plausible given the partial observation. ATAP therefore introduces an uncertainty-aware viewpoint planner that jointly optimizes expected entropy reduction over these competing hypotheses and expected affordance verification gain from real observations. This process iterates until the affordance is sufficiently verified for grasp execution. Experiments in simulation and real-world cluttered scenes show that ATAP substantially improves the functional grasp success rate over fixed-view TOG baselines, and outperforms reconstruction-based active perception with over 57% fewer NBV steps.
Chinese Translation
任务导向抓取(Task-oriented grasping, TOG)要求机器人抓取物体的功能性部位(例如,为了倒水而抓取马克杯的把手),然而这些功能可供性区域在杂乱场景中经常被遮挡。通过下一最优视角(next-best-view, NBV)规划实现的主动感知可以通过移动相机获取更具信息量的观测来消除此类遮挡。然而,现有的NBV方法通常针对将目标物体作为一个整体进行抓取来优化视角,而未区分哪个部位与任务相关。一种朴素的适应方法是在预测功能可供性之前对目标物体进行全面扫描,这会将大部分视角预算浪费在与任务无关的表面上(例如,倒水任务中的杯身)。为解决这一问题,我们提出了ATAP(Affordance-Targeted Active Perception,功能可供性目标主动感知)框架,该框架将视角规划从穷尽式目标扫描转变为针对性的功能可供性验证。ATAP通过生成式形状先验对被遮挡的目标几何结构进行假设,并在想象出的完整表面上预测功能可供性分布。在杂乱场景中,严重的遮挡可能使隐藏的功能可供性位置变得模糊,导致在仅有部分观测的情况下存在多个看似合理的位置。因此,ATAP引入了一种不确定性感知的视角规划器,该规划器联合优化对这些竞争性假设的预期熵减少量以及从真实观测中获得的预期功能可供性验证增益。这一过程不断迭代,直到功能可供性得到充分验证以执行抓取。在仿真和真实世界杂乱场景中的实验表明,ATAP相比固定视角的TOG基线方法大幅提升了功能性抓取的成功率,并且优于基于重建的主动感知方法,同时NBV步数减少了超过57%。
PINGU: Extending Air-Bearing Spacecraft Emulators with Open-Source Actuators and Learned Control for Contact-Rich Proximity Operations
PINGU:通过开源执行器与学习控制扩展气浮航天器模拟平台,面向接触密集型近距离操作
Castan, Ricard Marsal I, Uchida, Akiyoshi, Arora, Aman, Lima, Pedro, El-Hariry, Matteo, Orsula, Anrej, Grella, Francesco, Richard, Antoine, Pradalier, Cedric, Olivarez-Mendez, Miguel A.
Abstract
Low-cost planar air-bearing testbeds have matured into a standard proxy for free-flying spacecraft GNC, but they remain largely thruster-only and are rarely equipped for contact-rich, inertia-coupled manipulation. Building on the open-source ATMOS testbed, we contribute a reaction wheel and two force/torque-sensed robotic arms (LEVION) with interchangeable end-effectors, integrated as first-class control actuators through a unified ROS 2 abstraction layer. On top of the software stack we build a reinforcement-learning training environment and digital twin, and a controller that exploits these added degrees of freedom, letting classical optimal controllers and learned policies be swapped on the same hardware without modification. We validate the integrated system, PINGU, across four benchmark tasks: point-to-pose navigation (classical LQR vs. sim-to-real PPO), dynamic disturbance rejection under arm-induced center-of-mass shifts, reaction-wheel momentum stabilization, and force-controlled docking. The results show that these additions extend an ATMOS-class emulator into the contact-rich regime and bridge classical optimal control and reinforcement learning on one reproducible platform.
As AI agents become increasingly capable, agent-driven robotic control is emerging as a compelling paradigm. However, prevailing vision-language-action (VLA) models and world action models (WAMs) still rely on natural-language instructions to specify manipulation tasks, an ill-suited interface for agent-driven control: referentially ambiguous, spatially imprecise, redundant with the agent's inherent language understanding, and entangling intent with execution. We present AR-WAM, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt (a bounding box of the target) denoting the interaction object and location, and a learnable operation token dictating the atomic skill to execute. Our compact 0.5B-parameter model, with a frozen pretrained visual encoder and no language encoder, predicts scene evolution within compact latent states while decoding actions, exposing the policy's intent through explicit, supervisable reasoning signals. A model-agnostic compatibility layer provides three primitives (detect, execute, and query) so that local VLMs or online agent APIs can drive the policy directly, with long-horizon memory and closed-loop error recovery delegated to the agent side. On RoboTwin 2.0, RMBench, and a real Astribot S1 dual-arm platform, AR-WAM matches the strongest baselines on standard manipulation (87.2% average success) and outperforms them on memory-dependent and real-robot long-horizon tasks, improving success rates by 5.9% and 36.7%, respectively, while maintaining the lowest inference latency (14.1 ms).
TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation
TaskAnchor:在反应式视觉-语言-动作模型中为长时程操作引入任务状态锚定
Liu, Hengyan, Zhou, Wenlve, Yue, Bo, Su, Yongyi, Wang, Ruixiang, Zhang, Zhanqi, Lu, Dekun, Gao, Wei, Xing, Xiaofen, Jia, Kui
Abstract
Reactive vision--language--action (VLA) models struggle with long-horizon manipulation when visually similar observations can correspond to different actions depending on the task stage or interaction history. We refer to this ambiguity as task-state aliasing and introduce TaskAnchor, a lightweight adapter that grounds pretrained VLAs in execution history. TaskAnchor combines history-conditioned visual refinement with a milestone-supervised task-state coordinate, a scalar representing the semantic stage of execution. These signals are injected through the native visual and language interfaces, respectively, without introducing an explicit planner or modifying the action-generation mechanism. On RMBench, TaskAnchor achieves approximately 4.9--5.5$\times$ the average success rates of the published $\pi_{0.5}$ and X-VLA baselines, with consistent gains on RoboMemArena and real robots. The added latency is only 2.08\,ms per action chunk for $\pi_{0.5}$.
Simulation-trained humanoid proprioceptive odometry faces two transfer challenges: training trajectories generated by specific robot control policies intended for deployment cover only a limited range of motions, while sim-to-real mismatch can make unconstrained predictions unreliable. We address both with Prior-Informed Odometry from Human-Motion Tracking (PRIMO). On the data side, we generate odometry supervision by having the humanoid track diverse retargeted human motions in simulation, decoupling supervision from the deployment policies and broadening the training motion distribution. On the model side, a Prior-Informed estimator uses physics- and symmetry-informed priors to structure velocity and rotation prediction and a coarse raw-context pathway to preserve sensor context alongside encoded features, thereby strengthening sim-to-real generalization. Under a unified real-robot protocol, PRIMO reduces mean error by 31.6%-61.7% relative to the strongest evaluated external baseline in each domain-metric comparison. Across two locomotion-policy revisions, policy specialists exhibit symmetric crossover, whereas Tracking-Locomotion training reduces mean opposite-policy simulation error by 86.8%-94.6%. On real dynamic motion, Tracking-Locomotion training reduces mean error by 69.2%-81.7% relative to training on the union of both deployment policies. Across the tested motion compositions, the Prior-Informed estimator consistently lowers mean trajectory errors relative to its Unconstrained counterpart in both simulation and real-robot evaluation. Code is available at https://github.com/Agibot-Spatial-Intelligence/PRIMO.
CompVLA: A Variable Compliance Vision-Language-Action Model for Contact-rich Manipulation
CompVLA:一种面向接触密集型操作的柔顺可变视觉-语言-动作模型
Kim, Jongmin, Ha, Junsu, Park, Che-Sang, Song, Minchang, Jeong, Hyeokju, Hwang, Himchan, Fu, Jianlong, Park, Frank C.
Abstract
Contact-rich manipulation, requiring robots to regulate not only motion but also how they yield to external forces, has emerged as the next frontier for Vision-Language-Action (VLA) models. However, existing VLAs output purely kinematic commands, degrading performance on real-world contact-rich tasks. In this paper, we introduce CompVLA, a unified VLA framework that jointly predicts motion and stiffness matrix from RGB and language inputs. Our approach augments the conventional architecture with a dedicated Compliance Expert, which outputs time-varying stiffness and virtual displacement profiles executed via geometric impedance control. We demonstrate that CompVLA achieves the highest average success rate across diverse contact-rich tasks, outperforming both vanilla and compliance-aware VLA baselines, with ablations confirming each component is essential.
Spiking Neural Network Actor-Critic Proximal Policy Optimization Control for Autonomous UAV Navigation Through Constrained Openings in Civil Infrastructure and Buildings
Walugembe, Francis Noah, Wielgosz, Maciej, Goričan, Tomaž, Mertik, Matej
Abstract
Autonomous navigation of unmanned aerial vehicles in constrained three-dimensional environments has been a challenge in the robotics domain. The application of autonomous unmanned aerial vehicles in civil infrastructure inspection involves the use of such vehicles in bridge inspection, tunnel inspection, and structural inspection. The use of deep reinforcement learning in the autonomous navigation of unmanned aerial vehicles has been successful in constrained environments. However, the computational cost of the algorithm limits the application of the algorithm in the autonomous navigation of unmanned aerial vehicles. This paper proposes the use of the spiking neural network-based Proximal Policy Optimization algorithm in the autonomous navigation of unmanned aerial vehicles in constrained sequential environments. The proposed algorithm integrates the use of spike-based actor-critic reinforcement learning with the Proximal Policy Optimization algorithm. The proposed algorithm uses the stochastic Gaussian policy in the autonomous navigation of unmanned aerial vehicles. The proposed algorithm was implemented in the autonomous navigation of unmanned aerial vehicles in constrained 3D environments. The proposed algorithm was successful in completing 1913 episodes out of more than 3000. The proposed algorithm was successful in passing an average of 2.10 windows per episode. The proposed algorithm was successful in achieving a success rate of 63.77%. The proposed algorithm was successful in achieving success rates of more than 90% in the later stages of the algorithm.
Vision-language-action (VLA) models have achieved strong performance in embodied manipulation, but still lack a clear mechanism to balance behavioral stability with task-semantic sensitivity. We identify two complementary failure modes. Under task-preserving changes, where task semantics remain unchanged but scene appearance varies (e.g., style, illumination, clutter, or paraphrasing), policies often exhibit unnecessary action drift. Conversely, under semantic-breaking changes, where key task semantics such as the target object or constraint are altered, policies frequently fail to produce sufficiently distinct behaviors and instead follow the original trajectory. To address this gap, we propose BAS-VLA, a task-semantic action calibration framework built on top of a frozen base VLA. BAS-VLA adopts a breaking-centered calibration core as the default path, and introduces a selective evidence-gated preserving auxiliary that activates only when nuisance variation is detected while task semantics remain consistent. On the OpenPI-pi0.5 / LIBERO-Object Milk-Swap benchmark, BAS-VLA maintains high success on clean (98.0%) and semantics-preserving conditions (97.5%), while reducing clean-criterion success to 0.0% under deliberate target-object swaps, demonstrating strong stale-task suppression and task-semantic separation. On validated style-preserving shifts, it improves success from 42% to 70% without degrading clean performance. These results highlight that reliable VLA behavior requires moving beyond appearance robustness toward explicit task-semantic action calibration.
LiDAR-based unmanned aerial vehicle (UAV) exploration builds maps by continually selecting where to observe next. However, decisions based on the measured map provide limited foresight into spatial continuations behind occlusions, leaving potentially informative directions unrecognized. We present WOLF, a world-model-guided framework that predicts future observations to enhance autonomous exploration. In the training stage, a recurrent world model learns observation dynamics from exploration trajectories, with recurrent memory retaining the spatial context needed to interpret partial observations across successive views. Building on this context, the model combines observation history with candidate motions during exploration to predict local occupancy and visibility. To guide further sensing, a predictive frontier generation mechanism then aligns and fuses these predictions using confidence, branch agreement, and observation quality to identify promising regions. The resulting predictive frontiers join measured ones to guide geometric viewpoint selection and trajectory generation, while new scans update subsequent predictions. In simulations, our method reduces mean terminal time by 10.9% relative to EPIC in Garage at comparable coverage and increases mean coverage from 42.12% to 98.35% in Tunnel. Real-world experiments further demonstrate onboard deployment of the learned model for online inference during physical flight.
UniPoint: Unified Point-Level Sensor Fusion for Humanoid Locomotion Across Challenging Terrains
UniPoint:面向人形机器人跨越复杂地形的统一点级传感器融合方法
Li, Sicen, Chu, Zhen, Li, Chao, Zhu, Qiuguo, Wu, Jun
Abstract
Open-world deployment requires humanoid robots to cross highly heterogeneous terrain safely, with perception that simultaneously provides wide coverage, local accuracy, and redundancy against sensor failure. Existing approaches struggle to satisfy all three: one forward depth camera or nearby height sampling covers too little; odometry-corrected elevation maps drift under aggressive motion and miss thin vertical structures; image-level encoding costs grow with camera count. We present UniPoint, a humanoid whole-body locomotion framework built on multi-source point-level sensor fusion. Measurements from a 360{\deg} light detection and ranging (LiDAR) sensor and two depth cameras are early-fused into one base-frame point set. Voxelization resamples it to a fixed number of tokens encoded by linear self-attention and proprioception-queried cross-attention, decoupling forward cost from sensor count. The point set retains standing thin barriers; a single-modality failure removes only part of the tokens, so the policy degrades gracefully. A single training run with terrain-aware rewards, perception-degradation injection, and domain randomization produces one policy for all eight terrain types, deployed on an onboard RK3588 without fine-tuning. On a DR02 humanoid, 20 trials at each of nine real-world settings over seven terrain types validate the policy on 70-cm-high platforms, 100-cm gaps, thin barriers, and sparse or narrow footholds; it also generalizes zero-shot outdoors.
Marginal Calibration Does Not Compose: Hidden Dependence in Modular Robot Navigation
边际校准不可复合:模块化机器人导航中的隐含依赖性
Baral, Rista
Abstract
Robotic systems are typically composed of multiple independently developed modules that work together to perceive, predict, and act in the environment. Although each module may perform reliably in isolation, composing them does not necessarily preserve uncertainty calibration at the system level. In this work, we show that well-calibrated component interfaces do not necessarily produce calibrated downstream behavior after composition. Using a moving-obstacle prediction pipeline, we demonstrate that position and velocity estimators can each appear well calibrated individually, yet differences in how their error are correlated lead to substantially different estimates of future-state uncertainty. Consequently, assuming independence can make the system either overly confident or unnecessarily conservative, directly influencing downstream planning decisions and safety. Through simulations, we show that modeling the joint covariance restores downstream calibration and improves system performance, whereas dependence-robust uncertainty bounds enhance safety at the cost of increased conservatism. Our findings reveal a fundamental limitation of independently validating robotic modules and highlight the need for interfaces that communicate dependence information or support direct system-level calibration.
FlockDiffusion: Assignment-Conditioned Diffusion for Multi-Drone Task Allocation and Completion
FlockDiffusion:面向多无人机任务分配与执行的分配条件扩散模型
Zhura, Iana, Akopyan, Satenik, Khan, Roohan Ahmed, Cabrera, Miguel Altamirano, Fedoseev, Aleksey, Tsetserukou, Dzmitry
Abstract
Autonomous multi-drone navigation requires fleets to service distributed objectives in cluttered environments under tight computational budgets. Efficient coordination depends on task bundling, where each drone visits multiple objectives along its route. Separate solvers for cost estimation, assignment, and execution incur redundant graph search and produce long, abrupt paths. We propose FlockDiffusion, a learned framework combining a scene graph encoder, an explicit allocation head, an assignment conditioned diffusion transformer, and a closed form trajectory decoder. An autoregressive teacher provides offline supervision for parallel fleet trajectory generation. PyBullet ablations show that bundling increases task completion from 50% to 100%, while our complete teacher further reduces route cost by 8.4% relative to MAGNNET with bundling. In the optimized scalability benchmark, evaluated on 100 scenes per density with ten drones, FlockDiffusion achieves 6.2 to 7.6 times faster inference and approximately 37% shorter routes than the classical pipeline. As nominal task counts increase from 20 to 40, latency rises from 7.8 to 11.1 ms, compared with 48.0 to 75.8 ms for the baseline. In a separate evaluation across five Gazebo environments, FlockDiffusion achieves 100% planner coverage and reduces planned route cost by 15.4% relative to the baseline with bundling. These results demonstrate efficient planning under increasing task density in configurations that are demanding to reproduce with physical drone fleets.
Egocentric human data provide a principled source of supervision for learning dexterous robot manipulation. Unlike prior approaches that often collect such data in constrained or specially constructed environments, we collect in-the-wild egocentric demonstrations in real-world settings, including homes, factories, and pharmacies, etc., where people perform their ordinary tasks while wearing head-mounted cameras. This collection protocol captures diverse workflows and hand-object interactions across long-tailed object and skill distributions, but also yields visually challenging observations due to scene clutter and head-motion-induced viewpoint changes (a mean cumulative rotation of $15.93^{\circ}$/s). To address these issues, we introduce EgoWild2Dex, which transfers in-the-wild ego-human experience to dual-arm robots with dexterous hands by jointly aligning unstable egocentric views and human motions with robot observations and actions, respectively. This work offers three benefits. First, we introduce GeoFormer, a differentiable geometric transformer that warps noisy human observations toward robot observations. Second, we design a human-robot training scheme to bridge the embodiment gap, enabling high task success with limited robot supervision. Third, we release EgoWild, a 538.9-hour in-the-wild egocentric human dataset comprising 179,049 episodes, 125,961 unique task descriptions, and 1,282 object categories. On real robots, EgoWild2Dex achieves an average success rate of 96.7% across three long-horizon bimanual dexterous manipulation tasks and an average object-level zero-shot success rate of 33.3%. The data, models, and code will be released.
Robotic bin packing requires long-horizon sequential decision-making, as each object placement affects the available space for subsequent packing. Existing methods primarily rely on hand-crafted geometric heuristics that optimize predefined objectives or reinforcement learning policies learned through trial and error over predefined training configurations. Despite recent advances in multimodal large language models (MLLMs) for this task, their potential for closed-loop sequential decisions across heterogeneous packing configurations remains underexplored. To address this gap, we introduce PackLab, a comprehensive framework for developing, training, and evaluating MLLMs for closed-loop robotic bin packing. PackLab-Suite provides a physics-based simulation platform for scalable generation of diverse training packing trajectories and evaluation of their physical outcomes. PackLab-VLM is a packing-specialized MLLM that understands the evolving object and container states to jointly select objects and predict placements in a closed-loop manner. PackLab-Bench provides standardized packing scenarios at multiple difficulty levels for systematic evaluation. Extensive experiments demonstrate that, on average, PackLab-VLM outperforms conventional packing heuristics, traditional reinforcement learning methods, and general-purpose MLLMs across object sets and container configurations, highlighting the potential of MLLMs for long-horizon robotic packing. The code, model, dataset, and benchmark are available at https://github.com/Correr-Zhou/PackLab .
Risk-Aware Motion Planning and Control under Unknown Dynamics with Hybrid Observations
未知动力学与混合观测下的风险感知运动规划与控制
Zhang, Zhiquan, Ornik, Melkior
Abstract
We consider robotic motion planning and control under unknown dynamics with hybrid state observations, where state measurements are available only in parts of the state space. Existing work combines system identification, predicted reachability, graph search and controller synthesis in a hierarchical framework using local affine approximated models over polytopic state space partitioning, but requires state observations for identification and feedback control. Based on this framework, we address blind regions by selecting nominal dynamics and precomputing open-loop control sequences before observation is lost. Since the true dynamics may differ from the selected nominal model, the robot may exit a blind polytope through an unintended facet. We quantify this transition risk and incorporate the possible outcomes into a stochastic transition system. The high-level planning problem is formulated as a stochastic shortest path problem, whose policy guides controller synthesis. A case study demonstrates that the method guides the robot from an initial state to a target while balancing route efficiency and the risks associated with traversing blind regions.
ContactDP: Contact-Guided Diffusion Policy for Tight Insertion Tasks
ContactDP:面向紧凑插装任务的接触引导扩散策略
Xing, Chengyi, Yao, Shaoxiong, Romeres, Diego, Jha, Devesh K.
Abstract
High-precision connector insertion remains challenging for robotic systems due to tight mechanical tolerances, partial observability during contact, and multimodal uncertainty arising from occlusion and contact ambiguity. Successful insertion requires closed-loop contact guidance that continuously integrates global alignment cues with local contact feedback to produce stable corrective actions under interaction. In this work, we present ContactDP (Contact-Guided Diffusion Policy for Tight Insertion Tasks), a multimodal diffusion-policy framework for contact-rich insertion. ContactDP jointly integrates wrist RGB observations, fingertip tactile sensing, and wrist-mounted force-torque measurements to infer contact state and generate temporally consistent corrective motions during insertion. To ensure stable execution under contact, the learned policy operates together with a hybrid position-force controller that provides compliant low-level interaction. We evaluate our approach on a suite of industrial-grade connector insertion tasks with varying connector geometries, grasp conditions, and initial misalignment. Across all tasks, ContactDP significantly outperforms vision-only diffusion policies for performance, reliability and generalization.
Structured World-State Reasoning for Agentic Robotic Search
面向智能体机器人搜索的结构化世界状态推理
Holt, Finley R., Pabon, Luis A., Alora, John Irvin, Frey, Jonas, Pavone, Marco
Abstract
Long-horizon robotic search must resolve natural language against heterogeneous, incomplete, and often ambiguous evidence: textual information, prior maps, and observations arriving over time. The core challenge is to contextualize these streams and decide where to gather evidence before selecting a target. We present WORLDS: World-state Observation and Reasoning for Language-guided Discovery and Search, a framework that grounds reasoning in a persistent graph initialized from geospatial priors and updated by perception. Parallel Reasoners maintain competing candidate interpretations and request evidence to distinguish between them. We collect and process the requested observations with a multimodal Examiner, after which a Judge selects a grounded target or requests another pass. WORLDS achieves 51.8% navigation success across all 5,311 CityNav test episodes, the highest reported success rate, exceeding the previous published best by 15.7 percentage points under an OSM-only, high-resolution orthographic protocol. On 1,000 shared episodes, it achieves 50.0% versus 27.9% for the strongest adapted baseline using the same model, prior, sensing stack, and movement budget. Observation-based verification by the Examiner contributes 5.9 points of this success, and at a reduced reasoning-effort setting WORLDS still exceeds the adapted GeoNav baseline by 18.8 points while generating fewer tokens. We also demonstrate WORLDS on a quadrotor, which flies the generated sensing waypoints and grounds three language targets, including a vehicle absent from the map, from its onboard imagery.
Excessive force may damage tissue and increase the risk of anastomotic leakage in robotic colorectal surgery. Although the da Vinci 5 provides force sensing, this capability is unavailable on earlier da Vinci systems and many other surgical robotic platforms. In this work, we present a vision-based pipeline for estimating 3D interaction forces from soft-tissue deformation in stereo endoscopic video. We dynamically reconstruct the tissue point cloud in an object-centered coordinate frame, track tissue points with geometric constraints, and predict the 3D force vector with a neural network. We progressively evaluate the pipeline on rubber-glove phantoms, ex vivo porcine colons, and in vivo colorectal surgical video sequences. Under varying tissue orientations and positions within the endoscopic view, as well as different camera viewpoints, the proposed method achieves average root mean square error (RMSEs) of 0.77 N and 1.30 N on the phantom and porcine colon, respectively. Compared with the camera-frame representation, the object-centered representation reduces average RMSE by 51.3% and 56.7%, while geometry-constrained tracking reduces RMSE by 19.8% and 25.3% compared with CoTracker. We further qualitatively demonstrate the feasibility of vision-based force estimation on an in vivo colorectal surgical sequence, as a step toward clinical translation of vision-based, sensorless force estimation.
Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned using language, points, or box prompts, which are first transformed into a shared object-centric representation of the selected objects. This representation captures target-focused visual features and metric object geometry, which is mixed with robot state history through a multi-stream transformer to predict action chunks. Although GAMs can be run autonomously, they can also serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation. On RoboTwin 2.0, GAM achieves an average success rate of 55.3% across 50 tasks (vs. 52.0% for Spatial Forcing), including 47.6% under scene randomization (vs. 30.4% for Abot-M0), with its action policy trained only on clean-scene demonstrations. On LIBERO-PRO, it achieves a state-of-the-art average success rate of 61% (vs. 53% for $\pi_{0.5}$) across 16 perturbation settings, with the largest gains when targets are relocated or newly designated. On two real robots, GAM retains 17/20 successes under visual shift on a bimanual YAM versus 4/20 for $\pi_{0.5}$, while its composition with a Molmo2 planner on a Franka achieves 64.7% ID and 49.8% OOD step completion on long-horizon and memory-dependent tasks.
HumynexSurg-1: A Curated Expert Liposuction Dataset
HumynexSurg-1:一个经策划的专家级吸脂手术数据集
Huang, Rhea, Matlock, David L., Reich, Laurence
Abstract
Robot foundation models learn manipulation from large demonstration corpora, but surgery is missing from those corpora: across the 780-hour Open-H surgical collection, one dataset carries synchronized force and none covers an aesthetic procedure. Liposuction is the hard case, because the instrument works under the skin and the surgeon operates by feel and by judgment. Humynex Robotics builds curated expert datasets for this kind of procedure. HumynexSurg-1 is the first release: a master liposuction surgeon performing on porcine abdominal tissue while narrating every decision, recorded with synchronized suction pressure, six-axis hand force/torque, top-down RGB-D video, side video and a lavalier microphone -- 14 episodes, 42,738 frames, 35.6 minutes, 356 utterances of which 95% compile into a liposuction-specific label schema. The capture follows a patent-pending sensing plan organized around the quantities a policy needs, so a channel captured today by a model can be upgraded to a sensor tomorrow without changing the data format. This release captures the instrument motion as a tool-hand track in the side video and provides the force channel as state; the funded capture adds a measured 6-DoF handle pose, a validated force channel, ultrasound imaging of the fat layer, and palpation sensing. As a proof of concept, NVIDIA Isaac GR00T N1.7 fine-tunes on the dataset with no custom code in under an hour per run and learns the recorded sessions; scaling probes on the same episodes show where further gains come from: every new session lowers the error on an unseen session. The dataset, its label schema, its quality-assurance reports and its evaluation protocol are the product; the next capture, many short sessions across fat regions with the sensors named here, is what the probes point to.
Chinese Translation
机器人基础模型从大规模示教语料中学习操作技能,但这些语料中缺少外科手术的内容:在780小时的Open-H手术采集集中,仅有一个数据集包含同步的力信号,且没有任何数据集涵盖美容外科手术。吸脂手术是一个困难案例,因为器械在皮下操作,外科医生依靠触感和判断进行手术。Humynex Robotics正是为此类手术构建经策划的专家数据集。HumynexSurg-1是首个发布版本:由一位资深吸脂手术专家在猪腹部组织上操作,并对每一个决策进行口述讲解;采集数据包括同步的抽吸压力、六轴手部力/力矩、俯视RGB-D视频、侧视视频以及领夹式麦克风——共14个片段、42,738帧、35.6分钟、356条语音表述,其中95%可编译为吸脂手术专用的标签体系。数据采集遵循一项专利申请中的传感方案,围绕策略模型所需的各种物理量组织,使得今天由模型推测(捕获)的通道明天可以升级为真实传感器,而无需更改数据格式。本发布版本将器械运动作为侧视视频中的工具-手部轨迹提供,并将力信号作为状态提供;后续受资助的采集将增加实测的6自由度手柄位姿、经验证的力通道、脂肪层超声成像以及触觉感知。作为概念验证,NVIDIA Isaac GR00T N1.7无需任何自定义代码即可在该数据集上进行微调,每次训练耗时不到一小时,并能学会所记录的手术过程;在同一批片段上的扩展性探究实验显示了进一步提升的空间:每新增一个手术片段,都能降低模型在未见过的片段上的误差。本数据集、其标签体系、质量保证报告和评估协议即是本工作的成果;下一轮采集——在脂肪区域间开展多次使用本文所述传感器的短时程采集——正是这些探究实验所指明的方向。
Contact-rich manipulation requires estimating forces, slip and contact geometry that can remain ambiguous in scene images. Optical tactile sensors provide both visual observations of the contact surface and mechanical measurements, yet learning from these signals raises two challenges: representing contact beyond appearance and transferring its benefits to a policy that does not require fingertip observations at deployment. We introduce HapticWAM, a world-action model that combines heterogeneous tactile encoding, structured contact prediction and teacher-student distillation. Its teacher encodes gel images together with deformation, shear, distributed forces, resultant wrench and derived contact state into a frozen video backbone. Rather than predicting tactile pixels alone, the model jointly generates actions and a contact package describing future events and mechanics. Anticipatory Contact Coupling uses the previously imagined package to condition attention, preserving a contact-related input when direct tactile observations are unavailable. Haptic-Imagination Distillation transfers both contact futures and action predictions to a student that retains the generative contact head but removes its fingertip input branches. On a real-world setup, across three contact-rich pick-and-place tasks, HapticWAM Student achieves a 77% per-task mean success rate (41 of 50 starts, 82% pooled), reaching 95% on one of the tasks, outperforming the evaluated teacher and baseline configurations.
BarrierFormer: Transformer-Guided Predictive Barrier Enforcement for Safe Robot Control
BarrierFormer:基于Transformer引导的预测性障碍函数约束的安全机器人控制
Chauhan, Anandsingh, Garg, Kunal
Abstract
Control barrier functions (CBFs) have become one of the most popular tools for encoding and enforcing state constraints in safety-critical robotics. Standard CBF approaches are inherently myopic in nature as they enforce safety only at the current time step. Consequently, the system can be driven toward the boundary of the safe set where no feasible safe control exists at a future timestep. Model predictive control (MPC) based approaches address this by enforcing state constraints over a receding horizon. However, such approaches generally require the model to be known for solving a constrained optimization problem at every step, which is computationally expensive for real-time deployment. We propose BarrierFormer, a barrier-supervised transformer framework that addresses these limitations by encoding rollout-level CBF constraints in learning a model-free safe policy. A causal transformer encodes observation-action history, autoregressively generates a predictive rollout through the dynamics head to replace the model, and provides a residual correction to a nominal controller through the action head to replace the online computation. A barrier critic operating on local observations evaluates CBF constraint violations along this rollout, and a safety teacher computes barrier-consistent actions satisfying these constraints as direct supervision targets for the learned control policy. During inference, the policy maps observation-action history to control actions without any online optimization or model knowledge, enabling real-time model-free predictive safety enforcement. Evaluations across linear and nonlinear, 2D and 3D dynamical systems for safe goal-directed navigation demonstrate that BarrierFormer outperforms existing reinforcement learning (RL)-based, diffusion-based, MPC-based, and transformer-based approaches in safety rate and inference latency.
Simulation-based evaluation provides a scalable and repeatable alternative to real-world evaluation of vision-language-action (VLA) policies. However, reconstruction errors can cause simulated policy performance to diverge from real-world performance, motivating the need to assess reconstructed environments for downstream VLA policy evaluation. We present ReVeal, a real-to-sim assessment framework combining workspace reconstruction, reconstruction-level assessment, and matched closed-loop policy evaluation. Novel-View Mesh Fidelity (NVMF) and Annotated Planar Geometry Fidelity (APGF) assess observation and planar geometric fidelity, respectively. We also develop PGSR-D, a reconstruction pipeline incorporating monocular depth supervision to improve geometry where multi-view visual cues are limited. Across 8 assessment scenes, NVMF and APGF consistently distinguish the fidelity of 2DGS, PGSR, and PGSR-D. Matched evaluations of GR00T, SmolVLA, and pi0.5 across 8 humanoid manipulation tasks show consistent ordering between reconstruction fidelity and real-sim performance agreement across pipelines. Further analysis of the evaluation workspaces shows that higher fidelity is associated with stronger real-sim agreement.
Conflict scanning over synchronized robot paths requires detailed collision checking, potentially across every robot pair at every timestep, and may be repeated many times as conflicts are repaired. We present Multi-Robot SPITE (MR-SPITE), a conservative, motion-segment-based filter for accelerating these scans. MR-SPITE partitions each path into temporal intervals and assigns conservative bounds to each segment. An interval scheduler compares bounds for temporally overlapping motions: disjoint bounds certify the shared window as conflict-free, while unresolved windows are passed to the underlying collision checker. We integrate MR-SPITE into ARC and combine it with VAMP-based collision checking. For 16 Fetch robots, ARC with MR-SPITE achieves a paired median conflict scan speedup of 7.18x and reduces median planning time by 57% relative to the baseline ARC implementation with PRM+VAMP. These results demonstrate that motion-segment bounds complement configuration-level collision acceleration while preserving the behavior of the underlying discretized scanner.
Underwater robot learning relies on simulators that integrate high-fidelity hydrodynamics, convenient learning interfaces, and a credible transition to real scenarios. In this work, we present FinsSim, a reality-aligned integrated simulation platform for Sim-to-Real underwater robot learning. FinsSim first constructs high-fidelity simulation with selectable backends to adapt to diverse requirements. To facilitate underwater robot research, it further offers standard control baselines, alongside with unified robot learning workflows. For reliable Sim-to-Real transfer, FinsSim adopts a multi-sensor fusion scheme to provide low-cost yet precise localization. Moreover, it implements calibrated thruster-hydrodynamics models and a constrained wrench allocation algorithm. Bridging these modules by ROS~2, FinsSim establishes a complete Sim-to-Real transfer pipeline. Through matched simulations and experiments, it is demonstrated that reliable Sim-to-Real transfer of underwater robot control policies can be achieved with the FinsSim framework. Separate ablation studies also validate that the modules of FinsSim can address the pivotal issues of underwater Sim-to-Real from different aspects. Overall, this work aims to bridge the gap between theoretical research and practical applications, ultimately driving advancements in the field of underwater robotics.
Topology-Informed Visual Prompting For Vision Language Action Policies
面向视觉-语言-动作策略的拓扑引导视觉提示方法
Wu, Haoyang, Kumar, Abhinav, Berenson, Dmitry
Abstract
Vision-language-action (VLA) policies can struggle with manipulation tasks with complex obstacle geometries due to partial observability. These complex geometries can lead to similar visual observations or robot configurations requiring qualitatively different actions, a distinction that can be quantified using topological signatures. While motion planners with full knowledge of environment geometries and object states can reason about these signatures in planning, this information is often not known at deployment. To address this issue, we present a topology-guided visual-prompting framework that uses simulation-based planning to augment a nominal demonstration dataset and provides vision-based guidance at deployment. Our method uses a Gauss-Linking-Integral topological signature representation to capture important topological properties of the environment. Using privileged geometry information from a simulation approximation of our environment, we augment a VLA fine-tuning dataset with trajectories that move the system to a demonstrated signature and, from the new configuration, resume task execution. A vision-language model (VLM) is fine-tuned on the same dataset to both predict signatures from live camera observations and predict end-effector waypoints, which are rendered as visual prompts on the observations to guide the VLA. Across three simulated bimanual tasks and a real-world box pickup task, our method outperforms a VLA fine-tuned only on nominal demonstrations and a VLM-prompting baseline that can remove topology-relevant information from observations. On hardware, it exceeds the strongest baseline by 40% in task success. Project website: https://topology-vla.github.io.
Humanoid robots are expected to perform diverse human-level tasks in daily environments, many of which require precise regulation of interaction forces. While recent vision-language-action (VLA) models have shown promise for semantic planning and visuomotor control, existing humanoid systems primarily represent actions through geometric motion goals and rely on whole-body controllers focused on motion tracking, with limited explicit reasoning or control of interaction forces. This limitation is particularly relevant in contact-rich tasks, where geometrically similar motions may require different force regimes depending on the task context and where visual observations may become unreliable after contact. In this work, we present Opt2VLA, a force-aware VLA framework that introduces explicit force commands at the VLA-to-control interface for humanoid whole-body manipulation. A single multi-task VLA policy jointly predicts both geometric motion goals and continuous contact-force references, which are tracked by task-specific reinforcement learning (RL)-based whole-body controllers. To provide scalable and physically grounded supervision, we generate dynamically feasible and contact-consistent training data via whole-body trajectory optimization (TO) with explicit force references. We evaluate Opt2VLA on three contact-rich humanoid tasks and show that explicit force conditioning enables more accurate and consistent force regulation than motion-only control, while physically grounded torque supervision from TO further improves force tracking accuracy and stability. Closed-loop evaluations further demonstrate language-conditioned force modulation with Opt2VLA in simulation and on humanoid hardware.
Robots engaged in fast physical interactions often need to act before the intent of another agent is fully known. Anticipatory goalkeeping illustrates this challenge. Waiting provides more reliable information about the target but reduces the physical opportunity for interception, whereas acting early preserves reachability but requires initiating motion under uncertainty. Given a fixed closed-loop save controller, we formulate the decision of when to initiate motion as a policy-conditional finite-horizon optimal stopping problem. Building on this formulation, we propose monotone optimal stopping (MOS), a structured release-timing method for dynamic robotic interception. The quadruped save policy is trained with reinforcement learning, while MOS determines when the policy should be activated from the evolving robot state and target belief. Rather than predicting a release time or relying on confidence alone, MOS learns the return advantage of acting now over waiting for one more observation. We derive a direct Bellman recursion for this act-versus-wait margin and impose monotonicity only with respect to physical urgency, reflecting the irreversible loss of interception opportunity as time elapses. This structure enables early activation for dynamically demanding saves while preserving closed-loop adaptation when later observations change the predicted target. Under a single-crossing condition, MOS admits a threshold release boundary with a bounded approximation error. Extensive simulation studies show that MOS improves the mean save rate from 67.7% to 74.4% over a parameter-matched learned gate and increases reversal saves from 52.1% to 66.5%. Real-robot experiments further demonstrate rapid interception and post-release direction correction under human shot-direction feints.
Multi-robot collaboration could enable more efficient and scalable solutions to complex robotic tasks, but collaboration under partial observability remains challenging. Natural-language communication offers a promising approach to coordinating robots under partial observability. However, in decentralized manipulation, jointly learning explicit inter-robot communication and skill-level action selection from multimodal demonstrations remains underexplored for small vision-language models (VLMs) intended for on-device deployment. To address this gap, we introduce RoboTalk, a synthetic data-generation pipeline and dataset of 7,950 multimodal trajectories spanning 53 mobile-manipulation kitchen tasks for training small VLMs to communicate and coordinate. The dataset includes a leader-follower planning protocol, tool calls (perception, manipulation, navigation, and communication), rationale traces, and diversified natural-language communication. Fine-tuning open-source models on our dataset can reach 77% success on novel held-out tasks, a significant improvement over the untuned open source models, which had a success rate of around ~2%.
Reliable action evaluation in contact-rich manipulation requires looking beyond the current observation to future visual and contact consequences. Existing noise-space reinforcement learning efficiently steers a frozen Vision-Language-Action (VLA) policy, but its critics largely ignore these consequences. We present Imagine-RL, which augments noise-space VLA post-training with action-conditioned visual-torque imagination. For each candidate action chunk, a frozen visual-torque latent world model (VTLWM) autoregressively predicts compact future representations without pixel reconstruction. A current image-state-action query attends to observed histories and predicted futures, while previous-window prediction residuals provide token-wise confidence priors that suppress unreliable future tokens. By combining current evidence with predicted consequences, the action critic better evaluates candidate actions and supervises the actor, while the VLA and VTLWM remain frozen. Across four real-robot tasks with 50 evaluation trials per task, Imagine-RL uses only 100 RL trajectories and improves the average success rate by (23.6%) over DSRL and by (60%) over VLA baselines.
World Action Models (WAMs) have emerged as a promising paradigm for generalizable robot control. Despite the growing number of WAM systems, existing works often introduce unified systems that bundle together multiple design choices, such as architecture and training strategy, making it difficult to isolate individual contributions and systematically compare alternative designs. In this work, we present a controlled study that disentangles these design choices and analyzes not only their empirical effects, but also how and why they shape WAMs. More specifically, we focus on three fundamental questions in building WAMs: (1) what causal structure should govern the interaction between world modeling and action generation? (2) in which latent space should world modeling be performed? and (3) how do different world-action modeling objectives affect model behavior and performance? Through structurally controlled experiments on three representative benchmarks, RoboCasa-GR1, LIBERO, and LIBERO-Plus, we systematically compare six causal structures, eight latent representations, and four training objectives, covering popular design choices in existing WAMs. We further validate our key findings on real-robot data from the DROID dataset. We hope to provide a systematic understanding of how core design choices affect world-action modeling and what principles can guide the development of future WAM systems.
Intermittent visual loss disrupts target-relative feedback during underwater orbiting, making it difficult to maintain coordinated motion and reacquire a moving target. We present AquaOrbit, a reinforcement-learning controller with a recovery module for underwater target orbiting under interrupted visual feedback. During detection loss, the recovery module uses latched line-of-sight, roll, and depth references to support stabilization and target reacquisition. We train the controller in Isaac Sim with dynamics, observation, and vision-loss randomization. Evaluated without retraining in Gazebo/ROS2 under a different physics engine and perception perturbations, AquaOrbit completes 20/20 orbiting trials in each of the static- and moving-target conditions on an unseen variable-depth 3-D trajectory. In the moving-target condition, it reduces mean line-of-sight error by approximately 46% relative to a PID-based visual servoing controller with recovery while maintaining comparable path-tracking accuracy; removing the recovery module reduces completion to 9/20. Zero-shot physical deployment with fully onboard perception and control demonstrates elliptical, figure-eight, and variable-depth circular trajectories, including the latter two trajectory types absent from training. The robot maintains attitude stability during manual occlusions lasting up to 8s and reacquires the target within 2.5s in the reported attitude-induced field-of-view loss events.
Robots can recover from failures by asking bystanders for help, but effective human-in-the-loop recovery requires communication that accounts for differences in people's knowledge. Prior inverse-semantics work generates requests using a single listener model, leaving differences in listener knowledge untested. We introduce Listener Differences in Human-Robot Interaction (LD-HRI), a game, dataset, and benchmark that evaluates speakers through human listener performance. Our evaluation examines request properties, large language model (LLM) speakers, and inverse-semantics request-selection algorithms under controlled differences in listener information. The corpus contains 446 human-written requests and 1{,}302 listener trials. We additionally evaluated 24 frozen LLM-written requests with 70 human listeners across 560 trials. Novice success is descriptively higher with model-written requests across all four tasks, yet both request sources leave substantial expert--novice gaps, including 16 percentage points for LLM requests. LD-HRI makes these gaps measurable, providing a foundation for designing more robust communication in human-robot and human-agent interaction.
Chinese Translation
机器人可以通过向旁观者求助来从故障中恢复,但有效的人在环路(human-in-the-loop)恢复需要能够考虑到人们知识差异的沟通方式。先前的逆语义(inverse-semantics)研究使用单一听者模型生成请求,未能检验听者知识差异的影响。我们提出了人机交互中的听者差异(Listener Differences in Human-Robot Interaction, LD-HRI),这是一个通过人类听者表现来评估说话者的游戏、数据集与基准测试。我们的评估在受控的听者信息差异下,考察了请求属性、大语言模型(LLM)说话者以及逆语义请求选择算法。该语料库包含446条人类撰写的请求和1,302次听者试验。此外,我们使用70名人类听者对24条冻结的LLM撰写请求进行了560次试验评估。在所有四个任务中,新手在模型撰写请求下的成功率在描述性上更高,但两种来源的请求均留下了显著的专家—新手差距,其中LLM请求的差距达16个百分点。LD-HRI使这些差距可被量化测量,为人机交互和人—智能体交互中设计更鲁棒的沟通奠定了基础。
Automatic Labelling for Bimanual Mobile Manipulation
双臂移动操作(Bimanual Mobile Manipulation)的自动标注
Lu, Yupu, Pan, Jia
Abstract
Semantically meaningful subtask labels can provide useful contexts for long-horizon policies, but automatically identifying both reliable temporal boundaries and broad semantic descriptions for annotations remains difficult. We present an automatic labelling pipeline that assigns temporal localisation to deterministic trajectory analysis and semantic interpretation to vision-language (VL) reasoning. The pipeline segments synchronised kinematic signals into phases, performs phase-localised VL reasoning to describe the contents, and aggregates the outputs for the base, left arm, and right arm actions. We evaluate this pipeline primarily on 29 real Galaxea bimanual mobile-manipulation tasks. Repeating the VL reasoning three times first produces the same output value for 87.4% on selected tasks. A review by nine participants across all 29 tasks then judgements on the labelled phases and shows positive acceptance of temporal divisions (90.5%), body labels (90.7%), and arm labels (78.7%). The results indicate that the segmentation-VL design can produce structured annotations while preserving asynchronous bimanual behaviour, providing a basis for richer semantic subtask identification and state-based verification.
Hyper-redundant robots are well suited for confined-space manipulation due to their high dexterity, but safe operation in cluttered environments remains challenging. In addition, their slender structures often lead to uneven load distributions and nonuniform tracking errors along the body. To address these issues, this work proposes a weighted control barrier functions (W-CBFs) framework that enforces safety constraints while reducing tracking errors caused by uneven loading. The proposed controller was first evaluated on a circular path-following task under different obstacle configurations. With fixed weights, compared to the non-weighted method, the maximum reduction in root-mean-square (RMS) tracking error was 59.6\% in simulation and 87.7\% in physical experiments. An adaptive weighting strategy was then investigated based on the discrepancy between simulated and experimental performance under different mapping functions. The RMS errors were further reduced by 21.9\% and 8.5\%, respectively, although the error increases when obstacles were located close to the robot body. Finally, the robot was evaluated in a cleaning task requiring coverage of a rectangular area and compared with manual teleoperation. Although the controller was not explicitly optimized for area coverage, the autonomous strategy achieved comparable or better coverage performance while avoiding collisions with the surrounding frame, whereas collisions occurred during manual operation.
When Does Touch Matter? Charting the Vision-Interaction Gap in Cluttered Dexterous Grasping
触觉何时重要?描绘杂乱灵巧抓取中的视觉—交互差距
Jiang, Hao, Dominguez, Luis, Seita, Daniel
Abstract
Dexterous grasping in clutter poses a basic sensing question: when do tactile measurements and external wrench estimates improve on visual geometry? Occlusion and contact can obscure grasp quality, motivating a controlled evaluation of these interaction signals. We present a controlled real-world study over five tabletop scene conditions on a dexterous system that combines vision, per-finger and wrist wrench estimates, and distributed fingertip taxels. With demonstrations, visual observations, action space, and compliant control fixed, we compare vision-only, wrench, taxel, and combined policies plus representation and fusion baselines. The combined policy succeeds in 24/25 trials versus 14/25 for vision only, and 15/15 versus 6/15 across the three confined conditions. Ablations show that wrench and taxel feedback are complementary. Behavioral comparisons show that interaction feedback enables earlier rejection of inadequate contacts, regrasping before lift, and more stable grasps. To our knowledge, this is the first real-world study to combine and separately evaluate these interaction modalities for target-oriented dexterous grasping in clutter. These results chart a widening vision-interaction gap and position cluttered dexterous grasping as a benchmark for determining when the learned policy needs interaction sensing. Project website: https://interaction-dex-grasp.github.io/
Learning dexterous manipulation from demonstrations is bottlenecked by data: the contact forces that determine whether a grasp succeeds are absent from every scalable source of human demonstrations. This paper builds on two observations. First, what survives the change from a human hand to a robot hand is the contact structure of a demonstration - which finger regions touch which object locations, and in what order - rather than its joint motion. Second, physical consistency need not be engineered per task: a single residual reinforcement learning (RL) policy, trained once across diverse demonstrations, can repair kinematic recordings into physically consistent, contact-annotated trajectories, and the same residual formulation restores dynamic feasibility after retargeting. These observations yield a three-stage pipeline that converts human motion-capture recordings into dexterous robot policies with no real-robot training data: physics refinement with a simulated MANO hand recovers contacts and forces, contact-anchored retargeting transfers the demonstrated contact structure through an objective independent of hand morphology, and residual policy learning adapts the result to robot actuation. The pipeline reconstructs 25,454 single-hand trajectories (success 7.3% -> 59.3%) and 25 dual-hand tasks (16.0% -> 62.4%) with one shared policy per setting, transfers one human dataset to four morphologically distinct robot hands (+62.4 pp), and executes four contact-rich bimanual tasks on physical hardware with zero real-robot training data.
Ask Before It Tells: Benchmark-to-Robot Body-Cue Transfer for a Question-First Bedside Robot
先询问后告知:从基准数据集到机器人身体线索迁移的“询问优先”式床旁陪伴机器人
Yoon, Dongsik
Abstract
Body-cue recognition can support assistive robots, but benchmark accuracy does not guarantee reliable behavior under a robot-camera viewpoint. We present Nuni, a bedside robot prototype that treats a detected distress cue as a reason to ask rather than a reason to alert. We compare two X3D-UGT RGB appearance classifiers, which reach 97.7% and 94.8% six-way accuracy on NTU RGB+D, with a pose-centric hybrid pipeline on 28 single-actor scripted clips recorded from the robot camera. The hybrid path achieved 0.71 six-way macro recall, versus 0.25 and 0.29 for the fine-tuned and from-scratch RGB variants. More importantly for interaction, it produced a question-triggering distress cue in 12/16 distress clips and would have prompted unnecessarily in 2/8 normal clips; the RGB variants yielded a question-triggering cue in only 2/16 and 3/16 distress clips. We separately tested the question-first controller through event injection. All 13 state-transition trials passed: valid responses caused stand-down, two unanswered prompts produced one alert, and three boundary conditions were handled correctly. These results are a preliminary technical evaluation, not a user study or medical validation, but they show how interaction policy can limit the consequences of uncertain perception.
Vision-Language-Action (VLA) policies achieve strong performance in robotic manipulation but remain brittle once execution deviates from nominal trajectories. We propose CARE (Corrective Atomic Robotic Execution), a framework that improves recovery by learning from failures encountered during execution. Instead of generating corrective data from manually designed or random perturbations, CARE collects failed rollouts, models stage-conditioned post-failure deviations, and uses the resulting empirical distributions to synthesize representative failure states and corrective demonstrations. At inference time, CARE combines stage-wise planning with physically grounded 3D monitoring to trigger atomic adjustments or re-operations while preserving task progress. We further introduce the Failure State Recovery Benchmark (FSR-Bench), which evaluates recovery from intermediate failure states under local deviations and structural anomalies. Experiments across multiple VLA backbones, simulation benchmarks, and real-world dual-arm tasks show consistent improvements, with average task-success gains of 14.5 points in simulation and 15.9 points in the real world. Code, models, and data are available at https://github.com/xiaojunlan/care
Chinese Translation
视觉-语言-动作(VLA)策略在机器人操作中取得了优异的性能,但一旦执行偏离标称轨迹,其表现仍然十分脆弱。我们提出CARE(Corrective Atomic Robotic Execution,纠错原子机器人执行),这是一个通过从执行过程中遇到的失败中学习来提升恢复能力的框架。CARE并非通过人工设计或随机扰动来生成纠错数据,而是收集失败的执行轨迹,建模基于阶段的后失败偏差,并利用所得到的经验分布来合成具有代表性的失败状态与纠错示范。在推理阶段,CARE将分阶段规划与基于物理的3D监测相结合,在保持任务进度的同时触发原子化的调整或重新操作。我们还进一步提出了失败状态恢复基准(Failure State Recovery Benchmark,FSR-Bench),用于评估在局部偏差和结构性异常下从中间失败状态的恢复能力。在多种VLA骨干模型、仿真基准以及真实世界双臂任务上的实验表明,该方法带来了一致的性能提升,仿真环境中平均任务成功率达到14.5个百分点的提升,真实环境中达到15.9个百分点。代码、模型和数据可在 https://github.com/xiaojunlan/care 获取。
Active perception and manipulation are crucial for robots to interact with complex scenes. Existing benchmarks struggle to evaluate how robots effectively acquire and maintain information in memory in an active manner. To this end, we introduce ActiveArena-Sim, an active-perception simulator with controllable viewpoints and large-scale workspaces as the foundation. Built on this, we propose ActiveArena-Bench, which comprises 35 tasks across 5 fine-grained categories, covering visual exploration and interactive information acquisition. Each task is difficult to solve from passive observations alone, requiring multi-round evidence acquisition and memory-based reasoning. The benchmark provides rich memory annotations, standardized training data, and ID/OOD protocols featuring disjoint scenes, unseen distractor configurations, and novel backgrounds. Moreover, we present ActiveArena-VLA, a modular suite of 13 vision-language-action configurations for controlled studies of memory writing, memory capacity, proprioceptive state, subtask supervision, and high-level planning in active perception. Benchmark results reveal a substantial ID-OOD gap: uniform memory sampling, increased memory capacity under reliable write policies, proprioceptive inputs, and subtask supervision improve OOD generalization, while planner-guided memory management and decision-making achieve performance close to the best-performing configuration using only sparse memory. ActiveArena thus provides a unified testbed to develop and diagnose models for active perception and manipulation.
Recent advances in humanoid robotics and embodied intelligence have enabled robots to perform increasingly complex manipulation tasks. However, musical instrument performance remains a formidable benchmark, demanding not only collision-free trajectory execution but also precise contact timing, asymmetric bimanual coordination, and target acoustic outcomes on physical instruments. The guqin, a seven-string fretless zither, presents unique manipulation challenges due to its millimetric string spacing, transient right-hand plucking, and sustained left-hand harmonic contacts. In this work, we present a physical heterogeneous dual-arm robotic system for phrase-level autonomous guqin performance. We formulate guqin playing as a hybrid discrete--continuous execution problem and develop a hierarchical planning framework that coordinates working finger assignment, configuration continuity, obstacle avoidance, and tight bimanual contact schedules across consecutive musical events. The system integrates vision-guided instrument localization, tactile-based harmonic contact monitoring, and auditory feedback-informed plucking parameter calibration. Real-world experiments on a 25-event phrase demonstrate that the system reliably executes coordinated open-string and seventh-hui harmonic sequences on a physical guqin, achieving 93.6% and 96.8% event correctness across repeated trials.
BayesianGS-SLAM: Uncertainty-Aware Neural Rendering SLAM via Probabilistic Formulation
BayesianGS-SLAM:基于概率化公式的不确定性感知神经渲染SLAM
Kang, Kyeongsu, Ha, Seongbo, Lee, Sibaek, Yu, Hyeonwoo
Abstract
Neural-rendering-based SLAM relies on rendered RGB-D residuals for camera tracking and map optimization, but the reliability of these predictions can vary substantially because of sensor noise, limited observation coverage, and incomplete map representations. Without an explicit reliability estimate, unreliable residuals may adversely affect pose optimization, while frames already well explained by the current map may trigger redundant mapping updates. In this paper, we present BayesianGS-SLAM, an uncertainty-aware 3D Gaussian Splatting SLAM framework that estimates predictive color and depth uncertainty during mapping and consistently reuses it across the SLAM pipeline. Our tractable probabilistic formulation combines a sensor-noise uncertainty component with an opacity-induced map-representation component propagated through the rendering process. The resulting predictive uncertainty is used to augment mapping, normalize tracking residuals through a robust pose objective, and evaluate incoming frames using a predictive-surprise-based keyframe criterion. Unlike prior uncertainty-aware neural-rendering SLAM methods that primarily consider color uncertainty or use uncertainty only during mapping, our framework estimates predictive uncertainty for both color and depth and integrates it into mapping, tracking, and keyframe selection. Evaluations on real-world RGB-D datasets demonstrate substantially improved depth uncertainty-error ranking compared with existing uncertainty-aware SLAM methods. Moreover, the proposed keyframe-selection strategy reduces the number of selected keyframes and mapping calls while maintaining competitive tracking and rendering performance.
MimicAgent: Quadruped Skills via Prompt-to-Trajectory Generation
MimicAgent:基于提示到轨迹生成的四足机器人技能学习
Nayak, Lucky Kant, Parameswaran, Narayanan Palghat, Peri, Neehar, Ramanan, Deva
Abstract
We present MimicAgent, a prompt-to-trajectory generation framework for learning dynamic quadruped skills. Although reward shaping is extensively used when training quadruped policies, navigating the resulting reward landscape is notoriously difficult, requiring hours of "graduate student descent". Eureka attempts to automate reward design with LLMs, but we find that it struggles to generalize across diverse skills and morphologies. Our key observation is that it is far easier for a human - and by association, an LLM - to generate reference motions than to shape reward functions. Our hypothesis is motivated by the success of example-guided RL for humanoids, which exploits large-scale motion capture datasets as references for training locomotion policies. Unlike humanoids, quadrupeds lack such reference motion data. Towards this end, we propose MimicAgent, an agentic harness that, given a skill prompt, generates quadruped reference trajectories with coding agents. These coarse reference trajectories are then used to train example-guided RL policies that are deployable in simulation and in the real-world. Notably, we find that when prompting Claude Fable 5.1 within our agentic harness, 87% of prompts yield semantically aligned reference trajectories.
Robot visuomotor policies are commonly formulated as autoregressive, diffusion-based, or more recently, flow matching models. Among them, Action-to-Action (A2A) flow matching improves inference efficiency by initializing generation from historical action priors rather than stochastic noise. However, stale historical motion patterns and entangled global visual representations can jointly reduce robustness under spatial out-of-distribution (OOD) shifts and visual distractors. In this work, we propose SlotFlow, an object-centric flow matching policy for robust visuomotor manipulation. SlotFlow decouples scene observations into semantic ("what") features and lightweight image-plane spatial ("where") cues to provide object-aware policy conditioning and current-state grounding. The semantic representation suppresses irrelevant background correlations, while the spatial cue improves adaptation to shifted object configurations. Extensive simulation and real-world experiments demonstrate improved robustness under visual distractors and severe spatial perturbations while preserving the low-step inference efficiency of A2A. Controlled initialization and perception ablations further identify object-centric grounding as a major source of the gains and show that it complements, rather than replaces, useful historical motion priors.
Visual grasp proposal generation has advanced rapidly, yet converting a selected proposal into a stable physical grasp remains a central execution-stage challenge. This paper introduces GraspTune, a tactile-driven execution-stage refinement framework that starts from a nominal proposal and applies bounded residual TCP motions during approach, contact formation, and final grasp execution. GraspTune learns control-facing contact semantics from local depth, tactile signals, state, and history using state-conditioned expert contact queries and multi-task supervision for contact change, contact risk, and post-close readiness. The representation conditions a diffusion-pretrained residual policy and is aligned with PPO for closed-loop execution. Across more than 60,000 simulated executions over 20 object categories, GraspTune establishes an execution-layer benefit across four proposal generators, raising stable grasp success by +19.22, +9.55, +12.45, and +20.70 percentage points for GraspNet, Contact-GraspNet, AnyGrasp, and VGN. A four-fold held-out category study raises unseen-object execution from 54.58% to 70.33%, showing category-disjoint generalization of contact correction. Across more than 1,000 real-robot trials on a UR5e setup with Xense fingertip sensors, GraspTune raises GraspNet execution from 71.0% to 84.3%, validating direct transfer without realworld policy fine-tuning. Together, these results turn visually plausible proposals into stable physical grasps for downstream contact-rich manipulation. A supplementary video is available at https://youtu.be/kcq7fSLNtzU.
Autonomous endoscopic navigation requires the policy model to predict actions from texture-poor monocular observations, make safe control decisions, and retain evidence of lesions after they leave the field of view. Existing vision-language-action (VLA) models primarily rely on visual appearance and short-term context, limiting geometric grounding and episode-level reporting. We introduce StenoVLA-3D, a 3D-aware VLA framework for navigating through stenotic regions. We integrate point-maps into the Cosmos-Reason 2 backbone through learned geometry-gated fusion, and also propose a temporal state branch to model traversal progress. Our reasoning-and-action backbone predicts grounded reasoning with actions, while dedicated heads estimate stenosis shape and generate the final lesion report. We further introduce EndoCausal, an episode-level dataset with lesion annotations, actions, and temporally grounded reasoning. On 40 held-out recorded test episodes, StenoVLA-3D reaches 95.2\% semantic accuracy and 83.4\% action accuracy. On the physical 3-DoF endoscope, it attains 88.9\% and 77.8\% task success in esophageal and colonic phantoms (36 trials each), substantially outperforming the evaluated baselines.
A Topological Representation with Object-Path Graphs for Open-Vocabulary Instance Navigation
一种基于对象-路径图的开放式词汇实例导航拓扑表示方法
Zheng, Linwei, Peng, Daojie, Wang, Bingtao, Li, Haoang, Ma, Jun
Abstract
Vision-language navigation requires embodied agents to navigate environments using natural language instructions and visual observations. Existing approaches typically decompose navigation into sequential language-guided decisions or rely on online exploration without prior environmental knowledge. Scene graph representations offer compact semantic memory but remain decoupled from downstream navigation, which still depends on dense metric maps. To close this gap, we propose an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation. The proposed representation jointly supports semantic grounding, graph-based localization, and navigation within a single lightweight topological framework. Building on this graph, we introduce a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly over the graph without dense metric reconstruction. Experiments on HM3D and Replica demonstrate competitive performance in open-vocabulary object grounding through the proposed hierarchical graph structure, while achieving effective navigation performance. Real-world robot experiments further validate the practicality of the proposed framework.
Reliable perception is essential for underwater vehicles operating in complex environments, where light attenuation and scattering often degrade visibility and compromise optical sensing. Forward-looking sonar (FLS) offers an alternative by providing high-frame-rate acoustic imaging under poor optical conditions. However, real-time FLS mapping remains challenging due to unresolved target elevation, spatially non-uniform noise, and fragmented target boundaries, which hinder feature extraction and introduce geometric ambiguity during projection. To address these challenges, we propose a cascaded feature reconstruction pipeline combining fast Fourier transform (FFT)-based denoising, fast multiscale constant false alarm rate (MCFAR) detection, and gradient-adaptive boundary connection to extract geometric features from degraded sonar images with low latency. We integrate attitude-aware geometric projection with incremental occupancy accumulation to construct a depth-referenced 2.5D map for local mapping in confined underwater environments. The sonar's vertical position is referenced to an external sensor, while target elevation is assigned under an explicit geometric assumption rather than measured directly by FLS. Experiments in a 3 m X 5 m pool demonstrate centimeter-scale planar mapping accuracy, with a root-mean-square error (RMSE) below 3 cm across three sequences and an average processing time of 42.4 ms per frame.
Audio-based localization provides a low-cost and illumination-independent sensing solution for anti-UAV early warning. However, existing methods typically rely on a predefined fixed audio segment length, which limits temporal correspondence and creates a trade-off between sufficient acoustic evidence and timely localization. To address this issue, we propose an audio-based localization framework with adaptive temporal correspondence. A probe segment is first used to extract a compact acoustic state that characterizes the reliability and consistency of the observation. Guided by the state, a reinforcement learning controller dynamically determines the required audio window size for each localization decision. The selected audio segment is then processed by a Mamba-based localization network with adaptive temporal feature modulation for 3D position estimation. Extensive experiments demonstrate that our method achieves competitive 3D localization accuracy with substantially reduced temporal correspondence latency compared to SOTA methods and exhibits strong generalization across scenarios.
OpenFlyScan: A Quality-Guided Aerial Reconstruction System for Consumer Drones
OpenFlyScan:面向消费级无人机的质量引导式航拍重建系统
You, Zhongrui, Li, Zhen, Liu, Junli, Wang, Zhigang, Zhao, Bin
Abstract
3D Gaussian Splatting (3DGS) provides high-fidelity scenes for large-scale embodied simulation, but constructing large-scale urban assets remains constrained by expensive equipment and delayed quality feedback. Preset surveys can leave complex surfaces insufficiently observed, with defects discovered only after reconstruction, requiring return visits and repeated processing. We present OpenFlyScan, a quality-guided aerial reconstruction system for consumer drones that integrates a GS quality model, a reacquisition planner, and a custom-designed mobile app. The model learns from GS rendering errors to predict regional reconstruction quality. Based on these predictions, the planner then generates complementary reacquisition strips to be executed through the app, which also supports automated oblique surveys and data transfer without additional hardware on board. Across real aerial scenes, the model effectively identifies regions that are likely to be poorly reconstructed. In the Expo West field experiment, targeted reacquisition improves PSNR at additional views by 10.95 dB. With consumer drones, OpenFlyScan integrates capture, targeted reacquisition, and reconstruction to support rapid, low-cost urban asset creation. Code and models will be made publicly available at https://openflyscan.github.io/.
Current embodied systems largely rely on pretrained capabilities that remain fixed after deployment, limiting their ability to learn from physical interaction. We introduce MachEmbodied-Brain (ME-Brain), a self-evolving embodied system organized around a closed loop of action execution, experience acquisition, experience evolution, and improved execution. Evolvable Memory consolidates multimodal trajectories into hierarchical, reusable experience; Cognitive Core transforms physical experience into transferable skills; and the Action Model combines event-driven keyframes, EventCell local-world prediction, and action-conditioned memory modulation to focus computation on decision-critical moments, regions, and historical evidence. Together, these modules shift embodied intelligence from train-and-freeze to deploy-and-evolve without model retraining. Cognitive Core outperforms the strongest comparison models by 8.2 and 9.6 points on embodied and agent benchmarks. The Action Model achieves 47.88% mean success on RoboMME, a 3.26-point improvement over the strongest baseline. On RoboDojo, it reaches a 21.51 mean Score and 16.03% success rate, exceeding $\pi_{0.5}$ by 10.10 and 9.12 points. On the six-task ME-RealBench, ME-Brain achieves a 69.5 mean Score and 66.7% success rate, outperforming DM0.5 by 12.8 and 11.7 points, respectively.
vla.simd: Efficient CPU Inference for Language-Conditioned Manipulation
vla.simd:面向语言条件操控任务的高效 CPU 推理
Nguyen, Khanh D., Truong, Hoang M., Le, An T.
Abstract
Deploying language-conditioned manipulation without a dedicated GPU requires efficient inference and action chunks that cover the delay between policy queries. We present vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization. We relate query latency and execution horizon to action availability under lagged and time-aligned execution, distinguishing action supply from feedback frequency. Across six policies and four CPUs, vla.simd achieves approximately $1.4\times$ median speedup over compiled PyTorch references while preserving fp32 numerical fidelity. We also introduce IMPACT, an ACT-based policy with cached text representations and language-modulated visual features. IMPACT is the only language-conditioned policy in our evaluated set that supplies at least 30 actions/s on the Raspberry Pi 5: after a 90 s thermal soak, it supplies 33.5 actions/s in fp32 and 81.2 with int8. Separate GPU evaluations yield $76.4\%$ mean success across four LIBERO suites without robot pretraining; instruction-shuffling tests demonstrate selection among familiar goals. Trials with IMPACT on an SO-101 arm and SmolVLA on a UR10e with a Robotiq gripper demonstrate CPU deployment on two robot embodiments.
In social navigation, modeling the complex interactions between humans and robots is difficult, and deep reinforcement learning has therefore been actively studied. However, because simulation alone cannot fully reproduce diverse scenarios, robot dynamics, and the social conventions that vary across deployment environments, fine-tuning in the deployment environment is promising. In doing so, learning that preserves the base model's performance is required, so as not to compromise the primary objective of navigation, namely avoiding pedestrians and reaching the destination. In this study, we propose a method that applies diffusion steering via reinforcement learning (DSRL), which trains only the noise policy while keeping the diffusion policy fixed, thereby achieving learning that preserves performance. Furthermore, we integrate diffusion-based RL policies trained with multiple seeds to construct the base policy, improving learning performance. Our evaluation shows that, compared with other methods, the proposed method enables efficient learning while preserving performance, and we confirm flexible behavior control through adaptation to social conventions, as well as its effectiveness on a physical robot through hardware-in-the-loop simulation.
Robotic foundation models achieve impressive performance on standard manipulation benchmarks, yet these evaluations typically assume clean, timely, and consistent visual observations throughout execution. We introduce LIBERO-VPro, a benchmark for systematically evaluating the closed-loop visual robustness of robotic foundation models by perturbing the visual evidence available during execution. LIBERO-VPro covers four complementary dimensions, including Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, and Task-Relevant Scene Variation, spanning 12 challenge categories, 96 experimental settings, and 3,296 task-condition cases. We evaluate three vision-language-action models and three world-action models over approximately 196,000 simulated episodes, complemented by 200 real-world rollouts on a Franka Research 3. Our results reveal that strong nominal performance can mask substantial weaknesses in visual grounding and adaptation. Models often remain successful despite severe object-level occlusion, yet degrade sharply when local interaction cues are disrupted or familiar spatial priors are violated. They are also highly sensitive to stale or missing observations and struggle when changed task preconditions require behavioral adaptation. Finally, VLAs and WAMs exhibit distinct robustness profiles, showing that visual robustness is multi-dimensional and architecture-dependent. LIBERO-VPro provides a systematic diagnostic framework for developing robotic foundation models that can more reliably ground and adapt their actions under challenging visual conditions.
Chinese Translation
机器人基础模型在标准操作基准测试中取得了令人瞩目的性能,然而这些评估通常假设在整个执行过程中视觉观测是干净、及时且一致的。我们提出了LIBERO-VPro,这是一个通过扰动执行期间可用的视觉证据来系统性评估机器人基础模型闭环视觉鲁棒性的基准。LIBERO-VPro涵盖四个互补维度,包括视觉证据退化(Visual Evidence Degradation)、相机滞后(Camera Staleness)、视觉源一致性(Visual Source Consistency)和任务相关场景变化(Task-Relevant Scene Variation),共包含12个挑战类别、96种实验设置和3,296个任务条件案例。我们在约196,000次仿真回合中评估了三个视觉-语言-动作模型(VLA)和三个世界-动作模型(WAM),并在Franka Research 3机器人上补充了200次真实世界测试。我们的结果表明,强大的标称性能可能掩盖视觉接地和适应能力方面的重大缺陷。模型在严重的物体级遮挡下仍常能保持成功,但当局部交互线索被破坏或熟悉的空间先验被违背时,性能会急剧下降。模型对滞后或缺失的观测也高度敏感,并且在变化的任务前提条件需要行为适应时表现挣扎。最后,VLA和WAM表现出明显不同的鲁棒性特征,表明视觉鲁棒性是多维度且依赖于架构的。LIBERO-VPro为开发能够在具有挑战性的视觉条件下更可靠地进行视觉接地和动作适应的机器人基础模型提供了系统性的诊断框架。
Multi-Agent Transportation of Free-Flyers in Microgravity Via Pushing Interaction Under Human-in-the-Loop Control
人在回路控制下基于推挤交互的微重力自由飞行器多智能体运输
Marchesini, Gregorio, De Carli, Nicola, Cho, Sihyun, Kong, Youngkyoung, Krantz, Elias, Dhullipalla, Mani Hemanth, Dimarogonas, Dimos V., Kim, H. Jin
Abstract
We propose a safety-critical framework for the cooperative transportation of passive targets in microgravity, where a team of chaser robots acts through unilateral pushing contacts to track a human-provided desired twist while ensuring safe target motion. The pushing-only nature of the interaction introduces sparse, configuration-dependent actuation constraints requiring chasers to physically relocate on the target body when the desired pushing allocation changes. To address these challenges, we formulate a delay-aware feedback control architecture leveraging Control Lyapunov Function (CLF) and Control Barrier Function (CBF) constraints within a mixed-integer thrust allocation program to enforce stability and safety of the target, respectively. The proposed framework enables reference tracking while guaranteeing obstacle avoidance with a circular obstacle despite intermittent control authority, providing a foundation for human-supervised cooperative transportation of free-flyers in space environments. The proposed framework is validated through Gazebo simulations.
Tactile sensing is an essential modality for robots performing contact-rich, dexterous manipulation, particularly under visual occlusion. While pre-trained image encoders are standard in robot learning pipelines, tactile encoders are still commonly trained from scratch from raw, noisy signals, which might limit their expressivity. Existing self-supervised learning (SSL) approaches focus predominantly on vision-based tactile sensors, leaving distributed electronic skins largely unaddressed. These sensors, however, have a distinctive property: their sensing elements are sparse and irregularly arranged over the surface they cover, which makes direct reuse of visual SSL methods suboptimal. We present Tactile-JEPA, an efficient self-supervised pre-training method that uses the spatial arrangement of tactile sensors to learn topology-aware representations. Specifically, it is trained to predict the embeddings of masked sensing elements from the unmasked remainder, using the sensor connectivity graph to guide spatial masking. Our analysis shows that effective tactile representations require capturing both local contact details and the global state of the tactile surface, which we achieve through dual-scale masking. Across three diverse datasets spanning magnetic and piezoresistive sensors, different robot embodiments, and single- and paired-sensor configurations, Tactile-JEPA reduces force estimation error by 6.3% and in-hand orientation error by 20.8% over the prior state-of-the-art, with consistent gains in other downstream applications, including policy learning. Overall, our results demonstrate that the benefit of tactile sensing depends critically on the quality of encoder pre-training, a problem which Tactile-JEPA addresses directly. Code is available at https://github.com/E-Kovtun/tactile.
Egocentric video offers a scalable source of physical interaction experience, yet translating it into robot-executable knowledge and enabling continual adaptation remain challenging. We introduce Zeva-Ego, a unified framework that learns physical priors from human experience and evolves through robot interaction. An Action-Centric Encoder (ACE) converts egocentric visual transitions into action-centered supervision for VLA mid-training, while In-Context Causal Learning (ICCL) enables parameter-free adaptation from action-effect feedback at deployment. Scaling Ego data to 10K hours improves RoboTwin success from 63.8% to 75.3%, matching 2K hours of robot demonstrations (74.7%), corresponding to an empirical data ratio of roughly 4-5:1. With accumulated interaction experience, ICCL further improves success from 58% to 89% within four attempts without parameter updates. These results demonstrate a scalable path toward embodied intelligence that learns from human experience and continuously improves through its own interaction.
Robotic Valve Turning: Axial Misalignment Correction Using Reaction Torque Feedback
机器人阀门旋转:基于反作用力矩反馈的轴向偏差校正
Kumar, Amit, Turlapati, Sri Harsha, Golani, Gautami, Lin, Yang, Banavar, Ravi N., Campolo, Domenico
Abstract
In this work, we propose a haptic update control law that uses reaction torques to correct axial misalignment during robotic valve manipulation. Unlike vision-based estimates, which can be affected by calibration errors, occlusion, and uncertainty in the contact geometry, reaction torques arise directly from the physical interaction between the gripper and valve. A geometric relationship exists between the error (misalignment) vector and these torques. The primary aim of this work is to propose a stable controller exploiting this geometric property. Our control law is proven to be uniformly asymptotically stable. Simulations are performed for verification. Furthermore, we experimentally test the robustness of our method using a Kinova Gen3 robotic arm for initial misalignments ranging from $-15^\circ$ to $15^\circ$ at 3 different valve positions and report the resulting data distribution. The absolute value of the median misalignment across all 18 test cases is found to be within $2.46^\circ$ and that of reaction torques within $0.23\mathrm{Nm}$.
FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding
FoldQuantVLA:基于一致折叠的视觉-语言-动作模型原生低比特量化
Ho, Hung T., Nguyen, Khanh D., Nguyen, Quang D., Duong, Thanh Q., Le, Ngan, Guo, Meng, Ngo, Vien A., Le, An T.
Abstract
Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-training quantization framework that carries a consistent activation representation through calibration, weight rounding, and native integer execution. It combines channel scaling and block Hadamard transforms with dynamic per-token quantization, without policy retraining. Custom TensorRT plugins execute projections in both the language backbone and iterative action expert with four-bit weights and activations (W4A4) on Ada GPUs and Jetson AGX Orin. Evaluation spans LIBERO, SimplerEnv, and two robot platforms. Across three GR00T checkpoints and $\pi_{0.5}$, W4A4 achieves $1.20$ to $1.33\times$ speedups over floating-point TensorRT on Orin and $1.25$ to $1.52\times$ on desktop. Retaining language attention-output and feed-forward down projections at eight bits (W8A8) improves held-out action fidelity on all four checkpoints. Across four real-robot tasks, this configuration raises observed GR00T N1.7 success from $80.0\%$ with uniform W4A4 to $92.5\%$ over 80 trials per configuration, with a measured additional Orin latency of 1 ms.
Real-time motion planning remains challenging in large and high-dimensional environments. Prior acceleration of RRT* follows tree-centric state organization, which reduces per-query cost but preserves superlinear end-to-end complexity and limits parallelism through structural dependencies. This paper presents ScaleMPA, a motion-planning accelerator that rethinks RRT* with a grid-native representation. By replacing hierarchical traversal with direct grid-based access, ScaleMPA reduces the planner critical path and exposes fine-grained parallelism. To make this reformulation practical under sparse high-dimensional planning, ScaleMPA further proposes a multi-resolution grid search engine and a hash-grid memory system. Implemented in 28 nm CMOS, ScaleMPA achieves millisecond-level planning latency and delivers 4.7$\times$--44.4$\times$ speedup over state-of-the-art motion-planning accelerators.
Pneumatic proprioceptive actuators integrate actuation and sensing for soft robots that attract interest due to functional potential. Existing approaches often suffer from assembly errors or stress concentrations caused by heterogeneous materials. In this work, we propose the Monolithic Force-Proprioception Soft (MFPS) design and fabrication method that integrates an Asymmetric Origami Bending (AOB) chamber and a Force-Proprioception Soft (FPS) sensor with single material through one-step Fused Deposition Modeling (FDM) fabrication. Based on the resistance response to strain of conductive thermoplastic polyurethane (TPU), we design and analyze the structure of the FPS sensor, and conduct parametric analysis on the sensing characteristics. The FDM fabrication parameters of the MFPS actuator are analyzed, followed by actuator fabrication and characterization of the actuation and proprioception performance. Experimental results show that the MFPS actuator achieves a bending angle of 40{\deg}, an output force of 12.5 N, and a resistance change of 26.9% as the applied external force increased from 0 to 45 N. A two-finger force-proprioception gripper is developed based on the MFPS actuator. The grasping and force-proprioception capabilities are experimentally validated, proving that the MFPS design method provides a new approach for the development of self-sensing actuators.
Visuomotor policies trained from a few demonstrations may reproduce demonstrated trajectories without reliably following changes in object position. Existing approaches with explicit attention typically obtain spatial priors from human annotation or visual models. We introduce TACIT (tactile contact informs attention), which uses measured tactile contacts from teleoperated demonstrations to supervise spatial attention without additional point annotation. Gaussian targets over preceding camera point clouds supervise an attention head whose pooled output conditions a visuotactile diffusion policy. Targets are used only during training; tactile observations remain inputs at inference. In the primary real-robot benchmark, with ten demonstrations per task and five demonstrated placement regions, TACIT achieves 66.7% success on ball placement and 73.3% on peg insertion, compared with 10.0% and 20.0% for input-matched 3D visuotactile fusion and 20.0% and 43.3% for vision-only DP3. TACIT enters the 150 mm palm-to-object approach region within 12 seconds in all 30 trials per task; all remaining failures occur after arrival. Across three training seeds on real ball and simulated peg, TACIT outperforms input-matched fusion and an architecture-matched control without explicit attention supervision, supporting the contribution of supervision beyond branch capacity. Pre-contact and contact-time supervision show no consistent ordering. These results demonstrate that measured tactile contact provides effective spatial supervision for approach behavior from few demonstrations within the evaluated workspace.
Contact-rich precision insertion is a key manipulation skill in robotic assembly. Tight clearances make insertion more sensitive to alignment errors and prone to collisions and jamming, while variations in geometry and clearance across parts further complicate policy reuse. We present a reinforcement learning framework that trains insertion policies entirely in simulation for direct deployment without real-world demonstrations or policy fine-tuning. By combining target poses with compact three-dimensional fingertip force feedback, the policy learns to search for alignment and correct its motion despite errors in the estimated hole position. A decoupled gated reward coordinates alignment and insertion. Force-signal smoothing and state-independent standard deviations stabilize the learning process. The resulting policies perform real-world insertion across multiple hole geometries with a minimum nominal clearance of 0.02 mm and improve success while reducing peak contact forces under hole-position errors. Cross-clearance and cross-geometry evaluations further confirm policy generalization. The system achieved the first perfect score of 20/20 on ManipulationNet's peg-in-hole benchmark under its Human-in-the-Loop protocol, with fully autonomous insertion motions. A single policy trained only on a simulated hexagonal insertion task achieved an overall success rate of 95.0% across eight unseen real-world insertion tasks. These results show that learning entirely in simulation can yield precision insertion skills that can be deployed directly and reused across real-world tasks. The project website (https://mzhsoul.github.io/InsertAnything/) provides open-source simulation and real-robot experiment scripts, assets, and trained checkpoints.
Vision-Language-Action (VLA) models have demonstrated remarkable generalization in robotic manipulation via large-scale multimodal pretraining. However, VLA models are mainly trained on 2D-centric observations, which inherently constrains their capacity for precise spatial manipulation. Previous methods enhance 3D awareness by introducing implicit spatial priors, but still lack explicit geometry guidance. In this paper, we propose Bridge3D that integrates both implicit and explicit 3D geometry guidance into pre-trained 2D VLA models, enabling them to ''see'' and ''act'' in 3D. Bridge3D introduces two strategies: 1) Implicit Fusion, which enriches visual tokens with features from 3D foundation models to improve ''seeing'' in 3D; 2) Explicit Conditioning, which integrates action denoising with an explicit 3D semantic field to achieve ''acting'' in 3D. Furthermore, we utilize the proposed layer-wise linear probing to improve learning efficiency. Experiments show that Bridge3D achieves superior performance against state-of-the-art methods. On the RoboTwin 2.0 benchmark, Bridge3D exceeds $\pi_0$ by 14.0 percentage points, while in real-world experiments, it outperforms Spatial Forcing by 11.7 percentage points. These results demonstrate Bridge3D's strong capabilities in high-precision and spatial-sensitive manipulation tasks.
Chinese Translation
视觉-语言-动作(Vision-Language-Action,VLA)模型通过大规模多模态预训练在机器人操作任务中展现出卓越的泛化能力。然而,VLA模型主要以2D为中心的观测数据进行训练,这从根本上限制了其进行精确空间操作的能力。以往的方法通过引入隐式空间先验来增强3D感知能力,但仍缺乏显式的几何引导。本文提出Bridge3D,将隐式和显式的3D几何引导融合到预训练的2D VLA模型中,使其能够在3D中“感知”和“行动”。Bridge3D引入了两种策略:1)隐式融合(Implicit Fusion),通过融合来自3D基础模型的特征来丰富视觉token,从而提升3D“感知”能力;2)显式条件化(Explicit Conditioning),将动作去噪与显式的3D语义场相结合,实现3D“行动”。此外,我们还利用所提出的逐层线性探测(layer-wise linear probing)方法来提高学习效率。实验表明,Bridge3D相对于当前最先进的方法取得了更优的性能。在RoboTwin 2.0基准测试中,Bridge3D超过$\pi_0$模型14.0个百分点;在真实世界实验中,其性能优于Spatial Forcing 11.7个百分点。这些结果证明了Bridge3D在高精度和空间敏感操作任务中的强大能力。
Estimation and Control of Tensegrity Manipulator Kinematics based on Strut Inclination Angles
基于撑杆倾角的张拉整体机械臂运动学估计与控制
Bhat, Tufail Ahmad, Ikemoto, Shuhei
Abstract
Unlike conventional rigid-link robots defined by discrete joints, continuum robots pose a fundamental challenge for expressing their complex continuous bending configurations for closed-loop control. Several modelling approaches have been proposed for conventional continuum robots, but tensegrity-based continuum robots remain largely open. Moreover, many of these approaches assume a continuous elastic backbone and are therefore not directly applicable to tensegrity manipulators, whose bodies are networks of rigid struts and tensioned cables. This work presents a reduced-order model for shape and posture control of a tensegrity-based continuum manipulator. The manipulator is modelled as a serially connected parallel-link mechanism. The proposed method is formulated as an optimization problem that uses geometric constraints of the tensegrity structure together with information from the Inertial Measurement Unit (IMU) sensors embedded in the strut elements. To the best of our knowledge, this work presents the first experimental demonstration of a real-time IMU-based shape estimation method on a full-scale tensegrity manipulator and demonstrates posture control using a simple Proportional-Integral (PI) controller. The results show that the proposed method can estimate the shape of both single-module tensegrity structures and multi-module tensegrity manipulators from arbitrary static configurations and achieve desired postures.
Real-time embodied companion interaction requires a robot to infer user intent from streaming speech, generate timely responses, and execute expressive, interruptible motions. Existing systems typically decouple dialogue orchestration from gesture synthesis, relying on offline motion generation from complete audio. This separation leaves open how a deployed robot can dynamically synchronize response content, prosodic timing, and physical safety under incremental inputs and uncertain turn boundaries. We present MIRA, a unified framework for full-duplex embodied companion interaction. Given streaming user speech, dialogue history, and vocal affect, MIRA predicts both the response text and an explicit embodiment cue. Discrete social behaviors (\eg listening, greeting) are mapped to validated robot trajectories, while open-ended speaking is paired with streaming, co-speech motion. This generative motion is governed by a predict-more-than-commit sliding window that provides temporal look-ahead for motion continuity while limiting physical commitment to a short, cancellable prefix. Crucially, we design CORTEX, a dual-timescale interaction policy that manages low-latency streaming and deliberative turn decisions, backed by a robot-side execution layer that enforces physical safety constraints at the control rate. We deploy MIRA on an Astribot S1 humanoid robot. Quantitative evaluations demonstrate competitive audio-motion alignment relative to state-of-the-art motion-generation baselines, while real-robot deployment measurements characterize streaming responsiveness and interruption handling.
Embodied AI systems, particularly humanoid robots deployed in real world scenarios require whole-body control policies that are both task-responsive and physically smooth. However, smoothness is not uniform across the body: lower body must remain sufficiently reactive, while the upper body must be tightly regulated to preserve stability. Existing reinforcement learning approaches typically impose smoothness through auxiliary terms in the reward function, which compete with task objectives, treating the body as uniform and provide no direct control over the physical quantities responsible for smooth behavior. We introduce DeCap (Decoupled Constraint-aware policy), a constrained reinforcement learning algorithm that decouples whole-body smoothness into separate upper- and lower-body constraint groups, each formulates smoothness as explicit constraints on physical motion limits. To improve constraint satisfaction near feasibility boundaries, DeCap incorporates a bounded barrier penalty that activates proactively as limits are approached and remains bounded at the constraint limit. On real-world humanoid whole-body control task, DeCap reduces upper-body action rate by 2.50x and acceleration by 2.18x relative to reward-based smoothness policies, while also improving lower-body smoothness and reducing transient motion. We demonstrate that a fixed set of smoothness constraints transfers across diverse terrains, alleviating the need of extensive reward tuning.
ARSTAG: An Agentic Real2Sim2Real System for Task-Specific Robot Data Generation
ARSTAG:一种面向任务特定机器人数据生成的智能体化Real2Sim2Real系统
Li, Bowei, Zhang, Yuner, Liu, Changliu
Abstract
Adapting visuomotor policies to new manipulation tasks often requires substantial manual engineering or teleoperated data collection. Simulation can provide task-specific data at scale, but constructing the scene, designing expert behavior, and configuring data generation still require significant per-task effort. We present ARSTAG, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data. A hierarchy of language agents constructs a task-scoped simulation scene, generates robot-feasible demonstrations, and expands the training distribution through task-consistent randomization, while a coordinator agent manages cross-stage feedback and recovery. Across seven manipulation tasks spanning grasping, placement, and stacking, the ARSTAG-generated demonstrations enable sim-to-real transfer of three visuomotor policy architectures to a dual-arm robot, with pi0.5 achieving an average real-world success rate of 74.6%. Ablations show that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size. Project webpage: https://boweili666.github.io/ARSTAG/.
Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While this fusion of language instructions and visual observations allows multimodal reasoning, it obscures how information is routed across modalities or what mechanisms drive navigation decisions. Thus, it remains unclear whether VLN models ground their predictions in relevant semantic cues or can track task progress. In this work, we study the interpretability and steerability of VLN models. We use intervention-based metrics that measure how visual observations, instructions, and visual memory causally influence navigation decisions. Our results show that these navigation policies are sensitive to all input modalities and do not depend on a single one. We further show that these agents encode navigation progress and retain semantic structure from their VLM backbones, enabling concept-level steering through internal activations. Finally, we extract activation vectors for abstract behaviors to transfer them zero-shot to out-of-distribution real-world scenarios, improving performance without additional fine-tuning.
Learning tactile perception from high-bandwidth single-point sensing
从高带宽单点感知中学习触觉感知
Rigal, Joseph, Virot, Emmanuel, Pascal, Caroline
Abstract
Tactile sensing is increasingly being incorporated into learning-based robotic manipulation, yet many existing approaches rely on spatially distributed sensors. Here we introduce SpectRobot, a framework that transforms single-point tactile signals into compact time-frequency spectrograms. These spectrograms encode high-bandwidth tactile histories as fixed-size image-like representations. They can be processed by standard vision encoders and integrated into learning pipelines originally developed for vision, while preserving temporal and frequency information unavailable to conventional cameras. Rather than increasing spatial density through arrays of tactile elements, SpectRobot exploits the rich dynamics contained in sparse, high-bandwidth single-point measurements. In our implementation, the sensors are mounted away from the contact surface while remaining mechanically coupled to it, reducing direct exposure to wear and potentially improving robustness in harsh environments and for long-term deployment on dexterous robots. Our experiments demonstrate that: (1) a robot can exploit single-point vibration signals to solve a visually occluded manipulation task; (2) temporal history strongly influences policy performance, while sensing bandwidth controls the spectral information available, with measurements extending to 100~kHz; and (3) the same representation can be used across different tactile sensing technologies mediated by acceleration, force, or strain. We further show that capabilities previously associated with research-grade instrumentation can be accessed using readily available, off-the-shelf hardware. We believe that broader access to high-bandwidth tactile sensing could facilitate the integration of contact dynamics into embodied learning systems and, for some tasks, offer an alternative or complement to increasing the spatial density of tactile sensing.
From Semantic Decisions to Feasible Trajectories: Self-Evolving LLM-Guided Optimal Control for Narrow-Space Parking
从语义决策到可行轨迹:面向窄空间泊车的自演化LLM引导最优控制
Yao, Zhengbao, Luo, Yuanfu, Xue, Kehan
Abstract
Autonomous parking in nonconvex and narrow environments remains challenging. Although optimal-control methods can explicitly enforce vehicle dynamics and collision constraints, nonconvexity compromises solver robustness and can cause failures. Large language models (LLMs) exhibit strong semantic reasoning capabilities, but directly generating dense trajectories makes it difficult to guarantee physical feasibility. We introduce SE-LLM-OCP, a unified framework in which LLMs make high-level discrete maneuver decisions, while an optimal-control module enforces low-level vehicle dynamics and collision constraints. Online, the LLM proposes sparse maneuver plans, decomposing the parking task into a sequence of short-horizon trajectory-optimization problems. A low-level solver then sequentially solves optimal-control problems. If the solver fails, the LLM aggregates failure evidence from the solver and validation stages to guide replanning. Offline, SE-LLM-OCP automatically evolves a structured decision-making knowledge base from scratch, driven by accumulated online failures. We validate our proposed framework in simulation on a car-like vehicle model and on a differential-drive robot. Our experimental results show that SE-LLM-OCP enables safer autonomous parking in narrow scenarios and demonstrates transfer of the same maneuver representation to a different kinematic platform.
Feasibility Distance Fields for Heterogeneous Constraints in Robot Configuration Space
面向机器人构型空间中异构约束的可行性距离场
Cui, Xijing, Pu, Huayan, Luo, Jun, Wang, Gang
Abstract
Robot manipulators are monitored by constraint-specific indicators whose units and gradient scales are not comparable, so they do not provide a common measure of the configuration-space motion remaining before violation. We define the feasibility distance field (FDF) as the distance, under a fixed positive-definite joint-space metric, to the union of infeasible configuration sets. Classical distance-to-set theory gives 1-Lipschitz continuity, almost-everywhere differentiability, and unit dual-gradient norm wherever the nearest projection is unique. The robotics contribution is an admissibility analysis showing when practical constraints define non-empty closed sets. We derive admissible formulations for external and self-collision, joint limits, dexterity, Cartesian and task-projected compliance, joint torque under payload, and dynamic manipulability. Since every field uses the same metric, heterogeneous constraints compose by a pointwise minimum, conditioned constraints retain a fixed distance space, and multi-robot constraints produce block-sparse gradients that identify which robots must react. We generate projection-based labels and train neural approximations with a distance loss and an Eikonal penalty. Simulations on a UR5e and a dual-arm cell evaluate seven fields using value, projection, sign, gradient, composition, and moving-obstacle diagnostics. Across 8,000 configurations, the largest feasible-side secant ratio is 0.920, mean learned gradient norms range from 0.994 to 0.998, and projection residuals range from 0.011 to 0.034 rad. Across 24 random obstacle paths, the external and composed collision fields achieve 90.4% and 91.6% success within 3 cm, with sign-error rates below 2%. The results support a common configuration-space margin and identify approximation errors near medial axes and sparsely sampled boundaries.
Human demonstrations offer a scalable way to collect manipulation data, but their contacts may be unstable or infeasible when transferred to a robot hand. Collecting demonstrations directly on the target robot avoids this mismatch, but substantially increases the cost of data collection. To address this trade-off, we present \textbf{Touch2Robot}, a framework that lets humans collect demonstrations while seeing how the target robot hand would contact the object. We capture human hand motion, tactile-glove measurements, and object motion during human manipulation. These recordings guide object-specific RL policies to reproduce the demonstrated object motion while favoring contacts consistent with the recorded human touch. We distill the learned behaviors into a unified real-time retargeter that maps incoming human observations and object geometry to robot hand configurations. During collection, the predicted robot configuration is synchronized with the tracked object pose in simulation to reconstruct robot-object contacts, which are visualized to help the demonstrator adapt subsequent interactions to the target hand. Across four real-world tasks, Touch2Robot improves average real-robot replay completion from 37.9\% to 72.1\% over visual-only feedback, while reducing the collection time per replay-successful demonstration from 58.6~s to 18.2~s. Reconstructed target-hand contacts achieve 44.2\% F1 against real-robot tactile measurements, and policies trained on Touch2Robot demonstrations improve downstream Diffusion Policy performance by 29.1 percentage points over visual-only feedback. These results show that bringing robot touch into the human demonstration loop improves both the quality and efficiency of scalable dexterous data collection. \textit{Project webpage: \href{https://Touch2Robot.github.io/}{https://Touch2Robot.github.io/}.}
Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies
像世界模型一样思考,像VLA一样行动:将世界模型表示蒸馏进紧凑的机器人策略
Dao, Trung, Yamsani, Sankalp, Park, Jaden, Kim, Joohyung, Lee, Yong Jae
Abstract
Vision-Language-Action (VLA) models map observations to actions with no objective that accounts for how the world responds, so their robustness is bounded primarily by data coverage. World models carry precisely that missing objective and are better grounded for it, yet rolling the future forward costs seconds per decision and rules them out of the control loop. We show the two can be separated. What a world model knows about physical scenes lives in its \emph{internal features}; generating the future is merely the objective that produced them, so the grounding can be inherited while the generative machinery is left behind. We add one feature-alignment term to ordinary VLA training: a frozen world model is run over the training frames once and cached, and the student learns to agree with that cache. No teacher is loaded during training, the projector is discarded after it, and the deployed policy is identical to the undistilled baseline, running in $32$~ms and $1.86$~GB on a consumer RTX~5090, so every gain is attributable to the representation rather than to added capacity or test-time compute. A $0.8$B student reaches $97.9\%$ on LIBERO, improves from $48.2\%$ to $50.5\%$ on RoboCasa-GR1 humanoid manipulation, and the same objective carries over to real hardware, on both a single-arm and a bimanual platform. The gain survives changes of student scale, backbone, alignment layer, and teacher, indicating a broad representational prior rather than a fragile alignment between two particular networks. Project page: https://thaw-vla.trung-dt.com/.
AC-DC: Adaptive Communication for Scalable Dynamic Average Consensus in Multi-Robot Ergodic Search
AC-DC:面向多机器人遍历搜索中可扩展动态平均共识的自适应通信
Kee, Robin Inho, Cannataro, Begum, Tzoumas, Vasileios
Abstract
We study scalable peer-to-peer dynamic average consensus (DC) for multi-robot systems under finite-range, finite-rate, and interference-constrained communication. We introduce Adaptive Communication for Dynamic Average Consensus (AC-DC), which jointly adapts Who communicates with whom, When, and over What parts of the consensus state, using local inputs and successfully received neighbor information. Each robot's consensus state estimates the current average of the robots' local inputs. AC-DC updates these estimates as local inputs change and averages the values exchanged between robot pairs. In AC-DC, robot pairs update without waiting for every robot to complete a communication round, and the selected-state messages carry consensus state coordinates independent of team size for a fixed state representation. We apply AC-DC to dynamic-priority multi-robot ergodic search: one consensus stream estimates team visitation for motion coordination, while the other fuses regional measurement information to update uncertainty maps and search targets. Across twelve settings with up to 80 robots and 20 paired trials per setting, AC-DC has the lowest mean (i) normalized covariance-trace area under the curve (AUC) and (ii) attempted modeled communication payload among the compared decentralized methods. Averaged across settings, AC-DC achieves paired AUC reductions of 27.5% relative to state-of-the-art baselines, with 8.7x less communication traffic. As the number of robots increases, we observe that AC-DC's communication payload approaches that of the ideal centralized baseline (one ground compute-station communicating directly with all robots): with 120 robots in a fixed 600 x 600 m scaling test, AC-DC uses 19.3 MB versus 19.2 MB for the ideal centralized baseline, while remaining peer-to-peer.
Chinese Translation
我们研究了在有限范围、有限速率和受干扰约束的通信条件下,面向多机器人系统的可扩展点对点动态平均共识(Dynamic Average Consensus, DC)。我们提出了动态平均共识的自适应通信方法(Adaptive Communication for Dynamic Average Consensus, AC-DC),该方法利用本地输入和成功接收到的邻居信息,联合自适应地调整:与谁通信(Who)、何时通信(When)以及通信共识状态的哪些部分(What)。每个机器人的共识状态估计各机器人本地输入的当前平均值。AC-DC 在本地输入变化时更新这些估计,并对机器人对之间交换的值进行平均。在 AC-DC 中,机器人对无需等待所有机器人完成一轮通信即可进行更新,且在固定的状态表示下,所选状态消息携带的共识状态坐标数与团队规模无关。我们将 AC-DC 应用于动态优先级的多机器人遍历搜索(ergodic search):其中一条共识流估计团队的访问分布以实现运动协调,另一条融合区域测量信息以更新不确定性地图和搜索目标。在多达 80 个机器人、每个设置 20 组配对试验的十二种设置中,在所比较的分散式方法中,AC-DC 在以下两方面的均值最低:(i)归一化协方差迹的曲线下面积(AUC),(ii)尝试的建模通信负载。在各设置上取平均后,AC-DC 相对于最先进的基线方法实现了 27.5% 的配对 AUC 降低,通信流量减少 8.7 倍。随着机器人数量的增加,我们观察到 AC-DC 的通信负载趋近于理想的集中式基线(一个地面计算站与所有机器人直接通信):在固定的 600 x 600 米规模测试中,使用 120 个机器人时,AC-DC 的通信量为 19.3 MB,而理想集中式基线为 19.2 MB,同时 AC-DC 仍保持点对点通信方式。
SPARSER: Sparse Variable Projection by Exploiting Separable Structure in Robotic Perception
SPARSER:利用机器人感知中可分离结构的稀疏变量投影方法
Sanderson, Nikolas R., Fishberg, Andrew, Han, Haoyu, Yang, Heng, How, Jonathan P., Singh, Hanumant, Everett, Michael, Papalia, Alan
Abstract
Robotic perception often requires solving large nonlinear least-squares (NLS) problems. While sparsity has been widely exploited to scale solvers, a complementary and underused structure is \emph{separability}: some variables, such as visual landmarks, appear linearly in the residuals and admit a closed-form solution once the remaining variables, such as poses, are fixed. Variable projection (VarPro) exploits this structure by analytically eliminating the linear variables, yielding a reduced problem with favorable computational properties. However, its use in robotic perception has been limited by gauge symmetries, such as invariance to global translations and rotations, which introduce challenges for standard VarPro methods. We present SPARSER (\textbf{S}parsity \textbf{P}reserving \textbf{A}nalytic \textbf{R}eduction for \textbf{S}eparable \textbf{R}obotic \textbf{P}erception), a VarPro framework for gauge-symmetric problems that jointly exploits separability and sparsity. Our method constructs a \emph{matrix-free Schur complement operator} for efficient evaluation of reduced costs, gradients, and Hessian-vector products, enabling integration with iterative NLS solvers. We characterize the applicable problem class, identify common cases admitting further analytical simplifications, and show that IRLS-based robust costs preserve most of the exploitable structure. Across synthetic and real SLAM, SNL, and SfM benchmarks, SPARSER is on average $5\times$--$7\times$ faster than state-of-the-art baselines on CPU and GPU, with gains exceeding $40\times$ on individual datasets. On outlier-corrupted multi-robot SLAM data, the robust variant is $2\times$--$16\times$ faster than a state-of-the-art GNC solver. We release open-source C++ code and all datasets.
LLM-based Conversational AI Knowledge Assistant for MyBuddy Humanoid Robot
面向MyBuddy人形机器人的基于大语言模型(LLM)的对话式AI知识助手
Chen, Hanxiao
Abstract
Humanoid robots are increasingly being popular and developed for human-centered applications, yet their ability to provide intelligent conversations and natural interactive knowledge assistance remains constrained by traditional rule-based dialogue systems, pre-defined responses and limited knowledge repositories. Large language models (LLMs) have emerged as a powerful foundation for enabling natural, adaptive, and context-aware Human-Robot Interaction (HRI), which provides a significant opportunity to address such limitations by enabling robots to understand natural speech language, reason over complicated queries, maintain high-quality conversational context, and generate knowledge-rich responses. In this work, we originally present and implement an LLM-based versatile Conversational AI Knowledge Assistant for the Raspberry-Pi-powered 13-Axis MyBuddy humanoid robot, which integrates LLM-driven language understanding and AI reasoning with real-time speech recognition, knowledge retrieval via extensible access of internet engines (e.g., Wikipedia, arXiv), flexible dialogue management, and natural speech synthesis to enable much more intelligent multi-turn continuous conversations and advanced emotional-support Human-Robot Interaction.
Chinese Translation
人形机器人在以人为中心的应用中日益普及并快速发展,然而其在提供智能对话和自然交互式知识辅助方面的能力,仍然受限于传统的基于规则的对话系统、预定义的响应以及有限的知识库。大语言模型(Large Language Models, LLMs)已成为实现自然、自适应和情境感知的人机交互(Human-Robot Interaction, HRI)的强大基础,为解决上述局限提供了重要机遇,使机器人能够理解自然语言、对复杂问题进行推理、维护高质量的对话上下文并生成知识丰富的回复。在本工作中,我们原创性地提出并实现了一个基于大语言模型的多功能对话式AI知识助手,应用于由树莓派(Raspberry Pi)驱动、具有13个自由度的MyBuddy人形机器人。该系统将大语言模型驱动的语言理解与AI推理,同实时语音识别、通过可扩展的互联网引擎访问(如Wikipedia、arXiv)实现的知识检索、灵活的对话管理以及自然语音合成相集成,从而实现更加智能的多轮连续对话以及先进情感支持型人机交互。
World action models generate actions together with visual predictions of their consequences. These paired outputs create the potential for planning by sampling multiple actions from one state, comparing their imagined outcomes, and choosing the action with the most promising predicted outcome. However, how to use imagined futures to guide action selection remains unclear. We examine this planning potential empirically. First, we estimate an oracle upper bound on selection by choosing the sampled candidate whose realised outcome is best. In a controlled same-state analysis, this choice raises success from 68.9% under uniform random selection to 79.2%. We then test selectors based on visual quality, physical consistency, and task progression as controlled interventions. Some tested selectors yield higher observed success, but the gains are uneven and the matched selectors leave much of the measured opportunity unrecovered. To investigate this gap, we examine whether sampled actions lead to different outcomes, whether these differences are visible in the predictions, and whether a score recognises them. Counterfactual branching from the same states shows that selection opportunity is concentrated in relatively few decisions in the initial candidate sets. Action spread and outcome coverage need not increase together. In a further evaluation across trajectory phases with complete action execution, the tested scores again recover little of the available improvement despite a small gain from learned value. These findings distinguish producing consequential action choices from recognising them in generated futures, motivating the evaluation of WAM predictions through their usefulness for decisions rather than visual quality alone.
Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully. We identify a decision-local prediction gap: among the few futures competing for execution, a candidate predicted closer to the goal can produce a worse realized outcome than an available alternative. We introduce D-JEPA, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes. A bounded, permutation-equivariant operator jointly reasons over goal-relative predictive features and ordinal evidence, refining pretrained predictive geometry where action choices are most consequential. Restricted predictor adaptation and a shared ordinal interface extend this alignment across complementary predictive geometries. D-JEPA further realizes the learned decision structure in JEPA-compatible future representations, enabling deployment through native latent-distance planning. Evaluations across latent control, manipulation, pretrained action-producing models, physical robots and autonomous driving demonstrate improved action selection, including 87.89% success on PushT, a 15.04-point average gain on RoboTwin, and a 17-point gain on physical robot tasks. These results establish decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.
Aerial manipulators represent the forefront of aerial robotics. Although potentially capable of complex interaction tasks, controlling aerial manipulators throughout the dynamic transitions occurring during task execution presents significant challenges. Abrupt or discontinuous changes in system dynamics generated by the transitions suggest the use of a switched approach, yet the available aerial manipulation methods are not designed for coping with switched regimes. In addition, most available methods fall short in coping with the tight couplings between the aerial vehicle and the manipulator, as well as in coping with the state-dependent uncertainties arising from the difficulty in modeling such couplings. We propose a switched-based adaptive control framework for aerial manipulators not relying on a priori knowledge of the vehicle-manipulator couplings and of state-dependent uncertainties. To guarantee stable manipulation despite changes in system dynamics, the framework provides a class of switching signals characterizing those transition phases for which the system is guaranteed to remain stable. Comparative experiments further validate the effectiveness of the proposed switched-based framework over the state of the art.
H2RBench: A Real-to-Sim Benchmark for Evaluating Human-to-Robot Transfer
H2RBench:一个用于评估人到机器人迁移的真实到仿真基准
Xiao, Chuyang, Zhan, Haotian, Krishna, Sriram, Meng, Peilin, Irshad, Muhammad Zubair, Zakharov, Sergey, Held, David
Abstract
Learning robot manipulation policies from human video demonstrations constitutes a promising avenue for scalable robot learning. However, comparing different human-to-robot (H2R) transfer methods remains challenging, as existing approaches are evaluated under different settings, including differing task suites, scene layouts, object instances, and amounts of robot supervision. To address this challenge, we present H2RBench, a Real2Sim benchmark for evaluating H2R transfer methods. H2RBench provides a standardized protocol built on real human video demonstrations and simulated robot demonstrations, and includes four manipulation tasks spanning diverse interaction requirements. We evaluate multiple representative H2R transfer methods, each adopting a different strategy for bridging the embodiment gap. Using H2RBench, we systematically characterize how each method scales with the amount of human demonstrations, revealing that methods differ substantially in their ability to leverage additional human data. We further show that simulation performance is broadly predictive of real-world robot performance, with an overall Pearson correlation of r = 0.89, Spearman correlation of \r{ho} = 0.85 and Mean Maximum Rank Violation (MMRV) of 0.06 across method-task configurations. These results establish H2RBench as a practical and scalable benchmark for comparative H2R evaluation prior to real-world deployment.
Scalable simulation is essential for robot data generation, policy training, evaluation, and safe iteration, yet real-world interaction is costly and conventional simulators require labor-intensive construction. We present Uranus, a data-driven robot simulator built around a joint-trajectory-conditioned autoregressive diffusion model. Uranus offers three key capabilities: (1) streaming, open-ended rollout, which receives future joint-position trajectories online and autoregressively generates one latent frame per step, corresponding to four RGB frames, without a fixed horizon; (2) low-latency generation, achieving 24 FPS after inference optimization; and (3) scalable, extensible robot control, providing a unified interface for synchronized multi-view generation across diverse robot embodiments and camera configurations. We conduct comprehensive quantitative and qualitative evaluations on both in-distribution and out-of-distribution data, providing an objective assessment of Uranus and clearly identifying its current limitations. We release the code and model weights to empower the community with practical tools and insights.
Minimum Time Trajectories for a Car-Like Mobile Robot Moving with Rigid Wheels Under Non-Sliding Constraints
满足非滑动约束的刚性轮式类汽车移动机器人最短时间轨迹
Ben-Asher, Joseph, Rimon, Elon, Ravina, Leeor
Abstract
This paper studies the minimum time trajectoriesvof a car-like mobile robot navigating in an obstacle free environment. The robot, with forward and backward speeds, is controlled by bounded front-wheels acceleration and limited front-wheels steering rate. The paper extends previous results which solved this problem for the kinematic car-like robot. However, the kinematic model assumes pure rolling at the wheels ground contacts. This assumption requires non-sliding constraints for the front and rear wheels that can only be handled by the robot dynamics. This paper formulates the non-sliding constraints based on the robot dynamics then augments the kinematic model time-optimal path primitives with three new path primitives associated with the non-sliding constraints. The three non-sliding path primitives together with the kinematic model twelve path primitives form the car-like robot time optimal trajectories. Approximate analytic solutions for the non-sliding path primitives are also provided. Examples study the time-optimal path primitives along representative maneuvers, illustrating how the non-sliding constraints influence the time optimal trajectories of the car-like robot.
Diffusion models offer flexible motion generation, but translating this flexibility into feedback-responsive humanoid control remains challenging. Hierarchical systems steer motion through references that may exceed a separate tracker's capabilities, leaving recovery and physical execution largely to the tracker. Action-only diffusion generates actions directly but lacks an explicit future-state trajectory for test-time motion objectives. Joint state-action diffusion provides this representation, yet representative controllers often depend on privileged full-body states, and support for learned behavior selection and test-time motion steering remains fragmented. We present PredActor, a predictive action diffusion policy that brings these complementary steering capabilities into one directly executed policy using proprioceptive observations. Conditioned on proprioceptive history and optional task context, PredActor jointly generates executable actions and an internal future-state trajectory. Classifier-free guidance strengthens text-conditioned behavior, while classifier guidance steers predicted states toward test-time objectives. Only actions are executed, without a separate motion-reference tracker or externally estimated full-body states as policy inputs. In simulation, PredActor reaches all 15 destination targets and achieves a text retrieval score of 0.580, compared with 0.373 for conditional action diffusion, with similar observed disturbance survival. To make this guided policy practical onboard, rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, both below the 20 ms control period. We deploy PredActor on a Unitree G1; evaluations across simulation and physical hardware demonstrate text-conditioned motion, disturbance response, joystick control, and semantic interpolation.
Chinese Translation
扩散模型能够提供灵活的运动生成能力,但如何将这种灵活性转化为具备反馈响应能力的人形机器人控制仍然充满挑战。分层系统通过参考信号来引导运动,而这些参考信号可能超出独立跟踪器的执行能力,使得运动恢复和物理执行在很大程度上依赖于跟踪器本身。纯动作扩散模型可以直接生成动作,但缺乏用于测试时运动目标的显式未来状态轨迹。状态-动作联合扩散提供了这种表示形式,但代表性控制器往往依赖特权级的全身状态信息,并且对学习到的行为选择和测试时运动引导的支持仍较为零散。我们提出 PredActor,一种预测性动作扩散策略,它仅使用本体感知观测,将这些互补的引导能力整合到一个可直接执行的策略中。PredActor 以本体感知历史和可选的任务上下文为条件,联合生成可执行动作和内部未来状态轨迹。无分类器引导(classifier-free guidance)增强了文本条件下的行为表现,而分类器引导(classifier guidance)则将预测状态引导至测试时目标。策略仅执行动作,无需独立的运动参考跟踪器,也无需将外部估计的全身状态作为策略输入。在仿真中,PredActor 能够到达全部 15 个目标点,文本检索得分为 0.580,而条件动作扩散仅为 0.373,两者的扰动生存能力相近。为使该引导策略能够在机载环境中实际运行,滚动去噪和保持计算量的运行时优化将完整的回调时间在 Jetson Orin NX 上降低至中位数 16.790 ms 和 p95 19.383 ms,均低于 20 ms 的控制周期。我们将 PredActor 部署在 Unitree G1 人形机器人上;跨仿真与物理硬件的评估展示了文本条件运动、扰动响应、摇杆控制以及语义插值能力。
CAST: Collision-Aware Assembly with Construction Robots using Simultaneous Trajectory Estimation and Planning
CAST:基于同步轨迹估计与规划的碰撞感知建筑机器人装配方法
Shaji, Karthik, Kim, Chisung, D'Amato, John, Bruun, Edvard, Dellaert, Frank
Abstract
Multi-robot systems have shown increasing viability in construction due to their ability to execute high-precision actions while reducing human exposure to hazardous tasks. However, these environments have high-dimensional configuration spaces and possess substantial collision-avoidance constraints, which include other robots, assembly objects, and workspace boundaries. We utilize a single factor graph for trajectory estimation and planning that incorporates measured robot states together with explicit collision and learned cable constraints. This supports changing workspaces and enables synchronized, high-dimensional robot motion planning while accounting for the stiff, vibration-induced uncertainty of heavy robotic systems. We demonstrate the success of our framework on the construction of a post-and-lintel structure using one robot arm as a timber gripper, and a second robot as a nail-fastener.
Range-Aided SLAM Initialization Exploiting Accurate Heading Information
利用精确航向信息的测距辅助SLAM初始化方法
Lougheed, Isabel, Forbes, James Richard
Abstract
This paper presents a novel initialization method for range-aided simultaneous localization and mapping (RA-SLAM). The general SLAM problem has a well-known separable structure where landmark and robot positions can be solved for in a linear fashion given known robot headings. This paper considers the case where highly accurate heading information is available, which is typical of autonomous underwater vehicle (AUV) navigation, to solve for the range transponder and robot positions. The proposed approach consists of two steps. First, using a generalized trust region subproblem (GTRS), the positions of the transponders relative to the AUV are solved for given the nonlinear range measurements. Second, the relative transponder positions and the known heading of the AUV are used to estimate the transponder and AUV positions by solving a linear least-squares problem. These transponder and AUV position estimates, combined with the highly accurate heading information, provide a reliable initialization method for the general nonlinear RA-SLAM problem. The effectiveness of the proposed approach is tested on a real-world AUV dataset where long baseline (LBL) range measurements are provided in concert with highly accurate heading information provided by an inertial navigation system (INS).
Reaching a 6-DoF grasp pose in clutter requires a collision-free trajectory, conventionally obtained by reconstructing the scene in 3D and planning inside that reconstruction, at the cost of its accuracy and compute. Potential fields learned directly from images remove that dependency but inherit the classical weakness of artificial potential fields: where attractive and repulsive gradients cancel, the descent grazes the obstacle instead of going around it, and can stall short of the goal. We present an SE(3) neural potential field learned from posed RGB images and supervised with a navigation function, the geodesic distance to the grasp through free space recovered from those same images during training, which removes both failures. On two tabletop scenes, from obstacle-blocked starts executed on a UR10, the field converges within 3 cm of the grasp from every start and every path it executes is collision-free against the ground-truth geometry, against 25% and 0% under image supervision alone; mean clearance rises from under a centimeter to 8.6-8.8 cm and arm-link contacts fall from 20.6-50.4% to 2.7-5.5% of executed configurations. Executed grasp success is 90.0% and 40.0% on the two scenes, the residual failures being refusals of the Cartesian executor rather than of the field. Planning takes about 2 s against 67-133 s for RRT* on a reconstruction of the same images, though under a common offline harness the two are comparable: the deployed margin is the cost of collision-checking a dense reconstruction, not planner complexity.
World Action Models (WAMs) jointly generate robot actions and predict future world states, transferring priors from video pretraining to robot control. However, future visual prediction is computationally expensive, so existing WAMs often rely on long action chunks to amortize inference cost across control steps, at the cost of closed-loop responsiveness. We present \method, a dual-system WAM that preserves broader-horizon world-action generation while enabling high-frequency closed-loop action updates by decoupling global planning and local refinement. \systwo periodically performs high-noise bidirectional denoising over a broader world-action chunk to establish a global plan, while wrist-only \sysone extracts a temporally aligned short window from the intermediate denoising state and completes low-noise refinement using the latest wrist observations, which provide action-aligned cues about local geometry, motion, and contact during interaction. The two systems operate asynchronously along a shared denoising trajectory: each global plan is reused across multiple local updates, while \sysone repeatedly incorporates fresh interaction feedback. Across zero-shot manipulation tasks on Franka and Galbot, \method improves success over the strongest evaluated baseline by 4.5 percentage points on average, while achieving a 16.6$\times$ critical-path speedup. Further studies show that role-matched egocentric and UMI data improve success by 14 percentage points, and that the decoupled design naturally supports edge--cloud deployment with substantially lower communication overhead than the baseline.
Steerable and Reactive Grasping Through Modular Design with a Three-Point Interface
基于模块化设计与三点接口的可操控、可响应抓取
Nguyen, Andrew, Lee, Yonghyeon, Kim, Sangbae
Abstract
Dexterous grasping requires deciding where to grasp, reaching the target, and maintaining stable contact. We connect these stages through a compact three-point interface that separates global geometric reasoning from local contact control. Given object geometry and optional language commands, our framework samples contact triples from a precomputed grasp-affordance heatmap. A model-based reactive controller tracks the object, avoids collisions, and guides the hand toward the selected contacts. In the final centimeters, a Reinforcement Learning (RL) policy uses proprioceptive feedback to refine and stabilize the grasp despite reaching and perception errors. It observes only finger joint states and its recent actions, with no target points, visual observations, or object geometry, so a single policy is shared across objects and grasp configurations. In simulation, we compare grasp-and-lift success against squeeze and end-to-end baselines, characterize reaching convergence, and demonstrate grasp steering; hardware demonstrations on two training objects and one unseen object illustrate the full pipeline. Our modular framework uses geometry to guide the reach and local feedback to secure the grasp.
Visuomotor Robotic Pruning in Planar Orchards Using Hybrid Reinforcement Learning
基于混合强化学习的平面果园视觉运动机器人修剪
Jain, Abhinav, Grimm, Cindy, Lee, Stefan
Abstract
Dormant tree pruning is labor-intensive yet essential for maintaining modern high-productivity fruit orchards. In this work, we focus on pruning of modern planar tree training systems - V-Trellis apples and UFO cherries - where trunks and primary branches are trained into approximately planar walls. We introduce an end-to-end pipeline to learn a closed-loop visuomotor controller for robotic pruning. This controller is trained entirely using simulation and synthetically generated data and deployed in real orchards in a zero-shot manner. The pipeline comprises synthetic generation of planar orchard tree meshes, construction of a physics-based orchard simulator, automated collection of successful pruning trajectories via motion planning, and policy learning with a novel hybrid reinforcement-learning algorithm that combines offline demonstrations with online simulated rollouts. The controller uses optical-flow inputs from a wrist-mounted camera - avoiding the need for full 3D-reconstruction - and continuously guides the cutter through cluttered branch environments to a specified cutpoint with correct tool orientation. In exhaustive simulated task-space evaluations over 3,000 pruning points, the policy attains 49.9% success on V-Trellis apples and 46.0% on UFO cherries. We validate the learned controller across 38 physical trials - comprising 28 outdoor field trials in commercial and experimental orchards and 10 indoor laboratory tests - demonstrating zero-shot sim-to-real transfer. The learned policy also outperforms a classical RRT-Connect baseline on physical hardware in laboratory trials.
Autonomous navigation on Mars requires vehicles to distinguish between traversable terrains across diverse and visually challenging environments. However, progress in learning-based navigation for off-world environments has been limited by the lack of large-scale datasets. Since landing in Jezero Crater, the Mars 2020 Perseverance rover has traversed terrain ranging from sandy dunes, rocky patches, and flat bedrocks. As a result, this paper presents a dataset spanning 500 sols and 45km of trajectories driven by both human operators and the onboard planner, ENav. Our dataset contains grayscale stereo image pairs, poses, accelerometer readings, rocker-bogie angles, and estimates of tilt and wheel slip. Building on this dataset, we introduce an uncertainty-aware traversability-estimation framework that learns terrain representations from multimodal driving experience. We compare our proposed method against existing approaches on the Mars 2020 dataset and show that our method achieves an AUROC of 0.874 and an F1 score of 0.758, outperforming the strongest baseline by 0.058 and 0.156, respectively, while also achieving the highest average precision and recall. Finally, we show that the visual representations can be integrated into path planners, such as ENav, on a physical rover test bed. Videos, code, and the M2020 dataset will be available at https://darren-chiu.github.io/learning-to-drive-on-mars.
Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to modeling contact evolution as part of the predicted world state rather than tactile conditioning alone: removing tactile world modeling reduces the four-task mean from 74.7 to 26.6 while keeping the same tactile features and action expert. After four hours of tactile-encoder adaptation with a frozen pretrained vision VAE, our continual vision-to-touch learning extends the pretrained video model to touch using roughly 100 demonstrations per task without tactile midtraining, while retaining visual prediction quality within 0.5 dB of vision-only counterparts. The compressor retains 89.4% of pre-fusion contact recall while enabling 2.26x faster training and 1.29x faster inference. Together, these results show that pretrained video priors can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner.
MIGU: Multimodal Instruction Grounding under Uncertainty for Manipulation Planning
MIGU:不确定性下的多模态指令定位与操作规划
Lu, Mingke, Xiao, Anxing, Hsu, David
Abstract
Understanding natural human instructions is crucial for deploying robots in human-centric environments. We study multimodal instruction grounding, where language and gesture provide complementary but uncertain cues. We present MIGU, a modular framework that combines semantic and geometric evidence into a unified grounding belief and connects it to manipulation planning. MIGU constructs a 3D geometric likelihood by propagating viewing-direction and depth uncertainty through eye-finger geometry while accounting for hand-direction estimation error. A vision-language model (VLM) provides semantic priors over candidate objects and regions, which are combined with the geometric likelihood through Bayes-inspired fusion. The resulting belief supports behavior planning to either proceed directly to downstream planning or request clarification. Grounded targets then define goals for mobile manipulation and tabletop task-and-motion planning. On a real-world benchmark, MIGU outperforms all evaluated baselines, while ablations support the benefit of explicit multimodal uncertainty modeling. Project website: multimodal-instruction.github.io
Behavior cloning for robot manipulation relies on expert demonstrations. However, for tasks that require dynamic stability, precise contact timing, or dexterous coordination, human operators may find it hard or even impossible to collect data. We study this infeasible-demonstration regime and propose GLIDE: Guardrails for Learning from Infeasible Demonstrations Efficiently, a framework that infers task-specific failure modes and converts them into executable guardrails for data collection and policy deployment. Given a task description and the conditioning teleoperation code, GLIDE writes guardrails that use system states to filter teleoperation and policy commands, constrain failure-prone actions, and iteratively improve from trajectory feedback. Across three tasks, GLIDE discovers emergent guardrails that go beyond domain-expert hardcoded ones, improving data collection over naive VR teleoperation and domain-expert hardcoded guardrails. After refinement, GLIDE raises data-collection success from 0-10 percent to 70-90 percent across the three tasks. During policy execution, mixed-data guarded policies reach 70 percent, 60 percent, and 60 percent success on Tomato plate transfer, Marker handover and stand, and Wine serving tasks. These results show that GLIDE can support policy learning when direct demonstrations are infeasible. Project website: http://guardrail-policy.github.io/
Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how these data types differentially shape model capabilities remains unclear. We address this issue with token-matched experiments that vary the didactic-to-clinical ratio and analyze how data composition affects performance, capability profiles, and error patterns across knowledge-intensive and clinic-oriented tasks. We uncover an asymmetric transfer across task types: clinical data improves clinic-oriented tasks while remaining competitive on knowledge-intensive ones, whereas didactic data mainly improves knowledge-intensive tasks. Error analysis suggests a knowing-doing gap, where improvements in knowledge recall do not reliably generalize to clinical reasoning. We further observe that modest amounts of clinical data yield most of the gains on EHR-grounded tasks, while the optimal mixture ratio varies with the knowledge and clinical reasoning demands of downstream tasks. These findings suggest that medical LLM data curation should be application-driven, with higher proportions of clinical data preferred for reasoning-intensive use cases.
An Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users
一种面向视障用户多模式出行辅助的经济型人工智能集成智能盲杖
Akarma, Ali, Ahmad, Adeel, Syed, Toqeer Ali
Abstract
Visual impairment affects over 2.2 billion people worldwide, yet conventional white canes cannot detect elevated hazards or provide semantic environmental context. Existing AI-assisted navigation systems typically rely on expensive hardware or cloud connectivity, limiting accessibility in resource-constrained settings. This paper presents an affordable (\$88 USD), fully offline AI-integrated smart cane designed for multimodal mobility assistance on an ultra-low-power Raspberry Pi Zero 2W. The system fuses RGB vision sensing with Time-of-Flight (ToF) distance estimation, pairing an INT8-quantized SSD MobileNet V1 model with distance-aware vibrotactile feedback and real-time audio alerts. To ensure operational robustness on constrained hardware, a multiprocessing architecture isolates sensor acquisition, neural inference, and haptic feedback into independent processes with fail-safe sensing support. Experimental evaluation across indoor mobility scenarios demonstrates a macro-averaged F1-score of 0.82 (precision: 0.85, recall: 0.81), a mean end-to-end latency of 330\,ms, and a peak power draw of 2.8\,W. A preliminary usability study with 12 participants (SUS: 78.5, NASA-TLX) demonstrated positive user perception and enhanced obstacle awareness. The proposed prototype validates the feasibility of deploying privacy-preserving, edge-native assistive intelligence for cost-sensitive mobility assistance.
Chinese Translation
全球有超过22亿人受到视力障碍的影响,然而传统的白手杖无法检测高处障碍或提供语义化的环境信息。现有的AI辅助导航系统通常依赖昂贵的硬件或云连接,限制了其在资源受限环境中的可及性。本文提出了一种经济实惠(88美元)、完全离线运行的AI集成智能盲杖,旨在超低功耗的Raspberry Pi Zero 2W上实现多模式出行辅助。该系统将RGB视觉传感与飞行时间测距相结合,将INT8量化的SSD MobileNet V1模型与距离感知的振动触觉反馈和实时音频警报相结合。为确保在受限硬件上的运行稳健性,系统采用多进程架构,将传感器采集、神经网络推理和触觉反馈隔离为独立进程,并支持故障安全传感。在室内出行场景中的实验评估表明,系统的宏平均F1分数为0.82(精确率:0.85,召回率:0.81),平均端到端延迟为330毫秒,峰值功耗为2.8瓦。一项包含12名参与者的初步可用性研究(SUS:78.5,NASA-TLX)显示了积极的用户感知和增强的障碍物感知能力。所提出的原型验证了为成本敏感的出行辅助部署隐私保护、边缘原生辅助智能的可行性。
PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation
PAANI:面向河流机器人仿真的设备端视觉证据融合与可解释引导
Cardoz, Savio, Rajan, Santhiya
Abstract
Mobile river monitoring robots must interpret obstacles and water boundaries that geographic waypoints alone cannot describe. On resource constrained platforms, converting imperfect visual predictions into timely and inspectable guidance is a distinct challenge. An object label or steering command does not explain which evidence supports a decision or when that evidence is unreliable. We present PAANI, an on-device perception to guidance architecture that combines a project trained YOLO11n detector and a custom MobileNetV3 Small semantic segmenter with timestamp aligned evidence fusion on Arduino UNO Q. Bounded tracking supplies object persistence, while an explicit corridor policy combines surface labels, accepted detections, urgency and mask uncertainty. Each final advisory exposes its contributing evidence and policy reasons. ROS 2 interfaces connect the local AI pipeline to a separate Gazebo vessel, localization and control testbed. Training uses 10,000 WaterScenes images for four-class detection and 1,127 MaSTr1325 images for segmentation, including 198 segmentation validation images. The selected FP32 ONNX models occupy 14.817 MB. Detector checkpoint test mAP at 0.5 IoU is 0.7388, while the separately evaluated rectangular ONNX export achieves validation mAP at 0.5 IoU of 0.7367. Segmentation ONNX validation mIoU is 0.9750. A five-minute UNO Q recording produced median and 95th percentile pipeline latencies of 467.8 ms and 580.3 ms at a configured 0.5 Hz cadence. The evaluation also identifies black input misclassification and a sampling rate mismatch that prevents the diagnostic apparent motion estimator from collecting sufficient evidence. These results support an inspectable and reusable edge robotics foundation while clearly distinguishing model accuracy and on-board execution from validated on-water collision avoidance.
Social Influence and the Allocation of Scientific Attention in AI Populations
社会影响与人工智能群体中科学注意力的分配
Chupilkin, Maxim
Abstract
AI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human readers, but the collective consequences of these signals for artificial readers remain uncertain. This paper adapts the Music Lab design to a market for academic attention. In the first experiment, 1,000 AI agents choose papers from the titles and abstracts of all 114 regular research articles published in the American Economic Review in 2025. The experiment has five independent-choice communities and five social-influence communities, each with 100 sequential agents. Only agents in the social-influence condition observe earlier selections within their community. Agents may select any number of papers. Social-information communities select 17.2 percent fewer papers per agent, concentrate their choices more heavily, and collectively cover 73 papers, compared with 90 independently. Between-community variation is greater under social information. In a second experiment with 200 agents across twenty social communities, randomly assigning papers five initial selections raises their subsequent selection rate by 45.55 percentage points (95% CI: 41.20 to 49.90). Choices have modest correspondence with external citations and little correspondence with download counts. The results show how a simple information rule shapes the volume, breadth and distribution of scientific attention in an artificial population.
Learning 3D biophysical cell properties from 2D images and cell-population statistics
从二维图像和细胞群体统计学习三维细胞生物物理特性
Hernández-Orozco, Santiago, Zenil, Hector
Abstract
Inferring 3D cellular properties from 2D microscopy is difficult when a reference instrument reports only population statistics rather than labels for individual cells. Here we develop a population-supervised framework that maps single 2D red-cell images to latent biophysical quantities and aggregates them to mean corpuscular volume, red-cell distribution width and mean corpuscular haemoglobin. The model combines shared local inference, a biophysically structured decoder for volume and haemoglobin, learned instance weighting and device-specific calibration. We formalise conditions under which aggregate observations identify restricted instance predictors, show why population agreement does not by itself identify single-cell properties or 3D geometry, and derive the dispersion penalty induced by subset mean matching. The development dataset comprises 390 specimens and 1,105 acquisitions across six devices, with reported Pearson correlations of 0.86--0.98 against a Sysmex analyser. The framework provides a testable route from 2D images and population supervision to 3D cellular biophysics without claiming explicit 3D reconstruction.
Process discovery rarely yields a single coherent process structure. For analysis, a common step is to cluster process variants based on structural similarity and then assign business meaning to the resulting groups. Since these partitions are not derived from the organization's goals, analysts must manually interpret and consolidate variants into business-meaningful categories. This judgment-intensive step becomes increasingly difficult as the number and complexity of variants grow. In this paper, we propose a goal-driven approach to variant categorization that reverses this workflow. We first author an organization's goal model that predefines the categorization axis. Each variant is transformed into a textual narrative describing its behavior, and a Large Language Model (LLM) interprets it in the context of the goal model and assigns the variant to the most appropriate category. LLM-based semantic reasoning connects low-level process behavior with analyst-defined business goals. We instantiate this approach end-to-end and evaluate it on three public logs differing substantially in scale and behavioral diversity. Goal-model guidance yields partitions that differ from those produced by unguided induction and respond to controlled edits to the declared alternatives, at the cost of authoring a goal model.
Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
托管式大语言模型中无持久性的复现:行动时信念评估中的测量敏感性
Joshi, Bhushan Kashinath
Abstract
Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration (replication), whether the endpoint changes when the evaluation-and-inference configuration is rebuilt under the same identifier (measurement sensitivity), and whether the finding persists across subsequently tested identifiers under one common instrument (persistence). We study these questions in Regent Chess, a sequential environment in which a hidden, mutable state is recorded exactly, allowing stated beliefs to be scored against ground truth at action time; positive endpoint values mean worse performance than a matched-uniform comparator. The previously reported Gemini 3.1 Flash-Lite deficit recurs on fresh games under its historical configuration (+0.0530, 95% CI [+0.0329,+0.0714]). In a back-to-back same-day H/R comparison under the same public identifier, the model-minus-uniform endpoint is 0.0429 lower under the rebuilt configuration (95% CI for the H-minus-R contrast [+0.0182,+0.0667]); all six configuration components vary jointly, so no component is isolated. Under rebuilt R, the prospectively frozen, interleaved same-window 4K comparison reverses sign between Gemini 3.1 and Gemini 3.7, identifiers that differ in release and product tier; additional descriptive and exploratory cells show the same directional pattern. Any additional serving-period contribution remains unresolved (-0.0166, [-0.0483,+0.0157]). Replication, measurement sensitivity, and persistence can therefore yield different conclusions within one evaluation, motivating explicit indexing of hosted-model behavioural claims by tested identifier, serving period, measurement instrument, and inference configuration.
Chinese Translation
对托管式语言模型的行为评估可能因被评估服务、测量工具或两者在不同运行中存在差异而产生变化。我们区分了三个验证问题:先前发现在其历史配置下能否在新数据上重现(复现,replication);在相同标识符下重建评估与推理配置时端点指标是否发生变化(测量敏感性,measurement sensitivity);以及在同一测量工具下,该发现能否在后续测试的标识符间保持(持久性,persistence)。我们在 Regent Chess(一个顺序决策环境)中研究这些问题。在该环境中,隐藏且可变的状态被精确记录,使得模型声明的信念可以在行动时刻与真值(ground truth)进行比对打分;端点指标为正值意味着性能劣于匹配均匀分布的对照基线。先前报告的 Gemini 3.1 Flash-Lite 的性能缺陷在其历史配置下的新对局中重现(+0.0530,95% CI [+0.0329,+0.0714])。在相同公共标识符下的同日背靠背 H/R 对比中,重建配置下的“模型减均匀基线”端点降低了 0.0429(H 减 R 对比的 95% CI 为 [+0.0182,+0.0667]);由于六个配置组件同时联合变动,因此无法分离出任何单一组件的影响。在重建的 R 配置下,前瞻性冻结的、交错同时间窗口的 4K 对比在 Gemini 3.1 与 Gemini 3.7 之间出现了符号反转,而这两个标识符在发布版本和产品层级上均不同;其他描述性与探索性分析单元也显示出相同的方向性模式。服务期可能带来的额外贡献仍未得到解决(-0.0166,[-0.0483,+0.0157])。因此,复现、测量敏感性和持久性在同一评估中可能得出不同结论,这提示我们应对托管式模型的行为学结论进行显式索引,即按被测标识符、服务期、测量工具和推理配置进行标注。
The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as the wisdom of crowds. While this is usually attributed to the independence of estimates, an even stronger effect arises through deliberation: averaging the consensus estimates of small deliberating groups outperforms the classical wisdom of crowds, with individual judgments themselves also becoming more accurate after deliberation. Whether these improvements transfer to large language models deliberating amongst themselves is unknown. Here we adapt a three-stage deliberation paradigm previously used with human participants for use with large language models from three different families, and test it across four domains of increasing real-world stakes: visual numerical estimation (Study 1), peer review of machine-learning papers (Study 2), detection of hidden malicious behavior by an artificial intelligence agent (Study 3), and sports forecasting against a real prediction market (Study 4). Across domains, deliberation reduced collective error beyond passive aggregation of independent responses, and post-deliberation individual judgments retained this collective gain. Notably, the advantage required model diversity: groups composed of clones of a single model did not benefit from deliberating. These results establish machine deliberation as a general-purpose aggregation mechanism, and point to diversity as an active ingredient.
Chinese Translation
将众多非专业者的估计进行聚合,其表现往往优于个体专家的判断,这一现象被称为群体智慧(wisdom of crowds)。尽管这通常被归因于估计的独立性,但通过商讨(deliberation)会产生更强的效果:对小型商讨小组的共识估计取平均值,其表现优于经典的群体智慧,且个体判断本身在商讨后也变得更加准确。这些改进是否也能迁移到大型语言模型之间的相互商讨中,目前尚不清楚。本研究将此前用于人类被试的三阶段商讨范式加以改造,应用于来自三个不同家族的大型语言模型,并在四个现实风险逐级递增的领域中进行测试:视觉数值估计(研究1)、机器学习论文的同行评审(研究2)、检测人工智能智能体的隐藏恶意行为(研究3),以及与真实预测市场对比的体育赛事预测(研究4)。在各个领域中,商讨均降低了集体误差,效果超越了独立回答的被动聚合,且商讨后的个体判断保留了这一集体增益。值得注意的是,这种优势依赖于模型多样性:由单一模型克隆组成的群体无法从商讨中获益。这些结果确立了机器商讨作为一种通用的聚合机制,并表明多样性是其有效成分。
Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
一致性高估了证据强度:大语言模型裁判共识中的误差相关性
Hossain, Elias, Yousefi, Niloofar, Lim, Ser-Nam
Abstract
Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes. We study how this dependency affects the reliability of consensus. We find substantial error correlation across both open-weight and frontier LLM judges. In our main bank of ten judges, the average pairwise correlation between judge errors is 0.21. As a result, the ten judges only provide roughly as much statistical information as 3.5 independent judges. The dependency is even stronger among the high-accuracy frontier judges we evaluate, including judges from different providers. In up to 28% of our comparisons, ignoring shared errors leads to the conclusion that one system is significantly better, while accounting for them does not. We also find that the pattern of errors matters. Errors shared by most judges and errors concentrated among a smaller group affect consensus differently and favor different voting methods. Measuring the overall amount of correlation alone is therefore insufficient. Our results suggest a simple approach: use a small set of trusted examples to estimate judge accuracy and identify shared mistakes. These shared errors should then be considered when analyzing the results, and the voting method should be chosen using trusted examples before it is applied to new data.
IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law
IntLawNER:一个国际法领域的命名实体识别数据集与基准
Skura, Genis, Bouffanais, Roland, Wernli, Didier
Abstract
International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its texts remain without token-level named entity recognition (NER) resources. We introduce IntLawNER, a NER dataset and benchmark for codified sources of international law, covering 2,987 gold-annotated sentences and 8,094 entity spans from International Court of Justice (ICJ) decisions, UN Security Council resolutions, and European Court of Human Rights (ECtHR) judgments, annotated with seven institution-specific entity types. We construct IntLawNER with a cost-effective hybrid algorithmic-agentic pipeline that reduces 468k source sentences to a compact annotation set through candidate retrieval, LLM-based vetting, and human review, with 89.6% of gold spans accepted unchanged from the silver layer. However, the silver-to-gold analysis reveals that human-machine aggregate agreement metrics can be misleading in domain-specific NER: Cohen's kappa=0.964 on boundary-matched spans masks a macro-F1 of 0.753 when missing entities, boundary errors, and label corrections are included. The benchmark shows that zero-shot span-based GLiNER collapses on entity types dependent on institutional function rather than surface form (0.243 micro-F1), while fine-tuned transformers struggle on rare labels. Carefully selected few-shot examples that demonstrate label contrasts improve every LLM over zero-shot prompting, with Claude Opus 4.6 reaching the best score of 0.873 micro-F1. We release IntLawNER as a benchmark and reusable resource for extracting references in international legal texts.
Chinese Translation
国际法为国家协调行动、规范武装冲突和保护人权提供了规范框架,然而其文本至今仍缺乏词元级别的命名实体识别(NER)资源。我们提出了 IntLawNER,一个面向国际法成文来源的命名实体识别数据集与基准,涵盖 2,987 个黄金标注句子和 8,094 个实体跨度,数据来源包括国际法院(ICJ)判决、联合国安理会决议以及欧洲人权法院(ECtHR)判决,并标注了七种机构特定的实体类型。我们通过一个经济高效的算法-智能体混合流水线构建了 IntLawNER:通过候选检索、基于大语言模型(LLM)的筛选和人工审核,将 468,000 条源句子缩减为一个紧凑的标注集合,其中 89.6% 的黄金跨度直接沿用银层标注而未作修改。然而,银层到金层的分析表明,人机总体一致性指标在领域特定的 NER 任务中可能具有误导性:在边界匹配的跨度上,Cohen's kappa 达到 0.964,但若将实体遗漏、边界错误和标签修正纳入考量,宏观 F1 仅为 0.753。基准测试显示,零样本的基于跨度的 GLiNER 在依赖机构功能而非表面形式的实体类型上表现崩溃(微观 F1 仅为 0.243),而微调后的 Transformer 模型在稀有标签上表现不佳。经过精心挑选、能够展示标签对比的少样本示例能够使所有大语言模型的表现均优于零样本提示,其中 Claude Opus 4.6 取得了 0.873 微观 F1 的最佳成绩。我们将 IntLawNER 作为一个基准和可复用资源发布,用于国际法文本中的指称提取。
EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability
EvidenT:通过证据的忠实性与可追溯性构建可信的企业智能助手
Kabra, Anubha, Kim, Katie Jooyoung, Kou, Colin Zhiwei, Sajer, Helene, Fan, Yimei, Cisar, Radomir, Greenhalgh, Heather, Vidiri, Gabriel Martinez
Abstract
Enterprise AI assistants must produce responses that are verifiable and traceable to source evidence. However, retrieval augmented generation (RAG) over heterogeneous enterprise data can suffer from citation drift, unsupported content, and weak source traceability. We present EvidenT (T = Trust + Transparency + Traceability), a lightweight pipeline that verifies extracted evidence against retrieved documents before answer generation, without model retraining. EvidenT combines structured passage extraction with deterministic lexical alignment to filter unsupported content, correct citation drift, and preserve source-span traceability. On approximately 500 real enterprise queries, EvidenT improves gold-source hit rate by an average of 29% over prompting baselines, produces no citations to nonretrieved urls, and achieves near-saturated answer-to-source lexical coverage.
Training agents with reinforcement learning requires a gym, comprising a task, an executable environment in which the task can be attempted, and a verifier that reliably distinguishes success from failure. Constructing such gyms remains manual, expensive, and static. Task sets saturate as models improve and are increasingly exposed to contamination. Synthetic generation offers scale, but single-pass synthesis produces tasks whose difficulty is largely cosmetic. Models comparable in capability solve them despite convoluted phrasing, and correctness must be adjudicated post-hoc by unreliable LLM judges. We present AutoGym, a framework that generates complete gyms (tasks, executable environments, and verifiers) from a minimal domain seed or prior model trajectories. AutoGym introduces three mechanisms. (1) Blueprint-first generation specifies the valid solution space, environment requirements, and verification criteria before the environment is materialized, making solvability a construction prerequisite rather than a property verified after the fact. (2) Explicit generation parameters control task topology, interaction depth, capability axes, question obfuscation, and distractor composition, enabling fine-grained difficulty steering. (3) Active curriculum synthesis uses performance-informed calibration to adjust the distribution over these parameters as model capabilities evolve. Across productivity and temporal-reasoning settings, AutoGym generates gyms spanning the capability spectrum, including instances that challenge frontier models.
Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent outputs, but their verdicts can be sensitive to incidental changes in the evaluated response, judge instructions, and scoring rubric. Existing systems examine important subsets of these failure modes, but auditing a configured judge requires testing both the judge instrument and the items it evaluates. We introduce MAWILE, a developer-facing workbench for auditing judge sensitivity across four surfaces: the judge prompt, judge rubric, target-system input, and target-system output. Given a user-supplied judge and representative evaluation items, MAWILE constructs and validates controlled perturbations, re-executes the judge, and localizes the resulting sensitivity. Each perturbation declares whether the verdict should remain invariant or change in a specified direction, allowing the same system to measure both robustness to irrelevant variations and sensitivity to meaningful changes. MAWILE audits binary, ordinal, and pairwise judges without requiring gold labels. The code for this tool is available at: github.com/megagonlabs/mawile-judge.
GaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment
GaitVista:面向可及性纵向步态评估的可靠性感知AI测量方法
Jayasinghe, Nethmi, Parashar, Mihir, Trivedi, Amit Ranjan
Abstract
Tracking recovery of walking function requires detecting meaningful gait change across rehabilitation sessions, yet objective 3D measurement remains confined to specialized motion-capture laboratories. Small camera sets and body-worn inertial sensors broaden access, but reliability varies across joints and time, allowing sensing failures to masquerade as patient change. We present \textsc{GaitVista}, a reliability-aware measurement layer whose lightweight gate assigns joint- and frame-specific visual contributions using camera coverage, local visual quality, cross-modal disagreement, and root-motion continuity, and exposes them for inspection. Across seven clean and degraded sensing conditions on TotalCapture, \textsc{GaitVista} reduces average full-body and lower-body error by \textbf{27.7\%} and \textbf{27.8\%}, attains the lowest worst-condition error among fusion methods, and reduces the gap to a joint-frame oracle from $2.76$--$5.33$~cm for condition-blind baselines to $1.11$~cm. On MoVi with image-derived keypoints, it is the only deployable fusion method to improve over both unimodal streams, reducing marker-supported error by \textbf{6.4\%} relative to the strongest learned fusion baseline. On TotalCapture, it improves bilateral knee-flexion waveform accuracy by \textbf{18.9\%}. Raw inertial measurements from five TotalCapture participants show location- and time-varying magnetic disturbance, supporting the design's reliability premise. Both benchmarks contain neurologically healthy participants in controlled settings and retain participant-specific IMU calibration; we therefore report progress toward accessible gait assessment, not validated clinical deployment.
Scanned mail, uploaded PDFs, and consolidated attachments often arrive as page streams that must be split into individual documents before downstream classification, extraction, or routing. Zero-shot large language models can detect document boundaries without task-specific training, but standard Page Classification (PC) and Boundary Decision (BD) formulations resolve only one boundary per model call. We introduce Multi-Split Boundary Decision (MSBD), which predicts multiple boundaries within a page window in a single call, reducing the number of inference requests. We evaluate MSBD across multiple language models, document collections, input modalities, and window sizes. The results reveal a model- and corpus-dependent operating range in which MSBD preserves strong segmentation accuracy while substantially improving inference efficiency, followed by a sharp decline at larger windows. MSBD provided the strongest overall accuracy--efficiency trade-off, while large windows expose distinct over- and under-segmentation behavior across models. These findings show that multi-boundary prediction can make zero-shot page stream segmentation more efficient when the window size is selected for the target corpus.
Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this choice, holding the prompt, judge, and scoring pipeline fixed, across four commercial model endpoints, two corpora, and two context regimes (gold evidence pages and the full document). On documents that fit the image budget, page images lead on accuracy at every document length on both corpora, but this advantage carries a growing latency and cost premium: text latency stays roughly flat as documents lengthen while image latency rises steadily. Text and images also fail on different questions, with exactly one representation correct on 19--25% of items across the reported cells, so neither subsumes the other. Exploiting this complementarity, a lightweight TF-IDF router that reads only the question text gains 2.6 points over always-text while cutting median latency 30% relative to always-vision, on a document-disjoint held-out split.
Self-Organizing Agent Teams Learn to Reason Together
自组织智能体团队学会协同推理
Pappu, Aneesh, Suzgun, Mirac, Kwon, Yongchan, Bianchi, Federico, El, Batu, Kochenderfer, Mykel J., Cao, Hancheng, Zou, James
Abstract
Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasoning as it unfolds. Human teams routinely adapt this way, while existing AI agent teams rely on fixed protocols, explicit task decomposition, or routing. We introduce Self-Organizing Agent Teams (SAT), fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow. These strategies enable what we call collaborative computation: agents exchange, challenge, repair, and synthesize partial reasoning into solutions no member produced independently. In two independent settings, we learn teamwork strategies that transfer unchanged to unseen benchmarks, using only 15 mathematics and 25 graduate-level knowledge problems. Across five mathematics and physics benchmarks, self-organizing teams average 66.7% accuracy, versus 48.8% for their strongest member, 58.7% for compute-matched inference by that agent, and 59.0% for a perfect router over members' independent answers; on AIME 2026, they exceed this router by 13.4 points. Because gains vary across benchmarks, we ask when self-organizing collaboration helps. Across eight benchmarks, demonstrability (the organizational-psychology construct of whether a team can distinguish correct from incorrect reasoning) strongly tracks improvement over the strongest member (Spearman $\rho=0.90$, $p=0.005$): teams benefit most when correct reasoning can be recognized once it appears. More broadly, these results suggest that organization itself can become an agent capability: agent teams can learn how to reason together and produce solutions their members could not reach independently.
An enduring and richly elaborated dichotomy in cognitive neuroscience is that of human behavior control mechanisms, divided into habitual versus goal-directed. While existing human-like agent frameworks primarily focus on modeling goal- directed behavior, habitual behavior has been largely overlooked, though it plays a crucial role in human daily life. In this paper, we address this gap by studying multiple behavior control systems that jointly model goal-directed and habitual behaviors. We propose a human behavior control mechanism-inspired framework which the Habitual Controller retrieves cue-triggered behaviors from personal- ized habit memory, while the Goal-directed Controller employs a context-aware world model to predict action consequences and estimate their values. The Arbiter dynamically balances the influence of both systems according to individual differ- ences and momentary internal states. To reconstruct diverse human-level behavior instructions in 3D environments, we further develop a keyframe-guided 3D mo- tion generation module. Through extensive evaluation methods, human studies, and ablations studies, experimental results demonstrate that human-likeness per- formance is significantly improved by our approach. The efficacy of our approach indicates the benefits of leveraging habitual behavior and multiple behavior con- trol system coordination for believable embodied human-like agents.
Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review
面向可审计与经校准的痴呆相关事故严重程度预测AI:一种支持人工审查的选择性延迟决策框架
Chhetri, Gaurab, Baitullah, Anika, Das, Subasish
Abstract
Public crash databases increasingly support automated safety analysis, but crash severity prediction remains difficult to translate into public-sector decision workflows when models are evaluated primarily as ordinary classifiers. This study reframes dementia-related crash severity modeling as a decision-aware triage problem in which a system must classify crashes into no-injury/property-damage-only (O), minor or moderate injury (BC), and fatal or severe injury (KA), while also controlling outcome leakage, reporting severe under-triage, calibrating confidence, and preserving every raw prediction for audit. Using 4,781 Texas crash records with structured fields and police narratives, we evaluate structured, narrative, fusion, calibrated fusion, BERT-family, and local large-language-model baselines under a stratified 70/15/15 split. In the reported split, leakage-controlled Gemma obtains the highest observed macro-F1 (0.545; 95% bootstrap CI [0.507, 0.583]). The best calibrated fusion model obtains macro-F1 of 0.522 and expected calibration error of 0.033. Selective deferral improves performance among cases retained for automatic classification. At 70% coverage, macro-F1 rises to 0.573 and severity cost falls to 0.577, while deferred cases are treated as candidates for a proposed human-review process and are not further evaluated in the present experiment. The study contributes a reproducible, leakage-controlled, and uncertainty-aware evaluation framework for crash AI systems, emphasizing auditability and selective deferral rather than accuracy alone.
Lee, Sewoong, Canby, Marc E., Cho, Ikhyun, Hockenmaier, Julia
Abstract
The term "linear representation hypothesis" (LRH) has appeared across diverse subfields of artificial intelligence, neuroscience, and cognitive science. But previous works have not consistently treated the LRH as a falsifiable scientific hypothesis; we analyze these inconsistencies and examine their implications for how prior theoretical and methodological results should be interpreted. Based on this analysis, we argue that claims regarding linear representations become well-defined only through careful examination of the model, representation location, feature definition, and evaluation dataset. We therefore propose a more rigorous formalization of the LRH that makes these dependencies explicit and allows the hypothesis to be evaluated as a falsifiable scientific claim. Finally, we identify some non-trivial open problems that warrant further attention from the research community.
Building Trustworthy Mental Health Benchmarks on Bluesky: A Validation-Aware Weak-Supervision Framework
在Bluesky上构建可信的心理健康基准:一种注重验证的弱监督框架
Chhetri, Gaurab, Dutta, Anandi, Das, Subasish
Abstract
Decentralized social media platforms create new opportunities and challenges for computational mental health research because data access, moderation, labeling, and deployment responsibilities are distributed across multiple technical and governance layers. This paper presents a validation-aware weak-supervision system for constructing and evaluating suicidal ideation (SI) and broader mental health (MH) disclosure benchmarks on Bluesky, a decentralized social media platform built on the AT Protocol. The system integrates public firehose collection, task-specific lexicon filtering, Llama-3-8B-assisted binary annotation, human-adjudicated validation subsets, and transformer-based model benchmarking. Using this pipeline, we construct two task-specific corpora containing 8,346 SI-labeled posts and 9,988 MH-labeled posts. The evaluation shows that model performance depends strongly on both task definition and validation protocol. BERT+LSTM achieves the highest SI stratified cross-validation F1-score, RoBERTa achieves the strongest SI holdout F1-score, and DistilRoBERTa achieves the best MH cross-validation F1-score. Human validation reveals different weak-label failure modes across tasks, with SI labels dominated by false negatives and MH labels dominated by false positives. These findings show that decentralized social media can support reproducible mental health benchmarking, but only when system design, label provenance, validation strategy, and deployment constraints are evaluated together.
Hapi: A Multivariable Land-Surface Transformer for Medium-Range Hydrological Forecasting at Continental Scale
Hapi:面向大陆尺度中期水文预报的多变量陆面Transformer模型
Zhang, Hong, Hutchison, John K., Kotamarthi, Rao, Feinstein, Jeremy, Guan, Haiwen, Maulik, Romit, Ramalingam, Vijay P., Stock, Jason, Wall, Tom
Abstract
Accurate flood forecasts several days in advance are essential for flood control, water-resource management, and emergency response. Producing them at high resolution over a continental domain calls for local hydrological detail together with spatial context extending from river basins to synoptic weather systems. We developed Hapi, a U-Net Swin Transformer that uses fine three-dimensional patches and hierarchical shifted-window attention to forecast discharge, surface runoff, snow water equivalent, and soil wetness across the contiguous United States. The model produces 24--72-hour forecasts at $0.05^{\circ}$ resolution, with learned Laplacian task weights adjusting each variable's contribution to training. On 2024 test data using reconstructed weather and land-surface inputs from ERA5-Land, Hapi outperformed an operational physics-based model and a state-of-the-art AI model in flood detection. Independent validation against 3,881 U.S. Geological Survey gauges and a Hurricane Helene case study supported its advantage over the physics-based model in reproducing daily discharge. Controlled experiments showed that learned task weighting strengthens rare-flood detection, which is particularly sensitive to changes in precipitation inputs. Hapi produced a four-variable, 72-hour forecast across the contiguous United States with an average inference time of 0.11 seconds on a single A100 GPU.
Trustworthy Agentic AI: Failure Modes, Mitigation Strategies, and a Lifecycle Framework for Autonomous LLM Systems
可信赖的智能体AI:自主大语言模型系统的失效模式、缓解策略与生命周期框架
Syed, Fayeq Jeelani, Ahmad, Rehan, Bataineh, Ali Al, Adhikari, Aakriti
Abstract
Agentic AI systems built on large language models can plan over multiple steps, use external tools, retain information in memory, and coordinate with other agents. These capabilities make them more useful than static language models, but they also introduce new security and operational risks. Untrusted content from websites, emails, documents, and databases can enter the same context as system instructions; persistent memory can carry compromised information across sessions; and access to external tools can turn an incorrect model response into a consequential real-world action. This article reviews the trustworthiness of agentic AI across five interconnected dimensions: safety and robustness, alignment and human oversight, transparency and auditability, privacy and data governance, and regulatory compliance. It organizes key failure modes, including indirect prompt injection, backdoor triggers, goal misgeneralization, memory contamination, and cross-session data leakage, into a unified taxonomy. It also examines major mitigation approaches, such as instruction hierarchies, context isolation, spotlighting, process-based supervision, constrained tool use, and privacy-preserving memory, while distinguishing techniques supported by empirical evidence from those that remain largely conceptual. Building on this analysis, we introduce the Trustworthy Agent Development Lifecycle (TADL), a six-phase framework covering specification, design, training, evaluation, deployment, and monitoring. For each phase, TADL identifies relevant trust activities, expected evidence, and risk-based decision gates. Although TADL has not yet been empirically validated, it provides a structured foundation for developing and evaluating more secure and accountable agentic systems. The article concludes by identifying gaps in current benchmarks and outlining priorities for future research.
Chinese Translation
基于大语言模型构建的智能体AI系统能够进行多步骤规划、使用外部工具、在内存中保留信息,并与其他智能体协同工作。这些能力使其比静态语言模型更为实用,但同时也带来了新的安全与运行风险。来自网站、电子邮件、文档和数据库的不可信内容可能与系统指令进入同一上下文;持久化记忆可能将受污染的信息跨会话传递;而对外部工具的访问可能将错误的模型响应转化为产生实际后果的现实世界行为。本文从五个相互关联的维度审视智能体AI的可信性:安全性与鲁棒性、对齐与人类监督、透明性与可审计性、隐私与数据治理,以及合规性。文章将关键失效模式,包括间接提示注入、后门触发、目标错误泛化、记忆污染和跨会话数据泄露,整合为一个统一的分类体系。同时,本文考察了主要的缓解方法,如指令层级、上下文隔离、聚焦标注(spotlighting)、基于过程的监督、受约束的工具使用以及隐私保护记忆,并区分了有实证依据支持的技术与仍主要停留在概念层面的技术。基于上述分析,我们提出了可信赖智能体开发生命周期(Trustworthy Agent Development Lifecycle, TADL),这是一个涵盖规格定义、设计、训练、评估、部署和监控六个阶段的框架。针对每个阶段,TADL 明确了相关的可信性活动、预期证据以及基于风险的决策关卡。尽管 TADL 尚未经过实证验证,但它为开发和评估更安全、更负责任的智能体系统提供了结构化的基础。文章最后指出了当前基准测试中的不足,并展望了未来研究的优先方向。
Large Language Models (LLMs) have recently been introduced into traffic signal control (TSC) as decision agents due to their strengths in human-readable reasoning generation. Yet, existing LLM TSC methods optimize only from final outcomes and fail to distinguish valid from flawed reasoning steps, causing useful or misleading steps to be jointly updated and thus impairing the model's learning of effective reasoning. To bridge this gap, we propose an LLM-based framework ProcessLight to decompose signal decisions into verifiable semantic steps. Building on ProcessLight, we further develop Step-wise Traffic Process Policy Optimization (STeP-PO), a novel reinforcement learning framework that optimizes structured reasoning processes through step-level credit assignment. Specifically, STeP-PO uses step quality scores to evaluate local reasoning quality and step importance to measure each step's influence on the final action, and then assigns step-level advantages over a semantic step tree structure. The resulting step-level advantages are propagated to reasoning tokens, enabling fine-grained policy optimization beyond outcome-only rewards. Extensive experiments over multiple real-world datasets demonstrate the superiority of our methods. Our code is available at https://github.com/wenzhaoabc/processlight.
Chinese Translation
大语言模型(LLMs)因其在生成人类可读推理方面的优势,近来被引入交通信号控制(TSC)领域作为决策智能体。然而,现有的LLM交通信号控制方法仅从最终结果进行优化,无法区分有效与有缺陷的推理步骤,导致有用和误导性的步骤被同时更新,从而损害了模型对有效推理的学习。为弥补这一不足,我们提出了一个基于LLM的框架ProcessLight,将信号决策分解为可验证的语义步骤。在ProcessLight的基础上,我们进一步提出了逐步交通过程策略优化(Step-wise Traffic Process Policy Optimization, STeP-PO),这是一种通过步骤级信用分配来优化结构化推理过程的新型强化学习框架。具体而言,STeP-PO使用步骤质量分数来评估局部推理质量,使用步骤重要性来衡量每个步骤对最终动作的影响,然后在语义步骤树结构上分配步骤级优势。由此得到的步骤级优势被传播到推理token上,从而实现超越仅基于结果奖励的细粒度策略优化。在多个真实世界数据集上的大量实验证明了我们方法的优越性。我们的代码已发布于 https://github.com/wenzhaoabc/processlight。
Purpose: A vertebra at the lumbosacral junction is named by counting caudally from C2 on whole-spine imaging, but a lumbar case is planned on lumbar-only imaging (T12 to S1), without C2. Abdominopelvic CT holds that span plus the lowest ribs and pelvis. Where a lumbosacral transitional vertebra (LSTV) alters the count, the local anatomy is ambiguous: four rib-free vertebrae may be an L1 with a lumbar rib or an L5 assimilated to the sacrum, and six may be a sixth lumbar vertebra, a T12 with aplastic ribs, or a lumbarized S1. CTSpinoPelvic1K asks whether local morphology resolves it without the count. CTSpine1K's vertebrae and CTPelvic1K's pelvis covered these patients but were never joined; this release joins them on one series and adds the bones neither had. It provides 802 CT records with per-level ribs and femora, levels anchored on the lowest rib-bearing vertebra and S1, and classes for L6, T13, a separate S1 and lumbar ribs, so anomalies are recorded as such. Acquisition and Validation Methods: Records pair CTSpine1K and CTPelvic1K labels on each patient's bone-richest series under a VerSe-native scheme. Validation covered geometric invariants (802/802 pass), rib-vertebra incidence across 5,749 ribs (0.035% offset), and spinopelvic measures matching published values. Data Format and Usage Notes: NIfTI image/label pairs with patient-grouped LSTV-stratified five-fold splits and a loader; archived at https://doi.org/10.5281/zenodo.22139642. Potential Applications: Classifying a vertebra from local features; updating cadaveric morphometry; spinopelvic assessment; opportunistic screening; and, absent a public preoperative lumbar cohort, surgical planning research (377 records prone). Limitations: thoracic ground truth is field-of-view limited; postural angles supine; no held-out test set; ribs are triaged-review pseudolabels; Castellvi grades two-reader consensus on 33 records.
DVA-Neurons: Design and Verification of Adaptive LIF Neurons: From Single-Neuron Dynamics to Multi-Neuron Spiking Networks
DVA-神经元:自适应LIF神经元的设计与验证——从单神经元动力学到多神经元脉冲网络
Pham, Thanh, Islam, Riadul
Abstract
Spiking Neural Networks (SNNs) offer a promising path toward ultra-low-power artificial intelligence inference by emulating the event-driven computation of biological neurons. However, two challenges limit their practical deployment. First, fixed-parameter Leaky Integrate-and-Fire (LIF) neurons lack the adaptation mechanisms observed in biology, where neurons modulate their excitability based on firing history. Second, scaling from single neurons to multi-neuron networks introduces challenges in synaptic weight distribution and inter-neuron spike routing that are absent in isolated designs. This paper addresses both issues through the extension, verification, and physical implementation of adaptive LIF neurons at three architectural scales. This work contributes: a 2nd-order neuron with two-stage synaptic filtering for richer temporal dynamics; a fully-connected 6-neuron spiking network with configurable weights (100 to 5) demonstrating weight-based inter-neuron communication; and a direct verification methodology enabling per-cycle observation of all internal states. All designs were synthesized targeting Selected Area Electron Diffraction (SAED) 14 nm Complementary Metal-Oxide-Semiconductor (CMOS) technology at 1 GHz and verified with Cocotb-based Python testbenches under pulsed current stimuli (amplitude 80, ISI=3). The results show that adaptation effectively modulates firing: 31% suppression in the 2nd-order neuron (25 vs.\ 36 spikes) and 31% reduction in postsynaptic firing in the network (18 vs.\ 26 spikes). Physically, the 2nd-order neuron costs 1.77x more area and 1.52x more power than the 1st-order baseline, while the 6-neuron network demonstrates near-linear scaling (5.7x area, 5.3x power). Seven verification bugs spanning testbench connectivity, fixed-point overflow, and Verilog expression-width semantics are documented.
Chinese Translation
脉冲神经网络(SNN)通过模拟生物神经元的事件驱动计算,为实现超低功耗人工智能推理提供了一条有前景的路径。然而,两个挑战限制了其实际部署。其一,固定参数的泄漏积分发放(Leaky Integrate-and-Fire, LIF)神经元缺乏生物中观察到的自适应机制,即神经元会根据发放历史调节其兴奋性。其二,从单神经元扩展到多神经元网络会带来突触权重分布和神经元间脉冲路由等在孤立设计中不存在的问题。本文通过在三个架构层级上对自适应LIF神经元进行扩展、验证和物理实现,同时解决了这两个问题。本工作的贡献包括:一个具有两级突触滤波的二阶神经元,以实现更丰富的时间动力学;一个具有可配置权重(100至5)的全连接六神经元脉冲网络,展示了基于权重的神经元间通信;以及一种支持对所有内部状态进行逐周期观测的直接验证方法学。所有设计均基于Selected Area Electron Diffraction(SAED)14纳米互补金属氧化物半导体(CMOS)工艺在1 GHz下综合,并通过基于Cocotb的Python测试平台在脉冲电流激励(幅值80,ISI=3)下进行验证。结果表明,自适应机制能有效调节发放:二阶神经元的发放被抑制31%(25次对比36次脉冲),网络中突触后发放减少31%(18次对比26次脉冲)。在物理实现上,二阶神经元的面积和功耗分别为一阶基线的1.77倍和1.52倍,而六神经元网络表现出近线性扩展(面积5.7倍,功耗5.3倍)。此外,论文记录了七个验证缺陷,涉及测试平台连接、定点溢出以及Verilog表达式位宽语义等问题。
From Research Frontier to Laboratory Bench: Design of a Four-Tier Experimental Teaching System for Multimodal Medical Image Intelligent Diagnosis
从研究前沿到实验台:面向多模态医学影像智能诊断的四层次实验教学系统设计
Shan, Dongjing, Luo, Yamei, Li, Jin, Luo, Yong
Abstract
Undergraduate programmes in intelligent medical engineering are expanding, yet laboratory curricula lag behind the multimodal, long-tailed, and distributionally shifting realities of clinical AI. This design paper presents an advanced experimental teaching system that translates an ongoing multimodal deep learning research project on endometrial carcinoma into a structured undergraduate lab sequence. We identify three educational gaps (modality, authenticity, and deployment) and derive four pedagogical principles from constructive alignment, experiential learning, the research teaching nexus, and the CDIO framework. The curriculum comprises four progressive tiers plus an engineering layer, with 32 laboratory units over 64 contact hours, delivered via a custom virtual clinical workstation using de-identified multi-institutional data. Each tier maps to a specific technical bottleneck, prerequisite coursework, and criterion-referenced deliverables. Data governance, safety, and assessment protocols are specified. Learning outcome data will be collected across two implementation cycles.
ISA-Bench: A Benchmark for Computational Reasoning Across Instruction Set Architectures
ISA-Bench:一个面向跨指令集架构计算推理的基准测试
Pola, Aditya, Majumdar, Arkaprava, Balasubramanian, Vineeth N.
Abstract
Large language model code generation benchmarks primarily evaluate well-resourced languages like Python and Java, where models benefit from abundant training data. They provide limited evidence about reasoning in unfamiliar computational models: deriving arithmetic from a single subtract instruction, coordinating parallel programs across communicating nodes, or wiring logic gates into circuits. We present ISA-Bench, a benchmark of programming games with constrained instruction sets. For each game we provide a full execution stack (parser, VM, and verifier), enabling automated evaluation with structured feedback for iterative refinement. Reasoning models achieve higher average solve rates than code-specialized and general-purpose models, but unfamiliar syntax remains a major source of failure. Models solve more tasks with iterative feedback, though the gains vary substantially across architectures. We introduce a reasoning--execution gap (REG) analysis that reveals a recurring disconnect between identifying a plausible computational strategy and expressing it as a correct program in the target ISA. Code is open-sourced.
Vision-language agents that crop and zoom are trained with rewards that credit a successful tool call, yet a successful call does not show that the model needed to look or used the pixels it received. On our cold-start checkpoint only 10% to 12% of visual calls were both needed and used, and released agents make spurious calls 36% to 87% of the time on individual benchmarks. Outcome rewards, judge rewards, and branch probes each observe one side of this failure, and about two thirds of what an outcome reward pays goes to calls that were neither needed nor used. CounterCredit asks both questions of every image-returning call at its realized pre-call state, using the policy's own gold-answer score. A decision value compares the realized visual branch with answering immediately; an evidence value compares the returned crop with random same-size patches substituted into the same call. A call verified on both earns cashback and every other executed call pays rent; the price is bounded so that every correct trajectory outranks every wrong one, and a dual-channel GRPO advantage keeps the price in its own units. From the same cold start, prompt pool, and budget, CounterCredit reaches 89.5% on V*, 80.2% on HR-Bench-4K, and 76.4% on HR-Bench-8K, 6.3 to 9.4 points above outcome-only GRPO at 1.78 against 1.84 calls per question, and lowers the spurious-call rate to 31% to 36%, the lowest among the agents evaluated. The same recipe lifts a Qwen3-VL-8B base from 75.4 to 80.8 on average.
Beyond Linear Context: Graph-Guided Evidence Navigation for Long-Novel Reasoning with a Local 9B Language Model
超越线性语境:基于图引导证据导航的长篇小说推理方法(使用本地9B语言模型)
Fu, Wenji
Abstract
Long-context models read a novel the way a person reads a printout: one token after another, in narrative order, with the whole history competing for a fixed budget of attention. A detective does not work that way. They sort what happened when, and they keep a map of who relates to whom, so a clue from chapter one can meet a question asked at the end of the book. We test whether a frozen knowledge graph can give a small local model that same freedom. Thirty detective novels and 234 multiple-choice questions are answered by one fixed qwen3.5:9b reader under nine conditions: five graph routes, a recent-window baseline, whole-book compression, ordinary vector retrieval, and a question-only control. The strongest graph route reaches 53.85% (126/234) against 46.15% for the recent window, 51.28% for compression, 51.71% for vector retrieval and 40.17% for question-only. On the subset that no model can answer without the book, the graph route reaches 42.86%. None of the fifteen graph-baseline contrasts survives Holm correction, so we present the result as exploratory evidence about a design. Two structural findings survive scrutiny better than the headline number: annotated evidence concentrates in the topological core of these graphs (2.35x enrichment, pooled), and the two graph-building pipelines differ so much in annotation coverage (16% versus 73% of clue paragraphs) that pooled accuracy alone would hide which bottleneck is being measured.
AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows
AgentRouter:面向成本最优多步骤智能体工作流的异构模型路由
Paul, Rudrendu Kumar, Nandy, Sourav
Abstract
Enterprise agentic systems that route every trajectory step to a frontier model waste 60-80% of their inference budget on subtasks that smaller models handle equally well. Existing routing solutions optimize single-turn query assignment but ignore a property unique to agentic workflows: subtask complexity varies widely within a single trajectory. A planning step may require frontier-class reasoning while a subsequent formatting step needs only a 7B model. We formalize step-level model routing as a sequential assignment problem over agent trajectories and propose AgentRouter, a lightweight classifier (12M parameters, <5ms overhead per step on an A100 GPU) that maps each trajectory step to one of four model tiers using five features extractable at routing time. Trained on 50,000 annotated agent trajectory steps spanning planning, coding, research, and data analysis tasks, AgentRouter achieves 72% cost reduction relative to frontier-only baselines, retaining 97.3% of frontier-only quality (less than 3% degradation in end-to-end task completion); per-step routing accuracy reaches 91% on minimal-complexity steps and 85% on efficient-tier steps, with 76-82% on the harder mid-range and frontier tiers. On the same benchmarks, RouteLLM and FrugalGPT (applied per-step) achieve only 31% and 44% cost reduction respectively, because their single-turn training signal misses trajectory-level quality dependencies.
Parkinson's disease alters gait and bilateral coordination, but machine-learning performance also depends on how continuous gait signals are represented. This study investigates whether preserving anterior-posterior center-of-pressure (AP-COP) information at fixed locations across normalized stance provides a compact and informative representation of plantar-force gait signals. Bilateral vertical ground reaction force recordings from 165 participants in the Gait in Parkinson's Disease Database were evaluated using repeated fully nested participant-level cross-validation. We propose AP-COP10, comprising AP-COP position and bilateral asymmetry across five stance windows. AP-COP10 achieved an AUC of 0.894 and outperformed three harmonized literature-derived COP representations under the same evaluation pipeline. The complementary 25 non-AP-COP descriptors alone achieved an AUC of 0.856, while the complete 35-feature representation achieved 0.908. Removing AP-COP10 from the complete representation produced a statistically supported loss in discrimination, whereas adding the complementary descriptors to AP-COP10 yielded only a small, unsupported improvement. Feature competition indicated that the most informative stance-indexed descriptors were concentrated in early and early-mid stance, while source-study holdout and sensor-perturbation analyses supported the robustness of the representation. These findings indicate that stance-indexed AP-COP retains discriminative information that is not readily recovered by broader engineered gait descriptors, supporting compact and interpretable representations for machine-learning analysis of pathological gait.
R-GEAN: Regimen-Guided Edit Action Network for Within-Admission Medication Change Prediction
R-GEAN:面向住院期间用药变化预测的方案引导编辑动作网络
Mahat, Regan, Kim, Mansu
Abstract
The medications prescribed to a patient often change during a hospital admission as clinicians start, stop, or continue therapies. We study whether models can predict which medication classes are added or removed between 24 hours after admission and discharge. Metrics that compare the complete discharge regimen can reward models for copying medications that remain unchanged, even when they identify no actual changes. We therefore introduce a leakage-controlled benchmark that predicts net ATC3 additions and removals using only prior completed admissions and information available within the first 24 hours of the current admission. Addition candidates are classes not active at 24 hours, whereas removal candidates are classes active at that time. We also introduce R-GEAN, an asymmetric candidate-scoring network with independent addition and removal predictors. Across 240,480 admissions from 82,286 patients, R-GEAN achieves the highest predefined summary of addition, removal, changed-regimen, and action-pattern performance, termed the edit composite (0.464), compared with 0.435 for the strongest primary comparator. Reimplemented RETAIN, GAMENet, and MICRON baselines obtain 0.428, 0.420, and 0.288, respectively. R-GEAN's advantage is concentrated in correctly identifying medication classes no longer active at discharge, while rare additions and admissions with multiple medication changes remain difficult. Rankings based on micro-F1 over the reconstructed discharge regimen and the edit composite correlate weakly across the evaluated models (Spearman r = 0.20). The continuation baseline achieves the highest complete-regimen score despite predicting no additions or removals. These results show that complete-regimen and edit-level evaluation measure different aspects of medication prediction. The benchmark evaluates observed prescribing changes, not treatment appropriateness
Chinese Translation
患者在住院期间,医生会开始、停止或继续某些治疗,因此处方用药常常发生变化。我们研究模型能否预测从入院24小时后到出院期间哪些药物类别被新增或停用。比较完整出院方案的指标可能会奖励那些照抄未变化用药的模型,即使它们并未识别出任何实际变化。因此,我们提出了一个防止信息泄漏的基准任务,仅利用既往已完成的住院记录以及当前住院前24小时内可获得的信息,来预测ATC3层面的净新增与停用。新增候选为在24小时时点尚未使用的药物类别,而停用候选为该时点正在使用的药物类别。我们还提出了R-GEAN,一种具有独立新增预测器和停用预测器的非对称候选评分网络。在来自82,286名患者的240,480次住院记录上,R-GEAN在预先定义的新增、停用、变化方案及动作模式性能综合指标(称为编辑综合指标,edit composite)上取得最高分0.464,而最强的主要对比方法为0.435。重新实现的RETAIN、GAMENet和MICRON基线分别获得0.428、0.420和0.288。R-GEAN的优势集中在正确识别出院时已不再使用的药物类别上,而罕见的新增以及多次用药变化的住院案例仍然具有挑战性。在重构的出院方案上计算的micro-F1排名与编辑综合指标的排名在所评估的模型之间相关性较弱(Spearman r = 0.20)。仅预测无新增和停用的延续基线却在完整方案评分上取得最高分。这些结果表明,完整方案评估与编辑层面的评估衡量的是用药预测的不同方面。该基准评估的是观察到的处方变化,而非治疗 appropriateness(治疗是否恰当)。
Automated operations research (OR) modeling requires LLMs to translate natural-language decision problems into correct mathematical programs. Existing methods can improve individual formulations, but they often solve problems in isolation, retaining little reusable experience and repeating similar formulation errors. Prior memory-based approaches store examples, thoughts, or insights as references, while OR modeling requires reusable formulation skills that transfer across problem narratives and guide concrete modeling decisions. We propose OptiSkill, a skill-augmented framework that builds a hierarchical and evolving SkillBank for LLM-based OR modeling. SkillBank stores solver-verified experience as reusable skills, with Global Strategies for problem-level formulation skeletons and Step Experiences for local error-prevention rules. It is further refined through stable batch-level test-time evolution, where candidate skills are incorporated only after validation. Experiments on eight OR modeling benchmarks show that OptiSkill improves formulation accuracy across LLM backbones, outperforms strong agentic baselines, and gains further by expanding SkillBank coverage and reliability. Code and data are available at https://github.com/rachhhhing/OptiSkill
Physics-informed neural networks (PINNs) require coordinated choices over network representation, sampling, loss construction, and optimization, while effective configurations often vary substantially across partial differential equations (PDEs). Existing automated PINN design methods can search candidate configurations, but information revealed during actual training is still used mainly for evaluation rather than to improve subsequent design, leading to repeated trial-and-error and inefficient use of training budget. We propose PINNsForge, an LLM-driven evolutionary framework for execution-feedback-based automated PINN design. PINNsForge generates diverse candidate configurations from PDE-related prior knowledge, evaluates them through actual training, and feeds high-performing designs together with accumulated execution evidence back to the LLM. Guided by observed optimization behavior, the LLM then refines, recombines, and explores coupled PINN design components, forming a continual cycle of generation, execution, feedback, and evolution. Unlike one-shot search or evaluation-only feedback, PINNsForge progressively converts training experience into improved design decisions for the target PDE. Across 25 PDE benchmarks, PINNsForge achieves the lowest mean MSE on 24 tasks compared with RoPINN, PINNsFormer, and PINNsAgent. Ablation studies further confirm the importance of the PDE knowledge base, execution feedback, and evolutionary search: removing these components increases the mean MSE to 3.74$\times$, 12.10$\times$, and 10.10$\times$ that of the full PINNsForge, respectively.
Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial state, yet existing VLMs remain limited in both capabilities. Current spatial training primarily focuses on static questions about object attributes and spatial relations, providing limited direct supervision for state transitions; in contrast, interaction trajectories naturally connect a preceding observation, an action, and a subsequent observation, offering direct supervision for local state transitions, while complete trajectories reveal dependencies among consecutive transitions. We therefore introduce Spatial-Interactor, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories. Accordingly, we construct the Learning from Spatial Interaction dataset (LSI-108K) from simulated and real interaction trajectories, with tasks aligned with the objective of each level. Our two-stage training strategy applies Supervised Fine-Tuning (SFT) to L1 and L2 for local transition modeling, and On-Policy Distillation (OPD) then uses privileged self-distillation: a teacher branch given segment-level transition descriptions supervises the student's on-policy CoT, helping the student learn to integrate consecutive transitions over L3 long trajectories. Experiments across multiple VLMs and spatial benchmarks show consistent gains in local transition modeling and long-horizon integration.
Enforcing Narrative Reliability and Epistemic Pacing in LLM-Driven Detective Games via Structured Knowledge Trees
通过结构化知识树在LLM驱动的侦探游戏中强制保障叙事可靠性与认知节奏
Rahmati, Parsa, Zhao, Richard
Abstract
Large Language Models (LLMs) enable open-ended dialogue in interactive games, but their non-deterministic outputs make it difficult to preserve authorial control, factual consistency, and the intended sequence of information disclosure. These challenges are particularly significant in detective games, where premature revelation or fabricated details can undermine the logic of player progression. We present a Structured Knowledge Tree architecture coupled with a tri-agent LLM pipeline for controlling dialogue in an open-ended interrogation game. The system separates knowledge retrieval, dialogue generation, and response verification to ensure that the virtual suspect reveals only information permitted by the current narrative state. We evaluate the approach through The Interrogation of Adrian Gale, a playable detective-game testbed, and a formal user study examining hallucination reduction, adherence to authored disclosure sequences, and perceived logical progression. Our results demonstrate that the structured architecture reduces critical hallucinations by 64.78% and entirely prevents premature narrative disclosure. While the strict mechanical constraints introduced usability trade-offs regarding forced conversational reveals, the system successfully enforces rigorous epistemic pacing and provides players with a clear, subjective sense of progression toward solving the case.
Chinese Translation
大语言模型(LLM)为互动游戏中的开放式对话提供了可能,但其非确定性的输出使得维持作者控制、事实一致性以及预期的信息披露顺序变得困难。这些挑战在侦探游戏中尤为突出,因为过早揭示真相或虚构细节会破坏玩家推进游戏进程的逻辑。我们提出了一种结构化知识树(Structured Knowledge Tree)架构,结合三智能体LLM流水线,用于控制开放式审讯游戏中的对话。该系统将知识检索、对话生成与回复验证相互分离,以确保虚拟嫌疑人仅透露当前叙事状态所允许的信息。我们通过《审讯阿德里安·盖尔》(The Interrogation of Adrian Gale)——一个可玩的侦探游戏测试平台——以及一项正式的用户研究对该方法进行评估,考察幻觉减少程度、对预设信息披露顺序的遵循情况以及玩家感知到的逻辑推进。结果表明,该结构化架构将关键幻觉减少了64.78%,并完全杜绝了叙事信息的过早披露。尽管严格的机械性约束在强制对话揭示方面带来了可用性上的权衡,但该系统成功实现了严谨的认知节奏控制,并为玩家提供了清晰的主观破案推进感。
LazyAgent: Demand-Driven Materialization and Physical Optimization of Agentic Programs
LazyAgent:面向智能体程序的按需物化与物理优化
Heng, Xin
Abstract
Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present LazyAgent, a unified execution framework for agent-authored programs organized around a live, goal-derived demanded set. LazyAgent refreshes a backward closure from requested outputs as execution state changes and materializes a ready node only when the active goal requires it. This replaces repeated local judgments with one linear-time graph analysis followed by constant-time membership tests, allowing programs to remain broad while execution stays request-specific. On programs that describe more than the current request needs, LazyAgent consistently outperforms the strongest goal-stopping eager baseline by refusing unrelated work before it starts. Adding one unrelated product raises the eager bill by 22.5% and LazyAgent's by 0.0%. LazyAgent saves 42.0% of measured CPU on production scientific workflows and 51.7% of container time on a live release gate spanning four repositories. We also prove and verify exact equivalence when the request reaches the whole graph, leaving no unrelated work to avoid. Beyond permission, goal-relative output projection saves up to approximately 90% of a shared step on two third-party test suites while the identical eager control saves 0.0%; the advantage disappears when the omitted output has no other consumer or the request needs it. Ordering, reuse, and pruning can also save cost, but do not replace permission. Finally, we show that current public benchmarks are eager-shaped and contain almost no unrequested work. A pre-registered planning intervention did not broaden them. These findings motivate benchmarks built from standing programs and sequences.
Understanding the physical world requires more than object recognition, scene description, and short-term visual prediction, as real-world physical systems involve multiple continuous fields, latent causal mechanisms, partial observations, and intervention-sensitive dynamics. We propose FireWorldBench, a benchmark for evaluating complex physical world intelligence in multimodal large language models and agents through coupled-field fire dynamics. Fire provides a canonical stress-test environment, where multiple interacting physical fields jointly shape observable states and temporal dynamics. FireWorldBench is organized along two complementary axes, a physical capability axis and a fire scenario task axis, jointly covering physical-state understanding, temporal dynamics, causal mechanisms, and intervention reasoning. The benchmark comprises 520 fire-world entries, including 494 controlled simulation worlds and 26 real-world-aligned event groups, spanning 47 scene archetypes across 7 environment families. These entries combine structured textual observations, multiple 2D physical-field visualizations, and 3D event-level scene modeling, yielding 9,074 text-image interleaved question-answer pairs across choice-based and open-ended report-generation formats. FireWorldBench evaluates whether models can infer latent physical states, explain underlying mechanisms, forecast coupled-field evolution, and assess intervention consequences from multimodal partial observations, providing a challenging testbed for complex physical world intelligence.
Tutoring Large Language Models to be Domain-adaptive, Precise and Safe
引导大语言模型实现领域自适应、精准与安全
Banerjee, Somnath
Abstract
This thesis proposes a framework for "responsible intelligence" to address AI's critical challenges in safety, ethics, and cultural sensitivity. It advances three core areas: First, it improves domain adaptation in specialized fields using active learning and graph-based knowledge to reduce hallucinations. Second, it enhances ethical rigor via a novel decoding-time alignment mechanism that proactively blocks harmful text generation in real-time. Finally, it ensures cultural and multilingual safety through language-specific steering that respects diverse linguistic and social norms. Ultimately, this work provides a blueprint for building next-generation AI that is contextually knowledgeable, ethically sound, and culturally adaptable.
Forecasters often know an event is imminent but not the shape, size, or timing of its effect. We introduce Event Signature Transfer (EST), a training-free, model-agnostic operator that turns a completed past event into an explicit forecast scenario. EST removes a source event's own trend and seasonality, then scales and retimes the remaining event signature onto a native forecast, preserving the forecast's linked structure and reducing to it exactly at zero strength. Because it reads only output quantiles, EST applies to any quantile forecaster, with no training, no model internals, at transfer time. Across twelve real episodes and ten synthetic scenarios on Chronos-2, TimesFM-2.5 and Toto-2.0, manually configured EST reduces real-episode WQL by 21.7-90\% in-sample. On Chronos-2, it leads eleven of twelve matched comparisons against covariate conditioning, activation editing and raw replay. The operator builds a scenario; it does not estimate its likelihood.
From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving
从推理引擎到推理控制平面:连接 vLLM、llm-d 与高效分布式大语言模型推理服务的演进
Sisodia, Twinkll
Abstract
Large language model (LLM) inference is evolving from an engine-local optimization problem into a distributed control problem involving reusable state, phase placement, heterogeneous accelerators, networking, autoscaling, reliability, and service-level objectives. This paper connects that transition across peer-reviewed systems research, open-source implementations, and documented production studies. It treats vLLM and llm-d as complementary layers: model-serving engines optimize execution through mechanisms such as PagedAttention, continuous batching, kernels, quantization, and parallelism, while an inference control plane can optimize where, when, and under what policy execution occurs across a fleet. The contribution is synthesis rather than a new benchmark; all reported performance and deployment results remain attributed to their original sources. The combined evidence suggests that the scarce resource in modern inference is shifting from raw FLOPs alone toward managed state, placement, network movement, reliability, and decision quality. We propose an Inference Execution Planner that selects feasible execution plans rather than only endpoints, including aggregated versus disaggregated topology, KV source and transfer action, hardware variant, routing/admission policy, and slower scaling decisions. We also provide a source-local benchmark atlas, a bottleneck-migration taxonomy, practical deployment guidance, an evaluation framework based on SLO-goodput, and research questions for agentic, multimodal, heterogeneous, and resilient inference.
CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine
CraftBench-UE:面向虚幻引擎中编程智能体的确定性评估
Wu, Shutong, Calderone, Kevin, Tsen, Andy
Abstract
Building gameplay features in a game engine requires more than code, as code that compiles and runs does not necessarily implement the requested gameplay. We introduce CraftBenchUE, an evaluation harness that runs agents in an isolated Unreal Engine environment, reconstructs their saved submissions in fresh projects, and applies deterministic build, asset, and runtime checks without an LLM judge. Based on the harness, we built a benchmark consisting of 70 tasks spanning C++ source, Blueprint assets, and editor scripting. We evaluate seven models under two editor-tool configurations, with a file-and-shell baseline on C++ tasks. We further pair tasks that specify the same gameplay and use the same runtime tests, but require C++ and Blueprint as the deliverables. Across the 10 paired tasks, C++ completion rates exceed Blueprint by 30.0 and 42.9 percentage points in the two tool configurations. Among on-time Blueprint submissions in this paired set that pass asset checks, 42.2% and 50.0% fail explicit runtime assertions. These submissions satisfy asset requirements but fail the required gameplay tests. We will release the harness, task benchmark, and our trajectory findings with the report.
Chinese Translation
在游戏引擎中构建玩法功能不仅仅需要代码,因为能够编译和运行的代码并不一定实现了所要求的玩法。我们提出了 CraftBench-UE,这是一个评估框架,它在隔离的虚幻引擎(Unreal Engine)环境中运行智能体,在新项目中重建其保存的提交成果,并应用确定性的构建、资产和运行时检查,而无需 LLM 评判者。基于该框架,我们构建了一个包含 70 个任务的基准测试,涵盖 C++ 源代码、蓝图(Blueprint)资产和编辑器脚本。我们在两种编辑器工具配置下评估了七个模型,并在 C++ 任务上采用文件与命令行基线。我们进一步将指定相同玩法并使用相同运行时测试、但分别要求以 C++ 和蓝图作为交付成果的任务进行配对。在 10 个配对任务中,两种工具配置下 C++ 的完成率分别超出蓝图 30.0 和 42.9 个百分点。在这一配对集合中按时提交并通过资产检查的蓝图提交中,分别有 42.2% 和 50.0% 未通过明确的运行时断言。这些提交满足了资产要求,但未通过所需的玩法测试。我们将随本报告一并发布该框架、任务基准以及我们的轨迹分析结果。
Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation
不要相信基准测试:通用大语言模型排名的局限性及任务专用评估的必要性
Amin, Danial
Abstract
Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretable only in relation to the system tested, the questions included, and the conditions of evaluation. This perspective examines five connected limitations of general LLM rankings: differences between evaluated and publicly available systems; commercial incentives and dependencies in external evaluation; benchmark saturation, defective tests, and data contamination; models exploiting scoring procedures; and the limited relevance of general scores to users' tasks. Documented cases illustrate why these problems require different responses. I argue for evaluation procedures that disclose the tested configuration, validate questions and successful task completion, report performance alongside cost and execution time, and make the scope of generalization explicit. I then discuss \textbf{Isotanta}, a crowdsourced benchmarking platform, as a practical example of contributed questions and repeated evaluation. A larger question pool may improve task coverage, while repeated sampling can improve the stability of estimates on that pool; neither guarantees validity or personalization. The paper distinguishes the platform's current shared ranking from proposed task-specific and user-provided evaluations. Its central argument is that model selection requires evidence about performance on the intended work, not simply a high position on a general leaderboard.
Expansion Counts under Standard A* Tie-Breaking Strategies on the Final Plateau
标准A*平局打破策略在最终平台层上的节点扩展次数
Fukunaga, Alex
Abstract
In the A* search algorithm, the tie-breaking strategies for nodes with the same $f$-value determines which states A* expands on the final $f$-layer. For nine standard tie-breaking strategies, we show that under a consistent heuristic, every pair has positive-cost instances favoring each strategy over the other by an arbitrarily large additive expansion gap. A parameterized unit-cost grid example also gives unbounded expansion-count ratios between low-$h$ with FIFO and LIFO. In unit-cost search with $h > 0$ at non-goals, exact heuristic values near the goal lead to complementary extremal results: low-$h$ minimizes the number of remaining expansions from a common configuration within the perfect region, while high-$h$ maximizes the total number of expansions when every final-plateau state with $h=1$ is a goal predecessor. Finally, with the evaluation function $f_{\alpha} = g + \alpha h$, when $h>0$ at non-goals, every heuristic weight $0 \leq \alpha<1$ eliminates tie-breaking sensitivity, and all tie-breaking strategies expand the same set of states.
Chinese Translation
在A*搜索算法中,针对具有相同$f$值的节点的平局打破策略决定了A*在最终$f$层上扩展哪些状态。对于九种标准平局打破策略,我们证明了在一致性启发式函数下,对于任意一对策略,均存在正代价实例使得其中一种策略相对于另一种策略在扩展次数上具有任意大的加性优势。一个参数化的单位代价网格示例还表明,低$h$与FIFO组合和低$h$与LIFO组合之间的扩展次数比值是无界的。在非目标节点满足$h>0$的单位代价搜索中,目标附近精确的启发式值导致互补的极值结果:在完美区域内,从同一配置出发,低$h$最小化剩余扩展次数;而当每个$h=1$的最终平台层状态均为目标的前驱时,高$h$最大化总扩展次数。最后,对于评价函数$f_{\alpha} = g + \alpha h$,当非目标节点满足$h>0$时,任意启发式权重$0 \leq \alpha<1$均可消除对平局打破策略的敏感性,即所有平局打破策略扩展的状态集合相同。
Recent advances in large language models (LLMs) have led to the emergence of coding agents capable of performing complex engineering tasks, including register-transfer level (RTL) design and optimization. Existing RTL benchmarks mainly evaluate functional correctness and performance, power, and area (PPA) of the generated RTL designs, leaving agents' ability for \emph{timing closure} under-evaluated. We propose TicTacBench, a benchmark specifically designed to evaluate coding agents' capabilities for RTL-level timing closure under post-place-and-route (post-PnR) evaluation. TicTacBench contains 30 diverse tasks, each provided with a suboptimal RTL design, realistic timing constraints, functional equivalence verification, and timing reports. With over 300 runs of coding agents driven by 8 frontier LLMs, we find that even the best agent can only close 53.3\% of tasks with 7.18\% area-delay product (ADP) degradation and 8.83\% energy-delay-squared product (EDDP) improvement on average. We identify common failure categories that explain why agents fail to close timing. Then we propose TicTacSkill, a new method that guides agents to follow standard timing-closure procedures and improves the Timing Closure Rate by 9\%. These results suggest that while coding agents have made significant progress in RTL design, their timing-closure capability still has substantial room for improvement.
Leaky-integrator reconstruction: taming error accumulation in recursive differenced time-series forecasting
泄漏积分器重构:抑制递归差分时间序列预测中的误差累积
Yang, Zijiang
Abstract
We introduce leaky-integrator reconstruction, a training-free method that cures the error accumulation of recursive differenced forecasting. Our first contribution is diagnostic: predicting one-step changes and integrating them by cumulative summation, the standard remedy for non-stationarity, is a discrete integrator with a pole on the unit circle, and we show this makes recursive rollout of a nonlinear model diverge, its 336-step error reaching several times that of a well-behaved forecaster (normalised MAE 1.6-3.8 versus about 0.8) across every neural architecture tested. Our second, central contribution is the fix: move the pole inside the unit circle with a leaky integrator H(z) = 1/(1 - gamma z^-1), gamma < 1, which provably bounds the accumulated error variance. Applied at reconstruction time with a single fixed gamma=0.9 (no retraining, a two-line change to any deployed one-step or foundation-model forecaster), it shrinks error at every horizon, the mean gain over seven diverging architectures and twenty datasets growing from ~3% at H=24 to 23% at H=96, 37% at H=192 and 51% (43-74% across those architectures) at H=336 (78% with an oracle pole). Crucially, it is provably inert where no pathology exists (stable or joint predictors already at the irreducible rate), making it a safe, general default.
AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction
AgentBetta:通过选择性扩展与验证性收缩实现验证驱动的AI纳米智能体自适应配置
Babu, Md. Ashraful
Abstract
Large language model agents are typically deployed with predefined configurations, although the required model capability, context, tools, permissions, memory, and computational resources can vary substantially across tasks. This study develops and evaluates AgentBetta, an adaptive AI Nano-Agent framework that represents these factors as an executable configuration and updates them through verification-driven diagnosis, selective expansion, and verification-based counterfactual contraction. The evaluation distinguishes controlled mechanism validation from external agent comparisons. On the AB-ConfigBench benchmark, AgentBetta achieved 91.38% verified success while reducing median context allocation from 64,000 to 8,000 context characters and median tool exposure from five tools to zero compared with the fully provisioned configuration. The configuration-deficiency diagnosis achieved a macro-F1 score of 0.819 with precision of 1.000 across the evaluated dimensions, and selective expansion avoided unnecessary changes to unrelated configuration dimensions. Post-success contraction preserved verification outcomes in 56.41% of evaluated one-dimension contraction probes, indicating that some successful configurations contained removable capability under the tested conditions. External evaluations indicate that adaptive configuration can improve the balance between verified task completion and capability exposure; however, the results vary across benchmarks and agent families. In particular, the cross-family replication did not reproduce the primary-backbone accuracy ordering, and specialized systems remained advantageous for certain task domains. These results support interpreting AgentBetta as a configuration-adaptation mechanism that regulates capability allocation and inference expenditure rather than as a universal replacement for specialized agent architectures.
Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We distinguish alignment with human preferences from alignment with human behavior, and show that alignment with human preferences can make model behavior less human-like even when both preferences and responses come entirely from humans. We call this the Turing-test gap. We show that preference alignment preserves the human response distribution only under a restrictive condition, and find no consistent evidence that real human preferences satisfy it. Empirically, the loss of human-response likelihood increases with the strength of preference weighting, regardless of its direction, and the gap also appears under standard DPO. These results establish human-likeness as an explicit dimension of alignment rather than something assumed to follow from preference alignment.
Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as unmanned aerial vehicles (UAVs), which must perceive, reason, plan, and act in dynamic environments. Existing benchmarks assess physical perception, intuitive physics, embodied navigation, and collaborative reasoning, but rarely evaluate the agentic decision-making required for reliable autonomy. We introduce \textit{PhysAI-Bench}, a benchmark for evaluating this capability. It contains 10,178 standardized decision instances automatically extracted from conversational traces of autonomous UAV missions. Each instance preserves mission context, temporal dependencies, physical constraints, Model Context Protocol (MCP) tool calls, Agent-to-Agent (A2A) interactions, sensor observations, and AI-native 6G network conditions, including latency, packet loss, throughput, edge load, and network slicing. We expose only information preceding each decision, preventing future-event leakage and approximating online decision-making. We evaluate 29 foundation models using a two-stage protocol. We select model-specific configurations from 12 combinations of zero-, three-, and five-shot prompting and four temperatures, tested in three runs on a 35-instance, human-verified development set. We then freeze each selected configuration and evaluate it in three runs on a fixed, episode-disjoint set of 500 instances. GPT-5.3 achieves the highest accuracy (52.00%), followed by GPT-5.2 (49.40%) and Grok~4.5 (49.07%). Few-shot prompting generally improves performance, while temperature has limited influence. The results demonstrate that reliable agentic decision-making in Physical AI remains an open challenge. The dataset is available at https://github.com/maferrag/physai-bench
Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim assessment. Most existing systems, however, are organized around individual tasks: the same papers are repeatedly retrieved, segmented, and interpreted, and the understanding built in one task is difficult to reuse in the next. We present ScholarStack, a layered research asset framework that compiles a paper collection into reusable, versioned, and provenance-preserving assets at three complementary levels: source-grounded paper-level statements, domain-level organization, and evidence-grounded cross-paper syntheses. A common access interface returns task-specific views at the evidence granularity each task requires, preserving study conditions, source traceability, and verification status. We instantiate the framework on four task families spanning ten task settings, comparing agents that use the compiled assets with task-specific baselines under matched base models. Quality gains concentrate on tasks that require cross-paper evidence, such as multi-paper question answering and literature review generation, and query-time token cost falls on every task where it is measured, with assets compiled once and reused across tasks. These results suggest that layered research assets can serve as shared infrastructure for scientific agents, shifting literature-based assistance from isolated document processing toward cumulative, evidence-grounded workflows.
On Probabilistic Inference Through Parametric Tensor Decomposition in Base Tensor Networks
基于基张量网络中参数化张量分解的概率推断研究
Hamid, Sagad, Braun, Tanya
Abstract
Probabilistic inference is generally only tractable in low-treewidth graphical models, limiting its effective applicability in high-treewidth settings. Many existing methods improve efficiency by exploiting specific parametric structure, such as symmetries. However, they typically require such structure to be explicitly present, limiting their applicability to a broader range of graphical models. To address this limitation, we propose a framework where tractable inference is controlled by latent parametric structure exploitation, rather than requiring it to be explicitly present a priori. Our approach first reparameterises a graphical model as a specific tensor network representation, which we call a base tensor network. This representation yields two key properties that allow inference tractability to be controlled by parametric structure: 1) First, the complexity of inference is mainly determined by the parametric structure of a single tensor, called the base tensor. We characterise several tractable classes of base tensors for which the entire base tensor network can be contracted efficiently. 2) Second, decomposing the base tensor yields again a collection of base tensor networks. This allows inference to be naturally reduced to decomposing the base tensor into tractable components with sufficient parametric structure. We call this procedure parametric tensor decomposition. By exploiting parametric structure within the base tensor, our framework enables a novel view on inference beyond settings where such structure is explicitly present.
Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected tokens are billed as input tokens at the same per-token price as the system prompt and the user query. Production observability tools report total token cost but do not separate the tokens a node generates from the tokens it is handed, so this component of the bill is invisible to the teams paying it. We introduce the Total Cost of Agency (TCA), a decomposition of multi-agent workflow cost into base prompt, inference, memory injection, miss penalty and context-accumulation components, and an exact attribution method: a two-pass, non-billable token count that measures injected tokens directly rather than estimating them from word-count proxies. On a 200-task enterprise benchmark executed against real model APIs, memory injection accounts for 13.6 percent of the variable cost a compile-time optimizer can act on, about 12 percent of the full billed cost, and its share rises from a structural zero at workflow depth one to 27.6 percent at depth six. Injected tokens grow linearly with depth over the measured range (R^2 = 0.9974, depths two through six); a quadratic fit yields a negative leading coefficient, so the data do not exhibit convex growth at these depths. We show the component is controllable at fixed model tier: reducing the retrieval window capacity from 32 to 2 entries lowers injected tokens by 28.7 percent with an accuracy change within seed-level variation. We report in full that our graph-rewriting transforms are approximately cost-neutral in isolation, that two of the five decomposition terms are zero by construction in this harness, and that total workflow cost is dominated by model tier assignment, which we hold fixed and treat as prior work. Prompt caching is not evaluated; all figures are for the uncached case.
Chinese Translation
多智能体大语言模型(LLM)工作流中的每个节点都会从记忆中检索上下文并将其注入到提示词(prompt)中,这些被注入的令牌(token)与系统提示词和用户查询按相同的单价计费。生产环境的可观测性工具只报告总令牌成本,却不区分节点自身生成的令牌与被传递给它的令牌,因此账单中这一部分对支付团队而言是不可见的。我们提出代理总成本(Total Cost of Agency, TCA),将多智能体工作流成本分解为基础提示词、推理、记忆注入、未命中惩罚和上下文累积等组件,并提出一种精确归因方法:一种两遍式的、不计费令牌计数方法,直接测量注入的令牌,而非通过字数代理进行估算。在一个针对真实模型API执行、包含200个任务的企业基准测试中,记忆注入占编译期优化器可作用的可变成本的13.6%,约占全部计费成本的12%,且其占比从工作流深度为一时的结构性零上升到深度为六时的27.6%。在测量范围内,注入令牌随深度呈线性增长(R^2 = 0.9974,深度二至六);二次拟合的首项系数为负,因此数据在这些深度上并未表现出凸性增长。我们表明该成本组件在固定模型层级(model tier)下是可控的:将检索窗口容量从32条降低至2条,可使注入令牌减少28.7%,而准确率变化在随机种子波动范围之内。我们完整报告如下事实:我们的图重写变换在孤立情况下近似成本中性;五个分解项中有两项在该实验框架中由构造为零;工作流总成本由模型层级分配主导,我们将其固定并视为先前工作。本文未评估提示词缓存(prompt caching);所有数据均针对无缓存情形。
WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks
WorkWorlds:用于评估AI智能体职场任务的基础设施
Hua, Yining, Lian, Levi
Abstract
Many knowledge-work benchmarks are constructed around individual tasks, with the context needed for each task selected together with or after the task has been specified. This design measures performance on workplace-like tasks in an environment assembled for the task. When task specification guides which context is selected, the evaluation can encode task information into the environment and pre-complete part of the information-localization work that workplace performance normally requires. We introduce WorkWorlds, an evaluation infrastructure that separates organizational state from task specification. A world first fixes a revision, date, and employee seat and materializes the organizational state that employee can access; tasks are introduced only afterward. We implement WorkWorlds in a primary synthetic pharmaceutical company with 8 measured tasks across 6 employee seats, and construct additional organizational worlds. Across 192 matched evaluations, moving from task-curated context to the full role-visible workplace reduced evidence access from 90.4% to 74.5% and criterion pass from 79.4% to 68.2%, while pass conditional on evidence access remained nearly unchanged; most of the measured difference occurred before the agent reached sufficient evidence.
Pretraining of Medical Visual Encoders Toward Multi-modal Large Language Models
面向多模态大语言模型的医学视觉编码器预训练
Jiang, Tianyou
Abstract
Multimodal Large Language Models (MLLMs) commonly reuse visual encoders pretrained with CLIP, although the features of these ViTs are ultimately consumed by autoregressive LLMs. We refer to this mismatch as the semantic-interface gap and introduce MedMLIP, a framework that pretrains the visual encoder through report generation with a frozen LLM, while employing Local Relational Distillation (LRD) to preserve relationships among visual patches to avoid visual collapse. We pretrain MedMLIP on IU-Xray and Open-PMC-300K and evaluate the resulting encoders on VQA-RAD and SLAKE. Only the ViT is transferred, while the guiding LLM and projector are replaced, allowing us to assess cross-LLM transferability. Our cross-LLM transfer experiments demonstrate the value of pretraining visual encoders for their autoregressive LLM interface while trying to preserve more fine-grained visual information. Code and the pretrained model are available at https://github.com/SkyCol/MedMLIP
Modern music streaming platforms face a persistent tradeoff: exploiting familiar content versus driving the exploration of novel items. While users frequently desire discovery, they hesitate to select unknown artists over proven favorites. Providing transparent, natural language rationales that explain why an unexplored item is recommended lowers this barrier. However, while Large Language Models (LLMs) excel at this nuanced explainability, their real-time deployment is severely bottlenecked by prohibitive inference costs and computational overhead. In this paper, we present an industry case study of a decoupled recommendation architecture that successfully scales exploration without compromising latency. Our system isolates LLM inference asynchronously offline, pre-computing personalized candidate pools of undiscovered artists alongside tailored rationales. Large-scale online A/B experiments validate our design. We demonstrate that combining LLM-backed recommendations with these explanatory rationales significantly reduces the trust barrier for new content, yielding statistically significant improvements in both user exploration and overall engagement on the discovery surfaces.
Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer
提升技能水平会招募冻结国际象棋Transformer中更深的注意力层
Litman, David
Abstract
Chess involves complex reasoning in a deterministic environment, which makes it a useful setting for studying the mechanisms of computation inside transformers. The Maia-3 chess transformer takes Elo, a measure of competitive chess skill, as an input to the pre-trained network, so we can vary the skill the network is conditioned on with no change to its weights. Here we investigate how turning this skill dial affects self-attention. Ablating every attention head at every Elo from 700 to 2500, we find 1) increasing skill pushes the causal center of mass of the computation deeper, monotonically, for every chess piece and move type we measured; 2) the depth migration is much greater for specific tactics, especially knight forks, than for other move types; 3) the migration consists of deeper heads getting recruited for more specialized computations while one shared shallow head keeps a roughly constant contribution. These results may shed light on how conditioning inputs redistribute computation in larger transformers.
Echo State Network (ESN) for Signal Recovery in RF-Impaired IBFD MIMO Systems
面向射频损伤IBFD MIMO系统信号恢复的回声状态网络(ESN)
Prisby, Conrad, Li, Siyao, Xu, Chengtao, Yang, Thomas
Abstract
In-band full-duplex (IBFD) multiple-input multiple-output (MIMO) systems enable simultaneous transmission and reception on the same frequency band, improving spectral efficiency for next-generation wireless networks. However, IBFD-MIMO systems are susceptible to self-interference (SI), which may overpower signals of interest (SOI). In this scenario, blind source separation (BSS) algorithms can be adopted to remove SI and perform joint sensing and communication (JSAC), but BSS algorithms mostly assume an idealized linear and quasi-stationary signal model, which does not hold under realistic radio frequency (RF) impairments, such as I/Q imbalance, carrier frequency offset (CFO), phase noise, and power amplifier nonlinearity. This paper proposes a two-stage echo state network (ESN)-based scheme that is superior to BSS under these realistic conditions. A frozen ESN is trained offline to characterize the static SI path, while an adaptive ESN, updated online via recursive least squares, tracks the time-varying SOI path using sparse pilot symbols. We evaluate the proposed scheme's SOI recovery performance and acquisition speed with different block sizes, comparing it against other recurrent neural networks (RNN), such as long short-term memory (LSTM) and gated recurrent unit (GRU). Simulation results show that the proposed approach outperforms BSS, LSTM, and GRU in both efficiency and SOI recovery, demonstrating the viability of ESNs for real-time, nonlinear self-interference cancellation in realistic IBFD MIMO systems.
AI agents that carry a multi-step computer task through on their own became ordinary tools in the past year, and the same autonomy is available to anyone whose task is harmful. We ask what that means for a relying party -- an insurer, a lender, an auditor -- whose evidence is a filed PDF. AgentForge-Bench measures how reliably an off-the-shelf coding agent, driving one of seven open-weight models with a shell and the stock Python PDF stack, alters one dollar amount, date or address in a real filed financial document from a single sentence of intent, graded by rules rather than by a model. Across 1,750 cells, 1,419 (81.1%) satisfy the verifier, and 808 (46.2%) also survive every stricter filter: visible, localized, typeface-matched, original value gone document-wide. A deterministic script with no model in it solves 98 of the 125 documents; the agents solve 124, and none the script solves alone. Agents misreport 41% of their wrong edits as done, no model refused, and the cheapest verified forgery costs 2.4 cents. The raw rate overstates the threat by about a factor of two; the strict rate is still large.
Divergent strategies and convergent outcomes in autonomous materials discovery
自主材料发现中的策略分化与结果趋同
Kim, Jihan
Abstract
Scientific agents are mostly evaluated on whether they complete tasks or recover known results; we instead study variation across repeated open-ended campaigns. Sixteen separately initialized sessions of one model-harness configuration received a frozen database of 12,499 metal-organic frameworks, a methane-storage objective, a pinned protocol and a one-week budget. Strategies diverged into four approaches spanning 100--5,000 screened structures, and eight built 2,253 hypothetical structures. Yet the agents recovered the same materials frontier near 200 cm^3/cm^3, and an independent calculation of the database's porous region found its nine best structures all among their reports. Enforced checks on half the agents raised fresh-run reproduction from one of eight to eight of eight but could not detectably improve conclusion validity, because fifteen of sixteen agents selected the same audit-excluded entry, an incomplete structure whose missing anions created artificial pore volume. Replicated agents thus reveal both robust conclusions and common-mode errors from shared inputs.
UniK: Universal Knowledge Perception for Digital and Physical AI
UniK:面向数字AI与物理AI的通用知识感知
Desai, Nirmit, Sawarkar, Kunal, Mahakali, Aditya, Lee, Dongkon, Park, Kevin, Song, Eric
Abstract
Two transformative classes of AI systems are reshaping how organizations operate: \textit{digital AI}, which reasons over enterprise knowledge to power chatbots and agent workflows; and \textit{physical AI}, which learns to control robots and autonomous systems from video, gameplay, and sensor telemetry. Both face the same foundational bottleneck: raw knowledge at scale, spanning heterogeneous modalities, locked in private corpora that existing AI infrastructure cannot access reliably or efficiently. We propose \textit{Universal Knowledge Perception (UniK)} as a common platform for both classes, covering the full knowledge lifecycle (ingestion, enrichment, indexing, retrieval, and continuous evaluation) across modalities from rich text and video to molecular data and sensor telemetry. We present UniK, built on Polymath Retrieval (multi-index fusion over automatically enriched indices) with no task-specific fine-tuning. Across five digital AI domains (medical literature, open-domain QA, chemistry, legal video proceedings, and government open data) UniK combined with an open-source 70-billion-parameter model consistently matches or outperforms frontier proprietary LLMs that are orders of magnitude larger: 76\% RAG accuracy on government data versus 47\% for GPT-5; 77.9\% on medical QA without fine-tuning; topping all open-source chemistry pipelines. We show that the same infrastructure directly addresses the data curation, indexing, and retrieval challenges facing physical AI world model training, where the knowledge problem is harder but structurally identical.
Foundation models are endowing autonomous systems with greater intelligence, enabling a more comprehensive understanding of the environment through visual perception. A representative example is Human Mesh Recovery (HMR), which provides useful estimates of a target's 3D pose and shape that can benefit tactical missions. However, the size and power demands of such models make them difficult to run on edge platforms and limit their real-time performance, undermining the requirements of tactical edge deployment - especially for active perception, where a mobile robot must plan its next-best view on-board and cannot offload computation under contested communications. We present LEAP-NBV, a lightweight active-perception framework that runs foundation-model-driven Next-Best-View (NBV) planning on-board an edge device. To this end, we distill a family of large HMR teachers, each into a compact 32M student, with an offline mesh objective, then quantize the vision encoder to FP16 and characterize its on-device accuracy and latency. Within an occlusion-aware active perception loop, we evaluate all configurations on the same held-out benchmark and deploy the end-to-end pipeline on an NVIDIA Jetson Xavier NX, reporting measured on-device latency and energy. Distillation recovers 6-7 mm of Procrustes-aligned mean per-vertex position error (PA-MPVPE) over the undistilled student on the test set. Selecting the edge-optimal compression model brings the HMR engine to ~12 ms at a small accuracy cost and runs the full closed loop at 3.6 FPS and 2.6 J per frame, achieving a 2.0x speedup and 3.0x lower energy than the uncompressed model while nearly matching downstream task quality.
Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents
Jev-Mem:面向高效AI智能体的系统一控制式智能体记忆
Jiang, Dongming, Li, Yi, Li, Bingzhe
Abstract
Agentic memory is becoming essential for long-horizon AI agents, yet many existing systems rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expensive generation on the critical path of memory operations. We introduce \textbf{\method}, a new agentic memory architecture inspired by System-One/System-Two cognition. System One captures fast, lightweight decision-making, whereas System Two performs slower, deliberative reasoning. Jev-Mem brings this division of labor to agentic memory through a dedicated System-One control plane, a structured multi-relational memory plane, and a System-Two reasoning plane. The System-One controller governs memory typing and relational organization during construction, and dynamically performs query routing, retrieval-budget allocation, graph traversal, candidate scoring, and adaptive stopping during retrieval. System Two is invoked only for complex reasoning and answer synthesis. This design improves both memory effectiveness and system efficiency: on LoCoMo Jev-Mem achieves an overall LLM-as-a-Judge score of 0.777, an 11.0\% relative improvement over the strongest baseline, while reducing memory construction time to 158\,s, a 6.6$\times$ speedup over the fastest competing memory system, and lowering average query latency to 0.93\,s, a 36.7\% reduction.
Building general-purpose agents for industrial deployment requires integrating multiple capabilities, each typically acquired at a distinct stage of training. Yet there is currently no well-established recipe for Agent Continual Learning (ACL), with little understanding of the trade-offs among existing integration paradigms. To address this gap, we introduce ACLArena, a framework for comprehensively studying, analyzing, and evaluating ACL. We first build a sequential training pipeline and conduct an in-depth analysis that explains the mechanisms of forgetting and generalization from two complementary perspectives, the model level and the token level. Guided by these analyses, we systematically compare multi-teacher on-policy distillation, self-distilled fine-tuning, and model merging to assess their ability to recover previously learned capabilities while preserving newly acquired ones. Through extensive experiments, we develop a detailed understanding of how capabilities transfer across stages. Finally, we propose a new ACL recipe that combines offline replay over high-quality trajectories with a routed network of multiple LoRA experts each specialized via RL, substantially improving the agent's ability to learn across multiple domains. Comprehensive experiments on four reasoning and agentic tasks, evaluated under both in-domain and out-of-domain settings, demonstrate the value of our analysis and the effectiveness of our approach.
FinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering
FinInteract:面向模糊金融问答中澄清与意图融合的基准测试
Wang, Xinyu, Kwok, Tung Sum Thomas, Tai, Zhenghan, Cheng, Guang
Abstract
Large language model agents increasingly answer financial questions by searching regulatory filings. Such questions are often deceptively under-specified: Meta Platforms' "operating income" is $46.75B consolidated but $62.87B for the Family of Apps segment, and each reading is exactly verifiable against the filing. A capable agent should recognize the ambiguity and ask, rather than commit to a plausible but unintended reading. Existing financial benchmarks cannot measure this, because one gold answer per question cannot separate agents that resolve the ambiguity from those that guess the common reading, a blind spot we call the single-gold illusion. We release FinInteract, a bilingual (English/Chinese) benchmark of 173 instances that pairs each question with a default and an intended interpretation across a five-category ambiguity taxonomy, and grades whether an agent elicits the right clarification and then integrates it. Re-grading identical outputs against the default rather than the intended reading inflates GPT-4o's accuracy by 3.1 times, confirming the illusion. Beyond it, we find that models answer above 90% once the interpretation is supplied but at most 28.9% when they must elicit it themselves, that targeting is uneven across a taxonomy well powered for entity scope and metric definition and exploratory elsewhere, and that conditioning on the ambiguity category improves resolution at both inference and training time.
Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure
检验而非假定充分性:将生成式社会模拟器与涌现网络结构进行校准
Shao, Tengfei, Li, Chao, Wang, Xu, Goto, Masayuki
Abstract
Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check. We present an adequacy-aware calibration protocol that couples amortized posterior estimation with a synthetic identifiability assessment, a matched-sample-size adequacy check (prior-predictive reachability plus per-statistic posterior-predictive localization), a diagnosis-guided repair, and a statistic-held-out audit. We demonstrate it on a real second-hand luxury resale market with four channel-by-residency cells, each a bipartite buyer-brand network, using a forward model built from persona profiles elicited once, offline, by a language model. The behavioural parameters are recoverable in all four cells, though calibration is approximate and overconfident for one parameter. The observed summary falls outside the simulator's reachability reference in every cell, with the mean purchased tier as the pervasive discrepancy. The repair meets the value-block criterion in two of four cells but does not restore adequacy, and the held-out audit surfaces a buyer-breadth-dispersion miss no earlier diagnostic detected. A profile-source ablation finds the language-model profiles beat a flat rule baseline in all four cells, yet within-category brand relabelling causes no consistent degradation, so the profiles are a partially validated input whose value rests on structure, not brand identity. Making no causal claim, we conclude that an independent-aggregation account, without agent interaction or a buyer-breadth mechanism, cannot jointly reproduce the market's purchased-tier level, head-brand concentration, community structure and buyer-breadth heterogeneity.
Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech
情境感知的人工智能系统部署前评估:尼日利亚金融科技的监管框架
Uduimoh, Andrew Anogie, Yusuf, Hadiza Umar, Osho, Oluwafemi
Abstract
Commercial large language models are increasingly deployed across African fintech infrastructure for fraud detection and customer communication, yet no Nigerian or African continental regulatory instrument specifies what pre-deployment evaluation such systems must undergo before procurement. This paper reviews African fintech AI governance across global, continental, and Nigerian instruments, and shows that safety is affirmed as a principle while pre-deployment evaluation is operationally unspecified. Generic safety benchmarks cannot surface the failure modes most relevant to this domain, since none contain Nigerian institutional content or test for false positive misclassification of legitimate financial communications. These claims are demonstrated using SafeAlert, a purpose-built evaluation kit applied to six commercial models across three system prompt conditions. Results show that models resisting generic harmful content requests still produce complete fraud scripts under specific framing, and that several models misclassify most legitimate Nigerian bank communications as suspicious or fraudulent, a failure invisible to standard safety evaluation. The paper concludes with a regulatory framework proposing pre-deployment evaluation requirements for the CBN, NITDA, SEC, and the AU, arguing that the identified gap reflects an absence of regulatory specification, not a shortage of technical or financial resources.
Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search
基于程序化策略搜索的连续博弈反应式角色行为合成
Gumin, Maxim, Liu, Hsueh-Ti Derek, Zordan, Victor, Ritchie, Daniel
Abstract
We present a method for synthesizing reactive character behaviors for continuous games as compact, human-readable programs. Game AI practice still relies heavily on manually authored behavior trees, state machines, and scripts, while academic reinforcement learning typically produces opaque neural controllers that are expensive to train and difficult to edit. Our approach bridges this gap by searching directly over a domain-specific language for continuous-space game policies. The language is designed around reactive geometric decisions and includes higher-order constructs such as direction maximization. These constructs help discretize a continuous behavior space into enumerable program structures. To make program search practical, we introduce a large set of synthesis antipatterns that remove redundant program forms while preserving behavioral coverage. We further combine bottom-up symbolic enumeration with top-down guidance from a coding agent. Our resulting method, agentic sketching, has the agent propose high-level policy structure and call an enumerator to complete local program slots. We evaluate the method on a benchmark of 14 continuous games, ranging from classic control tasks to multi-agent football. We find that pure enumeration is often more efficient than using a coding agent alone, while the combined method substantially outperforms both. Our results suggest that programmatic policy search can be a practical authoring tool for game AI: designers specify reward functions, and the system discovers editable behaviors that are effective, portable, and often surprising.
Structured Decomposition for Reliable LLM-Generated Access Control Policies
面向可靠的LLM生成访问控制策略的结构化分解方法
Gupta, Vatsal, Sreenivasamurthy, Darshan
Abstract
This paper presents an LLM-based system that translates natural-language access control policies (NLACPs) into executable Rego code for Open Policy Agent (OPA). It provides a modular, end-to-end pipeline for policy detection, component extraction, schema validation, linting, compilation, and automated test generation and execution. The system is designed to bridge the gap between human-readable access requirements and machine-enforceable policy-as-code (PaC), with a focus on deployment reliability and security correctness. We evaluate the system on 372 ACRE-complete access control statements with non-null subject, action, and resource annotations against a direct single-prompt LLM baseline to isolate the contribution of structured decomposition and schema-aware validation. The system achieves a 50.3% end-to-end policy correctness rate, compared with 15.3% for the baseline, representing a 3.3x improvement. A policy is counted as correct only if it satisfies compilation, linting, and both positive and negative tests, making this a strict measure of deployable correctness. On security-critical patterns, the system generates correct deny semantics for 87.5% of deny policies (baseline: 37.5%), ownership conditions for 100% of ownership-qualified policies (baseline: 40%), and status-qualified conditions for 100% of status-qualified policies (baseline: 55.6%). These results indicate that structured decomposition and schema-aware validation play a critical role in improving the reliability of LLM-generated authorization policies.
Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires resource-intensive domain-specific fine-tuning. Here we introduce representation-guided in-context learning (RG-ICL), a training-free inference framework that retrieves query-aligned demonstrations using frozen encoders, without task-specific parameter updates. Across eight datasets spanning histopathology, radiology and retinal fundoscopy, RG-ICL improved classification (mean gain 20 percentage points) and visual question answering (VQA) (mean gain 13 percentage points) over no-context and conventional ICL, approaching or exceeding training-based comparators. Which cases were retrieved mattered more than how many: 6 query-aligned cases outperformed up to 32 randomly selected ones, whereas fixed or random cases often reduced accuracy below baseline. For VQA, aligning reference cases with both image content and question intent produced further gains. These findings indicate that for medical image interpretation, curating which reference cases an MLLM sees is a practical alternative to retraining it.
Long-horizon autonomous intelligent systems rely on heterogeneous components such as large language models, databases, external APIs, and rule engines, while their external states continuously change during execution. Re-executing the entire workflow after every change introduces substantial redundant computation. This paper proposes an incremental consistency execution method based on task fact contracts, field-level dependency masks, and state perturbation result invariant domains. After an initial verified execution, the system constructs conservative invariant domains for critical inputs and uses them to determine whether downstream results can be safely renewed without re-invoking expensive components. When re-execution is required, only the smallest affected output fields are recomputed, and an equivalence barrier prevents unnecessary downstream propagation. A submission-time version consistency gate further ensures the safety of side-effecting actions. Experiments on industrial fault diagnosis, enterprise analytics, and LLM-based multi-tool assistants show that the proposed method significantly reduces expensive component calls and end-to-end latency while maintaining high consistency and low incorrect-reuse rates.
DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents
DocMIDE:在视觉丰富文档中学习多跳隐式推导
Wang, Jeremy Cerwin, Wong, Wai Kit, Tang, Jeff Kai Tai
Abstract
Real-world document processing systems rely on rigid, predefined schemas, yet critical target fields often lack direct visual counterparts on the page. Extracting these implicit values requires multi-hop derivation, such as aggregating sub-categories or reasoning over visual marks. While existing methods handle explicit text spans or simple implicit queries, they fail at multi-hop visual reasoning even after standard fine-tuning: models retrieve incorrect visual evidence, or retrieve it correctly and then skip the intermediate steps of the derivation. To address this, we introduce DocMIDE, a fine-tuning framework that trains compact vision-language models to retrieve visual evidence explicitly before deriving an answer. DocMIDE constrains generation to a plan-retrieve-derive structure and optimizes it with Group Relative Policy Optimization under a four-component, rule-based reward that scores output format, the retrieved evidence block, every intermediate derivation step, and the final value against a verified reference trace. On a 4,151-pair implicit extraction benchmark, DocMIDE raises accuracy from 70.8% to 95.9% on Qwen3.5-4B from only a small set of annotated examples, and transfers to a second backbone architecture. Supervised demonstrations alone do not close this gap at any budget we tested; rewarding the intermediate steps is what does.
When More Evidence Hurts: Publication-Bias Drift and Principled Stopping for Biomedical Causal Search
当更多证据带来危害:发表偏倚漂移与生物医学因果搜索的原则性停止
Sun, Fred, Guo, Shangqi
Abstract
Automated biomedical evidence synthesis depends on retrieving published studies, but the biomedical literature is systematically skewed toward positive findings. Deeper retrieval can therefore make a system \emph{more} likely to falsely infer benefit when the true effect is null. We formalise this phenomenon as \emph{evidence drift} and prove that, under a standard publication-bias model, the false-positive probability on null-effect queries follows a strictly increasing large-sample envelope in retrieval depth, approaching one. Empirically, on a held-out test set of 140 Cochrane-derived queries, drift rises monotonically from 7.9\% to 15.7\% as the retrieval budget grows from 3 to 20 steps, and concentrates in the null-effect class. We present DACG-agent, a drift-aware causal-graph agent that incrementally builds a causal knowledge graph from PubMed abstracts and applies a two-layer stopping policy with complementary roles: a KL-divergence monitor that detects posterior convergence (the accuracy layer), and a Bradley--Terry process reward model (PRM) whose online decline detection halts retrieval once evidence quality peaks (the efficiency layer). Against full-budget retrieval, DACG-agent reduces evidence drift from 15.7\% to 6.4\% and improves null-effect accuracy by 21 percentage points (40.0\%$\to$61.4\%) while using 67\% fewer retrieval steps; overall accuracy rises from 61.4\% to 69.3\% (95\% CI 61--77). A simulation confirms the drift result transfers from the analysed vote-counting aggregator to the deployed noisy-OR one.
Chinese Translation
自动化生物医学证据综合依赖于对已发表研究的检索,但生物医学文献系统性地偏向阳性结果。因此,当真实效应为零时,更深度的检索反而可能使系统更容易错误地推断存在获益。我们将这一现象形式化为“证据漂移”(evidence drift),并证明在标准的发表偏倚模型下,零效应查询上的假阳性概率随检索深度遵循一个严格递增的大样本包络,趋于1。实证结果表明,在一个由140个Cochrane来源查询构成的保留测试集上,随着检索预算从3步增加到20步,漂移率从7.9%单调上升至15.7%,且集中于零效应类别。我们提出DACG-agent,一种漂移感知的因果图智能体,它从PubMed摘要中增量构建因果知识图谱,并采用具有互补作用的双层停止策略:一个检测后验收敛的KL散度监测器(准确性层),以及一个Bradley–Terry过程奖励模型(PRM),其在线下降检测在证据质量达到峰值时即停止检索(效率层)。与全额预算检索相比,DACG-agent将证据漂移从15.7%降至6.4%,将零效应准确率提升21个百分点(40.0%→61.4%),同时使用的检索步骤减少67%;总体准确率从61.4%提升至69.3%(95% CI 61–77)。仿真实验证实,漂移结果可从所分析的计票聚合器迁移到实际部署的noisy-OR聚合器。
Tool-calling LLM agents are increasingly deployed in enterprise applications. However, effective evaluation and optimization require high-quality, diverse task datasets that are often difficult to obtain due to privacy and other constraints. Existing synthetic task generation methods often produce generic tasks that ignore an agent's underlying state or database and fail to reflect real-world usage diversity. We propose EdgeGen, a synthetic task generation framework that extracts compliance rules from an agent's specification and uses them to generate database-grounded edge-case tasks designed to violate these rules. When combined with existing synthetic data generation techniques, EdgeGen enables agent improvement through finetuning and harness optimization. The resulting pipeline forms a fully automated closed-loop system that requires no human annotation. Finetuning on data generated by EdgeGen yields a consistent mean progress improvement of 2 percent to 42 percent on tau2bench airline domain, while other baseline methods show degradation for some models. On the other hand, for harness optimization, our method shows a mean progress improvement of 10 percent and 30 percent over the human-curated and base harnesses, respectively, for the Gemma-4-e4b model.
Self-Healing Harness for Runtime Oversight of Agent Self-Modification
用于智能体自我修改运行时监管的自愈框架(Self-Healing Harness)
Tayebati, Sina, Kumar, Divake, Darabi, Nastaran, Krishnan, Ranganath, Trivedi, Amit Ranjan
Abstract
LLM agents can change their own future behavior, raising a basic control question of which self-generated changes should be allowed to persist. We formulate this as admission control for self-modification. The agent may propose changes to its operating instructions, while an external runtime gate controls persistence. We implement this principle as a model-agnostic self-healing harness that runs a Detect, Notice, Heal, Validate loop around an otherwise unmodified agent. The agent authors candidate behavioral rules in an external workspace, where they receive provisional execution authority during evaluation and acquire persistent cross-episode authority only after measured improvement on the triggering failure without regression beyond a fixed margin on protected cases. Replay provides matched evidence when available, forward trials provide a weaker fallback, and a corpus-level guard re-tests the accumulated active rule set. Across 16 matched Baseline and Harness runs spanning AppWorld, Terminal-Bench, and $\tau^2$-Bench, the gate rejected 383 replay-decided proposals. Of these, 211 (55%) improved their triggering failure while degrading a case that previously worked. This shows that locally beneficial self-modifications can introduce collateral regressions often enough to materially affect gate decisions, providing direct empirical motivation for external admission control. Task-completion score is higher under the Harness in all 16 pairs, with two paired bootstrap intervals excluding zero, while repeated-trial reliability is higher in 12 pairs, tied in 4, and lower in none. Because adaptation modifies the policy-inducing context while leaving model weights fixed, admitted changes remain inspectable, reversible, and compatible with closed-weight models.
APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction
APEXA:面向同步辐射数据还原多智能体LLM自动化的执行完整性保障机制
Tripathi, Pawan K., Sharma, Hemant, Chuang, Andrew, Cherukara, Mathew J.
Abstract
Synchrotron data reduction, detector calibration followed by azimuthal integration of terabyte-scale diffraction series, is a multi-step, expert-bound bottleneck that increasingly limits the science rate of user facilities. LLM agents promise to collapse it, but driving a real pipeline with a stochastic model creates a failure mode chat benchmarks cannot see: an agent can report a calibration that was never computed. Correctness here is a property of what executed, not of the transcript. We present APEXA, a deployed multi-agent framework (61 tools over heterogeneous compute, run as a single reasoning loop) automating calibration and integration from natural language at a major light source. We make three contributions. First, execution-integrity enforcement: a deterministic tool-layer guard that refuses to surface any result not backed by an executed tool call, with a parser tolerant of cross-model tool-call format drift: in deployment, a frontier model fabricated a complete calibration-comparison report for commands that never ran, which the guard converts to an explicit non-result; the same code gates an optional motor-control surface at 0/200 adversarial violations against a simulated IOC, versus 15/200 for an equivalent safety prompt. Second, we release APEXA-Bench, an evaluation harness of 58 facility tasks (50 base plus an 8-task cross-detector slice) organized by a four-class physical-consequence taxonomy, the first benchmark axis we know of separating a wasted compute cycle from a damaged instrument; its cross-detector grading against NIST-traceable lattice constants surfaced two latent pipeline bugs. Large-scale agent scoring is left to a full-length study. Third, we validate APEXA on real beamline data: from one natural-language prompt it recovers detector geometry and integrates a full attenuation/exposure sweep. We release the framework, harness and traces.
CREDO: Variance-Guided Rubric Evolution for Replay-Corrected Credit Assignment
CREDO:基于方差引导的评分标准演化与重放校正的信用分配方法
Hu, Xuchun
Abstract
Long-horizon language agents receive sparse terminal feedback, while intermediate rubrics provide structured but potentially misspecified assessments of progress. In resettable training environments, counterfactual continuation rollouts can measure local credit, but exhaustive replay is costly. We propose Credo, a framework that couples evolving semantic rubrics with selective, execution-based credit correction. A frozen judge maps visible transitions to rubric features, and a credit head predicts the change in expected terminal reward associated with the realized transition. Independently sampled two-sided replays correct prediction residuals using their recorded inclusion probabilities. We derive conditional unbiasedness and a variance decomposition that connects two design choices: which rubric features to retain, and where to allocate a fixed expected replay budget. The resulting criterion weights prediction errors by policy-score sensitivity and missing replay coverage; its allocation rule additionally accounts for continuation cost. We also describe a practical mixture with terminal leave-one-out advantages and distinguish its clipped, token-normalized PPO implementation from the ideal policy-gradient estimator. This preliminary report provides the method, proofs, an exact finite-model audit, and a controlled evaluation protocol. It makes no claim of empirical superiority on language-agent benchmarks.
Large language models have achieved remarkable progress on Text-to-SQL through reasoning-enhanced fine-tuning, yet existing approaches predominantly rely on massive instruction corpora under the assumption that scale drives performance. We challenge this paradigm by investigating a fundamental question: what is the minimal data requirement for effective Text-to-SQL instruction tuning? We propose LIMIT(Less Is More for Instruction Tuning in Text-to-SQL), a data-centric framework that demonstrates strong database reasoning can emerge from an extremely compact training set when examples are strategically selected. LIMIT operates through four stages: difficulty-aware filtering that identifies samples within the model's learning frontier, chain-of-thought synthesis with consistency-based selection, multi-dimensional quality scoring via LLM-as-judge, and genetic algorithm optimization that jointly maximizes schema coverage and sample quality. On the BIRD and Spider benchmark, LIMIT selects only 796 and 863 samples while achieving 100% table coverage, enabling Qwen3-8B to reach 69.1% and 88.9% execution accuracy.This result surpasses methods trained on 20 times more data and establishes a new state-of-the-art among open-source approaches. Our findings suggest that careful data curation, rather than scale, is the key to efficient Text-to-SQL learning.
Chinese Translation
大型语言模型通过推理增强微调在Text-to-SQL任务上取得了显著进展,然而现有方法主要依赖大规模指令语料库,并默认规模是驱动性能的关键。我们对这一范式提出质疑,并研究了一个根本性问题:有效的Text-to-SQL指令微调所需的最小数据量是多少?我们提出了LIMIT(Less Is More for Instruction Tuning in Text-to-SQL),这是一个以数据为中心的框架,证明了当样本经过策略性选择时,强大的数据库推理能力可以从极其精简的训练集中涌现。LIMIT包含四个阶段:识别处于模型学习前沿的样本的难度感知过滤、结合一致性选择的思维链合成、基于LLM-as-judge的多维度质量评分,以及联合最大化模式覆盖和样本质量的遗传算法优化。在BIRD和Spider基准上,LIMIT仅选择796和863个样本即实现了100%的表格覆盖率,使Qwen3-8B分别达到69.1%和88.9%的执行准确率。该结果超越了使用20倍数据量训练的方法,并在开源方法中确立了新的最先进水平。我们的研究结果表明,精心的数据筛选而非规模,才是高效Text-to-SQL学习的关键。
SKstars at SHROOM: Visions Agreement-Guided Ensembling of Zero-Shot and LoRA-Adapted Vision--Language Models
SKstars在SHROOM:基于视觉一致性引导的零样本与LoRA适配视觉-语言模型集成方法
Athar, Ali, Ahsan, Imran, Jung, Joon-Yong
Abstract
This paper describes the SKstars submission to SHROOM-Visions 2026, a shared task on fine-grained hallucination detection in large vision-language model outputs. The task requires systems to identify hallucinated character spans, assign hallucination categories, and provide confidence estimates for their predictions. Our approach combines zero-shot predictions from Qwen2.5-VL-72B-Instruct with those of a LoRA-adapted Qwen2.5-VL-7B-Instruct model. The outputs of the two models are integrated through a lightweight ensemble procedure, followed by span refinement and confidence adjustment. We evaluate the main system components on a small internal development subset and report the performance of the submitted system on the official English test set. SKstars achieved a Cor+Lbl score of 0.2902, ranking 15th among 29 teams, and obtained Cor and IoU scores of 0.3642 and 0.3151, respectively, ranking 18th on both metrics. The results show that combining a large zero-shot model with a smaller adapted model provides a practical framework for multilingual and fine-grained hallucination localization, while also highlighting the difficulty of transferring development-set improvements to hidden test data. Code and predictions: https://github.com/aliathar1401/SK-Stars-shroom-visions-2026
Recovering Lost Details: Multi-Scale Frequency Compensation for Long-Term Time Series Forecasting
恢复丢失的细节:面向长期时间序列预测的多尺度频率补偿
Zou, Runmin, Xie, Siyi, Huang, Yaohui, Wang, Yun
Abstract
Long-term time series forecasting has made significant progress by leveraging multi-scale information to capture hierarchical temporal patterns and model long-range dependencies. However, temporal downsampling in existing multi-scale methods inevitably smooths detailed temporal fluctuations, and this information loss is further aggravated by their emphasis on dominant trends across scales, resulting in insufficiently expressive representations. To address this, we propose a Multi-Scale Wavelet Mixing (MWMixer) model, which incorporates a Bidirectional Frequency-Bands Mixing strategy to recover lost temporal details across scales, enabling complementary cross-scale information interactions. Then, a Dynamic Scale-Adaptive Fusion module learns time-varying weights for each scale to fuse multi-scale forecasts into the final prediction, enhancing the flexibility of multi-scale aggregation. In addition, a cross-scale consistency loss aligns each coarse-scale prediction with the interval-averaged fine-scale outputs, while a multi-scale supervision loss enforces prediction accuracy at each scale, promoting consistent learning across scales. Extensive experiments on seven real-world datasets demonstrate that MWMixer achieves competitive performance in long-term forecasting.
Reinforcement learning (RL) improves reasoning in vision-language models (VLMs) but can induce chain-of-thought (CoT) obfuscation: an operational, non-intentional outcome where task reward or accuracy rises while traces become less grounded and monitorable. Prior work largely documents this decay behaviorally, leaving its representation-level correlates and actionable controls unclear. We find that template- and ground-associated activations become less separable during RL; matched interventions support the contribution of selected features to monitorability degradation. Guided by this evidence, we propose Targeted Anti-obfuscation with Mechanistic Enforcement (TAME), which uses Sparse Autoencoders (SAEs) to combine behavioral feedback with targeted suppression of template-associated activations during RL. Its asymmetric constraint penalizes template activations only above their pre-RL baseline, anchoring the localized features while behavioral feedback promotes grounded refinements. Across VIRL-39k, SPA-VL, and two model families, TAME improves CoT monitorability by up to 30.9 and 16.7 percentage points over Group Relative Policy Optimization (GRPO), respectively. Blinded human evaluation finds higher human monitorability on both datasets, and two held-out monitor families reproduce the monitorability gains. Task accuracy changes are small and mixed, and general-capability benchmarks show task-specific trade-offs. These results provide a path from behavioral monitoring to representation-level oversight for more auditable RL-trained multimodal systems.
Chinese Translation
强化学习(RL)能够提升视觉语言模型(VLM)的推理能力,但也可能引发思维链(CoT)混淆:这是一种操作性、非意图性的现象,即任务奖励或准确率上升,而思维链轨迹却变得更难以溯源和监控。已有研究大多仅从行为层面记录这种退化现象,未能阐明其表征层面的关联机制以及可操作的干预手段。我们发现,在强化学习过程中,与模板相关的激活和与真实依据相关的激活变得越来越难以区分;匹配干预实验验证了特定特征对监控性退化的贡献。基于这一证据,我们提出了带有机制性约束的定向反混淆方法(Targeted Anti-obfuscation with Mechanistic Enforcement, TAME),该方法利用稀疏自编码器(SAE)将行为反馈与强化学习过程中对模板相关激活的定向抑制相结合。其非对称约束仅在模板激活超过强化学习前的基线水平时施加惩罚,从而锚定局部化特征,同时行为反馈促进基于真实依据的改进。在VIRL-39k和SPA-VL数据集以及两个模型家族上,TAME分别将CoT监控性相比组相对策略优化(GRPO)提升了最高30.9和16.7个百分点。盲评人类评估表明,在两个数据集上人类可监控性均有所提高,且两个保留的监控器家族均复现了监控性提升。任务准确率变化较小且不一致,通用能力基准测试显示出任务特定的权衡。这些结果为从行为监控走向表征级监督提供了一条路径,有助于构建更可审计的强化学习训练的多模态系统。
Unsupervised Brain Anomaly Detection as a Bayesian Inverse Problem with Diffusion Prior
基于扩散先验的贝叶斯逆问题无监督脑部异常检测
Roy, Hugues, Dorent, Reuben, Burgos, Ninon
Abstract
Unsupervised anomaly detection (UAD) aims to localize abnormal regions in medical scans without pixel-level annotations. A typical strategy seeks to reconstruct a pseudo-healthy image that preserves subject-specific anatomy. Recently, diffusion models have been proposed to perform UAD. However, these methods rely on heuristic noise schedules or synthetic corruptions to balance subject-specificity and anomaly removal. In this work, we propose an alternative formulation of UAD as a Bayesian inverse problem under a diffusion prior. First, we introduce a latent spatial anomaly mask that models pixel-wise consistency between a test image and its latent corresponding pseudo-healthy image. Then, we propose an approximation of the unknown generation process that links healthy anatomy, anomalies, and the observed image, enabling a well-defined likelihood within the Bayesian framework. Building on recent advances in diffusion-based inverse problem methods, we jointly infer the pseudo-healthy image and the anomaly mask via annealed posterior sampling. We evaluate our approach on FDG PET (ADNI) and FLAIR MRI (BraTS 2021), demonstrating improved anomaly localization performance compared to other diffusion-based approaches and validating the contribution of our introduced model. Our code is available at https://github.com/HuguesRoy/UAD_DAPS.
How Many Pixels Is a Digit Worth? Place-Aware Coordinate Entropy for GUI Agent Confidence Estimation
一个数字值多少像素?用于GUI智能体置信度估计的位置感知坐标熵
Li, Yunxiang, Wu, Xixin, Meng, Helen
Abstract
GUI agents predict click coordinates as digit-token sequences, but standard text-LLM confidence estimation methods rank correct clicks from wrong ones only weakly. GUI-specific alternatives use K samples or new supervision, but still leave room for improvement. We trace part of this to place-value asymmetry: bounding-box correctness often makes higher-place digits more important than lower-place digits, so uniform aggregation weakens the signal that determines correctness. The fix is to weight each digit's Shannon entropy by its place value. We call this Place-Aware Coordinate Entropy (PACE). Across fixed-scale agents on ScreenSpot-Pro and ScreenSpot-v2, PACE wins both AUROC and selective accuracy on all primary comparisons in a single forward pass, matching or outperforming K-sample baselines at a fraction of the cost. PACE provides a per-click confidence estimate that turns coordinate-token internals into a practical confidence signal for GUI agent deployment.
When and How Should an Agent Clarify? CIGAsk: Teaching LLMs to Clarify via Counterfactual Information Gain
智能体应在何时以及如何进行澄清?CIGAsk:通过反事实信息增益教会大语言模型进行澄清
Li, Yunxiang, Wu, Xixin, Meng, Helen
Abstract
Instruction-tuned LLMs faced with underspecified queries often commit to a single interpretation rather than ask for clarification, producing confidently wrong answers. In our experiments, prompting alone is insufficient: models either ask for clarification on every query or ask vague questions that fail to recover the missing information. Addressing this failure requires learning two coupled skills: when to ask rather than answer and how to ask a question that recovers the disambiguating information. Existing recipes either address only one of these skills or require a separately trained critic. We propose CIGAsk, an RL recipe that teaches both skills through two complementary reward signals within a multi-turn GRPO loop. Counterfactual Information Gain (CIG) compares the gold-answer log-likelihood under a frozen reference model with and without the user response, providing per-turn credit that guides how to ask. The Asymmetric Ambiguity Bonus assigns a signed reward at the terminal token based on the gold ambiguity label, guiding when to ask. Across three clarification benchmarks spanning table, passage, and open-domain QA, CIGAsk-7B outperforms the strongest external baseline despite using a smaller backbone. It also transfers across datasets without per-dataset tuning while preserving single-turn QA performance on out-of-distribution benchmarks.
Chinese Translation
面对欠明确的查询,经指令微调的大语言模型(LLM)往往倾向于采用单一解释而非请求澄清,从而生成自信但错误的答案。在我们的实验中,仅靠提示词并不足够:模型要么对每个查询都请求澄清,要么提出无法恢复缺失信息的模糊问题。解决这一失败需要学习两项相互耦合的技能:何时应该提问而非直接回答,以及如何提出能够恢复消歧信息的问题。现有方法要么只针对其中一项技能,要么需要单独训练的评判模型。我们提出 CIGAsk,这是一种强化学习方案,通过在多轮 GRPO 循环中引入两个互补的奖励信号来同时教授这两项技能。反事实信息增益(Counterfactual Information Gain,CIG)比较冻结参考模型在有和没有用户回复情况下对标准答案的对数似然,提供逐轮的信用分配,指导模型如何提问。非对称模糊奖励(Asymmetric Ambiguity Bonus)根据标准模糊标签在终止标记处赋予带符号的奖励,指导模型何时提问。在涵盖表格、段落和开放域问答的三个澄清基准上,CIGAsk-7B 尽管使用了更小的骨干模型,仍优于最强的外部基线。此外,它无需针对每个数据集进行调优即可跨数据集迁移,同时在分布外基准上保持了单轮问答性能。
Brain-Token Learning: Microstate-Based Tokenization and Multi-Scale Interaction for Long-Horizon EEG Sequence Modeling
脑token学习:基于微状态的分词与多尺度交互用于长时程脑电序列建模
Ye, Weishan, Pan, Yue, Zhang, Li, Huang, Gan, Liang, Zhen
Abstract
Electroencephalography (EEG) provides a non-invasive window into dynamic brain activity, yet modeling long-horizon EEG sequences remains challenging due to their high temporal complexity, substantial variability across subjects, and the lack of biologically meaningful sequence representations. Existing tokenization strategies, such as fixed-window and patch-based representations, discretize EEG signals according to artificial temporal boundaries, which may disrupt intrinsic brain-state dynamics. In this work, we propose Brain-Token Learning, a neuroscience-inspired framework that introduces Brain Tokenization for long-horizon EEG sequence modeling. Instead of partitioning EEG signals into predefined temporal segments, Brain Tokenization represents EEG as sequences of recurrent microstate-derived brain tokens, where each token corresponds to a quasi-stable large-scale brain state with variable temporal duration. Based on these biologically grounded tokens, we further develop a multi-scale token interaction module consisting of Latent State Aggregation and State Transition Modeling to jointly capture global brain-state context and local microstate transitions. We evaluate Brain-Token on five heterogeneous EEG datasets, including the newly collected long-horizon NeuroLong dataset and four affective or clinical EEG datasets (SEED, DEAP, MDD, and NSSI). Extensive experiments demonstrate that Brain-Token consistently outperforms conventional CNN/LSTM architectures, Transformer-based models, and domain adaptation methods across diverse EEG scenarios. Further analysis verifies the effectiveness of microstate-based tokenization and multi-scale interaction for learning robust and interpretable EEG representations. These results establish Brain-Token as a biologically grounded tokenization paradigm for long-horizon EEG sequence modeling.
Chinese Translation
脑电图(EEG)为观测动态大脑活动提供了一种无创窗口,然而由于其高度的时间复杂性、显著的被试间差异以及缺乏具有生物学意义的序列表示,长时程脑电序列建模仍然具有挑战性。现有的分词策略(如固定窗口和基于patch的表示)依据人为设定的时间边界对脑电信号进行离散化,这可能会破坏大脑状态的内在动态特性。在本工作中,我们提出Brain-Token Learning(脑token学习),一种受神经科学启发的框架,引入脑分词(Brain Tokenization)用于长时程脑电序列建模。与将脑电信号划分为预定义时间段不同,脑分词将脑电表示为由复现微状态(microstate)导出的脑token序列,其中每个token对应一个具有可变时间持续时间的准稳定大规模脑状态。基于这些具有生物学依据的token,我们进一步开发了一个多尺度token交互模块,由潜在状态聚合(Latent State Aggregation)和状态转移建模(State Transition Modeling)组成,以联合捕捉全局脑状态上下文和局部微状态转移。我们在五个异构脑电数据集上评估了Brain-Token,包括新收集的长时程NeuroLong数据集以及四个情感或临床脑电数据集(SEED、DEAP、MDD和NSSI)。大量实验表明,Brain-Token在多种脑电场景中始终优于传统的CNN/LSTM架构、基于Transformer的模型以及域适应方法。进一步分析验证了基于微状态的分词和多尺度交互在学习鲁棒且可解释的脑电表示方面的有效性。这些结果确立了Brain-Token作为一种具有生物学依据的长时程脑电序列建模分词范式。
Graph Retrieval-Augmented Generation (GraphRAG) has remarkably enhanced large language models on complex reasoning by leveraging structured entity topologies. However, existing frameworks heavily rely on standard autoregressive language models where the nature of inherent sequential generation severely hinders overall inference efficiency. Inspired by Diffusion Language Models (DLMs) that offer massive parallelism via continuous refine-in-parallel decoding, we aim to accelerate GraphRAG in the discrete space. However, it remains non-trivial for two challenges. First, partially denoised drafts are highly dynamic and uncertain, making dynamic graph grounding non-trivial. Second, raw denoising states are inherently noisy and unstable, making synchronous graph retrieval and multi-hop aggregation computationally prohibitive. To this end, we present LADDER, a novel framework that bridges diffusion language modeling with GraphRAG through graph-guided parallel decoding. Specifically, (i) we propose an event-driven self-clocking retrieval, inspired by our key insight that 88% of target entities emerge early in the partially denoised state, leading final commitment by an average of 5.7-9.6 steps. This mechanism dynamically triggers graph retrieval only when the set of graph-linkable entities expands, yielding an asynchronous self-clocking policy that bypasses learned gates or heuristic thresholds. (ii) An incomplete-query graph propagation module is designed to process the newly emerging entity queries using a specialized graph foundation model, continuously aggregating multi-hop evidence to sharpen parallel predictions and accelerate overall decoding convergence. Extensive experiments on three challenging multi-hop QA benchmarks show that LADDER raises average exact match from 39.6% to 45.2% while achieving a 4.1x latency reduction.
Chinese Translation
图检索增强生成(GraphRAG)通过利用结构化实体拓扑,显著提升了大语言模型在复杂推理任务上的表现。然而,现有框架高度依赖标准的自回归语言模型,其固有的顺序生成特性严重阻碍了整体推理效率。受扩散语言模型(Diffusion Language Models, DLMs)通过持续并行精炼解码实现大规模并行性的启发,我们致力于在离散空间中加速GraphRAG。然而,这面临两大非平凡挑战:其一,部分去噪的草稿具有高度动态性和不确定性,使得动态图对齐(graph grounding)难以实现;其二,原始去噪状态本身噪声大且不稳定,使得同步图检索与多跳聚合在计算上代价高昂。为此,我们提出了LADDER,一个通过图引导并行解码将扩散语言建模与GraphRAG相连接的新型框架。具体而言,(i)我们提出了事件驱动的自时钟检索机制,该机制源于我们的一个关键发现:88%的目标实体在部分去噪状态中即已早期涌现,并平均领先最终确定5.7至9.6步。该机制仅在可链接到图的实体集合扩展时才动态触发图检索,形成了一种异步自时钟策略,从而无需依赖学习得到的门控或启发式阈值。(ii)我们设计了不完整查询图传播模块,利用专门的图基础模型处理新涌现的实体查询,持续聚合多跳证据,以锐化并行预测并加速整体解码收敛。在三个具有挑战性的多跳问答基准上的大量实验表明,LADDER将平均精确匹配率从39.6%提升至45.2%,同时实现了4.1倍的延迟降低。
Large language models (LLMs), when acting as agents, are expected to take observed data in context, infer the latent state space underlying the world, and leverage it for downstream prediction. However, prior work demonstrated that LLMs struggle to use representations learned in context on a graph tracking task, where the model needs to construct a representation of the graph governing data generation process and use it for subsequent predictions. In this paper, we show that extending this to few-shot settings, where each demonstration is generated from a different world with either the same or different graph topologies, enhances its prediction on 6 models from 4 model families. To understand this improvement, we linearly probe a low-dimensional world representation that encodes graph information in the hidden states. Notably, we find that few-shot demonstrations relocate the world representation and increase its predictive use. Specifically, for each model, these world representations shift in directions nearly orthogonal to their original subspace, and interventions on these representations selectively impair performance more than interventions on other subspaces. Consistent with this insight, we show that few-shot demonstrations with observations from different worlds improve performance on ARC-AGI-1&2, web agent tasks, and Othello. Our findings elucidate the role and internal mechanisms of few-shot demonstrations in in-context world modeling. More broadly, our work advances our understanding of how LLM agents learn from in-context observations and provides implications for their further improvement.
VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning
VLM-in-Sandbox:面向智能体视觉推理的视觉工作区
Yang, Hexiong, Chen, Mingrui, Cao, Jie, He, Ran
Abstract
Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language models to vision-language models (VLMs) introduces a distinct state-management problem. Visual reasoning produces intermediate image-valued evidence---crops, masks, overlays, zoomed regions, and analytic renderings---that must remain addressable without accumulating unboundedly in multimodal context. We introduce VLM-in-Sandbox, a training-free framework for agentic multimodal reasoning in controlled computer environments. Its Visual Workspace registers generated artifacts in an image ledger, maintains a bounded active visual context, and lets the model explicitly promote selected evidence for subsequent inspection. This separates visual evidence generation, performed by sandbox tools, from visual evidence management. Across seven benchmarks and four base VLMs, VLM-in-Sandbox achieves the highest sample-weighted average accuracy among Vanilla VLM, Append-only Sandbox, and the proposed method. A compiler-matched $2\times2$ study on 1,260 examples further separates model-directed visibility from bounded retention: VLM-in-Sandbox reaches 66.27% accuracy with 18.6% fewer total tokens than the automatic, retain-all control. Over all 6,350 submitted GPT-4.1-mini examples, it produces 302 rescues and 142 regressions relative to Original Append-only. A local vLLM study with prefix caching confirms that the smaller request workload also reduces uncached tokens, time to first token, and end-to-end latency. These results identify explicit visual evidence state as a central abstraction for sandboxed VLM agents.
Predicting postprandial glycemic response (PPGR) is fundamental to personalized nutrition and type 2 diabetes management, yet existing approaches typically rely on manually reported dietary intake, limiting their scalability in free-living settings. We propose a multimodal framework that replaces manual dietary logging with image-derived macronutrient estimates and integrates them with clinical variables and gut microbiome information for personalized PPGR prediction. The framework jointly performs image-based macronutrient estimation and glucose prediction, while an attention-based prediction module models interactions between dietary and host-specific information. We evaluate the proposed approach on a real-world dataset comprising meal images, continuous glucose monitoring, clinical variables, and gut microbiome profiles. The proposed model outperforms existing PPGR baselines using image-derived nutritional inputs and approaches the performance of methods that rely on manually reported macronutrients despite using automatically estimated nutritional information. These results demonstrate that combining image-derived nutrition with complementary clinical and gut microbiome information provides a practical foundation for scalable personalized PPGR prediction.
Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary dimensions - diagnostic reasoning: the convergent, evidence-driven task of inferring a patient's condition from clinical data to produce a diagnosis, and clinical healthcare reasoning: the broader, navigational judgment required to communicate, plan, and adapt across multi-turn clinical interactions where a single correct answer may not exist. Recent benchmarks such as HealthBench and MedXpertQA reveal persistent weaknesses in both areas, exposing failures in complex diagnostic scenarios and limitations in contextual, patient-centered dialogue. We introduce a sequential training framework that targets these facets using synthetic data and rubric-based reinforcement learning. First, we improve diagnostic reasoning using MedBullets-derived questions with rule- and rubric-guided Reinforcement Learning (RL). We then shift to clinical reasoning by generating 5.3k synthetic multi-turn scenarios, each paired with multi-dimensional rubrics to comprehensively assess the response. This approach yields over 10% improvement on MedXpertQA, and our 30B model achieves 50.1% accuracy on HealthBench-Hard, surpassing proprietary baselines including GPT-5 (thinking). Our results show that targeted synthetic datasets and rubric-based training can systematically improve both diagnostic and interactive clinical reasoning in medical LLMs.
Not All Task Vectors Need Equal Rank: Energy-Proportional Allocation for Model Merging
并非所有任务向量都需要相同秩:面向模型合并的能量比例分配方法
Cho, Hyunjoong, Jang, Jinhyeok
Abstract
Model merging aims to combine multiple fine-tuned models derived from a common pretrained model into a single multi-task model without additional joint training. Recent spectral merging methods improve over simple weight averaging by exploiting low-rank structures of task-specific updates, but they commonly assign the same rank capacity to every task. This uniform allocation ignores that task vectors can have heterogeneous spectral complexity, causing the shared merging space to be used suboptimally. In this paper, we propose Spectral Energy-proportional Rank Allocation (SERA), a simple task-adaptive strategy that allocates ranks according to the singular-value energy structure of each task vector. By assigning richer spectral capacity to complex or isolated tasks and fewer directions to compact tasks, SERA extends SVD-based model merging from uniform-capacity merging to task-dependent capacity allocation. Experiments under standard vision model merging protocols show that SERA improves multi-task merging performance while preserving the same total rank budget as existing spectral merging methods. Further analysis demonstrates that task-level spectral concentration is closely related to the per-task effect of adaptive rank allocation, providing insight into when and why SERA is effective.
The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
无尽的考试:从当今模型通往超级智能的数学构造
Zhang, Muhan
Abstract
We introduce the Endless Exam, a benchmark for measuring mathematical progress from today's models toward artificial superintelligence through fourteen parameterised construction families. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at $1$. The families draw on open mathematical problems for long-term targets and generate new instances at larger parameters, where compact certificates keep large constructions verifiable. Across eight models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though no evaluated system surpasses a published frontier. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.
Ziletti, Angelo, D'Ambrosi, Leonardo, Tuchardt, Melanie, Kondziella, Tim
Abstract
Answering epidemiological questions from real-world clinical data requires medical coding, schema-aware SQL, and validation of implicit choices about populations, denominators, and time. We present Ascent, an agentic system that exposes medical coding, question answering, and cohort analysis through a shared Model Context Protocol tool surface for standardized and native schemas. We introduce EpiTrap, a dataset testing whether systems avoid recognized pharmacoepidemiological errors, and compare a fixed pipeline with agents across models and orchestrators. With capable models, agents improve accuracy over the fixed pipeline by an average of 27 and 20 percentage points on native and standardized schemas, respectively. These gains require more tool calls and longer runtimes. Experience from real projects highlights the system's value for feasibility assessment, diagnostic iteration, and expert-guided analysis.
How should natural language processing models be selected and adapted for global health literature in environments where annotated data and computational resources are limited? This thesis investigates these challenges through experiments on semantic tag discovery, named entity recognition (NER), and multi-label topic classification. First, skip-gram word2vec models trained on progressively larger specialized corpora are compared with BioWordVec to assess how corpus size and domain context influence tag discovery. Vocabulary coverage and qualitative evaluation indicate that broader coverage does not necessarily yield more useful domain-specific associations. The analysis then turns to entity extraction, comparing convolutional spaCy models with a RoBERTa-based transformer on 1,000 annotated sentences. Under a lenient scoring protocol, the transformer achieves 0.80 micro-F1 versus 0.65-0.69 for convolutional models, but takes 82 seconds rather than 5-6 seconds. This trade-off motivates fine-tuning convolutional models and integrating a disease recognizer that achieves 81.33% test F1 on the NCBI Disease Corpus. Combined with PDF preprocessing, entity filtering, and MeSH enrichment, the resulting pipeline supports document-level indexing. To complement entity extraction with thematic annotation, MiniLM-based few-shot classification is compared with BART-MNLI zero-shot inference across 50 topics and 1,000 handcrafted test sentences. BART-MNLI achieves 95.2% single-label accuracy versus 59%; reported multi-label accuracies are 88% and 32% under partly manual assessment. However, its higher inference cost limits practical integration. The results show where domain specialization and lightweight adaptation offer practical value, and where transformer accuracy justifies higher inference costs, providing an empirical basis for building knowledge systems under resource constraints.
LLM-based agents increasingly operate in environments where they interact with users, tools, and external systems. Yet most security evaluations assume passive users and static control, ignoring the interactive dynamics that shape real agent behavior. We introduce \textbf{DUMA-Bench}, a benchmark and evaluation protocol for measuring agent security under \emph{dual-control} interaction, where both the agent and the user can influence the shared environment state. DUMA-Bench extends $\tau^2$-bench ~\cite{barres2025tau} with adversarial environments covering eight vulnerability classes, including RAG poisoning, cross-agent manipulation, and unsafe output handling. We evaluate \textbf{14 models from five model families} (OpenAI, Anthropic, DeepSeek, Qwen, and Z.ai) across eight domains and multiple user-behavior regimes. Across our experiments, introducing dual-control interaction increases the attack success rate from \textbf{26.9\%} to \textbf{41.1\%}. These results show that agent security is not solely a property of the model but emerges from the interaction between the model, the user, and the environment. DUMA-Bench provides a missing evaluation layer for studying security in realistic agent deployments.
Self-evolving agents convert interaction feedback into persistent artifacts, such as memories or skills, which in turn guide subsequent decisions. As these artifacts are iteratively updated throughout an experience stream, the capabilities they support may evolve. Consequently, endpoint performance alone offers an incomplete view of self-evolution. Process-level evaluation is therefore essential to identify when a target capability emerges and whether later updates strengthen, preserve, or weaken it. Motivated by this, we propose \textsc{EvoPathBench}, a benchmark that tracks individual capabilities during artifact-level self-evolution. EvoPathBench fixes the base model, tools, freezes evolving artifacts at successive checkpoints, and evaluates the target capability on held-out episodes. This benchmark evaluates agent self-evolution using public trading data and calibrated trajectories. It tests three capabilities: generalization to unseen tasks, retention after unrelated learning, and rule adaptation to new evidence. Experimental results show that gains on similar unseen tasks often weaken under distribution shift, retention losses are concentrated in a minority of evolution paths, and no method achieves reliable rule adaptation. Moreover, while self-evolution enables agents to generate candidate artifacts with substantial held-out gains, the selected updates consistently fall short of realizing this potential. Together, these findings establish capability-level process evaluation as a foundation for analyzing self-evolution, identifying candidate evaluation and selection as key targets for improvement.
Large language models (LLMs) are increasingly used to make predictions from numerical time-series histories and textual events. Yet accuracy alone cannot reveal whether correct answers reflect effective integration of the two inputs or instead arise from event polarity, unimodal priors, or superficial cues. Likewise, plausible explanations may rationalize predictions without faithfully reflecting the evidence that drives model behavior. We introduce TimeLitmus, a diagnostic benchmark for cross-modal understanding and explanation faithfulness in event-conditioned time-series prediction. TimeLitmus contains 4,856 evaluation records across Finance and Traffic, combining natural prediction with controlled counterfactual and contrastive interventions, explanation-targeted faithfulness tests, and systematic shortcut controls. Across ten representative LLMs, standard prediction accuracy substantially overstates reliable cross-modal understanding: Hard Paired Contrast (HPC) pair correctness peaks at only 19.2% in Finance and 11.7% in Traffic, and all ten models show lower-than-expected consistency on Finance series-side controls. Models often recognize scenario relations explicitly yet fail to apply them during independent prediction. Explanation faithfulness shows a similar gap: in Traffic, most models cite the manipulated temporal factor in over 90% of cases, while behavioral support remains below 22%. Human annotators outperform LLMs on matched controlled and hard-pair diagnostics, confirming that these distinctions are recoverable from the inputs. Natural-only adaptation yields selective gains in evidence selection and input sensitivity, but not consistent gains in controlled or hard-pair behavior. The benchmark, evaluation suite, and supervised adaptation data will be released publicly.
Language agents solve complex tasks through plans and actions. A single step the world refuses puts the goal out of reach, and what the agent does next decides the task. Prompted planners fail at exactly this point, rewriting the refused step in new words, meeting the same refusal, and burning the attempt budget without moving. They fail because the plan was never tied to the world, so a refusal has nothing in the plan to attach to. A world is where a task runs, and it has its own rules, its own admissible actions, and its own constraints. We build synthetic worlds across 7 domains and extract training data from them. A program enforces each world's rules and grades its goal, and every world is admitted only if its goal is reachable from its initial state. Agents run inside and leave verified failures paired with repairs that carried the run to a state the world certified, a record of about 226K trajectories. On this record we train the World State Generator, a model that writes a plan as checkable states of the world and keeps that plan aligned with the world it runs in. That alignment is what a plan written in language lacks, since the world it runs in has physical limits, logical dependencies, and required orders the language never states, and the plan encounters these rules only when a state fails. WSG takes that failure as the rule the world has stated and rewrites the remaining states to obey it, so the plan bends to the world as the run goes on. Across 7 public benchmarks, WSG raises end-to-end success for two open models near 30B parameters over prompting and brings to the level of proprietary model.
Epi-Logic: A Conceptual Framework for Epistemic Runtime Control, Schema Validity Checking, and Controlled Accommodation in Autonomous AI Agents
Epi-Logic:自主AI智能体中认知运行时控制、模式有效性检验与受控顺应的概念框架
Wetzk, Boris
Abstract
Autonomous AI agents are increasingly deployed in areas where wrong decisions are hard to reverse. This paper examines schema mismatch: the condition in which an agent operates within an interpretive frame that no longer applies to the current context. Outputs produced under such a mismatch can appear internally consistent, linguistically plausible, and largely factually correct; output-quality metrics alone therefore capture the underlying loss of validity only partially. The paper introduces Epi-Logic, a conceptual framework for epistemic runtime control. It couples the detection of schema dissonance, a graduated reduction of autonomy, and the auditable switch to a validated schema. A schema is formalised as a tuple of variable space, expectation model, validity conditions, axioms, and metadata. The Epi-Score aggregates seven graded dimensions of epistemic dissonance; the temporal validity dimension D8, violations of the validity conditions G, and axiom violations are carried as separate categorical paths that are not offset against the aggregate. The architecture rests on a checking asymmetry: formalised validity conditions can be checked at runtime, whereas the correctness of many actions is established only ex post. The paper separates two architectural properties, a conditional result from sequential changepoint detection, and an empirical remainder. Eight falsifiable propositions with named baselines describe the transition to empirical validation. All propositions are empirically testable hypotheses, not established results.
When facing complex problems, humans tend to try various ideas for different issues. Human thinking patterns exhibit remarkable flexibility in adapting to diverse scenarios. GPT-o1, GPT-o3, and DeepSeek-R1 adopt long chain-of-thought models to address complex problems by increasing reasoning depth, which default to a forward reasoning mode. We conducted statistical analysis on the accuracy of different mathematical problem datasets on models of different scales, and found five reasons for errors: Insufficient solution-space coverage, Computational mistakes, Unverified assumptions, Ignoring constraint conditions, Maximum response length limitation. To address the above issues, we proposed a backward reasoning pattern construction method aimed at enhancing the model's reverse thinking ability and dynamic adaptability. First, we constructed an easy-hard two-stage Math dataset for training large models and gradually improving their inference ability at different difficulty levels. The dataset contains forward reasoning paths as well as backward reasoning paths. And a two-stage supervised fine-tuning process is applied to progressively train the model's backward reasoning capability. Furthermore, a fine-grained reward mechanism is developed, employing smoothed reward signals to strengthen the model's ability to autonomously select thinking modes during the reasoning process, thereby avoiding reward hacking. A linear-decay balanced sampling strategy is designed to maintain a balance between forward and backward reasoning path samples during training, enabling the model to converge quickly and stably. Experimental results show that our method significantly improves reasoning efficiency and accuracy in tasks such as mathematical proofs, offering a flexible and efficient reasoning paradigm for solving complex problems.
Convex AI Compositionality and the Governance of AI System Populations
凸性AI组合性与AI系统群体的治理
Ferrario, Andrea
Abstract
AI governance increasingly requires providers and public authorities to reason about multiple AI instantiations, alternative versions, and deployment configurations of multiple AI systems. Yet current regulation remains predominantly single-system-centric, acknowledging such multiplicity only sparsely without treating collections of related AI systems as governance objects. This creates an AI population governance problem: determining which instantiations can be meaningfully considered together and how their changing configurations can be represented and monitored. The first requirement has recently been addressed through trustworthiness-based accounts of AI identity. We address the second by introducing convex AI compositionality: a formal representation of the configurations generated by finite AI system populations that uses convex spaces. The core idea is that convex compositions of the operational states that a population of AI system instantiations may occupy over time are compatible with lifecycle reachability across the population and can preserve the formal identity relations between these systems. Well-known statistical and geometric constructions, such as weighted state distributions and convex hulls, become AI governance tools for distinguishing operational states, population weights, heterogeneity, and AI configuration change across different governance modes while remaining compatible, under stated conditions, with lifecycle reachability and AI identity. We illustrate our AI population governance framework through distributed healthcare deployments and controlled deployment of recruitment AI variants.
GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes
GRUET:量化智能体推理-行动过程的不确定性
Liang, Shuang, Hu, Xin-Yu, Zhang, Shao-Qun
Abstract
Agents have attracted considerably increasing attention due to the power of executing both Reasoning and Acting (ReAct) in open and dynamic environments. The ReAct process typically exhibits a multi-turn trajectory in which one drives Large Language Models (LLMs) to generate both reasoning chains and task-specific actions in an interleaved manner. However, agents often suffer from significant uncertainty, where identical tasks yield divergent trajectories; trajectories with higher uncertainty often produce incomprehensible behaviors, severely undermining agent credibility. This work conjectures that such trajectory-level uncertainty frequently stems from cumulative turn-level reasoning uncertainty induced by LLMs; the latter often exhibits a collection of branches of divergent reasoning chains and their resulting actions. Built upon this, we present the Graph-based Reasoning UncErtainty in Trajectories (GRUET) method for the uncertainty quantification of ReAct, comprising turn-level reasoning uncertainty quantification and trajectory-level uncertainty aggregation; the former precisely quantifies reasoning uncertainty via modeling the reasoning space spanned by potential reasoning branches as a graph and then approximating the reasoning space complexity with graph complexity, while the latter employs simple aggregation strategies for quantifying the overall trajectory credibility. Empirical evaluations across nine LLMs and five benchmarks validate the effectiveness of our proposed GRUET in terms of selective generation performance, measured by AUROC, AUPRC, and AUARC.
Medical agents increasingly combine general reasoning models with specialized clinical tools, yet their capabilities remain largely fixed by what clinicians and engineers design before deployment. Recursive self-improvement (RSI) offers a different paradigm in which agents learn from their own failures and autonomously expand their capabilities, but directly applying RSI to medicine introduces fundamental safety challenges. We introduce MedRSI, the first recursive self-improvement framework for medicine, which continuously transforms diagnostic failures into new clinical capabilities through tool composition and task-specific model training. Inspired by clinical practice, MedRSI introduces two mechanisms for clinically aligned self-evolution. Clinical-cost-aware failure prioritization directs improvement toward errors according to their potential clinical consequences rather than frequency alone. Fast discovery with slow registration separates rapid capability invention from conservative adoption, allowing new tools to enter the persistent agent only after demonstrating sustained benefit across subsequent patient cohorts. Across public glaucoma and heart disease benchmarks and two private clinical tasks, MedRSI progressively develops segmentation, measurement, prediction, multimodal reasoning, and generative capabilities, surpasses manually engineered medical agents, and autonomously discovers solutions to clinical problems not anticipated by its original designers. Our results show that medical agents need not remain constrained by capabilities specified before deployment: with clinically grounded mechanisms governing what to improve and what to retain, they can continuously construct, validate, and accumulate new capabilities from diagnostic experience. Code is available at https://github.com/ImprintLab/MedRSI.
Argumentative component detection (ACD) is a core subtask of Argument(ation) Mining (AM) and one of its most challenging aspects, as it requires jointly delimiting argumentative spans and classifying them into components such as claims and premises. While research on this subtask remains relatively limited compared to other AM tasks, most existing approaches formulate it as a simplified sequence labeling problem, component classification, or a pipeline of component segmentation followed by classification. In this paper, we propose ITFACD, a novel approach based on instruction-tuned Large Language Models (LLMs) using compact instruction-based prompts, and reframe ACD as a language generation task, enabling arguments to be identified directly from plain text without relying on pre-segmented components. Experiments on standard benchmarks show that our approach achieves higher performance compared to state-of-the-art systems. To the best of our knowledge, this is one of the first attempts to fully model ACD as a generative task, highlighting the potential of instruction tuning for complex AM problems. Our code and the datasets used are openly available in the following GitHub repository.
Partner-Specific Affective Precision in Social Active Inference
社会主动推断中的伙伴特异性情感精度
Shah, Harshil, Pashea, Andrew
Abstract
In multi-agent social settings, model reliability varies across relationships. Beyond inferring what others will do, an agent must calibrate how confidently those inferences should guide policy selection for each relationship. An agent may maintain a well-validated model of one partner, a fragile model of another, and a model under revision for a third; collapsing these into a single confidence estimate loses information relevant to policy selection. We therefore formalize affective precision as a relationship-specific metacognitive estimate of confidence in the current partner model. Each partner's behavioral evidence updates a local confidence estimate that modulates policy precision during selection, regulating how strongly current beliefs are expressed in policy rather than changing the content of those beliefs. Simulations in a multi-partner graded trust game show that partner-local affective precision influences behavior primarily through policy commitment rather than direct improvement of partner-state inference. Because the mechanism tracks partner-response predictability rather than realized payoff, greater confidence produces sharper policy commitment without necessarily producing higher rewards. Under abrupt shifts in social behavior, confidence accumulated from previously reliable predictions can remain behaviorally active after the relationship changes, showing that confidence revision can lag behind social change. Finally, varying precision gain and priors produce distinct trust-calibration dynamics, showing how confidence accumulation and revision depend on model parameters. Together, these results show how relationship-specific affective precision can distinguish social prediction from social policy commitment.
Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models
Pinocchio:面向黑盒语言模型的快速不确定性估计方法
Hayes, Kevin David, Pal, Arka, Zhang, Haosong, Goldstein, Tom, Goldblum, Micah
Abstract
In high-stakes decision-making applications of large language models (LLMs), practitioners require not only accurate LLMs but also uncertainty estimates for their predictions. Existing approaches to uncertainty estimation for LLMs require access to log-probabilities output by the model or require fine-tuning access. However, many industrial LLM products use closed-source API models, and many such API models like GPT do not return log-probabilities and may not allow fine-tuning. We introduce Pinocchio, an external calibrator that estimates the correctness of responses from black-box API models. Trained jointly on responses from seven LLMs, it achieves 0.862 AUROC predicting the correctness of held-out responses from those same models, and shows zero-shot transfer to thirteen unseen models across eight organizations. Our model needs only a single forward pass to generate an uncertainty estimate and requires no access to the target model's logits, weights, or internal states. A lightweight text only 0.8B checkpoint matches our largest model's AUROC. We release code for adding uncertainty estimation to existing repos in only two additional lines of code.
A Global Comparison of Schemas, Transparency, and Interoperability in Public-Sector AI Registers and Inventories
公共部门人工智能登记册与清单的模式、透明度与互操作性全球比较研究
Das, Dipto, Guha, Shion
Abstract
Artificial intelligence (AI) registers and inventories aim to make governmental AI visible, but their institutional scope, schemas, and reporting practices construct different representations of public-sector AI. We compare 8,368 records from country-specific and transnational inventories covering 72 countries. Across 23 harmonized fields, registers shared a descriptive core but rarely requested information about appeals, risks, legal bases, or external evaluation. We found that broad schemas often contained substantial missingness, schema similarity showed no significant patterned convergence, and multiple sources covering the same jurisdictions overlapped only selectively. Based on these findings, we synthesize a layered visibility framework that shows how register records reflect disclosure arrangements and why interoperability requires shared concepts, clear definitions, and preserved provenance.
Scientific weak signals are early, low-visibility research directions that later become central to mature scientific topics, yet existing resources such as trend tracking, citation forecasting, and foresight reports rarely provide validated reference sets that link concrete early precursors to later paradigms. We introduce BackTrend, a retrospective benchmark in which, given a mature target topic and a temporal evidence constraint, systems must recover two types of precursors: problem-space signals, underrecognized research problems, and solution-space signals, emerging methods for known problems. BackTrend contains 25 mature target topics in artificial intelligence and machine learning and 66 human-validated weak signals, reconstructed from large-scale literature by grounding each candidate in its 2019-2024 publication-frequency trajectory. We evaluate frontier LLMs, RAG systems, and agentic research systems using semantic matching and coverage-based metrics. Current systems often generate plausible but misaligned precursors, exhibiting topic drift, granularity mismatch, near-miss matching, and incomplete coverage; the strongest system achieves only 10.1% F1, while Coverage10 reaches at most 18.5% of the reference signals. Our budget analyses show that additional retrieval and web-search evidence can improve performance up to a moderate budget, but does not by itself close the substantial performance gap.
Personal AI agents make recommendations and take actions on people's behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access to the user's personal context, e.g., their email inbox and a structured profile of personal attributes, with the intention of making an optimal, personalized decision for the user. We show that by simply providing this personal context, the agent steers recommendations based on inferred wealth, without being explicitly instructed to do so. In a suite of 325K experiments on 13 agents across three types of economic decisions (flights, health insurance, and graduate programs), we find that 8 models systematically choose more expensive options for wealthier users when requests are identical. This steering continues even when it directly goes against the user's stated objective: when explicitly instructed to find the cheapest option, some agents still act on the wealth profile they have inferred. It also occurs when wealth is inferred from ambient data, such as emails unrelated to the task. And it persists under privacy controls that block specific attributes: blocking financial attributes largely removes the disparity, but blocking other attributes leaves it unchanged and can increase it by up to 40% for insurance, as agents rely on the remaining signals to infer wealth. Larger and more capable models are no better; Claude Opus 4.8 shows the largest effect. We term this misalignment "adversarial delegation", in which the very conditions that make a personal AI agent useful - access to personal information - enable it to act against the user's interests.
Chinese Translation
个人AI智能体(Personal AI agents)在高风险经济场景中代表用户做出推荐并采取行动,例如购买机票、选择健康保险或挑选研究生项目。智能体被授予访问用户个人上下文的权限,例如其电子邮件收件箱以及结构化的个人属性资料,目的是为用户做出最优的、个性化的决策。我们表明,仅仅提供这些个人上下文,智能体就会基于所推断的财富水平来引导推荐方向,而无需任何明确指令。在涵盖三类经济决策(机票、健康保险和研究生项目)、涉及13个智能体的32.5万次实验中,我们发现当请求完全相同时,有8个模型会系统性地为更富裕的用户选择更昂贵的选项。即使这种引导直接违背用户明确陈述的目标,它仍然存在:当被明确要求寻找最便宜的选项时,某些智能体仍会依据其推断出的财富画像采取行动。当财富是从与任务无关的环境数据(如电子邮件)中推断出来时,这种现象也会发生。而且在阻断特定属性的隐私控制下,这种现象依然持续:阻断金融属性基本可以消除这种差异,但阻断其他属性则使其保持不变,甚至可能使保险场景中的差异增加多达40%,因为智能体会依赖剩余的信号来推断财富。更大、能力更强的模型表现并不更好;Claude Opus 4.8表现出最大的效应。我们将这种对齐失当称为"对抗性委托"(adversarial delegation):使个人AI智能体有用的那些条件——即对个人信息的访问——恰恰使其能够做出违背用户利益的行为。
Emergent Collusion in Long-Horizon LLM Agent Interaction
长时程LLM智能体交互中的涌现性合谋
Shi, Xinrui, Zhang, Yanzhe, Yang, Diyi
Abstract
LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. We introduce realistic constraints that make compliance with the verification protocol incompatible with reward maximization, and find that agents increasingly deviate from the protocol over repeated interactions. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier. Controlled peer interventions show that collusion is shaped by peer behavior, while ablations reveal additional effects of reward structure, the verification feedback agents receive, and their interaction history. In particular, restricting the amount and scope of interaction history available to agents reduces collusion. Overall, our findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks.
Harness-Zero: Harness Distillation via Agent-as-Harness
Harness-Zero:基于Agent-as-Harness的脚手架蒸馏
Ye, Haoran, Lu, Yuxing, Dong, Haonan, Su, Zhaochen, Song, Guojie
Abstract
Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose agent must either settle for a suboptimal shared harness or route among an ever-growing set of specialized ones. We therefore study agent harness distillation: using a domain- or instance-optimized harness as training-time guidance and transferring the behaviors it induces into model weights, so that its gains survive under a single fixed target harness. The challenge is that the two harnesses differ in action space and available information, so guidance from the optimized harness cannot serve directly as supervision for the target one. We introduce Harness-Zero, which enables harness distillation through agent-as-harness. Guided by the optimized harness, a harnessing agent corrects student responses before execution in the target harness's action space, turning harness guidance into training demonstrations. Fine-tuning on the resulting trajectories internalizes harness-induced behavior into the model, so the specialized harness can be removed at deployment. Our experiments spanning knowledge work, tool use, and science domains show that: (1) For frontier LLMs using the same evolved harness, agent-as-harness outperforms code-as-harness. (2) With the specialized harness removed at deployment, Harness-Zero improves the base model's macro-average task success from 23.3% to 44.3%, even exceeding the 41.7% it reaches with that harness still attached. (3) Harness-Zero recovers harness-induced behaviors absent from the base model, with 82.3% average recovery across 28 patterns in the three domains.
Did You Steal My Shot? Pioneering Camera Motion Plagiarism Detection in Generative Videos
你窃取我的镜头了吗?生成式视频中相机运动抄袭检测的开创性研究
Zhang, Chengguo, Ping, Ping
Abstract
Camera motion often reflects directorial intent and requires professional equipment, making it a high value form of intellectual property. However, generative video models can imitate such high value camera motions with simple prompts, while existing similarity detection methods mainly operate on visual content and fail to capture deeper motion similarity. This is mainly because their training data entangles camera motion with visual content. Moreover, traditional optical flow is insufficient to represent complex camera motions. We therefore build the first benchmark for camera motion analysis, including a motion dataset with \textbf{11} motion styles and evaluation protocols. Furthermore, we propose a motion representation that augments optical flow with vorticity cues from fluid dynamics, thereby better capturing motions. Experiments show that our detector achieves a \textbf{3.02*} improvement in plagiarism detection over the strongest baseline and remains effective on generative videos. We believe our work extends copyright protection beyond static content to dynamic camera motion.
Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework
面向多模态卒中复发预测的视觉与跨模态学习:一种可解释的两步框架
Gapp, Christian, Tappeiner, Elias, Welk, Martin, Fritscher, Karl, Mangesius, Stephanie, Eisenschink, Constantin, Deisl, Philipp, Knoflach, Michael, Grams, Astrid E., Gizewski, Elke R., Schubert, Rainer
Abstract
Multimodal stroke recurrence prediction requires effective integration of heterogeneous clinical and imaging data, yet modality imbalance often causes models to over-rely on dominant modalities and underutilize complementary information. While self-supervised pretraining and selective parameter freezing are commonly employed to improve representation learning and fine-tuning stability, their effect on modality contributions and cross-modal behavior in multimodal medical models remains largely unexplored. In this work, we investigate whether image pretraining on 3D CTA scans reduces modality imbalance and improves cross-modal integration for stroke recurrence prediction, a clinically critical task we recently addressed. To this end, two multimodal neural networks are pretrained in a self-supervised manner and subsequently fine-tuned using two distinct freezing strategies. Their performance and modality utilization are compared against both the baseline model from our previous work and models trained entirely from scratch in this study. Our results demonstrate that self-supervised pretraining enables more effective utilization of the multimodal image-tabular dataset, outperforming both the prior baseline and all non-pretrained models. Notably, the best-performing Vision Transformer based neural network successfully overcomes unimodal collapse. Synergy analysis reveals significant interactions between vision and both gender and CHD, suggesting clinically relevant patterns for stroke recurrence. Overall, our findings demonstrate that self-supervised pretraining and strategic fine-tuning support more balanced modality utilization and enable meaningful cross-modal interactions. Code is publicly available at https://github.com/ChristianGappGit/SSL_Pretraining.
We formulate \emph{Artistic Intelligence} as exploration driven world realization, leaving space for creative possibility while preserving the semantic, artistic, and compositional structure that must remain true. Moonworks Lunara, a text-to-image model, implements this framework with a novel Diffusion Mixture Transformer architecture. A new training algorithm iteratively evolves the data distribution through informative sample acquisition and targeted injection of human-created art. We benchmark Lunara against seven image-generation models, including FLUX.2-Klein-4B, Qwen-Image (20B), and GPT-Image-1-Mini. With GPT-5.6 Sol as evaluator, Lunara ranks first in \emph{Aesthetic Quality (8.473 vs. 8.457 GPT-Image-1-mini)}, second in \emph{Emotional Resonance}, and remains competitive in \emph{Content Integrity}. A blind human evaluation over the same evaluation set corroborates the automated metrics, ranking Lunara first. It also stays among the strongest models under conventional measures including CLIPScore and LAION Aesthetic Predictor. On GenEval, Lunara achieves competitive performance against a broader set of 16 models, including GPT Image 2 and Seedream 4.0. These results place Lunara at the frontier with Artistic Intelligence while maintaining a sub-10B active-parameter footprint and sub-10-second inference latency. Lunara advances the general visual intelligence frontier by shifting the question from whether models can get images right to how deeply they can interpret meaning and realize it as imaginative, expressive worlds.
Foundation models have recently demonstrated strong capabilities across a wide range of medical imaging tasks. However, their performance in structured clinical interpretation settings remains insufficiently explored. In lung cancer screening, interpretative variability persists despite standardized frameworks such as Lung-RADS. In this study, we evaluate MedGemma, a medical general-purpose foundation model derived from Gemini and its fine-tuned version adapted for lung cancer detection and diagnosis, compared against radiologists performing Lung-RADS v2022 assessment on the NLST dataset. Twelve radiologists independently evaluated each case in a multi-reader design, enabling quantification of inter-reader variability. Radiologists achieved a mean AUC of 0.90, with substantial variability across readers (range: 0.80-0.94). The native foundation model achieved an AUC of 0.70, failing to reach clinically relevant performance. In contrast, fine-tuning significantly improved performance to an AUC of 0.83, placing the model within the lower range of individual radiologists performance. These findings highlight a trade-off between peak accuracy and prediction consistency. Unlike radiologists, under fixed conditions, the model produces deterministic outputs, removing inter-run variability under identical inputs, in contrast to inter-reader variability observed among radiologists. This supports the role of fine-tuned foundation models potential complementary tools for clinical decision support, particularly in settings with limited expertise. However, evaluation is performed on a case-enriched cohort from NLST and does not account for real-world prevalence or external validation, limiting direct clinical generalization.
Brain-to-Image Generation: Reconstructing Visual Stimuli from EEG using Generative Adversarial Networks
脑信号到图像生成:利用生成对抗网络从脑电信号重建视觉刺激
Goyal, Harshit
Abstract
Reconstructing visual stimuli from electroencephalography (EEG) is difficult because scalp measurements have high temporal but limited spatial resolution, and paired EEG-image datasets remain small relative to modern generative-model training corpora. We present a reproducible single-subject baseline on THINGS-EEG2 that first tests the more defensible question of whether EEG can retrieve the viewed stimulus in a visual embedding space. A compact temporal-spatial convolutional encoder maps repetition-averaged EEG (63 by 250) to provided 512-dimensional ViT-B/32 image features. Model selection uses a concept-disjoint validation split, and final evaluation uses the official 200-image, 200-concept test gallery. Across three training seeds, the model obtains 12.83 +/- 0.58%, 39.17 +/- 1.76%, and 58.00 +/- 1.73% image recall at 1, 5, and 10 (mean +/- sample standard deviation), compared with analytical chance levels of 0.5%, 2.5%, and 5.0%. A session-balanced ablation shows that averaging more test repetitions generally improves ranking. Applying the Subject 01 model to the other nine subjects without adaptation causes a sharp performance drop, exposing subject specificity. We further report exploratory stress tests of direct conditional generators trained without external visual weights: single-subject and ten-subject variants produce noise-dominated outputs, with early validation improvements reversing after one to four epochs. Finally, we distinguish direct reconstruction from semantic rendering with a pretrained diffusion prior. The results support above-chance coarse semantic decoding under a closed-set, repetition-averaged protocol, but do not support faithful recovery of stimulus pixels.
Rethinking Streaming Video Diffusion Model: Context, Execution, and Training
重新思考流式视频扩散模型:上下文、执行与训练
Zhang, Hongchen
Abstract
Understanding the design space of streaming video diffusion is essential to exploring its potential for generation quality and computational efficiency. We develop a unified analytical framework that relates model and sampler choices, historical conditioning, execution scheduling, and training strategies. The framework accommodates a broad family of causal context-selection policies and makes their computational dependencies and training-inference alignment explicit. Within this design space, we study three representative policies: clean, same-level, and progressive history. On the full VBench prompt set, same-level and progressive history achieve aggregate scores of 85.24 and 85.60, respectively, compared with 84.45 for the clean-history reference. Long-video comparisons further show improved subject consistency and more coherent motion with progressive history. By allowing multiple denoising nodes to be processed together, progressive-history pipelining achieves $1.57$-$2.83\times$ steady-state DiT speedups under our evaluated conditions. We additionally find that LoRA adaptation of the DMD fake-score network improves generation quality using only 2.15% as many trainable fake-score parameters as full-parameter adaptation. Together, these findings show that fully denoised history is not a prerequisite for high-quality streaming generation and motivate the joint design of historical conditioning, execution, and training.
Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection
基于rPPG衍生与唇部区域频率线索互补的说话人脸深度伪造检测
Harraq, Othmane, Aldwairi, Tamer
Abstract
Talking-face (TF) deepfakes are detected unevenly by rPPG-based methods across generators. We study two lightweight visual-only cues, rPPG-derived waveforms extracted by RhythmFormer and lip-region discrete cosine transform (DCT) coefficients, on the seven TF methods of Celeb-DF++ under a subject-independent protocol. In-domain, lip-region DCT matches or exceeds the rPPG-derived 1D ResNet on every method except SadTalker, and Concat fusion reaches AUC 0.891 against 0.824 and 0.827 for the unimodal baselines. Under leave-one-generator-out evaluation the cues split: each transfers clearly better to three held-out methods, and IP-LAP is near chance for both. Concat averages 0.798 but falls below rPPG alone where DCT transfers poorly, so static fusion only partly exploits this complementarity. Lip-region DCT outperforms full-face DCT on six of seven methods. We treat the rPPG-derived signal as an empirical cue and do not claim it is cardiac in origin.
This paper presents a comprehensive experimental evaluation and detailed analysis of state-of-the-art multi-object tracking algorithms, with an emphasis on quantifying the individual contributions of detection and association components to overall tracking performance. Unlike existing surveys that primarily offer theoretical categorizations or taxonomies of tracking methods, our work adopts a rigorous experimental perspective grounded in publicly available implementations, providing practical guidance for researchers and practitioners in method selection and system design. We introduce a unified pipeline diagram that consolidates the core components across the two main branches of visual multi-object tracking: tracking-by-detection and end-to-end deep learning paradigms, and systematically analyze the object detection, feature extraction, and data association modules. Through extensive empirical studies on standard benchmarks, including MOT16, MOT17, MOT20, SportsMOT, DanceTrack, and CrowdTrack datasets, we reveal critical insights: (1) detection quality dominates association strategy performance, with detector improvements yielding more than 10% gains compared to less than 5% from refined association strategies; (2) modern deep learning detectors paired with specialized re-identification models significantly outperform joint detection and embedding approaches; and (3) transformer-based end-to-end methods exhibit greater robustness to detection quality variations but at a substantial computational cost. Our findings from extensive experiments provide key insights into component-level effects in MOT, particularly the dominant influence of detection quality relative to association, while offering practical insights for designing and optimizing MOT systems under varying performance and robustness requirements. Code and experimental setups are available at github.com/linh-gist/VisualMOT.
Vision-language models (VLMs) and vision-language-action models (VLAs) are increasingly deployed in real-world applications. There, a small perturbation to the recorded camera image may change a decision significantly. However, existing benchmarks for these models only sample perturbations, which does not guarantee the absence of a failure in the untested region. We present the first robustness validation of six VLMs (drawn from the Gemma, InternVL, LLaVA, and Qwen families) and five VLAs (drawn from the GR00T, OpenVLA, and $\pi$ families) over entire continuous regions of photometric and geometric image perturbation: brightness shifts, camera rotations, and their composition. To this end, we build on the validation framework H$^2$V and introduce H$^2$V-M, a margin-aware convergence rule that makes validation affordable at the 32B parameter scale. We demonstrate that H$^2$V-M outperforms H$^2$V by an order of magnitude in model queries and that it finds counterexamples faster than random sampling while providing soundness guarantees. Our VLM and VLA robustness validation shows that robustness is mostly dependent on the perturbation type, rather than the model, and that VLMs are more robust to large camera rotations than VLAs. For VLAs, even perturbations as small as $\pm1^\circ$ can change the commanded action in many cases. We also show that robustness depends more on model family than on model size.
On The Robustness-Resolution Tradeoff In Temporal Quantization Of Event Streams
论事件流时间量化中的鲁棒性-分辨率权衡
Chowdhury, Sayeed Shafayet, Sharmin, Ruhi
Abstract
Event pipelines often discretize asynchronous timestamps before learning. This step looks harmless, but its stability depends directly on temporal resolution. We study this dependence at the representation level. We first show that hard temporal binning is discontinuous: an arbitrarily small timestamp shift near a boundary can move unit event mass between bins. We then define a class of nonnegative, mass-preserving, resolution-faithful continuous encoders and prove that every encoder in this class has global L1 sensitivity at least 2/Delta, where Delta denotes bin width. Linear two-bin interpolation attains this limit. Local support and first-moment preservation also make it unique. Experiments on SHD, N-MNIST, and DVS128 Gesture support the analysis. Across uniform timestamp budgets, linear interpolation lowers mean representation drift by 47-72% while keeping clean accuracy nearly unchanged. On DVS Gesture, it produces zero prediction flips across all tested budgets and three seeds. On SHD, measured drift follows 1/Delta with R^2 = 0.992.
Authority-Preserving Evaluation of Medical Vision-Language Assistants
医学视觉语言助手的权限保留式评估
Fan, Flint Xiaofeng, Tan, Cheston, Ong, Yew-Soon, Wattenhofer, Roger
Abstract
Medical vision-language models can propose how urgently a skin lesion should be reviewed, but the local service retains authority to accept or replace that proposal under referral policy, capacity, and locally held patient context. Proposal quality and selected-action quality are therefore distinct evaluation targets, and benchmark evidence transfers between them only when local review preserves the expected action score. We introduce AuthEval, a logging and evaluation framework that records both actions, scores the selected action under declared local criteria, and, where feasible, scores the declined proposal under the same rule. It reports the resulting authority gap only when the record supports it. Because the gap is the product of the proposal-change rate and the mean score change on changed cases, that rate alone determines neither its magnitude nor its sign. On ISIC 2019, with MedGemma and simulated local review, two constraint regimes with similar change rates produced an optimistic image-equal gap under capacity ($+0.744$ simulator units) but no detectable gap under safety. The declared evaluation unit also mattered: under mixed constraints the gap reversed from $+0.374$ to $-0.206$ when weighting shifted from image to lesion-aware cluster. AuthEval thus clarifies whether a study's records support claims about the model, the workflow, or both.
Coding-agent benchmarks usually evaluate implementation after the target behavior has been specified in text, code, or demonstrations. Existing research has extensively evaluated the ability of coding agents to generate programs from textual specifications. However, under black-box conditions where neither source code nor documentation is available, it remains underexplored whether an agent can induce the rules solely through visual observation and active interaction and reproduce the target system as a verifiable executable system. To this end, we present GameReplica, a closed-loop evaluation framework for end-to-end black-box game replication that covers the full perception, exploration, induction, reproduction, and verification pipeline. GameReplica comprises 125 tasks spanning 25 games across 5 core mechanism families, with each game instantiated at five difficulty levels. The tasks require an agent to access the target game only through screenshots and an action interface, induce the key visual elements and gameplay rules from pixel feedback and interaction outcomes, and generate a self-contained, runnable game replica that can be automatically verified by an external program. Experiments show that current coding agents still face substantial challenges in end-to-end black-box replication: the best-performing model (Claude Opus 4.8) achieves an overall score of 71.6\%, while the remaining models score only 4.0\%--42.9\%. Further analysis reveals a consistent pattern across all models: visual-fidelity scores are substantially higher than implementation- and rule-consistency scores, indicating that agents replicate visual appearance more readily than game mechanics. The difficulty levels further amplify the performance gap: from L1 to L5, the overall score of weaker agents drops sharply, whereas that of the best-performing agent declines only slightly.
Chinese Translation
编码智能体基准测试通常在目标行为已通过文本、代码或演示明确指定后评估其实现能力。已有研究广泛评估了编码智能体根据文本规范生成程序的能力。然而,在既无源代码也无文档可用的黑盒条件下,智能体能否仅通过视觉观察和主动交互归纳出规则,并将目标系统复刻为可验证的可执行系统,这一问题仍未得到充分探索。为此,我们提出了GameReplica,一个面向端到端黑盒游戏复刻的闭环评估框架,涵盖完整的感知、探索、归纳、复刻与验证流程。GameReplica包含125个任务,涵盖5大核心机制类别下的25款游戏,每款游戏设有五个难度级别。任务要求智能体仅通过截图和动作接口访问目标游戏,从像素反馈和交互结果中归纳关键视觉元素与游戏规则,并生成一个独立可运行、可由外部程序自动验证的游戏复刻版本。实验表明,当前编码智能体在端到端黑盒复刻任务上仍面临巨大挑战:表现最好的模型(Claude Opus 4.8)总分为71.6%,而其余模型的得分仅为4.0%–42.9%。进一步分析揭示了所有模型的一致规律:视觉保真度得分显著高于实现一致性和规则一致性得分,表明智能体更容易复刻视觉外观而非游戏机制。难度级别进一步放大了性能差距:从L1到L5,较弱智能体的总分急剧下降,而表现最好的智能体仅略有下降。
Automated segmentation of CT images has become increasingly important to enhance the reliability of simulations through the generation of high fidelity numerical models. This study addresses the challenging task of semi-automatically tracking textile reinforcements in fan blade dry preforms using X-ray CT images captured at coarse resolutions (i.e., above 140 $\mu$m). Our approach offers a scalable, slice-based analysis conducted on planes orthogonal to the main yarn directions, applied to a large-scale real industrial component. This enables accurate identification and tracking of yarn paths while requiring minimal training. The method models three key yarn properties statistically: their typical cross-section shape, their continuity and movement in the 3D space, and their spatial relative arrangement with respect to neighboring yarns. These statistical properties are integrated into a tracking framework via a variational formulation that optimizes all yarn center positions in successive cross-section planes. The method tracks more than 3,000 warp yarns across 1,500 slices and achieves a tracking success rate above 90%. Overall, this work demonstrates a promising approach toward large-scale, automated textile reinforcement annotation, paving the way for more efficient material characterization in complex composite structures.
ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning
ALPINE:面向参数与样本高效少样本学习的自适应定位方法
Yadav, Neeraj
Abstract
Few-shot learning research is predominantly evaluated on accuracy alone, with limited attention to the parameter and training-sample budgets required to reach that accuracy - a real constraint for practitioners without large-scale compute. We present an ultra-lightweight (22,249-34,917 parameter) spatial-relational architecture for few-shot image classification that combines fixed Gabor edge-energy guidance with a windowed, content-adaptive patch locator. Under a strictly matched, iso-episode-budget protocol (250 meta-training episodes, 5 canonical seeds, 600 evaluation episodes per seed), our architecture achieves 5-shot accuracy gains, consistent across all five seeds, over Prototypical Networks, Relation Networks, and MAML on both CIFAR-FS and MiniImageNet, while using 27-53% fewer parameters than any baseline. It also converges in fewer training episodes, generalizes better to an unseen fine-grained domain (CUB-200-2011 birds, zero retraining), and is more robust to 50% occlusion and 25% spatial translation than all three baselines. A series of falsification ablations - zeroing relational tokens at inference and retraining without them entirely - shows that the architecture's pairwise relational computation, while present, is not the primary driver of its performance; the content-adaptive patch locator is. We report this honestly, together with a capacity sweep showing a genuine accuracy plateau near 22-35k parameters, and release full seed-level results and checkpoint hashes for reproducibility.
Dimensionality reduction for AI based hyperspectral image classification based on XAI
基于可解释人工智能(XAI)的AI高光谱图像分类降维方法
Zeljković, Vladimir, Stojanović, Branka, Ganster, Harald, Nešković, Aleksandar
Abstract
This research addresses the challenge of limited material recycling in wood recycling processes by leveraging artificial intelligence (AI)-based dimensionality reduction. Our study explores the application of convolutional neural networks (CNNs) in multi-channel hyperspectral imaging (HSI), extending beyond RGB channels to over 200 spectral channels. Dimensionality reduction within this context involves streamlining the feature space for AI system training and inference. Focusing on explainable AI (XAI) methods, this paper contributes to a broader research initiative, presenting a solution framework that enhances the sustainability and efficiency of wood recycling processes.
Hierarchical Aggregation of Semantic Uncertainty in 3D Scene Graphs
三维场景图中语义不确定性的分层聚合
Zumaya, Carlos Cueto, Catalano, Iacopo, Bessa, Wallace Moreira, Placed, Julio A.
Abstract
Open-vocabulary 3D Scene Graphs (3DSGs) ground each object node in a vision-language embedding, yet they record every entry as equally certain, so a robot querying the map cannot tell which of its entries are unreliable. Estimators of semantic uncertainty could supply that distinction, but they require repeated sampling of a model, training, or held-out labels, none of which are available to a deployed system at query time. We present a framework that exploits the detector confidence and the embeddings a 3DSG already stores, converts them into a probability that an entry is correct, and propagates that probability through the containment hierarchy into a belief that a room contains a queried class. Four signals, each paired with the object-level error it indicates, are converted to probabilities at the logit scale learned by the vision-language model and combined in closed form with no additional perception or training. Objects sharing a detector and a vocabulary fail together, so the framework aggregates them in the fully correlated limit, where an aggregation under independence would treat one repeated error as repeated evidence. Evaluated on HM3DSem against a state-of-the-art 3DSG system, the framework improves object retrieval and lowers the error of the room-level assertions of the graph it reads.
Chinese Translation
开放词汇三维场景图(3D Scene Graphs, 3DSGs)将每个物体节点锚定于视觉-语言嵌入中,但其对所有条目均以同等确定性记录,因此查询地图的机器人无法判断哪些条目是不可靠的。语义不确定性估计器可以提供这种区分,但它们需要对模型进行重复采样、训练或预留标签,而这些对于部署系统在查询时均不可用。我们提出一个框架,利用三维场景图已经存储的检测器置信度和嵌入,将其转换为某条目正确的概率,并通过包含层级传播该概率,形成关于某房间包含所查询类别的信念。四种信号各自与它所指示的物体级错误配对,在视觉-语言模型学习到的logit尺度上被转换为概率,并以闭式形式组合,无需额外的感知或训练。共享同一检测器和同一词汇表的物体会一同出错,因此该框架在完全相关的极限下对它们进行聚合;而在独立性假设下的聚合则会将一个重复出现的错误视为重复的证据。在HM3DSem数据集上与最先进的三维场景图系统相比,该框架提升了物体检索性能,并降低了其读取的图中房间级断言的误差。
MarsRecon: Self-Supervised and Multimodal Surface Representations for Mars
MarsRecon:面向火星的自监督多模态表面表示学习
Naik, Akshay, Juston, Marius F. R., Mahajan, Jay
Abstract
High-resolution orbital imagery offers a rich record of the Martian surface, but sparse geological labels limit supervised representation learning. We present MarsRecon, a geospatially aware pipeline for learning visual and multimodal representations from HiRISE observations of Olympus Mons. The pipeline calibrates NASA Planetary Data System products, extracts valid georeferenced patches, and trains a masked autoencoder on unlabeled imagery. Increasing input resolution and filtering invalid tokens reduced held-out reconstruction loss from 0.1751 to 0.1342 in the principal Stage A model series. We then freeze the visual encoder and align its features with observation text, coordinates, and local--global image context. The strongest current local-primary model achieves image-to-text recall@10 of 0.3787, text-to-image recall@10 of 0.9161, and local-to-global recall@10 of 0.4350 on the held-out test split. These results establish a working Mars-specific pretraining and retrieval pipeline; further crop-overlap controls and downstream geological evaluations are needed to assess the broader utility of its embeddings.
Chinese Translation
高分辨率轨道影像为火星表面提供了丰富的记录,但稀疏的地质标注限制了有监督表示学习。我们提出了MarsRecon,一个具有地理空间感知能力的流水线,用于从奥林匹斯山(Olympus Mons)的HiRISE观测数据中学习视觉与多模态表示。该流水线对NASA行星数据系统(Planetary Data System)产品进行校准,提取有效的地理配准图像块,并在无标注影像上训练掩码自编码器(masked autoencoder)。在主要的Stage A模型系列中,提高输入分辨率并过滤无效token使留出集重建损失从0.1751降至0.1342。随后,我们冻结视觉编码器,并将其特征与观测文本、坐标以及局部—全局图像上下文进行对齐。当前最强的以局部为主的模型在留出测试集上实现了图像到文本recall@10为0.3787、文本到图像recall@10为0.9161,以及局部到全局recall@10为0.4350。这些结果建立了一个可运行的火星专用预训练与检索流水线;仍需进一步的裁剪重叠控制与下游地质评估来衡量其嵌入的更广泛实用性。
Image steganography hides secret message within normal images, with most existing works relying on cover-preserving transmission. However, such a paradigm becomes vulnerable once the original cover is exposed or can be reliably approximated. In this paper, we propose StyleStegaNet, a stylized image hiding framework that replaces cover matching with style-concealment transmission. Instead of transmitting a cover-like stego image, StyleStegaNet generates stylized stego images conditioned on publicly available style references, redefining steganography invisibility from cover-preserving concealment to behavior-level camouflage based on style transformation. Such a setting poses a substantial challenge to reliable secret recovery, since neural stylization can significantly alter the feature statistics exploited by deep hiding methods. To address this challenge, StyleStegaNet decouples the overall task into four coordinated stages: stego generation, stylized transmission, structure-preserving reconstruction, and secret recovery. Moreover, StyleStegaNet is optimized with a progressive three-stage training strategy, in which wavelet-domain constraints and perceptual supervision guide the recoverable information toward structural representations. We further provide an analysis showing that secret recoverability is largely restricted to the normalized structural subspace, offering a mechanistic explanation for why directly stylized baselines fail and why a reconstruction-guided recovery path is necessary. Extensive experiments on DIV2K and MS-COCO datasets demonstrate the effectiveness of StyleStegaNet. And few-shot image steganalysis with two deep detectors further shows detection accuracy near random guessing, approximately 51\%.
We address the problem of recovering high-speed videos from dynamic scenes under extreme photon sparsity. Existing methods rely on aggregating photon detections in local spatiotemporal windows to improve signal-to-noise ratio; however, this local grouping discards global structure and fails in low-light regimes where photon detections are sparse in space and time. In this work, we show that the information needed to recover both motion and illumination is encoded in correlations over the full space-time pattern of photon arrivals. Building on this insight, we develop a spatiotemporal flux probing theory and an algorithm that estimates the Fourier coefficients of the underlying intensity directly from the photon stream. We demonstrate that our approach (1) recovers fast motion and temporal illumination dynamics with substantially fewer photons than prior methods, (2) enables velocity-selective videography that automatically refocuses video onto specific detected motions, and (3) generalizes across sensing modalities including single-photon, event, and spike cameras.
Autonomous navigation requires precise and efficient semantic segmentation, yet existing frame-based approaches remain limited by motion blur, glare, latency, and the low temporal resolution (20-30 FPS) of conventional cameras, which leads to information loss between frames. Event cameras have emerged as an alternative sensing modality, capturing intensity changes asynchronously with high temporal resolution, high dynamic range, and sparse outputs. However, event-based algorithms still fall short of frame-based ones in accuracy, as most segmentation methods are designed for dense frame data. To overcome these limitations, we propose a hybrid vision architecture that combines conventional frame-based and event-based cameras. The system integrates two complementary components: (1) a compact Spiking Neural Network (SNN) with 42k parameters for motion estimation, and (2) a lightweight event-driven SNN with 0.84M parameters for frame-based semantic segmentation, which interpolates motion between frames to refine segmentation results. By predicting inter-frame segmentations, the framework achieves segmentation rates of up to 500 Hz with an energy consumption below 1.87 mJ per inference, while maintaining real-time GPU execution at frequencies up to 200 Hz. Additionally, our approach compensates for information loss in frames affected by blur or overexposure, enabling more robust perception in challenging conditions.
Rethinking Vision Architectures with Gated Linear Attention and KAN
基于门控线性注意力与KAN的视觉架构再思考
Mehizel, Ali, Khaldi, Oussama
Abstract
Vision Transformers allocate most parameters to multi-layer perceptrons (MLPs) for channel mixing, while token interactions usually rely on quadratic multi-head self-attention (MHSA). Linear attention reduces sequence complexity to O(N), but remains coupled with the same fixed-activation MLP as softmax Transformers. Kolmogorov-Arnold Networks (KANs) instead place learnable univariate maps on edges, yet prior vision KANs keep MHSA or omit attention entirely. We introduce LKAT (Linear Kolmogorov-Arnold Transformer), an isotropic ViT encoder that couples chunkwise Gated Linear Attention (GLA) with a two-layer KAN feed-forward, and we provide an I/O-aware fused RBF-KAN kernel for the radial-basis grid maps. Under a shared DeiT-style recipe we compare LKAT with ViT, ViT-5, and MLP-Mixer. LKAT-B exceeds ViT-B/16, ViT-5-B, and Mixer-B/16 on ImageNet-100. Tiny/Small/Base LKAT variants scale consistently on CIFAR-10/100, and ImageNet-100 pretraining transfers to CIFAR fine-tuning. The results support gated linear attention and KAN-based radial-basis functions as complementary inductive biases for mid-scale visual representation learning.
Chinese Translation
视觉Transformer将大部分参数分配给多层感知机(MLP)用于通道混合,而词元交互通常依赖于二次复杂度的多头自注意力(MHSA)。线性注意力将序列复杂度降低至O(N),但仍然与与softmax Transformer相同的固定激活MLP耦合在一起。Kolmogorov-Arnold网络(KAN)则将可学习的单变量映射置于网络边上,然而先前的视觉KAN要么保留MHSA,要么完全省略注意力机制。我们提出了LKAT(Linear Kolmogorov-Arnold Transformer),一种各向同性的ViT编码器,它将分块门控线性注意力(Gated Linear Attention, GLA)与两层KAN前馈网络相结合,并为径向基网格映射提供了一个I/O感知的融合RBF-KAN核。在共享的DeiT风格训练方案下,我们将LKAT与ViT、ViT-5和MLP-Mixer进行比较。在ImageNet-100上,LKAT-B超越了ViT-B/16、ViT-5-B和Mixer-B/16。Tiny/Small/Base版本的LKAT变体在CIFAR-10/100上表现出一致的扩展性,并且ImageNet-100的预训练可以迁移到CIFAR的微调任务中。结果表明,门控线性注意力与基于KAN的径向基函数可以作为互补的归纳偏置,用于中等规模的视觉表征学习。
AdaMerge: Tuning-Free Patch Compression for Multi-Vector Visual Document Retrieval
AdaMerge:面向多向量视觉文档检索的无调优补丁压缩方法
You, Jianxin, Ni, Kun
Abstract
Multi-vector visual document retrieval (VDR) models such as ColPali and ColNomic achieve strong accuracy by representing each document with hundreds to thousands of patch-level embeddings, at substantial storage and latency cost. Existing compression methods either prune unimportant patches or merge similar ones into clusters; the recent state-of-the-art merging method Prune-then-Merge (PtM) consistently outperforms pruning-only baselines at high compression, but requires a per-dataset cluster budget m to be tuned by grid search. We observe that the merge-cosine sequence produced by hierarchical clustering exhibits a sharp cliff separating mergeable redundancy from salient signal, and that the location of this cliff is concentrated in a narrow band across more than 11,000 documents from 14 datasets. This suggests the merge boundary can be detected per document rather than tuned per dataset. Building on this observation, we propose AdaMerge, a plug-and-play compression method that (i) detects each document's own cliff via gap analysis on the merge-cosine trajectory, and (ii) builds attention-weighted cluster centroids to preserve salient signal. On the long-document benchmark ViDoRe-V2 (4 datasets, two backbones), AdaMerge significantly outperforms tuned PtM across the operating range (p < 10^-4); on the short-document benchmark ViDoRe-V1 (10 datasets, two backbones), where all merging methods are already near-lossless, AdaMerge matches tuned PtM without any per-dataset tuning. AdaMerge adds only about 10 ms per document and exposes a single global hyperparameter shared across all datasets and backbones.
End-to-end and vision-language-action (VLA) driving policies are compared by leaderboard rank, but a rank reports an outcome, not the behaviour behind it, so it predicts poorly how a policy will behave at a new site. On six released policies, rank on nuScenes open-loop error or on NAVSIM's leaderboard does not carry over to scenes with a pedestrian near the ego corridor at a new site. We propose a counterfactual check-up: a few hundred real frames, each edited two ways (pedestrian removed, or re-lit by a night-style perturbation), every edit verified by an independent detector, and the change in the planned trajectory read as a diagnosis rather than a score. From these edits two causal axes are read, and five exams built on them separate what a score merges: how far the policy plans to drive, whether seeing the pedestrian buys safety, whether that response scales with danger, whether the plan moves when nothing requires it, and how much an irrelevant lighting change moves it. On 246 NAVSIM near-pedestrian scenes, in the cells where the pedestrian lies on the planned path only 1.9% of responses are genuine avoidance, and under our open-loop protocol the median clearance change is at most 0.03 m and the median change in planned distance at most 0.08 m for every policy. In a pre-registered test from left- to right-hand drive, the exposure and specificity orderings, the lighting verdict and the collision outcome transfer, while point values and the hazard-sensitivity verdict do not. Read as a selection report, the profiles say which policy is safe because it plans short, which covers a human-like distance without yielding, and which is unsteady under a change that requires no reaction, and they price each verdict: most settle within a few dozen frames, hazard sensitivity needs hundreds. Code and edited frames will be released.
Vision-language models (VLMs) perform strongly on visual question answering benchmarks, yet often make decisions that contradict visual evidence they have already identified correctly. We distinguish perceptual failure, where relevant evidence is not recognized, from process failure, where recognized evidence fails to constrain the final decision. We introduce VPAC-Bench, a benchmark spanning nine real-image process families, with each image annotated by its current activity stage and nearby stage transition. We also propose State-Relevance-Target (SRT), a family of structured process-prior interventions that requires models to connect visible evidence to the relevant process state before answering. Across multiple VLMs, process failure is widespread: models that correctly enumerate visual candidates still over-commit to a single answer in more than 95% of ambiguous cases. An explicit process-structured intervention reduces this rate to below 13% without degrading performance on unambiguous cases. However, the transfer of process priors is model-dependent, and generic SRT does not consistently outperform strong chain-of-thought baselines. When the relevant stage transition is known, boundary-aligned SRT substantially outperforms generic process prompting and all tested chain-of-thought baselines across assembly, physical state transition, navigation and traffic, and object-use affordance tasks. These results show that process priors are most useful when aligned with the scene's specific decision boundary, motivating boundary-aware prior selection for process-grounded visual reasoning.
X-Beat: An Explainable Framework for ECG Image Classification
X-Beat:一种面向心电图图像分类的可解释框架
Tahsin, Mohammad Sadman, Adarbah, Haitham Y., Noore, Afzel
Abstract
Accurate automated interpretation of electrocardio- grams (ECGs) is essential for early detection of cardiac condi- tions such as myocardial infarction and rhythm abnormalities. However, many high-performing deep learning models remain difficult to deploy in clinical settings due to limited transparency and lack of reliability validation. In this work, we present X- Beat, an explainable and reliability-aware benchmark framework for ECG image classification designed to support trustworthy AI systems in healthcare. The proposed framework combines transfer learning with post-hoc explainability and systematic reliability evaluation across four cardiac classes: Abnormal Heartbeat, History of Myocardial Infarction, Myocardial In- farction, and Normal Heartbeat. Multiple ImageNet-pretrained CNN backbones, including EfficientNet-B0, ResNet-50, DenseNet- 121, and MobileNetV3-Large, are evaluated under a unified training protocol. Beyond standard performance metrics, we incorporate Grad-CAM-based visual explanations together with additional analyses, including explanation stability, regional sen- sitivity, and confidence-based reliability assessment, to examine whether model predictions are supported by clinically meaningful evidence. Experimental results show that ResNet-50 achieves the best performance, reaching 91.94% accuracy and a macro F1- score of 0.9098, with strong class separability (AUC up to 0.995). Explanation analyses indicate that the model primarily focuses on waveform-relevant regions, while reliability evaluation suggests that most incorrect predictions occur with lower confidence. Overall, this work provides a structured and reproducible bench- mark for evaluating both predictive performance and explanation reliability in ECG image classification, contributing toward the development of trustworthy and interpretable AI components for clinical decision support systems.
Autoregressive video world models enable temporally coherent generation for a single observer. Extending them to multiple agents requires consistency across independently controlled views and temporal gaps under causal streaming. We present ConsistWorld, a multi-agent world model that generates camera-controlled video streams of a static scene from one shared image. We formulate consistency as routing evidence from committed multi-agent history and concurrently generated peer views to the tokens being generated. Pose Conditioned Memory Retrieval selects relevant historical observations from all agents, recovering evidence beyond the recent context window. Visibility-Gated Peer Sharing regulates current peer information according to estimated historical coverage and current-view overlap. Together, they determine which historical observations enter the context and where concurrent peer information contributes, supporting long-term recall and coordinated exploration. Both mechanisms use camera geometry and maintain a bounded active context for a fixed agent count and retrieval budget. Experiments on evidence sharing cases and video length and agent number generalizations show that ConsistWorld achieves a strong cross-time and cross-agent consistency while preserving competitive generation quality.
Math2Visual-X: A Modular Framework for Pedagogically Aligned Lower-Primary Math Visuals Generation
Math2Visual-X:一个面向教学对齐的小学低年级数学可视化图形生成的模块化框架
Maduranga, H. D. E., Munasinghe, S. K., Weerasekara, K. P. T. I., Ranathunga, Surangika, de Silva, Nisansa
Abstract
Visual representations can help lower-primary learners understand Math Word Problems, but generating classroom-usable visuals remains difficult. Existing symbolic systems are controllable but limited in coverage, while end-to-end text-to-image systems often fail to satisfy exact mathematical constraints. This paper presents a symbolic visual generation framework for lower-primary MWP generation with broader problem coverage and more scalable asset generation. The framework includes an LLM-based routing layer, three worksheet-oriented generation modules, and two fallback mechanisms for open-world SVG asset acquisition. A human evaluation comparing Math2Visual-X with Stable Diffusion XL, Nano Banana, and GPT Image showed that the proposed method achieved the strongest overall performance. The results indicate that the framework offers a scalable and pedagogically grounded approach for automatic MWP visual generation.
We present PanoSeg3R, a feed-forward framework for 3D panoramic semantic segmentation. Unlike existing methods designed for perspective inputs, PanoSeg3R jointly predicts 3D geometry and multi-view semantic segmentation in one single forward pass. Built upon a pretrained reconstruction backbone that supports panoramic images, our approach extends feed-forward 3D reconstruction with a query-based mask decoder. Furthermore, we introduce an automatic panorama data curation pipeline that leverages the complementary strengths of off-the-shelf foundation models to generate reliable pseudo semantic annotations, substantially expanding the training data and improving zero-shot generalization. PanoSeg3R achieves state-of-the-art performance on panoramic 3D semantic segmentation, improving 3D mIoU by up to 16.02 on ScanNet++, while the curated training data further improves zero-shot performance by up to 4.26 and 43.28 mIoU on Stanford2D3D and ToF-360, respectively. Website: https://harryyoon777.github.io/PanoSeg3R/
Generating parametric CAD models requires accurate geometry and stable feature dependencies. Existing methods face challenges in selecting geometric references, interpreting sketch-plane local coordinates, and establishing sketch constraints to projected external geometry. We present Vision2CAD, a visual agent harness that combines vision-language model (VLM) reasoning with deterministic CAD kernel operations. An ID-based interface supports explicit geometry selection, a local-coordinate bridge converts view coordinates into sketch coordinates, and projected-edge localization supports external sketch constraints. These mechanisms establish feature dependencies within the supported modeling operations and constraint types. We also introduce the Geometry Explicit Reference Dataset (GERD), which aligned commands, geometry states and IDs at every modeling step. On GERD-EVL and a DeepCAD test subset, Vision2CAD improves mIoU by 11.1\% and 5.6\% and reduces Chamfer distance by 17.3\% and 41.8\%, respectively. Parameter-editing experiments and ablation studies further proved the preservation of parametric dependencies.
DOA-SORT: Directional Occlusion-Aware Multi-Object Tracking with Distributional Observations
DOA-SORT:基于方向性遮挡感知与分布式观测的多目标跟踪
Wang, Hao
Abstract
Identity association in multi-object tracking (MOT) is vulnerable to partial occlusion, truncated detections, and fluctuating confidence scores. Existing motion-dominant trackers commonly represent occlusion as a scalar penalty. This treatment misses the directional observation bias caused by occlusion: left, right, top, and bottom occlusions distort the location and shape of a detection in different ways. We propose \ours{} (Directional Occlusion-Aware SORT), an online and training-free tracker that models these biases explicitly. First, it infers a soft front--back ordering from box overlap and relative bottom positions, and estimates directional occlusion coverage and depth. It then constructs a mixture of one clean and four directional occlusion observation components. The model uses a five-dimensional observation comprising box center, area, confidence, and aspect ratio, and adapts observation noise to predicted occlusion and detection confidence. The directional mixture likelihood is used in high-confidence association, low-confidence association, and track recovery; ambiguity penalties and local order-consistency swaps further reduce identity errors among nearby objects. On the DanceTrack validation split, \ours{} improves HOTA from 63.00 to 66.34, AssA from 45.10 to 49.57, and IDF1 from 62.19 to 65.28 over OA-SORT with the same detector and evaluation protocol. The gains are concentrated in association quality while detection accuracy remains stable. Additional local evaluations on MOT17 and MOT20 train splits characterize cross-dataset behavior under the same no-ReID tracking protocol.
Image-to-LiDAR registration estimates the camera pose of an image with respect to a LiDAR point cloud. It has diverse applications in autonomous driving, robot navigation etc. However, state-of-the-art (SOTA) methods still 1) mostly assume same-frame inputs, struggling with the image and point cloud from distant frames; 2) rely on domain-specific training, failing to generalize to unseen scenarios. We propose ZIL, the first foundation model for zero-shot non-synchronized image-to-LiDAR registration. ZIL encodes the input image and point cloud with the Vision and Point Transformers. In addition to regressing the relative pose, ZIL also learns to predict 3D coordinates, which substantially improves the pose accuracy without additional annotations. Interestingly, naive mix-data training cannot enable zero-shot generalization, which requires normalization on both camera intrinsics and the LiDAR vertical-axis origin. Trained on 7 public datasets with 1.4M LiDAR frames, ZIL consistently and significantly outperforms previous SOTA with a single model across 5 in-domain and zero-shot benchmarks, reducing the translation and rotation errors by up to 87% and 76% (shown in Fig. 1). Code and models are available at https://github.com/ZijunLi7/ZIL.
Manual attendance methods, such as paper or register-based systems, take a lot of time, can lead to errors, and are easy to falsify. Face recognition is more reliable, but it frequently struggles in classrooms because lighting and other conditions can vary. Face recognition datasets are designed for regulated environments and do not capture the actual challenges found in classrooms. To address this, a new face detection and recognition dataset, the Visage Face dataset, comprising 16,234 face samples, is proposed for the task of face detection and recognition. The photos are taken from different angles and under varying lighting conditions, with students showing a range of expressions, and some faces partly covered to reflect real-life situations. A YOLO-based system is used to detect faces and tested seven advanced face recognition models with thirteen configurations: LVFace, QCFace, FaceLiVTv2, TopoFR, EdgeFace, TransFace, and GhostFaceNets. Of these, FaceLiVTv2-M performed best, with 99.75% Top-1/Top-5 accuracy and an inference time of 6.459 ms. These results show that the Visage Face Dataset is a realistic and challenging benchmark for face recognition in classroom attendance.
Generative world-action models (WAMs) jointly generate future video and vehicle actions, while their action branches remain primarily optimized by expert imitation. Yet imitation provides no explicit closed-loop geometric verdict for generated trajectories, making verification important during both training and deployment. Closed-loop evaluators can check collision and drivable-area violations, but require privileged scene state unavailable at deployment. Existing approaches often close this gap by learning a verifier from sensor features. For these geometric checks, the rule itself is explicit. For example, collision is determined by whether the rolled-out ego footprint overlaps occupied vehicle space. What is unavailable at deployment is the scene state needed to apply the rule. We introduce DriveReferee, which uses a learned geometry readout to predict the scene representation from camera observations and executes the geometric safety rule directly rather than learning it. The resulting analytic referee evaluates collision and drivable-area safety from a scene state and candidate trajectory. During training, it scores self-sampled trajectories on ground-truth state and distills the resulting preferences into the WAM policy. At deployment, the same referee evaluates generated trajectories on this predicted state and selects a safer alternative when needed. The analytic referee requires no verdict-specific training, and its decisions follow an explicit geometric rule. Under matched candidates and inference budgets, it matches or outperforms all learned-verifier and heuristic baselines. Given the same predicted state and trajectory, learning the verdict provides no measurable downstream gain despite requiring tens of thousands of evaluator-labeled training examples. On the full NAVSIM navtest, DriveReferee reaches 92.02 PDMS with single-camera visual input and no external training data.
Video foundation models now reach human-level accuracy on physical-reasoning benchmarks, yet such tasks require predicting unobserved physical outcomes. Do these models perform human-like forward simulation, or do they exploit statistical regularities in visible scenes? Accuracy alone cannot distinguish these strategies. We introduce a distributional evaluation framework that treats model seeds and human raters as populations, enabling comparison of consensus, uncertainty, and strategy. On the Physion benchmark, we evaluate three ViT-L architectures (V-JEPA2, VideoMAEv2, DINOv2). V-JEPA2 narrows the accuracy gap to ~1 percentage point (73.2% vs. 74.2%), yet model-human disagreement reaches 26.4%, far exceeding human-human disagreement (4.8%), with substantially lower agreement (kappa ~ 0.48 vs. 0.91). The divergence follows forward-simulation demands: models outperform humans on geometric reasoning (linking, +11.8 pp) but underperform on gravitational dynamics (rolling, -11.8 pp) and causal chains (dominoes, -10.5 pp). Strategy fingerprinting confirms all three architectures share non-human strategies while none aligns with humans. Attribution analysis suggests that unobservable outcome features, rather than visible scene properties, predict this divergence, consistent with models relying more on scene-level statistical regularities than on explicit forward simulation, a systematic divergence that accuracy alone cannot reveal. Code is available at https://github.com/fanhong-li/model-human-divergence.
Chinese Translation
视频基础模型目前在物理推理基准测试上已达到人类水平的准确率,然而此类任务需要预测未被观察到的物理结果。这些模型是在进行类人的前向模拟,还是利用了可见场景中的统计规律?仅凭准确率无法区分这两种策略。我们提出了一种分布式评估框架,将模型种子和人类评分者视为总体,从而能够比较其共识性、不确定性和策略。在Physion基准上,我们评估了三种ViT-L架构(V-JEPA2、VideoMAEv2、DINOv2)。V-JEPA2将准确率差距缩小至约1个百分点(73.2% vs. 74.2%),但模型与人类的分歧高达26.4%,远超人类之间的分歧(4.8%),且一致性显著更低(kappa约为0.48 vs. 0.91)。这种分歧遵循前向模拟的需求:模型在几何推理(连接,+11.8个百分点)上优于人类,但在重力动力学(滚动,-11.8个百分点)和因果链(多米诺,-10.5个百分点)上表现不佳。策略指纹分析证实这三种架构均共享非人类策略,且无一与人类对齐。归因分析表明,预测这种分歧的是不可观察的结果特征,而非可见场景的属性,这与模型更多依赖场景级统计规律而非显式前向模拟的假设相一致——这是一种仅凭准确率无法揭示的系统性分歧。代码可在 https://github.com/fanhong-li/model-human-divergence 获取。
Image-to-layer decomposition converts a flattened image into editable RGBA layers, enabling element-level editing in design workflows. Existing diffusion-based systems typically adapt large pretrained text-to-image (T2I) models and introduce RGBA autoencoders or variable-layer architectural modules. We revisit this design choice and ask whether layer decomposition truly requires these heavyweight components. We introduce PixelART, a pixel-space rectified-flow Transformer trained from scratch for image-to-layer (I2L) decomposition. PixelART directly denoises regional RGBA pixel patches using a single-stream multi-modal diffusion Transformer, avoiding RGBA-VAEs, pretrained T2I backbones, and layer-specific decoders. We identify a key property of the task: high-noise timesteps determine layer assignment and coarse layer organization, while low-noise timesteps mainly refine color, alpha, texture, and boundaries. Based on this observation, we propose a terminal-boosted timestep sampling strategy to increase training coverage in the high-noise layer assignment regime. Trained on 4M multi-layer design templates, PixelART achieves state-of-the-art layer decomposition and composite reconstruction on Design-Multi-Layer-Bench and LICA with over 80% fewer parameters, 98% lower latency, and 85% lower memory than the recent Qwen-Image-Layered model. Ablation experiments show that pixel-space $\mathbf{x}$-prediction, terminal-boosted timestep sampling, and data/model scaling are critical, while T2I pretraining brings marginal benefits to the I2L task.
Implicit Neural Representation (INR) leverages neural networks to represent discrete signals such as images as continuous ones, where the network weights serve as a compact form of the signal itself. Most existing INR methods adopt Multi-Layer Perceptrons (MLPs) as their backbone. Since these models render each pixel independently, they inherently fail to exploit the spatial correlations that exist between neighboring pixels. In contrast,convolutional INRs can process pixels in parallel while inherently accounting for inter-pixel dependencies, making them a more natural fit for representing images. Nevertheless, convolutional INRs remain relatively underexplored, and the majority of them rely on fixed architectural settings, leaving little room for image-specific adaptation. In this paper, we investigate network customization for convolutional INRs. We replace conventional filters with irregular directional kernels, whose allocation is guided by the directional energy in the image spectrum, i.e., directions exhibiting stronger energy are assigned a larger number of kernels, enabling content-tailored convolution settings. These kernels are further reformulated via an orthogonal basis to achieve a superior sparse representation. Moreover, we introduce an annealed Gumbel-Softmax-based mechanism for kernel-level activation function selection, which gives the most suitable activation function for each convolution kernel. Extensive experiments demonstrate that our method, namely C$^{2}$-INR, achieves superior performance against state-of-the-art approaches under comparable parameter budgets across a wide range of image processing tasks, including representation, inpainting, and super-resolution.
SatOV: Restoring Spatial Priors for Training-Free Open-Vocabulary Segmentation in Remote Sensing Imagery
SatOV:面向免训练遥感影像开放词汇分割的空间先验恢复方法
Zhao, Changhao, Zeng, Linglin, Liu, Hai
Abstract
Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation to the spatial characteristics of remote sensing data. Although existing vision-language foundation models perform well in general domains, their image-level classification design weakens the spatial priors needed for high-resolution remote sensing segmentation: structural spatial relations are degraded during deep feature transformation, and fine-grained spatial details are lost during downsampling. To address these complementary deficiencies, we propose SatOV, a training-free framework for open-vocabulary remote sensing segmentation that restores spatial priors at two stages of the representation pipeline. Specifically, Residual QQ Attention (ResQQ) extracts Query-Key self-attention from an intermediate CLIP layer and fuses it with final-layer Query-Query attention via a residual combination, restoring structural spatial priors suppressed by the final-layer representation. Spatially Modulated Upsampling (SatUp) uses the original high-resolution RGB image as spatial guidance, combining spatial feature modulation with guided cross-attention to reconstruct pixel-level textures and boundaries. Extensive experiments on DOTA, UDD, LoveDA, and Vaihingen show that SatOV consistently improves training-free OVS and achieves competitive quantitative and qualitative results against state-of-the-art methods. These results validate the effectiveness of restoring spatial priors at both the representation and spatial-resolution stages for remote sensing open-vocabulary segmentation.
LINGO: Latent Initialization and Gradient Optimization for Sparse-view X-ray Novel View Synthesis and CT Reconstruction with 3D Gaussian Splatting
LINGO:基于3D高斯泼溅的稀疏视角X射线新视角合成与CT重建的隐式初始化与梯度优化方法
Xing, Lifeng, Jin, Dequan, Bu, Kunpeng, He, Peigeng, Ying, Shihui
Abstract
In novel view synthesis and Computed Tomography (CT) reconstruction with sparse-view X-ray imaging, insufficient angular coverage leads to structural ambiguity and accumulated noise. Integrating 3D Gaussian Splatting (3DGS) with X-ray absorption physics can achieve promising results, but it suffers from noisy initialization, positional insensitivity, and weak gradients in low-density regions. In this paper, we propose a unified Latent Initialization and Gradient Optimization (LINGO) framework to address these issues. LINGO combines latent mask-space initialization with dynamic gradient optimization to improve point cloud structural completeness while accelerating training. It constructs voxel-level 3D filters from X-ray masks to robustly suppress background noise and provide reliable geometric priors. By employing an adaptive voxel scaling strategy and dynamically scaling loss, LINGO can adjust spatial resolution and explicitly amplify gradients in low-density structures. To evaluate the quality of initialization, we introduce the Initialization Point Cloud Structural Deviation (IPSD) metric. Experiments on the X3D dataset indicate that for the novel view synthesis task, LINGO improves the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) by an average of 0.72 and 0.0039, respectively, over baselines under identical sparse-view settings, achieving comparable reconstruction quality within 5k steps to state-of-the-art models typically trained with 30k iterations. For the CT reconstruction task, LINGO also demonstrates consistent improvements, with average PSNR and SSIM gains of 0.36 and 0.0134. These results highlight LINGO's effectiveness in both accelerating training and enhancing reconstruction quality across different sparse-view imaging scenarios.
Chinese Translation
在稀疏视角X射线成像的新视角合成与计算机断层扫描(CT)重建中,角度覆盖不足会导致结构模糊和噪声累积。将3D高斯泼溅(3D Gaussian Splatting, 3DGS)与X射线吸收物理相结合可以取得较好的效果,但其存在初始化含噪、对位置变化不敏感以及低密度区域梯度微弱等问题。本文提出一个统一的隐式初始化与梯度优化(Latent Initialization and Gradient Optimization, LINGO)框架来解决这些问题。LINGO将隐式掩码空间初始化与动态梯度优化相结合,在提升点云结构完整性的同时加速训练。该方法从X射线掩码构建体素级3D滤波器,以稳健地抑制背景噪声并提供可靠的几何先验。通过采用自适应体素缩放策略和动态缩放损失,LINGO能够调整空间分辨率并显式放大低密度结构中的梯度。为评估初始化质量,我们引入了初始化点云结构偏差(Initialization Point Cloud Structural Deviation, IPSD)指标。在X3D数据集上的实验表明,对于新视角合成任务,在相同的稀疏视角设置下,LINGO相较基线方法的峰值信噪比(PSNR)和结构相似性指数(SSIM)平均分别提升0.72和0.0039,并且在仅5k步内即可达到通常需要30k次迭代训练的最先进模型相当的重建质量。对于CT重建任务,LINGO同样展现出一致的改进,PSNR和SSIM平均分别提升0.36和0.0134。这些结果凸显了LINGO在不同稀疏视角成像场景中在加速训练和提升重建质量两方面的有效性。
Dynamic object segmentation and ego-motion estimation are closely coupled problems in autonomous driving, as accurate ego-motion estimation typically requires static scene observations, while identifying static observations requires knowledge of the ego motion. We present Radar-Dot, a radar--RGB framework that exploits radar Doppler measurements to address this coupling. Radar returns are first used to estimate ego velocity through a linear Doppler constraint, with residual-based static/dynamic segmentation and robust estimation used to reduce the influence of moving objects. The estimated motion is then combined with metric depth and dense optical flow to identify image regions whose observed motion is inconsistent with the rigid scene motion. Experiments on 10 nuScenes scenes (part of nuscenes-mini) demonstrate that the resulting geometric pipeline achieves 20.24% dynamic IoU and 33.67% F1-score over 394 frame pairs, while radar-based static-point filtering improves ego-velocity estimation compared with using all radar returns. These results demonstrate the potential of radar as a modality for jointly improving ego-motion estimation and dynamic object segmentation.
Planning-Aligned Pretraining of BEV Representations with Sparse Action-Conditioned Targets for End-to-End Autonomous Driving
面向端到端自动驾驶的稀疏动作条件目标规划对齐BEV表征预训练
Song, Jaeha, Hwang, Soonmin
Abstract
End-to-end driving requires planning-relevant bird's-eye-view (BEV) representations, but existing pretraining approaches often rely on task annotations or dense scene reconstruction. We introduce PAVER, Planning-Aligned BEV Encoder Pretraining. From a single LiDAR sweep, PAVER constructs sparse risk and unknown targets describing occupied and unobserved evidence along rule-based ego motions. A 10K-parameter head predicts these targets from masked BEV features conditioned on the action state, directing supervision toward geometric constraints on candidate motions. Pretraining requires no driving-task annotations or dense reconstruction. Only the BEV encoder is transferred, preserving the downstream architecture and camera-only inference. On nuScenes, PAVER reduces VAD-Tiny's average collision rate from 0.51% to 0.19%, while improving planning L2, motion prediction, detection, and mapping. The selected VAD-Tiny and VAD-Base schedules use about 36% less estimated total training time than scratch training, including pretraining. On Bench2Drive Town05 Long, PAVER improves UniAD-Tiny's closed-loop Driving Score from 48.45 to 58.79. The project page is available at https://archiiive99.github.io/PAVER.
Combining Foundation Model Confidence and Monocular Depth for Training-Free Out-of-Distribution Segmentation
结合基础模型置信度与单目深度的免训练分布外分割方法
Varghese, Serin, Hüger, Fabian, Maag, Kira
Abstract
Autonomous vehicles operating in open-world scenarios are inevitably confronted with previously unknown objects, such as exotic animals or loose cargo. The reliable detection and segmentation of these out-of-distribution (OOD) objects is therefore crucial for a safe understanding of the environment and decision-making. Most existing approaches require access to OOD training samples, retraining of the segmentation backbone, or dedicated auxiliary architectures, limiting their practical applicability. We propose a training-free method that derives dense OOD scores directly from the confidence predictions of a foundation segmentation model, without any task-specific fine-tuning or access to anomalous data. To improve the robustness of our OOD segmentation, geometric information from monocular depth estimation is incorporated into the decision process, providing complementary cues to uncertainty-based predictions. We evaluate the proposed method on the SegmentMeIfYouCan benchmark and additionally assess its performance on OOD tracking in video sequences, reflecting the temporal nature of real-world perception systems. The method performs strongly on road-centered benchmarks.
Rastikerdar, Mohammad Mehdi, Guan, Hui, Ganesan, Deepak
Abstract
Large vision-language models (VLMs) enable recognition beyond a fixed class set, but their computational demands prevent them from running on many edge devices. Cloud offload makes this capability accessible, but sending every image consumes scarce bandwidth and communication energy. We ask how to bring the open-world recognition capability of VLMs to the edge while operating within tight compute, energy, and bandwidth budgets. Wildlife monitoring provides a natural setting for exploring this question because camera traps encounter species not known at deployment. We present Scout, an autonomous open-world recognition system that invokes a cloud VLM intermittently to teach new classes to a compact edge model. Given only the deployment location and empty site frames, Scout autonomously turns each species identified by the VLM into persistent, site-conditioned recognition capability in a resource-efficient edge model, without a predefined species list, human labeling, or manual tuning. Across 30 camera-trap deployments in three regions on an NVIDIA Jetson Orin Nano, the accuracy of Scout remains within 0.1-2.5% of a model given a predefined species list. On species outside its initial class set, Scout achieves 53.7-59.1% accuracy, compared with 56.5-65.1% for full cloud offload, while using 59-71% less deployment energy.
Chinese Translation
大型视觉语言模型(VLM)能够实现对固定类别集合之外的识别,但其计算需求使其无法在许多边缘设备上运行。云端卸载(cloud offload)使这一能力变得可行,但上传每张图像会消耗稀缺的带宽和通信能量。我们探讨如何在严格的计算、能量和带宽预算下,将VLM的开放世界识别能力带到边缘设备。野生动物监测为探索这一问题提供了天然的实验场景,因为红外触发相机(camera traps)在部署时会遇到未知的物种。我们提出了Scout,一个自主的开放世界识别系统,它间歇性地调用云端VLM来教会一个轻量级边缘模型识别新类别。仅需部署位置信息和无目标的空场景帧,Scout即可自主地将VLM识别出的每个物种转化为持久化的、针对特定地点的识别能力,并将其嵌入资源高效的边缘模型中,整个过程无需预定义物种列表、人工标注或手动调参。在NVIDIA Jetson Orin Nano平台上跨越三个地区的30次相机陷阱部署中,Scout的准确率与给定预定义物种列表的模型相比仅相差0.1-2.5%。对于其初始类别集合之外的物种,Scout达到53.7-59.1%的准确率(完全云端卸载为56.5-65.1%),同时部署能耗降低59-71%。
Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live system must synchronize the model's output with speech from an external voice agent, schedule video frames for playback, and handle interruptions. We introduce AVTR-1, an open stack for real-time interactive avatar conversations, built around a compact 153M-parameter autoregressive flow-matching motion generator conditioned on both participants' audio. We adapt its audio encoder for streaming through self-distillation. The stack turns the model's chunk-based generation into a continuous, synchronized audio-video stream driven by an external voice agent, and we analytically derive its contribution to the user-facing latencies and validate the resulting bounds with two commercial voice agents. Further experiments demonstrate that AVTR-1 leads the compared dyadic systems on all reported visual-quality metrics and most conventional listening-motion metrics while remaining competitive in lip synchronization. Its inference runtime operates in real time on data-center and consumer GPUs. However, conventional listening metrics do not establish whether the paired speaker's speech contributes to generated motion. We therefore introduce the Reference-Based Directed Granger Gain (R-DGG), which measures the additional predictive information carried by speaker speech after accounting for listener history and speaker motion. R-DGG finds statistically supported predictive dependence for recorded listeners and all evaluated dyadic systems, but not for talking-head generators without paired audio or mismatched speaker-listener pairs. We release the model weights, renderer, and serving backend under component-specific licenses.
Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation
规划与渲染协同:基于自回归布局与扩散模型的深度融合视觉文本生成
Chen, Guanqiao, Tan, Jingru, Mao, Dongxing, Chen, Catherine, Du, Zijian, Qin, Libo, Guo, Hu Jian, Wang, Alex Jinpeng
Abstract
Generating text-rich images from prompts requires both textual fidelity and the coherent integration of text into the surrounding image. An explicit layout can provide structured guidance about what text should appear and where, but a well-formed plan alone does not guarantee that the renderer will realize it faithfully. Existing layout-based AR-diffusion systems typically optimize planning and rendering separately, preventing the planner's representations from being adapted jointly with image synthesis. We introduce DuetGen, an autonomous visual text generator built on DeepFusion, which jointly learns autoregressive planning and continuous diffusion rendering. DeepFusion conditions a diffusion transformer on the planner's prompt and bbox-content hidden states, allowing rendering supervision to shape the representations connecting textual plans with visual outputs. Its joint objective combines autoregressive plan supervision, text-region-weighted diffusion learning, and auxiliary coordinate supervision to maintain structured planning, emphasize text-bearing regions, and improve the spatial precision of planner representations. During inference, Phase-Aware Attention Modulation strengthens the correspondence between image regions and their matched coordinate and content states, facilitating region-specific execution of the generated plan. With a 2B planner and a 4B single-stream DiT, DuetGen achieves 0.8293 word accuracy on CVTG-2K and 0.938 accuracy on LongText-Bench, closely matching the substantially larger Qwen-Image on both benchmarks. These results demonstrate the value of jointly learned planning representations and region-specific rendering for autonomous visual text generation.
Novel view synthesis from sparse inputs remains challenging for 3D Gaussian Splatting (3DGS) due to ambiguous geometry, cross-view inconsistency, and missing details in under-constrained regions, resulting in degraded reconstruction and unstable rendering. To tackle these issues, we propose D$^{3}$GS, a Depth-DINO-Diffusion guided sparse-view Gaussian reconstruction framework that jointly enhances geometry and appearance. D$^{3}$GS first recovers a high-resolution, metric depth map via diffusion-based completion and DPT (Dense Prediction Transformer) refinement, providing robust Gaussian initialization and geometric constraints. Then, a DINO-guided view-consistent learning is introduced to augment Gaussian attributes with structural features, improving multi-view consistency. Finally, a diffusion-based Gaussian refinement module injects generative priors into an iterative optimization strategy, enhancing high-frequency geometric and appearance details within the Gaussian representation. Experiments on DTU, LLFF, and Mip-NeRF 360 show that D$^{3}$GS achieves consistent and substantial improvements over strong baselines, with ablation studies validating the effectiveness and complementary roles of each component.
Generative models are rapidly expanding image quality assessment (IQA) beyond traditional fidelity factors to emerging dimensions such as physical plausibility and text-rendering correctness. However, existing IQA models rely on fixed definitions and heavy supervision, making them difficult to extend to open-ended perceptual dimensions. We identify holistic bias as an important limitation: when scoring an unseen dimension, models reuse generic quality priors, leading to scoring errors and rank inversion. To address this, we propose PACE (Perceptual Agentic Collaborative Evolution), a training-free multi-agent framework that formulates open-ended IQA as explicit protocol construction. Given a target dimension, PACE uses collaborative agents to construct an evaluation protocol composed of verifiable Visual Question Answering (VQA) probes, grounding evaluation in concrete visual evidence rather than holistic impressions. The resulting protocol is calibrated using only four human-annotated images per dimension, while a dual-track scoring mechanism aligns model perception with human scoring scales. Across traditional IQA, structural fidelity, context-aware aesthetics, and newly defined open-ended dimensions, PACE consistently improves its MLLM backbone, achieving competitive performance across diverse IQA settings, and reduces the Holistic Override Rate (HOR) from 44.4\% to 8.6\%.
Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing video reward models often produce unstable scalar scores because they directly map complex, subjective video quality into a single score without explicit evaluation criteria. This leads to scalar drift, where the scoring scale collapses or shifts across different prompts, making the reward unreliable for RL. Drawing inspiration from professional human annotation engineering, we address this problem with RewardVerse, a rubric-based video reward framework that introduces a dynamic rubric as an intermediate representation between the evaluation query and the scorer. Instead of unconstrained direct scoring, RewardVerse first generates explicit evaluation criteria and then performs rubric-guided scoring, providing a stable semantic anchor that mitigates scalar drift. To efficiently optimize this collaborative pipeline, we propose Rubric-Guided Policy Optimization (RGPO), a two-stage training algorithm. RGPO first warms up the scorer using self-evolving seed rubrics and then jointly optimizes the rubric generator to produce query-adaptive evaluation criteria while continuously aligning the scorer with human ratings. Extensive experiments on the 16-dimensional EvalVerse benchmark and external datasets demonstrate that RewardVerse mitigates scalar drift, achieves state-of-the-art performance on both pointwise and pairwise evaluation, and provides a robust and interpretable reward signal for RL in video generation.
While compressed-sensing regularizers enable interpretable reconstruction of 4D Flow CMR through transparent variational objectives, their hand-crafted nature is too restrictive under high acceleration. State-of-the-art learning-based approaches mitigate this, but typically encode regularization implicitly through unrolled network modules, which limits their interpretability. To address this limitation, we propose CLEAR, designed to combine the interpretability of compressed sensing with the flexibility of learned models. To the best of our knowledge, it is the first learned regularizer for a 4D reconstruction task. In the ultra-accelerated \(10\times\)--\(50\times\) regime of the CMRx4DFlow2026 challenge, CLEAR outperforms compressed sensing locally low-rank (LLR) and the popular variational network FlowVN, while using less than 10k parameters and preserving an interpretable regularization structure.
General Collaborative Intelligence: Architecting Cognition for Resilient Multi-Agent Ecosystems
通用协同智能:面向韧性多智能体生态系统的认知架构设计
Zhang, Lei, Ye, Chun, Yang, Le, Wang, Zhaozhong, Fan, Deng-Ping, Dai, Hang, Wang, Binglu
Abstract
Multi-agent unmanned systems are moving from isolated, ego-centric sensing toward collaborative intelligence, in which distributed agents exchange compact features to overcome a local observation trap that no single agent can escape: occlusions, finite sensor range, and environmental degradation. The field has matured across architectural, communication, embodied, resilience, and trust dimensions, yet existing surveys examine these dimensions in isolation and rarely expose their dependencies. This review offers a unified synthesis through two complementary lenses. The first is a five-dimensional taxonomy spanning collaboration stage, communication paradigm, fusion architecture, learning strategy, and application domain. The second is three cognitive synergy conditions, Semantic Disambiguation, Pragmatic Information Exchange, and Proactive Informational Foraging, that turn cognitive synergy into operational criteria. Across these lenses we survey collaboration architectures and topologies, neural-communication co-design that treats the channel as a differentiable pipeline component, embodied action-perception loops via multi-agent reinforcement learning, and resilience mechanisms for synchronization, uncertainty quantification, and label-efficient learning. We then map these advances onto four operational domains, V2X, unmanned aerial, industrial logistics, and smart cities, and onto the safety-privacy-utility triad. To counter benchmark saturation and evaluation fragmentation, we propose GCI-Bench, a five-pillar scoring protocol with a maturity model that makes the trade-offs of collaborative methods comparable across studies. A critical reflection on reproducibility, the sim-to-real gulf, and conditions under which collaboration degrades performance identifies open challenges and charts directions toward general collaborative intelligence under real-world uncertainty.
Chinese Translation
多智能体无人系统正从孤立的、以自我为中心的感知向协同智能演进。在协同智能中,分布式智能体通过交换紧凑特征来克服任何单一智能体都无法摆脱的局部观测困境:遮挡、有限的传感范围以及环境退化。该领域在架构、通信、具身、韧性和信任等维度上已趋于成熟,但现有综述往往孤立地考察这些维度,鲜少揭示它们之间的依赖关系。本综述通过两个互补的视角提供了统一的综合分析。第一个视角是涵盖协作阶段、通信范式、融合架构、学习策略和应用领域的五维分类体系。第二个视角是三项认知协同条件——语义消歧(Semantic Disambiguation)、语用信息交换(Pragmatic Information Exchange)和主动信息觅食(Proactive Informational Foraging)——它们将认知协同转化为可操作的评价标准。基于这些视角,我们调研了协作架构与拓扑结构、将信道视为可微流水线组件的神经-通信协同设计、通过多智能体强化学习实现的具身动作-感知回路,以及面向同步、不确定性量化和标签高效学习的韧性机制。随后,我们将这些进展映射到四个实际应用领域——车联网(V2X)、无人机、工业物流和智慧城市——以及安全-隐私-效用三元权衡之中。为应对基准饱和与评估碎片化问题,我们提出GCI-Bench,这是一个包含五大支柱的评分协议,并配有成熟度模型,使协作方法的权衡在不同研究之间具有可比性。最后,我们对可复现性、仿真到现实(sim-to-real)鸿沟以及协作导致性能退化的条件进行了批判性反思,指出了开放性挑战,并勾勒了在现实世界不确定性下迈向通用协同智能的方向。
We present M3GA-Wild, the first benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests. M3GA-Wild unifies and extends existing forest localisation datasets, providing a holistic benchmark with synchronised RGB imagery and LiDAR from ground traversals spanning 36 km, aligned high-resolution aerial imagery and multi-altitude LiDAR covering 370 hectares, and accurate geo-referenced 6-DoF poses for precise evaluation. M3GA-Wild captures diverse forest scenes with varying viewpoints, occlusion, and environmental conditions, enabling systematic evaluation of visual, LiDAR, cross-modal, and multi-modal methods. Baseline experiments show that LiDAR-based approaches significantly outperform vision-only methods under severe viewpoint differences, while current multi-modal fusion strategies yield limited gains due to poor cross-modal alignment. By pairing aerial RGB imagery with geo-referenced aerial LiDAR, M3GA-Wild also enables evaluation of foundation models for monocular depth estimation as a cheap source of 3D geometry from forest imagery, with initial experiments revealing shortfalls of current methods. These results highlight key challenges in cross-platform localisation, including modality misalignment and severe domain gaps. M3GA-Wild establishes a new benchmark to support research in robust multi-modal localisation and long-term autonomy in unstructured natural environments. The dataset and code will be available upon acceptance.
3D Gaussian Splatting (3DGS) enables high-quality novel view synthesis but incurs high storage and transmission costs due to dense Gaussian primitives. Recent anchor-based compression reduces per-primitive redundancy, yet redundancy across anchors remains largely unexploited. We propose CRP-GS (Cross-Representation Priors for Gaussian Splatting), a rate-distortion optimized compression framework that leverages cross-representation priors to improve anchor-level entropy modeling. First, a Correspondence-Oriented Hierarchical Structure (COHS) organizes anchors by feature correspondence rather than spatial proximity, constructing root-leaf dependencies so that selected anchors can act as informative priors to conditionally encode others, yielding more accurate likelihood prediction and lower conditional entropy. Second, Shared Feature Aggregation (SFA) extracts globally shared features from a contextual hash grid and injects them into anchor representations, factoring out scene-consistent low-frequency information that would otherwise be redundantly embedded in individual anchors. Both modules are trained under a unified rate-distortion objective to balance bitrate reduction and rendering fidelity. Experiments across multiple benchmarks show that CRP-GS achieves a favorable overall rate-distortion trade-off, yielding around 30% average bitrate reduction compared to anchor-based baselines while maintaining comparable rendering quality.
MixiMotion: One-Step Text-to-Motion Generation via Asymmetric Set Distillation
MixiMotion:基于非对称集合蒸馏的单步文本到动作生成
Dinh, Hung, Mai, Binh, Le, Tran Quoc Bao, Nguyen, Lam, Tran, Cong
Abstract
Iterative text-to-motion generation delivers high-quality and semantically aligned motions but requires multiple network evaluations, resulting in substantial inference latency. We present \textbf{MixiMotion}, a strict one-step text-to-motion generation framework based on offline set distillation. Instead of distilling a single teacher trajectory for each text prompt, MixiMotion constructs an offline bank of multiple teacher motions and aligns teacher and student sample sets through \textbf{asymmetric bidirectional matching}. The teacher-to-student direction promotes coverage of diverse teacher-supported motions, while the student-to-teacher direction suppresses unsupported generations. We further introduce differentiable decoded-space kinematic supervision to complement normalized representation matching with constraints in the decoded motion space. At inference, MixiMotion generates a complete motion sequence with a single network evaluation, without teacher queries, iterative sampling, or candidate ranking. On ViMoGen, MixiMotion achieves a semantic alignment score of $0.835$, outperforming the evaluated one-step baselines and approaching the $0.858$ score of its 50-step HY-Motion-1.0-Lite teacher. In blinded human evaluation, MixiMotion obtains an overall rating of $4.33$, compared with $4.50$ for the teacher, while outperforming the evaluated one-/few-step baselines. Meanwhile, generation latency is reduced from $829.58$\,ms to $9.30$\,ms, corresponding to an $89.2\times$ speedup. These results demonstrate an effective quality--efficiency trade-off for strict one-step text-to-motion generation.
CrowdCue: Specialist-Cue Conditioning for Vision-Language Crowd Counting
CrowdCue:面向视觉语言人群计数的专家线索条件化方法
Farazi, Moshiur, Ciftler, Bekir, Dandoush, Abdulhalim, Bendraou, Reda
Abstract
Generative vision-language models (VLMs) offer a counting paradigm in which one model produces both a count and a natural-language account of the scene, yet their raw counting accuracy sits in the range of sub-million-parameter specialist regressors. The open question is whether auxiliary guidance from a pretrained specialist can lift them into useful territory, and through which channel that guidance is best routed. We evaluate Qwen2.5-VL-7B on four widely used crowd counting benchmarks (ShanghaiTech A and B, UCF-QNRF, NWPU-Crowd). Zero-shot prompting rarely produces a parseable count, so LoRA supervised fine-tuning establishes the baseline at overall MAE 81.64. Conditioning on a P2PNet-derived density heatmap as an auxiliary visual signal fails in every encoding we tested, and an adversarial-swap protocol shows the model reads the heatmap but applies it counterproductively. We propose CrowdCue, a family that supplies the same specialist's already-integrated integer count to the VLM as a discrete symbol. The text-channel variant reaches MAE 72.04. The visual-channel variant, which renders the integer as printed digits and supplies it as a second image, reaches MAE 62.65, the strongest result in this paper and well ahead of the cue-supplying specialist alone (84.45 on the same split). In the late-fusion VLM we study, the binding constraint is not the channel but the abstraction level at which the specialist signal is delivered.
Chinese Translation
生成式视觉语言模型(VLM)提供了一种新的计数范式:单一模型既能输出人数,又能以自然语言描述场景,但其原始计数精度仍处于低于百万参数专用回归模型的水平。一个悬而未决的问题是:来自预训练专用模型的辅助引导能否将其提升到可用的精度范围,以及该引导通过哪种通道传递最为有效。我们在四个广泛使用的人群计数基准(ShanghaiTech A 和 B、UCF-QNRF、NWPU-Crowd)上评估了 Qwen2.5-VL-7B。零样本提示几乎无法产生可解析的计数结果,因此我们通过 LoRA 有监督微调建立了基线,总体 MAE 为 81.64。将 P2PNet 导出的密度热图作为辅助视觉信号进行条件化,在我们测试的所有编码方式中均告失败;对抗性交换协议表明模型虽然能读取热图,却以适得其反的方式加以应用。我们提出 CrowdCue 系列方法,将同一专用模型已整合的整数计数以离散符号的形式提供给 VLM。文本通道变体达到 MAE 72.04;视觉通道变体将整数渲染为印刷数字并作为第二张图像输入,达到 MAE 62.65,这是本文中最强的结果,且大幅领先于仅使用提供线索的专用模型(同一划分上为 84.45)。在我们研究的后期融合 VLM 中,关键制约因素并非通道本身,而是专用模型信号所传递的抽象层级。
Automated pollen analysis supports veterinary cytology, but brightfield microscopy is costlier and more complex than lens-less digital in-line holographic microscopy. We evaluate whether reconstructed holograms can narrow this gap and whether model explanations remain reliable under modality change. Six pollen species were imaged by brightfield and holographic microscopy. Raw, single back-propagation and iterative phase retrieval holograms were evaluated with YOLOv26s detection and MobileNetV4 classification after anchor-based annotation transfer. Six attribution methods were assessed for spatial grounding and faithfulness with the Attribution Health Inspection and Repair (AHIR) protocol, which tests model brittleness under weak noise and corrects attribution-map granularity when needed. Brightfield achieved 0.6890 mAP50-95 (0.8865 mAP50) for detection and 0.9687 macro-F1 (0.9705 accuracy) for classification. Reconstructed holograms narrowed the gap with a task-dependent split: p-type was strongest for detection at 0.5324 mAP50-95 (0.8229 mAP50), while r-type was strongest for classification at 0.7695 macro-F1 (0.7866 accuracy), both far above raw-hologram baselines. Activation-based explanations localized strongly on grains, and region-based methods retained ~60 to ~80% of faithfulness under holography. The holographic detector was highly brittle to weak perturbations, saturating deletion-based evaluation while insertion remained informative. Pixel-level gradient explanations approached random floor, yet spatial smoothing restored p-type gradient faithfulness from 0.05 to 0.51. For holographic classification, perturbation-based explanations remained faithful while gradient-based methods fell below random floor. Reconstruction improves low-cost holographic pollen analysis, while AHIR distinguishes genuine attribution failure from artifacts caused by model brittleness and map granularity.
Chinese Translation
自动化花粉分析可支持兽医细胞学,但明场显微镜相比无透镜数字同轴全息显微镜成本更高、系统更复杂。我们评估了重建全息图能否缩小这一差距,以及模型解释在模态变化下是否仍然可靠。使用明场和全息显微镜对六种花粉进行成像。在基于锚框的标注迁移之后,采用YOLOv26s检测和MobileNetV4分类对原始全息图、单次反向传播全息图和迭代相位恢复全息图进行评估。采用归因健康检查与修复(Attribution Health Inspection and Repair, AHIR)协议,从空间定位性和忠实性两方面评估了六种归因方法;该协议可测试模型在弱噪声下的脆弱性,并在必要时修正归因图的粒度。明场显微镜在检测上达到0.6890 mAP50-95(0.8865 mAP50),在分类上达到0.9687宏F1(0.9705准确率)。重建全息图以任务相关的方式缩小了差距:p型重建在检测上表现最佳,达到0.5324 mAP50-95(0.8229 mAP50);r型重建在分类上表现最佳,达到0.7695宏F1(0.7866准确率),二者均远高于原始全息图基线。基于激活的解释在花粉颗粒上表现出强空间定位能力,基于区域的方法在全息成像下保留了约60%至80%的忠实性。全息检测器对弱扰动高度脆弱,删除式评估趋于饱和,而插入式评估仍具信息量。像素级梯度解释接近随机下限,但空间平滑将p型梯度忠实性从0.05提升至0.51。对于全息分类,基于扰动的解释仍保持忠实性,而基于梯度的方法则低于随机下限。重建技术改善了低成本全息花粉分析,而AHIR协议能够区分真正的归因失败与由模型脆弱性和归因图粒度引起的伪影。
Brain lesion segmentation is a fundamental task in medical image analysis, playing a critical role in diagnosis, treatment planning, and longitudinal disease monitoring. Yet existing models still struggle to meet the demands of real clinical use, where deployments contain data distribution shifts, arising from differences in scanner hardware, imaging protocol (varying MRI modality sets), and new pathologies. We present BrainIAC (Brain lesion Interactive Adaptive Continuously learning segmentation), a unified framework that integrates (i) a multi-modal backbone network trained to segment multiple types of brain lesions and handle heterogeneous sets of modalities via zero-filling and random modality dropping; (ii) 3D interactive segmentation with bounding-box and click prompts that preserves fully automatic prediction when no prompt is given; and (iii) an online adaptation mechanism combining Mid-Interaction adaptation and Post-Interaction adaptation, supervised by the network's own predictions as pseudo labels and guided by an extra Click-Centered Gaussian loss. To our knowledge, this is one of the first 3D online adaptation methods for interactive segmentation, and the first to combine handling of heterogeneous modality sets with online adaptation. Experiments across seven brain MRI datasets demonstrate that the proposed components provide complementary and synergistic benefits. The method consistently outperforms existing approaches and generalizes well across heterogeneous imaging modalities, including those unseen during training, as well as previously unseen brain pathology types. The code and a 3D Slicer plug-in will be released at https://github.com/WenTXuL/BrainIAC upon publication.
Large-scale scene reconstruction is a critical foundational technology in robotic autonomous systems such as 3D mapping and autonomous driving. In recent years, 3D Gaussian Splatting (3DGS) has demonstrated remarkable advantages in both visual quality and computational efficiency, making it a promising representation for large-scale scene reconstruction. However, it still faces challenges in large-scale scenes, including excessive memory consumption and uneven viewpoint coverage caused by UAV acquisition, limiting its real-world applications. To address this, we propose VDGS, a novel 3DGS framework that incorporates camera distribution into scene modeling. VDGS introduces visibility-driven statistics for scene anchors to quantify supervision strength. These statistics are further leveraged for scene partitioning and for gradient compensation in under-optimized regions, thereby promoting balanced optimization across different regions. Extensive experiments on multiple large-scale aerial scene datasets demonstrate that, under imbalanced viewpoint distributions, VDGS consistently outperforms existing methods, while maintaining competitive performance in scenarios with more uniform view distributions.
HDMamba-YOLO: Efficient State-Space Perception and Local Spatial Reconstruction for UAV Small Object
HDMamba-YOLO:面向无人机小目标的高效状态空间感知与局部空间重建
Wei, Linduo, Fan, Junjie, Mai, Yijun, Qi, Yong
Abstract
Small-object detection in UAV imagery is challenged by weak visual evidence, ambiguous boundaries, dense object distributions, and complex backgrounds. Effective detection therefore requires long-range contextual information for target-background discrimination while preserving explicit local two-dimensional structures for accurate localization. These requirements arise at different stages of the detection pipeline and are not naturally addressed by a uniform feature-processing strategy. We propose Hybrid Dual-domain Mamba-YOLO (HDMamba-YOLO), a stage-wise heterogeneous SSM-CNN detector organized according to a perception-reconstruction-alignment-interaction rationale. EfficientVMamba-based EVSS establishes long-range contextual perception in the backbone, while PhasePatchMerging2D provides phase-aware hierarchical transitions. DST-Wrapper and Native C3k2-ASSAF then perform perception-to-reconstruction transition and repeated local two-dimensional reconstruction during FPN/PAN aggregation. DySample provides content-adaptive cross-scale resampling, while OS-CVTIA introduces macro-micro interaction and task-specific modulation for localization and classification. On VisDrone2019, HDMamba-YOLO-B achieves 42.737% mAP50 and 25.713% mAP50:95 with 10.042M parameters and 29.879 corrected GFLOPs. HDMamba-YOLO-Lite achieves 41.140% mAP50 and 24.741% mAP50:95 with 5.344M parameters. Under the unified AI-TOD evaluation protocol, HDMamba-YOLO-B obtains 21.621% AP and 47.881% AP50. Controlled ablations further support the stage-wise allocation of state-space perception, convolutional reconstruction, dynamic alignment, and task interaction for UAV small-object detection.
Referring surgical video instrument segmentation (RSVIS) aims at segmenting the instrument in a surgical video, given a textual description. Despite recent progress, current models are trained and assessed on relatively small-scale benchmarks, hindering the development of more general RSVIS. In addition, existing benchmarks support only the single-target expression that refers to one instrument in the video, while overlooking multi-target and no-target referring expressions, restricting the applicability of RSVIS in practical scenarios. Addressing these issues, we propose LD-RSVIS, a new benchmark aiming to facilitate more robust and general RSVIS. Specifically, LD-RSVIS consists of 3,536 surgical videos with 1.09 million frames and covers a broad set of 30 instrument classes from 25 various procedures. By including abundant videos and classes, LD-RSVIS could benefit large-scale training and evaluation of more general RSVIS methods. Besides, unlike existing datasets, LD-RSVIS offers diverse referring settings, including no-target, single-target, and multi-target expressions, which enables the development of more practical RSVIS models in real applications. In order to ensure high-quality annotations, all videos in LD-RSVIS are manually labeled with multiple rounds of inspection and refinement. To our knowledge, LD-RSVIS is the largest and most diverse benchmark for RSVIS. To analyze LD-RSVIS and to provide comparison for future research, we evaluate 12 representative methods, and the results reveal that more efforts are required for improvements. To encourage future research, we present a simple yet effective RSVIS method, dubbed Cascade-RSVIS, that first mines target-specific cues using the complementary multi-cue text information and then employs such cues and textual information for segmentation, achieving promising performance. Our benchmark and code will be released.
Chinese Translation
指代手术视频器械分割(Referring Surgical Video Instrument Segmentation, RSVIS)旨在根据给定的文本描述对手术视频中的器械进行分割。尽管近期取得了一定进展,但现有模型仍在相对小规模的基准数据集上进行训练和评估,这阻碍了更具通用性的RSVIS方法的发展。此外,现有基准仅支持指向视频中单个器械的单目标指代表达,而忽略了多目标和无目标指代表达,限制了RSVIS在实际场景中的应用。针对这些问题,我们提出了LD-RSVIS,一个旨在促进更鲁棒、更通用的RSVIS发展的新基准数据集。具体而言,LD-RSVIS包含3,536段手术视频,共109万帧,涵盖来自25种不同手术操作的30类器械。通过包含丰富的视频和器械类别,LD-RSVIS能够为更通用的RSVIS方法的大规模训练与评估提供支持。此外,与现有数据集不同,LD-RSVIS提供了多样化的指代设置,包括无目标、单目标和多目标表达,从而有助于开发在真实应用中更具实用性的RSVIS模型。为保证高质量的标注,LD-RSVIS中的所有视频均经过人工标注,并进行了多轮检查与修正。据我们所知,LD-RSVIS是迄今为止规模最大、多样性最高的RSVIS基准数据集。为了分析LD-RSVIS并为未来研究提供比较基准,我们评估了12种代表性方法,结果表明该方法仍需进一步改进。为促进未来研究,我们提出了一种简单而有效的RSVIS方法,称为Cascade-RSVIS,该方法首先利用互补的多线索文本信息挖掘目标特定线索,然后利用这些线索和文本信息进行分割,取得了良好的性能。我们的基准数据集和代码将会公开发布。
Multimodal Agentic Retrieval (MAR) requires agents to solve complex information-seeking tasks by iteratively invoking external tools. Typical frameworks such as ReAct maintain raw multimodal inputs and the accumulating interaction history in a single, ever-growing context, leading to the context explosion problem. While existing methods alleviate this issue by compressing redundant text, effective strategies for managing token-intensive visual content remain largely underexplored. To address this gap, we first conduct a systematic empirical study of approximately 10,000 trajectories. The results show that as visual cues are progressively extracted through external tools and textualized into the context, raw images become increasingly redundant. Continued image retention is associated with higher output entropy and can even degrade task accuracy. Motivated by these findings, we propose MM-ContextFold, a training-free framework that loads raw images only when needed. It maintains a persistent, text-only main context for high-level planning and spawns ephemeral branch contexts for image-dependent subtasks. Within each branch, the agent loads the relevant images, completes the subtask, and folds the result back into the main context as a concise textual summary; the images and branch trace are then discarded. Experiments on seven MAR benchmarks across five backbone models show that MM-ContextFold improves average accuracy by 6.3 percentage points over ReAct while reducing the working context length by 27.5\%.
Nguyen, Le Thien Phuc, Nguyen, Thien, Nguyen, Thanh-Huy, Hoang, Gia Minh, Vu, Anh Mai, Bagci, Ulas
Abstract
Most medical vision-language models (VLMs) excel at open-ended report generation and VQA but provide limited support for structured, fine-grained clinical perception within a unified interface. We present QwenVLConnector, a Qwen2.5-VL-based medical chatbot that unifies classification, multi-label classification, textualized detection, counting, regression, and free-form report generation under a single next-token objective. Our key component is a lightweight dense multi-layer Connector that aggregates low- and high-level visual features, aligns them through the pretrained vision Merger, and fuses them with the final visual representation without increasing sequence length. This design enriches visual tokens with complementary spatial and semantic cues while preserving efficiency. On FLARE-2D, QwenVLConnector improves detection F1 from 0.55 to 0.85, raises single-label classification from 0.37 to 0.51, and boosts report-generation GREEN by up to 18.3 points over the Qwen2.5-VL baseline. We further explore multimodal in-context learning for report generation, showing additional improvements without updating model parameters. Overall, QwenVLConnector offers a unified and efficient framework for combining structured medical perception with open-ended clinical text generation. Our code can be found at https://github.com/plnguyen2908/QwenConnector.
Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}: at extreme attention sparsity, step-local training losses keep decreasing while terminal generation quality stagnates or degrades. The trap is one of supervision: the dominant terminal errors originate in the high-noise structure-generation stage, and terminal-aligned training corrects terminal errors that substantially extended step-local training cannot. This yields a simple staging principle: \emph{first adapt the sparse architecture into a coarse prior, then correct the terminal distribution}. We instantiate the principle as \method, a unified acceleration framework for visual generation that combines a short sparse warm-up, few-step trajectory-mixed distillation, and FP8 quantization with fused kernels. \method sustains $97\%$ attention sparsity with strong visual quality on long-sequence 720P generation across Wan2.1/Wan2.2 backbones and T2V/I2V tasks, and $90\%$ sparsity on Wan2.1-T2V-1.3B-480P. With 3-step CFG-free inference, \method achieves a $265\times$ end-to-end speedup over the 50-step CFG dense baseline for Wan2.1-T2V-14B-720P on a single RTX~5090 ($220\times$ on H100), and denoises a Wan2.1-T2V-1.3B-480P video in $1.3$s.
A conventional camera uses millions of pixels to measure radiance from all directions within its field of view. We present an omnidirectional irradiance camera that measures the irradiance function---the illumination incident upon every point on a sphere. The irradiance function varies smoothly over the sphere and hence is bandlimited. We have analyzed this function in the frequency domain and have shown that it is well approximated by a weighted sum of the first seven degrees of spherical harmonics. As a result, the irradiance function can be accurately reconstructed from a small number of samples. This implies that an irradiance camera does not need millions of detectors (pixels)---just a handful of measurements suffice. This brings two major benefits. First, the camera consumes such little energy that it can be completely powered by the light falling on its detectors. Second, it does not capture the visual details needed to identify an individual, and hence privacy is preserved. We have built a prototype irradiance camera, called FluxCam, using 49 detectors arranged on the surface of a sphere. In a well-lit indoor environment, FluxCam can read out and wirelessly transmit its measurements at 30 frames per second using energy harvested from the light falling on it (i.e., without a battery, cable, or external power supply). We show how FluxCam can be used as an optical gyroscope for computing rotation, to monitor a workspace, as an untethered light probe for diffuse relighting, and as an omnidirectional pyranometer for estimating sky conditions and determining the best orientation of a solar panel.
High-quality texture generation is essential for creating realistic and production-ready 3D assets. Recent multi-view diffusion methods have shown promising results for image-guided 3D texturing, but they are typically constrained to low operating resolutions such as 512 or 768, making it difficult to preserve high-frequency details from high-resolution reference images. Scaling this paradigm to 2048 resolution is computationally prohibitive, as the unified multi-view sequence exceeds 212K tokens and incurs excessive memory and latency. In this paper, we present UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing. Our key observation is that object-centric multi-view renderings contain two major sources of redundancy: background-induced sequence redundancy and sparse token interactions within the foreground. To address them, we introduce Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence. To enable efficient foreground-only inference while avoiding reconstruction artifacts, we further design Foreground-Aware VAE Decoding to ensure the quality of the final high-resolution views. To satisfy the demanding data requirements of 2K-resolution multi-view diffusion training, we construct G-buffer TexVerse, a large-scale, ultra-high-resolution multi-view rendering dataset covering over 268,000 3D assets. Extensive experiments show that UltraTex generates visually faithful textures with rich fine-grained details, while substantially improving efficiency, achieving $20.6\times$--$91.1\times$ training speedup and $22.3\times$--$74.6\times$ end-to-end inference speedup over the baseline on common samples in our dataset. Code and data is at https://yiboz2001.github.io/UltraTex.
Feed-forward 3D Gaussian Splatting now reconstructs renderable scenes from unposed, uncalibrated images. Yet, most models supervise only photometric consistency and predict Gaussians pixel by pixel, which leaves global structure fragile and ties primitive count to image resolution and view count. To this end, GrapeSplat amalgamates multi-view cues into a voxel-aligned scene representation and decodes Gaussians directly from the learned grid, requiring no per-scene optimization or post-processing. An Atlas Encoder lifts all views into pixel-wise geometry-and-appearance features anchored at predicted 3D points. PEACH-Vox compands the unbounded scene into a bounded sparse grid through a smooth per-axis map with an exact closed-form inverse. The Sparse Decoder then consolidates the grid with sparse convolutions and decodes the full scene as multiple Gaussians per occupied cell. This amalgamated representation exploits sparse voxel occupancy, where the Gaussian count follows the occupied cells and saturates as views cover the scene, while grid resolution sets its ceiling. GrapeSplat turns unposed images into a renderable Gaussian scene in a single forward pass. Trained with 2D and 3D supervision on 8-view sequences, it generalizes zero-shot from 4 to 64 views across indoor and unbounded scenes. Code and trained weights are available at https://github.com/VAISR/GrapeSplat
Embodied world models learn to predict future physical dynamics from visual observations and control signals, where physical knowledge is implicitly entangled within latent representations. We introduce CausalWM, a 16B embodied world model that performs explicit causal chain-of-thought reasoning before future video prediction. CausalWM organizes useful variables into a reasoning trajectory, allowing the model to progressively capture causal dependencies underlying physical evolution. To train CausalWM, we collect 31K hours embodied data and develop a three-stage paradigm consisting of large-scale video pre-training, causal CoT mid-training, and multi-objective RL post-training. Despite using only a limited set of supervised CoT variables, CausalWM exhibits emergent in-context learning capabilities, enabling contextual visual feature guidance and efficient few-step generation. CausalWM achieves state-of-the-art performance across language-conditioned, action-conditioned, single-view and multi-view benchmarks, including Top-1 performance on TriWorldBench leaderboard.
SPACE: Semantic Projection and Alignment of CLIP Embeddings for Domain Adaptation
SPACE:面向领域自适应的CLIP嵌入语义投影与对齐方法
Manesco, João Renato Ribeiro, Jodas, Danilo Samuel, Rodrigues, Douglas, Passos, Leandro Aparecido, Papa, João Paulo
Abstract
A fundamental challenge in deploying vision models is domain shift, which arises when training and test data follow different distributions, leading to degraded performance. This challenge is amplified when the same semantic concept appears under distinct visual forms, such as photographs and sketches, where visual similarity is weak despite semantic correspondence. Existing unsupervised domain-adaptation methods aim to align distributions across domains but often ignore semantic relationships among samples of the same class. To address this issue, this paper introduces SPACE, a method that exploits the semantic structure of CLIP's vision-language space for domain adaptation. The key idea is to use text descriptions as semantic anchors by applying Singular Value Decomposition to CLIP embeddings of class descriptions, yielding an orthogonal basis that captures semantic relationships among categories. Visual features from both domains are projected into this semantic subspace, aligning images based on meaning rather than appearance.
Chinese Translation
视觉模型部署中的一个根本性挑战是领域偏移(domain shift),即当训练数据与测试数据遵循不同分布时,模型性能会随之下降。当同一语义概念以不同的视觉形式(如照片与素描)出现时,这一挑战尤为突出,因为此时视觉相似性较弱,而语义对应关系依然存在。现有的无监督领域自适应方法旨在对齐不同领域之间的分布,但往往忽略同一类别样本之间的语义关系。为解决这一问题,本文提出SPACE方法,该方法利用CLIP视觉-语言空间的语义结构来实现领域自适应。其核心思想是将文本描述作为语义锚点:通过对类别描述的CLIP嵌入应用奇异值分解(Singular Value Decomposition),得到一个能够刻画类别之间语义关系的正交基。随后,将来自两个领域的视觉特征投影到该语义子空间中,使图像基于语义含义而非外观进行对齐。
Exact Quotients of Fresnel-Kummer Surfaces and Certified Biaxial Refraction
菲涅耳-库默尔曲面的精确商与经验证的双轴折射
Shaska, Tanush
Abstract
The Fresnel wave surface governs the propagation of light in a transparent biaxial crystal. It is a special Kummer quartic, and we identify it exactly. Over the complex numbers the wave surface of a crystal is the Kummer surface of the Jacobian of an explicit genus-two curve branched at the signed square roots of the three principal permittivities. This Jacobian is isogenous, by an isogeny with kernel of order four, to a product of two elliptic curves. One elliptic curve carries the three permittivities, and the other carries the optic-axis angle. The physical family is Zariski dense in the locus of genus-two curves with an extra involution, and its automorphism strata are explicit. The identification instantiates a task-aware quotient, which identifies parameters that differ by a nuisance transformation and carries invariant coordinates and explicit strata. For biaxial crystals, two ratios of the permittivities form a complete invariant of the wave surface up to rotation and rescaling, and the four real nodes are given in closed form. At an interface the candidate transmitted waves are the roots of a quartic of exact degree four. Its real-root count, root order, and repeated-root events are decided by exact algebraic predicates, and along the generic single-node encounters of the paper its discriminant vanishes to second order. Floating-point solvers drop forward transmitted modes near the optic axes, and the certified solver does not. On exact equivalence classes, learned models on quotient coordinates are invariant and more accurate than models on raw tensors, while learned root-count predicates fail near the optic axes.
Blind Deconvolution of Binary and Pattern Images with Pixel Intensity Constraints and Sparse Gradient Prior
基于像素强度约束与稀疏梯度先验的二值图像与模式图像盲反卷积
Zhang, Qinghua, Yang, Xuesong, He, Liangtian, Deng, Liang-jian, Liu, Jun
Abstract
Blind image deconvolution (BID) is a prominent research topic in the field of imaging sciences, given its significant practical applications. Most existing model-based BID methods focus on natural images, incorporating appropriate prior knowledge about both the underlying image and the blur kernel. However, for certain classes of images, such as barcodes, text, and patterns, pixels can only take very limited values, a specific prior that is often overlooked in the literature. In this article, we introduce a novel pixel intensity constraint to leverage this important information, improving recovery performance for these specialized image classes. Specifically, we propose a unified framework for blind binary and pattern image deconvolution that incorporates both the pixel intensity constraint and a gradient sparsity regularizer. Numerical experiments demonstrate that our method outperforms many existing BID techniques, achieving superior results in terms of both visual quality and quantitative metrics.
3D reconstruction from a lengthy video stream input poses a dilemma for feed-forward reconstruction models (FFRMs), that a whole-stream inference context cannot be retained under limited GPU memory.Recent studies seek to resolve this problem via a trade-off between the integrity of inference context and GPU memory usage, which either suffer from a rapid memory inflation or degraded context integrity due to artificially capping memory usage.Driven by our key observation that the initial saliency of a token reliably dictates its long-term importance across the stream, we propose RegVGGT, a training-free token regulation method which aggressively regulates the tokens of incoming frames.By admitting at most 1% of tokens per frame to update the context memory, our method dramatically suppresses memory inflation as the stream progresses.Equipped with a FlashAttention-compatible token saliency estimation scheme, RegVGGT is capable of processing thousands of frames on a consumer-grade GPU with negligible compromise to reconstruction quality.Extensive experiments demonstrate that RegVGGT achieves state-of-the-art performance on long-horizon benchmarks across diverse FFRM prediction tasks, surpassing prior FFRM-based stream reconstruction baselines by a large margin.
Localizing and describing fine-grained differences between near-identical images is a critical yet underexplored capability for multimodal large language models (MLLMs). Existing benchmarks largely assess semantic comparison or single-image grounding in isolation, without jointly requiring faithful description and physical localization. To bridge this gap, we introduce MinCU, a benchmark for grounded minimal-change understanding, where each sample consists of an image pair differing by a single atomic variation in object category, attribute, count, or spatial position, and models are evaluated on their ability to describe the change, localize the changed regions, and identify the changed entity. We further propose Semantic-Guided Implicit Spatial Anchors (SG-ISA), a structured autoregressive method that decomposes prediction into a Think-Locate-Describe sequence. SG-ISA first predicts a semantic cue for the changed concept, then uses discrete spatial anchors as an implicit localization scaffold, and finally generates the change description together with the grounding box. Experiments reveal that even the strongest closed-source MLLMs and recent R1-style reasoning models struggle on MinCU, with most failing to jointly produce accurate descriptions and grounding boxes. Compared to the previous chain-of-thought method, fine-tuning with SG-ISA yields substantial joint improvements in grounding accuracy and description quality while reducing reasoning-token overhead by approximately 26%. These results suggest that an implicit intermediate spatial interface can be more effective than relying solely on model scale for grounded dual-image understanding.
As generative AI becomes increasingly used in anime-style image creation, distinguishing human-drawn, AI-inpainted, and text-to-image images is important for copyright attribution, visual provenance, and content governance. Existing AI-generated image detectors mainly target real-world photographs and often overlook anime-specific cues such as flat coloring, exaggerated structures, and artistic line control. To address this gap, we propose AniPrO, a multi-dimensional description-enhanced framework for interpretable anime image provenance. Built upon AnimeDL-2M, AniPrO contains 15,000 balanced samples from a 35,000-image candidate pool, covering Real, Inpainting, and Text2Image categories with structured five-dimensional descriptions. We further introduce AniPrO-SFD-Bench and AniPrO-MFR-Bench to evaluate provenance detection from statistical feature discrimination and multimodal fusion reasoning perspectives. Experiments show that structured semantic guidance reveals systematic AI-generation biases, such as the gap between global visual plausibility and local detail coherence, and improves the detection of challenging inpainting samples. The dataset and code will be released at: https://github.com/YAN-LIU05/AniPrO.
Bimanual interaction produces complementary tactile views of the same physical process, yet existing tactile representation learning largely models the two hands independently or combines them only for downstream prediction, leaving their cross-hand relationship unexplored. To exploit this overlooked structure, we introduce BiView-Touch, a tactile-only framework that completes masked target-hand latents from the remaining visible target-hand regions and the synchronized full contralateral hand. A student encoder with a geometry-conditioned directional decoder predicts full-view EMA latent targets, while temporal and layout counterfactuals encourage sensitivity to synchronized and anatomically organized source information. Controlled ablations and source-context interventions show that BiView-Touch learns structured cross-hand dependence on temporally aligned and anatomically organized contralateral tactile context, rather than benefiting from bilateral input alone. On the public HumanTouch dataset, its frozen representations consistently outperform representative self-supervised baselines across low-label settings. With only 5\% downstream labels, BiView-Touch achieves relative balanced-accuracy gains of 7.1\% on bilateral wrist-motion recognition and 14.1\% on force-derived interaction-phase recognition. We further introduce BVT-20, a 20-task bilateral tactile dataset, and demonstrate transfer across recording sessions and pretraining corpora, including transfer to a held-out bimanual task. Our code and dataset details are available on the anonymous project page: https://anonymous.4open.science/w/biview-touch-review-site-050C/.
Generation-free world action models (WAMs) retain future-video prediction during training but act from internal video features at inference, leaving unclear what these features should preserve for control. Our representation diagnostics show that representations with more predictable future changes need not make linear action decoding easier. Observed future changes provide additional action information beyond the present, and linearly readable action information is spatially concentrated. These findings motivate Action-Relevant Predictive States (ARPS), a compact predictive interface between the video and action experts. ARPS uses a horizon-conditioned state predictor to aggregate intermediate video features into a compact state that supplies all visual context to the action expert. Future-representation supervision trains different parts of this state to predict visual representations at different future times, together with their changes relative to the present. At inference, the supervision branch is removed, and the action expert only uses the learned predictive state computed from current observations. Controlled ablations show that future supervision substantially improves generalization under distribution shift. ARPS achieves 99.2% success on LIBERO and transfers to LIBERO-Plus without adaptation, reaching 87.3% and exceeding Fast-WAM by 39.2 percentage points.
Accurate multi-user localization is challenging in complex urban environments, where wireless measurements can become ambiguous under noise, blockage, and multipath, while visual observations provide complementary spatial context. This paper presents a vision-wireless fusion framework for multi-user localization using pilot-indexed channel state information (CSI). Orthogonal pilot indices preserve the identities of the communicating UEs in the CSI token sequence and localization outputs. The model encodes each pilot-indexed CSI observation as a query token and uses cross-attention to retrieve user-specific information from spatial visual memory. Self-attention among CSI tokens further captures inter-user interactions, while the resulting multimodal representations are used for user-wise localization. Experiments on different datasets show consistent improvements over model-based, CSI-only, and multimodal-fusion baselines. Further experiments evaluate the model under different wireless and visual conditions.
Gaussian Splatting has enabled real-time novel view synthesis, but its tightly coupled geometry and appearance representation often require a large number of primitives to reproduce high-frequency texture details, leading to substantial memory and optimization costs. Recent textured 2D Gaussian methods alleviate this limitation by attaching texture maps to Gaussian primitives. However, bridging the fundamental structural gap between discrete Gaussians and continuous 2D grids requires complex parameterizations that introduce severe computational overhead. This overhead fundamentally compromises the original efficiency of Gaussian Splatting, making the balance between detailed texturing and computational agility an unresolved challenge. To address these challenges, we propose LiteTex-GS, a fast and lightweight texturing framework for Gaussian Splatting. Our method initializes an extremely compact representation, assigning minimal local texture to each Gaussian and progressively allocates higher resolution only to primitives with significant reconstruction errors. To maintain a streamlined geometric scaffold, we introduce a contribution- and area-aware pruning strategy that eliminates low-utility Gaussians. Furthermore, to mitigate the gradient dilution caused by texture upsampling, we design a resolution-aware update rule that preserves rapid and stable convergence. Extensive experiments on standard novel view synthesis benchmarks demonstrate that our method achieves competitive or superior rendering quality while using substantially fewer parameters and less training time than existing textured Gaussian baselines.
ProxyBuild: Text-Guided Structured 3D Building Generation with Mesh-Anchored Procedural Proxies
ProxyBuild:基于网格锚定程序化代理的文本引导结构化三维建筑生成
Tang, Xiang, Li, Ruotong, Fan, Xiaopeng
Abstract
Text-guided 3D building generation holds tremendous application potential, yet existing generative models typically output inseparable single meshes or non-interactive rendered representations. While procedural modeling can generate editable buildings with hierarchical structures, rule authoring is laborious, and even with the aid of large language models (LLMs), it remains challenging to effectively solve procedural rules under geometric constraints. In this paper, we propose ProxyBuild, a novel hybrid framework for structured building generation. We introduce the Mesh-Anchored Procedural Proxy (MAPP) as a novel intermediate representation, which tightly anchors building components onto geometric shells, thereby decoupling the generation task into two phases: proxy prediction and proxy-to-asset instantiation. First, we construct a building dataset with MAPP annotations to train our designed face-edge bigraph encoder. By explicitly modeling the feature interactions of topological elements on heterogeneous mesh graphs, this encoder accurately infers the semantic roles of faces and edges. Subsequently, conditioned on textual styles and attribute parameters parsed by LLMs, we accomplish high-precision asset retrieval and assembly by integrating a spatial placement logic with hard constraints. Extensive experiments show that ProxyBuild not only significantly mitigates common issues in building generation such as over-smoothing, component collisions, and structural corruptions, but also accurately parses semantic-free shells from diverse sources. Outperforming prior baselines across various metrics, our method can robustly generate structurally clear, detail-rich, and post-editable 3D buildings from text, thereby providing a reliable and interactive content foundation for downstream applications such as virtual reality and digital twins.
Enhancing Shrimp Disease Detection via Deep Learning and Data Refinement for Resilient Aquaculture
通过深度学习与数据精炼提升虾病检测能力,助力韧性水产养殖
Truong, Vinh Canh-Thanh, Pham, Hai-Binh, Tran, Ngoc Hong
Abstract
Shrimp diseases continue to cause devastating losses in the aquaculture industry, driving a critical need for robust, automated detection. This work contributes the first application of Vision Transformers (ViT) and Self-Supervised Learning (SSL) to the shrimp farming domain, addressing both performance bottlenecks and data labeling challenges. We propose two deep learning pipelines to classify four key diseases: Healthy, Black Gill (BG), White Spot Syndrome Virus (WSSV), and a co-infection of both using a dataset of 4,348 images. First, our supervised transfer-learning approach leverages ImageNet-pretrained ViT-Small/16 and EfficientNet backbones. Second, we introduce a contrastive learning framework (SimCLR) with a ViT-Small encoder to extract robust representations from unlabeled images prior to fine-tuning. Our results establish strong new baselines for sustainable aquaculture monitoring. The supervised approach achieves an outstanding 96% accuracy with fast convergence, outperforming traditional generic models, while the label-efficient SSL approach reaches a highly competitive 85% validation accuracy.
Point cloud completion aims to infer a complete 3D shape from a partial point cloud and serves as a fundamental building block for downstream tasks such as reconstruction, editing, and simulation. Despite the recent progress, existing learning-based methods often implicitly rely on access to the ground-truth shape scale (GT-scale) during both training- and testing-time normalization, assuming privileged information that is unavailable in real-world inference. This hidden assumption limits practical deployment and can lead to severe completion artifacts, e.g., over- or under-completion and nested shells, once the oracle GT-scale cue is removed. We observe that the recent foundation image generation models exhibit a strong capability of understanding objects and geometries, and producing multi-view consistent renderings, making them promising priors for GT-scale-free 3D completion. Motivated by this insight, we propose ScaleBlind, a novel framework that leverages foundation-model-based image completion to recover global scale directly from partial inputs and then faithfully produces the 3D completion. Specifically, ScaleBlind dreams out complete multi-view appearances from rendered partial views, lifts the inferred missing regions back into 3D to obtain a geometry-aware coarse completion, and further refines it via a powerful cross-modal fusion network with the original partial point cloud. By harnessing 2D foundation priors, our method eliminates the need for accessing GT-scale information at inference. Moreover, it provides a principled bridge between 2D generative priors and 3D point cloud completion. Extensive experiments demonstrate the superiority of our framework, making ScaleBlind the new state-of-the-art for the point cloud completion task.
In frame interpolation tasks, motion ambiguity in the training set causes models to generate blurry intermediate frames. Moreover, the assumption of uniform motion between frames during inference further leads to inaccuracies in the generated intermediate frames. To tackle these challenges, we propose an Accurate motion estimation algorithm with B\'ezier Control point, ABC-Inter, for efficient frame Interpolation. Specifically, ABC-Inter designs an Accurate Flow estimation Module (AFM) by decoupling two-frame features and mapping to corresponding coordinates to better estimate the optical flow between the two frames. Furthermore, ABC-Inter eliminates motion ambiguity in the training set by introducing B\'ezier control points that are computed using the input frames and the intermediate ground-truth (gt) frames. This allows the model to estimate accurate optical flow between two frames during the training process, thereby solving the blurriness problem in the generated intermediate frames during inference. Benefiting from the more accurate flow estimation between two frames, we can introduce additional frames and directly use multiple flows to calculate B\'ezier control points for modeling non-uniform motion without retraining the model. Simultaneously, to realize the estimation of non-linear motion using only two frames, we also introduce a new B\'ezier control point estimation module which achieves better motion estimation between the two frames by performing fine-tuning on the model in the second stage. Experimental results demonstrate that our ABC-Inter achieves state-of-the-art performance on multiple benchmark datasets and exhibits excellent visual perception.
Tamim, Mahir Shahriar, Alim, Md. Samiul, Wasi, Azmine Toushik, Ridoy, Shahriyar Zaman, Nesa, Meharun, Yousuf, Mohammad Abu, Lamb, Alex, Moni, Mohammad Ali
Abstract
Cache-based test-time adaptation improves CLIP predictions by storing and retrieving examples from the target stream while keeping the model frozen. However, existing methods largely treat the feature space used for image-image retrieval as fixed, leaving open how much adaptation depends on the retrieval space itself. We study this question by fixing the memory and changing only the retrieval encoder, finding that the same memory can yield very different gains: across sixteen retrieval spaces, ImageNet-A cache gain ranges from at most +0.44 points for CLIP and MAE to +19.7 +/- 0.4 for DINOv2-L, while label-free retrieval-space selection retains 98% of oracle gain on ImageNet-V2. These results show that memory quality depends not only on which examples are stored, but also on how they are retrieved. Motivated by this finding, we propose MARC (Memory Augmented Retrieval for CLIP), a training-free system that uses frozen CLIP for prediction and DINOv2-B for retrieval with a single fusion weight. A single-view cache repairs 1074 +/- 21 baseline errors, compared with 878 +/- 4 for a 64-view ensemble, at roughly one seventh of the cost. Across four ImageNet distribution shifts, MARC reaches a 67.91% OOD average and, at matched DINOv2-B scale and eight views, achieves 64.17 +/- 0.31% versus 62.75 +/- 0.15% for a graph-based cache system while running 2.6 times faster. Overall, our results establish retrieval space as a first-order design choice for robust cache-based adaptation in remote sensing, scientific imaging, and changing visual environments.
Computational Fluid Dynamics (CFD) is widely used to evaluate ventilation and contaminant transport in occupied buildings, but deployment at scale is limited by three bottlenecks: acquiring room geometry without costly scanning hardware or manual CAD modeling, decomposing the scene into individually manipulable objects, and reconfiguring those objects for alternative layouts without re-capturing the room. We present a semi-automated workflow that converts a single 360-degree video of a room into individually editable, simulation-ready geometry assets. A dense point cloud is reconstructed using Neural Radiance Fields (NeRF), and 2D instance masks from text-prompted SAM 3 segmentation are lifted to 3D using multi-view consensus and depth-band filtering. Points are separated into object instances with an octree, and occlusion gaps are healed with a connectivity graph. Chair templates are fitted by Iterative Closest Point (ICP) alignment, and table geometry is generated procedurally. A browser-based editor supports quality assurance and rapid construction of alternative layout configurations. A steady Reynolds-averaged OpenFOAM solution then drives transient passive-scalar transport; the setup is verified using a mesh-sensitivity study and validated against an IEA Annex 20 benchmark. We apply the workflow to two university classrooms and a tiered lecture-hall auditorium. The capture-to-geometry pass takes two to five hours per room on a consumer workstation. In a controlled obstruction sequence in one classroom, the modeled half-clearance time varies non-monotonically as furniture is added, and a cross-room comparison indicates that clearance behavior cannot be reliably extrapolated between rooms, motivating per-room geometry acquisition. By making that acquisition low-cost, the workflow makes geometry-resolved comparative ventilation studies practical for spaces such as classrooms.
Vision foundation models targeting Earth observation (EO) tasks are commonly evaluated on clean downstream benchmarks, but operational EO products can already contain spatial, radiometric, alignment, noise, and harmonization defects before reaching the model. Existing robustness evaluations often use generic image corruptions or broad domain shifts, which do not isolate these product-level failure modes. We introduce \textbf{RSPDBench}, a physically grounded \textbf{r}emote-\textbf{s}ensing-\textbf{p}roduct \textbf{d}egradation \textbf{b}enchmark for vision foundation models. RSPDBench evaluates five EO datasets, seven foundation-model entries, and two supervised baselines under audited primitive degradations and compound product chains. Each model is evaluated under its clean-selected native protocol, with robustness measured as the drop from its own clean baseline. Our analysis reveals that degradation sensitivity is strongly structured: resolution-conditioned and channel-grouped encoders protect different failure axes, and the same physical defect can hurt one model while helping another. Compound chains expose failures that isolated degradations do not predict, with model-dependent amplification, saturation, or component dominance, and excess drops up to $38$ percentage points beyond the strongest component. These results show that EO robustness cannot be characterized by clean accuracy or generic perturbation tests alone; it must also be measured against the structured defects that remote-sensing products carry into deployment.
HOIBlender: Blending Lightweight Detection with Vision-Language Priors for Efficient Human-Object Interaction Detection
HOIBlender:融合轻量级检测与视觉-语言先验的高效人-物交互检测
Chen, Junwen, Yanai, Keiji
Abstract
Human-object interaction (HOI) detection requires grounding an interacting human-object pair and recognizing the verb that links them, often under severe long-tail supervision. Recent methods improve accuracy with stronger detectors and vision-language priors, but many still stack heavy transformer encoders, intricate denoising schedules, or post-hoc semantic calibration on top of the detector. We present \textbf{HOIBlender}, an efficient HOI detector named after its core design principle: blending detector-grounded visual tokens, spatial subject-object reasoning, and BLIP-2 semantic priors inside one lightweight decoding pipeline. HOIBlender builds on an RF-DETR/LW-DETR-style foundation with a DINOv2 backbone and selects top-$K$ image-conditioned tokens directly from the multi-scale projector as subject and object candidates, removing the dedicated encoder stage retained by prior HOI methods. A dual-stage decoder first stabilizes human-object geometry and then performs verb and HOI classification through progressive BLIP-2 prior fusion, with classifier weights initialized from BLIP-2 text embeddings for long-tail categories. Grouped-query training further enriches optimization without increasing inference cost. Across three model scales (Nano, Small, 2XL), HOIBlender consistently outperforms SOV-STG-VLA and Hybrid-SOV-VLA on HICO-DET, reaching $44.49$ Default Full mAP in only $9$ training epochs while maintaining competitive latency and parameter budgets. These results show that lightweight detection, structured spatial-semantic decoding, and deeply integrated vision-language priors can be blended into a single efficient HOI pipeline.
Chinese Translation
人-物交互(HOI)检测需要定位交互的人-物对并识别连接二者的动词,且通常面临严重的长尾监督问题。近期方法通过更强的检测器和视觉-语言先验提升了精度,但许多方法仍在检测器之上堆叠重型Transformer编码器、复杂的去噪调度或事后语义校准。我们提出了HOIBlender,一种高效的HOI检测器,其名称源于其核心设计原则:在一个轻量级解码流程中融合基于检测器的视觉词元、空间主客体推理以及BLIP-2语义先验。HOIBlender构建于RF-DETR/LW-DETR风格的基础架构之上,采用DINOv2骨干网络,直接从多尺度投影器中选取前K个图像条件化词元作为主语和宾语候选,从而去除了以往HOI方法保留的专用编码器阶段。双阶段解码器首先稳定人-物几何关系,随后通过渐进式BLIP-2先验融合完成动词与HOI分类,其中分类器权重由BLIP-2针对长尾类别的文本嵌入初始化。分组查询训练进一步丰富了优化过程而不增加推理成本。在三个模型规模(Nano、Small、2XL)上,HOIBlender在HICO-DET数据集上始终优于SOV-STG-VLA和Hybrid-SOV-VLA,仅用9个训练周期即达到44.49的Default Full mAP,同时保持了具有竞争力的延迟和参数预算。这些结果表明,轻量级检测、结构化的空间-语义解码以及深度集成的视觉-语言先验可以融合为单一高效的HOI处理流程。
Novel view synthesis from sparse observations is severely under-constrained. Although 3D Gaussian Splatting (3DGS) enables real-time rendering, it produces floaters, broken geometry, and washed-out backgrounds when trained with few views. We propose an alternating optimization framework that uses a pre-trained image diffusion model to generate geometrically consistent pseudo-views for additional 3DGS supervision. Generation is constrained by depth-conditioned ControlNet, IP-Adapter style transfer, LoRA scene adaptation, and img2img structural anchoring. We introduce Generative Active Pseudo-view Selection (GAPS) to balance reconstruction informativeness and generative reliability when choosing target views. Its annealing schedule shifts from conservative interpolation early in training to exploratory extrapolation later, gradually covering unobserved regions. A dual-criterion admission gate and uncertainty-weighted losses reject unreliable generations, while density-adaptive DropGaussian reduces overfitting in complex scenes. On LLFF with 3/6/9 views, our method improves average PSNR over vanilla 3DGS by 0.40/0.89/0.70 dB. On Mip-NeRF 360 with 12/24 views, the gains are 1.18/0.80 dB. SSIM improves and LPIPS decreases in every setting. Ablations show that active selection and density-adaptive regularization are both necessary; only the full method reduces LPIPS below the no-pseudo-view baseline on unbounded 360-degree scenes.
Diffusion models generate high-quality images, yet often violate the physical laws governing mirror reflections. Reflections often suffer from geometric aberrations, including positional offsets, directional misalignment, proportional imbalance, and structural distortion. These failures remain evident even in contemporary state-of-the-art generative systems. Existing methods itigate this problem through synthetic data scaling or auxiliary depth conditioning, yet their merely reliance on latent-space noise reconstruction losses as implicit supervision prevents direct enforcement of reflection-specific geometric and perceptual constraints. To bridge this gap, we present PhysReflect, a geometry and perception guided diffusion framework that decodes the predicted clean latent into pixel space at each training step and applies annealed supervision through two complementary differentiable objectives. The Geometric Loss enforces mirror-induced spatial consistency through sparse epipolar correspondence and dense boundary projection alignment, where a SAM2-based TwinTrack mechanism provides stable in-mirror localization for boundary-aware supervision. The Perceptual Loss preserves reflected appearance by combining Semantic Consistency Loss, which maintains reflected identity and appearance via DINOv2 features, and Lighting Consistency Loss, which regularizes depth, surface-normal, and illumination coherence under monocular geometry priors. Experiments on synthetic and real-world benchmarks show that PhysReflect outperforms prior mirror-reflection methods in geometric, perceptual, and physical-plausibility metrics, as well as qualitative visual results.
Rule-Constrained Assignment for Cue-Ball Identification in Broadcast Snooker
面向转播斯诺克比赛中母球识别的规则约束分配方法
Cao, Yuxin, Song, Wei, Wu, Yuezhong, Dong, Jin Song
Abstract
Accurate cue-ball identification is essential for metric analysis of broadcast snooker. Existing systems evaluate each candidate independently against a fixed white prototype and reject candidates above an appearance threshold. Under broadcast conditions, illumination changes can make colored balls appear white, while intrusions from players and equipment can obscure the cue ball or introduce competing candidates. We formulate cue-ball identification as a rule-constrained assignment problem that jointly assigns detected candidates to the bounded snooker inventory: one cue ball, up to 15 reds, and six colors with known spots. The cue ball is selected by the incremental cost of assigning each candidate to the white slot, and the conventional appearance test follows as the one-slot case. On 419 hand-annotated shots, our method improves identity accuracy from 88.1% to 95.5%, and from 80.5% to 95.2% on held-out venues. Within CueLift, our metric state-recovery system, the assignment expands coverage from 36.2% to 55.4% over 6,241 scorable shots. Assigning an estimate to every shot in a separate evaluation on 2,529 shots preserves this advantage.
Algebraic Consistency Alone Does Not Certify Temporal Structure in Latent Action Models
仅凭代数一致性无法认证潜在动作模型中的时间结构
Wen, Di, Zhang, Ruodi, Yang, Kailun, Peng, Kunyu
Abstract
Latent action models infer a code for the transition between two frames of action-free video. Recent methods regularise this code to compose additively and reverse antisymmetrically, and report order-of-magnitude reductions in the resulting errors as a label-free certificate that the code has captured temporal structure. We show that this conclusion does not follow. Reconstruction drives the decoded transition toward a difference of state features, for which both identities hold for any pairing, a solution the metric cannot distinguish from one encoding nuisance state or a coordinate convention. Across five source domains, a trained but unconstrained counterpart already achieves 83-97% of the reduction relative to an untrained anchor. The residual fold is governed as much by the decoder family as by what is learned. A constrained model retrained after its temporal pairing is destroyed still reaches, in each domain, a lower error than the unconstrained model on real data. Downstream, preserving the temporal pairing yields no consistent advantage on LIBERO-GOAL or LIBERO-SPATIAL, and across the tested arms the code's mean linear action decodability falls as the algebraic error improves. We also test the most direct repair, a violation-contrastive objective that requires the algebra to fail on destroyed pairings: in the tested configurations it yields only a marginal separation within the reconstruction budget, on training and test triples alike. We recommend a validation protocol that these methods currently lack: a baseline-corrected metric, retraining on destroyed pairings, and a seed-budget analysis.
Perception for embodied agents is video-based, often multi-view (ego, exo, or both), and inherently continual, with simultaneous task and viewpoint shifts. Yet continual learning (CL) remains dominated by exo-only recognition tasks, obscuring behavior under these real-world coupled shifts. We introduce Continual Ego, E}xo, and Ego-Exo Learning (CE$^4$L), a unified multi-view CL benchmark spanning four representative tasks: cross-view referenced skill assessment, temporal action segmentation, cross-view association, and action anticipation & planning. CE$^4$L highlights challenges largely absent in prior CL benchmarks, including cross-view correspondence, view-dependent asynchrony, and heterogeneous semantic objectives. To this end, we propose Video Incremental Subspace-routed Task Adapters (VISTA), a parameter-efficient baseline method that stores task-specific updates in lightweight adapters and performs training-free routing via residual distance to task-specific whitened subspaces estimated from second-order statistics. Extensive experiments demonstrate the significantly varied efficacy of representative CL methods across CE$^4$L settings, while VISTA is consistently competitive and achieves state-of-the-art overall performance. Our source code for benchmarks and methods is available at https://github.com/AnAppleCore/CE4L .
Failures of high-resolution MLLMs are commonly attributed to a visual problem, motivating zooming, cropping, and related visual interventions to recover fine-grained evidence or suppress interference. Yet recent studies suggest that relevant visual evidence is already encoded in intermediate representations, indicating that visual-side improvements alone insufficient. This raises a natural question: does the remaining bottleneck lie in the text that guides visual search? We identify a previously overlooked linguistic bottleneck: questions formulated for answering do not necessarily specify the visual evidence required for localization. To address this mismatch, we introduce EviSpec, a training-free compiler that derives complementary evidence specifications while preserving the original question for final reasoning. We further validate it through matched-control experiments that isolate the roles of evidence specification and localization. With the search budget fixed, structured evidence specifications yield an 8.6% relative gain over generic requests. With evidence geometry matched, the evidence localized by EviSpec yields a 14.8% relative gain over random evidence. Together, these controls isolate the benefit of specifying what evidence to seek rather than merely expanding visual access. Across all five MLLMs, EviSpec consistently improves upon the corresponding baseline on each of the three benchmarks, yielding average relative gains of \textbf{10.4%, 8.8%, and 12.4%} on V\textsuperscript{*}Bench, HR-Bench-4K, and HR-Bench-8K, respectively. Beyond high-resolution reasoning, EviSpec also achieves state-of-the-art performance on VQA and hallucination-focused benchmarks.
The increasing reliance on mobile phones has made phone-induced pedestrian distraction increasingly prevalent. Activities such as texting, watching videos, and making phone calls have become significant contributors to traffic accidents. Reliable detection of pedestrian distraction is essential for autonomous vehicles, as it improves situational awareness and enables timely risk assessment, thereby supporting safe motion planning and vehicle control. We propose a multimodal fusion Transformer (MFT) for detecting phone-induced pedestrian distraction. MFT jointly extracts skeletal dynamics from body pose keypoints and visual appearance features from pedestrian images, effectively leveraging the complementary information provided by the two modalities. A cross-modal attention module is proposed to capture inter-modal dependencies through multi-head cross-attention, facilitating effective fusion of complementary information across the two modalities. Then, a temporal attention fusion module, implemented with a Transformer encoder, is employed to capture temporal dependencies. MFT is trained and evaluated on a manually annotated dataset comprising 287 pedestrian instances with 20,741 images. Extensive experiments demonstrate that MFT attains an overall accuracy of 95%, exceeding the performance of six baseline approaches by 6%.
Novel view synthesis is a key task for dynamic scene reconstruction, where high rendering speed is essential for applications such as virtual reality. Existing deformable Gaussian Splatting methods achieve high-fidelity dynamic scene modeling, but still face limitations in memory usage and rendering efficiency due to the large number of redundant Gaussians. To address these challenges, we propose Geometry-Aware Redundancy Optimization (GARO), a unified redundancy measurement framework in the adaptive density control stage of the traditional dynamic scene reconstruction pipeline. This framework first selects low-gradient candidates using an optimization activity assessment strategy, and then evaluates geometric complexity through low curvature analysis to further filter and prune redundant points, resulting in a compact and expressive Gaussian representation. Extensive experiments on synthetic and real-world datasets demonstrate that GARO achieves robust trade-offs between quality and speed, with PSNR remaining stable and rendering speed improved by 2x, validating the efficiency and effectiveness of GARO.
Multimodal classifiers can converge to modality-dominant solutions in which one modality dominates the joint prediction, suppressing the learning of others. Existing balancing methods mainly adjust losses, gradients, or modality contributions, largely treating modality imbalance as an optimization problem while implicitly treating the weak modality as under-optimized but representationally intact. In this work, we find that this assumption does not always hold, as persistent modality dominance can induce a representation-level collapse of the weak modality, which we term \emph{manifold modality collapse} (MMC). MMC manifests as a coupled geometric degradation in which weak-modality representations collapse onto fewer directions within each class and become less separable across classes. Motivated by this observation, we propose \emph{GeoBalance}, a geometry-aware framework that monitors these two geometric properties and reconstructs the weak modality representation only when it exhibits signs of MMC. Once triggered, GeoBalance uses a fixed Simplex-ETF class scaffold and spectral regularization to restore class separation while preventing collapse onto a few feature directions. To preserve reconstruction during joint training, asymmetric gradient projection removes the joint-gradient component conflicting with reconstruction, leaving non-conflicting optimization unchanged. Extensive experiments across six multimodal benchmarks demonstrate great improvements over competitive balancing methods, validating its effectiveness.
PosEviLoc: Position-Conditioned Spatial Evidence for Language-Based 3D Localization
PosEviLoc:面向基于语言的3D定位的位置条件空间证据
Shang, Tianyi, Shi, Yike, Li, Zhenyu
Abstract
Language-based 3D localization retrieves the point-cloud submap containing a target position from descriptions of nearby objects and their spatial relations. Existing methods typically compress queries and submaps into global descriptors, potentially obscuring object-level semantics and cross-description spatial coherence. We propose Position-Conditioned Evidence Localization (PosEviLoc), a query-position-aware framework for coarse text-to-point-cloud localization. Instead of relying on global matching, PosEviLoc evaluates each candidate submap using explicit semantic and spatial evidence. It models direction as a relation jointly determined by an object position and a hypothetical query position. The resulting Query-Position Spatial Evidence Field (QSEF) measures the fraction of query descriptions supported at each hypothetical position, explicitly capturing their agreement without using the ground-truth query pose to construct the evidence field. A Multi-Level Evidence Readout (MER) summarizes this evidence in a compact representation, which a lightweight MLP converts into a retrieval score. Across five benchmarks, PosEviLoc outperforms MNCL by an average of 17 percentage points in Recall@1. When used as a plug-and-play reranker, it improves MNCL by an average of 16 percentage points. Moreover, PosEviLoc introduces substantially fewer parameters and achieves faster inference speed than existing methods.
Multimodal 3D object detection is fundamental to robust perception in autonomous driving because it integrates complementary information from LiDAR and camera sensors. However, existing methods often fail to maintain robustness under out-of-distribution (OOD) corruptions caused by sensor noise, adverse weather, and environmental changes. To address this problem, we propose RoboDistill, a robust and generalizable multimodal 3D object detection framework that leverages visual foundation models (VFMs), such as the Segment Anything Model (SAM). First, we introduce SAM-AD, a domain-specific pretraining strategy that fine-tunes SAM on autonomous-driving imagery to extract feature representations with rich semantic information. Second, we design the AD Feature Pyramid Network (AD-FPN) to refine and upsample SAM features at multiple scales for seamless fusion with LiDAR features. Third, we develop the Depth-Guided Wavelet Attention (DGWA) module, which suppresses high-frequency sensor noise while preserving critical contextual information. Finally, we introduce KD Fusion, in which the pretrained SAM-AD serves as a teacher that distills high-quality visual knowledge into a lightweight point-cloud network, thereby improving robustness under noisy conditions. Extensive experiments across 27 challenging OOD corruption settings show that RoboDistill generally delivers stronger or competitive detection performance and robustness relative to representative state-of-the-art methods. This work bridges the gap between VFMs and 3D object detection and advances robust multimodal perception for real-world autonomous-driving applications.
Generating sewing patterns from images and text requires modeling a heterogeneous representation composed of discrete topology and continuous geometry. Existing methods mainly follow two paradigms: diffusion-based methods enable holistic geometry generation by converting the entire pattern into a continuous representation, but weaken discrete topology modeling; in contrast, autoregressive methods preserve discrete topology through next-token prediction, but tie continuous geometry regression to token-level hidden states with limited panel-level context. To bridge this gap, we propose SewFusion, a unified autoregressive framework that adopts tailored generation mechanisms for discrete topology and panel-level continuous geometry, using next-token prediction for the former and flow matching for the latter. To support panel-level continuous geometry generation, we introduce a Panel Geometry VAE that learns a fixed-size latent space for variable-length panel geometry, together with Panel Geometry Flow for latent generation. We further propose Panel-Forcing to reduce the training--inference mismatch in topology context and improve robustness to topology prediction errors. Extensive experiments on SewFactory and GCD-MM demonstrate that SewFusion consistently outperforms previous state-of-the-art methods across various settings, achieving +6.36% Panel Accuracy, +11.30% Stitch Accuracy, and -1.90 Vertex L2 error in the image-text-based generation setting.
The SYNTAX score is a clinically established tool for assessing anatomical lesion complexity in coronary artery disease and guiding subsequent treatment. However, automated SYNTAX scoring is commonly formulated as a direct regression problem from coronary angiography videos to patient-level scores. In this work, we reformulate SYNTAX scoring as a vessel segment identity-preserving anatomical reasoning problem and propose a hierarchical modeling framework that explicitly aligns learning with the clinical workflow. Our approach maintains vessel segment identity across frames and views, estimates stenosis severity at the segment level, and aggregates evidence hierarchically according to coronary anatomy. Simultaneously, to address the scarcity of domain-specific data, we integrate and complete multiple public coronary angiography datasets, constructing a large-scale resource featuring completed vessel segmentation and derived structural annotations. Experiments demonstrate that vessel segment-level stenosis embedding enhances explanatory power and reduces prediction variability compared to baseline models, with the R^2 score improving by 0.201 and dev STD decreasing by 18.4%. These results highlight the necessity of structure-aligned modeling for reliable and stable automated SYNTAX scoring from multi-view coronary angiography videos. The GitHub link is https://github.com/VersaceSu7/SYNTAX_score_777.
Transferring Visual Explanations: How Cross-Architecture Knowledge Distillation Affects Model Interpretability
视觉解释的可迁移性:跨架构知识蒸馏如何影响模型可解释性
Czufarow, Aleks, Babin, Ihor
Abstract
Deploying efficient neural networks is essential in resource-constrained environments, yet compact models often sacrifice interpretability - a critical in safety-critical domains such as autonomous driving and medicine. This study investigates whether Knowledge Distillation transfers the spatial feature attribution of a large teacher network to a compact student. To assess the influence of the KD scheme on interpretability, we distill a ResNet-152 teacher into a ResNet-34 student on ImageNet-1K across five configurations by systematically varying the distillation temperature and soft-label loss weight. Models are evaluated on top-1 accuracy, along with two interpretability metrics: Relevance Mass Accuracy and Relevance Rank Accuracy. These metrics are computed via Grad-CAM heatmaps benchmarked against ground-truth object masks. Our results show that top-1 accuracy ranges from 71.6% to 74.0%. For Grad-CAM, RMA ranges from 7.7% to 9.7% and RRA from 7.3% to 10.1%; for Guided Grad-CAM, RMA ranges from 16.1% to 18.6% and RRA from 15.9% to 21.5%. Interpretability proves far more sensitive to the soft-label weight than to the temperature: keeping the student anchored to hard labels preserves both accuracy and coarse localization, whereas weighting the teacher heavily degrades both. Fine-grained attribution, however, fell below the undistilled baseline in every configuration tested, indicating that logit distillation transmits where a model attends more readily than the pixel-level structure of that attention. We evaluate 12 cross-architecture combinations of convolutional and transformer-based models, revealing that the inheritance of fine-grained spatial reasoning is fundamentally bottlenecked by the student's intrinsic structural biases. To our knowledge, this is the first application of this interpretability-aware evaluation framework - previously used for neural network pruning - to KD.
Chinese Translation
在资源受限的环境中部署高效的神经网络至关重要,然而紧凑模型往往牺牲了可解释性——这在自动驾驶和医疗等安全关键领域是极为关键的问题。本研究探究知识蒸馏(Knowledge Distillation)是否能将大型教师网络的空间特征归因迁移给紧凑的学生网络。为评估知识蒸馏方案对可解释性的影响,我们在ImageNet-1K数据集上,通过系统地改变蒸馏温度和软标签损失权重,以五种配置将ResNet-152教师模型蒸馏到ResNet-34学生模型中。模型以top-1准确率以及两个可解释性指标进行评估:相关性质量准确率(Relevance Mass Accuracy)和相关性排序准确率(Relevance Rank Accuracy)。这些指标通过Grad-CAM热力图相对于真实目标掩码进行计算。结果显示,top-1准确率介于71.6%至74.0%之间。对于Grad-CAM,RMA介于7.7%至9.7%,RRA介于7.3%至10.1%;对于Guided Grad-CAM,RMA介于16.1%至18.6%,RRA介于15.9%至21.5%。可解释性对软标签权重的敏感程度远高于对温度的敏感程度:让学生模型锚定于硬标签可以同时保持准确率和粗粒度定位能力,而过度偏向教师模型则会同时降低两者。然而,在所有测试配置中,细粒度归因均低于未蒸馏的基线,这表明logit蒸馏更易于传递模型关注的位置,而非该关注的像素级结构。我们评估了12种卷积模型与基于Transformer模型的跨架构组合,发现细粒度空间推理能力的继承从根本上受限于学生模型固有的结构偏置。据我们所知,这是该可解释性感知评估框架——此前用于神经网络剪枝——在知识蒸馏中的首次应用。
Vision-Language-Action (VLA) models integrate vision-language understanding with executable robot actions, enabling end-to-end learning for robot control. However, our empirical analysis reveals that existing models exhibit severe trajectory overfitting when finetuned on limited datasets. To guide the model in effectively utilizing wrist camera information, we propose MaskVLA, a masking-based fine-tuning strategy. By randomly masking a small portion of the main camera's visual information, the model is guided to autonomously learn more fine-grained, task-relevant, and effective visual features. This process leads to the emergence of robust policies, thereby enhancing the model's capability to tackle complex manipulation tasks and improving its generalization performance. Our method has been comprehensively evaluated on RoboTwin 2.0, achieving an average success rate improvement of 23.2% and 16.8% compared to $\pi_0$ and OpenVLA-OFT, respectively. Furthermore, experiments on real-world ALOHA robots also demonstrate the effectiveness of our approach.
6D object pose estimation is fundamental to robotic manipulation and automation. Recent zero-shot methods have significantly improved generalization to unseen objects, but most still rely on large-scale pretrained models with substantial GPU computation and memory demands. These requirements complicate deployment on robotic platforms where perception, planning, and control share limited computational resources, while learned intermediate representations offer limited geometric interpretability for task-specific adaptation. To address these limitations, we propose G6D, a learning-free, geometry-driven RGB-D 6D pose solver. Given an RGB-D observation, an object instance mask, camera intrinsics, and a CAD model, G6D generates pose hypotheses through template-based geometric matching and refines them using silhouette and depth consistency, forming a purely geometry-driven pose estimation paradigm. This paradigm requires neither pretrained visual models nor target-specific training and preserves interpretable geometric representations throughout pose estimation. Moreover, adjustable hypothesis counts provide flexible accuracy-computation trade-offs, while a CPU-only configuration supports deployment without GPU resources. Experiments on LineMOD and five BOP19 datasets demonstrate advanced performance. Real-world pick-and-place experiments further demonstrate G6D's applicability to robotic manipulation. The complete project is publicly available at https://ai4control.github.io/G6D-Project-Page .
The rapid development of video generative models (VGMs) has enabled the generation of highly realistic synthetic videos, raising concerns about the intellectual property rights (IPR) of these models. In particular, two closely related forensic tasks remain largely unaddressed: synthetic video verification (determining whether a video was generated by a protected VGM) and model ownership verification (determining whether a suspect VGM is an unauthorized copy of a protected VGM). In this paper, we propose a new in-generation watermarking scheme that can address the two verification tasks. First, a novel video watermarking network named VidMark is presented, which incorporates a two-scale discrete wavelet transform (DWT) decomposition and a global temporal attention block (GTAB) to enhance watermark robustness and imperceptibility. Second, we present a decoder-guided fine-tuning procedure. By leveraging the frozen VidMark decoder, this process enables VGMs to synthesize videos carrying an imperceptible, robust, and model-specific watermark. Finally, two verification frameworks are established to perform synthetic video verification and model ownership verification. Extensive experiments on representative VGMs demonstrate that the proposed scheme achieves over 99% watermark extraction accuracy and 100% verification accuracy on both tasks, with negligible impact on video generation quality. Furthermore, the watermarks exhibit strong robustness against a comprehensive range of video-level and model-level attacks.
Collapse, Not Complexity: Failure-Conditioned Decomposition Repair for End-to-End Document Parsing
是崩溃而非复杂性:面向端到端文档解析的失败条件化分解修复方法
Lin, Xingyu, Du, Dehui
Abstract
End-to-end document parsers increasingly offer an optional reasoning mode for complex pages. On a 180-page entropy-stratified discovery sample with one frozen 4B checkpoint, complexity is the wrong decision variable. Reasoning lowers mean quality by 2.21 Overall at 1.54x tokens; a preregistered input-only model cannot predict its signed benefit (held-out AUROC 0.47, indistinguishable from chance). The benefit concentrates on pages whose ordinary pass has already collapsed, and they do not look complex: shared collapses have lower layout entropy than healthy ones yet consume 19x the tokens as degenerate repetition that doubling the budget does not cure. Switching modes rarely repairs them: 83% recur under reasoning. We instead detect collapse from the ordinary-pass trace, decompose the page by projection, and re-parse each region. Repair gains 1.40 Overall (95% CI [0.68, 2.16]) at 1.13x tokens, replicates across three checkpoints, and, with all parameters frozen, gains 2.41 (CI [1.64, 3.46]) on the remaining 1,175 benchmark pages.
Chinese Translation
端到端文档解析器日益为复杂页面提供可选的推理模式。在一个包含180页、按熵分层的发现样本上,使用一个冻结的4B检查点,我们证明复杂性是错误的决策变量。推理模式以1.54倍的token开销使平均质量下降2.21 Overall;一个预先注册的仅基于输入的模型无法预测其有符号收益(留出集AUROC为0.47,与随机猜测无异)。收益集中在普通解析已经崩溃的页面上,而这些页面看起来并不复杂:发生共享性崩溃的页面其版面熵低于健康页面,却消耗19倍的token,表现为加倍的预算也无法消除的退化重复。切换模式很少能修复它们:83%的崩溃在推理模式下再次出现。我们转而从普通解析的轨迹中检测崩溃,通过投影分解页面,并对每个区域重新解析。修复方法以1.13倍的token开销提升1.40 Overall(95% CI [0.68, 2.16]),在三个检查点上均可复现,且在所有参数冻结的条件下,在其余1,175个基准页面上提升2.41(CI [1.64, 3.46])。
With face-recognition models now embedded in everyday authentication and surveillance, recent works have pinpointed a critical weakness: these models remain acutely vulnerable to adversarial semantic edits. I.e., adversarially produced semantic alterations to the input, such as slight aging or pose changes, can induce misclassifications. Certain existing attacks are powerful, but they can be computationally costly, rendering them inadequate for developing defenses (e.g., through adversarial training). To fill the gap, we introduce BoundStyle, a potent semantic attack operating in StyleGAN's rich latent space to maximize misclassification rates. Notably, BoundStyle achieves high attack success rates while being ${\sim}{\times}9.5$ faster than existing state-of-the-art attacks, making it suitable for adversarial training. Building on BoundStyle, we develop StyleAT, an efficient adversarial training scheme that incorporates low-budget attack variants yet defends against stronger and unseen semantic attacks. We evaluate on two datasets unseen during training and seven models, and find that StyleAT boosts robust accuracy against state-of-the-art attacks and outperforms common defenses in various settings.
PETR: Prompt Ensembling with Training-free Routing for Vision-Language Models
PETR:面向视觉-语言模型的无训练路由提示集成方法
Cai, Weihan, Tan, Hao, Gao, Xinping, Xu, Shibiao, Wan, Jun
Abstract
Prompt learning efficiently adapts vision-language models (VLMs) to downstream tasks, but gains on seen classes often come at the expense of generalization to unseen classes. To address this limitation, we propose prompt ensembling with training-free routing (PETR), whose key innovation is a carefully designed dual-prompt architecture: two complementary prompts are learned from different data and objectives to emphasize seen class discrimination and unseen-class generalization, respectively. During training, both prompts are fine-tuned using a shared frozen CLIP backbone, and statistical information is collected from the training set logits. At inference time, we determine the similarity of each test sample to seen data, and route the sample to the most appropriate prompt branch. To the best of our knowledge, this is the first prompt tuning framework that performs training-free adaptive routing based on statistical similarity. This design provides an interpretable routing signal and avoids common MoE-style routing pathologies, such as router training instability and load imbalance. Extensive experiments on 11 benchmark datasets demonstrate that our framework consistently outperforms previous methods on both seen and unseen classes, achieving new state-of-the-art results.
Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value (KV) caches. We introduce Prefix-Steered Recurrent Memory (PREM), a memory-token-free framework for frozen vision-language models (VLMs). PREM separates video ingestion from query answering: a recurrent writer distills visual streams into a compact 256 KiB multi-slot associative state, while a question-conditioned readout adds memory-derived key/value (K/V) steering modulations to existing non-visual prompt prefixes during prefill. This enables write-once, query-many inference without extra prompt tokens or decoding recurrence. Across six long-video benchmarks in offline and streaming end-of-stream settings, PREM consistently outperforms frozen baselines at every evaluated visual budget. Under a constrained budget of 16 frames, PREM improves macro-average accuracy by 3.06% on Qwen2.5-VL-3B, with gains of 11.0% on action antonym identification and 9.9% on localized needle retrieval. These gains require tuning 0.24% of backbone parameters at 0.03 GiB of peak GPU memory overhead.
Mesh texture compression typically relies on 2D UV atlases, whose chart discontinuities and mapping overhead can limit coding efficiency. To tackle this challenge, we introduce TexF, a surface-aligned texture field that organizes texture attributes in sparse voxels derived from the mesh surface. This representation supports high-resolution textures while preserving local 3D correlations for compression and enabling direct surface queries. For bitstream compression, TexF reuses established 3D attribute codecs, with voxel locations reconstructed from the decoded mesh without separate transmission. For GPU-resident compression, we develop 3DNTC, which combines quantized hash features with a lightweight decoder for random-access reconstruction at surface positions. Differentiable rendering enables image-space refinement of both voxel attributes and compressed neural fields. Experiments on the MPEG and AOM mesh compression benchmarks demonstrate improved average rate-distortion performance over representative UV-based methods for both bitstream and GPU-resident compression. 3DNTC also supports real-time rendering.
Snapshot hyperspectral imaging avoids sequential scanning, but systems that jointly achieve stable reconstruction, low cost, and compact optics remain limited. We present a snapshot hyperspectral imaging system based on angular-to-spectral diversity conversion. A tapered kaleidoscope creates replicated views with distinct incidence directions, and a directly attached birefringent filter converts them into view-channel-dependent spectral transmittances, yielding complementary measurements that better condition the inverse problem for more stable single-shot spectral reconstruction. The system preserves a simple pixel-wise linear model for fast non-learning-based reconstruction and uses only off-the-shelf components without relay optics or cascaded modules. We select the birefringent filter configuration using a condition-number-based criterion and validate the system on both synthetic and real data.
Imbalanced gradient magnitudes between the visual and language pathways of a vision-language model are often treated as a defect to be corrected. We test that premise for one family of correction, deliberately excluding adaptive, signal-driven schemes (e.g. BalGrad, OGM, PMR, CGGM), which are a mechanistically distinct class outside this study's scope. Measuring the language-to-visual gradient-norm ratio, reported in parameter-normalised form, across nine fine-tuning conditions, three seeds, and three captioning datasets spanning distinct physical domain shifts -- underwater, aerial, radiological -- we find imbalance magnitude varies markedly across domains with no predictable ordering. A plain learning-rate reduction cuts imbalance substantially and lands within a few BLEU points of the best method on every dataset. Staged freezing reduces the ratio on every domain yet never ranks first; a schedule-only control isolates freezing as the cause on one dataset but not the other two. Forcing the two gradient groups to equal magnitude drives per-parameter imbalance close to zero on every domain, yet is both the best result in the study and the worst placement among full fine-tuning methods, on different datasets, with identical settings. Reductions in gradient-norm ratio do not consistently predict captioning performance across domains, and how a given level of balance is reached matters as much as the level itself. As a secondary finding, a commonly reused LoRA configuration applied to BLIP silently adapts zero visual parameters; correcting it improves BLEU-4 on all three datasets.
Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-world physical laws. While existing solutions rely on external priors or specialized data, we investigate the root cause by exploring the internal mechanisms of these models. Specifically, we present the first interpretability study on the ''motion planning'' process of text-to-video diffusion models, revealing how motion trajectories form during early denoising stages. Building upon the ''first shape, then details'' finding, we combine cross-attention trajectory patterns with causal head contributions to identify a specific subset of attention heads driving motion planning. Further, our self-attention analysis shows that Rotary Position Embedding (RoPE) induces excessive spatial attention decay. This causes early candidate regions to prematurely lock into physically implausible positions, suppressing reasonable trajectories in adjacent frames and triggering generation failure modes. To address this fundamental flaw, we propose a lightweight architectural modification that scales the frequency of RoPE across different denoising steps. This strategy reduces excessive attention decay, helping the model explore better candidate regions to establish coherent physical motion. Finally, training-free and training-based experiments confirm the effectiveness of our approach in enhancing the physical commonsense of generated videos.
Which Terrain Is Better? Preference Learning with VLM Prototypes for Off-Road Traversability Ranking
哪种地形更好?基于VLM原型的偏好学习用于越野可通行性排序
Hwang, Ji-Hoon, Bae, Jisung, Son, E-In, Kim, Dong-Wook, Kim, Jung-Taak, Seo, Seung-Woo
Abstract
In vision-based off-road navigation, a robot needs to know not only which obstacles to avoid but also which terrain is better. The first is handled by freespace detection or semantic segmentation. The second is usually answered with a traversability score, but no universal ground truth exists for such a score, so perception falls back on a predefined value per semantic class or a freespace confidence. These scores say what a region is, not which region a robot should prefer. We therefore formulate this preference as visual traversability ranking, an ordering of visible terrain that can be supervised by comparisons between two regions. Standard annotations do not label preference, but they imply its direction. We present TravPro, which converts these annotations into ordered region pairs and fits a small readout on frozen vision--language model (VLM) patch tokens to these pairs. The tokens are clustered once into a fixed prototype bank, and the readout learns a preference score per prototype. The readout is then applied to every patch and serves as a teacher that turns sparse comparisons into dense preference pseudo-labels without pixel-wise annotation. An RGB student distills these maps into a dense terrain-preference map together with a non-ground mask that excludes obstacles and background from the ranking. On five unseen domains, TravPro reaches a mean pairwise accuracy of 0.915 against 0.783 for the strongest baseline, producing an ordering sensitive to surface condition that a per-class value cannot represent. The same VLM and the same supervision yield no such ordering when the VLM is prompted and the supervision is used as dense targets; what matters is how they are used.
Form Field Detection (FFD) is a fundamental component of document understanding systems, enabling applications ranging from large-scale industrial digitization to accessible form interaction for automated analysis. Unlike conventional object detection tasks, FFD is inherently challenging because fields are often defined by layout structure and whitespace rather than visible foreground content. Existing large-scale datasets frequently rely on heuristic annotation pipelines, resulting in noisy and inconsistent labels that hinder reliable evaluation. In this work, we introduce mini-CommonForms, a carefully curated FFD benchmark with consistent, high-quality annotations, and present a detailed evaluation of state-of-the-art detection approaches. The benchmark is designed to support reproducible research in document automation and accessibility-oriented applications. Dataset and code are available at https://github.com/moured/mini-commonforms
Chinese Translation
表单字段检测(Form Field Detection, FFD)是文档理解系统的基础组成部分,支撑着从大规模工业数字化到面向自动化分析的无障碍表单交互等多种应用。与传统的目标检测任务不同,FFD 本质上具有挑战性,因为表单字段通常由版面结构和空白区域定义,而非可见的前景内容。现有的大规模数据集往往依赖启发式标注流程,导致标签噪声大且不一致,妨碍了可靠的评估。在本工作中,我们提出了 mini-CommonForms,一个经过精心构建、标注一致且高质量的 FFD 基准,并对最先进的检测方法进行了详细评估。该基准旨在支持文档自动化及面向无障碍应用的可复现研究。数据集与代码可在 https://github.com/moured/mini-commonforms 获取。
Infectious bovine pinkeye is a contagious ocular disease that adversely affects cattle health, welfare, and agricultural productivity. Conventional diagnosis relies primarily on clinical observation, which can be subjective, time-consuming, and difficult to implement efficiently in large herds or remote settings. This study evaluated and compared You Only Look Once (YOLO) v11 and YOLOv26 for automated bovine pinkeye classification and investigated the effects of class-balancing strategies on model performance. Five variants (n, s, m, l, and x) of each architecture were trained and evaluated using the original imbalanced dataset, Random Minority Oversampling (RMO), and an adapted Synthetic Minority Oversampling Technique (SMOTE). Both YOLOv11 and YOLOv26 demonstrated strong classification performance, although the effects of class balancing varied across model variants. For YOLOv11, RMO-s achieved an accuracy of 0.99, a macro F1-score of 0.98, and a true positive rate (TPR) of 1.00, with no false-negative classifications. RMO-m also achieved a TPR of 1.00 with no false negatives. For YOLOv26, the original l, RMO-m, and RMO-l variants each achieved an accuracy of 0.99 and a macro F1-score of 0.98, with RMO-l attaining a TPR of 1.00 and no false negatives. Overall, RMO generally provided greater improvements in minority-class detection than adapted SMOTE, whereas the strong performance of the original YOLOv26-l demonstrates that oversampling was not necessary for all model variants. These findings demonstrate the potential of YOLOv11 and YOLOv26 for automated detection of bovine pinkeye and support further evaluation for livestock health monitoring.
Multimodal large language models (MLLMs) incur substantial computational overhead due to the reliance on hundreds of visual tokens to represent images. While token pruning has emerged as a promising approach to reduce the inference cost of MLLMs, existing methods typically reassign position embeddings to the retained tokens using either sparse or continuous position embeddings, each introducing distinct limitations. Sparse position embeddings tend to decrease the attention value allocated to visual tokens, thereby degrading the perception capability of MLLMs, whereas continuous position embeddings disrupt the original spatial correspondence of visual tokens, leading to weakened grounding capability. To mitigate this issue, we perform layer-wise analysis of the language decoder and observe that intermediate layers play a critical role for maintaining the grounding capability of MLLMs under token pruning. Based on this observation, we propose a layer-aware position embedding strategy, which switches to sparse position embeddings at grounding-sensitive layers while maintaining continuous position embeddings elsewhere. Extensive experiments across representative pruning methods and diverse benchmarks demonstrate that our approach improves the comprehensive multimodal performance of pruned MLLMs compared with standard sparse and continuous position embeddings.
Global vision--language similarities compress an image and a caption into one vector, preserving semantics but not which word corresponds to which region or how those regions are arranged; a model can recognize every word and object yet prefer a compositionally incorrect caption. We argue that a frozen encoder retains this association structure, so the problem is to read it rather than to rebuild it beside the pretrained similarity. We introduce BindCLIP, a pairwise scorer built on one latent object: a balanced token--patch--depth optimal-transport coupling that places both candidate captions and several visual depths in a single plan. Semantic, entity, order, and spatial evidence are read as energies of this state, and exchanging the candidates permutes the plan, making the score exactly antisymmetric. A geometric refinement inside the coupling contracts moves that the candidates and the visual depths do not support. No task label, parser, relation inventory, or detector is used. One checkpoint and one inference path improve the official What'sUp, ARO, and SugarCrepe benchmarks over frozen global CLIP, with the strongest transfer on the relation splits. Controls rule out patch access and caption-length shortcuts, and an inference-time lesion localizes spatial arrangement to the coupling.
VGGT-Prime: Compute-Adaptive Mixture-of-Heads for Efficient Visual Geometry Transformers
VGGT-Prime:面向高效视觉几何Transformer的计算自适应混合头模型
Arab, Abteen, Wu, Guile, Huang, Chengjie, Bai, Dongfeng
Abstract
Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-view images. Despite their promising performance, these models scale quadratically with the number of input views due to their global attention mechanism, resulting in substantial latency for long sequence inputs. There have been some recent efforts to accelerate VGGT, but they primarily focus on reducing \emph{token redundancy} through token merging or key/value sparsification. Our work resolves this bottleneck from a different perspective by investigating \emph{architectural redundancy} in visual geometry transformers. We show that the multi-head attention modules in VGGT's global-attention layers contain substantial architectural redundancy, with only a subset of heads carrying critical geometric information. In light of this observation, we propose VGGT-Prime, a compute-adaptive mixture-of-heads model that resolves this redundancy to accelerate visual geometry transformers while maintaining competitive reconstruction quality. The key idea of VGGT-Prime is to estimate the appropriate computation level for each global-attention head using a lightweight router and then dynamically assign each head to different computation modes. Extensive experiments on multiple datasets demonstrate that VGGT-Prime can achieve an {$8\times$} inference speedup over VGGT while maintaining competitive performance on camera pose, depth, and point-cloud predictions. We further show that VGGT-Prime is complementary to existing acceleration methods, such as token merging, further improving inference speed by up to $14{\times}$ over VGGT. An overview of our work is available on our \href{https://vggt-prime.github.io}{project page}.
Generative world models aim to predict future states conditioned on actions, where action controllability is fundamental for reliable dynamics modeling. While recent efforts leverage simulator-generated data to enhance this capability, existing training pipelines face two fundamental limitations. First, static offline data collection leads to a distribution misalignment between training sets and the model's evolving error patterns, failing to resolve critical long-tail scenarios where dynamics predictions remain unreliable. Second, the standard objective of minimizing observational discrepancy often encourages the model to exploit spurious correlations instead of capturing the underlying action-effect causality. To address these limitations, we propose OnlineWM, an online training framework that continuously improves world modeling through active simulator interaction and causality-aware optimization. OnlineWM introduces two key innovations: (1) Active Online Learning: Instead of using fixed datasets, OnlineWM adaptively queries the simulator for new interaction sequences that target the model's current predictive weaknesses, ensuring high-utility data acquisition. (2) Causality-Aware Fine-Tuning: We propose a counterfactual learning strategy that contrasts the outcomes of different actions from identical states, forcing the model to attribute state transitions to specific actions rather than ambient environmental evolution, thereby grounding its predictions in reliable causal mechanisms. By integrating active data acquisition with causal optimization, OnlineWM establishes a closed-loop refinement process that ensures the model is both robust to diverse scenarios and precise in its causal attribution. Extensive experiments demonstrate that OnlineWM significantly enhances action controllability and generalizes effectively to unseen domains.
Training-Free Spectral Transductive Refinement for Cross-Domain Few-Shot Classification
免训练的谱直推式精化方法用于跨域小样本分类
Rahman, Fahim, Rohan, S. M. Tanjeeb Meheran, Sayed, Md. Taimum Ibne, Herok, Asaduzzaman, Hasan, Md. Bakhtiar
Abstract
Few-shot recognition with frozen visual features is especially fragile under domain shift and one-shot supervision, where a single labelled image is an unreliable estimate of its class. We ask how far this fragility can be reduced purely at test time, without retraining the encoder or augmenting the source domain. We present Spectral Transductive Refinement (STR), a training-free transductive inference rule that exploits the geometry of the complete support-query episode. Given frozen embeddings, STR builds a joint k-nearest-neighbour graph, maps the episode into a normalized-Laplacian spectral coordinate system, initializes class representatives from the labelled support, and iteratively refines them using pseudo-labelled queries. We evaluate STR under two protocols. A controlled component study with frozen ResNet-18 features shows that spectral refinement consistently improves over single-prototype spectral initialization across five shifted domains, with the largest gains in the one-shot regime where support estimates are weakest. We then benchmark STR against recent Cross-Domain Few-Shot Learning (CD-FSL) methods using the standard miniImageNet-pretrained ResNet-10 backbone over eight established target domains. Operating entirely at inference time, STR attains the highest 1-shot average among compared methods and remains competitive at 5-shot, rivalling approaches relying on heavy source-domain meta-training augmentations. Because STR is transductive, we report its setting explicitly. Diagnostics attribute its gains to iterative refinement in spectral coordinates rather than added prototype capacity, which remains inactive in our configuration.
The disambiguation of semantically similar statutory text across jurisdictions is a retrieval problem that existing methods do not solve. This inter-context conflict can steer generative models toward confidently produced answers grounded in topically relevant but jurisdictionally incorrect sources. Tobacco and nicotine regulations vary by US jurisdiction, often sharing similar language, thus, robust reasoning requires identifying which jurisdiction's law governs a given product, not merely retrieving relevant text. Emerging products (e.g., pouches) exploit ambiguous definitions to evade regulation. State-of-the-art (SOTA) document retrieval-augmented generation (RAG) methods struggle to address this inter-context conflict, and thus struggle to connect image attributes (e.g., rich attribute captions) to the set of similar legislation texts. We introduce NicoPRISM (Nicotine Product and Regulation Image-and-Text Surveillance Multimodal), comprising 161,563 images, attribute captions, a knowledge base of product, health, and legislative documents spanning 13 US jurisdictions, and 1,495 validated question-answer pairs across two tasks: policy compliance QA and product knowledge QA. We also propose PRISM-RAG, a multimodal hypergraph RAG framework built over images, captions, and entities without any LLM calls at index time, grounding every query in a product image and routes retrieval through a jurisdiction-aware context assembly mechanism guaranteeing that statutory text from the queried jurisdiction reaches the language model by construction. PRISM-RAG retrieves passages from the correct jurisdiction in 93.9% of policy compliance queries, a 48.6 percentage point advantage over standard RAG (p<0.001), using zero LLM calls at index time and one at query time, and is competitive with or outperforms SOTA RAG frameworks across keyword, semantic, jurisdiction-, and compliance-accuracy metrics.
Single-image 3D object generation can now produce high-fidelity assets, yet accurately placing them into a coherent scene layout remains an open challenge. A central difficulty lies in how object layout is represented. Holistic methods absorb placement into a scene-level generation process, sacrificing object-level detail. Compositional methods preserve object fidelity by decoupling geometry from layout, but typically parameterize layout as sparse, unbounded pose variables that are difficult to learn and generalize poorly under scarce scene-level supervision.We present Mira-Scene, a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery. At its core is the Canonical Coordinate Map (CCM), a pixel-aligned field that maps each visible object pixel to a surface coordinate in the object's bounded canonical space. When paired with a scene-space Point Cloud Map (PCM) from monocular geometry estimation, CCM induces dense canonical-to-scene correspondences from which object transformations are recovered through robust geometric alignment. Because CCM operates in bounded canonical space, it provides a stable prediction target that can be trained from scalable object-level 3D data without requiring scene-level layout annotations. Mira-Scene further introduces a multimodal diffusion transformer that jointly generates object geometry and CCMs, using modality-specific expert streams with shared attention and positional encoding to promote geometry-layout consistency. Experiments on indoor, outdoor, synthetic, and in-the-wild scenes show that Mira-Scene substantially outperforms strong baselines in layout accuracy, achieving relative gains of 39.8% in 3D-IoU and 16.5% in 2D-IoU over SAM3D, using limited open-source training data.
Confidence-Aware Teacher-Student Distillation for 3D Medical Segmentation
面向三维医学分割的置信度感知师生蒸馏方法
Triantafyllou, Georgios, Iakovidis, Dimitris K.
Abstract
Medical image segmentation models typically rely on large amounts of densely annotated volumetric data, limiting their scalability across tasks and imaging modalities. This work addresses the challenge of predicting entire 3D anatomical structures from extreme annotation sparsity. An annotation-efficient student-teacher framework is proposed for automatic 3D medical segmentation that requires only a set of point prompts on a single 2D slice per volume, as input. A foundation model serves as an offline teacher, utilizing the provided point prompts from the selected slice to full-volume pseudo-annotations alongside their corresponding spatial confidence scores prior to student training. To mitigate the error propagation of noisy pseudo-annotations, a task-specific 3D student network is trained using a confidence-aware optimization strategy. By leveraging the teacher's pre-computed confidence scores, this strategy explicitly excludes statically uncertain regions of the pseudo-annotations from the loss calculation, while simultaneously emphasizing regions with higher confidence. Evaluated on 3D cardiac MRI datasets, our framework outperforms state-of-the-art semi-supervised methods, improving segmentation performance by up to 43.6%. Furthermore, it drastically reduces the manual annotation burden to just a few positive point prompts per volume, while improving surface boundary precision by up to 14.7% over the teacher and successfully recovering up to 34.1% of the performance gap toward the fully supervised upper bound.
Purkayastha, Monseej, Ghosh, Anindita, Slusallek, Philipp
Abstract
We present VISTA, a two-stage framework for generating stylized 3D human motion by fusing structural content from text prompts with expressive style from reference videos, without requiring jointly paired (text, video, stylized motion) triplets. A Dual-channel Autoencoder first maps motion sequences and video clips into a shared latent manifold. A masked autoregressive diffusion backbone then operates within this manifold, injecting video-derived style through a dedicated late-fusion Dual-AdaLN pathway while preserving text-conditioned content structure. A cross-batch unpaired training protocol with latent cycle consistency enables joint learning across separate semantically rich and stylistically diverse datasets. As a proof-of-concept for controllable animation synthesis, we validate VISTA on rendered motion-capture references: it achieves the highest style recognition accuracy among video-conditioned methods while preserving competitive content alignment, and its decomposed 3-way classifier-free guidance provides independent, user-controllable calibration of the content--style balance at inference time.
DFD-Lab: A Modular Audio-Visual Deepfake Detection Pipeline
DFD-Lab:一个模块化的音视频深度伪造检测流水线
Rybarczyk, Jan, Roszkowski, Mateusz, Komorowski, Jacek
Abstract
Comparing audio-visual deepfake detectors requires coordinating dataset adaptation, temporal input representation, model interfaces and experimental conditions. We present DFD-Lab, a modular pipeline that separates these responsibilities while supporting shared training and evaluation workflows. We integrate three implementations: Xception-based maximum-logit fusion, ResNet with temporal LSTM fusion, and our AVFF reimplementation. Experiments cover external testing, degradation-based training augmentation and evaluation-time corruption. On a filtered subset of Deepfake-Eval-2024, models trained on FakeAVCeleb attain baseline AUROC values of 0.504, 0.538 and 0.458. JPEG50 training augmentation raises these to 0.691, 0.605 and 0.570, respectively, while all three accuracies decrease. These results illustrate why training interventions, evaluation corruptions and metric-dependent outcomes should remain distinct within a common pipeline. The contribution is the integration of audio-visual processing, interchangeable detectors and configurable experimental workflows, supported by empirical case studies. The findings highlight the challenge of cross-dataset detection and the complementary information provided by ranking and classification metrics.
Automated grading of conjunctival photographs could reduce the cost and variability of trachoma prevalence surveys, but the relative value of modern pretrained visual representations, lightweight feature adaptation, and training-objective design has not been established under a common protocol. This study presents a controlled evaluation for binary classification of Trachomatous Inflammation-Follicular (TF) versus Normal using 1,546 images from the public UCSF/Lietman collection. Images are processed using the OPTED pipeline for zero-shot tarsal-conjunctiva segmentation, alignment, cropping, and standardization. We first compare six pretrained backbones using a common classification pipeline and then evaluate four lightweight adaptation mechanisms on DINOv2 ViT-B/14. Under stratified five-fold cross-validation, DINOv2 with Efficient Channel Attention (ECA) and focal-plus-center loss achieved 91.66 +/- 0.97% accuracy, 90.69 +/- 1.10% macro-F1, and 96.06 +/- 0.71% AUC. ECA introduces only five learnable parameters while matching the performance of substantially larger alternatives. Objective ablation further showed that ECA did not consistently improve plain DINOv2 across loss functions; the lowest-variance 91.66% accuracy was obtained with cross-entropy plus center loss. Overall, the fine-tuned DINOv2 representation provided most of the predictive performance, while ECA offered a highly parameter-efficient refinement whose effect depended on the training objective. The resulting workflow provides a reproducible benchmark for active trachoma image classification.
Joint Embedding Predictive Architectures (JEPAs) are a promising paradigm for learning task-agnostic latent world models without visual reconstruction. However, standard JEPA training exhibits a strong inductive bias towards slow features, causing feature suppression and the collapse of latent representation. While inverse dynamics provides temporal anti-collapse, it relies on action labels and offers little incentive to embed general, unlabeled dynamics. We introduce Difference Image and Single image embedding Regularization (DISReg), a novel regularizer that builds on an inverse-dynamics-style module that predicts temporal difference image embeddings without any pixel reconstruction loss, encouraging balanced static and dynamic feature learning. DISReg consists of a static term that shapes the distribution of the image embedding and encourages slow features, and a dynamic term, which, unlike direct regularization on the embedding, imposes no constraint on the image embedding's shape or distribution and instead only incentivizes that dynamic features be present. By integrating this regularizer into a standard JEPA, we establish our new architecture, MotionJEPA. Latent probing demonstrates that MotionJEPA produces more complete representations than other methods, and our trajectory analysis shows it maintains geometrically simple latent embeddings with low curvature. We further show that MotionJEPA improves downstream planning success under static-background distractors across four environments.
Chinese Translation
联合嵌入预测架构(Joint Embedding Predictive Architectures, JEPAs)是一种无需视觉重建即可学习任务无关的潜在世界模型的有前景范式。然而,标准的JEPA训练对缓慢变化的特征存在强烈的归纳偏置,导致特征抑制和潜在表示的坍缩。虽然逆动力学方法能够提供时间维度上的抗坍缩能力,但它依赖于动作标签,且几乎无法激励模型嵌入通用的、无标签的动态信息。我们提出了差分图像与单图像嵌入正则化(Difference Image and Single image embedding Regularization, DISReg),这是一种新颖的正则化方法,它基于一个逆动力学风格的模块,在不使用任何像素重建损失的情况下预测时间差分图像的嵌入,从而鼓励均衡的静态与动态特征学习。DISReg包含一个静态项,用于塑造图像嵌入的分布并鼓励学习缓慢变化的特征;以及一个动态项,与直接对嵌入进行正则化不同,它不对图像嵌入的形状或分布施加任何约束,而是仅激励模型保留动态特征。通过将这一正则化器集成到标准的JEPA中,我们构建了新的架构MotionJEPA。潜在探测实验表明,MotionJEPA相比其他方法能产生更完整的表示,且我们的轨迹分析显示它保持了具有低曲率的几何上简洁的潜在嵌入。我们进一步证明,在四种环境中,面对静态背景干扰物时,MotionJEPA提高了下游规划任务的成功率。
Monocular colonoscopic 3D reconstruction is important for surgical robotic colonoscopy, but remains challenging due to weak texture, specular reflections, limited view overlap, and non-rigid tissue motion. Conventional multi-view 3D reconstruction methods rely on stable correspondences and approximate rigidity, which are often violated in colonoscopy. Existing endoscopic methods often rely on domain-specific supervision, whereas there are not enough in-vivo labeled data available to adapt geometry foundation models to clinical colonoscopy. We present Colon3R, a cross-domain semi-supervised framework built on pretrained VGGT that transfers coupled camera, depth, and pointmap geometry from labeled phantom and simulated data to unlabeled in-vivo colonoscopy without requiring target-domain geometric annotations. Unlike source-only fine-tuning, which learns only from phantom and simulated data, Colon3R directly exploits unlabeled in-vivo video through teacher-derived cross-view supervision. Our proposed hierarchical quasi-rigid reliability selects reliable supervision at the sequence, directed-pair, and pixel levels, while source-preserving adaptation retains the learned coupled geometry during target-domain adaptation. Extensive experiments demonstrate that our method achieves superior overall performance over state-of-the-art approaches in depth, pointmap, and camera pose estimation. Qualitative comparisons on real in-vivo colonoscopy further show substantially more complete and geometrically consistent reconstructions than competing methods under clinical domain shift. The code will be public available after the paper is accepted.
Rethinking Diffusion Segmentation: When Does It Rely on Its Noisy State, and Does Diffusion Matter?
重新思考扩散模型分割:它何时依赖其噪声状态,扩散又是否重要?
Yang, Hengzhuo, Zeng, Yuming, Yang, Yuling
Abstract
Diffusion models are increasingly adapted from generation to conditional prediction, where a conditioning signal is combined with an evolving noisy representation of the target. In fully supervised segmentation, however, the conditioning image can already support direct target prediction, so endpoint performance alone establishes neither reliance on the added diffusion state nor a deterministic advantage over image-only prediction. For state reliance, we disrupt target-derived state content or correct image-state pairing during retraining of twelve published methods across three datasets, with ten matched seeds per setting. All 40 original-method comparisons whose evaluated-mask routes remained downstream of noised-quantity reconstruction exhibited state reliance, whereas all 30 comparisons with a segmentation-supervised bypass preserved reference performance. Rerouting five originally bypass-capable methods by forcing segmentation supervision through noise-to-mask reconstruction converted all 30 corresponding comparisons from preserved performance to state reliance. For deterministic utility, matched image-only counterparts achieved similar or better performance in 28 of 35 settings overall, including 16 of 20 whose native methods relied on both audited state properties. These results identify supervision path as a determinant of state reliance in the audited methods. Separately, matched image-only counterparts show that diffusion-specific computation often provides no deterministic endpoint advantage, including in methods that rely on the audited state properties. More generally, when conditioning already supports strong target prediction, diffusion-specific claims require additional evidence that the added state is used and that diffusion-specific computation improves the claimed capability beyond a matched condition-only counterpart.
Neuroimaging foundation models pretrained on large, predominantly western cohorts are increasingly proposed as general-purpose backbones for brain MRI analysis. Yet, their ability to generalize to underrepresented clinical populations remains largely untested. We evaluate four recent foundation models (BrainIAC, Neuro-JEPA, NeuroVFM, and Primus) on a three-way diagnostic classification task (Control, Dementia, Parkinson's disease) using a cohort of 88 subjects from a Nigerian clinical brain MRI dataset, across four modality configurations (T1w, T2w, T1w+T2w, FLAIR), and compare against an end-to-end trained ViT3D baseline. The frozen backbones collapse to majority-class predictions, while Neuro-JEPA on FLAIR shows modest but still limited discrimination. In contrast, the end-to-end trained ViT3D achieves higher accuracy and MCC on every task (up to 53.4% accuracy, MCC=0.27) and is the only model with non-trivial recall. Our findings suggest that these frozen neuroimaging foundation models are insufficient for fine-grained diagnostic classification in small, non-western clinical cohorts, motivating parameter-efficient adaptation and broader multi-site external validation for equitable deployment in global health settings.
Reliable quality control (QC) of magnetic resonance imaging (MRI) is essential for reliable diagnostic neuroimaging, yet standard manual assessment is subjective and time-consuming. MRIQC has established standardized automated extraction of image-quality metrics (IQMs), but its reliance on local computational imaging skills and capacity including high-performance computing, limits its adoption in resource-constrained settings (RCS). We present WebMRIQC (webmriqc.mailab.io), an open-source browser-based platform that wraps the validated MRIQC engine behind a zero-installation web interface. WebMRIQC automates the DICOM-to-BIDS conversion of de-identified MRI scans, executes the unmodified containerized MRIQC pipeline on a shared compute node governed by a fair-share job queue, and returns an interactive in-browser dashboard. The dashboard grounds every IQM in published quality thresholds, benchmarks each scan against the normative distribution of high-resource open datasets, and supports cross-site multicentre implementation of optimized scan protocols in RCS.We describe the system architecture and a validation framework establishing measurement equivalence between WebMRIQC and native MRIQC across thirteen IQMs on the BraTS-Africa and BraTS 2021 datasets. Preliminary results indicate strong agreement for contrast-, signal and noise-based metrics, demonstrating that web-based implementation lowers the barrier to standardized MRI QC and provides a foundation for harmonized, regionally adapted quality benchmarks across RCS imaging sites. The code is publicly available here https://github.com/CAMERA-MRI/WebMRIqc.
InterHier: Learning Interconnected Hierarchical Semantics for Open-Vocabulary Object Detection
Kim, Yeong-Jin, Kim, Ho-Joong, Lee, Seong-Whan
Abstract
In this paper, we investigate the limitations of fixed, hand-crafted connectors in hierarchical semantic representations for open-vocabulary object detection. Existing methods establish semantic relationships between base categories and unseen novel categories by placing a fixed connector between adjacent super-/sub-categories. However, such fixed connectors may not optimally capture the relationships within a semantic hierarchy. To address this limitation, we propose interconnected hierarchical semantic representations (InterHier), which utilize a prepended learnable context to globally guide the interpretation of prompts containing hierarchical relationships. InterHier operates in two main stages. First, it constructs a hierarchy-aware prompt by integrating super-/sub-categories and prepending a learnable context. Second, it optimizes this learnable context to align visual region embeddings and textual embeddings. InterHier consistently improves performance over methods that rely on fixed connectors and can be seamlessly integrated into existing open-vocabulary object detection models. Experiments on open-vocabulary object detection benchmarks demonstrate that InterHier achieves competitive performance against state-of-the-art methods.
In recent years, pre-training has become fundamental to learning effective video representations, enabling strong transfer to downstream tasks. A popular framework in pre-training involves aligning features of a video encoder with that of another modality, for example, language or audio. We introduce Video-STLayout pre-training, a novel strategy for obtaining rich video representations informed by spatio-temporal layout of object bounding boxes. Object layouts can easily be obtained by applying an off-the-shelf object detector on the video frames. Our method uses a contrastive loss to align video features with the layout features from a trained layout encoder. We show the effectiveness of our approach in the task of activity recognition in complex scenes.
U-PEN Mamba: Progressive Expansion with Selective State-Space Modeling for Efficient Retinal Vessel Segmentation
Reyes-Angulo, Abel A., Paheding, Sidike, Asari, Vijayan K., Alam, Mohammad, Devagiri, Jeevan
Abstract
Accurate retinal vessel segmentation is important for computer-aided ophthalmic analysis, yet thin vessels, low contrast, and severe foreground-background imbalance remain challenging for encoder-decoder networks. This paper presents U-PEN Mamba, a U-shaped retinal vessel segmentation architecture that couples progressive nonlinear feature expansion with selective state-space modeling. The proposed network enriches local vessel responses with progressive expansion, models long-range spatial dependencies through a Mamba Global Context (MGC) block with linear sequence complexity, and uses attention-based decoder fusion to recover fine vascular boundaries. We evaluate U-PEN Mamba on CHASE DB1 and DRIVE using a consistent patch-based preprocessing pipeline and compare it with convolutional, attention-based, transformer-based, and Mamba-based segmentation baselines. U-PEN Mamba obtains the best mean intersection over union among the compared methods, achieving 0.8394 on CHASE DB1 and 0.8221 on DRIVE, with Dice scores of 0.8187 and 0.8078, respectively, using 21.6M trainable parameters. Ablation studies show that the MGC block contributes the largest gain over the U-Net baseline, while projection dimension and state size provide practical accuracy-efficiency control. These results indicate that selective state-space modeling is a promising global-context mechanism for parameter-efficient retinal vessel segmentation. Code is available at: https://github.com/areyesan/UPEN_Mamba.
All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
基于文字系统感知混合专家的一体化多语言场景文本识别
Ye, Xingsong, Du, Yongkun, Zhang, Jiaxin, Li, Zhixian, Sun, Chong, Li, Chen, Lyu, Jing, Jin, Lianwen, Chen, Zhineng
Abstract
Multilingual scene text recognition (STR) remains challenging due to the scarcity of training data for most languages and the difficulty of serving diverse scripts within a single model. Existing solutions either deploy one recognizer per language, inflating cost and introducing error accumulation, or rely on massive vision-language models (VLMs) that are expensive and still inaccurate on many scripts. In this work, we pursue an all-in-one multilingual recognizer that is simpler than per-language experts, lighter than VLMs, and more accurate than both. First, we construct TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages. It provides balanced and sufficient supervision where real data is unavailable. Second, we propose ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture. It shares a single visual encoder and replaces the dense decoder with a sparse MoE block, which consists of an image-level router dispatches each image to the top-2 script-aligned experts and a shared expert absorbs cross-script knowledge. Extensive experiments on our assembled TextMuSS-Bench (10 scripts, 10,899 images) show that ScriptMoE achieves the highest accuracy of 82.06%, outperforming the strongest STR baseline by 1.31%. On the CC-OCR end-to-end multilingual task, replacing only the recognizer in PP-OCRv5 with ScriptMoE lifts F1 score from 65.71% to 80.89%, slightly surpassing the best VLM (80.73%) at a fraction of the parameter count.
Vision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea
New Guinea is the world's richest island flora (~2,856 orchid species), yet most species are represented by only a handful of photographs, far fewer than direct species-level classification requires. Methods for fine-grained identification in such species-rich, data-poor floras are needed, and it remains unclear which backbone architecture and pretraining strategy best support them. We built a two-stage system that first predicts the genus of a query photograph, then retrieves visually similar reference images of candidate species using FAISS. We compared four pretrained backbones -- two Vision Transformers (ViTs; DINOv2, BioCLIP 2) and two CNNs (ConvNeXt V2-L, EfficientNetV2-L) -- fine-tuned under an identical protocol on a fixed, species-stratified partition of 16,701 photographs spanning 120 genera and 1,350 species, assessing accuracy, calibration, error structure, species retrieval, and open-set detection of novel genera. DINOv2 attained the best genus performance (macro top-1 66.9%, 95% CI 63.7-70.6; global top-1 88.9%); both ViTs outranked both CNNs, and general-purpose self-supervised pretraining (DINOv2) outperformed domain-matched biological pretraining (BioCLIP 2) by 7.1 points of macro top-1. Errors concentrated on two abundant genera acting as error attractors. DINOv2 embeddings achieved species Recall@5 of 86.6% and genus Recall@5 of 98.7%; temperature scaling reduced every backbone's Expected Calibration Error to about 0.03; and a distance-based open-set gate flagged unseen genera (mean AUROC 0.958). A self-supervised Vision-Transformer backbone combined with embedding retrieval is an effective, deployable strategy for fine-grained identification in species-rich, data-poor floras. The system is released as an open web application (the New Guinea Orchid Identifier), offering a practical template for other hyperdiverse, under-documented taxa.
Chart reasoning agents are increasingly used to extract actionable insights in critical domains, achieving state-of-the-art performance on multiple benchmarks. Yet, high benchmark accuracy alone is insufficient for deployment, where stakeholders must be able to audit and verify how a model reaches its answer. Existing LVLM-based chart agents produce either answer-only predictions or free-form rationales that are hard to verify, obscuring whether an error arose from misreading the chart, extracting a wrong value, or miscomputing. We propose Chart-RVR, a reinforcement learning framework for training monitorable chart agents with verifiable process rewards. Chart-RVR decomposes chart reasoning into three auditable blocks: Structure, identifying the chart type; Evidence, reconstructing the underlying data table in JSON; and Derivation, exposing the stepwise trace that computes the answer. Across six in-domain and out-of-domain benchmarks, Chart-RVR attains state-of-the-art accuracy among comparable-sized LVLMs. Beyond accuracy, we assess monitorability using a triangulated protocol that combines ground-truth surrogate metrics, an oracle information-gain measure, and an LLM-as-auditor scoring Process Verifiability and Evidence Localization, showing that Chart-RVR yields rationales that are markedly more verifiable and evidence-grounded than those from CoT prompting, SFT, and existing chart-specific baselines.
Computer-vision systems used for construction monitoring can degrade under adverse environmental and visual conditions, yet such conditions remain underrepresented in existing construction image datasets. We present ConSynth-X, a paired synthetic construction-site image dataset containing 34,199 images derived from 3,109 real-world source scenes. The dataset comprises 11 condition-specific subsets spanning precipitation, fog, nighttime illumination, adverse weather at night, and small-object or long-distance views. Each synthetic image is linked to its corresponding source scene, enabling controlled comparison across environmental and visual conditions. ConSynth-X includes source-derived annotations, generation metadata, provenance information, and image-quality indicators, supporting object detection, image captioning, visual grounding, and visual question answering. Technical validation evaluates source-synthetic fidelity and alignment with real adverse-condition imagery using embedding-based similarity and distributional analyses. The dataset provides a structured resource for evaluating and improving the robustness of construction vision and vision-language models under challenging field conditions.
Bridging Reconstruction and Generation: A Latent Distribution Perspective on Evaluation and Improvement
连接重建与生成:基于潜在分布视角的评估与改进
Fang, Xianghong, Shu, Wenjie, Xu, Tongda, Mou, Wenlong, Kong, Dehan, Rudner, Tim G. J.
Abstract
In latent generative models, reconstruction quality is often assumed to correlate with generative performance. However, reconstruction FID (rFID) can exhibit weak or even negative correlation with generation FID (gFID). We attribute this discrepancy to a latent distribution mismatch: reconstruction evaluates the decoder on encoder-induced latents, whereas generation uses the same decoder on latents produced by the generative model. To characterize this shift, we introduce generation-aware reconstruction (GAR), which constructs a continuous trajectory from standard reconstruction toward generation by perturbing encoder latents with noise and denoising them through the generative model before decoding. GAR probes the decoder behavior along this trajectory, making the transition from encoder to generation-time latent distributions observable and diagnosable. The resulting trajectory-based diagnostic, GAR-FID, exhibits strong empirical correlation with gFID across diverse tokenizers and scales. Importantly, intermediate GAR latents become more generation-aware while preserving correspondence with their source images, thereby retaining paired supervision that is absent for fully generated latents. This correspondence enables decoder adaptation on intermediate GAR latents, consistently improving generative quality across model scales. Overall, latent distribution mismatch provides a useful perspective for evaluating and improving latent generative models.
While deep learning has enabled language decoding from intracranial brain recordings, extending this capability to non-invasive recordings remains an unresolved challenge. Decoding individual words from non-invasive brain recordings is particularly difficult, as word-level neural evidence is weak, temporally distributed, and entangled with acoustic, lexical, and semantic structure. Existing retrieval pipelines often collapse these factors into a single representation, potentially discarding information available at intermediate temporal scales. Here, we introduce Hierarchical Dynamic Neural Decoding (HDND), a hierarchical dynamic decoding framework that treats word decoding as structured refinement rather than flat label retrieval. HDND combines intermediate neural representations, contextual semantic predictions, and, for selected reading conditions, an auxiliary character-form objective. We evaluate HDND across seven electroencephalography (EEG) and magnetoencephalography (MEG) datasets spanning English, Dutch, Mandarin, and Cantonese listening, reading, and reading-aloud conditions. Across the nine-condition word-retrieval benchmark, the proposed HDND yields a higher participant-averaged balanced Top-10 point estimate than the matched contextual word-decoding baseline in every condition and achieves the highest mean among all compared methods in eight of nine conditions. Across the same nine matched conditions, HDND also yields higher token-micro and pooled word-macro Top-10 point estimates in every setting. Sentence retrieval favors HDND in eight of nine conditions, while auditory speech-segment retrieval is mixed across the six listening conditions. These results show that hierarchical residual refinement can improve multilingual word retrieval from heterogeneous non-invasive brain recordings.
Visual Question Answering (VQA) with Multimodal Large Language Models (MLLMs) requires not only producing safe and effective responses, but also grounding safety decisions in the multimodal evidence that determines risk. Recent safety-alignment methods improve refusal behavior and contextual risk awareness, yet correct safety outcomes may still rely on superficial textual or visual correlations, particularly when risk emerges from interactions between individually benign image and question content. To address this issue, we propose A$^2$Safe, a counterfactual evidence-aligned adaptive agent collaboration framework for safe and effective VQA. A$^2$Safe organizes localized visual observations, textual intent, and cross-modal risk relations through a Grounded Safety Evidence Board, making the basis of safety decisions explicit. Counterfactual safety evidence alignment enforces invariance to safety-irrelevant changes while requiring appropriate safety-state and response-mode transitions when risk-critical evidence is minimally altered. The resulting evidence state further supports adaptive collaboration, enabling direct answering when grounded evidence is sufficient and invoking policy critique and response revision when evidence is risky, uncertain, or conflicting. Under complementary safety-critical and general VQA protocols, A$^2$Safe achieves a 95.72 SIUO safety score, reduces the benign refusal rate on MOSSBench to 14.67%, and maintains an average general VQA score of 78.34 with 27.8% token overhead. These results support counterfactual evidence-aligned adaptive collaboration for safe and effective multimodal question answering.
Adaptive Cortically Constrained EEG-Vision Alignment for Zero-Shot Brain-to-Image Retrieval
面向零样本脑到图像检索的自适应皮层约束脑电-视觉对齐方法
Wang, Ye, Ren, Haokun, Wu, Wei, Wang, Guoyin, Yu, Zhuliang, Yu, Hong, Liu, Ke
Abstract
Zero-shot brain-to-image retrieval requires robust alignment between noisy EEG responses and visual representations. Existing EEG-vision alignment methods often operate in sensor space and apply fixed visual supervision to all responses, ignoring both spatial mixing in scalp EEG and response-wise variability in alignment reliability. We propose an adaptive cortically constrained EEG-vision alignment method for zero-shot brain-to-image retrieval. The method reconstructs EEG responses into predefined ROI-level source-pattern representations and encodes them with a Neuro-ROI Attention Encoder. To handle response-wise variability, we introduce an evidence-based adaptive visual supervision strategy that weights detail-controlled visual targets using model-based alignment evidence. On THINGS-EEG, the proposed method achieves strong 200-way zero-shot retrieval performance, with ROI-level attribution providing post hoc interpretability of the learned source-pattern representations. These results show that cortically constrained representation learning and adaptive supervision can jointly support EEG-vision alignment for zero-shot brain-to-image retrieval.
Patch-to-Global: Random Patch Diffusion for Globally Consistent Megapixel Artifact Inpainting in Whole Slide Images
Lee, Hyeseong, Kim, Eunsu, Bappy, D M, Kim, Ho Heon, Lee, Youngsuk, Chun, Se Young, Choi, Jang-Hwan, Lee, Sung Hak, Ahn, Sangjeong
Abstract
Although deep learning has advanced Whole Slide Image (WSI) Analysis, tissue artifacts like bubbles and folds often cause silent failures by concealing essential morphology. Current pathology image restoration methods are mostly restricted to small patches, struggling to maintain global structural coherence at a megapixel scale. We introduce RestorePath, a framework for globally consistent megapixel scale inpainting that reconstructs diagnostic structures in histological image to prevent incorrect high-confidence predictions and lower error rates. Our model utilizes a Latent Diffusion Model (LDM) conditioned on Pathology Foundation Model (PFM) embeddings, integrating Large Kernel Attention (LKA) to manage long-range dependencies during random patch diffusion. Enhanced by Distance-Weighted Interpolation (DWI) and an Adaptive Guidance Scale (AGS), RestorePath ensures structural consistency and fidelity by modulating information from surrounding patches. Evaluations across TCGA-BRCA, BACH, and Camelyon16 datasets for images ranging from 512 to 4608 pixels demonstrate state-of-the-art performance in maintaining histological consistency. RestorePath significantly improves downstream Computational Pathology (CP) tasks, outperforming both raw artifact images and the conventional Detect-and-Discard (D&D) approach. The code is available at https://github.com/PathfinderLab/RestorePath
Positive Pair Geometry Matters: Optimal Transport for Contrastive Learning of Visual Representations
Nanda, Akshit, Ahmad, Shahzad, Padhy, Ram Prasad
Abstract
Contrastive self-supervised learning has achieved strong performance by learning representations from multiple augmented views of the same image. However, most existing methods construct positive pairs using independently sampled stochastic augmentations, which may alter semantic content and ignore the intrinsic geometry of the data distribution. In this work, we propose OTCLR, an optimal transport-aware framework for contrastive learning representations that generates geometry-consistent positive samples. Instead of directly contrasting two randomly augmented views, we construct intermediate views between the original image and its augmented variants through entropic optimal-transport displacement interpolation. These transport-interpolated samples serve as positive views that better preserve image structure while explicitly modeling spatial distributional geometry. To further promote smooth representation learning, we evaluate auxiliary Sinkhorn regularization terms that encourage transport-interpolated views to remain consistent with their endpoint images. The proposed method can be incorporated into standard contrastive learning pipelines without modifying the encoder architecture. Experiments on multiple benchmark datasets show that our approach improves representation quality and transfer learning performance compared with conventional augmentation-based contrastive learning baselines.
Atomic activity understanding aims to recognize and localize structured traffic behaviors that jointly encode motion patterns and their grounding in road topology. Unlike conventional action recognition, atomic activities are multi-agent, multi-label, and topology-aware: multiple activities co-occur while many agents remain inactive. We introduce Action-Slot, a structured action-centric representation learning framework. Slot attention is widely used for object-centric decomposition, but its permutation-invariant design and object-level inductive bias are misaligned with atomic activity semantics. We reformulate slot learning as structured activity decomposition through three designs: (1) category-aligned action slots that anchor slots to predefined activity categories, (2) parallel spatio-temporal slot updating for holistic video-level reasoning, and (3) background and negative-slot regularization that enforces competition between foreground activities and irrelevant regions. Together these establish an activity-centric inductive bias that disentangles concurrent and asynchronous activities directly from raw video. Beyond recognition, the learned representations encode transferable spatio-temporal grounding signals. We further propose an attention-difference-based pseudo mask selection framework that suppresses false positives by measuring attention changes before and after candidate region removal, enabling weakly supervised localization without dense annotations. To support systematic evaluation, we introduce TACO, a balanced synthetic dataset with full atomic activity coverage and pixel-level annotations. Experiments on OATS, TACO, and annotated nuScenes show superior recognition, strong sim-to-real transfer, and state-of-the-art weakly supervised localization.
Brain-to-image retrieval seeks to identify the visual stimulus that elicited a non-invasive neural response. Candidate images are typically represented by pretrained vision models, whose internal representations vary in abstraction across depth. Existing methods usually train the neural encoder to recover a fixed final-layer visual target. Under this formulation, the visual hierarchy is reduced to a single prescribed endpoint, preventing representations at other depths from directly shaping the visual target. This limitation motivates learning how information across visual depths should contribute to the retrieval target. To this end, we introduce NeuroGlyph, which learns a trial-independent visual target from multiple depths of a frozen visual backbone. NeuroGlyph decomposes the target into factor-specific subspaces. Each subspace learns an image-conditioned allocation over visual depth. The resulting subspaces are fused into a single embedding for retrieval. Across THINGS-EEG and THINGS-MEG, NeuroGlyph outperforms final-layer supervision in all controlled comparisons. It also surpasses the post hoc best fixed-layer oracle in three of four comparisons. Parameter-matched ablations support both factorized target construction and image-conditioned depth allocation. Under comparable 200-way retrieval protocols, NeuroGlyph achieves the strongest system-level performance in six of eight reported metrics. These results support learning retrieval targets across the visual hierarchy rather than prescribing one visual depth.
Four-dimensional (4D) Radar has emerged as a key sensor for environmental perception, providing range, azimuth, elevation, and Doppler measurements while remaining robust to illumination changes and adverse weather conditions. However, conventional Radar preprocessing methods, such as constant false alarm rate (CFAR) detection, select measurements primarily based on signal-level criteria and may therefore discard information valuable for downstream perception during point cloud generation. In addition, existing 4D Radar perception pipelines typically optimize Radar data processing and downstream perception independently, preventing task objectives from directly guiding the preprocessing stage. To address these limitations, we propose a Scene- and Task-Aware Radar (STAR) Preprocessor together with an end-to-end training framework. The STAR Preprocessor incorporates scene context and downstream task objectives to generate task-relevant Radar points, enabling the Radar representation to be optimized directly for perception. On the K-Radar benchmark, the proposed method achieves 74.3 AP, outperforming the previous state of the art by 5.6 AP points. Furthermore, applying the task-relevant points generated by STAR to various existing 3D detectors improves detection performance in most evaluation settings and yields an overall positive average gain over point clouds produced by conventional preprocessing.
Reconstructing expressive and relightable 3D head avatars from monocular videos remains challenging in computer vision, as it requires accurate modeling of both non-rigid facial motion and illumination-dependent appearance. Existing Gaussian avatar methods commonly rely on globally coupled representations, in which Gaussian primitives share a unified motion or illumination response model. Such uniform modeling neglects the distinct motion patterns and material/reflectance properties of different facial semantic regions, thereby limiting fine-grained animation accuracy and reducing relighting plausibility. To address this limitation, we propose SAMIRA, a 3D Gaussian avatar framework for semantic-adaptive motion-illumination response modeling. For motion response modeling, the Semantic-Adaptive Motion Response module rasterizes current-to-reference mesh displacements into a topology-consistent UV space and leverages facial semantics to route displacement features through semantic-specific modulators, predicting localized Gaussian geometric residuals beyond coarse mesh binding. For illumination response modeling, the Semantic-Adaptive Illumination Response module learns compact diffuse and specular response factors for each facial region, allowing Gaussians in different regions to adapt their illumination responses to novel environment lighting. These response factors are incorporated into deferred physically based shading, providing a lightweight approximation of semantic-dependent illumination effects. Extensive experiments on self-reenactment, cross-reenactment, and relighting demonstrate that SAMIRA improves both fine-grained expression reconstruction and relighting realism over existing methods.
Embodied AI systems are often organized into System 1 and System 2. System 1 is typically a pretrained policy that generates actions at high frequency, whereas System 2 is often instantiated as a vision-enabled language model for high-level planning. We ask whether a large language model (LLM) can act as the policy for robot manipulation without task-specific finetuning. We call this setting LLM as policy. We evaluate three LLMs on all 42 RoboDojo tasks and compare their scores with 40 public policies. Astra and GPT-5.5 use the official 50-episode-per-task protocol; DeepSeek-Flash uses 10 episodes per task. GPT-6 Astra achieves 22.48% average success rate and 28.97 Score over 2,100 trials, ranking above every public entry. Yet GPT-5.5 and DeepSeek-Flash reach only 0.88% and 1.92% average success rate with the same post-processing. We find that Astra exhibits a sharply polarized capability profile. It generalizes well to tasks that require semantic understanding but not high-precision control. In contrast, it performs poorly on tasks that require precision, dynamic control, or complex bimanual coordination. In-context experiments show no aggregate benefit from one-shot demonstrations, while selected interaction traces show within-episode corrections under perturbations. Overall, the evaluated LLMs vary substantially in manipulation performance. Astra stands out and provides initial evidence for the potential of a general-purpose manipulation model, although reliable precision and dynamic control remain limitations in the evaluated setting.
Chinese Translation
具身人工智能系统通常被组织为系统1(System 1)和系统2(System 2)。系统1通常是一个以高频率生成动作的预训练策略,而系统2通常被实现为具备视觉能力的语言模型,用于高层规划。我们提出了这样一个问题:大型语言模型(LLM)能否在无需任务特定微调的情况下充当机器人操作(manipulation)的策略?我们将这一设置称为“LLM即策略”(LLM as policy)。我们在全部42个RoboDojo任务上评估了三个LLM,并将其得分与40个公开策略进行比较。其中,Astra和GPT-5.5采用官方的每任务50次试验(episode)协议,DeepSeek-Flash则采用每任务10次试验。GPT-6 Astra在2,100次试验中取得了22.48%的平均成功率和28.97的Score,排名超过所有公开的参赛方法。然而,在相同的后处理条件下,GPT-5.5和DeepSeek-Flash的平均成功率分别仅为0.88%和1.92%。我们发现Astra呈现出高度两极化的能力特征:它能够很好地泛化到那些需要语义理解但不需要高精度控制的任务;相比之下,它在需要精度、动态控制或复杂双手协调的任务上表现较差。上下文(in-context)实验表明,单次示范并未带来总体上的收益,而选取的交互轨迹则显示其在扰动下具有回合内的纠错行为。总体而言,被评估的各个LLM在操作性能上差异显著。Astra表现突出,为通用操作模型的潜力提供了初步证据,尽管在被评估的设置下,可靠的精度和动态控制仍是其局限所在。
Legends are fundamental to chart understanding, as reliable interpretation requires correctly binding legend entries to corresponding visual marks. While vision-language models (VLMs) are increasingly applied to chart understanding, their legend understanding is poorly diagnosed by aggregate accuracy, which can be satisfied by superficial shortcuts and confound legend-specific errors with other reasoning failures. To enable fine-grained diagnosis and controlled testing, we introduce LegendBench, a parametric benchmark and generation pipeline that produces targeted legend-centric test cases. LegendBench contributes (1) a capability-task taxonomy spanning legend parsing, legend grounding, legend-conditioned reasoning, and legend-aware abstention to localize failures, and (2) counterfactual group generation, where each base chart yields multiple variants under controlled legend interventions to probe model invariance and sensitivity. Using LegendBench, we evaluate both general-purpose VLMs and specialized chart models and generate their capability profiles, revealing persistent bottlenecks in reliable legend-to-mark binding and counterfactual consistency. We then use these capability profiles to guide targeted fine-tuning, demonstrating that bottleneck-specific interventions can effectively close the localized capability gaps and generalize to unseen data. We further leverage our counterfactual design to conduct fine-grained diagnostic experiments, analyzing encoding-channel effects, legend-order shortcuts, and abstention under varying visibility.
Benchmarking Off-the-Shelf Multimodal AI Models Against Dermatologists on Patient-Captured Skin Images
现成多模态人工智能模型与皮肤科医生在患者自摄皮肤图像上的基准对比
Dolphin, Rian, Knowles, Laura
Abstract
Artificial intelligence (AI) has advanced at a rapid pace in recent years. Initially, breakthroughs in large language models caught widespread attention. However, recent generations of frontier AI models have adopted multimodal capabilities as a first class citizen, with vision capabilities being central to that. In this paper, we evaluate three recently released models on the task of diagnosing dermatological conditions from patient-submitted images. The models chosen are at the low to mid tier in terms of pricing and thus represent a floor on current AI capabilities, not a ceiling. We evaluate AI performance relative to a panel of three certified dermatologists, who grade each image, and we present four interesting findings. Firstly, depending on the metric, the tested AI models are either on par or slightly trail humans in terms of inter-clinician agreement. Secondly, we find that asking AI models for a confidence rating produces poorly calibrated answers, meaning use of confidence thresholds should not be relied upon in a clinical setting. Thirdly, the effect of providing additional patient metadata is strongly model-specific, with one of the three models degrading on every metric considered. Finally, model cost is not predictive of performance. The best-performing model we tested costs on average $0.0045 per case.
Pedestrian head orientation recognition plays an important role in autonomous driving by providing valuable cues for understanding pedestrian attention and anticipating potential crossing behavior. However, reliable recognition in real-world traffic scenes remains challenging because pedestrian head regions are often captured at low resolution. To address this challenge, we propose a lightweight Low-Resolution Head Orientation Convolutional Neural Network (LRHO-CNN) for pedestrian head orientation recognition. We construct a new dataset by extracting pedestrian head images from multiple public datasets and manually annotating them into eight orientation categories. The collected images are systematically preprocessed and augmented to increase data diversity and better represent variations in illumination and image quality. The experimental analysis compares LRHO-CNN with three fine-tuned baseline models, namely ResNet-18, ResNet-34, and VGG-16. The results demonstrate that LRHO-CNN achieves the highest classification accuracy among the evaluated models. LRHO-CNN is further evaluated on the JAAD and PIE datasets, demonstrating its effectiveness in recognizing pedestrian head orientation in real-world traffic scenes and providing informative head-orientation cues that can support downstream pedestrian behavior and intention prediction.
SAFe: Segment-guided Aggregation of Feature Densities for Anomaly-aware Segmentation
SAFe:面向异常感知分割的分段引导特征密度聚合方法
Delić, Anja, Runtas, Jurica, Oršić, Marin, Marković, Ivan, Petrović, Ivan
Abstract
Visual segmentation systems encounter objects outside their training distribution during real-world deployment, hindering reliable autonomous systems that depend on scene parsing in the perception stage. Many recent methods address this by using self-supervised foundation models to train density estimators that yield low likelihood in anomalous image regions. Although promising, these methods suffer from poor feature semantics or they lack spatial consistency, both of which undermine critical downstream decisions. We address this problem with~\method, a generative method based on class-conditional density estimation over self-supervised representations. SAFe trains lightweight normalizing flows that produce class-conditional normalized likelihood estimates over frozen DINOv3 features. We combine density estimates from transformer features with density scores over multi-scale convolutional features to capture both global semantics and local detail. We introduce a method-agnostic post-processing step based on SAM3 that connects per-location likelihoods into spatially coherent segments while suppressing false positives, and enables instance-level anomaly detection without retraining. The post processing further distinguishes novel categories among anomalous objects by a similarity-based agglomerative clustering scheme. SAFe sets a new state of the art on the PANIC, OoDIS, SMIYC ObstacleTrack with strong performance on the ISSU benchmark.
CoaG: Cylinders on a Grid: Coarse 3D Layout Control for Video Generation
CoaG:网格上的圆柱体:面向视频生成的粗略3D布局控制
Yang, Zhangsihao, Shan, Mengyi
Abstract
We ask how little geometry a person has to draw to control both where people stand and where the camera moves in a generated video. Our answer is a ground plane and one cylinder per person. A user draws a grid on the ground, places one cylinder where each person should stand, moves the cylinders and the camera over 81 frames, and the model renders a photoreal video in which the people occupy the cylinders' positions, move as the cylinders move, and are seen from the drawn camera. Appearance comes from a text prompt and a background reference image; layout and motion come from the geometry. Because no dataset pairs such a signal with video, we build the pairs ourselves: an automatic engine writes 2000 captions from a combinatorial seed, generates a clip for each with a text-to-video model, and lifts every clip back to its geometry with person tracking, background inpainting, an agentic ground-mask loop, feed-forward multi-view reconstruction and a plane fit, with no real footage and no manual labels. A LoRA on Wan2.2-Fun-Control trained on 1935 such tuples follows drawn layouts and camera paths on hold-out clips: the generated people match the cylinders' count, order, position and height, the text changes who they are, the reference image changes where they are, and dolly-in, orbit, pan and crane paths are followed, dolly-out only weakly.
ChartJudgeBench: Evaluating LMM Judges for Chart-to-Code Generation
Wu, Lijian, Zhao, Henry Hengyuan, Zhang, Zijian, Tang, Jiahao, Wu, Jiajun, Wang, Alex Jinpeng
Abstract
Building strong chart-to-code systems increasingly relies on reinforcement learning, whose effectiveness depends critically on the quality of the reward signal. Large Multimodal Models (LMMs) play a natural critical role in jointly assessing chart visual appearance and task requirements. They are therefore increasingly used as visual critics and reward models, yet their reliability as judges remains largely unexplored. To this end, we introduce ChartJudgeBench, a diagnostic vision-language benchmark for assessing LMM judges in chart-to-code workflows. It includes 1,003 Chart Perception Alignment (CPA) instances for pairwise chart comparison and 650 Chart Reasoning Judgment (CRJ) instances for binary Accept/Reject verification in Chart Reproduction and Chart Editing. Together, these tasks emulate the core judging decisions required in agentic refinement and RL-based chart optimization. Our evaluation of strong LMMs reveals four systematic limitations: (i) positional bias in pairwise comparison, (ii) a strong tendency to overpredict Accept, (iii) difficulty in matching visual styles and aesthetics, and (iv) an unexpected leniency bias in RL-trained models. These findings show that current LMM judges require explicit reliability validation before being used as critics or reward models in chart-to-code optimization. The code and data are available on ChartJudgeBench.
Beyond Emotion Prompts: Fine-Grained Text-to-Image Generation Driven by Valence-Arousal-Dominance
超越情感提示词:由效价-唤醒度-支配度驱动的细粒度文生图生成
Li, Minglang, Fang, Yueyue, Gao, Xieping
Abstract
Although text-to-image models can accurately depict subjects and scenes, creators still struggle to specify the fine-grained emotions an image should convey without rewriting its content description. Natural language can suggest emotions, but it offers no control scale with stable meanings and ordered intensities. We propose EMOTRANS, which transforms psychologically grounded valence-arousal-dominance (VAD) coordinates into generation conditions that are independent of the content text and modulated across denoising stages, making emotional style a finely adjustable creative variable. To support this goal, we construct EMOVAD, an art-painting dataset that pairs objective content descriptions with separately collected emotional ratings from multiple annotators. We also coordinate emotional expression and content preservation through dual-branch training with a shared model. Objective and human evaluations show that the framework improves the accuracy of three-dimensional emotion control and produces perceptible, orderable continuous changes while maintaining competitive text alignment and image quality. This work provides a practical emotion-driven approach to image generation that extends objective content depiction to fine-grained emotional adjustment.
Document Retrieval-Aware Chunking (D-RAC): Universal Retrieval-Aware Ingestion of Enterprise Documents via PDF Normalization and Multimodal Markdown Conversion
Allu, Uday, Sivaprakash, Abhivanth, Singh, Pratik, Manocha, Aman
Abstract
Retrieval-Augmented Generation (RAG) systems over enterprise knowledge bases must ingest heterogeneous document formats -- PDFs, Word documents, presentations, and scans -- whose content is locked inside complex visual layouts, multi-column pages, and dense tables. Rule-based extraction and OCR destroy reading order, flatten tables, and lose heading hierarchy, while fully agentic chunking over extracted text incurs high token costs and hallucination risk. We present Document Retrieval-Aware Chunking (D-RAC), an extension of our Web Retrieval-Aware Chunking (W-RAC) framework to arbitrary document formats. D-RAC first normalizes any input document into PDF, exploiting the fact that virtually every format has a faithful, deterministic PDF rendering. A single multimodal LLM pass then converts rendered pages into retrieval-optimized Markdown -- rewriting tables as self-contained prose statements and preserving heading hierarchy -- after which chunking proceeds exactly as in W-RAC: deterministic parsing into ID-addressable units followed by lightweight LLM-based chunk planning over identifiers rather than text. Source text is never regenerated during chunking, preserving W-RAC's cost, determinism, and observability benefits while unlocking every renderable format as a first-class input. On the 236-document, 795-page PDF subset of the RAG-Multi-Corpus benchmark spanning five enterprise domains, D-RAC converts and chunks the entire corpus in 72 minutes with zero errors, producing 1,748 retrieval-ready chunks. Compared to agentic chunking with frontier LLMs, D-RAC reduces chunking-stage output tokens by 95.7%, cutting chunking cost by 77.8% (GPT-4.1 pricing) to 85.6% (Gemini 2.5 Pro pricing) and chunking time by 75%. D-RAC scales linearly to documents of 500+ pages.
Dynamic Time Warping (DTW) is the dominant approach for measuring similarity between time series, yet standard practice discards the optimal warping path after computing a single distance value, losing local alignment information most relevant to clinical diagnosis. We introduce DiaSeg, a framework that extracts diagonal segments from DTW paths with controlled breaks, characterizing each segment by five geometric features (effective length, interruption count, cost variation, temporal position, and path context), and enabling unsupervised pattern discovery without domain-specific feature engineering. Validated on 91 subjects across six clinical conditions (healthy aging, Parkinson's, Huntington's, ALS, brain tumor, and stroke), three findings emerge. First, diagonal segments form consistent unsupervised patterns (silhouette 0.33) aligned with biomechanical phase annotations, with label-based validation confirming near-perfect separation of healthy and pathological gait (ARI up to 0.986). Second, segments discriminate pathology at 69% (supervised) and 75% (patient-level clustering), with pathology manifesting through distributional shifts in segment length; combining segment and cycle-level features further improves classification to 91.7%. Third, while cycle-based methods achieve higher accuracy (91%), diagonal segments provide phase-specific interpretability unavailable in global representations, localizing where coordination breaks down within the gait cycle. DiaSeg thus transforms DTW from a black-box distance into a source of interpretable temporal features for neurodegenerative disease assessment.
SRPR-Net: Semantic and Relational Prompt Refinement for Automated SAM-based Instance Segmentation
SRPR-Net:面向自动化SAM实例分割的语义与关系提示精炼方法
Liu, Lufei, Li, Guojie, Xiang, Suncheng, Zhang, Fan
Abstract
Instance segmentation is a fundamental computer vision task with diverse real-world applications. Recently, prompt-driven foundation models have shown promising generalization. However, automated prompting remains limited by insufficient semantic guidance and inter-instance modeling. To address this challenge, we propose a novel architecture, named Semantic Relational Prompt Refinement Network (SRPR-Net), for automated SAM-based instance segmentation. A sequential prompt refinement mechanism is introduced to enrich detector geometry with visual-language semantics and then incorporate same-image instance dependencies, enabling context-aware box adjustment before SAM segmentation. Experiments on multiple standard benchmarks demonstrate that SRPR-Net achieves consistent improvements in segmentation performance over existing state-of-the-art approaches. The code is publicly available at https://github.com/JeremyXSC/SRPR-Net.
IMPLICIT-Bench: Measuring Implicit Bias in Text-to-Image Models under Neutral Prompts
Dai, Yue, Liu, Ziyang, Cheong, Marc, Han, Caren
Abstract
Text-to-image (T2I) models are typically evaluated for bias using slot-based templates such as ``a photo of a [profession]''. Such templates probe only \emph{explicit} demographic attributes (e.g., gender, skin tone) in isolation. They overlook a broader \emph{implicit} bias that arises in natural prompts: when stereotype-relevant attributes are left unspecified, models still default to stereotypical outputs. We introduce IMPLICIT-Bench, a benchmark for measuring implicit bias in T2I models under such prompts. The key design is a structured-knowledge-graph (KG) construction of controlled prompt triplets: neutral, stereotype, and anti-stereotype variants that differ only along a single bias dimension while preserving scene semantics. This enables precise attribution of bias effects that template benchmarks cannot achieve. IMPLICIT-Bench comprises 5,493 prompts across 11 bias categories, validated through multi-model agreement, CLIP-based verification, and human evaluation. Using this benchmark, we show that state-of-the-art T2I models exhibit systematic bias under neutral prompts, a failure mode largely invisible to existing evaluations. We then use IMPLICIT-Bench to evaluate debiasing methods, uncovering a fundamental trade-off between bias reduction and semantic fidelity.
Multimodal large language models (MLLMs) fail at fine-grained visual questions less because they cannot reason than because they never see the evidence: high-resolution images are downsampled before encoding, so the model answers from linguistic priors. The standard remedies are expensive: annotated answers (SFT), hand-engineered verifiers (RLVR), or a large external teacher (on-policy distillation). We ask whether the visual evidence itself can supply the signal for free. We formalize the contrastive evidence gap, the per-token log-likelihood ratio that a model assigns to its own output when conditioned on a question-relevant region versus an irrelevant one, and study it across Qwen2.5-VL-7B, Qwen3-VL-8B, and Qwen3-VL-30B-A3B on V*Bench. Our main positive result is training-free: selecting the candidate crop under which the model's answer distribution is most peaked, using a single-view, label-free criterion, discovers the answer-bearing region with no bounding boxes, training, or labels. It localizes the target 4.4 to 5.1 times better than chance and raises fine-grained accuracy from 70 percent to 85 percent at inference. We further show that the gap is complementary to the model's own confidence. Combining them predicts correctness better than either alone, with AUC up to 0.99, and flags confidently wrong answers, with AUC ranging from 0.97 to 1.00 within the high-confidence subset. All effects concentrate on perception-bottleneck questions and vanish on a global-context control. Finally, we report an honest negative result: converting the same signal into a training method, gated self-distillation (SEG-Distill), does not outperform the base model at pilot scale across three gate designs, while more aggressive gating degrades accuracy. The signal is real, but converting it into training gains remains an open problem.
Hyperbolic vision-language models (VLMs) represent image and text features in a geometry naturally suited to hierarchy, but their adaptation to downstream tasks has largely relied on fixed prompts. Existing prompt learning methods, meanwhile, treat class labels as a flat set and do not exploit available taxonomic structure. We address this gap with a hierarchical prompt learning plug-in for frozen hyperbolic VLMs. Given a fixed offline parent-class hierarchy, it augments a class prompt learner with a separate parent prompt learner, parent-level supervision, hyperbolic entailment regularization, and parent-feedback logit fusion. We instantiate the method with CoOp, CoCoOp and MaPLe, yielding HyPLO, CoHyPLO and MaHyPLO. Across the standard 11-dataset benchmark, all variants improve base-to-new generalization and cross-dataset transfer, and remain comparable to their prompt learning baselines under domain shift. Six hierarchical metrics and embedding analyses show that the method produces more taxonomically consistent predictions and induces a hierarchy-consistent organization of parent, class, and image embeddings in hyperbolic space. Its gains are largest when novel classes must be placed within a fixed taxonomy, and smallest for fine-grained confusions among sibling classes or shifts affecting only the image distribution.
Classifier-Free Guidance in Flow Matching: Non-Autonomous Potentials, Overshoot, and Posterior-Mean Control
流匹配中的无分类器引导:非自治势场、过冲与后验均值控制
Peng, Jishen, Ma, Zheng
Abstract
Classifier-free guidance (CFG) improves conditional generation in Flow Matching, but strong guidance can distort the generated distribution and reduce diversity. We provide a geometric account of this behavior by viewing Flow Matching as a time-varying gradient flow and characterizing how CFG reshapes its underlying potential. This view explains how stronger alignment can be accompanied by mean displacement and trajectory concentration, and motivates controlling guidance through the model-implied terminal posterior mean. We therefore propose Posterior-Mean-Capped CFG (PMC-CFG), a training-free, per-sample method that adaptively retains the strongest feasible guidance without additional network evaluations. Experiments on synthetic and large-scale image-generation benchmarks show that PMC-CFG limits guidance-induced distortion and concentration while improving the alignment--diversity trade-off, with particularly strong benefits when nominal guidance is large.
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.
AnalogDepth: Multi-view Geometry from FPV drones under Analog Video Transmission
AnalogDepth:基于模拟视频传输的FPV无人机多视角几何
Amorim, André, Proença, Pedro F.
Abstract
Analog video transmission (VTX) remains widespread in FPV drones due to low latency, weight and low cost. However analog VTX suffers from complex spatially structured image degradation which differ fundamentally from digital image corruption (e.g. AWGN) used in standard training augmentation. This work shows that this type of noise severely degrades the accuracy of Depth Anything 3 (DA3), a state-of-the-art feed forward visual geometry foundation model. To address this gap, we present AnalogDepth, a parameter-efficient training pipeline that adapts DA3 to analog FPV imagery using student-teacher knowledge distillation with Low-Rank Adaptation (LoRA) injected into the DINOv2 backbone. Rather than synthesizing noise analytically, we build a noise bank from static FPV recordings under diverse conditions and compare real-noise injection against PSD-matched Gaussian synthesis and AWGN as baselines. Experiments on six real FPV flight sequences across three indoor scenes show that training with our noise bank consistently reduces per-frame depth RMSE and 3D reconstruction Chamfer distance compared to the pretrained DA3 baseline and both Gaussian noise variants. These results demonstrate that replicating the spatial structure of real analog transmission noise is critical for effective adaptation.
NeuIDO: Neural Intrinsic Dynamics Operator for Physics-Informed 4D World Models
Lin, Jiajing, Zhang, Xin, Sun, Jianhua
Abstract
World models aim to capture environmental dynamics and predict future trajectories, showing growing potential for embodied intelligence. Physics-informed 4D generation integrates physical simulation to predict 3D object interactions, offering a promising pathway toward world models. However, this paradigm relies on manually imposed dynamical assumptions rather than internalizing world dynamics, and thus still leaves a gap toward a true world model. To bridge this gap, we propose NeuIDO, a novel world dynamics modeling framework that learns a unified intrinsic dynamics representation from visual observations, advancing physics-informed 4D generation toward a world model. Specifically, we formulate world modeling as a neural operator learning problem and introduce a two-stage training strategy to learn a generalizable mapping from the visual observation distribution to the intrinsic dynamics distribution. Building on this observation-dynamics mapping, NeuIDO enables zero-shot dynamics inference directly from videos and can be further aligned with complex real-world dynamics via few-shot adaptation. Extensive experiments demonstrate that NeuIDO effectively unifies the intrinsic dynamics underlying diverse visual observations into a shared representation and rapidly infers dynamics in novel scenes.
Image morphing aims to produce a smooth and semantically consistent transition between two input images. Existing diffusion-based morphing methods either require expensive per-pair optimization or rely on implicit spatial alignment, which easily fails under large layout discrepancies. To address these limitations, we propose AlignMorph, a novel tuning-free diffusion framework guided by the principle of transport-then-denoise. We explicitly decouple geometric alignment from generative denoising to avoid structural entanglement. Our framework consists of two core components. (1) Global Semantic Transport, which achieves diffusion-compatible semantic alignment via entropic optimal transport and reliability-aware latent warping; and (2) Coordinate-Aligned Generation, which uses a symmetric bi-phase attention handoff to maintain consistent spatial coordinates throughout denoising. Without any tuning, AlignMorph effectively eliminates ghosting and achieves superior structural coherence and temporal smoothness on morphing benchmarks. Code is available at https://github.com/51xOne/Alignmorph.
While Mamba-based models have shown strong potential for long sequence modeling, adapting them to vision is challenging due to the requirement of local neighborhood correlations and multi-directional spatial contexts for visual understanding. In this paper, we present LiAuto-MindViT, a novel hybrid vision backbone that synergizes the strengths of CNNs, Mamba, and Transformers. The core of our design is the Adaptive Bidirectional Mamba (ABM), which eliminates the directional bias of unidirectional SSMs through bidirectional selective scanning with learnable alpha blending, enabling content-adaptive directional fusion without the overhead of exhaustive multi-path routing. To further accelerate inference, we propose a deployment-friendly Reparameterized ConvSE (RepConvSE) module that leverages structural reparameterization to reduce latency and memory access overhead. Extensive experiments demonstrate that LiAuto-MindViT achieves state-of-the-art performance on image classification, object detection, and semantic segmentation while enabling efficient inference through reparameterization.
Image forensics is increasingly an open-world problem: manipulations range from fully synthetic images to localized edits, splicing and swapping, while most forensic detectors remain specialized to a single manipulation family. Agentic AI has recently emerged as a promising solution. In principle, such systems can assess the reliability of individual detectors, identify out-of-scope evidence, and arbitrate conflicting reports. However, it remains unclear which components actually drive performance and whether their benefits persist under distribution shift. To answer these questions, we study a training-free agentic framework built around specialist detectors, per-detector triage, and conflict-aware evidence arbitration. Using six configurations and three multimodal large language model backbones, we dissect the role of triage, prompting, and reasoning quality on both in-distribution and out-of-distribution data. Our results show that naive detector fusion suffers from severe false-positive rates on authentic images. Triage and prompting consistently improve performance by filtering unreliable evidence and exposing detector limitations. However, the dominant factor is represented by reasoning itself: A stronger judge substantially outperforms a weaker one, particularly under distribution shift. Most notably, manipulation recall is nearly saturated across all configurations, indicating that the main challenge of open-world image forensics is not detecting manipulations, but calibrating trust in specialized forensic tools and arbitrating conflicting evidence.
Fully supervised repetitive action counting (RAC) has achieved strong performance, but requires dense temporal annotations that are costly and difficult to scale. We propose TReViS, a self-supervised video synthesis framework that enables training RAC models without any repetition labels. TReViS estimates the underlying temporal repetition structure of an unlabeled video via a Temporal Self-Similarity Matrix, infers its cycle statistics, and synthesizes new training sequences that preserve realistic repetition patterns while introducing controlled temporal variability. These synthesized videos are paired with pseudo-labels and used to train existing RAC architectures from scratch. Across multiple datasets and backbones, TReViS consistently outperforms prior self-supervised methods and achieves performance competitive with several supervised baselines, while remaining fully label-free, demonstrating the effectiveness of structure-aware video synthesis for label-free RAC. The source code is available at https://github.com/yfqi/TReViS.
Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entangle many concepts in each neuron. Feature superposition is widely treated as the central obstacle to this decomposition, yet most mitigations (sparse autoencoders, dictionary learning) are post-hoc and leave the underlying network unchanged. We ask whether a spatial-locality training loss (TopoLoss) can act as a lightweight, training-time prior that improves interpretability of standard mech-interp tools. Training ViT on ImageNet-100 across multiple TopoLoss weights $\alpha$, we measure causal sufficiency of topographic clusters via activation patching and feature geometry via sparse autoencoders fit to the same residual stream. At $\alpha=1.0$, topographic clusters are 2.79$\times$ more causally sufficient than random unit sets of the same size, with the effect increasing monotonically in $\alpha$. SAE L0 sparsity decreases by 11% and dead-feature fraction rises 19-fold, yet standard neuron-level monosemanticity scores are unchanged, indicating that topographic pressure acts at circuit level, concentrating causal mass into spatially local structures without disentangling individual neurons. This dissociation suggests current neuron-level monosemanticity metrics are insensitive to a class of real interpretability gains, and positions cheap architectural priors as a viable training-time complement to post-hoc tooling.
A Lightweight Convolutional Neural Network for Real-Time Recognition of Hand-Drawn Geometric Shapes
Abdullah, Shahir
Abstract
Recognizing hand-drawn geometric shapes is a foundational sub-problem of sketch recognition, with applications in education, human-computer interaction, and diagram digitization. This paper presents the design, implementation, and evaluation of a desktop application that recognizes four basic hand-drawn geometric shapes, circle, square, rectangle, and triangle using a compact Convolutional Neural Network (CNN). A dataset of 2,000 labeled 28x28-pixel shape images was collected independently and released publicly. The classifier consists of three convolutional blocks (16, 32, and 64 filters) with max-pooling, an in-model data-augmentation stage (random horizontal flip, rotation, and zoom), a dropout-regularized dense layer of 128 units, and a 4-way linear output layer, totaling 97{,}956 trainable parameters. The network is trained with the Adam optimizer on a sparse categorical cross-entropy objective computed directly on logits. On an 80/20 train-validation split, the model achieves 94.80% training accuracy and 96.01% validation accuracy with a validation loss of 0.1437. A Tkinter-based graphical interface allows a user to draw a shape with the mouse and receive an immediate class prediction with a confidence score. We situate this system within the broader sketch and shape-recognition literature, compare its accuracy against related hand-drawn shape classification studies, and discuss the limitations inherent to a small, single-contributor dataset. The complete source code, trained model, and per-class datasets are released publicly to support reproducibility.
Biological visual systems achieve continuous, low-latency motion perception by processing sparse, asynchronous spiking signals, enabling real-time tracking under strict energy constraints. Event-based cameras, inspired by the mammalian retina, replicate this efficiency by capturing only local brightness changes as asynchronous events, offering a natural substrate for spiking neural networks (SNNs) to parallelise computation and adapt to fast-changing scenes. Pinball provides a controlled yet dynamic testbed, requiring precise motion estimation and fast reaction to a small, rapidly moving target. This work presents a fully spiking, real-time perception-to-action pipeline for closed-loop pinball gameplay. A dynamic vision sensor observes a small, fast-moving ball, and a network of spiking Time-Difference Encoders on the SpiNNaker neuromorphic platform jointly estimates its position, speed, and direction. The system is characterised across receptive field size, accumulation window, and angular tuning width for real-time operation, and benchmarked in closed loop against human players across two flipper regimes of increasing physical realism. It achieves a hit rate of 56.1%, nearly double the human average, reacting within 21.7 ms (5 ms network latency) and consuming an estimated 148 {\mu}W using fewer than 25k neurons, among the fastest and most energy-efficient event-based closed-loop demonstrators benchmarked. Under more realistic flipper dynamics, tuning a single interpretable policy parameter reproduces the full spectrum of human play styles, from cautious to aggressive, with no change to the perception pipeline. A physical demonstrator, tracking a real ball and actuating real flippers in closed loop, confirms the principle operates beyond simulation. Its fully spiking, learning-free design offers a compact, energy-efficient example of real-time neuromorphic perception-to-action.
Multi-task visual grounding requires models to jointly understand linguistic semantics and perform accurate visual localization and segmentation. Despite the success of multimodal large language models, effectively adapting them to multiple grounding objectives remains challenging. Existing methods commonly enforce task cooperation through shared representations, while overlooking the intrinsic conflict between task-oriented feature interests. In this paper, we introduce $\textbf{DeCo}$, an efficient $\textbf{De}$couple-to-$\textbf{Co}$uple learning framework that resolves this dilemma through a two-stage paradigm: task-specific representation decoupling followed by complementary prior coupling. Specifically, we first propose Task-aware Semantic Decoupling (TSD) to route shared visual cues into individual features under salient word-level guidance, alleviating representation interference between localization and segmentation. Furthermore, we observe that segmentation naturally provides informative localization priors due to dense supervision. Based on this insight, we introduce Hybrid Prior Coupling (HPC), which integrates sentence-level semantic prior with mask-derived spatial prior for enhanced grounding. Built upon a frozen multimodal encoder, DeCo requires lightweight trainable parameters while achieving strong generalization across multiple grounding objectives. Extensive experiments on RefCOCO/+, G-Ref, ReferIt, Flickr, DIOR-RSVG, SARVG1.0, RRSIS-D, RIS-LAD, and RefDIOR demonstrate that DeCo achieves state-of-the-art performance on both natural and remote sensing benchmarks. The code and models are available at https://github.com/xiaoqiang-lu/DeCo.
Estimating Accurate Hand Pose in Camera Space with Vision Transformer
基于视觉Transformer的相机空间精确手部姿态估计
Ren, Kaiwen, Jiang, Yiran, Ye, Yongjing, Xia, Shihong
Abstract
Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. However, this camera-space estimation confronts two fundamental challenges: (1) depth ambiguity in monocular settings, and (2) the coupling effect of hand local poses and global wrist positions in the perspective projections. In particular, this coupling reflects that the projections are jointly determined by local hand poses, wrist positions, and camera intrinsics. To overcome these challenges, our framework proposes two key innovations: Transformation-Isomorphism Supervision for hand-depth information extraction and Perspective Information Embedding for resolving above coupling effect of local pose and wrist position, both integrated within the mainstream encoder-decoder architecture. Besides, we propose a novel framerate-aware multi-dataset training strategy for sequential pose refinement. Our fully integrated approach achieves at most 37.1\% superiority in CS-MJE over SOTA on HO3D. Project page: https://github.com/Mine268/CS-ViT.
Chinese Translation
基于单目RGB的手部姿态估计已成为计算机视觉领域的重要研究前沿。局部手部姿态估计方法预测相对于手腕的手部姿态,而全局手部姿态估计还需要估计手腕在相机坐标系中的位置。然而,这种相机空间估计面临两个根本性挑战:(1)单目设置下的深度歧义性;(2)透视投影中手部局部姿态与全局手腕位置的耦合效应。具体而言,这种耦合表明投影是由局部手部姿态、手腕位置和相机内参共同决定的。为克服这些挑战,我们的框架提出了两项关键创新:用于手部深度信息提取的变换同构监督(Transformation-Isomorphism Supervision),以及用于解决上述局部姿态与手腕位置耦合效应的透视信息嵌入(Perspective Information Embedding),二者均集成于主流的编码器-解码器架构之中。此外,我们提出了一种新颖的帧率感知多数据集训练策略,用于序列姿态的精细化。我们的完整集成方法在HO3D数据集上的CS-MJE指标相比SOTA最高取得37.1%的优势。项目页面:https://github.com/Mine268/CS-ViT。
Recent 4D LiDAR language models aim to reason about objects and their evolving spatial relationships. Yet, in our evaluation, always selecting the same option nearly matches the multiple-choice accuracy of two B4DL-derived configurations. We introduce LiDAR-Hallu, a geometry-referenced benchmark and diagnostic protocol with 10,000 questions across 150 nuScenes scenes. It covers object existence, ego-relative position, distance ordering, relative motion, and temporal localization, with explicit rules for selecting objects, comparing times, and determining reference answers. Our protocol combines fixed-answer and candidate-content controls, cross-scene pairs with identical prompts but opposite reference answers, and relation-specific recall. Analysis of 100,000 recorded responses reveals failures hidden by aggregate accuracy. Candidate duration alone makes temporal answers predictable without observing LiDAR. On paired questions, the models frequently give the same answer to scenes requiring opposite answers. Relation-specific analysis further shows that both configurations miss every positive lateral-motion case across all tested conditions. Temporal-shuffle contrastive decoding provides little net improvement, as repairs are largely offset by new errors and the main failures persist. These results show that evaluating spatio-temporal reasoning requires testing whether models distinguish the queried physical relationships, rather than relying on individual-answer accuracy alone. The source code, checkpoints, and data are released at https://github.com/Awesome4D/4DMLLM_Hallucination_Bench.
Intelligent transportation systems require Incremental Learning (IL) to continually improve their overall performance in dynamic environments. However, most edge devices lack the computational resources to support on-device IL, requiring updates to be transmitted from centralized servers. We propose using this setup to obtain dense, specialized module coverage that adapts a fixed base model to specific spatiotemporal contexts, such as parking lots, gas stations, ferries, or construction sites. However, in order to reliably transmit these modules to the edge device, using TCP, UDP, and BTP over V2X, Wi-Fi, and 2G-5G hardware, we establish a strict limit of 14.6 KB per module to fit within the first TCP window and to minimize UDP/BTP fragmentation. We further introduce Mixture-of-Experts for Communication-Aware Incremental Learning (MECAIL), the first method that meets this strict requirement, in which each new domain or environment is served by a small expert network that adapts the base model. We validate MECAIL on D-RICO and ODinW-13, where it largely matches the performance of parameter-heavy approaches while enabling practical, bandwidth-efficient large-scale deployment. This allows comprehensive coverage by experts for highly specific, focused, and temporary situations.
Chinese Translation
智能交通系统需要增量学习(Incremental Learning, IL)以在动态环境中持续提升整体性能。然而,大多数边缘设备缺乏支持端侧增量学习的计算资源,需要从集中式服务器传输更新。我们提出利用这一设置来获得密集且专门化的模块覆盖,使固定的基础模型能够适应特定的时空场景,例如停车场、加油站、渡轮或建筑工地。然而,为了通过V2X、Wi-Fi以及2G-5G硬件,可靠地使用TCP、UDP和BTP将这些模块传输至边缘设备,我们将每个模块的大小严格限制在14.6 KB以内,以适配首个TCP窗口并最小化UDP/BTP的分片。我们进一步提出了面向通信感知增量学习的混合专家方法(Mixture-of-Experts for Communication-Aware Incremental Learning, MECAIL),这是首个满足这一严格限制的方法,其中每个新领域或新环境由一个小型专家网络提供服务,用以适配基础模型。我们在D-RICO和ODinW-13数据集上验证了MECAIL,其性能在很大程度上可媲美参数量庞大的方法,同时支持实用且带宽高效的大规模部署。这使得专家模块能够全面覆盖高度具体、集中且临时性的场景。
MIGA:Shared-Geometry Gaussian Representation with Implicit Amplitude Modeling for Accelerated 3D Multi-Echo MRI
Xu, Jingran, Liu, Yuanyuan, Zhu, Yanjie
Abstract
Three-dimensional multi-echo MRI provides rich anatomical and quantitative information, but repeated volumetric encoding prolongs acquisition and motivates k-space undersampling. Reconstructing undersampled multi-echo data requires exploiting shared anatomy while preserving echo-dependent signal variation; full-volume modeling also introduces substantial computational and memory demands. We propose MIGA, a scan-specific framework comprising shared anisotropic Gaussian geometry, a coordinate-conditioned multi-output amplitude network, and explicit echo-specific phase variables. The Gaussian geometry provides common spatial support across echoes, the implicit network models spatially structured amplitude variations, and the phase variables retain echo-specific complex signal information. All components are jointly optimized using only the acquired multi-coil k-space, requiring no fully sampled training data. Experiments showed that MIGA consistently outperformed the comparison methods across imaging tasks and acceleration factors, with larger improvements under stronger undersampling. MIGA also achieved a favorable quality-cost balance among the evaluated full-volume multi-echo methods. These results support the effectiveness of combining shared Gaussian geometry with implicit echo-dependent amplitude modeling for accelerated 3D multi-echo MRI reconstruction.
Spatial Action Review: A Visual Analytics Dashboard for Auditing Language-to-Action Hand-offs in Electron Microscopy
空间动作审查:用于审计电子显微镜中语言到动作交接的可视化分析仪表板
Mohinta, Samia, Cardona, Albert
Abstract
Multimodal large language models (MLLMs) are increasingly explored as interfaces for scientific image analysis, where a visual question-answering (VQA) response may be paired with a spatial output that guides a downstream stage. A supervisor reads the language answer, while a downstream workflow such as segmentation or region review consumes the point-set output. We call this transition from inspecting the answer to relying on its point action the language-to-action hand-off. A silent failure occurs when the answer is correct while the paired action misses annotated objects needed downstream, so answer-based oversight clears a region whose action is unreliable. We introduce Spatial Action Review, a visual analytics dashboard for auditing this failure mode in electron microscopy (EM) mitochondria analysis. It links paired answer-action records through an answer-action ledger, a task-by-dataset risk map, and an image-region audit view, connecting aggregate patterns to image evidence while an adjustable action-reliability gate supports re-audit. The review ends in a human-AI hand-off, where a supervisor records whether the action is accepted, escalated, held under a stricter gate, or flagged for model revision. Across 541 image regions from an EM-adapted Qwen3-VL case-study run, point actions fail the gate in 54.4% of records with a correct VQA response, and 27.4% of all records are silent failures. A correct answer is associated with only a 5.8-percentage-point higher probability of a reliable action, with a bootstrap interval spanning zero; the point-biserial correlation between answer correctness and object coverage is 0.061. This weak coupling persists across five model conditions on 753 matched image regions. Spatial Action Review makes answer-action mismatches visible and ties them to image evidence and a recorded decision before MLLM outputs enter autonomous scientific workflows.
Monocular 3D human pose estimation (HPE) remains challenging due to depth ambiguity, occlu- sions, and the need for temporal consistency. While multi-view methods provide superior accuracy over monocular approaches, they often require complex setups. We introduce STA-TFM, a transformer-based architecture that combines spatial and temporal information for multi-view pose estimation. The approach leverages DSTformer, a monocular feature extractor, to capture long-range pose dependencies within each view. A fusion transformer then aggregates information across views to produce coherent 3D estimates. To address training data scarcity, we use a data generation pipeline that transforms any existing 3D pose dataset into multi-view setups with controllable parameters. Experiments on various datasets demonstrate that STA-TFM outperforms existing camera-parameter-free multi-view methods. STA-TFM achieves 50.9% and 49.5% reductions in mean per joint position error (MPJPE) and mean per joint velocity error (MPJVE) on the DHP19 dataset. Furthermore, it achieves 6.7% and 7.7% respective reductions on HAA4D, and a 15.2% MPJPE reduction on TotalCapture. STA-TFM handles noisy and missing 2D inputs, supporting potential deployment in healthcare monitoring, athletic assessment, and immersive technologies. Code, training checkpoints, and data are available at https://zenodo.org/records/22832620.
Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We identify three key factors behind this degradation: text-guided selection bias, information loss from discarded tokens, and positional distortion caused by sequence compaction. Based on these observations, we propose \textbf{VPRune}, a training-free pre-LLM pruning framework consisting of visual-only diversity selection, similarity-guided token recycling, and position-preserving restoration. Experiments on FastVLM-1.5B across multiple vision-language benchmarks demonstrate that VPRune achieves a favorable accuracy--compression trade-off, with particularly pronounced advantages under aggressive compression. Furthermore, evaluations on edge-device show that VPRune effectively reduces end-to-end inference latency while maintaining superior task performance, demonstrating its practicality for resource-constrained LVLM deployment.
AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos
AgentSTAR:基于智能体的单目视频形状跟踪与重建
Mazur, Kirill, Karaev, Nikita, Chang, Matthew, Malik, Jitendra, Shafiullah, Nur Muhammad "Mahi''
Abstract
In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis. Unlike prior methods which first estimate dense pixel correspondences and then recover object motion from them, our method infers a structured 3D object model, including its geometry and kinematic structure, and uses this model to optimise object track estimates over time. In our optimisation loop, a Vision-Language Model (VLM) agent iteratively refines shape or generalised pose through a render-and-compare loop, combining coarse visual reasoning with numerical pose optimisation for precise state estimation. This structured formulation enables our method to track through large motion, articulation, and severe occlusion without relying on pixel-matching objectives. Quantitatively, on ARCTIC, our method substantially outperforms state-of-the-art 3D point-tracking baselines for articulated objects, and on HOT3D it outperforms all evaluated rigid-object tracking baselines.
Identity-Consistent Analysis of Long-Shot Windsurfing Video: A Domain-Specific Offline Tracking System
Braun, Bertil
Abstract
Long-shot windsurfing video combines small targets, large camera pans, prolonged overlaps, and rapidly changing backgrounds. The desired output is not a generic MOT trace but a separate, stable rider-relative video for each surfer; one false identity merge can invalidate an otherwise useful result. We present an offline analysis system that detects surfers, forms conservative local tracklets, links them globally with camera-compensated motion and a foreground-masked sail-color descriptor, and uses two pose keypoints on the rig to drive a rider-relative virtual camera. The tracking stage is evaluated on 21 manually reconstructed development videos containing 41,004 retained observations. On this fixed-observation protocol, the production system achieves 0.957 pairwise precision, 0.918 recall, and 0.937 F1, compared with 0.792 F1 for OC-SORT and 0.828 for BoT-SORT. Compared with OC-SORT, it reduces fragmentation excess from 845 to 42, but nine of its 95 output tracks mix rider identities and these errors affect seven of the 21 videos.
Accurate monocular depth estimation serves as a core enabler for single camera scene understanding. However, existing self-supervised monocular depth estimation methods generally suffer from the bottleneck of inefficient cross-scale information interaction and difficulty in balancing local and global spatial modeling. In this paper, we propose CMambaDepth, a self-supervised framework that achieves efficient multi-scale feature fusion and fine-grained contextual modeling via channel-wise selective state propagation. Specifically, Bidirectional Channel Mamba (Bi-CMamba) aligns encoder features across scales and enables bidirectional information exchange among ordered scale groups. Unidirectional Channel Mamba (Uni-CMamba) progressively aggregates decoder features and retains fine-grained scale groups through a group selection mechanism for subsequent fusion. Furthermore, a Hybrid Attention Module (HAM) is introduced to combine large-kernel local context and Manhattan self-attention for complementary spatial modeling. Experimental results demonstrate that our method achieves highly competitive performance. Specifically, our model achieves an AbsRel of 0.094 and an RMSE of 4.156 on KITTI, and an AbsRel of 0.140 on DDAD. In the zero-shot cross-dataset generalization test on NYUv2, it attains an AbsRel of 0.232, outperforming the baseline RA-Depth by 7.2%.
Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks. However, adapting VFMs to referring image segmentation (RIS) typically necessitates precise vision-language alignment via full fine-tuning, incurring substantial computational overhead and risking catastrophic forgetting. While existing parameter-efficient fine-tuning (PEFT) methods enable safe knowledge transfer with minimal training costs, they predominantly operate independently within individual modalities or focus exclusively on unidirectional guidance from language to vision, overlooking progressive cross-modal interaction and visual feedback for textual refinement. To address these limitations, we propose Bidirectional Reciprocal Learning (BRL), a novel adapter-based PEFT framework that facilitates hierarchical, bidirectional information flow within both token-mixing and channel-mixing layers of frozen foundation models. Specifically, BRL introduces two complementary lightweight modules. The Reciprocal Attention Adapter (RAA) performs cross-modal query-key exchanges at the token level, enabling visual and linguistic tokens to mutually attend to each other for fine-grained spatial grounding. The Reciprocal Gate Adapter (RGA) generates cross-modal gating signals at the channel level, allowing global semantic context from one modality to adaptively recalibrate channel activations of the other. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg benchmarks demonstrate the superiority of BRL over prior RIS methods, achieving state-of-the-art performance while requiring less than 0.5% backbone parameter updates. Code and models will be released at https://github.com/xiaoqiang-lu/BRL.
Preoperative Prediction of Microvascular Invasion in Hepatocellular Carcinoma by Integrating Multimodal Ultrasound and Clinical Data: A Multicenter Study
Background: Microvascular invasion (MVI) predicts recurrence and survival in hepatocellular carcinoma (HCC) but requires postoperative histopathology for diagnosis. We developed and validated a model integrating multimodal ultrasound and clinical data for preoperative MVI prediction. Methods: This multicenter study included 489 patients with HCC from eight centers. All patients had B-mode ultrasound (BUS), color Doppler flow imaging (CDFI), dynamic contrast-enhanced ultrasound (DCE-US), and clinical information. Data from seven centers (n = 421) were used for model development with five-fold cross-validation; data from the remaining center (n = 68) formed an independent external validation cohort. The proposed multimodal information fusion network used modality-specific encoders, a hemodynamic temporal change module for bidirectional DCE-US perfusion changes, and a representation consistency learning module to align heterogeneous ultrasound representations before Transformer-based fusion. Results: In external validation, DCE-US achieved the highest single-modality area under the receiver operating characteristic curve (AUC; 0.8545+/-0.0198), versus clinical information (0.6715+/-0.0156), CDFI (0.6435+/-0.0344), and BUS (0.6087+/-0.0417). Pixel-difference sampling and the proposed temporal module outperformed alternative sampling and video representation methods. The full model achieved the best performance, with an AUC of 0.8953+/-0.0180, accuracy of 81.18%+/-2.83%, sensitivity of 86.40%+/-6.69%, and specificity of 78.14%+/-6.28. Conclusions: Integrating multimodal ultrasound and clinical information enabled promising preoperative MVI prediction in HCC. DCE-US was the main source of predictive information, while BUS, CDFI, and clinical information provided complementary value. The proposed framework may support preoperative risk stratification and individualized clinical decision-making.
ME-VLM:A Unified VLM for Embodied Cognition and Agent Coordination
ME-VLM:面向具身认知与智能体协同的统一视觉语言模型
Model, Foundation, Inc, Li Auto
Abstract
Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and execution feedback. We introduce MachEmbodied-VLM (ME-VLM), a unified vision-language model with two variants, 4B and 35B-A3B, that brings together embodied cognition and multimodal agent capabilities. Our work emphasizes physical perception and spatiotemporal reasoning, together with planning, interaction, and outcome assessment in both digital and physical environments. We construct training data spanning embodied and multimodal agent tasks, including execution observations and feedback to support outcome assessment and decision refinement. The training pipeline comprises embodied capability injection, separate reinforcement learning of embodied and multimodal-agent experts, and multi-teacher on-policy distillation that consolidates their complementary capabilities into a single model. Experiments show competitive performance on both embodied and agent benchmarks, as well as on autonomous-driving and embodied-navigation tasks. For edge deployment, visual token compression, W4A8 quantization, and hardware--software co-optimization enable on-device inference of the 4B variant on the M100, reducing prefill latency from 400 ms to 188 ms. Project Page: https://machembodied.com/ME-Brain/ME-VLM.html Code Repository: https://github.com/MachEmbodied/ME-VLM
Thermography plays a vital role in military and broader thermal analysis applications. Recent progress in 3D thermal reconstruction has extended temperature analysis from 2D to 3D space, yet most existing works assume static temperature distributions, neglecting the temporal dynamics of heat transfer in real-world environments. To address this limitation, we propose the first dynamic RGB-Thermal reconstruction framework for complex scenes. Our method jointly models RGB appearance, thermal observations, and scene geometry as they change over time. Specifically, we introduce a multimodal dynamic scene representation that anchors both the color and thermal modalities to a shared geometric substrate, ensuring their consistency under spatiotemporal deformations. We further design multimodal embeddings to enhance the motion expressiveness for each modality, and propose a multimodal routing mechanism that retains a unified set of shared multimodal Gaussians as the geometric backbone while adaptively spawning modality-specific Gaussians to strengthen the representational capacity in detail-rich regions of each individual modality. In addition, we contribute a novel benchmark dataset featuring high-frequency temperature variations to facilitate the evaluation of 4D reconstruction. Extensive experiments demonstrate that our method achieves high-fidelity spatiotemporal reconstruction of both appearance and temperature. Our code and dataset are available at: https://github.com/LinLif1869/DTG.
MIRAGE: Full-Body Bystander Privacy for Smart Glasses with Consent-Based Restoration
MIRAGE:基于同意恢复机制的智能眼镜全身旁观者隐私保护
Umair, Muhammad, Maqbool, Muhammad Danial, Cheema, Fatima Arshad, Dev, Kapal, Alizai, Muhammad Hamad, Siddiqi, Muhammad Ali, Bhatti, Naveed Anwar
Abstract
Video recording on smart glasses exposes more than faces. Continuous capture reveals full-body biometric signatures, including gait, posture, and silhouette, that enable person re-identification (ReID) even after conventional face sanitization. We present MIRAGE, a three-tier architecture for privacy-preserving smart glasses that enforces full-body privacy, supports synthetic full-body replacement, and retains encrypted recovery material for consent-based restoration. We implement MIRAGE on a Raspberry Pi~5 (a CPU-only proxy for smart-glasses compute), companion phones, and a cloud generative backend. Compared to prior systems, MIRAGE achieves 0.948 AP and 0.976 AR while accurately detecting the complete visible body. Its bounding box masking reduces learned silhouette-based ReID to essentially random guessing, with 10.86% Rank-1 accuracy compared with an 11.12% measured chance level. Even against an adaptive adversary retrained on MIRAGE's sanitized pose signals, Rank-1 gait identification drops from 90.25% to 26.20%, removing 72.5% of the adversary's identification advantage.
Chinese Translation
智能眼镜上的视频录制所暴露的不仅仅是面部。持续拍摄会揭示包括步态、姿态和轮廓在内的全身生物特征信息,即使经过传统的人脸清洗处理,这些信息仍可实现行人重识别(ReID)。我们提出了MIRAGE,一种面向隐私保护智能眼镜的三层架构,能够执行全身隐私保护、支持合成的全身替换,并保留加密的恢复材料以实现基于同意的恢复。我们在Raspberry Pi 5(作为智能眼镜计算能力的纯CPU代理)、配套手机以及云端生成后端上实现了MIRAGE。与先前系统相比,MIRAGE在准确检测完整可见人体的同时,实现了0.948的AP和0.976的AR。其边界框掩码将基于轮廓学习的行人重识别能力降低至接近随机猜测水平,Rank-1准确率仅为10.86%,而实测的随机猜测水平为11.12%。即使面对在MIRAGE清洗后的姿态信号上重新训练的自适应攻击者,基于步态的Rank-1识别率也从90.25%降至26.20%,消除了攻击者72.5%的识别优势。
Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic representation, text descriptions have recently been incorporated as an additional modality. However, recent vision-language approaches often treat text descriptions as clean, deterministic signals and overlook their inherent noise, including modality-mismatched phrases and semantically ambiguous expressions. Moreover, prevailing methods lack explicit mechanisms to reconcile fine-grained structural discrepancies between modalities, even after high-level semantic alignment. To address these challenges, we propose a novel framework centered on Positive-Incentive Noise ({\pi}-noise) and structured prompt modulation. First, the Semantic Cross-Modal Modulator harnesses task-aware {\pi}-noise, sampled from a distribution conditioned on both visual and text inputs, to perturb global tokens and enable semantics-guided cross-modal compensation. Second, the Structure-Aware Prompt Adapter injects learnable geometric priors via prompts to enhance spatial consistency. Third, the Context-Aware Sparse Fusion module distills structural context to guide adaptive fusion while shielding identity features from noisy local details. Experiments on three multi-modal ReID benchmarks demonstrate the effectiveness and robustness of our approach. The code is available at https://github.com/zw-absin/INSPI.
Confocal Laser Endomicroscopy (CLE) provides real-time, cellular-resolution optical biopsy but has a narrow field of view, which image mosaicing can extend to provide anatomical context. Because of line-by-line acquisition, probe motion, and probe-tissue interaction, frame alignment generally requires a non-linear transformation whose accuracy is difficult to quantify: flexible transformation models can fit intensity features and noise, so appearance-based metrics such as Normalized Cross-Correlation (NCC) can improve without a genuine gain in geometric accuracy. We therefore establish a dataset of 132 frame pairs across fourteen pCLE sequences from 4 patients with manually annotated landmark correspondences, so that Target Registration Error (TRE) can serve as a geometrically grounded complement to NCC. We assess the effect of progressively increasing the transformation model's degrees of freedom, from translation to Thin Plate Spline (TPS), and of six feature-matching backends spanning classical (Shi-Tomasi, Lucas-Kanade) and learned (SuperPoint, SuperGlue, LightGlue, LoFTR, RoMa) approaches. Translation and rigid models prove insufficient under tissue deformation, while TPS with random sampling achieves the strongest landmark-derived alignment of the evaluated configurations; among the learned matchers, used without fine-tuning, only RoMa offers a robust, if modest, advantage over other methods. At the sequence level, pairwise registration quality proved an unreliable predictor of final mosaic quality, so mosaic quality must be evaluated directly rather than inferred from pairwise metrics.
CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encoder is known to preserve its generalization capability, recent studies show that fine-tuning both CLIP's text and image encoders jointly significantly enhances segmentation performance, especially for classes from open sets. In this work, we explain this phenomenon from the perspective of hierarchy alignment, since during fine-tuning, the hierarchical level of image embeddings shifts from image-level to pixel-level. We achieve this by leveraging hyperbolic space, which naturally encodes hierarchical structures. Our key observation is that, during fine-tuning, the hyperbolic radius of CLIP's text embeddings decreases, facilitating better alignment with the pixel-level granularity of visual data. Building on this, we propose HyperCLIP++, a novel and parameter-efficient adaptation strategy. HyperCLIP++ directly adjusts the hyperbolic radius of CLIP's embeddings via scaling transformations to achieve a hierarchy alignment to the target task, i.e., segmentation. To ensure this hierarchy alignment is effected consistently across both modalities and preserves their cross-modal alignment during training, HyperCLIP++ integrates a Dual Cross-Relation Communication (DCRC) module that synchronizes these adjustments between the vision and text pathways. Our experiments show that HyperCLIP++ achieves state-of-the-art performance across three benchmarks while fine-tuning only approximately 5% of CLIP's total parameters. More importantly, we observe that after adjustment, CLIP's text embeddings exhibit a relatively fixed hyperbolic radius across datasets, suggesting that the hierarchical level required for this segmentation task might be quantified using the hyperbolic radius.
Active Visual Sampling with a Connectome-Constrained Fly Model for One-Shot Hatch Recognition in Architectural Drawings
基于连接组约束果蝇模型的活动视觉采样用于建筑图纸中单样本填充图案识别
Kuklev, Dmitry
Abstract
Architectural drawings encode material classes through repeated hatch patterns. We test whether a connectome-constrained fly visual network, pretrained for motion, can be repurposed without task-specific weight updates as a descriptor for one-shot hatch matching. Each 64 x 64 patch is translated over eight scan trajectories and summarized across 57 cell types; query descriptors are then matched to one legend strip per class. On 400 development sheets from a synthetic benchmark built on CubiCasa5K geometry, the frozen fly pipeline reaches 0.857 area-weighted accuracy and 0.910 with an extended legend. On an equal-brightness orientation condition it reaches 0.840 versus 0.299 for eleven pixel statistics, while a Gabor bank reaches 0.900. Replacing drift with a repeated still frame lowers the combined equal-condition score by 0.089 [0.066, 0.112]. However, a receptors-only descriptor reaches 0.891 and a task-trained 5,888-parameter CNN averages 0.959, so the current evidence supports transfer and the usefulness of active sampling, but not an advantage of the biological wiring. We separate project-recorded results from recomputed checks and report a small real-drawing audit. The supported claim is therefore narrow: motion-oriented biological vision can be repurposed as a useful texture representation for architectural hatch matching, while the topology contribution and end-to-end BIM utility remain open questions.
Applications of Neural Cellular Automata: State of the Art, Challenges and Opportunities
神经元胞自动机的应用:现状、挑战与机遇
Lemke, Nick, Ihm, Niklas, Kalkhof, John, Konstantin, Mirko, Krumb, Henry J., Lang, Daniel M., Pajouheshagar, Ehsan, Sadafi, Ario, Kuijper, Arjan, Lekadir, Karim, Lorenzi, Marco, Marr, Carsten, Schnabel, Julia A., Mukhopadhyay, Anirban
Abstract
Neural Cellular Automata (NCAs) are a new type of neural network architecture which enable accurate and robust inference at extremely small model sizes. Recently, NCAs have advanced to become interesting low-resource alternatives to convolution- and attention-based architectures for various tasks such as image analysis, synthetic image generation, and simulation. The rapid development and increased research interest necessitate a comprehensive review of the emerging technology. This review provides an overview of the fundamentals of NCAs, applications to medical imaging, as well as insights into the state of the art. We analyze recent modifications to the originally proposed NCA architecture with respect to their efficiency and accuracy. Furthermore, we review practical applications in real-world scenarios with a focus on medical image analysis, segmentation, classification, registration, depth estimation, and image synthesis. Finally, we identify several advantages of NCAs, research gaps, and conclude with an analysis of future opportunities for NCAs in medical applications in confined settings or areas that have particular demands for robustness or efficient data processing.
Beyond Uniform Subspaces: Spectrum-Aware and Depth-Adaptive Fusion for Multi-Task Model Merging
超越均匀子空间:面向多任务模型合并的谱感知与深度自适应融合
Gu, Ruxi, Wang, Zilei, Wang, Wei
Abstract
Model merging aims to consolidate multiple task-specific models without access to extra training process. However, existing subspace-based methods largely rely on a uniform treatment of task updates, overlooking their intrinsic spectral and depth-wise heterogeneity. We identify two key deviations from this assumption: different tasks require different subspace capacity and exhibit different tolerance to spectral transformation, while subspace projection introduces depth-dependent distortion. Based on these observations, we propose SADA-Merging, a spectrum-aware and depth-adaptive framework for data-free model merging. SADA-Merging allocates task-specific subspace capacity according to spectral complexity, adapts spectral preservation according to task-wise plasticity, and applies depth-dependent anchoring to compensate for projection-induced distortion. This enables the fusion process to adapt to both the intrinsic geometry of each task and its sensitivity across network depth. SADA-Merging operates directly on task updates and is applicable to both full fine-tuning and LoRA settings. Extensive experiments demonstrate consistent improvements over existing data-free merging methods across different task scales and adaptation settings.
Video-based Surgical Skill Assessment Using Dynamics-and-Uncertainty-Aware Tree-based Gaussian Process Classifier
基于动态性与不确定性感知树状高斯过程分类器的视频手术技能评估
Rezaei, Arefeh, Ahmadi, Mohammad Javad, Molaei, Amir, Taghirad, Hamid D.
Abstract
The proposed pipeline integrates a representation-flow convolutional neural network with a dynamics- and uncertainty-aware tree-based Gaussian Process classifier. In this framework, latent motion dynamics are exploited both as discriminative representations and as a source of input uncertainty, enhancing robustness against temporal variations and abnormal motion transitions. Compared with conventional deep learning approaches, the proposed strategy requires less training data and offers improved computational efficiency. To further improve classification performance, we introduce novel semantic-aware compound kernels that effectively capture semantic, flow, and dynamic information embedded in surgical video features. In addition, uncertainty-aware kernels are developed to strengthen the robustness and practical applicability of the compound kernel framework. The proposed method is evaluated on two benchmark datasets, namely the JIGSAWS and the Cataract-LMM (Capsulorhexis) datasets. Experimental results demonstrate strong performance across both datasets, including the LOSO and LOUO evaluation protocols on JIGSAWS, including the subject-independent LOUO protocol on JIGSAWS, on which the framework attains a mean accuracy of \ph{96.9}\%; results under the within-subject LOSO protocol are reported for comparability with prior work, achieving competitive accuracy while substantially reducing computational cost. Overall, the proposed pipeline provides an efficient and accurate framework for video-based surgical skill assessment.
Latent world models learn predictive representations for autonomous driving, but the relational semantics these states preserve often remain implicit. We investigate whether traffic scene graphs can serve as privileged semantic supervision for latent world representations. Building on LAW, we construct actor-centric scene graphs from nuScenes 3D annotations, encode their serialized relational structure using a frozen text embedding model, and align the visual latent representations with this semantic target during training. We remove the supervision branch at inference, so it requires neither scene graphs nor 3D annotations and adds no test-time computation. On nuScenes, our method reduces average trajectory L2 error from 0.661 to 0.622 (5.9%) and collision rate from 0.456 to 0.217 (52.4%) relative to our retrained LAW baseline. It also outperforms an unstructured caption-style semantic target, supporting the benefit of explicit relational structure for latent world-model representation learning.
Chinese Translation
潜在世界模型为自动驾驶学习预测性表征,但这些状态所保留的关系语义往往仍是隐式的。我们研究了交通场景图能否作为潜在世界表征的特权语义监督信号。在 LAW(LAW)的基础上,我们基于 nuScenes 3D 标注构建以参与者为中心的场景图,使用冻结的文本嵌入模型对其序列化的关系结构进行编码,并在训练过程中将视觉潜在表征与该语义目标进行对齐。在推理阶段我们移除监督分支,因此该方法既不需要场景图也不需要 3D 标注,且不增加任何测试时计算量。在 nuScenes 数据集上,与重新训练的 LAW 基线相比,我们的方法将平均轨迹 L2 误差从 0.661 降至 0.622(降低 5.9%),碰撞率从 0.456 降至 0.217(降低 52.4%)。该方法还优于非结构化的描述式(caption-style)语义目标,这支持了显式关系结构对潜在世界模型表征学习的益处。
Multi-organ segmentation using deep learning requires large amounts of annotated patient data; however, institutions often lack sufficiently large and diverse annotated datasets. Privacy constraints further prevent institutions from sharing patient data to overcome this limitation. Moreover, due to the labor-intensive nature of annotation and the scarcity of diverse expertise, institutions typically have labels for only a small portion of their local data, leaving the larger unlabeled portion unused. In this work, we propose a flexible semi-supervised federated multi-task student-teacher framework that leverages federated learning (FL) to improve multi-organ segmentation using both labeled and unlabeled data across participating sites. At each communication round, the proposed framework initiates local training, where clients with labels for the same task form a federation to produce an aggregated teacher model. The resulting teachers generate task-specific features for all data at each client. Subsequently, all clients form a second federation to train a multi-task student model with a shared encoder and task-specific decoders that replicate the teacher-generated features across all segmentation tasks. The aggregated student model is then used to update the local teachers and initiate the next training round. Extensive experiments demonstrated the effectiveness of the proposed method compared with local and federated single-organ models, yielding an average performance gain of 13 percent across clients. The experiments also demonstrated the impact of multi-task learning and unlabeled data and the applicability of the framework in relaxing labeled-data requirements for client participation. The code is available at https://github.com/AshknMrd/FedMust.
High-resolution Nitrogen Dioxide Maps Reveal Exposure Limit Breaches across Europe
高分辨率二氧化氮地图揭示欧洲各地的暴露限值超标情况
Scheibenreif, Linus, Schindler, Konrad
Abstract
Nitrogen dioxide (NO2) is a common air pollutant, released into the atmosphere through the incomplete burning of fossil fuels, and associated with respiratory and cardiovascular diseases in humans. Ambient NO2 concentrations are regulated through air-quality limits assessed with a sparse network of fixed monitors. The revised EU Ambient Air Quality Directive (2024/2881) introduces a daily NO2 limit to be met from 2030. At present, neither the regulatory monitoring network nor existing coarse, annual-mean models can resolve NO2 concentrations at the spatio-temporal resolutions necessary to assess compliance. Here we map NO2 across Europe at hourly and 10m resolution with a machine-learning model that combines ground monitors with satellite, reanalysis, land-use, traffic and emission data and returns a calibrated predictive distribution at every location. Validated against held-out regulatory monitors and independent citizen-science campaigns, the maps resolve high-resolution spatiotemporal NO2 gradients for 110 metropolitan areas in Europe. We reconstruct the daily compliance statistic across those regions and find limit breaches in 91 EU air quality zones deemed compliant by the regulatory monitoring network, covering a population of approximately 135M. Beyond air quality zones and monitor locations, an estimated 9-9.4% (20M) of the population in mapped regions lives in areas where the daily NO2 limit is breached. The high-resolution maps offer a route to population-scale assessment of compliance with the 2030 limits.
Reliable uncertainty estimation is essential for deploying object detectors in autonomous systems operating in uncertain environments. Evidential Deep Learning (EDL) provides a principled framework for uncertainty-aware classification by representing network outputs as evidence and interpreting predictions through subjective logic. However, existing evidential object detectors typically combine evidential classification with regression uncertainty models that do not share the same theoretical foundation. In this work, we propose an evidential version of YOLOv8 in which both classification and bounding-box regression are formulated within a common evidential framework. Our approach exploits YOLOv8's distribution-based bounding-box representation, allowing the evidential formulation to be applied not only to classification but also to localisation. As a result, both tasks produce belief, uncertainty, and probability estimates that can be interpreted within the Dempster--Shafer framework. Experiments on KITTI, MUSES, and nuScenes show that the resulting detector remains broadly competitive with standard YOLOv8 in terms of detection accuracy while providing a localisation uncertainty that effectively discriminates between correct and erroneous detections. Moreover, this uncertainty becomes increasingly discriminative under domain shift.
Chinese Translation
可靠的不确定性估计对于在不确定环境中运行的自主系统中部署目标检测器至关重要。证据深度学习(Evidential Deep Learning, EDL)通过将网络输出表示为证据并借助主观逻辑(subjective logic)解释预测,为不确定性感知分类提供了一个有原则的框架。然而,现有的证据性目标检测器通常将证据性分类与不共享同一理论基础的不确定性回归模型相结合。在本工作中,我们提出了YOLOv8的证据性版本,其中分类和边界框回归均在统一的证据性框架内进行公式化。我们的方法利用了YOLOv8基于分布的边界框表示,使证据性公式化不仅能应用于分类,还能应用于定位。由此,这两个任务都能产生可在Dempster--Shafer框架内解释的信念(belief)、不确定性和概率估计。在KITTI、MUSES和nuScenes数据集上的实验表明,所得检测器在检测精度方面总体上与标准YOLOv8具有竞争力,同时其定位不确定性能够有效区分正确与错误的检测结果。此外,在域偏移(domain shift)条件下,这种不确定性的区分能力会不断增强。
Latent diffusion models now dominate medical image generation, and every such pipeline rests on a \emph{tokenizer} that compresses images into the latent codes for image generation to operate on. Thereby, the tokenizer choice bounds every downstream task from reconstruction fidelity and generation quality to the representations available for downstream analysis. Yet, medical imaging pipelines routinely utilize tokenizers from natural imaging on the hypothesis that their behavior carries over. However, this is an assumption never tested in the medical imaging regime, where datasets are orders of magnitude smaller and images exhibit far lower inter-sample variance. We present a systematic evaluation of medical image tokenizers evaluating thirty configurations across ten model families on twelve datasets at three compression factors, spanning reconstruction, generation, latent geometry, downstream classification, and memorization. We find that (1) performance on image reconstruction and generation strongly correlate, unlike prior reports on natural images; (2) modern tokenizers use nearly all of their codebook entries, but still leave most of the latent space unused; (3) training-set memorization is mild and is further suppressed by stronger latent space compression; and (4) discrete quantization can largely preserve downstream classification, with lookup-free schemes being the main exception.
Spatial--temporal benchmarks are valid only when their questions, source annotations, and answer options identify the same physical quantity. We audit STI-Bench against the official ScanNet, Waymo, and Omni6DPose sources and find systematic coordinate-system and timestamp errors, under-specified targets and times, and disagreements between keyed options and answer details. We introduce ReSTI, a source-backed revision that reconstructs every recoverable answer under an explicit target, time, coordinate system, physical quantity, and unit. Source reconstruction reveals task-level geometric failures: ScanNet Grounding omits the required alignment between annotation and raw camera coordinate systems, while Orientation measures camera rotation on the wrong plane. ReSTI replaces these labels with explicit, source-consistent geometric definitions and corrects other source-verifiable defects, including Waymo poses evaluated at the wrong timestamp. Across 2,064 legacy questions, ReSTI retains 1,782 questions and records 282 evidence-backed exclusions. ReSTI therefore provides a conservative and source-traceable basis for evaluating precise video spatial--temporal reasoning. Project page: https://github.com/pengzhansun/ReSTI.
GraphSVR: q-Space--Aware Graph-Based Slice-to-Volume Registration for Diffusion MRI
GraphSVR:面向扩散MRI的q空间感知的基于图的切片到体配准
Kertes, Noga, Sourani, Daphna Link, Bronstein, Alex M., Freiman, Moti
Abstract
Diffusion-weighted imaging (DWI) remains highly vulnerable to subject motion, particularly in time-efficient protocols and in motion-prone populations. While slice-to-volume registration (SVR) can mitigate inter-slice and inter-stack misalignment, diffusion MRI introduces additional complexity due to diffusion-direction-dependent contrast and the requirement to align dozens of measurements within a common reference frame, effectively yielding a 4D registration problem. Existing approaches rely primarily on sequential modeling or pairwise similarity and often degrade under sparse gradient sampling or severe motion. We introduce GraphSVR, a q-space-aware graph-based framework for 4D SVR registration in DWI. GraphSVR represents slice groups as nodes in an acquisition-structured graph, with edges encoding temporal proximity, spatial slice geometry and diffusion encoding relationships. A graph neural network predicts globally consistent stack-wise rigid motion, optimized in a self-supervised, zero-shot manner using only an anatomical reference image, without requiring paired ground-truth motion. We evaluate GraphSVR using both fully synthetic diffusion simulations and realistic recombination-based simulations from real acquisitions with controllable motion severity and gradient sparsity. Performance is quantified using grid error (mm) and rotation error relative to known ground-truth transforms. Under severe motion, GraphSVR reduces grid error and rotation error by 73% compared to FSL eddy, the standard DWI motion-correction method, with the largest gains observed in sparse-direction regimes. These results demonstrate that explicitly modeling acquisition structure through graph-based reasoning improves robustness and global consistency in 4D DWI motion estimation. Code is available at https://github.com/nogakertes/GraphSVR.git.
Vuong, Trinh T. L., Graham, Simon, Vu, Quoc Dang, Ho, Phat T. H., Lim, Jeewoo, Jahanifar, Mostafa, Rajpoot, Nasir, Kwak, Jin T.
Abstract
Mitotic figure (MF) analysis supports tumor grading and prognostic assessment, but automated models remain sensitive to differences in tissue type and image acquisition. We present MiTHras, a task-specific pretraining framework that combines pseudo-label-guided image- and token-level contrastive learning with masked reconstruction. We construct TCGA-MF-Pseudo, a corpus of 1.8 million cell-centered images from 14 TCGA cohorts spanning 11 organ sites. Comprehensive evaluation on MF classification, detection, count-based survival prediction, and subtype classification demonstrates the efficacy of MiTHras. It achieves the highest mean F1 on all three MF classification benchmarks and both subtype benchmarks. MiTHras also outperforms general-purpose and pathology foundation encoders by a larger margin under frozen-encoder linear probing than under full fine-tuning. Although detection gains are modest due to a shared candidate-detection stage, ablations confirm that token-level supervision improves typical-versus-atypical classification and linear probing. These findings establish that MiTHras yields robust, transferable representations for automated mitotic activity assessment.
We introduce Ananke, a representation-learning framework that scaffolds latent representations onto a structured product-torus prior, and its flagship visual backbone realization, Contractive Torus Attractor Networks (CTAN). By factorizing high-dimensional latent spaces into an orthogonal direct sum of two-dimensional phase planes ($\bigoplus_{k=1}^K \R^2$), Ananke coordinates feature updates via a decoupled dual-phase continuous flow: skew-symmetric Hamiltonian transport moves features tangentially along energy level sets to preserve semantic phase invariants, while signed gradient dissipation contracts transverse perturbations normally toward target invariant manifolds. For circular potential families with frozen parameters, logarithmic radial feedback yields the Exact Log-Symplectic Flow (ELSF), an analytical closed-form mapping with exact exponential decay of log-radius error that evaluates in a single forward pass without numerical integration. We establish local input-to-state bounds for level-set deviations and log-radius errors, and characterize the normal hyperbolicity and persistence of the ideal product torus under bounded perturbations. We further formulate the architecture through Lie--Trotter operator splitting, unifying spatial depthwise diffusion with local manifold contraction, and analyze both exact trigonometric flows and hardware-friendly symplectic dual-shear variants. Across natural image benchmarks (CIFAR-100) and clinically challenging endoscopy datasets (Kvasir-v2), CTAN demonstrates exceptional parameter efficiency: an ultra-compact hierarchical model with merely 0.27M parameters achieves 90.52\% accuracy on Kvasir-v2, outperforming 25M+ baselines (ResNet-50, DenseNet-161) by nearly two orders of magnitude in capacity, while scaled variants attain 80.32\% top-1 accuracy on CIFAR-100.
PrismGPT: Proxy-Guided Learning for Region-Aware Photo Editing with Self-Synthesized Reasoning
PrismGPT:基于代理引导学习的区域感知照片编辑与自合成推理
Zhao, Ke, Nguyen, Hue, Punnappurath, Abhijith, Wang, Zhongling, Mohomed, Iqbal, Brown, Michael S.
Abstract
Professional photo finishing relies on both global adjustments and region-specific local edits guided by semantic masks, yet current automated methods handle this workflow only partially. We present PrismGPT, a Vision-Language Model (VLM) framework that produces structured, region-aware editing plans from a single input image without relying on commercial black-box tools. Training a VLM to simultaneously diagnose aesthetic deficiencies at both global and local levels while predicting precise editing parameters is challenging due to the vast combinatorial decision space. We address this through proxy-guided learning: two simpler proxy tasks -- operation decomposition and region-aware aesthetic ranking -- teach the foundational skills the model needs, while a competence-based dynamic scheduler automatically rebalances the multi-task training ratio, progressively shifting emphasis from the proxy tasks to the primary editing task as each skill is mastered. Crucially, all reasoning traces used for supervised fine-tuning are self-synthesized by the same base model, eliminating the need for a stronger external teacher. Experiments on MIT-Adobe FiveK and SPIRE, a new professionally retouched benchmark we introduce, show that PrismGPT achieves state-of-the-art results while using only ~6% of the training data compared to the previous best method.
Brain Metastases Segmentation for BraTS 2026 Task 1: A Multi-Architecture Comparison
BraTS 2026 任务1的脑转移瘤分割:多种架构的比较研究
Islam, Mahdi, Tabassum, Musarrat
Abstract
Brain metastases are the most common intracranial malignancy, occurring in roughly 30% of patients with primary solid tumors and carrying a median survival near 5.9 months. Automated segmentation is critical for treatment planning and volumetric monitoring, but metastases are frequently small, numerous, and heterogeneous in size within a single patient. We compare a plain nnU-Net baseline, a Residual Encoder Large (ResEncL) variant, region-based training, and a Primus transformer model for BraTS-METS 2026 Task 1, using patient-grouped cross-validation to prevent leakage from the longitudinal UCSD subset. Primus (label-based) is our strongest individual model by aggregate DSC/NSD, achieving 0.710/0.761 (ET), 0.742/0.785 (TC), 0.683/0.689 (WT), and 0.531/0.436 (RC). ResEncL trails Primus on aggregate DSC/NSD but achieves substantially higher lesion-wise F1 (e.g. ET: 0.452 vs. 0.052); a probability-averaging ensemble of the two only partially preserves ResEncL's F1 advantage (ET lesion-wise F1: 0.064). We further report three postprocessing and label-reconstruction pitfalls we believe generalize beyond this challenge. Code is available at https://github.com/mahdiislam79/BraTS_METS_2026.
A new concept, termed virtual neural networks, is introduced, where the count of trainable parameters is kept constant, and scalability is attained purely through computational resources. This concept is an abstract framework that can be realized using any standard convolutional neural network. It merges siamese neural networks with a deep ensemble technique by generating numerous virtual models that share weights derived from a small set of physical models. The ensemble comprises up to hundreds of trained models simultaneously. All virtual networks take the same input, and their interconnected structure induces an internal distortion that boosts the entire ensemble robustness. The accuracy of the ensemble improves as the number of virtual networks increases, without changing the capacity. Virtual neural networks outperform larger capacity models, typical deep ensembles, and contemporary approaches like SWA and Masksembles. Additionally, the highest-performing individual model from the ensemble surpasses other models trained individually, even those with a greater number of parameters. Code: gitlab.com/EnginCZ/virtual-models-public
Forest inventories increasingly rely on artificial intelligence (AI) models to derive forest attributes from large-scale 3D point clouds. Current models are typically specialized to a single task, sensor, and forest type, making adaptation expensive in terms of annotations, computation, and expertise. We ask whether a single pretrained model can instead learn transferable representations across diverse forest inventory settings. Inspired by recent developments in language modelling and computer vision, we take a step toward a foundation model (FM) for 3D forestry. Using LitePT as backbone, we first establish a strong supervised baseline that sets a new state of the art on forest semantic and instance segmentation, tree species classification, and age regression benchmarks. We then curate a large-scale unlabelled corpus spanning airborne, UAV, and mobile laser scanning across diverse forest ecosystems, and pretrain the same backbone using self-supervised learning. We systematically evaluate representation learning strategies by comparing training from scratch, supervised pretraining, and self-supervised pretraining across four representative forestry tasks, under varying annotation budgets. Compared with training from scratch, self-supervised pretraining accelerates model convergence and consistently improves performance when annotations are scarce. Compared with task-specific supervised pretraining, self-supervised pretraining yields more transferable representations across downstream forestry tasks. These findings identify the practical regime in which pretrained representations are most valuable and suggest that instance discrimination, rather than forest semantics, is the main remaining obstacle to a general-purpose 3D forest foundation model. Code and models are available at: https://github.com/prs-eth/ForPT.
In this paper, we propose SVEET, a framework that requires merely training on a pretrained bidirectional video diffusion model but supports high-quality streaming video editing in an auto-regressive fashion. To tackle this problem, we first systematically revisit existing video-to-video diffusion approaches and identify two key principles for such streaming adaptation: backbone feature disentanglement and conditional frame independence. Building on these insights, we develop a novel paradigm for controllable video generation. At its core, an auxiliary model branch encodes source video inputs with temporally independent self-attention, and the intermediate features are injected into the corresponding backbone blocks for streaming-compatible control. Moreover, to bridge the discrepancy between the feature spaces of bidirectional and streaming models, we propose a decoupled training scheme that explicitly enforces the orthogonality between the optimization directions of video controllability and model causality. Such disentanglement ensures compatibility between the two objectives at inference and facilitates smooth zero-shot knowledge transfer across heterogeneous backbone architectures. Extensive experiments demonstrate that SVEET achieves superior editing quality while maintaining real-time performance, attaining 15 FPS on a single H100 GPU 17 without any auxiliary acceleration techniques. Codes are available at https://github.com/YujiaHu1109/SVEET.
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in multimodal tasks, yet they still exhibit poor ability in spatial reasoning. Existing training-dependent and training-free enhancement methods suffer from high computational costs with catastrophic forgetting and internal mechanism interference that compromises general capabilities, respectively. In this work, we first verify two key hypotheses: appropriate geometric image transformation and query-reversal transformation can recover incorrect spatial predictions, and correct predictions exhibit higher relation-token confidence than incorrect ones. Based on these findings, we propose INTCORT, a training-free spatial reasoning enhancement framework that constructs multiple inference views through input transformations and aggregates their predictions via relation-token confidence routing, without modifying the VLM's internal mechanisms. Experimental results on several commonly-used benchmarks demonstrate that INTCORT substantially improves spatial reasoning accuracy across diverse VLMs, achieving an average improvement of 10.01% over all models and benchmarks. Compared with prior works, INTCORT achieves superior performance with improvements of up to 25.01%.
Advances in processing power, camera technologies, and mobile image analysis have made smartphones and other mobile devices, such as laptops, increasingly suitable for medical diagnosis and healthcare applications. Researchers have developed low-cost solutions for the early detection and monitoring of various health conditions, including eye and ENT diseases, malnutrition, heart rate variability, skin and oral conditions, and injuries, using images captured by non-medical devices such as smartphones and webcams. This survey examines existing research on mobile image-based medical diagnosis, with an emphasis on its potential to enable low-cost and accessible healthcare. We comparatively analyze state-of-the-art solutions across different healthcare application categories, examining their advantages and limitations. Based on this analysis, we identify desirable characteristics of mobile image-based diagnostic tools and highlight areas where existing approaches have made progress as well as areas requiring further research. We also discuss application-specific and common challenges and outline directions for future research. Overall, this study provides a comprehensive overview of mobile image-based healthcare solutions and their potential to support low-cost disease diagnosis and monitoring, particularly for underserved populations in remote and resource-constrained settings.
ZVeC: A Zero-Shot Framework for Instance-Level Vehicle Extraction and Generative Point Cloud Completion
ZVeC:一种用于实例级车辆提取与生成式点云补全的零样本框架
Li, Daisy, Gao, Kyle, Wu, Quanyun, Jutzi, Boris, Zelek, John S., Li, Jonathan
Abstract
LiDAR point clouds acquired in underground environments exhibit severe geometric incompleteness due to occlusions and limited sensor viewpoints, making reliable point cloud completion challenging without large supervised datasets. We propose ZVeC, a zero-shot, instance-driven framework that reformulates scene-level completion as compositional object-level reconstruction. By decomposing a scene into semantic object instances, ZVeC reduces reconstruction ambiguity in cluttered environments while eliminating the need for scenario-specific training. Each segmented vehicle is completed independently using a depth- and 3D Gaussian-conditioned diffusion model that exploits generalized geometric priors before the reconstructed instances are recomposed into the original scene. To evaluate our approach, we construct a real-world dense LiDAR benchmark of underground parking environments. Experimental results demonstrate consistent improvements over representative scene-level baselines in both quantitative metrics and visual quality. The completed point cloud differs substantially from the measured input (average KL divergence ~ 2.1), yet reducing the input to only 1% of the original LiDAR measurements changes the completed reconstruction only marginally (KL divergence < 0.50). This demonstrates that ZVeC produces geometrically consistent completions even under extreme input sparsity.
When Wider Views Fail: Stress-Testing Feed-Forward 3D Reconstruction
当更宽的视角失效时:前馈式三维重建的压力测试
Li, Daisy, Gao, Kyle, Wu, Quanyun, Chomko, Hanna, Zelek, John S., Li, Jonathan
Abstract
Feed-forward 3D reconstruction models enable efficient geometry estimation from sparse images, but their pretrained nature can make them vulnerable to distribution shifts beyond their training data. Identifying these failure modes is important for understanding when such models can be reliably deployed in unconstrained imaging settings. We investigate viewpoint variation as a controlled distribution shift by varying the angular span of sparse image inputs while keeping the input budget fixed. Across multiple feed-forward reconstruction models, we observe substantial degradation as viewpoint span increases, with wide spans producing both incomplete surface coverage and geometry unsupported by the observed imagery. These results reveal that viewpoint variation can induce failure modes beyond conventional reconstruction incompleteness, highlighting the need to evaluate pretrained feed-forward models under distribution shifts that challenge their learned geometric priors.
Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera parameters are supplied. In this paper, we study the multi-view stereo (MVS) problem with known camera parameters and propose a novel approach that bridges conventional MVS and FF methods. Rather than casting MVS as a sequence-to-one mapping that predicts depth only for a single reference view, we reformulate it as a sequence-to-sequence task, akin to FF models, that jointly predicts geometry for all input views. We introduce a global transformer-based architecture with two components that explicitly exploit camera-induced priors: ray-map embeddings that inject camera parameters into image patch tokens, making the transformer camera-aware, and a unified global cost volume that replaces conventional per-view cost volumes to jointly capture 3D structure across all views. Extensive experiments on multiple public benchmarks show our approach achieves state-of-the-art performance, surpassing both MVS and FF reconstruction baselines.
DTKDP: A Dual Teacher Knowledge Distillation and Pruning Framework for Lightweight Oriented SAR Ship Detection
DTKDP:一种面向轻量化SAR舰船检测的双教师知识蒸馏与剪枝框架
Li, Yuming, Zhang, Fan, Achim, Alin M.
Abstract
Two-stage oriented detectors achieve high localization accuracy in synthetic aperture radar (SAR) ship detection, but their large backbones, feature pyramids, proposal modules, and heavy region of interest (RoI) heads hinder deployment. Existing lightweight SAR ship detectors typically use one-stage frameworks that lack proposal-level refinement for precise rotated localization. This paper presents a dual-teacher knowledge distillation and pruning (DTKDP) framework for lightweight oriented SAR ship detection. DTKDP introduces learnable gates into convolutional, normalization, and linear layers to prune convolutional channels and RoI-head neurons. Rotated proposal alignment (RPA) distills teacher and student predictions in a shared teacher-generated rotated proposal space, while a dual-teacher scheme combines classification and regression guidance from a homogeneous main teacher with complementary classification cues from a heterogeneous auxiliary teacher. Experiments on the SAR Ship Detection Dataset (SSDD) and Rotated Ship Detection Dataset in SAR Images (RSDD-SAR) show that DTKDP reduces the parameters of Oriented Region-based Convolutional Neural Network (Oriented R-CNN) and RoI Transformer equipped with ResNet-50 backbones by 87.5-91.8% and their floating-point operations (FLOPs) by 75.6-79.9%. In terms of average precision (AP) and mean average precision (mAP), the resulting Oriented R-CNN-slim and RoI Transformer-slim retain accuracy close to their full-scale counterparts. Relative changes across $\mathrm{AP}_{50}$, $\mathrm{AP}_{75}$, $\mathrm{mAP}_{50:75}$, and $\mathrm{mAP}_{50:95}$ range from a 2.38% decrease to a 0.65% improvement. Compared with RTMDet-tiny, they improve all four metrics on both datasets by 0.52-27.55% and consistently surpass representative distillation methods, demonstrating a favorable accuracy-efficiency trade-off.
Recent foundation models are moving toward native multimodal Vision-Language Models (VLMs), making VLMs a central form of next-generation foundation models. However, their large language backbones make edge deployment difficult due to high memory footprint and memory-bound autoregressive decoding. Weight-only post-training quantization is a practical solution, but pushing VLMs to extreme low bit-widths remains challenging: existing rotation-free methods suffer from outliers at 2-3 bits, while rotation-based methods improve accuracy at the cost of additional runtime overhead. We propose SPHQuant, a rotation-free spherical weight-only quantization framework for VLMs. Instead of quantizing weights directly in Cartesian coordinates, SPHQuant decomposes each 8D weight vector into coordinate signs, radius, and a positive unit direction. This representation isolates outlier magnitude into the radius while keeping directions bounded and statistically regular. Based on this insight, SPHQuant allocates extra precision to the radius to mitigate accuracy degradation induced by outliers. It further uses a compact positive-direction codebook and fine-tunes codebook entries through angular parameterization to preserve the unit-sphere constraint. We also design a hardware-friendly GEMV kernel that keeps the direction codebook small enough for shared-memory lookup and packs radial bits efficiently. Experiments show that SPHQuant matches the performance of state-of-the-art extreme low-bit quantization methods while improving decode throughput over QTIP by 30.3% on RTX A6000. Code will be released in https://github.com/Pushazf/SPHQuant.
Generating Chest X-Ray Counterfactuals by Specialising Foundation Image Models
通过专门化基础图像模型生成胸部X光反事实图像
Xing, Xiaodan, Rasal, Rajat R., Meister, Julia A., Ghorayeb, Sara, Khara, Galvin, Schrouff, Jessica
Abstract
Counterfactual image generation answers questions about how a subject would have looked under retrospective, hypothetical scenarios. Recent methods have improved perceptual quality, identity preservation and faithfulness to an underlying causal model, but their adoption in healthcare is limited by scarce annotated data, distribution shift between datasets, and mismatches between pretrained generative models and those required for counterfactual inference. We propose specialisation, a data and parameter-efficient framework for adapting pretrained, non-causal generative models into causal mechanisms under distribution shift. Based on this framework, we train a radiology counterfactual image generation model, called RadCF, using latent flow matching. We validate our approach on three chest X-ray datasets spanning different dataset shifts, data volumes, and counterfactual questions, associated with challenging, highly-localised interventions. Our results show that RadCF and specialisation improve counterfactual soundness over existing methods while being data and parameter efficient, and that the resulting counterfactuals can detect and mitigate shortcut learning in a downstream medical classifier. Code is available at https://github.com/GSK-AI/RadCF/.
SLICEChat: Progressive In-Encoder Token Pruning for Whole-Slide Pathology Language Models
SLICEChat:面向全切片病理语言模型的编码器内渐进式词元剪枝
Bozkurt, Ali Kerem, Bakay, Baris Cem, Kulac, Ibrahim, Gunduz-Demir, Cigdem, Erdem, Erkut, Erdem, Aykut
Abstract
Whole-slide pathology images (WSIs) contain gigapixel-scale visual content, creating a major scalability challenge for slide-level multimodal large language models (MLLMs). Existing approaches process thousands of patch tokens and typically apply compression only after slide encoding, leaving multimodal attention computationally expensive. We introduce SLICEChat, a slide-level MLLM that integrates progressive token pruning within a hybrid Mamba--Transformer slide encoder. Mamba layers enable efficient long-range propagation, while Transformer layers preserve global interactions as the sequence is progressively shortened. Between stages, language-supervised, region-aware pruning removes spatially coherent low-utility regions under a controlled keep-rate schedule, producing compact slide representations before multimodal fusion. On SlideBench VQA, SLICEChat achieves 79.84% accuracy on TCGA and 59.09% on BCNB cohorts, outperforming prior slide-level pathology MLLMs, and achieves the highest overall WSI-Bench metrics. It also provides competitive memory usage and the inference latency among the evaluated models. These results demonstrate accurate and computationally efficient multimodal reasoning over gigapixel WSIs.
Recent advances in pixel-space diffusion models have narrowed the image quality gap with latent-space diffusion, but still converge more slowly and lag behind in final image quality. We argue that a key reason is the lack of an explicit representation prior: unlike latent diffusion, which usually denoises in a compact and structured latent space, pixel diffusion needs to learn denoising-friendly representations and pixel generation simultaneously from raw RGB space. To address this problem, we propose PixelDiT2, an end-to-end pixel-space diffusion model designed to decouple representation learning from pixel generation without introducing an autoencoder or latent reconstruction bottleneck. We propose representation grounding that uses a frozen pretrained vision foundation model to provide explicit per-patch representation guidance throughout denoising, allowing the pixel diffusion transformer to focus more on pixel generation. On ImageNet-256x256, PixelDiT2 achieves an FID of 1.46 after 600 epochs; at 512x512 resolution, PixelDiT2 achieves an FID of 1.48 after 680 epochs.
Bone overlap can obscure abnormalities in chest radiographs, while scarce paired training data limit supervised bone suppression. We address this challenge with a digitally reconstructed radiograph (DRR) framework that converts chest computed tomography (CT) into paired supervision for component suppression. A novel bone segmentation algorithm enables CT decomposition into bone, non-lung soft-tissue, and lung components, which are projected separately. Their weighted combination yields synthetic radiographs with pixel-registered component images that sum exactly to the full DRR. Models trained on these data suppress bone or lung components by predicting the target component and recovering the remainder by subtraction, transferring to real radiographs without real paired training data. As an extension, their outputs on real radiographs provide target domains for unpaired, component-wise DRR translation, reducing the appearance gap while retaining anatomical details. Across multiple public datasets, downstream detection experiments demonstrate the utility of bone suppression, with gains concentrated on abnormalities with substantial bone overlap. Compared with open-source DRR engines applied to the same CTs, our unmodified DRRs achieve comparable realism and preservation of label-relevant anatomy, while translated DRRs achieve the best Fr\'echet inception distance (FID), lung-field sharpness, and agreement with source-CT anatomy among the evaluated methods. Models and inference code: https://huggingface.co/qureaiorg/bone-suppression; Translated projections: https://huggingface.co/datasets/qureaiorg/ct2xr-projections.
We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse generation tasks. In controlled comparisons that hold the generator and training protocol fixed, replacing the latent with GAE improves both visual quality and independently measured 3D coherence: FVD falls by $12.7\%$ and $23.1\%$ on RealEstate10K and DL3DV, and camera-trajectory error is halved on RealEstate10K. Together, these results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. In this paper, we present VideoGen-Agent, a multimodal agent trained through multitask agentic reinforcement learning to use external tools for video generation. The agent coordinates augmentation, generation, and verification tools through multi-turn interactions, using the prompt and intermediate observations to guide its decisions. We train a shared policy on a category-balanced dataset spanning six tasks. Supervised fine-tuning on teacher-generated trajectories establishes tool-use behavior, which is then refined through reinforcement learning. A category-aware hybrid reward evaluates tool-call validity, task-appropriate tool use, and generated video quality. We further introduce VABench, a held-out benchmark of 600 prompts covering procedural knowledge, single- and multi-entity identity preservation, physical consistency, scene composition, and multi-shot temporal structure. On VABench, VideoGen-Agent improves over its base text-to-video generator by 19.1 points, from 56.5 to 75.6. Upgrading the generation tools further raises the score to 86.1 without additional agent training. Human raters prefer the upgraded configuration over the strongest standalone baseline in 84.3% of comparisons. These results support learning tool use across video-generation tasks and show that the trained agent can benefit from subsequent advances in generation tools.
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. GameHorizon Suite consists of three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, utilizing the pipeline, we construct GameHorizon-Data, the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games, collected by 100 human expert players. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation using thousands of standardized questions organized into three primary tasks and a series of diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. Based on our GameHorizon Suite, we evaluate 47 models through more than one million model invocations, revealing a meaningful hierarchy of task difficulty and pronounced differences in model capabilities. Our work can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families. We will release our dataset, annotator, and benchmark to facilitate future research.
Accuracy of Low-bit quantization of linear layers is often dominated by a small number of outliers. Although existing methods, such as smoothing, rotation, or residual-based approaches, may mitigate this problem, they often introduce new accuracy bottlenecks to weights. Besides, most of these techniques are implemented as online approaches, which can result in heavy execution overheads. To address the afore-mentioned issues, We propose PRQuant (Permutation Residual Quantization), a training-free and low-overhead framework that combines channel reorganization with static weight-side residual compensation. After AWQ-style scaling, PRQuant identifies the input channels that contribute most to weight quantization error, permutes them into contiguous tail blocks, and constructs their residual weight sub-tensors offline. During inference, this contiguous structure enables the activation side to use tail blocks seamlessly without the expensive online gathering operation, and turns scattered residual compensation into a regular tail-augmented GEMM, substantially reducing latency. Experiments demonstrate that PRQuant effectively reduces down-projection reconstruction error. Ablation studies confirm that smoothing and residual compensation are the primary drivers of numerical improvement, while permutation provides a consistent marginal numerical benefit and, more importantly, enables a hardware-friendly contiguous layout that eliminates dynamic gathering overhead. Overall, PRQuant outperforms default MXFP4 and the evaluated PTQ baselines in average accuracy across five downstream benchmarks, improving over MXFP4 by 1.24 and 0.55 on Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507, respectively.
Making prediction with multimodal data is widely used in diverse scenarios. Existing multimodal fusion models, once deployed, can only handle predefined modalities (e.g., vision, text and audio) and single tasks, making it difficult to quickly adapt to new downstream applications. Therefore, a natural yet rather aggressive question arises, whether there exists a general multimodal fusion model that can be applied to arbitrary modality combinations and arbitrary prediction tasks. We argue that a unified multimodal fusion model should not depend on specific modalities and instead encode transferable patterns of multimodal correlation. To this end, we propose a simple and effective learning paradigm based on training over the generation of large-scale synthetic multimodal datasets with diverse causal structures that formally characterize the generative processes of multimodal data in real world. Building on this framework, we propose the generalized multimodal foundation model, a unified foundation model for generalized multimodal learning. By constructing large-scale synthetic multimodal datasets with diverse correlation patterns, our model encodes transferable multimodal correlations during training and activates appropriate associations through in-context examples during inference. Extensive experiments on 18 real-world datasets spanning 12 modalities and 11 prediction tasks demonstrate that our model achieves competitive performance with specialized models without task-specific adaptation.
Chinese Translation
利用多模态数据进行预测已广泛应用于各种场景。现有的多模态融合模型一旦部署,只能处理预定义的模态(如视觉、文本和音频)以及单一任务,难以快速适应新的下游应用。因此,一个自然却颇具挑战性的问题随之而来:是否存在一种通用的多模态融合模型,能够应用于任意的模态组合和任意的预测任务。我们认为,统一的多模态融合模型不应依赖于特定模态,而应编码可迁移的多模态关联模式。为此,我们提出了一种简单而有效的学习范式,其核心是在大规模合成的多模态数据集上进行训练,这些数据集具有多样的因果结构,能够形式化地刻画真实世界中多模态数据的生成过程。基于这一框架,我们提出了广义多模态基础模型(generalized multimodal foundation model),一个面向广义多模态学习的统一基础模型。通过构建具有多样关联模式的大规模合成多模态数据集,我们的模型在训练过程中编码可迁移的多模态关联,并在推理过程中通过上下文示例(in-context examples)激活相应的关联。在涵盖12种模态和11种预测任务的18个真实世界数据集上的大量实验表明,我们的模型无需针对特定任务进行适配即可取得与专用模型相当的性能。
Learning-enabled perception is important in many autonomous systems. Unlike traditional sensors, the boundary where ML perception does or does not work is poorly characterized. Incorrect perception can lead to unsafe or overtly conservative downstream control actions. In this paper, we propose a two-step strategy for correcting ML-based state estimation. First, an offline computation is used to characterize the uncertainties resulting from the ML module's state estimation, using preimages of perception contracts. Second, at runtime, a risk heuristic is used to choose particular states from the uncertain estimates to drive the control decisions. We perform extensive simulation-based evaluation of this runtime perception correction strategy on different vision-based adaptive cruise controllers (ACC modules), in different weather conditions, and road scenarios. Out of 45 ACC scenarios where the original perception-based control system using Yolo and LaneNet led to safety violations, in 73% of the scenarios, our runtime perception correction preserved safety; our method wouldn't be able to recover 27% of the scenarios where the construction of the preimages of perception contracts is not fully conformant. Further, our runtime perception correction strategy is not overly conservative---on the average only a 2.8% increase in completion time is experienced in the corrected scenarios, with mild interventions.
A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation
在选择性在线策略蒸馏中,共享学习率并非中性控制条件
Zhu, Chencheng
Abstract
Selective on-policy distillation trains a student only at the token positions a selector scores highest, and the literature compares selectors under a single shared learning rate--a control chosen to be neutral. We show it is not. Under LoRA on GSM8K (Qwen2.5-1.5B student, 7B teacher), across an 8x learning-rate grid, dense supervision is statistically flat (swing 1.8 pp, p=0.26) while every selective arm moves with the rate: 5.4 pp for a random 5% subset, 6.7 pp for a total-variation selector, up to 17.7 pp for a teachability selector. Consequently the dense-versus-selective verdict reads 10.1 pp at lr=1e-4 but 5.1 pp at 5e-5--a 2.0x difference decided by a parameter the protocol treats as scenery--and two of six pairwise significance calls between selectors flip between adjacent rates without any rank inversion. We call this selector-rate entanglement and trace it to selection itself rather than step size: AdamW update magnitudes track the rate to within 2.2% despite 15.5x gradient-norm differences across arms. A preregistered frozen-scoring ablation (selection scored by the initial student; criterion, budget, and on-policy rollouts unchanged; 12 seeds per cell) shows live scoring adds 3.79+/-1.69 pp of rate sensitivity (p=0.035) while the frozen arm remains significantly entangled (p=0.015): the feedback loop aggravates the phenomenon rather than causing it. Under full fine-tuning at the rates this literature actually uses (1e-6 to 1e-5) the pattern grows: dense itself swings 19.8 pp, the selective arm 49.5 pp, and the verdict ranges from a non-significant +3.6 pp at the published operating point to +34 pp (p=0.005) one notch hotter. On MATH-500 the rate dependence does not reproduce under LoRA, scoping that result, while the ~10 pp cost of selective training does. We prescribe reporting the arm x rate matrix, not a shared-rate column, as a precondition for selector comparisons.
Persistent low retention and completion rates in medications for opioid use disorder (MOUD) have driven the use of machine learning (ML) models to predict retention and identify patients at risk of premature discontinuation. However, the fairness of these models across patient populations remains largely unexplored, raising concerns about their application in treatment decision support. This study systematically assesses algorithmic fairness in ML models for predicting MOUD retention and premature discontinuation and investigates the effectiveness of bias mitigation techniques. Using the cross-sectional Treatment Episode Data Set-Discharges (TEDS-D), which includes treatment episodes for individuals in the U.S. discharged between 2015 and 2019, we trained four ML models to predict premature treatment discontinuation and retention beyond 180 days among individuals receiving outpatient MOUD. We evaluated overall performance and subgroup-level error rates across patient subgroups defined by race, ethnicity, age, and sex, complemented by model explanation analyses. We further assessed bias mitigation techniques and their effects on both fairness and predictive performance. Our findings demonstrate that ML models for MOUD outcome prediction can exhibit subgroup-level performance gaps even when overall predictive performance appears acceptable and that bias mitigation can reduce, but not fully eliminate, these gaps without trade-offs. By demonstrating the importance of fairness-aware evaluation and transparent reporting of subgroup performance, this study provides practical insights for the responsible and context-sensitive use of ML models for risk stratification and care prioritization in MOUD treatment settings.
ZoAQ: Adaptive Zeroth-Order Querying via Query-Reuse Coupling
ZoAQ:基于查询复用耦合的自适应零阶查询方法
Feng, Yangyang, Shu, Yao
Abstract
Zeroth-order optimization (ZOO) estimates updates from function evaluations, making perturbation queries a primary cost. Fixed budgets spend the same number of queries at every step, while adaptive controllers may offset their savings by using additional oracle calls to test estimator reliability. We introduce ZoAQ, an adaptive ZOO method built around query reuse. Rather than discarding past evaluations after each step, ZoAQ makes them useful for both the next update and the decision to query further. This enables adaptive query allocation without extra validation queries. Our analysis characterizes when this agreement identifies an update that supports descent and guides the controller to a sufficient query budget. On synthetic objectives, ZoAQ reduces queries by 43-48% relative to fixed baselines using 1.2M queries. In black-box attacks, it reaches 100% success with 320 and 625 average queries on MNIST and CIFAR-10, respectively. Across four OPT fine-tuning settings, ZoAQ saves 43-46% forward evaluations relative to fixed K=4, with accuracy changes within tasks ranging from -0.018 to +0.010.
Location representations provide mobility models with fundamental information about the spatial position, functional characteristics, and relationships of places. However, existing embeddings are often dependent on mobility observations, unable to represent unseen locations, and weakly constrained to retain geographic distance. This limits their reuse across datasets and mobility tasks. To address these limitations, we propose LE4Mob, an inductive, distance-aware, and geography-derived location embedding framework for mobility modelling. LE4Mob extends contrastive language-location pre-training while introducing a distance-aware regularisation objective that encourages the embedding space to preserve spatial relationships. Pre-trained from geographic context, LE4Mob can encode rich spatial-semantic information and generate embeddings for unseen locations inductively. Its independence from downstream mobility task supervision also makes it transferable across different mobility tasks. We evaluate LE4Mob on individual-level next location prediction and population-level commuter flow generation. Experiments across multiple datasets and study areas show that LE4Mob outperforms strong baselines, with particular advantages in inductive settings and when downstream models rely directly on interactions between location embeddings. These findings demonstrate the potential of distance-aware, geography-derived location representations as reusable foundations for human mobility modelling.
Test-time self-evolving agents improve by reusing past experience, yet sparse-reward trajectories contain failures, loops, and detours, while summaries often omit the state conditions and action dependencies needed for execution. We study executable Walkthrough induction from sparse-reward trajectories: extracting compact, state-conditioned, and verifiable procedures. Our key observation is that delayed credit identifies actions associated with progress but cannot determine whether they produce facts required by later actions. We propose Trace, a credit-guided, dependency-grounded framework that compiles noisy trajectories into executable Walkthrough Memory. It detects progress anchors from rewards and persistent state changes, propagates credit to identify valuable transitions, and estimates action prerequisites from cross-episode success and failure evidence. Backward dependency slicing then traces required facts to their producers, extracting dependency-consistent action chains while removing irrelevant loops and detours. The resulting Walkthroughs encode entry conditions, ordered state--action--effect steps, and completion and failure predicates, supporting reuse, intermediate-state resumption, and programmatic verification. Experiments on J-TTL, WebShop, and ScienceWorld with three open-source LLMs show that Trace consistently outperforms eight test-time learning and memory baselines. Compared with the strongest baseline, it improves average AUC and Final-$3$ by $30.0%$ and $40.5%$, respectively, while using fewer inference tokens. These results show that long-horizon interaction benefits more from state-conditioned executable procedures than from complete trajectories or abstract summaries.
Passively collected mobile phone location data provide large-scale, longitudinal observations of human mobility but do not directly reveal activity purposes. The functional characteristics of visited locations offer useful contextual information, yet their relationship with activity purpose remains uncertain, particularly in mixed-use urban environments. We conceptualise activity pattern mining as an integrated process of representation, clustering, and interpretation, and propose the Activity Chain Encoder (ACE) for the representation stage. ACE is a self-supervised model that combines pre-trained urban embeddings with visit timing and duration and uses a Transformer to model the sequential organisation of stays. It is trained using masked activity modelling and identity-guided contrastive learning without requiring deterministic activity purpose labels. Learned daily representations are aggregated into user-level profiles, clustered, and interpreted through temporal-functional patterns and Census-derived demographic context. Applied to mobile phone app location data from London and compared with three representative methods, ACE supports the identification of six differentiated weekday activity-pattern groups characterised by distinct daily rhythms, urban functional contexts, and demographic associations. These complementary forms of evidence further support the development of empirically grounded activity-pattern personas, establishing a holistic route for deriving behaviourally meaningful population patterns from unlabelled mobile phone location data. The source code for the entire analytical pipeline developed in this study is publicly available at https://github.com/xlwang233/ACE.
Rank Portability Does Not Imply Feasibility Portability: Target-Specific Evaluation of Joint Hardware Constraints
Shu, Wesley
Abstract
Cross-device hardware evaluation often assumes that if architecture rankings transfer across devices, a proxy device can support target-side model selection. We stress-test this assumption for joint latency-energy feasibility across two public architecture families. On NAS-Bench-201, cross-device rank correlations are moderate, while target-comparable feasible-set overlap remains incomplete. A faithful AdaProxy diagnostic substantially improves latency ranking, showing that the observed boundary failures are not simply due to weak adaptation. Exact finite-sample split-conformal analysis also exposes an evidence bottleneck: a finite one-sided 90% threshold requires at least nine calibration observations. We then replicate the phenomenon on 10,000 GPT architectures across 13 HW-GPT-Bench devices. Relative to an RTX3080 proxy, target latency SRCC ranges from 0.951 to 0.996, yet proxy-reuse violation risk ranges from 33.3% to 100% under matched joint constraints. These results show that rank portability, feasibility portability, and target-specific decision support are distinct evaluation objects. Cross-device evaluations should therefore report which target environments actually support the operating point being claimed.
Multi-station multivariate weather forecasting aims to forecast future weather variables at multiple weather stations from historical surface observations. Existing station forecasting models learn statistical dependencies among discrete stations, but lack explicit physical evolution. Meanwhile, PDE-based weather models provide interpretable physical dynamics, yet require continuous fields and upper-air variables unavailable in surface station data. To bridge this gap, we propose StationPDE, a station-oriented surface PDE learning model. StationPDE constructs a terrain-aware continuous surface field from discrete station observations and decomposes its physical evolution into surface wind transport and upper-air inference. Surface wind transport explicitly evolves observable weather variables, while upper-air inference uses learnable horizontal diffusion to approximate the missing influence of unavailable upper-air variables. A parallel data-driven diffusion branch captures complementary motion patterns, and an adaptive router integrates the two forecasts for station-level multivariate forecasting. Experiments on Weather2K and MeteoNet show that StationPDE consistently outperforms state-of-the-art baselines, reducing MSE by about $9.6\%$ on average compared with the strongest baseline. Code and implementation details are available at https://github.com/hnu-vis/StationPDE.
SolarFlowRefiner: Refinement-Aware Flow Matching for Surface Solar Radiation Downscaling
SolarFlowRefiner:面向地表太阳辐射降尺度的精化感知流匹配方法
Srivastava, Udbhav, Racheal, Antonita, Chen, Yiheng, Yu, Runlong, Ye, Xinyue
Abstract
High-resolution surface solar radiation (SSR) is important for solar forecasting and grid operation. However, physically consistent reanalysis products are too coarse to resolve localized cloud-driven variability. In this paper, we study a multisource downscaling task that reconstructs high-resolution SolarCube SSR fields from coarse ERA5 radiative variables and co-registered satellite channels. The task is challenging because a single ERA5 grid cell may contain both sunlit and cloud-shadowed regions. As a result, the missing high-resolution correction can be spatially sharp and inherently ambiguous. One-stage predictors often oversmooth these structures. Post-hoc refinement also introduces a stage-wise mismatch: the generator is optimized independently, even though its output determines the refiner's initial state. We introduce SolarFlowRefiner, a refinement-aware flow-matching framework for SSR downscaling. A conditional FlowMatch generator first predicts a normalized correction to an upsampled ERA5 baseline. The refiner is then trained on prediction-conditioned states between the current FlowMatch output and the target residual. This exposes the refiner to the structured errors produced by the generator. The refinement objective is also backpropagated through the FlowMatch sampler, allowing generation and correction to be jointly optimized for the final reconstruction. Experiments on a day-blocked ERA5--SolarCube benchmark show consistent improvements over standalone generation and post-hoc refinement. More broadly, SolarFlowRefiner provides a general strategy for coupling generative predictors with iterative correctors.
Helix-FNO: Spectral-Domain Operator Learning Coupled with a High-Fidelity Mechanistic Model for Fast Surrogate Simulation
Helix-FNO:谱域算子学习与高保真机理模型耦合的快速代理仿真方法
Zhao, Jiabao, Wang, Chuwei, Yang, Jinxi
Abstract
Mechanistic simulation models of full-scale treatment processes remain the only trustworthy, extrapolative description of the underlying physico-chemical dynamics, yet their runtime is far too slow to support the thousands of forward evaluations that a modern decision engine requires at a 5-minute decision cadence. The standard remedy-surrogate modelling-often produces a network that learns a single solution for a single configuration, so it generalises poorly to new influent profiles, control settings or plant layouts. This paper presents Helix-FNO, a teacher-student architecture that couples a thirty-two-state mechanistic teacher with a Fourier neural operator (FNO) student. The teacher supplies a high-fidelity dataset of input-field-to-solution pairs, curated by Latin-hypercube and uncertainty-based active learning to cover the boundary and overload regimes that matter in practice; the student learns, in the spectral domain, the solution operator itself rather than any single solution, thereby moving from learning one instance to learning an entire family of equations. We give the operator formulation, the spectral convolution definition, the weighted distillation loss and the active-learning criterion, and we analyse the approximation error of a truncated Fourier expansion with respect to the smoothness of the parametric solution manifold. An illustrative study compares Helix-FNO against a physics-informed network and a data-driven recurrent surrogate on accuracy, dataset efficiency and inference latency, and places the methods on a speed-accuracy Pareto front. The resulting operator is three orders of magnitude faster than the mechanistic teacher at millisecond inference, which is precisely the capability required for massive candidate screening and online decision support.
Developing innovative system architectures increasingly relies on advanced modeling and optimization techniques to frame the architecting process and define the corresponding computational problems. For complex System-of-Systems (SoS), high-fidelity multiphysics and multidisciplinary simulations are essential for capturing detailed behaviors. However, their computational expense and the risk of evaluation failures make direct optimization challenging. To overcome these limitations, surrogate-based approaches, like Bayesian optimization, have emerged as effective tools for managing expensive, black-box simulation tasks. This work introduces a hierarchical Bayesian optimization framework that leverages Gaussian process meta-modeling to handle discrete architectural choices, conditional dependencies, and heterogeneous design variables inherent to SoS problems. Results show that the hierarchical formulation improves search efficiency and robustness compared to conventional surrogate-based methods, enabling the exploration of large and structurally diverse design spaces with limited simulation budgets. We apply the approach to an aircraft-based multi-agent system for wildfire suppression, a use case developed within the EU-funded COLOSSUS project that illustrates how SoS principles can coordinate heterogeneous aerial platforms with complementary roles, supporting both sustainable mobility and emergency response missions. Our framework provides a scalable methodology for SoS architecting and model exploration, offering transferable insights for applications in aviation, sustainable mobility, and resilience-oriented system design. By combining hierarchical representations with surrogate-based optimization, this work is among the first practical demonstrations of hierarchical Bayesian optimization applied to real-world SoS problems, advancing both methodology and practice.
Diffusion large language models (dLLMs) offer a compelling alternative to autoregressive models, yet they may expose sensitive training data during denoising. Detecting such usage is challenging because dLLMs lack the efficient one-pass probability decomposition of causal architectures. Existing methods rely on random masking to obtain tractable token-wise detection signals under limited query budgets, but fail to control dependencies among masked tokens. We demonstrate that this token-wise approximation introduces a non-negative structural estimation error, which is theoretically characterized by the cumulative conditional mutual information (CMI) among masked tokens and can obscure subtle memorization signals. This insight suggests that reliable detection requires masked token sets with weak internal dependency. To avoid the prohibitive cost of directly estimating CMI over token combinations, we propose \textit{Independent Token Sampling} (ITS), a query-efficient framework that uses an attention-derived pairwise dependency proxy to approximate the CMI-aware selection criterion. ITS further incorporates a diversity-promoting strategy to improve token coverage across sampling rounds, yielding aggregated token-wise signals that are less affected by dependency-induced approximation error. Experiments on multiple datasets show that ITS consistently outperforms state-of-the-art baselines across different models and datasets, achieving an AUC improvement of 0.18 on the ArXiv dataset while maintaining strong performance under limited query budgets. The code is available at https://github.com/Chrisqcwx/DLLM-MIA .
GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training
GRRR:解码器大语言模型后训练中的重塑、旋转与路由几何学
Qi, Jianing, Tang, Hao, Zhu, Zhigang
Abstract
We study how post-training changes the weights of Large Language Models (LLMs) relative to their pretrained weights. Across 12 post-training chains with supervised fine-tuning (SFT) and reinforcement learning (RL), we express each weight update in the pretrained matrix's singular value decomposition (SVD) frame. This decomposition separates the changes of three geometrically distinct components: diagonal values, which reshapes singular values; off-diagonal values, which rotates the coupling between pretrained input and output directions; and null-space values, which routes outside the matrix's original nonzero SVD core. On a math evaluation suite, we find that removing the diagonal component usually preserves most of the gains from post-training. These results suggest that post-training gains are carried primarily by reconfiguring and extending pretrained pathways rather than by substantially changing singular values of pre-trained models.
Methods for addressing safety drift in fine-tuned Large Language Models (LLMs) are scattered across incompatible implementations, lifecycle stages, and evaluation protocols, making them difficult to adopt and compare. We introduce SafeTune, a source-available library that unifies four intervention paradigms: post-hoc weight recovery, safety-constrained fine-tuning, gradient-based unlearning, and inference-time steering, alongside shared interpretability, evaluation, and deployment utilities. SafeTune provides a consistent configuration-driven workflow while preserving the distinct inputs and intervention points each paradigm requires. Its modular registry supports new methods, benchmarks, judges, models, and fine-tuning domains without redesigning the surrounding pipeline. We demonstrate SafeTune through controlled comparisons and finance and medical deployment case studies, showing how it characterizes safety drift, evaluates feasible interventions on common refusal-behavior and capability evaluations, and supports calibrated or layered mitigation.
A Comparative Framework for Evaluating Foundation Models on Tabular Data: A Case Study in Healthcare
Delouee, Majid Lotfian, Veld, Sjors G. J. G. In 't, Schut, Martijn C.
Abstract
Tabular data is the most common format in clinical practice, encompassing laboratory results, medication records, diagnostic codes, and patient demographics. As foundation models for tabular data have grown in number and variety, a practical question has become harder to answer: which model should a clinician or data scientist actually choose for a given task, and why? Existing surveys catalogue what these models can do, but they stop short of providing a structured way to compare them against the specific demands of a real application. We introduce \system{}, a comparative evaluation framework that scores and ranks tabular foundation models (TFMs) across six clinically meaningful dimensions: how well a model generalizes to new datasets, how effectively it protects patient privacy, how much data it needs to perform well, how it scales with growing datasets and feature spaces, how interpretable its predictions are to clinicians, and how fairly it performs across patient subgroups. Each dimension is broken down into measurable sub-components, and groups of sub-components can optionally be combined into supplementary compound scores, called super-metrics, that provide a diagnostic view of how a model performs across several dimensions simultaneously. To show how the framework works in practice, we apply it to two healthcare use cases, screening for iron deficiency and predicting heart failure, demonstrating how the same set of metrics leads to different model rankings depending on what matters most in each clinical context. We also provide a taxonomy of 45 TFMs organized by their underlying architecture, which serves as a reference for researchers and practitioners looking to navigate this rapidly expanding field.
From Latent Biomarkers to Clinical Rules: Embedding-Guided Rule Mining and Attribution-Based Translation for Interpretable Tabular Learning
从潜在生物标志物到临床规则:面向可解释表格学习的嵌入引导规则挖掘与基于归因的翻译方法
Delouee, Majid Lotfian, Ayoobi, Hamed, Veld, Sjors G. J. G. In 't, Schut, Martijn C.
Abstract
Clinical decision support tools are most useful when accurate predictions are accompanied by understandable explanations. Rule-based models provide transparency, but rules derived directly from raw clinical measurements may miss patterns arising from interactions between multiple variables. We present a four-step pipeline that mines decision rules in the latent space of an FT-Transformer and translates them back into measurable clinical features. Embedding dimensions that consistently separate patient groups are treated as latent biomarkers, rules are mined using small decision trees, and selected rules are translated using gradient-input saliency and CLS attention attribution. We evaluate the framework on six public clinical and population health datasets at four embedding dimensions. Translated rules outperformed raw-feature rules in five of six datasets, with mean AUROC gains ranging from 0.04 to 0.23. On the heart disease dataset, embedding-space rules reached 0.98 AUROC, but translation reduced this to 0.72, showing that high-performing latent rules cannot always be represented by simple raw-feature conditions. These results show that latent-space rule discovery can uncover predictive patterns while translating them into clinically measurable features that can be evaluated by clinicians.
The Limits of Speculation: Bounding Speculative Decoding in Mixture-of-Experts
Amankulov, Aidar, Mamatin, Denis
Abstract
Speculative decoding in Mixture-of-Experts (MoE) models faces the problem of unstable verification cost caused by input-dependent expert loading. To study the physics of this process, we formulate speculation-budget selection as an offline Stochastic Shortest Path (SSP) problem over reference sequences and build a diagnostic Oracle that uses counterfactual simulation to account for MoE verification cost. A detailed analysis of the Oracle's decisions on the Qwen3-Coder and EAGLE-3 pairing, in the space of marginal deltas (Delta Space), shows that rejected candidates form a strict linear boundary. This result demonstrates that a complex global optimization is locally governed by a necessary condition balancing marginal cost against expected progress ($\frac{\Delta \mathbb{E}[Cost]}{\Delta \mathbb{E}[a]}$), providing a rigorous mathematical reference point for designing future adaptive online heuristics.
KV-cache eviction methods decide which tokens to keep but not whether to evict at all, so a benchmark mean can hide a class of inputs on which compression drives accuracy from 99\% to 0\%. We reframe eviction as a per-input admission decision and show that inputs separate into a capacity-bound class, where eviction is catastrophic at every budget, and a dilution-prone class, where eviction is safe or beneficial. A single label-free scalar computed from prefill attention, the early-to-late drop in pairwise top-$k$ head agreement, predicts this class before any decoding. PAGE thresholds this drop: it applies any base evictor when the drop is large and retains the full cache otherwise, with no training and no accuracy labels. The drop orders inputs by eviction safety consistently across four architecture families, and a per-model unlabeled pilot of about 100 inputs recalibrates the threshold for a new family. Used as a safeguard, PAGE cuts the harm rate on the capacity-bound regime from 0.75 to 0.026, a 29 $\times$ reduction, across four evictors, four models, and two benchmarks, turning a 99\% to 0\% collapse into a flat 89\% without retraining the evictor. The gate is inert wherever eviction is already safe, and the capacity-bound class it protects is a small, identifiable minority of inputs, so the benefit is a targeted safety gain rather than an average one. PAGE is a per-input safeguard, not a compressor: realized compression is $1.8 - 3.4 \times$ (mean 2.9$\times$) against a nominal 16$\times$ budget and decays toward unity by batch 16 under static provisioning, and a trained evictor wins at matched memory.
Key-value (KV) caching is essential for efficient autoregressive large language model (LLM) inference, but the cache grows linearly with context length, increasing storage and decoding costs. KV cache compression mitigates this cost by retaining only a subset of cached tokens. This challenge is particularly important for multi-step LLM agents, where a query expands into trajectories of reasoning, tool interactions, and retrieved observations. Existing pruning methods typically treat the cache as a flat token stream and rank tokens by recency or attention saliency. This creates a mismatch between the unit of compression and the unit of reasoning: token-level pruning removes individual entries, whereas useful information in multi-step agents is often organized into reasoning steps with uneven and delayed importance. Consequently, an early observation or intermediate decision may receive little recent attention yet remain essential for later evidence synthesis. We term this failure mode Reasoning Continuity Disruption.These observations motivate KV cache compression that jointly considers token- and reasoning-step-level information. StepKV addresses this goal by treating reasoning steps as first-class retention units. It associates cache entries with their generating steps, estimates step utility from trajectory-derived signals, and combines this utility with token-level saliency. The resulting scores globally rank prunable tokens, from which StepKV retains the top-scoring entries under a target budget. StepKV thus provides a step-centric perspective for agent KV cache compression. Across multi-hop QA and long-horizon web reasoning tasks, StepKV sustains accuracy under low KV budgets where token-level baselines degrade sharply, offering a more robust efficiency-accuracy trade-off for multi-step agent inference.
Uncertainty and Business-Aware Remaining Useful Life Estimation for Semiconductor Manufacturing
Frizzo, Davide, Borsatti, Francesco, Susto, Gian Antonio
Abstract
Semiconductor manufacturing relies on tightly interconnected components, so early identification of the assets most likely to fail is essential to prevent a single breakdown from disrupting the entire production pipeline. Maintenance planning must therefore balance unexpected failures against prematurely interrupted operating life. We present a Predictive Maintenance (PdM) framework combining Deep Learning (DL) sequence models and Simoultaneous Quantile Regression (SQR) for uncertainty-aware Remaining Useful Life (RUL) estimation and risk-aware maintenance decisions. Several architectures are compared on ion-milling data from the 2018 PHM Data Challenge (PHM18), including architectures based on State Space Models (SSM), using prediction and business metrics: Unexpected Breaks (UB), Unexploited Lifetime (UL), and a cost-weighted objective. Diagonal State Spaces (S4D) delivers the best Remaining Useful Life (RUL) estimates across quantiles and, relative to Preventive Maintenance (PvM) baselines, substantially lowers business cost by avoiding systematically early interventions. The results support uncertainty-aware, cost-sensitive maintenance planning in semiconductor production.
Automatic synthesis of analog and Radio Frequency (RF) circuits is an emerging area that requires an efficient circuit modeling method. In recent years, Machine Learning (ML) solutions have played a promising role in this regard. However, many existing ML approaches require separate training data for each circuit topology, even when a single circuit component is added or removed. In addition, they overlook circuit topology information, which limits their ability to capture complex component interactions. Furthermore, they rely on fully connected neural networks with flat feature representations, which require substantial amounts of training data. In this work, we propose an open-source topology-aware RF circuit modeling method, TARGet. Our model considers the circuit at two levels: sub-circuits and the overall circuit topology. At the sub-circuit level, TARGet leverages S-parameter representations to capture sub-circuit behavior rather than relying on individual circuit components, providing a reusable behavioral abstraction for RF building blocks. Moreover, TARGet explicitly incorporates circuit topology information into the model, enabling it to learn across multiple topologies. TARGet introduces a novel fusion-based architecture that integrates Graph Neural Networks (GNNs) and sub-circuit connectivity-aware neural networks to improve data efficiency. Experimental evaluation across multiple RF circuit topologies demonstrates that TARGet achieves sub-1% prediction error while reducing the required training data by up to 35.5x compared to state-of-the-art (SOTA) approaches. Furthermore, TARGet achieves 9.7x higher prediction accuracy under a strict 1% error threshold relative to SOTA models. A held-out matching-network evaluation further demonstrates zero-shot transfer to an unseen sub-circuit topology, where TARGet reduces NMAE by up to 45%.
The rapid advancement of Internet of Things (IoT) technology has led to the widespread deployment of smart, interconnected devices across a range of domains. However, this expansion has also resulted in a substantial increase in network traffic, creating more opportunities for malicious actors to launch cyberattacks and compromise sensitive information, thereby increasing the need for effective anomaly detection. The state-of-the-art in anomaly detection has predominantly focused on point anomalies. In contrast, the detection of collective anomalies remains relatively under-explored in the literature. In this paper, we introduce Unsupervised Graph Collective Anomaly Detection (UGCAD), a novel frame- work designed to identify collective anomalies in IoT network traffic. Unlike many existing methods, UGCAD operates on graph-structured data without any prior knowledge of group labels or membership. It leverages a variational graph autoencoder (VGAE) to learn the graph representation, which is subsequently used to enhance a clustering algorithm for effective grouping of nodes. To detect collective anomalies, clusters identified as normal are first aggregated and refined, after which anomaly scores are applied to detect collective anomalies. Extensive experiments conducted on the CICIoT2023 and ToN-IoT network datasets demonstrate the effectiveness of UGCAD in both clustering and collective anomaly detection (CAD). Furthermore, comparative evaluations against several traditional and state-of-the-art clustering-based CAD approaches confirm the superiority of UGCAD in accurately detecting collective anomalies.
Predicting reaction yield from molecular structure and reaction context can cut experimental trial-and-error and speed up condition screening in synthetic chemistry. Recent methods for this task use learned representations such as graph neural networks or Transformer encoders over reaction SMILES (Simplified Molecular Input Line Entry System), but these approaches carry heavy preprocessing overhead and can break when input formatting is inconsistent. We propose MFP, a reaction yield prediction method built on role-aware Morgan fingerprints where count-based circular fingerprints are computed for each reaction component, aggregated by chemical role (reactant, reagent, product), and combined with transformation-sensitive difference features into a fixed-length reaction descriptor fed to a feed-forward neural regressor. We test MFP against state of the art methods such as YieldBERT (with and without data augmentation) and GNAN (graph neural network) on the Suzuki-Miyaura and Buchwald-Hartwig benchmarks using a shared preprocessing and evaluation protocol. MFP reaches R2 = 0.878 on Suzuki-Miyaura and R2 = 0.969 on Buchwald-Hartwig while training an order of magnitude faster than graph- or Transformer-based alternatives. A formal complexity analysis confirms that MFP folds all representation cost into a one-time preprocessing step, removing the per-epoch message-passing overhead that graph methods carry. An ablation over fingerprint radius and folded vector length shows that radius-2 representations at nBits =2048 give the best balance of accuracy, speed, and cross-split stability on both datasets. These results establish MFP as an effective, reproducible, and efficient baseline for reaction yield prediction.
Multiple latent orderings better predict language model preferences
多重潜在排序能更好地预测语言模型的偏好
Chawla, Aviral, Thompson, William H. W., Young, Jean-Gabriel
Abstract
Language models are frequently employed in settings where they are asked to make value judgments and choices. These observed choices often exhibit intransitivity: A model may prefer item $A$ to $B$ and $B$ to $C$, while also preferring $C$ to $A$. Existing work that models LLM preferences treats such inconsistencies as sampling noise around a single latent ordering. We instead propose that intransitivity reflects the aggregation of multiple latent, internally consistent orderings. We first show that observed inconsistencies cannot be explained by a single ordering under any monotone link function. We then introduce a noise-augmented mixture Bradley-Terry (MBT) model that infers latent preference components from repeated pairwise comparisons. Across seven models and four tasks, a mixture of orderings often explains structural inconsistencies better than single-utility models. We find that aggregate preferences often hide underlying preference heterogeneity. A case study on Moral Machine dilemmas shows that models which disagree on aggregate orderings can still share latent components. Together, these results suggest that LLMs reflect plural preferences. Alignment and evaluation pipelines that treat LLM preferences as a single function, therefore, risk averaging over coherent orderings that different users may endorse differently.
Industrial Kinematic Trajectory Model (IKTM): Coordinate-Free Autoregressive Generator
工业运动学轨迹模型(IKTM):无坐标系自回归生成器
Amiri, Max, Eyers, David
Abstract
Mobility simulation supports logistics, safety, and communications planning in industrial environments such as ports, mines, and airports. Existing trajectory models, however, rely on absolute coordinates, road-network tokens, or semantic zones: representations that are site-specific and not well suited to unstructured industrial terrain. We introduce the Industrial Kinematic Trajectory Model (IKTM), a coordinate-free trajectory generator that represents industrial vehicle motion through kinematic sequences (speed and heading change) with no absolute spatial reference. IKTM uses an autoregressive causal transformer with probabilistic mixture heads and extends our prior coordinate-free Markovian model with deep sequence modelling and an explicit duration-conditioning signal. Trained on one site and evaluated zero-shot on three unseen sites, it matches the small-turn shape of the empirical turn-rate distributions of held-out telematics; across all four sites, the per-site mean Jensen-Shannon divergence over 100 sampling seeds spans approximately 0.035-0.050 bits under an oracle-length duration-target protocol and approximately 0.032-0.042 bits under a fully zero-shot prior-length protocol, with similar ranges whether out-of-distribution (OOD) duration targets are drawn from each held-out site's empirical length distribution or from the Site A prior. Both protocols stay above the metric's sampling-noise floor (<=0.0049 bits). Paired by site, the oracle-length values are 5.3-6.0x lower than those of a re-implementation of our prior Markovian model under the same 1 Hz protocol. Termination is duration-conditioned rather than spatial: rollouts stop on 100% of trials with a length-tracking error of +0.0 +/- 0.0 s against the sampled target (100 of 100 exactly on target; N=100, T=0.2, untouched in-distribution test split).
World models trained via pixel reconstruction can struggle in visually complex environments, where irrelevant information dominates the objective and distract the model from information relevant to planning and control. We present Contrastive World Models, an approach for learning latent dynamics models without pixel reconstruction. Building on Dreamer, we replace observation reconstruction in the standard world model objective with a Deep InfoMax-like lower bound that maximizes the mutual information between state-action sequences and local patch features of future observations, encouraging state representations to retain information that is predictive of the future without requiring the model to reconstruct visually irrelevant details. We evaluate our approach in small-scale experiments across three settings of increasing visual complexity. Our method matches Dreamer and a momentum prediction baseline in the default setting, and substantially outperforms both once distractors or natural video backgrounds are introduced, while also training more efficiently by removing the pixel decoder entirely. Our approach is general and makes minimal assumptions beyond access to state-action sequences and future observations. These results suggest that contrastive, infomax-based objectives are a principled and promising direction for building world models that are robust to visual nuisance factors, a property particularly relevant for transferring model-based RL agents to the real world.
OpenBlock: Constructive and Verified Content Generation for Adaptive Tile-Matching Games
OpenBlock:面向自适应方块消除类游戏的构造性可验证内容生成
Jun, Jiang
Abstract
Tile-matching puzzle games serve hundreds of millions of players, yet the content-generation algorithms that decide which pieces to present at each turn remain proprietary, and no open platform exists for studying adaptive difficulty in this genre. We present an adaptive tile-matching platform whose central algorithmic contribution is a dual-track content-generation architecture: a deterministic rule-based generator that is always available, and an optional learned generator, both subject to a common verification gate that establishes, by exhaustive sequential-placement search, that every delivered piece set is fully placeable so the learned track can never degrade the constructive-feasibility guarantee of the rule track. A self-play reinforcement-learning placement agent, supervised by auxiliary tasks that expose per-shape placeability to shared representations, is used to diagnose the game's dominant failure mode: at high board fill, long-bar pieces lose the majority of their legal placements. Across 234,000+ self-play episodes the agent reaches a 35.6\% win rate (median score 4,200), and controlled simulation shows that at board fill rates of 70--75\%, 33--56\% of long-bar pieces have no legal placement, while spawn difficulty distributions are statistically indistinguishable between won and lost games---evidence that board-state degeneration, not content difficulty, drives late-game failure. Head-to-head ablations show that per-shape placeability supervision---not aggregate difficulty features---drives the representation gain, and a 14-day online gray rollout (48,000 players; sample-ratio verified, CUPED-adjusted) lifts day-1 retention by 1.8 percentage points and session duration by 7\% over the rule track alone, quantifying the neural track's asymmetric upside in live play.
Computer-use agents (CUAs) have made rapid progress in completing complex tasks through graphical user interfaces, yet post-training centered on task success alone does not induce reliable safety behavior. A reliable CUA must condition its execution on risk: it should complete ordinary benign tasks, avoid environmental hazards and continue when a safe completion path remains, and refuse when the goal is harmful or no safe path exists. To learn this conditional policy, we develop Safety and Capability Optimization for Policy Execution (SCOPE), which jointly post-trains a CUA for task-execution capability and safety-aware decision making. To provide aligned training data for this joint objective, we further introduce SCOPE-Gen, an automated pipeline that synthesizes verifiable capability tasks and converts them into paired environment-risk variants while preserving their original goals. Using the resulting tasks, we construct SATraj-OS, a trajectory dataset comprising capability demonstrations, safe continuations, and explicit refusals. SCOPE first learns from all three trajectory types through supervised fine-tuning and then further improves task completion through online reinforcement learning. Starting from Qwen3.5-9B, SCOPE-RL achieves a 54.17% task success rate on OSWorld and a 64.30% attack-avoidance rate on OS-BLIND, yielding the best aggregate capability--safety score of 58.80% among the evaluated agents. Ablations reveal asymmetric but complementary roles for the two forms of safety supervision: refusal trajectories account for most of the attack-avoidance gain, whereas risk-handling trajectories preserve greater task utility at comparable attack-avoidance levels.
Gaussian Process Decorrelation for Spatiotemporal Deep Learning-Based Snow Water Equivalent Prediction
基于高斯过程去相关的时空深度学习雪水当量预测方法
Fenster, Colin, Marshall, Adrienne, Bandyopadhyay, Soutir, McKenzie, Daniel
Abstract
In the Western United States, snowmelt is essential to the agricultural industry in addition to being a key source of municipal drinking water. Consequently, accurate snowpack forecasting is critical for water policy and management. Automated Snow Telemetry (SNOTEL) stations provide accurate daily measurements of snow water equivalent (SWE) that exhibit strong correlations in space and in time. We tackle the problem of predicting future SWE values across the SNOTEL network. Specifically, we use a Gaussian Process-based linear transformation to remove spatial correlations before training a long short-term memory (LSTM) neural network on the decorrelated SWE data. This approach allows the LSTM to learn a clean temporal signal at each station. We show that this separation of spatial and temporal components yields better predictive success than multiple baseline models. Furthermore, we incorporate conformal prediction to quantify uncertainty in the resulting SWE forecasts, providing a distribution-free approach to illustrate a potential framework for establishing predictive intervals for spatiotemporal data. Together, accurate point forecasts and distribution-free uncertainty quantification provide a framework for SWE accumulation forecasting on subseasonal scales or projecting SWE with future data while motivating and supporting future work in predicting a large-scale, spatiotemporally complete SWE map.
CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds
CleanScore:基于阴性对照与敏感性边界的黑盒基准测试审计
Opoku, Jeffery, Banahene, David
Abstract
Public benchmark scores may reflect skill, prior exposure to the questions, or both, and for most models the training data are unknown. We present CleanScore, a black-box audit using scored outputs only. Each benchmark question becomes a parent item with one public form and two independently written fresh forms preserving its numbers, facts and answer. The audit reports an interval for the public-form advantage rather than a verdict, and a private negative-control bank with an explicit transport radius separates exposure from ordinary form mismatch. A registered controlled-exposure experiment detects planted exposure and stays quiet under fresh-form exposure. A registered audit of five open models on 200 GSM8K and 200 ARC-Challenge items finds no exposure-consistent advantage, bounding surface-form inflation below five points. Registered positive controls then bound what such a null can mean. Leaking an item raises accuracy on paraphrases the model never saw almost as much as on the leaked wording, leaving 52% to 110% of the effect invisible to a paraphrase audit. On ARC a planted 49-point advantage shows an observable gap of -0.020, and about 20 points survive rewriting stem and options, across four training seeds. A surface-form null bounds far less than the phrase contamination audit implies.
DPTM-DT: Dual-Pretrained Transformer Multitask Representation Learning for Drug-Target Prediction
DPTM-DT:面向药物-靶点预测的双预训练Transformer多任务表示学习
Kong, Ge
Abstract
Drug-target relation prediction supports candidate screening, drug repositioning, and mechanism analysis. Existing models often use incomplete drug or protein representations, model cross-modal interactions shallowly, or train affinity regression and interaction classification separately, although these tasks describe closely related views of the same drug-target pair. This paper presents DPTM-DT, a dual-pretrained Transformer framework for multitask drug-target prediction. DPTM-DT combines GROVER molecular graph embeddings, ESM protein language-model embeddings, and CTD physicochemical descriptors, then exchanges drug-target information through bidirectional cross-modal attention. A shared pair representation is used for continuous affinity regression, high-affinity binary classification, and six-level affinity classification. Experiments on Davis and KIBA cover random 80/20 and DeepDTA-style standard splits. On the random 80/20 split, DPTM-DT achieves MSE/CI values of 0.193/0.917 on Davis and 0.120/0.918 on KIBA. It also reports binary AUPR/MCC values of 0.727/0.654 and 0.798/0.689, and six-class Macro-F1/Top-2 values of 0.800/0.932 and 0.815/0.962 on Davis and KIBA, respectively. Across the reported regression, binary classification, and multiclass classification settings, DPTM-DT achieves the best overall performance among the compared methods. Results under the standard split show the same relative trend. Ablations indicate that dual target representation, gated fusion, and cross-modal attention each contribute to the final performance. Code and supplementary materials are available at: anonymous.4open.science/r/DPCM-DT-74E0.
Adaptive Physics-Informed Neural Networks for the Blasius Boundary-Layer Problem
面向Blasius边界层问题的自适应物理信息神经网络
Endalew, Mehari Fentahun, Zhang, Xiaoming John
Abstract
Physics-informed neural networks (PINNs) provide a mesh-free approach for solving differential equations, but their performance can depend strongly on loss weighting, collocation placement, and optimization strategy. This study develops an adaptive PINN framework for the Blasius boundary-layer equation using gradient-norm-based adaptive loss weighting, nonuniform and residual-based collocation, and sequential Adam--L-BFGS optimization. In the representative run using the architecture $[1,100,100,1]$, the model predicts $f''(0)=0.3320762918$, compared with the high-accuracy benchmark $0.332057336215$, giving an absolute error of $1.896\times10^{-5}$. The final weighted loss is $6.789\times10^{-8}$, and the predicted stream-function, velocity, and shear profiles agree closely with an independent numerical boundary-value solution. A separate full-training architecture study shows that the two-hidden-layer model achieves the smallest wall-shear error among the four tested architectures, $1.629\times10^{-6}$, whereas the deepest network attains the smallest weighted objective but a substantially larger wall-shear error. Compared with the previously reported PINN value $f''(0)=0.33165$, the representative run reduces the wall-shear error by approximately a factor of $21.5$. The results show that the combined adaptive training framework can achieve high accuracy for the Blasius problem and that weighted loss alone is insufficient for identifying the most physically accurate PINN. Because the adaptive components are applied jointly, their individual contributions cannot be isolated from the present results and would require a controlled ablation study for separate assessment.
Spatial transcriptomics (ST) provides spatially resolved gene expression profiling but remains expensive, motivating the prediction of ST from histology images. Generative models have emerged as a mainstream paradigm for ST prediction due to their ability to model the conditional distribution of gene expression and capture its inherent stochasticity. However, these methods typically treat genes as independent prediction targets and overlook the intrinsic gene-gene interactions in biological systems, which limits their ability to preserve biologically meaningful co-expression patterns. We argue that gene-gene interactions, which reflect shared pathways and regulatory mechanisms, are essential for generating numerically accurate and biologically coherent ST profiles. In this paper, we propose CorrFlow, a correlation-guided flow matching framework for histology-to-ST prediction that explicitly models gene-gene dependencies through two complementary mechanisms. First, we introduce an annealed masked flow matching strategy, where subsets of genes are progressively masked following a timestep-dependent annealing schedule, encouraging the model to infer masked genes conditioned on the remaining genes and promoting joint conditional modeling beyond per-gene marginal estimation. Second, we devise a gene graph-regularized optimization scheme that integrates prior knowledge from the STRING database and data-driven co-expression estimated by WGCNA to construct a gene affinity graph, which enforces both local consistency and global smoothness in the predicted expression. Extensive experiments across 12 datasets show that CorrFlow achieves the best average PCC and HPCC among evaluated methods, leading to more biologically coherent ST predictions.
WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction
Gopakumar, Arin, Pannozzo, Marco
Abstract
Machine learning is being increasingly used to predict where active wildfires will burn the following day, helping inform evacuation boundaries and containment lines. Most models are evaluated using Average Precision (AP), which summarizes performance across all decision thresholds, although acting on a forecast requires choosing one. We benchmarked five discriminative architectures and one generative model on WildfireSpreadTS using a shared evaluation pipeline and two input configurations. We found that model rankings varied depending on whether performance was measured by AP or by threshold-dependent metrics like F1 and IoU. The highest-AP model flagged 4 to 5 times the area that burned and ranked fifth of six on F1 and IoU, and the most recall-heavy model flagged 16 to 23 times. Models with more usable predictions had AP scores 24 to 37 lower. Across architectures, we identified three distinct prediction profiles: over-predicting, balanced, and under-predicting, which AP alone could not distinguish. Expanding the input from 7 to 23 channels changed AP by 0.03 on average, against a 0.21 to 0.24 spread across architectures. These results show AP alone can favor models whose predictions are poorly suited for operational wildfire forecasting.
SegTSim: A Big Data Driven Segmented Temporal Simulation Framework for Heterogeneous Multivariate Systems
Li, Xinhang, Geng, Chenxi, Sun, Yujia
Abstract
Heterogeneous multivariate time-series systems exhibit segment-specific nonlinear dynamics that challenge monolithic forecasting architectures. We propose SegTSim, a big-data-driven segmented temporal simulation framework that integrates segment-specific elasticity modeling with adaptive min-gating, dynamic production relocation optimization, multi-factor data fusion with exchange-rate propagation, and a deep ensemble validation pipeline. The framework is validated on US--Japan automotive trade data from USITC repositories spanning 2015 to 2025, comprising approximately 13000 annual records. Under a 25% perturbation scenario, Japanese import volume declines by 20.4% to 0.93 billion USD, while all output variables maintain coefficients of variation below 3.5% across 1000 ensemble inference runs.
Enterprise analytics agents solve long-horizon tool-use problems over distributed business data, requiring retrieval, reasoning, API calls, code execution, and adaptation to intermediate observations. Supervised fine-tuning (SFT) calibrates tool syntax and teacher-supported behavior, whereas reinforcement learning (RL) can explore reward-supported behaviors beyond demonstrations; applied uniformly, however, RL can perturb already-calibrated skills. We study how to balance SFT and RL under production-mirroring beta APIs. We observe that, in our controlled experiment, checkpoint trajectories retrospectively separated into three regimes: Imitation, where SFT captured reliable teacher behavior; Lift, where both stages helped; and Discovery, where useful reward-observable behavior lay outside reliable teacher support. We leverage this prospectively, using teacher support and reward-observable headroom to route features to SFT only, SFT then RL, increased RL allocation, or further environment development. Across 18 subsequent feature-specific experiments, the diagnostic predicted 15/18 observed trajectories. On GPT-OSS 120B, targeted SFT then RL produced positive point estimates on 7/8 advertiser skills relative to a frontier Control; five positive gains had paired 95% confidence intervals excluding zero, while one skill had a confidence-supported regression. The largest gain was non-disclosure (+11.27 points; 95% CI [+9.72, +12.82]). A separate SME audit surfaced that targeted RL reduces standard leakage from 11.8% to 2.9% and adversarial leakage from 22.9% to 6.8% relative to SFT while preserving actionability (86.2% to 85.7%). In a matched uniform-versus-targeted comparison with shared rewards and optimization, targeted RL improved the seven-skill mean delta from +1.62 to +3.57 while using 43% less incremental RL compute.
EvoRank: LLM-Guided Evolution of Multi-Objective Learning-to-Rank Pipelines
Patel, Rayhan, Patel, Shabaz
Abstract
We present EvoRank, an open autonomous ranking engineer: an LLM-guided evolutionary loop that discovers complete Learning-to-Rank pipelines (features, models, losses, ensembles) for multi-objective e-commerce search. On the Expedia ICDM 2013 dataset, with relevance, conversion, and revenue as competing objectives, three independent runs each converge within 50 iterations (about ten dollars) on interpretable pipelines that beat an Optuna-tuned LambdaMART on 60k held-out queries, an advantage that persists at full data scale and places in the top 6 percent of the original competition. A first campaign, evolving only training objectives, builds the central design rule: it appeared to work on its selection fold (the small dataset it uses to pick winners) while a transfer audit, re-scoring winners on held-out data, showed the gains were almost entirely fitness noise (the randomness of its own scoring), and neither seeded domain knowledge nor richer diagnostic feedback changed what transferred. The deciding quantity is measurable in advance: search-space headroom relative to fitness noise. We package this as a headroom gate that predicts, before any LLM spend, whether the loop will pay off, and we release the system, the auditing tools, and a catalog of failure modes with their guardrails, so teams can apply the procedure to their own ranking stacks.
Dissecting Hierarchical Reasoning Models: A Mechanistic Study
Rodrigues, Leo Raphael, Kang, Jian
Abstract
We study Hierarchical Reasoning Model (HRM), a representative hierarchical Transformer-based latent reasoning model with many variants, on Sudoku, Maze, and ARC-AGI-2. We mechanistically understand how HRM reasons and what information it encodes. Our analyses compare HRM against Transformer baselines with and without recurrent modules, apply causal interventions on recurrent states, and utilize linear probes against random-direction ablations, as well as sparse autoencoders with feature ablations. Our results reveal several key findings: recurrent models outperform one-pass baselines, while single-state recurrent Transformers are comparable to HRM. State interventions further show that the causal contributions of the high- and low-level states vary across task-specific checkpoints and inference stages. Selected task variables are linearly decodable from the recurrent states in HRM, yet ablating probe directions produce effects comparable to random controls. SAE ablations yield larger behavioral changes than probe-direction ablations. However, top-ranked SAE features show no stable advantage over size-matched random subsets at larger ablation sizes or across tasks; the same pattern persists in a Sudoku control with within-step BPTT. Together, we characterize that HRM is essentially implementing constraint-aware iterative refinement on a puzzle-specific solution state, in which the functional contributions of components at different levels vary without relying on a compact, causally important feature set. These results highlight the necessity of studying the different working mechanisms and the importance of developing mechanistic interpretability techniques better suited for latent-space, recursive reasoning models.
Transformer-based large language models often suffer from inter-layer parameter redundancy, where functional transformations are redundantly learned across network depths. We propose CS-MoE, a novel Transformer architecture featuring cross-layer expert sharing to address this inefficiency. Deviating from the widely used Mixture-of-Experts (MoE) architecture that terminates each Transformer block with layer-isolated experts, CS-MoE combines layer-independent experts with concurrent access to a centralized, globally shared expert pool. This \textit{Global Experts Sharing} mechanism enables elastic control over token-level parameter activation and computational consumption (FLOPs). Experiments demonstrate that CS-MoE achieves lower perplexity than equal-scale dense Transformers while activating only 55\% of parameters. Furthermore, its performance scales monotonically with an increased number of activated experts and approaches MoE counterparts that consume more FLOPs by expanding the shared pool with a fixed FLOPs budget. CS-MoE also establishes a flexible Pareto frontier between computational cost and model capacity, offering an efficient alternative for computation-constrained environments.
Reliable prediction of vehicle dynamics is essential for smart driving applications such as autonomous control and advanced driver-assistance systems. Off-road vehicles used in agricultural and construction settings are particularly prone to nonlinear behavior, including bifurcations and chaotic motion arising from intermittent loss of tire--road contact. Predicting such dynamics is challenging because it requires resolving both smooth nonlinearities and the discontinuous switching associated with contact loss. In this work, we investigate the feasibility of reservoir computing (RC) -- specifically an echo state network (ESN) -- for data-driven prediction of a jumping quarter-car model. The reservoir is trained on time-series data from a small number of points and evaluated on its ability to reconstruct bifurcation diagrams, phase-space attractors, and time trajectories across periodic and chaotic regimes. The trained reservoir qualitatively reproduces the period-doubling route to chaos, captures the geometric structure of periodic and chaotic attractors. These results demonstrate that reservoir computing is a feasible data-driven predictor of nonlinear dynamics in a practical, non-smooth vehicle system.
Chinese Translation
车辆动力学的可靠预测对于自动驾驶控制和高级驾驶辅助系统等智能驾驶应用至关重要。用于农业和建筑场景的非公路车辆特别容易出现非线性行为,包括由轮胎—路面接触间歇性丧失引起的分岔和混沌运动。预测这类动力学具有挑战性,因为它需要同时处理光滑非线性以及与接触丧失相关的不连续切换。本研究探讨了储备池计算(Reservoir Computing, RC)——具体为回声状态网络(Echo State Network, ESN)——用于跳跃式四分之一车模型数据驱动预测的可行性。储备池在少量点的时间序列数据上进行训练,并评估其重构分岔图、相空间吸引子以及周期和混沌状态下时间轨迹的能力。训练后的储备池在定性上重现了通向混沌的倍周期路径,并捕捉了周期吸引子和混沌吸引子的几何结构。这些结果表明,储备池计算是针对实际非光滑车辆系统中非线性动力学的可行数据驱动预测方法。
The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families
Twagirayezu, Leonard, Mitra, Prasenjit
Abstract
Quantization enables deployment of large language models on resource-constrained clinical edge devices, but its effect on clinical accuracy and safety remains understudied. We evaluate five 7-8B parameter models at FP16, GPTQ-INT8, and GPTQ-INT4 precision across five benchmarks: MedQA, MedMCQA, Med-HALT, a risk-stratified sample of HealthBench, and MedSafetyBench. The study jointly varies quantization bit width, model family, and clinical task type, with explicit risk stratification and safety measures. INT8 GPTQ is universally safe (max. degradation -1.9%-1.9%), while INT4 degradation is substantial and model-dependent: BioMistral-7B, clinically fine-tuned, loses 19.7% on MedMCQA, more than any general-purpose model, showing clinical fine-tuning does not confer compression robustness. MedMCQA degrades more than MedQA under INT4; Med-HALT is largely unaffected. On HealthBench's emergency-risk subgroup, Qwen2.5-7B degrades by 26.8% under INT4, suggesting high-risk scenarios are disproportionately vulnerable to compression. On MedSafetyBench, the model family dominates over precision (refusal rates range 10.2%-74.9% at FP16), though Qwen2.5-7B (-17.8%) and Meditron-7B (-28.3%) show substantial INT4 safety degradation; notably, Qwen2.5-7B is simultaneously the most accuracy-robust model, demonstrating that accuracy and safety robustness are independent properties. We additionally test two recovery methods, clinical calibration substitution and QLoRA fine-tuning, both producing the same trade-off: MedMCQA recovers while MedQA further degrades, indicating recovery strategies require task-specific validation rather than being assumed universally beneficial. These findings indicate INT8 is broadly safe for clinical deployment, while INT4 safety must be assessed per-model and per-task, and that safety alignment is determined primarily by instruction tuning rather than clinical domain adaptation.
Global In-situ Observation (GIO) provides fine-scale, direct records of the global weather system from sparse point stations, making it an indispensable source for capturing localized and transient dynamics beyond the reach of satellite gridded data, and playing a critical role in key fields such as numerical weather prediction, disaster prevention, and agriculture. However, GIO exhibits strong spatiotemporal incompleteness, severely impairing accurate and real-time in-situ weather modeling. Unlike existing methods waiting for completed AI-ready data with extra introduced errors, in this work, we explore UniGIO, a novel generative framework for directly modeling global in-situ weather dynamics from native incomplete GIO. By generating missing data from observed ones annotated by masks, it unifies the coexisting forecasting, imputation, and generation under arbitrary missing ratios. Between the missing and observed, UniGIO captures station and region level complementarity through the Observation Mixer and Event Aligner, which diffuse discrete observations into continuous spaces where weather processes naturally span multiple stations. We further establish temporal dependencies with pattern shifts using the Adaptive Temporal Mixer, and track extreme events in chaotic local weather systems through a Mixture-ofExperts structure. Steady and extreme events are adapted in decoder by a Local Refiner. Extensive experiments on the up-todate largest global station weather dataset Weather-5K validate its SOTA performance with 11%, 12%, and 5% advantages on accuracy, fidelity, and extreme event capture, delivering a novel holistic solution for weather modeling in GIO networks.
Toollery: Scaling LLM Agents to Thousands of Skills and Tools
Toollery:将LLM智能体扩展至数千种技能与工具
Tian, Xiangxi, Guan, Ran
Abstract
As LLM agents are exposed to hundreds to tens of thousands of skills, tools, and API functions, full-library prompting becomes costly, slow, and less reliable: each added candidate increases prompt tokens and latency, while longer candidate lists introduce more distractors for LLM selection. We present \textbf{Toollery}, a training-free candidate-compression framework for scalable LLM skill/tool selection. Following established document-side query expansion, Toollery generates user-intent queries from each skill/tool specification and builds a retrieval index that maps real user requests to compact candidate sets before final LLM decision-making. By treating high-level skills and atomic tools as selectable capabilities, Toollery can be applied to both skill libraries and tool registries. We evaluate Toollery on the roughly 79K-capability SkillRouter benchmark, BFCL-V4 with over 440 atomic tools, and 3,396 proprietary smart-cockpit requests over 220 tools. Across these settings, Toollery keeps online selection bounded to a compact top-$k$ candidate set and improves recall over ordinary specification retrieval. At a fixed top-10 budget, Toollery improves end-to-end selection on the cockpit dataset, and maintains comparable AST Accuracy on BFCL-V4. These results support Toollery as a practical candidate-compression framework for large and evolving agent capability libraries, while showing that quality and cost gains depend on workload coverage and provider caching.
Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
测量检测器:面向GPU内核基准测试判定器的变异分析
Du, Mingzhe, Luu, Anh Tuan, Huang, Dong, Ng, See-Kiong
Abstract
Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point tolerance, and their verdicts now feed leaderboards and reinforcement-learning rewards. Recent work agrees these checkers are weak and patches them by hand---extra input distributions, fuzzing recipes, tighter tolerances---with no way to \emph{measure} whether any patch suffices. We introduce mutation analysis as an adequacy metric for kernel-benchmark oracles: deterministic rules inject 10{,}303 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}384 of them with an independent kill witness; any test protocol is scored by the fraction it detects. The official check misses \textbf{one in six} witnessed faults (16.9%), deterministically, and the misses are skewed by family: 8.7% of arithmetic faults escape, but 78.6% of precision faults do. The metric explains why (a tolerance blind band growing with reduction size; a measured ceiling on input aggressiveness set by legitimate floating-point variance), audits the strongest existing patch (KernelBench-Verified's gain splits into $+4.0$ points from hidden inputs and $+4.5$ from tighter tolerance, a split its authors could not compute), and exposes a published fuzzing recipe that rejects \emph{correct} kernels 107 times. Optimizing suites over the kill matrix reaches 98.0% detection with two inputs per problem (94.8% held-out), and the measurement's fault taxonomy teaches a test generator more than the raw faults themselves. Across 48 whole architectures, the blindness grows with scale, concentrating in deep homogeneous pipelines, and two problems prove unrefereeable: their official references violate the benchmark's own tolerance against fp64. We release everything as \href{https://huggingface.co/datasets/Elfsong/KernelBench-M}{KernelBench-M}.
Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark
编码智能体能否复现官方统计数据?受控Eurostat基准测试中的元数据、重试预算与执行反馈的局限性
Necula, Sabina-Cristiana
Abstract
Large language models can generate executable data-analysis code, but successful execution is not equivalent to a valid official-statistics result. This study asks whether authoritative metadata and execution feedback improve the reproducibility of Eurostat answers produced by a coding agent, and isolates what execution feedback actually contributes. A benchmark of 30 natural-language tasks covering seven domains, seven Eurostat datasets and four difficulty tiers was run under four conditions: task only (A), task plus a frozen dataset metadata card (B), metadata plus a repair loop driven by sanitized execution feedback (C), and metadata plus the same attempt budget with no diagnostics of any kind (D). Claude Sonnet 5 generated Python through the Anthropic Messages API in three independent replicates, yielding 360 task-runs. Exact correctness required successful execution, the correct dataset, filters, output shape, values and unit. A companion experiment run under an under-specified output contract, in which the required ranking key and unit representation were never stated to the model, understated condition C by 23.4 points, showing that evaluator and contract design can dominate measured agent error. Reliable statistical coding agents need semantic validation against frozen specifications, a fully specified output contract, and a retry budget - not execution diagnostics.
A Synthetic Multivariate Refrigerator Time-Series Dataset for Predictive Maintenance
一个用于预测性维护的合成多变量冰箱时间序列数据集
Benamirouche, Islam, Fass, Feriel, Ziou, Djemel
Abstract
We generated synthetic multivariate time series for 27 refrigerators with a simplified physicsinspired simulator at one-minute resolution. The simulator includes ambient-temperature variation, door use, thermostat and compressor operation, heat exchange, defrost, electrical consumption, and six progressive degradation types. Each refrigerator provides 15 to 20 sensor outputs according to its configuration. The dataset contains 7,066,161 rows in 27 time-series files and 27 failure logs. The release also includes the Python generator, refrigerator configurations, and documentation. The data can support failure prediction, degradation analysis, and learning across refrigerators with different sensor-output sets.
Counting the items in a bracketed list looks trivial, yet open-weight chat models often get it wrong. Prior work usually blames input bottlenecks such as subword fragmentation or attention dilution, which predict that different models should fail in roughly the same way. Across seven instruct models on identical prompts, however, wrong answers form distinct modes: Qwen and Gemma 27B often flip odd lengths to a nearby even integer, OLMo concentrates errors on a few mid-sized integers, and Llama tends to under-count. These modes are useful labels rather than a stable family law (Gemma 9B does not reproduce Gemma 27B's odd-to-even drop), and heavier subword fragmentation does not make counting harder on our benchmark. When the model answers incorrectly, a linear probe can usually still recover the true count from the residual stream. Matching the same odd-to-even error also does not imply the same late-MLP magnitude fix: scaling a late MLP output helps Qwen modestly but is near null on Gemma 27B under the same protocol, while residual steering can move both only by trading odd gains for even losses. These results caution against transferring that magnitude fix across models without a transfer check.
Millions worldwide require Renal Replacement Therapy (RRT) as a treatment essential for survival. However, optimizing RRT strategies via AI is challenging due to heterogeneous patient dynamics, missing data, and the absence of an AI-oriented health assessment criterion. We propose an AI-Oriented Comprehensive Normalized Assessment (CNA) for healthy status and apply it to optimize RRT strategies by using offline reinforcement learning (RL). The key idea of CNA is transforming vital-sign distributions into a standard normal space, enabling a unified, data-driven health-status score defined by deviations from referent intervals, which also provides an AI-oriented criterion to assess strategy quality and supports RL termination. We further design a structured 23-dimensional state representation that integrates 19 indicators with 4 RRT descriptors, and employ matrix decomposition to reconstruct missing vital signs, improving data completeness for learning. These components are incorporated into multiple offline RL algorithms and validated via systematic ablation studies on RRT feature subsets. Compared with physicians' observed treatments, the best learned strategy reduces mortality from 13.2% to 5.0% (reducing 62.24%) and shortens average in-hospital stay from 308.5 to 250.1 hours (reducing 18.93%), demonstrating both methodological innovation and the potential of CNA-guided RL to improve RRT outcomes in nephrology.
SCALE: Simulation-Calibrated Amortized Learning for Energy Materials (A hybrid architecture connecting deterministic modeling, real-world data, and transformer-scale inference for accelerated energy-materials discovery)
Energy systems face converging pressures for security, affordability, resilience, and sustainability, creating a need for faster discovery of deployable energy materials. Here we introduce SCALE (Simulation-Calibrated Amortized Learning for Energy Materials), a physics-grounded, real-world-data-calibrated learning architecture that connects deterministic scientific operators, experimental calibration, expanded calibrated label generation, and transformer-scale inference. SCALE converts selected high-cost mechanistic computation and measured evidence into reusable models for rapid screening, ranking, inverse design, and active learning. We formulate the framework, identify ten method-based application regimes, and demonstrate SCALE for solid-state metal-hydride hydrogen-storage capacity prediction. In this implementation, a hydride phase-equilibrium capacity operator is calibrated against 381 measured ML-HydPARK capacity anchors and used to generate 5,000 candidate-condition-prototype teacher labels. A crystallographically anchored periodic-graph representation preserves atomic sites, periodic neighbor relationships, and local metal environments absent from formula-only encodings. An edge-biased graph transformer with 2.90 million parameters reproduces calibrated teacher labels with five-fold surrogate fidelity of MAE 0.0582 wt% H2, RMSE 0.0833 wt% H2, R2 = 0.9927, and Pearson r = 0.9963. Post hoc attention analysis suggests that SCALE learns chemically organized element groupings and metal-metal relationships consistent with established hydride chemistry, without chemistry-group labels as supervision. Once trained, SCALE shifts million-candidate evaluation from repeated deterministic workflow execution to batched learned inference, reducing per-candidate screening cost by approximately 10^7-10^8 while retaining links to simulation and experimental evidence.
Chinese Translation
能源系统面临安全、可负担性、韧性与可持续性等多重压力的交汇挑战,因而需要更快速地发现可部署的能源材料。本文提出SCALE(面向能源材料的模拟校准摊销学习),这是一种以物理为基础、经真实世界数据校准的学习架构,连接了确定性科学算子、实验校准、扩展的校准标签生成以及Transformer规模的推理。SCALE将精选的高成本机理计算与实测证据转化为可复用的模型,用于快速筛选、排序、逆向设计与主动学习。我们对该框架进行了系统阐述,识别出十种基于方法的应用模式,并以固态金属氢化物储氢容量预测为例演示了SCALE。在该实现中,氢化物相平衡容量算子基于381个实测的ML-HydPARK容量锚点数据进行校准,并用于生成5,000个候选条件原型教师标签。一种晶体学锚定的周期图表示保留了原子位点、周期性邻居关系以及仅用化学式编码无法体现的局域金属环境。一个具有290万参数的边偏置图Transformer以五折替代精度复现了校准的教师标签:MAE为0.0582 wt% H2,RMSE为0.0833 wt% H2,R2 = 0.9927,Pearson r = 0.9963。事后注意力分析表明,SCALE在没有化学基团标签监督的情况下,习得了与既有氢化物化学一致的化学组织的元素分组及金属-金属关系。训练完成后,SCALE将百万级候选材料的评估从重复执行确定性工作流转变为批量化的学习推理,使每个候选材料的筛选成本降低约10^7–10^8倍,同时保持与模拟和实验证据的关联。
Merging low-rank adapters (LoRAs) promises to eliminate the overhead of swapping task-specific weights at inference time. However, existing merging methods assume every layer needs the same rank budget. Further, some methods assume that rank budget needs to be split equally among the tasks too. We show this uniform-budget assumption is a major source of the performance gap between merged and per-task LoRAs. However, rank selection is an NP hard problem. To this end, we introduce Net Utility, a data free metric that first decomposes every task LoRA by its Singular Value Decomposition (SVD) and scores each of those singular directions by its task utility and its interference with other tasks directions. Next, we globally pool these scores to select singular directions with the highest values with a constraint on the total number of directions selected. The proposed Net Utility metric is applied on top of five different merging methods across three different merging spaces. The merging is done over two sets of tasks, vision and language tasks. Net utility based rank allocation outperforms its counterparts without that allocation. On average, over vision tasks it achieves +2.1% improvement in performance, and +2.2% improvement over the language tasks.
Task-Aware Hybrid QUBO Optimization for Structured Neural Network Pruning
面向结构化神经网络剪枝的任务感知混合QUBO优化方法
Orabi, Osama, Zagitov, Artur, Salloum, Hadi, Lobachev, Viktor A., Kholodov, Yaroslav
Abstract
Neural network pruning can be formulated as a combinatorial optimization problem, yet many existing approaches rely on independent filter-importance scores or simplified objective functions. In this work, we propose a Hybrid Quadratic Unconstrained Binary Optimization (QUBO) framework for structured filter pruning that combines task-aware sensitivity information with interactions between candidate filters. The formulation incorporates first-order Taylor sensitivity and Weight-Fisher sensitivity into the linear component of the objective and can additionally incorporate activation similarity into the quadratic interactions. To control the target pruning cardinality without introducing an explicit quadratic cardinality penalty, we use a binary search over the capacity incentive to identify a coefficient that empirically yields the target pruning cardinality. We further investigate a two-stage QUBO--Tensor-Train refinement strategy in which the QUBO solution initializes gradient-free probabilistic black-box optimization to search for improved pruning masks using the downstream metric. Experiments on the SIDD image denoising task and a Half-UNet model show that the Hybrid QUBO achieves higher PSNR and SSIM than the evaluated Taylor and L1-based QUBO baselines at the studied pruning target. Multi-seed experiments under a fixed dataset protocol are used to assess robustness, while controlled sub-problem experiments demonstrate that Tensor-Train refinement becomes increasingly valuable as the combinatorial problem size grows. The results support Hybrid QUBO as a task-aware structured pruning framework for the evaluated setting, while also highlighting the computational and deployment limitations of mask-based pruning.
Statistical Inference for Adversarial Training: Central Limit Theorems via Optimal Transport
Jakwang, Kim, Dohyun, Kwon
Abstract
The purpose of this paper is to rigorously quantify the statistical and learning-theoretic properties of adversarial training models for classification. Equivalently, we establish the statistical properties of empirical optimal partial transport. Precisely, first we provide two types of central limit theorems (CLT): CLT centered at the expected empirical value, and CLT centered at the population one with smoothing. These results are based on the uniqueness of optimal potential for various equivalent optimal transport formulations, and the empirical process theory argument. For the binary setting, we indeed prove the uniqueness of optimal potential by leveraging the connection between optimal partial transport and the derived multi-marginal optimal transport formula. As byproducts, we also obtain the stability of a saddle point of the adversarial training model, and the sample complexity and concentration probability of the generalization error.
Universal Observatory Graphs for Distributed Sky Coverage and Artificial Intelligence Based Interplanetary Routing
面向分布式天空覆盖与基于人工智能的行星际路由的通用观测台图模型
Razek, Mohammed Abdel
Abstract
This research proposes the Universal Observatory Graph (UOG), an AI-driven framework for distributed astronomical observation across the Solar System. The proposed architecture models autonomous observatories located at the Sun planet L2 Lagrange points as nodes in a weighted graph, while communication links are represented as graph edges characterized by multi-objective physical and operational metrics, including interplanetary distance, communication latency, transmission power, and link reliability. The resulting graph provides a unified mathematical representation of a cooperative interplanetary observatory network. This proposal examines a six-observatory Solar System configuration comprising Earth, Mars, Jupiter, Saturn, Uranus and Neptune. Instantaneous sky coverage is evaluated independently using a 200,000 direction Fibonacci sphere, a 2,000,000 direction fixed seed Monte Carlo calculation and deterministic spherical integration. All three methods yield complete network union coverage, approximately 0.43% complete six observatory intersection and approximately 24.96% mean pairwise Jaccard similarity under the adopted pointing model. Communication routing is subsequently formulated as a finite horizon Markov decision process and solved using tabular Q-learning. The reward balances node participation and a distance dependent reliability proxy against distance, light time latency and a distance squared transmission power proxy. The learned Earth-Saturn-Uranus-Neptune route is also the highest discounted return route among all 41 feasible simple paths under the four hop constraint. The framework provides a reproducible baseline for sequential coverage assessment and multi objective routing; time dependent ephemerides, mission specific visibility, calibrated link budgets and scalable graph policies remain future work.
CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents
CHART:面向工具链鲁棒搜索代理的工具链轮换课程训练方法
Zhang, Xinlu, Lin, Ying-Chun, Zhang, Zhihan, Fetahu, Besnik, Chen, Xi
Abstract
Search agents are usually trained under a single harness. But once an agent is deployed in a real application, its harness is frequently updated (e.g., a rewritten system prompt) to fit production needs. This exposes a fragility of post-trained agents: because a learned behavior is entangled with its training harness, even a harness update that leaves the task unchanged can fail to elicit the behavior. We train a search agent to perform parallel search, a popular strategy for improving both search efficiency and performance. We find that training under a fixed harness makes the behavior harness-local, overfit to that harness's surface form: when the harness changes, the model falls back to serial search. An intuitive fix is harness augmentation, but simply training on more harnesses does not resolve the problem. GRPO learns from the reward gap between parallel and serial rollouts of the same question: a small harness pool saturates that gap early, while a large pool dilutes the per-harness signal too thinly for any harness to consolidate. We therefore propose Curriculum HArness Rotation Training (CHART), a rotating curriculum that lets a search agent gradually consolidate parallel search across harnesses. At each periodic evaluation, CHART "graduates" the harnesses whose expected behavior is learned and replaces them with still-learnable ones, keeping the reward gap alive throughout training. Starting from the same harness pool, CHART makes the model learn parallel search on all harnesses, whereas static augmentation succeeds on at most half of them. The behavior also carries to held-out harnesses: CHART parallelizes on 89% of held-out turns, against at most 5% for the static pools. It further transfers to a new QA task and search environment, improving pass@1 by 5.6pp over the best static pool. Finally, CHART-trained agents benefit more from meta-harness search than baselines.
Forecasting karst aquifer dynamics is difficult because recharge responses are nonlinear, event-driven, and governed by strongly heterogeneous flow paths. This study develops and evaluates a deployment-aware framework for 1-12-week-ahead prediction of spring discharge and groundwater level using approximately 79 years of hydroclimatic observations from the Edwards Aquifer, Texas. Five model families were compared under a common temporal evaluation design: extreme gradient boosting, extremely randomized trees, long short-term memory, convolutional neural networks, and Transformers. Predictions were evaluated using coefficient of determination, Kling-Gupta efficiency, root-mean-square error, and agreement with operational drought thresholds. Extreme gradient boosting was consistently most reliable, with R2 at least 0.97, 0.96, and 0.94 across 1-4-, 5-8-, and 9-12-week horizons, respectively, and greater than 90% critical-stage agreement at the first three drought stages across all horizons. Deep models were competitive at short horizons but degraded progressively and exhibited isolated failures at longer lead times. We attribute this contrast to an alignment between tree partitioning and low-dimensional, axis-aligned hydroclimatic predictors, together with the tendency of neural models to smooth irregular extremes. The validated models were embedded in a five-agent operational architecture that automates data acquisition, model assignment, deterministic prediction, threshold monitoring, prospective verification, literature retrieval, and reporting. The contribution is therefore a transferable framework joining parsimonious model selection, leakage-aware multi-horizon evaluation, decision-relevant threshold skill, and auditable agentic automation.
CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation
CALM:一种用于基于活动的出行者模拟的校准大语言模型选择网络框架
Cheng, Yezhou
Abstract
We present CALM, a reproducible hybrid framework that integrates an optional large language model (LLM) activity planner with calibrated stochastic choice, shared network feedback, memory and habit, typed feasibility checks, and deterministic offline replay. Unlike trip-mode classifiers or diary-only generators, CALM executes a closed traveler-day loop and evaluates each generative module against an empirical, reproducible baseline. On the 2024 New York City Citywide Mobility Survey (CMS), 110,691 seven-mode trips are split by respondent into 78,487 training and 32,204 holdout trips. Training-only alternative-specific constant calibration reduces mean holdout mode Jensen-Shannon divergence from 0.15599 to 0.00394 across ten seeds. A matched live-LLM ablation then quantifies trade-offs among aggregate fit, temporal fit, behavioral persistence, and feasibility, while frozen prompt-response pairs support deterministic replay of downstream simulation. Controlled weather, delay, fare, and parking ladders further demonstrate consistent and interpretable responses under intervention. CALM contributes a reproducible protocol for integrating and evaluating generative planners in traveler simulation through person-disjoint calibration, matched module ablation, controlled stress testing, and end-to-end traceability.
Model merging has emerged as a promising paradigm for integrating multiple task-specific capabilities into a single large language model. However, existing methods predominantly focus on post-hoc processing of independently fine-tuned models, overlooking how the training phase itself impacts cross-task compatibility. Resolving parameter conflicts after fine-tuning is inherently sub-optimal. To address this, we propose CAMFT, a Conflict-Aware Mergeable Fine-Tuning method that makes task adaptation both efficient and mergeaware. CAMFT treats mergeability as a property shaped during fine-tuning, rather than only a problem to be solved after fine-tuning. By guiding each task to update sparse coordinates with lower cross-task conflict, CAMFT produces task updates that are efficient to train and more compatible for downstream model merging. Extensive experiments demonstrate that CAMFT outperforms standard finetuning baselines in multi-task merging scenarios. Codes are available at https://github.com/gyanchow/CAMFT-LLM.
On-policy distillation (OPD) is a promising approach for transferring knowledge between language models, where a student receives dense token-level supervision along its own generated trajectories. However, teacher supervision can be unreliable when conditioned on incomplete or low-quality student prefixes. We identify Teacher Uncertainty Contraction (TUC), a systematic phenomenon whereby the teacher's predictive uncertainty decreases as it continues from a student-generated prefix. We theoretically characterize this trade-off through a variance-bias decomposition of teacher-branch gradients, showing that uncertainty contraction reduces variance while teacher-student path divergence increases bias, thereby favoring a finite continuation. Guided by this insight, we propose Adaptive-Continuations On-Policy Distillation (AC-OPD), which augments informative states along student rollouts with teacher continuations and adaptively selects their effective supervision horizons. Experiments on mathematical reasoning and code generation across model scales demonstrate that AC-OPD consistently improves over standard OPD. Controlled-continuations and matched-budget analyses further validate the adaptive-continuations design, highlighting adaptive teacher continuations as an effective principle for reliable on-policy distillation.The code will be made publicly available upon publication.
Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning
面向自动化大语言模型微调的策略积累与引导执行
Zhao, Haoran, Du, Wei, Yang, Dingwen, Huang, Jixuan, Shang, Junlin, Fang, Lingyong, Guo, Ya, Gui, Tao, Zhang, Qi, Huang, Xuanjing
Abstract
Producing task-specific large language models requires discovering effective training strategies through experimentation. Automated fine-tuning systems have made this experimentation feasible with far less manual effort. However, these systems are stateless: each search discards its discovered strategies, dataset insights, and hyperparameter findings once it ends. Every new task must then repeat this costly search from a cold start. To address this, we propose Strategy Accumulation and Guided Execution (SAGE), a two-stage framework that makes automated fine-tuning search cumulative. In the first stage, a multi-agent pipeline performs Monte Carlo Tree Search-based exploration. A parallel Distillation Agent extracts task-specific exploration records and confidence-scored cross-task insights, which together constitute a structured experience repository. In the second stage, SAGE retrieves relevant experience from this repository and selects what applies to guide training on the new task. We evaluate SAGE on nine unseen tasks spanning both single- and cross-category settings. In single-round execution, SAGE's accumulated experience raises the average relative improvement over baseline from 3.2% to 15.6%, a 12.4-percentage-point gain over the same pipeline without it. These results show that persistent strategy experience provides effective guidance for automated fine-tuning on unseen tasks.
Chinese Translation
构建面向特定任务的大语言模型需要通过实验探索有效的训练策略。自动化微调系统使得这类实验所需的人工投入大幅减少。然而,这些系统是无状态的:每次搜索结束后,其发现的策略、数据集洞察以及超参数结论都会被丢弃。每个新任务都必须从冷启动开始重复这一昂贵的搜索过程。为解决这一问题,我们提出了策略积累与引导执行(Strategy Accumulation and Guided Execution,SAGE),这是一个使自动化微调搜索具有累积性的两阶段框架。在第一阶段,多智能体流水线执行基于蒙特卡洛树搜索(Monte Carlo Tree Search)的探索。并行运行的蒸馏智能体(Distillation Agent)提取任务特定的探索记录以及带有置信度评分的跨任务洞察,二者共同构成一个结构化的经验库。在第二阶段,SAGE从该经验库中检索相关经验,并选择适用于指导新任务训练的内容。我们在九个未见任务上评估了SAGE,涵盖单类别与跨类别两种设置。在单轮执行中,SAGE积累的经验将相对于基线的平均相对提升从3.2%提高到15.6%,相比不使用该经验的相同流水线提升了12.4个百分点。这些结果表明,持久化的策略经验能够为未见任务上的自动化微调提供有效指导。
Large language model-driven remote sensing (RS) agents offer a promising approach to automating geospatial analysis. However, lightweight RS agents based on compact language models struggle with multi-step interactive tasks due to loss of long-horizon states, inefficient environmental feedback utilization, and sparse optimization signals. We propose RS-Claw-Evolution, an environment-feedback-driven framework that progressively improves lightweight agents through three stages. Interaction evolution uses executable code to control observations, maintain intermediate states, and reduce context redundancy. Experience evolution combines failure-aware trajectory generation with error-turn masking to learn from informative failure-recovery experiences without imitating faulty actions. Decision evolution uses reinforcement learning with multi-dimensional environment rewards and turn-level advantage protection to optimize tool-use behaviors and improve credit assignment in long sequences. On Earth-Bench, the optimized Qwen3-4B-based agent achieves 65.9% accuracy in Autonomous Planning mode, outperforming the untrained Qwen3-32B baseline (43.8%) and DeepSeek-V3.1 (60.8%), while approaching GPT-5 (71.6%). These results demonstrate that learning from environmental feedback can improve lightweight agents and narrow their performance gap with larger models in long-horizon RS tasks.
Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch
抗拒、更新、拒绝:偏好优化安装了一个依赖先验的可靠性开关
Yang, Sen, Yeung, Yuen-Hei
Abstract
An aligned model asked to hold its answer against a manipulative source must still update on a reliable one and reject an unreliable one: resistance, reliable-update, and unreliable-source rejection are one three-way contract, not three independent behaviors. We show the objective most anti-sycophancy work optimizes is non-identifying with respect to source reliability: because no preference label depends on whether a source is actually reliable, any scalar mixture of the arms traces a single deference dial, and no point separates two same-template testimonies differing only in stated reliability. This fixation$\leftrightarrow$gullibility frontier is a property of the objective, not any model. We make reliability identifiable through data: a threshold benchmark where a source asserts the opposite answer while stating its reliability $r$, and the correct action is to flip iff $r$ exceeds the model's prior strength $p$. Preference optimization over balanced coverage installs a prior-dependent reliability switch: across three seeds on Qwen2.5-7B-Instruct the threshold $r^\star$ rises monotonically with the prior, decision accuracy reaches $0.84$ with a monotone flip curve (Spearman $0.56$), and the policy generalizes to unseen reliability values and a held-out notation, following stated reliability over role prestige. Three controls localize the cause: an unmatched variant installs the switch equally ($0.80$), a second preference optimizer (IPO) installs it just as well ($0.86$), whereas supervised imitation does not ($0.50$), so the cause is preference optimization over reliability-labeled coverage, not pairing, loss, or imitation. A confirmatory battery replicates the switch on a fresh test draw, bounds it honestly (it keys on reliability stated in the testimony, not a separately audited record), and transfers it to Llama-3.1-8B. The frontier is empirical, not a theorem.
Contrastive Siamese Representation Learning for Predictive Maintenance of Electrical Submersible Pumps
基于对比孪生网络表示学习的电潜泵预测性维护方法
Damarla, Seshu K., Zhu, Xiuli
Abstract
Electrical submersible pumps (ESPs) are essential in offshore oil production, where unexpected failures can result in significant operational and financial losses. Accurate predictive maintenance for ESP systems remains challenging due to nonlinear operating conditions, class imbalance, and variability among pump units. To address these issues, this study presents a fault diagnosis framework that incorporates class imbalance awareness by employing Siamese contrastive representation learning and prior-corrected k-nearest neighbor (KNN) classification. The method first extracts discriminative features relevant to fault detection from vibration-domain indicators and engineered harmonic relationships. A Siamese neural network is trained with class-balanced contrastive pairs to construct an embedding space that clusters samples of the same fault type and separates different fault classes. To further mitigate class imbalance during classification, a prior-corrected distance-weighted KNN is applied. The framework is validated using a Leave-One-ESP-Out (LOEO) strategy to evaluate generalization to previously unseen ESP units. Experimental results indicate that the proposed framework delivers robust and consistent fault classification performance under realistic industrial conditions, supporting its potential for reliable predictive maintenance and intelligent ESP system monitoring.
Common Cause, Not Cross-Attention: Blocking Visual Shortcuts in Audio-Video Generation
共同原因,而非交叉注意力:阻断音频-视频生成中的视觉捷径
Xu, Jian, Zeng, Delu, Paisley, John, Zhao, Qibin
Abstract
Joint audio--video generators are trained on data in which what an event looks like and what it sounds like are strongly, often spuriously, correlated: a particular material, texture, or object appearance co-occurs with a particular sound. This paper is a controlled causal study of the resulting failure mode. Building an AV structural causal model in which the audio is, by construction, independent of the video's nuisance appearance, we show that models which let audio read video directly-through cross-attention or a shared latent-learn a visual shortcut: they predict sound from appearance rather than from the causal event, and collapse when the appearance-event correlation is broken at test time, literally synthesizing the wrong event's sound. Crucially, the popular remedy of routing both modalities through a shared common-cause latent does not fix this: a bottleneck, an unsupervised shared/private factorization, and a faithful shared-prior model all grab the appearance proxy and fail like the direct model. Blocking the shortcut instead requires an intervention on the nuisance. Under the stated SCM and intervention assumptions we prove that counterfactual invariance is necessary and sufficient to identify the causal predictor, and we verify the mechanism across a feature-vector SCM, procedural pixel video, real images with spectrogram audio and a pretrained backbone, moving real digits, and a conditional generator. On a \emph{real, pretrained} video-to-audio generator, an input-intervention test shows the model is far from invariant to sound-irrelevant edits (recolouring or graying a video substantially changes the sound it generates).
Complex-valued Transformers have inherited softmax attention over the raw complex inner product. Outside natively complex domains this standard form stays near chance, and no complex attention had been shown to correct it. We show that the match must be a scaled cosine score: L2-normalise queries and keys, so the score reads their cosine similarity and ignores their magnitudes, and hold that score at order-one scale. With this the same models train on four diagnostic tasks under two different gates; without the normalisation they stay at chance on ListOps and Needle under both gates and fall far below on the other two, and a normalised score placed at too small a scale fails as well. The resulting family of phase-coherent Transformers (\PCT) matches or exceeds the strongest real-valued baseline across long-range memory, positional retrieval, hierarchical reasoning, frequency-domain classification and physical complex signals; it shows no degradation up to depth 20; and its loss decreases log-linearly over a 61-fold range of parameters. A member of the family, complex screening combined with a phase-coherent recurrence, is the first genuinely complex-valued neural network to solve Path-X, with 91.6% of its trainable parameters complex-valued against 38.2% for S4. We record these as signs of generalisation not previously seen in complex-valued neural networks.
In large-scale recommendation systems like the LinkedIn Feed, content generated by a member's network (connections and follows) makes up over 70% of impressions and engagement. It is therefore essential that the pre-ranking layer forwards the best possible few hundred candidates to the ranking layer. LinkedIn's professional knowledge graph carries engagement signals across both the first degree network (connections and follows) and the second-degree network: posts that a 1st-degree connection reacted to, commented on or reshared but did not author (a.k.a. stranger viral). Due to this fan out, the resulting candidate index exceeds one billion; selection of activities from the viewer's network narrows it down to roughly tens of thousands of activities that must be scored within a 120 ms p99 latency budget. We present Connected Content Retriever (CC Retriever), a pre-ranking system that scores these candidates with a full deep ranking model on GPUs at low latency. At its core is a sorted-search GPU primitive that joins dense graph affinity features (viewer to author) with document level features stored on the GPU at runtime in 5-10 ms. The shift to GPU served scoring enabled a 50x scale up of the ranking model's parameters and delivered a +2.5% lift in content time spent on the LinkedIn Feed in online experiments, significantly higher than the typical gains observed in LinkedIn Feed experiments. In this work, we describe the feature set we leverage from LinkedIn's economic graph and the model architecture used for scoring, with a particular emphasis on the online system that scales the stack.
Mixture-of-Experts (MoE) models are increasingly deployed alongside Speculative Decoding (SD) to accelerate inference, but combining the two is challenging. SD improves the inference speed of dense models by verifying groups of tokens in parallel. However, the inference speedup for SD with MoEs depends heavily on the number of tokens being verified. Using more verification tokens results in more experts being transferred from DRAM to the Neural Processing Unit (NPU), which increases the memory transfer cost. This negatively impacts model runtime, as memory transfer is typically the bottleneck in inference. In this work, we investigate the impact of MoE router design during training on the speed of MoEs with SD. We find that routers with high degrees of expert coactivation result in much faster runtimes, mitigating the impact of using more verification tokens. Motivated by this observation, we assess the impact of various router design choices on expert coactivation and runtime using billion-parameter transformer models. We find that combining a global load-balancing loss, shared experts, a consistency loss, and an autoregressive expert selection mechanism during training results in significantly stronger expert coactivation. This increased coactivation translates into higher overall runtime throughput: our exploration yields a model that improves throughput by 21% over MoE baselines, while maintaining on-par accuracy with the baseline MoE.
COREM: Cosine-Relation Momentum Reshaping with Stateful Writeback
COREM:具有状态回写的余弦关系动量重塑方法
Wang, Yan, Wang, Xiaochuan, Sun, Yuxiang
Abstract
Matrix-valued optimizer states may contain relational structure that is not captured by treating their entries independently. We study whether relations within matrix-valued optimizer states can be exploited to improve optimization. To this end, we introduce a unit-relation-transform abstraction and instantiate it as COREM, a Cosine-Relation Momentum Reshaping method with stateful writeback. COREM partitions the momentum state into update units, computes cosine relations among them, and uses these relations to reshape the momentum before writing the transformed state back to the optimizer. This stateful mechanism allows the reshaped momentum to affect not only the current update but also future optimization dynamics. We evaluate COREM on CIFAR-10 with an MLP and on enwik8 with a Transformer. Compared with Muon, COREM shows lower early-stage step efficiency but stronger improvement in the mid-to-late stages of training, achieving better final validation performance on CIFAR-10 and comparable final performance on enwik8. Spectral diagnostics on enwik8 show that COREM consistently increases entropy effective rank and reduces the concentration of singular energy in dominant modes, while preserving an anisotropic spectrum. For square matrix updates, COREM requires approximately 13.3% of the transformation FLOPs of Muon with five Newton-Schulz iterations.
EmbeddGAN: A Novel GAN Framework Using an Embedding Network and Gini Distance Correlation
EmbeddGAN:一种基于嵌入网络与基尼距离相关性的新型生成对抗网络框架
Caldwell, MaTais, Chen, Yixin, Dang, Xin, Walter, Charles
Abstract
Generative Adversarial Networks (GANs) have demonstrated strong performance in generating high-quality synthetic data. However, they are limited by no formal guarantees regarding convergence and the effectiveness of the learning process. In practice, this leads to training instability, mode collapse, and sensitivity to hyperparameters. To address this, we propose EmbeddGAN, a novel adversarial training framework based on a dependence-based objective. Instead of relying on a discriminator that classifies samples as real or fake, EmbeddGAN introduces an embedding network that learns a representation in which statistical dependence between samples and their real/fake labels is maximized, while the generator is trained to minimize this dependence. This objective is implemented using the Gini distance correlation (gCor), which equals zero if and only if the embeddings are statistically independent of the real/fake label. Minimizing this objective therefore encourages real and generated samples to become statistically indistinguishable in the learned embedding space. The embedding network projects both real and generated data into a shared low-dimensional space, where distributional discrepancies can be measured directly through pairwise distances. We adopt a minimax training strategy: the embedding network maximizes the Gini distance correlation (maximizing dependence), while the generator minimizes it (minimizing dependence). Experiments on the MNIST, CIFAR-10, and CelebA datasets demonstrate that EmbeddGAN achieves competitive performance relative to established baselines while exhibiting notably stable training dynamics on the evaluated datasets.
Backpropagation (BP) has driven the remarkable success of modern deep learning by enabling large hierarchical networks to learn complex functions end-to-end. Yet it does not by itself determine how parameters should be organized so that functional components can be reused and adapted selectively. For example, object recognition and motion prediction may depend on overlapping parameter sets, making them difficult to isolate or modify independently. We call this condition weight entanglement. Modern architectures dynamically select which parts of a network process each sample: nonlinearities gate units, attention selects interactions, and Mixture-of-Experts architectures route inputs to modules. Yet such selection does not ensure that the same functional component remains linked to an identifiable parameter set across samples. We propose weight operators: parameterized modules that implement reusable functional components and can be composed at inference to form the function required by each sample. Learning proceeds in two stages: the model first infers the required operator composition, then updates only the selected operators' parameter sets. Vector Networks (VNs) provide one implementation. They couple operator selection to local error-driven updates within each layer and show that learned operators can be reused in combinations absent from training while updates remain restricted to the selected parameter sets. This provides a basis for testing functional parameter identifiability: whether an operator remains linked to the same functional component during learning. We argue that functional parameter identifiability may provide an organizing principle for models that systematically reuse and recombine learned functions while adapting only the components that need to change.
Benchmarking Hybrid Deep Learning Architectures for Predictive Maintenance in Industry 4.0
面向工业4.0预测性维护的混合深度学习架构基准测试
Zhengyang, Gu, Hernandez, Joseph E., Cook, Thomas, Burtenshaw, John, Scott, Sean, Couch, Chris
Abstract
Predictive maintenance in Industry 4.0 refers to using data from sensors, machines, and production systems to estimate when equipment is likely to fail, so maintenance can be planned before a breakdown occurs [1]. However, a model that predicts maintenance may work perfectly in the lab but fail unexpectedly when applied to real factory data [2]. To solve this "reliability" gap, we evaluated six deep learning architectures across more than 700 experimental runs. We focused on the two dominant approaches in the field: Recurrent Neural Networks (RNNs), which process data step-by-step, like reading a sentence [3], and Transformers, a recent dominant approach, which look at the entire sequence at once to spot important connections [4]. We examined whether Transformers still outperform recurrent neural networks (RNNs) when the data includes noise [5]. We found that while Transformers excelled at tracking stable, slow-moving processes, they tend to overreact to chaotic data, mistakenly taking sensor noise for meaningful signals [6]. We also found that the hybrid method that combines a Long Short-Term Memory (LSTM) layer with a Transformer layer is more resilient to noisy data from factory shops [7]. Functioning as a noise filter, the LSTM smooths out data volatility, allowing the Transformer to focus on the bigger picture without being distracted [8]. The hybrid model did not just improve accuracy; it proved to be significantly more consistent than complex models, delivering reliable predictions regardless of how chaotic the underlying system became.
Augmenting PID Control with Deep Reinforcement Learning: A Hybrid Approach to the Industrial Benchmark
用深度强化学习增强PID控制:面向工业基准的混合方法
Zhengyang, Gu, Hernandez, Joseph E., Burtenshaw, John, Scott, Sean, Cook, Thomas, Couch, Chris
Abstract
As industrial processes grow in complexity, traditional Proportional-Integral-Derivative (PID) controllers are often insufficient for handling their non-linear, multi-input dynamics. We propose using advanced Deep Reinforcement Learning (DRL) to prove its advantages in these complex environments. To do this, we rely on the Industrial Benchmark (IB). The IB is a realistic simulation that tests DRL algorithms against the key challenges of industrial applications: high-dimensional state spaces, delayed effects, and conflicting multi-criterial objectives. This testbed highlights DRL's core trade-off: while its final policies can often be unstable, its unique strength is the ability to autonomously discover optimal, non-obvious policies in multi-dimensional spaces where simple controllers fail. In this paper, we propose a novel hybrid PID-RL controller that leverages DRL's discovery capability while ensuring Reliability. After developing a multi-objective reward function to make DRL viable, we use a twin-delayed deep deterministic (TD3) agent as a discovery tool to find the optimal, non-obvious settings for the IB's 'Gain' and 'Shift' parameters. By feeding these discovered parameters to a simple, tuned PID controller, our hybrid model successfully combines all three characteristics: it achieves the optimal Performance and Efficiency of the best DRL agent with the Reliability of a classical controller. This work demonstrates a practical methodology for using DRL to augment, rather than replace, trusted industrial control systems.
TWIG: A Time-Causal Wavelet Operator for Autoregressive Forecasting on Irregular Graphs
TWIG:一种用于不规则图自回归预测的时间因果小波算子
Venkatasubramanian, Subashree, Barajas-Solano, David A., Liu, Chuyang, Tartakovsky, Daniel M., Dwivedi, Dipankar
Abstract
We introduce TWIG (Time-Causal Wavelet Operator for Irregular Graphs), a graph-native neural operator for autoregressive surrogate modeling on static irregular graphs. TWIG transforms each node history into causal multiscale temporal features that separate recent variation from progressively slower memory components, then propagates these features through graph-wavelet operator blocks with gated pointwise channel mixing. The architecture is causal by construction and designed for closed-loop forecasting, where predictions are recursively reused as future inputs. We evaluate TWIG on three irregular-domain forecasting problems spanning regional diffusion, three-dimensional subsurface hydrology, and aerodynamic flow, with graphs ranging from 400 to 5,233 nodes and model capacities from approximately 70k to 10M parameters. TWIG achieves the lowest aggregate rollout errors on the subsurface-hydrology and regional-diffusion benchmarks and ranks second on the 10M-parameter aerodynamic-flow benchmark, behind the GPS Transformer. Across all three settings, TWIG consistently outperforms the corresponding non-time-causal Graph WNO baseline. These results demonstrate that TWIG provides an effective and scalable approach to stable autoregressive forecasting of dynamical fields on irregular graphs.
Chinese Translation
我们提出了TWIG(Time-Causal Wavelet Operator for Irregular Graphs,面向不规则图的时间因果小波算子),这是一种图原生的神经算子,用于在静态不规则图上进行自回归代理建模。TWIG将每个节点的历史序列转换为因果多尺度时间特征,将近期变化与逐渐变慢的记忆成分分离开来,然后通过带有门控逐点通道混合的图小波算子模块传播这些特征。该架构在构造上具有因果性,专为闭环预测而设计,其中预测结果被递归地用作未来的输入。我们在三个不规则域预测问题上评估了TWIG,涵盖区域扩散、三维地下水文和空气动力学流动,图规模从400到5,233个节点不等,模型容量约为70k到10M参数。TWIG在地下水文和区域扩散基准上取得了最低的总体滚动预测误差,并在10M参数的空气动力学流动基准上排名第二,仅次于GPS Transformer。在所有三种设置中,TWIG均始终优于对应的非时间因果的Graph WNO基线。这些结果表明,TWIG为不规则图上动力场的稳定自回归预测提供了一种有效且可扩展的方法。
This letter covers a broad comparison of methods for classification and regression applications for a user-level handover decision making in scenarios with adverse propagation conditions involving buildings, coverage holes, and shadowing effects. The simulation campaigns are based on network simulator ns-3. The comparison encompasses classical machine learning approaches, such as KNN, SVM, and neural networks, but also state-of-the-art fuzzy logic systems and latter boosting machines. The results indicate that SVM and MLP are the most suitable for the classification of the best handover target, although fuzzy system SOFL can perform similarly with lower processing time. Additionally, for the download time estimation, LightGBM provides the smallest error with short processing time, even in hard propagation scenarios.
Concurrency-Aware Process Model Forecasting with Causal Nets
基于因果网的并发感知过程模型预测
Yu, Yongbo, Peeperkorn, Jari, De Smedt, Johannes, De Weerdt, Jochen
Abstract
Process model forecasting (PMF) aims to predict the process model that will characterize a future period, thereby providing a process-level view of how behavior is expected to evolve. Existing PMF methods, however, forecast directly-follows graphs, which cannot explicitly represent concurrency. We extend PMF to causal nets by forecasting time series of relation and binding counts and using these forecasts to reconstruct future process models with AND/XOR semantics. To evaluate the resulting models, we introduce a protocol that accounts for partial traces and constructs the workflow nets required for conformance checking. Experiments on four event logs show that the forecasted models achieve conformance levels close to those of models re-mined from observations in the corresponding future windows. They also outperform static discovery baselines, which retain high precision on the structurally stable log but exhibit substantial precision losses on the other three logs. Filtering infrequent bindings improves most conformance metrics, although it also removes much of the concurrent behavior captured by the models.
Chinese Translation
过程模型预测(Process Model Forecasting, PMF)旨在预测能够刻画未来时期特征的过程模型,从而提供行为预期演化的过程级视角。然而,现有的PMF方法预测的是直接后继图(directly-follows graphs),无法显式表示并发关系。我们将PMF扩展到因果网(Causal Nets),通过预测关系计数和绑定计数的时间序列,并利用这些预测结果重构具有AND/XOR语义的未来过程模型。为了评估所得到的模型,我们提出了一种考虑部分轨迹并构建一致性检查所需工作流网(workflow nets)的评估协议。在四个事件日志上的实验表明,预测模型的一致性水平接近于从对应未来窗口的观测数据中重新挖掘所得的模型。预测模型还优于静态发现基线方法:后者在结构稳定的日志上保持较高的精确度,但在其他三个日志上出现了显著的精确度损失。过滤低频绑定可以提升大多数一致性指标,但同时也移除了模型所捕获的相当一部分并发行为。
Classification with Abstention Under Class-Conditional Error Constraints
类条件错误约束下的带弃判分类
Kalan, Mohammadreza M., Deng, Yuyang, Hamidi, Sanaz
Abstract
We study binary classification with abstention under separate class-conditional error constraints, with the objective of minimizing abstention while keeping both errors below prescribed thresholds. We characterize the distribution-free minimax rate of excess abstention risk, up to logarithmic factors, in terms of the complexity of the hypothesis class and the sample size. To make the framework amenable to computation with models such as neural networks, we introduce surrogate-loss formulations and derive finite-sample guarantees for excess surrogate ambiguity risk. We formulate the resulting learning task as a constrained optimization problem and characterize its computational complexity in the convex setting. Finally, we evaluate our approach on various datasets and compare its performance with a competing method for this problem.
Monotone-Constrained Diffusion Models for Long-Horizon Production Forecasting
面向长时程生产预测的单调约束扩散模型
Abraha, Temesgen Mikael, Lucet, Yves
Abstract
Forecasting a long horizon from only the first observations of a sequence is ill-posed: many trajectories are consistent with the same short history. We study this problem in oil and gas production forecasting, where forecasts made after roughly the first fifth of a well's producing life drive development and abandonment decisions, and where a usable forecast must describe a monotone decline. We present Physics-SIMS-TS, a conditional diffusion forecaster that combines negative guidance against synthetic artifacts, decline-curve constraints and an isotonic projection applied during sampling, spatial training augmentation, and an ensembled stochastic sampler yielding a full predictive distribution. Across three jurisdictions and more than 35,000 wells, under a shared-space, validation-frozen protocol, Physics-SIMS-TS is the most accurate diffusion forecaster in the comparison and is competitive with, but not superior to, ensembled transformer forecasters. Its forecasts are monotone by construction at a cost of at most 0.5% in mean squared error, and its trajectory ensemble yields calibrated intervals after one dispersion factor is fitted per jurisdiction. On six standard benchmarks a reversible-instance-normalization variant of the backbone is the leading diffusion baseline. We also quantify four protocol choices on which the measured ranking depends. Code and evaluation artifacts are released.
We develop an index policy for finite-horizon Bernoulli multi-armed bandits from minimax solutions to single-arm bandit (SAB) problems. Each SAB problem involves choosing between an unknown Bernoulli arm and a known reward. We show that minimizing worst-case regret of SAB problems over all non-anticipative policies admits an exact semi-infinite linear programming formulation. The resulting stopping policies offer a natural way to compare arms: the higher the known reward against which a policy continues sampling, the more promising the unknown arm. We turn this intuition into indices based on cumulative continuation probabilities, with a monotone adjustment and a reward-shortfall cap. By relating index errors to the regret of single-arm stopping policies, we establish a distribution-free regret bound of $4.45\sqrt{KT}+10.75K$ for $K$ arms and horizon $T$. This bound matches the minimax-optimal regret order established in the literature. The guarantee extends to rewards supported on $[0,1]$ through Bernoulli randomization. We also provide a finite-grid implementation with quantified approximation loss. In numerical experiments, the SAB-based index policy achieves lower worst-case regret than every tested benchmark policy across all evaluated numbers of arms and horizons, while closely matching the grid-based MAB minimax policy in the two-arm setting.
Autonomous Model Lifecycle Management for Digital Twin-Based Manufacturing Control
面向基于数字孪生的制造控制的自主模型生命周期管理
Zhengyang, Gu, Cook, Thomas, Rohrbaugh, Fredaljohn, Hernandez, Joseph E., Couch, Chris
Abstract
Manufacturing AI systems must autonomously adapt to continuous distributional shift from raw-material variability, ambient changes, and equipment aging, under strict safeguard and operator-trust requirements where model failures risk physical damage. This paper presents a closed-loop Cyber-Physical System (CPS) for autonomous model lifecycle management in automotive manufacturing, deployed since 2023. The system manages product-specialized model pairs: a sequence-to-sequence physics model (LPP) serving as a digital twin, and a deep Reinforcement Learning (RL) control policy (LCP) trained against it. Per retraining cycle, multiple model variants spanning architecture families and RL algorithms compete; only the best-scoring candidate advances. A Conductor orchestrator autonomously manages plant-wide model inventories with dependency-aware retraining and Proportional-Integral-Derivative (PID) fallback. Reflecting the principle of Human-Centric Intelligence, the LCP composite score embeds an operator-trust gate penalizing policies deviating from established practice; without it, 23% of policies are rejected by operators despite passing accuracy thresholds. Across multiple facilities, LCP-controlled processes achieve process stability improvements of 28-45% over uncontrolled baselines with zero safety incidents.
D-IMPL: A Diffusion-based Solver for Parameterized BBOs
D-IMPL:一种基于扩散模型的参数化黑盒优化求解器
Hu, Yang, Li, Na
Abstract
Diffusion models have demonstrated strong power in generative modeling tasks across multiple domains, exhibiting a remarkable capability of learning complex distributions from samples. In this paper, we leverage such capability to design an efficient universal diffusion-based solver for parameterized black-box optimizations (BBO), where the optimizer has only black-box access to queries of the objective function at the learning stage, yet is able to reduce the additional computational cost at the inference stage for each BBO instance while also capturing the potential multi-modal landscape of non-convex objectives. To cast our formulation as a compatible generative modeling task, we introduce the notion of minimization policy as a new solution concept, which defines a sampling distribution over the solutions that should concentrate around the minimizer set for each BBO instance. We then propose Diffusion-based Iterative Minimization Policy Learning (D-IMPL), a practical generative-model-based solver for solving parameterized BBOs that employs diffusion models to learn a minimization policy, whose density is proportional to the exponential of the negated objective values, thereby amortizing the computational costs across different BBO instances. Furthermore, we demonstrate the performance of our D-IMPL algorithm by establishing a sample complexity guarantee showing that a $\delta$-approximate minimization policy can be effectively learned within $O(\log(1/\delta))$ iterations, and by extensive empirical evaluations over a range of constrained and unconstrained BBO tasks.
Look Before You Steer: Geometry Predicts SAE Feature Steerability
三思而后行:几何特性预测SAE特征的可操控性
Khan, Muhammad, Channawar, Shlok, Gurugubelli, Akshaj, Gupta, Girish, Shah, Aditya
Abstract
Steering with SAE features requires per-feature coefficient tuning, which currently demands intervention sweeps. We ask whether properties of the SAE itself, computable before any forward pass, predict which features will be cheap or expensive to steer. We show that variation in SAE feature steerability is partially predicted by decoder-space geometry: neighbor density and maximum cosine similarity to nearby decoder directions, both computable from the SAE weight matrix before any intervention, rank features by how much steering they require for a fixed behavioral effect ($\rho$ up to $-0.546$, $p < 10^{-6}$, AUROC 0.610-0.822 across conditions; the signal is rank-based, consistent with grid discreteness). This geometry-steerability relationship replicates across two Gemma-2 model scales (2B and 9B), two SAE widths (16K and 65K), and is detectable cross-architecturally on Llama-3.1-8B-Instruct ($\rho = -0.266$, $n = 300$). On Qwen3-8B with BatchTopK SAEs, geometry predicts whether a feature is steerable at all but not the continuous ordering among responsive features, revealing a boundary condition tied to SAE training regime. The signal weakens at deep proportional layer depth in both models, where the cost of steering exceeds our intervention budget, a consistent depth boundary. These results provide preliminary evidence that pre-steering geometry can partially inform coefficient selection, offering a path toward screening features for controllability before deployment.
Improved Private Sparse Covariance Estimation with Multiscale Threshold Tests
基于多尺度阈值检验的改进型隐私稀疏协方差估计
Zhang, Zihan
Abstract
We study differentially private covariance estimation in operator norm for mean-zero sub-Gaussian distributions with unknown covariance support and at most $k$ nonzero entries per row. We develop a multiscale random-threshold algorithm with sample complexity $\ot(k^2/\alpha^2+k\sqrt d/(\alpha\varepsilon))$ for $(\varepsilon,\delta)$-differential privacy and error at most $\alpha\sigma^2$, where $d$ is the dimension and $\sigma$ is a known sub-Gaussian scale. The bound improves the privacy-dependent term of the existing $\ot(k^2/\alpha^2+k^{3/2}\sqrt d/(\alpha\varepsilon))$ \citep{kumar2026curse} upper bound by a factor of $\sqrt k$, and matches the lower bound of $\widetilde{\Omega}(k^2/\alpha^2 + k\sqrt{d}/(\alpha\varepsilon))$ in its applicable parameter regime. Our key technical ingredient is a direct operator-norm bound on the centered fluctuations of an ideal reconstruction, exploiting conditional independence rather than accumulating entrywise errors across each row. A multiscale allocation of threshold tests balances reconstruction variance against query sensitivity. Together, these ingredients sharpen the trade-off between approximation error and privacy protection, removing the additional $\sqrt{k}$ factor from the privacy-dependent sample complexity.
Robust Market Making with Hawkes Order Flow and Price Impact via Adversarial Reinforcement Learning
基于对抗强化学习的Hawkes订单流与价格影响的鲁棒做市
Yang, Hao, Xu, Zhenguo
Abstract
Market-making strategies in real limit order book markets face substantial model uncertainty and regime-shift risk. Existing adversarial reinforcement learning approaches improve robustness by formulating the Avellaneda--Stoikov market-making problem as a zero-sum game between a market maker and an environmental adversary. However, these approaches typically rely on Poisson order arrivals and neglect trade-induced price impact, limiting their ability to capture important high-frequency market microstructure effects such as clustered order flow, self-excitation, and post-trade price feedback. We extend adversarial reinforcement learning for market making to a more complex environment with Hawkes self-exciting order arrivals and trade-induced price impact. To mitigate the increased non-stationarity introduced by the expanded regime space, we incorporate an LSTM module that explicitly models the temporal structure of recent observations. We further characterize the equilibrium properties of the proposed framework through both game-theoretic analysis and numerical experiments, and introduce a robustness evaluation protocol focused on improvements in the left tail of the return distribution. Experimental results across a range of market regimes show that the proposed method achieves improved left-tail performance in most complex microstructure environments. In particular, the gains are pronounced in regimes with strong Hawkes excitation and low-to-moderate price impact. Bootstrap tests provide no evidence that these improvements are obtained through a stronger terminal directional inventory bias. These results suggest that combining adversarial training with temporal state representation can improve the robustness of reinforcement-learning-based market-making strategies under order-flow self-excitation, price impact, and regime uncertainty.
Reward-free latent world models can learn from offline videos and solve new image--goal tasks by optimizing actions through predicted latent futures. This setting places two demands on the planning state: its coordinates must be comparable with a goal image. Moreover, its dynamics must retain velocity, motion trend, contact, and other history--dependent information beyond those goal coordinates. Offline training creates a second mismatch: each recorded trajectory reveals one factual future, whereas a sampling--based planner compares many actions that were not taken from the same state. We introduce FIRM-WM (Factual--Interventional Recurrent World Model), a compact pixel world model designed around these two gaps. Its recurrent state separates a typed, goal--comparable configuration from a 128-dimensional dynamic fiber used for prediction but excluded from the terminal goal cost. Broad factual trajectories provide state coverage, while common--reset intervention branches provide observed outcomes for alternative action sequences. Before executing each branch, we reset the environment and restore the same recorded values exposed by the environment's state--setting interface. Under matched CEM planning and three independent full-pipeline seeds, FIRM-WM reaches 99.0$\pm$1.0% on TwoRoom, 92.7$\pm$2.1% on Reacher, and 88.0$\pm$3.0% on OGBench-Cube, compared with 89.0%, 88.0%, and 70.0% for LeWM. The deployed model uses 2.98--3.42M parameters and records 2.13--11.60$\times$ lower planning time on these tasks.
Chinese Translation
无奖励的潜在世界模型可以从离线视频中学习,并通过在预测的潜在未来上优化动作来解决新的图像-目标(image-goal)任务。这一设定对规划状态提出了两点要求:其坐标必须可与目标图像进行比较;同时,其动力学必须保留速度、运动趋势、接触以及其他超出目标坐标范围的历史依赖信息。离线训练带来了第二个不匹配:每条记录的轨迹只展示一种事实性的未来,而基于采样的规划器需要比较从同一状态出发的多种未被执行的动作。我们提出了FIRM-WM(事实-干预循环世界模型,Factual-Interventional Recurrent World Model),一个围绕这两个缺口设计的紧凑型像素世界模型。其循环状态将可类型化、可与目标比较的构型与一个用于预测但不计入终端目标代价的128维动态纤维分离开来。广泛的事实轨迹提供状态覆盖,而采用通用重置的干预分支则为备选动作序列提供可观测的结果。在执行每个分支之前,我们重置环境,并通过环境的状态设置接口恢复相同的已记录状态值。在匹配的CEM规划设置和三次独立的全流程随机种子下,FIRM-WM在TwoRoom上达到99.0±1.0%,在Reacher上达到92.7±2.1%,在OGBench-Cube上达到88.0±3.0%,相比之下LeWM分别为89.0%、88.0%和70.0%。部署的模型使用2.98–3.42M参数,并在这些任务上实现了2.13–11.60倍更低的规划时间。
Counterfactual Tool Ranking under Utility, Cost, and Privilege Constraints
效用、成本与权限约束下的反事实工具排序
Li, Jiapeng
Abstract
Counterfactual tool evaluation must distinguish authority, historical support, and what a comparison actually estimates. We study these distinctions with eleven executable enterprise-inspired tools, exact-propensity logs, and real local Model Context Protocol transport. An initial 45-run synthetic study is retained, then challenged by 30 realized-return control runs and 15 experiments on 1,930 independently released Berkeley Function Calling Leaderboard (BFCL) tasks. Full-return direct regression reverses an initially favorable doubly robust (DR) evaluation result in the linear setting: mean absolute errors are 0.0139 for direct regression and 0.0272 for DR. Under a shifted environment, DR retains an advantage, with errors 0.0227 versus 0.0948. On function-name-group-disjoint BFCL-derived splits, direct and DR selectors obtain balanced accuracies of 81.85% and 79.83%. Two pinned local Qwen2.5 models are evaluated on the same 200 held-out tasks, exposing a strong failure to abstain under the fixed prompt. We further characterize policy differences under missing support: unsupported actions shared by two policies cancel, allowing point identification of an incremental change when neither absolute value is identifiable. A disagreement-preserving fallback achieves this property in all five support-gap runs, but conservative sampling bounds do not certify deployment improvement. The contribution is a falsifiable evaluation method and independent public evidence, not a new DR estimator, official BFCL leaderboard score, or production-agent safety claim.
Chinese Translation
反事实工具评估必须区分权威性、历史支撑以及比较实际估计的内容。我们通过十一个可执行的企业风格工具、精确倾向性日志以及真实的本地模型上下文协议(Model Context Protocol)传输来研究这些区别。研究保留了一个初始的45次运行的合成实验,随后通过30次实际回报对照运行以及在1,930个独立发布的伯克利函数调用排行榜(Berkeley Function Calling Leaderboard, BFCL)任务上的15组实验对其加以检验。在设定为线性的场景中,全回报直接回归(direct regression)逆转了最初有利的双重稳健(doubly robust, DR)评估结果:直接回归的平均绝对误差为0.0139,而DR为0.0272。在环境发生偏移的情况下,DR保持优势,误差分别为0.0227与0.0948。在按函数名称组划分的BFCL派生互斥数据集上,直接选择器与DR选择器分别取得81.85%和79.83%的平衡准确率。两个固定的本地Qwen2.5模型在相同的200个保留任务上进行评估,暴露出在固定提示词下模型严重缺乏弃权能力。我们进一步刻画了缺失支撑下策略间的差异:两个策略共享的不受支撑的动作相互抵消,使得在两个绝对值均不可识别的情况下仍能对增量变化进行点识别。一个保持分歧的回退机制在全部五组支撑缺口运行中均实现了该性质,但保守的抽样边界并不能证明部署改进。本文的贡献是一种可证伪的评估方法以及独立的公开证据,而非新的DR估计器、官方BFCL排行榜分数或生产级智能体的安全性声明。
Average squared error cannot reveal whether forecasting performance degrades because the future becomes less predictable or because forecasts move farther from the conditional mean. We introduce paired, mechanism-controlled stress tests that decompose changes in expected squared error at each lead time into environmental risk and forecast-oracle distance, using an origin-conditioned predictive oracle unavailable to the evaluated models. Three end-to-end controls have known attribution. Specifically, the null, environmental-only, and information-gap controls verify that the pipeline assigns changes to the correct component. We then apply the benchmark to 24 deployable forecasters. Under frequent switching, 14 methods have higher realized MSE but lower oracle distance; under outlier-variance feedback, 19 have higher MSE but lower scale-standardized MSE. Short- and long-lead stress-response rankings have Spearman correlation 0.624, revealing substantial horizon-dependent reordering. We then study multivariate relation shifts. Across six models and three coupling severities, oracle distance accounts for only 0.7-3.9% of the decomposed expected-risk increase, and environmental-risk majority persists in an eight-channel system and a matched-difficulty audit of Ring, Block, and Hub relations. Finally, prespecified contrasts on independent data-generating process (DGP) realizations show that several visually compelling discovery profiles, including trend accumulation and the hypothesized switching reversal, do not replicate. The benchmark thus combines component-wise diagnosis with a held-out stability audit. It complements real-data out-of-distribution evaluation, which measures performance under realistic shifts when exact oracle attribution is unavailable.
We study personalized federated reinforcement learning, in which $n$ agents, each acting in its own Markov decision process, collaborate through a server to learn a shared MAML-style policy initialization that becomes effective for an individual agent once that agent adapts it with a single local policy-gradient step. We propose Per-FedAvg-PG, in which agents take $\tau$ local stochastic meta-policy-gradient steps between communication rounds, and prove that it reaches an $\varepsilon$-approximate first-order stationary point of the personalized objective in $K=\mathcal O(\varepsilon^{-3/2})$ rounds with $\tau=\Theta(\varepsilon^{-1/2})$ local steps. The analysis rests on a structural feature of the reinforcement learning setting: under standard policy-class regularity, the per-agent objectives have uniformly bounded gradients and Hessians with explicit constants, so the bounded-gradient and bounded-heterogeneity conditions imposed by the supervised theory hold automatically and no separate heterogeneity assumption is needed. The exact meta-gradient requires the inner-loop policy Hessian, which our experiments identify as the practical bottleneck. We therefore analyze the Hessian-free variant, bound its bias, and exhibit fixed points at which the meta-gradient is nonzero and of order $\alpha$, showing that the resulting stationarity floor is a property of the method rather than of the bound. Experiments on tabular and neural navigation confirm the predicted behavior and show transfer to unseen agents at an order of magnitude lower sample cost than independent training. Together these results identify the adaptation step size as a tunable personalization knob and the curvature estimate as the quantity that governs whether exact meta-gradients are affordable.
Time series foundation models (TSFMs) have recently delivered impressive zero-shot performance across diverse forecasting tasks. However, real-world decision-making frequently relies on \emph{irregular multivariate time series} (IMTS), where inconsistent inter-observation intervals and asynchronous sampling across variables coexist with informative missingness. Existing TSFMs handle such inputs either through imputation that injects spurious values or through index-based positional encodings that ignore continuous time. There is still a gap in the foundation model that follows the original IMTS patterns. In this paper, we propose a hybrid attention model that learns a unified time-aware patch representation for IMTS forecasting. We first design a \emph{time-aware patch encoding} that maps a variable number of intra-patch timestamps into a fixed-size embedding, producing a uniform format for irregular patches without resorting to imputation. We then introduce a \emph{time bias attention} mechanism that calibrates inter-patch temporal misalignment and asynchronous cross-channel dependencies as auxiliary attention offset. Finally, on top of a decoder-only Transformer backbone, we adopt a \emph{hybrid causal mask} that preserves a bidirectional full view over the historical context while keeping the forecast horizon strictly autoregressive. To support large-scale pretraining under irregular settings, we also curate VersaTSA, an archive of $30$B observations that retains the native sampling sparsity of its sources. Experiments on three IMTS benchmarks and a standard regular-MTS benchmark show that our model achieves state-of-the-art zero-shot performance on IMTS and remains competitive when transferred to regular forecasting.
Contrastive activation directions are often interpreted from what they decode or how strongly they steer behavior. But what evidence is sufficient to identify the construct represented by such a direction, rather than a correlated feature of the contrast used to extract it? We study this question for a good--bad outcome direction in a maze task, using controlled interventions that separate the realised outcome from the informational history through which it became known. Across multiple LLM checkpoints, directions fitted on one explicit outcome encoding transfer well to another, indicating that the readout is not tied to surface form. In contrast, when the same realised outcome is reached through announced and unannounced histories, transfer degrades substantially: even after both histories receive the same explicit outcome, the post-event readout remains strongly conditioned on the earlier announcement. In a matched maze-RL run, the post-RL direction becomes substantially more predictive of reference-MDP remaining return and the policy becomes more dependent on it at the tested sites, while this history dependence persists. These results support a functional, value-related interpretation of the direction, but not its identification with a history-invariant scalar valence state.
Graph neural networks are widely used for drug--target affinity (DTA) prediction, and discrete Ricci curvature has recently been used to characterize molecular graph geometry. Existing curvature-aware DTA approaches mainly use static curvature on the drug graph while representing proteins primarily with sequence-derived features. This leaves pair-adaptive use of graph geometry underexplored, which may limit adaptation to unseen entities in cold-start settings relevant to practical screening. We present CurvFlow-DTA, which replaces a single static curvature representation with weighted Forman curvature flow on both molecular and protein residue--residue contact graphs. A label-independent flow trajectory is precomputed for each entity, and a pair-conditioned selector determines the horizons read by a dual-branch Flow-GINE. A frozen ESM-2 supplies residue-level representations and contact scores used to construct the protein graph. Inference requires only SMILES strings and protein sequences, without a bound complex structure. On Davis and KIBA, CurvFlow-DTA improves on the protocol-matched Ricci-GraphDTA baseline in every warm and cold-start setting. Warm-split mean squared error (MSE) decreases by $19.9\%$ on Davis and $18.9\%$ on KIBA. Across the six cold-start comparisons, MSE decreases by $14.3$--$27.4\%$, with higher concordance index (CI) throughout. Within our compiled set of literature baselines, CurvFlow-DTA achieves the lowest MSE on both warm benchmarks and across four out of six cold-start evaluation settings.
We introduce Causilo, a tabular foundation model (TFM) that combines frontier predictive performance with exceptionally fast inference. On TabArena, Causilo achieves 1785.4 Elo, at a median inference time of 0.10 seconds per 1K test samples. It outperforms TabPFN-3.5-Fast with 31.6% less inference time, placing it on the performance--efficiency Pareto frontier. Causilo follows TabICL's column-then-row architecture but introduces another row-refinement module before row compression. This module exchanges information among cell representations within each row after column encoding. The refined cells then visit the context set again through an additional column stage before being compressed into row embeddings. For inference efficiency, both row stages use cross-attention through a fixed number of summary tokens, keeping their attention cost linear in the number of features. Pretrained on approximately 36M synthetic tables, Causilo delivers strong benchmark results across TabArena, BeyondArena, and ScoringBench, achieving frontier-level performance with substantially faster inference.
Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories
利用去噪轨迹全局调度以挖掘扩散模型的推理时计算潜力
Cao, Yuan, Tang, Yifu, Li, Hangqi, Zheng, Zeyu
Abstract
Diffusion models generate a sample by traversing a denoising trajectory, a sequence of stochastic noise-reduction steps that transforms pure noise into a draw from a target distribution. At deployment time, additional computation can improve sample quality without retraining: at each step, the sampler draws several candidate noise samples, scores the resulting predictions with a quality criterion called the verifier, and retains the best candidate at the cost of one network evaluation per candidate. This raises a resource allocation question: given a fixed budget of function evaluations, how should search effort be distributed across the steps of the denoising trajectory? We formulate this as a computational budget allocation problem. First, we show that, to leading order in the step size, the expected gain from evaluating $K$ candidates at a step factorizes into an endogenous, step-specific sensitivity parameter times a universal sample-size factor equal to the expected best of $K$ standard-normal draws. Second, for a fixed sensitivity profile, the optimal allocation solves a separable concave integer program with water-filling structure; at fixed total sensitivity, its advantage over uniform allocation increases with sensitivity dispersion in the majorization order. Third, we prove that when sensitivities vary across instances, no adaptive policy can avoid worst-case regret that grows linearly in the trajectory length, which motivates a design that anchors the allocation offline and adapts online only to recover instance-specific slack. We extend the analysis from independent random search to a broader family of local search operators, and instantiate it as an implementable algorithm. Experiments on three families of diffusion samplers show that the proposed allocation attains the quality of the uniform benchmark with 20 to 50 percent fewer function evaluations.
Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8 RL pipeline remains challenging. While previous works have focused on resolving train-inference mismatches using correction techniques like TIS, we reveal that full-pipeline FP8 RL still suffers from severe training instability, manifesting as anomalous mid-training entropy surges and garbled outputs. We trace this instability to a previously overlooked cause: compounded FP8 quantization noise distorts the importance ratio, disproportionately pushing negative-advantage tokens outside the trust region and erroneously zeroing out their gradients. As a result, pathological outputs are not properly penalized and accumulate over the course of training. To address this, we propose Calibrated Clipping, a dynamic method that aligns the FP8 clipping bounds with high-precision BF16 distributions by matching the lower-bound clipping quantile and rebalancing the upper bound accordingly. Extensive experiments across GRPO and DAPO algorithms, model scales from 8B to 32B, and multiple FP8 scaling granularities demonstrate that our approach successfully eliminates entropy surges and restores performance comparable to the BF16 baseline.
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, but fine-tuning them with reinforcement learning (RL) remains constrained by the cost of real-world robot interaction. Model-based reinforcement learning (MBRL) reduces this cost by using a learned world model to generate rollouts for policy optimization. However, it becomes computationally expensive as VLA policies and world models scale. Existing methods typically treat states equally, overlooking substantial differences in their utility for policy improvement. In this paper, we show that policy uncertainty helps identify states with greater potential for policy improvement. The policy exhibits high uncertainty at only a small subset of states, often during decision-sensitive stages where small action differences can alter task outcomes, suggesting that policy improvements at these states could be particularly valuable. Building on these findings, we introduce U-GROW, a lightweight, plug-and-play sampling layer that directs more model rollouts to these informative states. By modifying only the branched-start distribution, U-GROW can be integrated into existing MBRL pipelines without changing the policy optimization objective. Experiments in both simulated and real-world manipulation tasks demonstrate the efficiency and effectiveness of U-GROW, supporting the use of policy uncertainty to guide experience generation.
Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion
Merge++:基于免数据检查点反演的通用合并精炼方法
Pola, Aditya, Balasubramanian, Vineeth N.
Abstract
Model merging consolidates fine-tuned experts into one multi-task model without retraining. All existing data-free methods approach this problem entirely in weight space. Restricted to arithmetic on parameters, these methods never observe how each expert behaves, a signal that only emerges through forward evaluation. Accessing this behavioral signal requires inputs to evaluate on, which the data-free setting prohibits. We propose Merge++, a post-hoc method that addresses this by inverting the expert checkpoints to synthesize task-representative images, then distilling expert knowledge into the merged model using those images. Merge++ requires no additional data beyond the checkpoints themselves. It applies universally across merging algorithms and operates as a complementary refinement stage independent of the underlying weight-space method. The method consistently improves merging algorithms ranging from simple task arithmetic to state-of-the-art spectral methods, with average gains of +2 to +8 points and up to +25.9 on individual configurations.
Coreset selection picks a representative subset of the labeled training set to make training cheaper. However, it is usually evaluated by downstream accuracy at a fixed subset size, ignoring both the time spent selecting the subset and the training recipe behind each reported number. We introduce an end-to-end benchmark that standardizes downstream training and charges selection and training to the same auditable wall-clock budget, spanning 4 datasets from CIFAR-10 to ImageNet-1K, 11 selectors, 5 fractions, and 3 seeds, with over 1,500 released runs. Repeated-sampling work has shown that budget-aware evaluation already favors random strategies. Our two budget studies test whether that verdict survives when every selector is granted its most favorable operating point. Across eight wall-clock budget anchors on each of CIFAR-10 and Tiny ImageNet, no anchor is won by a sophisticated selector: every winner is class-balanced random sampling, repeated random sampling, or full-data training. In fixed-budget duels on ImageNet-1K, training on all data for fewer epochs beats every selection strategy we probe while also costing the least. A per-dataset cost audit shows that selection cost is dominated at every scale by a fixed full-dataset scan, so it cannot be amortized away by selecting a smaller fraction, and its absolute size does not extrapolate from one dataset to another. We further quantify when selection does pay back through subset reuse, and document 9 correctness fixes to a widely used codebase, one of which shifts a standard Herding baseline by nearly 6 points. Selection time is not free preprocessing, and an evaluation that ignores it measures the wrong quantity.
Computationally efficient safe exploration in reinforcement learning
强化学习中计算高效的安全探索
Murali, Shreeram, Deka, Shankar A., Baumann, Dominik
Abstract
Reinforcement learning in real-life applications requires safety guarantees during exploration. Typical reinforcement learning algorithms do not provide such guarantees, and many modifications that do rely on Gaussian processes (GPs), which have a large computational cost. We propose a computationally lightweight algorithm based on the Nadaraya-Watson estimator that safely explores and optimizes constrained Markov decision processes (MDPs). Our algorithm, \textsc{CoLSafe-MDP}, uses an estimator that scales in constant-time with bounds on the estimates, a significant improvement from its GP-based counterparts that scale cubically with the number of data points. We then evaluate its performance in a grid-based environment and on observational Martian terrain data.
Joint Domain-Class Modeling for Federated Learning Under Feature Skew
特征偏斜下联邦学习的域-类联合建模
Najafi, Sina, Tavassolipour, Mostafa, Shariatpanahi, Seyed Pooya
Abstract
Federated learning (FL) enables collaborative model training without centralizing private data, but performance often degrades under feature skew: clients share labels while the conditional input distributions $p_i(x\!\mid\!y)$ vary due to latent, client-specific appearance factors. We propose Joint Domain-Class Federated Learning (JDFL), a lightweight, optimizer-agnostic extension that makes this latent domain variation usable without sharing raw data. JDFL first infers domain clusters called pseudo-domains from brief local update signals. It then expands the classifier head to output $M\times C$, joint (domain-class) logits. This allows the model to represent domain-conditioned appearance while keeping a shared backbone. To train the expanded head we introduce two complementary supervision strategies based on simple intuitions: a similarity-aware soft-labeling that transfers evidence between nearby inferred domains while allowing domain-specific specialization, and a per-sample randomized target assignment that perturbs supervision across the joint outputs and serves as a low-cost training-time regularizer. JDFL integrates with existing standard FL methods (e.g., FedAvg, SCAFFOLD) with minimal changes. Empirically, both supervision modes consistently improve global test accuracy on standard domain-shifted image benchmarks; ablations and sensitivity studies show the gains stem from the proposed supervision and parametrization rather than mere capacity increase.
Token Utility Is Selection-Conditioned: Coupled Selection of Prompt Context and Response Supervision for Efficient Instruction Tuning
Token效用依赖于选择条件:面向高效指令微调的提示上下文与响应监督耦合选择
Wu, Can, Chen, Xinrui, Wu, Ou, Du, Yi
Abstract
Efficient large language model (LLM) instruction tuning requires selecting response supervision with supporting prompt context. Existing methods typically value both sides separately, risking selection-state mismatch between valuation and retained training subsets. BRIDGE (Budgeted Response-Prompt Interaction via Directional Gradient-guided Efficient Token Selection) captures selection-conditioned token utility through a shared validation-directed interaction surrogate valuing each side under the other's retained state. Budgeted alternating selection coordinates retained subsets by aggregating precomputed interactions over the current opposite-side subset to update conditional scores. Structure-aware projection converts conditional response scores into coherent supervision spans. Across three model families, BRIDGE leads compared selection methods overall in mathematical reasoning, code generation, and instruction following. In mathematical reasoning, its advantage over independent selection grows with compression.
Beyond Similarity: Coverage-Aware Prompt Selection for Time Series Forecasting with LLMs
超越相似性:面向大语言模型时间序列预测的覆盖感知提示选择方法
Ji, Daeun, Kim, Minkyoung, Kim, Dongkuk, Lee, Yohan, Kim, Beomsoo, Jang, Beakcheol
Abstract
Similarity-based retrieval is the dominant rule for conditioning large language models (LLMs) in in-context learning, retrieval-augmented generation, and prompt-based time series forecasting. The rule concentrates on near-duplicate candidates, an issue that has motivated diversity-aware retrieval but remains unexamined in other retrieval-conditioned pipelines. We study this issue using prompt-based time series forecasting as a test bed, where a learned prompt pool is retrieved by similarity. Dominant methods in this setting retrieve top-K entries by cosine similarity without redundancy control, producing a bias toward dominant temporal patterns while overlooking rare but informative events. We propose CASP-LLM, a coverage-aware semantic prompting framework that addresses this prompt selection bias by combining usage-tracking and saturating-gate techniques into a coverage regularizer that adds no learnable parameters. On six long-term benchmarks and the M4 short-term benchmark, CASP-LLM matches or improves on similarity-based LLM forecasters on most dataset-horizon settings, with the exceptions of Electricity, M4-Monthly, and the few-shot long-horizon setting. A controlled study locates the failure mode at the cross-batch usage level rather than per-retrieval redundancy: within-retrieval diversification such as MMR does not help, whereas regularizing anchor usage across training does.
LPINNs: First-Layer Gated Localization for Physics-Informed Neural Networks
LPINNs:面向物理信息神经网络的首层门控局部化方法
Chawla, Lakshay, Jain, Hardik
Abstract
Physics-informed neural networks (PINNs) use one shared representation over the computational domain, which can become difficult to optimize on long domains and for high-order operators. We study a minimal alternative: multiply the first hidden activation of an otherwise unchanged dense PINN by input-dependent localization functions, giving first-layer units receptive fields without partitioning the domain or adding interface losses. We screen 13 families of localization functions, in up to three parameterizations each, on a nonlinear harmonic oscillator (HO), a heat equation on a long spatial interval, and a manufactured four-dimensional (4D) fourth-order problem, with ten paired seeds throughout. Three configurations give large reductions in solution error at matched budgets: (i) Fixed Gaussian localization functions on the $2\pi$ HO domain cut mean solution RMSE from $4.8369\times10^{-1}$ to $8.83\times10^{-3}$ at 3k epochs. (ii) The inverse-quadratic family with learnable centers and widths cuts it from $2.896\times10^{-1}$ to $3.06\times10^{-2}$ on the $8\pi$ heat domain at 10k epochs. (iii) Fixed bump localization functions cut it from $1.75947\times10^{1}$ to $2.260\times10^{-1}$ on the $4\pi$ 4D domain at 10k epochs. Every paired seed improves in these three comparisons. The screen also shows that the mechanism is not a free win: on HO only 2 of 13 families beat the baseline, and 10 of the remaining 11 are 9 to 23 times worse; on 4D four families are non-finite and five are more than three orders of magnitude worse than the baseline. The inverse-quadratic family is the only one that beats the baseline on all three equations. Overall, these results show that first-layer localization can provide measurable improvements to baseline PINNs on long-domain and high-order problems.
We study the symmetric and antisymmetric parts of bilinear forms in the attention heads of trained large language models. We introduce an orthogonally invariant profile map from real bilinear forms to a three-dimensional simplex and observe that profiles of trained bilinear forms accumulate near profiles of rank-one bilinear forms. We prove that the symmetric part of a bilinear form in an attention head is the sum of a hyperbolic form and a zero form for a Zariski-dense subset of query-key matrices.
Multi-class open-set anomaly detection requires a model to characterize the normal acceptance domain formed by multiple heterogeneous subdistributions using only class-labeled samples from known normal classes, and to identify previously unseen anomalies at test time. Existing single-hypersphere methods cannot explicitly represent class-specific locations and acceptance ranges, while current multi-hypersphere or multi-class approaches do not fully integrate inter-class boundary constraints, learnable acceptance ranges, and interpretable decisions. To address these limitations, we propose Interpretable Multi-Hypersphere Deep Anomaly Detection (IMHD-AD). IMHD-AD constructs an independent hypersphere for each known normal class in a shared feature space. With target-inside and non-target-outside constraints, IMHD-AD embeds the class-specific hypersphere centers and radii directly into the final network layer and jointly optimizes them with the shared representation. The minimum signed boundary score across hyperspheres simultaneously determines open-set acceptance or rejection and provides a faithful geometric explanation of each decision. On MNIST, Fashion-MNIST, and CIFAR-10, IMHD-AD achieves the highest AUC in 28 of 30 open-set comparisons. A two-dimensional synthetic study further shows that model architecture must balance the compactness of known normal classes against the separability of unknown anomalies.
Chinese Translation
多类开放集异常检测要求模型仅利用已知正常类别的带标签样本,刻画由多个异质子分布构成的正常接受域,并在测试阶段识别未曾见过的异常。现有的单超球面方法无法显式表示各类别特定的位置与接受范围,而当前的多超球面或多类方法也未充分整合类间边界约束、可学习的接受范围以及可解释的决策。为解决这些局限,我们提出了可解释多超球面深度异常检测方法(Interpretable Multi-Hypersphere Deep Anomaly Detection,IMHD-AD)。IMHD-AD在共享特征空间中为每个已知正常类别构建一个独立的超球面。借助“目标在内”与“非目标在外”的约束,IMHD-AD将各类别特定的超球面中心与半径直接嵌入网络最后一层,并与共享表示进行联合优化。所有超球面中最小的带符号边界得分同时决定开放集的接受或拒绝,并为每个决策提供忠实的几何解释。在MNIST、Fashion-MNIST和CIFAR-10数据集上,IMHD-AD在30组开放集对比中有28组取得了最高的AUC。二维合成实验进一步表明,模型架构必须在已知正常类别的紧凑性与未知异常的可分性之间取得平衡。
WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models
波前解码:面向循环语言模型的并行化自推测解码
Ha, Hyeongju, Kim, Jae-Joon
Abstract
Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD), a training-free self-speculative decoding framework designed for looped language models. WFD exploits two properties of these architectures: intermediate recurrence outputs provide effective draft predictions, and weight sharing allows token states at different positions and recurrence depths to be processed in one batched recurrent-block call. WFD organizes these mixed-depth states into a diagonal wavefront, continuously drafting new positions at shallow depth while advancing earlier positions toward full-depth verification. Unlike the phase-separated draft-then-verify schedule, WFD therefore co-batches drafting and verification within the same recurrent calls, while rejected drafts are corrected using full-depth predictions. Across six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, consistently outperforming draft-then-verify. Cross-recurrence KV sharing further reduces wavefront KV traffic and increases WFD's speedup to 4.81x on Huginn-3.5B.
Optimizers for Diffusion Models: A Controlled Benchmark
扩散模型的优化器:一项受控基准研究
Bolatov, Arman, Shulgin, Egor, Li, David, Shtanchaev, Abduragim, Stich, Sebastian U., Panov, Maxim, Moulines, Eric, Richtárik, Peter, Takáč, Martin
Abstract
Discrete diffusion models now match autoregressive language models on several benchmarks, while the question of how best to train them has received far less attention: the optimizer is inherited from one paper to the next and never compared. New optimizers, meanwhile, are validated almost exclusively on autoregressive pretraining, a different objective on a different loss surface. We present a controlled optimizer benchmark across four diffusion formulations, to our knowledge the first for discrete diffusion: seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS-M, Schedule-Free) on masked diffusion (text8), uniform diffusion (QM9, and LM1B through the Gaussian duality) and Gaussian diffusion on images (CelebA-64), each on a task with published reference values. Every optimizer receives the same search protocol, and every winner is retrained at the full budget with three seeds. AdamW is a strong default but not always the right choice: it is beaten by a resolved margin on two of the four tasks, and the winner changes with the formulation, so the optimizer deserves the same care as the rest of the training recipe. Notably, methods validated on autoregressive language model pretraining transfer well: Muon, MARS-M and SOAP each beat the tuned AdamW on at least one diffusion formulation. The benchmark, all runs and every figure are reproducible end to end from the released code at https://github.com/armanbolatov/diffusion-baselines.
MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs
MolSC:利用取代基贡献增强大语言模型的细粒度分子理解能力
Park, Hyuntae, Kim, Sooyeon, Park, Jiwon, Lee, SangKeun
Abstract
Recent advances in natural language processing have led to molecular Large Language Models (LLMs) with strong performance across diverse chemistry tasks. However, they still struggle to capture fine-grained structure-property relationships, particularly how small, localized modifications alter a molecule's behavior. To address this limitation, we introduce MolSC, a dataset of substituent contributions, defined as property changes induced by attaching specific substituents to molecular scaffolds. Curated from manually annotated bioactivity records, MolSC spans structural-alert liability, target-specific bioactivity, and physicochemical descriptors, and contains 181K substituent-level examples for training. We further propose MolSC-Bench, a held-out evaluation benchmark of 1,541 examples disjoint from MolSC at the scaffold, substituent, and molecule levels. Our experiments show that existing molecular LLMs and strong proprietary models such as GPT-5.2 and Gemini-3-Flash show limited reliability in substituent contribution prediction. In contrast, training on MolSC substantially improves this ability and achieves strong performance across diverse downstream molecular tasks. These results highlight substituent contribution learning as a key component of fine-grained molecular understanding.
AirGC-CD: Gaussian-Circulant Precoding for Exactly Debiasable PAPR Reduction in Over-the-Air Federated Learning
AirGC-CD:面向空中联邦学习中可精确去偏PAPR降低的高斯-循环预编码
Jang, Jonggyu, Lyu, Hyeonsu, Yang, Hyun Jong
Abstract
Over-the-air federated learning lets edge devices transmit their local updates simultaneously, reducing the communication overhead. The resulting waveform, however, has a peak-to-average power ratio (PAPR) that grows with the model dimension, and keeping the amplifier in its linear range leaves two remedies: clipping the peaks or backing off the transmit power. Neither remedy is without cost: i) the clipping distortion appears at the receiver as a bias that cannot be removed, and ii) back-off keeps the signal intact but degrades the average signal-to-noise ratio (SNR). Independent of this trade-off, the transmission remains uncompressed, spending one channel use per model parameter, which keeps large-model training out of reach. To address these challenges, we propose AirGC-CD, an over-the-air scheme that precodes each local update with a partial Gaussian circulant matrix before clipping. In AirGC-CD, the precoder's output is exactly Gaussian regardless of the update's sparsity, so the clipping function is designed for a known distribution instead of inheriting it from the data. This enables the clipping to be inverted on average by a single scalar Bussgang gain in closed form, and we prove that the resulting aggregate is exactly unbiased, with clipping adding only variance. The clipping ratio is then the only free parameter left, trading the variance of the clipping against the SNR loss from back-off, and we derive its near-optimum in closed form. Since the precoder is linear, it also acts as a compressor, reducing the transmission from the model dimension d to the sketch dimension m at a cost of only O(dlog d) via two fast Fourier transforms, whereas a Gaussian sketch costs O(md). Experiments on five image datasets show that AirGC-CD outperforms baseline over-the-air FL schemes in most settings, particularly at low SNR, while using fewer channel uses per round.
Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone
神经谱容量:仅凭网络规格度量与设计架构
Zhu, Chenyu, Zhao, Ruoyu, Lu, Zhichao
Abstract
Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for these decisions, #Params and #FLOPs, capture size and compute but not architectural structure: two architectures with identical parameter budgets but different depth-width, head, or FFN allocations receive identical scores yet behave differently. We propose Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix. Under standard random initialization, the Marchenko-Pastur law renders NSC computable from the architectural specification alone, with no model instantiation, data, or gradients. Its layer-wise additive structure admits NSC-DP, an exact dynamic-programming solver returning the architecture globally maximizing NSC under resource constraints in seconds on a CPU -- a guarantee that black-box search over existing training-free proxies cannot provide. Empirically, NSC outperforms #Params, #FLOPs, and representative training-free proxies in ranking across seven Transformer and CNN families (on FlexiBERT, $\tau = 0.505$ on pairs differing in #Params by less than 10%, where #Params collapses to 0.082); NSC-DP discovers a Transformer-XL architecture on WikiText-103 that beats the human-designed baseline in 2 seconds; and prunes LLaMA-7B to the best 5.7B model across eight commonsense reasoning tasks without any calibration data, about 5900x faster than the strongest training-free proxy baseline.
The Role of Coordinates in Pareto Regret for Adversarial Multi-Objective Bandits
坐标在对抗性多目标老虎机的帕累托遗憾中的作用
Guan, Changkun, Xu, Mengfan
Abstract
Adversarial multi-objective bandits hold the potential to help us optimize choices (arms) whose reward is a multidimensional vector chosen by an adversary and whose performance is measured by Pareto regret. We define loss as one minus reward and measure the easiness of a coordinate by the smallest cumulative loss of the arms on it, and call the coordinate easier when this quantity is smaller. Existing work suggests that in theory an easier coordinate may reduce Pareto regret. However, in practice, one may not know which coordinate is easier. On the negative side, we show that this lack of information eliminates the possibility: a smaller cumulative loss does not improve the worst-case order of Pareto regret. Precisely, let \(L_d\) be the smallest cumulative loss along coordinate $d$ over $T$ rounds. For \(K\ge4\) arms, \(T\ge6\) rounds, and at least 2 coordinates, we prove that the minimax expected Pareto regret is \(\Omega(\min\{T-L_0,\sqrt{K(T-L_0)}\})\). It is monotonically decreasing in \(L_0\), even when \(L_0=\min_d L_d\) itself is known. On the positive side, this result motivates the possibility that other coordinates, not just the easy one, may suffice to attain the optimal rate of Pareto regret. When $L_0$ is known, we apply Poly-INF to a fixed coordinate and obtain an upper bound on Pareto regret that exhibits the same order and thus matches the lower bound. Without such knowledge, we develop a reward-doubling version of Poly-INF that adapts to this unknown quantity while still attaining the matching minimax rate. Another implication is that it has no extra \(\log T\) factor and is independent of the number of coordinates.
When Does Adversarial Refinement Help? A Negative Result and Open Problem in Adapting R3GAN to Time Series Imputation
对抗式精化何时有效?R3GAN适配时间序列插补的负面结果与开放问题
He, Yufeng
Abstract
Diffusion models and transformers have supplanted GANs for multivariate time series imputation, largely on grounds of GAN training instability. R3GAN (NeurIPS 2024) removes that instability via regularized relativistic losses with provable convergence, raising a natural question: do stable, modern GANs revive adversarial imputation? We adapt R3GAN to 1D temporal data with a coarse-to-fine refinement framework and a frequency-domain discriminator, and audit 14 saved configurations across 3 datasets. Because these are heterogeneous single runs, the evidence is descriptive rather than a matched causal ablation. We report a negative result. All five saved mean/zero-start configurations improve by 48.4-70.2%. Among eight eligible non-legacy linear-start configurations, the mean change is -0.7% (range -3.0% to +1.1%); a separate -21.9% legacy logging anomaly is retained for provenance but excluded from that aggregate. In a saved Weather comparison, standalone R3GAN-1D underperforms BRITS by 5.8x. Crucially, we argue the common explanation (that GANs optimize distributional rather than point-wise objectives) cannot be the whole story, since diffusion models also optimize distributional objectives yet achieve state-of-the-art imputation. Our saved reconstruction-weight sweep is consistent with the adversarial signal being inert or harmful, but cannot identify its causal contribution; a matched discriminator-removed ablation is the key next experiment. We frame the precise reason a learned discriminator fails to provide useful refinement gradients (where a learned diffusion denoiser succeeds) as an open problem, and offer practical guidance on when adversarial refinement is worthwhile.
Whitening Inverts the Hierarchy: What the Norm of a Whitened Embedding Measures
白化逆转层级结构:白化嵌入的范数究竟度量了什么
Ahnouch, Mohammed, Elaachak, Lotfi
Abstract
Whitening a foundation-model embedding and using its squared norm as a training-free likelihood surrogate is motivated by the observation that whitened coordinates often appear approximately standard normal. We show that this observation follows from the projection central limit theorem and therefore does not imply a Gaussian joint distribution. Across multiple encoders and three training objectives, we find systematic over-dispersion of the whitened radius relative to the Gaussian reference, including against distributional clones with identical mean and covariance. We further show that the commonly reported agreement between empirical and theoretical norm statistics is an algebraic consequence of in-sample whitening and does not constitute evidence for Gaussianity. We identify the mechanism behind this behavior: whitening reverses the encoder's spectral hierarchy, shifting the contribution to the squared norm toward near-degenerate directions that encode predominantly noise. In these directions, the dominant variability is governed by a single input-dependent scale. We estimate this scale from two moments and use it to predict, without additional free parameters, the cross-dependence between disjoint spectral halves. These results indicate that the squared whitened norm is better interpreted as a Mahalanobis measure of semantic atypicality than as a log-likelihood. This interpretation explains both its practical effectiveness and its calibration failures: the statistic can rank and detect atypical samples consistently with nonparametric density estimates and across encoders trained with different objectives, while Gaussian tail thresholds can be inaccurate by orders of magnitude. etc.
Perplexity Cost Understates What Activation Quantisation Breaks
困惑度代价低估了激活量化所破坏的能力
Sathyanarayanan, Anish
Abstract
Activation quantisation is usually evaluated with an aggregate metric, perplexity, averaged over every token a model predicts. We ask whether that average identifies which computations a quantiser damages. Perplexity turns out to be a reliable aggregate signal: across 12 models from four families and 780 within-model comparisons, the arm perplexity prefers also retains more induction and more retrieval in all but 2.1 and 4.0 percent of cases respectively. But where perplexity has risen by only a factor of 1.2 to 1.5, induction still keeps 0.959 of its intact accuracy while retrieval has already fallen to 0.554, a gap the aggregate number does not surface. This gap has structure, not just size: a matched Gaussian-noise control of the same per-channel magnitude leaves it largely intact, and randomising only the sign of the quantisation error, every magnitude held fixed, is nearly as harmless, so magnitude alone does not explain the damage. Quantising in a rotated basis, which changes coordinate alignment without changing error magnitude, restores induction from 0.001 to 0.980 at three average bits per token in a single-block intervention, though retrieval recovers less completely at the same setting (0.694); end-to-end at four average bits, induction reaches 0.968 and retrieval 0.534. The pattern holds on two further models up to 32B parameters and, in the deployed configurations we tested, under AWQ once activations are pushed to 4 bits. A perplexity target bounds the average cost of a transformation applied to the activation; it does not, by itself, show which computations survived.
Provably Efficient Reinforcement Learning in Continuous-Time Episodic MDPs with Poisson Decision Epochs
具有泊松决策时刻的连续时间片段式马尔可夫决策过程中可证明高效的强化学习
Guo, Kenny, Iverson, Valentio, Wijetunga, Sahan, Chang, William
Abstract
Many real-world reinforcement learning (RL) problems evolve in continuous time, where decisions occur at irregular, event-driven intervals rather than at fixed discrete steps. We study episodic continuous-time Markov Decision Processes (MDPs) in which decision epochs are governed by a homogeneous Poisson process and the reward and transition dynamics vary smoothly over time. We consider both a fixed number of jumps per episode and a fixed time budget with a random number of Poisson decision epochs. Under a Lipschitz continuity assumption in time, we exploit local smoothness through discretization and extend both UCRL (Auer and Ortner 2006) and Q-learning (Jin et al. 2018) to this setting, proving $\widetilde{O}(T^{2/3})$ regret bounds for both model-based and model-free algorithms. Finally, we establish matching $\widetilde{\Omega}(T^{2/3})$ minimax lower bounds, showing that the rate is optimal up to logarithmic factors. These results provide the first tight regret guarantees for Lipschitz-smooth continuous-time episodic MDPs with Poisson decision epochs.
Chinese Translation
许多现实世界的强化学习(RL)问题在连续时间中演化,其中决策发生在不规则的、由事件驱动的时刻,而非固定的离散步骤上。我们研究了片段式连续时间马尔可夫决策过程(MDP),其中决策时刻由齐次泊松过程决定,且奖励和转移动力学随时间平滑变化。我们考虑了每片段固定跳跃次数以及固定时间预算下随机数量泊松决策时刻两种情形。在时间上的Lipschitz连续性假设下,我们通过离散化利用局部平滑性,将UCRL(Auer and Ortner 2006)和Q-learning(Jin et al. 2018)扩展到该设定,并证明了基于模型和无模型算法的 $\widetilde{O}(T^{2/3})$ 遗憾界。最后,我们建立了匹配的 $\widetilde{\Omega}(T^{2/3})$ 极小极大下界,表明该速率在对数因子意义下是最优的。这些结果为具有泊松决策时刻的Lipschitz平滑连续时间片段式MDP提供了首个紧致的遗憾保证。
Non-intrusive load monitoring (NILM) estimates appliance-level consumption from a whole-home meter, but appliance-specific models and fixed output inventories make coverage costly to extend. We present FM4NILM (Foundation Model for NILM), a single prompt-programmable model that estimates a requested appliance's power trajectory from aggregate measurements, a natural-language description, and optional activation exemplars. A lightweight cadence-aware transformer is pretrained by masked reconstruction on 645k sequences from seven public corpora spanning 1-60 s sampling intervals, then aligned with appliance requests using observation-masked losses for partially labeled households. A Bernoulli-lognormal decoder separates activity detection from conditional power estimation. On held-out households and time periods from REDD, UK-DALE, and REFIT, one frozen text-prompted model serves twelve appliance-corpus requests, achieving 0.556 event F1, 0.625 AUPRC, and the lowest active-window MAE (251.8 W) among seven appliance-specific baselines. Streaming score aggregation raises event F1 to 0.582 with a 60 s aggregation delay. In a separate category-held-out evaluation, adding ten activation exemplars raises microwave AUPRC from 0.132 to 0.214 without parameter updates. Input-intervention ablations probe the model's dependence on appliance requests and aggregate measurements. These results demonstrate competitive disaggregation with one shared model and support extending appliance coverage through prompts and examples rather than additional specialist networks.
Inverse design of RF and electromagnetic (EM) circuits is challenging because the relationship between circuit layout and electrical response is non-unique, and full-wave simulation is computationally expensive. This paper presents K-TRAIL, a simulator-guided generative framework for automated EM/RF circuit synthesis. K-TRAIL combines diffusion-based layout generation with derivative-free ensemble Kalman guidance, allowing feedback from a black-box EM simulator to refine candidate layouts during generation without requiring adjoint sensitivities or differentiable solver models. The framework supports both synthesis from prescribed S-parameter responses and synthesis directly from RF performance constraints. Experiments on multi-layer RFIC structures show that simulator-guided generation improves agreement with target responses and can identify structurally distinct layouts that satisfy circuit-level design requirements. The proposed approach provides a practical path toward generative, verification-aware RF circuit design while retaining the flexibility to explore diverse layout topologies.
Lossy compression of scientific simulation data increasingly relies on learned, latent-space architectures such as Residual Vector Quantization (RVQ), which iteratively quantize a base representation and its residuals to progressively reduce reconstruction error. While effective, RVQ performs this residual modeling entirely in latent space, leaving the pixel-space error structure of the reconstruction largely unaddressed. In this work, we propose a post-processing pipeline that augments an RVQ-based compressor with a U-Net trained to predict and correct pixel-space residuals between the original volume and its RVQ reconstruction. We show that these residuals are spatially structured rather than driven by local intensity or gradient features, motivating the need for a deep spatial model rather than simple statistical correction. The U-Net-corrected reconstruction is then passed through a Guaranteed Autoencoder (GAE) stage, which projects the remaining residual onto a per-block PCA basis to enforce a user-specified block-wise error bound. To the best of our knowledge, this is the first pipeline to combine latent-space RVQ, explicit pixel-space residual correction via a deep spatial post-processing network, and GAE-based error-bound guarantees within a single framework for scientific data compression. We evaluate our approach on S3D, JHTDB and E3SM datasets, demonstrating consistent improvements in NRMSE, compression ratio] over RVQ-only and standard residual-correction baselines, while maintaining strict error guarantees required for scientific data fidelity.
LLM explainers are increasingly attached to autonomous agents as runtime oversight, with operators reading a generated account of the agent's beliefs and actions rather than its internal state. We audit the account itself, pairing an Active Inference (AIF) agent that tracks German grid demand and adjusts generation with an LLM explainer on three backends (GPT-4o, Claude-3-Opus, Gemini), and probing the pair with three black-box triggers. Corrupting the observation stream by 600 MW per step moves the agent's posterior by 490 MW, roughly 0.9% of grid capacity. None of the 30 explanations produced during the injection flag anything under a stated rubric, and each narrates the corrupted belief fluently. On timesteps where the agent takes an objectively wrong action, all three explainers produce a sycophantic rationalization 80-95% of the time (n = 20 per backend). Attacker-controlled text in the observation metadata field steers the explainer, with susceptibility differing by provider and data exfiltration succeeding on all three. We propose mitigations for each failure but do not evaluate them. In every failure we observed, the explanation was fluent and wrong. Moreover, nothing in the explainer architecture checks whether an explanation is true before an operator acts on it. Testing the explainer therefore belongs in any audit of an agentic deployment.
Causal Inference with Unobserved Confounding: A Mixture Learning Perspective
存在未观测混杂因素的因果推断:混合学习的视角
Sood, Mansi, Shah, Devavrat
Abstract
Unobserved confounding is a fundamental challenge in causal inference from observational data. This article develops a mixture-learning perspective, viewing latent confounders as sources of heterogeneity that induce mixture structure in observed data. Under suitable structural and identifiability assumptions, recovering the mixing distribution and component mechanisms enables estimation of interventional distributions and causal estimands. Using variants of Bernoulli mixtures as a running example, we contextualize mixture-learning techniques and their structural assumptions, and connect them to causal inference in panel-data settings, including latent factor models and synthetic interventions.We then consider high-dimensional exponential-family mixtures with dependent outcome trajectories, moving beyond counterfactual means to model counterfactual distributions. We situate this perspective relative to complementary approaches for unobserved confounding. Together, these ideas provide a bridge between mixture learning and causal inference, connecting recent advances in high-dimensional mixture learning to scalable identification and estimation of causal effects while raising new challenges for mixture learning.
We study two-timescale decision systems in which a planning layer periodically supplies a continuation-value function to a real-time optimizer that allocates arriving resources, with inventory placement as our motivating application. We propose an end-to-end reinforcement learning (RL) method for learning this function using \emph{proximal residual value functions}, which combine a strictly convex potential of post-decision inventory with a learned convex residual. This general form yields a well-posed optimization layer that supports end-to-end differentiation while preserving an explicit convex objective for real-time execution. We characterize the necessary and sufficient conditions under which a smooth value function yields decisions that are consistent across the planning and execution timescales. In an offline simulation using historical inventory arrival and demand patterns from a large e-commerce retailer, learned proximal residual value functions reduce total routing and transfer cost relative to a historical-production-system proxy by 5.0%.
Chinese Translation
我们研究了一类双时间尺度决策系统:其中规划层周期性地向实时优化器提供延续价值函数,实时优化器则对到达的资源进行分配,并以库存调拨作为我们的 motivating 应用。我们提出了一种端到端的强化学习(RL)方法,利用近端残差价值函数(proximal residual value functions)来学习该函数。该方法将决策后库存的严格凸势函数与一个可学习的凸残差相结合。这种一般化形式产生了良构的优化层,既支持端到端微分,又为实时执行保留了显式的凸目标函数。我们刻画了光滑价值函数在规划与执行两个时间尺度上产生一致决策的充要条件。在基于某大型电商零售商历史库存到达与需求模式的离线仿真中,学习得到的近端残差价值函数相对于历史生产系统代理方法,将总路由与调拨成本降低了5.0%。
The Price of Self-Calibration: Exact Evidence Budgets and Manufactured Blind Sets in Adaptive Monitoring
自校准的代价:自适应监测中的精确证据预算与人为制造的盲集
Atarmla, Abdou-Raouf
Abstract
Self-calibrating monitors adapt their threshold online to guarantee a prescribed long-run false-alarm rate under arbitrary drift. We compute the price of that guarantee, stating every law with its exact domain of validity. First, the guarantee is an accounting identity, insensitive to what the monitor is meant to detect. Two evidence identities make the cost exact for the online quantile tracker: a persistent step of height $\delta$ yields excess alarm mass within one alarm of $\delta/\eta$, and exactly $\delta/\eta$ pathwise when $\delta$ is a lattice multiple of the gain $\eta$; a ramp of slope $c$ yields a stationary excess rate of exactly $c/\eta$, independent of accumulated size, up to a boundary $c=\eta(1-\alpha)$ coinciding with the alarm-rate cap. Second, the certificate's own fluctuation obeys an exact law: the windowed alarm rate has standard deviation of order $1/L$, not the binomial $1/\sqrt{L}$, since the windowed mass telescopes to a difference of a tight internal state; the closed-form constant is validated with no fitted parameter. Detectors calibrated on the binomial scale are miscalibrated by $\sqrt{\eta\varphi(q_0)L}$, and correct calibration turns detection windows from quadratic to linear in the inverse fault speed. Third, any monitor required to tolerate a drift class $\mathcal{D}$ is blind, at any horizon and for any rule, to every fault in $\mathcal{D}-\mathcal{D}$; the proof is a deliberately elementary two-point argument and the contribution is the object it identifies: for speed-bounded classes the blind set is exactly the doubled-speed class, and the tracker absorbs a speed class fixed by its own gain, so that under a certification regime declaring absorbed drift normal, the monitor manufactures $\mathcal{D}$. An exact Gaussian projection bound, sharper than Pinsker and never vacuous, quantifies power outside it.
CTRL: Control-Based Time Series Forecasting with LLM-Guided Residual Learning
CTRL:基于控制的时间序列预测与LLM引导的残差学习
Kim, Minkyoung, Ji, Daeun, Lee, Yohan, Kim, Beomsoo, Jang, Beakcheol
Abstract
Time series forecasting underpins critical decision-making across diverse domains. While large language models (LLMs) offer promising reasoning capabilities, existing LLM-based time series forecasting approaches either reduce them to numerical predictors that bypass their strengths, or allow direct forecast generation that destabilizes predictions in non-stationary settings. We introduce CTRL, a framework that decouples semantic reasoning from quantitative prediction. A frozen backbone generates base forecasts, while specialized LLM agents function as controllers that analyze backbone prediction errors through decomposed trend, seasonal, and irregular components, grounding reasoning in interpretable temporal structure. Each agent outputs compact control signals that a lightweight residual decoder translates into forecast corrections. CTRL incorporates label-free test-time adaptation that detects distribution shift from input statistics alone and readapts control signals with only 3-24 LLM calls via caching. CTRL is explicitly designed to improve robustness under non-stationary temporal dynamics and distribution shift, while remaining competitive on highly stationary time series where adaptive correction provides limited additional benefit.
Subliminal Learning (SL) is a recently identified phenomenon in which a student model acquires downstream task capabilities by matching seemingly unrelated auxiliary outputs from a teacher, despite never observing task labels, task-specific outputs, or the original training data. While recent studies have identified where subliminal signals may reside, the optimization mechanism underlying this phenomenon remains poorly understood. In this work, we provide a mechanistic understanding of SL through the lens of learning dynamics. Specifically, we derive a chained cross-task kernel that explicitly links ghost-output supervision to changes in task predictions through shared backbone representations. Our unified analytical framework provides a rigorous mathematical explanation for three central empirical puzzles in SL: (i) under shared initialization, the transfer operator forms a strictly Positive Semi-Definite (PSD) structure, guaranteeing that ghost-output optimization aligns the student with the teacher's true task objective without explicit label exposure; (ii) the ghost-output dimensionality acts as an explicit rank bottleneck governing the transfer of task-relevant features; and (iii) synthetic, high-entropy inputs function as broadband probes that maximize cross-task kernel overlap, explaining why random noise consistently outperforms structured data for subliminal transfer. Experiments on the canonical ghost-output setting validate all three theoretical predictions, providing the first learning-dynamics-based theoretical explanation of how ghost-output supervision gives rise to subliminal learning.
Optimal No-Regret Learning for Repeated Prophet Inequality
重复先知不等式的最优无悔学习
Wang, Kun
Abstract
We study repeated prophet inequalities under prefix feedback. In each of $T$ rounds, a learner encounters fresh values drawn independently from $n$ boxes with unknown $[0,1]$-supported distributions in a fixed order and must irrevocably accept one, observing only the prefix up to its stopping box. Regret is measured against the optimal stopping policy that knows the distributions. We give an efficient algorithm achieving $\widetilde O(\sqrt{T})$ expected regret, matching the lower bound up to logarithmic factors. Our algorithm explores directly through near-optimal policies, combining empirical backward induction with box-specific reach bonuses. A relative-drop aggregation rule then exploits the nesting structure of observed prefixes to preserve exploration, thereby removing the polynomial dependence on the box number $n$. This resolves an open question posed by Liu et al. (2025).
Optimal Multi-way Decision Trees for Stratified Sampling in Online Controlled Experiments
在线受控实验中用于分层抽样的最优多路决策树
Takei, Tomoka, Ikeda, Shunnosuke, Takano, Yuichi
Abstract
Online controlled experiments, or A/B tests, are widely used to estimate causal effects on digital platforms. A central challenge is to improve experimental sensitivity, or statistical power, without increasing the experimental sample size. Stratified sampling is a classical variance reduction technique; however, its effectiveness depends critically on how the strata are constructed. We thus propose an optimization-based stratification framework for stratified sampling using optimal multi-way decision trees. Our method, called Optimal Multi-way Stratification Trees (OMST), formulates stratification as a path-selection problem over a feature graph. The selected paths define interpretable stratification rules and are optimized using an exact variance-minimizing binary optimization formulation under continuous proportional allocation and a Neyman-type optimal allocation. We incorporate supervised optimal binning to generate outcome-relevant candidate splits for numerical features. Furthermore, we introduce reduction procedures for redundant candidate paths and assignment constraints, substantially reducing the optimization problem size. Experiments on both a real-world and a simulated dataset demonstrate that OMST achieves comparable or superior variance reduction to existing methods while maintaining shallow and interpretable stratification trees.
ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs
ValueDiff:面向汇点抑制型大语言模型的价值几何KV缓存淘汰方法
Park, Junyoung, Choi, Jungwook, Lee, Mingu
Abstract
Modern LLMs with QK-normalization, gated attention, learned attention sinks, or logit softcapping exhibit weaker persistent attention sinks, on which existing KV cache eviction methods primarily rely. We observe that across these models, weaker sinks co-occur with greater value-vector dispersion relative to key-vector dispersion. Motivated by this value-side dispersion, we present ValueDiff, a value-geometric eviction that ranks tokens by the L2 deviation of their value vectors from the cache mean. The same score arises as the minimal-disturbance eviction under a max-entropy assumption about future attention. We evaluate under fixed cache budgets, with eviction at every block boundary during prefill and at every decoding step during generation. On RULER at a tight 2k token budget, ValueDiff retains 88--99\% of dense across seven sink-suppressed models (best on 6 out of 7). On LongBench at the 4k budget, ValueDiff averages 92\% retention across sink-suppressed models versus 83\% for the strongest prior baseline. On MATH-500, ValueDiff is the strongest non-dense method on every sink-suppressed model tested at the 25\% cache budget, outperforming prior methods by up to $\sim$20 points on gated-attention models. Across all three benchmarks, value geometry emerges as the more reliable query-invariant eviction signal for sink-suppressed models.
CSC: Calibrated Simplicity for Conflict-Aware Social Bot Detection in the LLM Era
CSC:面向LLM时代冲突感知社交机器人检测的校准简洁框架
Qian, Yipeng, Zhao, Pengjie, Niu, Chaoxi
Abstract
Social bot detection is essential for protecting online platforms from misinformation amplification, coordinated manipulation, and distorted public discourse. However, large language models have made social bots much harder to detect from text alone because semantic camouflage is now cheap, fluent, and scalable. The resulting challenge is modality conflict: an account may look human-like in semantics while remaining suspicious in graph structure, profile attributes, or cross-modal consistency. Recent graph-based detectors tackle this limitation by adding graph-side complexity, such as sparse prototype selection, adaptive gating, or architecture-specific control logic, yet our experiments suggest that complexity alone is not the most reliable way to resolve such conflict. We therefore propose CSC, a calibrated-simplicity framework for conflict-aware LLM-era social bot detection. The framework combines three design choices: a simplified prototype-guided graph expert that retains useful structural biases while removing unstable graph-side heuristics, calibrated simplex-constrained fusion that aligns heterogeneous confidence spaces before late fusion, and a lightweight inconsistency expert that models cross-modal disagreement. Experiments on TwiBot-22, TwiBot-20, and MGStBot-large show that \textsc{CSC} improves calibrated operating-point decision quality while remaining competitive across external benchmarks. Further analyses show that calibration improves confidence reliability, the inconsistency expert mainly provides localized corrections in high-conflict or near-threshold regions, and simplified graph-side control yields a better stability-cost trade-off. A targeted semantic-camouflage stress test further shows that replacing selected bot text with matched human text sharply degrades the standalone text expert while leaving graph and fused evidence stable on a balanced challenge set.
Rethinking Class Imbalance for Single-Cell Foundation Models: A Systematic Benchmark Across Architectures and Long-Tail Loss Functions
重新思考单细胞基础模型中的类别不平衡问题:跨架构与长尾损失函数的系统性基准测试
Dong, Zeyu, Zhong, Jiahui
Abstract
Single-cell foundation models (scGPT, scBERT, Geneformer) achieve cell-type classification accuracy up to 97.5% in our experiments, yet this aggregate accuracy can mask systematic failure on rare, often disease-relevant cell populations that long-tail loss functions are widely assumed to address. We present a systematic benchmark of six long-tail loss functions (cross-entropy, weighted CE, class-balanced loss, focal loss, LDAM, logit-adjusted softmax) across three architectures and three datasets (Multiple Sclerosis, Zheng68K, human Pancreas), totaling 162 controlled training runs (3 backbones x 3 datasets x 6 losses x 3 seeds). The gap between overall accuracy, Macro-F1, and rare-class recall under plain cross-entropy is consistent across all nine (architecture, dataset) settings, driven by dataset structure rather than pretraining. Rare-class failure itself splits into two regimes with distinct embedding-geometry signatures, visible before any loss is chosen: some classes are recoverable by the right loss, while others retain linear separability yet are absorbed into unrelated classes' neighborhoods under every evaluated loss and architecture. Among the recoverable classes, the efficacy of reweighting is predicted by a class's absolute training-set size, rather than its share of the dataset or the dataset's overall imbalance ratio. Class-balanced loss and LDAM are the most consistent choices across all nine settings, while logit adjustment trades rare-class precision for recall rather than improving both. Our results give both a reusable benchmark and mechanism-grounded practical guidelines for combining foundation models with imbalanced biological data.
A Patient World Model for Early Forecasting of Digital Health Campaign Outcomes: Capabilities and Limits
用于数字健康营销活动结果早期预测的患者世界模型:能力与局限
Wang, Yunlong
Abstract
Digital direct-to-consumer (DTC) health campaigns are usually measured after the fact. In-flight forecasting commonly relies on a separate classifier for every cutoff and horizon. We treat this task as a dynamic-system problem and build a compact patient world model. The architecture maintains a latent state per patient, learns exposure-conditioned state dynamics jointly with a weekly conversion hazard, and rolls forward into future conversion curves. We evaluate it on a US campaign dataset with 147{,}173 patients and 5.2 million at-risk person-weeks. In a retrospective evaluation conditioned on recorded future exposures, the model forecasts the remaining new-to-brand prescription volume through week 52 with a relative error of 2.9\% from a week-4 cutoff and 0.8--2.6\% from cutoffs at weeks 8--26. The strongest non-recurrent baseline, a pooled-hazard gradient boosting model given the same survival rollout and information, has relative errors of 13.6--33.1\%. Per-horizon classifiers perform substantially worse. A Fisher-information analysis motivates dense next-exposure supervision when conversions are rare. Removing this auxiliary objective increases prescription-volume error by approximately $2$--$14\times$, while providing no consistent disadvantage on the more common specialist-visit outcome. We also evaluate scenario simulation. Switching all future exposure off raises predicted conversion from 0.31 to 0.89, a pattern consistent with selection effects in observational exposure data. This result highlights the limits of interpreting exposure-conditioned rollouts causally.
Recurrent models must preserve information that changes future behavior while suppressing hidden-state error. These objectives conflict: contraction improves stability, but contraction along a future-distinguishing direction destroys memory. We formalize this boundary through the predictive quotient of a recurrent state space. Two hidden states are equivalent when they induce the same conditional future; their equivalence classes form predictive fibers. Every exact semantics-preserving corrector acts as the identity on this quotient. At a regular point with hidden dimension d and predictive dimension k, it can eliminate at most d - k independent directions. This establishes a discrete-continuous boundary: finite predictive states admit positive-radius exact correction basins, whereas an uncountable continuum of future-distinguishable states cannot be decoded after arbitrary positive-radius perturbations in finite-dimensional Euclidean space. To operationalize this principle, we develop an auditable finite-future framework. A compact deployment bank W is evaluated against an independent audit bank A (W subseteq A) on a declared correction domain. Under generative probe access and audit-metric coverage, finite stochastic rollouts furnish a high-probability certificate for the separation margin Omega_{W|A}(delta). Preserving learned W-predictions within this certified margin guarantees bounded audit-semantic distortion. For intrinsic audit dimension k, the required probe outcomes scale as O(M * Omega^{-(k+2)}), where M = |A|; a matching minimax lower bound proves this exponent is optimal. Extending guarantees to continuous futures is achieved via an explicit completeness modulus. Controlled experiments validate the certified margins, scaling laws, and automated probe refinement under a safety-first evaluation paradigm.
Chinese Translation
循环模型必须保留会改变未来行为的信息,同时抑制隐状态误差。这两个目标相互冲突:收缩可提升稳定性,但沿未来区分方向上的收缩会破坏记忆。我们通过循环状态空间的预测商(predictive quotient)来形式化这一边界。当两个隐状态诱导出相同的条件未来时,它们是等价的;其等价类构成预测纤维(predictive fibers)。任何保持语义的精确校正器在该商上均表现为恒等映射。在隐维度为 d、预测维度为 k 的正则点处,校正器至多能消除 d - k 个独立方向。这确立了一条离散-连续边界:有限的预测状态可以拥有正半径的精确校正盆地,而不可数的、可区分未来的连续状态在有限维欧几里得空间中无法在任意正半径扰动后被解码。为将该原理可操作化,我们开发了一个可审计的有限未来框架。在一个声明的校正域上,将紧凑的部署库 W 对照独立的审计库 A(W subseteq A)进行评估。在生成式探测访问与审计度量覆盖的条件下,有限随机展开可为分离裕度 Omega_{W|A}(delta) 提供高概率证书。在该认证裕度内保持学习到的 W-预测,可保证有界的审计语义失真。对于内在审计维度 k,所需探测结果的数量按 O(M * Omega^{-(k+2)}) 规模增长,其中 M = |A|;一个匹配的极小极大下界证明了该指数是最优的。通过一个显式的完备性模量(completeness modulus),我们将保证扩展至连续未来。受控实验在安全优先的评估范式下验证了认证裕度、标度律以及自动探测精化方法。
The Evidence Ladder for Reinforcement Learning in Healthcare: From Retrospective Policies to Trusted Interventions
医疗领域强化学习的证据阶梯:从回顾性策略到可信干预
Zhao, Yunfan
Abstract
Reinforcement learning (RL) offers a natural language for healthcare decisions whose conse- quences unfold over time, yet most reported progress remains far from routine intervention. Ex- isting surveys organize the field by algorithm or clinical application. We instead review healthcare RL through an evidence ladder: problem formulation, retrospective identification, policy estima- tion, stress testing, prospective evaluation, and lifecycle monitoring. This view connects clinical treatment, patient engagement, and health-system operations while exposing a recurring gap: evi- dence that a policy scores well in a historical dataset is not evidence that it will improve care. We synthesize the assumptions and failure modes at each rung, identify what evidence can and can- not transfer across settings, and propose reporting practices for cumulative evaluation. Restless bandits are included as one special case, not as the organizing framework. The central lesson is that healthcare RL should be evaluated as an intervention embedded in a changing sociotechnical system, rather than only as an optimizer of a retrospective reward.
Before a machine can discover a physical law, it must discover what its measurements are: which observations live on cells, which are intensive or extensive, which sectors are dual, and which distinctions are merely gauge. We introduce physical representation-language discovery, the problem of recovering this hidden ontology directly from anonymous controlled experiments. We give an identifiability theory and constructive polynomial-time procedure that recovers a carrier and differential sequence, measurement types and orientation twist, noninvertible refinement semantics, primal-dual Maxwell diagrams, and the residual equivalences that no permitted experiment can break. The theory turns material nuisance into a commutant, uses refinement to separate quantities from coordinates, and selects physics only after its representation has been recovered. For a certified finite experiment family, we prove an end-to-end two-stage measurement bound and a matching minimax rate in dimension, accuracy, and confidence. Blind Maxwell experiments recover complete primal/relative-dual ontologies on regular and unstructured carriers under jointly corrupted observations; an independent unstructured RLC system demonstrates that the result is not specific to Maxwell. The framework scales to tens of thousands of cells per carrier, while stress audits demonstrate robustness across severe physical regimes - including non-Markovian memory, nonlinearities, nonlocality, and complex constitutive hysteresis. A public FDTD audit demonstrates the emergence of anonymous curl structure from incomplete field data, while characterizing the informational prerequisites for complete recovery. The goal is to move scientific ML from learning laws in a human-supplied language to discovering the language in which laws become expressible, establishing exact theoretical limits on observational identifiability.
Blind Thermodynamic Ontology Discovery from Anonymous Experiments
基于匿名实验的盲热力学本体发现
Zhang, Linzhe, Xu, Changming
Abstract
Before a machine learning model can learn a thermodynamic equation of state, it must discover what its measurements represent: which channels scale with system size, which are intensive conjugates, how sectors pair through contact, and which potential governs stability. When sensors expose only an unknown linear mixture of extensive states and intensive responses, passive observations cannot disentangle physical quantities from coordinate artifacts. We formulate the problem of discovering this hidden thermodynamic ontology directly from anonymous controlled experiments. We present an operational identifiability theory and a constructive polynomial-time algorithm that extracts extensive and intensive scaling sectors from replication contrasts, recovers their dual cotangent pairing from thermal contact and reciprocity, verifies a globally admissible concave potential via discrete cyclic concavity, and determines an invariant matroid of reservoir ensembles. We prove that the residual observational equivalence is strictly (x, lambda) ~ (A x, a A^{-T} lambda + beta), establishing the sharp observational limit that no permitted experiment can break. Blind evaluations on van der Waals fluids and Curie-Weiss magnets confirm robust recovery under ill-conditioned mixing, correctly resolving anonymous Maxwell tie-lines while rejecting non-equilibrium continuations. External validation across six real fluids from the NIST WebBook demonstrates that operational ontology discovery transfers across real physical substances without coordinate leakage.
Chinese Translation
在机器学习模型学习热力学状态方程之前,它必须先发现其测量数据所代表的物理含义:哪些通道随系统规模伸缩,哪些是强度量共轭对,各扇区如何通过热接触配对,以及哪个势函数主导稳定性。当传感器仅暴露广延态与强度响应的未知线性混合时,被动观测无法将物理量与坐标伪影分离开来。我们提出了直接从匿名受控实验中发现这一隐藏热力学本体的问题。我们给出了一套可操作的可辨识性理论以及一个构造性的多项式时间算法:该算法从复制对比中提取广延与强度标度扇区,从热接触与互易性中恢复其对偶余切配对,通过离散循环凹性验证全局可容许的凹势函数,并确定储库集合的不变拟阵。我们证明,残余的观测等价性严格为 (x, lambda) ~ (A x, a A^{-T} lambda + beta),确立了任何被允许的实验都无法突破的精确观测极限。在范德华流体与居里-外斯磁体上的盲评估表明,即使在病态混合条件下也能稳健恢复,能够正确解析匿名的麦克斯韦双节线,同时拒绝非平衡延拓。基于 NIST WebBook 中六种真实流体的外部验证表明,可操作的本体发现能够在真实物理物质间迁移,而不发生坐标泄露。
Multi-omics sequences contain complex biological patterns, yet deciphering their mechanisms for automated scientific discovery remains challenging. As large language models (LLMs) interpret these sequences, evaluating both predictions and scientific reasoning is critical. However, existing benchmarks for multi-omics sequence tasks rely on classification and regression metrics, neglecting whether models grasp the underlying biological evidence. We introduce OmicsBench, the first reasoning benchmark for multi-omics sequences, comprising 1,160 expert-validated questions across six tasks spanning DNA regulation, RNA processing, and protein function. OmicsBench requires traceable evidence chains, evaluated using instance-specific rubrics developed with domain experts. Evaluating 17 LLMs reveals an inverse relationship: while scientific LLMs outperform general-purpose LLMs in sequence classification accuracy, they fail to provide valid evidence to support their predictions. One plausible interpretation is shortcut learning: specialized models may rely on statistical patterns rather than the biological mechanisms needed for scientific discovery. Motivated by this finding, we introduce tool-augmented on-policy distillation (TA-OPD), a post-training method to align sequence prediction with evidence-grounded biological reasoning. Across five Qwen3.5 models spanning 0.8B to 27B parameters, TA-OPD consistently strengthens biological evidence grounding while improving predictive performance on most tasks. These gains persist across model scales, indicating that stronger sequence reasoning does not arise solely from increased model capacity, but can be improved through evidence-aware training. Together, OmicsBench and TA-OPD provide a framework for diagnosing reasoning failures in multi-omics LLMs and a path toward models whose predictions are better grounded in biologically meaningful evidence.
Reinforcement Learning with Verifiable Rewards (RLVR) is expanding from tasks with well-defined correctness signals, such as mathematics and code, toward multifaceted quality requirements specified by multi-dimensional rubrics. Since policy optimization consumes one scalar per rollout, rubric-based pipelines must map multiple criterion scores into a scalar reward. This aggregation is often treated as score scaling, but it implicitly determines how quality dimensions trade off during training. The prevailing practice, normalizing each criterion and taking a linear combination, assumes that cardinal score differences are comparable across criteria and that gains on one criterion compensate for failures on another; both assumptions are unreliable when criteria are semantically heterogeneous. We propose Reinforcement Learning with Verifiable Rubric-based Ranking (RLVR$^2$), a verifiable ranking paradigm for rubric-based RLVR. For each criterion, RLVR$^2$ converts rubric scores into criterion-specific within-group ordinal outcomes, recovers a latent utility from the resulting comparison matrix, and merges these utilities into one training signal. By retaining only within-group ordering and discarding raw score magnitudes, RLVR$^2$ avoids calibrating heterogeneous rubric scales. It further supports objective-preserving attribute adjustment: auxiliary attributes that correlate with observed rankings but are not training objectives can enter the estimation without expanding the rubric or rewarding them directly. Across three model scales and 16 benchmarks, RLVR$^2$ consistently outperforms representative rubric-based baselines, achieving the best overall performance on most benchmarks at every scale. Analysis shows it controls systematic effects tied to reasoning efficiency and response formatting while preserving the quality objective.
Chinese Translation
可验证奖励强化学习(Reinforcement Learning with Verifiable Rewards, RLVR)正从具有明确正确性信号的任务(如数学和代码)扩展到由多维度评分准则(rubrics)规定的多方面质量要求。由于策略优化在每次采样(rollout)中只消耗一个标量,基于评分准则的流程必须将多个准则的得分映射为一个标量奖励。这种聚合通常被视为简单的分数缩放,但它实际上隐式地决定了训练过程中各质量维度之间的权衡。当前主流做法是对每个准则的分数进行归一化后取线性组合,这假设了不同准则之间的基数分数差异具有可比性,且在一个准则上的提升可以补偿在另一个准则上的失败;当各准则在语义上异质时,这两种假设都不可靠。我们提出了基于可验证评分准则排序的强化学习(Reinforcement Learning with Verifiable Rubric-based Ranking, RLVR$^2$),这是一种面向基于评分准则的RLVR的可验证排序范式。对于每个准则,RLVR$^2$将评分准则的分数转化为该准则特定的组内序数结果,从由此产生的比较矩阵中恢复潜在效用,并将这些效用合并为一个训练信号。通过仅保留组内排序信息而丢弃原始分数的数值大小,RLVR$^2$避免了对异质评分准则尺度的校准问题。它还进一步支持目标保持式属性调整:那些与观测到的排序相关但并非训练目标的辅助属性可以在不扩展评分准则或不对其进行直接奖励的情况下纳入估计过程。在三个模型规模和16个基准测试上,RLVR$^2$始终优于具有代表性的基于评分准则的基线方法,在每种规模下于大多数基准上取得了最佳综合性能。分析表明,该方法在保持质量目标的同时,能够控制与推理效率和回复格式相关的系统性效应。
Deep learning has advanced automated electrocardiogram (ECG) diagnosis, but the field's most accurate models, foundation models pretrained on millions of recordings, are not decision-pathway auditable: a clinician cannot trace a diagnosis to a physiological pathway or intervene on one. We propose TRACE, a Tractable Routing Autoencoder for Clinical ECG, whose 32-dimensional clinical latent space is specified in advance from domain knowledge rather than discovered by optimization. TRACE partitions this space into perfusion, structure, and conduction subspaces, routes each to its own diagnostic head by design, regularizes the partition with an orthogonality penalty, and reconstructs the ECG through a decoder that permits latent perturbation. On PTB-XL and Georgia, TRACE exceeds unconstrained classifiers and stays ahead of an ECG foundation model pretrained on ten million recordings, evaluated by linear probe on frozen features, at roughly an eighth of the parameter count. On the nine-label CPSC2018 cohort, which carries no structural class, the framework transfers with only the routing table re-specified to a perfusion/rhythm/conduction partition. Joint probe, erasure, and perturbation analyses verify the routing contract, and perturbing the depolarization and repolarization pathways modulates the reconstructed waveform. Removing the specified partition and its orthogonality penalty costs 1.70 AUC and 11.30 macro-F1 points on PTB-XL, and 2.76 AUC and 16.92 macro-F1 points on Georgia. A capacity-matched permutation control places arbitrary assignments within 0.34 AUC points of the ontology routing and leaves macro-F1 statistically level (p=0.619): the ontology supplies decision-pathway auditability at no macro-F1 cost.
Chinese Translation
深度学习推动了自动化心电图(ECG)诊断的发展,但该领域最精确的模型——在数百万条记录上预训练的基础模型——并不具备决策路径可审计性:临床医生无法将诊断追溯到某一生理通路,也无法对其进行干预。我们提出TRACE(Tractable Routing Autoencoder for Clinical ECG),一种面向临床心电图的易处理路由自编码器,其32维临床潜在空间由领域知识预先指定,而非通过优化发现。TRACE将该空间划分为灌注、结构和传导三个子空间,通过设计将每个子空间路由至各自的诊断头,并以正交惩罚对该划分进行正则化,同时通过一个允许潜在扰动的解码器重建心电图。在PTB-XL和Georgia数据集上,TRACE超越了无约束分类器,并在参数量仅约八分之一的情况下,通过线性探针评估冻结特征时领先于在千万条记录上预训练的ECG基础模型。在不包含结构类别的九标签CPSC2018队列上,该框架仅需将路由表重新指定为灌注/节律/传导划分即可实现迁移。联合探针、擦除和扰动分析验证了路由契约,且扰动去极化和复极化通路会相应调制重建波形。移除预指定的划分及其正交惩罚会使PTB-XL上的AUC下降1.70、宏平均F1下降11.30,在Georgia上AUC下降2.76、宏平均F1下降16.92。容量匹配的置换对照实验表明,任意指定的分配与本体论路由的AUC差距在0.34以内,且宏平均F1在统计上持平(p=0.619):本体论以零宏平均F1代价提供了决策路径可审计性。
ITSY: Causal Discovery From Irregular Time-Series Data
ITSY:从不规则时间序列数据中进行因果发现
Xu, Wenbo, He, Yue, Wang, Yunhai, Chen, Yueguo, Kuang, Kun
Abstract
Structural causal models for time series recover contemporaneous and lagged effects, but most methods require complete observation windows and become misspecified when samples are missing. We introduce ITSY, the first continuous-optimization method for causal discovery from irregular time series under a linear model. ITSY reformulates the structural equation so that prediction uses the nearest available history rather than the possibly missing current slice, and jointly imputes missing values while learning both graphs. A weighted reconstruction objective corrects the noise transformation induced by this reformulation. Across synthetic regimes varying missingness, scale, graph density, and noise, and on a real world benchmark, ITSY consistently improves graph recovery over representative SCM-based baselines, demonstrating the effectiveness of the proposed method. The results establish a focused solution for irregular linear first-order dynamics and clarify the assumptions required for nonlinear or higher-order extensions.
Feature Suppression and Differential Privacy for Residential Traffic Classification: A Two-Home Federated Study
住宅流量分类的特征抑制与差分隐私:一项双家庭联邦学习研究
Lipcsey-Magyar, Márton Pál, Pekar, Adrian
Abstract
Residential traffic classification supports service management, but learning across homes must account for heterogeneous traffic and privacy constraints. Privacy-aware training may impose uneven costs across traffic categories. We study this tradeoff in simulated two-client federated learning using 1.62 million preprocessed gateway-collected flows across six categories. We compare a full-feature baseline, feature suppression (FS), and differentially private stochastic gradient descent (DP-SGD) under one fixed record-level privacy setting. FS-mild excludes four timing features from 16 model inputs; it provides no formal privacy guarantee. With size-proportional aggregation, FS-mild achieves higher combined macro-F1 and worst-group F1 (the minimum per-class F1 across homes) than DP-SGD in all five seeds at both model capacities under stratified and temporal splits. The tested DP-SGD configuration incurs pronounced minority-category losses, especially in the smaller home, but FS-mild does not uniformly improve on the full-feature baseline. On stratified-split models, loss-based and shadow-model membership probes show near-chance aggregate discrimination without a consistent ranking across probes; this does not establish equivalent privacy. These findings support FS as an input-minimization baseline, not a substitute for formal privacy.
Predicting Out-of-Distribution Generalization of Neural Operators via Observable Spectral Error Decomposition
基于可观测谱误差分解预测神经算子的分布外泛化
Dong, Hang-Cheng, Cheng, Pengcheng
Abstract
Neural operators have emerged as powerful surrogates for solving partial differential equations (PDEs), yet their reliability under distribution shift remains a critical barrier to deployment. Existing approaches to out-of-distribution (OOD) generalization in operator learning are largely empirical and black-box: they report aggregate error metrics without explaining why errors arise or when they will grow. We propose a structure-preserving framework that makes OOD generalization predictable and auditable. Our key idea is to parameterize the learned solution operator as a spectral filter $h_\theta(\lambda)$ acting on the eigenvalues of the underlying elliptic operator, implemented via Chebyshev polynomial expansions and trained with a weak-form objective. This parameterization admits an exact decomposition of the energy-norm error into two observable components: a model-dependent spectral approximation term and a distribution-dependent spectral weighting term induced by the input. From this decomposition we derive three diagnostics: a conservative in-band supremum $\vareps_{\mathrm{sup}}$, a global RMS proxy $\vareps_{\mathrm{rms}}$, and a sample-dependent effective metric $\vareps_{\mathrm{eff}}(f)$. These diagnostics can be computed without access to ground-truth solutions. Through four controlled experiments, we show that $\vareps_{\mathrm{eff}}(f)\|f\|$ consistently predicts energy error under in-distribution, in-band spectral shift, out-of-band tail, and compound shifts, whereas global metrics can be systematically misleading. Our framework shifts OOD assessment of neural operators from black-box benchmarking to operator-structure diagnostics, providing a practical route to auditable scientific machine learning.
Causal discovery from observational data is a fundamental yet challenging task in scientific research. While existing approaches are primarily based on conditional independence tests, structure scores, or restrictive functional assumptions, we propose Decoupled Causal Discovery (DCD), a novel decoupling-based perspective that does not rely on these methodologies. DCD directly identifies the Markov boundary (MB) by decoupling non-target variables via weighting functions, such that only variables within the MB preserve dependence with the target under the decoupled distribution. Building on this, DCD iteratively constructs the Completed Partially Directed Acyclic Graph (CPDAG) by exploiting structural asymmetries within the MBs. We establish the theoretical identifiability, soundness, and completeness of DCD. Empirical evaluations demonstrate that DCD achieves strong performance, particularly excelling in challenging noise regimes.
Graph prompt learning enables parameter-efficient adaptation of frozen Graph Neural Networks to downstream tasks through lightweight prompt parameters. As routing becomes increasingly node-adaptive, however, independently optimized local decisions can collectively concentrate assignment mass on a small subset of a finite shared prompt bank, even when individual node--prompt matches remain locally meaningful. We propose MINT (Measure-INtegrity Transport), an entropically regularized optimal transport framework that formulates node-to-prompt adaptation as a globally coupled allocation problem. The transport cost favors local geometric compatibility, while a prescribed prompt-side marginal explicitly controls graph-wide prompt utilization. We further derive an exact variance decomposition that separates prompt-side geometric variance into retained prompt-update variation and within-node barycentric dispersion, together with a conditional stability bound for the frozen-encoder forward map. Across standard citation networks and additional heterophilic graphs, MINT remains competitive in few-shot adaptation. Controlled and end-to-end experiments further distinguish the roles of routing and topology: fixed-marginal routing controls graph-wide prompt utilization and has measurable end-to-end effects on citation networks, while topology augmentation provides a complementary, graph-dependent mechanism for addressing structural mismatch. Code is available at https://github.com/Ga1axy0051/MINT.
Physics-residual machine learning predicts oxygen-evolution catalyst activity beyond the training range from sparse polarization measurements
物理残差机器学习从稀疏极化测量中预测超出训练范围的析氧催化剂活性
Kim, Yong-Woon, Lee, Jihyeok, Park, Sungtae, Choi, Sooseok, Byun, Yung-Cheol
Abstract
Screening oxygen-evolution catalysts on combinatorial libraries requires deciding which candidates receive the remaining measurements. The deciding activity lies beyond each candidate's measured potential window and often above every activity recorded during fitting. We predict it by physics-residual machine learning: the Tafel equation extrapolates the candidate's own measured current and slope, a learned residual attenuated with feature-space distance corrects the magnitude, and an applicability-domain score identifies predictions above the training range before measurement. In a separately fabricated 322-candidate library, 282 above the training maximum, two measurements per candidate gave a mean absolute error of 0.203 mA cm$^{-2}$ against 1.330 for the selected data-driven machine-learning model. Errors inside the training range remained comparable, and 35 labelled catalysts were enough to fit it. In two independent datasets the same construction lowered the overpotential error by 29 to 52%. Campaigns can therefore shorten each measurement and still rank the most active compositions.
Chinese Translation
在组合材料库中筛选析氧催化剂需要决定哪些候选材料应获得后续测量。决定性的活性位于每个候选材料已测电位窗口之外,且往往高于拟合期间记录的所有活性。我们通过物理残差机器学习来预测该活性:Tafel方程基于候选材料自身测得的电流和斜率进行外推,一个随特征空间距离衰减的可学习残差对幅值进行校正,而适用域分数则在测量之前识别出超出训练范围的预测。在一个单独制备的包含322个候选材料的材料库中(其中282个高于训练最大值),每个候选材料仅需两次测量即可获得0.203 mA cm$^{-2}$的平均绝对误差,而所选的数据驱动机器学习模型的误差为1.330。训练范围内的误差保持相当水平,且35个有标注的催化剂足以完成拟合。在两个独立数据集上,同样的构建方法将过电位误差降低了29%至52%。因此,此类测量流程可以缩短每次测量的时间,同时仍能对最具活性的组分进行排序。
Global Ranks Survive, Selected Heads Shift: BOS-Sink Topology under 4-bit Weight-Only Quantization
全局排序得以保留,而被选中的注意力头发生偏移:4比特仅权重量化下的BOS-Sink拓扑
Chen, Kuanlin, Kuo, Chen-Wei, Ou, Cheng-En
Abstract
Sink-aware deployment may identify important first-token attention heads before a model is quantized, then reuse that map at the edge. We test when this shortcut is safe for 4-bit NF4 weight-only post-training quantization (PTQ). Our Sink Topology Consistency (STC) metrics separate global rank preservation, top-$k$ set overlap, and layerwise sink-mass shift, and distinguish per-input sensitivity from calibration-map transfer. Across Qwen2.5-0.5B, Qwen2.5-1.5B, and Llama-3.2-1B, global bf16-to-4-bit ranks remain high at 4,096 tokens ($\rho_s \geq 0.980$), yet top-$k$ Jaccard overlap is only 0.619-0.793, corresponding to 76.5-88.5% membership retention. The global statistic also masks local failures: terminal Qwen layers shift by 6.2-7.9x their model means, whereas Llama-3.2-1B shows low, nearly uniform drift. Under a C4-to-LongBench shift, cross-domain overlap degrades more than the within-domain precision comparison for both Qwen models, but not for Llama-3.2-1B. Matched-domain 4-bit recalibration reaches 90% of a split-half stability plateau at the smallest tested $n=8$ for both Qwen models and $n=32$ for Llama-3.2-1B, though not as a sharp threshold; for the two Qwen models, updating only selected layers does not reach the full-map stability criterion. On Jetson Orin NX, the 16-sample workload takes seconds for the two models with valid on-device sink measurements. The practical message is precise: global rankings often transfer, but discrete head sets, layer-local policies, and cross-domain calibration should be revalidated after quantization.
Chinese Translation
Sink感知的部署方式可以在模型量化之前识别重要的首token注意力头,然后在边缘端复用该映射。我们检验了这一捷径在4比特NF4仅权重训练后量化(PTQ)下何时是安全的。我们的Sink拓扑一致性(STC)度量将全局排序保持性、top-$k$集合重叠以及逐层sink质量偏移区分开来,并区分了逐输入敏感性与校准映射迁移。在Qwen2.5-0.5B、Qwen2.5-1.5B和Llama-3.2-1B上,从bf16到4比特的全局排序在4,096个token时仍然保持很高($\rho_s \geq 0.980$),然而top-$k$ Jaccard重叠仅为0.619-0.793,对应76.5-88.5%的成员保留率。全局统计量还掩盖了局部失效:Qwen末端各层的偏移达到其模型均值的6.2-7.9倍,而Llama-3.2-1B则表现出较低且近乎均匀的漂移。在从C4到LongBench的分布偏移下,两个Qwen模型的跨域重叠均比域内精度比较下降得更严重,但Llama-3.2-1B并非如此。在匹配域的4比特重新校准中,两个Qwen模型在最小测试规模$n=8$、Llama-3.2-1B在$n=32$时即可达到对半折稳定性平台的90%,尽管这并非一个明显的阈值;对于两个Qwen模型,仅更新选定层无法达到全映射的稳定性标准。在Jetson Orin NX上,16个样本的工作负载在两个模型上仅需数秒即可完成有效的设备端sink测量。实践结论是明确的:全局排序通常可以迁移,但离散的注意力头集合、层局部策略和跨域校准应在量化之后重新验证。
Cost-Aware Reinforcement Learning with Action Masking and Projection for Battery Energy Storage Dispatch under Suppressed-Spread Market Shifts
面向抑制价差市场变化下电池储能调度的成本感知强化学习:动作掩码与投影方法
Chen, Kuanlin, Kuo, Chen-Wei, Ou, Cheng-En
Abstract
Battery energy storage system (BESS) dispatch must preserve operational feasibility while declining price spreads reduce the margin available to pay for cycling. We study a proximal policy optimization (PPO) controller whose pre-selection physical action mask and emergency projection are separated from a causal, forecast-informed economic advisory. All forecast-dependent methods receive the same causal 24-step forecast and grid-side settlement. Across five PPO seeds, advice-on net profit is 30.59 and 18.04 USD per 336-hour T1 and T2 window, versus 36.77 and 22.94 USD for proxy-cost MPC; PPO remains below this reference in both periods. Advice raises T2 profit from 16.45 to 18.04 USD while reducing throughput, but is immaterial in T1. On disjoint weekly blocks, PPO is stable under daily, weekly, and blended seasonal forecasts, weakens under persistence, and remains below proxy-cost MPC. Paired diagnostics localize changes to the observed 5-10 USD/MWh regime with mixed SoC-dependent effects. An M0-M6 ablation shows that mask removal sends thousands of infeasible requests to projection, while removing both physical layers exposes ramp violations. The evidence separates economic screening from feasibility enforcement without claiming formal safety, lifecycle-optimal aging, or RL dominance.
Orthogonality in a LoRA factor does not by itself specify what the composed update protects: the answer depends on the task-start state, the parameterization, and the realized optimizer displacement. We formalize this question through Bilinear Optimization Divergence (BOD), an anchor-relative diagnostic of effective-update response on selected historical features. The finite-step analysis distinguishes two cases. In a shared adapter, protecting the routing displacement leaves a learned-anchor residual through the changing companion factor. In a fresh zero-output block, a feasible routing state can protect the composed update while both current factors remain trainable. These conditions yield Semi-Frozen Orthogonal Routing (SFOR) for shared adapters and current-block hard protection for cumulative O-LoRA; Weight Residual Projection (WRP) enforces the required displacement after the optimizer step. Controlled two-task traces verify the predicted residual paths, reducing normalized historical response from 19.12% to 0.005% in the shared family and from 7.72% to 0.002% in the cumulative family. Four-task experiments on Qwen3-8B characterize the resulting trade-offs: SFOR improves backward transfer (BWT) from -2.47 to -0.86 with nearly unchanged average accuracy (AA), while O-LoRA hard protection improves three-order mean AA from 80.27% to 81.30% and forgetting measure (FM) from 2.20 to 0.43. Component controls also show that stricter feasibility need not improve final task performance. Together, the analysis and evidence provide an architecture-conditioned account of which constraint to enforce, how to enforce it, and how to interpret its empirical value.
ETH-TraceBench: A Large-Scale Event-Stream Benchmark for Ethereum DeFi under Temporal, Protocol, and Contract Shift
ETH-TraceBench:面向时间、协议与合约分布偏移下以太坊DeFi的大规模事件流基准
Kirtac, Kemal, Maple, Carsten
Abstract
Ethereum decentralized finance (DeFi) provides a public, time-stamped record of transaction-level event streams, but the same public symbols can create strong machine-learning shortcuts. We introduce ETH-TraceBench, a benchmark for evaluating Ethereum DeFi representations under temporal, protocol, pool/infrastructure, and symbolic shift. The raw event universe covers January 2021-December 2025 and contains 1.35 billion transactions with logs and 5.01 billion raw log rows. Model evaluation uses a fixed 911,267-instance supervised sample, training on 2021-2024, selecting models on 2025H1, and testing on 2025H2. Simple models perform strongly on the aggregate temporal test: TraceStats-GB reaches 0.953 macro-F1 and TopicEmitterHashMLP 0.959 on the canonical DEX test set. Performance drops sharply under protocol novelty, with macro-F1 of 0.794, 0.743, and 0.766 for TraceStats-GB, TopicEmitterTrace-SGD, and TopicEmitterHashMLP, while strict unseen-pool scores remain 0.927, 0.897, and 0.935. Uniswap v4 and Ekubo v1, both absent from supervised training, are materially harder than the full test. Jointly masking emitter and topic identity reduces DEX macro-F1 to 0.916 and liquidation macro-F1 to 0.774 for TopicEmitterTrace-SGD. A standard Transformer over log-index-ordered events provides no consistent advantage over a deterministic shuffle of the same events, indicating that high aggregate scores can arise without sophisticated chronological modeling. A natural-prevalence audit estimates 2025H2 DEX prevalence among logged Ethereum transactions at about 22.5%, and a deterministic 400-transaction audit finds complete agreement with task label sources and independently re-queried raw-log counts. ETH-TraceBench therefore treats difficult transfer and controlled-input conditions, rather than a single aggregate score, as the main evaluation target.
Patch-based autoregressive time-series forecasting often ties input representation, learned transitions, and recursive execution to one patch length. We ask which of these roles can be adjusted separately. A supporting atomic-encoding study finds greater sensitivity to model width than to atom grouping on the evaluated grid. Our main finding is that a frozen parent's recursive trajectory is easier to fit than the observed future with lightweight parallel exits. Autoregressive Trajectory Distillation (ATD) turns this into selectable ATD-1/2/4/8 execution, with ATD-1 exactly recovering the parent. On a paired four-data-set comparison, ATD-8 reaches $5.54\times$ end-to-end speedup with stable quality across widths. Fewer calls do not automatically remove the parent's existing forecast error: ATD improves trajectory fidelity in all 21 seed runs but forecast accuracy in only 15 against matched clean-future supervision. We further find a correctable residual projection along a train-selected periodic history direction. Spectrum Tangent applies this correction without adding neural parameters or Transformer calls. At horizon 720, it reduces mean squared error (MSE) and mean absolute error (MAE) by 2.54% and 2.33% over seven data sets and two output widths, while remaining $3.24\times$ faster than recursive inference. Level and shape projections sometimes disagree. Trajectory compressibility, the fidelity-accuracy mismatch, and the correction recur across three public AR parents. Together these results separate representation, transition, and execution as AR design axes. Code is available at https://github.com/RowanFFF/ATD-Spectrum-Tangent.
A multi-temporal dataset for mapping burned areas in the Brazilian Cerrado using time series of remote sensing imagery
利用遥感影像时间序列绘制巴西塞拉多(Cerrado)火烧区域的多时相数据集
de Oliveira, Alisson Cleiton, Körting, Thales Sehn
Abstract
This paper introduces a multi-temporal tabular dataset derived from satellite images to map burned areas in the Chapada dos Veadeiros National Park, in Goi\'as, Brazil, covering the years 2020 to 2022. The dataset contains blue, green, red, and near-infrared bands, as well as the BAI, EVI, GEMI, NDVI, and NDWI spectral indices from the WFI sensor on the CBERS-4A, CBERS-4, and AMAZONIA-1 satellites, organized into a regular grid. We applied the Random Forest classifier to develop and validate models based on samples labeled as totally burned, partially burned, and non-burned. Two classification approaches were tested: one combining burned and non-burned areas into binary classes and another distinguishing between totally burned (TB), partially burned (PB), and non-burned (NB) classes. Seven validation approaches assessed different post-classification combinations, focusing on accuracy, precision, recall, and intersection over union (IoU) metrics. Results showed higher IoU when TB, PB, and NB were used as individual classes and TB was reclassified as burned area (BA) while PB and NB were grouped as non-burned. Comparing the annual results of this approach to the MCD64A1 product, the errors of omission for the BA class were 22% in 2020, 28% in 2021 and 59% in 2022, while the errors of commission were 46%, 43% and 46%, respectively. The study highlights the utility of the WFI sensor for burned area mapping without inter-satellite spectral calibration and suggests further exploration with other machine learning algorithms to evaluate the dataset potential and limitations.
Chinese Translation
本文介绍了一个由卫星影像提取的多时相表格数据集,用于绘制巴西戈亚斯州韦阿代鲁斯高地国家公园(Chapada dos Veadeiros National Park)2020年至2022年的火烧区域。该数据集包含蓝、绿、红和近红外波段,以及来自CBERS-4A、CBERS-4和AMAZONIA-1卫星WFI传感器的BAI、EVI、GEMI、NDVI和NDWI光谱指数,并以规则网格形式组织。我们应用随机森林(Random Forest)分类器,基于标注为完全火烧、部分火烧和非火烧的样本开发并验证了模型。测试了两种分类方法:一种将火烧与非火烧区域合并为二元类别,另一种则区分完全火烧(TB)、部分火烧(PB)和非火烧(NB)三个类别。采用七种验证方法评估了不同的分类后组合,重点关注精度、精确率、召回率和交并比(IoU)指标。结果表明,当TB、PB和NB作为独立类别,且将TB重新分类为火烧区域(BA),同时将PB和NB归为非火烧区域时,IoU更高。将该方法的年度结果与MCD64A1产品进行比较,BA类的漏检误差在2020年为22%,2021年为28%,2022年为59%,而错检误差分别为46%、43%和46%。本研究凸显了WFI传感器在无需星间光谱校准情况下进行火烧区域制图的实用性,并建议进一步探索其他机器学习算法,以评估该数据集的潜力与局限性。
Tail-Weight Control and Localized Generalization in Nearly Low-Rank Adversarial Classification
近似低秩对抗分类中的尾部权重控制与局部化泛化
Wang, Kunyu, Wang, Dehan, Chen, Wenjun
Abstract
We study norm-constrained linear classification under Eu clidean adversarial perturbations in a Gaussian model with a low-dimen sional informative subspace and an independent noise tail. For bounded ramp loss, we prove that a principal-space witness with risk below one half forces every near-optimal predictor to have small tail weight. A path-specific density bound yields constants without requiring positive tail variance. Under isotropic principal covariance, we establish a unique population minimizer and joint local growth. Boundary normalization then removes the common attack penalty from centered margins, giving localized finite-sample guarantees governed by principal dimension and total tail energy. Globalized growth removes the entrance condition at weaker constants; a model-aware comparison retains local guarantees. Experiments with twenty paired repetitions show decreasing excess risk and tail use with sample size, and nearly unchanged behavior when tail dimension grows at fixed total energy. Pure-noise controls and optimizer diagnostics clarify the scope and limitations of these conclusions.
GenVoid: Uncertainty-Aware Learning of Subsurface Material Defects with an Experimentally Validated Physics-Informed Generative Model
GenVoid:基于实验验证的物理信息生成模型的不确定性感知内部材料缺陷识别
Mondal, Trishit, Bharadwaj, Prajwal, Karanjgaokar, Nikhil, Jagtap, Ameya D.
Abstract
Internal voids are ubiquitous defects in manufactured structures, yet their characterization remains challenging because their geometry is hidden and can only be inferred indirectly from accessible measurements. Here we introduce \textit{GenVoid}, a physics-informed generative model-based framework for identifying internal voids in complex two- and three-dimensional solids from surface displacement measurements alone. By incorporating the governing mechanics into a generative inference framework, \textit{GenVoid} enables void identification across linear elastic, hyperelastic and plastic material behaviours and accommodates complex two- and three-dimensional structural geometries. Importantly, the framework explicitly accounts for uncertainty and noise in displacement measurements, producing probabilistic reconstructions of internal void geometry rather than a single deterministic estimate. We demonstrate the approach using high-fidelity synthetic datasets and experimentally measured displacement fields obtained from in-situ mechanical experiments, establishing its ability to infer hidden voids from realistic displacement measurements. To quantify the fundamental limits of such inference, we further introduce an observability measure that characterizes the sensitivity of boundary measurements to localized stiffness perturbations within the interior under an ensemble of applied loads. This framework provides a direct connection between defect location, sensor configuration and reconstruction fidelity, enabling systematic assessment of how the number and spatial distribution of boundary measurements govern void-identification accuracy. To this end, these results establish a physics-informed and uncertainty-aware approach for non-invasive characterization of hidden defects and provide a quantitative basis for designing measurement strategies for inverse problems in solid mechanics.
Statistical Convergence of Transformer Encoder-Accelerated Robust Reinforcement Learning
Transformer编码器加速的鲁棒强化学习的统计收敛性
Banerjee, Suman, Tsukamoto, Hiroyasu
Abstract
Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces. In this study, we present statistically rigorous convergence results for a robust reinforcement learning algorithm warm-started by a transformer-based action-value function prediction, where natural language prompts encode task specifications. Our framework adopts the R-contamination model to characterize uncertainty in the state transition kernel, and employs conformal prediction to certify convergence via trajectory-level nonconformity scores constructed from the contracting Bellman residual. The resulting conformal quantile bounds the gap between the running and optimal action-value functions simultaneously over all iterations, thereby yielding a pre-certified stopping rule that requires little knowledge of the true transition kernel. Numerical case studies on perturbed maze environments of varying size and contamination level confirm that the transformer-based warm start measurably reduces the initial error and accelerates convergence, while the proposed conformal bounds track the true error trajectory more tightly than existing guarantees.
Many real-world decisions require prioritizing high-risk cases, such as clinicians prioritizing high-risk patients before lower-risk ones. Falling rule lists (FRLs), which are ordered if--then rules with monotonically decreasing risks, provide an interpretable framework for such tasks; however, their single-path structure yields a highly restricted model class. We introduce falling trees, a new family of interpretable models that enforces the same monotonic risk constraint while permitting tree-structured branching. We present GRAVITree, a novel dynamic-programming-with-bounds algorithm for learning the Rashomon set of falling trees under depth and branching constraints. Our formulation can interpolate between rule lists and full decision trees, enabling user-desired model expressivity. In a new clinical dataset and in many public classification benchmarks, falling trees match or outperform FRLs and other interpretable baselines, often producing more sparse decisions for high-risk instances. Our results show that falling trees strike a practical balance between interpretability, expressiveness, and risk prioritization for high-stakes settings.
Belted Engression: Sufficient Dimension Reduction for Generative Distributional Regression
带束约束的Engression:面向生成式分布回归的充分降维方法
Tan, Wenxi, Li, Bing, Xue, Lingzhou
Abstract
Modern conditional generative models face significant challenges when learning complex covariate dependencies. While sufficient dimension reduction (SDR) provides a principled approach to compress these dependencies, traditional SDR frameworks were not formulated for conditional generation. To bridge this gap, we propose Belted Engression, a unified and architecturally parameter-efficient framework for generative distributional regression. Our approach establishes an end-to-end compress-then-generate paradigm driven by sufficient representation learning, embedding a structural bottleneck into the generative architecture. Theoretically, we prove that the standard SDR condition is equivalent to a law-preserving generative factorization, which is achieved at the global optimum of the population Belted Engression objective. Furthermore, by uncovering a localized Bernstein-type control for the energy-score loss, we establish finite-sample convergence rates that are sharper than those of existing results. We also prove that this belted architecture is strictly smaller, operating with an asymptotically vanishing parameter count relative to the unstructured baseline. Extensive simulations and real-world applications demonstrate that Belted Engression achieves superior distributional prediction and SDR recovery with fewer trainable parameters.
Dictionary learning seeks to recover an unknown dictionary $A$ from observations ${\bf y}_i = A{\bf x}_i$ with sparse coefficient vectors ${\bf x}_i$. We introduce the \emph{Iterative Atom Refinement} (IAR) algorithm, a simple procedure for recovering individual dictionary atoms. Starting from a random direction, IAR repeatedly selects the observations most strongly correlated with the current iterate and updates the direction by averaging the selected data. Our main contribution is a rigorous convergence theory of IAR. Using high-dimensional probabilistic estimates and a novel monotonicity principle for atom-selection probabilities, we show that a small initial advantage of one atom is amplified until that atom is isolated. Under our model assumptions, IAR identifies a generating atom after only three refinement steps. Numerical experiments support the theory and show that the resulting dynamics accurately capture the behavior observed in dictionary refinement.
Chinese Translation
字典学习旨在从观测数据 ${\bf y}_i = A{\bf x}_i$ 中恢复未知字典 $A$,其中系数向量 ${\bf x}_i$ 是稀疏的。我们提出了迭代原子精化(Iterative Atom Refinement, IAR)算法,这是一种用于恢复单个字典原子的简单方法。IAR 从一个随机方向出发,反复选择与当前迭代方向相关性最强的观测数据,并通过对所选数据的平均来更新方向。我们的主要贡献是为 IAR 建立了严格的收敛理论。利用高维概率估计以及一种针对原子选择概率的新型单调性原理,我们证明:某一原子的微小初始优势会被不断放大,直至该原子被成功分离。在本文的模型假设下,IAR 仅需三步精化即可识别出一个生成原子。数值实验验证了该理论,并表明所得动力学能够准确刻画字典精化过程中观察到的行为。
Real-time Generalizable Heart Valve Mechanics for Clinical Disease Assessment via a Physics-Conditioned Neural Operator
基于物理条件神经算子的实时可泛化心脏瓣膜力学临床疾病评估
Koohy, Shawn, Wu, Wensi, Jolley, Matthew A, Perdikaris, Paris
Abstract
Mitral regurgitation is the most common heart valve disorder worldwide, affecting over 2% of the global population, rising to at least 10% in adults over 75, and causing approximately 15% of valvular heart disease-related deaths. Yet only a minority of patients with severe disease undergo corrective surgery. Rapid assessment of valve mechanics could enable earlier, more precise intervention, but traditional finite element simulations remain too slow for clinical timelines and parameter sweeps. We introduce the Physics-Conditioned Neural Operator (PCNO), a transformer-based surrogate that predicts leaflet displacement, strain, and stress fields across mitral and tricuspid geometries, conditioned on systolic blood pressure and tissue properties. Trained on functional, regurgitated, and pathological valves, including tethering, P2 prolapse, and annular dilation, PCNO achieves up to a 15,260x speedup over fine mesh finite element simulations with comparable accuracy, identifies pathology class, and resolves diagnostic metrics within 3.5% error under out-of-distribution extrapolation.
Actionable Insights from Observational Data: The Case of Advanced Classes in K-12 Education
从观察性数据中获取可行动的洞见:以K-12教育中的高阶课程为例
Bajwa, Nabit, Hunter, Seth B., Das, Sanmay
Abstract
A fundamentally challenging question in K-12 education is about the effects of taking more advanced or challenging classes. It is particularly complex because students (and/or their parents) choose whether to enroll in these classes, making causal analysis challenging. In this paper, we begin to tackle this question by taking advantage of a novel dataset from a public school system in the US. This dataset records students' course enrollment decisions, prior academic histories, demographics, and subsequent outcomes around the time of a district-wide change that introduced optional open-enrollment advanced middle-school courses in subject areas. This is a rich observational dataset, but enrollment in advanced classes is driven by student characteristics and choices rather than random assignment. This creates a core identification challenge: the same factors that influence enrollment in advanced courses are also predictive of academic outcomes. As a result, simple comparisons between enrolled and non-enrolled students are confounded, and naive estimates may reflect underlying differences in student ability, motivation, or support rather than the impact of coursework itself. Our analysis shows that enrolling in advanced English courses has a net positive but modest effect on student achievement outcomes. However, these benefits are unevenly distributed: some students with relatively large predicted gains ("middle achievers" in prior years) are less likely to enroll than others. Some other groups (e.g. Black students and those with lower socio-economic status) also demonstrate significantly lower propensity to enroll. This gap between predicted benefit and observed enrollment illustrates how careful data analysis can extract actionable insights from large observational datasets, including identifying students who appear well-positioned to benefit but do not select into advanced options.
From Regional to Global: Transfer Learning for Atmospheric Transport Emulators
从区域到全球:面向大气传输代理模型的迁移学习
Clark, Jeff, Fillola, Elena, Keshtmand, Nawid, Santos-Rodriguez, Raul, Rigby, Matthew
Abstract
Greenhouse gas emissions estimates can be derived using inverse methods by combining atmospheric concentration observations with chemical transport models. The latter traditionally use physics-driven simulators such as Lagrangian Particle Dispersion Models (LPDMs), which are expensive to run and do not scale well to modern satellites' high resolution data. Previously we developed a performant atmospheric transport emulator that approximates LPDM outputs ("footprints") over South America ~1,000X faster than the UK Met Office's LPDM. Expanding towards global emulation is not straightforward, as atmospheric transport is regionally heterogeneous. This paper evaluates spatial transferability capabilities of models across four world regions: South America, East Asia, South Asia, North Africa using both region-specific and multi-region models, and leave-one-region-out experiments. Regional differences are characterised in the context of input variable and output footprint distributions. This work builds intuition in cross-region generalisation and transfer learning, aiding regional performance towards efficient global emissions estimates.
Chinese Translation
温室气体排放估算可通过反演方法实现,即将大气浓度观测数据与化学传输模型相结合。后者传统上使用物理驱动的模拟器,如拉格朗日粒子扩散模型(LPDM),其运行成本高昂,且难以扩展到现代卫星的高分辨率数据。此前,我们开发了一个高性能的大气传输代理模型(emulator),其近似拉格朗日粒子扩散模型输出("足迹")的速度比英国气象局(UK Met Office)的LPDM快约1000倍。然而,向全球范围扩展并非易事,因为大气传输具有区域异质性。本文评估了模型在南美洲、东亚、南亚和北非四个世界区域之间的空间可迁移性,采用了区域专用模型与多区域模型,并进行了留一区域(leave-one-region-out)实验。同时,基于输入变量和输出足迹分布对区域差异进行了刻画。这项工作为跨区域泛化与迁移学习建立了直观认识,有助于提升区域模型性能,以实现高效的全球排放估算。
Adaptive Determinantal Client Scheduling in Federated Learning
联邦学习中的自适应行列式客户端调度
Xu, Wen, Liang, Ben, Boudreau, Gary, Sokun, Hamza
Abstract
Scheduling clients for model training is critical in federated learning due to both data and system heterogeneity. Most previous works focus on the quality of the scheduled clients to achieve faster convergence, shorter wall-clock convergence time, or better average model performance. They rarely consider the diversity of clients, which is important to counter heterogeneity and improve performance for the worst-off clients. In this work, we advocate the use of determinantal point processes (DPPs) to model and enhance the diversity in client scheduling. We first design the kernel matrices of DPPs using gradient information and quality scores, which inherently enables a flexible quality-diversity trade-off. Applying fast MAP inference over DPPs, we propose Adaptive Determinantal Client Scheduling (ADCS) in FL. We further quantify the gradient approximation error of ADCS and develop convergence analysis for general biased client selection in FL with non-convex loss functions. We conduct comparative numerical experiments showing that ADCS outperforms state-of-the-art client scheduling algorithms, including both quality-based and diversity-based ones.
PROSE: A Theory of Optimal Stopping with Perishable Evidence for Peer Selection in Intermittently Connected Decentralised Learning
PROSE:间歇连接去中心化学习中面向同伴选择的易逝证据最优停止理论
Anagnostopoulos, Christos
Abstract
Decentralised federated learning removes the aggregation server but makes collaboration dependent on transient peer availability. In mobile and intermittently connected systems, evaluating a promising peer consumes contact time and may cause the exchange opportunity itself to vanish, so that the evidence a learner gathers about a peer is perishable: it decays because links expire and because peer models drift while old measurements age. This paper develops a self-contained theory of optimal stopping for the resulting peer-selection problem. We formalise a receiver's within-contact decision as a finite-horizon Markov optimal-stopping problem with costly information acquisition and a future-arrival outside option, and prove that it admits an optimal policy characterised by a reservation value (Snell-envelope structure). Around this formulation we prove: (i) stage-uniform, drift-aware concentration and a maximin certification rule that is correct with high probability together with a finite-sample identification bound; (ii) a mobility-aware value of-information stopping rule and comparative statics showing that higher link hazard lowers the value of continued probing and enlarges the stopping region; (iii) a closed-form value of waiting under marked-Poisson contact arrivals, together with a search-theoretic reservation value whose comparative statics we characterise; and (iv) a myopic-optimality theorem establishing that, in sufficiently volatile (monotone) mobility regimes, the one-step confidence-safe rule is a sound surrogate for the optimal policy and never stops prematurely. We instantiate the theory as PROSE (Perishable-evidence Reservation-value Optimal Stopping for Exchange), a lightweight, fully local policy, and delineate the static contact and drift-free limits in which classical sequential decision problems are recovered. The development is entirely analytical.
The rapid growth of resident space objects is increasing the complexity of space situational awareness sensor tasking, challenging classical optimization methods as they allocate finite, heterogeneous, and distributed sensing resources across ever-larger catalogues. Existing deep reinforcement learning approaches show promise in reduced settings, but fixed-dimensional state and action representations limit their ability to scale to large, dynamic catalogues and distributed sensing networks. We introduce VISTA (Variable-Entity Intelligent Sensor Tasking Architecture), a scalable deep reinforcement learning architecture for persistent uncertainty-driven catalogue maintenance across variable object populations and sensor configurations. VISTA combines physics- and mission-informed top-K retrieval with entity-centric attention, recurrent memory, and pointer-based action decoding, thereby keeping each agent's observation and action spaces independent of catalogue size. We evaluate VISTA across different scenarios, from fixed-size single-sensor benchmarks to large-scale space-based tasking and heterogeneous cooperative sensing. With 30 orbiting targets, VISTA recovers the catalogue 31.2% faster than the fixed-dimensional recurrent baseline. In the large-scale regime, VISTA reduces five-hour uncertainty by 97.5% relative to the strongest classical reference and by 99.3% relative to the recurrent learner. Zero-shot tests up to 20,000 objects reveal near-linear relations between sensing capacity, catalogue size, and recovery horizon. Learned policies also exhibit sensor modality adaptation and generalization to population and initial-uncertainty shifts. Together, these results demonstrate that VISTA provides a scalable framework for adaptive space situational awareness sensor tasking across large, distributed networks of heterogeneous ground- and space-based sensors.
Clinical multimodal models must often predict before all chest X-ray (CXR) and electronic health record (EHR) inputs are available. Existing approaches align observed representations, model missingness, or reconstruct across modalities, but do not jointly exploit within-patient and clinically similar inter-patient evidence. We propose GLR-MM, a Graph-Based Global-Local Reconstruction framework for early ICU mortality prediction. It maps five CXR-EHR modalities to a shared space, reconstructs missing embeddings through complementary local cross-modal and global graph-attention branches, adaptively fuses their estimates, and optimizes class-balanced prediction, reconstruction, and contrastive objectives. On 9,620 MIMIC-derived ICU stays, we evaluate 10%, 30%, and 50% random modality missingness with shared deterministic masks. MUSE performs better under mild and moderate missingness, whereas GLR-MM achieves higher AUROC and AUPRC at 50% by 0.0088 and 0.0249, respectively. These results indicate that graph-guided reconstruction is most useful when inputs are severely incomplete.
We present a collaborative streaming anomaly detection system for high-speed data streams that explicitly integrates human analysts into the decision loop. The system combines heterogeneous detectors and aggregates their outputs through a normalization-based weighted consensus, complemented by artifact-aware rules to stabilize anomaly scoring under deployment. To improve interpretability, it derives surrogate models that approximate the ensemble consensus and expose human-readable sensor conditions associated with anomalous behavior. Analysts can actively intervene by reviewing anomaly episodes, adjusting consensus behavior, and refining surrogate rules used for anomaly prediction, producing a human-adjusted ensemble. We evaluate the approach on an industrial stream with 260\,000 events and 3 anomalous episodes, showing robust detection and actionable human-AI interaction.
Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs
Circuit-Diff:一种基于事实性编辑的干预方法,用于在归因图中定位知识
Friedman, Edward G., Song, Xiangchen
Abstract
Mechanistic interpretability defines features as the fundamental units of a neural network and circuits as the weighted subgraphs that carry out its computation. Because individual neurons are polysemantic, Cross-Layer Transcoders (CLTs) were introduced as a way to approximate a model's circuits by generating an attribution graph. The nodes of that graph, however, are unlabeled features: reading a graph means pruning it and then working out by hand what each surviving node means. To make CLTs easier to use for circuit discovery, we introduce Circuit-Diff, which intervenes on the model itself with a low-rank factual edit and takes the features whose role in the attribution graph changes under that edit as related to the edited knowledge. On the edits we examine, the flagged nodes are not only detectors of the object token: read off the CLT's released feature dashboards, they include features for the history, geography and associations surrounding the old and new objects. We formalize the method, measure how reliable a frozen CLT remains after a factual edit, test the selected nodes causally by patching them on up to 24 CounterFact edits, give a case study, and release an open-source implementation built on the circuit-tracer package, together with two further tools (multi-prompt aggregation and rule-based supernode labeling).
GDN Tree-Scan: Served Tree Verification for Recurrent-Hybrid Language Models
GDN Tree-Scan:面向循环-混合语言模型的服务化树验证
Ma, Zhiyuan
Abstract
Tree speculative decoding verifies multiple candidate continuations in one target forward pass. For attention-only transformers, the verifier mainly needs an ancestry mask. Recurrent-hybrid language models break this assumption: a candidate row must also carry the recurrent state that native sequential decode would have produced along its root-to-node path. Otherwise, a verifier can use a correct attention mask while still conditioning on an impossible recurrent history. We present GDN Tree-Scan, a served verifier for Gated-DeltaNet hybrid language models integrated into vLLM. The system combines FlashAttention-2 tree-bias attention, branch-local GDN scan/replay, device-side multidraft commitment, and accepted-chain-only state publication. On the public Qwen3.6-27B-FP8 checkpoint, in a clean batch-one (B=1) SWE/Codex decode gate at temperature 0.6, a six-node root-branch tree increases committed tokens/event by 17.2% at near-native verify-forward time and reaches 23.88 token-weighted decode tokens/s versus 18.80 for native five-step MTP (E5), a 27.0% token-weighted decode-throughput gain. The per-request-equal latency view is +4.0%, and end-to-end task wall time remains prefill-heavy. Empirical equivalence evidence is scoped to recurrent-oracle probability-rescore (p-rescore) closure within the observed native flip floor, not a full distribution-distance proof.
Multivariate quantile regression via Kolmogorov-Arnold Networks
基于Kolmogorov-Arnold网络的多变量分位数回归
Polar, Andrew, Poluektov, Michael
Abstract
This paper introduces a novel algorithm for predicting conditional joint distributions of vector-valued targets in stochastic systems whose randomness is intrinsic rather than arising from observation errors or additive noise. Multivariate quantile regression also involves modeling conditional joint distributions but represents a less challenging task. It predicts the probability that vector-valued targets fall within predefined regions, identifies regions corresponding to predefined probability levels, or performs both tasks simultaneously. The proposed identification technique employs ensembles of Kolmogorov--Arnold networks (KANs) as flexible function approximators. Although the suggested technique is not theoretically restricted to KANs, KANs are particularly well suited to the proposed construction and are therefore used throughout this study. In addition to the training procedure, this work introduces a new discrepancy measure for joint distributions and a goodness-of-fit (GoF) test based on it. This GoF test was initially developed to validate and calibrate the proposed identification technique and is used here in an ad hoc manner. Although the test could be tabulated for broader use, such a tabulation is not pursued in this work. The test is also applicable more generally.
A discrete generative model of neuronal spiking activity on microelectrode arrays
一种用于微电极阵列神经元放电活动的离散生成模型
Tanveer, Md Sayed, Mostajo-Radji, Mohammed A., Wang, Ge
Abstract
Generative models of neural activity could help characterize tissue dynamics, compare experimental conditions, and simulate population activity for applications ranging from disease and drug-response studies to closed-loop experimentation. Existing approaches, however, typically assume a fixed set of sorted neurons, whereas high-density microelectrode arrays produce extremely sparse, array-wide binary spike volumes in which the observed subset of electrodes varies across assays. We introduce a discrete generative model that represents this activity using a shared vocabulary of spatiotemporal motifs. A residual vector-quantized autoencoder learns the motif vocabulary, while a factorized masked transformer predicts where activity occurs and which motif appears at each active location. We evaluate the model on 31 assays spanning human brain organoids and acute \emph{ex vivo} human hippocampal tissue. The learned motifs are broadly reused: assay identity explains only $9%$ of the entropy in motif use, and motif overlap across tissue types is comparable to overlap within them. When representation quality is evaluated independently of the generative prior, our approach achieves $5.2\times$ the voxel-level reconstruction average precision of a matched flat tokenizer. For masked completion and free generation, the full model achieves $1.4$--$2.6\times$ the site-level average precision of the matched generative baseline and outperforms it across all four families of generation metrics. These results establish a compact, reusable representation for array-wide spiking activity without learned assay-specific parameters, providing a scalable foundation for generative modeling across diverse neural preparations.
Matched-Input Estimates Differ in Sign Across Architectures: Auditing EEG Foundation Models on Motor Imagery
匹配输入的估计在不同架构间符号相反:对EEG基础模型在运动想象任务上的审计
Zhou, Kevin, Roy, Sparsh
Abstract
Pretrained EEG foundation models are increasingly proposed as general-purpose encoders for brain-computer interfaces, yet recent benchmarks disagree about when their representations transfer to downstream tasks. We audit LaBraM and CBraMod on motor imagery under a validation-locked protocol in which preprocessing, architecture, optimization, freeze depth, checkpoint, temperature, and method selection are determined using training-session data only. On four-class BCI Competition IV-2a, every supervised comparator evaluated here outperforms every foundation-model configuration, including validation-selected fine-tuning. We then examine a key confound: foundation models and task-specific decoders are normally evaluated using different input pipelines. Retraining three supervised architectures on the broadband arrays consumed by the foundation models produces matched-input accuracy differences of opposite sign across architectures: broadband input improves ATCNet by 0.078 accuracy while reducing EEG Conformer accuracy by 0.088. None of the three individual matched-input terms is significant after multiple-comparison correction at n = 9, so we treat the sign variation descriptively rather than as a formal architecture-by-pipeline interaction. These observed sign differences suggest that a single comparator may not provide an architecture-invariant decomposition of a pretrained-versus-supervised performance gap. The four-class deficit also does not reproduce uniformly across motor-imagery datasets: on two-class BNCI2014-004 we cannot detect the same separation between fine-tuned CBraMod and the supervised comparators. Finally, validation-fitted temperature scaling returns foundation-model calibration error to the supervised range despite substantially lower four-class accuracy.
The Neural Forcing for Three-Dimensional Incompressible Navier-Stokes finite time blowup
面向三维不可压缩Navier-Stokes方程有限时间爆破的神经强迫项方法
Li, Beibei
Abstract
We present a two-part neural framework for forced three-dimensional incompressible Navier--Stokes flow. Part~I develops the computational forcing system. A physics-informed neural model generates structured external-force trajectories, candidates are optimized through differentiable PDE rollouts or PPO-Clip, and selected forcings are frozen and checked by independent fixed-force replay. Part~II provides the mathematical certification layer. It separates neural candidate discovery from continuum analysis, derives integrated reciprocal-vorticity criteria that imply Riccati-type growth and finite-time loss of smooth continuation, develops a validated computational-to-continuum transfer strategy, and establishes a conditional positive-probability closure for a nondegenerate neural output law. The proof is complete at the continuum level.
MGRD: Compact morphology-gated residual diffusion for variance-aware cross-domain neurite forecasting
MGRD:用于方差感知跨域神经突预测的紧凑形态门控残差扩散模型
Hsieh, Tsung Yeh, Anitescu, Cosmin, Kim, Chunghwan, Webster-Wood, Victoria A., Zhang, Yongjie Jessica
Abstract
Tracking neurite morphology over time helps characterize structural changes during neuronal development and deterioration, but long-term time-lapse imaging is resource-intensive and difficult to scale. Forecasting future morphology could reduce this burden. Existing neurite digital-twin models such as gated spatiotemporal attention (gSTA) produce a single deterministic forecast without representing variability among plausible futures. We introduce Morphology-Gated Residual Diffusion (MGRD), a compact stochastic surrogate that jointly forecasts twenty future neurite-morphology frames from ten observed frames while conditioning on morphology features derived from the latest observation. On controlled phase-field trajectories, MGRD reduces trajectory-wise mean MAE by 9.7% relative to a matched control while updating 4.46 times fewer parameters. On human iPSC-derived neuron microscopy, MGRD improves all four reported metrics over gSTA, including a 39.6% reduction in trajectory-wise mean MAE and a 45.3% increase in skeleton F1. Without mouse-domain retraining or fine-tuning, MGRD also improves MAE and skeleton F1 on mouse cortical-neurosphere microscopy across 10-40-min sampling intervals and forecast horizons beyond 13 hours. Repeated sampling provides a case-level variance score for ranking forecast difficulty. Retaining approximately 60% of the lowest-variance cases reduces mean MAE by 17.6% on iPSC microscopy and 16.8% on simulation data. MGRD uses 1.01% of gSTA's parameters, requires less than one tenth of its training-update time, and generates a 50-step DDIM trajectory 7.9% faster when morphology features are cached. These results establish MGRD as a compact stochastic surrogate for neurite-morphology forecasting and case prioritization across simulation and microscopy datasets.
Chinese Translation
对神经突形态随时间变化的追踪有助于刻画神经元发育与退化过程中的结构性改变,但长期延时成像资源消耗大且难以规模化。预测未来形态可以减轻这一负担。现有的神经突数字孪生模型(如门控时空注意力机制 gSTA)只产生单一确定性预测,无法表征多种可能未来之间的变异性。我们提出了形态门控残差扩散模型(Morphology-Gated Residual Diffusion, MGRD),这是一种紧凑的随机代理模型,能够基于最近一次观测提取的形态特征作为条件,从十帧观测数据联合预测二十帧未来的神经突形态。在受控相场轨迹数据上,MGRD 相较于匹配的对照模型将逐轨迹平均 MAE 降低了 9.7%,同时更新的参数量减少至其 1/4.46。在人 iPSC 分化神经元显微图像上,MGRD 在全部四项报告指标上均优于 gSTA,包括逐轨迹平均 MAE 降低 39.6%,骨架 F1 提升 45.3%。无需在小鼠域上重新训练或微调,MGRD 在小鼠皮层神经球显微图像上,于 10–40 分钟的采样间隔以及超过 13 小时的预测时域范围内,同样提升了 MAE 和骨架 F1。重复采样可提供样本级的方差评分,用于对预测难度进行排序。保留约 60% 方差最低的样本,可使 iPSC 显微数据上的平均 MAE 降低 17.6%,仿真数据上降低 16.8%。MGRD 仅使用 gSTA 1.01% 的参数量,训练更新时间不足其十分之一,且在缓存形态特征后生成 50 步 DDIM 轨迹的速度提升 7.9%。这些结果确立了 MGRD 作为神经突形态预测与样本优先级排序的紧凑随机代理模型,可同时适用于仿真与显微图像数据集。
Simpler Methods Work Better for L1 Penalized Logistic Models and Large Datasets
更简单的方法在L1惩罚逻辑回归模型与大规模数据集上表现更佳
Raff, Edward, Holt, James
Abstract
Linear models with an $L_1$-norm penalty remain state-of-the-art for high-dimensional ($d > 1,000,000$) tasks, offering a straightforward method for solving real-world industry problems. Despite their widespread use in industry and utility, many $L_1$ solvers are not effective for general use, are prohibitively slow, and are ineffective in parallelization. This makes them difficult to train in an MLOps pipeline on large industry-scale corpora. In this work, we test several proposed ``state-of-the-art'' solutions from the literature and find that older methods are currently far superior for general use. We also identify several recommendations for academics to perform research that avoids erroneously overconfident results, which can prevent the transition to production use. Equally surprising, we find that a new and simple baseline, using LBFGS on a sub-gradient, is highly effective with minor tweaks, despite being dismissed in the literature for theoretical non-convergence. In practice, we find it is an easier-to-support and easier-to-scale method for production use.
Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models
开放权重大型语言模型中的临床风险分类错位与成本不对称性
Liu, Star S. D., Ding, Xiyu, Barrett, Robert B., Santamaria-Pang, Alberto, Dobbins, Nic, Lehmann, Harold P.
Abstract
How large language models (LLMs) integrate patient risk with clinical cost tradeoffs remains poorly understood. We investigated how four open-weight LLMs (Qwen-2.5-7B/32B and Llama-3.1-8B/70B) internally represent cost tradeoffs, how these representations relate to clinical predictions, and whether decisions shift as predicted by the specified cost direction and magnitude. Using a public diabetes dataset, we varied 11 false-negative (FN) to false-positive (FP) cost ratios across three phrasings and examined representations and behavioral outputs. Patient risk was linearly recoverable on par with conventional classifiers (AUC $\approx 0.83$), and cost direction was recoverable in every model. However, representational shifts in cost direction tracked output changes only in the two larger models, and responses to cost magnitude were predominantly direction-agnostic. Only 2 of 12 model-phrasings showed both opposing responses to increasing FN versus FP costs and cost-correct ordering. Representationally, a direction fitted on one cost side did not invert when transferred to the other, as expected under mirror-symmetric encoding. These findings suggest that LLMs encode risk and cost information but do not reliably integrate them into cost-correct decisions. Clinical evaluations should therefore include tradeoff tests, phrasing sensitivity, and default operating points alongside predictive performance.
Text-controlled time series generation aims to synthesize sequences that follow natural-language descriptions while remaining faithful to real data distributions. Existing paradigms often couple semantic understanding and sequence modeling in a single continuous latent space, lacking explicit local semantic anchors and separation between global continuous attributes and local discrete shapes. As a result, key local structures may be smoothed, missed, or misplaced. We propose Shape Lexicon (ShapeLex), which decouples text-to-sequence generation into discrete symbolization of local shapes and continuous modeling of global attributes. ShapeLex first induces a reusable vocabulary of discrete shape units, such as rises, spikes, and sharp drops, from training data, forming an interpretable symbolic space. An autoregressive generator then selects shapes according to the textual description, adjusts attributes such as position and duration, and composes them in temporal order into a shape skeleton. Finally, a mixture-density scale head models and samples the overall level and volatility to restore realistic global scale. Experiments on twelve public datasets, real user-written text, and downstream forecasting tasks show that ShapeLex generates series that better match real data distributions than existing methods. In addition, paired supervision is automatically synthesized from the learned vocabulary, avoiding annotation costs that grow with dataset size and improving scalability.
Graph-to-Grid (G2G): Continuous-Coordinate Feature Painting for Soccer Pass Surfaces
Graph-to-Grid(G2G):用于足球传球表面的连续坐标特征绘制
Günay, Kaan, Gun, Orhun
Abstract
Dense pass surfaces give, for every pitch cell, whether a pass played there would arrive, whether the carrier would choose it, and what the possession would then be worth. The networks that draw them read the state as a raster of per-cell counts, losing where inside a cell each player stands. LiDAR detectors, bird's-eye-view perception and graph weather models move entity features onto a grid, binning each entity to a cell or learning the transfer. We evaluate the interpolated form: each player's features are scattered bilinearly onto the grid at the player's measured coordinates, so the surface loss trains the per-player encoder end to end. Those systems adopt an interface; this paper measures one. On 53,628 passes from the 2022 World Cup, painting improves selection likelihood over the same core fed rasters alone by about a quarter of a nat: in every match of an eight-fold cross-validation, with every arm tuned over five seeds, and after retraining on seven Bundesliga and 2. Bundesliga matches from another provider. Thirteen pre-specified studies locate the gain: painting the nine raw player features with no encoder carries three quarters of it, and the learned encoder and message passing add a smaller, resolved increment. Painting also helps the original SoccerMap and a canonical U-Net, whereas offset channels, a finer raster, an attention painter and a raster-free decoder do not. Frozen across the provider boundary the likelihood advantage is lost; injected tracking error compresses it. These results concern observed-endpoint prediction, not calibrated evaluation of hypothetical passes.
Edge deployment motivates forecasting models with compact parameter storage and low-bit representations. Deep equilibrium models (DEQs) obtain implicit depth by repeatedly applying a shared layer, reducing the parameter cost of explicit layer stacking. Their usual Anderson solver, however, searches for update coefficients in the continuous real domain. We propose Q-DEQ, which formulates local updates in DEQ forward solving as discrete optimization problems. Candidate directions are constructed from the current state and iteration history, and a local quadratic residual model is used to evaluate their combinations. Binary encoding of the direction coefficients yields a quadratic unconstrained binary optimization (QUBO) problem that can be solved by simulated annealing (SA) or a coherent Ising machine (CIM). After fixed-point solving, a re-forward pass applies W8A8 fake quantization to the shared layer's weights and activations. We evaluate Q-DEQ with an iTransformer backbone on five multivariate time series forecasting datasets. Relative MSE differences from the explicit multi-layer baseline range from $-1.16\%$ to $+2.90\%$, with lower MSE on two datasets. DEQ parameter sharing reduces parameter counts by factors of $1.80\times$--$3.82\times$; combined with W8A8, static weight storage is reduced by factors of $4.3\times$--$12.8\times$. Local QUBO problems solved using CPU-based SA and the Kaiwu CIM physical backend produce closely matching downstream forecasts. These results establish local discrete solving as a viable component of DEQ time series forecasting and provide a route for executing fixed-point updates through different combinatorial optimization backends.
FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention
FlashBoB:面向Softmax注意力机制的高I/O效率精确二阶反向传播
Givans, Anthony, Crawshaw, Michael, Liu, Mingrui
Abstract
Transformer models built on the attention mechanism have become a central building block in modern deep learning, yet softmax attention remains a major bottleneck for long-context workloads. While FlashAttention makes the forward and first backward passes I/O-efficient, it does not support backward-over-backward (BoB), which enables exact differentiation through the backward pass for applications such as second-order optimization, test-time training, gradient-based memory, and meta-learning. Existing BoB implementations either materialize large intermediate tensors or exhaust GPU memory at long sequence lengths. We present FlashBoB, an exact, I/O-efficient algorithm for BoB in softmax attention that keeps computation within on-chip tiles and avoids all $N \times N$ intermediate tensors, where $N$ is the sequence length. The key insight is a hierarchical affine structure in the softmax double backward: two row-wise scalars determine all outputs through affine transformations. This yields a two-pass schedule with bounded on-chip static random-access memory (SRAM) usage and minimal off-chip high-bandwidth memory (HBM) traffic. FlashBoB achieves $\Theta(N^2 d^2/M)$ HBM traffic ($d$ is the head dimension and $M$ is the memory size) and, within the standard FlashAttention-style score-recomputation model, matches the inherited large-cache lower bound for exact forward attention. Empirically, it scales exact attention BoB to $N=262\text{K}$ on a single A100 80GB GPU, where prior PyTorch exact baselines fail by $N=16\text{K}$, and is up to $6.3\times$ faster than FlashBack. These results make exact second-order attention practical at long-context sequence lengths where prior implementations cannot run efficiently.
Reinforcement Learning under State and Outcome Uncertainty: A Foundational Distributional Perspective
状态与结果不确定性下的强化学习:一种基础性的分布式视角
Preuett, Larry, Zhang, Qiuyi, Ahmad, Muhammad Aurangzeb
Abstract
In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. We address both forms of uncertainty as a first step toward safer algorithms in partially observable settings. Specifically, we extend Distributional Reinforcement Learning (DistRL)-which models the entire return distribution for fully observable domains-to Partially Observable Markov Decision Processes (POMDPs), allowing an agent to learn the distribution of returns for each conditional plan. Concretely, we introduce new distributional Bellman operators for partial observability and prove their convergence under the supremum p-Wasserstein metric. We also propose a finite representation of these return distributions via psi-vectors, generalizing the classical alpha-vectors in POMDP solvers. Building on this, we develop Distributional Point-Based Value Iteration (DPBVI), which integrates psi-vectors into a standard point-based backup procedure-bridging DistRL and POMDP planning. By tracking return distributions, DPBVI lays the foundation for future risk-sensitive control in domains where rare, high-impact events must be carefully managed. We provide source code to foster further research in robust decision-making under partial observability.
SPeaR: Test-Time Adaptation with Steering Primitives for Realigning Representations
SPeaR:基于导向原语的表示重对齐测试时自适应方法
Dip, Muhammad Sudipto Siam, Etemad, Ali
Abstract
Test-time adaptation (TTA) addresses distribution shift using only unlabeled test data. Existing methods typically adapt pretrained models by updating their parameters, limiting both what is adapted and where adaptation can occur within the network. We instead keep the pretrained network frozen and steer its intermediate representations. We introduce SPeaR (Steering Primitive for Realigning Representations), which inserts lightweight learnable modules at stage boundaries and optimizes them directly from the test stream, requiring neither source data nor supervised warm-up. Each primitive is optimized using a gated objective that reduces uncertainty only when adaptation is beneficial, along with a diversity regularizer to prevent collapse, and a multi-depth anchor to stabilize adaptation. We show that steering early representations is the most effective strategy, and that the same primitive transfers across convolutional and Transformer architectures. Across CIFAR-10-C, CIFAR-100-C, and ImageNet-C, SPeaR consistently matches or outperforms methods that adapt orders of magnitude more parameters, remains robust across a wide range of batch sizes, and preserves source-domain performance during continual adaptation.
Chinese Translation
测试时自适应(Test-Time Adaptation, TTA)仅使用无标签测试数据来应对分布偏移。现有方法通常通过更新预训练模型的参数来进行自适应,这既限制了可自适应的内容,也限制了自适应在网络中可以发生的位置。我们转而保持预训练网络冻结,并对其中间表示进行导向调整。我们提出了SPeaR(Steering Primitive for Realigning Representations,用于重对齐表示的导向原语),它在网络的阶段边界插入轻量级可学习模块,并直接从测试数据流中对其进行优化,既不需要源数据,也不需要有监督的预热。每个原语通过一个门控目标函数进行优化,该目标函数仅在有助于自适应时才降低不确定性;同时还引入了防止坍塌的多样性正则化器,以及稳定自适应过程的多深度锚定机制。我们证明,对早期表示进行导向调整是最有效的策略,且同一原语可以在卷积架构和Transformer架构之间迁移。在CIFAR-10-C、CIFAR-100-C和ImageNet-C数据集上,SPeaR持续达到或超越那些自适应参数数量高出数个数量级的方法,在广泛的批量大小范围内保持稳健,并在持续自适应过程中保持源域性能。
PAC-Bayesian Meta-Learning for Few-Shot Identification of Linear Dynamical Systems
面向线性动力系统少样本辨识的PAC-贝叶斯元学习
Huang, Chenfeng, Michailidis, George
Abstract
Identifying linear time-invariant (LTI) dynamical systems is challenging when trajectories are short, noisy, or high-dimensional. Traditional system identification typically treats each system independently and cannot exploit shared structure across related systems. We propose PBML-LTI, a PAC-Bayesian meta-learning framework for few-shot LTI system identification that learns a transferable prior over task-specific dynamics while preserving task heterogeneity. Each task corresponds to an unknown LTI system, and the meta-learner uses training trajectories to learn a data-dependent prior over transition matrices. For a new system with limited data, PBML-LTI performs Bayesian adaptation under this prior to obtain a task-specific posterior, providing accurate estimates and principled uncertainty quantification. A key challenge is temporal dependence, since LTI trajectories violate the i.i.d. assumptions underlying most PAC-Bayes meta-learning analyses. We address this with a martingale PAC-Bayes analysis for dependent trajectory losses and derive a support-query predictive-risk bound that motivates a fit-KL meta-training objective. The bound clarifies the roles of empirical fit, posterior complexity, and prior quality in few-shot adaptation under sequential dependence. We further derive corollaries for transition-matrix recovery and multi-step trajectory prediction, connecting uncertainty-aware meta-identification with finite-sample guarantees for dependent dynamical data.
CLOOPD: Closing the Learner Loop in On-Policy Distillation
CLOOPD:在策略蒸馏中闭合学习器回路
Zheng, Keye, Li, Hanyu, Cheng, Zhan, Gao, Yuan
Abstract
On-policy distillation (OPD) pays twice for each fresh batch: the student generates trajectories and a stronger teacher scores them. Existing methods improve which trajectories are scored and how the teacher signal is constructed, but usually consume it with one actor update. We introduce CLOOPD, a closed-loop framework separating teacher-signal acquisition from student-side realization. CLOOPD selects an adaptive $\alpha$ waypoint inside a KL envelope, freezes the scored batch and its advantages, re-forwards the student after each actor pass, measures realization, and allocates actor work under a separate token budget. The framework includes deterministic two- and three-pass policies, token-priced CLOOPD-TPMR, and a budget-matched control. Across six 300-step runs on an 8-H20 node, every CLOOPD policy improves the one-pass TOP-D anchor at comparable teacher-token scale: macro accuracy rises from 15.41 to 17.78 with CLOOPD-Fixed2 and 19.36 with CLOOPD-Fixed3. At step 100, CLOOPD-Fixed3 reaches 15.35, nearly matching TOP-D at step 300 while using 67.2% fewer teacher-scored tokens and 28.0% fewer GPU-hours. Earlier 8-A100 ablations show adaptive $\alpha$ eliminates observed trust-envelope violations; a third pass adds headroom. These results position CLOOPD as a framework for budgeting how fully students learn from teacher-scored tokens.
Luck Is Not Skill: When Do Paired Rollouts Help Group-Relative RL of LLM Agents?
运气并非技巧:成对轨迹采样何时有助于大语言模型智能体的组相对强化学习?
Sakib, Nazmus
Abstract
Group-relative reinforcement learning compares rollouts of the same prompt, but independent environment noise can obscure these comparisons. We study paired rollouts, which share an event-keyed noise schedule within each group while preserving each rollout's marginal distribution. Pairing removes the between-schedule component of reward-contrast variance, but need not reduce gradient variance. For one-sided grader noise, we derive an exact condition for reduction and give a counterexample in which reward contrasts improve while gradient variance increases. A controlled study trains a 2B tool-use agent under tool faults and grader flips, with three seeds per design. The protocol was registered with a disclosed, previously completed pilot. Under tool faults, pairing improves final noisy-test success by +5.1 percentage points on average, with all three seed differences positive, but misses the registered learning-curve criterion. The criterion is also missed under grader flips: the validation-AUC difference is +0.003 (95% interval [-0.029, +0.033]). A gradient probe on eight distinct checkpoints from two fault-trained trajectories finds lower mean-centered covariance traces under both noise types: 21 to 30% for grader flips and 40 to 63% for tool faults. These finite-sample measurements support the variance mechanism without establishing a general learning-speed benefit. The results distinguish improving reward comparisons, reducing estimator variance, and improving learning.
Mind or Message? Auditing Theory of Mind in Multi-Agent Social Simulation
心智还是信息?多智能体社会模拟中心理理论(Theory of Mind)的审计研究
Li, Cong, Chen, Cheng, Fung, Thomas, Rossi, Alex, Li, Yi
Abstract
Language model agents are increasingly used to simulate social interaction, and the resulting transcripts read as though the agents understand one another. We ask whether that appearance rests on a model of the partner's mind or on the surface record of what the partner said. We build a social simulation in which both questions have exact answers: 40 multi-issue negotiations whose hidden preference weights and whose full Pareto frontier are known by construction. Two model families negotiate across 160 dyads, every transcript is frozen before any measurement, and 2880 counterfactual probes then hold the evidence byte identical while moving one factor at a time: the reader's own stake, the partner's tone, an identity label, and the order of recursion. The agents are socially fluent and economically poor. They reach agreement in 96.2% of dyads with 0 protocol failures, yet only 0.7% of deals land on the Pareto frontier, they leave 20.5% of the available joint value unclaimed, and they miss the one issue on which their interests are perfectly aligned in 76.6% of deals; on the frontier and on that aligned issue, a package drawn at random from the set both sides would accept does as well. The probes locate the failure. Swapping only the reader's own payoff sheet, while the partner's words and offers stay identical, moves the inferred top priority by 15.0 percentage points, which is egocentric projection rather than inference, while a tone rewrite moves it by 5.3 percentage points and an identity label by 0.0. Most tellingly, an agent predicts what its partner believes about it 72.5% of the time while that partner's belief is itself correct only 51.2% of the time: the agents track the conversation far better than they track the mind behind it.
Speculative decoding accelerates large language model (LLM) inference by using a lightweight draft model to generate multiple candidate tokens that are verified by the target model in a single forward pass. Its speedup is largely determined by the acceptance length, yet existing draft-model training methods mainly optimize cross-entropy or Kullback-Leibler (KL) divergence as proxies. These objectives encourage distribution matching but do not directly optimize acceptance length, and the acceptance mechanism also differs between greedy and sampling-based decoding. In this work, we propose acceptance-length-aware training losses that directly optimize the expected number of accepted tokens within a speculative window. For greedy verification, we derive an expected accepted length (EAL) loss that explicitly maximizes expected acceptance length. For sampling-based decoding, we introduce a window total variation (WTV) loss that optimizes the overlap between temperature-scaled draft and target distributions while accounting for sequential acceptance dependencies. Both objectives can be further combined with a group-relative reinforcement learning stage (GRPO) using simulated acceptance length as the reward. Experiments across different target and draft models, tasks, and decoding settings show that our losses consistently improve acceptance length over KL-based training. WTV provides particularly strong gains under sampling-based decoding, while EAL better matches greedy verification. These results show that directly optimizing the acceptance objective, with losses tailored to the decoding mode, is more effective than conventional distribution-matching objectives.
Speculative decoding losslessly accelerates large language model inference by having a lightweight draft model predict future tokens for verification by the target model. Recent block diffusion drafters further reduce drafting latency by predicting multiple tokens in parallel. However, existing block drafters project target hidden states at every input position into a separate drafter-side KV cache, incurring per-request memory and KV-write overhead that grow with concurrency; directly reusing target KVs in place removes this cache but fails to sustain draft quality throughout the block. We propose a hybrid target-context injection method that complements direct target KV reuse with target hidden states only at the last input position, requiring no separate drafter-side KV cache. Building on this design, we propose H-Spec, a hybrid Mamba-attention parallel drafter that consumes the two target-context sources through complementary modules. Mamba modules are initialized with projected last-token target hidden states, while attention modules reuse target KVs in place. Despite its recurrent formulation, Mamba's parallel scan allows H-Spec to preserve block-parallel drafting. Across three target models and diverse tasks, H-Spec improves over the best baseline by 5.0--13.3% in mean accepted length and 5.3--12.6% in batch-size-1 inter-token latency speedup. Under concurrent serving, H-Spec consistently achieves higher throughput while maintaining lower KV cache utilization than baselines across evaluated concurrency levels.
Opinion Leader Dynamics: How Sparse Attention Shapes Token Clustering
意见领袖动力学:稀疏注意力如何塑造词元聚类
Liu, Jingkun, Song, Yue
Abstract
Sparse attention reduces the quadratic cost of global self-attention while retaining strong empirical performance, but how its restricted interactions shape the evolution of token representations remains theoretically underexplored. Modeling tokens as particles on the unit sphere, we introduce opinion leader dynamics, a framework that identifies two mechanisms through which token groups converge internally while maintaining distinct limiting directions. In the explicit model, fixed representatives induce a potential that attracts tokens toward distinct local maxima. In the implicit model, disconnected interaction groups evolve toward separate consensus directions. We formulate both models as reverse Wasserstein gradient flows and establish exponential convergence under suitable conditions. We further connect these theoretical predictions to token evolution in frontier sparse-attention LLMs that motivate our framework. Across four benchmarks, Kimi-K3, MiniMax-M3, and DeepSeek-V4-Flash consistently exhibit clearer cluster separation and higher clustering scores than the dense-attention model GLM-4.7-Flash in projected token representations. These observations support the relevance of the predicted multiple-group structure to trained frontier LLMs, while finite-particle simulations illustrate the theoretical convergence behavior. Together, our results connect restricted token interactions to distinct group-level attractors, providing a dynamical account of how sparse attention can support alignment within groups while preserving separation between them.
Displacement Geometry Captures Platonic Shared Reality Across Models and Modalities
位移几何捕捉跨模型与跨模态的柏拉图共享现实
Shang, Chenming, Tang, Yujin, Yang, Jun Jie Ou, Xu, Ruize, Breuer, Adam, Singh, Nikhil
Abstract
The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent work finds only weak pointwise similarity between models. In this paper, we show that what models share is not the location of samples in representation space, but the directions (displacement vectors) between them. Under a single orthogonal alignment--rotation and reflection only--these displacement vectors are substantially preserved across 44 independently trained vision and language encoders spanning modalities and asymmetric capability pairs, consistent with the PRH evidence. The samples' absolute positions are not, consistent with recent counter-evidence. Both arise from a single decomposition: representations split into a shared semantic component that is linearly aligned across models, and a private capability component that is not. We trace this geometry to concept-level structure: within a model, parent concepts are orthogonal to their child variation vectors; across models, concept displacements are parallel. Our theory falsifiably predicts (and experiments confirm) that fine-tuning preserves pointwise similarity but collapses displacement, and that relational distillation does the opposite. A major implication is that, because semantics align linearly but capabilities do not, capabilities can be imported from one model to another using a single cached forward pass through the source. We call this Shadow Casting. As a proof of concept, our SHADOWCLIP instantiation outperforms strong fine-tuned baselines at orders of magnitude less compute. A cache can be released alongside open model weights, letting one model's capabilities be downloaded and imported into any number of other models without fine-tuning.
Electroencephalography (EEG) provides non-invasive monitoring of brain activity and is widely used in emotion recognition, motor imagery and sleep staging. Although within-subject decoding has achieved considerable progress, cross-subject generalization remains a central challenge in practical applications. EEG decoders are typically trained with Adam/AdamW under a fixed second-moment decay coefficient, even though cross-subject learning involves low signal-to-noise ratios, subject variability, and gradient nonstationarity. A fixed coefficient implicitly assumes that gradient statistics are homogeneous across layers and time, which can limit model's adaptability to cross-subject EEG signals and degrade generalization. To address these issues, we propose AFOR, a tensor-wise adaptive optimizer that converts the fixed second-moment decay coefficient into a dynamic coefficient estimated online from local gradient state. AFOR combines a Residual-Alignment Signal Scorer (RASS) and an Adaptive Forgetting Controller (AFC). RASS summarizes local gradient residuals and directional agreement into a signal-quality score, and AFC maps this score through self-referential normalization to a bounded per-step decay coefficient, with cumulative-product initialization correction maintaining consistency under time-varying decay. Under a strict cross-subject protocol on three EEG benchmarks that cover three representative fields, AFOR achieves the best average performance among the compared optimizers, improving the mean test accuracy over Adam by 3.00%, 2.07%, and 4.38%, respectively.
Uncovering latent variables and their causal relations from observed data is a fundamental yet challenging problem. Existing methods often rely on restrictive assumptions, such as linear relations or invertible mixing functions. To better address this problem under general nonlinear mixing procedures, we propose a condition called the cross-Hessian Rank Constraint (HRC), which serves as a primitive rank-based tool for nonlinear latent causal discovery. In particular, we show that a rank-based property arises from the cross-Hessian of the observed-data log-density in the nonlinear case, revealing information about the latent variables, and reduces to the Tetrad constraints in the linear Gaussian case. More specifically, when two groups of observed variables are d-separated by a set of lower-dimensional latent variables, the rank of this cross-Hessian is equal to the dimension of the latent variables, under a mild affine derivative assumption on the conditional log-density derivatives. This assumption can be naturally satisfied when the noise level is low or the relevant nonlinearity is moderate. As a downstream application, we instantiate HRC in the pure one-factor measurement setting for locating latent variables and recovering their causal structure up to Markov equivalence. Experimental results on synthetic and real-world datasets support the theoretical claims.
Reinforcement Learning Inspired Black-box Adversarial Attacks for Computer Vision
受强化学习启发的计算机视觉黑盒对抗攻击方法
Krone, Florian, Hoemann, Elena, Hallerbach, Sven
Abstract
Neural networks, both convolution or transformer based, are essential for modern computer vision systems. However, they are vulnerable to small perturbations, almost imperceptible to humans, which significantly alter the model's prediction. These adversarial attacks are often considered to be a significant threat to the implementation of neural networks in safety-critical applications. Most attacks utilize the white-box threat model and therefore require full access to the target model, making them unrealistic to use in practice. We propose a novel approach under the more realistic black-box threat model that utilizes concepts from reinforcement learning to optimize perturbations with a non-differentiable target model. Reinforcement learning algorithms have already been optimized to be query efficient, making them an ideal starting point when designing black-box adversarial attacks. We show the success of our reinforcement learning inspired black-box adversarial attack (RIBA) in generating adversarial perturbations using only a small number of queries to the target model, by comparing it to state of the art attacks on different models on the Cifar10 and ImageNet data sets. RIBA takes $25.4\%$ fewer median queries to generate attacked images against a ResNet-18 on Cifar10 and $22.5\%$ fewer median queries to fool a Vit-B/16 model on ImageNet. Additionally, we demonstrate that RIBA can match the performance of white-box attacks on an adversarially trained model.
Explainable Predictive Condition-based Maintenance of Naval-Propulsion Systems using Fuzzy Logic
基于模糊逻辑的船舶推进系统可解释预测性状态维护
Kalogeropoulos, Dionisis, Sovatzidi, Georgia, Kalozoumis, Panagiotis G., Iakovidis, Dimitris K.
Abstract
The shipping industry has a significant impact on the global economy, emphasizing the need for operational availability and safety through the use of effective maintenance techniques. During the last decades, predictive maintenance (PdM) has emerged as a promising solution compared to the existing conventional maintenance systems. This is because it offers several advantageous functions, such as damage predictions for vessel components, reduced downtime, improved and extended life of machinery, as well as higher safety during voyages. However, existing methodologies developed for performing PdM do not provide explanations of their results to users, so that they can understand the failures that may occur. To address this limitation, this paper proposes a novel framework based on a fuzzy decision tree and a deep residual neural network, aiming to perform explainable PdM on naval vessels. The proposed framework is able to generate fuzzy local rules based on the dataset used, and can provide explanations of its outcomes, using cause-and-effect relationships, in a way that are understandable to users, thereby gaining their trust. Experiments using a publicly available dataset demonstrate the effectiveness of the proposed framework, as it achieves an accuracy of 99.24%.
The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group relative policy optimization and on-policy self-distillation, further reveal a clear directional skew: trained models improve in one direction while deteriorating in the other. We therefore propose MemCalib-RL, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation. Results across model families and scales (Qwen3-8B, Ministral-3-8B-Instruct, and Qwen3.5-35B-A3B) show that MemCalib-RL achieves the best overall performance while better balancing over-use and under-use, with gains generalizing beyond MemCalib in external benchmark evaluation. Further experiments support its design choices and robustness and provide insight into its training dynamics.
Change point detection (CPD) identifies abrupt and significant changes in sequential data, with applications in human activity recognition, financial markets, cybersecurity, manufacturing, and autonomous systems. Traditional CPD methods often face computational challenges in high-dimensional settings and typically provide limited explanations for detected changes, which can restrict their practical usability. This paper introduces a CPD framework that improves scalability and interpretability by leveraging the Sliced Wasserstein (SW) distance. Our contributions are fourfold: (1) we transform multivariate sequential data into one-dimensional scores using the SW distance, making the resulting representation compatible with existing CPD methods; (2) we analyze the distributional behavior of random slices of the SW distance and show that, under suitable assumptions, they can be approximated by a Gamma distribution, providing a principled basis for threshold calibration; (3) we propose a self-adapting online CPD algorithm that combines this SW-based score with an adaptive quantile-based threshold; (4) we introduce a model-specific framework for generating contrastive explanations for annotated change points. Empirically, our method reduces false positives by at least $48\%$ on average compared with popular online and offline CPD baselines, while maintaining competitive or superior detection performance. Code is available at https://github.com/jsve96/SWCPD_Code. At the same time, it produces interpretable change-point annotations, making it practical for deployment in high-stakes applications.
Chinese Translation
变点检测(Change Point Detection, CPD)旨在识别序列数据中突然且显著的变化,其应用涵盖人体活动识别、金融市场、网络安全、制造业以及自主系统。传统CPD方法在高维场景下常面临计算挑战,且对检测到的变化通常只能提供有限的解释,这限制了其实际可用性。本文提出一种利用切片Wasserstein(Sliced Wasserstein, SW)距离来提升可扩展性与可解释性的CPD框架。我们的贡献有四点:(1)利用SW距离将多元序列数据变换为一维分数,使其表示形式与现有CPD方法兼容;(2)分析了SW距离随机切片的分布行为,并证明在适当假设下其可由Gamma分布近似,从而为阈值校准提供了有原则的依据;(3)提出一种自适应在线CPD算法,将基于SW的分数与自适应分位数阈值相结合;(4)引入一个针对特定模型的框架,为标注出的变点生成对比性解释。实验表明,与流行的在线和离线CPD基线方法相比,我们的方法平均可将误报至少降低48%,同时保持相当或更优的检测性能。代码可在 https://github.com/jsve96/SWCPD_Code 获取。同时,该方法能够生成可解释的变点标注,使其在高风险应用的实际部署中具有实用性。
As Large Language Model (LLM) agents are applied in continuously interactive environments, driving the evolution of their own capabilities becomes a core problem for achieving long-term autonomy. Currently, environmental knowledge is typically treated as an external fixed input rather than as part of the agent's ongoing evolution. Reinforcement learning methods usually optimize policies through environmental interaction but tend to adapt only to fixed task distributions or single environments. This paper proposes TTSE (Two-Track Self-Evolution), a dual-track online self-evolution framework that separates evolving knowledge into FACT (environmental facts, whose reliability is continuously verified through interaction evidence) and TIP (task-conditioned implementation procedures). From a decision-theoretic perspective, we decompose the agent's excess risk into environment-representation regret and conditional-execution regret, characterize the conditions under which environment-conditioned policies strictly outperform condition-agnostic policies, and bound the downstream risk in terms of FACT identification error and cross-condition mismatch cost. In practice, TTSE's ablation experiments on GDPevo validate the advantage of dual-track evolution. On the classic agent task benchmarks ALFWorld and ScienceWorld, TTSE further demonstrates superior task adaptation. Moreover, TTSE is broadly compatible with existing skill self-evolution methods; combined with the Bayesian-Agent algorithm, a single-track ablation validates the dual-track advantage, substantially improving the aggregate score across the five major domains of SOPBench over three independent repetitions. Finally, on the real end-to-end task benchmark PinchBench, TTSE is integrated into a general agent framework via retrieval-based injection and stably outperforms the baseline across three independent runs.
KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
KV-COBRA:基于协同优化的比特-秩分配的KV缓存压缩
Ha, Sihyeon, Lee, Jaeho, Jeon, Yo-Seb
Abstract
What limits KV-cache compression at extreme bit-rates? We argue that it is not the choice of compression scheme, but how its budget is allocated across attention heads. Existing methods apply rank and bit-width uniformly, ignoring that each head has a different optimal mix of rank truncation and quantization. We show that co-optimizing rank and bit-width per head, using only standard low-rank projection and scalar quantization, dominates uniform allocation, with the largest gains at low bit-rates. Our method, KV-COBRA (Co-Optimized Bit-Rank Allocation), formalizes this as a resource-allocation problem: it balances rank-truncation loss against quantization loss within each head, then redistributes budget across heads to minimize total distortion. A fused Hadamard rotation equalizes per-channel variance, and reordering the SVD basis by attention-KL importance makes the solver query-aware. The same allocator extends to joint $K{+}V$ compression. On perplexity, zero-shot, and long-context benchmarks from $0.5$ to $4$ bits per dimension (bpd), KV-COBRA shows the smallest accuracy degradation among evaluated methods at low bpd, with no per-token overhead.
Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident than their corresponding pretrained language models (PLMs). Because task-specific labeled calibration data can be costly or unavailable, the corresponding pretrained PLM provides a natural label-free reference for post-hoc calibration. Prior agreement-gated PLM-referenced calibration fits a scalar temperature using only examples on which the PoLM and its PLM reference agree, excluding disagreement examples because direct alignment can drive the fitted temperature excessively high and induce under-confidence. We revisit this binary treatment. A controlled reintroduction diagnostic reveals a non-monotonic aggregate effect: admitting a moderate fraction of disagreement examples can improve calibration, whereas the benefit diminishes as unit-weight inclusion approaches the full disagreement set. We introduce SupportCal, a label-free post-hoc method that retains agreement examples at unit weight and assigns disagreement examples continuous weights based on the own-base PLM's relative support and corroboration from pretrained references selected from a size-compatible candidate pool. We further characterize when the resulting weighted objective admits a finite optimal temperature. Across MedMCQA and MathQA, SupportCal yields lower ECE than the agreement-only baseline for nearly all evaluated target-model configurations; supplementary TweetEval Sentiment results show the same pattern on a fixed-label classification task.
We show that a quantized model that keeps its classification accuracy still changes $14$ to $46\%$ of its top-1 retrieval results, and that aggregate ranking metrics reveal only part of this damage. We tie this failure to the gap between the two highest scores and use that gap to decide when a quantized answer can be trusted and where additional precision should be spent. We show that the top-1 result is guaranteed to survive quantization only when this gap exceeds twice the largest rounding error. In classification, scores are the logits, and the loss function pushes the correct class away from other classes, encouraging this gap. In retrieval, scores are query-document scores, and nothing separates the top-1 item from the second. This gap can be measured without labels. Before deployment, it predicts which models will break under quantization, and at deployment time it tells, per input, whether the quantized answer still matches the full-precision answer. Most classification inputs have a gap wide enough to trust the quantized answer, but few retrieval queries do. That gap motivates a different fix in each task. In retrieval, spending extra bit-width on the layers whose quantization moves the gap most recovers up to three-quarters of an extra bit's benefit for half its cost. In classification, routing the few low-gap inputs to full precision recovers most of the lost accuracy at a fraction of the cost.
A Distributional Optimisation Perspective on Combining Models in Deep Learning
深度学习中模型组合的分布式优化视角
Wang, Congye, Lin, Yan, Shen, Zheyang, Fisher, Matthew A., Oates, Chris. J.
Abstract
Combining predictions from different models can improve performance at machine learning tasks, but the training of the individual models and the rule used to combine them are typically chosen separately, and by ad hoc means. Recent advances in distributional optimisation (i.e. where the optimisation occurs over the set of probability distributions) offer an opportunity for principled joint training, viewing the collection of models as a discrete distribution whose support points are to be optimised, but the potential of these methods is not well-understood. In this paper we (1) cast two standard combination strategies - ensembles and low-rank adapter averaging - as entropy-regularised distributional optimisation, observing that the resulting objective is convex in the ensemble case but not in the adapter-averaging case, so that existing convergence guarantees for mean field Langevin dynamics transfer only to the former; (2) assess existing and novel algorithms for this task, including a functional variant of variational gradient descent; and (3) report an empirical study spanning synthetic classification tasks and fine-tuning of large language models on a commonsense reasoning benchmark.
Chinese Translation
组合不同模型的预测可以提升机器学习任务的性能,但各个模型的训练以及用于组合它们的规则通常是分开选择的,且方式较为随意。分布式优化(即在概率分布集合上进行优化)的最新进展为原则性的联合训练提供了机会,即将模型的集合视为一个离散分布,其支撑点有待优化,但这类方法的潜力尚未得到充分理解。在本文中,我们:(1) 将两种标准的组合策略——集成(ensembles)和低秩适配器平均(low-rank adapter averaging)——表述为熵正则化的分布式优化,并观察到所得目标函数在集成情形下是凸的,而在适配器平均情形下不是凸的,因此均值场朗之万动力学(mean field Langevin dynamics)现有的收敛性保证仅适用于前者;(2) 评估了用于该任务的现有算法及新算法,其中包括变分梯度下降的一个泛函变体;(3) 报告了一项涵盖合成分类任务以及在常识推理基准上微调大语言模型的实证研究。
An Intraoperative Hypotension (IOH) event is a frequent complication during administration of general anaesthesia with serious downstream consequences, yet clinical management remains reactive and not predictive. Existing predictive models, however, ignore drug infusion history as a valuable signal for prediction despite its direct pharmacological relevance. Our model achieves an Area Under the Receiver Operating Characteristic curve (AUROC) of 0.7360 and an Area Under the Precision-Recall Curve (AUPRC) of 0.1794, representing a 2.73-fold lift over the random guessing AUPRC baseline (0.0657), with the removal of propofol and remifentanil effect-site concentrations resulting in a 13.9% AUPRC drop compared to the full model. This is consistent with the hypothesis that pharmacokinetic trajectories encode impending haemodynamic changes before they manifest in the Mean Arterial Pressure (MAP). Additionally, this paper shows that training without lead-gap filtering degraded AUROC by 16.7%, empirically confirming that unfiltered models learn to detect ongoing hypotension rather than predict future events. Finally, a Mamba-based architecture achieves the aforementioned high prediction performance while maintaining a constant memory footprint across a range of sequence lengths, unlike the quadratic VRAM overhead typical of vanilla Transformers, making it the more practical choice for continuous intraoperative deployment.
Explainable Neuro-Fuzzy Prediction for Trustworthy Decision-Making in Maritime
面向海事领域可信决策的可解释神经模糊预测
Kalogeropoulos, Dionisis, Sovatzidi, Georgia, Iakovidis, Dimitris K.
Abstract
Predicting when maritime systems require maintenance can be critical, avoiding hazards and costly consequences. To address this problem, this paper proposes an explainable decision-making framework that integrates a neuro-fuzzy prediction model with a two-stage explainable component. The first stage of this component produces feature-attribution explanations, using gradient-based saliency maps, and the second stage extracts local rules using a fuzzy decision tree. The proposed framework is generic and can be integrated into any deep learning-based approach, rendering it explainable. To the best of our knowledge, this is the first fuzzy logic-based framework enabling both feature-level and local rule-based explanations of black box models. This approach aims to foster trustworthiness in decision making through user-understandable machine inferences. The performance of the proposed framework using a deep residual-based neural backbone is evaluated on various general-purpose public benchmark datasets, and its utility in maritime is demonstrated in the context of early fault detection in a naval propulsion system dataset. The results indicate that it can provide predictions outperforming relevant state-of-the-art approaches, with an average AUC-ROC (Area Under the Receiver Operating Characteristic Curve) value, reaching up to 99%, while offering the advantage of explainability.
Prescriptive SVD-Inspired Attention via Spectral Energy Retention
基于谱能量保持的规范性SVD启发注意力机制
Arampatzakis, Vasileios, Sevetlidis, Vasileios, Pavlidis, George
Abstract
Self-attention is central to modern Transformer architectures, but its dense dot-product formulation makes it difficult to identify which internal directions are structurally important and which can be modified without disrupting the model. SVD-Inspired Attention (SVDA) addresses part of this problem by introducing a learned diagonal spectrum into the query-key score interaction, making latent attention directions explicitly inspectable through indicators such as spectral entropy, effective rank, sparsity, alignment, selectivity, and perturbation response. This paper examines the transition from diagnostic interpretation to operational intervention. A diagnosis--intervention--verification framework is proposed, and one intervention is evaluated: spectral energy retention in the attention-score pathway. Across FashionMNIST, CIFAR-10, CIFAR-100, and Food-101, the $\rho=0.90$ prescription removes 24.5--53.7\% of score directions, reduces parameters by 2.6--4.3\%, and reduces estimated MACs by 2.8--5.4\%. The paired mean accuracy change of the dimension-reduced model ranges from $-0.03$ to $+0.05$ percentage points over three seeds. These results support SVDA as an intrinsically interpretable attention mechanism whose learned spectrum exposes an operational coordinate system for deterministic and verifiable modification of attention-score formation.
RLVR has substantially improved the reasoning capabilities of LLMs. However, existing methods typically parameterize temporal progression in the Markov Decision Process by token-by-token generation, despite the highly non-uniform information flow along autoregressive trajectories. In this paper, we propose InfoPPO, which reparameterizes temporal progression using information density rather than raw token count. This reparameterization induces a common state-dependent structure for both temporal credit propagation and policy updates. InfoPPO restores the effectiveness of non-trivial discounting in long-horizon reasoning, retaining effective-horizon contraction while avoiding excessive attenuation of terminal supervision over long token sequences. Moreover, the information-time policy-improvement analysis naturally leads to a state-dependent update constraint, which we implement through adaptive clipping. By adapting the clipping threshold at each token position to the information density of its corresponding state, this mechanism enables more targeted policy updates while preserving proximal control. Theoretically, we extend performance-difference and policy-improvement analyses to the information-time MDP, deriving a policy-improvement lower bound when policy changes are regulated by information density. We further connect the general information-time analysis to practical LLM policy optimization by relating state-wise information density to local policy movement, while also providing theoretical grounding for the adaptive update mechanism. Experiments on Qwen3 models demonstrate consistent gains over competitive baselines across five challenging competition-style mathematical reasoning benchmarks. InfoPPO also maintains stable accuracy and response length across non-trivial discount settings under which token-time PPO deteriorates.
Credit Access is Associated with Improved Food Security in the Horn of Africa
信贷获取与非洲之角地区粮食安全状况的改善相关
Cerdà-Bautista, Jordi, Sitokonstantinou, Vasileios, López-Peña, José Manuel Veiga, Piovani, Duccio, Tárraga, José María, Camps-Valls, Gustau
Abstract
The intensification of climate change poses a growing threat to food security, especially in vulnerable communities. This study employs an observational machine-learning framework to estimate the causal association between access to credit and acute food insecurity in Somalia and across the Horn of Africa, drawing on a harmonized dataset spanning key environmental, socioeconomic, and conflict-related factors from 2015 to 2022. Results indicate that greater credit access is associated with a 2% reduction in acute food insecurity at the population level over the study period. Given that, on average, 16% of the population is in crisis, this effect represents a meaningful shift within the at-risk group. We interpret these estimates under explicit identification assumptions and complement them with robustness and refutation tests. The results provide context-specific evidence on how financial access correlates with food security outcomes in data-scarce, crisis-affected settings, and offer a transparent framework for integrating heterogeneous data sources when randomized evaluations are infeasible.
Machine Learning-Based Prediction of Childhood Stunting in Bangladesh: Fairness and Temporal Robustness Assessment
基于机器学习的孟加拉国儿童发育迟缓预测:公平性与时间稳健性评估
Haque, Md Ahshanul, Kabir, Muhammad Ashad
Abstract
Childhood stunting remains a major public health concern in Bangladesh and reflects long-term growth failure influenced by child, maternal, household, socioeconomic, and health-service factors. This study used nationally representative Bangladesh Demographic and Health Survey data from 2007 to 2022 to develop machine learning models for population-level prediction of childhood stunting and to assess temporal robustness and subgroup fairness. Children aged 0-59 months with complete anthropometric and predictor data were included. Data from the 2007, 2011, and 2014 survey rounds were used for model development, while the 2018 and 2022 rounds were retained as temporal test datasets. Twelve feature-selection approaches were assessed, and the KNN permutation importance-selected predictor set was used for final model evaluation. Eleven machine learning models were evaluated: ten conventional algorithms and one pretrained tabular foundation model, TabPFN. Performance was assessed using balanced accuracy, AUROC, F1-score, Brier score, and expected calibration error. Subgroup fairness was examined by child sex, place of residence, and socioeconomic status. The final analytic sample included 18,844 children, of whom 35.05% were stunted. In the development hold-out test dataset, TabPFN showed the highest observed balanced accuracy overall at 67.58%, while AdaBoost showed the highest observed balanced accuracy among conventional models at 67.51%. In temporal testing, the highest observed balanced accuracy was found for Gradient Boosting in BDHS 2018 and XGBoost in BDHS 2022. Model performance varied across survey rounds and subgroups, highlighting the importance of temporal validation, subgroup fairness assessment, and transparent interpretation in public health prediction modeling.
Voice-controlled interaction in industrial settings is hampered by acoustic noise, which severely degrades audio-only speech recognition. Audio-visual speech recognition (AVSR) addresses this by fusing lip-motion cues with the audio stream, but state-of-the-art pipelines rely on three-dimensional convolutions, recurrent units, and attention modules that exceed the budget of typical edge devices. We present NAVIR, an end-to-end AVSR system targeting the BrainChip Akida neuromorphic processor, which natively supports only sequential two-dimensional convolutional inference. The pipeline factorises spatial and temporal encoding into separate AkidaNet-based modules: a per-frame visual encoder, a temporal video encoder, and a spectrogram audio encoder, fused by a lightweight predictor head and decoded by constrained beam search. Models are trained with connectionist temporal classification on noise-augmented audio and then fine-tuned with quantization-aware training. On the GRID benchmark, the quantized audio-visual model reaches 14.0% word error rate (WER) under noise on the unseen-speaker split and 3.3% WER on the overlapped-speaker split, against 22.5% and 11.8% for audio-only baselines, and it attains 98.6% command accuracy at 1.5% WER on a task-specific industrial-command corpus. Operation-count analysis indicates a 13-fold energy advantage of the spiking formulation over its artificial neural network counterpart at 27.6% mean firing rate. On-board measurements show roughly 5-fold lower energy per inference than a Raspberry Pi central processing unit on the lip-reading model, and over 100-fold lower than a laptop graphics processing unit, while sustaining 14.5 inferences per second. To the best of our knowledge, this is the first complete multimodal AVSR pipeline running on neuromorphic hardware of this class.
Climate variability influences whether a market disruption escalates into a food crisis, yet broad climate patterns like El Ni\~no, tracked months before they alter hydro-climatic conditions, are still not incorporated as an early-warning component in food-security responses. We address this gap by introducing sensitivity regimes, a stratification of regions by the direction and strength of their vegetation response to the El Ni\~no Southern Oscillation, and using them to estimate how food price spikes affect acute food insecurity across sub-Saharan Africa. Integrating remote sensing, socioeconomic data, and causal machine learning, we find that in regions where ENSO systematically suppresses vegetation, a price spike raises the share of the population at acute risk by 5.4 percentage points in the following month. In regions where vegetation is unaffected by or positively linked to ENSO, the estimated effect is smaller (around 2 percentage points) and statistically insignificant. These results demonstrate that climate context is critical for understanding food security vulnerabilities. Sensitivity regimes can be combined with operational price-spike triggers to stage anticipatory action: the ENSO state flags vulnerable regions months ahead, and a pre-positioned response in those regions to a price spike would avert the largest jump in acute food insecurity.
Probabilistic Modelling of Operational Design Domains, A New Approach for Testing AI Systems
运行设计域的概率建模:一种测试人工智能系统的新方法
Wiesbrock, Hans-Werner
Abstract
The conventional testing process quickly fails when applied to ML-based systems such as obstacle detection in vehicles: if an obstacle is not detected in a test, classical bug fixing is impossible and an AI system will always retain shortcomings. Test results can therefore only be interpreted statistically, which in turn requires test sets that are not only complete with respect to the operational design domain (ODD) of the system, but also representative of it. To this end, we introduce probabilistically extended ontologies (PEONs): ontologies describing the ODD, augmented with a probability distribution over the partitioning they induce. Instead of unmaintainable conditional probability tables, only marginal distributions and functionally described dependencies need to be specified; algorithms based on couplings and optimal transport complete this specification to a Bayesian network. From a PEON we derive the sampling of representative test cases, rigorous end-of-test criteria for given quality targets and significance levels, and methods for re-evaluating existing test results and for assessing the balance of training data. We demonstrate the practical modelling of a complex ODD using the example of automatic train operation.
Artificial Structure Function Search: Preserving Artificial Functional Connectivity for Structured Pruning
人工结构-功能搜索:面向结构化剪枝的人工功能连接保持
Illeperuma, Mindula, Pina, Rafael, Herath, Charuka, Gabayre, Sharmarke A., De Silva, Varuna
Abstract
Structured pruning is a model compression technique that is used to reduce the computational cost of deploying deep neural networks on resource-constrained devices. Popular methods of pruning rely on opaque heuristics or weight-based criteria that give no indication as to the structural dependencies in the network. To address these limitations we present Artificial Structure Function Search (ASF-S): a novel structured pruning framework. ASF-S utilizes Principle Gradient Importance (PGI): a novel prune-candidate selection criteria that is inspired by structure-function relationships in the brain. By ensuring the pruned structure of the model respects topographical organization of the output layer, we define Artificial Functional Connectivity (AFC) for artificial neural networks. AFC provides evidence to demonstrate that accurate smaller networks can be found using careful prune candidate selection criteria. We present results for PGI as a selection criterion and for ASF-S as a pruning framework against recent benchmarks, demonstrating that our method yields model variants with 70\% parameter reduction, that can recover baseline accuracy without re-training the pruned layers.
Despite advances in long-context inference, large language models (LLMs) remain fundamentally limited by the key-value (KV) caching mechanisms that are necessary for stable computation. Techniques such as selective token eviction and pruning have vastly mitigated these issues, but often discard core information to manage the growing cache. In this paper, we propose Attention with Routed Memory (ARM) a novel KV caching structure that introduces a fully differentiable, fixed-size memory system organized as a hierarchical router. Via a Gumbel-Softmax, ARM learns to select memory slots and perform sigmoid-gated updates that softly combine new and stored information, avoiding hard eviction and reducing information loss. By further training a policy to dynamically select varying amounts of memory at inference, ARM adapts its accesses for both simple contexts and inputs that require deeper reasoning, enabling more scalable and effective retrieval on both short- and long-contexts. Experimental results on standard commonsense and long-context reasoning benchmarks demonstrate that ARM achieves superior performance and efficiency compared to fixed KV-caching approaches, while remaining efficient and scalable in terms of both memory and generation latency.
Prior-Amortized In-Context Bayesian Inference for Generalized Linear Mixed-Effects Models
面向广义线性混合效应模型的先验摊销式上下文贝叶斯推断
Kipnis, Alex, Binz, Marcel, Schulz, Eric
Abstract
Hierarchical data is ubiquitous in the empirical sciences and is most commonly analyzed with generalized linear mixed-effects models (GLMMs). Bayesian inference for GLMMs yields calibrated uncertainty but requires MCMC; the No-U-Turn Sampler (NUTS) is the gold standard but is slow and must restart from scratch for every new dataset, model and prior. We introduce metabeta, a pretrained neural network for prior-amortized in-context Bayesian inference over GLMMs. Unlike previous neural posterior estimators that fix the prior at training time, metabeta accepts prior families and hyperparameters as inputs at test time, enabling zero-shot generalization. Two set transformers and conditional normalizing flows mirror the posterior's two-level structure (global parameters shared across groups, local parameters per group). The model is trained on millions of realistic simulated datasets spanning continuous, binary, and count outcomes. By default, the flow posterior is refined by Independence Metropolis-Hastings against the unnormalized posterior, so its correctness rests on the sampler rather than the network; this yields tuning-free inference two to three orders of magnitude faster than NUTS. Alternatively, the flow can warm-start NUTS, giving nearly identical inference with substantially increased speed and stability. On controlled benchmarks with ground-truth parameters, metabeta matches NUTS in parameter recovery, calibration and out-of-sample prediction. On out-of-distribution real datasets, its posteriors closely match those of NUTS across all parameter types, and they remain faithful under misspecified likelihoods and priors, out-of-distribution predictors, collinear designs, and data-poor regimes. The model is open-source and open-weights and thus immediately deployable.
Sparse on-policy distillation (OPD) allocates teacher supervision to a small subset of tokens in student-generated trajectories. However, useful teacher guidance can yield a noisy update when its gradient is estimated from a sampled next token. We study this estimation problem at a fixed prefix in information geometry and propose an information-efficiency ratio (IER) based on a signal-to-noise decomposition. IER characterizes relative gradient estimation error under an optimal scalar baseline. A candidate-set approximation enables token selection based on IER and its combination with existing usefulness scores, while retaining the sampled reverse-KL training objective. On mathematical and medical reasoning tasks, adding IER improves existing selectors in multiple settings, with sparse configurations matching or exceeding full OPD without token selection at small token budgets of 0.1\%--1\%. These results support accounting for both usefulness and gradient-estimation reliability when allocating sparse supervision. Our code is available at https://github.com/BruceSheng1202/IER-OPD.
Comparing Latent Concept Formation in State Space Models and Transformers via Sparse Autoencoders
基于稀疏自编码器比较状态空间模型与Transformer中的潜在概念形成
Nagaraj, Rithin, Oruganti, Rupa Laalasa, Kunder, Prerna Subhashchandra, Joshi, Ashwini M
Abstract
The quadratic scaling of Transformer self-attention has driven the adoption of sub-quadratic Selective State Space Models (SSMs) like Mamba, which compress past context into a fixed-size recurrent hidden state. This strict informational bottleneck raises a foundational question for mechanistic interpretability: do SSMs and Transformers learn fundamentally distinct latent representations? In this work, we employ Sparse Autoencoders (SAEs) to conduct a large-scale, feature-level correspondence analysis between Mamba-130m and Pythia-70m over a 10-million token corpus. Contrary to hypotheses predicting widespread architectural divergence, we find no evidence of systematic representational divergence between architectures: across the observed Jaccard distribution, 99.98% of Mamba features cluster toward the upper alignment boundary, providing preliminary feature-level support for the Universality Hypothesis. We further identify and qualitatively characterize this microscopic fraction (0.02%) of diverging features, finding patterns consistent with the hypothesis that the recurrent bottleneck selectively limits the parsing of rigid syntax rather than broad semantic ontology. We demonstrate that while Pythia's unconstrained attention permits the monosemantic decomposition of distinct formatting edge-cases, Mamba is forced to compress unrelated syntactical anomalies into polysemantic "junk drawer" neurons to preserve state capacity. Collectively, these results suggest that architectural routing mechanisms may have negligible impact on core semantic understanding, with representational divergence confined to extreme structural margins.
Multivariate time-series forecasting is essential to many real-world applications. Recent large vision models (LVMs) offer a promising paradigm by transferring cross-domain visual priors to time-series forecasting. However, existing LVM-based methods face two key challenges: balancing independent visual representation spaces with cross-variable dependency modeling, and adapting vision backbones pretrained on natural images to the distinct temporal semantics of time-series images. To address these challenges, we propose MUSE, a dependency-aware adaptation framework built on a fully frozen pretrained MAE. First, the Variable Context Refinement Module (VCR) aggregates shared temporal information within each variable and models cross-variable contextual dependencies while preserving independent visual spaces. Second, the Temporal-Periodic Refinement Module (TPR) performs lightweight refinement at different encoder depths and explicitly models across-period temporal dependencies and within-period periodic dependencies. The two modules independently produce forecasts, which are fused through a learnable prediction-level gate. Experiments on 10 real-world datasets demonstrate that MUSE achieves state-of-the-art performance.
Accurate, reliable, and deployable wind power forecasting is critical for power system dispatch, renewable energy integration, and electricity market operations. Progress in this field hinges on the ability to empirically and comprehensively benchmark forecasting methods. Yet existing benchmarks fall short of supporting systematic evaluation in four key aspects: 1) limited coverage of wind power scenarios across turbine scale, variable composition, and spatial structure; 2) incomplete coverage of forecasting model families; 3) evaluation metrics misaligned with wind power requirements; and 4) limited structure-aware diagnostics beyond individual temporal patterns. To address these limitations, we propose WPBench, a comprehensive, fair, and extensible benchmark for wind power forecasting. WPBench integrates 26 public datasets organized by turbine scale and variable composition, spanning single-turbine, multi-turbine, univariate, and multivariate settings. Under unified processing, training, and evaluation protocols, it benchmarks 19 representative models covering traditional methods, deep temporal models, spatio-temporal models, and foundation models. Beyond point-wise errors, WPBench assesses forecast-curve fidelity and computational efficiency, and delivers structure-aware diagnostics across temporal, variable-dependency, and spatial-dependency perspectives. Together, these capabilities enable systematic model comparison across diverse wind scenarios and provide a reusable platform for future research.
Using a Large Language Model (LLM) as the clusterer at production scale is hard: prompts cannot hold the entire label space, and per-document serial processing does not deliver the throughput real workloads require. We present RAILS, a retrieval-augmented incremental LLM clusterer that turns clustering into a simple loop over a growing label pool and scales through document batching with bounded concurrency. On six public benchmarks RAILS exceeds the strongest prior LLM-clustering method on average, lifting accuracy from 51.2% to 59.3%, NMI from 67.2% to 74.8%, and ARI from 45.4% to 54.7%. We further report production-deployment evidence from a SaaS ticket-topic-discovery pipeline, where RAILS has replaced a traditional HDBSCAN stage with higher clustering quality, transparent prompt-driven control, and stateful incremental operation.
Music festival lineups emerge from complex relationships among artists, genres, releases, labels, and past performances, making the prediction of future lineups a natural fit for temporal knowledge graph (TKG) forecasting. In this work, we present a TKG covering 380 festivals over 55 years, comprising more than 90K festival performance quadruples along with information on festivals, artist tours, and artist metadata, and release it as a resource for TKG forecasting evaluation. We formalize festival lineup forecasting as temporal link prediction between artists and festivals at future timestamps. We evaluate six TKG forecasting models on this task, analyze their capabilities and limitations, and compare them against Large Language Models applied zero-shot. Our resource complements existing TKG benchmarks by grounding evaluation in a concrete, real-world application domain.
Lifted Bellman Linear Programming for Offline Reinforcement Learning
面向离线强化学习的提升贝尔曼线性规划
Yang, Hyukjun, Park, Jongchan, Jeong, Narim, Lee, Donghwan
Abstract
Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target networks with exponential moving average (EMA) updates. Multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead impose in-sample Bellman optimality on the critic through inequality constraints. We formulate the Lifted Bellman Linear Program (LBLP), which lifts the linear programming characterization of Bellman optimality to the joint $(Q,V)$ space so that every constraint involves only state-action pairs in the dataset. Its unique minimizer is the in-sample optimal pair, and constraints along $K$-step segments of dataset trajectories leave this minimizer unchanged for any rollout policy and horizon. Under deterministic dynamics, this minimizer lies between the best dataset return and the optimal value. Relaxing the constraints into hinge penalties recovers the same solution above a finite penalty coefficient in the tabular case. Approximate Lifted Bellman Unconstrained Minimization (ALBUM) implements this relaxation with neural networks and detaches the $K$-step rollout targets by stop gradient. Its objective contains no squared regression onto bootstrapped targets, so it can be trained without target networks or EMA updates. Under deterministic dynamics, the LBLP solution is a stationary point of the detached update under a coefficient condition independent of $\gamma$ and $K$, and the inequality constraints allow discounted returns along dataset trajectories to serve as lower bounds without off-policy correction or action chunking. On OGBench, ALBUM uses a single critic with a Gaussian policy, matches the average performance of FQL, and is comparable to recent action-chunking methods, while using the fewest parameters and the least peak GPU memory among all compared methods.
Fine-tuned checkpoints and adapters now fill public repositories, and the most common operation applied to these artifacts is model merging: arithmetic on their weights that assembles capabilities cheaply. We ask what this operation does to emergent capabilities: behaviors an artifact carries that were never an explicit training target. Studying two independent testbeds (activation oracles and emergent-misaligned models) across three model families, we find that the answer is threefold. First, merging preserves an emergent capability that both parents carry: merging two misaligned checkpoints retains most of their broad misalignment across the whole mixing range. Second, merging cannot create an emergent capability that is superadditive in its parents: no weighted merge of two single-task oracles reaches the jointly-trained oracle's auditing ability. Third, when only one parent carries the capability, merging dilutes it faster than the trained capability that accompanies it: the gap is significant in most settings. In short, emergent behaviors of an artifact do not compose the way its trained capability does.
We present $t_0$, a family of open-weights foundation models for forecasting with multivariate context. We release its first two members: $\texttt{t0-alpha}$ and $\texttt{t0-beta}$, respectively 102M and 256M parameters. Both condition their forecasts on target history, past covariates, and known-future covariates, without task-specific retraining. Their transformer layers alternate attention along time and across variates. They produce probabilistic forecasts through quantile predictions. Pretraining combines curated public data with synthetic generator families constructed to contain covariate-to-target dependencies. On GIFT-Eval, $\texttt{t0-alpha}$ reaches an aggregate CRPS of 0.4941, and $\texttt{t0-beta}$ a CRPS of 0.4738 and a MASE of 0.6865, third on both and within 4.0% of the best zero-shot TSFM. On fev-bench they score 42.2 and 46.7 in skill, the latter third again and 2.0 points behind the leader. We analyze $\texttt{t0-alpha}$ in depth. Known-future covariates raise its skill by 6.3 percentage points across 30 tasks. The report also examines its calibration, its rollout strategy on long horizons, and its robustness to missing data. On the Victoria electricity-demand benchmark, $\texttt{t0-beta}$ is among the most accurate models with a context of nearly a year. In an independent Macrocosm evaluation of hourly ERCOT prices over 29 months, both cut the MAE of the lagged-price baseline by 38%.
Universal Multi-Modal Traceformer: Integrating Heterogeneous Context for Process Event Prediction
通用多模态Traceformer:融合异构上下文的过程事件预测
Spaeh, Fabian, Fang, Jingxing, Zhe, Shandian, Shen, Bin
Abstract
Event logs arise in a wide range of real-world processes, capturing not only event activities and timestamps but also multi-modal contextual information. Existing event-sequence models, including many temporal point process approaches, primarily model event activities and timestamps while overlooking heterogeneous context, such as numerical measurements, categorical attributes, textual descriptions, and metadata associated with individual events and entire traces. In this paper, we propose Universal Multi-Modal Traceformer (UMT), a unified framework for incorporating heterogeneous process context into next-event prediction. Built on a Transformer backbone, UMT introduces a universal feature encoder that maps diverse feature types into a shared representation space and handles contextual information at both the event and trace levels. UMT further develops a per-event Perceiver module that dynamically weights contextual features and adaptively integrates them into event-token representations. To accommodate the heavy-tailed and potentially multi-modal distribution of inter-arrival times, UMT represents each interval at multiple temporal scales and jointly predicts the corresponding scale-specific quantities. Experiments on 13 real-world event logs show that UMT improves both next-event activity and time prediction over existing approaches.
Chinese Translation
事件日志产生于广泛的现实世界过程中,不仅记录事件活动和时间戳,还包含多模态的上下文信息。现有的事件序列模型,包括许多时间点过程(temporal point process)方法,主要对事件活动和时间戳进行建模,而忽视了异构上下文,例如数值测量、分类属性、文本描述以及与单个事件和整个轨迹(trace)相关的元数据。在本文中,我们提出了通用多模态Traceformer(Universal Multi-Modal Traceformer, UMT),这是一个将异构过程上下文融入下一事件预测的统一框架。UMT基于Transformer骨干网络构建,引入了一个通用特征编码器,将多种特征类型映射到共享的表示空间中,并在事件级和轨迹级两个层面处理上下文信息。UMT进一步开发了一个逐事件的Perceiver模块,该模块动态地对上下文特征进行加权,并将其自适应地整合到事件token表示中。为了适应事件间隔时间的重尾且可能呈多模态的分布,UMT在多个时间尺度上表示每个时间间隔,并联合预测相应的特定尺度量。在13个真实世界事件日志上的实验表明,UMT在下一事件活动预测和时间预测方面均优于现有方法。
Ngo, Long, Chamli, Mohammed Amine, Rivalan, Jonathan, Jaillon, Thomas
Abstract
Traditional evaluation metrics provides numerical values but often lack comprehensibility, hindering effective differentiation of model performances. Our work addresses this challenge by introducing overlay\_dx, a novel evaluation metric measuring the performance of time series prediction models. Overlay\_dx is a visual metric that represents the percentage of predictions falling within a confidence interval around actual values. Additionally, once evaluation results are plotted, overlay\_dx computes the area under the overlay curve, providing a quantitative measure of alignment between predicted and actual values across different thresholds and predictions. Through extensive experiments, we demonstrate that our approach offers a unified evaluation framework that combines both visual and numerical assessments, enabling improved model comparison and providing valuable insights for further research and optimization efforts in time series prediction.
Sea ice forecasts are issued several days ahead, allowing errors to accumulate while new, often sparse sea ice concentration (SIC) observations become available. We find that fixed-propagation errors concentrate near structured, high-gradient ice edges, whereas homogeneous interiors require limited propagation, suggesting that propagation distance should be state dependent. We therefore introduce ECHO (Evidence-guided Correction with Heterogeneous prOpagation), where ECHO-Scale adapts propagation distance while preserving correction geometry, and ECHO-Delta learns a bounded residual around fixed propagation. Across all 96 standard evaluation settings spanning diverse priors, observation times, sparsity levels, geometries, and noise conditions, both outperform fixed propagation. ECHO-Delta achieves the best average accuracy, while ECHO-Scale is more robust to geometry shifts. Code is available at https://github.com/yingtian22/TAKING-A-SECOND-LOOK.
Chinese Translation
海冰预报会提前数天发布,在此期间误差会不断累积,而新的、通常较为稀疏的海冰密集度(SIC)观测数据会逐渐可用。我们发现,固定传播的误差主要集中在具有结构性的高梯度冰缘附近,而均匀的冰区内部所需的传播范围有限,这表明传播距离应当依赖于状态。因此,我们提出了ECHO(基于证据引导的异构传播校正,Evidence-guided Correction with Heterogeneous prOpagation):其中ECHO-Scale在保持校正几何结构的同时自适应调整传播距离,而ECHO-Delta则在固定传播的基础上学习一个有界残差。在涵盖不同先验、观测时间、稀疏程度、几何形态和噪声条件共96个标准评估场景中,两种方法均优于固定传播方法。ECHO-Delta取得了最优的平均精度,而ECHO-Scale对几何形态变化更为稳健。代码可在 https://github.com/yingtian22/TAKING-A-SECOND-LOOK 获取。
Electricity forecasting often involves spatially related signals observed over regions, substations, and feeders, and Graph Neural Networks (GNNs) provide a natural way to represent these relations. Building a complete GNN forecasting experiment is nonetheless laborious, because graph construction, model selection, training, aggregation, and interpretation sit in incompatible tools. We present GraphToolbox, an open-source Python framework that unifies these stages in one configurationdriven pipeline built on PyTorch Geometric. It offers data-driven graph construction, an adapter that instantiates and trains 51 of the 65 PyTorch Geometric convolutions together with the recurrent cells of PyTorch Geometric Temporal, online expert aggregation, forecasting interpretability, and significance testing on cached forecasts. We evaluate the pipeline in two case studies. On French regional load, the 48 convolutions included in the complete forecasting sweep fall in a band from 1.14% to 1.60% error, online aggregation lowers this to 0.98%, and the graph models improve on classical additive and boosting baselines. On net-load, direct graph models are less accurate than a classical additive model, while forecasting each physical component separately improves them without closing that gap. Both comparisons use the same experimental interface, illustrating the role of GraphToolbox in systematic architectural evaluation.
Augmented Hypothesis Testing with Persona-Based LLM Simulations
基于角色化大语言模型仿真的增强假设检验
Benomar, Ziyad, Marjani, Aymen Al, Missault, Paul, Mansour, Saab
Abstract
A/B testing requires large sample sizes, long timelines, and significant costs. When auxiliary predictions of experimental outcomes are available from machine learning models, uncertain prediction quality precludes replacing human experiments entirely, yet these predictions may still contain useful signal. We propose a principled framework for learning-augmented hypothesis testing that leverages predictions of unknown quality to reduce sample sizes while maintaining statistical validity. Predictions naturally vary in granularity, from coarse aggregate signals to fine-grained individual-level estimates, and our framework addresses both ends of this spectrum: (1) for population-level directional predictions, where only a binary signal on the treatment effect sign is available, we use an asymmetric test and prove consistency and robustness bounds within the learning-augmented algorithms paradigm; (2) for individual-level predictions, we introduce Generalized PPI++ (GPPI), extending Prediction-Powered Inference to handle nonlinear prediction errors through higher-dimensional transformations. Both methods benefit from accurate predictions while remaining robust to inaccurate or adversarial ones. We validate our framework using persona-based LLM simulations, where AI agents equipped with user personas predict individual behavior, as a natural prediction source spanning both granularity levels. Experiments on four real-world datasets demonstrate that our methods, combined with persona-based predictions, substantially reduce experimental costs while preserving rigorous statistical validity.
iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs
iSDFT:面向大语言模型持续学习的信息近端自蒸馏方法
Khamis, Ahmed Khaled, Ji, Xiaotong, Jaber, Hassan, Tutunov, Rasul, Zimmer, Matthieu, Wang, Jun, Bou-Ammar, Haitham
Abstract
On-policy self-distillation fine-tuning (SDFT) learns new skills from demonstrations while reducing forgetting, but it always distils toward the full demonstration-conditioned teacher. This fixes teacher influence at the full-teacher endpoint, providing no control over how much demonstration information should be transferred at each prediction state. We introduce Information-Proximal SDFT (iSDFT), which instead treats the teacher as a budgeted source of information. At each token, iSDFT selects the distribution closest to the current student that satisfies a prescribed teacher-information constraint, yielding a closed-form exponential target with a locally determined tilt. To control cumulative drift, we further anchor the student to its frozen base policy. Across four heterogeneous LLM backbones and two specialisation tasks, iSDFT improves vanilla SDFT in 7 of 8 model-task settings and matches it in the remaining one. It also provides tighter retention on the original SDFT benchmark suite, with 73% of evaluations remaining within 0.5 points of the base model versus 52% for the strongest baseline, while achieving the largest mean improvement on all ten additional mathematics, coding, and competition-mathematics benchmarks. These results show that controlling how much and when teacher information is introduced improves specialisation while preserving broader capability.
Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement
校正强制:面向生成式语音增强中扩散模型与流模型的统一后训练方法
Yao, Qing, Gao, Lijian, Mao, Qirong
Abstract
Diffusion and flow models, as promising generative paradigms for speech enhancement, face a training--inference mismatch: training uses analytical path states, whereas inference recursively evaluates models on self-generated rollout states along discretized sampling trajectories. This mismatch causes prediction and discretization errors to accumulate. To address it, we introduce Corrective Forcing (CoF), a post-training paradigm that forces diffusion and flow models to learn from self-generated rollouts and correct their predictions. CoF corrects clean-speech predictions on rollout states toward the ground truth under dynamic sampling schedules, exposing the model to varying inference conditions. It further regularizes local evolution using locally corrected counterfactual transitions as references for factual transitions. By expressing model outputs through a shared clean-speech prediction parameterization, CoF applies the same post-training objective across diffusion and flow formulations. Experiments with SB-VE and OT-CFM demonstrate improvements in perceptual quality and reconstruction fidelity, together with robust performance across different numbers of sampling steps.
Muon Can Outperform Dedicated Continual Learning Methods
Muon可以超越专门的持续学习方法
Sincari, Sebastian George, Gheorghe, Bogdan Alexandru, Barbalau, Antonio
Abstract
Continual learning with Low-Rank Adapters (LoRA) typically mitigates forgetting by penalizing the overlap between a new update and the accumulated past weights, which discourages certain update directions without controlling how an update distributes its energy over the ones that remain. We ask whether that restriction has to be task-aware, or whether a generic one supplied by the optimizer is enough. We train a plain incremental LoRA (IncLoRA) with Muon, which orthogonalizes each update, and compare it against O-LoRA and ELLA over five seeds and three task orders on the Standard CL Benchmark and three seeds on TRACE. IncLoRA+Muon reaches the accuracy band of the dedicated methods on Standard CL and improves on every AdamW configuration on TRACE. One update-constraining mechanism is enough, whether it comes from the loss or from the optimizer; on Standard CL a second one does not help, and for the most restrictive method it costs 8.4 points of accuracy and the plasticity to fit each task. What separates the two optimizers is not the size of the update, which under Muon is 0.91 to 2.06 times that under AdamW, but how it is distributed. AdamW confines it to between 1.4 and 1.8 effective singular directions, Muon spreads it over 7.0, and the two do not overlap in any tracked run. Part of the advantage usually attributed to dedicated CL methods may therefore be explained by the geometry of the optimizer's updates.
Guaranteed Low-Rank Tensor Recovery from Modewise Measurements via Normalized Block-Weighted Riemannian Gradient Descent
基于归一化分块加权黎曼梯度下降的模态度量下低秩张量恢复的保证性研究
Zhou, Yushi, Zhang, Feng
Abstract
We consider the recovery of low-multilinear-rank tensors from linear measurements and propose an adaptive block-weighted modewise Riemannian gradient descent method. The method combines memory-efficient modewise measurements with a normalized adaptive weighting strategy for the core and factor components of the Riemannian gradient. The weighting improves convergence without increasing the multilinear-rank bound of the search direction or the size of the reduced core used for retraction. Under the tensor restricted isometry property and a suitable initialization, we establish local linear convergence and derive sampling guarantees for sub-Gaussian and subsampled orthogonal with random sign (SORS) measurements. Numerical experiments on synthetic low-Tucker-rank tensors show that the proposed method reduces iteration counts and computational time while maintaining reliable recovery performance, especially near the recovery threshold and for structured SORS measurements.
While a centralized approach involving patient consent to collect and analyze data centrally would theoretically offer the best data quality and predictive performance, it is not always feasible in practice. Federated Learning (FL) architectures have shown to be a very promising approach to use and access distributed disease related resources within the GDPR boundaries. In a previous case report, we described the preconditions at the participating sites and necessary administrative and process related steps to prepare data, people and infrastructure for improving subtype identification and assessing treatment options in pancreatic cancer. We update this report sharing our experience in tackling the challenges and show preliminary results of the actual federated learning AI pipelines. At the participating sites, we have to identify and annotate the data being accessible after extraction and transformation in a local FL hub - in our case a centrally developed and distributively deployed Docker container. This container comprises the FL scripts generating local models. We apply a newly developed FL algorithm considering all local features, including partial overlapping features specific to the local sites. Theoretically, an annotation in a cancer setting should succeed using the German oncology core data set (oBDS), which is already utilized for mandatory reporting to cancer registries, and can be sustained in the FL setting. The FL algorithms deal robustly with partially overlapping features as we showed with public data sets. Major roadblocks including straightening operational concepts for the infrastructures, ethics approval for such novel architectures and support for every site have been addressed. However, scaling up this approach in the future faces hurdles; while including broader multi-modal data sets should be feasible, large-scale deployment to more sites remains challenging.
An Exact Junction-Tree Extended Formulation for Optimal Classification Trees
基于联合树(Junction-Tree)扩展公式的最优分类树精确求解方法
TU, Jiancheng, WenqiFan
Abstract
We develop an exact linear programming (LP) formulation for bounded-depth classification trees with binary features, using a junction-tree representation. The formulation is integral and supports recursive subtree optimization. Exact reductions make the model smaller while preserving the optimal value and recovery of an optimal tree. The reduced model supports two solution methods: column generation and message passing. Column generation solves integral restricted LPs and uses bounds over the full feasible domain to certify optimality. Message passing recursively combines optimal subtree costs. Both methods solve common subtree problems that, once the preceding tree decisions are fixed, can be evaluated independently and in parallel. Computational experiments show that the exact reductions substantially reduce the size of the junction-tree formulation. The resulting linear programming formulation certifies instances for which the tested mixed-integer formulation does not establish optimality within the same computational budget, while the column-generation and message-passing methods certify more instances and achieve an order-of-magnitude reduction in geometric-mean runtime relative to an existing state-of-the-art exact method for optimal classification trees.
Enhancing Transformer Representations of Symbolic ODE Expressions
增强符号常微分方程表达式的Transformer表示
Fan, Xiyue, Prugel-Bennett, Adam, Middleton, Stuart E.
Abstract
Existing approaches to solving differential equations, such as symbolic regression, physics informed neural networks, and neural operators, typically focus on numerical approximations or blind symbolic search via fitting to numerical data. Less attention has been paid to learning structured representations of mathematical expressions that preserve commutative properties and could support mathematical reasoning in symbolic forms. Transformer models have shown strong capabilities in solving symbolic differential equations. However, standard positional embeddings in transformers are designed for sequence data. Symbolic differential equations are naturally represented by expression trees, so these positional embeddings may not efficiently capture their hierarchical structures. We investigate existing tree positional embeddings in symbolic ordinary differential equation (ODE) tasks. We systematically study their effectiveness under different settings. Our results show that tree positional embeddings aid learning in early epochs and continue to improve performance throughout, ultimately yielding consistent advantages across various data sizes and tasks. Based on learned structural representations, we apply contrastive learning to support the commutative property in mathematics. Ablation studies provide insight into how these methods interact in modelling symbolic mathematical structures.
Inference of Unknown Dynamical Components Using Next Generation Reservoir Computing: From Chaotic Systems to Climate Data
基于下一代储备池计算的未知动力学分量推断:从混沌系统到气候数据
Budnick, Jule, Keane, Andrew, Yanchuk, Serhiy
Abstract
We investigate next generation reservoir computing (NGRC) as a data-driven approach for inferring unseen components of dynamical systems. We compare NGRC with traditional reservoir computing (RC) using the Lorenz and R\"ossler system, where two unknown components are inferred from one given component. For both systems, NGRC achieves accurate results while requiring fewer training data and less computational time than RC. We identified an inverse proportional behavior between the number of time-delayed steps needed for NGRC and the temporal resolution, indicating that the physical time span covered by the delay interval is an important factor in determining the required number of delayed steps. Finally, we apply NGRC to the observational climate data of ENSO (El Ni\~no--Southern Oscillation) and infer one observable from the remaining variables. Despite the noise and complexity of the real-world data, the NGRC shows promising results. Our findings demonstrate the potential of NGRC for efficient inference of unseen components in both controlled dynamical systems and real-world data.
Linear RNNs based on the delta-rule enable efficient sequence modeling, but their linear updates with a low-rank correction constrain their expressivity. Prior work has shown that composing two delta-rule transitions in a single recurrent update can model a 2D rotation, but this increases the rank and the cost of the updates compared to a single transition. We show that Kimi Delta Attention (KDA) can realize 2D rotations by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate. This requires extending the parameter ranges of KDA by combining two existing range extensions: allowing gates in $[-1,1]$ and the delta-rule coefficient $\beta$ in $[0,2]$. We call the resulting model Complex KDA (CKDA). It preserves KDA's stability and efficiency, with transitions that remain diagonal-plus-rank-one and non-expansive, while reaching the state-tracking expressivity of DeltaProduct$_2$. We characterize the expressivity of CKDA and prove that every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix. A single CKDA layer can track every finite group isomorphic to a subgroup of $\mathrm{SO}(3)$, and many state-tracking results use one fewer layer for CKDA compared to other diagonal-plus-rank-one Linear RNNs. Empirically, combining both extensions yields the strongest length extrapolation among tested KDA range settings on $S_3$, $S_4$, and periodic audio continuation. In language modeling, CKDA outperforms Transformers and other linear RNNs, obtains similar results to a KDA baseline, and shows promising scaling behavior. Our code is open source at https://github.com/OpenEuroLLM/ComplexKDA and our models are available at https://huggingface.co/collections/openeurollm/complexkda.
G-NAC: Graph Neural Automata Clustering via Emergent Domain Formation
G-NAC:基于涌现域形成的图神经自动机聚类
Miller, Keith, Crawford, Tristan
Abstract
We introduce Graph Neural Automata Clustering (G-NAC), an unsupervised clustering method in which observations interact as cells on a fixed neighborhood graph. A shared recurrent graph-neural cellular rule evolves latent domain states through local interactions, which are converted into a rank-based spectral affinity for partitioning. Across 73 clustering tasks from 57 benchmark datasets, G-NAC achieved a mean adjusted Rand index (ARI) of 0.7951, comparable to Genie at 0.7941 and higher than the other evaluated baselines. Empirical training time and GPU memory scaled approximately linearly from 5,000 to 100,000 nodes. Learned transition rules also transferred from smaller source graphs to independent 100,000-node samples generated under matched conditions. These results demonstrate a recurrent graph-clustering formulation while identifying dependencies on graph quality, readout design, and source-target similarity.
Agentic time series forecasting concerns systems whose underlying mechanisms evolve, making the relative effectiveness of numerical models, reasoning strategies, and intervention rules inherently time-varying. Consequently, a time series agent must adapt the forecasts it produces and the orchestration policy that determines which components to trust and how to coordinate them. The deployment process naturally provides supervision for this adaptation as forecast horizons elapse and realized targets reveal the effectiveness of earlier decisions. Committing all numerical expert forecasts and candidate agent paths before target observation allows each realized outcome to evaluate the entire alternative set, providing delayed feedback without additional annotation. However, existing time series agents primarily incorporate prior experience through forecast refinement, reflection, or retrieval, without systematically converting realized outcomes into persistent updates to the joint orchestration policy governing later origins. To exploit this delayed feedback systematically, we introduce TimEvolve, a frozen-backbone time series agent that converts each realized outcome into persistent joint updates of expert trust, agent path selection, and intervention strength. A temporally ordered predict, reveal, and update protocol applies this feedback to subsequent forecasts. Experiments across eight Time-MMD domains show that TimEvolve achieves the best average MSE and MAE ranks among fifteen methods and the lowest errors on both metrics in seven domains. These results demonstrate the value of learning forecasting policies from the futures encountered during deployment.
Learning Prognostic Variables for AI Convective Parameterizations via Symbolic Distillation
通过符号蒸馏学习AI对流参数化方案的预报变量
Schönfeld, Jurij, Beucler, Tom, Savre, Julien, Sherwood, Steven, Eyring, Veronika
Abstract
Hybrid AI-physics climate modeling aims to improve coarse (~100km-resolution) Earth system models by learning to parameterize subgrid processes from high-fidelity data. However, this so far mostly involves local-in-time, diagnostic parameterizations, in which the subgrid state depends only on the current coarse state with no memory of previous states, which is unrealistic for processes such as convection that have intrinsic persistence. To address this, we enhance local-in-time parameterizations by learning prognostic variables that compactly carry important, additional past information where no explicit sub-grid information is available. First we compress past information into a low-dimensional latent space using an autoencoder, which then informs a neural network trained to parameterize targeted subgrid-scale processes. We then replace the autoencoder with symbolic equations that govern the time evolution of the latent variables, yielding additional prognostic memory variables that can be integrated alongside the resolved atmospheric state. We evaluate this approach on two systems: the Lorenz-96 model (online) and surface precipitation from high-resolution atmospheric simulations (offline). A forced multivariate linear ordinary differential equation recovers most of the added value achieved by the autoencoder-based approach in both experiments. Benchmarked against diagnostic parameterizations without memory, our memory-informed approach improves climate statistics and temporal structure, including a realistic diurnal cycle of tropical land precipitation.
Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization
推理处的精确性:一种面向分布外泛化的表征判据
Rocha, Filipe Marinho, Dutra, Inês, Costa, Vítor Santos, Reis, Luís Paulo
Abstract
A model generalizes outside its training distribution only when it computes a representation structurally equivalent to the generating mechanism, not an approximation fitted to it. Such equivalence is necessary for exactness in and out of distribution, and extrapolation is governed by this exactness at inference, whatever its realization. Tensor Logic shows this: a zero-temperature contraction is equivalent to discrete logic, deducing in place with no artefact extracted, its tensors Boolean, its embeddings orthonormal, only its arithmetic continuous. Lacking infinite recursion it reaches Datalog, not Prolog, and though exact over closed domains it needs external memory to bind a novel entity. The criterion needs neither a discrete representation nor an extracted expression, and constrains inference, not training: an exact marginal in $[0,1]$ passes, a Neural Network thresholded to a hard label does not. Logic Tensor Networks fail it, while differentiable ILP and Tensor Logic at $T=0$ pass. Piecewise-affine extrapolation divergence and an inability to bind novel entities are two faces of a shortfall in exact representability. For hybrid architectures, a propagation rule follows: the output inherits the bounds of every fitted estimator on its path, explaining which axes fail in equivariant models and the ARC-AGI induction/transduction split. Only an exact hypothesis class certifies what the training data leave underdetermined: on a law-derived partition it finds the $56.3\%$ of distant queries that are answerable, which ensembles meet with false confidence and distance metrics rank backwards. Common inductive biases, from symmetries to memory, reach exactness only because humans inject them, an argument for inducing exact representations rather than fitting surrogates whose residuals, even at the arithmetic floor in training, diverge outside the data and compound under composition.
Mousavi, S. Mohammad, Kadeethum, Teeratorn, Bouklas, Nikolaos, Goswami, Somdatta
Abstract
Neural operators evaluate parametric partial differential equations cheaply but degrade sharply outside their training distribution. Physics-informed neural networks avoid dependence on labeled data, yet their optimization can be basin-fragile: when the governing residual admits multiple solutions, a PINN trained from scratch may converge to a physically incorrect state despite achieving a small residual. We show that these failure modes can be addressed jointly: an imperfect NO provides the structural prior needed to place a PINN in the correct solution basin, while the PDE residual refines the solution beyond the operator's accuracy. We introduce a three-stage framework that freezes the spatial basis of a physics-informed NO, extrapolates its solution branch to an out-of-distribution parameter using a polynomial continuation prior, and distills the resulting field into a fresh PINN. The NO need not be accurate at the target; it transfers solution-branch information, while PDE residual minimization in the PINN governs convergence. We evaluate the framework on three nonlinear PDEs: 1D viscous Burgers, 2D steady Allen-Cahn near a pitchfork bifurcation, and 2D steady lid-driven cavity flow. For Allen-Cahn, where the trivial solution satisfies the PDE residual exactly, a standard PINN collapses to the trivial zero branch, whereas distillation from the crude extrapolated operator recovers the non-trivial branch that matches the finite-difference reference. For the lid-driven cavity, extrapolating to a Reynolds number of Re = 3200 accelerates convergence to the correct physical state, achieving competitive accuracy using fewer parameters and optimization steps than recent literature baselines. These results establish a simple principle: an NO need not accurately predict the solution to be useful; it only needs to identify the correct basin from which PINN optimization can recover it.
As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events can occur, but on how often they might. We study the problem of estimating the probability of rare events that arise from stochastic variation in the agent's own actions. Estimating this type of risk requires searching over the combinatorially vast space of trajectories. Naive Monte Carlo is computationally prohibitive in this regime, and constructing effective importance sampling (IS) proposals requires coordinated changes to a context-dependent chain of conditional distributions. We develop a new IS method that perturbs the original model's weights to construct the proposal. The proposal is itself a differentiably parameterized language model, enabling gradient-based search over weight space. We formulate an objective that combines a differentiable surrogate for event amplification and an adaptive regularization scheme that dynamically balances amplification against estimator stability. We evaluate our approach on $\sim$120M and $\sim$2.6B models across three event families spanning 300+ rare events as rare as $10^{-9}$, with reference probabilities computed with $<10\%$ relative standard error. In our most verifiable settings, we observe that our IS estimator achieves over $800\times$ compute-weighted efficiency gains over naive Monte Carlo for events with probabilities lower than $10^{-7}$. Our implementation is available at https://github.com/namkoong-lab/iterative-unalignment.
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.
On-device large language models (`LLMs'), e.g. running on mobile phones, are ripe for improvement via personalization. The limited compute resources of mobile devices impose limits on model scale and thus model quality, making any realizable quality gains highly impactful. At the same time, their personal nature (i.e., the close coupling to a particular user) means that a given on-device LLM tends to be used in similar, predictable patterns over the course of time. This paper presents a novel method for personalizing on-device LLMs. It trains a hypernetwork to map a user's context tokens to a low-rank adaptation (`LoRA') well-suited to that user. Once the trained common artifacts are deployed to users' devices, each user uses the hypernetwork to synthesize (entirely on device) a personalized LoRA. This approach blends the benefits while avoiding the drawbacks of two existing approaches to LLM customization: in-context learning (`ICL') and parameter-efficient fine-tuning (`PEFT'). Like ICL (and unlike PEFT), the on-device phase of our approach is computationally feasible, requiring only forward passes through neural networks. Like PEFT (and unlike ICL), our approach modifies the `target' base LLM via weights (the LoRA), avoiding negative consequences (e.g. increased latency) associated with extending the input sequence. Our approach is particularly well-suited to the mobile device regime. Apart from the on-device compute and latency benefits mentioned, it also requires minimal additional storage, as internally its architecture partly leverages the same LLM weights as belong to the target LLM to be personalized. We demonstrate the benefits of LoRA-generating hypernetworks on several representative personalization datasets, comparing against baselines like ICL and PEFT. Of note, our personalization experiments focus on more challenging and less studied long-form text generation tasks.
Multi-turn tool-use failures can hinge on a single model call, yet reward variation alone does not reveal which call would benefit from training. When rewards depend on later interactions, their variation can reflect downstream randomness rather than differences between the current actions. We introduce Critical-State RL to identify trainable states in multi-turn interactions. Given task-defined candidate calls and local rewards, the method assesses whether each reward captures the action's effect on task success and whether improvement over a reference policy is possible. It then uses nested sampling to separate action-dependent reward variation from continuation noise and optimizes the policy at the selected states using contextual-bandit training. Experiments on the Berkeley Function Calling Leaderboard (BFCL) v4 compare training at diagnostic-selected states with training at alternative states. For missing-function tasks, the diagnostic selects the response after the tool becomes available; for missing-argument tasks, it selects the response before the missing argument is supplied. Training the selected responses improves performance, including about 14 percentage points on the missing-function task, while training the alternatives leaves performance flat or worse. We further apply the recipe across models and tasks, including logged repeat-call avoidance and memory management.
Chinese Translation
多轮工具使用的失败往往取决于单次模型调用,然而仅凭奖励的变异性无法揭示哪一次调用能够从训练中获益。当奖励依赖于后续交互时,其变异性可能反映的是下游的随机性,而非当前动作之间的差异。我们提出临界状态强化学习(Critical-State RL)来识别多轮交互中可训练的状态。给定任务定义的候选调用和局部奖励,该方法评估每个奖励是否捕捉了动作对任务成功的影响,以及相对于参考策略是否存在改进空间。随后,该方法利用嵌套采样(nested sampling)将依赖于动作的奖励变异性与延续噪声分离,并通过情境老虎机(contextual-bandit)训练在所选状态上优化策略。在伯克利函数调用排行榜(Berkeley Function Calling Leaderboard, BFCL)v4上的实验将诊断所选状态的训练与备选状态的训练进行了比较。对于缺失函数任务,诊断方法选择在工具变为可用之后的那次响应;对于缺失参数任务,诊断方法选择在缺失参数被提供之前的那次响应。训练所选响应能够提升性能,例如在缺失函数任务上提升约14个百分点,而训练备选响应则使性能持平或更差。我们进一步将该方案应用于多种模型和任务,包括基于日志的重复调用规避和内存管理。