Forskningsradar

Science Journals

Peer-reviewade publikationer — 61549 artiklar

Bursty Arrivals, Smooth Sojourns: Non-Poissonian Temporal Dynamics in a Logistics Warehouse
arXiv:2607.04866v1 Announce Type: new Abstract: Warehouses are central nodes in logistics networks: they buffer material flows, synchronize heterogeneous actors, and absorb temporal mismatches between inbound and outbound operations. Yet most warehouse analyses still rely on aggregate performance indicators or on queueing assumptions in which event timing is stationary and approximately memoryless. Here we use one month of high-resolution pallet-level data from a large Spanish warehouse to characterize arrivals, departures, and outbound residence times from a statistical-physics perspective. Inter-arrival and inter-departure times are strongly heterogeneous and compatible with heavy-tailed, non-Poissonian behavior, whereas outbound sojourn times are more naturally described by a log-normal distribution, suggesting constrained service mechanisms with a characteristic operational scale. Disaggregation by logistics flow reveals systematic differences in burstiness, memory, and distributional similarity. A renewal-based aging analysis uncovers recurrent weekly accumulation and clearance cycles in the outbound buffer zone. Finally, a Little's-Law-inspired activity--sojourn scaling identifies two operational regimes: a near-linear baseline under regular turnover and a reproducible off-baseline branch associated with weekend accumulation and Monday dispatches. These results provide a compact diagnostic framework for temporal complexity in warehouse operations and show how limited but high-resolution industrial data can reveal operational structure invisible to aggregate throughput statistics.
Orbit recovery for spherical functions
arXiv:2508.02674v2 Announce Type: replace Abstract: Orbit recovery is a central problem in both mathematics and applied sciences, with important applications to structural biology. This paper focuses on recovering generic orbits of functions on ${\mathbb R}^{n}$ and the sphere $S^{n-1}$ under the rotation action of $SO(n)$. Specifically, we demonstrate that invariants of degree three (called the bispectrum) suffice to recover generic orbits of functions in finite-dimensional approximations of $L^2({\mathbb R}^n)$ obtained by band-limiting the spherical component and discretizing the radial direction. In particular, our main result explicitly bounds the number of samples in the radial direction required for recovery from the degree three invariants. From an application perspective, the most important case is $SO(3)$, which arises in many scientific fields, and in particular, plays a central role in leading structural biology applications such as cryo-electron tomography and cryo-electron microscopy. Our result for $SO(3)$ states that considering three spherical shells (i.e., samples in the radial direction) is sufficient to recover generic orbits, which verifies an implicit conjecture made in a paper of Bandeira et al. Our proof technique provides an explicit, computationally efficient algorithm to recover the signal by successively solving systems of linear equations. We implemented this algorithm and demonstrated its effectiveness on three protein structures.
Higher order PCA-like rotation-invariant features for detailed shape descriptors modulo rotation
arXiv:2601.03326v3 Announce Type: replace Abstract: PCA can be used for rotation invariant features, describing a shape with its $p_{ab}=E[(x_i-E[x_a])(x_b-E[x_b])]$ covariance matrix approximating shape by ellipsoid, allowing for rotation invariants like its traces of powers. However, real shapes are usually much more complicated, hence there is proposed its extension to e.g. $p_{abc}=E[(x_a-E[x_a])(x_b-E[x_b])(x_c-E[x_c])]$ order-3 or higher tensors describing central moments, or polynomial times Gaussian allowing decodable shape descriptors of arbitrarily high accuracy, and their analogous rotation invariants. Its practical applications could be rotation-invariant features to include shape modulo rotation e.g. for molecular shape descriptors, or for up to rotation object recognition in 2D images/3D scans maybe also for 3D scene understanding, or shape similarity metric allowing inexpensive comparison of objects modulo rotation avoiding costly optimization over rotations.
Open-Attribute Person Retrieval: Finding People Through Distinctive and Novel Attributes
arXiv:2508.01389v3 Announce Type: replace Abstract: Person retrieval in surveillance videos often depends on attributes described by witnesses or operators. However, the most useful cues in practice are not always common appearance descriptions (e.g., gender, clothing color), but rare and distinctive attributes that can sharply reduce the search space (e.g., holding a weapon, lying on the ground). Existing text-based person retrieval benchmarks and methods largely focus on identity-centric retrieval with common pedestrian descriptions, leaving such retrieval-critical attributes underexplored. In this paper, we introduce Open-Attribute Person Retrieval (OAPR), a practical retrieval setting that aims to retrieve all pedestrian instances matching a given attribute query, including rare or previously unseen visual concepts, regardless of identity. To support this task, we construct EPAD, an Expanded Pedestrian Attribute Dataset with 267,885 pedestrian images and a unified vocabulary of 65 attributes, including safety-critical actions, assistive devices, and object interactions that are rarely covered in prior benchmarks. We further propose GAP-CLIP, a lightweight CLIP-based framework that learns gated attribute-aware body-part representations for OAPR. Extensive experiments on EPAD demonstrate that GAP-CLIP achieves the strongest top-K retrieval performance on the full attribute space and on out-of-distribution attributes. The code and dataset are available at https://github.com/mlnjeongpark/Open-Attribute-Person-Retrieval.
GlaBoost: A Multimodal Structured Framework for Glaucoma Risk Stratification
arXiv:2508.03750v2 Announce Type: replace Abstract: Early and accurate glaucoma detection is critical to prevent irreversible vision loss, yet existing AI methods often rely on unimodal inputs and lack interpretability. We present GlaBoost, a multimodal gradient boosting framework that unifies three complementary signals for glaucoma risk prediction: fundus image embeddings from a pretrained convolutional encoder,free-text neuroretinal rim assessments encoded by a transformer-based language model, and structured ophthalmic biomarkers. These modalities are fused into a single representation and classified by an enhanced XGBoost model.On two real-world annotated datasets, GlaBoost consistently outperforms unimodal and generic multimodal baselines. Feature importance analysis highlights the cup-to-disc ratio, rim thinning, and the ISNT rule as the dominant predictors, yielding clinically consistent and interpretable decisions. GlaBoost offers a transparent and scalable foundation for multimodal decision support in ophthalmology.
Quasi-Trefftz spaces for a first-order formulation of the Helmholtz equation
arXiv:2509.08936v3 Announce Type: replace Abstract: This work is concerned with the development of quasi-Trefftz methods for first-order differential systems. It focuses on discrete quasi-Trefftz spaces, starting from their definition and including the construction of corresponding bases together with their computational aspect. This is the first attempt at constructing quasi-Trefftz bases for a problem governed by a first-order system without relying on an auxiliary scalar equation. A decoupling approach, with a second order scalar equation for the one unknown, is proposed here simply as a point of comparison to this new approach.
Industrial Data-Service-Knowledge Governance: Toward Integrated and Trusted Intelligence
arXiv:2601.04569v2 Announce Type: replace Abstract: The convergence of artificial intelligence, cyber-physical systems, and distributed networking has accelerated the evolution of industrial intelligence across edge, cloud, and cross-organizational communication environments. However, existing governance mechanisms remain fragmented across data management, service orchestration, and knowledge-based decision-making, making it difficult to ensure reliability, accountability, compliance, and explainability throughout the industrial intelligence stack. To address this gap, we present TRISK (TRusted Industrial Data-Service-Knowledge governance), a conceptual and taxonomic framework for trustworthy industrial intelligence. TRISK is grounded in a five-dimensional trust model covering quality, security, privacy, fairness, and explainability, and formalizes how trust is constructed, propagated, aggregated, and fed back across data, service, and knowledge layers in networked industrial systems. Through a structured synthesis of more than 100 representative studies, standards, and technical reports, we examine data governance as the foundation of trust construction, service governance as the mediation layer for trustworthy execution, and knowledge governance as the semantic anchor for reasoning, validation, and feedback adaptation. We further discuss industrial implementation patterns, cross-industry implications, and the role of emerging communication and computing technologies. Finally, we outline a future research roadmap toward adaptive, verifiable, and human-aligned industrial governance for Industry 5.0.
Ultra-high-speed line-scan Raman imaging
arXiv:2607.04877v1 Announce Type: new Abstract: Raman spectroscopic imaging has emerged as a potent tool due to its non-invasive nature and capability for chemical composition analysis. Line-scan Raman spectroscopy accelerates imaging speed by two orders of magnitude compared to point detection Raman methods. However, further enhancements in imaging speed were constrained by the readout speed of typically used charge-coupled device (CCD) spectroscopic detectors. We developed an ultra-fast line-scan Raman imaging technique based on recently available complementary metal-oxide-semiconductor (CMOS) detectors with low cost and read noise, and fast readout during exposure combined with a global shutter. Employing a high-efficiency transmissiongrating imaging spectrometer, we demonstrate imaging speeds up to two orders of magnitude faster than traditional line scan Raman imaging techniques and up to four orders of magnitude faster than point scan Raman methods, achieving Raman imaging up to 80 kHz spectral rate. We demonstrate that this technology is applicable to a variety of samples, including microplastics, biological cells, and tablets, creating images in an extremely short time frame, showcasing exceptional detection capabilities and the ability to reveal detailed information.
Symmetry all the way down
arXiv:2607.04887v1 Announce Type: new Abstract: Asymmetric trust generalizes classical symmetric quorum systems by allowing each process to specify its own failure assumptions. While this flexibility enables tolerance of strictly more failure scenarios, it is not known if, in these cases, it is actually possible to solve distributed tasks, and if so, which. We answer this question using the depth hierarchy for asymmetric trust (Amores-Sesar et al., OPODIS~'25), which characterizes how much a process must rely on others to solve a task. We prove that asymmetric trust does not increase the solvability of tasks requiring depth two or more, such as reliable broadcast or consensus. Specifically, for any Byzantine asymmetric quorum system, every failure scenario that permits solving a task requiring depth at least two can also be tolerated by a suitably constructed Byzantine symmetric quorum system. We show this via a compiler that transforms asymmetric quorum systems into symmetric ones. The additional failure patterns tolerated exclusively by asymmetric trust correspond to scenarios in which only simpler tasks requiring depth one or less (such as consistent broadcast) can be solved. We further prove that this result is tight in the depth hierarchy, meaning that there exist no compilers that produce symmetric quorum systems that are valid also in failure scenarios where correct processes have depths one or less. Our results clarify the precise power of asymmetric trust. While it strictly enlarges the set of tolerable failure patterns, it does not provide additional strength for solving tasks requiring depth two or higher.
Shadow Pricing of Static Voltage Stability Services within Unit Commitment for Inverter-Dominated Power Systems
arXiv:2607.05209v1 Announce Type: new Abstract: Modern power systems are increasingly dominated by Inverter-Based Resources (IBR), most of which work in Grid-following (GFL) mode. This implies that they do not directly control their terminal voltage, so the static voltage stability at these buses may be compromised, especially under constant-power-factor operation that lacks voltage-adaptive reactive support. In addition, weather-driven IBR are often installed in electrically remote areas with low Short-Circuit Ratio (SCR), further exacerbating voltage issues. To address this challenge, grid-forming control can be utilized to enhance low-SCR buses, while GFL-IBR could be explicitly required to provide voltage support through grid codes. As an alternative, a market mechanism could be devised that incentivizes relevant generators to proactively adjust their operating points as a service to maintain voltage stability, while the theoretical framework for such a market has not been developed. To fill this gap, this work adopts a second-order cone-based static voltage stability constraint for GFL-IBR buses within a unit commitment problem, and proposes a mechanism to assign shadow prices to this ancillary service. To determine appropriate price values under non-convex conditions, different pricing schemes are assessed. Using a modified IEEE 30-bus system, we demonstrate that both the dispatchable and restricted pricing methods can yield revenue-adequate service prices, though the former may deliver less efficient price signals and the latter may require well-defined uplift payments. This implies that, given differentiated pricing mechanisms and price signals, operators need to select a suitable pricing method in accordance with actual system conditions and market rules.
Efficient Perception in Automotive Detection and Tracking Using Neuromorphic Computing
arXiv:2607.04921v1 Announce Type: new Abstract: Deep learning algorithms are notorious for their high carbon footprint and computational demands that limit their deployment on edge devices and raise concerns about their long-term sustainability. Neuromorphic computing and Spiking Neural Networks (SNNs) offer a promising alternative to traditional Von Neumann architectures, providing energy-efficient performance, massively parallel computation, and on-chip learning capabilities. Autonomous machines represent a critical application domain where these advantages are particularly valuable. We present the first comprehensive evaluation of SNNs for real-world automotive multi-object detection and tracking. Using transfer learning with the SpikeYOLO architecture, we achieve mean Average Precision of 0.937 on the KITTI dataset and 0.771 on BDD100K MOT2020 dataset for object detection and a Higher Order Tracking Accuracy score of 0.701 (KITTI) and 0.445 (BDD100K MOT2020) for object tracking--results competitive with conventional deep learning methods. Our results demonstrate that SNNs can deliver high-performance object detection and tracking in an energy efficient manner, establishing their viability for perception in real-world autonomous systems.
Rational Inverse Reasoning: Few-Shot Imitation by Inferring Intent through Planning
arXiv:2508.08983v2 Announce Type: replace Abstract: Humans can learn a new manipulation task from one or two demonstrations and then perform it in a new room, with new objects, under new constraints. Modern robot imitation learning, in contrast, typically needs hundreds to thousands of demonstrations and still degrades under modest shifts in layout, geometry, object set or task constraints. We argue this gap is not just about data, but also about the level of abstraction at which learning occurs; generalization requires inferring the latent intent underlying why a demonstrator behaved in a certain way, rather than reproducing how they moved. We present Rational Inverse Reasoning (RIR), which casts few-shot imitation as inference over latent explanation programs: compact, executable descriptions of intent that map an object-centric scene to a structured task-and-motion-planning (TAMP) specification of goals, subgoals and constraints. A vision-language model proposes candidate programs, and a hierarchical planner supplies a bounded-rational likelihood. By combining VLM program proposals, and planner-grounded feedback, RIR iteratively refines the candidate set to approximate a posterior over concise, executable programs. On a 2D reasoning benchmark and a real Franka FR3, RIR recovers transferable task structure from as little as one demonstration. Generalizing to substantially new layouts and object sets, RIR outperforms VLM-planning baselines that lack explicit rationality and planning-grounded inference, increasing downstream success rate by $34$ and $28$ percentage points in the one- and three-shot settings.
Role of SKA in Advancing Remote Measurements of Magnetic Fields of Solar Coronal Mass Ejections
arXiv:2606.31440v2 Announce Type: replace-cross Abstract: Coronal Mass Ejections (CMEs) are large expulsions of magnetized plasma from the Sun into interplanetary space and are the primary drivers of extreme space weather variations. The strength and topology of CME magnetic fields largely determine their impact on Earth. Although visible-light coronagraphs routinely observe CMEs and provide their geometric and kinematic properties, they cannot directly measure CME vector magnetic fields. These fields evolve from initiation through the inner heliosphere due to interactions with other CMEs, coronal structures, and the ambient solar wind, leading to significant structural deformation. Such evolution complicates predictions of the CME magnetic field at Earth. Accurate measurements of CME magnetic fields in the corona and heliosphere are therefore essential for advancing space weather forecasting. Radio observations spanning MHz to GHz frequencies provide a powerful remote-sensing approach for measuring CME magnetic fields from the ground. Recent observations with Square Kilometre Array (SKA) precursors and pathfinder instruments, as well as other new-generation facilities, have demonstrated the potential of these radio techniques for CME magnetic-field diagnostics. At the same time, these studies have highlighted several limitations of current instruments. The higher sensitivity, wider instantaneous bandwidth, and broader frequency coverage of the SKA will open a new observational window, enabling these techniques to be fully exploited for constraining SpWx models and improving predictive accuracy. However, such observations are non-standard and require special consideration in scheduling, calibration, and imaging. Developments achieved with SKA precursors and pathfinders are paving the way for robust CME magnetic-field measurements with the SKA.
Predicting Biased Human Decision-Making with Large Language Models in Conversational Settings
arXiv:2601.11049v2 Announce Type: replace Abstract: We examine whether large language models (LLMs) can predict biased decision-making in conversational settings, and whether their predictions capture not only human cognitive biases but also how those effects change under cognitive load. In a pre-registered study (N = 1,648), participants completed six classic decision-making tasks via a chatbot with dialogues of varying complexity. Participants exhibited two well-documented cognitive biases: the Framing Effect and the Status Quo Bias. Increased dialogue complexity resulted in participants reporting higher mental demand. This increase in cognitive load selectively, but significantly, increased the effect of the biases, demonstrating the load-bias interaction. We then evaluated whether LLMs (GPT-4, GPT-5, and open-source models) could predict individual decisions given demographic information and prior dialogue. While results were mixed across choice problems, LLM predictions that incorporated dialogue context were significantly more accurate in several key scenarios. Importantly, their predictions reproduced the same bias patterns and load-bias interactions observed in humans. Across all models tested, the GPT-4 family consistently aligned with human behavior, outperforming GPT-5 and open-source models in both predictive accuracy and fidelity to human-like bias patterns. These findings advance our understanding of LLMs as tools for simulating human decision-making and inform the design of conversational agents that adapt to user biases.
When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On
arXiv:2603.05659v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) and Rubrics as Rewards (RaR) have driven strong gains in domains with clear correctness signals and even in subjective domains by synthesizing evaluation criteria from ideal reference answers. But many real-world tasks admit multiple valid outputs and lack the single ideal answer that rubric generation depends on. We identify this reference-free setting as a gap in current post-training methods and propose Implicit Error Counting (IEC) to fill it. Instead of checking what a response gets right against a rubric, IEC enumerates what it gets wrong, applying severity-weighted scores across task-relevant axes and converting them into calibrated per-aspect rewards. We show that na\"ive explicit enumeration is too noisy for stable optimization, and that two design choices: implicit score emission and group calibration are necessary to make error counting a reliable reward. As a case study, we validate IEC on virtual try-on (VTO), a domain that is simultaneously too constrained for holistic scoring and too permissive for rubric-based evaluation: subtle garment errors are unacceptable, yet many output variations are correct. We introduce Cascaded Error Counting (CEC) as an evaluation metric, which tracks human preferences well (60% top-1 vs. 30% others), and curate Mismatch-DressCode (MDressBench), a benchmark with maximal attribute mismatch to stress-test reward designs. On MDressBench, IEC outperforms RaR across all metrics (CEC: 5.31 vs. 5.60 on flat references; 5.20 vs. 5.53 on non-flat). On VITON-HD and DressCode, IEC matches or surpasses six baselines on 6 of 8 perceptual metrics. These results suggest that when ideal answers are unavailable, counting errors provide a stronger signal than constructing rubrics.
Graph Neural Networks are Heuristics
arXiv:2601.13465v4 Announce Type: replace Abstract: Graph neural networks are usually treated as auxiliaries for combinatorial optimization: they imitate algorithms, guide search, or supply scores to classical procedures. We show that this auxiliary role is not intrinsic. A GNN can itself be a heuristic. For the Euclidean Travelling Salesman Problem, we train a non-autoregressive GNN with no labels, rewards, sequential decoding, search, or local improvement. A differentiable Hamiltonian-cycle objective is the only supervision. The trained model produces a complete tour in one forward pass, while dropout and snapshots from a single training trajectory provide solution diversity without engineered moves. The heuristic is therefore learned, not programmed. It is also fast: batched inference remains in the millisecond regime on GPUs. Experiments on TSP100, TSP200, and TSP500 show that the model consistently improves over nearest-neighbor greedy baselines. These results identify unsupervised GNNs as a class of fast learned heuristics for combinatorial optimization.
MCLMR: A Model-Agnostic Causal Learning Framework for Multi-Behavior Recommendation
arXiv:2603.25126v2 Announce Type: replace Abstract: Multi-Behavior Recommendation (MBR) leverages multiple user interaction types (e.g., views, clicks, purchases) to enrich preference modeling and alleviate data sparsity issues in traditional single-behavior approaches. However, existing MBR methods face fundamental challenges: they lack principled frameworks to model complex confounding effects from user behavioral habits and item multi-behavior distributions, struggle with effective aggregation of heterogeneous auxiliary behaviors, and fail to align behavioral representations across semantic gaps while accounting for bias distortions. To address these limitations, we propose MCLMR, a novel model-agnostic causal learning framework that can be seamlessly integrated into various MBR architectures. MCLMR first constructs a causal graph to model confounding effects and performs interventions for unbiased preference estimation. Under this causal framework, it employs an Adaptive Aggregation module based on Mixture-of-Experts to dynamically fuse auxiliary behavior information and a Bias-aware Contrastive Learning module to align cross-behavior representations in a bias-aware manner. Extensive experiments on three real-world datasets demonstrate that MCLMR achieves significant performance improvements across various baseline models, validating its effectiveness and generality. All data and code will be made publicly available. For anonymous review, our code is available at the following the link: https://github.com/gitrxh/MCLMR.
Industrial3D: A Water-Treatment TLS Point Cloud Dataset and Cross-Paradigm Benchmark for MEP Scene Understanding
arXiv:2603.28660v2 Announce Type: replace Abstract: Automated semantic understanding of dense terrestrial laser scanning (TLS) point clouds is a prerequisite for Scan-to-BIM, digital twin maintenance, and as-built verifcation. Yet for operational industrial mechanical, electrical, and plumbing (MEP) facilities, this challenge remains largely unsolved: water-treatment TLS scans exhibit extreme geometric ambiguity, severe occlusion, and extreme class imbalance that architectural benchmarks such as S3DIS and ScanNet cannot adequately represent. We present Industrial3D, a terrestrial LiDAR dataset with 612.7 million expert-labeled points at 6 mm resolution from 20 room scenes, 13 dataset areas, and 7 operational water treatment facilities. At 6.6x the scale of the closest comparable MEP dataset, Industrial3D provides the largest industrial MEP testbed for within-domain scene understanding. We further establish a cross-paradigm benchmark of nine methods across fully supervised, weakly supervised, unsupervised, and foundation-model settings. The best supervised method reaches 55.74% mIoU, whereas zero-shot Point-SAM reaches 15.79%, a 39.95 percentage-point gap that quantifes unresolved domain transfer for industrial TLS data. Analysis attributes this gap to a dual crisis: 215:1 statistical rarity and cylindrical geometric ambiguity between tail classes and head-class pipes. The dataset, benchmark code, and pre-trained models will be publicly released at https://github.com/pointcloudyc/Industrial3D.
Palindrome complexity versus factor complexity
arXiv:2606.08127v3 Announce Type: replace-cross Abstract: Let ${\bf x} = (a_i)_{i \geq 0}$ be an infinite word over a finite alphabet $\Sigma$. Let $\rho (n)$ be the factor complexity function for $\bf x$ and ${\rm Pal}(n)$ be the palindrome complexity function for $\bf x$. We give a new relationship between these two quantities; namely, if $\bf x$ is not ultimately periodic, then $$ \lim_{n \rightarrow \infty} {{ {\rm Pal} (n) \log ({\rm Pal} (n) + 1)} \over {\rho (n)}} = 0. $$ Furthermore, we prove that the numerator in this result is essentially optimal.
ComplexMimic: Human-Scene Interaction Imitation in Complex 3D Environments
arXiv:2607.02034v2 Announce Type: replace Abstract: Physics-based Human-Scene Interaction (HSI) imitation learning is crucial for embodied intelligence as it bridges the gap between kinematic 3D motions and real-world dynamics. However, most existing methods focus on simplified scene settings, leaving complex environments largely unexplored, which limits their applicability in real-world scenarios. In this paper, we focus on HSI mimicry in complex environments. Under this complex setting, we observe an inherent trade-off between successfully performing interaction and maintaining natural, physically plausible motions. To address this challenge, we propose ComplexMimic, a framework that reconstructs diverse HSI by interpreting imperfect MoCap data. First, we introduce a Dual Flow Strategy, which learns two complementary experts: an imitation expert for accurate motion tracking and an interaction expert for collision-aware adaptation in complex scenes. Second, naive multi-expert distillation, which treats all experts equally, often under-samples challenging behaviors, limiting effective learning. To mitigate this issue, we propose a difficulty-aware distillation strategy that adaptively weights supervision and prioritizes hard-yet-learnable trajectories guided by failure statistics and learning progress signals. Extensive experiments on three benchmark datasets demonstrate that our approach outperforms current state-of-the-art methods.
Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
arXiv:2510.17801v2 Announce Type: replace Abstract: Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a central challenge. Recent embodied systems often follow a dual-system paradigm, where System 2 performs high-level reasoning and System 1 handles low-level control. We refer to System 2 as the embodied brain, the cognitive core for decision-making in manipulation. Although evaluating this embodied brain is crucial, existing benchmarks mainly measure execution success or cover only limited aspects of high-level cognition and task realism. We introduce RoboBench, a benchmark for evaluating multimodal large language models (MLLMs) as embodied brains. RoboBench covers five dimensions: Instruction Comprehension, Perception Reasoning, Generalized Planning, Affordance Prediction, and Failure Analysis. It spans 14 capabilities, 25 tasks, and 6,092 QA pairs. To improve realism, it draws from large-scale real robotic data and in-house collection across diverse embodiments, attribute-rich objects, multi-view scenes, and memory-driven navigation. For planning, RoboBench introduces an MLLM-as-world-simulator framework that assesses whether predicted plans can achieve critical object-state changes under physical and visual constraints, enabling more faithful evaluation of long-horizon reasoning than symbolic matching. Experiments on 18 state-of-the-art MLLMs reveal persistent limitations in implicit instruction understanding, spatiotemporal reasoning, cross-scenario planning, fine-grained affordance understanding, and failure diagnosis. We further analyze how embodied cognitive abilities relate to downstream robotic control. RoboBench offers a comprehensive scaffold for quantifying high-level cognition and guiding next-generation MLLMs toward more robust robotic intelligence.
Desirable Effort Fairness and Optimality Trade-offs in Strategic Learning
arXiv:2510.19098v2 Announce Type: replace Abstract: Strategic classification examines how decision rules interact with agents who strategically adapt their features. Most existing models focus on maximizing predictive performance, assuming agents best respond to the learned classifier. However, real decision-making systems are rarely optimized solely for accuracy: ethical, economic, and institutional considerations often make some feature changes more desirable than others. At the same time, principals may wish to incentivize these changes fairly across heterogeneous agents. While prior work has studied causal structure between features, notions of desirability, and information disparities in isolation, this work initiates a unified treatment of these components within a single framework. We frame the problem as a constrained optimization problem that captures the trade-offs between optimality, desirability, and fairness. We provide theoretical guarantees on the principal's optimality loss constrained to a particular desirability fairness tolerance for multiple broad classes of fairness measures. Finally, through experiments on real datasets, we show the explicit tradeoff between maximizing accuracy and fairness in desirability effort.
CableTract: A Co-Designed Cable-Driven Field Robot for Low-Compaction, Off-Grid Capable Agriculture
arXiv:2604.09938v2 Announce Type: replace Abstract: Conventional field operations spend most of their energy moving the tractor body, not the implement. Yet feasibility studies for novel agricultural vehicles rarely tie mechanics, energy harvest, draft, field geometry, economics, life-cycle CO2, and uncertainty quantification together on a single reproducible code path. This paper builds such a framework and applies it to CableTract, a two-module cable-driven field robot. A stationary Main Unit (winch + motor + battery + harvester module) (MU) and a lighter Anchor module (held by helical screw piles) tension a cable across a strip while a lightweight implement carriage rolls along it. The heavy bodies stay on the headland; only the carriage enters the field. The carriage runs a 10-implement library co-designed for the cable architecture. This co-design is the paper's central analytical lever. The framework is prototype-free. It chains a catenary cable model, a drivetrain efficiency chain, a stochastic draft model fitted to the co-designed library, an hourly solar + wind + battery simulator on six sites, a polygon coverage planner on a 50-field corpus, a contact-pressure compaction model, a discounted cash-flow economics engine with battery replacement and life-cycle CO2, and a global sensitivity analysis on 20 inputs. An operating-envelope sweep and an architectural-variant comparison close the loop. The full implementation is open source. Applied to the codesigned reference, the framework yields energy, compaction advantages and potential off-grid operation.
Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control
arXiv:2607.04978v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels, but every existing JEPA world model commits at training time to a single inference paradigm: either trajectory optimisation in a learned dynamics model, or direct behaviour cloning. A single checkpoint that serves both would defer this choice to inference, when deployment constraints (rollout cost, observation accessibility) determine which path wins. We present Qantara, an end-to-end JEPA whose joint training objective pairs a Brownian-bridge interpolant between consecutive clean latents on the state axis with noise-to-data flow matching on the action axis. The same checkpoint serves three inference paradigms without retraining: latent planning, behaviour-cloning action sampling, and inverse dynamics, which we query through a video-inverse composition that first predicts the next latent without action conditioning, then extracts the action. Training concentrates mass on the edges of the (action-time, state-time) noise square, where inference queries the predictor: replacing it with uniform interior sampling drops Push-T planning from 90.1 to 53.3 SR at matched compute. On the LeWM control suite, Qantara reaches a 91.2 SR three-train-seed average and sets new SOTA on OGBench-Cube (+7.7 SR over DINO-WM, +19.7 over LeWM). From the same weights, the behaviour-cloning and video-inverse paths reach 82-83 SR on Push-T and 71-73 SR on Cube. These results move JEPA world models from single-paradigm planners to multi-paradigm controllers.
CompLLM: Compression for Long Context Q&A
arXiv:2509.19228v2 Announce Type: replace Abstract: Large Language Models (LLMs) face significant computational challenges when processing long contexts due to the quadratic complexity of self-attention. While soft context compression methods, which map input text to smaller latent representations, have shown promise, their real-world adoption is limited. Existing techniques typically compress the context as a single unit, which leads to quadratic compression complexity and an inability to reuse computations across queries with overlapping contexts. In this work, we introduce CompLLM, a soft compression technique designed for practical deployment. Instead of processing the context holistically, CompLLM divides it into segments and compresses each one independently. This simple design choice yields three critical properties: efficiency, as the compression step scales linearly with the context length; scalability, enabling models trained on short sequences (e.g., 1k tokens) to generalize to contexts of 100k tokens; and reusability, allowing compressed segments to be cached and reused across different queries. Our experiments show that with a 2x compression rate, at high context lengths CompLLM speeds up Time To First Token (TTFT) by up to 4x and reduces the KV cache size by 50%. Furthermore, CompLLM achieves performance comparable to that obtained with the uncompressed context, and even surpasses it on very long sequences, demonstrating its effectiveness and practical utility.