Forskningsradar

Science Journals

Peer-reviewade publikationer — 55483 artiklar

Macroscopic Quantum Interference of the Center-of-Mass Motion of Levitated Superconducting Microparticles enabled by Magnetic Higher-Order Traps
arXiv:2607.03622v1 Announce Type: cross Abstract: We show how magnetostatic higher-order multipole traps can be used to generate macroscopic quantum interference of the motion of levitated superconducting microparticles. An appropriate combination of multipolar magnetic fields offers great versatility in constructing various trap potentials, including anharmonic trap potentials such as Duffing or double-well types. Crucially, the anharmonic trap potentials realize a nonlinearity on the order of hundred times the zero-point motion, i.e., on a length scale below nanometers. These anharmonic potentials allow for the generation of quantum features of the center-of-mass motion of a magnetically levitated superconducting microparticle. Importantly, they can be easily generated with a static arrangement of coils, requiring only that the current running through them is tunable. We propose protocols exploiting the versatility of the magnetic trap landscape to generate non-Gaussian motional states. We solve the dynamics of the center-of-mass motion of the particle in phase space and analyze its parameter dependence. Furthermore, we give a recipe to distinguish classical from quantum behavior in a statistically meaningful way through measurement of the position of the particle. Our results open a path to accessing the quantum regime of the center-of-mass motion of objects with masses larger than picogram, i.e., $10^{13}$ atomic mass units. This will enable fundamental physics experiments for studying the transition between quantum and classical behavior, exploring the intersection between quantum physics and gravity as well as probing of certain types of dark matter.
Machine Unlearning via Information Theoretic Regularization
arXiv:2502.05684v5 Announce Type: replace Abstract: How can we effectively remove or ``unlearn'' undesirable information, such as specific features or the influence of individual data points, from a learning outcome while minimizing utility loss and ensuring rigorous guarantees? We introduce a unified mathematical framework based on information-theoretic regularization to address both data-point unlearning and feature unlearning. For data-point unlearning, we introduce the \emph{Marginal Unlearning Principle}, an auditable and provable framework. Moreover, we provide an information-theoretic unlearning definition based on the proposed principle and provable guarantees on sufficiency and necessity of marginal unlearning. We then show that the proposed framework provides a natural solution to the marginal unlearning problem and yields auditable high-probability marginal-unlearning guarantees. For feature unlearning, the framework applies to deep learning with flexible training objectives. By combining flexibility in learning objectives with simplicity in regularization design, our approach is highly adaptable and practical for a wide range of machine learning and AI applications. From a mathematical perspective, we provide a unified analytic solution to the optimal feature unlearning problem with a variety of information-theoretic training objectives. Our theoretical analysis reveals intriguing connections between machine unlearning, information theory, optimal transport, and extremal sigma algebras. Numerical simulations support our theoretical findings.
CAP-CoT: Cycle Adversarial Prompt for Improving Chain of Thoughts in LLM Reasoning
arXiv:2604.23270v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has emerged as a simple and effective way to elicit step-by-step solutions from large language models (LLMs). However, CoT reasoning can be unstable across runs on long, multi-step problems, leading to inconsistent answers for unchanged task. Most prior work focuses on improving the forward reasoning chain within a single pass, with less attention to iterative and contrastive correction. To address this gap, we propose CAP-CoT, a Cycle Adversarial Prompt optimization framework designed to improve both CoT reasoning accuracy and stability of a single deployed solver. In each cycle, a forward solver generates candidate reasoning chains, an adversarial challenger constructs plausible but deliberately flawed chains using targeted error strategies, and a feedback agent contrasts the two chains and produces step-aligned structured feedback. This feedback closes the optimization loop in two directions, including updating the solver prompt based on errors exposed by the challenger, and updating the challenger prompt to generate increasingly targeted errors in subsequent cycles. Unlike safety-oriented adversarial prompting such as jailbreak or prompt-injection attacks, our adversarial component is task-semantic and aims to expose logical vulnerabilities in reasoning chains. Experiments across six benchmarks and four LLM backbones demonstrate that within two to three adversarial prompt optimization cycles, CAP-CoT consistently reduces variability across runs while improving reasoning accuracy and robustness to prompt perturbations.
A Real-Time Remote-Sensing-Guided Decision-Support Framework for Cloud-Seeding Operations: A Field Demonstration Using Himawari-9 and C-band Phased Array Weather Radar
arXiv:2607.05050v1 Announce Type: new Abstract: This study proposes a real-time remote-sensing-guided decision-support framework for cloud-seeding operations using high frequency geostationary satellite and ground weather radar observations. The framework integrates cloud assessment, human-in-the-loop decision support, and aircraft operation to translate high-frequency remote-sensing information into actionable guidance for seeding aircraft. We demonstrate the framework using 2.5-min Himawari-9 geostationary satellite observations and 60-s C-band phased-array weather radar (C-PAWR) observations during the preliminary dry-ice cloud-seeding field campaign conducted over Toyama Bay, Japan, in January 2026. In the 13 January case, the framework enabled the ground team to identify a developing cumulus cloud with a lifetime of approximately 20 min, communicate guidance to the aircraft, and conduct seeding immediately before the cloud began to dissipate naturally. Candidate seedable clouds were identified from Himawari-9 infrared indices, and their selection was supported by near-real-time C-PAWR observations of precipitation echoes. Because the released dry-ice amount was limited to 30 kg, this study does not attempt to attribute subsequent cloud evolution to seeding effects. Instead, the results demonstrate that rapid-scan satellite and ground radar observations can support real-time target selection and aircraft guidance for responsible, operationally feasible weather-intervention field experiments.
Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR
arXiv:2607.05051v1 Announce Type: new Abstract: End-to-end ASR models transcribe in a single pass, leaving no room for the decoder to revisit hard inputs. We propose LatentASR, a parameter-efficient method that adds continuous latent test-time scaling to a frozen ASR backbone. Two small trainable modules drive it: a Latent Adapter that iteratively refines a few latent prefix positions through bounded, stabilized updates, and a Value Head that predicts whether extra computation will help and halts the loop early. The Qwen3-ASR-0.6B backbone stays fully frozen, and we train only ~4M extra parameters. We activate this loop with a deliberately small, diverse 500-utterance training set. Under this minimal-data regime, standard adaptation methods all regress: full fine-tuning, LoRA, and prompt tuning each increase WER. LatentASR is the only tested method that reduces WER on both clean benchmarks (FLEURS -2.54% and VoxPopuli -0.47% relative). The reductions are concentrated on intrinsically hard inputs. On accented and code-switched speech (ASCEND), LatentASR achieves a 16.0% relative CER reduction. Across 30 FLEURS languages (23,049 utterances), the multilingual WER decreases uniformly across resource tiers, confirming that the adapter generalizes without overfitting. Dynamic halting preserves most of the clean-set reduction at a fraction of the compute, skipping roughly half of all utterances at the entry gate. Our results show that a small, carefully chosen activation set can switch on test-time scaling inside a frozen ASR model without corrupting the model itself, converting fixed per-utterance compute into input-dependent compute where it is most needed.
An AI-Assisted Solution to the Signed BAR Conjecture: Uniqueness in the Harrison--Reiman Class and a Completely-$\mathcal{S}$ Class Obstruction
arXiv:2607.03639v1 Announce Type: cross Abstract: For a multidimensional reflected diffusion, determining whether the associated basic adjoint relationship (BAR) uniquely characterizes the stationary distribution is a basic uniqueness problem in the BAR approach. The problem has remained unresolved for more than 35 years since the introduction of the BAR approach. In this paper, we resolve the finite-signed uniqueness problem for stable Harrison--Reiman data with a nonsingular $M$-matrix reflection matrix. The proof uses pathwise differentiability of the reflected diffusion implies feasible directional differentiability of the probabilistic resolvent to show that, at boundary points, its one-sided initial-state derivative factors through the tangent projection and vanishes along active reflection directions. An interior one-sided convolution then yields smooth test functions whose oblique derivatives are uniformly bounded and converge pointwise to zero on each closed face. The interior signed measure is consequently invariant for the reflected semigroup. The proof was discovered with the assistance of ChatGPT 5.5 Pro and subsequently verified by the authors. We also show that the nonsingular $M$-matrix assumption is structural. In the larger completely-$\mathcal{S}$ class, a nonsingular reflection matrix with a singular proper principal block admits boundary gauges supported on lower-dimensional strata. Under standard exponential ergodicity and a mild one-step regulator bound, these gauges produce nonzero zero-mass signed BAR tuples; indeed the zero-mass interior BAR coordinates contain an infinite-dimensional subspace. A four-parameter three-dimensional family, including an explicit rational example, verifies the obstruction. Thus the finite signed version of the Dai--Dieker question has a positive answer in the Harrison--Reiman $M$-matrix class and a negative answer in a natural completely-$\mathcal{S}$ extension.
Diffusion learning reveals viable parameter manifolds and compensation geometry in biological dynamical systems
arXiv:2607.03671v1 Announce Type: cross Abstract: Models of complex systems often have many parameters, yet are constrained by far fewer experimentally accessible observables: similar activity can emerge from coordinated parameter changes. We formalize these compatible parameter sets as \emph{viable parameter manifolds}: the inverse images of a system's target dynamical behaviors under a parameter-to-feature map. The relevant codimension is not the number of reported features, but the effective rank of that map at the target scale. Co-varying features lower the codimension, while poor conditioning, high curvature, or regime mixing degrade learnability. We train conditional score-based diffusion models on simulated parameter--feature pairs and use them as amortized samplers of prior-weighted viable sets. In the Lorenz system, scalar trajectory statistics generate thin viable sheets, and two-feature conditioning localizes a transition-adjacent corridor. In the Izhikevich neuron model, four firing descriptors lie close to a nearly two-dimensional family of features, and the learned inverse images reveal distinct regular and irregular compensation geometries. In a recent ODE reduction of finite spiking networks, the same framework reveals excitatory--inhibitory compensation, timescale--coupling tradeoffs, and input-dependent viable manifolds across 4--12 parameter dimensions. In this view, robustness, compensation, and hidden parameter dependencies are organized as inverse geometry, with diffusion models providing practical tools for sampling, visualizing, and interrogating that geometry.
Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection
arXiv:2607.05052v1 Announce Type: new Abstract: Human value detection is commonly formulated as sentence-level multi-label classification over the 19 refined Schwartz values, typically predicted as independent labels. Schwartz theory, however, describes them as a circular motivational continuum, in which adjacent values are compatible and opposing values are in tension. We ask whether this structure can be operationalized as an explicit output-space geometry and used as a soft bias rather than a hard constraint. On a DeBERTa-v3-base classifier, we compare two ways of injecting it: training-time geometry-aware objectives and a post-hoc Schwartz-aware energy decoder that scores whole label sets jointly. Across five seeds, training-time geometry gives only limited gains-no larger for the true continuum than for a random ordering-whereas the decoder makes label sets more coherent with the continuum-on theory-aware coherence metrics we introduce-at no cost to Macro-F1 or Micro-F1 (held fixed by its selection rule). The gain is specific to the true Schwartz ordering: it does not appear for a random permutation or an empirical co-occurrence graph through the identical decoder. A bounded Qwen2.5-72B-Instruct diagnostic shows that supplying the continuum at inference shifts behavior but does not match supervised structured prediction. Theory-aware decoding thus offers a lightweight, controllable way to make value detection faithful to its label space.
Rethinking Prototype-based Similarity Learning for Few-Shot Object Detection
arXiv:2606.23069v3 Announce Type: replace Abstract: Few-shot object detection aims to detect novel object categories from only a few labeled examples, avoiding costly large-scale annotation. Recent prototype-based similarity learning approaches enable training-free adaptation by matching query features with class prototypes. However, they suffer from two fundamental limitations: (i) class confusion arising from inter-class similarity margin collapse, and (ii) insufficient visual cues for precise localization, as similarity scores capture only class-level semantic affinity while providing limited spatial information. To address these issues, we introduce two complementary components. Text-Anchored Semantic Mask (TSMa) leverages class-level text features as semantic anchors to identify semantically aligned channels through channel-wise interaction between visual and text features. By suppressing style-induced spurious responses and emphasizing class-intrinsic signals, TSMa enlarges inter-class similarity margins and mitigates class confusion. We further propose Stage-Aligned Hierarchical Autoregressive Regression (SHARe), which reformulates localization as a hierarchical autoregressive process that progressively refines bounding boxes across multiple stages. SHARe leverages the layer-wise characteristics of ViT representations by aligning feature abstraction levels with regression stages: deeper layers guide early coarse localization, while shallower layers rich in edge and texture cues refine spatial details in later stages. Experiments on COCO demonstrate a new state of the art, outperforming the previous best by +10.1 nAP, with extensive analysis validating each component. The code is available at https://github.com/VisualScienceLab-KHU/ReSet.
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
arXiv:2503.06269v3 Announce Type: replace Abstract: Traditional white-box methods for creating adversarial perturbations against LLMs typically rely only on gradient computation from the targeted model, ignoring the internal mechanisms responsible for attack success or failure. Conversely, interpretability studies that analyze these internal mechanisms lack practical applications beyond runtime interventions. We bridge this gap by introducing a novel white-box approach that leverages mechanistic interpretability techniques to craft practical adversarial inputs. Specifically, we first identify acceptance subspaces - sets of feature vectors that do not trigger the model's refusal mechanisms - then use gradient-based optimization to reroute embeddings from refusal subspaces to acceptance subspaces, effectively achieving jailbreaks. This targeted approach significantly reduces computation cost, achieving attack success rates of 80-95\% on state-of-the-art models including Gemma2, Llama3.2, and Qwen2.5 within minutes or even seconds, compared to existing techniques that often fail or require hours of computation. We believe this approach opens a new direction for both attack research and defense development. Furthermore, it showcases a practical application of mechanistic interpretability where other methods are less efficient, which highlights its utility. The code and generated datasets are available at https://github.com/Sckathach/subspace-rerouting.
On the Physical Plausibility and Distribution Alignment for Sim-to-Real RF Positioning
arXiv:2607.04400v1 Announce Type: new Abstract: Reliable radio frequency (RF) positioning from cellular measurements is limited by the high cost and limited coverage of real drive-test data, especially when models must work on streets not seen during training. Previous work showed that ray tracing simulations can provide useful synthetic data for pretraining deep positioning models. In this paper, we focus on the simulation side and study how base-station calibration, physical realism, synthetic-data scale, and RSSI distribution alignment affect transfer to real data. Using a Sionna reconstruction of a Rome deployment, we calibrate each base station by adjusting its location, height, azimuth, and transmit power. We compare physically plausible calibrations with unconstrained ones that allow unrealistic base-station placements. We also compare deployment-specific synthetic data with much larger city-scale datasets. Although unconstrained calibration matches measured RSSI better, it does not always improve positioning accuracy. All synthetic pretraining approaches improve performance on known streets, with the best result obtained using city-scale unconstrained data. However, larger synthetic datasets alone do not improve performance on unseen streets. The best results on held-out streets are achieved only after normalizing simulated RSSI values to better match the real distribution. Overall, the results suggest that distribution alignment is more important than physical realism or dataset size for sim-to-real RF positioning.
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
arXiv:2602.24264v2 Announce Type: replace Abstract: Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems. Although modern models are trained on massive datasets, they still cover only a tiny fraction of the combinatorial space of possible inputs, raising the question of what structure representations must have to support generalization to unseen combinations. We formalize three desiderata for compositional generalization under standard training (divisibility, transferability, stability) and show they impose necessary geometric constraints: representations must decompose linearly into per-concept components, and these components must be orthogonal across concepts. This provides theoretical grounding for the Linear Representation Hypothesis: the linear structure widely observed in neural representations is a necessary consequence of compositional generalization. We further derive dimension bounds linking the number of composable concepts to the embedding geometry. Empirically, we evaluate these predictions across modern vision models (CLIP, SigLIP, DINO) and find that representations exhibit partial linear factorization with low-rank, near-orthogonal per-concept factors, and that the degree of this structure correlates with compositional generalization on unseen combinations. As models continue to scale, these conditions predict the representational geometry they may converge to. Code is available at https://github.com/oshapio/necessary-compositionality.
Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning
arXiv:2508.01651v2 Announce Type: replace Abstract: 3D affordance grounding aims to understand how diverse objects can be manipulated, making it a cornerstone of embodied interaction. However, prior works struggle to generalize to out-of-distribution, open-world scenarios, leaving a critical gap between limited dataset performance and real-world application needs. Inspired by the saying: \textit{\textbf{``What I can not create, I do not understand''}}, we find generative models can generate semantically valid HOI images, which indicates inherent encoding of affordance concepts. Building on this insight, we propose DAG, the first innovative diffusion-based 3D affordance grounding framework that extracts general affordance knowledge from text-to-image diffusion models for 3D affordance prediction. Specifically, we extract the affordance priors from a diffusion model to encode HOI priors, and design an affordance block with a multi-source affordance decoder for dense 3D affordance prediction. Extensive experiments show that DAG consistently outperforms state-of-the-art methods and exhibits strong open-world generalization, even in the challenging one-shot setting. The code of our method is released on \textcolor{blue}{\textit{https://github.com/hq-King/DAG}}.
A quadratic lower bound for 2DFAs against one-way liveness
arXiv:2602.24279v2 Announce Type: replace Abstract: We show that every two-way deterministic finite automaton (2DFA) that solves one-way liveness on height h has Omega(h^2) states. This implies a quadratic lower bound for converting one-way nondeterministic finite automata to 2DFAs, which asymptotically matches Chrobak's well-known lower bound for this conversion on unary languages. In contrast to Chrobak's simple proof, which relies on a 2DFA's inability to differentiate between any two sufficiently distant locations in a unary input, our argument can be applied to inputs over any alphabet and is structured around a main lemma that is general enough to potentially be reused elsewhere.
The Good, the Bad, and the Brittle: Benchmarking Robustness and Generalisation of Histopathology Foundation Models
arXiv:2607.04401v1 Announce Type: new Abstract: How robust and generalisable are pathology foundation models and have their scaling limites been reached? We benchmarked twelve pathology foundation models (PFMs) and ResNet baselines using our Robustness Evaluation and Enhancement Toolbox (REET) across eleven clinically realistic perturbations and a dissimilarity-driven Non-Redundant K-fold validation (NR-Kfold) protocol. We introduce a Perturbation Performance Index (PPI) to summarise accuracy trends under controlled perturbation sweeps and analyse robustness scaling with parameter count. We show that PFMs consistently outperform CNNs in both robustness and domain generalisation, yet model scaling shows diminishing returns: mid-sized models such (UNI2/Virchow-2 etc.) achieve comparable or greater resilience than larger systems. NR-Kfold analysis further reveals systematic accuracy loss and increased variability when training-test similarity is broken, underscoring the need for explicit distribution-shift evaluation. These findings suggest that the next generation of pathology foundation models must prioritise data quality, multimodality information and domain alignment over parameter count to achieve genuine clinical reliability.
Neuro-symbolic Weak Supervision: Theory and Semantics
arXiv:2503.18509v2 Announce Type: replace Abstract: Weak supervision enables machine learning models to learn from limited or noisy labels, but it introduces challenges in reliability and semantic clarity, particularly in multi-instance partial label learning (MI-PLL), where models must resolve both ambiguous supervision signals and uncertain instance-label mappings. This paper proposes a semantics for a neuro-symbolic framework that integrates inductive logic programming (ILP) to structure MI-PLL through relational constraints. In this formulation, ILP defines a hypothesis space over label transitions, formalizes the semantics of per-instance classifiers and provides a relational scaffold for reasoning about weak supervision. Two inductive tasks are studied in this framework: inferring the transition predicate from the observed and classifier predicates, and inferring instance-level classifier assignments from the observed and transition predicates. This formal semantics facilitates constraint specification, consistency checking and the diagnosis of semantic failure modes that bag-level accuracy alone may conceal.
FastCSP: Accelerated Molecular Crystal Structure Prediction with Universal Model for Atoms
arXiv:2508.02641v2 Announce Type: replace Abstract: Molecular crystal structure prediction (CSP) is essential for applications in pharmaceuticals and organic electronics. However, CSP remains challenging and computationally intensive due to the need to explore a large search space with sub-kJ/mol accuracy to distinguish between competing polymorphs. While dispersion-inclusive density functional theory (DFT) offers the necessary precision, its computational cost is impractical for a large number of putative structures. Here, we present FastCSP, an open-source, end-to-end CSP workflow driven entirely by a single pretrained universal machine learning interatomic potential (MLIP), the Universal Model for Atoms (UMA), without any system-specific fine-tuning or DFT calculations. FastCSP integrates conformer generation, random structure generation via Genarris 3, geometry optimization, free energy evaluation, and conformer energy corrections, all powered by UMA. Benchmarked on 28 semi-rigid and 10 flexible molecules spanning 74 experimental polymorphs, FastCSP reliably recovers all known structures, ranking them within 9 kJ/mol of the global minimum. UMA reproduces dispersion-inclusive DFT results with high fidelity across chemically diverse compounds. Conformer corrections are particularly beneficial for flexible compounds with conformational polymorphism, such as ROY. UMA's accuracy, transferability, and computational cost thus eliminate the need for classical force fields in early-stage screening and DFT-based re-ranking in CSP workflows. The open-source release of the entire FastCSP workflow lowers the barrier to accessing CSP, enabling both pharmaceutical-grade and high-throughput polymorph screening within practical computational reach.
Surrogate-based prioritization of sub-problems for Benders decomposition in energy planning
arXiv:2607.05063v1 Announce Type: new Abstract: Benders decomposition solves optimization problems by separating the first-stage master problem from one or more second-stage sub-problems. While the standard Benders decomposition solves all sub-problems in each iteration, solving only selected sub-problems still guarantees convergence and can reduce solution time, but raises the question of how to select. In this work, we introduce surrogate-based prioritization of sub-problems. The method leverages surrogates to estimate the sub-problems' objectives, assess the current error of the cutting-plane estimator, and then prioritize the sub-problem with the largest error. We implement surrogate-based prioritization within sequential and asynchronous Benders decomposition. Both these algorithms also leverage the surrogate to trigger convergence checks and implement regularization. Benchmarks for an energy planning problem with a few large sub-problems show that the applied prioritization strategy works. The reduction in solution time correlates with the surrogate's accuracy. In our case, geometric interpolation-based surrogates are more accurate than machine learning methods. As a result, prioritization consistently and significantly outperforms the standard algorithm in sequential Benders decomposition. The speed-up increases with the number of scenarios, reaching 33\% with four scenarios and 55% with ten scenarios. In the case of asynchronous parallelization, the impact on performance is less clear, and the average speed-up from prioritization is 19%.
SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
arXiv:2508.03177v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, they largely overlook the potential risks posed by stylized images, which play crucial roles in critical scenarios such as game scene understanding, art education, and medical analysis. In this work, we first construct a dataset comprising photographic images and their corresponding stylized versions with carefully annotated caption labels. We then conduct head-to-head comparisons on both discriminative and generative tasks by benchmarking 13 advanced LVLMs on the collected datasets. Our findings reveal that stylized images tend to induce significantly more hallucinations than their photographic counterparts. To address this issue, we propose Style-Aware Visual Early Revision SAVER, a novel mechanism that dynamically adjusts LVLMs' final outputs based on the token-level visual attention patterns, leveraging early-layer feedback to mitigate hallucinations caused by stylized images. Extensive experiments demonstrate that SAVER achieves state-of-the-art performance in hallucination mitigation across various models, datasets, and tasks.
A Spectral Generalisation of the Variance Ratio: Eigenstructure of Long-Horizon Portfolio Covariance and a Multi-Memory Factor Model of U.S. Equity Returns
arXiv:2607.03858v1 Announce Type: cross Abstract: We propose a multivariate generalisation of the Lo-MacKinlay (1988) variance ratio that decomposes long-horizon equity-return dynamics into separate return-channel and volatility-channel memory components across the cross-section of asset returns. The framework identifies a parsimonious five-factor model - capturing persistent, antipersistent, and multi-scale memory in returns and volatility - that fits four U.S. portfolio panels (the Fama-French 49-industry universe, its pre/post-1998 halves, and the Fama-French 100 size x book-to-market sort) and a European replication (Fama-French Europe 25), recovering seven stylised facts of long-horizon equity dynamics simultaneously across all five panels. Three findings carry economic content. (i) The same five-factor decomposition fits all five panels, indicating a cross-sectional structure robust to industry vs. size-and-value sorts, to sub-periods, and to U.S. vs. developed-European markets. (ii) U.S. equity volatility memory underwent a regime transition in the late 1980s - not at the static 1998 split-half boundary - with the slowest component of the volatility cascade lengthening from approximately two to four years; a 1000-replicate rolling-window bootstrap localises the transition with strictly non-overlapping 90% confidence bands separating pre- and post-transition windows. (iii) The cross-sectional loadings driving return-channel long memory are economically distinct from those driving volatility-channel cascade memory: a cross-channel beta-inversion test finds no panel with the positive alignment a single shared loading predicts, rejecting the shared-loading hypothesis toward anti-alignment on the two largest panels at Bonferroni p = 0.0004. Characteristics that predict return-momentum patterns therefore need not predict volatility-persistence patterns.
When AI Is Wrong on Purpose: How Students Respond to Buggy GenAI Code
arXiv:2607.05068v1 Announce Type: new Abstract: As Generative AI (GenAI) becomes increasingly central to software development, CS education is integrating prompt-centered workflows where students describe intended program behavior in natural language to elicit code. However, professional practice requires careful review and verification of GenAI-generated code that may appear correct while containing subtle faults. This creates a challenge for CS1-level activities, where current models often solve tasks correctly and reduce students' incentive to closely inspect generated outputs. We investigate how prompt-centered programming activities can be adapted to better foster these practices. Specifically, we explore an approach where realistic, runnable bugs are injected into otherwise correct solutions, thus requiring students to read and repair generated outputs. We analyzed 2,636 sessions from 917 students, and examined behavior across instances of naturally occurring prompt-related failures and deliberately injected bugs within each session. Our findings show that students responded differently across bug sources. Deliberately injected bugs more often led to direct code edits and higher next-attempt success, suggesting localized repair of near-miss solutions. Prompt-related failures instead more often led students to refine prompts by clarifying constraints, updating function signatures, adding edge cases, or reframing the task. Student reflections reinforce the emphasis on review and repair, describing useful practice in code understanding, code review, and debugging, as well as a more careful verification mindset and greater awareness of GenAI limitations. Ultimately, prompt-related failures and injected bugs together support a pedagogically useful GenAI workflow, where students practice both specification refinement through prompts and debugging through code editing.
Can Code Specify a System Precisely Enough to Formally Verify It?
arXiv:2607.05076v1 Announce Type: new Abstract: Formal verification is seldom applied to production software, because writing and maintaining a model has historically cost more than it returns. A companion study [1] extended SysMoBench [4] with a lower-cost alternative: specifications are graded against traces captured from the running system. It found that when large language models write the specifications, reliability is governed by the structure of the specification contract, not the language. This paper evaluates both on production software: the payment workflow of an operational restaurant point-of-sale system, which must keep the register, payment terminal, and payment processor in agreement. We report three results. First, the core protocol is correct relative to a hand-built, line-cited model under a precisely stated failure model. The audit found seven failure-handling gaps, nearly all with a common root cause; three were reproduced as real executions, and a patch closing them was re-checked with all failure gates enabled, after which a follow-up patch closed a defect the re-check itself exposed. Systematic extensions of the failure model (crash-restart, stale reads, two attempts) each found the windows they were designed to probe. Second, a single probe of the production payment sandbox exposed a response-shape divergence that makes an entire recovery ladder unreachable against the live API. The emulator-based audit could not detect it, because code and emulator share the same misreading: a correlated-oracle failure. Third, the companion study's central finding replicates across seven models from two vendors: contract structure, not language, governs what LLMs specify reliably. The replication concerns the ordering of contracts and the failure taxonomy, not the absolute level: only the strongest models reached the corpus ceiling, and the harder task restores discriminating power the benchmark had lost.
A Co-Design Framework for High-Performance Jumping of a Five-Bar Monoped with Actuator Optimization
arXiv:2604.06025v2 Announce Type: replace Abstract: The performance of legged robots depends strongly on both mechanical design and control, motivating co-design approaches that jointly optimize these parameters. However, most existing co-design studies focus on link dimensions and transmission ratios while neglecting detailed actuator design, particularly motor and gearbox parameter optimization, and are largely limited to serial open-chain mechanisms. In this work, we present a co-design framework for a planar closed-chain five-bar monoped that jointly optimizes mechanical design, motor and gearbox parameters, and control parameters for dynamic jumping. The objective is to maximize jump distance while minimizing mechanical energy consumption. The framework employs a two-stage optimization approach, where actuator optimization generates a mapping from gear ratio to actuator mass, efficiency, and peak torque, which is then incorporated into CMA-ES-based co-design optimization of the robot design and control parameters. Simulation results demonstrate an improvement of approximately 30.4% in jump distance and an 11.5% reduction in mechanical energy consumption compared to a nominal design, highlighting the effectiveness of the proposed framework for high-performance and energy-efficient planar jumping.
Multi Choice Min Prophet
arXiv:2607.05085v1 Announce Type: new Abstract: We study the minimization counterpart of the classic prophet inequality, often termed the min prophet or cost prophet inequality. Unlike the maximization setting, where simple threshold algorithms achieve half of the prophet's value, the minimization setting is significantly harder, with an exponential lower bound even for i.i.d.\ variables. We study a multi-choice relaxation in which the algorithm may select multiple variables and gets to choose the best amongst them (the minimum amongst those selected). Our goal is to minimize the expected number of selections while achieving a constant competitive ratio. For adversarial order, we show that a constant competitive ratio requires a nearly linear number of choices in expectation, ergo, $\Omega(n/\ln n)$. In contrast, we show that for the prophet secretary model (random order) one can attain constant competitiveness while requiring only an exponentially smaller expected number of choices i.e. $O(\ln n)$. We give a refined analysis and define $M$ to be the ratio of the minimum expected value of any single variable to the expected minimum value of all variables (the prophet's value) and present an algorithm that achieves a constant competitive ratio with $O(\min\{\ln \ln M, \ln n\})$ choices in expectation for the prophet secretary. We show that this is tight up to low order log factors even for the special case of the i.i.d. model. We also show that if we insist on a deterministic bound on the number of choices then every constant competitive algorithm requires $n$ choices. This holds even in the i.i.d.\ setting Finally, we consider a variant where both the algorithm and the adversary choose $r$ values and pay their sum, this is the minimization multi unit version. We extend our techniques to the multi-unit variant for i.i.d.\ variables, achieving a constant competitive ratio with a small expected number of choices.
His2Trans: A Knowledge-Guided Agentic Framework for Project-Level C-to-Rust Migration
arXiv:2603.02617v4 Announce Type: replace Abstract: C remains a major implementation language for operating systems, embedded platforms, and infrastructure software, but manual memory management continues to create security and maintenance costs. Rust is a practical migration target because it retains low-level control while enforcing stronger memory-safety checks. At project scale, especially under gradual C/Rust coexistence, migration is not a sequence of syntax-preserving function rewrites. A translator must preserve project interfaces, observable behavior, system interaction protocols, and low-level interoperability boundaries while staying consistent with migration choices already made in the codebase. We introduce His2Trans, a knowledge-guided agentic framework for project-level C-to-Rust migration. His2Trans reuses interface-level and fragment-level knowledge mined from historical C/Rust migrations to guide new translations toward Rust interfaces, wrapper choices, and local idioms already accepted in the evolving project. It then refines the assembled crate with project-level agentic feedback. On ten OpenHarmony modules, His2Trans reaches a 100.00\% incremental compilation pass rate, a 94.92\% Test Pass Rate, and a 16.35\% Unsafe Ratio. On eight open-source C projects, it reaches 100.00\% for both incremental compilation and Test Pass Rate, reducing Unsafe Ratio from 42.88\% under C2Rust to 8.59\%. These results support knowledge-guided migration and project-level agentic refinement as practical mechanisms for preserving observable behavior while reducing the unsafe burden of rule-based transpilation.