Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Approximation of Analytic Functions by ReLU Neural Networks with Adjustable Depth and Width
arXiv:2607.10589v1 Announce Type: cross Abstract: In contrast to most studies on neural network approximation theory that characterize results through a single parameter, such as the total number of network parameters, \cite{shen2020deep} pioneered the characterization of approximation rates as a joint function of the width parameter $N$ and the depth parameter $L$, thereby granting greater architectural flexibility. Existing works using the $(N,L)$-characterization focus on function classes with finite smoothness $s$, establishing a typical approximation rate of $\mathcal{O}\left(N^{-2s/d}L^{-2s/d}\right)$ with $d$ denoting the input dimension, which indicates that network depth and width play symmetric roles for these classes. In contrast, this paper establishes upper bounds for the approximation of analytic functions, which possess infinite smoothness, via ReLU networks under the $(N,L)$-characterization. Specifically, we derive approximation rates of $\mathcal{O}\left(N^{-C L^{\tau}}\right)$, where $C>0$ is some constant and $\tau>0$ is a parameter influenced by the relation between $L$ and $N$. In particular, $\tau=1$ if $N$ scales roughly as $L^d$. Our findings reveal that depth plays a more critical role than width in the context of analytic function approximation. The main technical difficulty of obtaining such upper bounds lies in the trade-off between the smoothness parameters and the approximation accuracy. To overcome this difficulty, we employ refined constructions of several ReLU networks to approximate power functions, multivariate multiplication, and polynomials, which may be of independent interest.
Recognizability equals CMSO-definability for graphs of rank-width at most two
arXiv:2607.10594v1 Announce Type: cross Abstract: We prove that, on finite graphs of rank-width at most two, VR-recognizability and counting monadic second-order definability coincide. This advances the recognizability-versus-definability problem from bounded linear clique-width to the first nontrivial bounded rank-width level beyond the rank-width-one split-decomposition case. The proof first treats split-prime graphs. The maximal partial-tree theory of Clark and Whittle organizes the non-sequential cut-rank-two separations, while a single strong separation orients all strong equivalence classes and yields a CMSO-definable laminar family of canonical cores. Although the auxiliary partial tree is not itself transduced, it proves that every canonical local piece has a port-contiguous layout of uniformly bounded linear rank-width. The width argument uses partition atoms and the branch-width-three display theorem of Hall, Oxley, Semple, and Whittle and does not assume that graph torsos remain prime. Coherent ordered rank-two frames then permit a finite-state bottom-up evaluation whose local transitions are definable by the bounded-linear-clique-width theorem of Boja\'nczyk, Grohe, and Pilipczuk. Finally, the CMSO-transducible canonical split decomposition lifts the result from prime graphs to arbitrary graphs of rank-width at most two.
On the Existence of Almost Periodic Solutions with Applications to Global Entrainment
arXiv:2607.10838v1 Announce Type: cross Abstract: This paper provides two results that are useful in the study of the existence and the stability properties of almost periodic solutions for a given dynamical system. The obtained results are generalizations of recent results for periodic systems and are applied to the global entrainment problem in nonlinear time-invariant control systems. It is shown that local exponential stability for the unforced system and input-to-state stability with respect to small inputs can guarantee global entrainment to small almost periodic inputs. In this way, global entrainment is shown in Lotka-Volterra systems with a Volterra-Lyapunov stable interaction matrix. All results can be extended to the uniformly recurrent case.
High-Mobility Ge-Doped $\beta$-Ga$_2$O$_3$ Growth on Sapphire by Low-Pressure Chemical Vapor Deposition
arXiv:2607.10908v1 Announce Type: cross Abstract: In this work, high-quality Ge-doped (-201) $\beta$-Ga$_2$O$_3$ thin films were heteroepitaxially grown on c-plane sapphire substrates with offcut angles of 0 deg, 2 deg, 6 deg, and 8 deg using low-pressure chemical vapor deposition (LPCVD). Increasing sapphire offcut promoted step-flow growth, resulting in improved terrace alignment, reduced surface roughness, and enhanced crystalline quality. Phase-pure monoclinic $\beta$-Ga$_2$O$_3$ with strong (-201) preferential orientation was confirmed by X-ray diffraction and Raman spectroscopy, while X-ray photoelectron spectroscopy revealed near-stoichiometric composition with an O/Ga ratio of 1.48. Electrical transport properties exhibited a strong dependence on substrate offcut angle, with room-temperature Hall mobility increasing from 15 to 117 cm$^2$/V s as the offcut angle increased from 0 deg to 6 deg, across carrier concentrations spanning $1.43 \times 10^{17}$ to $2.75 \times 10^{18}$ cm$^{-3}$. The 6 deg offcut sample achieved a room-temperature mobility of 117 cm$^2$/V s at a carrier concentration of $1.43 \times 10^{17}$ cm$^{-3}$ and a peak low-temperature mobility of 337 cm$^2$/V s at 128 K with a carrier concentration of $8.96 \times 10^{16}$ cm$^{-3}$, representing the highest reported room-temperature and low-temperature mobilities for Ge-doped $\beta$-Ga$_2$O$_3$ films grown on sapphire substrates. Carrier concentration and mobility data were analyzed using charge-neutrality and Boltzmann transport models incorporating donor activation together with polar optical phonon, ionized impurity, neutral impurity, acoustic deformation potential, and dislocation scattering mechanisms. The fitting revealed shallow donor activation energies of 12.5-19 meV, a deeper donor level at 80 meV, low acceptor compensation ($< 5 \times 10^{15}$ cm$^{-3}$), and threading dislocation densities on the order of $10^9$ cm$^{-2}$.
Quantum and Classical mechanics vs QFT
arXiv:2602.21243v2 Announce Type: replace Abstract: 15 years ago Dmitry Diakonov wrote the paper "Towards lattice-regularized Quantum Gravity", arXiv:1109.0091. In his approach, gravity with metric and tetrads arise from pre-geometric quantum fields leading to unusual dimensions of physical quantities. In particular, particle masses are dimensionless. We are trying to extend the Akama-Diakonov-Wetterich theory by introducing the Planck constants $\hbar$ and ${/\!\!h}=\hbar c$ as elements of the emergent metric. The inverse Planck constant $1/\hbar$ has the dimension of frequency, and, therefore, the mass $M$ of a particle, which has the dimension $\hbar\omega$, is dimensionless. In this extension, quantum mechanics emerges from the intrinsic quantum fields either in the symmetry breaking mechanism (GUT), or in the opposite mechanism of emergent symmetry in the low-energy corner (anti-GUT). In both cases, quantum mechanics (QM) serves as a bridge between the area of quantum fields (QFT) in the limit $1/\hbar \rightarrow 0$, and the area of classical physics (CM) in the limit $\hbar \rightarrow 0$. In the GUT scheme the inverse Planck constants, $1/\hbar$ and $1/{\\!\!h}$, play the role of the order parameter of the symmetry breaking phase transition from the pre-geometric QFT state to the QM state, in which the quantum mechanics emerges together with the space-time metric. In this phase transition, the integration over field variables in the QFT phase transforms to a path integral formulation of QM, which in turn yields the laws of classical mechanics in the limit $1/\hbar \rightarrow \infty$.
Probabilistic Wind Power Forecasting with Tree-Based Machine Learning and Weather Ensembles
arXiv:2602.13010v2 Announce Type: replace Abstract: Accurate production forecasts are essential for the integration of renewable energy sources into the power grid. This paper illustrates how to obtain probabilistic forecasts of wind power generation using gradient boosting trees and an ensemble of weather forecasts. To this end, we perform a comparative analysis across three state-of-the-art probabilistic prediction methods-conformalized quantile regression, natural gradient boosting and conditional diffusion models-all of which can be combined with tree-based machine learning. The methods are validated using four years of data for all Belgian offshore wind farms. We benchmark the models against the power curve and a calibrated wake model as well as a probabilistic method using stochastic variational Gaussian process regression. The tree-based models significantly reduce the mean absolute error in comparison to the deterministic baselines. Additionally, all three methods outperform the Gaussian process baseline in probabilistic skill, while two out of the three also improve point forecast accuracy. The conditional diffusion model attains the best performance, with improvements of 5% in mean absolute error and 12% in continuous rank probability score compared to the probabilistic baseline. Last, the results indicate an average improvement in point forecast accuracy of 17% by using an ensemble of weather forecasts instead of a single provider.
OmniX: Any-view and Any-time 4D Reconstruction via Feed-forward Trajectory Fields
arXiv:2607.10840v1 Announce Type: new Abstract: Previous feed-forward 4D reconstruction methods either predict per-frame static point clouds, ignoring foreground motion, or estimate point cloud trajectories while being limited to small camera motions. This restricts their ability to aggregate observations over time and reconstruct complete dynamic scenes under large viewpoint changes. To address this limitation, we propose OmniX, a feed-forward 4D reconstruction framework that predicts dense 3D point trajectories for every pixel from videos with large camera motion. OmniX decouples dynamic motion modeling from static geometry prediction and represents motion using a compact set of dynamic tokens. By leveraging the sparse and low-rank structure of 3D motion, these tokens generate trajectory fields for all pixels across all images while efficiently preserving global interactions. To facilitate training, we further build an automatic UE5-based 4D data engine and introduce a large-scale dataset containing 80K scenes and 1.28M multi-view videos with full geometric annotations. OmniX achieves state-of-the-art performance on dense 3D point trajectory prediction and 3D point tracking, while also demonstrating competitive results on video depth estimation and camera pose estimation.
From Sketch Prior to Trajectories: A Mission-Oriented Coordinated Navigation Framework for Indoor UAV Swarm
arXiv:2607.11386v1 Announce Type: new Abstract: UAV swarm for applications, such as indoor inspection, security patrol, and logistics delivery, are often mission-oriented rather than exploration-oriented. In these tasks, UAVs are required to visit task-relevant regions in a prescribed sequence, and such region-level mission information can often be obtained from pre-deployment sketch-map priors, such as floor plans, CAD layouts, or evacuation diagrams. Although these tasks are executed in three-dimensional space, UAVs usually fly within a specific altitude layer or a nearly fixed altitude range on each floor, making mission-level region transitions mainly governed by planar connectivity. Based on these observations, this paper proposes a mission-oriented coordinated navigation framework that exploits sketch-map priors for multi-UAV indoor operations. Onboard observations are used to perform topological alignment, and the aligned prior is fused with online observations to construct a mission-oriented traversability representation. A layered 2D--3D coordinated navigation framework is further developed, where 2D guided path planning generates mission-oriented guide paths and guide-driven 3D trajectory optimization produces dynamically feasible and collision-free trajectories. Simulation and real-world experiments validate the effectiveness of the proposed framework in structured multi-room indoor environments and further demonstrate its coordinated navigation capability under both communication-available and communication-loss conditions. Multi-floor simulation results show the scalability of the system to layered indoor structures.
Reliability Scaling Laws for Quantized Large Language Models
arXiv:2607.10855v1 Announce Type: new Abstract: Quantization is a powerful strategy to build capable and resource-efficient large language models (LLMs) by reducing the bitwidth of the parameters. While quantized LLMs achieve state-of-the-art performance on unperturbed inputs using standard predictive metrics, their performance on perturbed inputs, measured using reliability metrics, remains underexplored, despite its importance for reliable deployment. To address this gap, we first conduct a comprehensive reliability evaluation of quantized LLMs consisting of three key components: (1) Uncertainty: We assess the trustworthiness of LLMs quantized to 2, 3, 4, and 8 bits using six different quantization methods, employing established uncertainty metrics. (2) Calibration: We assess how well-calibrated the uncertainty estimates of quantized models are across model scales and bit precisions. (3) Robustness: We design character-level and word-level input perturbations to evaluate the reliability of quantized models under semantically-preserving variations in the inputs that arise in real-world applications. Second, we characterize how reliability scales with the total number of model bits. Our study reveals that while the performance scales monotonically with the total number of bits, the reliability scalings are nonlinear. A reliability peak occurs for 4-bit quantized models, indicating that quantizing moderately sized models offers the best reliability-efficiency trade-off. Additionally, our empirical findings reveal that quantization enhances the robustness of LLMs to natural input perturbations.
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
arXiv:2607.11523v1 Announce Type: new Abstract: When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches either wait passively for user queries or treat every detected event as requiring a response, without considering the user's history, current activity, or whether assistance would actually be welcome. We reframe proactive assistance as a context-dependent decision problem: the agent must not only perceive what is happening, but reason over accumulated temporal context to determine when and whether to intervene. To this end, we present Vinci2, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity. On the evaluation side, we present EgoServe, the first large-scale benchmark for proactive assistance in continuous egocentric video. EgoServe comprises over 3,000 service instances organized along 4 temporal memory horizons, ranging from immediate safety alerts to long-term habit coaching, across 10 service categories. On the modeling side, we propose EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives. At each timestep, EgoMemo performs retrieval-augmented reasoning to determine whether assistance is warranted and, if so, produces contextually grounded responses. Experiments demonstrate that EgoMemo establishes strong baselines on EgoServe while remaining competitive on existing egocentric benchmarks. Our benchmark and code are publicly available at \href{https://sitonggong.github.io/EgoServe-page/}{Vinci2}.
Implementation and representation of qudit multi-controlled unitaries and hypergraph states by N-body angular momentum couplings
arXiv:2506.21831v3 Announce Type: replace-cross Abstract: We construct a representation of qudit multi-controlled unitary operators in terms of N-body angular momentum interactions. The representation is particularly convenient for odd-dimensional systems, with interesting connections to the Pegg-Barnett phase formalism. We illustrate the main points in the special case of qutrits, where simplifications and connections to dipole-quadrupole and quadrupole-quadrupole interactions can be established. We describe the representation of the closely related set of qudit hypergraph states, identifying possible realizations and their main obstacles. Qutrit tripartite controlled unitaries are decomposed in terms of more familiar two-body angular momentum couplings, enabling their implementation in a variety of physical systems. We give then a concrete example of implementation of qutrit unitaries and hypergraph states in optical systems that employs single-photon sources, two-mode cross-Kerr interactions and linear optical operations. Moreover, we define a new set of states, called angular momentum hypergraph states, which are more directly related to the angular momentum representation.
Reliability-Aware CT-MRI Registration: A Quality Engineering Framework with Stability Analysis and Risk Classification
arXiv:2607.02585v2 Announce Type: replace Abstract: Multimodal CT-MRI registration is central to image-guided radiotherapy, surgical navigation, and diagnostic workflows, but most pipelines report only aggregate quality metrics without per-case reliability signals. We propose a reliability-aware framework that converts registration quality into Green/Yellow/Red risk categories using data-learned thresholds. CT images were registered to T1-weighted MRI using rigid and affine transformations on 90 paired slices from 18 patients across brain, abdominal, and neck anatomies. Reliability was assessed using Delta NMI, Delta SSIM, Dice overlap, registration stability, and inverse consistency error, combined into a single score R. Thresholds learned from training patients were applied unchanged to held-out test patients. Affine registration outperformed rigid registration on NMI and SSIM, yielding 44% Green classifications versus 33% for rigid. Reliability-filtered registrations improved the average alignment profile compared with unfiltered methods. Per-anatomy analysis showed substantial variation, with stronger reliability for abdominal registrations than brain registrations. Weight sensitivity analysis identified Dice overlap as the dominant reliability component. The proposed framework provides an interpretable quality-control layer for multimodal registration, while risk thresholds reflect statistical rather than clinical validation.
A Cheeger Inequality for Size-Specific Conductance
arXiv:2303.11452v3 Announce Type: replace Abstract: The $\mu$-conductance measure proposed by Lov\'asz and Simonovits is a size-specific conductance score that identifies the set with smallest conductance while disregarding those sets with volume smaller than a $\mu$ fraction of the whole graph. Using $\mu$-conductance enables us to study the network structures in new ways. In this manuscript we study a modified spectral cut for $\mu$-conductance that is a natural relaxation of the integer program of $\mu$-conductance and show that the optimum of this program has a two-sided Cheeger inequality with $\mu$-conductance.
Federated Topic Model and Model Pruning Based on Variational Autoencoder
arXiv:2311.00314v2 Announce Type: replace Abstract: Topic modeling has emerged as a valuable tool for discovering patterns and topics within large collections of documents. However, when cross-analysis involves multiple parties, data privacy becomes a critical concern. Federated topic modeling has been developed to address this issue, allowing multiple parties to jointly train models while protecting privacy. However, there are communication and performance challenges in the federated scenario. To solve these problems, this paper proposes a method to establish a federated topic model while ensuring the privacy of each node and uses neural network model pruning to accelerate the model. The client periodically sends cumulative neuron gradients and model weights to the server, and the server prunes the model. To address different requirements, two methods are proposed to determine the pruning rate. The first slowly prunes throughout training, which has limited acceleration during training but can ensure higher accuracy and significantly reduce inference time. The second quickly reaches the target pruning rate early in training and then continues training with a smaller model. This approach may lose more useful information but can complete training faster. Experimental results show that the proposed variational-autoencoder-based federated topic model pruning can greatly accelerate training while maintaining model performance.
Autonomous Close-Proximity Photovoltaic Panel Coating Using a Quadcopter
arXiv:2509.10979v3 Announce Type: replace Abstract: Photovoltaic (PV) panels are becoming increasingly widespread in the domain of renewable energy, and thus, small efficiency gains can have massive effects. Anti-reflective and self-cleaning coatings enhance panel performance but degrade over time, requiring periodic reapplication. Uncrewed Aerial Vehicles (UAVs) offer a flexible and autonomous way to apply protective coatings more often and at lower cost compared to traditional manual coating methods. We present a quadcopter-based system, equipped with a liquid dispersion mechanism, designed to automate such tasks. The localization stack only uses onboard sensors, relying on visual-inertial odometry and the relative position of the PV panel detected with respect to the quadcopter. The control relies on a model-based controller that accounts for the ground effect and the mass decrease of the quadcopter during liquid dispersion. We validate the autonomy capabilities of our system through extensive indoor and outdoor experiments.
Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion
arXiv:2506.22036v2 Announce Type: replace Abstract: With the increasing multimodal knowledge privatization requirements, multimodal knowledge graphs in different institutes are usually decentralized, lacking of effective collaboration system with both stronger reasoning ability and transmission safety guarantees. In this paper, we propose the Federated Multimodal Knowledge Graph Completion (FedMKGC) task, aiming at training over federated MKGs for better predicting the missing links in clients without sharing sensitive knowledge. We propose a framework named MMFeD3-HidE for addressing multimodal uncertain unavailability and multimodal client heterogeneity challenges of FedMKGC. (1) Inside the clients, our proposed Hyper-modal Imputation Diffusion Embedding model (HidE) recovers the complete multimodal distributions from incomplete entity embeddings constrained by available modalities. (2) Among clients, our proposed Multimodal FeDerated Dual Distillation (MMFeD3) transfers knowledge mutually between clients and the server with logit and feature distillation to improve both global convergence and semantic consistency. We propose a FedMKGC benchmark for a comprehensive evaluation, consisting of a general FedMKGC backbone named MMFedE, datasets with heterogeneous multimodal information, and three groups of constructed baselines. Experiments conducted on our benchmark validate the effectiveness, semantic consistency, and convergence robustness of MMFeD3-HidE.
Can Argus Judge Them All? Comparing VLMs Across Domains
arXiv:2507.01042v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) are increasingly used in industry VLM applications such as retrieval systems, content generation platforms, and decision-support workflows, where model selection is commonly guided by benchmark rankings. These rankings are largely determined by retrieval, captioning, and reasoning downstream tasks; however, models with similar task performance often show substantially different behavior across datasets. This creates a Capability-Reliability Gap between benchmark performance and observed model stability. We present ARGUS-EVAL, a capability-reliability-oriented evaluation framework for VLMs that characterizes model behavior through Benchmark Capability P(M), Cross-Dataset Consistency CDC(M), Robustness Retention RR(M), and Efficiency E(M). We evaluate CLIP, BLIP, LXMERT, Gemma-3-4B, and Qwen-2.5VL-3B-Instruct across retrieval, captioning, and reasoning downstream tasks. The results reveal notable differences between capability-oriented and reliability-oriented rankings. Qwen-2.5VL-3BInstruct achieves the strongest overall capability (R@1 = 82.7%, BLEU-4 = 47.2%, CIDEr = 141.6, CDC = 0.91), whereas CLIP records the lowest latency (31 ms) and memory footprint (0.9 GB).
Interaction Techniques that Encourage Longer Prompts Can Improve Psychological Ownership when Writing with AI
arXiv:2507.03670v3 Announce Type: replace Abstract: Writing longer prompts for an AI assistant to generate a story increases psychological ownership, a user's feeling that the writing belongs to them. To encourage users to write longer prompts, we evaluated two interaction techniques that modify the prompt entry interface of chat-based generative AI assistants: pressing and holding the prompt submission button, and continuously moving a slider up and down when submitting a short prompt. A within-subjects experiment investigated the effects of such techniques on prompt length and psychological ownership, and results showed that these techniques increased prompt length and led to higher psychological ownership than baseline techniques. A second experiment further augmented these techniques by showing AI-generated suggestions for how the prompts could be expanded. This further increased prompt length, but did not lead to improvements in psychological ownership. Our results show that simple interface modifications like these can elicit more writing from users and improve psychological ownership.
Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message
arXiv:2507.04673v2 Announce Type: replace Abstract: The rise of conversational interfaces has greatly enhanced LLM usability by leveraging dialogue history for sophisticated reasoning. However, this reliance introduces an unexplored attack surface. This paper introduces Trojan Horse Prompting, a novel jailbreak technique. Adversaries bypass safety mechanisms by forging the model's own past utterances within the conversational history provided to its API. A malicious payload is injected into a model-attributed message, followed by a benign user prompt to trigger harmful content generation. This vulnerability stems from Asymmetric Safety Alignment: models are extensively trained to refuse harmful user requests but lack comparable skepticism towards their own purported conversational history. This implicit trust in its "past" creates a high-impact vulnerability. Experimental validation on Google's Gemini-2.0-flash-preview-image-generation shows Trojan Horse Prompting achieves a significantly higher Attack Success Rate (ASR) than established user-turn jailbreaking methods. These findings reveal a fundamental flaw in modern conversational AI security, necessitating a paradigm shift from input-level filtering to robust, protocol-level validation of conversational context integrity.
Generalized and Unified Equivalences between Hardness and Pseudoentropy
arXiv:2507.05972v3 Announce Type: replace Abstract: Pseudoentropy characterizations give quantitatively precise formulations of the relationship between computational hardness and computational randomness. We prove a unified pseudoentropy characterization that generalizes and strengthens previous results in both uniform and nonuniform models of computation. Our characterization applies to a general family of entropy notions, including Shannon entropy and min-entropy as special cases. Moreover, the characterizations for these different entropy notions can be witnessed simultaneously by a single universal function, which captures both computational hardness and computational randomness. A key technical insight is that weight-restricted calibration, from the recent literature on algorithmic fairness, together with standard computational indistinguishability (known as multiaccuracy in the fairness literature), suffices for proving pseudoentropy characterizations for general entropy notions. To obtain this combination of properties, we prove an enhanced version of the Leakage Simulation Lemma (Jetchev and Pietrzak, 2014), which in turn extends the Complexity Theoretic-Regularity Lemma (Trevisan, Tulsiani, and Vadhan, 2009) from boolean functions to ones over a larger alphabet. Our Enhanced Regularity/Leakage-Simulation Lemma enables us to obtain an exponential improvement in the dependence on the alphabet size compared with the pseudoentropy characterizations of Casacuberta, Dwork, and Vadhan (2024), which are based on the stronger notion of multicalibration. We also show that this exponential dependence on the alphabet size is inevitable for multicalibration and even for the weaker notion of calibrated multiaccuracy.
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
arXiv:2507.11059v3 Announce Type: replace Abstract: The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, particularly the widely used SWE-bench dataset. Recent studies have uncovered severe data contamination issues, e.g., SWE-bench reports 32.67% of successful patches involve direct solution leakage and 31.08% pass due to inadequate test cases. We introduce SWE-MERA, a dynamic, continuously updated benchmark designed to address these fundamental challenges through an automated collection of real-world GitHub issues and rigorous quality validation. Our approach implements a reliable pipeline that ensures quality while minimizing contamination risks, resulting in approximately 10,000 potential tasks with 728 samples currently available. Evaluation using the Aider coding agent demonstrates strong discriminative power in state-of-the-art models. We report performance across a dozen recent LLMs evaluated on tasks collected between September 2024 and June 2025.
Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition
arXiv:2606.06065v4 Announce Type: replace Abstract: Second-language (L2) speech recognition often requires transcriptions of pronunciations and intended meanings. Multi-task learning (MTL) is a natural approach because it assumes that shared representations benefit both outputs. However, this paper shows that this assumption does not hold across Korean and English. MTL improves meaning but degrades surface transcription, especially in English, where the degradation scales with surface-meaning divergence measured by Levenshtein edit distance. Encoder analysis links these patterns to encoder-level entanglement, with Korean preserving disentangled representations while English produces nearly identical ones. Cross-output decoder analysis shows that the meaning dual-output decoder adapts with a unique representation, while the surface dual-output decoder remains constrained by the encoder. These findings motivate the design of MTL frameworks that mitigate encoder-level entanglement to reduce surface degradation in dual-output L2 automatic speech recognition.
Decompiling for Constant-Time Analysis
arXiv:2501.04183v4 Announce Type: replace Abstract: The CT programming discipline is commonly used to protect cryptographic libraries against side-channel attacks. However, it is hard to write CT code; moreover, compilers can introduce CT violations. Therefore, it is important to ensure that assembly code is CT. One approach is to show that source programs are CT, and that CT is preserved by compilation. In this paper, we explore the methodological soundness and scalability of the Decompile-then-Analyze (DtA) approach, a less conventional alternative that has been suggested in the broader setting of static analysis. Informally, the DtA approach uses decompilers a front-end for static analysis tools. As a motivation for our study, we show that current decompilers eliminate CT vulnerabilities before CT analysis, leading to non-CT programs being accepted as CT. Independently, we provide *constructed* examples of non-CT, exploitable, programs that are accepted by two popular CT analysis tools; in both cases the culprit are program transformations that are used internally prior to CT analysis and eliminate CT violations. While our examples do not invalidate the general approach of these tools, they emphasize the need for studying the DtA approach. On the methodological side, we define the notion of *CT transparency*. Informally, a program transformation is CT transparent if does not eliminate nor introduce CT violations. We also provide general methods for proving that a transformation is CT transparent, and show that several transformations of interest are transparent. We also sketch an extension of CT transparency to speculative CT, which is used by cryptographic software as a protection against Spectre attacks. On the practical side, we build a CT-transparent version of the popular LLVM-based decompiler RETDEC, and combine it with CT-LLVM, an existing CT verification...
Clapping propulsion and thin vortex rings: a computational study of vortex dynamics, energy equivalence, and core potential energy
arXiv:2507.11491v2 Announce Type: replace Abstract: We report a computational study of clapping propulsion using two thin rigid plates forming a 60-degree interplate cavity that generates a thrust-producing jet during closure. Plate kinematics are prescribed from experiments for two cases: dynamic and stationary, with forward motion constrained in the latter. The computations show that interplate pressure is higher in the stationary case compared to that in the dynamic case, resulting in differences in the thrust produced and in the evolution of wake vortices, with the stationary case forming triangular and Omega-shaped loops, while the dynamic case forms an elliptical loop for each plate. We examine the energy budget in the post-clapping phase, when the vortices are fully formed. The energy consists of kinetic energy and a component associated with vortex formation, whose sum approximately matches the work done on the fluid. This extra term, which we call the core potential energy, is found to be equal to the integral of pressure over the core volume of the vortex. This component is also checked using separate axisymmetric vortex ring simulations, where the kinetic energy is about 60% of the injected slug energy, and the remaining part is the core potential energy. Sullivan et al.(2008) had commented on this deficit for vortex rings and hypothesized the existence of a potential energy associated with the vortex structure.
Henri Poincare Saint Louis Lecture of 1904: Publication, Dissemination, and Historiographical Implications
arXiv:2603.23410v3 Announce Type: replace Abstract: Henri Poincare Saint Louis lecture of 1904 occupies an important place in the prehistory of relativity. In it, Poincare formulated the principle of relativity in general terms and presented it as one of the guiding principles of mathematical physics, together with the principles of least action and energy conservation. This article reconstructs the early publication and international dissemination of the lecture before the end of June 1905, through La Revue des idees, the Bulletin des sciences math\'ematiques, The Monist, and La valeur de la science. Drawing on library records, accession data, booksellers' advertisements, press notices, and correspondence, it shows that Poincare text circulated rapidly through scholarly, commercial, and institutional channels in Europe and North America. The significance of this circulation is not merely bibliographical. It bears directly on the documentary landscape within which Einstein 1905 relativity paper should be historically situated. The early availability of Poincare lecture and of La valeur de la science shifts attention toward the concrete conditions under which texts, concepts, and problems circulated in the months and weeks preceding Einstein's June 1905 paper. Evidence from Einstein Bern milieu further supports this view. Joseph Sauter later testimony and the possibility of redating the Habicht letter to 1 June 1905, on the basis of several converging chronological and documentary arguments, suggest that the intellectual environment of 1905 was denser than simplified narratives of solitary discovery imply. Availability is not influence; but without reconstructing availability, the historical problem of Einstein relation to Poincar\'e and Lorentz remains ill posed.