Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Optimization Is Not All You Need
arXiv:2607.11977v2 Announce Type: replace Abstract: In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produced their more fluent successors is usually regarded as an engineering achievement; we read it instead as the newest expression of optimization culture: the conviction, older than the technology, that measurable improvement along predefined axes exhausts the question of value. Tracing that conviction through the stack-pretraining, decoding, preference tuning, benchmarking, interface-and back through its genealogy in the audit society, we arrive at the limit: an optimization procedure can measure how improbable a piece of generated text is; it cannot tell whether that unlikelihood is error or invention. A procedure that cannot make that distinction has nonetheless, within half a decade, assumed the authority to set the protocols of legitimate language. Held for centuries by academies and schoolrooms, grammars and examiners, this authority has been given over to loss functions, reward models, benchmarks, and system prompts: an apparatus that executes the office of judgment with no capacity for judging.
Inversion of the Multiplicative Matrix Compound Operator
arXiv:2605.27682v3 Announce Type: replace-cross Abstract: We study the problem of determining a matrix whose $k$th multiplicative compound, with $k > 1$, is a prescribed matrix $M$. The cardinality of the set of matrices whose $k$th multiplicative compound equals $M$ is characterized in terms of $\rank(M)$. On the one hand, if $\rank(M)\le 1$, it is shown that there exist infinitely many such matrices for which a complete characterization is determined. On the other hand, if $\rank(M)>1$, then there exists a unique matrix -- up to an overall sign -- whose compound is $M$. An algorithm for finding a matrix whose compound equals $M$ is detailed, and its time complexity is analyzed.
CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference
arXiv:2607.13649v1 Announce Type: new Abstract: LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge devices. Weight precision plays a critical role in balancing inference accuracy, throughput, and energy consumption, while modern LLM workloads exhibit pronounced heterogeneity and tolerance that favors adaptive precision execution. This paper presents CIMERA, a reconfigurable-precision LLM inference accelerator that integrates compute-in-interconnect and memory to mitigate the memory wall and enable precision-aware execution. Compared to Nvidia H100, CIMERA delivers up to $25\times$ and $10\times$ higher energy efficiency for 1B and 13B models, respectively.
Calming traffic in unsignalized bidirectional streets by introducing one-lane bottlenecks
arXiv:2607.13650v1 Announce Type: new Abstract: A common strategy to calm traffic in an unsignalized, bidirectional street consists in narrowing the street to a single lane at one or more locations. These locations become bottlenecks. Two policies are commonly used there: first-in-first-out (FIFO), and directional priority (DP), which means one of the directions has priority. This paper examines the performance of these two policies on streets with one or more bottlenecks. Two measures of performance are considered: the expected delay (evaluated for typical situations with low demand) and the system capacity. The latter is defined as the pair of flows (d1, d2) that depart the street in the two opposite directions when the system is oversaturated with queues. For a street with a single bottleneck, the capacity under both FIFO and DP is not unique: capacity in one direction decreases with a higher flow in the other direction. As expected, we find that arrival flow pairs (a1, a2) below this boundary can get through the system without generating queues. Those above cannot. For a single bottleneck shows that FIFO yields lower expected delay than DP under low demand. For high demand, DP's capacity exceeds FIFO's when the flows are balanced, not lower otherwise. For multiple bottlenecks we consider low-demand delay, and capacity. In low demand, the expected total delay on the street is that of a single bottleneck multiplied by the number of bottlenecks. The multi-bottleneck street capacity under FIFO turns out to be the same as for a single bottleneck. For DP, traffic near capacity is complex and capacity points are numerous and disjointed. Formulas are derived for their values. Higher capacities are typically obtained for longer streets. Curiously, these capacity points are often above the capacity boundary of a street with a single FIFO or DP bottleneck, indicating that an extra bottleneck can increase capacity.
Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting
arXiv:2607.13808v1 Announce Type: new Abstract: Recent extensions of 3D Gaussian Splatting (3DGS) capture fine color details using hash-grid-based appearance parameterization but incur high computational cost during fragment rendering. We introduce a decoupled radiance representation that models low-frequency geometry and view dependent appearance features with 2D surfels while representing high-frequency textures via a view-independent spatial hash grid that is baked into a compact texture atlas. By including sparsity-enhancing optimizations that penalize semi-transparency and per-primitive falloff, our method aggressively prunes insignificant surfels and achieves significantly faster and sparser reconstructions than prior work. Exploiting geometric sparsity and efficient GPU texture mapping, our approach achieves up to a fivefold speedup over 3DGS while preserving state-of-the-art visual fidelity, enabling real-time 4K rendering at 60 FPS on consumer hardware.
Adaptive space-time BEM for the heat equation with Neumann boundary conditions
arXiv:2607.13578v1 Announce Type: new Abstract: We consider the space-time boundary element method (BEM) for the heat equation with prescribed initial and Neumann data. We propose a weighted-residual a posteriori error estimator that is an upper bound for the unknown BEM error. The possibly locally refined meshes are assumed to be parabolically scaled prismatic, i.e., their elements are tensor-products $J\times K$ of elements in time $J$ and space $K$ with $|J| \eqsim \text{diam}(K)^2$. In the considered numerical experiments on two-dimensional domains in space, an adaptive algorithm steered by the derived estimator yields significantly faster convergence compared to uniform refinement, achieving near-optimal rates even in the presence of strong singularities.
Accurate Solvation Properties in supercritical CO$_2$ with Molecular Density Functional Theory
arXiv:2607.13777v1 Announce Type: new Abstract: Supercritical CO$_2$ is a highly efficient solvent for the development of more environmentally benign chemical processes. It is crucial to predict its solvation properties -- the solvation free energy and the solvation structure -- both accurately and at low computational cost. We show here that classical density functional theory (cDFT) can reproduce the solvation properties obtained from conventional molecular simulations, while requiring a computational effort that is several orders of magnitude lower. This excellent agreement is achieved using a molecular cDFT formalism based on a density that depends on both the positions and orientations of CO$_2$ molecules in the vicinity of the solute. We further examine several levels of approximation for the excess free-energy functional in cDFT and demonstrate that the homogeneous reference fluid approximation is sufficient to recover the molecular dynamics (MD) benchmark results. These findings open the way to extending molecular cDFT to other thermodynamic conditions.
Auctions with Contract Design
arXiv:2607.13795v1 Announce Type: new Abstract: We consider a new auction model where the bidders' utilities and the auctioneer's revenue depend on a quality factor of the transaction determined by costly and strategic investments of the bidders. Applications of our model include ad auctions, government concessions and crowdsourcing contests. Crucially, these quality-enhancing efforts made by the bidders are often sunk costs incurred prior to the allocation, creating a fundamental moral hazard problem where the risk of losing the auction discourages investments. In this paper, we study the design of revenue-maximizing contracts integrated into auctions: the auctioneer commits to a transfer rule that rewards the winner for the ex-post realized quality of the transaction to incentivize higher effort. Our new framework is a natural generalization of both the auction theory and the principal-agent model. We consider both the second-price and the first-price auctions. We show that natural symmetric Bayes Nash equilibria exist in both auctions. Assuming these natural equilibria are played by the bidders and the number of bidders is large, we study linear contracts and derive the optimal reward factor of the transfer rule that maximizes the auctioneer's revenue. As the main result, we show that the optimal reward factor converges to the auctioneer's marginal benefit from the quality, as the number of bidders grows. That is, it is optimal for the auctioneer to fully pass through the quality value to the winner. This observation is largely independent of the auction rule used: we derive a revenue equivalence theorem showing that the revenue remains the same as long as symmetric Bayes Nash equilibria exist. Lastly, by quantitatively comparing with the standard auctions where no quality reward is used, we show that the use of contracts effectively improves the revenue by incentivizing high investments from the bidders.
Online Random Sampling with Real Probabilities
arXiv:2607.13828v1 Announce Type: new Abstract: We develop an efficient online algorithm to sample a sequence of discrete random variables using an entropy source of i.i.d. fair coin flips, in a standard model of real computation where real-valued probabilities are represented by rational approximations. For any sequence $F_1, F_2, \dots$ of probability distributions, our sampler generates $n$ outputs $X_1 \sim F_1, \dots, X_n \sim F_n$ using at most $\mathbb{E}\left[H(F_1) +\dots + H(F_n)\right] + O(\log n)$ coin flips in expectation while carrying $O(\log n)$ bits of persistent space, where $H$ is the Shannon entropy. Under standard assumptions, we prove that the space used by our sampler to achieve this information-theoretically optimal entropy rate is asymptotically optimal. The key idea is to replace the global arithmetic-decoding sampling scheme of Han and Hoshi (1997) with a local discrete uniform state, yielding an exponential reduction in space for a given entropy loss. Our approach applies to distributions with irrational probabilities and countably infinite supports, generalizing recent randomness-recycling methods beyond finite rational distributions with bounded denominator.
CAS I: A Geometric Coding Theorem
arXiv:2607.13796v1 Announce Type: new Abstract: This paper establishes a direct analogue of the classical Coding Theorem in the setting of symmetry groups. We consider computable bijections on the set of binary strings, called symmetries and define the symmetry prior of a string as the probability that a randomly chosen symmetry from a given group has the string as its unique fixed point. We show that for any fix-retractable symmetry group, a group admitting a computable section that selects an isolating symmetry for every string, the symmetry prior is a universal lower semi-computable semi-measure. In this case, the Geometric Coding Theorem holds. We also develop a Galois connection between subgroups of G and subsets of binary strings, characterizing closed points and maximal closed subgroups, and explore the join-semilattice of dense subgroups. Our results unify algorithmic information theory with group theory and provide a framework for studying symmetry-induced complexity measures. This paper is the first in a series on Computational Algorithmic Statistics (CAS).
The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model
arXiv:2607.13660v1 Announce Type: new Abstract: Contrastive Language-Image Pretraining (CLIP) representations form a semantic embedding space governed by cosine similarity, reflecting an intrinsic hyperspherical geometry. However, existing probabilistic interpretations typically rely on Gaussian assumptions, which fail to capture this directional and multimodal structure. We propose a principled density model for the CLIP latent space based on Mixtures of von Mises-Fisher (MovMF) distributions defined on the unit hypersphere. Using the Expectation-Maximization (EM) algorithm, we efficiently learn a probabilistic model in which each mixture component corresponds to a coherent semantic concept. This formulation yields a closed-form likelihood naturally aligned with hyperspherical geometry, enabling accurate and interpretable density estimation. Empirically, our model significantly improves long-tailed and out-of-distribution detection and provides a natural semantic decomposition, representing each embedding as a sparse probabilistic combination of interpretable concepts. These results suggest that CLIP latent space is more faithfully characterized as a hyperspherical semantic mixture rather than an isotropic Gaussian, establishing a simple and geometrically consistent probabilistic framework for modeling and understanding multimodal representations. Project page is available at https://xiaoyuzhizi.github.io/movmf-clip/.
RainDancer: RGB-Event Video Deraining with Rain-Oriented Spiking Dynamics
arXiv:2607.13802v1 Announce Type: new Abstract: Video deraining aims to recover clean visual content from rainy videos for reliable perception under adverse weather. Existing methods mainly rely on RGB sequences and temporal redundancy, but RGB-only restoration remains ambiguous in dynamic rainy scenes, where rain streaks, textures, boundaries, motion, and occlusions may share similar visual patterns. Event cameras provide complementary motion-sensitive cues with high temporal resolution, but event streams also contain sensor noise and background-triggered responses, so direct RGB-Event fusion may introduce cross-modal interference. To address this issue, we propose RainDancer, a progressive RGB-Event video deraining framework based on a decompose-before-interact paradigm. The core idea is to separate rain and background components within each modality before cross-modal interaction. In the RGB branch, frame features are progressively decomposed into rain and background representations. In the event branch, a rain-oriented spiking neural network module captures sparse and bursty event dynamics associated with rain motion. Component-level fusion is then performed between semantically aligned representations for structure preservation and rain suppression. We further introduce event-domain supervision to regularize sparse event reconstruction, structural consistency, and gradient orientation. Experiments on synthetic and real RGB-Event video deraining datasets demonstrate superior quantitative performance, visual quality, and downstream perception robustness. Code is available at https://github.com/AE86-plus/RainDancer.
Filon Methods for Highly Oscillatory Controlled Quantum Systems
arXiv:2607.13580v1 Announce Type: cross Abstract: Fast and accurate classical simulation of quantum systems is a central challenge in the design and control of quantum computers, but the highly oscillatory dynamics of these systems severely limit the efficiency of standard numerical methods. To address this, we adapt Filon quadrature for oscillatory integrals into two numerical methods, called Filon and Controlled Filon, for solving linear systems of ODEs with highly oscillatory solutions. We tailor both methods for efficient implementation in controlled quantum systems, and the Controlled Filon method additionally accounts for the oscillatory structure of the control pulses. We show by numerical experiments that these methods significantly reduce the computational cost of accurately simulating systems of superconducting transmon qubits by decreasing the number of timesteps needed to reach a given level of precision, with only a modest increase in the cost per timestep. For a realistic simulation of the dynamics of a CNOT gate, the Controlled Filon method is the most efficient method tested at every target accuracy, outperforming the best Hermite method by up to 6x and the Hermite method of the same order by up to 500x.
Uniform Approximation of Functions with Asymmetric Growth and Decay by Deep Weighted Polynomials
arXiv:2506.21306v2 Announce Type: replace Abstract: Functions that grow without bound on one side of the real line and decay to zero on the other cannot be approximated uniformly by ordinary polynomials on unbounded domains. Motivated by classical weighted polynomial approximation, we introduce a class of one-sided weighted \emph{deep} (composite) polynomial approximants for such asymmetric targets. The weight suppresses polynomial growth on the decaying side, while the composite polynomial remains free to capture growth on the other side. We prove that this mechanism reduces the half-line approximation problem to approximation on a compact interval whose length grows slowly with the degree, and we establish density and existence of best approximants in the appropriate closure of the model class. For computation, we first formulate the method as a trainable computational graph for \emph{deep} weighted polynomial approximation. However, direct end-to-end optimization becomes increasingly ill-conditioned at high composite degree and can suffer from local minima. To address this, we introduce a fine-tuning procedure in which a fixed inner composition of monotone polynomial self-maps supplies the effective degree, while only the outer polynomial and weight parameters are trained; the outer fit reduces to a linear program. Numerical experiments on Black--Scholes option-pricing functions show that the resulting fine-tuned weighted \emph{deep} polynomial achieves smaller uniform and \(L_2\) errors than matched-budget polynomial baselines and resolves the decaying tail to machine precision.
TMallGS: Scaling Unified Feature and Sequence Modeling for Generative E-commerce Search
arXiv:2607.13398v1 Announce Type: new Abstract: In industrial search and ranking systems, Click-Through Rate (CTR) prediction is shifting from traditional Deep Learning Recommendation Models (DLRM) toward unified, compute-intensive Transformer architectures. This transition is driven by the need to improve Model FLOPs Utilization (MFU) and achieve predictable gains through scaling laws. However, existing approaches such as OneTrans and Climber often adopt an all-in-tokenization strategy when adapting Large Language Model (LLM) architectures, overlooking the heterogeneous nature of ranking features. We propose TmallGS, a scalable ranking architecture for Tmall search. TmallGS includes five key components: (1) Hierarchical Distribution-Calibrated Tokenization, which combines Field-wise Saliency Reweighting (FSR) and Distribution-Calibrated Projection (DCP) to map diverse features into optimized subspaces; (2) a Field-Adaptive Gated Transformer Backbone with per-field QKV projections and noise-adaptive gating for refined semantic interaction; (3) Decoupled FiLM Late Fusion to preserve explicit high-frequency signals; (4) a Context-Aware Bias Net to decouple systemic bias from user intent; and (5) Error-Aware Progressive Training with dynamically weighted losses for robust learning. Extensive offline experiments and online A/B tests on Tmall Search show that TmallGS improves training throughput and achieves substantial gains in UCTCVR and GMV.
Designing Safety-Constrained LLM Systems for Public Health Information Access
arXiv:2607.13038v1 Announce Type: new Abstract: We present the design and implementation of a safety constrained large language model (LLM) system for public health information access, focusing on maternal and child health (MCH) resource navigation. While LLM based systems offer flexible and natural interfaces for information retrieval, their deployment in healthcare contexts introduces risks related to safety, trust, and uncontrolled generation. This work explores practical design patterns for constraining LLM behavior in safety critical environments. We introduce a multi-layered architecture that integrates domain-restricted retrieval augmented generation (RAG), strict boundary enforcement to prevent medical advice, anonymous multiuser session management, and comprehensive audit logging for monitoring and compliance. A key aspect of the design is a controlled data pipeline that grounds all responses in curated public health resources, avoiding reliance on the model pretrained medical knowledge. We implement the system in a real world public health setting and conduct scenario-based validation across in scope, out of scope, and emergency queries. Results show consistent enforcement of safety constraints, reliable resource grounding, and stable system performance, with an average response time of 5.3 seconds. Beyond the specific application, we discuss design trade offs and lessons learned in balancing safety, usability, and system flexibility. Our findings provide practical guidance for deploying LLM based systems in healthcare and other domains where strict information boundaries and accountability are required.
DeepStress: Stress-Testing Deep Search Agents
arXiv:2607.13920v1 Announce Type: new Abstract: While search agents demonstrate impressive capabilities in multi-step question answering, their robustness to poor-quality evidence remains under-explored. This phenomenon occurs rarely in realistic benchmarks but can lead to dramatic failure in real life applications. Therefore in this study we propose DeepStress, a stress testing framework that controls the frequency of challenging evidence by replacing the retrieval module of search agents with a controlled synthetic environment. We use this framework to control three dimensions that can affect document reliability: trustworthiness, relevance, and factuality. Testing several search agents on HotpotQA and BrowseCompPlus, we demonstrate that agents exhibit substantial differences in their ability to handle unreliable information and propose new metrics that better document systems outcomes as well as the interactions between conflicting parametric and retrieved knowledge.
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation
arXiv:2607.13124v1 Announce Type: new Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires. Two observations trace this gap. First, greedy \textsc{pass}@$1$ nearly vanishes after compression, yet \textsc{pass}@$k$ recovers substantially under repeated sampling: useful generations are demoted, not erased. Second, the recoverable regime fails mainly through suffix repetition. Recovery should therefore train on the compressed model's own on-policy states with dense token-level supervision, which On-Policy Distillation (OPD) provides by reusing the pre-compression model as a frozen teacher. However, long on-policy rollouts spend early recovery budget on low-information repetitive suffixes, delaying loss descent. To mitigate this waste, we propose \textbf{\shortopd}, a short-to-long OPD schedule that detects teacher-confirmed repetitive suffixes, treats the surviving prefix as each rollout's effective length, and allocates future rollout budgets to the effective lengths the policy can currently use. Across math, code, and open-ended generation, \shortopd\ raises the compressed model's score to about $9\times$ its unrecovered value and $1.6$--$4.4\times$ standard recovery recipes (SFT w/o KD, KD, and SeqKD), and it matches a fixed $8192$-token rollout horizon within two points using a quarter of the training time ($8.5$ vs.\ $35.9$ hours) and $71\%$ fewer rollout tokens. We hope this recipe helps move structured pruning beyond marginal gains on perplexity and multiple-choice benchmarks, a step closer to deployment-ready generation quality.
Microscopic constitutive theory of stress overshoot, yielding, and strain hardening in amorphous materials
arXiv:2607.13734v1 Announce Type: cross Abstract: We develop a microscopic constitutive theory for the nonlinear deformation of metallic and polymer glasses based on nonaffine elasticity coupled to irreversible many-body relaxation. The theory predicts the full stress--strain response, from linear elasticity through stress overshoot and yielding to steady plastic flow. We show that stress overshoot originates from the competition between a nonaffine elastic instability induced by strain-driven loss of mechanical connectivity at the atomic/molecular level, and viscous dissipation associated with structural relaxation. For polymer glasses, finite chain extensibility naturally accounts for strain hardening at large deformation. The stretched-exponential relaxation exponent is obtained independently from stress or modulus relaxation measurements and provides the primary dynamical input to the theory. Using a small set of physically meaningful parameters, the model quantitatively reproduces experimental stress--strain curves for metallic glasses, polycarbonate, PMMA, and epoxy resins over a broad range of strain rates. These results establish a unified microscopic framework linking relaxation dynamics, yielding, plastic flow, and strain hardening in amorphous solids.
Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning
arXiv:2607.13826v1 Announce Type: new Abstract: Accurate determination of pancreatic ductal adenocarcinoma (PDAC) resectability relies on evaluating how the tumor interacts with major peripancreatic vessels on CT imaging, yet expert assessment often shows substantial variability. We introduce a fully automated multimodal deep learning framework that jointly analyzes 3D contrast enhanced CT and structured clinical information to classify patients into the three National Comprehensive Cancer Network (NCCN) resectability categories (upfront resectable, borderline resectable, locally advanced). The approach uses a Swin-UNETR backbone to obtain anatomy aware image representations through auxiliary segmentation of pancreas, tumor, and vascular structures. These features are fused with a compact clinical embedding derived from 17 routinely collected variables and processed by a lightweight classification head. Model training is guided by a dynamic multitask objective that adapts the balance between segmentation and classification based on current tumor Dice performance, promoting feature representations that remain both anatomically informed and discriminative.
LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents
arXiv:2607.13041v1 Announce Type: new Abstract: Large Language Model (LLM) based AI educational content generation systems are increasingly being developed, yet no standardised benchmark exists to systematically evaluate them. This study introduces LessonBench-V1, a benchmark dataset comprising 647 human-written lessons paired with LLM-based reverse-engineered lesson plans across 240 STEM topics spanning mathematics, physics, chemistry, and computer science. The lessons are drawn from 97 trusted open sources, including LibreTexts, Brilliant.org and GeeksForGeeks. Each lesson plan is human-reviewed and produced through a pedagogically grounded methodology that synthesises Bloom's Taxonomy, Gagn\'e's Events, Merrill's First Principles, and the 5E Instructional Model. The lesson plans capture 3,620 learning objectives with pedagogical metadata, enabling systematic, reproducible evaluation of lesson-generation AI agents and supporting further research. The study further proposes a three-dimensional evaluation pipeline for use with the dataset.
Neural-network-based material identification in photon-counting computed tomography using region-of-interest spectral features
arXiv:2607.13835v1 Announce Type: new Abstract: This work investigates the feasibility of neural-network-based material identification based on photon-counting computed tomography (PCCT) data. The input data were spectral vectors extracted from regions of interest (ROIs) in reconstructed tomographic slices of a phantom containing La-, Nd-, and Gd-based samples, as well as air, water, bone, and polymethyl methacrylate (PMMA). The original tomographic slices were acquired at 12 detector threshold settings (THL = 45, 55, 65, 75, 85, 95, 105, 115, 125, 135, 145, 165). For model training, the spectra were interpolated onto a uniform grid from THL 45 to THL 165 with a step of one THL unit. Large ROIs, small ROIs, and their combined dataset were considered. To avoid an overestimated performance caused by correlated ROIs extracted from the same tomographic slice, the data were split into training, validation, and test subsets at the slice level.
Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation
arXiv:2607.13987v1 Announce Type: new Abstract: Reusable skills are becoming a fundamental building block of Large Language Model (LLM) agents, enabling capabilities to be packaged, shared, and reused across diverse applications. However, existing security research primarily focuses on prompt injection and runtime execution, leaving security risks throughout the broader skill lifecycle largely unexplored. In this paper, we present SkillSec-Eval, a lifecycle-aware framework for systematically evaluating the security of reusable agent skills. We first characterize the skill lifecycle and develop a threat taxonomy spanning repository admission, semantic retrieval, planner selection, execution, and skill evolution. We then instantiate this taxonomy in SkillSec-Eval and conduct a comprehensive empirical evaluation using a repository of 327 real-world skills. Our study demonstrates that vulnerabilities arise at multiple lifecycle stages beyond execution, highlighting the need for lifecycle-aware security analysis of reusable agent skills.
Is the Statistical Advantage Worth the Cost? An Empirical Comparison of KANs and MLPs for Structured Data Classification
arXiv:2607.13413v1 Announce Type: new Abstract: This study presents an empirical benchmarking comparison between Kolmogorov-Arnold Networks (KANs) and Multi-Layer Perceptrons (MLPs) on structured tabular classification tasks. Motivated by the growing interest in KANs as an alternative function-approximating architecture, we evaluate their out-of-the-box performance on twelve publicly available datasets spanning binary, multiclass, multilabel, and ordinal problems. Both models were trained under standardized preprocessing, architecture, and fixed hyperparameter settings, with performance assessed using test accuracy and F1-Score, paired hypothesis testing, and effect size analysis. Results show that KANs statistically outperform MLPs in binary and multiclass domains and achieve a significant aggregate advantage across all datasets. However, the observed medium effect size (d = -0.46) raises an important cost-benefit consideration: while KANs offer superior generalization through adaptive spline-based mappings, this advantage comes with substantially higher parameter and computational complexity relative to the MLP baseline. These findings suggest KANs are the preferred choice for high-precision applications, while MLPs remain a robust and efficient option for resource-constrained environments. Future work should extend this analysis to additional data modalities to further refine these architectural selection criteria.
Gaussian FSBP operators: Comparison and application to numerical methods for hyperbolic conservation laws
arXiv:2607.13224v1 Announce Type: new Abstract: Function-space summation-by-parts (FSBP) operators enable conservative and energy-stable numerical methods for hyperbolic conservation laws based on general, non-polynomial approximation spaces. Recent works show that using generalized Gaussian quadrature significantly reduces the number of grid points required compared to existing constructions that have mostly focused on equidistant grids. In this paper, we compare open and closed FSBP operators constructed with generalized Gaussian quadratures and apply them to numerically solve hyperbolic conservation laws. Furthermore, to support open node distributions, we extend the FSBP framework by introducing function-space exact extrapolation operators and operationalize them in numerical schemes for solving hyperbolic conservation laws. Our numerical experiments include the one-dimensional linear advection, non-viscous Burgers, and compressible Euler equations of gas dynamics. We observe that applying FSBP operators in numerical schemes can improve efficiency and accuracy. Notably, we demonstrate these advantages in more challenging time-dependent settings compared to other recent works on Gaussian FSBP operators.