Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
arXiv:2508.10123v3 Announce Type: replace Abstract: Advanced reasoning in LLMs on challenging domains like mathematical reasoning can be tackled using verifiable rewards based reinforced fine-tuning (ReFT). In standard ReFT frameworks, a behavior model generates multiple completions with answers per problem, for the answer to be then scored by a reward function. While such RL post-training methods demonstrate significant performance improvements across challenging reasoning domains, the computational cost of generating completions during training with multiple inference steps makes the training cost non-trivial. To address this, we draw inspiration from off-policy RL, and speculative decoding to introduce a novel ReFT framework, dubbed Nested-ReFT, where a subset of layers of the target model acts as the behavior model to generate off-policy completions during training. The behavior model configured with dynamic layer skipping per batch during training decreases the inference cost compared to the standard ReFT frameworks. Our theoretical analysis shows that Nested-ReFT yields unbiased gradient estimates with controlled variance. Our empirical analysis demonstrates improved computational efficiency measured as tokens/sec across multiple math reasoning benchmarks and model sizes. Additionally, we explore three variants of bias mitigation to minimize the off-policyness in the gradient updates that allows for maintaining performance that matches the baseline ReFT performance.
Graph Construction and Matching for Imperative Programs using Neural and Structural Methods
arXiv:2604.26578v3 Announce Type: replace Abstract: Reusing verification artefacts requires identifying structural and semantic similarities across programs and their specifications. In this paper, we focus on graph construction as a foundational step toward this goal. We present a pipeline that converts imperative programs and their annotations into typed, attributed graphs. Our experiments cover datasets including C with ACSL, Java with JML, and Dafny programs. The pipeline integrates abstract syntax tree parsing with semantic embeddings derived from models such as SentenceTransformer and CodeBERT. This enables the generation of graph representations that capture both structural relationships and semantic context. Our results show that consistent graph representations can be constructed across different languages and annotation styles. This work provides a practical basis for future steps in semantic enrichment and approximate graph matching for scalable verification artefact reuse.
Learning Tactile-Aware Quadrupedal Loco-Manipulation Policies
arXiv:2604.27224v3 Announce Type: replace Abstract: Quadrupedal loco-manipulation is commonly built on visual perception and proprioception. Yet reliable contact-rich manipulation remains difficult: vision and proprioception alone cannot resolve uncertain, evolving interactions with the environment. Tactile sensing offers direct contact observability, but scalable tactile-aware learning framework for quadrupedal loco-manipulation is still underexplored. In this paper, we present a tactile-aware loco-manipulation policy learning pipeline with a hierarchical structure. Our approach has two key components. First, we leverage real-world human demonstrations to train a tactile-conditioned visuotactile high-level policy. This policy predicts not only end-effector trajectories for manipulation, but also the evolving tactile interaction cues that characterize how contact should develop over time. Second, we perform large-scale reinforcement learning in simulation to learn a tactile-aware whole-body control policy that tracks diverse commanded trajectories and tactile interaction cues, and transfers zero-shot to the real world. Together, these components enable coordinated locomotion and manipulation under contact-rich scenarios. We evaluate the system on real-world contact-rich tasks, including in-hand reorientation with insertion, valve tightening, and delicate object manipulation. Compared to vision-only and visuotactile baselines, our method improves performance by 28.54% on average across these tasks.
Requential Coding: Pushing the Limits of Model Compression with Self-Generated Training Data
arXiv:2607.11883v1 Announce Type: new Abstract: Compression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization. Large neural networks may learn functions far simpler than their parameter counts suggest, but it is challenging to construct codes that realize this simplicity. Parameter-based methods such as quantization produce code lengths that scale with model size, insensitive to how much information the parameters store. Prequential coding bypasses this issue by compressing the training trajectory, but codes the exact data sequence regardless of how much the model learns, yielding large codes when the data has high entropy. We introduce requential coding, where a teacher model selects training samples drawn from the student's own distribution. The student's code records only these selections, which cost bits only where teacher and student disagree. The resulting code length is independent of parameter count and data entropy, and often orders of magnitude shorter than the prequential counterpart, with an advantage that grows with scale. This compression sheds light on phenomena inaccessible to prior compressors. Holding loss fixed, larger models and ensembles compress to much smaller sizes despite more parameters. Plugged into a PAC-Bayes bound, the requential code yields state-of-the-art generalization guarantees for billion-parameter LLMs, outperforming bounds built on aggressive post-training quantization even granted zero error. The bound tightens with scale in the compute-optimal regime, as models become increasingly compressible relative to dataset size. The same code predicts that models gradually overfit when trained for multiple epochs. It also isolates the learnable information in a dataset from its unpredictable, random content, revealing that lower-entropy text holds far more learnable structure than higher-entropy image data.
Quantum Multiscale Modeling: A Hierarchy of Algorithms for Complex Chemical Systems
arXiv:2607.11217v1 Announce Type: cross Abstract: Multiscale modeling of complex chemical systems requires algorithms that operate coherently across electronic, atomistic, mesoscopic, and continuum scales. While quantum algorithms have been proposed for each regime, no systematic framework exists to compose them across scale boundaries. Here, we identify the conditions under which fault-tolerant quantum algorithms might preserve scale-specific quantum advantages. We map quantum phase estimation, Hamiltonian simulation with Gibbs state preparation, quantum random walks, and quantum partial differential equation solvers onto electronic structure, molecular dynamics, mesoscopic kinetics, and continuum reactor physics, respectively. Crucially, these correspondences do not imply unconditional end-to-end quantum advantage; speedups depend heavily on state preparation, memory architectures, matrix conditioning, and classical readout costs. Six unresolved questions define this composition problem, illustrated via a quantum hierarchy for \ce{CO} oxidation over \ce{Pt(111)}. We propose viewing inter-scale transfer as a quantum channel composition problem at the interface of algorithm design and non-equilibrium statistical mechanics, and ask whether information loss at scale boundaries is intrinsic to multiscale modeling or merely a consequence of lossy classical transduction between algorithmic layers. The resulting roadmap suggests that multiscale quantum advantage is governed primarily by the structure of information transfer between algorithmic layers, rather than by performance at individual scales alone.
Edge transmission irregular graphs
arXiv:2607.10739v1 Announce Type: cross Abstract: The transmission of a vertex $v$ in a connected graph $G$ is the sum of distances from $v$ to all vertices in $G$. A transmission irregular (TI) graph is a connected graph in which any two distinct vertices have different transmissions. We extend the concept of transmission to edges by defining the transmission of an edge as the sum of the transmissions of its two endpoints. A connected graph can now be called edge transmission irregular (ETI) if any two distinct edges have different transmissions. We show that almost all graphs are not ETI and then investigate several related order realizability problems involving chemical ETI graphs. In particular, we prove that for every $n \ge 15$, there exists a subcubic tree of order $n$ that is both TI and ETI.
Anomalous Transverse Response and Multi-Field Ferrialtermagnetic-Ferroelectric Valve with CrSb Flakes
arXiv:2607.11222v1 Announce Type: cross Abstract: Altermagnets combine the zero-stray-field of antiferromagnets with the spin polarization of ferromagnets, showing great potential for spintronic applications. Here, we propose ferrialtermagnetism as a distinct subclass of altermagnetic family, where symmetry-inequivalent altermagnetic sublattices possess nonidentical Neel vectors, preventing mutual cancellation of alternating spin splitting and conferring intrinsic robustness against perturbations. This concept is realized in the three-atomic-layer CrSb (110) flakes, which exhibits spin splitting of 344 meV, moderate uniaxial magnetic anisotropy, and high Neel temperature of 657 K. The magneto-optical Kerr and the anomalous Hall effects are observed. Integrating this ferrialtermagnetic CrSb with ferroelectric Sc2CO2 and Cu spacer, we design an ferrialtermagnetic-ferroelectric valve. This device displays equilibrium tunneling magnetoresistance and electroresistance of ~10^3%, and non-equilibrium magnitudes under bias, thermal, or light field reaches ~10^4% with high spin filtering of 90%. The negative differential resistance and photogalvanic effects, and photocurrent extinction ratio of 283.8 are achieved. These findings establish ferrialtermagnetism as a fertile platform for multi-field-controlled, ultracompact, and self-powered spintronics and electronics.
Cosmology: 100 years after A. A. Friedmann
arXiv:2607.11791v1 Announce Type: new Abstract: A. A. Friedmann (04.06.1888 -- 16.09.1925) proposed the first physical cosmological models in 1920s. Despite the fact that Friedmann's works were very famous soon after their publication, the study of dynamic models of the Universe in the USSR was actually banned in the 30s - 50s of the last century, and Soviet philosophers and propagandists wrote that models of the evolving Universe were invented by Lemaitre on the demand of the Roman Pope, because according to Soviet philosophers such the birth of the universe is very similar to the divine creation of the world described in the Bible. Thus, in the USSR, Friedman's works were in oblivion in since 1930s until 1960s.
A Low-Latency Fraud Detection Layer for Detecting Adversarial Interaction Patterns in LLM-Powered Agents
arXiv:2605.01143v2 Announce Type: replace Abstract: Large Language Model (LLM)-powered agents demonstrate strong capabilities in autonomous task execution, tool use, and multi-step reasoning. However, their increasing autonomy also introduces a new attack surface: adversarial interactions can manipulate agent behavior through direct prompt injection, indirect content attacks, and multi-turn escalation strategies. Existing defense strategies focus on prompt-level filtering and rule-based guardrails, which are often insufficient when risk emerges gradually across interaction sequences. In this work, we propose a complementary defense mechanism: a low-latency fraud detection layer for detecting adversarial interaction patterns in LLM-powered agents. Instead of determining whether a single prompt is malicious, our approach models risk over interaction trajectories using structured runtime features derived from prompt characteristics, session dynamics, tool usage, execution context, and fraud-inspired signals. The detection layer can be implemented using lightweight models leading to low-latency real-time deployments. To evaluate the framework, we construct a synthetic corpus of 12,000 multi-turn agent interactions generated from parameterized templates that simulate realistic agentic workflows. Using 42 structured features and an XGBoost classifier, our detector achieves over 9 times faster than LLM-based detectors. Through the experiment and ablation studies, our work suggests that interaction-level behavioral detection should become a core component of deployment-time defense for LLM-powered agents.
PRiSM: Benchmarking Phone Realization in Speech Models
arXiv:2601.14046v2 Announce Type: replace Abstract: Phone recognition (PR) serves as the atomic interface for language-agnostic modeling for cross-lingual speech processing and phonetic analysis. Despite prolonged efforts in developing PR systems, current evaluations only measure surface-level transcription accuracy. We introduce PRiSM, the first open-source benchmark designed to expose blind spots in phonetic perception through intrinsic and extrinsic evaluation of PR systems. PRiSM standardizes transcription-based evaluation and assesses downstream utility in clinical, educational, and multilingual settings with transcription and representation probes. We find that diverse language exposure during training is key to PR performance, encoder-CTC models are the most stable, and specialized PR models still outperform Large Audio Language Models. PRiSM releases code, recipes, and datasets to move the field toward multilingual speech models with robust phonetic ability: https://github.com/changelinglab/prism.
Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO
arXiv:2607.11581v1 Announce Type: new Abstract: This paper introduces Actor as Its Own Critic, a unified reinforcement learning framework, Cycle Group Relative Policy Optimization (CycleGRPO), that jointly optimizes region understanding and localization for Multimodal Large Language Models (MLLMs). Unlike existing separate pipelines, we leverage the inherent duality between the two tasks to construct a self-evaluating reinforcement learning paradigm: "region $\to$ text $\to$ region''. Specifically, a single MLLM first acts as the actor to generate region captions, then immediately transitions to a critic to ground its generated text back in the spatial domain. Therefore, CycleGRPO requires only region inputs, e.g., masks or bounding boxes, entirely bypassing the need for textual ground truths. A quality-aware token-level cycle-consistency reward is employed to assess the semantic discriminability of text captions via their physical localization accuracy. Empirically, built upon SAMTok, our CycleGRPO framework successfully bootstraps both capabilities simultaneously. Without any task-specific fine-tuning, the framework yields consistent performance gains across a wide range of benchmarks, including region captioning, region VQA, grounded dialogue, and referring segmentation. Overall, CycleGRPO offers a straightforward and scalable way to advance pixel-level capabilities in MLLMs. Code and models are released at https://github.com/devinxzhang/CycleGRPO.
Overcoming Fourier Locking in Quantum Data Re-uploading Classifiers via Spectral Homotopy
arXiv:2607.11013v1 Announce Type: cross Abstract: Data re-uploading parameterized quantum circuits (DRU-PQCs) are universal function approximators, yet their expressivity produces oscillatory, non-convex loss landscapes that resist gradient-based optimization. We show that the primary optimization bottleneck in DRU-PQCs is not insufficient capacity but a structural failure mode we term Fourier locking (FL): because encoding weights and entangling layers are nonlinearly coupled, random initialization on high-frequency targets collapses the encoding parameters into spurious local minima. Two Fisher diagnostics characterize FL. The input-space quantum Fisher information $F_x$ measures the effective frequency content of the encoded state; the Fisher discriminant ratio of the measured features measures their alignment with the class labels. In two independent 50-seed experiments, the locking is literal: trapped circuits hold $F_x$ frozen for the entire run, while escaping circuits migrate their frequency content (direct training: $r_{pb} = -0.48$; curriculum: $d = 1.34$; both $p < 0.001$). The replicated signature is this spectral mobility, not any endpoint value of $F_x$, and trapped circuits retain a fully non-degenerate parameter-space QFIM ($r_{pb} \approx 0$): the failure is spectral misalignment of a responsive state, not a loss of geometric sensitivity. A frequency-staged homotopy protocol that paces the target frequency ($f: 1.0 \to 3.0$) convexifies the early loss landscape; escaping circuits raise $F_x$ in step with the curriculum, and the escape rate triples (18% vs. 6%). Fourier locking is a frequency-alignment problem, and its remedy is frequency pacing.
Maximum diversity and weighting for invariants of periodic time series
arXiv:2509.11146v2 Announce Type: replace-cross Abstract: Magnitude, obtained as a special case of Euler characteristic of enriched category, represents a sense of the size of metric spaces and is related to classical notions such as cardinality, dimension, and volume. While the studies have explained the meaning of magnitude from various perspectives, continuity also gives a valuable view of magnitude. Based on established results about continuity of magnitude and maximum diversity, this article focuses on continuity of weighting, a distribution whose totality is magnitude, and its variation corresponding to maximum diversity. Meanwhile, recent studies also illuminated the connection between magnitude and data analysis by applying magnitude theory to point clouds representing the data or the set of model parameters. This article will also provide an application for time series analysis by introducing a new kind of invariants of periodic time series, where the invariance follows directly from the continuity results. As a use-case, a simple machine learning experiment is conducted with real-world data, in which the suggested invariants improved the performance.
On the Boundary of the Robust Admissible Set in State and Input Constrained Nonlinear Systems
arXiv:2509.18825v2 Announce Type: replace-cross Abstract: In this paper, we consider nonlinear control systems subject to bounded disturbances and to both state and input constraints. We introduce the definition of robust admissible set - the set of all initial states from which the state and input constraints can be satisfied for all times against all admissible disturbances. We focus on its boundary that can be decomposed into the usable part on the state constraint boundary and the barrier, interior to the state constraints. We show that, at the intersection of these two components, the boundary of the robust admissible set must be tangent to the state constraint set and separate the interior of the robust admissible set and its complement, a property that we call the ultimate locally separating hyperplane condition. Moreover, we prove that the barrier must satisfy a saddle-point principle on a Hamiltonian, based on Pontryagin's maximum principle, whose final condition is precisely the ultimate locally separating condition, thus providing a set of differential equations made of the system and its adjoint for a direct construction of the barrier. Lastly, we illustrate our results by calculating the robust admissible set for an adaptive cruise control example.
People use fast and flat simulation to reason about new games
arXiv:2510.11503v2 Announce Type: replace-cross Abstract: Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence (AI), often focusing on expert-level or even super-human play. But real life also pushes human intelligence along a different frontier, requiring people to flexibly navigate decision-making problems that they have never thought about before. Here, we use novice gameplay to study how people reason about new problem settings. Through a series of large-scale behavioral studies with over 1000 participants and 121 two-player strategic board games (almost all novel to our participants), we show that people are systematic and adaptively rational in how they play a game for the first time, or evaluate a game (e.g., how fair or how fun it is likely to be) before they have played it even once. We explain these capacities via a computational cognitive model that we call the 'Intuitive Gamer', a model based on mechanisms of fast and flat (depth-limited) goal-directed probabilistic simulation. Our work offers new insights into how people rapidly evaluate, act, and make suggestions when encountering novel problems, and could inform the design of more flexible and human-like AI systems that can determine not just how to solve new tasks, but whether a task is worth thinking about at all.
IBPA: Real-time Free-form Manifold Mesh Reconstruction via Incremental Ball Pivoting with Integrated Hole Detection
arXiv:2607.11627v1 Announce Type: new Abstract: Both Remotely Operated underwater Vehicles (ROVs) and Autonomous Underwater Vehicles (AUVs) are frequently deployed to acquire geometric bathymetric data. However, it is often discovered post-survey that the acquired data coverage is incomplete. Given the high operational cost associated with underwater deployments, it is essential to incrementally visualize surface coverage in real-time to support informed decision-making by both the operators of ROVs and the AUVs during data collection. In addition, traditional incremental surface reconstruction methods, such as Digital Terrain Models (DTMs), are inherently limited in expressiveness: they represent surfaces as height fields, allows only one elevation value per $(x, y)$ coordinate and thus cannot capture overhangs or vertical structures. To overcome these limitations, we adapt the original Ball Pivoting Algorithm (BPA) into an incremental, real-time, and free-form surface reconstruction method, referred to as Incremental BPA (IBPA). Our method incrementally constructs an orientable, manifold mesh from streaming point cloud data without imposing assumptions regarding point cloud overlap or spatial distribution. Furthermore, we introduce a hole detection mechanism that identifies and highlights incomplete mesh regions. Compared to existing approaches, our method supports more complex surface topologies without prior structural assumptions. The source code of our reference implementation is available: https://github.com/Mauhing/Incremental-BPA
Proximity Measures for Classes of Phylogenetic Networks
arXiv:2607.11325v1 Announce Type: cross Abstract: Phylogenetic networks are used to represent the evolutionary history of species. Due to biological interpretations and computational advantages, researchers have focused on restricted classes of phylogenetic networks, such as tree-child, orchard, and tree-based. These classes capture different notions of tree-likeness: tree-child networks require every internal vertex to have a taxon reachable by a tree path, orchard networks are trees with horizontal arcs (for modelling histories rife with horizontal gene transfers), and tree-based networks are trees with additional (not-necessarily horizontal) arcs. A natural question to ask is ``how far is a given network from belonging to a particular class?'' This motivates the study of proximity measures, which measure the minimum number of graph modifications required to transform a network into one belonging to a particular class. In this paper, we consider three proximity measures based on leaf addition, valid arc deletion, and arc deletion. We study pairwise comparability of the proximity measures, prove complexity results, and derive extremal bounds for the classes of tree, tree-child, orchard, and tree-based networks.
Digital Engagement, Income Disparities, and Job Seeking in the United States since 2010
arXiv:2511.05294v2 Announce Type: replace Abstract: Surveys often record how frequently people use the internet without measuring the infrastructures, skills, and support systems that make digital participation possible. Using the U.S. National Longitudinal Survey of Youth 1997 cohort, we study how internet-use frequency relates to labor income, employment attachment, and job seeking after 2010. The main digital-engagement analysis uses the comparable 2011, 2013, and 2015 waves, with 2017 retained as later labor-market context. Across repeated cross sections, daily internet use consistently marks higher income and stronger employment attachment. Relative to daily use, less-than-daily use is associated with roughly 11 to 20 percent lower income, while nonuse is associated with about 18 to 21 percent lower income in 2011 and 2013. Respondents reporting no internet use are also 13 to 23 percentage points less likely to report full-year work. Job-search estimates reveal a distinct mechanism: active search is governed by employment status, search intensity, and application support, so a frequency item sorts respondents more sharply on durable labor-market attachment than on short-window search. Education accounts for a substantial share of the raw digital gradient, and pooled lagged-outcome and doubly robust transition estimates separate durable stratification from positive adoption margins. The results establish internet-use frequency as an informative behavioral marker of digitally mediated labor-market stratification and clarify why routine use should not be treated as a simple measure of digital access.
Controllably Efficient Language Models
arXiv:2511.05313v2 Announce Type: replace Abstract: The substantial inference costs of attention in transformers motivated the development of efficient sequence mixers: namely sparse and sliding window attention, convolutions and linear attention. Although these approaches result in impressive reductions in inference costs, they often trade-off with quality, specifically in-context recall. Apriori fixing this quality-cost tradeoff at training time means being suboptimal from the get-go: some downstream applications might fundamentally require more memory for in-context recall, while other tasks may require lower latency and memory. We propose a conceptually simple meta-sequence mixer with inference-cost controllability: the Compress & Attend Transformer (CAT). CAT decodes chunks of tokens by attending to compressed chunks of the sequence so far. Both compression and decoding can use any existing sequence mixer. Decoding from the compressed sequence yields compute and memory savings, with chunk size setting the operating point on the quality-cost trade-off. Importantly, training CAT across multiple chunk sizes at once unlocks test-time control of this trade-off without any retraining, all in a single model. Instantiated with the most basic choice, dense attention as the mixer, CAT surprisingly suffices to match 10 popular and diverse efficient models (linear, hybrids, sparse) on real-world long-context recall at comparable inference costs, all from a single trained model. CAT further performs competitively on long-context understanding benchmarks while providing 1.4-3.7x higher generation throughput than a dense transformer. Code is at: https://github.com/rajesh-lab/cat-transformer
Metacognition in LLMs: Foundations, Progress, and Opportunities
arXiv:2607.11881v1 Announce Type: new Abstract: Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities, nor how such abilities can be adapted to advance the fundamental capabilities, reliability, and intelligence of AI systems. This paper bridges this gap by presenting the first comprehensive overview of the current state of knowledge on metacognition for LLMs. We analyze and taxonomize the landscape of this emerging field and summarize recent technical advancements, including methods and benchmarks to measure and evaluate LLMs' metacognitive abilities, techniques to elicit, improve, and apply metacognition in LLMs, and findings and implications of ongoing research. We also discuss applications, open questions and challenges, and promising directions for future work. Our aim is to provide a detailed and up-to-date review of this topic and stimulate meaningful research and discussion. An organized list of papers can be found at https://github.com/yale-nlp/LLM-Metacognition.
Interaction Dynamics MPC for Knee Rehabilitation Exoskeletons: A Series-Elastic Instantiation
arXiv:2606.13485v2 Announce Type: replace Abstract: Safe rehabilitation is an interaction-dynamics problem: the controller must regulate a prescribed motion while absorbing involuntary spasm, voluntary effort, actuator compliance, and model mismatch as interaction disturbances. This paper instantiates the predictive interaction-dynamics framework of the base pHRI formulation on a series-elastic-actuated knee joint. SEA feedforward reduces the gravity-compensated knee to the same constant-coefficient scalar double integrator used in the base framework, while a dynamic-residual measurement from spring deflection supplies an interaction-disturbance observation. A steady-state target converts the estimated disturbance into a cancelling input, and a finite-horizon quadratic program regulates deviations from that target under range-of-motion, torque, and velocity constraints. The evaluation is stiffness- and damping-matched so improvements cannot be attributed to higher impedance. Under a motion-opposing $15\unit{Nm}$ step, classical impedance and MPC without estimation produce about $500\unit{mrad}$ steady-state error, whereas Kalman-augmented interaction MPC reduces this to $1.17\unit{mrad}$ at 100~Hz and $0.70\unit{mrad}$ at 500~Hz; the 500~Hz peak is $7.27\unit{mrad}$. In 30 randomized trials, the 95th-percentile peak is $21.57\unit{mrad}$. Bounded Assist-as-Needed scheduling, a corrective-channel energy tank, inequality-constrained OSQP stress cases, direct MuJoCo execution, and a posture-clamped MyoSuite knee-slice run are implemented. The results support the SEA-knee instantiation of the interaction-dynamics framework while separating it from clinical intent recognition, full-system passivity, safety certification, hardware trials, and free-standing multi-joint validation.
Exact Dynamics of Multi-class Stochastic Gradient Descent
arXiv:2510.14074v2 Announce Type: replace-cross Abstract: We develop a framework for analyzing the learning dynamics of high-dimensional problems trained using one-pass stochastic gradient descent (SGD) with data from multiple anisotropic classes. Our main theorem provides exact expressions for quantities of interest, including the risk and the overlap with the true signal, in terms of a deterministic system of ODEs, valid in the high-dimensional limit. The theorem holds for a broad class of optimization problems and extends to settings where the number of classes grows with dimension. To illustrate its utility, we investigate in detail the effect of the data's anisotropic structure on the problems of binary logistic regression and least-squares (LS) loss. We study the LS in a linear multiclass setup and derive a learning-rate threshold that depends on the average eigenvalue of the covariance matrices. In the binary logistic regression, we study three cases: isotropic covariances, data covariance matrices with a large fraction of zero eigenvalues (denoted as the zero-one model), and covariance matrices with power-law spectra. We show that a structural phase transition occurs. In particular, for the zero-one model and the power-law model with sufficiently large power, SGD aligns more closely with values of the class mean that are projected onto the ``clean directions'' (i.e., directions of smaller variance). This is supported by analytical studies and numerical simulations, which show the exact asymptotic behavior of the loss in the high-dimensional limit. The effects of data anisotropy that we demonstrate are likely to hold beyond these examples and illustrate one application of the broader theorem that we prove.
Calibrated Hybrid CNN-Transformer for Retinal OCT Classification
arXiv:2607.09809v1 Announce Type: cross Abstract: Deep models for retinal optical coherence tomography (OCT) classification report high accuracy but rarely report whether their confidence can be trusted -- a gap that matters when a wrong-but-confident reading delays sight-saving treatment. We pair a hybrid convolutional-Transformer encoder with a gradient-boosting (XGBoost) classification head and a three-part clinical safety layer: confidence calibration, out-of-distribution (OOD) rejection, and per-prediction uncertainty flagging. On four-class OCT (84,495 scans) the model reaches 95.4% accuracy while cutting calibration error twelve-fold (expected calibration error, ECE = 0.0024), so the confidence it reports tracks its true accuracy. To our knowledge this is the first OCT classifier to validate all three safety mechanisms jointly, with public weights and reproducible multi-seed evaluation.
Inf-Sup Neural Networks for High Dimensional PDEs
arXiv:2607.11718v1 Announce Type: new Abstract: Solving partial differential equations (PDEs) in high dimensions remains challenging due to the curse of dimensionality. We propose a neural-network-based framework that reformulates PDEs as inf--sup optimization problems through the introduction of a Lagrange multiplier. The primal solution and the associated Lagrange multiplier are parameterized by two networks and are computed via an iterative saddle-point optimization procedure. We prove the theoretical equivalence between the proposed optimization formulation and the original PDE problem, and we derive rigorous error estimates that quantify the total approximation error in terms of the network approximation error, statistical (sampling) error, and optimization error. Numerical experiments demonstrate the accuracy, stability, and efficiency of the proposed method for solving high-dimensional PDEs.
Enhancing Adversarial Transferability through Block Stretch and Shrink
arXiv:2511.17688v2 Announce Type: replace Abstract: Input transformation-based attacks improve adversarial transferability by aggregating gradients over transformed inputs. Existing analyses mainly explain their efficacy from image diversity, semantic preservation, attention variance or hypothesis space augmentation, yet overlook the critical role of model frontend responses. In this paper, we revisit transformation-based attacks from an implicit ensemble perspective: each transformation can be viewed as a pre-processing operator before the surrogate model, inducing a distinct frontend response for gradient aggregation. Based on this view, we propose FRO, a Frontend Response-Oriented input transformation method that enriches such responses through two complementary operators. The Local Scaling Operator perturbs local content sampling via block-wise stretch-and-shrink operations, while the Projection Operator modifies global spatial organization through coherent perspective deformation. Together, they produce structured transformed views to optimize transferable adversarial perturbations. Experiments on an ImageNet subset show that FRO consistently improves black-box transferability across diverse CNN and Vision Transformer models. We further analyze the effect of implicit ensemble size and evaluate different transformation-based methods under a unified ensemble scale, demonstrating the superiority of designing input transformations from the perspective of front-end response ensembles.