Forskningsradar

Science Journals

Peer-reviewade publikationer — 57198 artiklar

Connections Between Pairs of Filters Improve the Accuracy of Convolutional Neural Networks
arXiv:2606.13736v1 Announce Type: new Abstract: While researchers continue to find new and improved network structures for CNNs, most of the newly invented architectures still rely on the traditional pattern of stacking convolutional blocks and separating them with pointwise activation functions. However, there are drawbacks to a network purely building on pointwise nonlinearities. One alternative is to introduce a pairwise connection between two filters of a network. Typical connection functions use multiplications or the minimum operation to realize logical AND connections. In this paper, we go one step further by demonstrating that CNNs can benefit from more general connections, which include parameters that are learned. With such parameters, the network is able to implement different connections in different network layers and better adapt the connection function to the task at hand.
Right or Wrong, Models Comply: Directional Blindness in LLM Moral Judgment
arXiv:2606.14037v1 Announce Type: new Abstract: As language models take integrated roles across many domains, the response of LLMs to user pushback becomes a critical alignment property. Yet many existing evaluations treat compliance as unidirectional, measuring whether models resist pressure but not whether they resist it selectively. We introduce Compliance Asymmetry (A = BCR/HCR), a bidirectional diagnostic that compares beneficial output change under helpful nudges with harmful change under misleading nudges. Across 9 models and 972,000 nudge-condition responses, we find that this selectivity differs in factual and moral judgments: models follow helpful nudges more than harmful ones on factual questions (A = 1.58), but follow both directions at nearly identical rates on moral questions (A = 1.04). This phenomenon persists across model families, capability levels, and nudging types. Interestingly, we also find that chain-of-thought prompting amplifies helpful and harmful compliance together, while identity-based prompting suppresses both by nearly identical margins. These results identify direction-blind moral compliance as a distinct failure mode in current LLMs and suggest that alignment should target directionally calibrated updating rather than lower compliance alone.
Machine learning for rarefied gas transport in vacuum and micro/nano systems: promise, pitfalls, and a verification agenda
arXiv:2606.14039v1 Announce Type: new Abstract: Machine learning is beginning to influence rarefied-gas modeling at multiple levels, including equation-solving, operator learning, learned collision physics, moment closures, direct simulation Monte Carlo (DSMC) field surrogates, and gas--surface models. This Perspective argues that the central challenge is not demonstration-level success, but trustworthy use under realistic deployment conditions: multiregime Knudsen behavior, stochastic DSMC labels, sharp nonequilibrium structures, uncertain gas--surface interaction, and scarce direct experimental anchors. I classify the main method families by what is learned, distinguish soft physics penalties from structure-preserving designs, and propose evaluation standards based on extrapolation tests, noise-aware metrics, end-to-end cost accounting, and a three-level validation hierarchy. Most current evidence is solver-facing: it demonstrates surrogate fidelity to a teacher solver more often than direct physical fidelity to experiment. The aim is not to dismiss ML for rarefied and vacuum-related gas transport, but to separate what is already credible from what remains provisional, and to define a reporting standard that makes future claims auditable.
Rethinking One-Step Image Editing through ChordEdit: Reproduction, Simplification, and New Insights
arXiv:2606.14042v1 Announce Type: new Abstract: One-step image editing is important for making text-guided editing fast, practical, and easy to deploy, but its underlying mechanism is still not fully understood. We revisit ChordEdit through reproduction, ablation, and simplification. Our analysis shows that a) the chord window $\delta$ largely acts as an effective timestep shift from $t$ to $t - \delta$; b) chord transport acts on high-noise images and mainly performs low-frequency semantic editing; and c) proximal alignment acts on low-noise images and complements it by adding high-frequency target details. In this view, ChordEdit naturally decomposes editing into a coarse low-frequency transport stage and a fine high-frequency alignment stage. These findings suggest a path toward prompt-conditioned dynamic timestep selection for adaptive image editing. All code and results can be found at \href{https://github.com/Harvard-AI-and-Robotics-Lab/ChordEdit-Reproduction}{link}.
BigPower: Hierarchical Source-Level Module Power Estimation for CPUs with Large Language Models
arXiv:2606.13747v1 Announce Type: new Abstract: Accurate power estimation is important for understanding and optimizing CPU power behavior, yet practical workflows often rely on simulation-derived information or post-silicon analysis. In this work, we present BigPower, a hierarchical source-level surrogate model for fine-grained module-level power estimation during CPU design. BigPower leverages large language model-based representations together with architectural hierarchy, module connectivity, configuration parameters, and workload context to estimate module-level power consumption directly from source-level design information, without requiring additional simulation during inference. Experimental results in the open-source XiangShan processor family demonstrate practical fine-grained power estimation across diverse configurations and workloads, offering an efficient alternative to conventional simulation-based workflows.
Extreme-Scale Atomistic Simulation of Real-Temperature Magnetic Skyrmion Dynamics by Coupled Spin-Lattice Modeling
arXiv:2606.14073v1 Announce Type: new Abstract: Real-temperature topological magnetic dynamics in functional materials is governed by coupled lattice and spin evolution, yet remains inaccessible to predictive simulation at device-relevant scales. As a flagship example, thermally driven helix-to-skyrmion transformation in FeGe requires atomistic resolution, explicit lattice motion, and micrometer-scale domains to resolve device-scale topological texture formation. We combine a spin-constrained density-functional-theory-trained neuro-evolution potential with a structure-preserving spin-lattice integrator within one machine-learned framework. Architecture-specific optimizations, kernel fusion, SVE2 vectorization, and NUMA-aware data layout deliver a seven orders-of-magnitude speedup over prior spin-aware methods. Deployed on LineShine exascale supercomputer, the full application scales to 12.45 million CPU cores with 89.7% weak-scaling efficiency, enabling simulations of 1.34 trillion atoms and an equal number of spins while reaching 48.5 PFLOPS in double precision. The simulations directly resolve real-temperature skyrmion nucleation and reorganization at previously inaccessible scales, establishing a new regime for predictive simulation of coupled spin-lattice topological magnetic dynamics.
Numbers Already Carry Their Own Embeddings
arXiv:2606.14108v1 Announce Type: new Abstract: We introduce Adelic operation-preserved embeddings (AOE), a training-free representation that captures both a number's real value and its modular (p-adic) signatures. This construction preserves additive and multiplicative structure by design, turning numerical input into embeddings that "speak in the language of mathematics." Unlike prior approaches that rely on task-specific retraining, AOE is plug-and-play and drops seamlessly into existing architectures. On algebraic combinatorics benchmarks, it delivers consistent gains including the first-ever perfect accuracy on the Weaving Pattern task-while suggesting a principled path forward for overcoming the long-standing "number problem" in AI.
Robin-Neumann Coupling of PINN and FEM Solvers: A Steklov-Poincar\'e View, with Application to Fluid-Structure Interaction with Contact
arXiv:2606.14181v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) are meshless and carry moving geometry and topology change through resampling of collocation points; the finite-element method (FEM) is the workhorse for boundary-fitted discretisations. Coupling the two across a shared interface promises the best of both, yet existing PINN-FEM schemes are validated only empirically. We put the coupling on a domain-decomposition footing: viewing each solver as a Steklov-Poincar\'e (trace-to-flux) operator, we transfer the classical Dirichlet-Neumann (DN) divergence diagnosis and its Robin-Neumann (RN) cure, including a closed-form, sweep-free interface impedance, and prove a PINN-specific contraction theorem: a trained network realises only a perturbed Steklov operator with a per-step training residual, and RN still contracts, with no shared-eigenbasis hypothesis, to a floor set by the achieved training loss. Because a PINN has no stiffness matrix, we introduce a Fourier-mode interface probe that recovers the network's resolvable Steklov eigenvalues to within 0.5% and doubles as a diagnostic of the network's spectral cap. The theory predicts measured PINN-FEM contraction rates to within 7% on 1D and 2D Poisson couplings, and a two-slab analogue of the large-added-mass regime shows RN's per-mode impedance matching winning decisively where tuned scalar relaxation saturates. We demonstrate the framework on a Stokes/rigid-disc problem with Alart-Curnier contact: the meshless PINN fluid absorbs the topology change at contact by collocation exclusion alone, no remeshing and no cut cells, and the static-equilibrium contact reaction matches the submerged weight to 0.4% under mesh refinement. We quantify remaining limitations: the warm-started PINN drifts off the Stokes manifold over long horizons, and matched FEM-FEM benchmarks attribute pre-impact squeeze-film signatures to PINN under-resolution.
SOS-based Stability Verification for Saturated INDI Control of Hybrid-VTOL Aircraft Pitch Rate Dynamics
arXiv:2606.14198v1 Announce Type: new Abstract: Incremental nonlinear dynamic inversion (INDI) is a prominent flight-control strategy valued for its robust disturbance rejection; however, its formal stability verification has traditionally been limited to linearized dynamical models. This paper presents a formal nonlinear stability certificate for a saturated INDI pitch-rate controller for a hybrid vertical take-off and landing (VTOL) aircraft by representing the INDI controller via an equivalent recurrent equilibrium network (REN). By casting the saturated INDI architecture as a REN, the closed-loop dynamics are exactly mapped to an augmented state-feedback system. This structural equivalence enables the use of sum of squares (SOS) programming to synthesize a locally valid Lyapunov function without relying on conservative bounding approximations. The resulting certificate yields an inner estimate of the region of attraction (RoA) that explicitly accounts for actuator saturation, formally verifying the controller's stability in operating regimes where standard linear margins lose their validity.
Fast and accurate simulation of Raman spectra of gold-organic systems
arXiv:2606.14207v1 Announce Type: new Abstract: Resolving the spectral Raman signature of molecules grafted on a metallic support is often a difficult task, in which quantum chemistry methods allow for precious additional rationalization and signal attributions, especially to probe the formation of a bond with the support. In the specific case of gold-organic architectures based on a Au-C bond, only a limited amount of experimental and theoretical reference data are available in the literature, and Raman simulations based on quantum mechanics quickly become unaffordable with the size of the system. In this work, we evaluate the precision of a cost-efficient DFTB method to simulate Raman spectra of gold-organic systems at different scales, from gold complexes to functionalized gold surfaces. After a validation of the method through a careful comparison of DFTB Raman spectra of organometallic gold(I) and (III) complexes to DFT and experimental reference data, we discuss the case of molecules grafted on gold aggregates. For these simulations, the choice of the model (cluster or periodic surface) appears to be critical, and significant differences arise (positions and intensities of the peaks) when considering a full metallic slab, as allowed by the low computational cost of the method.
Closing the Reflection Gap: A Free Calibration Bonus for Agentic RL
arXiv:2606.14211v1 Announce Type: new Abstract: LLMs are increasingly deployed as agents that interact with external environments and observe feedback such as execution results, error messages, and tool outputs. A well-functioning agent should be able to leverage this feedback to accurately assess its own performance. Yet we find a persistent reflection gap: LLM agents tend to mis-assess their own outputs after observing concrete environment feedback -- even for questions they correctly answered -- and standard RL barely helps due to a credit-assignment mismatch. To close this gap, we propose RefGRPO, a simple yet effective fix that augments standard RL algorithms with two key ingredients: a free calibration bonus computed by contrasting the agent's own reflection with the actual outcome (requiring no additional reward model, LLM judge, or external annotation), and a dynamic schedule on its coefficient. Compared to standard RL baselines, our method simultaneously improves reflection calibration (e.g., reduces underconfidence rate $44.4\% \to 7.7\%$) and task accuracy (e.g., $75.1\% \to 76.5\%$) on text-to-SQL across five benchmarks. The resulting calibrated reflection turns the agent into its own verifier grounded in environment feedback, which further enables (i) better self-improvement that uses reflections as pseudo-rewards without outcome supervision, and (ii) more effective test-time selective prediction by committing only to rollouts flagged as correct.
ForestBack: Breadcrumb-Based Pedestrian Dead Reckoning for Infrastructure-Free Return Navigation
arXiv:2606.14421v1 Announce Type: new Abstract: Reliable return navigation remains an important challenge in GPS-denied environments where external positioning infrastructure may be unavailable or unreliable. This paper presents ForestBack, an infrastructure-free pedestrian return navigation framework based on breadcrumb-based pedestrian dead reckoning (PDR). The system records a user's walking route as a sequence of reversible breadcrumb nodes and generates reverse-path guidance without requiring GPS, Wi-Fi, Bluetooth beacons, or pre-installed infrastructure. ForestBack integrates acceleration-based step detection, adaptive step-length estimation, magnetometer-assisted heading estimation, barometric-altitude correction, and bidirectional breadcrumb path reconstruction. The system was evaluated using an indoor obstacle-avoidance route with five checkpoints, where the user navigated around a central obstacle. A dataset of 36 walking trials and 42,474 time-series samples was used for evaluation, including IMU signals, magnetometer readings, barometric variables, turn-event labels, ground-truth trajectories, baseline PDR outputs, proposed ForestBack outputs, and power-related measurements. Experimental results show that ForestBack reduced the mean RMSE from 1.129 m to 0.965 m compared with traditional PDR, corresponding to a 15.76% improvement. The mean final-position error was reduced from 1.781 m to 1.388 m, while turn-event detection consistency reached approximately 99.90%. These results indicate that ForestBack improves trajectory reconstruction and route-preserving return guidance in obstacle-avoidance scenarios. The released dataset and analysis notebook support reproducibility and future benchmarking of infrastructure-free PDR-based return navigation systems.
Breaking TinyML: Why Quantized Neural Networks Need Domain-Specific Security Analysis
arXiv:2606.14427v1 Announce Type: new Abstract: Most TinyML hardware accelerators focus on supporting Quantized Neural Networks (QNNs) to meet stringent constraints on power consumption and size. Despite this, the security aspects of quantization within TinyML hardware remain largely unexplored. Although previous studies indicate that QNNs demonstrate similar or enhanced robustness when compared to full-precision Deep Neural Networks (DNNs) against typical evasion attacks, no attack strategies tailored specifically for TinyML hardware have been proposed yet. This paper addresses this shortfall by demonstrating how a two-step attack pipeline can surpass the current state-of-the-art in the QNN context and shows the need for more hardware-aware security research.
Evaluating LLMs for Obfuscation Detection and Classification in Android Apps
arXiv:2606.14233v1 Announce Type: new Abstract: Android applications (apps) developers increasingly rely on code obfuscation techniques to hinder reverse engineering and protect intellectual property. However, obfuscation also reduces the effectiveness of static analysis and vulnerability detection tools, creating challenges for Android security analysis. Existing approaches for detecting obfuscation in Android apps predominantly rely on handcrafted heuristics, engineered features, or task-specific learning pipelines, which may struggle to generalize across evolving obfuscation strategies. This paper presents a large-scale empirical study investigating the capability of Large Language Models (LLMs) to detect obfuscation in Android apps through semantic reasoning. Our study evaluates whether off-the-shelf LLMs can identify obfuscated code without relying on handcrafted rules, predefined signatures, or dedicated model training. The empirical evaluation is conducted on both a controlled benchmark containing an app obfuscated with multiple techniques and a real-world dataset of Android apps collected from Google Play. The study further examines the impact of prompt design, model selection, and decision thresholds across several open-weight and proprietary LLMs. Finally, the analysis compares LLM-based reasoning with existing SAST-based obfuscation-detection approaches and discusses the broader implications and limitations of applying LLMs to Android security analysis.
Security in a Workflow: Exploring Role-Based Agentic Architectures for Vulnerability Handling
arXiv:2606.14261v1 Announce Type: new Abstract: Secure software engineering in practice is a multi-stage workflow involving vulnerability analysis, remediation, and fix verification. However, current LLM-based software security approaches often focus on isolated tasks such as detection or patch generation, with limited attention to agentic architectures reflecting industrial workflow. This creates a gap between existing LLM-based vulnerability-handling methods and real-world practices. In this paper, we study a role-based agentic workflow for vulnerability analysis and mitigation consisting of Planner, Analyzer, Fixer, and Verifier roles. To explore the effect of static analysis tool, the analyzer agent was integrated with the CodeQL in one of the workflows. The models used include nemotron-cascade-2:30b, qwen3-coder-next, and gpt-oss:120b. Our evaluation uses 25 real-world C/C++ vulnerabilities. The study reports 44% vulnerability detection accuracy comparable to GPT 5.5 and 19% fix accuracy. We also list implications from this study in context of software security practitioners.
The Good, the Bad, and the Ugly -- Living with Priors in Bayesian confirmation
arXiv:2606.14281v1 Announce Type: new Abstract: Bayesian confirmation faces a classic problem: where do initial priors come from? In cases with abundant data and repeated updating, different priors tend to converge to the same posterior. However, in frontier research this convergence often fails, and confirmation remains sensitive to priors. We examine how physics practice in the case of gravitational wave research deals with such cases and, normatively, when prior sensitivity should be regarded as epistemically problematic. We offer a practice-based account of the prior problem, so far absent from the philosophical literature on Bayesianism. One upshot is a clearer diagnosis of so-called 'analogue confirmation'.
Hybrid Dynamics of Rocking Blocks Beyond Overturning: Saltation Analysis, Bifurcations, and Stability Characterization
arXiv:2606.14288v1 Announce Type: new Abstract: This work investigates how restitution modeling affects the dynamics of rocking blocks subjected to harmonic excitation. While several studies have reported discrepancies between experimentally observed impact behavior and the predictions obtained using the classical Housner restitution coefficient, the implications of adopting alternative restitution formulations on the global dynamics of rocking systems remain largely unexplored. The system is formulated as a hybrid non-smooth dynamical model and analyzed through bifurcation diagrams, Lyapunov exponents, and basins of attraction for different slenderness ratios. By comparing the classical restitution model proposed by Housner with the alternative formulation of Mao et al., we show that the choice of restitution model strongly influences the predicted system response. The alternative formulation leads to an earlier onset and greater prevalence of complex oscillations, as well as changes in the type, stability, and accessibility of attractors compared to the classical model. However, as the slenderness ratio increases, the dynamical features produced by both formulations progressively converge, indicating a reduced sensitivity to the restitution model for taller blocks. These results provide a dynamical perspective on why alternative restitution formulations, which predict impact responses closer to experimental observations, can produce markedly different behaviors from those obtained using the classical Housner model.
A Robust Point Cloud Analysis Framework Inspired By Primary Visual Cortex
arXiv:2606.14292v1 Announce Type: new Abstract: Despite significant advancements in point cloud analysis, reducing energy consumption and improving robustness remain understudied, largely due to the inherent limitations of Convolutional Neural Networks (CNNs). To address this issue, we draw inspiration from the primary visual cortex and propose a Dendritic-Connected Continuous-Coupled Neural Network (DC-CCNN), a novel Brain-Inspired Neural Network (BINN) architecture for point cloud analysis. By combining discrete and continuous encoding, our design replaces traditional Multilayer Perceptrons (MLPs) with more efficient and robust BINNs. Building upon this framework, we further propose an extended model, DC-CCNN++, to improve robustness under complex corruption conditions. Specifically, we introduce a Neuro-Inspired Robust Modulation-and-Readout Module (NRMR) to enhance feature stability and decision robustness through global-context gain modulation and dual-code evidence integration. We also design a Cortically Inspired Progressive Variability Training (CPVT) strategy, which progressively exposes the model to structured environmental variability while preserving stable clean-sample anchors during training. Experimental results show that DC-CCNN++ improves the performance of brain-inspired networks on point cloud analysis while maintaining performance comparable to state-of-the-art methods. Compared with the original DC-CCNN, it achieves stronger results on both classification and part segmentation, and exhibits enhanced robustness against sparsity, occlusion, Gaussian noise, salt-and-pepper noise, and spatial transformations. With its efficiency, robustness, and biologically grounded design, DC-CCNN++ provides a promising alternative to traditional deep learning methods for point cloud analysis. Code is available at https://anonymous.4open.science/r/DC-CCNNpp-44E3.
What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective
arXiv:2606.14299v1 Announce Type: new Abstract: Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain vulnerable to distribution shifts encountered at deployment. Test-Time Adaptation (TTA) has recently been extended to CLIP as a lightweight solution, leading to a rapidly growing body of TTA4CLIP methods. However, empirical progress in this area has largely outpaced our understanding of what truly drives adaptation, where their gains originate, and under which shifts they remain reliable. In this paper, we take a step back from the pursuit of state-of-the-art accuracy and conduct a systematic controlled study of TTA4CLIP. We first organize existing methods into three unified paradigms according to what is updated at test time. We then introduce TTABC, an open-source TTA Benchmark for CLIP, which standardizes evaluation protocols and integrates more than 20 representative methods. Our controlled empirical analysis focuses on three key areas. First, we determine the driving factors in parameter-based methods, revealing that adaptation gains are primarily driven by test-time evidence and reliable proxies rather than heavy optimization. Second, we explore evidence utilization beyond heavy parameter tuning, showing that competitive and efficient performance can be achieved through cross- or current-sample evidence and lightweight prototype updates. Finally, we demonstrate that there is no silver bullet for TTA: no single adaptation paradigm is universally optimal, and the preferred paradigm depends on the nature of shift. We hope our benchmark and study provide a clearer understanding of the current TTA4CLIP landscape and establish a foundation for further research.
PIConGPU modeling of nanoplasma formation in helium nanodroplets irradiated by intense femtosecond laser pulses
arXiv:2606.14300v1 Announce Type: new Abstract: Helium nanodroplets provide a unique and versatile platform for investigating strong-field-driven nanoplasma dynamics. In this work, we present large-scale, GPU-accelerated particle-in-cell simulations using \textsc{PIConGPU} to study the interaction of pure helium nanodroplets containing up to $10^{6}$ atoms with intense near-infrared femtosecond laser pulses, and compare the results with single-shot velocity-map electron imaging and ion measurements. The simulations describe the plasma evolution from the first ionization events to collective electron motion, nanoplasma formation, and early expansion. We show that the calculated electron and ion observables reproduce the main features of the measured spectra in systems with similar cluster sizes and laser intensities. Our results demonstrate that \textsc{PIConGPU} captures the essential physics of nanoplasma formation previously addressed mainly with molecular-dynamics or TDDFT approaches, while remaining computationally efficient and applicable to much larger systems. This establishes \textsc{PIConGPU} as a powerful and scalable tool for connecting nanoplasma theory with experimentally accessible observables.
Taming Slivers: A Robust TFEM Framework for Reliable Computations on Degenerate Tetrahedral Meshes
arXiv:2606.14301v1 Announce Type: new Abstract: Sliver elements are an intrinsic difficulty of three-dimensional tetrahedral mesh generation and remain costly, and sometimes impractical, to eliminate completely. Although isolated degenerate elements do not necessarily prevent finite element convergence, connected clusters or sheets of slivers may impose artificial constraints on the discrete solution, leading to locking and severe loss of accuracy. In this work, we revisit the effect of slivers from the viewpoint of the finite element solution and propose a robust solver-side treatment based on the Tempered Finite Element Method (TFEM). The method limits the singular contribution of degenerate elements by introducing a lower bound on the Jacobian determinant, which can be interpreted as a vanishing added-volume correction. The resulting formulation prevents the effective element volume from falling below a threshold while preserving the relevant physical modes of the solution. We analyze the stiffness matrices of degenerate tetrahedra, identify the mechanisms responsible for locking in sliver bands, and assess the method on a range of representative physical problems, including incompressible flow, Cahn--Hilliard phase-field dynamics, transient wave propagation, and vibro-acoustic fluid--structure interaction. The numerical results show that TFEM consistently recovers accurate and physically meaningful solutions on meshes for which standard FEM exhibits locking or loss of convergence, providing a simple and broadly applicable alternative to exhaustive geometric sliver removal.
Retrospective Progress-Aware Self-Refinement for LLM Agent Training
arXiv:2606.14302v1 Announce Type: new Abstract: LLM-based agents trained with reinforcement learning optimize step-wise action prediction but lack metacognitive awareness of task progress, inducing a gap that hinders long-horizon scaling. A pilot study reveals that online progress prompting hurts performance while retrospective demonstrations help, yet this capability cannot emerge from outcome-reward training alone. We present RePro, Retrospective Progress-Aware Training, a framework that trains agents to self-generate progress signals via a forward-then-reflect rollout paradigm: the agent executes actions online, then retrospectively reassesses its step-wise progress given the completed trajectory and known outcome. RePro initializes with a Retrospection Warmup that teaches reflection format from minimal external demonstrations, then further trains through RePro-PO with a composite reward that produces self-generated signals without continuous external supervision. Experiments on WebShop, ALFWorld, and Sokoban show that RePro enhances the Qwen family's performance, with up to $12\%$ absolute success rate gains.
Instantaneous Pitch Estimation via Wave-U-Net-Based Fundamental Waveform Enhancement
arXiv:2606.14324v1 Announce Type: new Abstract: Instantaneous pitch estimation plays an important role in analyzing steep pitch variations such as speech prosody and singing techniques. Conventional approaches estimate instantaneous frequency after isolating the fundamental waveform from signals that contain harmonics and noise, which makes the accuracy sensitive to imperfect fundamental filtering. In this study, we formulate fundamental waveform filtering as a speech enhancement problem. Specifically, we train a Wave-U-Net model to extract a fundamental waveform from an input speech signal. The instantaneous pitch is then obtained by computing the instantaneous frequency from the analytic signal of the estimated fundamental waveform. Experimental results show that the proposed method outperforms conventional deterministic approaches and provides accurate and robust instantaneous pitch estimation across diverse domains, including speech, singing voice, musical instruments, and degraded speech signals.
iGLU 5.0: A Novel, Non-invasive and Intelligent HbA1c Measurement Device using Glucose values and Physiological Parameter for Smart Healthcare
arXiv:2606.14326v1 Announce Type: new Abstract: The laboratory test process of HbA1c measurement is a time-consuming and invasive method. The HbA1c parameter is the most important feature to predict the level of diabetes. Although invasive methods are irritating in the case of frequent measurements. Moreover, HbA1c measurement is only possible at the diagnostic centre, followed by medical protocol. Hence, it is still challenging to measure the HbA1c frequently at remote locations, where diagnostic centres are not easily available. Therefore, an intelligent and new non-invasive HbA1c measurement system, iGLU 5.0, is proposed for instant diagnosis of the HbA1c value without prior measurement setup. The proposed measurement device is based on optical spectroscopy for the collection of glucose values in different formats. The glucose values have been collected in fasting, postprandial, and random formats. The glucose value has also been collected using the oral glucose tolerance test (OGTT), along with the average blood pressure value, correspondingly. These four formats of glucose values, along with blood pressure, were used to predict the estimated average glucose (eAG) using an optimized prediction model. Further, the predicted average glucose is converted into an HbA1c value using a standard formula. The eAG prediction models have been trained and validated using 2000 samples of healthy, prediabetic and diabetic people to analyze the optimized prediction model. 94% and 96% accuracy have been examined during training and cross-validation of optimized DNN model, respectively. A 0.3 mean absolute difference has been identified from predicted HbA1c values using the proposed DNN model. The novel non-invasive HbA1c prediction system is useful for instant diagnosis without irritation for smart healthcare.
Can Deep Neural Networks Improve Compression of Very Large Scientific Data?
arXiv:2606.14353v1 Announce Type: new Abstract: Error-bounded lossy compression is a fundamental technique for managing the rapidly growing volumes of scientific data produced by modern simulations and observational instruments. Most state-of-the-art-compressors follow a prediction-residual paradigm, where compression effectiveness depends on the quality of the predictor: more accurate predictions generate smaller residuals that are easier to compress. This observation raises a question: can modern machine learning models serve as superior predictors for scientific data compression? Answering this question directly is challenging because developing compression-specific ML predictors requires substantial resources. Instead, we leverage the climate domain where highly accurate pretrained weather forecasting foundation models already exist, making them an ideal testbed. We present a framework that integrates spatial and temporal deep learning models into a conventional error-bounded compression pipeline. The framework supports auto-regressive forecasting models and avoids error accumulation. Using ERA5 climate data as a representative large-scale scientific dataset, we evaluate three distinct ML predictors: a VAEformer-based codec (CRA5), a graph neural network forecaster (GraphCast), and a vision-transformer forecaster (Aurora), against the state-of-the-art compressor SZ3.1 under identical quantization and entropy-coding backends. Our evaluation over approximately 1.7 TB of data reveals a surprising result: although ML predictors generate more accurate predictions and can improve reconstruction quality by up to 91% while achieving up to 9.6x higher compression ratios for highly predictable variables, they do not improve overall dataset-level compression ratio. We show that prediction accuracy alone is insufficient: the spatial structure of the resulting residuals plays a decisive role in entropy coding efficiency.