Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Game Theory in Formula 1: From Physical to Strategic Interactions
arXiv:2503.05421v5 Announce Type: replace Abstract: This paper presents an optimization framework to model multi-agent racing dynamics. By incorporating physically accurate interaction models and accounting for the optimal responses of competing agents, our approach reveals strategic behaviors typical of motorsport. Aerodynamic wake effects, trajectory optimization, and energy management are captured and evaluated on a representative case study, based on a Formula 1 scenario. We describe the minimum lap time problem with two agents as either a Nash or a Stackelberg game, and by employing the Karush-Kuhn-Tucker conditions during the problem formulation, we recover the structure of a nonlinear program. In addition, we introduce an algorithm to refine local Stackelberg solutions, using the Nash costs as upper bounds. The resulting strategies are analyzed through case studies. We examine the impact of slipstreaming on trajectory selection in corners, straights, and high-speed sections, while also identifying optimal overtaking locations based on energy allocation strategies. Exploiting the structural similarities of the game formulations, we are able to compare symmetric and hierarchical strategies to analyze competitive racing dynamics. The proposed methodology closes the gap between theoretical game theory and practical applications, with relevance in multi-agent systems with coupled nonlinear dynamics.
Optically-powered Low Power Low Noise Amplifiers for MRI
arXiv:2607.10019v1 Announce Type: new Abstract: Purpose: Fully optical receive coils can potentially allow dense receiver arrays with a large channel count, reduced channel crosstalk, and less cable clutter. The power requirements of conventional low-noise amplifiers (LNAs) are prohibitive for simultaneously driving many coils through optical means, as opto-electric power conversion efficiencies can only reach about 50%. The goal is to develop low-power LNAs (LPLNA) with substantially lower power consumption without compromising noise figure (NF) and gain. Methods: A LPLNA was designed as a two-stage cascaded amplifier using an MR-compatible E-pHEMT (Enhancement-mode Pseudomorphic High Electron Mobility Transistor) transistor. The design was implemented on a single-sided printed circuit board (PCB), and its performance was compared with a commercial LNA. A four-channel shielded loop resonator array was constructed, and the signal-to-noise ratio (SNR), noise covariance, and preamplifier decoupling performance were evaluated. Results: The LPLNA had a five-fold lower electrical power consumption (40 mW) than the commercial LNA and provided comparable SNR in phantom measurements. In vivo experiments further confirmed that the LPLNA operates reliably under realistic MRI conditions. Additionally, four-channel receiver array measurements demonstrated comparable SNR within 2% of the commercial LNA and lower inter-channel noise correlation with 0.26 vs 0.3 on average. Conclusion: This study demonstrates the feasibility of LPLNAs for optically-powered RF receiver coil arrays. The LPLNA could also be applied in power-constrained or remote MRI environments.
Mid-Infrared Single-Photon Detection via Enhanced Cross-Phase Modulation in Topology-Optimized Epsilon-Near-Zero Dual-Wavelength Nanocavities
arXiv:2607.10472v1 Announce Type: new Abstract: We use the Green's tensor quantization theory for open resonant nanostructures with absorption losses to study the cross-phase modulation (XPM) process at the single photon level in nanoscale Kerr-type epsilon-near-zero (ENZ) materials with an effective nonlinear susceptibility $\chi^{(3)}(\omega)$ integrated inside dual-wavelength nanocavities. We obtain general analytical formulas for the achievable XPM frequency shift in a hybrid nanocavity that simultaneously traps a classical probe (signal) beam at 1.5 $\mu$m and single photon pump at 3 $\mu$m wavelengths. By focusing on mid-infrared photon detection at room temperature, we present a comprehensive analysis of the fundamental limits for single photon detection in the quantum nondemolition modality for a nanoscale region of high mobility cadmium oxide (CdO) with ENZ-enhanced Kerr-type nonlinearity embedded in a surrounding silicon (Si) environment inverse designed by free-form topology optimization. We numerically implement our theoretical results using finite element simulations within the rigorous framework of quasi-normal modes, demonstrating a single photon XPM frequency shift $\Delta f_s \approx 18.4 \text{ GHz}$ with fractional shift (i.e., frequency pulling) $\Delta f_s / f_s \approx 9.23 \times 10^{-5}$ and addressing the feasibility of detection in the proposed hybrid Si-CdO dual-wavelength nanocavity, either with a classical probe beam or a squeezed probe state, beyond the traditional limitations from self-phase modulation noise, thermorefractive noise, shot noise, and electronic jitter effects. This work establishes a robust benchmark for the engineering of mid-infrared single-photon nonlinear devices such as nondemolition quantum detectors, sensors, and all-optical gates on a solid state photonic platform.
Learning behavior accounts for background-related advantage in AI-assisted education
arXiv:2607.10101v1 Announce Type: new Abstract: Generative AI has been found, and will likely be found increasingly, useful in education. However, existing AI-for-education studies provide inconsistent evidence on its average effects. More broadly, research on prior educational technologies shows that average effects often mask substantial heterogeneity across student populations. Motivated by this evidence, this study examines heterogeneity in students' learning behavior with AI, which students benefit from AI assistance, and how learner profiles and learning behavior shape these patterns. To this end, we recruited 318 university students to participate in structured learning experiments lasting up to 125 minutes. Our findings indicate that students' learning behavior is strongly associated with learning outcomes, with behaviors characterized by proactive and critical engagement, rather than limited engagement, associated with significantly better performance. These behavioral differences are related to learner profiles, with students from higher-ranking universities and those with greater prior knowledge tending to benefit more, consistent with their greater likelihood of adopting proactive interaction strategies. Accounting for learning behavior substantially weakens or eliminates the associations between learner profiles and learning outcomes, suggesting that how students use AI is a key pathway through which background differences are linked to learning gains. Overall, this work provides a deeper understanding of AI assistance in education by showing how differences in learner profiles and learning behavior shape who benefits from AI-supported learning. These insights can help educators and students better navigate and integrate AI into educational practices.
Comparing algebraic cubature rules on spline curved elements
arXiv:2607.10485v1 Announce Type: new Abstract: We compare four methods for the construction of algebraic cubature rules on planar elements, whose boundary is tracked by splines. The methods, that we have developed over the last two decades, are based on Green theorem together with some cornerstones of polynomial approximation theory: Gaussian quadrature,Tchakaloff theorem, discretized Chebyshev expansion (hyperinterpolation), Fekete-like interpolation. We discuss their advantages and drawbacks in view of the application to curved polytopal element methods. We have also made freely available at a single site the corresponding open-source Matlab codes.
Measuring AI Ability to Complete Long Software Tasks
arXiv:2503.14499v4 Announce Type: replace Abstract: Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon. This is the time humans typically take to complete tasks that AI models can complete with 50% success rate. We first timed humans with relevant domain expertise on a combination of RE-Bench, HCAST, and 66 novel shorter tasks. On these tasks, current frontier AI models such as Claude 3.7 Sonnet have a 50% time horizon of around 50 minutes. Furthermore, frontier AI time horizon has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024. The increase in AI models' time horizons seems to be primarily driven by greater reliability and ability to adapt to mistakes, combined with better logical reasoning and tool use capabilities. We discuss the limitations of our results -- including their degree of external validity -- and the implications of increased autonomy for dangerous capabilities. If these results generalize to real-world software tasks, extrapolation of this trend predicts that within 5 years, AI systems will be capable of automating many software tasks that currently take humans a month.
Approximate Colorwise Tensorization of Entropy and Optimal Mixing of the Wang-Swendsen-Koteck\'y Dynamics
arXiv:2607.10119v1 Announce Type: new Abstract: We study the mixing time of Wang-Swendsen-Koteck\'{y} (WSK) dynamics for uniformly sampling proper $q$-colorings. The WSK dynamics is widely used in statistical physics for sampling from the antiferromagnetic Potts model and can be considered a global counterpart of the flip dynamics, which currently yields the state-of-the-art bounds for sampling colorings in general graphs (Carlson and Vigoda, SODA 2025). However, despite its importance, the tools for analyzing such dynamics remain limited. We develop new tools that enable us to analyze the mixing time of the WSK dynamics through the lens of relative entropy contraction. We introduce new criteria for multi-spin distributions: approximate colorwise tensorization of entropy (ACTE) and approximate colorwise subadditivity of entropy (ACSE). These criteria provide a colorwise counterpart to standard vertex-wise entropy factorization principles, and expose a form of color symmetry beyond coordinate-wise analyses. We also develop new inductive approaches for establishing such criteria on specific types of graphs, which can be viewed as local-to-global arguments for proving high-dimensional functional inequalities in a graph-theoretic sense. As concrete applications, we establish an optimal $O_q(\log n)$ mixing time for the WSK dynamics on chordal and outerplanar graphs, down to the optimal number of colors. Because trees and line graphs of trees are chordal, the result covers both vertex and edge colorings of trees. Our results work in a regime that bypasses the irreducibility threshold for Glauber dynamics while also improving the best known mixing time bounds (Carlson, Chen, Feng and Vigoda, SODA 2025).
PTEI: Integrating Personality Traits to Enhance Emotional Intelligence in Large Language Models
arXiv:2607.10245v1 Announce Type: new Abstract: Despite advances in Emotional Intelligence (EI), Large Language Models (LLMs) still significantly underperform humans in complex emotional reasoning. This gap originates partly from the limited incorporation of individual differences, particularly personality traits, which are fundamental to human emotional inference. To address this, we propose PTEI, a novel framework for integrating Personality Traits into Emotional Intelligence tasks using LLMs. In PTEI, MBTI and OCEAN personality traits are first extracted directly from the given emotional scenarios and then utilized as contextual knowledge within personality-aware prompts, guiding LLMs to accurately infer emotions and their underlying causes. To ensure optimal contextual grounding, we employ Contrastive Learning to construct an optimized retrieval system that surfaces emotionally and personally aligned scenarios, enhancing reasoning quality. Extensive experiments on established EI benchmarks show that PTEI enhances the Emotional Understanding (EU) capabilities of various LLMs, with the strongest improvement observed in GPT models. Combining PTEI with Chain-of-Thought (CoT) reasoning yields an additional 4 percent increase in accuracy. These findings underscore PTEI's contribution toward advancing AI systems with more sophisticated social and psychological grounding.
Diversify Diffusion with Temperature Sampling and Variance-Corrective Time Shifting
arXiv:2607.10853v1 Announce Type: new Abstract: Diffusion models faithfully reproduce their training distribution, but also inherit its imbalances and leave rare or under-represented modes hard to reach. A natural inference-time remedy is to sample from the high-temperature target $p^{(\gamma)}_0(x) \propto p_0(x)^{\gamma}$ for $0 < \gamma < 1$, which flattens dominant modes and lifts rare ones. However, naive score scaling while correctly reweighting modes also inflates the per-mode variance, breaking the reverse diffusion process and degrading sample quality. We introduce variance-corrective time shifting, a training-free fix that queries the network at a shifted timestep and scales the resulting score by $\gamma$, canceling the variance inflation while preserving the mode reweighting. The correction turns simple temperature sampling into a practical diversity knob for pretrained diffusion and flow-matching backbones with no retraining, and we demonstrate consistent gains at minimal cost to sample quality and condition fidelity across DiT, Stable Diffusion and Motion Diffusion models. We further show that the timing of the temperature intervention enables coarse-to-fine control: high-noise stages drive compositional diversity across modes, while low-noise stages drive local appearance variation under a fixed composition.
Active Noise Floor Estimation for Reliability-Optimal POMDPs: A Value-of-Noise-Information Approach
arXiv:2607.11822v1 Announce Type: new Abstract: Finite Reliability Representations (FRR) certify when a cell-constant policy is sufficient for reliable decision-making in a partially observed system with a known physical noise floor. In practice, however, sensing and execution noise can be latent and context-dependent. This paper develops a certificate-aware active disambiguation framework for an unknown physical noise parameter theta = (sigma_y, sigma_u), with the sensor-only case obtained by fixing sigma_u. We define the Value of Noise Information (VoNI) as the expected excess FRR certificate gap caused by using a reliability cover calibrated to the current estimate rather than to the realized noise parameter. We bound VoNI using action-value model mismatch and FRR radius inflation, showing that noise estimation has low decision value in sub-crossover regimes where the FRR certificate is insensitive to theta, but becomes valuable when posterior uncertainty can invalidate the current cover. A bi-level decision maker uses a posterior over theta, obtained from innovation statistics, execution residuals, or another online estimator, and triggers diagnostic probing only when uncertainty threatens the FRR certificate. We also interpret VoNI as a tractable, certificate-aware approximation to a high-level finite POMDP for latent sensing-execution regime disambiguation. Under stationary, identifiable, and persistently exciting regimes, we establish posterior consistency and convergence of the induced policy loss to the FRR approximation floor. Closed-loop UGV simulations with EKF-based innovation residuals show earlier detection of abrupt sensing-noise jumps, lower drift-tracking error, and substantially fewer probing actions than posterior-entropy exploration over 50 Monte Carlo trials.
FIERO: Empowering Creative Writing Through Collaborative Game Play
arXiv:2607.11837v1 Announce Type: new Abstract: Creativity often flourishes in collaboration, such as when designers brainstorm a new app together, or storytellers collectively build a world with elements of each person's narrative. However, collaborative storytelling can have challenges for its participants, such as when they disagree about the plot proposed, or when different ideas become fragmented when voiced individually. While current tools for creative collaboration focus on synchronous online text sharing, they often neglect the social dynamics of in-person collaboration critical to creative synergy. To address this, we created FIERO, a multiplayer web-based card game. Physical cards provide tangible scaffolding and social interaction, while the digital interface generates contextual visuals, facilitate group decisions, ensure narrative coherence, and synthesize different idea contributions using generative AI. Compared against online collaborative writing alone, the game significantly enhanced intuitive stimulation, idea fluency, and novelty generation, and also improved the content of the stories produced, leading to greater plot coherence (N=60). The cards provided creative structure and social engagement, while the interface provided contextualized augmentation without affecting player agency. This work shows how collaborative play can be utilized to foster creative support.
A Glimpse into Long-term Physical Coexistence with Intelligent Robots
arXiv:2607.11377v1 Announce Type: new Abstract: Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must support diverse user-facing interfaces, maintain long-horizon memory of people and preferences, coordinate across robot embodiments, and translate human intent into safe physical execution. We introduce PHILIA, a multi-robot agent built around a robot gateway abstraction. PHILIA retains the rich interaction and tool ecosystem of OpenClaw while exposing robot-local runtimes, onboard perception, navigation, speaker, and robot policies through a unified capability interface. This design decouples low-frequency, high-semantic agent reasoning from high-frequency, low-level robot execution, enabling plug-and-play integration of user interfaces, robot embodiments, and policy backends. As a result, the user experience becomes compositional: advances in user interfaces, robot embodiments, robot policies, navigation, or interaction algorithms can improve the overall experience without redesigning the system. We validate the architecture on Astribot S1 robots while designing the robot gateway contract to support future heterogeneous robot platforms through a shared capability interface for observation, task execution, navigation, speech playback, status monitoring, and task cancellation. We present representative use cases in which agent memory and scene understanding are grounded in robot actions. These span interactive household scenarios, ranging from simple organization to challenging long-horizon and dexterous service tasks, such as packing a backpack and lifting a garbage bag. We highlight the human-robot interaction flow, where contextual understanding of user intent and preferences, together with human-in-the-loop confirmation or adjustment during execution, is essential for effective assistance.
Semiotic problem framing: a new framework to guide students and teachers in conceptual understanding and teaching of physics
arXiv:2503.17040v2 Announce Type: replace Abstract: Problem solving in physics requires more than applying formulas: it involves describing and modeling phenomena, connecting mathematics with physics, and justifying reasoning choices. This process, known as problem framing, has been extensively studied in its cognitive and epistemic dimensions, but its semiotic aspects - how visuals, symbols, language, and metaphors shape understanding - remain underexplored. Physics relies on multiple representational modes that must be coordinated to construct meaning, and semiotics plays a central role in this integration. In this theoretical paper, we propose a new framework - the Semiotic Problem Framing - that explicitly incorporates semiotics into existing problem framing in physics. SPF highlights how students mobilize and shift across linguistic, visual, symbolic, and metaphorical resources in problem solving. For students, it offers a guide to structure reasoning and develop representational fluency; for teachers, it provides a diagnostic tool to scaffold and monitor learning processes. SPF enables analysis of reasoning patterns and error types not captured in previous frameworks, and suggests new directions for instructional design in physics education.
TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues
arXiv:2607.10130v1 Announce Type: new Abstract: Gaze target estimation aims to infer the position of a person's gaze within a scene. Within mainstream design logic, multi-branch methods require extra supervision and annotations, while streamlined designs prioritize low-level visual saliency over true gaze intent. The former leads to a high annotation burden and hinders domain transfer, whereas the latter causes misalignment between predicted attention and actual gaze targets. To address this issue, we propose TextGaze, a unified cross-modal architecture that leverages a Large Vision-Language Model (LVLM) as scalable semantic guidance to balance the two design paradigms. The model extracts visual features from a frozen encoder and utilizes an LVLM to obtain gaze-aligned textual cues. We design a transformer-based fusion module with hierarchical text supervision to preserve task semantics. Lightweight decoding heads enable the joint prediction of gaze heatmaps and in-/out-of-frame status. We evaluate our method on four mainstream datasets, and the results show competitive performance across key metrics with robust cross-dataset generalisation without extra fine-tuning. Overall, we provide a streamlined alternative to traditional designs and highlight the potential of LVLMs as accessible auxiliary guidance for gaze estimation.
Is Model Instability just Noise to be Tolerated or a Property that can be Managed?
arXiv:2607.10420v1 Announce Type: new Abstract: In software analytics, rerunning the same analysis twice often yields different models and conclusions. This reduces trust in the model and limits its use. We find that model instability is a major problem. Across 127 multi-objective SE optimization problems (12,700 test cases), repeated runs of a state-of-the-art optimizer agree on only 13.7% of test cases, even under improved settings. We argue that this instability is not merely noise to tolerate, but a property that can be measured and managed. By adjusting how labels are spent, how complex the models become, and how splits are scored, we obtain models that agree 4.8 times as often as the default configuration. The standard deviation of optimization error falls by 22% on average (mean std 17.4 to 13.6), while recommendation quality improves rather than degrades. In terms of quality, the refined settings are statistically top-ranked on 119 of 127 datasets, compared to 74 for the defaults. We then test causal and data-locality interventions and find that they help only partially, suggesting a residual stability floor. Our evidence suggests there are fundamental limits to stability set by the data itself (noise, scarce labels, proxy objectives, and the many near-equivalent models a dataset admits). We conclude that instability should be treated as a standard evaluation axis in SE optimization, which should be routinely measured, reported alongside performance, and used to calibrate trust in any single run. The methods in this paper provide a baseline against which future efforts to reduce SBSE instability can be judged. To support open science, we offer the following reproduction package: https://tinyurl.com/Model-Instability
Unlocking Innate Computing Abilities in Electric Grids
arXiv:2505.10382v2 Announce Type: replace Abstract: Electric power grids are engineered energy systems whose forward electrical responses embody high-dimensional and memory-bearing transformations of input signals. In this work, we reveal that these transformations-inherent in electric circuit elements, power flows and network topologies-can be conveniently harnessed for computation without modifying physical grid architectures. By encoding structured input data into the operational setpoints of power electronic converters inside grids, we demonstrate how forward grid dynamics are interpreted into physical representations comprising system variables by showcasing through an affine transformation example implemented on a direct-current (DC) grid, which justifies the capability of grids performing information processing tasks concurrently alongside normal power flows. Our work not only underscores the computation capability intrinsic to grid physics, but also opens a new perspective on how energy networks can function as sustainable computational substrate. This positions them as flexible assets where several computing tasks from data centers can be sustainably outsourced.
From ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argumentation
arXiv:2607.09664v1 Announce Type: new Abstract: To provide a structured and interpretable assessment, we decompose the image-based diagnosis into components following the Toulmin model of argumentation. This model consists of a claim, grounds, warrant, qualifier, rebuttal, and backing. Consider a claim generated by a machine learning (ML) model for retinal diagnosis. Rather than accepting this claim at face value, one could either apply explainable AI (XAI) methods or adopt an argumentation-based approach. In our framework, a model specialized in biomarker extraction from images provides the grounds. The warrant-linking the grounds to the claim - is analyzed by an agent equipped with medical knowledge; in our architecture, this role is fulfilled by a MedGemma agent. The qualifier is determined based on the overall quantitative evaluation of both the warrant and grounds models. Finally, a rebuttal is constructed using image similarity measures computed with MedSigLip. All these components are presented to the human expert, enabling a more informed and critical assessment of the ML-generated diagnosis.
Channel Knowledge Empowered Finite-Blocklength Rate-Splitting Transmission for High-Mobility Autonomous Driving
arXiv:2607.10138v1 Announce Type: new Abstract: To meet the extended ultra-low latency and high reliability (xURLLC) requirements for autonomous driving systems, multiple access schemes must operate reliably in high-mobility and complex propagation environments. Recently, rate-splitting multiple access (RSMA) has emerged as a promising multi-user transmission framework, showing robustness in dynamic situations where imperfect and outdated channel state information (CSI) is prevalent.Moreover, the advanced sensing, localization, and on-board computation capabilities of autonomous driving vehicles facilitate the construction of a channel knowledge map (CKM), which is a key enabler for environment-aware communications in future 6G networks.In this work, we propose a CKM empowered finite-blocklength (FBL) RSMA for downlink autonomous driving system. The location-dependent large-scale channel information provided by CKM is exploited in RSMA to develop a refined rate-splitting design. The min-rate performance of FBL rate splitting is analyzed explicitly to ensure user fairness. We derive a new and tight closed-form bound for the private-stream ergodic rate. Combined with the closed-form common-stream expression, an efficient optimization design of rate-splitting ratios has been formulated. Numerical results show that the CKM empowered FBL RSMA outperforms space-division multiple access (SDMA) and non-orthogonal multiple access (NOMA), particularly in high-mobility scenarios. Its performance is improved by a data-based CKM, which provides more accurate large-scale channel information than model-based approaches and enables more precise common-stream allocation. The results also reveal that RSMA is sensitive to errors in large-scale channel knowledge, emphasizing the importance of accurate CKM information for optimal rate-splitting.
Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Understanding
arXiv:2505.13353v5 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed for understanding large codebases, but whether they understand operational semantics of long code context or rely on pattern matching shortcuts remains unclear. We distinguish between lexical recall (retrieving code verbatim) and semantic recall (understanding operational semantics). Evaluating 10 state-of-the-art LLMs, we find that while frontier models achieve near-perfect, position-independent lexical recall, semantic recall degrades severely when code is centrally positioned in long contexts. We introduce semantic recall sensitivity to measure whether tasks require understanding of code's operational semantics vs. permit pattern matching shortcuts. Through a novel counterfactual measurement method, we show that models rely heavily on pattern matching shortcuts to solve existing code understanding benchmarks. We propose a new task SemTrace, which achieves high semantic recall sensitivity through unpredictable operations; LLMs' accuracy exhibits severe positional effects, with median accuracy drops of 92.73% versus CRUXEval's 53.36% as the relevant code snippet approaches the middle of the input code context. Our findings suggest current evaluations substantially underestimate semantic recall failures in long context code understanding.
Gender Gap Analysis in News and Talk Online Radio Broadcast
arXiv:2607.09675v1 Announce Type: new Abstract: Radio broadcasting remains a dominant medium of communication, reaching 82% of Americans ages 12 and older weekly. Given its broad media impact, gender representation on radio news and talk stations may play an important role in shaping social and cultural perceptions. In this study, we examined patterns of gender representation in radio broadcasts, focusing on gender-based differences in total speaking time, air-time allocation across the day, and participation across broadcast topics. The dataset comprises filtered recordings from 74 US news and talk radio stations, collected over a 24-hour period and yielding more than 1,400 hours of content. We analyzed the data using VANPY, an in-house voice-analysis framework that combines multi-channel radio recording with AI-based speaker diarization, gender classification, speech-to-text transcription, and topic analysis. The results revealed consistent gender differences in allocated broadcast time, with male speakers accounting for 77% of total speaking time (SD = 6.8%). This gap remained evident during commute periods, when female speaking time was approximately 7.5-9.5 minutes per hour, compared with approximately 30.6-32.9 minutes for male speakers. Topic analysis further showed male dominance across all examined content categories. Female representation was lowest in "Talk Show Segments", where women accounted for only 10.8% of speaking time. Even in "Entertainment News", where female representation was highest, women accounted for only 35.8% of speaking time. Beyond these findings on gender dynamics in broadcast media, VANPY is publicly available and provides a systematic approach for analyzing large-scale audio data. It can be applied not only to radio broadcasts, but also to other audio domains such as human-computer interaction, corporate communication, and security applications.
Transcript-Free Lightweight Detection of Alzheimer's Disease from Spontaneous Speech Using Handcrafted MFCC-Dominant Acoustic Biomarkers
arXiv:2607.10168v1 Announce Type: new Abstract: It is still hard to find Alzheimer's disease (AD) early, especially when neuroimaging is expensive or tools that depend on language are not available. Spontaneous speech provides a non-invasive signal; however, numerous current methodologies depend on transcripts/ASR or computationally intensive deep models. We offer a simple, audio-only baseline for detecting AD using 176 Cookie Theft recordings from the DementiaBank Pitt corpus (88 AD, 88 controls). WebRTC voice activity detection (VAD) is used to separate speech from non-speech. We take out 99 hand-crafted acoustic-temporal features, including pause and fluency statistics, spectral/prosodic descriptors, and MFCC summaries with {\Delta} and {\Delta}{\Delta}. Evaluation is performed using a stringent speaker-independent GroupShuffleSplit,documenting performance across 30 iterations. A lightweight SVM with an RBF kernel gets an average AUC of 0.674 across runs. For example, a single split has an AUC of 0.742 and an accuracy of 0.657. We also present an exploratory compact-feature analysis utilizing a Top-20 subset ranked by Random Forest importance; since selection is not nested within training splits, these results may be overly optimistic and are not employed for primary conclusions (AUC 0.719). The results indicate that transcript-free spectro-temporal and fluency-related cues can facilitate speaker-independent Alzheimer's disease screening from raw audio, establishing a practical foundation for deployment-oriented research.
LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans
arXiv:2607.10878v1 Announce Type: new Abstract: AI agents are evolving from answer engines into persistent teams that use tools, delegate work, learn from experience, and modify the artifacts that shape their future behavior. The defining question for deployment is no longer merely what agents can do, but who controls what they are allowed to become. We introduce logos, a pluggable layer for self-evolution and governance that strengthens existing multiagent frameworks rather than replacing them. logos compiles heterogeneous multimodal inputs, including documents, images, audio, tables, databases, APIs, and human instructions into versioned agent packs containing agents, tools, knowledge, tests, permissions, and policies. During operation, it transforms agent activity into portable, auditable event traces and applies fail-closed verification across frameworks and backends. Every learned prompt, memory, skill, tool, role, or workflow remains an untrusted release candidate until held-out execution evidence, human-controlled policy, and explicit authorization permit its promotion. This architecture enables "verifiable human-agent loop engineering": agents can act, ask, learn, and propose improvements, while humans can steer objectives, permissions, approvals, and irreversible actions without interrupting continuous operation. logos provides a living logic for accountable automation. Agents may evolve at machine speed, but only evidence and human authority can close the loop.
A Better Analysis For PPSZ For 3-SAT
arXiv:2607.10697v1 Announce Type: new Abstract: We revisit Scheder's analysis of the original PPSZ algorithm. Keeping his regular and irregular estimates unchanged, we express them in common structural coordinates and replace only their final recombination by an explicit linear-programming dual certificate. The old and new running-time bounds are \[ \begin{array}{c|cc} & \text{Unique-$3$-SAT} & \text{general $3$-SAT} \\ \hline \text{Scheder's analysis} & O^*(1.306972377^n) & O^*(1.307031594^n) \\ \text{this work} & O^*(1.306969598^n) & O^*(1.307031578^n). \end{array} \] In both rows, the general-case bound is obtained by applying the same existing Scheder--Steinberger unique-to-general lifting theorem to the corresponding Unique-$3$-SAT analysis. To the best of our knowledge, $O^*(1.307031578^n)$ is the best currently known worst-case randomized running-time bound for general $3$-SAT. Neither PPSZ nor the lifting theorem is modified. The numerical inequalities are certified by exact rational interval computation.
Anisotropy and intermittency in drift-wave turbulence with zonal flows: a two-dimensional continuous wavelet analysis
arXiv:2607.10136v1 Announce Type: new Abstract: We examine anisotropy and spatial intermittency at small scales in drift-wave turbulence with zonal flows. We use a two-dimensional directional continuous wavelet transform, which allows simultaneous localization in scale, position, and direction. This wavelet analysis is applied to vorticity fields obtained from numerical simulations of the modified Hasegawa--Wakatani model, a reduced model of resistive drift-wave turbulence in magnetized plasmas with zonal flows. Directional wavelet statistics characterize the anisotropy of the turbulence. The second-order moment is enhanced around directions perpendicular to the zonal flow. Spatial intermittency, characterized by scale-dependent flatness, is more pronounced around directions along the zonal flow.
Low-Rank Attention Residuals
arXiv:2607.09694v1 Announce Type: new Abstract: Attention Residuals replace the fixed residual sum with depthwise attention over previous sub-layer outputs in large language models (LLMs), but use each output as both a full-dimensional key and value. This couples routing with representation and makes depth-routing scores scale with the hidden width $d$. We propose Low-Rank Attention Residuals (LR-AttnRes), which keep full-dimensional residual values while using $r$-dimensional keys, with $r \ll d$, for routing. Projected LR-AttnRes emits learned low-rank keys from existing output projections, decoupling routing from residual content and achieving the best validation loss among the variants tested. Sliced LR-AttnRes uses the last $r$ dimensions of each value as the routing key, removing the auxiliary key-projection path and reducing residual-side FLOPs while still improving performance. Comprehensive sweeps show that depthwise routing can be effective with far fewer dimensions than the model width. We release code and models to facilitate future research.