Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

Direct Rotor Thrust Sensing and Feedback Control for Disturbance Rejection of Multirotors Using Load-cells
arXiv:2607.10099v1 Announce Type: new Abstract: Gust disturbances, dynamic vertical inflow and ground effect are key adverse aerodynamic phenomena that induce variations in the forces acting on a multirotor and complicate its flight control. Miniature rotorcraft typically rely on simplified modelling of such effects to compute adjustments in thrust to counteract these forces. In the most basic case, disturbance force estimations are derived from the aircraft's motion and the generated thrust is assumed to exactly match that requested by the controller. However, such systems rely on the aircraft's trajectory to be affected before disturbances can be sensed and compensated. Numerous approaches presented over the last 15-20 years aim to reject external disturbances more quickly, but challenges remain. This paper presents a new approach in this category by measuring the instantaneous force of the rotors directly at the point of generation using load-cells and implementing high-speed control to accurately track the desired thrust. Measurements from load-cells were previously considered too noisy to provide meaningful input, but the experiments presented in the paper using purpose-built hardware from low-cost commodity components in single- and dual rotor see-saw models and a flying aircraft demonstrate both the feasibility and the effectiveness of the approach in the presence of complex aerodynamic phenomena.
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
arXiv:2502.20681v3 Announce Type: replace Abstract: Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Counterfact dataset, the answers progress from syntactically incorrect to syntactically correct to semantically correct. However, existing theoretical analyses hardly account for this feature-level two-stage phenomenon, which could be conceptually attributed to disentangled two-type features like syntax and semantics. In this paper, we theoretically demonstrate how the two-stage training dynamics potentially occur in transformers. Specifically, we analyze the feature learning dynamics induced by the aforementioned disentangled two-type feature structure, grounding our analysis in a simplified yet illustrative setting that comprises normalized ReLU self-attention and structured data. Such disentanglement of feature structure is general in practice, e.g., natural languages contain syntax and semantics, and proteins contain primary and secondary structures. To our best knowledge, this is the first rigorous result regarding a feature-level two-stage optimization process in transformers within this theoretical framework. A corollary further indicates that such a two-stage process is closely related to the spectral properties of attention weights.
Game Theory in Formula 1: From Physical to Strategic Interactions
arXiv:2503.05421v5 Announce Type: replace Abstract: This paper presents an optimization framework to model multi-agent racing dynamics. By incorporating physically accurate interaction models and accounting for the optimal responses of competing agents, our approach reveals strategic behaviors typical of motorsport. Aerodynamic wake effects, trajectory optimization, and energy management are captured and evaluated on a representative case study, based on a Formula 1 scenario. We describe the minimum lap time problem with two agents as either a Nash or a Stackelberg game, and by employing the Karush-Kuhn-Tucker conditions during the problem formulation, we recover the structure of a nonlinear program. In addition, we introduce an algorithm to refine local Stackelberg solutions, using the Nash costs as upper bounds. The resulting strategies are analyzed through case studies. We examine the impact of slipstreaming on trajectory selection in corners, straights, and high-speed sections, while also identifying optimal overtaking locations based on energy allocation strategies. Exploiting the structural similarities of the game formulations, we are able to compare symmetric and hierarchical strategies to analyze competitive racing dynamics. The proposed methodology closes the gap between theoretical game theory and practical applications, with relevance in multi-agent systems with coupled nonlinear dynamics.
Optically-powered Low Power Low Noise Amplifiers for MRI
arXiv:2607.10019v1 Announce Type: new Abstract: Purpose: Fully optical receive coils can potentially allow dense receiver arrays with a large channel count, reduced channel crosstalk, and less cable clutter. The power requirements of conventional low-noise amplifiers (LNAs) are prohibitive for simultaneously driving many coils through optical means, as opto-electric power conversion efficiencies can only reach about 50%. The goal is to develop low-power LNAs (LPLNA) with substantially lower power consumption without compromising noise figure (NF) and gain. Methods: A LPLNA was designed as a two-stage cascaded amplifier using an MR-compatible E-pHEMT (Enhancement-mode Pseudomorphic High Electron Mobility Transistor) transistor. The design was implemented on a single-sided printed circuit board (PCB), and its performance was compared with a commercial LNA. A four-channel shielded loop resonator array was constructed, and the signal-to-noise ratio (SNR), noise covariance, and preamplifier decoupling performance were evaluated. Results: The LPLNA had a five-fold lower electrical power consumption (40 mW) than the commercial LNA and provided comparable SNR in phantom measurements. In vivo experiments further confirmed that the LPLNA operates reliably under realistic MRI conditions. Additionally, four-channel receiver array measurements demonstrated comparable SNR within 2% of the commercial LNA and lower inter-channel noise correlation with 0.26 vs 0.3 on average. Conclusion: This study demonstrates the feasibility of LPLNAs for optically-powered RF receiver coil arrays. The LPLNA could also be applied in power-constrained or remote MRI environments.
Mid-Infrared Single-Photon Detection via Enhanced Cross-Phase Modulation in Topology-Optimized Epsilon-Near-Zero Dual-Wavelength Nanocavities
arXiv:2607.10472v1 Announce Type: new Abstract: We use the Green's tensor quantization theory for open resonant nanostructures with absorption losses to study the cross-phase modulation (XPM) process at the single photon level in nanoscale Kerr-type epsilon-near-zero (ENZ) materials with an effective nonlinear susceptibility $\chi^{(3)}(\omega)$ integrated inside dual-wavelength nanocavities. We obtain general analytical formulas for the achievable XPM frequency shift in a hybrid nanocavity that simultaneously traps a classical probe (signal) beam at 1.5 $\mu$m and single photon pump at 3 $\mu$m wavelengths. By focusing on mid-infrared photon detection at room temperature, we present a comprehensive analysis of the fundamental limits for single photon detection in the quantum nondemolition modality for a nanoscale region of high mobility cadmium oxide (CdO) with ENZ-enhanced Kerr-type nonlinearity embedded in a surrounding silicon (Si) environment inverse designed by free-form topology optimization. We numerically implement our theoretical results using finite element simulations within the rigorous framework of quasi-normal modes, demonstrating a single photon XPM frequency shift $\Delta f_s \approx 18.4 \text{ GHz}$ with fractional shift (i.e., frequency pulling) $\Delta f_s / f_s \approx 9.23 \times 10^{-5}$ and addressing the feasibility of detection in the proposed hybrid Si-CdO dual-wavelength nanocavity, either with a classical probe beam or a squeezed probe state, beyond the traditional limitations from self-phase modulation noise, thermorefractive noise, shot noise, and electronic jitter effects. This work establishes a robust benchmark for the engineering of mid-infrared single-photon nonlinear devices such as nondemolition quantum detectors, sensors, and all-optical gates on a solid state photonic platform.
Learning behavior accounts for background-related advantage in AI-assisted education
arXiv:2607.10101v1 Announce Type: new Abstract: Generative AI has been found, and will likely be found increasingly, useful in education. However, existing AI-for-education studies provide inconsistent evidence on its average effects. More broadly, research on prior educational technologies shows that average effects often mask substantial heterogeneity across student populations. Motivated by this evidence, this study examines heterogeneity in students' learning behavior with AI, which students benefit from AI assistance, and how learner profiles and learning behavior shape these patterns. To this end, we recruited 318 university students to participate in structured learning experiments lasting up to 125 minutes. Our findings indicate that students' learning behavior is strongly associated with learning outcomes, with behaviors characterized by proactive and critical engagement, rather than limited engagement, associated with significantly better performance. These behavioral differences are related to learner profiles, with students from higher-ranking universities and those with greater prior knowledge tending to benefit more, consistent with their greater likelihood of adopting proactive interaction strategies. Accounting for learning behavior substantially weakens or eliminates the associations between learner profiles and learning outcomes, suggesting that how students use AI is a key pathway through which background differences are linked to learning gains. Overall, this work provides a deeper understanding of AI assistance in education by showing how differences in learner profiles and learning behavior shape who benefits from AI-supported learning. These insights can help educators and students better navigate and integrate AI into educational practices.
Comparing algebraic cubature rules on spline curved elements
arXiv:2607.10485v1 Announce Type: new Abstract: We compare four methods for the construction of algebraic cubature rules on planar elements, whose boundary is tracked by splines. The methods, that we have developed over the last two decades, are based on Green theorem together with some cornerstones of polynomial approximation theory: Gaussian quadrature,Tchakaloff theorem, discretized Chebyshev expansion (hyperinterpolation), Fekete-like interpolation. We discuss their advantages and drawbacks in view of the application to curved polytopal element methods. We have also made freely available at a single site the corresponding open-source Matlab codes.
Measuring AI Ability to Complete Long Software Tasks
arXiv:2503.14499v4 Announce Type: replace Abstract: Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon. This is the time humans typically take to complete tasks that AI models can complete with 50% success rate. We first timed humans with relevant domain expertise on a combination of RE-Bench, HCAST, and 66 novel shorter tasks. On these tasks, current frontier AI models such as Claude 3.7 Sonnet have a 50% time horizon of around 50 minutes. Furthermore, frontier AI time horizon has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024. The increase in AI models' time horizons seems to be primarily driven by greater reliability and ability to adapt to mistakes, combined with better logical reasoning and tool use capabilities. We discuss the limitations of our results -- including their degree of external validity -- and the implications of increased autonomy for dangerous capabilities. If these results generalize to real-world software tasks, extrapolation of this trend predicts that within 5 years, AI systems will be capable of automating many software tasks that currently take humans a month.
Approximate Colorwise Tensorization of Entropy and Optimal Mixing of the Wang-Swendsen-Koteck\'y Dynamics
arXiv:2607.10119v1 Announce Type: new Abstract: We study the mixing time of Wang-Swendsen-Koteck\'{y} (WSK) dynamics for uniformly sampling proper $q$-colorings. The WSK dynamics is widely used in statistical physics for sampling from the antiferromagnetic Potts model and can be considered a global counterpart of the flip dynamics, which currently yields the state-of-the-art bounds for sampling colorings in general graphs (Carlson and Vigoda, SODA 2025). However, despite its importance, the tools for analyzing such dynamics remain limited. We develop new tools that enable us to analyze the mixing time of the WSK dynamics through the lens of relative entropy contraction. We introduce new criteria for multi-spin distributions: approximate colorwise tensorization of entropy (ACTE) and approximate colorwise subadditivity of entropy (ACSE). These criteria provide a colorwise counterpart to standard vertex-wise entropy factorization principles, and expose a form of color symmetry beyond coordinate-wise analyses. We also develop new inductive approaches for establishing such criteria on specific types of graphs, which can be viewed as local-to-global arguments for proving high-dimensional functional inequalities in a graph-theoretic sense. As concrete applications, we establish an optimal $O_q(\log n)$ mixing time for the WSK dynamics on chordal and outerplanar graphs, down to the optimal number of colors. Because trees and line graphs of trees are chordal, the result covers both vertex and edge colorings of trees. Our results work in a regime that bypasses the irreducibility threshold for Glauber dynamics while also improving the best known mixing time bounds (Carlson, Chen, Feng and Vigoda, SODA 2025).
PTEI: Integrating Personality Traits to Enhance Emotional Intelligence in Large Language Models
arXiv:2607.10245v1 Announce Type: new Abstract: Despite advances in Emotional Intelligence (EI), Large Language Models (LLMs) still significantly underperform humans in complex emotional reasoning. This gap originates partly from the limited incorporation of individual differences, particularly personality traits, which are fundamental to human emotional inference. To address this, we propose PTEI, a novel framework for integrating Personality Traits into Emotional Intelligence tasks using LLMs. In PTEI, MBTI and OCEAN personality traits are first extracted directly from the given emotional scenarios and then utilized as contextual knowledge within personality-aware prompts, guiding LLMs to accurately infer emotions and their underlying causes. To ensure optimal contextual grounding, we employ Contrastive Learning to construct an optimized retrieval system that surfaces emotionally and personally aligned scenarios, enhancing reasoning quality. Extensive experiments on established EI benchmarks show that PTEI enhances the Emotional Understanding (EU) capabilities of various LLMs, with the strongest improvement observed in GPT models. Combining PTEI with Chain-of-Thought (CoT) reasoning yields an additional 4 percent increase in accuracy. These findings underscore PTEI's contribution toward advancing AI systems with more sophisticated social and psychological grounding.
Diversify Diffusion with Temperature Sampling and Variance-Corrective Time Shifting
arXiv:2607.10853v1 Announce Type: new Abstract: Diffusion models faithfully reproduce their training distribution, but also inherit its imbalances and leave rare or under-represented modes hard to reach. A natural inference-time remedy is to sample from the high-temperature target $p^{(\gamma)}_0(x) \propto p_0(x)^{\gamma}$ for $0 < \gamma < 1$, which flattens dominant modes and lifts rare ones. However, naive score scaling while correctly reweighting modes also inflates the per-mode variance, breaking the reverse diffusion process and degrading sample quality. We introduce variance-corrective time shifting, a training-free fix that queries the network at a shifted timestep and scales the resulting score by $\gamma$, canceling the variance inflation while preserving the mode reweighting. The correction turns simple temperature sampling into a practical diversity knob for pretrained diffusion and flow-matching backbones with no retraining, and we demonstrate consistent gains at minimal cost to sample quality and condition fidelity across DiT, Stable Diffusion and Motion Diffusion models. We further show that the timing of the temperature intervention enables coarse-to-fine control: high-noise stages drive compositional diversity across modes, while low-noise stages drive local appearance variation under a fixed composition.
Semiotic problem framing: a new framework to guide students and teachers in conceptual understanding and teaching of physics
arXiv:2503.17040v2 Announce Type: replace Abstract: Problem solving in physics requires more than applying formulas: it involves describing and modeling phenomena, connecting mathematics with physics, and justifying reasoning choices. This process, known as problem framing, has been extensively studied in its cognitive and epistemic dimensions, but its semiotic aspects - how visuals, symbols, language, and metaphors shape understanding - remain underexplored. Physics relies on multiple representational modes that must be coordinated to construct meaning, and semiotics plays a central role in this integration. In this theoretical paper, we propose a new framework - the Semiotic Problem Framing - that explicitly incorporates semiotics into existing problem framing in physics. SPF highlights how students mobilize and shift across linguistic, visual, symbolic, and metaphorical resources in problem solving. For students, it offers a guide to structure reasoning and develop representational fluency; for teachers, it provides a diagnostic tool to scaffold and monitor learning processes. SPF enables analysis of reasoning patterns and error types not captured in previous frameworks, and suggests new directions for instructional design in physics education.
Active Noise Floor Estimation for Reliability-Optimal POMDPs: A Value-of-Noise-Information Approach
arXiv:2607.11822v1 Announce Type: new Abstract: Finite Reliability Representations (FRR) certify when a cell-constant policy is sufficient for reliable decision-making in a partially observed system with a known physical noise floor. In practice, however, sensing and execution noise can be latent and context-dependent. This paper develops a certificate-aware active disambiguation framework for an unknown physical noise parameter theta = (sigma_y, sigma_u), with the sensor-only case obtained by fixing sigma_u. We define the Value of Noise Information (VoNI) as the expected excess FRR certificate gap caused by using a reliability cover calibrated to the current estimate rather than to the realized noise parameter. We bound VoNI using action-value model mismatch and FRR radius inflation, showing that noise estimation has low decision value in sub-crossover regimes where the FRR certificate is insensitive to theta, but becomes valuable when posterior uncertainty can invalidate the current cover. A bi-level decision maker uses a posterior over theta, obtained from innovation statistics, execution residuals, or another online estimator, and triggers diagnostic probing only when uncertainty threatens the FRR certificate. We also interpret VoNI as a tractable, certificate-aware approximation to a high-level finite POMDP for latent sensing-execution regime disambiguation. Under stationary, identifiable, and persistently exciting regimes, we establish posterior consistency and convergence of the induced policy loss to the FRR approximation floor. Closed-loop UGV simulations with EKF-based innovation residuals show earlier detection of abrupt sensing-noise jumps, lower drift-tracking error, and substantially fewer probing actions than posterior-entropy exploration over 50 Monte Carlo trials.
FIERO: Empowering Creative Writing Through Collaborative Game Play
arXiv:2607.11837v1 Announce Type: new Abstract: Creativity often flourishes in collaboration, such as when designers brainstorm a new app together, or storytellers collectively build a world with elements of each person's narrative. However, collaborative storytelling can have challenges for its participants, such as when they disagree about the plot proposed, or when different ideas become fragmented when voiced individually. While current tools for creative collaboration focus on synchronous online text sharing, they often neglect the social dynamics of in-person collaboration critical to creative synergy. To address this, we created FIERO, a multiplayer web-based card game. Physical cards provide tangible scaffolding and social interaction, while the digital interface generates contextual visuals, facilitate group decisions, ensure narrative coherence, and synthesize different idea contributions using generative AI. Compared against online collaborative writing alone, the game significantly enhanced intuitive stimulation, idea fluency, and novelty generation, and also improved the content of the stories produced, leading to greater plot coherence (N=60). The cards provided creative structure and social engagement, while the interface provided contextualized augmentation without affecting player agency. This work shows how collaborative play can be utilized to foster creative support.
Unlocking Innate Computing Abilities in Electric Grids
arXiv:2505.10382v2 Announce Type: replace Abstract: Electric power grids are engineered energy systems whose forward electrical responses embody high-dimensional and memory-bearing transformations of input signals. In this work, we reveal that these transformations-inherent in electric circuit elements, power flows and network topologies-can be conveniently harnessed for computation without modifying physical grid architectures. By encoding structured input data into the operational setpoints of power electronic converters inside grids, we demonstrate how forward grid dynamics are interpreted into physical representations comprising system variables by showcasing through an affine transformation example implemented on a direct-current (DC) grid, which justifies the capability of grids performing information processing tasks concurrently alongside normal power flows. Our work not only underscores the computation capability intrinsic to grid physics, but also opens a new perspective on how energy networks can function as sustainable computational substrate. This positions them as flexible assets where several computing tasks from data centers can be sustainably outsourced.
A Glimpse into Long-term Physical Coexistence with Intelligent Robots
arXiv:2607.11377v1 Announce Type: new Abstract: Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must support diverse user-facing interfaces, maintain long-horizon memory of people and preferences, coordinate across robot embodiments, and translate human intent into safe physical execution. We introduce PHILIA, a multi-robot agent built around a robot gateway abstraction. PHILIA retains the rich interaction and tool ecosystem of OpenClaw while exposing robot-local runtimes, onboard perception, navigation, speaker, and robot policies through a unified capability interface. This design decouples low-frequency, high-semantic agent reasoning from high-frequency, low-level robot execution, enabling plug-and-play integration of user interfaces, robot embodiments, and policy backends. As a result, the user experience becomes compositional: advances in user interfaces, robot embodiments, robot policies, navigation, or interaction algorithms can improve the overall experience without redesigning the system. We validate the architecture on Astribot S1 robots while designing the robot gateway contract to support future heterogeneous robot platforms through a shared capability interface for observation, task execution, navigation, speech playback, status monitoring, and task cancellation. We present representative use cases in which agent memory and scene understanding are grounded in robot actions. These span interactive household scenarios, ranging from simple organization to challenging long-horizon and dexterous service tasks, such as packing a backpack and lifting a garbage bag. We highlight the human-robot interaction flow, where contextual understanding of user intent and preferences, together with human-in-the-loop confirmation or adjustment during execution, is essential for effective assistance.
Sense and Sensitivity: Examining the Influence of Semantic Recall on Long Context Code Understanding
arXiv:2505.13353v5 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed for understanding large codebases, but whether they understand operational semantics of long code context or rely on pattern matching shortcuts remains unclear. We distinguish between lexical recall (retrieving code verbatim) and semantic recall (understanding operational semantics). Evaluating 10 state-of-the-art LLMs, we find that while frontier models achieve near-perfect, position-independent lexical recall, semantic recall degrades severely when code is centrally positioned in long contexts. We introduce semantic recall sensitivity to measure whether tasks require understanding of code's operational semantics vs. permit pattern matching shortcuts. Through a novel counterfactual measurement method, we show that models rely heavily on pattern matching shortcuts to solve existing code understanding benchmarks. We propose a new task SemTrace, which achieves high semantic recall sensitivity through unpredictable operations; LLMs' accuracy exhibits severe positional effects, with median accuracy drops of 92.73% versus CRUXEval's 53.36% as the relevant code snippet approaches the middle of the input code context. Our findings suggest current evaluations substantially underestimate semantic recall failures in long context code understanding.
TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues
arXiv:2607.10130v1 Announce Type: new Abstract: Gaze target estimation aims to infer the position of a person's gaze within a scene. Within mainstream design logic, multi-branch methods require extra supervision and annotations, while streamlined designs prioritize low-level visual saliency over true gaze intent. The former leads to a high annotation burden and hinders domain transfer, whereas the latter causes misalignment between predicted attention and actual gaze targets. To address this issue, we propose TextGaze, a unified cross-modal architecture that leverages a Large Vision-Language Model (LVLM) as scalable semantic guidance to balance the two design paradigms. The model extracts visual features from a frozen encoder and utilizes an LVLM to obtain gaze-aligned textual cues. We design a transformer-based fusion module with hierarchical text supervision to preserve task semantics. Lightweight decoding heads enable the joint prediction of gaze heatmaps and in-/out-of-frame status. We evaluate our method on four mainstream datasets, and the results show competitive performance across key metrics with robust cross-dataset generalisation without extra fine-tuning. Overall, we provide a streamlined alternative to traditional designs and highlight the potential of LVLMs as accessible auxiliary guidance for gaze estimation.
Is Model Instability just Noise to be Tolerated or a Property that can be Managed?
arXiv:2607.10420v1 Announce Type: new Abstract: In software analytics, rerunning the same analysis twice often yields different models and conclusions. This reduces trust in the model and limits its use. We find that model instability is a major problem. Across 127 multi-objective SE optimization problems (12,700 test cases), repeated runs of a state-of-the-art optimizer agree on only 13.7% of test cases, even under improved settings. We argue that this instability is not merely noise to tolerate, but a property that can be measured and managed. By adjusting how labels are spent, how complex the models become, and how splits are scored, we obtain models that agree 4.8 times as often as the default configuration. The standard deviation of optimization error falls by 22% on average (mean std 17.4 to 13.6), while recommendation quality improves rather than degrades. In terms of quality, the refined settings are statistically top-ranked on 119 of 127 datasets, compared to 74 for the defaults. We then test causal and data-locality interventions and find that they help only partially, suggesting a residual stability floor. Our evidence suggests there are fundamental limits to stability set by the data itself (noise, scarce labels, proxy objectives, and the many near-equivalent models a dataset admits). We conclude that instability should be treated as a standard evaluation axis in SE optimization, which should be routinely measured, reported alongside performance, and used to calibrate trust in any single run. The methods in this paper provide a baseline against which future efforts to reduce SBSE instability can be judged. To support open science, we offer the following reproduction package: https://tinyurl.com/Model-Instability
BiomechGPT: Extending Motion-Language Models to Clinical Motion Understanding
arXiv:2505.18465v2 Announce Type: replace Abstract: Advances in markerless motion capture are making high-quality biomechanical data increasingly accessible, creating a growing need for scalable downstream analytics. Building a bespoke pipeline for each analysis task is time-consuming, motivating models that can flexibly handle diverse clinical questions within a single framework. Recent work has shown that fine-tuning language models to accept tokenized motion as an additional modality enables descriptive captioning of movement, raising the question of whether these models are also capable of clinically relevant motion understanding, where diverse tasks and annotations provide a natural testbed. We investigate whether such a multimodal motion--language model can answer detailed, clinically meaningful questions about movement. We collected 71 hours of biomechanical data from 750 participants, many with movement impairments, performing tasks commonly used in clinical assessment. To further expand the training dataset, we designed a cross-format tokenizer that directly encodes motion data from heterogeneous formats into a shared latent space without paired data, allowing a second dataset to be incorporated and enabling pooling annotations across datasets. From these tokenized representations, we constructed a multimodal dataset of motion-related question--answer pairs and used it to train BiomechGPT, a multimodal biomechanics--language model. BiomechGPT achieves competitive performance across a range of clinically relevant tasks, with performance scaling with both dataset and model size. It offers a new way for clinicians and researchers to interact with biomechanical data and represents a promising direction for rehabilitation-focused movement analysis. Project page: https://intelligentsensingandrehabilitation.github.io/BiomechGPT/
From ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argumentation
arXiv:2607.09664v1 Announce Type: new Abstract: To provide a structured and interpretable assessment, we decompose the image-based diagnosis into components following the Toulmin model of argumentation. This model consists of a claim, grounds, warrant, qualifier, rebuttal, and backing. Consider a claim generated by a machine learning (ML) model for retinal diagnosis. Rather than accepting this claim at face value, one could either apply explainable AI (XAI) methods or adopt an argumentation-based approach. In our framework, a model specialized in biomarker extraction from images provides the grounds. The warrant-linking the grounds to the claim - is analyzed by an agent equipped with medical knowledge; in our architecture, this role is fulfilled by a MedGemma agent. The qualifier is determined based on the overall quantitative evaluation of both the warrant and grounds models. Finally, a rebuttal is constructed using image similarity measures computed with MedSigLip. All these components are presented to the human expert, enabling a more informed and critical assessment of the ML-generated diagnosis.
Channel Knowledge Empowered Finite-Blocklength Rate-Splitting Transmission for High-Mobility Autonomous Driving
arXiv:2607.10138v1 Announce Type: new Abstract: To meet the extended ultra-low latency and high reliability (xURLLC) requirements for autonomous driving systems, multiple access schemes must operate reliably in high-mobility and complex propagation environments. Recently, rate-splitting multiple access (RSMA) has emerged as a promising multi-user transmission framework, showing robustness in dynamic situations where imperfect and outdated channel state information (CSI) is prevalent.Moreover, the advanced sensing, localization, and on-board computation capabilities of autonomous driving vehicles facilitate the construction of a channel knowledge map (CKM), which is a key enabler for environment-aware communications in future 6G networks.In this work, we propose a CKM empowered finite-blocklength (FBL) RSMA for downlink autonomous driving system. The location-dependent large-scale channel information provided by CKM is exploited in RSMA to develop a refined rate-splitting design. The min-rate performance of FBL rate splitting is analyzed explicitly to ensure user fairness. We derive a new and tight closed-form bound for the private-stream ergodic rate. Combined with the closed-form common-stream expression, an efficient optimization design of rate-splitting ratios has been formulated. Numerical results show that the CKM empowered FBL RSMA outperforms space-division multiple access (SDMA) and non-orthogonal multiple access (NOMA), particularly in high-mobility scenarios. Its performance is improved by a data-based CKM, which provides more accurate large-scale channel information than model-based approaches and enables more precise common-stream allocation. The results also reveal that RSMA is sensitive to errors in large-scale channel knowledge, emphasizing the importance of accurate CKM information for optimal rate-splitting.
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
arXiv:2505.18610v2 Announce Type: replace Abstract: Recently, significant progress has been made in developing reasoning-capable Large Language Models (LLMs) through long Chain-of-Thought (CoT) techniques. However, this long-CoT reasoning process imposes substantial memory overhead due to the large Key-Value (KV) Cache memory overhead. Post-training KV Cache quantization has emerged as a promising compression technique and has been extensively studied in short-context scenarios. However, directly applying existing methods to long-CoT LLMs causes significant performance degradation due to the following two reasons: (1) Large cumulative error: Existing methods fail to adequately leverage available memory, and they directly quantize the KV Cache during each decoding step, leading to large cumulative quantization error. (2) Short-context calibration: Due to Rotary Positional Embedding (RoPE), the use of short-context data during calibration fails to account for the distribution of less frequent channels in the Key Cache, resulting in performance loss. We propose Progressive Mixed-Precision KV Cache Quantization (PM-KVQ) for long-CoT LLMs to address the above issues in two folds: (1) To reduce cumulative error, we design a progressive quantization strategy to gradually lower the bit-width of KV Cache in each block. Then, we propose block-wise memory allocation to assign a higher bit-width to more sensitive transformer blocks. (2) To increase the calibration length without additional overhead, we propose a new calibration strategy with positional interpolation that leverages short calibration data with positional interpolation to approximate the data distribution of long-context data. Extensive experiments on 7B-70B long-CoT LLMs show that PM-KVQ improves reasoning benchmark performance by up to 8% over SOTA baselines under the same memory budget and achieves 2.73-5.18x throughput over the original 16-bit LLMs.
Gender Gap Analysis in News and Talk Online Radio Broadcast
arXiv:2607.09675v1 Announce Type: new Abstract: Radio broadcasting remains a dominant medium of communication, reaching 82% of Americans ages 12 and older weekly. Given its broad media impact, gender representation on radio news and talk stations may play an important role in shaping social and cultural perceptions. In this study, we examined patterns of gender representation in radio broadcasts, focusing on gender-based differences in total speaking time, air-time allocation across the day, and participation across broadcast topics. The dataset comprises filtered recordings from 74 US news and talk radio stations, collected over a 24-hour period and yielding more than 1,400 hours of content. We analyzed the data using VANPY, an in-house voice-analysis framework that combines multi-channel radio recording with AI-based speaker diarization, gender classification, speech-to-text transcription, and topic analysis. The results revealed consistent gender differences in allocated broadcast time, with male speakers accounting for 77% of total speaking time (SD = 6.8%). This gap remained evident during commute periods, when female speaking time was approximately 7.5-9.5 minutes per hour, compared with approximately 30.6-32.9 minutes for male speakers. Topic analysis further showed male dominance across all examined content categories. Female representation was lowest in "Talk Show Segments", where women accounted for only 10.8% of speaking time. Even in "Entertainment News", where female representation was highest, women accounted for only 35.8% of speaking time. Beyond these findings on gender dynamics in broadcast media, VANPY is publicly available and provides a systematic approach for analyzing large-scale audio data. It can be applied not only to radio broadcasts, but also to other audio domains such as human-computer interaction, corporate communication, and security applications.
Transcript-Free Lightweight Detection of Alzheimer's Disease from Spontaneous Speech Using Handcrafted MFCC-Dominant Acoustic Biomarkers
arXiv:2607.10168v1 Announce Type: new Abstract: It is still hard to find Alzheimer's disease (AD) early, especially when neuroimaging is expensive or tools that depend on language are not available. Spontaneous speech provides a non-invasive signal; however, numerous current methodologies depend on transcripts/ASR or computationally intensive deep models. We offer a simple, audio-only baseline for detecting AD using 176 Cookie Theft recordings from the DementiaBank Pitt corpus (88 AD, 88 controls). WebRTC voice activity detection (VAD) is used to separate speech from non-speech. We take out 99 hand-crafted acoustic-temporal features, including pause and fluency statistics, spectral/prosodic descriptors, and MFCC summaries with {\Delta} and {\Delta}{\Delta}. Evaluation is performed using a stringent speaker-independent GroupShuffleSplit,documenting performance across 30 iterations. A lightweight SVM with an RBF kernel gets an average AUC of 0.674 across runs. For example, a single split has an AUC of 0.742 and an accuracy of 0.657. We also present an exploratory compact-feature analysis utilizing a Top-20 subset ranked by Random Forest importance; since selection is not nested within training splits, these results may be overly optimistic and are not employed for primary conclusions (AUC 0.719). The results indicate that transcript-free spectro-temporal and fluency-related cues can facilitate speaker-independent Alzheimer's disease screening from raw audio, establishing a practical foundation for deployment-oriented research.